PAM diversification of CAS9 variants

WO2026180998A1PCT designated stage Publication Date: 2026-09-03CRISPR THERAPEUTICS AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2026/051852
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-26
Filing Date
2026-02-26
Publication Date
2026-09-03

Smart Images

  • Figure IB2026051852_03092026_PF_FP_ABST
    Figure IB2026051852_03092026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein include methods, compositions, and kits suitable for use in gene editing. In some embodiments, novel Cas9 polynucleotides are provided having an diversified range of PAM recognition profiles. Editors, base editors, polymerase-based editors, and epigenetic editors comprising all or a portion of a variant Cas9 polynucleotide are provided in some embodiments.
Need to check novelty before this filing date? Find Prior Art

Description

CT242-PCT1 / 80EM-800002-WO PATENT PAM DIVERSIFICATION OF CAS9 VARIANTSRELATED APPLICATIONS

[0001] This application claims the benefit under 35 U.S.C. §119(e) of U.S. Provisional Patent Application Ser. No. 63 / 763,733, filed February 26, 2025, the content of this related application is incorporated herein by reference in its entirety for all purposes.REFERENCE TO SEQUENCE LISTING

[0002] The present application is being filed along with a Sequence Listing in electronic format. The Sequence Listing is provided as a file entitled 80EM-800002-WO, created February 20, 2026, which is 2,429,953 bytes in size. The information in the electronic format of the Sequence Listing is incorporated herein by reference in its entirety.BACKGROUNDField

[0003] The present disclosure relates generally to the field of gene editing.Description of the Related Art

[0004] CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats)-Cas (CRISPR associated) systems have been used for genome editing, and require a Cas polypeptide or variant thereof guided by a customizable guide RNA (gRNA) for programmable DNA targeting. Cas9 requires the presence of a protospacer-adjacent motif (PAM) immediately adjacent to the 3 ’-end of the targeted nucleic acid sequence, which effectively limits the nucleotide sequences which can be efficiently targeted by Cas9. There is a need for variant Cas9 polypeptides with diversified PAM sequence compatibilities.SUMMARY

[0005] Disclosed herein include engineered Cas9 proteins. An engineered Cas9 protein can comprise a substitution, insertion or deletion at one or more amino acid residues in the Recognition (REC) Lobe Domain and / or the PAM-Interacting Domain (PID) as compared to a parent Cas9 protein comprising the sequence selected from SEQ ID Nos: 728-747. In some embodiments, the substitution, insertion or deletion at one or more amino acid residues in the Recognition (REC) Lobe Domain. In some embodiments, the substitution, insertion or deletion is at one or more than one of amino acid residues selected from amino acid positions 73-466 of SEQ ID Nos: 728-747. In some embodiments, the substitution, insertion or deletion is at one or more of amino acid residues selected from amino acid positions 218-269 of SEQ ID Nos: 728-747. In some embodiments, the substitution, insertion or deletion is at one or more of amino acid residues selected from amino acid positions 240-263 of SEQ ID Nos: 728-747. The engineered Cas9protein can comprise an insertion at an amino acid position selected from amino acid positions 240-263 of SEQ ID Nos: 728-747, wherein the insertion is about 1-50 amino acids. The engineered Cas9 protein can comprise a deletion at an amino acid position selected from amino acid positions 240-263 of SEQ ID Nos: 728-747, wherein the deletion is about 1-20 amino acids. In some embodiments, the substitution, insertion or deletion at one or more amino acid residues in the PID. In some embodiments, the substitution, insertion or deletion is at one or more of amino acid residues selected from the 185 most C-terminal amino acids of SEQ ID Nos: 728-747. In some embodiments, the substitution, insertion or deletion is at one or more of amino acid residues selected from amino acid positions 943-1093 of SEQ ID Nos: 728-747. The engineered Cas9 protein can comprise an insertion at an amino acid position selected from amino acid positions 943-1093 of SEQ ID Nos: 728-747, wherein the insertion is about 1-50 amino acids. The engineered Cas9 protein can comprise a deletion at an amino acid position selected from amino acid positions 943-1093 of SEQ ID Nos: 728-747, wherein the deletion is about 1-20 amino acids. In some embodiments, the Cas9 protein is a nuclease. In some embodiments, the Cas9 protein is a nickase. In some embodiments, the Cas9 protein is a deactivated (dead) Cas9 protein.

[0006] Disclosed herein include nucleases. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 1, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 1, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 2, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 2, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 3, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 3, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 4, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 4, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNGNAANN, NNGAAAGN, NNGAAAAN, NNGAAAGN, NNGGAATN, and / or NNGGAAGN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 5, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 5, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 6, optionally an amino acidsequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 6, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN, NNAGAAAN, NNAGCAAN, and / or NNAGGAAN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 7, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 7, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 8, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 8, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 9, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 9, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 10, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 10, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNRNNNNN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 11, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 11, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNANGANN, NNAAGACN, NNAGGATN, NNAAGATN, NNAAGCAN, NNAAGAAN, NNAAGCGN, and / or NNAAGCCN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 12, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 12, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 13, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 13, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAYAAMN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 14, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 14, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 15, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 15, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity ofNNAGAAAN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 16, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 16, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 17, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 17, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 18, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 18, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 19, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 19, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNGAAANN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 20, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 20, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNGGNNCN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 21, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 21, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNRGAACN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 22, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 22, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNNNAANN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 23, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 23, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 24, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 24, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 25, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 25, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity ofNNANAANN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 26, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 26, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 27, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 27, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAYAACN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 28, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 28, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAYAACN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 29, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 29, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAYAACN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 30, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 30, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 31, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 31, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAAAGNN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 32, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 32, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNGGNNNN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 33, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 33, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNGGNNNN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 34, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 34, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNGGNNCN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 35, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 35, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity ofNNAYAANN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 36, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 36, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAAAGNN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 37, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 37, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNNNAANN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 38, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 38, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNNNAANN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 39, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 39, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 40, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 40, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNGNNNNN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 41, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 41, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 42, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 42, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAAAGNN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 43, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 43, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAAANNN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 44, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 44, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAYAANN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 45, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 45, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity ofNNAAAGNN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 46, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 46, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN. The nuclease can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 47, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 47, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN.

[0007] Disclosed herein include nickases. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 48, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 48, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 49, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 49, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 50, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 50, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 51, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 51, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNGNAANN, NNGAAAGN, NNGAAAAN, NNGAAAGN, NNGGAATN, and / or NNGGAAGN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 52, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 52, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 53, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 53, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN, NNAGAAAN, NNAGCAAN, and / or NNAGGAAN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 54, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 54, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nickase can comprise an amino acid sequence that is at least 80% identical toSEQ ID NO: 55, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 55, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 56, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 56, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 57, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 57, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNRNNNNN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 58, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 58, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNANGANN, NNAAGACN, NNAGGATN, NNAAGATN, NNAAGCAN, NNAAGAAN, NNAAGCGN, and / or NNAAGCCN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 59, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 59, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 60, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 60, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAYAAMN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 61, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 61, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 62, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 62, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 63, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 63, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 64, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 64, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity ofNNAGAAAN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 65, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 65, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 66, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 66, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNGAAANN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 67, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 67, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNGGNNCN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 68, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 68, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNRGAACN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 69, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 69, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNNNAANN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 70, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 70, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 71, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 71, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 72, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 72, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 73, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 73, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 74, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 74, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity ofNNAYAACN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 75, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 75, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAYAACN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 76, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 76, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAYAACN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 77, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 77, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 78, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 78, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAAAGNN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 79, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 79, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNGGNNNN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 80, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 80, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNGGNNNN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 81, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 81, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNGGNNCN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 82, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 82, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAYAANN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 83, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 83, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAAAGNN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 84, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 84, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity ofNNNNAANN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 85, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 85, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNNNAANN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 86, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 86, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 87, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 87, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNGNNNNN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 88, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 88, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 89, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 89, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAAAGNN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 90, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 90, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAAANNN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 91, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 91, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAYAANN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 92, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 92, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAAAGNN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 93, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 93, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN. The nickase can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 94, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 94, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity ofNNAGAANN. The nickase can comprise an amino acid sequence that is at least 80% identical to any one of the sequences of SEQ ID NOs: 748-770, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to any one of the sequences of SEQ ID NOs: 748-770.

[0008] Disclosed herein include editors. The editor can comprise: an engineered Cas9 protein disclosed herein; a nuclease disclosed herein; or a nickase disclosed herein; or a deactivated (dead) Cas9 protein disclosed herein. Disclosed herein include editing systems. The editing system can comprise: an editor comprising the engineered Cas9 protein disclosed herein, a nuclease disclosed herein or a nickase disclosed herein or a deactivated (dead) Cas9 protein disclosed herein; and a guide RNA (gRNA). In some embodiments, the editor comprises an effector domain. In some embodiments, the effector domain comprises nuclease activity, nickase activity, recombinase activity, deaminase activity, methyltransferase activity, methylase activity, acetylase activity, acetyltransferase activity, transcriptional activation activity, transcriptional repression activity, and / or polymerase activity; RNA binding activity, DNA binding activity. In some embodiments, the effector domain comprises a non-long-terminal-repeat (non-LTR) retrotransposable element enzyme or a portion thereof, optionally derived from CRE, R2, Randl / Dualen, R4, NeSL, Hero, Proto 1, LI, Txl, Proto2, RTETP, RTEX, RTE, I, Outcast, Nimb, Ingi, Jockey, Rl, Loa, Tadl, Rexl, CR1, L2A, L2B, L2, Daphne, Crack, Vingis, or any combination thereof. In some embodiments, the effector domain is a nucleic acid editing domain. In some embodiments, the nucleic acid editing domain comprises a deaminase domain.

[0009] Disclosed herein include base editors. In some embodiments, the base editor comprises: an engineered Cas9 protein disclosed herein or a nickase disclosed herein; and a deaminase domain. Disclosed herein include base editing systems. The base editing system can comprise: a base editor comprising a deaminase domain and (i) an engineered Cas9 protein disclosed herein or (ii) a nickase disclosed herein; and a guide RNA (gRNA). In some embodiments, the deaminase domain is an adenosine deaminase domain. In some embodiments, the adenosine deaminase domain is an E. coli Tad A (ecTadA) deaminase domain. In some embodiments, the deaminase domain is a cytosine deaminase domain. In some embodiments, the cytosine deaminase domain is an apolipoprotein B mRNA-editing complex (APOBEC) family deaminase domain.

[0010] Disclosed herein include reverse transcriptase (RT) editors. In some embodiments, the RT editor comprises: a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, wherein the DNA polymerase domain, the DNA binding domain, and the DNA endonuclease domain are fused or linked to form a fusion protein, wherein the DNApolymerase domain comprises a reverse transcriptase, and wherein the DNA endonuclease domain comprises an engineered Cas9 protein disclosed herein or a nickase disclosed herein.

[0011] Disclosed herein include reverse transcriptase (RT) editing systems. The RT editing system can comprise: a template armed guide RNA (tagRNA), or a nucleic acid encoding the tagRNA, wherein the tagRNA comprises: a spacer that is complementary to a search target sequence on a first strand of a nucleic acid molecule; an editing template that comprises a region of complementarity to an editing target sequence on a second strand of the nucleic acid molecule; and a scaffold sequence that associates with a reverse transcriptase (RT) editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain; and a reverse transcriptase (RT) editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, or a nucleic acid encoding the RT editor, wherein the DNA polymerase domain comprises a reverse transcriptase; and wherein the DNA endonuclease domain comprises an engineered Cas9 protein disclosed herein or a nickase disclosed herein.

[0012] In some embodiments, the tagRNA comprises a flap binding sequence at least partially complementary to the spacer. In some embodiments, the scaffold sequence is between the spacer and the editing template. The tagRNA can comprise from 5’ to 3’: the spacer, the scaffold sequence, the editing template, and the flap binding sequence. In some embodiments, the spacer, the scaffold sequence, the editing template, and the flap binding sequence form a contiguous sequence in a single molecule. In some embodiments, the editing template comprises an intended nucleotide edit compared to the double stranded target DNA. In some embodiments, the tagRNA guides the RT editor to incorporate the intended nucleotide edit into the double stranded target DNA when the tagRNA is contacted with the double stranded target DNA. In some embodiments, the RT editor synthesizes a single stranded DNA encoded by the editing template, wherein the single stranded DNA replaces the editing target sequence and results in incorporation of the intended nucleotide edit into a region corresponding to the editing target in the double stranded target DNA. In some embodiments, the search target sequence is complementary to a protospacer sequence in the double stranded target DNA, and wherein the protospacer sequence is adjacent to a protospacer adjacent motif (PAM) in the double stranded target DNA. In some embodiments, the tagRNA results in incorporation of a nucleotide edit in the PAM when contacted with the double stranded target DNA. In some embodiments, the spacer of the tagRNA is from 16 to 25 nucleotides in length, optionally 20 nucleotides in length or 21-23 nucleotides in length. In some embodiments, the flap binding sequence is about 2 to 20 nucleotides in length, optionally about 8 to 16 nucleotides in length or 6 nucleotides in length. In some embodiments, the editing template is about 4 to 30 nucleotides in length, optionally about 10 to 30 nucleotides in length, further optionally 6 to 9 nucleotides in length. In some embodiments, the tagRNA results inincorporation of the intended nucleotide edit about 0 to 30 base pairs downstream of the nickase cleavage site. In some embodiments, the intended nucleotide edit comprises a single nucleotide substitution compared to the region corresponding to the editing target in the double stranded target DNA. In some embodiments, the intended nucleotide edit comprises an insertion compared to the region corresponding to the editing target in the double stranded target DNA, optionally an insertion of a nucleotide sequence at least 50, at least 45, at least 40, at least 35, at least 30, at least 25, at least 20, at least 15, at least 10, or at least 5, nucleotides in length. In some embodiments, the intended nucleotide edit comprises a deletion compared to the region corresponding to the editing target in the double stranded target DNA. In some embodiments, the editing template comprises one or more silent nucleotide edits compared to the region corresponding to the editing target in the double stranded target DNA, optionally said silent nucleotide edits do not alter the amino acid sequence of the protein encoded by the double stranded target DNA, further optionally said one or more silent nucleotide edits comprise a substitution of 2 to 5 contiguous nucleotides. In some embodiments, the editing template comprises a wild type double stranded target DNA sequence. In some embodiments, the tagRNA results in correction of a mutation when contacted with the double stranded target DNA. In some embodiments, the reverse transcriptase is a retrovirus reverse transcriptase. In some embodiments, the reverse transcriptase is a Moloney murine leukemia virus (MMLV) reverse transcriptase. In some embodiments, the RT editor comprises an amino acid sequence that is at least 80% identical to any one of SEQ ID NOs: 771-796, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of any one of SEQ ID NOs: 771-796. In some embodiments, the DNA polymerase domain, the DNA binding domain, and the DNA endonuclease domain are fused or linked to form a fusion protein.

[0013] Disclosed herein include ribonucleoprotein (RNP) complexes. The RNP complex can comprise an editing system disclosed herein, base editing system disclosed herein, an RT editing system disclosed herein, or a component thereof. Disclosed herein include lipid nanoparticles (LNPs). The LNP can comprise an editing system disclosed herein, base editing system disclosed herein, an RT editing system disclosed herein, or a component thereof. The LNP can comprise (iii) the gRNA and the nucleic acid encoding the editor (ii) the gRNA and the nucleic acid encoding the base editor, or (iii) the tagRNA and the nucleic acid encoding the RT editor. In some embodiments, the nucleic acid encoding the editor, base editor, or RT editor is mRNA. The LNP can comprise the egRNA.

[0014] Disclosed herein include systems for editing SERPINA1. The system for editing SERPINA1 can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 95-99 and 305-309, or a sequence that exhibits at least about 85%identity to any one of the sequences of SEQ ID NOs: 95-99 and 305-309; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 515-519, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 515-519; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 4, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 4; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 51, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 51. The system for editing SERPINA1 can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 100-119 and 310-329, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 100-119 and 310-329; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 520-539, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 520-539; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 11, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 11; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 58, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 58. The system for editing SERPINA1 can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 120-127 and 330-337, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 120-127 and 330-337; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 540-547, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 540-547; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 13, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 13; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 60, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 60. The system for editing SERPINA1 can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 128-133 and 338-343, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 128-133 and 338-343; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 548-553, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 548-553; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQID NO: 19, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 19; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 66, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 66. The system for editing SERPINA1 can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 134-153 and 344-363, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 134-153 and 344-363; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 554-573, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 554-573; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 20, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 20; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 67, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 67. The system for editing SERPINA1 can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 154-159 and 364-369, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 154-159 and 364-369; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 574-579, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 574-579; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 10, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 10; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 57, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 57. The system for editing SERPINA1 can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 160-163 and 370-373, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 160-163 and 370-373; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 580-583, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 580-583; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 6, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 6; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 53, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 53.

[0015] Disclosed herein include systems for editing PAH. The system for editing PAH can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 164-182 and 374-392, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 164-182 and 374-392; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 584-602, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 584-602; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 4, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 4; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 51, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 51. The system for editing PAH can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 183-209 and 393-419, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 183-209 and 393-419; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 603-629, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 603-629; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 11, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 11; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 58, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 58. The system for editing PAH can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 210-219 and 420-429, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 210-219 and 420-429; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 630-639, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 630-639; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 13, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 13; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 60, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 60. The system for editing PAH can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 220-226 and 430-436, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 220-226 and 430-436; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 640-646, or a sequencethat exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 640-646; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 19, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 19; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 66, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 66. The system for editing PAH can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 227-248 and 437-458, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 227-248 and 437-458; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 647-668, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 647-668; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 20, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 20; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 67, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 67. The system for editing PAH can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 249-259 and 459-469, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 249-259 and 459-469; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 669-679, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 669-679; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 10, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 10; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 57, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 57. The system for editing PAH can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 260-271 and 470-481, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 260-271 and 470-481; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 680-691, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 680-691; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 6, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 6; or a nickase comprising an amino acid sequence thatis at least 80% identical to SEQ ID NO: 53, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 53.

[0016] Disclosed herein include systems for editing BCL11 A. The system for editing BCL11 A can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 272-277 and 482-487, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 272-277 and 482-487; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 692-697, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 692-697; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 4, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 4; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 51, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 51. The system for editing BCL11 A can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 278-285 and 488-495, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 278-285 and 488-495; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 698-705, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 698-705; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 11, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 11; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 58, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 58. The system for editing BCL11 A can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 286-290 and 496-500, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 286-290 and 496-500; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 706-710, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 706-710; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 13, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 13; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 60, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 60. The system for editing BCL11 A can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 291-292 and 501-502, or a sequence that exhibits atleast about 85% identity to any one of the sequences of SEQ ID NOs: 291-292 and 501-502; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 711-712, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 711-712; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 19, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 19; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 66, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 66. The system for editing BCL11 A can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 293-300 and 503-510, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 293-300 and 503-510; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 713-720, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 713-720; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 20, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 20; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 67, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 67. The system for editing BCL11A can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 301 and 511, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 301 and 511; and / or a protospacer comprising the sequence of SEQ ID NO: 721, or a sequence that exhibits at least about 85% identity to the sequence of SEQ ID NO: 721; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 10, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 10; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 57, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 57. The system for editing BCL11 A can comprise: a gRNA or a tagRNA, comprising: a sequence of any one of the sequences of SEQ ID NOs: 302-304 and 512-514, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 302-304 and 512-514; and / or a protospacer comprising any one of the sequences of SEQ ID NOs: 722-724, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 722-724; and an editor, a base editor, or RT editor, comprising: a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 6, optionally anamino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 6; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 53, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 53.

[0017] Disclosed herein include polynucleotides. The polynucleotide can encode the editor or editing system disclosed herein, the base editor or base editing system disclosed herein, the RT editor or RT editing system disclosed herein, or the system disclosed herein. The polynucleotide can comprise any one of the sequences of SEQ ID NOs: 797-822, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 797-822. In some embodiments, the polynucleotide is an mRNA. In some embodiments, the polynucleotide is operably linked to a regulatory element, optionally the regulatory element is an inducible regulatory element. Disclosed herein include vectors comprising a polynucleotide disclosed herein. In some embodiments, the vector is an AAV vector.

[0018] Disclosed herein include isolated cell(s) comprising an editor or editing system disclosed herein, a base editor or base editing system disclosed herein, a RT editor or RT editing system disclosed herein, a RNP disclosed herein, a LNP disclosed herein, a system disclosed herein, a polynucleotide disclosed herein, or a vector disclosed herein. In some embodiments, the cell is a mammalian cell, optionally a human cell. In some embodiments, the cell is a primary cell. In some embodiments, the cell is a hepatocyte. In some embodiments, the cell is from a subject having a disease or disorder, optionally the disease or disorder is selected from the group comprising an autoimmune disease, a neurological disease or disorder, a cancer, an inflammatory disease, a cardiovascular disease, an infectious disease, a genetic disease, a trinucleotide repeat expansion disorder, a metabolic disease, or any combination thereof, optionally P -thalassemia, sickle cell disease, alphal antitrypsin deficiency disease, phenylketonuria, or hyperphenylalaninemia, further optionally the subject is a human.

[0019] Disclosed herein include pharmaceutical compositions. The pharmaceutical composition can comprise: (i) an editor or editing system disclosed herein, a base editor or base editing system disclosed herein, or a RT editor or RT editing system disclosed herein, a RNP disclosed herein, a LNP disclosed herein, a system disclosed herein, a polynucleotide disclosed herein, a vector disclosed herein, or a cell disclosed herein; and (ii) a pharmaceutically acceptable carrier.

[0020] Disclosed herein include methods for editing a double stranded target DNA. The method can comprise: contacting the double stranded target DNA with an editor or editing system disclosed herein, a base editor or base editing system disclosed herein, a RT editor or RT editing system disclosed herein, or a system disclosed herein, thereby editing the double strandedtarget DNA. In some embodiments, the double stranded target DNA is in a cell. In some embodiments, the cell is a mammalian cell, optionally a human cell. In some embodiments, the cell is a primary cell. In some embodiments, the cell is a hepatocyte. In some embodiments, the cell is a stem cell; optionally, an embryonic stem cell, an induced pluripotent stem cell, or an adult stem cell. In some embodiments, the cell is in a subject, optionally the subject is a human. In some embodiments, the cell is from a subject having a disease or disorder, optionally the disease or disorder is selected from the group comprising an autoimmune disease, a neurological disease or disorder, a cancer, an inflammatory disease, a cardiovascular disease, an infectious disease, a genetic disease, a trinucleotide repeat expansion disorder, a metabolic disease, or any combination thereof, optionally P-thalassemia, sickle cell disease, alphal antitrypsin deficiency disease, phenylketonuria, or hyperphenylalaninemia. The method can further comprise administering the cell to the subject after incorporation of the intended nucleotide edit.

[0021] Disclosed herein include cell(s) generated by a method disclosed herein. Disclosed herein include populations of cells generated by a method disclosed herein. Disclosed herein include methods for treating or preventing a disease or disorder in a subj ect in need thereof. The method can comprise administering to the subject an editor or editing system disclosed herein, a base editor or base editing system disclosed herein, a RT editor or RT editing system disclosed herein, a system disclosed herein, a RNP disclosed herein, a LNP disclosed herein, a pharmaceutical composition disclosed herein, a cell disclosed herein, or a population of cells disclosed herein, thereby treating or preventing the disease or disorder in the subject, optionally the disease or disorder is selected from the group comprising an autoimmune disease, a neurological disease or disorder, a cancer, an inflammatory disease, a cardiovascular disease, an infectious disease, a genetic disease, a trinucleotide repeat expansion disorder, a metabolic disease, or any combination thereof.BRIEF DESCRIPTION OF THE DRAWINGS

[0022] FIGS. 1 A-1D depict non-limiting exemplary schematics and alignments related to the design of novel Cas9 variants. FIG. 1A depicts the engineering of SsaCas9_AR00_12 through Ancestral Sequence Reconstruction (ASR) in the REC lobe, where the YTTKKDSEDEYI region of SsaCas9_0 is replaced by the aligned sequence of ScrCas9_26 or SorCas9_36, FRTDGT, to generate SsaCas9_AR00_12. Ten parental wild-type sequences were aligned. SEQ ID NOs shown are the full-length sequences. FIG. IB depicts an overall protein sequence alignment of 47 Cas9 variants disclosed herein. FIG. 1C depicts a detailed protein sequence alignment of 47 Cas9 variants disclosed herein. SEQ ID NOs shown are the full-length sequences. FIG. ID depicts a phylogenetic tree of 47 Cas9 variants disclosed herein. Blue colored variants are representative Cas9s that show high activity in human cells.

[0023] FIGS. 2A-2AU depict PAM-related heatmaps and bar plots related to a first PAM depletion study of SsaCas9_0 (FIG. 2A, favored PAM is NNAGAAAN), SsaCas9_AR00_l (FIG. 2B, favored PAM is NNAGAANN), SsaCas9_AR00_2 (FIG. 2C, favored PAM is NNAGAAAN), SsaCas9_AR00_4 (FIG. 2D, favored PAM is NNGNAANN), SsaCas9_AR00 11 (FIG. 2E, favored PAM is NNAGAANN), SsaCas9_AR00_12 (FIG. 2F, favored PAM is NNAGAAAN), SsaCas9_AR00_14 (FIG. 2G, favored PAM is NNAGAAAN), SsaCas9_AR01_17 (FIG. 2H, favored PAM is NNANAANN), SeqCas9_22 (FIG. 21, favored PAM is NNAGAAAN), SeqCas9_23 (FIG. 2J, favored PAM is NNRNNNNN), SveCas9_24 (FIG. 2K, favored PAM is NNANGANN), SveCas9_25 (FIG. 2L, favored PAM is NNAGAAAN), ScrCas9_26 (FIG. 2M, favored PAM is NNAYAAMN), ScrCas9_27 (FIG. 2N, favored PAM is NNAGAAAN), ScrCas9_29 (FIG. 20, favored PAM is NNAGAAAN), ScrCas9_31 (FIG. 2P, favored PAM is NNAGAAAN), ScrCas9_33 (FIG. 2Q, favored PAM is NNAGAAAN), SorCas9_35 (FIG. 2R, favored PAM is NNAGAAAN), SorCas9_36 (FIG. 2S, favored PAM is NNGAAANN), EntCas9_60 (FIG. 2T, favored PAM is NNGGNNCN), SsaCas9_AR00_3 (FIG. 2U, favored PAM is NNRGAACN), SsaCas9_AR00 _10 (FIG. 2V, favored PAM is NNNNAANN), SsaCas9_AR00_15 (FIG. 2W, favored PAM is NNANAANN), SsaCas9_AR00_16 (FIG. 2X, favored PAM is NNAGAAAN), SsaCas9_AR02_19 (FIG. 2Y, favored PAM is NNANAANN), SeqCas9_21 (FIG. 2Z, favored PAM is NNANAANN), ScrCas9_28 (FIG. 2AA, favored PAM is NNAYAACN), ScrCas9_30 (FIG. 2AB, favored PAM is NNAYAACN), ScrCas9_32 (FIG. 2AC, favored PAM is NNAYAACN), SorCas9_34 (FIG.2AD, favored PAM is NNAGAANN), SubCas9_41 (FIG. 2AE, favored PAM is NNAAAGNN), EntCas9_53 (FIG. 2AF, favored PAM is NNGGNNNN), EntCas9_58 (FIG. 2AG, favored PAM is NNGGNNNN), EntCas9_62 (FIG. 2AH, favored PAM is NNGGNNCN), VpeCas9_67 (FIG.2AI, favored PAM is NNAYAANN), SsaCas9_AR00_5 (FIG. 2AJ, favored PAM is NNAAAGNN), SsaCas9_AR00 _7 (FIG. 2AK, favored PAM is NNNNAANN), SsaCas9_AR00 _8 (FIG. 2AL, favored PAM is NNNNAANN), SsaCas9_AR00 _13 (FIG. 2AM, favored PAM is NNANAANN), SmiCas9_37 (FIG. 2AN, favored PAM is NNGNNNNN), SmiCas9_38 (FIG.2AO, favored PAM is NNAGAAAN), SubCas9_42 (FIG. 2AP, favored PAM is NNAAAGNN), SubCas9_44 (FIG. 2AQ, favored PAM is NNAAANNN), EdiCas9_64 (FIG. 2AR, favored PAM is NNAYAANN), SubCas9_46 (FIG. 2AS, favored PAM is NNAAAGNN), VpeCas9_68 (FIG.2AT, favored PAM is NNAGAANN), and VpeCas9_69 (FIG. 2AU, favored PAM is NNAGAANN).

[0024] FIGS. 3 A-3B depict non-limiting exemplary schematics of the workflow of the cell-based PAM validation assay (FIG. 3 A) and the Nuclease Cas9-RT construct employed (FIG.3B).

[0025] FIGS. 4A-4B depict non-limiting exemplary schematics related to gRNA design strategy based on the selected PAM. SEQ ID NOs shown are the full-length sequences.

[0026] FIG. 5 depicts combined data at SERPINA1, PAH and BCL11A loci with regards to the evaluation of Nuclease SsaCas9_AR00 12-MMLV RT activity in Huh7 cells.

[0027] FIGS. 6A-6D depict data at the PAH (FIG. 6A), SERPINA1 (FIG. 6B), and BCL11A (FIG. 6C) loci, as well as combined data at SERPINA1, PAH and BCL11A loci (FIG.6D), with regards to the evaluation of Nuclease SveCas9_24-MMLV_RT activity in Huh7 cells.

[0028] FIG. 7 depicts combined data at SERPINA1, PAH and BCL11A loci with regards to the evaluation of Nuclease SsaCas9_AR00 4-MMLV RT activity in Huh7 cells.

[0029] FIG. 8 depicts data related to the RT editing efficiency of Cas9 variants.

[0030] FIGS. 9A-9W depict PAM-related heatmaps and bar plots related to a second PAM depletion study of SveCas9_25 (FIG. 9A), SveCas9_24 (FIG. 9B), SsaCas9_AR01_17 (FIG. 9C), SsaCas9_AR00_14 (FIG. 9D), SsaCas9_AR00_12 (FIG. 9E), SsaCas9_AR00_l 1 (FIG. 9F), SsaCas9_AR00_4 (FIG. 9G), SsaCas9_AR00_2 (FIG. 9H), SsaCas9_AR00_l (FIG.91), SsaCas9_0 (PC) (FIG. 9 J), SorCas9_36 (FIG. 9K), SorCas9_35 (FIG. 9L), SorCas9_34 (FIG.9M), SmiCas9_38 (FIG. 9N), SeqCas9_23 (FIG. 90), SeqCas9_22 (FIG. 9P), ScrCas9_33 (FIG.9Q), ScrCas9_31 (FIG. 9R), ScrCas9_29 (FIG. 9S), ScrCas9_27 (FIG. 9T), ScrCas9_26 (FIG.9U), EntCas9_60 (FIG. 9V), and VpeCas9_69 (FIG. 9W).

[0031] FIGS. 10A-10B display editing efficiencies of further engineered nSveCas9_25 variants (FIG. 10A) and nSsaCas9 variants (FIG. 10B).

[0032] FIG. 11 displays editing efficiencies of nSsaCas9 variants engineered with iGeoCas9 mutations.

[0033] FIG. 12 displays mutation sites in Cas9 to reduce indels.

[0034] FIGS. 13A-13B displays frequencies of on-target editing (FIG. 13A) and indels (FIG. 13B) using the indicated variants.

[0035] FIG. 14 shows non-limiting exemplary data related to further indel analysis using the indicated Cas9 variants.

[0036] FIGS. 15A-15B displays editing frequencies (FIG. 15 A) and indels (FIG. 15B) using the indicated variants in primary human hepatocytes. The s23_F6E10 1F guide sequence C0mprises:mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmU mGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGrArUrArArGrGm CmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrUrUrUrCrCrCrGrUrCrGrArUrG*mG*mU*m C (SEQ ID NO: 825).

[0037] FIG. 16 displays exemplary data of editing efficiencies for a triple nickase Cas9 mutant.

[0038] FIG. 17 depicts a non-limiting exemplary alignment related to the design of novel Cas9 variants disclosed herein. SEQ ID NOs shown are the full-length sequences.

[0039] FIG. 18 depicts a non-limiting exemplary alignment related to the design of novel Cas9 variants disclosed herein. SEQ ID NOs shown are the full-length sequences.

[0040] FIG. 19 depicts a non-limiting exemplary alignment related to the design of novel Cas9 variants disclosed herein. SEQ ID NOs shown are the full-length sequences.DETAILED DESCRIPTION

[0041] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the Figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein and made part of the disclosure herein.

[0042] All patents, published patent applications, other publications, and sequences from GenBank, and other databases referred to herein are incorporated by reference in their entirety with respect to the related technology.

[0043] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs. See, e.g. Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For purposes of the present disclosure, the following terms are defined below.

[0044] As used herein, the term “about” can mean plus or minus 5% of the provided value.

[0045] As used herein, the term “gene editing” (including genomic editing) is a type of genetic engineering in which nucleotide(s) / nucleic acid(s) is / are inserted, deleted, and / or substituted in a DNA sequence, such as in the genome of a targeted cell. Targeted gene editing enables insertion, deletion, and / or substitution at pre-selected sites in the genome of a targeted cell (e.g., in a targeted gene or targeted DNA sequence). When a sequence of an endogenous gene is edited, for example by deletion, insertion or substitution of nucleotide(s) / nucleic acid(s), the endogenous gene comprising the affected sequence can be knocked-out or knocked-down due to the sequence alteration. Therefore, targeted editing can be used to disrupt endogenous geneexpression. In some embodiments, gene editing is precise editing, and comprises an intended nucleotide edit at an intended location (e.g., a single nucleotide substitution, such as an E342K correction). In some embodiments, precise gene editing involves the generation of a single nick or the generation of two or more nicks. Said nicks can be simultaneous or sequential, can occur on the same strand on or opposite strands, and can occur near each other (leading to a doublestrand break) or separated from each such that a double-strand break does not occur. Precise editing does not comprise a double-strand break in some embodiments.

[0046] As used herein, the term “RNA-guided endonuclease” refers to a polypeptide capable of binding a RNA (e.g., a gRNA) to form a complex targeted to a specific DNA sequence (e.g., in a target DNA). A non-limiting example of RNA-guided endonuclease is a Cas polypeptide (e.g., a Cas endonuclease, such as a Cas9 endonuclease). In some embodiments, the RNA-guided endonuclease as described herein is targeted to a specific DNA sequence in a target DNA by an RNA molecule to which it is bound. The RNA molecule can include a sequence that is complementary to and capable of hybridizing with a target sequence within the target DNA, thus allowing for targeting of the bound polypeptide to a specific location within the target DNA.

[0047] The term “deaminase” refers to an enzyme that catalyzes a deamination reaction. Base editors disclosed herein include deaminase-based base editors (dBEs) and deaminase-free glycosylase-based base editosr (gBEs). In some embodiments, the deaminase is a cytidine deaminase, catalyzing the hydrolytic deamination of cytidine or deoxycytidine to uracil or deoxyuracil, respectively. In some embodiments, the deaminase is an adenosine deaminase. In some embodiments, adenosine is converted to inosine after deamination. The systems, methods, compositions, and kits provided herein can, in some embodiments, be employed in concert with the systems, methods, compositions, and kits for base editing described in U.S. Patent Publication Nos. US20230235309A1, US20220411777A1, US20220170013A1, US20220282275A1, US20180312828A1, US20190093099A1, and in Tong, Huawei, et al. ("Development of deaminase-free T-to-S base editor and C-to-G base editor by engineered human uracil DNA glycosylase." Nature Communications 15.1 (2024): 4897), Ye, Lijun, et al. ("Glycosylase-based base editors for efficient T-to-G and C-to-G editing in mammalian cells." Nature biotechnology (2024): 1-10), He, Yan, et al. ("Protein language model s-assisted optimization of a uracil-N-glycosylase variant enables programmable T-to-G and T-to-C base editing." Molecular Cell 84.7 (2024): 1257-1270), the entire contents of which are incorporated herein by reference.

[0048] As used herein, the term “invariable region” of a gRNA refers to the nucleotide sequence of the gRNA that associates with the RNA-guided endonuclease. In some embodiments, the gRNA comprises a crRNA and a transactivating crRNA (tracrRNA), wherein the crRNA and tracrRNA hybridize to each other to form a duplex. In some embodiments, the crRNA comprises5’ to 3’: a spacer sequence and minimum CRISPR repeat sequence (also referred to as a “crRNA repeat sequence” herein); and the tracrRNA comprises a minimum tracrRNA sequence complementary to the minimum CRISPR repeat sequence (also referred to as a “tracrRNA antirepeat sequence” herein) and a 3’ tracrRNA sequence. In some embodiments, the invariable region of the gRNA refers to the portion of the crRNA that is the minimum CRISPR repeat sequence and the tracrRNA.

[0049] As used herein, the term “fusion protein” refers to a hybrid polypeptide which comprises protein domains from at least two different proteins. One protein may be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxy -terminal (C-terminal) protein thus forming an “amino-terminal fusion protein” or a “carboxy-terminal fusion protein,” respectively. A protein may comprise different domains, for example, a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9 that directs the binding of the protein to a target site) and a nucleic acid cleavage domain or a catalytic domain of a recombinase. In some embodiments, a protein comprises a proteinaceous part, e.g., an amino acid sequence constituting a nucleic acid binding domain, and an organic compound, e.g., a compound that can act as a nucleic acid cleavage agent. In some embodiments, a protein is in a complex with, or is in association with, a nucleic acid, e.g., RNA. Any of the proteins provided herein can be produced by any method known in the art. For example, the proteins provided herein may be produced via recombinant protein expression and purification, which is especially suited for fusion proteins comprising a peptide linker. Methods for recombinant protein expression and purification are well known, and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4thed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), the entire contents of which are incorporated herein by reference.

[0050] As used herein, the terms “DNA binding domain” and “nucleic acid programmable DNA binding protein” or “napDNAbp,” of which a Cas protein is an example, may be used interchangeably herein, and can refer to proteins that use RNA:DNA hybridization to target and bind to specific sequences in a DNA molecule. Each napDNAbp is associated with at least one guide nucleic acid (e.g., guide RNA), which localizes the napDNAbp to a DNA sequence that comprises a DNA strand (i.e., a target strand) that is complementary to the guide nucleic acid, or a portion thereof (e.g., the protospacer of a guide RNA). In other words, the guide nucleic-acid “programs” the napDNAbp (e.g., Cas9 or equivalent) to localize and bind to a complementary sequence. DNA binding domains can refer to proteins that bind DNA in a sequence-specific manner, such as, for example, TALENs and Zinc finger proteins. In some embodiments DNA binding domains can operate without a guide nucleic acid. In some embodiments, Cas9 comprises a napDNAbp and a DNA endonuclease domain. In some embodiments, a Cas9 polypeptidepossess both napDNAbp and DNA endonuclease activities. The term “endonuclease” can refer to an enzyme that can cleave the phosphodiester bond within a polynucleotide chain. The polynucleotide may be double-stranded DNA (dsDNA), single-stranded DNA (ssDNA), RNA, double-stranded hybrids of DNA and RNA, and synthetic DNA (for example, containing bases other than A, C, G, and T). An endonuclease may cut a polynucleotide symmetrically, leaving “blunt” ends, or in positions that are not directly opposing, creating overhangs, which may be referred to as “sticky ends.” Examples of endonucleases include Cas, Argonaute protein (AGO), TAL Effector Nuclease” (TALEN), and meganucleases such as MegaTAL, or a fusion protein comprising a domain of an endonuclease, for example, Cas9, Ago, TALEN, or MegaTAL, or one or more portions thereof. An endonuclease can be an RNA-guided endonuclease.

[0051] As used herein, the term “guide RNA” or “gRNA” can refer to a site-specific targeting RNA that can bind an RNA-guided endonuclease or napDNAbp to form a complex, and direct the activities of the bound RNA-guided endonuclease (such as a Cas endonuclease) or napDNAbp to a specific sequence within a target nucleic acid (e.g., a specific gene or region within a gene). The guide RNA can include one or more RNA molecules. In some embodiments, the gRNA is a template armed gRNA (tagRNA). In some embodiments, the gRNA is an enhancer gRNA (egRNA).

[0052] As used herein, the term “target DNA” can refer to the specific region on a double-stranded DNA within a subject’s genome intended for editing by a gene editing system. In certain embodiments, the gene editing system is an RT editing system.

[0053] As used herein, the term “search target sequence” can refer to the sequence on target strand that is complementary or substantially complementary to the spacer sequence of the tagRNA.

[0054] As used herein, the term “editing target sequence” can refer to the sequence being edited on the non-target strand.

[0055] As used herein, the term “silent mutation” can refer to a nucleotide change or nucleotide changes (for e.g., substitution or substitutions) in a DNA sequence that does not result in a change to the amino acid sequence of the protein the DNA sequence encodes.

[0056] As used herein, the term “protospacer” refers to the sequence in DNA adjacent to the PAM (protospacer adjacent motif) sequence. In some embodiments, the protospacer is 20 nucleotides long. The protospacer shares the same sequence as the spacer sequence of the guide RNA. The guide RNA anneals to the complement of the protospacer sequence on the target DNA (specifically, one strand thereof, i.e., the “target strand” versus the “non-target strand” of the target DNA sequence). In order for Cas9 to function it also requires a specific protospacer adjacent motif (PAM) that varies depending on the bacterial species of the Cas9 gene. The skilled person willappreciate that the literature in the state of the art sometimes refers to the “protospacer” as the target-specific guide sequence on the guide RNA itself, rather than referring to it as a “spacer.” Thus, in some cases, the term “protospacer” as used herein may be used interchangeably with the term “spacer.” The context of the description surrounding the appearance of either “protospacer” or “spacer” will help inform the reader as to whether the term is in reference to the gRNA or the DNA target.

[0057] As used herein, the terms “upstream” and “downstream” define relevant positions at least two regions or sequences in a nucleic acid molecule orientated in a 5'-to-3' direction. For example, a first sequence is upstream of a second sequence in a DNA molecule where the first sequence is positioned 5’ to the second sequence. Accordingly, the second sequence is downstream of the first sequence.

[0058] As used herein, the term “protospacer adjacent sequence” or “PAM” refers to a DNA sequence that is an important targeting component of a Cas9 nuclease. The PAM sequence can be on either strand, and is downstream in the 5' to 3' direction of the Cas9 cut site. Different PAM sequences can be associated with different Cas9 nucleases or equivalent proteins from different organisms.

[0059] As used herein, the term “spacer sequence” in connection with a guide RNA or a tagRNA refers to the portion of the guide RNA or tagRNA which contains a nucleotide sequence that shares the same sequence as the protospacer sequence in the target DNA sequence. The spacer sequence anneals to the complement of the protospacer sequence to form a ssRNA / ssDNA hybrid structure at the target site and a corresponding R loop ssDNA structure of the endogenous DNA strand.

[0060] As used herein, a “secondary structure” of a nucleic acid molecule (e.g., an RNA fragment, or a gRNA) refers to the base pairing interactions within the nucleic acid molecule.

[0061] As used herein, the term “Cas endonuclease” or “Cas nuclease” refers to an RNA-guided DNA endonuclease associated with and / or derived from the CRISPR adaptive immunity system. The term “nickase” refers to a Cas9 or other endonuclease with one of the two nuclease domains inactivated. This enzyme is capable of cleaving only one strand of a target DNA.

[0062] Unless otherwise indicated “nuclease” and “endonuclease” are used interchangeably herein to refer to an enzyme which possesses endonucleolytic catalytic activity for polynucleotide cleavage.

[0063] The terms “polynucleotide” and “nucleic acid” are used interchangeably herein and refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. A polynucleotide can be single-, double-, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids / triple helices, or a polymer including purine andpyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases. Any of the RNA sequences disclosed herein may also be DNA (either single-stranded or double-stranded), e.g., wherein “U” is converted to “T.” Any of the DNA sequences disclosed herein may also be RNA, e.g., wherein “T” is converted to “U ”

[0064] A “functional variant” or “functional mutant”, as used herein, refers to any variant or mutant of a reference protein (e.g., a wild-type protein) that encompasses one or more alterations to the amino acid sequence of the reference protein while retaining one or more of the functions, e.g., catalytic or binding functions. In some embodiments, the one or more alterations to the amino acid sequence comprises amino acid substitutions, insertions or deletions, or any combination thereof. In some embodiments, the one or more alterations to the amino acid sequence comprises amino acid substitutions. For example, a functional variant of a reverse transcriptase may comprise one or more amino acid substitutions compared to the amino acid sequence of a wild-type reverse transcriptase but retains the ability under at least one set of conditions to catalyze the polymerization of a polynucleotide. When the reference protein is a fusion of multiple functional domains, a functional variant thereof may retain one or more of the functions of at least one of the functional domains. For example, in some embodiments, a functional fragment of a Cas9 may comprise one or more amino acid substitutions in a nuclease domain, e.g., an H840A amino acid substitution, compared to the amino acid sequence of a wild type Cas9, but retains the DNA binding ability and lacks the nuclease activity partially or completely.

[0065] As used herein, the term “binding” refers to a non-covalent interaction between macromolecules (e.g., between a protein and a nucleic acid). While in a state of non-covalent interaction, the macromolecules are said to be “associated” or “interacting” or “binding” (e.g., when a molecule X is said to interact with a molecule Y, it means that the molecule X binds to molecule Y in a non-covalent manner). Binding interactions can be characterized by a dissociation constant (Kd), for example a Kd of, or a Kd less than, 10'6M, 10'7M, 10'8M, 10'9M, 10'10M, 10"11M, 10'12M, 10'13M, 10'14M, 10'15M, or a number or a range between any two of these values. Kd can be dependent on environmental conditions, e.g., pH and temperature. “Affinity” refers to the strength of binding, and increased binding affinity is correlated with a lower Kd.

[0066] As used herein, the term “hybridizing” or “hybridize” refers to the pairing of substantially complementary or complementary nucleic acid sequences within two different molecules. Pairing can be achieved by any process in which a nucleic acid sequence joins with a substantially or fully complementary sequence through base pairing to form a hybridization complex. “Hybridizing” or “hybridize” can comprise denaturing the molecules to disrupt the intramolecular structure(s) (e.g., secondary structure(s)) in the molecule. In some embodiments,denaturing the molecules comprises heating a solution comprising the molecules to a temperature sufficient to disrupt the intramolecular structures of the molecules. In some instances, denaturing the molecules comprises adjusting the pH of a solution comprising the molecules to a pH sufficient to disrupt the intramolecular structures of the molecules. For purposes of hybridization, two nucleic acid sequences or segments of sequences are “substantially complementary” if at least 80% of their individual bases are complementary to one another. The complementary portion of each sequence can be referred to herein as a “segment”, and the segments are substantially complementary if they have 80% or greater identity.

[0067] The terms “complementarity” and “complementary” mean that a nucleic acid can form hydrogen bond(s) with another nucleic acid based on traditional Watson-Crick base paring rule, that is, adenine (A) pairs with thymine (T, or uracil (U) in RNA) and guanine (G) pairs with cytosine (C). Complementarity can be perfect (e.g., complete complementarity) or imperfect (e.g. partial complementarity). Perfect or complete complementarity indicates that each and every nucleic acid base of one strand is capable of forming hydrogen bonds according to Watson-Crick canonical base pairing with a corresponding base in another, antiparallel nucleic acid sequence. Partial complementarity indicates that only a percentage of the contiguous residues of a nucleic acid sequence can form Watson-Crick base pairing with the same number of contiguous residues in another, antiparallel nucleic acid sequence. In some embodiments, the complementarity can be at least 70%, 80%, 90%, 100% or a number or a range between any two of these values. In some embodiments, the complementarity is perfect, z.e., 100%. For example, the complementary candidate sequence segment is perfectly complementary to the candidate sequence segment, whose sequence can be deduced from the candidate sequence segment using the Watson-Crick base pairing rules.

[0068] As used herein, the terms “nucleic acid" and “polynucleotide” are interchangeable and refer to any nucleic acid, whether composed of phosphodiester linkages or modified linkages such as phosphotriester, phosphoramidate, siloxane, carbonate, carboxymethylester, acetamidate, carbamate, thioether, bridged phosphoramidate, bridged methylene phosphonate, bridged phosphoramidate, bridged phosphoramidate, bridged methylene phosphonate, phosphorothioate, methylphosphonate, phosphorodithioate, bridged phosphorothioate or sultone linkages, and combinations of such linkages. The terms “nucleic acid” and “polynucleotide” also specifically include nucleic acids composed of bases other than the five biologically occurring bases (adenine, guanine, thymine, cytosine and uracil).

[0069] The terms “DNA editing efficiency,” or “editing efficiency” may be used interchangeably herein and can refer to the number or proportion of intended target sequences that are edited. In some embodiments, the efficiency can be reported as % indel, e.g., the proportionof insertions and / or deletions detected in the target sequence. Indels (e.g., insertion-deletions) can result from repair of double-stranded DNA breaks caused by Cas9 cleavage by processes including, but not limited to, non-homologous end joining (NHEJ) repair.

[0070] The term “off-target editing frequency,” as used herein, refers to the number or proportion of unintended DNA sequences that are edited. On-target and off-target editing frequencies may be measured by the methods and assays described herein, further in view of techniques known in the art, including high-throughput sequencing reads. As used herein, high-throughput sequencing involves the hybridization of nucleic acid primers (e.g., DNA primers) with complementarity to nucleic acid (e.g., DNA) regions just upstream or downstream of the target sequence or off-target sequence of interest. Since many of the Cas9-dependent off-target sites have high sequence identity to the target site of interest, nucleic acid primers with sufficient complementarity to regions upstream or downstream of the Cas9-dependent off-target site may be designed using techniques and kits known in the art. These kits make use of polymerase chain reaction (PCR) amplification, which produces amplicons as intermediate products. The target and off-target sequences may comprise genomic loci that further comprise protospacers and PAMs. Accordingly, the term “amplicons,” as used herein, may refer to nucleic acid molecules that constitute the aggregates of genomic loci, protospacers and PAMs. High-throughput sequencing techniques used herein may further include Sanger sequencing and / or whole genome sequencing (WGS).

[0071] As used herein, the terms “transfection” or “infection” refer to the introduction of a nucleic acid into a host cell, such as by contacting the cell with liposomes or nanoparticles (e.g., lipid nanoparticles) as described herein.

[0072] As used herein, “treatment” refers to a clinical intervention made in response to a disease, disorder or physiological condition manifested by a patient or to which a patient may be susceptible. The aim of treatment includes, but is not limited to, the alleviation or prevention of symptoms, slowing or stopping the progression or worsening of a disease, disorder, or condition and / or the remission of the disease, disorder or condition. “Treatments” refer to one or both of therapeutic treatment and prophylactic or preventative measures. Subjects in need of treatment include those already affected by a disease or disorder or undesired physiological condition as well as those in which the disease or disorder or undesired physiological condition is to be prevented.

[0073] As used herein, the terms “effective amount” or “pharmaceutically effective amount” or “therapeutically effective amount” refer to an amount sufficient to effect beneficial or desirable biological and / or clinical results.

[0074] The term “pharmaceutically acceptable excipient” as used herein refers to anysuitable substance that provides a pharmaceutically acceptable carrier, additive or diluent for administration of a compound(s) of interest to a subject. Pharmaceutically acceptable excipients can encompass substances referred to as pharmaceutically acceptable diluents, pharmaceutically acceptable additives, and pharmaceutically acceptable carriers.

[0075] As used herein, a “subject” refers to an animal for whom a diagnosis, treatment, or therapy is desired. In some embodiments, the subject is a mammal. “Mammal,” as used herein, refers to an individual belonging to the class Mammalia and includes, but not limited to, humans,. In some embodiments, the mammal is a primate. In some embodiments, the mammal is a human. In some embodiments, the mammal is not a human. In some embodiments, the subject has or is suspected of having a disease that can be corrected via gene editing. In some embodiments, the gene editing system is an RT editing system. In some embodiments, the subject has or is suspected of having P-thalassemia or sickle cell disease. In some embodiments, the subject has or is suspected of having alphal antitrypsin deficiency disease. In some embodiments, the subject has or is suspected of having phenylketonuria or hyperphenylalaninemia.Gene Editing or Genome Editing

[0076] Genome editing generally refers to the process of modifying the nucleotide sequence of a genome, preferably in a precise or pre-determined manner. Examples of methods of genome editing described herein include methods of using site-directed nucleases to cut deoxyribonucleic acid (DNA) at precise target locations in the genome, thereby creating singlestrand or double-strand DNA breaks at particular locations within the genome. Such breaks can be and regularly are repaired by natural, endogenous cellular processes, such as homology-directed repair (HDR) and non-homologous end joining (NHEJ), as described in Cox et al., “Therapeutic genome editing: prospects and challenges,”, Nature Medicine, 2015, 21(2), 121-31. These two main DNA repair processes consist of a family of alternative pathways. NHEJ directly joins the DNA ends resulting from a double-strand break, sometimes with the loss or addition of nucleotide sequence, which may disrupt or enhance gene expression. HDR utilizes a homologous sequence, or donor sequence, as a template for inserting a defined DNA sequence at the break point. The homologous sequence can be in the endogenous genome, such as a sister chromatid. Alternatively, the donor sequence can be an exogenous polynucleotide, such as a plasmid, a singlestrand oligonucleotide, a double-stranded oligonucleotide, a duplex oligonucleotide or a virus, that has regions (e.g., left and right homology arms) of high homology with the nuclease-cleaved locus, but which can also contain additional sequence or sequence changes including deletions that can be incorporated into the cleaved target locus. A third repair mechanism can be microhomology-mediated end joining (MMEJ), also referred to as "Alternative NHEJ,” in which the genetic outcome is similar to NHEJ in that small deletions and insertions can occur at thecleavage site. MMEJ can make use of homologous sequences of a few base pairs flanking the DNA break site to drive a more favored DNA end joining repair outcome, and recent reports have further elucidated the molecular mechanism of this process; see, e.g., Cho and Greenberg, Nature, 2015, 518, 174-76; Kent et al., Nature Structural and Molecular Biology, 2015, 22(3):230-7; Mateos-Gomez et al., Nature, 2015, 518, 254-57; Ceccaldi et al., Nature, 2015, 528, 258-62. In some instances, it may be possible to predict likely repair outcomes based on analysis of potential microhomologies at the site of the DNA break.

[0077] Each of these genome editing mechanisms can be used to create desired genetic modifications. A step in the genome editing process can be to create one or two DNA breaks, the latter as double-strand breaks or as two single-stranded breaks, in the target locus as near the site of intended mutation. This can be achieved via the use of endonucleases, as described herein.

[0078] There are provided, in some embodiments, editors and editing systems. The editor can comprise an engineered Cas9 protein disclosed herein, a nuclease disclosed herein (e.g., a nickase comprising an amino acid sequence that is at least 80% identical to any one of the sequences of SEQ ID NOs: 1-93, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to any one of the sequences of SEQ ID NOs: 1-93), a nickase disclosed herein (e.g., a nickase comprising an amino acid sequence that is at least 80% identical to any one of the sequences of SEQ ID NOs: 48-94 and 748-770, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to any one of the sequences of SEQ ID NOs: 48-94 and 748-770). Disclosed herein include editing systems. The editing system can comprise: an editor disclosed herein and a guide RNA (gRNA). In some embodiments, the editor comprises an effector domain. In some embodiments, the effector domain comprises nuclease activity, nickase activity, recombinase activity, deaminase activity, methyltransferase activity, methylase activity, acetylase activity, acetyltransferase activity, transcriptional activation activity, transcriptional repression activity, and / or polymerase activity. In some embodiments, the effector domain comprises a non-long-terminal-repeat (non-LTR) retrotransposable element enzyme or a portion thereof, optionally derived from CRE, R2, Randl / Dualen, R4, NeSL, Hero, Proto 1, LI, Txl, Proto2, RTETP, RTEX, RTE, I, Outcast, Nimb, Ingi, Jockey, Rl, Loa, Tadl, Rexl, CR1, L2A, L2B, L2, Daphne, Crack, Vingis, or any combination thereof. In some embodiments, the effector domain is a nucleic acid editing domain. In some embodiments, the nucleic acid editing domain comprises a deaminase domain.CRISPR Endonuclease System

[0079] The CRISPR-endonuclease system is a naturally occurring defense mechanism in prokaryotes that has been repurposed as a RNA-guided DNA-targeting platform used for gene editing. Accordingly, in some embodiments, a CRISPR-endonuclease system is utilized tointroduce a gene edit into a cell. CRISPR systems include Types I, II, III, IV, V, and VI systems. In some embodiments, the CRISPR system is a Type II CRISPR / Cas9 system. In some embodiments, the CRISPR system is a Type V CRISPR / Cprf system. CRISPR systems rely on a DNA endonuclease, e.g., Cas9, and two noncoding RNAs - crisprRNA (crRNA) and transactivating RNA (tracrRNA) - to target the cleavage of DNA.

[0080] The crRNA drives sequence recognition and specificity of the CRISPR-endonuclease complex through Watson-Crick base pairing, typically with a ~20 nucleotide (nt) sequence in the target DNA. Changing the sequence of the 5’ 20 nt in the crRNA allows targeting of the CRISPR-endonuclease complex to specific loci. The CRISPR-endonuclease complex only binds DNA sequences that contain a sequence match to the first 20 nt of the single-guide RNA (sgRNA) if the target sequence is followed by a specific short DNA motif (with the sequence NGG) referred to as a protospacer adjacent motif (PAM).

[0081] TracrRNA hybridizes with the 3’ end of crRNA to form an RNA-duplex structure that is bound by the endonuclease to form the catalytically active CRISPR-endonuclease complex, which can then cleave the target DNA.

[0082] Once the CRISPR-endonuclease complex is bound to DNA at a target site, two independent nuclease domains within the endonuclease each cleave one of the DNA strands three bases upstream of the PAM site, leaving a double-strand break (DSB) where both strands of the DNA terminate in a base pair (a blunt end).

[0083] In some embodiments, the endonuclease is a Cas9 (CRISPR associated protein 9). In some embodiments, the Cas9 endonuclease is from Streptococcus pyogenes, although other Cas9 homologs may be used, e.g., S. aureus Cas9, N. meningitidis Cas9, S. thermophilus CRISPR 1 Cas9, S. thermophilus CRISPR 3 Cas9, or T. denticola Cas9. In some embodiments, the CRISPR endonuclease is Cpfl, e.g, L. bacterium ND2006 Cpfl or Acidaminococcus sp. BV3L6 Cpfl. In some embodiments, the endonuclease is Casl, CaslB, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csxl2), CaslOO, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4, or Cpfl endonuclease. In some embodiments, wild-type variants may be used. In some embodiments, modified versions (e.g, a homolog thereof, a recombination of the naturally occurring molecule thereof, codon-optimized thereof, or modified versions thereof) of the preceding endonucleases may be used.

[0084] The CRISPR nuclease can be linked to at least one nuclear localization signal (NLS). The at least one NLS can be located at or within 50 amino acids of the aminoterminus of the CRISPR nuclease and / or at least one NLS can be located at or within 50 amino acids of thecarboxy -terminus of the CRISPR nuclease.

[0085] Exemplary CRISPR / Cas polypeptides include the Cas9 polypeptides as published in Fonfara et al., “Phylogeny of Cas9 determines functional exchangeability of dual-RNA and Cas9 among orthologous type II CRISPR-Cas systems,” Nucleic Acids Research, 2014, 42: 2577-2590. The CRISPR / Cas gene naming system has undergone extensive rewriting since the Cas genes were discovered. Fonfara et al. also provides PAM sequences for the Cas9 polypeptides from various species.RNA-Guided Endonucleases

[0086] The RNA-guided endonuclease systems as used herein can comprise an amino acid sequence having at least 10%, at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% amino acid sequence identity to a wild-type exemplary endonuclease, e.g., Cas9 from S. pyogenes, US2014 / 0068797 Sequence ID No. 8 or Sapranauskas et al., Nucleic Acids Res, 39(21): 9275-9282 (2011). The endonuclease can comprise at least 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity to a wild-type endonuclease (e.g., Cas9 from S. pyogenes, supra) over 10 contiguous amino acids. The endonuclease can comprise at most: 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity to a wild-type endonuclease (e.g., Cas9 from S. pyogenes, supra) over 10 contiguous amino acids. The endonuclease can comprise at least: 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity to a wild-type endonuclease (e.g., Cas9 from S. pyogenes, supra) over 10 contiguous amino acids in a HNH nuclease domain of the endonuclease. The endonuclease can comprise at most: 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity to a wild-type endonuclease (e.g., Cas9 from S. pyogenes, supra) over 10 contiguous amino acids in a HNH nuclease domain of the endonuclease. The endonuclease can comprise at least: 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity to a wildtype endonuclease (e.g., Cas9 from S. pyogenes, supra) over 10 contiguous amino acids in a RuvC nuclease domain of the endonuclease. The endonuclease can comprise at most: 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity to a wild-type endonuclease (e.g., Cas9 from S. pyogenes, supra) over 10 contiguous amino acids in a RuvC nuclease domain of the endonuclease.

[0087] The endonuclease can comprise a modified form of a wild-type exemplary endonuclease. The modified form of the wild-type exemplary endonuclease can comprise a mutation that reduces the nucleic acid-cleaving activity of the endonuclease. The modified form of the wild-type exemplary endonuclease can have less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nucleic acid-cleaving activity of the wild-type exemplary endonuclease (e.g., Cas9 from S. pyogenes, supra). The modified form of the endonuclease canhave no substantial nucleic acid-cleaving activity. When an endonuclease is a modified form that has no substantial nucleic acid-cleaving activity, it is referred to herein as "enzymatically inactive."

[0088] Mutations contemplated can include substitutions, additions, and deletions, or any combination thereof. The mutation converts the mutated amino acid to alanine. The mutation converts the mutated amino acid to another amino acid (e.g., glycine, serine, threonine, cysteine, valine, leucine, isoleucine, methionine, proline, phenylalanine, tyrosine, tryptophan, aspartic acid, glutamic acid, asparagine, glutamine, histidine, lysine, or arginine). The mutation converts the mutated amino acid to a non-natural amino acid (e.g., selenomethionine). The mutation converts the mutated amino acid to amino acid mimics (e.g., phosphomimics). The mutation can be a conservative mutation. For example, the mutation converts the mutated amino acid to amino acids that resemble the size, shape, charge, polarity, conformation, and / or retainers of the mutated amino acids (e.g., cysteine / serine mutation, lysine / asparagine mutation, histidine / phenylalanine mutation). The mutation can cause a shift in reading frame and / or the creation of a premature stop codon. Mutations can cause changes to regulatory regions of genes or loci that affect expression of one or more genes.Guide RNAs

[0089] In some embodiments, a guide RNA (gRNA) that can direct the activities of an associated endonuclease to a specific target site within a polynucleotide is used to introduce a gene edit into a cell. A guide RNA can comprise at least a spacer sequence that hybridizes to a target nucleic acid sequence of interest, and a CRISPR repeat sequence. In CRISPR Type II systems, the gRNA also comprises a second RNA called the tracrRNA sequence. In the CRISPR Type II guide RNA (gRNA), the CRISPR repeat sequence and tracrRNA sequence hybridize to each other to form a duplex. In CRISPR Type V systems, the gRNA comprises a crRNA that forms a duplex. In some embodiments, a gRNA can bind an endonuclease, such that the gRNA and endonuclease form a complex. The gRNA can provide target specificity to the complex by virtue of its association with the endonuclease. The genome-targeting nucleic acid thus can direct the activity of the endonuclease.

[0090] Exemplary guide RNAs include a spacer sequence that comprises 15-200 nucleotides wherein the gRNA targets a genome location based on the GRCh38 human genome assembly. As is understood by the person of ordinary skill in the art, each gRNA can be designed to include a spacer sequence complementary to its genomic target site or region. See Jinek et al., Science, 2012, 337, 816-821 andDeltcheva et al., Nature, 2011, 471, 602-607.

[0091] The gRNA can be a double-molecule guide RNA. The gRNA can be a single molecule guide RNA.

[0092] A double -molecule guide RNA can comprise two strands of RNA. The first strand comprises in the 5' to 3' direction, an optional spacer extension sequence, a spacer sequence and a minimum CRISPR repeat sequence. The second strand can comprise a minimum tracrRNA sequence (complementary to the minimum CRISPR repeat sequence), a 3’ tracrRNA sequence and an optional tracrRNA extension sequence.

[0093] A single -molecule guide RNA (sgRNA) can comprise, in the 5' to 3' direction, an optional spacer extension sequence, a spacer sequence, a minimum CRISPR repeat sequence, a single-molecule guide linker, a minimum tracrRNA sequence, a 3’ tracrRNA sequence and an optional tracrRNA extension sequence. The optional tracrRNA extension can comprise elements that contribute additional functionality (e.g., stability) to the guide RNA. The single-molecule guide linker can link the minimum CRISPR repeat and the minimum tracrRNA sequence to form a hairpin structure. The optional tracrRNA extension can comprise one or more hairpins.

[0094] In some embodiments, a sgRNA comprises a 20 nucleotide spacer sequence at the 5’ end of the sgRNA sequence. In some embodiments, a sgRNA comprises a less than a 20 nucleotide spacer sequence at the 5’ end of the sgRNA sequence. In some embodiments, a sgRNA comprises a more than 20 nucleotide spacer sequence at the 5’ end of the sgRNA sequence. In some embodiments, a sgRNA comprises a variable length spacer sequence with 17-30 nucleotides at the 5’ end of the sgRNA sequence. In some embodiments, a sgRNA comprises a spacer extension sequence with a length of more than 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, or 200 nucleotides. In some embodiments, a sgRNA comprises a spacer extension sequence with a length of less than 3, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides.

[0095] In some embodiments, a sgRNA comprises a spacer extension sequence that comprises another moiety (e.g., a stability control sequence, an endoribonuclease binding sequence, or a ribozyme). The moiety can decrease or increase the stability of a nucleic acid targeting nucleic acid. The moiety can be a transcriptional terminator segment (i.e., a transcription termination sequence). The moiety can function in a eukaryotic cell. The moiety can function in a prokaryotic cell. The moiety can function in both eukaryotic and prokaryotic cells. Non-limiting examples of suitable moi eties include: a 5' cap (e.g., a 7-methylguanylate cap (m7 G)), a riboswitch sequence (e.g., to allow for regulated stability and / or regulated accessibility by proteins and protein complexes), a sequence that forms a dsRNA duplex (i.e., a hairpin), a sequence that targets the RNA to a subcellular location (e.g., nucleus, mitochondria, chloroplasts, and the like), a modification or sequence that provides for tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, a sequence that allows for fluorescent detection, etc.), and / or a modification or sequence that provides a binding site forproteins e.g., proteins that act on DNA, including transcriptional activators, transcriptional repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, and the like).

[0096] In some embodiments, a sgRNA comprises a spacer sequence that hybridizes to a sequence in a target polynucleotide. The spacer of a gRNA can interact with a target polynucleotide in a sequence-specific manner via hybridization (i.e., base pairing). The nucleotide sequence of the spacer can vary depending on the sequence of the target nucleic acid of interest.

[0097] In a CRISPR-endonuclease system, a spacer sequence can be designed to hybridize to a target polynucleotide that is located 5' of a PAM of the endonuclease used in the system. The spacer may perfectly match the target sequence or may have mismatches. Each endonuclease, e.g., Cas9 nuclease, has a particular PAM sequence that it recognizes in a target DNA. For example, S. pyogenes Cas9 recognizes a PAM that comprises the sequence 5'-NRG-3', where R comprises either A or G, where N is any nucleotide and N is immediately 3' of the target nucleic acid sequence targeted by the spacer sequence.

[0098] A target polynucleotide sequence can comprise 20 nucleotides. The target polynucleotide can comprise less than 20 nucleotides. The target polynucleotide can comprise more than 20 nucleotides. The target polynucleotide can comprise at least: 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30 or more nucleotides. The target polynucleotide can comprise at most: 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30 or more nucleotides. The target polynucleotide sequence can comprise 20 bases immediately 5' of the first nucleotide of the PAM.

[0099] A spacer sequence that hybridizes to a target polynucleotide can have a length of at least about 6 nucleotides (nt). The spacer sequence can be at least about 6 nt, at least about 10 nt, at least about 15 nt, at least about 18 nt, at least about 19 nt, at least about 20 nt, at least about 25 nt, at least about 30 nt, at least about 35 nt or at least about 40 nt, from about 6 nt to about 80 nt, from about 6 nt to about 50 nt, from about 6 nt to about 45 nt, from about 6 nt to about 40 nt, from about 6 nt to about 35 nt, from about 6 nt to about 30 nt, from about 6 nt to about 25 nt, from about 6 nt to about 20 nt, from about 6 nt to about 19 nt, from about 10 nt to about 50 nt, from about 10 nt to about 45 nt, from about 10 nt to about 40 nt, from about 10 nt to about 35 nt, from about 10 nt to about 30 nt, from about 10 nt to about 25 nt, from about 10 nt to about 20 nt, from about 10 nt to about 19 nt, from about 19 nt to about 25 nt, from about 19 nt to about 30 nt, from about 19 nt to about 35 nt, from about 19 nt to about 40 nt, from about 19 nt to about 45 nt, from about 19 nt to about 50 nt, from about 19 nt to about 60 nt, from about 20 nt to about 25 nt, from about 20 nt to about 30 nt, from about 20 nt to about 35 nt, from about 20 nt to about 40 nt, from about 20 nt to about 45 nt, from about 20 nt to about 50 nt, or from about 20 nt to about 60 nt. In some examples, the spacer sequence can comprise 20 nucleotides. In some examples, thespacer can comprise 19 nucleotides. In some examples, the spacer can comprise 18 nucleotides. In some examples, the spacer can comprise 22 nucleotides.

[0100] In some examples, the percent complementarity between the spacer sequence and the target nucleic acid is at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, or 100%. In some examples, the percent complementarity between the spacer sequence and the target nucleic acid is at most about 30%, at most about 40%, at most about 50%, at most about 60%, at most about 65%, at most about 70%, at most about 75%, at most about 80%, at most about 85%, at most about 90%, at most about 95%, at most about 97%, at most about 98%, at most about 99%, or 100%. In some examples, the percent complementarity between the spacer sequence and the target nucleic acid is 100% over the six contiguous 5 '-most nucleotides of the target sequence of the complementary strand of the target nucleic acid. The percent complementarity between the spacer sequence and the target nucleic acid can be at least 60% over about 20 contiguous nucleotides. The length of the spacer sequence and the target nucleic acid can differ by 1 to 6 nucleotides, which may be thought of as a bulge or bulges.

[0101] A tracrRNA sequence can comprise nucleotides that hybridize to a minimum CRISPR repeat sequence in a cell. A minimum tracrRNA sequence and a minimum CRISPR repeat sequence may form a duplex, i.e. a base-paired double-stranded structure. Together, the minimum tracrRNA sequence and the minimum CRISPR repeat can bind to an RNA-guided endonuclease. At least a part of the minimum tracrRNA sequence can hybridize to the minimum CRISPR repeat sequence. The minimum tracrRNA sequence can be at least about 30%, about 40%, about 50%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or 100% complementary to the minimum CRISPR repeat sequence.

[0102] The minimum tracrRNA sequence can have a length from about 7 nucleotides to about 100 nucleotides. For example, the minimum tracrRNA sequence can be from about 7 nucleotides (nt) to about 50 nt, from about 7 nt to about 40 nt, from about 7 nt to about 30 nt, from about 7 nt to about 25 nt, from about 7 nt to about 20 nt, from about 7 nt to about 15 nt, from about 8 nt to about 40 nt, from about 8 nt to about 30 nt, from about 8 nt to about 25 nt, from about 8 nt to about 20 nt, from about 8 nt to about 15 nt, from about 15 nt to about 100 nt, from about 15 nt to about 80 nt, from about 15 nt to about 50 nt, from about 15 nt to about 40 nt, from about 15 nt to about 30 nt or from about 15 nt to about 25 nt long. The minimum tracrRNA sequence can be approximately 9 nucleotides in length. The minimum tracrRNA sequence can be approximately 12 nucleotides. The minimum tracrRNA can consist of tracrRNA nt 23-48 described in Jinek et al., supra.

[0103] The minimum tracrRNA sequence can be at least about 60% identical to a reference minimum tracrRNA (e.g., wild type, tracrRNA from S. pyogenes) sequence over a stretch of at least 6, 7, or 8 contiguous nucleotides. For example, the minimum tracrRNA sequence can be at least about 65% identical, about 70% identical, about 75% identical, about 80% identical, about 85% identical, about 90% identical, about 95% identical, about 98% identical, about 99% identical or 100% identical to a reference minimum tracrRNA sequence over a stretch of at least 6, 7, or 8 contiguous nucleotides.

[0104] The duplex between the minimum CRISPR RNA and the minimum tracrRNA can comprise a double helix. The duplex between the minimum CRISPR RNA and the minimum tracrRNA can comprise at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotides. The duplex between the minimum CRISPR RNA and the minimum tracrRNA can comprise at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotides.

[0105] The duplex can comprise a mismatch (i.e., the two strands of the duplex are not 100% complementary). The duplex can comprise at least about 1, 2, 3, 4, or 5 or mismatches. The duplex can comprise at most about 1, 2, 3, 4, or 5 or mismatches. The duplex can comprise no more than 2 mismatches.

[0106] In some embodiments, a tracrRNA may be a 3' tracrRNA. In some embodiments, a 3’ tracrRNA sequence can comprise a sequence with at least about 30%, about 40%, about 50%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or 100% sequence identity to a reference tracrRNA sequence (e.g., a tracrRNA from S. pyogenes).

[0107] In some embodiments, a gRNA may comprise a tracrRNA extension sequence. A tracrRNA extension sequence can have a length from about 1 nucleotide to about 400 nucleotides. The tracrRNA extension sequence can have a length of more than 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, or 200 nucleotides. The tracrRNA extension sequence can have a length from about 20 to about 5000 or more nucleotides. The tracrRNA extension sequence can have a length of less than 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides. The tracrRNA extension sequence can comprise less than 10 nucleotides in length. The tracrRNA extension sequence can be 10-30 nucleotides in length. The tracrRNA extension sequence can be 30-70 nucleotides in length.

[0108] The tracrRNA extension sequence can comprise a functional moiety (e.g., a stability control sequence, ribozyme, endoribonuclease binding sequence). The functional moiety can comprise a transcriptional terminator segment (i.e., a transcription termination sequence). The functional moiety can have a total length from about 10 nucleotides (nt) to about 100 nucleotides, from about 10 nt to about 20 nt, from about 20 nt to about 30 nt, from about 30 nt to about 40 nt,from about 40 nt to about 50 nt, from about 50 nt to about 60 nt, from about 60 nt to about 70 nt, from about 70 nt to about 80 nt, from about 80 nt to about 90 nt, or from about 90 nt to about 100 nt, from about 15 nt to about 80 nt, from about 15 nt to about 50 nt, from about 15 nt to about 40 nt, from about 15 nt to about 30 nt, or from about 15 nt to about 25 nt.

[0109] In some embodiments, a sgRNA may comprise a linker sequence with a length from about 3 nucleotides to about 100 nucleotides. In Jinek et al., supra, for example, a simple 4 nucleotide "tetraloop" (-GAAA-) was used (Jinek et al., Science, 2012, 337(6096): 816-821). An illustrative linker has a length from about 3 nucleotides (nt) to about 90 nt, from about 3 nt to about 80 nt, from about 3 nt to about 70 nt, from about 3 nt to about 60 nt, from about 3 nt to about 50 nt, from about 3 nt to about 40 nt, from about 3 nt to about 30 nt, from about 3 nt to about 20 nt, from about 3 nt to about 10 nt. For example, the linker can have a length from about 3 nt to about 5 nt, from about 5 nt to about 10 nt, from about 10 nt to about 15 nt, from about 15 nt to about 20 nt, from about 20 nt to about 25 nt, from about 25 nt to about 30 nt, from about 30 nt to about 35 nt, from about 35 nt to about 40 nt, from about 40 nt to about 50 nt, from about 50 nt to about 60 nt, from about 60 nt to about 70 nt, from about 70 nt to about 80 nt, from about 80 nt to about 90 nt, or from about 90 nt to about 100 nt. The linker of a single-molecule guide nucleic acid can be between 4 and 40 nucleotides. The linker can be at least about 100, 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, or 7000 or more nucleotides. The linker can be at most about 100, 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, or 7000 or more nucleotides.

[0110] Linkers can comprise any of a variety of sequences, although in some examples the linker will not comprise sequences that have extensive regions of homology with other portions of the guide RNA, which might cause intramolecular binding that could interfere with other functional regions of the guide. In Jinek et al., supra, a simple 4 nucleotide sequence -GAAA- was used (Jinek et al., Science, 2012, 337(6096):816-821), but numerous other sequences, including longer sequences can likewise be used.[oni] The linker sequence can comprise a functional moiety. For example, the linker sequence can comprise one or more features, including an aptamer, a ribozyme, a proteininteracting hairpin, a protein binding site, a CRISPR array, an intron, or an exon. The linker sequence can comprise at least about 1, 2, 3, 4, or 5 or more functional moieties. In some examples, the linker sequence can comprise at most about 1, 2, 3, 4, or 5 or more functional moieties.

[0112] In some embodiments, a sgRNA does not comprise a uracil, e.g., at the 3 ’end of the sgRNA sequence. In some embodiments, a sgRNA does comprise one or more uracils, e.g., at the 3’end of the sgRNA sequence. In some embodiments, a sgRNA comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 uracils (U) at the 3’ end of the sgRNA sequence.

[0113] A sgRNA may be chemically modified. In some embodiments, a chemically modified gRNA is a gRNA that comprises at least one nucleotide with a chemical modification, e.g., a 2'-O-methyl sugar modification. In some embodiments, a chemically modified gRNA comprises a modified nucleic acid backbone. In some embodiments, a chemically modified gRNA comprises a 2'-O-methyl-phosphorothioate residue. In some embodiments, chemical modifications enhance stability, reduce the likelihood or degree of innate immune response, and / or enhance other attributes, as described in the art.

[0114] In some embodiments, a modified gRNA may comprise a modified backbones, for example, phosphor othi oates, phosphotriesters, morpholines, methyl phosphonates, short chain alkyl or cycloalkyl intersugar linkages or short chain heteroatomic or heterocyclic intersugar linkages.

[0115] Morpholino-based compounds are described in Braasch and David Corey, Biochemistry, 2002, 41(14): 4503-4510; Genesis, 2001, Volume 30, Issue 3; Heasman, Dev. Biol., 2002, 243: 209-214; Nasevicius et al., Nat. Genet., 2000, 26:216-220; Lacerra et al., Proc. Natl. Acad. Sci., 2000, 97: 9591-9596.; and U.S. Pat. No. 5,034,506, issued Jul. 23, 1991.

[0116] Cyclohexenyl nucleic acid oligonucleotide mimetics are described in Wang et al., J. Am. Chem. Soc., 2000, 122: 8595-8602.

[0117] In some embodiments, a modified gRNA may comprise one or more substituted sugar moieties, e.g., one of the following at the 2' position: OH, SH, SCH3, F, OCN, OCH3, OCH3 O(CH2)n CH3, O(CH2)n NH2, or O(CH2)n CH3, where n is from 1 to about 10; Cl to CIO lower alkyl, alkoxyalkoxy, substituted lower alkyl, alkaryl or aralkyl; Cl; Br; CN; CF3; OCF3; O-, S-, orN-alkyl; O-, S-, orN-alkenyl; SOCH3; SO2 CH3; ONO2; NO2; N3; NH2; heterocycloalkyl; heterocycloalkaryl; aminoalkylamino; polyalkylamino; substituted silyl; an RNA cleaving group; a reporter group; an intercalator; 2'-O-(2 -methoxy ethyl); 2'-methoxy (2'-O-CH3); 2'-propoxy (2'-OCH2 CH2CH3); and 2'-fluoro (2'-F). Similar modifications may also be made at other positions on the gRNA, particularly the 3' position of the sugar on the 3' terminal nucleotide and the 5' position of 5' terminal nucleotide. In some examples, both a sugar and an inter-nucleoside linkage, i.e., the backbone, of the nucleotide units can be replaced with novel groups.

[0118] Guide RNAs can also include, additionally or alternatively, nucleobase (often referred to in the art simply as "base") modifications or substitutions. As used herein, "unmodified" or "natural" nucleobases include adenine (A), guanine (G), thymine (T), cytosine (C), and uracil (U). Modified nucleobases include nucleobases found only infrequently or transiently in natural nucleic acids, e.g., hypoxanthine, 6-methyladenine, 5-Me pyrimidines, particularly 5 -methylcytosine (also referred to as 5 -methyl -2' deoxy cytosine and often referredto in the art as 5-Me-C), 5-hydroxymethylcytosine (HMC), glycosyl HMC and gentobiosyl HMC, as well as synthetic nucleobases, e.g., 2-aminoadenine, 2-(methylamino)adenine, 2-(imidazolylalkyl)adenine, 2-(aminoalklyamino)adenine or other heterosubstituted alkyladenines, 2-thiouracil, 2-thiothymine, 5 -bromouracil, 5-hydroxymethyluracil, 8-azaguanine, 7-deazaguanine, N6 (6-aminohexyl)adenine, and 2,6-diaminopurine. Kornberg, A., DNA Replication, W. H. Freeman & Co., San Francisco, pp75-77, 1980; Gebeyehu et al., Nucl. Acids Res. 1997, 15:4513. A "universal" base known in the art, e.g., inosine, can also be included. 5-Me-C substitutions have been shown to increase nucleic acid duplex stability by 0.6-1.2 °C. (Sanghvi, Y. S., in Crooke, S. T. and Lebleu, B., eds., Antisense Research and Applications, CRC Press, Boca Raton, 1993, pp. 276-278) and are aspects of base substitutions.

[0119] Modified nucleobases can comprise other synthetic and natural nucleobases, such as 5-methylcytosine (5-me-C), 5 -hydroxymethyl cytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl uracil and cytosine, 6-azo uracil, cytosine and thymine, 5-uracil (pseudo-uracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo particularly 5-bromo, 5 -trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylquanine and 7-methyladenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, and 3 -deazaguanine and 3 -deazaadenine.Complexes of a Genome-targeting Nucleic Acid and an Endonuclease

[0120] A gRNA interacts with an endonuclease (e.g., a RNA-guided nuclease such as Cas9), thereby forming a complex. The gRNA guides the endonuclease to a target polynucleotide.

[0121] The endonuclease and gRNA can each be administered separately to a cell or a subject. In some embodiments, the endonuclease can be pre-complexed with one or more guide RNAs, or one or more crRNA together with a tracrRNA. The pre-complexed material can then be administered to a cell or a subject. Such pre-complexed material is known as a ribonucleoprotein particle (RNP). The endonuclease in the RNP can be, for example, a Cas9 endonuclease or a Cpfl endonuclease. The endonuclease can be flanked at the N-terminus, the C-terminus, or both the N-terminus and C-terminus by one or more nuclear localization signals (NLSs). For example, a Cas9 endonuclease can be flanked by two NLSs, one NLS located at the N-terminus and the second NLS located at the C-terminus. The NLS can be any NLS known in the art, such as a SV40 NLS. The weight ratio of genome-targeting nucleic acid to endonuclease in the RNP can be 1:1. For example, the weight ratio of sgRNA to Cas9 endonuclease in the RNP can be 1:1.Base editing

[0122] In some embodiments, a gene can be edited using base editing. Base editing is a genome editing method that directly generates point mutations within a specific region of the genomic DNA without causing double-stranded breaks (DSB). DNA base editors (BEs) comprise fusions between a catalytically impaired Cas nuclease and a base-modification enzyme. Nucleobase editors typically include a polynucleotide programmable nucleotide binding domain and a nucleobase editing domain (e.g., adenosine deaminase, cytidine deaminase). A polynucleotide programmable nucleotide binding domain, when in conjunction with a bound guide polynucleotide (e.g., gRNA), can specifically bind to a target polynucleotide sequence and thereby localize the base editor to the target nucleic acid sequence desired to be edited. In some embodiments, base editing can be used to introduce a loss-of-function mutation (e.g., premature stop codons, destabilizing mutations, altering splicing, etc.). In other embodiments, base editing can be used to correct, a mutation (e.g., a disease-causing mutation). The systems, methods, compositions, and kits provided herein can, in some embodiments, be employed in concert with the systems, methods, compositions, and kits for base editing described in U. S. Patent Application No. 18 / 877,212, entitled, “BASE EDITING PROTEINS AND USES THEREOF,” filed December 19, 2024, the content of which is incorporated herein by reference in its entirety.

[0123] In some embodiments, base editors comprising a polynucleotide programmable nucleotide binding domain comprise all or a portion (e.g, a functional portion) of a CRISPR protein. In some embodiments, the polynucleotide programmable nucleotide binding domain comprises a nickase domain. Herein the term “nickase” shall be given its ordinary meaning, and shall also refer to a polynucleotide programmable nucleotide binding domain comprising a nuclease domain that is capable of cleaving only one strand of the two strands in a double-stranded nucleic acid molecule (e.g, DNA). For example, where a polynucleotide programmable nucleotide binding domain comprises a nickase domain derived from Cas9, the Cas9-derived nickase domain can include a D10A mutation and a histidine at position 840. In another example, a Cas9-derived nickase domain comprises an H840A mutation, while the amino acid residue at position 10 remains a D. In some embodiments, a Cas9 nuclease has an inactive (e.g., an inactivated) DNA cleavage domain (e.g., the Cas9 is a nickase, referred to as an “nCas9” protein). Suitable Cas9 nickases will be apparent to those of skill in the art based on this disclosure and knowledge in the field and are within the scope of this disclosure. In some embodiments, base editors comprise a polynucleotide programmable nucleotide binding domain which is catalytically dead (e.g., incapable of cleaving a target polynucleotide sequence). For example, in the case of a base editor comprising a Cas9 domain, the Cas9 can comprise both a D10A mutation and an H840A mutation. In further embodiments, a catalytically dead polynucleotide programmable nucleotide binding domain comprises a point mutation (e.g., D10A or H840A) as well as a deletionof all or a portion e.g., a functional portion) of a nuclease domain. Disclosed herein include base editors comprising: an engineered Cas9 protein disclosed herein or a nickase disclosed herein (e.g., a nickase comprising an amino acid sequence that is at least 80% identical to any one of the sequences of SEQ ID NOs: 48-94 and 748-770, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to any one of the sequences of SEQ ID NOs: 48-94 and 748-770); and a deaminase domain.

[0124] Disclosed herein are engineered Cas9 proteins. There are provided, in some embodiments, variant Cas9 polypeptides (e.g. nickases and nucleases) having alternative PAM recognition sequences (Table 1) derived from the parental Cas9 protein sequences shown in Table 2A.TABLE 1: CAS9 VARIANT PROTEIN SEQUENCESTABLE 2A: PARENTAL CAS9 PROTEIN SEQUENCESTABLE 2B: SsaCas9 VARIANTS

[0125] In some embodiments, a base editor comprises an adenosine deaminase domain. Such an adenosine deaminase domain of a base editor can facilitate the editing of an adenine (A) nucleobase to a guanine (G) nucleobase by deaminating the A to form inosine (I), which exhibits base pairing properties of G. In some embodiments, an A-to-Gbase editor further comprises an inhibitor of inosine base excision repair, for example, a uracil glycosylase inhibitor (UGI) domain or a catalytically inactive inosine specific nuclease. Without wishing to be bound by any particular theory, the UGI domain or catalytically inactive inosine specific nuclease can inhibit or prevent base excision repair of a deaminated adenosine residue (e.g., inosine), which can improve the activity or efficiency of the base editor. The adenosine deaminase can be derived from any suitable organism (e.g., E. coli, e.g., ecTadA deaminase). In some embodiments, the adenine deaminase is a naturally-occurring adenosine deaminase that includes one or more mutations. Details of A to G nucleobase editing proteins are described W02018 / 027078 and Gaudelli, N.M., et al., “Programmable base editing of A»T to G»C in genomic DNA without DNA cleavage” Nature, 551, 464-471 (2017), the entire contents of which are hereby incorporated by reference.

[0126] In some embodiments, a base editor comprises a fusion protein or complex comprising cytidine deaminase capable of deaminating a target cytidine (C) base of a polynucleotide to produce uridine (U), which has the base pairing properties of thymine. In some embodiments, for example where the polynucleotide is double-stranded (e.g., DNA), the uridine base can then be substituted with a thymidine base (e.g., by cellular repair machinery) to give rise to a C:G to a T:A transition. In other embodiments, deamination of a C to U in a nucleic acid by a base editor cannot be accompanied by substitution of the U to a T. The deamination of a target C in a polynucleotide to give rise to a U is a non-limiting example of a type of base editing that can be executed by a base editor described herein. In another example, a base editor comprising a cytidine deaminase domain can mediate conversion of a cytosine (C) base to a guanine (G) base. For example, a U of a polynucleotide produced by deamination of a cytidine by a cytidine deaminase domain of a base editor can be excised from the polynucleotide by a base excision repair mechanism (e.g., by a uracil DNA glycosylase (UDG) domain), producing an abasic site. The nucleobase opposite the abasic site can then be substituted (e.g., by base repair machinery) with another base, such as a C, by for example a translesion polymerase. Although it is typical for a nucleobase opposite an abasic site to be replaced with a C, other substitutions (e.g., A, G or T) can also occur.

[0127] Accordingly, in some embodiments a base editor described herein comprises a deamination domain (e.g., cytidine deaminase domain) capable of deaminating a target C to a U in a polynucleotide. Further, as described below, the base editor can comprise additional domainswhich facilitate conversion of the U resulting from deamination to, in some embodiments, a T or a G. For example, a base editor comprising a cytidine deaminase domain can further comprise a uracil glycosylase inhibitor (UGI) domain to mediate substitution of a U by a T, completing a C-to-T base editing event. In another example, the base editor can comprise a uracil stabilizing protein as described herein. In another example, a base editor can incorporate a translesion polymerase to improve the efficiency of C-to-G base editing, since a translesion polymerase can facilitate incorporation of a C opposite an abasic site (i.e., resulting in incorporation of a G at the abasic site, completing the C-to-G base editing event). A base editor comprising a cytidine deaminase as a domain can deaminate a target C in any polynucleotide, including DNA, RNA and DNA-RNA hybrids.

[0128] In some embodiments, a cytidine deaminase of a base editor comprises all or a portion (e.g., a functional portion) of an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. APOBEC is a family of evolutionarily conserved cytidine deaminases. Members of this family are C-to-U editing enzymes. The N-terminal domain of APOBEC like proteins is the catalytic domain, while the C-terminal domain is a pseudocatalytic domain. More specifically, the catalytic domain is a zinc dependent cytidine deaminase domain and is important for cytidine deamination. APOBEC family members include APOBEC 1, AP0BEC2, AP0BEC3A, AP0BEC3B, APOBEC3C, AP0BEC3D (“AP0BEC3E” now refers to this), APOBEC3F, AP0BEC3G, AP0BEC3H, AP0BEC4, and Activation-induced (cytidine) deaminase. In some embodiments, the deaminases are activation-induced deaminases (AID). In some embodiments, an APOBEC deaminase incorporated into a base editor can comprise one or more mutations selected from the group consisting of H121R, H122R, R126A, R126E, R118A, W90A, W90Y, and R132E of rAPOBECl; D316R, D317R, R320A, R320E, R313A, W285A, W285Y, and R326E of hAPOBEC3G; and any alternative mutation at the corresponding position, or one or more corresponding mutations in another APOBEC deaminase. A number of modified cytidine deaminases are commercially available, including, but not limited to, SaBE3, SaKKH-BE3, VQR-BE3, EQR-BE3, VRER-BE3, YE1-BE3, EE-BE3, YE2-BE3, and YEE-BE3, which are available from Addgene (plasmids 85169, 85170, 85171, 85172, 85173, 85174, 85175, 85176, 85177). In some embodiments, a deaminase incorporated into a base editor comprises all or a portion (e.g., a functional portion) of an APOBEC 1 deaminase.

[0129] Details of C to T nucleobase editing proteins are described in W02017 / 070632 and Komor, A.C., et al., “Programmable editing of a target base in genomic DNA without doublestranded DNA cleavage” Nature 533, 420-424 (2016), the entire contents of which are hereby incorporated by reference.

[0130] A polynucleotide programmable nucleotide binding domain, when inconjunction with a bound guide polynucleotide e.g., gRNA), can specifically bind to a target polynucleotide sequence (i.e., via complementary base pairing between bases of the bound guide nucleic acid and bases of the target polynucleotide sequence) and thereby localize the base editor to the target nucleic acid sequence desired to be edited (e.g., a double-stranded DNA target). In one embodiment, the guide polynucleotide is a gRNA. In some embodiments, the guide polynucleotide is at least one single guide RNA (“sgRNA” or “gRNA”). In some embodiments, the methods described herein can utilize an engineered Cas protein. A guide RNA (gRNA) is a short synthetic RNA composed of a scaffold sequence necessary for Cas-binding and a user-defined ~20 nucleotide spacer that defines the genomic target to be modified. Thus, the specificity of the Cas protein for the genomic target of the Cas protein is partially determined by how specific the gRNA targeting sequence is for the genomic target compared to the rest of the genome. In some embodiments, the spacer is about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25, or more nucleotides in length. The spacer of a gRNA can be or can be about 19, 20, or 21 nucleotides in length.

[0131] In some embodiments, the deamination of the target nucleobase results in the correction of a genetic defect, e.g., in the correction of a point mutation that leads to a loss of function in a gene product. The genetic defect can be associated with a disease or disorder. In some embodiments, the methods provided herein are used to introduce a deactivating point mutation into a gene or allele that encodes a gene product that is associated with a disease or disorder (e.g., an oncogene). In some embodiments, the sequence associated with a disease or disorder is a gene encoding a protein. The deamination results in a disrupted or deactivated gene such that expression of a functional protein from the gene is reduced or inhibited. For example the deamination can generate a premature stop codon in a coding sequence, which results in the expression of a truncated gene product, e.g., a truncated protein lacking the function of the full-length protein.

[0132] In some embodiments, the methods described herein can restore the function of dysfunctional gene via base editing. Generally, any single nucleotide polymorphisms (SNPs) involving a T to C point mutation while having a nearby Cas recognizable PAM sequence can be corrected using the method herein described. Deamination of the mutant C back to U corrects the mutation. Examples of genes comprising a T to C point mutation associated with a disease or disorder include, but are not limited to, CBS, DPYS, AGA, ALDOB, RFX6, TMEM67, ERCC6, GJC2, PC, ADSL, TPP1, BEST1, ACAT1, SMPD1, LDLR, TLL1, SLC26A2, GBA, EIF2B3, GAN, ANKH, FBLN5, ROBO2, KLF6, DDX11, FOXF1, NOTCH3, MMP13, ITGB2, ABCA1, MT-TL2, MT-TI, MT-ND1, PRPS1, F8, F9, TAZ, BTK, FOXP3, CASK, MECP2, MVK, TEAD1, S0S1, FGFR2, GP9, GPI, PTH, KIT, MAPT, LMNA, KRT5, HBG2, GH1, GRN, GPD2,NR3C1, SMC1A, SGSH, MRPS22, CLN6, MFN2, RHBDF2, UVSSA, P0C1A, HINT1, COL4A5, SMN1, HCFC1, MVK, SLC29A3, DLD, BRAF, RYR1, MYBPC3, MYH7, USH2A, BRCA2, BRCA1, KARS, KAT6A, KCNQ1, SCN5A, SCN1A, SCN10A, MSH2, GBA, DYSF, ATP6V50A2, FOLR1, STX11, IFT140, OTC, PRPH2, ITGB2, SAMHD1, ACTB, RRM2B, PGM3, HNRNPA1, DNMT3A, CDKN1C, MTCO3, GATA1, DNAL4, SCN11 A, TRNT1, DCX, LMNA, MYBPC3, MTHFR, PTCHI, AGXT, SGCA, ACTN2, SCN1A, LAMA2, EDA, COL6A1, KCNH2, SCN2A, CLCN1, NDUFS2, and TP53. Detailed information about the above gene symbols, including genotypes, sequences, and associated genetic diseases, can be found in public domains such as NCBI ClinVar database as will be understood to a person skilled in the art.Deaminase Domain

[0133] The base editor (e.g., base editing fusion protein) described herein can comprise a nucleic acid editing domain such as a deaminase or deaminase domain. The term “deaminase” refers to an enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase or deaminase domain belongs to the deaminase superfamily. The deaminase superfamily encompasses zinc-dependent enzymes catalyzing the deamination of bases in free nucleotides and nucleic acids. The deaminase superfamily displays a conserved P-sheet with five P-strands arranged in 2- 1-3 -4-5 order interleaved with three a-helices forming an a / p-fold. The active site comprises two zinc-chelating motifs, respectively represented by a motif of HxE / CxE / DxE at the end of helix 2 and CxnC (where x is any amino acid and n is >2) located in loop 5 and the beginning of helix 3. In some embodiments, the zinc ion is coordinated by the side chains of residues His and Cys, such as His257, Cys291 and Cys288 in APOBEC3G.

[0134] In some embodiments, the deaminase is a cytidine deaminase, catalyzing the hydrolytic deamination of cytidine or deoxy cytidine to uracil or deoxyuracil, respectively. In some embodiments, the deaminase is an apolipoprotein B mRNA-editing complex (APOBEC) family deaminase. In some embodiments, an APOBEC family deaminase contains a conserved a / p-fold core domain at the N-terminal and a C-terminal domain. The APOBEC deaminase core domain comprises three active loops (e.g., loops 1, 3 and 7 in APOBEC1 and loops 1, 5 and 7 in APOBEC3) containing amino acid residues known to form interactions with a nucleic acid upon its binding to the deaminase. In some embodiments, the C-terminal domain of an APOBEC family deaminase comprises a P-hairpin and three small helical domains. In some embodiments, the deaminase herein described is a dimer formed by two APOBEC deaminase monomers mediated through interactions between the C-terminal domains. In some embodiments, amino acid residues in the active loops of the deaminase core domain and / or residues in the C-terminal domain can be mutated to generate deaminase variants with desired properties such as low spurious off-targetediting activity and improved on-target editing precision (e.g., by narrowing editing window). The APOBEC family of cytosine deaminase enzymes encompasses eleven proteins that serve to initiate mutagenesis in a controlled and beneficial manner as will be understood by a person skilled in the art.

[0135] In some embodiments, the deaminase is an APOBEC 1 family deaminase. In some embodiments, the deaminase is an APOBEC3 family deaminase. In some embodiments, the deaminase is an activation-induced cytidine deaminase (AID), which is responsible for the maturation of antibodies by converting cytosines in ssDNA to uracils in a transcription-dependent, strand-biased fashion. In some embodiments, the deaminase is an ACF1 / ASE deaminase. Additional suitable nucleic acid-editing enzymes or domains will be apparent to the skilled artisan based on this disclosure.

[0136] The deaminase or deaminase domain can comprise, or can be, a naturally-occurring deaminase from an organism, mammals, fungus, reptiles, amphibians and birds. In some embodiments, the deaminase or deaminase domain is a variant of a naturally-occurring deaminase from an organism, that does not occur in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75% at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally-occurring deaminase from an organism. In some embodiments, the deaminase or deaminase domain is a variant of a naturally-occurring deaminase comprising one or more mutations or a deletion or insertion of one or more residues within a sequence. In some embodiments, the one or more mutations / deletions / insertions result in altered catalytic deaminase activity. For example, the one or more mutations / deletions / insertions result in reduced catalytic deaminase activity such that the deaminase or deaminase domain is less likely to catalyze the deamination of a residue adjacent to a target residue, thereby narrowing the deamination window. Various methods for making the amino acid substitutions (mutations) provided herein are well known in the art, and are provided by, for example, Green and Sambrook, Molecular Cloning: A Laboratory Manual (4thed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)).

[0137] The deaminase or deaminase domain can be from a mammal, for example, rat, bat, armadillo, or orangutan; or be a variant thereof. In some embodiments, the deaminase or deaminase domain is a rat deaminase. In some embodiments, the deaminase or deaminase domain is a bat deaminase. In some embodiments, the deaminase or deaminase domain is an armadillo deaminase. In some embodiments, the deaminase or deaminase domain is an orangutan deaminase. In some embodiments, the deaminase or deaminase domain is from Dasypus novemcinctus (nine-banded armadillo), Meriones unguiculatus (Mongolian gerbil) or Myotislucifugus (little brown bat). In some embodiments, the deaminase or deaminase domain is from Dasypus novemcinctus .Uracil glycosylase inhibitor (UGI)

[0138] In some embodiments, the base editor comprises one or more uracil glycosylase inhibitor (UGI) domain. The term “uracil glycosylase inhibitor” or “UGI,” as used herein, refers to a protein that is capable of inhibiting a uracil-DNA glycosylase base-excision repair enzyme. Without wishing to be bound by any particular theory, cellular DNA-repair response to the presence of U:G heteroduplex DNA may be responsible for the decrease in nucleobase editing efficiency in cells. For example, uracil DNA glycosylase (UDG) catalyzes removal of U from DNA in cells, which may initiate base excision repair, with reversion of the U:G pair to a C:G pair as the most common outcome. In some embodiments, fusion proteins comprising a UGI domain can inhibit human UDG activity, enabling increased deamination efficiency. A UGI domain can comprise a wild-type UGI or a UGI variant including a fragment of UGI or a protein homologous to a UGI or UGI fragment.

[0139] In some embodiments, the present disclosure includes a fusion protein comprising a deaminase and Cas9 nickase domain further fused to at least one UGI domain. The UGI domain can be fused to the Cas9 domain and / or the deaminase domain either directly or via an optional linker. In some embodiments, the base editor comprises two UGI domains. The one or more UGI domains, the deaminase domain, and the Cas9 domain can be connected to one another in any suitable configuration. For example, in some embodiments, the base editor comprises the structure of [deaminase domain]-[Cas9 domain]-[UGI]-[UGI] connected either directly or via an optional linker.

[0140] A UGI domain can be any protein or fragment thereof capable of inhibiting (e.g., sterically blocking) a uracil-DNA glycosylase base-excision repair enzyme. Additionally, any proteins that block or inhibit base-excision repair are also within the scope of this disclosure. In some embodiments, a UGI is a protein that binds uracil. In some embodiments, a UGI is a protein that binds uracil in DNA. In some embodiments, a UGI is a catalytically inactive uracil DNA-glycosylase protein. In some embodiments, a UGI is a catalytically inactive uracil DNA-glycosylase protein that does not excise uracil from the DNA. Suitable UGI protein and nucleotide sequences are known to those in the art, and include, for example, those published in Wang et al., Uracil-DNA glycosylase inhibitor gene of bacteriophage PBS2 encodes a binding protein specific for uracil-DNA glycosylase. J. Biol. Chem. 264:1163-1171(1989); Lundquist et al., Site-directed mutagenesis and characterization of uracil-DNA glycosylase inhibitor protein. Role of specific carboxylic amino acids in complex formation with Escherichia coli uracil-DNA glycosylase. J. Biol. Chem. 272:21408-21419(1997); Ravishankar et al., X-ray analysis of a complexof Escherichia coli uracil DNA glycosylase (EcUDG) with a proteinaceous inhibitor. The structure elucidation of a prokaryotic UDG. Nucleic Acids Res. 26:4880-4887(1998); and Putnam et al., Protein mimicry of DNA from crystal structures of the uracil-DNA glycosylase inhibitor protein and its complex with Escherichia coli uracil-DNA glycosylase. J. Mol. Biol. 287:331-346(1999), the entire contents of which are incorporated herein by reference.

[0141] In some embodiments, the base editors comprising a UGI further comprise a nuclear targeting sequence, for example a nuclear localization sequence. In some embodiments, fusion proteins provided herein further comprise a nuclear localization sequence (NLS). The NLS can be fused to the N-terminus or C-terminus of the base editor. In some embodiments, the NLS can be fused to the N-terminus or C-terminus of the UGI, the N-terminus or C-terminus of the Cas9 domain, or the N-terminus or C-terminus of the deaminase domain. The fusion can be either direct or via a linker.

[0142] The Cas9 domain, the deaminase domain, and optionally the UGI and NLS can be fused to one another via a linker. The term “linker,” as used herein, refers to a chemical group or a molecule linking two molecules or moieties. In some embodiments, a linker joins a Cas9 domain (e.g., a Cas9 nickase) and a deaminase domain. In some embodiments, a linker joins a Cas9 domain with a UGI domain. In some embodiments, a linker join the deaminase domain with a UGI domain. In some embodiments, a linker joins one UGI domain with another UGI domain. A linker can be an amino acid, a peptide or protein, an organic molecule, a polymer or chemical moiety. The sequence, length and flexibility of the linker can vary in different embodiments. In some embodiments, the linker can have about 3-100 amino acids in length. Suitable linker motifs and configurations are described in published literatures (see, e.g., Chen et al., Fusion protein linkers: property, design and functionality. Adv Drug Deliv Rev. 2013; 65(10): 1357-69 and Guilinger J P, Thompson D B, Liu D R. Fusion of catalytically inactive Cas9 to FokI nuclease improves the specificity of genome modification. Nat. BiotechnoL 2014; 32(6): 577-82, the contents of which are incorporated herein by reference) and will be apparent to those of skill in the art. In some embodiments, the linker comprises a (GGS)n, (G)n, ((GGGGS)n motif, wherein n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 19, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. In some embodiments, the optional linker comprises a (GGS)n motif, wherein n is 1, 3, or 7.

[0143] The deaminase domain, the Cas9 domain and optionally one or more UGI and NSL can be connected to one another in a variety of suitable configurations. In some embodiments, the deaminase domain is fused to the N-terminus of the Cas9 domain. In some embodiments, the deaminase domain is fused to the C-terminus of the Cas9 domain. In some embodiments, a base editing fusion protein described herein comprises the structure from the N-terminus to the C-terminus: [deaminase domain]-[Cas9 domain], in which the deaminase domainand the Cas9 domain are connected directly or via a linker. In some embodiments, a base editing fusion protein described herein comprises a structure from the N-terminus to the C-terminus selected from the following: [deaminase domain]-[Cas9 domain]-[UGI], [UGI]-[deaminase domain]-[Cas9 domain] or [deaminase domain]-[UGI]-[Cas9 domain]. In some embodiments, more than one UGI domain can be fused to the deaminase and / or Cas9 domain. For example, a base editing fusion protein described herein can have a structure of [deaminase domain]-[Cas9 domain]-[UGI]-[UGI],

[0144] The base editing fusion protein used herein can be a portion or variant of BE1 base editor, BE2 base editor, BE3 base editor, HF-BE3, BE4, BE4-GAM, YE1-BE3, EE-BE3, YE2-BE3, YEE-BE3, VQR-BE3, VRER-BE3, VRER-BE3, Sa-BE3, Sa-BE4, SaBE4-Gam, SaKKH-BE3, Casl2a-BE, xBE3 described in Rees and Liu, Nat Rev Genet. 2018 December; 19 (12): 770-788, the contents of which are incorporated herein by reference.

[0145] The base editing fusion protein described herein can efficiently generate an intended mutation, such as a point mutation from C to T, in a nucleic acid without generating an unintended mutation such as an unintended point mutation. In some embodiments, the base editing fusion protein described herein can modify a specific nucleotide base without generating undesired byproducts including indels, translocations and other DNA rearrangements. The term “indel”, as used herein, refers to the insertion or deletion of one or more nucleotide base within a nucleic acid. In some embodiments, the base editor described herein can generate fewer than 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.2%, 0.1%, 0.01%, 0.001%, or less indels. In some embodiments, the base editor described herein can generate a ratio of intended point mutations to indels that is at least 5:1, 10:1, 20:1, 30:1, 40:1, 50:1, 100:1, 200:1, 400:1, 600:1, 800:1, 1000:1 or higher. In some embodiments, the base editor described herein can generate fewer than 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.2%, 0.1%, 0.01%, 0.001%, or less translocations or other DNA rearrangement. In some embodiments, the base editor described herein does not generate any indel, DNA translocation or other DNA rearrangement.Reverse Transcriptase (RT) editing

[0146] RT editing methods and compositions disclosed herein are directed to correction of a disease mutation using reverse transcriptase (RT) editing. Compositions disclosed herein comprise and RT editor or an mRNA encoding a reverse transcriptase (RT) editor, a long guide RNA encoding the edit designated as ‘tagRNA’, and optionally, a second guide RNA designated as ‘egRNA’ (opposite strand gRNA). In some embodiments, the composition induces programmable editing of a target DNA using an RT editor complexed with a tagRNA to incorporate an intended nucleotide edit (also referred to herein as a nucleotide change) into the target DNA. In some embodiments, the composition comprises a second guide RNA (egRNA).The systems, methods, compositions, and kits provided herein can, in some embodiments, be employed in concert with the systems, methods, compositions, and kits for RT editing described in PCT Application No. PCT / IB2025 / 052079, entitled, “RT EDITING COMPOSITIONS AND METHODS,” filed February 26, 2025, the content of which is incorporated herein by reference in its entirety. The systems, methods, compositions, and kits provided herein can, in some embodiments, be employed in concert with the reverse transcriptases described in the U.S. Provisional Patent Application No. 63 / 895,913, entitled, “SYNTHETIC NUCLEOTIDE TEMPLATE POLYMERASE (SYNTASE) EDITING,” filed October 8, 2025, the content of which is incorporated herein by reference in its entirety. The systems, methods, compositions, and kits provided herein can, in some embodiments, be employed in concert with the systems, methods, compositions, and kits for prime editing describedin W02023015309, WO2022150790, W02022067130, WO2020191233, WO2020191234, WO2020191239, W02020191241, WO2020191242, WO2020191243, WO2020191245, WO2020191246, WO2020191248, WO2020191249, W02020191153, and W02020191171, the contents of which are incorporated herein by reference in their entireties. The systems, methods, compositions, and kits provided herein can, in some embodiments, be employed in concert with the systems, methods, compositions, and kits for click editing described in WO2024211883A1, the content of which is incorporated herein by reference in its entirety. The systems, methods, compositions, and kits provided herein can, in some embodiments, be employed in concert with the systems, methods, compositions, and kits for DNA polymerase editing (DPE) described in WO2023235501A1, the content of which is incorporated herein by reference in its entirety. The systems, methods, compositions, and kits provided herein can, in some embodiments, be employed in concert with the systems, methods, compositions, and kits for writing / rewriting described in WO2021188840A1, the content of which is incorporated herein by reference in its entirety.

[0147] A target gene of the RT editing may comprise a double stranded DNA molecule having two complementary strands: a first strand that may be referred to as a “target strand” or a “non-edit strand”, and a second strand that may be referred to as a “non-target strand,” or an “edit strand” or an “opposite strand”. The RT editors provided herein can be employed to introduce more than one nucleotide change. In some embodiments, two (or more) RT editors provided herein can be employed to edit two (or more) different genes or two (or more) different sites of a gene. RT editors provided herein can correct more than one mutation. Two disclosed RT editors can be used to target two different genes or two different segments of a gene.RT editor

[0148] The term “RT editor” refers to the polypeptide or polypeptide components involved in RT editing, or any polynucleotide(s) encoding the polypeptide or polypeptidecomponents. In various embodiments, a RT editor includes a polypeptide domain having DNA endonuclease activity and a polypeptide domain having DNA polymerase activity. In some embodiments, the RT editor further comprises a polypeptide domain having nuclease activity. In some embodiments, the polypeptide domain having DNA binding activity comprises a nuclease domain or nuclease activity. In some embodiments, the polypeptide domain having nuclease activity comprises a nickase, or a fully active nuclease. As used herein, the term “nickase” refers to a nuclease capable of cleaving only one strand of a double-stranded DNA target. In some embodiments, the RT editor comprises a polypeptide domain that is an inactive nuclease, in some embodiments, the polypeptide domain having comprises a nucleic acid guided DNA endonuclease domain, for example, a CRISPR-Cas protein, for example, a Cas9 nickase, a Cpfl nickase, or another CRISPR-Cas nuclease. In some embodiments, the polypeptide domain having DNA polymerase activity comprises a template-dependent DNA polymerase, for example, a DNA-dependent DNA polymerase or an RNA-dependent DNA polymerase. In some embodiments, the DNA polymerase is a reverse transcriptase, in some embodiments, the RT editor comprises additional polypeptides involved in RT editing, for example, a polypeptide domain having a 5’ endonuclease activity, e.g., a 5’ endogenous DNA flap endonucleases (e.g., FEN1), for helping to drive the RT editing process towards the edited product formation. In some embodiments, the RT editor further comprises an RNA-protein recruitment polypeptide, for example, a MS2 coat protein.

[0149] A RT editor may be engineered. In some embodiments, the polynucleotide or polypeptide components of a RT editor do not naturally occur in the same organism or cellular environment. In some embodiments, the polynucleotide or polypeptide components of a RT editor may be of different origins or from different organisms. In some embodiments, a RT editor comprises a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain that are derived from different species. In some embodiments, a RT editor comprises a Cas polypeptide (DNA endonuclease domain) and a reverse transcriptase polypeptide (DNA polymerase) that are derived from different species. The DNA endonuclease domain can comprise an engineered Cas9 protein disclosed herein or a nickase disclosed herein (e.g., a nickase comprising an amino acid sequence that is at least 80% identical to any one of the sequences of SEQ ID NOs: 48-94 and 748-770, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to any one of the sequences of SEQ ID NOs: 48-94 and 748-770).

[0150] In some embodiments, polypeptide domains of a RT editor may be fused or linked by a peptide linker to form a fusion protein. In other embodiments, a RT editor comprises one or more polypeptide domains provided in trans as separate proteins, which are capable of being associated to each other through non-peptide linkages or through aptamers or recruitmentsequences. For example, a RT editor may comprise a DNA binding domain and a reverse transcriptase domain associated with each other by an RNA-protein recruitment aptamer, e.g., an MS2 aptamer, which may be linked to a tagRNA. RT editor polypeptide components may be encoded by one or more polynucleotides in whole or in part, in some embodiments, a single polynucleotide, construct, or vector encodes the RT editor fusion protein. In some embodiments, multiple polynucleotides, constructs, or vectors each encode a polypeptide domain or portion of a domain of a RT editor, or a portion of a RT editor fusion protein. For example, a RT editor fusion protein may comprise an N-terminal portion fused to an intein-N and a C-terminal portion fused to an intein-C, each of which is individually encoded by an AAV vector

[0151] In some embodiments, the RT Editor is transcribed from an mRNA. In some embodiments, the mRNA comprises a 5'-cap structure. In some embodiments, the mRNA comprises a nuclear localization sequence (NLS). In some embodiments, the mRNA sequence comprises a dead Cas9 sequence. In some embodiments, the Cas9 is a Cas nickase sequence. In some embodiments, the RT Editor mRNA sequence comprises a 5’-UTR. In some embodiments, the RT Editor mRNA sequence comprises a 3’-UTR. In some embodiments, the RT Editor comprises a sequence of a viral element. In some embodiments, the viral element is a woodchuck hepatitis virus post-transcriptional regulatory element (WPRE) sequence. In some embodiments, the RT editor mRNA sequence has a structure comprising or consisting of a sequence comprising from 5’ to 3’ as [5’UTR]-[NLS]-[nCas9]-[linker]-[RT]-[NLS]-[3’UTR_and / or_viral_element]-[polyA sequence], In some embodiments, any nucleoside of the RT editor mRNA sequence may be chemically modified.

[0152] Compositions disclosed herein comprise a long guide RNA designated herein as “template armed guide RNA” or “tagRNA”. The tagRNA comprises a spacer sequence, a scaffold sequence, an editing template and a flap binding sequence. In some embodiments, the tagRNA comprises in 5’ to 3’ order: spacer sequence, scaffold sequence, editing template, and a flap binding sequence.RT Editor Nucleotide Polymerase Domain

[0153] In some embodiments, a RT editor comprises a nucleotide polymerase domain, e.g., a DNA polymerase domain. The DNA polymerase domain may be a wild-type DNA polymerase domain, a full-length DNA polymerase protein domain, or may be a functional mutant, a functional variant, or a functional fragment thereof. In some embodiments, the polymerase domain is a template dependent polymerase domain. For example, the DNA polymerase may rely on a template polynucleotide strand, e.g., the editing template sequence, for new strand DNA synthesis. In some embodiments, the RT editor comprises a DNA-dependent DNA polymerase. For example, a RT editor having a DNA-dependent DNA polymerase cansynthesize a new single stranded DNA using a tagRNA editing template that comprises a DNA sequence as a template. In such cases, the tagRNA is a chimeric or hybrid tagRNA, and comprising an extension arm comprising a DNA strand. The chimeric or hybrid tagRNA may comprise an RNA portion (including the spacer and the gRNA core) and a DNA portion (the extension arm comprising the editing template that includes a strand of DNA).

[0154] In some embodiments, a Cas protein, e.g., Cas9, can be a wild type or a modified form of a Cas protein. In some embodiments, a Cas protein, e.g., Cas9, can be a nuclease active variant, nuclease inactive variant, a nickase, or a functional variant or functional fragment of a wild-type Cas protein. In some embodiments, a Cas protein, e.g., Cas9, can be a wild type or a modified form of a Cas protein. A Cas protein, e.g., Cas9, can be a nuclease active variant, nuclease inactive variant, a nickase, or a functional variant or functional fragment of a wild-type Cas protein. A Cas protein, e.g., Cas9, can comprise an amino acid change such as a deletion, insertion, substitution, fusion, chimera, or any combination thereof relative to a corresponding wild-type version of the Cas protein. In some embodiments, a Cas protein can be a polypeptide with at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or sequence similarity to a wild type exemplary Cas protein.

[0155] A Cas protein, e.g., Cas9, may comprise one or more domains. Non-limiting examples of Cas domains include, guide nucleic acid recognition and / or binding domain, nuclease domains (e.g., DNase or RNase domains, RuvC, HNH), DNA binding domain, a DNA endonuclease domain, RNA binding domain, helicase domains, protein-protein interaction domains, and dimerization domains. In various embodiments, a Cas protein comprises a guide nucleic acid recognition and / or binding domain that can interact with a guide nucleic acid, and one or more nuclease domains that comprise catalytic activity for nucleic acid cleavage.

[0156] In some embodiments, a Cas protein, e.g., Cas9, comprises one or more nuclease domains. A Cas protein can comprise an amino acid sequence having at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a nuclease domain (e.g., RuvC domain, HNH domain) of a wild-type Cas protein. In some embodiments, a Cas protein comprises a single nuclease domain. For example, a Cpfl may comprise a RuvC domain but lacks HNH domain. In some embodiments, a Cas protein comprises two nuclease domains, e.g., a Cas9 protein can comprise an HNH nuclease domain and a RuvC nuclease domain.

[0157] In some embodiments, aRT editor comprises a Cas protein, e.g., Cas9, wherein all nuclease domains of the Cas protein are active. In some embodiments, a RT editor comprises a Cas protein having one or more inactive nuclease domains. One or a plurality of the nuclease domains (e.g., RuvC, HNH) of a Cas protein can be deleted or mutated so that they are no longerfunctional or comprise reduced nuclease activity. In some embodiments, a Cas protein, e.g., Cas9, comprising mutations in a nuclease domain has reduced (e.g., nickase) or abolished nuclease activity while maintaining its ability to target a nucleic acid locus at a search target sequence when complexed with a guide nucleic acid, e.g., a tagRNA.

[0158] In some embodiments, a RT editor comprises a Cas nickase that can bind to the target gene in a sequence-specific manner and generate a single-strand break at a protospacer within double-stranded DNA in the target gene, but not a double-strand break. For example, the Cas nickase can cleave the edit strand or the non-edit strand of the target gene, but may not cleave both. In some embodiments, a RT editor comprises a Cas nickase comprising two nuclease domains (e.g., Cas9), with one of the two nuclease domains modified to lack catalytic activity or deleted. In some embodiments, the Cas nickase of a RT editor comprises a nuclease inactive RuvC domain and a nuclease active HNH domain. In some embodiments, the Cas nickase of a RT editor comprises a nuclease inactive HNH domain and a nuclease active RuvC domain. In some embodiments, a RT editor comprises a Cas9 nickase having an amino acid substitution in the RuvC domain e.g., an amino acid substitution that reduces or abolishes nuclease activity of the RuvC domain. In some embodiments, the Cas9 nickase comprises a D10X amino acid substitution compared to a wild type S. pyogenes Cas9, wherein X is any amino acid other than D. In some embodiments, a RT editor comprises a Cas9 nickase having an amino acid substitution in the HNH domain, e.g., an amino acid substitution that reduces or abolishes nuclease activity of the HNH domain. In some embodiments, the Cas9 nickase comprises a H840X amino acid substitution compared to a wild type S. pyogenes Cas9, wherein X is any ammo acid other than H.

[0159] In some embodiments, a RT editor comprises a Cas protein that can bind to the target gene in a sequence-specific manner but lacks or has abolished nuclease activity and may not cleave either strand of a double stranded DNA in a target gene. Abolished activity or lacking activity can refer to an enzymatic activity less than 1%, less than 2%, less than 3%, less than 4%, less than 5%, less than 6%, less than 7%, less than 8%, less than 9%, or less than 10% activity compared to a wild-type exemplary activity (e.g., wild-type Cas9 nuclease activity). In some embodiments, a Cas protein of a RT editor completely lacks nuclease activity. A nuclease, e.g., Cas9, that lacks nuclease activity may be referred to as nuclease inactive or “nuclease dead” (abbreviated by “d”). A nuclease dead Cas protein (e.g., dCas, dCas9) can bind to a target polynucleotide but may not cleave the target polynucleotide. In some aspects, a dead Cas protein is a dead Cas9 protein. In some embodiments, a RT editor comprises a nuclease dead Cas protein wherein all of the nuclease domains (e.g., both RuvC and HNH nuclease domains in a Cas9 protein; RuvC nuclease domain in a Cpfl protein) are mutated to lack catalytic activity, or are deleted.

[0160] A Cas protein can be modified. A Cas protein, e.g., Cas9, can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, and / or enzymatic activity. Cas proteins can also be modified to change any other activity or property of the protein, such as stability. For example, one or more nuclease domains of the Cas protein can be modified, deleted, or inactivated, or a Cas protein can be truncated to remove domains that are not essential for the function of the protein or to optimize (e.g., enhance or reduce) the activity of the Cas protein.

[0161] A Cas protein can be a fusion protein. For example, a Cas protein can be fused to a cleavage domain, an epigenetic modification domain, a transcriptional regulation domain, or a polymerase domain. A Cas protein can also be fused to a heterologous polypeptide providing increased or decreased stability. The fused domain or heterologous polypeptide can be located at the N-terminus, the C-terminus, or internally within the Cas protein.

[0162] In some embodiments, the Cas protein of a RT editor is a Class 2 Cas protein. In some embodiments, the Cas protein is a type II Cas protein. In some embodiments, the Cas protein is a Cas9 protein, a modified version of a Cas9 protein, a Cas9 protein homolog, mutant, variant, or a functional fragment thereof. As used herein, a Cas9, Cas9 protein, Cas9 polypeptide or a Cas9 nuclease refers to an RNA guided nuclease comprising one or more Cas9 nuclease domains and a Cas9 gRNA binding domain having the ability to bind a guide polynucleotide, e.g., a tagRNA. A Cas9 protein may refer to a wild-type Cas9 protein from any organism or a homolog, ortholog, or paralog from any organisms; any functional mutants or functional variants thereof; or any functional fragments or domains thereof. In some embodiments, a RT editor comprises a full-length Cas9 protein. In some embodiments, the Cas9 protein can generally comprises at least about 50%, 60%, 70%, 80%, 90%, 100% sequence identity to a wild-type reference Cas9 protein (e.g., Cas9 from S. pyogenes). In some embodiments, the Cas9 comprises an amino acid change such as a deletion, insertion, substitution, fusion, chimera, or any combination thereof as compared to a wild-type reference Cas9 protein.Reverse transcriptases

[0163] In some embodiments, a RT editor comprises an RNA-dependent DNA polymerase domain, for example, a reverse transcriptase (RT). An RT or an RT domain may be a wild-type RT domain, a full-length RT domain, or may be a functional mutant, a functional variant, or a functional fragment thereof. An RT or an RT domain of a RT editor may comprise a wild-type RT, or may be engineered or evolved to contain specific amino acid substitutions, truncations, or variants. An engineered RT may comprise sequences or amino acid changes different from a naturally occurring RT. In some embodiments, the engineered RT may have improved reverse transcription activity over a naturally occurring RT or RT domain. In someembodiments, the engineered RT may have improved features over a naturally occurring RT, for example, improved thermostability, reverse transcription efficiency, or target fidelity. In some embodiments, a RT editor comprising the engineered RT has improved RT editing efficiency over a RT editor having a reference naturally occurring RT.

[0164] In some embodiments, a RT editor comprises a virus RT, for example, a retrovirus RT. Nonlimiting examples of virus RT include Moloney murine leukemia virus (M-MLV or MLVRT or M-MLV RT); human T-cell leukemia virus type 1 (HTLV-1) RT; bovine leukemia virus (BLV) RT; Rous Sarcoma Virus (RSV) RT; human immunodeficiency virus (HIV) RT, M-MFV RT, Avian Sarcoma-Leukosis Virus (ASLV) RT, Rous Sarcoma Virus (RSV) RT, Avian Myeloblastosis Virus (AMV) RT, Avian Erythroblastosis Virus (AEV) Helper Virus MCAV RT, Avian Myelocytomatosis Virus MC29 Helper Virus MCAV RT, Avian Reticuloendotheliosis Virus (REV-T) Helper Virus REV-A RT, Avian Sarcoma Virus UR2 Helper Virus (LJR2AV) RT, Avian Sarcoma Virus Y73 Helper Virus YAV RT, Rous Associated Virus (RAV) RT, and Myeloblastosis Associated Virus (MAV) RT, all of which may be suitably used in the methods and composition described herein.

[0165] In some embodiments, the RT editor comprises a wild-type M-MLV RT, a functional mutant, a functional variant, or a functional fragment thereof. In some embodiments, the RT editor comprises a reference M-MLV RT, a functional mutant, a functional variant, or a functional fragment thereof.

[0166] In some embodiments, the RT is an RT (e.g., a retro-transposon RT) from an avian genome. The avian may be of the order Galliformes, Aseriformes. Passeriformes, Gruiformes, Slriilhioniformes, Pheiformes, Casuariformes, Apyerygiformes, Olidiformes, Columbiformes, Sphenisciformes, Cathartiformes, Accipilriformes, Slrigiformes, Psittaciformes, Charadriiformes, or Falconiformes. Exemplary avians include those of Galliformes (e.g. chicken, quails, and turkey), Anseriformes (e.g. duck, and goose), Charadriiformes (e.g. gull, barred button quail, and plover), Columbiformes (e.g. pigeon), Struthioniformes (e.g. ostrich), Passeriformes (e.g. crow, finch, sparrow, starling, and swallow), Psittaciformes (e.g. parrot), Falconiformes (e.g. eagle, and falcon), Strigiformes (e.g. owl), Sphenisciformes (e.g. penguin), and Psittaciformes (e.g. parakeet, and parrot). The avian can belong to the order Passeriformes. The avian can belong to the family Passerellidae. The avian can belong to the genus Melospiza. In some embodiments, the RT is an RT from a mammalian genome. In some embodiments, the mammal may be of the order Chiroptera. In some embodiments, the RT is from the genome of M. georgiana. In some embodiments, thegeorgiana A is an engineered RT. In some embodiments, the mutation sites on the M. georgiana RT may be one or more of G206N, M312K, W319F, E336P, and L611W.

[0167] In some embodiments, the RT is an RT from a mammalian genome. In some embodiments, the mammal may be of the order Chiroptera (bat). That bat can be of the family Phyllostomidae (e.g., leaf-nosed bats), Noctilionidae (bulldog bats), Cistugidae, Thyropteridae (e.g., disk-winged bats), Molossidae (e.g., free-tailed bats), Miniopteridae (e.g., long winged bat), Mormoopidae, Mystacinidae (e.g., New Zealand short-tailed bats), Myzopodidae (e.g., suckerfooted bats), Natalidae (e.g., funnel-eared bats), Emballonuridae (e.g., sheath-tailed bats), Nycteridae (e.g., slit-faced bats), Furipteridae (e.g., smoky bats), Vespertilionidae (e.g., vesper bats), Craseonycteridae, Megadermatidae (e.g., false vampire bats), Rhinolophidae (e.g., horseshoe bats), Pteropodidae (e.g., Old World fruit bats), Hipposideridae (e.g., Old World leaf-nosed bats), or Rhinopomatidae. In some embodiments, the bat is of the genus Myotis, Pipistrelles, or Eptesicus. In some embodiments, the bat species is M. brandtii, M. daubentonii, P. kuhlii, or E. nilssonii.Additional sequence elements o f the RT editor mRNA

[0168] The present disclosure provides optimized mRNAs encoding a RT editor, that provide effective genome editing of a target cell population when administered with one or more gRNAs (e.g., tagRNA and egRNA). In some embodiments, additional segmented polyA sequences increase mRNA stability and half-life of the RT editor. In some embodiments, the disclosure provides an mRNA comprising (i) a 5'-cap, (ii) a 5 ’-untranslated region (UTR); (ii) an open reading frame (ORF) comprising a nucleotide sequence that encodes a RT editor; and (iv) a 3' untranslated region (UTR). In some embodiments, the mRNA further comprises viral element sequences 3’ of the 3 ’-UTR.tagRNAs

[0169] Disclosed herein include target priming RNAs (tagRNAs). The term “target priming RNA”, or “tagRNA”, refers to a guide polynucleotide that comprises one or more intended nucleotide edits for incorporation into the target DNA. In some embodiments, the tagRNA associates with and directs a RT editor to incorporate the one or more intended nucleotide edits into the target gene via RT editing. “Nucleotide edit” or “intended nucleotide edit” refers to a specified deletion of one or more nucleotides at one specific position, insertion of one or more nucleotides at one specific position, substitution of a single nucleotide, or other alterations at one specific position to be incorporated into the sequence of the target gene. Intended nucleotide edit may refer to the edit on the editing template as compared to the sequence on the target strand of the target gene or may refer to the edit encoded by the editing template on the newly synthesized single stranded DNA that replaces the editing target sequence, as compared to the editing target sequence. In some embodiments, a tagRNA comprises a spacer sequence that is complementary or substantially complementary to a search target sequence on a target strand of the target gene,in some embodiments, the tagRNA comprises a gRNA core that associates with a DNA endonuclease domain, e.g., a CRISPR-Cas protein domain, of a RT editor. In some embodiments, the tagRNA further comprises an extended nucleotide sequence comprising one or more intended nucleotide edits compared to the endogenous sequence of the target gene, wherein the extended nucleotide sequence may be referred to as an extension arm.

[0170] In some embodiments, in a template armed gRNA (tagRNA), a spacer sequence is complementary or substantially complementary to a specific sequence on the target strand, which may be referred to as a “search target sequence.” In some embodiments, the spacer sequence anneals with the target strand at the search target sequence. The target strand may also be referred to as the “non-Protospacer Adjacent Motif (non-PAM strand).” In some embodiments, the non-target strand may also be referred to as the “PAM strand.” In some embodiments, the PAM strand comprises a protospacer sequence and optionally a protospacer adjacent motif (PAM) sequence. In RT editing using a Cas-protein-based RT editor, a PAM sequence refers to a short DNA sequence immediately adjacent to the protospacer sequence on the PAM strand of the target gene. A PAM sequence may be specifically recognized by a programmable DNA binding protein, e.g., a Cas nickase or a Cas nuclease, in some embodiments, a specific PAM is characteristic of a specific programmable DNA binding protein, e.g., a Cas nickase or a Cas nuclease. A protospacer sequence refers to a specific sequence in the PAM strand of the target gene that is complementary to the search target sequence. In a tagRNA, a spacer sequence may have a substantially identical sequence as the protospacer sequence on the edit strand of a target gene, except that the spacer sequence may comprise Uracil (U) and the protospacer sequence may comprise Thymine (T).

[0171] In some embodiments, the double stranded target DNA comprises a nick site on the PAM strand (or non-target strand). As used herein, a “nick site" refers to a specific position in between two nucleotides or two base pairs of the double stranded target DNA. In some embodiments, the position of a nick site is determined relative to the position of a specific PAM sequence. In some embodiments, the nick site is the particular position where a nick will occur when the double stranded target DNA is contacted with a nickase, for example, a Cas nickase, that recognizes a specific PAM sequence. In some embodiments, the nick site is upstream of a specific PAM sequence on the PAM strand of the double stranded target DNA. In some embodiments, the nick site is downstream of a specific PAM sequence on the PAM strand of the double stranded target DNA. In some embodiments, the nick site is upstream of a PAM sequence recognized by a Cas9 nickase, wherein the Cas9 nickase comprises a nuclease active RuvC domain and a nuclease inactive HNH domain. In some embodiments, the nick site is 3 nucleotides upstream of the PAM sequence, and the PAM sequence is recognized by a Streptococcus pyogenes Cas9 nickase.

[0172] In some embodiments, the nick site is 3 base pairs upstream of the PAM sequence, and the PAM sequence is recognized by a Cas9 nickase, wherein the Cas9 nickase comprises a nuclease active HNH domain and a nuclease inactive RuvC domain. In some embodiments, the nick site is 2 nucleotides upstream of the PAM sequence, and the PAM sequence is recognized by a S. thermophilus Cas9 nickase that comprises a nuclease active RuvC domain and a nuclease inactive HNH domain.

[0173] An “editing template” of a tagRNA is a single-stranded portion of the tagRNA that is 5' of the FB sequence and comprises a region of complementarity to the PAM strand (i.e., the non-target strand or the edit strand), and comprises one or more intended nucleotide edits compared to the endogenous sequence of the double stranded target DNA. In some embodiments, the editing template and the FB sequence are immediately adjacent to each other. Accordingly, in some embodiments, a tagRNA in RT editing comprises a single-stranded portion that comprises the editing template sequence and the FB sequence immediately adjacent to each other. In some embodiments, the single stranded portion of the tagRNA comprising both the editing template sequence and the flap binding sequence is complementary or substantially complementary to an endogenous sequence on the PAM strand (i.e., the non-target strand or the edit strand) of the double stranded target DNA except for one or more non-complementary nucleotides at the intended nucleotide edit positions. As used herein, regardless of relative 5 -3' positioning in other contexts, the relative positions as between the FB sequence and the editing template, and the relative positions as among elements of a tagRNA, are determined by the 5' to 3' order of the tagRNA as a single molecule regardless of the position of sequences in the double stranded target DNA that may have complementarity or identity to elements of the tagRNA. In some embodiments, the editing template is complementary or substantially complementary to a sequence on the PAM strand that is immediately downstream of the nick site, except for one or more non-complementary nucleotides at the intended nucleotide edit positions. The endogenous, e.g., genomic, sequence that is complementary or substantially complementary to the editing template, except for the one or more non-complementary nucleotides at the position corresponding to the intended nucleotide edit, may be referred to as an “editing target sequence." In some embodiments, the editing template has identity or substantial identity to a sequence on the target strand that is complementary to, or having the same position in the genome as, the editing target sequence, except for one or more insertions, deletions, or substitutions at the intended nucleotide edit positions. In some embodiments, the editing template encodes a single stranded DNA, wherein the single stranded DNA has identity or substantial identity to the editing target sequence except for one or more insertions, deletions, or substitutions at the positions of the one or more intended nucleotide edits.

[0174] A “flap binding (FB) sequence” is a single-stranded portion of the tagRNA that comprises a region of complementarity to the PAM strand (i.e., the non-target strand or the edit strand). The FB sequence is complementary or substantially complementary to a sequence on the PAM strand of the double stranded target DNA that is immediately upstream of the nick site. In some embodiments, in the process of RT editing, the tagRNA complexes with and directs a RT editor to bind the search target sequence on the target strand of the double stranded target DNA and the RT editor generates a nick at the nick site on the non-target strand (e.g., the PAM strand) of the double stranded target DNA. In some embodiments, the FB sequence is complementary to or substantially complementary to, and can anneal to, a free 3' end on the non-target strand of the double stranded target DNA at the nick site. In some embodiments, the FB sequence annealed to the free 3' end on the non-target strand can initiate target-primed DNA synthesis. In some embodiments, the FB sequence is about 2 to 20 nucleotides in length. In some embodiments, the FB sequence is about 8 to 16 nucleotides in length. In some embodiments, the FBS is 6 nucleotides in length.Nucleic acid modifications

[0175] In some embodiments, any of the nucleic acids of the disclosure can comprise one or more modifications (e.g., gRNA, tagRNA, egRNA, or any nucleic acid encoding any component of the RT editor systems disclosed herein). In some embodiments, the gRNA (e.g., egRNA) or tagRNA is a chemically modified gRNA or tagRNA. Various types of RNA modifications can be introduced to the gRNAs or tagRNAs to enhance stability, reduce the likelihood or degree of innate immune response, and / or enhance other attributes as described in the art. The gRNAs or tagRNAs described herein can comprise one or more modifications including intemucleoside linkages, purine or pyrimidine bases, or sugar. In some embodiments, a modification is introduced at the terminal of a gRNA or tagRNA with chemical synthesis or with a polymerase enzyme. Examples of modified nucleic acids and their synthesis are disclosed in WO2013 / 052523. Synthesis of modified polynucleotides is also described in Verma and Eckstein, Annual Review of Biochemistry, vol. 76, 99-134 (1998).

[0176] In some embodiments, programmable editing of a target DNA comprises a template armed gRNA (tagRNA). In some embodiments, one or more nucleoside of the tagRNA is chemically modified. In some embodiments, the chemical modifications can be any one of an LNA, 2’ -fluoro, DNA, 2’-0Me, and 2’ MethoxyEthoxy (2’ -MOE)- chemical modification. In some embodiments, each nucleotide of the editing template comprises an LNA modification. In some embodiments, each nucleotide of the flap binding sequence comprises an LNA modification, each nucleotide of both the editing template and the flap binding sequences comprises an LNA modification. In some embodiments, every second nucleotide of the editing template comprisesan LNA modification. In some embodiments, every second nucleotide of the flap binding sequence comprises an LNA modification, every second nucleotide of both the editing template and the flap binding sequences comprises an LNA modification. In some embodiments, every third nucleotide of the editing template comprises an LNA modification. In some embodiments, every third nucleotide of the flap binding sequence comprises an LNA modification, every third nucleotide of both the editing template and the flap binding sequences comprises an LNA modification. In some embodiments, each nucleotide of the editing template comprises a 2’-fluoro modification. In some embodiments, each nucleotide of the flap binding sequence comprises a 2’-fluoro modification, each nucleotide of both the editing template and the flap binding sequences comprises a 2’-fluoro modification. In some embodiments, every second nucleotide of the editing template comprises a 2’ -fluoro modification. In some embodiments, every second nucleotide of the flap binding sequence comprises a 2’-fluoro modification, every second nucleotide of both the editing template and the flap binding sequences comprises a 2’ -fluoro modification. In some embodiments, every third nucleotide of the editing template comprises a 2’-fluoro modification. In some embodiments, every third nucleotide of the flap binding sequence comprises a 2’ -fluoro modification, every third nucleotide of both the editing template and the flap binding sequences comprises a 2’ -fluoro modification. In some embodiments, each nucleotide of the editing template comprises a 2’-0Me modification. In some embodiments, each nucleotide of the flap binding sequence comprises a 2’-0Me modification, each nucleotide of both the editing template and the flap binding sequences comprises a 2’-0Me modification. In some embodiments, every second nucleotide of the editing template comprises a 2’-0Me modification. In some embodiments, every second nucleotide of the flap binding sequence comprises a 2’-0Me modification, every second nucleotide of both the editing template and the flap binding sequences comprises a 2’-0Me modification. In some embodiments, every third nucleotide of the editing template comprises a 2’-0Me modification. In some embodiments, every third nucleotide of the flap binding sequence comprises a 2’-0Me modification, every third nucleotide of both the editing template and the flap binding sequences comprises a 2’-0Me modification. In some embodiments, each nucleotide of the editing template comprises a 2’ -MOE modification. In some embodiments, each nucleotide of the flap binding sequence comprises a 2’ -MOE modification, each nucleotide of both the editing template and the flap binding sequences comprises a 2’ -MOE modification. In some embodiments, every second nucleotide of the editing template comprises a 2’ -MOE modification. In some embodiments, every second nucleotide of the flap binding sequence comprises a 2’ -MOE modification, every second nucleotide of both the editing template and the flap binding sequences comprises a 2’ -MOE modification. In some embodiments, every third nucleotide of the editing template comprises a 2’-M0E modification. In some embodiments, every third nucleotide of theflap binding sequence comprises a 2’ -MOE modification, every third nucleotide of both the editing template and the flap binding sequences comprises a 2’-M0E modification. In some embodiments, the total number of tagRNA nucleotides comprising a chemical modification provided herein (e.g., an LNA, 2’-fluoro, DNA, 2’-0Me, and / or 2’-M0E chemical modification) can be at least, or can be at most, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100, nucleotides.

[0177] In some embodiments, a tagRNA complexes with and directs a RT editor to bind to the search target sequence of the target gene. In some embodiments, the bound RT editor generates a nick on the edit strand (PAM strand) of the target gene at the nick site. In some embodiments, a flap binding site (FB sequence) of the tagRNA anneals with a free 3’ end formed at the nick site, and the RT editor initiates DNA synthesis from the nick site, using the free 3’ end as a primer. Subsequently, a single-stranded DNA encoded by the editing template of the tagRNA is synthesized. In some embodiments, the newly synthesized single-stranded DNA comprises one or more intended nucleotide edits compared to an endogenous target gene sequence. Accordingly, in some embodiments, the editing template of a tagRNA is complementary to a sequence in the edit strand except for one or more mismatches at the intended nucleotide edit positions in the editing template. The endogenous, e.g., genomic, sequence that is partially complementary to the editing template may be referred to as an “editing target sequence.” Accordingly, in some embodiments, the newly synthesized single stranded DNA has identity or substantial identity to a sequence in the editing target sequence, except for one or more insertions, deletions, or substitutions at the intended nucleotide edit positions.

[0178] In some embodiments, the newly synthesized single-stranded DNA equilibrates with the editing target on the edit strand of the target gene for pairing with a target strand of a target gene. In some embodiments, an editing target sequence of a target gene is excised by a flap endonuclease (FEN), for example, FEN1. In some embodiments, the FEN is an endogenous FEN, for example, in a cell comprising a target gene. In some embodiments, the FEN is provided as part of the RT editor, either linked to other components of the RT editor or provided in trans. In some embodiments, the newly synthesized single stranded DNA, which comprises the intended nucleotide edit, replaces the endogenous single stranded editing target sequence on the edit strand of the target gene. In some embodiments, the newly synthesized single stranded DNA and the endogenous DNA on the target strand form a heteroduplex DNA structure at the region corresponding to the editing target sequence of the target gene. In some embodiments, the newlysynthesized single-stranded DNA comprising the nucleotide edit is paired in the heteroduplex with the target strand of the target DNA that does not comprise the nucleotide edit, thereby creating a mismatch between the two otherwise complementary strands. In some embodiments, the mismatch is recognized by DNA repair machinery, e.g., an endogenous DNA repair machinery. In some embodiments, through DNA repair, the intended nucleotide edit is incorporated into the target gene.

[0179] In some embodiments, a RT editor comprises a Cas9 functional variant that is of smaller molecular weight than a wild-type SPYCas9 protein. In some embodiments, a smaller-sized Cas9 functional variant may facilitate delivery to cells, e.g., by an expression vector, nanoparticle, or other means of delivery. In some embodiments, a smaller-sized Cas9 functional variant is a Class 2 Type II Cas protein. In some embodiments, a smaller-sized Cas9 functional variant is a Class 2 Type V Cas protein. In some embodiments, a smaller-sized Cas9 functional variant is a Class 2 Type VI Cas protein.Nuclear Localization Sequences and linkers

[0180] In some embodiments, a RT editor further comprises one or more nuclear localization sequence (NLS). In some embodiments, the NLS helps promote translocation of a protein into the cell nucleus. In some embodiments, a RT editor comprises a fusion protein, e.g., a fusion protein comprising a DNA endonuclease domain and a DNA polymerase, that comprises one or more NLSs. In some embodiments, one or more polypeptides of the RT editor are fused to or linked to one or more NLSs. In some embodiments, the RT editor comprises a DNA endonuclease domain and a DNA polymerase domain that are provided in trans, wherein the DNA endonuclease domain and / or the DNA polymerase domain is fused or linked to one or more NLSs. In some embodiments the RT editor mRNA comprises or consists of the structure: [5’UTR]-[NLS]-[nCas9]-[linker]-[RT]-[NLS]-[3’UTR and / or viral element]-[polyA sequence]. The RT editor mRNA can comprises or consist of the structure [5’-cap]-[5’UTR]-[NLS]-[nCas9]-[linker and / or NLS]-[RT]-[NLS]-[3’UTR_and / or_viral_element]-[secondary structure motif and / or polyA sequence]. In some cases, the positions of nCas9 and RT are swapped to e.g. improve editing efficiency.

[0181] In some embodiments, a RT editor or RT editing complex comprises at least one NLS. In some embodiments, a RT editor or RT editing complex comprises at least two NLSs. In embodiments with at least two NLSs, the NLSs can be the same NLS, or they can be different NLSs.

[0182] In addition, the NLSs can be expressed as part of a RT editor complex. The location of the NLS fusion can be at the N-terminus, the C-terminus, or positioned anywhere within a sequence of a RT editor or a component thereof (e.g., inserted between the DNAendonuclease domain and the DNA polymerase domain of a RT editor fusion protein, between the DNA endonuclease domain and a linker sequence, between a DNA polymerase and a linker sequence, between two linker sequences of a RT editor fusion protein or a component thereof, in either N-terminus to C-terminus or C-terminus to N-terminus order).

[0183] Any NLSs that are known in the art are also contemplated herein. The NLSs may be any naturally occurring NLS, or any non-naturally occurring NLS (e.g., an NLS with one or more mutations relative to a wild-type NLS). In some embodiments, the one or more NLSs of a RT editor comprise bipartite NLSs. In some embodiments, a nuclear localization signal (NLS) is predominantly basic. In some embodiments, the one or more NLSs of a RT editor are rich in lysine and arginine residues. In some embodiments, the one or more NLSs of a RT editor comprise proline residues.

[0184] In some embodiments, components of a RT editor are directly fused to each other. In some embodiments, components of a RT editor are associated to each other via a linker. As used herein, a linker can be any chemical group or a molecule linking two molecules or moieties, e.g., a DNA binding domain, a DNA endonuclease domain and a polymerase domain of a RT editor. In some embodiments, a linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker comprises a non-peptide moiety. The linker may be a covalent bond (e.g., a carbon-carbon bond, disulfide bond, carbon-heteroatom bond, etc.), or it may be a polymeric linker, for example, a polynucleotide sequence.

[0185] In some embodiments, two or more components of a RT editor are linked to each other by a peptide linker. In some embodiments, a peptide linker is 5-100 amino acids in length, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. In some embodiments, the peptide linker is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140,150, 160, 175, 180, 190, or 200 amino acids in length. In some embodiments, the peptide linker is 5-100 amino acids in length. In some embodiments, the peptide linker is 10-80 amino acids in length. In some embodiments, the peptide linker is 15-70 amino acids in length. In some embodiments, the peptide linker is 16 amino acids in length, 24 amino acids in length, 64 amino acids in length, or 96 amino acids in length, in some embodiments, the peptide linker is at least 50 amino acids in length, in some embodiments, the peptide linker is at least 40 amino acids in length, in some embodiments, the peptide linker is at least 30 amino acids in length. In some embodiments, the peptide linker is 46 amino acids in length. In some embodiments, the peptide linker is 92 amino acids in length. In some embodiments, the peptidelinker is 16 amino acids in length, 24 amino acids in length, 64 amino acids in length, or 96 amino acids in length.tagRNAs

[0186] Disclosed herein include template armed gRNAs (tagRNAs). In some embodiments, the tagRNA comprises: a spacer that is complementary to a search target sequence on a first strand of a target gene; a scaffold sequence, an editing template that comprises a region of complementarity to an editing target sequence on a second strand of the target gene; and a flap binding sequence, wherein the first strand and the second strand are complementary to each other. In some embodiments, the editing target sequence is within the mutation site of the target gene. In some embodiments, the editing target sequence is within exon 5 of the SERPINA1 gene. In some embodiments, the editing target sequence is within exon 12 of the PAH gene.

[0187] The tagRNA can comprise from 5’ to 3’: the spacer, the scaffold sequence, the editing template, and the FB sequence. In some embodiments, spacer, the scaffold sequence, the editing template, and the FB sequence form a contiguous sequence in a single molecule. The editing template can comprise an intended nucleotide edit compared to the target gene. In some embodiments, the tagRNA guides the RT editor to incorporate the intended nucleotide edit into the target gene when the tagRNA is contacted with the target gene. In some embodiments, the RT editor synthesizes a single stranded DNA encoded by the editing template, wherein the single stranded DNA replaces the editing target sequence and results in incorporation of the intended nucleotide edit into a region corresponding to the editing target in the target gene. The search target sequence can be complementary to a protospacer sequence in the target gene. The protospacer sequence can be adjacent to a protospacer adjacent motif (PAM) in the target gene. In some embodiments, the tagRNA results in incorporation of the intended nucleotide edit in the PAM when contacted with the target gene. In some embodiments, the target gene is the SERPINA1 gene. In some embodiments, the target gene is the BCL11A gene. In some embodiments, the target gene is the PAH gene.

[0188] In some embodiments, the extension arm comprises a flap binding site sequence (FB sequence) that can initiate target-primed DNA synthesis. In some embodiments, the FB sequence is complementary or substantially complementary to a free 3’ end on the edit strand of the target gene at a nick site generated by the RT editor. In some embodiments, the extension arm further comprises an editing template that comprises one or more intended nucleotide edits to be incorporated in the target gene by RT editing. In some embodiments, the editing template is a template for an RNA-dependent DNA polymerase domain or polypeptide of the RT editor, for example, a reverse transcriptase domain. The reverse transcriptase editing template may also be referred to herein as an editing template. In some embodiments, the editing template comprisespartial complementarity to an editing target sequence in the target gene, e.g., an BCL11 A gene, a PAH gene, or a SERPINA1 gene. In some embodiments, the editing template comprises substantial or partial complementarity to the editing target sequence except at the position of the intended nucleotide edits to be incorporated into the target gene.

[0189] In some embodiments, a tagRNA includes only RNA nucleotides and forms an RNA polynucleotide. In some embodiments, a tagRNA is a chimeric polynucleotide that includes both RNA and DNA nucleotides. For example, a tagRNA can include DNA in the spacer sequence, the scaffold sequence, or the extension arm. In some embodiments, a tagRNA comprises DNA in the spacer sequence. In some embodiments, the entire spacer sequence of a tagRNA is a DNA sequence. In some embodiments, the tagRNA comprises DNA in the scaffold sequence, for example, in a stem region of the scaffold sequence. In some embodiments, the tagRNA comprises DNA in the extension arm, for example, in the editing template. An editing template that comprises a DNA sequence may serve as a DNA synthesis template for a DNA polymerase in a RT editor, for example, a DNA-dependent DNA polymerase. Accordingly, the tagRNA may be a chimeric polynucleotide that comprises RNA in the spacer, scaffold sequence, and / or the FB sequence sequences and DNA in the editing template.

[0190] In some embodiments, a spacer sequence comprises a region that has substantial complementarity to a search target sequence on the target strand of a double stranded target DNA, e.g., an AT7B gene, a PAH gene, or a SERPINA1 gene. In some embodiments, the spacer sequence of a tagRNA is identical or substantially identical to a protospacer sequence on the edit strand of the target gene (except that the protospacer sequence comprises thymine and the spacer sequence may comprise uracil). In some embodiments, the spacer sequence is at least about 70%, 75%, 80%, 85%, 90%, 95%, or 100% complementary to a search target sequence in the target gene. In some embodiments, the spacer comprises is substantially complementary to the search target sequence.

[0191] The spacer of the tagRNA can be from 16 to 25 nucleotides in length. The spacer can be 20 nucleotides in length. The spacer can be 21-23 nucleotides in length. In some embodiments, the length of the spacer varies from about 10 to about 100 nucleotides. In some embodiments, the spacer is 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, 20 nucleotides, 21 nucleotides, 22 nucleotides, 23 nucleotides, 24 nucleotides, or 25 nucleotides in length. In some embodiments, the spacer is from 15 nucleotides to 30 nucleotides in length, 15 to 25 nucleotides in length, 18 to 22 nucleotides in length, 10 to 20 nucleotides in length, or 20 to 30 nucleotides in length. In some embodiments, the spacer is 16 to 22 nucleotides in length. In some embodiments, the spacer is 16 to 20 nucleotides in length. In some embodiments, the spacer is 17 to 18 nucleotides in length.

[0192] As used herein in a gRNA (e.g., a tagRNA or a enhancer guide egRNA sequence), or fragments thereof such as a spacer, FB sequence, or editing template sequence, unless indicated otherwise, it should be appreciated that the letter “T” or “thymine” indicates a nucleobase in a DNA sequence that encodes the tagRNA or guide RNA sequence, and is intended to refer to a uracil (U) nucleobase of the tagRNA or guide RNA or any chemically modified uracil nucleobase known in the art, such as 5 -methoxyuracil.

[0193] The extension arm of a tagRNA may comprise a flap binding sequence (FB sequence; FBS) and an editing template (e.g., an ET). The extension arm may be partially complementary to the spacer. In some embodiments, the editing template (e.g., ET) is partially complementary to the spacer. In some embodiments, the editing template (e.g., ET) and the flap binding sequence (FB sequence; FBS) are each partially complementary to the spacer. An extension arm of a tagRNA may comprise a flap binding sequence sequence (FB sequence, or FBS) that comprises complementarity to and can hybridize with a free 3’ end of a single stranded DNA in the target gene (e.g., the BCL11A gene, PAH gene, or SERPINA1 gene) generated by nicking with a RT editor at the nick site on the PAM strand.

[0194] The length of the FB sequence may vary depending on, e.g., the RT editor components, the search target sequence and other components of the tagRNA. The FB sequence can be about 2 to 20 nucleotides in length. The FB sequence can be about 8 to 16 nucleotides in length. In some embodiments, the FB sequence is 6 nucleotides in length In some embodiments, the FB sequence is about 3 to 19 nucleotides in length, or about 3 to 17 nucleotides in length. In some embodiments, the FB sequence is about 4 to 16 nucleotides, about 6 to 16 nucleotides, about 6 to 18 nucleotides, about 6 to 20 nucleotides, about 8 to 20 nucleotides, about 10 to 20 nucleotides, about 12 to 20 nucleotides, about 14 to 20 nucleotides, about 16 to 20 nucleotides, or about 18 to 20 nucleotides in length. In some embodiments, the FB sequence is 8 to 17 nucleotides in length. In some embodiments, the FB sequence is 8 to 16 nucleotides in length. In some embodiments, the FB sequence is 8 to 15 nucleotides in length. In some embodiments, the FB sequence is 8 to 14 nucleotides in length. In some embodiments, the FB sequence is 8 to 13 nucleotides in length. In some embodiments, the FB sequence is 8 to 12 nucleotides in length. In some embodiments, the FB sequence is 8 to 11 nucleotides in length. In some embodiments, the FB sequence is 8 to 10 nucleotides in length. In some embodiments, the FB sequence is 8 or 9 nucleotides in length. In some embodiments, the FB sequence is 16 or 17 nucleotides in length, in some embodiments, the FB sequence is 15 to 17 nucleotides in length. In some embodiments, the FB sequence is 14 to 17 nucleotides in length. In some embodiments, the FB sequence is 13 to 17 nucleotides in length. In some embodiments, the FB sequence is 12 to 17 nucleotides in length. In some embodiments, the FB sequence is 11 to 17 nucleotides in length. In some embodiments, theFB sequence is 10 to 17 nucleotides in length. In some embodiments, the FB sequence is 9 to 17 nucleotides in length. In some embodiments, the FB sequence is about 7 to 15 nucleotides in length. In some embodiments, the FB sequence is 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 nucleotides in length. In some embodiments, the FB sequence is 8, 9, 10, 11, 12, 13, or 14 nucleotides in length.

[0195] The FB sequence may be complementary or substantially complementary to a DNA sequence in the edit strand of the target gene. By annealing with the edit strand at a free hydroxy group, e.g., a free 3’ end generated by RT editor nicking activity, the FB sequence may initiate synthesis of a new single stranded DNA encoded by the editing template at the nick site. In some embodiments, the FB sequence is at least about 70%, 75%, 80%, 85%, 90%, 95%, or 100% complementary to a region of the edit strand of the target gene (e.g., the BCL11A gene, PAH gene, or SERPINA1 gene).

[0196] An extension arm of a tagRNA may comprise an editing template that serves as a DNA synthesis template for the DNA polymerase in a RT editor during RT editing. The length of an editing template may vary depending on, e.g., the RT editor components, the search target sequence and other components of the tagRNA. In some embodiments, the editing template serves as a DNA synthesis template for a reverse transcriptase. The editing template can be about 4 to 30 nucleotides in length. The editing template can be about 10 to 30 nucleotides in length. The editing template can be 6 to 9 nucleotides in length. In some embodiments, the editing template is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length.

[0197] In some embodiments, the editing template sequence is about 70%, 75%, 80%, 85%, 90%, 95%, or 99% complementary to the editing target sequence on the edit strand of the target gene. In some embodiments, the editing template sequence is complementary to the editing target sequence except at positions of the intended nucleotide edits to be incorporated int the target gene. In some embodiments, the editing template comprises a nucleotide sequence comprising about 85% to about 95% complementarity to (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values, complementarity to) an editing target sequence in the edit strand in the target gene. In some embodiments, the editing template comprises about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% complementarity to an editing target sequence in the edit strand of the target gene (e.g., the BCL11 A gene, PAH gene, or SERPINA1 gene).

[0198] An intended nucleotide edit or intended nucleotide edits in an editing template of a tagRNA may comprise various types of alterations as compared to the target gene sequence, in some embodiments, the nucleotide edit or edits is nucleotide substitution(s) as compared to the target gene sequence. In some embodiments, the nucleotide edit comprise deletion(s) as compared to the target gene sequence. In some embodiments, the nucleotide edit comprise an insertion as compared to the target gene sequence. In some embodiments, the editing template comprises one to ten intended nucleotide edits as compared to the target gene sequence, in some embodiments, the editing template comprises one or more intended nucleotide edits as compared to the target gene sequence. In some embodiments, the editing template comprises two or more intended nucleotide edits as compared to the target gene sequence. In some embodiments, the editing template comprises three or more intended nucleotide edits as compared to the target gene sequence. In some embodiments, the editing template comprises four or more, five or more, or six or more intended nucleotide edits as compared to the target gene sequence. In some embodiments, the editing template comprises two or more single nucleotide substitutions, insertions, deletions, or any combination thereof, as compared to the target gene sequence. In some embodiments, the editing template comprises three single nucleotide substitutions, insertions, deletions, or any combination thereof as compared to the target gene sequence. In some embodiments, the editing template comprises four, five, or six single nucleotide substitutions, insertions, deletions, or any combination thereof, as compared to the target gene sequence. In some embodiments, a nucleotide substitution comprises an adenine (A)-to-thymine (T) substitution. In some embodiments, a nucleotide substitution comprises an A-to-guanine (G) substitution. In some embodiments, a nucleotide substitution comprises an A-to-cytosine (C) substitution. In some embodiments, a nucleotide substitution comprises a T-A substitution. In some embodiments, a nucleotide substitution comprises a T-G substitution, in some embodiments, a nucleotide substitution comprises a T-C substitution. In some embodiments, a nucleotide substitution comprises a G-to-A substitution. In some embodiments, a nucleotide substitution comprises a G-to-T substitution. In some embodiments, a nucleotide substitution comprises a G-to-C substitution. In some embodiments, a nucleotide substitution comprises a C-to-A substitution, in some embodiments, a nucleotide substitution comprises a C-to-T substitution. In some embodiments, a nucleotide substitution comprises a C-to-G substitution.

[0199] In some embodiments, a nucleotide insertion is at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 22 nucleotides, at least 24nucleotides, at least 26 nucleotides, at least 28 nucleotides, at least 30 nucleotides, at least 32 nucleotides, at least 34 nucleotides, at least 36 nucleotides, at least 38 nucleotides, at least 40 nucleotides, at least 42 nucleotides, at least 44 nucleotides, at least 46 nucleotides, at least 48 nucleotides, or at least 50 nucleotides, in length. In some embodiments, a nucleotide insertion is from 1 to 2 nucleotides, from 1 to 3 nucleotides, from 1 to 4 nucleotides, from 1 to 5 nucleotides, form 2 to 5 nucleotides, from 3 to 5 nucleotides, from 3 to 6 nucleotides, from 3 to 8 nucleotides, from 4 to 9 nucleotides, from 5 to 10 nucleotides, from 6 to 11 nucleotides, from 7 to 12 nucleotides, from 8 to 13 nucleotides, from 9 to 14 nucleotides, from 10 to 15 nucleotides, from 11 to 16 nucleotides, from 12 to 17 nucleotides, from 13 to 18 nucleotides, from 14 to 19 nucleotides, from 15 to 20 nucleotides in length. In some embodiments, a nucleotide insertion is a single nucleotide insertion. In some embodiments, a nucleotide insertion comprises insertion of two nucleotides.

[0200] The editing template of a tagRNA may comprise one or more intended nucleotide edits, compared to the target gene to be edited. Position of the intended nucleotide edit(s) relevant to other components of the tagRNA, or to particular nucleotides (e.g., mutations) in the target gene may vary. In some embodiments, the nucleotide edit is in a region of the tagRNA corresponding to or homologous to the protospacer sequence. In some embodiments, the nucleotide edit is in a region of the tagRNA corresponding to a region of the target gene outside of the protospacer sequence.

[0201] In some embodiments, the position of a nucleotide edit incorporation in the target gene may be determined based on position of the protospacer adjacent motif (PAM). For instance, the intended nucleotide edit may be installed in a sequence corresponding to the protospacer adjacent motif (PAM) sequence. In some embodiments, a nucleotide edit in the editing template is at a position corresponding to the 5’ most nucleotide of the PAM sequence. In some embodiments, a nucleotide edit in the editing template is at a position corresponding to the 3’ most nucleotide of the PAM sequence, in some embodiments, position of an intended nucleotide edit in the editing template may be referred to by aligning the editing template with the partially complementary edit strand of the target gene, and referring to nucleotide positions on the editing strand where the intended nucleotide edit is incorporated, in some embodiments, a nucleotide edit is incorporated at a position corresponding to about 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides upstream of the 5’ most nucleotide of the PAM sequence in the edit strand of the target gene. By 0 base pair upstream or downstream of a reference position, it is meant that the intended nucleotide is immediately upstream or downstream of the reference position. In some embodiments, a nucleotide edit is incorporated at a position corresponding toabout 0 to 2 nucleotides, 0 to 4 nucleotides, 0 to 6 nucleotides, 0 to 8 nucleotides, 0 to 10 nucleotides, 2 to 4 nucleotides, 2 to 6 nucleotides, 2 to 8 nucleotides, 2 to 10 nucleotides, 2 to 12 nucleotides, 4 to 6 nucleotides, 4 to 8 nucleotides, 4 to 10 nucleotides, 4 to 12 nucleotides, 4 to 14 nucleotides, 6 to 8 nucleotides, 6 to 10 nucleotides, 6 to 12 nucleotides, 6 to 14 nucleotides, 6 to 16 nucleotides, 8 to 10 nucleotides, 8 to 12 nucleotides, 8 to 14 nucleotides, 8 to 16 nucleotides, 8 to 18 nucleotides, 10 to 12 nucleotides, 10 to 14 nucleotides, 10 to 16 nucleotides, 10 to 18 nucleotides, 10 to 20 nucleotides, 12 to 14 nucleotides, 12 to 16 nucleotides, 12 to 18 nucleotides, 12 to 20 nucleotides, 12 to 22 nucleotides, 14 to 16 nucleotides, 14 to 18 nucleotides, 14 to 20 nucleotides, 14 to 22 nucleotides, 14 to 24 nucleotides, 16 to 18 nucleotides, 16 to 20 nucleotides, 16 to 22 nucleotides, 16 to 24 nucleotides, 16 to 26 nucleotides, 18 to 20 nucleotides, 18 to 22 nucleotides, 18 to 24 nucleotides, 18 to 26 nucleotides, 18 to 28 nucleotides, 20 to 22 nucleotides, 20 to 24 nucleotides, 20 to 26 nucleotides, 20 to 28 nucleotides, or 20 to 30 nucleotides upstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, the nucleotide edit is incorporated at a position corresponding to 3 nucleotides upstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, the nucleotide edit in is incorporated at a position corresponding to 4 nucleotides upstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, the nucleotide edit is incorporated at a position corresponding to 5 nucleotides upstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, the nucleotide edit in the editing template is at a position corresponding to 6 nucleotides upstream of the 5’ most nucleotide of the PAM sequence.

[0202] In some embodiments, an intended nucleotide edit is incorporated at a position corresponding to about 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides downstream of the 5’ most nucleotide of the PAM sequence in the edit strand of the target gene. In some embodiments, a nucleotide edit is incorporated at a position corresponding to about 0 to 2 nucleotides, 0 to 4 nucleotides, 0 to 6 nucleotides, 0 to 8 nucleotides, 0 to 10 nucleotides, 2 to 4 nucleotides, 2 to 6 nucleotides, 2 to 8 nucleotides, 2 to 10 nucleotides, 2 to 12 nucleotides, 4 to 6 nucleotides, 4 to 8 nucleotides, 4 to 10 nucleotides, 4 to 12 nucleotides, 4 to 14 nucleotides, 6 to 8 nucleotides, 6 to 10 nucleotides, 6 to 12 nucleotides, 6 to 14 nucleotides, 6 to 16 nucleotides, 8 to 10 nucleotides, 8 to 12 nucleotides, 8 to 14 nucleotides, 8 to 16 nucleotides, 8 to 18 nucleotides, 10 to 12 nucleotides, 10 to 14 nucleotides, 10 to 16 nucleotides, 10 to 18 nucleotides, 10 to 20 nucleotides, 12 to 14 nucleotides, 12 to 16 nucleotides, 12 to 18 nucleotides, 12 to 20 nucleotides, 12 to 22 nucleotides, 14 to 16 nucleotides, 14 to 18 nucleotides, 14 to 20 nucleotides, 14 to 22 nucleotides, 14 to 24 nucleotides, 16 to 18 nucleotides, 16 to 20 nucleotides, 16 to 22 nucleotides, 16 to 24 nucleotides, 16 to 26 nucleotides, 18 to 20 nucleotides, 18 to 22 nucleotides, 18 to 24-n-nucleotides, 18 to 26 nucleotides, 18 to 28 nucleotides, 20 to 22 nucleotides, 20 to 24 nucleotides, 20 to 26 nucleotides, 20 to 28 nucleotides, or 20 to 30 nucleotides downstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, a nucleotide edit is incorporated at a position corresponding to 3 nucleotides downstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, a nucleotide edit is incorporated at a position corresponding to 4 nucleotides downstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, a nucleotide edit is incorporated at a position corresponding to 5 nucleotides downstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, a nucleotide edit is incorporated at a position corresponding to 6 nucleotides downstream of the 5’ most nucleotide of the PAM sequence.

[0203] In some embodiments, the position of a nucleotide edit incorporation in the target gene can be determined based on position of the nick site. In some embodiments, position ofan intended nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, or 150 nucleotides apart from the nick site. In some embodiments, position of an intended nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, or 150 nucleotides downstream of the nick site on the PAM strand (or the non-target strand, or the edit strand) of the double stranded target DNA. In some embodiments, position of the intended nucleotide edit in the editing template can be referred to by aligning the editing template with the partially complementary editing target sequence on the edit strand and referring to nucleotide positions on the editing strand where the intended nucleotide edit is incorporated. Accordingly, in some embodiments, a nucleotide edit in an editing template is at a position corresponding to a position about 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, or 150 nucleotides apart from the nick site. In some embodiments, a nucleotide edit in an editing template is at a position corresponding to a position about 0 to 2 nucleotides, 0 to 4 nucleotides, 0 to 6 nucleotides, 0 to 8 nucleotides, 0 to 10 nucleotides, 2 to 4 nucleotides, 2 to 6 nucleotides, 2 to 8 nucleotides, 2 to 10 nucleotides, 2 to 12 nucleotides, 4 to 6 nucleotides, 4 to 8 nucleotides, 4 to 10 nucleotides, 4 to 12 nucleotides, 4 to 14 nucleotides, 6 to 8 nucleotides, 6 to 10 nucleotides, 6 to 12 nucleotides, 6 to 14 nucleotides, 6 tol6 nucleotides, 8 to 10 nucleotides, 8 to 12 nucleotides, 8 to 14 nucleotides, 8 to 16 nucleotides, 8 to 18 nucleotides, 10 to 12 nucleotides, 10 to 14 nucleotides, 10 to 16 nucleotides, 10 to 18 nucleotides, 10 to 20 nucleotides, 12 to 14 nucleotides, 12 to 16 nucleotides, 12 to 18 nucleotides, 12 to 20 nucleotides, 12 to 22nucleotides, 14 to 16 nucleotides, 14 to 18 nucleotides, 14 to 20 nucleotides, 14 to 22 nucleotides, 14 to 24 nucleotides, 16 to 18 nucleotides, 16 to 20 nucleotides, 16 to 22 nucleotides, 16 to 24 nucleotides, 16 to 26 nucleotides, 18 to 20 nucleotides, 18 to 22 nucleotides, 18 to 24 nucleotides, 18 to 26 nucleotides, 18 to 28 nucleotides, 20 to 22 nucleotides, 20 to 24 nucleotides, 20 to 26 nucleotides, 20 to 28 nucleotides, 20 to 30 nucleotides, 30 to 40 nucleotides, 40 to 50 nucleotides, 50 to 60 nucleotides, 60 to 70 nucleotides, 70 to 80 nucleotides, 80 to 90 nucleotides, 90 to 100 nucleotides, 100 to 110 nucleotides, 110 to 120 nucleotides, 120 to 130 nucleotides, 130 to 140 nucleotides, or 140 to 150 nucleotides apart from the nick site. In some embodiments, when referred to in the context of the PAM strand (or the non-target strand, or the edit strand), a nucleotide edit in an editing template is at a position corresponding to a position about 0 to 2 nucleotides, 0 to 4 nucleotides, 0 to 6 nucleotides, 0 to 8 nucleotides, 0 to 10 nucleotides, 2 to 4 nucleotides, 2 to 6 nucleotides, 2 to 8 nucleotides, 2 to 10 nucleotides, 2 to 12 nucleotides, 4 to 6 nucleotides, 4 to 8 nucleotides, 4 to 10 nucleotides, 4 to 12 nucleotides, 4 to 14 nucleotides, 6 to 8 nucleotides, 6 to 10 nucleotides, 6 to 12 nucleotides, 6 to 14 nucleotides, 6 tol6 nucleotides, 8 to 10 nucleotides, 8 to 12 nucleotides, 8 to 14 nucleotides, 8 to 16 nucleotides, 8 to 18 nucleotides, 10 to 12 nucleotides, 10 to 14 nucleotides, 10 to 16 nucleotides, 10 to 18 nucleotides, 10 to 20 nucleotides, 12 to 14 nucleotides, 12 to 16 nucleotides, 12 to 18 nucleotides, 12 to 20 nucleotides, 12 to 22 nucleotides, 14 to 16 nucleotides, 14 to 18 nucleotides, 14 to 20 nucleotides, 14 to 22 nucleotides, 14 to 24 nucleotides, 16 to 18 nucleotides, 16 to 20 nucleotides, 16 to 22 nucleotides, 16 to 24 nucleotides, 16 to 26 nucleotides, 18 to 20 nucleotides, 18 to 22 nucleotides, 18 to 24 nucleotides, 18 to 26 nucleotides, 18 to 28 nucleotides, 20 to 22 nucleotides, 20 to 24 nucleotides, 20 to 26 nucleotides, 20 to 28 nucleotides, 20 to 30 nucleotides, 30 to 40 nucleotides, 40 to 50 nucleotides, 50 to 60 nucleotides, 60 to 70 nucleotides, 70 to 80 nucleotides, 80 to 90 nucleotides, 90 to 100 nucleotides, 100 to 110 nucleotides, 110 to 120 nucleotides, 120 to 130 nucleotides, 130 to 140 nucleotides, or 140 to 150 nucleotides downstream from the nick site. The relative positions of the intended nucleotide edit(s) and nick site may be referred to by numbers. For example, in some embodiments, the nucleotide immediately downstream of the nick site on a PAM strand (or the non-target strand, or the edit strand) may be referred to as at position 0. The nucleotide immediately upstream of the nick site on the PAM strand (or the non-target strand, or the edit strand) may be referred to as at position -1. The nucleotides downstream of position 0 on the PAM strand can be referred to as at positions +1, +2, +3, +4, ... +n, and the nucleotides upstream of position -1 on the PAM strand may be referred to as at positions -2, -3, -4, .. -n. Accordingly, in some embodiments, the nucleotide in the editing template that corresponds to position 0 when the editing template is aligned with the partially complementary editing target sequence by complementarity can also be referred to as position 0 in the editing template, the nucleotides inthe editing template corresponding to the nucleotides at positions +1, +2, +3, +4, +n on the PAM strand of the double stranded target DNA can also be referred to as at positions +1, +2, +3, +4, -in in the editing template, and the nucleotides in the editing template corresponding to the nucleotides at positions -1, -2, -3, -4, -n on the PAM strand on the double stranded target DNA may also be referred to as at positions -1, -2, -3, -4 -n on the editing template, even though when the tagRNA is viewed as a standalone nucleic acid, positions +1, +2, +3, +4, ..., +n are 5' of position 0 and positions -1, -2, -3, -4, ...-n are 3' of position 0 in the editing template. In some embodiments, an intended nucleotide edit is at position +n of the editing template relative to position 0. Accordingly, the intended nucleotide edit may be incorporated at position +n of the PAM strand of the double stranded target DNA (and subsequently, the target strand of the double stranded target DNA) by RT editing. The number n may be referred to as the nick to edit distance.

[0204] When referred to within the tagRNA, positions of the one or more intended nucleotide edits may be referred to relevant to components of the tagRNA. For example, an intended nucleotide edit may be 5’ or 3’ to the FB sequence, in some embodiments, a tagRNA comprises the structure, from 5’ to 3’: a spacer, a gRNA core, an editing template, and a FB sequence. In some embodiments, the intended nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides upstream to the 5’ most nucleotide of the FB sequence, in some embodiments, the intended nucleotide edit is 0 to 2 nucleotides, 0 to 4 nucleotides, 0 to 6 nucleotides, 0 to 8 nucleotides, 0 to 10 nucleotides, 2 to 4 nucleotides, 2 to 6 nucleotides, 2 to 8 nucleotides, 2 to 10 nucleotides, 2 to 12 nucleotides, 4 to 6 nucleotides, 4 to 8 nucleotides, 4 to 10 nucleotides, 4 to 12 nucleotides, 4 to 14 nucleotides, 6 to 8 nucleotides, 6 to 10 nucleotides, 6 to 12 nucleotides, 6 to 14 nucleotides, 6 tol6 nucleotides, 8 to 10 nucleotides, 8 to 12 nucleotides, 8 to 14 nucleotides, 8 to 16 nucleotides, 8 to 18 nucleotides, 10 to 12 nucleotides, 10 to 14 nucleotides, 10 to 16 nucleotides, 10 to 18 nucleotides, 10 to 20 nucleotides, 12 to 14 nucleotides, 12 to 16 nucleotides, 12 to 18 nucleotides, 12 to 20 nucleotides, 12 to 22 nucleotides, 14 to 16 nucleotides, 14 to 18 nucleotides, 14 to 20 nucleotides, 14 to 22 nucleotides, 14 to 24 nucleotides, 16 to 18 nucleotides, 16 to 20 nucleotides, 16 to 22 nucleotides, 16 to 24 nucleotides, 16 to 26 nucleotides, 18 to 20 nucleotides, 18 to 22 nucleotides, 18 to 24 nucleotides, 18 to 26 nucleotides, 18 to 28 nucleotides, 20 to 22 nucleotides, 20 to 24 nucleotides, 20 to 26 nucleotides, 20 to 28 nucleotides, or 20 to 30 nucleotides upstream to the 5’ most nucleotide of the FB sequence.

[0205] The corresponding positions of the intended nucleotide edit incorporated in the target gene may also be referred to based on the nicking position (i.e., the nick site) generated by a RT editor based on sequence homology and complementarity. For example, in some embodiments, the distance between the intended nucleotide edit to be incorporated into the targetgene and the nick site (also referred to as the “nick to edit distance”) may be determined by the position of the nick site and the position of the nucleotide(s) corresponding to the intended nucleotide edit(s), for example, by identifying sequence complementarity between the spacer and the search target sequence and sequence complementarity between the editing template and the editing target sequence. In some embodiments, the position of the nucleotide edit can be in any position downstream of the nick site on the edit strand (or the PAM strand) generated by the RT editor, such that the distance between the nick site and the intended nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. In some embodiments, the position of the nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides upstream of the nick site on the edit strand. In some embodiments, the position of the nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides downstream of the nick site on the edit strand. In some embodiments, the position of the nucleotide edit is 0 base pair from the nick site on the edit strand, that is, the editing position is at the same position as the nick site. As used herein, the distance between the nick site and the nucleotide edit, for example, where the nucleotide edit comprises an insertion or deletion, refers to the 5’ most position of the nucleotide edit for a nick that creates a 3’ free end on the edit strand (i.e., the “near position” of the nucleotide edit to the nick site). Similarly, as used herein, the distance between the nick site and a PAM position edit, for example, where the nucleotide edit comprises an insertion, deletion, or substitution of two or more contiguous nucleotides, refers to the 5’ most position of the nucleotide edit and the 5’ most position of the PAM sequence.

[0206] In some embodiments, the editing template extends beyond a nucleotide edit to be incorporated into the target gene sequence. For example, in some embodiments, the editing template comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80 nucleotides 3’ to the nucleotide edit to be incorporated into the target gene sequence. In some embodiments, the editing template comprises 1 to 80 nucleotides 3’ to the nucleotide edit to be incorporated into the target gene sequence.

[0207] In some embodiments, the editing template can comprise a second editing sequence comprising a second mutation relative to a target sequence. The second mutation can be designed to mutate or otherwise silence a PAM sequence such that a corresponding nucleic acid guided nuclease or CRISPR nuclease is no longer able to cleave the target sequence. In some embodiments, this mutation or silencing of a PAM can serve as a method for selectingtransformants in which the first editing sequence has been incorporated. In some embodiments, the mutation is in at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleic acids in a PAM motif.

[0208] The editing template of a tagRNA may encode a new single stranded DNA (e.g., by reverse transcription) to replace a editing target sequence in the target gene. In some embodiments, the editing target sequence in the edit strand of the target gene is replaced by the newly synthesized strand, and the nucleotide edit(s) are incorporated into the region of the target gene.

[0209] In some embodiments, the target gene is SERPINA1 gene. In some embodiments, the editing template of the tagRNA encodes a newly synthesized single stranded DNA that comprises a wild type SERPINA1 gene sequence. In some embodiments, the newly synthesized DNA strand replaces the editing target sequence in the target SERPINA1 gene, wherein the editing target sequence (or the endogenous sequence complementary to the editing target sequence on the target strand of the SERPINA1 gene) comprises a mutation compared to a wild-type SERPINA1 gene. In some embodiments, the mutation is associated with A1ATD disease or disorder.

[0210] In some embodiments, the target gene is the PAH gene. In some embodiments, the editing template of the tagRNA encodes a newly synthesized single stranded DNA that comprises a wild type PAH gene sequence. In some embodiments, the newly synthesized DNA strand replaces the editing target sequence in the target PAH gene, wherein the editing target sequence (or the endogenous sequence complementary to the editing target sequence on the target strand of the PAH gene) comprises a mutation compared to a wild-type PAH gene. In some embodiments, the mutation is associated with phenylketonuria or hyperphenylalaninemia.

[0211] In some embodiments, the target gene is the BCL11A gene. In some embodiments, the editing template of the tagRNA encodes a newly synthesized single stranded DNA that comprises a wild type BCL11A gene sequence. In some embodiments, the newly synthesized DNA strand replaces the editing target sequence in the target BCL11 A gene, wherein the editing target sequence (or the endogenous sequence complementary to the editing target sequence on the target strand of the BCL11 A gene) comprises a mutation compared to a wild-type BCL11A gene. In some embodiments, the mutation is associated with sickle cell disease and P-thalassemia.

[0212] In some embodiments, the mutation is the S allele mutation or the Z allele mutation in SERPINA1 gene associated with alpha-1 antitrypsin deficiency disease (Al ATD). In some embodiments, the mutation is the Z allele mutation on exon 5 of SERPINA1. In some embodiment, the mutation is an E342K mutation commonly associated with A1ATD. E342K mutation leads to misfolding of the alpha- 1 -antitrypsin (AAT) protein leading to polymers andliver damage. In some embodiments, the mutation is the R408W mutation in PAH gene associated with phenylketonuria or hyperphenylalaninemia.

[0213] In some embodiments, the tagRNA results in incorporation of the intended nucleotide edit about 0 to 27 base pairs downstream (e.g., 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27 base pairs downstream) of the 5’ end of the PAM when contacted with the target gene. The intended nucleotide edit can comprise a single nucleotide substitution compared to the region corresponding to the editing target in the target gene. The single nucleotide substitution can be a T>G substitution. The intended nucleotide edit can comprise an insertion compared to the region corresponding to the editing target in the target gene. The intended nucleotide edit can comprise a deletion compared to the region corresponding to the editing target in the target gene. The editing target sequence can comprise a mutation associated with a disease or disorder. In some embodiments, the mutation encodes an amino acid substitution. In some embodiments, the mutation is the S allele mutation or the Z allele mutation in SERPINA1 gene associated with alpha-1 antitrypsin deficiency disease (Al ATD). The Z allele mutation on exon 5 of SERPINA1 is an E342K mutation commonly associated with A1ATD. E342K mutation leads to misfolding of the alpha- 1 -antitrypsin (AAT) protein leading to polymers and liver damage. The editing template can comprise a wild-type SERPINA1 gene sequence. In some embodiments, the tagRNA results in correction of the mutation when contacted with the SERPINA1 gene.

[0214] A scaffold sequence (also referred to herein as the gRNA core, gRNA scaffold, or gRNA backbone sequence) of a tagRNA may contain a polynucleotide sequence that binds to a DNA binding domain, a DNA endonuclease domain (e.g., Cas9) of a RT editor. The scaffold sequence may interact with a RT editor as described herein, for example, by association with a DNA binding domain, a DNA endonuclease domain, such as a DNA nickase of the RT editor.

[0215] One of skill in the art will recognize that different RT editors having different DNA binding domain and a DNA endonuclease domains from different DNA binding domain and a DNA endonuclease proteins may require different gRNA core sequences specific to the DNA binding domain and a DNA endonuclease protein. In some embodiments, the scaffold sequence is capable of binding to a Cas9-based RT editor. In some embodiments, the scaffold sequence is capable of binding to a Cpfl -based RT editor. In some embodiments, the scaffold sequence is capable of binding to a Casl2b-based RT editor.

[0216] A tagRNA may also comprise optional modifiers, e.g., 3’ end modifier region and / or an 5' end modifier region. In some embodiments, a tagRNA comprises at least one nucleotide that is not part of a spacer, a scaffold sequence, or an extension arm. The optional sequence modifiers can be positioned within or between any of the other regions of the tagRNA,and not limited to being located at the 3' and 5' ends. In some embodiments, the tagRNA comprises secondary RNA structure, such as, but not limited to, aptamers, hairpins, stem / loops, toeloops, and / or RNA-binding protein recruitment domains (e.g., the MS2 aptamer which recruits and binds to the MS2cp protein). In some embodiments, a tagRNA comprises a short stretch of uracil at the 5’ end or the 3’ end. For example, in some embodiments, a tagRNA comprising a 3’ extension arm comprises a “UUU” sequence at the 3’ end of the extension arm. In some embodiments, a tagRNA comprises a toeloop sequence at the 3 ’ end. In some embodiments, the tagRNA comprises a 3’ extension arm and a toeloop sequence at the 3’ end of the extension arm. In some embodiments, the tagRNA comprises a 5’ extension arm and a toeloop sequence at the 5’ end of the extension arm. In some embodiments, the tagRNA comprises a toeloop element having the sequence S’-GAAANNNNN-3’, wherein N is any nucleobase. In some embodiments, the secondary RNA structure is positioned within the spacer. In some embodiments, the secondary structure is positioned within the extension arm. in some embodiments, the secondary structure is positioned within the gRNA core. In some embodiments, the secondary structure is positioned between the spacer and the gRNA core, between the gRNA core and the extension arm, or between the spacer and the extension arm. In some embodiments, the secondary structure is positioned between the FB sequence and the editing template. In some embodiments the secondary structure is positioned at the 3’ end or at the 5’ end of the tagRNA. in some embodiments, the tagRNA comprises a transcriptional lamination signal at the 3' end of the tagRNA. In addition to secondary RNA structures, the tagRNA may comprise a chemical linker or a poly(N) linker or tail, where “N” can be any nucleobase. In some embodiments, the chemical linker may function to prevent reverse transcription of the gRNA core.

[0217] In some embodiments, appending one or more RNA structural motifs to a tagRNA can protect against degradation of the tagRNA. Such RNA structural motifs can include, but are not limited to (i) a prequeosinel-1 riboswitch aptamer (evopreQl) and variants thereof, (ii) a frameshifting pseudoknot from Moloney murine leukemia virus (MMLV), hereafter referred to as “mpknot,” and variants thereof (iii) G-quadruplexes, (iv) hairpin structures (e.g., 15-bp hairpins), (v) xrRNA, and (vi) a P4-P6 domain of the group I intron.

[0218] In one embodiment, the modified tagRNAs include a nucleic acid moiety at the 3' end of the tagRNA. Optionally, the 3' end of the tagRNA is fused to the nucleic acid moiety through a nucleotide linker. In various embodiments, it will be appreciated that a wide variety of nucleotide sequences will work reasonably well for each genomic target site. Linker length can also be variable. In some cases, linkers ranging in length from 3-18 nucleotides can be used. In other cases, the linker may be at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, or at least 30 nucleotides.

[0219] In general, the nucleic acid moieties that may be used to modify a tagRNA, for example, by attaching it to the 3' end of a tagRNA, may include any nucleic acid moiety, including, for instance, a nucleic acid molecule comprising or which forms a double-helix moiety, toeloop moiety, hairpin moiety, stem-loop moiety, pseudoknot moiety, aptamer moiety, G quadraplex moiety, tRNA moiety, or a ribozyme moiety. The nucleic acid moiety may be characterized as forming a secondary nucleic acid structure, a tertiary nucleic acid structure, or a quadruple nucleic acid structure. In other words, the nucleic acid moiety may form any two dimensional or three dimensional structure known to be formed by such structures. The nucleic acid moiety may be DNA or RNA.

[0220] In some embodiments, a tagRNA or a nick guide RNA (egRNA) can be chemically synthesized, or can be assembled or cloned and transcribed from a DNA sequence, e.g., a plasmid DNA sequence, or by any RNA oligonucleotide synthesis method known in the art. In some embodiments, DNA sequence that encodes a tagRNA (or egRNA) can be designed to append one or more nucleotides at the 5' end or the 3' end of the tagRNA (or nick guide RNA) encoding sequence to enhance tagRNA transcription. For example, in some embodiments, a DNA sequence that encodes a tagRNA (or an egRNA) can be designed to append a nucleotide G at the 5' end. Accordingly, in some embodiments, the tagRNA (or nick guide RNA) can comprise an appended nucleotide G at the 5' end. In some embodiments, a DNA sequence that encodes a tagRNA (or nick guide RNA) can be designed to append a sequence that enhances transcription, e.g., a Kozak sequence, at the 5' end. In some embodiments, a DNA sequence that encodes a tagRNA (or nick guide RNA) can be designed to append the sequence CACC or CCACC at the 5' end. Accordingly, in some embodiments, the tagRNA (or nick guide RNA) can comprise an appended sequence CACC or CCACC at the 5' end. in some embodiments, a DNA sequence that encodes a tagRNA (or nick guide RNA) can be designed to append the sequence TTT, TTTT, TTTTT, TTTTTT, TTTTTTT at the 3'enc[ Accordingly, in some embodiments, the tagRNA (or nick guide RNA) can comprise an appended sequence UUU, UUUU, UUUUU, UUUUUU, or UUUUUUU at the 3' end.

[0221] Disclosed herein include RT editing systems. In some embodiments, the RT editing system comprises: any of the tagRNAs disclosed herein, or a nucleic acid encoding thetagRNA; and a RT editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, or a nucleic acid encoding the RT editor. In some embodiments, the intended nucleotide edit incorporation rate of the RT editing system is greater than at least about 30%, about 40%, about 50%, about 60%, about 70%, or about 80% (e.g., about 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values).

[0222] Disclosed herein include RT editing complexes. In some embodiments, the RT editing complex comprises: (i) any of the tagRNA disclosed herein and a RT editor comprising a DNA endonuclease domain and a DNA polymerase domain; or (ii) any of the RT editing systems of the disclosure. In some embodiments, the intended nucleotide edit incorporation rate of the RT editing complex is greater than at least about 30%, about 40%, about 50%, about 60%, about 70%, or about 80% (e.g., about 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values).Enhancer gRNA (egRNA)

[0223] Disclosed herein include RT editing systems. In some embodiments, the RT editing system comprises: any of the tagRNAs disclosed herein, or a nucleic acid encoding the tagRNA; and a RT editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, or a nucleic acid encoding the RT editor. The RT editing system can comprise: an enhancer guide RNA (egRNA), or a nucleic acid encoding the egRNA, wherein the egRNA comprises a egRNA spacer that is complementary to a second search target sequence in the target gene.

[0224] In some embodiments, a RT editing system or composition further comprises an enhancer guide polynucleotide, such as an enhancer guide RNA (egRNA). Without wishing to be bound by any particular theory, the non-edit strand of a double stranded target DNA in the target gene may be nicked by a CRISPR-Cas nickase directed by an egRNA. In some embodiments, the nick on the non-edit strand directs endogenous DNA repair machinery to use the edit strand as a template for repair of the non-edit strand, which may increase efficiency of RT editing. In some embodiments, the non-edit strand is nicked by a RT editor localized to the non-edit strand by the egRNA. Accordingly, also provided herein are tagRNA systems comprising at least one tagRNA and at least one egRNA.

[0225] In some embodiments, the egRNA is a guide RNA which contains a variable spacer sequence and a guide RNA scaffold or core region that interacts with the DNA binding domain, a DNA endonuclease domain, e.g., Cas9 of the RT editor, in some embodiments, the egRNA comprises a spacer sequence (referred to herein as an ng spacer, or a second spacer) that is substantially complementary to a second search target sequence (or ng search target sequence), which is located on the edit strand, or the non-target strand. Thus, in some embodiments, the egRNA search target sequence recognized by the egRNA spacer and the search target sequence recognized by the spacer sequence of the tagRNA are on opposite strands of the double stranded target DNA of target gene, e.g., the BCL11A gene, PAH gene, or SERPINA1 gene. In some embodiments, an egRNA spacer sequence is complementary to, and may hybridize with the second search target sequence only after an intended nucleotide edit has been incorporated on the edit strand, by the editing template of a tagRNA.

[0226] In some embodiments, the egRNA search target sequence is located on the nontarget strand, within 10 base pairs to 100 base pairs of an intended nucleotide edit incorporated by the tagRNA on the edit strand, in some embodiments, the egRNA target search target sequence is within 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 91 bp, 92 bp, 93 bp, 94 bp, 95 bp, 96 bp, 97 bp, 98 bp, 99 bp, or 100 bp of an intended nucleotide edit incorporated by the tagRNA on the edit strand. In some embodiments, the 5’ ends of the egRNA search target sequence and the tagRNA search target sequence are within 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bp apart from each other. In some embodiments, the 5’ ends of the egRNA search target sequence and the tagRNA search target sequence are within 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 91 bp, 92 bp, 93 bp, 94 bp, 95 bp, 96 bp, 97 bp, 98 bp, 99 bp, or 100 bp apart from each other.

[0227] In some embodiments, an egRNA spacer sequence is complementary to, and may hybridize with the second search target sequence only after an intended nucleotide edit has been incorporated on the edit strand, by the editing template of a tagRNA. In some embodiments, the egRNA comprises a spacer sequence that matches only the edit strand after incorporation of the nucleotide edits, but not the endogenous target gene sequence on the edit strand. Accordingly, in some embodiments, an intended nucleotide edit is incorporated within the egRNA search target sequence. In some embodiments, the intended nucleotide edit is incorporated within about 1-10 nucleotides of the position corresponding to the PAM of the egRNA search target sequence.

[0228] The RT editing system can comprise: a nick guide RNA (egRNA), or a nucleic acid encoding the egRNA, wherein the egRNA comprises an egRNA spacer that is complementaryto a second search target sequence in the target gene. The second search target sequence can be on the second strand of the target gene. The egRNA spacer can be from 16 to 22 nucleotides in length. In some embodiments, the length of the spacer varies from about 10 to about 100 nucleotides. In some embodiments, the spacer is 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, 20 nucleotides, 21 nucleotides, 22 nucleotides, 23 nucleotides, 24 nucleotides, or 25 nucleotides in length. In some embodiments, the spacer is from 15 nucleotides to 30 nucleotides in length, 15 to 25 nucleotides in length, 18 to 22 nucleotides in length, 10 to 20 nucleotides in length, or 20 to 30 nucleotides in length. In some embodiments, the spacer is 16 to 22 nucleotides in length. In some embodiments, the spacer is 16 to 20 nucleotides in length. In some embodiments, the spacer is 17 to 18 nucleotides in length. In some embodiments, the egRNA spacer is 20 nucleotides in length.Nucleic acids

[0229] In some embodiments, the gRNAs or tagRNAs described herein can be produced by in vitro transcription (IVT), synthetic and / or chemical synthesis methods, or a combination thereof. One or more of enzymatic IVT, solid-phase, liquid-phase, combined synthetic methods, small region synthesis, and ligation methods can be utilized. In some embodiments, the gRNAs or tagRNAs are made using IVT enzymatic synthesis methods. Methods of making polynucleotides by IVT are known in the art and are described in WO2013 / 151666. Polynucleotides constructs and vectors can be used to in vitro transcribe a gRNA or tagRNA described herein.

[0230] In some embodiments, a nucleic acid encoding a RT editor is administered to the subject. In some embodiments, the nucleic acid can be generated by an in vitro transcription reaction. In some embodiments, generating in vitro transcribed RNA comprises incubating a linear DNA template with an RNA polymerase and a nucleotide mixture under conditions to allow (runoff) RNA in vitro transcription. The nucleotide mixture can be part of an in vitro transcription mix (IVT-mix). In some embodiments, the RNA polymerase is a T7 RNA polymerase.

[0231] The nucleotide mixture used in RNA in vitro transcription can additionally contain modified nucleotides as defined below. In some embodiments, the nucleotide mixture (e.g., the fraction of each nucleotide in the mixture) used for RNA in vitro transcription reactions can be optimized for the given RNA sequence (optimized NTP mix). Such methods are described, for example in WO2015 / 188933. RNA obtained by a process using an optimized NTP mix is, in some embodiments, characterized by reduced immune stimulatory properties.

[0232] In various embodiments the nucleotide mixture in an in vitro transcription reaction comprises a cap analog. Accordingly, in some embodiments the cap analog is a capO, capl, cap2, a modified capO or a modified capl analog, or a capl analog as described below.

[0233] The term “cap analog” or “5 ’-cap structure” as used herein can refer to the 5’ structure of the RNA, particularly a guanine nucleotide, positioned at the 5 ’-end of an RNA, e.g., an mRNA. In some embodiments, the 5’-cap structure is connected via a 5 ’-5 ’-triphosphate linkage to the RNA. In some embodiments, a “5’-cap structure” or a “cap analogue” is not considered to be a “modified nucleotide” or “chemically modified nucleotides”. 5 ’-cap structures which may be suitable include capO (methylation of the first nucleobase, e.g., m7GpppN), capl (additional methylation of the ribose of the adjacent nucleotide of m7GpppN), cap2 (additional methylation of the ribose of the 2nd nucleotide downstream of the m7GpppN), cap3 (additional methylation of the ribose of the 3rd nucleotide downstream of the m7GpppN), cap4 (additional methylation of the ribose of the 4th nucleotide downstream of the m7GpppN), ARCA (anti-reverse cap analogue), modARCA (e.g., phosphothioate modARCA), inosine, Nl-methyl-guanosine, 2’-fluoro-guanosine, 7-deaza-guanosine, 8-oxo-guanosine, 2-amino-guanosine, LNA-guanosine, and 2-azido-guanosine.

[0234] A 5’-cap (capO or capl) structure can be formed in chemical RNA synthesis, using capping enzymes, or in RNA in vitro transcription (co-transcriptional capping) using cap analogs. The term “cap analog” as used herein can refer to a non-polymerizable di-nucleotide or tri -nucleotide that has cap functionality in that it facilitates translation or localization, and / or prevents degradation of the RNA when incorporated at the 5 ’ -end of the RNA. Non-polymerizable means that the cap analogue will be incorporated only at the 5 ’-terminus because it does not have a 5’ triphosphate and therefore cannot be extended in the 3 ’-direction by a template-dependent polymerase, (e.g., a DNA-dependent RNA polymerase). Examples of cap analogues include m7GpppG, m7GpppA, m7GpppC; unmethylated cap analogues (e.g., GpppG); dimethylated cap analogue (e.g., m2,7GpppG), trimethylated cap analogue (e.g. m2,2,7GpppG), dimethylated symmetrical cap analogues (e.g. m7Gpppm7G), or anti reverse cap analogues (e.g., ARCA; m7,2’OmeGpppG, m7,2’dGpppG, m7,3’OmeGpppG, m7,3’dGpppG and their tetraphosphate derivatives). Further cap analogues have been described previously, e.g, W02008 / 016473, W02008 / 157688, WO2009 / 149253, WO2011 / 015347, and WO2013 / 059475. Further suitable cap analogues in that context are described in, e.g, WO2017 / 066793, WO2017 / 066781, WO2017 / 066791, WO2017 / 066789, WO2017 / 053297, WO2017 / 066782, WO2018 / 075827 and WO2017 / 066797 wherein the disclosures relating to cap analogues are incorporated herewith by reference.

[0235] In some embodiments, a capl structure is generated using tri-nucleotide cap analogue as disclosed in WO2017 / 053297, WO2017 / 066793, WO2017 / 066781, WO2017 / 066791, WO2017 / 066789, WO2017 / 066782, WO2018 / 075827 and WO2017 / 066797. For example, any cap analog derivable from the structure disclosed in claim 1-5 ofWO2017 / 053297 may be suitably used to co-transcriptionally generate a capl structure. In some embodiments, any cap analog derivable from the structure described in WO2018 / 075827 can be suitably used to co-transcriptionally generate a capl structure. In some embodiments, the capl analog is a capl trinucleotide cap analog. In some embodiments, the capl structure of the in vitro transcribed RNA is formed using co-transcriptional capping using tri -nucleotide cap analog m7G(5’)ppp(5')(2’OMeA)pG or m7G(5’)ppp(5’)(2’OMeG)pG. In some embodiments, the capl analog is m7G(5’)ppp(5’)(2’OMeA)pG.

[0236] In some embodiments, the RNA (e.g., mRNA) comprises a 5 ’-cap structure, e.g., a capl structure. In some embodiments, the 5’ cap structure can improve stability and / or expression of the mRNA. A capl structure comprising mRNA (produced by, e.g., in vitro transcription) has several advantageous features including an increased translation efficiency and a reduced stimulation of the innate immune system. In some embodiments, the in vitro transcribed RNA comprises at least one coding sequence encoding at least one peptide or protein. In some embodiments, the protein is an RNA-guided endonuclease. In some embodiments, the RNA-guided endonuclease is Cas9 or a derivative thereof.

[0237] In some embodiments, the mRNA can comprise at least one chemically modified nucleoside and / or nucleotide. In some embodiments, the chemically modified nucleoside and / or nucleotide is selected from pseudouridine, N1 -methylpseudouridine, and 5-methoxyuridine. In some embodiments, the chemically modified nucleoside is Nl-methylpseudouridine (e.g., 1 -methylpseudouridine). In some embodiments, at least about 80% or more (e.g., about 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%) of uridines in the mRNA are modified or replaced with N1 -methylpseudouridine. In some embodiments, 100% of the uridines (e.g. , uracils) in the mRNA are modified or replaced with N 1 -methylpseudouridine.Pharmaceutical Compositions and Therapeutic Applications

[0238] Provided herein include pharmaceutical compositions for carrying out the methods disclosed herein and related methods of using the systems, RNPs, LNP, vectors, polynucleotides, and cells described herein to prevent or treat a disease or disorder.

[0239] Disclosed compositions can, for example, comprise a nucleic acid encoding the editor, base editor, and / or RT editor, and one or more guide RNAs or tagRNAs targeting one or more target genes. In some embodiments, a pharmaceutical composition is provided, the pharmaceutical composition comprising the composition and a pharmaceutical acceptable carrier or excipient.

[0240] In some embodiments, a composition described above can further have one or more additional reagents, where such additional reagents are selected from a buffer, a buffer for introducing a polypeptide or polynucleotide into a cell, a wash buffer, a control reagent, a controlvector, a control RNA polynucleotide, a reagent for in vitro production of the polypeptide from DNA, adaptors for sequencing and the like. A buffer can be a stabilization buffer, a reconstituting buffer, a diluting buffer, or the like. In some embodiments, a composition can also include one or more components that can be used to facilitate or enhance the on-target binding or the cleavage of DNA by the endonuclease, or improve the specificity of targeting.

[0241] In some embodiments, any components of a composition are formulated with pharmaceutically acceptable excipients such as carriers, solvents, stabilizers, adjuvants, diluents, etc., depending upon the particular mode of administration and dosage form. In embodiments, guide RNA compositions are generally formulated to achieve a physiologically compatible pH, and range from a pH of about 3 to a pH of about 11, about pH 3 to about pH 7, depending on the formulation and route of administration. In some embodiments, the pH is adjusted to a range from about pH 5.0 to about pH 8.

[0242] Suitable excipients can include, for example, carrier molecules that include large, slowly metabolized macromolecules such as proteins, polysaccharides, polylactic acids, polyglycolic acids, polymeric amino acids, amino acid copolymers, and inactive virus particles. Other exemplary excipients include antioxidants (for example and without limitation, ascorbic acid), chelating agents (for example and without limitation, EDTA), carbohydrates (for example and without limitation, dextrin, hydroxyalkylcellulose, and hydroxyalkylmethylcellulose), stearic acid, liquids (for example and without limitation, oils, water, saline, glycerol and ethanol), wetting or emulsifying agents, pH buffering substances, and the like.

[0243] Physiologically tolerable carriers are well known in the art. Exemplary liquid carriers are sterile aqueous solutions that contain no materials in addition to the active ingredients and water or contain a buffer such as sodium phosphate at physiological pH value, physiological saline or both, such as phosphate-buffered saline. Aqueous carriers can contain more than one buffer salt, as well as salts such as sodium and potassium chlorides, dextrose, polyethylene glycol and other solutes. Liquid compositions can also contain liquid phases in addition to and to the exclusion of water. Exemplary of such additional liquid phases are glycerin, vegetable oils such as cottonseed oil, and water-oil emulsions. The amount of an active compound used in the cell compositions that is effective in the treatment of a particular disorder or condition will depend on the nature of the disorder or condition and can be determined by standard clinical techniques.

[0244] In some embodiments, the components of a disclosed composition can be delivered via transfection such as calcium phosphate transfection, DEAE-dextran mediated transfection, cationic lipid-mediated transfection, electroporation, electrical nuclear transport, chemical transduction, electrotransduction, Lipofectamine-mediated transfection, Effectene-mediated transfection, lipid nanoparticle (LNP)-mediated transfection, or any combinationthereof. In some embodiments, the composition is introduced to the cells via lipid-mediated transfection using a lipid nanoparticle.

[0245] In some embodiments, the components of a disclosed composition can be formulated in a liposome or lipid nanoparticle. In some embodiments, the compounds of the composition are formulated in a lipid nanoparticle (LNP). LNP is a non-viral delivery system that safely and effectively deliver nucleic acids to target organs (e.g., liver). The term “lipid nanoparticle” refers to a nanoscopic particle composed of lipids having a size measured in nanometers (e.g., 1-5,000 nm). In some embodiments, the lipids comprised in the lipid nanoparticles comprise cationic lipids and / or ionizable lipids. Any suitable cationic lipids and / or ionizable lipids known in the art can be used to formulate LNPs for delivery of gRNA and Cas endonuclease to the cells. Exemplary cationic lipids include one or more amine group(s) bearing positive charge. In some embodiments, the cationic lipids are ionizable such that they can exist in a positively charged or neutral from depending on pH. In some embodiments, the cationic lipid of the lipid nanoparticle comprises a protonatable tertiary amine head group that shows positive charge at low pH. The lipid nanoparticles can further comprise one or more neutral lipids (e.g., Distearoylphosphatidylcholine (DSPC), 1, 2-Dioleoyl-sn-glycero-3 -phosphocholine (DOPC), 1,2-Dimyristoyl-sn-glycero-3-phosphoethanolamine (DMPE), l,2-Dipalmitoyl-sn-glycero-3-phosphorylethanolamine (DPPE) etc. as a helper lipid), charged lipids, steroids, and polymers conjugated lipids.

[0246] The lipid nanoparticles can have a mean diameter of about, at least, at least about, at most or at most about 30 nm, 40 nm, 50 nm, 60 nm, 70 nm, 80 nm, 90 nm, 100 nm, 110 nm, 120 nm, 130 nm, 140 nm, 150 nm, or a number or a range between any of these values. In some embodiments, the lipid nanoparticle particle size is about 50 to about 100 nm in diameter, or about 70 to about 90 nm in diameter, or about 55 to about 95 nm in diameter.

[0247] The compounds of the composition described herein can be encapsulated in the lipid portion of the lipid nanoparticle or an aqueous space enveloped by some or all of the lipid portion of the lipid nanoparticle. The encapsulation can be full encapsulation, partial encapsulation, or both. In some embodiments, the nucleic acid and / or polypeptides are fully or substantially encapsulated (e.g., greater than 90% of the RNA) in the lipid nanoparticle.

[0248] In some embodiments, one or more compounds herein described are associated with a liposome or lipid nanoparticle via a covalent bond or non-covalent bond. In some embodiments, any of the compounds in the composition can be separately or together contained in a liposome or lipid nanoparticle.

[0249] The compositions and / or pharmaceutical compositions provided herein described can be administered to a subject in need thereof to prevent or treat a disease or disorder.Accordingly, the present disclosure also provides a method of preventing or treating a disease or disorder in a subject in need thereof. The method can comprise administering to the subject a therapeutically effective amount of the composition or pharmaceutical composition described herein, wherein the one or more target gene is associated with the disease or disorder.

[0250] A subject can be any subject for whom diagnosis, treatment, or therapy is desired. In some embodiments, the subject is a mammal. In some embodiments, the subject is a human. In some embodiments, the subject is suspected to have, having or diagnosed to have a disease or disorder.

[0251] Any suitable administration route capable of delivering the composition can be used herein. In some embodiments, the pharmaceutical composition thereof can be administered by aerosol delivery, nasal delivery, vaginal delivery, rectal delivery, buccal delivery, ocular delivery, local delivery, topical delivery, intracistemal delivery, intraperitoneal delivery, oral delivery, intramuscular injection, intravenous injection, subcutaneous injection, intranodal injection, intratumoral injection, intracardiac injection, intraperitoneal injection, intrathecal injection, intraventricular injection, intracerebroventricular injection, intradermal injection, or any combination thereof. The administration can be local or systemic. Systemic administration refers to the administration of a population of cells other than directly into a target site, tissue, or organ, such that it enters the subject's circulatory system and, thus, is subject to metabolism and other like processes. Systemic administration can include enteral and parenteral administration. In some embodiments, more than one administration can be employed to achieve the desired level of gene expression over a period of various intervals, e.g., daily, weekly, monthly, or yearly. In some embodiments, the route is intravenous. The pharmaceutical composition thereof can be administered to a subject in need thereof at a pharmaceutically effective amount. The amount of the pharmaceutical composition can result in a desired reduction or loss of function in one or more gene products.

[0252] In some embodiments, the method can comprise administering to a subject an engineered cell herein described or a population of the engineered cells. The step of administering can include introducing e.g., transplantation) the cells, e.g., an engineered cell or a population thereof described herein, into a subject, by a method or route that results in at least partial localization of the introduced cells at a desired site, such as tumor, such that a desired effect(s) is produced. Engineered cells can be administered by any appropriate route that results in delivery to a desired location in the subject where at least a portion of the implanted cells or components of the cells remain viable. The period of viability of the cells after administration to a subject can be as short as a few hours, e.g., twenty -four hours, to a few days, to as long as several years, or even the life time of the subject, i.e., long-term engraftment. For example, in some aspectsdescribed herein, an effective amount of engineered cells is administered via a systemic route of administration, such as an intraperitoneal or intravenous route.

[0253] In some embodiments, an engineered cell population being administered according to the methods described herein comprises allogeneic cells obtained from one or more donors. Allogeneic refers to a cell, cell population, or biological samples comprising cells, obtained from one or more different donors of the same species, where the genes at one or more loci are not identical to the recipient. For example, an engineered cell population, being administered to a subject can be derived from one or more unrelated donors, or from one or more non-identical siblings. In some embodiments, syngeneic cell populations can used, such as those obtained from genetically identical donors, (e.g., identical twins). In some embodiments, the cells are autologous cells; that is, the engineered cells are obtained or isolated from a subject and administered to the same subject, z.e., the donor and recipient are the same. A donor as used herein is an individual who is not the subj ect being treated. A donor is an individual who is not the patient. In some embodiments, a donor is an individual who does not have or is not suspected of having the cancer being treated. In some embodiments, multiple donors, e.g., two or more donors, are used.

[0254] In some embodiments, an engineered cell population being administered according to the methods described herein does not induce toxicity in the subject, e.g., the engineered cells do not induce toxicity in non-cancer cells. In some embodiments, an engineered cell population being administered does not trigger complement mediated lysis, or does not stimulate antibody-dependent cell mediated cytotoxicity (ADCC).

[0255] An effective amount refers to the amount of a population of engineered cells needed to prevent or alleviate at least one or more signs or symptoms of a medical condition, and relates to a sufficient amount of a composition to provide the desired effect, e.g, to treat a subject having a medical condition.

[0256] In some embodiments, an effective amount of cells (e.g, engineered cells) comprises about, at least or at least about 102cells, 5 * 102cells, 103cells, 5 * 103cells, 104cells, 5 x 104cells, 105cells, 2 x io5cells, 3 x io5cells, 4 x io5cells, 5 x io5cells, 6 x io5cells, 7 x 105cells, 8 x io5cells, 9 x io5cells, 1 x io6cells, 2 x io6cells, 3 x io6cells, 4 x io6cells, 5 x 106cells, 6 x io6cells, 7 x io6cells, 8 x io6cells, 9 x io6cells, 10 x io6cells, 12 x io6cells, 14 x 106cells, 16 x io6cells, 18 x io6cells, 20 x io6cells, 25 x io6cells, 30 x io6cells, or a number between any two of the values. The cells are derived from one or more donors, or are obtained from an autologous source. In some embodiments described herein, the cells are expanded in culture prior to administration to a subject in need thereof.

[0257] In some embodiments, the disease or disorder is cancer. Non-limiting examplesof cancers that can be treated as provided herein include: breast cancer, e.g., estrogen receptorpositive breast cancer, prostate cancer, squamous tumors, e.g., of the skin, bladder, lung, cervix, endometrium, head neck, and biliary tract, and neuronal tumors. The compositions and methods provided herein can be used to treat various types of cancer, including but are not limited to, melanoma (e.g., metastatic malignant melanoma), renal cancer (e.g., clear cell carcinoma), prostate cancer (e.g., hormone refractory prostate adenocarcinoma), pancreatic adenocarcinoma, breast cancer, colon cancer, lung cancer (e.g., non-small cell lung cancer (NSCLC) and small-cell lung cancer (SCLC)), esophageal cancer, squamous cell carcinoma of the head and neck, liver cancer, ovarian cancer, cervical cancer, thyroid cancer, glioblastoma, glioma, leukemia, lymphoma, and other neoplastic malignancies. Additionally, the disease or condition provided herein includes refractory or recurrent malignancies whose growth may be inhibited using the methods and compositions disclosed herein. In some embodiments, the cancer is carcinoma, squamous carcinoma, adenocarcinoma, sarcomata, endometrial cancer, breast cancer, ovarian cancer, cervical cancer, fallopian tube cancer, primary peritoneal cancer, colon cancer, colorectal cancer, squamous cell carcinoma of the anogenital region, melanoma, renal cell carcinoma, lung cancer, non-small cell lung cancer, squamous cell carcinoma of the lung, stomach cancer, bladder cancer, gall bladder cancer, liver cancer, thyroid cancer, laryngeal cancer, salivary gland cancer, esophageal cancer, head and neck cancer, glioblastoma, glioma, squamous cell carcinoma of the head and neck, prostate cancer, pancreatic cancer, mesothelioma, sarcoma, hematological cancer, leukemia, lymphoma, neuroma, or a combination thereof. In some embodiments, the cancer is carcinoma, squamous carcinoma (e.g., cervical canal, eyelid, tunica conjunctiva, vagina, lung, oral cavity, skin, urinary bladder, tongue, larynx, and gullet), and adenocarcinoma (for example, prostate, small intestine, endometrium, cervical canal, large intestine, lung, pancreas, gullet, rectum, uterus, stomach, mammary gland, and ovary). In some embodiments, the cancer is sarcomata (e.g., myogenic sarcoma), leukosis, neuroma, melanoma, and lymphoma.

[0258] The cancer can include pancreatic cancer, gastric cancer, ovarian cancer, uterine cancer, breast cancer, prostate cancer, testicular cancer, thyroid cancer, nasopharyngeal cancer, non-small cell lung (NSCLC), glioblastoma, neuronal, soft tissue sarcomas, leukemia, lymphoma, melanoma, colon cancer, colon adenocarcinoma, brain glioblastoma, hepatocellular carcinoma, liver hepatocholangiocarcinoma, osteosarcoma, gastric cancer, esophagus squamous cell carcinoma, advanced stage pancreas cancer, lung adenocarcinoma, lung squamous cell carcinoma, lung small cell cancer, renal carcinoma, intrahepatic biliary cancer, and a combination thereof. In some embodiments, the cancer is breast cancer, prostate cancer, squamous tumor cancer, neuronal tumor cancer, or a combination thereof. In some embodiments, the cancer comprises cancer cells expressing LIV1.

[0259] The cancer can be a solid tumor, a liquid tumor, or a combination thereof. In some embodiments, the cancer is a solid tumor, including but are not limited to, melanoma, renal cell carcinoma, lung cancer, bladder cancer, breast cancer, cervical cancer, colon cancer, gall bladder cancer, laryngeal cancer, liver cancer, thyroid cancer, stomach cancer, salivary gland cancer, prostate cancer, pancreatic cancer, Merkel cell carcinoma, brain and central nervous system cancers, and any combination thereof. In some embodiments, the cancer is a liquid tumor. In some embodiments, the cancer is a hematological cancer. Non-limiting examples of hematological cancer include diffuse large B cell lymphoma (“DLBCL”), Hodgkin's lymphoma (“HL”), Non-Hodgkin's lymphoma (“NHL”), Follicular lymphoma (“FL”), acute myeloid leukemia (“AML”), and multiple myeloma (“MM”).

[0260] In some embodiments, the compositions and methods provided herein can be used to treat arteriosclerosis, atherosclerosis, cardiovascular diseases, coronary heart disease, diabetes, diabetes mellitus, non-insulin-dependent diabetes mellitus, fatty liver, hyperinsulinism, hyperlipidemia, hypertriglyceridemia, hypobetalipoproteinemias, inflammation, insulin resistance, metabolic diseases, obesity, malignant neoplasm of mouth, lipid metabolism disorders, lip and oral cavity carcinoma, dyslipidemias, metabolic syndrome x, hypotriglyceridemia, opitz trigonocephaly syndrome, ischemic stroke, hypertriglyceridemia result, hypobetalipoproteinemia familial 2, familial hypobetalipoproteinemia, and ischemic cerebrovascular accident.

[0261] In some embodiments, the compositions and methods provided herein can be used to treat a metabolic disease, a lipid metabolism disease, obesity, atherosclerosis, hyperfattyacidemia, metabolic syndrome, dyslipidemia, hypobetalipoproteinemia, familial hypercholesterolemia (including homozygous familial hypercholesterolemia (HoFH) and heterozygous familial hypercholesterolemia (HeFH)), hypertriglyceridemia, familial combined hyperlipidemia, familial chylomicronemia syndrome, multifactorial chylomicronemia syndrome, familial combined hyperlipidemia (FCHL), metabolic syndrome (MetS), nonalcoholic fatty liver disease (NAFLD), elevated lipoprotein (a), elevated lipids such as total cholesterol, triglycerides, LDLs, HDLs, and / or other non-HDLs in the blood, or a combination thereof. NAFLD can be hepatic steatosis or steatohepatitis. The diabetes can be type 2 diabetes or type 2 diabetes with dyslipidemia. Dyslipidemia can be hyperlipidemia, for example hypercholesterolemia, hypertriglyceridemia, or both.

[0262] In some embodiments, the methods and compositions herein described can be used to treat a subject with a blood disorder such as sickle cell disease. Sickle cell disease refers to a group of inherited red blood cell disorders when a child receives two sickle cell genes, one from each parent. With sickle cell disease, red blood cells contort into a sickle shape. The cells die early, leaving a shortage of healthy red blood cells, and can block blood flow causing pain.Symptoms of sickle cell disease include infections, pain, and fatigue. Sickle cell disease can be diagnosed with a blood test. Common types of sickle cell disease include HbSS also referred to as sickle cell anemia, HbSC, HbS beta thalassemia, HbSD, HbSE and HbSO. Patients with HbSS inherit two sickle cell genes (“S”), one from each parent. Patients with HbSC inherit a sickle cell gene (“S”) from one parent and from the other parent a gene for an abnormal hemoglobin called “C”. Patients with HbS beta thalassemia inherit one sickle cell gene (“S”) from one patent and one gene of beta thalassemia from the other parent. There are two types of beta thalassemia: “0” and Those with HbS beta 0-thalassemia usually have a severe form of sickle cell disease. Those with HbS beta +-thalassemia tend to have a milder form of sickle cell disease. Patients with HbSD, HbSE or HbSO inherit one sickle cell gene (“S”) and one gene from an abnormal type of hemoglobin (“D”, “E”, or “O”). In some embodiments, the subject in need is a subject having sickle cell trait. A subject having sickle cell trait inherit one sickle cell gene (“S”) from one parent and one normal gene from the other parent. Subjects with sickle cell trait typically do not have any of the signs of the disease, but can pass the trait on to their children.

[0263] Provided herein include methods for the treatment of diseases or disorders, e.g., diseases or disorders that are associated or caused by a point mutation(s) that can be corrected by gene editing. Some such diseases are described herein, and additional suitable diseases that can be treated with the compositions and methods provided herein will be apparent to those of skill in the art based on the instant disclosure. It will be understood that the numbering of the specific positions or residues in the respective sequences depends on the particular protein and numbering scheme used. Numbering might be different, e.g., in precursors of a mature protein and the mature protein itself, and differences in sequences from species to species may affect numbering. One of skill in the art will be able to identify the respective residue in any homologous protein and in the respective encoding nucleic acid by methods well known in the art, e.g., by sequence alignment and determination of homologous residues.Kits

[0264] Provided herein also includes kits for use in producing the editors, base editors, and / or RT editors alone or in complex with a gRNA or tagRNA, the compositions, and the engineered cells and carrying out the methods described herein for therapeutic uses.

[0265] In some embodiments, a kit provide herein comprises components for performing editing, base editing, and / or RT editing of one or more target sequences. A kit can comprise an editor, base editor, and / or RT editor herein described or a nucleic acid sequence encoding the editor, base editor, and / or RT editor, and one or more guide RNAs or tagRNAs targeting the one or more target sequences. A kit can also comprises a nucleic acid construct comprising (a) a nucleotide sequence encoding a editor, base editor, and / or RT editor as providedherein; and (b) a heterologous promoter that drives expression of the sequence of (a). A kit can comprise a complex of the editor, base editor, and / or RT editor bound with a gRNA or tagRNA. A kit can also comprise cells comprising a deaminase protein, a fusion protein, a nucleic acid molecule encoding the fusion protein, a complex comprising the fusion protein bound with a gRNA, and / or a vector as provided herein. Components of a kit can be in separate containers, or combined in a single container.

[0266] Any kit described above can further comprise one or more additional reagents selected from a buffer, a buffer for introducing a polypeptide or polynucleotide into a cell, a wash buffer, a control reagent, a control vector, a control RNA polynucleotide, a reagent for in vitro production of the polypeptide from DNA, adaptors for sequencing and the like. A buffer can be a stabilization buffer, a reconstituting buffer, a diluting buffer, or the like. A kit can also comprise one or more components that can be used to facilitate or enhance the on-target binding or the cleavage of DNA by the endonuclease, or improve the specificity of targeting.

[0267] In some embodiments, a kit can further include instructions for using the components of the kit to practice the methods described herein. The instructions for practicing the methods are generally recorded on a suitable recording medium. For example, the instructions can be printed on a substrate, such as paper or plastic, etc. The instructions can be present in the kits as a package insert, in the labeling of the container of the kit or components thereof (i.e., associated with the packaging or subpackaging), etc. The instructions can be present as an electronic storage data file present on a suitable computer readable storage medium, e.g., CD-ROM, diskette, flash drive, etc. In some instances, the actual instructions are not present in the kit, but means for obtaining the instructions from a remote source (e.g., via the Internet), can be provided. An example of this embodiment is a kit that includes a web address where the instructions can be viewed and / or from which the instructions can be downloaded. As with the instructions, this means for obtaining the instructions can be recorded on a suitable substrate.EXAMPLES

[0268] Some aspects of the embodiments discussed above are disclosed in further detail in the following examples, which are not in any way intended to limit the scope of the present disclosure.Example 1Evaluation of novel Cas9s and determination of their PAMs

[0269] Provided in this Example are methods and compositions related to evaluation of PAM of novel Cas9s in vitro.

[0270] Systematic metagenomic searches across 47 naturally occurring Cas9 sequences combined with structure-based analyses identified conserved clusters within the PAM-interacting domain (PID) and recognition (REC) lobe. Alignments of all Cas9s sequences are shown in Fig IB and 1C. From such alignment, several conserved clusters of mutations were identified. The resulting phylogenetic tree from these alignments is shown in FIG ID. These conserved clusters provided the foundation for designing novel Cas9 variants for comprehensive PAM evaluation For instance, FIG. 1 A shows one such cluster which is found in the REC lobe of ScrCas9 and SorCas9. This cluster was imported into SsaCas9_0 resulting in SsaCas9_AR00-12. REC lobe generally “senses” nucleic acids, regulates HNH conformation transitions into active and inactive states.

[0271] Twenty -three highly active Cas9 variants (out of 89 total designed variants) with AATD PAM or novel PAMs were selected from the first PAM depletion screen for PAM depletion assay optimization, followed by further evaluation in human cells (Table 3).TABLE 3: LIST OF CAS9 VARIANTS FOR PAM DEPLETION ASSAY

[0272] An in vitro PAM depletion assay was performed to evaluate the PAM recognition profiles of various Cas9 variants. First, Cas9 variants were cloned into plasmids containing a T7 promoter and ribosome binding site (RBS) upstream, and a T7 terminator downstream. Similarly, the gRNA was cloned into a plasmid with a T7 promoter. Cell-free transcription and translation (TXTL) enables simultaneous production of both Cas9 protein and guide RNA (gRNA) in a single reaction. For the TXTL reaction, linearized templates for Cas9 — comprising the T7 promoter, RBS, Cas9 coding sequence, and T7 terminator — and for gRNA with T7 promoters were used. The resulting Cas9 ribonucleoprotein (RNP) targets a PAM library withan 8N randomized sequence (NNNNNNNN, where N represents any nucleotide: A, T, G, or C) positioned immediately downstream of three protospacers. These protospacers correspond to sequences within the SERPINA1, PAH, and ATP7B genes. The depletion assay ran for 1 minute before quenching, followed by deep sequencing to assess the extent of PAM depletion. Computational analysis was performed to generate PAM patterns of each Cas9 variant.

[0273] FIGS. 2A-2AU depict PAM-related heatmaps and bar plots PAM depletion, of all 47 Cas9 variants, with a depletion scale ranging from 0 to 6. Darker colors indicate greater depletion at the corresponding PAMs specified by the X and Y axes. Higher depletion reflects higher Cas9 activity. The bar plot displays the overall PAM preference for each Cas9 variant. The X-axis contains 8 bases of the PAM sequence, with 4 different colors representing A, T, G, and C. A larger portion of a color within a bar indicates a stronger preference for that base at the corresponding position. Table 4 shows the consensus PAM motif of those Cas9s identified through this PAM ID assay. Furthermore, the PAM bar plots which are calculated using PAM depletion data, show preference of individual nucleotide at each PAM position for all the Cas9s tested. FIGS. 9A-9W, shows 2ndround of PAM depletion with a subset of 23 Cas9s under more sensitive condition of 1 min depletion. The resulting PAM depletion heat maps and PAM bar plots are shown.TABLE 4: PAM ID ASSAY RESULTSExample 2Validation of Novel Cas9s in Human Cells

[0274] Provided in this Example are methods and compositions related to editing of SERPINA1, PAH and BCL11A for validating novel Cas9s in human cells. To validate a subset of novel Cas9 variants in human cells, the top PAM sequences identified from the in vitro PAM depletion assay were used to design chemically synthesized gRNAs targeting three regions: SERPINA1, PAH, and BCL11 A (Table 5). FIG. 3 A-3B depict non-limiting exemplary schematics of the workflow of the cell-based PAM validation assay (FIG. 3A) and the Nuclease Cas9-RT construct employed (FIG. 3B). Cas9 nuclease was fused with MMLV-RT and the mRNA was prepared using in vitro transcription (IVT).TABLE 5: GUIDE SEQUENCES-Ill-

[0275] Disclosed herein include methods of generation of...

Claims

WHAT IS CLAIMED IS:

1. An engineered Cas9 protein comprising a substitution, insertion or deletion at one or more amino acid residues in the Recognition (REC) Lobe Domain and / or the PAM-Interacting Domain (PID) as compared to a parent Cas9 protein comprising the sequence selected from SEQ ID Nos: 728-747.

2. The engineered Cas9 protein of claim 1, wherein the substitution, insertion or deletion is at one or more amino acid residues in the Recognition (REC) Lobe Domain.

3. The engineered Cas9 protein of claim 2, wherein the substitution, insertion or deletion is at one or more of amino acid residues selected from amino acid positions 73-466 of SEQ ID Nos: 728-747.

4. The engineered Cas9 protein of claim 3, wherein the substitution, insertion or deletion is at one or more of amino acid residues selected from amino acid positions 218-269 of SEQ ID Nos: 728-747.

5. The engineered Cas9 protein of claim 4, wherein the substitution, insertion or deletion is at one or more of amino acid residues selected from amino acid positions 240-263 of SEQ ID Nos: 728-747.

6. The engineered Cas9 protein of claim 5, comprising an insertion at an amino acid position selected from amino acid positions 240-263 of SEQ ID Nos: 728-747, wherein the insertion is about 1-50 amino acids.

7. The engineered Cas9 protein of any one of claims 5-6, comprising a deletion at an amino acid position selected from amino acid positions 240-263 of SEQ ID Nos: 728-747, wherein the deletion is about 1-20 amino acids.

8. The engineered Cas9 protein of claim 1, wherein the substitution, insertion or deletion at one or more amino acid residues in the PID.

9. The engineered Cas9 protein of claim 8, wherein the substitution, insertion or deletion is at one or more of amino acid residues selected from the 185 most C-terminal amino acids of SEQ ID Nos: 728-747.

10. The engineered Cas9 protein of claim 8, wherein the substitution, insertion or deletion is at one or more of amino acid residues selected from amino acid positions 943-1093 of SEQ ID Nos: 728-747.

11. The engineered Cas9 protein of claim 10, comprising an insertion at an amino acid position selected from amino acid positions 943-1093 of SEQ ID Nos: 728-747, wherein the insertion is about 1-50 amino acids.

12. The engineered Cas9 protein of claim 10, comprising a deletion at an amino acid position selected from amino acid positions 943-1093 of SEQ ID Nos: 728-747, wherein the deletion is about 1-20 amino acids.

13. The engineered Cas9 protein of any one of claims 1-12, wherein the Cas9 protein is a nuclease.

14. The engineered Cas9 protein of any one of claims 1-12, wherein the Cas9 protein is a nickase.

15. A nuclease comprising:an amino acid sequence that is at least 80% identical to SEQ ID NO: 1, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 1, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 2, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 2, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 3, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 3, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 4, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 4, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNGNAANN, NNGAAAGN, NNGAAAAN, NNGAAAGN, NNGGAATN, and / or NNGGAAGN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 5, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 5, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 6, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 6, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN, NNAGAAAN, NNAGCAAN, and / or NNAGGAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 7, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ IDNO: 7, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 8, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 8, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 9, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 9, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 10, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 10, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNRNNNNN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 11, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 11, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNANGANN, NNAAGACN, NNAGGATN, NNAAGATN, NNAAGCAN, NNAAGAAN, NNAAGCGN, and / or NN A AGO CN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 12, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 12, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 13, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 13, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAYAAMN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 14, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 14, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 15, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 15, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 16, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 16, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 17, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 17, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 18, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 18, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 19, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 19, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNGAAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 20, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 20, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNGGNNCN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 21, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 21, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNRGAACN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 22, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 22, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNNNAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 23, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 23, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 24, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence ofSEQ ID NO: 24, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 25, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 25, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 26, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 26, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 27, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 27, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAYAACN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 28, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 28, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAYAACN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 29, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 29, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAYAACN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 30, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 30, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 31, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 31, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAAAGNN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 32, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 32, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNGGNNNN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 33, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 33, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNGGNNNN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 34, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 34, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNGGNNCN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 35, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 35, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAYAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 36, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 36, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAAAGNN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 37, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 37, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNNNAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 38, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 38, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNNNAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 39, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 39, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 40, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 40, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNGNNNNN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 41, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence ofSEQ ID NO: 41, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 42, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 42, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAAAGNN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 43, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 43, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAAANNN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 44, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 44, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAYAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 45, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 45, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAAAGNN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 46, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 46, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN; and / oran amino acid sequence that is at least 80% identical to SEQ ID NO: 47, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 47, further optionally said nuclease has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN.

16. A nickase comprising:an amino acid sequence that is at least 80% identical to SEQ ID NO: 48, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 48, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 49, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 49, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 50, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 50, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 51, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 51, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNGNAANN, NNGAAAGN, NNGAAAAN, NNGAAAGN, NNGGAATN, and / or NNGGAAGN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 52, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 52, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 53, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 53, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN, NNAGAAAN, NNAGCAAN, and / or NNAGGAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 54, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 54, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 55, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 55, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 56, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 56, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 57, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 57, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNRNNNNN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 58, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence ofSEQ ID NO: 58, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNANGANN, NNAAGACN, NNAGGATN, NNAAGATN, NNAAGCAN, NNAAGAAN, NNAAGCGN, and / or NNAAGCCN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 59, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 59, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 60, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 60, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAYAAMN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 61, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 61, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 62, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 62, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 63, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 63, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 64, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 64, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 65, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 65, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 66, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 66, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNGAAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 67, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 67, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNGGNNCN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 68, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 68, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNRGAACN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 69, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 69, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNNNAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 70, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 70, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 71, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 71, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 72, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 72, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 73, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 73, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 74, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 74, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAYAACN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 75, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence ofSEQ ID NO: 75, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAYAACN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 76, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 76, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAYAACN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 77, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 77, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 78, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 78, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAAAGNN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 79, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 79, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNGGNNNN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 80, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 80, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNGGNNNN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 81, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 81, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNGGNNCN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 82, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 82, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAYAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 83, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 83, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAAAGNN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 84, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 84, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNNNAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 85, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 85, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNNNAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 86, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 86, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNANAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 87, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 87, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNGNNNNN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 88, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 88, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAAAN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 89, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 89, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAAAGNN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 90, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 90, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAAANNN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 91, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 91, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAYAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 92, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence ofSEQ ID NO: 92, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAAAGNN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 93, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 93, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN;an amino acid sequence that is at least 80% identical to SEQ ID NO: 94, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 94, further optionally said nickase has a Protospacer Adjacent Motif (PAM) specificity of NNAGAANN; and / oran amino acid sequence that is at least 80% identical to any one of the sequences of SEQ ID NOs: 748-770, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to any one of the sequences of SEQ ID NOs: 748-770.

17. An editor, comprising:the engineered Cas9 protein of any one of claims 1-14;the nuclease of claim 15; orthe nickase of claim 16.

18. An editing system, comprising:an editor comprising the engineered Cas9 protein of any one of claims 1-14, the nuclease of claim 15 or the nickase of claim 16; anda guide RNA (gRNA).

19. The editor or editing system of any one of claims 17-18, wherein the editor comprises an effector domain.

20. The editor or editing system of any one of claims 17-19, wherein the effector domain comprises nuclease activity, nickase activity, recombinase activity, deaminase activity, methyltransferase activity, methylase activity, acetylase activity, acetyltransferase activity, transcriptional activation activity, transcriptional repression activity, and / or polymerase activity.

21. The editor or editing system of any one of claims 17-20, wherein the effector domain comprises a non-long-terminal-repeat (non-LTR) retrotransposable element enzyme or a portion thereof, optionally derived from CRE, R2, Randl / Dualen, R4, NeSL, Hero, Proto 1, LI, Txl, Proto2, RTETP, RTEX, RTE, I, Outcast, Nimb, Ingi, Jockey, Rl, Loa, Tadl, Rexl, CR1, L2A, L2B, L2, Daphne, Crack, Vingis, or any combination thereof.

22. The editor or editing system of any one of claims 17-21, wherein the effector domain is a nucleic acid editing domain.

23. The editor or editing system of any one of claims 17-22, wherein the nucleic acid editing domain comprises a deaminase domain.

24. A base editor, comprising:the engineered Cas9 protein of claim 14 or the nickase of claim 16; and a deaminase domain.

25. A base editing system, comprising:a base editor comprising a deaminase domain and (i) the engineered Cas9 protein of claim 14 or (ii) the nickase of claim 16; anda guide RNA (gRNA).

26. The editor, editing system, base editor, or base editing system of any one of claims 23-25, wherein the deaminase domain is an adenosine deaminase domain.

27. The editor, editing system, base editor, or base editing system of any one of claims 23-26, wherein the adenosine deaminase domain is an E. coli Tad A (ecTadA) deaminase domain.

28. The editor, editing system, base editor, or base editing system of any one of claims 23-27, wherein the deaminase domain is a cytosine deaminase domain.

29. The editor, editing system, base editor, or base editing system of any one of claims 23-28, wherein the cytosine deaminase domain is an apolipoprotein B mRNA-editing complex (APOBEC) family deaminase domain.

30. A reverse transcriptase (RT) editor comprising:a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain,wherein the DNA polymerase domain, the DNA binding domain, and the DNA endonuclease domain are fused or linked to form a fusion protein,wherein the DNA polymerase domain comprises a reverse transcriptase, and wherein the DNA endonuclease domain comprises the engineered Cas9 protein of claim 14 or the nickase of claim 16.

31. A reverse transcriptase (RT) editing system comprising:a template armed guide RNA (tagRNA), or a nucleic acid encoding the tagRNA, wherein the tagRNA comprises:a spacer that is complementary to a search target sequence on a first strand of a nucleic acid molecule;an editing template that comprises a region of complementarity to an editing target sequence on a second strand of the nucleic acid molecule; anda scaffold sequence that associates with a reverse transcriptase (RT) editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain; anda reverse transcriptase (RT) editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, or a nucleic acid encoding the RT editor,wherein the DNA polymerase domain comprises a reverse transcriptase; andwherein the DNA endonuclease domain comprises the engineered Cas9 protein of claim 14 or the nickase of claim 16.

32. The RT editor or RT editing system of any one of claims 30-31, wherein the tagRNA comprises a flap binding sequence at least partially complementary to the spacer.

33. The RT editor or RT editing system of any one of claims 30-32, wherein the scaffold sequence is between the spacer and the editing template.

34. The RT editor or RT editing system of any one of claims 30-33, comprising from 5’ to 3’: the spacer, the scaffold sequence, the editing template, and the flap binding sequence.

35. The RT editor or RT editing system of any one of claims 30-34, wherein the spacer, the scaffold sequence, the editing template, and the flap binding sequence form a contiguous sequence in a single molecule.

36. The RT editor or RT editing system of any one of claims 30-35, wherein the editing template comprises an intended nucleotide edit compared to the double stranded target DNA.

37. The RT editor or RT editing system of any one of claims 30-36, wherein the tagRNA guides the RT editor to incorporate the intended nucleotide edit into the double stranded target DNA when the tagRNA is contacted with the double stranded target DNA.

38. The RT editor or RT editing system of any one of claims 30-37, wherein the RT editor synthesizes a single stranded DNA encoded by the editing template, wherein the single stranded DNA replaces the editing target sequence and results in incorporation of the intended nucleotide edit into a region corresponding to the editing target in the double stranded target DNA.

39. The RT editor or RT editing system of any one of claims 30-38, wherein the search target sequence is complementary to a protospacer sequence in the double stranded target DNA, and wherein the protospacer sequence is adjacent to a protospacer adjacent motif (PAM) in the double stranded target DNA.

40. The RT editor or RT editing system of any one of claims 30-39, wherein the tagRNA results in incorporation of a nucleotide edit in the PAM when contacted with the double stranded target DNA.

41. The RT editor or RT editing system of any one of claims 30-40, wherein the spacer of the tagRNA is from 16 to 25 nucleotides in length, optionally 20 nucleotides in length or 21-23 nucleotides in length.

42. The RT editor or RT editing system of any one of claims 30-41, wherein the flap binding sequence is about 2 to 20 nucleotides in length, optionally about 8 to 16 nucleotides in length or 6 nucleotides in length.

43. The RT editor or RT editing system of any one of claims 30-42, wherein the editing template is about 4 to 30 nucleotides in length, optionally about 10 to 30 nucleotides in length, further optionally 6 to 9 nucleotides in length.

44. The RT editor or RT editing system of any one of claims 30-43, wherein the tagRNA results in incorporation of the intended nucleotide edit about 0 to 30 base pairs downstream of the nickase cleavage site.

45. The RT editor or RT editing system of any one of claims 30-44, wherein the intended nucleotide edit comprises a single nucleotide substitution compared to the region corresponding to the editing target in the double stranded target DNA.

46. The RT editor or RT editing system of any one of claims 30-45, wherein the intended nucleotide edit comprises an insertion compared to the region corresponding to the editing target in the double stranded target DNA, optionally an insertion of a nucleotide sequence at least 50, at least 45, at least 40, at least 35, at least 30, at least 25, at least 20, at least 15, at least 10, or at least 5, nucleotides in length.

47. The RT editor or RT editing system of any one of claims 30-46, wherein the intended nucleotide edit comprises a deletion compared to the region corresponding to the editing target in the double stranded target DNA.

48. The RT editor or RT editing system of any one of claims 30-47, wherein the editing template comprises one or more silent nucleotide edits compared to the region corresponding to the editing target in the double stranded target DNA, optionally said silent nucleotide edits do not alter the amino acid sequence of the protein encoded by the double stranded target DNA, further optionally said one or more silent nucleotide edits comprise a substitution of 2 to 5 contiguous nucleotides.

49. The RT editor or RT editing system of any one of claims 30-48, wherein the editing template comprises a wild type double stranded target DNA sequence.

50. The RT editor or RT editing system of any one of claims 30-49, wherein the tagRNA results in correction of a mutation when contacted with the double stranded target DNA.

51. The RT editor or RT editing system of any one of claims 30-50, wherein the reverse transcriptase is a retrovirus reverse transcriptase.

52. The RT editor or RT editing system of any one of claims 30-51, wherein the reverse transcriptase is a Moloney murine leukemia virus (MMLV) reverse transcriptase.

53. The RT editor or RT editing system of any one of claims 30-52, wherein the RT editor comprises an amino acid sequence that is at least 80% identical to any one of SEQ ID NOs: 771-796, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of any one of SEQ ID NOs: 771-796.

54. The RT editor or RT editing system of any one of claims 30-52, wherein the DNA polymerase domain, the DNA binding domain, and the DNA endonuclease domain are fused or linked to form a fusion protein.

55. A ribonucleoprotein (RNP) complex comprising the editing system, base editing system, or RT editing system of any one of claims 18-23, 25-29, and 31-53 or a component thereof.

56. A lipid nanoparticle (LNP) comprising the editing system, base editing system, or RT editing system of any one of claims 18-23, 25-29, and 31-53, or a component thereof.

57. The LNP of claim 56, comprising (iii) the gRNA and the nucleic acid encoding the editor (ii) the gRNA and the nucleic acid encoding the base editor, or (iii) the tagRNA and the nucleic acid encoding the RT editor.

58. The LNP of claim 57, wherein the nucleic acid encoding the editor, base editor, or RT editor is mRNA.

59. The LNP of any one of claims 56-58, further comprising the egRNA.

60. A system for editing SERPINA1, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 95-99 and 305- 309, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 95-99 and 305-309; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 515- 519, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 515-519; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 4, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 4; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 51, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 51.

61. A system for editing SERPINA1, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 100-119 and 310- 329, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 100-119 and 310-329; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 520- 539, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 520-539; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 11, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 11; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 58, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 58.

62. A system for editing SERPINA1, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 120-127 and 330- 337, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 120-127 and 330-337; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 540- 547, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 540-547; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 13, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 13; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 60, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 60.

63. A system for editing SERPINA1, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 128-133 and 338- 343, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 128-133 and 338-343; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 548- 553, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 548-553; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 19, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 19; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 66, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 66.

64. A system for editing SERPINA1, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 134-153 and 344- 363, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 134-153 and 344-363; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 554- 573, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 554-573; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 20, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 20; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 67, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 67.

65. A system for editing SERPINA1, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 154-159 and 364- 369, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 154-159 and 364-369; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 574- 579, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 574-579; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 10, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 10; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 57, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 57.

66. A system for editing SERPINA1, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 160-163 and 370- 373, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 160-163 and 370-373; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 580- 583, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 580-583; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 6, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 6; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 53, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 53.

67. A system for editing PAH, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 164-182 and 374- 392, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 164-182 and 374-392; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 584- 602, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 584-602; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 4, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 4; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 51, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 51.

68. A system for editing PAH, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 183-209 and 393- 419, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 183-209 and 393-419; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 603- 629, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 603-629; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 11, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 11; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 58, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 58.

69. A system for editing PAH, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 210-219 and 420- 429, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 210-219 and 420-429; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 630- 639, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 630-639; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 13, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 13; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 60, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 60.

70. A system for editing PAH, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 220-226 and 430- 436, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 220-226 and 430-436; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 640- 646, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 640-646; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 19, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 19; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 66, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 66.

71. A system for editing PAH, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 227-248 and 437- 458, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 227-248 and 437-458; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 647- 668, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 647-668; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 20, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 20; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 67, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 67.

72. A system for editing PAH, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 249-259 and 459- 469, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 249-259 and 459-469; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 669- 679, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 669-679; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 10, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 10; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 57, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 57.

73. A system for editing PAH, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 260-271 and 470- 481, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 260-271 and 470-481; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 680- 691, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 680-691; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 6, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 6; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 53, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 53.

74. A system for editing BCL11 A, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 272-277 and 482- 487, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 272-277 and 482-487; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 692- 697, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 692-697; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 4, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 4; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 51, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 51.

75. A system for editing BCL11 A, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 278-285 and 488- 495, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 278-285 and 488-495; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 698- 705, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 698-705; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 11, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 11; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 58, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 58.

76. A system for editing BCL11 A, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 286-290 and 496- 500, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 286-290 and 496-500; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 706- 710, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 706-710; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 13, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 13; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 60, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 60.

77. A system for editing BCL11 A, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 291-292 and 501- 502, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 291-292 and 501-502; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 711- 712, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 711-712; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 19, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 19; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 66, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 66.

78. A system for editing BCL11 A, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 293-300 and 503- 510, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 293-300 and 503-510; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 713- 720, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 713-720; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 20, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 20; ora nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 67, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 67.

79. A system for editing BCL11 A, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 301 and 511, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 301 and 511; and / ora protospacer comprising the sequence of SEQ ID NO: 721, or a sequence that exhibits at least about 85% identity to the sequence of SEQ ID NO: 721; and an editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 10, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 10; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 57, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 57.

80. A system for editing BCL11 A, comprising:a gRNA or a tagRNA, comprising:a sequence of any one of the sequences of SEQ ID NOs: 302-304 and 512- 514, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 302-304 and 512-514; and / ora protospacer comprising any one of the sequences of SEQ ID NOs: 722- 724, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 722-724; andan editor, a base editor, or RT editor, comprising:a nuclease comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 6, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 6; or a nickase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 53, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 53.

81. A polynucleotide encoding the editor or editing system of any one of claims 17-23, the base editor or base editing system of any one of claims 24-29, the RT editor or RT editing system of any one of claims 30-53, or the system of any one of claims 60-80, optionally a polynucleotide comprising any one of the sequences of SEQ ID NOs: 797-822, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 797-822.

82. The polynucleotide of claim 81, wherein the polynucleotide is an mRNA.

83. The polynucleotide of any one of claims 81-82, wherein the polynucleotide is operably linked to a regulatory element, optionally the regulatory element is an inducible regulatory element.

84. A vector comprising the polynucleotide of any one of claims 81-83.

85. The vector of claim 84, wherein the vector is an AAV vector.

86. An isolated cell comprising the editor or editing system of any one of claims 17-23, the base editor or base editing system of any one of claims 24-29, or the RT editor or RT editing system of any one of claims 30-53, the RNP of claim 55, the LNP of any one of claims 56-59, the system of any one of claims 60-80, the polynucleotide of any one of claims 81-83, or the vector of any one of claims 84-85.

87. The cell of claim 86, wherein the cell is a mammalian cell, optionally a human cell.

88. The cell of any one of claims 86-87, wherein the cell is a primary cell.

89. The cell of any one of claims 86-88, wherein the cell is a hepatocyte.

90. The cell of any one of claims 86-89, wherein the cell is from a subject having a disease or disorder, optionally the disease or disorder is selected from the group comprising an autoimmune disease, a neurological disease or disorder, a cancer, an inflammatory disease, a cardiovascular disease, an infectious disease, a genetic disease, a trinucleotide repeat expansion disorder, a metabolic disease, or any combination thereof, optionally P-thalassemia, sickle cell disease, alphal antitrypsin deficiency disease, phenylketonuria, or hyperphenylalaninemia, further optionally the subject is a human.

91. A pharmaceutical composition comprising:(i) the editor or editing system of any one of claims 17-23, the base editor or base editing system of any one of claims 24-29, or the RT editor or RT editing system of any one of claims 30-53, the RNP of claim 55, the LNP of any one of claims 56-59, the system of any one of claims 60-80, the polynucleotide of any one of claims 81-83, the vector of any one of claims 84-85, or the cell of any one of claims 86-90; and(ii) a pharmaceutically acceptable carrier.

92. A method for editing a double stranded target DNA, the method comprising contacting the double stranded target DNA with the editor or editing system of any one of claims 17-23, the base editor or base editing system of any one of claims 24-29, the RT editor or RT editing system of any one of claims 30-53, or the system of any one of claims 60-80, thereby editing the double stranded target DNA.

93. The method of claim 92, wherein the double stranded target DNA is in a cell.

94. The method of claim 93, wherein the cell is:a mammalian cell, optionally a human cell;a primary cell;a hepatocyte;a stem cell; optionally, an embryonic stem cell, an induced pluripotent stem cell, or an adult stem cell; and / orin a subject, optionally the subject is a human.

95. The method of any one of claims 93-94, wherein the cell is from a subject having a disease or disorder, optionally the disease or disorder is selected from the group comprising an autoimmune disease, a neurological disease or disorder, a cancer, an inflammatory disease, a cardiovascular disease, an infectious disease, a genetic disease, a trinucleotide repeat expansion disorder, a metabolic disease, or any combination thereof, optionally P-thalassemia, sickle cell disease, alphal antitrypsin deficiency disease, phenylketonuria, or hyperphenylalaninemia.

96. The method of claim 95, further comprising administering the cell to the subject after incorporation of the intended nucleotide edit.

97. A cell generated by the method of any one of claims 93-96.

98. A population of cells generated by the method of any one of claims 93-96.

99. A method for treating or preventing a disease or disorder in a subject in need thereof, the method comprising administering to the subject the editor or editing system of any one of claims 17-23, the base editor or base editing system of any one of claims 24-29, the RT editor or RT editing system of any one of claims 30-53, the system of any one of claims 60-80, the RNP of claim 55, the LNP of any one of claims 56-59, the pharmaceutical composition of claim 91, the cell of claim 97, or the population of cells of claim 98, thereby treating or preventing the disease or disorder in the subject, optionally the disease or disorder is selected from the group comprising an autoimmune disease, a neurological disease or disorder, a cancer, an inflammatory disease, a cardiovascular disease, an infectious disease, a genetic disease, a trinucleotide repeat expansion disorder, a metabolic disease, or any combination thereof.