Adenosine deaminase base editor and method for modifying nucleic acid bases in a target sequence using the same

The introduction of a novel adenine base editor (ABE8) with specific modifications enhances the specificity and efficiency of targeted nucleic acid editing, overcoming the limitations of current base editors.

JP7693552B2Active Publication Date: 2025-06-17BEAM THERAPEUTICS INC

Patent Information

Application Number
JP2021546894
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-01-27
Filing Date
2020-02-13
Publication Date
2025-06-17
Estimated Expiration
2040-02-13

AI Technical Summary

Technical Problem

Current base editors for targeted nucleic acid editing lack specificity and efficiency, necessitating the development of improved adenine base editors.

Method used

The development of a novel adenine base editor (ABE8) with increased efficiency, incorporating specific modifications at amino acid positions 82 and/or 166 of the adenosine deaminase variant, and a polynucleotide-programmable DNA binding domain.

Benefits of technology

The novel adenine base editor (ABE8) achieves higher specificity and efficiency in editing target sequences within genomic DNA, addressing the limitations of existing base editors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693552000100
    Figure 0007693552000100
  • Figure 0007693552000101
    Figure 0007693552000101
  • Figure 0007693552000102
    Figure 0007693552000102
Patent Text Reader

Abstract

The present disclosure provides compositions comprising novel adenosine base editors (e.g., ABE8) with increased efficiency and methods of using these adenosine deaminase variants to edit target sequences.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Cross - reference to Related Applications This application is an international PCT application claiming priority and benefit of U.S. Provisional Application No. 62 / 805,271, filed on February 13, 2019; No. 62 / 805,238, filed on February 13, 2019; No. 62 / 805,277, filed on February 13, 2019; No. 62 / 852,228, filed on May 23, 2019; No. 62 / 852,224, filed on May 23, 2019; No. 62 / 873,138, filed on July 11, 2019; No. 62 / 873,140, filed on July 11, 2019; No. 62 / 873,144, filed on July 11, 2019; No. 62 / 876,354, filed on July 19, 2019; No. 62 / 888,867, filed on August 19, 2019; No. 62 / 912,992, filed on October 9, 2019; No. 62 / 931,722, filed on November 6, 2019; No. 62 / 931,747, filed on November 6, 2019; No. 62 / 941,523, filed on November 27, 2019; No. 62 / 941,569, filed on November 27, 2019; and No. 62 / 966,526, filed on January 27, 2020, and the entire contents of all of these are incorporated herein by reference in their entirety.

[0002] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. Absent any other indication, the publications, patents, and patent applications mentioned in this specification are incorporated herein by reference in their entirety.

[0003] Background of the Disclosure Targeted editing of nucleic acid sequences, such as targeted cleavage or targeted modification of genomic DNA, is a very promising approach for studying gene function and also has the potential to provide new therapies for human genetic diseases. Currently available base editors include cytidine base editors (e.g., BE4) that convert target C·G base pairs to T·A and adenine base editors (e.g., ABE7.10) that convert A·T to G·C. There is a need in the art for improved base editors that can induce modifications within a target sequence with higher specificity and efficiency.

Summary of the Invention

[0004] The present invention provides compositions comprising a novel adenine base editor (e.g., ABE8) having increased efficiency, and methods of using a base editor comprising an adenosine deaminase variant for editing a target sequence.

[0005] In one aspect, the present invention provides a polynucleotide-programmable DNA binding domain, and modifications at amino acid positions 82 and / or 166 of MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSST Provided is a fusion protein comprising at least one base editor domain that is an adenosine deaminase variant comprising a modification corresponding to position 82 and / or 166 in the variant. In some embodiments, the invention provides a polynucleotide programmable DNA binding domain, as well as MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD, or an adenosine deaminase variant comprising a modification at amino acid positions 82 and / or 166 of an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the adenosine deaminase variant is any of the following adenosine deaminases, or in any variant thereof MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD comprising a modification at amino acid positions 82 and / or 166 corresponding to: Staphylococcus aureus (S. aureus) TadA: MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN Bacillus subtilis (B. subtilis) TadA: MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLSE Salmonella typhimurium (S. typhimurium) TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFRMRRQEIKALKKADRAEGAGPAV Shewanella putrefaciens (S. putrefaciens) TadA: MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE Haemophilus influenzae F3031 (H. influenzae) TadA: MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTΑΗAEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRREEKKIEKALLKSLSDK Caulobacter crescentus (C. crescentus) TadA: MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLTDLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAKI Geobacter sulfurreducens (G. sulfurreducens) TadA: MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRKKAKATPALFIDERKVPPEP

[0006] In some embodiments, the polynucleotide programmable DNA binding domain has the following sequence: TIFF0007693552000001.tif171162, where the bold sequence indicates the sequence derived from Cas9, the italic sequence indicates the linker sequence, and the underlined sequence indicates the bipartite nuclear localization sequence.

[0007] In various embodiments of the above aspects, the adenosine deaminase variant comprises modifications at amino acid positions 82 and 166. In various embodiments of the above aspects, the adenosine deaminase variant comprises the V82S modification. In various embodiments of the above aspects, the adenosine deaminase variant comprises the T166R modification. In various embodiments of the above aspects, the adenosine deaminase variant comprises the V82S and T166R modifications. In various embodiments of the above aspects, the adenosine deaminase variant further comprises one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, and Q154R. In another embodiment, the adenosine deaminase variant comprises a combination of the following modifications: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R + Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; or I76Y + V82S + Y123H + Y147R + Q154R. In various embodiments of the above aspects, the adenosine deaminase variant comprises a C-terminal deletion starting at a residue selected from the group consisting of 149, 150, 151, 152, 153, 154, 155, 156, and 157.

[0008] In some embodiments, the base editor domain comprises an adenosine deaminase variant monomer. In some embodiments, the base editor domain is ABE8.1-m, ABE8.2-m, ABE8.3-m, ABE8.4-m, ABE8.5-m, ABE8.6-m, ABE8.7-m, ABE8.8-m, ABE8.9-m, ABE8.10-m, ABE8.11-m, ABE8.12-m, ABE8.13-m, ABE8.14-m, ABE8.15-m, ABE8.16-m, ABE8.17-m, ABE8.18-m, ABE8.19-m, ABE8.20-m, ABE8.21-m, ABE8.22-m, ABE8.23-m, or ABE8.24-m. In various embodiments, the base editor domain comprises an adenosine deaminase heterodimer comprising a wild-type adenosine deaminase domain and an adenosine deaminase variant. In some embodiments, the base editor domain is ABE8.1-d, ABE8.2-d, ABE8.3-d, ABE8.4-d, ABE8.5-d, ABE8.6-d, ABE8.7-d, ABE8.8-d, ABE8.9-d, ABE8.10-d, ABE8.11-d, ABE8.12-d, ABE8.13-d, ABE8.14-d, ABE8.15-d, ABE8.16-d, ABE8.17-d, ABE8.18-d, ABE8.19-d, ABE8.20-d, ABE8.21-d, ABE8.22-d, ABE8.23-d, or ABE8.24-d. In various embodiments of the above aspects, the base editor comprises a heterodimer comprising a TadA7.10 domain and an adenosine deaminase variant domain. In some embodiments, the adenosine deaminase variant is TadA * 8.1, TadA * 8.2, TadA * 8.3, TadA * 8.4, TadA * 8.5, TadA * 8.6, TadA * 8.7, TadA * 8.8, TadA * 8.9, TadA * 8.10, TadA *8.11, TadA * 8.12, TadA * 8.13, TadA * 8.14, TadA * 8.15, TadA * 8.16, TadA * 8.17, TadA * 8.18, TadA * 8.19, TadA * 8.20, TadA * 8.21, TadA * 8.22, TadA * 8.23, or TadA * 8.24.

[0009] In various embodiments, the adenosine deaminase variant comprises or consists essentially of the following sequence or a fragment thereof having adenosine deaminase activity: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD

[0010] In another embodiment, the adenosine deaminase variant lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 N-terminal amino acid residues compared to the full-length adenosine deaminase. In another embodiment, the adenosine deaminase variant lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 C-terminal amino acid residues compared to the full-length adenosine deaminase. In some embodiments, the adenosine deaminase variant comprises Y147T+Q154R. In some embodiments, the adenosine deaminase variant comprises Y147T+Q154S. In some embodiments, the adenosine deaminase variant comprises Y147R+Q154S. In some embodiments, the adenosine deaminase variant comprises V82S+Q154S; V82S+Y147R. In some embodiments, the adenosine deaminase variant comprises V82S+Q154R. In some embodiments, the adenosine deaminase variant comprises V82S+Y123H. In some embodiments, the adenosine deaminase variant comprises I76Y+V82S. In some embodiments, the adenosine deaminase variant comprises V82S+Y123H+Y147T. In some embodiments, the adenosine deaminase variant comprises V82S+Y123H+Y147R. In some embodiments, the adenosine deaminase variant comprises V82S+Y123H+Q154R. In some embodiments, the adenosine deaminase variant comprises Y147R+Q154R +Y123H. In some embodiments, the adenosine deaminase variant comprises Y147R+Q154R+I76Y. In some embodiments, the adenosine deaminase variant comprises Y147R+Q154R+T166R. In some embodiments, the adenosine deaminase variant comprises Y123H+Y147R+Q154R+I76Y. In some embodiments, the adenosine deaminase variant comprises V82S+Y123H+Y147R+Q154R.In some embodiments, the adenosine deaminase variant comprises I76Y+V82S+Y123H+Y147R+Q154R.

[0011] In some embodiments, the polynucleotide programmable DNA binding domain is Cas9. In some embodiments, the Cas9 polypeptide has the following amino acid sequence (Cas9 reference sequence): TIFF0007693552000002.tif168163 (single underline: HNH domain; double underline: RuvC domain; (Cas9 reference sequence)), or a corresponding region thereof.

[0012] In various embodiments, the polynucleotide-programmable DNA binding domain is Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1Cas9 (St1Cas9), Streptococcus pyogenes Cas9 (SpCas9), or variants thereof. In various embodiments, the polynucleotide-programmable DNA binding domain includes variants of SpCas9 having modified protospacer adjacent motif (PAM) specificities or specificities for non-G PAMs. In various embodiments of the above aspects, the modified PAM has specificity for the nucleic acid sequence 5'-NGC-3'. In various embodiments of the above aspects, the modified SpCas9 includes the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R, or corresponding amino acid substitutions thereof. In various embodiments of the above aspects, the polynucleotide-programmable DNA binding domain is nuclease-inactive or a nickase variant. In various embodiments of the above aspects, the nickase variant includes the amino acid substitution D10A or a corresponding amino acid substitution thereof. In various embodiments of the above aspects, the base editor further includes a zinc finger domain. In various embodiments of the above aspects, the adenosine deaminase domain is capable of deaminating adenosine in deoxyribonucleic acid (DNA).

[0013] In various embodiments, the adenosine deaminase variant is TadA deaminase. In various embodiments of the above aspects, the TadA deaminase is TadA * 7.10. In various embodiments, the TadA deaminase is TadA *There are 8 variants. The adenosine deaminase variant can deaminate adenine in deoxyribonucleic acid (DNA). In some embodiments, the adenosine deaminase variant is Staphylococcus aureus TadA, Bacillus subtilis TadA, Salmonella typhimurium TadA, Shewanella putrefaciens TadA, Haemophilus influenzae F3031 TadA, Caulobacter crescentus (C. crescentus) TadA, or Geobacter sulfurreducens TadA, or a fragment thereof. In some embodiments, the adenosine deaminase variant is a non-naturally occurring adenosine deaminase.

[0014] In various embodiments of the above aspects, the fusion protein contains a linker between the polynucleotide programmable DNA binding domain and the adenosine deaminase domain. In various embodiments of the above aspects, the linker contains the amino acid sequence: SGGSSGGSSGSETPGTSESATPES. In various embodiments of the above aspects, it contains one or more nuclear localization signals. In various embodiments of the above aspects, the nuclear localization signal is a bipartite nuclear localization signal.

[0015] In various embodiments of the above aspects, Cas9 is StCas9 or SaCas9. In various embodiments of the above aspects, Cas9 is a modified SaCas9. In various embodiments of the above aspects, the modified SaCas9 contains the amino acid substitutions E781K, N967K, and R1014H, or their corresponding amino acid substitutions. In various embodiments of the above aspects, the modified SaCas9 has the amino acid sequence: comprises

[0016] an adenosine deaminase variant domain, MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSST comprising the amino acid sequence thereof, wherein the amino acid sequence comprises at least one modification, an adenosine deaminase variant domain, and a Cas9 or Cas12 polypeptide, wherein the adenosine deaminase variant domain is inserted into the Cas9 or Cas12 polypeptide, a fusion protein comprising the Cas9 or Cas12 polypeptide.

[0017] In some embodiments, the adenosine deaminase variant domain is TadA * 8 adenosine deaminase monomer comprising an adenosine deaminase variant domain. In some embodiments, the adenosine deaminase variant domain is a wild-type adenosine deaminase domain and TadA * 8 adenosine deaminase heterodimer comprising an adenosine deaminase variant domain. In some embodiments, the adenosine deaminase variant domain is a TadA domain and TadA *An adenosine deaminase heterodimer comprising an 8 adenosine deaminase variant domain. In some embodiments, the adenosine deaminase variant domain is inserted within a flexible loop, an alpha helix region, an unstructured portion, or a solvent accessible portion of a Cas9 or Cas12 polypeptide. In some embodiments, the flexible loop comprises a portion of the alpha helix structure of the Cas9 or Cas12 polypeptide. In some embodiments, the adenosine deaminase variant domain is flanked by an N-terminal fragment and a C-terminal fragment of the Cas9 polypeptide. In some embodiments, the fusion protein comprises the structure: NH2-[N-terminal fragment of Cas9]-[adenosine deaminase variant]-[C-terminal fragment of Cas9]-COOH, where each instance of "]-[“ is an optional linker. In some embodiments, the N-terminal fragment or C-terminal fragment of the Cas9 or Cas12 polypeptide binds to a target polynucleotide sequence.

[0018] In some embodiments, the C-terminus of the N-terminal fragment or the N-terminus of the C-terminal fragment comprises a flexible loop portion of the Cas9 or Cas12 polypeptide. In some embodiments, the flexible loop comprises an amino acid that is proximal to the target nucleobase when the fusion protein deaminates the target nucleobase. In some embodiments, the target nucleobase is separated from the protospacer adjacent motif (PAM) sequence in the target polynucleotide sequence by 1 to 20 nucleobases. In some embodiments, the target nucleobase is located 2 to 12 nucleobases upstream of the PAM sequence. In some embodiments, the N-terminal fragment or the C-terminal fragment comprises the RuvC domain, the N-terminal fragment or the C-terminal fragment comprises the HNH domain, neither the N-terminal fragment nor the C-terminal fragment comprises the HNH domain, or neither the N-terminal fragment nor the C-terminal fragment comprises the RuvC domain. In some embodiments, the Cas9 or Cas12 polypeptide comprises a partial or complete deletion in one or more structural domains, and the adenosine deaminase is inserted at the position of the partial or complete deletion of the Cas9 or Cas12 polypeptide. In some embodiments, the deletion is within the RuvC domain, the deletion is within the HNH domain, or the deletion bridges the RuvC domain and the C-terminal domain.

[0019] In some embodiments, the adenosine deaminase variant domain is inserted into the Cas12 polypeptide. In some embodiments, the Cas9 polypeptide is Streptococcus pyogenes Cas9 (SpCas9), Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), or a variant thereof. In some embodiments, Cas9 has the following amino acid sequence (Cas9 reference sequence): TIFF0007693552000003.tif169163 (single underline: HNH domain; double underline: RuvC domain; ("Cas9 reference sequence")), or a corresponding region thereof.

[0020] In one aspect, the present invention provides any fusion protein provided herein, wherein the Cas9 polypeptide comprises a deletion of amino acids 1017 - 1069 or their corresponding amino acids in the Cas9 polypeptide reference sequence numbering, the Cas9 polypeptide comprises a deletion of amino acids 792 - 872 or their corresponding amino acids in the Cas9 polypeptide reference sequence numbering, or the Cas9 polypeptide comprises a deletion of amino acids 792 - 906 or their corresponding amino acids in the Cas9 polypeptide reference sequence numbering. In one embodiment, the adenosine deaminase variant domain is inserted within a flexible loop of the Cas9 polypeptide. In one embodiment, the flexible loop comprises a region selected from the group consisting of amino acid residues at positions 530 - 537, 569 - 579, 686 - 691, 768 - 793, 943 - 947, 1002 - 1040, 1052 - 1077, 1232 - 1248, and 1298 - 1300 in the Cas9 reference sequence numbering, or their corresponding amino acid positions. In one embodiment, the adenosine deaminase variant domain is inserted between amino acid positions 768 - 769, 791 - 792, 792 - 793, 1015 - 1016, 1022 - 1023, 1026 - 1027, 1029 - 1030, 1040 - 1041, 1052 - 1053, 1054 - 1055, 1067 - 1068, 1068 - 1069, 1247 - 1248, or 1248 - 1249 in the Cas9 reference sequence numbering, or their corresponding amino acid positions. In one embodiment, the adenosine deaminase variant domain is inserted between amino acid positions 768 - 769, 792 - 793, 1022 - 1023, 1026 - 1027, 1040 - 1041, 1068 - 1069, or 1247 - 1248 in the Cas9 reference sequence numbering, or their corresponding amino acid positions. In one embodiment, the adenosine deaminase variant domain is inserted between amino acid positions 1016 - 1017, 1023 - 1024, 1029 - 1030, 1040 - 1041, 1069 - 1070, or 1247 - 1248 in the Cas9 reference sequence numbering, or their corresponding amino acid positions.In one embodiment, the deaminase variant domain is inserted into the Cas9 polypeptide at the locus identified in Table 10A. In one embodiment, the N-terminal fragment comprises amino acid residues 1-529, 538-568, 580-685, 692-942, 948-1001, 1026-1051, 1078-1231, and / or 1248-1297 of the Cas9 reference sequence, or corresponding residues thereof. In one embodiment, the C-terminal fragment comprises amino acid residues 1301-1368, 1248-1297, 1078-1231, 1026-1051, 948-1001, 692-942, 580-685, and / or 538-568 of the Cas9 reference sequence, or corresponding residues thereof.

[0021] In some embodiments, the Cas9 polypeptide is a nickase or the Cas9 polypeptide is nuclease-inactive. In one embodiment, the Cas9 polypeptide is a modified SpCas9 polypeptide having specificity for a modified PAM or specificity for a non-G PAM. In one embodiment, the modified SpCas9 polypeptide comprises the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (SpCas9-MQKFRAER) and has specificity for the modified PAM 5'-NGC-3'.

[0022] In one embodiment, the adenosine deaminase variant domain is inserted into the Cas12 polypeptide. In one embodiment, the Cas12 polypeptide is Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, or Cas12i. In one embodiment, the Cas12 polypeptide has at least about 85% amino acid sequence identity to Bacillus hisashii Cas12b, Bacillus thermoamylovorans Cas12b, Bacillus sp. V3-13 Cas12b, or Alicyclobacillus acidiphilus Cas12b. In one embodiment, the Cas12 polypeptide comprises or consists essentially of a fragment of Bacillus hisashii Cas12b, Bacillus thermoamylovorans Cas12b, Bacillus sp. V3-13 Cas12b, or Alicyclobacillus acidiphilus Cas12b. In one embodiment, the adenosine deaminase variant domain is inserted between amino acid positions: a) 153-154, 255-256, 306-307, 980-981, 1019-1020, 534-535, 604-605, or 344-345 of BhCas12b, or the corresponding amino acid residues of Cas12a, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, or Cas12i, b) 147 and 148, 248 and 249, 299 and 300, 991 and 992, or 1031 and 1032 of BvCas12b, or the corresponding amino acid residues of Cas12a, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, or Cas12i, or c) 157 and 158, 258 and 259, 310 and 311, 1008 and 1009, or 1044 and 1045 of AaCas12b, or the corresponding amino acid residues of Cas12a, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, or Cas12i.In one embodiment, the adenosine deaminase variant domain is inserted at the locus identified in Table 10B.

[0023] In some embodiments, the Cas12 polypeptide is Cas12b. In one embodiment, the Cas12 polypeptide comprises a BhCas12b domain, a BvCas12b domain, or an AACas12b domain. The Cas12b polypeptide comprises a mutation that silences the catalytic activity of the RuvC domain. In one embodiment, the Cas12b polypeptide comprises D574A, D829A and / or D952A mutations.

[0024] In one embodiment, the adenosine deaminase variant domain comprises Staphylococcus aureus TadA, Bacillus subtilis TadA, Salmonella typhimurium TadA, Shewanella putrefaciens TadA, Haemophilus influenzae F3031 TadA, Caulobacter crescentus (C. crescentus) TadA, or Geobacter sulfurreducens TadA, or a variant or fragment thereof. In one embodiment, the adenosine deaminase is a non-naturally occurring adenosine deaminase.

[0025] In various embodiments, the fusion protein further comprises a cytidine deaminase. In one aspect, the present invention provides a fusion protein comprising the structure: NH2-[TadA * 8]-[Cas9]-[cytidine deaminase]-COOH, where each instance of "]-[ " is an optional linker. In another aspect, the present invention provides a fusion protein comprising the structure: NH2-[cytidine deaminase]-[Cas9]-[TadA * 8]-COOH, where each instance of "]-[ " is an optional linker.

[0026] In one aspect, the present invention provides a fusion protein comprising the structure: NH2-[Cas9 (TadA * 8)]-[cytidine deaminase]-COOH, wherein each description of "]-[ " is an optional linker (e.g., TadA * 8 is fused inside Cas9, and cytidine deaminase is fused to the C-terminus). In one aspect, the present invention provides a fusion protein comprising the structure: NH2-[cytidine deaminase]-[Cas9 (TadA * 8)]-COOH, wherein each description of "]-[ " is an optional linker (e.g., TadA * 8 is fused inside Cas9, and cytidine deaminase is fused to the N-terminus). In another aspect, the present invention provides a fusion protein comprising the structure: NH2-[Cas9(cytidine deaminase)]-[TadA * 8]-COOH, wherein each description of "]-[ " is an optional linker. In yet another aspect, the present invention provides a fusion protein comprising the structure: NH2-[TadA * 8]-[Cas9(cytidine deaminase)]-COOH, wherein each description of "]-[ " is an optional linker.

[0027] In one aspect, the present invention provides a fusion protein comprising the structure: NH2-[Cas12 (adenosine deaminase)]-[cytidine deaminase]-COOH, where each instance of "]-[ " is an optional linker (e.g., adenosine deaminase is fused internally to Cas12 and cytidine deaminase is fused to the C-terminus). In one aspect, the present invention provides a fusion protein comprising the structure: NH2-[cytidine deaminase]-[Cas12 (adenosine deaminase)]-COOH, where each instance of "]-[ " is an optional linker (e.g., adenosine deaminase is fused internally to Cas12 and cytidine deaminase is fused to the N-terminus). In another aspect, the present invention provides a fusion protein comprising the structure: NH2-[Cas12 (cytidine deaminase)]-[adenosine deaminase]-COOH, where each instance of "]-[ " is an optional linker. In yet another aspect, the present invention provides a fusion protein comprising the structure: NH2-[adenosine deaminase]-[Cas12 (cytidine deaminase)]-C, where each instance of "]-[ " is an optional linker.

[0028] In another aspect, the present invention provides any of the fusion proteins provided herein in complex with one or more guide nucleic acid sequences that effect deamination of a target nucleobase. In some embodiments, the fusion protein is further complexed with a target polynucleotide.

[0029] In one aspect, the present invention provides a polynucleotide encoding any of the fusion proteins provided herein. In another aspect, the present invention provides an expression vector comprising any of the polynucleotides provided herein. In some embodiments, the expression vector is a mammalian expression vector. In some embodiments, the vector is a viral vector selected from the group consisting of adeno-associated virus (AAV), retroviral vector, adenoviral vector, lentiviral vector, Sendai virus vector, and herpesvirus vector. In some embodiments, the vector comprises a promoter.

[0030] In one aspect, the present invention provides a cell comprising any fusion protein provided herein. In one aspect, the present invention provides a cell comprising any polynucleotide provided herein. In one aspect, the present invention provides a cell comprising any vector provided herein. In various embodiments, the cell is a bacterial cell, a plant cell, an insect cell, a human cell, or a mammalian cell.

[0031] In one aspect, the present invention provides a base editor comprising any fusion polypeptide provided herein in complex with one or more guide polynucleotides. In one aspect, the present invention provides a pharmaceutical composition comprising any fusion protein provided herein and a pharmaceutically acceptable excipient. In one aspect, the present invention provides a pharmaceutical composition comprising any polynucleotide provided herein and a pharmaceutically acceptable excipient. In one aspect, the present invention provides a pharmaceutical composition comprising any vector provided herein and a pharmaceutically acceptable excipient. In one aspect, the present invention provides a pharmaceutical composition comprising any cell provided herein and a pharmaceutically acceptable excipient. In one aspect, the present invention provides a pharmaceutical composition comprising any base editor provided herein and a pharmaceutically acceptable excipient.

[0032] In one aspect, the present invention provides a kit comprising any fusion protein provided herein. In one aspect, the present invention provides a kit comprising any polynucleotide provided herein. In one aspect, the present invention provides a kit comprising any vector provided herein. In one aspect, the present invention provides a kit comprising any base editor provided herein.

[0033] In another aspect, the present invention provides a method comprising contacting a polynucleotide sequence with any of the fusion proteins provided herein, wherein the adenosine deaminase variant domain of the fusion protein deaminates the nucleobases in the polynucleotide, thereby editing the polynucleotide sequence. In some embodiments, the method further comprises contacting a target polynucleotide sequence with one or more guide polynucleotides to effect deamination of a target nucleobase.

[0034] In one aspect, the present invention provides a method of editing a target polynucleotide, the method comprising contacting the target polynucleotide with the base editor of claim 116 to effect a modification from A·T to G·C in the target polynucleotide. In some embodiments, the method further comprises contacting in a cell, eukaryotic cell, mammalian cell, or human cell. In some embodiments, the cell is in vivo. In some embodiments, the cell is ex vivo.

[0035] In one aspect, the present invention provides a method of treating a genetic defect in a subject, the method comprising administering to the subject a base editor comprising or consisting essentially of any of the fusion proteins provided herein, or a polynucleotide encoding the base editor and one or more guide polynucleotides that direct the base editor to deaminate a target nucleobase in the target nucleotide sequence of the subject, thereby treating the genetic defect.

[0036] In some embodiments, the guide polynucleotide comprises a nucleic acid sequence selected from the group consisting of: a) GACCUAGGCGAGGCAGUAGG; b) CCAGUAUGGACACUGUCCAAA; c) CAGUAUGGACACUGUCCAAA; and d) AGUAUGGACACUGUCCAAAG. In various embodiments, the gRNA has the nucleic acid sequence GUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG Further comprising. In various embodiments, the guide RNA comprises a CRISPR RNA (crRNA) and a trans-activating crRNA (tracrRNA).

[0037] In some embodiments, the method further comprises delivering a base editor, or a polynucleotide encoding said base editor, and one or more guide polynucleotides to a target cell. In various embodiments, the subject is a mammal or a human. In various embodiments, the deamination of the target nucleobase replaces the target nucleobase with a wild-type nucleobase. In various embodiments, the deamination of the target nucleobase replaces the target nucleobase with a non-wild-type nucleobase, and the deamination of the target nucleobase improves the symptoms of a genetic condition. In various embodiments, the target polynucleotide sequence comprises a mutation associated with a genetic condition at a nucleobase other than the target nucleobase.

[0038] The description and examples herein explain the embodiments of the present disclosure in detail. It should be understood that the present disclosure is not limited to the specific embodiments described herein and that they can be modified. Those skilled in the art will recognize that there are numerous variations and modifications of the present disclosure, and these are included within its scope.

[0039] The practice of some embodiments disclosed herein, unless otherwise indicated, uses conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genetics, and recombinant DNA, which are within the skill of the art. See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (F. M. Ausubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (M.J. MacPherson, B.D. Hames and G.R. Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (R.I. Freshney, ed. (2010)).

[0040] The section headings used herein are for editorial purposes only and should not be construed as limiting the subject matter described.

[0041] Although various features of the disclosure may be described in the context of a single embodiment, the features may also be provided separately or in any suitable combination. Conversely, although the disclosure may be described herein in the context of separate embodiments for clarity, the disclosure may also be implemented in a single embodiment. The section headings used herein are for editorial purposes only and should not be construed as limiting the subject matter described.

[0042] The features of the present disclosure are described in detail in the appended claims. A better understanding of the features and advantages of the present disclosure can be obtained by referring to the following detailed description which describes exemplary embodiments in which the principles of the present disclosure are utilized, and by considering the appended drawings as described hereinafter in this specification.

[0043] [Definitions] The following definitions supplement those of the art and are for the present application and do not pertain to related or unrelated cases, such as patents or applications under common ownership. Any methods and materials similar or equivalent to those described herein can be used in the practice of the present disclosure, but the preferred materials and methods are described herein. Accordingly, the terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting.

[0044] Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. The following references provide to one of ordinary skill in the art general definitions of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991).

[0045] In this application, the use of the singular form includes the plural, unless otherwise specified. It should be noted that as used herein, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" include references to the plural. In this application, unless stated otherwise, the use of "or" means "and / or" and is understood to be inclusive. Furthermore, the use of the term "including", as well as other forms such as "include", "includes", and "included", is non-limiting.

[0046] As used in this specification and the claims, the terms "comprising" (and any form thereof such as "comprise" and "comprises"), "having" (any form thereof such as "have" and "has"), "including" (any form thereof such as "include" and "includes"), or "containing" (any form thereof such as "contains" and "contain") are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. Any embodiment discussed herein is contemplated to be applicable to any method or composition of the present disclosure, and vice versa. Furthermore, the methods of the present disclosure can be achieved using the compositions of the present disclosure.

[0047] The terms "about" or "approximately" mean within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within one standard deviation or within more than one standard deviation according to the practice in the relevant art. Alternatively, "about" can mean within a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. As another alternative, especially with respect to biological systems or processes, this term can mean within the same order of magnitude, e.g., within a factor of five, within a factor of two. When a particular value is recited in the application and claims, unless otherwise stated, the term "about" meaning within an acceptable error range for that particular value should be presumed.

[0048] The ranges provided herein are to be understood as shorthand for all of the values within the range. For example, the range of 1 to 50 is to be understood to include any number, combination of numbers, or sub-ranges from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.

[0049] References in the specification to "some embodiments", "an embodiment", "one embodiment", or "another embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least some embodiments of the present disclosure, but not necessarily all embodiments.

[0050] An "abasic base editor" refers to an agent that can excise a nucleic acid base and insert a DNA nucleic acid base (A, T, C, or G). The abasic base editor includes a nucleic acid glycosylase polypeptide or a fragment thereof. In one embodiment, the nucleic acid glycosylase contains Asp at amino acid 204 of the following sequence or at the corresponding position of uracil DNA glycosylase (e.g., by substituting Asn at amino acid 204), and is a mutant human uracil DNA glycosylase or an active fragment thereof having cytosine-DNA glycosylase activity. In one embodiment, the nucleic acid glycosylase contains Ala, Gly, Cys, or Ser at amino acid 147 of the following sequence or at the corresponding position of uracil DNA glycosylase (e.g., by substituting Tyr at amino acid 147), and is a mutant human uracil DNA glycosylase or an active fragment thereof having thymine-DNA glycosylase activity. The sequence of an exemplary human uracil-DNA glycosylase, isoform 1, is as follows: 1 mgvfclgpwg lgrklrtpgk gplqllsrlc gdhlqaipak kapagqeepg tppssplsae 61 qldriqrnka aallrlaarn vpvgfgeswk khlsgefgkp yfiklmgfva eerkhytvyp 121 pphqvftwtq mcdikdvkvv ilgqdp y hgp nqahglcfsv qrpvppppsl eniykelstd 181 iedfvhpghg dlsgwakqgv lll n avltvr ahqanshker gweqftdavv swlnqnsngl 241 vfllwgsyaq kkgsaidrkr hhvlqtahps p l svy r gffg crhfsktnel lqksgkkpid 301 wkel

[0051] The sequence of human uracil-DNA glycosylase, isoform 2 is as follows. 1 migqktlysf fspsparkrh apspepavqg tgvagvpees gdaaaipakk apagqeepgt 61 ppssplsaeq ldriqrnkaa allrlaarnv pvgfgeswkk hlsgefgkpy fiklmgfvae 121 erkhytvypp phqvftwtqm cdikdvkvvi lgqdp y hgpn qahglcfsvq rpvppppsle 181 niykelstdi edfvhpghgd lsgwakqgvl ll n avltvra hqanshkerg weqftdavvs 241 wlnqnsnglv fllwgsyaqk kgsaidrkrh hvlqtahpsp l svy r gffgc rhfsktnell 301 qksgkkpidw kel

[0052] In other embodiments, the base editor is any of the base editors described in PCT / JP205 / 080958 and US20170321210, which are incorporated herein by reference. In certain embodiments, the base editor comprises a mutation at the positions shown in bold and underlined in the above sequences, or at the corresponding amino acids in any other base editor or uracil glycosylase known in the art. In one embodiment, the base editor comprises a mutation at Y147, N204, L272, and / or R276 or corresponding positions. In another embodiment, the base editor comprises the Y147A or Y147G mutation or corresponding mutations. In another embodiment, the base editor comprises the N204D mutation or corresponding mutations. In another embodiment, the base editor comprises the L272A mutation or corresponding mutations. In another embodiment, the base editor comprises the R276E or R276C mutation or corresponding mutations.

[0053] "Adenosine deaminase" means a polypeptide or fragment thereof that can catalyze the hydrolytic deamination of adenine or adenosine. In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine to inosine, or deoxyadenosine to deoxyinosine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., genetically engineered adenosine deaminases, evolved adenosine deaminases) can be from any organism, such as bacteria.

[0054] In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is a TadA variant. In some embodiments, the TadA variant is TadA *It is 8. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism such as a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain does not occur naturally. For example, in some embodiments, the deaminase or deaminase domain has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identity to a naturally occurring deaminase. For example, the deaminase domain is described in International PCT Application Nos. PCT / 2017 / 045381 (WO 2018 / 027078) and PCT / US2016 / 058344 (WO 2017 / 070632), each of which is hereby incorporated by reference in its entirety.Komor, A.C., et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, N.M., et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, A.C., et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity”Science Advances 3:eaao4774 (2017) ), and Rees, H.A., et al., “Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec;19(12):770-788. doi: 10.1038 / s41576-018-0059-1 (the entire contents of which are incorporated herein by reference). See also.

[0055] Wild-type TadA (wt) adenosine deaminase has the following sequence (also referred to as the TadA reference sequence; SEQ ID NO: 2): MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD

[0056] In some embodiments, adenosine deaminase includes modifications in the following sequence: MSEVEFSHEY WMRHALTLAK RARDEREVPV GAVLVLNNRV IGEGWNRAIG LHDPTAHAEI MALRQGGLVM QNYRLIDATL YVTFEPCVMC AGAMIHSRIG RVVFGVRNAK TGAAGSLMDV LHYPGMNHRV EITEGILADE CAALLCYFFR MPRQVFNAQK KAQSSTD (TadA * (also referred to as 7.10).

[0057] In some embodiments, TadA * 7.10 includes at least one modification. In some embodiments, TadA * 7.10 includes modifications at amino acids 82 and / or 166. In certain embodiments, variants of the above-referenced sequence include one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. The modification Y123H may also be referred to herein as H123H (wherein the modification H123Y in TadA * 7.10 has reverted to Y123H (wt)). In other embodiments, variants of the TadA * 7.10 sequence include combinations of modifications selected from the group consisting of Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R.

[0058] In other embodiments, the invention provides adenosine deaminase variants that contain deletions, such as TadA * 7.10, having a C-terminal deletion starting at residue 149, 150, 151, 152, 153, 154, 155, 156, or 155 compared to the TadA reference sequence, or corresponding mutations in another TadA 7 * 8. In other embodiments, the adenosine deaminase variant is a TadA * 7.10, having one or more of the following modifications compared to the TadA reference sequence: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R, or corresponding mutations in another TadA (e.g., TadA * 8) monomer. In other embodiments, the adenosine deaminase variant is a TadA * 7.10, having a combination of modifications selected from the group consisting of Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R compared to the TadA reference sequence, or a monomer containing corresponding mutations in another TadA.

[0059] In yet other embodiments, the adenosine deaminase variant is a TadA * 7.10, having one or more of the following modifications Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R respectively, compared to the TadA reference sequence, or corresponding mutations in another TadA, in two adenosine deaminase domains (e.g., TadA *It is a homodimer containing (8). In other embodiments, the adenosine deaminase variant is TadA * 7.10, compared to the TadA reference sequence, combinations of mutations selected from the group of Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R, or corresponding mutations in another TadA, and two adenosine deaminase domains each having (e.g., TadA * 8) and is a homodimer.

[0060] In other embodiments, the adenosine deaminase variant is a wild-type TadA adenosine deaminase domain, and TadA * 7.10, an adenosine deaminase variant domain containing one or more of the following modifications compared to the TadA reference sequence: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R, or corresponding mutations in another TadA (e.g., TadA * 8) and is a heterodimer. In other embodiments, the adenosine deaminase variant is a wild-type TadA adenosine deaminase domain, and TadA *7.10. A combination of modifications selected from the group of Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R, or corresponding mutations in another TadA, including an adenosine deaminase variant domain (e.g., TadA * 8) is a heterodimer.

[0061] In other embodiments, the adenosine deaminase variant is TadA * 7.10 domain, and TadA * 7.10. An adenosine deaminase variant domain (e.g., TadA * 8) including one or more of the following modifications Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R compared to the TadA reference sequence, or corresponding mutations in another TadA, is a heterodimer. In other embodiments, the adenosine deaminase variant is TadA * 7.10 domain, and TadA *7.10. Modifications compared to the TadA reference sequence: Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; or a combination of I76Y+V82S+Y123H+Y147R+Q154R, or corresponding mutations in another TadA, an adenosine deaminase variant domain (e.g., TadA * 8) is a heterodimer containing.

[0062] In one embodiment, the adenosine deaminase comprises or consists essentially of the following sequence or a fragment thereof having adenosine deaminase activity, TadA * 8 is: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD

[0063] In some embodiments, TadA * 8 is truncated. In some embodiments, the truncated TadA * 8 lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues compared to the full-length TadA * 8. In some embodiments, the truncated TadA * 8 is the full-length TadA *It lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues compared to 8. In some embodiments, the adenosine deaminase variant is full-length TadA * is 8.

[0064] In certain embodiments, the adenosine deaminase heterodimer is TadA * 8 domain and an adenosine deaminase domain selected from one of the following:

[0065] Staphylococcus aureus (S. aureus) TadA: MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN

[0066] Bacillus subtilis (B. subtilis) TadA: MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLSE

[0067] Salmonella typhimurium (S. typhimurium) TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFRMRRQEIKALKKADRAEGAGPAV

[0068] Shewanella putrefaciens (S. putrefaciens) TadA: MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE

[0069] Haemophilus influenzae F3031 (H. influenzae) TadA: MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTΑΗAEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRREEKKIEKALLKSLSDK

[0070] Caulobacter crescentus (C. crescentus) TadA: MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLTDLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAKI

[0071] Geobacter sulfurreducens (G. sulfurreducens) TadA: MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRKKAKATPALFIDERKVPPEP

[0072] TadA*7.10 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD

[0073] The "Adenosine Deaminase Base Editor 8 (ABE8) polypeptide" or "ABE8" means a base editor as defined herein that includes an adenosine deaminase variant that contains a modification at amino acid positions 82 and / or 166 of the following reference sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD In some embodiments, ABE8 includes additional modifications as described herein compared to the reference sequence.

[0074] The "Adenosine Deaminase Base Editor 8 (ABE8) polynucleotide" means a polynucleotide encoding ABE8.

[0075] "Administering" is referred to herein as providing one or more of the compositions described herein to a patient or subject. By way of example, and not limitation, administration of a composition, such as an injection, can be performed by intravenous (i.v.) injection, subcutaneous (s.c.) injection, intradermal (i.d.) injection, intraperitoneal (i.p.) injection, or intramuscular (i.m.) injection. One or more such routes can be used. Parenteral administration can be performed, for example, by bolus injection or by gradually perfusing over time. Alternatively, or simultaneously, administration can be by the oral route.

[0076] "Agent" means any small molecule compound, antibody, nucleic acid molecule, or polypeptide, or fragment thereof.

[0077] "Modification" means a change (e.g., an increase or decrease) in the structure, expression level, or activity of a gene or polypeptide, detected by standard methods known in the art such as those described herein. As used herein, a modification includes a change in the polynucleotide or polypeptide sequence or a change in the expression level, such as a 25% change, a 40% change, a 50% change, or a greater change.

[0078] "Ameliorate" means to reduce, suppress, attenuate, diminish, arrest, or stabilize the development or progression of a disease.

[0079] "Analog" means a molecule that has similar functional or structural features, although not identical. For example, a polynucleotide or polypeptide analog retains the biological activity of the corresponding naturally occurring polynucleotide or polypeptide while having certain modifications that enhance the analog's function compared to the naturally occurring polynucleotide or polypeptide. Such modifications can increase the affinity, efficiency, specificity, protease or nuclease resistance, membrane permeability, and / or half-life of the analog for DNA, for example, without modifying ligand binding. An analog can contain unnatural nucleotides or amino acids.

[0080] "Base editor (BE)" or "nucleic acid base editor (NBE)" means an agent that binds to a polynucleotide and has nucleic acid base modification activity. In various embodiments, a base editor comprises a nucleic acid base-modifying polypeptide (such as a deaminase) and a nucleic acid programmable nucleotide-binding domain in combination with a guide polynucleotide (such as guide RNA). In various embodiments, the agent is a biomolecular complex comprising a protein domain having base editing activity, i.e., a domain capable of modifying a base (such as A, T, C, G, U) within a nucleic acid molecule (such as DNA). In some embodiments, the polynucleotide programmable DNA-binding domain is fused or linked to a deaminase domain. In one embodiment, the agent is a fusion protein comprising a domain having base editing activity. In another embodiment, the protein domain having base editing activity is linked to the guide RNA (such as via an RNA-binding motif on the guide RNA and an RNA-binding domain fused to the deaminase). In some embodiments, the domain having base editing activity can deaminate a base within a nucleic acid molecule. In some embodiments, the base editor can deaminate one or more bases within a DNA molecule. In some embodiments, the base editor can deaminate adenosine (A) in DNA. In some embodiments, the base editor is an adenosine base editor (ABE).

[0081] As an example, the cytidine base editor (CBE) used in the base editing compositions, systems, and methods described herein has the following nucleic acid sequence (8877 base pairs) (Addgene, Watertown, MA.; Komor AC, et al., 2017, Sci Adv., 30;3(8):eaao4774. doi: 10.1126 / sciadv.aao4774). Also included are polynucleotide sequences having at least 95% or more identity to the BE4 nucleic acid sequence. 1 atatgccaag tacgccccct attgacgtca atgacggtaa atggcccgcc tggcattatg 61 cccagtacat gaccttatgg gactttccta cttggcagta catctacgta ttagtcatcg 121 ctattaccat ggtgatgcgg ttttggcagt acatcaatgg gcgtggatag cggtttgact 181 cacggggatt tccaagtctc caccccattg acgtcaatgg gagtttgttt tggcaccaaa 241 atcaacggga ctttccaaaa tgtcgtaaca actccgcccc attgacgcaa atgggcggta 301 ggcgtgtacg gtgggaggtc tatataagca gagctggttt agtgaaccgt cagatccgct 361 agagatccgc ggccgctaat acgactcact atagggagag ccgccaccat gagctcagag 421 actggcccag tggctgtgga ccccacattg agacggcgga tcgagcccca tgagtttgag 481 gtattcttcg atccgagaga gctccgcaag gagacctgcc tgctttacga aattaattgg 541 gggggccggc actccatttg gcgacataca tcacagaaca ctaacaagca cgtcgaagtc 601 aacttcatcg agaagttcac gacagaaaga tatttctgtc cgaacacaag gtgcagcatt 661 acctggtttc tcagctggag cccatgcggc gaatgtagta gggccatcac tgaattcctg 721 tcaaggtatc cccacgtcac tctgtttatt tacatcgcaa ggctgtacca ccacgctgac 781 ccccgcaatc gacaaggcct gcgggatttg atctcttcag gtgtgactat ccaaattatg 841 actgagcagg agtcaggata ctgctggaga aactttgtga attatagccc gagtaatgaa 901 gcccactggc ctaggtatcc ccatctgtgg gtacgactgt acgttcttga actgtactgc 961 atcatactgg gcctgcctcc ttgtctcaac attctgagaa ggaagcagcc acagctgaca 1021 ttctttacca tcgctcttca gtcttgtcat taccagcgac tgcccccaca cattctctgg 1081 gccaccgggt tgaaatctgg tggttcttct ggtggttcta gcggcagcga gactcccggg 1141 acctcagagt ccgccacacc cgaaagttct ggtggttctt ctggtggttc tgataaaaag 1201 tattctattg gtttagccat cggcactaat tccgttggat gggctgtcat aaccgatgaa 1261 tacaaagtac cttcaaagaa atttaaggtg ttggggaaca cagaccgtca ttcgattaaa 1321 aagaatctta tcggtgccct cctattcgat agtggcgaaa cggcagaggc gactcgcctg 1381 aaacgaaccg ctcggagaag gtatacacgt cgcaagaacc gaatatgtta cttacaagaa 1441 atttttagca atgagatggc caaagttgac gattctttct ttcaccgttt ggaagagtcc 1501 ttccttgtcg aagaggacaa gaaacatgaa cggcacccca tctttggaaa catagtagat 1561 gaggtggcat atcatgaaaa gtacccaacg atttatcacc tcagaaaaaa gctagttgac 1621 tcaactgata aagcggacct gaggttaatc tacttggctc ttgcccatat gataaagttc 1681 cgtgggcact ttctcattga gggtgatcta aatccggaca actcggatgt cgacaaactg 1741 ttcatccagt tagtacaaac ctataatcag ttgtttgaag agaaccctat aaatgcaagt 1801 ggcgtggatg cgaaggctat tcttagcgcc cgcctctcta aatcccgacg gctagaaaac 1861 ctgatcgcac aattacccgg agagaagaaa aatgggttgt tcggtaacct tatagcgctc 1921 tcactaggcc tgacaccaaa ttttaagtcg aacttcgact tagctgaaga tgccaaattg 1981 cagcttagta aggacacgta cgatgacgat ctcgacaatc tactggcaca aattggagat 2041 cagtatgcgg acttattttt ggctgccaaa aaccttagcg atgcaatcct cctatctgac 2101 atactgagag ttaatactga gattaccaag gcgccgttat ccgcttcaat gatcaaaagg 2161 tacgatgaac atcaccaaga cttgacactt ctcaaggccc tagtccgtca gcaactgcct 2221 gagaaatata aggaaatatt ctttgatcag tcgaaaaacg ggtacgcagg ttatattgac 2281 ggcggagcga gtcaagagga attctacaag tttatcaaac ccatattaga gaagatggat 2341 gggacggaag agttgcttgt aaaactcaat cgcgaagatc tactgcgaaa gcagcggact 2401 ttcgacaacg gtagcattcc acatcaaatc cacttaggcg aattgcatgc tatacttaga 2461 aggcaggagg atttttatcc gttcctcaaa gacaatcgtg aaaagattga gaaaatccta 2521 acctttcgca taccttacta tgtgggaccc ctggcccgag ggaactctcg gttcgcatgg 2581 atgacaagaa agtccgaaga aacgattact ccatggaatt ttgaggaagt tgtcgataaa 2641 ggtgcgtcag ctcaatcgtt catcgagagg atgaccaact ttgacaagaa tttaccgaac 2701 gaaaaagtat tgcctaagca cagtttactt tacgagtatt tcacagtgta caatgaactc 2761 acgaaagtta agtatgtcac tgagggcatg cgtaaacccg cctttctaag cggagaacag 2821 aagaaagcaa tagtagatct gttattcaag accaaccgca aagtgacagt taagcaattg 2881 aaagaggact actttaagaa aattgaatgc ttcgattctg tcgagatctc cggggtagaa 2941 gatcgattta atgcgtcact tggtacgtat catgacctcc taaagataat taaagataag 3001 gacttcctgg ataacgaaga gaatgaagat atcttagaag atatagtgtt gactcttacc 3061 ctctttgaag atcgggaaat gattgaggaa agactaaaaa catacgctca cctgttcgac 3121 gataaggtta tgaaacagtt aaagaggcgt cgctatacgg gctggggacg attgtcgcgg 3181 aaacttatca acgggataag agacaagcaa agtggtaaaa ctattctcga ttttctaaag 3241 agcgacggct tcgccaatag gaactttatg cagctgatcc atgatgactc tttaaccttc 3301 aaagaggata tacaaaaggc acaggtttcc ggacaagggg actcattgca cgaacatatt 3361 gcgaatcttg ctggttcgcc agccatcaaa aagggcatac tccagacagt caaagtagtg 3421 gatgagctag ttaaggtcat gggacgtcac aaaccggaaa acattgtaat cgagatggca 3481 cgcgaaaatc aaacgactca gaaggggcaa aaaaacagtc gagagcggat gaagagaata 3541 gaagagggta ttaaagaact gggcagccag atcttaaagg agcatcctgt ggaaaatacc 3601 caattgcaga acgagaaact ttacctctat tacctacaaa atggaaggga catgtatgtt 3661 gatcaggaac tggacataaa ccgtttatct gattacgacg tcgatcacat tgtaccccaa 3721 tcctttttga aggacgattc aatcgacaat aaagtgctta cacgctcgga taagaaccga 3781 gggaaaagtg acaatgttcc aagcgaggaa gtcgtaaaga aaatgaagaa ctattggcgg 3841 cagctcctaa atgcgaaact gataacgcaa agaaagttcg ataacttaac taaagctgag 3901 aggggtggct tgtctgaact tgacaaggcc ggatttatta aacgtcagct cgtggaaacc 3961 cgccaaatca caaagcatgt tgcacagata ctagattccc gaatgaatac gaaatacgac 4021 gagaacgata agctgattcg ggaagtcaaa gtaatcactt taaagtcaaa attggtgtcg 4081 gacttcagaa aggattttca attctataaa gttagggaga taaataacta ccaccatgcg 4141 cacgacgctt atcttaatgc cgtcgtaggg accgcactca ttaagaaata cccgaagcta 4201 gaaagtgagt ttgtgtatgg tgattacaaa gtttatgacg tccgtaagat gatcgcgaaa 4261 agcgaacagg agataggcaa ggctacagcc aaatacttct tttattctaa cattatgaat 4321 ttctttaaga cggaaatcac tctggcaaac ggagagatac gcaaacgacc tttaattgaa 4381 accaatgggg agacaggtga aatcgtatgg gataagggcc gggacttcgc gacggtgaga 4441 aaagttttgt ccatgcccca agtcaacata gtaaagaaaa ctgaggtgca gaccggaggg 4501 ttttcaaagg aatcgattct tccaaaaagg aatagtgata agctcatcgc tcgtaaaaag 4561 gactgggacc cgaaaaagta cggtggcttc gatagcccta cagttgccta ttctgtccta 4621 gtagtggcaa aagttgagaa gggaaaatcc aagaaactga agtcagtcaa agaattattg 4681 gggataacga ttatggagcg ctcgtctttt gaaaagaacc ccatcgactt ccttgaggcg 4741 aaaggttaca aggaagtaaa aaaggatctc ataattaaac taccaaagta tagtctgttt 4801 gagttagaaa atggccgaaa acggatgttg gctagcgccg gagagcttca aaaggggaac 4861 gaactcgcac taccgtctaa atacgtgaat ttcctgtatt tagcgtccca ttacgagaag 4921 ttgaaaggtt cacctgaaga taacgaacag aagcaacttt ttgttgagca gcacaaacat 4981 tatctcgacg aaatcataga gcaaatttcg gaattcagta agagagtcat cctagctgat 5041 gccaatctgg acaaagtatt aagcgcatac aacaagcaca gggataaacc catacgtgag 5101 caggcggaaa atattatcca tttgtttact cttaccaacc tcggcgctcc agccgcattc 5161 aagtattttg acacaacgat agatcgcaaa cgatacactt ctaccaagga ggtgctagac 5221 gcgacactga ttcaccaatc catcacggga ttatatgaaa ctcggataga tttgtcacag 5281 cttgggggtg actctggtgg ttctggagga tctggtggtt ctactaatct gtcagatatt 5341 attgaaaagg agaccggtaa gcaactggtt atccaggaat ccatcctcat gctcccagag 5401 gaggtggaag aagtcattgg gaacaagccg gaaagcgata tactcgtgca caccgcctac 5461 gacgagagca ccgacgagaa tgtcatgctt ctgactagcg acgcccctga atacaagcct 5521 tgggctctgg tcatacagga tagcaacggt gagaacaaga ttaagatgct ctctggtggt 5581 tctggaggat ctggtggttc tactaatctg tcagatatta ttgaaaagga gaccggtaag 5641 caactggtta tccaggaatc catcctcatg ctcccagagg aggtggaaga agtcattggg 5701 aacaagccgg aaagcgatat actcgtgcac accgcctacg acgagagcac cgacgagaat 5761 gtcatgcttc tgactagcga cgcccctgaa tacaagcctt gggctctggt catacaggat 5821 agcaacggtg agaacaagat taagatgctc tctggtggtt ctcccaagaa gaagaggaaa 5881 gtctaaccgg tcatcatcac catcaccatt gagtttaaac ccgctgatca gcctcgactg 5941 tgccttctag ttgccagcca tctgttgttt gcccctcccc cgtgccttcc ttgaccctgg 6001 aaggtgccac tcccactgtc ctttcctaat aaaatgagga aattgcatcg cattgtctga 6061 gtaggtgtca ttctattctg gggggtgggg tggggcagga cagcaagggg gaggattggg 6121 aagacaatag caggcatgct ggggatgcgg tgggctctat ggcttctgag gcggaaagaa 6181 ccagctgggg ctcgataccg tcgacctcta gctagagctt ggcgtaatca tggtcatagc 6241 tgtttcctgt gtgaaattgt tatccgctca caattccaca caacatacga gccggaagca 6301 taaagtgtaa agcctagggt gcctaatgag tgagctaact cacattaatt gcgttgcgct 6361 cactgcccgc tttccagtcg ggaaacctgt cgtgccagct gcattaatga atcggccaac 6421 gcgcggggag aggcggtttg cgtattgggc gctcttccgc ttcctcgctc actgactcgc 6481 tgcgctcggt cgttcggctg cggcgagcgg tatcagctca ctcaaaggcg gtaatacggt 6541 tatccacaga atcaggggat aacgcaggaa agaacatgtg agcaaaaggc cagcaaaagg 6601 ccaggaaccg taaaaaggcc gcgttgctgg cgtttttcca taggctccgc ccccctgacg 6661 agcatcacaa aaatcgacgc tcaagtcaga ggtggcgaaa cccgacagga ctataaagat 6721 accaggcgtt tccccctgga agctccctcg tgcgctctcc tgttccgacc ctgccgctta 6781 ccggatacct gtccgccttt ctcccttcgg gaagcgtggc gctttctcat agctcacgct 6841 gtaggtatct cagttcggtg taggtcgttc gctccaagct gggctgtgtg cacgaacccc 6901 ccgttcagcc cgaccgctgc gccttatccg gtaactatcg tcttgagtcc aacccggtaa 6961 gacacgactt atcgccactg gcagcagcca ctggtaacag gattagcaga gcgaggtatg 7021 taggcggtgc tacagagttc ttgaagtggt ggcctaacta cggctacact agaagaacag 7081 tatttggtat ctgcgctctg ctgaagccag ttaccttcgg aaaaagagtt ggtagctctt 7141 gatccggcaa acaaaccacc gctggtagcg gtggtttttt tgtttgcaag cagcagatta 7201 cgcgcagaaa aaaaggatct caagaagatc ctttgatctt ttctacgggg tctgacgctc 7261 agtggaacga aaactcacgt taagggattt tggtcatgag attatcaaaa aggatcttca 7321 cctagatcct tttaaattaa aaatgaagtt ttaaatcaat ctaaagtata tatgagtaaa 7381 cttggtctga cagttaccaa tgcttaatca gtgaggcacc tatctcagcg atctgtctat 7441 ttcgttcatc catagttgcc tgactccccg tcgtgtagat aactacgata cgggagggct 7501 taccatctgg ccccagtgct gcaatgatac cgcgagaccc acgctcaccg gctccagatt 7561 tatcagcaat aaaccagcca gccggaaggg ccgagcgcag aagtggtcct gcaactttat 7621 ccgcctccat ccagtctatt aattgttgcc gggaagctag agtaagtagt tcgccagtta 7681 atagtttgcg caacgttgtt gccattgcta caggcatcgt ggtgtcacgc tcgtcgtttg 7741 gtatggcttc attcagctcc ggttcccaac gatcaaggcg agttacatga tcccccatgt 7801 tgtgcaaaaa agcggttagc tccttcggtc ctccgatcgt tgtcagaagt aagttggccg 7861 cagtgttatc actcatggtt atggcagcac tgcataattc tcttactgtc atgccatccg 7921 taagatgctt ttctgtgact ggtgagtact caaccaagtc attctgagaa tagtgtatgc 7981 ggcgaccgag ttgctcttgc ccggcgtcaa tacgggataa taccgcgcca catagcagaa 8041 ctttaaaagt gctcatcatt ggaaaacgtt cttcggggcg aaaactctca aggatcttac 8101 cgctgttgag atccagttcg atgtaaccca ctcgtgcacc caactgatct tcagcatctt 8161 ttactttcac cagcgtttct gggtgagcaa aaacaggaag gcaaaatgcc gcaaaaaagg 8221 gaataagggc gacacggaaa tgttgaatac tcatactctt cctttttcaa tattattgaa 8281 gcatttatca gggttattgt ctcatgagcg gatacatatt tgaatgtatt tagaaaaata 8341 aacaaatagg ggttccgcgc acatttcccc gaaaagtgcc acctgacgtc gacggatcgg 8401 gagatcgatc tcccgatccc ctagggtcga ctctcagtac aatctgctct gatgccgcat 8461 agttaagcca gtatctgctc cctgcttgtg tgttggaggt cgctgagtag tgcgcgagca 8521 aaatttaagc tacaacaagg caaggcttga ccgacaattg catgaagaat ctgcttaggg 8581 ttaggcgttt tgcgctgctt cgcgatgtac gggccagata tacgcgttga cattgattat 8641 tgactagtta ttaatagtaa tcaattacgg ggtcattagt tcatagccca tatatggagt 8701 tccgcgttac ataacttacg gtaaatggcc cgcctggctg accgcccaac gacccccgcc 8761 cattgacgtc aataatgacg tatgttccca tagtaacgcc aatagggact ttccattgac 8821 gtcaatgggt ggagtattta cggtaaactg cccacttggc agtacatcaa gtgtatc

[0082] In some embodiments, the cytidine base editor is BE4 having a nucleic acid sequence selected from one of the following:

[0083] Original BE4 nucleic acid sequence:

[0084] BE4 codon-optimized 1 nucleic acid sequence:

[0085] BE4 codon-optimized 2 nucleic acid sequence:

[0086] In some embodiments, the base editor is generated by cloning an adenosine deaminase variant (e.g., TadA * 8) onto a scaffold comprising a circularly permuted Cas9 (e.g., spCAS9 or saCAS9) and a bipartite nuclear localization sequence (e.g., ABE8). Circularly permuted Cas9s are known in the art and are described, for example, in Oakes et al., Cell 176, 254-267, 2019. Exemplary circularly permuted Cas9s follow the convention where the bolded sequences indicate sequences derived from Cas9, the italicized sequences indicate linker sequences, and the underlined sequences indicate bipartite nuclear localization sequences. CP5 (MSP “NGC = Pam Variant with mutations Regular Cas9 likes NGG” PID = protein interaction domain and “D10A” nickase): TIFF0007693552000004.tif183162

[0087] In some embodiments, ABE8 is selected from base editors from Table 7, Table 9, Table 14, or Table 15 below. In some embodiments, ABE8 contains an adenosine deaminase variant evolved from TadA. In some embodiments, the adenosine deaminase variant of ABE8 is TadA * 8 variant as described in Table 7, Table 9, Table 14 or Table 15 below. In some embodiments, the adenosine deaminase variant comprises one or more of the modifications selected from the group of Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R TadA * 7.10 variant (e.g., TadA *8). In various embodiments, ABE8 has a combination of modifications selected from the group consisting of Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R of TadA * 7.10 variant (e.g., TadA * 8). In some embodiments, ABE8 is a monomeric construct. In some embodiments, ABE8 is a heterodimeric construct. In some embodiments, ABE8 has the sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD .

[0088] In some embodiments, the polynucleotide-programmable DNA binding domain is a CRISPR-associated (e.g., Cas or Cpf1) enzyme. In some embodiments, the base editor is catalytically inactive Cas9 (dCas9) fused to a deaminase domain. In some embodiments, the base editor is Cas9 nickase (nCas9) fused to a deaminase domain. Details of base editors are described in International PCT Application Nos. PCT / 2017 / 045381 (WO 2018 / 027078) and PCT / US2016 / 058344 (WO 2017 / 070632), each of which is incorporated herein by reference in its entirety. See also Komor, A.C., et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, N.M., et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, A.C., et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao4774 (2017), and Rees, H.A., et al., “Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec;19(12):770-788. doi: 10.1038 / s41576-018-0059-1 (the entire contents of which are incorporated herein by reference).

[0089] As an example, the adenine base editor (ABE) used in the base editing compositions, systems, and methods described herein has a nucleic acid sequence (8877 base pairs) as provided below (Addgene, Watertown, MA.; Gaudelli NM, et al., Nature. 2017 Nov 23;551(7681):464-471. doi: 10.1038 / nature24644; Koblan LW, et al., Nat Biotechnol. 2018 Oct;36(9):843-846. doi: 10.1038 / nbt.4172.). Polynucleotide sequences having at least 95% or higher identity to the ABE nucleic acid sequence are also included. ATATGCCAAGTACGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTATGCCCAGTACAT GACCTTATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATGGTGATGCGG TTTTGGCAGTACATCAATGGGCGTGGATAGCGGTTTGACTCACGGGGATTTCCAAGTCTCCACCCCATTG ACGTCAATGGGAGTTTGTTTTGGCACCAAAATCAACGGGACTTTCCAAAATGTCGTAACAACTCCGCCCC ATTGACGCAAATGGGCGGTAGGCGTGTACGGTGGGAGGTCTATATAAGCAGAGCTGGTTTAGTGAACCGT CAGATCCGCTAGAGATCCGCGGCCGCTAATACGACTCACTATAGGGAGAGCCGCCACCATGAAACGGACA GCCGACGGAAGCGAGTTCGAGTCACCAAAGAAGAAGCGGAAAGTCTCTGAAGTCGAGTTTAGCCACGAGT ATTGGATGAGGCACGCACTGACCCTGGCAAAGCGAGCATGGGATGAAAGAGAAGTCCCCGTGGGCGCCGT GCTGGTGCACAACAATAGAGTGATCGGAGAGGGATGGAACAGGCCAATCGGCCGCCACGACCCTACCGCA CACGCAGAGATCATGGCACTGAGGCAGGGAGGCCTGGTCATGCAGAATTACCGCCTGATCGATGCCACCC TGTATGTGACACTGGAGCCATGCGTGATGTGCGCAGGAGCAATGATCCACAGCAGGATCGGAAGAGTGGT GTTCGGAGCACGGGACGCCAAGACCGGCGCAGCAGGCTCCCTGATGGATGTGCTGCACCACCCCGGCATG AACCACCGGGTGGAGATCACAGAGGGAATCCTGGCAGACGAGTGCGCCGCCCTGCTGAGCGATTTCTTTA GAATGCGGAGACAGGAGATCAAGGCCCAGAAGAAGGCACAGAGCTCCACCGACTCTGGAGGATCTAGCGG AGGATCCTCTGGAAGCGAGACACCAGGCACAAGCGAGTCCGCCACACCAGAGAGCTCCGGCGGCTCCTCC GGAGGATCCTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGCCAAGAGGG CACGCGATGAGAGGGAGGTGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGCGAGGGCTG GAACAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATGGCCCTGAGACAGGGCGGCCTG GTCATGCAGAACTACAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCG GCGCCATGATCCACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACGCAAAAACCGGCGCCGCAGG CTCCCTGATGGACGTGCTGCACTACCCCGGCATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCA GATGAATGTGCCGCCCTGCTGTGCTATTTCTTTCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAAGG CCCAGAGCTCCACCGACTCCGGAGGATCTAGCGGAGGCTCCTCTGGCTCTGAGACACCTGGCACAAGCGA GAGCGCAACACCTGAAAGCAGCGGGGGCAGCAGCGGGGGGTCAGACAAGAAGTACAGCATCGGCCTGGCC ATCGGCACCAACTCTGTGGGCTGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGG TGCTGGGCAACACCGACCGGCACAGCATCAAGAAGAACCTGATCGGAGCCCTGCTGTTCGACAGCGGCGA AACAGCCGAGGCCACCCGGCTGAAGAGAACCGCCAGAAGAAGATACACCAGACGGAAGAACCGGATCTGC TATCTGCAAGAGATCTTCAGCAACGAGATGGCCAAGGTGGACGACAGCTTCTTCCACAGACTGGAAGAGT CCTTCCTGGTGGAAGAGGATAAGAAGCACGAGCGGCACCCCATCTTCGGCAACATCGTGGACGAGGTGGC CTACCACGAGAAGTACCCCACCATCTACCACCTGAGAAAGAAACTGGTGGACAGCACCGACAAGGCCGAC CTGCGGCTGATCTATCTGGCCCTGGCCCACATGATCAAGTTCCGGGGCCACTTCCTGATCGAGGGCGACC TGAACCCCGACAACAGCGACGTGGACAAGCTGTTCATCCAGCTGGTGCAGACCTACAACCAGCTGTTCGA GGAAAACCCCATCAACGCCAGCGGCGTGGACGCCAAGGCCATCCTGTCTGCCAGACTGAGCAAGAGCAGA CGGCTGGAAAATCTGATCGCCCAGCTGCCCGGCGAGAAGAAGAATGGCCTGTTCGGAAACCTGATTGCCC TGAGCCTGGGCCTGACCCCCAACTTCAAGAGCAACTTCGACCTGGCCGAGGATGCCAAACTGCAGCTGAG CAAGGACACCTACGACGACGACCTGGACAACCTGCTGGCCCAGATCGGCGACCAGTACGCCGACCTGTTT CTGGCCGCCAAGAACCTGTCCGACGCCATCCTGCTGAGCGACATCCTGAGAGTGAACACCGAGATCACCA AGGCCCCCCTGAGCGCCTCTATGATCAAGAGATACGACGAGCACCACCAGGACCTGACCCTGCTGAAAGC TCTCGTGCGGCAGCAGCTGCCTGAGAAGTACAAAGAGATTTTCTTCGACCAGAGCAAGAACGGCTACGCC GGCTACATTGACGGCGGAGCCAGCCAGGAAGAGTTCTACAAGTTCATCAAGCCCATCCTGGAAAAGATGG ACGGCACCGAGGAACTGCTCGTGAAGCTGAACAGAGAGGACCTGCTGCGGAAGCAGCGGACCTTCGACAA CGGCAGCATCCCCCACCAGATCCACCTGGGAGAGCTGCACGCCATTCTGCGGCGGCAGGAAGATTTTTAC CCATTCCTGAAGGACAACCGGGAAAAGATCGAGAAGATCCTGACCTTCCGCATCCCCTACTACGTGGGCC CTCTGGCCAGGGGAAACAGCAGATTCGCCTGGATGACCAGAAAGAGCGAGGAAACCATCACCCCCTGGAA CTTCGAGGAAGTGGTGGACAAGGGCGCTTCCGCCCAGAGCTTCATCGAGCGGATGACCAACTTCGATAAG AACCTGCCCAACGAGAAGGTGCTGCCCAAGCACAGCCTGCTGTACGAGTACTTCACCGTGTATAACGAGC TGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAAGCCCGCCTTCCTGAGCGGCGAGCAGAAAAAGGC CATCGTGGACCTGCTGTTCAAGACCAACCGGAAAGTGACCGTGAAGCAGCTGAAAGAGGACTACTTCAAG AAAATCGAGTGCTTCGACTCCGTGGAAATCTCCGGCGTGGAAGATCGGTTCAACGCCTCCCTGGGCACAT ACCACGATCTGCTGAAAATTATCAAGGACAAGGACTTCCTGGACAATGAGGAAAACGAGGACATTCTGGA AGATATCGTGCTGACCCTGACACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTGAAAACCTATGCC CACCTGTTCGACGACAAAGTGATGAAGCAGCTGAAGCGGCGGAGATACACCGGCTGGGGCAGGCTGAGCC GGAAGCTGATCAACGGCATCCGGGACAAGCAGTCCGGCAAGACAATCCTGGATTTCCTGAAGTCCGACGG CTTCGCCAACAGAAACTTCATGCAGCTGATCCACGACGACAGCCTGACCTTTAAAGAGGACATCCAGAAA GCCCAGGTGTCCGGCCAGGGCGATAGCCTGCACGAGCACATTGCCAATCTGGCCGGCAGCCCCGCCATTA AGAAGGGCATCCTGCAGACAGTGAAGGTGGTGGACGAGCTCGTGAAAGTGATGGGCCGGCACAAGCCCGA GAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCACCCAGAAGGGACAGAAGAACAGCCGCGAGAGA ATGAAGCGGATCGAAGAGGGCATCAAAGAGCTGGGCAGCCAGATCCTGAAAGAACACCCCGTGGAAAACA CCCAGCTGCAGAACGAGAAGCTGTACCTGTACTACCTGCAGAATGGGCGGGATATGTACGTGGACCAGGA ACTGGACATCAACCGGCTGTCCGACTACGATGTGGACCATATCGTGCCTCAGAGCTTTCTGAAGGACGAC TCCATCGACAACAAGGTGCTGACCAGAAGCGACAAGAACCGGGGCAAGAGCGACAACGTGCCCTCCGAAG AGGTCGTGAAGAAGATGAAGAACTACTGGCGGCAGCTGCTGAACGCCAAGCTGATTACCCAGAGAAAGTT CGACAATCTGACCAAGGCCGAGAGAGGCGGCCTGAGCGAACTGGATAAGGCCGGCTTCATCAAGAGACAG CTGGTGGAAACCCGGCAGATCACAAAGCACGTGGCACAGATCCTGGACTCCCGGATGAACACTAAGTACG ACGAGAATGACAAGCTGATCCGGGAAGTGAAAGTGATCACCCTGAAGTCCAAGCTGGTGTCCGATTTCCG GAAGGATTTCCAGTTTTACAAAGTGCGCGAGATCAACAACTACCACCACGCCCACGACGCCTACCTGAAC GCCGTCGTGGGAACCGCCCTGATCAAAAAGTACCCTAAGCTGGAAAGCGAGTTCGTGTACGGCGACTACA AGGTGTACGACGTGCGGAAGATGATCGCCAAGAGCGAGCAGGAAATCGGCAAGGCTACCGCCAAGTACTT CTTCTACAGCAACATCATGAACTTTTTCAAGACCGAGATTACCCTGGCCAACGGCGAGATCCGGAAGCGG CCTCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGTGGGATAAGGGCCGGGATTTTGCCACCGTGC GGAAAGTGCTGAGCATGCCCCAAGTGAATATCGTGAAAAAGACCGAGGTGCAGACAGGCGGCTTCAGCAA AGAGTCTATCCTGCCCAAGAGGAACAGCGATAAGCTGATCGCCAGAAAGAAGGACTGGGACCCTAAGAAG TACGGCGGCTTCGACAGCCCCACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAAGGGCAAGT CCAAGAAACTGAAGAGTGTGAAAGAGCTGCTGGGGATCACCATCATGGAAAGAAGCAGCTTCGAGAAGAA TCCCATCGACTTTCTGGAAGCCAAGGGCTACAAAGAAGTGAAAAAGGACCTGATCATCAAGCTGCCTAAG TACTCCCTGTTCGAGCTGGAAAACGGCCGGAAGAGAATGCTGGCCTCTGCCGGCGAACTGCAGAAGGGAA ACGAACTGGCCCTGCCCTCCAAATATGTGAACTTCCTGTACCTGGCCAGCCACTATGAGAAGCTGAAGGG CTCCCCCGAGGATAATGAGCAGAAACAGCTGTTTGTGGAACAGCACAAGCACTACCTGGACGAGATCATC GAGCAGATCAGCGAGTTCTCCAAGAGAGTGATCCTGGCCGACGCTAATCTGGACAAAGTGCTGTCCGCCT ACAACAAGCACCGGGATAAGCCCATCAGAGAGCAGGCCGAGAATATCATCCACCTGTTTACCCTGACCAA TCTGGGAGCCCCTGCCGCCTTCAAGTACTTTGACACCACCATCGACCGGAAGAGGTACACCAGCACCAAA GAGGTGCTGGACGCCACCCTGATCCACCAGAGCATCACCGGCCTGTACGAGACACGGATCGACCTGTCTC AGCTGGGAGGTGACTCTGGCGGCTCAAAAAGAACCGCCGACGGCAGCGAATTCGAGCCCAAGAAGAAGAG GAAAGTCTAACCGGTCATCATCACCATCACCATTGAGTTTAAACCCGCTGATCAGCCTCGACTGTGCCTT CTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCAC TGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGT GGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCT CTATGGCTTCTGAGGCGGAAAGAACCAGCTGGGGCTCGATACCGTCGACCTCTAGCTAGAGCTTGGCGTA ATCATGGTCATAGCTGTTTCCTGTGTGAAATTGTTATCCGCTCACAATTCCACACAACATACGAGCCGGA AGCATAAAGTGTAAAGCCTAGGGTGCCTAATGAGTGAGCTAACTCACATTAATTGCGTTGCGCTCACTGC CCGCTTTCCAGTCGGGAAACCTGTCGTGCCAGCTGCATTAATGAATCGGCCAACGCGCGGGGAGAGGCGG TTTGCGTATTGGGCGCTCTTCCGCTTCCTCGCTCACTGACTCGCTGCGCTCGGTCGTTCGGCTGCGGCGA GCGGTATCAGCTCACTCAAAGGCGGTAATACGGTTATCCACAGAATCAGGGGATAACGCAGGAAAGAACA TGTGAGCAAAAGGCCAGCAAAAGGCCAGGAACCGTAAAAAGGCCGCGTTGCTGGCGTTTTTCCATAGGCT CCGCCCCCCTGACGAGCATCACAAAAATCGACGCTCAAGTCAGAGGTGGCGAAACCCGACAGGACTATAA AGATACCAGGCGTTTCCCCCTGGAAGCTCCCTCGTGCGCTCTCCTGTTCCGACCCTGCCGCTTACCGGAT ACCTGTCCGCCTTTCTCCCTTCGGGAAGCGTGGCGCTTTCTCATAGCTCACGCTGTAGGTATCTCAGTTC GGTGTAGGTCGTTCGCTCCAAGCTGGGCTGTGTGCACGAACCCCCCGTTCAGCCCGACCGCTGCGCCTTA TCCGGTAACTATCGTCTTGAGTCCAACCCGGTAAGACACGACTTATCGCCACTGGCAGCAGCCACTGGTA ACAGGATTAGCAGAGCGAGGTATGTAGGCGGTGCTACAGAGTTCTTGAAGTGGTGGCCTAACTACGGCTA CACTAGAAGAACAGTATTTGGTATCTGCGCTCTGCTGAAGCCAGTTACCTTCGGAAAAAGAGTTGGTAGC TCTTGATCCGGCAAACAAACCACCGCTGGTAGCGGTGGTTTTTTTGTTTGCAAGCAGCAGATTACGCGCA GAAAAAAAGGATCTCAAGAAGATCCTTTGATCTTTTCTACGGGGTCTGACACTCAGTGGAACGAAAACTC ACGTTAAGGGATTTTGGTCATGAGATTATCAAAAAGGATCTTCACCTAGATCCTTTTAAATTAAAAATGA AGTTTTAAATCAATCTAAAGTATATATGAGTAAACTTGGTCTGACAGTTACCAATGCTTAATCAGTGAGG CACCTATCTCAGCGATCTGTCTATTTCGTTCATCCATAGTTGCCTGACTCCCCGTCGTGTAGATAACTAC GATACGGGAGGGCTTACCATCTGGCCCCAGTGCTGCAATGATACCGCGAGACCCACGCTCACCGGCTCCA GATTTATCAGCAATAAACCAGCCAGCCGGAAGGGCCGAGCGCAGAAGTGGTCCTGCAACTTTATCCGCCT CCATCCAGTCTATTAATTGTTGCCGGGAAGCTAGAGTAAGTAGTTCGCCAGTTAATAGTTTGCGCAACGT TGTTGCCATTGCTACAGGCATCGTGGTGTCACGCTCGTCGTTTGGTATGGCTTCATTCAGCTCCGGTTCC CAACGATCAAGGCGAGTTACATGATCCCCCATGTTGTGCAAAAAAGCGGTTAGCTCCTTCGGTCCTCCGA TCGTTGTCAGAAGTAAGTTGGCCGCAGTGTTATCACTCATGGTTATGGCAGCACTGCATAATTCTCTTAC TGTCATGCCATCCGTAAGATGCTTTTCTGTGACTGGTGAGTACTCAACCAAGTCATTCTGAGAATAGTGT ATGCGGCGACCGAGTTGCTCTTGCCCGGCGTCAATACGGGATAATACCGCGCCACATAGCAGAACTTTAA AAGTGCTCATCATTGGAAAACGTTCTTCGGGGCGAAAACTCTCAAGGATCTTACCGCTGTTGAGATCCAG TTCGATGTAACCCACTCGTGCACCCAACTGATCTTCAGCATCTTTTACTTTCACCAGCGTTTCTGGGTGA GCAAAAACAGGAAGGCAAAATGCCGCAAAAAAGGGAATAAGGGCGACACGGAAATGTTGAATACTCATAC TCTTCCTTTTTCAATATTATTGAAGCATTTATCAGGGTTATTGTCTCATGAGCGGATACATATTTGAATG TATTTAGAAAAATAAACAAATAGGGGTTCCGCGCACATTTCCCCGAAAAGTGCCACCTGACGTCGACGGA TCGGGAGATCGATCTCCCGATCCCCTAGGGTCGACTCTCAGTACAATCTGCTCTGATGCCGCATAGTTAA GCCAGTATCTGCTCCCTGCTTGTGTGTTGGAGGTCGCTGAGTAGTGCGCGAGCAAAATTTAAGCTACAAC AAGGCAAGGCTTGACCGACAATTGCATGAAGAATCTGCTTAGGGTTAGGCGTTTTGCGCTGCTTCGCGAT GTACGGGCCAGATATACGCGTTGACATTGATTATTGACTAGTTATTAATAGTAATCAATTACGGGGTCAT TAGTTCATAGCCCATATATGGAGTTCCGCGTTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCC CAACGACCCCCGCCCATTGACGTCAATAATGACGTATGTTCCCATAGTAACGCCAATAGGGACTTTCCAT TGACGTCAATGGGTGGAGTATTTACGGTAAACTGCCCACTTGGCAGTACATCAAGTGTATC

[0090] "Base editing activity" means acting to chemically modify a base within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is cytidine deaminase activity, such as the activity to convert a target C·G to T·A. In another embodiment, the base editing activity is cytidine deaminase activity, such as the activity to convert a target C·G to T·A, or adenosine or adenine deaminase activity, such as the activity to convert A·T to G·C (see PCT / US2019 / 044935, PCT / US2020 / 016288; each of these is hereby incorporated by reference in its entirety).

[0091] In some embodiments, base editing activity is evaluated by the efficiency of editing. Base editing efficiency can be measured by any suitable average, e.g., by Sanger sequencing or next-generation sequencing. In some embodiments, base editing efficiency is measured by the percentage of all sequencing reads having a nucleic acid base conversion effected by a base editor, e.g., the percentage of all sequencing reads having a target A·T base pair converted to a G·C base pair. In some embodiments, base editing efficiency is measured by the percentage of total cells having a nucleic acid base conversion effected by a base editor when the base editing is performed in a population of cells.

[0092] The term "base editor system" refers to a system for editing nucleic acid bases of a target nucleotide sequence. In various embodiments, a base editor system comprises: (1) a polynucleotide-programmable nucleotide binding domain (e.g., Cas9); (2) a deaminase domain for deaminating said nucleic acid base (e.g., adenosine deaminase and / or cytidine deaminase; see PCT / US2019 / 044935, PCT / US2020 / 016288, each of which is incorporated herein by reference in its entirety); and (3) one or more guide polynucleotides (e.g., guide RNA). In some embodiments, the polynucleotide-programmable nucleotide binding domain is a polynucleotide-programmable DNA binding domain. In some embodiments, the base editor is an adenine or adenosine base editor (ABE). In some embodiments, the base editor system is ABE8.

[0093] In some embodiments, a base editor system can include more than one base editing component. For example, a base editor system can include more than one deaminase. In some embodiments, a base editor system can include one or more adenosine deaminases. In some embodiments, single guide polynucleotides can be utilized to target different deaminases to a target nucleic acid sequence. In some embodiments, a single pair of guide polynucleotides can be utilized to target different deaminases to a target nucleic acid sequence.

[0094] The deaminase domain and polynucleotide-programmable nucleotide-binding component of the base editor system can be associated with each other covalently or non-covalently, or in any combination of these associations and interactions. For example, in some embodiments, the deaminase domain can be targeted to a target nucleotide sequence by a polynucleotide-programmable nucleotide-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain can be fused or linked to the deaminase domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain can target the deaminase domain to a target nucleotide sequence by interacting or associating non-covalently with the deaminase domain. For example, in some embodiments, the deaminase domain can interact with, associate with, or form a complex with an additional heterologous moiety or domain that is part of the polynucleotide-programmable nucleotide-binding domain. In some embodiments, the additional heterologous moiety can be one that can bind to, interact with, associate with, or form a complex with a polypeptide. In some embodiments, the additional heterologous moiety can be one that can bind to, interact with, associate with, or form a complex with a polynucleotide. In some embodiments, the additional heterologous moiety can be one that can bind to a guide polynucleotide. In some embodiments, the additional heterologous moiety can be one that can bind to a polypeptide linker. In some embodiments, the additional heterologous moiety can be one that can bind to a polynucleotide linker. The additional heterologous moiety can be a protein domain.In some embodiments, the additional heterologous moieties can be K homology (KH) domains, MS2 coat protein domains, PP7 coat protein domains, SfMu Com coat protein domains, sterile alpha motifs, telomerase Ku-binding motifs and Ku proteins, telomerase Sm7-binding motifs and Sm7 proteins, or RNA recognition motifs.

[0095] The base editor system may further comprise a guide polynucleotide component. It should be understood that the components of the base editor system may associate with each other via covalent bonds, non-covalent interactions, or any combination of these associations and interactions. In some embodiments, the deaminase domain may be targeted to a target nucleotide sequence by a guide polynucleotide. For example, in some embodiments, the deaminase domain may interact with, associate with, or form a complex with a portion or segment of the guide polynucleotide (e.g., a polynucleotide motif), or an additional heterologous portion or domain (e.g., a polynucleotide binding domain such as an RNA or DNA binding protein). In some embodiments, the additional heterologous portion or domain (e.g., a polynucleotide binding domain such as an RNA or DNA binding protein) may be fused or linked to the deaminase domain. In some embodiments, the additional heterologous portion may be one that can bind to, interact with, associate with, or form a complex with a polypeptide. In some embodiments, the additional heterologous portion may be one that can bind to, interact with, associate with, or form a complex with a polynucleotide. In some embodiments, the additional heterologous portion may be one that can bind to a guide polynucleotide. In some embodiments, the additional heterologous portion may be one that can bind to a polypeptide linker. In some embodiments, the additional heterologous portion may be one that can bind to a polynucleotide linker. The additional heterologous portion may be a protein domain. In some embodiments, the additional heterologous portion may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif.

[0096] In some embodiments, the base editor system can further include an inhibitor of a base excision repair (BER) component. It should be understood that the components of the base editor system can associate with each other via covalent bonds, non-covalent interactions, or any combination of these associations and interactions. The inhibitor of the BER component can include a BER inhibitor. In some embodiments, the inhibitor of BER can be a uracil DNA glycosylase inhibitor (UGI). In some embodiments, the inhibitor of BER can be an inosine BER inhibitor. In some embodiments, the inhibitor of BER can be targeted to a target nucleotide sequence by a polynucleotide-programmable nucleotide binding domain. In some embodiments, the polynucleotide-programmable nucleotide binding domain can be fused or linked to the inhibitor of BER. In some embodiments, the polynucleotide-programmable nucleotide binding domain can be fused or linked to a deaminase domain and an inhibitor of BER. In some embodiments, the polynucleotide-programmable nucleotide binding domain can target the inhibitor of BER to the target nucleotide sequence by non-covalently interacting with or associating with the inhibitor of BER. For example, in some embodiments, the inhibitor of the BER component can interact with, associate with, or form a complex with an additional heterologous moiety or domain that is part of the polynucleotide-programmable nucleotide binding domain.

[0097] In some embodiments, the BER inhibitor can be targeted to the target nucleotide sequence by a guide polynucleotide. For example, in some embodiments, the BER inhibitor can interact with, associate with, or form a complex with a portion or segment of the guide polynucleotide (e.g., a polynucleotide motif), an additional heterologous portion or domain (e.g., a polynucleotide binding domain such as an RNA or DNA binding protein). In some embodiments, an additional heterologous portion or domain of the guide polynucleotide (e.g., a polynucleotide binding domain such as an RNA or DNA binding protein) can be fused or linked to the BER inhibitor. In some embodiments, the additional heterologous portion can be one that can bind to, interact with, associate with, or form a complex with a polynucleotide. In some embodiments, the additional heterologous portion can be one that can bind to the guide polynucleotide. In some embodiments, the additional heterologous portion can be one that can bind to a polypeptide linker. In some embodiments, the additional heterologous portion can be one that can bind to a polynucleotide linker. The additional heterologous portion can be a protein domain. In some embodiments, the additional heterologous portion can be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif.

[0098] The term "Cas9" or "Cas9 domain" refers to an RNA-guided nuclease that includes the Cas9 protein or a fragment thereof (e.g., a protein containing the active, inactive, or partially active DNA cleavage domain of Cas9, and / or the gRNA binding domain of Cas9). The Cas9 nuclease may also be referred to as the casnl nuclease or the CRISPR (clustered regularly interspaced short palindromic repeat)-associated nuclease. CRISPR is an adaptive immune system that provides defense against mobile genetic elements (viruses, transposable elements, conjugative plasmids). The CRISPR cluster contains spacers, sequences complementary to the preceding mobile element, and the target invading nucleic acid. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, correct processing of pre-crRNA requires trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. TracrRNA serves as a guide for the processing of pre-crRNA by ribonuclease 3. Subsequently, Cas9 / crRNA / tracrRNA cleaves the linear or circular dsDNA target complementary to the spacer with an endonuclease. The target strand not complementary to the crRNA is first cleaved endonucleolytically and then trimmed 3'-5' exonucleolytically. Natively, DNA binding and cleavage typically require the protein and both RNAs. However, single guide RNA ("sgRNA" or simply "gRNA") can be engineered to incorporate both sides of crRNA and tracrRNA into a single RNA species. See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012) (the entire content of which is incorporated herein by reference). Cas9 recognizes a short motif (PAM or protospacer adjacent motif) in the CRISPR repeat sequence to help distinguish self from non-self.The sequences and structures of Cas9 nucleases are well known to those of ordinary skill in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C., Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma C.M., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Vogel J., Charpentier E., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in a variety of species including, but not limited to, S. pyogenes and S. thermophilus.Additional suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on the present disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737, the entire contents of which are incorporated herein by reference.

[0099] An exemplary Cas9 is Streptococcus pyogenes Cas9 (spCas9), and its amino acid sequence is provided below. TIFF0007693552000005.tif169162 (single underline: HNH domain; double underline: RuvC domain)

[0100] The nuclease-inactivated Cas9 protein can be alternatively referred to as the "dCas9" protein (meaning nuclease- "dead" Cas9) or catalytically inactive Cas9. Methods for generating a Cas9 protein (or a fragment thereof) having an inactive DNA cleavage domain are known (see, for example, Jinek et al, Science. 337:816-821 (2012); Qi et al, "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression" (2013) Cell. 28;152(5):1173-83 (each content is incorporated herein by reference)). For example, the DNA cleavage domain of Cas9 is known to include two subdomains: the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, and the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can suppress the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al, Science. 337:816-821 (2012); Qi et al, Cell. 28;152(5):1173-83 (2013)). In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is a nickase called the "nCas9" protein (meaning "nickase" Cas9). In some embodiments, a protein comprising a fragment of Cas9 is provided. For example, in some embodiments, the protein comprises one of the following two Cas9 domains: (1) the gRNA binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, a protein comprising Cas9 or a fragment thereof is referred to as a "Cas9 variant". The Cas9 variant shares homology with Cas9 or a fragment thereof.For example, a Cas9 variant has at least about 70% identity, at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity with wild-type Cas9. In some embodiments, the Cas9 variant can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type Cas9. In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain), and the fragment has at least about 70% identity, at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity with the corresponding fragment of wild-type Cas9. In some embodiments, the fragment is at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.5% the amino acid length of the corresponding wild-type Cas9.

[0101] In some embodiments, the fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids in length.

[0102] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_017053.1, nucleotide and amino acid sequences are as follows). TIFF0007693552000006.tif169162(Underline once: HNH domain; Underline twice: RuvC domain)

[0103] In some embodiments, wild-type Cas9 corresponds to, or comprises, the following nucleotide and / or amino acid sequences: TIFF0007693552000007.tif168163(Underline once: HNH domain; Underline twice: RuvC domain)

[0104] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_002737.2 (the following nucleotide sequence); and Uniprot reference sequence: Q99ZW2 (the following amino acid sequence). TIFF0007693552000008.tif169162 (SEQ ID NO: 1; single underline: HNH domain; double underline: RuvC domain)

[0105] In some embodiments, Cas9 refers to Cas9 from Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquisI (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref: YP_820832.1), Listeria innocua (NCBI Ref: NP_472073.1), Campylobacter jejuni (NCBI Ref: YP_002344900.1) or Neisseria meningitidis (NCBI Ref: YP_002342100.1), or Cas9 from any other organism.

[0106] In some embodiments, Cas9 is from Neisseria meningitidis (Nme). In some embodiments, Cas9 is Nme1, Nme2, or Nme3. In some embodiments, the PAM interaction domains of Nme1, Nme2, or Nme3 are N4GAT, N4CC, and N4CAAA, respectively (see, for example, Edraki, A., et al., A Compact, High-Accuracy Cas9 with a Dinucleotide PAM for In Vivo Genome Editing, Molecular Cell (2018)). The exemplary Neisseria meningitidis Cas9 protein, Nme1Cas9 (NCBI Reference: WP_002235162.1; type II CRISPR RNA-guided endonuclease Cas9) has the following amino acid sequence: 1 maafkpnpin yilgldigia svgwamveid edenpiclid lgvrvferae vpktgdslam 61 arrlarsvrr ltrrrahrll rarrllkreg vlqaadfden glikslpntp wqlraaaldr 121 kltplewsav llhlikhrgy lsqrkneget adkelgallk gvadnahalq tgdfrtpael 181 alnkfekesg hirnqrgdys htfsrkdlqa elillfekqk efgnphvsgg lkegietllm 241 tqrpalsgda vqkmlghctf epaepkaakn tytaerfiwl tklnnlrile qgserpltdt 301 eratlmdepy rkskltyaqa rkllgledta ffkglrygkd naeastlmem kayhaisral 361 ekeglkdkks plnlspelqd eigtafslfk tdeditgrlk driqpeilea llkhisfdkf 421 vqislkalrr ivplmeqgkr ydeacaeiyg dhygkkntee kiylppipad eirnpvvlra 481 lsqarkving vvrrygspar ihietarevg ksfkdrkeie krqeenrkdr ekaaakfrey 541 fpnfvgepks kdilklrlye qqhgkclysg keinlgrlne kgyveidhal pfsrtwddsf 601 nnkvlvlgse nqnkgnqtpy eyfngkdnsr ewqefkarve tsrfprskkq rillqkfded 661 gfkernlndt ryvnrflcqf vadrmrltgk gkkrvfasng qitnllrgfw glrkvraend 721 rhhaldavvv acstvamqqk itrfvrykem nafdgktidk etgevlhqkt hfpqpweffa 781 qevmirvfgk pdgkpefeea dtpeklrtll aeklssrpea vheyvtplfv srapnrkmsg 841 qghmetvksa krldegvsvl rvpltqlklk dlekmvnrer epklyealka rleahkddpa 901 kafaepfyky dkagnrtqqv kavrveqvqk tgvwvrnhng iadnatmvrv dvfekgdkyy 961 lvpiyswqva kgilpdravv qgkdeedwql iddsfnfkfs lhpndlvevi tkkarmfgyf 1021 aschrgtgni nirihdldhk igkngilegi gvktalsfqk yqidelgkei rpcrlkkrpp 1081 vr

[0107] Another exemplary Neisseria meningitidis Cas9 protein, Nme2Cas9 (NCBI Reference: WP_002230835; type II CRISPR RNA-guided endonuclease Cas9) has the following amino acid sequence: 1 maafkpnpin yilgldigia svgwamveid eeenpirlid lgvrvferae vpktgdslam 61 arrlarsvrr ltrrrahrll rarrllkreg vlqaadfden glikslpntp wqlraaaldr 121 kltplewsav llhlikhrgy lsqrkneget adkelgallk gvannahalq tgdfrtpael 181 alnkfekesg hirnqrgdys htfsrkdlqa elillfekqk efgnphvsgg lkegietllm 241 tqrpalsgda vqkmlghctf epaepkaakn tytaerfiwl tklnnlrile qgserpltdt 301 eratlmdepy rkskltyaqa rkllgledta ffkglrygkd naeastlmem kayhaisral 361 ekeglkdkks plnlsselqd eigtafslfk tdeditgrlk drvqpeilea llkhisfdkf 421 vqislkalrr ivplmeqgkr ydeacaeiyg dhygkkntee kiylppipad eirnpvvlra 481 lsqarkving vvrrygspar ihietarevg ksfkdrkeie krqeenrkdr ekaaakfrey 541 fpnfvgepks kdilklrlye qqhgkclysg keinlvrlne kgyveidhal pfsrtwddsf 601 nnkvlvlgse nqnkgnqtpy eyfngkdnsr ewqefkarve tsrfprskkq rillqkfded 661 gfkecnlndt ryvnrflcqf vadhilltgk gkrrvfasng qitnllrgfw glrkvraend 721 rhhaldavvv acstvamqqk itrfvrykem nafdgktidk etgkvlhqkt hfpqpweffa 781 qevmirvfgk pdgkpefeea dtpeklrtll aeklssrpea vheyvtplfv srapnrkmsg 841 ahkdtlrsak rfvkhnekis vkrvwlteik ladlenmvny kngreielye alkarleayg 901 gnakqafdpk dnpfykkggq lvkavrvekt qesgvllnkk naytiadngd mvrvdvfckv 961 dkkgknqyfi vpiyawqvae nilpdidckg yriddsytfc fslhkydlia fqkdekskve 1021 fayyincdss ngrfylawhd kgskeqqfri stqnlvliqk yqvnelgkei rpcrlkkrpp 1081 vr

[0108] In some embodiments, dCas9 corresponds to a Cas9 amino acid sequence having one or more mutations that inactivate Cas9 nuclease activity, or comprises a part or all thereof. For example, in some embodiments, the dCas9 domain comprises the D10A and H840A mutations or corresponding mutations in another Cas9. In some embodiments, dCas9 comprises the following amino acid sequence of dCas9 (D10A and H840A): TIFF0007693552000009.tif168162 (single underline: HNH domain; double underline: RuvC domain).

[0109] In some embodiments, the Cas9 domain contains a D10A mutation, while the residue at position 840 in the amino acid sequence provided above, or the residue at the corresponding position in any of the amino acid sequences provided herein, remains histidine.

[0110] In other embodiments, dCas9 variants are provided that have mutations other than D10A and H840A, which result in, for example, nuclease-inactivated Cas9 (dCas9). Such mutations include, for example, other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). In some embodiments, variants or homologs of dCas9 are provided that have at least about 70% identity, at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity. In some embodiments, variants of dCas9 are provided that have an amino acid sequence that is shorter or longer by about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids, or more.

[0111] In some embodiments, the Cas9 fusion proteins provided herein include the full-length amino acid sequence of the Cas9 protein, for example, one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein do not include the full-length Cas9 sequence and include only one or more fragments thereof. Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and additional suitable sequences of Cas9 domains and fragments will be apparent to those of skill in the art. Additional Cas9 proteins (e.g., nuclease-dead Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9), including variants and homologs thereof, are to be understood to be within the scope of the present disclosure. Exemplary Cas9 proteins include, but are not limited to, those provided below. In some embodiments, the Cas9 protein is nuclease-dead Cas9 (dCas9). In some embodiments, the Cas9 protein is Cas9 nickase (nCas9). In some embodiments, the Cas9 protein is nuclease-active Cas9.

[0112] Exemplary catalytically inactive Cas9 (dCas9):

[0113] Exemplary catalytic Cas9 nickase (nCas9):

[0114] Exemplary catalytically active Cas9:

[0115] In some embodiments, Cas9 refers to Cas9 from archaea (e.g., Nanoarchaea) that constitute the domain and kingdom of single-celled prokaryotic microorganisms. In some embodiments, Cas9 refers to, for example, CasX or CasY as described in Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res. 2017 Feb 21. doi:10.1038 / cr.2017.21, the entire content of which is incorporated herein by reference. Using genomic resolution metagenomics, many CRISPR-Cas systems have been identified, including Cas9 first reported in the archaeal domain of life. This divergent Cas9 protein was discovered as part of an active CRISPR-Cas system in the little-studied Nanoarchaea. In bacteria, two previously unknown systems, CRISPR-CasX and CRISPR-CasY, were discovered, and they belong to the most compact systems discovered so far. In some embodiments, Cas9 represents CasX or a variant of CasX. In some embodiments, Cas9 represents CasY or a variant of CasY. Other RNA-guided DNA-binding proteins can also be used as nucleic acid programmable DNA binding proteins (napDNAbp), and it should be understood that they are within the scope of the present disclosure.

[0116] In certain embodiments, napDNAbps useful in the methods of the present disclosure are known in the art and include, for example, the circular permutants described in Oakes et al., Cell 176, 254-267, 2019. Exemplary circular permutants follow the convention where the bold sequences represent sequences derived from Cas9, the italic sequences represent linker sequences, and the underlined sequences represent bipartite nuclear localization sequences.

[0117] CP5 (MSP 「NGC = Pam Variant with mutations Regular Cas9 likes NGG」, PID = Protein Interaction Domain and having 「D10A」 Nickase): TIFF0007693552000010.tif185162

[0118] Non-limiting examples of polynucleotide programmable nucleotide binding domains that can be incorporated into base editors include domains derived from CRISPR proteins, restriction nucleases, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases (ZFNs).

[0119] In some embodiments, any nucleic acid programmable DNA binding protein (napDNAbp) of the fusion proteins provided herein can be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the napDNAbp is a CasY protein. In some embodiments, the napDNAbp comprises an amino acid sequence having at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to a naturally occurring CasX or CasY protein. In some embodiments, the napDNAbp is a naturally occurring CasX or CasY protein. In some embodiments, the napDNAbp comprises an amino acid sequence having at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to any of the CasX or CasY proteins described herein. It should be understood that Cas12b / C2c1, CasX, and CasY from other bacterial species can also be used in accordance with the present disclosure.

[0120] Cas12b / C2c1 (uniprot.org / uniprot / T0D7A2#2) sp|T0D7A2|C2C1_ALIAG CRISPR-associated endo- nuclease C2c1 OS = Alicyclobacillus acido-terrestris (strain ATCC 49025 / DSM 3922 / CIP 106132 / NCIMB 13137 / GD3B) GN=c2c1 PE=1 SV=1

[0121] CasX (uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) >tr|F0NN87|F0NN87_SULIH CRISPR-associated Casx protein OS = Sulfolobus islandicus (strain HVE10 / 4) GN = SiH_0402 PE=4 SV=1 MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAVGQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG

[0122] >tr|F0NH53|F0NH53_SULIR CRISPR associated protein, Casx OS = Sulfolobus islandicus (strain REY15A) GN=SiRe_0771 PE=4 SV=1 MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAVGQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG

[0123] Deltaproteobacteria CasX MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMTDDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAILQVYWQEFKDDHVGLMCKFAQPASKKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFVYKLEQVSEKGKAYTNYFGRCNVAEHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESTHPVKPLAQIAGNRYASGPVGKALSDACMGTIASFLSKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDAYNEVIARVRMWVNLNLWQKLKLSRDDAKPLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENPKKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEMDEKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFTDGTDIKKSGKWQGLLYGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKIGRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQAAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAKLAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKELSAELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSYKSGKQPFVGAWQAFYKRRLKEVWKPNA

[0124] CasY (ncbi.nlm.nih.gov / protein / APG80656.1) >APG80656.1 CRISPR-associated protein CasY [uncultured Parcubacteria group bacterium]

[0125] The term "Cas12" or "Cas12 domain" refers to an RNA-guided nuclease that includes a Cas12 protein or a fragment thereof (e.g., a protein comprising an active, inactive, or partially active DNA cleavage domain of Cas12, and / or a gRNA binding domain of Cas12). Cas12 belongs to the Class 2, Type V CRISPR / Cas system. Cas12 nucleases are sometimes referred to as CRISPR (clustered regularly interspaced short palindromic repeat)-associated nucleases. The sequence of an exemplary Bacillus hisashii Cas 12b (BhCas12b) Cas 12 domain is provided below:

[0126] Amino acid sequences having at least 85% or higher identity to the BhCas12b amino acid sequence are also useful in the methods of the present disclosure.

[0127] The terms "conservative amino acid substitution" or "conservative mutation" refer to the replacement of one amino acid with another amino acid having common characteristics. A functional way to define the common properties between individual amino acids is to analyze the normalized frequency of amino acid changes between corresponding proteins of homologous organisms (Schulz, G. E. and Schirmer, R. H., Principles of Protein Structure, Springer-Verlag, New York (1979)). According to such an analysis, groups of amino acids can be defined in which the amino acids within the group are preferentially exchanged with each other and are therefore most similar to each other in their effects on the overall protein structure (Schulz, G. E. and Schirmer, R. H. supra). Non-limiting examples of conservative mutations include, for example, the substitution of the amino acid arginine, which can maintain a positive charge, with lysine and vice versa; the substitution of aspartic acid, which can maintain a negative charge, with glutamic acid and vice versa; the substitution of threonine, in which the free -OH is maintained, with serine; and the substitution of asparagine, which can maintain a free NH2, with glutamine.

[0128] As used interchangeably herein, the term "coding sequence" or "protein coding sequence" refers to a segment of a polynucleotide that encodes a protein. This region or sequence is bounded by a start codon closer to the 5' end and a stop codon closer to the 3' end. The coding sequence is also referred to as an open reading frame.

[0129] "Cytidine deaminase" means a polypeptide or a fragment thereof that can catalyze a deamination reaction that converts an amino group to a carbonyl group. In one embodiment, cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine. PmCDA1 (Petromyzon marinus cytosine deaminase 1) from Petromyzon marinus, or AID (activation-induced cytidine deaminase; AICDA) from mammals (e.g., human, pig, cow, horse, monkey, etc.), and APOBEC are exemplary cytidine deaminases.

[0130] As used herein, the terms "deaminase" or "deaminase domain" refer to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or adenine (A) to inosine (I). In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., genetically engineered adenosine deaminases, evolved adenosine deaminases) can be from any organism such as bacteria. In some embodiments, the adenosine deaminase is from a bacterium such as Escherichia coli, Staphylococcus aureus, Salmonella typhimurium, Shewanella putrefaciens, Haemophilus influenzae, or Caulobacter crescentus.

[0131] In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is a TadA variant. In some embodiments, the TadA variant is TadA *It is 8. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism such as a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain does not occur naturally. For example, in some embodiments, the deaminase or deaminase domain has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identity to a naturally occurring deaminase. For example, the deaminase domain is described in International PCT Application Nos. PCT / 2017 / 045381 (WO 2018 / 027078) and PCT / US2016 / 058344 (WO 2017 / 070632), each of which is incorporated herein by reference in its entirety.Komor, A.C., et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, N.M., et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, A.C., et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity”Science Advances 3:eaao4774 (2017) ), and Rees, H.A., et al., “Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec;19(12):770-788. doi: 10.1038 / s41576-018-0059-1 are also incorporated by reference in their entirety (all of these are incorporated herein by reference).

[0132] "Detecting" refers to identifying the presence, absence, or amount of an analyte to be detected. In one embodiment, a sequence change in a polynucleotide or polypeptide is detected. In another embodiment, the presence of an indel is detected.

[0133] "Detectable label" means a composition that, when linked to a molecule of interest, makes the latter detectable via spectroscopic, photochemical, biochemical, immunochemical, or chemical means. For example, useful labels include radioisotopes, magnetic beads, metal beads, colloidal particles, fluorescent dyes, high electron density reagents, enzymes (such as those commonly used in ELISA), biotin, digoxigenin, or haptens.

[0134] "Disease" means any condition or disorder that impairs or interferes with the normal function of cells, tissues, or organs.

[0135] As used herein, the term "effective amount" refers to the amount of a bioactive agent sufficient to induce a desired biological response. The effective amount of the active agent used to practice the present disclosure for the therapeutic treatment of a disease will vary depending on the mode of administration, the age, weight, and general health of the subject. Ultimately, the attending physician or veterinarian will determine the appropriate amount and dosage. Such an amount is referred to as an "effective" amount. In one embodiment, the effective amount is the amount of the base editor of the present disclosure (e.g., a fusion protein comprising a programmable DNA-binding protein, a nucleic acid base editor, and a gRNA) sufficient to introduce a modification into a gene of interest in a cell (e.g., an in vitro or in vivo cell). In one embodiment, the effective amount is the amount of the base editor necessary to achieve a therapeutic effect (e.g., to reduce or control a disease or its symptoms or condition). Such a therapeutic effect need not be sufficient to modify the gene of interest in all cells, tissues, or organs of the subject, but may only modify the gene of interest in about 1%, 5%, 10%, 25%, 50%, 75%, or more of the cells present in the subject, tissue, or organ.

[0136] "Fragment" means a portion of a polypeptide or nucleic acid molecule. This portion comprises at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the full length of the reference nucleic acid molecule or polypeptide. The fragment can comprise 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids.

[0137] "Guide RNA" or "gRNA" means a polynucleotide that can form a complex with a polynucleotide-programmable nucleotide-binding domain protein (e.g., Cas9 or Cpf1) that can be specific for a target sequence. In certain embodiments, the guide polynucleotide is a guide RNA (gRNA). The gRNA may exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule may sometimes be referred to as a single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to guide RNAs that exist as a single molecule or as a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species includes two domains: (1) a domain that shares homology with the target nucleic acid (e.g., that directs binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to tracrRNA as provided in Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those including domain 2) can be found in U.S. Provisional Patent Application No. 61 / 874,682, filed Sep. 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application No. 61 / 874,746, filed Sep. 6, 2013, entitled "Delivery System For Functional Nucleases", the entire contents of each of which are incorporated herein by reference. In some embodiments, the gRNA includes two or more of domains (1) and (2) and may be referred to as an "extended gRNA." The extended gRNA binds to two or more Cas9 proteins and binds to the target nucleic acid in two or more different regions, as described herein.The gRNA contains a nucleotide sequence complementary to the target site, which mediates the binding of the nuclease / RNA complex to the target site and provides the sequence specificity of the nuclease:RNA complex.

[0138] The term "heterodimer" refers to a fusion protein containing two domains such as a wild-type TadA domain and a variant of the TadA domain (e.g., TadA * 8) or two variant TadA domains (e.g., TadA * 7.10 and TadA * 8 or two TadA * 8 domains).

[0139] "Hybridization" means hydrogen bonding between complementary nucleobases and can be Watson-Crick, Hoogsteen or reverse Hoogsteen hydrogen bonds. For example, adenine and thymine are complementary nucleobases that form hydrogen bonds to form a pair.

[0140] The term "inhibitor of base repair" or "IBR" refers to a protein that can inhibit the activity of a nucleic acid repair enzyme, such as a base excision repair (BER) enzyme. In certain embodiments, the IBR is an inhibitor of inosine base excision repair. Examples of inhibitors of base repair include inhibitors of APE1, Endo III, Endo IV, Endo V, Endo VIII, Fpg, hOGGl, hNEILl, T7 Endol, T4 PDG, UDG, hSMUGL and hAAG. In some embodiments, the IBR is an inhibitor of Endo V or hAAG. In some embodiments, the IBR is catalytically inactive EndoV or catalytically inactive hAAG. In some embodiments, the base repair inhibitor is an inhibitor of Endo V or hAAG. In some embodiments, the base repair inhibitor is catalytically inactive EndoV or catalytically inactive hAAG.

[0141] In some embodiments, the base repair inhibitor is uracil glycosylase inhibitor (UGI). UGI refers to a protein that can inhibit the base excision repair enzyme of uracil-DNA glycosylase. In some embodiments, the UGI domain comprises wild-type UGI or a fragment thereof. In some embodiments, the UGI proteins provided herein include fragments of UGI and proteins homologous to UGI or UGI fragments. In some embodiments, the base repair inhibitor is an inhibitor of inosine base excision repair. In some embodiments, the base repair inhibitor is a "catalytically inactive inosine-specific nuclease" or a "dead inosine-specific nuclease". Without wishing to be bound by any particular theory, a catalytically inactive inosine glycosylase (e.g., alkyladenine glycosylase (AAG)) can bind to inosine but cannot create an abasic site or remove inosine, thereby sterically blocking newly formed inosine moieties from the DNA damage / repair machinery. In some embodiments, the catalytically inactive inosine-specific nuclease can bind to inosine in a nucleic acid but does not cleave the nucleic acid. Non-limiting examples of representative catalytically inactive inosine-specific nucleases include catalytically inactive alkyladenine glycosylase (AAG nuclease) (e.g., of human origin), and catalytically inactive endonuclease V (EndoV nuclease) (e.g., of E. coli origin). In some embodiments, the catalytically inactive AAG nuclease comprises the E125Q mutation or a corresponding mutation in another AAG nuclease.

[0142] "Increase" means a positive change of at least 10%, 25%, 50%, 75%, or 100%.

[0143] An "intein" is a protein fragment that can excise itself and ligate the remaining fragments (exteins) by peptide bonds in a process known as protein splicing. An intein is also referred to as a "protein intron". The process by which an intein excises itself and ligates the remaining part of the protein is referred to herein as "protein splicing" or "intein-mediated protein splicing". In some embodiments, the intein of a precursor protein (the intein-containing protein prior to intein-mediated protein splicing) is derived from two genes. Such inteins are referred to herein as split inteins (e.g., split intein-N and split intein-C). For example, in cyanobacteria, DnaE, which is the catalytic subunit a of DNA polymerase III, is encoded by two separate genes, dnaE-n and dnaE-c. The intein encoded by the dnaE-n gene may be referred to herein as "intein N". The intein encoded by the dnaE-c gene may be referred to herein as "intein C".

[0144] Other intein systems can also be used. As an example, synthetic inteins based on the dnaE intein, i.e., the intein pair of Cfa-N (e.g., split intein-N) and Cfa-C (e.g., split intein-C), have been described (e.g., Stevens et al., J Am Chem Soc. 2016 Feb. 24; 138(7):2162-5, which is incorporated herein by reference). Non-limiting examples of intein pairs that can be used in accordance with the present disclosure include the Cfa DnaE intein, the Ssp GyrB intein, the Ssp DnaX intein, the Ter DnaE3 intein, the Ter ThyX intein, the Rma DnaB intein, and the Cne Prp8 intein (e.g., as described in U.S. Patent No. 8,394,604, which is incorporated herein by reference).

[0145] Provide exemplary nucleotide and amino acid sequences of inteins. DnaE intein-N DNA: TGCCTGTCATACGAAACCGAGATACTGACAGTAGAATATGGCCTTCTGCCAATCGGGAAGATTGTGGAGAAACGGATAGAATGCACAGTTTACTCTGTCGATAACAATGGTAACATTTATACTCAGCCAGTTGCCCAGTGGCACGACCGGGGAGAGCAGGAAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAGGGCCACTAAGGACCACAAATTTATGACAGTCGATGGCCAGATGCTGCCTATAGACGAAATCTTTGAGCGAGAGTTGGACCTCATGCGAGTTGACAACCTTCCTAAT DnaE intein-N protein: CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDR GEQEVFEYCLEDGSLIRATKDHKFMTVDGQMLPIDEIFERELDLMRVDNLPN DnaE intein-C DNA: ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGA TATTGGAGTCGAAAGAGATCACAACTTTGCTCTGAAGAACGGATTCATAGCTTCTAAT Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN Cfa-N DNA: TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGCCTATTGGAAAGATTGTCGAAGAGAGAATTGAATGCACAGTATATACTGTAGACAAGAATGGTTTCGTTTACACACAGCCCATTGCTCAATGGCACAATCGCGGCGAACAAGAAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACGAGCAACTAAAGATCATAAATTCATGACCACTGACGGGCAGATGTTGCCAATAGATGAGATATTCGAGCGGGGCTTGGATCTCAAACAAGTGGATGGATTGCCA Cfa-N protein: CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQMLPIDEIFERGLDLKQVDGLP Cfa-C DNA: ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAGGAAAGTAAAGATAATATCTCGAAAAAGTCTTGGTACCCAAAATGTCTATGATATTGGAGTGGAGAAAGATCACAACTTCCTTCTCAAGAACGGTCTCGTAGCCAGCAAC Cfa-C protein: MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLVASN

[0146] To join the N-terminal portion of split Cas9 and the C-terminal portion of split Cas9, intein N and intein C can be fused to the N-terminal portion of split Cas9 and the C-terminal portion of split Cas9, respectively. For example, in some embodiments, intein-N is fused to the C-terminus of the N-terminal portion of split Cas9, i.e., forming a structure of N--[N-terminal portion of split Cas9]-[intein-N]--C. In some embodiments, intein-C is fused to the N-terminus of the C-terminal portion of split Cas9, i.e., forming a structure of N-[intein-C]--[C-terminal portion of split Cas9]-C. The mechanism of intein-mediated protein splicing for ligating the proteins (such as split Cas9) where the intein is fused is known in the art, as described, for example, in Shah et al., Chem Sci. 2014; 5(1):446-461, which is hereby incorporated by reference in its entirety. Methods for designing and using inteins are known in the art and are described, for example, by WO2014004336, WO2017132580, US20150344549, and US20180127780, each of which is hereby incorporated by reference in its entirety

[0147] The terms "isolated," "purified," or "biologically pure" refer to a substance from which the components that are normally associated with it in its natural state have been removed to varying degrees. "Isolation" indicates the degree of separation from the original source or surrounding environment. "Purification" indicates a higher degree of separation than isolation. A "purified" or "biologically pure" protein has other substances sufficiently removed so that impurities do not substantially affect the biological properties of the protein or cause other adverse results. That is, the nucleic acids or peptides of the present disclosure are purified when they are produced by recombinant DNA technology and substantially free of cellular material, viral material, or medium, or when chemically synthesized and substantially free of chemical precursors or other chemicals. Purity and homogeneity are typically determined using analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high performance liquid chromatography. The term "purified" can mean that a nucleic acid or protein gives rise to essentially one band on an electrophoretic gel. For proteins that can undergo modifications, such as phosphorylation or glycosylation, different modifications can give rise to different isolated proteins that can be purified separately.

[0148] "Isolated polynucleotide" means a nucleic acid (e.g., DNA) that does not contain the genes adjacent to the gene in the natural genome of the organism from which the nucleic acid molecule of the present disclosure is derived. Thus, the term includes, for example, recombinant DNA incorporated into a vector; incorporated into an autonomously replicating plasmid or virus; incorporated into the genomic DNA of a prokaryote or eukaryote; or existing as a separate molecule independent of other sequences (e.g., cDNA or genomic or cDNA fragments generated by PCR or restriction endonuclease digestion). Further, the term includes RNA molecules transcribed from DNA molecules, as well as recombinant DNA that is part of a hybrid gene encoding additional polypeptide sequences.

[0149] "Isolated polypeptide" means a polypeptide of the present disclosure that has been separated from the components naturally associated therewith. Typically, a polypeptide is isolated when it does not have at least 60% by weight of the proteins and naturally occurring organic molecules with which it is naturally associated. Preferably, the preparation is at least 75%, more preferably at least 90%, and most preferably at least 99% by weight of the polypeptide of the present disclosure. The isolated polypeptides of the present disclosure can be obtained, for example, by extraction from a natural source, by expression of a recombinant nucleic acid encoding such a polypeptide, or by chemically synthesizing the protein. Purity can be measured by any suitable method, such as column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.

[0150] As used herein, the term "linker" can refer to two molecules or moieties, such as two components of a protein complex or ribonucleo complex, or two domains of a fusion protein, such as a polynucleotide programmable DNA binding domain (e.g., dCas9) and a deaminase domain (e.g., adenosine deaminase, cytidine deaminase, or a covalent linker (e.g., a covalent bond) that links adenosine deaminase and cytidine deaminase), a non-covalent linker, a chemical group, or a molecule; see PCT / US2019 / 044935, PCT / US2020 / 016288, each of which is hereby incorporated by reference in its entirety). A linker can connect different components or different parts of a base editor system. For example, in some embodiments, a linker can connect the guide polynucleotide binding domain of a polynucleotide programmable nucleotide binding domain and the catalytic domain of a deaminase. In some embodiments, a linker can bind a CRISPR polypeptide and a deaminase. In some embodiments, a linker can connect Cas9 and a deaminase. In some embodiments, a linker can connect dCas9 and a deaminase. In some embodiments, a linker can connect nCas9 and a deaminase. In some embodiments, a linker can connect a guide polynucleotide and a deaminase. In some embodiments, a linker can connect the deaminating component of a base editor system and a polynucleotide programmable nucleotide binding component. In some embodiments, a linker can connect the RNA binding portion of the deaminating component of a base editor system with a polynucleotide programmable nucleotide binding component. In some embodiments, a linker can bind the RNA binding portion of the deaminating component of a base editor system to the RNA binding portion of a polynucleotide programmable nucleotide binding component.A linker is disposed between, or sandwiched by, two groups, molecules, or other moieties, and is linked to each via covalent or non-covalent interaction, thus enabling the linking of the two. In some embodiments, the linker can be an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker can be a polynucleotide. In some embodiments, the linker can be a DNA linker. In some embodiments, the linker can be an RNA linker. In some embodiments, the linker can include an aptamer that can bind to a ligand. In some embodiments, the ligand can be a carbohydrate, peptide, protein, or nucleic acid. In some embodiments, the linker can include an aptamer derived from a riboswitch. The riboswitch from which the aptamer is derived can be selected from a theophylline riboswitch, thiamine pyrophosphate (TPP) riboswitch, adenosylcobalamin (AdoCbl) riboswitch, S-adenosylmethionine (SAM) riboswitch, SAH riboswitch, flavin mononucleotide (FMN) riboswitch, tetrahydrofolate riboswitch, lysine riboswitch, glycine riboswitch, purine riboswitch, GlmS riboswitch, or prequeuosine 1 (PreQ1) riboswitch. In some embodiments, the linker can include an aptamer bound to a protein domain such as a polypeptide or polypeptide ligand. In some embodiments, the polypeptide ligand can be a K homology (KH) domain, MS2 coat protein domain, PP7 coat protein domain, SfMu Com coat protein domain, sterile alpha motif, telomerase Ku binding motif and Ku protein, telomerase Sm7 binding motif and Sm7 protein, or RNA recognition motif. In some embodiments, the polypeptide ligand can be part of a base editor system component. For example, a nucleic acid base editing component can include a deaminase domain and an RNA recognition motif.

[0151] In some embodiments, the linker can be an amino acid or multiple amino acids (e.g., a peptide or a protein). In some embodiments, the linker can be about 5 to 100 amino acids in length, e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20 - 30, 30 - 40, 40 - 50, 50 - 60, 60 - 70, 70 - 80, 80 - 90, or 90 - 100 amino acids in length. In some embodiments, the linker can be about 100 - 150, 150 - 200, 200 - 250, 250 - 300, 300 - 350, 350 - 400, 400 - 450, or 450 - 500 amino acids in length. Longer or shorter linkers are also contemplated. In some embodiments, the linker connects the gRNA binding domain of an RNA programmable nuclease that includes a Cas9 nuclease domain to the catalytic domain of a nucleic acid editing protein (e.g., adenosine deaminase). In some embodiments, the linker connects dCas9 and a nucleic acid editing protein. For example, the linker is positioned between or adjacent to two groups, molecules, or other moieties and is covalently linked to each through a covalent bond, thus linking the two. In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or a protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5 - 200 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140, 150, 160, 175, 180, 190, or 200 amino acids in length. Longer or shorter linkers are also contemplated.

[0152] In some embodiments, the domains of the nucleobase editor are fused via a linker comprising the amino acid sequence SGGSSGSETPGTSESATPESSGGS, SGGSSGGSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTE PSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS. In some embodiments, the domains of the nucleobase editor are fused via a linker comprising the amino acid sequence SGSETPGTSESATPES, which may also be referred to as an XTEN linker. In some embodiments, the linker comprises the amino acid sequence SGGS. In some embodiments, the linker is (SGGS) n , (GGGS) n , (GGGGS) n , (G) n , (EAAAK) n , (GGS) n , SGSETPGTSESATPES, or (XP) n motif, or any combination thereof, where n is independently an integer from 1 to 30 and X is any amino acid. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15.

[0153] In some embodiments, the linker is 24 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In some embodiments, the linker is 40 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTESATPESSGGSSGGSSGSSGGS. In some embodiments, the linker is 64 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGSSGGS. In some embodiments, the linker is 92 amino acids in length. In some embodiments, the linker comprises the amino acid sequence PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS.

[0154] "Marker" means any protein or polynucleotide having a change in expression level or activity associated with a disease or disorder.

[0155] As used herein, the term "mutation" refers to the substitution of a residue within a sequence, e.g., within a nucleic acid or amino acid sequence, with another residue, or the deletion or insertion of one or more residues within the sequence. Mutations are typically described herein by identifying the original residue, then identifying the position of the residue within the sequence, and then identifying the newly substituted residue. Various methods for making amino acid substitutions (mutations) provided herein are well known in the art and are provided, for example, by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)). In some embodiments, the base editors of the present disclosure can efficiently generate "intended mutations" such as point mutations in nucleic acids (e.g., nucleic acids within a target genome) without generating a significant number of unintended mutations, e.g., unintended point mutations. In some embodiments, an intended mutation is a mutation caused by a specific base editor (e.g., an adenosine base editor) that is specifically designed to cause that intended mutation and that binds to a guide polynucleotide (e.g., a gRNA).

[0156] Generally, mutations made or identified in a sequence (e.g., an amino acid sequence described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain that mutation. One of ordinary skill in the art will readily understand methods for determining the position of mutations in amino acid and nucleic acid sequences relative to a reference sequence.

[0157] The term "non-conservative mutation" relates to amino acid substitutions between different groups, for example, from tryptophan to lysine, or from serine to phenylalanine, etc. In this case, it is preferred that the non-conservative amino acid substitution does not interfere with or inhibit the biological activity of the functional variant. The non-conservative amino acid substitution can enhance the biological activity of the functional variant such that the biological activity of the functional variant is increased compared to the wild-type protein.

[0158] The term "nuclear localization sequence", "nuclear localization signal" or "NLS" means an amino acid sequence that promotes the import of a protein into the cell nucleus. Nuclear localization sequences are known in the art and are described, for example, in Plank et al. of international PCT application PCT / EP 2000 / 011690, filed on November 23, 2000 and published as WO / 2001 / 038547 on May 31, 2001, the content of which is incorporated herein by reference for the disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, for example, as described by Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS is the amino acid sequence KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC comprises.

[0159] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to a compound containing a nucleobase and an acidic moiety, such as a nucleoside, nucleotide, or polymer of nucleotides. Typically, a polymeric nucleic acid, such as a nucleic acid molecule containing three or more nucleotides, is a linear molecule in which adjacent nucleotides are linked to each other via phosphodiester bonds. In some embodiments, "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides). In some embodiments, "nucleic acid" refers to an oligonucleotide chain containing three or more individual nucleotide residues. As used herein, the terms "oligonucleotide" and "polynucleotide" may be used interchangeably to refer to a polymer of nucleotides (e.g., a chain of at least three nucleotides). In some embodiments, "nucleic acid" encompasses RNA as well as single-stranded and / or double-stranded DNA. Nucleic acids can occur naturally, for example, in association with a genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, chromatid, or other naturally occurring nucleic acid molecule. On the other hand, a nucleic acid molecule can be a non-naturally occurring molecule, such as a non-naturally occurring molecule, recombinant DNA or RNA, artificial chromosome, engineered genome, or fragment thereof, or synthetic DNA, RNA, DNA / RNA hybrid, or a molecule containing non-naturally occurring nucleotides or nucleosides. Further, the terms "nucleic acid", "DNA", "RNA", and / or similar terms include nucleic acid analogs, such as analogs having other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems and optionally purified, chemically synthesized, etc. In the case of chemically synthesized molecules, nucleic acids can include, where appropriate, nucleoside analogs, such as chemically modified bases or sugars, and analogs having backbone modifications. Nucleic acid sequences are shown in the 5' to 3' direction unless otherwise indicated.In some embodiments, the nucleic acid is a natural nucleoside (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); a nucleoside analog (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanine, O(6)-methylguanine, and 2-thiocytidine); a chemically modified base; a biologically modified base (e.g., a methylated base); an inserted base; a modified sugar (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or a modified phosphate group (e.g., phosphorothioate and 5'-N-phosphoramidite linkages), or comprises them.

[0160] The term "programmable nucleic acid-binding protein" or "napDNAbp" can be used interchangeably with "polynucleotide-programmable nucleotide-binding domain" and refers to a protein that associates with a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid or guide polynucleotide (e.g., gRNA) that guides the napDNAbp to a specific nucleic acid sequence. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable DNA-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable RNA-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a Cas9 protein. The Cas9 protein can bind to a guide RNA that directs the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the napDNAbp is a Cas9 domain, such as nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Non-limiting examples of programmable nucleic acid-binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i.Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also called Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V Cas effector proteins, type VI Cas effector proteins, CARF, DinG, their homologs, or their modified or engineered versions. Other nucleic acid programmable DNA-binding proteins may not be specifically listed in this disclosure but are within the scope of this disclosure. See, for example, Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct;1:325-336. doi: 10.1089 / crispr.2018.0033; Yan et al., “Functionally diverse type V CRISPR-Cas systems” Science. 2019 Jan 4;363(6422):88-91. doi: 10.1126 / science.aav7271 (the entire contents of each are incorporated herein by reference).

[0161] The terms "nucleobase", "nitrogenous base", or "base" are used interchangeably herein and refer to nitrogen-containing biological compounds that form nucleosides, which are components of nucleotides. The ability of nucleobases to form base pairs and stack on top of each other directly results in long-chain helical structures such as ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). The five nucleobases, adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), are called primary bases or standard bases. Adenine and guanine are derived from purines, while cytosine, uracil, and thymine are derived from pyrimidines. DNA and RNA may also contain other (non-primary) modified bases. Non-limiting exemplary modified nucleobases include hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydroxymethylcytosine. Hypoxanthine and xanthine can be generated by the presence of mutagens and are both produced by deamination (substitution of an amine group with a carbonyl group). Hypoxanthine can be modified from adenine. Xanthine can be modified from guanine. Uracil can result from the deamination of cytosine. A "nucleoside" consists of a nucleobase and a pentose sugar (ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxyuridine, and deoxycytidine. Nucleosides with modified nucleobases include inosine (I), xanthosine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), pseudouridine (Ψ), etc. A nucleotide consists of a nucleobase, a pentose sugar (ribose or deoxyribose), and at least one phosphate group.

[0162] As used herein, the term "nucleobase editing domain" or "nucleobase editing protein" refers to a protein or enzyme that can catalyze nucleobase modifications in RNA or DNA, such as deamination of cytosine (or cytidine) to uracil (or uridine) or thymine (or thymidine), and of adenine (or adenosine) to hypoxanthine (or inosine), as well as non-template nucleotide addition and insertion. In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., an adenine deaminase or adenosine deaminase). In some embodiments, the nucleobase editing domain is a plurality of deaminase domains (e.g., an adenine deaminase or adenosine deaminase and a cytidine or cytosine deaminase). In some embodiments, the nucleobase editing domain can be a naturally occurring nucleobase editing domain. In some embodiments, the nucleobase editing domain can be a nucleobase editing domain engineered or evolved from a naturally occurring nucleobase editing domain. The nucleobase editing domain can be derived from any organism, such as bacteria, human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse.

[0163] As used herein, "obtaining" as in "obtaining an agent" includes synthesizing, purchasing, or otherwise acquiring the agent.

[0164] As used herein, "patient" or "subject" refers to a mammalian subject or individual diagnosed with, at risk of developing, or suspected of having or developing a disease or disorder. In some embodiments, the term "patient" refers to a mammalian subject at higher than average risk of developing a disease or disorder. Exemplary patients can be humans, non-human primates, cats, dogs, pigs, cows, cats, horses, camels, llamas, goats, sheep, rodents (e.g., mice, rabbits, rats, guinea pigs) and other mammals that can benefit from the treatments disclosed herein. Exemplary human patients can be male and / or female.

[0165] "A patient in need thereof" or "a subject in need thereof" is herein referred to as a patient diagnosed with, at risk of having, or determined or suspected of having a disease or disorder.

[0166] The terms "pathogenic variant", "pathogenic mutation", "disease-causing variant", "disease-causing mutation", "deleterious variant" or "predisposing mutation" refer to a genetic change or mutation that increases an individual's susceptibility or predisposition to a particular disease or disorder. In some embodiments, a pathogenic variant includes one in which at least one wild-type amino acid in a protein encoded by a gene is replaced by at least one pathogenic amino acid.

[0167] The terms "protein", "peptide", "polypeptide", and their grammatical equivalents are used interchangeably herein and refer to polymers of amino acid residues linked together by peptide (amide) bonds. This term refers to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide is at least 3 amino acids in length. A protein, peptide, or polypeptide can refer to an individual protein or a group of proteins. One or more amino acids in a protein, peptide, or polypeptide can be modified by the addition of chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups, linkers for attachment, functionalization, or other modifications. A protein, peptide, or polypeptide can also be a single molecule or a multimolecular complex. A protein, peptide, or polypeptide can be a mere fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide can be naturally occurring, recombinant, or synthetic, or any combination thereof. As used herein, the term "fusion protein" refers to a hybrid polypeptide that contains protein domains derived from at least two different proteins. One protein can be located at the amino-terminal (N-terminal) portion or the carboxy-terminal (C-terminal) portion of the fusion protein and thus forms an amino-terminal fusion protein or a carboxy-terminal fusion protein, respectively. A protein can contain different domains, such as a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9 that induces binding of the protein to a target site) and a nucleic acid cleavage domain, or a catalytic domain of a nucleic acid editing protein. In some embodiments, a protein contains a proteinaceous portion, such as an amino acid sequence that constitutes a nucleic acid binding domain, and an organic compound, such as a compound that can act as a nucleic acid cleavage agent. In some embodiments, a protein is complexed with or associated with a nucleic acid (e.g., RNA or DNA).Any protein provided herein can be produced by any method known in the art. For example, the proteins provided herein can be produced via recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for the expression and purification of recombinant proteins are well known and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), the entire content of which is incorporated herein by reference.

[0168] The polypeptides and proteins disclosed herein (including functional portions and functional variants thereof) can contain synthetic amino acids in place of one or more naturally occurring amino acids. Such synthetic amino acids are known in the art and include, for example, aminocyclohexanecarboxylic acid, norleucine, α-amino-n-decanoic acid, homoserine, S-acetylaminomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine, β-phenylserine β-hydroxyphenylalanine, phenylglycine, α-naphthylalanine, cyclohexylalanine, cyclohexylglycine, indoline-2-carboxylic acid, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, aminomalonic acid monoamide, N'-benzyl-N'-methyllysine, N',N'-dibenzyl-lysine, 6-hydroxylysine, ornithine, α-aminocyclopentanecarboxylic acid, aminocyclohexanecarboxylic acid, aminocyclohexanecarboxylic acid, α-aminocycloheptanecarboxylic acid, α-(2-amino-2-norbornane)-carboxylic acid, α,γ-diaminobutyric acid, α,β-diaminopropionic acid, homophenylalanine, and α-tert-butylglycine. Polypeptides and proteins can be conjugated to post-translational modifications of one or more amino acids of the polypeptide construct. Non-limiting examples of post-translational modifications include acylation including phosphorylation, acetylation and formylation, glycosylation (including N-link and O-link), amidation, hydroxylation, alkylation including methylation and ethylization, ubiquitination, addition of pyrrolidonecarboxylic acid, formation of disulfide bridges, sulfation, myristoylation, palmitoylation, isoprenylation, farnesylation, geranylation, glypiation, lipoylation, and iodination.

[0169] As used herein in connection with proteins or nucleic acids, the term "recombinant" refers to a protein or nucleic acid that does not exist in nature but is the product of human engineering. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that contains at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7 mutations as compared to any naturally occurring sequence.

[0170] "Decrease" means a negative change of at least 10%, 25%, 50%, 75%, or 100%.

[0171] "Reference" means a standard or control condition. In one embodiment, the reference is a wild-type or healthy cell. In other embodiments, without limitation, the reference is an untreated cell that has not been exposed to the test conditions or has been exposed to a placebo or normal saline, medium, buffer, and / or a control vector that does not carry the polynucleotide of interest.

[0172] "Reference sequence" is a defined sequence used as a basis for sequence comparison. The reference sequence can be a subset or the entirety of a particular sequence; for example, a segment of a full-length cDNA or gene sequence, or a full cDNA or gene sequence. For polypeptides, the length of the reference polypeptide sequence is generally at least about 16 amino acids, at least about 20 amino acids, at least about 25 amino acids, about 35 amino acids, about 50 amino acids, or about 100 amino acids. For nucleic acids, the length of the reference nucleic acid sequence is generally at least about 50 nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides or about 300 nucleotides or any integer around or between them. In some embodiments, the reference sequence is the wild-type sequence of the protein of interest. In other embodiments, the reference sequence is the polynucleotide sequence encoding the wild-type protein.

[0173] The terms "RNA-programmable nuclease" and "RNA-guided nuclease" are used with (e.g., bound to or associated with) one or more RNAs that are not the target of cleavage. In some embodiments, an RNA-programmable nuclease, when in complex with an RNA, may be referred to as a nuclease:RNA complex. Typically, the bound RNA is called a guide RNA (gRNA). The gRNA may exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule may sometimes be called a single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to guide RNAs that exist as a single molecule or as a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species comprises two domains: (1) a domain that shares homology with the target nucleic acid (e.g., that directs binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to tracrRNA as provided in Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those including domain 2) can be found in U.S. Provisional Patent Application U.S.S.N. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application U.S.S.N. 61 / 874,746, filed September 6, 2013, entitled "Delivery System For Functional Nucleases", the entire contents of each of which are incorporated herein by reference. In some embodiments, the gRNA comprises two or more of domains (1) and (2) and may be referred to as an "extended gRNA".As an example, an extended gRNA binds to, for example, two or more Cas9 proteins and binds to a target nucleic acid in two or more different regions, as described herein. The gRNA contains a nucleotide sequence complementary to the target site, which mediates the binding of the nuclease / RNA complex to the target site and provides sequence specificity for the nuclease:RNA complex.

[0174] In some embodiments, the RNA-programmable nuclease is a (CRISPR-associated system) Cas9 endonuclease, e.g., Cas9 (Csnl) from Streptococcus pyogenes (see, e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes." Ferretti J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C, Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663(2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chylinski K., Sharma CM., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Vogel J., Charpentier E., Nature 471:602-607(2011)).

[0175] RNA-programmable nucleases (e.g., Cas9) use RNA:DNA hybridization to target DNA cleavage sites, so in principle these proteins can target any sequence specified by the guide RNA. Methods of using RNA-programmable nucleases such as Cas9 for site-specific cleavage (e.g., for modifying the genome) are known in the art (see, e.g., Cong, L. et al., Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, W.Y. et al., Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); DiCarlo, J.E. et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic acids research (2013); Jiang, W. et al., RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature biotechnology 31, 233-239 (2013). The entire contents of each of these are incorporated herein by reference).

[0176] The term "single nucleotide polymorphism (SNP)" refers to a variation of a single nucleotide that occurs at a specific position in the genome, where each variation exists to a certain extent (e.g., >1%) that can be recognized within a population. For example, at a specific base position in the human genome, the C nucleotide may appear in most individuals, but in a small number of individuals, that position is occupied by A. This means that there is an SNP at this specific position, and the two nucleotide variations, C or A, are the alleles at this position. SNPs underlie differences in susceptibility to diseases. The severity of a disease and the body's response to treatment are also manifestations of genetic variations. SNPs can exist in the coding region of a gene, the non-coding region of a gene, or the intergenic region (the region between genes). In some embodiments, SNPs within the coding sequence do not necessarily change the amino acid sequence of the protein produced due to the degeneracy of the genetic code. There are two types of SNPs in the coding region: synonymous SNPs and non-synonymous SNPs. Synonymous SNPs do not affect the protein sequence, while non-synonymous SNPs change the amino acid sequence of the protein. There are two types of non-synonymous SNPs: missense and nonsense. SNPs that are not in the region encoding the protein may affect gene splicing, transcription factor binding, messenger RNA degradation, or the sequence of non-coding RNA. Gene expression affected by this type of SNP is called eSNP (expressed SNP) and can be upstream or downstream of the gene. A single nucleotide variant (SNV) is a single nucleotide variation without frequency limitation and can occur in somatic cells. Somatic single nucleotide variations can also be called single nucleotide modifications.

[0177] "Specifically binds" means a nucleic acid molecule, polypeptide, or complex thereof (e.g., a nucleic acid programmable DNA binding domain and a guide nucleic acid), compound, or molecule that recognizes and binds to the polypeptides and / or nucleic acid molecules of the present disclosure but does not substantially recognize and bind to other molecules in a sample (e.g., a biological sample).

[0178] Nucleic acid molecules useful in the methods of the present disclosure include any nucleic acid molecule encoding a polypeptide of the present disclosure or a fragment thereof. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but typically exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence can typically hybridize to at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present disclosure include any nucleic acid molecule encoding a polypeptide of the present disclosure or a fragment thereof. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but typically exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence can typically hybridize to at least one strand of a double-stranded nucleic acid molecule. "Hybridize" means to form pairs between complementary polynucleotide sequences (e.g., the genes described herein) or portions thereof to form double-stranded molecules under various stringency conditions. (See, e.g., Wahl, G. M. and S. L. Berger (1987) Methods Enzymol. 152:399; Kimmel, A. R. (1987) Methods Enzymol. 152:507).

[0179] For example, stringent salt concentrations are typically less than about 750 mM NaCl and 75 mM trisodium citrate, preferably less than about 500 mM NaCl and 50 mM trisodium citrate, more preferably less than about 250 mM NaCl and 25 mM trisodium citrate. Low stringency hybridization can be obtained in the absence of an organic solvent such as formamide, while high stringency hybridization can be obtained in the presence of at least about 35% formamide, more preferably at least about 50% formamide. Stringent temperature conditions will typically include a temperature of at least about 30°C, more preferably at least about 37°C, most preferably at least about 42°C. Various additional parameters such as hybridization time, concentration of surfactant (e.g., sodium dodecyl sulfate (SDS)), and inclusion or exclusion of carrier DNA are well known to those of skill in the art. By combining these various conditions as appropriate, various levels of stringency are achieved. In one embodiment, hybridization occurs at 30°C in 750 mM NaCl, 75 mM trisodium citrate and 1% SDS. In another embodiment, hybridization occurs at 37°C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide and 100 μg / ml denatured salmon sperm DNA (ssDNA). In another embodiment, hybridization occurs at 42°C in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide and 200 μg / ml ssDNA. Useful variations of these conditions will be readily apparent to those of skill in the art.

[0180] For most applications, the washing steps following hybridization also vary in stringency. The washing stringency conditions can be defined by salt concentration and temperature. As noted above, the washing stringency can be increased by decreasing the salt concentration or increasing the temperature. For example, a stringent salt concentration for a washing step is preferably less than about 30 mM NaCl and 3 mM trisodium citrate, and most preferably can be less than about 15 mM NaCl and 1.5 mM trisodium citrate. Stringent temperature conditions for a washing step typically include a temperature of at least about 25°C, more preferably at least about 42°C, even more preferably at least about 68°C. In certain embodiments, the washing step is carried out at 25°C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the washing step is carried out at 42°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the washing step is carried out at 68°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Further variations of these conditions will be readily apparent to those skilled in the art.Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.

[0181] "Split" means being split into two or more fragments.

[0182] The "split Cas9 protein" or "split Cas9" refers to a Cas9 protein provided as an N-terminal fragment and a C-terminal fragment encoded by two separate nucleotide sequences. The polypeptides corresponding to the N-terminal portion and the C-terminal portion of the Cas9 protein can be spliced to form a "reconstituted" Cas9 protein. In certain embodiments, the Cas9 protein is split into two fragments within the disordered region of the protein, as described, for example, in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014, or as described in Jiang et al. (2016) Science 351: 867-871. PDB file: 5F9R (each incorporated herein by reference). In some embodiments, the protein is split into two fragments at any C, T, A, or S within the region between approximately amino acids A292-G364, F445-K483, or E565-T637 of SpCas9, or at the corresponding position in any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or other napDNAbp. In some embodiments, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of splitting the protein into two fragments is referred to as "splitting" the protein.

[0183] In other embodiments, the N-terminal portion of the Cas9 protein comprises amino acids 1-573 or 1-637 of S. pyogenes Cas9 wild type (SpCas9) (NCBI reference sequence: NC_002737.2, Uniprot reference sequence: Q99ZW2), and the C-terminal portion of the Cas9 protein comprises the portion of amino acids 574-1368 or 638-1368 of SpCas9 wild type, or the corresponding portion thereof.

[0184] The C-terminal portion of split Cas9 can be linked to the N-terminal portion of split Cas9 to form a complete Cas9 protein. In some embodiments, the C-terminal portion of the Cas9 protein begins where the N-terminal portion of the Cas9 protein ends. Thus, in some embodiments, the C-terminal portion of split Cas9 comprises the portion of spCas9 amino acids (551-651)-1368. By "(551-651)-1368" it is meant starting from the amino acids between amino acids 551 to 651 (both ends included) and ending at amino acid 1368.For example, the C-terminal portion of split Cas9 may include any one of the portions of spCas9 amino acids 551-1368, 552-1368, 553-1368, 554-1368, 555-1368, 556-1368, 557-1368, 558-1368, 559-1368, 560-1368, 561-1368, 562-1368, 563-1368, 564-1368, 565-1368, 566-1368, 567-1368, 568-1368, 569-1368, 570-1368, 571-1368, 572-1368, 573-1368, 574-1368, 575-1368, 576-1368, 577-1368, 578-1368, 579-1368, 580-1368, 581-1368, 582-1368, 583-1368, 584-1368, 585-1368, 586-1368, 587-1368, 588-1368, 589-1368, 590-1368, 591-1368, 592-1368, 593-1368, 594-1368, 595-1368, 596-1368, 597-1368, 598-1368, 599-1368, 600-1368, 601-1368, 602-1368, 603-1368, 604-1368, 605-1368, 606-1368, 607-1368, 608-1368, 609-1368, 610-1368, 611-1368, 612-1368, 613-1368, 614-1368, 615-1368, 616-1368, 617-1368, 618-1368, 619-1368, 620-1368, 621-1368, 622-1368, 623-1368, 624-1368, 625-1368, 626-1368, 627-1368, 628-1368, 629-1368, 630-1368, 631-1368, 632-1368, 633-1368, 634-1368, 635-1368, 636-1368, 637-1368, 638-1368, 639-1368, 640-1368, 641-1368, 642-1368, 643-1368, 644-1368, 645-1368, 646-1368, 647-1368, 648-1368, 649-1368, 650-1368, or 651-1368.In some embodiments, the C-terminal portion of the split Cas9 protein comprises the portion of SpCas9 amino acids 574-1368 or 638-1368.

[0185] "Subject" means a mammal, including but not limited to a human or non-human mammal such as a cow, horse, dog, sheep or cat. The subject includes livestock and breeding animals raised to produce goods such as labor and food, including but not limited to cows, goats, chickens, horses, pigs, rabbits, and sheep.

[0186] "Substantially identical" means a polypeptide or nucleic acid molecule that exhibits at least 50% identity to a reference amino acid sequence (e.g., any one of the amino acid sequences described herein) or nucleic acid sequence (e.g., any one of the nucleic acid sequences described herein). In one embodiment, such a sequence has identity with the sequence used for comparison at the amino acid level or nucleic acid of at least 60%, 80% or 85%, 90%, 95% or 99%.

[0187] Sequence identity is typically measured using sequence analysis software (e.g., the Sequence Analysis Software Package from Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, GAP, or the PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning a degree of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; phenylalanine, tyrosine. In an exemplary approach for determining the degree of identity, the BLAST program can be used, and the probability score between e -3 and e -100 indicates sequences that are closely related. COBALT is used, for example, with the following parameters: a) Alignment parameters: Gap penalties -11, -1 and End-Gap penalties -5, -1, b) CDD parameters: Use RPS BLAST on; Blast E-value 0.003; Find Conserved columns and Recompute on c) Query clustering parameters: Use query clusters on; Word Size 4; Max cluster distance 0.8; Alphabet Regular. EMBOSS Needle is used, for example, with the following parameters. a) Matrix: BLOSUM62; b) GAP OPEN: 10; c) GAP EXTEND: 0.5; d) OUTPUT FORMAT: pair; e) END GAP PENALTY: false; f) END GAP OPEN: 10; and g) END GAP EXTEND: 0.5.

[0188] As used herein, the term "target site" refers to a sequence within a nucleic acid molecule that is modified by a nucleic acid base editor. In one embodiment, the target site is deaminated by a fusion protein comprising a deaminase or a deaminase (e.g., adenine deaminase).

[0189] As used herein, the terms "treat", "treating", "treatment", etc. refer to reducing or ameliorating a disorder and / or its associated symptoms, or obtaining a desired pharmacological and / or physiological effect. It will be understood that treating a disorder or condition does not necessarily require complete elimination of the associated disorder, condition or symptoms (although complete elimination is not excluded). In some embodiments, the effect is therapeutic, i.e., without limitation, the effect is to partially or completely reduce, decrease, eliminate, alleviate, mitigate, attenuate, or cure a disease and / or its deleterious symptoms caused thereby. In some embodiments, the effect is prophylactic, i.e., the effect protects or prevents the occurrence or recurrence of a disease or condition. For this purpose, the methods of the present disclosure include administering a therapeutically effective amount of a composition as described herein.

[0190] "Uracil glycosylase inhibitor", or "UGI" for short, refers to a factor that inhibits the uracil excision repair system. In one embodiment, the factor is a protein or a fragment thereof that binds to the host uracil-DNA glycosylase and prevents the removal of uracil residues from DNA. In certain embodiments, UGI is a protein, a fragment thereof, or a domain that can inhibit the uracil-DNA glycosylase base excision repair enzyme. In some embodiments, the UGI domain includes the wild-type UGI or a modified version thereof. In some embodiments, the UGI domain includes a fragment of the exemplary amino acid sequence presented below. In some embodiments, the UGI fragment includes an amino acid sequence that comprises at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the exemplary UGI sequence provided below. In some embodiments, UGI includes an amino acid sequence that is homologous to the exemplary UGI amino acid sequence or a fragment thereof as described below. In some embodiments, UGI or a portion thereof has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, or 100% identity to the wild-type UGI or the UGI sequence or a portion thereof as described below. Exemplary UGI includes the following amino acid sequence: >splP14739IUNGI_BPPB2 Uracil-DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSD APEYKPWALVIQDSNGENKIKML.

[0191] The term "vector" refers to a means of introducing a nucleic acid sequence into a cell, resulting in a transformed cell. Examples of vectors include plasmids, transposons, phages, viruses, liposomes, and episomes. An "expression vector" is a nucleic acid sequence that contains a nucleotide sequence to be expressed in a recipient cell. An expression vector can contain additional nucleic acid sequences that promote and / or facilitate the expression of the introduced sequence, such as initiation, termination, enhancer, promoter, and secretion sequences.

[0192] The compositions or methods provided herein can be combined with one or more of the other compositions and methods provided herein.

[0193] DNA editing has emerged as a viable means to alter disease states by correcting pathogenic mutations at the gene level. Until recently, all DNA editing platforms functioned by inducing DNA double-strand breaks (DSBs) at specific genomic sites and relying on endogenous DNA repair pathways to determine product outcomes in a semi-random fashion, resulting in complex gene product populations. Precise, user-defined repair outcomes can be achieved via the homology-directed repair (HDR) pathway, but numerous challenges have hampered efficient repair using HDR in therapeutically relevant cell types. In fact, this pathway is insufficient compared to the competing, error-prone non-homologous end joining pathway. Additionally, HDR is severely limited to the G1 and S phases of the cell cycle, which precludes accurate repair of DSBs in post-mitotic cells. As a result, it has proven difficult or impossible to modify genomic sequences in these populations in a highly efficient, user-defined, programmable manner.

Brief Description of the Drawings

[0194]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Figure 44

Figure 45

Figure 46

Figure 47

Figure 48

Figure 49

Figure 50

Figure 51

Figure 52

Figure 53

Figure 54

Figure 55

Figure 56

Figure 57

Figure 58

Figure 59

Figure 60

Figure 61

Figure 62

[0195] The present disclosure provides compositions comprising novel adenine base editors (e.g., ABE8) having increased efficiency, and methods of using these to effect modifications in a target nucleic acid base sequence.

[0196] [Nucleic acid base editor] Disclosed herein are base editors or nucleic acid base editors for editing, modifying or altering a target nucleotide sequence of a polynucleotide. Described herein are nucleic acid base editors or base editors comprising a polynucleotide programmable nucleotide binding domain and a nucleic acid base editing domain. The polynucleotide programmable nucleotide binding domain (e.g., adenosine deaminase) can specifically bind to a target polynucleotide sequence (through complementary base pairing between the bases of the bound guide polynucleotide (e.g., gRNA) and the bases of the target polynucleotide sequence), thereby localizing the base editor to the target nucleic acid sequence that is desired to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid.

[0197] [Polynucleotide programmable nucleotide binding domain] It should be understood that the polynucleotide-programmable nucleotide binding domain can also include nucleic acid-programmable proteins that bind to RNA. For example, a polynucleotide-programmable nucleotide binding domain can be bound to a nucleic acid that guides the polynucleotide-programmable nucleotide binding domain to RNA. Other nucleic acid-programmable DNA binding proteins are also within the scope of the present disclosure, but they are not specifically listed in the present disclosure.

[0198] The polynucleotide-programmable nucleotide binding domain of a base editor can itself include one or more domains. For example, a polynucleotide-programmable nucleotide binding domain can include one or more nuclease domains. In some embodiments, the nuclease domain of the polynucleotide-programmable nucleotide binding domain can include an endonuclease or an exonuclease. As used herein, the term "exonuclease" refers to a protein or polypeptide that can digest a nucleic acid (e.g., RNA or DNA) from a free end, and the term "endonuclease" refers to a protein or polypeptide that can catalyze (e.g., cleave) an internal region of a nucleic acid (e.g., DNA or RNA). In some embodiments, the endonuclease can cleave one strand of a double-stranded nucleic acid. In some embodiments, the endonuclease can cleave both strands of a double-stranded nucleic acid molecule. In some embodiments, the polynucleotide-programmable nucleotide binding domain can be a deoxyribonuclease. In some embodiments, the polynucleotide-programmable nucleotide binding domain can be a ribonuclease.

[0199] In some embodiments, the nuclease domain of the polynucleotide-programmable nucleotide-binding domain can cleave zero, one, or two strands of the target polynucleotide. In some embodiments, the polynucleotide-programmable nucleotide-binding domain can include a nickase domain. As used herein, the term "nickase" refers to a polynucleotide-programmable nucleotide-binding domain that includes a nuclease domain capable of cleaving only one strand of a double-strand in a double-stranded nucleic acid molecule (e.g., DNA). In some embodiments, the nickase can be derived from a fully catalytically active (e.g., native) form of the polynucleotide-programmable nucleotide-binding domain by introducing one or more mutations into the active polynucleotide-programmable nucleotide-binding domain. For example, when the polynucleotide-programmable nucleotide-binding domain includes a nickase domain derived from Cas9, the nickase domain derived from Cas9 can include a D10A mutation and a histidine at position 840. In such embodiments, residue H840 retains catalytic activity and can thereby cleave one strand of the nucleic acid duplex. In another example, the Cas9-derived nickase domain can include an H840A mutation, while the amino acid residue at position 10 remains as D. In some embodiments, the nickase can be derived from a fully catalytically active (e.g., native) form of the polynucleotide-programmable nucleotide-binding domain by removing all or a portion of the nuclease domain that is not required for nickase activity. For example, when the polynucleotide-programmable nucleotide-binding domain includes a nickase domain derived from Cas9, the nickase domain derived from Cas9 can include a deletion of all or a portion of the RuvC domain or the HNH domain.

[0200] The amino acid sequence of exemplary catalytically active Cas9 is as follows:

[0201] Thus, a base editor comprising a polynucleotide-programmable nucleotide binding domain that includes a nickase domain can generate a single-stranded DNA break (nick) at a specific polynucleotide target sequence (e.g., determined by the complementary sequence of a bound guide nucleic acid). In some embodiments, the strand of the nucleic acid duplex target polynucleotide sequence that is cleaved by a base editor comprising a nickase domain (e.g., a nickase domain derived from Cas9) is the strand that is not edited by the base editor (i.e., the strand cleaved by the base editor is the strand opposite the strand that contains the base to be edited). In other embodiments, a base editor comprising a nickase domain (e.g., a nickase domain derived from Cas9) can cleave the strand of the DNA molecule targeted for editing. In such embodiments, the non-target strand is not cleaved.

[0202] Base editors that include a catalytically dead (i.e., unable to cleave a target polynucleotide sequence) polynucleotide-programmable nucleotide binding domain are also provided herein. As used herein, the terms "catalytically dead" and "nuclease-inactive" are used interchangeably to refer to a polynucleotide-programmable nucleotide binding domain having one or more mutations and / or deletions that result in the inability to cleave a nucleic acid strand. In some embodiments, a catalytically dead polynucleotide-programmable nucleotide binding domain base editor can lack nuclease activity as a result of specific point mutations in one or more nuclease domains. For example, in the case of a base editor that includes a Cas9 domain, Cas9 can include both the D10A mutation and the H840A mutation. Such mutations inactivate both nuclease domains, resulting in the loss of nuclease activity. In other embodiments, a catalytically dead polynucleotide-programmable nucleotide binding domain can include one or more deletions of all or part of the catalytic domain (e.g., the RuvC1 and / or HNH domains). In further embodiments, a catalytically dead polynucleotide-programmable nucleotide binding domain includes point mutations (e.g., D10A or H840A) as well as deletions of all or part of the nuclease domain.

[0203] Also contemplated herein are mutations that can generate catalytically dead polynucleotide programmable nucleotide binding domains from a previously functional version of the polynucleotide programmable nucleotide binding domain. For example, in the case of catalytically dead Cas9 (“dCas9”), variants are provided that have mutations other than D10A and H840A that result in nuclease-inactivated Cas9. Such mutations include, for example, other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). Further suitable nuclease-inactive dCas9 domains may be apparent to those skilled in the art based on the present disclosure and knowledge in the art and are within the scope of the present disclosure. Such further exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013; 31(9): 833-838, the entire contents of which are incorporated herein by reference).

[0204] Non-limiting examples of polynucleotide-programmable nucleotide binding domains that can be incorporated into base editors include CRISPR protein-derived domains, restriction nucleases, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases (ZFNs). In some embodiments, the base editor comprises a polynucleotide-programmable nucleotide binding domain comprising a native or modified protein or a portion thereof that can bind to a nucleic acid sequence via a binding guide nucleic acid during CRISPR (i.e., Clustered Regularly Interspaced Short Palindromic Repeats)-mediated modification of the nucleic acid. Such proteins are referred to herein as "CRISPR proteins." Accordingly, disclosed herein is a base editor comprising a polynucleotide-programmable nucleotide binding domain comprising all or a portion of a CRISPR protein (i.e., a base editor comprising all or a portion of a CRISPR protein as a domain, which is also referred to as the "CRISPR protein-derived domain" of the base editor). The CRISPR protein-derived domain incorporated into the base editor can be modified compared to the wild-type or native CRISPR protein. For example, as described below, the CRISPR protein-derived domain can include one or more mutations, insertions, deletions, rearrangements, and / or recombinations compared to the wild-type or native CRISPR protein.

[0205] CRISPR is an adaptive immune system that provides defense against mobile genetic elements (viruses, transposable elements, conjugative plasmids). CRISPR clusters contain spacers, sequences complementary to preceding mobile elements, and target invading nucleic acids. The CRISPR clusters are transcribed and processed into CRISPR RNAs (crRNAs). In type II CRISPR systems, correct processing of pre-crRNA requires trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. TracrRNA serves as a guide for the processing of pre-crRNA by ribonuclease 3. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to the crRNA is first endonucleolytically cleaved and then exonucleolytically trimmed 3'-5'. In nature, both a protein and both RNAs are required for DNA binding and cleavage. However, a single guide RNA (the "sgRNA", or simply "gRNA") can be engineered to incorporate both sides of crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J. A., Charpentier E. Science 337:816-821 (2012) (the entire content of which is incorporated herein by reference). Cas9 recognizes a short motif (PAM or protospacer adjacent motif) in the CRISPR repeat sequences to help distinguish "self" from "non-self".

[0206] In some embodiments, the methods described herein can utilize engineered Cas proteins. Guide RNAs (gRNAs) are short synthetic RNAs consisting of a scaffold sequence necessary for Cas binding and a user-defined approximately 20-base spacer that defines the genomic target to be modified. Thus, one of ordinary skill in the art can change the genomic target of Cas protein specificity, which is partially determined by how specific the gRNA targeting sequence is for the genomic target as compared to other parts of the genome.

[0207] In some embodiments, the gRNA scaffold sequence is as follows: GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU

[0208] In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is an endonuclease (e.g., deoxyribonuclease or ribonuclease) that can bind to the target polynucleotide when combined with the binding guide nucleic acid. In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is a nickase that can bind to the target polynucleotide when combined with the binding guide nucleic acid. In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is a catalytically dead domain that can bind to the target polynucleotide when combined with the binding guide nucleic acid. In some embodiments, the target polynucleotide that binds to the CRISPR protein-derived domain of the base editor is DNA. In some embodiments, the target polynucleotide that binds to the CRISPR protein-derived domain of the base editor is RNA.

[0209] The CAs proteins that can be used in this specification include Class 1 and Class 2. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also called Csn1 or Csx12), Cas10, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5, Csn1, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Cssx16, Cx16, Cx, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i, CARF, DinG, their homologs, or their variants. Unmodified CRISPR enzymes can have DNA cleavage activity having two functional endonuclease regions, such as RuvC and HNH, like Cas9. The CRISPR enzyme can induce cleavage of one or both strands at a target sequence, such as within the target sequence and / or within the complementary strand of the target sequence. For example, the CRISPR enzyme can induce cleavage of one or both strands that are about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500 base pairs or more from the first or last nucleotide of the target sequence.

[0210] A vector encoding a CRISPR enzyme mutated with respect to a corresponding wild-type enzyme can be used so as to lack the ability to cleave one or both strands of a target polynucleotide containing a target array. Cas9 can refer to a polypeptide having at least, or at least approximately, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology with an exemplary wild-type Cas9 polypeptide (e.g., Cas9 from S. pyogenes). Cas9 can refer to a polypeptide having at most, or at most approximately, about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology with an exemplary wild-type Cas9 polypeptide (e.g., from S. pyogenes). Cas9 can refer to a modified form of a Cas9 protein that can include amino acid changes such as wild-type, or deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof.

[0211] In some embodiments, the CRISPR protein-derived domain of the base editor can comprise all or a portion of Cas9 from Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquis (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref: YP_820832.1); Listeria innocua (NCBI Ref: NP_472073.1); Campylobacter jejuni (NCBI Ref: YP_002344900.1); Neisseria meningitidis (NCBI Ref: YP_002342100.1), Streptococcus pyogenes, or Staphylococcus aureus.

[0212] [Cas9 Domain of Nucleic Acid Base Editor] The sequences and structures of Cas9 nucleases are well known to those of ordinary skill in the art (see, e.g., “Complete genome sequence of an Ml strain of Streptococcus pyogenes.” Ferretti et al., J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C, Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663(2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma CM., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Vogel J., Charpentier E., Nature 471:602-607(2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821(2012). The entire contents of which are hereby incorporated by reference in their entirety.). Cas9 orthologs have been described in various species including, but not limited to, S. pyogenes and S. thermophilus.Additional suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on the present disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737. The entire content is hereby incorporated by reference into this specification.

[0213] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is a Cas9 domain. Non-limiting exemplary Cas9 domains are provided herein. The Cas9 domain can be a nuclease-active Cas9 domain, a nuclease-inactive Cas9 domain (dCas9), or a Cas9 nickase (nCas9). In some embodiments, the Cas9 domain is a nuclease-active domain. For example, the Cas9 domain can be a Cas9 domain that cleaves both strands of a double-stranded nucleic acid (e.g., both strands of a double-stranded DNA molecule). In some embodiments, the Cas9 domain comprises any one of the amino acid sequences described herein. In some embodiments, the Cas9 domain comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the amino acid sequences described herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any one of the amino acid sequences described herein.In some embodiments, the Cas9 domain comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical consecutive amino acid residues as compared to any one of the amino acid sequences described herein.

[0214] In some embodiments, proteins comprising fragments of Cas9 are provided. For example, in some embodiments, the protein comprises one of the following two Cas9 domains: (1) the gRNA binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, a protein comprising Cas9 or a fragment thereof is referred to as a "Cas9 variant". A Cas9 variant shares homology with Cas9 or a fragment thereof. For example, a Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to wild-type Cas9. In some embodiments, a Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type Cas9. In some embodiments, a Cas9 variant comprises a fragment of Cas9 (e.g., the gRNA binding domain or the DNA cleavage domain), and the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9.In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid length of the corresponding wild-type Cas9. In some embodiments, the fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids in length.

[0215] In some embodiments, the Cas9 fusion proteins provided herein include the full-length amino acid sequence of the Cas9 protein, e.g., one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein do not include the full-length Cas9 sequence and include only one or more fragments thereof. Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and additional suitable sequences of Cas9 domains and fragments will be apparent to those skilled in the art.

[0216] The Cas9 protein can bind to a guide RNA that guides the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the polynucleotide programmable nucleotide binding domain is a Cas9 domain, such as nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Examples of nucleic acid programmable DNA binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpfl, Cas12b / C2Cl, and Cas12c / C2C3.

[0217] In some embodiments, the wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_017053.1, the following nucleotide and amino acid sequences). TIFF0007693552000011.tif169163(Underlined once: HNH domain; Underlined twice: RuvC domain)

[0218] In some embodiments, wild-type Cas9 corresponds to, or comprises, the following nucleotide and / or amino acid sequences: TIFF0007693552000012.tif168162(Underlined once: HNH domain; Underlined twice: RuvC domain).

[0219] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_002737.2 (the following nucleotide sequence); and Uniprot reference sequence: Q99ZW2 (the following amino acid sequence): TIFF0007693552000013.tif168161(Underlined once: HNH domain; Underlined twice: RuvC domain)

[0220] In some embodiments, Cas9 refers to Cas9 from Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquisI (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref: YP_820832.1), Listeria innocua (NCBI Ref: NP_472073.1), Campylobacter jejuni (NCBI Ref: YP_002344900.1) or Neisseria meningitidis (NCBI Ref: YP_002342100.1), or Cas9 from any other organism.

[0221] Additional Cas9 proteins (e.g., nuclease-dead (dead) Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9), including their variants and homologs, are to be understood as being within the scope of the present disclosure. Exemplary Cas9 proteins include, but are not limited to, those provided below. In some embodiments, the Cas9 protein is nuclease-dead Cas9 (dCas9). In some embodiments, the Cas9 protein is Cas9 nickase (nCas9). In some embodiments, the Cas9 protein is nuclease-active Cas9.

[0222] In some embodiments, the Cas9 domain is a nuclease-inactive Cas9 domain (dCas9). For example, the dCas9 domain can bind to a double-stranded nucleic acid molecule (e.g., via a gRNA molecule) without cleaving either strand of the double-stranded nucleic acid molecule. In some embodiments, the nuclease-inactive dCas9 domain comprises the D10X and H840X mutations of the amino acid sequences described herein, or corresponding mutations in any of the amino acid sequences provided herein, where X is any amino acid change. In some embodiments, the nuclease-inactive dCas9 domain comprises the D10A and H840A mutations of the amino acid sequences described herein, or corresponding mutations in any of the amino acid sequences described herein. As an example, the nuclease-inactive Cas9 domain comprises the following amino acid sequence provided in the cloning vector pPlatTET-gRNA 2 (accession number BAV54124):

[0223] The amino acid sequence of an exemplary catalytically inactive Cas9 (dCas9) is as follows: (See, e.g., Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression.” Cell. 2013; 152(5):1173-83 (the entire content of which is incorporated herein by reference).)

[0224] Additional suitable nuclease-inactive dCas9 domains will be apparent to those skilled in the art based on the present disclosure and knowledge in the art and are within the scope of the present disclosure. Such additional exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013; 31(9):833-838 (the entire content of which is incorporated herein by reference)).

[0225] In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is a nickase referred to as the “nCas9” protein (short for “nickase” Cas9). The nuclease-inactivated Cas9 protein may also be referred to interchangeably as the “dCas9” protein (short for nuclease-“dead” Cas9) or catalytically inactive Cas9. Methods for generating a Cas9 protein (or a fragment thereof) having an inactive DNA cleavage domain are known (see, e.g., Jinek et al, Science. 337:816-821 (2012); Qi et al, “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression” (2013) Cell. 28; 152(5):1173-83 (each of which is incorporated herein by reference)). For example, the DNA cleavage domain of Cas9 is known to include two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, and the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can suppress the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al, Science. 337:816-821 (2012); Qi et al, Cell. 28;152(5):1173-83 (2013)).

[0226] In some embodiments, the dCas9 domain comprises an amino acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to any of the dCas9 domains provided herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the amino acid sequences described herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical consecutive amino acid residues compared to any of the amino acid sequences described herein.

[0227] In some embodiments, dCas9 corresponds to, or comprises a part or all of, a Cas9 amino acid sequence having one or more mutations that inactivate Cas9 nuclease activity. For example, in some embodiments, the dCas9 domain comprises the D10A and H840A mutations or corresponding mutations in another Cas9.

[0228] In some embodiments, dCas9 comprises the amino acid sequence of dCas9 (D10A and H840A): TIFF0007693552000014.tif168162(Underline once: HNH domain; Underline twice: RuvC domain)

[0229] In some embodiments, the Cas9 domain contains a D10A mutation, while the residue at position 840 in the amino acid sequence provided above, or the residue at the corresponding position in any of the amino acid sequences provided herein, remains histidine.

[0230] In other embodiments, dCas9 variants are provided that have mutations other than D10A and H840A, which result in, for example, nuclease-inactivated Cas9 (dCas9). Such mutations include, for example, other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). In some embodiments, variants or homologs of dCas9 are provided that have at least about 70% identity, at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity. In some embodiments, variants of dCas9 are provided that have an amino acid sequence that is about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids or more shorter or longer.

[0231] In some embodiments, the Cas9 domain is a Cas9 nickase. A Cas9 nickase can be a Cas9 protein that can cleave only one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 nickase cleaves the target strand of the double-stranded nucleic acid molecule, which means that the Cas9 nickase cleaves the strand that is base-pairing (complementary) with the gRNA (e.g., sgRNA) bound to the Cas9. In some embodiments, the Cas9 nickase contains a D10A mutation and has a histidine at position 840. In some embodiments, the Cas9 nickase cleaves the non-target, non-base-editing strand of the double-stranded nucleic acid molecule, which means that the Cas9 nickase cleaves the strand that is not base-pairing with the gRNA (e.g., sgRNA) bound to the Cas9. In some embodiments, the Cas9 nickase contains an H840A mutation and has an aspartic acid residue at position 10, or a corresponding mutation. In some embodiments, the Cas9 nickase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the Cas9 nickases provided herein. Further suitable Cas9 nickases will be apparent to those skilled in the art based on the present disclosure and knowledge in the art and are within the scope of the present disclosure.

[0232] The amino acid sequence of an exemplary catalytic Cas9 nickase (nCas9) is as follows:

[0233] In some embodiments, Cas9 refers to Cas9 from archaea (e.g., Nanoarchaea) that constitute the domain and kingdom of single-celled prokaryotic microorganisms. In some embodiments, the programmable nucleotide-binding protein can be, for example, the CasX or CasY protein described in Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res. 2017 Feb 21. doi: 10.1038 / cr.2017.21, the entire content of which is incorporated herein by reference. Using genomic resolution metagenomics, many CRISPR-Cas systems have been identified, including Cas9 first reported in the archaeal domain of life. This divergent Cas9 protein was discovered as part of an active CRISPR-Cas system in the little-studied Nanoarchaea. In bacteria, two previously unknown systems, CRISPR-CasX and CRISPR-CasY, were discovered, and they fall into the most compact systems discovered so far. In some embodiments, in the base editor systems described herein, Cas9 is replaced by CasX or a variant of CasX. In some embodiments, in the base editor systems described herein, Cas9 is replaced by CasY or a variant of CasY. Other RNA-guided DNA-binding proteins can also be used as nucleic acid programmable DNA-binding proteins (napDNAbp), and it should be understood that they are within the scope of the present disclosure.

[0234] In some embodiments, any nucleic acid programmable DNA binding protein (napDNAbp) of the fusion proteins provided herein can be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the napDNAbp is a CasY protein. In some embodiments, the napDNAbp comprises an amino acid sequence having at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to a naturally occurring CasX or CasY protein. In some embodiments, the programmable nucleotide binding protein is a naturally occurring CasX or CasY protein. In some embodiments, the programmable nucleotide binding protein comprises an amino acid sequence having at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to any CasX or CasY protein described herein. It should be understood that CasX and CasY from other bacterial species can also be used in accordance with the present disclosure.

[0235] The amino acid sequence of exemplary CasX ((uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) tr|F0NN87|F0NN87_SULIHCRISPR-associated Casx protein OS = Sulfolobus islandicus (strain HVE10 / 4) GN = SiH_0402 PE=4 SV=1) is as follows: MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAVGQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG

[0236] The amino acid sequence of exemplary CasX (>tr|F0NH53|F0NH53_SULIR CRISPR associated protein, Casx OS = Sulfolobus islandicus (strain REY15A) GN=SiRe_0771 PE=4 SV=1) is as follows: MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAVGQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG

[0237] Deltaproteobacteria CasX MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMTDDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAILQVYWQEFKDDHVGLMCKFAQPASKKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFVYKLEQVSEKGKAYTNYFGRCNVAEHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESTHPVKPLAQIAGNRYASGPVGKALSDACMGTIASFLSKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDfAYNEVIARVRMWVNLNLWQKLKLSRDDAKPLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENPKKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEMDEKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFTDGTDIKKSGKWQGLLYGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKIGRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQAAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAKLAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKELSAELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSYKSGKQPFVGAWQAFYKRRLKEVWKPNA

[0238] The amino acid sequence of exemplary CasY ((ncbi.nlm.nih.gov / protein / APG80656.1) >APG80656.1 CRISPR-associated protein CasY [uncultured Parcubacteria group bacterium]) is as follows:

[0239] Cas9 nuclease has two functional endonuclease domains, RuvC and HNH. When Cas9 binds to the target DNA, it undergoes a conformational change that positions the nuclease domain and cleaves the strand opposite the target DNA. The final result of DNA cleavage via Cas9 is a double-strand break (DSB) within the target DNA (about 3-4 nucleotides upstream of the PAM sequence). The resulting DSB is repaired by one of two general repair pathways: (1) the error-prone but efficient non-homologous end joining (NHEJ) pathway; or (2) the high-fidelity but less efficient homology-directed repair (HDR) pathway.

[0240] The "efficiency" of non-homologous end joining (NHEJ) and / or homology-directed repair (HDR) can be calculated by any convenient method. For example, in some embodiments, efficiency can be expressed as the percentage of successful HDR. For example, cleavage products can be generated using an assay nuclease assay, and the percentage can be calculated using the ratio of product to substrate. For example, a survey nuclease enzyme that directly cleaves DNA containing a newly incorporated restriction sequence as a result of successful HDR can be used. The higher the number of substrates cleaved, the higher the percentage of HDR (the higher the efficiency of HDR). As an illustrative example, the percentage of HDR can be calculated using the following formula: [(cleavage product) / (substrate + cleavage product)] (e.g., (b + c) / (a + b + c) where "a" is the band intensity of the DNA substrate and "b" and "c" are the cleavage products).

[0241] In some embodiments, efficiency can be represented by the success rate of NHEJ. For example, the T7 endonuclease I assay can be used to generate cleavage products, and the percentage of NHEJ can be calculated using the ratio of the product to the substrate. T7 endonuclease I cleaves mismatched heteroduplex DNA resulting from the hybridization of wild-type and mutant DNA strands (NHEJ results in small random insertions or deletions (indels) at the initial cleavage site). More cleavage indicates a higher proportion of NHEJ (higher efficiency of NHEJ). As an illustrative example, the proportion (percentage) of NHEJ can be calculated using the formula (1-(1-(b + c) / (a + b + c)) 1 / 2 ) × 100, where "a" is the band intensity of the DNA substrate, and "b" and "c" are the cleavage products (Ran et. al., Cell. 2013 Sep. 12; 154(6):1380 - 9; and Ran et al., Nat Protoc. 2013 Nov.; 8(11):2281 - 2308).

[0242] The NHEJ repair pathway is the most active repair mechanism and frequently causes small nucleotide insertions or deletions (indels) at the DSB site. The randomness of NHEJ-mediated DSB repair has important practical implications because a cell population expressing Cas9 and gRNA or guide polynucleotide can result in a diverse array of mutations. In most embodiments, NHEJ causes small indels in the target DNA, resulting in amino acid deletions, insertions, or frameshift mutations that introduce premature stop codons within the open reading frame (ORF) of the target gene. The ideal end result is a loss-of-function mutation within the target gene.

[0243] NHEJ-mediated DSB repair often disrupts the open reading frame of genes, while homology-directed repair (HDR) can be used to generate specific nucleotide changes ranging from single nucleotide changes to large insertions such as the addition of fluorophores or tags. To utilize HDR for gene editing, a DNA repair template containing the desired sequence can be delivered to the target cell type along with gRNA and Cas9 or Cas9 nickase. The repair template can contain the desired edit as well as additional homologous sequences immediately upstream and downstream of its target (referred to as left and right homology arms). The length of each homology arm can depend on the size of the change to be introduced, with larger insertions requiring longer homology arms. The repair template can be a single-stranded oligonucleotide, a double-stranded oligonucleotide, or a double-stranded DNA plasmid. The efficiency of HDR is generally low (less than 10% modified alleles) even in cells expressing Cas9, gRNA, and the exogenous repair template. Since HDR occurs during the S and G2 phases of the cell cycle, the efficiency of HDR can be increased by synchronizing the cells. Chemically or genetically inhibiting genes involved in NHEJ can also increase HDR frequency.

[0244] In some embodiments, Cas9 is a modified Cas9. A given gRNA target sequence can have additional sites of partial homology across the genome. These sites are called off-targets and need to be considered when designing the gRNA, but in addition to optimizing the design of the gRNA, the specificity of CRISPR can also be enhanced by modifying Cas9. Cas9 generates double-strand breaks (DSBs) through the combined activity of two nuclease domains, RuvC and HNH. Cas9 nickase, which is a D10A mutant of SpCas9, retains one nuclease domain and generates a DNA nick instead of a DSB. Nickases can also be combined with HDR-mediated gene editing for specific gene editing.

[0245] In some embodiments, Cas9 is a variant Cas9 protein. The variant Cas9 polypeptide has an amino acid sequence that differs by one amino acid unit (e.g., having a deletion, insertion, substitution, fusion) compared to the amino acid sequence of the wild-type Cas9 protein. In some examples, the variant Cas9 polypeptide has an amino acid change (e.g., a deletion, insertion, or substitution) that reduces the nuclease activity of the Cas9 polypeptide. For example, in some examples, the variant Cas9 polypeptide has less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nuclease activity of the corresponding wild-type Cas9 protein. In some embodiments, the variant Cas9 protein has no substantial nuclease activity. If the subject Cas9 protein is a variant Cas9 protein that has no substantial nuclease activity, it can be referred to as "dCas9".

[0246] In some embodiments, the variant Cas9 protein reduces nuclease activity. For example, the variant Cas9 protein exhibits less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% of the endonuclease activity of the wild-type Cas9 protein (e.g., wild-type Cas9 protein).

[0247] In some embodiments, the variant Cas9 protein can cleave the complementary strand of the guide-target sequence, but has a reduced ability to cleave the non-complementary strand of the double-stranded guide-target sequence. For example, the variant Cas9 protein can have a mutation (amino acid substitution) that reduces the function of the RuvC domain. As a non-limiting example, in some embodiments, the variant Cas9 protein has D10A (from aspartic acid to alanine at amino acid position 10), and thus can cleave the complementary strand of the double-stranded guide-target sequence, but has a reduced ability to cleave the non-complementary strand of the double-stranded guide-target sequence (thus, when this variant Cas9 protein cleaves a double-stranded target nucleic acid, a single-strand break (SSB) occurs instead of a double-strand break (DSB)) (see, e.g., Jinek et al., Science. 2012 Aug. 17; 337(6096):816-21).

[0248] In some embodiments, the variant Cas9 protein can cleave the non-complementary strand of the double-stranded guide-target sequence, but has a reduced ability to cleave the complementary strand of the guide-target sequence. For example, the variant Cas9 protein can have a mutation (amino acid substitution) that reduces the function of the HNH domain (RuvC / HNH / RuvC domain motif). As a non-limiting example, in some embodiments, the variant Cas9 protein has an H840A (from histidine to alanine at amino acid position 840) mutation, and thus can cleave the non-complementary strand of the guide-target sequence, but has a reduced ability to cleave the complementary strand of the guide-target sequence (thus, when this variant Cas9 protein cleaves a double-stranded guide-target sequence, an SSB occurs instead of a DSB). Such Cas9 proteins have a reduced ability to cleave a guide-target sequence (e.g., a single-stranded guide-target sequence), but retain the ability to bind to a guide-target sequence (e.g., a single-stranded guide-target sequence).

[0249] In some embodiments, the variant Cas9 protein has a reduced ability to cleave both the complementary and non-complementary strands of double-stranded target DNA. As a non-limiting example, in some embodiments, the variant Cas9 protein has both D10A and H840A mutations, such that the polypeptide has a reduced ability to cleave both the complementary and non-complementary strands of double-stranded target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA).

[0250] As another non-limiting example, in some embodiments, the variant Cas9 protein has W476A and W1126A mutations, such that the polypeptide has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA).

[0251] As another non-limiting example, in some embodiments, the variant Cas9 protein has P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, such that the polypeptide has a reduced ability to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA).

[0252] As another non-limiting example, in some embodiments, the variant Cas9 protein has H840A, W476A, and W1126A mutations, such that the polypeptide has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some embodiments, the variant Cas9 protein has H840A, D10A, W476A, and W1126A mutations, such that the polypeptide has a reduced ability to cleave target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, the variant Cas9 has a restored catalytic His residue at position 840 of the Cas9 HNH domain (A840H).

[0253] As another non-limiting example, in some embodiments, the variant Cas9 protein has the H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, such that the polypeptide has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some embodiments, the variant Cas9 protein has the D10A, H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, such that the polypeptide has a reduced ability to cleave target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, when the variant Cas9 protein has the W476A and W1126A mutations, or when the variant Cas9 protein has the P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, the variant Cas9 protein does not efficiently bind to the PAM sequence. Thus, in some such embodiments, when such a variant Cas9 protein is used in a binding method, the method does not require a PAM sequence. In other words, in some embodiments, when such a variant Cas9 protein is used in a binding method, the method may include a guide RNA, but the method can be performed in the absence of a PAM sequence (thus, the binding specificity is provided by the target segment of the guide RNA). To achieve the above effects, other residues can be mutated (i.e., inactivate one or the other nuclease moiety). As non-limiting examples, the residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 can be modified (i.e., substituted). Also, mutations other than alanine substitutions are suitable.

[0254] In some embodiments, a variant Cas9 protein having reduced catalytic activity (e.g., when the Cas9 protein has D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 mutations, such as D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A) can bind to target DNA in a site-specific manner as long as it retains its ability to interact with the guide RNA (because it is directed to the target DNA sequence by the guide RNA).

[0255] In some embodiments, the variant Cas protein can be spCas9, spCas9-VRQR, spCas9-VRER, xCas9 (sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK, or spCas9-LRVSQL.

[0256] In some embodiments, a modified SpCas9 containing the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (SpCas9-MQKFRAER) and having specificity for the modified PAM 5'-NGC-3' was used.

[0257] As an alternative to S. pyogenes Cas9, RNA-guided endonucleases derived from the Cpf1 family that exhibit cleavage activity in mammalian cells can be mentioned. CRISPR from Prevotella and Francisella 1 (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. Cpf1 is an RNA-guided endonuclease of the class II CRISPR / Cas system. This acquired immune mechanism is found in Prevotella and Francisella bacteria. The Cpf1 gene is associated with the CRISPR locus and encodes an endonuclease that uses guide RNA to find and cleave viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9 and overcomes some of the limitations of the CRISPR / Cas9 system. Unlike the Cas9 nuclease, the result of DNA cleavage via Cpf1 is a double-strand break with a short 3' overhang. The staggered cleavage pattern of Cpf1 can open up the possibility of directional gene insertion similar to traditional restriction enzyme cloning, which can enhance the efficiency of gene editing. Similar to the variants and orthologs of Cas9 described above, Cpf1 can also expand the number of sites that CRISPR can target to AT-rich regions lacking the NGG PAM site preferred by SpCas9 or to AT-rich genomes. The Cpf1 locus contains an alpha / beta mixed domain, a RuvC-I followed by a helical region, a RuvC-II, and a zinc finger-like domain. The Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9. Furthermore, Cpf1 does not have an HNH endonuclease region, and the N-terminus of Cpf1 does not have the alpha-helix recognition lobe of Cas9. The Cpf1 CRISPR-Cas domain composition indicates that Cpf1 is functionally unique and is classified as a class 2, type V CRISPR system. The Cpf1 locus encoded Cas1, Cas2, and Cas4 proteins that were more similar to type I and type III than to type II systems. Functional Cpf1 does not require trans-activating CRISPR RNA (tracrRNA); thus, only CRISPR (crRNA) is required.Not only is Cpf1 smaller than Cas9, but it also has a smaller sgRNA molecule (about half the number of nucleotides of Cas9), which is beneficial for genome editing. In contrast to the G-rich PAM targeted by Cas9, the Cpf1-crRNA complex cleaves target DNA or RNA by the identification of a protospacer adjacent to the motif 5'-YTN-3'. After the identification of the PAM, Cpf1 introduces a sticky-end-like DNA double-strand break with a 4- or 5-nucleotide overhang.

[0258] In some embodiments, Cas9 is a Cas9 variant that is specific for a modified PAM sequence. In some embodiments, additional Cas9 variants and PAM sequences are described in Miller, S.M., et al. Continuous evolution of SpCas9 variants compatible with non-G PAMs, Nat. Biotechnol. (2020) (which is incorporated herein by reference in its entirety). In some embodiments, the Cas9 variant does not have a specific PAM requirement. In some embodiments, the Cas9 variant, e.g., the SpCas9 variant, is specific for an NRNH PAM, where R is A or G and H is A, C, or T. In some embodiments, the SpCas9 variant is specific for the PAM sequences AAA, TAA, CAA, GAA, TAT, GAT, or CAC. In some embodiments, the SpCas9 variant comprises an amino acid substitution at position 1114, 1134, 1135, 1137, 1139, 1151, 1180, 1188, 1211, 1218, 1219, 1221, 1249, 1256, 1264, 1290, 1318, 1317, 1320, 1321, 1323, 1332, 1333, 1335, 1337, or 1339, or the corresponding position in the numbering of SEQ ID NO: 1. In some embodiments, the SpCas9 variant comprises an amino acid substitution at position 1114, 1135, 1218, 1219, 1221, 1249, 1320, 1321, 1323, 1332, 1333, 1335, or 1337, or the corresponding position in the numbering of SEQ ID NO: 1. In some embodiments, the SpCas9 variant comprises an amino acid substitution at position 1114, 1134, 1135, 1137, 1139, 1151, 1180, 1188, 1211, 1219, 1221, 1256, 1264, 1290, 1318, 1317, 1320, 1323, 1333, or the corresponding position in the numbering of SEQ ID NO: 1.In some embodiments, the SpCas9 variant comprises amino acid substitutions at positions 1114, 1131, 1135, 1150, 1156, 1180, 1191, 1218, 1219, 1221, 1227, 1249, 1253, 1286, 1293, 1320, 1321, 1332, 1335, 1339, or their corresponding positions as numbered in SEQ ID NO: 1. In some embodiments, the SpCas9 variant comprises amino acid substitutions at positions 1114, 1127, 1135, 1180, 1207, 1219, 1234, 1286, 1301, 1332, 1335, 1337, 1338, 1349, or their corresponding positions as numbered in SEQ ID NO: 1. Exemplary amino acid substitutions and PAM specificities of the SpCas9 variant are shown in Tables 1A - 1D.

[0259] Table 1A

Table 1-1

[0260] Table 1B

Table 1-2

[0261] Table 1C

Table 1-3

[0262] Table 1D

Table 1-4

[0263] In some embodiments, Cas9 is Neisseria menigitidis Cas9 (NmeCas9) or a variant thereof. In some embodiments, NmeCas9 is specific for the NNNNGAYW PAM, where Y is C or T and W is A or T. In some embodiments, NmeCas9 is specific for the NNNNGYTT PAM, where Y is C or T. In some embodiments, NmeCas9 is specific for the NNNNGTCT PAM. In some embodiments, NmeCas9 is Nme1 Cas9. In some embodiments, NmeCas9 is specific for the NNNNGATT PAM, NNNNCCTA PAM, NNNNCCTC PAM, NNNNCCTT PAM, NNNNCCTG PAM, NNNNCCGT PAM, NNNNCCGGPAM, NNNNCCCA PAM, NNNNCCCT PAM, NNNNCCCC PAM, NNNNCCAT PAM, NNNNCCAG PAM, NNNNCCAT PAM, or NNNGATT PAM. In some embodiments, Nme1Cas9 is specific for the NNNNGATT PAM, NNNNCCTA PAM, NNNNCCTC PAM, NNNNCCTT PAM, or NNNNCCTG PAM. In some embodiments, NmeCas9 is specific for the CAA PAM, CAAA PAM, or CCA PAM. In some embodiments, NmeCas9 is Nme2 Cas9. In some embodiments, NmeCas9 is specific for the NNNNCC (N4CC) PAM, where N is any one of A, G, C, or T. In some embodiments, NmeCas9 is specific for the NNNNCCGT PAM, NNNNCCGGPAM, NNNNCCCA PAM, NNNNCCCT PAM, NNNNCCCC PAM, NNNNCCAT PAM, NNNNCCAG PAM, NNNNCCAT PAM, or NNNGATT PAM. In some embodiments, NmeCas9 is Nme3Cas9.In some embodiments, NmeCas9 is specific for the NNNNCAAA PAM, NNNNCC PAM, or NNNNCNNN PAM. Additional NmeCas9 features and PAM sequences as described in Edraki et al. Mol. Cell. (2019) 73(4): 714-726 are hereby incorporated by reference in their entirety. Exemplary amino acid sequences of Nme1Cas9 are provided below:.

[0264] Type II CRISPR RNA-guided endonuclease Cas9 [Neisseria meningitidis] WP_002235162.1 1 maafkpnpin yilgldigia svgwamveid edenpiclid lgvrvferae vpktgdslam 61 arrlarsvrr ltrrrahrll rarrllkreg vlqaadfden glikslpntp wqlraaaldr 121 kltplewsav llhlikhrgy lsqrkneget adkelgallk gvadnahalq tgdfrtpael 181 alnkfekesg hirnqrgdys htfsrkdlqa elillfekqk efgnphvsgg lkegietllm 241 tqrpalsgda vqkmlghctf epaepkaakn tytaerfiwl tklnnlrile qgserpltdt 301 eratlmdepy rkskltyaqa rkllgledta ffkglrygkd naeastlmem kayhaisral 361 ekeglkdkks plnlspelqd eigtafslfk tdeditgrlk driqpeilea llkhisfdkf 421 vqislkalrr ivplmeqgkr ydeacaeiyg dhygkkntee kiylppipad eirnpvvlra 481 lsqarkving vvrrygspar ihietarevg ksfkdrkeie krqeenrkdr ekaaakfrey 541 fpnfvgepks kdilklrlye qqhgkclysg keinlgrlne kgyveidhal pfsrtwddsf 601 nnkvlvlgse nqnkgnqtpy eyfngkdnsr ewqefkarve tsrfprskkq rillqkfded 661 gfkernlndt ryvnrflcqf vadrmrltgk gkkrvfasng qitnllrgfw glrkvraend 721 rhhaldavvv acstvamqqk itrfvrykem nafdgktidk etgevlhqkt hfpqpweffa 781 qevmirvfgk pdgkpefeea dtpeklrtll aeklssrpea vheyvtplfv srapnrkmsg 841 qghmetvksa krldegvsvl rvpltqlklk dlekmvnrer epklyealka rleahkddpa 901 kafaepfyky dkagnrtqqv kavrveqvqk tgvwvrnhng iadnatmvrv dvfekgdkyy 961 lvpiyswqva kgilpdravv qgkdeedwql iddsfnfkfs lhpndlvevi tkkarmfgyf 1021 aschrgtgni nirihdldhk igkngilegi gvktalsfqk yqidelgkei rpcrlkkrpp 1081 vr

[0265] Exemplary amino acid sequences of Nme2Cas9 are provided below: Type II CRISPR RNA-guided endonuclease Cas9 [Neisseria meningitidis] WP_002230835.1 1 maafkpnpin yilgldigia svgwamveid eeenpirlid lgvrvferae vpktgdslam 61 arrlarsvrr ltrrrahrll rarrllkreg vlqaadfden glikslpntp wqlraaaldr 121 kltplewsav llhlikhrgy lsqrkneget adkelgallk gvannahalq tgdfrtpael 181 alnkfekesg hirnqrgdys htfsrkdlqa elillfekqk efgnphvsgg lkegietllm 241 tqrpalsgda vqkmlghctf epaepkaakn tytaerfiwl tklnnlrile qgserpltdt 301 eratlmdepy rkskltyaqa rkllgledta ffkglrygkd naeastlmem kayhaisral 361 ekeglkdkks plnlsselqd eigtafslfk tdeditgrlk drvqpeilea llkhisfdkf 421 vqislkalrr ivplmeqgkr ydeacaeiyg dhygkkntee kiylppipad eirnpvvlra 481 lsqarkving vvrrygspar ihietarevg ksfkdrkeie krqeenrkdr ekaaakfrey 541 fpnfvgepks kdilklrlye qqhgkclysg keinlvrlne kgyveidhal pfsrtwddsf 601 nnkvlvlgse nqnkgnqtpy eyfngkdnsr ewqefkarve tsrfprskkq rillqkfded 661 gfkecnlndt ryvnrflcqf vadhilltgk gkrrvfasng qitnllrgfw glrkvraend 721 rhhaldavvv acstvamqqk itrfvrykem nafdgktidk etgkvlhqkt hfpqpweffa 781 qevmirvfgk pdgkpefeea dtpeklrtll aeklssrpea vheyvtplfv srapnrkmsg 841 ahkdtlrsak rfvkhnekis vkrvwlteik ladlenmvny kngreielye alkarleayg 901 gnakqafdpk dnpfykkggq lvkavrvekt qesgvllnkk naytiadngd mvrvdvfckv 961 dkkgknqyfi vpiyawqvae nilpdidckg yriddsytfc fslhkydlia fqkdekskve 1021 fayyincdss ngrfylawhd kgskeqqfri stqnlvliqk yqvnelgkei rpcrlkkrpp 1081 vr

[0266] [Cas12 Domain of Nucleic Acid Base Editor] Typically, microbial CRISPR-Cas systems are divided into class 1 and class 2 systems. Class 1 systems have multi-subunit effector complexes, and class 2 systems have a single protein effector. For example, Cas9 and Cpf1 are class 2 effectors, although of different types (type II and type V, respectively). In addition to Cpf1, class 2, type V CRISPR-Cas systems also include Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i (see, for example, Shmakov et al., “Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems,” Mol. Cell, 2015 Nov. 5; 60(3): 385-397; Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR Journal, 2018, 1(5): 325-336; and Yan et al., “Functionally Diverse Type V CRISPR-Cas Systems,” Science, 2019 Jan. 4; 363: 88-91, the entire contents of each of which are incorporated herein by reference). Type V Cas proteins contain a RuvC (or RuvC-like) endonuclease domain. The production of mature CRISPR RNA (crRNA) is generally tracrRNA-independent, although, for example, Cas12b / C2c1 requires tracrRNA for crRNA production. Cas12b / C2c1 depends on both crRNA and tracrRNA for DNA cleavage.

[0267] Nucleic acid programmable DNA binding proteins contemplated in the present disclosure include Cas proteins classified as Class 2, Type V (Cas12 proteins). Non-limiting examples of Cas Class 2, Type V proteins include Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i, homologs thereof, or modified versions thereof. As used herein, Cas12 proteins may also be referred to as Cas12 nucleases, Cas12 domains, or Cas12 protein domains. In some embodiments, the Cas12 proteins of the present disclosure include amino acid sequences flanked by internally fused protein domains such as deaminase domains.

[0268] In some embodiments, the Cas12 domain is a nuclease-inactive Cas12 domain or a Cas12 nickase. In some embodiments, the Cas12 domain is a nuclease-active domain. For example, the Cas12 domain can be a Cas12 domain that nicks one strand of a double-stranded nucleic acid (e.g., a double-stranded DNA molecule). In some embodiments, the Cas12 domain comprises any one of the amino acid sequences described herein. In some embodiments, the Cas12 domain comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the amino acid sequences described herein. In some embodiments, the Cas12 domain comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any one of the amino acid sequences described herein. In some embodiments, the Cas12 domain comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical consecutive amino acid residues compared to any one of the amino acid sequences described herein.

[0269] In some embodiments, proteins comprising fragments of Cas12 are provided. For example, in some embodiments, the protein comprises one of two Cas12 domains: (1) the gRNA binding domain of Cas12, or (2) the DNA cleavage domain of Cas12. In some embodiments, a protein comprising Cas12 or a fragment thereof is referred to as a "Cas12 variant." A Cas12 variant shares homology to Cas12, or a fragment thereof. For example, a Cas12 variant has at least about 70% identity, at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity to wild-type Cas12. In some embodiments, a Cas12 variant can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type Cas12. In some embodiments, a Cas12 variant comprises a fragment of Cas12 (e.g., the gRNA binding domain or the DNA cleavage domain), and the fragment has at least about 70% identity, at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity to the corresponding fragment of wild-type Cas12.In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid length of the corresponding wild-type Cas12. In some embodiments, the fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids in length.

[0270] In some embodiments, Cas12 corresponds to, or partially or wholly comprises, a Cas12 amino acid sequence having one or more mutations that modify Cas12 nuclease activity. Examples of such mutations include amino acid substitutions within the RuvC nuclease domain of Cas12. In some embodiments, variants or homologs of Cas12 are provided that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to wild-type Cas12. In some embodiments, variants of Cas12 are provided that have amino acid sequences that are shorter or longer by about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids or more.

[0271] In some embodiments, the Cas12 fusion proteins provided herein include the full-length amino acid sequence of a Cas12 protein, e.g., one of the Cas12 sequences provided herein. In other embodiments, however, the fusion proteins provided herein do not include the full-length Cas12 sequence, but rather include only one or more fragments thereof. Exemplary amino acid sequences of suitable Cas12 domains are provided herein, and additional suitable sequences of Cas12 domains and fragments will be apparent to those of skill in the art.

[0272] Generally, class 2, type V Cas proteins have a single functional RuvC endonuclease domain (see, e.g., Chen et al., “CRISPR-Cas12a target binding unleashes indiscriminate single-stranded DNase activity,” Science 360:436-439 (2018)). In some cases, the Cas12 protein is a variant Cas12b protein. (See Strecker et al., Nature Communications, 2019, 10(1): Art. No.: 212). In one embodiment, the variant Cas12 polypeptide has an amino acid sequence that differs by 1, 2, 3, 4, 5, or more amino acids (e.g., having deletions, insertions, substitutions, fusions) when compared to the amino acid sequence of the wild-type Cas12 protein. In some cases, the variant Cas12 polypeptide has an amino acid change (e.g., a deletion, insertion, or substitution) that reduces the activity of the Cas12 polypeptide. For example, in some cases, the variant Cas12 is a Cas12b polypeptide having less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nickase activity of the corresponding wild-type Cas12b protein. In some cases, the variant Cas12b protein has substantially no nickase activity.

[0273] In some cases, the variant Cas12b protein has reduced nickase activity. For example, the variant Cas12b protein exhibits less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% of the nickase activity of the wild-type Cas12b protein.

[0274] In some embodiments, the Cas12 protein comprises an RNA-guided endonuclease from the Cas12a / Cpf1 family that is active in mammalian cells. CRISPR from Prevotella and Francisella 1 (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. Cpf1 is an RNA-guided endonuclease of the class II CRISPR / Cas system. This acquired immune mechanism is found in Prevotella and Francisella bacteria. The Cpf1 gene is associated with the CRISPR locus and encodes an endonuclease that uses guide RNA to find and cleave viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9 and overcomes some of the limitations of the CRISPR / Cas9 system. Unlike the Cas9 nuclease, the result of DNA cleavage via Cpf1 is a double-strand break with short 3' overhangs. The staggered cleavage pattern of Cpf1 can open up the possibility of directional gene insertion similar to traditional restriction enzyme cloning, which can enhance the efficiency of gene editing. Similar to the variants and orthologs of Cas9 described above, Cpf1 can also expand the number of sites that CRISPR can target to AT-rich regions lacking the NGG PAM site preferred by SpCas9 or to AT-rich genomes. The Cpf1 locus contains an alpha / beta mixed domain, an RuvC-I followed by a helical region, an RuvC-II, and a zinc finger-like domain. The Cpf1 protein has an RuvC-like endonuclease domain similar to the RuvC domain of Cas9. Furthermore, unlike Cas9, Cpf1 does not have an HNH endonuclease region, and the N-terminus of Cpf1 does not have the alpha helix recognition lobe of Cas9. The Cpf1 CRISPR-Cas domain composition indicates that Cpf1 is functionally unique and is classified as a class 2, type V CRISPR system. The Cpf1 locus encoded Cas1, Cas2, and Cas4 proteins that were more similar to type I and type III than to type II systems.Functional Cpf1 does not require trans-activating CRISPR RNA (tracrRNA); thus, it only requires CRISPR (crRNA). Cpf1 is not only smaller than Cas9 but also has a smaller sgRNA molecule (about half the number of nucleotides of Cas9), which is beneficial for genome editing. In contrast to the G-rich PAM targeted by Cas9, the Cpf1-crRNA complex cleaves target DNA or RNA by the identification of a protospacer adjacent to the motif 5'-YTN-3' or 5'-TTTN-3'. After the identification of the PAM, Cpf1 introduces a sticky-end-like DNA double-strand break with a 4- or 5-nucleotide overhang.

[0275] In some aspects of the present disclosure, a vector encoding a CRISPR enzyme mutated relative to a corresponding wild-type enzyme can be used to lack the ability to cleave one or both strands of a target polynucleotide containing a target sequence. Cas12 can refer to a polypeptide having at least, or at least approximately, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology with an exemplary wild-type Cas12 polypeptide (e.g., Cas12 from Bacillus hisashii). Cas12 can refer to a polypeptide having at most, or at most approximately, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology with an exemplary wild-type Cas12 polypeptide (e.g., derived from Bacillus hisashii (BhCas12b), Bacillus sp. V3-13 (BvCas12b), and Alicyclobacillus acidiphilus (AaCas12b)). Cas12 can refer to a modified form of a Cas12 protein that may include amino acid changes such as wild-type, or deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof.

[0276] [Nucleic acid programmable DNA binding protein] Some aspects of the present disclosure provide fusion proteins that include a domain that acts as a nucleic acid programmable DNA binding protein, which can be used to direct a protein, such as a base editor, to a specific nucleic acid (e.g., DNA or RNA) sequence. In certain embodiments, the fusion protein includes a nucleic acid programmable DNA binding protein domain and a deaminase domain. Non-limiting examples of nucleic acid programmable DNA binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i. Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V Cas effector proteins, type VI Cas effector proteins, CARF, DinG, homologs thereof, or modified or engineered versions thereof. Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, although they may not be specifically listed in the present disclosure.See, for example, Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct;1:325-336. doi: 10.1089 / crispr.2018.0033; Yan et al., “Functionally diverse type V CRISPR-Cas systems” Science. 2019 Jan 4;363(6422):88-91. doi: 10.1126 / science.aav7271 (the entire contents of each are incorporated herein by reference).

[0277] An example of a nucleic acid programmable DNA binding protein having a PAM specificity different from Cas9 is Clustered Regularly Interspaced Short Palindromic Repeats (Cpf1) from Prevotella and Francisella 1. Similar to Cas9, Cpf1 is also a class 2 CRISPR effector. Cpf1 has been shown to mediate robust DNA interference with characteristics different from Cas9. Cpf1 is a single RNA-guided endonuclease lacking tracrRNA and utilizes a protospacer adjacent motif rich in T (TTN, TTTN, or YTN). Furthermore, Cpf1 cleaves DNA with staggered DNA double-strand breaks. Among the 16 Cpf1 family proteins, two enzymes from Acidaminococcus and Lachnospiraceae have been shown to have efficient genome editing activity in human cells. Cpf1 proteins are known in the art and have been described, for example, in the past by Yamano et al., “Crystal structure of Cpf1 in complex with guide RNA and target DNA.” Cell (165) 2016, p. 949-962, the entire contents of which are incorporated herein by reference.

[0278] Nuclease-inactive Cpf1 (dCpf1) variants that can be used as guide nucleotide sequence-programmable DNA-binding protein domains are also useful in the compositions and methods of the present invention. The Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9, but does not have an HNH endonuclease domain, and the N-terminus of Cpf1 does not have the alpha-helix recognition lobe of Cas9. Zetsche et al., Cell, 163, 759-771, 2015 (incorporated herein by reference) showed that the RuvC-like domain of Cpf1 is responsible for cleavage of both DNA strands and inactivation of the RuvC-like domain inactivates Cpf1 nuclease activity. For example, mutations corresponding to D917A, E1006A, or D1255A in Francisella novicida Cpf1 inactivate Cpf1 nuclease activity. In some embodiments, the dCpf1 of the present disclosure comprises mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A. It should be understood that any mutation that inactivates the RuvC domain of Cpf1, such as a substitution mutation, deletion, or insertion, can be used according to the present disclosure.

[0279] In some embodiments, any nucleic acid programmable DNA binding protein (napDNAbp) of the fusion proteins provided herein can be a Cpf1 protein. In some embodiments, the Cpf1 protein is a Cpf1 nickase (nCpf1). In some embodiments, the Cpf1 protein is a nuclease-inactive Cpf1 (dCpf1). In some embodiments, the Cpf1, nCpf1, or dCpf1 comprises an amino acid sequence having at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to the Cpf1 sequences disclosed herein. In some embodiments, the dCpf1 comprises an amino acid sequence having at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to the Cpf1 sequences disclosed herein and comprises a mutation corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A. It should be understood that Cpf1s from other bacterial species can also be used in accordance with the present disclosure.

[0280] Wild-type Francisella novicida Cpf1 (with D917, E1006, and D1255 in bold and underlined) MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDDKAIKENKGEGYKKIVYKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISESYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI DRGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF E DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDYKNFGDKAAKGKWTIASFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA D ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN

[0281] Francisella novicida Cpf1 D917A (with A917, E1006, and D1255 in bold and underlined) MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDDKAIKENKGEGYKKIVYKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISESYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI ARGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF E DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDYKNFGDKAAKGKWTIASFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA D ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN

[0282] Francisella novicida Cpf1 E1006A (with D917, A1006, and D1255 in bold and underlined) MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDDKAIKENKGEGYKKIVYKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISESYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI DRGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF A DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDYKNFGDKAAKGKWTIASFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA D ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN

[0283] Francisella novicida Cpf1 D1255A (D917, E1006, and A1255 are in bold and underlined) MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDDKAIKENKGEGYKKIVYKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISESYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI DRGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF E DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDYKNFGDKAAKGKWTIASFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA A ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN

[0284] Francisella novicida Cpf1 D917A / E1006A (A917, A1006, and D1255 are in bold and underlined) MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDDKAIKENKGEGYKKIVYKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISESYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI ARGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF A DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDYKNFGDKAAKGKWTIASFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA D ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN

[0285] Francisella novicida Cpf1 D917A / D1255A (bold and underlined are A917, E1006, and A1255) MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDDKAIKENKGEGYKKIVYKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISESYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI ARGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF E DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDYKNFGDKAAKGKWTIASFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA A ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN

[0286] Francisella novicida Cpf1 E1006A / D1255A (D917, A1006, and A1255 are in bold and underlined) MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDDKAIKENKGEGYKKIVYKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISESYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI DRGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF A DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDYKNFGDKAAKGKWTIASFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA A ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN

[0287] Francisella novicida Cpf1 D917A / E1006A / D1255A (A917, A1006, and A1255 are in bold and underlined) MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDDKAIKENKGEGYKKIVYKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISESYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI ARGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF A DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDYKNFGDKAAKGKWTIASFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA A ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN

[0288] In some embodiments, one of the Cas9 domains present in the fusion protein can be replaced with a guide nucleotide sequence programmable DNA binding protein domain that does not require a PAM sequence.

[0289] In some embodiments, the Cas9 domain is a Cas9 domain from Staphylococcus aureus (SaCas9). In some embodiments, the SaCas9 domain is nuclease-active SaCas9, nuclease-inactive SaCas9 (SaCas9d), or SaCas9 nickase (SaCas9n). In some embodiments, SaCas9 includes the N579A mutation, or a corresponding mutation in any of the amino acid sequences provided herein.

[0290] In some embodiments, the SaCas9 domain, the SaCas9d domain, or the SaCas9n domain can bind to a nucleic acid sequence having a non-standard PAM, and in some embodiments, the SaCas9 domain, the SaCas9d domain, or the SaCas9n domain can bind to a nucleic acid sequence having an NNGRRT or NNGRRT PAM sequence. In some embodiments, the SaCas9 domain includes one or more of the E781X, N967X, and R1014X mutations, or corresponding mutations in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SaCas9 domain includes one or more of the E781K, N967K, and R1014H mutations, or one or more corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SaCas9 domain includes the E781K, N967K, or R1014H mutation, or a corresponding mutation in any of the amino acid sequences provided herein.

[0291] Exemplary SaCas9 sequences KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHSELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKDGEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKEWYEMLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAINLILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVSFDNSFNNKVLVKQEE NSKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDFINRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLIVNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIKYYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAEFIASFYNNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKKYSTDILGNLYEVKSKKHPQIIKKG The above residue N579, which is underlined and in bold, can be mutated (e.g., to A579) to yield SaCas9 nickase.

[0292] Exemplary SaCas9n sequences KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHSELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKDGEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKEWYEMLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAINLILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVSFDNSFNNKVLVKQEE ASKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDFINRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLIVNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIKYYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAEFIASFYNNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKKYSTDILGNLYEVKSKKHPQIIKKG The residue A579 above, which can be mutated from N579 to result in SaCas9 nickase, is underlined and shown in bold.

[0293] Exemplary SaKKH Cas9 TIFF0007693552000019.tif129162The residue A579 above, which can be mutated from N579 to result in SaCas9 nickase, is underlined and shown in bold. The residues K781, K967, and H1014 above, which can be mutated from E781, N967, and R1014 to result in SaKKH Cas9, are underlined and shown in italics.

[0294] In some embodiments, napDNAbp is a circular permutant. In the following sequences, plain text indicates the adenosine deaminase sequence, bold sequences indicate sequences derived from Cas9, italic sequences indicate linker sequences, and underlined sequences indicate bipartite nuclear localization sequences. CP5 (with MSP "NGC" PID and "D10A" nickase): TIFF0007693552000020.tif179164

[0295] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is a single effector of the microbial CRISPR-Cas system. Non-limiting examples of single effectors of the microbial CRISPR-Cas system include Cas9, Cpf1, Cas12b / C2c1, and Cas12c / C2c3. Typically, the microbial CRISPR-Cas system is divided into class 1 and class 2 systems. Class 1 systems have multi-subunit effector complexes, and class 2 systems have single protein effectors. For example, Cas9 and Cpf1 are class 2 effectors. In addition to Cas9 and Cpf1, three different class 2 CRISPR-Cas systems (Cas12b / C2c1 and Cas12c / C2c3) are described by Shmakov et al., “Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems”, Mol. Cell, 2015 Nov. 5; 60(3): 385-397 (the entire content of which is incorporated herein by reference). The effectors of two systems, Cas12b / C2c1 and Cas12c / C2c3, contain a RuvC-like endonuclease region related to Cpf1. The third system includes an effector with two predicted HEPN RNase domains. Unlike the production of CRISPR RNA by Cas12b / C2c1, the production of mature CRISPR RNA is tracrRNA-independent. Cas12b / C2c1 depends on both CRISPRRNA and tracrRNA for DNA cleavage.

[0296] The crystal structure of Alicyclobaccillus acidoterrastris Cas12b / C2c1 (AacC2c1) has been reported as a complex with a chimeric single-guide RNA (sgRNA). See, for example, Liu et al., “C2c1-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism”, Mol. Cell, Jan. 19, 2017; 65(2):310-322 (incorporated herein by reference in its entirety). Also, the crystal structure has been reported for Alicyclobacillus acidoterrestris C2c1 bound to target DNA as a ternary complex. See, for example, Yang et al., “PAM-dependent Target DNA Recognition and Cleavage by C2C1 CRISPR-Cas endonuclease”, Cell, Dec. 15, 2016; 167(7):1814-1828 (the entire content of which is incorporated herein by reference). Together with both the target DNA strand and the non-target DNA strand, the catalytically competent conformation of AacC2c1 is independently captured within one RuvC catalytic pocket, and Cas12b / C2c1-mediated cleavage results in the cleavage of a 7-nucleotide stagger of the target DNA. A structural comparison between the Cas 12b / C2c1 ternary complex and previously identified Cas9 and Cpf1 counterparts shows the diversity of the mechanisms used by the CRISPR-Cas9 system.

[0297] In some embodiments, any nucleic acid programmable DNA binding protein (napDNAbp) of the fusion proteins provided herein can be a Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a Cas12b / C2c1 protein. In some embodiments, the napDNAbp is a Cas12c / C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence having at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence having at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to any of the napDNAbp sequences provided herein. It should be understood that Cas12b / C2c1 or Cas12c / C2c3 from other bacterial species can also be used according to the present disclosure.

[0298] The amino acid sequence of Cas12b / C2c1 ((...

Claims

1. A nucleic acid programmable DNA binding protein (napDNAbp) domain, and the following TadA*7.10 amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD An adenosine deaminase variant comprising the above TadA*7.10 amino acid sequence or a fragment of the above TadA*7.10 amino acid sequence lacking only the N-terminal methionine, wherein the fusion protein provided that the adenosine deaminase variant comprises an amino acid modification consisting of a single amino acid modification selected from Y147T, Y147R, Q154S, V82S, and T166R, or the adenosine deaminase variant Y123H, Y147R, and Q154R; I76Y, Y147R, and Q154R; Y147R, Q154R, and T166R; Y147T and Q154R; Y147T and Q154S; I76Y, Y123H, Y147R, and Q154R; I76Y and V82S; V82S and Y147R; V82S, Y123H, and Y147R; V82S and Q154R; V82S, Y123H, and Q154R; V82S, Y123H, Y147R, and Q154R; and I76Y, V82S, Y123H, Y147R, and Q154R comprising a plurality of amino acid modifications consisting of a single amino acid modification combination selected from the group consisting of, wherein the numbering of said modifications is based on the TadA*7.10 amino acid sequence, Fusion protein.

2. said napDNAbp domain has the following sequence: comprising, wherein the bold sequence indicates the sequence derived from Cas9, the italicized sequence indicates the linker sequence, and the underlined sequence indicates the bipartite nuclear localization sequence, the fusion protein according to claim 1.

3. said napDNAbp domain comprises a Cas9 domain comprising an active, inactive, or partially active DNA cleavage domain and a gRNA binding domain, the fusion protein according to claim 1.

4. said napDNAbp domain has the following amino acid sequence: (wherein the HNH domain is indicated by a single underline and the RuvC domain is indicated by a double underline), the fusion protein according to claim 3.

5. said Cas9 domain is Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), or Streptococcus pyogenes Cas9 (SpCas9), the fusion protein according to claim 3 or 4.

6. (i) said Cas9 domain is SpCas9 having a modified protospacer adjacent motif (PAM) specificity or specificity for non-G PAM; or (ii) said Cas9 domain is SpCas9 having protospacer adjacent motif (PAM) specificity for the nucleic acid sequence 5'-NGC-3', The fusion protein according to any one of claims 3 to 5.

7. (i) whether the napDNAbp domain is a nuclease-inactive variant; or (ii) whether the napDNAbp domain is a nickase variant; or (iii) whether the napDNAbp domain is a nickase variant containing the amino acid substitution D10A, The fusion protein according to any one of claims 1 to 6.

8. (i) whether it contains a linker between the napDNAbp domain and the adenosine deaminase variant; or (ii) whether it contains a linker between the napDNAbp domain and the adenosine deaminase variant, and the linker contains the amino acid sequence: SGGSSGGSSGSETPGTSESATPES, The fusion protein according to any one of claims 1 to 7.

9. (i) whether it contains one or more nuclear localization signals; or (ii) whether it contains one or more nuclear localization signals, and the nuclear localization signal is a bipartite nuclear localization signal, The fusion protein according to any one of claims 1 to 8.

10. A polynucleotide encoding the fusion protein according to any one of claims 1 to 9.

11. An expression vector containing the polynucleotide according to claim 10.

12. It is a mammalian expression vector; It is a viral vector; It is a viral vector selected from the group consisting of adeno-associated virus (AAV), retroviral vector, adenoviral vector, lentiviral vector, Sendai virus vector, and herpesvirus vector; and / or It contains a promoter, The expression vector according to claim 11.

13. A cell comprising the fusion protein according to any one of claims 1 to 9, the polynucleotide according to claim 10, or the expression vector according to claim 11 or 12.

14. The cell according to claim 13, wherein the cell is a bacterial cell, a plant cell, an insect cell, or a mammalian cell.

15. A base editor comprising the fusion protein according to any one of claims 1 to 9 in a complex with one or more guide polynucleotides.

16. A pharmaceutical composition comprising the fusion protein according to any one of claims 1 to 9, the polynucleotide according to claim 10, the expression vector according to claim 11 or 12, the cell according to claim 13 or 14, or the base editor according to claim 15, and a pharmaceutically acceptable excipient.

17. An in vitro or ex vivo base editing method comprising contacting a target polynucleotide sequence with the fusion protein according to any one of claims 1 to 9, wherein the adenosine deaminase variant deaminates a nucleobase in the target polynucleotide, thereby editing the polynucleotide sequence.

18. An in vitro or ex vivo method of editing a target polynucleotide sequence, comprising contacting the target polynucleotide sequence with the base editor according to claim 15 to effect deamination of a nucleobase in the target polynucleotide, thereby causing a modification from A to G in the target polynucleotide sequence.

19. The method according to claim 17 or 18, further comprising contacting the target polynucleotide sequence with one or more guide polynucleotides to effect deamination of the nucleobase.

20. The method according to any one of claims 17 to 19, wherein the contacting is performed in a cell.

21. A pharmaceutical composition for correcting genetic defects in a subject, comprising a base editor containing or consisting of the fusion protein according to any one of claims 1 to 9, or a polynucleotide encoding the base editor, and one or more guide polynucleotides for inducing the base editor to deaminate a target nucleobase in a target nucleotide sequence.

22. The one or more guide polynucleotides are a) GACCUAGGCGAGGCAGUAGG; b) CCAGUAUGGACACUGUCCAAA; c) CAGUAUGGACACUGUCCAAA; d) AGUAUGGACACUGUCCAAAG; and e) GUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG The pharmaceutical composition according to claim 21, comprising a nucleic acid sequence selected from the group consisting of.

23. The base editor, or the polynucleotide encoding the base editor, and the one or more guide polynucleotides are delivered to a cell, and the pharmaceutical composition according to claim 21 or 22.

24. The pharmaceutical composition according to claim 23, wherein the cell is a mammalian cell or a human cell.

25. The pharmaceutical composition according to any one of claims 21 to 24, wherein the deamination of the target nucleobase replaces the target nucleobase with a wild-type nucleobase.

26. The pharmaceutical composition according to any one of claims 21 to 24, wherein the deamination of the target nucleobase replaces the target nucleobase with a non-wild-type nucleobase, and the deamination of the target nucleobase improves the symptoms of the genetic state associated with the genetic defect.

27. The one or more guide polynucleotides are a) GACCUAGGCGAGGCAGUAGG; b) CCAGUAUGGACACUGUCCAAA; c) CAGUAUGGACACUGUCCAAA; d) AGUAUGGACACUGUCCAAAG; and e) GUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG The method according to claim 19, comprising a nucleic acid sequence selected from the group consisting of

Citation Information

Patent Citations

  • Adenosine nucleobase editors and uses thereof

    WO2018027078A1

Cited By

  • Adenosine deaminase base editor and methods for modifying nucleobase in target sequence using the same

    JP2025143279A