Novel Nucleic Acid Base Editor and Method of Using the Same

The ABE9 base editor addresses the inefficiencies of current base editors by incorporating specific amino acid modifications, enabling precise and efficient conversion of A·T to G·C base pairs for therapeutic applications.

JP7717684B2Active Publication Date: 2025-08-04BEAM THERAPEUTICS INC

Patent Information

Application Number
JP2022514994
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-02-13
Filing Date
2020-09-09
Publication Date
2025-08-04
Estimated Expiration
2040-09-09

AI Technical Summary

Technical Problem

Current base editors lack the specificity and efficiency needed for targeted nucleic acid modifications, particularly in the conversion of A·T to G·C base pairs for therapeutic applications.

Method used

A novel programmable nucleic acid base editor, ABE9, is developed with specific amino acid modifications at defined positions, combined with a DNA binding domain, enhancing the specificity and efficiency of converting A·T to G·C base pairs.

Benefits of technology

ABE9 achieves precise and efficient editing of polynucleotides, particularly in addressing pathogenic mutations, with reduced indel formation and improved targeting capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007717684000116
    Figure 0007717684000116
  • Figure 0007717684000117
    Figure 0007717684000117
  • Figure 0007717684000118
    Figure 0007717684000118
Patent Text Reader

Abstract

The present invention features novel programmable nucleobase editors comprising an adenosine deaminase domain and methods of use thereof for polynucleotide editing. In some embodiments, the programmable nucleobase editors edit pathogenic mutations associated with genetic diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications This application is an international PCT application claiming priority and benefit of U.S. Provisional Application No. 62 / 897,777, filed on September 9, 2019, and claiming priority of International PCT Application No. PCT / US2020 / 018195, filed on February 13, 2020, the entire contents of which are hereby incorporated by reference in their entirety.

Background Art

[0002] Targeted editing of nucleic acid sequences, for example, targeted cleavage of genomic DNA or targeted introduction of specific modifications (alterations), is a very promising approach for studying gene function and can provide new therapeutic methods for human genetic diseases. Currently available base editors include cytidine base editors (such as BE4) that convert target C·G base pairs to T·A, and adenine base editors (such as ABE7.10) that convert A·T to G·C. In the art, there is a need for improved base editors that can induce modifications within a target sequence with higher specificity and efficiency.

Summary of the Invention

[0003] As described below, the present invention features a novel programmable nucleic acid base editor comprising an adenosine deaminase domain (e.g., TadA*9 or ABE9), and methods of using the same for polynucleotide editing. In some embodiments, ABE9 of the present invention edits polynucleotides, for example, polynucleotides containing pathogenic mutations associated with genetic diseases.

[0004] In one aspect, an adenosine deaminase comprising a modification at an amino acid position selected from the group consisting of 21, 23, 25, 38, 51, 54, 70, 71, 72, 73, 94, 124, 133, 139, 146, and 158 of SEQ ID NO: 1, or a corresponding modification in another adenosine deaminase: Provides JPEG0007717684000001.jpg51165. In one embodiment, adenosine deaminase comprises a modification selected from the group consisting of R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, M94V, P124W, T133K, D139L, D139M, C146R, and A158K of SEQ ID NO: 1, or a corresponding modification in another adenosine deaminase. In one embodiment, adenosine deaminase further comprises the V82T modification of SEQ ID NO: 1, or a corresponding modification in another adenosine deaminase. In one embodiment, adenosine deaminase comprises a modification at an amino acid position selected from the group consisting of 21, 23, 25, 38, 51, 54, 70, 71, 72, 73, 94, 124, 133, 139, 146, and 158 of SEQ ID NO: 1, or a corresponding modification in another adenosine deaminase. In one embodiment, the adenosine deaminase of this aspect and its embodiments comprises two or more modifications. In one embodiment, the adenosine deaminase of this aspect and its embodiments comprises three or more of the above modifications. In one embodiment, the adenosine deaminase of this aspect and its embodiments further comprises one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, and Q154R. In one embodiment, the adenosine deaminase of this aspect and its embodiments comprises any one of the following groups of modifications: E25F + V82S + Y123H; T133K + Y147R + Q154R; E25F + V82S + Y123H + Y147R + Q154R; L51W + V82S + Y123H + C146R + Y147R + Q154R; Y73S + V82S + Y123H + Y147R + Q154R; P54C + V82S + Y123H + Y147R + Q154R; N38G + V82T + Y123H + Y147R + Q154R; N72K + V82S + Y123H + D139L + Y147R + Q154R; E25F + V82S + Y123H + D139M + Y147R + Q154R; Q71M + V82S + Y123H + Y147R + Q154R; E25F + V82S + Y123H + T133K + Y147R + Q154R; E25F + V82S + Y123H + Y147R + Q154R; V82S + Y123H + P124W + Y147R + Q154R; L51W + V82S + Y123H + C146R + Y147R + Q154R; P54C + V82S + Y123H + Y147R + Q154R; Y73S + V82S + Y123H + Y147R + Q154R; N38G + V82T + Y123H + Y147R + Q154R; R23H + V82S + Y123H + Y147R + Q154R; R21N + V82S + Y123H + Y147R + Q154R; V82S + Y123H + Y147R + Q154R + A158K; N72K + V82S + Y123H + D139L + Y147R + Q154R; E25F + V82S + Y123H + D139M + Y147R + Q154R; M70V + V82S + M94V + Y123H + Y147R + Q154R; Q71M + V82S + Y123H + Y147R + Q154R; E25F + I76Y + V82S + Y123H + Y147R + Q154R; I76Y + V82T + Y123H + Y147R + Q154R; N38G + I76Y + V82S + Y123H + Y147R + Q154R; R23H + I76Y + V82S + Y123H + Y147R + Q154R; P54C + I76Y + V82S + Y123H + Y147R + Q154R; R21N + I76Y + V82S + Y123H + Y147R + Q154R; I76Y + V82S + Y123H + D139M + Y147R + Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; E25F+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82T+Y123H+Y147R+Q154R; N38G+I76Y+V82S+Y123H+Y147R+Q154R; R23H+I76Y+V82S+Y123H+Y147R+Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; V82S+Q154R; N72K+V82S+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R; N72K_V82S+Y123H+Y147R+Q154R; Q71M_V82S+Y123H+Y147R+Q154R; M70V+V82S+M94V+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; or M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R. In one embodiment, the adenosine deaminase variant comprises any modification or group of modifications described in Table 14 or 18. In one embodiment, the adenosine deaminase of this aspect and its embodiments comprises a C-terminal deletion starting with a residue selected from the group consisting of 149, 150, 151, 152, 153, 154, 155, 156, and 157. In one embodiment, the adenosine deaminase of this aspect and its embodiments further comprises a modification selected from the group consisting of Y147T, Y147R, Q154S, Y123H, V82S, T166R, and Q154R. In one embodiment, the adenosine deaminase of this aspect and its embodiments is an adenosine deaminase variant described in Table 14, Table 18, or FIGS. 3A - 3C.

[0005] In another aspect, a fusion protein is provided, the fusion protein comprising a polynucleotide-programmable DNA binding domain and at least one base editor domain comprising a modification at an amino acid position selected from the group consisting of 21, 23, 25, 38, 51, 54, 70, 71, 72, 73, 94, 124, 133, 139, 146, and 158 of SEQ ID NO: 1, or a corresponding modification in another adenosine deaminase: JPEG0007717684000002.jpg41165

[0006] In one embodiment, the adenosine deaminase variant comprises a modification selected from the group consisting of R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, M94V, P124W, T133K, D139L, D139M, C146R, and A158K of SEQ ID NO: 1, or a corresponding modification in another adenosine deaminase.

[0007] In another aspect, a fusion protein is provided, the fusion protein comprising a polynucleotide programmable DNA binding domain and at least one base editor domain that is an adenosine deaminase variant comprising a modification selected from the group consisting of R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, M94V, P124W, T133K, D139L, D139M, C146R, and A158K of SEQ ID NO: 1, or a corresponding modification in another adenosine deaminase.

[0008] In any embodiment of any of the fusion proteins of the above aspects and their embodiments, the adenosine deaminase variant further comprises the V82T modification of SEQ ID NO: 1, or a corresponding modification in another adenosine deaminase.

[0009] In another aspect, a fusion protein is provided, the fusion protein comprising a polynucleotide programmable DNA binding domain and at least one base editor domain that is an adenosine deaminase variant comprising the V82T modification and one or more modifications selected from the group consisting of R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, M94V, P124W, T133K, D139L, D139M, C146R, and A158K of SEQ ID NO: 1, or a corresponding modification in another adenosine deaminase.

[0010] In an embodiment of the fusion protein according to any of the above-described embodiments and their embodiments, the adenosine deaminase variant comprises modifications at two or more amino acid positions selected from the group consisting of 21, 23, 25, 38, 51, 54, 70, 71, 72, 73, 94, 124, 133, 139, 146, and 158 of SEQ ID NO: 1, or corresponding modifications in another adenosine deaminase. In one embodiment, the adenosine deaminase variant comprises two or more modifications. In one embodiment, the adenosine deaminase variant comprises three or more modifications. In one embodiment, the adenosine deaminase variant further comprises one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, and Q154R. In one embodiment, the adenosine deaminase variant comprises a C-terminal deletion starting with a residue selected from the group consisting of 149, 150, 151, 152, 153, 154, 155, 156, and 157.

[0011] In an embodiment of the above-described fusion protein and its embodiments, the base editor domain comprises an adenosine deaminase variant monomer, and the adenosine deaminase monomer comprises one or more modifications selected from the group consisting of R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, V82T, M94V, P124W, T133K, D139L, D139M, C146R, and A158K of SEQ ID NO: 1. In one embodiment, the base editor domain comprises an adenosine deaminase heterodimer comprising a wild-type adenosine deaminase domain and an adenosine deaminase variant. In one embodiment, the adenosine deaminase variant further comprises a modification selected from the group consisting of Y147T, Y147R, Q154S, Y123H, V82S, T166R, and Q154R. In one embodiment, the base editor domain comprises an adenosine deaminase heterodimer comprising a TadA*7.10 domain and an adenosine deaminase variant domain. In one embodiment, the adenosine deaminase variant comprises two or more modifications.

[0012] In another embodiment of the fusion protein of any of the above aspects and their embodiments, the adenosine deaminase variant is ABE9 (TadA*9 deaminase variant) described in Table 14, Table 18, or Figures 3A - 3C.

[0013] In another embodiment of the fusion protein of any of the above aspects and their embodiments, the adenosine deaminase variant is a truncated ABE8 or ABE9 with 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 C - terminal amino acid residues deleted compared to full - length ABE9.

[0014] In another embodiment of the fusion protein of any of the above aspects and their embodiments, the polynucleotide - programmable DNA - binding domain is a Cas9, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ domain.

[0015] In another aspect, a fusion protein is provided, the fusion protein comprising a polynucleotide - programmable DNA - binding domain comprising the following sequence: JPEG0007717684000003.jpg166160In this array, the bolded array represents the sequence derived from Cas9, the italicized array represents the linker sequence, the underlined array represents the bipartite (two-part) nuclear localization sequence, and at least one base editor domain includes an adenosine deaminase variant with modifications at amino acid positions selected from the group consisting of 21, 23, 25, 38, 51, 54, 70, 71, 72, 73, 94, 124, 133, 138, 139, 146, and 158 of SEQ ID NO: 1. In one embodiment, the adenosine deaminase variant includes modifications selected from the group consisting of R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, M94V, P124W, T133K, D138M, D139L, D139M, C146R, and A158K of SEQ ID NO: 1. In another embodiment, the adenosine deaminase variant includes the modification V82T of SEQ ID NO: 1. In one embodiment, the adenosine deaminase variant includes two or more of the above modifications. In one embodiment, the adenosine deaminase variant includes three or more of the above modifications. In one embodiment, the adenosine deaminase variant further includes modifications selected from the group consisting of Y147T, Y147R, Q154S, Y123H, V82S, T166R, and Q154R. In one embodiment, the adenosine deaminase variant includes two or more of the following modifications: Y147T, Y147R, Q154S, Y123H, and Q154R.

[0016] In any embodiment of the above fusion protein and its embodiments, the adenosine deaminase variant includes any one of the following groups of modifications: E25F + V82S + Y123H; T133K + Y147R + Q154R; E25F + V82S + Y123H + Y147R + Q154R; L51W + V82S + Y123H + C146R + Y147R + Q154R; Y73S + V82S + Y123H + Y147R + Q154R; P54C + V82S + Y123H + Y147R + Q154R; N38G + V82T + Y123H + Y147R + Q154R; N72K + V82S + Y123H + D139L + Y147R + Q154R; E25F + V82S + Y123H + D139M + Y147R + Q154R; Q71M + V82S + Y123H + Y147R + Q154R; E25F + V82S + Y123H + T133K + Y147R + Q154R; E25F + V82S + Y123H + Y147R + Q154R; V82S + Y123H + P124W + Y147R + Q154R; L51W + V82S + Y123H + C146R + Y147R + Q154R; P54C + V82S + Y123H + Y147R + Q154R; Y73S + V82S + Y123H + Y147R + Q154R; N38G + V82T + Y123H + Y147R + Q154R; R23H + V82S + Y123H + Y147R + Q154R; R21N + V82S + Y123H + Y147R + Q154R; V82S + Y123H + Y147R + Q154R + A158K; N72K + V82S + Y123H + D139L + Y147R + Q154R; E25F + V82S + Y123H + D139M + Y147R + Q154R; M70V + V82S + M94V + Y123H + Y147R + Q154R; Q71M + V82S + Y123H + Y147R + Q154R; E25F + I76Y + V82S + Y123H + Y147R + Q154R; I76Y + V82T + Y123H + Y147R + Q154R; N38G + I76Y + V82S + Y123H + Y147R + Q154R; R23H + I76Y + V82S + Y123H + Y147R + Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; E25F+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82T+Y123H+Y147R+Q154R; N38G+I76Y+V82S+Y123H+Y147R+Q154R; R23H+I76Y+V82S+Y123H+Y147R+Q154R; P54C+I76Y+V82S+Y123H+Y147R+Q154R; R21N+I76Y+V82S+Y123H+Y147R+Q154R; I76Y+V82S+Y123H+D139M+Y147R+Q154R; Y73S+I76Y+V82S+Y123H+Y147R+Q154R; V82S+Q154R; N72K+V82S+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R; N72K+V82S+Y123H+Y147R+Q154R; Q71M+V82S+Y123H+Y147R+Q154R; M70V+V82S+M94V+Y123H+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R; V82S+Y123H+T133K+Y147R+Q154R+A158K; M70V + Q71M + N72K + V82S + Y123H + Y147R + Q154R In one embodiment, the adenosine deaminase variant comprises any other modification or group of modifications described in Table 14 or 18, or FIGS. 3A - 3C.

[0017] In embodiments of the fusion proteins of any of the above aspects and their embodiments, the polynucleotide programmable DNA binding domain is Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), Streptococcus pyogenes Cas9 (SpCas9) or a variant thereof.

[0018] In embodiments of the fusion proteins of any of the above aspects and their embodiments, the polynucleotide programmable DNA binding domain comprises a modified SaCas9 having a modified protospacer adjacent motif (PAM) specificity. In one embodiment, the modified SaCas9 comprises the amino acid substitutions E782K, N968K, and R1015H, or corresponding amino acid substitutions thereto.

[0019] In an embodiment of the fusion protein of any of the above aspects and their embodiments, the polynucleotide-programmable DNA binding domain comprises a variant of SpCas9 having a modified protospacer adjacent motif (PAM) specificity. In one embodiment, the modified PAM has specificity for the nucleic acid sequences 5'-NGA-3', 5'-NGC-3', 5'-NGG-3', 5'-NGT-3', or 5''-NGN-3'. In one embodiment, the variant SpCas9 comprises an amino acid substitution selected from: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R, or an amino acid substitution corresponding thereto; I322V, S409I, E427G, R654L, R753G (MQKFRAER) or an amino acid substitution corresponding thereto; I322V, S409I, E427G, R654L, R753G, R1114G, or an amino acid substitution corresponding thereto; or the amino acid substitutions described in FIGS. 3A-3C.

[0020] In an embodiment of the fusion protein of any of the above aspects and their embodiments, the polynucleotide-programmable DNA binding domain is a nuclease-inactive or nickase variant. In one embodiment, the nickase variant comprises the amino acid substitution D10A or an amino acid substitution corresponding thereto.

[0021] In an embodiment of the fusion protein of any of the above aspects and their embodiments, the adenosine deaminase domain is capable of deaminating adenosine in deoxyribonucleic acid (DNA).

[0022] In an embodiment of the fusion protein of any of the above aspects and their embodiments, the adenosine deaminase is a modified adenosine deaminase not found in nature.

[0023] In the embodiments of adenosine deaminase of the above-described aspects and their embodiments, the adenosine deaminase is TadA deaminase. In the embodiments of any fusion protein of the above-described aspects and their embodiments, the adenosine deaminase is TadA deaminase. In one embodiment, the TadA deaminase is the TadA*7.10 variant.

[0024] In the embodiments of any fusion protein of the above-described aspects and their embodiments, the fusion protein includes a linker between the polynucleotide programmable DNA-binding domain and the adenosine deaminase domain. In one embodiment, the linker includes the amino acid sequence: SGGSSGGSSGSETPGTSESATPES.

[0025] In the embodiments of any fusion protein of the above-described aspects and their embodiments, the fusion protein includes one or more nuclear localization signals. In one embodiment, the nuclear localization signal is a bipartite nuclear localization signal.

[0026] In the embodiments of any fusion protein of the above-described aspects and their embodiments, Cas9 is StCas9.

[0027] In the embodiments of any fusion protein of the above-described aspects and their embodiments, Cas9 is SaCas9 or SpCas9.

[0028] In the embodiments of any fusion protein of the above-described aspects and their embodiments, Cas9 is a modified SaCas9 or a modified SpCas9. In one embodiment, the modified SaCas9 includes the amino acid substitutions E782K, N968K, and R1015H, or amino acid substitutions corresponding thereto. In one embodiment, the modified SaCas9 includes the following amino acid sequence:

[0029] In another aspect, provided is a polynucleotide encoding any one of the above-described aspects and fusion proteins of any one of its embodiments.

[0030] In another aspect, provided is a cell produced by introducing into the cell, or its progenitor cell, a polynucleotide encoding any one of the above-described aspects and fusion proteins of any one of its embodiments, and one or more guide polynucleotides that target a base editor to effect a change from A·T to G·C of an SNP associated with a genetic disease. In one embodiment, the cell is a human cell. In one embodiment, the cell is in vitro or in vivo. In one embodiment, the genetic disease is α-1 antitrypsin deficiency (A1AD). In one embodiment, the fusion protein and the one or more guide polynucleotides form a complex intracellularly.

[0031] In another aspect, provided is an isolated cell or a population of cells grown or expanded from the cells of the above-described aspects and their embodiments.

[0032] In one aspect, provided is a method for treating a genetic disease in a subject in need thereof, the method comprising administering to the subject any one of the cells, isolated cells, or cell populations of the above-described aspects and their embodiments. In one embodiment of the method, the cells, isolated cells, or cell populations are autologous, allogeneic, or xenogeneic to the subject.

[0033] In one aspect, provided is a base editor system, the base editor system comprising a polynucleotide-programmable DNA binding domain and at least one base editor domain that is an adenosine deaminase variant comprising a modification at an amino acid position selected from the group consisting of 21, 23, 25, 38, 51, 54, 70, 71, 72, 73, 82, 94, 124, 133, 139, 146, and 158 of SEQ ID NO: 1 below, and corresponding modifications in another adenosine deaminase: In an embodiment of the JPEG0007717684000004.jpg 49160 base editor system, the adenosine deaminase variant comprises a modification selected from the group consisting of R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, V82T, M94V, P124W, T133K, D139L, D139M, C146R, and A158K of SEQ ID NO: 1, or a corresponding modification in another adenosine deaminase. In one embodiment, the base editor system further comprises one or more guide polynucleotides that target the base editor domain to effect a change from A·T to G·C of an SNP associated with a genetic disease. In one embodiment of the base editor system, the adenosine deaminase variant is capable of deaminating adenine in deoxyribonucleic acid (DNA). In one embodiment of the base editor system, the guide polynucleotide comprises ribonucleic acid (RNA), or deoxyribonucleic acid (DNA). In one embodiment of the base editor system, the guide polynucleotide comprises a CRISPR RNA (crRNA) sequence, a trans-activating CRISPR RNA (tracrRNA) sequence, or a combination thereof. In one embodiment, the base editor system further comprises a second guide polynucleotide. In one embodiment, the second guide polynucleotide comprises ribonucleic acid (RNA), or deoxyribonucleic acid (DNA). In one embodiment, the second guide polynucleotide comprises a CRISPR RNA (crRNA) sequence, a trans-activating CRISPR RNA (tracrRNA) sequence, or a combination thereof. In embodiments of the above base editor system and its embodiments, the polynucleotide-programmable DNA-binding domain comprises a Cas9, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ domain. In one embodiment, the polynucleotide-programmable DNA-binding domain is nuclease-inactive. In one embodiment, the polynucleotide-programmable DNA-binding domain is a nickase.In one embodiment, the polynucleotide-programmable DNA binding domain comprises a Cas9 domain. In one embodiment, the Cas9 domain comprises nuclease-inactive (dead) Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9. In one embodiment, the Cas9 domain comprises Cas9 nickase. In one embodiment, the polynucleotide-programmable DNA binding domain is a modified or engineered polynucleotide-programmable DNA binding domain. In embodiments of the base editor system and its embodiments described above, the genetic disease is α-1 antitrypsin deficiency (A1AD).

[0034] In another aspect, provided is a method of modifying a single nucleotide polymorphism (SNP) in a polynucleotide, the method comprising contacting a target nucleotide sequence, at least a portion of which is located in the polynucleotide or its reverse complement, with any one of the fusion proteins of any one of the above aspects and its embodiments, or any one of the base editor systems of any one of the above aspects and its embodiments; editing the SNP by deaminating the SNP or its complementary nucleobase when targeting the base editor to the target nucleotide sequence, wherein the SNP is modified by deaminating the SNP or its complementary nucleobase. In one embodiment, the SNP is associated with α-1 antitrypsin deficiency (A1AD). In one embodiment, the SNP is in the SERPINA1 gene and the modification comprises the alteration of E342K (PiZ allele).

[0035] In one aspect, a method for editing a polynucleotide is provided, the method comprising contacting a target nucleotide sequence with a fusion protein of any one of the above aspects and embodiments, or a base editor system of any one of the above aspects and embodiments, thereby editing the polynucleotide. In one embodiment of the method, the editing results in less than 20% indel formation, less than 15% indel formation, less than 10% indel formation; less than 5% indel formation; less than 4% indel formation; less than 3% indel formation; less than 2% indel formation; less than 1% indel formation; less than 0.5% indel formation; or less than 0.1% indel formation. In one embodiment of the method, the editing does not result in a translocation.

[0036] In another aspect, ABE9 (TadA*9 deaminase variant) comprising a TadA*7.10 adenosine deaminase variant domain selected from the following, and a Cas9 endonuclease domain: monoTadA*7.10 having the mutations I76Y+V82T+Y147T+Q154S+A109S of SEQ ID NO: 1, and spCas9 (MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G; monoTadA*7.10 having the mutations I76Y+V82T+Y147T+Q154S+T111R of SEQ ID NO: 1, and spCas9 (MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G; monoTadA*7.10 having the mutations I76Y+V82T+Y147T+Q154S+D119N of SEQ ID NO: 1, and spCas9 (MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G; monoTadA*7.10 having the mutations I76Y+V82T+Y147T+Q154S+H122N of SEQ ID NO: 1, and spCas9 (MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G; monoTadA*7.10 having the mutation I76Y+V82T+Y147d+Q154S of SEQ ID NO: 1, and spCas9(MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G; monoTadA*7.10 having the mutation I76Y+V82T+Y147T+Q154S+F149Y of SEQ ID NO: 1, and spCas9(MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G; monoTadA*7.10 having the mutation I76Y+V82T+Y147T+Q154S+T166I of SEQ ID NO: 1, and spCas9(MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G; and monoTadA*7.10 having the mutation I76Y+V82T+Y147T+Q154S+D167N of SEQ ID NO: 1, and spCas9(MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G. monoTadA*7.10 having the mutation I76Y+V82T+Y147T+Q154S+L36H+N157K of SEQ ID NO: 1, and spCas9(MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G, R1114G; monoTadA*7.10 having the mutation I76Y+V82T+Y147D+Q154S+F149Y+D167N+L36H+N157K of SEQ ID NO: 1, and SpCas9(MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G, R1114G; monoTadA*7.10 having the mutation I76Y+V82T+Y147D+Q154S+F149Y+D167N+L36H+N157K+V106W of SEQ ID NO: 1, and SpCas9(MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G, R1114G; monoTadA*7.10 having the mutation A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N of SEQ ID NO: 1, and SpCas9, MQKFRAER having the mutations I322V, S409I, E427G, R654L, R753G, R1114G; and monoTadA*7.10 having the mutation A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N+V106W of SEQ ID NO: 1, and SpCas9 (MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G, R1114G; and one or more guide polynucleotides that target the adenosine deaminase variant domain to effect a change from A·T to G·C of an SNP associated with a genetic disease A base editor is provided that includes the above. In one embodiment of the base editor, the SNP is associated with alpha-1 antitrypsin deficiency (A1AD).

[0037] In another aspect, a vector is provided that includes one or more polynucleotides encoding an ABE9 base editor that includes a TadA adenosine deaminase domain and a SpCas9 endonuclease domain selected from the following. monoTadA*7.10 having the mutations I76Y+V82T+Y147T+Q154S+A109S and spCas9 (MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G; monoTadA*7.10 having the mutations I76Y+V82T+Y147T+Q154S+T111R and spCas9 (MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G; monoTadA*7.10 having the mutations I76Y+V82T+Y147T+Q154S+D119N and spCas9 (MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G; monoTadA*7.10 having mutation I76Y+V82T+Y147T+Q154S+H122N and spCas9(MQKFRAER) having mutations I322V, S409I, E427G, R654L, R753G; monoTadA*7.10 having mutation I76Y+V82T+Y147d+Q154S and spCas9(MQKFRAER) having mutations I322V, S409I, E427G, R654L, R753G; monoTadA*7.10 having mutation I76Y+V82T+Y147T+Q154S+F149Y and spCas9(MQKFRAER) having mutations I322V, S409I, E427G, R654L, R753G; monoTadA*7.10 having mutation I76Y+V82T+Y147T+Q154S+T166I and spCas9(MQKFRAER) having mutations I322V, S409I, E427G, R654L, R753G; and monoTadA*7.10 having mutation I76Y+V82T+Y147T+Q154S+D167N and spCas9(MQKFRAER) having mutations I322V, S409I, E427G, R654L, R753G. monoTadA*7.10 having mutation I76Y+V82T+Y147T+Q154S+L36H+N157K, spCas9(MQKFRAER) having mutations I322V, S409I, E427G, R654L, R753G, R1114G; monoTadA*7.10 having mutation I76Y+V82T+Y147D+Q154S+F149Y+D167N+L36H+N157K and SpCas9(MQKFRAER) having mutations I322V, S409I, E427G, R654L, R753G, R1114G; monoTadA*7.10 having mutation I76Y+V82T+Y147D+Q154S+F149Y+D167N+L36H+N157K+V106W and SpCas9(MQKFRAER) having mutations I322V, S409I, E427G, R654L, R753G, R1114G; monoTadA*7.10 having the mutations A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N and SpCas9(MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G, R1114G; and monoTadA*7.10 having the mutations A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N+V106W, SpCas9(MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G, R1114G. In one embodiment, the vector is a plasmid, virus, or mRNA vector.

[0038] In another aspect, provided is a composition comprising the fusion protein of any one of the above aspects and its embodiments, or the base editor system of any one of the above aspects and its embodiments. In one embodiment, the composition further comprises a pharmaceutically acceptable excipient, diluent, or carrier.

[0039] In another aspect, provided is a composition comprising the fusion protein of any one of the above aspects and its embodiments bound to a guide RNA, wherein the guide RNA comprises a nucleic acid sequence complementary to the SERPINA1 gene associated with alpha-1 antitrypsin deficiency (A1AD).

[0040] In another aspect, provided is a composition comprising the base editor system of any one of the above aspects and its embodiments bound to a guide RNA, wherein the guide RNA comprises a nucleic acid sequence complementary to the SERPINA1 gene associated with alpha-1 antitrypsin deficiency (A1AD).

[0041] In an embodiment of the composition of any one of the above aspects and its embodiments, the adenosine deaminase variant is capable of deaminating adenine in deoxyribonucleic acid (DNA).

[0042] In an embodiment of any one of the above-described embodiments and compositions thereof, the fusion protein or base editor system is (i) comprising a Cas9 nickase; (ii) comprising a nuclease-inactive Cas9; (iii) comprising an SpCas9 variant comprising a combination of amino acid substitutions shown in FIGS. 3A-3C; or (iv) an SpCas9 variant comprising a combination of amino acid sequence substitutions selected from I322V, S409I, E427G, R654L, R753G (MQKFRAER); or I322V, S409I, E427G, R654L, R753G, R1114G (MQKFRAER).

[0043] In an embodiment of any one of the above-described embodiments and compositions thereof, the composition further comprises a pharmaceutically acceptable excipient, diluent, or carrier, i.e., it is a pharmaceutical composition.

[0044] In one aspect, there is provided a pharmaceutical composition for the treatment of a disease or disorder, comprising a composition further comprising a pharmaceutically acceptable excipient, diluent, or carrier. In one embodiment of the pharmaceutical composition, the disease or disorder is α-1 antitrypsin deficiency (A1AD). In one embodiment of the pharmaceutical composition, the fusion protein or base editor system is bound to a guide RNA, and the guide RNA comprises a nucleic acid sequence complementary to the SERPINA1 gene associated with α-1 antitrypsin deficiency (A1AD). In one embodiment of the pharmaceutical composition, the gRNA and the base editor are formulated together or separately. In an embodiment of the above pharmaceutical composition and its embodiments, the gRNA is, from 5' to 3', 5’-ACCAUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU-3’; 5’-CCAUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU-3’; 5’-CAUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU-3’; 5’-AUCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU-3’; 5’-UCGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU-3’; or It comprises a nucleic acid sequence selected from one or more of 5’-CGACAAGAAAGGGACUGA GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU-3’, or a 5’-truncated fragment thereof of 1, 2, 3, 4, or 5 nucleotides. In the above pharmaceutical composition and its embodiments, the pharmaceutical composition further comprises a vector suitable for expression in mammalian cells, and the vector comprises a polynucleotide encoding a base editor. In one embodiment of the pharmaceutical composition, the polynucleotide encoding the base editor is mRNA. In one embodiment of the pharmaceutical composition, the vector is a viral vector. In one embodiment of the pharmaceutical composition, the viral vector is a retroviral vector, an adenoviral vector, a lentiviral vector, a herpes viral vector, or an adeno-associated viral vector (AAV). In the above aspect and any one of its embodiments of the pharmaceutical composition, the pharmaceutical composition further comprises ribonucleoprotein particles suitable for expression in mammalian cells. In any one of the above aspects and its embodiments of the pharmaceutical composition, the pharmaceutical composition further comprises a lipid.

[0045] In another aspect, a method for treating α-1 antitrypsin deficiency (A1AD) is provided, the method comprising administering to a subject in need thereof a pharmaceutical composition according to any one of the above aspects and its embodiments.

[0046] In another aspect, the use of a pharmaceutical composition according to any one of the above aspects and its embodiments in the treatment of α-1 antitrypsin deficiency (A1AD) in a subject is provided.

[0047] In an embodiment of the above method or use, the subject is human.

[0048] In any one of the above aspects and its embodiments of the fusion protein or base editor system, the adenosine deaminase variant comprises any one of the following groups of modifications: E25F + V82S + Y123H; T133K + Y147R + Q154R; E25F + V82S + Y123H + Y147R + Q154R; L51W + V82S + Y123H + C146R + Y147R + Q154R; Y73S + V82S + Y123H + Y147R + Q154R; P54C + V82S + Y123H + Y147R + Q154R; N38G + V82T + Y123H + Y147R + Q154R; N72K + V82S + Y123H + D139L + Y147R + Q154R; E25F + V82S + Y123H + D139M + Y147R + Q154R; Q71M + V82S + Y123H + Y147R + Q154R; E25F + V82S + Y123H + T133K + Y147R + Q154R; E25F + V82S + Y123H + Y147R + Q154R; V82S + Y123H + P124W + Y147R + Q154R; L51W + V82S + Y123H + C146R + Y147R + Q154R; P54C + V82S + Y123H + Y147R + Q154R; Y73S + V82S + Y123H + Y147R + Q154R; N38G + V82T + Y123H + Y147R + Q154R; R23H + V82S + Y123H + Y147R + Q154R; R21N + V82S + Y123H + Y147R + Q154R; V82S + Y123H + Y147R + Q154R + A158K; N72K + V82S + Y123H + D139L + Y147R + Q154R; E25F + V82S + Y123H + D139M + Y147R + Q154R; M70V + V82S + M94V + Y123H + Y147R + Q154R; Q71M + V82S + Y123H + Y147R + Q154R; E25F + I76Y + V82S + Y123H + Y147R + Q154R; I76Y + V82T + Y123H + Y147R + Q154R; N38G + I76Y + V82S + Y123H + Y147R + Q154R; R23H + I76Y + V82S + Y123H + Y147R + Q154R; P54C + I76Y + V82S + Y123H + Y147R + Q154R; R21N + I76Y + V82S + Y123H + Y147R + Q154R; I76Y + V82S + Y123H + D139M + Y147R + Q154R; Y73S + I76Y + V82S + Y123H + Y147R + Q154R; E25F + I76Y + V82S + Y123H + Y147R + Q154R; I76Y + V82T + Y123H + Y147R + Q154R; N38G + I76Y + V82S + Y123H + Y147R + Q154R; R23H + I76Y + V82S + Y123H + Y147R + Q154R; P54C + I76Y + V82S + Y123H + Y147R + Q154R; R21N + I76Y + V82S + Y123H + Y147R + Q154R; I76Y + V82S + Y123H + D139M + Y147R + Q154R; Y73S + I76Y + V82S + Y123H + Y147R + Q154R; V82S + Q154R; N72K + V82S + Y123H + Y147R + Q154R; Q71M + V82S + Y123H + Y147R + Q154R; V82S + Y123H + T133K + Y147R + Q154R; V82S + Y123H + T133K + Y147R + Q154R + A158K; M70V + Q71M + N72K + V82S + Y123H + Y147R + Q154R; N72K + V82S + Y123H + Y147R + Q154R; Q71M + V82S + Y123H + Y147R + Q154R; M70V + V82S + M94V + Y123H + Y147R + Q154R; V82S + Y123H + T133K + Y147R + Q154R; V82S + Y123H + T133K + Y147R + Q154R + A158K; M70V + Q71M + N72K + V82S + Y123H + Y147R + Q154R。 In one embodiment, an adenosine deaminase variant, such as a TadA*9 deaminase variant, comprises any modification or group of modifications described in Table 14 or 18.

[0049] As would be understood by one of ordinary skill in the art in connection with the above-described embodiments and adenosine deaminases thereof, amino acid modifications in other adenosine deaminases corresponding to the amino acid modifications set forth in SEQ ID NO: 1 can be readily determined by performing a normal sequence alignment and assessing the relatedness and / or identity of the amino acid sequence of SEQ ID NO: 1 with the sequence of other adenosine deaminase(s), such as TadA deaminase, or a related protein portion thereof, as described above. In one embodiment, the amino acid sequence of another adenosine deaminase has at least 85% sequence identity to SEQ ID NO: 1. In one embodiment, the amino acid sequence of another adenosine deaminase has at least 90% sequence identity to SEQ ID NO: 1. In one embodiment, the amino acid sequence of another adenosine deaminase has at least 95% sequence identity to SEQ ID NO: 1. In one embodiment, the amino acid sequence of another adenosine deaminase has at least 98% sequence identity to SEQ ID NO: 1. In one embodiment, the amino acid sequence of another adenosine deaminase has at least 99% sequence identity to SEQ ID NO: 1.

[0050] In another aspect, provided are the above adenosine deaminase, fusion protein, base editor, or base editor system and embodiments thereof, comprising adenosine deaminase, or an adenosine deaminase variant that is a TadA*7.10 variant comprising any one of the following amino acid modifications or groups of modifications: V82T; I76Y+V82T; or I76Y+V82T+Y147T+Q154S.

[0051] In another aspect, provided is an adenosine deaminase variant that is a TadA*7.10 variant comprising any one of the following amino acid modifications or groups of modifications: V82T; I76Y+V82T; or I76Y+V82T+Y147T+Q154S.

[0052] In another aspect, provided is a fusion protein, the fusion protein comprising a polynucleotide-programmable DNA binding domain and at least one base editor domain that is a TadA*7.10 adenosine deaminase variant comprising any one of the following amino acid modifications or groups of modifications: V82T; I76Y+V82T; or I76Y+V82T+Y147T+Q154S. In one embodiment of the fusion protein, the polynucleotide-programmable DNA binding domain comprises a Cas9 endonuclease domain. In one embodiment of the fusion protein, the Cas9 endonuclease domain comprises spCas9(MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G.

[0053] In embodiments of the above adenosine deaminase variant and its embodiments, or the above fusion protein and its embodiments, TadA7*10 is monomeric.

[0054] In another aspect, provided is a nucleic acid base editor, the nucleic acid base editor comprising a TadA*7.10 adenosine deaminase variant domain and a Cas9 endonuclease domain selected from: monoTadA*7.10 having the mutation V82T and spCas9(MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G; monoTadA*7.10 having the mutations I76Y+V82T and spCas9(MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G; or monoTadA*7.10 having the mutations I76Y+V82T+Y147T+Q154S and spCas9(MQKFRAER) having the mutations I322V, S409I, E427G, R654L, R753G.

[0055] Definitions The following definitions supplement those in the art and are for the purposes of this application and do not belong to related or unrelated cases such as patents or applications related to common ownership. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of this disclosure, the preferred materials and methods are described herein. Accordingly, the terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting.

[0056] Unless otherwise defined, all technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. The following references provide one of the technologies including general definitions of many terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them below, unless otherwise specified.

[0057] In this application, the use of the singular form includes the plural unless specifically stated otherwise. It should be noted that, as used herein, the singular forms "a", "an", and "the" include the plural referents unless the context clearly dictates otherwise. In this application, the use of "or" means "and / or" unless otherwise stated. Further, the use of the terms "including", as well as other forms such as "include", "includes", and "included", is not limiting.

[0058] As used in this specification and the claims (if any), the word "comprising" (and any form of "comprising", such as "comprise" and "comprises"), "having" (and any form of "having", such as "have" and "has"), "including" (and any form of "including", such as "includes" and "include") or "containing" (and any form of "containing", such as "contains" and "contain") is inclusive or open-ended and does not exclude additional, unrecited elements or method steps. Any embodiment recited herein is contemplated to be able to be practiced with respect to any method or composition of the present disclosure and vice versa. Further, the methods of the present disclosure can be achieved using the compositions of the present disclosure.

[0059] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within one or more standard deviations, per the convention in the art. Alternatively, "about" can mean within a range of up to 20%, 10%, 5%, or 1% of a given value. Alternatively, especially with respect to biological systems or processes, the term can mean within the same order of magnitude, such as within 5-fold, or 2-fold, of a value. Unless otherwise specified, when a particular value is recited in this application and the claims, the term "about" should be assumed to mean within an acceptable error range for that particular value.

[0060] References herein to "some embodiments", "an embodiment", "one embodiment" or "another embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least some embodiments, but not necessarily all embodiments of the present disclosure.

[0061] The term "adenosine deaminase" refers to a polypeptide or a fragment thereof that can catalyze the hydrolytic deamination of adenine or adenosine. In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine to inosine or deoxyadenosine to deoxyinosine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., modified adenosine deaminases, evolved adenosine deaminases) can be derived from any organism such as bacteria.

[0062] In some embodiments, the deaminase or deaminase domain is a variant of a native deaminase derived from an organism such as a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain does not exist in nature. For example, in some embodiments, the deaminase or deaminase domain has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to a native deaminase. In some embodiments, the adenosine deaminase is derived from a bacterium such as E. coli, S. aureus, S. typhi, S. putrefaciens, H. influenzae, or C. crescentus. In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is E. coli TadA (ecTadA) deaminase or a fragment thereof.

[0063] For example, deaminase domains are described in International PCT Application Nos. PCT / 2017 / 045381 (WO2018 / 027078) and PCT / US2016 / 058344 (WO2017 / 070632), each of which is hereby incorporated by reference in its entirety. See also Komor, A. C., et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, N. M., et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, A. C., et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao4774 (2017)), and Rees, H. A., et al., “Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec; 19 (12): 770-788. doi: 10.1038 / s41576-018-0059-1 (the entire contents of which are hereby incorporated by reference).

[0064] Wild-type TadA (wt) adenosine deaminase has the following sequence (also referred to as the TadA reference sequence): MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD。

[0065] In some embodiments, adenosine deaminase comprises a modification of the following sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (Also known as TadA*7.10).

[0066] The present invention features a novel nucleobase editor with modifications to the TadA*7.10 reference sequence.

[0067] In some embodiments, TadA*7.10 includes at least one modification. In some embodiments, TadA*7.10 includes modifications at amino acids 82 and / or 166. In certain embodiments, variants of the above reference sequences include one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. The modification Y123H refers to the modification H123Y of TadA*7.10 being reverted to Y123H TadA (wt). In other embodiments, variants of the TadA*7.10 sequence include one or more of the following modifications of SEQ ID NO: 1: R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, M94V, P124W, T133K, D139L, D139M, C146R, and A158K. In some embodiments, variants of the TadA*7.10 sequence include a combination of modifications selected from the group consisting of: Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R.

[0068] In other embodiments, the present invention provides deletion-containing adenosine deaminase variants, such as TadA*8, that include a C-terminal deletion starting at residue 149, 150, 151, 152, 153, 154, 155, 156, or 157 as compared to the corresponding mutation in TadA*7.10, the TadA reference sequence, or another TadA.

[0069] In yet other embodiments, the adenosine deaminase variant is a homodimer comprising two adenosine deaminase domains, each having one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R, as compared to TadA*7.10, the TadA reference sequence, or the corresponding variant of another TadA. In other embodiments, the adenosine deaminase variant is, as compared to TadA*7.10, the TadA reference sequence, or the corresponding variant of another TadA, each having a combination of modifications selected from the group consisting of: Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R, and is a homodimer comprising two adenosine deaminase domains (e.g., TadA*8).

[0070] In other embodiments, the adenosine deaminase variant is a heterodimer comprising an adenosine deaminase variant domain (e.g., TadA*8) that includes one or more of the following modifications compared to the wild-type TadA adenosine deaminase domain and the corresponding mutations of TadA*7.10, the TadA reference sequence, or another TadA, namely, Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. In other embodiments, the adenosine deaminase variant, compared to the wild-type TadA adenosine deaminase domain and the corresponding mutations of TadA*7.10, the TadA reference sequence, or another TadA, includes a combination of modifications selected from the group consisting of: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R + Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R, and is a heterodimer comprising an adenosine deaminase variant domain (e.g., TadA*8).

[0071] In other embodiments, the adenosine deaminase variant is a heterodimer comprising an adenosine deaminase variant domain (e.g., TadA*8) that includes one or more of the following modifications as compared to the TadA*7.10 domain and the corresponding mutations of TadA*7.10, the TadA reference sequence, or another TadA, namely, Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. In other embodiments, the adenosine deaminase variant is a heterodimer comprising an adenosine deaminase variant domain (e.g., TadA*8) that includes a combination of the following modifications as compared to the TadA*7.10 domain and the corresponding mutations of TadA*7.10, the TadA reference sequence, or another TadA: Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; or I76Y+V82S+Y123H+Y147R+Q154R. In one embodiment, the adenosine deaminase is TadA*8 that comprises or consists essentially of the following sequence or a fragment thereof having adenosine deaminase activity: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD。

[0072] In some embodiments, TadA*8 is truncated. In some embodiments, the truncated TadA*8 lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 N-terminal amino acid residues relative to the full-length TadA*8. In some embodiments, the truncated TadA*8 lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 C-terminal amino acid residues relative to the full-length TadA*8. In some embodiments, the adenosine deaminase variant is the full-length TadA*8.

[0073] In certain embodiments, the adenosine deaminase heterodimer comprises a TadA*8 domain and an adenosine deaminase domain selected from one of the following: Staphylococcus aureus (S. aureus) TadA: MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAH AEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGS LMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN Bacillus subtilis (B.subtilis) TadA: MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEML VIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMN LLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLSE Salmonella typhimurium (S. typhimurium) TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFRMRRQEIKALKADRAEGAGPAV Shewanella putrefaciens (S. putrefaciens) TadA: MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE Haemophilus influenzae F3031 (H. influenzae) TadA: MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTΑΗAEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRREEKKIEKALLKSLSDK Caulobacter crescentus (C. crescentus) TadA: JPEG0007717684000005.jpg29168Geobacter sulfurreducens (G. sulfurreducens) TadA: MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSN DPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERVVFGCYDP KGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRKKAKATPALF IDERKVPPEP TadA*7.10 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD。

[0074] The "Adenosine deaminase base editor 8 (ABE8) polynucleotide" means a polynucleotide encoding ABE8.

[0075] The "Adenosine deaminase base editor 9 (ABE9) polypeptide" or "ABE9" means a base editor as defined herein that includes an adenosine deaminase variant (TadA*9) that contains one or more modifications at the positions of sssssss in the following sequences. In one embodiment, the adenosine deaminase variant (TadA*9) includes the following modifications: R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, V82T, M94V, P124W, T133K, D139L, D139M, C146R, and A158K in the following reference sequences: The relevant bases to be modified in the JPEG0007717684000006.jpg45167 reference sequence are shown underlined and in bold. In some embodiments, ABE9 includes additional modifications as described herein, compared to the reference sequence.

[0076] The "Adenosine Deaminase Base Editor 9 (ABE9) polynucleotide" means a polynucleotide encoding ABE9.

[0077] The "α-1 antitrypsin (A1AT) protein" means a polypeptide or a fragment thereof having at least about 95% amino acid sequence identity to UniProt accession number P01009. In certain embodiments, the A1AT protein comprises one or more modifications as compared to the following reference sequence. In one particular embodiment, the A1AT protein associated with A1AD comprises the E342K mutation. An exemplary A1AT amino acid sequence is >sp|P01009|A1AT_HUMAN α-1-antitrypsin OS=Homo sapiens OX=9606 GN=SERPINA1 PE=1 SV=3 and has the following amino acid sequence: JPEG0007717684000007.jpg55168 In this A1AT protein sequence, the first 24 amino acids constitute the signal peptide (underlined). The position 342 of the sequence that is mutated in A1AD (i.e., E342K) is determined based on setting the amino acid residue "E" following the signal sequence as amino acid "1".

[0078] "Administering" is referred to herein as providing one or more of the compositions described herein to a patient or subject. By way of example, and not limitation, administration of the composition, such as an injection, can be carried out by intravenous (i.v.) injection, subcutaneous (s.c.) injection, intradermal (i.d.) injection, intraperitoneal (i.p.) injection, or intramuscular (i.m.) injection. One or more such routes can be used. Parenteral administration can be effected, for example, by bolus injection or by gradual perfusion over time. Alternatively, or concurrently, administration can be effected by the oral route.

[0079] "Agent" means any small molecule chemical compound, antibody, nucleic acid molecule, or polypeptide, or fragment thereof.

[0080] "Alteration" means a change (increase or decrease) in the sequence, expression level or activity of a gene or polypeptide detected by standard methods known in the art, such as the methods described herein. As used herein, alterations include a 10% change in expression level, a 25% change in expression level, a 40% change, and a change of 50% or more.

[0081] "Remission" means, for example, reducing, suppressing, attenuating, decreasing, arresting, or stabilizing the onset or progression of a disease.

[0082] "Analog" means a molecule having similar functional or structural features but not being identical. For example, a polypeptide analog has certain biochemical modifications that enhance the function of the analog relative to the native polypeptide while retaining the biological activity of the corresponding native polypeptide. Such biochemical modifications can increase the protease resistance, membrane permeability, or half-life of the analog, for example, without changing ligand binding. An analog may contain non-natural amino acids.

[0083] "Base editor (BE)" or "nucleic acid base editor (NBE)" means an agent that binds to a polynucleotide and has nucleic acid base modification activity. In various embodiments, a base editor comprises a polynucleotide-programmable nucleotide binding domain combined with a nucleic acid base-modifying polypeptide (e.g., a deaminase) and a guide polynucleotide (e.g., guide RNA). In various embodiments, the agent is a biomolecular complex comprising a protein domain having base editing activity, i.e., a domain capable of modifying a base (e.g., A, T, C, G, or U) within a nucleic acid molecule (e.g., DNA). In some embodiments, the polynucleotide-programmable DNA binding domain is fused or linked to a deaminase domain. In one embodiment, the agent is a fusion protein comprising one or more domains having base editing activity. In another embodiment, the protein domain having base editing activity is linked to a guide RNA (e.g., via an RNA binding motif on the guide RNA and an RNA binding domain fused to a deaminase). In some embodiments, the domain having base editing activity can deaminate a base within a nucleic acid molecule. In some embodiments, the base editor can deaminate one or more bases within a DNA molecule. In some embodiments, the base editor can deaminate cytosine (C) or adenosine (A) in DNA. In some embodiments, the base editor can deaminate both cytosine (C) and adenosine (A) in DNA. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenosine base editor (ABE). In some embodiments, the base editor is both an adenosine base editor (ABE) and a cytidine base editor (CBE). In some embodiments, the base editor is nuclease-inactive Cas9 (dCas9) fused to an adenosine deaminase. In some embodiments, Cas9 is a circular permutant Cas9 (e.g., spCas9 or saCas9).Circularly permuted Cas9 is known in the art and is described, for example, in Oakes et al., Cell 176, 254-267, 2019. In some embodiments, the base editor is fused to an inhibitor of base excision repair, such as the UGI domain, or the dISN domain. In some embodiments, the fusion protein comprises a Cas9 nickase fused to a deaminase and an inhibitor of base excision repair, such as the UGI or dISN domain. In other embodiments, the base editor is a base editor for base dropout.

[0084] In some embodiments, adenosine deaminase is evolved from TadA. In some embodiments, the polynucleotide programmable DNA binding domain is a CRISPR-associated (e.g., Cas or Cpf1) enzyme. In some embodiments, the base editor is catalytically inactive Cas9 (dCas9) fused to a deaminase domain. In some embodiments, the base editor is Cas9 nickase (nCas9) fused to a deaminase domain. In some embodiments, the base editor is fused to an inhibitor of base excision repair (BER). In some embodiments, the inhibitor of base excision repair is uracil DNA glycosylase inhibitor (UGI). In some embodiments, the inhibitor of base excision repair is an inosine base excision repair inhibitor. Details of base editors are described in International PCT applications PCT / 2017 / 045381 (WO2018 / 027078) and PCT / US2016 / 058344 (WO2017 / 070632), each of which is hereby incorporated by reference in its entirety.Also see Komor, A. C., et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533,420-424 (2016); Gaudelli, N. M., et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, A. C., et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3: eaao 4774 (2017), and Rees, H. A., et al., “Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec; 19 (12): 770-788. doi: 10.1038 / s41576-018-0059-1 (the entire contents of which are incorporated herein by reference).

[0085] In some embodiments, base editors are generated by cloning an adenosine deaminase variant (e.g., TadA*8) into a backbone that includes a circularly permuted Cas9 (e.g., spCAS9) and a bipartite nuclear localization sequence (e.g., ABE8 or ABE9). Circularly permuted Cas9 is known in the art and is described, for example, in Oakes et al., Cell 176, 254-267, 2019. Exemplary circularly permuted sequences are described below, where the bold sequences indicate sequences derived from Cas9, the italic sequences indicate linker sequences, and the underlined sequences indicate bipartite nuclear localization sequences. CP5 (Pam variant with MSP “NGC = mutation, normal Cas9 prefers NGG” PID = protein interaction domain and “D10A” nickase): JPEG0007717684000008.jpg184169

[0086] In some embodiments, ABE8 is selected from the base editors of Tables 10, 11, or 13 below. In some embodiments, ABE8 comprises an adenosine deaminase variant evolved from TadA. In some embodiments, the adenosine deaminase variant of ABE8 is a TadA*8 variant as described in Tables 8, 10, 11, or 13 below. In some embodiments, the adenosine deaminase variant is a TadA*7.10 variant (e.g., TadA*8) comprising one or more modifications selected from the group consisting of Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. In various embodiments, ABE8 comprises a TadA*7.10 variant (e.g., TadA*8) comprising a combination of modifications selected from the group of Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R.

[0087] In some embodiments, ABE8 is a monomeric construct comprising one copy of a TadA deaminase, e.g., one TadA*8 variant. In some embodiments, ABE8 is a dimeric or heterodimeric construct comprising multiple, e.g., two copies of the same or different TadA deaminases, e.g., wild-type TadA and TadA*8 variants.

[0088] In some embodiments, ABE9 is selected from the base editors of Table 14 below. In some embodiments, ABE9 comprises an adenosine deaminase variant evolved from TadA. In some embodiments, the adenosine deaminase variant of ABE9 is the TadA*7.10 variant described in Table 14. In some embodiments, the adenosine deaminase variant comprises TadA*7.10 with one or more modifications selected from the group consisting of Y147T, Y147R, Q154S, Y123H, V82S, T166R, Q154R. In various embodiments, ABE9 comprises TadA*7.10 having modifications selected from the following in addition to those described in Table 14: Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y147T+Q154R; Y147T+Q154S; V82S+Q154S; V82T+Q154S and Y123H+Y147R+Q154R+I76Y. In some embodiments, ABE9 is a monomeric construct comprising one copy of the TadA deaminase, e.g., the TadA*9 variant. In some embodiments, ABE9 is a dimeric or heterodimeric construct comprising multiple, e.g., two copies of the same or different TadA deaminases, e.g., wild-type TadA and the TadA*9 variant.

[0089] In some embodiments, the ABE9 base editor comprises the following sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD.

[0090] As an example, the adenine base editor ABE used in the base editing compositions, systems and methods described herein has the following nucleic acid sequence (8877 base pairs) (Addgene, Watertown, MA.; Gaudelli N M, et al., Nature. 2017 Nov 23; 551 (7681): 464-471. doi: 10.1038 / nature 24644; Koblan LW, et al., Nat Biotechnol. 2018 Oct; 36 (9): 843-846. doi: 10.1038 / nbt.4172.). Polynucleotide sequences having at least 95% identity to the ABE nucleic acid sequence are also included. ATATGCCAAGTACGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTATGCCCAGTACAT GACCTTATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATGGTGATGCGG TTTTGGCAGTACATCAATGGGCGTGGATAGCGGTTTGACTCACGGGGATTTCCAAGTCTCCACCCCATTG ACGTCAATGGGAGTTTGTTTTGGCACCAAAATCAACGGGACTTTCCAAAATGTCGTAACAACTCCGCCCC ATTGACGCAAATGGGCGGTAGGCGTGTACGGTGGGAGGTCTATATAAGCAGAGCTGGTTTAGTGAACCGT CAGATCCGCTAGAGATCCGCGGCCGCTAATACGACTCACTATAGGGAGAGCCGCCACCATGAAACGGACA GCCGACGGAAGCGAGTTCGAGTCACCAAAGAAGAAGCGGAAAGTCTCTGAAGTCGAGTTTAGCCACGAGT ATTGGATGAGGCACGCACTGACCCTGGCAAAGCGAGCATGGGATGAAAGAGAAGTCCCCGTGGGCGCCGT GCTGGTGCACAACAATAGAGTGATCGGAGAGGGATGGAACAGGCCAATCGGCCGCCACGACCCTACCGCA CACGCAGAGATCATGGCACTGAGGCAGGGAGGCCTGGTCATGCAGAATTACCGCCTGATCGATGCCACCC TGTATGTGACACTGGAGCCATGCGTGATGTGCGCAGGAGCAATGATCCACAGCAGGATCGGAAGAGTGGT GTTCGGAGCACGGGACGCCAAGACCGGCGCAGCAGGCTCCCTGATGGATGTGCTGCACCACCCCGGCATG AACCACCGGGTGGAGATCACAGAGGGAATCCTGGCAGACGAGTGCGCCGCCCTGCTGAGCGATTTCTTTA GAATGCGGAGACAGGAGATCAAGGCCCAGAAGAAGGCACAGAGCTCCACCGACTCTGGAGGATCTAGCGG AGGATCCTCTGGAAGCGAGACACCAGGCACAAGCGAGTCCGCCACACCAGAGAGCTCCGGCGGCTCCTCC GGAGGATCCTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGCCAAGAGGG CACGCGATGAGAGGGAGGTGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGCGAGGGCTG GAACAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATGGCCCTGAGACAGGGCGGCCTG GTCATGCAGAACTACAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCG GCGCCATGATCCACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACGCAAAAACCGGCGCCGCAGG CTCCCTGATGGACGTGCTGCACTACCCCGGCATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCA GATGAATGTGCCGCCCTGCTGTGCTATTTCTTTCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAAGG CCCAGAGCTCCACCGACTCCGGAGGATCTAGCGGAGGCTCCTCTGGCTCTGAGACACCTGGCACAAGCGA GAGCGCAACACCTGAAAGCAGCGGGGGCAGCAGCGGGGGGTCAGACAAGAAGTACAGCATCGGCCTGGCC ATCGGCACCAACTCTGTGGGCTGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGG TGCTGGGCAACACCGACCGGCACAGCATCAAGAAGAACCTGATCGGAGCCCTGCTGTTCGACAGCGGCGA AACAGCCGAGGCCACCCGGCTGAAGAGAACCGCCAGAAGAAGATACACCAGACGGAAGAACCGGATCTGC TATCTGCAAGAGATCTTCAGCAACGAGATGGCCAAGGTGGACGACAGCTTCTTCCACAGACTGGAAGAGT CCTTCCTGGTGGAAGAGGATAAGAAGCACGAGCGGCACCCCATCTTCGGCAACATCGTGGACGAGGTGGC CTACCACGAGAAGTACCCCACCATCTACCACCTGAGAAAGAAACTGGTGGACAGCACCGACAAGGCCGAC CTGCGGCTGATCTATCTGGCCCTGGCCCACATGATCAAGTTCCGGGGCCACTTCCTGATCGAGGGCGACC TGAACCCCGACAACAGCGACGTGGACAAGCTGTTCATCCAGCTGGTGCAGACCTACAACCAGCTGTTCGA GGAAAACCCCATCAACGCCAGCGGCGTGGACGCCAAGGCCATCCTGTCTGCCAGACTGAGCAAGAGCAGA CGGCTGGAAAATCTGATCGCCCAGCTGCCCGGCGAGAAGAAGAATGGCCTGTTCGGAAACCTGATTGCCC TGAGCCTGGGCCTGACCCCCAACTTCAAGAGCAACTTCGACCTGGCCGAGGATGCCAAACTGCAGCTGAG CAAGGACACCTACGACGACGACCTGGACAACCTGCTGGCCCAGATCGGCGACCAGTACGCCGACCTGTTT CTGGCCGCCAAGAACCTGTCCGACGCCATCCTGCTGAGCGACATCCTGAGAGTGAACACCGAGATCACCA AGGCCCCCCTGAGCGCCTCTATGATCAAGAGATACGACGAGCACCACCAGGACCTGACCCTGCTGAAAGC TCTCGTGCGGCAGCAGCTGCCTGAGAAGTACAAAGAGATTTTCTTCGACCAGAGCAAGAACGGCTACGCC GGCTACATTGACGGCGGAGCCAGCCAGGAAGAGTTCTACAAGTTCATCAAGCCCATCCTGGAAAAGATGG ACGGCACCGAGGAACTGCTCGTGAAGCTGAACAGAGAGGACCTGCTGCGGAAGCAGCGGACCTTCGACAA CGGCAGCATCCCCCACCAGATCCACCTGGGAGAGCTGCACGCCATTCTGCGGCGGCAGGAAGATTTTTAC CCATTCCTGAAGGACAACCGGGAAAAGATCGAGAAGATCCTGACCTTCCGCATCCCCTACTACGTGGGCC CTCTGGCCAGGGGAAACAGCAGATTCGCCTGGATGACCAGAAAGAGCGAGGAAACCATCACCCCCTGGAA CTTCGAGGAAGTGGTGGACAAGGGCGCTTCCGCCCAGAGCTTCATCGAGCGGATGACCAACTTCGATAAG AACCTGCCCAACGAGAAGGTGCTGCCCAAGCACAGCCTGCTGTACGAGTACTTCACCGTGTATAACGAGC TGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAAGCCCGCCTTCCTGAGCGGCGAGCAGAAAAAGGC CATCGTGGACCTGCTGTTCAAGACCAACCGGAAAGTGACCGTGAAGCAGCTGAAAGAGGACTACTTCAAG AAAATCGAGTGCTTCGACTCCGTGGAAATCTCCGGCGTGGAAGATCGGTTCAACGCCTCCCTGGGCACAT ACCACGATCTGCTGAAAATTATCAAGGACAAGGACTTCCTGGACAATGAGGAAAACGAGGACATTCTGGA AGATATCGTGCTGACCCTGACACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTGAAAACCTATGCC CACCTGTTCGACGACAAAGTGATGAAGCAGCTGAAGCGGCGGAGATACACCGGCTGGGGCAGGCTGAGCC GGAAGCTGATCAACGGCATCCGGGACAAGCAGTCCGGCAAGACAATCCTGGATTTCCTGAAGTCCGACGG CTTCGCCAACAGAAACTTCATGCAGCTGATCCACGACGACAGCCTGACCTTTAAAGAGGACATCCAGAAA GCCCAGGTGTCCGGCCAGGGCGATAGCCTGCACGAGCACATTGCCAATCTGGCCGGCAGCCCCGCCATTA AGAAGGGCATCCTGCAGACAGTGAAGGTGGTGGACGAGCTCGTGAAAGTGATGGGCCGGCACAAGCCCGA GAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCACCCAGAAGGGACAGAAGAACAGCCGCGAGAGA ATGAAGCGGATCGAAGAGGGCATCAAAGAGCTGGGCAGCCAGATCCTGAAAGAACACCCCGTGGAAAACA CCCAGCTGCAGAACGAGAAGCTGTACCTGTACTACCTGCAGAATGGGCGGGATATGTACGTGGACCAGGA ACTGGACATCAACCGGCTGTCCGACTACGATGTGGACCATATCGTGCCTCAGAGCTTTCTGAAGGACGAC TCCATCGACAACAAGGTGCTGACCAGAAGCGACAAGAACCGGGGCAAGAGCGACAACGTGCCCTCCGAAG AGGTCGTGAAGAAGATGAAGAACTACTGGCGGCAGCTGCTGAACGCCAAGCTGATTACCCAGAGAAAGTT CGACAATCTGACCAAGGCCGAGAGAGGCGGCCTGAGCGAACTGGATAAGGCCGGCTTCATCAAGAGACAG CTGGTGGAAACCCGGCAGATCACAAAGCACGTGGCACAGATCCTGGACTCCCGGATGAACACTAAGTACG ACGAGAATGACAAGCTGATCCGGGAAGTGAAAGTGATCACCCTGAAGTCCAAGCTGGTGTCCGATTTCCG GAAGGATTTCCAGTTTTACAAAGTGCGCGAGATCAACAACTACCACCACGCCCACGACGCCTACCTGAAC GCCGTCGTGGGAACCGCCCTGATCAAAAAGTACCCTAAGCTGGAAAGCGAGTTCGTGTACGGCGACTACA AGGTGTACGACGTGCGGAAGATGATCGCCAAGAGCGAGCAGGAAATCGGCAAGGCTACCGCCAAGTACTT CTTCTACAGCAACATCATGAACTTTTTCAAGACCGAGATTACCCTGGCCAACGGCGAGATCCGGAAGCGG CCTCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGTGGGATAAGGGCCGGGATTTTGCCACCGTGC GGAAAGTGCTGAGCATGCCCCAAGTGAATATCGTGAAAAAGACCGAGGTGCAGACAGGCGGCTTCAGCAA AGAGTCTATCCTGCCCAAGAGGAACAGCGATAAGCTGATCGCCAGAAAGAAGGACTGGGACCCTAAGAAG TACGGCGGCTTCGACAGCCCCACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAAGGGCAAGT CCAAGAAACTGAAGAGTGTGAAAGAGCTGCTGGGGATCACCATCATGGAAAGAAGCAGCTTCGAGAAGAA TCCCATCGACTTTCTGGAAGCCAAGGGCTACAAAGAAGTGAAAAAGGACCTGATCATCAAGCTGCCTAAG TACTCCCTGTTCGAGCTGGAAAACGGCCGGAAGAGAATGCTGGCCTCTGCCGGCGAACTGCAGAAGGGAA ACGAACTGGCCCTGCCCTCCAAATATGTGAACTTCCTGTACCTGGCCAGCCACTATGAGAAGCTGAAGGG CTCCCCCGAGGATAATGAGCAGAAACAGCTGTTTGTGGAACAGCACAAGCACTACCTGGACGAGATCATC GAGCAGATCAGCGAGTTCTCCAAGAGAGTGATCCTGGCCGACGCTAATCTGGACAAAGTGCTGTCCGCCT ACAACAAGCACCGGGATAAGCCCATCAGAGAGCAGGCCGAGAATATCATCCACCTGTTTACCCTGACCAA TCTGGGAGCCCCTGCCGCCTTCAAGTACTTTGACACCACCATCGACCGGAAGAGGTACACCAGCACCAAA GAGGTGCTGGACGCCACCCTGATCCACCAGAGCATCACCGGCCTGTACGAGACACGGATCGACCTGTCTC AGCTGGGAGGTGACTCTGGCGGCTCAAAAAGAACCGCCGACGGCAGCGAATTCGAGCCCAAGAAGAAGAG GAAAGTCTAACCGGTCATCATCACCATCACCATTGAGTTTAAACCCGCTGATCAGCCTCGACTGTGCCTT CTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCAC TGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGT GGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCT CTATGGCTTCTGAGGCGGAAAGAACCAGCTGGGGCTCGATACCGTCGACCTCTAGCTAGAGCTTGGCGTA ATCATGGTCATAGCTGTTTCCTGTGTGAAATTGTTATCCGCTCACAATTCCACACAACATACGAGCCGGA AGCATAAAGTGTAAAGCCTAGGGTGCCTAATGAGTGAGCTAACTCACATTAATTGCGTTGCGCTCACTGC CCGCTTTCCAGTCGGGAAACCTGTCGTGCCAGCTGCATTAATGAATCGGCCAACGCGCGGGGAGAGGCGG TTTGCGTATTGGGCGCTCTTCCGCTTCCTCGCTCACTGACTCGCTGCGCTCGGTCGTTCGGCTGCGGCGA GCGGTATCAGCTCACTCAAAGGCGGTAATACGGTTATCCACAGAATCAGGGGATAACGCAGGAAAGAACA TGTGAGCAAAAGGCCAGCAAAAGGCCAGGAACCGTAAAAAGGCCGCGTTGCTGGCGTTTTTCCATAGGCT CCGCCCCCCTGACGAGCATCACAAAAATCGACGCTCAAGTCAGAGGTGGCGAAACCCGACAGGACTATAA AGATACCAGGCGTTTCCCCCTGGAAGCTCCCTCGTGCGCTCTCCTGTTCCGACCCTGCCGCTTACCGGAT ACCTGTCCGCCTTTCTCCCTTCGGGAAGCGTGGCGCTTTCTCATAGCTCACGCTGTAGGTATCTCAGTTC GGTGTAGGTCGTTCGCTCCAAGCTGGGCTGTGTGCACGAACCCCCCGTTCAGCCCGACCGCTGCGCCTTA TCCGGTAACTATCGTCTTGAGTCCAACCCGGTAAGACACGACTTATCGCCACTGGCAGCAGCCACTGGTA ACAGGATTAGCAGAGCGAGGTATGTAGGCGGTGCTACAGAGTTCTTGAAGTGGTGGCCTAACTACGGCTA CACTAGAAGAACAGTATTTGGTATCTGCGCTCTGCTGAAGCCAGTTACCTTCGGAAAAAGAGTTGGTAGC TCTTGATCCGGCAAACAAACCACCGCTGGTAGCGGTGGTTTTTTTGTTTGCAAGCAGCAGATTACGCGCA GAAAAAAAGGATCTCAAGAAGATCCTTTGATCTTTTCTACGGGGTCTGACACTCAGTGGAACGAAAACTC ACGTTAAGGGATTTTGGTCATGAGATTATCAAAAAGGATCTTCACCTAGATCCTTTTAAATTAAAAATGA AGTTTTAAATCAATCTAAAGTATATATGAGTAAACTTGGTCTGACAGTTACCAATGCTTAATCAGTGAGG CACCTATCTCAGCGATCTGTCTATTTCGTTCATCCATAGTTGCCTGACTCCCCGTCGTGTAGATAACTAC GATACGGGAGGGCTTACCATCTGGCCCCAGTGCTGCAATGATACCGCGAGACCCACGCTCACCGGCTCCA GATTTATCAGCAATAAACCAGCCAGCCGGAAGGGCCGAGCGCAGAAGTGGTCCTGCAACTTTATCCGCCT CCATCCAGTCTATTAATTGTTGCCGGGAAGCTAGAGTAAGTAGTTCGCCAGTTAATAGTTTGCGCAACGT TGTTGCCATTGCTACAGGCATCGTGGTGTCACGCTCGTCGTTTGGTATGGCTTCATTCAGCTCCGGTTCC CAACGATCAAGGCGAGTTACATGATCCCCCATGTTGTGCAAAAAAGCGGTTAGCTCCTTCGGTCCTCCGA TCGTTGTCAGAAGTAAGTTGGCCGCAGTGTTATCACTCATGGTTATGGCAGCACTGCATAATTCTCTTAC TGTCATGCCATCCGTAAGATGCTTTTCTGTGACTGGTGAGTACTCAACCAAGTCATTCTGAGAATAGTGT ATGCGGCGACCGAGTTGCTCTTGCCCGGCGTCAATACGGGATAATACCGCGCCACATAGCAGAACTTTAA AAGTGCTCATCATTGGAAAACGTTCTTCGGGGCGAAAACTCTCAAGGATCTTACCGCTGTTGAGATCCAG TTCGATGTAACCCACTCGTGCACCCAACTGATCTTCAGCATCTTTTACTTTCACCAGCGTTTCTGGGTGA GCAAAAACAGGAAGGCAAAATGCCGCAAAAAAGGGAATAAGGGCGACACGGAAATGTTGAATACTCATAC TCTTCCTTTTTCAATATTATTGAAGCATTTATCAGGGTTATTGTCTCATGAGCGGATACATATTTGAATG TATTTAGAAAAATAAACAAATAGGGGTTCCGCGCACATTTCCCCGAAAAGTGCCACCTGACGTCGACGGA TCGGGAGATCGATCTCCCGATCCCCTAGGGTCGACTCTCAGTACAATCTGCTCTGATGCCGCATAGTTAA GCCAGTATCTGCTCCCTGCTTGTGTGTTGGAGGTCGCTGAGTAGTGCGCGAGCAAAATTTAAGCTACAAC AAGGCAAGGCTTGACCGACAATTGCATGAAGAATCTGCTTAGGGTTAGGCGTTTTGCGCTGCTTCGCGAT GTACGGGCCAGATATACGCGTTGACATTGATTATTGACTAGTTATTAATAGTAATCAATTACGGGGTCAT TAGTTCATAGCCCATATATGGAGTTCCGCGTTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCC CAACGACCCCCGCCCATTGACGTCAATAATGACGTATGTTCCCATAGTAACGCCAATAGGGACTTTCCAT TGACGTCAATGGGTGGAGTATTTACGGTAAACTGCCCACTTGGCAGTACATCAAGTGTATC

[0091] "Base editing activity" means acting to chemically change a base within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, base editing activity is cytidine deaminase activity, for example, conversion from a target C·G to T·A. In another embodiment, base editing activity is adenosine or adenine deaminase activity, for example, conversion from A·T to G·C. In another embodiment, base editing activity is cytidine deaminase activity, for example, conversion from a target C·G to T·A, and adenosine or adenine deaminase activity, for example, conversion from A·T to G·C.

[0092] The term "base editor system" refers to a system for editing nucleic acid bases of a target nucleotide sequence. In various embodiments, a base editing (BE) system comprises (1) a polynucleotide-programmable nucleotide binding domain, a deaminase domain (e.g., a cytidine deaminase or an adenosine deaminase) for deaminating a nucleic acid base in a target nucleotide sequence; and (2) one or more guide polynucleotides (e.g., guide RNA) combined with the polynucleotide-programmable nucleotide binding domain. In various embodiments, a base editing (BE) system comprises a nucleic acid base editor domain selected from an adenosine deaminase or a cytidine deaminase, and a domain having nucleic acid sequence-specific binding activity. In some embodiments, a base editor system comprises (1) a base editor (BE) comprising a polynucleotide-programmable DNA binding domain and a deaminase domain for deaminating one or more nucleic acid bases in a target nucleotide sequence; and (2) one or more guide RNAs combined with the polynucleotide-programmable DNA binding domain. In some embodiments, the polynucleotide-programmable nucleotide binding domain is a polynucleotide-programmable DNA binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE) or a cytidine base editor (CBE).

[0093] The term "Cas9" or "Cas9 domain" refers to an RNA-guided nuclease that includes a Cas9 protein or a fragment thereof (e.g., a protein containing an active, inactive, or partially active DNA cleavage domain of Cas9 and / or the gRNA-binding domain of Cas9). The Cas9 nuclease may also be referred to as the casnl nuclease or a CRISPR (clustered regularly interspaced short palindromic repeat)-associated nuclease. An exemplary Cas9 is Streptococcus pyogenes Cas9 (spCas9), the amino acid sequence of which is shown below: JPEG0007717684000009.jpg182169

[0094] The term "Cas12b" or "Cas12b domain" refers to an RNA-guided nuclease that includes a Cas12b / C2c1 protein or a fragment thereof (e.g., a protein containing an active, inactive, or partially active DNA cleavage domain of Cas12b and / or the gRNA-binding domain of Cas12b), the respective contents of which are incorporated herein by reference). Cas12b orthologs have been shown in various species including, but not limited to, Alicyclobacillus acidoterrestris, Alicyclobacillus acidophilus (Teng et al., Cell Discov. 2018 Nov 27; 4: 63), Bacillus hisashi, and Bacillus species V3-13. Additional suitable Cas12b nucleases and sequences will be apparent to those skilled in the art based on the present disclosure.

[0095] In some embodiments, a protein comprising Cas12b or a fragment thereof is referred to as a "Cas12b variant". A Cas12b variant shares homology with Cas12b or a fragment thereof. For example, a Cas12b variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to wild-type Cas12b. In some embodiments, a Cas12b variant can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type Cas12b. In some embodiments, a Cas12b variant comprises a fragment of Cas12b (e.g., a gRNA binding domain or a DNA cleavage domain) such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild-type Cas12b. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid length of the corresponding wild-type Cas12b. Exemplary Cas12b polypeptides are described below. Cas12b / C2c1 (uniprot.org / uniprot / T0D7A2#2) sp|T0D7A2|C2C1_ALIAG CRISPR-associated endonuclease C2c1 OS=Alicyclobacillus acido-terrestris (strain ATCC49025 / DSM3922 / CIP106132 / NCIMB13137 / GD3B) GN=c2c1 PE=1 SV=1 AacCas12b (Alicyclobacillus acidiphilus)-WP_067623834 BhCas12b (Bacillus hisashii) NCBI reference sequence: WP_095142515 JPEG0007717684000010.jpg A variant called BvCas12b V4 contains changes of S893R, K846R, and E837G compared to the above wild-type sequence. BvCas12b (Bacillus species V3-13) NCBI reference sequence: WP_101661451.1 MAIRSIKLKMKTNSGTDSIYLRKALWRTHQLINEGIAYYMNLLTLYRQEAIGDKTKEAYQAELINIIRNQQRNNGSSEEHGSDQEILALLRQLYELIIPSSIGESGDANQLGNKFLYPLVDPNSQSGKGTSNAGRKPRWKRLKEEGNPDWELEKKKDEERKAKDPTVKIFDNLNKYGLLPLFPLFTNIQKDIEWLPLGKRQSVRKWDKDMFIQAIERLLSWESWNRRVADEYKQLKEKTESYYKEHLTGGEEWIEKIRKFEKERNMELEKNAFAPNDGYFITSRQIRGWDRVYEKWSKLPESASPEELWKVVAEQQNKMSEGFGDPKVFSFLANRENRDIWRGHSERIYHIAAYNGLQKKLSRTKEQATFTLPDAIEHPLWIRYESPGGTNLNLFKLEEKQKKNYYVTLSKIIWPSEEKWIEKENIEIPLAPSIQFNRQIKLKQHVKGKQEISFSDYSSRISLDGVLGGSRIQFNRKYIKNHKELLGEGDIGPVFFNLVVDVAPLQETRNGRLQSPIGKALKVISSDFSKVIDYKPKELMDWMNTGSASNSFGVASLLEGMRVMSIDMGQRTSASVSIFEVVKELPKDQEQKLFYSINDTELFAIHKRSFLLNLPGEVVTKNNKQQRQERRKKRQFVRSQIRMLANVLRLETKKTPDERKKAIHKLMEIVQSYDSWTASQKEVWEKELNLLTNMAAFNDEIWKESLVELHHRIEPYVGQIVSKWRKGLSEGRKNLAGISMWNIDELEDTRRLLISWSKRSRTPGEANRIETDEPFGSSLLQHIQNVKDDRLKQMANLIIMTALGFKYDKEEKDRYKRWKETYPACQIILFENLNRYLFNLDRSRRENSRLMKWAHRSIPRTVSMQGEMFGLQVGDVRSEYSSRFHAKTGAPGIRCHALTEEDLKAGSNTLKRLIEDGFINESELAYLKKGDIIPSQGGELFVTLSKRYKKDSDNNELTVIHADINAAQNLQKRFWQQNSEVYRVPCQLARMGEDKLYIPKSQTETIKKYFGKGSFVKNNTEQEVYKWEKSEKMKIKTDTTFDLQDLDGFEDISKTIELAQEQQKKYLTMFRDPSGYFFNNETWRPQKEYWSIVNNIIKSCLKKKILSNKVEL

[0096] The term "conservative amino acid substitution" or "conservative mutation" refers to the replacement of one amino acid with another amino acid having common characteristics. A functional way to define the common characteristics among individual amino acids is to analyze the normalized frequency of amino acid changes between corresponding proteins of homologous organisms (Schulz, G. E. and Schirmer, R. H., Principles of Protein Structure, Springer-Verlag, New York (1979)). According to such an analysis, groups of amino acids can be defined when the amino acids within the group are preferentially exchanged with each other and thus have the most similar effects on the overall protein structure (Schulz, G. E. and Schirmer, R. H., supra). Non-limiting examples of conservative mutations include amino acid substitutions, such as the substitution of arginine with lysine and vice versa, so as to be able to maintain a positive charge; the substitution of aspartic acid with glutamic acid and vice versa, so as to be able to maintain a negative charge; the substitution of threonine with serine so as to be able to maintain a free -OH; and the substitution of asparagine with glutamine so as to be able to maintain a free -NH2.

[0097] The terms "coding sequence" or "protein coding sequence", which are used interchangeably herein, refer to a segment of a polynucleotide that encodes a protein. The coding sequence may also be referred to as an open reading frame. The region or sequence is bounded by a start codon near the 5' end and a stop codon near the 3' end. Examples of stop codons useful in the base editors described herein include: JPEG0007717684000011.jpg38170

[0098] "Cytidine deaminase" means a polypeptide or a fragment thereof that can catalyze a deamination reaction that converts an amino group to a carbonyl group. In one embodiment, cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine. PmCDA1 derived from Petromyzon marinus (Petromyzon marinus cytidine deaminase 1, "PmCDA1"), AID (activation-induced cytidine deaminase; AICDA) derived from mammals (e.g., humans, pigs, cows, horses, monkeys, etc.), and APOBEC are exemplary cytidine deaminases.

[0099] As used herein, the term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes the hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. In some embodiments, the deaminase or deaminase domain is a cytosine deaminase that catalyzes the hydrolytic deamination of cytosine to uracil. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or adenine (A) to inosine (I). In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., modified adenosine deaminases, evolved adenosine deaminases) can be derived from any organism such as bacteria. In some embodiments, the adenosine deaminase is derived from a bacterium such as E. coli, S. aureus, S. typhi, S. putrefaciens, H. influenzae, or C. crescentus. In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the deaminase or deaminase domain is a variant of a natural deaminase derived from an organism such as a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain does not exist in nature.For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical to a native deaminase.

[0100] "Detecting" refers to identifying the presence, absence, or amount of an analyte to be detected. In one embodiment, changes in the sequence of a polynucleotide or polypeptide are detected. In another embodiment, the presence of an indel is detected.

[0101] "Detectable label" means a composition that enables a target molecule to be detected via spectroscopic, photochemical, biochemical, immunochemical, or chemical means when linked to the target molecule. For example, useful labels include radioisotopes, magnetic beads, metal beads, colloidal particles, fluorescent dyes, high electron density reagents, enzymes (e.g., enzymes commonly used in enzyme-linked immunosorbent assay (ELISA)), biotin, digoxigenin, or haptens.

[0102] "Disease" means a pathological condition or disorder that impairs or interferes with the normal function of cells, tissues, or organs.

[0103] "Effective amount" means the amount of a drug or active compound, e.g., the amount of the base editor described herein, or the amount of a drug or active compound sufficient to induce a desired biological response, that is required to relieve the symptoms of a disease as compared to an untreated patient or an individual without the disease, i.e., a healthy individual. For the therapeutic treatment of a disease, the effective amount of the active compound(s) used to practice the present invention will vary depending on the mode of administration, the age, weight, and general health of the subject. Ultimately, the attending physician or veterinarian will determine the appropriate amount and dosing regimen. Such amount is referred to as an "effective" amount. In one embodiment, the effective amount is the amount of the base editor of the present invention sufficient to introduce a change into a gene of interest in a cell (e.g., an in vitro or in vivo cell). In one embodiment, the effective amount is the amount of the base editor required to achieve a therapeutic effect. Such a therapeutic effect need not be sufficient to alter the pathogenic gene in all cells of the subject, tissue, or organ, and may be sufficient to alter the pathogenic gene in about 1%, 5%, 10%, 25%, 50%, 75% or more of the cells present in the subject, tissue, or organ. In one embodiment, the effective amount is sufficient to improve one or more symptoms of the disease.

[0104] In some embodiments, the effective amount of a nucleic acid base editor provided herein, e.g., a fusion protein comprising an nCas9 domain and a deaminase domain (e.g., adenosine deaminase, cytidine deaminase), refers to the amount sufficient to induce editing of a target site at which the nucleic acid base editor described herein specifically binds and edits. As will be understood by those skilled in the art, the effective amount of a drug, e.g., a fusion protein, can vary depending on various factors such as, for example, the desired biological response, e.g., the specific allele, genome or target site to be edited, the cell or tissue to be targeted, and / or the drug being used.

[0105] In some embodiments, an effective amount of a fusion protein provided herein, such as a fusion protein comprising an nCas9 domain and a deaminase domain, can refer to an amount of the fusion protein sufficient to induce editing of a target site where the fusion protein specifically binds and edits. As will be appreciated by those skilled in the art, the effective amount of an agent, such as a fusion protein, nuclease, hybrid protein, protein dimer, complex of a protein (or protein dimer) and a polynucleotide, or polynucleotide, can vary depending on, for example, various factors such as the desired biological response, such as a particular allele, genome or target site to be edited, the cell or tissue to be targeted, and / or the agent being used.

[0106] "Fragment" means a portion of a polypeptide or nucleic acid molecule. This portion preferably contains at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the full length of the reference nucleic acid molecule or polypeptide. Fragments can contain 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids.

[0107] "Guide RNA" or "gRNA" means a polynucleotide that is specific for a target sequence and can form a complex with a polynucleotide-programmable nucleotide-binding domain protein (e.g., Cas9 or Cpf1). In one embodiment, the guide polynucleotide is guide RNA (gRNA). The gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule may be referred to as a single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to a guide RNA that exists as a single molecule or as a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species includes two domains: (1) a domain that shares homology with the target nucleic acid (e.g., the domain that induces binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to tracrRNA as described in Jinek et al., Science 337: 816-821 (2012), the entire content of which is incorporated herein by reference. Other examples of gRNAs (e.g., those containing domain 2) can be found in US20160208288 entitled "Switchable Cas9 Nucleases and Uses Thereof" and US 9,737,604 entitled "Delivery System For Functional Nucleases", the entire contents of each of which are incorporated herein by reference in their entirety. In some embodiments, the gRNA includes two or more domains (1) and (2) and can be referred to as an "extended gRNA". An extended gRNA binds to two or more Cas9 proteins and binds to the target nucleic acid at two or more different regions, as described herein. The gRNA includes a nucleotide sequence complementary to the target site, which mediates the binding of the nuclease / RNA complex to the target site and provides sequence specificity of the nuclease:RNA complex.

[0108] "Heterodimer" means a fusion protein comprising two domains such as a wild-type TadA domain and a variant of the TadA domain (e.g., TadA*8 or TadA*9) or two variant TadA domains (e.g., TadA*7.10 and TadA*8 or two TadA*8 domains; or TadA*7.10 and TadA*9 or two TadA*9 domains).

[0109] "Hybridization" means being hydrogen-bonded, which may be a Watson-Crick type, Hoogsteen type, or reverse Hoogsteen type hydrogen bond between complementary nucleobases. For example, adenine and thymine are complementary nucleobases that pair through the formation of a hydrogen bond.

[0110] "Increase" means a positive change of at least 10%, 25%, 50%, 75%, or 100%.

[0111] The term "inhibitor of base repair", "base repair inhibitor", "IBR" or their grammatical equivalents refers to a protein that can inhibit the activity of nucleic acid repair enzymes, such as base excision repair enzymes. In some embodiments, IBR is an inhibitor of inosine base excision repair. Exemplary inhibitors of base repair include inhibitors of APE1, Endo III, Endo IV, Endo V, Endo VIII, Fpg, hOGG1, hNEIL1, T7 Endol, T4PDG, UDG, hSMUG1, and hAAG. In some embodiments, the base repair inhibitor is an inhibitor of Endo V or hAAG. In some embodiments, IBR is an inhibitor of Endo V or hAAG. In some embodiments, IBR is catalytically inactive EndoV or catalytically inactive hAAG. In some embodiments, the base repair inhibitor is catalytically inactive EndoV or catalytically inactive hAAG. In some embodiments, the base repair inhibitor is uracil glycosylase inhibitor (UGI). UGI is a protein that can inhibit the base excision repair enzyme of uracil DNA glycosylase. In some embodiments, the UGI domain comprises wild-type UGI or a fragment of wild-type UGI. In some embodiments, the UGI proteins provided herein include fragments of UGI and proteins homologous to UGI or UGI fragments. In some embodiments, the base repair inhibitor is an inhibitor of inosine base excision repair. In some embodiments, the base repair inhibitor is a "catalytically inactive inosine-specific nuclease" or an "inactive inosine-specific nuclease". Without wishing to be bound by any particular theory, catalytically inactive inosine glycosylase (e.g., alkyladenine glycosylase (AAG)) can bind to inosine but cannot create an abasic site or remove inosine, and as a result, the newly formed inosine moiety is sterically blocked from the DNA damage / repair machinery. In some embodiments, the catalytically inactive inosine-specific nuclease can bind to inosine in a nucleic acid but does not cleave the nucleic acid.Non-limiting exemplary catalytically inactive inosine-specific nucleases include, for example, a catalytically inactive alkyladenosine glycosylase (AAG nuclease) derived from humans, and, for example, a catalytically inactive endonuclease V (EndoV nuclease) derived from E. coli. In some embodiments, the catalytically inactive AAG nuclease includes the E125Q mutation or a corresponding mutation in another AAG nuclease.

[0112] An "intein" is a protein fragment that can excise itself and ligate the remaining fragments (exteins) via a peptide bond in a process called protein splicing. An intein is also referred to as a "protein intron". The process by which an intein excises itself and ligates the remaining portion of the protein is referred to herein as "protein splicing" or "intein-mediated protein splicing". In some embodiments, the intein of a precursor protein (the intein-containing protein prior to intein-mediated protein splicing) is derived from two genes. Such an intein is referred to herein as a split (divided) intein (e.g., split intein-N and split intein-C). For example, in cyanobacteria, DnaE, which is the catalytic subunit a of DNA polymerase III, is encoded by two separate genes, dnaE-n and dnaE-c. The intein encoded by the dnaE-n gene may be referred to herein as "intein-N". The intein encoded by the dnaE-c gene may be referred to herein as "intein-C".

[0113] Other inteins can also be used. For example, synthetic inteins based on dnaE intein, Cfa-N (e.g., split intein-N) and Cfa-C (e.g., split intein-C) intein pairs are described (e.g., Stevens et al., J Am Chem Soc. 2016 Feb. 24; 138 (7): 2162-5, which is incorporated herein by reference). Non-limiting examples of intein pairs that can be used according to the present disclosure include: Cfa DnaE intein, Spp GyrB intein, Spp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB intein, and Cne Prp8 intein (e.g., as described in U.S. Patent No. 8,394,604, which is incorporated herein by reference).

[0114] Exemplary nucleotide and amino acid sequences of inteins are shown.

[0115] DnaE intein-N DNA: TGCCTGTCATACGAAACCGAGATACTGACAGTAGAATATGGCCTTCTGCCAATCGGGAAGATTGTGGAGAAACGGATAGAATGCACAGTTTACTCTGTCGATAACAATGGTAACATTTATACTCAGCCAGTTGCCCAGTGGCACGACCGGGGAGAGCAGGAAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAGGGCCACTAAGGACCACAAATTTATGACAGTCGATGGCCAGATGCTGCCTATAGACGAAATCTTTGAGCGAGAGTTGGACCTCATGCGAGTTGACAACCTTCCTAAT DnaE intein-N protein: CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDR GEQEVFEYCLEDGSLIRATKDHKFMTVDGQMLPIDEIFERELDLMRVDNL PN DnaE intein-C DNA: ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGA TATTGGAGTCGAAAGAGATCACAACTTTGCTCTGAAGAACGGATTCATAG CTTCTAAT Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN Cfa-N DNA: TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGCCTATTGGAAAGATTGTCGAAGAGAGAATTGAATGCACAGTATATACTGTAGACAAGAATGGTTTCGTTTACACACAGCCCATTGCTCAATGGCACAATCGCGGCGAACAAGAAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACGAGCAACTAAAGATCATAAATTCATGACCACTGACGGGCAGATGTTGCCAATAGATGAGATATTCGAGCGGGGCTTGGATCTCAAACAAGTGGATGGATTGCCA Cfa-N protein: CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQMLPIDEIFERGLDLKQVDGLP Cfa-C DNA: ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAGGAAAGTAAAGATAATATCTCGAAAAAGTCTTGGTACCCAAAATGTCTATGATATTGGAGTGGAGAAAGATCACAACTTCCTTCTCAAGAACGGTCTCGTAGCCAGCAAC Cfa-C protein: MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLVASN

[0116] To join the N-terminal portion of split Cas9 and the C-terminal portion of split Cas9, intein-N and intein-C may be fused to the N-terminal portion of split Cas9 and the C-terminal portion of split Cas9, respectively. For example, in some embodiments, intein-N is fused to the C-terminus of the N-terminal portion of split Cas9, i.e., to form a structure of N--[N-terminal portion of split Cas9]-[intein-N]--C. In some embodiments, intein-C is fused to the N-terminus of the C-terminal portion of split Cas9, i.e., to form a structure of N-[intein-C]--[C-terminal portion of split Cas9]-C. Intein-mediated protein splicing mechanisms for joining proteins (e.g., split Cas9) to which an intein is fused are known in the art, as described, for example, in Shah et al., Chem Sci. 2014; 5 (1): 446-461, which is hereby incorporated by reference in its entirety. Methods for designing and using inteins are known in the art and are described, for example, in WO2014004336, WO2017132580, US20150344549, and US20180127780, each of which is hereby incorporated by reference in its entirety.

[0117] The terms "isolated", "purified", or "biologically pure" refer to substances that have been removed to varying degrees from the components that are normally associated with them as found in their natural state. "Isolating" indicates the degree of separation from the original source or surroundings. "Purifying" indicates a higher degree of separation than isolation. A "purified" or "biologically pure" protein is sufficiently free of other substances such that impurities do not substantially affect the biological properties of the protein or cause other adverse effects. That is, a nucleic acid or peptide of the present invention is purified if, when produced by recombinant DNA technology, it substantially lacks cellular material, viral material, or culture medium, or if chemically synthesized, it substantially lacks chemical precursors or other chemicals. Purity and homogeneity are usually measured using analytical chemistry techniques such as polyacrylamide gel electrophoresis or high performance liquid chromatography. The term "purified" can indicate that a nucleic acid or protein yields essentially one band on an electrophoretic gel. In the case of proteins that can be subjected to modifications such as phosphorylation or glycosylation, different modifications can yield different isolated proteins that can be purified separately.

[0118] "Isolated polynucleotide" means a nucleic acid (e.g., DNA) that does not contain the genes adjacent to the gene in the natural genome of the organism from which the nucleic acid molecule of the present invention is derived. Thus, this term includes, for example, recombinant DNA incorporated into a vector; recombinant DNA incorporated into an autonomously replicating plasmid or virus; or recombinant DNA incorporated into the genomic DNA of a prokaryote or eukaryote; or recombinant DNA that exists as a separate molecule independent of other sequences (e.g., cDNA or genomic or cDNA fragments produced by PCR or restriction endonuclease digestion). Further, this term includes RNA molecules transcribed from a DNA molecule, and recombinant DNA that is part of a hybrid gene encoding additional polypeptide sequences.

[0119] "Isolated polypeptide" means a polypeptide of the present invention that is separated from the components naturally associated therewith. Generally, a polypeptide is isolated when at least 60% by weight of the proteins and natural organic molecules where it is naturally associated have been removed. Preferably, the preparation is at least 75% by weight, more preferably at least 90% by weight, and most preferably at least 99% by weight of the polypeptide of the present invention. The isolated polypeptide of the present invention may be obtained, for example, by extraction from a natural source, by expression of a recombinant nucleic acid encoding such a polypeptide; or by chemically synthesizing the protein. The purity can be measured by appropriate methods, such as column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.

[0120] As used herein, the term "linker" can refer to a covalent linker (e.g., a covalent bond), a non-covalent linker, a chemical group, or a molecule that connects two molecules or moieties, such as two components of a protein complex or ribonucleo complex, or two domains of a fusion protein, such as a polynucleotide-programmable DNA binding domain (e.g., dCas9) and a deaminase domain (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase). The linker can connect various components or different parts of a base editor system. For example, in some embodiments, the linker can connect the guide polynucleotide binding domain of a polynucleotide-programmable nucleotide binding domain and the catalytic domain of a deaminase. In some embodiments, the linker can connect a CRISPR polypeptide and a deaminase. In some embodiments, the linker can connect Cas9 and a deaminase. In some embodiments, the linker can connect dCas9 and a deaminase. In some embodiments, the linker can connect nCas9 and a deaminase. In some embodiments, the linker can connect a guide polynucleotide and a deaminase. In some embodiments, the linker can connect the deaminating component of a base editor system and a polynucleotide-programmable nucleotide binding component. In some embodiments, the linker can connect the RNA-binding portion of the deaminating component of a base editor system and a polynucleotide-programmable nucleotide binding component. In some embodiments, the linker can connect the RNA-binding portion of the deaminating component of a base editor system and the RNA-binding portion of a polynucleotide-programmable nucleotide binding component. The linker is located between or adjacent to two groups, molecules, or other moieties and connects each through covalent or non-covalent interaction, thus being able to connect the two. In some embodiments, the linker can be an organic molecule, group, polymer, or chemical moiety.In some embodiments, the linker can be a polynucleotide. In some embodiments, the linker can be a DNA linker. In some embodiments, the linker can be an RNA linker. In some embodiments, the linker can include an aptamer that can bind to a ligand. In some embodiments, the ligand can be a carbohydrate, a peptide, a protein, or a nucleic acid. In some embodiments, the linker can include an aptamer derived from a riboswitch. The riboswitch from which the aptamer is derived can be selected from a theophylline riboswitch, a thiamine pyrophosphate (TPP) riboswitch, an adenosylcobalamin (AdoCbl) riboswitch, an S-adenosylmethionine (SAM) riboswitch, an SAH riboswitch, a flavin mononucleotide (FMN) riboswitch, a tetrahydrofolate riboswitch, a lysine riboswitch, a glycine riboswitch, a purine riboswitch, a GlmS riboswitch, or a prequeuosine 1 (PreQ1) riboswitch. In some embodiments, the linker can include an aptamer bound to a polypeptide or protein domain such as a polypeptide ligand. In some embodiments, the polypeptide ligand can be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile α motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the polypeptide ligand can be part of a base editor system component. For example, a nucleic acid base editing component can include a deaminase domain and an RNA recognition motif.

[0121] In some embodiments, the linker can be one amino acid or multiple amino acids (e.g., a peptide or a protein). In some embodiments, the linker can have a length of about 5 to 100 amino acids, such as about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20 - 30, 30 - 40, 40 - 50, 50 - 60, 60 - 70, 70 - 80, 80 - 90, or 90 - 100 amino acids. In some embodiments, the linker can have a length of about 100 - 150, 150 - 200, 200 - 250, 250 - 300, 300 - 350, 350 - 400, 400 - 450, or 450 - 500 amino acids. Longer or shorter linkers can also be used. Longer or shorter linkers are also contemplated. In some embodiments, the linker comprises the amino acid sequence SGSETPGTSESATPES, which can also be referred to as an XTEN linker. In some embodiments, the linker comprises the amino acid sequence SGGS. In some embodiments, the linker comprises (SGGS)n, (GGGS)n, (GGGGS)n, (G)n, (EAAAK)n, (GGS)n, SGSETPGTSESATPES, or (XP)n motifs, or any combination thereof, where in the sequence, n is independently an integer from 1 to 30 and X is any amino acid. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. In some embodiments, the linker comprises multiple proline residues and has a length of 5 - 21, 5 - 14, 5 - 9, 5 - 7 amino acids, such as PAPAP, PAPAPA, PAPAPAP, PAPAPAPA, P(AP)4, P(AP)7, P(AP) 10 and so on. Such proline - rich linkers are also referred to as "rigid" linkers.

[0122] In some embodiments, the linker binds the gRNA binding domain of an RNA-programmable nuclease, which includes a Cas9 nuclease domain, to the catalytic domain of a nucleic acid editing protein (e.g., cytidine deaminase or adenosine deaminase). In some embodiments, the linker binds dCas9 to the nucleic acid editing protein. For example, the linker is located between or adjacent to two groups, molecules, or other moieties and connects to each via a covalent bond, thus connecting the two. In some embodiments, the linker is one amino acid or multiple amino acids (e.g., a peptide or a protein). In some embodiments, the linker is an organic molecule, a group, a polymer, or a chemical moiety. In some embodiments, the linker has a length of 5 to 200 amino acids, e.g., a length of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140, 150, 160, 175, 180, 190, or 200 amino acids.

[0123] In some embodiments, the domain of the base editor is fused via a linker comprising the amino acid sequence SGGSSGSETPGTSESATPESSGGS, SGGSSGGSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTE PSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS. In some embodiments, the domain of the base editor is fused via a linker comprising the amino acid sequence SGSETPGTSESATPES, which may also be referred to as the XTEN linker. In some embodiments, the linker is 24 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In some embodiments, the linker is 40 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS. In some embodiments, the linker is 64 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGS SGGS. In some embodiments, the linker is 92 amino acids in length. In some embodiments, the linker comprises the amino acid sequence PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS.

[0124] "Marker" means any protein or polynucleotide whose expression level or activity changes in relation to a disease or disorder.

[0125] As used herein, the term "mutation" refers to a substitution of a residue within a sequence (e.g., a nucleic acid or amino acid sequence) by another residue, or a deletion or insertion of one or more residues within the sequence. Mutations are typically described herein by identifying the original residue, followed by identifying the position of the residue within the sequence, and then identifying the newly substituted residue. Various methods for performing amino acid substitutions (mutations) provided herein are well known in the art and are described, for example, in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)). In some embodiments, the base editors of the present disclosure can efficiently generate "intended mutations", such as point mutations, within a nucleic acid (e.g., a nucleic acid within the genome of a subject) without generating a significant number of unintended mutations, such as unintended point mutations. In some embodiments, an intended mutation is a mutation generated by a specific base editor (e.g., a cytidine base editor or an adenosine base editor) bound to a guide polynucleotide (e.g., a gRNA) that is specifically designed to generate that intended mutation.

[0126] Generally, mutations created or identified within a sequence (e.g., an amino acid sequence described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain the mutation. One of ordinary skill in the art will readily understand methods for determining the position of mutations in amino acid and nucleic acid sequences relative to a reference sequence.

[0127] The term "non-conservative mutation" includes, for example, amino acid substitutions between different groups, such as from tryptophan to lysine, from serine to phenylalanine, etc. In this case, it is preferred that the non-conservative amino acid substitution does not interfere with or inhibit the biological activity of the functional variant. The non-conservative amino acid substitution can enhance the biological activity of the functional variant, such that the biological activity of the functional variant is increased compared to the wild-type protein.

[0128] The term "nuclear localization sequence", "nuclear localization signal", or "NLS" refers to an amino acid sequence that promotes the import of a protein into the cell nucleus. Nuclear localization sequences are known in the art and are described, for example, in Plank et al., International PCT Application, PCT / EP2000 / 011690, filed on November 23, 2000 and published as WO / 2001 / 038547 on May 31, 2001, the content of which is incorporated herein by reference for the purpose of disclosing exemplary nuclear localization sequences. In other embodiments, the NLS is, for example, an optimized NLS described by Koblan et al., Nature Biotech. 2018 doi: 10.1038 / nbt.4172. In some embodiments, the NLS comprises the amino acid sequence KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.

[0129] As used herein, the terms "nucleobase", "nitrogenous base", or "base" refer to nitrogen-containing biological compounds that form nucleosides, which are components of nucleotides. The ability of nucleobases to form base pairs and stack on top of each other directly results in long-chain helical structures such as ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). The five nucleobases (adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U)) are referred to as primary nucleobases or standard nucleobases. Adenine and guanine are derived from purines, while cytosine, uracil, and thymine are derived from pyrimidines. DNA and RNA can also include other (non-primary) bases that are modified. Non-limiting exemplary modified nucleobases can include hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydroxymethylcytosine. Hypoxanthine and xanthine are generated by the presence of mutagens and can both be generated by deamination (substitution of an amine group with a carbonyl group). Hypoxanthine can be modified from adenine. Xanthine can be modified from guanine. Uracil can result from the deamination of cytosine. A "nucleoside" consists of a nucleobase and a five-carbon sugar (either ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxyuridine, and deoxycytidine. Examples of nucleosides with modified nucleobases include inosine (I), xanthosine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), and pseudouridine (Ψ). A "nucleotide" consists of a nucleobase, a five-carbon sugar (either ribose or deoxyribose), and at least one phosphate group.

[0130] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to compounds containing nucleobases and acidic moieties, such as nucleosides, nucleotides, or polymers of nucleotides. Usually, polymeric nucleic acids, such as nucleic acid molecules containing three or more nucleotides, are linear molecules in which adjacent nucleotides are linked to each other via phosphodiester bonds. In some embodiments, "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides). In some embodiments, "nucleic acid" refers to an oligonucleotide chain containing three or more individual nucleotide residues. As used herein, the terms "oligonucleotide" and "polynucleotide" can be used interchangeably to refer to a polymer of nucleotides (e.g., a string of at least three nucleotides). In some embodiments, "nucleic acid" encompasses RNA as well as single-stranded and / or double-stranded DNA. Nucleic acids can occur naturally, for example, in the context of a genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, chromatid, or other natural nucleic acid molecule. On the other hand, a nucleic acid molecule can be a non-natural molecule, such as recombinant DNA or RNA, artificial chromosome, modified genome, or a fragment thereof, or a synthetic DNA, RNA, DNA / RNA hybrid, or can contain non-natural nucleotides or nucleosides. Further, "nucleic acid", "DNA", "RNA", and / or similar terms include nucleic acid analogs, such as analogs having other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems, and can be, for example, optionally purified or chemically synthesized. If desired, for example, in the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, such as those having chemically modified bases or sugars, and backbone modifications. Nucleic acid sequences are presented in the 5' to 3' direction unless otherwise specified.In some embodiments, the nucleic acid is a natural nucleoside (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); a nucleoside analog (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine); a chemically modified base; a biologically modified base (e.g., a methylated base); an intercalated base; a modified sugar (e.g., 2'-fluoro-ribose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or a modified phosphate group (e.g., phosphorothioate and 5'-N-phosphoramidite linkages) or comprises them.

[0131] The term "nucleic acid programmable DNA-binding protein" or "napDNAbp" may be used interchangeably with "polynucleotide programmable nucleotide-binding domain" and refers to a protein that binds to a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid or guide polynucleotide (e.g., gRNA) that directs the napDNAbp to a specific nucleic acid sequence. In some embodiments, the polynucleotide programmable nucleotide-binding domain is a polynucleotide programmable DNA-binding domain. In some embodiments, the polynucleotide programmable nucleotide-binding domain is a polynucleotide programmable RNA-binding domain. In some embodiments, the polynucleotide programmable nucleotide-binding domain is a Cas9 protein. The Cas9 protein can bind to a guide RNA that guides the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the napDNAbp is a Cas9 domain, such as nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Non-limiting examples of nucleic acid programmable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ.Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Cas12j / CasΦ, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V Cas effector proteins, type VI Cas effector proteins, CARF, DinG, homologs thereof, or modified or altered versions thereof. Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, although they may not be specifically described herein. See, for example, Makarova et al. “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct; 1:325-336. doi: 10.1089 / crispr.2018.0033; Yan et al., “Functionally diverse type V CRISPR-Cas systems” Science. 2019 Jan 4; 363 (6422): 88-91. doi:10.1126 / science.aav7271 (the entire contents of each are incorporated herein by reference).

[0132] As used herein, the term "nucleobase editing domain" or "nucleobase editing protein" refers to a protein or enzyme that can catalyze nucleobase modifications in RNA or DNA, such as deamination of cytosine (or cytidine) to uracil (or uridine) or thymine (or thymidine), and adenine (or adenosine) to hypoxanthine (or inosine), as well as addition and insertion of nucleotides in the absence of a template. In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., adenine deaminase or adenosine deaminase; or cytidine deaminase or cytosine deaminase). In some embodiments, the nucleobase editing domain is a plurality of deaminase domains (e.g., adenine deaminase or adenosine deaminase or cytidine deaminase or cytosine deaminase). In some embodiments, the nucleobase editing domain can be a native nucleobase editing domain. In some embodiments, the nucleobase editing domain can be a modified or evolved nucleobase editing domain from a native nucleobase editing domain. The nucleobase editing domain can be derived from any organism such as bacteria, human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse.

[0133] As used herein, "obtaining" in "obtaining an agent" includes synthesizing, purchasing, or obtaining the agent by other means.

[0134] As used herein, "patient" or "subject" refers to a mammalian subject or individual that has, is at risk of developing, or is diagnosed as having or suspected of having a disease or disorder. In some embodiments, the term "patient" refers to a mammalian subject that has a higher than average likelihood of developing a disease or disorder. Exemplary patients can be humans, non-human primates, cats, dogs, pigs, cows, cats, horses, camels, llamas, goats, sheep, rodents (e.g., mice, rabbits, rats, or guinea pigs), and other mammals that can benefit from the therapies disclosed herein. Exemplary human patients can be male and / or female.

[0135] As used herein, "a patient in need thereof" or "a subject in need thereof" refers to a patient who has, is at risk of, or is diagnosed as having, pre-determined as having, or suspected of having a disease or disorder.

[0136] The terms "pathogenic variant", "pathogenic mutation", "pathogenic change", "pathogenic alteration", "harmful variant", or "predisposing mutation" refer to a genetic change or mutation that increases an individual's susceptibility or predisposition to a particular disease or disorder. In some embodiments, a pathogenic variant includes one in which at least one wild-type amino acid in a protein encoded by a gene is replaced by at least one pathogenic amino acid.

[0137] The terms "protein", "peptide", "polypeptide", and their grammatical equivalents are used interchangeably herein and refer to polymers of amino acid residues linked together by peptide (amide) bonds. These terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Generally, a protein, peptide, or polypeptide is at least 3 amino acids in length. A protein, peptide, or polypeptide can refer to an individual protein or a collection of proteins. One or more amino acids in a protein, peptide, or polypeptide can be modified by the addition of chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups, linkers for conjugation, functionalization, or other modifications. A protein, peptide, or polypeptide can also be a single molecule or a multimolecular complex. A protein, peptide, or polypeptide can be a mere fragment of a native protein or peptide. A protein, peptide, or polypeptide can be native, recombinant, synthetic, or any combination thereof. As used herein, the term "fusion protein" refers to a hybrid polypeptide containing protein domains derived from at least two different proteins. One protein can be placed at the amino-terminal (N-terminal) portion or carboxy-terminal (C-terminal) portion of the fusion protein and thus, an amino-terminal fusion protein or a carboxy-terminal fusion protein can be formed, respectively. A protein can contain different domains, such as a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9 that mediates binding of the protein to a target site) and a nucleic acid cleavage domain, or the catalytic domain of a nucleic acid editing protein. In some embodiments, a protein contains a proteinaceous portion, such as an amino acid sequence that constitutes a nucleic acid binding domain, and an organic compound, such as a compound that can act as a nucleic acid cleavage agent. In some embodiments, a protein is complexed with or associated with a nucleic acid, such as RNA or DNA. Any of the proteins provided herein can be produced by any method known in the art.For example, the proteins provided herein can be generated via recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for the expression and purification of recombinant proteins are well known and include those described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), the entire content of which is incorporated herein by reference.

[0138] The polypeptides and proteins disclosed herein (including functional portions and functional variants thereof) can contain synthetic amino acids in place of one or more natural amino acids. Such synthetic amino acids are known in the art and include, for example, aminocyclohexanecarboxylic acid, norleucine, α-amino-n-decanoic acid, homoserine, S-acetylaminomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine, β-phenylserine β-hydroxy-phenylalanine, phenylglycine, α-naphthylalanine, cyclohexylalanine, cyclohexylglycine, indoline-2-carboxylic acid, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, aminomalonic acid monoamide, N'-benzyl-N'-methyl-lysine, N',N'-dibenzyl-lysine, 6-hydroxylysine, ornithine, α-aminocyclopentanecarboxylic acid, α-aminocyclohexanecarboxylic acid, α-aminocycloheptanecarboxylic acid, α-(2-amino-2-norbornane)-carboxylic acid, α,γ-diaminobutyric acid, α,β-diaminopropionic acid, homophenylalanine, and α-tert-butylglycine. Polypeptides and proteins can be accompanied by post-translational modification of one or more amino acids of the polypeptide construct. Non-limiting examples of post-translational modification include acylation including phosphorylation, acetylation and formylation, glycosylation (including N-linked and O-linked), amidation, hydroxylation, alkylation including methylation and ethylization, ubiquitination, addition of pyrrolidonecarboxylic acid, formation of disulfide bridges, sulfation, myristoylation, palmitoylation, isoprenylation, farnesylation, geranylation, glypiation, lipoylation and iodination.

[0139] As used herein with respect to a protein or nucleic acid, the term "recombinant" refers to a protein or nucleic acid that does not exist in nature but is a product of human engineering. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that contains at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7 mutations as compared to any natural sequence.

[0140] "Decrease" means a negative change of at least 10%, 25%, 50%, 75%, or 100%.

[0141] "Reference" means a standard or control condition. In one embodiment, the reference is a wild-type or healthy cell. In other embodiments, without limitation, the reference is an untreated cell that is not subjected to the test conditions or is subjected to a placebo or normal saline, medium, buffer, and / or a control vector that does not have the polynucleotide of interest.

[0142] "Reference sequence" is a defined sequence used as a basis for sequence comparison. The reference sequence may be a subset or the entirety of the designated sequence; for example, a segment of a full-length cDNA or gene sequence, or a complete cDNA or gene sequence. In the case of a polypeptide, the length of the reference polypeptide sequence is generally at least about 16 amino acids, at least about 20 amino acids, at least about 25 amino acids, about 35 amino acids, about 50 amino acids, or about 100 amino acids. In the case of a nucleic acid, the length of the reference nucleic acid sequence is generally at least about 50 nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides, or about 300 nucleotides, or any integer therearound or between them. In some embodiments, the reference sequence is the wild-type sequence of the protein of interest. In other embodiments, the reference sequence is the polynucleotide sequence encoding the wild-type protein.

[0143] The terms “RNA-programmable nuclease” and “RNA-guided nuclease” are used (e.g., bound or associated) with one or more RNAs (plural) that are not the target of cleavage. In some embodiments, an RNA-programmable nuclease, when in complex with an RNA, may be referred to as a nuclease:RNA complex. Typically, the bound RNA is called a guide RNA (gRNA). In some embodiments, the RNA-programmable nuclease is a (CRISPR-associated) Cas9 endonuclease, e.g., Cas9 (Csn1) from Streptococcus pyogenes (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti J. J. et al., Proc. Natl. Acad. Sci. U.S.A. 98: 4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E.et al., Nature 471: 602-607 (2011)).

[0144] RNA-programmable nucleases (e.g., Cas9) use RNA:DNA hybridization to target DNA cleavage sites, and thus these proteins can in principle target any sequence specified by a guide RNA. Methods for using RNA-programmable nucleases such as Cas9 for site-specific cleavage (e.g., for modifying a genome) are known in the art (see, e.g., Cong, L. et al, Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al, RNA-guided human genome engineering via Cas9. Science 339,823-826(2013); Hwang, W.Y. et al, Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al, RNA-programmed genome editing in human cells. eLife 2, e00471(2013); Dicarlo, J. E. et al, Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic acids research (2013); Jiang, W. et al RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature biotechnology 31, 233-239 (2013), each of which is incorporated herein by reference in its entirety).

[0145] The term "single nucleotide polymorphism (SNP)" refers to a variation of a single nucleotide that occurs at a specific position in the genome, and each variation is present to a significant extent (e.g., >1%) within a population. For example, at a specific base position in the human genome, the C nucleotide may occur in many individuals, but in a small number of individuals, that position is occupied by an A. This means that an SNP exists at this specific position, and the two possible nucleotide variations, C or A, are referred to as alleles at this position. SNPs underlie differences in susceptibility to diseases. The severity of diseases and the manner in which the body responds to treatment are also manifestations of genetic variations. SNPs can exist in the coding region of a gene, the non-coding region of a gene, or the intergenic region (the region between genes). In some embodiments, SNPs within the coding sequence do not necessarily change the amino acid sequence of the protein produced due to the degeneracy of the genetic code. There are two types of SNPs in the coding region: synonymous SNPs and non-synonymous SNPs. Synonymous SNPs do not affect the protein sequence, but non-synonymous SNPs change the amino acid sequence of the protein. Non-synonymous SNPs have two types: missense and nonsense. SNPs that are not in the protein-coding region can further affect gene splicing, transcription factor binding, messenger RNA degradation, or the sequence of non-coding RNAs. Gene expression affected by this type of SNP is called eSNP (expressed SNP) and can exist upstream or downstream of the gene. A single nucleotide variant (SNV) is a variation of a single nucleotide without a frequency limit and can occur in somatic cells. Somatic single nucleotide variations are also called single nucleotide modifications.

[0146] "Specifically binds" means a nucleic acid molecule, polypeptide, or complex thereof (e.g., a nucleic acid programmable DNA binding domain and a guide nucleic acid), compound, or molecule that recognizes and binds to the polypeptide and / or nucleic acid molecule of the present invention but does not substantially recognize and bind to other molecules in a sample (e.g., a biological sample).

[0147] Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but generally exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence can generally hybridize to at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but generally exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence can generally hybridize to at least one strand of a double-stranded nucleic acid molecule.

[0148] "Hybridizing" means annealing to a complementary polynucleotide sequence (e.g., a gene described herein), or a portion thereof, to form a double-stranded molecule under various stringency conditions. (See, e.g., Wahl, G. M. and S. L. Berger (1987) Methods Enzymol. 152: 399; Kimmel, A. R. (1987) Methods Enzymol. 152: 507).

[0149] For example, stringent salt concentrations are typically less than about 750 mM NaCl and 75 mM trisodium citrate, preferably less than about 500 mM NaCl and 50 mM trisodium citrate, more preferably less than about 250 mM NaCl and 25 mM trisodium citrate. In the absence of an organic solvent such as formamide, low stringency hybridization can be obtained, while in the presence of at least about 35% formamide, more preferably at least about 50% formamide, high stringency hybridization can be obtained. Stringent temperature conditions typically include a temperature of at least about 30°C, more preferably at least about 37°C, and most preferably at least about 42°C. Various additional parameters such as hybridization time, concentration of surfactant (e.g., sodium dodecyl sulfate (SDS)), and inclusion or exclusion of carrier DNA are well known to those skilled in the art. If desired, various levels of stringency can be achieved by combining these various conditions. In a preferred embodiment, hybridization is carried out at 30°C in 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS. In a more preferred embodiment, hybridization is carried out at 37°C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA). In the most preferred embodiment, hybridization is carried out at 42°C in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 μg / ml ssDNA. Useful variations of these conditions will be readily apparent to those skilled in the art.

[0150] In most applications, the stringency of the washing step following hybridization also varies widely. The stringency conditions of washing can be defined by salt concentration and temperature. As described above, the stringency of washing can be increased by lowering the salt concentration or raising the temperature. For example, a stringent salt concentration for the washing step is preferably less than about 30 mM NaCl and 3 mM trisodium citrate, most preferably less than about 15 mM NaCl and 1.5 mM trisodium citrate. Stringent temperature conditions for the washing step typically include a temperature of at least about 25°C, more preferably at least about 42°C, even more preferably at least about 68°C. In one embodiment, the washing step is carried out at 25°C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In another embodiment, the washing step is carried out at 42°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the washing step is carried out at 68°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Additional variations of these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science 196: 180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72: 3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.

[0151] "Split" means being divided into two or more fragments.

[0152] "Split Cas9 protein" or "split Cas9" refers to the Cas9 protein provided as an N-terminal fragment and a C-terminal fragment encoded by two separate nucleotide sequences. The polypeptides corresponding to the N-terminal and C-terminal portions of the Cas9 protein can be spliced to form a "reconstituted" Cas9 protein. In certain embodiments, the Cas9 protein is split into two fragments within the disordered region of the protein, as described, for example, in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014, or as described in Jiang et al. (2016) Science 351: 867-871. The PDB files: 5F9R are each incorporated herein by reference. In some embodiments, the protein is split into two fragments at any C, T, A, or S within the region of SpCas9 between approximately amino acids A292-G364, F445-K483, or E565-T637, or at the corresponding position of any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or other napDNAbp. In some embodiments, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of splitting the protein into two fragments is referred to as "splitting" the protein.

[0153] In other embodiments, the N-terminal portion of the Cas9 protein comprises amino acids 1-573 or 1-637 of S. pyogenes Cas9 wild-type (SpCas9) (NCBI reference sequence: NC_002737.2, Uniprot reference sequence: Q99ZW2), and the C-terminal portion of the Cas9 protein comprises the portion of amino acids 574-1368 or 638-1368 of SpCas9 wild-type.

[0154] The C-terminal portion of split Cas9 can be combined with the N-terminal portion of split Cas9 to form a complete Cas9 protein. In some embodiments, the C-terminal portion of the Cas9 protein starts at the position where the N-terminal portion of the Cas9 protein ends. Thus, in some embodiments, the C-terminal portion of split Cas9 includes a part of amino acids (551-651)-1368 of spCas9. "(551-651)-1368" means starting with the amino acids between amino acids 551-651 (including both ends) and ending with amino acid 1368.For example, the C-terminal portion of split Cas9 may include a part of any one of amino acids 551-1368, 552-1368, 553-1368, 554-1368, 555-1368, 556-1368, 557-1368, 558-1368, 559-1368, 560-1368, 561-1368, 562-1368, 563-1368, 564-1368, 565-1368, 566-1368, 567-1368, 568-1368, 569-1368, 570-1368, 571-1368, 572-1368, 573-1368, 574-1368, 575-1368, 576-1368, 577-1368, 578-1368, 579-1368, 580-1368, 581-1368, 582-1368, 583-1368, 584-1368, 585-1368, 586-1368, 587-1368, 588-1368, 589-1368, 590-1368, 591-1368, 592-1368, 593-1368, 594-1368, 595-1368, 596-1368, 597-1368, 598-1368, 599-1368, 600-1368, 601-1368, 602-1368, 603-1368, 604-1368, 605-1368, 606-1368, 607-1368, 608-1368, 609-1368, 610-1368, 611-1368, 612-1368, 613-1368, 614-1368, 615-1368, 616-1368, 617-1368, 618-1368, 619-1368, 620-1368, 621-1368, 622-1368, 623-1368, 624-1368, 625-1368, 626-1368, 627-1368, 628-1368, 629-1368, 630-1368, 631-1368, 632-1368, 633-1368, 634-1368, 635-1368, 636-1368, 637-1368, 638-1368, 639-1368, 640-1368, 641-1368, 642-1368, 643-1368, 644-1368, 645-1368, 646-1368, 647-1368, 648-1368, 649-1368, 650-1368, or 651-1368.In some embodiments, the C-terminal portion of the split Cas9 protein comprises the portion of SpCas9 from amino acid 574 to 1368 or 638 to 1368.

[0155] "Subject" means a human or non-human mammal, such as, but not limited to, a non-human primate (monkey), cow, horse, dog, sheep, or cat. In some embodiments, the subject described herein comprises a pathogenic mutation in the polynucleotide sequence.

[0156] "Substantially identical" means that a polypeptide or nucleic acid molecule exhibits at least 50% identity to a reference amino acid sequence (e.g., any one of the amino acid sequences described herein) or nucleic acid sequence (e.g., any one of the nucleic acid sequences described herein). In one embodiment, such a sequence is at least 60%, 80%, 85%, 90%, 95%, or even 99% identical at the amino acid level or nucleic acid level to the sequence used for comparison.

[0157] Sequence identity is typically measured using sequence analysis software (e.g., the sequence analysis software package from Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, COBALT, EMBOSS Needle, GAP, or the PILEUP / PRETTYBOX program). Such software matches identical or similar sequences by assigning a degree of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine, valine, isoleucine, leucine, aspartic acid, glutamic acid, asparagine, glutamine, serine, threonine, lysine, arginine, and phenylalanine, tyrosine. In an exemplary method of measuring the degree of identity, a BLAST program with a probability score of e-3 to e-100 indicating closely related sequences can be used. COBALT, for example, is used with the following parameters: a) Alignment parameters: Gap penalties -11, -1 and End-Gap penalties -5, -1, b) CDD parameters: Use RPS BLAST on; Blast E-value 0.003; Find Conserved columns and Recompute on, and c) Query clustering parameters: Use query clusters on; Word Size 4; Max cluster distance 0.8; Alphabet Regular. EMBOSS Needle, for example, is used with the following parameters: a) Matrix: BLOSUM62, b) GAP OPEN: 10, c) GAP EXTEND: 0.5, d) OUTPUT FORMAT: pair, e) END GAP PENALTY: false, f) END GAP OPEN: 10, and g) END GAP EXTEND: 0.5.

[0158] The term "target site" refers to a sequence within a nucleic acid molecule that is deaminated by a deaminase (e.g., cytidine deaminase or adenine deaminase) or a fusion protein comprising a deaminase (e.g., a dCas9-adenosine deaminase fusion protein or a base editor disclosed herein).

[0159] As used herein, the terms "treat," "treating," and "treatment," etc., refer to reducing or alleviating a disorder and / or an associated symptom thereof, or obtaining a desired pharmacological and / or physiological effect. It will be understood that treating a disorder or condition does not necessarily require that the disorder, condition, or an associated symptom thereof be completely eliminated, although this is not excluded. In some embodiments, the effect is therapeutic, i.e., the effect, without limitation, partially or completely reduces, decreases, inhibits, alleviates, mitigates, attenuates, or cures a disease and / or a deleterious symptom resulting therefrom. In some embodiments, the effect is prophylactic, i.e., the effect prevents or precludes the onset or recurrence of a disease or condition. For this purpose, the methods disclosed herein include administering a therapeutically effective amount of a composition as described herein.

[0160] "Uracil glycosylase inhibitor" or "UGI" means a drug that inhibits the uracil excision repair system. In one embodiment, the drug is a protein or a fragment thereof that binds to the host's uracil DNA glycosylase and prevents the removal of uracil residues from DNA. In one embodiment, UGI is a protein, a fragment thereof, or a domain that can inhibit the base excision repair enzyme of uracil DNA glycosylase. In some embodiments, the UGI domain includes wild-type UGI or a modified version thereof. In some embodiments, the UGI domain includes a fragment of the exemplary amino acid sequence described below. In some embodiments, the UGI fragment includes an amino acid sequence that includes at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the exemplary UGI sequence shown below. In some embodiments, UGI includes an amino acid sequence that is homologous to the exemplary UGI amino acid sequence or a fragment thereof as described below. In some embodiments, UGI or a portion thereof has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, or 100% identity to the wild-type UGI or UGI sequence, or a portion thereof, as described below. Exemplary UGI includes an amino acid sequence as follows: >splP14739IUNGI_BPPB2 Uracil-DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSD APEYKPWALVIQDSNGENKIKML。

[0161] The ranges provided herein are to be understood as a shorthand for all values within the range. For example, a range of 1 to 50 is to be understood to include any number, combination of numbers, or sub-range from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.

[0162] The recitation of a list of chemical groups in any definition of a variable herein includes the definition of that variable as any single group or combination of the recited groups. The recitation of embodiments for a variable or aspect herein includes that embodiment as any single embodiment or in combination with any other embodiment or portion thereof.

[0163] Any composition or method provided herein can be combined with one or more of any other compositions and methods provided herein.

[0164] The description and examples herein illustrate embodiments of the present disclosure in detail. It is to be understood that the present disclosure is not limited to the specific embodiments described herein and can, therefore, vary in many different ways. Those skilled in the art will recognize that many variations and modifications of the present disclosure exist and that they are within the scope of the present disclosure.

[0165] All terms are intended to be understood as would be understood by one of ordinary skill in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0166] The practice of some embodiments disclosed herein uses conventional techniques in immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, which are within the skill of the art unless otherwise indicated. See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (F.M. Ausubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (M.J. MacPherson, B.D. Hames and G.R. Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (R.I. Freshney, ed. (2010)).

[0167] The various features of the present disclosure may be described in the context of a single embodiment, but they may also be provided separately or in any suitable combination. Conversely, although the present disclosure may be described herein in the context of separate embodiments for clarity, the present disclosure may also be implemented in a single embodiment. The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described.

[0168] The features of the present disclosure are particularly recited in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description, which describes exemplary embodiments that utilize the principles of the present disclosure, and by considering the appended drawings described below.

Brief Description of the Drawings

[0169]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Mode for Carrying Out the Invention

[0170] The present invention features novel adenine base editors (e.g., ABE9) for editing target sequences and methods of using their adenosine deaminase variants.

[0171] Nucleic acid base editor Disclosed herein are novel base editors (e.g., ABE8 and ABE9) or nucleic acid base editors for editing, modifying, or altering a target nucleotide sequence of a polynucleotide. In particular, the novel ABE9 base editor and its component adenosine deaminase are described in Tables 14 and 18 below. Disclosed herein are nucleic acid base editors or base editors comprising a polynucleotide-programmable nucleotide binding domain and a nucleic acid base editing domain (e.g., adenosine deaminase). The polynucleotide-programmable nucleotide binding domain, when combined with a bound guide polynucleotide (e.g., gRNA), specifically binds to a target polynucleotide sequence (i.e., via complementary base pairing between the bases of the bound guide nucleic acid and the bases of the target polynucleotide sequence), thereby localizing the base editor to the target nucleic acid sequence that is desired to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid.

[0172] Polynucleotide-programmable nucleotide binding domain It should be understood that the polynucleotide-programmable nucleotide binding domain can also include a nucleic acid-programmable protein that binds to RNA. For example, the polynucleotide-programmable nucleotide binding domain can be associated with a nucleic acid that guides it to RNA. Other nucleic acid-programmable DNA binding proteins are also within the scope of the present disclosure, but they are not specifically described in the present disclosure.

[0173] The polynucleotide-programmable nucleotide-binding domain of a base editor can itself comprise one or more domains. For example, the polynucleotide-programmable nucleotide-binding domain can comprise one or more nuclease domains. In some embodiments, the nuclease domain of the polynucleotide-programmable nucleotide-binding domain can comprise an endonuclease or an exonuclease. As used herein, the term "exonuclease" refers to a protein or polypeptide capable of digesting a nucleic acid (e.g., RNA or DNA) from a free end, and the term "endonuclease" refers to a protein or polypeptide capable of catalyzing (e.g., cleaving) an internal region of a nucleic acid (e.g., DNA or RNA). In some embodiments, the endonuclease can cleave a single strand of a double-stranded nucleic acid. In some embodiments, the endonuclease can cleave both strands of a double-stranded nucleic acid molecule. In some embodiments, the polynucleotide-programmable nucleotide-binding domain can be a deoxyribonuclease. In some embodiments, the polynucleotide-programmable nucleotide-binding domain can be a ribonuclease.

[0174] In some embodiments, the nuclease domain of a polynucleotide-programmable nucleotide-binding domain can cleave 0, 1, or 2 strands of a target polynucleotide. In some cases, the polynucleotide-programmable nucleotide-binding domain can include a nickase domain. As used herein, the term "nickase" refers to a polynucleotide-programmable nucleotide-binding domain that includes a nuclease domain capable of cleaving only one of the two strands in a double-stranded nucleic acid molecule (e.g., DNA). In some embodiments, the nickase can be derived from a fully catalytically active (e.g., native) form of the polynucleotide-programmable nucleotide-binding domain by introducing one or more mutations into the active polynucleotide-programmable nucleotide-binding domain. For example, if the polynucleotide-programmable nucleotide-binding domain includes a nickase domain derived from Cas9, the Cas9-derived nickase domain can include the D10A mutation and histidine at position 840. In such cases, residue H840 retains catalytic activity and can thereby cleave one strand of the nucleic acid duplex. In another example, the Cas9-derived nickase domain can include the H840A mutation, while the amino acid residue at position 10 remains as D. In some embodiments, the nickase can be derived from a fully catalytically active (e.g., native) form of the polynucleotide-programmable nucleotide-binding domain by removing all or a portion of the nuclease domain that is not required for nickase activity. For example, if the polynucleotide-programmable nucleotide-binding domain includes a nickase domain derived from Cas9, the Cas9-derived nickase domain can include a deletion of all or a portion of the RuvC domain or the HNH domain.

[0175] The amino acid sequence of exemplary catalytically active Cas9 is as follows:

[0176] Also provided herein are base editors that include a catalytically inactive (i.e., unable to cleave a target polynucleotide sequence) polynucleotide-programmable nucleotide binding domain. As used herein, the terms “catalytically inactive” and “nuclease inactive” are used interchangeably to refer to a polynucleotide-programmable nucleotide binding domain having one or more mutations and / or deletions that result in the inability to cleave a nucleic acid strand. In some embodiments, a catalytically inactive polynucleotide-programmable nucleotide binding domain base editor may lack nuclease activity as a result of specific point mutations in one or more nuclease domains. For example, in the case of a base editor that includes a Cas9 domain, Cas9 may include both a D10A mutation and an H840A mutation. Such mutations inactivate both nuclease domains, thereby resulting in a loss of nuclease activity. In other embodiments, a catalytically inactive polynucleotide-programmable nucleotide binding domain may include one or more deletions of all or part of a catalytic domain (e.g., the RuvC1 and / or HNH domains). In further embodiments, a catalytically inactive polynucleotide-programmable nucleotide binding domain includes a point mutation (e.g., D10A or H840A), and a deletion of all or part of a nuclease domain.

[0177] Also, herein, mutations are contemplated that can generate catalytically inactive polynucleotide-programmable nucleotide binding domains from previous functional versions of polynucleotide-programmable nucleotide binding domains. For example, in the case of catalytically inactive Cas9 (“dCas9”), variants are provided that have mutations other than D10A and H840A, resulting in nuclease-inactivated Cas9. Such mutations include, by way of example, other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). Additional suitable nuclease-inactive dCas9 domains may be apparent to those skilled in the art based on the present disclosure and knowledge in the art and are within the scope of the present disclosure. Such additional exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013; 31 (9): 833-838, the entire contents of each of which are incorporated herein by reference).

[0178] Non-limiting examples of polynucleotide-programmable nucleotide-binding domains that can be incorporated into base editors include domains derived from CRISPR proteins, restriction nucleases, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases (ZFNs). In some cases, a base editor includes a polynucleotide-programmable nucleotide-binding domain that includes a native or modified protein or a portion thereof, and can bind to a nucleic acid sequence during CRISPR (i.e., Clustered Regularly Interspaced Short Palindromic Repeats)-mediated nucleic acid modification via a bound guide nucleic acid. Such proteins are referred to herein as "CRISPR proteins." Thus, disclosed herein are base editors that include a polynucleotide-programmable nucleotide-binding domain that includes all or a portion of a CRISPR protein (i.e., a base editor that includes all or a portion of a CRISPR protein as a domain, also referred to as the "CRISPR protein-derived domain" of the base editor). The domain derived from a CRISPR protein incorporated into a base editor can be modified as compared to the wild-type or native version of the CRISPR protein. For example, as described below, the domain derived from a CRISPR protein can include one or more mutations, insertions, deletions, rearrangements, and / or recombinations as compared to the wild-type or native version of the CRISPR protein.

[0179] CRISPR is an adaptive immune system that provides defense against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain spacers, sequences complementary to the aforementioned mobile elements, and target invading nucleic acids. The CRISPR clusters are transcribed and processed into CRISPR RNAs (crRNAs). In type II CRISPR systems, the correct processing of pre-crRNA requires trans-encoded small RNAs (tracrRNAs), the endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA functions as a guide for ribonuclease 3-assisted processing of the pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand that is not complementary to the crRNA is first endonucleolytically cleaved and then trimmed 3'-5' exonucleolytically. In nature, both a protein and both RNAs are usually required for DNA binding and cleavage. However, a single guide RNA ("sgRNA", or simply "gNRA") can be created to incorporate both aspects of crRNA and tracrRNA into a single RNA species. See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J. A., Charpentier E. Science 337: 816-821 (2012) (the entire contents of which are incorporated herein by reference). Cas9 recognizes a short motif (PAM or protospacer adjacent motif) of the CRISPR repeat sequence and helps to distinguish self from non-self.

[0180] In some embodiments, the methods described herein can utilize modified Cas proteins. A guide RNA (gRNA) is a short synthetic RNA consisting of a scaffold sequence necessary for Cas binding and a user-defined approximately 20-nucleotide spacer that defines the genomic (or polynucleotide, e.g., DNA or RNA) target to be modified. Thus, one of ordinary skill in the art can change the genomic or polynucleotide target of the Cas protein by changing the target sequence present in the gRNA. The specificity of the Cas protein is partially determined by how specific the gRNA targeting sequence is for the genomic polynucleotide target sequence compared to other parts of the genome. In one embodiment, the Cas protein is SpCas9.

[0181] In some embodiments, the gRNA scaffold sequence is as follows: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU.

[0182] In some embodiments, the gRNA scaffold sequence is as follows: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGGACCGAGUCGGUGCUUUU.

[0183] In one embodiment, the terminal uracil (U) of the above gRNA scaffold may optionally include "mU*mU*mU*U", which indicates 2'OMe and has phosphorothioate bonds.

[0184] In one embodiment, the RNA scaffold includes a stem-loop. In one embodiment, the RNA scaffold includes the nucleic acid sequence: GUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG.

[0185] In one embodiment, the sgRNA scaffold polynucleotide sequence of S. pyrogenes is as follows: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC。

[0186] In one embodiment, the sgRNA scaffold polynucleotide sequence of S. aureus is as follows: GUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGCGAGA。

[0187] In one embodiment, the sgRNA scaffold of BhCas12b has the following polynucleotide sequence: GUUCUGTCUUUUGGUCAGGACAACCGUCUAGCUAUAAGUGCUGCAGGGUGUGAGAAACUCCUAUUGCUGGACGAUGUCUCUUACGAGGCAUUAGCAC。

[0188] In one embodiment, the BvCas12b sgRNA scaffold has the following polynucleotide sequence: GACCUAUAGGGUCAAUGAAUCUGGGCGUGUGCCAUAAGUAAUUAAAAAUUACCCACCACAGGAGCACCUGAAAACAGGUGCUUGGCAC。

[0189] In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is an endonuclease (e.g., deoxyribonuclease or ribonuclease) that can bind to a target polynucleotide when combined with a bound guide nucleic acid. In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is a nickase that can bind to a target polynucleotide when combined with a bound guide nucleic acid. In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is a catalytically inactive domain that can bind to a target polynucleotide when combined with a bound guide nucleic acid. In some embodiments, the target polynucleotide to which the CRISPR protein-derived domain of the base editor binds is DNA. In some embodiments, the target polynucleotide to which the CRISPR protein-derived domain of the base editor binds is RNA.

[0190] The Cas proteins that can be used in this specification include Class 1 and Class 2. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ, CARF, DinG, their homologs, or their modified or altered versions. An unmodified CRISPR enzyme such as Cas9, which has two functional endonuclease domains, RuvC and HNH, can have DNA cleavage activity. The CRISPR enzyme can induce cleavage of one or both strands at a target sequence, for example, within the target sequence and / or within the complementary strand sequence of the target sequence. For example, the CRISPR enzyme can induce cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence.

[0191] A vector encoding a CRISPR enzyme that has been mutagenized with respect to a corresponding wild-type enzyme can be used such that the mutant CRISPR enzyme lacks the ability to cleave one or both strands of a target polynucleotide containing a target sequence. Cas9 can refer to a polypeptide having at least (or at least about) 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology to an exemplary wild-type Cas9 polypeptide (e.g., Cas9 from S. pyogenes). Cas9 can refer to a polypeptide having at most (or at most about) 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology to an exemplary wild-type Cas9 polypeptide (e.g., from S. pyogenes). Cas9 can refer to a wild-type or modified form of a Cas9 protein that can include amino acid changes such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof.

[0192] In some embodiments, the CRISPR protein-derived domain of the base editor can comprise all or part of Cas9 from Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquis (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref: YP_820832.1); Listeria innocua (NCBI Ref: NP_472073.1); Campylobacter jejuni (NCBI Ref: YP_002344900.1); Neisseria meningitidis (NCBI Ref: YP_002342100.1), Streptococcus pyogenes, or Staphylococcus aureus.

[0193] The Cas9 domain of the nucleic acid base editor The sequences and structures of Cas9 nucleases are well known to those of ordinary skill in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., Proc. Natl. Acad. Sci. U.S.A. 98: 4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E. et al., Nature 471: 602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M. et al., Science 337: 816-821 (2012) (the entire contents of each of which are incorporated herein by reference)). Cas9 orthologs have been described in a variety of species, including but not limited to S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those of ordinary skill in the art based on the present disclosure, and such Cas9 nucleases and sequences include Cas9 sequences derived from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10: 5, 726-737; the entire contents of which are incorporated herein by reference).

[0194] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is a Cas9 domain. Non-limiting exemplary Cas9 domains are provided herein. The Cas9 domain can be a nuclease-active Cas9 domain, a nuclease-inactive Cas9 domain, or a Cas9 nickase. In some embodiments, the Cas9 domain is a nuclease-active domain. For example, the Cas9 domain can be a Cas9 domain that cleaves both strands of a double-stranded nucleic acid (e.g., both strands of a double-stranded DNA molecule). In some embodiments, the Cas9 domain comprises any one of the amino acid sequences described herein. In some embodiments, the Cas9 domain has at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to any one of the amino acid sequences described herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any one of the amino acid sequences described herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical consecutive amino acid residues compared to any one of the amino acid sequences described herein.

[0195] In some embodiments, a protein comprising a fragment of Cas9 is provided. For example, in some embodiments, the protein comprises one of two Cas9 domains: (1) the gRNA-binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, a protein comprising Cas9 or a fragment thereof is referred to as a "Cas9 variant". A Cas9 variant shares homology with Cas9 or a fragment thereof. For example, a Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to wild-type Cas9. In some embodiments, a Cas9 variant can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type Cas9. In some embodiments, a Cas9 variant comprises a fragment of Cas9 (e.g., the gRNA-binding domain or the DNA cleavage domain), such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% the amino acid length of the corresponding wild-type Cas9.In some embodiments, the fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids in length.

[0196] In some embodiments, the Cas9 fusion proteins provided herein include the full-length amino acid sequence of the Cas9 protein, e.g., one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein do not include the full-length Cas9 sequence and include only one or more fragments thereof. Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and additional suitable sequences of Cas9 domains and fragments will be apparent to those of ordinary skill in the art.

[0197] The Cas9 protein can be bound to a guide RNA that guides the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the polynucleotide programmable nucleotide binding domain is a Cas9 domain, such as nuclease active Cas9, Cas9 nickase (nCas9), or nuclease inactive Cas9 (dCas9). Examples of nucleic acid programmable DNA binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ. Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Cas12j / CasΦ, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V Cas effector proteins, type VI Cas effector proteins, CARF, DinG, their homologs, or modified or altered versions thereof.

[0198] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_017053.1, the following nucleotide and amino acid sequences). JPEG0007717684000012.jpg42169JPEG0007717684000013.jpg228157JPEG0007717684000014.jpg239157JPEG0007717684000015.jpg154167

[0199] In some embodiments, wild-type Cas9 corresponds to or comprises the following nucleotide sequence and / or amino acid sequence: JPEG0007717684000016.jpg188167

[0200] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_002737.2 (the following nucleotide sequence); and Uniprot reference sequence: Q99ZW2 (the following amino acid sequence): JPEG0007717684000017.jpg195168

[0201] In some embodiments, Cas9 refers to Cas9 derived from Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquisI (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref: YP_820832.1); Listeria innocua (NCBI Ref: NP_472073.1); Campylobacter jejuni (NCBI Ref: YP_002344900.1); or Neisseria meningitidis (NCBI Ref: YP_002342100.1), or Cas9 derived from any other organism.

[0202] It should be understood that additional Cas9 proteins, including variants and homologs (e.g., nuclease-inactive Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9), are within the scope of the present disclosure. Exemplary Cas9 proteins include, but are not limited to, the proteins shown below. In some embodiments, the Cas9 protein is nuclease-inactive Cas9 (dCas9). In some embodiments, the Cas9 protein is Cas9 nickase (nCas9). In some embodiments, the Cas9 protein is nuclease-active Cas9.

[0203] In some embodiments, the Cas9 domain is a nuclease-inactive Cas9 domain (dCas9). For example, the dCas9 domain can bind to a double-stranded nucleic acid molecule (e.g., via a gRNA molecule) without cleaving either strand of the double-stranded nucleic acid molecule. In some embodiments, the nuclease-inactive dCas9 domain comprises the D10X and H840X mutations of the amino acid sequences described herein, or corresponding mutations in any amino acid sequence provided herein, where X is any amino acid change. In some embodiments, the nuclease-inactive dCas9 domain comprises the D10A and H840A mutations of the amino acid sequences described herein, or corresponding mutations in any amino acid sequence provided herein. As an example, the nuclease-inactive Cas9 domain comprises the amino acid sequence described in the cloning vector pPlatTET-gRNA2 (registration number BAV54124).

[0204] The amino acid sequence of an exemplary catalytically inactive Cas9 (dCas9) is as follows: (See, e.g., Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression.” Cell. 2013; 152(5): 1173-83, the entire content of which is incorporated herein by reference).

[0205] Additional suitable nuclease-inactive dCas9 domains will be apparent to those skilled in the art based on the present disclosure and the knowledge in the art, and are within the scope of the present disclosure. Such additional exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013; 31 (9): 833-838, the entire content of which is incorporated herein by reference).

[0206] In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is a nickase called the "nCas9" protein ("nickase" Cas9). The nuclease-inactivated Cas9 protein can be referred to in the same sense as the "dCas9" protein (nuclease "inactive" Cas9) or catalytically inactive Cas9. Methods for generating Cas9 proteins (or fragments thereof) having an inactive DNA cleavage domain are known (see, e.g., Jinek et al., Science. 337: 816-821 (2012); Qi et al., “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression” (2013) Cell. 28; 152 (5): 1173-83, the entire contents of each of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to include two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al., Science. 337: 816-821 (2012); Qi et al., Cell. 28; 152 (5): 1173-83 (2013)).

[0207] In some embodiments, the dCas9 domain comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the dCas9 domains provided herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any one of the amino acid sequences described herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical contiguous amino acid residues compared to any one of the amino acid sequences described herein.

[0208] In some embodiments, dCas9 corresponds to or partially or wholly comprises a Cas9 amino acid sequence having one or more mutations that inactivate Cas9 nuclease activity. In some embodiments, the nuclease-inactive dCas9 domain comprises the D10X and H840X mutations of the amino acid sequences described herein, or corresponding mutations in any of the amino acid sequences provided herein, where X is any amino acid change. In some embodiments, the nuclease-inactive dCas9 domain comprises the D10A and H840A mutations of the amino acid sequences described herein, or corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the nuclease-inactive Cas9 domain comprises the amino acid sequence described in the cloning vector pPlatTET-gRNA2 (registration number BAV54124).

[0209] In some embodiments, dCas9 comprises the amino acid sequence of dCas9 (D10A and H840A): JPEG0007717684000018.jpg189167

[0210] In some embodiments, the amino acid sequence of an exemplary catalytically inactive Cas9 (dCas9) is as follows: (See, e.g., Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression.” Cell. 2013; 152 (5): 1173-83, the entire contents of which are incorporated herein by reference).

[0211] In some embodiments, the exemplary amino acid sequence of catalytically inactive Cas9 (dCas9) is as follows:

[0212] In some embodiments, the Cas9 domain comprises the D10A mutation, while the residue at position 840 remains a histidine in the amino acid sequence provided above or at the corresponding position of any of the amino acid sequences provided herein.

[0213] In other embodiments, dCas9 variants having mutations other than D10A and H840A are provided, which result in, for example, nuclease-inactivated Cas9 (dCas9). Such mutations include, by way of example, other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). In some embodiments, variants or homologs of dCas9 that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical are provided. In some embodiments, variants of dCas9 having amino acid sequences that are about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids or more shorter or longer are provided.

[0214] Additional suitable nuclease-inactive dCas9 domains will be apparent to those skilled in the art based on the present disclosure and the knowledge in the art and are within the scope of the present disclosure. Such additional exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013; 31 (9): 833-838, the entire content of which is incorporated herein by reference).

[0215] In some embodiments, the Cas9 domain is a Cas9 nickase. The Cas9 nickase can be a Cas9 protein that can cleave only one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 nickase cleaves the target strand of the double-stranded nucleic acid molecule, which means that the Cas9 nickase cleaves the strand that base pairs (is complementary) with the gRNA (e.g., sgRNA) that binds to Cas9. In some embodiments, the Cas9 nickase contains a D10A mutation and has a histidine at position 840. In some embodiments, the Cas9 nickase cleaves the non-target non-base-editing strand of the double-stranded nucleic acid molecule, which means that the Cas9 nickase cleaves the strand that does not base pair with the gRNA (e.g., sgRNA) that binds to Cas9. In some embodiments, the Cas9 nickase contains an H840A mutation and has an aspartic acid residue at position 10, or has a corresponding mutation. In some embodiments, the Cas9 nickase has an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the Cas9 nickases provided herein. Further suitable Cas9 nickases will be apparent to those skilled in the art based on the present disclosure and the knowledge in the art and are within the scope of the present disclosure.

[0216] The amino acid sequence of an exemplary catalytic Cas9 nickase (nCas9) is as follows:

[0217] In some embodiments, Cas9 refers to Cas9 from archaea (e.g., nanoarchaea) that constitute the domain and kingdom of unicellular prokaryotic microorganisms. In some embodiments, the programmable nucleotide binding protein may be, for example, the CasX or CasY protein described in Burstein et al., “New CRISPR-Cas systems from uncultivated microbes.” Cell Res. 2017 Feb 21. doi: 10.1038 / cr.2017.21, the entire content of which is incorporated herein by reference. Using genome-resolved metagenomics, multiple CRISPR-Cas systems have been identified, including Cas9 first reported in the archaea life domain. This diverse set of Cas9 proteins was discovered in nanoarchaea, which are rarely studied as part of active CRISPR-Cas systems. In bacteria, two previously unknown systems, CRISPR-CasX and CRISPR-CasY, have been discovered, and these are one of the most compact systems discovered so far. In some embodiments, in the base editor systems described herein, Cas9 is replaced by CasX, or a variant of CasX. In some embodiments, in the base editor systems described herein, Cas9 is replaced by CasY, or a variant of CasY. Other RNA-guided DNA-binding proteins may be used as nucleic acid programmable DNA-binding proteins (napDNAbp), and it should be understood that they are within the scope of the present disclosure.

[0218] In some embodiments, any nucleic acid programmable DNA binding protein (napDNAbp) of the fusion proteins provided herein can be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the napDNAbp is a CasY protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a native CasX or CasY protein. In some embodiments, the programmable nucleotide binding protein is a native CasX or CasY protein. In some embodiments, the programmable nucleotide binding protein comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any CasX or CasY protein described herein. It should be understood that CasX and CasY from other bacterial species may also be used in accordance with the present disclosure.

[0219] The exemplary CasX ((uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) tr|F0NN87|F0NN87_SULIHCRISPR-related Casx protein OS=Sulfolobus islandicus (HVE10 / 4 strain) GN=SiH_0402 PE=4 SV=1) amino acid sequence is as follows: MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAVGQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG。

[0220] The exemplary CasX (>tr|F0NH53|F0NH53_SULIR CRISPR - associated protein, Casx OS=Sulfolobus islandicus (REY15A strain) GN=SiRe_0771 PE=4 SV=1) amino acid sequence is as follows: MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAVGQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG。

[0221] Deltaproteobacteria CasX MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMTDDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAILQVYWQEFKDDHVGLMCKFAQPASKKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFVYKLEQVSEKGKAYTNYFGRCNVAEHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESTHPVKPLAQIAGNRYASGPVGKALSDACMGTIASFLSKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDfAYNEVIARVRMWVNLNLWQKLKLSRDDAKPLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENPKKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEMDEKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFTDGTDIKKSGKWQGLLYGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKIGRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQAAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAKLAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKELSAELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSYKSGKQPFVGAWQAFYKRRLKEVWKPNA

[0222] The exemplary CasY ((ncbi.nlm.nih.gov / protein / APG80656.1)>APG80656.1 CRISPR - associated protein CasY [uncultured Parcubacteria group bacterium]) amino acid sequence is as follows:

[0223] Cas9 nuclease has two functional endonuclease domains, RuvC and HNH. Cas9 undergoes a conformational change upon target binding, whereby the nuclease domains are positioned to cleave the strands on opposite sides of the target DNA. The final result of DNA cleavage via Cas9 is a double-strand break (DSB) within the target DNA (approximately 3 to 4 nucleotides upstream of the PAM sequence). The resulting DSBs are repaired by either of the following two general repair pathways: (1) the non-homologous end joining (NHEJ) pathway, which is efficient but error-prone; or (2) the homologous recombination repair (HDR) pathway, which is less efficient but highly faithful.

[0224] The "efficiency" of non-homologous end joining (NHEJ) and / or homologous recombination repair (HDR) can be calculated in any convenient way. For example, in some cases, efficiency can be expressed as the percentage of successful HDR. For example, a Surveyor nuclease assay can be used to generate cleavage products, and the ratio of products to substrate can be used to calculate the percentage. For example, a Surveyor nuclease enzyme can be used that directly cleaves DNA containing a newly incorporated restriction enzyme sequence as a result of successful HDR. The more substrate that is cleaved, the higher the percentage of HDR (the higher the efficiency of HDR). As an example for illustration, the percentage of HDR can be calculated using the following formula: [(cleavage products) / (substrate + cleavage products)] (e.g., (b + c) / (a + b + c), where "a" is the band intensity of the DNA substrate and "b" and "c" are the cleavage products).

[0225] In some cases, efficiency can be represented by the percentage of successful NHEJ. For example, the T7 endonuclease I assay can be used to generate cleavage products, and the ratio of products to substrate can be used to calculate the rate of NHEJ. T7 endonuclease I cleaves mismatched heteroduplex DNA resulting from hybridization of wild-type and mutant DNA strands (NHEJ generates small random insertions or deletions (indels) at the original cleavage site). More cleavage indicates a higher rate of NHEJ (higher efficiency of NHEJ). As an example for illustration, the rate (percentage) of NHEJ can be calculated using the following formula: (1 - (1 - (b + c) / (a + b + c)) 1 / 2 ) × 100, where "a" is the band intensity of the DNA substrate, and "b" and "c" are the cleavage products (Ran et. al., Cell. 2013 Sep. 12; 154 (6): 1380-9; and Ran et al., Nat Protoc. 2013 Nov.; 8 (11): 2281-2308).

[0226] The NHEJ repair pathway is the most active repair mechanism and frequently causes small nucleotide insertions or deletions (indels) at the DSB site. Since a population of cells expressing Cas9 and gRNA or guide polynucleotide can generate diverse mutations, the randomness of DSB repair via NHEJ has important practical implications. In most cases, NHEJ causes small indels in the target DNA, resulting in amino acid deletions, insertions, or frameshift mutations, and generating premature stop codons within the open reading frame (ORF) of the target gene. The ideal end result is a loss-of-function mutation within the target gene.

[0227] DSB repair via NHEJ often disrupts the open reading frame of genes, but homologous recombination repair (HDR) can be used to generate specific nucleotide changes ranging from single nucleotide changes to large insertions such as the addition of fluorophores or tags. To utilize HDR for gene editing, a DNA repair template containing the sequence of interest can be delivered to the target cell type along with gRNA(s) and Cas9 or Cas9 nickase. The repair template can include the desired edit as well as additional homologous sequences immediately upstream and downstream of the target (referred to as left and right homology arms). The length of each homology arm can depend on the size of the change to be introduced, with larger insertions requiring longer homology arms. The repair template can be a single-stranded oligonucleotide, a double-stranded oligonucleotide, or a double-stranded DNA plasmid. Also, in cells expressing Cas9, gRNA, and the exogenous repair template, the efficiency of HDR is generally low (less than 10% of the alleles to be modified). Since HDR occurs during the S and G2 phases of the cell cycle, the efficiency of HDR can be increased by synchronizing the cells. Chemically or genetically inhibiting genes involved in NHEJ can also increase the HDR frequency.

[0228] In some embodiments, Cas9 is a modified Cas9. A given gRNA targeting sequence can have additional sites where there is partial homology across the genome. These sites are called off-targets and need to be considered when designing the gRNA. In addition to optimizing gRNA design, the specificity of CRISPR can also be enhanced through modification of Cas9. Cas9 generates double-strand breaks (DSBs) through the combined activity of two nuclease domains, RuvC and HNH. Cas9 nickase, which is a D10A mutant of SpCas9, retains one nuclease domain and generates a DNA nick instead of a DSB. The nickase system can also be combined with gene editing via HDR for specific gene editing.

[0229] In some cases, Cas9 is a variant Cas9 protein. The variant Cas9 polypeptide has an amino acid sequence that differs by one amino acid (e.g., has a deletion, insertion, substitution, fusion) when compared to the amino acid sequence of the wild-type Cas9 protein. In some cases, the variant Cas9 polypeptide has an amino acid change (e.g., a deletion, insertion, or substitution) that reduces the nuclease activity of the Cas9 polypeptide. For example, in some cases, the variant Cas9 polypeptide has less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nuclease activity of the corresponding wild-type Cas9 protein. In some cases, the variant Cas9 protein has substantially no nuclease activity. If the subject Cas9 protein is a variant Cas9 protein that has substantially no nuclease activity, it can be referred to as "dCas9".

[0230] In some cases, the variant Cas9 protein has low nuclease activity. For example, the variant Cas9 protein exhibits less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% of the endonuclease activity of the wild-type Cas9 protein (e.g., wild-type Cas9 protein).

[0231] In some cases, variant Cas9 proteins can cleave the complementary strand of the guide-target sequence, but have a reduced ability to cleave the non-complementary strand of the double-stranded guide-target sequence. For example, a variant Cas9 protein can have a mutation (amino acid substitution) that reduces the function of the RuvC domain. As a non-limiting example, in some embodiments, the variant Cas9 protein has D10A (from aspartic acid to alanine at amino acid position 10), and thus can cleave the complementary strand of the double-stranded guide-target sequence, but has a reduced ability to cleave the non-complementary strand of the double-stranded guide-target sequence (thus, when the variant Cas9 protein cleaves a double-stranded target nucleic acid, it results in a single-strand break (SSB) instead of a double-strand break (DSB)) (see, for example, Jinek et al., Science. 2012 Aug. 17; 337 (6096): 816-21).

[0232] In some cases, variant Cas9 proteins can cleave the non-complementary strand of the double-stranded guide-target sequence, but have a reduced ability to cleave the complementary strand of the guide-target sequence. For example, a variant Cas9 protein can have a mutation (amino acid substitution) that reduces the function of the HNH domain (RuvC / HNH / RuvC domain motif). As a non-limiting example, in some embodiments, the variant Cas9 protein has the H840A (from histidine to alanine at amino acid position 840) mutation, and thus can cleave the non-complementary strand of the guide-target sequence, but has a reduced ability to cleave the complementary strand of the guide-target sequence (thus, when the variant Cas9 protein cleaves the double-stranded guide-target sequence, it results in an SSB rather than a DSB). Such Cas9 proteins have a reduced ability to cleave the guide-target sequence (e.g., a single-stranded guide-target sequence), but retain the ability to bind to the guide-target sequence (e.g., a single-stranded guide-target sequence).

[0233] In some cases, the variant Cas9 protein has a reduced ability to cleave both the complementary and non-complementary strands of double-stranded target DNA. As a non-limiting example, in some cases, the variant Cas9 protein contains both the D10A and H840A mutations, whereby this polypeptide has a reduced ability to cleave both the complementary and non-complementary strands of double-stranded target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA).

[0234] As another non-limiting example, in some cases, the variant Cas9 protein has the W476A and W1126A mutations, whereby this polypeptide has a reduced ability to cleave target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA).

[0235] As another non-limiting example, in some cases, the variant Cas9 protein has the P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, whereby this polypeptide has a reduced ability to cleave target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA).

[0236] As another non-limiting example, in some cases, the variant Cas9 protein has H840A, W476A, and W1126A mutations, whereby this polypeptide has a reduced ability to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some cases, the variant Cas9 protein has H840A, D10A, W476A, and W1126A mutations, whereby this polypeptide has a reduced ability to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, the variant Cas9 has restored the catalytic His residue at position 840 of the Cas9 HNH domain (A840H).

[0237] As another non-limiting example, in some cases, a variant Cas9 protein has the H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, whereby this polypeptide has a reduced ability to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some cases, a variant Cas9 protein has the D10A, H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, whereby this polypeptide has a reduced ability to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). In some cases, when the variant Cas9 protein contains the W476A and W1126A mutations, or when the variant Cas9 protein contains the P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, the variant Cas9 protein does not efficiently bind to the PAM sequence. Thus, in some such cases, when using such a variant Cas9 protein in a binding method, the method does not require a PAM sequence. In other words, in some cases, when using such a variant Cas9 protein in a binding method, the method can include a guide RNA, but the method can be carried out without a PAM sequence (and the specificity of binding is thus provided by the targeting segment of the guide RNA). Other residues can be mutated (i.e., inactivated one or the other nuclease moiety) to achieve the above effects. As non-limiting examples, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 can be modified (i.e., substituted). Also, mutations other than alanine substitutions are suitable.

[0238] In some embodiments, a variant Cas9 protein with reduced catalytic activity (e.g., when the Cas9 protein has D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 mutations, such as D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A) can still bind to target DNA in a site-specific manner as long as it retains the ability to interact with the guide RNA (since it is guided to the target DNA sequence by the guide RNA).

[0239] In some embodiments, the variant Cas protein can be spCas9, spCas9-VRQR, spCas9-VRER, xCas9(sp), saCas9, saCas9-KKH, SpCas9-MQKFRAER, spCas9-MQKSER, spCas9-LRKIQK, or spCas9-LRVSQL.

[0240] In some embodiments, a modified SpCas9 containing the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (SpCas9-MQKFRAER) and having specificity for the modified PAM 5'-NGC-3' is used.

[0241] Alternatives to S. pyogenes Cas9 may include RNA-guided endonucleases of the Cpf1 family that exhibit cleavage activity in mammalian cells. CRISPR from Prevotella and Francisella 1 (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. Cpf1 is an RNA-guided endonuclease of the class II CRISPR / Cas system. This acquired immune mechanism is observed in Prevotella and Francisella bacteria. The Cpf1 gene is associated with the CRISPR locus and encodes an endonuclease that uses guide RNA to find and cleave viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9 and overcomes some of the limitations of the CRISPR / Cas9 system. Unlike the Cas9 nuclease, the result of DNA cleavage via Cpf1 is a double-strand break with short 3' overhangs. The staggered cleavage pattern of Cpf1 can enable directional gene insertion, similar to conventional restriction enzyme cloning, and can enhance the efficiency of gene editing. Similar to the above-mentioned Cas9 variants and orthologs, Cpf1 can also expand the number of sites that can be targeted by CRISPR to AT-rich regions or AT-rich genomes lacking the NGG PAM site preferred by SpCas9. The Cpf1 locus contains a mixed α / β domain, followed by the RuvC-I helical region, RuvC-II, and a zinc finger-like domain. The Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9. Furthermore, Cpf1 does not have an HNH endonuclease domain and lacks the α-helical recognition lobe of Cas9 at its N-terminus. The Cpf1 CRISPR-Cas domain architecture indicates that Cpf1 is functionally unique and is classified as a class 2, type V CRISPR system. The Cpf1 locus encodes Cas1, Cas2, and Cas4 proteins that are more similar to type I and type III than type II systems.Functional Cpf1 does not require a trans-activating CRISPR RNA (tracrRNA), and thus only requires the CRISPR (crRNA). This is effective for genome editing because Cpf1 is not only smaller than Cas9, but also has a smaller sgRNA molecule (about half the nucleotides of Cas9). The Cpf1-crRNA complex cleaves target DNA or RNA by identifying the protospacer adjacent motif 5'-YTN-3', as opposed to the G-rich PAM targeted by Cas9. After identification of the PAM, Cpf1 introduces a sticky-end-like DNA double-strand break with a 4- or 5-nucleotide overhang.

[0242] In some embodiments, Cas9 is a Cas9 variant having specificity for a modified PAM sequence. In some embodiments, additional Cas9 variants and PAM sequences are described in Miller, S. M. et al. Continuous evolution of SpCas9 variants compatible with non-G PAMs, Nat. Biotechnol. (2020), which is hereby incorporated by reference in its entirety. In some embodiments, the Cas9 variant does not have specific PAM requirements. In some embodiments, the Cas9 variant, e.g., the SpCas9 variant, has specificity for an NRNH PAM (wherein in the sequence, R is A or G, and H is A, C, or T). In some embodiments, the SpCas9 variant has specificity for the PAM sequences AAA, TAA, CAA, GAA, TAT, GAT, or CAC. In some embodiments, the SpCas9 variant has amino acid substitutions at positions 1114, 1134, 1135, 1137, 1339, 1151, 1180, 1188, 1211, 1218, 1219, 1221, 1249, 1256, 1264, 1290, 1318, 1317, 1320, 1321, 1323, 1332, 1333, 1335, 1337, or 1339 or corresponding positions thereto. In some embodiments, the SpCas9 variant comprises amino acid substitutions at positions 1114, 1135, 1218, 1219, 1221, 1249, 1320, 1321, 1323, 1332, 1333, 1335, or 1337 or corresponding positions thereto. In some embodiments, the SpCas9 variant has amino acid substitutions at positions 1114, 1134, 1135, 1137, 1339, 1151, 1180, 1188, 1211, 1219, 1221, 1256, 1264, 1290, 1318, 1317, 1320, 1323, 1333 or corresponding positions thereto.In some embodiments, the SpCas9 variant has an amino acid substitution at position 1114, 1131, 1135, 1150, 1156, 1180, 1191, 1218, 1219, 1221, 1227, 1249, 1253, 1286, 1293, 1320, 1321, 1332, 1335, 1339 or the corresponding position. In some embodiments, the SpCas9 variant comprises an amino acid substitution at position 1114, 1227, 1135, 1180, 1207, 1219, 1234, 1286, 1301, 1332, 1335, 1337, 1338, 1349 or the corresponding position. Exemplary amino acid substitutions and PAM specificities of the SpCas9 variant are shown in Tables 1A - 1D.

Table 1-1

Table 1-2

Table 1-3

Table 1-4

[0243] In some embodiments, Cas9 is Neisseria meningitidis Cas9 (NmeCas9) or a variant thereof. In some embodiments, NmeCas9 has specificity for the NNNNGAYW PAM, where in the sequence, Y is C or T and W is A or T. In some embodiments, NmeCas9 has specificity for the NNNNGYTT PAM, where in the sequence, Y is C or T. In some embodiments, NmeCas9 has specificity for the NNNNGTCT PAM. In some embodiments, NmeCas9 is Nme1 Cas9. In some embodiments, NmeCas9 has specificity for the NNNNGATT PAM, NNNNCCTA PAM, NNNNCCTC PAM, NNNNCCTT PAM, NNNNCCTG PAM, NNNNCCGT PAM, NNNNCCGGPAM, NNNNCCCA PAM, NNNNCCCT PAM, NNNNCCCC PAM, NNNNCCAT PAM, NNNNCCAG PAM, NNNNCCAT PAM, or NNNGATT PAM. In some embodiments, Nme1Cas9 has specificity for the NNNNGATT PAM, NNNNCCTA PAM, NNNNCCTC PAM, NNNNCCTT PAM, or NNNNCCTG PAM. In some embodiments, NmeCas9 has specificity for the CAA PAM, CAAA PAM, or CCA PAM. In some embodiments, NmeCas9 is Nme2 Cas9. In some embodiments, NmeCas9 has specificity for the NNNNCC(N4CC)PAM, where in the sequence, N is any one of A, G, C, or T. In some embodiments, NmeCas9 has specificity for the NNNNCCGT PAM, NNNNCCGGPAM, NNNNCCCA PAM, NNNNCCCT PAM, NNNNCCCC PAM, NNNNCCAT PAM, NNNNCCAG PAM, NNNNCCAT PAM, or NNNGATT PAM. In some embodiments, NmeCas9 is Nme3Cas9.In some embodiments, NmeCas9 is specific for the NNNNCAAA PAM, NNNNCC PAM, or NNNNCNNN PAM. Additional NmeCas9 properties and PAM sequences described in Edraki et al. Mol. Cell. (2019) 73 (4): 714-726 are hereby incorporated by reference in their entirety. An exemplary amino acid sequence of Nme1Cas9 is shown below: Type II CRISPR RNA-guided endonuclease Cas9 [Neisseria meningitidis] WP_002235162.1 JPEG0007717684000023.jpg131167

[0244] An exemplary amino acid sequence of Nme2Cas9 is shown below: Type II CRISPR RNA-guided endonuclease Cas9 [Neisseria meningitidis] WP_002230835.1 JPEG0007717684000024.jpg130170

[0245] The Cas12 domain of the nucleobase editor Generally, the microbial CRISPR-Cas systems are divided into Class 1 systems and Class 2 systems. Class 1 systems have multi-subunit effector complexes, and Class 2 systems have a single protein effector. For example, Cas9 and Cpf1 are Class 2 effectors, although they are of different types (Type II and Type V, respectively). In addition to Cpf1, the Class 2 Type V CRISPR-Cas systems also include Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ. See, for example, Shmakov et al., “Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems,” Mol. Cell, 2015 Nov. 5; 60 (3): 385-397; Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR Journal, 2018, 1 (5): 325-336; and Yan et al., “Functionally Diverse Type V CRISPR-Cas Systems, “Science, 2019 Jan. 4; 363: 88-91 (the entire contents of each are incorporated herein by reference). Type V Cas proteins have an RuvC (or RuvC-like) endonuclease domain. The generation of mature CRISPR RNA (crRNA) is generally tracrRNA-independent, although, for example, Cas12b / C2c1 requires tracrRNA for crRNA generation. Cas12b / C2c1 depends on both crRNA and tracrRNA for DNA cleavage.

[0246] Nucleic acid programmable DNA binding proteins contemplated by the present invention include Cas proteins (Cas12 proteins) classified as Class 2, Type V. Non-limiting examples of Cas Class 2, Type V proteins include Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ homologs, or modified versions thereof. As used herein, a Cas12 protein may be referred to as a Cas12 nuclease, Cas12 domain, or Cas12 protein domain. In some embodiments, the Cas12 proteins of the present invention include amino acid sequences interrupted by internally fused protein domains such as a deaminase domain.

[0247] In some embodiments, the Cas12 domain is a nuclease-inactive Cas12 domain or a Cas12 nickase. In some embodiments, the Cas12 domain is a nuclease-active domain. For example, the Cas12 domain can be a Cas12 domain that nicks one strand of a double-stranded nucleic acid (e.g., a double-stranded DNA molecule). In some embodiments, the Cas12 domain comprises any one of the amino acid sequences described herein. In some embodiments, the Cas12 domain has at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to any one of the amino acid sequences described herein. In some embodiments, the Cas12 domain comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any one of the amino acid sequences described herein. In some embodiments, the Cas12 domain comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical consecutive amino acid residues compared to any one of the amino acid sequences described herein.

[0248] In some embodiments, a protein comprising a fragment of Cas12 is provided. For example, in some embodiments, the protein comprises one of two Cas12 domains: (1) the gRNA binding domain of Cas12; or (2) the DNA cleavage domain of Cas12. In some embodiments, a protein comprising Cas12 or a fragment thereof is referred to as a "Cas12 variant". A Cas12 variant shares homology with Cas12 or a fragment thereof. For example, a Cas12 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to wild-type Cas12. In some embodiments, a Cas12 variant can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type Cas12. In some embodiments, a Cas12 variant comprises a fragment of Cas12 (e.g., the gRNA binding domain or the DNA cleavage domain), such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild-type Cas12.In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid length of the corresponding wild-type Cas12. In some embodiments, the fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids in length.

[0249] In some embodiments, Cas12 corresponds to, or comprises a part or all of, a Cas12 amino acid sequence having one or more mutations that alter Cas12 nuclease activity. Such mutations include, for example, amino acid substitutions within the RuvC nuclease domain of Cas12. In some embodiments, a variant or homolog of Cas12 is provided that is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to wild-type Cas12. In some embodiments, variants of Cas12 are provided having amino acid sequences that are about 5, about 10, about 15, about 20, about 25, about 30, about 40, about 50, about 75, about 100 amino acids or more shorter or longer.

[0250] In some embodiments, the Cas12 fusion proteins provided herein include the full-length amino acid sequence of a Cas12 protein, e.g., one of the Cas12 sequences provided herein. However, in other embodiments, the fusion proteins provided herein do not include the full-length Cas12 sequence and include only one or more fragments thereof. Exemplary amino acid sequences of suitable Cas12 domains are provided herein, and additional suitable sequences of Cas12 domains and fragments will be apparent to those skilled in the art.

[0251] Generally, class 2 type V Cas proteins have a single functional RuvC endonuclease domain (see, e.g., Chen et al., “CRISPR-Cas12a target binding unleashes indiscriminate single-stranded DNase activity,” Science 360: 436-439 (2018)). In some cases, the Cas12 protein is a variant Cas12b protein (see Strecker et al., Nature Communications, 2019, 10 (1): Art. No.: 212). In one embodiment, the variant Cas12 polypeptide has an amino acid sequence that differs by 1, 2, 3, 4, 5, or more amino acids (e.g., has deletions, insertions, substitutions, fusions) when compared to the amino acid sequence of the wild-type Cas12 protein. In some examples, the variant Cas12 polypeptide has an amino acid change (e.g., deletion, insertion, or substitution) that reduces the activity of the Cas12 polypeptide. For example, in some examples, the variant Cas12 is a Cas12b polypeptide having less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nickase activity of the corresponding wild-type Cas12b protein. In some cases, the variant Cas12b protein has substantially no nickase activity.

[0252] In some cases, the variant Cas12b protein has reduced nickase activity. For example, the variant Cas12b protein exhibits less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% of the nickase activity of the wild-type Cas12b protein.

[0253] In some embodiments, the Cas12 protein comprises an RNA-guided endonuclease from the Cas12a / Cpf1 family that is active in mammalian cells. CRISPR from Prevotella and Francisella 1 (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. Cpf1 is an RNA-guided endonuclease of the class II CRISPR / Cas system. This acquired immune mechanism is found in the bacteria Prevotella and Francisella 1. The Cpf1 gene is associated with the CRISPR locus and encodes an endonuclease that uses guide RNA to find and cut viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9 and overcomes some of the limitations of the CRISPR / Cas9 system. Unlike the Cas9 nuclease, the result of DNA cleavage via Cpf1 is a double-strand break with short 3' overhangs. The sticky-end type cleavage pattern of Cpf1 can expand the possibility of directional gene transfer and enhance the efficiency of gene editing, similar to conventional restriction enzyme cloning. Similar to the above-described Cas9 variants and orthologs, Cpf1 can also expand the number of sites that can be targeted by CRISPR to AT-rich regions or AT-rich genomes lacking the NGG PAM site preferred by SpCas9. The Cpf1 locus contains a mixed α / β domain, a RuvC-I followed by a helical region, a RuvC-II, and a zinc finger-like domain. The Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9. Furthermore, unlike Cas9, Cpf1 does not have an HNH endonuclease domain and does not have the α-helical recognition lobe of Cas9 at its N-terminus. The Cpf1 CRISPR-Cas domain architecture indicates that Cpf1 is functionally unique and is classified as a class 2, type V CRISPR system. The Cpf1 locus encodes Cas1, Cas2, and Cas4 proteins and is more similar to type I and type III than to type II systems.Functional Cpf1 does not require a trans-activating CRISPR RNA (tracrRNA) and thus only requires a CRISPR (crRNA). This is useful for genome editing because Cpf1 is not only smaller than Cas9, but also the sgRNA molecule is smaller (about half the number of nucleotides of Cas9). The Cpf1-crRNA complex cleaves target DNA or RNA by recognizing the protospacer adjacent motif 5'-YTN-3' or 5'-TTTN-3', as opposed to the G-rich PAM targeted by Cas9. After PAM recognition, Cpf1 introduces a sticky-end-like DNA double-strand break with a 4- or 5-nucleotide overhang.

[0254] In some aspects of the invention, a vector encoding a CRISPR enzyme that has been mutagenized relative to the corresponding wild-type enzyme (such that the mutagenized CRISPR enzyme lacks the ability to cleave one or both strands of a target polynucleotide containing the target sequence) can be used. Cas12 can refer to a polypeptide having at least (or at least about) 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology to an exemplary wild-type Cas12 polypeptide (e.g., Cas12 from Bacillus hisashii). Cas12 can refer to a polypeptide having at most or at most about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology to an exemplary wild-type Cas12 polypeptide (e.g., from Bacillus hisashii (BhCas12b), Bacillus species V3-13 (BvCas12b), and Alicyclobacillus acidiphilus (AaCas12b)). Cas12 can refer to a wild-type or modified form of a Cas12 protein that can include amino acid changes such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof.

[0255] In some embodiments, the BhCas12b guide polynucleotide has the following sequence: BhCas12b sgRNA scaffold (underlined) + 20 nt to 23 nt guide sequence (N n as shown) JPEG0007717684000025.jpg20170

[0256] In some embodiments, the BvCas12b and AaCas12b guide polynucleotides have the following sequences: BvCas12b sgRNA scaffold (underlined) + 20 nt to 23 nt guide sequence (N n as shown) JPEG0007717684000026.jpg19170AaCas12b sgRNA scaffold (underlined) + 20 nt to 23 nt guide sequence (N n as shown) JPEG0007717684000027.jpg22169

[0257] Nucleic acid programmable DNA binding protein Some aspects of the present disclosure provide fusion proteins that include domains that act as nucleic acid programmable DNA binding proteins, which can be used to guide proteins such as base editors to specific nucleic acid (e.g., DNA or RNA) sequences. In certain embodiments, the fusion protein includes a nucleic acid programmable DNA binding protein domain and a deaminase domain. Non-limiting examples of nucleic acid programmable DNA binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ. Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Cas12j / CasΦ, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V Cas effector proteins, type VI Cas effector proteins, CARF, DinG, their homologs, or modified or altered versions thereof. Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, although they may not be specifically described herein.See, for example, Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct; 1: 325-336. doi: 10.1089 / crispr.2018.0033; Yan et al., “Functionally diverse type V CRISPR-Cas systems” Science. 2019 Jan 4; 363 (6422): 88-91. doi: 10.1126 / science.aav7271 (the entire contents of each are incorporated herein by reference).

[0258] An example of a nucleic acid programmable DNA binding protein having a PAM specificity different from Cas9 is Clustered Regularly Interspaced Short Palindromic Repeat from Prevotella and Francisella 1 (Cpf1). Similar to Cas9, Cpf1 is also a class 2 CRISPR effector. Cpf1 has been shown to mediate robust DNA interference with functions different from Cas9. Cpf1 is a single RNA-guided endonuclease lacking tracrRNA and utilizes a T-rich protospacer adjacent motif (TTN, TTTN, or YTN). Furthermore, Cpf1 cleaves DNA via staggered DNA double-strand breaks. Of the 16 Cpf1 family proteins, two enzymes from Acidaminococcus and Lachnospiraceae have been shown to have efficient genome editing activity in human cells. Cpf1 proteins are known in the art and have been previously described, for example, in Yamano et al., “Crystal structure of Cpf1 in complex with guide RNA and target DNA.” Cell (165) 2016, p. 949-962, the entire contents of which are incorporated herein by reference.

[0259] Nuclease-inactive Cpf1 (dCpf1) variants that can be used as DNA-binding protein domains programmable with guide nucleotide sequences are useful in the present compositions and methods. The Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9, but lacks an HNH endonuclease domain and does not have the α-helical recognition lobe of Cas9 at its N-terminus. Zetsche et al., Cell, 163, 759-771, 2015 (incorporated herein by reference) showed that the RuvC-like domain of Cpf1 is involved in cleavage of both DNA strands and that inactivation of the RuvC-like domain inactivates Cpf1 nuclease activity. For example, mutations corresponding to D917A, E1006A, or D1255A of Francisella novicida Cpf1 inactivate Cpf1 nuclease activity. In some embodiments, the dCpf1 of the present disclosure comprises mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A. It should be understood that any mutation that inactivates the RuvC domain of Cpf1, such as a substitution mutation, deletion, or insertion, can be used in accordance with the present disclosure.

[0260] In some embodiments, any of the nucleic acid programmable DNA binding proteins (napDNAbps) of the fusion proteins provided herein can be a Cpf1 protein. In some embodiments, the Cpf1 protein is a Cpf1 nickase (nCpf1). In some embodiments, the Cpf1 protein is a nuclease-inactive Cpf1 (dCpf1). In some embodiments, the Cpf1, nCpf1, or dCpf1 has an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the Cpf1 sequences disclosed herein. In some embodiments, the dCpf1 has an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the Cpf1 sequences disclosed herein and includes mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A. It should be understood that Cpf1s from other bacterial species may also be used in accordance with the present disclosure.

[0261] Wild-type Francisella novicida Cpf1 (D917, E1006, and D1255 are in bold and underlined) JPEG0007717684000028.jpg162169

[0262] Francisella novicida Cpf1 D917A (A917, E1006, and D1255 are in bold and underlined) JPEG0007717684000029.jpg164170

[0263] Francisella novicida Cpf1 E1006A (D917, A1006, and D1255 are in bold and underlined) JPEG0007717684000030.jpg164169

[0264] Francisella novicida Cpf1 D1255A (D917, E1006, and A1255 are in bold and underlined) JPEG0007717684000031.jpg167170

[0265] Francisella novicida Cpf1 D917A / E1006A (A917, A1006, and D1255 are in bold and underlined) JPEG0007717684000032.jpg167167

[0266] Francisella novicida Cpf1 D917A / D1255A (A917, E1006, and A1255 are in bold and underlined) JPEG0007717684000033.jpg162169

[0267] Francisella novicida Cpf1 E1006A / D1225A (D917, A1006, and A1255 are in bold and underlined) JPEG0007717684000034.jpg167170

[0268] Francisella novicida Cpf1 D917A / E1006A / D1255A (A917, A1006, and A1255 are in bold and underlined) JPEG0007717684000035.jpg164169

[0269] In some embodiments, one of the Cas9 domains present in the fusion protein may be replaced with a DNA-binding protein domain programmable with a guide nucleotide sequence that does not require a PAM sequence.

[0270] In some embodiments, the Cas9 domain is a Cas9 domain derived from Staphylococcus aureus (SaCas9). In some embodiments, the SaCas9 domain is nuclease-active SaCas9, nuclease-inactive SaCas9 (SaCas9d), or SaCas9 nickase (SaCas9n). In some embodiments, SaCas9 includes the N579A mutation, or a corresponding mutation in any of the amino acid sequences provided herein.

[0271] In some embodiments, the SaCas9 domain, the SaCas9d domain, or the SaCas9n domain can bind to a nucleic acid sequence having a non-standard PAM. In some embodiments, the SaCas9 domain, the SaCas9d domain, or the SaCas9n domain can bind to a nucleic acid sequence having an NNGRRT or NNGRRT PAM sequence. In some embodiments, the SaCas9 domain includes one or more of the E781X, N967X, and R1014X mutations, or a corresponding mutation in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SaCas9 domain includes one or more of the E781K, N967K, and R1014H mutations, or one or more corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SaCas9 domain includes the E781K, N967K, or R1014H mutation, or a corresponding mutation in any of the amino acid sequences provided herein.

[0272] Exemplary SaCas9 sequences Mutating the above residue N579 that is underlined and shown in bold (e.g., to A579) may result in a SaCas9 nickase.

[0273] Exemplary SaCas9n sequences JPEG0007717684000037.jpg131169 The above residue A579 from which mutation of residue N579 can result in a SaCas9 nickase is underlined and shown in bold.

[0274] Exemplary SaKKH Cas9 JPEG0007717684000038.jpg121159 The above residue A579 from which mutation of residue N579 can result in a SaCas9 nickase is underlined and shown in bold. The above residues K781, K967, and H1014 from which mutation of E781, N967, and R1014 can result in a SaKKH Cas9 are underlined and shown in italics.

[0275] In some embodiments, the napDNAbp is a circular permutant. In the following sequences, the plain text shows the adenosine deaminase sequence, the bold sequence shows the sequence derived from Cas9, the italic sequence shows the linker sequence, the underlined sequence shows the bipartite nuclear localization sequence, and the double-underlined sequence shows the mutation. CP5 (having MSP “NGC” PID and “D10A” nickase): JPEG0007717684000039.jpg180169

[0276] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is a single effector of the microbial CRISPR-Cas system. Single effectors of the microbial CRISPR-Cas system include, but are not limited to, Cas9, Cpf1, Cas12b / C2c1, and Cas12c / C2c3. Typically, the microbial CRISPR-Cas system is divided into class 1 systems and class 2 systems. Class 1 systems have multi-subunit effector complexes, and class 2 systems have a single protein effector. For example, Cas9 and Cpf1 are class 2 effectors. In addition to Cas9 and Cpf1, three different class 2 CRISPR-Cas systems (Cas12b / C2c1 and Cas12c / C2c3) are described in Shmakov et al., “Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems”, Mol. Cell, 2015 Nov. 5; 60 (3): 385-397 (the entire content of which is incorporated herein by reference). Two effectors of the system, Cas12b / C2c1 and Cas12c / C2c3, contain a RuvC-like endonuclease domain related to Cpf1. The third system includes an effector having two predicted HEPN RNase domains. The generation of mature CRISPR RNA is independent of tracrRNA, unlike the generation of CRISPR RNA by Cas12b / C2c1. Cas12b / C2c1 depends on both CRISPR RNA and tracrRNA for DNA cleavage.

[0277] The crystal structure of Alicyclobaccillus acidoterrastris Cas12b / C2c1 (AacC2c1) complexed with a chimeric single-guide RNA (sgRNA) has been reported. See, e.g., Liu et al., “C2c1-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism”, Mol. Cell, 2017 Jan. 19; 65 (2): 310-322 (the entire content of which is incorporated herein by reference). The crystal structure has also been reported for Alicyclobacillus acidoterrestris C2c1 bound to target DNA as a ternary complex. See, e.g., Yang et al., “PAM-dependent Target DNA Recognition and Cleavage by C2C1 CRISPR-Cas endonuclease”, Cell, 2016 Dec. 15; 167 (7): 1814-1828 (the entire content of which is incorporated herein by reference). Catalytically competent conformations of AacC2c1 have been captured, independently positioned within a single RuvC catalytic pocket, on both the target and non-target DNA strands, where cleavage via Cas12b / C2c1 results in a staggered cleavage of the target DNA by 7 nucleotides. Structural comparisons of the Cas12b / C2c1 ternary complex with previously identified counterparts of Cas9 and Cpf1 show the diversity of the mechanisms used in the CRISPR-Cas9 system.

[0278] In some embodiments, any of the nucleic acid programmable DNA binding proteins (napDNAbps) of the fusion proteins provided herein can be a Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a Cas12b / C2c1 protein. In some embodiments, the napDNAbp is a Cas12c / C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a native Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a native Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the napDNAbp sequences provided herein. It should be understood that Cas12b / C2c1 or Cas12c / C2c3 from other bacterial species may also be used in accordance with the present disclosure.

[0279] The amino acid sequence of Cas12b / C2c1 ((uniprot.org / uniprot / T0D7A2#2) sp|T0D7A2|C2C1_ALIAG CRISPR-associated endonuclease C2c1 OS=Alicyclobacillus acido-terrestris (strain ATCC49025 / DSM3922 / CIP106132 / NCIMB13137 / GD3B) GN=c2c1 PE=1 SV=1) is as follows: AacCas12b (Alicyclobacillus acidiphilus)-WP_067623834 BhCas12b (Bacillus hisashii) NCBI reference sequence: WP_095142515

[0280] In some embodiments, Cas12b is a variant of BhCas12b and is BvCas12b (V4) including the following changes compared to BhCas12b: S893R, K846R, and E837G. BhCas12b (V4) is represented as follows: 5’mRNA Cap---5’UTR---bhCas12b---STOP sequence---3’UTR---120 polyA tail. 5’UTR: GGGAAATAAGAGAGAAAAGAAGAGTAAGAAGAAATATAAGAGCCACC 3’UTR (TriLink standard UTR) GCTGGAGCCTCGGTGGCCATGCTTCTTGCCCCTTGGGCCTCCCCCCAGCCCCTCCTCCCCTTCCTGCACCCGTACCCCCGTGGTCTTTGAATAAAGTCTGA Nucleic acid sequence of bhCas12b (V4)

[0281] In some embodiments, Cas12b is BvCas12B. In some embodiments, Cas12b comprises the amino acid substitutions S893R, K846R, and E837G numbered in the exemplary sequences of BvCas12b shown below. BvCas12b (Bacillus species V3-13) NCBI reference sequence: WP_101661451.1

[0282] In some embodiments, Cas12b is BTCas12b. BTCas12b (Bacillus thermoamylovorans) NCBI reference sequence: WP_041902512 MATRSFILKIEPNEEVKKGLWKTHEVLNHGIAYYMNILKLIRQEAIYEHHEQDPKNPKKV SKAEIQAELWDFVLKMQKCNSFTHEVDKDVVFNILRELYEELVPSSVEKKGEANQLSNKF LYPLVDPNSQSGKGTASSGRKPRWYNLKIAGDPSWEEEKKKWEEDKKKDPLAKILGKLAE YGLIPLFIPFTDSNEPIVKEIKWMEKSRNQSVRRLDKDMFIQALERFLSWESWNLKVKEE YEKVEKEHKTLEERIKEDIQAFKSLEQYEKERQEQLLRDTLNTNEYRLSKRGLRGWREII QKWLKMDENEPSEKYLEVFKDYQRKHPREAGDYSVYEFLSKKENHFIWRNHPEYPYLYAT FCEIDKKKKDAKQQATFTLADPINHPLWVRFEERSGSNLNKYRILTEQLHTEKLKKKLTV QLDRLIYPTESGGWEEKGKVDIVLLPSRQFYNQIFLDIEEKGKHAFTYKDESIKFPLKGT LGGARVQFDRDHLRRYPHKVESGNVGRIYFNMTVNIEPTESPVSKSLKIHRDDFPKFVNF KPKELTEWIKDSKGKKLKSGIESLEIGLRVMSIDLGQRQAAAASIFEVVDQKPDIEGKLF FPIKGTELYAVHRASFNIKLPGETLVKSREVLRKAREDNLKLMNQKLNFLRNVLHFQQFE DITEREKRVTKWISRQENSDVPLVYQDELIQIRELMYKPYKDWVAFLKQLHKRLEVEIGK EVKHWRKSLSDGRKGLYGISLKNIDEIDRTRKFLLRWSLRPTEPGEVRRLEPGQRFAIDQ LNHLNALKEDRLKKMANTIIMHALGYCYDVRKKKWQAKNPACQIILFEDLSNYNPYEERS RFENSKLMKWSRREIPRQVALQGEIYGLQVGEVGAQFSSRFHAKTGSPGIRCSVVTKEKL QDNRFFKNLQREGRLTLDKIAVLKEGDLYPDKGGEKFISLSKDRKLVTTHADINAAQNLQ KRFWTRTHGFYKVYCKAYQVDGQTVYIPESKDQKQKIIEEFGEGYFILKDGVYEWGNAGK LKIKKGSSKQSSSELVDSDILKDSFDLASELKGEKLMLYRDPSGNVFPSDKWMAAGVFFG KLERILISKLTNQYSISTIEDDSSKQSM

[0283] In some embodiments, napDNAbp refers to Cas12c. In some embodiments, the Cas12c protein is Cas12c1 or a variant of Cas12c1. In some embodiments, the Cas12 protein is Cas12c2 or a variant of Cas12c2. In some embodiments, the Cas12 protein is a Cas12c protein from the Oleiphilus species HI0009 (i.e., OspCas12c) or a variant of OspCas12c. These Cas12c molecules are described in Yan et al., “Functionally Diverse Type V CRISPR-Cas Systems,” Science, 2019 Jan. 4; 363: 88-91, the entire content of which is incorporated herein by reference. In some embodiments, napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a native Cas12c1, Cas12c2, or OspCas12c protein. In some embodiments, napDNAbp is a native Cas12c1, Cas12c2, or OspCas12c protein. In some embodiments, napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the Cas12c1, Cas12c2, or OspCas12c proteins described herein. It should be understood that Cas12c1, Cas12c2, or OspCas12c from other bacterial species may also be used in accordance with the present disclosure. Cas12c1 Cas12c2 OspCas12c

[0284] In some embodiments, napDNAbp refers to Cas12g, Cas12h, or Cas12i, which are described, for example, in Yan et al., “Functionally Diverse Type V CRISPR-Cas Systems,” Science, 2019 Jan. 4; 363: 88-91; the entire content of each is incorporated herein by reference. By aggregating over 10 terabytes of sequence data, a new classification of type V Cas proteins was identified that shows weak similarity to previously characterized class V proteins, including Cas12g, Cas12h, and Cas12i. In some embodiments, the Cas12 protein is Cas12g or a variant of Cas12g. In some embodiments, the Cas12 protein is Cas12h or a variant of Cas12h. In some embodiments, the Cas12 protein is Cas12i or a variant of Cas12i. Other RNA-guided DNA-binding proteins may be used as napDNAbp, and it should be understood that they are within the scope of the present disclosure. In some embodiments, napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a native Cas12g, Cas12h, or Cas12i protein. In some embodiments, napDNAbp is a native Cas12g, Cas12h, or Cas12i protein. In some embodiments, napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any Cas12g, Cas12h, or Cas12i protein described herein. It should be understood that Cas12g, Cas12h, or Cas12i from other bacterial species may also be used in accordance with the present disclosure.In some embodiments, Cas12i is Cas12i1 or Cas12i2. Cas12g1 MAQASSTPAVSPRPRPRYREERTLVRKLLPRPGQSKQEFRENVKKLRKAFLQFNADVSGVCQWAIQFRPRYGKPAEPTETFWKFFLEPETSLPPNDSRSPEFRRLQAFEAAAGINGAAALDDPAFTNELRDSILAVASRPKTKEAQRLFSRLKDYQPAHRMILAKVAAEWIESRYRRAHQNWERNYEEWKKEKQEWEQNHPELTPEIREAFNQIFQQLEVKEKRVRICPAARLLQNKDNCQYAGKNKHSVLCNQFNEFKKNHLQGKAIKFFYKDAEKYLRCGLQSLKPNVQGPFREDWNKYLRYMNLKEETLRGKNGGRLPHCKNLGQECEFNPHTALCKQYQQQLSSRPDLVQHDELYRKWRREYWREPRKPVFRYPSVKRHSIAKIFGENYFQADFKNSVVGLRLDSMPAGQYLEFAFAPWPRNYRPQPGETEISSVHLHFVGTRPRIGFRFRVPHKRSRFDCTQEELDELRSRTFPRKAQDQKFLEAARKRLLETFPGNAEQELRLLAVDLGTDSARAAFFIGKTFQQAFPLKIVKIEKLYEQWPNQKQAGDRRDASSKQPRPGLSRDHVGRHLQKMRAQASEIAQKRQELTGTPAPETTTDQAAKKATLQPFDLRGLTVHTARMIRDWARLNARQIIQLAEENQVDLIVLESLRGFRPPGYENLDQEKKRRVAFFAHGRIRRKVTEKAVERGMRVVTVPYLASSKVCAECRKKQKDNKQWEKNKKRGLFKCEGCGSQAQVDENAARVLGRVFWGEIELPTAIP Cas12h1 MKVHEIPRSQLLKIKQYEGSFVEWYRDLQEDRKKFASLLFRWAAFGYAAREDDGATYISPSQALLERRLLLGDAEDVAIKFLDVLFKGGAPSSSCYSLFYEDFALRDKAKYSGAKREFIEGLATMPLDKIIERIRQDEQLSKIPAEEWLILGAEYSPEEIWEQVAPRIVNVDRSLGKQLRERLGIKCRRPHDAGYCKILMEVVARQLRSHNETYHEYLNQTHEMKTKVANNLTNEFDLVCEFAEVLEEKNYGLGWYVLWQGVKQALKEQKKPTKIQIAVDQLRQPKFAGLLTAKWRALKGAYDTWKLKKRLEKRKAFPYMPNWDNDYQIPVGLTGLGVFTLEVKRTEVVVDLKEHGKLFCSHSHYFGDLTAEKHPSRYHLKFRHKLKLRKRDSRVEPTIGPWIEAALREITIQKKPNGVFYLGLPYALSHGIDNFQIAKRFFSAAKPDKEVINGLPSEMVVGAADLNLSNIVAPVKARIGKGLEGPLHALDYGYGELIDGPKILTPDGPRCGELISLKRDIVEIKSAIKEFKACQREGLTMSEETTTWLSEVESPSDSPRCMIQSRIADTSRRLNSFKYQMNKEGYQDLAEALRLLDAMDSYNSLLESYQRMHLSPGEQSPKEAKFDTKRASFRDLLRRRVAHTIVEYFDDCDIVFFEDLDGPSDSDSRNNALVKLLSPRTLLLYIRQALEKRGIGMVEVAKDGTSQNNPISGHVGWRNKQNKSEIYFYEDKELLVMDADEVGAMNILCRGLNHSVCPYSFVTKAPEKKNDEKKEGDYGKRVKRFLKDRYGSSNVRFLVASMGFVTVTTKRPKDALVGKRLYYHGGELVTHDLHNRMKDEIKYLVEKEVLARRVSLSDSTIKSYKSFAHV Cas12i1 Cas12i2

[0285] Representative nucleic acid and protein sequences of base editors are as follows: BhCas12b GGSGGS-ABE8-Xten20 (P153) JPEG0007717684000040.jpg215169JPEG0007717684000041.jpg244164JPEG0007717684000042.jpg212169BhCas12b GGSGGS-ABE8-Xten20 (K255) JPEG0007717684000043.jpg30169JPEG0007717684000044.jpg243168JPEG0007717684000045.jpg234157JPEG0007717684000046.jpg133163BhCas12b GGSGGS-ABE8-Xten20 (D306) JPEG0007717684000047.jpg108169JPEG0007717684000048.jpg221157JPEG0007717684000049.jpg227153JPEG0007717684000050.jpg57165BhCas12b GGSGGS-ABE8-Xten20 (D980) JPEG0007717684000051.jpg193169JPEG0007717684000052.jpg229160JPEG0007717684000053.jpg214163BhCas12b GGSGGS-ABE8-Xten20 (K1019) JPEG0007717684000054.jpg232157JPEG0007717684000055.jpg229153JPEG0007717684000056.jpg147168

[0286] In the above array, the Kozak sequence is underlined in bold; the dashed line indicates the N-terminal nuclear localization signal (NLS) following the Kozak sequence; the lowercase letters indicate the GGGSGGS linker; the dotted line indicates the sequence encoding ABE8; the unmodified sequence encodes BhCas12b; the double underline indicates the Xten20 linker; the single underline indicates the C-terminal NLS; the dotted-underlined GGATCC indicates the GS linker; and the italicized letters represent the coding sequence of the 3× hemagglutinin (HA) tag.

[0287] In some embodiments, any nucleic acid programmable DNA binding protein (napDNAbp) of the fusion proteins provided herein may be a Cas12j / CasΦ protein. Cas12j / CasΦ is described in Pausch et al., “CRISPR-CasΦ from huge phages is a hypercompact genome editor,” Science, 17 July 2020, Vol. 369, Issue 6501, pp. 333-337, which is incorporated herein by reference in its entirety. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a native Cas12j / CasΦ protein. In some embodiments, the napDNAbp is a native Cas12j / CasΦ protein. In some embodiments, the napDNAbp is a nuclease-inactive (“dead”) Cas12j / CasΦ protein. It should be understood that Cas12j / CasΦ from other species may also be used in accordance with the present disclosure. Exemplary Cas12j / CasΦ amino acid sequences are as follows: >CasΦ-1 MADTPTLFTQFLRHHLPGQRFRKDILKQAGRILANKGEDATIAFLRGKSEESPPDFQPPVKCPIIACSRPLTEWPIYQASVAIQGYVYGQSLAEFEASDPGCSKDGLLGWFDKTGVCTDYFSVQGLNLIFQNARKRYIGVQTKVTNRNEKRHKKLKRINAKRIAEGLPELTSDEPESALDETGHLIDPPGLNTNIYCYQQVSPKPLALSEVNQLPTAYAGYSTSGDDPIQPMVTKDRLSISKGQPGYIPEHQRALLSQKKHRRMRGYGLKARALLVIVRIQDDWAVIDLRSLLRNAYWRRIVQTKEPSTITKLLKLVTGDPVLDATRMVATFTYKPGIVQVRSAKCLKNKQGSKLFSERYLNETVSVTSIDLGSNNLVAVATYRLVNGNTPELLQRFTLPSHLVKDFERYKQAHDTLEDSIQKTAVASLPQGQQTEIRMWSMYGFREAQERVCQELGLADGSIPWNVMTATSTILTDLFLARGGDPKKCMFTSEPKKKKNSKQVLYKIRDRAWAKMYRTLLSKETREAWNKALWGLKRGSPDYARLSKRKEELARRCVNYTISTAEKRAQCGRTIVALEDLNIGFFHGRGKQEPGWVGLFTRKKENRWLMQALHKAFLELAHHRGYHVIEVNPAYTSQTCPVCRHCDPDNRDQHNREAFHCIGCGFRGNADLDVATHNIAMVAITGESLKRARGSVASKTPQPLAAE* >CasΦ-2 MPKPAVESEFSKVLKKHFPGERFRSSYMKRGGKILAAQGEEAVVAYLQGKSEEEPPNFQPPAKCHVVTKSRDFAEWPIMKASEAIQRYIYALSTTERAACKPGKSSESHAAWFAATGVSNHGYSHVQGLNLIFDHTLGRYDGVLKKVQLRNEKARARLESINASRADEGLPEIKAEEEEVATNETGHLLQPPGINPSFYVYQTISPQAYRPRDEIVLPPEYAGYVRDPNAPIPLGVVRNRCDIQKGCPGYIPEWQREAGTAISPKTGKAVTVPGLSPKKNKRMRRYWRSEKEKAQDALLVTVRIGTDWVVIDVRGLLRNARWRTIAPKDISLNALLDLFTGDPVIDVRRNIVTFTYTLDACGTYARKWTLKGKQTKATLDKLTATQTVALVAIDLGQTNPISAGISRVTQENGALQCEPLDRFTLPDDLLKDISAYRIAWDRNEEELRARSVEALPEAQQAEVRALDGVSKETARTQLCADFGLDPKRLPWDKMSSNTTFISEALLSNSVSRDQVFFTPAPKKGAKKKAPVEVMRKDRTWARAYKPRLSVEAQKLKNEALWALKRTSPEYLKLSRRKEELCRRSINYVIEKTRRRTQCQIVIPVIEDLNVRFFHGSGKRLPGWDNFFTAKKENRWFIQGLHKAFSDLRTHRSFYVFEVRPERTSITCPKCGHCEVGNRDGEAFQCLSCGKTCNADLDVATHNLTQVALTGKTMPKREEPRDAQGTAPARKTKKASKSKAPPAEREDQTPAQEPSQTS >CasΦ-3 MEKEITELTKIRREFPNKKFSSTDMKKAGKLLKAEGPDAVRDFLNSCQEIIGDFKPPVKTNIVSISRPFEEWPVSMVGRAIQEYYFSLTKEELESVHPGTSSEDHKSFFNITGLSNYNYTSVQGLNLIFKNAKAIYDGTLVKANNKNKKLEKKFNEINHKRSLEGLPIITPDFEEPFDENGHLNNPPGINRNIYGYQGCAAKVFVPSKHKMVSLPKEYEGYNRDPNLSLAGFRNRLEIPEGEPGHVPWFQRMDIPEGQIGHVNKIQRFNFVHGKNSGKVKFSDKTGRVKRYHHSKYKDATKPYKFLEESKKVSALDSILAIITIGDDWVVFDIRGLYRNVFYRELAQKGLTAVQLLDLFTGDPVIDPKKGVVTFSYKEGVVPVFSQKIVPRFKSRDTLEKLTSQGPVALLSVDLGQNEPVAARVCSLKNINDKITLDNSCRISFLDDYKKQIKDYRDSLDELEIKIRLEAINSLETNQQVEIRDLDVFSADRAKANTVDMFDIDPNLISWDSMSDARVSTQISDLYLKNGGDESRVYFEINNKRIKRSDYNISQLVRPKLSDSTRKNLNDSIWKLKRTSEEYLKLSKRKLELSRAVVNYTIRQSKLLSGINDIVIILEDLDVKKKFNGRGIRDIGWDNFFSSRKENRWFIPAFHKAFSELSSNRGLCVIEVNPAWTSATCPDCGFCSKENRDGINFTCRKCGVSYHADIDVATLNIARVAVLGKPMSGPADRERLGDTKKPRVARSRKTMKRKDISNSTVEAMVTA* >CasΦ-4 MYSLEMADLKSEPSLLAKLLRDRFPGKYWLPKYWKLAEKKRLTGGEEAACEYMADKQLDSPPPNFRPPARCVILAKSRPFEDWPVHRVASKAQSFVIGLSEQGFAALRAAPPSTADARRDWLRSHGASEDDLMALEAQLLETIMGNAISLHGGVLKKIDNANVKAAKRLSGRNEARLNKGLQELPPEQEGSAYGADGLLVNPPGLNLNIYCRKSCCPKPVKNTARFVGHYPGYLRDSDSILISGTMDRLTIIEGMPGHIPAWQREQGLVKPGGRRRRLSGSESNMRQKVDPSTGPRRSTRSGTVNRSNQRTGRNGDPLLVEIRMKEDWVLLDARGLLRNLRWRESKRGLSCDHEDLSLSGLLALFSGDPVIDPVRNEVVFLYGEGIIPVRSTKPVGTRQSKKLLERQASMGPLTLISCDLGQTNLIAGRASAISLTHGSLGVRSSVRIELDPEIIKSFERLRKDADRLETEILTAAKETLSDEQRGEVNSHEKDSPQTAKASLCRELGLHPPSLPWGQMGPSTTFIADMLISHGRDDDAFLSHGEFPTLEKRKKFDKRFCLESRPLLSSETRKALNESLWEVKRTSSEYARLSQRKKEMARRAVNFVVEISRRKTGLSNVIVNIEDLNVRIFHGGGKQAPGWDGFFRPKSENRWFIQAIHKAFSDLAAHHGIPVIESDPQRTSMTCPECGHCDSKNRNGVRFLCKGCGASMDADFDAACRNLERVALTGKPMPKPSTSCERLLSATTGKVCSDHSLSHDAIEKAS* >CasΦ-5 MSSLPTPLELLKQKHADLFKGLQFSSKDNKMAGKVLKKDGEEAALAFLSERGVSRGELPNFRPPAKTLVVAQSRPFEEFPIYRVSEAIQLYVYSLSVKELETVPSGSSTKKEHQRFFQDSSVPDFGYTSVQGLNKIFGLARGIYLGVITRGENQLQKAKSKHEALNKKRRASGEAETEFDPTPYEYMTPERKLAKPPGVNHSIMCYVDISVDEFDFRNPDGIVLPSEYAGYCREINTAIEKGTVDRLGHLKGGPGYIPGHQRKESTTEGPKINFRKGRIRRSYTALYAKRDSRRVRQGKLALPSYRHHMMRLNSNAESAILAVIFFGKDWVVFDLRGLLRNVRWRNLFVDGSTPSTLLGMFGDPVIDPKRGVVAFCYKEQIVPVVSKSITKMVKAPELLNKLYLKSEDPLVLVAIDLGQTNPVGVGVYRVMNASLDYEVVTRFALESELLREIESYRQRTNAFEAQIRAETFDAMTSEEQEEITRVRAFSASKAKENVCHRFGMPVDAVDWATMGSNTIHIAKWVMRHGDPSLVEVLEYRKDNEIKLDKNGVPKKVKLTDKRIANLTSIRLRFSQETSKHYNDTMWELRRKHPVYQKLSKSKADFSRRVVNSIIRRVNHLVPRARIVFIIEDLKNLGKVFHGSGKRELGWDSYFEPKSENRWFIQVLHKAFSETGKHKGYYIIECWPNWTSCTCPKCSCCDSENRHGEVFRCLACGYTCNTDFGTAPDNLVKIATTGKGLPGPKKRCKGSSKGKNPKIARSSETGVSVTESGAPKVKKSSPTQTSQSSSQSAP* >CasΦ-6 MNKIEKEKTPLAKLMNENFAGLRFPFAIIKQAGKKLLKEGELKTIEYMTGKGSIEPLPNFKPPVKCLIVAKRRDLKYFPICKASCEIQSYVYSLNYKDFMDYFSTPMTSQKQHEEFFKKSGLNIEYQNVAGLNLIFNNVKNTYNGVILKVKNRNEKLKKKAIKNNYEFEEIKTFNDDGCLINKPGINNVIYCFQSISPKILKNITHLPKEYNDYDCSVDRNIIQKYVSRLDIPESQPGHVPEWQRKLPEFNNTNNPRRRRKWYSNGRNISKGYSVDQVNQAKIEDSLLAQIKIGEDWIILDIRGLLRDLNRRELISYKNKLTIKDVLGFFSDYPIIDIKKNLVTFCYKEGVIQVVSQKSIGNKKSKQLLEKLIENKPIALVSIDLGQTNPVSVKISKLNKINNKISIESFTYRFLNEEILKEIEKYRKDYDKLELKLINEA >CasΦ-7 MSNTAVSTREHMSNKTTPPSPLSLLLRAHFPGLKFESQDYKIAGKKLRDGGPEAVISYLTGKGQAKLKDVKPPAKAFVIAQSRPFIEWDLVRVSRQIQEKIFGIPATKGRPKQDGLSETAFNEAVASLEVDGKSKLNEETRAAFYEVLGLDAPSLHAQAQNALIKSAISIREGVLKKVENRNEKNLSKTKRRKEAGEEATFVEEKAHDERGYLIHPPGVNQTIPGYQAVVIKSCPSDFIGLPSGCLAKESAEALTDYLPHDRMTIPKGQPGYVPEWQHPLLNRRKNRRRRDWYSASLNKPKATCSKRSGTPNRKNSRTDQIQSGRFKGAIPVLMRFQDEWVIIDIRGLLRNARYRKLLKEKSTIPDLLSLFTGDPSIDMRQGVCTFIYKAGQACSAKMVKTKNAPEILSELTKSGPVVLVSIDLGQTNPIAAKVSRVTQLSDGQLSHETLLRELLSNDSSDGKEIARYRVASDRLRDKLANLAVERLSPEHKSEILRAKNDTPALCKARVCAALGLNPEMIAWDKMTPYTEFLATAYLEKGGDRKVATLKPKNRPEMLRRDIKFKGTEGVRIEVSPEAAEAYREAQWDLQRTSPEYLRLSTWKQELTKRILNQLRHKAAKSSQCEVVVMAFEDLNIKMMHGNGKWADGGWDAFFIKKRENRWFMQAFHKSLTELGAHKGVPTIEVTPHRTSITCTKCGHCDKANRDGERFACQKCGFVAHADLEIATDNIERVALTGKPMPKPESERSGDAKKSVGARKAAFKPEEDAEAAE* >CasΦ-8 MIKPTVSQFLTPGFKLIRNHSRTAGLKLKNEGEEACKKFVRENEIPKDECPNFQGGPAIANIIAKSREFTEWEIYQSSLAIQEVIFTLPKDKLPEPILKEEWRAQWLSEHGLDTVPYKEAAGLNLIIKNAVNTYKGVQVKVDNKNKNNLAKINRKNEIAKLNGEQEISFEEIKAFDDKGYLLQKPSPNKSIYCYQSVSPKPFITSKYHNVNLPEEYIGYYRKSNEPIVSPYQFDRLRIPIGEPGYVPKWQYTFLSKKENKRRKLSKRIKNVSPILGIICIKKDWCVFDMRGLLRTNHWKKYHKPTDSINDLFDYFTGDPVIDTKANVVRFRYKMENGIVNYKPVREKKGKELLENICDQNGSCKLATVDVGQNNPVAIGLFELKKVNGELTKTLISRHPTPIDFCNKITAYRERYDKLESSIKLDAIKQLTSEQKIEVDNYNNNFTPQNTKQIVCSKLNINPNDLPWDKMISGTHFISEKAQVSNKSEIYFTSTDKGKTKDVMKSDYKWFQDYKPKLSKEVRDALSDIEWRLRRESLEFNKLSKSREQDARQLANWISSMCDVIGIENLVKKNNFFGGSGKREPGWDNFYKPKKENRWWINAIHKALTELSQNKGKRVILLPAMRTSITCPKCKYCDSKNRNGEKFNCLKCGIELNADIDVATENLATVAITAQSMPKPTCERSGDAKKPVRARKAKAPEFHDKLAPSYTVVLREAV* >CasΦ-9 MRSSREIGDKILMRQPAEKTAFQVFRQEVIGTQKLSGGDAKTAGRLYKQGKMEAAREWLLKGARDDVPPNFQPPAKCLVVAVSHPFEEWDISKTNHDVQAYIYAQPLQAEGHLNGLSEKWEDTSADQHKLWFEKTGVPDRGLPVQAINKIAKAAVNRAFGVVRKVENRNEKRRSRDNRIAEHNRENGLTEVVREAPEVATNADGFLLHPPGIDPSILSYASVSPVPYNSSKHSFVRLPEEYQAYNVEPDAPIPQFVVEDRFAIPPGQPGYVPEWQRLKCSTNKHRRMRQWSNQDYKPKAGRRAKPLEFQAHLTRERAKGALLVVMRIKEDWVVFDVRGLLRNVEWRKVLSEEAREKLTLKGLLDLFTGDPVIDTKRGIVTFLYKAEITKILSKRTVKTKNARDLLLRLTEPGEDGLRREVGLVAVDLGQTHPIAAAIYRIGRTSAGALESTVLHRQGLREDQKEKLKEYRKRHTALDSRLRKEAFETLSVEQQKEIVTVSGSGAQITKDKVCNYLGVDPSTLPWEKMGSYTHFISDDFLRRGGDPNIVHFDRQPKKGKVSKKSQRIKRSDSQWVGRMRPRLSQETAKARMEADWAAQNENEEYKRLARSKQELARWCVNTLLQNTRCITQCDEIVVVIEDLNVKSLHGKGAREPGWDNFFTPKTENRWFIQILHKTFSELPKHRGEHVIEGCPLRTSITCPACSYCDKNSRNGEKFVCVACGATFHADFEVATYNLVRLATTGMPMPKSLERQGGGEKAGGARKARKKAKQVEKIVVQANANVTMNGASLHSP* >CasΦ-10 MDMLDTETNYATETPAQQQDYSPKPPKKAQRAPKGFSKKARPEKKPPKPITLFTQKHFSGVRFLKRVIRDASKILKLSESRTITFLEQAIERDGSAPPDVTPPVHNTIMAVTRPFEEWPEVILSKALQKHCYALTKKIKIKTWPKKGPGKKCLAAWSARTKIPLIPGQVQATNGLFDRIGSIYDGVEKKVTNRNANKKLEYDEAIKEGRNPAVPEYETAYNIDGTLINKPGYNPNLYITQSRTPRLITEADRPLVEKILWQMVEKKTQSRNQARRARLEKAAHLQGLPVPKFVPEKVDRSQKIEIRIIDPLDKIEPYMPQDRMAIKASQDGHVPYWQRPFLSKRRNRRVRAGWGKQVSSIQAWLTGALLVIVRLGNEAFLADIRGALRNAQWRKLLKPDATYQSLFNLFTGDPVVNTRTNHLTMAYREGVVNIVKSRSFKGRQTREHLLTLLGQGKTVAGVSFDLGQKHAAGLLAAHFGLGEDGNPVFTPIQACFLPQRYLDSLTNYRNRYDALTLDMRRQSLLALTPAQQQEFADAQRDPGGQAKRACCLKLNLNPDEIRWDLVSGISTMISDLYIERGGDPRDVHQQVETKPKGKRKSEIRILKIRDGKWAYDFRPKIADETRKAQREQLWKLQKASSEFERLSRYKINIARAIANWALQWGRELSGCDIVIPVLEDLNVGSKFFDGKGKWLLGWDNRFTPKKENRWFIKVLHKAVAELAPHRGVPVYEVMPHRTSMTCPACHYCHPTNREGDRFECQSCHVVKNTDRDVAPYNILRVAVEGKTLDRWQAEKKPQAEPDRPMILIDNQES* The asterisk (*) in the above sequence indicates a STOP codon. Alternatively, CasΦ-1 is also referred to as Cas12j ortholog 1. Thus, CasΦ-1 - CasΦ-10 can be referred to as Cas12j orthologs 1 - 10, respectively.

[0288] Guide polynucleotide In one embodiment, the guide polynucleotide is a guide RNA. As used herein, the term "guide RNA (gRNA)" and its grammatical equivalents can refer to an RNA that can be specific for a target DNA and can form a complex with a Cas protein. The RNA / Cas complex can assist in "guiding" the Cas protein to the target DNA. Cas9 / crRNA / tracrRNA endonucleolytically cleaves a linear or circular dsDNA target complementary to the spacer. The target strand not complementary to the crRNA is first endonucleolytically cleaved and then trimmed 3'-5' exonucleolytically. In nature, both a protein and both RNAs are usually required for DNA binding and cleavage. However, a single guide RNA ("sgRNA", or simply "gRNA") can be engineered to incorporate both aspects of crRNA and tracrRNA into a single RNA species. See, for example, Jinek M. et al, Science 337: 816-821 (2012) (the entire content of which is incorporated herein by reference). Cas9 recognizes a short motif (PAM or protospacer adjacent motif) within the CRISPR repeat sequence to assist in distinguishing self from non-self.The sequences and structures of Cas9 nucleases are well known to those of ordinary skill in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti, J. J. et al., Proc. Natl. Acad. Sci. U.S.A. 98: 4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E.et al., Nature 471: 602-607 (2011); and “Programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M. et al, Science 337: 816-821 (2012) (the entire contents of each of which are incorporated herein by reference)). Cas9 orthologs have been described in a variety of species, including but not limited to S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences may be apparent to those of ordinary skill in the art based on the present disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10: 5, 726-737; the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is a nickase.

[0289] In some embodiments, the guide polynucleotide is at least one single guide RNA (“sgRNA” or “gRNA”). In some embodiments, the guide polynucleotide is at least one tracrRNA. In some embodiments, the guide polynucleotide does not require a PAM sequence to guide a polynucleotide programmable DNA binding domain (e.g., Cas9 or Cpf1) to a target nucleotide sequence.

[0290] The polynucleotide-programmable nucleotide binding domains (e.g., CRISPR-derived domains) of the base editors disclosed herein can recognize a target polynucleotide sequence by associating with a guide polynucleotide. The guide polynucleotide (e.g., gRNA) is generally single-stranded and is programmed to bind site-specifically (i.e., via complementary base pairs) to the target sequence of the polynucleotide, thereby directing the base editor combined with the guide nucleic acid to the target sequence. The guide polynucleotide can be DNA. The guide polynucleotide can be RNA. As will be understood by those skilled in the art, thymine (T) in the sequence is replaced by uracil (U) in the guide polynucleotide sequence. In some cases, the guide polynucleotide contains natural nucleotides (e.g., adenosine). In some cases, the guide polynucleotide contains non-natural (or unnatural) nucleotides (e.g., peptide nucleic acids or nucleotide analogs). In some cases, the target region of the guide nucleic acid sequence can be at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. The target region of the guide nucleic acid can be between 10 and 30 nucleotides in length, or between 15 and 25 nucleotides in length, or between 15 and 20 nucleotides in length. In some embodiments, the guide polynucleotide can be truncated by 1, 2, 3, 4, etc. nucleotides, particularly at the 5' end. By way of non-limiting example, a guide polynucleotide that is 20 nucleotides in length can be truncated by 1, 2, 3, 4, etc. nucleotides, particularly at the 5' end.

[0291] In some embodiments, the guide polynucleotide comprises two or more individual polynucleotides, which can interact with each other, for example, via complementary base pairing (e.g., a dual guide polynucleotide). For example, the guide polynucleotide can comprise a CRISPR RNA (crRNA) and a trans-activating CRISPR RNA (tracrRNA). For example, the guide polynucleotide can comprise one or more trans-activating CRISPR RNAs (tracrRNAs).

[0292] In type II CRISPR systems, targeting of nucleic acids by a CRISPR protein (such as Cas9) generally requires complementary base pairing between a first RNA molecule (crRNA) containing a sequence that recognizes the target sequence and a second RNA molecule (trRNA) containing repeat sequences that form a scaffold region that stabilizes the guide RNA-CRISPR protein complex. Such a dual guide RNA system can be used as a guide polynucleotide to direct the base editors disclosed herein to a target polynucleotide sequence.

[0293] In some embodiments, the base editors provided herein utilize a single guide polynucleotide (e.g., sgRNA). In some embodiments, the base editors provided herein utilize a dual guide polynucleotide (e.g., dual gRNA). In some embodiments, the base editors provided herein utilize one or more guide polynucleotides (e.g., multiplex gRNA). In some embodiments, a single guide polynucleotide is utilized for different base editors described herein. For example, a single guide polynucleotide can be utilized for a cytidine base editor and an adenosine base editor.

[0294] In other embodiments, the guide polynucleotide can include both the polynucleotide targeting portion of the nucleic acid and the scaffold portion of the nucleic acid in a single molecule (i.e., a single molecule guide nucleic acid). For example, the single molecule guide polynucleotide can be a single guide RNA (sgRNA or gRNA). As used herein, the term guide polynucleotide sequence contemplates any single, double, or multi-molecule nucleic acid that can interact with a base editor and direct it to a target polynucleotide sequence.

[0295] Generally, a guide polynucleotide (e.g., a crRNA / trRNA complex or a gRNA) includes a “polynucleotide targeting segment” that contains a sequence capable of recognizing and binding to a target polynucleotide sequence, and a “protein binding segment” that stabilizes the guide polynucleotide within the polynucleotide programmable nucleotide binding domain component of a base editor. In some embodiments, the polynucleotide targeting segment of the guide polynucleotide recognizes and binds to a DNA polynucleotide, thereby facilitating editing of a base in the DNA. In other cases, the polynucleotide targeting segment of the guide polynucleotide recognizes and binds to an RNA polynucleotide, thereby facilitating editing of a base in the RNA. As used herein, “segment” refers to a section or region of a molecule, e.g., a continuous stretch of nucleotides in a guide polynucleotide. A segment can also represent a region / section of a complex, and thus a segment can include multiple regions of multiple molecules. For example, if a guide polynucleotide includes multiple nucleic acid molecules, the protein binding segment can include all or a portion of multiple separate molecules hybridized along complementary regions. In some embodiments, the protein binding segment of a DNA-targeting RNA that includes two separate molecules can include (i) base pairs 40-75 of a first RNA molecule that is 100 base pairs in length and; (ii) base pairs 10-25 of a second RNA molecule that is 50 base pairs in length. The definition of “segment” is not limited to a particular number of total base pairs, not limited to a particular number of base pairs from a given RNA molecule, not limited to a particular number of separate molecules within a complex, can include regions of any full-length RNA molecule, and can include regions complementary to other molecules, unless otherwise specifically defined in a particular context.

[0296] A guide RNA or guide polynucleotide can comprise two or more RNAs, such as CRISPR RNA (crRNA) and trans-activating RNA (tracrRNA). The guide RNA or guide polynucleotide can optionally comprise a single-stranded RNA formed by fusion of a portion of the crRNA (e.g., the functional portion) and the tracrRNA, or a single guide RNA (sgRNA). The guide RNA or guide polynucleotide can also be a double-stranded RNA comprising the crRNA and the tracrRNA. Further, the crRNA can hybridize to a target DNA.

[0297] As described above, the guide RNA or guide polynucleotide can be an expression product. For example, the DNA encoding the guide RNA can be a vector comprising the sequence encoding the guide RNA. The guide RNA or guide polynucleotide can be introduced into cells by transfecting the cells with an isolated guide RNA or with plasmid DNA comprising the sequences encoding the guide RNA and a promoter. The guide RNA or guide polynucleotide can also be introduced into cells by other methods, such as using virus-mediated gene delivery.

[0298] The guide RNA or guide polynucleotide can be isolated. For example, the guide RNA can be transfected into cells or organisms in the form of isolated RNA. The guide RNA can be prepared by in vitro transcription using any in vitro transcription system known in the art. The guide RNA can be introduced into cells in the form of isolated RNA rather than in the form of a plasmid comprising the coding sequence of the guide RNA.

[0299] A guide RNA or guide polynucleotide can include three regions: a first region at the 5' end that can be complementary to a target site in a chromosomal sequence, a second internal region that can form a stem-loop structure, and a third 3' region that can be single-stranded. The first regions of each guide RNA may differ such that each guide RNA guides a fusion protein to a specific target site. Further, the second and third regions of each guide RNA can be the same in all guide RNAs.

[0300] The first region of a guide RNA or guide polynucleotide can be complementary to the sequence at a target site in a chromosomal sequence such that the first region of the guide RNA can form base pairs with the target site. In some cases, the first region of the guide RNA can include about 10 nucleotides to 25 nucleotides (i.e., 10 nucleotides to nucleotides; or about 10 nucleotides to about 25 nucleotides; or 10 nucleotides to about 25 nucleotides; or about 10 nucleotides to 25 nucleotides) or more. For example, the region of base pair formation between the first region of the guide RNA and the target site in the chromosomal sequence can be about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25, or more nucleotides in length. In some embodiments, the first region of the guide RNA can be 19, 20, or 21 nucleotides in length or can be about 19, 20, or 21 nucleotides in length.

[0301] The guide RNA or guide polynucleotide may also include a second region that forms a secondary structure. For example, the secondary structure formed by the guide RNA may include a stem (or hairpin) and a loop. The lengths of the loop and the stem can vary. For example, the loop can range in length from about 3 to 10 nucleotides, and the stem can range in length from about 6 to 20 base pairs. The stem can include one or more bulges of 1 to 10 or about 10 nucleotides. The total length of the second region can range in length from about 16 to 60 nucleotides. For example, the loop can be 4 nucleotides in length or about 4 nucleotides in length, and the stem can be 12 base pairs in length or about 12 base pairs in length.

[0302] The guide RNA or guide polynucleotide may also include a third region at the 3' end that can be essentially single-stranded. For example, the third region may or may not be complementary to the chromosomal sequence in the target cell and may or may not be complementary to the rest of the guide RNA. Furthermore, the length of the third region can vary. The third region can be more than 4 nucleotides or about more than 4 nucleotides in length. For example, the length of the third region can range from about 5 to 60 nucleotides in length.

[0303] The guide RNA or guide polynucleotide can target any exon or intron of a gene target. In some cases, the guide can target exon 1 or 2 of a gene, and in other cases, the guide can target exon 3 or 4 of a gene. The composition can include a multiplex guide RNA that all targets the same exon or, in some cases, a multiplex guide RNA that can target different exons. The exons and introns of a gene can be targeted.

[0304] A guide RNA or guide polynucleotide can target a nucleic acid sequence of 20 nucleotides or about 20 nucleotides. The target nucleic acid can be less than 20 nucleotides or about less than 20 nucleotides. The target nucleic acid can be at least or at least about 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, or any length from 1 to 100 nucleotides. The target nucleic acid can be at most 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, or at most about 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, or any length from 1 to 100 nucleotides. The target nucleic acid sequence can be 20 bases or about 20 bases immediately 5' to the first nucleotide of the PAM. The guide RNA can target a nucleic acid sequence. The target nucleic acid can be at least (or at least about) 1 to 10, 1 to 20, 1 to 30, 1 to 40, 1 to 50, 1 to 60, 1 to 70, 1 to 80, 1 to 90, or 1 to 100 nucleotides.

[0305] A guide polynucleotide, e.g., a guide RNA, can refer to a nucleic acid that can hybridize to another nucleic acid, e.g., a target nucleic acid or protospacer in a cell's genome. The guide polynucleotide can be RNA. The guide polynucleotide can be DNA. The guide polynucleotide can be programmed or designed to bind site-specifically to a nucleic acid sequence. The guide polynucleotide can comprise a single polynucleotide strand and can be referred to as a single guide polynucleotide. The guide polynucleotide can comprise two polynucleotide strands and can be referred to as a dual guide polynucleotide. The guide RNA can be introduced into a cell or embryo as an RNA molecule. For example, the RNA molecule can be transcribed in vitro and / or chemically synthesized. The RNA can be transcribed from a synthetic DNA molecule, e.g., a gBlocks® gene fragment. The guide RNA can then be introduced into the cell or embryo as an RNA molecule. The guide RNA can also be introduced into the cell or embryo in the form of a non-RNA nucleic acid molecule, e.g., a DNA molecule. For example, the DNA encoding the guide RNA can be operably linked to a promoter control sequence for expressing the guide RNA in a cell or embryo of interest. The RNA coding sequence can be operably linked to a promoter sequence recognized by RNA polymerase III (Pol III). Examples of plasmid vectors that can be used to express the guide RNA include, but are not limited to, the px330 vector and the px333 vector. In some cases, the plasmid vector (e.g., the px333 vector) can comprise a DNA sequence encoding at least two guide RNAs.

[0306] Methods for selecting, designing, and validating guide polynucleotides, such as guide RNAs, and target sequences are described herein and are known to those of skill in the art. For example, to minimize the potential impact of substrate non-specificity of deaminase domains (e.g., AID domains) in nucleic acid base editor systems, the number of residues that could potentially be targets of deamination inadvertently (e.g., off-target C residues that could potentially be present in ssDNA within the target nucleic acid locus) can be minimized. Additionally, software tools can be used to optimize the gRNA corresponding to the target nucleic acid sequence, e.g., to minimize the overall off-target activity across the genome. For example, for each selection of a possible targeting domain using S. pyogenes Cas9, all off-target sequences (e.g., preceded by a selected PAM such as NAG or NGG) containing a maximum of a specified number (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) of mismatched base pairs can be identified across the genome. A first region of the gRNA complementary to the target site can be identified, and all first regions (e.g., crRNAs) can be ranked according to the sum of their predicted off-target scores; the top targeting domains represent domains that are likely to have the highest on-target activity and the lowest off-target activity. Candidate targeting gRNAs can be functionally evaluated by using methods known in the art and / or methods described herein.

[0307] As a non-limiting example, the target DNA hybridizing sequence in the crRNA of a guide RNA for use with Cas9 may be identified using a DNA sequence search algorithm. gRNA design may be performed using custom gRNA design software based on the public tool cas-offinder, as described in Bae S., Park J., & Kim J.-S. Cas-OFFinder: A fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases. Bioinformatics 30, 1473-1475 (2014). This software calculates the off-target tendency across the genome and then scores the guides. For guides in the length range of 17 to 24, usually, matches in the range of perfect match to 7 mismatches are considered. After determining the off-target sites computationally, an aggregate score is calculated for each guide and summarized in tabular output using a web interface. In addition to identifying potential target sites adjacent to the PAM sequence, the software also identifies all PAM-adjacent sequences that differ from the selected target site by 1, 2, 3, or more than 3 nucleotides. A genomic DNA sequence for a target nucleic acid sequence (e.g., a target gene) can be obtained, and repetitive elements can be screened using publicly available tools such as the RepeatMasker program. RepeatMasker searches the input DNA sequence for repetitive elements and low-complexity regions. The output is a detailed annotation of the repeats present in a given query sequence.

[0308] Following identification, the first region of a guide RNA, e.g., crRNA, can be ranked hierarchically based on the distance to the target site, their orthogonality, and the presence of 5' nucleotides for close matches with the associated PAM sequence (e.g., for the associated PAM (e.g., NGG PAM for S. pyogenes, NNGRRT or NNGRRV PAM for S. aureus)) for close matches within the human genome containing 5'G). As used herein, orthogonality refers to the number of sequences in the human genome that contain a minimum number of mismatches to the target sequence. "High level of orthogonality" or "good orthogonality" can refer to, for example, a 20mer targeting domain that does not have the same sequence in the human genome outside of the intended target or does not have a sequence that contains one or two mismatches within the target sequence. To minimize off-target DNA cleavage, a targeting domain with good orthogonality can be selected.

[0309] In some embodiments, a reporter system may be used to detect base editing activity and test candidate guide polynucleotides. In some embodiments, the reporter system may include a reporter gene-based assay where base editing activity results in the expression of a reporter gene. For example, the reporter system may include a reporter gene that contains an inactivated start codon on the template strand, e.g., a mutation from 3'-TAC-5' to 3'-CAC-5'. When deamination of the target C is successful, the corresponding mRNA is transcribed as 5'-AUG-3' instead of 5'-GUG-3', enabling translation of the reporter gene. Suitable reporter genes will be apparent to those skilled in the art. Non-limiting examples of reporter genes include genes encoding green fluorescent protein (GFP), red fluorescent protein (RFP), luciferase, secreted alkaline phosphatase (SEAP), or any other gene whose expression is detectable and apparent to those skilled in the art. The reporter system can be used to test many different gRNAs, for example, to determine which residue(s) each deaminase targets with respect to a target DNA sequence. sgRNAs targeting the non-template strand can also be tested to evaluate the off-target effects of specific base editing proteins such as Cas9 deaminase fusion proteins. In some embodiments, such gRNAs can be designed such that the mutated start codon does not form base pairs with the gRNA. The guide polynucleotide can include standard ribonucleotides, modified ribonucleotides (e.g., pseudouridine), ribonucleotide isomers, and / or ribonucleotide analogs. In some embodiments, the guide polynucleotide can include at least one detectable label. The detectable label can be a fluorophore (e.g., FAM, TMR, Cy3, Cy5, Texas Red, Oregon Green, Alexa Fluors, Halo tag, or a suitable fluorescent dye), a detection tag (e.g., biotin, digoxigenin, etc.), a quantum dot, or a gold particle.

[0310] Guide polynucleotides can be chemically synthesized, enzymatically synthesized, or a combination thereof. For example, guide RNAs can be synthesized using standard phosphoramidite-based solid-phase synthesis methods. Alternatively, guide RNAs can be synthesized in vitro by operably linking DNA encoding the guide RNA to a promoter control sequence recognized by a phage RNA polymerase. Examples of suitable phage promoter sequences include T7, T3, SP6 promoter sequences, or variants thereof. In embodiments where the guide RNA comprises two separate molecules (e.g., crRNA and tracrRNA), the crRNA can be chemically synthesized and the tracrRNA can be enzymatically synthesized.

[0311] In some embodiments, a base editor system can include multiple guide polynucleotides, e.g., gRNAs. For example, the gRNAs can target one or more target loci included in the base editor system (e.g., at least 1 gRNA, at least 2 gRNAs, at least 5 gRNAs, at least 10 gRNAs, at least 20 gRNAs, at least 30 gRNAs, at least 50 gRNAs). The multiple gRNA sequences can be arranged in tandem, preferably separated by direct repeats.

[0312] The DNA sequence encoding the guide RNA or guide polynucleotide can also be part of a vector. Further, the vector can include additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), selectable marker sequences (e.g., antibiotic resistance genes such as GFP or puromycin), an origin of replication, etc. The DNA molecule encoding the guide RNA can also be linear. The DNA sequence encoding the guide RNA or guide polynucleotide can also be circular.

[0313] In some embodiments, one or more components of the base editor system can be encoded by a DNA sequence. Such DNA sequences can be introduced into an expression system, such as a cell, together or separately. For example, DNA sequences encoding a polynucleotide-programmable nucleotide binding domain and a guide RNA can be introduced into a cell, and each DNA sequence can be part of a separate molecule (e.g., one vector containing the polynucleotide-programmable nucleotide binding domain coding sequence and a second vector containing the guide RNA coding sequence), or both can be part of the same molecule (e.g., one vector containing the coding (and regulatory) sequences for both the polynucleotide-programmable nucleotide binding domain and the guide RNA).

[0314] One or more modifications can be included in the guide polynucleotide to provide new or enhanced features to the nucleic acid. The guide polynucleotide can include a nucleic acid affinity tag. The guide polynucleotide can include synthetic nucleotides, synthetic nucleotide analogs, nucleotide derivatives, and / or modified nucleotides.

[0315] In some cases, the gRNA or guide polynucleotide can include modifications. The modifications can be made at any position of the gRNA or guide polynucleotide. Multiple modifications can be added to a single gRNA or guide polynucleotide. After modification, the gRNA or guide polynucleotide can be subjected to quality control. In some cases, quality control includes PAGE, HPLC, MS, or any combination thereof.

[0316] Modifications of the gRNA or guide polynucleotide can be substitutions, insertions, deletions, chemical modifications, physical modifications, stabilization, purification, or any combination thereof.

[0317] The gRNA or guide polynucleotide can also be modified by 5'-adenylate, 5'-guanosine-triphosphate cap, 5'-N7-methylguanosine-triphosphate cap, 5'-triphosphate cap, 3'-phosphate, 3'-thiophosphate, 5'-phosphate, 5'-thiophosphate, cis-Syn thymidine dimer, trimer, C12 spacer, C3 spacer, C6 spacer, dSpacer, PC spacer, rSpacer, Spacer 18, Spacer 9, 3'-3' modification, 5'-5' modification, abasic, acridine, azobenzene, biotin, biotin BB, biotin TEG, cholesteryl TEG, desthiobiotin TEG, DNP TEG, DNP-X, DOTA, dT-biotin, dual biotin, PC biotin, psoralen C2, psoralen C6, TINA, 3'-DABCYL, black hole quencher 1, black hole quencher 2, DABCYL SE, dT-DABCYL, IRDye QC-1, QSY-21, QSY-35, QSY-7, QSY-9, carboxyl linker, thiol linker, 2'-deoxyribonucleoside analog purine, 2'-deoxyribonucleoside analog pyrimidine, ribonucleoside analog, 2'-O-methylribonucleoside analog, sugar-modified analog, wobble / universal base, fluorescent dye label, 2'-fluoro RNA, 2'-O-methyl RNA, methyl phosphonate, phosphodiester DNA, phosphodiester RNA, phosphorothioate DNA, phosphorothioate RNA, UNA, pseudouridine-5'-triphosphate, 5'-methylcytidine-5'-triphosphate, or any combination thereof.

[0318] In some cases, the modification is permanent. In other cases, the modification is transient. In some cases, multiple modifications are made to the gRNA or guide polynucleotide. gRNA or guide polynucleotide modifications can change the physicochemical properties of the nucleotides such as their conformation, polarity, hydrophobicity, chemical reactivity, base pair interactions, or any combination thereof.

[0319] The modification can also be a phosphorothioate substitute. In some cases, the native phosphodiester bond may be susceptible to rapid degradation by cellular nucleases; modification of the internucleotide bond using a phosphorothioate (PS) bond substitute may be more stable to hydrolysis by cellular degradation. The modification can enhance the stability of the gRNA or guide polynucleotide. The modification can also enhance biological activity. In some cases, an RNA gRNA fortified with phosphorothioate can inhibit RNase A, RNase T1, bovine serum nuclease, or any combination thereof. Due to these properties, PS-RNA gRNAs can be used in applications where exposure to nucleases occurs with high probability in vivo or in vitro. For example, a phosphorothioate (PS) bond can be introduced between the last 3-5 nucleotides at the 5'- or 3'-end of the gRNA, which can inhibit exonuclease degradation. In some cases, phosphorothioate bonds can be added throughout the gRNA to reduce endonuclease attack.

[0320] Protospacer Adjacent Motif The term "protospacer adjacent motif (PAM)" or PAM-like motif refers to a 2-6 base pair DNA sequence immediately following the DNA sequence targeted by the Cas9 nuclease of the CRISPR bacterial adaptive immune system. In some embodiments, the PAM can be a 5' PAM (i.e., located upstream of the 5' end of the protospacer). In other embodiments, the PAM can be a 3' PAM (i.e., located downstream of the 5' end of the protospacer).

[0321] The PAM sequence is essential for target binding, but the exact sequence varies depending on the type of Cas protein. The PAM sequence can be any PAM sequence known in the art. Suitable PAM sequences include, but are not limited to, NGG, NGA, NGC, NGN, NGT, NGTT, NGCG, NGAG, NGAN, NGNG, NGCN, NGCG, NGTN, NNGRRT, NNNRRT, NNGRR(N), TTTV, TYCV, TYCV, TATV, NNNNGATT, NNAGAAW, or NAAAAC. Y is a pyrimidine; N is any nucleotide base; and W is A or T.

[0322] The base editors provided herein may include a domain derived from a CRISPR protein that can bind to a nucleotide sequence containing a standard or non-standard protospacer adjacent motif (PAM) sequence. The PAM site is a nucleotide sequence adjacent to the target polynucleotide sequence. Some aspects of the present disclosure provide base editors that include all or part of a CRISPR protein having different PAM specificities. For example, typically, the Cas9 protein, e.g., Cas9 from S. pyogenes (spCas9), requires a standard NGG PAM sequence to bind to a specific nucleic acid region, where the "N" in "NGG" is adenine (A), thymine (T), guanine (G), or cytosine (C), and G is guanine. The PAM is specific to the CRISPR protein and can vary between different base editors that include different domains derived from CRISPR proteins. The PAM can be located 5' or 3' of the target sequence. The PAM can be upstream or downstream of the target sequence. The PAM can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides or longer. In many cases, the length of the PAM is 2-6 nucleotides.

[0323] In some embodiments, for example, as described in R. T. Walton et al., 2020, Science, 10.1126 / science.aba8853 (2020) (the entire content of which is incorporated herein by reference), the PAM is a "NRN" PAM, where the "N" of "NRN" is adenine (A), thymine (T), guanine (G), or cytosine (C), and R is adenine (A) or guanine (G); or the PAM is a "NYN" PAM, where the "N" of NYN is adenine (A), thymine (T), guanine (G), or cytosine (C), and Y is cytosine (C) or thymine (T). Some PAM variants are shown in Table 1E.

Table 1-5

[0324] In some embodiments, the PAM is NGC. In some embodiments, the NGC PAM is recognized by a Cas9 variant, such as a SpCas9 variant. In some embodiments, the NGC PAM variant includes one or more amino acid substitutions selected from D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (collectively referred to as "MQKFRAER").

[0325] In some embodiments, the PAM is NGT. In some embodiments, the NGT PAM is recognized by a Cas9 variant. In some embodiments, the NGT PAM variant is generated by a target mutation at one or more of residues 1335, 1337, 1135, 1136, 1218, and / or 1219. In some embodiments, the NGT PAM variant is created by a target mutation at one or more of residues 1219, 1335, 1337, 1218. In some embodiments, the NGT PAM variant is created by a target mutation at one or more of residues 1135, 1136, 1218, 1219, and 1335. In some embodiments, the NGT PAM variant is selected from the sets of target mutations shown in Tables 2 and 3 below. [Table 2] [Table 3]

[0326] In some embodiments, the NGT PAM variant is selected from variants 5, 7, 28, 31, or 36 of Tables 2 and 3. In some embodiments, the variant has improved NGT PAM recognition.

[0327] In some embodiments, the NGT PAM variant has a mutation at residue 1219, 1335, 1337, and / or 1218. In some embodiments, the NGT PAM variant is selected using mutations to improve recognition from the variants shown in Table 4 below. [Table 4]

[0328] In some embodiments, the NGT PAM is selected from the variants shown in Table 5 below. [Table 5]

[0329] In some embodiments, the NGTN variant is Variant 1. In some embodiments, the NGTN variant is Variant 2. In some embodiments, the NGTN variant is Variant 3. In some embodiments, the NGTN variant is Variant 4. In some embodiments, the NGTN variant is Variant 5. In some embodiments, the NGTN variant is Variant 6.

[0330] In some embodiments, the Cas9 domain is a Cas9 domain derived from Streptococcus pyogenes (SpCas9). In some embodiments, the SpCas9 domain is nuclease-active SpCas9, nuclease-inactive SpCas9 (SpCas9d), or SpCas9 nickase (SpCas9n). In some embodiments, SpCas9 comprises a D9X mutation, or a corresponding mutation in any of the amino acid sequences provided herein, where X is any amino acid other than D. In some embodiments, SpCas9 comprises a D9A mutation, or a corresponding mutation in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 domain, the SpCas9d domain, or the SpCas9n domain can bind to a nucleic acid sequence having a non-standard PAM. In some embodiments, the SpCas9 domain, the SpCas9d domain, or the SpCas9n domain can bind to a nucleic acid sequence having an NGG, NGA, or NGCG PAM sequence.

[0331] In some embodiments, the SpCas9 domain comprises one or more of the D1135X, R1335X, and T1337X mutations, or corresponding mutations in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of the D1135E, R1335Q, and T1337R mutations, or corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 domain comprises the D1135E, R1335Q, and T1337R mutations, or corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 domain comprises one or more of the D1135X, R1335X, and T1337X mutations, or corresponding mutations in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of the D1135V, R1335Q, and T1337R mutations, or corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 domain comprises the D1135V, R1335Q, and T1337R mutations, or corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 domain comprises one or more of the D1135X, G1218X, R1335X, and T1337X mutations, or corresponding mutations in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of the D1135V, G1218R, R1335Q, and T1337R mutations, or corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 domain comprises the D1135V, G1218R, R1335Q, and T1337R mutations, or corresponding mutations in any of the amino acid sequences provided herein.

[0332] In some embodiments, any Cas9 domain of the fusion proteins provided herein comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the Cas9 polypeptides described herein. In some embodiments, any Cas9 domain of the fusion proteins provided herein comprises the amino acid sequence of any Cas9 polypeptide described herein. In some embodiments, any Cas9 domain of the fusion proteins provided herein consists of the amino acid sequence of any Cas9 polypeptide described herein.

[0333] In some examples, the PAM recognized by the CRISPR protein-derived domain of the base editors disclosed herein can be provided to cells on an oligonucleotide separate from the insert (e.g., AAV insert) encoding the base editor. In such embodiments, providing the PAM on a separate oligonucleotide can enable cleavage of target sequences that would not otherwise be cleavable because the PAM is not adjacent on the same polynucleotide as the target sequence.

[0334] In one embodiment, S. pyogenes Cas9 (SpCas9) can be used as a CRISPR endonuclease for genome modification. However, other ones can be used. In some embodiments, different endonucleases can be used to target specific genomic targets. In some embodiments, synthetic SpCas9-derived variants having non-NGG PAM sequences can be used. Furthermore, other Cas9 orthologs derived from various species have been identified, and these "non-SpCas9" can bind to various PAM sequences that may also be useful in the present disclosure. For example, the relatively large-sized SpCas9 (encoding sequence of approximately 4 kb) can result in a plasmid carrying an SpCas9 cDNA that cannot be efficiently expressed intracellularly. Conversely, the coding sequence of Staphylococcus aureus Cas9 (SaCas9) is approximately 1 kilobase shorter than SpCas9 and may be efficiently expressed intracellularly. Similar to SpCas9, the SaCas9 endonuclease can modify target genes in mammalian cells in vitro and in mice in vivo. In some embodiments, the Cas protein can target different PAM sequences. In some embodiments, the target gene can be adjacent to, for example, the Cas9 PAM, 5'-NGG. In other embodiments, other Cas9 orthologs can have different PAM requirements. For example, other PAMs such as those of S. thermophilus (5'-NNAGAA for CRISPR1 and 5'-NGGNG for CRISPR3) and Neisseria meningitidis (5'-NNNNGATT) can also be found adjacent to the target gene.

[0335] In some embodiments, in the case of S. pyogenes, the target gene sequence can precede the 5'-NGG PAM (i.e., to the 5' side), and the 20-nt guide RNA sequence can base pair with the opposite strand to mediate Cas9 cleavage adjacent to the PAM. In some embodiments, the adjacent cut can be 3 base pairs or about 3 base pairs upstream of the PAM. In some embodiments, the adjacent cut can be 10 base pairs or about 10 base pairs upstream of the PAM. In some embodiments, the adjacent cut can be 0 to 20 base pairs or about 0 to 20 base pairs upstream of the PAM. For example, the adjacent cut can be adjacent to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 base pairs upstream of the PAM. The adjacent cut can also be 1 to 30 base pairs downstream of the PAM. The sequence of an exemplary SpCas9 protein that can bind to the PAM sequence is as follows:

[0336] The amino acid sequence of an exemplary PAM-binding SpCas9 is as follows:

[0337] The amino acid sequence of exemplary PAM-binding SpCas9n is as follows:

[0338] The amino acid sequence of exemplary PAM-binding SpEQR Cas9 is as follows: JPEG0007717684000062.jpg168168 In this sequence, the residues E1135, Q1335, and R1337, which can be mutated from D1135, R1335, and T1337 to generate SpEQR Cas9, are underlined and shown in bold.

[0339] The amino acid sequence of exemplary PAM-binding SpVQR Cas9 is as follows: JPEG0007717684000063.jpg186169 In this sequence, the residues V1135, Q1335, and R1337, which can be mutated from D1135, R1335, and T1337 to generate SpVQR Cas9, are underlined and shown in bold.

[0340] The amino acid sequence of exemplary PAM-binding SpVRER Cas9 is as follows: JPEG0007717684000064.jpg177170 In the above sequence, the residues V1135, R1218, Q1335, and R1337, which can be mutated from D1134, G1218, R1335, and T1337 to generate SpVRER Cas9, are underlined and shown in bold.

[0341] In some embodiments, the modified SpCas9 variant can recognize a protospacer adjacent motif (PAM) sequence adjacent to 3’H (non-G PAM) (see Tables 1A-1E). In some embodiments, the SpCas9 variant recognizes an NRNH PAM (wherein in the sequence, R is A or G, and H is A, C, or T). In some embodiments, the non-G PAM is NRRH, NRTH, or NRCH (see, for example, Miller, S. M., et al. Continuous evolution of SpCas9 variants compatible with non-G PAMs, Nat. Biotechnol. (2020) (the entire content of which is incorporated herein by reference)).

[0342] In some embodiments, the Cas9 domain is a recombinant Cas9 domain. In some embodiments, the recombinant Cas9 domain is a SpyMacCas9 domain. In some embodiments, the SpyMacCas9 domain is nuclease-active SpyMacCas9, nuclease-inactive SpyMacCas9 (SpyMacCas9d), or SpyMacCas9 nickase (SpyMacCas9n). In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to a nucleic acid sequence having a non-standard PAM. In some embodiments, the SpyMacCas9 domain, SpCas9d domain, or SpCas9n domain can bind to a nucleic acid sequence having an NAA PAM sequence.

[0343] The sequence of an exemplary Spy Cas9 Cas9A homolog in Streptococcus macacae having a native 5’-NAAN-3’ PAM specificity is known in the art and is described, for example, in Jakimo et al., (www.biorxiv.org / content / biorxiv / early / 2018 / 09 / 27 / 429654.full.pdf) and is shown below. SpyMacCas9

[0344] In some cases, the variant Cas9 protein has the H840A, P475A, W476A, N477A, D1125A, W1126A, and D1218A mutations, whereby the polypeptide has a reduced ability to cleave target DNA or target RNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some cases, the variant Cas9 protein has the D10A, H840A, P475A, W476A, N477A, D1125A, W1126A, and D1218A mutations, whereby the polypeptide has a reduced ability to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). In some cases, when the variant Cas9 protein contains the W476A and W1126A mutations, or when the variant Cas9 protein contains the P475A, W476A, N477A, D1125A, W1126A, and D1218A mutations, the variant Cas9 protein does not efficiently bind to the PAM sequence. Thus, in some such cases, when using such a variant Cas9 protein in a binding method, the method does not require a PAM sequence. In other words, in some cases, when using such a variant Cas9 protein in a binding method, the method can include a guide RNA, but the method can be carried out without a PAM sequence (and the specificity of binding is thus provided by the targeting segment of the guide RNA). Other residues can be mutated (i.e., substituted) to achieve the above effects (i.e., inactivate one or the other nuclease moiety). As non-limiting examples, the residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 can be modified (i.e., substituted). Also, mutations other than alanine substitutions are suitable.

[0345] In some embodiments, the CRISPR protein-derived domain of the base editor may comprise all or part of a Cas9 protein having a standard PAM sequence (NGG). In other embodiments, the Cas9-derived domain of the base editor can use a non-standard PAM sequence. Such sequences are described in the art and will be apparent to those skilled in the art. For example, Cas9 domains that bind to non-standard PAM sequences are described in Kleinstiver, B. P., et al., “Engineered CRISPR-Cas9 nucleases with altered PAM specificities” Nature 523, 481-485 (2015); and Kleinstiver, B. P., et al., “Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition” Nature Biotechnology 33, 1293-1298 (2015); R. T. Walton et al. “Unconstrained genome targeting with near-PAMless engineered CRISPR-Cas9 variants” Science 10.1126 / science.aba8853 (2020); Hu et al. “Evolved Cas9 variants with broad PAM compatibility and high DNA specificity,” Nature, 2018 Apr. 5, 556 (7699), 57-63; S. Miller et al., “Co...

Claims

1. An adenosine deaminase comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 1, wherein said adenosine deaminase comprises the V82T modification of SEQ ID NO: 1, and SEQ ID NO: 1 has the following amino acid sequence: The adenosine deaminase is capable of deaminating adenine in deoxyribonucleic acid, adenosine deaminase.

2. The adenosine deaminase according to claim 1, further comprising a modification selected from the group consisting of R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, M94V, P124W, T133K, D139L, D139M, C146R, and A158K of SEQ ID NO:

1.

3. The adenosine deaminase according to claim 1 or 2, further comprising the Y147T, Y147R, Q154S, Y123H, and / or Q154R modification of SEQ ID NO:

1.

4. The adenosine deaminase, wherein said adenosine deaminase is a group of the following modifications of SEQ ID NO: 1: N38G+V82T+Y123H+Y147R+Q154R; N38G+V82T+Y123H+Y147R+Q154R; or I76Y+V82T+Y123H+Y147R+Q154R The adenosine deaminase according to any one of claims 1 to 3, comprising any one of the above.

5. The adenosine deaminase according to any one of claims 1 to 4, comprising a C-terminal deletion starting with a residue selected from the group consisting of 152, 153, 154, 155, 156, and 157 with respect to SEQ ID NO:

1.

6. A fusion protein comprising a napDNAbp domain capable of associating with a nucleic acid that directs a nucleic acid-programmable DNA-binding protein (napDNAbp) to a specific nucleic acid sequence, and at least one base editor domain comprising the adenosine deaminase according to any one of claims 1 to 5.

7. The fusion protein according to claim 6, wherein said base editor domain comprises an adenosine deaminase heterodimer comprising a wild-type adenosine deaminase domain and the adenosine deaminase according to any one of claims 1 to 5.

8. The fusion protein according to claim 6 or 7, wherein the adenosine deaminase lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 C-terminal amino acid residues as compared to SEQ ID NO:

1.

9. The fusion protein according to any one of claims 6 to 8, wherein the napDNAbp domain is a Cas9, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ domain.

10. The napDNAbp domain is Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), Streptococcus pyogenes Cas9 (SpCas9), or variants thereof, or a modified SaCas9 or SpCas9 having a modified protospacer adjacent motif (PAM) specificity. The fusion protein according to any one of claims 6 to 9.

11. The fusion protein according to any one of claims 6 to 10, wherein the napDNAbp domain is a nuclease-inactive variant or a nickase variant.

12. The fusion protein according to any one of claims 6 to 11, comprising a linker between the napDNAbp domain and the adenosine deaminase domain.

13. The fusion protein according to any one of claims 6 to 12, further comprising one or more nuclear localization signals.

14. A polynucleotide encoding the fusion protein according to any one of claims 6 to 13.

15. A base editor system comprising the fusion protein according to any one of claims 6 to 13.

16. The base editor system according to claim 15, further comprising one or more guide polynucleotides that target the base editor domain to effect a modification of an A·T to a G·C of an SNP associated with a genetic disease.

17. An in vitro or ex vivo method for correcting a single nucleotide polymorphism (SNP) in a polynucleotide, comprising: Contacting a target nucleotide sequence, at least a part of which is located in the polynucleotide or its reverse complementary strand, with the fusion protein according to any one of claims 6 to 13 or the base editor system according to claim 15 or 16; editing the SNP by deaminating the SNP or its complementary nucleobase when targeting the base editor to the target nucleotide sequence, wherein the deamination of the SNP or its complementary nucleobase corrects the SNP, said method.

18. The method according to claim 17, wherein the SNP is related to α-1 antitrypsin deficiency (A1AD), or the SNP is in the SERPINA1 gene and the correction includes the modification of E342K (PiZ allele).

19. An in vitro or ex vivo method for editing a polynucleotide, said method comprising contacting a target nucleotide sequence with the fusion protein according to any one of claims 6 to 13 or the base editor system according to claim 15 or 16, thereby editing the polynucleotide.

20. The method according to claim 19, wherein the editing results in less than 20% indel formation, less than 15% indel formation, less than 10% indel formation; less than 5% indel formation; less than 4% indel formation; less than 3% indel formation; less than 2% indel formation; less than 1% indel formation; less than 0.5% indel formation; or less than 0.1% indel formation, and / or the editing does not result in translocation.

21. A composition comprising the fusion protein according to any one of claims 6 to 13 or the base editor system according to claim 15 or 16.

22. The composition according to claim 21, further comprising a pharmaceutically acceptable excipient, diluent, or carrier.

23. The adenosine deaminase according to any one of claims 1 to 5, the fusion protein according to any one of claims 6 to 13, or the base editor system according to claim 15 or 16, wherein the adenosine deaminase comprises any one of the following amino acid modifications or groups of modifications of SEQ ID NO: 1: V82T; I76Y+V82T; or I76Y+V82T+Y147T+Q154S ​

Citation Information

Patent Citations

  • Adenosine nucleobase editors and uses thereof

    WO2018027078A1

Cited By

  • Novel nucleic acid base editor and method of use thereof

    JP2025170240A