Compositions and methods for treating glycogen storage disease type 1a

JP2025087691A5Active Publication Date: 2025-11-26BEAM THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025017262
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-01-27
Filing Date
2025-02-05
Publication Date
2025-11-26
Estimated Expiration
2040-02-13

AI Technical Summary

Technical Problem

Current genome editing techniques, such as CRISPR, are inefficient for correcting point mutations in genetic diseases like glycogen storage disease type 1a (GSD1a), often resulting in random insertions or deletions at the target genetic locus.

Method used

The use of programmable nucleobase editors, specifically adenosine deaminase base editor 8 (ABE8), to precisely correct single nucleotide polymorphisms in the G6PC gene associated with GSD1a, converting A·T to G·C base pairs.

Benefits of technology

This approach allows for precise and efficient correction of pathogenic amino acids in the G6PC gene, potentially offering a more effective treatment for GSD1a compared to existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025087691000001
    Figure 2025087691000001
Patent Text Reader

Abstract

To provide methods of using base editors comprising adenosine deaminase variants for altering mutations associated with Glycogen Storage Disease Type 1a (GSD1a).SOLUTION: The present invention provides a method of editing a glucose-6-phosphatase (G6PC) polynucleotide comprising a single nucleotide polymorphism (SNP) associated with Glycogen Storage Disease Type 1a (GSD1a). The method comprises contacting the G6PC polynucleotide with an adenosine deaminase base editor 8 (ABE8) in a complex with one or more guide polynucleotides, wherein the adenosine deaminase base editor 8 (ABE8) comprises a polynucleotide programmable DNA binding domain and an adenosine deaminase domain, and wherein the one or more of guide polynucleotides target the base editor to effect an A T to G C alteration of the SNP associated with the GSD1a.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a joint venture of U.S. Provisional Application No. 62 / 805,271, filed February 13, 2019, and U.S. Provisional Application No. 62 / 805,271, filed May 23, 2019. No. 62 / 852,228, filed May 23, 2019; No. 4, filed July 19, 2019; U.S. Provisional Application No. 62 / 876,354, filed October 9, 2019; No. 62 / 912,992, filed on November 6, 2019; U.S. Provisional Application No. 62 / 931,722, filed on November 6, 2019; No. 62 / 941,569, filed November 27, 2019, and U.S. Provisional Application No. 62 / 941,569, filed January 27, 2020. International PCT applications claiming priority to and the benefit of U.S. Provisional Application No. 62 / 966,526. and all of which are incorporated herein by reference in their entireties.

[0002] Incorporation by Reference All publications, patents, and patent applications mentioned herein are hereby incorporated by reference in their entirety. Any patents or patent applications are specifically and individually indicated to be incorporated by reference. All rights reserved. All publications, patents, and patent applications referred to herein are hereby incorporated by reference in their entirety. is incorporated into. [Background technology]

[0003] For most known genetic diseases, there is no research to determine or address the underlying cause of the disease. To address this issue, correction of point mutations at targeted loci is required, rather than stochastic disruption of genes. is required. Current genome editing techniques that utilize clustered regularly interspaced short palindromic repeat (CRISPR) systems introduce double-strand DNA breaks at target genetic loci as a first step for gene correction. In response to the double-strand DNA breaks, the intracellular DNA repair process typically results in random insertions or deletions (indels) at the DNA cleavage site by non-homologous end joining. Most genetic diseases are caused by point mutations, but current approaches for point mutation correction are inefficient and typically induce numerous random insertions and deletions (indels) at the target genetic locus due to the cellular response to dsDNA breaks. Therefore, there is a need for improved genome editing that is more efficient and produces far fewer undesirable products such as probabilistic insertions or deletions (indels) or translocations. Current genome editing techniques that utilize clustered regularly interspaced short palindromic repeat (CRISPR) systems introduce double-strand DNA breaks at target genetic loci as a first step for gene correction. In response to the double-strand DNA breaks, the intracellular DNA repair process typically results in random insertions or deletions (indels) at the DNA cleavage site by non-homologous end joining. Most genetic diseases are caused by point mutations, but current approaches for point mutation correction are inefficient and typically induce numerous random insertions and deletions (indels) at the target genetic locus due to the cellular response to dsDNA breaks. Therefore, there is a need for improved genome editing that is more efficient and produces far fewer undesirable products such as probabilistic insertions or deletions (indels) or translocations. Most genetic diseases are caused by point mutations, but current approaches for point mutation correction are inefficient and typically induce numerous random insertions and deletions (indels) at the target genetic locus due to the cellular response to dsDNA breaks. Therefore, there is a need for improved genome editing that is more efficient and produces far fewer undesirable products such as probabilistic insertions or deletions (indels) or translocations. Most genetic diseases are caused by point mutations, but current approaches for point mutation correction are inefficient and typically induce numerous random insertions and deletions (indels) at the target genetic locus due to the cellular response to dsDNA breaks.

[0004] Glycogen storage disease type 1 (GSD1, also known as Von Gierke disease) is a hereditary disorder that causes deficiencies in glycogenolysis and gluconeogenesis, resulting in the accumulation of glycogen and lipids in tissues, life-threatening hypoglycemia and lactic acidosis, and potential CNS damage as well as long-term liver and kidney complications (e.g., steatosis, hepatic adenomas, and hepatocellular carcinoma). Glycogen storage disease type 1 (GSD1, also known as Von Gierke disease) is a hereditary disorder that causes deficiencies in glycogenolysis and gluconeogenesis, resulting in the accumulation of glycogen and lipids in tissues, life-threatening hypoglycemia and lactic acidosis, and potential CNS damage as well as long-term liver and kidney complications (e.g., steatosis, hepatic adenomas, and hepatocellular carcinoma). Glycogen storage disease type 1 (GSD1, also known as Von Gierke disease) is a hereditary disorder that causes deficiencies in glycogenolysis and gluconeogenesis, resulting in the accumulation of glycogen and lipids in tissues, life-threatening hypoglycemia and lactic acidosis, and potential CNS damage

[0005] There are two types of GSD1, type 1a (GSD1a) and type 1b (GSD1b), which are caused by different gene mutations. GSD1a is caused by a mutation in the glucose-6-phosphatase (G6PC) gene and accounts for approximately 80% of GSD1 patients. There are two types of GSD1, type 1a (GSD1a) and type 1b (GSD1b), which are caused by different gene mutations. GSD1a is caused by a mutation in the glucose-6-phosphatase (G6PC) gene and accounts for approximately 80% of GSD1 patients. There are two types of GSD1, type 1a (GSD1a) and type 1b (GSD1b), which are caused by different gene mutations. GSD1a is caused by a mutation in the glucose-6-phosphatase (G6PC) gene and accounts for approximately 80% of GSD1 patients. In the country, approximately 1 in 100,000 newborns has GSD1a, and about 22% of patients have the recessive mutation Q3 47*, and 37% of patients have the recessive mutation R83C.

[0006] There is no approved drug treatment for GSD1a. Liver transplantation is curative but there is no approved treatment method, and the current treatment regimen includes almost continuous intake of corn starch. If not treated chronically, patients may develop severe lactic acidosis, progress to renal failure, and may die in infancy or childhood. GSD1a is an area of important medical need that has not yet been addressed. Therefore, new compositions and methods for treating patients with GSD 1a are needed. 1a.

[0007] Incorporation by reference All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the same extent as if each individual publication, patent, or patent application were specifically and individually indicated to be incorporated by reference. Unless otherwise specified, the publications, patents, and patent applications mentioned in this specification are incorporated herein by reference in their entirety. SUMMARY OF THE INVENTION

[0008] The present invention features compositions and methods for the precise correction of pathogenic amino acids using programmable nucleobase editors. In particular, the compositions and methods of the present invention are useful for the treatment of glycogenosis type 1a (GSD1a). Therefore, the present invention uses adenosine to precisely correct single nucleotide polymorphisms in the endogenous G6PC gene to correct harmful mutations (e.g., Q347X, R83C). ​​​​ (A) Compositions and methods for treating GSD1a using a base editor (ABE) (e.g., ABE8).

[0009] In one aspect, the present invention is a method of editing a G 6PC polynucleotide comprising a single nucleotide polymorphism (SNP) associated with glycogen storage disease type 1a (GSD1a), the method comprising contacting the G6PC polynucleotide with an adenosine deaminase base editor 8 ( ABE) complexed with one or more guide polynucleotides, wherein ABE8 comprises a polynucleotide-programmable DNA binding domain and an adenosine deaminase domain, and one or more of the guide polynucleotides target the base editor to effect a modification of the A·T of the SNP associated with GSD1a to G·C. In another aspect, the present invention provides a cell comprising a polynucleotide-programmable DNA binding domain, and an adenosine deaminase domain-containing adenosine deaminase base editor 8 (ABE8), or a polynucleotide encoding said base editor, and one or more guide polynucleotides that target the base editor to effect a modification of the A·T of the SNP associated with GSD1a to G·C. In another aspect, the present invention is a method of treating GSD1a in a subject, the method comprising administering to the subject an adenosine deaminase base editor 8 (ABE8) comprising a polynucleotide-programmable DNA binding domain and an adenosine deaminase domain, or a polynucleotide encoding said base editor, and an adenosine deaminase base one or more of which target the base editor to effect a modification of the A·T of the SNP associated with GSD1a to G·C. In another aspect, the present invention provides a cell comprising a polynucleotide-programmable DNA binding domain, and an adenosine deaminase domain-containing adenosine deaminase base editor 8 (ABE8), or a polynucleotide encoding said base editor, and one or more guide polynucleotides that target the base editor to effect a modification of the A·T of the SNP associated with GSD1a to G·C. In another aspect, the present invention is a method of treating GSD1a in a subject, the method comprising administering to the subject an adenosine deaminase base editor 8 (ABE8) comprising a polynucleotide-programmable DNA binding domain and an adenosine deaminase domain, or a polynucleotide encoding said base editor, and an adenosine deaminase base editor 8 (ABE8) complexed with one or more guide polynucleotides that target the base editor to effect a modification of the A·T of the SNP associated with GSD1a to G·C. In another aspect, the present invention provides a cell comprising a polynucleotide-programmable DNA binding domain, and an adenosine deaminase domain-containing adenosine deaminase base editor 8 (ABE8), or a polynucleotide encoding said base editor, and one or more guide polynucleotides that target the base editor to effect a modification of the A·T of the SNP associated with GSD1a to G·C. In another aspect, the present invention is a method of treating GSD1a in a subject, the method comprising administering to the subject an adenosine deaminase base editor 8 (ABE8) comprising a polynucleotide-programmable DNA binding domain and an adenosine deaminase domain, or a polynucleotide encoding said base editor, and an adenosine deaminase base editor 8 (ABE8) complexed with one or more guide polynucleotides that target the base editor to effect a modification of the A·T of the SNP associated with GSD1a to G·C. In another aspect, the present invention provides a cell comprising a polynucleotide-programmable DNA binding domain, and an adenosine deaminase domain-containing adenosine deaminase base editor 8 (ABE8), or a polynucleotide encoding said base editor, and one or more guide polynucleotides that target the base editor to effect a modification of the A·T of the SNP associated with GSD1a to G·C. In another aspect, the present invention is a method of treating GSD1a in a subject, the method comprising administering to the subject an adenosine deaminase base editor 8 (ABE8) comprising a polynucleotide-programmable DNA binding domain and an adenosine deaminase domain, or a polynucleotide encoding said base editor, and an adenosine deaminase base editor 8 (ABE8) complexed with one or more guide polynucleotides that target the base editor to effect a modification of the A·T of the SNP associated with GSD1a to G·C. In another aspect, the present invention provides a cell comprising a polynucleotide-programmable DNA binding domain, and an adenosine deaminase domain-containing adenosine deaminase base editor 8 (ABE8), or a polynucleotide encoding said base editor, and one or more guide polynucleotides that target the base editor to effect a modification of the A·T of the SNP associated with GSD1a to G·C. In another aspect, the present invention is a method of treating GSD1a in a subject, the method comprising administering to the subject an adenosine deaminase base editor 8 (ABE8) comprising a polynucleotide-programmable DNA binding domain and an adenosine deaminase domain, or a polynucleotide encoding said base editor, and an adenosine deaminase base editor 8 (ABE8) complexed with one or more guide polynucleotides that target the base editor to effect a modification of the A·T of the SNP associated with GSD1a to G·C. In another aspect, the present invention provides a cell comprising a polynucleotide-programmable DNA binding domain, and an adenosine deaminase domain-containing adenosine deaminase base editor 8 (ABE8), or a polynucleotide encoding said base editor, and one or more guide polynucleotides that target the base editor to effect a modification of the A·T of the SNP associated with GSD1a to G·C. In another aspect, the present invention is a method of treating GSD1a in a subject, the method comprising administering to the subject an adenosine deaminase base editor 8 (ABE8) comprising a polynucleotide-programmable DNA binding domain and an adenosine deaminase domain, or a polynucleotide encoding said base editor, and an adenosine deaminase base editor 8 (ABE8) complexed with one or more Targeting deaminase 8 (ABE8) and administering one or more guide polynucleotides that also effect a modification of the A·T to G·C of the SNP associated with GSD1a A method is provided that includes . In another aspect, the present invention is a method of generating hepatocytes or precursors thereof, comprising: a) introducing into induced pluripotent stem cells or hepatocyte precursors containing an SNP associated with GSD1a, a polynucleotide programmable nucleotide binding domain and an adenosine deaminase domain Adenosine deaminase base editor 8 (ABE8) containing, or a polynucleotide encoding adenosine deaminase base editor 8 (ABE8), and one or more guide polynucleotides that target the base editor to effect a modification of the A·T to G·C of the SNP associated with GSD1a And b) differentiating the induced pluripotent stem cells or hepatocyte precursors into hepatocytes . A method is provided that includes . In one aspect, the present invention is a method of editing a glucose-6-phosphatase (G6PC) polynucleotide containing a single nucleotide polymorphism (SNP) associated with glycogen storage disease type 1a (GSD1a), the method comprising contacting the G6PC polynucleotide with an adenosine deaminase base editor 8 (ABE8) complexed with one or more guide polynucleotides, wherein the adenosine deaminase base editor 8 (ABE8) contains an adenosine deaminase variant domain inserted within a Cas9 or Cas12 polypeptide, and one or more of the guide polynucleotides target the base editor to effect a modification of the A·T to G·C of the SNP associated with GSD1a

[0010] In one aspect, the present invention is a ​​​​​​​​​Provided is a method that brings about a change. In another aspect, the present invention is a method for treating glycogen storage disease type 1a (GSD1a) in a subject, comprising administering to the subject an adenosine deaminase base editor containing an adenosine deaminase variant inserted into a Cas9 or Cas12 polypeptide 8 (ABE8), or a polynucleotide encoding the base editor, and one or more guide polynucleotides that target the adenosine deaminase base editor 8 (ABE8) to effect a change from A·T to G·C of a SNP associated with GSD1a, thereby treating GSD1a in the subject. In yet another aspect, the present invention is a method for treating glycogen storage disease type 1a (GSD1a) in a subject, comprising administering to the subject a fusion protein containing an adenosine deaminase variant inserted into a Cas9 or Cas 12 polypeptide, or a polynucleotide encoding the fusion protein, and one or more guide polynucleotides that target the fusion protein to effect a change from A·T to G·C of a single nucleotide polymorphism (SNP) associated with GSD1a, thereby treating GSD1a in the subject. In one aspect, the present invention provides a pharmaceutical composition for the treatment of glycogen storage disease type 1a (GSD1a) comprising an effective amount of an adenosine deaminase base editor 8 (ABE8) comprising a polynucleotide-programmable DNA binding domain

[0011] and an adenosine deaminase variant domain. In some embodiments, the pharmaceutical composition targets the adenosine deaminase base editor 8 (ABE8) to effect a change from A·T to G·C of a SNP associated with GSD1a and thereby treat GSD1a. In some embodiments, the pharmaceutical composition targets the adenosine deaminase base editor 8 (ABE8) to effect a change from A·T to G·C of a SNP associated with GSD1a comprises one or more guide polynucleotides capable of doing so. In another aspect, the present invention provides a pharmaceutical composition for the treatment of glycogen storage disease type 1a (GSD1a), comprising an effective amount of any of the cells provided herein. In some embodiments, the pharmaceutical composition comprises a pharmaceutically acceptable excipient.

[0012] In another aspect, the present invention is a kit for the treatment of glycogen storage disease type 1a (GSD1a), comprising an adenosine deaminase base editor 8 (ABE8) comprising a polynucleotide programmable DNA binding domain and an adenosine deaminase domain, and one or more guide polynucleotides capable of targeting the adenosine deaminase base editor 8 (ABE8) to effect the modification of the A·T to G·C of the SNP associated with GSD1a. In yet another aspect, the present invention is a kit for the treatment of glycogen storage disease type 1a (GSD1a), comprising a kit comprising any of the cells provided herein. In some embodiments, the contacting occurs in a cell, eukaryotic cell, mammalian cell, or human cell. In some embodiments, the cell is in vivo. In some embodiments, the cell is ex vivo. In some embodiments, the cell is a hepatocyte, a hepatocyte precursor, or an iPSC-derived hepatocyte. In some

[0013] embodiments, the cell expresses the G6PC polypeptide. In some embodiments, the cell or hepatocyte precursor is derived from a subject having GSD1a. In some embodiments, In one form, the subject is a mammal or a human. In some embodiments, this hepatocyte or hepatocyte precursor is a mammalian cell or a human cell. In some embodiments In one form, the adenosine deaminase base editor 8 (ABE8) or a polynucleotide encoding the adenosine deaminase base editor 8 (ABE8) and the one or more guide polynucleotides are delivered to the cells of the subject.

[0014] In various embodiments of the above aspects or any other aspect of the invention described herein the SNP associated with GSD1a is located in the glucose-6-phosphatase (G6PC) gene. In one embodiment, the change from A·T to G·C at the SNP associated with glycogen storage disease type 1a (GSD1a) results in a change from glutamine (Q) to a non-glutamine (X) amino acid. In one embodiment the change from A·T to G·C at the SNP associated with glycogen storage disease type 1a (GSD1a) results in a change from arginine (R) to a non-arginine (X) in the G6PC polypeptide. In one embodiment the SNP associated with GSD1a results in the expression of a G6PC polypeptide having a non-glutamine (X) amino acid at position 347 or a non-arginine (X) amino acid at position 83. In one embodiment base editor correction replaces the non-glutamine amino acid (X) at position 347 with glutamine. In another embodiment base editor correction replaces the non-arginine amino acid (X) at position 83 with arginine. In one embodiment the change from A·T to G·C at the SNP associated with GSD1a terminates prematurely at amino acid position 347 or at position 83 In another embodiment, In one embodiment, the change from A·T to G·C at the SNP associated with GSD1a results in premature termination at amino acid position 347 or at position 83 Results in the expression of the G6PC polypeptide encoding cysteine. In some embodiments the modification at the SNP is one or more of Q347X and / or R83C.

[0015] In various embodiments of the above aspects or any other aspect of the invention described herein the adenosine deaminase variant is inserted within a flexible loop, alpha helix region, unstructured portion, or solvent-accessible portion of a Cas9 or Cas12 polypeptide In some embodiments, the adenosine deaminase variant is flanked by N-terminal and C-terminal fragments of a Cas9 or Cas12 polypeptide. In some embodiments the fusion protein or adenosine deaminase base editor 8 (ABE8) has the structure NH -[N-terminal fragment of a Cas9 or Cas12 polypeptide]-[adenosine deaminase variant ant]-[C-terminal fragment of a Cas9 or Cas12 polypeptide]-COOH, where each instance of "]-[ " 2 is an optional linker. In one embodiment, the C-terminus of the N-terminal fragment or the N-terminus of the C-terminal fragment includes a portion of a flexible loop of a Cas9 or Cas12 polypeptide. In one embodiment this flexible loop includes amino acids proximal to the target nucleobase. In some embodiments one or more guide polynucleotides direct the fusion protein or adenosine deaminase base editor 8 (ABE8) to effect deamination of a target nucleobase. In some embodiments deamination of the SNP target nucleobase results in replacement of the target nucleobase with a non-wild-type nucleobase and deamination of the target nucleobase results in symptoms of GSD1a In some embodiments, and deamination of the target nucleobase results in symptoms of GSD1a. In some embodiments, deamination of the SNP target nucleobase causes the target nucleobase to be replaced with a non-wild-type nucleobase, and deamination of the target nucleobase causes the symptoms of GSD1a to be alleviated. is improved. In one embodiment, the target nucleobase is 1 to 20 nucleobases away from the PAM sequence in the target polynucleotide sequence. In one embodiment, the target nucleobase is upstream of 2 to 12 nucleobases of the PAM sequence.

[0016] In one embodiment, the N-terminal fragment or C-terminal fragment of the Cas polypeptide or Cas12 polypeptide binds to the target polynucleotide sequence. In one embodiment, the N-terminal fragment or C-terminal fragment contains the RuvC domain, or the N-terminal fragment or C-terminal fragment contains the HNH domain, or neither the N-terminal nor the C-terminal fragment contains the HNH domain, or neither the N-terminal nor the C-terminal fragment contains the RuvC domain. In one embodiment, the Cas9 or Cas12 polypeptide contains partial or complete deletions in one or more structural domains, and the deaminase is inserted at the position of the partial or complete deletion of the Cas9 or Cas12 polypeptide. In one embodiment, this deletion is within the RuvC domain, this deletion is within the HNH domain, or this deletion bridges the RuvC domain and the C-terminal domain, the L-I domain and the HNH domain, or the RuvC domain and the L-I domain. In various embodiments, the polynucleotide-programmable DNA-binding domain is a Cas 9 polypeptide. In some embodiments, the fusion protein or adenosine

[0017] deaminase base editor 8 (ABE8) contains an adenosine deaminase variant domain inserted into the Cas9 polypeptide. In some embodiments, the Cas9 polypeptide is In some embodiments, the fusion protein or adenosine deaminase base editor 8 (ABE8) contains an adenosine deaminase variant domain inserted into the Cas9 polypeptide. In some embodiments, the Cas9 polypeptide contains partial or complete deletions in one or more structural domains, and the deaminase is inserted at the position of the partial or complete deletion of the Cas9 polypeptide. In one embodiment, this deletion is within the RuvC domain, this deletion is within the HNH domain, or this deletion bridges the RuvC domain and the C-terminal domain, the L-I domain and the HNH domain, or the RuvC domain and the L-I domain. , Streptococcus pyogenes Cas9 (SpCas9), Staphylococcus aureus Cas9 (SaCas9), Str eptococcus thermophilus 1 Cas9 (St1Cas9), or variants thereof. In some embodiments, the Cas9 polypeptide has the following amino acid sequence (Cas9 reference sequence): JPEG2025087691000001.jpg171164 (single underline: HNH domain; double underline: RuvC domain; (Cac9 reference sequence), or its corresponding region.

[0018] In some embodiments, the Cas9 polypeptide contains a deletion of amino acids 1017 - 1069 or their corresponding amino acids numbered in the Cas9 polypeptide reference sequence, or the Cas9 po lypeptide contains a deletion of amino acids 792 - 872 or their corresponding amino acids numbered in the Cas9 polypeptide reference sequence, or the Cas9 polypeptide contains a deletion of amino acids 792 - 906 or their corresponding amino acids numbered in the Cas9 polypeptide ref erence sequence. In some embodiments, the adenosine deaminase variant is inserted within a flexible loop of the Cas9 polypeptide. In some embodiments, this flexible loop contains a region selected from the group consisting of amino acid residues at positions 530 - 537, 569 - 579, 686 - 691, 768 - 793, 94 3 - 947, 1002 - 1040, 1052 - 1077, 1232 - 1248, and 1298 - 1300, numbered in the Cas9 reference sequence, or their corresponding amino acid positions. In some embodiments, the deaminase is at amino acid positions 768 - 7 numbered in the Cas9 reference sequence. is numbered in the Cas9 reference sequence at positions 530 - 537, 569 - 579, 686 - 691, 768 - 793, 94 3 - 947, 1002 - 1040, 1052 - 1077, 1232 - 1248, and 1298 - 1300, or their corresponding amino acid positions. In some embodiments, the deaminase is at amino acid positions 768 - 7 numbered in the Cas9 reference sequence. 69, 791 - 792, 792 - 793, 1015 - 1016, 1022 - 1023, 1026 - 1027, 1029 - 1030, 1040 - 10 41, 1052 - 1053, 1054 - 1055, 1067 - 1068, 1068 - 1069, 1247 - 1248, or 1248 - 12 49, or inserted between these corresponding amino acid positions. In some embodiments the deaminase is inserted between amino acid positions 768 - 769, 792 - 793, 1022 - 1023, 1026 - 1027, 1040 - 1041, 1068 - 1069, or 1247 - 1248, as numbered in the Cas9 reference sequence, or between these corresponding amino acid positions. In some embodiments, the deaminase is inserted between amino acid positions 1016 - 1017, 1023 - 1024, 1029 - 1030, 1040 - 1041, 1069 - 1070, or 1247 - 1248, or between these corresponding amino acid positions. In some embodiments, the adenosine deaminase variant is inserted into the Cas9 polypeptide at the loci identified in Table 10A. In some embodiments, the N - terminal fragment comprises amino acid residues 1 - 529, 538 - 568, 580 - 685, 692 - 942, 948 - 1001, 1026 - 1051, 1078 - 1231, and / or 1248 - 1297, or their corresponding residues in the Cas9 reference sequence. In some embodiments, the C - terminal fragment comprises amino acid residues 1301 - 1368, 1248 - 1297, 1078 - 1231, 1026 - 1051, 948 - 1001, 692 - 942, 580 - 685, and / or 538 - 568, or their corresponding residues.

[0019] In some embodiments, the Cas9 polypeptide is a nickase or the Cas9 polypeptide is nuclease-inactive. In some embodiments, the Cas9 polypeptide is a modified SpCas9 and has specificity for a modified PAM or specificity for a non-G PAM. In some embodiments, this modified SpCas9 polypeptide contains the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (SpCas9-MQKFRAER) and has specificity for the modified

[0020] PAM 5'-NGC-3'. In various embodiments, the polynucleotide-programmable DNA binding domain is a modified Streptococcus pyogenes Cas9 (SpCas9) or a variant thereof. In various embodiments of any other aspect of the invention described herein, the polynucleotide-programmable DNA binding domain comprises a modified SpCas9 having modified protospacer adjacent motif (PAM) specificity or specificity for a non-G PAM. In one embodiment, this modified SpCas9 has specificity for the nucleic acid sequence 5'-NGA-3'. In one embodiment, this modified SpCas9 has specificity for the nucleic acid sequence 5'-AGA-3' or 5'-GGA-3'. In one

[0021] embodiment, this modified SpCas9 has specificity for an NGA PAM variant. It is Staphylococcus aureus Cas9 (SaCas9) or a variant thereof. In one embodiment, this SaCas9 has specificity for the nucleic acid sequence 5'-NNGRRT-3'. In one embodiment, this SaCas9 has specificity for the nucleic acid sequence 5'-GAGAAT-3'. In one embodiment, this SaCas9 has specificity for the NNGRRT PAM variant.

[0022] In various embodiments, the polynucleotide programmable DNA binding domain is Cas 12 polypeptide. In one embodiment, the adenosine deaminase variant is Ca s12 polypeptide is inserted. In one embodiment, the Cas12 polypeptide is Cas12a , Cas12b, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, or Cas12i. In one embodiment it is inserted between the following amino acid positions: a) 153-154, 255-256, 306-307, 980-981, 1019-1020, 534-535, 604-605, or 344-345 of BhCas12b, or the corresponding amino acid residues of Cas12a, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h , or Cas12i; b) 147 and 148, 248 and 249, 299 and 300, 991 and 992, or 1031 and 1032 of BvCas12b, or the corresponding amino acid residues of Cas12a, Cas12c, Cas12d , Cas12e, Cas12g, Cas12h, or Cas12i; or c) 157 and 158, 258 and 259, 310 and 311, 1008 and 1009, or 1044 and 10 of AaCas12b 31, or the corresponding amino acid residues of Cas12a, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, or Cas12i; or c) 157 and 158, 258 and 259, 310 and 311, 1008 and 1009, or 1044 and 10 45, or a corresponding amino acid residue of Cas12a, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, or Cas12i. In one embodiment, the adenosine deaminase variant is inserted into the Cas12 polypeptide at the locus identified in Table 10B. In one embodiment, the Cas12 polypeptide is Cas12b. In one embodiment, the Cas12 polypeptide comprises a BhCas12b domain, a BvCas12b domain, or an AACas12b domain. In various embodiments, the polynucleotide programmable DNA binding domain is a nuclease-inactive variant. In other embodiments, the polynucleotide programmable DNA binding domain is a nickase variant. In one embodiment, this nickase variant comprises the amino acid substitution D10A or its corresponding amino acid substitution. In some embodiments, the adenosine deaminase domain can deaminate adenosine in deoxyribonucleic acid (DNA). In some embodiments, the adenosine deaminase domain is a monomer comprising an adenosine deaminase variant. In some embodiments, the adenosine deaminase domain is a heterodimer comprising a wild-type adenosine deaminase domain and an adenosine deaminase variant. In some embodiments, the adenosine deaminase variant has the amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL

[0023] In some embodiments, the polynucleotide programmable DNA binding domain is a nuclease-inactive variant. In other embodiments, the polynucleotide programmable DNA binding domain is a nickase variant. In one embodiment, this nickase variant comprises the amino acid substitution D10A or its corresponding amino acid substitution. In some embodiments, the adenosine deaminase domain can deaminate adenosine in deoxyribonucleic acid (DNA). In some embodiments, the adenosine deaminase domain is a monomer comprising an adenosine deaminase variant. In some embodiments, the adenosine deaminase domain is a heterodimer comprising a wild-type adenosine deaminase domain and an adenosine deaminase variant. In some embodiments, the adenosine deaminase variant has the amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL In some embodiments, the adenosine deaminase domain can deaminate adenosine in deoxyribonucleic acid (DNA). In some embodiments, the adenosine deaminase domain is a monomer comprising an adenosine deaminase variant. In some embodiments, the adenosine deaminase domain is a heterodimer comprising a wild-type adenosine deaminase domain and an adenosine deaminase variant. In some embodiments, the adenosine deaminase variant has the amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL In some embodiments, the adenosine deaminase domain can deaminate adenosine in deoxyribonucleic acid (DNA). In some embodiments, the adenosine deaminase domain is a monomer comprising an adenosine deaminase variant. In some embodiments, the adenosine deaminase domain is a heterodimer comprising a wild-type adenosine deaminase domain and an adenosine deaminase variant. In some embodiments, the adenosine deaminase variant has the amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL

[0024] In some embodiments, the adenosine deaminase variant has the amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD comprising, and this amino acid sequence comprises at least one modification. In some embodiments the adenosine deaminase variant comprises a modification at amino acid positions 82 and / or 166 relative to the above sequence. In some embodiments, this at least one modification is V82S, Y147T, Y147R, Q154S, Y123H, and / or Q154R relative to the above sequence. In some embodiments, this at least one modification is Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V8 2S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R + Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T 166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R. In some embodiments, this at least one modification is Y147T + Q154S relative to the above sequence.

[0025] In some embodiments, the adenosine deaminase variant comprises a C-terminal deletion starting with a residue selected from the group consisting of 149, 150, 151, 152, 153, 154, 155, 156, and 157. In some embodiments, the adenosine deaminase variant is TadA An adenosine deaminase monomer containing an *8 adenosine deaminase variant domain There is. In some embodiments, the adenosine deaminase variant is a wild-type adeno An adenosine deaminase heterodimer containing a nosine deaminase domain and a TadA*8 adenosine deaminase variant domain There is. In some embodiments, the adenosine An adenosine deaminase heterodimer containing a TadA domain and a TadA*8 adenosine deaminase variant domain Domain. In some embodiments, the guide polynucleotide is a) GACCUAGGCGAGGCAGUAGG; b) CCAGUAUGGACACUGUCCAAA; c) CAGUAUGGACACUGUCCAAA; and d) AGUAUGGACACUGUCCAAAG. Contains a nucleic acid sequence selected from the group of

[0026] In some embodiments, one or more guide RNAs include CRISPR RNA (crRNA) and also And trans-encoded small RNA (tracrRNA), and this crRNA contains a SNP associated with GSD1a and is a G6P Contains a nucleic acid sequence complementary to the C nucleic acid sequence. In some embodiments, adenosine deaminase Base editor 8 (ABE8) forms a complex with a single guide RNA (sgRNA) containing a nucleic acid sequence complementary to the G6PC nucleic acid sequence containing the SNP associated with GSD1a. Acid sequence. In some embodiments, the adenosine deaminase is a TadA deaminase. In one embodiment, the TadA deaminase is a TadA*8 variant. In some embodiments In this state, the TadA*8 variant is selected from the group consisting of TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5 , TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA* 8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.2 0, TadA*8.21, TadA*8.22, TadA*8.23, TadA*8.24. In some embodiments, the adenosine deaminase base editor 8 (ABE8) is ABE8.1-m, ABE 8.2-m, ABE8.3-m, ABE8.4-m, ABE8.5-m, ABE8.6-m, ABE8.7-m, ABE8.8-m, ABE8.9-m, ABE 8.10-m, ABE8.11-m, ABE8.12-m, ABE8.13-m, ABE8.14-m, ABE8.15-m, ABE8.16-m, ABE8.1 7-m, ABE8.18-m, ABE8.19-m, ABE8.20-m, ABE8.21-m, ABE8.22-m, ABE8.23-m, ABE8.24-m , ABE8.1-d, ABE8.2-d, ABE8.3-d, ABE8.4-d, ABE8.5-d, ABE8.6-d, ABE8.7-d, ABE8.8-d , ABE8.9-d, ABE8.10-d, ABE8.11-d, ABE8.12-d, ABE8.13-d, ABE8.14-d, ABE8.15-d, AB E8.16-d, ABE8.17-d, ABE8.18-d, ABE8.19-d, ABE8.20-d, ABE8.21-d, ABE8.22-d, ABE8. 23-d, or ABE8.24-d.

[0027] In some embodiments, adenosine deaminase base editor 8 (ABE8) is an adenosine having the following sequence or a fragment thereof having adenosine deaminase activity: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQK KAQSSTD comprising or consisting essentially of.

[0028] In some embodiments, the gRNA has the following sequence: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU and comprises a scaffold having.

[0029] In some embodiments, the gRNA has the following sequence: GUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGC GAGAUUUU and comprises a scaffold having.

[0030] In one aspect, provided herein is a base editor comprising adenosine deaminase base editor 8 (ABE8) complexed with one or more guide polynucleotides, wherein the adenosine deaminase base editor 8 (ABE8) comprises a polynucleotide programmable DNA binding domain and an adenosine deaminase domain, and one or more of the guide polynucleotides target the base editor wherein the adenosine deaminase base editor 8 (ABE8) comprises a polynucleotide programmable DNA binding domain and an adenosine deaminase domain, and one or more of the guide polynucleotides target the base editor A base editor that results in the modification of A·T to G·C of SNPs related to GSD1a. In some embodiments, the adenosine deaminase variant includes the V82S modification and / or the T166R modification. In some embodiments, the adenosine deaminase variant further includes one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, and Q154R. In some embodiments, the base editor domain includes an adenosine deaminase heterodimer comprising a wild-type adenosine deaminase domain and an adenosine deaminase variant. In some embodiments, the adenosine deaminase variant is a truncated TadA8 lacking 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to full-length TadA8. In some embodiments, the adenosine deaminase variant is a truncated TadA8 lacking 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to full-length TadA8. In some embodiments, the polynucleotide-programmable DNA binding domain is a modified Staphylococcus aure us Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), a modified Streptococcus pyogenes Cas9 (SpCas9), or variants thereof. In some embodiments, the polynucleotide-programmable DNA binding domain is a modified protospacer In some embodiments, the adenosine deaminase variant is a truncated TadA8 lacking 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to full-length TadA8. In some embodiments, the adenosine deaminase variant is a truncated TadA8 lacking 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to full-length TadA8. In some embodiments, the polynucleotide-programmable DNA binding domain is a modified Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), a modified Streptococcus pyogenes Cas9 (SpCas9), or variants thereof. In some embodiments, the polynucleotide-programmable DNA binding domain is a modified protospacer relative to full-length TadA8. In some embodiments, the adenosine deaminase variant is a truncated TadA8 lacking 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to full-length TadA8. In some embodiments, the polynucleotide-programmable DNA binding domain is a modified Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), a modified Streptococcus pyogenes Cas9 (SpCas9), or variants thereof. In some embodiments, the polynucleotide-programmable DNA binding domain is a modified protospacer In some embodiments, the polynucleotide-programmable DNA binding domain is a modified Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), a modified Streptococcus pyogenes Cas9 (SpCas9), or variants thereof. In some embodiments, the polynucleotide-programmable DNA binding domain is a modified protospacer In some embodiments, the polynucleotide-programmable DNA binding domain is a modified Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), a modified Streptococcus pyogenes Cas9 (SpCas9), or variants thereof. In some embodiments, the polynucleotide-programmable DNA binding domain is a modified protospacer In some embodiments, the polynucleotide-programmable DNA binding domain is a modified Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), a modified Streptococcus pyogenes Cas9 (SpCas9), or variants thereof. In some embodiments, the polynucleotide-programmable DNA binding domain is a modified protospacer In some embodiments, the polynucleotide-programmable DNA binding domain is a modified Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), a modified Streptococcus pyogenes Cas9 (SpCas9), or variants thereof. In some embodiments, the polynucleotide-programmable DNA binding domain is a modified protospacer In some embodiments, the polynucleotide-programmable DNA binding domain is a modified Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), a modified Streptococcus pyogenes Cas9 (SpCas9), or variants thereof. In some embodiments, the polynucleotide-programmable DNA binding domain is a modified protospacer Variants of SpCas9 that have adjacent motif (PAM) specificity or specificity for non-G PAMs. In some embodiments, the polynucleotide-programmable DNA binding domain is nuclease-inactive Cas9. In some embodiments, the po lynucleotide-programmable DNA binding domain is a Cas9 nickase.

[0031] In one aspect, provided herein are one or more guide RNAs, as well as the fo llowing sequences: EIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSK ESILPKRNSDKLIARKKDWDPKKYGGFMQPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGY KEVKKDLIIKLPKYSLFELENGRKRMLASAKFLQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLD EIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPRAFKYFDTTIARKEYRSTKEVLDATL IHQSITGLYETRIDLSQLGGDGGSGGSGGSGGSGGSGGSGGMDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTD RHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIF GNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEEN PINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLL AQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGY AGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREK IEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFT VYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLK IIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTI LDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENI VIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVD HIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKR QLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIK KYPKLESEFVYGDYKVYDVRKMIAKSEQEGADKRTADGSEFESPKKKRKV* (Here, the boldface sequences indicate sequences derived from Cas9, the italicized sequences indicate linker sequences, and the underlined sequences indicate bipartite nuclear localization sequences) comprising a polynucleotide-programmable DNA binding domain, and MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD an adenosine deaminase variant comprising a modification at amino acid positions 82 and / or 166 of and at least one base editor domain, and a fusion protein comprising the base editor system wherein one or more of the guide polynucleotides target the base editor to effect a modification from A·T to G·C of an SNP associated with GSD1a, a cis system.

[0032] In one aspect, a cell comprising any one of the base editor systems described above is provided. In some embodiments, the cell is a human cell or a mammalian cell. In some embodiments, the cell is ex vivo, in vivo, or in vitro.

[0033] The description and examples herein detail embodiments of the present disclosure. It should be understood that the present disclosure is not limited to the specific embodiments described herein and can therefore vary. Those skilled in the art will recognize that there are numerous variations and will recognize that.

[0034] The practice of several embodiments disclosed herein, unless otherwise indicated, is within the skill of those in the art and uses conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA. For example, see Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (F.M. Ausubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (M.J. MacPherson, B .D. Hames and G.R. Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (R.I. Freshney, ed. (2010)). See also the series Current Protocols in Molecular Biology (F.M. Ausubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (M.J. MacPherson, B .D. Hames and G.R. Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (R.I. Freshney, ed. (2010)). A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (R.I. Freshney, ed. (2010)). See also the series Current Protocols in Molecular Biology (F.M. Ausubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (M.J. MacPherson, B .D. Hames and G.R. Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies,

[0035] Although the various features of the present disclosure may be described in the context of a single embodiment, these features may also be provided separately or in any suitable combination. Conversely, although the present disclosure may be described in the context of separate embodiments for clarity, the present disclosure may also be in a single embodiment or in any suitable combination. Conversely, although the present disclosure may be described in the context of separate embodiments for clarity, the present disclosure may also be in a single embodiment or in any suitable combination. Conversely, although the present disclosure may be described in the context of separate embodiments for clarity, the present disclosure may also be in a single embodiment may also be implemented. The section headings used in this specification are for the sole purpose of summarization and should not be construed as limiting the subject matter being described.

[0036] The features of the present disclosure are particularly set forth in the appended claims. By referring to the following detailed description which shows exemplary embodiments in which the principles of the present disclosure are utilized, and considering the accompanying drawings described below, a better understanding of the features and advantages of the present disclosure will be obtained.

[0037] Definitions The following definitions supplement those of the art and are for the purposes of the present application and do not pertain to related or unrelated cases, such as patents or applications under common ownership. Any methods and materials similar or equivalent to those described herein can be used in the practice of the tests of the present disclosure, but the preferred materials and methods are described herein. Accordingly, the terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting. Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. The following references provide to one of ordinary skill in the art a general definition of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of

[0038] Science and Technology (Singleton et al., eds., 1993); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). Singleton et al., eds., 1993); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). n et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Science and Technology (Singleton et al., eds., 1993); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). f Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991).

[0039] In this application, the use of the singular form includes the plural form unless otherwise specified. As used herein, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. It should be noted that in this application, the use of "or" means "and / or" and is understood to be inclusive unless otherwise stated. Further, the term "including", as well as other forms (e.g., "include", "includes", and "included"), is used in a non-limiting sense. (For example, "include", "includes", and "included")[[]] The use thereof is non-limiting.

[0040] As used in this specification and the claims, the term "comprising" (and any form thereof such as "comprise" and "comprises"), "having" (any form thereof such as "have" and "has"), "including" (any form thereof such as "include" and "includes"), or "containing" (any form thereof such as "contains" and "contain") is inclusive or open-ended and does not exclude additional, unrecited elements or method steps. Any embodiment discussed herein is any method or composition of the present disclosure is inclusive or open-ended and does not exclude additional, unrecited elements or method steps. Any embodiment discussed herein is any method or composition of the present disclosure can be carried out with respect to, and vice versa, it is considered to be the same. Further, the compositions of the present disclosure can be used to achieve the methods of the present disclosure.

[0041] The term "about" or "approximately" means, as would be determined by one of ordinary skill in the art, to be within an acceptable error range for a particular value, which depends in part on how the value is measured or determined, i.e., on the limitations of the measurement system. For example, "about" can mean within one standard deviation or more than one standard deviation, according to practices in the relevant art. Alternatively, "about" can mean within up to 20%, up to 10%, up to 5%, or up to 1% of a given value. As another alternative, particularly with respect to biological systems or processes, the term can mean a value within the same order of magnitude, e.g., within five-fold or within two-fold of the value. When a particular value is recited in the claims and the specification, unless otherwise stated, the term "about" is presumed to mean that the value is within an acceptable error range for that particular value.

[0042] The ranges provided herein are to be understood as all shorthand representations of values within that range. For example, a range of 1 to 50 is understood to include any number, combination of numbers, or subrange from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.

[0043] References to "some embodiments", "an embodiment", "one embodiment", or "other embodiments" in the specification mean that the specific features, structures, or characteristics described in connection with that embodiment are included in at least some embodiments of the present disclosure, but not necessarily all

[0044] "Adenosine deaminase" means a polypeptide or a fragment thereof that can catalyze the hydrolytic deamination of adenine or adenosine. In certain aspects, the de aminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine to inosine, or deoxyadenosine to deoxyinosine. In certain aspects, adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., genetically engineered adenosine deaminases, evolved adenosine deaminases) can be from any organism such as bacteria.

[0045] In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is a TadA variant. In some embodiments, the TadA variant is TadA*8. In some embodiments, the de aminase or deaminase domain is a variant of a naturally occurring deaminase from an organism (e.g., human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse). In some embodiments, the deaminase or deaminase domain does not exist naturally. For example, in some embodiments, the deaminase or de aminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% , at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical. For example, the deaminase domain is described in International PCT Application No. PCT / 2017 / 045381 (International Publication No. WO 2018 / 027078) and PCT / US2016 / 058344 (International Publication No. WO 2017 / 070632), each of which is hereby incorporated by reference in its entirety. Similarly, Ko mor, A.C., et al., “Programmable editing of a target base in genomic DNA withou t double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, N.M., et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA clea vage” Nature 551, 464-471 (2017); Komor, A.C., et al., “Improved base excision repair by adenine base editing” Nature 568, 413-419 (2019); and Gaudelli, N.M., et al., “Cytosine base editing converts C·G to T·A in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); each of these references is hereby incorporated by reference in its entirety. Repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base edito rs with higher efficiency and product purity” Science Advances 3:eaao4774 (2017 ), and Rees, H.A., et al., “Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec;19(12):770-788. d oi: 10.1038 / s41576-018-0059-1 (the entire contents of which are incorporated herein by reference ; see).

[0046] Wild-type TadA (wt) adenosine deaminase has the following sequence (also referred to as the TadA reference sequence) : MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQK KAQSSTD and has.

[0047] In some embodiments, the adenosine deaminase has the following sequence: MSEVEFSHEY WMRHALTLAK RARDEREVPV GAVLVLNNRV IGEGWNRAIG LHDPTAHAEI MALRQGGLVM QNY RLIDATL YVTFEPCVMC AGAMIHSRIG RVVFGVRNAK TGAAGSLMDV LHYPGMNHRV EITEGILADE CAALLC YFFR MPRQVFNAQK KAQSSTD including modifications therein (also referred to as TadA*7.10).

[0048] In some embodiments, TadA*7.10 includes at least one modification. In some embodiments, TadA*7.10 includes modifications at amino acids 82 and / or 166. In certain embodiments, variants of the sequences referred to above include one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. The modification Y1 23H is also referred to herein as H123H (the modification H123Y in TadA*7.10 reverted to Y123H (wt)). In other embodiments, variants of the TadA*7.10 sequence include Y147T+Q154R;Y147T+Q154S;Y147R+Q154S;V82S+Q154S;V82S+Y147R;V82S+Q154R;V82S+Y123H; I76Y+V82S;V82S+Y123H+Y147T;V82S+Y123H+Y147R;V82S+Y123H+Q154R;Y147R+Q154R+Y123H;Y 147R+Q154R+I76Y;Y147R+Q154R+T166R;Y123H+Y147R+Q154R+I76Y;V82S+Y123H+Y147R+Q154R; and combinations of modifications selected from the group of I76Y+V82S+Y123H+Y147R+Q154R .

[0049] In other embodiments, the invention provides pairs in TadA*7.10, the TadA reference sequence, or another TadA beginning at residue 149, 150, 151, 152, 153, 154, 155, 156, or 157 in response to a corresponding mutation An adenosine deaminase variant comprising a deletion that includes a complete C-terminal deletion (e.g., TadA*8). In other embodiments, the adenosine deaminase variant is a TadA (e.g., TadA*8) monomer that includes one or more of the following modifications : Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R relative to the corresponding mutation in TadA*7.10, the TadA reference sequence, or a corresponding mutation in another TadA. In other embodiments, the adenosine deaminase variant is a monomer having a combination of modifications selected from the group of Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y 147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y + V82S + Y123H + Y147R + Q154R relative to the corresponding mutation in TadA*7.10, the TadA reference sequence, or a corresponding mutation in another TadA. In yet other embodiments, the adenosine deaminase variant is a monomer having one or more of the following modifications Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R relative to the corresponding mutation in TadA*7.10, the TadA reference

[0050] sequence, or a corresponding mutation in another TadA, respectively. In still other embodiments, the adenosine deaminase variant is a dimer having one or more of the following modifications Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R relative to the corresponding mutation in TadA*7.10, the TadA reference It is a homodimer containing one adenosine deaminase domain (e.g., TadA*8). In other embodiments, the adenosine deaminase variant is for the corresponding mutation in TadA*7.10, the TadA reference sequence, or another TadA Y147T+Q154R;Y147T+Q154S;Y147R+Q154S;V82S+Q154S;V82S+Y147R;V82S+Q154R;V82S+Y123H; I76Y+V82S;V82S+Y123H+Y147T;V82S+Y123H+Y147R;V82S+Y123H+Q154R;Y147R+Q154R+Y123H;Y 147R+Q154R+I76Y;Y147R+Q154R+T166R;Y123H+Y147R+Q154R+I76Y;V82S+Y123H+Y147R+Q154R; and is a homodimer containing two adenosine deaminase domains (e.g., TadA*8) each having a combination of modifications selected from the group of I76Y + V82S + Y123H + Y147R + Q154R respectively.

[0051] In other embodiments, the adenosine deaminase variant is a heterodimer comprising a wild-type TadA adenosine deaminase domain and an adenosine deaminase variant domain (e.g., TadA*8) containing one or more of the following modifications Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R for the corresponding mutation in TadA*7.10, the TadA reference sequence, or another TadA In other embodiments, the adenosine deaminase variant is a heterodimer comprising a wild-type TadA adenosine deaminase domain and an adenosine deaminase variant domain (e.g., TadA*8) containing one or more of the following modifications Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R for the corresponding mutation in TadA*7.10, the TadA reference sequence, or another TadA ​​​​​​Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R + Y123H; Y 147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R, a heterodimer comprising a combination of modifications selected from the group of for example TadA*8).

[0052] In other embodiments, the adenosine deaminase variant is a heterodimer comprising a TadA*7.10 domain and an adenosine deaminase variant domain (e.g., TadA*8) comprising one or more of the following modifications Y 147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R relative to the corresponding mutation in TadA*7.10, TadA*7.10, the TadA reference sequence, or another TadA In other embodiments, the adenosine deaminase variant is a heterodimer comprising a TadA*7.10 domain and an adenosine deaminase variant domain (e.g., TadA*8) comprising one or more of the following modifications relative to the corresponding mutation in TadA*7.10, TadA*7.10, the TadA reference sequence, or another TadA : The following modifications: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R + Y123H; Y 147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; or an adenosine deaminase variant comprising a combination of I76Y + V82S + Y123H + Y147R + Q154R is a heterodimer comprising a catalytic domain (e.g., TadA*8).

[0053] In one embodiment, the adenosine deaminase has the following sequence or a fragment thereof having adenosine deaminase activity: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQK KAQSSTD and is TadA*8 comprising or consisting essentially of the foregoing.

[0054] In some embodiments, TadA*8 is a truncated form. In some embodiments, the truncated TadA*8 lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 1 3, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to full-length TadA*8. In some embodiments, the truncated TadA*8 lacks 1, 2, 3, 4, 5, 6, 7 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to full-length TadA*8. In some embodiments, the adenosine deaminase variant is full-length TadA*8.

[0055] ​​​ In certain embodiments, the adenosine deaminase heterodimer comprises a TadA*8 domain and an adenosine deaminase domain selected from one of the following: Staphylococcus aureus (S. aureus) TadA: MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTL YVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN Bacillus subtilis (B. subtilis) TadA: MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTL EPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLS E Salmonella typhimurium (S. typhimurium) TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLV LQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFF RMRRQEIKALKKADRAEGAGPAV Shewanella putrefaciens (S. putrefaciens) TadA: MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEP CAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQ QGIE Haemophilus influenzae F3031 (H. influenzae) TadA: MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTΑΗAEIIALRNGAKNIQNYR LLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRRE EKKIEKALLKSLSDK Caulobacter crescentus (C. crescentus) TadA: MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLT DLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAK I Geobacter sulfurreducens (G. sulfurreducens) TadA: MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLT GATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRK KAKATPALFIDERKVPPEP TadA*7.10 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD

[0056] "Adenosine deaminase base editor 8 (ABE8) polypeptide" or "ABE8" refers to the following reference sequences: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD and includes an adenosine deaminase variant that contains modifications at amino acid positions 82 and / or 166 of the following: which means a base editor as defined herein. In some embodiments, ABE 8 includes further modifications as described herein relative to the reference sequence.

[0057] "Adenosine deaminase base editor 8 (ABE8) polynucleotide" means a polynucleotide that encodes ABE8.

[0058] "Administering" is referred to herein as providing one or more of the compositions described herein to a patient or subject. By way of example, but not limitation, administration of a composition ​​, for example, injection can be performed by intravenous (i.v.) injection, subcutaneous (s.c.) injection, intradermal (i.d.) injection, intraperitoneal ( i.p.) injection, or intramuscular (i.m.) injection. One or more such routes can be used. Parenteral administration can be performed, for example, by bolus injection or by perfusion over time. Alternatively, or simultaneously, administration can be by the oral route.

[0059] "Agent" means any small molecule compound, antibody, nucleic acid molecule, or polypeptide, or fragment thereof.

[0060] "Modification" means a change in the structure, expression level, or activity of a gene or polypeptide, such as can be detected by known methods of standard techniques as described herein (e.g., an increase or decrease). As used herein, modification includes a change in the sequence of a polynucleotide or polypeptide, or a change in the expression level (e.g., a 25% change, a 40%

[0061] change, a 50% change, or a higher change). "Improve" means a decrease, suppression, attenuation, reduction, arrest, or

[0062] stabilization of the occurrence or progression of a disease. "Analog" means a molecule that has similar functional or structural characteristics but is not identical. For example, a polynucleotide analog or polypeptide analog retains the biological activity of the corresponding have certain modifications. Such modifications can, for example, increase the affinity, efficiency, specificity, protease resistance or nuclease resistance, membrane permeability, and / or half-life of the analog without changing ligand binding. The analog can contain unnatural nucleotides or amino acids. , the affinity, efficiency, specificity, protease resistance or nuclease resistance, membrane permeability, and / or half-life for the analog resistance, membrane permeability, and / or half-life of the analog can be increased. The analog can contain unnatural nucleotides or amino acids.

[0063] "Base editor (BE)" or "nucleic acid base editor (NBE)" means an agent that binds to a polynucleotide and has nucleic acid base modification activity. In various embodiments, the base editor includes a nucleic acid base modifying polypeptide (e.g., deaminase) and a polynucleotide programmable nucleotide binding domain together with a guide polynucleotide (e.g., guide RNA). In various embodiments, the agent is a biomolecular complex that includes a protein domain having base editing activity, i.e., a domain that can modify bases (e.g., A, T, C, G, U) within a nucleic acid molecule (e.g., DNA). In some embodiments the polynucleotide programmable DNA binding domain is fused or linked to a deaminase domain. In one embodiment, the agent is a fusion protein that includes a domain having base editing activity . In another aspect, the protein domain having base editing activity is linked to a guide RNA (e.g., via an RNA binding motif on the guide RNA and an RNA binding domain fused to a deaminase ). In one aspect, the domain having base editing activity can deaminate bases within a nucleic acid molecule. In one aspect the base editor can deaminate one or more bases within a DNA molecule. In one aspect, the polynucleotide programmable DNA binding domain is fused or linked to a deaminase domain. In one embodiment, the agent is a fusion protein that includes a domain having base editing activity . In another aspect, the protein domain having base editing activity is linked to a guide RNA (e.g., via an RNA binding motif on the guide RNA and an RNA binding domain fused to a deaminase ). In one aspect, the domain having base editing activity can deaminate bases within a nucleic acid molecule. In one aspect the domain having base editing activity is linked to a guide RNA (e.g., via an RNA binding motif on the guide RNA and an RNA binding domain fused to a deaminase ). In one aspect, the domain having base editing activity can deaminate bases within a nucleic acid molecule. In one aspect the domain having base editing activity can deaminate bases within a nucleic acid molecule. In one aspect the base editor can deaminate one or more bases within a DNA molecule. In one In some embodiments, the base editor can deaminate adenosine (A). In one embodiment, the base editor is an adenosine base editor (ABE).

[0064] In some embodiments, the base editor is generated by cloning a circular permutant of Cas9 (e.g., spCAS9 or saCAS9) and a bipartite nuclear localization sequence to a scaffold containing an adenosine deaminase variant (e.g., TadA*8) (e.g., ABE8). Circular permutant Cas9 is known in the art and is described, for example, in Oakes et al., Cell 176, 254-267, 2019. Exemplary circular permutants are shown below, where the bold sequence indicates the sequence derived from Cas9, the italicized sequence indicates the linker sequence, and the underlined sequence indicates the bipartite nuclear localization sequence. CP5 (and MSP “NGC = Pam Variant with mutations Regular Cas9 likes NGG” PID = Protein Interaction Domain, and “D10A” Nickase):

[0065] In some embodiments, ABE8 is selected from the base editors in Table 7 or Table 9 below. In some embodiments, ABE8 contains an adenosine deaminase variant evolved from TadA. In some embodiments, the adenosine deaminase variant of ABE8 is the TadA*8 variant described in Table 7 or Table 9 below. In some embodiments, the adenosine deaminase variant is Y147T, Y147R, Q154S, Y12 One or more of the modifications selected from the group consisting of 3H, V82S, T166R, and / or Q154R A TadA*7.10 variant (e.g., TadA*8) that contains. In various embodiments, ABE8 is Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R + Y123H; Y 147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; And a TadA*7.10 variant (e.g., TadA*8) having a combination of modifications selected from the group consisting of I76Y + V82S + Y123H + Y147R + Q154R. In some embodiments, ABE8 is a mo Nomer construct. In some embodiments, ABE8 is a heterodimer construct . In some embodiments, adenosine deaminase base editor 8 (ABE8) has the sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQK KAQSSTD Including. KAQSSTD Including.

[0066] In one aspect, the polynucleotide programmable DNA binding domain is CRISP It is an R-related (e.g., Cas or Cpf1) enzyme. In certain embodiments, the base editor is catalytically dead Cas9 (dCas9) fused to a deaminase domain. In certain embodiments, the base editor is Cas9 nickase (nCas9) fused to a deaminase domain. Details of the base editors are described in International PCT Application No. PCT / 2017 / 045381 (International Publication No. WO 2018 / 027078) and PCT / US 2016 / 058344 (International Publication No. WO 2017 / 070632), each of which is incorporated herein by reference in its entirety. Komor, A.C., et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, N.M., et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, A.C., et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao4774 (2017), and Rees, H.A., et al., “Base It is catalytically dead Cas9 (dCas9) fused to a deaminase domain. In certain embodiments, the base editor is Cas9 nickase (nCas9) fused to a deaminase domain. Details of the base editors are described in International PCT Application No. PCT / 2017 / 045381 (International Publication No. WO 2018 / 027078) and PCT / US 2016 / 058344 (International Publication No. WO 2017 / 070632), each of which is incorporated herein by reference in its entirety. Komor, A.C., et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, N.M., et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, A.C., et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao4774 (2017), and Rees, H.A., et al., “Base editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, N.M., et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); ature 533, 420-424 (2016); Gaudelli, N.M., et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); ature 533, 420-424 (2016); Gaudelli, N.M., et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, A.C., et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao4774 (2017), and Rees, H.A., et al., “Base editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, N.M., et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); uct purity” Science Advances 3:eaao4774 (2017), and Rees, H.A., et al., “Base editing: precision chemistry on the genome and transcriptome of living cells.” See also Nat Rev Genet. 2018 Dec;19(12):770-788. doi: 10.1038 / s41576-018-0059-1 which is hereby incorporated by reference in its entirety).

[0067] For example, the adenine base editors (ABEs) used in the base editing compositions, systems, and methods described herein have the nucleic acid sequence (8,877 base pairs) provided below (Addgene, Watertown, MA.; Gaudelli NM, et al., Nature. 2017 Nov 23;551(7681):464-4 71. doi: 10.1038 / nature24644; Koblan LW, et al., Nat Biotechnol. 2018 Oct;36(9 ):843-846. doi: 10.1038 / nbt.4172.). Polynucleotide sequences having at least 95% identity to the ABE nucleic acid sequence are also included. ATATGCCAAGTACGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTATGCCCAGTACAT GACCTTATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATGGTGATGCGG TTTTGGCAGTACATCAATGGGCGTGGATAGCGGTTTGACTCACGGGGATTTCCAAGTCTCCACCCCATTG ACGTCAATGGGAGTTTGTTTTGGCACCAAAATCAACGGGACTTTCCAAAATGTCGTAACAACTCCGCCCC ​​ATTGACGCAAATGGGCGGTAGGCGTGTACGGTGGGAGGTCTATATAAGCAGAGCTGGTTTAGTGAACCGT CAGATCCGCTAGAGATCCGCGGCCGCTAATACGACTCACTATAGGGAGAGCCGCCACCATGAAACGGACA GCCGACGGAAGCGAGTTCGAGTCACCAAAGAAGAAGCGGAAAGTCTCTGAAGTCGAGTTTAGCCACGAGT ATTGGATGAGGCACGCACTGACCCTGGCAAAGCGAGCATGGGATGAAAGAGAAGTCCCCGTGGGCGCCGT GCTGGTGCACAACAATAGAGTGATCGGAGAGGGATGGAACAGGCCAATCGGCCGCCACGACCCTACCGCA CACGCAGAGATCATGGCACTGAGGCAGGGAGGCCTGGTCATGCAGAATTACCGCCTGATCGATGCCACCC TGTATGTGACACTGGAGCCATGCGTGATGTGCGCAGGAGCAATGATCCACAGCAGGATCGGAAGAGTGGT GTTCGGAGCACGGGACGCCAAGACCGGCGCAGCAGGCTCCCTGATGGATGTGCTGCACCACCCCGGCATG AACCACCGGGTGGAGATCACAGAGGGAATCCTGGCAGACGAGTGCGCCGCCCTGCTGAGCGATTTCTTTA GAATGCGGAGACAGGAGATCAAGGCCCAGAAGAAGGCACAGAGCTCCACCGACTCTGGAGGATCTAGCGG AGGATCCTCTGGAAGCGAGACACCAGGCACAAGCGAGTCCGCCACACCAGAGAGCTCCGGCGGCTCCTCC GGAGGATCCTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGCCAAGAGGG CACGCGATGAGAGGGAGGTGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGCGAGGGCTG GAACAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATGGCCCTGAGACAGGGCGGCCTG GTCATGCAGAACTACAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCG GCGCCATGATCCACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACGCAAAAACCGGCGCCGCAGG CTCCCTGATGGACGTGCTGCACTACCCCGGCATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCA GATGAATGTGCCGCCCTGCTGTGCTATTTCTTTCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAAGG CCCAGAGCTCCACCGACTCCGGAGGATCTAGCGGAGGCTCCTCTGGCTCTGAGACACCTGGCACAAGCGA GAGCGCAACACCTGAAAGCAGCGGGGGCAGCAGCGGGGGGTCAGACAAGAAGTACAGCATCGGCCTGGCC ATCGGCACCAACTCTGTGGGCTGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGG TGCTGGGCAACACCGACCGGCACAGCATCAAGAAGAACCTGATCGGAGCCCTGCTGTTCGACAGCGGCGA AACAGCCGAGGCCACCCGGCTGAAGAGAACCGCCAGAAGAAGATACACCAGACGGAAGAACCGGATCTGC TATCTGCAAGAGATCTTCAGCAACGAGATGGCCAAGGTGGACGACAGCTTCTTCCACAGACTGGAAGAGT CCTTCCTGGTGGAAGAGGATAAGAAGCACGAGCGGCACCCCATCTTCGGCAACATCGTGGACGAGGTGGC CTACCACGAGAAGTACCCCACCATCTACCACCTGAGAAAGAAACTGGTGGACAGCACCGACAAGGCCGAC CTGCGGCTGATCTATCTGGCCCTGGCCCACATGATCAAGTTCCGGGGCCACTTCCTGATCGAGGGCGACC TGAACCCCGACAACAGCGACGTGGACAAGCTGTTCATCCAGCTGGTGCAGACCTACAACCAGCTGTTCGA GGAAAACCCCATCAACGCCAGCGGCGTGGACGCCAAGGCCATCCTGTCTGCCAGACTGAGCAAGAGCAGA CGGCTGGAAAATCTGATCGCCCAGCTGCCCGGCGAGAAGAAGAATGGCCTGTTCGGAAACCTGATTGCCC TGAGCCTGGGCCTGACCCCCAACTTCAAGAGCAACTTCGACCTGGCCGAGGATGCCAAACTGCAGCTGAG CAAGGACACCTACGACGACGACCTGGACAACCTGCTGGCCCAGATCGGCGACCAGTACGCCGACCTGTTT CTGGCCGCCAAGAACCTGTCCGACGCCATCCTGCTGAGCGACATCCTGAGAGTGAACACCGAGATCACCA AGGCCCCCCTGAGCGCCTCTATGATCAAGAGATACGACGAGCACCACCAGGACCTGACCCTGCTGAAAGC TCTCGTGCGGCAGCAGCTGCCTGAGAAGTACAAAGAGATTTTCTTCGACCAGAGCAAGAACGGCTACGCC GGCTACATTGACGGCGGAGCCAGCCAGGAAGAGTTCTACAAGTTCATCAAGCCCATCCTGGAAAAGATGG ACGGCACCGAGGAACTGCTCGTGAAGCTGAACAGAGAGGACCTGCTGCGGAAGCAGCGGACCTTCGACAA CGGCAGCATCCCCCACCAGATCCACCTGGGAGAGCTGCACGCCATTCTGCGGCGGCAGGAAGATTTTTAC CCATTCCTGAAGGACAACCGGGAAAAGATCGAGAAGATCCTGACCTTCCGCATCCCCTACTACGTGGGCC CTCTGGCCAGGGGAAACAGCAGATTCGCCTGGATGACCAGAAAGAGCGAGGAAACCATCACCCCCTGGAA CTTCGAGGAAGTGGTGGACAAGGGCGCTTCCGCCCAGAGCTTCATCGAGCGGATGACCAACTTCGATAAG AACCTGCCCAACGAGAAGGTGCTGCCCAAGCACAGCCTGCTGTACGAGTACTTCACCGTGTATAACGAGC TGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAAGCCCGCCTTCCTGAGCGGCGAGCAGAAAAAGGC CATCGTGGACCTGCTGTTCAAGACCAACCGGAAAGTGACCGTGAAGCAGCTGAAAGAGGACTACTTCAAG AAAATCGAGTGCTTCGACTCCGTGGAAATCTCCGGCGTGGAAGATCGGTTCAACGCCTCCCTGGGCACAT ACCACGATCTGCTGAAAATTATCAAGGACAAGGACTTCCTGGACAATGAGGAAAACGAGGACATTCTGGA AGATATCGTGCTGACCCTGACACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTGAAAACCTATGCC CACCTGTTCGACGACAAAGTGATGAAGCAGCTGAAGCGGCGGAGATACACCGGCTGGGGCAGGCTGAGCC GGAAGCTGATCAACGGCATCCGGGACAAGCAGTCCGGCAAGACAATCCTGGATTTCCTGAAGTCCGACGG CTTCGCCAACAGAAACTTCATGCAGCTGATCCACGACGACAGCCTGACCTTTAAAGAGGACATCCAGAAA GCCCAGGTGTCCGGCCAGGGCGATAGCCTGCACGAGCACATTGCCAATCTGGCCGGCAGCCCCGCCATTA AGAAGGGCATCCTGCAGACAGTGAAGGTGGTGGACGAGCTCGTGAAAGTGATGGGCCGGCACAAGCCCGA GAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCACCCAGAAGGGACAGAAGAACAGCCGCGAGAGA ATGAAGCGGATCGAAGAGGGCATCAAAGAGCTGGGCAGCCAGATCCTGAAAGAACACCCCGTGGAAAACA CCCAGCTGCAGAACGAGAAGCTGTACCTGTACTACCTGCAGAATGGGCGGGATATGTACGTGGACCAGGA ACTGGACATCAACCGGCTGTCCGACTACGATGTGGACCATATCGTGCCTCAGAGCTTTCTGAAGGACGAC TCCATCGACAACAAGGTGCTGACCAGAAGCGACAAGAACCGGGGCAAGAGCGACAACGTGCCCTCCGAAG AGGTCGTGAAGAAGATGAAGAACTACTGGCGGCAGCTGCTGAACGCCAAGCTGATTACCCAGAGAAAGTT CGACAATCTGACCAAGGCCGAGAGAGGCGGCCTGAGCGAACTGGATAAGGCCGGCTTCATCAAGAGACAG CTGGTGGAAACCCGGCAGATCACAAAGCACGTGGCACAGATCCTGGACTCCCGGATGAACACTAAGTACG ACGAGAATGACAAGCTGATCCGGGAAGTGAAAGTGATCACCCTGAAGTCCAAGCTGGTGTCCGATTTCCG GAAGGATTTCCAGTTTTACAAAGTGCGCGAGATCAACAACTACCACCACGCCCACGACGCCTACCTGAAC GCCGTCGTGGGAACCGCCCTGATCAAAAAGTACCCTAAGCTGGAAAGCGAGTTCGTGTACGGCGACTACA AGGTGTACGACGTGCGGAAGATGATCGCCAAGAGCGAGCAGGAAATCGGCAAGGCTACCGCCAAGTACTT CTTCTACAGCAACATCATGAACTTTTTCAAGACCGAGATTACCCTGGCCAACGGCGAGATCCGGAAGCGG CCTCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGTGGGATAAGGGCCGGGATTTTGCCACCGTGC GGAAAGTGCTGAGCATGCCCCAAGTGAATATCGTGAAAAAGACCGAGGTGCAGACAGGCGGCTTCAGCAA AGAGTCTATCCTGCCCAAGAGGAACAGCGATAAGCTGATCGCCAGAAAGAAGGACTGGGACCCTAAGAAG TACGGCGGCTTCGACAGCCCCACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAAGGGCAAGT CCAAGAAACTGAAGAGTGTGAAAGAGCTGCTGGGGATCACCATCATGGAAAGAAGCAGCTTCGAGAAGAA TCCCATCGACTTTCTGGAAGCCAAGGGCTACAAAGAAGTGAAAAAGGACCTGATCATCAAGCTGCCTAAG TACTCCCTGTTCGAGCTGGAAAACGGCCGGAAGAGAATGCTGGCCTCTGCCGGCGAACTGCAGAAGGGAA ACGAACTGGCCCTGCCCTCCAAATATGTGAACTTCCTGTACCTGGCCAGCCACTATGAGAAGCTGAAGGG CTCCCCCGAGGATAATGAGCAGAAACAGCTGTTTGTGGAACAGCACAAGCACTACCTGGACGAGATCATC GAGCAGATCAGCGAGTTCTCCAAGAGAGTGATCCTGGCCGACGCTAATCTGGACAAAGTGCTGTCCGCCT ACAACAAGCACCGGGATAAGCCCATCAGAGAGCAGGCCGAGAATATCATCCACCTGTTTACCCTGACCAA TCTGGGAGCCCCTGCCGCCTTCAAGTACTTTGACACCACCATCGACCGGAAGAGGTACACCAGCACCAAA GAGGTGCTGGACGCCACCCTGATCCACCAGAGCATCACCGGCCTGTACGAGACACGGATCGACCTGTCTC AGCTGGGAGGTGACTCTGGCGGCTCAAAAAGAACCGCCGACGGCAGCGAATTCGAGCCCAAGAAGAAGAG GAAAGTCTAACCGGTCATCATCACCATCACCATTGAGTTTAAACCCGCTGATCAGCCTCGACTGTGCCTT CTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCAC TGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGT GGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCT CTATGGCTTCTGAGGCGGAAAGAACCAGCTGGGGCTCGATACCGTCGACCTCTAGCTAGAGCTTGGCGTA ATCATGGTCATAGCTGTTTCCTGTGTGAAATTGTTATCCGCTCACAATTCCACACAACATACGAGCCGGA AGCATAAAGTGTAAAGCCTAGGGTGCCTAATGAGTGAGCTAACTCACATTAATTGCGTTGCGCTCACTGC CCGCTTTCCAGTCGGGAAACCTGTCGTGCCAGCTGCATTAATGAATCGGCCAACGCGCGGGGAGAGGCGG TTTGCGTATTGGGCGCTCTTCCGCTTCCTCGCTCACTGACTCGCTGCGCTCGGTCGTTCGGCTGCGGCGA GCGGTATCAGCTCACTCAAAGGCGGTAATACGGTTATCCACAGAATCAGGGGATAACGCAGGAAAGAACA TGTGAGCAAAAGGCCAGCAAAAGGCCAGGAACCGTAAAAAGGCCGCGTTGCTGGCGTTTTTCCATAGGCT CCGCCCCCCTGACGAGCATCACAAAAATCGACGCTCAAGTCAGAGGTGGCGAAACCCGACAGGACTATAA AGATACCAGGCGTTTCCCCCTGGAAGCTCCCTCGTGCGCTCTCCTGTTCCGACCCTGCCGCTTACCGGAT ACCTGTCCGCCTTTCTCCCTTCGGGAAGCGTGGCGCTTTCTCATAGCTCACGCTGTAGGTATCTCAGTTC GGTGTAGGTCGTTCGCTCCAAGCTGGGCTGTGTGCACGAACCCCCCGTTCAGCCCGACCGCTGCGCCTTA TCCGGTAACTATCGTCTTGAGTCCAACCCGGTAAGACACGACTTATCGCCACTGGCAGCAGCCACTGGTA ACAGGATTAGCAGAGCGAGGTATGTAGGCGGTGCTACAGAGTTCTTGAAGTGGTGGCCTAACTACGGCTA CACTAGAAGAACAGTATTTGGTATCTGCGCTCTGCTGAAGCCAGTTACCTTCGGAAAAAGAGTTGGTAGC TCTTGATCCGGCAAACAAACCACCGCTGGTAGCGGTGGTTTTTTTGTTTGCAAGCAGCAGATTACGCGCA GAAAAAAAGGATCTCAAGAAGATCCTTTGATCTTTTCTACGGGGTCTGACACTCAGTGGAACGAAAACTC ACGTTAAGGGATTTTGGTCATGAGATTATCAAAAAGGATCTTCACCTAGATCCTTTTAAATTAAAAATGA AGTTTTAAATCAATCTAAAGTATATATGAGTAAACTTGGTCTGACAGTTACCAATGCTTAATCAGTGAGG CACCTATCTCAGCGATCTGTCTATTTCGTTCATCCATAGTTGCCTGACTCCCCGTCGTGTAGATAACTAC GATACGGGAGGGCTTACCATCTGGCCCCAGTGCTGCAATGATACCGCGAGACCCACGCTCACCGGCTCCA GATTTATCAGCAATAAACCAGCCAGCCGGAAGGGCCGAGCGCAGAAGTGGTCCTGCAACTTTATCCGCCT CCATCCAGTCTATTAATTGTTGCCGGGAAGCTAGAGTAAGTAGTTCGCCAGTTAATAGTTTGCGCAACGT TGTTGCCATTGCTACAGGCATCGTGGTGTCACGCTCGTCGTTTGGTATGGCTTCATTCAGCTCCGGTTCC CAACGATCAAGGCGAGTTACATGATCCCCCATGTTGTGCAAAAAAGCGGTTAGCTCCTTCGGTCCTCCGA TCGTTGTCAGAAGTAAGTTGGCCGCAGTGTTATCACTCATGGTTATGGCAGCACTGCATAATTCTCTTAC TGTCATGCCATCCGTAAGATGCTTTTCTGTGACTGGTGAGTACTCAACCAAGTCATTCTGAGAATAGTGT ATGCGGCGACCGAGTTGCTCTTGCCCGGCGTCAATACGGGATAATACCGCGCCACATAGCAGAACTTTAA AAGTGCTCATCATTGGAAAACGTTCTTCGGGGCGAAAACTCTCAAGGATCTTACCGCTGTTGAGATCCAG TTCGATGTAACCCACTCGTGCACCCAACTGATCTTCAGCATCTTTTACTTTCACCAGCGTTTCTGGGTGA GCAAAAACAGGAAGGCAAAATGCCGCAAAAAAGGGAATAAGGGCGACACGGAAATGTTGAATACTCATAC TCTTCCTTTTTCAATATTATTGAAGCATTTATCAGGGTTATTGTCTCATGAGCGGATACATATTTGAATG TATTTAGAAAAATAAACAAATAGGGGTTCCGCGCACATTTCCCCGAAAAGTGCCACCTGACGTCGACGGA TCGGGAGATCGATCTCCCGATCCCCTAGGGTCGACTCTCAGTACAATCTGCTCTGATGCCGCATAGTTAA GCCAGTATCTGCTCCCTGCTTGTGTGTTGGAGGTCGCTGAGTAGTGCGCGAGCAAAATTTAAGCTACAAC AAGGCAAGGCTTGACCGACAATTGCATGAAGAATCTGCTTAGGGTTAGGCGTTTTGCGCTGCTTCGCGAT GTACGGGCCAGATATACGCGTTGACATTGATTATTGACTAGTTATTAATAGTAATCAATTACGGGGTCAT TAGTTCATAGCCCATATATGGAGTTCCGCGTTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCC CAACGACCCCCGCCCATTGACGTCAATAATGACGTATGTTCCCATAGTAACGCCAATAGGGACTTTCCAT TGACGTCAATGGGTGGAGTATTTACGGTAAACTGCCCACTTGGCAGTACATCAAGTGTATC

[0068] "Base editing activity" means acting to chemically modify a base within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, base editing activity is cytidine deaminase activity, such as the activity of converting a target C·G to T·A. In another embodiment, base editing activity is adenosine or adenine deaminase activity, such as the activity of converting A·T to G·C. In another embodiment, base editing activity is cytidine deaminase activity, such as the activity of converting a target C·G to T·A, and adenosine or adenine deaminase activity, such as the activity of converting A·T to G·C. In some embodiments, base editing activity is evaluated by the efficiency of editing. Base editing efficiency can be measured by any suitable means, for example, by Sanger sequencing or next-generation sequencing. In some embodiments, base editing efficiency is measured by the proportion of total sequencing reads in which the conversion of nucleic acid bases is effected by a base editor, for example, the proportion of total sequencing reads in which the target A.T base pair is converted to a G.C base pair. In some embodiments, base editing efficiency is measured by the proportion of total cells in a population of cells in which the conversion of nucleic acid bases is effected by a base editor when base editing is performed. In some embodiments, base editing efficiency is measured by the proportion of total cells in which the conversion of nucleic acid bases is effected by a base editor when base editing is performed in a population of cells. In some embodiments, base editing efficiency is measured by the proportion of total cells in which the conversion of nucleic acid bases is effected by a base editor when base editing is performed in a population of cells. Base editing efficiency can be measured by any suitable means, for example, by Sanger sequencing or next-generation sequencing. In some embodiments, base editing efficiency is measured by the proportion of total sequencing reads in which the conversion of nucleic acid bases is effected by a base editor, for example, the proportion of total sequencing reads in which the target A.T base pair is converted to a G.C base pair. In some embodiments, base editing efficiency is measured by the proportion of total cells in a population of cells in which the conversion of nucleic acid bases is effected by a base editor when base editing is performed. In some embodiments, base editing efficiency is measured by the proportion of total cells in which the conversion of nucleic acid bases is effected by a base editor when base editing is performed in a population of cells. In some embodiments, base editing efficiency is measured by the proportion of total cells in which the conversion of nucleic acid bases is effected by a base editor when base editing is performed in a population of cells. In some embodiments, base editing efficiency is measured by the proportion of total cells in which the conversion of nucleic acid bases is effected by a base editor when base editing is performed in a population of cells.

[0069] The term "base editor system" is for editing the nucleic acid bases of a target nucleotide sequence refers to the system. In various embodiments, the base editor system comprises (1) a polynu cleotide-programmable nucleotide binding domain (e.g., Cas9); (2) a deaminase domain for deaminating the nucleic acid base (e.g., adenosine deaminase); and ( 3) one or more guide polynucleotides (e.g., guide RNA). In some embodiments, the polynucleotide-programmable nucleotide binding domain is a polynu cleotide-programmable DNA binding domain. In some embodiments, the base editor is an adenine or adenosine base editor (ABE). In some embodiments, the base editor system is ABE8.

[0070] In some embodiments, the base editor system comprises a plurality of base editing components. For example, the base editor system may comprise a plurality of deaminases. In some embodiments, the base editor system may comprise one or more adenosine deaminases. In some embodiments, a single guide polynucleotide can be used to target different deaminases to a target nucleic acid sequence. In some embodiments, a pair of guide polynucleotides can be used to target different deaminases to a target nucleic acid sequence.

[0071] The deaminase domain and the polynucleotide-programmable nucleotide binding components of the base editor system can be associated with each other covalently or non-covalently, or can be associated by any combination of these associations In some embodiments, the deaminase domain can be targeted to a target nucleotide sequence by a polynucleotide-programmable nucleotide-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain can be fused or linked to the deaminase domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain can target the deaminase domain to the target nucleotide sequence by non-covalently interacting or associating with the deaminase domain. For example, in some embodiments, the deaminase domain can interact, associate, or form a complex with a further heterologous moiety or domain that is part of the polynucleotide-programmable nucleotide-binding domain. In some embodiments, this further heterologous moiety can bind, interact, associate, or form a complex with a polypeptide. In some embodiments, this further heterologous moiety can bind, interact, associate, or form a complex with a polynucleotide. In some embodiments, this further heterologous moiety can bind to a guide polynucleotide. In some embodiments, this further heterologous moiety can bind to a polypeptide linker. In some embodiments, this further heterologous moiety can bind to a polynucleotide linker. This further heterologous moiety can be a protein domain. In some embodiments, this further heterologous moiety is a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain. In some embodiments, the deaminase domain can interact, associate, or form a complex with a further heterologous moiety or domain that is part of the polynucleotide-programmable nucleotide-binding domain. In some embodiments, this further heterologous moiety can bind, interact, associate, or form a complex with a polypeptide. In some embodiments, this further heterologous moiety can bind, interact, associate, or form a complex with a polynucleotide. In some embodiments, this further heterologous moiety can bind to a guide polynucleotide. In some embodiments, this further heterologous moiety can bind to a polypeptide linker. In some embodiments, this further heterologous moiety can bind to a polynucleotide linker. This further heterologous moiety can be a protein domain. In some embodiments, this further heterologous moiety is a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain. In some embodiments, the deaminase domain can interact, associate, or form a complex with a further heterologous moiety or domain that is part of the polynucleotide-programmable nucleotide-binding domain. N, SfMu Com coat protein domain, sterile alpha motif, telomerase Ku binding motif and Ku protein, telomerase Sm7 binding motif and Sm7 protein, or may be an RNA recognition motif.

[0072] The base editor system may further include a guide polynucleotide component. The components of the base editor system should be understood to be associated with each other by covalent bonds, non-covalent interactions, or any combination of these associations and interactions. In some embodiments, the deaminase domain may be targeted to a target nucleotide sequence by a guide polynucleotide. For example, in some embodiments, the deaminase domain may interact with, associate with, or form a complex with a part or segment of the guide polynucleotide (e.g., a polynucleotide motif), or may include a further heterologous moiety or domain (e.g., a polynucleotide binding domain such as an RNA-binding protein or a DNA-binding protein). In some embodiments, this further heterologous moiety or domain (e.g., a polynucleotide binding domain such as an RNA-binding protein or a DNA-binding protein) may be fused to or linked to the deaminase domain. In some embodiments, this further heterologous moiety may bind to, interact with, associate with, or form a complex with a polypeptide. In some embodiments, this further heterologous moiety may bind to, interact with, associate with, or form a complex with a polynucleotide. In some embodiments, this heterologous moiety may bind to, interact with, associate with, or form a complex with a polynucleotide. In some embodiments, this heterologous moiety may bind to, interact with, associate with, or form a complex with a polynucleotide. In some embodiments, this heterologous moiety may bind to, interact with, associate with, or form a complex with a polynucleotide. In some embodiments, this heterologous moiety may bind to, interact with, associate with, or form a complex with a polynucleotide. In some embodiments, this heterologous moiety may bind to, interact with, associate with, or form a complex with a polynucleotide. In some embodiments, this Additional heterologous moieties can bind to the guide polynucleotide. In some embodiments, this additional heterologous moiety can bind to a polypeptide linker. In some embodiments this additional heterologous moiety can bind to a polynucleotide linker. This additional heterologous moiety can be a protein domain. In some embodiments, this additional heterologous moiety can be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif.

[0073] In some embodiments, the base editor system can further comprise an inhibitor of a base excision repair (BER) component. The components of the base editor system can associate with each other by covalent bonds, non-covalent interactions, or any combination of these associations and interactions. It should be understood that the inhibitor of the BER component can include a BER inhibitor. In some embodiments, the inhibitor of BER can be a uracil DNA glycosylase inhibitor (UGI). In some embodiments, the inhibitor of BER can be an inosine BER inhibitor. In some embodiments, the inhibitor of BER can be targeted to a target nucleotide sequence by a polynucleotide programmable nucleotide binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain can be fused or linked to the inhibitor of BER. In some embodiments, and the polynucleotide-programmable nucleotide binding domain can be fused or linked to deaminase domains and inhibitors of BER. In some embodiments the polynucleotide-programmable nucleotide binding domain targets the inhibitor of BER to a target nucleotide sequence by non-covalent interaction or association with the inhibitor of BER . For example, in some embodiments, the inhibitor of a BER component can interact with, associate with, or form a complex with a further heterologous moiety or domain that is part of the polynucleotide-programmable nucleotide binding domain . For example, in some embodiments, the inhibitor of a BER component can interact with, associate with, or form a complex with a further heterologous moiety or domain that is part of the polynucleotide-programmable nucleotide binding domain . For example, in some embodiments, the inhibitor of a BER component can interact with, associate with, or form a complex with a further heterologous moiety or domain that is part of the polynucleotide-programmable nucleotide binding domain and can include a further heterologous moiety or domain that can interact with, associate with, or form a complex with the polynucleotide-programmable nucleotide binding domain .

[0074] In some embodiments, the inhibitor of BER can be targeted to a target nucleotide sequence by a guide polynucleotide. For example, in some embodiments, the inhibitor of BER can interact with, associate with, or form a complex with a further heterologous moiety or domain (e.g., a polynucleotide binding domain such as an RNA binding protein or a DNA binding protein) that is part of or a segment (e.g., a polynucleotide motif) of the guide polynucleotide . For example, in some embodiments, the inhibitor of BER can interact with, associate with, or form a complex with a further heterologous moiety or domain (e.g., a polynucleotide binding domain such as an RNA binding protein or a DNA binding protein) that is part of or a segment (e.g., a polynucleotide motif) of the guide polynucleotide . For example, in some embodiments, the inhibitor of BER can interact with, associate with, or form a complex with a further heterologous moiety or domain (e.g., a polynucleotide binding domain such as an RNA binding protein or a DNA binding protein) that is part of or a segment (e.g., a polynucleotide motif) of the guide polynucleotide and can include a further heterologous moiety or domain (e.g., a polynucleotide binding domain such as an RNA binding protein or a DNA binding protein) that can interact with, associate with, or form a complex with the polynucleotide-programmable nucleotide binding domain . In some embodiments, the further heterologous moiety or domain (e.g., a polynucleotide binding domain such as an RNA binding protein or a DNA binding protein) of the guide polynucleotide can be fused or linked to the inhibitor of BER . In some embodiments, this further heterologous moiety can bind to, interact with, associate with, or form a complex with the polynucleotide . In some embodiments, the further heterologous moiety or domain (e.g., a polynucleotide binding domain such as an RNA binding protein or a DNA binding protein) of the guide polynucleotide can be fused or linked to the inhibitor of BER . In some embodiments, this further heterologous moiety can bind to, interact with, associate with, or form a complex with the polynucleotide and can include a further heterologous moiety or domain (e.g., a polynucleotide binding domain such as an RNA binding protein or a DNA binding protein) that can interact with, associate with, or form a complex with the polynucleotide and can include a further heterologous moiety or domain (e.g., a polynucleotide binding domain such as an RNA binding protein or a DNA binding protein) that can interact with, associate with, or form a complex with the polynucleotide 。In some embodiments, this additional heterologous moiety can bind to the guide polynucleotide. In some embodiments, this additional heterologous moiety can bind to the polypeptide linker. In some embodiments, this additional heterologous moiety can bind to the polynucleotide linker. This additional heterologous moiety can be a protein domain. In some embodiments, this additional heterologous moiety can be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif. The terms "Cas9" or "Cas9 domain" refer to an RNA-guided nuclease comprising a Cas9 protein or a fragment thereof (e.g., an active, inactive, or partially active DNA cleavage domain of Cas9, and / or a protein comprising a gRNA binding domain of Cas9). The Cas9 nuclease may also be referred to as a Cas1 nuclease or a CRISPR (clustered regularly interspaced short palindromic repeat) binding nuclease. CRISPR is an adaptive immune system that provides defense against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). A CRISPR cluster comprises spacers, sequences complementary to a preceding mobile element, and a target invading nucleic acid. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). Type II CRIS The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). Type II CRISPR

[0075] The terms "Cas9" or "Cas9 domain" refer to an RNA-guided nuclease comprising a Cas9 protein or a fragment thereof (e.g., an active, inactive, or partially active DNA cleavage domain of Cas9, and / or a protein comprising a gRNA binding domain of Cas9). The Cas9 nuclease may also be referred to as a Cas1 nuclease or a CRISPR (clustered regularly interspaced short palindromic repeat) binding nuclease. CRISPR is an adaptive immune system that provides defense against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). A CRISPR cluster comprises spacers, sequences complementary to a preceding mobile element, and a target invading nucleic acid. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). Type II CRISPR The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). Type II CRISPR The terms "Cas9" or "Cas9 domain" refer to an RNA-guided nuclease comprising a Cas9 protein or a fragment thereof (e.g., an active, inactive, or partially active DNA cleavage domain of Cas9, and / or a protein comprising a gRNA binding domain of Cas9). The Cas9 nuclease may also be referred to as a Cas1 nuclease or a CRISPR (clustered regularly interspaced short palindromic repeat) binding nuclease. CRISPR is an adaptive immune system that provides defense against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). A CRISPR cluster comprises spacers, sequences complementary to a preceding mobile element, and a target invading nucleic acid. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). Type II CRISPR The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). Type II CRISPR The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). Type II CRISPR The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). Type II CRISPR In the PR system, correct processing of pre-crRNA requires trans-encoded small RNA (tracrRNA ), endogenous ribonuclease 3 (rnc), and Cas9 protein. TracrRNA serves as a guide for the processing of pre-crRNA by ribonuclease 3. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to crRNA is first cleaved by the endonuclease and then trimmed 3'-5' by the exonuclease. In nature, DNA binding and cleavage typically require both proteins and both RNAs. However, a single guide RNA (abbreviated as "sgRNA" or simply "gRNA") can be engineered to incorporate both sides of crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012) (the entire content of which is incorporated herein by reference). Cas9 recognizes short motifs (PAM or protospacer adjacent motif) in the CRISPR repeat sequences to help distinguish self from non-self. The sequence and structure of Cas9 nuclease are known to those of ordinary skill in the art (e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti i et al., J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux See also, for example, Savic D.J., Savic G., Lyon K., Primeaux However, both sides of crRNA and tracrRNA can be incorporated into a single RNA species to form a single guide RNA (abbreviated as "sgRNA" or simply "gRNA"). For example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012) (the entire content of which is incorporated herein by reference). Cas9 recognizes short motifs (PAM or protospacer adjacent motif) in the CRISPR repeat sequences to help distinguish self from non-self. The sequence and structure of Cas9 nuclease are known to those of ordinary skill in the art (e.g., uer M., Doudna J.A., Charpentier E. Science 337:816-821(2012)(the entire content of which is incorporated herein by reference). Cas9 recognizes short motifs (PAM or protospacer adjacent motif) in the CRISPR repeat sequences to help distinguish self from non-self. The sequence and structure of Cas9 nuclease are known to those of ordinary skill in the art (e.g., See also, for example, "Complete genome sequence of an M1 strain of Streptococcus pyogenes." Ferretti i et al., J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux in the CRISPR repeat sequences to help distinguish self from non-self. The sequence and structure of Cas9 nuclease are known to those of ordinary skill in the art (e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti i et al., J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C., Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663(2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma C.M., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Voge l J., Charpentier E., Nature 471:602-607(2011); and “A programmable dual-RNA-g uided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821(2012)(the entire contents of which are hereby incorporated by reference). See also). Cas9 ortho logs have been described in various species including, but not limited to, S. pyogenes and S. thermophilus. Further suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on the present disclosure, and such Cas9 nucleases and sequences include, for example, Chylinski, R as well as those disclosed in the following references: (incorporated herein by reference in their entirety). Further suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on the present disclosure, and such Cas9 nucleases and sequences include, for example, Chylinski, R as well as those disclosed in the following references: (incorporated herein by reference in their entirety). Further suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on the present disclosure, and such Cas9 nucleases and sequences include, for example, Chylinski, R hun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas imm unity systems”(2013) RNA Biology 10:5, 726-737 (the entire content of which is incorporated herein by reference) include Cas9 sequences from organisms and loci disclosed in . .

[0076] An exemplary Cas is Streptococcus pyogenes Cas9 (spCas9), and its amino acid sequence is shown below: JPEG2025087691000003.jpg166164 (single underline: HNH domain; double underline: RuvC domain)

[0077] The nuclease-inactivated Cas9 protein can be interchangeably referred to as the “dCas9” protein (nuclease “inactive” Cas9) or catalytically inactive CAs9. Methods for generating a Cas9 protein (or a fragment thereof) having an inactive DNA cleavage domain are known (see, e.g., Jinek et al., Science. 337:816-821 (2012); Qi et al., “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression” (2013) Cell . 28;152(5):1173-83 (the entire content of each of which is incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to include the following two subdomains: H NH nuclease subdomain and RuvC1 subdomain. The HNH sub domain and RuvC1 subdomain are known to be involved in the DNA cleavage activity of Cas9. Methods for inactivating these subdomains to generate nuclease-inactive Cas9 proteins are also known (see, e.g., Jinek et al., Science. 337:816-821 (2012); Qi et al., “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression” (2013) Cell . 28;152(5):1173-83 (the entire content of each of which is incorporated herein by reference). For example, the DNA cleavage domain of Cas9 can be inactivated by introducing mutations into the amino acid residues involved in DNA binding and cleavage, such as the amino acid residues in the HNH nuclease subdomain and RuvC1 subdomain. In one embodiment, the mutations can be introduced into the amino acid residues corresponding to the following positions in the amino acid sequence of Cas9: . For example, the DNA cleavage domain of Cas9 can be inactivated by introducing mutations into the amino acid residues involved in DNA binding and cleavage, such as the amino acid residues in the HNH nuclease subdomain and RuvC1 subdomain. In one embodiment, the mutations can be introduced into the amino acid residues corresponding to the following positions in the amino acid sequence of Cas9: The HNH nuclease subdomain and RuvC1 subdomain are known to be involved in the DNA cleavage activity of Cas9. Methods for inactivating these subdomains to generate nuclease-inactive Cas9 proteins are also known (see, e.g., Jinek et al., Science. 337:816-821 (2012); Qi et al., “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression” (2013) Cell ​The budomain cleaves the strand complementary to the gRNA, and the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can suppress the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (Jinek et al., Science. 337:816-821(2012); Qi et al., Cell. 28;152(5):1 173-83 (2013)). In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is a nickase referred to as the "nCas9" protein (the "nickase" Cas9). In some embodiments, a protein comprising a fragment of Cas9 is provided. For example, in some embodiments, the protein comprises one of the following two Cas9 domains: (1) the gRNA binding domain of Cas9; (2) the DNA cleavage domain of Cas9. In some embodiments, a protein comprising Cas9 or a fragment thereof is referred to as a "Cas9 variant". The Cas9 variant shares homology with Cas9 or a fragment thereof. For example, the Cas9 variant is at least about 70% identical to wild-type Cas9, or at least about 80% identical, or at least about 90% identical, or at least about 95% identical, or at least about 96% identical, or at least about 97% identical, or at least about 98% identical, or at least about 99% identical, or at least about 99.5% identical, or at least about 99.9% identical. In some embodiments, the Cas9 variant is wild-type C as9. ​​​​compared to as9, having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes. In some embodiments, the Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild - type Cas9, and includes a fragment of Cas9 (e.g., the gRNA - binding domain or the DNA - cleavage domain). In some embodiments, this fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid length of the corresponding wild - type Cas9. In some embodiments, this fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid length of the corresponding wild - type Cas9. In some embodiments, this fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid length of the corresponding wild - type Cas9. In some embodiments, this fragment is at least 100 amino acids in length.

[0078] In some embodiments, this fragment is at least 100, 150, 200, 250, 300, 350 amino acids in length. In some embodiments, this fragment is at least 100, 150, 200, 250, 300, 350 , 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, It is 1150, 1200, 1250, or at least 1300 amino acids in length.

[0079] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_017053.1, the nucleotide and amino acid sequences are as follows ). ATGGATAAGAAATACTCAATAGGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGATTATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG GCAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGCAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAATCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTAGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAGAAATGGCTTGTTTGGGAAT CTCATTGCTTTGTCATTGGGATTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATAGTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAGCGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAGGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGCGCCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGGGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGATATTCAAAAAGCACAGGTGTCTGGACAAGGCCATAGTTTA CATGAACAGATTGCTAACTTAGCTGGCAGTCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAATTGTTGATGAACT GGTCAAAGTAATGGGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCCAGA AAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCTGTT GAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTACAAAATGGAAGAGACATGTATGTGGACCAAGAATT AGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCATTAAAGACGATTCAATAGACAATA AGGTACTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAAAAC TATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGGTTT GAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAATTT TGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCTAAA TTAGTTTCTGACTTCCGAAAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGCGTA TCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATAAAG TTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCTAAT ATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGGGGA AACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATATTG TCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTACCAAAAAGAAATTCGGACAAGCTTATTGCT CGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGCTAA GGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCTTTG AAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAAAGACTTAATCATTAAACTACCTAAATAT AGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAAAAGGAAATGAGCTGGCTCT GCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAACAAA AACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTTATT TTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGAAAA TATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTAAAC GATATACGTCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATTGAT TTGAGTCAGCTAGGAGGTGACTGA JPEG2025087691000004.jpg168164(Underlined once: HNH domain; Underlined twice: RuvC domain)

[0080] In some embodiments, wild-type Cas9 corresponds to or comprises the following nucleotide and / or amino acid sequences: ATGGATAAAAAGTATTCTATTGGTTTAGACATCGGCACTAATTCCGTTGGATGGGCTGTCATAACCGATGAATACAAAGT ACCTTCAAAGAAATTTAAGGTGTTGGGGAACACAGACCGTCATTCGATTAAAAAGAATCTTATCGGTGCCCTCCTATTCG ATAGTGGCGAAACGGCAGAGGCGACTCGCCTGAAACGAACCGCTCGGAGAAGGTATACACGTCGCAAGAACCGAATATGT TACTTACAAGAAATTTTTAGCAATGAGATGGCCAAAGTTGACGATTCTTTCTTTCACCGTTTGGAAGAGTCCTTCCTTGT CGAAGAGGACAAGAAACATGAACGGCACCCCATCTTTGGAAACATAGTAGATGAGGTGGCATATCATGAAAAGTACCCAA CGATTTATCACCTCAGAAAAAAGCTAGTTGACTCAACTGATAAAGCGGACCTGAGGTTAATCTACTTGGCTCTTGCCCAT ATGATAAAGTTCCGTGGGCACTTTCTCATTGAGGGTGATCTAAATCCGGACAACTCGGATGTCGACAAACTGTTCATCCA GTTAGTACAAACCTATAATCAGTTGTTTGAAGAGAACCCTATAAATGCAAGTGGCGTGGATGCGAAGGCTATTCTTAGCG CCCGCCTCTCTAAATCCCGACGGCTAGAAAACCTGATCGCACAATTACCCGGAGAGAAGAAAAATGGGTTGTTCGGTAAC CTTATAGCGCTCTCACTAGGCCTGACACCAAATTTTAAGTCGAACTTCGACTTAGCTGAAGATGCCAAATTGCAGCTTAG TAAGGACACGTACGATGACGATCTCGACAATCTACTGGCACAAATTGGAGATCAGTATGCGGACTTATTTTTGGCTGCCA AAAACCTTAGCGATGCAATCCTCCTATCTGACATACTGAGAGTTAATACTGAGATTACCAAGGCGCCGTTATCCGCTTCA ATGATCAAAAGGTACGATGAACATCACCAAGACTTGACACTTCTCAAGGCCCTAGTCCGTCAGCAACTGCCTGAGAAATA TAAGGAAATATTCTTTGATCAGTCGAAAAACGGGTACGCAGGTTATATTGACGGCGGAGCGAGTCAAGAGGAATTCTACA AGTTTATCAAACCCATATTAGAGAAGATGGATGGGACGGAAGAGTTGCTTGTAAAACTCAATCGCGAAGATCTACTGCGA AAGCAGCGGACTTTCGACAACGGTAGCATTCCACATCAAATCCACTTAGGCGAATTGCATGCTATACTTAGAAGGCAGGA GGATTTTTATCCGTTCCTCAAAGACAATCGTGAAAAGATTGAGAAAATCCTAACCTTTCGCATACCTTACTATGTGGGAC CCCTGGCCCGAGGGAACTCTCGGTTCGCATGGATGACAAGAAAGTCCGAAGAAACGATTACTCCATGGAATTTTGAGGAA GTTGTCGATAAAGGTGCGTCAGCTCAATCGTTCATCGAGAGGATGACCAACTTTGACAAGAATTTACCGAACGAAAAAGT ATTGCCTAAGCACAGTTTACTTTACGAGTATTTCACAGTGTACAATGAACTCACGAAAGTTAAGTATGTCACTGAGGGCA TGCGTAAACCCGCCTTTCTAAGCGGAGAACAGAAGAAAGCAATAGTAGATCTGTTATTCAAGACCAACCGCAAAGTGACA GTTAAGCAATTGAAAGAGGACTACTTTAAGAAAATTGAATGCTTCGATTCTGTCGAGATCTCCGGGGTAGAAGATCGATT TAATGCGTCACTTGGTACGTATCATGACCTCCTAAAGATAATTAAAGATAAGGACTTCCTGGATAACGAAGAGAATGAAG ATATCTTAGAAGATATAGTGTTGACTCTTACCCTCTTTGAAGATCGGGAAATGATTGAGGAAAGACTAAAAACATACGCT CACCTGTTCGACGATAAGGTTATGAAACAGTTAAAGAGGCGTCGCTATACGGGCTGGGGACGATTGTCGCGGAAACTTAT CAACGGGATAAGAGACAAGCAAAGTGGTAAAACTATTCTCGATTTTCTAAAGAGCGACGGCTTCGCCAATAGGAACTTTA TGCAGCTGATCCATGATGACTCTTTAACCTTCAAAGAGGATATACAAAAGGCACAGGTTTCCGGACAAGGGGACTCATTG CACGAACATATTGCGAATCTTGCTGGTTCGCCAGCCATCAAAAAGGGCATACTCCAGACAGTCAAAGTAGTGGATGAGCT AGTTAAGGTCATGGGACGTCACAAACCGGAAAACATTGTAATCGAGATGGCACGCGAAAATCAAACGACTCAGAAGGGGC AAAAAAACAGTCGAGAGCGGATGAAGAGAATAGAAGAGGGTATTAAAGAACTGGGCAGCCAGATCTTAAAGGAGCATCCT GTGGAAAATACCCAATTGCAGAACGAGAAACTTTACCTCTATTACCTACAAAATGGAAGGGACATGTATGTTGATCAGGA ACTGGACATAAACCGTTTATCTGATTACGACGTCGATCACATTGTACCCCAATCCTTTTTGAAGGACGATTCAATCGACA ATAAAGTGCTTACACGCTCGGATAAGAACCGAGGGAAAAGTGACAATGTTCCAAGCGAGGAAGTCGTAAAGAAAATGAAG AACTATTGGCGGCAGCTCCTAAATGCGAAACTGATAACGCAAAGAAAGTTCGATAACTTAACTAAAGCTGAGAGGGGTGG CTTGTCTGAACTTGACAAGGCCGGATTTATTAAACGTCAGCTCGTGGAAACCCGCCAAATCACAAAGCATGTTGCACAGA TACTAGATTCCCGAATGAATACGAAATACGACGAGAACGATAAGCTGATTCGGGAAGTCAAAGTAATCACTTTAAAGTCA AAATTGGTGTCGGACTTCAGAAAGGATTTTCAATTCTATAAAGTTAGGGAGATAAATAACTACCACCATGCGCACGACGC TTATCTTAATGCCGTCGTAGGGACCGCACTCATTAAGAAATACCCGAAGCTAGAAAGTGAGTTTGTGTATGGTGATTACA AAGTTTATGACGTCCGTAAGATGATCGCGAAAAGCGAACAGGAGATAGGCAAGGCTACAGCCAAATACTTCTTTTATTCT AACATTATGAATTTCTTTAAGACGGAAATCACTCTGGCAAACGGAGAGATACGCAAACGACCTTTAATTGAAACCAATGG GGAGACAGGTGAAATCGTATGGGATAAGGGCCGGGACTTCGCGACGGTGAGAAAAGTTTTGTCCATGCCCCAAGTCAACA TAGTAAAGAAAACTGAGGTGCAGACCGGAGGGTTTTCAAAGGAATCGATTCTTCCAAAAAGGAATAGTGATAAGCTCATC GCTCGTAAAAAGGACTGGGACCCGAAAAAGTACGGTGGCTTCGATAGCCCTACAGTTGCCTATTCTGTCCTAGTAGTGGC AAAAGTTGAGAAGGGAAAATCCAAGAAACTGAAGTCAGTCAAAGAATTATTGGGGATAACGATTATGGAGCGCTCGTCTT TTGAAAAGAACCCCATCGACTTCCTTGAGGCGAAAGGTTACAAGGAAGTAAAAAAGGATCTCATAATTAAACTACCAAAG TATAGTCTGTTTGAGTTAGAAAATGGCCGAAAACGGATGTTGGCTAGCGCCGGAGAGCTTCAAAAGGGGAACGAACTCGC ACTACCGTCTAAATACGTGAATTTCCTGTATTTAGCGTCCCATTACGAGAAGTTGAAAGGTTCACCTGAAGATAACGAAC AGAAGCAACTTTTTGTTGAGCAGCACAAACATTATCTCGACGAAATCATAGAGCAAATTTCGGAATTCAGTAAGAGAGTC ATCCTAGCTGATGCCAATCTGGACAAAGTATTAAGCGCATACAACAAGCACAGGGATAAACCCATACGTGAGCAGGCGGA AAATATTATCCATTTGTTTACTCTTACCAACCTCGGCGCTCCAGCCGCATTCAAGTATTTTGACACAACGATAGATCGCA AACGATACACTTCTACCAAGGAGGTGCTAGACGCGACACTGATTCACCAATCCATCACGGGATTATATGAAACTCGGATA GATTTGTCACAGCTTGGGGGTGACGGATCCCCCAAGAAGAAGAGGAAAGTCTCGAGCGACTACAAAGACCATGACGGTGA TTATAAAGATCATGACATCGATTACAAGGATGACGATGACAAGGCTGCAGGA JPEG2025087691000005.jpg167162(Underlined once: HNH domain; Underlined twice: RuvC domain)

[0081] In some embodiments, wild-type Cas9 is Cas9 from Streptococcus pyogenes (NC BI reference sequence: NC_002737.2 (nucleotide sequence is as follows); and Uniprot reference sequence: Q99ZW corresponds to 2 (amino acid sequence is as follows). ATGGATAAGAAATACTCAATAGGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGAATATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG ACAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGTAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAACCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTGGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAAAAATGGCTTATTTGGGAAT CTCATTGCTTTGTCATTGGGTTTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATACTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAACGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAAGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGTACCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGAGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGACATTCAAAAAGCACAAGTGTCTGGACAAGGCGATAGTTTA CATGAACATATTGCAAATTTAGCTGGTAGCCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAGTTGTTGATGAATT GGTCAAAGTAATGGGGCGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCC AGAAAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCT GTTGAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTCCAAAATGGAAGAGACATGTATGTGGACCAAGA ATTAGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCCTTAAAGACGATTCAATAGACA ATAAGGTCTTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAA AACTATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGG TTTGAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAA TTTTGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCT AAATTAGTTTCTGACTTCCGAAAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGC GTATCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATA AAGTTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCT AATATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGG GGAAACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATA TTGTCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTACCAAAAAGAAATTCGGACAAGCTTATT GCTCGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGC TAAGGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCT TTGAAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAAAGACTTAATCATTAAACTACCTAAA TATAGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAAAAGGAAATGAGCTGGC TCTGCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAAC AAAAACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTT ATTTTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGA AAATATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTA AACGATATACGTCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATT GATTTGAGTCAGCTAGGAGGTGACTGA JPEG2025087691000006.jpg168164(SEQ ID NO: 1. Single underline: HNH domain; double underline: RuvC domain)

[0082] In some embodiments, Cas9 refers to Cas9 derived from the following or Cas9 derived from any other organism: Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1 ); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_01786 1.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torq uisI (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref: YP_820832.1) ​, Listeria innocua (NCBI Ref: NP_472073.1), Campylobacter jejuni (NCBI Ref: YP_0 02344900.1), or Neisseria meningitidis (NCBI Ref: YP_002342100.1).

[0083] In some embodiments, Cas9 is derived from Neisseria meningitidis (Nme). In some embodiments, Cas9 is Nme1, Nme2, or Nme3. In some embodiments, the PAM interaction domains for Nme1, Nme2, or Nme3 are N 4 GAT, N 4 CC, and N 4 CAAA, respectively (see, for example, Edraki, A., et al., A Compact, High-Accuracy C as9 with a Dinucleotide PAM for In Vivo Genome Editing, Molecular Cell (2018)). The exemplary Neisseria meningitidis Cas9 protein Nme1Cas9 (NCBI reference: WP_002235162.1; type II CRISPR RNA-guided endonuclease Cas9) has the following amino acid sequence: 1 maafkpnpin yilgldigia svgwamveid edenpiclid lgvrvferae vpktgdslam 61 arrlarsvrr ltrrrahrll rarrllkreg vlqaadfden glikslpntp wqlraaaldr 121 kltplewsav llhlikhrgy lsqrkneget adkelgallk gvadnahalq tgdfrtpael 121 kltplewsav llhlikhrgy lsqrkneget adkelgallk gvadnahalq tgdfrtpael 181 alnkfekesg hirnqrgdys htfsrkdlqa elillfekqk efgnphvsgg lkegietllm 241 tqrpalsgda vqkmlghctf epaepkaakn tytaerfiwl tklnnlrile qgserpltdt 301 eratlmdepy rkskltyaqa rkllgledta ffkglrygkd naeastlmem kayhaisral 361 ekeglkdkks plnlspelqd eigtafslfk tdeditgrlk driqpeilea llkhisfdkf 421 vqislkalrr ivplmeqgkr ydeacaeiyg dhygkkntee kiylppipad eirnpvvlra 481 lsqarkving vvrrygspar ihietarevg ksfkdrkeie krqeenrkdr ekaaakfrey 541 fpnfvgepks kdilklrlye qqhgkclysg keinlgrlne kgyveidhal pfsrtwddsf 601 nnkvlvlgse nqnkgnqtpy eyfngkdnsr ewqefkarve tsrfprskkq rillqkfded 661 gfkernlndt ryvnrflcqf vadrmrltgk gkkrvfasng qitnllrgfw glrkvraend 721 rhhaldavvv acstvamqqk itrfvrykem nafdgktidk etgevlhqkt hfpqpweffa 781 qevmirvfgk pdgkpefeea dtpeklrtll aeklssrpea vheyvtplfv srapnrkmsg 841 qghmetvksa krldegvsvl rvpltqlklk dlekmvnrer epklyealka rleahkddpa 901 kafaepfyky dkagnrtqqv kavrveqvqk tgvwvrnhng iadnatmvrv dvfekgdkyy 961 lvpiyswqva kgilpdravv qgkdeedwql iddsfnfkfs lhpndlvevi tkkarmfgyf 1021 aschrgtgni nirihdldhk igkngilegi gvktalsfqk yqidelgkei rpcrlkkrpp 1081 vr

[0084] Another exemplary Neisseria meningitidis Cas9 protein, Nme2Cas9 (NCBI reference: WP_00223083 5; type II CRISPR RNA-guided endonuclease Cas9) has the following amino acid sequence: 1 maafkpnpin yilgldigia svgwamveid eeenpirlid lgvrvferae vpktgdslam 61 arrlarsvrr ltrrrahrll rarrllkreg vlqaadfden glikslpntp wqlraaaldr 121 kltplewsav llhlikhrgy lsqrkneget adkelgallk gvannahalq tgdfrtpael 181 alnkfekesg hirnqrgdys htfsrkdlqa elillfekqk efgnphvsgg lkegietllm 241 tqrpalsgda vqkmlghctf epaepkaakn tytaerfiwl tklnnlrile qgserpltdt 301 eratlmdepy rkskltyaqa rkllgledta ffkglrygkd naeastlmem kayhaisral 361 ekeglkdkks plnlsselqd eigtafslfk tdeditgrlk drvqpeilea llkhisfdkf 421 vqislkalrr ivplmeqgkr ydeacaeiyg dhygkkntee kiylppipad eirnpvvlra 481 lsqarkving vvrrygspar ihietarevg ksfkdrkeie krqeenrkdr ekaaakfrey 541 fpnfvgepks kdilklrlye qqhgkclysg keinlvrlne kgyveidhal pfsrtwddsf 601 nnkvlvlgse nqnkgnqtpy eyfngkdnsr ewqefkarve tsrfprskkq rillqkfded 661 gfkecnlndt ryvnrflcqf vadhilltgk gkrrvfasng qitnllrgfw glrkvraend 721 rhhaldavvv acstvamqqk itrfvrykem nafdgktidk etgkvlhqkt hfpqpweffa 781 qevmirvfgk pdgkpefeea dtpeklrtll aeklssrpea vheyvtplfv srapnrkmsg 841 ahkdtlrsak rfvkhnekis vkrvwlteik ladlenmvny kngreielye alkarleayg 901 gnakqafdpk dnpfykkggq lvkavrvekt qesgvllnkk naytiadngd mvrvdvfckv 961 dkkgknqyfi vpiyawqvae nilpdidckg yriddsytfc fslhkydlia fqkdekskve 1021 fayyincdss ngrfylawhd kgskeqqfri stqnlvliqk yqvnelgkei rpcrlkkrpp 1081 vr

[0085] In some embodiments, dCas9 comprises one or more nucleotides that inactivate Cas9 nuclease activity. corresponds to a Cas9 amino acid sequence having multiple mutations, or a portion or the entire For example, in some embodiments, the dCas9 domain comprises the D10A and H840A aberrations. In some embodiments, the dCas9 comprises a mutation, or a corresponding mutation in another Cas9. 9 includes the amino acid sequence of dCas9 (D10A and H840A): JPEG2025087691000007.jpg168163 (single underline: HNH domain; double underline: RuvC domain).

[0086] In some embodiments, the Cas9 domain comprises a D10A mutation and is or position 840 of the amino acid sequence provided herein. The residue at the corresponding position in either of the above remains a histidine.

[0087] In other embodiments, D10A and D20A, which result in, for example, nuclease-inactivated Cas9 (dCas9), dCas9 variants having mutations other than H840A and H840A are provided. Such mutations include , e.g., other amino acid substitutions at D10 and H840, or within the nuclease domain of Cas9. Other substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain In some embodiments, a variant or homolog of dCas9 is and wherein the sequence is at least about 70% identical, or at least about 80% identical, or at least At least about 90% identical, at least about 95% identical, at least about 98% identical, or at least is at least about 99% identical, at least about 99.5% identical, or at least about 99.9 % identical variants or homologs are provided. In some embodiments, a variant of dCa s9 having an amino acid sequence that is about 5 amino acids, about 10 amino acids, about 15 amino acids, about 2 0 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids, or more amino acids shorter or longer is provided.

[0088] In some embodiments, the Cas9 fusion protein provided herein comprises the full-length amino acid sequence of the Cas9 protein, e.g., one of the Cas9 sequences provided herein . However, in other embodiments, the fusion protein provided herein does not comprise the full-length Cas9 sequence and comprises only one or more fragments thereof. Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and additional suitable sequences of Cas9 domains and fragments will be apparent to those of skill in the art.

[0089] Additional Cas9 proteins (e.g., nuclease-inactive Cas9 (dCas9), Cas9 nickase (nCas 9), or nuclease-active Cas9) (including this variant and homolog) are within the scope of the present disclosure . Exemplary Cas9 proteins include, but are not limited to, those provided below. In some embodiments, the Cas9 protein is nuclease-inactive Cas9 (dCas9). In some embodiments, the Cas9 protein is nuclease-inactive Cas9 (dCas9). In some embodiments, the Cas9 protein The protein is Cas9 nickase (nCas9). In some embodiments, the Cas9 protein is nuclease-active Cas9.

[0090] Exemplary, catalytically inactive Cas9 (dCas9): DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD

[0091] Exemplary catalytic Cas9 nickase (nCas9): DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD

[0092] Exemplary, catalytically active Cas9: DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD.

[0093] In some embodiments, Cas9 refers to Cas9 from archaea (e.g., Nanoarchaea) that constitute the domain and kingdom of unicellular prokaryotic microorganisms. In some embodiments, Cas9 refers to, for example, CasX or CasY as described in Burstein et al., "New CRISPR-Cas systems from uncultivated microbes. " Cell Res. 2017 Feb 21. doi: 10.1038 / cr.2017.21 (the entire content of which is incorporated herein by reference ). Many CRISPR-Cas systems, including Cas9, first reported in the archaeal domain of life, have been identified using genome-resolved metagenomics . This branched Cas9 protein was found in Nanoarchaea, which are little studied as part of active CRISPR-Cas systems. In bacteria, two previously unknown systems were discovered: CRISPR-CasX and CRISPR-CasY, which belong to the most compact systems discovered so far . In some embodiments, Cas9 refers to CasX or a variant of CasX. In some embodiments, Cas9 refers to CasY or a variant of CasY. Other RNA-guided DNA-binding proteins are nucleic acid-programmable DNA -binding proteins that have been engineered to recognize specific target sequences and cleave DNA at those sites . The following two systems, previously unknown in bacteria: CRISPR-CasX and CRISPR-CasY, were discovered, and these belong to the most compact systems already discovered. In some embodiments, Cas9 refers to CasX or a variant of CasX. In some embodiments, Cas9 refers to CasY or a variant of CasY . In some embodiments, Cas9 refers to CasX or a variant of CasX. In some embodiments, Cas9 refers to CasY or a variant of CasY . Other RNA-guided DNA-binding proteins are nucleic acid-programmable DNA It can be used as a binding protein (napDNAbp) and is to be understood as being within the scope of the present disclosure. It should be.

[0094] In certain embodiments, as a napDNAbp useful in the methods of the invention, those known in the art and, for example, the cyclic substitutions described by Oakes et al., Cell 176, 254-267, 2019. Exemplary cyclic substitutions are shown below, where the bold sequences are those derived from Cas9, the italicized sequences are linker sequences, and the underlined sequences are bipartite nuclear localization sequences. CP5 (and MSP "NGC = Pam Variant with mutations Regular Cas9 likes NGG" PID = protein interaction domain, and "D10A" nickase):

[0095] Non-limiting examples of polynucleotide-programmable nucleotide-binding domains that can be incorporated into base editors include domains derived from CRISPR proteins, restriction nucleases, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases (ZFNs). Examples include.

[0096] In some embodiments, any of the fusion proteins provided herein The nucleic acid-programmable DNA-binding protein (napDNAbp) can be a CasX protein or a CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the napDNAbp is a CasY protein. In some embodiments In one state, napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 9 4%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring CasX protein or CasY protein. In some embodiments, n apDNAbp is a naturally occurring CasX protein or CasY protein. In some embodiments, napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 9 8%, at least 99%, or at least 99.5% identical to any CasX protein or Ca sY protein described herein. It should be understood that Cas12b / C2c1, CasX, and CasY from other bacterial species can also be used in accordance with the present disclosure.

[0097] Cas12b / C2c1 (uniprot.org / uniprot / T0D7A2#2) sp|T0D7A2|C2C1_ALIAG CRISPR-associated endo-nuclease C2c1 OS = Alicyclobacillus acido-terrestris (strain ATCC 49025 / DSM 3922 / CIP 106132 / NCIMB 13137 / GD3B) GN= c2c1 PE=1 SV=1 MAVKSIKVKLRLDDMPEIRAGLWKLHKEVNAGVRYYTEWLSLLRQENLYRRSPNGDGEQECDKTAEECKAELLERLRARQ ​VENGHRGPAGSDDELLQLARQLYELLVPQAIGAKGDAQQIARKFLSPLADKDAVGGLGIAKAGNKPRWVRMREAGEPGWE EEKEKAETRKSADRTADVLRALADFGLKPLMRVYTDSEMSSVEWKPLRKGQAVRTWDRDMFQQAIERMMSWESWNQRVGQ EYAKLVEQKNRFEQKNFVGQEHLVHLVNQLQQDMKEASPGLESKEQTAHYVTGRALRGSDKVFEKWGKLAPDAPFDLYDA EIKNVQRRNTRRFGSHDLFAKLAEPEYQALWREDASFLTRYAVYNSILRKLNHAKMFATFTLPDATAHPIWTRFDKLGGN LHQYTFLFNEFGERRHAIRFHKLLKVENGVAREVDDVTVPISMSEQLDNLLPRDPNEPIALYFRDYGAEQHFTGEFGGAK IQCRRDQLAHMHRRRGARDVYLNVSVRVQSQSEARGERRPPYAAVFRLVGDNHRAFVHFDKLSDYLAEHPDDGKLGSEGL LSGLRVMSVDLGLRTSASISVFRVARKDELKPNSKGRVPFFFPIKGNDNLVAVHERSQLLKLPGETESKDLRAIREERQR TLRQLRTQLAYLRLLVRCGSEDVGRRERSWAKLIEQPVDAANHMTPDWREAFENELQKLKSLHGICSDKEWMDAVYESVR RVWRHMGKQVRDWRKDVRSGERPKIRGYAKDVVGGNSIEQIEYLERQYKFLKSWSFFGKVSGQVIRAEKGSRFAITLREH IDHAKEDRLKKLADRIIMEALGYVYALDERGKGKWVAKYPPCQLILLEELSEYQFNNDRPPSENNQLMQWSHRGVFQELI NQAQVHDLLVGTMYAAFSSRFDARTGAPGIRCRRVPARCTQEHNPEPFPWWLNKFVVEHTLDACPLRADDLIPTGEGEIF VSPFSAEEGDFHQIHADLNAAQNLQQRLWSDFDISQIRLRCDWGEVDGELVLIPRLTGKRTADSYSNKVFYTNTGVTYYE RERGKKRRKVFAQEKLSEEEAELLVEADEAREKSVVLMRDPSGIINRGNWTRQKEFWSMV NQRIEGYLVKQIRSRVPLQ DSACENTGDI

[0098] CasX (uniprot.org / uniprot / F0NN87;uniprot.org / uniprot / F0NH53) >tr|F0NN87|F0NN87_SULIH CRISPR-associated Casx protein OS = Sulfolobus islandicu s(strain HVE10 / 4) GN = SiH_0402 PE=4 SV=1 MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAV GQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNL NSDDGKVRDLKLISAYVNGELIRGEG

[0099] >tr|F0NH53|F0NH53_SULIR CRISPR associated protein, Casx OS = Sulfolobus island icus (strain REY15A) GN=SiRe_0771 PE=4 SV=1 MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAV GQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLN SDDGKVRDLKLISAYVNGELIRGEG

[0100] Deltaproteobacteria CasX MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMTDDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAIL QVYWQEFKDDHVGLMCKFAQPASKKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFVYKLEQVSEKGKAYTNYFGRCNVA EHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESTHPVKPLAQIAGNRYASGPVGKALSDACMGTIASFL SKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDAYNEVIARVRMWVNLNLWQKLKLSRDDAK PLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENPK KPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEMD EKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFTD GTDIKKSGKWQGLLYGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKIG RDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQA AKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAKL AYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKELS AELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSYK SGKQPFVGAWQAFYKRRLKEVWKPNA

[0101] CasY (ncbi.nlm.nih.gov / protein / APG80656.1) >APG80656.1 CRISPR-associated protein CasY [uncultured Parcubacteria group bacte rium] MSKRHPRISGVKGYRLHAQRLEYTGKSGAMRTIKYPLYSSPSGGRTVPREIVSAINDDYVGLYGLSNFDDLYNAEKRNEE KVYSVLDFWYDCVQYGAVFSYTAPGLLKNVAEVRGGSYELTKTLKGSHLYDELQIDKVIKFLNKKEISRANGSLDKLKKD IIDCFKAEYRERHKDQCNKLADDIKNAKKDAGASLGERQKKLFRDFFGISEQSENDKPSFTNPLNLTCCLLPFDTVNNNR NRGEVLFNKLKEYAQKLDKNEGSLEMWEYIGIGNSGTAFSNFLGEGFLGRLRENKITELKKAMMDITDAWRGQEQEEELE KRLRILAALTIKLREPKFDNHWGGYRSDINGKLSSWLQNYINQTVKIKEDLKGHKKDLKKAKEMINRFGESDTKEEAVVS SLLESIEKIVPDDSADDEKPDIPAIAIYRRFLSDGRLTLNRFVQREDVQEALIKERLEAEKKKKPKKRKKKSDAEDEKET IDFKELFPHLAKPLKLVPNFYGDSKRELYKKYKNAAIYTDALWKAVEKIYKSAFSSSLKNSFFDTDFDKDFFIKRLQKIF SVYRRFNTDKWKPIVKNSFAPYCDIVSLAENEVLYKPKQSRSRKSAAIDKNRVRLPSTENIAKAGIALARELSVAGFDWK DLLKKEEHEEYIDLIELHKTALALLLAVTETQLDISALDFVENGTVKDFMKTRDGNLVLEGRFLEMFSQSIVFSELRGLA GLMSRKEFITRSAIQTMNGKQAELLYIPHEFQSAKITTPKEMSRAFLDLAPAEFATSLEPESLSEKSLLKLKQMRYYPHY FGYELTRTGQGIDGGVAENALRLEKSPVKKREIKCKQYKTLGRGQNKIVLYVRSSYYQTQFLEWFLHRPKNVQTDVAVSG SFLIDEKKVKTRWNYDALTVALEPVSGSERVFVSQPFTIFPEKSAEEEGQRYLGIDIGEYGIAYTALEITGDSAKILDQN FISDPQLKTLREEVKGLKLDQRRGTFAMPSTKIARIRESLVHSLRNRIHHLALKHKAKIVYELEVSRFEEGKQKIKKVYA TLKKADVYSEIDADKNLQTTVWGKLAVASEISASYTSQFCGACKKLWRAEMQVDETITTQELIGTVRVIKGGTLIDAIKD FMRPPIFDENDTPFPKYRDFCDKHHISKKMRGNSCLFICPFCRANADADIQASQTIALLRYVKEEKKVEDYFERFRKLKN IKVLGQMKKI

[0102] The term "Cas12" or "Cas12 domain" refers to an RNA-guided nuclease comprising a Cas12 protein or a fragment thereof (e.g., a DNA cleavage domain of Cas12 that is active, inactive, or partially active, and / or a Cas 12 gRNA binding domain-containing protein). Cas12 belongs to the class 2, type V CRISPR / Cas system. The Cas12 nuclease is also sometimes referred to as a CRISPR (clustere d regularly interspaced short palindromic repeat)-binding nuclease. There are also exemplary ones. The sequence of an exemplary Bacillus hisashii Cas 12b (BhCas12b) Cas12 domain is provided below as follows: as follows: MAPKKKRKVGIHGVPAAATRSFILKIEPNEEVKKGLWKTHEVLNHGIAYYMNILKLIRQEAIYEHHEQDPKNPKKVSKAE IQAELWDFVLKMQKCNSFTHEVDKDEVFNILRELYEELVPSSVEKKGEANQLSNKFLYPLVDPNSQSGKGTASSGRKPRW YNLKIAGDPSWEEEKKKWEEDKKKDPLAKILGKLAEYGLIPLFIPYTDSNEPIVKEIKWMEKSRNQSVRRLDKDMFIQAL ERFLSWESWNLKVKEEYEKVEKEYKTLEERIKEDIQALKALEQYEKERQEQLLRDTLNTNEYRLSKRGLRGWREIIQKWL KMDENEPSEKYLEVFKDYQRKHPREAGDYSVYEFLSKKENHFIWRNHPEYPYLYATFCEIDKKKKDAKQQATFTLADPIN HPLWVRFEERSGSNLNKYRILTEQLHTEKLKKKLTVQLDRLIYPTESGGWEEKGKVDIVLLPSRQFYNQIFLDIEEKGKH AFTYKDESIKFPLKGTLGGARVQFDRDHLRRYPHKVESGNVGRIYFNMTVNIEPTESPVSKSLKIHRDDFPKVVNFKPKE LTEWIKDSKGKKLKSGIESLEIGLRVMSIDLGQRQAAAASIFEVVDQKPDIEGKLFFPIKGTELYAVHRASFNIKLPGET LVKSREVLRKAREDNLKLMNQKLNFLRNVLHFQQFEDITEREKRVTKWISRQENSDVPLVYQDELIQIRELMYKPYKDWV AFLKQLHKRLEVEIGKEVKHWRKSLSDGRKGLYGISLKNIDEIDRTRKFLLRWSLRPTEPGEVRRLEPGQRFAIDQLNHL NALKEDRLKKMANTIIMHALGYCYDVRKKKWQAKNPACQIILFEDLSNYNPYEERSRFENSKLMKWSRREIPRQVALQGE IYGLQVGEVGAQFSSRFHAKTGSPGIRCSVVTKEKLQDNRFFKNLQREGRLTLDKIAVLKEGDLYPDKGGEKFISLSKDR KCVTTHADINAAQNLQKRFWTRTHGFYKVYCKAYQVDGQTVYIPESKDQKQKIIEEFGEGYFILKDGVYEWVNAGKLKIK KGSSKQSSSELVDSDILKDSFDLASELKGEKLMLYRDPSGNVFPSDKWMAAGVFFGKLERILISKLTNQYSISTIEDDSS KQSMKRPAATKKAGQAKKKK.

[0103] Amino acids having at least 85% or higher identity to the BhCas12b amino acid sequence sequences are also useful in the methods of the present invention.

[0104] The term "conservative amino acid substitution" or "conservative mutation" refers to the replacement of one amino acid with another amino acid having common characteristics. Functional methods for defining the common characteristics between individual amino acids are to analyze the normalized frequencies of amino acid changes between corresponding proteins of homologous organisms (Schulz, G. E. and Schirmer, R. H., Principles o f Protein Structure, Springer-Verlag, New York (1979)). According to such an analysis, amino acids within a group are preferentially exchanged with each other, and thus groups of amino acids that are most similar to each other in their effects on the overall protein structure are defined (Schulz, G. E. and Schirmer, R. H., Principles of Protein Structure, Springer-Verlag, New York (1979)). According to such an analysis, amino acids within a group are preferentially exchanged with each other, and thus groups of amino acids that are most similar to each other in their effects on the overall protein structure are defined by analyzing the normalized frequencies of amino acid changes between corresponding proteins of homologous organisms (Schulz, G. E. and Schirmer, R. H., Principles o can be (Schulz, G. E. and Schirmer, R. H. supra). Non-limiting examples of conservative mutations include, for example, from the amino acid arginine, which can maintain a positive charge, to lysine and its reverse; from aspartic acid, which can maintain a negative charge, to glutamic acid and its reverse; from threonine, in which the free -OH is maintained, to serine; and from asparagine, which can maintain a free NH to amino acid substitutions such as glutamine. As used interchangeably herein, the term "coding sequence" or "protein - coding sequence" refers to a segment of a polynucleotide that encodes a protein. This region or sequence is bounded by a start codon nearer the 5' end and a stop codon nearer the 3' end. The coding sequence is also referred to as an open reading frame. 2 As used herein, the term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In certain embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenine to hypoxanthine. In certain embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or adenine (A) to inosine (I). In certain embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In certain embodiments, the adenosine deaminase

[0105] or its reverse; from threonine, in which the free -OH is maintained, to serine; and from asparagine, which can maintain a free NH to glutamine. This region or sequence is bounded by a start codon nearer the 5' end and a stop codon nearer the 3' end. The coding sequence is also referred to as an open reading frame. As used interchangeably herein, the term "coding sequence" or "protein - coding sequence" refers to a segment of a polynucleotide that encodes a protein.

[0106] As used herein, the term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In certain embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenine to hypoxanthine. In certain embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or adenine (A) to inosine (I). In certain embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In certain embodiments, the adenosine deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In certain embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In certain embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In certain embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. Ze catalyzes the hydrolytic deamination of adenosine in deoxyribonucleic acid (DNA). This The adenosine deaminases provided in the specification (e.g., genetically engineered adenosine de aminases, evolved adenosine deaminases) can be derived from any organism such as bacteria and can be obtained. In some embodiments, the adenosine deaminase is derived from bacteria such as Escherichia coli, Staphyloco ccus aureus, Salmonella typhimurium, Shewanella putrefaciens, Haemophilus influe nzae, or Caulobacter crescentus.

[0107] In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments the TadA deaminase is a TadA variant. In some embodiments, the TadA variant is TadA*8. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from organisms such as humans, chimpanzees, gorillas, monkeys, cows, dogs, rats, or mice. In some embodiments, the deamin ase or deaminase domain does not occur naturally. For example, in some embodiments, the de aminase or deaminase domain has at least 50 %, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 9 8%, at least 99% sequence identity with a naturally occurring deaminase. 8%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, and is at least 99.9% identical. For example, the deaminase domain is described in International PCT Application No. PCT / 20 17 / 045381 (WO 2018 / 027078) and PCT / US2016 / 058344 (WO 2017 / 070632), each of which is hereby incorporated by reference in its entirety. Also, the content of which is hereby incorporated by reference in its entirety, Komor, A.C., et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” N ature 533, 420-424 (2016), Gaudelli, N.M., et al., “Programmable base editing o f A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017) , Komor, A.C., et al., “Improved base excision repair inhibition and bacterioph , et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and pro duct purity” Science Advances 3:eaao4774 (2017), and Rees, H.A., et al., “Ba s, H.A., et al., “Base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017), Komor, A.C., et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and pro See editing: precision chemistry on the genome and transcriptome of living cells. ” Nat Rev Genet. 2018 Dec;19(12):770-788. doi: 10.1038 / s41576-018-0059-1 should also be referred to. For reference.

[0108] "Detection" refers to identifying the presence, absence or amount of an analyte to be detected. In one embodiment, sequence changes in a polynucleotide or polypeptide are detected. In another embodiment, the presence of an indel is detected.

[0109] "Detectable label" means a composition that enables the latter to be detected via spectroscopic, photochemical, biochemical chemical, immunochemical, or chemical means when linked to a molecule of interest. For example, useful labels include radioisotopes, magnetic beads, metal beads, colloidal particles, fluorescent dyes, high electron density reagents, enzymes (e.g., those commonly used in ELISA), biotin , digoxigenin, or haptens.

[0110] "Disease" means any condition or disorder that impairs or interferes with the normal function of a cell, tissue or organ. An example of a disease is glycogen storage disease type 1 (also known as GSD1 or Von Gierke disease). In some embodiments, GSD1 is type 1a (GSD1a).

[0111] "Effective amount" means the amount necessary to improve the symptoms of a disease compared to an untreated patient. The active compound(s) used to practice the present invention for the therapeutic treatment of a disease ​The effective amount of [substance] varies depending on the mode of administration, the age, weight, and general health status of the subject. Ultimately, the attending physician or veterinarian will determine the appropriate amount and dosage. Such an amount is referred to as an "effective" amount. In one embodiment, the effective amount is sufficient to introduce a modification into a gene of interest (e.g., G6PC) in a cell (e.g., an in vitro or in vivo cell) of the present invention's base editor. In one embodiment, the effective amount is the amount of the base editor necessary to achieve a therapeutic effect (e.g., reducing or controlling GSD1a or its symptoms or condition). Such a therapeutic effect need not be sufficient to alter G6PC in all cells of the subject, tissue, or organ, and may only alter G6PC in about 1%, 5%, 10%, 25%, 50%, 75%, or more of the cells present in the subject, tissue, or organ. In one embodiment, the effective amount is sufficient to improve one or more symptoms of GSD1a. "Fragment" means a portion of a polypeptide or nucleic acid molecule. This portion comprises at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% of the full length of the reference nucleic acid

[0112] molecule or polypeptide. A fragment may contain 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 3 00, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids. The "glucose-6-phosphatase (G6PC) polypeptide" means a polypeptide having at least about 95% amino acid sequence identity with NCBI accession number AAA16222. 1 or a fragment thereof.

[0113] ​​​Yes. In certain embodiments, the present invention relates to a single nucleotide polymorphism (SNP) associated with glycogen storage disease type 1a (GSD1a). Provided is a method of editing a G6PC polynucleotide comprising the same. In one embodiment, modification of A·T to G·C at the SNP associated with GSD1a results in a change from glutamine (Q) to a non-glutamine (X) amino acid in the G6PC polypeptide. In another embodiment, modification of A·T to G·C at the SNP associated with GSD1a results in a change from arginine (R) to a non-arginine (X) amino acid in the G6PC polypeptide. In one embodiment, the SNP associated with GSD1a results in the expression of a G6PC polypeptide having a non-glutamine (X) amino acid at position 347 or a non-arginine (X) amino acid at position 83. In one embodiment, base editor correction replaces the glutamine at position 347 with a non-glutamine amino acid (X). In another embodiment, base editor correction replaces the arginine at position 83 with a non-arginine amino acid (X). In certain embodiments, G6PC comprises one or more modifications as compared to the following reference sequence. In certain embodiments, G6PC associated with GSD1a comprises one or more mutations selected from Q347X and R83C. An exemplary G6PC amino acid sequence from Homo Sapiens is provided below: 1 MEEGMNVLHD FGIQSTHYLQ VNYQDSQDWF ILVSVIADLR NAFYVLFPIW FHLQEAVGIK 61 LLWVAVIGDW LNLVFKWILF GQRPYWWVLD TDYYSNTSVP LIKQFPVTCE TGPGSPSGHA 121 MGTAGVYYVM VTSTLSIFQG KIKPTYRFRC LNVILWLGFW AVQLNVCLSR IYLAAHFPHQ 181 YFESLTPDIL LAGLGTLGLL VGLVLLLAVG GVLLLTVLLL LLLLLLLLLL LLLLLLLLLL 1 MEEGMNVLHD FGIQSTHYLQ VNYQDSQDWF ILVSVIADLR NAFYVLFPIW FHLQEAVGIK 61 LLWVAVIGDW LNLVFKWILF GQRPYWWVLD TDYYSNTSVP LIKQFPVTCE TGPGSPSGHA 121 MGTAGVYYVM VTSTLSIFQG KIKPTYRFRC LNVILWLGFW AVQLNVCLSR IYLAAHFPHQ 181 VVAGVLSGIA VAETFSHIHS IYNASLKKYF LITFFLFSFA IGFYLLLKGL GVDLLWTLEK 241 AQRWCEQPEW VHIDTTPFAS LLKNLGTLFG LGLALNSSMY RESCKGKLSK WLPFRLSSIV 301 ASLVLLHVFD SLKPPSQVEL VFYVLSFCKS AVVPLASVSV IPYCLAQVLG QPHKKSL

[0114] The "glucose-6-phosphatase polynucleotide" means a polynucleotide encoding the G6PC polypeptide. An exemplary G6PC nucleotide sequence derived from Homo Sapiens is provided below (GenBank: U01120.1): 1 ATAGCAGAGC AATCACCACC AAGCCTGGAA TAACTGCAAG GGCTCTGCTG ACATCTTCCT 61 GAGGTGCCAA GGAAATGAGG ATGGAGGAAG GAATGAATGT TCTCCATGAC TTTGGGATCC 121 AGTCAACACA TTACCTCCAG GTGAATTACC AAGACTCCCA GGACTGGTTC ATCTTGGTGT 181 CCGTGATCGC AGACCTCAGG AATGCCTTCT ACGTCCTCTT CCCCATCTGG TTCCATCTTC 241 AGGAAGCTGT GGGCATTAAA CTCCTTTGGG TAGCTGTGAT TGGAGACTGG CTCAACCTCG 301 TCTTTAAGTG GTAAGAACCA TATAGAGAGG AGATCAGCAA GAAAAGAGGC TGGCATTCGC 361 TCTCGCAATG TCTGTCCATC AGAAGTTGCT TTCCCCAGGC TATTCAGGAA GCCACGGGCT The "glucose-6-phosphatase polynucleotide" means a polynucleotide encoding the G6PC polypeptide. An exemplary G6PC nucleotide sequence derived from Homo Sapiens is provided below (GenBank: U01120.1): 361 TCTCGCAATG TCTGTCCATC AGAAGTTGCT TTCCCCAGGC TATTCAGGAA GCCACGGGCT 421 ACTCATGCTT CCAACCCCTC TCTCTGACTT TGGATCATCT ACATAAAGGG GGAAGACAGA 481 AAAAATCCTA CCAGTGAGTT GAAAATACAG GAAAGCCTAT TTCATATGGG TTAAAGGGTA 541 GGACAGTTGA ATTTCGTGAA AAGTCTGAGT TATATAGGCT TTGAGCAAAG AGTTTTATTA 601 GTATGAAGCA GAAGAGGTAA CATAAAGAAA GATGTATGGG GCCAGGCATG GTGGCTCACA 661 CCTGTAATCC CAGCACTTTG GGAGGCCGAG GTGGGCGAAT CACTCCTGGG TGAACTCAGG 721 AGTTCAAGAC CAGCCTGGGC AACATGGCGA AACTCCATCT CTACAAAAAC ATTACGAAAA 781 TTAGCTGGGC GTGTTGGTGC TGTAGTCCCA GCTACTCAGG AGGCTGAGGT GAGAGGCGGA 841 GGAGGTTGCA GTGAGTCAAG ATCATGCCAC TGCACTCCAG CCTGGGCAAC AGAGTAAGAC 901 CCTGTCTCAA AAAAAAAAAA AAGATAGATG ATGTATGCTG TATGAAAAAA GGAAACACAC 961 AGATGATTCA ACAGCCTGTT TTGTGGGGTA ATGAAAAGTC ACCCTGGGAA CTGGGCTCCA 1021 GCCCTCGTTC TGCCACCCAC CAACTACATG TCCTTGGCAA GTCATATCAA TTATCTGAGT 1081 TTCTGTTTTA TAATCTACAA ATAGGTTATC TCTGGCAGCT TAATAATAAT CAGGGTTAAC 1141 ATTTATTAAA CAGTGTGTGC CAGTCCATGT GCTATGTGCT TTTCTGTGAG GTAGTTACTG 1201 CTATTTACAG AAACAGTAGA TGCAGAGACC AAGGTGCTGA GTTAAATGAT TAGGCCAACA 1261 AGGTTAGTAC ATGCCGAGCC AGGATGGAAG CCCAGGTAGG CAGGCTGGCT TCCGCGGCAA 1321 TGCTCTTATG AACTATGTTA CGTCCAGTGC TGATAAACTG ACTCTCTGGG GAGCAGGGGA 1381 AAGCCCTGAG TTTAGCATTT GCCAATTTCT ATCACGTAAA CATTCCCATT CTGGCCACTT 1441 TCTTTCTTTC TTTCTTTTGT TTGTTTGTTT GAGATGGAGT CTCGCACTGT TGCCTGGCTG 1501 GAGTGCAATG GTGCAATCTC AGCTCACTGC AACCTCTGCC TCTCCGGTTC AAGTGATTCT 1561 CCTGCCTCAG CCTCCCAAGT AGCTGGGATT ACAGGTGCCC GCCACCATGC CCAGCTAATT 1621 TTTTTTGTAT TTTTAGTAGA GACATGGTTT CACTATGTTG ACTAGGCTGG TCTCGAACTC 1681 CTGACCTCAT GATCTGCCTG CCTTGGCCTC CCTAAGTGCT AGGATTACAG GCGTGAGCCA 1741 CTACACCCAG CCGCATGATT CTAAAAAATA AAAAGATGAA GTGTTATTCC AAACATCTGA 1801 TCTCCATTGA AGAACCATGC AATCTCTCTG GGTTGATAGA GGCCAGAGTT AGTGGCTCTC 1861 CCTGATTTCG GTGAGAAATC ACTATTCCAC CATCACGGGA TAAAAGGCAT CCTGACTGGC 1921 GGTTGACACC TATTTCCACA GTGAAAGATA TATCTAGTAC TTTTAAAGGG GAAGTGGTTT 1981 GTCTGAGATA CTCTGTTTCA AAGTAGAGAG GATACAGAAC AAGCATCTGA AGCTATATAC 2041 ATCCTTACAG AGAGCAATTC TGATGGAAAT GCAGGCCATG TTTCCCTGGG GGGGGCTCGT 2101 CCTAGGGGCT GGAGTGCATT CTCTGATGTC AGAGGAAATG CAAGATTCCC TGAGGCCTGA 2161 GGGAACCCAT GGTATATGCA AGTCCAAGTT TCAAACTGTA GTTCCATATG CATTCTTCCA 2221 GGACAAATAC TTCTTGAGGT TAAAAAAAAA AAGTCACATA GCTGCCATTT TATGGATTTC 2281 AGGATTTTTT TTTTTTTTTT TTTGAGATGG AGTCTTGCTC TGTCACCCAG CCTGTAGTGC 2341 AGTGGCATAA TCTCGGCTCA CGGCAACCTC CGCCTCCCAG GTTCAAGCGA TTCTCTTGCC 2401 TTAGCCTCCC GAGTAGCTGG GATTACAGTC ACGCACCACC ACATCTGGCT AATTCTTTAT 2461 ATTTTTTGGT AGAAACGGTG TTTCACCATG TTGGCCAGGC TGGTCTCAAA CTCCTGACCT 2521 CATGTGATCT GCCTGCCTTG GCCTCCCAAA GTGCTGAGAT TACAGGTGTG AGCCACCGCG 2581 CCTGCCTGGA GTTCAGAATC TTGGGCTTCA TTATTTGTGT TTAAATAGAT CATACAGTCA 2641 GGCACGGTGG CTCATGCCTG TAATCCCAGC ACTTTGGGAG GCTGAGGTGG GAGGATTGCC 2701 TGAGTTCAGG AGATGGAGAC CAGCCTGGGC AACATGGTGA AACCCCGTCT CTACTAAAAA 2761 TACAAAAACT AGCTGGATGT GGTGGCACAC ACCTGTAGTC CCAGCTATTC AGGAGGCTGA 2821 GGTGGGAGGA TCCCAGGAGG TAGAGGTCAC AATGAGCCGA GATTGCGCCA CTGCACTCCA 2881 GGCTGGGTTA CTGAGCCAGA TCCTGTCTCA AAAAAAAAAA AGATAATACA TTCAAACAGT 2941 TCAAAATGCA AAAGTTACAT ACATAAGGAA GTGTCATGAA ATATCTCCCT CTCACACTTC 3001 TCCCCAGCCA CCCAGTTCTC CCTTCTAGAG GCAACATGTG AAATCCTTCT CAGGCTACAC 3061 TCTTCTTGAA GGTGTAGGCT TTGGGCAAAA GCATTCATTC AGTAACCCCA GAAACTTGTT 3121 CTGTTTTTCC ATAGGATTCT CTTTGGACAG CGTCCATACT GGTGGGTTTT GGATACTGAC 3181 TACTACAGCA ACACTTCCGT GCCCCTGATA AAGCAGTTCC CTGTAACCTG TGAGACTGGA 3241 CCAGGTAAGC GTCCCAGCCC CTGCAGACAG AAGCTGAGTG GACCTCGTTT ACCTGTTATG 3301 GATGAAACTG ACCTTGAGGG GACATGAGGA GAGCCATTCC TTTGTACTTT TGTCATGCTC 3361 TTCAATTGGC ACAAATTAAT TCACTTCTGC AATACTTTCC TGAATAGCAC AGTAGTATTG 3421 GAAATCTGCC TATTACAGAA CCTGGATGGA GTCCAGAGAG GCACGGGCAT CCATGGGCAA 3481 AGGGCTCGTG AGAGTCACCG CCCTGCAGCG CTGTGTCCTG AGAAAGGAGG GGGCAGAAGC 3541 CTGAGCTTCT GGGGGTCCTT CCCAATGGCC TGGCCCACTG GATGTGCCCT CCTGAGCTGA 3601 CCGTCCAATC CCTTGCCCTC TCTGTGCCTA CGTTTTATTA GTTACAGCCA GATGGTTACT 3661 GTCAAATCAA ATGATAGATT TCATTTTCAG TATGTAATAG GAAGCCCCTC CCTCACCCTA 3721 AAGTCTCAGC TGCCCTCTAA GACTAGTACT CTCTAAGGTA CTAGTATCCC TTCCTCAGAG 3781 ACCCTTTCCC TGACCCCAAA ACTAGGGAAG GTCCCTTAGT TATTTGCTCT CACAGACCAC 3841 GCATTTACCT CAGAGCATAT TCACTCATTC AGCTGTTACT TACCAAGCAC CTACTGGGAG 3901 CTATACACTG TTCTATGTGC TAGGGATACC TCTGTCAGTG AACAACACAG ACACAAAGAT 3961 CCCTGCCCTT GTGGAGCTGA AATCTGAATA GAGGAGGTGA AATATACAAA AATTATAATA 4021 AATAAGTAAA CTAGGCCAGT TGTGGTTGCT CATGCCTGTA ATCCCAGCAC TTTGGGAAGC 4081 CAAGGTAGGT AGATCACCTG AGGTCAGGAG TTCAAAACCA GCCTGGCCAA CATTGCAAAA 4141 TCCTGTCTTT ACTAAAAATG GAAAAATTGG TCAGGCGTGA TGGCACACGC CTGTAGTCTC 4201 AGCTACCTGG GAGGCTGAGG CAGGAGAATC GCTTGAACCT GGGAGGCAGA GGTTGCAGTG 4261 AACCGAGATC GGACCACTGC ACTCCAGCCT GAATGACAGA ACGAGACTCT GTCTCAAAAA 4321 AAAAGTAAAC TATTAATATG TAGGATAGGC CAGGCACGGT GGCTCACCCT GTAATCCCAG 4381 CACTTTGGGA GGCTGAGGCG GGTGGATCAC CTGAGGTGAG GAGTTCAAGA CCAGCCTGGC 4441 CAACATGGCA AAACCCTGTC TCTACTAAAA ATACAAAAAT TAGCTGGGTG TCCTGGTGCA 4501 TGCCTGTAAT CTGAGCTACT CAGGAGGCTA AGGCAGGAGA ATCGCTTGAA CCTGGGAGGT 4561 GGTGAGCCAA GATTGCGCCA TTGCACTCCA GCCTGGGCGA CAAAATGAGA CACCATCTGA 4621 AAAAAAAAAA AAAATATATA TATATATACA CACACACACA CACACACACA CACACACACA 4681 TATAATACTA GAAAATGATT GTTTATAGGC AAAAAAAAAA AAAAAGAAGA AGAAGAAGAA 4741 AAGGAAAGGA GAAGGAAAGA AGGACCAAAC ATCTTTTGTA GAAATATGTT TGCTTTCATC 4801 ATAACAGCTT GTTATCAAGG ATGAATTTCT CCCTGAAATT AATGGAGGCA CAGACTGGAA 4861 AGTTTAAAGT GGCTTTAAGA GGTTATTTTA TTTAGTCCTC TGTCTTAATA GAAGCAAATT 4921 ATTATCTCTG CTCCTTAGGT AGAGTAGCTA AGGCTCAGAA AGTAGGCCGG GCGCGGTGGC 4981 TCACGCCTGT AATCCTAGCA CTTTGGGAGG CCAACGCAGG TGGATCACCT GAGGTCAGGA 5041 GTTTGAGACC AGCCTGGCCA ACATGGTGAA ACCTCGTCAC TAATAAAAAA ATACAAAAAC 5101 TTAGCCAGGC ATGGTGGCGG GCGCCTGTAA TCCCAGCTAC CCAGGAGGCT GCGGCAGGAG 5161 AATCACTTCA ACCCGGGAGG CAGAGGTTGC AGTGAGCTGA AATCACACCA CTGCACTCCA 5221 GCCTTGGTGA CAGAGAAAGA TTCTGTCAGG AAAAAAAAAA AAAAGTTTAA ATGAATTACC 5281 CAAGGTATAT AATTGTTAGT GTTAGAAGGA AGAAGAAGGG AGGGAGGAAG GAAGGGAGAA 5341 AGAAAGGGAA GGAGGAAGGG AGGGAGGGAA GAAAGCCTTT ATTTATCTAT GGGGTTCCCT 5401 GGAAAGCAGG CTGAAATGGA GATTCACGTG CAGGAGTTTA GATACTCTGG GGAACTATAC 5461 TTGTAGAAGG GAAGGAACAG GAACAGGGCA GAAGGAGAGG TCCGGTTGTG ATTCTGCCTC 5521 ATCCAACCCC ACAGCGAGCT CTGAAGCTGG GGATGGCTCC TCAGAGTTGG TCCAAGTTGG 5581 GACAAGGGAA TCAGACCCTG GGGAGAGCGT AACCTTGATC AAGGCGACTC TCTTTAGCCC 5641 AGGGCAATGC CAGGAGAAGG CTGAGAGCAG AAAGCCATCT ACCATCACAC TCTCAACAGC 5701 TACGAAATAA GTCCTGCAGT TCAGGAGGGA GGTCTGGGCG GCACATCTCA GGACCCTCTA 5761 TCTCTCAGGG TAGAGGAATT AAGAATGGGA TGGGAACCAG ACGGGCCATG GTGGCTCACA 5821 CCTATAATCC CAACACTTTG GGAGGCCAAG GGTAGGAGGA TTGCTTGAGC CCAAGAGTTC 5881 AAAACCAGCC TGGGCAAAAA CAATCAAACA AACAAACAAA ACACATTTAA AAAATTTGCT 5941 GTGTGTGGTG GTGTGCACCT GTGGTCCCAG CTACTCAGGG GGCTGAGGTG GGAGGATTGC 6001 TTGAGTCCAG GAGGTCGAGG CTGCAGTGAG CTATGATCAT GGCACTGCAT TGCAGCCTAG 6061 GAGACAAAGC AAGACACTGT CTCTAAAAAA ACAAAAAACA AACAAATAAA AAAACGGAAC 6121 CGGTTGCAAG CAGGGTTAAA TAGCGTGGTC AGAGTAGGAC TCACTGAGAA TATGAGATCT 6181 GAGTCAAGTC TTCAAGGATG TGAGGAAGTA AGTTTCTGGC AGAAGAGCTG TGAAGGGCTG 6241 TCTGGCCAGA GAAGATTGCA ATGCAAAAGC CCTGAGGTGG GAACGTGTTT GGTGTGTTTA 6301 AAGGAAAGCA ATGAGGCCAG TGTAGCCAGA ACAGAGTGTG CAAGGAGAGA AGGAACAGAA 6361 GATGTGGAGG GCAGATCAGT TTGTAATTGT ACGCCCAGTA TGCTGATTCT TTGTGTAATC 6421 TCCAGACTGT ATTAAACTGC AAGAGCAGGG CCCCTCTCTG GCTTTGCTCA TCATTGTATT 6481 CCCAGAGCCT TGCACAATGC TTGGTGCATA GGAGATGGAA ATTTGTTAAA TAAATGAATT 6541 ATGGATAACG AATGGATGGT AAGATGGGTG GATGGATGGG GGGTGAACGG ATGGATGGGG 6601 GGTGAATGGA TGGATGAATG GGTAGATGGG TGGATAGGGG GATGGCTGGG TGGCTGGGTA 6661 GATGATGCAC TGTCTCCCAG ATGAGGACCT TTTCACCTTT ACTCCATTCT CTTTCCTGCC 6721 CTTTAGGGAG CCCCTCTGGC CATGCCATGG GCACAGCAGG TGTATACTAC GTGATGGTCA 6781 CATCTACTCT TTCCATCTTT CAGGGAAAGA TAAAGCCGAC CTACAGATTT CGGTAAGAAC 6841 TCACCACTGG GGTGTAGGTG GTGGAGGGCA GGAGGCAGCT CTCTCTGTAG CTGACACACC 6901 ACGTATTCTT CCTCACATCC CCCTAGCCCG CTCCCACACC TGGGCAGCCG CTGATTAAGA 6961 GTTGTGGCAC TTTGGATAGG GATAAACCTC AGAGTCAGGG AATGTTTGGG CTGAAAGGGA 7021 TCCAGTAGTG CAATCCGTTG TTTTACAGAT AAGGAAACAA AGCCCAACAC CATGAAGGGA 7081 CTTATAAAAA TAAGGTAGTG AAGTAGCAGC AGGGCTTAAA TAAAAACCCA TGTCTGTACC 7141 AACCACAGAG TCACCCATCC AGGTTAAAAT AACCAGAGAA ACAGAAGATA TTCCTACTAC 7201 AGAGAATTCC GGGTGTGCAG CCACAGTGCA AATCCTTTTT ATTTTTATTT TTGAGATGCA 7261 GTCTCGCTCT GTCATCCAGG CTGAAGTGCA GTGGCACGAT CATGTCTCGC TGCAACCTCT 7321 GCCTCCCAGG CTCAAGCGAT CCTCCCACCT CAGCCATCTG AGTAGCTGGG ACCACAGGCC 7381 ACACACCACA CCCAGCTAAT TTCTCGTATC TTTTTGTAGA GACAGAGTTC TGCTATGTTG 7441 CCCAGGCTCA GGCTGGTCTT GATCTCAAGC AATTGGCTTG CCTCAGCCTC CTAAAATATT 7501 GGGATTACAG GCATGAGCCA CCGCGCCAGC CATGCAAATC CTTAATTATC AAACAGATAA 7561 AATAGGGAAG TTAAAATTCA TATACACAAG GGTTAACCAC TTGCCACAGG CATTTTTTTT 7621 TTTTTTTTGA GACGGAATCT CGCTCTGTTG CCCAGGCTGG AGTGCAGTGG CGCCATCTCG 7681 CCTCACTGCA ACCTCCGCTT CCTGGGTTCA AGCTATTCTT CTGCCTCAGC CTACCGAGTA 7741 GCTGGGACTA CAGGCACGTG CCACCACACC TGGCTAATTT TTTTATTTTT AGTAGAGATG 7801 GGGTTTCACC ATATTGGCCA GGCTGGTCTT GAACTCCTGA CCTAGTGATC CATCCGCCTC 7861 AGCCTCCCAA AGTGCTGGGA TTGCAGGCAT GAGCCACCGC GCCTGGCCTT TTTTTTTTTT 7921 TTTTGAGACG GAGTTTTGCT CTTGTTGCCC AGGCTAGAGT GCAGTGGCGC AGTCTCGGCT 7981 CACTGTAACC TCCACCTCCT GAGTTCAAGC AATTCTCCTG CCTCAGCCTC TCAAATAGCT 8041 GGGATTACAG GCGTGAGCCA CCCCACCTGG CTAATTTTGT AATTTTTTTT TTAGTAGAGA 8101 TGGGGTTTCA CCTGTTGATC AGGCTGGTCT CAAACTCCTG ACCTCAAGTG ATCCACCCAC 8161 CTCGGCCTCC CAAAGTGCTG GGATTACAAG CATAAGCCAC CGTGCCTGGT CAATTTTGAT 8221 CTTTTTTAAA GAGACAGGGG TCTTGCTATG TTGCCCAGAC TAGTCTTGAA CTCCTGGCCT 8281 CAAGTGATCC TCTCACCTCG GCCTCCCAAA GTATTGGGAT TACAGGTCTG AGCCGCTGCA 8341 CCCAGCCCCC AACAGGCATC TTTGGACTTT TGAGTACTGG CTTTAATTTA CAAAAATTCC 8401 ACTGAGAGCA CCTAAGTTTG CCAGGCTCCA ACATTTCTGC AGGGGCTGTT TTCTTTGCTG 8461 AAGGATCTGC ACCTGTGTTC TGTTATGGTT GCCTCTTCTG TTGCAGGTGC TTGAATGTCA 8521 TTTTGTGGTT GGGATTCTGG GCTGTGCAGC TGAATGTCTG TCTGTCACGA ATCTACCTTG 8581 CTGCTCATTT TCCTCATCAA GTTGTTGCTG GAGTCCTGTC AGGTATGGGC TGATCTGACT 8641 CCCTTCCTTC TCCCCCAAAC CCCATTCCGT TTCTCTCCCT AATCAGGACA AAATCCCAGC 8701 ATTCCAGCCA CATCCTGTGT GTAATCAGTA CTGTTAGCAT TTCTGTGGGT TGAAAGTCAA 8761 GAATGAGCAA CTTGAAATGA TTAATTTCTA TAAGAGTGCC CAGATCTATA GAATGAATTG 8821 TGTAGAAGTT ACCATACATC AAATTAACGC ACCAAATTGA ATTAGCTTGA AATCTCAGAG 8881 CTTTTTACAA TCTTTATTTC TTACTGGTCT TCAACAGGCC CTAATTTACT TTTCAGGGAA 8941 TCTGCCAAAT TTAACAAATT AACACGATGT CCTAGGAAAG CTGTTCATTT AAATACATTC 9001 ATTTGCAAAC CTAATAGATA ACTGCAGTTG ATCTCTTTTA TAGGTTCAGA GTTTTGAATA 9061 TGTTTTTTTT TGTTTTTTTT TTTTGAGATG GAGTCTCGCT CTGTGACCCA GGCTAGAGTG 9121 CAGTGGTGCG ATCTCGGCTC ACTGCAAGCT CCACCTCCTG GGTTCACGCC ATTCTCCTGC 9181 CTCAGCCTCT CCGAGTAGCT GGGACTACAG GCGCCCGCCA CCATGCCCGG CTAATTTTTT 9241 GTATTTTTAG CAGAGACGGG GTTTCACCGT GGTCTTGATC TCCTGACCTC GTGATCCGCC 9301 CGCCTCGGCC TCCCAAAGCG CTGGGATTAC AAGGGTGAGC CACCGCACCC TGCCTGAATA 9361 TGTGTTTTCT TAGATCCAAT TAACAAGGGT AAGACAAGAT TTAAGTTAAG CATAAGAAAG 9421 ATTTTGTGGG AGGCACTGGA ATATAAGACC TTAACAAAAC TGTGGAATTT CTCCCCTGGA 9481 GATTTGTAAG AACGGAACAT AGCAGCATTC AAAGAAGAAT GTTGAGAACA AGGGAGATAA 9541 TGGTTTCATG GTAATCACAA AAGTAACACA GCATTTAGTA CTGGGTTCCA TGTTTGAGGA 9601 AGAACCTGGA AGCCATATCA CATGAAAAAC CTGGGAATGT TTAGGTTAGA GAGAATAACT 9661 GTGTTCAAAT GTGTGACAGA GGGACTAGAT TCATCACTTA CTAACTCCTG CAGAAAGAAC 9721 TGAGAAAAAT AGACAGTATT AGAGGGGGAC CAGTTTCACA CAGACAAGGA AGAACTATTC 9781 AGCAATCAAT TCCGTTCAAA GATAAAATGG ACTGTTATAG TGGGGGTGAG CTCCCTACCT 9841 CTGAGGGTAT TTCAAGTAGA GATAGGAGGA CCTCCTGGTA GGAAATTTGC ATACGGTGGG 9901 AGATTGTACG TGATATGGCA CCTCCATCTG AAAGAGTCTA TATTGAGGGC AGGCTGGAGT 9961 CACACATGGG AATAAGCCAG GCGACCCTCC CATCTGCCAT CTGTGATTTA ATTCCACAGT 10021 CGCAGAACGG ATGGCATGTC ACCCACTCCT CCAAACCCAC CTCTAGCAAA GGTCCCAAAT 10081 CCTTCCTATC TCTCACAGTC ATGCTTTCTT CCACTCAGGC ATTGCTGTTA CAGAAACTTT 10141 CAGCCACATC CACAGCATCT ATAATGCCAG CCTCAAGAAA TATTTTCTCA TTACCTTCTT 10201 CCTGTTCAGC TTCGCCATCG GATTTTATCT GCTGCTCAAG GGACTGGGTG TAGACCTCCT 10261 GTGGACTCTG GAGAAAGCCC AGAGGTGGTG CGAGCAGCCA GAATGGGTCC ACATTGACAC 10321 CACACCCTTT GCCAGCCTCC TCAAGAACCT GGGCACGCTC TTTGGCCTGG GGCTGGCTCT 10381 CAACTCCAGC ATGTACAGGG AGAGCTGCAA GGGGAAACTC AGCAAGTGGC TCCCATTCCG 10441 CCTCAGCTCT ATTGTAGCCT CCCTCGTCCT CCTGCACGTC TTTGACTCCT TGAAACCCCC 10501 ATCCCAAGTC GAGCTGGTCT TCTACGTCTT GTCCTTCTGC AAGAGTGCGG TAGTGCCCCT 10561 GGCATCCGTC AGTGTCATCC CCTACTGCCT CGCCCAGGTC CTGGGCCAGC CGCACAAGAA 10621 GTCGTTGTAA GAGATGTGGA GTCTTCGGTG TTTAAAGTCA ACAACCATGC CAGGGATTGA 10681 GGAGGACTAC TATTTGAAGC AATGGGCACT GGTATTTGGA GCAAGTGACA TGCCATCCAT 10741 TCTGCCGTCG TGGAATTAAA TCACGGATGG CAGATTGGAG GGTCGCCTGG CTTATTCCCA 10801 TGTGTGACTC CAGCCTGCCC TCAGCACAGA CTCTTTCAGA TGGAGGTGCC ATATCACGTA 10861 CACCATATGC AAGTTTCCCG CCAGGAGGTC CTCCTCTCTC TACTTGAATA CTCTCACAAG 10921 TAGGGAGCTC ACTCCCACTG GAACAGCCCA TTTTATCTTT GAATGGTCTT CTGCCAGCCC 10981 ATTTTGAGGC CAGAGGTGCT GTCAGCTCAG GTGGTCCTCT TTTACAATCC TAATCATATT 11041 GGGTAATGTT TTTGAAAAGC TAATGAAGCT ATTGAGAAAG ACCTGTTGCT AGAAGTTGGG 11101 TTGTTCTGGA TTTTCCCCTG AAGACTTACT TATTCTTCCG TCACATATAC AAAAGCAAGA 11161 CTTCCAGGTA GGGCCAGCTC ACAAGCCCAG GCTGGAGATC CTAACTGAGA ATTTTCTACC 11221 TGTGTTCATT CTTACCGAGA AAAGGAGAAA GGAGCTCTGA ATCTGATAGG AAAAGAAGGC 11281 TGCCTAAGGA GGAGTTTTTA GTATGTGGCG TATCATGCAA GTGCTATGCC AAGCCATGTC 11341 TAAATGGCTT TAATTATATA GTAATGCACT CTCAGTAATG GGGGACCAGC TTAAGTATAA 11401 TTAATAGATG GTTAGTGGGG TAATTCTGCT TCTAGTATTT TTTTTACTGT GCATACATGT 11461 TCATCGTATT TCCTTGGATT TCTGAATGGC TGCAGTGACC CAGATATTGC ACTAGGTCAA 11521 AACATTCAGG TATAGCTGAC ATCTCCTCTA TCACATTACA TCATCCTCCT TATAAGCCCA 11581 GCTCTGCTTT TTCCAGATTC TTCCACTGGC TCCACATCCA CCCCACTGGA TCTTCAGAAG 11641 GCTAGAGGGC GACTCTGGTG GTGCTTTTGT ATGTTTCAAT TAGGCTCTGA AATCTTGGGC 11701 AAAATGACAA GGGGAGGGCC AGGATTCCTC TCTCAGGTCA CTCCAGTGTT ACTTTTAATT 11761 CCTAGAGGGT AAATATGACT CCTTTCTCTA TCCCAAGCCA ACCAAGAGCA CATTCTTAAA 11821 GGAAAAGTCA ACATCTTCTC TCTTTTTTTT TTTTTTTGAG ACAGGGTCTC ACTATGTTGC 11881 CCAGGCTGCT CTTGAATTCC TGGGCTCAAG CAGTCCTCCC ACCCTACCAC AGCGTCCCGC 11941 GTAGCTGGGA CTACAGGTGC AAGCCACTAT GTCCAGCTAG CCAACTCCTC CTTGCCTGCT 12001 TTTCTTTTTT TTTCTTTTTT TGAGACGGCG CACCTATCAC CCAGGCTGGA GTGGAGTGGC 12061 ACGATCTTGG CTCACTGCAA CCTCTTCCTC CTGGTTCAAG CGATTCTCAT GTCTCAGCCT 12121 CCTCAGTAGC TAGGACTACC GGCGTGCACC ACCATGCCAG GCTAATTTTT ATATTTTTAG 12181 AATTTTAGAA GAGATGGGAT TTCATCATGT TGGCCAGGCT GGTCTCGAAC TCCTGACCTC 12241 AAGTGATCCA CCTGCCTTGG CCTCCCAAGG TGCTAGGATT ACAGGCATGA GCCACCGCAC 12301 CGGGCCCTCC TTGCCTGTTT TTCAATCTCA TCTGATATGC AGAGTATTTC TGCCCCACCC 12361 ACCTACCCCC CAAAAAAAGC TGAAGCCTAT TTATTTGAAA GTCCTTGTTT TTGCTACTAA 12421 TTATATAGTA TACCATACAT TATCATTCAA AACAACCATC CTGCTCATAA CATCTTTGAA 12481 AAGAAAAATA TATATGTGCA GTATTTTATT AAAGCAACAT TTTATTTAAG AATAAAGTCT 12541 TGTTAATTAC TATATTTTAG ATGCAATGTG ATCTGAAGTT TCTAATTCTG GCCCAACTAA 12601 ATTTCTAGCT CTGTTTCCCT AAACAAATAA TTTGGTTTCT CTGTGCCTGC ATTTTCCCTT 12661 TGGAGAAGAA AAGTGCTCTC TCTTGAGTTG ACCGAGAGTC CCATTAGGGA TAGGGAGACT 12721 TAAATGCATC CACAGGGGCA CAGGCAGAGT TGAGCACATA AACGGAGGCC CAAAATCAGC 12781 ATAGAACCAG AAAGATTCAG AGTTGGCCAA GAATGAACAT TGGCTACCAG ACCACAAGTC 12841 AGCATGAGTT GCTCTATGGC ATCAAATTGC AACTTGAGAG TAGATGGGCA GGGTCACTAT 12901 CAAATTAAGC AATCAGGGCA CACAAGTTGC AGTAACACAA CAAGACTAGG CCAGCTCTGG 12961 AATCCAGTAA CTCAGTGTCA GCAAGGTTTT GGGTTATAGT TCAAGAAAGT CTAAACAGAG 13021 CCAGTCACAG CACCAAGGAA TGCTCAAGGG AGCTATTGCA GGTTTCTCTG CTAAGAGATT 13081 TATTTCATCC TGGGTGCAGG GTTCGACCTC CAAAGGCCTC AAATCATCAC CGTATCAATG 13141 GATTTCCTGA GGGTAAGCTC CGCTATTTCA CACCTGAACT CCGGAGTCTG TATATTCAGG 13201 GAAGATTGCA TTCTCCTACT GGATTTGGGC TCTCAGAGGG CGTTGTGGGA ACCAGGCCCC 13261 TCACAGAATC AAATGGTCCC AACCAGGGAG AAAGAAAATA GTCTTTTTTT TTTTTTTAAT 13321 AGAGATGGGG GTCTCACTAT GCTGCCCAGG CTGGTCTTGA ACTCCTGGGT TCAAGTGATC 13381 CTCCTGCCTC AGCCTCCCAA AGTGCTGGGA TTACAGTGTG AGCCACTGCG CTTGGCCAGA 13441 AATGGTTTTG ATCTGTCTGA ACTGAACCCT ACTGCTTAGG CATAGCCCCA TCCTTGATAA 13501 TCTATTTGCT CCCAAGGACC AAGTCCAAGA TCCTTACAAG AAAGGTCTGC CAGAAAGTAA 13561 ATACTGCCCC CACTCCCTGA AGTTTATGAG GTTGATAAGA AAACATAACA GATAAAGTTT 13621 ATTGAGTGCT AACTTTA

[0115] "Guide RNA" or "gRNA" can be specific to a target sequence and means a polynucleotide that can form a complex with a programmable nucleotide binding domain protein (such as Cas9 or Cpf1). In one embodiment, the guide polynucleotide is guide RNA (gRNA). The gRNA may exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule may be referred to as single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to guide RNA that exists as a single molecule or as a complex of two or more molecules. ​​​​​​is used. Typically, a gRNA that exists as a single RNA species has (1) a domain that is homologous to the target nucleic acid (e.g., that directs binding of the Cas9 complex to the target); and (2) two domains, a domain that binds to the Cas9 protein. In certain embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to a tracrRNA as provided in Jinek et al., Science 337:816-821 (2012) (the entire contents of which are incorporated herein by reference). Other examples of gRNAs (e.g., those that include domain 2) can be found in U.S. Provisional Patent Application No. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application No. 61 / 874,746, filed September 6, 2013, entitled "Delivery System For Functional Nucleases", the entire contents of each of which are incorporated herein by reference in their entirety. In some embodiments, a gRNA includes two or more of domains (1) and (2) and can be referred to as an "extended gRNA". An extended gRNA binds to two or more Cas9 proteins and binds to the target nucleic acid at two or more distinct regions, as described herein. The gRNA includes a nucleotide sequence complementary to the target site, which mediates binding of the nuclease / RNA complex to the target site and provides sequence specificity for the nuclease:RNA complex. As will be understood by those of skill in the art, RN contains a domain that is common to the target nucleic acid (e.g., that directs the binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In certain embodiments, domain (2 contains a domain that is common to the target nucleic acid (e.g., that directs the binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In certain embodiments, domain (2 ) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to a tracrRNA as provided in Jinek et al., Science 337:816-821 (2012) (the entire contents of which are incorporated herein by reference). Other examples of gRNAs (e.g., those that include domain 2) can be found in U.S. Provisional Patent Application No. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application No. 61 / 874,746, filed September 6, 2013, entitled "Delivery System For Functional Nucleases", the entire contents of each of which are incorporated herein by reference in their entirety. In some embodiments, a gRNA includes two or more of domains (1) and (2) and can be referred to as an "extended gRNA". An extended gRNA binds to two or more Cas9 proteins and binds to the target nucleic acid at two or more distinct regions, as described herein. The gRNA includes a nucleotide sequence complementary to the target site, which mediates binding of the nuclease / RNA complex to the target site and provides sequence specificity for the nuclease:RNA complex. As will be understood by those of skill in the art, RN contains a domain that is common to the target nucleic acid (e.g., that directs the binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In certain embodiments, domain (2 ) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to a tracrRNA as provided in Jinek et al., Science 337:816-821 (2012) (the entire contents of which are incorporated herein by reference). Other examples of gRNAs (e.g., those that include domain 2) can be found in U.S. Provisional Patent Application No. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application No. 61 / 874,746, filed September 6, 2013, entitled "Delivery System For Functional Nucleases", the entire contents of each of which are incorporated herein by reference in their entirety. In some embodiments, a gRNA includes two or more of domains (1) and (2) and can be referred to as an "extended gRNA". An extended gRNA binds to two or more Cas9 proteins and binds to the target nucleic acid at two or more distinct regions, as described herein. The gRNA includes a nucleotide sequence complementary to the target site, which mediates binding of the nuclease / RNA complex to the target site and provides sequence specificity for the nuclease:RNA complex. As will be understood by those of skill in the art, RN contains a domain that is common to the target nucleic acid (e.g., that directs the binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In certain embodiments, domain (2 contains a domain that is common to the target nucleic acid (e.g., that directs the binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In certain embodiments, domain (2 ) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to a tracrRNA as provided in Jinek et al., Science 337:816-821 (2012) (the entire contents of which are incorporated herein by reference). Other examples of gRNAs (e.g., those that include domain 2) can be found in U.S. Provisional Patent Application No. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application No. 61 / 874,746, filed September 6, 2013, entitled "Delivery System For Functional Nucleases", the entire contents of each of which are incorporated herein by reference in their entirety. In some embodiments, a gRNA includes two or more of domains (1) and (2) and can be referred to as an "extended gRNA". An extended gRNA binds to two or more Cas9 proteins and binds to the target nucleic acid at two or more distinct regions, as described herein. The gRNA includes a nucleotide sequence complementary to the target site, which mediates binding of the nuclease / RNA complex to the target site and provides sequence specificity for the nuclease:RNA complex. As will be understood by those of skill in the art, RN contains a domain that is common to the target nucleic acid (e.g., that directs the binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In certain embodiments, domain (2 contains a domain that is common to the target nucleic acid (e.g., that directs the binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In certain embodiments, domain (2 contains a domain that is common to the target nucleic acid (e.g., that directs the binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In certain embodiments, domain (2 contains a domain that is common to the target nucleic acid (e.g., that directs the binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In certain embodiments, domain (2 contains a domain that is common to the target nucleic acid (e.g., that directs the binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In certain embodiments, domain (2 contains a domain that is common to the target nucleic acid (e.g., that directs the binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In certain embodiments, domain (2 contains a domain that is common to the target nucleic acid (e.g., that directs the binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In certain embodiments, domain (2 A polynucleotide sequence, such as a gRNA sequence, contains the nucleobase uracil (U), a pyrimidine derivative, instead of the nucleobase thymine (T) contained in a DNA polynucleotide sequence. In RNA, uracil base pairs with adenine and replaces thymine during DNA transcription.

[0116] The term "heterodimer" refers to a fusion protein containing two domains, such as a wild-type TadA domain and a variant of the TadA domain (e.g., TadA *8) or two variant TadA domains (e.g., TadA*7.10 and TadA*8 or two TadA*8 domains).

[0117] "Hybridization" means hydrogen bonding between complementary nucleobases and can be Watson-Crick, Hoogsteen or reverse Hoogsteen hydrogen bonds. For example, adenine and thymine are complementary nucleobases that form hydrogen bonds to form a pair.

[0118] The term "inhibitor of base repair" or "IBR" refers to a protein that can inhibit the activity of a nucleic acid repair enzyme, such as a base excision repair (BER) enzyme. In some embodiments, the IBR is an inhibitor of inosine base excision repair. Exemplary inhibitors of base repair include inhibitors of APE1, Endo III, Endo IV, Endo V, Endo VIII, Fpg, hOGGl, hNEILl, T7 Endol, T4PDG, UDG, hSMUGl, and hAAG. In some embodiments, the IBR is an inhibitor of Endo V or hAAG. In some embodiments, the IBR is catalytically inactive EndoV or catalytically inactive hAAG. In some embodiments, the base repair inhibitor is Endo V or hAAG It is an inhibitory factor. In some embodiments, the base repair inhibitory factor is catalytically inactive EndoV or is catalytically inactive hAAG.

[0119] In certain embodiments, the base repair inhibitory factor is uracil glycosylase inhibitor (UGI). UGI refers to a protein that can inhibit the base excision repair enzyme of uracil-DNA glycosylase. In some embodiments, the UGI domain includes wild-type UGI or a fragment thereof. In some embodiments, the UGI protein provided herein includes a fragment of UGI and a protein homologous to UGI or a UGI fragment. In one aspect, the base repair inhibitory factor is an inhibitor of inosine base excision repair. In some aspects, the base repair inhibitory factor is a "catalytically inactive inosine-specific nuclease" or a "dead inosine-specific nuclease". Without wishing to be bound by any particular theory, catalytically inactive inosine glycosylase (e.g., alkyladenine glycosylase (AAG)) can bind to inosine but cannot create an abasic site or remove inosine, thereby sterically blocking the newly formed inosine moiety from the DNA damage / repair mechanism. In some embodiments, the catalytically inactive inosine-specific nuclease can bind to inosine in nucleic acids but cannot cleave the nucleic acids. Non-limiting examples of representative catalytically inactive inosine-specific nucleases include, (e.g., human-derived) catalytically inactive alkyladenine glycosylase (AAG nuclease), and (e.g., E. coli-derived) catalytically inactive endonuclease V (EndoV nuclease). ​​​​​​​Examples include (clease). In some embodiments, the catalytically inactive AAG nuclease contains the E125Q mutation or a corresponding mutation in another AAG nuclease.

[0120] "Increase" means a positive change of at least 10%, 25%, 50%, 75%, or 100%.

[0121] "Intein" is a protein fragment that can excise itself and link the remaining fragments (exteins) by peptide bonds in a process known as protein splicing. Intein is also called "protein intron". The process by which an intein excises itself and ligates the remaining part of the protein is referred to herein as "protein splicing" or "intein-mediated protein splicing". In one aspect, the intein of a precursor protein (the intein-containing protein before intein-mediated protein splicing) is derived from two genes. Such inteins are referred to herein as split inteins (e.g., split intein-N and split intein-C). For example, in Synechocystis, DnaE, which is the catalytic subunit a of DNA polymerase III, is encoded by two separate genes, dnaE-n and dnaE-c. The intein encoded by the dnaE-n gene can be referred to herein as "intein N". The intein encoded by the dnaE-c gene can be referred to herein as "intein C".

[0122] Other intein systems can also be used. As an example, the dnaE intein, i.e., Cfa-N (e.g.,​​​ based on intein pairs of split intein-N) and Cfa-C (e.g., split intein-C) Synthetic inteins are described (e.g., Steven et al., J Am Chem Soc. 2016 Feb. 24; 138(7):2162-5, which is incorporated herein by reference). Non-limiting examples of intein pairs that can be used according to the present disclosure include Cfa DnaE intein, Ssp GyrB intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB intein, and Cne Prp8 intein (e.g., as described in U.S. Patent No. 8,394,604, which is incorporated herein by reference). tein, and Cne Prp8 intein (e.g., as described in U.S. Patent No. 8,394,604, which is incorporated herein by reference).

[0123] Exemplary nucleotide and amino acid sequences of the intein are provided.

[0124] DnaE intein-N DNA: TGCCTGTCATACGAAACCGAGATACTGACAGTAGAATATGGCCTTCTGCCAATCGGGAAGATTGTGGAGAAACGGATAGA ATGCACAGTTTACTCTGTCGATAACAATGGTAACATTTATACTCAGCCAGTTGCCCAGTGGCACGACCGGGGAGAGCAGG AAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAGGGCCACTAAGGACCACAAATTTATGACAGTCGATGGCCAG ATGCTGCCTATAGACGAAATCTTTGAGCGAGAGTTGGACCTCATGCGAGTTGACAACCTTCCTAAT

[0125] DnaE intein-N protein: ​​​​CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDR GEQEVFEYCLEDGSLIRATKDHKFMTVDG QMLPIDEIFERELDLMRVDNLPN

[0126] DnaE intein-C DNA: ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGA TATTGGAGTCGAAAGAGATCACAACTTTG CTCTGAAGAACGGATTCATAGCTTCTAAT

[0127] Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN

[0128] Cfa-N DNA: TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGCCTATTGGAAAGATTGTCGAAGAGAGAATTGA ATGCACAGTATATACTGTAGACAAGAATGGTTTCGTTTACACACAGCCCATTGCTCAATGGCACAATCGCGGCGAACAAG AAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACGAGCAACTAAAGATCATAAATTCATGACCACTGACGGGCAG ATGTTGCCAATAGATGAGATATTCGAGCGGGGCTTGGATCTCAAACAAGTGGATGGATTGCCA

[0129] Cfa-N protein: CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQ MLPIDEIFERGLDLKQVDGLP

[0130] Cfa-C DNA: ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAGGAAAGTAAAGATAATATCTCGAAAAAGTCT TGGTACCCAAAATGTCTATGATATTGGAGTGGAGAAAGATCACAACTTCCTTCTCAAGAACGGTCTCGTAGCCAGCAAC

[0131] Cfa-C protein: MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLVASN

[0132] To join the N-terminal part of split Cas9 and the C-terminal part of split Cas9, intein N and intein C can be fused to the N-terminal part of split Cas9 and the C-terminal part of split Cas9, respectively. For example , in some embodiments, intein-N is fused to the C-terminus of the N-terminal portion of split Cas9 , i.e., forming a structure of N--[N-terminal portion of split Cas9]-[intein-N]--C . In some embodiments, intein-C is fused to the N-terminus of the C-terminal portion of split Cas9 , i.e., forming a structure of N-[intein-C]--[C-terminal portion of split Cas9]-C . The mechanism of intein-mediated protein splicing for ligating the protein (such as split Cas9) where the intein is fused is, for example, as described in Shah et al., Chem Sci. 2014; 5(1):446-461, which is incorporated herein by reference . The mechanism of intein-mediated protein splicing for ligating the protein (such as split Cas9) where the intein is fused is, for example, as described in Shah et al., Chem Sci. 2014; 5(1):446-461, which is incorporated herein by reference and is known in the art. Methods for designing and using inteins are known in the art . The mechanism of intein-mediated protein splicing for ligating the protein (such as split Cas9) where the intein is fused is, for example, as described in Shah et al., Chem Sci. 2014; 5(1):446-461, which is incorporated herein by reference and are described, for example, in WO2014004336, WO2017132580, US20150344549 and US20180127780, each of which is hereby incorporated by reference in its entirety. and each of these is hereby incorporated by reference in its entirety.

[0133] The terms "isolated", "purified", or "biologically pure" refer to a substance from which the components that are normally associated with it in its natural state have been removed to varying degrees. "Isolation" indicates the degree of separation from the original source or surrounding environment. "Purification" indicates a higher degree of separation than isolation. A "purified" or "biologically pure" protein has other substances sufficiently removed so that impurities do not substantially affect the biological properties of the protein or cause other adverse results. That is, the nucleic acids or peptides of the present invention are purified if they are produced by recombinant DNA technology and substantially free of cellular material, viral material or medium, or if they are chemically synthesized and substantially free of chemical precursors or other chemical substances. Purity and homogeneity are typically determined using analytical chemistry techniques such as polyacrylamide gel electrophoresis or high performance liquid chromatography. The term "purified" can mean that the nucleic acid or protein gives rise to essentially one band on an electrophoretic gel. For proteins that can undergo modifications such as phosphorylation or glycosylation, different modifications can give rise to different isolated proteins that can be purified separately. when found in its natural state, have been removed to varying degrees. "Isolation" indicates the degree of separation from the original source or surrounding environment. "Purification" indicates a higher degree of separation than isolation. A "purified" or "biologically pure" protein has other substances sufficiently removed so that impurities do not substantially affect the biological properties of the protein or cause other adverse results. That is, the nucleic acids or peptides of the present invention are purified if they are produced by recombinant DNA technology and substantially free of cellular material, viral material or medium, or if they are chemically synthesized and substantially free of chemical precursors or other chemical substances. when produced by recombinant DNA technology, are substantially free of cellular material, viral material or medium, or when chemically synthesized, are substantially free of chemical precursors or other chemical substances. Purity and homogeneity are typically determined using analytical chemistry techniques such as polyacrylamide gel electrophoresis or high performance liquid chromatography. The term "purified" can mean that the nucleic acid or protein gives rise to essentially one band on an electrophoretic gel. For proteins that can undergo modifications such as phosphorylation or glycosylation, different modifications can give rise to different isolated proteins that can be purified separately. when produced by recombinant DNA technology, are substantially free of cellular material, viral material or medium, or when chemically synthesized, are substantially free of chemical precursors or other chemical substances. For proteins that can undergo modifications such as phosphorylation or glycosylation, different modifications can give rise to different isolated proteins that can be purified separately.

[0134] An "isolated polynucleotide" refers to a nucleic acid molecule of the present invention that is not It means a nucleic acid (e.g., DNA) that does not contain genes adjacent to the gene. Therefore, this term includes, for example, recombinant DNA incorporated into a vector; incorporated into an autonomously replicating plasmid or virus ; incorporated into the genomic DNA of prokaryotes or eukaryotes; or existing as a separate molecule independent of other sequences (e.g., cDNA or genomic or cDNA fragments generated by PCR or restriction endonuclease digestion). Further, this term includes RNA molecules transcribed from DNA molecules, as well as recombinant DNA that is part of a hybrid gene encoding additional polypeptide sequences.

[0135] The term "isolated polypeptide" means a polypeptide of the present invention separated from the components that accompany it in its natural state. Typically, a polypeptide is considered isolated when it is at least 60% free by weight from the proteins and natural organic molecules with which it associates in its natural state. Preferably, the preparation is at least 75%, more preferably at least 90%, and even more preferably at least 99% by weight of the polypeptide of the present invention. The isolated polypeptide of the present invention can be obtained, for example, by extraction from a natural source, expression of recombinant nucleic acid encoding such a polypeptide; or chemical synthesis of the protein. Purity can be measured by any suitable method, such as column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis. The term "linker" as used herein refers to two molecules or moieties (e.g., two components of a protein complex or ribonucleoprotein complex, or two domains of a fusion protein

[0136] A polynucleotide having a programmable DNA binding domain (e.g., dCas9) and a deamino acid A covalent linker (e.g., a covalent linker) that links the enzyme domains (e.g., adenosine deaminase) A linker can refer to a covalently (covalently) bonded, non-covalently bonded linker, chemical group, or molecule. Connecting different components or different parts of a component of a base editor system For example, in some embodiments, the linker can be a polynucleotide. Guide polynucleotide binding domain for programmable nucleotide binding domain and the catalytic domain of the deaminase. is capable of binding a CRISPR polypeptide and a deaminase. In one embodiment, the linker can link the Cas9 and the deaminase. The linker can link dCas9 and the deaminase. can link the nCas9 and the deaminase. In some embodiments, the linker is The guide polynucleotide and the deaminase can be linked. In the present invention, the linker comprises a deamination component of the base editor system and a polynucleotide protease. In some embodiments, the nucleotide binding moieties can be linked to In this case, the linker is a linker between the RNA-binding portion of the deamination component of the base editor system and the polynucleotide A nucleotide can be linked to a programmable nucleotide binding component. In some embodiments, the linker comprises an RNA binding portion of a deaminating component of a base editor system. and the RNA-binding portion of the polynucleotide programmable nucleotide-binding component. can be combined. A linker is placed between two groups, molecules, or other moieties, or is sandwiched by them, and is linked to each via covalent or non-covalent interaction, thus enabling the linking of two. In certain embodiments, the linker can be an organic molecule, group, polymer, or chemical moiety. In certain embodiments, the linker can be a polynucleotide. In certain embodiments, the linker can be a DNA linker. In certain embodiments, the linker can be an RNA linker. In certain embodiments, the linker can include an aptamer that can bind to a ligand. In certain embodiments, the ligand can be a carbohydrate, peptide, protein, or nucleic acid. In some embodiments, the linker can include an aptamer derived from a riboswitch. The riboswitch from which the aptamer is derived can be selected from a theophylline riboswitch, thiamine pyrophosphate (TPP) riboswitch, adenosylcobalamin (AdoCbl) riboswitch, S-adenosylmethionine (SAM) riboswitch, SAH riboswitch, flavin mononucleotide (FMN) riboswitch, tetrahydrofolate riboswitch, lysine riboswitch, glycine riboswitch, purine riboswitch, GlmS riboswitch, or prequeuosine 1 (PreQ1) riboswitch. In certain embodiments, the linker can include an aptamer that binds to any protein domain of a polypeptide or polypeptide ligand. In certain embodiments, the polypeptide ligand is a K- homologous (KH) domain, MS2 coat protein domain, PP7 coat protein domain, Ku binding motif and Ku protein, telomerase Sm7 binding motif and Sm7 protein , or may be an RNA recognition motif. In certain embodiments, the polypeptide ligand is a salt may be part of a base editor system component. For example, a nucleic acid base editing component may include a deaminase domain and an RNA recognition motif.

[0137] In certain embodiments, the linker may be an amino acid or multiple amino acids (e.g., a peptide or a protein). In certain embodiments, the linker is about 5-100 amino acids in length, e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20-30 , 30-40, 40-50, 50-60, 60-70, 70-80, 80-90 or 90-100 amino acids in length. In certain embodiments, the linker is about 100-150, 150-200, 200-250, 250-300, 300-350, 350-400, 400-450, or 450-500 amino acids in length. Longer or shorter linkers are also contemplated.

[0138] In certain embodiments, the linker connects the gRNA binding domain of an RNA programmable nuclease comprising a Cas9 nuclease domain to the catalytic domain of a nucleic acid editing protein (e.g., adenosine deaminase ). In certain embodiments, the linker connects dCas9 and a nucleic acid editing protein. For example, the linker is positioned between two groups, molecules, or other moieties, or is flanked by two groups, molecules, or other moieties and is covalently linked to each via a covalent bond, thus linking the two. In certain embodiments, the linker is an amino group and is covalently linked to each via a covalent bond, thus linking the two. In certain embodiments, the linker is an amino group and is covalently linked to each via a covalent bond, thus linking the two. In certain embodiments, the linker is an amino It is an amino acid or a plurality of amino acids (e.g., a peptide or a protein). In certain embodiments the linker is an organic molecule, group, polymer, or chemical moiety. In certain embodiments the linker is 5 to 200 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13 14, 15, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140, 150, 160, 175, 1 80, 190, or 200 amino acids in length. Longer or shorter linkers are also contemplated.

[0139] In some embodiments, the domains of the nucleic acid base editor are fused via a linker comprising the amino acid sequence SGGSSGSETPGTSESATPESSGGS, SGGSSGGSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEG SAPGSPAGSPTSTEEGTSTE PSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS. In some embodiments, the domains of the nucleic acid base editor are fused via a linker comprising the amino acid sequence SGSETPGTSESATPES, which may also be referred to as an XTEN linker. In some embodiments, the linker comprises the amino acid sequence SGGS. In some embodiments, the linker comprises (SGGS) . In some embodiments, the linker is (SGGS) n , (GGGS) n , (GGGGS) n , (G) n、 (EAAAK) n , (G GS) n , SGSETPGTSESATPES, or (XP) ncomprises a motif, or any combination thereof, wherein n is independently an integer from 1 to 30 and X is any amino acid. In some embodiments , n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15.

[0140] In one embodiment, the linker is 24 amino acids in length. In one embodiment, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In one embodiment, the linker is 40 amino acids in length. In one embodiment, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTES ATPESSGGSSGGSSGSSGGS. In one embodiment, the linker is 64 amino acids in length. In one embodiment, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS SGSETPGTSESATPESSGGSSGGS. In one embodiment, the linker is 92 amino acids in length In some embodiments, the linker comprises the amino acid sequence PGSPAGSPTSTEEGTSESATPESGPG TSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS.

[0141] As used herein, the term "marker" means any protein or polynucleotide having a change in expression level or activity associated with a disease or disorder.

[0142] As used herein, the term "mutation" refers to the substitution of a residue within a sequence, e.g., a nucleic acid or amino acid sequence, with another residue, or the deletion or insertion of one or more residues within the sequence. Mutations are typically identified herein by identifying the original residue and then ​ by identifying the positions of residues within the array and identifying the newly substituted residues, are described. Various methods for making amino acid substitutions (mutations) provided herein are well known in the art and are provided, for example, by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)). In some embodiments, the base editors disclosed herein generate a significant number of unintended mutations, such as unintended point mutations, without generating "intended mutations" such as point mutations in nucleic acids (e.g., nucleic acids within a target genome). In some embodiments, the intended mutations are mutations caused by a specific base editor (e.g., an adenosine base editor) that is specifically designed to cause that intended mutation and that binds to a guide polynucleotide (e.g., gRNA). (e.g., an adenosine base editor) that binds to a specific guide polynucleotide (e.g., gRNA). In general, mutations made or identified in a sequence (e.g., an amino acid sequence described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain that mutation. One of ordinary skill in the art will readily understand how to determine the positions of mutations in amino acid and nucleic acid sequences relative to a reference sequence.

[0143] In general, mutations made or identified in a sequence (e.g., the amino acid sequences described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain that mutation. One of ordinary skill in the art will readily understand how to determine the positions of mutations in amino acid and nucleic acid sequences relative to a reference sequence. The term "non-conservative mutation" refers to amino acid substitutions between different groups, e.g., from tryptophan to lysine, or from serine to phenylalanine, etc. In this case, One of ordinary skill in the art will readily understand how to determine the positions of mutations in amino acid and nucleic acid sequences relative to a reference sequence.

[0144] The term "non-conservative mutation" refers to amino acid substitutions between different groups, e.g., from tryptophan to lysine, or from serine to phenylalanine, etc. In this case, Non-conservative amino acid substitutions that do not interfere with or inhibit the biological activity of the functional variant are preferred. Non-conservative amino acid substitutions can enhance the biological activity of the functional variant such that the biological activity of the functional variant is increased compared to the wild-type protein. is preferred. Non-conservative amino acid substitutions can enhance the biological activity of the functional variant such that the biological activity of the functional variant is increased compared to the wild-type protein. is increased compared to the wild-type protein. can be

[0145] The terms "nuclear localization sequence", "nuclear localization signal", or "NLS" refer to an amino acid sequence that promotes the translocation of a protein into the cell nucleus. Nuclear localization sequences are known in the art and are described, for example, in Plank et al. of international PCT application PCT / EP 2000 / 011690 filed on November 23, 2000 and published as WO / 2001 / 038547 on May 31, 2001, the content of which is incorporated herein by reference for the disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, for example, as described by Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises an amino acid sequence selected from KRTADGS EFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAI VVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC. into the cell nucleus. Nuclear localization sequences are known in the art and are described, for example, in Plank et al. of international PCT application PCT / EP 2000 / 011690 filed on November 23, 2000 and published as WO / 2001 / 038547 on May 31, 2001, the content of which is incorporated herein by reference for the disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, for example, as described by Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises an amino acid sequence selected from KRTADGS EFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAI VVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC. and published as WO / 2001 / 038547 on May 31, 2001, the content of which is incorporated herein by reference for the disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, for example, as described by Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises an amino acid sequence selected from KRTADGS EFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAI VVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC. In some embodiments, the NLS comprises an amino acid sequence selected from KRTADGS EFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAI VVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC. In some embodiments, the NLS comprises an amino acid sequence selected from KRTADGS EFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAI VVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC. In other embodiments, the NLS is an optimized NLS, for example, as described by Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises an amino acid sequence selected from KRTADGS EFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAI VVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC. In other embodiments, the NLS is an optimized NLS, for example, as described by Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises an amino acid sequence selected from KRTADGS EFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAI VVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC. EFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAI VVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC. In some embodiments, the NLS comprises an amino acid sequence selected from KRTADGS EFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAI VVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.

[0146] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to a compound containing nucleobases and an acidic moiety, such as a nucleoside, nucleotide, or polymer of nucleotides. Typically, a polymeric nucleic acid, such as a nucleic acid molecule containing three or more nucleotides. As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to a compound containing nucleobases and an acidic moiety, such as a nucleoside, nucleotide, or polymer of nucleotides. Typically, a polymeric nucleic acid, such as a nucleic acid molecule containing three or more nucleotides. Typically, a polymeric nucleic acid, such as a nucleic acid molecule containing three or more nucleotides. is a linear molecule in which adjacent nucleotides are linked to each other via phosphodiester bonds. In certain embodiments, "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides). In certain embodiments, "nucleic acid" refers to an oligonucleotide chain comprising three or more individual nucleotide residues. As used herein, the terms "oligonucleotide" and "polynucleotide" may be used interchangeably to refer to a polymer of nucleotides (e.g., a chain of at least three nucleotides). In certain embodiments, "nucleic acid" encompasses RNA as well as single-stranded and / or double-stranded DNA. Nucleic acids can occur naturally, for example, in association with a genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, chromatin, or other naturally occurring nucleic acid molecule. On the other hand, nucleic acid molecules can be non-naturally occurring molecules, such as recombinant DNA or RNA, artificial chromosomes, engineered genomes, or fragments thereof, or synthetic DNA, RNA, DNA / RNA hybrids, or non-naturally occurring molecules containing non-naturally occurring nucleotides or nucleosides. Furthermore, the terms "nucleic acid", "DNA", "RNA", and / or similar terms include nucleic acid analogs, such as analogs having a backbone other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems and optionally purified, chemically synthesized, etc. In the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, such as chemically modified bases or sugars, and analogs having backbone modifications, as appropriate. Nucleic acid sequences are shown in the 5' to 3' direction unless otherwise indicated. e.g., genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, chromatin, chromatid, or other naturally occurring nucleic acid molecule. Alternatively, a nucleic acid molecule can be a non-naturally occurring molecule, such as a non-naturally occurring molecule, recombinant DNA or RNA, artificial chromosome, engineered genome, or fragment thereof, or synthetic DNA, RNA, DNA / RNA hybrid, or non-naturally occurring molecule containing non-naturally occurring nucleotides or nucleosides. Furthermore, the terms "nucleic acid", "DNA", "RNA", and / or similar terms include nucleic acid analogs, such as analogs having a backbone other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems and optionally purified, chemically synthesized, etc. In the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, such as chemically modified bases or sugars, and analogs having backbone modifications, as appropriate. Nucleic acid sequences are shown in the 5' to 3' direction unless otherwise indicated. e.g., non-naturally occurring molecule, recombinant DNA or RNA, artificial chromosome, engineered genome, or fragment thereof, or synthetic DNA, RNA, DNA / RNA hybrid, or non-naturally occurring molecule containing non-naturally occurring nucleotides or nucleosides. Furthermore, the terms "nucleic acid", "DNA", "RNA", and / or similar terms include nucleic acid analogs, such as analogs having a backbone other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems and optionally purified, chemically synthesized, etc. In the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, such as chemically modified bases or sugars, and analogs having backbone modifications, as appropriate. Nucleic acid sequences are shown in the 5' to 3' direction unless otherwise indicated. Furthermore, the terms "nucleic acid", "DNA", "RNA", and / or similar terms include nucleic acid analogs, such as analogs having a backbone other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems and optionally purified, chemically synthesized, etc. In the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, such as chemically modified bases or sugars, and analogs having backbone modifications, as appropriate. Nucleic acid sequences are shown in the 5' to 3' direction unless otherwise indicated. Furthermore, the terms "nucleic acid", "DNA", "RNA", and / or similar terms include nucleic acid analogs, such as analogs having a backbone other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems and optionally purified, chemically synthesized, etc. In the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, such as chemically modified bases or sugars, and analogs having backbone modifications, as appropriate. Nucleic acid sequences are shown in the 5' to 3' direction unless otherwise indicated. Furthermore, the terms "nucleic acid", "DNA", "RNA", and / or similar terms include nucleic acid analogs, such as analogs having a backbone other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems and optionally purified, chemically synthesized, etc. In the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, such as chemically modified bases or sugars, and analogs having backbone modifications, as appropriate. Nucleic acid sequences are shown in the 5' to 3' direction unless otherwise indicated. Furthermore, the terms "nucleic acid", "DNA", "RNA", and / or similar terms include nucleic acid analogs, such as analogs having a backbone other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems and optionally purified, chemically synthesized, etc. In the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, such as chemically modified bases or sugars, and analogs having backbone modifications, as appropriate. Nucleic acid sequences are shown in the 5' to 3' direction unless otherwise indicated. Furthermore, the terms "nucleic acid", "DNA", "RNA", and / or similar terms include nucleic acid analogs, such as analogs having a backbone other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems and optionally purified, chemically synthesized, etc. In the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, such as chemically modified bases or sugars, and analogs having backbone modifications, as appropriate. Nucleic acid sequences are shown in the 5' to 3' direction unless otherwise indicated. Furthermore, the terms "nucleic acid", "DNA", "RNA", and / or similar terms include nucleic acid analogs, such as analogs having a backbone other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems and optionally purified, chemically synthesized, etc. In the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, such as chemically modified bases or sugars, and analogs having backbone modifications, as appropriate. Nucleic acid sequences are shown in the 5' to 3' direction unless otherwise indicated. In certain embodiments, the nucleic acid is a natural nucleoside (e.g., adenosine, thymidine, guanosine , cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); a nucleoside analog (e.g., 2-aminoadenosine, 2-thiothym idine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2-ami noadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-pro pynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanine, O( 6)-methylguanine, and 2-thiocytidine); a chemically modified base; a biologically modified base (e.g., a methylated base); an inserted base; a modified sugar (e.g., 2’-fluoro-ribose, ribose, 2 ’-deoxyribose, arabinose, and hexose); and / or a modified phosphate group ( e.g., phosphorothioate and 5’-N-phosphoramidite linkages) or comprises one or more of them.

[0147] The term “nucleic acid programmable DNA binding protein” or “napDNAbp” is used interchangeably with “polynucleotide programmable nucleotide binding domain” and refers to a protein that associates with a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid or guide polynucleotide (e.g., gRNA) that guides the napDNAbp to a specific nucleic acid sequence. In certain embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable DNA binding domain. In certain embodiments, ​And the polynucleotide-programmable nucleotide binding domain is a polynu cleotide-programmable RNA binding domain. In certain embodiments, the po lynucleotide-programmable nucleotide binding domain is a Cas9 protein . The Cas9 protein can bind to a guide RNA that directs the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In certain embodiments, napDNAbp is a Cas9 domain , such as nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Non-limiting examples of nucleic acid-programmable DNA binding proteins include Cas9 (such as dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas 12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i. Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a , Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also called Csn1 or Csx12), Ca s10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Ca s12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc 1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4 . ​​, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Cs x1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa 2, Csa3, Csa4, Csa5, Type II Cas effector protein, Type V Cas effector protein, Type VI Cas effector protein, CARF, DinG, their homologs, or their modified or engineered versions. Other nucleic acid-programmable possible DNA-binding proteins may not be specifically listed in this disclosure, but are within the scope of this disclosure. For example, see Makarova et al., “Classification and Nomenclature of CRISPR R-Cas Systems: Where from Here?” CRISPR J. 2018 Oct;1:325-336. doi: 10.1089 / cris pr.2018.0033; Yan et al., “Functionally diverse type V CRISPR-Cas systems” Sci ence. 2019 Jan 4;363(6422):88-91. doi: 10.1126 / science.aav7271 (the entire contents of each are incorporated herein by reference).

[0148] The terms “nucleobase,” “nitrogenous base,” or “base” are used interchangeably herein and refer to nitrogen-containing biological compounds that form nucleosides, which are components of nucleotides. The ability of nucleobases to form base pairs and stack on top of each other is ribo ​​​​directly result in long-chain helical structures such as nucleic acid (RNA) and deoxyribonucleic acid (DNA). Adeni ne (A), cytosine (C), guanine (G), thymine (T), and uracil (U) are five nucle ic acid bases, which are called primary bases or standard bases. Adenine and guanine are derived from purine , and cytosine, uracil, and thymine are derived from pyrimidine. DNA and RNA can also contain other (non-primary) modified bases. Non-limiting exemplary modified nucleic acid bases include hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydroxymethylcytosine. Hypoxanthine and xanthine can be generated by the presence of mutagens, and both are generated by deamination (replacing an amine group with a carbonyl group). Hypoxanthine can be modified from adenine. Xanthine can be modified from guanine. Uracil can be generated from the deamination of cytosine. A "nucleoside" consists of one nucleic acid base and one pentose sugar (either ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine , 5-methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, de oxyuridine, and deoxycytidine. Examples of nucleosides with modified nucleic acid bases include inosine (I), xanthosine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytosine (m5C), and pseudouridine (Ψ). A "nucleotide" consists of one nucleic acid base, one pentose sugar (ribose or deoxy ribose), and one phosphate group. Examples of nucleotides include adenosine monophosphate (AMP), guanosine monophosphate (GMP), uridine monophosphate (UMP), cytidine monophosphate (CMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), thymidine monophosphate (dTMP), It consists of either ribose) and at least one phosphate group.

[0149] As used herein, the term "nucleobase editing domain" or "nucleobase editing protein" refers to a protein or enzyme that can catalyze nucleobase modifications in RNA or DNA, such as the deamination of cytosine (or cytidine) to uracil (or uridine) or thymine (or thymidine), and the deamination of adenine (or adenosine) to hypoxanthine (or inosine), as well as non-template nucleotide addition and insertion. In certain embodiments, the nucleobase editing domain is a deaminase domain (e.g., adenine deaminase or adenosine deaminase). In certain embodiments, the nucleobase editing domain can be a naturally occurring nucleobase editing domain. In some embodiments, the nucleobase editing domain can be a nucleobase editing domain engineered or evolved from a naturally occurring nucleobase editing domain. The nucleobase editing domain can be derived from any organism, such as bacteria, human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse.

[0150] As used herein, "obtaining" as in "obtaining an agent" includes synthesizing, purchasing, or otherwise acquiring that agent.

[0151] As used herein, "patient" or "subject" refers to a mammalian subject or individual diagnosed with, at risk of developing, or suspected of having or developing a disease or disorder. In certain embodiments, the term "patient" refers to an individual who develops a disease or disorder Refers to mammalian subjects with a higher than average likelihood. Exemplary patients include humans, non-human primates, cats, dogs, pigs, cows, cats, horses, camels, llamas, goats, sheep, rodents (e.g., mice, rabbits, rats, guinea pigs) and other mammals that can benefit from the treatments disclosed herein. Exemplary human patients can be male and / or female.

[0152] As used herein, "a patient in need thereof" or "a subject in need thereof" refers to a patient diagnosed with or suspected of having a disease or disorder, e.g., but not limited to, glycogen storage disease type 1 (GSD1 or Von Gierke disease ).

[0153] The terms "pathogenic variant", "pathogenic mutation", "disease-causing variant", "disease-causing mutation", "harmful variant" or "predisposing mutation" refer to genetic changes or mutations that increase an individual's susceptibility or predisposition to a particular disease or disorder. In certain embodiments, a pathogenic variant includes those in which at least one wild-type amino acid in the protein encoded by the gene is replaced by at least one pathogenic amino acid.

[0154] The terms "protein", "peptide", "polypeptide", and their grammatical equivalents are used interchangeably herein and refer to polymers of amino acid residues linked to each other by peptide (amide) bonds. This term refers to proteins, peptides or polypeptides of any size, structure or function. Typically, a protein, peptide, or polypeptide is at least 3 amino acids in length. A protein, peptide, or polypeptide ​A peptide can refer to an individual protein or a group of proteins. One or more amino acids in a protein, peptide, or polypeptide can be modified by the addition of chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups, linkers for attachment, functionalization, or other modifications. A protein, peptide, or polypeptide can also be a single molecule or a multimolecular complex. A protein, peptide, or polypeptide can be merely a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide can be naturally occurring, recombinant, or synthetic, or any combination thereof. As used herein, the term "fusion protein" refers to a hybrid polypeptide that contains protein domains derived from at least two different proteins. One protein can be located at the amino-terminal (N-terminal) portion or the carboxy-terminal (C-terminal) protein of the fusion protein, thus forming an amino-terminal fusion protein or a carboxy-terminal fusion protein, respectively. A protein can include different domains, such as a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9 that induces binding of the protein to a target site) and a nucleic acid cleavage domain, or the catalytic domain of a nucleic acid editing protein. In certain embodiments, a protein includes a proteinaceous portion, such as an amino acid sequence that constitutes a nucleic acid binding domain, and an organic compound, such as a compound that can act as a nucleic acid cleavage agent. In certain embodiments, a protein includes a nucleic acid (e.g., RNA or DNA) One or more amino acids in a protein, peptide, or polypeptide can be modified by the addition of chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups, linkers for attachment, functionalization, or other modifications One or more amino acids in a protein, peptide, or polypeptide can be modified by the addition of chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups, linkers for attachment, functionalization, or other modifications One or more amino acids in a protein, peptide, or polypeptide can be modified by the addition of chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups, linkers for attachment, functionalization, or other modifications A protein, peptide, or polypeptide can also be a single molecule or a multimolecular complex A protein, peptide, or polypeptide can also be a single molecule or a multimolecular complex A protein, peptide, or polypeptide can be merely a fragment of a naturally occurring protein or peptide A protein, peptide, or polypeptide can be naturally occurring, recombinant, or synthetic, or any combination thereof A protein, peptide, or polypeptide can be naturally occurring, recombinant, or synthetic, or any combination thereof As used herein, the term "fusion protein" refers to a hybrid polypeptide that contains protein domains derived from at least two different proteins As used herein, the term "fusion protein" refers to a hybrid polypeptide that contains protein domains derived from at least two different proteins One protein can be located at the amino-terminal (N-terminal) portion or the carboxy-terminal (C-terminal) protein of the fusion protein, thus forming an amino-terminal fusion protein or a carboxy-terminal fusion protein, respectively One protein can be located at the amino-terminal (N-terminal) portion or the carboxy-terminal (C-terminal) protein of the fusion protein, thus forming an amino-terminal fusion protein or a carboxy-terminal fusion protein, respectively A protein can include different domains, such as a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9 that induces binding of the protein to a target site) and a nucleic acid cleavage domain, or the catalytic domain of a nucleic acid editing protein A protein can include different domains, such as a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9 that induces binding of the protein to a target site) and a nucleic acid cleavage domain, or the catalytic domain of a nucleic acid editing protein In certain embodiments, a protein includes a proteinaceous portion, such as an amino acid sequence that constitutes a nucleic acid binding domain, and an organic compound, such as a compound that can act as a nucleic acid cleavage agent In certain embodiments, a protein includes a proteinaceous portion, such as an amino acid sequence that constitutes a nucleic acid binding domain, and an organic compound, such as a compound that can act as a nucleic acid cleavage agent In certain embodiments, a protein includes a proteinaceous portion, such as an amino acid sequence that constitutes a nucleic acid binding domain, and an organic compound, such as a compound that can act as a nucleic acid cleavage agent ) is complexed with or associated with a nucleic acid. Any of the proteins provided herein can be produced by any method known in the art. For example, the proteins provided herein can be produced via recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for the expression and purification of recombinant proteins are well known and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), the entire contents of which are incorporated herein by reference. The polypeptides and proteins (including functional portions and functional variants thereof) disclosed herein can contain synthetic amino acids in place of one or more naturally occurring amino acids. Such synthetic amino acids are known in the art and include, for example, aminocyclohexanecarboxylic acid, norleucine, α-amino-n-decanoic acid, homoserine, S-acetylaminomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine, β-phenylserine β-hydroxyphenylalanine, phenylglycine, α-naphthylalanine, cyclohexylalanine, cyclohexylglycine, indoline-2-carboxylic acid, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, a nd the like. Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press , Cold Spring Harbor, N.Y. (2012)) and are incorporated herein by reference in their entirety.

[0155] The polypeptides and proteins disclosed herein (including functional portions and functional variants thereof) can contain synthetic amino acids in place of one or more naturally occurring amino acids. Such synthetic amino acids are known in the art and include, for example, aminocyclohexanecarboxylic acid, norleucine, α-amino-n-decanoic acid, homoserine, S-acetylaminomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine, β-phenylserine β-hydroxyphenylalanine, phenylglycine, α-naphthylalanine, cyclohexylalanine, cyclohexylglycine, indoline-2-carboxylic acid, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, a nd the like. Such synthetic amino acids are known in the art and include, for example, aminocyclohexanecarboxylic acid, norleucine, α-amino-n-decanoic acid, homoserine, S-acetylaminomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine, β-phenylserine β-hydroxyphenylalanine, phenylglycine, α-naphthylalanine, cyclohexylalanine, cyclohexylglycine, indoline-2-carboxylic acid, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, a nd the like. Such synthetic amino acids are known in the art and include, for example, aminocyclohexanecarboxylic acid, norleucine, α-amino-n-decanoic acid, homoserine, S-acetylaminomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine, β-phenylserine β-hydroxyphenylalanine, phenylglycine, α-naphthylalanine, cyclohexylalanine, cyclohexylglycine, indoline-2-carboxylic acid, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, a nd the like. Minomalonic acid monoamide, N'-benzyl-N'-methyllysine, N',N'-dibenzyl-lysine, 6- Hydroxylysine, ornithine, α-aminocyclopentanecarboxylic acid, aminocyclohe xanecarboxylic acid, aminocyclohexanecarboxylic acid, α-aminocycloheptanecarbo nic acid, α-(2-amino-2-norbornane)-carboxylic acid, α,γ-diaminobutyric acid, α,β-diamin opropionic acid, homophenylalanine, and α-tert-butylglycine. Polypeptides and proteins can be conjugated to post-translational modifications of one or more amino acids of the polypeptide construct. Non-limiting examples of post-translational modifications include acylation, including phosphorylation, acetylation and formylation, glycosylation (including N-linked and O-linked), amidation, hydroxylation, alkylation, including methylation and ethylization, ubiquitination, pyro lidone carboxylic acid addition, disulfide bridge formation, sulfation, myristoylation, palmito ylation, isoprenylation, farnesylation, geranylation, glypiation, lipoylation and iodination.

[0156] As used herein in connection with a protein or nucleic acid, the term "recombinant" refers to a protein or nucleic acid that is not found in nature but is the product of human engineering. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that contains at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7 mutations as compared to any naturally occurring sequence.

[0157] "Decrease" means a negative change of at least 10%, 25%, 50%, 75%, or 100%. 。

[0158] "Reference" means a standard or control condition. In one embodiment , the reference is a wild-type or healthy cell. In other embodiments, but not limited to, the reference is a non-treated cell that has not been exposed to the test conditions or has been exposed to a placebo or normal saline, medium, buffer , and / or a control vector that does not carry the polynucleotide of interest.

[0159] "Reference sequence" is a defined sequence used as a basis for sequence comparison. The reference sequence can be a subset or the entirety of a particular sequence; for example, a segment of a full-length cDNA or gene sequence, or a full cDNA or gene sequence. For polypeptides, the length of the reference polypeptide sequence is generally at least about 16 amino acids, at least about 20 amino acids, at least about 25 amino acids, about 35 amino acids, about 50 amino acids, or about 100 amino acids. For nucleic acids, the length of the reference nucleic acid sequence is generally at least about 50 nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides or about 300 nucleotides or any integer between or around them. In some embodiments , the reference sequence is the wild-type sequence of the protein of interest. In other embodiments, the reference sequence is the polynucleotide sequence encoding the wild-type protein.

[0160] The terms "RNA-programmable nuclease" and "RNA-guided nuclease" refer to cleavage Used with one or more RNAs that are not the target (e.g., bound to or associated with it) . In certain embodiments, when the RNA programmable nuclease is complexed with RNA, it can be referred to as a nuclease:RNA complex. Typically, the bound RNA is called a guide RNA (gRNA). The gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule can be referred to as a single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to a guide RNA that exists either as a single molecule or as a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species contains two domains: (1) a domain that shares homology with the target nucleic acid (e.g., and directs the binding of the Cas9 complex to the target), and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to tracrRNA as provided in Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those containing domain 2) are provided in U.S. Provisional Patent Application No. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application No. 61 / 874,746, filed September 6, 2013, entitled "Delivery System F can do. A gRNA that exists as a single RNA molecule can be referred to as a single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to a guide RNA that exists either as a single molecule or as a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species contains two domains: (1) a domain that shares homology with the target nucleic acid (e.g., and directs the binding of the Cas9 complex to the target), and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to tracrRNA as provided in Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those containing domain 2) are provided in U.S. Provisional Patent Application No. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application No. 61 / 874,746, filed September 6, 2013, entitled "Delivery System F can do. A gRNA that exists as a single RNA molecule can be referred to as a single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to a guide RNA that exists either as a single molecule or as a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species contains two domains: (1) a domain that shares homology with the target nucleic acid (e.g., and directs the binding of the Cas9 complex to the target), and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to tracrRNA as provided in Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those containing domain 2) are provided in U.S. Provisional Patent Application No. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application No. 61 / 874,746, filed September 6, 2013, entitled "Delivery System F can do. A gRNA that exists as a single RNA molecule can be referred to as a single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to a guide RNA that exists either as a single molecule or as a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species contains two domains: (1) a domain that shares homology with the target nucleic acid (e.g., and directs the binding of the Cas9 complex to the target), and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to tracrRNA as provided in Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those containing domain 2) are provided in U.S. Provisional Patent Application No. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application No. 61 / 874,746, filed September 6, 2013, entitled "Delivery System F can do. A gRNA that exists as a single RNA molecule can be referred to as a single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to a guide RNA that exists either as a single molecule or as a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species contains two domains: (1) a domain that shares homology with the target nucleic acid (e.g., and directs the binding of the Cas9 complex to the target), and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to tracrRNA as provided in Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those containing domain 2) are provided in U.S. Provisional Patent Application No. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application No. 61 / 874,746, filed September 6, 2013, entitled "Delivery System F can do. A gRNA that exists as a single RNA molecule can be referred to as a single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to a guide RNA that exists either as a single molecule or as a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species contains two domains: (1) a domain that shares homology with the target nucleic acid (e.g., and directs the binding of the Cas9 complex to the target), and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to tracrRNA as provided in Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those containing domain 2) are provided in U.S. Provisional Patent Application No. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application No. 61 / 874,746, filed September 6, 2013, entitled "Delivery System F e.g., and directs the binding of the Cas9 complex to the target), and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to tracrRNA as provided in Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those containing domain 2) are provided in U.S. Provisional Patent Application No. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application No. 61 / 874,746, filed September 6, 2013, entitled "Delivery System F e.g., and directs the binding of the Cas9 complex to the target), and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to tracrRNA as provided in Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those containing domain 2) are provided in U.S. Provisional Patent Application No. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application No. 61 / 874,746, filed September 6, 2013, entitled "Delivery System F e.g., and directs the binding of the Cas9 complex to the target), and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to tracrRNA as provided in Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those containing domain 2) are provided in U.S. Provisional Patent Application No. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application No. 61 / 874,746, filed September 6, 2013, entitled "Delivery System F e.g., and directs the binding of the Cas9 complex to the target), and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to tracrRNA as provided in Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those containing domain 2) are provided in U.S. Provisional Patent Application No. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application No. 61 / 874,746, filed September 6, 2013, entitled "Delivery System F e.g., and directs the binding of the Cas9 complex to the target), and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to tracrRNA as provided in Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those containing domain 2) are provided in U.S. Provisional Patent Application No. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application No. 61 / 874,746, filed September 6, 2013, entitled "Delivery System F e.g., and directs the binding of the Cas9 complex to the target), and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to tracrRNA as provided in Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those containing domain 2) are provided in U.S. Provisional Patent Application No. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application No. 61 / 874,746, filed September 6, 2013, entitled "Delivery System F .S.S.N. 61 / 874,682, entitled "Switchable Cas9 Nucleases and Uses Thereof" and U.S. Provisional Patent Application No. 61 / 874,746, filed September 6, 2013, entitled "Delivery System F 3 years September 6, 2013, entitled "Delivery System F can be found in 「or Functional Nucleases」, and the entire contents of each of these are hereby incorporated by reference in their entirety In some embodiments, the gRNA includes two or more of domains (1) and (2) and may be referred to as an 「extended gRNA」. For example, an extended gRNA will bind to two or more Cas9 proteins and bind to the target nucleic acid at two or more distinct regions, as described herein. The gRNA includes a nucleotide sequence complementary to the target site, which mediates the binding of the nuclease / RNA complex to the target site and provides the sequence specificity of the nuclease:RNA complex and provides the sequence specificity of the nuclease:RNA complex. and provides the sequence specificity of the nuclease:RNA complex. and provides the sequence specificity of the nuclease:RNA complex.

[0161] In one aspect, the RNA programmable nuclease is a (CRISPR associated system) Cas9 endonuclease, such as Cas9 (Csnl) from Streptococcus pyogenes (e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes." Ferretti J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C, Sezat e S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z. e S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z. , Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663(2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chylinski K ., Sharma CM., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Vogel J., Charpe ntier E., Nature 471:602-607(2011) reference).

[0162] RNA-programmable nucleases (such as Cas9) target DNA cleavage sites and use RNA:DNA hybridization, so these proteins can, in principle, target any sequence specified by the guide RNA. Site-specific cleavage (e.g., for modifying the genome) using RNA-programmable nucleases such as Cas9 is known in the art (e.g., Cong, L. et al., Multipl ex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013), Ma li, P. et ah, RNA-guided human genome engineering via Cas9. Science 339, 823-82 6 (2013), Hwang, W.Y. et al., Efficient genome editing in zebrafish using a CRIS PR-Cas system. Nature biotechnology 31, 227-229 (2013), Jinek, M. et ah, RNA-pr ogrammed genome editing in human cells. eLife 2, e00471 (2013), Dicarlo, J.E. et et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic acids research (2013), Jiang, W. et al. RNA-guided editing of bacterial g enomes using CRISPR-Cas systems. Nature biotechnology 31, 233-239 (2013). See also , the entire contents of each of which are incorporated herein by reference).

[0163] The term "single nucleotide polymorphism (SNP)" refers to a variation in a single nucleotide that occurs at a specific position in the genome where each variation exists to a certain extent (e.g., >1%) that can be recognized within a population. For example at a specific base position in the human genome, the C nucleotide may occur in most individuals, but in a few individuals that position is occupied by A. This indicates that there is an SNP at this specific position, meaning that the two nucleotide variations, C or A, are the alleles at this position . SNPs underlie differences in susceptibility to diseases. The severity of a disease and the body's response to treatment are also manifestations of genetic variations. SNPs can be present in the coding region of a gene, the non-coding region of a gene, or the intergenic region (the region between genes). In certain embodiments , SNPs within the coding sequence do not necessarily change the amino acid sequence of the protein produced due to the degeneracy of the genetic code . There are two types of SNPs in the coding region: synonymous SNPs and non-synonymous SNPs . Synonymous SNPs do not affect the protein sequence, while non-synonymous SNPs change the amino acid sequence of the protein . There are two types of non-synonymous SNPs: missense and nonsense. Coding for a protein SNPs that are not in the region to be done may affect gene splicing, transcription factor binding, messenger RNA degradation, or the sequence of non-coding RNA. Gene expression affected by this type of SNP is called eSNP (expressed SNP) and can be upstream or downstream of the gene. A single nucleotide variant (SNV) is a single nucleotide variation without frequency limitation and can occur in somatic cells. Somatic single nucleotide variations can also be called single nucleotide modifications. "Specifically binds" means recognizing and binding to the polypeptide and / or nucleic acid molecule of the present invention, but not substantially recognizing and binding to other molecules in a sample (e.g., a biological sample), such as a nucleic acid molecule, polypeptide, or their complex (e.g., a nucleic acid programmable DNA binding domain and a guide nucleic acid), a compound, or a molecule.

[0164] Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding the polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but typically exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence can typically hybridize to at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding the polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but typically exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence can typically hybridize to at least one strand of a double-stranded nucleic acid molecule. nucleic acid molecule, polypeptide, or their complex (e.g., a nucleic acid programmable DNA binding domain and a guide nucleic acid), a compound, or a molecule.

[0165] Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding the polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but typically exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence can typically hybridize to at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding the polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but typically exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence can typically hybridize to at least one strand of a double-stranded nucleic acid molecule. nucleic acid sequence, but typically exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence can typically hybridize to at least one strand of a double-stranded nucleic acid molecule. A polynucleotide having "substantial identity" to an endogenous sequence can typically hybridize to at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding the polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but typically exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence can typically hybridize to at least one strand of a double-stranded nucleic acid molecule. polypeptide or a fragment thereof. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but typically exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence can typically hybridize to at least one strand of a double-stranded nucleic acid molecule. to an endogenous sequence having "substantial identity" can typically hybridize to at least one strand of a double-stranded nucleic acid molecule. can hybridize with at least one strand. "Hybridize" means forming a pair to form a double-stranded molecule between complementary polynucleotide sequences (e.g., the genes described herein) or between parts thereof under various stringency conditions. (See, for example, Wahl, G. M. and S. L. Berger (1987) Methods Enzymol. 152:399; Kimmel, A. R. (1987) Methods Enzymol. 152:507). For example, stringent salt concentrations are usually less than about 750 mM NaCl and 75 mM trisodium citrate, preferably less than about 500 mM NaCl and 50 mM trisodium citrate, more preferably less than about 250 mM NaCl and 25 mM trisodium citrate. Low stringency hybridization can be obtained in the absence of an organic solvent, such as formamide, while high stringency hybridization can be obtained in the presence of at least about 35% formamide, more preferably at least about 50% formamide. Stringent temperature conditions will usually include a temperature of at least about 30°C, more preferably at least about 37°C, most preferably at least about 42°C. During hybridization, various additional parameters such as the concentration of a surfactant (e.g., sodium dodecyl sulfate (SDS)) and the inclusion or exclusion of carrier DNA are well known to those skilled in the art. By combining these various conditions as needed, various levels of stringency can be achieved. In one embodiment, hybridization is performed at 30°C with 750 mM NaCl, 75 mM trisodium citrate. For example, see Wahl, G. M. and S. L. Berger (1987) Methods Enzymol. 152:399; Kimmel, A. R . (1987) Methods Enzymol. 152:507.

[0166] trisodium citrate. trisodium citrate, preferably less than about 500 mM NaCl and 50 mM trisodium citrate, more preferably less than about 250 mM NaCl and 25 mM trisodium citrate. Low stringency hybridization can be obtained in the absence of an organic solvent, such as formamide, while high stringency hybridization can be obtained in the presence of at least about 35% formamide more preferably at least about 50% formamide. Stringent temperature conditions will usually include a temperature of at least about 30°C, more preferably at least about 37°C most preferably at least about 42°C. During hybridization, the concentration of a surfactant (e.g., sodium dodecyl sulfate (SDS)), and the inclusion or exclusion of carrier DNA and other various additional parameters are well known to those skilled in the art. By combining these various conditions as needed various levels of stringency can be achieved. In one embodiment, hybridization is performed at 30°C with 750 mM NaCl, 75 mM trisodium citrate. or exclusion of carrier DNA and other various additional parameters are well known to those skilled in the art. By combining these various conditions as needed various levels of stringency can be achieved. In one embodiment, hybridization is performed at 30°C with 750 mM NaCl, 75 mM trisodium citrate. It occurs in trisodium citrate and 1% SDS. In another embodiment, the hybridization occurs at 37°C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide and 100 μg / ml denatured salmon sperm DNA (ssDNA). In another embodiment, the hybr idization occurs at 42°C in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS , 50% formamide and 200 μg / ml ssDNA. Useful variations of these conditions will be readily apparent to those skilled in the art.

[0167] In most applications, the washing step following hybridization is also stringen cy different. The washing stringency conditions can be defined by salt concentrat ion and temperature. As described above, the washing stringency can be increased by decreasing the salt concentration or increasing the temperature. For example, the salt concentration of the washing agent is preferably less than about 30 mM NaCl and 3 mM trisodium citrate, and most preferably less than about 15 mM NaCl and 1.5 mM trisodium citrate. The stringent temperature conditions for the washing step usually include a temperature of at least about 25°C, more pref erably at least about 42°C, even more preferably at least about 68°C. In one em bodiment, the washing step is carried out at 25°C in 30 mM NaCl, 3 mM trisod ium citrate, and 0.1% SDS. In a more preferred embodiment, the washing step is The washing step is carried out at 68°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Further variations of these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Bent on and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. S ci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Clon ing Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecula r Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.

[0168] "Split" means being split into two or more fragments.

[0169] "Split Cas9 protein" or "split Cas9" refers to a Cas9 protein provided as an N-terminal fragment and a C-terminal fragment encoded by two separate nucleotide sequences. The polypeptides corresponding to the N-terminal and C-terminal portions of the Cas9 protein can be spliced to form a "reconstituted" Cas9 protein. In certain embodiments, The Cas9 protein is described, for example, in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-9 49, 2014, or as described in Jiang et al. (2016) Science 351: 867-871. PDB file: 5F9R (each incorporated herein by reference), and is split into two fragments within the disordered region of the protein In some embodiments, the protein is split into two fragments at any C, T, A, or S within the region of approximately amino acids A292-G364, F445-K483, or E565-T637 of SpCas9 or at any corresponding position in any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or other napDNAbp In one aspect, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574 In some embodiments, the process of splitting the protein into two fragments is referred to as "splitting" the protein

[0170] In other embodiments, the N-terminal portion of the Cas9 protein comprises amino acids 1-573 or 1-637 of S. pyogenes Cas9 wild type (SpCas9) (NCBI reference sequence: NC_002737.2, Uniprot reference sequence: Q99ZW2), and the C-terminal portion of the Cas9 protein comprises amino acids 574-1368 or a portion of 638-1368 of SpCas9 wild type, or the corresponding positions thereof (NCBI reference sequence: NC_002737.2, Uniprot reference sequence: Q99ZW2), and the C-terminal portion of the Cas9 protein comprises amino acids 574-1368 or a portion of 638-1368 of SpCas9 wild type, or the corresponding positions thereof (NCBI reference sequence: NC_002737.2, Uniprot reference sequence: Q99ZW2), and the C-terminal portion of the Cas9 protein comprises amino acids 574-1368 or a portion of 638-1368 of SpCas9 wild type, or the corresponding positions thereof

[0171] The C-terminal portion of the split Cas9 is linked to the N-terminal portion of the split Cas9 to form a complete Cas9 protein ​​​​It can form proteins. In some embodiments, the C-terminus of the Cas9 protein starts where the N-terminal part of the Cas9 protein ends. Thus, in some embodiments, the C-terminal part of split Cas9 contains the amino acids (551-651)-1368 of spCas9 " (551-651)-1368" means starting from the amino acids between amino acids 551 to 651 (including both ends) and ending at amino acid 1368. For example, the C-terminal part of split Cas9 is, for spCas9, amino acids 551-1368, 552-1368, 553-1368, 554-1368, 555-1368, 556-1368, 55 7-1368, 558-1368, 559-1368, 560-1368, 561-1368, 562-1368, 563-1368, 564-1368, 56 5-1368, 566-1368, 567-1368, 568-1368, 569-1368, 570-1368, 571-1368, 572-1368, 57 3-1368, 574-1368, 575-1368, 576-1368, 577-1368, 578-1368, 579-1368, 580-1368, 58 1-1368, 582-1368, 583-1368, 584-1368, 585-1368, 586-1368, 587-1368, 588-1368, 58 9-1368, 590-1368, 591-1368, 592-1368, 593-1368, 594-1368, 595-1368, 596-1368, 59 7-1368, 598-1368, 599-1368, 600-1368, 601-1368, 602-1368, 603-1368, 604-1368, 60 5-1368, 606-1368, 607-1368, 608-1368, 609-1368, 610-1368, 611-1368, 612-1368, 61 5-1368, 606-1368, 607-1368, 608-1368, 609-1368, 610-1368, 611-1368, 612-1368, 61 3-1368, 614-1368, 615-1368, 616-1368, 617-1368, 618-1368, 619-1368, 620-1368, 62 1-1368, 622-1368, 623-1368, 624-1368, 625-1368, 626-1368, 627-1368, 628-1368, 62 9-1368, 630-1368, 631-1368, 632-1368, 633-1368, 634-1368, 635-1368, 636-1368, 63 7-1368, 638-1368, 639-1368, 640-1368, 641-1368, 642-1368, 643-1368, 644-1368, 64 5-1368, 646-1368, 647-1368, 648-1368, 649-1368, 650-1368, or any one of 651-1368 may include one part. In some embodiments, the C-terminal portion of the split Cas9 protein includes the portion of amino acids 574-1368 or 638-1368 of SpCas9.

[0172] "Subject" means a mammal, including but not limited to, a human, or a non-human mammal such as a cow , horse, dog, sheep, or cat. As subjects, there are included, but not limited to, livestock such as beef cattle, goats, chickens, horses, pigs, rabbits, and sheep that are raised to produce labor and provide articles such as food. Livestock are mentioned.

[0173] "Substantially identical" means a polypeptide or nucleic acid molecule that exhibits at least 50% identity to a reference amino acid sequence (e.g., any one of the amino acid sequences described herein) or a nucleic acid sequence (e.g., any one of the nucleic acid sequences described herein) . In one embodiment In a form, such an array is at least 60%, 80%, 85%, 90%, 95%, or even 99% identical at the amino acid level or nucleic acid with the array used for comparison. to be at least 60%, 80%, 85%, 90%, 95%, or even 99% identical.

[0174] Sequence identity is typically measured using sequence analysis software (e.g., the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Mad ison, Wis. 53705's Sequence Analysis Software Package, BLAST, BESTFIT, COBALT, EM BOSS Needle, GAP, or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning a degree of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; as partic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine , arginine; phenylalanine, tyrosine. In an exemplary approach for determining the degree of identity, the BLAST program can be used, and the probability score between e and e indicates closely related sequences. COBALT is used, for example, with the following parameters : -3 and e -100 The probability score between indicates closely related sequences. COBALT is used, for example, with the following parameters : a) Alignment parameters: Gap penalties -11, -1 and End-Gap penalties -5, -1, b) CDD parameters: Use RPS BLAST on; Blast E-value 0.003; Find Conserved col mns and Recompute on c) Query clustering parameters: Use query clusters on; Word Size 4; M ax cluster distance 0.8; Alphabet Regular. The EMBOSS Needle is used, for example, with the following parameters: a) Matrix: BLOSUM62; b) GAP OPEN: 10; c) GAP EXTEND: 0.5; d) OUTPUT FORMAT: pair; e) END GAP PENALTY: false; f) END GAP OPEN: 10; and g) END GAP EXTEND: 0.5.

[0175] The term "target site" refers to a sequence within a nucleic acid molecule that is modified by a nucleobase editor. In one embodiment, the target site is a sequence that is targeted by a deaminase or a fusion enzyme comprising the same. It is deaminated by a synthase (e.g., adenine deaminase).

[0176] As used herein, the terms "treat," "treating," and "treatment" refer to "Treatment" and the like are intended to alleviate or ameliorate a disorder and / or its associated symptoms. It refers to the process of treating or obtaining a desired pharmacological and / or physiological effect. Treating a condition requires the complete elimination of the associated disorder, condition or symptom. It will be understood that there is not (nor is complete removal excluded). In some aspects, the effect is therapeutic, i.e., without limitation, the effect is the disease and / or harmful symptoms resulting therefrom are partially or completely reduced, decreased, removed, alleviated , mitigated, attenuated, or cured. In one aspect, the effect is prophylactic, i.e., the effect protects or prevents the occurrence or recurrence of a disease or condition. For this purpose, the methods of the present disclosure involve administering a therapeutically effective amount of the composition as described herein .

[0177] "Uracil glycosylase inhibitor", or "UGI" means a factor that inhibits the uracil excision repair system. In one embodiment, the factor is a protein or fragment thereof that binds to the host uracil - DNA glycosylase and prevents the removal of uracil residues from DNA. In one embodiment, UGI is a protein, fragment thereof, or domain that can inhibit the uracil - DNA glycosylase base excision repair enzyme . In some aspects, the UGI domain includes wild - type UGI or a modified version thereof. In some embodiments , the UGI domain includes a fragment of the exemplary amino acid sequences presented below. In some embodiments, the UGI fragment includes an amino acid sequence that includes at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the exemplary UGI sequences provided below. In some embodiments, UGI is as follows . In some embodiments, the UGI fragment includes an amino acid sequence that includes at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the exemplary UGI sequences provided below. In some embodiments, UGI is as follows . In some embodiments, UGI is as follows As described below, it includes an amino acid sequence homologous to the exemplary UGI amino acid sequence or a fragment thereof. In some embodiments, UGI or a portion thereof has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9% or 100% identity to the wild-type UGI or UGI sequence or a portion thereof, as described below. An exemplary UGI includes the following amino acid sequence: As described below, it includes an amino acid sequence homologous to the exemplary UGI amino acid sequence or a fragment thereof. In some embodiments, UGI or a portion thereof has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9% or 100% identity to the wild-type UGI or UGI sequence or a portion thereof, as described below. An exemplary UGI includes the following amino acid sequence: As described below, it includes an amino acid sequence homologous to the exemplary UGI amino acid sequence or a fragment thereof. In some embodiments, UGI or a portion thereof has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9% or 100% identity to the wild-type UGI or UGI sequence or a portion thereof, as described below. An exemplary UGI includes the following amino acid sequence: As described below, it includes an amino acid sequence homologous to the exemplary UGI amino acid sequence or a fragment thereof. In some embodiments, UGI or a portion thereof has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9% or 100% identity to the wild-type UGI or UGI sequence or a portion thereof, as described below. An exemplary UGI includes the following amino acid sequence: As described below, it includes an amino acid sequence homologous to the exemplary UGI amino acid sequence or a fragment thereof. In some embodiments, UGI or a portion thereof has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9% or 100% identity to the wild-type UGI or UGI sequence or a portion thereof, as described below. An exemplary UGI includes the following amino acid sequence: As described below, it includes an amino acid sequence homologous to the exemplary UGI amino acid sequence or a fragment thereof. In some embodiments, UGI or a portion thereof has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9% or 100% identity to the wild-type UGI or UGI sequence or a portion thereof, as described below. An exemplary UGI includes the following amino acid sequence: >splP14739IUNGI_BPPB2 Uracil-DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGEN KIKML.

[0178] The term "vector" refers to a means for introducing a nucleic acid sequence into a cell and resulting in a transformed cell. Vectors include plasmids, transposons, phages, viruses, liposomes, and episomes. An "expression vector" is a nucleic acid sequence that contains a nucleotide sequence to be expressed in a recipient cell. The expression vector may include additional nucleic acid sequences, such as initiation, termination, enhancer, promoter, and secretion sequences, to facilitate and / or enable the expression of the introduced sequence. The term "vector" refers to a means for introducing a nucleic acid sequence into a cell and resulting in a transformed cell. Vectors include plasmids, transposons, phages, viruses, liposomes, and episomes. An "expression vector" is a nucleic acid sequence that contains a nucleotide sequence to be expressed in a recipient cell. The expression vector may include additional nucleic acid sequences, such as initiation, termination, enhancer, promoter, and secretion sequences, to facilitate and / or enable the expression of the introduced sequence. The term "vector" refers to a means for introducing a nucleic acid sequence into a cell and resulting in a transformed cell. Vectors include plasmids, transposons, phages, viruses, liposomes, and episomes. An "expression vector" is a nucleic acid sequence that contains a nucleotide sequence to be expressed in a recipient cell. The expression vector may include additional nucleic acid sequences, such as initiation, termination, enhancer, promoter, and secretion sequences, to facilitate and / or enable the expression of the introduced sequence. The term "vector" refers to a means for introducing a nucleic acid sequence into a cell and resulting in a transformed cell. Vectors include plasmids, transposons, phages, viruses, liposomes, and episomes. An "expression vector" is a nucleic acid sequence that contains a nucleotide sequence to be expressed in a recipient cell. The expression vector may include additional nucleic acid sequences, such as initiation, termination, enhancer, promoter, and secretion sequences, to facilitate and / or enable the expression of the introduced sequence. The term "vector" refers to a means for introducing a nucleic acid sequence into a cell and resulting in a transformed cell. Vectors include plasmids, transposons, phages, viruses, liposomes, and episomes. An "expression vector" is a nucleic acid sequence that contains a nucleotide sequence to be expressed in a recipient cell. The expression vector may include additional nucleic acid sequences, such as initiation, termination, enhancer, promoter, and secretion sequences, to facilitate and / or enable the expression of the introduced sequence. The term "vector" refers to a means for introducing a nucleic acid sequence into a cell and resulting in a transformed cell. Vectors include plasmids, transposons, phages, viruses, liposomes, and episomes. An "expression vector" is a nucleic acid sequence that contains a nucleotide sequence to be expressed in a recipient cell. The expression vector may include additional nucleic acid sequences, such as initiation, termination, enhancer, promoter, and secretion sequences, to facilitate and / or enable the expression of the introduced sequence.

[0179] Any composition or method provided herein can be combined with any one or more of the other compositions and methods provided herein. Any composition or method provided herein can be combined with any one or more of the other compositions and methods provided herein.

[0180] DNA editing has emerged as a viable means to modify disease states by correcting pathogenic mutations at the gene level. Until recently, all DNA editing platforms induced double-stranded breaks (DSBs) at specified genomic sites and relied on endogenous DNA repair pathways to determine the outcome of the product in a semi-random fashion, resulting in a complex population of gene products. The homologous recombination repair (HDR) pathway can achieve accurate, user-defined repair outcomes, but several challenges have hindered efficient repair using HDR in cell types relevant to therapy. In fact, this pathway is inefficient compared to the competing error-prone non-homologous end joining pathway. Furthermore, HDR is severely restricted in the G1 and S phases of the cell cycle, preventing accurate repair of DSBs in post-mitotic cells. As a result, it has proven difficult or impossible to modify genomic sequences efficiently in a user-defined programmable manner in these populations. Until recently, all DNA editing platforms induced double-stranded breaks (DSBs) at specified genomic sites and relied on endogenous DNA repair pathways to determine the outcome of the product in a semi-random fashion, resulting in a complex population of gene products. Until recently, all DNA editing platforms induced double-stranded breaks (DSBs) at specified genomic sites and relied on endogenous DNA repair pathways to determine the outcome of the product in a semi-random fashion, resulting in a complex population of gene products. Until recently, all DNA editing platforms induced double-stranded breaks (DSBs) at specified genomic sites and relied on endogenous DNA repair pathways to determine the outcome of the product in a semi-random fashion, resulting in a complex population of gene products. The homologous recombination repair (HDR) pathway can achieve accurate, user-defined repair outcomes, but several challenges have hindered efficient repair using HDR in cell types relevant to therapy. The homologous recombination repair (HDR) pathway can achieve accurate, user-defined repair outcomes, but several challenges have hindered efficient repair using HDR in cell types relevant to therapy. In fact, this pathway is inefficient compared to the competing error-prone non-homologous end joining pathway. In fact, this pathway is inefficient compared to the competing error-prone non-homologous end joining pathway. Furthermore, HDR is severely restricted in the G1 and S phases of the cell cycle, preventing accurate repair of DSBs in post-mitotic cells. As a result, it has proven difficult or impossible to modify genomic sequences efficiently in a user-defined programmable manner in these populations. As a result, it has proven difficult or impossible to modify genomic sequences efficiently in a user-defined programmable manner in these populations.

[0181] The features of the present disclosure are described in detail in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description of exemplary embodiments in which the principles of the present disclosure are utilized, and to the accompanying drawings, the contents of which are as follows: A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description of exemplary embodiments in which the principles of the present disclosure are utilized, and to the accompanying drawings, the contents of which are as follows: A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description of exemplary embodiments in which the principles of the present disclosure are utilized, and to the accompanying drawings, the contents of which are as follows:

Brief Description of the Drawings

[0182]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

[0183] The present invention provides compositions comprising novel adenosine base editors (e.g., ABE8) with increased efficiency. , and adenosine deaminase to modify mutations associated with glycogen storage disease type 1a (GSD1a) The present invention provides methods of using base editors that include enzyme variants.

[0184] The present invention relates, at least in p...

Claims

1. 1. An in vitro or ex vivo method for editing a glucose-6-phosphatase (G6PC) polynucleotide comprising a single nucleotide polymorphism (SNP) associated with glycogen storage disease type 1a (GSD1a), comprising contacting the G6PC polynucleotide with a base editor complexed with one or more guide polynucleotides, wherein the base editor comprises a polynucleotide-programmable DNA-binding domain and an adenosine deaminase domain, wherein the adenosine deaminase domain has the following amino acid sequence: TadA*7.10 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD or a fragment thereof lacking only the N-terminal methionine, wherein the adenosine deaminase domain comprises at least one alteration selected from the group consisting of I76Y, V82G, and F149Y, as numbered according to the TadA*7.10 amino acid sequence; one or more of the guide polynucleotides target the base editor to result in an A.T to G.C modification of the SNP associated with GSD1a, wherein the SNP associated with GSD1a results in expression of a G6PC polypeptide that terminates prematurely at amino acid position 347 or contains a cysteine ​​at position 83. method.

2. the polynucleotide-programmable DNA-binding domain is Streptococcus pyogenes Cas9 (SpCas9); the polynucleotide-programmable DNA-binding domain comprises a modified SpCas9 with altered protospacer adjacent motif (PAM) specificity or with specificity for a non-G PAM; the polynucleotide-programmable DNA-binding domain is Staphylococcus aureus Cas9 (SaCas9); the polynucleotide-programmable DNA-binding domain is nuclease-inactive or a nickase; or the polynucleotide-programmable DNA-binding domain is a nickase containing the amino acid substitution D10A or a corresponding amino acid substitution; The method of claim 1.

3. the adenosine deaminase domain is numbered according to the TadA*7.10 amino acid sequence as follows: I76Y, V82G, F149Y, and Q154S; Y147T and Q154R; Y147T and Q154S; Y147R and Q154S; V82S and Q154S; V82S and Y147R; V82S and Q154R; V82S and Y123H; I76Y and V82S; V82S, Y123H, and Y147T; V82S, Y123H, and Y147R; V82S, Y123H, and Q154R; Y147R, Q154R, and Y123H; Y147R, Q154R, and I76Y; Y147R, Q154R, and T166R; Y123H, Y147R, Q154R, and I76Y; V82S, Y123H, Y147R, and Q154R; and I76Y, V82S, Y123H, Y147R, and Q154R 3. The method of claim 1 or 2, comprising or further comprising a combination of modifications selected from the group consisting of:

4. 4. The method of any one of claims 1 to 3, wherein the base editor further comprises a wild-type adenosine deaminase domain.

5. A polynucleotide-programmable DNA-binding domain and the following TadA*7.10 amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD or a fragment thereof lacking only the N-terminal methionine, wherein the adenosine deaminase domain comprises at least one alteration selected from the group consisting of I76Y, V82G, and F149Y, as numbered according to the TadA*7.10 amino acid sequence; or a polynucleotide encoding said base editor; one or more guide polynucleotides that target the base editor to result in an A.T to G.C modification of a GSD1a-associated SNP, wherein the GSD1a-associated SNP results in expression of a G6PC polypeptide that terminates prematurely at amino acid position 347 or contains a cysteine ​​at position 83; and In vitro or ex vivo cells comprising:

6. The cell of claim 5 , which is a mammalian hepatocyte or hepatocyte precursor.

7. The cell of claim 6 , wherein the hepatocyte is an iPSC-derived hepatocyte.

8. the guide polynucleotide a) GACCUAGGCGAGGCAGUAGG; b) CCAGUAUGGACACUGUCCAAA; c) CAGUAUGGACACUGUCCAAA; and d) AGUAUGGACACUGUCCAAAG 7. The cell of claim 5 or 6, comprising a nucleic acid sequence selected from the group consisting of:

9. 1. A polynucleotide-programmable base editor comprising a DNA-binding domain and an adenosine deaminase domain, wherein the adenosine deaminase domain has the following TadA*7.10 amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD or a fragment thereof lacking only the N-terminal methionine, wherein the adenosine deaminase domain comprises at least one alteration selected from the group consisting of I76Y, V82G, and F149Y, as numbered according to the TadA*7.10 amino acid sequence; and one or more guide polynucleotides that target the base editor to result in an A.T to G.C modification of a GSD1a-associated SNP, wherein the GSD1a-associated SNP results in expression of a glucose-6-phosphatase (G6PC) polypeptide that terminates prematurely at amino acid position 347 or contains a cysteine ​​at position 83; and A pharmaceutical composition for treating GSD1a in a subject, comprising:

10. (a) inducing pluripotent stem cells or hepatocyte precursors containing a SNP associated with glycogen storage disease type 1a (GSD1a); 1. A polynucleotide-programmable base editor comprising a nucleotide-binding domain and an adenosine deaminase domain, wherein said adenosine deaminase domain has the following TadA*7.10 amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD or a fragment thereof lacking only the N-terminal methionine, wherein the adenosine deaminase domain comprises at least one alteration selected from the group consisting of I76Y, V82G, and F149Y, as numbered according to the TadA*7.10 amino acid sequence; and one or more guide polynucleotides that target the base editor to result in an A.T to G.C modification of the SNP associated with GSD1a, wherein the SNP associated with GSD1a results in expression of a glucose-6-phosphatase (G6PC) polypeptide that terminates prematurely at amino acid position 347 or contains a cysteine ​​at position 83; and and (b) differentiating the induced pluripotent stem cells into hepatocytes or their precursors.

1. An in vitro or ex vivo method for generating hepatocytes, comprising:

11. The method of claim 10, wherein the hepatocyte precursors are obtained from a subject with GSD1a.

12. The adenosine deaminase domain is numbered according to the TadA*7.10 amino acid sequence. I76Y, V82G, F149Y, and Q154S; Y147T and Q154R; Y147T and Q154S; Y147R and Q154S; V82S and Q154S; V82S and Y147R; V82S and Q154R; V82S and Y123H; I76Y and V82S; V82S, Y123H, and Y147T; V82S, Y123H, and Y147R; V82S, Y123H, and Q154R; Y147R, Q154R, and Y123H; Y147R, Q154R, and I76Y; Y147R, Q154R, and T166R; Y123H, Y147R, Q154R, and I76Y; V82S, Y123H, Y147R, and Q154R; and I76Y, V82S, Y123H, Y147R, and Q154R 12. The method of claim 10 or 11, comprising or further comprising a combination of modifications selected from the group consisting of:

13. 13. The method of any one of claims 10-12, wherein the base editor further comprises a wild-type adenosine deaminase domain.

14. 1. An in vitro or ex vivo method of editing a glucose-6-phosphatase (G6PC) polynucleotide comprising a single nucleotide polymorphism (SNP) associated with glycogen storage disease type 1a (GSD1a), comprising contacting the G6PC polynucleotide with a base editor complexed with one or more guide polynucleotides; wherein the base editor comprises an adenosine deaminase domain inserted within a Cas9 or Cas12 polypeptide, the adenosine deaminase domain having the following TadA*7.10 amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD or an amino acid sequence having at least 90% sequence identity to a fragment thereof lacking only the N-terminal methionine, wherein the adenosine deaminase domain comprises at least one alteration selected from the group consisting of I76Y, V82G, and F149Y, numbered according to the TadA*7.10 amino acid sequence; one or more of the guide polynucleotides target the base editor to result in an A.T to G.C modification of the SNP associated with GSD1a, wherein the SNP associated with GSD1a results in expression of a G6PC polypeptide that terminates prematurely at amino acid position 347 or contains a cysteine ​​at position 83. method.

15. a) a polynucleotide-programmable base editor comprising a DNA-binding domain and an adenosine deaminase domain, wherein the adenosine deaminase domain has the following TadA*7.10 amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD or a fragment thereof lacking only the N-terminal methionine, wherein the adenosine deaminase domain comprises at least one modification selected from the group consisting of I76Y, V82G, and F149Y, numbered according to the TadA*7.10 amino acid sequence; and b) one or more guide polynucleotides that target the base editor to result in an A-T to G-C modification of a GSD1a-associated SNP, wherein the GSD1a-associated SNP results in expression of a G6PC polypeptide that terminates prematurely at amino acid position 347 or contains a cysteine ​​at position 83; and 10. A pharmaceutical composition for treating glycogen storage disease type 1a (GSD1a), comprising an effective amount of

16. A pharmaceutical composition for the treatment of glycogen storage disease type 1a (GSD1a), comprising an effective amount of the cells according to any one of claims 5 to 8.

17. A kit for treating glycogen storage disease type 1a (GSD1a), comprising: a) a polynucleotide-programmable base editor comprising a DNA-binding domain and an adenosine deaminase domain, wherein the adenosine deaminase domain has the following TadA*7.10 amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD or a fragment thereof lacking only the N-terminal methionine, wherein the adenosine deaminase domain comprises at least one modification selected from the group consisting of I76Y, V82G, and F149Y, numbered according to the TadA*7.10 amino acid sequence; and b) one or more guide polynucleotides capable of targeting the base editor to result in an A-T to G-C modification of a GSD1a-associated SNP, wherein the GSD1a-associated SNP results in expression of a G6PC polypeptide that terminates prematurely at amino acid position 347 or contains a cysteine ​​at position 83; and Kit including:

18. A kit for treating glycogen storage disease type 1a (GSD1a), comprising the cells of any one of claims 5 to 8.

19. 1. A base editor system comprising a base editor complexed with one or more guide polynucleotides, wherein the base editor comprises a polynucleotide-programmable DNA-binding domain and the following TadA*7.10 amino acid sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD or a fragment thereof lacking only the N-terminal methionine, wherein said adenosine deaminase domain comprises at least one alteration selected from the group consisting of I76Y, V82G, and F149Y, numbered according to the TadA*7.10 amino acid sequence; and one or more of said guide polynucleotides targets said base editor to result in an A.T to G.C alteration of a GSD1a-associated SNP, wherein said GSD1a-associated SNP results in expression of a G6PC polypeptide that terminates prematurely at amino acid position 347 or contains a cysteine ​​at position 83. Base editor system.

20. the polynucleotide-programmable DNA-binding domain is a modified Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), or a modified Streptococcus pyogenes Cas9 (SpCas9); and / or the polynucleotide-programmable DNA-binding domain is SpCas9 with engineered protospacer adjacent motif (PAM) specificity or with specificity for a non-G PAM; and / or the polynucleotide-programmable DNA-binding domain is a nuclease-inactive Cas9 or a Cas9 nickase; 20. The base editor system of claim 19.

21. 21. An in vitro or ex vivo cell comprising the base editor system of claim 19 or 20.

22. the cell is a mammalian cell, and / or The cell is ex vivo or in vitro. The cell of claim 21.