Methods of editing a disease-associated gene using adenosine deaminase base editors, including for the treatment of genetic disease

JP2025032080A5Pending Publication Date: 2025-06-16BEAM THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024193416
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-01-27
Filing Date
2024-11-05
Publication Date
2025-06-16
Patent Text Reader

Abstract

To provide methods for treating neurological disorders.SOLUTION: The invention provides methods of treating a disease or disorder, (e.g., Parkinson's disease, Hurler syndrome, Rett syndrome, or Stargardt disease) in a subject by administering to the subject a programmable adenosine base editor system (e.g., ABE8) that have increased efficiency.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of U.S. Provisional Application No. 62 / 805,271, filed February 13, 2019, and U.S. Provisional Application No. 62 / 805,271, filed May 23, 2019. U.S. Provisional Application No. 62 / 852,228 filed on May 23, 2019; U.S. Provisional Application No. 62 / 852,228 filed on May 23, 2019; / 852,224, filed July 11, 2019; U.S. Provisional Application No. 62 / 873,138, filed August 19, 2019; U.S. Provisional Application No. 62 / 888,867 filed on November 6, 2019 / 931,722, filed November 27, 2019; U.S. Provisional Application No. 62 / 941,569, filed January 27, 2020; This application claims the benefit of U.S. Provisional Application No. 62 / 966,526, filed on 2006 / 01 / 14, the disclosure of which is incorporated herein by reference. is incorporated herein in its entirety.

[0002] Incorporation by Reference All publications, patents, and patent applications mentioned herein are hereby incorporated by reference in their entirety. The patent or patent application is specifically and individually indicated to be incorporated by reference; To the same extent, the invention is incorporated herein by reference. All cited publications, patents, and patent applications are incorporated herein by reference in their entirety. can be. [Background technology]

[0003] Targeted editing of nucleic acid sequences, such as targeted cleavage or targeted editing of genomic DNA. These modifications are a very promising approach for studying gene function and may be useful in the treatment of human genetic diseases. Currently available base editors include target C·G Cytidine base editors (e.g., BE4) convert base pairs to T·A and A·T to G·C Includes adenine base editors (e.g., ABE7.10) that target targets with higher specificity and efficiency. There is a need in the art for improved base editors that can introduce modifications into sequences. It is being done. Summary of the Invention

[0004] The present invention provides compositions comprising novel adenine base editors (e.g., ABE8) with increased efficiency. and a base enzyme containing adenosine deaminase variants for editing target sequences. A method for using the editor is provided.

[0005] In some embodiments, provided herein is a method of treating a neurological disorder in a subject, comprising: (i) an adenosine base editor or a nucleic acid sequence encoding the same; and (ii) a guide polynucleotide. and administering to a subject a nucleic acid sequence encoding the same, The syn-base editor contains a programmable DNA-binding domain and an adenosine deaminase domain. and a domain comprising the adenosine deaminase domain, the adenosine deaminase domain being represented by the numbering in SEQ ID NO: 2. and an amino acid substitution at amino acid position 82 or 166 or a position corresponding thereto, The guide polynucleotide induces the adenosine base editor to induce the neurotransmitter Inducing an A to G nucleobase alteration in a target gene or its regulatory element associated with the disorder and thereby treating said neurological disorder in a subject. In another embodiment, the target gene is the alpha-L-iduronidase (IDUA) gene. and the neurological disease is Hurler syndrome. is the leucine-rich repeat kinase 2 (LRRK2) gene, and the neurological disease is Parkinson's disease. In one embodiment of this aspect, the target gene is methyl-CpG binding protein 2 (MECP 2) is a gene and the neurological disease is Rett syndrome. The target gene is the ATP-binding cassette subfamily member 4 (ABCA4) gene, and The disease is Stargardt disease.

[0006] In some embodiments, provided herein are methods of treating Hurler syndrome in a subject. (i) an adenosine base editor or a nucleic acid sequence encoding the same; and (ii) a guide and administering to a subject a polynucleotide or a nucleic acid sequence encoding the same, Adenosine base editors combine programmable DNA-binding domains with adenosine deamina and an adenosine deaminase domain, the adenosine deaminase domain being the sequence shown in SEQ ID NO: 2. containing an amino acid substitution at amino acid position 82 or 166 or a position corresponding thereto in the code. and the guide polynucleotide directs the adenosine base editor to encode the adenosine base of interest. A to G in the IDUA gene or its regulatory elements Methods are provided for producing nucleobase modifications, thereby treating Hurler syndrome in a subject. can be.

[0007] In one embodiment, the administration ameliorates at least one symptom associated with Hurler syndrome. In one embodiment, the administration improves the activity of the amino acid in adenosine deaminase. Compared with treatment with base editors without substitutions, resulting in more rapid improvement of at least one symptom.

[0008] In one embodiment, the IDUA gene or a regulatory element thereof is associated with Hurler syndrome. In one embodiment, the A to G nucleobase modification is a SNP associated with Hurler syndrome. In one embodiment, the SNP associated with Hurler syndrome is , W402X or W401X amino acids in the IDUA polypeptide according to the numbering in SEQ ID NO:4 resulting in a mutant or variant thereof encoded by the IDUA gene, where X is the terminal In one embodiment, the A to G nucleobase modification is a stop codon associated with Hurler syndrome. The associated SNP is changed to a wild-type nucleobase. In one embodiment, the A to G nucleobase The base modification is a SNP associated with Hurler syndrome that results in the amelioration of one or more symptoms of Hurler syndrome. In one embodiment, the SN associated with Hurler syndrome is altered to a non-wild-type nucleobase that results in a The A to G modification in P converts the stop codon encoded by the IDUA gene into the IDUA polypeptide. This converts it to tryptophan in the peptide.

[0009] In one embodiment, the guide polynucleotide comprises a SNP associated with Hurler syndrome. In one embodiment, the nucleic acid sequence is complementary to the IDUA gene or a regulatory element thereof. wherein the adenosine base editor is a gene encoding an IDUA gene containing a SNP associated with Hurler syndrome. or complexes with a single guide RNA (sgRNA) containing a nucleic acid sequence complementary to its regulatory element. In one embodiment, the sgRNA is 5′-GACUCUAGGCAGAGGUCUCAA-3′, 5′-ACUCUAGGCAG AGGUCUCAA-3′, 5′- CUCUAGGCCGAAGUGUCGC -3′, and 5′-GCUCUAGGCCGAAGUGUCGC-3 ',

[0010] In some embodiments, provided herein are methods of treating Parkinson's disease in a subject. (i) an adenosine base editor or a nucleic acid sequence encoding the same; and (ii) a guide and administering to a subject a polynucleotide or a nucleic acid sequence encoding the same, Adenosine base editors combine programmable DNA-binding domains with adenosine deamina and an adenosine deaminase domain, the adenosine deaminase domain being the sequence shown in SEQ ID NO: 2. containing an amino acid substitution at amino acid position 82 or 166 or a position corresponding thereto in the code. and the guide polynucleotide directs the adenosine base editor to encode a locus of interest. A or B in the isin-rich repeat kinase 2 (LRRK2) gene or its regulatory elements to G, thereby treating Parkinson's disease in a subject. is provided.

[0011] In one embodiment, the administration improves at least one symptom associated with Parkinson's disease. In one embodiment, the administration reduces the amino acid substitution in adenosine deaminase. Compared with treatment with base editors that do not involve substitutions, Both result in more rapid improvement of a single symptom.

[0012] In one embodiment, the LRRK2 gene or a regulatory element thereof is associated with Parkinson's disease. In one embodiment, the A to G nucleobase modification comprises a SNP associated with Parkinson's disease. In one embodiment, the SNP associated with Parkinson's disease is in a SNP associated with A419V, R1441C, R1441H, or have the G2019S amino acid variant or its variant encoded by the LRRK2 gene. Glass.

[0013] In one embodiment, the A to G nucleobase modification is a SNP associated with Parkinson's disease. to a wild-type nucleobase. In one embodiment, the A to G nucleobase modification is SNPs associated with Parkinson's disease were compared with non-field-specific SNPs that result in improvement of one or more symptoms of Parkinson's disease. In one embodiment, the A to G nucleobase modification is The cysteine ​​or histidine in the LRRK2 polypeptide encoded by the 2 gene is replaced with alpha-cysteine. In one embodiment, the A to G modification is mediated by the LRRK2 gene. In one embodiment, the serine is changed to a glycine in the encoded LRRK2 polypeptide. In the case of SEQ ID NO: 3, the A to G modification is the LRRK2 polypeptide or LRRK2 polypeptide according to the numbering in SEQ ID NO: 3. The variant encoded by the RK2 gene contains a cysteine ​​(C) or histidine at position 144. Substitution of arginine (R) for arginine (H) or glycine (G) for serine at position 2019 Replace with.

[0014] In some embodiments, provided herein are methods of treating Parkinson's disease in a subject. (i) an adenosine base editor or a nucleic acid sequence encoding the same; and (ii) a guide and administering to a subject a polynucleotide or a nucleic acid sequence encoding the same, Adenosine base editors combine programmable DNA-binding domains with adenosine deamina and a guide polynucleotide comprising an adenosine base editor domain. A to G nucleotide substitution in the LRRK2 gene SNP associated with Parkinson's disease The SNP results in a base alteration, and the SNP is located at the position in the LRRK2 polypeptide as numbered in SEQ ID NO: 3. The method does not encode the G2019S mutant or a variant thereof.

[0015] In one embodiment, the adenosine deaminase domain is and containing an amino acid substitution at amino acid position 82 or 166 or a position corresponding thereto. In one embodiment, the guide polynucleotide is a SNP associated with Parkinson's disease. In one embodiment, the nucleic acid sequence is complementary to the LRRK2 gene or a regulatory element thereof, comprising: In this study, the adenosine base editor was used to identify the LRRK2 gene containing a SNP associated with Parkinson's disease. complex with a single guide RNA (sgRNA) containing a nucleic acid sequence complementary to the promoter or its regulatory element In one embodiment, the sgRNA comprises the nucleic acid sequence 5'-AAGCGCAAGCCUGGAGGGAA-3'; contains 5'-ACUACAGCAUUGCUCAGUAC-3'.

[0016] In one aspect, provided herein is a method of treating Rett Syndrome in a subject, comprising: (i) an adenosine base editor or a nucleic acid sequence encoding the same; and (ii) a guide protein. and administering to a subject a nucleic acid sequence encoding the same. The adenosine base editor combines a programmable DNA-binding domain and an adenosine deaminase and an adenosine deaminase domain, the adenosine deaminase domain being the number 1 in SEQ ID NO: 2. and containing an amino acid substitution at amino acid position 82 or 166 or a position corresponding thereto. The guide polynucleotide directs the adenosine base editor to produce a methylation product of interest. Nucleotide A to G transitions in the methyl CpG-binding protein 2 (MECP2) gene or its regulatory elements Methods are provided for effecting acid-base alterations, thereby treating Rett syndrome in a subject. .

[0017] In one embodiment, the administration ameliorates at least one symptom associated with Rett Syndrome. In one embodiment, the administration reduces the amino acid substitution in adenosine deaminase. Compared with treatment with a base editor that does not involve substitution, In one embodiment, the MECP2 gene or The regulatory element comprises a SNP associated with Rett syndrome. The nucleobase alteration from A to G is in a SNP associated with Rett Syndrome. In the present invention, the SNP associated with Rett syndrome is located in the MECP2 polypeptide as numbered in SEQ ID NO: 5. R106W or T158M amino acid variant in the thyroid gland or the thyroid gland encoded by the MECP2 gene In one embodiment, the Rett syndrome-associated SNP is R255X or R270X amino acid in the MECP2 polypeptide encoded by the ECP2 gene This results in a mutant where X is a stop codon.

[0018] In one embodiment, the A to G nucleobase modification is a SNP associated with Rett Syndrome. In one embodiment, the A to G nucleobase modification is The syndrome-associated SNPs were altered to non-wild-type nucleobases, resulting in improvement of Rett syndrome symptoms. In one embodiment, the A to G nucleobase in the SNP associated with Rett Syndrome The modification changes the stop codon to a tryptophan in the MECP2 polypeptide.

[0019] In one embodiment, the guide polynucleotide comprises a SNP associated with Rett Syndrome. The nucleic acid sequence may be complementary to the MECP2 gene or a regulatory element thereof. In one embodiment, the adenosine base editor is a gene encoding an MECP2 gene containing a SNP associated with Rett syndrome. complex with a single guide RNA (sgRNA) containing a nucleic acid sequence complementary to the promoter or its regulatory element In one embodiment, the guide polynucleotide comprises 5'-CUUUUCACUUCCUGCCGGG G-3′, 5′-AGCUUCCAUGUCCAGCCUUC-3′, 5′- ACCAUGAAGUCAAAAUCAUU-3′, and 5′- G CUUUCAGCCCCGUUUCUUG-3′.

[0020] In some embodiments, the present disclosure provides a method for treating Stargardt disease in a subject. (i) an adenosine base editor or a nucleic acid sequence encoding the same; and (ii) a guide and administering to a subject a polynucleotide or a nucleic acid sequence encoding the same, Adenosine base editors combine programmable DNA-binding domains with adenosine deamina and an adenosine deaminase domain, the adenosine deaminase domain being the sequence shown in SEQ ID NO: 2. containing an amino acid substitution at amino acid position 82 or 166 or a position corresponding thereto in the code. and the guide polynucleotide directs the adenosine base editor to encode an AT polypeptide of interest. In the P-binding cassette subfamily member 4 (ABCA4) gene or its regulatory elements and resulting in an A to G nucleobase alteration, thereby treating Stargardt disease in a subject. This provides a method for

[0021] In one embodiment, the administering is to treat at least one symptom associated with Stargardt disease. In one embodiment, the administration improves the amino acid sequence of adenosine deaminase. Compared with treatment with a base editor without acid substitution, resulting in more rapid improvement of at least one symptom.

[0022] In one embodiment, the ABCA4 gene contains a SNP associated with Stargardt disease. In some embodiments, the A to G nucleobase alteration is the same as that in a SNP associated with Stargardt disease. In one embodiment, the SNP associated with Stargardt disease is the SNP in SEQ ID NO: 6. A1038V or G1961E amino acid variant in the ABCA4 polypeptide or ABC A4 gene. The SNP associated with Gardt's disease is identified as the SNP in the ABCA4 polypeptide as numbered in SEQ ID NO: 6. This results in a G1961E amino acid mutation or variant thereof.

[0023] In one embodiment, the A to G nucleobase alteration is a SNP associated with Stargardt disease. to a wild-type nucleobase. In one embodiment, the A to G nucleobase modification SNPs associated with Stargardt disease are identified as non-coding genes that result in the improvement of one or more symptoms of Stargardt disease. In one embodiment, the guide polynucleotide is Nucleic acids complementary to the ABCA4 gene or its regulatory elements containing SNPs associated with Talgart disease Contains arrays.

[0024] In one embodiment, the adenosine base editor is an SN associated with Stargardt disease. A single guide RNA comprising a nucleic acid sequence complementary to the ABCA4 gene or its regulatory element, including P In one embodiment, the sgRNA is complexed with the sequence 5′-CUCCAGGGCGAACUUCGA Contains CACACAGC-3′.

[0025] In various embodiments, the treatments described herein are directed to an adenovirus that does not involve said amino acid substitution. compared with treatment with a base editor containing the nosine deaminase domain results in symptom improvement.

[0026] In one embodiment, the present invention relates to a target gene or its regulatory element associated with a neurological disorder. a method for editing a target gene or a regulatory element thereof, the method comprising: (i) (ii) contacting a denosine base editor with a guide polynucleotide; The adenosine base editor comprises a programmable DNA binding domain and an adenosine deoxyribonucleic acid (ADC) domain. and an adenosine deaminase domain, the adenosine deaminase domain being the sequence shown in SEQ ID NO: 2. Amino acid substitution at amino acid position 82 or 166 or their corresponding positions in the numbering system wherein the guide polynucleotide guides the adenosine base editor to A to G nucleobase modification in target genes or their regulatory elements associated with neurological disorders In one embodiment of this aspect, the target gene is leucine. The gene is nucleotide-rich repeat kinase 2 (LRRK2), and the neurological disorder is Parkinson's disease. In another embodiment of this aspect, the target gene is alpha-L-iduronidase (IDUA). In one embodiment of this aspect, the target The gene is the methyl-CpG binding protein 2 (MECP2) gene, and the neurological disease is Rett syndrome. In another embodiment of this aspect, the target gene is a gene of the ATP-binding cassette subfamily. -member 4 (ABCA4) gene, and the neurological disorder is Stargardt disease.

[0027] In one embodiment, the leucine-rich repeat kinase 2 (LRRK2) gene is used herein. or a regulatory element thereof, The compound is a compound selected from the group consisting of (i) an adenosine base editor or a nucleic acid sequence encoding the same, and (ii) a guide and contacting the antigen with a nucleic acid sequence encoding the antigen. The adenosine base editor combines a programmable DNA-binding domain and an adenosine deaminase and an adenosine deaminase domain, the adenosine deaminase domain being the number 1 in SEQ ID NO: 2. and containing an amino acid substitution at amino acid position 82 or 166 or a position corresponding thereto. wherein the guide polynucleotide induces the adenosine base editor to encode the LRRK2 gene. Methods are provided for producing an A to G nucleobase modification in a nucleic acid sequence or its regulatory element. can be.

[0028] In one embodiment, the A to G nucleobase modification is at a SNP associated with Parkinson's disease. In one embodiment, the SNP associated with Parkinson's disease is SEQ ID NO: 3 A419V, R1441C, R1441H, or G2019S in the LRRK2 polypeptide according to the numbering in This results in an amino acid variant or variant thereof encoded by the LRRK2 gene. In some embodiments, the A to G nucleobase modification is a mutation that alters a SNP associated with Parkinson's disease in the wild. In one embodiment, the A to G nucleobase modification is a non-wild-type nucleic acid that results in the improvement of one or more symptoms of Parkinson's disease; Change to a base.

[0029] In one embodiment, the A to G nucleobase modification is The cysteine ​​or histidine in the LRRK2 polypeptide to be treated is changed to arginine. In one embodiment, the A to G modification is present in the LRRK2 polypeptide encoded by the LRRK2 gene. In one embodiment, the A to G modification changes the serine to glycine in the peptide. is encoded by the LRRK2 polypeptide or LRRK2 gene as numbered in SEQ ID NO: 3 The variants replace the cysteine ​​(C) or histidine (H) at position 144 with arginine (R) or by substituting the serine at position 2019 with glycine (G).

[0030] In one embodiment, the leucine-rich repeat kinase 2 (LRRK2) gene is used herein. or a method for editing the LRRK2 gene or its regulatory element, (i) an adenosine base editor or a nucleic acid sequence encoding the same; and (ii) contacting the guide polynucleotide or a nucleic acid sequence encoding the same with the The adenosine base editor comprises a programmable DNA binding domain and an adenosine base. aminase domain, and the guide polynucleotide is -induced to result in a nucleobase alteration from A to G at a SNP in the LLRK2 gene, said SNP being LRRK2 polypeptide G2019S mutant or variant thereof according to the numbering in SEQ ID NO: 3 A method is provided that does not require coding.

[0031] In one embodiment, the adenosine deaminase domain is and containing an amino acid substitution at amino acid position 82 or 166 or a position corresponding thereto. In one embodiment, the guide polynucleotide is a SNP associated with Parkinson's disease. The present invention comprises a nucleic acid sequence complementary to the LRRK2 gene or a regulatory element thereof, comprising:

[0032] In one embodiment, the adenosine base editor is A single guide RNA ( In one embodiment, the sgRNA is complexed with the nucleic acid sequence 5′-AAGCGCAAGCCUGGAG GGAA-3′; or 5′-ACUACAGCAUUGCUCAGUAC-3′.

[0033] In some embodiments, the alpha-L-iduronidase (IDUA) gene or A method for editing the IDUA gene or its regulatory element, comprising: (i) an adenosine base editor or a nucleic acid sequence encoding the same; and (ii) a guide polypeptide. contacting said adenosine nucleotide with a nucleic acid sequence encoding the same; The base editor contains a programmable DNA binding domain and an adenosine deaminase domain. the adenosine deaminase domain comprises a nucleotide sequence as numbered in SEQ ID NO:2; containing an amino acid substitution at amino acid position 82 or 166 or a position corresponding thereto, The guide polynucleotide directs the adenosine base editor to encode the IDUA gene or or a regulatory element thereof, resulting in an A to G nucleobase modification. .

[0034] In one embodiment, the IDUA gene or a regulatory element thereof is associated with Hurler syndrome. In one embodiment, the A to G nucleobase modification is a SNP associated with Hurler syndrome. In one embodiment, the SNP associated with Hurler syndrome is NP is a W402X or W401X amino acid in the IDUA polypeptide according to the numbering in SEQ ID NO: 4 resulting in a nucleotide variant or variant thereof encoded by the IDUA gene, wherein X is a stop codon.

[0035] In one embodiment, the A to G nucleobase modification identifies a SNP associated with Hurler syndrome. In one embodiment, the A to G nucleobase modification is A non-wild type SNP associated with Hurler syndrome that results in amelioration of one or more symptoms of Hurler syndrome. In one embodiment, the A in the SNP associated with Hurler syndrome is changed to a nucleobase. The modification of α-G to G results in a termination codon encoded by the IDUA gene in the IDUA polypeptide. It converts it into tryptophan, which is found in

[0036] In one embodiment, the guide polynucleotide comprises a SNP associated with Hurler syndrome. In one embodiment, the nucleic acid sequence is complementary to the IDUA gene or a regulatory element thereof. wherein the adenosine base editor is a gene encoding an IDUA gene containing a SNP associated with Hurler syndrome. or complexes with a single guide RNA (sgRNA) containing a nucleic acid sequence complementary to its regulatory element. In one embodiment, the sgRNA has the sequence 5′-GACUCUAGGCAGAGGUCUCAA-3′, 5′-ACUCUAG GCAGAGGUCUCAA-3′, 5′- CUCUAGGCCGAAGUGUCGC -3′, and 5′-GCUCUAGGCCGAAGUGUCGC -3'.

[0037] In one embodiment, the methyl-CpG binding protein 2 (MECP2) gene or A method for editing the regulatory element, comprising: (i) an adenosine base editor or (ii) a guide polynucleotide or a nucleic acid encoding the same; administering a sequence to a subject, wherein the adenosine base editor is programmable. a DNA binding domain and an adenosine deaminase domain, The nase domain is located at amino acid position 82 or 166 or thereabouts in the numbering in SEQ ID NO:2. and the guide polynucleotide comprises an amino acid substitution at a position corresponding to the Inducing a syn-base editor to convert A to G in the MECP2 gene or its regulatory elements

[0023] Methods are provided that result in nucleobase modifications of:

[0038] In one embodiment, the MECP2 gene or a regulatory element thereof is associated with Rett syndrome. In one embodiment, the A to G nucleobase modification is associated with Rett syndrome. In one embodiment, the SNP associated with Rett syndrome is , R106W or T158M amino acid in the MECP2 polypeptide according to the numbering in SEQ ID NO:5 This results in a mutant or variant thereof encoded by the MECP2 gene. In this embodiment, the SNP associated with Rett syndrome is located in the MECP2 polypeptide encoded by the MECP2 gene. resulting in an R255X or R270X amino acid variant in the polypeptide, where X is a stop codon is.

[0039] In one embodiment, the A to G nucleobase modification is a SNP associated with Rett Syndrome. In one embodiment, the A to G nucleobase modification is The syndrome-associated SNPs are compared with non-wild-type nucleic acid sequences that result in amelioration of one or more symptoms of Rett syndrome. In one embodiment, the A to G in the SNP associated with Rett Syndrome is changed to The nucleobase modification changes the stop codon to tryptophan in the MECP2 polypeptide. do.

[0040] In one embodiment, the guide polynucleotide comprises a SNP associated with Rett Syndrome. In one embodiment, the nucleic acid sequence is complementary to the MECP2 gene or a regulatory element thereof. wherein the adenosine base editor is a gene encoding an MECP2 gene containing a SNP associated with Rett syndrome or is complexed with a single guide RNA (sgRNA) that contains a nucleic acid sequence complementary to the regulatory element. In one embodiment, the guide polynucleotide has the sequence 5'-CUUUUCACUUCCUGCCGGGG-3' , 5'-AGCUUCCAUGUCCAGCCUUC-3', 5'- ACCAUGAAGUCAAAAUCAUU-3', and 5'- GCUUU CAGCCCCGUUUCUUG-3'

[0041] In some embodiments, the present disclosure provides a method for the detection of ATP-binding cassette subfamily member 4 (ABCA4) A method for editing a gene or its regulatory element, comprising: The regulatory element may be: (i) an adenosine base editor or a nucleic acid sequence encoding same; (ii) contacting the guide polynucleotide or a nucleic acid sequence encoding the same. and wherein the adenosine base editor comprises a programmable DNA-binding domain and an adenosine and an adenosine deaminase domain, the adenosine deaminase domain being represented by SEQ ID NO: 2. Amino acid position 82 or 166, or a position corresponding thereto, in the numbering system in adenosine base substitution, and the guide polynucleotide induces the adenosine base editor. resulting in an A to G nucleobase alteration in the ABCA4 gene or its regulatory elements , a method is provided.

[0042] In one embodiment, the administering is to treat at least one symptom associated with Stargardt disease. In one embodiment, the administration improves the amino acid sequence of adenosine deaminase. Associated with Stargardt disease compared with treatment with base editors without acid substitution resulting in more rapid improvement of at least one symptom.

[0043] In one embodiment, the ABCA4 gene contains a SNP associated with Stargardt disease. In some embodiments, the A to G nucleobase alteration is the same as that in a SNP associated with Stargardt disease. In one embodiment, the SNP associated with Stargardt disease is the SNP in SEQ ID NO: 6. A1038V or G1961E amino acid variant in the ABCA4 polypeptide or ABC A4 gene. The SNP associated with Gardt's disease is identified as the SNP in the ABCA4 polypeptide as numbered in SEQ ID NO: 6. This results in a G1961E amino acid mutation or variant thereof.

[0044] In one embodiment, the A to G nucleobase alteration is a SNP associated with Stargardt disease. to a wild-type nucleobase. In one embodiment, the A to G nucleobase modification is SNPs associated with Stargardt disease that result in amelioration of one or more symptoms of Stargardt disease In one embodiment, the guide polynucleotide is Nucleic acid complementary to the ABCA4 gene or its regulatory element containing a SNP associated with Schuttgardt disease Contains the acid sequence.

[0045] In one embodiment, the adenosine base editor is an SN associated with Stargardt disease. A single guide RNA comprising a nucleic acid sequence complementary to the ABCA4 gene or its regulatory element, including P In one embodiment, the sgRNA is complexed with the sequence 5′-CUCCAGGGCGAACUUCGACACA Contains CAGC-3′.

[0046] In various embodiments of the above aspects, the contacting occurs intracellularly. In embodiments, the contacting results in fewer than 10% indels in the genome of the cell. , where the indel rate is the difference between the sequence adjacent to a single nucleotide alteration and the unaltered sequence. In one embodiment, the contacting is performed by detecting a match frequency in the cell. resulting in less than 5% indels in the genome, where the indel rate is adjacent to a single nucleotide alteration. In one embodiment, the mismatch frequency is measured between the adjacent sequences and the unmodified sequence. wherein said contacting results in less than 1% indels in the genome of the cell, The mismatch rate is determined by the mismatch frequency between the sequence adjacent to the single nucleotide alteration and the unaltered sequence. Therefore, it is measured.

[0047] In various embodiments of the above aspects, the cell is a neuron. In one embodiment, the contacting occurs within a population of cells. and wherein after said contacting step, at least 40% of said population of cells contain an A to G nucleic acid salt. In some embodiments, the contacting step results in a pre-treatment after the contacting step. resulting in an A to G nucleobase modification in at least 50% of the population of said cells. In some embodiments, the contacting step results in at least 70% of the population of cells being positive after the contacting step. In one embodiment, the at least one In one embodiment, the population of cells is 90% viable after the contacting step. In one embodiment, the population of cells is enriched for neurons after the contacting step. In one embodiment, the contacting occurs in vivo or ex vivo.

[0048] In the various aspects and embodiments described above, the polynucleotide-programmable In one embodiment, the DNA binding domain is Cas9. In one embodiment, the Cas9 is SpCas9, SaCas9, or In one embodiment, the polynucleotide The tunable DNA-binding domain exhibits engineered protospacer adjacent motif (PAM) specificity. In one embodiment, the Cas9 comprises a modified SpCas9 having one of the following amino acids: NGG, NGA, NGCG, NGN , NNGRRT, NNNRRT, NGCG, NGCN, NGTN, and NGC. wherein N is A, G, C, or T and R is A or G. wherein the polynucleotide-programmable DNA binding domain binds to a nuclease In one embodiment, the polynucleotide is programmable. In one embodiment, the DNA-binding domain is a nickase variant. The variant contains the amino acid substitution D10A or a corresponding amino acid substitution. In aspects and embodiments provided herein, the adenosine deaminase domain is TadA. In one embodiment, the adenosine deaminase comprises a V82S modification and / or a V82S domain. or a TadA deaminase containing the T166R modification.

[0049] In various aspects and embodiments described above, the adenosine deaminase is Y147T, Y147R, It further comprises one or more of the following modifications: Q154S, Y123H, Q154R, or a combination thereof. In aspects and embodiments provided herein, the adenosine deaminase is Y1 47R + Q154R +Y123H;Y147R + Q154R + I76Y;Y147R + Q154R + T166R;Y147T + Q154R; Y147T + Q154S; and Y123H + Y147R + Q154R + I76Y In aspects and embodiments provided herein, the adenosine The base editor domain comprises an adenosine deaminase monomer. In aspects and embodiments, the adenosine base editor is an adenosine deaminase In one embodiment, the TadA deaminase is a TadA*8 variant. In some embodiments, the TadA*8 variant is TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.31 dA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12 and TadA*8.13. In one embodiment, the adenosine base Editors include ABE8.1, ABE8.2, ABE8.3, ABE8.4, ABE8.5, ABE8.6, ABE8.7, ABE8.8, AB ABE8 base enzyme selected from the group consisting of E8.9, ABE8.10, ABE8.11, ABE8.12, and ABE8.13 It's a dieter.

[0050] In certain aspects, the present specification provides a method for manufacturing a cellular phone according to various aspects and embodiments disclosed herein. Cells produced by the methods described are provided. , produced by the methods described in various aspects and embodiments disclosed herein. A population of cells is provided.

[0051] In some embodiments, provided herein are: (i) an adenosine base editor or a compound encoding the same; (ii) a guide polynucleotide or a nucleic acid sequence encoding the same; a base editor system, the adenosine base editor comprising: adenosine deaminase domain and adenosine binding domain, The deaminase domain is located at amino acid position 82 or 166 or contains amino acid substitutions at the corresponding positions, and the guide polynucleotide Induction of adenosine base editors to target genes or their regulatory elements associated with neurological disorders A base editor system is provided that results in an A to G nucleobase modification in the nucleotide sequence. In one embodiment of the above aspect, the target gene is leucine-rich repeat kinase 2. (LRRK2) gene and the neurological disease is Parkinson's disease. In some cases, the target gene is the alpha-L-iduronidase (IDUA) gene, and the target gene is a gene encoding a neurological disease. In one embodiment of this aspect, the target gene is a methyl-CpG The neurological disorder is Rett syndrome. In embodiments, the target gene is the ATP-binding cassette subfamily member 4 (ABCA4) gene. It is a genetic disorder and the neurological disease is Stargardt disease.

[0052] In some embodiments, provided herein are: (i) an adenosine base editor or a compound encoding the same; and (ii) a guide polynucleotide or a nucleic acid sequence encoding the same. a base editor system, wherein the adenosine base editor is programmable. adenosine deaminase domain capable of binding to DNA, The aminase domain is located at amino acid position 82 or 166 in the numbering of SEQ ID NO:2, or and the guide polynucleotide comprises an amino acid substitution at the corresponding position. Induction of a denosine base editor to A in the LRRK2 gene or its regulatory elements A base editor system is provided that results in nucleobase modifications to G.

[0053] In one embodiment, the A to G nucleobase modification is In one embodiment, the SNP is a SNP associated with Parkinson's disease in the The SNPs associated with Parkinson's disease are identified in the LRRK2 polypeptide as numbered in SEQ ID NO:3. A419V, R1441C, R1441H, or G2019S amino acid variants in the LRRK2 gene resulting in that variant being encoded.

[0054] In one embodiment, the A to G nucleobase modification is In one embodiment, the A to G nucleobase modification is The SNP associated with Parkinson's disease is a non-wild type that results in an improvement in the symptoms of Parkinson's disease. In one embodiment, the A to G nucleobase modification is The cysteine ​​or histidine in the LRRK2 polypeptide encoded by the gene is replaced with arginine. In one embodiment, the A to G modification is a modification of a amino acid sequence encoded by the LRRK2 gene. In one embodiment, the serine is changed to a glycine in the LRRK2 polypeptide to be loaded. wherein the A to G modification is a LRRK2 polypeptide or an LRRK polypeptide according to the numbering in SEQ ID NO: 3. The variants encoded by the 2 genes contain a cysteine ​​(C) or histidine at position 144. Substitution of arginine (R) for arginine (H) or glycine (G) for serine at position 2019 In one embodiment, the adenosine deaminase domain is substituted with SEQ ID NO:2. Amino acid at amino acid position 82 or 166 or a position corresponding thereto in the numbering system Contains acid substitution.

[0055] In one embodiment, the guide polynucleotide comprises a SNP associated with Parkinson's disease. In one embodiment, the nucleic acid sequence is complementary to the LRRK2 gene or a regulatory element thereof. The present invention relates to a method for the preparation of a gene encoding an adenosine base editor, the method comprising: Complexed with a single guide RNA (sgRNA) containing a nucleic acid sequence complementary to a gene or its regulatory element In one embodiment, the sgRNA comprises the nucleic acid sequence 5′-AAGCGCAAGCCUGGAGGGAA-3′; or 5′-ACUACAGCAUUGCUCAGUAC-3′.

[0056] In some embodiments, provided herein are: (i) an adenosine base editor or a compound encoding the same; (ii) a guide polynucleotide or a nucleic acid sequence encoding the same; a base editor system, the adenosine base editor comprising: adenosine deaminase domain and adenosine binding domain, The deaminase domain is located at amino acid position 82 or 166 or contains an amino acid substitution at the corresponding position, and the guide polynucleotide Induction of adenosine base editors to encode the alpha-L-iduronidase (IDUA) gene or its A base editor system that produces A to G nucleobase modifications in regulatory elements of is provided.

[0057] In one embodiment, the IDUA gene or a regulatory element thereof is associated with Hurler syndrome. In one embodiment, the A to G nucleobase modification is associated with Hurler syndrome. In one embodiment, the Hurler syndrome-associated SNP is the W402X or W401X amino acid sequence in the IDUA polypeptide according to the numbering in SEQ ID NO: 4 resulting in an acid variant or variant thereof encoded by the IDUA gene, where X is It is a stop codon.

[0058] In one embodiment, the A to G nucleobase modification identifies a SNP associated with Hurler syndrome. In one embodiment, the A to G nucleobase modification is The SNPs associated with Hurler syndrome are compared with non-wild-type SNPs that result in improvement of one or more symptoms of Hurler syndrome. In one embodiment, the A in the SNP associated with Hurler syndrome is changed to a nucleotide base. The modification of the IDUA gene to G results in a termination codon in the IDUA polypeptide encoded by the IDUA gene. converts it to tryptophan.

[0059] In one embodiment, the guide polynucleotide comprises a SNP associated with Hurler syndrome. In one embodiment, the nucleic acid sequence is complementary to the IDUA gene or a regulatory element thereof. the adenosine base editor is an IDUA gene or a gene encoding a SNP associated with Hurler syndrome. or complexes with a single guide RNA (sgRNA) containing a nucleic acid sequence complementary to its regulatory element. In one embodiment, the sgRNA has the sequence 5′-GACUCUAGGCAGAGGUCUCAA-3′, 5′-ACUCUAG GCAGAGGUCUCAA-3′, 5′- CUCUAGGCCGAAGUGUCGC -3′, and 5′-GCUCUAGGCCGAAGUGUCGC -3' of a nucleic acid sequence selected from the group consisting of:

[0060] In some embodiments, provided herein are: (i) an adenosine base editor or (ii) a guide polynucleotide or a nucleic acid sequence encoding the same; and a sequence, wherein the adenosine base editor is a a grammable DNA binding domain and an adenosine deaminase domain, The nosine deaminase domain is located at amino acid position 82 or 1 in the numbering of SEQ ID NO:2. 66 or a corresponding amino acid substitution at position 66, , by inducing the adenosine base editor to encode the methyl-CpG-binding protein 2 (MECP2) gene. or a base editor that results in an A to G nucleobase modification in the regulatory element thereof. A system is provided.

[0061] In one embodiment, the MECP2 gene or a regulatory element thereof is associated with Rett syndrome. In one embodiment, the A to G nucleobase modification is associated with Rett syndrome. In one embodiment, the SNP associated with Rett Syndrome is in R106W or T158M amino acid variants in the MECP2 polypeptide with numbering in column no. 5 or a variant thereof encoded by the MECP2 gene. , a SNP associated with Rett syndrome is found in the MECP2 polypeptide encoded by the MECP2 gene. This results in an R255X or R270X amino acid variant in the sequence SEQ ID NO: 1, where X is a stop codon.

[0062] In one embodiment, the A to G nucleobase modification is a SNP associated with Rett Syndrome. In one embodiment, the A to G nucleobase modification is The syndrome-associated SNPs are compared with non-wild-type nucleic acid sequences that result in amelioration of one or more symptoms of Rett syndrome. In one embodiment, the A to G nucleotide at the SNP associated with Rett Syndrome is changed to a The acid-base modification changes the stop codon to tryptophan in the MECP2 polypeptide.

[0063] In one embodiment, the guide polynucleotide comprises a SNP associated with Rett Syndrome. In one embodiment, the nucleic acid sequence is complementary to the MECP2 gene or a regulatory element thereof. wherein the adenosine base editor is a gene encoding an MECP2 gene containing a SNP associated with Rett syndrome or is complexed with a single guide RNA (sgRNA) that contains a nucleic acid sequence complementary to the regulatory element. In one embodiment, the guide polynucleotide has the sequence 5'-CUUUUCACUUCCUGCCGGGG-3' , 5′-AGCUUCCAUGUCCAGCCUUC-3′, 5′- ACCAUGAAGUCAAAAUCAUU-3′, and 5′- GCUUUC AGCCCCGUUUCUUG-3'.

[0064] In some embodiments, provided herein are: (i) an adenosine base editor or a compound encoding the same; (ii) a guide polynucleotide or a nucleic acid sequence encoding the same; 1. A base editor system comprising contacting said adenosine base editor contains a programmable DNA binding domain and an adenosine deaminase domain wherein the adenosine deaminase domain is at amino acid position 8 in the numbering of SEQ ID NO:2 2 or 166 or a corresponding amino acid substitution at said guide polynucleotide; Otides induce the adenosine base editor to bind to the ATP-binding cassette subfamily members. A to G nucleobase alteration in the ABCA4 gene or its regulatory elements Thus, a base editor system is provided.

[0065] In one embodiment, the administration improves at least one symptom associated with Stargardt disease. In one embodiment, the administration induces the amino acid substitution in adenosine deaminase. Compared with treatment with base editors without genomic DNA fragments, In one embodiment, the ABCA4 gene is a steroid hormone. In one embodiment, the A to G nucleobase modification comprises a SNP associated with Lugart's disease. In one embodiment, the SNP is associated with Stargardt disease. The SNP associated with Gardt's disease is identified as A10 in the ABCA4 polypeptide as numbered in SEQ ID NO: 6. 38V or G1961E amino acid variant or variants thereof encoded by the ABCA4 gene In one embodiment, the SNP associated with Stargardt disease is SEQ ID NO: 6. The G1961E amino acid variant or variant thereof in the ABCA4 polypeptide is This brings about

[0066] In one embodiment, the A to G nucleobase alteration is a SNP associated with Stargardt disease. to a wild-type nucleobase. In one embodiment, the A to G nucleobase modification SNPs associated with Stargardt disease are identified as non-wild-type nucleic acids that result in improvement of Stargardt disease symptoms. In one embodiment, the guide polynucleotide is a Stargardt polynucleotide. The nucleic acid sequence is complementary to the ABCA4 gene or its regulatory element containing the disease-associated SNP. .

[0067] In one embodiment, the adenosine base editor is a polypeptide associated with Stargardt disease. A single guide comprising a nucleic acid sequence complementary to the ABCA4 gene or its regulatory element containing the SNP. In one embodiment, the sgRNA has the sequence 5′-CUCCAGGGCGAACUUC Contains GACACACAGC-3′.

[0068] In various aspects and embodiments provided herein, the polynucleotide The reprogrammable DNA binding domain is Cas9. In one embodiment, the Cas9 is SpCas 9, SaCas9, or a variant thereof. In one embodiment, the polynucleotide A more programmable DNA-binding domain is constructed with an engineered protospacer adjacent motif (P In one embodiment, the Cas9 comprises a modified SpCas9 with specificity for NGG, NGA, or GA. , NGCG, NGN, NNGRRT, NNNRRT, NGCG, NGCN, NGTN, and NGC It has specificity for PAM sequences, where N is A, G, C, or T, and R is A or G. In one embodiment, the polynucleotide-programmable DNA binding domain is a nucleic acid. In one embodiment, the polynucleotide is a cleavage-inactive variant. In one embodiment, the programmable DNA binding domain is a nickase variant. The nickase variant comprises the amino acid substitution D10A or a corresponding amino acid substitution. .

[0069] In various aspects and embodiments provided herein, the adenosine deaminase In one embodiment, the adenosine deaminase is and TadA deaminase containing the V82S modification and / or the T166R modification.

[0070] In various aspects and embodiments provided herein, the adenosine deaminase Ze is one or more of Y147T, Y147R, Q154S, Y123H, Q154R, or a combination thereof In various aspects and embodiments provided herein, the above Denosine deaminase is Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q154 R + T166R; Y147T + Q154R; Y147T + Q154S; and Y123H + Y147R + Q154R + I76Y In one embodiment, the adenosine base The editor domain comprises an adenosine deaminase monomer. The adenosine base editor comprises an adenosine deaminase dimer.

[0071] In various aspects and embodiments provided herein, the TadA deaminase is Ta In one embodiment, the TadA*8 variant is TadA*8.1, TadA*8 .2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8 .10, TadA*8.11, TadA*8.12, and TadA*8.13. In the example, the adenosine base editor is ABE8.1, ABE8.2, ABE8.3, ABE8.4, ABE8.5, AB Group consisting of E8.6, ABE8.7, ABE8.8, ABE8.9, ABE8.10, ABE8.11, ABE8.12, and ABE8.13 and an ABE8 base editor selected from:

[0072] In some embodiments, the adenosine base editors described herein are provided herein. In one embodiment, a vector is provided herein that comprises a nucleic acid sequence encoding the The adenosine base editors and guide polynucleotides encoding the polynucleotides described herein In one embodiment, the vector is a viral vector. vector, lentiviral vector, or AAV vector.

[0073] In some embodiments, the base editor systems described herein or In some embodiments, the cell is a central nervous system cell. In some embodiments, the cells are neurons. In some embodiments, the cells are photoreceptors. In certain embodiments, the cell is in vitro, in vivo, or ex vivo.

[0074] In some embodiments, the base editors, vectors, or cells and a pharmaceutically acceptable carrier. In another embodiment, the pharmaceutical compositions described herein further comprise a lipid. The pharmaceutical compositions described herein further comprise a virus.

[0075] In some embodiments, the base editors or vectors described herein A kit is provided comprising:

[0076] In various embodiments of the methods described herein, the guide polynucleotide At least one nucleotide contains a non-natural modification. In certain embodiments, at least one nucleotide of the nucleic acid sequence comprises a non-natural modification. In various embodiments, at least one nucleotide of the nucleic acid sequence is non-naturally occurring. In one embodiment, the non-natural modification is a chemical modification. In one embodiment, the chemical modification is 2'-O-methylation. Contains oate.

[0077] The description and examples herein particularly illustrate embodiments of the present disclosure. It is understood that the present invention is not limited to the particular embodiments described in the specification, as such may vary. Those skilled in the art will recognize that this disclosure contains numerous variations and modifications that are encompassed within its scope. You will realize there are fixes.

[0078] Implementation of the embodiments disclosed herein is within the skill of those of ordinary skill in the art, unless otherwise indicated. Immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics and and conventional techniques of recombinant DNA. See, e.g., Sambrook and Green, Molecular Cloning : A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molec ular Biology (FM Ausubel, et al. eds.); the series Methods In Enzymology (Aca (demic Press, Inc.), PCR 2: A Practical Approach (MJ MacPherson, BD Hames and GR Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialize See Applications, 6th Edition (R.I. Freshney, ed. (2010)).

[0079] The section headings used herein are for organizational purposes only. and should not be construed as limiting the subject matter described.

[0080] Various features of the disclosure may be described in the context of a single embodiment, but these features may also be combined in a single embodiment. They may also be provided separately or in any suitable combination. Although for clarity, may be described in the context of separate embodiments, the present disclosure also provides a single embodiment. The section headings used herein are for organizational purposes only. The information is for illustrative purposes only and should not be construed as limiting the subject matter described.

[0081] The features of the present disclosure are set forth with particularity in the appended claims. , which presents an illustrative embodiment in which the principles of the present disclosure are utilized. A better understanding of the features and advantages of the present disclosure will be provided in light of the accompanying drawings, which are described below. is obtained.

[0082] definition The following definitions supplement those in the art and are intended for the present application: Related or unrelated matters, such as those resulting from commonly owned patents or applications Any methods and materials similar or equivalent to those described herein are not intended to be limiting. Although the materials and methods described herein may be used in carrying out the tests shown, preferred materials and methods are described herein. Therefore, the terminology used herein is for the purpose of describing particular embodiments only. It is for illustrative purposes only and is not intended to be limiting.

[0083] Unless otherwise defined, all technical and scientific terms used herein are defined by the It has the meaning commonly understood by one of ordinary skill in the art to which the invention pertains. , provides those skilled in the art with general definitions of many of the terms used in this invention: n et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The C ambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary o f Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991).

[0084] In this application, the use of the singular includes the plural unless specifically stated otherwise. As used herein, the singular forms "a," "an," and "the" are used unless the context clearly indicates otherwise. It should be noted that unless specifically indicated, plural referents are included. In this context, the use of "or" means and includes "and / or" unless otherwise stated. It will be understood that the terms "including," "include," "includes," and "i The use of "includes" and other forms such as "included" is non-limiting.

[0085] As used in this specification and claims, the terms "comprising" (and " "comprise" and "comprises" and all its forms), "having" g) (any of its forms, such as "have" and "has"), "include "including" ("include" and "includes" and any of its forms) or "including containing (any of its forms, such as "contains" and "contain") ) is inclusive or open-ended and does not exclude additional, unrecited elements or method steps. Any embodiment discussed herein may be used with respect to any method or composition of the present disclosure. It is believed that the same can be done, and vice versa. The method of the present disclosure can be achieved by

[0086] The terms "about" or "approximately" refer to a range of values ​​as determined by one of ordinary skill in the art. This means that the value is within an acceptable margin of error for a particular value, which indicates how How it is measured or determined depends in part on the limitations of the measurement system. For example, "about" means, according to practice in the art, within 1 or more than 1 standard deviation. Alternatively, "about" can mean up to 20%, up to 10%, up to 5%, or Alternatively, it may refer to a range of up to 1% of the total mass of a biological system or process. Therefore, the term can mean values ​​within the same order of magnitude, e.g., within 5 times, or within 2 times. Where specific values ​​are described in the application and claims, unless otherwise stated, The term "about" means within an acceptable range of error for that particular value. It should be estimated.

[0087] Ranges provided herein are understood to be shorthand for all values ​​within the range. For example, the range 1 to 50 is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16. , 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36 , 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 It is understood to include any number, combination of numbers, or subrange.

[0088] In the specification, "some embodiments," "an embodiment," "one embodiment," or Reference to "another embodiment" may include any particular features, structures, or feature is included in at least some embodiments of the present disclosure, but not necessarily all. This means that the embodiments are not necessarily included.

[0089] "Abasic base editor" is a tool that extracts nucleobases and converts them into DNA nucleobases (A, T, Abasic base editors refer to agents that can insert nucleic acid groups (C, G, or G). In one embodiment, the nucleic acid glycosylase comprises a glycosylase polypeptide or a fragment thereof. at amino acid 204 of the following sequence or the corresponding position in uracil DNA glycosylase: containing Asp (e.g., substituting Asn at amino acid 204) and cytosine-DNA glycosylation A mutant human uracil DNA glycosylase or an active fragment thereof having enzyme activity. In one embodiment, the nucleic acid glycosylase comprises amino acid 147 or uracil of the following sequence: containing Ala, Gly, Cys, or Ser at the corresponding positions of a DNA glycosylase (e.g., a mutation in which Tyr is substituted at amino acid 147, and which has thymine-DNA glycosylase activity. An exemplary human uracil-DNA glycosylase is a human uracil-DNA glycosylase or an active fragment thereof. The sequence of A glycosylase, isoform 1 is as follows: 1 mgvfclgpwg lgrklrtpgk gplqllsrlc gdhlqaipak kapagqeepg tppssplsae 61 qldriqrnka aallrlaarn vpvgfgeswk khlsgefgkp yfiklmgfva eerkhytvyp 121 pphqvftwtq mcdikdvkvv ilgqdp y hgp nqahglcfsv qrpvppppsl eniykelstd 181 iedfvhpghg dlsgwakqgv lll n avltvr ahqanshker gweqftdavv swlnqnsngl 241 vfllwgsyaq kkgsaidrkr hhvlqtahps p l svy r gffg crhfsktnel lqksgkkpid 301 wkel

[0090] The sequence of human uracil-DNA glycosylase, isoform 2 is as follows: 1 migqktlysf fspsparkrh apspepavqg tgvagvpees gdaaaipakk apagqeepgt 61 ppssplsaeq ldriqrnkaa allrlaarnv pvgfgeswkk hlsgefgkpy fiklmgfvae 121 erkhytvypp phqvftwtqm cdikdvkvvi lgqdp y hgpn qahglcfsvq rpvppppsle 181 niykelstdi edfvhpghgd lsgwakqgvl ll n avltvra hqanshkerg weqftdavvs 241 wlnqnsnglv fllwgsyaqk kgsaidrkrh hvlqtahpsp l svy r gffgc rhfsktnell 301 qksgkkpidw kel

[0091] In other embodiments, the abasic editor is a base editor as described in PCT / JP205 / 080958 and US20170321210. The base editor may be any of the base editors listed in the In certain embodiments, the abasic editor is shown in bold and underlined in the sequence above. or any other abasic editor or uracil degumming known in the art. In one embodiment, the abasic enzyme comprises a mutation at the corresponding amino acid in the glycosylase. The diter contains mutations at Y147, N204, L272, and / or R276 or corresponding positions. In another embodiment, the abasic editor is a Y147A or Y147G mutation or the corresponding In another embodiment, the abasic editor comprises a N204D mutation or a corresponding In another embodiment, the abasic editor comprises a L272A mutation or a corresponding In another embodiment, the abasic editor comprises a R276E or R276C mutation. or corresponding mutations.

[0092] "Adenosine deaminase" refers to the enzyme that hydrolyzes the deamination of adenine or adenosine. In one embodiment, the term "de" refers to a polypeptide or fragment thereof that is capable of catalyzing the deactivation of a protein. The aminase or deaminase domain converts adenosine to inosine or deoxyribonucleic acid. Adenosine deamina catalyzes the hydrolytic deamination of adenosine to deoxyinosine. In some embodiments, the adenosine deaminase is a deoxyribonucleic acid (DNA) The compounds provided herein catalyze the hydrolytic deamination of adenine or adenosine in the ribozyme. Adenosine deaminases (e.g., engineered adenosine deaminases, evolved The adenosine deaminase (enzyme-activated adenosine deaminase) can be from any organism, such as a bacterium.

[0093] In one embodiment, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is a TadA variant. The dA variant is TadA*8. In one embodiment, the deaminase or deaminase The domains may be, for example, human, chimpanzee, gorilla, monkey, cow, dog, rat, or mammal. In some embodiments, the deaminase is a variant of a naturally occurring deaminase from an organism such as a mouse. The enzyme or deaminase domain is non-naturally occurring. For example, in some embodiments In this embodiment, the deaminase or deaminase domain is At least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least At least 92%, at least 93%, at least, at least 94%, at least 95%, at least 96%, At least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, At least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99 0.7%, at least 99.8%, or at least 99.9% identity. For example, Zedaima is a registered trademark of International PCT Application Nos. PCT / 2007 / 045381 (WO2018 / 027078) and PCT / US201 6 / 058344 (WO 2017 / 070632), each of which is incorporated by reference in its entirety. See also Komor, AC, et al., “Programmable editing of fa target base in genomic DNA without double-stranded DNA cleavage” Nature 533 , 420-424 (2016); Gaudelli, NM, et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, AC, et al., “Improved base excision repair inhibition and bacteriophage Mu Ga m protein yields C:G-to-T:A base editors with higher efficiency and product puri ty” Science Advances 3:eaao4774 (2017) ), and Rees, HA, et al., “Base edit ing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec;19(12):770-788. See also doi: 10.1038 / s41576-018-0059-1 (see also the entire contents of which are incorporated herein by reference).

[0094] The wild-type TadA (wt) adenosine deaminase has the following sequence (also known as the TadA reference sequence): called): MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHEIMALRQGGLVMQNYRLIDATL YVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQK KAQSSTD (SEQ ID NO: 2).

[0095] In some embodiments, the adenosine deaminase has a modification in the following sequence: include: MSEVEFSHEY WMRHALTLAK RARDEREVPV GAVLVLNNRV IGEGWNRAIG LHDPTAHAEI MALRQGGLVM QNY RLIDATL YVTFEPCVMC AGAMIHSRIG RVVFGVRNAK TGAAGSLMDV LHYPGMNHRV EITEGILADE CAALLC YFFR MPRQVFNAQK KAQSSTD (also known as TadA*7.10).

[0096] In some embodiments, TadA*7.10 comprises at least one modification. In embodiments, TadA*7.10 contains modifications at amino acids 82 and / or 166. In embodiments, variants of the reference sequence include one or more of the following modifications: Y1 47T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. The modification Y123H is It is also called H123H in the literature (the modification H123Y in TadA*7.10 has been reverted to Y123H(wt)). In other embodiments, the variant of the TadA*7.10 sequence is selected from the group consisting of: Includes combinations of changes: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q1 54S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T ; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154 R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y1 47R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R.

[0097] In other embodiments, the present invention provides a method for the preparation of a nucleic acid comprising a nucleic acid sequence selected from residues 149, 150, 151, 152, 153, 154, 155, 156, or or adenosine deaminase variants containing deletions, including C-terminal deletions beginning at 157 In another embodiment, adenosine deaminase variants, such as TadA*8, are provided. The target is a TadA (e.g., TadA*8) monomer containing one or more of the following modifications: Y147T , Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. The adenosine deaminase variants are those that contain a combination of modifications selected from the group consisting of: TadA monomers containing the following combinations (e.g., TadA*8): Y147T + Q154R; Y147T + Q154S; Y147 R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76 Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R.

[0098] In yet other embodiments, the adenosine deaminase variants each have the following modifications: One or more of the following variants: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R It is a homodimer containing two adenosine deaminase domains (e.g., TadA*8) In other embodiments, the adenosine deaminase variants are each selected from the following: two adenosine deaminase domains having a combination of modifications selected from the group consisting of: (e.g., TadA*8) are homodimers containing: Y147T + Q154R; Y147T + Q154S; Y147R + Q1 54S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H ; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82 S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R.

[0099] In other embodiments, the adenosine deaminase variant is a wild-type TadA adenosine deaminase variant. The deaminase domain and the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and adenosine deaminase variant domains containing one or more of Q154R and / or Q154R (e.g. In another embodiment, the adenosine deaminase The variant comprises a wild-type TadA adenosine deaminase domain and a variant selected from the group consisting of: adenosine deaminase variant domains (e.g., TadA) containing a combination of modifications to be *8) and heterodimers containing: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V 82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147 R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y1 23H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R.

[0100] In another embodiment, the adenosine deaminase variant comprises a TadA*7.10 domain and: one of the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R adenosine deaminase variant domains containing TadA*8 or more (e.g., TadA*8) and heterodimers containing TadA*8 or more In another embodiment, the adenosine deaminase variant is TadA*7.1 0 domain and adenosine deaminase variant domains containing a combination of the following modifications: and heterodimers containing the following amino acids (e.g., TadA*8): Y147T + Q154R; Y147T + Q154S; Y147 R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76 Y; V82S + Y123H + Y147R + Q154R; or I76Y + V82S + Y123H + Y147R + Q154R.

[0101] In one embodiment, the adenosine deaminase has the following sequence or adenosine deaminase TadA*8 comprises or consists essentially of a fragment thereof having enzyme activity: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQK KAQSSTD.

[0102] In some embodiments, TadA*8 is truncated. , the truncated TadA*8 has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 1 Lacking 4, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues. In terms of morphology, the truncated TadA*8 has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 , lacking 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues. In some embodiments, the adenosine deaminase variant is full-length TadA*8. be.

[0103] In certain embodiments, the adenosine deaminase heterodimer comprises a TadA*8 domain and an adenosine deaminase domain selected from one of the following:

[0104] In certain embodiments, the adenosine deaminase heterodimer comprises a TadA*8 domain and an adenosine deaminase domain selected from one of the following:

[0105] Escherichia coli TadA: MRRAFITGVFFLSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHEIMALRQGGLV MQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFF RMRRQEIKAQKKAQSSTD

[0106] E. coli TadA (N-terminal truncated): MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQK KAQSSTD

[0107] Staphylococcus aureus (S. aureus) TadA: MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTL YVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN

[0108] Bacillus subtilis (B. subtilis) TadA: MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTL EPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLS E

[0109] Salmonella typhimurium (S. typhimurium) TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLV LQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFF RMRRQEIKALKKADRAEGAGPAV

[0110] Shewanella putrefaciens (S. putrefaciens) TadA: MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEP CAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQ QGIE

[0111] Haemophilus influenzae F3031 (H. influenzae) TadA: MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTΑΗAEIIALRNGAKNIQNYR LLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRRE EKKIEKALLKSLSDK

[0112] Caulobacter crescentus (C. crescentus) TadA: MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLT DLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAK I

[0113] Geobacter sulfurreducens (G. sulfurreducens) TadA: MSSLKKTPIRDDAYWMGKAIREAAKAAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLT GATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRK KAKATPALFIDERKVPPEP

[0114] TadA*7.10 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD

[0115] An additional TadA7.10 or TadA7.10 is proposed as a component of the heterodimer with TadA*8. TadA7.10 variants include:

[0116] GSSGSETPGTSESATPESSGSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIM ALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADEC AALLCYFFRMPRQVFNAQKKAQSSTD

[0117] TadA7.10 CP65 TAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITE GILADECAALLCYFFRMPRQVFNAQKKAQSSTDGSSGSETPGTSESATPESSGSEVEFSHEYWMRHALTLAKRARDEREV PVGAVLVLNNRVIGEGWNRAIGLHDP

[0118] TadA7.10 CP83 YRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMP RQVFNAQKKAQSSTDGSSGSETPGTSESATPESSGSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWN RAIGLHDPTAHAEIMALRQGGLVMQN

[0119] TadA7.10 CP136 MNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTDGSSGSETPGTSESATPESSGSEVEFSHEYWMRHALTLA KRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRI GRVVFGVRNAKTGAAGSLMDVLHYPG

[0120] TadA7.10 C-truncated GSSGSETPGTSESATPESSGSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIM ALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADEC AALLCYFFRMPRQVFN

[0121] TadA7.10 C-truncated 2 GSSGSETPGTSESATPESSGSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIM ALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADEC AALLCYFFRMPRQ

[0122] TadA7.10 delta59-66+C-truncated GSSGSETPGTSESATPESSGSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNNRVIGEGWNRAHAEIMALRQGGLV MQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFF RMPRQVFN

[0123] TadA7.10 delta 59-66 GSSGSETPGTSESATPESSGSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNNRVIGEGWNRAHAEIMALRQGGLV MQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFF RMPRQVFNAQKKAQSSTD.

[0124] In some embodiments, the adenosine deaminase variant is in TadA7.10. In some embodiments, TadA7.10 comprises a modification at amino acid 82 or 166. In certain embodiments, the variant of the reference sequence includes any of the following modifications: Includes one or more of: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and Q154R. In this state, the adenosine deaminase variants are Y147R + Q154R + Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y147T + Q154R; Y147T + Q154S; and Y123H + Y147R + Q154R + I76Y.

[0125] In other embodiments, the present invention provides adenosine deaminase variants comprising deletions, e.g. For example, C-terminal deletions beginning at residues 149, 150, 151, 152, 153, 154, 155, 156, or 157 In another embodiment, the adenosine deaminase variant TadA7.10 is provided. is a TadA monomer containing one or more of the following modifications: Y147T, Y147R, Q154S, Y12 3H, V82S, T166R, Q154R. In other embodiments, the adenosine deaminase variant is Monomers containing the following modifications: Y147R + Q154R + Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y147T + Q154R; Y147T + Q154S; and Y123H + Y147R + Q154R + I7 In yet other embodiments, the adenosine deaminase variants are each selected from the group consisting of: Two with one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, Q154R In another embodiment, the adenosine deaminase domain is a homodimer containing the adenosine deaminase domain. The adenosine deaminase variants are expressed in either the wild-type adenosine deaminase domain or the TadA7.10 domain. and an adenosine deaminase variant domain comprising one or more of the following modifications: It is a heterodimer containing: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and Q154R In another embodiment, the adenosine deaminase variant comprises a TadA7.10 domain and and a heterodimer containing an adenosine deaminase variant of TadA7.10 containing the modification Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y147T + Q154R; Y147T + Q154S; and Y123H + Y147R + Q154R + I76Y.

[0126] "Administering" refers to providing one or more compositions described herein to a patient or subject. By way of example and not limitation, administration of a composition For example, injections can be intravenous (iv), subcutaneous (sc), intradermal (id), or intraperitoneal (i. p.) injection or intramuscular (i.m.) injection. Using more than one such route. Parenteral administration can be, for example, by bolus injection or by gradual infusion over time. In some embodiments, parenteral administration can be achieved by infusing the Intraductal, intravenous, intramuscular, intraarterial, intrathecal, intratumoral, intradermal, intraperitoneal, transtracheal, subcutaneous, subcorneal , including intra-articular, intracapsular, intrathecal and intrasternal infusion or injection; or Concurrently, administration can be by the oral route.

[0127] An "agent" is any small molecule compound, antibody, nucleic acid molecule, or polypeptide. , or fragments thereof.

[0128] "Alteration" refers to any alteration that can be detected by standard art known methods such as those described herein. Such a change in the structure, expression level or activity of a gene or polypeptide (e.g., an increase or As used herein, modification means a change in a polynucleotide or A change in the sequence of a polypeptide or a change in expression level, for example, a 10% change, a 25% change, This includes a 40% change, a 50% change, or a greater change in expression level.

[0129] "Ameliorate" means to reduce, inhibit, or attenuate the occurrence or progression of a disease. "To cause, reduce, stop, or stabilize" means to cause, reduce, stop, or stabilize.

[0130] "Analog" means a molecule that is not identical but has similar functional or structural characteristics. For example, a polynucleotide or polypeptide analog is a polynucleotide or polypeptide analog that is similar to the corresponding naturally occurring polynucleotide. A naturally occurring polynucleotide or polypeptide can be synthesized while retaining the biological activity of the nucleotide or polypeptide. or polypeptides, and have specific modifications that enhance the function of the analogs. Such modifications can improve the DNA affinity, efficiency, etc. of the analogs without altering, for example, ligand binding. affinity, specificity, protease or nuclease resistance, membrane permeability, and / or half-life Analogs can be used to increase the frequency of non-naturally occurring polynucleotides or amino acids. It may include.

[0131] A "base editor (BE)" or "nucleobase editor (NBE)" is a In various embodiments, the term "agent" refers to an agent that binds to an oxidase and has nucleobase modifying activity. Base editors include nucleobase-modifying polypeptides (e.g., deaminases) and nucleic acid promoters. The ramifiable nucleotide-binding domain is attached to a guide polynucleotide (e.g., a guide RNase A). In various embodiments, the agent comprises a protein having base editing activity. domains, i.e., bases (e.g., A, T, C, G, U) within a nucleic acid molecule (e.g., DNA), In some embodiments, the biomolecular complex comprises a domain capable of binding to a polypeptide. A nucleotide-programmable DNA-binding domain fused to a deaminase domain In one embodiment, the agent is a domain having base editing activity. In another embodiment, the fusion protein comprises a protein domain having base editing activity. The deaminase is ligated to the guide RNA (e.g., by ligating the RNA-binding motif and deaminase on the guide RNA). In some embodiments, the nucleic acid molecule has base editing activity (e.g., via an RNA-binding domain fused to an enzyme). In one embodiment, the domain is capable of deaminating a base within a nucleic acid molecule. The base editor can deaminate one or more bases in a DNA molecule. In , the base editor can deaminate adenosines (A) in DNA. In some embodiments, the base editor is an adenosine base editor (ABE).

[0132] "Cytidine deaminase" is a enzyme that converts amino groups into carbonyl groups through a deamination reaction. In some embodiments, the term "catalyzed polypeptide" refers to a polypeptide or fragment thereof that is capable of catalyzing In this case, the cytidine deaminase has at least about 85% identity to APOBEC or AID. In one embodiment, the cytidine deaminase converts cytosine to uracil and converts 5-methylcytosine to thymine. PmCDA1 from Petromyzon marinus (Petromyzon marinus cytosine deaminase 1, “PmCDA1”), or mammalian (e.g., human, porcine) AID (activation-induced cytidine deaminase; AICDA) derived from mammals such as cattle, horses, and monkeys, and APOBECs are exemplary cytidine deaminases.

[0133] In some embodiments, the base editor is a deaminase (e.g., adenosine Reprogrammable base editor fused to deaminase or cytidine deaminase In some embodiments, the base editor is a deaminase (e.g., Cas9 fused to adenosine deaminase or cytidine deaminase. In embodiments, the base editor is a deaminase (e.g., adenosine deaminase or nuclease-inactivated Cas9 (dCas9) fused to a cytidine deaminase (CdCas9). In some embodiments, the Cas9 is a circular permutant Cas9 (e.g., spCas9 or saCas9). Circularly permuted Cas9s are known in the art and are described, for example, in Oakes et al. et al., Cell 176, 254-267, 2019. In some embodiments, the salt Base editors are fused to inhibitors of base excision repair, e.g., UGI domains or dISN domains. In one embodiment, the fusion protein comprises a deaminase and a UGI domain or or a Cas9 nickase fused to an inhibitor of base excision repair, such as a dISN domain. In other embodiments, the base editor is an abasic base editor.

[0134] In some embodiments, the base editor is an adenosine base editor (ABE). In some embodiments, the adenosine deaminase is evolved from TadA. In some embodiments, the base editors of the invention are internally fused It contains a napDNAbp domain with a catalytic (e.g., deaminase) domain. In embodiments, napDNAbp comprises a Cas12a deaminase domain fused internally. In some embodiments, the napDNAbp is an internally fused deamidation In some embodiments, the napDNAbp is a Cas12b(c2c1) having a cleavage domain. is Cas12c(c2c3) with an internally fused deaminase domain. In one embodiment, napDNAbp comprises a Ca 2+ deaminase domain fused internally to the napDNAbp. In some embodiments, the napDNAbp is an internally fused deoxyribonucleotide. In some embodiments, napD is Cas12e (CasY) with an aminase domain. NAbp is a Cas12g with an internally fused deaminase domain. In an embodiment, napDNAbp is a Cas12h gene having an internally fused deaminase domain. In some embodiments, the napDNAbp is an internally fused deaminase In some embodiments, the base editor is a Cas12i having a deamidating domain. A catalytically dead Cas12 (dCas12) fused to a nase domain. In this study, the base editor is a Cas12 nickase (nCa) fused to a deaminase domain. s12).

[0135] In some embodiments, the base editor is an adenosine deaminase variant (e.g., TadA*8) with a circularly permuted Cas9 (e.g., spCAS9 or saCAS9) and a bipartite nuclear-localized Circular permutation is generated by cloning into a scaffold containing the targeting sequence (e.g., ABE8). The Cas9 enzyme is known in the art and is described, for example, in Oakes et al., Cell 176, 254-267, 2019. Exemplary circular permutations are described below, where the bolded sequences are derived from Cas9. The italicized sequences indicate linker sequences, and the underlined sequences indicate bipartite nuclear localization. The sequence is shown.

[0136] JPEG2025032080000002.jpg188168

[0137] In some embodiments, ABE8 is a base editor from Tables 6-9, 13, or 14 below. In some embodiments, ABE8 is an adenosine triphosphate (ABT)-dependent ATPase (ATA ... In some embodiments, the adenosine deaminase variant of ABE8 is The amino acid sequence variant is a TadA*8 variant as set forth in Tables 7, 9, 13, or 14 below. In some embodiments, the adenosine deaminase variant is Y147T, Y14 7R, Q154S, Y123H, V82S, T166R, and / or Q154R. In various embodiments, the ABE is a TadA*7.10 variant (e.g., TadA*8) that includes the mutation. 8 is Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V8 2S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147 R; V82S + Y123H + Q154R; Y147R + Q154R +Y123H; Y147R + Q154R + I76Y; Y147R + Q15 4R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I7 6Y + V82S + Y123H + Y147R + Q154R In some embodiments, ABE8 is a monomeric construct and includes TadA*7.10 (e.g., TadA*8). In some embodiments, ABE8 is a heterodimeric construct. In the ABE8 base editor, the sequence is: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQK KAQSSTD.

[0138] In some embodiments, the polynucleotide programmable DNA binding domain is CRISP In some embodiments, the base editor is a deamidating enzyme. In one embodiment, the Cas9 domain is a catalytically dead Cas9 (dCas9) fused to a cleavage domain. The base editor is a Cas9 nickase (nCas9) fused to a deaminase domain. In some embodiments, the base editor is fused to an inhibitor of base excision repair (BER). In some embodiments, the inhibitor of base excision repair is a uracil DNA glycosylase inhibitor. In some embodiments, the inhibitor of base excision repair is an inosine base excision repair inhibitor. It is a reversible inhibitor.

[0139] For more information on the base editor, see International PCT Application No. PCT / 2017 / 045381 (International Publication No. WO 2018 / 027078). ) and PCT / US 2016 / 058344 (International Publication No. 2017 / 070632), Komor, AC, et al., “Programmable editing ing of a target base in genomic DNA without double-stranded DNA cleavage” Natur e 533, 420-424 (2016); Gaudelli, NM, et al., “Programmable base editing of A ·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); K omor, AC, et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and produc t purity” Science Advances 3:eaao4774 (2017), and Rees, HA, et al., “Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec;19(12):770-788. See also doi: 10.1038 / s41576-018-0059-1 (the entire contents of which are incorporated herein by reference).

[0140] By way of example, base editing compositions, systems, and methods described herein may be used. The cytidine base editor used is the following nucleic acid sequence (8877 base pairs) (Addgene, Watertown, MA) .; Komor AC, et al., 2017, Sci Adv., 30;3(8):eaao4774. doi: 10.1126 / sciadv.aao47 74) A polynucleotide having at least 95% identity to the BE4 nucleic acid sequence. Arrangements of the amino acid sequence are also included.

[0141] 1 atatgccaag tacgccccct attgacgtca atgacggtaa atggcccgcc tggcattatg 61 cccagtacat gaccttatgg gactttccta cttggcagta catctacgta ttagtcatcg 121 ctattaccat ggtgatgcgg ttttggcagt acatcaatgg gcgtggatag cggtttgact 181 cacggggatt tccaagtctc caccccattg acgtcaatgg gagtttgttt tggcaccaaa 241 atcaacggga ctttccaaaa tgtcgtaaca actccgcccc attgacgcaa atgggcggta 301 ggcgtgtacg gtgggaggtc tatataagca gagctggttt agtgaaccgt cagatccgct 361 agagatccgc ggccgctaat acgactcact atagggagag ccgccaccat gagctcagag 421 actggcccag tggctgtgga ccccacattg agacggcgga tcgagcccca tgagtttgag 481 gtattcttcg atccgagaga gctccgcaag gagacctgcc tgctttacga aattaattgg 541 gggggccggc actccatttg gcgacataca tcacagaaca ctaacaagca cgtcgaagtc 601 aacttcatcg agaagttcac gacagaaaga tatttctgtc cgaacacaag gtgcagcatt 661 acctggttc tcagctggag cccatgcggc gaatgtagta gggccatcac tgaattcctg 721 tcaaggtatc cccacgtcac tctgtttat tacatcgcaa ggctgtacca ccacgctgac 781 ccccgcaatc gacaaggcct gcgggatttg atctcttcag gtgtgactat ccaaattatg 841 actgagcagg agtcaggata ctgctggaga aactttgtga attatagccc gagtaatgaa 901 gcccactggc ctaggtatcc ccatctgtgg gtacgactgt acgttcttga actgtactgc 961 atcatactgg gcctgcctcc ttgtctcaac attctgagaa ggaagcagcc acagctgaca 1021 ttctttacca tcgctcttca gtcttgtcat taccagcgac tgcccccaca cattctctgg 1081 gccaccgggt tgaaatctgg tggttcttct ggtggttcta gcggcagcga gactcccggg 1141 acctcagagt ccgccacacc cgaaagttct ggtggttctt ctggtggttc tgataaaaag 1201 tattctattg gtttagccat cggcactaat tccgttggat gggctgtcat aaccgatgaa 1261 tacaaagtac cttcaaagaa atttaaggtg ttgggggaaca cagaccgtca ttcgattaaa 1321 aagaatctta tcggtgccct cctattcgat agtggcgaaa cggcagaggc gactcgcctg 1381 aaacgaaccg ctcggagaag gtatacacgt cgcaagaacc gaatatgtta cttacaagaa 1441 atttttagca atgagatggc caagttgac gattctttct ttcaccgttt ggaagagtcc 1501 ttccttgtcg agaggacaa gaaacatgaa cggcacccca tctttggaa catagtagat 1561 gaggtggcat atcatgaaaa gtaccaacg atttacacc tcagaaaaaa gctagttgac 1621 tcaacgata aagcggacct gaggttaatc tacttggctc ttacccatat gataagttc 1681 cgtgggcact ttctcattga gggtgatcta aatccggaca actcggatgt cgacaactg 1741 ttcatccagt tagtacaac ctataatcag ttgtttgaag agaaccctat aaatgcaagt 1801 ggcgtggatg cgaaggctat tcttagccc cgcctctcta atcccgacg gctagaaac 1861 ctgatcgcac aattacccgg agagagaaa aatggttgt tcggtaacct tatagcgctc 1921 tcactaggcc tgacaccaaa tttagtcg aacttcgact tagctgaga tgccaattg 1981 cagcttagta aggacacgta cgatgacgat ctcgacaatc tactggcaca attggagat 2041 cagtagcgg acttatttt ggctgccaa aaccttagcg atgcaatcct cctactcgac 2101 atactgagag ttatactga gattaccaag gcgccgtta ccgctcaat gatcaaagg 2161 tacgatgac atcaccaaga cttgacactt ctcaaggccc tagtccgtca gcaactgcct 2221 gagaatata aggaatatt ctttgatcag tcgaaaaacg ggtacgcagg ttatattgac 2281 ggcggagcga gtcaagagga attctacaag tttatcaaac ccatattaga gaagatggat 2341 gggacggaag agttgcttgt aaaactcaat cgcgaagatc tactgcgaaa gcagcggact 2401 ttcgacaacg gtagcattcc acatcaaatc cacttaggcg aattgcatgc father 2461 aggcaggagg atttttatcc gttcctcaaa gacaatcgtg aaaagattga gaaaatccta 2521 acctttcgca taccttacta tgtgggaccc ctggcccgag ggaactctcg gttcgcatgg 2581 atgacaaga agtccgaaga aacgattact ccatggaatt ttgaggaagt tgtcgataaa 2641 ggtgcgtcag ctcaatcgtt catcgagagg atgaccaact ttgacaagaa tttaccgaac 2701 gaaaaagtat tgcctaagca cagtttactt tcgagtatt tcacagtgta caatgaactc 2761 acgaaagtta agtatgtcac tgagggcatg cgtaaacccg cctttctaag cggagaacag 2821 aagaaagcaa tagtagtct gttattcaag accaaccgca aagtgacagt tagcaattg 2881 aaagaggact actttaagaa aattgaatgc ttcgattctg tcgagatctc cggggtagaa 2941 gatcgattta atgcgtcact tggtacgtat catgacctcc taaagataat taaagataag 3001 gacttcctgg ataacgaaga gaatgaagat atcttagaag atatagtgtt gactcttacc 3061 ctctttgaag atcgggaaat gattgaggaa agactaaaaa catacgctca cctgttcgac 3121 gataaggtta tgaaacagtt aaagaggcgt cgctatacgg gctggggacg attgtcgcgg 3181 aaacttatca acgggataag agacaagcaa agtggtaaaa ctattctcga tttcttaaag 3241 agcgacggct tcgccaatag gaactttatg cagctgatcc atgatgactc tttaaccttc 3301 aaagaggata tacaaaaggc acaggtttcc ggacaagggg actcattgca cgaacatatt 3361 gcgaatcttg ctggttcgcc agccatcaaa aagggcatac tccagacagt caaagtagtg 3421 gatgagctag ttaaggtcat gggacgtcac aaaccggaaa acattgtaat cgagatggca 3481 cgcgaaaatc aaacgactca gaaggggcaa aaaaacagtc gagagcggat gaagagaata 3541 gaagagggta ttaaagaact gggcagccag atcttaaagg agcatcctgt ggaaaatacc 3601 cattgcaga acgagaact tacctactat tacctacaa atggaaggga catgtatgtt 3661 gatcaggaac tggacataaa ccgtttatct gattacgacg tcgatcacat tgtacccca 3721 tccttttga aggacgattc atcgacaat aaagtgctta cacgctcgga taagaaccga 3781 gggaaagtg acatgttcc aagcgaggaa gtcgtaaga aaatgaagaa ctattggcgg 3841 cagctcctaa atgcgaact gataacgca aggaagttcg attackac taagctgag 3901 agggtggct tgtgaact tgacaaggcc ggatttatta aacgtcagct cgtggaaacc 3961 cgccaaatca caagcatgt tgcacagata ctagattccc gatgaatac gaatacgac 4021 gagaacgata agctgattcg ggaagtcaa gtaatcactt taagtcaa attggtgtcg 4081 gacttcagaa aggatttca attcttaaa gttagggaga taaataacta ccaccatgcg 4141 cacgacgctt atcttaatgc cgtcgtaggg accgcactca ttaagaata cccgaagcta 4201 gaagtgagt ttgtgtatgg tgattacaaa gtttatgacg tccgtagat gatcgcgaaa 4261 agcgacagg agataggcaa ggctacagcc aaatactct tttattctaa cattatgaat 4321 ttctttaaga cggaaatcac tctggcaaac ggagagatac gcaaacgacc tttaattgaa 4381 accaatgggg agacaggtga aatcgtatgg gataagggcc gggacttcgc gacggtgaga 4441 aaagttttgt ccatgcccca agtcaacata gtaaagaaaa ctgaggtgca gaccggaggg 4501 tttcaaagg aatcgattct tccaaaagg aatagtgata agctcatcgc tcgtaaaaag 4561 gactgggacc cgaaaaagta cggtggcttc gatagcccta cagttgccta ttctgccta 4621 gtagtggcaa aagttgagaa gggaaaatcc aagaaactga agtcagtcaa agaattattg 4681 gggataacga tttggagcg ctcgtctttt gaaaagaacc ccatcgactt cttgaggcg 4741 aaaggttaca aggaagtaaa aaaggatctc ataattaaac taccaaagta tagtctgttt 4801 gagttagaaa atggccgaaa acggatgttg gctagcgccg gagagcttca aaaggggaac 4861 gaactcgcac taccgtctaa atacgtgaat ttcctgtatt tagcgtccca ttacgagaag 4921 ttgaaaggtt cacctgaaga taacgaacag aagcaacttt ttgttgagca gcacaaacat 4981 tatctcgacg aaatcataga gcaaatttcg gaattcagta agagagtcat cctagctgat 5041 gccaatctgg acaaagtatt aagcgcatac aacaagcaca gggataaacc catacgtgag 5101 caggcggaaa atattatcca tttgtttact cttaccaacc tcggcgctcc agccgcattc 5161 aagtattttg acacaacgat agatcgcaaa cgatacactt ctaccaagga ggtgctagac 5221 gcgacactga ttcaccaatc catcacggga ttatatgaaa ctcggataga tttgtcacag 5281 cttgggggtg actctggtgg ttctggagga tctggtggtt ctactaatct gtcagatatt 5341 attgaaaagg agaccggtaa gcaactggtt atccaggaat ccatcctcat gctcccagag 5401 gaggtggaag aagtcattgg gaacaagccg gaaagcgata tactcgtgca caccgcctac 5461 gacgagagca ccgacgagaa tgtcatgctt ctgactagcg acgcccctga atacaagcct 5521 tgggctctgg tcatacagga tagcaacggt gagaacaaga ttaagatgct ctctggtggt 5581 tctggaggat ctggtggttc tactaatctg tcagatatta ttgaaaagga gaccggtaag 5641 caactggtta tccaggaatc catcctcatg ctcccagagg aggtggaaga agtcattggg 5701 aacaagccgg aaagcgatat actcgtgcac accgcctacg acgagagcac cgacgagaat 5761 gtcatgcttc tgactagcga cgcccctgaa tacaagcctt gggctctggt catacaggat 5821 agcaacggtg agaacaagat taagatgctc tctggtggtt ctcccaagaa gaagaggaaa 5881 gtctaaccgg tcatcatcac catcaccatt gagtttaaac ccgctgatca gcctcgactg 5941 tgccttctag ttgccagcca tctgttgttt gcccctcccc cgtgccttcc ttgaccctgg 6001 aaggtgccac tcccactgtc ctttcctaat aaaatgagga aattgcatcg cattgtctga 6061 gtaggtgtca ttctattctg gggggtgggg tggggcagga cagcaagggg gaggattggg 6121 aagacaatag caggcatgct ggggatgcgg tgggctctat ggcttctgag gcggaaagaa 6181 ccagctgggg ctcgataccg tcgacctcta gctagagctt ggcgtaatca tggtcatagc 6241 tgtttcctgt gtgaaattgt tatccgctca caattccaca caacatacga gccggaagca 6301 taaagtgtaa agcctagggt gcctaatgag tgagctaact cacattaatt gcgttgcgct 6361 cactgcccgc tttccagtcg ggaaacctgt cgtgccagct gcattaatga atcggccaac 6421 gcgcgggg aggcggttg cgtattgggc gctcttcgc ttctcgctc actgactcgc 6481 tgcgctcggt cgttcggctg cggcgagcgg tatcagctca ctcaaaggcg gtaatacggt 6541 tatccacaga atcaggggat aacgcaggaa agaacatgtg agcaaaggc cagcaaaagg 6601 ccaggaaccg taaaaggcc gcgttgctgg cgtttttcca taggctccgc ccccctgacg 6661 agcatcacaa aaatcgacgc tcaagtcaga ggtggcgaaa cccgacagga ctataaagat 6721 accaggcgtt tccccctgga agctccctcg tgcgctctc tgttccgacc ctgccgctta 6781 ccggatacct gtccgccttt ctcccttcgg gaagcgtggc gctttctcat agctcacgct 6841 gtaggtatct cagttcggtg tagtcgttc gctccaagct gggctgtgtg cacgaacccc 6901 ccgttcagcc cgaccgctgc gccttatccg gtaactatcg tcttgagtcc aacccggtaa 6961 gacacgactt atcgccactg gcagcagcca ctggtaacag gattagcaga gcgaggtatg 7021 taggcggtgc tacagagttc ttgaagtggt ggcctaacta cggctacact agaagaacag 7081 tatttggtat ctgcgctctg ctgaagccag ttaccttcgg aaaaagagt ggtagctctt 7141 gatccggcaa acaaaccacc gctggtagcg gtggtttttt tgtttgcaag cagcagatta 7201 cgcgcagaaa aaaaggatct caagaagatc ctttgatctt ttctacgggg tctgacgctc 7261 aatgggaacga aaactcacgt taagggattt tggtcatgag attatcaaaa aggatcttca 7321 7381 cttggtctga cagttaccaa tgcttaatca gtgaggcacc tatctcagcg atctgtctat 7441 ttcgttcatc catagttgcc tgactccccg tcgtgtagat aactacgata cgggagggct 7501 taccatctgg ccccagtgct gcaatgatac cgcgagaccc acgctcaccg gctccagatt 7561 tatcagcaat aaaccagcca gccggaaggg ccgagcgcag aagtggtcct gcaactttat 7621 ccgcctccat ccagtctatt aattgttgcc gggaagctag agtaagtagt tcgccagtta 7681 atagtttgcg caacgttgtt gccattgcta caggcatcgt ggtgtcacgc tcgtcgtttg 7741 gtatggcttc attcagctcc ggttcccaac gatcaaggcg agttacatga tcccccatgt 7801 tgtgcaaaaa agcggttagc tccttcggtc ctccgatcgt tgtcagaagt aagttggccg 7861 cagtgttatc actcatggtt atggcagcac tgcataattc tcttactgtc atgccatccg 7921 taagatgctt ttctgtgact ggtgagtact caaccaagtc attctgagaa tagtgtatgc 7981 ggcgaccgag ttgctcttgc ccggcgtcaa tacgggataa taccgcgcca catagcagaa 8041 ctttaaaagt gctcatcatt ggaaaacgtt cttcggggcg aaaactctca aggatcttac 8101 cgctggtgag atccagttcg atgtaaccca ctcgtgcacc caactgatct tcagcatctt 8161 ttactttcac cagcgtttct gggtgagcaa aaacaggaag gcaaaatgcc gcaaaaaag 8221 gaataagggc gacacggaaa tgttgaatac tcatactctt cctttttcaa tattattgaa 8281 gcatttatca gggttatgt ctcatgagcg gatacatatt tgaatgtatt tagaaaaata 8341 aacaaatagg ggttccgcgc acatttcccc gaaaagtgcc acctgacgtc gacggatcgg 8401 gagatcgatc tcccgatccc ctagggtcga ctctcagtac aatctgctct gatgccgcat 8461 agttaagcca gtatctgctc cctgcttgtg tgttggaggt cgctgagtag tgcgcgagca 8521 aaatttaagc tacaacaagg caaggcttga ccgacaattg catgaagaat ctgcttaggg 8581 ttaggcgttt tgcgctgctt cgcgatgtac gggccagata tacgcgttga cattgattat 8641 tgactagtta ttaatagtaa tcaattacgg ggtcattagt tcatagccca tatatggagt 8701 tccgcgttac ataacttacg gtaaatggcc cgcctggctg accgcccaac gacccccgcc 8761 cattgacgtc aataatgacg tatgttccca tagtaacgcc aatagggact ttccattgac 8821 gtcaatgggt ggagtatta cggtaaactg cccacttggc agtacatcaa gtgtatc

[0142] BE4 amino acid sequence: MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNT RCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYS PSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPPHILWATGLKSGGSSGGSSGS ETPGTSESATPESSGGSSGGSDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAE ATRLKRTARRRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRK KLVDSTDKADLRIYLALAMHIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSR RLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDLNLLAQIGDQYADLFLAAKNLSDAI LLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPIL EKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNS RFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFL SGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIV LTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDD SLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRER MKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRS DKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMN TKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRK MIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEV QTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPID FLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTK EVLDATLIHQSITGLYETRIDLSQLGGDSGGSGGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILV HTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLSGGSGGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVE EVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLSGGSPKKKRK

[0143] By way of example, adenine dinucleotides used in the base editing compositions, systems, and methods described herein can be used. The Additive Base Editor (ABE) has the nucleic acid sequence (8877 base pairs) provided below (Addge ne, Watertown, MA.; Gaudelli NM, et al., Nature. 2017 Nov 23;551(7681):464-471. doi: 10.1038 / nature24644; Koblan LW, et al., Nat Biotechnol. 2018 Oct;36(9):8 43-846. doi: 10.1038 / nbt.4172.) At least 95% identity to the ABE nucleic acid sequence Also encompassed are polynucleotide sequences having the following structure:

[0144] ATATGCCAAGTACGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTATGCCCAGTACAT GACCTTATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATGGTGATGCGG TTTTGGCAGTACATCAATGGGCGTGGATAGCGGTTTGACTCACGGGATTTCCAAGTCTCCACCCCATTG ACGTCAATGGGAGTTTGTTTTGGCACCAAAATCAACGGGACTTTCCAAAATGTCGTAACAACTCCGCCCC ATTGACGCAAATGGGCGGTAGGCGTGTACGGTGGGAGGTCTATATAAGCAGAGCTGGTTTAGTGAACCGT CAGATCCGCTAGAGATCCGCGGCCGCTAATACGACTCACTATAGGGAGAGCCGCCACCATGAAACGGACA GCCGACGGAAGCGAGTTCGAGTCACCAAAGAAGAAGCGGAAAGTCTCTGAAGTCGAGTTTAGCCACGAGT ATTGGATGAGGCACGCACTGACCCTGGCAAAGCGAGCATGGGATGAAAGAGAAGTCCCCGTGGGCGCCGT GCTGGTGCACAACAATAGAGTGATCGGAGAGGGATGGAACAGGCCAATCGGCCGCCACGACCCTACCGCA CACGCAGAGATCATGGCACTGAGGCAGGGAGGCCTGGTCATGCAGAATTACCGCCTGATCGATGCCACCC TGTATGTGACACTGGAGCCATGCGTGATGTGCGCAGGAGCAATGATCCACAGCAGGATCGGAAGAGTGGT GTTCGGAGCACGGGACGCCAAGACCGGCGCAGCAGGCTCCCTGATGGATGTGCTGCACCACCCCGGCATG AACCACCGGGTGGAGATCACAGAGGGAATCCTGGCAGACGAGTGCGCCGCCCTGCTGAGCGATTTCTTTA GAATGCGGAGACAGGAGATCAAGGCCCAGAAGAAGGCACAGAGCTCCACCGACTCTGGAGGATCTAGCGG AGGATCCTCTGGAAGCGAGACACCAGGCACAAGCGAGTCCGCCACACCAGAGAGCTCCGGCGGCTCCTCC GGAGGATCCTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGCCAAGAGGG CACGCGATGAGAGGGAGGTGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGCGAGGGCTG GAACAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATGGCCCTGAGACAGGGCGGCCTG GTCATGCAGAACTACAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCG GCGCCATGATCCACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACGCAAAAACCGGCGCCGCAGG CTCCCTGATGGACGTGCTGCACTACCCCGGCATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCA GATGAATGTGCCGCCCTGCTGTGCTATTTCTTTCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAAGG CCCAGAGCTCCACCGACTCCGGAGGATCTAGCGGAGGCTCCTCTGGCTCTGAGACACCTGGCACAAGCGA GAGCGCAACACCTGAAAGCAGCGGGGGCAGCAGCGGGGGGTCAGACAAGAAGTACAGCATCGGCCTGGCC ATCGGCACCAACTCTGTGGGCTGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGG TGCTGGGCAACACCGACCGGCACAGCATCAAGAAGAACCTGATCGGAGCCCTGCTGTTCGACAGCGGCGA AACAGCCGAGGCCACCCGGCTGAAGAGAACCGCCAGAAGAAGATACACCAGACGGAAGAACCGGATCTGC TATCTGCAAGAGATCTTCAGCAACGAGATGGCCAAGGTGGACGACAGCTTCTTCCACAGACTGGAAGAGT CCTTCCTGGTGGAAGAGGATAAGAAGCACGAGCGGCACCCCATCTTCGGCAACATCGTGGACGAGGTGGC CTACCACGAGAAGTACCCCACCATCTACCACCTGAGAAAGAAACTGGTGGACAGCACCGACAAGGCCGAC CTGCGGCTGATCTATCTGGCCCTGGCCCACATGATCAAGTTCCGGGGCCACTTCCTGATCGAGGGCGACC TGAACCCCGACAACAGCGACGTGGACAAGCTGTTCATCCAGCTGGTGCAGACCTACAACCAGCTGTTCGA GGAAAACCCCATCAACGCCAGCGGCGTGGACGCCAAGGCCATCCTGTCTGCCAGACTGAGCAAGAGCAGA CGGCTGGAAAATCTGATCGCCCAGCTGCCCGGCGAGAAGAAGAATGGCCTGTTCGGAAACCTGATTGCCC TGAGCCTGGGCCTGACCCCAACTTCAAGAGCAACTTCGACCTGGCCGAGGATGCCAAACTGCAGCTGAG CAAGGACACCTACGACGACGACGACCTGGACAACCTGCTGGCCCAGATCGGCGACCAGTACGGCCGACCTGTTTT CTGGCCGCCAAGAACCTGCTCGACGCCATCCTGCTGAGCGACATCCTGAGAGTGAACACCGAGATCACCA AGGCCCCCCTGAGCGCCTCTATGATCAAGAGATACGACGAGCACCACCAGGACCTGACCCTGCTGAAAGC TCTCGTGCGGCAGCAGCTGCCTGAGAAGTACAAAGAGATTTTCTTCGACCAGAGCAAGAACGGCTACGCC GGCTACATTGACGGCGGAGCCAGCCAGGAAGAGTTCTACAAGTTCATCAAGCCCATCCTGGAAAAGATGG ACGGCACCGAGGAACTGCTCGTGAAGCTGAACAGAGGACCTGCTGCGGAAGCAGCGGACCTTCGACAA CGGCAGCATCCCCCACCAGATCCACCTGGGAGAGCTGCACGCCATTCTGCGGCGGCAGGAAGATTTTTAC CCATTCCTGAAGGACAACCGGGAAAGATCGAGAAGATCCTGACCTTCCGCATCCCCTACTACGTGGGCC CTCTGGCCAGGGGAAACAGCAGATTCGCCTGGATCCAGAAAAGGCGAGGAAACCATCACCCCCTGGAA CTTCGAGGAAGTGGTGGACAAGGGCGCTTCCGCCCCAGAGCTTCATCGAGCGGATGACCACTTCGATAAG AACCTGCCCAACGAGAAGGTGCTGCCCAAGCACAGCCTGCTGTACGAGTACTTCACCGTGTATAACGAGC TGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAAGCCCGCCTTCCTGAGCGGCGAGCAGAAAAAGGC CATCGTGGACCTGCTGTTCAAGACCAACCGGAAAGTGACCGTGAAGCAGCTGAAAGAGGACTACTTCAAG AAAATCGAGTGCTTCGACTCCGTGGAAATCTCCGGCGTGGAAGATCGGTTCAACGCCTCCCTGGGCACAT ACCACGATCTGCTGAAAATTATCAAGGACAAGGACTTCCTGGACAATGAGGAAAACGAGGACATTCTGGA AGATATCGTGCTGACCCTGACACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTGAAAACCTATGCC CACCTGTTCGACGACAAAGTGATGAAGCAGCTGAAGCGGCGGAGATACACCGGCTGGGGCAGGCTGAGCC GGAAGCTGATCAACGGCATCCGGGACAAGCAGTCCGGCAAGACAATCCTGGATTTCCTGAAGTCCGACGG CTTCGCCAACAGAAACTTCATGCAGCTGATCCACGACGACAGCCTGACCTTTAAAGAGGACATCCAGAAA GCCCAGGTGTCCGGCCAGGGCGATAGCCTGCACGAGCACATTGCCAATCTGGCCGGCAGCCCCGCCATTA AGAAGGGCATCCTGCAGACAGTGAAGGTGGTGGACGAGCTCGTGAAAGTGATGGGCCGGCACAAGCCCGA GAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCACCCAGAAGGGACAGAAGAACAGCCGCGAGAGA ATGAAGCGGATCGAAGAGGGCATCAAAGAGCTGGGCAGCCAGATCCTGAAAGAACACCCCGTGGAAAACA CCCAGCTGCAGAACGAGAAGCTGTACCTGTACTACCTGCAGAATGGGCGGGATATGTACGTGGACCAGGA ACTGGACATCAACCGGCTGTCCGACTACGATGTGGACCATATCGTGCCTCAGAGCTTTCTGAAGGACGAC TCCATCGACAACAAGGTGCTGACCAGAAGCGACAAGAACCGGGGCAAGAGCGACAACGTGCCCTCCGAAG AGGTCGTGAAGAAGATGAAGAACTACTGGCGGCAGCTGCTGAACGCCAAGCTGATTACCCAGAGAAAGTT CGACAATCTGACCAAGGCCGAGAGAGGCGGCCTGAGCGAACTGGATAAGGCCGGCTTCATCAAGAGACAG CTGGTGGAAACCCGGCAGATCACAAAGCACGTGGCACAGATCCTGGACTCCCGGATGAACACTAAGTACG ACGAGAATGACAAGCTGATCCGGGAAGTGAAAGTGATCACCCTGAAGTCCAAGCTGGTGTCCGATTTCCG GAAGGATTTCCAGTTTTACAAAGTGCGCGAGATCAACAACTACCACCACGCCCACGACGCCTACCTGAAC GCCGTCGTGGGAACCGCCCTGATCAAAAAGTACCCTAAGCTGGAAAGCGAGTTCGTGTACGGCGACTACA AGGTGTACGACGTGCGGAAGATGATCGCCAAGAGCGAGCAGGAAATCGGCAAGGCTACCGCCAAGTACTT CTTCTACAGCAACATCATGAACTTTTTCAAGACCGAGATTACCCTGGCCAACGGCGAGATCCGGAAGCGG CCTCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGTGGGATAAGGGCCGGGATTTTGCCACCGTGC GGAAAGTGCTGAGCATGCCCCAAGTGAATATCGTGAAAAAGACCGAGGTGCAGACAGGCGGCTTCAGCAA AGAGTCTATCCTGCCCAAGAGGAACAGCGATAAGCTGATCGCCAGAAAGAAGGACTGGGACCCTAAGAAG TACGGCGGCTTCGACAGCCCCACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAAGGGCAAGT CCAAGAAACTGAAGAGTGTGAAAGAGCTGCTGGGGATCACCATCATGGAAAGAAGCAGCTTCGAGAAGAA TCCCATCGACTTTCTGGAAGCCAAGGGCTACAAAGAAGTGAAAAAGGACCTGATCATCAAGCTGCCTAAG TACTCCCTGTTCGAGCTGGAAAACGGCCGGAAGAGAATGCTGGCCTCTGCCGGCGAACTGCAGAAGGGAA ACGAACTGGCCCTGCCCTCCAAATATGTGAACTTCCTGTACCTGGCCAGCCACTATGAGAAGCTGAAGGG CTCCCCCGAGGATAATGAGCAGAAACAGCTGTTTGTGGAACAGCACAAGCACTACCTGGACGAGATCATC GAGCAGATCAGCGAGTTCTCCAAGAGAGTGATCCTGGCCGACGCTAATCTGGACAAAGTGCTGTCCGCCT ACAACAAGCACCGGGATAAGCCCATCAGAGAGCAGGCCGAGAATATCATCCACCTGTTTACCCTGACCAA TCTGGGAGCCCCTGCCGCCTTCAAGTACTTTGACACCACCATCGACCGGAAGAGGTACACCAGCACCAAA GAGGTGCTGGACGCCACCCTGATCCACCAGAGCATCACCGGCCTGTACGAGACACGGATCGACCTGTCTC AGCTGGGAGGTGACTCTGGCGGCTCAAAAAGAACCGCCGACGGCAGCGAATTCGAGCCCAAGAAGAAGAG GAAAGTCTAACCGGTCATCATCACCATCACCATTGAGTTTAAACCCGCTGATCAGCCTCGACTGTGCCTT CTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCAC TGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGT GGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCT CTATGGCTTCTGAGGCGGAAAGAACCAGCTGGGGCTCGATACCGTCGACCTCTAGCTAGAGCTTGGCGTA ATCATGGTCATAGCTGTTTCCTGTGTGAAATTGTTATCCGCTCACAATTCCACACAACATACGAGCCGGA AGCATAAAGTGTAAAGCCTAGGGTGCCTAATGAGTGAGCTAACTCACATTAATTGCGTTGCGCTCACTGC CCGCTTTCCAGTCGGGAAACCTGTCGTGCCAGCTGCATTAATGAATCGGCCAACGCGCGGGGAGAGGCGG TTTGCGTATTGGGCGCTCTTCCGCTTCCTCGCTCACTGACTCGCTGCGCTCGGTCGTTCGGCTGCGGCGA GCGGTATCAGCTCACTCAAAGGCGGTAATACGGTTATCCACAGAATCAGGGGATAACGCAGGAAAGAACA TGTGAGCAAAAGGCCAGCAAAAGGCCAGGAACCGTAAAAAGGCCGCGTTGCTGGCGTTTTTCCATAGGCT CCGCCCCCCTGACGAGCATCACAAAAATCGACGCTCAAGTCAGAGGTGGCGAAACCCGACAGGACTATAA AGATACCAGGCGTTTCCCCCTGGAAGCTCCCTCGTGCGCTCTCCTGTTCCGACCCTGCCGCTTACCGGAT ACCTGTCCGCCTTTCTCCCTTCGGGAAGCGTGGCGCTTTCTCATAGCTCACGCTGTAGGTATCTCAGTTC GGTGTAGGTCGTTCGCTCCAAGCTGGGCTGTGTGCACGAACCCCCCGTTCAGCCCGACCGCTGCGCCTTA TCCGGTAACTATCGTCTTGAGTCCAACCCGGTAAGACACGACTTATCGCCACTGGCAGCAGCCACTGGTA ACAGGATTAGCAGAGCGAGGTATGTAGGCGGTGCTACAGAGTTCTTGAAGTGGTGGCCTAACTACGGCTA CACTAGAAGAACAGTATTTGGTATCTGCGCTCTGCTGAAGCCAGTTACCTTCGGAAAAAGAGTTGGTAGC TCTTGATCCGGCAAACAAACCACCGCTGGTAGCGGTGGTTTTTTTGTTTGCAAGCAGCAGATTACGCGCA GAAAAAAAGGATCTCAAGAAGATCCTTTGATCTTTTCTACGGGGTCTGACACTCAGTGGAACGAAAACTC ACGTTAAGGGATTTTGGTCATGAGATTATCAAAAAGGATCTTCACCTAGATCCTTTTAAATTAAAAATGA AGTTTTAAATCAATCTAAAGTATATATGAGTAAACTTGGTCTGACAGTTACCAATGCTTAATCAGTGAGG CACCTATCTCAGCGATCTGTCTATTTCGTTCATCCATAGTTGCCTGACTCCCCGTCGTGTAGATAACTAC GATACGGGAGGGCTTACCATCTGGCCCCAGTGCTGCAATGATACCGCGAGACCCACGCTCACCGGCTCCA GATTTATCAGCAATAAACCAGCCAGCCGGAAGGGCCGAGCGCAGAAGTGGTCCTGCAACTTTATCCGCCT CCATCCAGTCTATTAATTGTTGCCGGGAAGCTAGAGTAAGTAGTTCGCCAGTTAATAGTTTGCGCAACGT TGTTGCCATTGCTACAGGCATCGTGGTGTCACGCTCGTCGTTTGGTATGGCTTCATTCAGCTCCGGTTCC CAACGATCAAGGCGAGTTACATGATCCCCCATGTTGTGCAAAAAAGCGGTTAGCTCCTTCGGTCCTCCGA TCGTTGTCAGAAGTAAGTTGGCCGCAGTGTTATCACTCATGGTTATGGCAGCACTGCATAATTCTCTTAC TGTCATGCCATCCGTAAGATGCTTTTCTGTGACTGGTGAGTACTCAACCAAGTCATTCTGAGAATAGTGT ATGCGGCGACCGAGTTGCTCTTGCCCGGCGTCAATACGGGATAATACCGCGCCACATAGCAGAACTTTAA AAGTGCTCATCATTGGAAAACGTTCTTCGGGGCGAAAACTCTCAAGGATCTTACCGCTGTTGAGATCCAG TTCGATGTAACCCACTCGTGCACCCAACTGATCTTCAGCATCTTTTACTTTCACCAGCGTTTCTGGGTGA GCAAAAACAGGAAGGCAAAATGCCGCAAAAAAGGGAATAAGGGCGACACGGAAATGTTGAATACTCATAC TCTTCCTTTTTCAATATTATTGAAGCATTTATCAGGGTTATTGTCTCATGAGCGGATACATATTTGAATG TATTTAGAAAAATAAACAAATAGGGGTTCCGCGCACATTTCCCCGAAAAGTGCCACCTGACGTCGACGGA TCGGGAGATCGATCTCCCGATCCCCTAGGGTCGACTCTCAGTACAATCTGCTCTGATGCCGCATAGTTAA GCCAGTATCTGCTCCCTGCTTGTGTGTTGGAGGTCGCTGAGTAGTGCGCGAGCAAAATTTAAGCTACAAC AAGGCAAGGCTTGACCGACAATTGCATGAAGAATCTGCTTAGGGTTAGGCGTTTTGCGCTGCTTCGCGAT GTACGGGCCAGATATACGCGTTGACATTGATTATTGACTAGTTATTAATAGTAATCAATTACGGGGTCAT TAGTTCATAGCCCATATATGGAGTTCCGCGTTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCC CAACGACCCCCGCCCATTGACGTCAATAATGACGTATGTTCCCATAGTAACGCCAATAGGGACTTTCCAT TGACGTCAATGGGTGGAGTATTTACGGTAAACTGCCCACTTGGCAGTACATCAAGTGTATC

[0145] "Base editing activity" refers to the ability to chemically modify bases within a polynucleotide. In one embodiment, the first base is converted to the second base. base editing activity is cytidine deaminase activity, e.g., the activity of converting the target C·G to T·A. In another embodiment, the base editing activity is adenosine or adenine deaminase activity. In another embodiment, the base editing activity is a cytogenetic activity, for example, the activity of converting A·T to G·C. Deaminase activity, e.g., converting the target C·G to T·A, resulting in adenosine or Adenine deaminase activity, for example, converting A·T to G·C. In this embodiment, base editing activity is evaluated by editing efficiency. Base editing efficiency can be measured using any suitable method. by suitable means, e.g., Sanger sequencing or next-generation sequencing. In some embodiments, the base editing efficiency can be measured by the base editor. The percentage of all sequencing reads with nucleobase conversions resulting from, e.g., G. by the percentage of all sequencing reads with the target AT base pair converted to a C base pair. In some embodiments, base editing efficiency is measured by determining the number of bases in a population of cells. Whole cells with nucleobase changes effected by base editors when editing occurs It is measured by the percentage of

[0146] The term "base editor system" refers to a system that edits nucleic acid bases in a target nucleotide sequence. In various embodiments, the base editor system comprises: (1) (2) a polynucleotide-programmable nucleotide-binding domain (e.g., Cas9); a deaminase domain for deaminating the nucleic acid base (e.g., adenosine deaminase) and (3) one or more guide polynucleases. In some embodiments, the polynucleotide primer comprises a nucleic acid sequence (e.g., a guide RNA). The programmable nucleotide binding domain is a polynucleotide programmable DNA In some embodiments, the base editor is an adenine or In some embodiments, the base editor system is an adenosine base editor (ABE). The stem is ABE8.

[0147] In some embodiments, the base editor system comprises two or more base editing components. For example, a base editor system can include multiple deaminases. In some embodiments, the base editor system comprises one or more adenosine deoxyribonucleotides. In some embodiments, a single guide polynucleotide may be utilized. can be used to target different deaminases to a target nucleic acid sequence. In some embodiments, a single pair of guide polynucleotides is utilized to identify different determinants. The enzyme can be targeted to a target nucleic acid sequence.

[0148] Deaminase domains of base editor systems and polynucleotide programmability The nucleotide-binding moieties may be covalently or non-covalently bound to one another or by any of the bonds. The molecules may be linked by any combination of binding and interaction. In the present invention, the deaminase domain is capable of forming programmable nucleotide bonds in polynucleotides. The target nucleotide sequence can be targeted by the binding domain. In this case, the polynucleotide programmable nucleotide binding domain is a deaminator. In some embodiments, the polynucleotide may be fused or linked to a nucleotide domain. The programmable nucleotide-binding domain binds non-covalently to the deaminase domain By selectively interacting or binding to the deaminase domain, the deaminase domain is directed to the target nucleotide sequence. For example, in some embodiments, Deamina The enzyme domain is one of the polynucleotide-programmable nucleotide binding domains. interacting with, associating with, or forming a complex with a further heterologous moiety or domain that is part of the In some embodiments, the polypeptide may contain additional heterologous moieties or domains. In the form, the additional heterologous moiety binds to, interacts with, associates with, or In some embodiments, the additional heterologous moiety can form a complex with a poly It is capable of binding to, interacting with, associating with, or forming a complex with a nucleotide. In some embodiments, the additional heterologous moiety is capable of binding to the guide polynucleotide. In some embodiments, the additional heterologous moiety is attached to the polypeptide linker. In some embodiments, the additional heterologous moiety can be a polynucleotide. The additional heterologous moiety can be a protein domain or a In some embodiments, the additional heterologous moiety is a K homology (KH) domain, an MS2 domain, or a nucleotide sequence. coat protein domain, PP7 coat protein domain, SfMu Com coat protein domain In, steryl α motif, telomerase Ku binding motif and Ku protein, telomerase The motif may be a Sm7-binding motif and an Sm7 protein, or an RNA recognition motif.

[0149] The base editor system can further comprise a guide polynucleotide component. The components of the base editor system may be linked by covalent bonds, non-covalent interactions, or both. It should be understood that the compounds may be coupled to one another via any combination of bonds and interactions. In some embodiments, the deaminase domain is selected from the group consisting of: It can be targeted to a target nucleotide sequence. For example, in some embodiments The deaminase domain may be a portion or segment of the guide polynucleotide (e.g., a polynucleotide Further variants that can interact, bind, or form complexes with the nucleotide motif A species moiety or domain (e.g., a polynucleotide binding protein, such as an RNA or DNA binding protein) In some embodiments, additional heterologous moieties or domains ( Polynucleotide-binding domains, e.g., RNA- or DNA-binding proteins, are deaminants of In some embodiments, the additional heterologous moiety may be fused or linked to the enzyme domain. The polypeptide binds to, interacts with, associates with, or forms a complex with the polypeptide. In some embodiments, the additional heterologous moiety can be a polynucleotide. capable of binding to, interacting with, associating with, or forming a complex with a polynucleotide In some embodiments, the additional heterologous moiety is attached to the guide polynucleotide. In some embodiments, the additional heterologous moiety can be a polypeptide linker. In some embodiments, the additional heterologous moiety can be linked to a polynucleotide. The additional heterologous moiety can be a protein domain or a nucleotide linker. In some embodiments, the additional heterologous moiety may be a K homology (KH) domain, MS2 Coat protein domain, PP7 coat protein domain, SfMu Com coat protein domain, sterile alpha motif, telomerase Ku binding motif and Ku protein, Te It may be a chromosome enzyme Sm7 binding motif and Sm7 protein, or an RNA recognition motif.

[0150] In one embodiment, the base editor system comprises an inhibitor of base excision repair (BER). The components of the base editor system can further include a shared Covalent bonds, non-covalent interactions, or any combination of these bonds and interactions. It should be understood that the BER components can be linked to each other via a BER inhibitor. In one embodiment, the inhibitor of BER is uracil DNA glycosylase In one embodiment, the inhibitor of BER can be an inosine BER inhibitor (UGI). In one embodiment, the inhibitor of BER is a polynucleotide protease inhibitor. Targeting to target nucleotide sequences via a rammable nucleotide-binding domain In some embodiments, the polynucleotide may be a programmable nucleotide. The binding domain may be fused or linked to an inhibitor of BER. The programmable nucleotide-binding domain is a deaminase domain. In some embodiments, the polynucleotides may be fused or linked to an inhibitor of BER. Nucleotide-programmable nucleotide-binding domains are not covalent inhibitors of BER. BER inhibitors by covalently interacting with or associating with BER inhibitors. The molecule can be targeted to a target nucleotide sequence. In embodiments, the inhibitor of the BER component is a polynucleotide programmable nucleotide. interacting with, associating with, or having associated with further heterologous moieties or domains that are part of the octide-binding domain or may contain further heterologous moieties or domains capable of complexing.

[0151] In one embodiment, the inhibitor of BER is a nucleotide sequence that is mediated by a guide polynucleotide to target nucleotides. For example, in some embodiments, the inhibitor may be targeted to a nucleotide sequence that inhibits BER. The deleterious agent may be a portion or segment of a guide polynucleotide (e.g., a polynucleotide Further heterologous moieties or domains (e.g., nucleotides) that can interact with, associate with, or complex with the nucleotide motif (e.g., nucleotides) For example, a polynucleotide binding domain such as an RNA or DNA binding protein. In some embodiments, the guide polynucleotide may further comprise a heterologous moiety or a The main (e.g., polynucleotide-binding domain, such as an RNA- or DNA-binding protein) In some embodiments, the additional heterologous nucleotide may be fused or linked to an inhibitor of BER. The moiety is capable of binding, interacting, associating, or complexing with a polynucleotide In some embodiments, the additional heterologous moiety is attached to the guide polynucleotide. In some embodiments, the additional heterologous moiety can be a polypeptide linker. In some embodiments, the additional heterologous moiety can be linked to a polynucleotide. The additional heterologous moiety can be a protein domain or a nucleotide linker. In some embodiments, the additional heterologous moiety may be a K homology (KH) domain, MS2 Coat protein domain, PP7 coat protein domain, SfMu Com coat protein domain, sterile alpha motif, telomerase Ku binding motif and Ku protein, Te It may be a chromosome enzyme Sm7 binding motif and Sm7 protein, or an RNA recognition motif.

[0152] The term "Cas9" or "Cas9 domain" refers to a Cas9 protein or a fragment thereof (e.g., Cas9 Active, inactive, or partially active DNA cleavage domains of Cas9 and / or gRNA binding of Cas9 Cas9 nuclease refers to an RNA-guided nuclease containing a c domain. asnl nuclease or CRISPR (clustered regularly interspaced short palindromic CRISPR is also known as a repeat-associated nuclease. The adaptive immune system provides defense against transposable elements (transposable elements, conjugative plasmids). CRIS comprises a spacer, a sequence complementary to the preceding mobile element, and a target invader nucleic acid. The PR cluster is transcribed and processed into CRISPR RNA (crRNA). In the endothelial cell line, the correct processing of pre-crRNA is mediated by a small trans-coding RNA (tracrRNA), an endogenous RNA. tracrRNA requires ribonuclease 3 (rnc) and Cas9 protein. This guides the processing of pre-crRNA by Cas9 / crRNA / tracrRNA. A endonucleolytically cleaves a linear or circular dsDNA target complementary to the spacer. The target strand that is not complementary to the crRNA is first endonucleolytically cleaved and then exonucleolytically cleaved. In nature, DNA binding and cleavage are carried out by proteins. However, both crRNA and tracrRNA aspects are typically required. A single guide RNA ("sgRNA," or simply "gRNA") is engineered to incorporate into a single RNA species. See, for example, Jinek M. et al., Science 337:816-821 (2012). Cas9 is a CRISPR repeat sequence. It recognizes a short motif (PAM or protospacer adjacent motif) in the spacer to distinguish self from non-self. The sequence and structure of Cas9 nuclease are well known to those skilled in the art. (e.g., "Complete genome sequence of an M1 strain of Streptococcus pyog" Ferretti et al., Proc. Natl. Acad. Sci. USA 98:4658-4663(2001); SPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltc heva E. et al., Nature 471:602-607(2011); and “A programmable dual-RNA-guide d DNA endonuclease in adaptive bacterial immunity.” Jinek M. et al., Science 33 7:816-821 (2012), the entire contents of which are incorporated herein by reference. Orthologs include, but are not limited to, S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences are described in the present disclosure. Such Cas9 nucleases and sequences will be apparent to those skilled in the art based on the disclosures of Chylinsk i, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas Organisms and genes disclosed in “immunity systems” (2013) RNA Biology 10:5, 726-737 The Cas9 sequences from the locus are included, the entire contents of which are incorporated herein by reference.

[0153] An exemplary Cas9 is Streptococcus pyogenes Cas9 (spCas9), the amino acid sequence of which is shown below. The following information is provided. JPEG2025032080000003.jpg161169 (single underline: HNH domain; double underline: RuvC domain)

[0154] The nuclease-inactivated Cas9 protein is interchangeably referred to as the “dCas9” protein (nuclease-“ It may also be referred to as "dead" Cas9 or catalytically inactive Cas9. Methods for generating a Cas9 protein (or a fragment thereof) having the following structure are known (see, e.g., Jinek et al. al, Science. 337:816-821(2012); Qi et al, “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression”(2013) Cell. 28; 152 (5): 1173-83, the contents of each of which are incorporated herein by reference. For example, the DN of Cas9 The A cleavage domain is composed of two subdomains: the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA and binds to RuvC1. The subdomains cleave the non-complementary strand. Mutations within these subdomains allow Cas9 to For example, mutations D10A and H840A can inhibit the nuclease activity of S. pyogenes Cas9. completely inactivates the ATPase activity (Jinek et al., Science. 337:816-821(2012); Qi et al. l, Cell. 28;152(5): 1173-83 (2013)). In some embodiments, the Cas9 nuclease is inactivated. The Cas9 has an active (e.g., inactivated) DNA cleavage domain, i.e., the Cas9 is referred to as an "nCas9" protein. The nickase is called Cas9 (meaning "nickase"). In some embodiments, Cas9 For example, in some embodiments, proteins comprising fragments of the protein The gene contains one of two Cas9 domains: (1) the gRNA-binding domain of Cas9; (2) the C In some embodiments, a protein comprising Cas9 or a fragment thereof is Cas9 variants are those that share homology with Cas9 or its fragments. For example, a Cas9 variant may have at least about 70% identity to wild-type Cas9, at least about 80% identity to wild-type Cas9, or at least about 80% identity to wild-type Cas9. 0% identity, at least about 90% identity, at least about 95% identity, at least about 96% Identity, at least about 97% identity, at least about 98% identity, at least about 99% identity have at least about 99.5% identity, or at least about 99.9% identity. In some embodiments, the Cas9 mutant has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 2 9, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 4 In some embodiments, the Cas9 vector may have 9, 50, or more amino acid changes. The variant contains a fragment of Cas9 (e.g., the gRNA binding domain or the DNA cleavage domain) and The fragments have at least about 70% identity to the corresponding fragment of wild-type Cas9, and at least about 80% identity to the corresponding fragment of wild-type Cas9. has at least about 90% identity, has at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity have at least about 99% identity, have at least about 99.5% identity, and In some embodiments, the fragment has at least about 99.9% identity to the corresponding wild-type Cas At least 30%, at least 35%, at least 40%, at least 45%, at least 9 amino acids in length at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least At least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, is 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% .

[0155] In some embodiments, the fragment is at least 100 amino acids in length. In the above, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 60 0, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or It is at least 1300 amino acids in length.

[0156] In one embodiment, the wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes. (NCBI reference sequence: NC_017053.1, nucleotide and amino acid sequences are as follows): ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGATTATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG GCAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGCAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAATCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTAGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAGAAATGGCTTGTTTGGGAAT CTCATTGCTTTGTCATTGGGATTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATAGTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAGCGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAGGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGCGCCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGGGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGATATTCAAAAAGCACAGGTGTCTGGACAAGGCCATAGTTTA CATGAACAGATTGCTAACTTAGCTGGCAGTCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAATTGTTGATGAACT GGTCAAAGTAATGGGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCCAGA AAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCTGTT GAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTACAAAATGGAAGAGACATGTATGTGGACCAAGAATT AGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCATTAAAGACGATTCAATAGACAATA AGGTACTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAAAAC TATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGGTTT GAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAATTT TGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCTAAA TTAGTTTCTGACTTCCGAAAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGCGTA TCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATAAAG TTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCTAAT ATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGGGGA AACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATATTG TCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTACCAAAAAGAAATTCGGACAAGCTTATTGCT CGTAAAAAGACTGGGATCCAAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGCTAA GGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCTTTG AAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATAAGGAAGTTAAAAGACTTAATCATTAAACTACCTAATAT AGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAATTACAAAAAGGAAATGAGCTGGCTCT GCCAAGCAAATGTGAATTTTTTATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAACAAA AACAATTGTTTGTGGAGCAGCATAAGCATTATTTGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTTATT TTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGAAAA TATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTAAAC GATATACGTCTACAAAAGAAGTTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATTGAT TTGAGTCAGCTAGGAGGTGACTGA JPEG2025032080000004.jpg165169 (single underline: HNH domain; double underline: RuvC domain)

[0157] In one embodiment, the wild-type Cas9 has the following nucleotide and / or amino acid sequence: Corresponding to or containing columns: ATGGATAAAAAGTATTCTATTGGTTTAGACATCGGCACTAATTCCGTTGGATGGGCTGTCATAACCGATGAATACAAAGT ACCTTCAAAGAAATTTAAGGTGTTGGGGAACACAGACCGTCATTCGATTAAAAAAGAATCTTATCGGTGCCCTCCTATTCG ATAGTGGCGAAACGGCAGAGGCGACTCGCCTGAAACGAACCGCTCGGAGAAGGTATACACGTCGCAAGAACCGAATATGT TACTTACAAGAAATTTTTAGCAATGAGATGGCCAAAGTTGACGATTCTTTCTTTCACCGTTTGGAAGAGTCCTTCCTTGT CGAAGAGGACAAGAAACATGAACGGCACCCCATCTTTGGAAACATAGTAGATGAGGTGGCATATCATGAAAAGTACCCAA CGATTTATCACCTCAGAAAAAAGCTAGTTGACTCAACTGATAAAGCGGACCTGAGGTTAATCTACTTGGCTCTTGCCCAT ATGATAAAGTTCCGTGGGCACTTTCTCATTGAGGGTGATCTAAATCCGGACAACTCGGATGTCGACAAACTGTTCATCCA GTTAGTACAAACCTATAATCAGTTGTTTGAAGAACCCTATAAATGCAAGTGGCGTGGATGCGAAGGCTATTCTTAGCG CCCGCCTCTCTAAATCCCGACGGCTAGAAAAACCTGATCGCACAATTACCCGGAGAGAAAAAAATGGTTGTTCGGTAAC CTTATAGCGCTCTCACTAGGCCTGACACCAAATTTAAGTCGAACTTCGACTTAGCTGAAGATGCCAAATTGCAGCTTAG TAAGGACACGTACGATGACGATCTCGCAATCTACTGGCACAAATTGGAGATCAGTATGCGGACTTATTTTTGGCTGCCA AAAACCTTAGCGATGCAATCCTCCTATCTGACATACTGAGAGTTAATACTGAGATTACCAAGGCGCCGTTATCCCGCTTCA ATGATCAAAAGGTACGATGAACATCACCAAGACTTGACACTTCTCAAGGCCCTAGTCCGTCAGCAACTGCCTGAGAAATA TAAGGAAATATTCTTTGATCAGTCGAAAACGGGTACGCAGGTTATATTGACGGGCGGAGCGAGTCAAGAGGAATTCTACA AGTTTATCAAACCCATATTAGAGAATGGATGGGACGGAAGAGTTGCTTGTAAAACTCAATCGCGAAGATCTACTGCGA AAGCAGCGGACTTTCGACAACGGTAGCATTCCACATCAAATCCACTTAGGCGAATTGCATGCTATACTTAGAAGGCAGGA GGATTTTTATCCGTTCCTCAAAGACAATCGTGAAAAGATTGAGAAAATCCTAACCTTTCGCATACCTACTATGTGGGAC CCCTGGCCCGAGGGAACTCTCGGTTCGCATGGATGACAAGAAAGTCCGAAGAAACGATTACTCCATGGAATTTTGAGGAA GTTGTCGATAAAGGTGCGTCAGCTCAATCGTTCATCGAGAGGATGACCAACTTTGACAAGAATTTACCGAACGAAAAAGT ATTGCCTAAGCACAGTTTACTTTACGAGTATTTCACAGTGTACAATGAACTCACGAAAGTTAAGTATGTCACTGAGGGCA TGCGTAAACCCGCCTTTCTAAGCGGAGAACAGAAGAAAGCAATAGTAGATCTGTTATTCAAGACCAACCGCAAAGTGACA GTTAAGCAATTGAAAGAGGACTACTTTAAGAAAATTGAATGCTTCGATTCTGTCGAGATCTCCGGGGTAGAAGATCGATT TAATGCGTCACTTGGTACGTATCATGACCTCCTAAAGATAATTAAAGATAAGGACTTCCTGGATAACGAAGAGAATGAAG ATATCTTAGAAGATATAGTGTTGACTCTTACCCTCTTTGAAGATCGGGAAATGATTGAGGAAAGACTAAAAACATACGCT CACCTGTTCGACGATAAGGTTATGAAACAGTTAAAGAGGCGTCGCTATACGGGCTGGGGACGATTGTCGCGGAAACTTAT CAACGGGATAAGAGACAAGCAAAGTGGTAAAACTATTCTCGATTTTCTAAAGAGCGACGGCTTCGCCATAGGAACTTTA TGCAGCTGATCCATGATGACTCTTTAACCTTCAAAGAGGATATACAAAAGGCACAGGTTTCCGGACAAGGGGACTCATTG CACGAACATATTGCGAATCTTGCTGGTTCGCCAGCCATCAAAAAGGGCATACTCCAGACAGTCAAAGTAGTGGATGAGCT AGTTAAGGTCATGGGACGTCACAAACCGGAAAACATTGTAATCGAGATGGCACGCGAAAATCAAACGACTCAGAAGGGGC AAAAAAACAGTCGAGAGCGGATGAAGAGAATAGAAGAGGGTATTAAAGAACTGGGCAGCCAGATCTTAAAGGAGCATCCT GTGGAAAATACCCAATTGCAGAACGAGAAACTTTACCTCTATTACCTACAAAATGGAAGGGACATGTATGTTGATCAGGA ACTGGACATAAACCGTTTATCTGATTACGACGTCGATCACATTGTACCCCAATCCTTTTTGAAGGACGATTCAATCGACA ATAAAGTGCTTACACGCTCGGATAAGAACCGAGGGAAAAGTGACAATGTTCCAAGCGAGGAAGTCGTAAAGAAAATGAAG AACTATTGGCGGCAGCTCCTAAATGCGAAACTGATAACGCAAAGAAAGTTCGATAACTTAACTAAAGCTGAGAGGGGTGG CTTGTCTGAACTTGACAAGGCCGGATTTATTAAACGTCAGCTCGTGGAAACCCGCCAAATCACAAAGCATGTTGCACAGA TACTAGATTCCCGAATGAATACGAAATACGACGAGAACGATAAGCTGATTCGGGAAGTCAAAGTAATCACTTTAAAGTCA AAATTGGTGTCGGACTTCAGAAAGGATTTTCAATTCTATAAAGTTAGGGAGATAAATAACTACCACCATGCGCACGACGC TTATCTTAATGCCGTCGTAGGGACCGCACTCATTAAGAAATACCCGAAGCTAGAAAGTGAGTTTGTGTATGGTGATTACA AAGTTTATGACGTCCGTAAGATGATCGCGAAAGCGAACAGGAGATAGGCAAGGCTACAGCCAAATACTTCTTTTATTCT AACATTATGAATTTCTTTAAGACGGAAATCACTCTGGCAAACGGAGAGATACGCAAACGACCTTTAATTGAAACCAATGG GGAGACAGGTGAAATCGTATGGGATAAGGGCCGGGACTTCGCGACGGTGAGAAAAGTTTTGTCCATGCCCCAAGTCAACA TAGTAAAGAAAACTGAGGTGCAGACCGGAGGGTTTTCAAAGGAATCGATTCTTCCAAAAAGGAATAGTGATAAGCTCATC GCTCGTAAAAAGGACTGGGACCCGAAAAAGTACGGTGGCTTCGATAGCCCTACAGTTGCCTATTCTGTCCTAGTAGTGGC AAAAGTTGAGAAGGGAAATCCAAGAAACTGAAGTCAGTCAAAGAATTATTGGGGATAACGATTATGGAGCGCTCGTCTT TTGAAAAGACCCCATCGACTTCCTTGAGGCGAAAGGTTACAAGGAAGTAAAAAAGGATCTCATAATTAAACTACCAAAG TATAGTCTGTTTGAGTTAGAAATGGCCGAAAACGGATGTTGGCTAGCGCCGGAGAGCTTCCAAAGGGGAACGAACTCGC ACTACCGTCTAAATACGTGAATTTCCTGTATTTAGCGTCCCATTACGAGAAGTTGAAAGGTTCACCTGAAGATAACGAAC AGAAGCAACTTTTTGTTGAGCAGCACAAACATTATCTCGACGAAATCATAGAGCAAATTTCGGAATTCAGTAAGAGAGTC ATCCTAGCTGATGCCAATCTGGACAAAGTATTAAGCGCATACAACAAGCACAGGGATAAACCCATACGTGAGCAGGCGGA AAATATTATCCATTTGTTTACTCTTACCAACCTCGGCGCTCCAGCCGCATTCAAGTATTTTGACACAACGATAGATCGCA AACGATACACTTCTACCAAGGAGGTGCTAGACGCGACACTGATTCACCAATCCATCACGGGATTATATGAAACTCGGATA GATTTGTCACAGCTTGGGGGTGACGGATCCCCAAGAAGAAGAGGAAAGTCTCGAGCGACTACAAAGACCATGACGGTGA TTATAAAGATCATGACATCGATTACAAGGATGACGATGACAAGGCTGCAGGA JPEG2025032080000005.jpg163169 (single underline: HNH domain; double underline: RuvC domain)

[0158] In some embodiments, the wild-type Cas9 is Cas9 from Streptococcus pyogenes (NC BI reference sequence: NC_002737.2 (nucleotide sequence below) and Uniprot reference sequence: Q99ZW2 (nucleotide sequence below) The amino acid sequence corresponds to the following: ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGAATATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG ACAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGTAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAACCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTGGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAAAAATGGCTTATTTGGGAAT CTCATTGCTTTGTCATTGGGTTTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATACTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAACGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAAGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGTACCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGAGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGACATTCAAAAAGCACAAGTGTCTGGACAAGGCGATAGTTTA CATGAACATATTGCAAATTTAGCTGGTAGCCCTGCTATTAAAAAGGTATTTTACAGACTGTAAAAGTTGTTGATGAATT GGTCAAAGTAATGGGGCGGCATAAGCCAGAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCC AGAAAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAAGAGCATCCT GTTGAAAATACTCAAATTGCAAATGAAAAGCTCTATCTCTATTATCTCCAAAATGGAAGAGACATGTATGTGGACCAAGA ATTAGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCCTTAAAGACGATTCAATAGACA ATAAGGTCTTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAA AACTATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGG TTTGAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAA TTTTGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCT AATTAGTTTCTGACTTCCGAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGC GTATCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATA AAGTTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCT AATATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGG GGAAACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATA TTGTCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTTACCAAAAGAAATTCGGACAAGCTTATT GCTCGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGC TAAGGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCT TTGAAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAAAGACTTAATCATTAAACTACCTAAA TATAGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAAAAGGAAATGAGCTGGC TCTGCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAAC AAAAACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTT ATTTTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGA AAATATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTA AACGATATACGTCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATT GATTTGAGTCAGCTAGGAGGTGACTGA JPEG2025032080000006.jpg164168(SEQ ID NO:1) (Single underline: HNH domain; double underline: RuvC domain)

[0159] In one embodiment, Cas9 is isolated from Corynebacterium ulcerans (NCBI Refs: NC_015683.1 , NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1) ; Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Streptococcu s iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psyc hroflexus torquisI (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref : YP_820832.1), Listeria innocua (NCBI Ref: NP_472073.1), Campylobacter jejuni ( NCBI Ref: YP_002344900.1) or Neisseria meningitidis (NCBI Ref: YP_002342100 .1), or Cas9 from any other organism.

[0160] In some embodiments, dCas9 contains one or more nucleotides that inactivate Cas9 nuclease activity. It corresponds to, or contains part or all of, a mutated Cas9 amino acid sequence. For example, in some embodiments, the dCas9 domain is D10 according to the numbering in SEQ ID NO: 1. In some embodiments, the Cas9 includes the A and H840A mutations or corresponding mutations in another Cas9. wherein dCas9 comprises the amino acid sequence of dCas9 (D10A and H840A): JPEG2025032080000007.jpg163168 (single underline: HNH domain; double underline: RuvC domain)

[0161] In some embodiments, the Cas9 domain comprises a D10A mutation, while the the residue at position 840 in the amino acid sequence provided herein, or The residue at the corresponding position in either sequence remains a histidine.

[0162] In other embodiments, D10A, e.g., resulting in nuclease-inactivated Cas9 (dCas9), and dCas9 variants having mutations other than H840A. For example, other amino acid substitutions at D10 and H840, or the nuclease domain of Cas9 may be used. Other substitutions within the domain (e.g., HNH nuclease subdomain and / or RuvC1 subdomain) In some embodiments, a variant or homolog of dCas9 is at least about 70% identity, at least about 80% identity, at least about 90% identity , at least about 95% identity, at least about 98% identity, at least about 99% identity, Those having at least about 99.5% identity, or at least about 99.9% identity are provided. In some embodiments, the amino acid sequence is about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, or Amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids Variants of dCas9 having amino acid sequences shorter or longer than 1000 or 10000 are provided. do.

[0163] In some embodiments, the Cas9 fusion proteins provided herein comprise a Cas9 protein. The full-length amino acid sequence of the protein, for example, one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein contain a full-length Cas9 sequence. Examples of Suitable Cas9 Domains and Cas9 Fragments Suitable amino acid sequences are provided herein, and further suitable sequences for Cas9 domains and fragments are , will be apparent to those skilled in the art.

[0164] Additional Cas9 proteins (e.g., nuclease-dead Cas9 (dCas9), Cas9 nCas9, or nuclease-active Cas9, its variants and homologs It should be understood that within the scope of this disclosure are any and all Cas9 proteins, including: In some embodiments, Cas9 includes, but is not limited to, those provided below. The protein is a nuclease-inactive Cas9 (dCas9). In some embodiments, the Cas9 protein In some embodiments, the Cas9 protein is a Cas9 nickase (nCas9). The quality is nuclease-active Cas9.

[0165] Exemplary catalytically inactive Cas9 (dCas9): DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKLVDSTDKADLRIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD

[0166] Exemplary catalyst of Cas9ニッカーゼ (nCas9): DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD

[0167] Exemplary catalyst activityCas9: DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPCKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD.

[0168] In one embodiment, Cas9 is directed against archaeal organisms that comprise the domain and kingdom of unicellular prokaryotic microorganisms. In one embodiment, the Cas9 protein is derived from, for example, a fungus (e.g., nanoarchaea). , Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res . 2017 Feb 21. doi: 10.1038 / cr.2017.21 refers to CasX or CasY, and The entire contents of which are incorporated herein by reference. Many CRISPR-Cas systems, including Cas9, which was first reported in the archaeal domain of life, have been identified. This branched Cas9 protein was identified in the little-studied nanoarchaea. , discovered as part of an active CRISPR-Cas system, previously unknown in bacteria. Two systems, CRISPR-CasX and CRISPR-CasY, have been discovered that are the most efficient known to date. In some embodiments, Cas9 is a CasX or CasX-like protein. In some embodiments, Cas9 represents a variant of CasY or a variant of CasY. It represents nucleic acid programmable DNA binding protein (napDNAbp). Other RNA-guided DNA binding proteins may also be used, such as RNA-guided DNA binding proteins (RNA-guided DNA binding proteins), as disclosed herein. It should be understood that the range is

[0169] In certain embodiments, napDNAbp useful in the methods of the invention contain circular substitutions. , which is known and described, for example, in Oakes et al., Cell 176, 254-267, 2019 Below are exemplary circular permutations, where bolded sequences indicate Cas9-derived sequences and italicized sequences: The sequence represents the linker sequence and the underlined sequence represents the bipartite nuclear localization sequence. JPEG2025032080000008.jpg189168

[0170] Polynucleotide programmable nucleotides that can be incorporated into base editors Non-limiting examples of binding domains include domains derived from CRISPR proteins, restriction nucleases, and the like. Enzymes, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases Examples include ZFNs.

[0171] In some embodiments, the nucleic acid of any of the fusion proteins provided herein The programmable DNA binding protein (napDNAbp) can be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the napDNAbp is a CasY protein. is at least 85%, at least 90%, or At least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% In some embodiments, the napDNAbp comprises a naturally occurring amino acid sequence having identity to the napDNAbp. In some embodiments, the napDNAbp is a CasX or CasY protein present. At least 85%, at least at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least Contains amino acid sequences with at least 99.5% identity to Cas12b / C2c1 and CasX from other bacterial species. It should be understood that CasY and CasC may also be used in accordance with the present disclosure.

[0172] Cas12b / C2c1 (uniprot.org / uniprot / T0D7A2#2) sp|T0D7A2|C2C1_ALIAG CRISPR-associated endonuclease C2c1 OS = Alicyclobacillus acido- terrestris (strain ATCC 49025 / DSM 3922 / CIP 106132 / NCIMB 13137 / GD3B) GN=c2c1 PE=1 SV=1 MAVKSIKVKLRLDDMPEIRAGLWKLHKEVNAGVRYYTEWLSLLRQENLYRRSPNGDGEQECDKTAEECKAELLERLRARQ VENGHRGPAGSDDELLQLARQLYELLVPQAIGAKGDAQQIARKFLSPLADKDAVGGLGIAKAAGNKPRWVRMREAGEPGWE EEKEKAETRKSADRTADVLRALADFGLKPLMRVYTDSEMSSVEWKPLRKGQAVRTWDRDMFQQAIERMMSWESWNQRVGQ EYAKLVEQKNRFEQKNFVGQEHLVHLVNQLQQDMKEASPGLESKEQTAHYVTGRALRGSDKVFEKWGKLAPDAPFDLYDA EIKNVQRRNTRRFGSHDLFAKLAEPEYQALWREDASFLTRYAVYNSILRKLNHAKMFATTFTLPDATAHPIWTRFDKLGGN LHQYTFLFNEFGERRHAIRFHKLLKVENGVAREVDDVTVPISMSEQLDNLLPRDPNEPIALYFRDYGAEQHFTGEFGGAK IQCRRDQLAHMHRRRGARDVYLNVSVRVQSQSEARGERRPPYAAVFRLVGDNHRAFVHFDKLSDYLAEHPDDGKLGSEGL LSGLRVMSVDLGLRTSASISVFRVARKDELKPNSKGRVPFFFPIKGNDNLVAVHERSQLLKLPGETESKDLRAIREERQR TLRQLRTQLAYLRLLVRCGSEDVGRRERSWAKLIEQPVDAANHMTPDWREAFENELQKLKSLHGICSDKEWMDAVYESVR RVWRHMGKQVRDWRKDVRSGERPKIRGYAKDVVGGNSIEQIEYLERQYKFLKSWSFFGKVSGQVIRAEKGSRFAITLREH IDHAKEDRLKKLADRIIMEALGYVYALDERGKGKWVAKYPPCQLILLEELSEYQFNNDRPPSENNQLMQWSHRGVFQELI NQAQVHDLLVGTMYAAFSSRFDARTGAPGIRCRRVPARCTQEHNPEPFPWWLNKFVVEHTLDACPLRADDLIPTGEGEIF VSPFSAEEGDFHQIHADLNAAQNLQQRLWSDFDISQIRLRCDWGEVDGELVLIPRLTGKRTADSYSNKVFYTNTGVTYYE RERGKKRRKVFAQEKLSEEEAELLVEADEAREKSVVLMRDPSGIINRGNWTRQKEFWSMV NQRIEGYLVKQIRSRVPLQ DSACENTGDI

[0173] CasX (uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) >tr|F0NN87|F0NN87_SULIH CRISPR-associated Casx protein OS = Icelandic Sulfolobus s (strain HVE10 / 4) GN = SiH_0402 PE=4 SV=1 MEVPLYNIFGDNYIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAA AGWVLTRLGKAKVSEGDYVGVNVFTRPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVSSVTNPNVSVVRIYTISDAV GQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSSNMRERYIVLANIIYEYLTG SKRLEDLLYFANRDLIMNL NSDDGKVRDLKLISAIVNGELIRGEG

[0174] >tr|F0NH53|F0NH53_SULIR CRISPR associated protein, Casx OS = Sulfolobus islandic us (strain REY15A) GN=SiRe_0771 PE=4 SV=1 MEVPLYNIFGDNYIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAA AGWVLTRLGKAKVSEGDYVGVNVFTRPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVSSVTNPNVSVVSIYTISDAV GQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLN SDDGKVRDLKLISAYVNGELIRGEG

[0175] Deltaproteobacteria CasX MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMDTDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAIL QVYWQEFKDDHVGLMCKFAQPASKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFWYKLEQVSEKGKAITNYFGRCNVA EHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESHTPVKPLAQIAGNRYASGPVGKALSDACMGTIASFL SKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDAYNEVIARVRMWVNLNLWQKLKLSRDDAK PLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENPK KPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEMD EKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFTD GTDIKKSGKWQGLLYGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKIG RDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQA AKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAKL AYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKELS AELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSYK SGKQPFVGAWQAFYKRRLKEVWKPNA

[0176] CasY (ncbi.nlm.nih.gov / protein / APG80656.1) >APG80656.1 CRISPR-associated protein CasY [uncultured Parcubacteria group bacte rium] MSKRHPRISGVKGYRLHAQRLEYTGKSGAMRTIKYPLYSSPSGGRTVPREIVSAINDDYVGLYGLSNFDDLYNAEKRNEE KVYSVLDFWYDCVQYGAVFSYTAPGLLKNVAEVRGGSYELTKTLKGSHLYDELQIDKVIKFLNKKEISRANGSLDKLKKD IIDCFKAEYRERHKDQCNKLADDIKNAKKDAGASLGERQKKLFRDFFGISEQSENDKPSFTNPLNLTCCLLPFDTVNNNR NRGEVLFNKLKEYAQKLDKNEGSLEMWEYIGIGNSGTAFSNFLGEGFLGRLRENKITELKKAMMDITDAWRGQEQEEELE KRLRILAALTIKLREPKFDNHWGGYRSDINGKLSSWLQNYINQTVKIKEDLKGHKKDLKKAKEMINRFGESDTKEEAVVS SLLESIEKIVPDDSADDEKPDIPAIAIYRRFLSDGRLTLNRFVQREDVQEALIKERLEAEKKKKPKKRKKKSDAEDEKET IDFKELFPHLAKPLKLVPNFYGDSKRELYKKYKNAAIYTDALWKAVEKIYKSAFSSSLKNSFFDTDFDKDFFIKRLQKIF SVYRRFNTDKWKPIVKNSFAPYCDIVSLAENEVLYKPKQSRSRKSAAIDKNRVRLPSTENIAKAGIALARELSVAGFDWK DLLKKEEHEEYIDLIELHKTALALLLAVTETQLDISALDFVENGTVKDFMKTRDGNLVLEGRFLEMFSQSIVFSELRGLA GLMSRKEFITRSAIQTMNGKQAELLYIPHEFQSAKITTPKEMSRAFLDLAPAEFATSLEPESLSEKSLLKLKQMRYYPHY FGYELTRTGQGIDGGVAENALRLEKSPVKKREIKCKQYKTLGRGQNKIVLYVRSSYYQTQFLEWFLHRPKNVQTDVAVSG SFLIDEKKVKTRWNYDALTVALEPVSGSERVFVSQPFTIFPEKSAEEEGQRYLGIDIGEYGIAYTALEITGDSAKILDQN FISDPQLKTLREEVKGLKLDQRRGTFAMPSTKIARIRESLVHSLRRNRIHHLALKHKAKIVYELEVSRFEEGKQKIKKVYA TLKKADVYSEIDADKNLQTTVWGKLAVASEISASYTSQFCGACKKLWRAEMQVDETITTQELIGTVRVIKGGTLIDAIKD FMRPPIFDENDTPFPKYRDFCDKHHISKKMRGNSCLFICPFCRANADADIQASQTIALLRYVKEEKKVEDYFERFRKLKN IKVLGQMKKI

[0177] The term "Cas12" or "Cas12 domain" refers to the Cas12 protein or a fragment thereof (e.g., For example, an active, inactive, or partially active DNA cleavage domain of Cas12 and / or Cas Cas12 refers to an RNA-guided nuclease containing a gene encoding ... Cas12 nuclease belongs to the type V CRISPR / Cas system. d) regularly interspaced short palindromic repeat (REG)-associated nucleases An exemplary Bacillus hisashii Cas 12b (BhCas12b) Cas 12 domain sequence is shown below: Provided to. MAPKKKRKVGIHGVPAAATRSFILKIEPNEEVKKGLWKTHEVLNHGIAYYMNILKLIRQEAIYEHHEQDPKNPKKVSKAE IQAELWDFVLKMQKCNSFTHEVDKDEVFNILRELYEELVPSSVEKKGEANQLSNKFLYPLVDPNSQSGKGTASSGRKPRW YNLKIAGDPSWEEEKKKWEEDKKKDPLAKILGKLAEYGLIPLFIPYTDSNEPIVKEIKWMEKSRNQSVRRLDKDMFIQAL ERFLSWESWNLKVKEEYEKVEKEYKTLEERIKEDIQALKALEQYEKERQEQLLRDTLNTNEYRLSKRGLRGWREIIQKWL KMDENEPSEKYLEVFKDYQRKHPREAGDYSVYEFLSKKENHFIWRNHPEYPYLYATFCEIDKKKKDAKQQATFTLADPIN HPLWVRFEERSGSNLNKYRILTEQLHTEKLKKKLTVQLDRLIYPTESGGWEEKGKVDIVLLPSRQFYNQIFLDIEEKGKH AFTYKDESIKFPLKGTLGGARVQFDRDHLRRYPHKVESGNVGRIYFNMTVNIEPTESPVSKSLKIHRDDFPKVVNFKPKE LTEWIKDSKGKKLKSGIESLEIGLRVMSIDLGQRQAAAASIFEVVDQKPDIEGKLFFPIKGTELYAVHRASFNIKLPGET LVKSREVLRKAREDNLKLMNQKLNFLRNVLHFQQFEDITEREKRVTKWISRQENSDVPLVYQDELIQIRELMYKPYKDWV AFLKQLHKRLEVEIGKEVKHWRKSLSDGRKGLYGISLKNIDEIDRTRKFLLRWSLRPTEPGEVRRLEPGQRFAIDQLNHL NALKEDRLKKMANTIIMHALGYCYDVRKKKWQAKNPACQIILFEDLSNYNPYEERSRFENSKLMKWSRREIPRQVALQGE IYGLQVGEVGAQFSSRFHAKTGSPGIRCSVVTKEKLQDNRFFKNLQREGRLTLDKIAVLKEGDLYPDKGGEKFISLSKDR KCVTTHADINAAQNLQKRFWTRTHGFYKVYCKAYQVDGQTVYIPESKDQKQKIIEEFGEGYFILKDGVYEWVNAGKLKIK KGSSKQSSSELVDSDILKDSFDLASELKGEKLMLYRDPSGNVFPSDKWMAAGVFFGKLERILISKLTNQYSISTIEDDSS KQSMKRPAATKKAGQAKKKK.

[0178] An amino acid sequence having at least 85% identity to the BhCas12b amino acid sequence is also included. Also useful in the methods of the present invention.

[0179] "Cytidine deaminase" is a enzyme that converts amino groups into carbonyl groups through a deamination reaction. In one embodiment, the term "catalyzed polypeptide" refers to a polypeptide or fragment thereof that is capable of catalyzing , cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine PmCDA1 (Petromyzon marinus cytosine deaminase) from Petromyzon marinus 1, "PmCDA1"), or A derived from a mammal (e.g., human, pig, cow, horse, monkey, etc.) ID (activation-induced cytidine deaminase; AICDA), and APOBEC are exemplary cytidine deaminase It is minase.

[0180] The term "conservative amino acid substitution" or "conservative variation" refers to a mutation in which an amino acid is replaced by another amino acid that shares a common characteristic. It refers to the substitution of an amino acid with another amino acid that has the same properties. It defines the common properties between individual amino acids. A functional method for determining the normality of amino acid changes between corresponding proteins of homologous organisms is The goal is to analyze the frequency of the results (Schulz, GE and Schirmer, RH, Principles of f Protein Structure, Springer-Verlag, New York (1979). According to such analysis, Amino acids within a group are preferentially exchanged with each other, thus affecting the overall protein structure. Define the group of amino acids that are most similar to each other in their effect on the (Schulz, GE and Schirmer, RH, supra). Non-limiting examples of conservative mutations include: For example, the amino acids arginine to lysine and the like that can maintain a positive charge are used. Reverse; aspartic acid to glutamic acid and vice versa, which can maintain a negative charge; free threonine to serine, which maintains the -OH; and asparagine to asparagine, which maintains the free NH Examples include amino acid substitutions such as glutamine.

[0181] The terms "coding sequence" or "protein-coding sequence" are used interchangeably herein. " refers to a segment of a polynucleotide that encodes a protein. The sequence is bounded by a start codon near the 5' end and a stop codon near the 3' end. The coding sequence is also called an open reading frame.

[0182] As used herein, the terms "deaminase" or "deaminase domain" and "deaminase domain" refer to refers to a protein or enzyme that catalyzes a deamination reaction. The enzyme is an adenosine dehydrogenase that catalyzes the hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase is an adenosine or adenine (A) deaminase. adenosine deaminase, which catalyzes the hydrolytic deamination of α- and β-inosine(I) In some embodiments, the deaminase or deaminase domain is an adenosine deaminase. enzymes that convert adenosine or deoxyinosine to inosine or deoxyinosine, respectively. In some embodiments, adenosine deamination This enzyme catalyzes the hydrolytic deamination of adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., genetically engineered adenosine deaminases) aminase, evolved adenosine deaminase) can be derived from any organism, including bacteria. In some embodiments, the adenosine deaminase is obtained from Escherichia coli, Staphylococcus aureus, us aureus, Salmonella typhimurium, Shewanella putrefaciens, Haemophilus influenzae ae, or Caulobacter crescentus.

[0183] In one embodiment, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is a TadA variant. The dA variant is TadA*8. In one embodiment, the deaminase or deaminase The domains may be, for example, human, chimpanzee, gorilla, monkey, cow, dog, rat, or mammal. In some embodiments, the deaminase is a variant of a naturally occurring deaminase from an organism such as a mouse. The enzyme or deaminase domain is non-naturally occurring. For example, in some embodiments In this embodiment, the deaminase or deaminase domain is At least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least At least 92%, at least 93%, at least, at least 94%, at least 95%, at least 96%, At least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, At least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99 0.7%, at least 99.8%, or at least 99.9% identity. For example, Zedaima is registered in international PCT application numbers PCT / 2017 / 045381 (WO2018 / 027078) and PCT / US2016 / 0583 44 (WO2017 / 070632), each of which is incorporated herein by reference in its entirety. Komor, AC, et al., “Programmable editing of a target base n genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, NM, et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage”Nature 551, 464-471 (2017); Komor, AC, et al., “Im proved base excision repair inhibition and bacteriophage Mu Gam protein yields C :G-to-T:A base editors with higher efficiency and product purity”Science Advanc es 3:eaao4774 (2017) and Rees, HA, et al., “Base editing: precision chemist ry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec; 19(12):770-788. See also doi: 10.1038 / s41576-018-0059-1 (the entire contents of which are hereby incorporated by reference). (Incorporated into the specification).

[0184] "Detection" refers to identifying the presence, absence, or amount of an analyte to be detected. In an embodiment, sequence variations in a polynucleotide or polypeptide are detected. In another embodiment, the presence of indels is detected.

[0185] A "detectable label" means a label that, when attached to a molecule of interest, is detectable by spectroscopic, photochemical, or biochemical means. "detectable" refers to a composition that renders the latter detectable through biological, immunochemical, or chemical means. For example, useful labels include radioisotopes, magnetic beads, metal beads, colloidal particles, fluorescent Photochromic dyes, electron-dense reagents, enzymes (e.g., commonly used in enzyme-linked immunosorbent assays (ELISAs)) These include hydroxybenzoates, biotin, digoxigenin, or haptens.

[0186] "Disease" means any condition that damages or interferes with the normal function of a cell, tissue, or organ. Or it means disability.

[0187] As used herein, the term "effective amount" refers to an amount sufficient to induce a desired biological response. The term "amount of biologically active agent" refers to the amount of biologically active agent used to practice the present invention for the therapeutic treatment of disease. The effective amount of active agent(s) administered will depend on the mode of administration, the age, weight, and general health of the subject. Ultimately, your doctor or veterinarian will determine the appropriate amount and dosage. Such an amount is referred to as an "effective" amount. In one embodiment, an effective amount is a dose that is administered to a cell (e.g., an iPSC). A gene of interest in a cell (in vitro or in vivo) is prepared by subjecting the gene to a gene encoding the gene of interest to a gene encoding the gene of interest. Base editors (e.g., programmable DNA binding proteins, nucleobase editors, and In some embodiments, the fusion proteins provided herein are An effective amount of a protein, e.g., an nCas9 domain and a deaminase domain (e.g., an adenosine triphosphate deaminase domain) is administered. An effective amount of a nucleobase editor comprising a nucleobase editor (e.g., cytidine deaminase or cytidine deaminase) is Induce editing of target sites that are specifically bound and edited by the nucleobase editor of In one embodiment, an effective amount refers to an amount of a fusion protein sufficient to have a therapeutic effect (e.g., (e.g., reducing or controlling a disease or its symptoms or conditions) Such a therapeutic effect may be felt in all cells of a subject, tissue, or organ. It need not be sufficient to alter the gene of interest in the subject, tissue, or organ, but rather to Alter the gene of interest in approximately 1%, 5%, 10%, 25%, 50%, 75% or more of the cells present. All you have to do is

[0188] In some embodiments, the fusion proteins provided herein (e.g., nCas9 domains) are Main and deaminase domains (e.g., adenosine deaminase, cytidine deaminase) An effective amount of a nucleobase editor, including a nucleobase editor (e.g., a nucleobase editor enzyme), is a fusion protein sufficient to induce editing of a target site that is specifically bound and edited by As will be appreciated by those skilled in the art, the term "quantity" refers to the amount of an agent (e.g., a fusion protein, a nucleic acid, ase, hybrid protein, protein dimer, protein (or protein An effective amount of a complex of a dimer and a polynucleotide, or a polynucleotide, can be, for example, , a desired biological response, e.g., a specific allele, genome, or target region to be edited. Various factors, such as the location, the cell or tissue being targeted, and / or the agent being used, may be involved. This may vary depending on the child.

[0189] By "fragment" is meant a portion of a polypeptide or nucleic acid molecule, which portion is identical to that of a reference nucleic acid At least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% of the entire length of the molecule or polypeptide , or 90%. The fragments may be 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 3 nucleotides or amino acids. do.

[0190] "Guide RNA" or "gRNA" is a polynucleotide that can be specific for a target sequence. a programmable nucleotide-binding domain protein (e.g., Cas9 or Cpf1) In one embodiment, the term "polynucleotide" refers to a polynucleotide that can form a complex with a The guide polynucleotide is a guide RNA (gRNA). gRNA is a complex of two or more RNAs. It can exist as a single RNA molecule or as a single RNA molecule. The gRNAs present are sometimes called single guide RNAs (sgRNAs), but "gRNA" refers to a single Interchangeable to refer to guide RNAs that exist as a single molecule or as a complex of two or more molecules. Typically, gRNAs exist as a single RNA species, and are used in a variety of ways. (1) They are homologous to the target nucleic acid. (2) a domain that shares a common property (e.g., directs binding of the Cas9 complex to the target); and In one embodiment, the domains are the s9 protein-binding domains. In (2), the sequence corresponds to the sequence known as tracrRNA and contains a stem-loop structure. In some embodiments, domain (2) is selected from the group consisting of the amino acid sequence of Jinek et al., Science 337:816-821(201 2) tracrRNA as provided in (the entire contents of which are incorporated herein by reference) Other examples of gRNAs (e.g., those containing domain 2) are listed on September 6, 2013. In a filed U.S. provisional patent application entitled "Switchable Cas9 Nucleases and Uses Thereof," U.S. SSN 61 / 874,682 and "Delivery System For Functional Numerical System" filed September 6, 2013 The entire disclosure of each of these patent applications can be found in U.S. Provisional Patent Application No. USSN 61 / 874,746 entitled "Patent Application No. 61 / 874,746," which is hereby incorporated by reference. The contents of which are incorporated herein by reference. In some embodiments, the gRNA A gRNA containing two or more of the main (1) and (2) sequences may be referred to as an "extended gRNA." The gRNA binds to two or more Cas9 proteins and expresses two or more Cas9 proteins as described herein. The gRNA binds to the target nucleic acid at a different region on the target site. sequence, which mediates binding of the nuclease / RNA complex to the target site, :Provides sequence specificity of the RNA complex.

[0191] "Hybridization" means hydrogen bonding between complementary nucleobases, as defined by Watson- It can be a Crick, Hoogsteen or reversed Hoogsteen hydrogen bond. For example, Adenine and thymine are complementary nucleobases that form hydrogen bonds to form pairs.

[0192] The term "inhibitor of base repair" or "IBR" refers to a nucleic acid Inhibiting the activity of repair enzymes, such as base excision repair (BER) enzymes In one embodiment, IBR refers to a protein capable of inhibiting inosine base excision repair. Examples of inhibitors of base repair include APE1, Endo III, Endo IV, Endo V, and Endo o VIII, Fpg, hOGGl, hNEILl, T7 Endol, T4 PDG, UDG, hSMUGL, and hAAG inhibitors In one embodiment, the IBR is an inhibitor of Endo V or hAAG. In some embodiments, the IBR is catalytically inactive EndoV or catalytically inactive hAAG. In one embodiment, the base repair inhibitor is an inhibitor of Endo V or hAAG. , the base repair inhibitor is catalytically inactive EndoV or catalytically inactive hAAG.

[0193] In some embodiments, the base repair inhibitor is a uracil glycosylase inhibitor (UGI). UGI is a gene that can inhibit the base excision repair enzyme uracil-DNA glycosylase. In some embodiments, the UGI domain refers to a protein that is a wild-type UGI or In some embodiments, the UGI proteins provided herein comprise a fragment of U In one embodiment, the present invention includes a fragment of GI and a protein homologous to UGI or a UGI fragment. The base repair inhibitor is an inhibitor of inosine base excision repair. In the present study, base repair inhibitors were identified as "catalytically inactive inosine-specific nucleases" or "dead" nucleases. Without wishing to be bound by any particular theory, However, catalytically inactive inosine glycosylases (e.g., alkyladenine glycosylases) Abasic amino acid sequence (AAG) can bind to inosine but also creates an abasic site. It is also unable to remove the inosine moiety, thereby blocking the newly formed inosine moiety from DNA damage / repair. In some embodiments, catalytically inactive inositols are sterically blocked from the cleavage mechanism. Inosine-specific nucleases can bind to inosine in nucleic acids but do not cleave the nucleic acid. Representative, non-limiting examples of catalytically inactive inosine-specific nucleases include: Catalytically inactive alkyl adenosine glycosylase (AAG nuclease) (e.g., from humans) and catalytically inactive endonuclease V (EndoV nuclease) (e.g., from E. coli). In some embodiments, catalytically inactive AAG nucleases are used. The AAG nuclease contains an E125Q mutation or a corresponding mutation in another AAG nuclease.

[0194] By "increase" is meant a positive change of at least 10%, 25%, 50%, 75%, or 100%.

[0195] An "intein" excises itself and releases the remaining fragment (the extein) (extein)) to form peptide bonds in a process known as protein splicing Inteins are fragments of proteins that can be linked together using a "protein intron." The intein excises itself and links the rest of the protein. The process is referred to herein as "protein splicing" or "intein-mediated In some embodiments, precursor proteins (integrins) are spliced. Inteins (intein-containing proteins before intein-mediated protein splicing) Such inteins are referred to herein as split inteins. These are called split inteins (e.g., split intein-N and split intein-C). In Bacteria, DnaE, ​​the catalytic subunit a of DNA polymerase III, is expressed in two separate genes. Encoded by the genes dnaE-n and dnaE-c. Encoded by the dnaE-n gene The intein may be referred to herein as "intein N." The loaded intein may be referred to herein as "intein C."

[0196] Other intein systems can also be used. For example, the dnaE intein, i.e., Cfa-N (e.g., Based on the intein pair Cfa-C (e.g., split intein-N) and Cfa-C (e.g., split intein-C) Synthetic inteins have been described (see, e.g., Stevens, J. Med. Chem. Soc. 1999, 144:111-112, which is incorporated herein by reference). s et al., J Am Chem Soc. 2016 Feb. 24; 138(7):2162-5). Use in accordance with this disclosure Non-limiting examples of intein pairs that can be used include Cfa DnaE intein, Ssp GyrB intein, and Intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB intein Cne intein, and Cne Prp8 intein (see, e.g., U.S. Pat. No. 6,223,629, incorporated herein by reference). Examples include those described in Patent No. 8,394,604.

[0197] Exemplary nucleotide and amino acid sequences of inteins are provided. DnaE Intein-N DNA: TGCCTGTCATACGAAACCGAGATACTGACAGTAGAATATGGCCTTCTGCCAATCGGG AAGATTGTGGAGAAACGGATAGAATGCACAGTTTACTCTGTCGATAACAATGGTAACATTTATACTCAGCCAGTTGCCCA GTGGCACGACCGGGGAGAGCAGGAAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAGGGCCACTAAGGACCACA AATTTATGACAGTCGATGGCCAGATGCTGCCTATAGACGAAATCTTTGAGCGAGAGTTGGACCTCATGCGAGTTGACAAC CTTCCTAAT DnaE Intein-N Protein: CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDR GEQEVFEYCLEDGSLIRATKDHKFMTVDGQMLPIDEIFERELDLMRVDNLPN DnaE Intein-C DNA: ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGA TATTGG AGTCGAAAGAGATCACAACTTTGCTCTGAAGAACGGATTCATAGCTTCTAAT Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN Cfa-N DNA: TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGCCTATTGGAAAGATTGTCGAA GAGAGAATTGAATGCACAGTATATACTGTAGACAAGAATGGTTTCGTTTACACACAGCCCATTGCTCAATGGCACAATCG CGGCGAACAAGAAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACGAGCAACTAAAGATCATAAATTCATGACCA CTGACGGGCAGATGTTGCCAATAGATGAGATATTCGAGCGGGGCTTGGATCTCAAACAAGTGGATGGATTGCCA Cfa-N protein: CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQ MLPIDEIFERGLDLKQVDGLP Cfa-C DNA: ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAGGAAAGTAAAGATAATATCT CGAAAAAGCTTGGTACCCAAAATGTCTATGATATTGGAGTGGAGAAAGATCACAACTTCCTTCTCAAGAACGGTCTCGT AGCCAGCAAC Cfa-C protein: MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLVASN

[0198] To join the N-terminal part of the split Cas9 with the C-terminal part of the split Cas9, intein N and intein B were used. The tein C can be fused to the N-terminal end of the split-Cas9 and the C-terminal end of the split-Cas9, respectively. For example, In some embodiments, intein-N is C-terminal to the N-terminal portion of the split Cas9. The split-Cas9 is fused to the intein-N domain, forming the structure N--[N-terminal portion of split-Cas9]-[intein-N]--C. In some embodiments, the intein-C is located at the N-terminus of the C-terminal portion of the split Cas9. The intein is fused to the C-terminal portion of the split Cas9, forming the structure N-[intein-C]-[C-terminal portion of the split Cas9]-C. Intein for linking the protein (e.g., split-Cas9) to which the intein is fused The mechanism of tein-mediated protein splicing is described, for example, in the publications As described in Shah et al., Chem Sci. 2014; 5(1):446-461, the present invention Methods for designing and using inteins are known in the art. and are disclosed, for example, in WO2014004336, WO2017132580, US20150344549 and US20180127780. and US Pat. No. 6,263,999, each of which is incorporated herein by reference in its entirety.

[0199] The terms "isolated," "purified," or "biologically pure" refer to a substance that is in its native state. Substances that have been removed to varying degrees from components normally associated with them when found in the natural state. "Isolated" refers to the degree of separation from the original source or surrounding environment. "Purified" refers to the degree of separation from the original source or surrounding environment. A "purified" or "biologically pure" protein is one that is free from impurities. that the substance does not materially affect the biological properties of the protein or cause other adverse consequences. In other words, the nucleic acid or peptide of the present invention is or, if produced by recombinant DNA technology, cellular material, viral material, or culture medium. If the substance is not naturally contained in the substance, or if it is chemically synthesized, chemical precursors or other chemical substances Purity and homogeneity are typically determined by analytical methods. Chemical techniques, such as polyacrylamide gel electrophoresis or high performance liquid chromatography The term "purified" refers to the degree to which a nucleic acid or protein is purified by electrophoresis. This can mean that the resulting protein essentially produces one band. For proteins that can undergo glycosylation, the different modifications are purified separately. This can result in different isolated proteins that can be isolated.

[0200] An "isolated polynucleotide" is a nucleic acid molecule that does not occur in the naturally occurring genome of the organism from which the nucleic acid molecule of the invention is derived. It means nucleic acid (e.g., DNA) that does not contain the genes adjacent to the gene. The term refers to, for example, those incorporated into vectors; autonomously replicating plasmids or viruses. integrated into the genomic DNA of prokaryotes or eukaryotes; or independent of other sequences another molecule created (e.g., cDNA generated by PCR or restriction endonuclease digestion) or genomic or cDNA fragments). RNA molecules transcribed from DNA molecules, as well as hybrids encoding additional polypeptide sequences. It contains recombinant DNA that is part of the hybrid gene.

[0201] An "isolated polypeptide" is a polypeptide of the invention separated from components that naturally accompany it. Typically, a polypeptide is a polypeptide that is a protein with which it is naturally associated. A substance is isolated if it is at least 60% by weight free from substances and naturally occurring organic molecules. Preferably, the preparation comprises at least 75% by weight of soluble fiber, more preferably at least 90% by weight of soluble fiber, most preferably at least 10% by weight of soluble fiber. or at least 99% of the isolated polypeptide of the present invention. can be prepared by, for example, extraction from a natural source, expression of a recombinant nucleic acid encoding such a polypeptide, or the like. or by chemically synthesizing the protein. Purity can be achieved by any suitable method. Suitable methods, such as column chromatography, polyacrylamide gel electrophoresis, or can be measured by HPLC analysis.

[0202] As used herein, the term "linker" refers to a molecule that binds two molecules or moieties (e.g., a Two components of a protein or ribonucleocomplex, or two domains of a fusion protein In, for example, a polynucleotide programmable DNA binding domain (e.g., dCas9) Deaminase domains (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase), or the napDNAbp domain (e.g. Cas12b) and a deaminase domain (e.g., adenosine deaminase or cytidine deaminase) Covalent linkers (e.g., covalent bonds), non-covalent linkers, chemical bonds, etc., linking the In certain embodiments, the linker can refer to a chemical group, a molecule, or a molecule that binds to the Cas protein. The linker is adjacent to the deaminase domain inserted within the protein or a fragment thereof. Connecting different components or parts of a component in a editor system For example, in some embodiments, the linker can be a polynucleotide protease. a guide polynucleotide binding domain for a grammable nucleotide binding domain; and In some embodiments, the linker can link the catalytic domains of the deaminase. The ISPR polypeptide and the deaminase can be linked. The linker can link the Cas9 and the deaminase. In some embodiments, the linker is In some embodiments, the linker can link the dCas9 and the deaminase. The s9 and the deaminase can be linked. For example, in one embodiment, the linker can be Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h In some embodiments, the linker may link the guidepoietin and the deaminase. In some embodiments, the oligonucleotide and the deaminase can be linked. The linker is a base editor system having a deaminating component and a polynucleotide programming component. In some embodiments, the linkable nucleotide binding moieties can be linked. The anchor connects the RNA-binding moiety of the deaminating component of the base editor system with the napDNAbp component. In some embodiments, the linker can be used to decouple the base editor system. The RNA-binding portion of the amination component and the polynucleotide-programmable nucleotide bond In some embodiments, the linker can be a base element. The RNA-binding moiety of the deaminating component of the RNA-binding protein system and the polynucleotide programmable nuclease A linker is a molecule that can link two groups, molecules, or a combination of two or more ... or other moieties, and are disposed between or sandwiched by them and are bonded covalently or non-covalently They are connected to each other through binding interactions, thus allowing the two to be linked together. In various embodiments, the linker can be an organic molecule, group, polymer, or chemical moiety. In embodiments, the linker can be a polynucleotide. In some embodiments, the linker may be a DNA linker. In some embodiments, the linker may be an RNA linker. In such a case, the linker can include an aptamer capable of binding to the ligand. In some embodiments, the ligand is a carbohydrate, peptide, protein, or nucleic acid. In some embodiments, the linker can be an aptamer derived from a riboswitch. The riboswitch from which the aptamer is derived can be theophylline riboswitch. , thiamine pyrophosphate (TPP) riboswitch, adenosine cobalamin (AdoCbl) riboswitch switch, S-adenosylmethionine (SAM) riboswitch, SAH riboswitch, flavin mononucleotide Nucleotide (FMN) riboswitch, tetrahydrofolate riboswitch, lysine riboswitch, Glycine riboswitch, purine riboswitch, GlmS riboswitch, or prequeosine 1 (PreQ1) riboswitch. In some embodiments, the linker can be selected from a polypeptide The aptamer may comprise an aptamer bound to a protein domain, such as a polypeptide ligand. In some embodiments, the polypeptide ligand comprises a K homology (KH) domain, an MS2 coat protein, or a phosphodiesterase (MPS) domain. Protein domain, PP7 coat protein domain, SfMu Com coat protein domain, Sterile α motif, telomerase Ku binding motif and Ku protein, telomerase Sm7 The Sm7 protein binding motif may be an Sm7 protein binding motif or an RNA recognition motif. Thus, the polypeptide ligand can be part of a base editor system component. The base editing component may comprise a deaminase domain and an RNA recognition motif.

[0203] In some embodiments, the linker is an amino acid or a plurality of amino acids (e.g., a peptide or In some embodiments, the linker may be about 5 to 100 amino acids in length. For example, lengths of approximately 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20-30 , 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, or 90-100 amino acids. In some embodiments, the linker has a length of about 100-150, 150-200, 200-250, 250-300, It can be 300-350, 350-400, 400-450, or 450-500 amino acids. Shorter linkers are also contemplated.

[0204] In some embodiments, the linker comprises an RNA programmable The gRNA binding domain of a nuclease and a nucleic acid editing protein (e.g., cytidine or adenine) In some embodiments, the linker connects the catalytic domain of dCas 9 and the nucleic acid editing protein. For example, the linker may be a linker that connects two groups, molecules, or other Located between or flanked by two groups, molecules, or other moieties, are linked to each other via a covalent bond, thus linking the two. The anchor is an amino acid or multiple amino acids (eg, a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5 to 200 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 70, 80, 95, 100, 150, 100, 250, 350, 400, 550, 600, 700, 800, 950, 1000, 1500, 1000, 2500, 3500, 4000, 5 ...6000, 7000, 8000, 9500, 1000, 1 0, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 7 0, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140, 150 , 160, 175, 180, 190, or 200 amino acids. is also contemplated.

[0205] In some embodiments, the nucleobase editor domain is SGGSSGSETPGTSESATPESSGGS, SG GSSGGSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSA The amino acid sequence of PGSPAGSPTSTEEGTSTE PSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS In some embodiments, the nucleobase editor domain is fused via a linker comprising The XTEN linker is fused via a linker containing the amino acid sequence SGSETPGTSESATPES, which may also be referred to as an XTEN linker. In some embodiments, the linker comprises the amino acid sequence SGGS. , the linker is (SGGS) n , (GGGS) n , (GGGGS) n , (G) n、 (EAAAK) n , (GGS) n ,SGSETPGTSESAT PES or (XP) n motif, or any combination thereof, where n is a unique In some embodiments, n is an integer between 1 and 30, and X is any amino acid. is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15.

[0206] In some embodiments, the linker is 24 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In some embodiments, the linker has the amino acid sequence SGGSSGGSSSGSETP In some embodiments, the linker comprises GTSESATPESSGGSSGGSSGGSSGGS. In some embodiments, the linker has the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSG In some embodiments, the linker comprises GSSGGSSGGSSGSETPGTSESATPESSGGS SGGS. In some embodiments, the linker has the amino acid sequence PGSPAGSPTSTEEGTSESATP. Contains ESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS.

[0207] A "marker" is any molecule that has an altered expression level or activity that is associated with a disease or disorder. The term "protein" refers to any protein or polynucleotide.

[0208] As used herein, the term "mutation" refers to a change in a sequence, e.g., a nucleic acid or amino acid sequence. The substitution of a residue in a sequence of amino acids by another residue, or the deletion of one or more residues in the sequence. Mutations, as used herein, are typically made by identifying the original residue and then and identifying the newly substituted residue, Various methods for making amino acid substitutions (mutations) are provided herein. are well known in the art and are described, for example, in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012). In some embodiments, this The disclosed base editors do not generate a significant number of unintended mutations, e.g., unintended point mutations. without creating an "intended mutation" in a nucleic acid (e.g., a nucleic acid in a subject's genome), e.g. Point mutations can be efficiently generated. In some embodiments, the intended The mutations are generated by a guide polynucleotide sequence specifically designed to produce the intended mutation. A specific base editor (e.g., a cytidine base editor or is a mutation caused by an adenosine base editor.

[0209] Generally, the amino acid sequence of the present invention is made or identified in a sequence (e.g., an amino acid sequence described herein). The mutations detected are numbered relative to the reference (or wild-type) sequence, i.e., the sequence that does not contain the mutation. Those skilled in the art will recognize mutations in amino acid and nucleic acid sequences relative to a reference sequence. It will be easy to understand how to determine the position.

[0210] The term "non-conservative mutation" refers to amino acid substitutions between different groups, e.g. For example, tryptophan to lysine, or serine to phenylalanine. The non-conservative amino acid substitutions do not disrupt or inhibit the biological activity of the functional variant. Non-conservative amino acid substitutions are preferred because they improve the biological activity of the functional variant compared to the wild-type protein. The biological activity of the functional variant can be enhanced so that it is increased compared to the protein. do.

[0211] The term "nuclear localization sequence," "nuclear localization signal," or "NLS" refers to the localization of a protein. The term "nuclear localization sequence" refers to an amino acid sequence that promotes import into the cell nucleus. For example, WO / 2001 / 038547 filed on November 23, 2000 and published on May 31, 2001 is incorporated herein by reference. and Plank et al. in published international PCT application PCT / EP 2000 / 011690, in which No. 6,239,493, which is incorporated herein by reference for its disclosure of exemplary nuclear localization sequences. In embodiments, the NLS may be any of the NLSs described, for example, in Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4 172. In some embodiments, the NLS is Sequence KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR , RKSGKIAAIVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.

[0212] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to a nucleic acid molecule comprising a nucleobase and Compounds containing an acidic moiety, such as nucleosides, nucleotides, or polynucleotides Typically, a polymeric nucleic acid, e.g., a nucleic acid molecule containing three or more nucleotides is a linear sequence in which adjacent nucleotides are linked to each other via phosphodiester bonds. In some embodiments, a "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or or nucleosides). In some embodiments, a "nucleic acid" refers to three or more individual nucleosides. As used herein, the term "oligonucleotide" refers to an oligonucleotide chain containing nucleotide residues. "Nucleotide" and "polynucleotide" refer to a polymer of nucleotides (e.g., may be used interchangeably to refer to a chain of at least three nucleotides. In this context, "nucleic acid" encompasses RNA and single- and / or double-stranded DNA. For example, genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome It may naturally occur in the context of a chromatid, or other naturally occurring nucleic acid molecule. On the other hand, nucleic acid molecules can be, for example, non-naturally occurring molecules, recombinant DNA or RNA, artificial chromosomes, Engineered genomes, or fragments thereof, or synthetic DNA, RNA, DNA / RNA hybrids, or or non-naturally occurring molecules containing non-naturally occurring nucleotides or nucleosides Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms may be used interchangeably. Nucleic acids include nucleic acid analogs, e.g., analogs having other than a phosphodiester backbone. purified from the source, produced using a recombinant expression system and optionally purified, or chemically synthesized In the case of chemically synthesized molecules, the nucleic acid may, where appropriate, be Nucleotides such as analogs with chemically modified bases or sugars, and backbone modifications, are also suitable. Nucleic acid sequences are shown in the 5' to 3' direction unless otherwise indicated. In some embodiments, nucleic acids contain natural nucleosides (e.g., adenosine, thymidine, guanine, adenosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine cytidine, and deoxycytidine; nucleoside analogs (e.g., 2-aminoadenosine, 2-thiocytidine); Othymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine, 2- Aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5- Propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanine , O(6)-methylguanine, and 2-thiocytidine); chemically modified bases; biologically modified bases (e.g., methylated bases); inserted bases; modified sugars (e.g., 2'-fluororibose, ribose , 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages) or Includes these.

[0213] The term "nucleic acid programmable DNA binding protein" or "napDNAbp" refers to Used interchangeably with "polynucleotide programmable nucleotide binding domain" A guide nucleic acid or guide polynucleotide that guides the napDNAbp to a specific nucleic acid sequence. It refers to a protein that associates with a nucleic acid (e.g., DNA or RNA) such as a nucleic acid (e.g., gRNA). In embodiments, a polynucleotide-programmable nucleotide binding domain is a polynucleotide-programmable DNA binding domain. In the present invention, the polynucleotide-programmable nucleotide binding domain is In one embodiment, the RNA-binding domain is programmable by oligonucleotides. The polynucleotide-programmable nucleotide binding domain is a Cas9 protein. The Cas9 protein binds to a specific DNA sequence complementary to the guide RNA. In some embodiments, the napDNAbp can bind to a guide RNA that guides the Cas9 domain. Main, e.g., nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease Inactive Cas9 (dCas9). Non-limiting examples of nucleic acid programmable DNA binding proteins. Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, and Cas12c / C2c3 Cas enzymes include Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i. Non-limiting examples include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12) ), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / Ca sX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5 e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3 , Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Cs x1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa 1, Csa2, Csa3, Csa4, Csa5, type II Cas effector protein, type V Cas effector protein Cas effector proteins, type VI Cas effector proteins, CARF, DinG, and their homologs Other nucleic acid programming may be used. Also included are DNA binding proteins that may not be specifically listed in this disclosure. It is within the scope of this disclosure, e.g., Makarova et al. CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct;1:325-336. doi: 10.10 89 / crispr.2018.0033; Yan et al., “Functionally diverse type V CRISPR-Cas system s” Science. 2019 Jan 4;363(6422):88-91. See doi: 10.1126 / science.aav7271 (see the entire contents of each of which are incorporated herein by reference).

[0214] The terms "nucleobase," "nitrogenous base," or "base" are used interchangeably herein. refers to nitrogen-containing biological compounds that form nucleosides, which are nucleotides The ability of nucleobases to form base pairs and stack with each other is directly related to the synthesis of ribonucleic acid ( Adenine (A) is a nucleotide that gives rise to long helical structures such as RNA and deoxyribonucleic acid (DNA). The five nucleic acid bases are cytosine (C), guanine (G), thymine (T), and uracil (U). They are called primary or standard bases. Adenine and guanine are derived from purines, while cytosine, Uracil and thymine are derived from pyrimidines. DNA and RNA contain modified other (non-primary) bases. Non-limiting exemplary modified nucleobases include hypoxanthine, xanthine, , 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydroxybenzoates. Hypoxanthine and xanthine are produced in the presence of mutagens. Both can be produced by deamination (replacement of an amine group with a carbonyl group). Hypoxanthine can be modified from adenine. Xanthine can be modified from guanine. Uracil can be generated by deamination of cytosine. A "nucleoside" is a nucleic acid base and a five-carbon atom. They consist of a sugar (ribose or deoxyribose). Examples of nucleosides include adenosine , guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, These include deoxyguanosine, thymidine, deoxyuridine, and deoxycytidine. Nucleosides with modified nucleobases include inosine (I), xanthosine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), psoriasis Nucleotides consist of a nucleic acid base, a pentose sugar (ribose or is deoxyribose), and at least one phosphate group.

[0215] The term "nucleic acid programmable DNA binding protein" or "napDNAbp" refers to napDNA Associate with a nucleic acid (e.g., DNA or RNA) such as a guide nucleic acid that directs the Abp to a specific nucleic acid sequence For example, the Cas12 protein is a protein that binds the Cas12 protein to a guide RNA complementary to the guide RNA. In some embodiments, the nucleic acid sequence can be linked to a guide RNA that directs the nucleic acid sequence to a specific DNA sequence. , napDNAbp is a Cas12 domain, e.g., a nuclease-active Cas12 domain. Examples of bps include Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, and Ca Other napDNAbps are also specifically listed in this disclosure. Although this may not be possible, it is within the scope of this disclosure. on and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct ;1:325-336. doi: 10.1089 / crispr.2018.0033; Yan et al., “Functionally diverse ty pe V CRISPR-Cas systems” Science. 2019 Jan 4;363(6422):88-91. doi: 10.1126 / scie See nce.aav7271, the entire contents of each of which are incorporated herein by reference.

[0216] The terms "nucleobase editing domain" or "nucleobase editing protein" are used herein. When used in the formula, cytosine (or cytidine) to uracil (or uridine) or hypoxia to thymine (or thymidine) and from adenine (or adenosine) Deamination to sansine (or inosine), and non-templated nucleotide addition and insertion Proteins or enzymes that can catalyze nucleobase modifications in RNA or DNA, such as refers to an enzyme. In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., For example, adenine deaminase or adenosine deaminase; or cytidine deaminase In some embodiments, the nucleobase editing domain is a nucleobase editing domain (e.g., a nucleobase editing domain ... The enzymes may contain multiple deaminase domains (e.g., adenine deaminase or adenosine deaminase). aminase and cytidine or cytosine deaminase). The acid-base editing domain may be a naturally occurring nucleobase-editing domain. In embodiments, the nucleobase-editing domain is engineered from a naturally occurring nucleobase-editing domain. The nucleobase editing domain may be engineered or evolved in bacteria, Any living organism, such as a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse For example, the nucleobase-editing protein can be derived from the protein described in International PCT Application No. PCT / 2017 / 0453 81 (WO 2018 / 027078) and PCT / US2016 / 058344 (WO 2017 / 070632), which each of which is incorporated herein by reference in its entirety. “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, NM, et al., “Programmabl e base editing of A·T to G·C in genomic DNA without DNA cleavage”Nature 551, 464-471 (2017); and Komor, AC, et al., “Improved base excision repair inhi bition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with high See also “Efficiency and product purity” Science Advances 3:eaao4774 (2017) the entire contents of which are incorporated herein by reference).

[0217] As used herein, "obtaining," as in "obtaining a drug," means This includes synthesizing, purchasing, or otherwise obtaining the drug.

[0218] As used herein, a "patient" or "subject" refers to a person who has been diagnosed with a disease or disorder. Mammals at risk of developing them or suspected of having or developing them In certain embodiments, the term "patient" refers to a subject or individual suffering from a disease or disorder. refers to a mammalian subject with a higher than average likelihood of developing a disease. Exemplary patients include humans, non-human primates, cats, dogs, pigs, cattle, horses, camels, llamas, goats, sheep, rodents (e.g. mice, rabbits, rats, guinea pigs) and will benefit from the treatments disclosed herein. Exemplary human patients include males and / or females. could be.

[0219] A "patient in need thereof" or a "subject in need thereof" is used herein to refer to a person with a disease or have been diagnosed with, are at risk of having, or may have a disability or disorder It is referred to as a patient who has been predetermined or is suspected of having it.

[0220] "Pathogenic mutation," "pathogenic variant," "disease-causing mutation," "disease-causing variant" The terms "deleterious mutation" or "predisposing mutation" refer to a mutation that is associated with a particular disease or disorder. Refers to a genetic change or mutation that increases an individual's susceptibility or predisposition. A pathogenic variant is a mutation that results in at least one wild-type mutation in the protein encoded by the gene. This includes those in which an amino acid has been replaced with at least one pathogenic amino acid.

[0221] The term "pharmaceutically acceptable carrier" refers to a liquid or solid filler, diluent, excipient, manufacturing agent, or Auxiliaries (e.g. lubricants, magnesium talc, calcium stearate or stearic acid zinc, or stearic acid), or solvent encapsulating materials, etc., in certain parts of the body (e.g. transport or transport of a compound from one site (e.g., a delivery site) to another site (e.g., an organ, tissue, or part of the body) Pharmaceutically acceptable materials, compositions, or vehicles involved in the delivery or transport of pharmaceuticals. A commercially acceptable carrier is one that is "compatible" with the other ingredients of the formulation and is not deleterious to the tissues of the subject. "Acceptable" (e.g., physiological compatibility, sterility, physiological pH, etc.). Terms such as "pharmaceutically acceptable carrier," "vehicle," and the like are used interchangeably herein. do.

[0222] The term "pharmaceutical composition" may refer to a composition formulated for pharmaceutical use.

[0223] The terms "protein," "peptide," "polypeptide," and their grammatical equivalents are used interchangeably herein and are linked to each other by a peptide (amide) bond. The term refers to a polymer of amino acid residues bound together in a single chain. It typically refers to a protein, peptide, or polypeptide. A protein, peptide, or polypeptide is at least three amino acids in length. A polypeptide can refer to an individual protein or a group of proteins. One or more amino acids in a protein, peptide, or polypeptide may have, for example, a carbohydrate group, a hydrophobic group, xyl group, phosphate group, farnesyl group, isofarnesyl group, fatty acid group, linkage They may be modified by the addition of chemical entities such as anchors, functionalization, or other modifications. Proteins, peptides, or polypeptides may also be single molecules or multi-molecular. The protein, peptide, or polypeptide may be a naturally occurring It may be merely a fragment of a protein or peptide. Peptides may be naturally occurring, recombinant, or synthetic, or any of the foregoing. As used herein, the term "fusion protein" refers to a combination of: Hybrid polypeptides containing protein domains from at least two different proteins One protein is the amino-terminal (N-terminal) portion of the fusion protein or the capsid. can be located at the carboxy-terminus (C-terminus) of a protein, thus The protein forms a terminal or carboxy-terminal fusion protein. a nucleic acid binding domain (e.g., a nucleic acid binding domain that induces binding of a protein to a target site) Cas9 gRNA binding domain and nucleic acid cleavage domain, or nucleic acid editing protein In some embodiments, the protein may comprise a proteinaceous moiety, e.g., a catalytic domain of For example, an amino acid sequence constituting a nucleic acid binding domain and an organic compound, such as a nucleic acid cleaving agent In some embodiments, proteins include compounds that can act as nucleic acids (e.g., RNA). or DNA) are complexed with or associated with nucleic acids. Any protein can be produced by any method known in the art. For example, the proteins provided herein can be prepared via recombinant protein expression and purification. This is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and are described in Green and Samb rook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Labora and those described by the University of California Press, Cold Spring Harbor, NY (2012). the entire contents of which are incorporated herein by reference.

[0224] The polypeptides and proteins disclosed herein (including functional portions thereof and their functional variants) (including riboflavin) contains synthetic amino acids in place of one or more naturally occurring amino acids Such synthetic amino acids are known in the art and include, for example, aminocyclo Hexanecarboxylic acid, norleucine, α-amino n-decanoic acid, homoserine, S-acetylacetone aminomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenyl phenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine phenylalanine, β-phenylserine, β-hydroxyphenylalanine, phenylglycine, α-Naphthylalanine, cyclohexylalanine, cyclohexylglycine, indoline 2-carboxylic acid, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, Aminomalonic acid monoamide, N'-benzyl-N'-methyllysine, N',N'-dibenzyl-lysine, 6- Hydroxylysine, ornithine, α-aminocyclopentanecarboxylic acid, aminocyclohexyl cyclohexanecarboxylic acid, aminocyclohexanecarboxylic acid, α-aminocycloheptanecarboxylic acid carboxylic acid, α-(2-amino-2-norbornane)-carboxylic acid, α,γ-diaminobutyric acid, α,β-diamino These include α-tert-butylglycine, α-propionic acid, homophenylalanine, and α-tert-butylglycine. Polypeptides and proteins are derived from post-translational modifications of one or more amino acids of the polypeptide construct. Non-limiting examples of post-translational modifications include phosphorylation, acetylation, and the like. and acylation, including formylation, glycosylation (including N-linked and O-linked), amino Derivatization, hydroxylation, alkylation including methylation and ethylation, ubiquitination, pyrolysis, Addition of lidocaine carboxylic acid, formation of disulfide bridges, sulfation, myristoylation, palmitoylation ylation, isoprenylation, farnesylation, geranylation, glypiation, lipoylation and iodination.

[0225] "Polynucleotide programmable nucleotide binding domain" or "nucleic acid programmable nucleotide binding domain" The term "programmable DNA binding protein (napDNAbp)" refers to a polynucleotide programmable DNA binding protein. A guide polynucleotide (e.g., a nucleotide-binding domain) that directs the targetable nucleotide binding domain to a specific nucleic acid sequence. It refers to a protein that binds to nucleic acids (e.g., DNA or RNA) such as guide RNA. In some embodiments, the polynucleotide-programmable nucleotide binding domain The DNA is a polynucleotide programmable DNA binding domain. In one embodiment, the polynucleotide programmable nucleotide binding domain is In some embodiments, the polynucleotide is a programmable RNA binding domain. The nucleotide-programmable nucleotide-binding domain is the Cas12 protein.

[0226] The term "recombinant" as used herein with respect to a protein or nucleic acid means that the protein or nucleic acid is It refers to proteins or nucleic acids that do not occur in nature but are the product of human engineering. For example, In some embodiments, the recombinant protein or nucleic acid molecule is any naturally occurring At least one, at least two, at least three, at least four, or at least at least five, at least six, or at least seven amino acid or nucleotide mutations Contains an octide sequence.

[0227] "Decrease" means a negative change of at least 10%, 25%, 50%, 75%, or 100% .

[0228] "Reference" means a standard or control condition. , the reference is a wild-type or healthy cell. In other embodiments, including but not limited to, the reference are not exposed to the test conditions or are exposed to placebo or normal saline, medium, or buffer. and / or exposed to a control vector that does not carry the polynucleotide of interest, It is a physiological cell.

[0229] A "reference sequence" is a defined sequence used as a basis for sequence comparison. It can be a subset or the entirety of a particular sequence; for example, a full-length cDNA or gene sequence. A segment of a gene, or the complete cDNA or gene sequence. For polypeptides, see the reference polypeptide. The length of the peptide sequence is generally at least about 16 amino acids, at least about 20 amino acids, At least about 25 amino acids, about 35 amino acids, about 50 amino acids, or about 100 amino acids. For nucleic acids, the length of a reference nucleic acid sequence is generally at least about 50 nucleotides, At least about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides, or about 300 nucleotides thiazolinone, ... In one embodiment, the reference sequence is the wild-type sequence of the protein of interest. The sequence is a polynucleotide sequence that encodes the wild-type protein.

[0230] The terms "RNA programmable nuclease" and "RNA-guided nuclease" refer to Used in conjunction with (e.g., bound to or associated with) one or more non-target RNAs In some embodiments, the RNA-programmable nuclease, when complexed with RNA, Typically, the bound RNA is called a guide RNA (gRNA). gRNAs can exist as a complex of two or more RNAs, or as a single RNA molecule. A gRNA that exists as a single RNA molecule is called a single guide RNA (sgRNA). Although sometimes referred to as "gRNA," "gRNA" can occur as a single molecule or as a complex of two or more molecules. are used interchangeably to refer to guide RNAs present as a single RNA species. The gRNA present in the target nucleic acid has (1) a domain that shares homology with the target nucleic acid (e.g., a domain that facilitates Cas9 duplication to the target). and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) is directed against a sequence known as tracrRNA. For example, in some embodiments, domain (2) comprises: Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. The gRNA (e.g., domain) is identical to or homologous to the tracrRNA provided in the Other examples of compounds containing 2 include "Switchable Cas9 Nucleases and Us" filed on September 6, 2013. No. 61 / 874,682, filed September 6, 2013, entitled "Methods Thereof," U.S. Provisional Patent Application USSN 080202009 entitled "Delivery System For Functional Nucleases" No. 61 / 874,746, the entire contents of each of which are incorporated herein by reference. In some embodiments, the gRNA comprises two or more of domains (1) and (2), For example, an extended gRNA may be referred to as an "extended gRNA" as described herein. For example, two or more Cas9 proteins can be bound to target nuclei in two or more different regions. The gRNA contains a nucleotide sequence complementary to the target site, which binds to the target site. It mediates the binding of the nuclease / RNA complex and provides sequence specificity for the nuclease:RNA complex. Provide.

[0231] In some embodiments, the RNA programmable nuclease is a Cas9 enzyme (CRISPR-associated system). endonuclease, such as Cas9 (Casnl) from Streptococcus pyogenes (e.g., For example, "Complete genome sequence of an Ml strain of Streptococcus pyogenes." Ferret ti JJ et al., Proc. Natl. Acad. Sci. USA 98:4658-4663(2001); "CRISPR RNA ma turation by trans-encoded small RNA and host factor RNase III." Deltcheva E. et al., Nature 471:602-607(2011)).

[0232] RNA-programmable nucleases (e.g., Cas9) use RNA:D to target DNA cleavage sites. Since NA hybridization is used, these proteins can, in principle, be used as guide RNAs. Any sequence specified by A can be targeted. For site-specific cleavage, C Methods using RNA-programmable nucleases such as as9 (e.g., to modify genomes) (for example, Cong, L. et al., Multiplex gene ome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013 ); Hwang, WY et al., Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programme d genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., G enome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic acids research (2013); Jiang, W. et al., RNA-guided editing of bacterial genomes using CRISPR-Cas systems. See Nature biotechnology 31, 233-239 (2013). the entire contents of each of which are incorporated herein by reference).

[0233] The term "single nucleotide polymorphism (SNP)" refers to a single nucleotide variation that occurs at a specific position in the genome. where each mutation is present in the population to a noticeable extent (e.g., >1%). For example, At certain base positions in the human genome, C nucleotides can occur in most individuals, In a small number of individuals, the position is occupied by A. This means that there is a SNP at this particular position, and C or A, meaning that two nucleotide variations are alleles at this position. SNPs underlie differences in susceptibility to disease, affecting the severity of the disease and the body's response to treatment. SNPs are found in the coding region of genes, non-coding regions of genes, and so on. It may be present in a gene region, or in an intergenic region (the region between genes). In this case, SNPs within the coding sequence may affect the identity of the protein produced due to the degeneracy of the genetic code. SNPs in the coding region are of two types: synonymous SNPs and non-synonymous SNPs. Synonymous SNPs do not affect the protein sequence, whereas non-synonymous SNPs affect the amino acid sequence of a protein. There are two types of nonsynonymous SNPs: missense and nonsense. SNPs not in the coding region affect gene splicing, transcription factor binding, and messenger receptor (MR) activity. These types of SNPs can affect the degradation of RNA or the sequence of non-coding RNA. The gene expression that is affected is called an eSNP (expressed SNP) and can be upstream or downstream of the gene. Single nucleotide variants (SNVs) are single nucleotide variations with unlimited frequency and occur in somatic Somatic single base variations may also be referred to as single base modifications.

[0234] "Specifically bind" means to recognize and bind to the polypeptide and / or nucleic acid molecule of the present invention. and bind to other molecules in the sample (e.g., biological sample) but do not substantially recognize or bind to other molecules in the sample. Nucleic acid molecules, polypeptides, or complexes thereof (e.g., nucleic acid programmable DNA binding domain and guide nucleic acid), compound, or molecule.

[0235] Nucleic acid molecules useful in the methods of the present invention may encode a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules include any nucleic acid molecule that is 100% identical to the endogenous nucleic acid sequence. Typically, but not necessarily, substantial identity is shown. A polynucleotide having a nucleotide sequence typically hybridizes with at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention can be modified to The present invention also includes any nucleic acid molecule encoding an endogenous nucleotide sequence or a fragment thereof. It need not be 100% identical to the nucleic acid sequence, but typically will show substantial identity. Polynucleotides having "substantial identity" to a double-stranded nucleic acid molecule typically have a small number of double-stranded nucleic acid molecules. "Hybridize" means to hybridize with at least one of the strands of a given molecule. Complementary polynucleotide sequences (e.g., those described herein) are synthesized under various stringency conditions. This means that a double-stranded molecule is formed between the two genes (or a part of them). For example, Wahl, GM and SL Berger (1987) Methods Enzymol. 152:399; Kimmel, A. R (1987) Methods Enzymol. 152:507).

[0236] For example, a stringent salt concentration is typically less than about 750 mM NaCl and 75 mM citrate. trisodium, preferably less than about 500 mM NaCl and 50 mM trisodium citrate, more preferably Preferably, the concentration is less than about 250 mM NaCl and 25 mM trisodium citrate. See hybridization can be obtained in the absence of organic solvents, e.g., formamide. whereas high stringency hybridization requires at least about 35% formaldehyde. It can be obtained in the presence of at least about 50% formamide. Stringent temperature conditions are typically at least about 30°C, more preferably at least about 37°C. 0 C, and most preferably at least about 42°C. The concentration of detergent (e.g., sodium dodecyl sulfate (SDS)) and carrier DNA content Various additional parameters, such as inclusion or exclusion, are well known to those skilled in the art. By combining these various conditions, various levels of stringency can be achieved. In one embodiment, hybridization is achieved at 30° C. in 750 mM NaCl, 75 In another embodiment, hybridization occurs in 10 mM trisodium citrate and 1% SDS. The solution was incubated at 37°C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide. In another embodiment, the hybridization occurs in 100 μg / ml denatured salmon sperm DNA (ssDNA). Hybridization was performed at 42°C in 250 mM NaCl, 25 mM trisodium citrate, 1% S The reaction takes place in 500 μg / ml ssDNA, 50% formamide, and 200 μg / ml ssDNA. Useful variations of these conditions are The alternatives will be readily apparent to those skilled in the art.

[0237] For most applications, the washing steps that follow hybridization are also stringent. Wash stringency conditions are defined by salt concentration and temperature. As mentioned above, wash stringency can be increased by decreasing the salt concentration or increasing the temperature. This can be increased by increasing the temperature. The appropriate salt concentration is preferably less than about 30 mM NaCl and 3 mM trisodium citrate. and most preferably less than about 15 mM NaCl and 1.5 mM trisodium citrate. Stringent temperature conditions for the process are typically at least about 25°C, more preferably In one embodiment, the temperature is at least about 42°C, and even more preferably at least about 68°C. Washing steps were performed at 25°C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the wash steps are carried out in 15 mM NaCl, 1.5 mM NaCl, at 42°C. 10 mM trisodium citrate, and 0.1% SDS. Washing steps were performed at 68°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Further variations of these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton et al. nd Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wil ey Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular C loning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York It has been done.

[0238] "Split" means divided into two or more pieces.

[0239] A "split Cas9 protein" or "split-Cas9" is a protein that is composed of two separate nucleotides. Cas9 protein provided as N-terminal and C-terminal fragments encoded by the sequences The polypeptides corresponding to the N-terminal and C-terminal parts of the Cas9 protein are spliced ​​together. In certain embodiments, the Cas9 protein can be reconstituted to form a "reconstituted" Cas9 protein. The Cas9 protein is described, for example, in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-9 49, 2014, or Jiang et al. (2016) Science 351: 867-871. PDB As described in file: 5F9R (each of which is incorporated herein by reference), In some embodiments, the protein is split into two fragments within a disordered region of the protein. The protein binds to the SpCas9 protein between approximately amino acids A292-G364, F445-K483, or E565-T637. At any C, T, A, or S within the region, or any other Cas9, Cas9 variant ( For example, nCas9, dCas9), or other nap DNA fragments at corresponding positions in the two fragments. In some embodiments, the protein is split into SpCas9 T310, T313, A456, S469, or is split into two fragments at C574. In some embodiments, the protein is split into two fragments. The process of dividing a protein into multiple fragments is called "splitting" the protein. do.

[0240] In other embodiments, the N-terminal portion of the Cas9 protein is S. pyogenes Cas9 wild type (SpC as9) (NCBI Reference Sequence: NC_002737.2, Uniprot Reference Sequence: Q99ZW2) amino acids 1 to 573 The C-terminal part of the Cas9 protein contains amino acids 1 to 637, and the C-terminal part of the Cas9 protein contains amino acids 574 to 1368 of the wild-type SpCas9. includes the part 638 to 1368.

[0241] The C-terminal part of the split Cas9 is ligated with the N-terminal part of the split Cas9 to form the complete Cas9 tag. In some embodiments, the C-terminus of the Cas9 protein can form a protein. The end portion begins where the N-terminal portion of the Cas9 protein ends. In some embodiments, the C-terminal portion of the split Cas9 is amino acids (551-651)-1368 of spCas9. "(551-651) -1368" refers to the amino acids between 551 and 651 (inclusive). This means that the C-terminal portion of the split Cas9 begins with amino acid 1368 and ends with amino acid 1368. The amino acids 551-1368, 552-1368, 553-1368, 554-1368, 555-1368, and 556-1368 of spCas9 are , 557-1368, 558-1368, 559-1368, 560-1368, 561-1368, 562-1368, 563-1368, 564-1368 , 565-1368, 566-1368, 567-1368, 568-1368, 569-1368, 570-1368, 571-1368, 572-1368 , 573-1368, 574-1368, 575-1368, 576-1368, 577-1368, 578-1368, 579-1368, 580-1368 , 581-1368, 582-1368, 583-1368, 584-1368, 585-1368, 586-1368, 587-1368, 588-1368 , 589-1368, 590-1368, 591-1368, 592-1368, 593-1368, 594-1368, 595-1368, 596-1368 , 597-1368, 598-1368, 599-1368, 600-1368, 601-1368, 602-1368, 603-1368, 604-1368 , 605-1368, 606-1368, 607-1368, 608-1368, 609-1368, 610-1368, 611-1368, 612-1368 , 613-1368, 614-1368, 615-1368, 616-1368, 617-1368, 618-1368, 619-1368, 620-1368 , 621-1368, 622-1368, 623-1368, 624-1368, 625-1368, 626-1368, 627-1368, 628-1368 , 629-1368, 630-1368, 631-1368, 632-1368, 633-1368, 634-1368, 635-1368, 636-1368 , 637-1368, 638-1368, 639-1368, 640-1368, 641-1368, 642-1368, 643-1368, 644-1368 , 645-1368, 646-1368, 647-1368, 648-1368, 649-1368, 650-1368, or 651-1368 In some embodiments, the split Cas9 protein may comprise either one of the following portions: The C-terminal portion of SpCas9 includes amino acids 574-1368 or 638-1368 of SpCas9.

[0242] "Subject" means a mammal, including a human, or a bovine, equine, canine, ovine, or Subjects include, but are not limited to, non-human mammals such as cats. Subjects include livestock, labor Domestic animals (cows, goats, chickens) that are raised to produce power and provide goods such as food , horses, pigs, rabbits, and sheep).

[0243] "Substantially identical" means that the amino acid sequence of a reference amino acid sequence (e.g., an amino acid sequence described herein) is substantially identical to the amino acid sequence of a reference amino acid sequence (e.g., an amino acid sequence described herein). any one of the sequences) or nucleic acid sequence (e.g., any one of the nucleic acid sequences described herein) It means a polypeptide or nucleic acid molecule that exhibits at least 50% identity to In this form, such sequences are identified at the amino acid or nucleic acid level as sequences used for comparison. The sequence may have at least 60%, 80%, or 85%, 90%, 95% or even 99% identity in the sequence.

[0244] Sequence identity is typically determined using sequence analysis software (e.g., Genetics Computer Group , University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705 Sequence Analysis Software Package, BLAST, BESTFIT, GAP, or PILE UP / PRETTYBOX program). Such software is available with various permutations. By assigning degrees of homology to the sequences, deletions, and / or other modifications, it is possible to determine whether the sequences are identical or Similar sequences are matched. Conservative substitutions typically include substitutions within the following groups: leucine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, Paragine, glutamine; serine, threonine; lysine, arginine; phenylalanine In an exemplary approach to determining the degree of identity, the BLAST program You can use the RAM, -3 and e -100 Probability scores between indicate closely related sequences COBALT is used, for example, with the following parameters: a) Alignment parameters: Gap penalties -11, -1 and End-Gap penalties -5, -1 b) CDD parameters: Use RPS BLAST on; Blast E-value 0.003; Find Conserved column s and Recompute on c) Query clustering parameters: Use query clusters on; Word Size 4; M ax cluster distance 0.8; Alphabet Regular. The EMBOSS Needle is used, for example, with the following parameters: a) Matrix: BLOSUM62; b) GAP OPEN: 10; c) GAP EXTEND: 0.5; d) OUTPUT FORMAT: pair; e) END GAP PENALTY: false; f) END GAP OPEN: 10; and g) END GAP EXTEND: 0.5.

[0245] The term "target site" refers to a sequence within a nucleic acid molecule that is modified by a nucleobase editor. In one embodiment, the target site is a sequence that is targeted by a deaminase (e.g., a cytidine or or adenine deaminase) or a fusion protein containing it .

[0246] As used herein, the terms "treat," "treating," and "treatment" refer to "Treatment" and the like are intended to alleviate or improve a disorder and / or its associated symptoms. refers to the process of achieving a desired pharmacological and / or physiological effect. Treating a condition requires the complete elimination of the associated disorder, condition, or symptom. It will be understood that some aspects of the In some instances, the effect is therapeutic, i.e., the effect is, but is not limited to, the effect of treating a disease. Partially or completely reduce, diminish, eliminate or alleviate the disease and / or adverse symptoms resulting therefrom. In some embodiments, the effect is preventative, i.e., The effect is to protect or prevent the occurrence or recurrence of a disease or condition. The disclosed methods comprise administering a therapeutically effective amount of a composition as described herein. include.

[0247] "Uracil glycosylase inhibitor," or "UGI," is a protein that inhibits the uracil excision repair system. In one embodiment, the agent inhibits host uracil-DNA glycosylation. A protein or fragment thereof that binds to uracil and prevents the removal of uracil residues from DNA. In one embodiment, the UGI inhibits uracil-DNA glycosylase base excision repair enzyme. In some embodiments, the target protein is a protein, fragment, or domain thereof that can inhibit the target protein. In some embodiments, the UGI domain comprises wild-type UGI or a modified version thereof. In the present invention, the UGI domain comprises a fragment of the exemplary amino acid sequence provided below. In some embodiments, the UGI fragment comprises at least 60% of the exemplary UGI sequences provided below: At least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least In some embodiments, the UGI comprises an amino acid sequence that is at least 99%, or 100%, of the amino acid sequence of As described below, amino acids homologous to exemplary UGI amino acid sequences or fragments thereof may be used. In some embodiments, the UGI or a portion thereof comprises a sequence as described below. For example, at least 70%, at least 75% of the wild-type UGI or UGI sequence or a portion thereof , at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9% or have 100% identity. Exemplary UGIs include the following amino acid sequences: >splP14739IUNGI_BPPB2 Uracil-DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSD APEYKPWALVIQDSNGEN KIKML.

[0248] The term "vector" refers to a nucleic acid sequence that is introduced into a cell, resulting in a transformed cell. Vectors include plasmids, transposons, phages, viruses, and ribosomal An "expression vector" includes a vector that is expressed in a recipient cell. An expression vector is a nucleic acid sequence containing a nucleotide sequence that encodes a gene encoding a gene. Enhances expression of introduced sequences, such as promoter, secretory, and other sequences. and / or may contain additional nucleic acid sequences for facilitating

[0249] The compositions and methods provided herein may be used in combination with other compositions and methods provided herein. It can be combined with any one or more of the methods.

[0250] DNA editing is a method to correct disease states by correcting pathogenic mutations at the genetic level. Until recently, all DNA editing platforms induces DNA double-strand breaks (DSBs) at specific genomic sites and is partially repaired depending on endogenous DNA repair pathways. It functions by determining product outcomes in a probabilistic manner, resulting in a complex population of genetic products. Achieving accurate and user-defined repair outcomes via the homology-directed repair (HDR) pathway Although it is possible to achieve high resolution imaging using HDR in therapeutically relevant cell types, many challenges remain. Efficient repair is hindered. In reality, this pathway involves competing, error-prone, and non-correlated pathways. HDR is less efficient than the homo-end joining pathway. Furthermore, HDR is strictly confined to the G1 and S phases of the cell cycle. This limits the ability of DSBs to be repaired accurately in post-mitotic cells. Highly efficient genome sequencing in these populations in a user-defined and programmable manner It has proven difficult or impossible to alter

[0251] The compositions and methods provided herein may be used in combination with other compositions and methods provided herein. It can be combined with any one or more of the methods.

[0252] DNA editing is a method to correct disease states by correcting pathogenic mutations at the genetic level. Until recently, all DNA editing platforms induces DNA double-strand breaks (DSBs) at specific genomic sites and is partially repaired depending on endogenous DNA repair pathways. It functions by determining product outcomes in a probabilistic manner, resulting in a complex population of genetic products. Achieving accurate and user-defined repair outcomes via the homology-directed repair (HDR) pathway Although it is possible to achieve high resolution imaging using HDR in therapeutically relevant cell types, many challenges remain. Efficient repair is hindered. In reality, this pathway involves competing, error-prone, and non-correlated pathways. HDR is less efficient than the homo-end joining pathway. Furthermore, HDR is strictly confined to the G1 and S phases of the cell cycle. This limits the ability of DSBs to be repaired accurately in post-mitotic cells. Highly efficient genome sequencing in these populations in a user-defined and programmable manner It has proven difficult or impossible to alter [Brief explanation of the drawings]

[0253] [Figure 1] Figures 1A-1C show plasmids. Figure 1A is an expression vector encoding the TadA7.10-dCas9 base editor. Figure 1B is a plasmid containing a nucleic acid molecule encoding a protein that confers chloramphenicol resistance (CamR) and spectinomycin resistance (SpectR). This plasmid also contains a kanamycin resistance gene disabled by two point mutations. Figure 1C is a plasmid containing a nucleic acid molecule encoding a protein that confers chloramphenicol resistance (CamR) and spectinomycin resistance (SpectR). This plasmid also contains a kanamycin resistance gene disabled by three point mutations. [Figure 2]Figure 2 shows an image of a bacterial colony transduced with the expression vector shown in Figure 1, which contains a nonfunctional kanamycin resistance gene. The vector contained an ABE7.10 variant generated using error-prone PCR. Bacterial cells expressing these "evolved" ABE7.10 variants were selected for kanamycin resistance using increasing concentrations of kanamycin. Bacteria expressing ABE7.10 variants with adenosine deaminase activity were able to correct the mutation introduced into the kanamycin resistance gene and restore kanamycin resistance. Kanamycin-resistant cells were selected for further analysis. [Figure 3] Figures 3A and 3B show editing of the regulatory region of the hemoglobin subunit gamma (HGB1) locus, a therapeutically relevant site for upregulation of fetal hemoglobin. Figure 3A is a diagram of a portion of the regulatory region of the HGB1 gene. Figure 3B quantifies the efficiency and specificity of adenosine deaminase variants. Editing was assayed at the hemoglobin subunit gamma 1 (HGB1) locus in HEK293T cells, which is a therapeutically relevant site for upregulation of fetal hemoglobin. The top panel shows nucleotide residues in the targeted region of the regulatory sequence of the HGB1 gene. A5, A8, A9, and A11 indicate the edited adenosine residues in HGB1. [Figure 4] Figure 4 shows the relative efficacy of adenosine base editors, including dCas9, that recognize non-canonical PAM sequences. The top panel shows the coding sequence for a hemoglobin subunit. The bottom panel shows the efficiency of adenosine deaminase variant base editors with guide RNAs of various lengths. [Figure 5] Figure 5 is a graph showing the efficiency and specificity of the ABE8 base editor, quantitating the percent editing at the intended target nucleotide and at unintended target nucleotides (bystanders). [Figure 6]Figure 6 is a graph showing the efficiency and specificity of the ABE8 base editor, quantitating the percent editing at intended target nucleotides and unintended target nucleotides (bystanders). [Figure 7] Figures 7A–7D show that eighth-generation adenine base editors mediate superior A·T to G·C conversion in human cells. Figure 7A shows an overview of adenine base editing: i) ABE8 creates an R-loop at the sgRNA targeting site in the genome; ii) TadA* deaminase chemically converts adenine to inosine through hydrolytic deamination of the ss-DNA portion of the R-loop; iii) Cas9 D10A nickase nicks the strand opposite the inosine-containing strand; iv) The inosine-containing strand can be used as a template during DNA replication; v) In the case of DNA polymerase, inosine preferentially base pairs with cytosine; and vi) After replication, inosine is replaced with guanosine. Figure 7B shows the architecture of ABE8.xm and ABE8.xd. Figure 7C shows three perspective views of E. coli TadA deaminase (PDB 1Z3A) aligned with S. aureus TadA (not shown) in complex with tRNA Arg2 (PDB 2B3J). Mutations identified over eight rounds of evolution are highlighted. Figure 7D is a graph showing the A·T to G·C base editing efficiency of the core ABE8 construct relative to the ABE7.10 construct across eight genomic sites in Hek293T cells. Values ​​and error bars reflect the mean and SD of three independent biological replicates performed on different days.

[0254] [Figure 8]Figures 8A–8C show that the Cas9 PAM variant ABE8 and the catalytically inactive Cas9 ABE8 variant mediate higher A·T to G·C conversion in human cells than the corresponding ABE7.10 variant. Values ​​and error bars reflect the mean and standard deviation (SD) of three independent biological replicates performed on different days. Figure 8A shows A·T to G·C conversion in Hek293T cells harboring NG-Cas9 ABE8s (-NG PAM). Figure 8B shows A·T to G·C conversion in Hek293T cells harboring Sa-Cas9 ABE8s (-NNGRRT PAM). Figure 8C shows A·T to G·C conversion in Hek293T cells harboring catalytically inactive dCas9-ABE8s (S. pyogenes Cas9 D10A, H840A). [Figure 9] Figures 9A-9E show a comparison of on-target and off-target editing frequencies between ABE7.10, ABEmax, and ABEmax with one BPNLS in Hek293T cells. Individual data points are shown for n=3 independent biological replicates performed on different days, and error bars represent standard deviation (SD). Figures 9A and 9B are graphs showing on-target DNA editing frequencies. Figures 9B and 9C are graphs showing the frequency of sgRNA-induced DNA off-target editing. Figure 9E is a graph showing RNA off-target editing frequencies. [Figure 10] Figures 10A-10B show the median A·T to G·C conversions and corresponding indel formation of TadA, C-terminal α-helical truncated ABE constructs in HEK293T cells. Figure 10A is a heatmap showing the median A·T to G·C editing conversions across eight genomic sites. Figure 10B is a heatmap showing indel formation. Delta residue values ​​correspond to the deletion position in TadA. Median values ​​generated from n=3 biological replicates. [Figure 11]Figure 11 is a heatmap showing the median A·T to G·C conversions of 40 ABE8 constructs across eight genomic sites in HEK293T cells. Medians were determined from two or more biological replicates. [Figure 12] Figure 12 is a heatmap showing the median % indels of 40 ABE8 constructs across eight genomic sites in HEK293T cells. Medians were determined from two or more biological replicates. [Figure 13] Figure 13 is a graph showing fold change in editing (ABE8:ABE7). Representation of average ABE8:ABE7 A·T→G·C editing across all A positions within targets at eight different genomic sites in Hek293T cells. Positions 2-12 indicate the location of the target adenine within the 20 nt protospacer, with position 20 immediately 5' to the -NGG PAM. [Figure 14] Figure 14 shows the ABE8 dendrogram, with the core ABE8 constructs selected for further study highlighted in black. [Figure 15] Figure 15 is a heatmap showing the median A⋅T to G⋅C conversions of the eight core ABE8 constructs across eight genomic sites in HEK293T cells. Medians were determined from three or more biological replicates.

[0255] [Figure 16] FIG. 16 is a heatmap showing the median indel frequencies of the eight core ABE8s tested at eight genomic sites in HEK293T cells. [Figure 17] Figure 17 is a heatmap showing the median A·T to G·C conversion of core NG-ABE8 construct 9 (-NG PAM) at six genomic sites in HEK293T cells. Median values ​​generated from n=3 biological replicates. [Figure 18]Figure 18 is a heatmap showing the median indel frequencies of core NG-ABE8 tested at six genomic sites in HEK293T cells. Median values ​​generated from n=3 biological replicates. [Figure 19] Figure 19 is a heatmap showing the median A·T to G·C conversion of the core Sa-ABE8 construct (-NNGRRT PAM) at six genomic sites in HEK293T cells. Site positions within the 22 nt protospacer are numbered from -2 to 20 (5' to 3'). Position 20 is 5' to the NNGRRT PAM. Median values ​​derived from n=3 biological replicates. [Figure 20] Figure 20 is a heatmap showing the median indel frequencies of core sa-ABE8 tested at 8 genomic sites in HEK293T cells. Median values ​​generated from n=3 biological replicates. [Figure 21] Figure 21 is a heatmap showing median A T to G C conversions of the core dC 9-ABE8-m construct at eight genomic sites in HEK293T cells. Dead Cas9 (dC 9) is defined as the D10A and H840A mutations in S. pyogenes Cas9. Median across three or more biological replicates. [Figure 22] Figure 22 is a heatmap showing median A T to G C conversions of the core dC9-ABE8-d construct at eight genomic sites in HEK293T cells. Dead Cas9 (dC9) is defined as the D10A and H840A mutations in S. pyogenes Cas9. Median values ​​generated from n≧3 biological replicates. [Figure 23]Figures 23A and 23B show the median indel frequencies of core dC9-ABE8 tested at eight genomic sites in HEK293T cells. Median values ​​generated from n≧3 biological replicates. Figure 23A is a heatmap showing indels indicated for the dC9-ABE8-m variant compared to ABE7.10. Figure 23B is a heatmap showing indels indicated for the dC9-ABE8-d variant compared to ABE7.10. [Figure 24] Figure 24 shows C·G to T·A editing by Hek293T cells treated with ABE8 and ABE7.10. Editing frequency for each site averaged across all C positions within the target. Cytosines within the protospacer are shaded. [Figure 25] Figures 25A-25H show DNA on-target editing and sgRNA-mediated DNA off-target editing by the ABE8 construct and an ABE8 construct with TadA mutations to improve DNA specificity. Individual data points are shown, and error bars represent the standard deviation (sd) for n=3 independent biological replicates performed on different days. Figures 25A and 25B are graphs showing the on-target DNA editing frequency for the core ABE8 construct compared to ABE7. Figures 25C and 25D are graphs showing the on-target DNA editing frequency for ABE8 with mutations that improve RNA off-target editing. Figures 25E and 25F are graphs showing the sgRNA-guided DNA off-target editing frequency for the core ABE8 construct compared to ABE7. Figures 25G and 25H are graphs showing the gRNA-guided DNA off-target editing frequency for ABE8 constructs with mutations that improve RNA off-target editing.

[0256] [Figure 26]Figure 26 is a graph showing indel frequencies at 12 previously identified sgRNA-dependent Cas9 off-target loci in human cells; individual data points are shown and error bars represent the s.d. for n=3 independent biological replicates performed on different days. [Figure 27] Figures 27A and 27B show A·T to G·C conversion in primary cells and the phenotypic results. Figure 27A is a graph showing A·T to G·C conversion at the -198 HBG1 / 2 site in ABE-treated CD34+ cells from two separate donors. NGS analysis performed 48 and 144 hours after treatment. The -198 HBG1 / 2 target sequence is shown with A7 highlighted. Percent A·T→G·C plotted for A7. Figure 27B is a graph showing the percentage of gamma globin formed as a percentage of alpha globin. Values ​​shown are from two different donors after ABE treatment and erythroid differentiation. [Figure 28] Figures 28A and 28B show A·T→G·C transversions at the −198 promoter site upstream of HBG1 / 2 in CD34+ cells treated with ABE8. Figure 28A is a heatmap showing the frequency of ABE8 A→G edits at 48 and 144 hours after editor treatment in CD34+ cells from two donors, where donor 2 is heterozygous for sickle cell disease. Figure 28B is a graphical representation of the distribution of total sequencing reads containing either the A7 edit alone or the (A7+A8) combined edit. [Figure 29] Figure 29 is a heat map showing indel frequencies at -198 in the gamma globin promoter in ABE8-treated CD34+ cells. Frequencies shown are from two donors at 48 and 144 hours. [Figure 30] FIG. 30 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of untreated differentiated CD34+ cells (donor 1). [Figure 31] FIG. 31 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 1) treated with ABE7.10-m. [Figure 32] FIG. 32 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 1) treated with ABE7.10-d. [Figure 33] Figure 33 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 1) treated with ABE8.8-m.

[0257] [Figure 34] FIG. 34 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 1) treated with ABE8.8-d. [Figure 35] FIG. 35 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 1) treated with ABE8.13-m. [Figure 36] FIG. 36 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 1) treated with ABE8.13-d. [Figure 37] FIG. 37 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 1) treated with ABE8.17-m. [Figure 38] FIG. 38 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 1) treated with ABE8.17-d. [Figure 39] FIG. 39 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 1) treated with ABE8.20-m. [Figure 40]FIG. 40 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 1) treated with ABE8.20-d. [Figure 41] Figure 41 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of untreated differentiated CD34+ cells (donor 2). Note: Donor 2 is heterozygous for sickle cell disease. [Figure 42] Figure 42 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 2) treated with ABE7.10-m. Note: Donor 2 is heterozygous for sickle cell disease. [Figure 43] Figure 43 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 2) treated with ABE7.10-d. Note: Donor 2 is heterozygous for sickle cell disease. [Figure 44] Figure 44 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 2) treated with ABE8.8-m. Note: Donor 2 is heterozygous for sickle cell disease. [Figure 45] Figure 45 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 2) treated with ABE8.8-d. Note: Donor 2 is heterozygous for sickle cell disease. [Figure 46] Figure 46 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 2) treated with ABE8.13-m. Note: Donor 2 is heterozygous for sickle cell disease.

[0258] [Figure 47]Figure 47 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 2) treated with ABE8.13-d. Note: Donor 2 is heterozygous for sickle cell disease. [Figure 48] Figure 48 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 2) treated with ABE8.17-m. Note: Donor 2 is heterozygous for sickle cell disease. [Figure 49] Figure 49 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 2) treated with ABE8.17-d. Note: Donor 2 is heterozygous for sickle cell disease. [Figure 50] Figure 50 shows the UHPLC UV-Vis trace (220 nm) and integration of globin chain levels of differentiated CD34+ cells (donor 2) treated with ABE8.20-m. Note: Donor 2 is heterozygous for sickle cell disease. [Figure 51] Figures 51A-51E show that editing by ABE8.8 at two independent sites reached over 90% editing before enucleation on day 11 after erythroid differentiation and approximately 60% gamma globin relative to alpha globin or total beta family globins on day 18 after erythroid differentiation. Figure 51A is a graph showing the average of ABE8.8 editing in two healthy donors in two independent experiments. Editing efficiency was measured with primers that distinguish between HBG1 and HBG2. Figure 51B is a graph showing the average of one healthy donor in two independent experiments. Editing efficiency was measured with primers that recognize both HBG1 and HBG2. Figure 51C is a graph showing ABE8.8 editing in a donor with a heterozygous E6V mutation. Figures 51D and 51E are graphs showing the increase in gamma globin in ABE8.8-edited cells. [Figure 52]Figures 52A and 52B show the percent editing (%Editing) using ABE variants to correct sickle cell mutations. Figure 52A is a graph showing screening of different editor variants with approximately 70% editing in SCD patient fibroblasts. Figure 52B is a graph showing CD34 cells from a healthy donor edited with the lead ABE variant, targeting a synonymous mutation A13 at the adjacent proline that lies within the editing window and serves as a proxy for editing of the SCD mutation. The ABE8 variant showed an average editing frequency of approximately 40% at proxy A13. [Figure 53] Figures 53A and 53B show RNA amplicon sequencing to detect intracellular A→I editing in RNA associated with ABE treatment. Individual data points are shown, and error bars represent the standard deviation (sd) for n=3 independent biological replicates performed on different days. Figure 53A is a graph showing the frequency of A→I editing in target RNA amplicons for the core ABE8 construct compared to ABE7 and Cas9 (D10A) nickase controls. Figure 53B is a graph showing the frequency of A→I editing in target RNA amplicons for ABE8 with mutations reported to improve RNA off-target editing.

[0259] [Figure 54] FIG. 54 is a schematic diagram showing the loss of dopamine resulting from dopaminergic neuron loss in Parkinson's disease. [Figure 55] FIG. 55 is a schematic diagram showing guide RNA and target sequences for correction of the R1441C and R1441H mutations in LRRK2 associated with Parkinson's disease. [Figure 56] FIG. 56 is a schematic diagram showing target sequences for correction of the Y1699C, G2019S, and I2020 mutations in LRRK2 associated with Parkinson's disease. [Figure 57]Figures 57A-C provide graphs, schematic diagrams, and tables. Figure 57A quantifies the percent A to G conversion at nucleic acid position 7 of the LRRK2 target sequence. The editors used are referred to as PV1-PV14 and are described below. pCMV indicates the CMV promoter; bpNLS indicates the bipartite nuclear localization signal; and monoABE8.1 indicates the monomeric form of the ABE8.1 base editor. Figure 57B shows the target sequence and guide RNA for correction of the R1441C mutation in LRRK2 associated with Parkinson's disease. Figure 57C shows the percent A to G conversion at nucleic acid position 7 of the LRRK2 target sequence. LRRK2 R1441C was edited using editors PV1-14. G2109 was edited using editors (15-28). The editors (PV1-28) used to correct the LRRK2 mutations are as follows: PV1 (also called PV15). pCMV_monoABE8.1_bpNLS + Y147TMSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD PV2 (also called PV16). pCMV_monoABE8.1_bpNLS + Y147RMSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCRFFRMPRQVFNAQKKAQSSTD PV3 (also called PV17). pCMV_monoABE8.pCMV_monoABE8.1_bpNLS + Q154SMS EVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRSVFNAQKKAQSSTD PV4 (also known as PV18). pCMV_monoABE8.1_bpNLS + Y123HMS EVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD PV5 (also known as PV19). pCMV_monoABE8.1_bpNLS + V82SMSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYSTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD PV6 (also called PV20). pCMV_monoABE8.1_bpNLS + T166RMSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSRD PV7 (also known as PV21). pCMV_monoABE8.PV8 (also called PV22). pCMV_monoABE8.1_bpNLS + Y147R_Q154R_Y123HMSEVEF SHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRRVFNAQKKAQSSTD PV9 (also called PV23). pCMV_monoABE8.1_bpNLS + Y147R_Q154R_I76YMSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLYDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCRFFRMPRRVFNAQKKAQSSTD PV10 (also called PV24). pCMV_monoABE8.1_bpNLS + Y147R_Q154R_T166RMSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCRFFRMPRRVFNAQKKAQSSRD PV11 (also called PV25). pCMV_monoABE8.PV12 (also called PV26). pCMV_monoABE8.1_bpNLS + Y147T_Q154RMSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRRVFNAQKKAQSSTD PV13 (also called PV27). pCMV_monoABE8.1_bpNLS + H123Y123H_Y147R_Q154R_I76YMSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLYDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLCRFFRMPRRVFNAQKKAQSSTD PV14 (also called PV28). pCMV_monoABE8.1_bpNLS + V82S + Q154RMSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYSTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRRVFNAQKKAQSSTD.

[0260] [Figure 58]Figures 58A-C provide graphs, schematic diagrams, and tables. Figure 58A quantifies the percent A to G conversion at nucleic acid position 6 of the LRRK2 target sequence. The editors used were PV15-PV28, the descriptions of which are provided above. pCMV indicates the CMV promoter; bpNLS indicates the bipartite nuclear localization signal; and monoABE8.1 indicates the monomeric form of the ABE8.1 base editor. Figure 58B shows the target sequence and guide RNA for correction of the G2019S mutation in LRRK2 associated with Parkinson's disease. Figure 58C shows the percent A to G conversion at nucleic acid positions 4 and 6 of the LRRK2 target sequence. The A to G transition at position 4 is a bystander effect. [Figure 59] Figures 59A-59L show the sequence read for the A to G transition at position 7 of the LRRK2 target sequence encoding R1441C (see Figures 57A-57C). Editors are displayed (PV1-14). Descriptions of PV1-28 are provided in Figure 56. [Figure 60] Figures 60 A-60 W show the sequence reads for the A to G transition at positions 4 and 6 of the LRRK2 target sequence encoding G2019S (see Figures 58 A-58 C). [Figure 61]Figure 61A provides a schematic diagram showing the target sequence for correction of the pathogenic mutation A419V in LRRK2, which is encoded by an antisense strand G>A mutation. The mutation is corrected using an ABE targeting A at position 12 using an SpCas9 variant with specificity for the TGG PAM. Figure 61B is a schematic diagram showing the target sequence for correction of the pathogenic mutation L1114L in LRRK2 associated with Parkinson's disease. This mutation is antisense strand T>C and is corrected using a base editor (CBE) with cytidine deaminase activity. Figure 61C is a schematic diagram showing the target sequence for correction of the pathogenic mutation I1122V in LRRK2 associated with Parkinson's disease. This mutation is antisense strand T>C and is corrected using a base editor (CBE) with cytidine deaminase activity. Figure 61D provides a schematic diagram showing the target sequence for correction of the pathogenic mutation M1869V in LRRK2 associated with Parkinson's disease. This mutation is T>C in the antisense strand and is corrected using a base editor with cytidine deaminase activity (CBE). [Figure 62] Figures 62A and 62B show precise base editing correction of the Mus musculus IDUA W401X mutation in HEK293T cells. Figure 62A is a graph showing the percentage of base editing of the Mus musculus IDUA W401X mutation using an ABE8 base editor variant with a 21-nucleotide guide RNA. Figure 62B is a graph showing the percentage of indels for an ABE8 base editor variant with a 21-nucleotide guide RNA. [Figure 63] Figure 63 is a graph showing the rate of base editing of the Mus musculus IDUA W401X mutation using ABE8 base editor variants using either a 20-nucleotide guide RNA or a 21-nucleotide guide RNA. [Figure 64]Figure 64 shows a schematic diagram of the Homo sapiens IDUA genomic nucleic acid and amino acid sequence as a target for A-to-G nucleotide base editing to correct the W402X mutation. The figure also shows the nucleic acid sequence of the corresponding guide RNA (gRNA). The figure also shows the target adenosine (A) nucleobase (boxed) in the IDUA nucleic acid sequence. [Figure 65] Figures 65A and 65B show precise base editing correction of the Homo sapiens IDUA W402X mutation in HEK293T cells. Figure 65A is a graph showing the percentage of base edits of the Homo sapiens IDUA W402X mutation using an ABE8 base editor variant with a 20-nucleotide guide RNA. Figure 65B is a graph showing the percentage of indels for an ABE8 base editor variant with a 20-nucleotide guide RNA. [Figure 66] Figures 66A-66O are tables showing the percentage efficiency of A to G nucleotide changes in IDUA nucleic acid sequences using ABE8 base editor variants, as detected by PCR of genomic DNA in base-edited cells followed by deep sequencing (MySeq). Figures 66A-66M show the percent A to G base editing at position 6 of the IDUA nucleic acid target site using three samples each of the ABE8 base editor variants ABE8.1-ABE8.13. Figure 66N shows the percent A to G base editing at position 6 of the IDUA nucleic acid target site using three samples of the positive control base editor ABE7.10. Figure 66O shows the percent A to G base editing at position 6 of the IDUA nucleic acid target site using two samples of the negative control. [Figure 67] Figure 67 shows Rett / MECP2: mutation correction. MECP2 loss of function can result from many different de novo mutations. X-linked: XX patients are mosaic for MECP2 deficiency; XY usually results in infantile lethality. [Figure 68]Figure 68 shows the Rett syndrome R106W mutation correction for the top 3 Guide sequences. [Figure 69] Figure 69 shows Rett syndrome R255X mutation correction using an editor with NGTT PAM optimization. [Figure 70] Figures 70A-C: Hurler / IDUA mutation correction. Figure 70A shows the experimental design for IDUA W402X mutation correction. Figure 70B shows the percent editing of each editor construct. Figure 70C shows the specific activity (nmol / mg / h) for edited and unedited constructs. [Figure 71] Figure 71 shows in vivo base editing by ABE8.8. For each sample, from left to right, guide 11 (AAV9), guide 12 (AAV9), guide 11 (PHP.eB), guide 12 (PHP.eB), and control.

[0261] [Figure 72] Figures 72A-72B. A·T to G·C conversion by the ABE7.10 and ABE8 variants at the ABCA4 G1961E allele in model cell lines. Figure 72A: A·T to G·C conversion at the integrated disease allele and wobble base at the ABCA4 G1961E codon in HEK293T cells after plasmid lipofection of a 21-nt spacer sgRNA and base editor variant. Cells incubated for 5 days after lipofection were assessed for editing. Figure 72B: DNA sequence at the site of interest, including the ABCA4 G1961E disease allele, the wobble base at the codon, and the -NGG PAM used by the 21-nt spacer sgRNA. Error bars represent the standard deviation of three replicates. In each dataset, the disease allele is on the left and the wobble base is on the right. [Figure 73]A·T to G·C conversion by sgRNA spacer length variants at the ABCA4 G1961E allele in a model cell line. A·T to G·C conversion at the integrated disease allele and wobble base of the ABCA4 G1961E codon in HEK293T cells after lipofection of sgRNAs with various spacer lengths and the ABE7.10 plasmid. Cells incubated for 5 days after lipofection were assessed for editing. hRz = sgRNA containing a self-cleaving hammerhead ribozyme at the 5' end. Error bars represent the standard deviation of three replicates. In each dataset, the disease allele is on the left and the wobble base is on the right. [Figure 74] Schematic of dual AAV delivery of a split base editor using split intein reconstitution. Two AAV particles are separately packaged with the components required for base editing. One virus encodes the C-terminal region of the base editor with an N-terminal split-intein fusion, and the complementary virus encodes the N-terminal region of the base editor with a C-terminal split-intein fusion along with an sgRNA. Co-transduction of the complementary viruses transcribes the sgRNA, and each half of the base editor is expressed and recombined through split-intein-mediated protein trans-splicing. [Figure 75]Figures 75A-75B. A·T to G·C conversion at ABCA4 G1961 in wild-type cells by dual AAV delivery of the split ABE variant. Figure 75A: A·T → G·C and C·G → T·A transversions at the wild-type ABCA4 G1961 target site in wild-type ARPE-19 cells, where editing at position 8A serves as a surrogate target for editing in these cells. Cells were infected at an MOI of 5E+4 viral genomes / virus / cell. Cells were incubated for 2 weeks post-infection and assessed for editing. Error bars represent the standard deviation of six replicates. For each data point, samples treated with the position 8 (A>G)-surrogate site are shown on the left, and position 5 (C>T) is shown on the right. Figure 75 B: DNA sequence at the wild-type target site, including the ABCA4 G 1961 allele and the -NGG PAM used by the 21 nt spacer sgRNA targeting the wild-type sequence. [Figure 76] Figures 76A-76B. Off-target base editing in wild-type ARPE-19 cells dually infected with AAV2 expressing split ABE7.10 and sgRNA targeting the ABCA4 G1961E disease allele. Figure 76A: Maximum A·T → G·C conversion across the on-target (On Targ) or off-target (OT) protospacer 2 weeks after co-infection with dual AAV (blue) compared to untreated control (gray). Figure 76B: Maximum non-A·T → G·C conversion across the on-target (On Targ) or off-target (OT) protospacer 2 weeks after co-infection with dual AAV (blue) compared to untreated control (gray). For each data point, treated wild-type (wt) ARPE-19 cells are shown on the left, and untreated wt ARPE-19 cells are shown on the right. [Figure 77]Indel formation by base editing in wild-type ARPE-19 cells co-infected with AAV2 expressing split ABE7.10 and an sgRNA targeting the disease allele of ABCA4 G1961E. Percentage of indels formed within or proximal to the on-target or off-target protospacer 2 weeks after co-infection with the dual AAV (blue) compared to untreated controls (gray). For each data point, treated wild-type (wt) ARPE-19 cells are shown on the left, and untreated wt ARPE-19 cells are shown on the right. [Figure 78] Primate retinal integrity and GFP expression at 22 days after culture. Sections were immunolabeled overnight at 4°C with anti-rhodopsin, anti-GFP, and biotinylated peanut agglutinin (PNA) antibodies. Anc80L65.hGRK.eGFP demonstrated that GFP was observed exclusively in the photoreceptor-containing outer nuclear layer (ONL), confirming photoreceptor-specific activity of the GRK promoter. The top row represents untransduced cells at day 0. The second row represents untransduced cells at day 22. The third row represents day 22 GRK. The fourth row represents day 22 CMB. Columns represent unstained (1st column), DAPI (2nd column), GFP (3rd column), PNA (4th column), and rhodopsin (5th column). [Figure 79] Cas9 expression in NHPs. Cas9 expression is detected in primate retina as early as 6 days after culture. Results are shown for ABE7.10 (columns 1 and 2), ABE8.5 (columns 2 and 3), and ABE8.9 (columns 3 and 4). Top row: Day 6 after culture. Bottom row: Day 17 after culture. Results demonstrate that the AAV system delivers split intein expressing Cas9. Scale bar: 100 μm. DETAILED DESCRIPTION OF THE INVENTION

[0262] The present invention provides novel adenine nucleotide sequences with increased efficiency for generating targeted nucleic acid base sequence modifications. Compositions comprising base editors (e.g., ABE8) and methods of using them are provided.

[0263] [Nucleobase Editor] A base editor for editing, modifying or altering a target nucleotide sequence of a polynucleotide. Disclosed herein are nucleic acid base editors or nucleobase editors. Polynucleotide-programmable nucleotide-binding domains (e.g., Cas9) and nucleic acids Nucleic acid base editors or salts containing a base editing domain (e.g., adenosine deaminase) Polynucleotide programmable nucleotide binding domains (e.g., Cas9), when combined with an attached guide polynucleotide (e.g., gRNA), (Complementary base pairing between the bases of the bound guide nucleic acid and the bases of the target polynucleotide sequence capable of specifically binding to a target polynucleotide sequence (through formation of a and localizing the base editor to the target nucleic acid sequence desired to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. Polynucleotide sequences include DNA-RNA hybrids.

[0264] Polynucleotide-programmable nucleotide-binding domains Polynucleotide programmable nucleotide binding domains also bind to nuclear RNA. It is understood that the present invention may include an acid programmable protein. The polynucleotide-programmable nucleotide binding domain is The nucleotide-binding domain can be linked to a nucleic acid that guides the RNA. Other DNA-binding proteins are also within the scope of this disclosure, although they are not specifically listed in this disclosure. It has not been done.

[0265] The polynucleotide-programmable nucleotide-binding domain of the base editor is The polynucleotide programmable vector itself can contain one or more domains. The nucleotide-binding domain capable of binding to the nuclease may include one or more nuclease domains. In some embodiments, the nucleic acid of a polynucleotide-programmable nucleotide binding domain The nuclease domain can comprise an endonuclease or an exonuclease. As used herein, the term "exonuclease" refers to an enzyme that liberates nucleic acids (e.g., RNA or DNA). The term "endonuclease" refers to a protein or polypeptide that can be digested from its termini. A "clease" is a nucleic acid that can catalyze (e.g., cleave) an internal region of a nucleic acid (e.g., DNA or RNA). In some embodiments, an endonuclease refers to a protein or polypeptide that It is capable of cleaving a single strand of double-stranded nucleic acid. Both strands of a double-stranded nucleic acid molecule can be cleaved. The programmable nucleotide binding domain can be a deoxyribonuclease. In some embodiments, the polynucleotide programmable nucleotide binding domain is a ribonucleotide. It may be a nuclease.

[0266] In one embodiment, the nucleotide of the polynucleotide-programmable nucleotide binding domain The cleavage domain can cleave zero, one, or two strands of a target polynucleotide. In one embodiment, the polynucleotide-programmable nucleotide binding domain The amino acid may comprise a nickase domain. " refers to a nucleic acid that can cleave only one of the two strands in a double-stranded nucleic acid molecule (e.g., DNA). a polynucleotide-programmable nucleotide-binding domain containing a cleavage domain; In some embodiments, the nickase is an active polynucleotide programmable nuclease. By introducing one or more mutations into the nucleotide binding domain, polynucleotide protease activity can be increased. Derived from a fully catalytically active (e.g., native) form of a programmable nucleotide-binding domain For example, polynucleotide programmable nucleotide binding domains can be used. If the gene contains a nickase domain derived from Cas9, the nickase domain derived from Cas9 The protein may contain a D10A mutation and a histidine at position 840. In embodiments, residue H840 retains catalytic activity, thereby cleaving a single strand of a nucleic acid duplex. In another example, the nickase domain from Cas9 contains an H840A mutation. while the amino acid residue at position 10 remains D. In some embodiments, the nickase may comprise all of the nuclease domain that is not required for nickase activity. or by removing a portion of the polynucleotide, programmable nucleotide binding It can be derived from a fully catalytically active (e.g., native) form of the domain. A nucleotide-programmable nucleotide-binding domain derived from Cas9 is used to identify nickases If the domain is included, the nickase domain from Cas9 is a RuvC domain or an HNH domain. It may contain a deletion of all or part of the domain.

[0267] The amino acid sequence of an exemplary catalytically active Cas9 is as follows: MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD.

[0268] Thus, polynucleotides containing nickase domains can be used as programmable nucleotides. Base editors containing a base-binding domain can be used to bind to specific polynucleotide target sequences (e.g., binding domains). The DNA fragments are then ligated to the target nucleic acid, which then generates a single-stranded DNA break (a nick) at the target site (determined by the complementary sequence of the guide nucleic acid). In some embodiments, a nickase domain (e.g., a nickase domain derived from Cas9) can be used. Nucleic acid double-stranded target polynucleotides cleaved by base editors containing a base editor domain (a cleavage enzyme domain) The strand of the base sequence is the strand that is not edited by the base editor (i.e., (The strand cleaved by the cleavage is the opposite strand to the strand containing the base to be edited.) and base editors containing a nickase domain (e.g., a nickase domain derived from Cas9). can cleave the strand of a DNA molecule targeted for editing. In this state, the non-target strand is not cleaved.

[0269] Catalytically dead (i.e., unable to cleave target polynucleotide sequence) polynucleotides Base editors comprising nucleotide-programmable nucleotide-binding domains are also described herein. As used herein, the terms "catalytically dead" and "nuclease inactive" has one or more mutations and / or deletions that result in the inability to cleave a strand of nucleic acid. Polynucleotides with deletions are replaced to refer to programmable nucleotide binding domains. In some embodiments, catalytically dead polynucleotide proteases are used. A gram-capable nucleotide-binding domain base editor binds one or more nuclease domains Nuclease activity can be lost as a result of specific point mutations in the In the case of base editors containing the Cas9 domain, Cas9 is able to reverse the D10A and H840A mutations. Such a mutation would inactivate both nuclease domains. In another embodiment, catalytically dead polynucleotides are The nucleotide-programmable nucleotide-binding domain is a catalytic domain (e.g., RuvC1 and and / or the HNH domain). In embodiments, catalytically dead polynucleotide programmable nucleotide binding domains The mutants contain point mutations (e.g., D10A or H840A) as well as all or part of the nuclease domain. Contains some deletions.

[0270] Also provided herein are the following polynucleotide-programmable nucleotide binding domains: Catalytically dead polynucleotide programmable nucleic acid from a previously functioning version Mutations that can generate peptide-binding domains are also contemplated. In the case of dead Cas9 ("dCas9"), mutations other than D10A and H840A are present, leading to the nuclease Variants that result in inactive Cas9 are provided. Such mutations include, for example, D10 and other amino acid substitutions at H840 or other substitutions within the nuclease domain of Cas9. (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain) Additional suitable nuclease-inactive dCas9 domains are described in this disclosure and in the art. Such modifications would be apparent to those skilled in the art based on their knowledge and are within the scope of the present disclosure. Exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to: Contains D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (e.g., Prashant et al., CAS9 transcriptional activators for target specific ficity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013; 31(9): 833-838, the entire contents of which are incorporated herein by reference. incorporated into the book).

[0271] Polynucleotide programmable nucleotides that can be incorporated into base editors Non-limiting examples of binding domains include domains derived from CRISPR proteins, restriction nucleases, and the like. Enzymes, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases In some embodiments, the base editor is a CRISP gene for nucleic acids. R (i.e., Clustered Regularly Interspaced Short Palindromic Repeats)-mediated modifications In this case, natural or modified nucleic acids that can bind to a nucleic acid sequence via a binding guide nucleic acid are used. Polynucleotide-programmable nucleotide-binding domains containing proteins or portions thereof Such proteins are referred to herein as "CRISPR proteins." Thus, disclosed herein are methods for detecting CRISPR proteins, including all or part of the proteins. Polynucleotide base editors containing programmable nucleotide binding domains (i.e. That is, a base editor that contains all or part of the CRISPR protein as a domain (which is It is also called the "CRISPR protein-derived domain" of the base editor. The CRISPR protein-derived domain integrated into the CRISPR receptor is expressed as a wild-type or naturally occurring CRISPR protein. For example, as described below, CRISPR proteins can be modified relative to the CRISPR protein. The CRISPR domain may have one or more mutations, as compared to a wild-type or naturally occurring CRISPR protein, It may involve insertions, deletions, rearrangements and / or recombinations.

[0272] CRISPR provides defense against mobile genetic elements (viruses, transposable elements, conjugative plasmids) CRISPR clusters are composed of spacers, sequences that are complementary to the preceding mobile element. The CRISPR cluster contains a sequence of target nucleic acids, and the target invader nucleic acid. The CRISPR cluster is transcribed into CRISPR RNA (crRNA). In type II CRISPR systems, correct processing of pre-crRNA is essential for transcription. transcoding small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein tracrRNA guides the processing of pre-crRNA by ribonuclease 3. Cas9 / crRNA / tracrRNA then binds to a linear or circular dsDNA target complementary to the spacer. The target strand that is not complementary to the crRNA is first cleaved by the endonuclease. It is cleaved protease-wise and then exonucleolytically trimmed to 3'-5'. In the case of mitochondrial DNA, both the protein and the RNA are required for DNA binding and cleavage. A single guide RNA ("sgRNA", For example, Jinek M. et al., Science 337: 816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 targets a short motif (PAM or protospacer adjacent motif) in the CRISPR repeats. It helps us to recognize and distinguish between "self" and "non-self."

[0273] In some embodiments, the methods described herein involve recombinantly engineered The guide RNA (gRNA) is required for Cas binding. The required scaffold sequence and a user-defined approximately 20-base spacer that defines the genomic target to be modified. Therefore, those skilled in the art can identify the genomic target of Cas protein specificity. The gRNA targeting sequence can be varied to target the genome relative to the rest of the genome. This is determined in part by how specific it is to the target.

[0274] In some embodiments, the gRNA scaffold sequence is: GUUUUAGAGC UAGAAAU AGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAAGU GGCACCGAGU CGGUGCUUUU.

[0275] In some embodiments, the base editor comprises a CRISPR protein-derived domain. The main component is capable of binding to a target polynucleotide when combined with a binding guide nucleic acid. Endonucleases (e.g., deoxyribonucleases or ribonucleases) that can In some embodiments, the CRISPR protein incorporated into the base editor The protein-derived domain binds to the target polynucleotide when combined with the binding guide nucleic acid. In some embodiments, the base editor is a nickase that can bind to The CRISPR protein-derived domains integrated into the CRISPR protein bind to the CRISPR protein when combined with the guide nucleic acid. A catalytically dead domain is capable of binding to a target polynucleotide when In embodiments, a target polynucleotide that binds to a CRISPR protein-derived domain of a base editor is The nucleic acid is DNA, and in some embodiments, the nucleic acid comprises a CRISPR protein-derived domain of a base editor. The target polynucleotide that binds to the in is RNA.

[0276] The CAs proteins that can be used herein include class 1 and class 2. Non-limiting examples of proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5d, Cas5t, as5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also called Csn1 or Csx12), Cas10, Csy1, C sy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5, Csn1, Csn2, Csm2, Csm3, Csm4, Csm 5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Cssx16, Cx16, Cx, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1 , Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i, CARF, DinG, and their homologs Unmodified CRISPR enzymes, like Cas9, are composed of two It has functional endonuclease domains, RuvC and HNH, and can have DNA cleavage activity. CRISPR enzymes target sequences, such as within the target sequence and / or the complementary strand of the target sequence. For example, CRISPR enzymes can induce cleavage of one or both strands of a target sequence. Approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 25 nucleotides from the first or last nucleotide of the string Induce breaks in one or both strands at 50, 100, 200, 500 base pairs or more. It is possible.

[0277] Loss of ability to cleave one or both strands of a target polynucleotide containing the target sequence To achieve this, vectors were used to encode CRISPR enzymes that were mutated relative to the corresponding wild-type enzyme. Cas9 can be a wild-type exemplary Cas9 polypeptide (e.g., Cas9 from S. pyogenes). Cas9) and at least approximately 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93% %, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology Cas9 can refer to a polypeptide having a wild-type exemplary Cas9 polypeptide ( For example, from S. pyogenes), at most, or at most, approximately, about 50%, 60%, 70% , 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or polypeptides having sequence homology. or deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any of these It can refer to modified forms of the Cas9 protein that may contain amino acid changes, such as combinations.

[0278] In some embodiments, the CRISPR protein-derived domain of the base editor is Corynebacterium diphtheria (NCBI Refs: NC_015683.1, NC_017317.1); NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021 284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (N CBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Belliella ba ltica (NCBI Ref: NC_018010.1); Psychroflexus torquis (NCBI Ref: NC_018721.1); St reptococcus thermophilus (NCBI Ref: YP_820832.1); Listeria innocua (NCBI Ref: NP _472073.1); Campylobacter jejuni (NCBI Ref: YP_002344900.1); Neisseria meningiti dis (NCBI Ref: YP_002342100.1), Streptococcus pyogenes, or Staphylococcus au The Cas9 fragment may comprise all or part of the Cas9 fragment derived from Cas9.

[0279] [Cas9 domain, a nucleobase editor] The sequence and structure of Cas9 nuclease are well known to those skilled in the art (see, e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes.” Ferretti et al., Pr oc. Natl. Acad. Sci. USA 98:4658-4663(2001); “CRISPR RNA maturation by trans -encoded small RNA and host factor RNase III.” Deltcheva E. et al., Nature 471: 602-607(2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptation See Jinek M. et al., Science 337:816-821 (2012). The entire contents of which are incorporated herein by reference.) Cas9 orthologs include, but are not limited to: It has been described in various species, including S. pyogenes and S. thermophilus, although not exclusively in the Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on the present disclosure. Such Cas9 nucleases and sequences are described in Chylinski, Rhun, and Charpentier, “ The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) R NA Biology 10:5, 726-737. The entire contents of which are incorporated herein by reference.

[0280] In some embodiments, nucleic acid programmable DNA binding proteins (napDNAbp) is a Cas9 domain. Non-limiting exemplary Cas9 domains are provided herein. The 9 domains are divided into nuclease-active Cas9 domains, nuclease-inactive Cas9 domains (dCas9 ), or Cas9 nickase (nCas9). In some embodiments, the Cas9 domain can be Nuclease activity domains. For example, the Cas9 domain binds both strands of a double-stranded nucleic acid (e.g., For example, it may be a Cas9 domain that cleaves both strands of a double-stranded DNA molecule. In some embodiments, the Cas9 domain comprises any one of the amino acid sequences described herein. In some embodiments, the Cas9 domain comprises any of the amino acid sequences described herein. At least 60%, at least 65%, at least 70%, at least 75%, or at least at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least or at least 97%, at least 98%, at least 99%, or at least 99.5% identical amino acid sequence. In some embodiments, the Cas9 domain comprises an amino acid sequence described herein. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 5, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 3 5, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more In some embodiments, the Cas9 domain comprises an amino acid sequence having a mutation. At least 10, at ... At least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 20 0, at least 250, at least 300, at least 350, at least 400, at least 500, a few At least 600, at least 700, at least 800, at least 900, at least 1000, at least amino acid sequences that have at least 1100 or at least 1200 identical stretches of amino acid residues Includes.

[0281] In some embodiments, proteins comprising fragments of Cas9 are provided. In embodiments, the protein includes one of the following two Cas9 domains: (1) Cas9 (2) a gRNA binding domain of Cas9; and (3) a DNA cleavage domain of Cas9. In some embodiments, Cas9 or a fragment thereof Proteins containing the fragments are referred to as "Cas9 variants." Cas9 variants are Cas9 or For example, a Cas9 variant may share at least about 70% homology with a wild-type Cas9. % identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, In some embodiments, the sequences are about 99.5% identical, or at least about 99.9% identical to each other. Cas9 mutants have the following advantages compared to wild-type Cas9: 4, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 3 4, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more In some embodiments, the Cas9 variant may have an amino acid change in the fragment of Cas9. The fragment contains a fragment of wild-type Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain), At least about 70% identical to the corresponding fragment, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least In some embodiments, the fragment is about 99.5% identical, or at least about 99.9% identical. , at least 30%, at least 35%, at least 40% of the amino acid length of the corresponding wild-type Cas9; At least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least or at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least In some embodiments, the fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 50 0, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 12 50, or at least 1300 amino acids in length.

[0282] In some embodiments, the Cas9 fusion proteins provided herein comprise a Cas9 protein. The full-length amino acid sequence of the protein, for example, one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein contain a full-length Cas9 sequence. Examples of Suitable Cas9 Domains and Cas9 Fragments Suitable amino acid sequences are provided herein, and further suitable sequences for Cas9 domains and fragments are , will be apparent to those skilled in the art.

[0283] The Cas9 protein guides the protein to a specific DNA sequence complementary to its guide RNA. In one embodiment, the polynucleotide protease is capable of binding to a guide RNA. The activatable nucleotide binding domain can be a Cas9 domain, e.g., a nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Examples of tunable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), CasX, C These include, but are not limited to, asY, Cpf1, Cas12b / C2C1, and Cas12c / C2C3. In one embodiment, the wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes ( NCBI reference sequence: NC_017053.1, nucleotide and amino acid sequences are as follows:

[0284] ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGATTATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG GCAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGCAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAATCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTAGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAGAAATGGCTTGTTTGGGAAT CTCATTGCTTTGTCATTGGGATTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATAGTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAGCGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAGGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGCGCCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGGGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGATATTCAAAAAGCACAGGTGTCTGGACAAGGCCATAGTTTA CATGAACAGATTGCTAACTTAGCTGGCAGTCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAATTGTTGATGAACT GGTCAAAGTAATGGGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCCAGA AAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCTGTT GAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTACAAAATGGAAGAGACATGTATGTGGACCAAGAATT AGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCATTAAAGACGATTCAATAGACAATA AGGTACTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAAAAC TATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGGTTT GAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAATTT TGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCTAAA TTAGTTTCTGACTTCCGAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGCGTA TCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATAAAG TTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCTAAT ATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGGGGA AACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATATTG TCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTACCAAAGAAATTCGGACAAGCTTATTGCT CGTAAAAAGACTGGGATCCAAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGCTAA GGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCTTTG AAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATAAGGAAGTTAAAAGACTTAATCATTAAACTACCTAATAT AGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAATTACAAAAAGGAAATGAGCTGGCTCT GCCAAGCAAATGTGAATTTTTTATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAACAAA AACAATTGTTTGTGGAGCAGCATAAGCATTATTTGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTTATT TTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGAAAA TATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTAAAC GATATACGTCTACAAAGAAGTTTTAGATGCCACTCTTATCCATCCATCCATGGTCTTTGAAACACGCATTGAT TTGAGTCAGCTAGGAGGTGACTGA JPEG2025032080000009.jpg168169(Picture:HNHドメイン;Picture:RuvCドメイン)

[0285] In some embodiments, wild-type Cas9 contains the following nucleotides and / or amino acids: corresponding to or including the amino acid sequence: ATGGATAAAAAGTATTCTATTGGTTTAGACATCGGCACTAATTCCGTTGGATGGGCTGTCATAACCGATGAATACAAAGT ACCTTCAAAGAAATTTAAGGTGTTGGGGAACACAGACCGTCATTCGATTAAAAAAGAATCTTATCGGTGCCCTCCTATTCG ATAGTGGCGAAACGGCAGAGGCGACTCGCCTGAAACGAACCGCTCGGAGAAGGTATACACGTCGCAAGAACCGAATATGT TACTTACAAGAAATTTTTAGCAATGAGATGGCCAAAGTTGACGATTCTTTCTTTCACCGTTTGGAAGAGTCCTTCCTTGT CGAAGAGGACAAGAAACATGAACGGCACCCCATCTTTGGAAACATAGTAGATGAGGTGGCATATCATGAAAAGTACCCAA CGATTTATCACCTCAGAAAAAAGCTAGTTGACTCAACTGATAAAGCGGACCTGAGGTTAATCTACTTGGCTCTTGCCCAT ATGATAAAGTTCCGTGGGCACTTTCTCATTGAGGGTGATCTAAATCCGGACAACTCGGATGTCGACAAACTGTTCATCCA GTTAGTACAAACCTATAATCAGTTGTTTGAAGAGAACCCTATAAAATGCAAGTGGCGTGGATGCGAAGGCTATTCTTAGCG CCCGCCTCTCTAAATCCCGACGGCTAGAAAACCTGATCGCACAATTACCCGGAGAGAAGAAAAAATGGGTTGTTCGGTAAC CTTATAGCGCTCTCACTAGGCCTGACACCAAATTTAAGTCGAACTTCGACTTAGCTGAAGATGCCAAATTGCAGCTTAG TAAGGACACGTACGATGACGATCTCGCAATCTACTGGCACAAATTGGAGATCAGTATGCGGACTTATTTTTGGCTGCCA AAAACCTTAGCGATGCAATCCTCCTATCTGACATACTGAGAGTTAATACTGAGATTACCAAGGCGCCGTTATCCCGCTTCA ATGATCAAAAGGTACGATGAACATCACCAAGACTTGACACTTCTCAAGGCCCTAGTCCGTCAGCAACTGCCTGAGAAATA TAAGGAAATATTCTTTGATCAGTCGAAAACGGGTACGCAGGTTATATTGACGGGCGGAGCGAGTCAAGAGGAATTCTACA AGTTTATCAAACCCATATTAGAGAATGGATGGGACGGAAGAGTTGCTTGTAAAACTCAATCGCGAAGATCTACTGCGA AAGCAGCGGACTTTCGACAACGGTAGCATTCCACATCAAATCCACTTAGGCGAATTGCATGCTATACTTAGAAGGCAGGA GGATTTTTATCCGTTCCTCAAAGACAATCGTGAAAAGATTGAGAAAATCCTAACCTTTCGCATACCTACTATGTGGGAC CCCTGGCCGAGGGAACTCTCGGTTCGCATGGATGACAAGAAAGTCCGAAGAAACGATTACTCCATGGAATTTTGAGGAA GTTGTCGATAAAGGTGCGTCAGCTCAATCGTTCATCGAGAGGATGACCACTTTGACAAGAATTTACCGAACGAAAAAGT ATTGCCTAAGCACAGTTTACTTTACGAGTATTTCACAGTGTACAATGAACTCACGAAAGTTAAGTATGTCACTGAGGGCA TGCGTAAACCCGCCTTTCTAAGCGGAGAACAGAAGAAAGCAATAGTAGATCTGTTATTCAAGACCAACCGCAAAGTGACA GTTAAGCAATTGAAAGAGGACTACTTTAAGAAAATTGAATGCTTCGATTCTGTCGAGATCTCCGGGGTAGAAGATCGATT TAATGCGTCACTTGGTACGTATCATGACCTCCTAAAGATAATTAAAGATAAGGACTTCCTGGATAACGAAGAGAATGAAG ATATCTTAGAAGATATAGTGTTGACTCTTACCCTCTTTGAAGATCGGGAAATGATTGAGGAAAGACTAAAAACATACGCT CACCTGTTCGACGATAAGGTTATGAAACAGTTAAAGAGGCGTCGCTATACGGGCTGGGGACGATTGTCGCGGAAACTTAT CAACGGGAT...

Claims

1. 1. A pharmaceutical composition for treating a neurological disorder in a subject, comprising: (i) an adenosine base editor, or a nucleic acid sequence encoding same; and (ii) a guide polynucleotide, or a nucleic acid sequence encoding same; the adenosine base editor comprises a programmable DNA binding domain and an adenosine deaminase domain; The adenosine deaminase domain comprises an amino acid modification selected from the group consisting of I76Y, V82T, Y147T, Y147R, Q154S, and T166R, or a corresponding amino acid modification, as numbered in SEQ ID NO: 20; wherein the guide polynucleotide induces the adenosine base editor to effect an A to G nucleobase modification in a target gene or a regulatory element thereof associated with the neurological disorder, thereby treating the neurological disorder, i) the target gene is the alpha-L-iduronidase (IDUA) gene and the neurological disorder is Hurler syndrome; or ii) the target gene is the leucine-rich repeat kinase 2 (LRRK2) gene and the neurological disorder is Parkinson's disease; or iii) the target gene is the methyl-CpG binding protein 2 (MECP2) gene and the neurological disorder is Rett syndrome; or iv) the target gene is the ATP-binding cassette subfamily member 4 (ABCA4) gene and the neurological disorder is Stargardt's disease; Pharmaceutical compositions.

2. (i) the IDUA gene or a regulatory element thereof contains a SNP associated with Hurler syndrome; or (ii) the LRRK2 gene or a regulatory element thereof contains a SNP associated with Parkinson's disease; or (iii) the MECP2 gene or a regulatory element thereof contains a SNP associated with Rett Syndrome; or (iv) the ABCA4 gene or a regulatory element thereof contains a SNP associated with Stargardt disease; The pharmaceutical composition of claim 1.

3. 3. The pharmaceutical composition of claim 2, wherein the A to G nucleobase modification is in a SNP associated with Hurler syndrome, Parkinson's disease, Rett syndrome, or Stargardt disease.

4. (i) the Hurler syndrome-associated SNP results in a W402X or W401X amino acid variant in the IDUA polypeptide as numbered in SEQ ID NO:4, or a variant thereof encoded by the IDUA gene, where X is a stop codon; or (ii) the SNP associated with Parkinson's disease results in an A419V, R1441C, R1441H, or G2019S amino acid variant in the LRRK2 polypeptide as numbered in SEQ ID NO:3, or a variant thereof encoded by the LRRK2 gene; or (iii) the Rett Syndrome-associated SNP results in a R106W or T158M amino acid variant in an MECP2 polypeptide as numbered in SEQ ID NO:5, or a variant thereof encoded by the MECP2 gene; or (iv) the SNP associated with Stargardt disease results in an A1038V or G1961E amino acid variant in the ABCA4 polypeptide as numbered in SEQ ID NO: 6, or a variant thereof encoded by the ABCA4 gene; or (v) the Rett Syndrome-associated SNP results in a R255X or R270X amino acid mutation in an MECP2 polypeptide encoded by the MECP2 gene, where X is a stop codon. The pharmaceutical composition according to claim 2 or 3.

5. a) the A to G nucleobase modification changes a SNP associated with Hurler syndrome, Parkinson's disease, Rett syndrome, or Stargardt disease to a wild-type nucleobase; or b) the A to G nucleobase alteration changes the SNP associated with Hurler Syndrome to a non-wildtype nucleobase that results in amelioration of one or more symptoms of Hurler Syndrome; or c) an A to G modification in a SNP associated with Hurler syndrome changes a stop codon encoded by the IDUA gene to a tryptophan in the IDUA polypeptide; or d) the A to G nucleobase alteration in a SNP associated with Rett Syndrome changes a stop codon to a tryptophan in an MECP2 polypeptide; or e) the A to G nucleobase modification changes a cysteine ​​or histidine to an arginine in the LRRK2 polypeptide encoded by the LRRK2 gene; or f) said A to G nucleobase modification changes a serine to a glycine in an LRRK2 polypeptide encoded by the LRRK2 gene; or g) the A to G nucleobase modification substitutes arginine for cysteine ​​or histidine at position 144 or substitutes glycine for serine at position 2019 of the LRRK2 polypeptide or a variant thereof encoded by the LRRK2 gene as numbered in SEQ ID NO: 3; The pharmaceutical composition according to claim 2.

6. The guide polynucleotide the group consisting of 5′- GACUCUAGGCAGAGGUCUCAA -3′, 5′- ACUCUAGGCAGAGGUCUCAA-3′, 5′- CUCUAGGCCGAAGUGUCGC -3′, and 5′-GCUCUAGGCCGAAGUGUCGC-3′ for Hurler syndrome; the group consisting of 5′-AAGCGCAAGCCUGGAGGGAA -3′, 5′-ACUACAGCAUUGCUCAGUAC-3′ for Parkinson's disease; the group consisting of 5′- CUUUUCACUUCCUGCCGGGG-3′, 5′-AGCUUCCAUGUCCAGCCUUC-3′, 5′- ACCAUGAAGUCAAAAUCAUU-3′, and 5′- GCUUUCAGCCCCGUUUCUUG-3′ for Rett syndrome; and 5′-CUCCAGGGCGAACUUCGACACACAGC-3′ for Stargardt disease The pharmaceutical composition according to any one of claims 1 to 5, which is an sgRNA comprising a nucleic acid sequence selected from the following:

7. The method of claim 1, wherein the adenosine deaminase domain comprises an amino acid sequence having at least 90% identity to SEQ ID NO: 20; and / or b) the adenosine deaminase domain further comprises the amino acid modification V82S; The pharmaceutical composition of claim 1.

8. 1. An in vitro or ex vivo method of editing a target gene or a regulatory element thereof associated with a neurological disorder, comprising contacting the target gene or a regulatory element thereof with (i) an adenosine base editor and (ii) a guide polynucleotide; the adenosine base editor comprises a programmable DNA binding domain and an adenosine deaminase domain; The adenosine deaminase domain comprises an amino acid modification selected from the group consisting of I76Y, V82T, Y147T, Y147R, Q154S, and T166R, or a corresponding amino acid modification, as numbered in SEQ ID NO: 20; wherein the guide polynucleotide directs the adenosine base editor to effect an A to G nucleobase modification in a target gene associated with the neurological disorder or a regulatory element thereof, i) the target gene is the alpha-L-iduronidase (IDUA) gene and the neurological disorder is Hurler syndrome; or ii) the target gene is the leucine-rich repeat kinase 2 (LRRK2) gene and the neurological disorder is Parkinson's disease; or iii) the target gene is the methyl-CpG binding protein 2 (MECP2) gene and the neurological disorder is Rett syndrome; or iv) the target gene is the ATP-binding cassette subfamily member 4 (ABCA4) gene and the neurological disorder is Stargardt's disease; method.

9. a) said A to G nucleobase alteration changes said neurological disorder-associated SNP to a wildtype nucleobase; or b) the A to G nucleobase alteration changes the SNP associated with the neuropathy to a non-wild-type nucleobase that results in amelioration of one or more symptoms of the neuropathy. The method according to claim 8.

10. 1. An in vitro or ex vivo method of editing a Parkinson's disease-associated leucine-rich repeat kinase 2 (LRRK2) gene, or a regulatory element thereof, comprising contacting the LRRK2 gene, or a regulatory element thereof, with: (i) an adenosine base editor, or a nucleic acid sequence encoding same; and (ii) a guide polynucleotide, or a nucleic acid sequence encoding same; the adenosine base editor comprises a programmable DNA binding domain and an adenosine deaminase domain; The adenosine deaminase domain comprises an amino acid modification selected from the group consisting of I76Y, V82T, Y147T, Y147R, Q154S, and T166R, or a corresponding amino acid modification, as numbered in SEQ ID NO: 20; wherein said guide polynucleotide directs said adenosine base editor to result in an A to G nucleobase alteration at a SNP in the LRRK2 gene, and said SNP does not encode a G2019S mutation in an LRRK2 polypeptide as numbered in SEQ ID NO:3, or a variant thereof.

11. The method of claim 10 , wherein the contacting occurs intracellularly.

12. The guide polynucleotide is an sgRNA comprising the nucleic acid sequence 5'-AAGCGCAAGCCUGGAGGGAA-3' or 5'-ACUACAGCAUUGCUCAGUAC-3', 12. The method according to claim 10 or 11.

13. (i) the SNP results in an A419V, R1441C, R1441H, or G2019S amino acid variant in the LRRK2 polypeptide as numbered in SEQ ID NO:3, or a variant thereof encoded by the LRRK2 gene; or (ii) the A to G nucleobase modification changes a cysteine ​​or histidine to an arginine in the LRRK2 polypeptide encoded by the LRRK2 gene; or (iii) the A to G nucleobase modification changes a serine to a glycine in an LRRK2 polypeptide encoded by the LRRK2 gene; or (iv) the A to G nucleobase modification replaces the cysteine ​​or histidine at position 144 with arginine or replaces the serine at position 2019 with glycine of the LRRK2 polypeptide or a variant thereof encoded by the LRRK2 gene as numbered in SEQ ID NO: 3; The method according to any one of claims 10 to 12.

14. 1. An in vitro or ex vivo method for editing an alpha-L-iduronidase (IDUA) gene associated with Hurler syndrome, or a regulatory element thereof, comprising contacting the IDUA gene, or a regulatory element thereof, with: (i) an adenosine base editor, or a nucleic acid sequence encoding same; and (ii) a guide polynucleotide, or a nucleic acid sequence encoding same; the adenosine base editor comprises a programmable DNA binding domain and an adenosine deaminase domain; The adenosine deaminase domain comprises an amino acid substitution or a corresponding amino acid modification selected from the group consisting of I76Y, V82T, Y147T, Y147R, Q154S, and T166R, as numbered in SEQ ID NO: 20; the guide polynucleotide induces the adenosine base editor to effect an A to G nucleobase modification in the IDUA gene or a regulatory element thereof; method.

15. The method of claim 14 , wherein the contacting occurs intracellularly.

16. (i) the IDUA gene or a regulatory element thereof contains a SNP associated with Hurler syndrome; or (ii) the A to G nucleobase alteration is in a SNP associated with Hurler syndrome; or (iii) the IDUA gene or a regulatory element thereof contains a SNP associated with Hurler syndrome that results in a W402X or W401X amino acid variant in the IDUA polypeptide or a variant thereof encoded by the IDUA gene, as numbered in SEQ ID NO:4, where X is a stop codon; or (iv) the IDUA gene or a regulatory element thereof comprises a SNP associated with Hurler syndrome, and the A to G nucleobase modification changes the SNP associated with Hurler syndrome to a wild-type nucleobase; or (v) the IDUA gene or a regulatory element thereof contains a SNP associated with Hurler Syndrome, and the A to G nucleobase modification changes the SNP associated with Hurler Syndrome to a non-wild-type nucleobase that results in amelioration of one or more symptoms of Hurler Syndrome; or (vi) the IDUA gene or a regulatory element thereof contains a SNP associated with Hurler syndrome, and the A to G modification in the SNP associated with Hurler syndrome changes a stop codon encoded by the IDUA gene to tryptophan in the IDUA polypeptide; or (vii) the IDUA gene or a regulatory element thereof comprises a SNP associated with Hurler syndrome, and the guide polynucleotide comprises a nucleic acid sequence complementary to the IDUA gene or a regulatory element thereof comprising a SNP associated with Hurler syndrome; or (viii) the IDUA gene, or a regulatory element thereof, comprises a SNP associated with Hurler Syndrome, and the adenosine base editor is complexed with a single guide RNA comprising a nucleic acid sequence that is complementary to the IDUA gene, or a regulatory element thereof, that comprises a SNP associated with Hurler Syndrome; or (ix) the guide polynucleotide is an sgRNA comprising a nucleic acid sequence selected from the group consisting of 5'-GACUCUAGGCAGAGGUCUCAA-3', 5'-ACUCUAGGCAGAGGUCUCAA-3', 5'-CUCUAGGCCGAAGUGUCGC-3', and 5'-GCUCUAGGCCGAAGUGUCGC-3'; 16. The method according to claim 14 or 15.

17. 1. An in vitro or ex vivo method for editing a methyl-CpG binding protein 2 (MECP2) gene associated with Rett syndrome, or a regulatory element thereof, comprising contacting the MECP2 gene, or a regulatory element thereof, with (i) an adenosine base editor, or a nucleic acid sequence encoding same, and (ii) a guide polynucleotide, or a nucleic acid sequence encoding same; the adenosine base editor comprises a programmable DNA binding domain and an adenosine deaminase domain; The adenosine deaminase domain comprises an amino acid modification selected from the group consisting of I76Y, V82T, Y147T, Y147R, Q154S, and T166R, or a corresponding amino acid modification, as numbered in SEQ ID NO: 20; the guide polynucleotide directs the adenosine base editor to effect an A to G nucleobase modification in the MECP2 gene or a regulatory element thereof. method.

18. The method of claim 17 , wherein the contacting occurs intracellularly.

19. (i) the MECP2 gene or a regulatory element thereof contains a SNP associated with Rett Syndrome; and / or (ii) the MECP2 gene or a regulatory element thereof comprises a SNP associated with Rett Syndrome, and the A to G nucleobase alteration is in a SNP associated with Rett Syndrome; and / or (iii) the MECP2 gene or a regulatory element thereof comprises a SNP associated with Rett Syndrome, wherein the SNP associated with Rett Syndrome results in a R106W or T158M amino acid variant in an MECP2 polypeptide, or a variant thereof encoded by the MECP2 gene, as numbered in SEQ ID NO:5; or results in a R255X or R270X amino acid variant in an MECP2 polypeptide encoded by the MECP2 gene, where X is a stop codon; and / or (iv) the MECP2 gene or a regulatory element thereof comprises a SNP associated with Rett Syndrome, and the A to G nucleobase modification changes the SNP associated with Rett Syndrome to a wild-type nucleobase; and / or (v) the MECP2 gene or a regulatory element thereof comprises a SNP associated with Rett Syndrome, and the A to G nucleobase modification changes the SNP associated with Rett Syndrome to a non-wild-type nucleobase that results in amelioration of one or more symptoms of Rett Syndrome; and / or (vi) the MECP2 gene or a regulatory element thereof comprises a SNP associated with Rett Syndrome, and the A to G nucleobase alteration in the SNP associated with Rett Syndrome changes a stop codon to tryptophan in an MECP2 polypeptide; and / or (vii) the MECP2 gene or a regulatory element thereof comprises a SNP associated with Rett Syndrome, and the guide polynucleotide comprises a nucleic acid sequence complementary to the MECP2 gene or a regulatory element thereof comprising a SNP associated with Rett Syndrome; and / or (viii) the MECP2 gene, or a regulatory element thereof, comprises a SNP associated with Rett Syndrome, and the adenosine base editor is complexed with a single guide RNA comprising a nucleic acid sequence that is complementary to the MECP2 gene, or a regulatory element thereof, that comprises a SNP associated with Rett Syndrome; and / or (ix) the guide polynucleotide comprises a nucleic acid sequence selected from the group consisting of 5'-CUUUUCACUUCCUGCCGGGG-3', 5'-AGCUUCCAUGUCCAGCCUUC-3', 5'-ACCAUGAAGUCAAAAUCAUU-3', and 5'-GCUUUCAGCCCCGUUUCUUG-3'; 19. The method of claim 17 or 18.

20. 1. An in vitro or ex vivo method for editing a Stargardt disease-associated ATP-binding cassette subfamily member 4 (ABCA4) gene or a regulatory element thereof, comprising contacting the ABCA4 gene or a regulatory element thereof with: (i) an adenosine base editor or a nucleic acid sequence encoding same; and (ii) a guide polynucleotide or a nucleic acid sequence encoding same; the adenosine base editor comprises a programmable DNA binding domain and an adenosine deaminase domain; The adenosine deaminase domain comprises an amino acid substitution or a corresponding amino acid modification selected from the group consisting of I76Y, V82T, Y147T, Y147R, Q154S, and T166R, as numbered in SEQ ID NO: 20; the guide polynucleotide induces the adenosine base editor to effect an A to G nucleobase modification in the ABCA4 gene or a regulatory element thereof. method.

21. 21. The method of claim 20, wherein the contacting occurs intracellularly.

22. (i) the ABCA4 gene contains a SNP associated with Stargardt disease; and / or (ii) the ABCA4 gene comprises a SNP associated with Stargardt disease, and the A to G nucleobase modification is in the SNP associated with Stargardt disease; and / or (iii) the ABCA4 gene contains a SNP associated with Stargardt disease, the SNP associated with Stargardt disease resulting in an A1038V or G1961E amino acid variant in the ABCA4 polypeptide, as numbered in SEQ ID NO: 6, or a variant thereof encoded by the ABCA4 gene; and / or (iv) the ABCA4 gene contains a SNP associated with Stargardt disease, and said A to G nucleobase modification changes the SNP associated with Stargardt disease to a wild-type nucleobase; and / or (v) the ABCA4 gene contains a SNP associated with Stargardt disease, and said A to G nucleobase modification changes the SNP associated with Stargardt disease to a non-wild-type nucleobase that results in amelioration of one or more symptoms of Stargardt disease; and / or (vi) the ABCA4 gene comprises a SNP associated with Stargardt disease, and the guide polynucleotide comprises a nucleic acid sequence complementary to the ABCA4 gene or a regulatory element thereof comprising the SNP associated with Stargardt disease; and / or (vii) the ABCA4 gene comprises a SNP associated with Stargardt disease, and the adenosine base editor is complexed with a single guide RNA comprising a nucleic acid sequence complementary to the ABCA4 gene or a regulatory element thereof that comprises the SNP associated with Stargardt disease; and / or (viii) the guide polynucleotide is an sgRNA comprising the sequence 5'-CUCCAGGGCGAACUUCGACACACAGC-3'; 22. The method of claim 20 or 21.

23. The method of claim 23, wherein a) the adenosine deaminase domain comprises an amino acid sequence having at least 90% identity to SEQ ID NO: 20; and / or b) the adenosine deaminase domain further comprises the amino acid modification V82S; 21. The method of any one of claims 8, 10, 14, 17, and 20.

24. A cell produced by the method according to any one of claims 8 to 23.

25. A base editor system comprising: (i) an adenosine base editor or a nucleic acid sequence encoding same; and (ii) a guide polynucleotide or a nucleic acid sequence encoding same, the adenosine base editor comprises a programmable DNA binding domain and an adenosine deaminase domain; The adenosine deaminase domain comprises an amino acid substitution or a corresponding amino acid modification selected from the group consisting of I76Y, V82T, Y147T, Y147R, Q154S, and T166R, as numbered in SEQ ID NO: 20; wherein the guide polynucleotide directs the adenosine base editor to effect an A to G nucleobase modification in a target gene or a regulatory element thereof associated with a neurological disorder selected from the group consisting of: LRRK2, IDUA, MECP2, and ABCA4, i) the target gene is the alpha-L-iduronidase (IDUA) gene and the neurological disorder is Hurler syndrome; or ii) the target gene is the leucine-rich repeat kinase 2 (LRRK2) gene and the neurological disorder is Parkinson's disease; or iii) the target gene is the methyl-CpG binding protein 2 (MECP2) gene and the neurological disorder is Rett syndrome; or iv) the target gene is the ATP-binding cassette subfamily member 4 (ABCA4) gene and the neurological disorder is Stargardt's disease; Base editor system.

26. A method for the preparation of a polypeptide comprising the steps of: a) administering to a subject a polypeptide programmable by a DNA-binding domain comprising administering to said subject a polypeptide programmable by a DNA-binding domain; b) the polynucleotide-programmable DNA-binding domain is SpCas9 or SaCas9; or c) the polynucleotide-programmable DNA-binding domain comprises SpCas9 with protospacer adjacent motif specificity for a nucleotide sequence selected from the group consisting of NGG, NGA, NGCG, NGN, NNGRRT, NNNRRT, NGCG, NGCN, NGTN, and NGC, where N is A, G, C, or T and R is A or G; or d) the polynucleotide-programmable DNA binding domain comprises: (i) is nuclease inactive; or (ii) nickase; 26. The base editor system of claim 25. (iii) the adenosine deaminase comprises a combination of modifications selected from the group consisting of: Y147R + Q154R + Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y147T + Q154R; Y147T + Q154S; and Y123H + Y147R + Q154R + I76Y; or (iv) the adenosine base editor domain comprises an adenosine deaminase monomer; or (v) the adenosine base editor comprises an adenosine deaminase dimer; or (vi) the adenosine deaminase further comprises the amino acid modification V82S; 27. A base editor system according to claim 25 or 26.

28. The base editor system of claim 25, wherein the adenosine deaminase domain comprises an amino acid sequence having at least 90% identity to SEQ ID NO:

20.

29. An in vitro or ex vivo cell comprising the base editor system of any one of claims 25 to 28.

30. A pharmaceutical composition comprising the base editor system of any one of claims 25 to 28, or the cell of claim 29, and a pharma- ceutically acceptable carrier.

31. 31. The pharmaceutical composition of claim 30, further comprising a lipid.

32. A kit comprising the base editor system of any one of claims 25 to 28, or the cell of claim 29.