Compositions and methods for treating hepatitis B virus infection

JP2024534639A5Pending Publication Date: 2025-10-06BEAM THERAPEUTICS INC +4
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024518996
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-16
Filing Date
2022-09-27
Publication Date
2025-10-06

AI Technical Summary

Technical Problem

Current therapeutic approaches for hepatitis B virus (HBV) infection, such as antiviral drugs, do not cure the infection and are costly, and liver transplants are necessary for severe cases, posing significant health and financial burdens.

Method used

Introducing modifications to the HBV genome using a base editor system comprising a programmable DNA binding protein, nucleobase editor, and guide RNA to modify nucleobases, specifically targeting premature stop codons or deaminating nucleobases in covalently closed circular DNA (cccDNA) to alter viral replication.

Benefits of technology

This method potentially leads to a cure for HBV infection by reducing viral replication and decreasing the risk of chronic hepatitis, cirrhosis, and hepatocellular carcinoma, offering a more effective and less costly treatment alternative to existing therapies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Hepatitis B is a serious liver infection caused by the hepatitis B virus (HBV). Current therapeutic approaches to HBV infection have significant limitations. Antiviral drugs, such as the nucleotidic reverse transcriptase inhibitor tenofovir, can reduce viral replication but do not cure HBV-infected patients. The extent of liver damage caused by HBV may necessitate transplantation in some cases. In addition to the inherent risks of organ transplantation, costs can be high. Thus, improved methods for treating HBV infection are urgently needed. The present invention features compositions and methods for introducing mutations into the hepatitis B virus (HBV) genome.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 371,634, filed August 16, 2022, U.S. Provisional Application No. 63 / 357,623, filed June 30, 2022, and U.S. Provisional Application No. 63 / 248,938, filed September 27, 2021, the entire contents of which are incorporated herein by reference.

[0002] Electronic Sequence Listing Reference The contents of the electronic sequence listing (180802_053004_PCT_SL.xml; size: 10,086,238 bytes, and creation date: September 26, 2022) are incorporated herein by reference in their entirety. [Background technology]

[0003] Hepatitis B is a serious liver infection caused by the hepatitis B virus (HBV). HBV is a small DNA hepadnavirus that replicates via an RNA intermediate and can persist within infected cells by integrating into the host genome. Approximately 257 million people worldwide (including 850,000–2.2 million in the United States) are chronically infected with HBV. Chronic HBV infection manifests as chronic hepatitis, cirrhosis, and / or hepatocellular carcinoma. 20%–30% of adults with chronic HBV infection develop hepatocellular carcinoma or cirrhosis. 600,000–1 million people die from HBV infection annually.

[0004] Current treatment approaches for HBV infection have significant limitations. Antiviral medications, such as the nucleotide reverse transcriptase inhibitor tenofovir, can reduce viral replication but do not cure HBV-infected patients. Patients can pay as much as $500 to $1,500 per month for these antiviral therapies. Depending on the extent of liver damage caused by HBV, liver transplantation may be necessary. In addition to the inherent risks of organ transplantation, costs can be high. Therefore, improved methods for treating HBV infection are urgently needed. Summary of the Invention

[0005] As described below, the present invention features compositions and methods for treating hepatitis B virus (HBV) infection by introducing modifications into the HBV genome. In certain embodiments, the present invention provides base editor systems (e.g., fusion proteins comprising a programmable DNA-binding protein, a nucleobase editor, and a gRNA) for modifying the HBV genome to introduce changes, such as premature stop codons or changes in the coding sequence of HBV, or deamination of nucleobases in HBV covalently closed circular DNA (cccDNA).

[0006] In one aspect, the invention of this disclosure features a method for editing nucleobases of a hepatitis B virus (HBV) genome. The method includes contacting the HBV genome with one or more guide RNAs and a base editor containing a polynucleotide-programmable DNA-binding domain and an adenosine deaminase. The guide RNA targets the base editor, resulting in modification of the nucleobases of the HBV genome. The one or more guide RNAs may be selected from the group consisting of ACAAGAAUCCUCACAAUACC (SEQ ID NO: 405), UCGCUGGAUGUGUCUGCGGC (SEQ ID NO: 406), CACCUGUAUUCCCAUCCCAU (SEQ ID NO: 407), GGGGCCAAGUCUGUACAGCA (SEQ ID NO: 408), GCUGACGCAACCCCCACUGG (SEQ ID NO: 409), GGAGCUACUGUGGAGUUACU (SEQ ID NO: 410), UUCUUCUAGGGGACCUGCCU (SEQ ID NO: 411), GAGGACAAACGGGCAACAU A (SEQ ID NO: 412), UUGUCAACAAGAAAAACCCC (SEQ ID NO: 413), CCCAAGGUCUUACAUAAGAG (SEQ ID NO: 414), CCGGGCAACGGGGUAAAGGU (SEQ ID NO: 415), ACACAGAAAGGCCUUGUAAG (SEQ ID NO: 416), and CACGCACGCGCUGAUGGCCC (SEQ ID NO: 417), or a 1, 2, 3, 4, or 5 nucleotide 5' and / or 3' truncation fragment thereof.

[0007] In another aspect, the invention of this disclosure features a method for editing nucleobases in a hepatitis B virus (HBV) genome. The method includes (i) contacting a cell containing an HBV genome with one or more guide RNAs and a base editor containing a polynucleotide-programmable DNA-binding domain and an adenosine deaminase or cytidine deaminase. The guide RNA targets the base editor, resulting in modification of a nucleobase in the HBV genome, thereby editing the nucleobase in the HBV genome. The method further includes (ii) contacting the cell with an antiretroviral drug. Contacting the cell with the antiretroviral drug is associated with increased base editing efficiency compared to reference cells not contacted with the antiretroviral drug.

[0008] In another aspect, the invention of this disclosure features a method for editing nucleobases in a hepatitis B virus (HBV) genome. The method includes (i) contacting a cell containing an HBV genome with mRNA encoding BE4 and guide RNAs gRNA37 and gRNA40. The guide RNA targets a base editor to modify the nucleobases in the HBV genome, thereby editing the nucleobases in the HBV genome. The method further includes (ii) contacting the cell with lamivudine. Contacting the cell with an antiretroviral drug is associated with increased base editing efficiency compared to reference cells not contacted with the antiretroviral drug.

[0009] In another aspect, the invention of this disclosure features a method for editing nucleobases in a hepatitis B virus (HBV) genome. The method includes contacting a cell containing the HBV genome with an mRNA encoding a base editor comprising a polynucleotide-programmable DNA-binding domain and an adenosine deaminase or cytidine deaminase, and a guide RNA comprising the sequences GAAAGCCCAGGAUGAUGGGA (SEQ ID NO: 657) and CCAUGCCCCAAAGCCACCCA (SEQ ID NO: 662). The guide RNA targets the base editor, resulting in modification of a nucleobase in the HBV genome, thereby editing the nucleobase in the HBV genome.

[0010] In another aspect, the invention of this disclosure features a method of treating hepatitis B virus (HBV) infection in a subject. The method includes administering to a subject in need thereof a fusion protein, protein complex, or polynucleotide encoding the fusion protein or protein complex. The fusion protein or protein complex contains a polynucleotide-programmable DNA-binding domain, a base editor domain that is an adenosine deaminase domain, and one or more guide RNAs that target the base editor domain to modify a nucleobase in the HBV genome from A·T to G·C. The one or more guide RNAs may be selected from the group consisting of ACAAGAAUCCUCACAAUACC (SEQ ID NO: 405), UCGCUGGAUGUGUCUGCGGC (SEQ ID NO: 406), CACCUGUAUUCCCAUCCCAU (SEQ ID NO: 407), GGGGCCAAGUCUGUACAGCA (SEQ ID NO: 408), GCUGACGCAACCCCCACUGG (SEQ ID NO: 409), GGAGCUACUGUGGAGUUACU (SEQ ID NO: 410), UUCUUCUAGGGGACCUGCCU (SEQ ID NO: 411), GAGGACAAACGGGCAACAU A (SEQ ID NO: 412), UUGUCAACAAGAAAAACCCC (SEQ ID NO: 413), CCCAAGGUCUUACAUAAGAG (SEQ ID NO: 414), CCGGGCAACGGGGUAAAGGU (SEQ ID NO: 415), ACACAGAAAGGCCUUGUAAG (SEQ ID NO: 416), and CACGCACGCGCUGAUGGCCC (SEQ ID NO: 417), or a 1, 2, 3, 4, or 5 nucleotide 5' and / or 3' truncation fragment thereof.

[0011] In another aspect, the invention of this disclosure features a method of treating hepatitis B virus (HBV) infection in a subject. The method includes administering to a subject in need thereof one or more polynucleotides encoding a base editor domain that is a polynucleotide-programmable DNA-binding domain and an adenosine deaminase domain, and one or more guide RNAs that target the base editor domain and result in an A·T to G·C modification of a nucleobase in the HBV genome. The one or more guide RNAs may be ACAAGAAUCCUCACAAUACC (SEQ ID NO: 405), UCGCUGGAUGUGUCUGCGGC (SEQ ID NO: 406), CACCUGUAUUCCCAUCCCAU (SEQ ID NO: 407), GGGGCCAAGUCUGUACAGCA (SEQ ID NO: 408), GCUGACGCAACCCCCACUGG (SEQ ID NO: 409), GGAGCUACUGUGGAGUUACU (SEQ ID NO: 410), UUCUUCUAGGGGACCUGCCU (SEQ ID NO: 411), GAGGACAAACGGGCAACAU A (SEQ ID NO: 412), UUGUCAACAAGAAAAACCCC (SEQ ID NO: 413), CCCAAGGUCUUACAUAAGAG (SEQ ID NO: 414), CCGGGCAACGGGGUAAAGGU (SEQ ID NO: 415), ACACAGAAAGGCCUUGUAAG (SEQ ID NO: 416), and CACGCACGCGCUGAUGGCCC (SEQ ID NO: 417), or a 1, 2, 3, 4, or 5 nucleotide 5' and / or 3' truncation fragment thereof.

[0012] In another aspect, the invention of this disclosure features a method of treating hepatitis B virus (HBV) infection in a subject. The method includes administering to a subject in need thereof one or more guide RNAs and a base editor containing a polynucleotide-programmable DNA-binding domain and an adenosine deaminase or cytidine deaminase. The one or more guide RNAs contain the sequences GAAAGCCCAGGAUGAUGGGA (SEQ ID NO: 657) and CCAUGCCCCAAAGCCACCCA (SEQ ID NO: 662), which target the base editor to modify a nucleobase in the HBV genome, thereby editing the nucleobase in the HBV genome.

[0013] In another aspect, the invention of this disclosure features a composition containing base editor(s) bound to guide RNA(s), wherein the guide RNA(s) are selected from the group consisting of ACAAGAAUCCUCACAAUACC (SEQ ID NO: 405), UCGCUGGAUGUGUCUGCGGC (SEQ ID NO: 406), CACCUGUAUUCCCAUCCCAU (SEQ ID NO: 407), GGGGCCAAGUCUGUACAGCA (SEQ ID NO: 408), GCUGACGCAACCCCCACUGG (SEQ ID NO: 409), GGAGCUACUGUGGAGUUACU (SEQ ID NO: 410), UUCUUCUAGGGGACCUGCCU (SEQ ID NO: 411), GAGGACAAACGGGCAACAU A (SEQ ID NO: 412), UUGUCAACAAGAAAAACCCC (SEQ ID NO: 413), CCCAAGGUCUUACAUAAGAG (SEQ ID NO: 414), CCGGGCAACGGGGUAAAGGU (SEQ ID NO: 415), ACACAGAAAGGCCUUGUAAG (SEQ ID NO: 416), and CACGCACGCGCUGAUGGCCC (SEQ ID NO: 417), or a 1, 2, 3, 4, or 5 nucleotide 5' and / or 3' truncation fragment thereof.

[0014] In another aspect, the invention of this disclosure features a pharmaceutical composition for treating HBV infection. The composition includes (i) a base editor, or a nucleic acid encoding a base editor, and one or more guide RNAs (gRNAs) containing a nucleic acid sequence complementary to an HBV gene in a pharmaceutically acceptable excipient. The one or more guide RNAs may be selected from the group consisting of ACAAGAAUCCUCACAAUACC (SEQ ID NO: 405), UCGCUGGAUGUGUCUGCGGC (SEQ ID NO: 406), CACCUGUAUUCCCAUCCCAU (SEQ ID NO: 407), GGGGCCAAGUCUGUACAGCA (SEQ ID NO: 408), GCUGACGCAACCCCCACUGG (SEQ ID NO: 409), GGAGCUACUGUGGAGUUACU (SEQ ID NO: 410), UUCUUCUAGGGGACCUGCCU (SEQ ID NO: 411), GAGGACAAACGGGCAACAU A (SEQ ID NO: 412), UUGUCAACAAGAAAAACCCC (SEQ ID NO: 413), CCCAAGGUCUUACAUAAGAG (SEQ ID NO: 414), CCGGGCAACGGGGUAAAGGU (SEQ ID NO: 415), ACACAGAAAGGCCUUGUAAG (SEQ ID NO: 416), and CACGCACGCGCUGAUGGCCC (SEQ ID NO: 417), or a 1, 2, 3, 4, or 5 nucleotide 5' and / or 3' truncation fragment thereof.

[0015] In another aspect, the invention of this disclosure features a method of treating HBV infection, the method comprising administering to a subject in need thereof the composition of any one of the above aspects.

[0016] In another aspect, the invention of this disclosure features a method of treating HBV infection, the method comprising administering to a subject in need thereof the pharmaceutical composition of any one of the above aspects.

[0017] In another aspect, the invention of this disclosure features the use of a composition of any one of the above aspects in treating an HBV infection in a subject.

[0018] In another aspect, the invention of this disclosure features the use of the pharmaceutical composition of any one of the above aspects in treating an HBV infection in a subject.

[0019] In another aspect, the invention of the present disclosure provides a guide RNA, comprising: ACAAGAAUCCUCACAAUACC (SEQ ID NO: 405), UCGCUGGAUGUGUCUGCGGC (SEQ ID NO: 406), CACCUGUAUUCCCAUCCCAU (SEQ ID NO: 407), GGGGCCAAGUCUGUACAGCA (SEQ ID NO: 408), GCUGACGCAACCCCCACUGG (SEQ ID NO: 409), GGAGCUACUGUGGAGUUACU (SEQ ID NO: 410), UUCUUCUAGGGGACCUGCCU (SEQ ID NO: 411), GAGGACAAACGGGCAACA The present invention features guide RNAs that contain a 5' to 3' nucleotide sequence selected from one or more of UA (SEQ ID NO: 412), UUGUCAACAAGAAAAACCCC (SEQ ID NO: 413), CCCAAGGUCUUACAUAAGAG (SEQ ID NO: 414), CCGGGCAACGGGGUAAAGGU (SEQ ID NO: 415), ACACAGAAAGGCCUUGUAAG (SEQ ID NO: 416), and CACGCACGCGCUGAUGGCCC (SEQ ID NO: 417), or a 1, 2, 3, 4, or 5 nucleotide 5' and / or 3' truncation fragment thereof.

[0020] In another aspect, the invention of this disclosure features a pharmaceutical composition containing (i) a nucleic acid encoding a base editor; and (ii) a guide RNA of any one of the above aspects.

[0021] In any of the above embodiments, the antiretroviral agent is selected from one or more of lamivudine, entecavir, tenofovir, interferon, and PEG-interferon. In any of the above embodiments, the retroviral agent is lamivudine.

[0022] In any of the above embodiments, step (i) precedes step (ii), or step (ii) precedes step (i).

[0023] In any of the above aspects, the cells are first contacted with the antiretroviral agent about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 days or at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 days prior to step (i). In any of the above aspects, the cells have not previously been contacted with the antiretroviral agent. In any of the above aspects, the cells are in a subject. In any of the above aspects, the contacting is in a eukaryotic cell, a mammalian cell, or a human cell. In any of the above aspects, the cell is in vitro or in vivo. In any of the above aspects, or embodiments thereof, the subject is a mammal or a human. In any of the above aspects, or embodiments thereof, the subject is a mammal. In any of the above aspects, or embodiments thereof, the subject is a human.

[0024] In any of the above embodiments, the HBV genome is contacted with two or more guide RNAs simultaneously.

[0025] In any of the above aspects, the guide RNA is selected from one or more of ACAAGAAUCCUCACAAUACC (SEQ ID NO: 405), UCGCUGGAUGUGUCUGCGGC (SEQ ID NO: 406), CACCUGUAUUCCCAUCCCAU (SEQ ID NO: 407), GGGGCCAAGUCUGUACAGCA (SEQ ID NO: 408), GCUGACGCAACCCCCACUGG (SEQ ID NO: 409), GGAGCUACUGUGGAGUUACU (SEQ ID NO: 410), UUCUUCUAGGGGACCUGCCU (SEQ ID NO: 411), GAGGACAAACGGGCAACAUA (SEQ ID NO: 412), UUGUCAACAAGAAAAACCCC (SEQ ID NO: 413), CCCAAGGUCUUACAUAAGAG (SEQ ID NO: 414), CCGGGCAACGGGGUAAAGGU (SEQ ID NO: 415), ACACAGAAAGGCCUUGUAAG (SEQ ID NO: 416), and CACGCACGCGCUGAUGGCCC (SEQ ID NO: 417).

[0026] In any of the above embodiments, the one or more guide RNAs contain a spacer sequence selected from the sequences listed in Table 2A, Table 2B, Table 2C, and / or SEQ ID NOs: 3105-5485 and 8220-10830.

[0027] In any of the above embodiments, the one or more guide RNAs contain a spacer containing the nucleotide sequences GAAAGCCCAGGAUGAUGGGA (SEQ ID NO: 657) and CCAUGCCCCAAAGCCACCCA (SEQ ID NO: 662). In any of the above embodiments, the one or more guide RNAs contain the nucleotide sequences mG*mA*mA*AGCCCAGGAUGAUGGGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUmU*mU*mU* (SEQ ID NO: 424) and mC*mC*mA*UGCCCCAAAGCCACCCAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUmU*mU*mU* (SEQ ID NO: 422).

[0028] In any of the above embodiments, base editing efficiency is improved by at least about 5%, 10%, 15%, or 20%.

[0029] In any of the above embodiments, the adenosine deaminase converts a target A·T to G·C in the hepatitis B virus (HBV) genome.

[0030] In any of the above aspects, the nucleobase is in a polynucleotide encoding an HBV protein. In any of the above aspects, the nucleobase is in a polynucleotide sequence encoding an HBV protein selected from HBV core protein (h), HBV polymerase (Pol), HBV surface protein, or HBV protein X. In embodiments of any of the above aspects, the nucleobase modification in the polynucleotide encoding the HBV protein results in a missense mutation. In embodiments of any of the above aspects, the nucleobase modification is associated with decreased transcription of the polynucleotide sequence encoding the HBV protein. In embodiments of any of the above aspects, the HBV protein is HBV polymerase (Pol) and / or HBV protein X. In embodiments of any of the above aspects, the missense mutation is in the HBV pol gene. In embodiments of any of the above aspects, the missense mutation results in E24G, L25F, P26F, R27C, V48A, V48I, S382F, V378I, V378A, V379I, V379A, L377F, D380G, D380N, F381P, R376G, A422T, F423P, A432V, M433V, P434S, D540G, A688V, D689G, A717T, E718K, P713S, P713L, or L719P in the HBV polymerase protein encoded by the HBV pol gene. In embodiments of any of the above aspects, the missense mutation is in the HBV core gene. In embodiments of any of the above aspects, the missense mutation results in T160A, T160A, P161F, S162L, C183R, or *184Q in the HBV core protein encoded by the HBV core gene. In embodiments of any of the above aspects, the missense mutation is in the HBV X gene. In embodiments of any of the above aspects, the missense mutation results in H86R, W120R, E122K, E121K, or L141P in the HBV X protein encoded by the HBV X gene. In embodiments of any of the above aspects, the missense mutation is in the HBV S gene.In embodiments of any of the above aspects, the missense mutation results in S38F, L39F, W35R, W36R, T37I, T37A, R78Q, S34L, F80P, or D33G in the HBV S protein encoded by the HBV S gene.

[0031] In any of the above aspects, the nucleobase is associated with a transcription site. In embodiments of any of the above aspects, the transcription site is selected from one or more of enhancer II box A, enhancer I, and the HBX promoter.

[0032] In any of the above embodiments, the polynucleotide programmable DNA-binding domain contains a Cas12 polypeptide. In any of the above embodiments, the polynucleotide programmable DNA-binding domain contains Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, or Cas12i. In any of the above embodiments, the polynucleotide programmable DNA-binding domain contains a sequence having at least about 85% amino acid sequence identity to Bacillus hisashii Cas12b, Bacillus thermoamylovorans Cas12b, Bacillus sp. V3-13 Cas12b, or Alicyclobacillus acidiphilus Cas12b. In any of the above embodiments, the polynucleotide programmable DNA-binding domain contains a sequence having at least about 85% amino acid sequence identity to Bacillus hiasashii Cas12b (bhCas12b). In any of the above aspects, the polynucleotide-programmable DNA-binding domain contains a nuclease-inactive variant or a nickase variant. In embodiments of any of the above aspects, the nuclease-inactive variant or nickase variant contains a nuclease-inactivated bhCas12b containing the amino acid substitutions D952A, S893R, K846R, and E837G, or their corresponding amino acid substitutions. In any of the above aspects, the base editor contains bhCas12b or a bhCas12b variant containing the amino acid substitutions D952A, S893R, K846R, and E837G.

[0033] In any of the above aspects, the base editor contains an adenosine deaminase. In any of the above aspects, or embodiments thereof, the adenosine deaminase domain is capable of deaminating adenine in deoxyribonucleic acid (DNA). In any of the above aspects, or embodiments thereof, the deaminase domain is a TadA domain. In some embodiments, the TadA deaminase is TadA*7.10, TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23, or TadA*8.24. In an embodiment of any of the above aspects, the TadA deaminase is TadA*8.13.

[0034] In any of the above embodiments, the one or more guide RNAs comprise a CRISPR RNA (crRNA) and a transcoding small RNA (tracrRNA), wherein the crRNA comprises a nucleic acid sequence complementary to an HBV nucleic acid sequence. In any of the above embodiments, the one or more guide RNAs target 3, 4, or 5 nucleobases of the HBV genome.

[0035] In any of the above embodiments, the base editor is complexed with a single guide RNA (sgRNA) that contains a nucleic acid sequence complementary to an HBV nucleic acid sequence.

[0036] In any of the above embodiments, the method comprises editing two or more nucleobases.

[0037] In any of the above aspects, the method comprises contacting the HBV genome with two or more guide RNAs targeting two or more HBV nucleic acid sequences. In any of the above aspects, the method comprises delivering a fusion protein, a polynucleotide encoding the fusion protein, or one or more polynucleotides encoding a polynucleotide programmable DNA-binding domain and a base editor domain, and one or more guide RNAs to a cell of the subject. In any of the above aspects, or embodiments thereof, the cell is a hepatocyte.

[0038]

[0039] In any of the above aspects, the composition or pharmaceutical composition further comprises a lipid. In embodiments of any of the above aspects, the lipid is a cationic lipid.

[0040] In any of the above aspects, the composition contains a pharmaceutically acceptable excipient.

[0041] In any of the above embodiments, the one or more gRNAs and the base editor are formulated together. In any of the above embodiments, the one or more gRNAs and the base editor are formulated separately.

[0042] In any of the above aspects, the composition or pharmaceutical composition further comprises a vector suitable for expression in a mammalian cell, the vector containing a polynucleotide encoding a base editor. In any of the above aspects, or embodiments thereof, the vector is a viral vector. In any of the above aspects, or embodiments thereof, the viral vector is a retroviral vector, an adenoviral vector, a lentiviral vector, a herpesvirus vector, or an adeno-associated viral vector (AAV). In any of the above aspects, or embodiments thereof, the composition or pharmaceutical composition further comprises a ribonucleopartite suitable for expression in a mammalian cell.

[0043] In any of the above aspects, the HBV is of genotype A or of genotype D. In any of the above aspects, the method reduces the level of a marker selected from one or more of hepatitis B surface antigen, hepatitis B E antigen, hepatitis B virus total DNA, hepatitis B 3.5 kb RNA, and hepatitis B covalently closed circular DNA. In any of the above aspects, the method edits the hepatitis B S antigen site and the PreCore site. In any of the above aspects, or embodiments thereof, the editing efficiency is about 30% at the S antigen site and about 60% at the PreCore site.

[0044] In any of the above embodiments, the nucleic acid encoding the base editor is an mRNA.

[0045] In any of the above embodiments, the one or more guide RNAs comprise the following scaffold sequence: gUUUUAGagcuagaaauagcaaGUUaAaAuAaggcuaGUccGUUAucAAcuugaaaaagugGcaccgagucggugcusususu (SEQ ID NO: 317), where a, c, u, or g represent bases with a 2'O-methyl (M) modification and as, cs, us, or gs represent bases with a 2'-O-methyl 3'-phosphorothioate (MS) modification. In any of the above aspects, one or more guide RNAs comprise the nucleotide sequences gsasasAGCCCAGGAUGAUGGGAgUUUUAGagcuagaaauagcaaGUUaAaAuAaggcuaGUccGUUAucAAcuugaaaaagugGcaccgagucggugcusususu (SEQ ID NO: 424) and cscsasUGCCCCAAAGCCACCCAgUUUUAGagcuagaaauagcaaGUUaAaAuAaggcuaGUccGUUAucAAcuugaaaaagugGcaccgagucggugcusususu (SEQ ID NO: 422), where a, c, u, or g represent A, C, U, or G nucleotides containing a 2'O-methyl (M) modification, respectively, and as, cs, us, or gs represent A, C, U, or G nucleotides containing a 2'-O-methyl 3'-phosphorothioate (MS) modification.

[0046] In any of the above aspects, the cells are contacted with one or more guide RNAs and base editors at a first and second time point. In embodiments, the interval between the first time and the second time the cells are contacted is at least one week. In embodiments, the interval between the first time and the second time the cells are contacted is at least two weeks.

[0047] In any of the above embodiments, the base editor is BE4. In any of the above embodiments, the gRNA comprises GAAAGCCCAAGAUGAUGGGA (SEQ ID NO: 10834).

[0048] The compositions and articles defined by the present invention have been isolated or otherwise prepared in connection with the examples provided below. Other features and advantages of the present invention will be apparent from the detailed description and claims.

[0049] definition Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by those skilled in the art to which this invention belongs. The following references provide those skilled in the art with general definitions of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994), The Cambridge Dictionary of Science and Technology (Walker ed., 1988), The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991), and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the following meanings unless otherwise specified.

[0050] By "BE4 polypeptide" is meant a polypeptide having at least about 85% identity to the following amino acid sequence, or a fragment thereof that has cytidine deaminase activity:

[0051] By "BE4 polynucleotide" is meant a polynucleotide that encodes a BE4 polypeptide.

[0052] By "HBV polymerase (pol) protein" is meant a polypeptide or fragment thereof having at least about 95% identity to the amino acid sequence of wild-type HBV polymerase that functions in hepatitis B virus infection. In one embodiment, the HBV polymerase is encoded by an HBV A, B, C, D, E, F, G, or H genotype. In one embodiment, the HBV polymerase amino acid sequence is provided under UniPro accession number Q8B5R0-1, which is reproduced below. MPLSYQHFRRLLLLDDEAGPLEEELPRLADEGLNRRVAEDLNLGNLNVSIPWTHKVGNFTGLYSSTVPVFNPHWKTPSFPNIHLHQDIIKKCEQFVGPLTVNEKRRLQLIMPARFYPKVTKYLPLDKGIKPYYPEHLVNHYFQTRHYLHTLWKAGILYKRETTHSASFCGSPYSWEQDLQHGAESFHQQS (SEQ ID NO: 370).

[0053] Mutations in HBV polymerase include, in any of the above aspects or embodiments thereof, E24G, L25F, P26F, R27C, V48A, V48I, S382F, V378I, V378A, V379I, V379A, L377F, D380G, D380N, F381P, R376G, A422T, F423P, A432V, M433V, P434S, D540G, A688V, D689G, A717T, E718K, P713S, P713L, or L719P.

[0054] Other exemplary HBV DNA polymerases include, for example, NCBI accession number AAB59972.1, which has the following sequence: (SEQ ID NO: 371).

[0055] "HBV polymerase gene" means a polynucleotide that encodes HBV polymerase.

[0056] By "hepatitis B surface antigen (HBsAg) polypeptide" or "HBV surface protein (S)" is meant an antigenic protein or fragment thereof having at least about 85% identity to NCBI Accession No. AAB59969.1 that functions in HBV viral infection. An exemplary HBsAg amino acid sequence is provided below. MENITSGFLGPLLVLQAGFFLLTRILTIPQSLDSWWTSLNFLGGTTVCLGQNSQSPTSNHSPTSCPPTCPGYRWMCLRRFIIFLFILLLCLIFLLVLLDYQGMLPVCPLIPGSSTTSTGPCRTCMTTAQGTSMYPSCCCTKPSDGNCTCIPIPSSWAFGKFLWEWASARFSWLSLLVPFVQWFVGLSPTVWLSVIWMMWYWGPSLYSILSPFLPLLPIFFCLWVYI (SEQ ID NO: 372).

[0057] By "HbsAg polynucleotide" is meant a polynucleotide that encodes the HBsAg protein.

[0058] "HBV X protein" or "HBV protein X" refers to a polypeptide or fragment thereof having at least about 85% identity to NCBI Accession No. AAB59970.1 that functions in HBV viral infection. Exemplary amino acid sequences are provided below: MAARLCCQLDPARDVLCLRPVGAESCGRPFSGSLGTLSSPSPSAVPTDHGAHLSLRGLPVCAFSSAGPCALRFTSARRMETTVNAHRMLPKVLHKRTLGLSAMSTTDLEAYFKDCLFKDWEELGEEIRLKVFVLGGCRHKLVCAPAPCNFFTSA (SEQ ID NO: 373).

[0059] By "core antigen precursor," "precore protein," or "hepatitis B e antigen (HBeAg)" is meant a polypeptide or fragment thereof having at least about 85% identity to NCBI Accession No. AAB59971.1 that functions in HBV viral infection. MQLFHLCLIISCSCPTVQASKLCLGWLWGMDIDPYKEFGATVELLSFLPSDFFPSVRDLLDTASALYREALESPEHCSPHHTALRQAILCWGELMTLATWVGVNLEDPASRDLVVSYVNTNMGLKFRQLLWFHISCLTFGRETVIEYLVSFGVWIRTPPAYRPPNAPILSTLPETTVVRRRGRSPRRRTPSPRRRRSQSPRRRRSQSREPQC (SEQ ID NO: 10831).

[0060] "HBV core protein," "HBc," or "core protein" refers to a polypeptide having at least about 95% identity to the amino acid sequence of wild-type HBV core protein, or a fragment thereof. In one embodiment, the HBV core protein functions in hepatitis B virus infection. In one embodiment, the HBV core protein is encoded by HBV A, B, C, D, E, F, G, or H genotypes. In one embodiment, the HBV core protein amino acid sequence is provided at NCBI GenBank accession number AXG50928.1 and is provided below: MDIDPYKEFGASVELLSFLPSDFFPSIRDLLDTASALYREALESPEHCSPHHTALRQAILCWGELMNLATWVGSNLEDPASRELVVSYVNVNMGLKIRQLLWFHISCLTFGRETVLEYLVSFGVWIRTPPAYRPPNAPILSTLPETTVVRRRGRSPRRRTPSPRRRRSQSPRRRRSQSRESQC (SEQ ID NO: 374). "HBV X protein" means a polynucleotide encoding the HBV X protein.

[0061] "HBV X protein (genotype B)" refers to a polypeptide having at least about 95% identity to the amino acid sequence of a wild-type HBV genotype BX protein, or a fragment thereof. In one embodiment, the HBV X protein functions in hepatitis B virus infection. In one embodiment, the HBV genotype BX protein amino acid sequence is provided at NCBI GenBank accession number BAQ95575.1 and is provided below: MAARLCCQLDPARDVLCLRPVGAESRGRPLPGPLGALPPASPPVVPSDHGAHLSLRGLPVCAFSSXGPCALRFTSARRMETTVNAHRNLPKVLHKRTLGLSAMSTTDLEAYFKDCVFXEWEELGEEXRLKVFVLGGCRHKLVCSPAPCNFFTSA (SEQ ID NO: 375).

[0062] "HBV X protein (genotype C)" refers to a polypeptide having at least about 95% identity to the amino acid sequence of a wild-type HBV genotype CX protein, or a fragment thereof. In one embodiment, the HBV X protein functions in hepatitis B virus infection. In one embodiment, the HBV genotype CX protein amino acid sequence is provided at NCBI GenBank accession number BAQ95563.1 and is provided below: MAARVCCQLDPARDVLCLRPVGAESRGRPVSGPFGPLPSPSSSAVPADYGAHLSLRGLPVCAFSSAGPCALRFTSARRMETTVNAHQVLPKLLHKRTLGLSAMSTTDLEAYFKDCLFKDWEELGEEIRLKVFVLGGCRHKLVCSPAPCNFFTSA (SEQ ID NO: 376).

[0063] "HBV S protein" refers to a polypeptide having at least about 95% identity to a wild-type HBV S protein amino acid sequence or a fragment thereof. In one embodiment, the HBV S protein functions in hepatitis B virus infection. In one embodiment, the HBV S protein is encoded by an HBV A, B, C, D, E, F, G, or H genotype. In one embodiment, the HBV S protein amino acid sequence is provided at NCBI GenBank accession number ABV02793.1 and is provided below: MENTTSGFLGPLLVLQAGFFLLTRNLTIPQSLDSWWTSLNFLGGAPTCPGQNSQSPTSNHSPTSCPPICPGYRWMCLRRFIIFLFILLLCLIFLLVLLDYQGMLPVCPLLPGTSTTSTGPCKTCTIPAQGTSMFPSCCCTKPSDGNCTCIPIPSSWAFARFLWEWASVRFSWLSLLVPFVQWFVGLSPTVWLSVIWMMWYWGPSLYNILSPFLPLLPIFFCLWVYI (SEQ ID NO: 377).

[0064] The complete genome of hepatitis B virus subtype ayw (complete genome), including polynucleotides encoding HBV polymerase, HBsAg protein, HBV X protein, and core antigen precursor, is provided under GenBank accession number U95551.1 and is reproduced below. The nucleotide locations of regions corresponding to regulatory elements (e.g., enhancer I, enhancer II box A, HBX promoter) and polypeptide-encoding sequences within the hepatitis B virus genome are known in the art (see, e.g., Panjaworayan, et al. "HBVRegDB: Annotation, comparison, detection, and visualization of regulatory elements in hepatitis B virus sequences," Virol. J., 4:136, DOI:10.1186 / 1743-422X-4-136).

[0065] "Adenine" or "9H-purin-6-amine" has the molecular formula C5H5N5 and the structure [ka] and refers to the purine nucleobase having the formula:

[0066] "Adenosine" or "4-amino-1-[(2R,3R,4S,5R)-3,4-dihydroxy-5-(hydroxymethyl)oxolan-2-yl]pyrimidin-2(1H)-one" is a compound attached to a ribose sugar via a glycosidic bond and has the structure [ka] It refers to the adenine molecule having the formula C and corresponding to the CAS number 65-46-3. 10 H 13 It is N5O4.

[0067] "Adenosine deaminase" or "adenine deaminase" refers to a polypeptide or fragment thereof that can catalyze the hydrolytic deamination of adenine or adenosine. In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine to inosine or deoxyadenosine to deoxyinosine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases (e.g., engineered adenosine deaminases, evolved adenosine deaminases) provided herein can be derived from any organism (e.g., eukaryotes, prokaryotes), including, but not limited to, algae, bacteria, fungi, plants, invertebrates (e.g., insects), and vertebrates (e.g., amphibians, mammals). In some embodiments, the adenosine deaminase is an adenosine deaminase variant having one or more modifications and capable of deaminating both adenine and cytosine in a target polynucleotide (e.g., DNA, RNA). In some embodiments, the target polynucleotide is single-stranded or double-stranded. In some embodiments, the adenosine deaminase variant is capable of deaminating both adenine and cytosine in DNA. In some embodiments, the adenosine deaminase variant is capable of deaminating both adenine and cytosine in single-stranded DNA. In some embodiments, the adenosine deaminase variant is capable of deaminating both adenine and cytosine in RNA.

[0068] "Adenosine deaminase activity" means catalyzing the deamination of adenine or adenosine to guanine in a polynucleotide. In some embodiments, the adenosine deaminase variants provided herein maintain adenosine deaminase activity (e.g., at least about 30%, 40%, 50%, 60%, 70%, 80%, 90%, or more of the activity of a reference adenosine deaminase (e.g., TadA*8.20 or TadA*8.19)).

[0069] "Adenosine base editor (ABE)" means a base editor that includes an adenosine deaminase.

[0070]

[0047] An "adenosine base editor (ABE) polynucleotide" refers to a polynucleotide encoding an ABE. An "adenosine deaminase base editor 8 (ABE8) polypeptide" or "ABE8" refers to a base editor as defined herein, including an adenosine deaminase variant, that comprises one or more of the modifications listed in Table 15, one of the combinations of modifications listed in Table 15, or a modification at one or more of the sites listed in Table 15 (e.g., 82 and / or 166), where such modifications are relative to the following reference sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 1).

[0071] In some embodiments, ABE8 comprises further modifications relative to the reference sequence, as described herein.

[0072] By "adenosine deaminase base editor 8 (ABE8) polynucleotide" is meant a polynucleotide that encodes ABE8.

[0073] "Administering" is referred to herein as providing one or more compositions described herein to a patient or subject.

[0074] By "agent" is meant any small molecule chemical compound, antibody, nucleic acid molecule, or polypeptide, or fragment thereof.

[0075] "Alteration" refers to a change (increase or decrease) in the level, structure, or activity of an analyte, gene, or polypeptide, as detected by standard methods known in the art, such as those described herein. As used herein, alteration includes a 10% change in expression levels, a 25% change, a 40% change, and a 50% or greater change in expression levels. In some embodiments, alteration includes an insertion, deletion, or substitution of a nucleic acid base or amino acid.

[0076] "Ameliorate" means to reduce, suppress, attenuate, decrease, arrest, or stabilize the onset or progression of a disease.

[0077] "Analog" refers to a molecule that has similar, but not identical, functional or structural characteristics. For example, a polypeptide analog retains the biological activity of the corresponding naturally occurring polypeptide while possessing certain biochemical modifications that enhance the analog's function relative to the naturally occurring polypeptide. Such biochemical modifications may increase the analog's protease resistance, membrane permeability, or half-life, for example, without altering ligand binding. Analogs may also contain unnatural amino acids.

[0078] "Base editor (BE)" or "nucleobase editor polypeptide (NBE)" refers to an agent that binds to a polynucleotide and has nucleobase-modifying activity. In various embodiments, a base editor comprises a polynucleotide-programmable nucleotide-binding domain (e.g., Cas9 or Cpf1) in combination with a nucleobase-modifying polypeptide (e.g., a deaminase) and a guide polynucleotide (e.g., a guide RNA (gRNA)). Representative nucleic acid and protein sequences of base editors are provided in the Sequence Listing as SEQ ID NOs: 2-11.

[0079] "Base editing activity" refers to acting to chemically modify a base within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is cytidine deaminase activity, e.g., converting a targeted C·G to T·A. In another embodiment, the base editing activity is adenosine or adenine deaminase activity, e.g., converting A·T to G·C.

[0080] The term "base editor system" refers to an intermolecular complex for editing nucleobases of a target nucleotide sequence. In various embodiments, a base editor (BE) system comprises: (1) a polynucleotide-programmable nucleotide-binding domain, a deaminase domain (e.g., cytidine deaminase or adenosine deaminase) for deaminating nucleobases in a target nucleotide sequence; and (2) one or more guide polynucleotides (e.g., guide RNAs) in combination with the polynucleotide-programmable nucleotide-binding domain. In various embodiments, the base editor (BE) system comprises a nucleobase editor domain selected from adenosine deaminase or cytidine deaminase, and a domain with nucleic acid sequence-specific binding activity. In some embodiments, the base editor system comprises: (1) a base editor (BE) comprising a polynucleotide-programmable DNA-binding domain and a deaminase domain for deaminating one or more nucleobases in a target nucleotide sequence; and (2) one or more guide RNAs in combination with the polynucleotide-programmable DNA-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable DNA-binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE) or a cytidine base editor (CBE).

[0081] "Base editing activity" refers to acting to chemically modify a base within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is cytidine deaminase activity, e.g., converting a target C·G to T·A. In another embodiment, the base editing activity is adenosine deaminase activity, e.g., converting A·T to G·C.

[0082] The term "Cas9" or "Cas9 domain" refers to an RNA-guided nuclease that includes a Cas9 protein or a fragment thereof (e.g., a protein that includes an active, inactive, or partially active DNA-cleavage domain of Cas9 and / or a gRNA-binding domain of Cas9). Cas9 nucleases are also sometimes referred to as casnl nucleases or CRISPR (clustered regularly interspaced short palindromic repeats)-associated nucleases.

[0083] The term "conservative amino acid substitution" or "conservative mutation" refers to the substitution of one amino acid for another amino acid that shares common properties. A functional method for defining common properties between individual amino acids is to analyze the normalized frequency of amino acid changes between corresponding proteins of homologous organisms (Schulz, GE and Schirmer, RH, Principles of Protein Structure, Springer-Verlag, New York (1979)). According to such analysis, groups of amino acids can be defined when amino acids within the group are preferentially exchanged with each other, and therefore have the most similar effects on the overall protein structure (Schulz, GE and Schirmer, RH, supra). Non-limiting examples of conservative mutations include amino acid substitutions, such as arginine to lysine, which can maintain a positive charge, and vice versa, aspartic acid to glutamic acid, which can maintain a negative charge, threonine to serine, which can maintain a free -OH, and asparagine to glutamine, which can maintain a free -NH.

[0084] The terms "coding sequence" or "protein-coding sequence," as used interchangeably herein, refer to a segment of a polynucleotide that encodes a protein. A coding sequence may also be referred to as an open reading frame. A region or sequence is bounded proximal to the 5' end by a start codon and proximal to the 3' end by a stop codon. Stop codons useful in the base editors described herein include the following:

[0085] TIFF2024534639000004.tif51165

[0086] A "complex" refers to a combination of two or more molecules whose interaction relies on intermolecular forces. Non-limiting examples of intermolecular forces include covalent and non-covalent interactions. Non-limiting examples of non-covalent interactions include hydrogen bonds, ionic bonds, halogen bonds, hydrophobic bonds, van der Waals interactions (e.g., dipole-dipole interactions, dipole-induced dipole interactions, and London dispersion forces), and the π effect. In one embodiment, a complex comprises a polypeptide, a polynucleotide, or a combination of one or more polypeptides and one or more polynucleotides. In one embodiment, a complex comprises one or more polypeptides that associate to form a base editor (e.g., a nucleic acid-programmable DNA-binding protein such as Cas9, and a base editor comprising a deaminase) and a polynucleotide (e.g., a guide RNA). In one embodiment, the complex is held together by hydrogen bonds. It should be understood that one or more components of a base editor (e.g., a deaminase, or a nucleic acid-programmable DNA-binding protein) can be covalently or non-covalently associated. As an example, a base editor can include a deaminase covalently linked (e.g., by a peptide bond) to a nucleic acid-programmable DNA-binding protein. Alternatively, a base editor can include a deaminase and a nucleic acid-programmable DNA-binding protein that are non-covalently associated (e.g., when one or more components of the base editor are provided in trans and associated directly or via another molecule, such as a protein or nucleic acid). In one embodiment, one or more components of the complex are held together by hydrogen bonds.

[0087] "Cytosine" or "4-aminopyrimidin-2(1H)-one" has the molecular formula C4H5N3O and the structure [ka] and refers to the purine nucleobase having the formula:

[0088] "Cytidine" is attached to the ribose sugar via a glycosidic bond and has the structure [ka] and corresponds to CAS number 65-46-3. Its molecular formula is CH 13 It is N3O5.

[0089] "Cytidine base editor (CBE)" means a base editor that includes a cytidine deaminase.

[0090] By "cytidine base editor (CBE) polynucleotide" is meant a polynucleotide that comprises a CBE.

[0091] "Cytidine deaminase" or "cytosine deaminase" refers to a polypeptide or fragment thereof capable of deaminating cytidine or cytosine. In one embodiment, cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine. The terms "cytidine deaminase" and "cytosine deaminase" are used interchangeably throughout this application. Petromyzon marinus cytosine deaminase 1 (PmCDA1) (SEQ ID NOS: 13-14), activation-induced cytidine deaminase (AICDA) (SEQ ID NOS: 15-21), and APOBEC (SEQ ID NOS: 12-61) are exemplary cytidine deaminases. Further exemplary cytidine deaminase (CDA) sequences are set forth in the Sequence Listing as SEQ ID NOS: 62-66 and 67-189. In an embodiment, the cytidine deaminase is BE4, rrA3F, or ppApo.

[0092] By "rra3F polypeptide" is meant a cytidine deaminase having cytidine deaminase activity and having an amino acid sequence having about or at least about 85%, 90%, 95%, 99%, or 100% sequence identity to the following sequence: MKPQIRDHRPNPMEAMYPHIFYFHFENLEKAYGRNETWLCFTVEIIKQYLPVPWKKGVFRNQVDPETHCHAEKCFLSWFCNNTLSPKKNYQVTWYTSWSPCPECAGEVAEFLAEHSNVKLTIYTARLYYFWDTDYQEGLRSLSEEGASVEIMDYEDFQYCWENFVYDDGEPFKRWKGLKYNFQSLTRRLREILQ (SEQ ID NO: 67).

[0093] By "ppAPOBEC-1 (ppApo)" is meant a cytidine deaminase having cytidine deaminase activity and having an amino acid sequence having about or at least about 85%, 90%, 95%, 99%, or 100% sequence identity to the following sequence: MTSEKGPSTGDPTLRRRIESWEFDVFYDPRELRKETCLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERRFHSSISCSITWFLSWSPCWECSQAIREFLSQHPGVTLVIYVARLFWHMDQRNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYPPGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLAFFRLHLQNCHYQTIPPHILLATGLIHPSVTWR (SEQ ID NO: 23).

[0094] "Cytosine" means a pyrimidine nucleobase having the molecular formula C4H5N3O.

[0095] "Cytosine deaminase activity" means catalyzing the deamination of cytosine or cytidine. In one embodiment, a polypeptide having cytosine deaminase activity converts an amino group to a carbonyl group. In one embodiment, a cytosine deaminase converts cytosine to uracil (i.e., C to U) or 5-methylcytosine to thymine (i.e., 5mC to T). In some embodiments, the cytosine deaminase provided herein has increased cytosine deaminase activity (e.g., at least 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, or more) relative to a reference cytosine deaminase.

[0096] As used herein, the term "deaminase" or "deaminase domain" refers to a protein or fragment thereof that catalyzes a deamination reaction.

[0097] "Detection" refers to identifying the presence, absence, or amount of an analyte being detected. In one embodiment, a sequence alteration in a polynucleotide or polypeptide is detected. In another embodiment, the presence of an indel is detected.

[0098] "Detectable label" refers to a composition that, when linked to a molecule of interest, renders the molecule of interest detectable by spectroscopic, photochemical, biochemical, immunochemical, or chemical means. For example, useful labels include radioisotopes, magnetic beads, metal beads, colloidal particles, fluorescent dyes, electron-dense reagents, enzymes (e.g., enzymes commonly used in enzyme-linked immunosorbent assays (ELISAs)), biotin, digoxigenin, or haptens.

[0099] "Disease" means any condition or disorder that damages or interferes with the normal function of a cell, tissue, or organ. Examples of diseases include HBV infection and related diseases and disorders, including cirrhosis, hepatocellular carcinoma (HCC), and any other disease associated with or resulting from HBV infection.

[0100] An "effective amount" refers to the amount required to alleviate the symptoms of a disease compared to an untreated patient. The effective amount of the active compound(s) used to practice the present invention for the therapeutic treatment of a disease will vary depending on the mode of administration, the age, weight, and general health of the subject. Ultimately, the attending physician or veterinarian will determine the appropriate amount and dosing regimen. Such an amount is referred to as an "effective" amount. In one embodiment, an effective amount is the amount of a base editor of the present invention that is sufficient to introduce a modification in the HBV genome in a cell (e.g., a cell in vitro or in vivo). In one embodiment, an effective amount is the amount of a base editor required to achieve a therapeutic effect (e.g., to reduce or control HBV infection). Such a therapeutic effect need not be sufficient to modify the HBV genome in all cells of a subject, tissue, or organ, but may be sufficient to modify the HBV genome in about 1%, 5%, 10%, 25%, 50%, 75% or more of the cells present in the subject, tissue, or organ. In one embodiment, an effective amount is sufficient to ameliorate one or more symptoms of HBV.

[0101] "Entecavir" is a compound with the structure corresponding to CAS number 142217-69-4. [ka] and pharmaceutically acceptable salts thereof. Entecavir has antiretroviral activity.

[0102] By "fragment" is meant a portion of a polypeptide or nucleic acid molecule, which portion contains at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the entire length of the reference nucleic acid molecule or polypeptide. Fragments can contain 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids.

[0103] "Guide polynucleotide" refers to a polynucleotide or polynucleotide complex that is specific to a target sequence and can form a complex with a polynucleotide-programmable nucleotide-binding domain protein (e.g., Cas9 or Cpf1). In one embodiment, the guide polynucleotide is a guide RNA (gRNA). The gRNA can exist as a complex of two or more RNAs or as a single RNA molecule.

[0104] "Hybridization" refers to hydrogen bonding, which may be Watson-Crick, Hoogsteen, or reversed Hoogsteen hydrogen bonding, between complementary nucleobases. For example, adenine and thymine are complementary nucleobases that pair through the formation of hydrogen bonds.

[0105] By "increase" is meant a positive change of at least 10%, 25%, 50%, 75%, or 100%.

[0106] The terms "inhibitor of base repair," "base repair inhibitor," "IBR," or grammatical equivalents thereof, refer to proteins capable of inhibiting the activity of nucleic acid repair enzymes, e.g., base excision repair enzymes.

[0107] An "intein" is a fragment of a protein that can excise itself and join the remaining fragment (an extein) with a peptide bond in a process known as protein splicing.

[0108] The terms "isolated," "purified," or "biologically pure" refer to material that is free, to varying degrees, from components that normally accompany it as found in its native state. "Isolated" refers to a degree of separation from the original source or surroundings. "Purified" refers to a degree of separation greater than isolation. A "purified" or "biologically pure" protein is sufficiently free from other substances so that any impurities do not substantially affect the biological properties of the protein or cause other adverse events. That is, the nucleic acids or peptides of the invention, if produced by recombinant DNA technology, are purified to be substantially free of cellular material, viral material, or culture medium, or, if chemically synthesized, are purified to be substantially free of chemical precursors or other chemicals. Purity and homogeneity are typically determined using analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high-performance liquid chromatography. The term "purified" can indicate that the nucleic acid or protein gives rise to essentially one band in an electrophoretic gel. For example, in the case of proteins that can be subject to modifications such as phosphorylation or glycosylation, different modifications can give rise to different isolated proteins that can be separately purified.

[0109] An "isolated polynucleotide" refers to a nucleic acid (e.g., DNA) that is free of the genes that flank it in the naturally occurring genome of the organism from which the nucleic acid molecule of the present invention is derived. Thus, the term includes, for example, a recombinant DNA that is integrated into a vector, an autonomously replicating plasmid, or a virus, or into the genomic DNA of a prokaryote or eukaryote, or that exists as a separate molecule independent of other sequences (e.g., cDNA, or a genomic or cDNA fragment produced by PCR or restriction endonuclease digestion). In addition, the term includes RNA molecules transcribed from DNA molecules, as well as recombinant DNA that is part of a hybrid gene encoding additional polypeptide sequences.

[0110] By "isolated polypeptide" is meant a polypeptide of the invention separated from components which naturally accompany it. Generally, a polypeptide is isolated when it is at least 60%, by weight, free from the proteins and naturally occurring organic molecules with which it is naturally associated. Preferably, a preparation is at least 75%, by weight, more preferably at least 90%, and most preferably at least 99%, by weight, the polypeptide of the invention. Isolated polypeptides of the invention can be obtained, for example, by extraction from a natural source, by expression of a recombinant nucleic acid encoding such a polypeptide, or by chemically synthesizing the protein. Purity can be measured by any appropriate method, for example, column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.

[0111] "Lamivudine" or "3TC" refers to the structure corresponding to CAS number 134678-17-4 and having the IUPAC name 2',3'-dideoxy-3'-thiacytidine 4-amino-1-[(2R,5S)-2-(hydroxymethyl)-1,3-oxathiolan-5-yl]-1,2-dihydropyrimidin-2-one, and pharmaceutically acceptable salts thereof. [ka] Lamivudine has antiretroviral activity and is classified as a nucleoside / nucleotide reverse transcriptase inhibitor (NRTI).

[0112] As used herein, the term "linker" refers to a molecule that connects two moieties. In embodiments, the term "linker" refers to a covalent linker (e.g., a covalent bond) or a non-covalent linker.

[0113] "Marker" refers to any protein or polynucleotide having an altered expression level or activity associated with a disease or disorder. Examples of diseases include HBV infection and related diseases and disorders, including cirrhosis, hepatocellular carcinoma (HCC), and any other disease associated with or resulting from HBV infection. Markers can be HBV polynucleotides and / or polypeptides. Non-limiting examples of hepatitis B virus markers include HBsAg, HBeAg, core protein, 3.5 kb viral RNA, and cccDNA (see, e.g., Bai, et al. "Quantification of Pregenomic RNA and Covalently Closed Circular DNA in Hepatitis B Virus-Related Hepatocellular Carcinoma," Int J Hepatol, 2013:849290 (2013), DOI:10.1155 / 2013 / 849290).

[0114] As used herein, the term "mutation" refers to the substitution of a residue in a sequence, e.g., a nucleic acid or amino acid sequence, with another residue, or the deletion or insertion of one or more residues in a sequence. Mutations are typically described herein by identifying the original residue, followed by the position of the residue in the sequence, and then identifying the newly substituted residue. Various methods for making amino acid substitutions (mutations) provided herein are well known in the art and are described, for example, in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4 th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012).

[0115] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to a compound comprising a nucleobase and an acidic moiety, e.g., a nucleoside, a nucleotide, or a polymer of nucleotides. Typically, polymeric nucleic acids, e.g., nucleic acid molecules comprising three or more nucleotides, are linear molecules in which adjacent nucleotides are linked to each other via phosphodiester bonds. In some embodiments, "nucleic acid" refers to an individual nucleic acid residue (e.g., a nucleotide and / or a nucleoside). In some embodiments, "nucleic acid" refers to an oligonucleotide chain comprising three or more individual nucleotide residues. As used herein, the terms "oligonucleotide" and "polynucleotide" can be used interchangeably to refer to a polymer of nucleotides (e.g., a string of at least three nucleotides). In some embodiments, "nucleic acid" encompasses RNA and single- and / or double-stranded DNA. Nucleic acids can occur naturally, for example, in the context of a genome, a transcript, mRNA, tRNA, rRNA, siRNA, snRNA, a plasmid, a cosmid, a chromosome, a chromatid, or other naturally occurring nucleic acid molecule. Alternatively, a nucleic acid molecule may be a non-naturally occurring molecule, e.g., recombinant DNA or RNA, an artificial chromosome, an engineered genome or fragment thereof, or a synthetic DNA, RNA, or DNA / RNA hybrid, or may contain non-naturally occurring nucleotides or nucleosides. Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms include nucleic acid analogs, e.g., analogs having other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems, and optionally purified or chemically synthesized. Where appropriate, for example, in the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, such as analogs having chemically modified bases or sugars, and backbone modifications. Nucleic acid sequences are presented in the 5' to 3' direction unless otherwise indicated.In some embodiments, nucleic acids are selected from natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), nucleotide analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C ...bromouridine, C5-bromouridine, C5-bromouridine, C5-bromouridine, C5-bromouridine, C5-brom -aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalated bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose), and / or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages).

[0116] The term "nuclear localization sequence," "nuclear localization signal," or "NLS" refers to an amino acid sequence that promotes the import of proteins into the cell nucleus. Nuclear localization sequences are known in the art and are described, for example, in International PCT Application PCT / EP2000 / 011690, filed November 23, 2000 by Plank et al. (published May 31, 2001 as WO / 2001 / 038547), the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, for example, as described in Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises the amino acid sequence KRTADGSEFESPKKKRKV (SEQ ID NO: 190), KRPAATKKAGQAKKKK (SEQ ID NO: 191), KKTELQTTNAENKTKKL (SEQ ID NO: 192), KRGINDRNFWRGENGRKTR (SEQ ID NO: 193), RKSGKIAAIVVKRPRK (SEQ ID NO: 194), PKKKRKV (SEQ ID NO: 195), or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 196).

[0117] The terms "nucleobase," "nitrogenous base," or "base" are used interchangeably herein to refer to nitrogen-containing biological compounds that form nucleosides, the building blocks of nucleotides. The ability of nucleobases to base pair and stack with one another directly leads to long-chain helical structures such as ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). The five nucleobases, adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), are referred to as primary or canonical. Adenine and guanine are derived from purines, while cytosine, uracil, and thymine are derived from pyrimidines. DNA and RNA can also contain other (non-primary) modified bases. Non-limiting exemplary modified nucleobases include hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydromethylcytosine. Both hypoxanthine and xanthine can be produced through deamination (replacing an amine group with a carbonyl group) in the presence of mutagens. Hypoxanthine can be modified from adenine. Xanthine can be modified from guanine. Uracil can result from the deamination of cytosine. A "nucleoside" consists of a nucleobase and a five-carbon sugar (either ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxyuridine, and deoxycytidine. Examples of nucleosides having modified nucleobases include inosine (I), xanthosine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), and pseudouridine (Ψ). A "nucleotide" consists of a nucleobase, a five-carbon sugar (either ribose or deoxyribose), and at least one phosphate group.

[0118] The term "nucleic acid programmable DNA binding protein" or "napDNAbp" may be used interchangeably with "polynucleotide programmable nucleotide binding domain" and may refer to a protein that associates with a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid or guide polynucleotide (e.g., gRNA), that guides the napDNAbp to a specific nucleic acid sequence. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable DNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable RNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a Cas9 protein. The Cas9 protein can associate with a guide RNA that guides the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the napDNAbp is a Cas9 domain, e.g., a nuclease-active Cas9, a Cas9 nickase (nCas9), or a nuclease-inactive Cas9 (dCas9). Non-limiting examples of nucleic acid programmable DNA binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ (Cas12j / Casphi).Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Cas12j / CasΦ, Cpf1, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csn3, Csn4, Csn5, Csn6, Csn7, Csn8, Csn9, Csn10, Csn11, Csn12, Csn13, Csn14, Csn15, Csn16, Csn17, Csn18, Csn19, Csn11, Csn11, Csn12, Csn13, Csn14, Csn15, Csn16, Csn17, Csn18, Csn19 ...9, Csn11, Csn12, Csn13, Csn14, Csn15, Csn Examples include sn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V Cas effector proteins, type VI Cas effector proteins, CARF, DinG, homologs thereof, or modified or engineered versions thereof. Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, although they may not be specifically listed herein. See, for example, Makarova et al., "Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?" CRISPRJ.2018 Oct;1:325-336.doi:10.1089 / crispr.2018.0033; Yan et al., "Functionally diverse type V CRISPR-Cas systems" Science.2019 Jan 4;363(6422):88-91.doi:10.1126 / science.aav7271 (the entire contents of each are incorporated herein by reference).Exemplary nucleic acid programmable DNA binding proteins and nucleic acid sequences encoding the nucleic acid programmable DNA binding proteins are provided in the Sequence Listing as SEQ ID NOs: 197-230.

[0119] As used herein, the term "nucleobase editing domain" or "nucleobase editing protein" refers to a protein or enzyme that can catalyze nucleobase modifications in RNA or DNA, such as the deamination of cytosine (or cytidine) to uracil (or uridine) or thymine (or thymidine), and the deamination of adenine (or adenosine) to hypoxanthine (or inosine), as well as non-templated nucleotide addition and insertion. In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., adenine deaminase or adenosine deaminase; or cytidine deaminase or cytosine deaminase).

[0120] As used herein, "obtaining" in "obtaining a drug" includes synthesizing, purchasing, or otherwise obtaining a drug.

[0121] "Subject" means a mammal, including, but not limited to, a human or non-human mammal (e.g., bovine, equine, canine, bovine, rodent, or feline). In one embodiment, "patient" refers to a mammalian subject who has a higher-than-average likelihood of developing a disease or disorder. Exemplary patients can be humans, non-human primates, felines, canines, porcines, bovines, felines, equines, camels, llamas, goats, sheep, rodents (e.g., mice, rabbits, rats, or guinea pigs), and other mammals that can benefit from the therapies disclosed herein. Exemplary human patients can be male and / or female. Exemplary diseases include HBV infection and associated diseases and disorders, including cirrhosis, hepatocellular carcinoma (HCC), and any other disease associated with or resulting from HBV infection.

[0122] A "patient in need thereof" or "subject in need thereof" as used herein refers to a patient who has been diagnosed with, is at risk of or has, has been predetermined to have, or is suspected of having a disease or disorder.

[0123] The terms "pathogenic mutation," "pathogenic variant," "disease casing mutation," "pathogenic variant," "deleterious mutation," or "predisposing mutation" refer to a genetic change or mutation that increases an individual's susceptibility or predisposition to a particular disease or disorder. In some embodiments, a pathogenic mutation comprises at least one wild-type amino acid substituted with at least one pathogenic amino acid in a protein encoded by the gene.

[0124] The terms "protein," "peptide," "polypeptide," and their grammatical equivalents are used interchangeably herein to refer to a polymer of amino acid residues linked together by peptide (amide) bonds. A protein, peptide, or polypeptide may be naturally occurring, recombinant, or synthetic, or any combination thereof.

[0125] As used herein, the term "fusion protein" refers to a hybrid polypeptide that contains protein domains derived from at least two different proteins.

[0126] The term "recombinant" as used herein in the context of a protein or nucleic acid refers to a protein or nucleic acid that does not occur in nature but is the product of human engineering. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that contains at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations compared to any naturally occurring sequence.

[0127] By "reduction" is meant a negative change of at least 10%, 25%, 50%, 75%, or 100%.

[0128] "Reference" refers to a standard or control condition. In one embodiment, the reference is a wild-type or healthy cell. In another embodiment, the reference is an untreated cell that is not subjected to the test condition, or is subjected to a placebo or saline, medium, buffer, and / or a control vector that does not carry the polynucleotide of interest. The reference can be a cell or subject without HBV infection. The reference can be a subject before undergoing treatment or before modifying the treatment.

[0129] A "reference sequence" is a defined sequence used as the basis for sequence comparison. A reference sequence may be a subset of a particular sequence or its entirety, for example, a segment of a full-length cDNA or gene sequence, or a complete cDNA or gene sequence. For polypeptides, the length of a reference polypeptide sequence is generally at least about 16 amino acids, at least about 20 amino acids, at least about 25 amino acids, at least about 35 amino acids, at least about 50 amino acids, or at least about 100 amino acids. For nucleic acids, the length of a reference nucleic acid sequence is generally at least about 50 nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, at least about 100 nucleotides, or at least about 300 nucleotides, or any integer thereabout or therebetween. In some embodiments, the reference sequence is the wild-type sequence of a protein of interest. In other embodiments, the reference sequence is a polynucleotide sequence encoding a wild-type protein.

[0130] The terms "RNA-programmable nuclease" and "RNA-guided nuclease" are used in conjunction with (e.g., bound to or associated with) one or more RNA(s) that are not targets for cleavage. In some embodiments, when an RNA-programmable nuclease is complexed with an RNA, it may be referred to as a nuclease:RNA complex. Typically, the bound RNA(s) are referred to as guide RNAs (gRNAs). In some embodiments, the RNA-programmable nuclease is a (CRISPR-associated system) Cas9 endonuclease, for example, Cas9 (Csnl) from Streptococcus pyogenes.

[0131] The term "single nucleotide polymorphism (SNP)" refers to a single nucleotide variation that occurs at a specific location in the genome, with each variation occurring to some degree (eg, >1%) in the population.

[0132] By "specifically binds" is meant that in a sample, e.g., a biological sample, a nucleic acid molecule, polypeptide, polypeptide / polynucleotide complex, compound or molecule recognizes and binds to a polypeptide and / or nucleic acid molecule of the invention, but does not substantially recognize or substantially bind to other molecules.

[0133] "Substantially identical" refers to a polypeptide or nucleic acid molecule that exhibits at least 50% identity to a reference amino acid sequence. In one embodiment, the reference sequence is a wild-type amino acid or nucleic acid sequence. In another embodiment, the reference sequence is any one of the amino acid or nucleic acid sequences described herein. In one embodiment, such a sequence is at least 60%, 80%, 85%, 90%, 95%, or even 99% identical at the amino acid or nucleic acid level to the sequence used for comparison.

[0134] Sequence identity is typically measured using sequence analysis software (e.g., the BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs in the sequence analysis software package of the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705). Such software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. An exemplary approach for determining the degree of identity may use the BLAST program, with a probability score of e-3 to e-100 indicating closely related sequences.

[0135] COBALT is used, for example, with the following parameters: a) Alignment parameters: gap penalty -11, -1 and end gap penalty -5, -1; b) CDD parameters: Use RPS BLAST, Blast E-value 0.003, find conserved columns and recalculate, and c) Query clustering parameters: Use query cluster, word size 4, maximum cluster distance 0.8, alphabet normal. EMBOSS Needle is used, for example, with the following parameters: a) Matrix: BLOSUM62; b) GAP OPEN: 10, c) GAP EXTEND: 0.5; d) OUTPUT FORMAT: vs. e) END GAP PENALTY: FALSE, f) END GAP OPEN: 10, and g)END GAP EXTEND:0.5.

[0136] Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to an endogenous nucleic acid sequence, but will typically exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence will typically be able to hybridize with at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to an endogenous nucleic acid sequence, but will typically exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence will typically be able to hybridize with at least one strand of a double-stranded nucleic acid molecule. "Hybridizing" means that the pair forms a double-stranded molecule between complementary polynucleotide sequences (e.g., genes described herein) or portions thereof under various stringency conditions. (See, e.g., Wahl, GM and SL Berger (1987) Methods Enzymol. 152:399; Kimmel, AR (1987) Methods Enzymol. 152:507).

[0137] For example, stringent salt concentrations are typically less than about 750 mM NaCl and less than 75 mM trisodium citrate, preferably less than about 500 mM NaCl and less than 50 mM trisodium citrate, and more preferably less than about 250 mM NaCl and less than 25 mM trisodium citrate. Low stringency hybridization can be achieved in the absence of organic solvents, such as formamide, while high stringency hybridization can be achieved in the presence of at least about 35% formamide, more preferably at least about 50% formamide. Stringent temperature conditions typically include a temperature of at least about 30°C, more preferably at least about 37°C, and most preferably at least about 42°C. Various additional parameters, such as hybridization time, detergent concentration, e.g., sodium dodecyl sulfate (SDS), and the inclusion or exclusion of carrier DNA, are well known to those skilled in the art. Various levels of stringency can be achieved by combining these various conditions as needed. In a preferred embodiment, hybridization is carried out in 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS at 30° C. In a more preferred embodiment, hybridization is carried out in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA) at 37° C. In a most preferred embodiment, hybridization is carried out in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 μg / ml ssDNA at 42° C. Useful variations on these conditions will be readily apparent to those of skill in the art.

[0138] In most applications, the washing steps following hybridization will also vary in stringency. Stringency conditions for washing can be defined by salt concentration and temperature. As noted above, washing stringency can be increased by decreasing salt concentration or increasing temperature. For example, stringent salt concentrations for washing steps are preferably less than about 30 mM NaCl and less than 3 mM trisodium citrate, and most preferably less than about 15 mM NaCl and less than 1.5 mM trisodium citrate. Stringent temperature conditions for washing steps typically include temperatures of at least about 25°C, more preferably at least about 42°C, and even more preferably at least about 68°C. In one embodiment, washing steps are performed at 25°C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In another embodiment, washing steps are performed at 42°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the wash step is carried out in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS at 68°C. Additional variations in these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science 196:180, 1977), Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975), Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001), Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York), and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.

[0139] "Split" means divided into two or more pieces.

[0140] "Split Cas9 protein" or "split Cas9" refers to a Cas9 protein that is provided as an N-terminal fragment and a C-terminal fragment that are encoded by two separate nucleotide sequences. Polypeptides corresponding to the N-terminal and C-terminal portions of the Cas9 protein can be spliced ​​to form a "reassembled" Cas9 protein.

[0141] The term "target site" refers to a sequence within a nucleic acid molecule that is deaminated by a deaminase (e.g., a cytidine deaminase or an adenine deaminase) or a fusion protein that includes a deaminase (e.g., a dCas9-adenosine deaminase fusion protein or a base editor disclosed herein).

[0142] "Tenofovir" is a compound with the structure corresponding to CAS number 147127-20-6. [ka] and pharmaceutically acceptable salts thereof. Tenofovir has antiretroviral activity. Tenofovir is classified as a nucleoside / nucleotide reverse transcriptase inhibitor (NRTI).

[0143] As used herein, terms such as "treat," "treating," and "treatment" refer to reducing or ameliorating a disorder and / or its associated symptoms, or achieving a desired pharmacological and / or physiological effect. It will be understood that treating a disorder or condition does not necessarily require, but does not exclude, that the disorder, condition, or its associated symptoms be completely eliminated. In some embodiments, the effect is therapeutic, i.e., without limitation, the effect partially or completely reduces, diminishes, suppresses, alleviates, alleviates, reduces the intensity of, or cures, the disease and / or adverse symptoms resulting from the disease. In some embodiments, the effect is prophylactic, i.e., the effect protects against or prevents the occurrence or recurrence of the disease or condition. To this end, the methods disclosed herein comprise administering a therapeutically effective amount of a composition as described herein. In one embodiment, the present invention provides treatment of HBV infection.

[0144] "Uracil glycosylase inhibitor" or "UGI" refers to an agent that inhibits the uracil excision repair system. Base editors containing cytidine deaminase convert cytosine to uracil, which is then converted to thymine during DNA replication or repair. Inclusion of an inhibitor of uracil DNA glycosylase (UGI) in a base editor prevents base excision repair, which changes U back to C. An exemplary UGI comprises the following amino acid sequence: >splP14739IUNGI_BPPB2 Uracil DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSD APEYKPWALVIQDSNGENKIKML (SEQ ID NO: 231).

[0145] Ranges provided herein are understood to be shorthand for all values ​​within the range. For example, a range of 1 to 50 is understood to include any number, combination of numbers, or subrange from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.

[0146] The recitation of a list of chemical groups within any definition of a variable herein includes the definition of that variable as any single group or combination of listed groups. The recitation of an embodiment for a variable or aspect herein includes that embodiment as any single embodiment or in combination with any other embodiment or portion thereof.

[0147] All terms are intended to be understood as understood by one of ordinary skill in the art. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0148] In this application, the use of the singular includes the plural unless specifically stated otherwise. It should be noted that as used herein, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. In this application, the use of "or" means "and / or" unless specifically stated otherwise. Furthermore, the use of the term "including," as well as other forms such as "include," "includes," and "included," is not limiting.

[0149] As used in this specification and claims, the words "comprising" (and any of its forms, such as "comprise" and "comprises"), "having" (and any of its forms, such as "have" and "has"), "including" (and any of its forms, such as "includes" and "include"), or "containing" (and any of its forms, such as "contains" and "contain") are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. It is contemplated that any embodiment discussed herein can be implemented with respect to any method or composition of the disclosure, and vice versa. Furthermore, the compositions of the disclosure can be used to achieve the methods of the disclosure.

[0150] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined (i.e., the limitations of the measurement system). For example, "about" can mean within one standard deviation or within more than one standard deviation, in accordance with practice in the art. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, e.g., within 5-fold or 2-fold of a value. When a particular value is described in this application and claims, unless otherwise specified, it means within an acceptable error range for the particular value.

[0151] References herein to "some embodiments," "embodiments," "one embodiment," or "other embodiments" mean that the particular feature, structure, or characteristic described in connection with the embodiment is included in at least some embodiments of the present disclosure, but not necessarily in all embodiments. [Brief explanation of the drawings]

[0152] [Figure 1]

[0023] Figure 1 shows the partially double-stranded, overlapping open reading frames (ORFs) for the hepatitis B surface antigen (HBsAg) gene, polymerase gene, protein X gene, and core gene. The HBsAg gene contains ORFs PreS1, PreS2, and S, which encode the large, middle, and small surface proteins, respectively. ORFs core and PreC encode the capsid protein. [Figure 2] Figure 1 shows the HBV life cycle. The term "ER" refers to the endoplasmic reticulum. The term "HBsAg" refers to the hepatitis B surface antigen. "HBx transcriptional activator" is an HBV-specific transcriptional activator of the polymerase II and III promoters. [Figure 3A] Maps and presentation slides are presented. A map of the geographic distribution of Hepatitis B virus genotypes around the world. [Figure 3B] We present maps and presentation slides. We provide a summary of base editing strategies to introduce stop codons into viral genes and generate abasic sites to treat chronic HBV. [Figure 4] A bar graph showing the %A-G editing efficiency of the indicated guide RNA sequences used in combination with the indicated base editors is provided. The HEK293 lenti-HBV system was used in an RNA transfection format. Several of the guide RNA sequences were associated with editing efficiencies greater than 50%. The guide RNAs targeted HBV transcriptional elements or introduced missense mutations into HBV genes. Listed on the x-axis are the guide RNA sequences used in combination with the indicated base editors (see Table 1). [Figure 5A]

[0023] Figures 5A-5C provide bar graphs showing the antiviral efficacy of the indicated base editors used in combination with the indicated guide RNAs (listed on the x-axis) in the HBV-primary human hepatocyte (HBV-PHH) system. In Figures 5A-5C, the term "HBsAG" stands for "hepatitis B surface antigen," and the term "HBeAg" stands for "hepatitis B e antigen." In Figure 5, the y-axis represents the fold change relative to an untreated sample. [Figure 5B]

[0023] Figures 5A-5C provide bar graphs showing the antiviral efficacy of the indicated base editors used in combination with the indicated guide RNAs (listed on the x-axis) in the HBV-primary human hepatocyte (HBV-PHH) system. In Figures 5A-5C, the term "HBsAG" stands for "hepatitis B surface antigen," and the term "HBeAg" stands for "hepatitis B e antigen." In Figure 5, the y-axis represents the fold change relative to an untreated sample. [Figure 5C]

[0023] Figures 5A-5C provide bar graphs showing the antiviral efficacy of the indicated base editors used in combination with the indicated guide RNAs (listed on the x-axis) in the HBV-primary human hepatocyte (HBV-PHH) system. In Figures 5A-5C, the term "HBsAG" stands for "hepatitis B surface antigen," and the term "HBeAg" stands for "hepatitis B e antigen." In Figure 5, the y-axis represents the fold change relative to an untreated sample. [Figure 6] A schematic diagram outlining the experimental procedure for functional gRNA screening in HBV-infected HepG2-NTCP cells is provided. [Figure 7] 1 provides a schematic map of the HBV genome, annotated to show the locations of base edits associated with gRNA37 and gRNA40 used in combination with BE4. [Figure 8]A set of bar graphs shows the results of functional gRNA screening in HBV-infected HepG2pNTCP cells. Two lead gRNAs (i.e., gRNA37 and gRNA40) targeting the HBsAg and Core genes were identified. Guide gRNA37 (Stop-S) was associated with a reduction in HBsAg, and gRNA40 (Stop-PreCore) was associated with a reduction in HBeAg. Both guide RNAs, gRNA37 and gRNA40, were associated with a reduction in total HBV DNA and 3.5 kb RNA levels. The x-axis lists the gRNAs used in combination with BE4 (i.e., gRNA12, gRNA37, gRNA19, gRNA190, and gRNA40). The negative control (-) was untreated cells, and the control (contr) was a gRNA targeting the PCSK9 gene. [Figure 9] This figure provides a series of bar graphs showing that gRNA multiplexing simultaneously reduced each HBV viral parameter in HepG2-NTCP. The x-axis lists the gRNAs (i.e., gRNA37 and gRNA40) used in combination with BE4. Both gRNAs targeted the same covalently closed circular DNA strand (cccDNA), with target sequences spaced approximately 1 kb apart. Multiplexed gRNA37 and gRNA40 reduced HBsAg and HbBeAg, as well as total HBV DNA and 3.5 kb of RNA. The base editor used was BE4. The negative control (-) was untreated cells, and the control (contr) was a gRNA targeting the PCSK9 gene. [Figure 10]

[0023] Figure 1 provides a bar graph showing that a base editor system containing BE4 and guide RNA(s) performed cccDNA editing without reducing the level of cccDNA. The negative control (-) was untreated cells, and the control (contr) was a gRNA targeting the PCSK9 gene. [Figure 11] A schematic diagram summarizing the experimental protocol for assessing the impact of lamivudine pretreatment on editing is provided. [Figure 12]A bar graph showing that pretreatment with lamivudine was associated with increased editing efficiency. Base editing efficiency was improved by approximately 20% in HepG2-NTCP cells by pretreatment with lamivudine. High cccDNA editing in lamivudine-pretreated cells suggested that CBE directly targets cccDNA. [Figure 13] A set of bar graphs is provided showing that base editing using gRNA37 combined with gRNA40 combined with lamivudine pretreatment (i.e., combination / multiple gRNA treatment) was associated with a robust reduction in hepatitis B virus (HBV) markers. The negative control (-) was untreated cells, and the control (contr) was a gRNA targeting the PCSK9 gene. [Figure 14] A schematic summarizing the experimental protocol for evaluating the antiviral activity of base editing on lamivudine is provided. Figure 14 lists the HBV parameters measured. [Figure 15] A series of bar graphs are provided showing that base editing with gRNA37 in combination with gRNA40 (i.e., combination / multiple gRNA treatment) resulted in a 70% to 80% reduction in all HBV viral markers at 25 days post-treatment, whereas 8 days of treatment with lamivudine (LAM) did not. The negative control (-) was untreated cells, and the control (contr) was gRNA targeting the PCSK9 gene. [Figure 16] Plots are provided showing the time course of extracellular HBV DNA levels for the indicated treatments. Base editing prevented rebound in primary hepatocyte co-cultures. HBV rebounded after cessation of treatment with lamivudine alone. Control gRNA targeted the PCSK9 gene. "HBV" samples were untreated samples. No HBV rebound was observed two weeks after the second transfection with the base editing reagent. HBV replication was assessed by HBV DNA qPCR at s / n. [Figure 17]This figure shows bar graphs showing the individual base editing efficiencies at the sites targeted by gRNA37 and gRNA40 when the two gRNAs were used in combination with BE4 to edit HBV cccDNA. Approximately 30% editing efficiency was observed at the S antigen site, and approximately 60% editing efficiency was observed at the PreCore gene. These editing efficiencies were sufficient to enable high antiviral efficacy and prevent rebound in primary human hepatocytes (PHHs). [Figure 18] FIG. 1 is a schematic diagram providing an overview of an experiment evaluating combination treatment with lamivudine and a base editing system comprising BE4 and gRNAs 37 and 40. [Figure 19] Plots are provided showing the time course of extracellular HBV DNA levels for the indicated treatments. "HBV" samples were untreated samples. Pretreatment with lamivudine improved editing in primary hepatocytes. Combination treatment with the base editor system and lamivudine prevented HBV rebound. HBV replication was assessed by HBV DNA qPCR at s / n. [Figure 20] A bar graph showing base editing efficiency at sites targeted by gRNA37 and gRNA40 is provided. Cells were contacted with BE4 in combination with gRNA37 and gRNA40 and treated with lamivudine (LAM). Pretreatment with lamivudine improved editing in primary hepatocytes. Control cells (-) were not treated with lamivudine. [Figure 21] A bar graph showing that base editing in combination with lamivudine treatment in primary hepatocytes resulted in a reduction of all HBV markers by approximately 70% to approximately 80%. The designation "HBV" indicates the untreated control. [Figure 22]A bar graph showing base editing at sites targeted by gRNA37 and gRNA40 is provided. Cells were contacted with BE4, ppApo, or rrA3F in combination with gRNA37 and gRNA40. The use of ppApo and rrA3F base editors, which have reduced off-target activity and reduced nonspecific effects, instead of BE4 did not worsen base editing efficiency compared to BE4. All base editors were associated with comparable editing efficiency. [Figure 23]

[0039] Figure 10 provides a bar graph showing that contacting cells with base editors BE4, ppApo, or rrA3F in combination with gRNA37 and gRNA40 resulted in a reduction in HBsAg levels in HBV-infected cells. Thus, all three base editor systems had antiviral activity. The negative control (HBV) was untreated cells, and the control (contr) was a gRNA targeting the PCSK9 gene. [Figure 24A]Figure 24 provides plots showing in vivo levels of Hepatitis B Virus (HBV) antigen measured in a replication-competent HBV circle mouse model (n=5) on day 1 after injection of lipid nanoparticles containing a polynucleotide encoding BE4 and the guide RNAs gRNA37 and gRNA40. Figure 24A plots mean %HBsAg as IU / ml over time. Figure 24B plots HBeAg as PEI (Paul Elrich Institute International Standard Serum) U / ml over time. Figure 24C plots HBV DNA as copies / ml over time. In Figures 24A-24C, the term "HM" refers to gRNA37-HM and gRNA40-HM-containing scaffolds with "heavily modified" (i.e., heavily modified guide RNAs), the term "SOC" indicates "standard of care," and "SEM" indicates "standard error of the mean." In Figures 24A-24C, "1x" indicates only the first injection ("LNP-1st inj."), and "2x" indicates both the first injection ("LNP-1st inj.") and the 14th injection ("LNP 2nd inj."). [Figure 24B]Figure 24 provides plots showing in vivo levels of Hepatitis B Virus (HBV) antigen measured in a replication-competent HBV circle mouse model (n=5) on day 1 after injection of lipid nanoparticles containing a polynucleotide encoding BE4 and the guide RNAs gRNA37 and gRNA40. Figure 24A plots mean %HBsAg as IU / ml over time. Figure 24B plots HBeAg as PEI (Paul Elrich Institute International Standard Serum) U / ml over time. Figure 24C plots HBV DNA as copies / ml over time. In Figures 24A-24C, the term "HM" refers to gRNA37-HM and gRNA40-HM-containing scaffolds with "heavily modified" (i.e., heavily modified guide RNAs), the term "SOC" indicates "standard of care," and "SEM" indicates "standard error of the mean." In Figures 24A-24C, "1x" indicates only the first injection ("LNP-1st inj."), and "2x" indicates both the first injection ("LNP-1st inj.") and the 14th injection ("LNP 2nd inj."). [Figure 24C]Figure 24 provides plots showing in vivo levels of Hepatitis B Virus (HBV) antigen measured in a replication-competent HBV circle mouse model (n=5) on day 1 after injection of lipid nanoparticles containing a polynucleotide encoding BE4 and the guide RNAs gRNA37 and gRNA40. Figure 24A plots mean %HBsAg as IU / ml over time. Figure 24B plots HBeAg as PEI (Paul Elrich Institute International Standard Serum) U / ml over time. Figure 24C plots HBV DNA as copies / ml over time. In Figures 24A-24C, the term "HM" refers to gRNA37-HM and gRNA40-HM-containing scaffolds with "heavily modified" (i.e., heavily modified guide RNAs), the term "SOC" indicates "standard of care," and "SEM" indicates "standard error of the mean." In Figures 24A-24C, "1x" indicates only the first injection ("LNP-1st inj."), and "2x" indicates both the first injection ("LNP-1st inj.") and the 14th injection ("LNP 2nd inj."). [Figure 25A]Figure 25 provides plots showing in vivo levels of hepatitis B virus (HBV) antigen measured in a replication-competent HBV circle mouse model (n=5) after day 1 injection of lipid nanoparticles containing a polynucleotide encoding BE4 and the guide RNAs gRNA37 and gRNA40. Figure 25A plots mean %HBsAg as IU / mL over time. Figure 25B plots HBeAg as PEI (Paul Elrich Institute International Standard Serum) U / ml over time. Figure 25C plots HBV DNA as copies / ml over time. In Figures 25A-25C, the term "HM" refers to gRNA37-HM and gRNA40-HM-containing scaffolds with "heavily modified" (i.e., heavily modified guide RNAs), the term "ETV" refers to "entecavir," and "SD" refers to "standard deviation." In Figures 25A-25C, "1x" refers to the first injection only ("LNP-1st inj."), and "2x" refers to the first injection ("LNP-1st inj.") and the 14th injection ("LNP 2nd inj."). [Figure 25B]Figure 25 provides plots showing in vivo levels of hepatitis B virus (HBV) antigen measured in a replication-competent HBV circle mouse model (n=5) after day 1 injection of lipid nanoparticles containing a polynucleotide encoding BE4 and the guide RNAs gRNA37 and gRNA40. Figure 25A plots mean %HBsAg as IU / mL over time. Figure 25B plots HBeAg as PEI (Paul Elrich Institute International Standard Serum) U / ml over time. Figure 25C plots HBV DNA as copies / ml over time. In Figures 25A-25C, the term "HM" refers to gRNA37-HM and gRNA40-HM-containing scaffolds with "heavily modified" (i.e., heavily modified guide RNAs), the term "ETV" refers to "entecavir," and "SD" refers to "standard deviation." In Figures 25A-25C, "1x" refers to the first injection only ("LNP-1st inj."), and "2x" refers to the first injection ("LNP-1st inj.") and the 14th injection ("LNP 2nd inj."). [Figure 25C]Figure 25 provides plots showing in vivo levels of hepatitis B virus (HBV) antigen measured in a replication-competent HBV circle mouse model (n=5) after day 1 injection of lipid nanoparticles containing a polynucleotide encoding BE4 and the guide RNAs gRNA37 and gRNA40. Figure 25A plots mean %HBsAg as IU / mL over time. Figure 25B plots HBeAg as PEI (Paul Elrich Institute International Standard Serum) U / ml over time. Figure 25C plots HBV DNA as copies / ml over time. In Figures 25A-25C, the term "HM" refers to gRNA37-HM and gRNA40-HM-containing scaffolds with "heavily modified" (i.e., heavily modified guide RNAs), the term "ETV" refers to "entecavir," and "SD" refers to "standard deviation." In Figures 25A-25C, "1x" refers to the first injection only ("LNP-1st inj."), and "2x" refers to the first injection ("LNP-1st inj.") and the 14th injection ("LNP 2nd inj."). [Figure 26A] A schematic diagram illustrating the strategy for gRNA selection and screening is provided. This strategy involves targeting conserved HBV regions, focusing on genotype D, which is the most common worldwide and for which both cellular and animal models exist. In silico gRNA selection was based on the potential of the gRNA to silence HBV genes. One hundred gRNAs (NGG, NGA, NNNRRT PAM) were designed to introduce stop codons into viral genes. Twenty-four conserved gRNAs were predicted to introduce missense mutations into the HBV genome. The final step involved detecting edits in the Hek293-Lenti-HBV cell line. [Figure 26B]A schematic diagram illustrating the strategy for gRNA selection and screening is provided. This strategy involves targeting conserved HBV regions, focusing on genotype D, which is the most common worldwide and for which both cellular and animal models exist. In silico gRNA selection was based on the potential of the gRNA to silence HBV genes. One hundred gRNAs (NGG, NGA, NNNRRT PAM) were designed to introduce stop codons into viral genes. Twenty-four conserved gRNAs were predicted to introduce missense mutations into the HBV genome. The final step involved detecting edits in the Hek293-Lenti-HBV cell line. [Figure 27] Base editing prevents HBV rebound in primary human hepatocytes (PHH). A provides a schematic diagram showing the experimental schedule used in PHH. B is a graph showing HBV replication assessed by HBV DNA qPCR in PHH supernatants at different days after transfection. While 3TC withdrawal leads to HBV rebound, base editing prevents this rebound. C is a graph showing that base editing results in efficient reduction of HBsAg, HBeAg, 3.5kb RNA, and HBV DNA, further improving HBV replication inhibition in PHH. D is a graph showing that approximately 30% editing of the S antigen and approximately 60% editing of the PreCore gene were sufficient to enable high antiviral efficacy and prevent rebound in PHH. [Figure 28A] Figure 1 shows that multiplexing two lead gRNAs reduces HBV parameters in a hepatoma cell line (HepG2-NTCP). The experimental schedule is shown. [Figure 28B] Figure 1 shows that multiplexing two lead gRNAs reduces HBV parameters in a hepatoma cell line (HepG2-NTCP). Includes several graphs showing that base editing results in efficient reduction of viral extracellular (HBsAg, HBeAg) and intracellular (3.5 kb RNA, and HBV DNA) parameters. [Figure 28C]Figure 1 shows that multiplexing two lead gRNAs reduces HBV parameters in a hepatoma cell line (HepG2-NTCP). A series of graphs showing the results observed for HBsAg, HBeAg, and 3.5 kb RNA upon 3TC treatment are provided. [Figure 28D] Figure 1 shows that multiplexing two lead gRNAs reduces HBV parameters in a hepatoma cell line (HepG2-NTCP). The combination of gRNA B + gRNA H also inhibited all HBV isoforms, as observed by Western blotting. [Figure 29] Figure 1 shows that base editing reduced HBsAg from naturally integrated HBV. Figure 2A provides the experimental protocol used for PLC / PRF5 cells. Figure 2B is a graph showing that extracellular HBsAg levels were determined by ELISA 6 days after transfection of BE4 mRNA and gRNA B*. Figure 2C is a graph showing that editing approximately 50% of the S antigen site was sufficient to enable robust reduction of HBsAg. [Figure 30] This shows that the base editor functions through cccDNA editing without reducing cccDNA levels. (A) Graph. cccDNA levels were assessed by qPCR on DNA samples pretreated with ExoI / III to eliminate viral replication intermediates. No reduction in cccDNA was observed upon base editing with gRNA B + gRNA H in the absence or presence of 3TC. (B) Graph showing functional editing assessed by NGS on cccDNA-enriched samples. Higher editing efficiency was detected in the presence of 3TC. [Figure 31A]Figure 31A shows that multiplexing two gRNAs with the BE4 base editor simultaneously reduced HBV viral parameters in HepG2-NTCP. Figure 31A provides a series of graphs showing that base editing resulted in efficient reduction of viral extracellular (HBsAg, HBeAg) and intracellular (3.5 kb RNA, and HBV DNA) parameters compared to control samples treated with a base editing reagent targeting the unrelated PCSK9 gene. BE4 / gRNA (S1 + C2) treatment inhibited all HBV isoforms, as observed by Western blot. Figure 31B provides an experimental schedule for pretreatment with lamivudine. Figure 31C shows that combination therapy with lamivudine resulted in robust reduction of HBV viral markers, similar to the results shown in panel A. Figure 31D is a graph showing that base editing does not reduce cccDNA levels in HepG2-NTCP. Figure 31E shows that pretreatment with lamivudine increased base editing rates by 20% in HepG2-NTCP. Without intending to be bound by theory, the high cccDNA editing under lamivudine pretreatment conditions indicates that CBE likely targets cccDNA directly. [Figure 31B]Figure 31A shows that multiplexing two gRNAs with the BE4 base editor simultaneously reduced HBV viral parameters in HepG2-NTCP. Figure 31A provides a series of graphs showing that base editing resulted in efficient reduction of viral extracellular (HBsAg, HBeAg) and intracellular (3.5 kb RNA, and HBV DNA) parameters compared to control samples treated with a base editing reagent targeting the unrelated PCSK9 gene. BE4 / gRNA (S1 + C2) treatment inhibited all HBV isoforms, as observed by Western blot. Figure 31B provides an experimental schedule for pretreatment with lamivudine. Figure 31C shows that combination therapy with lamivudine resulted in robust reduction of HBV viral markers, similar to the results shown in panel A. Figure 31D is a graph showing that base editing does not reduce cccDNA levels in HepG2-NTCP. Figure 31E shows that pretreatment with lamivudine increased base editing rates by 20% in HepG2-NTCP. Without intending to be bound by theory, the high cccDNA editing under lamivudine pretreatment conditions indicates that CBE likely targets cccDNA directly. [Figure 31C]Figure 31A shows that multiplexing two gRNAs with the BE4 base editor simultaneously reduced HBV viral parameters in HepG2-NTCP. Figure 31A provides a series of graphs showing that base editing resulted in efficient reduction of viral extracellular (HBsAg, HBeAg) and intracellular (3.5 kb RNA, and HBV DNA) parameters compared to control samples treated with a base editing reagent targeting the unrelated PCSK9 gene. BE4 / gRNA (S1 + C2) treatment inhibited all HBV isoforms, as observed by Western blot. Figure 31B provides an experimental schedule for pretreatment with lamivudine. Figure 31C shows that combination therapy with lamivudine resulted in robust reduction of HBV viral markers, similar to the results shown in panel A. Figure 31D is a graph showing that base editing does not reduce cccDNA levels in HepG2-NTCP. Figure 31E shows that pretreatment with lamivudine increased base editing rates by 20% in HepG2-NTCP. Without intending to be bound by theory, the high cccDNA editing under lamivudine pretreatment conditions indicates that CBE likely targets cccDNA directly. [Figure 31D]Figure 31A shows that multiplexing two gRNAs with the BE4 base editor simultaneously reduced HBV viral parameters in HepG2-NTCP. Figure 31A provides a series of graphs showing that base editing resulted in efficient reduction of viral extracellular (HBsAg, HBeAg) and intracellular (3.5 kb RNA, and HBV DNA) parameters compared to control samples treated with a base editing reagent targeting the unrelated PCSK9 gene. BE4 / gRNA (S1 + C2) treatment inhibited all HBV isoforms, as observed by Western blot. Figure 31B provides an experimental schedule for pretreatment with lamivudine. Figure 31C shows that combination therapy with lamivudine resulted in robust reduction of HBV viral markers, similar to the results shown in panel A. Figure 31D is a graph showing that base editing does not reduce cccDNA levels in HepG2-NTCP. Figure 31E shows that pretreatment with lamivudine increased base editing rates by 20% in HepG2-NTCP. Without intending to be bound by theory, the high cccDNA editing under lamivudine pretreatment conditions indicates that CBE likely targets cccDNA directly. [Figure 31E]Figure 31A shows that multiplexing two gRNAs with the BE4 base editor simultaneously reduced HBV viral parameters in HepG2-NTCP. Figure 31A provides a series of graphs showing that base editing resulted in efficient reduction of viral extracellular (HBsAg, HBeAg) and intracellular (3.5 kb RNA, and HBV DNA) parameters compared to control samples treated with a base editing reagent targeting the unrelated PCSK9 gene. BE4 / gRNA (S1 + C2) treatment inhibited all HBV isoforms, as observed by Western blot. Figure 31B provides an experimental schedule for pretreatment with lamivudine. Figure 31C shows that combination therapy with lamivudine resulted in robust reduction of HBV viral markers, similar to the results shown in panel A. Figure 31D is a graph showing that base editing does not reduce cccDNA levels in HepG2-NTCP. Figure 31E shows that pretreatment with lamivudine increased base editing rates by 20% in HepG2-NTCP. Without intending to be bound by theory, the high cccDNA editing under lamivudine pretreatment conditions indicates that CBE likely targets cccDNA directly. [Figure 32A]

[00130] Figure 32A shows that base editing prevents viral rebound in PHH. Figure 32A is a graph showing HBV replication as assessed by HBV DNA qPCR in PHH supernatants. Discontinuation of lamivudine leads to HBV rebound, but base editing prevented viral rebound. [Figure 32B]

[0023] Figure 1 shows that base editing prevents viral rebound in PHH. Includes four graphs showing that base editing leads to efficient reduction of HBsAg, HBeAg, 3.5kb RNA, and HBV DNA. [Figure 32C]

[0023] Figure 1 shows that base editing prevents viral rebound in PHH.

[0024] Figure 1 shows that approximately 55% of the edited S antigen and approximately 80% of the edited PreCore gene are sufficient to enable high antiviral efficacy and prevent rebound in PHH. [Figure 33-1]Figure 33 shows that LNP-mediated delivery of BE4 mRNA and gRNA S1 / C2 results in a sustained reduction of viral markers in the HBV minicircle model. (A) Graph showing that there was a >2-log reduction in mean HBsAg. Significantly, 5 / 9 mice showed a reduction in HBsAg below the limit of detection. (B) Graph showing that HBV rebounds in the entecavir-treated group (positive control). In the base-edited treatment group, there was a sustained >3-log reduction in serum HBV DNA. No HBV rebound was observed in the base-edited group. (C) Loss of HBeAg expression in all mice below the limit of detection 2 weeks after the first LNP injection. Data presented as mean + / - SEM, n = 4 or 5 per group. [Figure 33-2] Figure 1 shows that LNP-mediated delivery of BE4 mRNA and gRNA S1 / C2 results in a sustained reduction of viral markers in the HBV minicircle model. C shows loss of HBeAg expression in all mice below the limit of detection 2 weeks after the first LNP injection. Data presented as mean + / - SEM, n = 4 or 5 per group. [Figure 34] 1A and 1B are graphs showing that BE4, in combination with gRNA EMSbeam12 having the spacer sequence of SEQ ID NO:578 and MSPbeam37-PLC having the spacer sequence of SEQ ID NO:10834, edited the naturally integrated HBV genome present in the Alexander hepatoma cell line PLC / PRF / 5 (A), which reduced HBsAg secretion (B). The protospacer sequence is also included. The MSPbeam37 protospacer corresponds to SEQ ID NO:508, and the MSPbeam37-PLC protospacer (SEQ ID NO:10833) contains a single nucleobase change relative to SEQ ID NO:508. [Figure 35] A schematic diagram showing the experimental design for evaluating base editing in HepG2.2.15 cells in vitro with or without lamivudine (LAM) pretreatment is provided. In Figure 35, "D0," "D1," etc., represent day 0, day 1, etc., from the start of the experiment, "SN" indicates "supernatant," and "6 dpt" indicates "6 days post-transduction." [Figure 36A]Figure 36A provides bar graphs showing HBsAg and HBeAg polypeptide levels in HepG2.2.15 cells after base editing using the BE4 base editor in combination with guide gRNA37 (g37) or gRNA12 (g12) (for base editing to reduce HBsAg expression) and / or in combination with guide gRNA40 (g40) (for base editing to reduce HBeAg expression). As a control, cells were contacted with the BE4 base editor in combination with a guide targeting the PCSK9 gene. Cells were base edited with and without pretreatment with lamivudine (LAM). Figure 36A shows HBsAg expression levels in HepG2.2.15 cells edited with the BE4 base editor in combination with the indicated guide, with and without pretreatment with LAM. Figure 36B shows HBeAg expression levels in HepG2.2.15 cells edited with and without LAM pretreatment using the BE4 base editor in combination with the indicated guides. Expression levels were compared using ANOVA / non-parametric test / no matched pairs / multiple comparisons with PCSK9. [Figure 36B]Figure 36A provides bar graphs showing HBsAg and HBeAg polypeptide levels in HepG2.2.15 cells after base editing using the BE4 base editor in combination with guide gRNA37 (g37) or gRNA12 (g12) (for base editing to reduce HBsAg expression) and / or in combination with guide gRNA40 (g40) (for base editing to reduce HBeAg expression). As a control, cells were contacted with the BE4 base editor in combination with a guide targeting the PCSK9 gene. Cells were base edited with and without pretreatment with lamivudine (LAM). Figure 36A shows HBsAg expression levels in HepG2.2.15 cells edited with the BE4 base editor in combination with the indicated guide, with and without pretreatment with LAM. Figure 36B shows HBeAg expression levels in HepG2.2.15 cells edited with and without LAM pretreatment using the BE4 base editor in combination with the indicated guides. Expression levels were compared using ANOVA / non-parametric test / no matched pairs / multiple comparisons with PCSK9. DETAILED DESCRIPTION OF THE INVENTION

[0153] The disclosed invention features compositions and methods for editing the HBV genome. For example, the compositions contemplated herein, in some embodiments, can include a base editor, a guide nucleic acid that targets a specific nucleotide within an HBV gene. In some embodiments, the editing introduces a premature stop codon in the coding sequence of one of the viral proteins. In other embodiments, the editing introduces one or more substitutions (e.g., missense mutations) in the coding sequence of one or more HBV proteins. In one embodiment, the editing modifies a nucleic acid base within a transcription element (e.g., a polyA site, an enhancer, or a promoter).

[0154] Recent studies have shown that newer generation ABEs (e.g., ABE8) can induce higher editing rates. Furthermore, they have lower gRNA-independent off-target editing rates compared to earlier generation cytidine base editors (CBEs), making them attractive as potential therapeutic approaches. ABEs do not allow for the generation of stop codons, but they do allow for targeting of AT-rich regions within the HBV genome.

[0155] The use of base editing to treat HBV advantageously prevents HBV rebound by safely introducing permanent mutations into cccDNA, irreversibly silencing HBsAg expression from integrated HBV DNA that does not contain DSBs.

[0156] As described in more detail below, the present disclosure provides a method for targeting established HBV covalently closed circular DNA (cccDNA) pools using a cytosine base editor (CBE, C to T conversion). Infected HepG2-NTCP cells and long-term primary human hepatocyte cultures (PHH) were cotransfected with selected HBV-targeting gRNAs and mRNA encoding the CBE. Base editing efficiency was assessed by DNA amplicon sequencing, and the consequences of base editing on viral replication were determined by analyzing different viral parameters. Without affecting the integrity of cccDNA, base editing introduced nonsense mutations into the HB or HBe / HBc open reading frame (ORF), efficiently inhibiting the release of total HBV DNA and HBs / HBe antigens. Furthermore, base editing rates remained high in the presence of pretreatment with nucleoside analogs (NAs), which reduced the levels of HBV DNA replication intermediates, indicating that CBEs can directly target cccDNA. Multiplexing two gRNAs with CBE resulted in greater reductions of HBsAg, HBeAg, total HBV DNA, and the 3.5 kb viral pregenomic RNA. Importantly, dual gRNA / CBE treatment prevented HBV rebound. Overall, these results demonstrate that CBE can directly target cccDNA and generate nonsense mutations in the HB and HBe / HBc ORFs that interfere with HBV replication. Furthermore, these effects persisted after CBE degradation, indicating permanent functional alterations of HBV cccDNA.

[0157] Targeting cccDNA transcriptional activity using base editing The regions involved in cccDNA transcriptional activity are AT-rich regions, and a surrogate to cccDNA degradation for HBV cure would be permanent disruption of cccDNA transcriptional activity, which is currently unattainable.

[0158] One example of a region involved in cccDNA transcriptional activity is the polyadenylation signal (PAS), which is unique to all HBV viral RNAs and also functions as a regulatory element (TATA box). Targeting this region in cccDNA can destabilize all viral RNA species, including its genome component pgRNA, resulting in cccDNA silencing.

[0159] As described in the Examples provided herein, sgRNAs were identified by in silico analysis targeting regulatory regions of cccDNA (e.g., enhancer I, enhancer II, basal core promoter, cryptic and canonical polyadenylation signals (PAS), S and X gene promoters). Over 250 sgRNAs were identified, these candidates were curated, and several candidates with high in silico prediction efficiency were tested in the Hek293-lentiHBV system.

[0160] As further described in the Examples, some embodiments of the present disclosure involve selecting highly conserved ABE gRNAs predicted to introduce missense mutations into HBV genes. For example, gRNAs were selected based on high conservation across HBV genotypes, focusing on genotypes D and A. High sequence conservation allows for targeting a broader patient population and is necessary to avoid the emergence of viral mutations that allow the HBV virus to escape base editing treatment. Another example of a selection criterion was the ability of the gRNA to generate edits in the HBV genome that would result in missense mutations. For amino acid changes generated by base editing reagents, gRNAs were selected that would generate missense amino acid changes with sequences rarely detected in naturally occurring HBV genotypes (amino acid changes with a frequency of <0.05%). In this way, the likelihood that the gRNA would introduce mutations that lead to disruption of the HBV protein / genome was increased.

[0161] HBV genome The HBV genome is approximately 3.2 kb of partial double-stranded DNA and open reading frames (ORFs) encoding seven proteins. Referring to Figure 1, open reading frame (ORF) P encodes the viral polymerase. ORF C / PreC encodes the capsid proteins. ORF PreS1, ORF PreS2, and ORF S encode the large (L), middle (M), and small (S) surface proteins, respectively. ORF X encodes the secreted X protein.

[0162] The partially double-stranded HBV genome is converted by host factors into covalently closed circular DNA (cccDNA). The cccDNA is transcribed by host RNA polymerase to produce viral mRNAs, including pregenomic RNA (pgRNA). The pgRNA is then transcribed by HBV polymerase into genomic HBV DNA, which can be converted to cccDNA, packaged into virions, or integrated into the host cell genome (Figure 2). cccDNA, a critical component of the HBV life cycle, is a stable molecule involved in chronic HBV infection. Editing the HBV genome can disrupt the formation of cccDNA, thereby reducing viral pathogenicity.

[0163] There are 10 distinct HBV genotypes (A-J) (Figure 3A). A "genotype" is characterized by less than 92% sequence identity with any other genome, while a subgenotype is characterized by less than 96% to 92% sequence identity. HBV genotype D is the most common in the United States (Figure 3A). Research models of HBV genotype D are available, including virus stocks (e.g., genotype D, subgenotype ayw (Imquest)) and mouse models (e.g., humanized mouse model (Phoenixbio)). Editing of target polynucleotides

[0164] The disclosed invention provides methods and compositions for targeting HBV ORFs for editing. These compositions can include a nucleic acid base editor having a Cas9 or other nucleic acid programmable DNA-binding protein domain and an adenosine or cytosine deaminase domain. In some embodiments, the base editor introduces one or more modifications into the HBV ORF. In some embodiments, the changes result in mutations in conserved portions of HBV proteins. In certain embodiments, the modifications introduce one or more stop codons. Throughout this specification, the introduction of a stop codon that results in premature termination of the protein is represented by the amino acid symbol, amino acid position, and the term stop (e.g., R87STOP indicates that the codon encoding arginine at amino acid position 87 is replaced by a stop codon). Advantageously, the methods of the present invention do not introduce double-strand breaks into the HBV genome. The present invention provides a strategy for treating chronic HBV using base editing (Figure 3B). Using the methods and compositions described herein, introduction of a stop codon into a viral gene can be achieved without generating a double-strand break, thereby eliminating or reducing the risk of HBV cleaving the host's genetic material after integration into the host's genome. Additionally, the compositions can utilize deaminases that are natural HBV antiviral restriction factors. For example, induction of APOBEC cytidine deaminase by interferon alpha or lymphotoxin β receptor (LTBR) promotes abasic site formation and cccDNA degradation (Figure 3B). Furthermore, base editors lacking a uracil glycosylase inhibitor domain can be used to target cellular uracil glycosylase to cccDNA and promote its degradation.

[0165] In embodiments, HBV ORFs are edited using a base editor containing adenosine deaminase (i.e., an adenosine base editor (ABE)). Compared to first-generation rAPOBEC1-based CBEs, ABEs have several advantages, including robust on-target editing and low gRNA-independent off-target editing. ABEs do not induce the uracil N-glycosylase (UNG) response. A>I cccDNA deamination has not been described for HBV. ABEs can be used to edit polymerase active sites and / or cccDNA transcription sites, including the RNA polyadenylation site TATAAA, which is common to HBV transcripts. Key regions involved in cccDNA transcription activity are AT-rich regions and surrogates for cccDNA degradation.

[0166] In some aspects, methods and compositions are provided for editing HBV cccDNA using base editors comprising a cytidine deaminase or an adenosine deaminase domain.

[0167] Exemplary guide RNAs are provided below in Table 1. Further exemplary guide RNAs include guide RNAs containing a spacer sequence provided in any of Tables 1, 2A, 2B, and / or 2C or listed in the Sequence Listing (e.g., in SEQ ID NOs: 3105-5485 and 8220-10830). In various embodiments, multiple guide RNAs, optionally in combination with one or more nucleobase editor polypeptides, can be administered to a subject or cell simultaneously. In embodiments, the multiple guide RNAs include 2, 3, 4, 5, 6, 7, 8, 9, 10, or more guides, optionally wherein the guide RNAs are selected from two or more of casl2b-4, casl2b-5, casl2b-11, casl2b-17, casl2b-30, casl2b-42, casl2b-100, casl2b-147, casl2b-154, casl2b-35, casl2b-124, casl2b-127, and casl2b-122b (see Table 1). In embodiments, the multiple guide RNAs administered to a subject or cell contain two guide RNAs (e.g., gRNA37 and gRNA40; see Table 1). A guide can optionally be administered simultaneously or sequentially with one or more nucleobase editor polypeptide(s) (e.g., base editor(s) encoded by an mRNA molecule(s)). Non-limiting examples of spacer sequences suitable for use in guide RNA molecules of the present disclosure include any spacer sequence that targets any portion of the Hepatitis B virus genome. Such spacer sequences can be designed and selected using methods available to those of skill in the art (e.g., design of gRNAs targeting specific sequences is available as a commercial service from Invitrogen). Various strategies can be used for gRNA selection, including, by way of non-limiting example, 1) targeting HBV transcriptional elements, and 2) selecting highly conserved ABE gRNAs predicted to introduce missense mutations into HBV genes.In embodiments, guide RNAs can be used to target base editors for the introduction of missense mutations into HBV genes (e.g., casl2b-4, casl2b-5, casl2b-11, casl2b-17, casl2b-30, casl2b-42, casl2b-100, casl2b-147, or casl2b-154). Guide RNAs listed in Table 1, or optional guide RNAs comprising spacer sequences listed in Table 2A, Table 2B, Table 2C, or listed within or anywhere in SEQ ID NOs: 3105-5485 and 8220-10830 of the Sequence Listing, can also be used to target base editors for modifying nucleobases at transcription sites (e.g., casl2b-35, casl2b-124, casl2b-127, or casl2b-122b). Transcription sites that can be targeted for base editing by guide RNAs in Table 1 or any guide RNA containing a spacer sequence listed in Tables 2A, 2B, and 2C, or anywhere in the Sequence Listing (e.g., SEQ ID NOS: 3105-5485 and 8220-10830), include enhancer II box A (e.g., casl2b-35), enhancer I (e.g., casl2b-124 or casl2b-127), and the HBX promoter (e.g., casl2b-122b). Non-limiting examples of transcription sites include the basal core promoter (BCP), enhancer I, enhancer II, HBV polyA site, HBX promoter, and S promoter-CCAATC-region II. Guide RNAs of the invention can be used in combination with any base editor (e.g., Cas12b-ABE (ABE-bhCas12b); SEQ ID NO: 418). Base editors can target the PAM sequences listed in Table 2A or Table 2B below (e.g., RTTN sequences).

[0168] In embodiments, two or more guide RNAs are introduced into a subject, optionally simultaneously. In embodiments, the two or more guide RNAs include casl2b-5 and casl2b-4; casl2b-11 and casl2b-4; casl2b-17 and casl2b-4; casl2b-30 and casl2b-4; casl2b-42 and casl2b-4; casl2b-100 and casl2b-4; casl2b-147 and casl2b-4; casl2b-154 and casl2b-4; casl2b-35 and casl2b-4; casl2b-124 and casl2b-4; casl2b-127 and cas12b-4;cas12b-122b and cas12b-4;cas12b-11 and cas12b-5;cas12b-17 and cas12b-5;cas12b-30 and cas12b-5;cas12b-42 and cas12b-5;cas12b-100 and cas12b-5;cas12b-147 and cas12b-5;cas12b-154 and cas12b-5;cas12b-35 and cas12b-5;cas12b-124 and cas12b-5;cas12b-127 and cas12b-5;ca s12b-122b and cas12b-5; cas12b-17 and cas12b-11; cas12b-30 and cas12b-11; cas12b-42 and cas12b-11; cas12b-100 and cas12b-11; cas12b-147 and cas12b-11; cas12b-154 and cas12b-11; cas12b-35 and cas12b-11; cas12b-124 and cas12b-11; cas12b-127 and cas12b-11; cas12b-122b and cas12b-11; c as12b-30 and cas12b-17;cas12b-42 and cas12b-17;cas12b-100 and cas12b-17;cas12b-147 and cas12b-17;cas12b-154 and cas12b-17;cas12b-35 and cas12b-17;cas12b-124 and cas12b-17;cas12b-127 and cas12b-17;cas12b-122b and cas12b-17;cas12b-42 and cas12b-30;cas12b-100 and cas12b-30;cas12b-147 and cas12b-30; cas12b-154 and cas12b-30; cas12b-35 and cas12b-30; cas12b-124 and cas12b-30; cas12b-127 and cas12b-30; cas12b-122b and cas12b-30; cas12b-100 and cas12b-42; cas12b-147 and cas12b-42; cas12b-154 and cas12b-42 ;cas12b-35 and cas12b-42;cas12b-124 and cas12b-42;cas12b-127 and cas12b-42;cas12b-122b and cas12b-42;cas12b-147 and cas12b-100;cas12b-154 and cas12b-100;cas12b-35 and cas12b-100;cas12b-124 and cas12b-100;cas12b-127 and cas12 b-100;cas12b-122b and cas12b-100;cas12b-154 and cas12b-147;cas12b-35 and cas12b-147;cas12b-124 and cas12b-147;cas12b-127 and cas12b-147;cas12b-122b and cas12b-147;cas12b-35 and cas12b-154;cas12b-124 and cas12b-154;cas12b- 127 and ca12b-154; ca12b-122b and ca12b-154; ca12b-124 and ca12b-35; ca12b-127 and ca12b-35; ca12b-122b and ca12b-35; ca12b-127 and ca12b-124; ca12b-122b and ca12b-124; or ca12b-122b and ca12b-127 (see Table 1).

[0169] [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5] [Table 1-6] [Table 1-7]

[0170] To generate the gene edits described above, cells (e.g., hepatocytes) are contacted with two or more guide RNAs and a nucleobase editor polypeptide (comprising a nucleic acid-programmable DNA-binding protein (napDNAbp) and a cytidine deaminase or an adenosine deaminase). In some embodiments, the cell to be edited is contacted with at least one nucleic acid, where the at least one nucleic acid encodes two or more guide RNAs and a nucleobase editor polypeptide (comprising a nucleic acid-programmable DNA-binding protein (napDNAbp) and a cytidine deaminase). In some embodiments, the gRNA comprises nucleotide analogs (e.g., mA*, mC*, mG*, and / or mU*). These nucleotide analogs can inhibit degradation of the gRNA by cellular processes. Tables 2A-2C provide exemplary target and spacer sequences for use in gRNAs. Further exemplary target and spacer sequences are set forth in the Sequence Listing as SEQ ID NOS: 724-3104 and 5609-8219, and SEQ ID NOS: 3105-5485 and 8220-10830, respectively. gRNAs can target all or a portion of all HBV genotypes (e.g., A, B, C, D, E, F, G, and / or H). gRNAs can target 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 85%, or all of the Hepatitis B viruses of a particular genotype or set of genotypes.

[0171] In various embodiments, a gRNA of the present disclosure (e.g., gRNA37-HM or gRNA40-HM) contains the following "heavily modified" (HM) scaffold: gUUUUAGagcuagaaauagcaaGUUaAaAuAaggcuaGUccGUUAucAAcuugaaaaagugGcaccgagucggugcusususu (SEQ ID NO: 317), where "a," "c," "u," or "g" represent A, C, U, or G nucleotides containing a 2'-O-methyl (M) modification, respectively, and as, cs, us, or gs represent A, C, U, or G nucleotides containing a 2'-O-methyl 3'-phosphorothioate (MS) modification, respectively. In embodiments, a guide RNA of the present disclosure contains an HM scaffold and a spacer sequence listed in any one of Tables 2A-2C.

[0172] [Table 2A-1] [Table 2A-2]

[0173] [Table 2B-1] [Table 2B-2] [Table 2B-3]

[0174] [Table 2C-1] [Table 2C-2] [Table 2C-3] [Table 2C-4] [Table 2C-5] [Table 2C-6] [Table 2C-7] [Table 2C-8]

[0175] The gRNA can target the HBV core protein, HBV polymerase gene, HBV surface protein, or HBV X protein. In embodiments, the PAM sequence associated with the sequence targeted by the gRNA is not limiting. Other targets include the polyadenylation signal (PAS), which is involved in cccDNA transcriptional activity and is inherent in all essential RNAs. The PAS also functions as a regulatory element (TATA box). Targeting the PAS in cccDNA may destabilize all viral RNA species, including pgRNA, resulting in cccDNA silencing. The sgRNA can target the main regulatory regions of cccDNA (e.g., enhancer I, enhancer II, basal core promoter, cryptic and canonical polyadenylation signals (PAS), and S and X gene promoters). The gRNA can be designed and evaluated in silico before being tested in cells (e.g., the HEK293-lentiHBV system).

[0176] In embodiments, any one of guides E1-E24 is used in combination with a nucleobase editor polypeptide, including SaCas9, and / or targets a sequence related to an NNRRT, NGG, or NGA PAM sequence. In embodiments, a guide RNA comprising any of the spacer sequences provided herein (e.g., any of Tables 2A-2C or those listed in the Sequence Listing), or a variant thereof, is used in combination with any of the base editors and / or nucleic acid-programmable DNA-binding proteins provided herein (e.g., BE4 or Cas12b).

[0177] In certain embodiments, the fusion proteins provided herein comprise one or more features that improve the base editing activity of the fusion protein. For example, any of the fusion proteins provided herein may contain a Cas9 domain with reduced nuclease activity. In some embodiments, any of the fusion proteins provided herein may have a Cas9 domain without nuclease activity (dCas9) or a Cas9 domain that cleaves one strand of a double-stranded DNA molecule (referred to as Cas9 nickase (nCas9)). Without wishing to be bound by any particular theory, the presence of a catalytic residue (e.g., H840) maintains the activity of Cas9 to cleave the unedited (e.g., unmethylated) strand opposite the target nucleobase. Mutation of the catalytic residue (e.g., D10 to A10) prevents cleavage of the edited strand containing the target A residue. Such Cas9 variants can generate single-stranded DNA breaks (nicks) at specific positions based on the target sequence defined by the gRNA, resulting in repair of the unedited strand and ultimately in a nucleobase change in the unedited strand.

[0178] In some embodiments, precise modifications of HBV genes reduce viral virulence and / or reduce the virus's ability to grow in vitro. The modifications can be premature stop codons or other mutations that impair or otherwise reduce the activity of HBV proteins. In some embodiments, the HBV genes targeted for modification are pol, X, S, pre-S1, pre-S2, the core of the pre-core gene, the transcription site, or a combination thereof.

[0179] Nucleic acid base editor Nucleobase editors that edit, modify, or alter a target nucleotide sequence of a polynucleotide are useful in the methods and compositions described herein. The nucleobase editors described herein generally comprise a polynucleotide-programmable nucleotide-binding domain and a nucleobase-editing domain (e.g., adenosine deaminase or cytidine deaminase). The polynucleotide-programmable nucleotide-binding domain, when combined with an associated guide polynucleotide (e.g., gRNA), can specifically bind to the target polynucleotide sequence, thereby localizing the base editor to the target nucleic acid sequence to be edited.

[0180] Polynucleotide Programmable Nucleotide Binding Domains The polynucleotide-programmable nucleotide binding domain binds to a polynucleotide (e.g., RNA, DNA). The polynucleotide-programmable nucleotide binding domain of a base editor may itself comprise one or more domains (e.g., one or more nuclease domains). In some embodiments, the nuclease domain of the polynucleotide-programmable nucleotide binding domain may comprise an endonuclease or an exonuclease. An endonuclease can cleave one strand of a double-stranded nucleic acid molecule or both strands of a double-stranded nucleic acid molecule. In some embodiments, the nuclease domain of the polynucleotide-programmable nucleotide binding domain can cleave zero, one, or two strands of a target polynucleotide.

[0181] Non-limiting examples of polynucleotide-programmable nucleotide-binding domains that can be incorporated into base editors include CRISPR protein-derived domains, restriction nucleases, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases (ZFNs). In some embodiments, the base editor comprises a polynucleotide-programmable nucleotide-binding domain comprising a natural or modified protein or portion thereof, and can bind to a nucleic acid sequence via an attached guide nucleic acid during CRISPR (i.e., clustered regularly interspaced short palindromic repeats)-mediated nucleic acid modification. Such proteins are referred to herein as "CRISPR proteins." Accordingly, disclosed herein are base editors comprising a polynucleotide-programmable nucleotide-binding domain comprising all or a portion of a CRISPR protein (i.e., a base editor comprising all or a portion of a CRISPR protein as a domain, also referred to as the "CRISPR protein-derived domain" of the base editor). The CRISPR protein-derived domain incorporated into the base editor can be modified compared to a wild-type or natural version of the CRISPR protein. For example, as described below, a domain derived from a CRISPR protein can include one or more mutations, insertions, deletions, rearrangements, and / or modifications relative to a wild-type or naturally occurring version of the CRISPR protein.

[0182] Cas proteins that may be used herein include class 1 and class 2. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Cs x17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf1, Cas12b / C2c1 (e.g., SEQ ID NO: 232), Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ, CARF, DinG, homologs thereof, or modified versions thereof. CRISPR enzymes can direct cleavage of one or both strands at a target sequence, e.g., within the target sequence and / or within a sequence complementary to the target sequence. For example, CRISPR enzymes can direct cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence.

[0183] A vector encoding a CRISPR enzyme mutated relative to the corresponding wild-type enzyme can be used, such that the mutant CRISPR enzyme lacks the ability to cleave one or both strands of a target polynucleotide containing a target sequence. A Cas protein (e.g., Cas9, Cas12) or Cas domain (e.g., Cas9, Cas12) can refer to a polypeptide or domain that has at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or homology to a wild-type exemplary Cas polypeptide or Cas domain. Cas (e.g., Cas9, Cas12) can refer to wild-type or modified forms of Cas proteins that can include amino acid changes such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof.

[0184] In some embodiments, the base editor CRISPR protein-derived domains are derived from Corynebacterium ulcerans (NCBI References: NC_015683.1, NC_017317.1), Corynebacterium diphtheria (NCBI References: NC_016782.1, NC_016786.1), Spiroplasma syrphidicola (NCBI Reference: NC_021284.1), Prevotella intermedia (NCBI Reference: NC_017861.1), Spiroplasma taiwanense (NCBI Reference: NC_021846.1), Streptococcus iniae (NCBI Reference: NC_021314.1), Belliella baltica (NCBI Reference: NC_018010.1), Psychroflexus The Cas9 may comprise all or a portion of Cas9 from S. torquis (NCBI Reference: NC_018721.1), Streptococcus thermophilus (NCBI Reference: YP_820832.1), Listeria innocua (NCBI Reference: NP_472073.1), Campylobacter jejuni (NCBI Reference: YP_002344900.1), Neisseria meningitidis (NCBI Reference: YP_002342100.1), Streptococcus pyogenes, or Staphylococcus aureus.

[0185] The sequence and structure of Cas9 nuclease are well known to those of skill in the art (see, e.g., "Complete genome sequence of an M1 strain of Streptococcus pyogenes," Ferretti et al., Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III," Deltcheva E., et al., Nature 471:602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity," Jinek M., et al., Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including, but not limited to, S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on the present disclosure, including Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737, the entire contents of which are incorporated herein by reference.

[0186] High-fidelity Cas9 domain Some aspects of the present disclosure provide high-fidelity Cas9 domains. High-fidelity Cas9 domains are known in the art and are described, for example, in Kleinstiver, B.P., et al., "High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects," Nature 529, 490-495 (2016) and Slaymaker, I.M., et al., "Rationally engineered Cas9 nucleases with improved specificity," Science 351, 84-88 (2015), the entire contents of each of which are incorporated herein by reference. An exemplary high-fidelity Cas9 domain is set forth in the Sequence Listing as SEQ ID NO: 233. In some embodiments, the high-fidelity Cas9 domain is a modified Cas9 domain containing one or more mutations relative to the corresponding wild-type Cas9 domain that reduce the electrostatic interaction between the Cas9 domain and the sugar-phosphate backbone of DNA. High-fidelity Cas9 domains with reduced electrostatic interactions with the sugar phosphate backbone of DNA have fewer off-target effects. In some embodiments, the Cas9 domain (e.g., a wild-type Cas9 domain (SEQ ID NOS: 197 and 200)) comprises one or more mutations that reduce the association between the Cas9 domain and the sugar phosphate backbone of DNA. In some embodiments, the Cas9 domain comprises one or more mutations that reduce the association between the Cas9 domain and the sugar phosphate backbone of DNA by at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, or at least 70%.

[0187] In some embodiments, any of the Cas9 fusion proteins provided herein includes one or more of the following mutations: D10A, N497X, R661X, Q695X, and / or Q926X, or corresponding mutations in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the high-fidelity Cas9 enzyme is SpCas9(K855A), eSpCas9(1.1), SpCas9-HF1, or a hyper-precise Cas9 variant (HypaCas9). In some embodiments, the modified Cas9, eSpCas9(1.1), contains an alanine substitution that weakens the interaction between the HNH / RuvC groove and non-target DNA strands, preventing strand separation and cleavage at off-target sites. Similarly, SpCas9-HF1 reduces off-target editing via an alanine substitution that disrupts the interaction of Cas9 with the DNA phosphate backbone. HypaCas9 contains mutations in the REC3 domain (SpCas9 N692A / M694A / Q695A / H698A) that enhance Cas9 proofreading and target discrimination. All three high-fidelity enzymes generate fewer off-target edits than wild-type Cas9.

[0188] Reduced exclusivity Cas9 domains Typically, Cas9 proteins, such as Cas9 from S. pyogenes (spCas9), require a "protospacer adjacent motif (PAM)" or PAM-like motif, a 2-6 base pair DNA sequence immediately following the DNA sequence targeted by the Cas9 nuclease in the CRISPR bacterial adaptive immune system. The presence of the NGG PAM sequence is required to bind to specific nucleic acid regions, where the "N" in "NGG" is adenosine (A), thymidine (T), or cytosine (C), and the "G" is guanosine. This can limit the ability to edit desired bases within the genome. In some embodiments, the base-editing fusion proteins provided herein may need to be placed at a precise location (e.g., a region containing the target base upstream of the PAM). See, e.g., Komor, A.C., et al., "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage," Nature 533, 420-424 (2016), the entire contents of which are incorporated herein by reference. Exemplary polypeptide sequences of spCas9 proteins capable of binding to PAM sequences are set forth in the Sequence Listing as SEQ ID NOS: 197, 201, and 234-237. Thus, in some embodiments, any of the fusion proteins provided herein can contain a Cas9 domain capable of binding to a nucleotide sequence that does not contain a canonical (e.g., NGG) PAM sequence. Cas9 domains that bind to non-canonical PAM sequences have been described in the art and would be apparent to one of skill in the art.For example, Cas9 domains that bind to non-canonical PAM sequences are described in Kleinstiver, B.P., et al., "Engineered CRISPR-Cas9 nucleases with altered PAM specificities," Nature 523, 481-485 (2015), and Kleinstiver, B.P., et al., "Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition," Nature Biotechnology 33, 1293-1298 (2015), the entire contents of each of which are incorporated herein by reference.

[0189] Nickase In some embodiments, a polynucleotide-programmable nucleotide-binding domain can comprise a nickase domain. As used herein, the term "nickase" refers to a polynucleotide-programmable nucleotide-binding domain, including a nuclease domain, that can cleave only one of the two strands of a double-stranded nucleic acid molecule (e.g., DNA). In some embodiments, a nickase can be derived from a fully catalytically active (e.g., native) form of a polynucleotide-programmable nucleotide-binding domain by introducing one or more mutations into the active polynucleotide-programmable nucleotide-binding domain. For example, if a polynucleotide-programmable nucleotide-binding domain comprises a nickase domain derived from Cas9, the Cas9-derived nickase domain can comprise a D10A mutation and a histidine at position 840. In such an embodiment, residue H840 retains catalytic activity, thereby enabling cleavage of one strand of a nucleic acid duplex. In another example, a Cas9-derived nickase domain can comprise an H840A mutation, but the amino acid residue at position 10 remains D. In some embodiments, a nickase can be derived from a fully catalytically active (e.g., native) form of a polynucleotide-programmable nucleotide-binding domain by removing all or part of a nuclease domain that is not required for nickase activity. For example, if the polynucleotide-programmable nucleotide-binding domain comprises a nickase domain derived from Cas9, the Cas9-derived nickase domain can comprise a deletion of all or part of the RuvC domain or the HNH domain.

[0190] In some embodiments, the wild-type Cas9 corresponds to or comprises the following amino acid sequence: [ka]

[0191] In some embodiments, the strand of a nucleic acid duplex target polynucleotide sequence that is cleaved by a base editor comprising a nickase domain (e.g., a Cas9-derived nickase domain, a Cas12-derived nickase domain) is the strand that is not edited by the base editor (i.e., the strand cleaved by the base editor is the opposite strand that contains the base to be edited). In other embodiments, a base editor comprising a nickase domain (e.g., a Cas9-derived nickase domain, a Cas12-derived nickase domain) can cleave the strand of a DNA molecule that is targeted for editing. In such embodiments, the non-target strand is not cleaved.

[0192] In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., (in the case of a "nickase" Cas9), the Cas9 is a nickase and is referred to as an "nCas9" protein. The Cas9 nickase can be a Cas9 protein that can cleave only one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 nickase cleaves the target strand of the double-stranded nucleic acid molecule, meaning that the Cas9 nickase cleaves the strand that is base-paired (complementary) to a gRNA (e.g., an sgRNA) bound to the Cas9. In some embodiments, the Cas9 nickase comprises a D10A mutation and has a histidine at position 840. In some embodiments, the Cas9 nickase cleaves the non-targeted, unbase-edited strand of the double-stranded nucleic acid molecule, meaning that the Cas9 nickase cleaves the strand that is not base-paired to a gRNA (e.g., an sgRNA) bound to the Cas9. In some embodiments, the Cas9 nickase comprises an H840A mutation and an aspartic acid residue at position 10, or a corresponding mutation. In some embodiments, the Cas9 nickase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the Cas9 nickases provided herein. Additional suitable Cas9 nickases will be apparent to those of skill in the art based on this disclosure and knowledge in the art, and are within the scope of this disclosure.

[0193] The amino acid sequence of an exemplary catalytic Cas9 nickase (nCas9) is as follows:

[0194] The Cas9 nuclease has two functional endonuclease domains (RuvC and HNH). Upon target binding, Cas9 undergoes a conformational change that positions the nuclease domains to cleave opposite strands of the target DNA. The end result of Cas9-mediated DNA cleavage is a double-strand break (DSB) within the target DNA (approximately 3–4 nucleotides upstream of the PAM sequence). The resulting DSB is then repaired by one of two general repair pathways: (1) the efficient but error-prone non-homologous end joining (NHEJ) pathway, or (2) the less efficient but high-fidelity homology-directed repair (HDR) pathway.

[0195] The "efficiency" of non-homologous end joining (NHEJ) and / or homology-directed repair (HDR) can be calculated by any convenient method. For example, in some embodiments, efficiency can be expressed as the percentage of successful HDR. For example, a Surveyor nuclease assay can be used to generate cleavage products, and the percentage can be calculated using the ratio of product to substrate. For example, a Surveyor nuclease enzyme can be used to directly cleave DNA containing the newly incorporated restriction sequence resulting from successful HDR. Cleavage of more substrate indicates a higher percentage of HDR (higher efficiency of HDR). As an illustrative example, the rate (percentage) of HDR can be calculated using the following formula [(cleavage product) / (substrate + cleavage product)] (e.g., (b + c) / (a + b + c), where "a" is the band intensity of the DNA substrate, and "b" and "c" are the band intensities of the cleavage products).

[0196] In some embodiments, efficiency may be expressed as the percentage of successful NHEJ. For example, a T7 endonuclease I assay can be used to generate cleavage products, and the ratio of products to substrates can be used to calculate the percentage of NHEJ. T7 endonuclease I cleaves mismatched heteroduplex DNA resulting from hybridization of wild-type and mutant DNA strands (NHEJ generates small random insertions or deletions (indels) at the original cleavage site). More cleavages indicate a higher percentage of NHEJ (more efficient NHEJ). As an illustrative example, the percentage of NHEJ can be calculated using the following formula: (1-(1-(b+c) / (a+b+c)) 1 / 2 ) × 100 (where "a" is the band intensity of the DNA substrate, and "b" and "c" are the cleavage products) (Ran et al., Cell. 2013 Sep. 12;154(6):1380-9 and Ran et al., Nat Protoc. 2013 Nov.;8(11):2281-2308).

[0197] The NHEJ repair pathway is the most active repair mechanism and frequently causes small nucleotide insertions or deletions (indels) at DSB sites. The random nature of NHEJ-mediated DSB repair has important practical implications, as cell populations expressing Cas9 and gRNA or guide polynucleotides can result in diverse mutations. In most embodiments, NHEJ generates small indels within the target DNA, resulting in amino acid deletions, insertions, or frameshift mutations that lead to premature stop codons within the open reading frame (ORF) of the target gene. The ideal end result is a loss-of-function mutation within the target gene.

[0198] While NHEJ-mediated DSB repair often disrupts the open reading frame of a gene, homology-directed repair (HDR) can be used to generate specific nucleotide changes ranging from single nucleotide changes to large insertions (e.g., addition of fluorophores or tags).

[0199] To utilize HDR for gene editing, a DNA repair template containing the desired sequence can be delivered to the cell type of interest using gRNA(s) and Cas9 or Cas9 nickase. The repair template can contain the desired edit as well as additional homologous sequences immediately upstream and downstream of the target (referred to as the left and right homology arms). The length of each homology arm can depend on the size of the alteration being introduced; larger insertions require longer homology arms. The repair template can be a single-stranded oligonucleotide, a double-stranded oligonucleotide, or a double-stranded DNA plasmid. Even in cells expressing Cas9, gRNA, and an exogenous repair template, HDR efficiency is generally low (less than 10% of modified alleles). Because HDR occurs during the S and G2 phases of the cell cycle, HDR efficiency can be enhanced by synchronizing cells. Chemical or genetic inhibition of genes involved in NHEJ can also increase HDR frequency.

[0200] In some embodiments, the Cas9 is a modified Cas9. A given gRNA targeting sequence may have additional sites throughout the genome where partial homology exists. These sites are called off-targets and must be considered when designing the gRNA. In addition to optimizing the gRNA design, CRISPR specificity can also be enhanced through modifications to Cas9. Cas9 generates double-strand breaks (DSBs) through the combined activity of two nuclease domains, RuvC and HNH. Cas9 nickase, a D10A mutant of SpCas9, retains one nuclease domain and generates DNA nicks instead of DSBs. The nickase system can also be combined with HDR-mediated gene editing for specific gene editing.

[0201] Catalytically inactive nucleases Also provided herein are base editors comprising a catalytically inactive (i.e., incapable of cleaving a target polynucleotide sequence) polynucleotide-programmable nucleotide-binding domain. As used herein, the terms "catalytically dead" and "nuclease dead" are used interchangeably to refer to a polynucleotide-programmable nucleotide-binding domain having one or more mutations and / or deletions that result in an inability to cleave a strand of nucleic acid. In some embodiments, a catalytically inactive polynucleotide-programmable nucleotide-binding domain base editor may lack nuclease activity as a result of specific point mutations in one or more nuclease domains. For example, in the case of a base editor comprising a Cas9 domain, Cas9 may contain both the D10A and H840A mutations. Such mutations inactivate both nuclease domains, thereby resulting in loss of nuclease activity. In other embodiments, a catalytically inactive polynucleotide-programmable nucleotide-binding domain may contain one or more deletions of all or part of a catalytic domain (e.g., the RuvC1 and / or HNH domain). In further embodiments, the catalytically inactive polynucleotide programmable nucleotide binding domain comprises a point mutation (e.g., D10A or H840A) and a deletion of all or part of the nuclease domain. dCas9 domains are known in the art and are described, for example, in Qi et al., "Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression." Cell. 2013;152(5):1173-83, the entire contents of which are incorporated herein by reference.

[0202] Additional suitable nuclease-inactive dCas9 domains will be apparent to those of skill in the art based on this disclosure and knowledge in the art and are within the scope of this disclosure. Such additional exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013;31(9):833-838, the entire contents of which are incorporated herein by reference).

[0203] In some embodiments, the dCas9 corresponds to, or partially or entirely comprises, a Cas9 amino acid sequence having one or more mutations that inactivate Cas9 nuclease activity. In some embodiments, the nuclease-inactive dCas9 domain comprises a D10X mutation and a H840X mutation in the amino acid sequence described herein, or a corresponding mutation in any amino acid sequence provided herein, where X is any amino acid change. In some embodiments, the nuclease-inactive dCas9 domain comprises a D10A mutation and a H840A mutation in the amino acid sequence described herein, or a corresponding mutation in any amino acid sequence provided herein. In some embodiments, the nuclease-inactive Cas9 domain comprises the amino acid sequence described in the cloning vector pPlatTET-gRNA2 (accession number BAV54124).

[0204] In some embodiments, the variant Cas9 protein is capable of cleaving the complementary strand of a guide target sequence but has a reduced ability to cleave the non-complementary strand of a double-stranded guide target sequence. For example, the variant Cas9 protein may have a mutation (amino acid substitution) that reduces the function of the RuvC domain. As a non-limiting example, in some embodiments, the variant Cas9 protein has D10A (aspartic acid to alanine at amino acid position 10), and thus is capable of cleaving the complementary strand of a double-stranded guide target sequence but has a reduced ability to cleave the non-complementary strand of a double-stranded guide target sequence (thus generating a single-strand break (SSB) instead of a double-strand break (DSB) when the variant Cas9 protein cleaves a double-stranded target nucleic acid) (see, e.g., Jinek et al., Science. 2012 Aug. 17; 337(6096): 816-21).

[0205] In some embodiments, the variant Cas9 protein can cleave the non-complementary strand of a double-stranded guide target sequence, but has a reduced ability to cleave the complementary strand of the guide target sequence. For example, the variant Cas9 protein can have a mutation (amino acid substitution) that reduces the function of the HNH domain (RuvC / HNH / RuvC domain motif). As a non-limiting example, in some embodiments, the variant Cas9 protein has H840A (histidine to alanine at amino acid position 840), and therefore can cleave the non-complementary strand of a guide target sequence, but has a reduced ability to cleave the complementary strand of a guide target sequence (thus, when the variant Cas9 protein cleaves a double-stranded guide target sequence, an SSB occurs instead of a DSB). Such a Cas9 protein has a reduced ability to cleave a guide target sequence (e.g., a single-stranded guide target sequence), but retains the ability to bind to a guide target sequence (e.g., a single-stranded guide target sequence).

[0206] As another non-limiting example, in some embodiments, a variant Cas9 protein comprises the W476A and W1126A mutations, thereby reducing the ability of the polypeptide to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA).

[0207] As another non-limiting example, in some embodiments, a variant Cas9 protein has P475A, W476A, N477A, ​​D1125A, W1126A, and D1127A mutations, thereby reducing the ability of the polypeptide to cleave target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA).

[0208] As another non-limiting example, in some embodiments, a variant Cas9 protein has H840A, W476A, and W1126A mutations, thereby reducing the ability of the polypeptide to cleave target DNA. Such a Cas9 protein has reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some embodiments, a variant Cas9 protein has H840A, D10A, W476A, and W1126A mutations, thereby reducing the ability of the polypeptide to cleave target DNA. Such a Cas9 protein has reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, a variant Cas9 has a catalytic His residue restored to position 840 of the Cas9 HNH domain (A840H).

[0209] As another non-limiting example, in some embodiments, a variant Cas9 protein has H840A, P475A, W476A, N477A, ​​D1125A, W1126A, and D1127A mutations, thereby reducing the ability of the polypeptide to cleave target DNA. Such Cas9 proteins have reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some embodiments, a variant Cas9 protein has D10A, H840A, P475A, W476A, N477A, ​​D1125A, W1126A, and D1127A mutations, thereby reducing the ability of the polypeptide to cleave target DNA. Such Cas9 proteins have reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, when a variant Cas9 protein has the W476A and W1126A mutations, or when a variant Cas9 protein has the P475A, W476A, N477A, ​​D1125A, W1126A, and D1127A mutations, the variant Cas9 protein does not efficiently bind to a PAM sequence. Thus, in some such embodiments, when such a variant Cas9 protein is used in a binding method, the method does not require a PAM sequence. In other words, in some embodiments, when such a variant Cas9 protein is used in a binding method, the method can include a guide RNA, but can be performed in the absence of a PAM sequence (thus, binding specificity is provided by the targeting segment of the guide RNA). Other residues can be mutated to achieve the above effects (i.e., to inactivate one or other nuclease moieties). As non-limiting examples, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 can be modified (i.e., substituted). Mutations other than alanine substitutions are also suitable.

[0210] In some embodiments, variant Cas9 proteins with reduced catalytic activity (e.g., when the Cas9 protein has a D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 mutation (e.g., D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A)) can still bind to target DNA in a site-specific manner (because it is still guided to the target DNA sequence by the guide RNA), so long as it retains the ability to interact with the guide RNA.

[0211] In some embodiments, the variant Cas protein can be spCas9, spCas9-VRQR, spCas9-VRER, xCas9(sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK, or spCas9-LRVSQL.

[0212] In some embodiments, the Cas9 domain is a Cas9 domain from Staphylococcus aureus (SaCas9). In some embodiments, the SaCas9 domain is a nuclease-active SaCas9, a nuclease-inactive SaCas9 (SaCas9d), or a SaCas9 nickase (SaCas9n). In some embodiments, the SaCas9 comprises an N579A mutation or a corresponding mutation in any of the amino acid sequences provided in the Sequence Listing submitted herewith.

[0213] In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to a nucleic acid sequence having a non-canonical PAM. In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to a nucleic acid sequence having an NNGRRT or NNGRRV PAM sequence. In some embodiments, the SaCas9 domain comprises one or more of the E781X, N967X, and R1014X mutations, or corresponding mutations in any of the amino acid sequences provided herein (wherein X is any amino acid). In some embodiments, the SaCas9 domain comprises one or more of the E781K, N967K, and R1014H mutations, or corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SaCas9 domain comprises the E781K, N967K, or R1014H mutations, or corresponding mutations in any of the amino acid sequences provided herein.

[0214] In some embodiments, one of the Cas9 domains present in the fusion protein can be replaced with a guide nucleotide sequence-programmable DNA-binding protein domain that does not require a PAM sequence. In some embodiments, the Cas9 is SaCas9. Residue A579 of SaCas9 can be mutated from N579 to obtain SaCas9 nickase. Residues K781, K967, and H1014 can be mutated from E781, N967, and R1014 to obtain SaKKH Cas9.

[0215] In some embodiments, a modified SpCas9 was used that contains the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (SpCas9-MQKFRAER) and has specificity for the engineered PAM 5'-NGC-3'.

[0216] An alternative to S. pyogenes Cas9 is an RNA-guided endonuclease from the Cpf1 family that exhibits cleavage activity in mammalian cells. CRISPR from Prevotella and Francisella 1 (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. Cpf1 is an RNA-guided endonuclease of the class II CRISPR / Cas system. This acquired immune mechanism is found in Prevotella and Francisella bacteria. The Cpf1 gene is associated with the CRISPR locus and encodes an endonuclease that uses guide RNA to find and cleave viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9, overcoming some of the limitations of the CRISPR / Cas9 system. Unlike Cas9 nuclease, Cpf1-mediated DNA cleavage results in double-strand breaks with short 3' overhangs. The staggered cleavage pattern of Cpf1 opens up the possibility of directional gene transfer similar to traditional restriction enzyme cloning, potentially increasing the efficiency of gene editing. Like the above-described Cas9 variants and orthologs, Cpf1 can also expand the number of sites that can be targeted by CRISPR, even to AT-rich regions or AT-rich genomes lacking NGG PAM sites suitable for SpCas9. The Cpf1 locus contains a mixed alpha / beta domain, RuvC-I, followed by a helical region, RuvC-II, and a zinc finger-like domain. The Cpf1 protein possesses a RuvC-like endonuclease domain similar to the RuvC domain of Cas9.

[0217] Furthermore, unlike Cas9, Cpf1 does not possess an HNH endonuclease domain, and the N-terminus of Cpf1 lacks the alpha-helical recognition lobe of Cas9. The Cpf1 CRISPR-Cas domain architecture indicates that Cpf1 is functionally unique and is classified as a Class 2, Type V CRISPR system. The Cpf1 locus encodes Cas1, Cas2, and Cas4 proteins, which are more similar to Type I and Type III systems than Type II systems. Functional Cpf1 does not require trans-activating CRISPR RNA (tracrRNA) and therefore only requires CRISPR (crRNA). This is beneficial for genome editing because Cpf1 is not only smaller than Cas9, but also has a smaller sgRNA molecule (roughly half the number of nucleotides as Cas9). The Cpf1-crRNA complex cleaves target DNA or RNA by identifying the protospacer adjacent motif 5'-YTN-3' or 5'-TTN-3', in contrast to the G-rich PAM targeted by Cas9. After PAM recognition, Cpf1 introduces sticky-end-like DNA double-strand breaks with 4- or 5-nucleotide overhangs.

[0218] In some embodiments, the Cas9 is a Cas9 variant with specificity for engineered PAM sequences. In some embodiments, additional Cas9 variants and PAM sequences are described in Miller, SM, et al. Continuous evolution of SpCas9 variants compatible with non-G PAMs, Nat. Biotechnol. (2020), the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 variant does not have a specific PAM requirement. In some embodiments, the Cas9 variant, e.g., an SpCas9 variant, has specificity for the NRNH PAM, where R is A or G and H is A, C, or T. In some embodiments, the SpCas9 variant has specificity for the PAM sequence AAA, TAA, CAA, GAA, TAT, GAT, or CAC. In some embodiments, the SpCas9 variant comprises an amino acid substitution at, or a corresponding position of, 1114, 1134, 1135, 1137, 1139, 1151, 1180, 1188, 1211, 1218, 1219, 1221, 1249, 1256, 1264, 1290, 1318, 1317, 1320, 1321, 1323, 1332, 1333, 1335, 1337, or 1339. In some embodiments, the SpCas9 variant comprises an amino acid substitution at, or a corresponding position of, 1114, 1135, 1218, 1219, 1221, 1249, 1320, 1321, 1323, 1332, 1333, 1335, or 1337. In some embodiments, the SpCas9 variant comprises an amino acid substitution at positions 1114, 1134, 1135, 1137, 1139, 1151, 1180, 1188, 1211, 1219, 1221, 1256, 1264, 1290, 1318, 1317, 1320, 1323, 1333, or their corresponding positions.In some embodiments, the SpCas9 variant comprises an amino acid substitution at positions 1114, 1131, 1135, 1150, 1156, 1180, 1191, 1218, 1219, 1221, 1227, 1249, 1253, 1286, 1293, 1320, 1321, 1332, 1335, 1339, or their corresponding positions. In some embodiments, the SpCas9 variant comprises an amino acid substitution at positions 1114, 1127, 1135, 1180, 1207, 1219, 1234, 1286, 1301, 1332, 1335, 1337, 1338, 1349, or their corresponding positions. Tables 3A-3D show exemplary amino acid substitutions and PAM specificities of SpCas9 variants.

[0219] [Table 3A]

[0220] [Table 3B-1] [Table 3B-2]

[0221] [Table 3C]

[0222] [Table 3D]

[0223] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) is a single effector of a microbial CRISPR-Cas system. Single effectors of microbial CRISPR-Cas systems include, but are not limited to, Cas9, Cpf1, Cas12b / C2c1, and Cas12c / C2c3. Microbial CRISPR-Cas systems are generally divided into class 1 and class 2 systems. Class 1 systems have a multi-subunit effector complex, while class 2 systems have a single protein effector. For example, Cas9 and Cpf1 are class 2 effectors. In addition to Cas9 and Cpf1, three distinct Class 2 CRISPR-Cas systems (Cas12b / C2c1 and Cas12c / C2c3) were described by Shmakov et al., "Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems," Mol. Cell, 2015 Nov. 5;60(3):385-397, the entire contents of which are incorporated herein by reference. Two of the systems' effectors, Cas12b / C2c1 and Cas12c / C2c3, contain a RuvC-like endonuclease domain related to Cpf1. The third system contains an effector with two predicted HEPN RNase domains. Unlike CRISPR RNA production by Cas12b / C2c1, mature CRISPR RNA production is independent of tracrRNA. Cas12b / C2c1 relies on both CRISPR RNA and tracrRNA for DNA cleavage.

[0224] In some embodiments, the napDNAbp is a circular permutation (eg, SEQ ID NO: 238).

[0225] The crystal structure of Alicyclobacillus acidoterrastris Cas12b / C2c1 (AacC2c1) complexed with a chimeric single-molecule guide RNA (sgRNA) has been reported. See, e.g., Liu et al., "C2c1-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism," Mol. Cell, 2017 Jan. 19;65(2):310-322, the entire contents of which are incorporated herein by reference. A crystal structure has also been reported for Alicyclobacillus acidoterrestris C2c1 bound to target DNA as a ternary complex. See, e.g., Yang et al., "PAM-dependent Target DNA Recognition and Cleavage by C2C1 CRISPR-Cas endonuclease," Cell, 2016 Dec. 15;167(7):1814-1828, the entire contents of which are incorporated herein by reference. A catalytically competent conformation of AacC2c1 has been captured in both the target and non-target DNA strands, independently positioned within a single RuvC catalytic pocket, resulting in a staggered seven-nucleotide cleavage of the target DNA upon Cas12b / C2c1-mediated cleavage. Structural comparison of the Cas12b / C2c1 ternary complex with previously identified Cas9 and Cpf1 counterparts demonstrates the diversity of the mechanisms employed by the CRISPR-Cas9 system.

[0226] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) of any of the fusion proteins provided herein can be a Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a Cas12b / C2c1 protein. In some embodiments, the napDNAbp is a Cas12c / C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the napDNAbp sequences provided herein. It is understood that Cas12b / C2c1 or Cas12c / C2c3 from other bacterial species may also be used in accordance with the present disclosure.

[0227] In some embodiments, napDNAbp refers to Cas12c. In some embodiments, the Cas12c protein is Cas12c1 (SEQ ID NO: 239) or a variant of Cas12c1. In some embodiments, the Cas12 protein is Cas12c2 (SEQ ID NO: 240) or a variant of Cas12c2. In some embodiments, the Cas12 protein is a Cas12c protein from Oleiphilus species HI0009 (i.e., OspCas12c, SEQ ID NO: 241) or a variant of OspCas12c. These Cas12c molecules are described in Yan et al., "Functionally Diverse Type V CRISPR-Cas Systems," Science, 2019 January 4;363:88-91, the entire contents of which are incorporated herein by reference. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring Cas12c1, Cas12c2, or OspCas12c protein. In some embodiments, the napDNAbp is a naturally occurring Cas12c1, Cas12c2, or OspCas12c protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any Cas12c1, Cas12c2, or OspCas12c protein described herein. It should be understood that Cas12c1, Cas12c2, or OspCas12c from other bacterial species may also be used in accordance with the present disclosure.

[0228] In some embodiments, napDNAbp refers to Cas12g, Cas12h, or Cas12i, e.g., as described in Yan et al., "Functionally Diverse Type V CRISPR-Cas Systems," Science, 2019 January 4;363:88-91, the entire contents of each of which are incorporated herein by reference. Exemplary Cas12g, Cas12h, and Cas12i polypeptide sequences are set forth in the Sequence Listing as SEQ ID NOs: 242-245. By aggregating over 10 terabytes of sequence data, new classes of type V Cas proteins, such as Cas12g, Cas12h, and Cas12i, have been identified that show weak similarity to previously characterized class V proteins. In some embodiments, the Cas12 protein is Cas12g or a variant of Cas12g. In some embodiments, the Cas12 protein is Cas12h or a variant of Cas12h. In some embodiments, the Cas12 protein is Cas12i or a variant of Cas12i. It is understood that other RNA-guided DNA-binding proteins may be used as napDNAbp and are within the scope of the present disclosure. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring Cas12g, Cas12h, or Cas12i protein. In some embodiments, the napDNAbp is a naturally occurring Cas12g, Cas12h, or Cas12i protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any Cas12g, Cas12h, or Cas12i protein described herein.It should be understood that Cas12g, Cas12h, or Cas12i from other bacterial species may also be used in accordance with the present disclosure. In some embodiments, the Cas12i is Cas12i1 or Cas12i2.

[0229] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) of any of the fusion proteins provided herein can be a Cas12j / CasΦ protein. Cas12j / CasΦ is described in Pausch et al., "CRISPR-CasΦ from huge phages is a hypercompact genome editor," Science, July 17, 2020, Vol. 369, Issue 6501, pp. 333-337, which is incorporated herein by reference in its entirety. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring Cas12j / CasΦ protein. In some embodiments, the napDNAbp is a naturally occurring Cas12j / CasΦ protein. In some embodiments, the napDNAbp is a nuclease-inactive ("dead") Cas12j / CasΦ protein. It is understood that Cas12j / CasΦ from other species may also be used in accordance with the present disclosure.

[0230] Fusion proteins with internal insertions Provided herein are fusion proteins comprising a heterologous polypeptide fused to a nucleic acid-programmable nucleic acid binding protein, e.g., napDNAbp. The heterologous polypeptide can be a polypeptide not found in the native or wild-type napDNAbp polypeptide sequence. The heterologous polypeptide can be fused to the napDNAbp at the C-terminus of the napDNAbp, the N-terminus of the napDNAbp, or inserted at an internal position of the napDNAbp. In some embodiments, the heterologous polypeptide is a deaminase (e.g., cytidine adenosine deaminase) or a functional fragment thereof. For example, the fusion protein can include a deaminase flanked by N-terminal and C-terminal fragments of a Cas9 or Cas12 (e.g., Cas12b / C2c1) polypeptide. In some embodiments, the cytidine deaminase is an APOBEC deaminase (e.g., APOBEC1). In some embodiments, the adenosine deaminase is TadA (e.g., TadA*7.10 or TadA*8). In some embodiments, TadA is TadA*8 or TadA*9. The TadA sequences described herein (e.g., TadA7.10 or TadA*8) are suitable deaminases for the above-described fusion proteins.

[0231] In some embodiments, the fusion protein has the following structure: NH2-[N-terminal fragment of napDNAbp]-[deaminase]-[C-terminal fragment of napDNAbp]-COOH, NH2-[N-terminal fragment of Cas9]-[adenosine deaminase]-[C-terminal fragment of Cas9]-COOH, NH2-[N-terminal fragment of Cas12]-[adenosine deaminase]-[C-terminal fragment of Cas12]-COOH, NH2-[N-terminal fragment of Cas9]-[cytidine deaminase]-[C-terminal fragment of Cas9]-COOH, NH2-[N-terminal fragment of Cas12]-[cytidine deaminase]-[C-terminal fragment of Cas12]-COOH, where each instance of "]-[" is an optional linker.

[0232] The deaminase can be a circular permutant deaminase. For example, the deaminase can be a circular permutant adenosine deaminase. In some embodiments, the deaminase is a circularly permuted TadA that is circularly permuted at amino acid residues 116, 136, or 65 numbered in the TadA reference sequence.

[0233] The fusion protein may contain multiple deaminases. The fusion protein may contain, for example, 1, 2, 3, 4, 5, or more deaminases. In some embodiments, the fusion protein contains one or two deaminases. The two or more deaminases of the fusion protein may be adenosine deaminase, cytidine deaminase, or a combination thereof. The two or more deaminases may be homodimers or heterodimers. The two or more deaminases may be inserted in tandem in the napDNAbp. In some embodiments, the two or more deaminases may not be in tandem in the napDNAbp.

[0234] In some embodiments, the napDNAbp in the fusion protein is a Cas9 polypeptide or a fragment thereof. The Cas9 polypeptide can be a variant Cas9 polypeptide. In some embodiments, the Cas9 polypeptide is a Cas9 nickase (nCas9) polypeptide or a fragment thereof. In some embodiments, the Cas9 polypeptide is a nuclease-inactive Cas9 (dCas9) polypeptide or a fragment thereof. The Cas9 polypeptide in the fusion protein can be a full-length Cas9 polypeptide. In some cases, the Cas9 polypeptide in the fusion protein may not be a full-length Cas9 polypeptide. The Cas9 polypeptide can be truncated, for example, at the N-terminus or C-terminus compared to a naturally occurring Cas9 protein. The Cas9 polypeptide can be a circularly permuted Cas9 protein. The Cas9 polypeptide can be a fragment, portion, or domain of a Cas9 polypeptide, but it can still bind to a target polynucleotide and a guide nucleic acid sequence.

[0235] In some embodiments, the Cas9 polypeptide is Streptococcus pyogenes Cas9 (SpCas9), Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), or a fragment or variant of any of the Cas9 polypeptides described herein.

[0236] In some embodiments, the fusion protein comprises an adenosine deaminase domain and a cytidine deaminase domain inserted into Cas9. In some embodiments, the adenosine deaminase is fused within Cas9 and the cytidine deaminase is fused to the C-terminus. In some embodiments, the adenosine deaminase is fused within Cas9 and the cytidine deaminase is fused to the N-terminus. In some embodiments, the cytidine deaminase is fused within Cas9 and the adenosine deaminase is fused to the C-terminus. In some embodiments, the cytidine deaminase is fused within Cas9 and the adenosine deaminase is fused to the N-terminus.

[0237] Exemplary structures of fusion proteins with adenosine deaminase and cytidine deaminase and Cas9 are provided below: NH2-[Cas9(adenosine deaminase)]-[cytidine deaminase]-COOH, NH2-[cytidine deaminase]-[Cas9(adenosine deaminase)]-COOH, NH2-[Cas9(cytidine deaminase)]-[adenosine deaminase]-COOH, or NH2-[adenosine deaminase]-[Cas9 (cytidine deaminase)]-COOH.

[0238] In some embodiments, the "-" used in the general architecture above indicates the presence of an optional linker.

[0239] In various embodiments, the catalytic domain has a DNA-modifying activity (e.g., deaminase activity), such as adenosine deaminase activity. In some embodiments, the adenosine deaminase is TadA (e.g., TadA*7.10). In some embodiments, TadA is TadA*8. In some embodiments, TadA*8 is fused within Cas9 and a cytidine deaminase is fused to the C-terminus. In some embodiments, TadA*8 is fused within Cas9 and a cytidine deaminase is fused to the N-terminus. In some embodiments, a cytidine deaminase is fused within Cas9 and TadA*8 is fused to the C-terminus. In some embodiments, a cytidine deaminase is fused within Cas9 and TadA*8 is fused to the N-terminus. Exemplary structures of fusion proteins having TadA*8 and cytidine deaminase and Cas9 are provided below: NH2-[Cas9(TadA*8)]-[cytidine deaminase]-COOH, NH2-[cytidine deaminase]-[Cas9(TadA*8)]-COOH, NH2-[Cas9(cytidine deaminase)]-[TadA*8]-COOH or NH2-[TadA*8]-[Cas9(cytidine deaminase)]-COOH.

[0240] In some embodiments, the "-" used in the general architecture above indicates the presence of an optional linker.

[0241] A heterologous polypeptide (e.g., a deaminase) can be inserted into the napDNAbp (e.g., Cas9 or Cas12 (e.g., Cas12b / C2c1)) at a suitable position, for example, so that the napDNAbp retains the ability to bind to a target polynucleotide and a guide nucleic acid. The deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted into the napDNAbp without impairing the function of the deaminase (e.g., base editing activity) or the napDNAbp (e.g., the ability to bind to a target nucleic acid and a guide nucleic acid). The deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted into the napDNAbp, for example, in a disordered region or a region containing a high temperature factor or B factor, as shown by crystallographic studies. Ordered, disordered, or unstructured regions of proteins (e.g., solvent-exposed regions and loops) can be used for insertion without compromising structure or function. Deaminases (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted into napDNAbp in flexible loop regions or solvent-exposed regions. In some embodiments, deaminases (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) are inserted into flexible loops of Cas9 or Cas12b / C2c1 polypeptides.

[0242] In some embodiments, the insertion location of the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is determined by B-factor analysis of the crystal structure of the Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted into a region of the Cas9 polypeptide containing a higher than average B-factor (e.g., a higher B-factor compared to the whole protein or a protein domain containing disordered regions). The B-factor or temperature factor may indicate the variation of atoms from their average positions (e.g., as a result of temperature-dependent atomic vibrations or static disorder in the crystal lattice). A high B-factor (e.g., a higher-than-average B-factor) of backbone atoms may indicate a region with relatively high local mobility. Such a region can be used to insert the deaminase without compromising structure or function. The deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted at a position containing a residue with a Cα atom that has a B factor that is 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, or more than 200% of the average B factor of the entire protein. A deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted at a position containing a residue having a B-factor that is 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, or greater than the average B-factor of the Cas9 protein domain containing that residue. Positions in the Cas9 polypeptide containing higher than average B-factors can include, for example, residues 768, 792, 1052, 1015, 1022, 1026, 1029, 1067, 1040, 1054, 1068, 1246, 1247, and 1248, as numbered in the Cas9 reference sequence above.Regions of a Cas9 polypeptide containing a higher than average B factor can include, for example, residues 792-872, 792-906, and 2-791, as numbered in the above Cas9 reference sequence.

[0243] A heterologous polypeptide (e.g., a deaminase) can be inserted into the napDNAbp at an amino acid residue selected from the group consisting of residues 768, 791, 792, 1015, 1016, 1022, 1023, 1026, 1029, 1040, 1052, 1054, 1067, 1068, 1069, 1246, 1247, and 1248, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues of another Cas9 polypeptide. In some embodiments, the heterologous polypeptide is inserted between numbered amino acid positions 768-769, 791-792, 792-793, 1015-1016, 1022-1023, 1026-1027, 1029-1030, 1040-1041, 1052-1053, 1054-1055, 1067-1068, 1068-1069, 1247-1248, or 1248-1249 in the above Cas9 reference sequence, or their corresponding amino acid positions. In some embodiments, the heterologous polypeptide is inserted between amino acid positions 769-770, 792-793, 793-794, 1016-1017, 1023-1024, 1027-1028, 1030-1031, 1041-1042, 1053-1054, 1055-1056, 1068-1069, 1069-1070, 1248-1249, or 1249-1250, as numbered in the above Cas9 reference sequence or its corresponding amino acid positions. In some embodiments, the heterologous polypeptide replaces an amino acid residue selected from the group consisting of residues 768, 791, 792, 1015, 1016, 1022, 1023, 1026, 1029, 1040, 1052, 1054, 1067, 1068, 1069, 1246, 1247, and 1248 as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. It should be understood that reference to the above Cas9 reference sequence with respect to insertion positions is for exemplary purposes.Insertions discussed herein are not limited to the Cas9 polypeptide sequences of the above Cas9 reference sequences, but include insertions at corresponding locations in variant Cas9 polypeptides (e.g., Cas9 nickase (nCas9), nuclease-inactive Cas9 (dCas9), Cas9 variants lacking a nuclease domain, truncated Cas9, or a Cas9 domain partially or completely lacking the HNH domain).

[0244] A heterologous polypeptide (e.g., a deaminase) may be inserted into the napDNAbp at an amino acid residue selected from the group consisting of amino acid residues 768, 792, 1022, 1026, 1040, 1068, and 1247, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the heterologous polypeptide is inserted between amino acid positions 768-769, 792-793, 1022-1023, 1026-1027, 1029-1030, 1040-1041, 1068-1069, or 1247-1248, as numbered in the above Cas9 reference sequence, or their corresponding amino acid positions. In some embodiments, the heterologous polypeptide is inserted between numbered amino acid positions 769-770, 793-794, 1023-1024, 1027-1028, 1030-1031, 1041-1042, 1069-1070, or 1248-1249 in the above Cas9 reference sequence, or their corresponding amino acid positions. In some embodiments, the heterologous polypeptide replaces an amino acid residue selected from the group consisting of amino acid residues 768, 792, 1022, 1026, 1040, 1068, and 1247 in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide.

[0245] A heterologous polypeptide (e.g., a deaminase) can be inserted into the napDNAbp at an amino acid residue described herein or a corresponding amino acid residue in another Cas9 polypeptide. In one embodiment, a heterologous polypeptide (e.g., a deaminase) can be inserted into the napDNAbp at an amino acid residue selected from the group consisting of amino acid residues 1002, 1003, 1025, 1052-1056, 1242-1247, 1061-1077, 943-947, 686-691, 569-578, 530-539, and 1060-1077, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide. The deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted at the N-terminus or C-terminus of the residue or replace the residue. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at the C-terminus of the residue.

[0246] In some embodiments, an adenosine deaminase (e.g., TadA) may be inserted at an amino acid residue selected from the group consisting of: amino acid residues 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246 numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, an adenosine deaminase (e.g., TadA) is inserted in place of residues 792-872, 792-906, or 2-791 numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the adenosine deaminase is inserted N-terminally at an amino acid selected from the group consisting of amino acid residues 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246 numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the adenosine deaminase is inserted C-terminally at an amino acid selected from the group consisting of amino acid residues 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246 numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the adenosine deaminase is inserted to replace an amino acid selected from the group consisting of amino acid residues 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide.

[0247] In some embodiments, the cytidine deaminase (e.g., APOBEC1) is inserted at an amino acid residue selected from the group consisting of amino acid residues 1016, 1023, 1029, 1040, 1069, and 1247 numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the cytidine deaminase is inserted N-terminally at an amino acid selected from the group consisting of amino acid residues 1016, 1023, 1029, 1040, 1069, and 1247 numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the cytidine deaminase is inserted C-terminally at an amino acid selected from the group consisting of amino acid residues 1016, 1023, 1029, 1040, 1069, and 1247 numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the cytidine deaminase is inserted to replace an amino acid selected from the group consisting of amino acid residues 1016, 1023, 1029, 1040, 1069, and 1247, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide.

[0248] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 768 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminally to amino acid residue 768 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminally to amino acid residue 768 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted to replace amino acid residue 768, numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.

[0249] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 791 or amino acid residue 792, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminally to amino acid residue 791 or amino acid residue 792, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminally at amino acid 791 or N-terminally at amino acid 792, as numbered in the above Cas9 reference sequence, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted to replace amino acid 791 or amino acid 792, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.

[0250] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1016 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminally to amino acid residue 1016 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminally to amino acid residue 1016 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted to replace amino acid residue 1016, numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.

[0251] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1022 or amino acid residue 1023, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminal to amino acid residue 1022 or amino acid residue 1023, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminal to amino acid residue 1022 or amino acid residue 1023, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted to replace amino acid residue 1022 or amino acid residue 1023, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.

[0252] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1026 or amino acid residue 1029, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminal to amino acid residue 1026 or amino acid residue 1029, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminal to amino acid residue 1026 or amino acid residue 1029, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted to replace amino acid residue 1026 or amino acid residue 1029, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.

[0253] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1040 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminally to amino acid residue 1040 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminally to amino acid residue 1040 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted to replace amino acid residue 1040, numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.

[0254] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1052 or amino acid residue 1054, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminally to amino acid residue 1052 or amino acid residue 1054, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminal to amino acid residue 1052 or amino acid residue 1054, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted to replace amino acid residue 1052 or amino acid residue 1054, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.

[0255] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1067, or amino acid residue 1068, or amino acid residue 1069, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminal to amino acid residue 1067, or amino acid residue 1068, or amino acid residue 1069, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminal to amino acid residue 1067, or amino acid residue 1068, or amino acid residue 1069, as numbered in the above Cas9 reference sequence, or to the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted to replace amino acid residue 1067, or amino acid residue 1068, or amino acid residue 1069, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide.

[0256] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1246, or amino acid residue 1247, or amino acid residue 1248, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminal to amino acid residue 1246, or amino acid residue 1247, or amino acid residue 1248, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminal to amino acid residue 1246, or amino acid residue 1247, or amino acid residue 1248, as numbered in the above Cas9 reference sequence, or to the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted to replace amino acid residue 1246, or amino acid residue 1247, or amino acid residue 1248, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide.

[0257] In some embodiments, the heterologous polypeptide (e.g., a deaminase) is inserted into a flexible loop of a Cas9 polypeptide. The flexible loop portion can be selected from the group consisting of amino acid residues 530-537, 569-570, 686-691, 943-947, 1002-1025, 1052-1077, 1232-1247, or 1298-1300, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. The flexible loop portion can be selected from the group consisting of amino acid residues 1-529, 538-568, 580-685, 692-942, 948-1001, 1026-1051, 1078-1231, or 1248-1297, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide.

[0258] A heterologous polypeptide (e.g., adenine deaminase) can be inserted into a region of a Cas9 polypeptide corresponding to the following amino acid residues: 1017-1069, 1242-1247, 1052-1056, 1060-1077, 1002-1003, 943-947, 530-537, 568-579, 686-691, 1242-1247, 1298-1300, 1066-1077, 1052-1056, or 1060-1077, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide.

[0259] A heterologous polypeptide (e.g., adenine deaminase) can be inserted in place of the deleted region of a Cas9 polypeptide. The deleted region can correspond to the N-terminal or C-terminal portion of the Cas9 polypeptide. In some embodiments, the deleted region corresponds to residues 792-872 as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the deleted region corresponds to residues 792-906 as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the deleted region corresponds to residues 2-791 as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the deleted region corresponds to residues 1017-1069 as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues.

[0260] Exemplary internal fusion base editors are provided in Table 4 below. [Table 4]

[0261] A heterologous polypeptide (e.g., a deaminase) can be inserted within a structural or functional domain of a Cas9 polypeptide. A heterologous polypeptide (e.g., a deaminase) can be inserted between two structural or functional domains of a Cas9 polypeptide. A heterologous polypeptide (e.g., a deaminase) can be inserted in place of a structural or functional domain of a Cas9 polypeptide, for example, after deleting the domain from the Cas9 polypeptide. A structural or functional domain of a Cas9 polypeptide can include, for example, RuvCI, RuvCII, RuvCIII, Rec1, Rec2, PI, or HNH.

[0262] In some embodiments, the Cas9 polypeptide lacks one or more domains selected from the group consisting of: RuvCI, RuvCII, RuvCIII, Rec1, Rec2, PI, or HNH domain. In some embodiments, the Cas9 polypeptide lacks a nuclease domain. In some embodiments, the Cas9 polypeptide lacks an HNH domain. In some embodiments, the Cas9 polypeptide lacks a portion of the HNH domain, such that the Cas9 polypeptide has reduced or eliminated HNH activity. In some embodiments, the Cas9 polypeptide comprises a deletion of the nuclease domain, and a deaminase is inserted to replace the nuclease domain. In some embodiments, the HNH domain is deleted and a deaminase is inserted in its place. In some embodiments, one or more of the RuvC domains are deleted and a deaminase is inserted in its place.

[0263] A fusion protein comprising a heterologous polypeptide can be flanked by N- and C-terminal fragments of a napDNAbp. In some embodiments, the fusion protein comprises a deaminase flanked by N- and C-terminal fragments of a Cas9 polypeptide. The N- or C-terminal fragment can bind to a target polynucleotide sequence. The C-terminus of the N- or C-terminal fragment can comprise a portion of a flexible loop of a Cas9 polypeptide. The C-terminus of the N- or C-terminal fragment can comprise a portion of an alpha-helical structure of a Cas9 polypeptide. The N- or C-terminal fragment can comprise a DNA-binding domain. The N- or C-terminal fragment can comprise a RuvC domain. The N- or C-terminal fragment can comprise an HNH domain. In some embodiments, either the N- or C-terminal fragment does not comprise an HNH domain.

[0264] In some embodiments, the C-terminus of the N-terminal Cas9 fragment comprises amino acids that are proximal to the target nucleobase when the fusion protein deaminates the target nucleobase. In some embodiments, the N-terminus of the C-terminal Cas9 fragment comprises amino acids that are proximal to the target nucleobase when the fusion protein deaminates the target nucleobase. The insertion locations of different deaminases can be different to allow for proximity between the target nucleobase and the C-terminus of the N-terminal Cas9 fragment or the amino acids at the N-terminus of the C-terminal Cas9 fragment. For example, the insertion locations of deaminases can be at amino acid residues selected from the group consisting of amino acid residues 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246 numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide.

[0265] The N-terminal Cas9 fragment of the fusion protein (i.e., the N-terminal Cas9 fragment adjacent to the deaminase of the fusion protein) can comprise the N-terminus of the Cas9 polypeptide. The N-terminal Cas9 fragment of the fusion protein can comprise a length of at least about: 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, or 1300 amino acids. The N-terminal Cas9 fragment of the fusion protein can comprise a sequence corresponding to the following amino acid residues: 1-56, 1-95, 1-200, 1-300, 1-400, 1-500, 1-600, 1-700, 1-718, 1-765, 1-780, 1-906, 1-918, or 1-1100, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues of another Cas9 polypeptide. An N-terminal Cas9 fragment can comprise amino acid residues 1-56, 1-95, 1-200, 1-300, 1-400, 1-500, 1-600, 1-700, 1-718, 1-765, 1-780, 1-906, 1-918, or 1-1100, as numbered in the above Cas9 reference sequences, or a sequence comprising at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to the corresponding amino acid residues of another Cas9 polypeptide.

[0266] The C-terminal Cas9 fragment of the fusion protein (i.e., the C-terminal Cas9 fragment adjacent to the deaminase of the fusion protein) can comprise the C-terminus of the Cas9 polypeptide. The C-terminal Cas9 fragment of the fusion protein can comprise a length of at least about: 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, or 1300 amino acids. The C-terminal Cas9 fragment of the fusion protein can comprise a sequence corresponding to the following amino acid residues: 1099-1368, 918-1368, 906-1368, 780-1368, 765-1368, 718-1368, 94-1368, or 56-1368, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide. An N-terminal Cas9 fragment can comprise amino acid residues 1099-1368, 918-1368, 906-1368, 780-1368, 765-1368, 718-1368, 94-1368, or 56-1368, as numbered in the above Cas9 reference sequences, or a sequence comprising at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to the corresponding amino acid residues in another Cas9 polypeptide.

[0267] The N-terminal Cas9 fragment and the C-terminal Cas9 fragment of the fusion protein taken together may not correspond to a native full-length Cas9 polypeptide sequence, e.g., as set forth in the Cas9 reference sequence above.

[0268] The fusion proteins described herein can provide targeted deamination with reduced deamination at non-target sites (e.g., off-target sites) (e.g., reduced genome-wide spurious deamination). The fusion proteins described herein can provide targeted deamination with reduced bystander deamination at non-target sites. Unwanted or off-target deamination can be reduced by at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%, for example, compared to end terminus fusion proteins comprising a deaminase fused to the N- or C-terminus of a Cas9 polypeptide. Unwanted or off-target deamination can be reduced by, for example, at least 1-fold, at least 2-fold, at least 3-fold, at least 4-fold, at least 5-fold, at least 10-fold, at least 15-fold, at least 20-fold, at least 30-fold, at least 40-fold, at least 50-fold, at least 60-fold, at least 70-fold, at least 80-fold, at least 90-fold, or at least 100-fold compared to a terminal fusion protein comprising a deaminase fused to the N-terminus or C-terminus of a Cas9 polypeptide.

[0269] In some embodiments, the deaminase of the fusion protein (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) deaminates no more than two nucleobases within the R-loop. In some embodiments, the deaminase of the fusion protein deaminates no more than three nucleobases within the R-loop. In some embodiments, the deaminase of the fusion protein deaminates no more than 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleobases within the R-loop. An R-loop is a triple-stranded nucleic acid structure, comprising a DNA:RNA hybrid, a DNA:DNA, or an RNA:RNA complementary structure, associated with single-stranded DNA. As used herein, an R-loop can be formed when a target polynucleotide contacts a CRISPR complex or a base editing complex, where a portion of the guide polynucleotide (e.g., guide RNA) hybridizes to and displaces a portion of the target polynucleotide (e.g., target DNA). In some embodiments, the R-loop comprises a hybridized region of a spacer sequence and a target DNA-complementary sequence. The R-loop region can be a region of about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleobase pairs in length. In some embodiments, the R-loop region is about 20 nucleobase pairs in length. It should be understood that, as used herein, the R-loop region is not limited to the target DNA strand that hybridizes with the guide polynucleotide. For example, editing of target nucleobases in an R-loop region may be on the DNA strand that contains the complementary strand to the guide RNA, or on the DNA strand that is the opposite strand to the complementary strand to the guide RNA. In some embodiments, editing in the R-loop region involves editing nucleobases on the non-complementary strand to the guide RNA (protospacer strand) in the target DNA sequence.

[0270] The fusion proteins described herein can effect targeted deamination in an editing window distinct from canonical base editing. In some embodiments, the target nucleobase is located about 1 to about 20 bases upstream of the PAM sequence in the target polynucleotide sequence. In some embodiments, the target nucleobase is located about 2 to about 12 bases upstream of the PAM sequence in the target polynucleotide sequence. In some embodiments, the target nucleobase is located about 1 to 9 base pairs, about 2 to 10 base pairs, about 3 to 11 base pairs, about 4 to 12 base pairs, about 5 to 13 base pairs, about 6 to 14 base pairs, about 7 to 15 base pairs, about 8 to 16 base pairs, about 9 to 17 base pairs, about 10 to 18 base pairs, about 11 to 19 base pairs, about 12 to 20 base pairs, about 1 to 7 base pairs, about 2 to 8 base pairs, about 3 to 9 ... base pairs, about 4 to 10 base pairs, about 5 to 11 base pairs, about 6 to 12 base pairs, about 7 to 13 base pairs, about 8 to 14 base pairs, about 9 to 15 base pairs, about 10 to 16 base pairs, about 11 to 17 base pairs, about 12 to 18 base pairs, about 13 to 19 base pairs, about 14 to 20 base pairs, about 1 to 5 base pairs, about 2 to 6 base pairs, about 3 to 7 base pairs, about 4 to 8 base pairs, about 5 to 9 base pairs, about 6-10 base pairs, about 7-11 base pairs, about 8-12 base pairs, about 9-13 base pairs, about 10-14 base pairs, about 11-15 base pairs, about 12-16 base pairs, about 13-17 base pairs, about 14-18 base pairs, about 15-19 base pairs, about 16-20 base pairs, about 1-3 base pairs, about 2-4 base pairs, about 3-5 base pairs, about 4-6 base pairs, about 5-7 base pairs, about 6-8 The target nucleobase is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more base pairs away from the PAM sequence or upstream of the PAM sequence. In some embodiments, the target nucleobase is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more base pairs away from the PAM sequence or upstream of the PAM sequence. In some embodiments, the target nucleobase is about 1, 2, 3, 4, 5, 6, 7, 8, or 9 base pairs upstream of the PAM sequence. In some embodiments, the target nucleobase is about 2, 3, 4, or 6 base pairs upstream of the PAM sequence.

[0271] A fusion protein can contain multiple heterologous polypeptides. For example, a fusion protein can further contain one or more UGI domains and / or one or more nuclear localization signals. Two or more heterologous domains can be inserted in tandem. Two or more heterologous domains can be inserted in positions in the NapDNAbp where they are not tandem.

[0272] The fusion protein can include a linker between the deaminase and the napDNAbp polypeptide. The linker can be a peptide or a non-peptide linker. For example, the linker can be XTEN, (GGGS)n (SEQ ID NO:246), (GGGGS)n (SEQ ID NO:247), (G)n, (EAAAK)n (SEQ ID NO:248), (GGS)n, or SGSETPGTSESATPES (SEQ ID NO:249). In some embodiments, the fusion protein includes a linker between the N-terminal Cas9 fragment and the deaminase. In some embodiments, the fusion protein includes a linker between the C-terminal Cas9 fragment and the deaminase. In some embodiments, the N-terminal and C-terminal fragments of the napDNAbp are connected to the deaminase by a linker. In some embodiments, the N-terminal and C-terminal fragments are linked to the deaminase domain without a linker. In some embodiments, the fusion protein includes a linker between the N-terminal Cas9 fragment and the deaminase, but does not include a linker between the C-terminal Cas9 fragment and the deaminase. In some embodiments, the fusion protein comprises a linker between the C-terminal Cas9 fragment and the deaminase, but does not comprise a linker between the N-terminal Cas9 fragment and the deaminase.

[0273] In some embodiments, the napDNAbp in the fusion protein is a Cas12 polypeptide (e.g., Cas12b / C2c1) or a fragment thereof. The Cas12 polypeptide can be a variant Cas12 polypeptide. In other embodiments, the N-terminal or C-terminal fragment of the Cas12 polypeptide comprises a nucleic acid programmable DNA-binding domain or a RuvC domain. In other embodiments, the fusion protein contains a linker between the Cas12 polypeptide and the catalytic domain. In other embodiments, the amino acid sequence of the linker is GGSGGS (SEQ ID NO: 250) or GSSGSETPGTSESATPESSG (SEQ ID NO: 251). In other embodiments, the linker is a rigid linker. In other embodiments of the above aspects, the linker is encoded by GGAGGCTCTGGAGGAAGC (SEQ ID NO: 252) or GGCTCTTCTGGATCTGAAACACCTGGCACAAGCGAGAGCGCCACCCCTGAGAGCTCTGGC (SEQ ID NO: 253).

[0274] Fusion proteins containing heterologous catalytic domains flanking the N- and C-terminal fragments of a Cas12 polypeptide are also useful for base editing in the methods described herein. Fusion proteins containing Cas12 and one or more deaminase domains (e.g., adenosine deaminase, or an adenosine deaminase domain flanking a Cas12 sequence) are also useful for highly specific and efficient base editing of target sequences. In one embodiment, a chimeric Cas12 fusion protein contains a heterologous catalytic domain (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) inserted within a Cas12 polypeptide. In some embodiments, the fusion protein contains an adenosine deaminase domain and a cytidine deaminase domain inserted within Cas12. In some embodiments, the adenosine deaminase is fused within Cas12 and the cytidine deaminase is fused to the C-terminus. In some embodiments, adenosine deaminase is fused within Cas12 and cytidine deaminase is fused to the N-terminus. In some embodiments, cytidine deaminase is fused within Cas12 and adenosine deaminase is fused to the C-terminus. In some embodiments, cytidine deaminase is fused within Cas12 and adenosine deaminase is fused to the N-terminus. Exemplary structures of fusion proteins having adenosine deaminase and cytidine deaminase and Cas12 are provided below: NH2-[Cas12(adenosine deaminase)]-[cytidine deaminase]-COOH, NH2-[cytidine deaminase]-[Cas12(adenosine deaminase)]-COOH, NH2-[Cas12(cytidine deaminase)]-[adenosine deaminase]-COOH, or NH2-[adenosine deaminase]-[Cas12 (cytidine deaminase)]-COOH, In some embodiments, the "-" used in the general architecture above indicates the presence of an optional linker.

[0275] In various embodiments, the catalytic domain has a DNA-modifying activity (e.g., deaminase activity), such as adenosine deaminase activity. In some embodiments, the adenosine deaminase is TadA (e.g., TadA*7.10). In some embodiments, TadA is TadA*8. In some embodiments, TadA*8 is fused within Cas12 and a cytidine deaminase is fused to the C-terminus. In some embodiments, TadA*8 is fused within Cas12 and a cytidine deaminase is fused to the N-terminus. In some embodiments, a cytidine deaminase is fused within Cas12 and TadA*8 is fused to the C-terminus. In some embodiments, a cytidine deaminase is fused within Cas12 and TadA*8 is fused to the N-terminus. Exemplary structures of fusion proteins having TadA*8 and cytidine deaminase and Cas12 are provided below: N-[Cas12(TadA*8)]-[cytidine deaminase]-C, N-[cytidine deaminase]-[Cas12(TadA*8)]-C, N-[Cas12(cytidine deaminase)]-[TadA*8]-C, or N-[TadA*8]-[Cas12(cytidine deaminase)]-C. In some embodiments, the "-" used in the general architecture above indicates the presence of an optional linker.

[0276] In other embodiments, the fusion protein contains one or more catalytic domains. In other embodiments, at least one of the one or more catalytic domains is inserted within a Cas12 polypeptide or fused at the N-terminus or C-terminus of Cas12. In other embodiments, at least one of the one or more catalytic domains is inserted within a loop, alpha-helical region, unstructured portion, or solvent-accessible portion of a Cas12 polypeptide. In other embodiments, the Cas12 polypeptide is Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ. In other embodiments, the Cas12 polypeptide has at least about 85% amino acid sequence identity to Bacillus hisashii Cas12b, Bacillus thermoamylovorans Cas12b, Bacillus species V3-13 Cas12b, or Alicyclobacillus acidiphilus Cas12b (SEQ ID NO: 254). In other embodiments, the Cas12 polypeptide has at least about 90% amino acid sequence identity to Bacillus hisashii Cas12b (SEQ ID NO:255), Bacillus thermoamylovorans Cas12b, Bacillus sp. V3-13 Cas12b, or Alicyclobacillus acidiphilus Cas12b. In other embodiments, the Cas12 polypeptide has at least about 95% amino acid sequence identity to Bacillus hisashii Cas12b, Bacillus thermoamylovorans Cas12b (SEQ ID NO:256), Bacillus sp. V3-13 Cas12b, or Alicyclobacillus acidiphilus Cas12b (SEQ ID NO:257).In other embodiments, the Cas12 polypeptide contains or consists essentially of a fragment of Bacillus hisashii Cas12b, Bacillus thermoamylovorans Cas12b, Bacillus species V3-13 Cas12b, or Alicyclobacillus acidiphilus Cas12b. In embodiments, the Cas12 polypeptide contains BvCas12b (V4), which in some embodiments is represented as: 5' mRNA Cap---5' UTR---bhCas12b---termination sequence---3' UTR---120 polyA tail (SEQ ID NOs:258-260).

[0277] In other embodiments, the catalytic domain is inserted between amino acid positions 153-154, 255-256, 306-307, 980-981, 1019-1020, 534-535, 604-605, or 344-345 of BhCas12b, or between the corresponding amino acid residues of Cas12a, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ. In other embodiments, the catalytic domain is inserted between amino acids P153 and S154 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acids K255 and E256 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acids D980 and G981 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acids K1019 and L1020 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acids F534 and P535 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acids K604 and G605 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acids H344 and F345 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acid positions 147 and 148, 248 and 249, 299 and 300, 991 and 992, or 1031 and 1032 of BvCas12b, or between the corresponding amino acid residues of Cas12a, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ. In other embodiments, the catalytic domain is inserted between amino acids P147 and D148 of BvCas12b. In other embodiments, the catalytic domain is inserted between amino acids G248 and G249 of BvCas12b. In other embodiments, the catalytic domain is inserted between amino acids P299 and E300 of BvCas12b. In other embodiments, the catalytic domain is inserted between amino acids G991 and E992 of BvCas12b. In other embodiments, the catalytic domain is inserted between amino acids K1031 and M1032 of BvCas12b.In other embodiments, the catalytic domain is inserted between amino acid positions 157 and 158, 258 and 259, 310 and 311, 1008 and 1009, or 1044 and 1045 of AaCas12b, or between the corresponding amino acid residues of Cas12a, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ. In other embodiments, the catalytic domain is inserted between amino acids P157 and G158 of AaCas12b. In other embodiments, the catalytic domain is inserted between amino acids V258 and G259 of AaCas12b. In other embodiments, the catalytic domain is inserted between amino acids D310 and P311 of AaCas12b. In other embodiments, the catalytic domain is inserted between amino acids G1008 and E1009 of AaCas12b. In other embodiments, the catalytic domain is inserted between amino acids G1044 and K1045 of AaCas12b.

[0278] In other embodiments, the fusion protein contains a nuclear localization signal (e.g., a bipartite nuclear localization signal). In other embodiments, the amino acid sequence of the nuclear localization signal is MAPKKKRKVGIHGVPAA (SEQ ID NO: 261). In other embodiments of the above aspects, the nuclear localization signal is encoded by the following sequence: ATGGCCCCAAAGAAGAAGCGGAAGGTCGGTATCCACGGAGTCCCAGCAGCC (SEQ ID NO: 262). In other embodiments, the Cas12b polypeptide contains a mutation that silences the catalytic activity of the RuvC domain. In other embodiments, the Cas12b polypeptide contains a D574A, D829A, and / or D952A mutation. In other embodiments, the fusion protein further contains a tag (e.g., an influenza hemagglutinin tag).

[0279] In some embodiments, the fusion protein comprises a napDNAbp domain (e.g., a Cas12-derived domain) with an internally fused nucleobase editing domain (e.g., all or part of a deaminase domain (e.g., an adenosine deaminase domain)). In some embodiments, the napDNAbp is Cas12b. In some embodiments, the base editor comprises a BhCas12b domain with an internally fused TadA*8 domain inserted into a locus provided in Table 5 below.

[0280] [Table 5]

[0281] As a non-limiting example, an adenosine deaminase (e.g., TadA*8.13) can be inserted into BhCas12b to generate a fusion protein (e.g., TadA*8.13-BhCas12b) that effectively edits nucleic acid sequences.

[0282] In some embodiments, the base editor systems described herein are ABEs in which TadA has been inserted into Cas9. The polypeptide sequences of relevant ABEs with TadA inserted into Cas9 are provided as SEQ ID NOs: 263-308 in the accompanying sequence listing.

[0283] In some embodiments, an adenosine deaminase base editor was generated to insert TadA, or a variant thereof, into a Cas9 polypeptide at a specified position.

[0284] Exemplary, but non-limiting, fusion proteins are described in International PCT Application No. PCT / US2020 / 016285 and U.S. Provisional Application Nos. 62 / 852,228 and 62 / 852,224, the contents of which are incorporated by reference in their entireties.

[0285] Editing A to G In some embodiments, the base editors described herein comprise an adenosine deaminase domain. Such an adenosine deaminase domain of a base editor can facilitate editing of an adenine (A) nucleobase to a guanine (G) nucleobase by deaminating A to form inosine (I), which exhibits G base pairing properties. Adenosine deaminase can deaminate (i.e., remove the amine group from) the adenine of a deoxyadenosine residue in deoxyribonucleic acid (DNA). In some embodiments, an A-to-G base editor further comprises an inhibitor of inosine base excision repair, such as a uracil glycosylase inhibitor (UGI) domain or a catalytically inactive inosine-specific nuclease. Without wishing to be bound by any particular theory, the UGI domain or catalytically inactive inosine-specific nuclease can inhibit or prevent base excision repair of deaminated adenosine residues (e.g., inosine), which can improve the activity or efficiency of the base editor.

[0286] Base editors comprising adenosine deaminase can act on any polynucleotide, including DNA, RNA, and DNA-RNA hybrids. In certain embodiments, base editors comprising adenosine deaminase can deaminate target A in polynucleotides comprising RNA. For example, a base editor can comprise an adenosine deaminase domain capable of deaminating target A in RNA polynucleotides and / or DNA-RNA hybrid polynucleotides. In one embodiment, the adenosine deaminase incorporated into the base editor comprises all or a portion of an adenosine deaminase that acts on RNA (ADAR, e.g., ADAR1 or ADAR2) or tRNA (ADAT). Base editors comprising an adenosine deaminase domain can also deaminate A nucleobases in DNA polynucleotides. In one embodiment, the adenosine deaminase domain of the base editor comprises all or a portion of ADAT containing one or more mutations that enable ADAT to deaminate target A in DNA. For example, a base editor can comprise all or a portion of ADAT from Escherichia coli (EcTadA) containing one or more of the following mutations: D108N, A106V, D147Y, E155V, L84F, H123Y, I156F, or a corresponding mutation in another adenosine deaminase. Exemplary ADAT homolog polypeptide sequences are provided in the Sequence Listing as SEQ ID NOs: 1 and 309-315.

[0287] The adenosine deaminase can be derived from any suitable organism (e.g., E. coli). In some embodiments, the adenosine deaminase is derived from a prokaryote. In some embodiments, the adenosine deaminase is derived from a bacterium. In some embodiments, the adenosine deaminase is derived from Escherichia coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis. In some embodiments, the adenosine deaminase is derived from E. coli. In some embodiments, the adenosine deaminase is a naturally occurring adenosine deaminase that includes one or more mutations corresponding to any of the mutations provided herein (e.g., mutations in ecTadA). Corresponding residues in any homologous proteins can be identified, for example, by sequence alignment and determination of homologous residues. Mutations in any naturally occurring adenosine deaminase (e.g., with homology to ecTadA) that correspond to any of the mutations described herein (e.g., any of the mutations identified in ecTadA) can be generated accordingly.

[0288] In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the amino acid sequences described in any of the adenosine deaminases provided herein. It should be understood that the adenosine deaminases provided herein can comprise one or more mutations (e.g., any of the mutations provided herein). The present disclosure provides any deaminase domain that has a specific percent identity, as well as any of the mutations described herein or combinations thereof. In some embodiments, the adenosine deaminase comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to a reference sequence or any of the adenosine deaminases provided herein. In some embodiments, the adenosine deaminase comprises an amino acid sequence having at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, or at least 170 identical contiguous amino acid residues compared to any one of the amino acid sequences known in the art or described herein.

[0289] It should be understood that any of the mutations provided herein (e.g., based on the TadA reference sequence) can be introduced into other adenosine deaminases, such as E. coli TadA (ecTadA), S. aureus TadA (saTadA), or other adenosine deaminases (e.g., bacterial adenosine deaminases). It will be apparent to one of skill in the art that additional deaminases can be similarly aligned to identify homologous amino acid residues and mutated as provided herein. Thus, any of the mutations identified in the TadA reference sequence can be made in other adenosine deaminases (e.g., ecTadA) that have homologous amino acid residues. It should also be understood that any of the mutations provided herein can be made in the TadA reference sequence or another adenosine deaminase, individually or in any combination.

[0290] In some embodiments, the adenosine deaminase comprises a D108X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a D108G, D108N, D108V, D108A, or D108Y mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase. However, it should be understood that additional deaminases can be similarly aligned to identify homologous amino acid residues and can be mutated as provided herein.

[0291] In some embodiments, the adenosine deaminase comprises an A106X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A106V mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).

[0292] In some embodiments, the adenosine deaminase comprises an E155X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where the presence of X indicates any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an E155D, E155G, or E155V mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).

[0293] In some embodiments, the adenosine deaminase comprises a D147X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where the presence of X indicates any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises D147Y, a mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).

[0294] In some embodiments, the adenosine deaminase comprises an A106X, E155X, or D147X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA), where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an E155D, E155G, or E155V mutation. In some embodiments, the adenosine deaminase comprises D147Y.

[0295] It should also be understood that any of the mutations provided herein can be made in ecTadA or another adenosine deaminase, individually or in any combination. For example, an adenosine deaminase can contain D108N, A106V, E155V, and / or D147Y mutations in the TadA reference sequence, or the corresponding mutations in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase includes the following mutations in the TadA reference sequence (mutations are separated by ";"), or corresponding mutations in another adenosine deaminase: D108N and A106V; D108N and E155V; D108N and D147Y; A106V and E155V; A106V and D147Y; E155V and D147Y; D108N, A106V, and E155V; D108N, A106V, and D147Y; D108N, E155V, and D147Y; A106V, E155V, and D147Y; and D108N, A106V, E155V, and D147Y. However, it should be understood that any combination of the corresponding mutations provided herein can be made in an adenosine deaminase (eg, ecTadA).

[0296] In some embodiments, the adenosine deaminase comprises one or more of the H8X, T17X, ​​L18X, W23X, L34X, W45X, R51X, A56X, E59X, E85X, M94X, I95X, V102X, F104X, A106X, R107X, D108X, K110X, M118X, N127X, A138X, F149X, M151X, R153X, Q154X, I156X and / or K157X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein the occurrence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the H8Y, T17S, L18E, W23L, L34S, W45L, R51H, A56E or A56S, E59G, E85K or E85G, M94L, I95L, V102A, F104L, A106V, R107C or R107H or R107P, D108G or D108N or D108V or D108A or D108Y, K110I, M118K, N127S, A138V, F149Y, M151V, R153C, Q154L, I156D, and / or K157R mutations in a TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.

[0297] In some embodiments, the adenosine deaminase comprises one or more of an H8X, D108X, and / or N127X mutation in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, where X indicates the presence of any amino acid. In some embodiments, the adenosine deaminase comprises one or more of an H8Y, D108N, and / or N127S mutation in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.

[0298] In some embodiments, the adenosine deaminase comprises one or more of the H8X, R26X, M61X, L68X, M70X, A106X, D108X, A109X, N127X, D147X, R152X, Q154X, E155X, K161X, Q163X and / or T166X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, where X represents the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of H8Y, R26W, M61I, L68Q, M70V, A106T, D108N, A109T, N127S, D147Y, R152C, Q154H or Q154R, E155G or E155V or E155D, K161Q, Q163H, and / or T166P in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.

[0299] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8X, D108X, N127X, D147X, R152X, and Q154X in the TadA reference sequence, or corresponding mutations in another adenosine deaminase (e.g., ecTadA), where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8X, M61X, M70X, D108X, N127X, Q154X, E155X, and Q163X in the TadA reference sequence, or the corresponding mutations in another adenosine deaminase (e.g., ecTadA), where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8X, D108X, N127X, E155X, and T166X in the TadA reference sequence, or the corresponding mutation(s) in another adenosine deaminase (e.g., ecTadA), where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase.

[0300] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8X, A106X, and D108X, or a corresponding mutation(s) in another adenosine deaminase (where X represents the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase). In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8X, R26X, L68X, D108X, N127X, D147X, and E155X, or a corresponding mutation(s) in another adenosine deaminase (where X represents the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase).

[0301] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, or seven mutations selected from the group consisting of H8X, R126X, L68X, D108X, N127X, D147X, and E155X in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase). In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8X, D108X, A109X, N127X, and E155X in the TadA reference sequence, or a corresponding mutation(s) in another adenosine deaminase (where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase).

[0302] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8Y, D108N, N127S, D147Y, R152C, and Q154H in the TadA reference sequence, or a corresponding mutation(s) in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8Y, M61I, M70V, D108N, N127S, Q154R, E155G, and Q163H in the TadA reference sequence, or a corresponding mutation(s) in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8Y, D108N, N127S, E155V, and T166P in the TadA reference sequence, or corresponding mutation(s) in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8Y, A106T, D108N, N127S, E155D, and K161Q in the TadA reference sequence, or corresponding mutation(s) in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8Y, R26W, L68Q, D108N, N127S, D147Y, and E155V in the TadA reference sequence, or corresponding mutation(s) in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8Y, D108N, A109T, N127S, and E155G in the TadA reference sequence, or corresponding mutation(s) in another adenosine deaminase (e.g., ecTadA).

[0303] In some embodiments, the adenosine deaminase comprises one or more of the above mutations or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises a D108N, D108G, or D108V mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A106V and D108N mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R107C and D108N mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises H8Y, D108N, N127S, D147Y, and Q154H mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises H8Y, D108N, N127S, D147Y, and E155V mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises D108N, D147Y, and E155V mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises H8Y, D108N, and N127S mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises the A106V, D108N, D147Y, and E155V mutations in the TadA reference sequence, or the corresponding mutations in another adenosine deaminase (e.g., ecTadA).

[0304] In some embodiments, the adenosine deaminase comprises one or more of the S2X, H8X, I49X, L84X, H123X, N127X, I156X, and / or K160X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, where the presence of X denotes any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the S2A, H8Y, I49F, L84F, H123Y, N127S, I156F, and / or K160S mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase (e.g., ecTadA).

[0305] In some embodiments, the adenosine deaminase comprises an L84X mutant adenosine deaminase (where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase). In some embodiments, the adenosine deaminase comprises an L84F mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).

[0306] In some embodiments, the adenosine deaminase comprises an H123X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an H123Y mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.

[0307] In some embodiments, the adenosine deaminase comprises an I156X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an I156F mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.

[0308] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, or seven mutations selected from the group consisting of L84X, A106X, D108X, H123X, D147X, E155X, and I156X in the TadA reference sequence, or corresponding one or more mutations in another adenosine deaminase, where X represents the presence of any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of S2X, I49X, A106X, D108X, D147X, and E155X in the TadA reference sequence, or corresponding mutations in another adenosine deaminase, where X represents the presence of any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8X, A106X, D108X, N127X, and K160X in the TadA reference sequence, or the corresponding mutation(s) in another adenosine deaminase (where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase).

[0309] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, or seven mutations selected from the group consisting of L84F, A106V, D108N, H123Y, D147Y, E155V, and I156F in the TadA reference sequence, or the corresponding one or more mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of S2A, I49F, A106V, D108N, D147Y, and E155V in the TadA reference sequence.

[0310] In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8Y, A106T, D108N, N127S, and K160S in the TadA reference sequence, or the corresponding one or more mutations in another adenosine deaminase.

[0311] In some embodiments, the adenosine deaminase comprises one or more of the E25X, R26X, R107X, A142X, and / or A143X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein the occurrence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of E25M, E25D, E25A, E25R, E25V, E25S, E25Y, R26G, R26N, R26Q, R26C, R26L, R26K, R107P, R107K, R107A, R107N, R107W, R107H, R107S, A142N, A142D, A142G, A143D, A143G, A143E, A143L, A143W, A143M, A143S, A143Q, and / or A143R in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the mutations described herein corresponding to the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.

[0312] In some embodiments, the adenosine deaminase comprises an E25X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an E25M, E25D, E25A, E25R, E25V, E25S, or E25Y mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).

[0313] In some embodiments, the adenosine deaminase comprises an R26X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R26G, R26N, R26Q, R26C, R26L, or R26K mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).

[0314] In some embodiments, the adenosine deaminase comprises a R107X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in a wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a R107P, R107K, R107A, R107N, R107W, R107H, or R107S mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).

[0315] In some embodiments, the adenosine deaminase comprises an A142X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A142N, A142D, A142G mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).

[0316] In some embodiments, the adenosine deaminase comprises an A143X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in a wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A143D, A143G, A143E, A143L, A143W, A143M, A143S, A143Q, and / or A143R mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).

[0317] In some embodiments, the adenosine deaminase comprises one or more of the H36X, N37X, P48X, I49X, R51X, M70X, N72X, D77X, E134X, S146X, Q154X, K157X and / or K161X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein the occurrence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the following mutations in the TadA reference sequence: H36L, N37T, N37S, P48T, P48L, I49V, R51H, R51L, M70L, N72S, D77G, E134G, S146R, S146C, Q154H, K157N, and / or K161T, or one or more corresponding mutations in another adenosine deaminase (e.g., ecTadA).

[0318] In some embodiments, the adenosine deaminase comprises an H36X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an H36L mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.

[0319] In some embodiments, the adenosine deaminase comprises an N37X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an N37T or N37S mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.

[0320] In some embodiments, the adenosine deaminase comprises a P48X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a P48T or P48L mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.

[0321] In some embodiments, the adenosine deaminase comprises an R51X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R51H or R51L mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.

[0322] In some embodiments, the adenosine deaminase comprises a S146X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a S146R or S146C mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.

[0323] In some embodiments, the adenosine deaminase comprises a K157X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a K157N mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.

[0324] In some embodiments, the adenosine deaminase comprises a P48X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a P48S, P48T, or P48A mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.

[0325] In some embodiments, the adenosine deaminase comprises an A142X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A142N mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.

[0326] In some embodiments, the adenosine deaminase comprises a W23X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a W23R or W23L mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.

[0327] In some embodiments, the adenosine deaminase comprises a R152X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a R152P or R52H mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.

[0328] In one embodiment, the adenosine deaminase may comprise the mutations H36L, R51L, L84F, A106V, D108N, H123Y, S146C, D147Y, E155V, I156F, and K157N. In some embodiments, the adenosine deaminase comprises the following combination of mutations relative to the TadA reference sequence (where each mutation in the combination is separated by a "_" and each combination of mutations is between parentheses: (A106V_D108N), (R107C_D108N), (H8Y_D108N_N127S_D147Y_Q154H), (H8Y_D108N_N127S_D147Y_E155V), (D108N_D147Y_E155V), (H8Y_D108N_N127S), (H8Y_D108N_N127S_D147Y_Q154H), (A106V_D108N_D147Y_E155V), (D108Q_D147Y_E155V), (D108M_D147Y_E155V), (D108L_D147Y_E155V), (D108K_D147Y_E155V), (D108I_D147Y_E155V), (D108F_D147Y_E155V), (A106V_D108N_D147Y), (A106V_D108M_D147Y_E155V), (E59A_A106V_D108N_D147Y_E155V)、 (E59A cat dead_A106V_D108N_D147Y_E155V)、 (L84F_A106V_D108N_H123Y_D147Y_E155V_I156Y)、 (L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (D103A_D104N)、 (G22P_D103A_D104N)、 (D103A_D104N_S138A)、 (R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F)、 (E25G_R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F)、 (E25D_R26G_L84F_A106V_R107K_D108N_H123Y_A142N_A143G_D147Y_E155V_I156F)、(R26Q_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F)、 (E25M_R26G_L84F_A106V_R107P_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F)、(R26C_L84F_A106V_R107H_D108N_H123Y_A142N_D147Y_E155V_I156F)、(L84F_A106V_D108N_H123Y_A142N_A143L_D147Y_E155V_I156F)、 (R26G_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F)、 (E25A_R26G_L84F_A106V_R107N_D108N_H123Y_A142N_A143E_D147Y_E155V_I156F)、(R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F)、 (A106V_D108N_A142N_D147Y_E155V)、 (R26G_A106V_D108N_A142N_D147Y_E155V)、 (E25D_R26G_A106V_R107K_D108N_A142N_A143G_D147Y_E155V)、(R26G_A106V_D108N_R107H_A142N_A143D_D147Y_E155V)、 (E25D_R26G_A106V_D108N_A142N_D147Y_E155V)、 (A106V_R107K_D108N_A142N_D147Y_E155V)、 (A106V_D108N_A142N_A143G_D147Y_E155V)、 (A106V_D108N_A142N_A143L_D147Y_E155V)、 (H36L_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、 (N37T_P48T_M70L_L84F_A106V_D108N_H123Y_D147Y_I49V_E155V_I156F)、 (N37S_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K161T)、 (H36L_L84F_A106V_D108N_H123Y_D147Y_Q154H_E155V_I156F)、 (N72S_L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F)、 (H36L_P48L_L84F_A106V_D108N_H123Y_E134G_D147Y_E155V_I156F)、 (H36L_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K157N)(H36L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F)、 (L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F_K161T)、 (N37S_R51H_D77G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (R51L_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K157N)、 (D24G_Q71R_L84F_H96L_A106V_D108N_H123Y_D147Y_E155V_I156F_K160E)、 (H36L_G67V_L84F_A106V_D108N_H123Y_S146T_D147Y_E155V_I156F)、 (Q71L_L84F_A106V_D108N_H123Y_L137M_A143E_D147Y_E155V_I156F)、 (E25G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_Q159L)、 (L84F_A91T_F104I_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (N72D_L84F_A106V_D108N_H123Y_G125A_D147Y_E155V_I156F)、 (P48S_L84F_S97C_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (W23G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (D24G_P48L_Q71R_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_Q159L)、 (L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F)、 (H36L_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_E155V_I156F_K157N)、(N37S_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F_K161T)、 (L84F_A106V_D108N_D147Y_E155V_I156F)、 (R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N_K161T)、 (L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K161T)、 (L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N_K160E_K161T)、 (L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N_K160E)、(R74Q_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (R74A_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (R74Q_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (L84F_R98Q_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (L84F_A106V_D108N_H123Y_R129Q_D147Y_E155V_I156F)、 (P48S_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F)、 (P48S_A142N)、 (P48T_I49V_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F_L157N)、 (P48T_I49V_A142N)、 (H36L_P48S_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、(H36L_P48S_R51L_L84F_A106V_D108N_H123Y_S146C_A142N_D147Y_E155V_I156F(H36L_P48T_I49V_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、 (H36L_P48T_I49V_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_E155V_I156F_K157N)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_E155V_I156F_K157N)、(H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_A142N_D147Y_E155V_I156F_K157N)、 (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、(W23R_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、(W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F_K161T)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152H_E155V_I156F_K157N)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152P_E155V_I156F_K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152P_E155V_I156F_K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142A_S146C_D147Y_E155V_I156F_K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142A_S146C_D147Y_R152P_E155V_I156F_K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F_K161T), (W23R_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152P_E155V_I156F_ K157N), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_R152P_E155 V_I156F_K157N).

[0329] In some embodiments, the TadA deaminase is a TadA variant. In some embodiments, the TadA variant is TadA*7.10. In certain embodiments, the fusion protein comprises a single TadA*7.10 domain (e.g., provided as a monomer). In other embodiments, the fusion protein comprises TadA*7.10 and TadA(wt), which can form a heterodimer. In one embodiment, a fusion protein of the invention comprises wild-type TadA bound to TadA*7.10, which is bound to a Cas9 nickase.

[0330] In some embodiments, TadA*7.10 comprises at least one modification. In some embodiments, the adenosine deaminase comprises a modification in the following sequence: TadA*7.10 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 1).

[0331] In some embodiments, TadA*7.10 comprises modifications at amino acids 82 and / or 166. In certain embodiments, TadA*7.10 comprises one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. In other embodiments, variants of TadA*7.10 comprise a combination of modifications selected from the following group: Y147T+Q154R, Y147T+Q154S, Y147R+Q154S, V82S+Q154S, V82S+Y147R, V82S+Q154R, V82S+Y123H, I76Y+V82S, V82S+Y123H+Y147T, V82 S+Y123H+Y147R, V82S+Y123H+Q154R, Y147R+Q154R+Y123H, Y147R+Q154R+I76Y, Y147R+Q154R+T166R, Y123H+Y147R+Q154R+I76Y, V82S+Y123H+Y147R+Q154R, and I76Y+V82S+Y123H+Y147R+Q154R.

[0332] In some embodiments, the adenosine deaminase variant (e.g., TadA*8) comprises a deletion. In some embodiments, the adenosine deaminase variant comprises a C-terminal deletion. In particular embodiments, the adenosine deaminase variant comprises TadA*7.10, a C-terminal deletion starting at residues 149, 150, 151, 152, 153, 154, 155, 156, and 157 relative to the TadA reference sequence, or a corresponding mutation in another TadA.

[0333] In other embodiments, the adenosine deaminase variant (e.g., TadA*8) is a monomer that includes one or more of the following modifications: TadA*7.10, Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R relative to the TadA reference sequence, or corresponding mutations in another TadA. In other embodiments, the variant of adenosine deaminase (TadA*8) is a monomer comprising a combination of modifications selected from the following group: TadA*7.10, Y147T+Q154R, Y147T+Q154S, Y147R+Q154S, V82S+Q154S, V82S+Y147R, V82S+Q154R, V82S+Y123H, I76Y+V82S, V82S+Y123H+ Y147T, V82S+Y123H+Y147R, V82S+Y123H+Q154R, Y147R+Q154R+Y123H, Y147R+Q154R+I76Y, Y147R+Q154R+T166R, Y123H+Y147R+Q154R+I76Y, V82S+Y123H+Y147R+Q154R, and I76Y+V82S+Y123H+Y147R+Q154R, or the corresponding mutations in another TadA.

[0334] In other embodiments, the adenosine deaminase variant is a homodimer comprising two adenosine deaminase domains (e.g., TadA*8), each having one or more of the following modifications relative to the TadA*7.10, TadA reference sequence: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R, or corresponding mutations in another TadA. In other embodiments, the adenosine deaminase variant is a monomer comprising two adenosine deaminase domains (e.g., TadA*8), each having a combination of modifications selected from the following group: TadA*7.10, Y147T+Q154R, Y147T+Q154S, Y147R+Q154S, V82S+Q154S, V82S+Y147R, V82S+Q154R, V82S+Y123H, I76Y+V relative to the TadA reference sequence. 82S, V82S+Y123H+Y147T, V82S+Y123H+Y147R, V82S+Y123H+Q154R, Y147R+Q154R+Y123H, Y147R+Q154R+I76Y, Y147R+Q154R+T166R, Y123H+Y147R+Q154R+I76Y, V82S+Y123H+Y147R+Q154R, and I76Y+V82S+Y123H+Y147R+Q154R, or the corresponding mutations in another TadA.

[0335] In other embodiments, the adenosine deaminase variant is a heterodimer of a wild-type adenosine deaminase domain and an adenosine deaminase variant domain containing one or more of the following modifications relative to the TadA reference sequence: TadA*7.10, Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R, or corresponding mutations in another TadA (e.g., TadA*8). In other embodiments, the adenosine deaminase variant is a heterodimer of a wild-type adenosine deaminase domain and an adenosine deaminase variant domain (e.g., TadA*8) that includes a combination of modifications selected from the following group: TadA*7.10, Y147T+Q154R, Y147T+Q154S, Y147R+Q154S, V82S+Q154S, V82S+Y147R, V82S+Q154R, V82S+Y12 relative to the TadA reference sequence. 3H, I76Y+V82S, V82S+Y123H+Y147T, V82S+Y123H+Y147R, V82S+Y123H+Q154R, Y147R+Q154R+Y123H, Y147R+Q154R+I76Y, Y147R+Q154R+T166R, Y123H+Y147R+Q154R+I76Y, V82S+Y123H+Y147R+Q154R, and I76Y+V82S+Y123H+Y147R+Q154R, or the corresponding mutations in another TadA.

[0336] In other embodiments, the adenosine deaminase variant is a heterodimer of the TadA*7.10 domain and an adenosine deaminase variant domain containing one or more of the following modifications relative to the TadA*7.10, TadA reference sequence: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R, or corresponding mutations in another TadA (e.g., TadA*8). In other embodiments, the adenosine deaminase variant is a heterodimer of a TadA*7.10 domain and an adenosine deaminase variant domain (e.g., TadA*8), wherein the adenosine deaminase variant domain comprises a combination of modifications relative to the TadA reference sequence TadA*7.10 selected from the following group or corresponding mutations in another TadA: Y147T+Q154R, Y147T+Q154S, Y147R+Q154S, V82S ... 2S+Y147R, V82S+Q154R, V82S+Y123H, I76Y+V82S, V82S+Y123H+Y147T, V82S+Y123H+Y147R, V82S+Y123H+Q154R, Y147R+Q154R+Y123H, Y147R+Q154R+I76Y, Y147R+Q154R+T166R, Y123H+Y147R+Q154R+I76Y, V82S+Y123H+Y147R+Q154R and I76Y+V82S+Y123H+Y147R+Q154R.

[0337] In certain embodiments, the adenosine deaminase heterodimer comprises a TadA*8 domain and an adenosine deaminase domain selected from Staphylococcus aureus (S. aureus) TadA, Bacillus subtilis (B. subtilis) TadA, Salmonella typhimurium (S. typhimurium) TadA, Shewanella putrefaciens (S. putrefaciens) TadA, Haemophilus influenzae F3031 (H. influenzae) TadA, Caulobacter crescentus (C. crescentus) TadA, Geobacter sulfurreducens (G. sulfurreducens) TadA, or TadA*7.10.

[0338] In some embodiments, the adenosine deaminase is TadA*8. In one embodiment, the adenosine deaminase is TadA*8 comprising or consisting essentially of the following sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 316).

[0339] In some embodiments, TadA*8 is truncated. In some embodiments, the truncated TadA*8 lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to full-length TadA*8. In some embodiments, the truncated TadA*8 lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to full-length TadA*8. In some embodiments, the variant of adenosine deaminase is full-length TadA*8.

[0340] In some embodiments, TadA*8 is TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23, or TadA*8.24.

[0341] In other embodiments, base editors of the disclosure comprise a variant of adenosine deaminase (e.g., TadA*8) monomers, comprising one or more of the following modifications: TadA*7.10, R26C, V88A, A109S, T111R, D119N, H122N, Y147D, F149Y, T166I and / or D167N relative to the TadA reference sequence, or corresponding mutations in another TadA. In other embodiments, the variant adenosine deaminase (TadA*8) monomer comprises a combination of modifications selected from the following group: TadA*7.10, R26C+A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N, V88A ... 22N+F149Y+T166I+D167N, R26C+A109S+T111R+D119N+H122N+F149Y+T166I+D167N, V88A+T111R+D119N+F149Y, and A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N, or the corresponding mutations in another TadA.

[0342] In other embodiments, the base editor comprises a heterodimer of a wild-type adenosine deaminase domain and an adenosine deaminase variant domain that includes TadA*7.10, one or more of the following modifications relative to the TadA reference sequence: R26C, V88A, A109S, T111R, D119N, H122N, Y147D, F149Y, T166I, and / or D167N, or a corresponding mutation in another TadA (e.g., TadA*8). In other embodiments, the base editor comprises a heterodimer of a wild-type adenosine deaminase domain and an adenosine deaminase variant domain (e.g., TadA*8) that comprises a combination of modifications selected from the following group: TadA*7.10, R26C+A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N, V88A relative to the TadA reference sequence +A109S+T111R+D119N+H122N+F149Y+T166I+D167N, R26C+A109S+T111R+D119N+H122N+F149Y+T166I+D167N, V88A+T111R+D119N+F149Y, and A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N, or the corresponding mutations in another TadA.

[0343] In other embodiments, the base editor is a heterodimer of a TadA*7.10 domain and an adenosine deaminase variant domain that includes one or more of the following modifications relative to the TadA*7.10, TadA reference sequence: R26C, V88A, A109S, T111R, D119N, H122N, Y147D, F149Y, T166I, and / or D167N, or corresponding mutations in another TadA (e.g., TadA*8). In other embodiments, the base editor comprises a heterodimer of a TadA*7.10 domain and an adenosine deaminase variant domain (e.g., TadA*8) comprising a combination of modifications selected from the following group: TadA*7.10, R26C+A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N, V88A+A relative to the TadA reference sequence 109S+T111R+D119N+H122N+F149Y+T166I+D167N, R26C+A109S+T111R+D119N+H122N+F149Y+T166I+D167N, V88A+T111R+D119N+F149Y, and A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N, or the corresponding mutations in another TadA.

[0344] In some embodiments, TadA*8 is a variant shown in Table 6. Table 6 shows the numbers of specific amino acid positions in the TadA amino acid sequence and the amino acids present at those positions in TadA-7.10 adenosine deaminase. Table 6 also shows the amino acid changes of TadA variants relative to TadA-7.10 after phage-assisted non-persistent evolution (PANCE) and phage-assisted persistent evolution (PACE) as described in M. Richter et al., 2020, Nature Biotechnology, doi.org / 10.1038 / s41587-020-0453-z, the entire contents of which are incorporated herein by reference. In some embodiments, TadA*8 is TadA*8a, TadA*8b, TadA*8c, TadA*8d, or TadA*8e. In some embodiments, TadA*8 is TadA*8e.

[0345] [Table 6]

[0346] In one embodiment, a fusion protein of the invention comprises wild-type TadA linked to a variant of adenosine deaminase described herein (e.g., TadA*8) and linked to a Cas9 nickase. In certain embodiments, the fusion protein comprises a single TadA*8 domain (e.g., provided as a monomer). In other embodiments, the fusion protein comprises TadA*8 and TadA(wt), which are capable of forming a heterodimer.

[0347] In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the amino acid sequences described in any of the adenosine deaminases provided herein. It should be understood that the adenosine deaminases provided herein can comprise one or more mutations (e.g., any of the mutations provided herein). The present disclosure provides any deaminase domain that has a specific percent identity, as well as any of the mutations described herein or combinations thereof. In some embodiments, the adenosine deaminase comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to a reference sequence or any of the adenosine deaminases provided herein. In some embodiments, the adenosine deaminase comprises an amino acid sequence having at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, or at least 170 identical contiguous amino acid residues compared to any one of the amino acid sequences known in the art or described herein.

[0348] In certain embodiments, TadA*8 comprises one or more mutations at any of the following positions shown in bold: In other embodiments, TadA*8 comprises one or more mutations at any of the following underlined positions: TIFF2024534639000039.tif32165 (SEQ ID NO: 1)

[0349] For example, TadA*8 includes, relative to TadA*7.10, the TadA reference sequence, modifications at amino acid positions 82 and / or 166 (e.g., V82S, T166R) alone or in combination with any one or more of the following: Y147T, Y147R, Q154S, Y123H, and / or Q154R, or corresponding mutations in another TadA. In certain embodiments, the combination of modifications is selected from the following group: TadA*7.10, Y147T+Q154R, Y147T+Q154S, Y147R+Q154S, V82S+Q154S, V82S+Y147R, V82S+Q154R, V82S+Y123H, I76Y+V82S, V82S+Y123H+Y147T, V82S+Y123H+ Y147R, V82S+Y123H+Q154R, Y147R+Q154R+Y123H, Y147R+Q154R+I76Y, Y147R+Q154R+T166R, Y123H+Y147R+Q154R+I76Y, V82S+Y123H+Y147R+Q154R, and I76Y+V82S+Y123H+Y147R+Q154R, or the corresponding mutations in another TadA.

[0350] In some embodiments, TadA*8 is truncated. In some embodiments, the truncated TadA*8 lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to full-length TadA*8. In some embodiments, the truncated TadA*8 lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to full-length TadA*8. In some embodiments, the variant of adenosine deaminase is full-length TadA*8.

[0351] In one embodiment, a fusion protein of the invention comprises wild-type TadA, linked to a variant of adenosine deaminase described herein (e.g., TadA*8), and linked to a Cas9 nickase. In certain embodiments, the fusion protein comprises a single TadA*8 domain (e.g., provided as a monomer). In other embodiments, the base editor comprises TadA*8 and TadA(wt), which can form a heterodimer.

[0352] In certain embodiments, the fusion protein comprises a single (e.g., provided as a monomer) TadA*8. In some embodiments, TadA*8 is linked to a Cas9 nickase. In some embodiments, the fusion protein of the invention comprises wild-type TadA (TadA(wt)) linked to TadA*8 as a heterodimer. In other embodiments, the fusion protein of the invention comprises TadA*7.10 linked to TadA*8 as a heterodimer. In some embodiments, the base editor is ABE8 comprising a TadA*8 variant monomer. In some embodiments, the base editor is ABE8 comprising a heterodimer of TadA*8 and TadA(wt). In some embodiments, the base editor is ABE8 comprising a heterodimer of TadA*8 and TadA*7.10. In some embodiments, the base editor is ABE8 comprising a heterodimer of TadA*8. In some embodiments, the TadA*8 is selected from Table 6, 12, or 13. In some embodiments, ABE8 is selected from Table 12, 13, or 15.

[0353] In some embodiments, the adenosine deaminase is a TadA*9 variant. In some embodiments, the adenosine deaminase is a TadA*9 variant, selected from the variants described below and referring to the following sequence (referred to as TadA*7.10): [ka]

[0354] In some embodiments, the adenosine deaminase comprises one or more of the following modifications: R21N, R23H, E25F, N38G, L51W, P54C, M70V, Q71M, N72K, Y73S, V82T, M94V, P124W, T133K, D139L, D139M, C146R, and A158K. The one or more modifications are shown in the sequence above in underlined and bold.

[0355] In some embodiments, the adenosine deaminase comprises one or more of the following combinations of modifications: V82S+Q154R+Y147R, V82S+Q154R+Y123H, V82S+Q154R+Y147R+Y123H, Q154R+Y147R+Y123H+I76Y+V82S, V82S+I76Y, V82S+Y147R, V82S+Y147R+Y123H, V82S+Q154R+Y123H, Q154R+Y147R+Y 123H+I76Y, V82S+Y147R, V82S+Y147R+Y123H, V82S+Q154R+Y123H, V82S+Q154R+Y147R, V82S+Q154R+Y147R, Q154R+Y147R+Y123H+I76Y, Q154R+Y147R+Y123H+I76Y+V82S, I76Y_V82S_Y123H_Y147R_Q154R, Y147R+Q154R+H123H, and V82S+Q154R.

[0356] In some embodiments, the adenosine deaminase comprises one or more of the following combinations of modifications: E25F+V82S+Y123H, T133K+Y147R+Q154R, E25F+V82S+Y123H+Y147R+Q154R, L51W+V82S+Y123H+C146R+Y147R+Q154R, Y73S+V82S+Y123H+Y147R+Q154R, P54C+V82S +Y123H+Y147R+Q154R, N38G+V82T+Y123H+Y147R+Q154R, N72K+V82S+Y123H+D139L+Y147R+Q154R, E25F+V82S +Y123H+D139M+Y147R+Q154R, Q71M+V82S+Y123H+Y147R+Q154R, E25F+V82S+Y123H+T133K+Y147R+Q154R, E25 F+V82S+Y123H+Y147R+Q154R, V82S+Y123H+P124W+Y147R+Q154R, L51W+V82S+Y123H+C146R+Y147R+Q154R, P5 4C+V82S+Y123H+Y147R+Q154R, Y73S+V82S+Y123H+Y147R+Q154R, N38G+V82T+Y123H+Y147R+Q154R, R23H+V82 S+Y123H+Y147R+Q154R, R21N+V82S+Y123H+Y147R+Q154R, V82S+Y123H+Y147R+Q154R+A158K, N72K+V82S+Y123H+D139L+Y147R+Q154R, E25F+V82S+Y123H+D139M+Y147R+Q154R, and M70V+V82S+M94V+Y123H+Y147R+Q154R.

[0357] In some embodiments, the adenosine deaminase comprises one or more of the following combinations of modifications: Q71M+V82S+Y123H+Y147R+Q154R, E25F+I76Y+V82S+Y123H+Y147R+Q154R, I76Y+V82T+Y123H+Y147R+Q154R, N38G+I76Y+V82S+Y123H+Y147R+Q154R, R23H+I76Y+V82S+Y123H+Y147R+Q154R, P54C+I76Y+V82S+Y123H+Y147R+Q154R, R21N+ I76Y+V82S+Y123H+Y147R+Q154R, I76Y+V82S+Y123H+D139M+Y147R+Q154 R, Y73S+I76Y+V82S+Y123H+Y147R+Q154R, E25F+I76Y+V82S+Y123H+Y147 R+Q154R, I76Y+V82T+Y123H+Y147R+Q154R, N38G+I76Y+V82S+Y123H+Y14 7R+Q154R, R23H+I76Y+V82S+Y123H+Y147R+Q154R, P54C+I76Y+V82S+Y12 3H+Y147R+Q154R, R21N+I76Y+V82S+Y123H+Y147R+Q154R, I76Y+V82S+Y 123H+D139M+Y147R+Q154R, Y73S+I76Y+V82S+Y123H+Y147R+Q154R, and V8 2S+Q154R, N72K_V82S+Y123H+Y147R+Q154R, Q71M_V82S+Y123H+Y147R+Q 154R, V82S+Y123H+T133K+Y147R+Q154R, V82S+Y123H+T133K+Y147R+Q15 4R+A158K, M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R, N72K_V82S+Y123H+Y147R+Q154R, Q71M_V82S+Y123H+Y147R+Q154R, M70V+V82S+M94V+Y123H+Y147R+Q154R, V82S+Y123H+T133K+Y147R+Q154R, V82S+Y123H+T133K+Y147R+Q154R+A158K, and M70V+Q71M+N72K+V82S+Y123H+Y147R+Q154R. In some embodiments, the adenosine deaminase is expressed as a monomer.In other embodiments, the adenosine deaminase is expressed as a heterodimer. In some embodiments, the deaminase or other polypeptide sequence lacks methionine, for example, when included as a component of a fusion protein. This may result in a different numbering of positions. However, one skilled in the art will understand that such corresponding mutations refer to the same mutation (e.g., Y73S and Y72S, and D139M and D138M).

[0358] In some embodiments, the TadA*9 variant comprises the modifications set forth in Table 16, as described herein. In some embodiments, the TadA*9 variant is a monomer. In some embodiments, the TadA*9 variant is a heterodimer with a wild-type TadA adenosine deaminase. In some embodiments, the TadA*9 variant is a heterodimer with another TadA variant (e.g., TadA*8, TadA*9). Further details of the TadA*9 adenosine deaminase are described in International PCT Application No. 2020 / 049975, the entire contents of which are incorporated herein by reference.

[0359] Any of the mutations provided herein, and any additional mutations (e.g., based on the ecTadA amino acid sequence), can be introduced into any other adenosine deaminase. Any of the mutations provided herein can be made individually or in any combination in the TadA reference sequence or another adenosine deaminase (e.g., ecTadA).

[0360] Details of the A to G nucleobase editing protein are described in International PCT Application No. 2017 / 045381 (WO2018 / 027078) and Gaudelli, NM, et al., "Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage," Nature, 551, 464-471 (2017), the entire contents of which are incorporated herein by reference.

[0361] Editing C to T In some embodiments, a base editor disclosed herein comprises a fusion protein comprising a cytidine deaminase and can deaminate a targeted cytidine (C) base of a polynucleotide to generate a uridine (U), which has the base pairing properties of thymine. In some embodiments, for example, when the polynucleotide is double-stranded (e.g., DNA), the uridine base can then be replaced (e.g., by a cellular repair mechanism) with a thymidine base, resulting in a C:G to T:A transition. In other embodiments, deamination of a nucleic acid by a base editor from C to U cannot be accompanied by a U to T substitution.

[0362] Deamination of a targeted C in a polynucleotide to generate a U is a non-limiting example of a type of base editing that can be performed by the base editors described herein. In another example, a base editor comprising a cytidine deaminase domain can mediate the conversion of a cytosine (C) base to a guanine (G) base. For example, a U in a polynucleotide produced by deamination of a cytidine by the cytidine deaminase domain of a base editor can be removed from the polynucleotide by a base excision repair mechanism (e.g., by a uracil DNA glycosylase (UDG) domain) to generate an abasic site. The nucleobase opposite the abasic site can then be replaced with another base (e.g., C) by, for example, a translesion polymerase (e.g., by a base repair mechanism). Typically, the nucleobase opposite the abasic site is replaced with C, although other substitutions (e.g., A, G, or T) can also occur.

[0363] Thus, in some embodiments, the base editors described herein comprise a deamination domain (e.g., a cytidine deaminase domain) that can deaminate a targeted C in a polynucleotide to U. Furthermore, as described below, base editors, in some embodiments, can comprise additional domains that facilitate the conversion of U to T or U to G resulting from deamination. For example, a base editor comprising a cytidine deaminase domain can further comprise a uracil glycosylase inhibitor (UGI) domain that mediates the substitution of U with T, completing a C to T base editing event. In another example, a base editor can incorporate a translesion polymerase to improve the efficiency of C to G base editing, because the translesion polymerase can facilitate the incorporation of a C opposite the abasic site (i.e., resulting in the incorporation of a G at the abasic site, completing a C to G base editing event).

[0364] Base editors that include a cytidine deaminase domain can deaminate a target C in any polynucleotide, including DNA, RNA, and DNA-RNA hybrids. Typically, cytidine deaminase catalyzes a C nucleobase that is located in the context of a single-stranded portion of a polynucleotide. In some embodiments, the entire polynucleotide that includes the target C can be single-stranded. For example, a cytidine deaminase incorporated into a base editor can deaminate a target C in a single-stranded RNA polynucleotide. In other embodiments, a base editor that includes a cytidine deaminase domain can act on a double-stranded polynucleotide, but the target C can be located in a portion of the polynucleotide that is single-stranded during the deamination reaction. For example, in embodiments in which the NAGPB domain includes a Cas9 domain, some nucleotides can remain unpaired during formation of the Cas9-gRNA target DNA complex, resulting in the formation of an "R-loop complex" of Cas9. These unpaired nucleotides can form bubbles of single-stranded DNA and serve as substrates for single-strand-specific nucleotide deaminase enzymes (eg, cytidine deaminase).

[0365] In some embodiments, the base editor cytidine deaminase can comprise all or part of an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. APOBEC is an evolutionarily conserved family of cytidine deaminases. Members of this family are C to U editing enzymes. The N-terminal domain of APOBEC-like proteins is the catalytic domain, and the C-terminal domain is the pseudocatalytic domain. More specifically, the catalytic domain is a zinc-dependent cytidine deaminase domain and is important for cytidine deamination. APOBEC family members include APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D (now referred to as "APOBEC3E"), APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, and activation-induced (cytidine) deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of an APOBEC1 deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of an APOBEC2 deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of an APOBEC3 deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of an APOBEC3A deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of an APOBEC3B deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of an APOBEC3C deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of an APOBEC3D deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of an APOBEC3E deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of an APOBEC3F deaminase, hi some embodiments, the deaminase incorporated into the base editor comprises all or a portion of an APOBEC3G deaminase.In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of an APOBEC3H deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of an APOBEC4 deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of an activation-induced deaminase (AID). In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of cytidine deaminase 1 (CDA1). It should be understood that a base editor can comprise a deaminase from any suitable organism (e.g., human or rat). In some embodiments, the deaminase domain of the base editor is from a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase domain of the base editor is from a rat (e.g., rat APOBEC1). In some embodiments, the deaminase domain of the base editor is human APOBEC1. In some embodiments, the deaminase domain of the base editor is pmCDA1.

[0366] Other exemplary deaminases that can be fused to Cas9 according to aspects of the present disclosure are provided below. In embodiments, the deaminase is activation-induced deaminase (AID). It should be understood that in some embodiments, the active domain of each sequence can be used, for example, a domain without a localization signal (nuclear localization sequence, no nuclear transport signal, cytoplasmic localization signal).

[0367] Some aspects of the present disclosure are based on the recognition that modulating the catalytic activity of the deaminase domain of any of the fusion proteins described herein, for example, by creating point mutations in the deaminase domain, affects the processivity of the fusion protein (e.g., a base editor). For example, mutations that reduce, but do not eliminate, the catalytic activity of the deaminase domain within a base-editing fusion protein can reduce the likelihood that the deaminase domain will catalyze the deamination of residues adjacent to a target residue, thereby narrowing the deamination window. The ability to narrow the deamination window can prevent undesired deamination of residues adjacent to a particular target residue, reducing or preventing off-target effects.

[0368] For example, in some embodiments, an APOBEC deaminase incorporated into a base editor can comprise one or more mutations selected from the group consisting of H121X, H122X, R126X, R126X, R118X, W90X, W90X, and R132X (where X is any amino acid) of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, an APOBEC deaminase incorporated into a base editor can comprise one or more mutations selected from the group consisting of H121R, H122R, R126A, R126E, R118A, W90A, W90Y, and R132E of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase.

[0369] In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise one or more mutations selected from the group consisting of D316X, D317X, R320X, R320X, R313X, W285X, W285X, R326X (where X is any amino acid) of hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase that includes one or more mutations selected from the group consisting of D316R, D317R, R320A, R320E, R313A, W285A, W285Y, R326E of hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase.

[0370] In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise the H121R and H122R mutations of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises the R126A mutation of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises the R126E mutation of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises the R118A mutation of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises a W90A mutation in rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises a W90Y mutation in rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises a R132E mutation in rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises a W90Y and R126E mutations in rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase.In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises the R126E and R132E mutations of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises the W90Y and R132E mutations of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises the W90Y, R126E, and R132E mutations of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase.

[0371] In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises the D316R and D317R mutations in hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any of the fusion proteins provided herein comprises an APOBEC deaminase that comprises the R320A mutation in hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises the R320E mutation in hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises the R313A mutation in hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises a W285A mutation in hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises a W285Y mutation in hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises a R326E mutation in hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises a W285Y and R320E mutations in hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase.In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises the R320E and R326E mutations in hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises the W285Y and R326E mutations in hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise an APOBEC deaminase that comprises the W285Y, R320E, and R326E mutations in hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase.

[0372] Several modified cytidine deaminases are commercially available, including, but not limited to, SaBE3, SaKKH-BE3, VQR-BE3, EQR-BE3, VRER-BE3, YE1-BE3, EE-BE3, YE2-BE3, and YEE-BE3, which are available from Addgene (plasmids 85169, 85170, 85171, 85172, 85173, 85174, 85175, 85176, 85177). In some embodiments, the deaminase incorporated into the base editor comprises all or a portion of an APOBEC1 deaminase.

[0373] In some embodiments, the fusion proteins of the present invention comprise one or more cytidine deaminase domains. In some embodiments, the cytidine deaminases provided herein are capable of deaminating cytosine or 5-methylcytosine to uracil or thymine. In some embodiments, the cytidine deaminases provided herein are capable of deaminating cytosine in DNA. The cytidine deaminase can be derived from a...