Compositions and methods for treating hepatitis b

JP2025102768A5Pending Publication Date: 2025-08-13BEAM THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025033734
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-10-29
Filing Date
2025-03-04
Publication Date
2025-08-13

AI Technical Summary

Technical Problem

Current therapeutic approaches for Hepatitis B virus (HBV) infection are limited and costly, with antiviral drugs like tenofovir being expensive and liver transplants being risky and costly, necessitating the need for improved treatment methods.

Method used

A method involving modifications to the HBV genome using base editing systems, such as programmable DNA-binding proteins, to introduce premature stop codons or missense mutations in HBV genes, specifically targeting nucleobases in the HBV genome through compositions comprising fusion proteins with polypeptides, base editors, and guide RNAs to alter nucleotide sequences, thereby disrupting essential HBV protein function.

Benefits of technology

This approach effectively disrupts HBV protein function, potentially reducing viral replication and disease progression, offering a more affordable and less invasive treatment option than existing therapies.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

To provide a method for editing a nucleobase of a hepatitis B virus (HBV) genome.SOLUTION: The method comprises contacting the HBV genome with one more guide RNAs and a base editor comprising a polynucleotide programmable DNA domain and an adenosine deaminase or cytidine deaminase domain, where the guide RNA targets the base editor to effect an alteration of the nucleobase of the HBV genome.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This international PCT application is a continuation of U.S. Provisional Application No. 62 / 846,422, filed May 10, 2019, and U.S. Provisional Application No. 62 / 846,422, filed May 10, 2019. This application claims priority to and the benefit of U.S. Provisional Application No. 62 / 927,585, filed October 29, 2007, each of which is hereby incorporated by reference. The contents of which are incorporated herein by reference in their entirety. [Background technology]

[0002] Hepatitis B is a serious liver infection caused by the hepatitis B virus (HBV). BV is a small DNA hepadnavirus that replicates through an RNA intermediate and integrates into the host genome. The virus can survive inside infected cells by being absorbed into the body. Approximately 257 million people worldwide are chronically infected with HBV. Chronic HBV infection is a chronic hepatitis HBV infection manifests as hepatocellular carcinoma, cirrhosis, and / or hepatocellular carcinoma. 20%-30% of adults with chronic HBV infection develop hepatocellular carcinoma. HBV infection causes 600,000 to 1 million deaths annually. There are.

[0003] Current therapeutic approaches to HBV infection have severe limitations. Antiviral drugs such as the enzyme inhibitor tenofovir can reduce viral replication. These antiviral therapies cost patients up to 500,000 yen per month. Depending on the extent of liver damage caused by HBV, a transplant may be necessary. In addition to the risks inherent in organ transplantation, the costs can be prohibitive. Therefore, there is an urgent need for improved methods for treating HBV infection. Summary of the Invention

[0004] As described below, the present invention provides a method for the prevention and treatment of hepatitis B virus by introducing modifications into the HBV genome. In certain embodiments, the present invention features compositions and methods for treating Hepatitis B virus (HBV) infection, comprising: The present invention relates to a method for modifying the HBV genome to eliminate a premature stop codon or a coding sequence of HBV or a coding sequence of HBV. Deamidation of nucleobases in covalently closed circular DNA (cccDNA) Base editing systems (e.g., programmable DNA-binding proteins) to introduce changes such as cleavage or oxidation The present invention provides a fusion protein comprising a polypeptide, a base editor, and a gRNA.

[0005] Provided herein are methods and compositions for editing the Hepatitis B (HBV) genome, as well as related In one aspect, the present invention provides a method for treating hepatitis B virus ( The present invention provides a method for editing nucleobases of an HBV genome, the method comprising: and polynucleotide programmable DNA binding domains and adenosine deamins. contacting the nucleic acid with a base editor comprising a cytidine deaminase or cytidine deaminase domain; The guide RNA targets the base editor to modify the nucleobase of the HBV genome. In some embodiments, the nucleobases of the HBV genome are polynucleotides encoding HBV proteins. In some embodiments, the contacting is with a eukaryotic cell, a mammalian cell, or a human cell. In some embodiments, the contacting is performed in a human cell. In some embodiments, the cytidine deaminase converts a target C to a U in the HBV genome. Convert. In some embodiments, cytidine deaminase converts target C·G in a polynucleotide encoding an HBV protein to T·A. In some embodiments, adenine deaminase converts target A·T in a polynucleotide encoding an HBV protein to G·C. In some embodiments, adenine deaminase converts target A·T in a polynucleotide encoding an HBV protein to G·C. Convert.

[0006] In some embodiments of the above method, the modification of the nucleobases in the polynucleotide encoding the HBV protein results in a premature stop codon. In some embodiments, the modification of the nucleobases results in R87* or W120* termination in the HBV X protein. In some embodiments, the modification of the nucleobases results in W35* or W36* in the HBV S protein. In some embodiments, the modification of the HBV polynucleotide is a missense mutation. In some embodiments, the missense mutation is within the HBV pol gene. In some embodiments, the missense mutation results in E24G, L25F, P26F, R27C, V48A, V48I, S382F, V378I, V378A, V379I, V379A, L377F, D380G, D380N, F381P, R376G, A422T, F423P, A432V, M433V, P434S, D540G, A688V, D689G, A717T, E718K, P713S, P713L, or L719P in the HBV polymerase protein encoded by the HBV pol gene. In some embodiments, the missense mutation is within the HBV core gene. In some embodiments, the missense mutation results in T160A, T160A, P161F, S162L, C183R, or *184Q in the HBV core protein encoded by the HBV core gene. In some embodiments, the modification of the nucleobases in the HBV genome within the polynucleotide encoding the HBV protein results in a premature stop codon. In some embodiments, the modification of the nucleobases results in R87* or W120* termination in the HBV X protein. In some embodiments, the modification of the nucleobases results in W35* or W36* in the HBV S protein. In some embodiments, the modification of the HBV polynucleotide is a missense mutation. In some embodiments, the missense mutation is within the HBV pol gene. In some embodiments, the missense mutation results in E24G, L25F, P26F, R27C, V48A, V48I, S382F, V378I, V378A, V379I, V379A, L377F, D380G, D380N, F381P, R376G, A422T, F423P, A432V, M433V, P434S, D540G, A688V, D689G, A717T, E718K, P713S, P713L, or L719P in the HBV polymerase protein encoded by the HBV pol gene. In some embodiments, the missense mutation is within the HBV core gene. In some embodiments, the missense mutation results in T160A, T160A, P161F, S162L, C183R, or *184Q in the HBV core protein encoded by the HBV core gene. In some embodiments, the modification of the nucleobases results in R87* or W120* termination in the HBV X protein. In some embodiments, the modification of the nucleobases results in W35* or W36* in the HBV S protein. In some embodiments, the modification of the HBV polynucleotide is a missense mutation. In some embodiments, the missense mutation is within the HBV pol gene. In some embodiments, the missense mutation is within the HBV pol gene. In some embodiments, the missense mutation results in E24G, L25F, P26F, R27C, V48A, V48I, S382F, V378I, V378A, V379I, V379A, L377F, D380G, D380N, F381P, R376G, A422T, F423P, A432V, M433V, P434S, D540G, A688V, D689G, A717T, E718K, P713S, P713L, or L719P in the HBV polymerase protein encoded by the HBV pol gene. In some embodiments, the missense mutation results in E24G, L25F, P26F, R27C, V48A, V48I, S382F, V378I, V378A, V379I, V379A, L377F, D380G, D380N, F381P, R376G, A422T, F423P, A432V, M433V, P434S, D540G, A688V, D689G, A717T, E718K, P713S, P713L, or L719P in the HBV polymerase protein encoded by the HBV pol gene. In some embodiments, the missense mutation results in E24G, L25F, P26F, R27C, V48A, V48I, S382F, V378I, V378A, V379I, V379A, L377F, D380G, D380N, F381P, R376G, A422T, F423P, A432V, M433V, P434S, D540G, A688V, D689G, A717T, E718K, P713S, P713L, or L719P in the HBV polymerase protein encoded by the HBV pol gene. In some embodiments, the missense mutation is within the HBV core gene. In some embodiments, the missense mutation is within the HBV core gene. In some embodiments, the missense mutation results in T160A, T160A, P161F, S162L, C183R, or *184Q in the HBV core protein encoded by the HBV core gene. In an embodiment, the missense mutation is within the HBV X gene. In some embodiments, the missense mutation results in H86R, W120R, E122K, E121K, or L141P in the HBV X protein encoded by the HBV X gene. In certain embodiments, the missense mutation results in S38F, L39F, W35R, W3 6R, T37I, T37A, R78Q, S34L, F80P, or D33G in the HBV S protein encoded by the HBV S gene.

[0007] In some embodiments of the above method, the polynucleotide programmable DNA binding domain provided herein is Streptococcus pyogenes Cas9 (SpCas9), Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), Steptococcus canis Cas9 (ScCas9), or a Cas9 selected from variants thereof. In some embodiments, Cas9 has protospacer adjacent motif (PAM) specificity for a nucleic acid sequence selected from 5'-NGG-3', 5'-NAG-3', 5'-NGA-3', 5'-NAA-3', 5'-NNAG GA-3', or 5'-NNACCA-3'. In some embodiments, the polynucleotide programmable DNA binding domain includes a modified Cas9 having a modified protospacer adjacent motif (PAM) specificity. In some embodiments, the modified PAM is 5'-NNNRRT-3' , NGA-3', 5'-NGCG-3', 5'-NGN-3', NGCN-3', 5'-NGTN-3', or 5'-NAA-3' ​​selected from. In some embodiments, the polynucleotide-programmable DNA binding domain is a nuclease-inactive variant or a nickase variant. In some embodiments, the nuclease-inactive variant or nickase variant is an amino acid substitution D10A or a nuclease-inactivated Cas9 (dCas9) comprising the corresponding amino acid substitution.

[0008] In some embodiments of the above method, the adenosine deaminase domain can deaminate adenine in deoxyribonucleic acid (DNA). In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, TadA de aminase is TadA*7.10, TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA* 8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, T adA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA *8.21, TadA*8.22, TadA*8.23, or TadA*8.24. In some embodiments, the cyt idine deaminase domain can deaminate cytosine in DNA. In some embodiments, the cytidine deaminase is APOBEC or a derivative thereof. In some embodiments, the base editor further comprises a uracil glycosylase inhibitor (UGI). In some embodiments, the base editor does not include a uracil glycos ylase inhibitor (UGI).

[0009] In some embodiments of the above method, one or more guide RNAs for editing nucleobases in the HBV genome include CRISPR RNA (crRNA) and trans-encoded small RNA (tracrRNA), wherein the crRNA contains a nucleic acid sequence complementary to the HBV nucleic acid sequence. In some embodiments, the base editor forms a complex with a single guide RNA (sgRNA)

[0010] containing a nucleic acid sequence complementary to the HBV nucleic acid sequence. In some embodiments, the HBV protein is the HBV S, polymerase (pol), core, or X protein. In some aspects, the above method for editing nucleobases of the hepatitis B virus (HBV) genome involves editing one or more nucleobases. In some embodiments, the method includes two or more guide RNAs targeting two or more HBV nucleic acid sequences. In some embodiments, the guide RNAs areAUGUUGCCCA;GACUUCUCUCAAUUUUCUAG;GUUCCGCAGUAUGGAUCGGC;UACUAACAUUGAGGUUCCCG;UC CGCAGUAUGGAUCGGCAG;UCCUCUGCCGAUCCAUACUG;GUAGCUCCAAAUUCUUUAUA; or AAUCCACACU A 5’-to-3’ sequence selected from one or more of CCGAAAGACA, or a 1-, 2-, 3-, 4-, or 5-nucleotide 5’ truncation fragment thereof.

[0011] In another aspect, there is provided a method of treating hepatitis B virus (HBV) infection in a subject, the method comprising administering to a subject in need thereof a fusion protein or a polynucleotide encoding said fusion protein, the fusion protein comprising a polynucleotide-programmable DNA-binding domain and a base editor domain that is an adenosine deaminase or cytidine deaminase domain and one or more guide polynucleotides that target the base editor domain to effect a change in the HBV polypeptide encoding nucleic acid sequence from A·T to G·C, C·G to T·A, or C·G to U·A.

[0012] In another aspect, there is provided a method of treating hepatitis B virus (HBV) infection in a subject, the method comprising administering to a subject in need thereof one or more polynucleotides encoding a polynucleotide-programmable DNA-binding domain and a base editor domain that is an adenosine deaminase or cytidine deaminase domain and one or more guide polynucleotides that target the base editor domain to effect a change in the HBV polypeptide encoding nucleic acid sequence from A·T to G·C, C·G to T·A, or C·G to U· A. ​​Administering one or more guide polynucleotides that cause a modification to A.

[0013] In some embodiments of the above treatment methods, the subject is a mammal or a human. In some embodiments, the method includes delivering a fusion protein, a polynucleotide encoding the fusion protein, or one or more polynucleotides encoding a polynucleotide-programmable DNA binding domain and a base editor, as well as the one or more guide polynucleotides, to the cells of the subject. In some embodiments, the cells are hepatocytes. In some embodiments, the polynucleotide-programmable DNA binding domain is a Cas9 selected from Streptococcus pyogenes Cas9 (SpCas9), Staphylococcus aureus Cas9 (SaCas9), S treptococcus thermophilus 1 Cas9 (St1Cas9), Steptococcus canis Cas9 (ScCas9), or a variant thereof. In some embodiments, Cas9 has protospacer adjacent motif (PAM) specificity for a nucleic acid sequence selected from 5'-NGG-3', 5'-NAG-3', 5'-NGA-3', 5'-NAA-3', 5'-NNAGGA-3', or 5'-NNACCA -3'. In some embodiments, the polynucleotide-programmable DNA binding domain includes a modified Cas9 having a modified protospacer adjacent motif (PAM) specificity. In some embodiments, the nucleic acid sequence of the modified PAM is selected from 5'-NNNRRT-3', NGA-3', 5 '-NGCG-3', 5'-NGN-3', NGCN-3', 5'-NGTN-3', or 5'-NAA-3' . ​ In some embodiments, the polynucleotide-programmable DNA binding domain is a nuclease-inactive variant or a nickase variant. In some embodiments the nuclease-inactive variant or nickase variant is nuclease-inactivated Cas9 (dCas9) containing the amino acid substitution D10A or a corresponding amino acid substitution. In some embodiments the adenosine deaminase domain can deaminate adenosine in deoxyribonucleic acid (DNA). In some embodiments the adenosine deaminase is TadA deaminase. In some embodiments the TadA deaminase is TadA*7.10, TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, T adA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8 .15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22 .23, or TadA*8.24. In some embodiments, the cytidine deaminase domain can deaminate cytidine in DNA. In some embodiments the cytidine deaminase is APOBEC or a derivative thereof. In some embodiments the base editor further comprises one or more uracil glycosylase inhibitors (UGIs). In some embodiments, the base editor does not contain a uracil glycosylase inhibitor (UGI). In some embodiments, one or more guide RNAs are CRISPR RNA (crRNA) and trans-encoded RNA (tracrRNA). It contains the generated small interfering RNA (tracrRNA), and the crRNA contains a nucleic acid sequence complementary to the HBV nucleic acid sequence. In some embodiments, the base editor forms a complex with a single guide RNA (sgRNA) that contains a nucleic acid sequence complementary to the HBV nucleic acid sequence. In some embodiments, the sgRNA contains a nucleic acid sequence that contains at least 10 consecutive nucleotides complementary to the HBV nucleic acid sequence. In some embodiments, the sgRNA contains a nucleic acid sequence that contains 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 consecutive

[0014] nucleotides complementary to the HBV nucleic acid sequence. In some embodiments, the above method includes editing one or more nucleic acid bases. In some embodiments, the above method includes two or more guide RNAs that target two or more HBV nucleic acid sequences. In some embodiments, the above method includes two or more guide RNAs that target 3, 4, or 5 HBV nucleic acid sequences. In some embodiments of the above method, the HBV nucleic acid sequence encodes one or more HBV proteins selected from HBV polymerase, HBV core protein, HBV S protein, HBV X protein, or combinations thereof. In some embodiments of the above method, one or more guide RNAs are UCAAUCCCAACAAGGACACC; GG GAACAAGAUCUACAGCAU; AAGCCCAGGAUGAUGGGAUG; CUGCCAACUGGAUCCUGCGC; GACACAUCCAGCGAUA ACCA; GCUGCCAACUGGAUCCUGCG; UAUGGAUGAUGUGGUAUUGG; CCAUGCCCCAAAGCCACCCA; AAGCCACC ACCA; GCUGCCAACUGGAUCCUGCG; UAUGGAUGAUGUGGUAUUGG; CCAUGCCCCAAAGCCACCCA; AAGCCACC CAAGGCACAGCU;GAGAAGUCCACCACGAGUCU;CUUCUCUCAAUUUUCUAGGG;GACGACGAGGCAGGUCCCCU; CCCAACAAGGACACCUGGCC;UGCCAACUGGAUCCUGCGCG;AGGAGUUCCGCAGUAUGGAU;CCGCAGUAUGGAUC GGCAGA;CCUCUGCCGAUCCAUACUGC;CGCCCACCGAAUGUUGCCCA;GACUUCUCUCAAUUUUCUAG;GUUCCG CAGUAUGGAUCGGC;UACUAACAUUGAGGUUCCCG;UCCGCAGUAUGGAUCGGCAG;UCCUCUGCCGAUCCAUACUG ;GUAGCUCCAAAUUCUUUAUA; or a sequence from 5’ to 3’ selected from one or more of AAUCCACACUCCGAAAGACA, or a 5’ truncation fragment of 1, 2, 3, 4, or 5 nucleotides thereof is included. In some embodiments, the modification of the polynucleotide encoding the HBV protein is a premature stop codon. In some embodiments, the modification of the nucleic acid sequence results in R87* or W120* in the HBV X protein encoded by the nucleic acid. In some embodiments, the modification of the nucleic acid sequence results in W35* or W36* in the HBV S protein encoded by the nucleic acid. In some embodiments, the modification of the polynucleotide encoding the HBV protein is a missense mutation. In some embodiments, the missense mutation is within the HBV pol gene. In some embodiments, the missense mutation within the HBV pol gene is E24G, L25F, P26F, R27C, V4 in the HBV polymerase encoded by the HBV pol gene is included. In some embodiments, the modification of the polynucleotide encoding the HBV protein is a premature stop codon. In some embodiments, the modification of the nucleic acid sequence results in R87* or W120* in the HBV X protein encoded by the nucleic acid. In some embodiments, the modification of the nucleic acid sequence results in W35* or W36* in the HBV S protein encoded by the nucleic acid. In some embodiments, the modification of the polynucleotide encoding the HBV protein is a missense mutation. In some embodiments, the missense mutation is within the HBV pol gene. In some embodiments, the missense mutation within the HBV pol gene is E24G, L25F, P26F, R27C, V4 in the HBV polymerase encoded by the HBV pol gene is included. In some embodiments, the modification of the polynucleotide encoding the HBV protein is a premature stop codon. In some embodiments, the modification of the nucleic acid sequence results in R87* or W120* in the HBV X protein encoded by the nucleic acid. In some embodiments, the modification of the nucleic acid sequence results in W35* or W36* in the HBV S protein encoded by the nucleic acid. In some embodiments, the modification of the polynucleotide encoding the HBV protein is a missense mutation. In some embodiments, the missense mutation is within the HBV pol gene. In some embodiments, the missense mutation within the HBV pol gene is E24G, L25F, P26F, R27C, V4 in the HBV polymerase encoded by the HBV pol gene is included. In some embodiments, the modification of the polynucleotide encoding the HBV protein is a premature stop codon. In some embodiments, the modification of the nucleic acid sequence results in R87* or W120* in the HBV X protein encoded by the nucleic acid. In some embodiments, the modification of the nucleic acid sequence results in W35* or W36* in the HBV S protein encoded by the nucleic acid. In some embodiments, the modification of the polynucleotide encoding the HBV protein is a missense mutation. In some embodiments, the missense mutation is within the HBV pol gene. In some embodiments, the missense mutation within the HBV pol gene is E24G, L25F, P26F, R27C, V4 in the HBV polymerase encoded by the HBV pol gene is included. In some embodiments, the modification of the polynucleotide encoding the HBV protein is a premature stop codon. In some embodiments, the modification of the nucleic acid sequence results in R87* or W120* in the HBV X protein encoded by the nucleic acid. In some embodiments, the modification of the nucleic acid sequence results in W35* or W36* in the HBV S protein encoded by the nucleic acid. In some embodiments, the modification of the polynucleotide encoding the HBV protein is a missense mutation. In some embodiments, the missense mutation is within the HBV pol gene. In some embodiments, the missense mutation within the HBV pol gene is E24G, L25F, P26F, R27C, V4 in the HBV polymerase encoded by the HBV pol gene is included. In some embodiments, the modification of the polynucleotide encoding the HBV protein is a premature stop codon. In some embodiments, the modification of the nucleic acid sequence results in R87* or W120* in the HBV X protein encoded by the nucleic acid. In some embodiments, the modification of the nucleic acid sequence results in W35* or W36* in the HBV S protein encoded by the nucleic acid. In some embodiments, the modification of the polynucleotide encoding the HBV protein is a missense mutation. In some embodiments, the missense mutation is within the HBV pol gene. In some embodiments, the missense mutation within the HBV pol gene is E24G, L25F, P26F, R27C, V4 in the HBV polymerase encoded by the HBV pol gene is included. In some embodiments, the modification of the polynucleotide encoding the HBV protein is a premature stop codon. In some embodiments, the modification of the nucleic acid sequence results in R87* or W120* in the HBV X protein encoded by the nucleic acid. In some embodiments, the modification of the nucleic acid sequence results in W35* or W36* in the HBV S protein encoded by the nucleic acid. In some embodiments, the modification of the polynucleotide encoding the HBV protein is a missense mutation. In some embodiments, the missense mutation is within the HBV pol gene. In some embodiments, the missense mutation within the HBV pol gene is E24G, L25F, P26F, R27C, V4 in the HBV polymerase encoded by the HBV pol gene is included. In some embodiments, the modification of the polynucleotide encoding the HBV protein is a premature stop codon. In some embodiments, the modification of the nucleic acid sequence results in R87* or W120* in the HBV X protein encoded by the nucleic acid. In some embodiments, the modification of the nucleic acid sequence results in W35* or W36* in the HBV S protein encoded by the nucleic acid. In some embodiments, the modification of the polynucleotide encoding the HBV protein is a missense mutation. In some embodiments, the missense mutation is within the HBV pol gene. In some embodiments, the missense mutation within the HBV pol gene is E24G, L25F, P26F, R27C, V4 in the HBV polymerase encoded by the HBV pol gene is included. In some embodiments, the modification of the polynucleotide encoding the HBV protein is a premature stop codon. In some embodiments, the modification of the nucleic acid sequence results in R87* or W120* in the HBV X protein encoded by the nucleic acid. In some embodiments, the modification of the nucleic acid sequence results in W35* or W36* in the HBV S protein encoded by the nucleic acid. In some embodiments, the modification of the polynucleotide encoding the HBV protein is a missense mutation. In some embodiments, the missense mutation is within the HBV pol gene. In some embodiments, the missense mutation within the HBV pol gene is E24G, L25F, P26F, R27C, V4 in the HBV polymerase encoded by the HBV pol gene 8A, V48I, S382F, V378I, V378A, V379I, V379A, L377F, D380G, D380N, F381P, R376G, A422T, F423P, A432V, M433V, P434S, D540G, A688V, D689G, A717T, E718K, P713S, P71 3L, or results in L719P. In some embodiments, the missense mutation is within the HBV core gene. In some embodiments, the missense mutation within the HBV core gene results in T160A, T160A, P161F, S162L, C183R, or *184Q in the HBV core protein encoded by the HBV core gene. In some embodiments, the missense mutation is within the HBV X gene. In some embodiments, the missense mutation results in H86R, W120R, E122K, E121K, or L141P in the HBV X protein encoded by the HBV X gene. In some embodiments, the missense mutation is within the HBV S gene. In some embodiments, the missense mutation results in S38F, L39F, W35R, W36R, T37I, T37A, R78Q, S34L, F80P, or D33G in the HBV S protein encoded by the HBV S gene. In some embodiments, the base editor is BE4 or a variant of BE4, APOBEC-1 is replaced with the sequence of APOBEC-3A, and / or Cas9 is replaced with a Cas9 variant (referred to as SpCas9-VRQR) comprising V1134, R1217, Q1334, and R1336.

[0015] In one aspect, a composition is provided, for example, for treating HBV infection. In one Provided is a composition, which comprises a base editor bound to a guide RNA, and the guide RNA comprises a nucleic acid sequence complementary to the HBV gene. In one embodiment, the base editor is adenosine deaminase or cytidine deaminase. In one embodiment, the adenosine deaminase can deaminate adenine in deoxyribonucleic acid (DNA). In one embodiment, the adenosine deaminase is TadA deaminase. In one embodiment, the TadA deamin ase is TadA*7.10, TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, Tad A*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14 , TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, T adA*8.22, TadA*8.23, or TadA*8.24. In one embodiment, the cytidine deaminase domain can deaminate cytidine in DNA. In one embodiment, the cytidi ne deaminase is APOBEC or a derivative thereof. In one embodiment, the base editor further comprises one or more uracil glycosylase inhibitors (UGI). In one embodiment, the base editor does not comprise a uracil glycosylase inhibitor (UGI). In one embodiment, the base editing body: (i) comprises a Cas9 nickase; (ii) comprises a nuclease-inactive Cas9; (iii) does not comprise a UGI domain; (iv) comprises APOBEC-1 or APOBEC-3A cytidine deaminase; (v) MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNT RCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYS PSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLKSGGSSGGSSGS ETPGTSESATPESSGGSSGGSDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAE ATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRK KLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSR RLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAI LLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPIL EKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNS RFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFL SGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIV LTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDD SLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRER MKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRS DKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMN TKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRK MIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEV QTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPID FLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTK At least 80%, 85%, 9 for EVLDATLIHQSITGLYETRIDLSQLGGDSGGSKRTADGSEFESPKKKRKVE comprises an amino acid sequence having an identity of 0%, 95%, 96%, 97%, 98%, 99%, or 100%; or (vi) MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNT RCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYS PSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLKSGGSSGGSSGS ETPGTSESATPESSGGSSGGSDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAE ATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRK KLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSR RLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAI LLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPIL EKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNS RFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFL SGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIV LTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDD SLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRER MKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRS DKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMN TKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRK MIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEV QTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPID FLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTK comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to EVLDATLIHQSITGLYETRIDLSQLGGDSGGSKRTADGSEFESPKKKRKVE.

[0016] In one embodiment, the guide RNA of the composition comprises a nucleic acid sequence complementary to an HBV gene encoding an HBV polymerase, an HBV core protein, an HBV S protein, or an HBV X protein. In one embodiment, the guide RNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 consecutive nucleotides that are completely complementary to the HBV gene encoding the HBV X protein. In one embodiment, the guide RNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 consecutive nucleotides that are completely complementary to the HBV gene encoding the HBV S protein. In one embodiment, the guide RNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 consecutive nucleotides that are completely complementary to the HBV gene encoding the HBV polymerase. In one embodiment, the guide RNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 consecutive nucleotides that are completely complementary to the HBV gene encoding the HBV core protein. In one embodiment, the guide RNA is, from 5' to 3', UCAAUCCCAACAAGGACACC; GGGAACAAGAUCUACAGCAU; AAGCCCAGGAUGAUGGGAUG;CUGCCAACUGGAUCCUGCGC; GACACAUCCAGCGAUAACCA;GCUGCCAACUGGAUCCUGCG; UAUGGAUGAUGUGGUAUUGG;CCAUGCCCCAAAGCCACCCA; AAGCCACCCAAGGCACAGCU;GAGAAGUCCACCACGAGUCU; CUUCUCUCAAUUUUCUAGGG;GACGACGAGGCAGGUCCCCU; CCCAACAAGGACACCUGGCC;UGCCAACUGGAUCCUGCGCG; AGGAGUUCCGCAGUAUGGAU;CCGCAGUAUGGAUCGGCAGA; CCUCUGCCGAUCCAUACUGC;CGCCCACCGAAUGUUGCCCA; GACUUCUCUCAAUUUUCUAG;GUUCCGCAGUAUGGAUCGGC; UACUAACAUUGAGGUUCCCG;UCCGCAGUAUGGAUCGGCAG; UCCUCUGCCGAUCCAUACUG;GUAGCUCCAAAUUCUUUAUA, and selected from the group consisting of 1, 2, 3, 4, or 5 nucleic acid truncations from the 5' end of the nucleic acid. In one embodiment, the guide RNA is from 5' to 3' UCAAUCCCAACAAGGACACC;GGGAACAAGAUCUACAGCAU; AAGCCCAGGAUGAUGGGAUG;CUGCCAACUGGAUCCUGCGC; GACACAUCCAGCGAUAACCA;GCUGCCAACUGGAUCCUGCG; UAUGGAUGAUGUGGUAUUGG;CCAUGCCCCAAAGCCACCCA; AAGCCACCCAAGGCACAGCU;GAGAAGUCCACCACGAGUCU; CUUCUCUCAAUUUUCUAGGG;GACGACGAGGCAGGUCCCCU; CCCAACAAGGACACCUGGCC;UGCCAACUGGAUCCUGCGCG; AGGAGUUCCGCAGUAUGGAU;CCGCAGUAUGGAUCGGCAGA; CCUCUGCCGAUCCAUACUGC;CGCCCACCGAAUGUUGCCCA; GACUUCUCUCAAUUUUCUAG;GUUCCGCAGUAUGGAUCGGC; UACUAACAUUGAGGUUCCCG;UCCGCAGUAUGGAUCGGCAG; UCCUCUGCCGAUCCAUACUG;GUAGCUCCAAAUUCUUUAUA, and selected from the group consisting of nucleic acids selected from AAUCCACACUCCGAAAGACA. In one embodiment, the above composition further comprises a lipid. In one embodiment, the lipid is a cationic lipid. In one embodiment the composition further comprises a pharmaceutically acceptable excipient.

[0017] In another aspect, there is provided a pharmaceutical composition, the pharmaceutical composition comprising, in a pharmaceutically acceptable excipient, a base editor, or a nucleic acid encoding a base editor, and one or more guide RNAs (gRNAs) comprising a nucleic acid sequence complementary to the HBV gene In one embodiment, the base editor (i) comprises a Cas9 nickase or; (ii) comprises a nuclease-inactive Cas9 (iii) does not comprise a UGI domain (iv) comprises APOBEC-1 or APOBEC-3A cytidine deaminase (v) MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNT RCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYS PSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLKSGGSSGGSSGS ETPGTSESATPESSGGSSGGSDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAE ATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRK KLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSR RLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAI LLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPIL EKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNS RFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFL SGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIV LTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDD SLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRER MKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRS DKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMN TKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRK MIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEV QTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPID FLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTK At least 80%, 85%, 9 with respect to EVLDATLIHQSITGLYETRIDLSQLGGDSGGSKRTADGSEFESPKKKRKVE comprises an amino acid sequence having an identity of 0%, 95%, 96%, 97%, 98%, 99%, or 100%; or (vi) MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNT RCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYS PSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLKSGGSSGGSSGS ETPGTSESATPESSGGSSGGSDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAE ATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRK KLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSR RLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAI LLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPIL EKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNS RFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFL SGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIV LTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDD SLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRER MKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRS DKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMN TKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRK MIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEV QTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPID FLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTK at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to EVLDATLIHQSITGLYETRIDLSQLGGDSGGSKRTADGSEFESPKKKRKVE, and comprises an amino acid sequence. In one embodiment the base editor comprises Cas9, or a Cas9 variant (SpCas9-VRQR) comprising V1134, R1217, Q1334, and R1336. In one embodiment, the gRNA and the base editor are formulated together or separately In one embodiment, the gRNA is, from 5' to 3': UCAAUCCCAACAAGGACACC;GGGAACAAGAUCUACAGCAU AAGCCCAGGAUGAUGGGAUG;CUGCCAACUGGAUCCUGCGC GACACAUCCAGCGAUAACCA;GCUGCCAACUGGAUCCUGCG UAUGGAUGAUGUGGUAUUGG;CCAUGCCCCAAAGCCACCCA AAGCCACCCAAGGCACAGCU;GAGAAGUCCACCACGAGUCU CUUCUCUCAAUUUUCUAGGG;GACGACGAGGCAGGUCCCCU CCCAACAAGGACACCUGGCC;UGCCAACUGGAUCCUGCGCG AGGAGUUCCGCAGUAUGGAU;CCGCAGUAUGGAUCGGCAGA CCUCUGCCGAUCCAUACUGC;CGCCCACCGAAUGUUGCCCA GACUUCUCUCAAUUUUCUAG;GUUCCGCAGUAUGGAUCGGC UACUAACAUUGAGGUUCCCG;UCCGCAGUAUGGAUCGGCAG UACUAACAUUGAGGUUCCCG;UCCGCAGUAUGGAUCGGCAG UACUAACAUUGAGGUUCCCG;UCCGCAGUAUGGAUCGGCAG UCCUCUGCCGAUCCAUACUG; GUAGCUCCAAAUUCUUUAUA; or one or more of AAUCCACACUCCGAAAGACA a nucleic acid sequence selected from, or a 5' truncated fragment of 1, 2, 3, 4, or 5 nucleotides thereof is included. In one embodiment, the pharmaceutical composition further comprises a vector suitable for expression in mammalian cells and the vector comprises a polynucleotide encoding a base editor. In one embodiment the vector is a viral vector. In one embodiment, the viral vector is , a retroviral vector, an adenoviral vector, a lentiviral vector, a herpes virus vector, or an adeno-associated virus vector (AAV). In one embodiment the pharmaceutical composition further comprises a ribonucleoparticle suitable for expression in mammalian cells .

[0018] In another aspect, a method for treating HBV infection is provided, the method comprising administering to a subject in need thereof the above composition or pharmaceutical composition.

[0019] Another aspect provides an HBV genome comprising a modification selected from the group consisting of: an immature stop codon introducing R87STOP or W120STOP into the X gene. an immature stop codon introducing W35STOP or W36STOP into the S gene. into HBV polymerase: E24G, L25F, P26F, R27C, V48A, V48I, S382F, V378I, V378A, V3 79I, V379A, L377F, D380G, D380N, F381P, R376G, A422T, F423P, A432V, M433V, P434S , HBV pol into which D540G, A688V, D689G, A717T, E718K, P713S, P713L, or L719P is introduced Missense mutations within the gene; Introduce T160A, T160A, P161F, S162L, C183R, or STOP184Q into the HBV core polypeptide Missense mutations within the HBV core gene to do so; Within the X gene that introduces H86R, W120R, E122K, E121K, or L141P into the HBV X polypeptide Missense mutations of, and Within the S gene that introduces S38F, L39F, W35R, W36R, T37I, T37A, R78Q, S34L, F80P, or D33G into the HBV S polypeptide. Missense mutations within the S gene.

[0020] In one embodiment, the HBV genome comprises two or more of the above modifications.

[0021] In the above method, or in the embodiment of the above HBV genome, HBV is genotype C or genotype D.

[0022] In another aspect, there is provided the use of any of the above aspects and embodiments in the treatment of HBV infection in a subject. In another aspect, there is provided the use of any of the above aspects and embodiments in the treatment of HBV infection in a subject.

[0023] In another aspect, there is provided the use of any of the above aspects and embodiments in the treatment of HBV infection in a subject. In another aspect, there is provided the use of any of the above aspects and embodiments in the treatment of HBV infection in a subject.

[0024] In one embodiment of the above use, the subject is a mammal. In an embodiment of the above use, the subject is a human.

[0025] In one embodiment of the above method or pharmaceutical composition, one or more guide RNAs are as listed in Table 26. as follows.

[0026] In another aspect, a guide RNA (gRNA) is provided that includes a nucleic acid sequence complementary to the HBV gene. In one embodiment, the guide RNA includes a nucleic acid sequence complementary to an HBV gene encoding an HBV polymerase, an HBV core protein, an HBV S protein, or an HBV X protein. In one embodiment, the guide RNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 consecutive nucleotides that are completely complementary to the HBV gene encoding the HBV X protein. In one embodiment, the guide RNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 consecutive nucleotides that are completely complementary to the HBV gene encoding the HBV S protein. In one embodiment, the guide RNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 consecutive nucleotides that are completely complementary to the HBV gene encoding the HBV polymerase. In one embodiment, the guide RNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 consecutive nucleotides that are completely complementary to the HBV gene encoding the HBV core protein. In one embodiment, the guide RNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 consecutive nucleotides that are completely complementary to the HBV gene encoding the HBV core protein. In one embodiment, the guide RNA is, from 5' to 3', UCAAUCCCAACAAGGACACC; GGGAACAAGAU CUACAGCAU; AAGCCCAGGAUGAUGGGAUG; CUGCCAACUGGAUCCUGCGC; GACACAUCCAGCGAUAACCA; GCUGCCAACUGGAUCCUGCG; In one embodiment, the guide RNA includes 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 consecutive nucleotides that are completely complementary to the HBV gene encoding the HBV core protein. In one embodiment, the guide RNA is, from 5' to 3', UCAAUCCCAACAAGGACACC; GGGAACAAGAU CUACAGCAU; AAGCCCAGGAUGAUGGGAUG; CUGCCAACUGGAUCCUGCGC; GACACAUCCAGCGAUAACCA; GCUGCCAACUGGAUCCUGCG; UAUGGAUGAUGUGGUAUUGG;CCAUGCCCCAAAGCCACCCA; AAGCCACCCAAGGCACAGCU;GAGAAGUCCACCACGAGUCU; CUUCUCUCAAUUUUCUAGGG;GACGACGAGGCAGGUCCCCU; CCCAACAAGGACACCUGGCC;UGCCAACUGGAUCCUGCGCG; AGGAGUUCCGCAGUAUGGAU;CCGCAGUAUGGAUCGGCAGA; CCUCUGCCGAUCCAUACUGC;CGCCCACCGAAUGUUGCCCA; GACUUCUCUCAAUUUUCUAG;GUUCCGCAGUAUGGAUCGGC; UACUAACAUUGAGGUUCCCG;UCCGCAGUAUGGAUCGGCAG; UCCUCUGCCGAUCCAUACUG;GUAGCUCCAAAUUCUUUAUA, and 1, 2, 3, 4, or 5 truncations of the nucleic acid selected from the group consisting of AAUCCACACUCCGAAAGACA. In one embodiment, the guide RNA is from 5' to 3' and UCAAUCCCAACAAGGACACC;GGGAACAAGAUCUACAGCAU; AAGCCCAGGAUGAUGGGAUG;CUGCCAACUGGAUCCUGCGC; GACACAUCCAGCGAUAACCA;GCUGCCAACUGGAUCCUGCG; UAUGGAUGAUGUGGUAUUGG;CCAUGCCCCAAAGCCACCCA; AAGCCACCCAAGGCACAGCU;GAGAAGUCCACCACGAGUCU; CUUCUCUCAAUUUUCUAGGG;GACGACGAGGCAGGUCCCCU; CCCAACAAGGACACCUGGCC;UGCCAACUGGAUCCUGCGCG; AGGAGUUCCGCAGUAUGGAU;CCGCAGUAUGGAUCGGCAGA; CCUCUGCCGAUCCAUACUGC;CGCCCACCGAAUGUUGCCCA; GACUUCUCUCAAUUUUCUAG;GUUCCGCAGUAUGGAUCGGC; UACUAACAUUGAGGUUCCCG;UCCGCAGUAUGGAUCGGCAG; UCCUCUGCCGAUCCAUACUG;GUAGCUCCAAAUUCUUUAUA, and a nucleic acid selected from the group consisting of AAUCCACACUCCGAAAGACA.

[0027] In another aspect, a pharmaceutical composition is provided, the pharmaceutical composition comprising (i) a nucleic acid encoding a base editor; and (ii) a guide RNA of any of the above aspects and embodiments. In one embodiment, the pharmaceutical composition further comprises a lipid. In one embodiment of the pharmaceutical composition, the nucleic acid encoding the base editor is mRNA.

[0028] The compositions and articles defined by the present invention were isolated or otherwise manufactured in connection with the examples shown below. Other features and advantages of the present invention will become apparent from the description of the embodiments for carrying out the invention and from the claims.

[0029] Definitions The following definitions supplement those in the art and are intended for this application and, where relevant or not, for example, generally owned patents or applications not attributable to them. Methods and materials similar or equivalent to those described herein are within the scope of this application. Materials that can be used in the practice or testing of the present disclosure, as well as preferred materials and methods, are described herein. Accordingly, the terms used herein are for the purpose of describing certain embodiments only and are not intended to be limiting.

[0030] Unless otherwise defined, all technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. The following references provide one of the arts including general definitions of many of the terms used in the present invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed.1994);T he Cambridge Dictionary of Science and Technology (Walker ed., 1988);The Glossar y of Genetics,5th Ed., R. Rieger et al. (eds.),Springer Verlag (1991), and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein in the context of the present application, the following terms have the meanings ascribed to them below, unless otherwise specified.

[0031] In this application, the use of the singular includes the plural unless specifically stated otherwise. It should be noted that, as used in this specification, the singular forms "a", "an", and "the" include the plural referents unless the context clearly dictates otherwise. In The term "including", as well as other forms such as "include", "includes", and " "included" is used in a non-limiting sense.

[0032] When used in this specification and the claims (if any), the word "comprising" " (and any form of "comprise" and "comprises" such as "comprise" and "comprises") ), "having" (and any form of "have" and "has" such as "have" and "has") ), "including" (and any form of "include", "includes", and "include" such as "including") or "containing" (and any form of "contain", "contains", and "contain" such as "containing") is inclusive or non-restrictive and does not exclude additional elements or method steps not recited. Any embodiment recited herein is contemplated to be practiced with respect to any method or composition of the present disclosure, and vice versa. Further, the methods of the present disclosure can be achieved using the compositions of the present disclosure.

[0033] The term "about" or "approximately" means within an acceptable error range for a particular value as measured by one of ordinary skill in the art and depends in part on the method of measurement or determination of the value, i.e., the limitations of the measuring system. For example, "about" can mean within one or up to one It may mean a range of up to 5% or up to 1%. Alternatively, particularly with respect to biological systems or processes this term may mean, for example, on the order of within 5-fold or within 2-fold of a value. Unless otherwise stated when a particular value is recited in this application and the claims, the term "about" should be regarded as meaning being within an acceptable error range with respect to the particular value .

[0034] References in this specification to "some embodiments", "an embodiment", "one embodiment" or " other embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least some embodiments but not necessarily in all embodiments of the disclosure .

[0035] "Adenosine deaminase" means a polypeptide or fragment thereof that can catalyze the hydrolytic deamination of adenine or adenosine. In some embodiments the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine to inosine or deoxyadenosine to deoxyinosine. In some embodiments, adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided in this specification (e.g., modified adenosine deaminase, evolved adenosine deaminase) can be derived from any organism such as bacteria. In some embodiments, the deaminase or deaminase domain is from a human, chimpanzee

[0036] ​A variant of a deaminase derived from an organism such as yeast, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain does not exist in nature. For example, in some embodiments, the deaminase or deaminase domain has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity with a natural deaminase. In some embodiments the adenosine deaminase is derived from a bacterium such as E. coli, S. aureus, S. typhi, S. putrefaciens, H . influenzae, or C. crescentus. In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA de aminase is E. coli TadA (ecTadA) deaminase or a fragment thereof. For example, the deaminase domain is described in International PCT Application No. PCT / 2017 / 045381 (WO2018 / 027078)

[0037] and PCT / US2016 / 058344 (WO2017 / 070632), each of which is incorporated herein by reference in its entirety. Also, Komor, A.C., et al., ”Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” ​​Nature 533, 420 - 424 (2016); Gaudelli, N. M., et al., ”Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464 - 471 (2017) ; Komor, A. C., et al., ”Improved base excision repair inhibition and bacteriop hage Mu Gam protein yields C:G - to - T:A base editors with higher efficiency and pr oduct purity” Science Advances 3: eaa04774 (2017)), and Rees, H. A., et al., ” Base editing: precision chemistry on the genome and transcriptome of living cell s.” Nat Rev Genet. 2018 Dec; 19 (12): 770 - 788. doi: 10.1038 / s41576 - 018 - 0059 - 1 are also referred to (the entire content of which is incorporated herein by reference).

[0038] Wild - type TadA (wt) adenosine deaminase has the following sequence (also referred to as the TadA reference sequence) and has: MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQK KAQSSTD.

[0039] In some embodiments, adenosine deaminase comprises the following sequence modifications: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD (also referred to as TadA*7.10).

[0040] In some embodiments, TadA*7.10 comprises at least one modification. In some embodiments TadA*7.10 comprises modifications at amino acids 82 and / or 166. In certain embodiments the variant of the reference sequence above comprises one or more of the following modifications: Y147T, Y147R, Q1 54S, Y123H, V82S, T166R, and / or Q154R. The modification Y123H refers to the modification H123Y of TadA*7.10 being reverted to Y123H TadA (wt). In other embodiments, variants of the TadA*7.10 sequence are: Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R ; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y 147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V 82S+Y123H+Y147R+Q154R, and combinations of modifications selected from the group of I76Y+V82S+Y123H+Y147R+Q154R are included.

[0041] In other embodiments, the present invention has a deletion starting at residue 149, 150, 151, 152, 153, 154, 155, 156, or 157, compared to the corresponding mutation in TadA*7.10, the TadA reference sequence, or another TadA, e.g., a deletion including the C-terminal deletion including TadA*8. In other embodiments, the adenosine deaminase variant is a TadA (e.g., TadA *8) monomer including one or more of the following modifications: Y147T, Y147R, Q1 54S, Y123H, V82S, T166R, and / or Q154R, compared to the corresponding mutation in TadA*7.10, the T adA reference sequence, or another TadA. In other embodiments, the adenosine deaminase variant is a TadA*7.10 , the TadA reference sequence, or a TadA (e.g., TadA *8) monomer including one or more of the following modifications: Y147T+Q154R ; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y +V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q1 54R, and I76Y+V82S+Y123H+Y147R+Q154R, compared to the corresponding mutation in TadA*7.10, the TadA reference sequence, or another TadA. In yet other embodiments, the adenosine deaminase variant is a TadA*7.10, the TadA reference sequence, or another TadA, each having the following modifications: Y147T, Y1

[0042] 147R, Q154S, Y123H, V82S, T166R, and / or Q154R, compared to the corresponding mutation in TadA*7.10, the TadA reference sequence, or another TadA. In other embodiments, the adenosine deaminase variant is a TadA*7.10, the TadA reference sequence, or another TadA, each having the following modifications: Y147T, Y1 Two adenines having one or more of 47R, Q154S, Y123H, V82S, T166R, and / or Q154R is a homodimer containing an adenosine deaminase domain. In other embodiments, the adenosine de aminase variant is compared with the corresponding mutation of TadA*7.10, the TadA reference sequence, or another TadA, and each is: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y1 47R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R + Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147 R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R, and a homodimer containing two adenosine deaminase domains (e.g., TadA *8) having a combination of modifications selected from the group of I76Y + V82S + Y123H + Y147R + Q154R. In other embodiments, the adenosine deaminase variant is a heterodimer containing an adenosine deaminase domain (e.g., TadA*8) that includes one or more of the following modifications compared to the wild-type TadA adenosine deaminase domain and the corresponding mutation of TadA *7.10, the TadA reference sequence, or another TadA: Y147T, Y147R, Q154S, Y12 3H, V82S, T166R, and / or Q154R. In other embodiments, the adenosine deaminase variant is compared to the wild-type TadA adenosine deaminase domain and the corresponding mutation of TadA *7.10, the TadA reference sequence, or another TadA: Y147T + Q154R; Y147T + a heterodimer containing an adenosine deaminase domain (e.g., TadA*8) that includes one or more of the following modifications: Y147T, Y147R, Q154S, Y12 3H, V82S, T166R, and / or Q154R. In other embodiments, the adenosine deaminase variant is compared to the wild-type TadA adenosine deaminase domain and the corresponding mutation of TadA Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V 82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R + Y123H; Y147R + Q1 54R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R, and combinations of modifications selected from the group consisting of I76Y + V82S + Y123H + Y147R + Q154R A heterodimer comprising an adenosine deaminase variant domain (e.g., TadA*8).

[0043] In other embodiments, the adenosine deaminase variant is the TadA*7.10 domain, and T compared to the corresponding mutations of TadA*7.10, the TadA reference sequence, or another TadA, the following modifications, namely one or more of Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R A heterodimer comprising an adenosine deaminase variant domain (e.g., TadA*8) containing the same. In other embodiments, the adenosine deaminase variant is the TadA*7.10 domain, and compared to the corresponding mutations of TadA*7.10, the TadA reference sequence, or another TadA, the following combinations of modifications : Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82 S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q 154R; Y147R + Q154R + Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; or an adenosine containing I76Y + V82S + Y123H + Y147R + Q154R is a heterodimer containing a deaminase variant domain (e.g., TadA*8). In one embodiment the adenosine deaminase is TadA*8 containing the following sequence or a fragment thereof having adenosine deaminase activity, or consisting essentially of them: : MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQK KAQSSTD.

[0044] In some embodiments, TadA*8 is truncated. In some embodiments, the truncated form of TadA*8 lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 1 5, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to full-length TadA*8. In some embodiments the truncated form of TadA*8 lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 1 2, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to full-length TadA*8. In some embodiments, the adenosine deaminase variant is full-length TadA*8. In certain embodiments, the adenosine deaminase heterodimer is the TadA*8 domain, and among the following Selected from one of the following adenosine deaminase domains: Staphylococcus aureus (S. aureus) TadA: MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAH AEHIAIERAAKVLGSWRLEGCT LYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGS LMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKS TN Bacillus subtilis (B. subtilis) TadA: MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEML VIDEACKALGTWRLEGATLYVT LEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMN LLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKN LSE Salmonella typhimurium (S. typhimurium) TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEG WNRPIGRHDPTAHAEIMALRQGGL VLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIG RVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSD FFRMRRQEIK ALKKADRAEGAGPAV Shewanella putrefaciens (S. putrefaciens) TadA: MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEI LCLRSAGKKLENYRLLDATLYITLE PCAMCAGAMVHSRIARVVYGARDEKTGAAGT VVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQR AQQGIE Haemophilus influenzae F3031 (H. influenzae) TadA: MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTΑΗ AEIIALRNGAKNIQNY RLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYK TGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKR REEKKIEKALLKSLSDK Caulobacter crescentus (C. crescentus) TadA: MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAH DPTAHAEIAAMRAAAAKLGNYRL TDLTLVVTLEPCAMCAGAISHARIGRVVFGADD PKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRK AKI Geobacter sulfurreducens (G. sulfurreducens) TadA: MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSN DPSAHAEMIAIRQAARRSANWRL TGATLYVTLEPCLMCMGAIILARLERVVFGCYDP KGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRR RKKAKATPALF IDERKVPPEP TadA*7.10 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD。

[0045] The "adenosine deaminase base editor 8 (ABE8) polypeptide" or "ABE8" means, as defined in this specification, a base editor, and includes an adenosine deaminase variant that contains a modification at amino acid position 82 and / or 166 of the following reference sequence: MSEVEFSHEYWMRHALT LAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHS RIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD。

[0046] In some embodiments, ABE8 contains additional modifications as described herein relative to the reference sequence.

[0047] The "adenosine deaminase base editor 8 (ABE8) polynucleotide" means a polynucleotide encoding ABE8.

[0048] "Administer" means to provide one or more of the compositions described herein to a patient or subject, as referred to herein. By way of example, and not limitation, administration of a composition, such as an injection, can be intravenous (i.v.), subcutaneous (s.c.), intradermal (i.d.), Alternatively, it can be carried out by intramuscular (i.m.) injection. One or more such routes can be used. Parenteral administration can be carried out, for example, by bolus injection or by progressive perfusion over time. Alternatively, or simultaneously, administration can be carried out by the oral route.

[0049] "Agent" means any small molecule chemical compound, antibody, nucleic acid molecule, or polypeptide, or a fragment thereof.

[0050] "Modification" means a change (increase or decrease) in the expression level or activity of a gene or polypeptide detected by standard methods known in the art, for example, by the methods described herein. As used herein, modification includes a 10% change in the expression level, a 25% change in the expression level, a 40% change, and a change of 50% or more.

[0051] "Remit" means, for example, reducing, suppressing, attenuating, decreasing, arresting, and / or stabilizing the onset or progression of a disease.

[0052] "Analog" means a molecule that has similar functional or structural characteristics but is not identical. For example, a polypeptide analog retains the biological activity of the corresponding native polypeptide while having specific biochemical modifications that enhance the function of the analog relative to the native polypeptide. Such biochemical modifications can increase the protease resistance, membrane permeability, or half-life of the analog, for example, without changing ligand binding. An analog may contain non-natural amino acids.

[0053] ​​​​​​​​​​"Base editor (BE)" or "nucleic acid base editor (NBE)" refers to an agent that binds to a polynucleotide and has nucleic acid base modification activity. In various embodiments, the base editor is a polynucleotide-programmable nucleotide binding domain in combination with a nucleic acid base-modifying polypeptide (e.g., deaminase) and a guide polynucleotide (e.g., guide RNA). In various embodiments, the agent is a biomolecular complex comprising a protein domain having base editing activity , i.e., a domain capable of modifying a base (e.g., A, T, C, G, or U) within a nucleic acid molecule (e.g., DNA). In some embodiments, the polynucleotide -programmable DNA binding domain is fused or linked to a deaminase domain. In one embodiment, the agent is a fusion protein comprising one or more domains having base editing activity . In another embodiment, a protein domain having base editing activity is linked to a guide RNA (e.g., via an RNA binding motif on the guide RNA and an RNA binding domain fused to a deaminase). In some embodiments, the domain having base editing activity can deaminate a base within a nucleic acid molecule. In some embodiments, the base editor can deaminate one or more bases within a DNA molecule. In some embodiments , the base editor can deaminate cytosine (C) or adenosine (A) in DNA. In some embodiments, the base editor can deaminate both cytosine (C) and adenosine (A) in DNA. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenosine base editor (ABE) . In some embodiments, the base editor can deaminate cytosine (C) or adenosine (A) in DNA. In some embodiments, the base editor can deaminate both cytosine (C) and adenosine (A) in DNA. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenosine base editor (ABE) . In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenosine base editor (ABE) E). In some embodiments, the base editor is an adenosine base editor (ABE) and a cytidine base editor (CBE). In some embodiments, the base editor is a nuclease-inactive Cas9 (dCas9) fused to an adenosine deaminase. In some embodiments, Cas9 is a circular permutant Cas9 (e.g., spCas9 or saCas9). Circular permutant Cas9 is known in the art and is described, for example, in Oakes et al., Cell 176, 254-267, 2019. In some embodiments, the base editor is fused to an inhibitor of base excision repair, such as the UGI domain or the dISN domain. In some embodiments, the fusion protein comprises a Cas9 nickase fused to a deaminase and an inhibitor of base excision repair, such as the UGI or dISN domain. In other embodiments, the base editor is a base editor for base dropout. In some embodiments, the adenosine deaminase is evolved from TadA. In some embodiments, the polynucleotide-programmable DNA binding domain is a CRISPR-related (e.g., Cas or Cpf1) enzyme. In some embodiments, the base editor is a catalytically inactive Cas9 (dCas9) fused to a deaminase domain. In some embodiments, the base editor is a Cas9 nickase (nCas9) fused to a deaminase domain. In some embodiments, the base editor is fused to an inhibitor of base excision repair (BER). In some embodiments, the inhibitor of base excision repair is a uracil DNA glycosylase inhibitor.

[0054] In some embodiments, the adenosine deaminase is evolved from TadA. In some embodiments, the polynucleotide-programmable DNA binding domain is a CRISPR-related (e.g., Cas or Cpf1) enzyme. In some embodiments, the base editor is a catalytically inactive Cas9 (dCas9) fused to a deaminase domain. In some embodiments, the base editor is a Cas9 nickase (nCas9) fused to a deaminase domain. In some embodiments, the base editor is fused to an inhibitor of base excision repair (BER). In some embodiments, the inhibitor of base excision repair is a uracil DNA glycosylase inhibitor. In some embodiments, the base editor is a catalytically inactive Cas9 (dCas9) fused to a deaminase domain. In some embodiments, the base editor is a Cas9 nickase (nCas9) fused to a deaminase domain. In some embodiments, the base editor is a Cas9 nickase (nCas9) fused to a deaminase domain. In some embodiments, the base editor is fused to an inhibitor of base excision repair (BER). In some embodiments, the inhibitor of base excision repair is a uracil DNA glycosylase inhibitor. ​​​is a uracil glycosylase inhibitor (UGI). In some embodiments, the base excision repair inhibitor is an inosine base excision repair inhibitor. Details of the base editor are described in International PCT Application No. PCT / 2017 / 045381 (WO2018 / 027078) and PCT / US2016 / 058344 (WO2017 / 070632), each of which is incorporated herein by reference in its entirety. Also, Komor, A. C., et al., ”Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533,420-424 (2016);Gaudelli, N. M., et al., ”Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551,464-471 (2017);Komor, A. C., et al., ”Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao 4774 (2017), and Rees, H. A., et al., ”Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet.2018 Dec;19 (12):770-788.doi:10.1038 / s41576-018-0059-1 is a uracil glycosylase inhibitor (UGI). In some embodiments, the base excision repair inhibitor is an inosine base excision repair inhibitor. Details of the base editor are described in International PCT Application No. PCT / 2017 / 045381 (WO2018 / 027078) and PCT / US2016 / 058344 (WO2017 / 070632), each of which is incorporated herein by reference in its entirety. Also, Komor, A. C., et al., ”Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533,420-424 (2016);Gaudelli, N. M., et al., ”Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551,464-471 (2017);Komor, A. C., et al., ”Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao 4774 (2017), and Rees, H. A., et al., ”Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet.2018 Dec;19 (12):770-788.doi:10.1038 / s41576-018-0059-1 027078) and PCT / US2016 / 058344 (WO2017 / 070632), each of which is incorporated herein by reference in its entirety. Also, Komor, A. C., et al., ”Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533,420-424 (2016);Gaudelli, N. M., et al., ”Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551,464-471 (2017);Komor, A. C., et al., ”Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao 4774 (2017), and Rees, H. A., et al., ”Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet.2018 Dec;19 (12):770-788.doi:10.1038 / s41576-018-0059-1 ogrammable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533,420-424 (2016);Gaudelli, N. M., et al., ”Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551,464-471 (2017);Komor, A. C., et al., ”Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao 4774 (2017), and Rees, H. A., et al., ”Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet.2018 Dec;19 (12):770-788.doi:10.1038 / s41576-018-0059-1 ogrammable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533,420-424 (2016);Gaudelli, N. M., et al., ”Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551,464-471 (2017);Komor, A. C., et al., ”Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao 4774 (2017), and Rees, H. A., et al., ”Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet.2018 Dec;19 (12):770-788.doi:10.1038 / s41576-018-0059-1 ogrammable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533,420-424 (2016);Gaudelli, N. M., et al., ”Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551,464-471 (2017);Komor, A. C., et al., ”Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao 4774 (2017), and Rees, H. A., et al., ”Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet.2018 Dec;19 (12):770-788.doi:10.1038 / s41576-018-0059-1 editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551,464-471 (2017);Komor, A. C., et al., ”Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao 4774 (2017), and Rees, H. A., et al., ”Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet.2018 Dec;19 (12):770-788.doi:10.1038 / s41576-018-0059-1 editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551,464-471 (2017);Komor, A. C., et al., ”Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao 4774 (2017), and Rees, H. A., et al., ”Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet.2018 Dec;19 (12):770-788.doi:10.1038 / s41576-018-0059-1 acteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao 4774 (2017), and Rees, H. A., et al., ”Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet.2018 Dec;19 (12):770-788.doi:10.1038 / s41576-018-0059-1 acteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao 4774 (2017), and Rees, H. A., et al., ”Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet.2018 Dec;19 (12):770-788.doi:10.1038 / s41576-018-0059-1 al., ”Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet.2018 Dec;19 (12):770-788.doi:10.1038 / s41576-018-0059-1 al., ”Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet.2018 Dec;19 (12):770-788.doi:10.1038 / s41576-018-0059-1 See also (the entire content of which is incorporated herein by reference).

[0055] In some embodiments, the base editor is generated by cloning an adenosine deaminase variant (e.g., , TadA*8) into a backbone comprising a circular permutation variant Cas9 (e.g., spCAS9) and a bipartite nuclear localization sequence (e.g., ABE8). Circular permutation variant Cas9 is known in the art and is described, for example, in Oakes et al., Cell 176, 254-267, 2019 . Exemplary circular permutation variant sequences are set forth below, where the bold sequences represent sequences derived from Cas9, the italic sequences represent linker sequences, and the underlined sequences represent bipartite nuclear localization sequences. CP5 (Pam variant with MSP “NGC= mutation, normal Cas9 prefers NGG” PID= protein interaction domain and “D10A” nickase): JPEG2025102768000001.jpg171165

[0056] In some embodiments, ABE8 is selected from the base editors in Table 8 below. In some embodiments, ABE8 comprises an adenosine deaminase variant evolved from TadA . In some embodiments, the adenosine deaminase variant of ABE8 is a TadA*8 variant as described in Table 8 below. In some embodiments, the adenosine deaminase variant is a TadA*7.10 variant (e.g., TadA*8) comprising one or more modifications selected from the group consisting of Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R . ​。In various embodiments, ABE8 is Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q1 54S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H +Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R ; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R, and I76Y+V82S+Y123H+Y147R+Q154 R, including a combination of modifications selected from the group of TadA*7.10 variants (e.g., TadA*8). contain.

[0057] In some embodiments, ABE8 is a monomeric construct. In some embodiments, ABE8 is a heterodimeric construct. In some embodiments, the ABE8 base editor comprises the following sequence including: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQK KAQSSTD

[0058] As an example, the adenine base editor ABE used in the base editing compositions, systems, and methods described herein has the following nucleic acid sequence (8877 base pairs) (Addgene, Watertown, MA.; G as shown below (8877 base pairs) (Addgene, Watertown, MA.; G Audelli NM, et al., Nature. 2017 Nov 23;551(7681):464-471. doi:10.1038 / nature2464 4; Koblan LW, et al., Nat Biotechnol. 2018 Oct;36(9):843-846. doi:10.1038 / nbt.4172 .)。Polynucleotide sequences having at least 95% or more identity to the ABE nucleic acid sequence are also included. ATATGCCAAGTACGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTATGCCCAGTACAT GACCTTATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATGGTGATGCGG TTTTGGCAGTACATCAATGGGCGTGGATAGCGGTTTGACTCACGGGGATTTCCAAGTCTCCACCCCATTG ACGTCAATGGGAGTTTGTTTTGGCACCAAAATCAACGGGACTTTCCAAAATGTCGTAACAACTCCGCCCC ATTGACGCAAATGGGCGGTAGGCGTGTACGGTGGGAGGTCTATATAAGCAGAGCTGGTTTAGTGAACCGT CAGATCCGCTAGAGATCCGCGGCCGCTAATACGACTCACTATAGGGAGAGCCGCCACCATGAAACGGACA GCCGACGGAAGCGAGTTCGAGTCACCAAAGAAGAAGCGGAAAGTCTCTGAAGTCGAGTTTAGCCACGAGT ATTGGATGAGGCACGCACTGACCCTGGCAAAGCGAGCATGGGATGAAAGAGAAGTCCCCGTGGGCGCCGT GCTGGTGCACAACAATAGAGTGATCGGAGAGGGATGGAACAGGCCAATCGGCCGCCACGACCCTACCGCA CACGCAGAGATCATGGCACTGAGGCAGGGAGGCCTGGTCATGCAGAATTACCGCCTGATCGATGCCACCC TGTATGTGACACTGGAGCCATGCGTGATGTGCGCAGGAGCAATGATCCACAGCAGGATCGGAAGAGTGGT GTTCGGAGCACGGGACGCCAAGACCGGCGCAGCAGGCTCCCTGATGGATGTGCTGCACCACCCCGGCATG AACCACCGGGTGGAGATCACAGAGGGAATCCTGGCAGACGAGTGCGCCGCCCTGCTGAGCGATTTCTTTA GAATGCGGAGACAGGAGATCAAGGCCCAGAAGAAGGCACAGAGCTCCACCGACTCTGGAGGATCTAGCGG AGGATCCTCTGGAAGCGAGACACCAGGCACAAGCGAGTCCGCCACACCAGAGAGCTCCGGCGGCTCCTCC GGAGGATCCTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGCCAAGAGGG CACGCGATGAGAGGGAGGTGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGCGAGGGCTG GAACAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATGGCCCTGAGACAGGGCGGCCTG GTCATGCAGAACTACAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCG GCGCCATGATCCACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACGCAAAAACCGGCGCCGCAGG CTCCCTGATGGACGTGCTGCACTACCCCGGCATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCA GATGAATGTGCCGCCCTGCTGTGCTATTTCTTTCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAAGG CCCAGAGCTCCACCGACTCCGGAGGATCTAGCGGAGGCTCCTCTGGCTCTGAGACACCTGGCACAAGCGA GAGCGCAACACCTGAAAGCAGCGGGGGCAGCAGCGGGGGGTCAGACAAGAAGTACAGCATCGGCCTGGCC ATCGGCACCAACTCTGTGGGCTGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGG TGCTGGGCAACACCGACCGGCACAGCATCAAGAAGAACCTGATCGGAGCCCTGCTGTTCGACAGCGGCGA AACAGCCGAGGCCACCCGGCTGAAGAGAACCGCCAGAAGAAGATACACCAGACGGAAGAACCGGATCTGC TATCTGCAAGAGATCTTCAGCAACGAGATGGCCAAGGTGGACGACAGCTTCTTCCACAGACTGGAAGAGT CCTTCCTGGTGGAAGAGGATAAGAAGCACGAGCGGCACCCCATCTTCGGCAACATCGTGGACGAGGTGGC CTACCACGAGAAGTACCCCACCATCTACCACCTGAGAAAGAAACTGGTGGACAGCACCGACAAGGCCGAC CTGCGGCTGATCTATCTGGCCCTGGCCCACATGATCAAGTTCCGGGGCCACTTCCTGATCGAGGGCGACC TGAACCCCGACAACAGCGACGTGGACAAGCTGTTCATCCAGCTGGTGCAGACCTACAACCAGCTGTTCGA GGAAAACCCCATCAACGCCAGCGGCGTGGACGCCAAGGCCATCCTGTCTGCCAGACTGAGCAAGAGCAGA CGGCTGGAAAATCTGATCGCCCAGCTGCCCGGCGAGAAGAAGAATGGCCTGTTCGGAAACCTGATTGCCC TGAGCCTGGGCCTGACCCCCAACTTCAAGAGCAACTTCGACCTGGCCGAGGATGCCAAACTGCAGCTGAG CAAGGACACCTACGACGACGACCTGGACAACCTGCTGGCCCAGATCGGCGACCAGTACGCCGACCTGTTT CTGGCCGCCAAGAACCTGTCCGACGCCATCCTGCTGAGCGACATCCTGAGAGTGAACACCGAGATCACCA AGGCCCCCCTGAGCGCCTCTATGATCAAGAGATACGACGAGCACCACCAGGACCTGACCCTGCTGAAAGC TCTCGTGCGGCAGCAGCTGCCTGAGAAGTACAAAGAGATTTTCTTCGACCAGAGCAAGAACGGCTACGCC GGCTACATTGACGGCGGAGCCAGCCAGGAAGAGTTCTACAAGTTCATCAAGCCCATCCTGGAAAAGATGG ACGGCACCGAGGAACTGCTCGTGAAGCTGAACAGAGAGGACCTGCTGCGGAAGCAGCGGACCTTCGACAA CGGCAGCATCCCCCACCAGATCCACCTGGGAGAGCTGCACGCCATTCTGCGGCGGCAGGAAGATTTTTAC CCATTCCTGAAGGACAACCGGGAAAAGATCGAGAAGATCCTGACCTTCCGCATCCCCTACTACGTGGGCC CTCTGGCCAGGGGAAACAGCAGATTCGCCTGGATGACCAGAAAGAGCGAGGAAACCATCACCCCCTGGAA CTTCGAGGAAGTGGTGGACAAGGGCGCTTCCGCCCAGAGCTTCATCGAGCGGATGACCAACTTCGATAAG AACCTGCCCAACGAGAAGGTGCTGCCCAAGCACAGCCTGCTGTACGAGTACTTCACCGTGTATAACGAGC TGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAAGCCCGCCTTCCTGAGCGGCGAGCAGAAAAAGGC CATCGTGGACCTGCTGTTCAAGACCAACCGGAAAGTGACCGTGAAGCAGCTGAAAGAGGACTACTTCAAG AAAATCGAGTGCTTCGACTCCGTGGAAATCTCCGGCGTGGAAGATCGGTTCAACGCCTCCCTGGGCACAT ACCACGATCTGCTGAAAATTATCAAGGACAAGGACTTCCTGGACAATGAGGAAAACGAGGACATTCTGGA AGATATCGTGCTGACCCTGACACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTGAAAACCTATGCC CACCTGTTCGACGACAAAGTGATGAAGCAGCTGAAGCGGCGGAGATACACCGGCTGGGGCAGGCTGAGCC GGAAGCTGATCAACGGCATCCGGGACAAGCAGTCCGGCAAGACAATCCTGGATTTCCTGAAGTCCGACGG CTTCGCCAACAGAAACTTCATGCAGCTGATCCACGACGACAGCCTGACCTTTAAAGAGGACATCCAGAAA GCCCAGGTGTCCGGCCAGGGCGATAGCCTGCACGAGCACATTGCCAATCTGGCCGGCAGCCCCGCCATTA AGAAGGGCATCCTGCAGACAGTGAAGGTGGTGGACGAGCTCGTGAAAGTGATGGGCCGGCACAAGCCCGA GAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCACCCAGAAGGGACAGAAGAACAGCCGCGAGAGA ATGAAGCGGATCGAAGAGGGCATCAAAGAGCTGGGCAGCCAGATCCTGAAAGAACACCCCGTGGAAAACA CCCAGCTGCAGAACGAGAAGCTGTACCTGTACTACCTGCAGAATGGGCGGGATATGTACGTGGACCAGGA ACTGGACATCAACCGGCTGTCCGACTACGATGTGGACCATATCGTGCCTCAGAGCTTTCTGAAGGACGAC TCCATCGACAACAAGGTGCTGACCAGAAGCGACAAGAACCGGGGCAAGAGCGACAACGTGCCCTCCGAAG AGGTCGTGAAGAAGATGAAGAACTACTGGCGGCAGCTGCTGAACGCCAAGCTGATTACCCAGAGAAAGTT CGACAATCTGACCAAGGCCGAGAGAGGCGGCCTGAGCGAACTGGATAAGGCCGGCTTCATCAAGAGACAG CTGGTGGAAACCCGGCAGATCACAAAGCACGTGGCACAGATCCTGGACTCCCGGATGAACACTAAGTACG ACGAGAATGACAAGCTGATCCGGGAAGTGAAAGTGATCACCCTGAAGTCCAAGCTGGTGTCCGATTTCCG GAAGGATTTCCAGTTTTACAAAGTGCGCGAGATCAACAACTACCACCACGCCCACGACGCCTACCTGAAC GCCGTCGTGGGAACCGCCCTGATCAAAAAGTACCCTAAGCTGGAAAGCGAGTTCGTGTACGGCGACTACA AGGTGTACGACGTGCGGAAGATGATCGCCAAGAGCGAGCAGGAAATCGGCAAGGCTACCGCCAAGTACTT CTTCTACAGCAACATCATGAACTTTTTCAAGACCGAGATTACCCTGGCCAACGGCGAGATCCGGAAGCGG CCTCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGTGGGATAAGGGCCGGGATTTTGCCACCGTGC GGAAAGTGCTGAGCATGCCCCAAGTGAATATCGTGAAAAAGACCGAGGTGCAGACAGGCGGCTTCAGCAA AGAGTCTATCCTGCCCAAGAGGAACAGCGATAAGCTGATCGCCAGAAAGAAGGACTGGGACCCTAAGAAG TACGGCGGCTTCGACAGCCCCACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAAGGGCAAGT CCAAGAAACTGAAGAGTGTGAAAGAGCTGCTGGGGATCACCATCATGGAAAGAAGCAGCTTCGAGAAGAA TCCCATCGACTTTCTGGAAGCCAAGGGCTACAAAGAAGTGAAAAAGGACCTGATCATCAAGCTGCCTAAG TACTCCCTGTTCGAGCTGGAAAACGGCCGGAAGAGAATGCTGGCCTCTGCCGGCGAACTGCAGAAGGGAA ACGAACTGGCCCTGCCCTCCAAATATGTGAACTTCCTGTACCTGGCCAGCCACTATGAGAAGCTGAAGGG CTCCCCCGAGGATAATGAGCAGAAACAGCTGTTTGTGGAACAGCACAAGCACTACCTGGACGAGATCATC GAGCAGATCAGCGAGTTCTCCAAGAGAGTGATCCTGGCCGACGCTAATCTGGACAAAGTGCTGTCCGCCT ACAACAAGCACCGGGATAAGCCCATCAGAGAGCAGGCCGAGAATATCATCCACCTGTTTACCCTGACCAA TCTGGGAGCCCCTGCCGCCTTCAAGTACTTTGACACCACCATCGACCGGAAGAGGTACACCAGCACCAAA GAGGTGCTGGACGCCACCCTGATCCACCAGAGCATCACCGGCCTGTACGAGACACGGATCGACCTGTCTC AGCTGGGAGGTGACTCTGGCGGCTCAAAAAGAACCGCCGACGGCAGCGAATTCGAGCCCAAGAAGAAGAG GAAAGTCTAACCGGTCATCATCACCATCACCATTGAGTTTAAACCCGCTGATCAGCCTCGACTGTGCCTT CTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCAC TGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGT GGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCT CTATGGCTTCTGAGGCGGAAAGAACCAGCTGGGGCTCGATACCGTCGACCTCTAGCTAGAGCTTGGCGTA ATCATGGTCATAGCTGTTTCCTGTGTGAAATTGTTATCCGCTCACAATTCCACACAACATACGAGCCGGA AGCATAAAGTGTAAAGCCTAGGGTGCCTAATGAGTGAGCTAACTCACATTAATTGCGTTGCGCTCACTGC CCGCTTTCCAGTCGGGAAACCTGTCGTGCCAGCTGCATTAATGAATCGGCCAACGCGCGGGGAGAGGCGG TTTGCGTATTGGGCGCTCTTCCGCTTCCTCGCTCACTGACTCGCTGCGCTCGGTCGTTCGGCTGCGGCGA GCGGTATCAGCTCACTCAAAGGCGGTAATACGGTTATCCACAGAATCAGGGGATAACGCAGGAAAGAACA TGTGAGCAAAAGGCCAGCAAAAGGCCAGGAACCGTAAAAAGGCCGCGTTGCTGGCGTTTTTCCATAGGCT CCGCCCCCCTGACGAGCATCACAAAAATCGACGCTCAAGTCAGAGGTGGCGAAACCCGACAGGACTATAA AGATACCAGGCGTTTCCCCCTGGAAGCTCCCTCGTGCGCTCTCCTGTTCCGACCCTGCCGCTTACCGGAT ACCTGTCCGCCTTTCTCCCTTCGGGAAGCGTGGCGCTTTCTCATAGCTCACGCTGTAGGTATCTCAGTTC GGTGTAGGTCGTTCGCTCCAAGCTGGGCTGTGTGCACGAACCCCCCGTTCAGCCCGACCGCTGCGCCTTA TCCGGTAACTATCGTCTTGAGTCCAACCCGGTAAGACACGACTTATCGCCACTGGCAGCAGCCACTGGTA ACAGGATTAGCAGAGCGAGGTATGTAGGCGGTGCTACAGAGTTCTTGAAGTGGTGGCCTAACTACGGCTA CACTAGAAGAACAGTATTTGGTATCTGCGCTCTGCTGAAGCCAGTTACCTTCGGAAAAAGAGTTGGTAGC TCTTGATCCGGCAAACAAACCACCGCTGGTAGCGGTGGTTTTTTTGTTTGCAAGCAGCAGATTACGCGCA GAAAAAAAGGATCTCAAGAAGATCCTTTGATCTTTTCTACGGGGTCTGACACTCAGTGGAACGAAAACTC ACGTTAAGGGATTTTGGTCATGAGATTATCAAAAAGGATCTTCACCTAGATCCTTTTAAATTAAAAATGA AGTTTTAAATCAATCTAAAGTATATATGAGTAAACTTGGTCTGACAGTTACCAATGCTTAATCAGTGAGG CACCTATCTCAGCGATCTGTCTATTTCGTTCATCCATAGTTGCCTGACTCCCCGTCGTGTAGATAACTAC GATACGGGAGGGCTTACCATCTGGCCCCAGTGCTGCAATGATACCGCGAGACCCACGCTCACCGGCTCCA GATTTATCAGCAATAAACCAGCCAGCCGGAAGGGCCGAGCGCAGAAGTGGTCCTGCAACTTTATCCGCCT CCATCCAGTCTATTAATTGTTGCCGGGAAGCTAGAGTAAGTAGTTCGCCAGTTAATAGTTTGCGCAACGT TGTTGCCATTGCTACAGGCATCGTGGTGTCACGCTCGTCGTTTGGTATGGCTTCATTCAGCTCCGGTTCC CAACGATCAAGGCGAGTTACATGATCCCCCATGTTGTGCAAAAAAGCGGTTAGCTCCTTCGGTCCTCCGA TCGTTGTCAGAAGTAAGTTGGCCGCAGTGTTATCACTCATGGTTATGGCAGCACTGCATAATTCTCTTAC TGTCATGCCATCCGTAAGATGCTTTTCTGTGACTGGTGAGTACTCAACCAAGTCATTCTGAGAATAGTGT ATGCGGCGACCGAGTTGCTCTTGCCCGGCGTCAATACGGGATAATACCGCGCCACATAGCAGAACTTTAA AAGTGCTCATCATTGGAAAACGTTCTTCGGGGCGAAAACTCTCAAGGATCTTACCGCTGTTGAGATCCAG TTCGATGTAACCCACTCGTGCACCCAACTGATCTTCAGCATCTTTTACTTTCACCAGCGTTTCTGGGTGA GCAAAAACAGGAAGGCAAAATGCCGCAAAAAAGGGAATAAGGGCGACACGGAAATGTTGAATACTCATAC TCTTCCTTTTTCAATATTATTGAAGCATTTATCAGGGTTATTGTCTCATGAGCGGATACATATTTGAATG TATTTAGAAAAATAAACAAATAGGGGTTCCGCGCACATTTCCCCGAAAAGTGCCACCTGACGTCGACGGA TCGGGAGATCGATCTCCCGATCCCCTAGGGTCGACTCTCAGTACAATCTGCTCTGATGCCGCATAGTTAA GCCAGTATCTGCTCCCTGCTTGTGTGTTGGAGGTCGCTGAGTAGTGCGCGAGCAAAATTTAAGCTACAAC AAGGCAAGGCTTGACCGACAATTGCATGAAGAATCTGCTTAGGGTTAGGCGTTTTGCGCTGCTTCGCGAT GTACGGGCCAGATATACGCGTTGACATTGATTATTGACTAGTTATTAATAGTAATCAATTACGGGGTCAT TAGTTCATAGCCCATATATGGAGTTCCGCGTTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCC CAACGACCCCCGCCCATTGACGTCAATAATGACGTATGTTCCCATAGTAACGCCAATAGGGACTTTCCAT TGACGTCAATGGGTGGAGTATTTACGGTAAACTGCCCACTTGGCAGTACATCAAGTGTATC

[0059] As an example, the cytidine salts used in the base editing compositions, systems, and methods described herein base editor (CBE) has the following nucleic acid sequence (8877 base pairs) (Addgene, Watertown, MA.; K omor AC, et al., 2017, Sci Adv., 30;3 (8):eaao 4774.doi:10.1126 / sciadv.aao4774) Also included are polynucleotide sequences having at least 95% identity to the BE4 nucleic acid sequence. . 1 ATATGCCAAG TACGCCCCCT ATTGACGTCA ATGACGGTAA ATGGCCCGCC TGGCATTATG 61 CCCAGTACAT GACCTTATGG GACTTTCCTA CTTGGCAGTA CATCTACGTA TTAGTCATCG 121 CTATTACCAT GGTGATGCGG TTTTGGCAGT ACATCAATGG GCGTGGATAG CGGTTTGACT 181 CACGGGGATT TCCAAGTCTC CACCCCATTG ACGTCAATGG GAGTTTGTTT TGGCACCAAA 241 ATCAACGGGA CTTTCCAAAA TGTCGTAACA ACTCCGCCCC ATTGACGCAA ATGGGCGGTA 301 GGCGTGTACG GTGGGAGGTC TATATAAGCA GAGCTGGTTT AGTGAACCGT CAGATCCGCT 361 AGAGATCCGC GGCCGCTAAT ACGACTCACT ATAGGGAGAG CCGCCACCAT GAGCTCAGAG 421 ACTGGCCCAG TGGCTGTGGA CCCCACATTG AGACGGCGGA TCGAGCCCCA TGAGTTTGAG 481 GTATTCTTCG ATCCGAGAGA GCTCCGCAAG GAGACCTGCC TGCTTTACGA AATTAATTGG 541 GGGGGCCGGC ACTCCATTTG GCGACATACA TCACAGAACA CTAACAAGCA CGTCGAAGTC 601 AACTTCATCG AGAAGTTCAC GACAGAAAGA TATTTCTGTC CGAACACAAG GTGCAGCATT 661 ACCTGGTTTC TCAGCTGGAG CCCATGCGGC GAATGTAGTA GGGCCATCAC TGAATTCCTG 721 TCAAGGTATC CCCACGTCAC TCTGTTTATT TACATCGCAA GGCTGTACCA CCACGCTGAC 781 CCCCGCAATC GACAAGGCCT GCGGGATTTG ATCTCTTCAG GTGTGACTAT CCAAATTATG 841 ACTGAGCAGG AGTCAGGATA CTGCTGGAGA AACTTTGTGA ATTATAGCCC GAGTAATGAA 901 GCCCACTGGC CTAGGTATCC CCATCTGTGG GTACGACTGT ACGTTCTTGA ACTGTACTGC 961 ATCATACTGG GCCTGCCTCC TTGTCTCAAC ATTCTGAGAA GGAAGCAGCC ACAGCTGACA 1021 TTCTTTACCA TCGCTCTTCA GTCTTGTCAT TACCAGCGAC TGCCCCCACA CATTCTCTGG 1081 GCCACCGGGT TGAAATCTGG TGGTTCTTCT GGTGGTTCTA GCGGCAGCGA GACTCCCGGG 1141 ACCTCAGAGT CCGCCACACC CGAAAGTTCT GGTGGTTCTT CTGGTGGTTC TGATAAAAAG 1201 TATTCTATTG GTTTAGCCAT CGGCACTAAT TCCGTTGGAT GGGCTGTCAT AACCGATGAA 1261 TACAAAGTAC CTTCAAAGAA ATTTAAGGTG TTGGGGAACA CAGACCGTCA TTCGATTAAA 1321 AAGAATCTTA TCGGTGCCCT CCTATTCGAT AGTGGCGAAA CGGCAGAGGC GACTCGCCTG 1381 AAACGAACCG CTCGGAGAAG GTATACACGT CGCAAGAACC GAATATGTTA CTTACAAGAA 1441 ATTTTTAGCA ATGAGATGGC CAAAGTTGAC GATTCTTTCT TTCACCGTTT GGAAGAGTCC 1501 TTCCTTGTCG AAGAGGACAA GAAACATGAA CGGCACCCCA TCTTTGGAAA CATAGTAGAT 1561 GAGGTGGCAT ATCATGAAAA GTACCCAACG ATTTATCACC TCAGAAAAAA GCTAGTTGAC 1621 TCAACTGATA AAGCGGACCT GAGGTTAATC TACTTGGCTC TTGCCCATAT GATAAAGTTC 1681 CGTGGGCACT TTCTCATTGA GGGTGATCTA AATCCGGACA ACTCGGATGT CGACAAACTG 1741 TTCATCCAGT TAGTACAAAC CTATAATCAG TTGTTTGAAG AGAACCCTAT AAATGCAAGT 1801 GGCGTGGATG CGAAGGCTAT TCTTAGCGCC CGCCTCTCTA AATCCCGACG GCTAGAAAAC 1861 CTGATCGCAC AATTACCCGG AGAGAAGAAA AATGGGTTGT TCGGTAACCT TATAGCGCTC 1921 TCACTAGGCC TGACACCAAA TTTTAAGTCG AACTTCGACT TAGCTGAAGA TGCCAAATTG 1981 CAGCTTAGTA AGGACACGTA CGATGACGAT CTCGACAATC TACTGGCACA AATTGGAGAT 2041 CAGTATGCGG ACTTATTTTT GGCTGCCAAA AACCTTAGCG ATGCAATCCT CCTATCTGAC 2101 ATACTGAGAG TTAATACTGA GATTACCAAG GCGCCGTTAT CCGCTTCAAT GATCAAAAGG 2161 TACGATGAAC ATCACCAAGA CTTGACACTT CTCAAGGCCC TAGTCCGTCA GCAACTGCCT 2221 GAGAAATATA AGGAAATATT CTTTGATCAG TCGAAAAACG GGTACGCAGG TTATATTGAC 2281 GGCGGAGCGA GTCAAGAGGA ATTCTACAAG TTTATCAAAC CCATATTAGA GAAGATGGAT 2341 GGGACGGAAG AGTTGCTTGT AAAACTCAAT CGCGAAGATC TACTGCGAAA GCAGCGGACT 2401 TTCGACAACG GTAGCATTCC ACATCAAATC CACTTAGGCG AATTGCATGC TATACTTAGA 2461 AGGCAGGAGG ATTTTTATCC GTTCCTCAAA GACAATCGTG AAAAGATTGA GAAAATCCTA 2521 ACCTTTCGCA TACCTTACTA TGTGGGACCC CTGGCCCGAG GGAACTCTCG GTTCGCATGG 2581 ATGACAAGAA AGTCCGAAGA AACGATTACT CCATGGAATT TTGAGGAAGT TGTCGATAAA 2641 GGTGCGTCAG CTCAATCGTT CATCGAGAGG ATGACCAACT TTGACAAGAA TTTACCGAAC 2701 GAAAAAGTAT TGCCTAAGCA CAGTTTACTT TACGAGTATT TCACAGTGTA CAATGAACTC 2761 ACGAAAGTTA AGTATGTCAC TGAGGGCATG CGTAAACCCG CCTTTCTAAG CGGAGAACAG 2821 AAGAAAGCAA TAGTAGATCT GTTATTCAAG ACCAACCGCA AAGTGACAGT TAAGCAATTG 2881 AAAGAGGACT ACTTTAAGAA AATTGAATGC TTCGATTCTG TCGAGATCTC CGGGGTAGAA 2941 GATCGATTTA ATGCGTCACT TGGTACGTAT CATGACCTCC TAAAGATAAT TAAAGATAAG 3001 GACTTCCTGG ATAACGAAGA GAATGAAGAT ATCTTAGAAG ATATAGTGTT GACTCTTACC 3061 CTCTTTGAAG ATCGGGAAAT GATTGAGGAA AGACTAAAAA CATACGCTCA CCTGTTCGAC 3121 GATAAGGTTA TGAAACAGTT AAAGAGGCGT CGCTATACGG GCTGGGGACG ATTGTCGCGG 3181 AAACTTATCA ACGGGATAAG AGACAAGCAA AGTGGTAAAA CTATTCTCGA TTTTCTAAAG 3241 AGCGACGGCT TCGCCAATAG GAACTTTATG CAGCTGATCC ATGATGACTC TTTAACCTTC 3301 AAAGAGGATA TACAAAAGGC ACAGGTTTCC GGACAAGGGG ACTCATTGCA CGAACATATT 3361 GCGAATCTTG CTGGTTCGCC AGCCATCAAA AAGGGCATAC TCCAGACAGT CAAAGTAGTG 3421 GATGAGCTAG TTAAGGTCAT GGGACGTCAC AAACCGGAAA ACATTGTAAT CGAGATGGCA 3481 CGCGAAAATC AAACGACTCA GAAGGGGCAA AAAAACAGTC GAGAGCGGAT GAAGAGAATA 3541 GAAGAGGGTA TTAAAGAACT GGGCAGCCAG ATCTTAAAGG AGCATCCTGT GGAAAATACC 3601 CAATTGCAGA ACGAGAAACT TTACCTCTAT TACCTACAAA ATGGAAGGGA CATGTATGTT 3661 GATCAGGAAC TGGACATAAA CCGTTTATCT GATTACGACG TCGATCACAT TGTACCCCAA 3721 TCCTTTTTGA AGGACGATTC AATCGACAAT AAAGTGCTTA CACGCTCGGA TAAGAACCGA 3781 GGGAAAAGTG ACAATGTTCC AAGCGAGGAA GTCGTAAAGA AAATGAAGAA CTATTGGCGG 3841 CAGCTCCTAA ATGCGAAACT GATAACGCAA AGAAAGTTCG ATAACTTAAC TAAAGCTGAG 3901 AGGGGTGGCT TGTCTGAACT TGACAAGGCC GGATTTATTA AACGTCAGCT CGTGGAAACC 3961 CGCCAAATCA CAAAGCATGT TGCACAGATA CTAGATTCCC GAATGAATAC GAAATACGAC 4021 GAGAACGATA AGCTGATTCG GGAAGTCAAA GTAATCACTT TAAAGTCAAA ATTGGTGTCG 4081 GACTTCAGAA AGGATTTTCA ATTCTATAAA GTTAGGGAGA TAAATAACTA CCACCATGCG 4141 CACGACGCTT ATCTTAATGC CGTCGTAGGG ACCGCACTCA TTAAGAAATA CCCGAAGCTA 4201 GAAAGTGAGT TTGTGTATGG TGATTACAAA GTTTATGACG TCCGTAAGAT GATCGCGAAA 4261 AGCGAACAGG AGATAGGCAA GGCTACAGCC AAATACTTCT TTTATTCTAA CATTATGAAT 4321 TTCTTTAAGA CGGAAATCAC TCTGGCAAAC GGAGAGATAC GCAAACGACC TTTAATTGAA 4381 ACCAATGGGG AGACAGGTGA AATCGTATGG GATAAGGGCC GGGACTTCGC GACGGTGAGA 4441 AAAGTTTTGT CCATGCCCCA AGTCAACATA GTAAAGAAAA CTGAGGTGCA GACCGGAGGG 4501 TTTTCAAAGG AATCGATTCT TCCAAAAAGG AATAGTGATA AGCTCATCGC TCGTAAAAAG 4561 GACTGGGACC CGAAAAAGTA CGGTGGCTTC GATAGCCCTA CAGTTGCCTA TTCTGTCCTA 4621 GTAGTGGCAA AAGTTGAGAA GGGAAAATCC AAGAAACTGA AGTCAGTCAA AGAATTATTG 4681 GGGATAACGA TTATGGAGCG CTCGTCTTTT GAAAAGAACC CCATCGACTT CCTTGAGGCG 4741 AAAGGTTACA AGGAAGTAAA AAAGGATCTC ATAATTAAAC TACCAAAGTA TAGTCTGTTT 4801 GAGTTAGAAA ATGGCCGAAA ACGGATGTTG GCTAGCGCCG GAGAGCTTCA AAAGGGGAAC 4861 GAACTCGCAC TACCGTCTAA ATACGTGAAT TTCCTGTATT TAGCGTCCCA TTACGAGAAG 4921 TTGAAAGGTT CACCTGAAGA TAACGAACAG AAGCAACTTT TTGTTGAGCA GCACAAACAT 4981 TATCTCGACG AAATCATAGA GCAAATTTCG GAATTCAGTA AGAGAGTCAT CCTAGCTGAT 5041 GCCAATCTGG ACAAAGTATT AAGCGCATAC AACAAGCACA GGGATAAACC CATACGTGAG 5101 CAGGCGGAAA ATATTATCCA TTTGTTTACT CTTACCAACC TCGGCGCTCC AGCCGCATTC 5161 AAGTATTTTG ACACAACGAT AGATCGCAAA CGATACACTT CTACCAAGGA GGTGCTAGAC 5221 GCGACACTGA TTCACCAATC CATCACGGGA TTATATGAAA CTCGGATAGA TTTGTCACAG 5281 CTTGGGGGTG ACTCTGGTGG TTCTGGAGGA TCTGGTGGTT CTACTAATCT GTCAGATATT 5341 ATTGAAAAGG AGACCGGTAA GCAACTGGTT ATCCAGGAAT CCATCCTCAT GCTCCCAGAG 5401 GAGGTGGAAG AAGTCATTGG GAACAAGCCG GAAAGCGATA TACTCGTGCA CACCGCCTAC 5461 GACGAGAGCA CCGACGAGAA TGTCATGCTT CTGACTAGCG ACGCCCCTGA ATACAAGCCT 5521 TGGGCTCTGG TCATACAGGA TAGCAACGGT GAGAACAAGA TTAAGATGCT CTCTGGTGGT 5581 TCTGGAGGAT CTGGTGGTTC TACTAATCTG TCAGATATTA TTGAAAAGGA GACCGGTAAG 5641 CAACTGGTTA TCCAGGAATC CATCCTCATG CTCCCAGAGG AGGTGGAAGA AGTCATTGGG 5701 AACAAGCCGG AAAGCGATAT ACTCGTGCAC ACCGCCTACG ACGAGAGCAC CGACGAGAAT 5761 GTCATGCTTC TGACTAGCGA CGCCCCTGAA TACAAGCCTT GGGCTCTGGT CATACAGGAT 5821 AGCAACGGTG AGAACAAGAT TAAGATGCTC TCTGGTGGTT CTCCCAAGAA GAAGAGGAAA 5881 GTCTAACCGG TCATCATCAC CATCACCATT GAGTTTAAAC CCGCTGATCA GCCTCGACTG 5941 TGCCTTCTAG TTGCCAGCCA TCTGTTGTTT GCCCCTCCCC CGTGCCTTCC TTGACCCTGG 6001 AAGGTGCCAC TCCCACTGTC CTTTCCTAAT AAAATGAGGA AATTGCATCG CATTGTCTGA 6061 GTAGGTGTCA TTCTATTCTG GGGGGTGGGG TGGGGCAGGA CAGCAAGGGG GAGGATTGGG 6121 AAGACAATAG CAGGCATGCT GGGGATGCGG TGGGCTCTAT GGCTTCTGAG GCGGAAAGAA 6181 CCAGCTGGGG CTCGATACCG TCGACCTCTA GCTAGAGCTT GGCGTAATCA TGGTCATAGC 6241 TGTTTCCTGT GTGAAATTGT TATCCGCTCA CAATTCCACA CAACATACGA GCCGGAAGCA 6301 TAAAGTGTAA AGCCTAGGGT GCCTAATGAG TGAGCTAACT CACATTAATT GCGTTGCGCT 6361 CACTGCCCGC TTTCCAGTCG GGAAACCTGT CGTGCCAGCT GCATTAATGA ATCGGCCAAC 6421 GCGCGGGGAG AGGCGGTTTG CGTATTGGGC GCTCTTCCGC TTCCTCGCTC ACTGACTCGC 6481 TGCGCTCGGT CGTTCGGCTG CGGCGAGCGG TATCAGCTCA CTCAAAGGCG GTAATACGGT 6541 TATCCACAGA ATCAGGGGAT AACGCAGGAA AGAACATGTG AGCAAAAGGC CAGCAAAAGG 6601 CCAGGAACCG TAAAAAGGCC GCGTTGCTGG CGTTTTTCCA TAGGCTCCGC CCCCCTGACG 6661 AGCATCACAA AAATCGACGC TCAAGTCAGA GGTGGCGAAA CCCGACAGGA CTATAAAGAT 6721 ACCAGGCGTT TCCCCCTGGA AGCTCCCTCG TGCGCTCTCC TGTTCCGACC CTGCCGCTTA 6781 CCGGATACCT GTCCGCCTTT CTCCCTTCGG GAAGCGTGGC GCTTTCTCAT AGCTCACGCT 6841 GTAGGTATCT CAGTTCGGTG TAGGTCGTTC GCTCCAAGCT GGGCTGTGTG CACGAACCCC 6901 CCGTTCAGCC CGACCGCTGC GCCTTATCCG GTAACTATCG TCTTGAGTCC AACCCGGTAA 6961 GACACGACTT ATCGCCACTG GCAGCAGCCA CTGGTAACAG GATTAGCAGA GCGAGGTATG 7021 TAGGCGGTGC TACAGAGTTC TTGAAGTGGT GGCCTAACTA CGGCTACACT AGAAGAACAG 7081 TATTTGGTAT CTGCGCTCTG CTGAAGCCAG TTACCTTCGG AAAAAGAGTT GGTAGCTCTT 7141 GATCCGGCAA ACAAACCACC GCTGGTAGCG GTGGTTTTTT TGTTTGCAAG CAGCAGATTA 7201 CGCGCAGAAA AAAAGGATCT CAAGAAGATC CTTTGATCTT TTCTACGGGG TCTGACGCTC 7261 AGTGGAACGA AAACTCACGT TAAGGGATTT TGGTCATGAG ATTATCAAAA AGGATCTTCA 7321 CCTAGATCCT TTTAAATTAA AAATGAAGTT TTAAATCAAT CTAAAGTATA TATGAGTAAA 7381 CTTGGTCTGA CAGTTACCAA TGCTTAATCA GTGAGGCACC TATCTCAGCG ATCTGTCTAT 7441 TTCGTTCATC CATAGTTGCC TGACTCCCCG TCGTGTAGAT AACTACGATA CGGGAGGGCT 7501 TACCATCTGG CCCCAGTGCT GCAATGATAC CGCGAGACCC ACGCTCACCG GCTCCAGATT 7561 TATCAGCAAT AAACCAGCCA GCCGGAAGGG CCGAGCGCAG AAGTGGTCCT GCAACTTTAT 7621 CCGCCTCCAT CCAGTCTATT AATTGTTGCC GGGAAGCTAG AGTAAGTAGT TCGCCAGTTA 7681 ATAGTTTGCG CAACGTTGTT GCCATTGCTA CAGGCATCGT GGTGTCACGC TCGTCGTTTG 7741 GTATGGCTTC ATTCAGCTCC GGTTCCCAAC GATCAAGGCG AGTTACATGA TCCCCCATGT 7801 TGTGCAAAAA AGCGGTTAGC TCCTTCGGTC CTCCGATCGT TGTCAGAAGT AAGTTGGCCG 7861 CAGTGTTATC ACTCATGGTT ATGGCAGCAC TGCATAATTC TCTTACTGTC ATGCCATCCG 7921 TAAGATGCTT TTCTGTGACT GGTGAGTACT CAACCAAGTC ATTCTGAGAA TAGTGTATGC 7981 GGCGACCGAG TTGCTCTTGC CCGGCGTCAA TACGGGATAA TACCGCGCCA CATAGCAGAA 8041 CTTTAAAAGT GCTCATCATT GGAAAACGTT CTTCGGGGCG AAAACTCTCA AGGATCTTAC 8101 CGCTGTTGAG ATCCAGTTCG ATGTAACCCA CTCGTGCACC CAACTGATCT TCAGCATCTT 8161 TTACTTTCAC CAGCGTTTCT GGGTGAGCAA AAACAGGAAG GCAAAATGCC GCAAAAAAGG 8221 GAATAAGGGC GACACGGAAA TGTTGAATAC TCATACTCTT CCTTTTTCAA TATTATTGAA 8281 GCATTTATCA GGGTTATTGT CTCATGAGCG GATACATATT TGAATGTATT TAGAAAAATA 8341 AACAAATAGG GGTTCCGCGC ACATTTCCCC GAAAAGTGCC ACCTGACGTC GACGGATCGG 8401 GAGATCGATC TCCCGATCCC CTAGGGTCGA CTCTCAGTAC AATCTGCTCT GATGCCGCAT 8461 AGTTAAGCCA GTATCTGCTC CCTGCTTGTG TGTTGGAGGT CGCTGAGTAG TGCGCGAGCA 8521 AAATTTAAGC TACAACAAGG CAAGGCTTGA CCGACAATTG CATGAAGAAT CTGCTTAGGG 8581 TTAGGCGTTT TGCGCTGCTT CGCGATGTAC GGGCCAGATA TACGCGTTGA CATTGATTAT 8641 TGACTAGTTA TTAATAGTAA TCAATTACGG GGTCATTAGT TCATAGCCCA TATATGGAGT 8701 TCCGCGTTAC ATAACTTACG GTAAATGGCC CGCCTGGCTG ACCGCCCAAC GACCCCCGCC 8761 CATTGACGTC AATAATGACG TATGTTCCCA TAGTAACGCC AATAGGGACT TTCCATTGAC 8821 GTCAATGGGT GGAGTATTTA CGGTAAACTG CCCACTTGGC AGTACATCAA GTGTATC

[0060] In some embodiments, the cytidine base editor is BE4 having a nucleic acid sequence selected from one of the following: :

[0061] Original BE4 nucleic acid sequence: ATGagctcagagactggcccagtggctgtggaccccacattgagacggcggatcgagccccatgagtttgaggtattctt cgatccgagagagctccgcaaggagacctgcctgctttacgaaattaattgggggggccggcactccatttggcgacata catcacagaacactaacaagcacgtcgaagtcaacttcatcgagaagttcacgacagaaagatatttctgtccgaacaca aggtgcagcattacctggtttctcagctggagccgcgaatgtagtagggccatcactgaattcctgtcaaggtatcccca cgtcactctgtttatttacatcgcaaggctgtaccaccacgctgacccccgcaatcgacaaggcctgcgggatttgatct cttcaggtgtgactatccaaattatgactgagcaggagtcaggatactgctggagaaactttgtgaattatagcccgagt aatgaagcccactggcctaggtatccccatctgtgggtacgactgtacgttcttgaactgtactgcatcatactgggcct gcctccttgtctcaacattctgagaaggaagcagccacagctgacattctttaccatcgctcttcagtcttgtcattacc agcgactgcccccacacattctctgggccaccgggttgaaatctggtggttcttctggtggttctagcggcagcgagact cccgggacctcagagtccgccacacccgaaagttctggtggttcttctggtggttctgataaaaagtattctattggttt agccatcggcactaattccgttggatgggctgtcataaccgatgaatacaaagtaccttcaaagaaatttaaggtgttgg ggaacacagaccgtcattcgattaaaaagaatcttatcggtgccctcctattcgatagtggcgaaacggcagaggcgact cgcctgaaacgaaccgctcggagaaggtatacacgtcgcaagaaccgaatatgttacttacaagaaatttttagcaatga gatggccaaagttgacgattctttctttcaccgtttggaagagtccttccttgtcgaagaggacaagaaacatgaacggc accccatctttggaaacatagtagatgaggtggcatatcatgaaaagtacccaacgatttatcacctcagaaaaaagcta gttgactcaactgataaagcggacctgaggttaatctacttggctcttgcccatatgataaagttccgtgggcactttct cattgagggtgatctaaatccggacaactcggatgtcgacaaactgttcatccagttagtacaaacctataatcagttgt ttgaagagaaccctataaatgcaagtggcgtggatgcgaaggctattcttagcgcccgcctctctaaatcccgacggcta gaaaacctgatcgcacaattacccggagagaagaaaaatgggttgttcggtaaccttatagcgctctcactaggcctgac accaaattttaagtcgaacttcgacttagctgaagatgccaaattgcagcttagtaaggacacgtacgatgacgatctcg acaatctactggcacaaattggagatcagtatgcggacttatttttggctgccaaaaaccttagcgatgcaatcctccta tctgacatactgagagttaatactgagattaccaaggcgccgttatccgcttcaatgatcaaaaggtacgatgaacatca ccaagacttgacacttctcaaggccctagtccgtcagcaactgcctgagaaatataaggaaatattctttgatcagtcga aaaacgggtacgcaggttatattgacggcggagcgagtcaagaggaattctacaagtttatcaaacccatattagagaag atggatgggacggaagagttgcttgtaaaactcaatcgcgaagatctactgcgaaagcagcggactttcgacaacggtag cattccacatcaaatccacttaggcgaattgcatgctatacttagaaggcaggaggatttttatccgttcctcaaagaca atcgtgaaaagattgagaaaatcctaacctttcgcataccttactatgtgggacccctggcccgagggaactctcggttc gcatggatgacaagaaagtccgaagaaacgattactccatggaattttgaggaagttgtcgataaaggtgcgtcagctca atcgttcatcgagaggatgaccaactttgacaagaatttaccgaacgaaaaagtattgcctaagcacagtttactttacg agtatttcacagtgtacaatgaactcacgaaagttaagtatgtcactgagggcatgcgtaaacccgcctttctaagcgga gaacagaagaaagcaatagtagatctgttattcaagaccaaccgcaaagtgacagttaagcaattgaaagaggactactt taagaaaattgaatgcttcgattctgtcgagatctccggggtagaagatcgatttaatgcgtcacttggtacgtatcatg acctcctaaagataattaaagataaggacttcctggataacgaagagaatgaagatatcttagaagatatagtgttgact cttaccctctttgaagatcgggaaatgattgaggaaagactaaaaacatacgctcacctgttcgacgataaggttatgaa acagttaaagaggcgtcgctatacgggctggggacgattgtcgcggaaacttatcaacgggataagagacaagcaaagtg gtaaaactattctcgattttctaaagagcgacggcttcgccaataggaactttatgcagctgatccatgatgactcttta accttcaaagaggatatacaaaaggcacaggtttccggacaaggggactcattgcacgaacatattgcgaatcttgctgg ttcgccagccatcaaaaagggcatactccagacagtcaaagtagtggatgagctagttaaggtcatgggacgtcacaaac cggaaaacattgtaatcgagatggcacgcgaaaatcaaacgactcagaaggggcaaaaaaacagtcgagagcggatgaag agaatagaagagggtattaaagaactgggcagccagatcttaaaggagcatcctgtggaaaatacccaattgcagaacga gaaactttacctctattacctacaaaatggaagggacatgtatgttgatcaggaactggacataaaccgtttatctgatt acgacgtcgatcacattgtaccccaatcctttttgaaggacgattcaatcgacaataaagtgcttacacgctcggataag aaccgagggaaaagtgacaatgttccaagcgaggaagtcgtaaagaaaatgaagaactattggcggcagctcctaaatgc gaaactgataacgcaaagaaagttcgataacttaactaaagctgagaggggtggcttgtctgaacttgacaaggccggat ttattaaacgtcagctcgtggaaacccgccaaatcacaaagcatgttgcacagatactagattcccgaatgaatacgaaa tacgacgagaacgataagctgattcgggaagtcaaagtaatcactttaaagtcaaaattggtgtcggacttcagaaagga ttttcaattctataaagttagggagataaataactaccaccatgcgcacgacgcttatcttaatgccgtcgtagggaccg cactcattaagaaatacccgaagctagaaagtgagtttgtgtatggtgattacaaagtttatgacgtccgtaagatgatc gcgaaaagcgaacaggagataggcaaggctacagccaaatacttcttttattctaacattatgaatttctttaagacgga aatcactctggcaaacggagagatacgcaaacgacctttaattgaaaccaatggggagacaggtgaaatcgtatgggata agggccgggacttcgcgacggtgagaaaagttttgtccatgccccaagtcaacatagtaaagaaaactgaggtgcagacc ggagggttttcaaaggaatcgattcttccaaaaaggaatagtgataagctcatcgctcgtaaaaaggactgggacccgaa aaagtacggtggcttcgatagccctacagttgcctattctgtcctagtagtggcaaaagttgagaagggaaaatccaaga aactgaagtcagtcaaagaattattggggataacgattatggagcgctcgtcttttgaaaagaaccccatcgacttcctt gaggcgaaaggttacaaggaagtaaaaaaggatctcataattaaactaccaaagtatagtctgtttgagttagaaaatgg ccgaaaacggatgttggctagcgccggagagcttcaaaaggggaacgaactcgcactaccgtctaaatacgtgaatttcc tgtatttagcgtcccattacgagaagttgaaaggttcacctgaagataacgaacagaagcaactttttgttgagcagcac aaacattatctcgacgaaatcatagagcaaatttcggaattcagtaagagagtcatcctagctgatgccaatctggacaa agtattaagcgcatacaacaagcacagggataaacccatacgtgagcaggcggaaaatattatccatttgtttactctta ccaacctcggcgctccagccgcattcaagtattttgacacaacgatagatcgcaaacgatacacttctaccaaggaggtg ctagacgcgacactgattcaccaatccatcacgggattatatgaaactcggatagatttgtcacagcttgggggtgactc tggtggttctggaggatctggtggttctactaatctgtcagatattattgaaaaggagaccggtaagcaactggttatcc aggaatccatcctcatgctcccagaggaggtggaagaagtcattgggaacaagccggaaagcgatatactcgtgcacacc gcctacgacgagagcaccgacgagaatgtcatgcttctgactagcgacgcccctgaatacaagccttgggctctggtcat acaggatagcaacggtgagaacaagattaagatgctctctggtggttctggaggatctggtggttctactaatctgtcag atattattgaaaaggagaccggtaagcaactggttatccaggaatccatcctcatgctcccagaggaggtggaagaagtc attgggaacaagccggaaagcgatatactcgtgcacaccgcctacgacgagagcaccgacgagaatgtcatgcttctgac tagcgacgcccctgaatacaagccttgggctctggtcatacaggatagcaacggtgagaacaagattaagatgctctctg gtggttctAAAAGGACGGCGGACGGATCAGAGTTCGAGAGTCCGAAAAAAAAACGAAAGGTCGAAtaa

[0062] BE4 codon-optimized 1 nucleic acid sequence: ATGTCATCCGAAACCGGGCCAGTGGCCGTAGACCCAACACTCAGGAGGCGGATAGAACCCCATGAGTTTGAAGTGTTCTT CGACCCCAGAGAGCTGCGCAAAGAGACTTGCCTCCTGTATGAAATAAATTGGGGGGGTCGCCATTCAATTTGGAGGCACA CTAGCCAGAATACTAACAAACACGTGGAGGTAAATTTTATCGAGAAGTTTACCACCGAAAGATACTTTTGCCCCAATACA CGGTGTTCAATTACCTGGTTTCTGTCATGGAGTCCATGTGGAGAATGTAGTAGAGCGATAACTGAGTTCCTGTCTCGATA TCCTCACGTCACGTTGTTTATATACATCGCTCGGCTTTATCACCATGCGGACCCGCGGAACAGGCAAGGTCTTCGGGACC TCATATCCTCTGGGGTGACCATCCAGATAATGACGGAGCAAGAGAGCGGATACTGCTGGCGAAACTTTGTTAACTACAGC CCAAGCAATGAGGCACACTGGCCTAGATATCCGCATCTCTGGGTTCGACTGTATGTCCTTGAACTGTACTGCATAATTCT GGGACTTCCGCCATGCTTGAACATTCTGCGGCGGAAACAACCACAGCTGACCTTTTTCACGATTGCTCTCCAAAGTTGTC ACTACCAGCGATTGCCACCCCACATCTTGTGGGCTACTGGACTCAAGTCTGGAGGAAGTTCAGGCGGAAGCAGCGGGTCT GAAACGCCCGGAACCTCAGAGAGCGCAACGCCCGAAAGCTCTGGAGGGTCAAGTGGTGGTAGTGATAAGAAATACTCCAT CGGCCTCGCCATCGGTACGAATTCTGTCGGTTGGGCCGTTATCACCGATGAGTACAAGGTCCCTTCTAAGAAATTCAAGG TTTTGGGCAATACAGACCGCCATTCTATAAAAAAAAACCTGATCGGCGCCCTTTTGTTTGACAGTGGTGAGACTGCTGAA GCGACTCGCCTGAAGCGAACTGCCAGGAGGCGGTATACGAGGCGAAAAAACCGAATTTGTTACCTCCAGGAGATTTTCTC AAATGAAATGGCCAAGGTAGATGATAGTTTTTTTCACCGCTTGGAAGAAAGTTTTCTCGTTGAGGAGGACAAAAAGCACG AGAGGCACCCAATCTTTGGCAACATAGTCGATGAGGTCGCATACCATGAGAAATATCCTACGATCTATCATCTCCGCAAG AAGCTGGTCGATAGCACGGATAAAGCTGACCTCCGGCTGATCTACCTTGCTCTTGCTCACATGATTAAATTCAGGGGCCA TTTCCTGATAGAAGGAGACCTCAATCCCGACAATTCTGATGTCGACAAACTGTTTATTCAGCTCGTTCAGACCTATAATC AACTCTTTGAGGAGAACCCCATCAATGCTTCAGGGGTGGACGCAAAGGCCATTTTGTCCGCGCGCTTGAGTAAATCACGA CGCCTCGAGAATTTGATAGCTCAACTGCCGGGTGAGAAGAAAAACGGGTTGTTTGGGAATCTCATAGCGTTGAGTTTGGG ACTTACGCCAAACTTTAAGTCTAACTTTGATTTGGCCGAAGATGCCAAATTGCAGCTGTCCAAAGATACCTATGATGACG ACTTGGATAACCTTCTTGCGCAGATTGGTGACCAATACGCGGATCTGTTTCTTGCCGCAAAAAATCTGTCCGACGCCATA CTCTTGTCCGATATACTGCGCGTCAATACTGAGATAACTAAGGCTCCCCTCAGCGCGTCCATGATTAAAAGATACGATGA GCACCACCAAGATCTCACTCTGTTGAAAGCCCTGGTTCGCCAGCAGCTTCCAGAGAAGTATAAGGAGATATTTTTCGACC AATCTAAAAACGGCTATGCGGGTTACATTGACGGTGGCGCCTCTCAAGAAGAATTCTACAAGTTTATAAAGCCGATACTT GAGAAAATGGACGGTACAGAGGAATTGTTGGTTAAGCTCAATCGCGAGGACTTGTTGAGAAAGCAGCGCACATTTGACAA TGGTAGTATTCCACACCAGATTCATCTGGGCGAGTTGCATGCCATTCTTAGAAGACAAGAAGATTTTTATCCGTTTCTGA AAGATAACAGAGAAAAGATTGAAAAGATACTTACCTTTCGCATACCGTATTATGTAGGTCCCCTGGCTAGAGGGAACAGT CGCTTCGCTTGGATGACTCGAAAATCAGAAGAAACAATAACCCCCTGGAATTTTGAAGAAGTGGTAGATAAAGGTGCGAG TGCCCAATCTTTTATTGAGCGGATGACAAATTTTGACAAGAATCTGCCTAACGAAAAGGTGCTTCCCAAGCATTCCCTTT TGTATGAATACTTTACAGTATATAATGAACTGACTAAAGTGAAGTACGTTACCGAGGGGATGCGAAAGCCAGCTTTTCTC AGTGGCGAGCAGAAAAAAGCAATAGTTGACCTGCTGTTCAAGACGAATAGGAAGGTTACCGTCAAACAGCTCAAAGAAGA TTACTTTAAAAAGATCGAATGTTTTGATTCAGTTGAGATAAGCGGAGTAGAGGATAGATTTAACGCAAGTCTTGGAACTT ATCATGACCTTTTGAAGATCATCAAGGATAAAGATTTTTTGGACAACGAGGAGAATGAAGATATCCTGGAAGATATAGTA CTTACCTTGACGCTTTTTGAAGATCGAGAGATGATCGAGGAGCGACTTAAGACGTACGCACATCTCTTTGACGATAAGGT TATGAAACAATTGAAACGCCGGCGGTATACTGGCTGGGGCAGGCTTTCTCGAAAGCTGATTAATGGTATCCGCGATAAGC AGTCTGGAAAGACAATCCTTGACTTTCTGAAAAGTGATGGATTTGCAAATAGAAACTTTATGCAGCTTATACATGATGAC TCTTTGACGTTCAAGGAAGACATCCAGAAGGCACAGGTATCCGGCCAAGGGGATAGCCTCCATGAACACATAGCCAACCT GGCCGGCTCACCAGCTATTAAAAAGGGAATATTGCAAACCGTTAAGGTTGTTGACGAACTCGTTAAGGTTATGGGCCGAC ACAAACCAGAGAATATCGTGATTGAGATGGCTAGGGAGAATCAGACCACTCAAAAAGGTCAGAAAAATTCTCGCGAAAGG ATGAAGCGAATTGAAGAGGGAATCAAAGAACTTGGCTCTCAAATTTTGAAAGAGCACCCGGTAGAAAACACTCAGCTGCA GAATGAAAAGCTGTATCTGTATTATCTGCAGAATGGTCGAGATATGTACGTTGATCAGGAGCTGGATATCAATAGGCTCA GTGACTACGATGTCGACCACATCGTTCCTCAATCTTTCCTGAAAGATGACTCTATCGACAACAAAGTGTTGACGCGATCA GATAAGAACCGGGGAAAATCCGACAATGTACCCTCAGAAGAAGTTGTCAAGAAGATGAAAAACTATTGGAGACAATTGCT GAACGCCAAGCTCATAACACAACGCAAGTTCGATAACTTGACGAAAGCCGAAAGAGGTGGGTTGTCAGAATTGGACAAAG CTGGCTTTATTAAGCGCCAATTGGTGGAGACCCGGCAGATTACGAAACACGTAGCACAAATTTTGGATTCACGAATGAAT ACCAAATACGACGAAAACGACAAATTGATACGCGAGGTGAAAGTGATTACGCTTAAGAGTAAGTTGGTTTCCGATTTCAG GAAGGATTTTCAGTTTTACAAAGTAAGAGAAATAAACAACTACCACCACGCCCATGATGCTTACCTCAACGCGGTAGTTG GCACAGCTCTTATCAAAAAATATCCAAAGCTGGAAAGCGAGTTCGTTTACGGTGACTATAAAGTATACGACGTTCGGAAG ATGATAGCCAAATCAGAGCAGGAAATTGGGAAGGCAACCGCAAAATACTTCTTCTATTCAAACATCATGAACTTCTTTAA GACGGAGATTACGCTCGCGAACGGCGAAATACGCAAGAGGCCCCTCATAGAGACTAACGGCGAAACCGGGGAGATCGTAT GGGACAAAGGACGGGACTTTGCGACCGTTAGAAAAGTACTTTCAATGCCACAAGTGAATATTGTTAAAAAGACAGAAGTA CAAACAGGGGGGTTCAGTAAGGAATCCATTTTGCCCAAGCGGAACAGTGATAAATTGATAGCAAGGAAAAAAGATTGGGA CCCTAAGAAGTACGGTGGTTTCGACTCTCCTACCGTTGCATATTCAGTCCTTGTAGTTGCGAAAGTGGAAAAGGGGAAAA GTAAGAAGCTTAAGAGTGTTAAAGAGCTTCTGGGCATAACCATAATGGAACGGTCTAGCTTCGAGAAAAATCCAATTGAC TTTCTCGAGGCTAAAGGTTACAAGGAGGTAAAAAAGGACCTGATAATTAAACTCCCAAAGTACAGTCTCTTCGAGTTGGA GAATGGGAGGAAGAGAATGTTGGCATCTGCAGGGGAGCTCCAAAAGGGGAACGAGCTGGCTCTGCCTTCAAAATACGTGA ACTTTCTGTACCTGGCCAGCCACTACGAGAAACTCAAGGGTTCTCCTGAGGATAACGAGCAGAAACAGCTGTTTGTAGAG CAGCACAAGCATTACCTGGACGAGATAATTGAGCAAATTAGTGAGTTCTCAAAAAGAGTAATCCTTGCAGACGCGAATCT GGATAAAGTTCTTTCCGCCTATAATAAGCACCGGGACAAGCCTATACGAGAACAAGCCGAGAACATCATTCACCTCTTTA CCCTTACTAATCTGGGCGCGCCGGCCGCCTTCAAATACTTCGACACCACGATAGACAGGAAAAGGTATACGAGTACCAAA GAAGTACTTGACGCCACTCTCATCCACCAGTCTATAACAGGGTTGTACGAAACGAGGATAGATTTGTCCCAGCTCGGCGG CGACTCAGGAGGGTCAGGCGGCTCCGGTGGATCAACGAATCTTTCCGACATAATCGAGAAAGAAACCGGCAAACAGTTGG TGATCCAAGAATCAATCCTGATGCTGCCTGAAGAAGTAGAAGAGGTGATTGGCAACAAACCTGAGTCTGACATTCTTGTC CACACCGCGTATGACGAGAGCACGGACGAGAACGTTATGCTTCTCACTAGCGACGCCCCTGAGTATAAACCATGGGCGCT GGTCATCCAAGATTCCAATGGGGAAAACAAGATTAAGATGCTTAGTGGTGGGTCTGGAGGGAGCGGTGGGTCCACGAACC TCAGCGACATTATTGAAAAAGAGACTGGTAAACAACTTGTAATACAAGAGTCTATTCTGATGTTGCCTGAAGAGGTGGAG GAGGTGATTGGGAACAAACCGGAGTCTGATATACTTGTTCATACCGCCTATGACGAATCTACTGATGAGAATGTGATGCT TTTaACGTCAGACGCTCCCGAGTACAAACCCTGGGCTCTGGTGATTCAGGACAGCAATGGTGAGAATAAGATTAAAATGT TGAGTGGGGGCTCAAAGCGCACGGCTGACGGTAGCGAATTTGAGAGCCCCAAAAAAAAACGAAAGGTCGAAtaa

[0063] BE4 codon-optimized 2 nucleic acid sequence: ATGAGCAGCGAGACAGGCCCTGTGGCTGTGGATCCTACACTGCGGAGAAGAATCGAGCCCCACGAGTTCGAGGTGTTCTT CGACCCCAGAGAGCTGCGGAAAGAGACATGCCTGCTGTACGAGATCAACTGGGGCGGCAGACACTCTATCTGGCGGCACA CAAGCCAGAACACCAACAAGCACGTGGAAGTGAACTTTATCGAGAAGTTTACGACCGAGCGGTACTTCTGCCCCAACACC AGATGCAGCATCACCTGGTTTCTGAGCTGGTCCCCTTGCGGCGAGTGCAGCAGAGCCATCACCGAGTTTCTGTCCAGATA TCCCCACGTGACCCTGTTCATCTATATCGCCCGGCTGTACCACCACGCCGATCCTAGAAATAGACAGGGACTGCGCGACC TGATCAGCAGCGGAGTGACCATCCAGATCATGACCGAGCAAGAGAGCGGCTACTGCTGGCGGAACTTCGTGAACTACAGC CCCAGCAACGAAGCCCACTGGCCTAGATATCCTCACCTGTGGGTCCGACTGTACGTGCTGGAACTGTACTGCATCATCCT GGGCCTGCCTCCATGCCTGAACATCCTGAGAAGAAAGCAGCCTCAGCTGACCTTCTTCACAATCGCCCTGCAGAGCTGCC ACTACCAGAGACTGCCTCCACACATCCTGTGGGCCACCGGACTTAAGAGCGGAGGATCTAGCGGCGGCTCTAGCGGATCT GAGACACCTGGCACAAGCGAGTCTGCCACACCTGAGAGTAGCGGCGGATCTTCTGGCGGCTCCGACAAGAAGTACTCTAT CGGACTGGCCATCGGCACCAACTCTGTTGGATGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGG TGCTGGGCAACACCGACCGGCACAGCATCAAGAAGAATCTGATCGGCGCCCTGCTGTTCGACTCTGGCGAAACAGCCGAA GCCACCAGACTGAAGAGAACCGCCAGGCGGAGATACACCCGGCGGAAGAACCGGATCTGCTACCTGCAAGAGATCTTCAG CAACGAGATGGCCAAGGTGGACGACAGCTTCTTCCACAGACTGGAAGAGTCCTTCCTGGTGGAAGAGGACAAGAAGCACG AGCGGCACCCCATCTTCGGCAACATCGTGGATGAGGTGGCCTACCACGAGAAGTACCCCACCATCTACCACCTGAGAAAG AAACTGGTGGACAGCACCGACAAGGCCGACCTGAGACTGATCTACCTGGCTCTGGCCCACATGATCAAGTTCCGGGGCCA CTTTCTGATCGAGGGCGATCTGAACCCCGACAACAGCGACGTGGACAAGCTGTTCATCCAGCTGGTGCAGACCTACAACC AGCTGTTCGAGGAAAACCCCATCAACGCCTCTGGCGTGGACGCCAAGGCTATCCTGTCTGCCAGACTGAGCAAGAGCAGA AGGCTGGAAAACCTGATCGCCCAGCTGCCTGGCGAGAAGAAGAATGGCCTGTTCGGCAACCTGATTGCCCTGAGCCTGGG ACTGACCCCTAACTTCAAGAGCAACTTCGACCTGGCCGAGGATGCCAAACTGCAGCTGAGCAAGGACACCTACGACGACG ACCTGGACAATCTGCTGGCCCAGATCGGCGATCAGTACGCCGACTTGTTTCTGGCCGCCAAGAACCTGTCCGACGCCATC CTGCTGAGCGATATCCTGAGAGTGAACACCGAGATCACAAAGGCCCCTCTGAGCGCCTCTATGATCAAGAGATACGACGA GCACCACCAGGATCTGACCCTGCTGAAGGCCCTCGTTAGACAGCAGCTGCCAGAGAAGTACAAAGAGATTTTCTTCGATC AGTCCAAGAACGGCTACGCCGGCTACATTGATGGCGGAGCCAGCCAAGAGGAATTCTACAAGTTCATCAAGCCCATCCTG GAAAAGATGGACGGCACCGAGGAACTGCTGGTCAAGCTGAACAGAGAGGACCTGCTGCGGAAGCAGCGGACCTTCGACAA TGGCTCTATCCCTCACCAGATCCACCTGGGAGAGCTGCACGCCATTCTGCGGAGACAAGAGGACTTTTACCCATTCCTGA AGGACAACCGGGAAAAGATCGAGAAGATCCTGACCTTCAGGATCCCCTACTACGTGGGACCACTGGCCAGAGGCAATAGC AGATTCGCCTGGATGACCAGAAAGAGCGAGGAAACCATCACACCCTGGAACTTCGAGGAAGTGGTGGACAAGGGCGCCAG CGCTCAGTCCTTCATCGAGCGGATGACCAACTTCGATAAGAACCTGCCTAACGAGAAGGTGCTGCCCAAGCACTCCCTGC TGTATGAGTACTTCACCGTGTACAACGAGCTGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAAGCCCGCCTTTCTG AGCGGCGAGCAGAAAAAGGCCATTGTGGATCTGCTGTTCAAGACCAACCGGAAAGTGACCGTGAAGCAGCTGAAAGAGGA CTACTTCAAGAAAATCGAGTGCTTCGACAGCGTGGAAATCAGCGGCGTGGAAGATCGGTTCAATGCCAGCCTGGGCACAT ACCACGACCTGCTGAAAATTATCAAGGACAAGGACTTCCTGGACAACGAAGAGAACGAGGACATTCTCGAGGACATCGTG CTGACCCTGACACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTGAAAACATACGCCCACCTGTTCGACGACAAAGT GATGAAGCAACTGAAGCGGAGGCGGTACACAGGCTGGGGCAGACTGTCTCGGAAGCTGATCAACGGCATCCGGGATAAGC AGTCCGGCAAGACAATCCTGGATTTCCTGAAGTCCGACGGCTTCGCCAACAGAAACTTCATGCAGCTGATCCACGACGAC AGCCTGACCTTTAAAGAGGACATCCAGAAAGCCCAGGTGTCCGGCCAAGGCGATTCTCTGCACGAGCACATTGCCAACCT GGCCGGATCTCCCGCCATTAAGAAGGGCATCCTGCAGACAGTGAAGGTGGTGGACGAGCTTGTGAAAGTGATGGGCAGAC ACAAGCCCGAGAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCACACAGAAGGGCCAGAAGAACAGCCGCGAGAGA ATGAAGCGGATCGAAGAGGGCATCAAAGAGCTGGGCAGCCAGATCCTGAAAGAACACCCCGTGGAAAACACCCAGCTGCA GAACGAGAAGCTGTACCTGTACTACCTGCAGAATGGACGGGATATGTACGTGGACCAAGAGCTGGACATCAACCGGCTGA GCGACTACGATGTGGACCATATCGTGCCCCAGAGCTTTCTGAAGGACGACTCCATCGATAACAAGGTCCTGACCAGAAGC GACAAGAACCGGGGCAAGAGCGATAACGTGCCCTCCGAAGAGGTGGTCAAGAAGATGAAGAACTACTGGCGACAGCTGCT GAACGCCAAGCTGATTACCCAGCGGAAGTTCGATAACCTGACCAAGGCCGAGAGAGGCGGCCTGAGCGAACTTGATAAGG CCGGCTTCATTAAGCGGCAGCTGGTGGAAACCCGGCAGATCACCAAACACGTGGCACAGATTCTGGACTCCCGGATGAAC ACTAAGTACGACGAGAATGACAAGCTGATCCGGGAAGTGAAAGTCATCACCCTGAAGTCTAAGCTGGTGTCCGATTTCCG GAAGGATTTCCAGTTCTACAAAGTGCGGGAAATCAACAACTACCATCACGCCCACGACGCCTACCTGAATGCCGTTGTTG GAACAGCCCTGATCAAGAAGTATCCCAAGCTGGAAAGCGAGTTCGTGTACGGCGACTACAAGGTGTACGACGTGCGGAAG ATGATCGCCAAGAGCGAACAAGAGATCGGCAAGGCTACCGCCAAGTACTTTTTCTACAGCAACATCATGAACTTTTTCAA GACAGAGATCACCCTGGCCAACGGCGAGATCCGGAAAAGACCCCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGT GGGATAAGGGCAGAGATTTTGCCACAGTGCGGAAAGTGCTGAGCATGCCCCAAGTGAATATCGTGAAGAAAACCGAGGTG CAGACAGGCGGCTTCAGCAAAGAGTCTATCCTGCCTAAGCGGAACAGCGATAAGCTGATCGCCAGAAAGAAGGACTGGGA CCCTAAGAAGTACGGCGGCTTCGATAGCCCTACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAAGGGCAAGT CCAAAAAGCTCAAGAGCGTGAAAGAGCTGCTGGGGATCACCATCATGGAAAGAAGCAGCTTTGAGAAGAACCCGATCGAC TTTCTGGAAGCCAAGGGCTACAAAGAAGTCAAGAAGGACCTCATCATCAAGCTCCCCAAGTACAGCCTGTTCGAGCTGGA AAATGGCCGGAAGCGGATGCTGGCCTCAGCAGGCGAACTGCAGAAAGGCAATGAACTGGCCCTGCCTAGCAAATACGTCA ACTTCCTGTACCTGGCCAGCCACTATGAGAAGCTGAAGGGCAGCCCCGAGGACAATGAGCAAAAGCAGCTGTTTGTGGAA CAGCACAAGCACTACCTGGACGAGATCATCGAGCAGATCAGCGAGTTCTCCAAGAGAGTGATCCTGGCCGACGCTAACCT GGATAAGGTGCTGTCTGCCTATAACAAGCACCGGGACAAGCCTATCAGAGAGCAGGCCGAGAATATCATCCACCTGTTTA CCCTGACCAACCTGGGAGCCCCTGCCGCCTTCAAGTACTTCGACACCACCATCGACCGGAAGAGGTACACCAGCACCAAA GAGGTGCTGGACGCCACACTGATCCACCAGTCTATCACCGGCCTGTACGAAACCCGGATCGACCTGTCTCAGCTCGGCGG CGATTCTGGTGGTTCTGGCGGAAGTGGCGGATCCACCAATCTGAGCGACATCATCGAAAAAGAGACAGGCAAGCAGCTCG TGATCCAAGAATCCATCCTGATGCTGCCTGAAGAGGTTGAGGAAGTGATCGGCAACAAGCCTGAGTCCGACATCCTGGTG CACACCGCCTACGATGAGAGCACCGATGAGAACGTCATGCTGCTGACAAGCGACGCCCCTGAGTACAAGCCTTGGGCTCT CGTGATTCAGGACAGCAATGGGGAGAACAAGATCAAGATGCTGAGCGGAGGTAGCGGAGGCAGTGGCGGAAGCACAAACC TGTCTGATATCATTGAAAAAGAAACCGGGAAGCAACTGGTCATTCAAGAGTCCATTCTCATGCTCCCGGAAGAAGTCGAG GAAGTCATTGGAAACAAACCCGAGAGCGATATTCTGGTCCACACAGCCTATGACGAGTCTACAGACGAAAACGTGATGCT CCTGACCTCTGACGCTCCCGAGTATAAGCCCTGGGCACTTGTTATCCAGGACTCTAACGGGGAAAACAAAATCAAAATGT TGTCCGGCGGCAGCAAGCGGACAGCCGATGGATCTGAGTTCGAGAGCCCCAAGAAGAAACGGAAGGTgGAGtaa

[0064] "Base editing activity" means acting to chemically change the bases within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, base editing activity is cytidine deaminase activity, for example, the conversion of a target C·G to T·A. In another embodiment, base editing activity is adenosine or adenine deaminase activity, for example, the conversion of A·T to G·C. In another embodiment, base editing activity is cytidine deaminase activity, for example, the conversion of a target C·G to T·A, and adenosine or adenine deaminase activity, for example, the conversion of A·T to G·C. The term "base editing system" refers to a system for editing the nucleic acid bases of a target nucleotide sequence. In various embodiments, a base editing (BE) system comprises (1) a polynucleotide-programmable nucleotide binding domain, a deaminase domain, and a cytidine deaminase domain for deaminating the nucleic acid bases in a target nucleotide sequence; and (2) one or more guide polynucleotides combined with the polynucleotide-programmable nucleotide binding domain.

[0065] ​​​​​​​​​It contains nucleotides (e.g., guide RNA). In various embodiments, the base editing (BE) system is an Nucleic acid base editor domain selected from adenosine deaminase or cytidine deaminase And a domain having nucleic acid sequence-specific binding activity. In some embodiments, the Base editing system is for (1) deaminating one or more nucleic acid bases in a target nucleotide sequence A base editor (BE) containing a polynucleotide programmable DNA binding domain and a deaminase domain; And (2) one or more guide RNAs combined with a polynucleotide programmable DNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is mainly a polynucleotide programmable DNA binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE ). In some embodiments, the base editor is an adenine or adenosine base editor (ABE ) or a cytidine base editor (CBE). The term "Cas9" or "Cas9 domain" refers to an RNA-guided nuclease containing a Cas9 protein or a fragment thereof (e.g., a protein containing the activity, inactivity, or partially active DNA cleavage domain of Cas9 and / or the gRNA binding domain of Cas9). The Cas9 nuclease

[0066] May also be called a casnl nuclease or a CRISPR (clustered regularly interspaced short palindromic repeat)-associated nuclease. An exemplary Cas9 is Streptococcus pyog And proteins containing the active, inactive, or partially active DNA cleavage domain of the gRNA binding domain of Cas9). The Cas9 nuclease May also be called a casnl nuclease or a CRISPR (clustered regularly interspaced short palindromic repeat)-associated nuclease. An exemplary Cas9 is Streptococcus pyog is enes Cas9 (spCas9), and its amino acid sequence is shown below: JPEG2025102768000002.jpg166162

[0067] The term "conservative amino acid substitution" or "conservative mutation" refers to replacing an amino acid with another amino acid that has common characteristics . A functional way to define the common characteristics between individual amino acids is to analyze the normalized frequency of amino acid changes between corresponding proteins of homologous organisms (Schulz, G.E. and Schirmer, R.H., Principles of Protein Structure, Springer-Verlag, New York (1979)). According to such an analysis, groups of amino acids can be defined when the amino acids within the group are preferentially exchanged with each other, and thus the overall effects on the protein structure are most similar to each other (Schulz, G.E. and Sch hirmer, R.H., supra). Non-limiting examples of conservative mutations include amino acid substitutions, for example, an amino acid substitution from arginine to lysine and vice versa such that the positive charge can be maintained; from aspartic acid to glutamic acid and vice versa so that the negative charge can be maintained; from threonine to serine so that the free -OH can be maintained; and substitutions from asparagine to glutamine such that the free -NH2 can be maintained.

[0068] The terms "coding sequence" or "protein coding sequence", which are used interchangeably herein, refer to a segment of a polynucleotide that encodes a protein. A region or sequence is bounded proximally at the 5' end by a start codon and proximally at the 3' end by a stop codon are provided. Examples of stop codons useful in the base editors described herein include the following: The JPEG2025102768000003.jpg37165 coding sequence can also be referred to as an open reading frame.

[0069] The term "cytidine deaminase" refers to a polypeptide or a fragment thereof that can catalyze a deamination reaction that converts an amino group to a carbonyl group. In one embodiment, cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine. PmCDA1 derived from Petromyzon marinus (Petromyzon marinus cytidine deaminase 1, "PmCDA1"), AID (activation-induced cytidine deaminase; AICDA) derived from mammals (e.g., humans, pigs, cows, horses, monkeys, etc.), and APOBEC are exemplary cytidine deaminases. means a polypeptide or a fragment thereof that can catalyze a deamination reaction that converts an amino group to a carbonyl group. In one embodiment, cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine. PmCDA1 derived from Petromyzon marinus (Petromyzon marinus cytidine deaminase 1, "PmCDA1"), AID (activation-induced cytidine deaminase; AICDA) derived from mammals (e.g., humans, pigs, cows, horses, monkeys, etc.), and APOBEC are exemplary cytidine deaminases. converts cytosine to uracil or 5-methylcytosine to thymine. PmCDA1 derived from Petromyzon marinus (Petromyzon marinus cytidine deaminase 1, "PmCDA1"), AID (activation-induced cytidine deaminase; AICDA) derived from mammals (e.g., humans, pigs, cows, horses, monkeys, etc.), and APOBEC are exemplary cytidine deaminases. from Petromyzon marinus (Petromyzon marinus cytidine deaminase 1, "PmCDA1"), AID (activation-induced cytidine deaminase; AICDA) derived from mammals (e.g., humans, pigs, cows, horses, monkeys, etc.), and APOBEC are exemplary cytidine deaminases. from Petromyzon marinus (Petromyzon marinus cytidine deaminase 1, "PmCDA1"), AID (activation-induced cytidine deaminase; AICDA) derived from mammals (e.g., humans, pigs, cows, horses, monkeys, etc.), and APOBEC are exemplary cytidine deaminases. from Petromyzon marinus (Petromyzon marinus cytidine deaminase 1, "PmCDA1"), AID (activation-induced cytidine deaminase; AICDA) derived from mammals (e.g., humans, pigs, cows, horses, monkeys, etc.), and APOBEC are exemplary cytidine deaminases. are provided. Examples of stop codons useful in the base editors described herein include

[0070] As used herein, the term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes hydrolytic deamination of cytosine to uracil. In some embodiments, the deaminase is an adenosine deaminase that catalyzes hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase As used herein, the term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes hydrolytic deamination of cytosine to uracil. In some embodiments, the deaminase is an adenosine deaminase that catalyzes hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes hydrolytic deamination of cytosine to uracil. In some embodiments, the deaminase is an adenosine deaminase that catalyzes hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes hydrolytic deamination of cytosine to uracil. In some embodiments, the deaminase is an adenosine deaminase that catalyzes hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes hydrolytic deamination of cytosine to uracil. In some embodiments, the deaminase is an adenosine deaminase that catalyzes hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes hydrolytic deamination of cytosine to uracil. In some embodiments, the deaminase is an adenosine deaminase that catalyzes hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes hydrolytic deamination of cytosine to uracil. In some embodiments, the deaminase is an adenosine deaminase that catalyzes hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes hydrolytic deamination of cytosine to uracil. In some embodiments, the deaminase is an adenosine deaminase that catalyzes hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase Minase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or adenine (A) to inosine (I). In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., modified adenosine deaminases, evolved adenosine deaminases) can be derived from any organism such as bacteria. In some embodiments, the adenosine deaminase is derived from bacteria such as E. coli, S. aureus, S. typhi, S. putrefaciens, H. influenzae, or C. crescentus. In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the deaminase or deaminase domain is a variant of a natural deaminase derived from an organism such as human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain does not exist in nature. For example, in some embodiments, the deaminase or deaminase domain has at least 50%, at least 55%, at least 60%, at least 65%, at least 70% identity with a natural deaminase. In some embodiments, the deaminase or deaminase domain has at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with a natural deaminase. In some embodiments, the deaminase or deaminase domain is a non-naturally occurring variant of a natural deaminase. For example, in some embodiments, the deaminase or deaminase domain is a non-naturally occurring variant of a natural deaminase that has been modified by rational design, directed evolution, or other mutagenesis techniques to improve its catalytic activity, substrate specificity, stability, or other properties. In some embodiments, the deaminase or deaminase domain is a non-naturally occurring variant of a natural deaminase that has been modified to have altered substrate specificity, such that it can catalyze the deamination of substrates other than adenosine or deoxyadenosine. at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identity.

[0071] "Detecting" refers to identifying the presence, absence, or amount of an analyte to be detected. In one embodiment, changes in the sequence of a polynucleotide or polypeptide are detected. In another embodiment the presence of an indel is detected.

[0072] "Detectable label" means a composition that enables a target molecule to be detected via spectroscopic, photochemical, biochemical, immunochemical, or chemical means when linked to the target molecule. Examples of useful labels include radioisotopes, magnetic beads, metal beads, colloidal particles, fluorescent dyes, high electron density reagents, enzymes (e.g., enzymes commonly used in enzyme-linked immunosorbent assay (ELISA)), biotin, digoxigenin, or haptens.

[0073] "Disease" means a pathological condition or disorder that impairs or interferes with the normal function of cells, tissues, or organs. Examples of diseases include HBV infection, as well as cirrhosis, hepatocellular carcinoma (HCC), and related diseases and disorders including those associated with or resulting from HBV infection.

[0074] "Effective amount" means the amount required to alleviate the symptoms of a disease as compared to an untreated patient. For the therapeutic treatment of a disease, the active compound used to practice The effective amount of the plurality of [objects] varies depending on the mode of administration, the age, weight, and general health of the subject. Ultimately, the attending physician or veterinarian will determine the appropriate amount and dosing regimen. Such an amount is referred to as an "effective" amount. In one embodiment, the effective amount is sufficient to introduce a change into the HBV genome within cells (e.g., in vitro or in vivo cells) of the base editor of the present invention. In one embodiment, the effective amount is the amount of the base editor required to achieve a therapeutic effect (e.g., to reduce or control HBV infection). Such a therapeutic effect need not be sufficient to alter the HBV genome in all cells of the subject, tissue, or organ, but may be sufficient to alter the HBV genome in about 1%, 5%, 10%, 25%, 50%, 75% or more of the cells present in the subject, tissue, or organ. In one embodiment, the effective amount is sufficient to improve one or more symptoms of HBV.

[0075] In some embodiments, the effective amount of the fusion proteins provided herein, e.g., the nCas9 domain and the deaminase domain (e.g., adenosine deaminase, cytidine deaminase) comprising the nucleic acid base editor refers to an amount sufficient for the nucleic acid base editor described herein to specifically bind and induce editing of the target site to be edited. As will be understood by those skilled in the art, the effective amount of an agent, e.g., a fusion protein, can vary depending on various factors such as, for example, the specific genome or target site to be edited, the cells or tissues targeted, and / or the agent being used.

[0076] In some embodiments, the fusion proteins provided herein, e.g., the nCas9 An effective amount of the fusion protein comprising the zinc finger and deaminase domain can refer to an amount of the fusion protein sufficient for the fusion protein to specifically bind and induce editing of the target site to be edited. As will be understood by those skilled in the art, the effective amount of an agent, e.g., the fusion protein, can vary depending on, for example, the desired biological response, e.g., the specific genome or target site to be edited, the cell or tissue to be targeted, and / or various factors such as the agent being used.

[0077] "Fragment" means a portion of a polypeptide or nucleic acid molecule. This portion preferably contains at least 10%, 20%, 30%, 40%, 50%, 6 0%, 70%, 80%, or 90% of the full length of the reference nucleic acid molecule or polypeptide. The fragment can contain 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids.

[0078] "Guide RNA" or "gRNA" means a polynucleotide that is specific for a target sequence and can complex with a polynucleotide programmable nucleotide-binding domain protein (e.g., Cas9 or Cpf1). In one embodiment, the guide polynucleotide is a guide RNA (gRNA). The gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule may be referred to as a single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to guide RNAs that exist as a single molecule or as a complex of two or more molecules. Typically, as a single RNA species it can exist, and the gRNA that exists as a single RNA molecule may be called a single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to guide RNAs that exist as a single molecule or as a complex of two or more molecules. Usually, as a single RNA species it can exist, and the gRNA that exists as a single RNA molecule may be called a single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to guide RNAs that exist as a single molecule or as a complex of two or more molecules. Usually, as a single RNA species it can exist, and the gRNA that exists as a single RNA molecule may be called a single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to guide RNAs that exist as a single molecule or as a complex of two or more molecules. Usually, as a single RNA species The gRNA that exists has two domains: (1) a domain that shares homology with the target nucleic acid (e.g., C a domain that induces binding of the Cas9 complex to the target), and (2) a domain that binds to the Cas9 protein . In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to tracrRNA as described in J inek et al., Science 337:816-821 (2012), the entire content of which is incorporated herein by reference. Other examples of gRNA (e.g., those containing domain 2) can be found in US20160208288 entitled "Switchable Cas9 Nucleases and Uses Thereof" and US9 ,737,604 entitled "Delivery System For Functional Nucleases", the entire content of each of which is incorporated herein by reference in its entirety . In some embodiments, the gRNA contains two or more domains (1) and (2) and can be referred to as "extended g RNA". Extended gRNA binds to two or more Cas9 proteins and binds to the target nucleic acid at two or more different regions as described herein. The gRNA contains a nucleotide sequence complementary to the target site, which mediates the binding of the nuclease / RNA complex to the target site and provides the sequence specificity of the nuclease:RNA complex . The "HBV polymerase protein" has at least about 95% identity to the wild-type HBV polymerase amino acid sequence or a fragment thereof that functions in hepatitis B virus infection . . . . .

[0079] . . It means polypeptide. In one embodiment, the HBV polymerase is encoded by HBV genotypes A, B, C, D, E, F , G, or H. In one embodiment, the HBV polymerase amino acid sequence is provided by UniPro accession number Q8B5R0-1, and its sequence is shown below. JPEG2025102768000004.jpg78165

[0080] Mutations within the HBV polymerase include E24G, L25F, P26F, R27C, V48A, V48I, S382F, V378 I, V378A, V379I, V379A, L377F, D380G, D380N, F381P, R376G, A422T, F423P, A432V, M433V, P434S, D540G, A688V, D689G, A717T, E718K, P713S, P713L, or L719P, among others.

[0081] Other exemplary HBV DNA polymerases include, for example, NCBI accession number A B59972.1 having the following sequence. MPLSYQHFRKLLLLDDEAGPLEEELPRLADEGLNRRVAEDLNLG NLNVSIPWTHKVGNFTGLYSSTVPVFNPHWKTPSFPNIHLHQDIIKKCEQFVGPLTVN EKRRLQLIMPARFYPKVTKYLPLDKGIKPYYPEHLVNHYFQTRHYLHTLWKAGILYKR ETTHSASFCGSPYSWEQDLQHGAESFHQQSSGILSRPPVGSSLQSKHSKSRLGLQSQQ GHLARRQQGRSWSIRAGFHPTARRPFGVEPSGSGHTTNFASKSASCLHQSPDRKAAYP AVSTFEKHSSSGHAVEFHNLSPNSARSQSERPVFPCWWLQFRSSKPCSDYCLSLIVNL​ LEDWGPCAEHGEHHIRIPRTPSRVTGGVFLVDKNPHNTAESRLVVDFSQFSRGNYRVS WPKFAVPNLQSLTNLLSSNLSWLSLDVSAAFYHLPLHPAAMPHLLVGSSGLSRYVARL SSNSRILNHQHGTMPNLHDYCSRNLYVSLLLLYQTFGRKLHLYSHPIILGFRKIPMGV GLSPFLLAQFTSAICSVVRRAFPHCLAFSYMDDVVLGAKSVQHLESLFTAVTNFLLSL GIHLNPNKTKRWGYSLNFMGYVIGSYGSLPQEHIIQKIKECFRKLPINRPIDWKVCQR IVGLLGFAAPFTQCGYPALMPLYACIQSKQAFTFSPTYKAFLCKQYLNLYPVARQRPG LCQVFADATPTGWGLVMGHQRVRGTFSAPLPIHTAELLAACFARSRSGANIIGTDNSV VLSRKYTSYPWLLGCAANWILRGTSFVYVPSALNPADDPSRGRLGLSRPLLRLPFRPT TGRTSLYADSPSVPSHLPDRVHFASPLHVAWRPP

[0082] The "HBV polymerase gene" means a polynucleotide encoding HBV polymerase. It means.

[0083] The "hepatitis B surface antigen (HBsAg) polypeptide" means an antigenic protein that functions in HBV virus infection and has at least about 85% identity to NCBI accession number AAB59969.1 or a fragment thereof. An exemplary HBsAg amino acid sequence is shown below: or a fragment thereof. An exemplary HBsAg amino acid sequence is shown below: or a fragment thereof. An exemplary HBsAg amino acid sequence is shown below: MENITSGFLGPLLVLQAGFFLLTRILTIPQSLDSWWTSLNFLGGTTVCLGQNSQSPTSNHSPTSCPPTCPGYRWMCLRRF IIFLFILLLCLIFLLVLLDYQGMLPVCPLIPGSSTTSTGPCRTCMTTAQGTSMYPSCCCTKPSDGNCTCIPIPSSWAFGK FLWEWASARFSWLSLLVPFVQWF

[0084] "HbsAg polynucleotide" means a polynucleotide encoding the HBsAg protein. means.

[0085] "HBV X protein" means a polypeptide or a fragment thereof having at least about 85% identity to NCBI accession number AAB59970.1 and functioning in HBV virus infection. For example, the exemplary amino acid sequences are shown below: 1 maarlccqld pardvlclrp vgaescgrpf sgslgtlssp spsavptdhg ahlslrglpv 61 cafssagpca lrftsarrme ttvnahrmlp kvlhkrtlgl samsttdlea yfkdclfkdw 121 eelgeeirlk vfvlggcrhk lvcapapcnf ftsa

[0086] "Core antigen precursor" means a polypeptide or a fragment thereof having at least about 85% identity to NCBI accession number AAB59971.1 and functioning in HBV virus infection. means.

[0087] "HBV core protein" means a polypeptide having at least about 95% identity to the amino acid sequence of the wild-type HBV core protein or a fragment thereof. In one embodiment, the HBV core protein functions in hepatitis B virus infection. In one embodiment, the HBV core protein functions in hepatitis B virus infection. In one embodiment, the HBV core protein functions in hepatitis B virus infection. ​​The protein is encoded by HBV genotype A, B, C, D, E, F, G, or H. In one embodiment, the amino acid sequence of the HBV core protein is shown by the following NCBI GenBank accession number A XG50928.1: 1 mdidpykefg asvellsflp sdffpsirdl ldtasalyre alespehcsp hhtalrqail 61 cwgelmnlat wvgsnledpa srelvvsyvn vnmglkirql lwfhiscltf gretvleylv 121 sfgvwirtpp ayrppnapil stlpettvvr rrgrsprrrt psprrrrsqs prrrrsqsre 181 sqc

[0088] "HBV X protein" means a polynucleotide encoding the HBV X protein.

[0089] "HBV X protein (genotype B)" means a polypeptide having at least about 95% identity to the amino acid sequence of the wild-type HBV genotype B X protein or a fragment thereof. In one embodiment, the HBV X protein functions in hepatitis B virus infection. In one embodiment, the amino acid sequence of the HBV genotype B X protein is shown by the following NCBI GenBank accession number BAQ95575.1: 1 maarlccqld pardvlclrp vgaesrgrpl pgplgalppa sppvvpsdhg ahlslrglpv 61 cafssagpca lrftsarrme ttvnahrmlp kvlhkrtlgl samsttdlea yfkdclfkdw 121 eelgeeirlk vfvlggcrhk lvcapapcnf ftsa​​

[0090] "HBV X protein (genotype C)" refers to the amino acid sequence of the wild-type HBV genotype CX protein. By this is meant a polypeptide having at least about 95% identity to the sequence or a fragment thereof. In one embodiment, the HBV X protein functions in Hepatitis B virus infection. In the form of the amino acid sequence of the HBV genotype CX protein, the following NCBI GenBank accession numbers are used: Provided in No. BAQ95563.1: 1 maarvccqld pardvlclrp vgaesrgrpv sgpfgplpsp sssavpadyg ahlslrglpv 61 cafssagpca lrftsarrme ttvnahrmlp kvlhkrtlgl samsttdlea yfkdclfkdw 121 eelgeeirlk vfvlggcrhk lvcapapcnf ftsa

[0091] "HBV S protein" refers to an amino acid sequence corresponding to the wild-type HBV S protein or a fragment thereof. In one embodiment, the HBV S polypeptide has at least about 95% identity to the HBV S polypeptide. The protein functions in Hepatitis B virus infection. In one embodiment, the HBV S protein The protein is encoded by an HBV A, B, C, D, E, F, G, or H genotype. The amino acid sequence of HBV S protein is shown below under NCBI GenBank accession number ABV02793.1. Offered in: 1 menttsgflg pllvlqagff lltrnltipq sldswwtsln flggaptcpg qnsqsptsnh 61 sptscppicp gyrwmclrrf iiflfilllc lifllvlldy qgmlpvcpll pgtsttstgp 121 cktctipaqg tsmfpsccct kpsdgnctci pipsswafar flwewasvrf swlsllvpfv 181 qwfvglsptv wlsviwmmwy wgpslynils pflpllpiff clwvyi

[0092] The complete genome of hepatitis B virus subtype ayw, HBV polymerase, HBsAg protein, HBV X protein, and a polynucleotide encoding the core antigen precursor, a complete geno m is provided by GenBank accession number U95551.1 and its sequence is shown below: 1 aattccacaa cctttcacca aactctgcaa gatcccagag tgagaggcct gtatttccct 61 gctggtggct ccagttcagg agcagtaaac cctgttccga ctactgcctc tcccttatcg 121 tcaatcttct cgaggattgg ggaccctgcg ctgaacatgg agaacatcac atcaggattc 181 ctaggacccc ttctcgtgtt acaggcgggg tttttcttgt tgacaagaat cctcacaata 241 ccgcagagtc tagactcgtg gtggacttct ctcaattttc tagggggaac taccgtgtgt 301 cttggccaaa attcgcagtc cccaacctcc aatcactcac caacctcctg tcctccaact 361 tgtcctggtt atcgctggat gtgtctgcgg cgttttatca tcttcctctt catcctgctg 421 ctatgcctca tcttcttgtt ggttcttctg gactatcaag gtatgttgcc cgtttgtcct 481 ctaattccag gatcctcaac caccagcacg ggaccatgcc gaacctgcat gactactgct 541 caaggaacct ctatgtatcc ctcctgttgc tgtaccaaac cttcggacgg aaattgcacc 601 tgtattccca tcccatcatc ctgggctttc ggaaaattcc tatgggagtg ggcctcagcc 661 cgtttctcct ggctcagttt actagtgcca tttgttcagt ggttcgtagg gctttccccc 721 actgtttggc tttcagttat atggatgatg tggtattggg ggccaagtct gtacagcatc 781 ttgagtccct ttttaccgct gttaccaatt ttcttttgtc tttgggtata catttaaacc 841 ctaacaaaac aaagagatgg ggttactctc tgaattttat gggttatgtc attggaagtt 901 atgggtcctt gccacaagaa cacatcatac aaaaaatcaa agaatgtttt agaaaacttc 961 ctattaacag gcctattgat tggaaagtat gtcaacgaat tgtgggtctt ttgggttttg 1021 ctgccccatt tacacaatgt ggttatcctg cgttaatgcc cttgtatgca tgtattcaat 1081 ctaagcaggc tttcactttc tcgccaactt acaaggcctt tctgtgtaaa caatacctga 1141 acctttaccc cgttgcccgg caacggccag gtctgtgcca agtgtttgct gacgcaaccc 1201 ccactggctg gggcttggtc atgggccatc agcgcgtgcg tggaaccttt tcggctcctc 1261 tgccgatcca tactgcggaa ctcctagccg cttgttttgc tcgcagcagg tctggagcaa 1321 acattatcgg gactgataac tctgttgtcc tctcccgcaa atatacatcg tatccatggc 1381 tgctaggctg tgctgccaac tggatcctgc gcgggacgtc ctttgtttac gtcccgtcgg 1441 cgctgaatcc tgcggacgac ccttctcggg gtcgcttggg actctctcgt ccccttctcc 1501 gtctgccgtt ccgaccgacc acggggcgca cctctcttta cgcggactcc ccgtctgtgc 1561 cttctcatct gccggaccgt gtgcacttcg cttcacctct gcacgtcgca tggagaccac 1621 cgtgaacgcc caccgaatgt tgcccaaggt cttacataag aggactcttg gactctctgc 1681 aatgtcaacg accgaccttg aggcatactt caaagactgt ttgtttaaag actgggagga 1741 gttgggggag gagattagat taaaggtctt tgtactagga ggctgtaggc ataaattggt 1801 ctgcgcacca gcaccatgca actttttcac ctctgcctaa tcatctcttg ttcatgtcct 1861 actgttcaag cctccaagct gtgccttggg tggctttggg gcatggacat cgacccttat 1921 aaagaatttg gagctactgt ggagttactc tcgtttttgc cttctgactt ctttccttca 1981 gtacgagatc ttctagatac cgcctcagct ctgtatcggg aagccttaga gtctcctgag 2041 cattgttcac ctcaccatac tgcactcagg caagcaattc tttgctgggg ggaactaatg 2101 actctagcta cctgggtggg tgttaatttg gaagatccag catctagaga cctagtagtc 2161 agttatgtca acactaatat gggcctaaag ttcaggcaac tcttgtggtt tcacatttct 2221 tgtctcactt ttggaagaga aaccgttata gagtatttgg tgtctttcgg agtgtggatt 2281 cgcactcctc cagcttatag accaccaaat gcccctatcc tatcaacact tccggaaact 2341 actgttgtta gacgacgagg caggtcccct agaagaagaa ctccctcgcc tcgcagacga 2401 aggtctcaat cgccgcgtcg cagaagatct caatctcggg aacctcaatg ttagtattcc 2461 ttggactcat aaggtgggga actttactgg tctttattct tctactgtac ctgtctttaa 2521 tcctcattgg aaaacaccat cttttcctaa tatacattta caccaagaca ttatcaaaaa 2581 atgtgaacag tttgtaggcc cacttacagt taatgagaaa agaagattgc aattgattat 2641 gcctgctagg ttttatccaa aggttaccaa atatttacca ttggataagg gtattaaacc 2701 ttattatcca gaacatctag ttaatcatta cttccaaact agacactatt tacacactct 2761 atggaaggcg ggtatattat ataagagaga aacaacacat agcgcctcat tttgtgggtc 2821 accatattct tgggaacaag atctacagca tggggcagaa tctttccacc agcaatcctc 2881 tgggattctt tcccgaccac cagttggatc cagccttcag agcaaacaca gcaaatccag 2941 attgggactt caatcccaac aaggacacct ggccagacgc caacaaggta ggagctggag 3001 cattcgggct gggtttcacc ccaccgcacg gaggcctttt ggggtggagc cctcaggctc 3061 agggcatact acaaactttg ccagcaaatc cgcctcctgc ctccaccaat cgccagacag 3121 gaaggcagcc taccccgctg tctccacctt tgagaaacac tcatcctcag gccatgcagt 3181 gg

[0093] The term "heterodimer" refers to the wild-type TadA domain and variants of the TadA domain (e.g., T adA*8) or two variant TadA domains (e.g., TadA*7.10 and TadA*8, or two It means a fusion protein containing two domains such as the TadA*8 domain.

[0094] "Hybridization" means a hydrogen bond, which is a Watson-Crick type, Hoogsteen type, or reverse Hoogsteen type hydrogen bond. between complementary nucleobases. For example, adenine and thymine are complementary nucleobases that pair through the formation of a hydrogen bond.

[0095] The terms "isolated," "purified," or "biologically pure" refer to substances that are free to varying degrees from the components that are normally associated with them in their natural state. "Isolating" indicates the degree of separation from the original source or surroundings. "Purifying" indicates a higher degree of separation than isolation. A "purified" or "biologically pure" protein contains other substances to such a limited extent that impurities do not substantially affect or cause other adverse effects on the protein's biological properties. That is, the nucleic acids or peptides of the present invention are purified when they are produced by recombinant DNA technology and do not contain cell material, viral material, or culture medium, or when they are chemically synthesized and do not substantially contain chemical precursors or other chemical substances. Purity and homogeneity are usually measured using analytical chemistry techniques such as polyacrylamide gel electrophoresis or high performance liquid chromatography. The term "purified" may indicate that a nucleic acid or protein yields essentially one band in an electrophoretic gel. In the case of proteins that can be subjected to modifications such as phosphorylation or glycosylation, different modifications can yield different isolated proteins that can be purified separately. ​​​​​​​​​​​​​​​

[0096] "Isolated polynucleotide" means a nucleic acid (e.g., DNA) that does not contain genes adjacent to the gene in the natural genome of the organism from which the nucleic acid molecule of the present invention is derived. Thus, this term includes, for example, recombinant DNA incorporated into a vector; recombinant DNA incorporated into an autonomously replicating plasmid or virus; or recombinant DNA incorporated into the genome of a prokaryotic or eukaryotic organism; or recombinant DNA that exists as a separate molecule independent of other sequences (e.g., cDNA or genomic or cDNA fragments generated by PCR or restriction endonuclease digestion). Further, this term includes RNA molecules transcribed from a DNA molecule, and recombinant DNA that is part of a hybrid gene encoding additional polypeptide sequences. "Increase" means a positive change of at least 10%, 25%, 50%, 75%, or 100%. The terms "inhibitor of base repair", "base repair inhibitor", "IBR" or their grammatical equivalents refer to proteins that can inhibit the activity of nucleic acid repair enzymes, such as base excision repair enzymes. In some embodiments, IBR is an inhibitor of inosine base excision repair. Exemplary inhibitors of base repair include inhibitors of APE1, Endo III, Endo IV, Endo V, Endo VIII, Fpg, hOGG1, hNEIL1, T7 Endol, T4PDG, UDG, hSMUG1, and hAAG. In some embodiments, the base repair inhibitor is an inhibitor of Endo V or hAAG.

[0097]

[0098] 。In some embodiments, the IBR is an inhibitor of Endo V or hAAG. In some embodiments, the IBR is catalytically inactive Endo V or catalytically inactive hAAG. In some embodiments, the base repair inhibitor is catalytically inactive Endo V or catalytically inactive hAAG. In some embodiments, the base repair inhibitor is uracil glycosylase inhibitor (UGI). UGI is a protein that can inhibit the base excision repair enzyme of uracil DNA glycosylase. In some embodiments, the UGI domain comprises wild-type UGI or a fragment of wild-type UGI. In some embodiments, the UGI proteins provided herein include fragments of UGI and proteins homologous to UGI or UGI fragments. In some embodiments, the base repair inhibitor is an inhibitor of inosine base excision repair. In some embodiments, the base repair inhibitor is a "catalytically inactive inosine-specific nuclease" or an "inactive inosine-specific nuclease". Without wishing to be bound by a particular theory, catalytically inactive inosine glycosylase (e.g., alkyladenine glycosylase (AAG)) can bind to inosine, but cannot create abasic sites or remove inosine, and as a result, the newly formed inosine moiety is sterically blocked from the DNA damage / repair machinery. In some embodiments, the catalytically inactive inosine-specific nuclease can bind to inosine in nucleic acids but does not cleave the nucleic acids. Non-limiting exemplary catalytically inactive inosine-specific nucleases include, for example, Claise), and, for example, catalytically inactive endonuclease V (EndoV) derived from E. coli Nuclease). In some embodiments, the catalytically inactive AAG nuclease contains the E125Q mutation or a corresponding mutation in another AAG nuclease.

[0099] An "intein" is a fragment of a protein that can excise itself and ligate the remaining fragments (exteins) by peptide bonds in a process called protein splicing. An intein is also referred to as a "protein intron". The process by which an intein excises itself and ligates the remaining part of the protein is referred to herein as "protein splicing" or "intein-mediated protein splicing". In some embodiments, the intein of the precursor protein (the intein-containing protein before intein-mediated protein splicing) is derived from two genes. Such an intein is referred to herein as a split intein (e.g., split intein-N and split intein-C). For example, in cyanobacteria, DnaE, which is the catalytic subunit a of DNA polymerase III, is encoded by two separate genes, dnaE-n and dnaE-c. The intein encoded by the dnaE-n gene can be referred to herein as "intein-N". The intein encoded by the dnaE-c gene can be referred to herein as "intein-C". Other intein systems can also be used. For example, dnaE intein, Cfa-N ( ... ... ... ... ... ... ...

[0100] ... For example, split intein-N) and Cfa-C (e.g., split intein-C) inte Synthetic inteins based on impairment are described (e.g., incorporated herein by reference Stevens et al., J Am Chem Soc.2016 Feb.24;138 (7):2162-5). In accordance with this disclosure Non-limiting examples of intein pairs that can be used include: Cfa DnaE intein, Spp GyrB inte tein, Spp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB i tein and Cne Prp8 intein (e.g., described in U.S. Patent No. 8 ,394,604, which is incorporated herein by reference).

[0101] Exemplary nucleotide and amino acid sequences of inteins are shown. DnaE intein-N DNA: TGCCTGTCATACGAAACCGAGATACTGACAGTAGAATATGGCCTTCTGCCAATCGGGAAGATTGTGGAGAAACGGATAGA ATGCACAGTTTACTCTGTCGATAACAATGGTAACATTTATACTCAGCCAGTTGCCCAGTGGCACGACCGGGGAGAGCAGG AAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAGGGCCACTAAGGACCACAAATTTATGACAGTCGATGGCCAG ATGCTGCCTATAGACGAAATCTTTGAGCGAGAGTTGGACCTCATGCGAGTTGACAACCTTCCTAAT DnaE intein-N protein: CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDR GEQEVFEYCLEDGSLIRATKDHKFMTVDG QMLPIDEIFERELDLMRVDNL PN DnaE intein - cDNA: ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGA TATTGGAGTCGAAAGAGATCACAACTTTG CTCTGAAGAACGGATTCATAG CTTCTAAT Intein - C:MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN Cfa - N DNA: TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGCCTATTGGAAAGATTGTCGAAGAGAGAATTGA ATGCACAGTATATACTGTAGACAAGAATGGTTTCGTTTACACACAGCCCATTGCTCAATGGCACAATCGCGGCGAACAAG AAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACGAGCAACTAAAGATCATAAATTCATGACCACTGACGGGCAG ATGTTGCCAATAGATGAGATATTCGAGCGGGGCTTGGATCTCAAACAAGTGGATGGATTGCCA Cfa - N protein: CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQ MLPIDEIFERGLDLKQVDGLP Cfa - C DNA: ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAGGAAAGTAAAGATAATATCTCGAAAAAGTCT TGGTACCCAAAATGTCTATGATATTGGAGTGGAGAAAGATCACAACTTCCTTCTCAAGAACGGTCTCGTAGCCAGCAAC Cfa-C protein: MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLVASN

[0102] In order to join the N-terminal part of split Cas9 and the C-terminal part of split Cas9, Intein-N and intein-C are the N-terminal portion of split Cas9 and split Cas1, respectively. 9. For example, in some embodiments, intein-N may be fused to the C-terminal portion of The N-terminal portion of split Cas9 was fused to the C-terminus, i.e., N--[N-terminus of split Cas9 In some embodiments, the intein-C is formed as follows: is fused to the N-terminus of the C-terminal portion of split Cas9, i.e., N-[intein-C]-[ssp The C-terminal part of the split Cas9 is used to form a C structure. For example, intein-mediated protein splicing to bind split Cas9 The binding mechanism is described, for example, in Shah et al., Chem Sci. 2014;5 ( Inteins are known in the art, as described in, for example, U.S. Pat. No. 6,333,446. Methods for using the method are known in the art, see, for example, WO2014004336, WO2017132580, US No. 20150344549, and US20180127780, each of which is incorporated herein by reference in its entirety. No. 6,399,433, which is incorporated herein by reference.

[0103] An "isolated polypeptide" refers to a polypeptide of the invention that is separated from components that naturally accompany it. Generally, a polypeptide is at least 60% by weight of the and isolated when it does not contain proteins and natural organic molecules that are naturally associated with it is present. Preferably, the preparation is at least 75% by weight, more preferably at least 90% by weight, most preferably at least 99% by weight of the polypeptide of the present invention. The isolated polypeptide of the present invention can be obtained, for example, by extraction from a natural source, by expression of a recombinant nucleic acid encoding such a polypeptide; or by chemically synthesizing the protein . The purity can be measured by appropriate methods, such as column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis . As used herein, the term "linker" refers to a covalent linker (e.g., a covalent bond), a non-covalent linker, a chemical group, or a molecule that links two molecules or moieties, such as two components of a protein complex

[0104] or ribonucleo complex, or two domains of a fusion protein, e.g., a polynucleotide-programmable DNA binding domain (e.g., dCas9) and a deaminase domain ((e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase)). The linker can link various components or different parts of a base editing system. For example, in some embodiments, the linker can link the guide polynucleotide binding domain of a polynucleotide-programmable nucleotide binding domain to the catalytic domain of a deaminase . In some embodiments, the linker can link a CRISPR polypeptide to a deaminase . In some embodiments, the linker can link Cas9 to a deaminase . ​​​-ase can be conjugated. In some embodiments, the linker can conjugate dCas9 and deaminase -ase can be conjugated. In some embodiments, the linker can conjugate nCas9 and deaminase -ase can be conjugated. In some embodiments, the linker can conjugate the guide polynu cleotide and deaminase. In some embodiments, the linker can connect the deaminating component of the base editing system and the polynucleotide programmable nucleotide binding component In some embodiments, the linker can connect the deamino part of the deaminating component of the base editing system and the polynucleotide programmable nucleotide binding component In some embodiments, the linker can connect the deaminating component of the base editing system RNA binding part of the deaminating component and the RNA binding part of the polynucleotide programmable nucleotide binding component The linker is located between, adjacent to, or connected to two groups, molecules, or other moieties via covalent or non-covalent interactions, respectively, and thus can connect the two In some embodiments, the linker can be an organic molecule , group, polymer, or chemical moiety. In some embodiments, the linker can be a poly nucleotide. In some embodiments, the linker can be a DNA linker . In some embodiments, the linker can be an RNA linker. In some embodiments the linker can include an aptamer that can bind to a ligand. In some embodiments, the ligand can be a carbohydrate, peptide, protein, or nucleic acid . In some embodiments, the linker can include an aptamer derived from a riboswitch . In some embodiments, the linker can include an aptamer derived from a riboswitch It is. The riboswitch from which the aptamer is derived is the theophylline riboswitch, thiamine pyrophosphate (TPP) riboswitch, adenosylcobalamin (AdoCbl) riboswitch, S-adenosylmethionine (SAM) riboswitch, SAH riboswitch, flavin mononucleotide (FMN) riboswitch, tetrahydrofolate riboswitch, lysine riboswitch, glycine riboswitch, purine riboswitch, GlmS riboswitch, or prequeuosine 1 (PreQ1) riboswitch. In some embodiments, the linker may include an aptamer bound to a protein domain such as a polypeptide or a polypeptide ligand. In some embodiments, the polypeptide ligand may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a stella alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the polypeptide ligand may be part of a base editing system component. For example, the nucleic acid base editing component may include a deaminase domain and an RNA recognition motif. ine (TPP) riboswitch, adenosylcobalamin (AdoCbl) riboswitch, S-adenosylmethionine (SAM) riboswitch, SAH riboswitch, flavin mononucleotide (FMN) riboswitch, tetrahydrofolate riboswitch, lysine riboswitch, glycine riboswitch, purine riboswitch, GlmS riboswitch, or prequeuosine 1 (PreQ1) riboswitch. In some embodiments, the linker may include an aptamer bound to a protein domain such as a polypeptide or a polypeptide ligand. In some embodiments, the polypeptide ligand may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a stella alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the polypeptide ligand may be part of a base editing system component. For example, the nucleic acid base editing component may include a deaminase domain and an RNA recognition motif. ionine (SAM) riboswitch, SAH riboswitch, flavin mononucleotide (FMN) riboswitch, tetrahydrofolate riboswitch, lysine riboswitch, glycine riboswitch, purine riboswitch, GlmS riboswitch, or prequeuosine 1 (PreQ1) riboswitch. In some embodiments, the linker may include an aptamer bound to a protein domain such as a polypeptide or a polypeptide ligand. In some embodiments, the polypeptide ligand may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a stella alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the polypeptide ligand may be part of a base editing system component. For example, the nucleic acid base editing component may include a deaminase domain and an RNA recognition motif. switch, SAH riboswitch, flavin mononucleotide (FMN) riboswitch, tetrahydrofolate riboswitch, lysine riboswitch, glycine riboswitch, purine riboswitch, GlmS riboswitch, or prequeuosine 1 (PreQ1) riboswitch. In some embodiments, the linker may include an aptamer bound to a protein domain such as a polypeptide or a polypeptide ligand. In some embodiments, the polypeptide ligand may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a stella alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the polypeptide ligand may be part of a base editing system component. For example, the nucleic acid base editing component may include a deaminase domain and an RNA recognition motif. e riboswitch, tetrahydrofolate riboswitch, lysine riboswitch, glycine riboswitch, purine riboswitch, GlmS riboswitch, or prequeuosine 1 (PreQ1) riboswitch. In some embodiments, the linker may include an aptamer bound to a protein domain such as a polypeptide or a polypeptide ligand. In some embodiments, the polypeptide ligand may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a stella alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the polypeptide ligand may be part of a base editing system component. For example, the nucleic acid base editing component may include a deaminase domain and an RNA recognition motif. riboswitch, lysine riboswitch, glycine riboswitch, purine riboswitch, GlmS riboswitch, or prequeuosine 1 (PreQ1) riboswitch. In some embodiments, the linker may include an aptamer bound to a protein domain such as a polypeptide or a polypeptide ligand. In some embodiments, the polypeptide ligand may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a stella alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the polypeptide ligand may be part of a base editing system component. For example, the nucleic acid base editing component may include a deaminase domain and an RNA recognition motif. riboswitch, glycine riboswitch, purine riboswitch, GlmS riboswitch, or prequeuosine 1 (PreQ1) riboswitch. In some embodiments, the linker may include an aptamer bound to a protein domain such as a polypeptide or a polypeptide ligand. In some embodiments, the polypeptide ligand may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a stella alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the polypeptide ligand may be part of a base editing system component. For example, the nucleic acid base editing component may include a deaminase domain and an RNA recognition motif. riboswitch, purine riboswitch, GlmS riboswitch, or prequeuosine 1 (PreQ1) riboswitch. In some embodiments, the linker may include an aptamer bound to a protein domain such as a polypeptide or a polypeptide ligand. In some embodiments, the polypeptide ligand may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a stella alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the polypeptide ligand may be part of a base editing system component. For example, the nucleic acid base editing component may include a deaminase domain and an RNA recognition motif. riboswitch, GlmS riboswitch, or prequeuosine 1 (PreQ1) riboswitch. In some embodiments, the linker may include an aptamer bound to a protein domain such as a polypeptide or a polypeptide ligand. In some embodiments, the polypeptide ligand may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a stella alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the polypeptide ligand may be part of a base editing system component. For example, the nucleic acid base editing component may include a deaminase domain and an RNA recognition motif. riboswitch, or prequeuosine 1 (PreQ1) riboswitch. In some embodiments, the linker may include an aptamer bound to a protein domain such as a polypeptide or a polypeptide ligand. In some embodiments, the polypeptide ligand may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a stella alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the polypeptide ligand may be part of a base editing system component. For example, the nucleic acid base editing component may include a deaminase domain and an RNA recognition motif. riboswitch. In some embodiments, the linker may include an aptamer bound to a protein domain such as a polypeptide or a polypeptide ligand. In some embodiments, the polypeptide ligand may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a stella alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the polypeptide ligand may be part of a base editing system component. For example, the nucleic acid base editing component may include a deaminase domain and an RNA recognition motif. In some embodiments, the polypeptide ligand may be part of a base editing system component. For example, the nucleic acid base editing component may include a deaminase domain and an RNA recognition motif. In some embodiments, the polypeptide ligand may be part of a base editing system component. For example, the nucleic acid base editing component may include a deaminase domain and an RNA recognition motif.

[0105] In some embodiments, the linker may be one amino acid or a plurality of amino acids (e.g., a peptide or a protein). In some embodiments, the linker may have a length of about 5 to 100 amino acids, such as about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 90, or 90 to 100 amino acids. In some embodiments, the linker may be one amino acid or a plurality of amino acids (e.g., a peptide or a protein). In some embodiments, the linker may have a length of about 5 to 100 amino acids, such as about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 90, or 90 to 100 amino acids. In some embodiments, the linker may be one amino acid or a plurality of amino acids (e.g., a peptide or a protein). In some embodiments, the linker may have a length of about 5 to 100 amino acids, such as about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 90, or 90 to 100 amino acids. In some embodiments, the linker may be one amino acid or a plurality of amino acids (e.g., a peptide or a protein). In some embodiments, the linker may have a length of about 5 to 100 amino acids, such as about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 90, or 90 to 100 amino acids. It can be the length of the amino acid. In some embodiments, the linker has a length of about 100-150, 150 -200, 200-250, 250-300, 300-350, 350-400, 400-450, or 450-500 amino acids can be. Longer or shorter linkers are also contemplated.

[0106] In some embodiments, the linker binds the gRNA binding domain of an RNA programmable nuclease comprising a Cas9 nuclease domain to the catalytic domain of a nucleic acid editing protein (e.g., cytidine deaminase or adenosine deaminase). In some embodiments, the linker binds dCas9 and a nucleic acid editing protein. For example, the linker is located between, adjacent to, or connected to two groups, molecules, or other moieties via a covalent bond, thus connecting the two. In some embodiments, the linker is one amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker has a length of 5-200 amino acids, e.g., a length of 5, 6, 7, 8 rammable nuclease's gRNA binding domain and the catalytic domain of a nucleic acid editing protein (e.g., cytidine deaminase or adenosine deaminase). In some embodiments, the linker binds dCas9 and a nucleic acid editing protein. For example, the linker is located between, adjacent to, or connected to two groups, molecules, or other moieties via a covalent bond, thus connecting the two. In some embodiments, the linker is located between, adjacent to, or connected to two groups, molecules, or other moieties via a covalent bond, thus connecting the two. In some embodiments, the linker is located between, adjacent to, or connected to two groups, molecules, or other moieties via a covalent bond, thus connecting the two. In some embodiments, the linker is one amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is one amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker has a length of 5-200 amino acids, e.g., a length of 5, 6, 7, 8 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 1

[0107] In some embodiments, the domain of the base editor is SGGSSGSETPGTSESATPESSGGS, SGGS SGGSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPG It contains the amino acid sequence of SPAGSPTSTEEGTSTE PSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS Fuse through a linker. In some embodiments, the domain of the nucleic acid base assembly is fused through a linker containing the amino acid sequence SGSETPGTSESATPES, which may also be referred to as an XTEN linker. In some embodiments, the linker is 24 amino acids in length. In some embodiments, the linker contains the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In some embodiments, the linker is 40 amino acids in length. In some embodiments, the linker contains the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS. In some embodiments, the linker is 64 amino acids in length. In some embodiments, the linker contains the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGS SGGS In some embodiments, the linker is 92 amino acids in length. In some embodiments, the linker contains the amino acid sequence PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSP TSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS.

[0108] "Marker" means any protein or polynucleotide whose expression level or activity changes in relation to a disease or disorder.

[0109] ​ As used herein, the term "mutation" refers to a substitution of a residue within a sequence, e.g., a nucleic acid or an amino acid, by another residue of the sequence, or a deletion or insertion of one or more residues within the sequence. A mutation is typically described herein by identifying the original residue, followed by identifying the position of the residue within the sequence, and then identifying the newly substituted residue. Various methods for making the amino acid substitutions (mutations) provided herein are well known in the art and are described, for example, in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4 th ed., Cold Spring th Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)). In some embodiments, the base editors of the present disclosure can efficiently generate any "intended mutation" within a nucleic acid (e.g., a nucleic acid within a target genome) without generating a significant number of unintended mutations, e.g., unintended point mutations. In some embodiments, the intended mutation is a mutation generated by a specific base editor (e.g., a cytidine base editor or an adenosine base editor) that binds to a guide polynucleotide (e.g., a gRNA) specifically designed to generate the intended mutation. g Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)). In some embodiments, the base editors of the present disclosure can efficiently generate any "intended mutation" within a nucleic acid (e.g., a nucleic acid within a target genome) without generating a significant number of unintended mutations, e.g., unintended point mutations. In some embodiments, the intended mutation is a mutation generated by a specific base editor (e.g., a cytidine base editor or an adenosine base editor) that binds to a guide polynucleotide (e.g., a gRNA) specifically designed to generate the intended mutation. Typically, mutations created or identified within a sequence (e.g., an amino acid sequence described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain the mutation. A person of ordinary skill in the art will readily understand how to determine the position of mutations in amino acid and nucleic acid sequences relative to a reference sequence. Typically, mutations created or identified within a sequence (e.g., an amino acid sequence described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain the mutation. A person of ordinary skill in the art will readily understand how to determine the position of mutations in amino acid and nucleic acid sequences relative to a reference sequence. Typically, mutations created or identified within a sequence (e.g., an amino acid sequence described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain the mutation. A person of ordinary skill in the art will readily understand how to determine the position of mutations in amino acid and nucleic acid sequences relative to a reference sequence. Typically, mutations created or identified within a sequence (e.g., an amino acid sequence described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain the mutation. A person of ordinary skill in the art will readily understand how to determine the position of mutations in amino acid and nucleic acid sequences relative to a reference sequence. Typically, mutations created or identified within a sequence (e.g., an amino acid sequence described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain the mutation. A person of ordinary skill in the art will readily understand how to determine the position of mutations in amino acid and nucleic acid sequences relative to a reference sequence. Typically, mutations created or identified within a sequence (e.g., an amino acid sequence described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain the mutation. A person of ordinary skill in the art will readily understand how to determine the position of mutations in amino acid and nucleic acid sequences relative to a reference sequence.

[0110] Typically, mutations created or identified within a sequence (e.g., an amino acid sequence described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain the mutation. A person of ordinary skill in the art will readily understand how to determine the position of mutations in amino acid and nucleic acid sequences relative to a reference sequence. Typically, mutations created or identified within a sequence (e.g., an amino acid sequence described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain the mutation. A person of ordinary skill in the art will readily understand how to determine the position of mutations in amino acid and nucleic acid sequences relative to a reference sequence. Typically, mutations created or identified within a sequence (e.g., an amino acid sequence described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain the mutation. A person of ordinary skill in the art will readily understand how to determine the position of mutations in amino acid and nucleic acid sequences relative to a reference sequence. Typically, mutations created or identified within a sequence (e.g., an amino acid sequence described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain the mutation. A person of ordinary skill in the art will readily understand how to determine the position of mutations in amino acid and nucleic acid sequences relative to a reference sequence.

[0111] The term "non-conservative mutation" includes, for example, amino acid substitutions between different groups, such as from tryptophan to lysine, or from serine to phenylalanine. In this case, it is preferred that the non-conservative amino acid substitution does not interfere with or inhibit the biological activity of the functional variant. The non-conservative amino acid substitution can enhance the biological activity of the functional variant, and as a result, the biological activity of the functional variant is increased compared to the wild-type protein.

[0112] In some embodiments, the base editor of the present disclosure can efficiently generate "intended mutations" such as point mutations in a nucleic acid (e.g., a nucleic acid within the genome of a subject) without generating a significant number of unintended mutations, such as unintended point mutations. In some embodiments, the intended mutations are mutations generated by a specific base editor (e.g., a cytidine base editor or an adenosine base editor) that binds to a guide polynucleotide (e.g., gRNA) specifically designed to generate the intended mutations.

[0113] Generally, mutations created or identified within a sequence (e.g., an amino acid sequence described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain the mutation. One of ordinary skill in the art will readily understand how to determine the positions of amino acid sequence and nucleic acid sequence mutations relative to the reference sequence.

[0114] The term "non-conservative mutation" includes, for example, amino acid substitutions between different groups, such as from tryptophan to lysine, or from serine to phenylalanine. In this case, the non-conservative amino Preferably, the acid substitutions do not interfere with or inhibit the biological activity of the functional variant. Non-conservative amino acid substitutions can enhance the biological activity of functional variants, As a result, the biological activity of the functional variant is increased compared to the wild-type protein.

[0115] The term "nuclear localization sequence," "nuclear localization signal," or "NLS" refers to a sequence that directs a protein to the cell nucleus. Nuclear localization sequences refer to amino acid sequences that facilitate the import of proteins. For example, P.O. 2000 / 11 / 23, published as WO / 2001 / 038547 on May 31, 2001 Lank et al., International PCT Application No. PCT / EP2000 / 011690, the contents of which are incorporated by reference in their entirety. In another embodiment, the nucleic acid sequence of the present invention is a nucleic acid sequence that is incorporated herein by reference to disclose a suitable nuclear localization sequence. NLSs are described, for example, by Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS is the optimized NLS described above. SEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAA Contains IVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.

[0116] The terms "nucleobase," "nitrogenous base," or "base" are used interchangeably herein. " refers to nitrogen-containing biological compounds that form nucleosides, the building blocks of nucleotides. The ability of nucleic acid bases to form base pairs and stack with each other is what gives rise to ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). Directly generate long-chain helical structures such as acids (DNA). Five nucleobases (adenine (A), cyto sine (C), guanine (G), thymine (T), and uracil (U)) are referred to as primary nucleobases or standard nucleic bases. Adenine and guanine are derived from purines, and cytosine, uracil, and thym ine are derived from pyrimidines. DNA and RNA can also include other (non-primary) modified bases. Non-limiting exemplary modified nucleobases can include hypoxanthine, xanthine, 7-methyl guanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydroxymethylcyt osine. Hypoxanthine and xanthine are generated by the presence of mutagens and can both be generated by deamination (substitution of an amine group with a carbonyl group). Hypox anthine can be modified from adenine. Xanthine can be modified from guanine. Uracil can result from the deamination of cytosine. A "nucleoside" consists of a nucleobase and a five-carbon sugar (either ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-methyluridine (m 5U), deoxyadenosine, deoxyguanosine, thymidine, deoxyuridine, and de oxycytidine. Examples of nucleosides with modified nucleobases include inosine (I), xanthosine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcyt idine (m5C), and pseudouridine (Ψ). A "nucleotide" consists of a nucleobase, a five -carbon sugar (either ribose or deoxyribose), and at least one phosphate group.

[0117] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to nucleic acids that contain nucleobases and acidic moieties. The term refers to a compound containing a nucleotide, such as a nucleoside, a nucleotide, or a polymer of nucleotides. Generally, large nucleic acids, e.g., nucleic acid molecules containing three or more nucleotides, are linear molecules. In this case, adjacent nucleotides are linked to each other via phosphodiester bonds. In some embodiments, a "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or nucleotides). In some embodiments, a "nucleic acid" refers to a nucleic acid sequence consisting of three or more individual nucleotides. As used herein, the term "oligonucleotide" refers to an oligonucleotide chain that includes 5-amino-1, 5-amino-2, 5-diamino-1, 5-di ... "Nucleotide" and "polynucleotide" refer to a polymer of nucleotides (e.g., at least 3 They can be used interchangeably to refer to a string of 1 or 2 nucleotides. In general terms, "nucleic acid" includes RNA as well as single-stranded and / or double-stranded DNA. For example, genomes, transcripts, mRNA, tRNA, rRNA, siRNA, snRNA, plasmids, cosmids, It may occur naturally in the context of a chromosome, chromatid, or other naturally occurring nucleic acid molecule. Acid molecules can be non-naturally occurring molecules, such as recombinant DNA or RNA, artificial chromosomes, modified genomes, etc. It may be a nucleic acid sequence, or a fragment thereof, or a synthetic DNA, RNA, or DNA / RNA hybrid, includes non-natural nucleotides or nucleosides. " and / or similar terms include nucleic acid analogs, e.g., nucleic acids having other than a phosphodiester backbone. Nucleic acids can be purified from natural sources and can be expressed using recombinant expression systems. It can be produced using, for example, optionally purified or chemically synthesized. Depending on, for example, in the case of a chemically synthesized molecule, the nucleic acid may include nucleoside analogs such as bases or sugars with chemically modified bases or analogs, and may include backbone modifications. The nucleic acid sequence is , unless otherwise specified, presented in the 5' to 3' direction. In some embodiments, the nucleic acid is , natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine , deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine ); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, py rolopyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-b romouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5 -propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine , 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine ; chemically modified bases; biologically modified bases (e.g., methylated bases); intercalated bases; modified sugars (2'-e.g., fluoro ribose, ribose, 2'-deoxyribose, arabinose, and hexose), and / or or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages ), or includes them.

[0118] The term "nucleic acid programmable DNA binding protein" or "napDNAbp" may be used interchangeably with "polynucleotide programmable nucleotide binding domain", and a guide nucleic acid or guide polynucleotide that directs napDNAbp to a specific nucleic acid sequence (e.g., ​For example, it refers to a protein that binds to a nucleic acid (e.g., DNA or RNA) such as gRNA. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a poly nucleotide-programmable DNA-binding domain. In some embodiments, the poly nucleotide-programmable nucleotide-binding domain is a polynucleotide-program mable RNA-binding domain. In some embodiments, the polynucleotide-program mable nucleotide-binding domain is a Cas9 protein. The Cas9 protein can bind to a guide RNA that guides the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, napDNAbp is a Cas9 domain, such as nuclease active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Non-limiting examples of nucleic acid-programmable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g , Cas12h, and Cas12i. Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2 , Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, , Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12 , b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy , Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12 , b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy 2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1 , Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf 4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Type II Cas effector protein, Type V Cas effector protein, Type VI Cas effector protein , CARF, DinG, its homologs, or its modified or altered versions. Other programmable DNA-binding proteins are also within the scope of the present disclosure, but they may not be specifically described in the present disclosure. For example, see Makarova et al. ”Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct;1:325-336.doi:10.1089 / crispr.2018.0033; Yan et al., ”Functionally diverse type V CRISPR-Cas systems” Science. 2019 Jan 4;363 (6422):88-91.doi:10.1126 / science.aav7271 (the entire contents of each are hereby incorporated by reference).

[0119] The term "nucleobase editing domain" or "nucleobase editing protein" as used herein " refers to a nucleobase modification in RNA or DNA, e.g., a cytosine (or cytidine) to cytosine (or cytidine) to uracil (or uridine) or thymine (or thymidine), and to adenine (or adenine Deamination of nucleosine (HnO) to hypoxanthine (or inosine), and template-free nucleosine synthesis. It refers to a protein or enzyme that can catalyze the addition and insertion of nucleotides. In embodiments, the nucleobase editing domain is a deaminase domain (e.g., an adenine deaminase adenosine deaminase or adenosine deaminase; or cytidine deaminase or cytosine deaminase In some embodiments, the nucleobase editing domain is a multiple deaminases. A enzyme domain (e.g., adenine deaminase or adenosine deaminase or In some embodiments, the enzyme is a cytidine deaminase or a cytosine deaminase. In some embodiments, the nucleobase-editing domain may be a naturally occurring nucleobase-editing domain. In some embodiments, the nucleobase editing domain is modified from a naturally occurring nucleobase editing domain, or The nucleobase editing domain may be an evolved nucleobase editing domain. The nucleobase editing domain may be a bacterial, human, Derived from any organism, such as a leopard, gorilla, monkey, cow, dog, rat, or mouse possible.

[0120] As used herein, the term "obtaining" in "obtaining a drug" includes synthesizing, purchasing, or obtained by other means.

[0121] As used herein, a "patient" or "subject" refers to a person who has a disease or disorder or Mammals at risk of developing the disease or diagnosed as being affected or suspected of developing the disease In some embodiments, the term "patient" refers to a subject or individual suffering from a disease or disorder. Refers to mammalian subjects with a higher than average likelihood of developing. Exemplary patients can be humans, non-human primates, cats, dogs, pigs, cows, horses, donkeys, llamas, goats, sheep, rodents (e.g., mice, rabbits, rats, or guinea pigs) and other mammals who can benefit from the therapies disclosed herein. Exemplary human patients can be male and / or female. "A patient in need thereof" or "a subject in need thereof" refers, herein, to a patient who has a disease or disorder, is at risk, or has been diagnosed as having, pre-determined as having, or suspected of having. The terms "protein", "peptide", "polypeptide", and their grammatical equivalents are used interchangeably herein and refer to polymers of amino acid residues linked together by peptide (amide) bonds. These terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Usually, a protein, peptide, or

[0122] polypeptide is at least 3 amino acids in length. A protein, peptide, or polypeptide can refer to an individual protein or a collection of proteins. One or more amino acids in a protein, peptide, or polypeptide can be modified by the addition of chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isoprenyl groups, fatty acid groups, linkers for conjugation, functionalization, or other modifications. A protein, peptide, or

[0123] polypeptide can also be a single molecule or a multimolecular complex. are used herein in the same sense and refer to polymers of amino acid residues linked together by peptide (amide) bonds. These terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Usually, a protein, peptide, or polypeptide is at least 3 amino acids in length. A protein, peptide, or polypeptide can refer to an individual protein or a collection of proteins. One or more amino acids in a protein, peptide, or polypeptide can be modified by the addition of chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isoprenyl groups, fatty acid groups, linkers for conjugation, functionalization, or other modifications. A protein, peptide, or polypeptide can also be a single molecule or a multimolecular complex. polypeptide can be modified by the addition of chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isoprenyl groups, fatty acid groups, linkers for conjugation, functionalization, or other modifications. A protein, peptide, or polypeptide can also be a single molecule or a multimolecular complex. polypeptide can be a single molecule or a multimolecular complex. It can be a protein, peptide, or polypeptide, which can be a natural protein or merely a fragment of a peptide. The protein, peptide, or polypeptide can be natural, recombinant, or synthetic, or any combination thereof. As used herein, the term "fusion protein" refers to a hybrid polypeptide containing protein domains derived from at least two different proteins. One protein can be placed at the amino-terminal (N-terminal) portion or carboxy-terminal (C-terminal) protein of the fusion protein, thus forming an amino-terminal fusion protein or a carboxy-terminal fusion protein, respectively. The protein can include different domains, such as a nucleic acid-binding domain (e.g., the gRNA-binding domain of Cas9 that directs binding of the protein to the target site of the nucleic acid) and a nucleic acid cleavage domain, or the catalytic domain of a nucleic acid editing protein. In some embodiments, the protein can include a proteinaceous portion, such as an amino acid sequence constituting a nucleic acid-binding domain, and an organic compound, such as a compound that can act as a nucleic acid cleaving agent. In some embodiments, the protein forms a complex with or associates with nucleic acids, such as RNA or DNA. Any of the proteins provided herein can be generated by any method known in the art. For example, the proteins provided herein can be generated via expression and purification of recombinant proteins, which is particularly suitable for fusion proteins containing peptide linkers. Methods for nual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012 ) includes the methods described therein, the entire contents of which are incorporated herein by reference.

[0124] The polypeptides and proteins disclosed herein (including their functional parts and functional variants ) can contain synthetic amino acids in place of one or more natural amino acids. Such synthetic amino acids are known in the art and include, for example, aminocyclohexanecarboxylic acid, norleucine, α-amino-n-decanoic acid, homoserine, S-acetylaminomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine, 4- nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine, β-phenylserine β-hydroxyphenylalanine, phenylglycine, α-naphthylalanine, cyclohexylalanine, cyclohexylglycine, indoline-2-carboxylic acid, , 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, aminomalonic acid monoamide, N'-benzyl-N'-methyl-lysine, N',N'-dibenzyl-lysine, 6-hydroxy lysidine, ornithine, α-aminocyclopentanecarboxylic acid, α-aminocyclohexane carboxylic acid, α-aminocycloheptanecarboxylic acid, α-(2-amino-2-norbornane)-carboxylic acid, α,γ-diaminobutyric acid, α,β-diaminopropionic acid, homophenylalanine, and α-tert-butylglycine. Polypeptides and proteins can be associated with post-translational modifications of one or more amino acids of the polypeptide construct. Post-translational modifications of non- By way of limited example, acylation including phosphorylation, acetylation and formylation, glycosylation (including N-linked and O-linked), amidation, hydroxylation, methylation and alkylation including methylation and ethylylation, ubiquitination, addition of pyrrolidone carboxylic acid, formation of disulfide bridges, sulfation, myristoylation, palmitoylation, isoprenylation, farnesylation, geranylation, glypiation, lipoylation and iodination are included.

[0125] As used herein, the term "recombinant" with respect to a protein or nucleic acid refers to a protein or nucleic acid that is a product of human engineering and does not exist in nature. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that contains at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7 mutations compared to any natural sequence.

[0126] "Decrease" means a negative change of at least 10%, 25%, 50%, 75%, or 100%.

[0127] "Reference" means a standard or control condition. In one embodiment, the amount of virus present in cells treated with the base editing system described herein is compared to the level of HBV infection present in untreated control cells, where the control functions as a reference therein. In another embodiment, the sequence of the HBV genome present in cells contacted with the base editing system described herein is compared to the sequence of the HBV genome

[0128] present in untreated control cells. A subset or the entirety of a specified array; for example, a segment of a full-length cDNA or gene sequence, or it may be a complete cDNA or gene sequence. In the case of a polypeptide, the length of the reference polypeptide sequence is generally at least about 16 amino acids, at least about 20 amino acids, at least about 25 amino acids, about 35 amino acids, about 50 amino acids, or about 100 amino acids. In the case of a nucleic acid, the length of the reference nucleic acid sequence is generally at least about 50 nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides, or about 300 nucleotides, or any integer in between or around them. In some embodiments, the reference sequence is the wild-type sequence of the protein of interest. In other embodiments, the reference sequence is a polynucleotide sequence encoding the wild-type protein. tides, or it may be a complete cDNA or gene sequence. In the case of a polypeptide, the length of the reference polypeptide sequence is generally at least about 16 amino acids, at least about 20 amino acids at least about 25 amino acids, about 35 amino acids, about 50 amino acids, or about 100 amino acids. In the case of a nucleic acid, the length of the reference nucleic acid sequence is generally at least about 50 nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides, or about 300 nucleotides, or any integer in between or around them. In some embodiments, the reference sequence is the wild-type sequence of the protein of interest. In other embodiments, the reference sequence is a polynucleotide sequence encoding the wild-type protein. The terms "RNA-programmable nuclease" and "RNA-guided nuclease" are used (e.g., bound or associated) with one or more RNAs (plural) that are not the target of cleavage. In some embodiments, an RNA-programmable nuclease, when in a complex with RNA, may be referred to as a nuclease:RNA complex. Typically, the bound RNA(s) is / are called guide RNA (gRNA). In some embodiments, the RNA-programmable nuclease is a (CR

[0129] ISPR association system) Cas9 endonuclease, e.g., Cas9 (Csn1) from Streptococcus pyogenes (e.g., "Complete genome sequence of an Ml strain of Streptococcus pyo genes (e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes," PLoS Pathogens 8(1): e1002318 (2012)). In some embodiments, the RNA-programmable nuclease is in a complex with RNA and may be referred to as a nuclease:RNA complex. Typically, the bound RNA(s) is / are called guide RNA (gRNA). In some embodiments, the RNA-programmable nuclease is a (CR ISPR association system) Cas9 endonuclease, e.g., Cas9 (Csn1) from Streptococcus pyogenes (e.g., "Complete genome sequence of an Ml strain of Streptococcus pyo genes (e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes," PLoS Pathogens 8(1): e1002318 (2012)). genes (e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes," PLoS Pathogens 8(1): e1002318 (2012)). genes.”Ferretti J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., P rimeaux C,Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H .G., Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B .A., McLaughlin R.E., Proc.Natl.Acad.Sci.U.S.A.98:4658-4663 (2001);”CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma CM., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Voge l J., Charpentier E., Nature 471:602-607 (2011). See also).

[0130] "Specifically binds" means recognizing and binding to the polypeptide and / or nucleic acid molecule of the present invention, but not substantially recognizing and binding to other molecules in the sample, such as nucleic acid molecules, polypeptides, or complexes thereof (e.g., nucleic acid programmable DNA binding domains and guide nucleic acids), compounds, or molecules in a biological sample. The nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding the polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules are 100% identical to the endogenous nucleic acid sequence molecules, polypeptides, or complexes thereof (e.g., nucleic acid programmable DNA binding domains and guide nucleic acids), compounds, or molecules in a biological sample. and guide nucleic acids), compounds, or molecules.

[0131] The nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding the polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules are 100% identical to the endogenous nucleic acid sequence and 100% identical to the endogenous nucleic acid sequence It is not necessary, but generally shows substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence can generally hybridize to at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding the polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but generally show substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence can generally hybridize to at least one strand of a double-stranded nucleic acid molecule. "Hybridizing" means annealing to form a double-stranded molecule between complementary polynucleotide sequences (e.g., the genes described herein), or a portion thereof, under various stringency conditions. (See, e.g., Wahl, G.M. and S.L. Berger (1987) Methods Enzymol. 152: 399; Kimmel, A.R. (1987) Methods Enzymol. 152:507). For example, stringent salt concentrations are usually less than about 750 mM NaCl and 75 mM trisodium citrate, preferably less than about 500 mM NaCl and 50 mM trisodium citrate, more preferably less than about 250 mM NaCl

[0132] and 25 mM trisodium citrate. In the absence of organic solvents, such as formamide, low-stringency hybridization can be obtained, while in the presence of at least about 35 % formamide, more preferably at least about 50% formamide, high-stringency hybridization can be obtained. Stringent temperature conditions usually include at least Temperatures of at least about 30°C, more preferably at least about 37°C, and most preferably at least about 42°C are included. Various additional parameters, such as hybridization time, the concentration of a surfactant, for example, sodium dodecyl sulfate (SDS), and the inclusion or exclusion of carrier DNA, are well known to those skilled in the art. Optionally, by combining these various conditions, various levels of stringency can be achieved. In a preferred embodiment, hybridization is performed at 30°C in 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS. In a more preferred embodiment, hybridization is performed at 37°C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA). In the most preferred embodiment, hybridization is performed at 42°C in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 μg / ml ssDNA. Useful variations of these conditions will be readily apparent to those skilled in the art. For most applications, the stringency of the wash step following hybridization also varies. The conditions for wash stringency can be defined by salt concentration and temperature. As noted above, wash stringency can be increased by lowering the salt concentration or raising the temperature. For example, stringent salt concentrations for the wash step are less than about 30 mM NaCl and 3 mM trisodium citrate, or less than about 15 mM NaCl and 1.5 mM trisodium citrate. Stringent temperature conditions for the wash step typically include at least about 25°C, at least about 42°C, and even at least about 68°C.

[0133] The temperature of ℃ is included. In one embodiment, the washing step is carried out at 25 °C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In another embodiment, the washing step is carried out at 42 °C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the washing step is carried out at 68 °C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Additional variations of these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); A usubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New Y ork, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic c Press, New York), and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York. "Split" means being divided into two or more fragments.

[0134] "Split Cas9 protein" or "split Cas9" refers to the Cas9 protein provided as an N-terminal fragment and a C-terminal fragment encoded by two separate nucleotide

[0135] sequences. sequences. Refers. Polypeptides corresponding to the N-terminal and C-terminal portions of the Cas9 protein can be spliced to form a "reconstituted" Cas9 protein. In certain embodiments, the Cas9 protein is split into two fragments within the denaturation region of the protein, as described, for example, in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014, or as described in Jiang et al. (2016) Science 351:867-871. The PDB file: 5F9R, each of which is incorporated herein by reference. In some embodiments, the protein is split into two fragments at any C, T, A, or S within the region of SpCas9 between amino acids A292 - G364, F445 - K483, or E565 - T637, or at the corresponding position of any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or any other napDNAbp. In some embodiments, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of splitting the protein into two fragments is referred to as "fragmentation" of the protein. and can form a "reconstituted" Cas9 protein. In certain embodiments, the Cas9 protein is split into two fragments within the denaturation region of the protein, as described, for example, in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014, or as described in Jiang et al. (2016) Science 351:867-871. The PDB file: 5F9R, each of which is incorporated herein by reference. In some embodiments, the protein is split into two fragments at any C, T, A, or S within the region of SpCas9 between amino acids A292 - G364, F445 - K483, or E565 - T637, or at the corresponding position of any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or any other napDNAbp. In some embodiments, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of splitting the protein into two fragments is referred to as "fragmentation" of the protein. protein is split into two fragments within the denaturation region of the protein, as described, for example, in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014, or as described in Jiang et al. (2016) Science 351:867-871. The PDB file: 5F9R, each of which is incorporated herein by reference. In some embodiments, the protein is split into two fragments at any C, T, A, or S within the region of SpCas9 between amino acids A292 - G364, F445 - K483, or E565 - T637, or at the corresponding position of any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or any other napDNAbp. In some embodiments, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of splitting the protein into two fragments is referred to as "fragmentation" of the protein. In other embodiments, the N-terminal portion of the Cas9 protein includes amino acids 1 - 573 or 1 - 637 of S. pyogenes Cas9 wild-type (SpCas9) (NCBI reference sequence: NC_002737.2, Uniprot reference sequence: Q99ZW2), and the C-terminal portion of the Cas9 protein includes a portion of SpCas9 wild-type amino acids 574 - 1368 or 638 - 1368. as described, for example, in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014, or as described in Jiang et al. (2016) Science 351:867-871. The PDB file: 5F9R, each of which is incorporated herein by reference. In some embodiments, the protein is split into two fragments at any C, T, A, or S within the region of SpCas9 between amino acids A292 - G364, F445 - K483, or E565 - T637, or at the corresponding position of any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or any other napDNAbp. In some embodiments, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of splitting the protein into two fragments is referred to as "fragmentation" of the protein. In some embodiments, the protein is split into two fragments at any C, T, A, or S within the region of SpCas9 between amino acids A292 - G364, F445 - K483, or E565 - T637, or at the corresponding position of any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or any other napDNAbp. In some embodiments, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of splitting the protein into two fragments is referred to as "fragmentation" of the protein. In some embodiments, the protein is split into two fragments at any C, T, A, or S within the region of SpCas9 between amino acids A292 - G364, F445 - K483, or E565 - T637, or at the corresponding position of any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or any other napDNAbp. In some embodiments, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of splitting the protein into two fragments is referred to as "fragmentation" of the protein. In some embodiments, the protein is split into two fragments at any C, T, A, or S within the region of SpCas9 between amino acids A292 - G364, F445 - K483, or E565 - T637, or at the corresponding position of any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or any other napDNAbp. In some embodiments, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of splitting the protein into two fragments is referred to as "fragmentation" of the protein. In some embodiments, the protein is split into two fragments at any C, T, A, or S within the region of SpCas9 between amino acids A292 - G364, F445 - K483, or E565 - T637, or at the corresponding position of any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or any other napDNAbp. In some embodiments, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of splitting the protein into two fragments is referred to as "fragmentation" of the protein. In some embodiments, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of splitting the protein into two fragments is referred to as "fragmentation" of the protein. In some embodiments, the process of splitting the protein into two fragments is referred to as "fragmentation" of the protein. In some embodiments, the process of splitting the protein into two fragments is referred to as "fragmentation" of the protein.

[0136] In other embodiments, the N-terminal portion of the Cas9 protein includes amino acids 1 - 573 or 1 - 637 of S. pyogenes Cas9 wild-type (SpCas9) (NCBI reference sequence: NC_002737.2, Uniprot reference sequence: Q99ZW2), and the C-terminal portion of the Cas9 protein includes a portion of SpCas9 wild-type amino acids 574 - 1368 or 638 - 1368. In other embodiments, the N-terminal portion of the Cas9 protein includes amino acids 1 - 573 or 1 - 637 of S. pyogenes Cas9 wild-type (SpCas9) (NCBI reference sequence: NC_002737.2, Uniprot reference sequence: Q99ZW2), and the C-terminal portion of the Cas9 protein includes a portion of SpCas9 wild-type amino acids 574 - 1368 or 638 - 1368. In other embodiments, the N-terminal portion of the Cas9 protein includes amino acids 1 - 573 or 1 - 637 of S. pyogenes Cas9 wild-type (SpCas9) (NCBI reference sequence: NC_002737.2, Uniprot reference sequence: Q99ZW2), and the C-terminal portion of the Cas9 protein includes a portion of SpCas9 wild-type amino acids 574 - 1368 or 638 - 1368. In other embodiments, the N-terminal portion of the Cas9 protein includes amino acids 1 - 573 or 1 - 637 of S. pyogenes Cas9 wild-type (SpCas9) (NCBI reference sequence: NC_002737.2, Uniprot reference sequence: Q99ZW2), and the C-terminal portion of the Cas9 protein includes a portion of SpCas9 wild-type amino acids 574 - 1368 or 638 - 1368.

[0137] The C-terminal portion of split Cas9 can be combined with the N-terminal portion of split Cas9 to form a complete Cas9 protein. In some embodiments, the C-terminal portion of the Cas9 protein starts at the position where the N-terminal portion of the Cas9 protein ends. Thus, in some embodiments, the C-terminal portion of split Cas9 includes a part of amino acids (551-651)-1368 of spCas9. By "(551-651)-1368", it means starting with amino acids between 551 and 651 (including both ends) and ending with amino acid 1368. For example, the C-terminal portion of split Cas9 is amino acids 551-1368, 552-1368, 553-1368, 554-1368, 555-1368, 556 -1368, 557-1368, 558-1368, 559-1368, 560-1368, 561-1368, 562-1368, 563-1 368, 564-1368, 565-1368, 566-1368, 567-1368, 568-1368, 569-1368, 570-1368 , 571-1368, 572-1368, 573-1368, 574-1368, 575-1368, 576-1368, 577-1368, 5 78-1368, 579-1368, 580-1368, 581-1368, 582-1368, 583-1368, 584-1368, 585 -1368, 586-1368, 587-1368, 588-1368, 589-1368, 590-1368, 591-1368, 592-1 368, 593-1368, 594-1368, 595-1368, 596-1368, 597-1368, 598-1368, 599-1368 , 600-1368, 601-1368, 602-1368, 603-1368, 604-1368, 605-1368, 606-1368, 6 07-1368, 608-1368, 609-1368, 610-1368, 611-1368, 612-1368, 613-1368, 614-1368, 6 07 - 1368, 608 - 1368, 609 - 1368, 610 - 1368, 611 - 1368, 612 - 1368, 613 - 1368, 614 - 1368, 615 - 1368, 616 - 1368, 617 - 1368, 618 - 1368, 619 - 1368, 620 - 1368, 621 - 1 368, 622 - 1368, 623 - 1368, 624 - 1368, 625 - 1368, 626 - 1368, 627 - 1368, 628 - 1368 , 629 - 1368, 630 - 1368, 631 - 1368, 632 - 1368, 633 - 1368, 634 - 1368, 635 - 1368, 6 36 - 1368, 637 - 1368, 638 - 1368, 639 - 1368, 640 - 1368, 641 - 1368, 642 - 1368, 643 - 1368, 644 - 1368, 645 - 1368, 646 - 1368, 647 - 1368, 648 - 1368, 649 - 1368, 650 - 1 368, or may include a part of any one of 651 - 1368. In some embodiments , the C - terminal portion of the split Cas9 protein includes amino acids 574 - 1368 of SpCas9 or a part of 638 - 136 8.

[0138] "Subject" means a human or non - human mammal, such as a non - human primate (monkey), cow, horse, dog, sheep, or cat, and includes, but is not limited to, mammals. In some embodiments, the subject described herein is infected with HBV or has a tendency to develop HBV .

[0139] "Substantially identical" means that a polypeptide or nucleic acid molecule exhibits at least 50% identity to a reference amino acid sequence (e.g., any one of the amino acid sequences described herein) or a nucleic acid sequence (e.g., any one of the nucleic acid sequences described herein ). In one embodiment In an embodiment, such an array is at least 60%, 80%, 85%, 90%, 95%, or even 99% identical at the amino acid level or at the nucleic acid level to the array used for comparison.

[0140] Sequence identity is typically measured using sequence analysis software (e.g., Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705 sequence analysis software package, BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX program). Such software aligns identical or similar sequences by assigning a degree of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine, valine, isoleucine, leucine, aspartic acid, glutamic acid, asparagine, glutamine, serine, threonine, lysine, arginine, and phenylalanine, tyrosine. ersity of Wisconsin Biotechnology Center,1710 University Avenue,Madison,Wis.5370 5 sequence analysis software package, BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX program). Such software aligns identical or similar sequences by assigning a degree of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine, valine, isoleucine, leucine, aspartic acid, glutamic acid, asparagine, glutamine, serine, threonine, lysine, arginine, and phenylalanine, tyrosine. In an exemplary method of measuring the degree of identity, the BLAST program may be used with an e~e probability score for closely related sequences. COBALT is used, for example, with the following parameters: a) Alignment parameters: Gap penalties -11, -1 and End-Gap penalties -5, -1, b) CDD parameters: Use RPS BLAST on; Blast E-value 0.003; Find Conserved columns and Recompute on, and In an exemplary method of measuring the degree of identity, the BLAST program may be used with an e~e probability score for closely related sequences. COBALT is used, for example, with the following parameters: -3 ~e -100 probability score. COBALT is used, for example, with the following parameters: a) Alignment parameters: Gap penalties -11, -1 and End-Gap penalties -5, -1, b) CDD parameters: Use RPS BLAST on; Blast E-value 0.003; Find Conserved columns and Recompute on, and c) Query clustering parameters: Use query clusters on; Word Size 4; Max cluster distance 0.8;Alphabet Regular. EMBOSS Needle can be used, for example, with the following parameters: a) Matrix: BLOSUM62; b) GAP OPEN: 10; c) GAP EXTEND: 0.5; d)OUTPUT FORMAT:pair; e)END GAP PENALTY:false; f) END GAP OPEN: 10, and g) END GAP EXTEND: 0.5.

[0141] The term "target site" refers to a site that is a site for a deaminase or a deaminase (e.g., a cytidine deaminase Adenine deaminase or adenine deaminase fusion protein or a base-editing The term "deaminated" refers to a sequence within a nucleic acid molecule that is deaminated by a fusion protein comprising the nucleotide sequence of the nucleic acid molecule.

[0142] RNA-programmable nucleases (e.g., Cas9) are capable of catalyzing RNA:DNA hybridization These proteins, in principle, can be used to target DNA cleavage sites using guide R Any sequence specified by the NA can be targeted. Site-specific cleavage (e.g., genomic To modify the genome, we use RNA-programmable nucleases such as Cas9. Methods are known in the art (e.g., Cong, L. et al., Multiplex genome engineering using CRISPR / Cas systems.Science 339,819-823 (2013);Mali,P.et ah,RNA-guided huma n genome engineering via Cas9. Science 339,823-826 (2013);Hwang,W.Y.et ah,Effici ent genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31,227-229 (2013);Jinek,M.et ah,RNA-programmed genome editing in human cells. eL ife 2,e00471 (2013);Dicarlo,J.E.et ah,Genome engineering in Saccharomyces cerevi siae using CRISPR-Cas systems. Nucleic acids research (2013);Jiang,W.et ah RNA-g uided editing of bacterial genomes using CRISPR-Cas systems. Nature biotechnolog y 31,233-239 (2013); see also, each of which is hereby incorporated by reference in its entirety.

[0143] As used herein, the terms "treating," "treatment," and "treat" and the like refer to reducing and / or alleviating a disorder and / or associated symptoms or obtaining a desired pharmacological and / or physiological effect. Treating a disorder or condition is understood not necessarily to mean that the disorder, condition, or associated symptoms are completely eliminated, although this is not excluded. In some embodiments, the effect is therapeutic, i.e., the effect is, without limitation, Reduce, decrease, inhibit, alleviate, mitigate, or lower the intensity of, or cure, in part or in whole, the symptoms. In some embodiments, the effect is preventive, i.e., the effect defends against or prevents the onset or recurrence of a disease or medical condition. For this purpose, the methods disclosed herein include administering a therapeutically effective amount of

[0144] the composition as described herein. In one embodiment, the invention provides a treatment for HBV infection. "Uracil glycosylase inhibitor" or "UGI" means an agent that inhibits the uracil excision repair system. In one embodiment, the agent is a protein or a fragment thereof that binds to the host uracil DNA glycosylase and prevents the removal of uracil residues from DNA. In one embodiment, UGI is a protein, a fragment thereof, or a domain that can inhibit the base excision repair enzyme of uracil DNA glycosylase. In some embodiments, the UGI domain includes the wild-type UGI or a modified version thereof. In some embodiments, the UGI domain includes a fragment of the exemplary amino acid sequences described below. In some embodiments, the UGI fragment includes an amino acid sequence that comprises at least 60%, at least 65%, at least 70%, at least 75%, at least or a part thereof, and having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, or 100% identity. Exemplary U GIs include the following amino acid sequences: >splP14739IUNGI_BPPB2 uracil DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLT S D APE YKPW ALVIQDS NGENKIKML.

[0145] The ranges provided herein are to be understood as shorthand for all values within the range. For example, a range of 1-50 is understood to include any number, combination of numbers, or sub-range from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50. The recitation of a list of chemical groups in any definition of a variable herein includes the definition of that variable with any single group or combination of the recited groups. The recitation of embodiments for a variable or aspect herein includes that embodiment as any single embodiment, or in combination with any other embodiment or

[0146] a part thereof.

[0147] ​​​​Any composition or method provided herein can be combined with one or more of any other compositions and methods provided herein.

[0148] The description and examples herein illustrate embodiments of the disclosure in detail. The disclosure is not limited to the specific embodiments described herein and is understood to be capable of various differences yes. Those skilled in the art will recognize that many variations and modifications of the disclosure exist and that they are within the scope of the disclosure.

[0149] All terms are intended to be understood as would be understood by those skilled in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0150] The practice of some embodiments disclosed herein, unless otherwise indicated, is within the scope of the art of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics ics and recombinant DNA using conventional techniques. For example, Sambrook and Green, Molecular Clo ning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Mol ecular Biology (F.M. Ausubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (M.J. MacPherson, B.D. Hames and G. R. Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Man ual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Ap plications, 6th Edition (R.I. Freshney, ed. (2010)). Please refer to it.

[0151] Although various features of the present disclosure can be described in the context of a single embodiment, the features can also be provided separately or in any suitable combination. Conversely, although the present disclosure can be described herein in the context of separate embodiments for clarity, the present disclosure can also be implemented in a single embodiment. The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described. The features of the present disclosure are particularly described in the appended claims. A better understanding of the features and advantages of this specification can be obtained by referring to the following detailed description of exemplary embodiments that utilize the principles of the present disclosure and considering the appended drawings described below.

[0152] The features of the present disclosure are particularly recited in the appended claims. A better understanding of the features and advantages of this specification can be obtained by referring to the following detailed description of exemplary embodiments that utilize the principles of the present disclosure and considering the appended drawings described below.

Brief Description of the Drawings

[0153]

Figure 1

Figure 2

Figure 3A

Figure 3B

Figure 3C

Figure 3D

Figure 3E

Figure 3F

Figure 3G

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

[0154]

Figure 11

Figure 12A

Figure 12B

Figure 12C

Figure 13A

Figure 13B

Figure 13C

Figure 13D

Figure 13E

Figure 13F

Figure 14

Figure 15

Figure 16A

Figure 16B

Figure 17A

Figure 17B

Figure 18

Figure 19

Figure 20

[0155]

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

[0156]

Figure 26A

Figure 26B

Figure 26C

Figure 27A

Figure 27B

[0157] The present invention features compositions and methods for editing the HBV genome. For example, as described herein The compositions contemplated in the written description can, in some embodiments, include a base editor and a guide nucleic acid that targets specific nucleotides in the HBV gene. In some embodiments, editing introduces a premature stop codon into the coding sequence of one of the viral proteins. In another embodiment, editing introduces one or more functional substitutions into the coding sequences of one or more HBV proteins.

[0158] The HBV genome consists of a 3.2 kb partially double-stranded DNA and an open reading frame (ORF) that encodes seven proteins. As shown in Figure 1, the open reading frame (O RF) P encodes the viral polymerase. ORF C / PreC encodes the capsid protein. ORF preS1, ORF preS2, and ORF S encode the large (L), medium (M), and small (S) surface proteins, respectively. ORF X encodes the secreted X protein.

[0159] The partially double-stranded HBV genome is converted by host factors into covalently closed circular DNA (cccDNA). The cccDNA is transcribed by host RNA polymerase to generate viral mRNAs, including the pregenomic RNA (pgRN A). The pgRNA is reverse transcribed by the HBV polymerase into genomic HBV DNA, converted into cccDNA, packaged into virions, or integrated into the host cell genome (Figure 2). CccDNA, an important component of the HBV life cycle, is a stable molecule that causes chronic HBV infection. Editing of the HBV genome can interfere with the formation of cccDNA and thereby reduce the pathogenicity of the virus.

[0160] ​​​​​ There are 10 different HBV genotypes (A - J) (Figure 3A). A "genotype" is characterized by having less than 92% sequence identity with any other genome, and a sub - genotype is characterized by having a sequence identity of 96 - 92% less. HBV of genotype D is the most common in the United States (Figure 3A). Research models of HBV genotype D can be utilized, including virus stocks (e.g., genotype D, sub - genotype ayw (Imquest)) and mouse models (e.g., humanized mouse model (Phoenixbio)). Therefore, in some embodiments, methods and compositions for targeting and editing the HBV ORF are provided. These compositions can include Cas9 or other nucleic acid programmable DNA - binding protein domains and nucleic acid base editors having an adenosine deaminase domain or a cytosine deaminase domain. In some embodiments, the base editor introduces one or more modifications into the HBV ORF. In some embodiments, the modification results in a mutation in a conserved portion of the HBV protein. In certain embodiments, the modification introduces one or more stop codons. Throughout this specification, the introduction of a stop codon that results in premature termination of a protein is represented by amino acid symbols, amino acid positions, and the term STOP (e.g., R87STOP indicates that the codon encoding arginine at amino acid position 87 is replaced with a stop codon). Advantageously, the methods of the present invention do not introduce double - strand breaks into the HBV genome. The present invention provides a strategy for treating chronic HBV using base editing (Figure 3B). In this specification, guide RNAs that introduce stop codons or functional mutations into the HBV gene are identified.

[0161] ​​​​​​​​​​​​ for, or to identify gRNAs that generate abasic sites in the hyperconserved regions of the HBV genome cleaning is described (Figure 3C). Using the methods and compositions described herein to introduce a stop codon into a viral gene can be achieved without generating double-strand breaks, thereby eliminating or reducing the risk of cleaving host genetic material after HBV is integrated into the host genome. Furthermore, the compositions use deaminases, which are natural HBV antiviral restriction factors. For example, induction of APOBEC cytidine deaminase with interferon α or lymphotoxin β receptor (LTBR) promotes the formation of abasic sites and cccDNA degradation (Figure 3B). Furthermore, using base editors in the absence of the uracil glycosylase inhibitor domain can target cellular uracil glycosylase to cccDNA and promote its degradation (Figure 3B). Furthermore, using base editors in the absence of the uracil glycosylase inhibitor domain can target cellular uracil glycosylase to cccDNA and promote its degradation (Figure 3B). Furthermore, using base editors in the absence of the uracil glycosylase inhibitor domain can target cellular uracil glycosylase to cccDNA and promote its degradation (Figure 3B). Furthermore, using base editors in the absence of the uracil glycosylase inhibitor domain can target cellular uracil glycosylase to cccDNA and promote its degradation (Figure 3B). Furthermore, using base editors in the absence of the uracil glycosylase inhibitor domain can target cellular uracil glycosylase to cccDNA and promote its degradation (Figure 3B). Furthermore, using base editors in the absence of the uracil glycosylase inhibitor domain can target cellular uracil glycosylase to cccDNA and promote its degradation (Figure 3B).

[0162] Another screening provided identifies conserved gRNAs that can be used to generate abasic sites in cccDNA. As shown in Figure 3D, when lentivirus was used to introduce a base editor and gRNA (Lenti-HBV), seven guide RNAs with editing efficiencies above 20% were identified (Figure 3D). Guide RNAs targeting the conserved regions are shown in Figure 3E. Some of the gRNAs had editing efficiencies of at least 45% (Figure 3D). Guide RNAs targeting the conserved regions are shown in Figure 3E. Some of the gRNAs had editing efficiencies of at least 45% (Figure 3D). Guide RNAs targeting the conserved regions are shown in Figure 3E. Some of the gRNAs had editing efficiencies of at least 45% (Figure 3F and 3G).

[0163] In some embodiments, provided are methods and compositions for editing HBV cccDNA using base editors comprising a cytidine deaminase domain or an adenosine deaminase domain (Figure 3D). Guide RNAs targeting the conserved regions are shown in Figure 3E. Some of the gRNAs had editing efficiencies of at least 45% . In one embodiment, the base editor comprises an APOBEC cytidine deaminase domain, a Cas9 domain , and optionally, one or more uracil glycosylase inhibitor (UGI) domains (Figure 4 A, 4B).

[0164] Nucleic acid base editor Disclosed herein are base editors or nucleic acid base editors for editing, modifying, or altering the target nucleotide sequence of a polynucleotide. As used herein, a nucleic acid base editor or base editor comprising a polynucleotide-programmable nucleotide binding domain and a nucleic acid base editing domain (e.g., , adenosine deaminase, cytidine deaminase). The polynucleotide-programmable nucleotide binding domain specifically binds to a target polynucleotide sequence (i.e., via complementary base pairing between the bases of the bound guide polynucleotide (e.g., gRNA) and the bases of the target polynucleotide sequence) when combined with the bound guide polynucleotide (e.g., gRNA), thereby localizing the base editor to the target nucleic acid sequence to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid. , adenosine deaminase, cytidine deaminase). The polynucleotide-programmable nucleotide binding domain specifically binds to a target polynucleotide sequence (i.e., via complementary base pairing between the bases of the bound guide polynucleotide (e.g., gRNA) and the bases of the target polynucleotide sequence) when combined with the bound guide polynucleotide (e.g., gRNA), thereby localizing the base editor to the target nucleic acid sequence to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid. , adenosine deaminase, cytidine deaminase). The polynucleotide-programmable nucleotide binding domain specifically binds to a target polynucleotide sequence (i.e., via complementary base pairing between the bases of the bound guide polynucleotide (e.g., gRNA) and the bases of the target polynucleotide sequence) when combined with the bound guide polynucleotide (e.g., gRNA), thereby localizing the base editor to the target nucleic acid sequence to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid. collective is described. The polynucleotide-programmable nucleotide binding domain binds to a target polynucleotide sequence (i.e., via complementary base pairing between the bases of the bound guide polynucleotide (e.g., gRNA) and the bases of the target polynucleotide sequence) when combined with the bound guide polynucleotide (e.g., gRNA), thereby localizing the base editor to the target nucleic acid sequence to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid. to a target polynucleotide sequence (i.e., via complementary base pairing between the bases of the bound guide polynucleotide (e.g., gRNA) and the bases of the target polynucleotide sequence) when combined with the bound guide polynucleotide (e.g., gRNA), thereby localizing the base editor to the target nucleic acid sequence to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid. to a target polynucleotide sequence (i.e., via complementary base pairing between the bases of the bound guide polynucleotide (e.g., gRNA) and the bases of the target polynucleotide sequence) when combined with the bound guide polynucleotide (e.g., gRNA), thereby localizing the base editor to the target nucleic acid sequence to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid. to a target polynucleotide sequence (i.e., via complementary base pairing between the bases of the bound guide polynucleotide (e.g., gRNA) and the bases of the target polynucleotide sequence) when combined with the bound guide polynucleotide (e.g., gRNA), thereby localizing the base editor to the target nucleic acid sequence to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid. to a target polynucleotide sequence (i.e., via complementary base pairing between the bases of the bound guide polynucleotide (e.g., gRNA) and the bases of the target polynucleotide sequence) when combined with the bound guide polynucleotide (e.g., gRNA), thereby localizing the base editor to the target nucleic acid sequence to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid. to a target polynucleotide sequence (i.e., via complementary base pairing between the bases of the bound guide polynucleotide (e.g., gRNA) and the bases of the target polynucleotide sequence) when combined with the bound guide polynucleotide (e.g., gRNA), thereby localizing the base editor to the target nucleic acid sequence to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid. to a target polynucleotide sequence (i.e., via complementary base pairing between the bases of the bound guide polynucleotide (e.g., gRNA) and the bases of the target polynucleotide sequence) when combined with the bound guide polynucleotide (e.g., gRNA), thereby localizing the base editor to the target nucleic acid sequence to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid.

[0165] Polynucleotide-programmable nucleotide binding domain It should be understood that the polynucleotide-programmable nucleotide binding domain may also include a nucleic acid-programmable protein that binds to RNA. For example, polynucle otide-programmable nucleotide binding domain may also include a nucleic acid-programmable protein that binds to RNA. For example, polynucle A nucleotide-programmable nucleotide-binding domain can be associated with a nucleic acid that guides the nucleotide-binding domain programmable by the polynucleotide. Other nucleic acid-programmable DNA-binding proteins are also within the scope of the present disclosure, but they are not specifically described in the present disclosure.

[0166] The polynucleotide-programmable nucleotide-binding domain of the base editor may itself contain one or more domains. For example, the polynucleotide-programmable nucleo tide-binding domain may contain one or more nuclease domains. In some embodiments the nuclease domain of the polynucleotide-programmable nucleotide-binding domain may contain an endonuclease or an exonuclease. As used herein, the term "exonuclease" refers to a protein or polypeptide capable of digesting a nucleic acid (e.g., RNA or DNA) from a free end, and the term "endonuclease" refers to a protein or polypeptide capable of catalyzing (e.g., cleaving) an internal region of a nucleic acid (e.g., DNA or RNA). In some embodiments, the endonuclease can cleave one strand of a double-stranded nucleic acid. In some embodiments, the endonuclease can cleave both strands of a double-stranded nucleic acid molecule. In some embodiments, the polynucleotide-programmable nucleotide-binding domain can be a deoxyribonuclease. In some embodiments, the polynucleotide-programmable nucleotide-binding domain can be a ribonuclease.

[0167] In some embodiments, the nuclease domain of the polynucleotide-programmable nucleotide binding domain can cleave 0, 1, or 2 strands of the target polynucleotide. In some cases, the polynucleotide-programmable nucleotide binding domain can include a nickase domain. As used herein, the term "nickase" refers to a polynucleotide-programmable nucleotide binding domain that includes a nuclease domain capable of cleaving only one of the two strands in a double-stranded nucleic acid molecule (e.g., DNA). In some embodiments, the nickase can be derived from a fully catalytically active (e.g., native) form of the polynucleotide-programmable nucleotide binding domain by introducing one or more mutations into the active polynucleotide-programmable nucleotide binding domain. For example, if the polynucleotide-programmable nucleotide binding domain includes a nickase domain derived from Cas9, the Cas9-derived nickase domain can include the D10A mutation and histidine at position 840. In such cases, residue H840 retains catalytic activity and can thereby cleave one strand of the nucleic acid duplex. In another example, the Cas9-derived nickase domain can include the H840A mutation, while the amino acid residue at position 10 remains D. In some embodiments, the nickase can be derived from a fully catalytically active (e.g., native) form of the polynucleotide-programmable nucleotide binding domain by removing all or part of the nuclease domain that does not require nickase activity. For example, the polynucleotide-programmable nucleotide ... ... ... ... ... ... ... ... ... When the binding domain includes a nickase domain derived from Cas9, the nickase domain derived from Cas9 may include a deletion of all or part of the RuvC domain or the HNH domain.

[0168] The amino acid sequence of an exemplary catalytically active Cas9 is as follows: MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA ​HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD。

[0169] Therefore, a polynucleotide programmable nucleotide containing a nickase domain base editor containing a binding domain binds to a specific polynucleotide target sequence (e.g., generating a single-stranded DNA break (nick) (determined by the complementary sequence of the guide nucleic acid) is possible. In some embodiments, the nucleic acid double-stranded target polynucleotide sequence cleaved by a base editor comprising a nickase domain (e.g., a nickase domain derived from Cas9) strand is the strand that is not edited by the base editor (i.e., the strand cleaved by the base editor is on the opposite side of the strand containing the base to be edited). In other embodiments, a base editor comprising a nickase domain (e.g., a nickase domain derived from Cas9) can cleave the strand of the DNA molecule that is the target of editing. In such cases, the non-target strand is not cleaved.

[0170] Base editors comprising a catalytically inactive (i.e., unable to cleave the target polynucleotide sequence) polynucleotide programmable nucleotide binding domain are also provided herein. As used herein, the terms "catalytically inactive" and "nuclease inactive" are used interchangeably to refer to a polynucleotide programmable nucleotide binding domain having one or more mutations and / or deletions that result in the inability to cleave a strand of nucleic acid. In some embodiments, a catalytically inactive polynucleotide programmable nucleotide binding domain base editor may lack nuclease activity as a result of specific point mutations in one or more nuclease domains. For example, in the case of a base editor comprising a Cas9 domain, Cas9 may contain both a D10A mutation and an H840A mutation. Such mutations inactivate both nuclease domains, thereby resulting in a loss of nuclease activity. In other embodiments, a catalytically inactive polynucleotide otide programmable nucleotide binding domain base editor may lack nuclease activity as a result of specific point mutations in one or more nuclease domains. For example, in the case of a base editor comprising a Cas9 domain, Cas9 may contain both a D10A mutation and an H840A mutation. Such mutations inactivate both nuclease domains, thereby resulting in a loss of nuclease activity. In other embodiments, a catalytically inactive polynucleotide ​​​​​​​​​​​​​The programmable nucleotide binding domain can include one or more deletions of all or part of the catalytic domain (e.g., the RuvC1 and / or the HNH domain). In further embodiments the catalytically inactive polynucleotide programmable nucleotide binding domain is point mutations (e.g., D10A or H840A), and deletions of all or part of the nuclease domain .

[0171] Also, contemplated herein are mutations that can generate a catalytically inactive polynucleotide programmable nucleotide binding domain from a previous functional version of the polynucleotide programmable nucleotide binding domain. For example, in the case of catalytically inactive Cas9 (“dCas9”), mutants having mutations other than D10A and H840A are provided, resulting in nuclease-inactivated Cas9. Such mutations include, by way of example, other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). Additional suitable nuclease-inactive dCas9 domains will be apparent to those skilled in the art based on the present disclosure and the knowledge in the art, and are within the scope of the present disclosure. Such additional exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H84 0A / N863A mutant domains (e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nick , which are incorporated herein by reference). Included are the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H84 0A / N863A mutant domains, but are not limited thereto (e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nick Cases for cooperative genome engineering. Nature Biotechnology.2013;31 (9):833-83 See 8; the entire contents of each are hereby incorporated by reference.

[0172] Polynucleotide-programmable nucleotide-binding domains that can be incorporated into base editors Non-limiting examples of polynucleotide-programmable nucleotide-binding domains include domains derived from CRISPR proteins, restriction nucleases, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases (ZFNs). In some cases, the base editor includes a polynucleotide-programmable nucleotide-binding domain that includes a native or modified protein or a portion thereof, and can bind to a nucleic acid sequence during modification of the nucleic acid via a combined guide nucleic acid, via CRISPR (i.e., Clustered Regularly Interspaced Short Palindromic Repeats). Such proteins are referred to herein as "CRISPR proteins." Thus, in the present specification, a base editor that includes a polynucleotide-programmable nucleotide-binding domain that includes all or a portion of a CRISPR protein (i.e., a base editor that is also referred to as the "CRISPR protein-derived domain" of the base editor, and includes all or a portion of a CRISPR protein as a domain) is disclosed. The domain derived from the CRISPR protein incorporated into the base editor can be modified compared to the wild-type or native version of the CRISPR protein. For example, as described below, the CRISPR protein-derived domain can be a CRISPR protein derived domain that is modified compared to the wild-type or native version of the CRISPR protein. For example, as described below, the CRISPR protein-derived domain can be a CRISPR protein derived domain that is modified compared to the wild-type or native version of the CRISPR protein. For example, as described below, the CRISPR protein-derived domain can be a CRISPR protein derived domain that is modified compared to the wild-type or native version of the CRISPR protein. For example, as described below, the CRISPR protein-derived domain can be a CRISPR protein derived domain that is modified compared to the wild-type or native version of the CRISPR protein. For example, as described below, the CRISPR protein-derived domain can be a CRISPR protein derived domain that is modified compared to the wild-type or native version of the CRISPR protein. For example, as described below, the CRISPR protein-derived domain can be a CRISPR protein It may contain one or more mutations, insertions, deletions, rearrangements and / or recombinations compared to the wild-type or native version of the protein. and / or recombination.

[0173] CRISPR is an adaptive immune system that provides defense against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). The CRISPR cluster contains spacers, sequences complementary to the aforementioned mobile elements, and target invading nucleic acids. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, to correctly process pre-crRNA, trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein are required. TracrRNA functions as a guide for ribonuclease 3-assisted processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to crRNA is first endonucleolytically cleaved and then trimmed exonucleolytically in the 3'-5' direction. In nature, DNA binding and cleavage usually require both a protein and two RNAs. However, a single guide RNA (abbreviated as "sgRNA" or simply "gRNA") can be engineered to incorporate both aspects of crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012) (the entire contents of which are hereby incorporated by reference). are included. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, to correctly process pre-crRNA, trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein are required. TracrRNA functions as a guide for ribonuclease 3-assisted processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to crRNA is first endonucleolytically cleaved and then trimmed exonucleolytically in the 3'-5' direction. In nature, DNA binding and cleavage usually require both a protein and two RNAs. However, a single guide RNA (abbreviated as "sgRNA" or simply "gRNA") can be engineered to incorporate both aspects of crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012) (the entire contents of which are hereby incorporated by reference). are included. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, to correctly process pre-crRNA, trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein are required. TracrRNA functions as a guide for ribonuclease 3-assisted processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to crRNA is first endonucleolytically cleaved and then trimmed exonucleolytically in the 3'-5' direction. In nature, DNA binding and cleavage usually require both a protein and two RNAs. However, a single guide RNA (abbreviated as "sgRNA" or simply "gRNA") can be engineered to incorporate both aspects of crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012) (the entire contents of which are hereby incorporated by reference). are included. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, to correctly process pre-crRNA, trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein are required. TracrRNA functions as a guide for ribonuclease 3-assisted processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to crRNA is first endonucleolytically cleaved and then trimmed exonucleolytically in the 3'-5' direction. In nature, DNA binding and cleavage usually require both a protein and two RNAs. However, a single guide RNA (abbreviated as "sgRNA" or simply "gRNA") can be engineered to incorporate both aspects of crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012) (the entire contents of which are hereby incorporated by reference). are included. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, to correctly process pre-crRNA, trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein are required. TracrRNA functions as a guide for ribonuclease 3-assisted processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to crRNA is first endonucleolytically cleaved and then trimmed exonucleolytically in the 3'-5' direction. In nature, DNA binding and cleavage usually require both a protein and two RNAs. However, a single guide RNA (abbreviated as "sgRNA" or simply "gRNA") can be engineered to incorporate both aspects of crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012) (the entire contents of which are hereby incorporated by reference). are included. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, to correctly process pre-crRNA, trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein are required. TracrRNA functions as a guide for ribonuclease 3-assisted processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to crRNA is first endonucleolytically cleaved and then trimmed exonucleolytically in the 3'-5' direction. In nature, DNA binding and cleavage usually require both a protein and two RNAs. However, a single guide RNA (abbreviated as "sgRNA" or simply "gRNA") can be engineered to incorporate both aspects of crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012) (the entire contents of which are hereby incorporated by reference). are included. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, to correctly process pre-crRNA, trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein are required. TracrRNA functions as a guide for ribonuclease 3-assisted processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to crRNA is first endonucleolytically cleaved and then trimmed exonucleolytically in the 3'-5' direction. In nature, DNA binding and cleavage usually require both a protein and two RNAs. However, a single guide RNA (abbreviated as "sgRNA" or simply "gRNA") can be engineered to incorporate both aspects of crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012) (the entire contents of which are hereby incorporated by reference). are included. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, to correctly process pre-crRNA, trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein are required. TracrRNA functions as a guide for ribonuclease 3-assisted processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to crRNA is first endonucleolytically cleaved and then trimmed exonucleolytically in the 3'-5' direction. In nature, DNA binding and cleavage usually require both a protein and two RNAs. However, a single guide RNA (abbreviated as "sgRNA" or simply "gRNA") can be engineered to incorporate both aspects of crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012) (the entire contents of which are hereby incorporated by reference). are included. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, to correctly process pre-crRNA, trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein are required. TracrRNA functions as a guide for ribonuclease 3-assisted processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to crRNA is first endonucleolytically cleaved and then trimmed exonucleolytically in the 3'-5' direction. In nature, DNA binding and cleavage usually require both a protein and two RNAs. However, a single guide RNA (abbreviated as "sgRNA" or simply "gRNA") can be engineered to incorporate both aspects of crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012) (the entire contents of which are hereby incorporated by reference). are included. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, to correctly process pre-crRNA, trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein are required. TracrRNA functions as a guide for ribonuclease 3-assisted processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to crRNA is first endonucleolytically cleaved and then trimmed exonucleolytically in the 3'-5' direction. In nature, DNA binding and cleavage usually require both a protein and two RNAs. However, a single guide RNA (abbreviated as "sgRNA" or simply "gRNA") can be engineered to incorporate both aspects of crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012) (the entire contents of which are hereby incorporated by reference). are included. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, to correctly process pre-crRNA, trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein are required. TracrRNA functions as a guide for ribonuclease 3-assisted processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to crRNA is first endonucleolytically cleaved and then trimmed exonucleolytically in the 3'-5' direction. In nature, DNA binding and cleavage usually require both a protein and two RNAs. However, a single guide RNA (abbreviated as "sgRNA" or simply "gRNA") can be engineered to incorporate both aspects of crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012) (the entire contents of which are hereby incorporated by reference). are included. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, to correctly process pre-crRNA, trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein are required. TracrRNA functions as a guide for ribonuclease 3-assisted processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to crRNA is first endonucleolytically cleaved and then trimmed exonucleolytically in the 3'-5' direction. In nature, DNA binding and cleavage usually require both a protein and two RNAs. However, a single guide RNA (abbreviated as "sgRNA" or simply "gRNA") can be engineered to incorporate both aspects of crRNA and tracrRNA into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012) (the entire contents of which are hereby incorporated by reference). See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012) (the entire contents of which are hereby incorporated by reference). (which is incorporated herein by reference). Cas9 recognizes short motifs (PA M or protospacer adjacent motif) of the CRISPR repeat sequence and helps to distinguish self from non-self .

[0174] In some embodiments, the methods described herein can utilize a modified Cas protein . A guide RNA (gRNA) is a short synthetic RNA consisting of a scaffold sequence necessary for Cas binding and a user-defined ~20 nucleotide spacer that defines the genomic ( or polynucleotide, e.g., DNA or RNA) target to be modified. Thus, one of ordinary skill in the art can change the genomic or polynucleotide target of the Cas protein by changing the target sequence present in the gRNA. The specificity of the Cas protein depends in part on how specific the gRNA target sequence is for the genomic polynucleotide target as compared to other parts of the genome .

[0175] In some embodiments, the gRNA scaffold sequence is as follows: GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGG UGCUUUU.

[0176] In one embodiment, the RNA scaffold contains a stem-loop. In one embodiment, the RNA scaffold is the nucleic acid sequence: GUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAG . In one embodiment, the RNA scaffold contains AUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG. In one embodiment, the RNA scaffold is the nucleic acid sequence: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAG contains UCGGUGCUUUU.

[0177] In one embodiment, the sgRNA scaffold polynucleotide sequence of S. pyrogenes is as follows: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC.

[0178] In one embodiment, the sgRNA scaffold polynucleotide sequence of S. aureus is as follows: GUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGC GAGA.

[0179] In one embodiment, the sgRNA scaffold of BhCas12b has the following polynucleotide sequence: GUUCUGTCUUUUGGUCAGGACAACCGUCUAGCUAUAAGUGCUGCAGGGUGUGAGAAACUCCUAUUGCUGGACGAUGUCUC UUACGAGGCAUUAGCAC.

[0180] In one embodiment, the BvCas12b sgRNA scaffold has the following polynucleotide sequence: GACCU AUAGGGUCAAUGAAUCUGUGCGUGUGCCAUAAGUAAUUAAAAAUUACCCACCACAGGAGCACCUGAAAACAGGUGCUUGG CAC.

[0181] In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is and can bind to a target polynucleotide when combined with a bound guide nucleic acid is an endonuclease (e.g., deoxyribonuclease or ribonuclease) In some embodiments, a CRISPR protein-derived domain incorporated into a base editor is a nickase that can bind to a target polynucleotide when combined with a bound guide nucleic acid In some embodiments, a CRISPR protein-derived domain incorporated into a base editor is a catalytically inactive domain that can bind to a target polynucleotide when combined with a bound guide nucleic acid In some embodiments, the target polynucleotide to which the CRISPR protein-derived domain of the base editor binds is DNA In some embodiments, the target polynucleotide to which the CRISPR protein-derived domain of the base editor binds is RNA

[0182] Cas proteins that can be used herein include Class 1 and Class 2. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, C as5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10, Cs y1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn 2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2 , Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Cs f4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf 1. Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12 i, CARF, DinG, their homologs, or modified or altered versions thereof. Two Functional endonuclease domains: Unmodified CRISPR enzymes such as Cas9 with RuvC and HNH can have DNA cleavage activity. CRISPR enzymes can direct cleavage of one or both strands at a target sequence, e.g., within the target sequence and / or within the complementary sequence of the target sequence. For example, a CRISPR enzyme can direct cleavage of one or both strands within about 1, 2, 3, 4, 5, 6 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence. A mutant CRISPR enzyme can be encoded by a vector in which mutations have been introduced with respect to the corresponding wild - type enzyme so that it lacks the ability to cleave one or both strands of a target polynucleotide containing the target sequence. Cas9 can refer to a polypeptide having at least or at least about 50%, 60%, 70%

[0183] 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology to an exemplary wild - type Cas9 polypeptide (e.g., Cas9 from S. pyogenes). Cas9 can refer to a polypeptide having at most or at most about 50%, 6 to an exemplary wild - type Cas9 polypeptide (e.g., from S. pyogenes) and having at most or at most about 50%, 6 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology to an exemplary wild - type Cas9 polypeptide (e.g., from S. pyogenes). and / or sequence homology to an exemplary wild - type Cas9 polypeptide (e.g., from S. pyogenes). Cas9 can refer to a polypeptide having at most or at most about 50%, 6 to an exemplary wild - type Cas9 polypeptide (e.g., from S. pyogenes). an array of 0%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% can refer to a polypeptide having identity and / or sequence homology. Cas9 is a deletion, insertion, substitution, variant, mutation, fusion, chimera, or any combination thereof, such as an amino acid change, of a wild-type or modified form of the Cas9 protein.

[0184] In some embodiments, the CRISPR protein-derived domain of the base editor is Corynebact erium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (N CBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1 ); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquis (NCBI Ref: NC_018721.1); Streptococcus ther mophilus (NCBI Ref: YP_820832.1); Listeria innocua (NCBI Ref: NP_472073.1); Campylob acter jejuni (NCBI Ref: YP_002344900.1); Neisseria meningitidis (NCBI Ref: YP_00234 2100.1), may include all or a part of Cas9 derived from Streptococcus pyogenes or Staphylococcus aureus.

[0185] The Cas9 domain of the nucleobase editor The sequences and structures of Cas9 nucleases are well-known to those skilled in the art (e.g., ”Complete genome s equence of an Ml strain of Streptococcus pyogenes.” Ferretti et al., J.J., McSh an W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C,Sezate S., Suvoro v A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z., Ren Q., Zh u H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl.Acad.Sci.U.S.A.98:4658 - 4663 (2001); ”CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma CM., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Vogel J., Charpentier E., Natur e 471:602 - 607 (2011), and ”A programmable dual-RNA-guided DNA endonuclease in a "adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doud na J.A., Charpentier E. Science 337:816-821 (2012) (see these in their entirety, which are incorporated herein by reference)). Cas9 orthologs have been described in various species, including but not limited to S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on the present disclosure, and such Cas9 nucleases and sequences include, for example, Cas9 sequences derived from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The trac rRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5,726-737; the entire contents of which are incorporated herein by reference. In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is a Cas9 domain in. Non-limiting exemplary Cas9 domains are provided herein. A Cas9 domain can be a nuclease-active Cas9 domain, a nuclease-inactive Cas9 domain, or a Cas9 nickase. In some embodiments, the Cas9 domain is a nuclease-active domain main. For example, the Cas9 domain can be a Cas9 domain that cleaves both strands of a double-stranded nucleic acid (e.g., both strands of a double-stranded DNA molecule). In some embodiments, the Cas9 domain

[0186] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is a Cas9 domain in. Non-limiting exemplary Cas9 domains are provided herein. A Cas9 domain can be a nuclease-active Cas9 domain, a nuclease-inactive Cas9 domain, or a Cas9 nickase. In some embodiments, the Cas9 domain is a nuclease-active domain main. For example, the Cas9 domain can be a Cas9 domain that cleaves both strands of a double-stranded nucleic acid (e.g., both strands of a double-stranded DNA molecule). In some embodiments, the Cas9 domain can be a Cas9 domain that cleaves both strands of a double-stranded nucleic acid (e.g., both strands of a double-stranded DNA molecule). In some embodiments, the Cas9 domain It contains any one of the amino acid sequences described in this specification. In some embodiments, C The Cas9 domain has at least 60 %, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% amino acid sequence identity with any one of the amino acid sequences described in this specification. In some embodiments The Cas9 domain contains an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any one of the amino acid sequences described in this specification. In some embodiments, the Cas9 domain contains an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 10 0, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical consecutive amino acid residues compared to any one of the amino acid sequences described in this specification. In some embodiments, proteins containing fragments of Cas9 are provided. For example, some

[0187] embodiments provide proteins containing fragments of Cas9. For example, some In an embodiment, the protein comprises one of two Cas9 domains: (1) the gRNA binding domain of Cas9; and (2) the DNA cleavage domain of Cas9. In some embodiments, the protein comprising Cas9 or a fragment thereof is referred to as a "Cas9 variant". The Cas9 variant shares homology with Cas9 or a fragment thereof. For example, the Cas9 variant is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to wild-type Cas9. In some embodiments , the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 1 2, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 3 2, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type Cas9. In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., the gRNA binding domain or the DNA cleavage domain), such that the fragment is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99. 9% identical to the corresponding fragment of wild-type Cas9. In some embodiments, the fragment is at least 30 of the corresponding wild-type Cas9 amino acids in length. In some embodiments, the fragment is at least 30 amino acids in length. amino acids in length. %, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% amino acid length. In some embodiments the fragment is at least 100 amino acids in length. In some embodiments the fragment is , at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750 , 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 a mino acid length.

[0188] In some embodiments, the Cas9 fusion proteins provided herein comprise the full-length amino acid sequence of the Cas9 protein, e.g., one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein comprise only one or more fragments thereof. Appropriate Cas9 domains and exemplary amino acid sequences of Cas9 fragments are provided herein, and additional appropriate sequences of Cas9 domains and fragments will be apparent to those of skill in the art.

[0189] The Cas9 protein can bind to a guide RNA that guides the Cas9 protein to a specific DNA sequence that is complementary to the guide RNA. In some embodiments, the polynucleotide programmable nucleotide binding domain is a Cas9 domain, e.g., a nuclease active domain, for example, a nuclease active It is wild-type Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Nucleic acid Examples of programmable DNA-binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), CasX , CasY, Cpf1, Cas12b / C2C1, and Cas12c / C2C3.

[0190] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_017053.1, the following nucleotide and amino acid sequences). ATGGATAAGAAATACTCAATAGGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGATTATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG GCAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGCAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAATCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTAGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAGAAATGGCTTGTTTGGGAAT CTCATTGCTTTGTCATTGGGATTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATAGTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAGCGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAGGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGCGCCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGGGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGATATTCAAAAAGCACAGGTGTCTGGACAAGGCCATAGTTTA CATGAACAGATTGCTAACTTAGCTGGCAGTCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAATTGTTGATGAACT GGTCAAAGTAATGGGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCCAGA AAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCTGTT GAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTACAAAATGGAAGAGACATGTATGTGGACCAAGAATT AGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCATTAAAGACGATTCAATAGACAATA AGGTACTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAAAAC TATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGGTTT GAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAATTT TGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCTAAA TTAGTTTCTGACTTCCGAAAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGCGTA TCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATAAAG TTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCTAAT ATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGGGGA AACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATATTG TCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTACCAAAAAGAAATTCGGACAAGCTTATTGCT CGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGCTAA GGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCTTTG AAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAAAGACTTAATCATTAAACTACCTAAATAT AGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAAAAGGAAATGAGCTGGCTCT GCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAACAAA...

Claims

1. 1. A method for in vitro or ex vivo editing of nucleobases of a Hepatitis B virus (HBV) genome, the method comprising contacting the HBV genome with one or more guide RNAs and a base editor comprising a nuclease-inactive or nickase Cas9 domain and an adenosine deaminase domain, wherein the one or more guide RNAs target the base editor to modify the nucleobases of the HBV genome, and wherein the one or more guide RNAs are one or more of the following: CACCACGAGUCUAGACUCUG;UCAAUCCCAACAAGGACACC;GGGAACAAGAUCUACAGCAU; AAGCCCAGGAUGAUGGGAUG; CUGCCAACUGGAUCCUGCGC; GACACAUCCAGCGAUAACCA; GCUGCCAACUGGAUCCUGCG; UAUGGAUGAUGUGGUAUUG; CCAUGCCCCAAAGCCACCCA; AAGCCACCCAAGGCACAGCU;GAGAAGUCCACCACGAGUCU; CUUCUCUCAAUUUUCUAGG;GACGACGAGGCAGGUCCCCU; AGGAGUUCCGCAGUAUGGAU;CCGCAGUAUGGAUCGGCAGA; CCUCUGCCGAUCCAUACUGC; CGCCCACCGAAUGUUGCCCA; GACUUCUCUCAAUUUUCUAG; GUUCCGCAGUAUGGAUCGGC; UACUAACAUUGAGGUUCCCG;UCCGCAGUAUGGAUCGGCAG; and UCCUCUGCCGAUCCAUACUG and a sequence selected from the modification of the nucleobase in the polynucleotide encoding the HBV protein results in a missense mutation in the HBV pol gene, HBV core gene, HBV X gene, or HBV S gene; The editing method.

2. The method described in claim 1, wherein the adenosine deaminase converts target A and T to G and C in a polynucleotide encoding an HBV protein.

3. The method of claim 1 or 2, wherein the missense mutation is in the HBV S gene.

4. the Cas9 domain is selected from Streptococcus pyogenes Cas9, Staphylococcus aureus Cas9, Streptococcus thermophilus 1 Cas9, and Streptococcus canis Cas9; or the Cas9 domain has protospacer adjacent motif specificity for a nucleic acid sequence selected from 5'-NGG-3', 5'-NAG-3', 5'-NGA-3', 5'-NAA-3', 5'-NNAGGA-3', and 5'-NNACCA-3'; or the Cas9 domain has an altered protospacer adjacent motif specificity; or the Cas9 domain has an altered protospacer adjacent motif specificity, and the nucleic acid sequence of the altered protospacer adjacent motif is selected from 5'-NNNRRT-3', NGA-3', 5'-NGCG-3', 5'-NGN-3', NGCN-3', 5'-NGTN-3', and 5'-NAA-3'; The method according to any one of claims 1 to 3.

5. (i) a fusion protein, or a polynucleotide encoding said fusion protein, comprising a nuclease-inactive or nickase Cas9 domain and a base editor domain that is an adenosine deaminase domain, and one or more guide polynucleotides that target the base editor domain to result in modification of a polynucleotide encoding an HBV protein; or (ii) one or more polynucleotides encoding a nuclease-inactive or nickase Cas9 domain and a base editor domain that is an adenosine deaminase domain, and one or more guide polynucleotides that target the base editor domain to result in modification of a polynucleotide encoding an HBV protein.

1. A pharmaceutical composition for treating a hepatitis B virus (HBV) infection in a subject, comprising: The one or more guide polynucleotides are one or more of the following: CACCACGAGUCUAGACUCUG;UCAAUCCCAACAAGGACACC;GGGAACAAGAUCUACAGCAU; AAGCCCAGGAUGAUGGGAUG; CUGCCAACUGGAUCCUGCGC; GACACAUCCAGCGAUAACCA; GCUGCCAACUGGAUCCUGCG; UAUGGAUGAUGUGGUAUUG; CCAUGCCCCAAAGCCACCCA; AAGCCACCCAAGGCACAGCU;GAGAAGUCCACCACGAGUCU; CUUCUCUCAAUUUUCUAGG;GACGACGAGGCAGGUCCCCU; AGGAGUUCCGCAGUAUGGAU;CCGCAGUAUGGAUCGGCAGA; CCUCUGCCGAUCCAUACUGC; CGCCCACCGAAUGUUGCCCA; GACUUCUCUCAAUUUUCUAG;GUUCCGCAGUAUGGAUCGGC; UACUAACAUUGAGGUUCCCG;UCCGCAGUAUGGAUCGGCAG; and UCCUCUGCCGAUCCAUACUG and the modification of the nucleobase in the polynucleotide encoding the HBV protein results in a missense mutation in the HBV pol gene, HBV core gene, HBV X gene, or HBV S gene; Pharmaceutical compositions.

6. The pharmaceutical composition described in claim 5, wherein the adenosine deaminase converts the target A / T in the polynucleotide encoding the HBV protein to G / C.

7. 7. The pharmaceutical composition of claim 5 or 6, which is delivered to cells of a mammalian subject and / or delivered to hepatocytes.

8. the Cas9 domain is selected from Streptococcus pyogenes Cas9, Staphylococcus aureus Cas9, Streptococcus thermophilus 1 Cas9, and Streptococcus canis Cas9; and / or the Cas9 domain has protospacer adjacent motif specificity for a nucleic acid sequence selected from 5'-NGG-3', 5'-NAG-3', 5'-NGA-3', 5'-NAA-3', 5'-NNAGGA-3', and 5'-NNACCA-3'; and / or the Cas9 domain has altered protospacer adjacent motif specificity; and / or the Cas9 domain has an altered protospacer adjacent motif specificity, and the nucleic acid sequence of the altered protospacer adjacent motif is selected from 5'-NNNRRT-3', NGA-3', 5'-NGCG-3', 5'-NGN-3', NGCN-3', 5'-NGTN-3', and 5'-NAA-3'; The pharmaceutical composition according to any one of claims 5 to 7.

9. the Cas9 domain comprises the amino acid substitution D10A or a corresponding amino acid substitution; The pharmaceutical composition according to any one of claims 5 to 8.

10. A pharmaceutical composition described in any one of claims 5 to 9, wherein the missense mutation is in the HBV S gene.

11. 1. A composition comprising a base editor bound to a guide RNA, wherein the guide RNA comprises a nucleic acid sequence complementary to an HBV gene, the base editor comprises a Cas9 domain that is nuclease-inactive or a nickase and an adenosine deaminase domain, and the guide RNA comprises one or more of the following: CACCACGAGUCUAGACUCUG;UCAAUCCCAACAAGGACACC;GGGAACAAGAUCUACAGCAU; AAGCCCAGGAUGAUGGGAUG; CUGCCAACUGGAUCCUGCGC; GACACAUCCAGCGAUAACCA; GCUGCCAACUGGAUCCUGCG; UAUGGAUGAUGUGGUAUUG; CCAUGCCCCAAAGCCACCCA; AAGCCACCCAAGGCACAGCU;GAGAAGUCCACCACGAGUCU; CUUCUCUCAAUUUUCUAGG;GACGACGAGGCAGGUCCCCU; AGGAGUUCCGCAGUAUGGAU;CCGCAGUAUGGAUCGGCAGA; CCUCUGCCGAUCCAUACUGC; CGCCCACCGAAUGUUGCCCA; GACUUCUCUCAAUUUUCUAG; GUUCCGCAGUAUGGAUCGGC; UACUAACAUUGAGGUUCCCG;UCCGCAGUAUGGAUCGGCAG; and UCCUCUGCCGAUCCAUACUG; and a sequence selected from the modification of the nucleobase in the polynucleotide encoding the HBV protein results in a missense mutation in the HBV pol gene, HBV core gene, HBV X gene, or HBV S gene; The composition.

12. further comprising a lipid, optionally wherein the lipid is a cationic lipid; and / or further comprising a pharmaceutically acceptable excipient, The composition of claim 11.

13. A pharmaceutical composition for the treatment of HBV infection, comprising: In a pharmaceutically acceptable excipient: (i) a base editor comprising a nuclease-inactive or nickase Cas9 domain and an adenosine deaminase domain, or a nucleic acid encoding the base editor; and (ii) one or more guide RNAs (gRNAs) comprising nucleic acid sequences complementary to an HBV gene, wherein the one or more gRNAs are one or more of the following: CACCACGAGUCUAGACUCUG;UCAAUCCCAACAAGGACACC;GGGAACAAGAUCUACAGCAU; AAGCCCAGGAUGAUGGGAUG; CUGCCAACUGGAUCCUGCGC; GACACAUCCAGCGAUAACCA; GCUGCCAACUGGAUCCUGCG; UAUGGAUGAUGUGGUAUUG; CCAUGCCCCAAAGCCACCCA; AAGCCACCCAAGGCACAGCU;GAGAAGUCCACCACGAGUCU; CUUCUCUCAAUUUUCUAGG;GACGACGAGGCAGGUCCCCU; AGGAGUUCCGCAGUAUGGAU;CCGCAGUAUGGAUCGGCAGA; CCUCUGCCGAUCCAUACUGC; CGCCCACCGAAUGUUGCCCA; GACUUCUCUCAAUUUUCUAG; GUUCCGCAGUAUGGAUCGGC; UACUAACAUUGAGGUUCCCG;UCCGCAGUAUGGAUCGGCAG; and UCCUCUGCCGAUCCAUACUG; and a sequence selected from the modification of the nucleobase in the polynucleotide encoding the HBV protein results in a missense mutation in the HBV pol gene, HBV core gene, HBV X gene, or HBV S gene; The pharmaceutical composition.

14. 14. The pharmaceutical composition of claim 13 for the treatment of HBV infection.

15. an mRNA encoding the base edit, and a 5' to 3' nucleic acid sequence selected from the group consisting of: CACCACGAGUCUAGACUCUG;AAGCCCAGGAUGAUGGGAUG;GACACAUCCAGCGAUAACCA;GAGAAGUCCACCACGAGUCU;CUUCUCUCAAUUUUCUAGGG;GACGACGAGGCAGGUCCCCU;CCGCAGUAUGGAUCGGCAGA;CCUCUGCCGAUCCAUACUGC;GUUCCGCAGUAUGGAUCGGC;and UACUAACAUUGAGGUUCCCG; A pharmaceutical composition comprising a guide RNA comprising: The pharmaceutical composition is for use in targeting the base editor to modify a nucleic acid base of the HBV genome, wherein the base editor comprises a nuclease-inactive or nickase Cas9 domain and an adenosine deaminase domain; the modification of the nucleobase in the polynucleotide encoding the HBV protein results in a missense mutation in the HBV pol gene, HBV core gene, or HBV S gene; Pharmaceutical compositions.

16. 16. The pharmaceutical composition of claim 15, further comprising a lipid.