Deaminase and variant thereof for base editing
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ACCUREDIT THERAPEUTICS (SUZHOU) CO LTD
- Filing Date
- 2025-11-14
- Publication Date
- 2026-05-21
Smart Images

Figure CN2025135157_21052026_PF_FP_ABST
Abstract
Description
Deaminases and their variants used for base editing
[0001] Priority information
[0002] This application claims priority to PCT / CN2024 / 132069, filed on November 14, 2024, which is incorporated herein by reference in its entirety. Technical Field
[0003] This application relates to deaminases for use with, for example, nucleic acid-guided nucleases to modify target nucleic acid sequences. More specifically, this application relates to adenosine deaminases and cytidine deaminases with high base editing efficiency. Background Technology
[0004] Targeted editing of nucleic acid sequences, such as targeted cleavage or modification of genomic DNA, is an effective method for studying gene function and has the potential to provide new treatments for human genetic diseases. Currently available base editors include cytidine base editors (e.g., CBE4) that convert target C·G base pairs to T·A base pairs, and adenosine base editors (e.g., ABE7.10) that convert A·T base pairs to G·C base pairs. There is a need in the art for improved base editors capable of inducing modifications within target sequences with greater specificity and efficiency. Summary of the Invention
[0005] In a first aspect, this application provides an adenosine deaminase comprising:
[0006] (i) an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any one of SEQ ID NO:1-47 (e.g., SEQ ID NO:2-47); or
[0007] (ii) Select one or more changes from the group consisting of D12X, H21X, A61X, V95X, S155X, S156X, C159X, E164X, E167X, and K169X relative to the position of SEQ ID NO:1, wherein X is any amino acid. Preferably, select one or more substitutions from the group consisting of D12H, H21D or H21Q, A61C, V95T, S155R, S156N or S156K, C159R, E164V or E164W, E167R or E167W, and K169N relative to the position of SEQ ID NO:1.
[0008] In some embodiments, the adenosine deaminase comprises:
[0009] (i) amino acids having 80% to 99.5% (e.g., 85% to 99.5%, 90% to 99.5%, or 95% to 99.5%) sequence identity with SEQ ID NO:1, and
[0010] (ii) Select one or more changes from the group consisting of D12X, H21X, A61X, V95X, S155X, S156X, C159X, E164X, E167X, and K169X relative to the position of SEQ ID NO:1, wherein X is any amino acid. Preferably, select one or more substitutions from the group consisting of D12H, H21D or H21Q, A61C, V95T, S155R, S156N or S156K, C159R, E164V or E164W, E167R or E167W, and K169N relative to the position of SEQ ID NO:1.
[0011] In some embodiments, adenosine deaminase can deaminate adenosine in DNA. In other embodiments, adenosine deaminase can deaminate cytosine in DNA, but with lower efficiency.
[0012] In some embodiments, the adenosine deaminase comprises, or consists of, the amino acid sequence of any one of SEQ ID NO:1-47 (e.g., SEQ ID NO:2-47).
[0013] In some embodiments, the adenosine deaminase comprises, or consists of, the amino acid sequence of any one of SEQ ID NO:1-19 (e.g., SEQ ID NO:2-19).
[0014] In some embodiments, the adenosine deaminase comprises, or consists of, the amino acid sequence of any one of SEQ ID NO:20-47.
[0015] In a second aspect, this application provides a cytidine deaminase comprising:
[0016] (i) an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any one of SEQ ID NO:48-63, 64-99, or 118-136 (e.g., SEQ ID NO:48-63, 65-99, or 118-136); or
[0017] (ii) Selecting one or more changes from the group consisting of A17X, H34X, D36X, I46X, G64X, G68X, Y70X, L73X, V79X, F81X, M91X, H93X, G102X, A103X, A106X, D112X, A119X, N124X, R130X, S139X, A140X, C143X, R151X, and K153X relative to SEQ ID NO:64, wherein X is any amino acid, preferably selected relative to SEQ ID One or more substitutions from the group consisting of A17N, H34N, D36L, I46Q, G64C, G68T, Y70G or Y70H or Y70K, L73F or L73R, V79T, F81L, M91L, H93N, G102S, A103V, A106R, D112N, A119G, N124V or N124P, R130V or R130T, S139K, A140R, C143A or C143S, R151M, K153G or K153R of SEQ ID NO:64, wherein the amino acid position is relative to the number of SEQ ID NO:64.
[0018] In some embodiments, the cytidine deaminase comprises:
[0019] (i) amino acids having 80% to 99.5% (e.g., 85% to 99.5%, 90% to 99.5%, or 95% to 99.5%) sequence identity with SEQ ID NO:64, and
[0020] (ii) Selecting one or more changes from the group consisting of A17X, H34X, D36X, I46X, G64X, G68X, Y70X, L73X, V79X, F81X, M91X, H93X, G102X, A103X, A106X, D112X, A119X, N124X, R130X, S139X, A140X, C143X, R151X, and K153X relative to SEQ ID NO:64, wherein X is any amino acid, preferably selected relative to SEQ ID One or more substitutions from the group consisting of A17N, H34N, D36L, I46Q, G64C, G68T, Y70G or Y70H or Y70K, L73F or L73R, V79T, F81L, M91L, H93N, G102S, A103V, A106R, D112N, A119G, N124V or N124P, R130V or R130T, S139K, A140R, C143A or C143S, R151M, K153G or K153R of SEQ ID NO:64, wherein the amino acid position is relative to the number of SEQ ID NO:64.
[0021] In some embodiments, cytidine deaminase can deaminate cytidine in DNA. In some embodiments, cytidine deaminase can deaminate adenosine in DNA, but with lower efficiency.
[0022] In some embodiments, the cytidine deaminase comprises, or is composed of, the amino acid sequence of any one of SEQ ID NO:48-63, 64-99 or 118-136 (e.g., SEQ ID NO:48-63, 65-99 or 118-136).
[0023] In a third aspect, this application provides a nucleobase editor comprising:
[0024] (i) a multinucleotide programmable DNA binding domain; and
[0025] (ii) a deaminase, wherein the deaminase is the adenosine deaminase described in the first aspect or the cytidine deaminase described in the second aspect.
[0026] In some implementations, the nucleobase editor also includes one or more uracil glycosylase inhibitor (UGI) domains.
[0027] In some embodiments, the polynucleotide-programmable DNA-binding domain and the deaminase can form a fusion protein. In other embodiments, the polynucleotide-programmable DNA-binding domain, the deaminase, and the UGI domain can form a fusion protein. For example, the nucleotide sequence encoding the deaminase, the nucleotide sequence encoding the polynucleotide-programmable DNA-binding domain, and optionally the nucleotide sequence encoding the UGI domain can be operatively linked together in a reading frame conformal manner to enable expression of the fusion protein. In some embodiments, the nucleotide sequence encoding the fusion protein can be encapsulated in lipid nanoparticles (LNPs), or can be mixed with a corresponding gRNA and then encapsulated in lipid nanoparticles (LNPs).
[0028] In some embodiments, the nucleotide sequences encoding the polynucleotide programmable DNA-binding domain and the deaminase are mixed with the nucleotide sequence encoding the UGI domain and then encapsulated into lipid nanoparticles (LNPs). That is, the LNP contains at least two nucleotide sequences: a first nucleotide sequence encoding the deaminase and the polynucleotide programmable DNA-binding domain, and a second nucleotide sequence encoding the UGI domain, which are mixed in a specific mass ratio (e.g., 7:1) and then encapsulated into the lipid nanoparticles (LNPs). Optionally, the LNP also contains a corresponding gRNA. For example, the first nucleotide sequence, the second nucleotide sequence, and the gRNA are mixed and then encapsulated into the LNP.
[0029] In some embodiments, the nucleobase editor is an adenosine base editor (ABE), and the deaminase is an adenosine deaminase. In a preferred embodiment, the adenosine deaminase comprises:
[0030] (i) an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any one of SEQ ID NO:1-47 (e.g., SEQ ID NO:2-47); or
[0031] (ii) Select one or more changes from the group consisting of D12X, H21X, A61X, V95X, S155X, S156X, C159X, E164X, E167X, and K169X relative to the position of SEQ ID NO:1, wherein X is any amino acid. Preferably, select one or more substitutions from the group consisting of D12H, H21D or H21Q, A61C, V95T, S155R, S156N or S156K, C159R, E164V or E164W, E167R or E167W, and K169N relative to the position of SEQ ID NO:1.
[0032] In some embodiments, the nucleobase editor is a cytidine base editor (CBE). In a preferred embodiment, the cytidine deaminase comprises:
[0033] (i) an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any one of SEQ ID NO:48-63, 64-99, or 118-136 (e.g., SEQ ID NO:48-63, 65-99, or 118-136); or
[0034] (ii) Selecting one or more changes from the group consisting of A17X, H34X, D36X, I46X, G64X, G68X, Y70X, L73X, V79X, F81X, M91X, H93X, G102X, A103X, A106X, D112X, A119X, N124X, R130X, S139X, A140X, C143X, R151X, and K153X relative to SEQ ID NO:64, wherein X is any amino acid, preferably selected relative to SEQ ID One or more substitutions from the group consisting of A17N, H34N, D36L, I46Q, G64C, G68T, Y70G or Y70H or Y70K, L73F or L73R, V79T, F81L, M91L, H93N, G102S, A103V, A106R, D112N, A119G, N124V or N124P, R130V or R130T, S139K, A140R, C143A or C143S, R151M, K153G or K153R of SEQ ID NO:64, wherein the amino acid position is relative to the number of SEQ ID NO:64.
[0035] In some embodiments, the polynucleotide-programmable DNA-binding domain comprises a domain selected from the group consisting of a Cas9 domain, a Cas12 domain, a TnpB domain, an IscB domain, a homing meganuclease domain, a zinc finger DNA-binding domain, and a transcription activator-like effector (TALE) DNA-binding domain. In some embodiments, the polynucleotide-programmable DNA-binding domain comprises a Cas9 domain. In some embodiments, the Cas9 domain is Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), Streptococcus pyogenes Cas9 (SpCas9), or a variant thereof. In some embodiments, the polynucleotide-programmable DNA-binding domain is an inactive nuclease or a nickase variant.
[0036] In some embodiments, the polynucleotide programmable DNA-binding domain is a Cas9 domain, preferably a mutated Cas9 containing one or two amino acid substitutions selected from D700R or E1243R, more preferably a Cas9 containing D700R or a Cas9 containing E1243R. In a preferred embodiment, the mutated Cas9 contains, or is composed of, the amino acid sequence of SEQ ID NO: 115 or 116.
[0037] In some embodiments, the base editor comprises a nucleic acid-guided nuclease protein and an embedded nucleotide base editor (NBE) domain. In some embodiments, the embedded NBE domain is an adenosine base editor (ABE) domain. In some embodiments, the embedded ABE domain is an embedded adenosine deaminase protein domain. In some embodiments, the polynucleotide programmable DNA-binding domain comprises a linker flanked by at least one embedded domain protein. In some embodiments, the embedded NBE domain is a cytidine base editor (CBE) domain. In some embodiments, the embedded CBE domain is an embedded cytidine deaminase protein domain.
[0038] In some embodiments, the polynucleotide programmable DNA-binding domain associates with the deaminase, for example, by direct linkage or by linkage via a adapter. Typically, those skilled in the art can select a suitable adapter. In one embodiment, the adapter comprises, or consists of, any of the amino acid sequences in SEQ ID NO:104-107. Preferably, the adapter is SEQ ID NO:106.
[0039] In some implementations, the nucleobase editor may be in the form of a kit.
[0040] In a fourth aspect, this application provides a fusion protein comprising a polynucleotide-programmable DNA-binding domain and at least one nucleobase editor domain containing a deaminase, linked together. The polynucleotide-programmable DNA-binding domain and the at least one nucleobase editor domain containing a deaminase can be directly linked or linked via a adapter.
[0041] Typically, those skilled in the art can select a suitable adapter. In one embodiment, the adapter comprises, or consists of, any of the amino acid sequences in SEQ ID NO:104-107. Preferably, the adapter is SEQ ID NO:106.
[0042] In some embodiments, the polynucleotide programmable DNA-binding domain is a Cas9 domain, preferably a mutated Cas9 containing one or more amino acid mutations (e.g., amino acid substitutions), such as a Cas9 containing one or two amino acid substitutions selected from D700R or E1243R, more preferably a Cas9 containing D700R or a Cas9 containing E1243R. In a preferred embodiment, the mutated Cas9 contains, or is composed of, the amino acid sequence of SEQ ID NO: 115 or 116.
[0043] In some embodiments, the deaminase is the adenosine deaminase or cytidine deaminase described in this invention.
[0044] In a fifth aspect, this application provides a polynucleotide encoding an adenosine deaminase, said adenosine deaminase comprising:
[0045] (i) an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any one of SEQ ID NO:1-47 (e.g., SEQ ID NO:2-47); or
[0046] (ii) Select one or more changes from the group consisting of D12X, H21X, A61X, V95X, S155X, S156X, C159X, E164X, E167X, and K169X relative to the position of SEQ ID NO:1, wherein X is any amino acid. Preferably, select one or more substitutions from the group consisting of D12H, H21D or H21Q, A61C, V95T, S155R, S156N or S156K, C159R, E164V or E164W, E167R or E167W, and K169N relative to the position of SEQ ID NO:1.
[0047] In some embodiments, the adenosine deaminase comprises:
[0048] (i) amino acids having 80% to 99.5% (e.g., 85% to 99.5%, 90% to 99.5%, or 95% to 99.5%) sequence identity with SEQ ID NO:1, and
[0049] (ii) Select one or more changes from the group consisting of D12X, H21X, A61X, V95X, S155X, S156X, C159X, E164X, E167X, and K169X relative to the position of SEQ ID NO:1, wherein X is any amino acid. Preferably, select one or more substitutions from the group consisting of D12H, H21D or H21Q, A61C, V95T, S155R, S156N or S156K, C159R, E164V or E164W, E167R or E167W, and K169N relative to the position of SEQ ID NO:1.
[0050] In some implementations, the polynucleotide may be codon-optimized for better expression in host cells.
[0051] In a sixth aspect, this application provides a polynucleotide encoding a cytidine deaminase, said cytidine deaminase comprising:
[0052] (i) an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any one of SEQ ID NO:48-63, 64-99, or 118-136 (e.g., SEQ ID NO:48-63, 65-99, or 118-136); or
[0053] (ii) Selecting one or more changes from the group consisting of A17X, H34X, D36X, I46X, G64X, G68X, Y70X, L73X, V79X, F81X, M91X, H93X, G102X, A103X, A106X, D112X, A119X, N124X, R130X, S139X, A140X, C143X, R151X, and K153X relative to SEQ ID NO:64, wherein X is any amino acid, preferably selected relative to SEQ ID One or more substitutions from the group consisting of A17N, H34N, D36L, I46Q, G64C, G68T, Y70G or Y70H or Y70K, L73F or L73R, V79T, F81L, M91L, H93N, G102S, A103V, A106R, D112N, A119G, N124V or N124P, R130V or R130T, S139K, A140R, C143A or C143S, R151M, K153G or K153R of SEQ ID NO:64, wherein the amino acid position is relative to the number of SEQ ID NO:64.
[0054] In some embodiments, the cytidine deaminase comprises:
[0055] (i) amino acids having 80% to 99.5% (e.g., 85% to 99.5%, 90% to 99.5%, or 95% to 99.5%) sequence identity with SEQ ID NO:64, and
[0056] (ii) Selecting one or more changes from the group consisting of A17X, H34X, D36X, I46X, G64X, G68X, Y70X, L73X, V79X, F81X, M91X, H93X, G102X, A103X, A106X, D112X, A119X, N124X, R130X, S139X, A140X, C143X, R151X, and K153X relative to SEQ ID NO:64, wherein X is any amino acid, preferably selected relative to SEQ ID One or more substitutions from the group consisting of A17N, H34N, D36L, I46Q, G64C, G68T, Y70G or Y70H or Y70K, L73F or L73R, V79T, F81L, M91L, H93N, G102S, A103V, A106R, D112N, A119G, N124V or N124P, R130V or R130T, S139K, A140R, C143A or C143S, R151M, K153G or K153R of SEQ ID NO:64, wherein the amino acid position is relative to the number of SEQ ID NO:64.
[0057] In some implementations, the polynucleotide may be codon-optimized for better expression in host cells.
[0058] In a seventh aspect, this application provides an expression vector comprising a polynucleotide encoding the adenosine deaminase of the present invention. In some embodiments, the expression vector is a viral vector selected from the group consisting of adeno-associated virus (AAV), retroviral vectors, adenovirus vectors, lentiviral vectors, Sendai virus vectors, and herpesvirus vectors.
[0059] In an eighth aspect, this application provides an expression vector comprising a polynucleotide encoding the cytidine deaminase of the present invention. In some embodiments, the expression vector is a viral vector selected from the group consisting of adeno-associated virus (AAV), retroviral vectors, adenovirus vectors, lentiviral vectors, Sendai virus vectors, and herpesvirus vectors.
[0060] In a ninth aspect, this application provides a polynucleotide encoding a nucleobase editor, the nucleobase editor comprising: (i) a polynucleotide programmable DNA-binding domain; and (ii) a deaminase, wherein the deaminase is the adenosine deaminase described in the first aspect or the cytidine deaminase described in the second aspect.
[0061] In some embodiments, the deaminase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any one of SEQ ID NO:1-47, SEQ ID NO:48-63, 64-99, or 118-136 (e.g., SEQ ID NO:2-47, SEQ ID NO:48-63, 65-99, or 118-136).
[0062] In some embodiments, the deaminase is an adenosine deaminase, and the adenosine deaminase comprises:
[0063] (i) amino acids having 80% to 99.5% (e.g., 85% to 99.5%, 90% to 99.5%, or 95% to 99.5%) sequence identity with SEQ ID NO:1, and
[0064] (ii) Select one or more changes from the group consisting of D12X, H21X, A61X, V95X, S155X, S156X, C159X, E164X, E167X, and K169X relative to the position of SEQ ID NO:1, wherein X is any amino acid. Preferably, select one or more substitutions from the group consisting of D12H, H21D or H21Q, A61C, V95T, S155R, S156N or S156K, C159R, E164V or E164W, E167R or E167W, and K169N relative to the position of SEQ ID NO:1.
[0065] In some embodiments, the deaminase is a cytidine deaminase, and the cytidine deaminase comprises:
[0066] (i) amino acids having 80% to 99.5% (e.g., 85% to 99.5%, 90% to 99.5%, or 95% to 99.5%) sequence identity with SEQ ID NO:64, and
[0067] (ii) Selecting one or more changes from the group consisting of A17X, H34X, D36X, I46X, G64X, G68X, Y70X, L73X, V79X, F81X, M91X, H93X, G102X, A103X, A106X, D112X, A119X, N124X, R130X, S139X, A140X, C143X, R151X, and K153X relative to SEQ ID NO:64, wherein X is any amino acid, preferably selected relative to SEQ ID One or more substitutions from the group consisting of A17N, H34N, D36L, I46Q, G64C, G68T, Y70G or Y70H or Y70K, L73F or L73R, V79T, F81L, M91L, H93N, G102S, A103V, A106R, D112N, A119G, N124V or N124P, R130V or R130T, S139K, A140R, C143A or C143S, R151M, K153G or K153R of SEQ ID NO:64, wherein the amino acid position is relative to the number of SEQ ID NO:64.
[0068] In some embodiments, the polynucleotide may be codon-optimized for better expression in host cells. In a tenth aspect, this application provides an expression vector comprising a polynucleotide encoding the nucleobase editor described herein. In some embodiments, the vector is a viral vector selected from the group consisting of adeno-associated virus (AAV), retroviral vectors, adenovirus vectors, lentiviral vectors, Sendai virus vectors, and herpesvirus vectors.
[0069] In an eleventh aspect, this application provides a polynucleotide encoding the fusion protein of the present invention and an expression vector comprising said polynucleotide.
[0070] In some implementations, the polynucleotide may be codon-optimized for better expression in host cells.
[0071] Those skilled in the art will readily understand that the expression vector described in this invention may contain transcriptional regulatory elements such as promoters and enhancers.
[0072] In a twelfth aspect, this application provides cells comprising the expression vector described in this invention.
[0073] In some embodiments, the cell can be any suitable type of cell, such as a prokaryotic or eukaryotic cell, including, but not limited to, bacterial cells, fungal cells, plant cells, insect cells, or mammalian cells.
[0074] In a thirteenth aspect, this application provides a base editing method comprising contacting a polynucleotide sequence with a nucleobase editor, the nucleobase editor comprising: (i) a polynucleotide programmable DNA-binding domain; and (ii) a deaminase, wherein the deaminase is an adenosine deaminase as described in the first aspect or a cytidine deaminase as described in the second aspect, wherein the deaminase deaminates nucleosides in the polynucleotide, thereby editing the polynucleotide sequence.
[0075] In some embodiments, the deaminase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any one of SEQ ID NO:1-47, SEQ ID NO:48-63, 64-99, or 118-136 (e.g., SEQ ID NO:2-47, SEQ ID NO:48-63, 65-99, or 118-136).
[0076] In a fourteenth aspect, this application provides a method for correcting a genetic defect in a subject, the method comprising administering a nucleobase editor to the subject or administering a polynucleotide encoding a nucleobase editor, the nucleobase editor comprising: (i) a polynucleotide programmable DNA-binding domain; and (ii) a deaminase, wherein the deaminase is an adenosine deaminase as described in the first aspect or a cytidine deaminase as described in the second aspect, to deaminate a target nucleobase in a target nucleotide sequence of the subject, thereby correcting the genetic defect.
[0077] In some embodiments, the deaminase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any one of SEQ ID NO:1-47, SEQ ID NO:48-63, 64-99, or 118-136 (e.g., SEQ ID NO:2-47, SEQ ID NO:48-63, 65-99, or 118-136).
[0078] In some embodiments, the method includes delivering the nucleobase editor or a polynucleotide encoding the nucleobase editor, and one or more guide polynucleotides, to the cells of the subject. In some embodiments, the subject is a mammal, including non-human mammals or humans. In some embodiments, deamination of the target nucleobase causes the target nucleobase to be replaced by a wild-type nucleobase.
[0079] In a fifteenth aspect, the present invention provides pharmaceutical uses of the adenosine deaminase or cytidine deaminase, nucleobase editor or fusion protein or their encoding polynucleotides as described herein.
[0080] In some embodiments, this application provides the use of the adenosine deaminase or cytidine deaminase or their encoding polynucleotides described herein in the preparation of a medicament or kit for base editing or correction of genetic defects in a subject.
[0081] In some embodiments, this application provides the use of the fusion protein described herein or its encoded polynucleotide in the preparation of a medicament or kit for base editing or correction of genetic defects in a subject.
[0082] In some embodiments, this application provides the use of the nucleobase editor of the present invention or the polynucleotide encoding the nucleobase editor in the preparation of a medicament or kit for base editing or correction of genetic defects in a subject.
[0083] In a sixteenth aspect, this application provides a molecular complex comprising the base editor described herein and one or more guide RNA sequences, tracrRNA sequences, or target DNA sequences.
[0084] In a seventeenth aspect, this application provides a molecular complex comprising the fusion protein described herein and one or more guide RNA sequences, tracrRNA sequences, or target DNA sequences.
[0085] In an eighteenth aspect, this application provides a pharmaceutical composition comprising the base editor or fusion protein of the present invention, or an expression vector or molecular complex comprising a polynucleotide encoding the base editor or fusion protein, and a pharmaceutically acceptable excipient.
[0086] In a nineteenth aspect, this application provides a kit comprising: the base editor or fusion protein of the present invention, or an expression vector or molecular complex comprising a polynucleotide encoding the base editor or fusion protein, and instructions for use.
[0087] The pharmaceutical compositions or kits of the present invention can be used for base editing or correction of genetic defects in subjects.
[0088] In a twentieth aspect, this application provides a mutated Cas9 comprising one or two amino acid substitutions selected from D700R or E1243R.
[0089] In some embodiments, the mutated Cas contains D700R, and in other embodiments, the mutated Cas contains E1243R. In a preferred embodiment, the mutated Cas9 contains, or is composed of, the amino acid sequence of SEQ ID NO: 115 or 116.
[0090] In this invention, the terms "mutated Cas" and "Cas variant" are used interchangeably.
[0091] In a twentieth aspect, this application provides a composition comprising: (i) a first nucleotide sequence comprising a first open reading frame encoding a polypeptide comprising a deaminase and a polynucleotide programmable DNA binding domain; and (ii) a second nucleotide sequence comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI) domain; wherein the first nucleotide sequence differs from the second nucleotide sequence, and optionally, wherein the composition is encapsulated in lipid nanoparticles (LNPs).
[0092] In some embodiments, the composition further comprises (iii) a third nucleotide sequence comprising a gRNA sequence, wherein the first nucleotide sequence, the second nucleotide sequence, and the third nucleotide sequence are each a different nucleotide sequence.
[0093] In some embodiments, this application provides a lipid nanoparticle (LNP) encapsulating: (i) a first nucleotide sequence comprising a first open reading frame encoding a polypeptide containing a deaminase and a polynucleotide programmable DNA-binding domain; and (ii) a second nucleotide sequence comprising a second open reading frame encoding a uracil glycosylation inhibitor (UGI) domain; wherein the first nucleotide sequence is different from the second nucleotide sequence. In some embodiments, the LNP further comprises (iii) a third nucleotide sequence comprising a gRNA sequence, wherein the first, second, and third nucleotide sequences are each different nucleotide sequences.
[0094] In some embodiments, the deaminase is adenosine deaminase or cytidine deaminase, for example, but not limited to the adenosine deaminase or cytidine deaminase described in this application.
[0095] In some embodiments, the polynucleotide-programmable DNA-binding domain is selected from Cas9, Cas12, TnpB, IscB, meganuclease, zinc finger DNA-binding domains, and transcription activator-like effector (TALE) DNA-binding domains. In some embodiments, the polynucleotide-programmable DNA-binding domain is a Cas9 domain. In some embodiments, the polynucleotide-programmable DNA-binding domain is an RNA-guided nickase. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The methods and materials used in this invention are described herein; other suitable methods and materials known in the art may also be used. Materials, methods, and examples are illustrative only and are not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references mentioned herein are incorporated herein by reference in their entirety. In case of conflict, this specification (including definitions) shall prevail.
[0096] Other features and advantages of the invention will be apparent from the following detailed description, drawings, and claims. Attached Figure Description
[0097] Figure 1 shows the A to G editing efficiencies of different TadA variants in sgRNA library cell lines verified in Example 1.
[0098] Figure 2 shows the fold change in the percentage of mutations in surviving colonies before and after selecting E. coli cells transformed with the cytidine deaminase variant library in Example 7 (16 or 64 μg / ml chloramphenicol).
[0099] Figure 3 shows the percentage of mutations identified by next-generation sequencing (NGS) in surviving colonies on selective plates for the engineered CBE variant of the present invention in Example 7.
[0100] Figure 4 shows the C to T editing efficiency of variants T88.74-V10, T88.74-V14, T88.74-V16 and T88.74-V17 relative to T88.74-V0 detected in Example 8.
[0101] Figure 5 shows the fold change in the percentage of mutations in surviving colonies before and after selecting E. coli cells transformed with the cytidine deaminase variant library in Example 8 (64 or 128 μg / ml chloramphenicol).
[0102] Figure 6 shows the C to T editing efficiency of variants T88.74-V21 and T88.74-V22 relative to variant T88.74-V14, as detected in Example 8.
[0103] Figure 7 shows the C to T editing efficiency of variants T88.74-V27, T88.74-V29, T88.74-V32 and T88.74-V33 relative to variant T88.74-V17, as detected in Example 8.
[0104] Figure 8 shows the C to T editing efficiency of variants T88.74-V39, T88.74-V40, and T88.74-V41 relative to variant T88.74-V27.
[0105] Figure 9 shows the C-to-T editing efficiency of variants T88.74-V44 and T88.74-V51 in HEK293T sgRNA library cells. Detailed Implementation
[0106] Base Editor
[0107] This document discloses nucleobase editors, such as adenosine base editors (ABEs) or cytidine base editors (CBEs), for editing, modifying, or altering target nucleotide sequences of polynucleotides. The document describes a nucleobase editor comprising a programmable nucleotide-binding domain (e.g., a polynucleotide programmable nucleotide-binding domain (e.g., Cas9), a zinc finger protein DNA-binding domain, or a TALE DNA-binding domain) and at least one nucleobase editing domain (e.g., an adenosine deaminase or a cytidine deaminase). When the polynucleotide programmable nucleotide-binding domain (e.g., Cas9 or Cas12) is present in the cell and binds to a bound guide polynucleotide (e.g., gRNA), it can specifically bind to the target polynucleotide sequence (i.e., through complementary base pairing between the bases of the bound guide nucleic acid and the bases of the target polynucleotide sequence) and thereby position the base editor to the target nucleic acid sequence to be edited.
[0108] In some implementations, base editing activity is assessed by editing efficiency. Base editing can be determined by any suitable method, such as Sanger sequencing or next-generation sequencing. In some implementations, base editing efficiency is measured as the percentage of total sequencing reads resulting from nucleobase conversions achieved by the base editor; for example, for adenosine deaminases, the percentage of total sequencing reads measured from target AT base pairs converted to GC base pairs; for cytidine deaminases, the percentage of total sequencing reads measured from target CG base pairs converted to TA base pairs. In some implementations, when base editing is performed in a cell population, base editing efficiency is measured as the percentage of cells with nucleobase conversions achieved by the base editor.
[0109] The term "base editor system" refers to the components required to edit the nucleosides of a target nucleotide sequence. In various embodiments, a base editor system comprises: (1) a polynucleotide-programmable nucleotide-binding domain (e.g., Cas9); (2) a deaminase domain for deaminating the nucleosides (e.g., adenosine deaminase and / or cytidine deaminase; see PCT / US2019 / 044935, PCT / US2020 / 016288, the entire contents of each of which are incorporated herein by reference); and (3) one or more guide polynucleotides (e.g., guide RNA). In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable DNA-binding domain. In some embodiments, the base editor is an adenosine base editor (ABE). In other embodiments, the base editor is a cytidine base editor (CBE).
[0110] In some embodiments, the base editor system may contain more than one base editing component. For example, the base editor system may contain more than one deaminase. In some embodiments, the base editor system may contain one or more adenosine deaminases or one or more cytidine deaminases. In some embodiments, a single guide polynucleotide may be used to target different deaminases to the target nucleic acid sequence. In some embodiments, a pair of guide polynucleotides may be used to target different deaminases to the target nucleic acid sequence.
[0111] The deaminase domain and the polynucleotide-programmable nucleotide binding component of the base editor system can associate with each other covalently or nonvalently, or with any combination of their association and interactions. For example, in some embodiments, the deaminase domain can target a target nucleotide sequence via the polynucleotide-programmable nucleotide binding domain. In some embodiments, the polynucleotide-programmable nucleotide binding domain can be fused to or linked to the deaminase domain. In some embodiments, the polynucleotide-programmable nucleotide binding domain can target the deaminase domain to a target nucleotide sequence through nonvalent interaction or association with the deaminase domain. For example, in some embodiments, the deaminase domain may include additional heterologous portions or domains that are capable of interacting, associating, or forming complexes with additional heterologous portions or domains that are part of the polynucleotide-programmable nucleotide binding domain.
[0112] In some embodiments, the additional heterologous motif may be capable of binding, interacting, associating, or forming a complex with the peptide. In some embodiments, the additional heterologous motif may be capable of binding, interacting, associating, or forming a complex with a polynucleotide. In some embodiments, the additional heterologous motif may be capable of binding to a guide polynucleotide. In some embodiments, the additional heterologous motif may be capable of binding to a peptide linker. In some embodiments, the additional heterologous motif may be capable of binding to a polynucleotide linker. The additional heterologous motif may be a protein domain. In some embodiments, the additional heterologous motif may be a K homology (KH) domain, an MS2 capsid protein domain, a PP7 capsid protein domain, an SfMuCom capsid protein domain, a sterile α motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif.
[0113] The base editor system may also include guide polynucleotide components. It should be understood that components of the base editor system can associate with each other through covalent bonds, non-covalent interactions, or any combination of association and interaction thereof. In some embodiments, the deaminase domain can target a target nucleotide sequence via the guide polynucleotide. For example, in some embodiments, the deaminase domain may include an additional heterologous portion or domain (e.g., a polynucleotide-binding domain, such as an RNA or DNA-binding protein) capable of interacting with, associating with, or forming a complex with a portion or segment (e.g., a polynucleotide motif) of the guide polynucleotide. In some embodiments, the additional heterologous portion or domain (e.g., a polynucleotide-binding domain, such as an RNA or DNA-binding protein) may be fused to or linked to the deaminase domain. In some embodiments, the additional heterologous portion may be capable of binding to, interacting with, associating with, or forming a complex with a polypeptide. In some embodiments, the additional heterologous portion may be capable of binding to, interacting with, associating with, or forming a complex with a polynucleotide. In some embodiments, the additional heterologous portion may be capable of binding to the guide polynucleotide. In some embodiments, the additional heterologous portion may be capable of binding to a polypeptide linker. In some embodiments, the additional heterologous portion may be capable of binding to a polynucleotide linker. The additional heterologous portion may be a protein domain. In some embodiments, the additional heterologous portion may be a K homology (KH) domain, an MS2 capsid protein domain, a PP7 capsid protein domain, an SfMuCom capsid protein domain, a sterile α motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif.
[0114] Guide polynucleotides
[0115] In some implementations, the guide polynucleotide is a guide RNA. The RNA / Cas complex helps “guide” the Cas protein to the target DNA. Cas9 / crRNA / tracrRNA cleaves the linear or circular dsDNA target complementary to the spacer via endonuclease. First, the target strand not complementary to the crRNA is cleaved via endonuclease, followed by 3'-5' trimming via exonuclease. In nature, DNA binding and cleavage typically require both proteins and two RNAs. However, a single guide RNA (“sgRNA” or simply “gRNA”) can be engineered to integrate aspects of both crRNA and tracrRNA into a single RNA species. See, for example, Jinek M. et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes short motifs (PAMs or protospacer adjacent motifs) in CRISPR repeat sequences to help distinguish itself from non-self.
[0116] The sequence and structure of the Cas9 nuclease are well known to those skilled in the art (see, for example, "Complete genome sequence of an Ml strain of Streptococcus pyogenes." Ferretti, JJ et al., Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E. et al., Nature 471:602-607 (2011); and "Programmable dual-RNA-guided DNAendonuclease in adaptive bacterial immunity." Jinek M. et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference). Cas9 orthologs have been described in a variety of species, including but not limited to Streptococcus pyogenes and Streptococcus thermophilus. Based on this application, those skilled in the art will recognize other suitable Cas9 nucleases and sequences, including Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737 (the entire contents of which are incorporated herein by reference). In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is a nicking enzyme. In some embodiments, the guide polynucleotide is at least one single guide RNA (“sgRNA” or “gNRA”). In some embodiments, the guide polynucleotide is at least one tracrRNA. In some embodiments, the guide polynucleotide does not require a protospacer adjacent motif (PAM) sequence to guide the polynucleotide programmable DNA-binding domain (e.g., Cas9 or Cpfl) to the target nucleotide sequence. The multinucleotide programmable nucleotide-binding domains (e.g., CRISPR-derived domains) of the base editor disclosed herein can recognize target polynucleotide sequences by associating with guide polynucleotides. Guide polynucleotides (e.g., gRNAs) are typically single-stranded and can be programmed to bind specifically (i.e., via complementary base pairing) to the target sequence of the polynucleotide, thereby guiding the base editor bound to the guide nucleic acid to the target sequence.The guide polynucleotide can be DNA. The guide polynucleotide can be RNA. In some embodiments, the guide polynucleotide comprises a natural nucleotide (e.g., adenosine). In some embodiments, the guide polynucleotide comprises a non-natural (or not naturally occurring) nucleotide (e.g., peptide nucleic acid or nucleotide analogue). In some embodiments, the length of the target region of the guide nucleic acid sequence can be at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. The length of the target region of the guide nucleic acid can be 10-30 nucleotides, or 15-25 nucleotides, or 15-20 nucleotides.
[0117] Programmable nucleotide binding domain
[0118] The programmable nucleotide-binding domain of the base editor may itself contain one or more domains. For example, a polynucleotide programmable nucleotide-binding domain may contain one or more nuclease domains. In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide-binding domain may contain an endonuclease or an exonuclease. The term "exonuclease" as used herein refers to a protein or polypeptide capable of digesting nucleic acids (e.g., RNA or DNA) from their free ends, and the term "endonuclease" refers to a protein or polypeptide capable of catalyzing (e.g., cleaving) internal regions of nucleic acids (e.g., DNA or RNA). In some embodiments, an endonuclease can cleave single strands in a double-stranded nucleic acid.
[0119] Any DNA destabilizing molecule may be used in any combination of the compositions described herein, including but not limited to Cas9 or Cas12 nickases, Cas9 or Cas12 proteins operatively linked to a single guide RNA (sgRNA) (e.g., dCas), any RNA programmable system, zinc finger nuclease nickases (ZFN nickases), TALEN nickases, and / or one or more nucleotides (e.g., one or more peptide nucleic acids (PNAs), locked nucleic acids (LNAs), and / or bridging nucleic acids (BNAs)). In some embodiments, the base editing composition comprises more than one DNA destabilizing molecule, such as one or more proteins (e.g., nickases, etc.) and / or one or more nucleotides. In some embodiments, the composition comprises a ZFN nickase and one or more additional protein and / or nucleotide DNA destabilizing molecules (e.g., one or more nucleotides as described herein). In some aspects, the base editing composition does not contain a Cas9 protein but may contain other Cas proteins (e.g., non-Cas9 RNA programmable systems). In some embodiments, the DNA destabilizing molecule includes a zinc finger nuclease (ZFN) nickase.
[0120] In some implementations, the nuclease is a zinc finger nuclease (ZFN) or a TALE DNA-binding domain-nuclease fusion (TALEN). ZFNs and TALENs comprise a DNA-binding domain (zinc finger protein or TALE DNA-binding domain) and a cleavage domain or cleavage hemidomain, said DNA-binding domain being engineered to bind to a target site in a selected gene.
[0121] At least one zinc finger protein (ZFP) DNA-binding domain of the base editing composition may be operatively linked to one or more other components of the base editing composition, such as to one or more DNA destabilizing molecules (e.g., to Cas9 nickase, dCas9, etc.) and / or to at least one adenine or cytosine deaminase. In some embodiments, at least one ZFP DNA-binding domain is operatively linked to an adenine or cytosine deaminase. In other embodiments, the base editing composition comprises first and second ZFP DNA-binding domains, wherein the first ZFP DNA-binding domain is operatively linked to a Cas9 nickase. The ZFP DNA-binding domain may comprise 3, 4, 5, 6, or more fingers and may bind to a target site on either side (5' or 3') of the targeted base to be edited. In some embodiments, the ZFP-binding target site is located at 1 to 100 nucleotides (or any number in between) on either side of the targeted base. In other embodiments, the target site for ZFP binding is located at 1 to 50 (or any number in between) nucleotides on either side of the targeted base.
[0122] In some embodiments, the endonuclease can cleave both strands of a double-stranded nucleic acid molecule. In some embodiments, the polynucleotide-programmable nucleotide-binding domain can be a deoxyribonuclease. In some embodiments, the polynucleotide-programmable nucleotide-binding domain can be a ribonuclease.
[0123] In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide binding domain can cleave zero, one, or both strands of the target polynucleotide. In some embodiments, the polynucleotide programmable nucleotide binding domain may include a nicking enzyme domain. The term "nicking enzyme" as used herein refers to a polynucleotide programmable nucleotide binding domain that includes a nuclease domain capable of cleaving only one strand of a double-stranded nucleic acid molecule (e.g., DNA). In some embodiments, the nicking enzyme can be derived from the fully catalytically active (e.g., native) form of the polynucleotide programmable nucleotide binding domain by introducing one or more mutations into the active polynucleotide programmable nucleotide binding domain. For example, when the polynucleotide programmable nucleotide binding domain includes a nicking enzyme domain derived from Cas9, the Cas9-derived nicking enzyme domain may contain a D10A mutation and a histidine residue at position 840. In such an embodiment, residue H840 retains catalytic activity, thereby enabling the cleavage of a single strand of the nucleic acid duplex. In another embodiment, the Cas9-derived nicking enzyme domain may contain an H840A mutation, while the amino acid residue at position 10 remains D. In some implementations, the nicking enzyme can be derived from the fully catalytically active (e.g., native) form of the polynucleotide programmable nucleotide binding domain by removing all or part of the nuclease domains that are not required for nicking enzyme activity. For example, when the polynucleotide programmable nucleotide binding domain includes a nicking enzyme domain derived from Cas9, the Cas9-derived nicking enzyme domain may contain all or part of the RuvC domain or the HNH domain.
[0124] Therefore, a base editor containing a polynucleotide programmable nucleotide-binding domain (including a nicking enzyme domain) can generate single-strand DNA breaks (nicks) at specific polynucleotide target sequences (e.g., determined by the complementary sequence of the bound guide nucleic acid). In some embodiments, the strand of the double-stranded target polynucleotide sequence cleaved by a base editor containing a nicking enzyme domain (e.g., a Cas9-derived nicking enzyme domain) is the unedited strand (i.e., the cleaved strand is opposite the strand containing the base to be edited). In other embodiments, a base editor containing a nicking enzyme domain (e.g., a Cas9-derived nicking enzyme domain) can cleave the strand of the DNA molecule targeted for editing. In such embodiments, untargeted strands are not cleaved.
[0125] This document also provides a base editor comprising a catalytically inactivated (i.e., incapable of cleaving the target polynucleotide sequence) polynucleotide programmable nucleotide-binding domain. In this document, the terms “catalytic inactivation” and “nuclease inactivation” are used interchangeably to refer to a polynucleotide programmable nucleotide-binding domain having one or more mutations and / or deletions, resulting in its inability to cleave the nucleic acid chain while retaining its ability and specificity to bind to the target polynucleotide. In some embodiments, the catalytically inactivated polynucleotide programmable nucleotide-binding domain base editor may lack nuclease activity due to specific point mutations in one or more nuclease domains. For example, in the case of a base editor containing a Cas9 domain, Cas9 may contain both the D10A and H840A mutations. Such mutations inactivate both nuclease domains, resulting in a loss of nuclease activity. In other embodiments, the catalytically inactivated polynucleotide programmable nucleotide-binding domain may contain one or more deletions of all or part of the catalytic domain (e.g., the RuvCl and / or HNH domains). In a further embodiment, the catalytically inactivated polynucleotide programmable nucleotide-binding domain includes point mutations (e.g., D10A or H840A) and the deletion of all or part of the nuclease domain.
[0126] Some aspects of this application provide fusion proteins comprising domains that act as polynucleotide-programmable DNA-binding proteins, which can be used to guide proteins (such as base editors) to specific nucleic acid (e.g., DNA or RNA) sequences. In particular embodiments, the fusion protein comprises a nucleic acid-programmable DNA-binding protein domain and one or more deaminase domains. Non-limiting examples of polynucleotide-programmable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), Casl2a / Cpfl, Casl2b / C2cl, Casl2c / C2c3, Casl2d / CasY, Casl2e / CasX, Casl2g, Casl2h, and Casl2i. Unrestricted examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csnl or Csxl2), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, and Ca. sl2h, Casl2i, Csyl, Csy2, Csy3, Csy4, Csel, Cse2, Cse3, Cse4, Cse5e, Cscl, Csc2, Csa5, Csnl, Csn2, Csml, Csm2, Csm3, Csm 4. Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, Csxl0, Csxl6, CsaX, Csx3, Csxl, CsxlS, Csxl 1. Csfl, Csf2, CsO, Csf4, Csdl, Csd2, Cstl, Cst2, Cshl, Csh2, Csal, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V Cas effector proteins, type VI Cas effector proteins, CARF, DinG, their homologs, or their modified or engineered forms. Other nucleic acid-programmable DNA-binding proteins are also within the scope of this application, although they may not be specifically listed herein.For example, see Makarova et al., "Classification and Nomenclature of CRISPR-Cas Systems: Where from Here" CRISPR J. 2018 Oct; 1:325-336. doi:10.1089 / crispr.2018.0033; Yan et al., "Functionally diverse type V CRISPR-Cas systems" Science. 2019 Jan 4; 363(6422):88-91. doi:10.l126 / science.aav7271. The full content of each reference is incorporated into this paper by citation.
[0127] In some embodiments, the polynucleotide programmable DNA-binding domain is a Cas9 domain, preferably a mutated Cas9 containing one or two amino acid substitutions selected from D700R or E1243R, more preferably a Cas9 containing D700R or a Cas9 containing E1243R. In a preferred embodiment, the mutated Cas9 contains, or is composed of, the amino acid sequence of SEQ ID NO: 115 or 116.
[0128] In some embodiments, this application provides fusion proteins comprising type V CRISPR / Cas effector proteins. Type V CRISPR / Cas effector proteins are subtypes of class 2 CRISPR / Cas effector proteins. Examples of type V CRISPR / Cas systems and their effector proteins (e.g., Cas12 family proteins, such as Cas12a) can be found, for example, Shmakov et al., Nat Rev Microbial. 2017 March; 15(3):169-182: "Diversity and evolution of class 2 CRISPR-Cas systems." Examples include, but are not limited to, the Cas12 family (Cas12a, Cas12b, Cas12c), C2c4, C2c8, C2c5, C2c10, and C2c9; and CasX (Cas12e) and CasY (Cas12d). See also, for example, Koonin et al., Curr Opin Microbial. 2017 June; 37:67-78: "Diversity, classification and evolution of CRISPR-Cas systems." In some implementations, the CBE disclosed herein comprises a type V CRISPR / Cas effector protein.
[0129] In some embodiments, this application provides a TALE base editor that may include a TALE domain, a deaminase domain, and / or a cofactor protein (e.g., FokI endonuclease) domain. The TALE base editor comprises a fusion protein having the universal structures NH2-[TALE]-[deaminase domain]-COOH, NH2-[deaminase domain]-[TALE]-COOH, NH2-[TALE]-[deaminase domain]-[cofactor protein]-COOH, NH2-[cofactor protein]-[deaminase domain]-[TALE]-COOH, NH2-[cofactor protein]-[TALE]-[deaminase]-COOH, or NH2-[deaminase domain]-[TALE]-[cofactor protein]-COOH; wherein each instance of "]-[" includes an optional linker, such as a peptide linker.
[0130] In some embodiments, the disclosed method involves transducing (e.g., by transfection) cells with multiple complexes, each complex containing a fusion protein and a cofactor protein, the fusion protein containing a TAL effector domain and a deaminase domain, wherein each cofactor protein localizes the fusion protein to a different target sequence. See Yang L. et al., Engineering and optimizing deaminase fusions for genome editing, Nature Comms., 2016. In a particular embodiment, the method disclosed herein involves a TAL effector domain that binds to the target site not via Watson-Crick hybridization but by binding to the major groove of the DNA double helix. In some embodiments, the method involves transfecting nucleic acid constructs (e.g., plasmids), each nucleic acid construct (or together) encoding components of multiple complexes of a TALE base editor and a cofactor protein, the TALE base editor containing a TALE domain and a deaminase domain. In some embodiments, the disclosed fusion protein contains a cofactor protein domain, i.e., this domain is integrated into the fusion protein construct. In other implementations, the TALE base editor includes a TALE domain and a deaminase domain, and the cofactor protein is introduced into the cell separately from the base editor.
[0131] In some embodiments of the disclosed method, the construct encoding the TALE base editor is transfected into cells separately from the construct encoding the cofactor protein. In some embodiments, these components are encoded on a single construct and transfected together. In a particular embodiment, these single constructs encoding the TALE base editor and the cofactor protein can be iteratively transfected into cells, each iteration associated with a subset of the target sequences. In a particular embodiment, these single constructs can be transfected into cells within a few days. In other embodiments, they can be transfected into cells within several weeks.
[0132] A to G editors
[0133] In some embodiments, the base editor described herein may include a deaminase domain comprising adenosine deaminase. This adenosine deaminase domain of the base editor facilitates the editing of adenine (A) nucleotides to guanine (G) nucleotides by deaminating adenosine or adenine (A) to form inosine (I). Inosine (I) exhibits similar base-pairing properties to G; that is, inosine can pair with cytosine, thereby introducing guanine at the adenine site deamination by adenosine deaminase during subsequent transcription. Adenosine deaminase is capable of deaminating (i.e., removing the amino group) adenine from deoxyadenosine residues in deoxyribonucleic acid (DNA).
[0134] In some embodiments, the nucleobase editors provided herein can be prepared by fusing one or more protein domains together to produce fusion proteins. In some embodiments, the fusion proteins provided herein include one or more features that improve the base editing activity (e.g., efficiency, selectivity, and specificity) of the fusion protein. For example, the fusion proteins provided herein may include a Cas9 domain with reduced nuclease activity. In some embodiments, the fusion proteins provided herein may have a Cas9 domain without nuclease activity (dCas9), or a Cas9 domain that cleaves one strand of a double-stranded DNA molecule (referred to as Cas9 cleavage enzyme (nCas9)). Without being bound by any particular theory, the presence of a catalytic residue (e.g., H840) maintains the activity of Cas9 to cleave the unedited (e.g., undeamination) strand containing a T opposite to the targeted A. Mutations in the catalytic residues of Cas9 (e.g., D10 to A10) prevent the cleavage of the edited strand containing the targeted A residue. These Cas9 variants can generate single-strand DNA breaks (gaps) at specific locations based on gRNA-defined target sequences, leading to the repair of the unedited strand and ultimately resulting in T-to-C changes on the unedited strand. In some embodiments, the A-to-G base editor also includes an inhibitor of inosine base excision repair, such as a uracil glycosylase inhibitor (UGI) domain or a catalytically inactive inosine-specific nuclease. Without being bound by any particular theory, the UGI domain or the catalytically inactive inosine-specific nuclease can inhibit or prevent base excision repair of deamination of adenosine residues (e.g., inosine), which can improve the activity or efficiency of the base editor.
[0135] A base editor containing adenosine deaminase can act on any polynucleotide, including DNA, RNA, and DNA-RNA hybrids. In some embodiments, the base editor containing adenosine deaminase can deaminate target A of a polynucleotide (including RNA). For example, the base editor may contain an adenosine deaminase domain capable of deaminating target A of RNA polynucleotides and / or DNA-RNA hybrid polynucleotides. In one embodiment, the adenosine deaminase integrated into the base editor comprises all or part of an adenosine deaminase (ADAR, e.g., ADAR1 or ADAR2) that acts on RNA. In another embodiment, the adenosine deaminase integrated into the base editor comprises all or part of an adenosine deaminase (ADAT) that acts on tRNA. A base editor containing an adenosine deaminase domain may also be capable of deaminating the A nucleobase of a DNA polynucleotide. In one embodiment, the adenosine deaminase domain of the base editor comprises all or part of ADAT, which contains one or more mutations that allow ADAT to deaminate target A in DNA.
[0136] Adenosine deaminases can be derived from any suitable organism (e.g., *Escherichia coli*). In some embodiments, the adenosine deaminase is a naturally occurring adenosine deaminase comprising one or more mutations corresponding to the amino acid site mutations provided herein (e.g., mutations occurring in any of SEQ ID NO:1-47 (e.g., SEQ ID NO:2-47)). The corresponding residues in any homologous protein can be identified, for example, by sequence alignment and determination of homologous residues. Mutations in any naturally occurring adenosine deaminase corresponding to any of the mutations described herein can be generated accordingly.
[0137] adenosine deaminase
[0138] Some aspects of this application provide adenosine deaminases. In some embodiments, the adenosine deaminases provided herein are capable of deaminating adenine. In some embodiments, the adenosine deaminases provided herein are capable of deaminating adenine from deoxyadenosine residues of DNA. As mentioned herein, the term "adenosine deaminase" refers to a deaminase capable of deaminating adenine from deoxyadenosine residues of DNA, and in some cases, said adenosine deaminase can also deaminate cytosine from deoxycytidine residues of DNA. In some embodiments, the adenosine deaminases provided herein are capable of deaminating cytosine. In some embodiments, the adenosine deaminases provided herein are capable of deaminating cytosine from deoxycytidine residues of DNA, but with lower efficiency. Adenosine deaminases can be derived from any suitable organism (e.g., *Escherichia coli*). In some embodiments, the adenine deaminase is a naturally occurring adenosine deaminase containing one or more mutations corresponding to the amino acid site mutations provided herein. Those skilled in the art will be able to identify any homologous protein and the corresponding residues in its respective encoding nucleic acid using methods well-known in the art (e.g., by sequence alignment and determination of homologous residues). Therefore, those skilled in the art will be able to generate mutations in any naturally occurring adenosine deaminase corresponding to any of the mutations described herein. In some embodiments, the adenosine deaminase is derived from prokaryotes. In some embodiments, the adenosine deaminase is derived from bacteria. In some embodiments, the adenosine deaminase is derived from *Escherichia coli*, *Staphylococcus aureus*, *Salmonella typhi*, *Shewanella putrefaciens*, *Haemophilus influenzae*, *Strychnos nucleus*, or *Bacillus subtilis*. In some embodiments, the adenosine deaminase is derived from *Escherichia coli*.
[0139] In some embodiments, the adenosine deaminase is a TadA deaminase. In some embodiments, the TadA deaminase is a TadA variant.
[0140] In some embodiments, the adenosine deaminase comprises any amino acid sequence shown in any of SEQ ID NO:1-47 (e.g., SEQ ID NO:2-47) or at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any adenosine deaminase provided herein. It should be understood that the adenosine deaminase provided herein may contain one or more mutations (e.g., any of the mutations provided herein). This application provides any deaminase domain having a certain percentage of identity plus any mutation described herein or a combination thereof. In some embodiments, the adenosine deaminase comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the amino acid sequences shown in SEQ ID NO:1-47 or any of the adenosine deaminases provided herein. In some embodiments, the adenosine deaminase comprises an amino acid sequence having at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, or at least 170 identical consecutive amino acid residues compared to any amino acid sequence shown in SEQ ID NO:1-47 or any adenosine deaminase provided herein.
[0141] In some embodiments, the deaminases provided herein are capable of deaminating adenine. In some embodiments, the adenosine deaminases provided herein are capable of deaminating adenine from deoxyadenosine residues in DNA.
[0142] In some embodiments, the deaminases provided herein are capable of deaminating cytosine or 5-methylcytosine to uracil or thymine. In some embodiments, the deaminases provided herein are capable of deaminating cytosine in DNA.
[0143] In some embodiments, the adenosine deaminase comprises one or more amino acid substitutions selected from the group consisting of D12H, H21D or H21Q, A61C, V95T, S156K, C159R, E164V or E164W, and E167W, corresponding to the numbers in SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having an amino acid sequence having any of the amino acid sequences in SEQ ID NO:1-47 (e.g., SEQ ID NO:2-47). In one embodiment, the adenosine deaminase consists of an amino acid sequence having any of the amino acid sequences in SEQ ID NO:1-47 (e.g., SEQ ID NO:2-47).
[0144] In some embodiments, the adenosine deaminase comprises a V95T amino acid substitution relative to the number in SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO:2.
[0145] In some embodiments, the adenosine deaminase comprises the V95T and A61C amino acid substitutions relative to SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO:3.
[0146] In some embodiments, the adenosine deaminase comprises amino acid substitutions for V95T and C159R relative to SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO:4.
[0147] In some embodiments, the adenosine deaminase comprises the V95T and D12H amino acid substitutions relative to the number in SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO:5.
[0148] In some embodiments, the adenosine deaminase comprises amino acid substitutions of V95T and E164V relative to SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO:6.
[0149] In some embodiments, the adenosine deaminase comprises the amino acid substitutions V95T and E164W relative to SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO:7.
[0150] In some embodiments, the adenosine deaminase comprises amino acid substitutions for V95T and E167R relative to SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO:8.
[0151] In some embodiments, the adenosine deaminase comprises amino acid substitutions of V95T and E167W relative to SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO:9.
[0152] In some embodiments, the adenosine deaminase comprises amino acid substitutions of V95T, E164V, and E167W relative to SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO:10.
[0153] In some embodiments, the adenosine deaminase comprises the amino acid substitutions V95T and H21D relative to SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO:11.
[0154] In some embodiments, the adenosine deaminase comprises amino acid substitutions selected from V95T, C159R, E164V, and E167W relative to the number in SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO:12.
[0155] In some embodiments, the adenosine deaminase comprises amino acid substitutions corresponding to A61C, E164V, and E167W relative to SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO:13.
[0156] In some embodiments, the adenosine deaminase comprises amino acid substitutions of C159R, E164V, and E167W relative to SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO:14.
[0157] In some embodiments, the adenosine deaminase comprises amino acid substitutions of C159R and E167W relative to SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO:15.
[0158] In some embodiments, the adenosine deaminase comprises H21Q and A61C amino acid substitutions relative to SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO:16.
[0159] In some embodiments, the adenosine deaminase comprises amino acid substitutions of S155R and K169N relative to SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO:17.
[0160] In some embodiments, the adenosine deaminase comprises one or more amino acid substitutions relative to the group consisting of S156N and E164V of SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO:18.
[0161] In some embodiments, the adenosine deaminase comprises one or more amino acid substitutions relative to the group consisting of V95T, S156K, C159R, E164V, and E167W of SEQ ID NO:1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO:19.
[0162] It should be recognized that any mutations provided herein can be introduced into other adenosine deaminases, such as Staphylococcus aureus TadA (saTadA), or other adenosine deaminases (e.g., bacterial adenosine deaminases). Those skilled in the art will understand how to identify amino acid residues homologous to the mutated residues in SEQ ID NO:1-47 from other adenosine deaminases. Therefore, any mutation in SEQ ID NO:1-47 can be made in other adenosine deaminases. It should also be understood that any mutations provided herein can be made alone or in any combination in another adenosine deaminase.
[0163] Table A. Amino acid sequence of adenosine deaminase
[0164] Cytidine deaminase and C to T(U) editing
[0165] "Cytidine deaminase" refers to a polypeptide or fragment thereof capable of catalyzing a deamination reaction that converts an amino group to a carbonyl group. In one embodiment, the cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine. The cytidine deaminases provided herein (e.g., engineered cytidine deaminases) can be derived from any organism, such as bacteria.
[0166] In some implementations, the cytidine deaminase of the base editor may include all or part of the apolipoprotein B mRNA editing complex (APOBEC) family of deaminases or variants thereof. APOBEC is an evolutionarily conserved family of cytidine deaminases. Members of this family are C-to-U editing enzymes. In some embodiments, the cytidine deaminase includes, but is not limited to, members of the APOBEC family, including, but not limited to, APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D (now called "APOBEC3E"), APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, activation-induced (cytidine) deaminase (AID), hAPOBEC1 (derived from Homo sapiens), rAPOBEC1 (derived from Rattus norvegicus), ppAPOBEC1 (derived from Pongo pygmaeus), AmAPOBEC1 (BEM3.31) (derived from Alligator mississippiensis), and ocAPOBEC1 (derived from Orangutan). cuniculus), SsAPOBEC2 (BEM3.39) (derived from Eurasian wild boar (Sus scrofa)), hAPOBEC3A (derived from Homo sapiens), maAPOBEC1 (derived from golden hamster (Mesocricetus auratus)), mdAPOBEC1 (derived from grey opossum (Monodelphis domestica)); cytidine deaminase 1 (CDA1), hA3A (i.e., APOBEC3A derived from Homo sapiens), RrA3F (BEM3.14) (which is APOBEC3F derived from Sichuan golden monkey (Rhinopithecus roxellana)); PmCDA1 (derived from sea lamprey (Petromyzon)). marinus); AID (activation-induced cytidine deaminase; AICDA), which is derived from mammals (e.g., humans, pigs, cattle, horses, monkeys, etc.); hAID (derived from Homo sapiens); and FENRY.
[0167] The terms "deaminase" or "deaminase domain," as used herein, refer to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes the hydrolytic deamination reaction of cytidine to uridine or deoxycytidine to deoxyuridine. In some embodiments, the deaminase or deaminase domain is a cytosine deaminase that catalyzes the hydrolytic deamination reaction of cytosine to uracil, thereby enabling editing from C:G to T:A (i.e., C to T editing).
[0168] In some embodiments, the cytidine deaminase is a de novo-designed TadA variant with cytidine deaminase activity. In some embodiments, the cytidine deaminase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any one of SEQ ID NO:48-63, 65-99, or 118-136.
[0169] In some embodiments, the cytidine deaminase is generated through directed evolution based on SEQ ID NO:64. In some embodiments, the cytidine deaminase comprises one or more substitutions selected from the group consisting of A17N, H34N, D36L, I46Q, G64C, G68T, Y70G or Y70H or Y70K, L73F or L73R, V79T, F81L, M91L, H93N, G102S, A103V, A106R, D112N, A119G, N124V or N124P, R130V or R130T, S139K, A140R, C143A or C143S, R151M, K153G or K153R, wherein the amino acid positions are relative to the numbering in SEQ ID NO:64. In some embodiments, the cytidine deaminase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any one of SEQ ID NO:64-99 or 118-136.
[0170] In some embodiments, the cytidine deaminase comprises the amino acid substitutions Y70H and M91L relative to SEQ ID NO:64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO:74.
[0171] In some embodiments, the cytidine deaminase comprises amino acid substitutions of Y70H, M91L, and F81L relative to SEQ ID NO:64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO:78.
[0172] In some embodiments, the cytidine deaminase comprises amino acid substitutions of Y70H, M91L, F81L, C143A, G64C, and L73F relative to SEQ ID NO:64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO:80.
[0173] In some embodiments, the cytidine deaminase comprises amino acid substitutions of Y70H, M91L, F81L, G64C, and L73F relative to SEQ ID NO:64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO:81.
[0174] In some embodiments, the cytidine deaminase comprises amino acid substitutions of Y70H, M91L, F81L, and N124V relative to SEQ ID NO:64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO:85.
[0175] In some embodiments, the cytidine deaminase comprises amino acid substitutions of Y70H, M91L, F81L, G64C, L73F, and N124V relative to SEQ ID NO:64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO:91.
[0176] In some embodiments, the cytidine deaminase comprises amino acid substitutions of Y70H, M91L, F81L, G64C, L73F, R130V, A140R, D112N, and N124V relative to SEQ ID NO:64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO:97.
[0177] In some embodiments, the cytidine deaminase comprises an R130T amino acid substitution relative to SEQ ID NO:64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO:118.
[0178] In some embodiments, the cytidine deaminase comprises an H34N amino acid substitution relative to SEQ ID NO:64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO:119.
[0179] In some embodiments, the cytidine deaminase comprises amino acid substitutions of Y70H, M91L, F81L, G64C, L73F, N124V, R130T, and H34N relative to SEQ ID NO:64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO:123.
[0180] In some embodiments, the cytidine deaminase comprises amino acid substitutions corresponding to Y70H, M91L, F81L, G64C, L73F, N124V, R130T, H34N, and A106R relative to SEQ ID NO:64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO:126.
[0181] In some embodiments, the cytidine deaminase comprises amino acid substitutions corresponding to Y70H, M91L, F81L, G64C, L73F, and A106R relative to SEQ ID NO:64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO:127.
[0182] In some embodiments, the cytidine deaminase comprises amino acid substitutions corresponding to Y70H, M91L, F81L, G64C, L73F, N124V, R130T, H34N, A106R, and C143S relative to SEQ ID NO:64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO:131.
[0183] In some embodiments, the cytidine deaminase comprises amino acid substitutions corresponding to Y70H, M91L, F81L, G64C, L73F, N124V, R130T, H34N, A106R, and I46Q relative to SEQ ID NO:64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO:132.
[0184] In some embodiments, the cytidine deaminase comprises amino acid substitutions corresponding to Y70H, M91L, F81L, G64C, L73F, N124V, R130T, H34N, A106R, and V79T relative to SEQ ID NO:64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO:133.
[0185] In some embodiments, the cytidine deaminase comprises the amino acid sequence of any one of SEQ ID NO:48-63, SEQ ID NO:64-99, or SEQ ID NO:118-136. Preferably, the cytidine deaminase consists of the amino acid sequence of any one of SEQ ID NO:48-63, SEQ ID NO:64-99, or SEQ ID NO:118-136.
[0186] Table B-1. Amino acid sequences of de novo designed TadA variants with cytidine deaminase activity
[0187] Amino acid sequences of cytidine deaminases derived from SEQ ID NO:64 through directed evolution.
[0188] Amino acid sequences of cytidine deaminases derived from SEQ ID NO:64 through directed evolution. (Table B-3.)
[0189] Table C. Encoding nucleotide sequences of adenosine deaminase and cytidine deaminase
[0190] Other sequences used in this application, such as adapter sequences or sgRNA sequences, are listed in Tables D and E.
[0191] Table D. Connector Sequence
[0192] Table E. sgRNA sequence
[0193] Table F. Other Sequences
[0194] Table G-1. Amino acid sequences of fused and split forms of CBE and UGI
[0195] Table G-2. Nucleotide sequences of fusion and fragmentation forms of CBE and UGI
[0196] Table H. Cas9_ortholog1 nickase amino acid and nucleotide sequences
[0197] Split system
[0198] In some embodiments, the nucleobase editor disclosed herein includes a “splitting” system in which a Cas protein, deaminase, or both is split into two components and can be reconstructed to form a functional protein to perform base editing. “Split Cas9 protein,” “split Cas9,” or “split deaminase” refers to a Cas9 or deaminase protein provided as an N-terminal and C-terminal fragment encoded by two separate nucleotide sequences. Polypeptides corresponding to the N-terminal and C-terminal portions of the Cas9 protein or deaminase can be spliced to form a “reconstructed” Cas9 protein or deaminase. In some embodiments, the Cas9 protein or deaminase is split into two fragments within a disordered region of the protein, for example, as described in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014, or as described in Jiang et al. (2016) Science 351:867-871. PDB file:5F9R, both of which are incorporated herein by reference.
[0199] Embedded systems
[0200] In some embodiments, the nucleobase editor system disclosed herein includes an embedded system in which a deaminase (e.g., adenosine deaminase or cytidine deaminase) is inserted into Cas9. In some embodiments, this application provides a fusion protein comprising a nucleic acid-guided nuclease protein (e.g., Cas9 or any nuclease disclosed herein) and an embedded nucleotide deaminase protein domain. In some embodiments, the embedded nucleotide deaminase protein domain is an embedded adenosine deaminase protein domain. In some embodiments, the embedded adenosine deaminase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any one of SEQ ID NO:1-47 (e.g., SEQ ID NO:2-47). In some embodiments, the embedded nucleotide deaminase protein domain is an embedded cytidine deaminase protein domain. In some embodiments, the embedded cytidine deaminase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any one of SEQ ID NO:48-63, 64-99, or 118-136 (e.g., SEQ ID NO:48-63, 65-99, or 118-136).
[0201] In some implementations, the fusion protein further includes a UGI domain.
[0202] Example
[0203] The compositions and methods disclosed herein are further described in the following examples, which are merely illustrative and do not imply that the scope of the invention is limited thereto. The scope of the invention is defined by the claims.
[0204] Example 1. Evaluation of engineered TadA variants in sgRNA library cells
[0205] TadA deaminase originates from nature. Through rational design, mutations were introduced into it, resulting in a highly active DNA adenosine deaminase mutant, T8.4-V0 (SEQ ID NO:1). Subsequently, directed evolution was used to further enhance its activity by introducing single or multiple mutations, yielding a series of TadA mutants (Table 1).
[0206] Table 1. Mutations contained in different adenosine deaminase variants relative to T8.4-V0 (SEQ ID NO:1)
[0207] To systematically evaluate the base editing capabilities of each mutant in mammalian cells, we used an sgRNA library cell line (Max W. Shen et al., Nature, 2018 Nov; 563(7733):646-651. doi:10.1038 / s41586-018-0686-x, cloned and constructed by Genewiz Biotechnology Co., Ltd.). This library cell line was derived from HEK293T and integrated over 10,000 target sites paired with homologous sgRNAs.
[0208] The nucleotide sequence encoding ecTadA8e in the ABE8e expression vector (Michelle F. Richter et al., Nat Biotechnol, 2020 Jul; 38(7):883-891. doi:10.1038 / s41587-020-0453-z, gene cloning was completed by GenScript Biotech) was replaced with the nucleotide sequence encoding the TadA variant designed by the inventors. In the experiment, a series of mRNAs encoding ABE were generated by in vitro transcription and their activity was tested in expression library cell lines. 20-24 hours before transfection, HEK293T library cells (1x10⁻¹²) were... 5 (Number of cells) were seeded in 24-well plates (Corning). Following the manufacturer's protocol, 400 ng of mRNA encoding the ABE variant was transfected into cells using Lipofectamine MessengerMAX (ThermoFisher, catalog number LMRNA003). After 72 hours, genomic DNA was extracted from the cells and editing efficiency was analyzed by next-generation sequencing (NGS). sgRNAs detected in all samples were defined as effective sgRNAs, and the editing efficiency of these paired target sites (A through G) was analyzed. To quantify the potency of each variant, its average editing efficiency value was calculated and normalized to the average editing efficiency value for T8.4-V0. As shown in Table 2 and Figure 1, the editing efficiency of T8.4-V0 is slightly higher than that of ABE8.8 (a known ABE variant, as a control; Gaudelli NM et al., Nat Biotechnol. 2020 Jul; 38(7):892-900. doi:10.1038 / s41587-020-0491-6.), and the average editing efficiency of many engineered variants of the present invention is higher than that of T8.4-V0, demonstrating that the engineered variants of the present invention are suitable as an activity-enhanced adenine base editor.
[0209] Table 2. Editing efficiency of different TadA variants for A to G in library cells.
[0210] Figure 1 shows the overall A-to-G editing efficiency (i.e., A:T to G:C conversion rate) of some adenosine deaminase variants of the present invention (T8.4-V1, T8.4-V2, T8.4-V4, T8.4-V7, T8.4-V10, T8.4-V12) and the known ABE variant ABE8.8 (SEQ ID NO: 113) compared to T8.4-V0. Compared to T8.4-V0, some variants of the present invention (T8.4-V1, T8.4-V2, T8.4-V4, T8.4-V10) show similar editing windows but higher editing efficiency, while other variants (T8.4-V7, T8.4-V12) show wider editing windows and higher editing efficiency. Here, "editing window" refers to the position of editable A within the original spacer region. "Wider editing window" means that there are more editable A positions within the original spacer region.
[0211] Example 2. Evaluation of the activity of adenine base editor (ABE) variants in mice.
[0212] Wild-type mice (strain C57BL / 6, purchased from Biocytogen, approximately 7-9 weeks old and weighing 20-25g at the time of the experiment) were used to further evaluate the in vivo editing effect of the base editor constructed according to the adenosine deaminase variant selected in Example 1. The inventors designed a specific guide RNA (gRNA_1: CCCATACCTTGGAGCAACGG, SEQ ID NO: 108) to disrupt the mouse Pcsk9 gene. The mRNA encoding the ABE variant obtained through in vitro transcription and the chemically synthesized gRNA_1 were combined and encapsulated into lipid nanoparticles (LNPs) according to the method described in patent application WO / 2023 / 185697A2. These formulated LNPs were then injected into mice via tail vein at a dose of 0.15 mg / kg body weight (0.15 mg / kg body weight (mpk)).
[0213] One week after injection, mouse liver samples were collected, and DNA was extracted. PCR amplification and NGS sequencing analysis were then performed on the region near the editing site. As shown in Table 3, mice treated with the engineered mutant of this invention exhibited higher A to G editing efficiency compared to T8.4-V0.
[0214] Table 3. DNA editing efficiency (%) achieved by the adenosine deaminase variant of the present invention in wild-type mice.
[0215] Example 3. Effect of different adapter sequences on the activity of adenine base editor (ABE) variants
[0216] The adenosine deaminase variants tested in Examples 1 and 2 were linked to nCas9 (D10A) via a 32-amino acid linker (SGGSSGGSSGSETPGTSESATPESSGGSSGGS, SEQ ID NO:104). To test the effect of different linker sequences on the editing activity of ABE variants, we tested the effect of changing the length and type of the linker or omitting the linker on the activity of T8.4-V0. The same experimental system as in Example 1 was used for the tests, and the results are shown in Table 4. Using the editor (T8.4-L0) obtained by linking T8.4-V0 and nCas9 (D10A) via the linker shown in SEQ ID NO:104 as a control, the change of linker did not reduce the base editing activity of T8.4-V0. Among them, T8.4-L3, obtained by linking T8.4-V0 and nCas9 using the linker sequence PAPAPAP (SEQ ID NO:106), showed a significant increase in editing activity.
[0217] Table 4. Editing efficiency of different adapter sequences from A to G in library cells.
[0218] Example 4. Modifying Cas9 to enhance base editing activity
[0219] The editing activity of an adenine base editor (ABE) variant was enhanced by modifying nCas9 (nickase D10A, SEQ ID NO: 114). Based on a rational design of the structure, mutations were introduced into Cas9, and activity was tested using the same experimental system as in Example 1. The editor T8.4-C0, composed of T8.4-V0 and unmutated nCas9, was used as a control. The results are shown in Table 5. Introducing D700R or E1243R into nCas9 can improve the editing activity and efficiency of the adenine base editor.
[0220] Table 5. A to G editing efficiencies of T8.4-V0 and mutant Cas9 in library cells.
[0221] Example 5. De novo design of a highly active adenosine deaminase
[0222] Based on structural and sequence information, novel TadA variants with adenosine deaminase activity were designed. The polynucleotide sequence encoding ecTadA8e was replaced in the ABE8e expression vector with a polynucleotide sequence encoding the newly designed TadA variant, generating a series of ABE variants. The editing activity of these variants was measured in HEK293T cells, as described below.
[0223] 20-24 hours before transfection, HEK293T cells (1.5 x 10⁻⁶) were... 4Cells were seeded in 96-well plates (Corning). For targeted editing experiments, cells were transfected according to the manufacturer's protocol using 80 ng of an expression vector encoding an ABE variant, 40 ng of spCas9gRNA_2 (CCCGCACCTTGGCGCAGCGG, SEQ ID NO:109) expression plasmid, and 3.6 μL of FuGENE HD transfection reagent (E2312, Promega). After 72 hours, genomic DNA was extracted from the cells, and the editing efficiency of the ABE variants was analyzed by NGS. The results are shown in Table 6. Many newly designed ABE variant sequences exhibited A to G editing activity, with A-DN-8, A-DN-9, A-DN-10, A-DN-11, and A-DN-12 showing A to G editing efficiencies exceeding 60%.
[0224] Table 6. Editing efficiency of A to G under different ABEs
[0225] Example 6. De novo design of TadA variants with cytidine deaminase activity
[0226] TadA was modified to mediate C-to-T editing and, compared to natural cytidine deaminases such as rAPOBEC used in BE4max, has advantages such as smaller size and fewer off-target effects. Therefore, based on structural and sequence information, a new TadA with cytidine deaminase activity was designed. The polynucleotide sequence encoding rAPOBEC (SEQ ID NO: 103) was replaced in the BE4max expression vector (encoding nucleotide sequence SEQ ID NO: 102) with a polynucleotide sequence encoding the newly designed TadA variant to generate a series of CBE variants. Using the experimental method described in Example 5, their editing activity was measured in the HEK293T cell line. The sgRNA sequence used was sgRNA_3 (GGAATCCCTTCTGCAGCACC, SEQ ID NO: 110), and its A-to-G and C-to-T editing activities were detected by NGS sequencing. As shown in Table 7, many of the newly designed CBE variants exhibited C-to-T editing activity, with variants C-DN-7 and C-DN-8 showing editing activities comparable to the control BE4max. Furthermore, none of these CBE variants possess A to G editing activity, demonstrating that they are specific cytosine deaminases.
[0227] Table 7. Editing efficiency of A to G and C to T for different CBE variants
[0228] Example 7. Enhancing the C-to-T editing activity of T88.74 through directed evolution.
[0229] To enhance the activity of the cytidine deaminase T88.74-V0 (SEQ ID NO:64) constructed by the inventors, we constructed a site-saturated mutant library of T88.74 and performed directed evolution on it. The experimental system used three plasmids: a selection plasmid containing the chloramphenicol resistance gene (CamR) with two inactivation mutations (L158P and H193R); an sgRNA expression plasmid generating two sgRNAs targeting the mutation sites in the CamR gene; and an editor (i.e., CBE variant) plasmid library expressing the T88.74-dCas9-2xUGI variant from the library. Library members capable of introducing C:G to T:A editing at the two mutation sites restored chloramphenicol resistance, enabling bacteria to survive in the presence of chloramphenicol.
[0230] We co-transformed all three plasmids into *E. coli* strain DH10B. After overnight editor library expression induction, the bacteria were inoculated onto culture plates containing chloramphenicol (selective; 16 or 64 μg / ml) or without chloramphenicol (non-selective). After overnight incubation at 37°C, plasmids from colonies grown on selective and non-selective plates were extracted and analyzed using next-generation sequencing (NGS). We calculated the percentage of mutations and the fold enrichment on selective plates. The top 50 variants are shown in Figures 3 and 2, respectively.
[0231] Example 8. Validation of beneficial mutations in bacterial and mammalian sgRNA library cells.
[0232] We first evaluated the CBE activity of enriched T88.74 variants in *E. coli* using the system described in Example 7, and measured and compared the editing efficiency of each variant at two target sites on the selection plasmid by NGS (CAM-gRNA1: TGATCCGAACGTGGCCAATA, SEQ ID NO: 111; CAM-gRNA2: TACGGCGCGGTGCACCTGGA, SEQ ID NO: 112). The 11 highly active variants and their corresponding editing efficiencies are listed in Table 8.
[0233] Table 8. C-to-T editing efficiency of different variants of T88.74 in bacteria
[0234] We then selected mutations from the first seven T88.74 variants validated in *E. coli* for combination and further tested their CBE activity in HEK293T sgRNA library cells. In this HEK293T sgRNA library, over 10,000 different sgRNAs and their corresponding target sequences were inserted into hotspot regions of the HEK293T genome, allowing for the analysis of CBE activity on multiple sgRNAs in a single, simple experiment.
[0235] To evaluate the CBE activity of selected T88.74 variants, 400 ng of T88.74 variant-based CBE mRNA was transfected into sgRNA library cells using Lipofectamine MessengerMAX (ThermoFisher, catalog number LMRNA003) according to the manufacturer's instructions. Genomic DNA was extracted from the cells after 72 hours, and editing efficiency was analyzed by NGS. sgRNA detected in all samples was defined as effective sgRNA. C-to-T editing efficiencies at these paired target sites were analyzed. The mean editing efficiency value for each variant in the sgRNA library was calculated. As shown in Table 9 and Figure 4, the mean editing efficiency of many engineered CBE variants was significantly higher than that of the original T88.74-V0.
[0236] Table 9. C-to-T editing efficiency of different variants of T88.74 in library cells.
[0237] We then introduced three enhancing mutations (T88.74-V14) validated in bacterial and mammalian cells into the original directed evolution library and selected again using higher concentrations of chloramphenicol (64 or 128 μg / ml) as described in Example 7. The top 50 variants from the second round of directed evolution are shown in Figure 5.
[0238] Next, we selected several enhancing mutations and evaluated their activity in HEK293T sgRNA library cells using the methods described above. Variants T88.74-V21 and T88.74-V22 showed significantly enhanced C-to-T editing efficiency compared to the T88.74-V14 variant, as shown in Table 10 and Figure 6.
[0239] Table 10. C-to-T editing efficiency of different variants of T88.74 in library cells.
[0240] We then added a subset of novel enhancing mutations (Table 10) confirmed based on T88.74-V14 to T88.74-V17, one of the previously screened strongest variants, and evaluated their activity in HEK293T sgRNA library cells as described above. Variants T88.74-V27 and T88.74-V33 showed significantly enhanced C-to-T editing efficiency compared to variant T88.74-V17, as shown in Table 11 and Figure 7.
[0241] Table 11. C-to-T editing efficiency of different variants of T88.74 in library cells.
[0242] Example 9. Further screening and validation of highly active T88.74 mutations in bacteria.
[0243] From the enriched variants obtained in Example 7, we selected a new T88.74 variant and evaluated its cytosine base editing (CBE) activity in *E. coli*. The editing efficiency of each variant was compared by performing NGS sequencing on two target sites on the selected plasmid (CAM-gRNA1: TGATCCGAACGTGGCCAATA, SEQ ID NO: 111; CAM-gRNA2: TACGGCGCGGTGCACCTGGA, SEQ ID NO: 112). Two additional highly active variants, T88.74-V36 (R130T) and T88.74-V37 (H34N), were identified; the editing efficiencies of these variants are listed in Table 12.
[0244] Table 12. C-to-T editing efficiency of different variants of T88.74 in bacteria
[0245] Furthermore, we further increased the screening pressure based on the screening system described in Example 7: a third inactivating mutation (C31R) was introduced into the chloramphenicol resistance gene (CamR), and a third guide RNA (CAM-gRNA3: TCAACGCACCTATAACCAGA, SEQ ID NO: 145) was designed to restore its activity via CBE. Using this optimized screening system, we obtained a batch of new mutation sites. Subsequently, representative mutations were selected from the T88.74 variants with high enrichment folds or large final proportions, and their editing efficiency at the three target sites (CAM-gRNA1, CAM-gRNA2, and CAM-gRNA3) was compared in E. coli. The results showed that another mutation, A106R, could improve activity, and the editing efficiency of its corresponding variants is listed in Table 13.
[0246] Table 13. C-to-T editing efficiency of different variants of T88.74 in bacteria
[0247] Example 10. Validation of beneficial mutations in mammalian sgRNA library cells
[0248] Beneficial mutations validated in Example 9 and previous rounds of screening were introduced into the current variants, and their CBE activity was further evaluated in HEK293T sgRNA library cells according to the previously described methods. As shown in Figure 8 and Table 14, the variants obtained by introducing the new mutations further enhanced the C to T editing activity level compared to the previous effective T88.74 variants, validating the applicability of these new variants in mammalian systems.
[0249] Table 14. C-to-T editing efficiency of different variants of T88.74 in library cells.
[0250] Example 11. Designing new mutations by combining structural and evolutionary information and validating them in multiple systems.
[0251] Based on the CBE domain and sequence alignment results and evolutionary conservation analysis, we designed a series of mutation sites that may further enhance CBE activity. First, we evaluated the editing efficiency of these mutations in an *E. coli* system (experimental procedure as in Example 9), and the results are shown in Table 15. Simultaneously, we validated some variants in mammalian systems (including HEK293T sgRNA library cells, Huh7 cell line, and human primary hepatocytes PHH).
[0252] The transfection protocol for the Huh7 cell line was as follows: 24 hours prior to transfection, cells were seeded in 96-well plates at a density of 8000 cells / well. For each sample, 100 ng of base editor mRNA and 100 ng of terminally modified synthetic sgRNA were transfected using Lipofectamine MessengerMAX (ThermoFisher, catalog number LMRNA003) according to the manufacturer's protocol. Genomic DNA was harvested from the Huh7 cells 72 hours later, and editing efficiency was analyzed by amplicon sequencing using NGS.
[0253] The transfection protocol for PHH cells was as follows: Cells were counted and seeded into 48-well plates (ThermoFisher, catalog number 877272) coated with Bio-coat collagen I at a density of 132,000 cells / well. The seeded cells were allowed to settle and adhere in a tissue culture incubator at 37°C and a 5% CO2 atmosphere for 4–6 hours. Following the manufacturer's protocol, 125 ng of base editor mRNA and 125 ng of terminally modified sgRNA were transfected into each well using RNAiMax. After 24 hours, the culture medium was replaced with fresh medium. 72 hours after transfection, genomic DNA was extracted from the cells, and editing efficiency was analyzed using NGS via amplicon sequencing.
[0254] The results of C-to-T editing efficiency in library cells and in HepG2 and PHH cells are listed in Tables 16-17 and Figure 9, showing that multiple variants exhibited effects that further enhanced CBE activity.
[0255] Table 15. C-to-T editing efficiency of different variants of T88.74 in bacteria
[0256] Table 16. C-to-T editing efficiency of different variants of T88.74 in library cells.
[0257] Table 17. C-to-T editing efficiency of different variants of T88.74 in HepG2 and PHH cells.
[0258] Example 12. Evaluation of the effect of dissociated UGI functional elements on CBE activity in mice.
[0259] Wild-type mice of strain C57BL / 6 (purchased from Biocytogen) were used to compare the effects of UGI functional elements expressed in fused or fragmented forms on CBE activity.
[0260] A guide RNA (gRNA_4:CAGGTTCCATGGGATGCTCT, SEQ ID NO:146) specifically targeting the mouse Pcsk9 gene was designed, and the CBE-encoded mRNA obtained by in vitro transcription was prepared in two forms:
[0261] (1) Fusion protein form: CBE-UGI-fused; and
[0262] (2) Protein splitting: CBE-split and UGI encoding nucleotide sequences are mixed at a mass ratio of 7:1.
[0263] Exemplary forms of the protein and their encoding nucleotide sequences are shown in Tables G-1 and G-2. One fusion protein, as shown in SEQ ID NO:137, is a fusion protein of deaminase T88.74-V41 (SEQ ID NO:123), nCas9 (SEQ ID NO:114), and UGI (SEQ ID NO:139), wherein the first amino acid (i.e., M) of nCas9 and UGI is removed. The split protein, as shown in SEQ ID NO:138, is a fusion protein of deaminase T88.74-V41 (SEQ ID NO:123) and nCas9 (SEQ ID NO:114), wherein the first amino acid (i.e., M) of nCas9 is removed.
[0264] The above mRNA was mixed with chemically synthesized gRNA_4 and encapsulated into lipid nanoparticles (LNPs) using the method described in WO / 2023 / 185697A2. The mixture was administered to mice via tail vein injection at a dose of 0.5 mg / kg body weight.
[0265] One week after injection, liver tissue samples were collected from mice, DNA was extracted, and the target region was amplified by PCR and NGS sequencing analysis. The C-to-T editing results in mice are shown in Table 18. The results show that the split form of the CBE system exhibited significantly higher editing activity than the fusion form in vivo, indicating that the split UGI strategy can effectively improve the in vivo editing efficiency of CBE.
[0266] Table 18. Effects of UGI component separation on CBE activity in mice
[0267] Additionally, it should be noted that the CBE used in this embodiment is T88.74-V41 (SEQ ID NO:123), which was randomly selected from the CBE series variants of the present invention and serves as an example. This result indicates that the CBE variants of the present invention can be fused with UGI or not during use. For example, before use, the encoding nucleotide sequences of both can be mixed in a certain mass ratio and encapsulated into lipid nanoparticles (LNPs). For the fusion protein form, the nucleotide sequence encoding a deaminase, the nucleotide sequence encoding a polynucleotide programmable DNA-binding domain (e.g., Cas9), and optionally the nucleotide sequence encoding UGI can be operably linked together in a reading frame conformal manner, thereby enabling the expression of the fusion protein; and the nucleotide sequence encoding the fusion protein can be encapsulated into lipid nanoparticles (LNPs) or mixed with gRNA and then encapsulated into LNPs. For the fragmented protein form, there are usually two different nucleotide sequences: the first nucleotide sequence encodes a deaminase and a polynucleotide programmable DNA-binding domain (e.g., Cas9), and the second nucleotide sequence encodes UGI. The first and second nucleotide sequences are mixed in a certain mass ratio (e.g., 7:1) and optionally encapsulated together with gRNA into lipid nanoparticles (LNPs).
[0268] Example 13. Testing the activity of T88.74 mutant fusion with other Cas proteins
[0269] To test whether the T88.74 mutant could function with other Cas-binding proteins, the spCas9 nickase was replaced with Cas9_ortholog1 nickase (amino acid sequence SEQ ID NO:143, nucleotide sequence SEQ ID NO:144), and its editing activity was assessed by fusion with T88.74-V27. After expression as mRNA, it was co-transfected with the following sgRNAs into PHH cells to test activity. Transfection was performed according to the procedure in Example 11, and the editing activity was detected by NGS (as shown in Table 19), indicating that the T88.74 variant can fuse with other Cas proteins to produce highly efficient editing.
[0270] Table 19. C-to-T editing efficiency of T88.74-V27 fused with Cas9_ortholog1 nickase in PHH cells.
[0271] Other implementation plans
[0272] It should be understood that although the invention has been described in conjunction with its detailed description, the foregoing description is intended to be illustrative and not to limit the scope of the invention, which is defined by the appended claims. Other aspects, advantages, and improvements are within the scope of the following claims.
Claims
An adenosine deaminase, said adenosine deaminase comprising: (i) an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any one of SEQ ID NO:2-47; or (ii) Select one or more substitutions from the group consisting of D12H, H21D or H21Q, A61C, V95T, S155R, S156N or S156K, C159R, E164V or E164W, E167R or E167W, K169N relative to SEQ ID NO:1, wherein the amino acid position is relative to the number in SEQ ID NO:
1. A cytidine deaminase, said cytidine deaminase comprising: (i) an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any one of SEQ ID NO: 48-63, 65-99, or 118-136; or (ii) Select one or more substitutions from the group consisting of A17N, H34N, D36L, I46Q, G64C, G68T, Y70G or Y70H or Y70K, L73F or L73R, V79T, F81L, M91L, H93N, G102S, A103V, A106R, D112N, A119G, N124V or N124P, R130V or R130T, S139K, A140R, C143A or C143S, R151M, K153G or K153R relative to SEQ ID NO:64, wherein the amino acid position is relative to the numbering of SEQ ID NO:
64. A nucleobase editor comprising: (i) a polynucleotide programmable DNA binding domain; and (ii) a deaminase, wherein the deaminase is the adenosine deaminase of claim 1 or the cytidine deaminase of claim 2. According to claim 3, the nucleobase editor further comprises one or more uracil glycosylation inhibitor (UGI) domains. According to claim 4, the nucleobase editor wherein the polynucleotide programmable DNA binding domain, the deaminase, and the UGI domain form a fusion protein; or the nucleotide sequence encoding the fusion protein is encapsulated in a lipid nanoparticle (LNP) or mixed with gRNA and then encapsulated in an LNP; or the nucleotide sequence encoding the polynucleotide programmable DNA binding domain and the deaminase is mixed with the nucleotide sequence encoding the UGI domain and then encapsulated in an LNP; or the nucleotide sequence encoding the polynucleotide programmable DNA binding domain and the deaminase, the nucleotide sequence encoding the UGI domain, and gRNA are mixed and then encapsulated in an LNP. According to claim 3, the nucleobase editor is an adenosine base editor (ABE) and / or a cytidine base editor (CBE). According to any one of claims 3-6, the nucleobase editor wherein the polynucleotide programmable DNA binding domain comprises a domain selected from the group consisting of the Cas9 domain, Cas12 domain, TnpB domain, IscB domain, homing endonuclease domain, zinc finger DNA binding domain, and transcription activator-like effector (TALE) DNA binding domain. According to claim 7, the nucleobase editor wherein the polynucleotide programmable DNA binding domain comprises a Cas9 domain. According to claim 8, the nucleobase editor wherein the Cas9 domain is Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus lCas9 (St1Cas9), Streptococcus pyogenes Cas9 (SpCas9) or a variant thereof. According to claim 8, the Cas9 domain is a mutated Cas9 containing one or two amino acid substitutions selected from D700R or E1243R, preferably, the mutated Cas9 contains the amino acid sequence of SEQ ID NO: 115 or 116, or is composed of the amino acid sequence of SEQ ID NO: 115 or 116. The nucleobase editor according to any one of claims 3-6, wherein the polynucleotide programmable DNA binding domain is an inactive nuclease or nickase variant. Nucleobase editor according to any one of claims 3-6, wherein the nucleobase editor comprises a nucleic acid-guided nuclease protein and an embedded nucleotide base editor domain. According to claim 12, the embedded nucleotide base editor domain is an embedded adenosine base editor (ABE) domain or an embedded cytidine base editor (CBE) domain. According to claim 13, the embedded adenosine base editor (ABE) domain is an embedded adenosine deaminase protein domain, or the embedded cytidine base editor (CBE) domain is an embedded cytidine deaminase protein domain. According to claim 14, the nucleobase editor wherein the polynucleotide programmable DNA binding domain comprises a linker that is side-attached to at least one embedded domain protein. A fusion protein comprising a multinucleotide programmable DNA-binding domain linked together and at least one nucleobase editor domain comprising a deaminase, wherein the deaminase is the adenosine deaminase of claim 1 or the cytidine deaminase of claim 2. A polynucleotide encoding the adenosine deaminase according to claim 1 or the cytidine deaminase according to claim 2. A polynucleotide encoding a nucleobase editor according to any one of claims 3-15. A polynucleotide encoding the fusion protein according to claim 16. An expression vector comprising a polynucleotide according to any one of claims 17-19. The expression vector of claim 20, wherein the vector is a viral vector selected from the group consisting of an adeno-associated virus (AAV), a retroviral vector, an adenoviral vector, a lentiviral vector, a Sendai virus vector, and a herpes virus vector. A cell comprising the expression vector of claim 20 or 21. A method of base editing, the method comprising contacting a polynucleotide sequence with the nucleobase editor of any one of claims 3-15, wherein the deaminase deaminates a nucleobase in the polynucleotide, thereby editing the polynucleotide sequence. A method of correcting a genetic deficiency in a subject, the method comprising administering to the subject: the nucleobase editor of any one of claims 3-15 or a polynucleotide encoding the nucleobase editor of any one of claims 3-15, and one or more guide polynucleotides that direct the nucleobase editor to deaminate a target nucleobase in a target nucleotide sequence of the subject, thereby correcting the genetic deficiency. The method of claim 24, comprising delivering the nucleobase editor or polynucleotide encoding the nucleobase editor, and one or more guide polynucleotides to a cell of the subject. The method of claim 24, wherein the subject is a non-human mammal or a human. The method of any one of claims 24-26, wherein deamination of the target nucleobase results in the target nucleobase being replaced with a wild-type nucleobase. A molecular complex comprising: the nucleobase editor of any one of claims 3-15 or the fusion protein of claim 16, and one or more guide RNA sequences, tracrRNA sequences, or target DNA sequences. A kit comprising: the nucleobase editor of any one of claims 3-15 or the fusion protein of claim 16, or the expression vector of claim 20, or the molecular complex of claim 28, and instructions for use. A pharmaceutical composition comprising: the nucleobase editor of any one of claims 3-15 or the fusion protein of claim 16, or the expression vector of claim 20, or the molecular complex of claim 28, and a pharmaceutically acceptable excipient. A mutated Cas9 comprising one or two amino acid substitutions selected from D700R or E1243R, preferably wherein the mutated Cas9 comprises or consists of the amino acid sequence of SEQ ID NO: 115 or 116. A composition comprising: (i) a first nucleotide sequence comprising a first open reading frame encoding a polypeptide comprising a deaminase and a polynucleotide programmable DNA binding domain; (ii) a second nucleotide sequence comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI) domain; wherein the first nucleotide sequence is different from the second nucleotide sequence, optionally, wherein the composition is encapsulated in a lipid nanoparticle (LNP). The composition of claim 32, wherein the composition further comprises (iii) a third nucleotide sequence comprising a gRNA sequence, wherein the first nucleotide sequence, the second nucleotide sequence, and the third nucleotide sequence are each different nucleotide sequences.