Compact proteins

Truncated nucleic acid base editors, optimized for size and efficiency, address the limitations of multiple AAV delivery in CRISPR systems, enhancing editing efficacy and safety by fitting within a single AAV vector.

JP2026516500APending Publication Date: 2026-05-25VERTEX PHARMACEUTICALS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
VERTEX PHARMACEUTICALS INC
Filing Date
2024-05-10
Publication Date
2026-05-25

AI Technical Summary

Technical Problem

Existing CRISPR-based genome editing technologies, such as base editors, are limited by the size constraints of AAV vectors, requiring multiple AAVs for delivery and increasing viral dose, which poses safety concerns and production burdens.

Method used

Development of truncated nucleic acid base editors with specific amino acid lengths and linkers, allowing for a single AAV to encapsulate the base editing components, reducing the overall size and maintaining or enhancing editing efficiency.

Benefits of technology

The truncated base editors achieve efficient genome editing with reduced size, potentially lowering viral doses and safety risks, while maintaining or improving editing activity compared to wild-type editors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026516500000030
    Figure 2026516500000030
  • Figure 2026516500000031
    Figure 2026516500000031
  • Figure 2026516500000032
    Figure 2026516500000032
Patent Text Reader

Abstract

In some embodiments, compact protein compositions, such as compact nucleic acid base editors that can be used in gene editing technologies, are disclosed herein.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Reference to Electronic Sequence Listing The content of the electronic sequence listing (41816_PCT_SequenceListing.xml; size: 172 KB; and creation date: May 6, 2024) is hereby incorporated by reference in its entirety into this specification.

[0002] Cross - Reference to Related Applications This application claims the priority of U.S. Patent Provisional Application No. 63 / 501,793, filed on May 12, 2023, the disclosure of which is hereby incorporated by reference in its entirety into this specification.

Background Art

[0003] Clustered regularly interspaced short palindromic repeats (CRISPR) together with CRISPR - associated (Cas) proteins comprise an RNA - guided adaptive immune system in archaea and bacteria. These systems target and inactivate nucleic acids derived from foreign genetic elements to provide immunity. Utilizing technologies such as CRISPR can facilitate the correction of genetic diseases by single - base mutations.

[0004] Base editors are genome editing technologies that leverage the ability of CRISPR-Cas proteins to find specific sequences of DNA guided by modular guide RNA sequences. CRISPR base editor platforms (BEs) have a unique ability to produce precise, user-defined genome editing events without requiring donor DNA molecules. BEs are a class of gene editing enzymes that often contain Cas niccas fused to a nucleotide deaminase domain. In principle, BEs localize to target regions in the genome guided by gRNA. Once bound, the Cas9 complex may displace the unbound strand to form an ssDNA R loop. The R loop may become reachable to the tethered deaminase domain, thereby allowing cytidine deaminase base editors (CBEs, C:G to T:A) to deaminate C to U, which then base-pairs as T, and adenosine deaminase base editors (ABEs, A:T to G:C) to deaminate A to I, which then base-pairs as G. Next, when the non-edited strand is simultaneously cleaved by core Cas9 nickase, DNA repair is stimulated, using the newly deaminated bases as a template for DNA polymerization, thereby preserving the edits on both strands of DNA. In contrast to traditional nuclease-dependent genome editing approaches, BE does not produce double-strand DNA breaks, does not require a DNA donor template, and is more efficient at editing non-dividing cells, making BE an attractive drug for in vivo therapeutic genome editing.

[0005] Vectors such as AAVs can be used to pack genetic information-coding cargo such as BEs. In addition to the base editor itself, it is desirable that the AAV encoding the base editor also encodes a guide RNA, a promoter that drives the expression of the base editor and single guide RNA, and a cis-regulatory element. Base editors are generally too large to fit into a single AAV carrying cargo with a size limit of approximately 4.7 kb without inverted end sequences (ITRs). Delivery of BEs by AAVs has generally been addressed by relying on the use of intein trans-splicing to split the BE between two AAVs and assemble it into a full-length effector. While effective, this approach requires successful transduction of target cells by both AAVs and successful trans-splicing of the bipartite inteins. The requirement to deliver two AAV vectors also increases the viral dose required for treatment, raising safety concerns and burdening AAV production. There is a need for smaller, more compact CRISPR-based cargoes (e.g., BEs) that can be encoded in nucleic acids and encapsulated by a single AAV. [Overview of the project] [Problems that the invention aims to solve]

[0006] This disclosure pertains to truncated proteins (e.g., truncated nucleic acid base editors) and nucleic acids encoding such truncated proteins. [Means for solving the problem]

[0007] In a first embodiment, the disclosure relates to a nucleic acid base editor comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, wherein the amino acid lengths are between 1175 and 1249. In some embodiments, the nucleic acid base editor is divided into groups of 1175-1249, 1180-1249, 1190-1249, 1200-1249, 1210-1249, 1220-1249, 1230-1249, 1240-1249, 1180-1200, 1180-1190, 1200-1249, 1200-1240, and 1 The amino acid lengths are between 200-1230, 1200-1220, 1200-1210, 1210-1249, 1210-1240, 1210-1230, 1210-1220, 1220-1249, 1220-1240, 1220-1230, 1230-1249, 1230-1240, or 1240-1249. In some embodiments, the linker has an amino acid length of 2-10, 3-7, 3-5, or 3 amino acids. In some embodiments, the nucleotide deaminase has an amino acid length of 145-166, 145-162, 145-158, 145-154, 145-151, 150-166, 150-162, 150-156, 150-153, 152-166, 152-158, 152-154, 157-166, 157-162, 161-166, or 161-163 amino acids. In some embodiments, the Cas9 nuclease has an amino acid length between 1010-1052, 1010-1040, 1010-1025, 1010-1015, 1020-1052, 1020-1040, 1020-1030, 1030-1052, 1030-1040, 1030-1035, 1035-1046, 1040-1052, 1040-1046, 1045-1052, or 1045-1048. In some embodiments, the linker comprises one or more prolines and one or more alanines. In some embodiments, the nucleotide deaminase comprises an amino acid sequence or a functional fragment thereof that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 1.In some embodiments, the nucleotide deaminase lacks the amino acids corresponding to the first 1, 2, 3, 4, 5, 6, 7, or 8 amino acids of SEQ ID NO: 1. In some embodiments, the nucleotide deaminase lacks the amino acids corresponding to the last 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids of SEQ ID NO: 1. In some embodiments, the nucleotide deaminase lacks the amino acids corresponding to the first 6 amino acids of SEQ ID NO: 1, and also lacks the amino acids corresponding to the last 8 amino acids of SEQ ID NO: 1. In some embodiments, the Cas9 nuclease includes an amino acid sequence or a functional fragment thereof that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 1-4, 1-6, 77-85, 124-126, 124-129, 124-143, 124-145, 131-134, 135-138, 122-129, 505-622, 687-688, 716-720, 718-719, 718-730, 765, and / or 865-873 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks amino acids corresponding to amino acids 1-4, 1-6, 77-85, 124-126, 124-129, 124-143, 124-145, 131-134, 135-138, 122-129, 505-622, 687-688, 716-720, 718-719, 718-730, 765, and / or 865-873 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks amino acids corresponding to: 718-719 and 765 of SEQ ID NO: 2; 718-719 and 687-688 of SEQ ID NO: 2; 716-720 and 687-688 of SEQ ID NO: 2; 1-4 and 718-730 of SEQ ID NO: 2; 1-4, 687-688, and 716-720 of SEQ ID NO: 2; 1-4 and 124-145 of SEQ ID NO: 2; or 1-4, 505-622, and 718-730 of SEQ ID NO: 2.In some embodiments, the Cas9 nuclease is no exception to one or more amino acids from the range of amino acids corresponding to: SEQ ID NO: 2-1 to 2-8, SEQ ID NO: 2-1 to 2-10, SEQ ID NO: 2-81 to 2-94, SEQ ID NO: 2-163 to 2-164, SEQ ID NO: 2-181 to 2-184; SEQ ID NO: 2-244 to 2-424, SEQ ID NO: 2-714 to 2-769, SEQ ID NO: 2-736 to 2-742, SEQ ID NO: 2-978 to 2-979, and / or SEQ ID NO: 2-1039 to 2-1052. In some embodiments, the Cas9 nuclease does not lack a range of amino acids corresponding to: amino acids 1-8 of SEQ ID NO: 2, 1-10 of SEQ ID NO: 2, 81-94 of SEQ ID NO: 2, 163-164 of SEQ ID NO: 2, 181-184 of SEQ ID NO: 2, 244-424 of SEQ ID NO: 2, 714-769 of SEQ ID NO: 2, 736-742 of SEQ ID NO: 2, 978-979 of SEQ ID NO: 2, and / or 1039-1052 of SEQ ID NO: 2. In some embodiments, the nucleic acid base editor has at least 50%, 60%, 70%, 80%, or 90% of the base editing activity compared to a nucleic acid base editor containing the sequence of SEQ ID NO: 3. In some embodiments, the base editing activity is evaluated using a polynucleotide containing the nucleotide sequence of SEQ ID NO: 4. In some embodiments, the base editing activity is evaluated at positions A5, A8, or A13 of SEQ ID NO: 4. In some embodiments, the nucleic acid base editor further comprises one or more nuclear localization signals (NLS) and one or more linkers that optionally fuse one or more NLS to the nucleic acid base editor. In some embodiments, one or more NLS are selected from c-myc NLS (e.g., an NLS containing the amino acid sequence of SEQ ID NO: 5), SV40 NLS (e.g., an NLS containing the amino acid sequence of SEQ ID NO: 7), and / or nucleoplasmin NLS (e.g., an NLS containing the amino acid sequence of SEQ ID NO: 9). In some embodiments, the nucleic acid base editor comprises a) a first NLS conjugated to the N-terminus of the nucleic acid base editor using optionally a first NLS linker, and b) a second NLS conjugated to the C-terminus of the nucleic acid base editor using optionally a second NLS linker.In some embodiments, the nucleic acid base editor further comprises a third NLS conjugated to a second NLS using a third NLS linker as appropriate. In some embodiments, the first NLS comprises the amino acid sequence of SEQ ID NO: 5, any first NLS linker comprises the amino acid sequence of SEQ ID NO: 6, the second NLS comprises the amino acid sequence of SEQ ID NO: 7, any second NLS linker comprises the amino acid sequence of SEQ ID NO: 8, the third NLS comprises the amino acid sequence of SEQ ID NO: 9, and any third NLS linker comprises the amino acid sequence of SEQ ID NO: 10. In some embodiments, one or more NLSs and one or more linkers that fuse one or more NLSs to the nucleic acid base editor have a combined total of 40, 37, 35, 32, 28, or 26 or fewer amino acids. In some embodiments, one or more NLSs and one or more linkers that fuse one or more NLSs to the nucleic acid base editor have a combined total of 32 or fewer amino acids.

[0008] In some embodiments, the linker amino acid sequence is SEQ ID NO: 11. In some embodiments, the linker amino acid sequence is SEQ ID NO: 12. In some embodiments, the linker amino acid sequence is SEQ ID NO: 13. In some embodiments, up to 8 amino acids are deleted from the amino terminus of the nucleotide deaminase. In some embodiments, up to 8 amino acids are deleted from the carboxy terminus of the nucleotide deaminase. In some embodiments, up to 8 amino acids are deleted from the carboxy terminus of the nucleotide deaminase. In some embodiments, up to 22 amino acids are deleted from the REC-lobe domain of the Cas9 nuclease. In some embodiments, the linker amino acid sequence is selected from the group SEQ ID NO: 11, SEQ ID NO: 12, and SEQ ID NO: 13; up to 8 amino acids are deleted from the amino terminus of the RNA adenine deaminase; up to 8 amino acids are deleted from the carboxy terminus of the RNA adenine deaminase; and up to 22 amino acids are deleted from the REC-lobe domain of the Cas9 nuclease. In some embodiments, the truncated fusion protein has a total coding size of less than 5 kb. In some embodiments, the truncated fusion protein has a total coding size of less than 4.9 kb. In some embodiments, the truncated fusion protein has a total coding size of less than 4.85 kb.

[0009] In some embodiments, the amino acid sequence encoding the nucleic acid base editor is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30, SEQ ID NO: 31, SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51, SEQ ID NO: 52, SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, or SEQ ID NO: 56. In some embodiments, the amino acid sequence encoding the nucleic acid base editor is selected from the group SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, and 56.

[0010] In some embodiments, the total activity of the nucleic acid base editor is at least 30% or more greater than the total activity of the wild-type nucleic acid base editor. In some embodiments, the total activity of the nucleic acid base editor is at least 50% or more greater than the total activity of the wild-type nucleic acid base editor. In some embodiments, the total activity of the nucleic acid base editor is at least 75% or more greater than the total activity of the wild-type nucleic acid base editor. In some embodiments, the total activity of the nucleic acid base editor is at least 80% or more greater than the total activity of the wild-type nucleic acid base editor. In some embodiments, the total activity of the nucleic acid base editor is at least 85% or more greater than the total activity of the wild-type nucleic acid base editor. In some embodiments, the total activity of the nucleic acid base editor is at least 90% or more greater than the total activity of the wild-type nucleic acid base editor. In some embodiments, the nucleotide deaminase is RNA adenine deaminase. In some embodiments, the RNA adenine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is TadA-8e. In some embodiments, the Cas9 nuclease is Cas9 nickase. In some embodiments, the Cas9 nickase is nSaCas9.

[0011] In some embodiments, the nucleic acid base editor is a truncated nucleic acid base editor comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, wherein the nucleic acid base editor has an amino acid length between 1175 and 1249. a) Nucleotide deaminase lacking the amino acids corresponding to the first 1, 2, 3, 4, 5, 6, 7, or 8 amino acids of SEQ ID NO: 1 and / or nucleotide deaminase lacking the amino acids corresponding to the last 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids of SEQ ID NO: 1; b) Cas9 nucleases lack one or more amino acids from the range of amino acids corresponding to amino acids 1-4, 1-6, 77-85, 124-126, 124-129, 124-143, 124-145, 131-134, 135-138, 122-129, 505-622, 687-688, 716-720, 718-719, 718-730, 765, and / or 865-873 of SEQ ID NO: 2 c) The linker between the deaminase and the Cas9 nuclease was selected from the group of SEQ ID NOs: 11, 12, and 13; d) The nucleic acid base editor has at least 50%, 60%, 70%, 80%, or 90% of the base editing activity compared to the nucleic acid base editor containing the sequence of Sequence ID No. 3; e) Nucleic acid base editors are i. Further comprising one or more nuclear localization signals (NLS), one or more NLS selected from c-myc NLS (e.g., an NLS containing the amino acid sequence of SEQ ID NO: 5), SV40 NLS (e.g., an NLS containing the amino acid sequence of SEQ ID NO: 7), and / or nucleoplasmin NLS (e.g., an NLS containing the amino acid sequence of SEQ ID NO: 9); ii. One or more NLSs are fused to a nucleic acid base editor via one or more linkers, as appropriate; iii. One or more NLSs and one or more linkers fusing one or more NLSs to a nucleic acid base editor have a combined total of 32 amino acids or less; and f) The amino acid sequences encoding the nucleic acid base editor are selected from the group of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, and 56.

[0012] In some embodiments, the truncated nucleic acid base editor has substantially the same base editing activity as the nucleic acid base editor containing the sequence of SEQ ID NO: 3.

[0013] Some aspects of this disclosure relate to proteins containing the Cas9 protein, wherein the Cas9 protein has an amino acid length between 1010-1052, 1010-1040, 1010-1025, 1010-1015, 1020-1052, 1020-1040, 1020-1030, 1030-1052, 1030-1040, 1030-1035, 1035-1046, 1040-1052, 1040-1046, 1045-1052, or 1045-1048. In some embodiments, the protein comprises a Cas9 protein, which comprises an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 2, or a functional fragment of SEQ ID NO: 2. In some embodiments, the Cas9 protein lacks one or more amino acids from the range of amino acids corresponding to amino acids 1-4, 1-6, 77-85, 124-126, 124-129, 124-143, 124-145, 131-134, 135-138, 122-129, 505-622, 687-688, 716-720, 718-719, 718-730, 765, and / or 865-873 of SEQ ID NO: 2. In some embodiments, the Cas9 protein lacks amino acids corresponding to amino acids 1-4, 1-6, 77-85, 124-126, 124-129, 124-143, 124-145, 131-134, 135-138, 122-129, 505-622, 687-688, 716-720, 718-719, 718-730, 765, and / or 865-873 of SEQ ID NO: 2 (i.e., lacking any one or combination of the aforementioned amino acid ranges). In some embodiments, the Cas9 protein lacks amino acids corresponding to: 718-719 and 765 of SEQ ID NO: 2; 718-719 and 687-688 of SEQ ID NO: 2; 716-720 and 687-688 of SEQ ID NO: 2; 1-4 and 718-730 of SEQ ID NO: 2; 1-4, 687-688, and 716-720 of SEQ ID NO: 2; 1-4 and 124-145 of SEQ ID NO: 2; or 1-4, 505-622, and 718-730 of SEQ ID NO: 2.In some embodiments, the Cas9 protein is not lacking one or more amino acids from the range of amino acids corresponding to: SEQ ID NO: 1-8; SEQ ID NO: 1-10; SEQ ID NO: 81-94; SEQ ID NO: 163-164; SEQ ID NO: 181-184; SEQ ID NO: 244-424; SEQ ID NO: 714-769; SEQ ID NO: 736-742; SEQ ID NO: 978-979; and / or SEQ ID NO: 1039-1052. In some embodiments, the Cas9 protein does not lack the amino acid ranges corresponding to: 1-8 of SEQ ID NO: 2; 1-10 of SEQ ID NO: 2; 81-94 of SEQ ID NO: 2; 163-164 of SEQ ID NO: 2; 181-184 of SEQ ID NO: 2; 244-424 of SEQ ID NO: 2; 714-769 of SEQ ID NO: 2; 736-742 of SEQ ID NO: 2; 978-979 of SEQ ID NO: 2; and / or 1039-1052 of SEQ ID NO: 2. In some embodiments, the Cas9 protein lacks one or more amino acids from the amino acid ranges corresponding to 81-94, 123-145, 244-424, 714-769, and / or 717-730. In some embodiments, the Cas9 protein lacks the amino acids corresponding to 81-94, 123-145, 244-424, 714-769, and / or 717-730. In some embodiments, the Cas9 protein is fused to a heterologous region. In some embodiments, the heterologous region is a deaminase. In some embodiments, the heterologous region is a protein useful for gene silencing (e.g., the KRAB domain, Dnmt3A, and / or Dnmt3L protein domains). In some embodiments, the heterologous region is a DNA methyltransferase, e.g., Dnmt1, Dnmt3A, or Dnmt3B. In some embodiments, the heterologous region is a regulator of a DNA methyltransferase factor [e.g., a catalytically inactive regulator of DNA methyltransferase (e.g., Dnmt3L)]. In some embodiments, the heterologous region is a protein useful for reversing gene silencing (e.g., the TET1 DNA demethylase catalytic domain). In some embodiments, the heterologous region is a transcriptional repressor. In some embodiments, the heterologous region is a transcriptional activator.

[0014] Some aspects of this disclosure relate to nucleic acids encoding a nucleic acid base editor having an amino acid length between 1175 and 1249, comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease. In some embodiments, the nucleic acid further comprises an inverted end sequence (ITR). In some embodiments, the nucleic acid further comprises a gRNA (e.g., sgRNA). In some embodiments, the nucleic acid further comprises a promoter sequence. In some embodiments, the total nucleic acid is less than 5 kb. In some embodiments, the amino acid sequence encoding the nucleic acid base editor is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30, SEQ ID NO: 31, SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51, SEQ ID NO: 52, SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, or SEQ ID NO: 56. In some embodiments, the amino acid sequence encoding the nucleic acid base editor is selected from the group SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, and 56.

[0015] Some aspects of this disclosure relate to vectors comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, and comprising a nucleic acid encoding a nucleic acid base editor having an amino acid length between 1175 and 1249. In some embodiments, the vector is selected from the group of adeno-associated viruses, adenoviruses, lentiviruses, bacteriophages, and virus-like particles. In some embodiments, the vector is an adeno-associated virus. In some embodiments, the total nucleic acid is less than 5 kb.

[0016] Certain embodiments of this disclosure relate to isolated cells comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, and a nucleic acid base editor having an amino acid length between 1175 and 1249. In some embodiments, the isolated cells comprise a nucleic acid encoding a nucleic acid base editor comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, and having an amino acid length between 1175 and 1249. In some embodiments, the total nucleic acid is less than 5 kb. In some embodiments, the isolated cells comprise a vector comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, and a nucleic acid encoding a nucleic acid base editor having an amino acid length between 1175 and 1249. In some embodiments, the total nucleic acid is less than 5 kb.

[0017] The patent or application file must include at least one color drawing. A copy of the published patent or patent application, accompanied by the color drawing(s), will be provided by the Patent Office upon request and payment of the required fees. [Brief explanation of the drawing]

[0018] [Figure 1]Figures 1A and 1B show the base editor activity of the manipulated base editor as a percentage of the wild-type base editor activity, where the linker between the deaminase and Cas9 nuclease was replaced with a 3-amino acid linker (SEQ ID NO: 11), a 5-amino acid linker (SEQ ID NO: 12), or a 7-amino acid linker (SEQ ID NO: 13). A) Percentage activity of pBE0015 (SEQ ID NO: 14) and pBE0016 (SEQ ID NO: 15) of wild-type nucleic acid base editor activity measured at position A5 of SEQ ID NO: 4, for nucleic acid base editor activity evaluation polynucleotides. B) Percentage activity of pBE0119 (SEQ ID NO: 16) of wild-type nucleic acid base editor activity measured at position A5 of SEQ ID NO: 4. [Figure 2] Figures 2A-C are graphs showing the base editor activity of manipulated base editors as a percentage of wild-type base editor activity in which a portion of TadA deaminase is deleted. A) pBE0019 (SEQ ID NO: 17), pBE0020 (SEQ ID NO: 18), and pBE0021 (SEQ ID NO: 19) were shown to have more than 100% base editor activity against SEQ ID NO: 3 measured at position A8 of SEQ ID NO: 4. B) Percentage base editor activity of pBE0057 (SEQ ID NO: 20), pBE0058 (SEQ ID NO: 21), and pBE0059 (SEQ ID NO: 22) against SEQ ID NO: 3 measured at position A8 of SEQ ID NO: 4. C) Percentage base editor activity of pBE0015 (SEQ ID NO: 14) and pBE0117 (SEQ ID NO: 40) against SEQ ID NO: 3 measured at position A8 of SEQ ID NO: 4. [Figure 3]Figures 3A-B are graphs showing manipulated base editor combinations that remove a portion of the TadA deaminase and replace the linker between the deaminase and the Cas9 nuclease with SEQ ID NO: 11, SEQ ID NO: 12, or SEQ ID NO: 13. A) pBE0105 (SEQ ID NO: 34) has 108% base editor activity compared to the wild-type base editor activity at position A5 of SEQ ID NO: 4. pBE0105 (SEQ ID NO: 34) has TadA with the first 4 amino acids deleted, TadA with the last 8 amino acids deleted, the linker of SEQ ID NO: 12, and the Cas9 nuclease with the first 4 amino acids deleted. pBE0107 (SEQ ID NO: 35) shows 184% base editor activity compared to the wild-type activity at position A5. pBE0107 (SEQ ID NO: 35) has a TadA molecule with the first 8 amino acids deleted, a TadA molecule with the last 4 amino acids deleted, the linker from SEQ ID NO: 12, and a Cas9 nuclease with the first 4 amino acids deleted. pBE0108 (SEQ ID NO: 36) exhibits 125% base editor activity compared to wild-type activity at position A5. pBE0108 (SEQ ID NO: 36) has a TadA molecule with the first 8 amino acids deleted, a TadA molecule with the last 6 amino acids deleted, and the linker from SEQ ID NO: 12. pBE0109 (SEQ ID NO: 37) exhibits 118% base editor activity compared to wild-type activity at position A5. pBE0109 (SEQ ID NO: 37) has a TadA molecule with the first 8 amino acids deleted, a TadA molecule with the last 6 amino acids deleted, the linker from SEQ ID NO: 12, and a Cas9 nuclease with the first 4 amino acids deleted. B) shows that pBE0110 (SEQ ID NO: 38) has 65% of the base-editing activity at position A8 compared to the wild-type activity. pBE0110 (SEQ ID NO: 38) has a TadA with the first 8 amino acids deleted, a TadA with the last 8 amino acids deleted, and the linker of SEQ ID NO: 12. [Figure 4]Figures 4A and 4B are graphs showing the base editor activity of manipulated base editors as a percentage of wild-type base editor activity when amino acids are deleted from either or both ends of the Cas9 nuclease. A) pBE0042 (SEQ ID NO: 30) has 27% of the base editor activity compared to the wild-type base editor activity at position A8 of SEQ ID NO: 4, and pBE0042 (SEQ ID NO: 30) deletes amino acids K2-Y5 from the Cas9 nuclease. pBE0043 (SEQ ID NO: 31) has 120% of the base editor activity compared to the wild-type base editor activity at position A8 of SEQ ID NO: 4, and pBE0043 (SEQ ID NO: 31) deletes amino acids E1040-G1053 from the carboxyl terminus of the Cas9 nuclease. B) pBE0120 (SEQ ID NO: 42) has 112% of the base editor activity at position A8 of SEQ ID NO: 4, and pBE0120 (SEQ ID NO: 42) has the first 6 amino acids from the amino terminus of the Cas9 nuclease deleted. pBE0121 (SEQ ID NO: 43) has 13% of the base editor activity at position A8 of SEQ ID NO: 4, and pBE0121 (SEQ ID NO: 43) has the first 8 amino acids from the amino terminus of the Cas9 nuclease deleted. pBE0122 (SEQ ID NO: 44) has 0% of the base editor activity at position A8 of SEQ ID NO: 4, and pBE0122 (SEQ ID NO: 44) has the first 10 amino acids from the amino terminus of the Cas9 nuclease deleted. [Figure 5]Figures 5A-C show graphs of the base activity of manipulated base editors as a percentage of wild-type base editor activity when amino acids are deleted from the REC lobe region of Cas9 nuclease. A) pBE0022 (SEQ ID NO: 23) shows 115% of the base editor activity compared to the wild-type base editor activity at position A8 of SEQ ID NO: 4, and pBE0022 (SEQ ID NO: 23) deletes amino acids E125-D127 from Cas9 nuclease. pBE0024 (SEQ ID NO: 24) shows 82% of the base editor activity compared to the wild-type base editor activity at position A8 of SEQ ID NO: 4, and pBE0024 (SEQ ID NO: 24) deletes amino acids E125-N130 from Cas9 nuclease. pBE0025 (SEQ ID NO: 25) exhibits 110% of the base editor activity at position A8 of SEQ ID NO: 4, and pBE0025 (SEQ ID NO: 25) deletes amino acids L132-K135 from the Cas9 nuclease. pBE0026 (SEQ ID NO: 26) exhibits 136% of the base editor activity at position A8 of SEQ ID NO: 4, and pBE0026 (SEQ ID NO: 26) deletes amino acids E136-S139 from the Cas9 nuclease. pBE0027 (SEQ ID NO: 57) exhibits 5% of the base editor activity at position A8 of SEQ ID NO: 4, and pBE0027 (SEQ ID NO: 57) deletes amino acids V164-R165 from the Cas9 nuclease. B) pBE0028 (SEQ ID NO: 58) shows 1% of the base editor activity at position A8 of SEQ ID NO: 4, and pBE0028 (SEQ ID NO: 58) deletes amino acids Q182-K185 from the Cas9 nuclease. Del3 (SEQ ID NO: 50) shows 99% of the base editor activity at position A8 of SEQ ID NO: 4, and Del3 (SEQ ID NO: 50) deletes amino acids E123-N130 from the Cas9 nuclease. Del4 (SEQ ID NO: 51) shows 25% of the base editor activity at position A8 of SEQ ID NO: 4, and Del4 (SEQ ID NO: 51) deletes amino acids R245-P425 from the Cas9 nuclease.Del5 (SEQ ID NO: 52) exhibits 53% of the base editor activity at position A8 of SEQ ID NO: 4, and Del5 (SEQ ID NO: 52) deletes amino acids T78-I86 from the Cas9 nuclease. C) pBE0038 (SEQ ID NO: 28) exhibits 92% of the base editor activity at position A8 of SEQ ID NO: 4, and pBE0038 (SEQ ID NO: 28) deletes amino acids E125-E146 from the Cas9 nuclease. pBE0039 (SEQ ID NO: 61) exhibits 0% of the base editor activity at position A8 of SEQ ID NO: 4, and pBE0039 (SEQ ID NO: 61) deletes amino acids E82-G95 from the Cas9 nuclease. [Figure 6]Figures 6A-C show graphs of the base activity of manipulated base editors as a percentage of wild-type base editor activity when amino acids are deleted from the RuvC region of Cas9 nuclease. A) pBE0029 (SEQ ID NO: 27) shows 96% of the base editor activity compared to the wild-type base editor activity at position A8 of SEQ ID NO: 4, and pBE0029 (SEQ ID NO: 27) deletes amino acids K719-Q731 from Cas9 nuclease. Del2 (SEQ ID NO: 59) shows 3% of the base editor activity compared to the wild-type base editor activity at position A8 of SEQ ID NO: 4, and Del2 (SEQ ID NO: 59) deletes amino acids F715-K770 from Cas9 nuclease. B) Del6 (SEQ ID NO: 53) exhibits 132% of the base editor activity at position A8 of SEQ ID NO: 4, and Del6 (SEQ ID NO: 53) deletes amino acids K689, K719, and K720 from the Cas9 nuclease. B) Del7 (SEQ ID NO: 54) exhibits 124% of the base editor activity at position A8 of SEQ ID NO: 4, and Del7 (SEQ ID NO: 54) deletes amino acids K719 and K720 from the Cas9 nuclease. Del8 (SEQ ID NO: 55) exhibits 116% of the base editor activity at position A8 of SEQ ID NO: 4, and Del8 (SEQ ID NO: 55) deletes amino acids W688, K689, K719, and K720 from the Cas9 nuclease. Del9 (SEQ ID NO: 56) exhibits 102% of the base editor activity at position A8 of SEQ ID NO: 4, and deletes amino acids W688, K689, and E717-L721 from the Cas9 nuclease. C) pBE0041 (SEQ ID NO: 62) exhibits 0% of the base editor activity at position A8 of SEQ ID NO: 4, and deletes amino acids Q737-E743 from the Cas9 nuclease. [Figure 7]Figure 7 is a graph of the base activity of engineered base editors as a percentage of wild-type base editor activity when amino acids are deleted from the PAM interaction domain of Cas9 nuclease. Del7 (SEQ ID NO: 54) shows 4% base editor activity compared to wild-type base editor activity at position A8 of SEQ ID NO: 4, and Del7 (SEQ ID NO: 54) deletes amino acids Y979 and R980 from Cas9 nuclease. [Figure 8] Figure 8 is a graph of the base activity of engineered base editors as a percentage of wild-type base editor activity when amino acids are deleted from the wedge domain of Cas9 nuclease. pBE0040 (SEQ ID NO: 29) shows 81% base editor activity compared to wild-type base editor activity at position A8 of SEQ ID NO: 4, and pBE0040 (SEQ ID NO: 29) deletes amino acids T866 - G847 from Cas9 nuclease. [Figure 9]Figures 9A-E show graphs of the base activity of manipulated base editors as a percentage of wild-type base editor activity when combinations of manipulations are performed on the base editor. A) pBE0060 (SEQ ID NO: 32) shows 117% base editor activity compared to the wild-type base editor activity at position A8 of SEQ ID NO: 4. pBE0060 (SEQ ID NO: 32) has the first and last four amino acids deleted from TadA deaminase and has a linker with the amino acid sequence of SEQ ID NO: 12. pBE0061 (SEQ ID NO: 33) shows 109% base editor activity compared to the wild-type base editor activity at position A8 of SEQ ID NO: 4. pBE0061 (SEQ ID NO: 33) has the first and last four amino acids deleted from TadA deaminase and has a linker with the amino acid sequence of SEQ ID NO: 12, and has amino acids E717-L721, W688, and K689 deleted from Cas9 nuclease. B) pBE0125 (SEQ ID NO: 47) exhibits 84% ​​of the base editor activity compared to the wild-type base editor activity at position A8 of SEQ ID NO: 4, and 121% of the base editor activity compared to the wild-type base editor activity at position A5 of SEQ ID NO: 4. pBE0125 (SEQ ID NO: 47) has the first 8 amino acids and the last 6 amino acids deleted from the TadA deaminase, has a linker with the amino acid sequence of SEQ ID NO: 11, and has the first 4 amino acids deleted from the Cas9 nuclease. C) pBE0123 (SEQ ID NO: 45) exhibits 80% of the base editor activity at position A5 of SEQ ID NO: 4 and 57% of the base editor activity at position A8 of SEQ ID NO: 4. pBE0123 (SEQ ID NO: 45) has the first 8 amino acids and the last 6 amino acids deleted from TadA deaminase and possesses a linker with the amino acid sequence of SEQ ID NO: 12, and the first 4 amino acids and amino acids K719-Q731 deleted from Cas9 nuclease.pBE0124 (SEQ ID NO: 46) exhibits 48% base editor activity compared to the wild-type base editor activity at position A5 of SEQ ID NO: 4, and 38% base editor activity compared to the wild-type base editor activity at position A8 of SEQ ID NO: 4. pBE0124 (SEQ ID NO: 46) lacks the first 8 amino acids and the last 6 amino acids from TadA deaminase, has a linker with the amino acid sequence of SEQ ID NO: 12, and lacks the first 4 amino acids from Cas9 nuclease as well as amino acids E717 - L721, W688, and K689. D) pBE0126 (SEQ ID NO: 48) exhibits 84% base editor activity compared to the wild-type base editor activity at position A5 of SEQ ID NO: 4. pBE0126 (SEQ ID NO: 48) lacks the first 8 amino acids and the last 6 amino acids from TadA deaminase, has a linker with the amino acid sequence of SEQ ID NO: 12, and lacks the first 4 amino acids from Cas9 nuclease as well as amino acids E125 - E146. E) pBE0127 (SEQ ID NO: 49) exhibits 38% base editor activity compared to the wild-type base editor activity at position A5 of SEQ ID NO: 4. pBE0127 (SEQ ID NO: 49) lacks the first 8 amino acids and the last 6 amino acids from TadA deaminase, has a linker with the amino acid sequence of SEQ ID NO: 12, and lacks the first 4 amino acids from Cas9 nuclease as well as amino acids E125 - E146 and K719 - Q731. pBE0129 (SEQ ID NO: 64) exhibits 0% base editor activity compared to the wild-type base editor activity at position A5 of SEQ ID NO: 4. pBE0129 (SEQ ID NO: 64) lacks the first 8 amino acids and the last 6 amino acids from TadA deaminase, has a linker with the amino acid sequence of SEQ ID NO: 11, and lacks the first 4 amino acids from Cas9 nuclease as well as amino acids E125 - E146, E717 - L721, and W688 - K689.

[0019] Incorporation by reference All publications, patents, and patent applications referenced herein are incorporated by reference to the same extent as if each individual publication, patent, or patent application were explicitly indicated separately as being incorporated by reference. It is intended that this specification supersedes and / or takes precedence over any such conflicting material to the extent that the incorporated publications and patents or patent applications contradict the disclosures contained in the specification. [Modes for carrying out the invention]

[0020] The following description and examples illustrate embodiments of the present disclosure in detail. It should be understood that the present disclosure is not limited to the specific embodiments described herein and is therefore subject to change. Those skilled in the art will recognize that there are numerous variations and changes to the present disclosure, which are included within the scope of the present disclosure.

[0021] All terms are intended to be understood as they would be understood by those skilled in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as they would be generally understood by those skilled in the art in which this disclosure is made.

[0022] The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described.

[0023] While various features of this disclosure may be described in the context of a single embodiment, the features may also be provided separately or in any suitable combination. Conversely, while this disclosure may be described herein in the context of separate embodiments for clarity, this disclosure may also be implemented in a single embodiment.

[0024] The following definitions supplement those in the technical field and apply to this application; they should not be attributed to any related or unrelated cases, such as any generally owned patents or applications. Any methods and materials similar to or equivalent to those described herein may be used in the practice for testing of this disclosure, but preferred materials and methods are described herein. Therefore, the terminology used herein is intended solely to describe specific embodiments and is not intended to be limiting.

[0025] In this application, the use of the singular includes the plural unless otherwise specified. It should be noted that, as used herein, the singular forms "a," "an," and "the" include plural references unless the context explicitly states otherwise.

[0026] In this application, the use of “or” means “and / or” unless otherwise specified. The terms “and / or” and “any combination thereof” and their grammatical equivalents as used herein can be used interchangeably. These terms can convey that any combination is specifically assumed. For illustrative purposes only, the following phrases “A, B, and / or C” or “A, B, C, or any combination thereof” can mean “A individually; B individually; C individually; A and B; B and C; A and C; as well as A, B, and C.” The term “or” can be used conjunctively or separately unless the context specifically refers to separate use.

[0027] Furthermore, the use of the term "including," as well as other forms such as "include," "includes," and "included," is not limited.

[0028] References to “some embodiments,” “an embodiment,” “one embodiment,” or “other embodiments” in this specification mean that certain features, structures, or characteristics described in relation to that embodiment are included in at least some embodiments of this disclosure, but not necessarily in all embodiments.

[0029] As used herein and in the claims, the words “comprising” (and any form of “comprising,” such as “comprise” and “comprises”), “having” (and any form of “having,” such as “have” and “has”), “including” (and any form of “including,” such as “includes” and “include”), or “containing” (and any form of “containing,” such as “contains” and “contain”) are inclusive or non-limiting and do not exclude additional unlisted elements or method steps. Any embodiment discussed herein can be carried out with respect to any method or composition of the Disclosure, and vice versa. Furthermore, the compositions of the Disclosure can be used to achieve the methods of the Disclosure.

[0030] The terms "about" or "approximately" mean within an acceptable margin of error for a particular value, as determined by those skilled in the art, which will depend in part on how that value is measured or determined, for example, on the limits of the measurement system.

[0031] As used herein, the terms “edit,” “editing,” or “edited” refer to a method of altering a polynucleotide nucleic acid sequence (e.g., a wild-type naturally occurring nucleic acid sequence or a mutated naturally occurring sequence) by selective modification (e.g., deletion, insertion, or substitution) of one or more nucleotides in a particular genomic target. Such particular genomic targets include, but are not limited to, chromosomal regions, genes, promoters, open reading frames, or any nucleic acid sequence.

[0032] As used herein, the terms “target” or “target site” refer to a pre-specified nucleic acid sequence of any composition and / or length. Such target sites include, but are not limited to, chromosomal regions, genes, promoters, open reading frames, or any nucleic acid sequence. In some embodiments, the present invention examines these specific genomic target sequences with complementary gRNA sequences.

[0033] As used herein, the term “effective dose” refers to a specific amount of a pharmaceutical composition containing a therapeutic agent that achieves a clinically beneficial outcome (i.e., a reduction in symptoms). The toxicity and therapeutic effect of such compositions can be determined, for example, by standard pharmaceutical procedures in cell cultures or experimental animals to determine the LD50 (50% lethal dose) and ED50 (50% therapeutic effective dose) of the population. The dose ratio between the toxic effect and the therapeutic effect is a therapeutic index, which can be expressed as the ratio LD50 / ED50. Compounds exhibiting a large therapeutic index are preferred. Data obtained from these cell culture assays and additional animal studies can be used to formulate a range of doses for human use. Doses of such compounds are preferably within the range of circulating concentrations containing an ED50 that is little to no toxicity. Doses vary within this range, depending on the dosage form used, patient sensitivity, and route of administration.

[0034] As used herein, the terms “pharmaceutically acceptable” or “pharmacologically acceptable” refer to molecular entities and compositions that, when administered to animals or humans, do not produce any harmful, allergic, or other adverse reactions.

[0035] The term “pharmaceutically acceptable carrier” as used herein includes water, ethanol, polyols (e.g., glycerol, propylene glycol, and liquid polyethylene glycol, and the like), suitable mixtures thereof, and any solvent or dispersion medium, including but not limited to vegetable oils, coatings, isotonic and absorption retardants, liposomes, commercial cleansers, and the like. Co-active ingredients may also be incorporated into such carriers.

[0036] The term “viral vector” encompasses any nucleic acid construct derived from a viral genome that can incorporate a heterologous nucleic acid sequence for expression in a host organism. Examples of such viral vectors include, but are not limited to, adeno-associated virus vectors, lentiviral vectors, SV40 virus vectors, retroviral vectors, or adenovirus vectors. Viral vectors are sometimes produced from pathogenic viruses, but can be modified to minimize their overall health risk. In some embodiments, this involves the deletion of a portion of the viral genome involved in viral replication. Such viruses can efficiently infect cells, but after infection, they may require a helper virus to provide the missing protein for the production of new virions. Preferably, viral vectors should have minimal effect on the physiological function of the cells they infect and exhibit genetically stable properties (e.g., not undergoing spontaneous genomic rearrangement). Most viral vectors are engineered to infect as wide a range of cell types as possible. Even so, viral receptors can be modified to target specific types of cells. Viruses modified in this manner are said to be pseudotyped. Viral vectors are often engineered to incorporate specific genes that help identify which cells have taken up the viral gene. These genes are called marker genes. For example, a common marker gene confers antibiotic resistance to a particular antibiotic.

[0037] As used herein, the term “Cas9” refers to a protein derived from the type II CRISPR system. In some embodiments, Cas9 is an enzyme specialized to produce double-strand breaks in DNA, having two active cleavage sites (HNH and RuvC domains), one for each strand of the double helix. In some embodiments, the Cas9 protein lacks the ability to cleave both strands of DNA; i.e., the Cas9 protein is a nickase. In some embodiments, the Cas9 protein lacks the ability to cleave either strand of DNA; i.e., the Cas9 protein is a “nuclease dead.” In some embodiments, to produce a nickase or nuclease dead Cas9, Cas9 has one or more inactivating mutations in the HNH nuclease domain (e.g., the mutation corresponding to H840A in SpCas9) and / or in the RuvC nuclease domain (e.g., the mutation corresponding to D10A in SpCas9). In some embodiments, Cas9 contains D at the amino acid position corresponding to position 9 of SEQ ID NO: 2. In some embodiments, Cas9 contains D at the amino acid position corresponding to position 9 of SEQ ID NO: 65. In some embodiments, Cas9 contains an amino acid other than D (e.g., A) at the amino acid position corresponding to position 9 of SEQ ID NO: 2. In some embodiments, Cas9 contains an amino acid other than D (e.g., A) at the amino acid position corresponding to position 9 of SEQ ID NO: 65. In some embodiments, Cas9 contains N at the amino acid position corresponding to position 579 of SEQ ID NO: 2. In some embodiments, Cas9 contains N at the amino acid position corresponding to position 581 of SEQ ID NO: 65. In some embodiments, Cas9 contains an amino acid other than N (e.g., A) at the amino acid position corresponding to position 579 of SEQ ID NO: 2. In some embodiments, Cas9 contains an amino acid other than N (e.g., A) at the amino acid position corresponding to position 581 of SEQ ID NO: 65. In preferred embodiments, any of the Cas9 proteins disclosed herein can bind to DNA (e.g., when complexed with sgRNA).In some embodiments, tracrRNA and spacer RNA combine to form a “single guide RNA” (sgRNA) molecule, which, when mixed with Cas9, can locate and cleave DNA targets through Watson-Crick pairing between the guide sequence and target DNA sequence within the sgRNA. Jinek et al., “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity” Science 337(6096):816-821 (2012). In some embodiments, the Cas9 protein may be fused to a heterologous moiety (e.g., TadA). In some embodiments, the heterologous moiety is a protein useful for gene silencing (e.g., the KRAB domain, Dnmt3A, and / or the Dnmt3L protein domain). In some embodiments, the heterologous moiety is a DNA methyltransferase, e.g., Dnmt1, Dnmt3A, or Dnmt3B. In some embodiments, the heterologous portion is a regulator of a DNA methyltransferase factor (e.g., a catalytically inactive regulator of DNA methyltransferase (e.g., Dnmt3L)). In some embodiments, the heterologous portion is a protein useful for reversing gene silencing (e.g., the TET1 DNA demethylase catalytic domain). In some embodiments, the heterologous portion is a transcriptional repressor. In some embodiments, the heterologous portion is a transcriptional activator.

[0038] As used herein, the terms “KRAB domain” or “Kruppel-associated box domain” refer to a category of transcriptional repression domains that are typically 45–75 amino acids long. See, for example, Ecco, G., Imbeault, M., Trono, D., KRAB zinc finger proteins, Development 144, 2017; Lambert et al. The human transcription factors, Cell 172, 2018. In some embodiments, the KRAB domain is the KRAB domain of Kox1. In some embodiments, the KRAB domain contains the sequence of Sequence ID No. 73. In some embodiments, the KRAB domain includes an amino acid sequence that is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 73:DAKSLTAWSRTLVTFKDVFVDFTREEWKLLDTAQQIVYRNVMLENYKNLVSLGYQLT KPDVILRLEKGEEP.

[0039] As used herein, the term "Dnmt3A" refers to the "DNA(cytosine-5)-methyltransferase 3A" or "DNA methyltransferase 3a" protein or a fragment thereof. In some embodiments, Dnmt3A includes the sequence of SEQ ID NO: 74. In some embodiments, Dnmt3A is: SEQ ID NO: 74:NHDQEFDPPKVYPPVPAEKRKPIRVLSLFDGIATGLLVLKDLGIQVDRYIASEVCEDSITVGMVRHQGKIMYVGDVRSVTQKHIQEWGPFDLVIGGSPCNDLSIVNPARKGLYEGTGRLFFEFYRLLHDARPKEGDDRPFFWLFENVVAMGVSDKRDISRFLESNPVMIDAKEVSAAHRARYFWGNL Contains an amino acid sequence that is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to PGMNRPLASTVNDKLELQECLEHGRIAKFSKVRTITTRSNSIKQGKDQHFPVFMNEKEDILWCTEMERVFGFPVHYTDVSNMSRLARQRLLGRSWSVPVIRHLFAPLKEYFACV.

[0040] As used herein, the term "Dnmt3L" refers to the "DNA(cytosine-5)-methyltransferase 3L" or "DNA methyltransferase 3L" protein or a fragment thereof. In some embodiments, Dnmt3L includes the sequence of SEQ ID NO: 75. In some embodiments, Dnmt3L contains an amino acid sequence that is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 75:MGPMEIYKTVSAWKRQPVRVLSLFRNIDKVLKSLGFLESGSGSGGGTLKYVEDVTNVVRRDVEKWGPFDLVYGSTQPLGSSCDRCPGWYMFQFHRILQYALPRQESQRPFFWIFMDNLLLTEDDQETTTRFLQTEAVTLQDVRGRDYQNAMRVWSNIPGLKSKHAPLTPKEEEYLQAQVR SRSKLDAPKVDLLVKNCLLPLREYFKYFSQNSLPL.

[0041] As used herein, the term “protospacer adjacent motif” (or PAM) refers to a DNA sequence that Cas9 / sgRNA may need to form an R-loop to examine a specific DNA sequence through Watson-Crick pairing with its guide RNA in the genome. PAM specificity may be a function of the DNA binding specificity of the Cas9 protein (e.g., the “protospacer adjacent motif recognition domain” at the C-terminus of Cas9).

[0042] The terms “guide RNA,” “gRNA,” and simply “guide” are used interchangeably herein and refer to crRNA (also known as CRISPR RNA), or a combination of crRNA and trRNA (also known as tracrRNA). crRNA and trRNA may associate as a single RNA molecule (single guide RNA, sgRNA) or as two separate RNA molecules (dual guide RNA, dgRNA). “Guide RNA” refers to each of these types. trRNA may be a naturally occurring sequence or a trRNA sequence that is modified or variant compared to a naturally occurring sequence. For clarity, as used herein and unless otherwise specified, the terms “guide RNA” or “guide” may refer to an RNA molecule (containing A, C, G, and U nucleotides) or a DNA molecule (containing A, C, G, and T nucleotides) that encodes such an RNA molecule, or a complementary sequence thereof. In general, for a DNA nucleic acid construct encoding a guide RNA, any U residue in the RNA sequence described herein may be replaced with a T residue, and for a guide RNA construct encoded by any of the DNA sequences described herein, a T residue may be replaced with a U residue.

[0043] As used herein, the term “sgRNA” refers to a single guide RNA used in conjunction with CRISPR-associated systems (Cas). An sgRNA is a fusion of crRNA and tracrRNA containing nucleotides in a sequence complementary to the desired target site. Watson-crick pairing of the sgRNA with the target site enables R-loop formation, which in conjunction with functional PAM allows for DNA cleavage or, in the case of nuclease-deficient Cas9, binding to the DNA at its locus.

[0044] The term “spacer sequence,” which may also be referred to herein and in the literature as “spacer,” “protospacer,” “guide sequence,” or “target sequence,” as used herein, refers to a sequence within guide RNA that is complementary to the target sequence and functions to direct the guide RNA towards the target sequence for Cas9 cleavage. For clarity, as used herein and unless otherwise specified, the terms “spacer sequence,” “spacer,” “protospacer,” “guide sequence,” or “target sequence” may refer to an RNA molecule (including A, C, G, and U nucleotides) or a DNA molecule (including A, C, G, and T nucleotides) encoding such an RNA molecule or its complementary sequence. The guide sequence may have a base pair length of 24, 23, 22, 21, 20, or fewer, for example, Staphylococcus lugdunensis (i.e., SluCas9) or Staphylococcus aureus (i.e., SaCas9) and associated Cas9 homologs / orthologues. In preferred embodiments, the guide / spacer sequence for SluCas9 or SaCas9 is at least 20 base pairs long, or more specifically, within 20–25 base pairs long (see, e.g., Schmidt et al., 2021, Nature Communications, “Improved CRISPR genome editing using small highly active and specific engineered RNA-guided nucleases”). Shorter or longer sequences, such as 15-, 16-, 17-, 18-, 19-, 20-, 21-, 22-, 23-, 24-, or 25-nucleotide lengths, can also be used as guides.

[0045] As used herein, the term “fluorescent protein” refers to a protein domain containing at least one organic compound moiety that emits fluorescence (fluorescent light) in response to a suitable wavelength. For example, fluorescent proteins may emit red, blue, and / or green light. Such proteins are readily available commercially and include, but are not limited to, i) mCherry (Clonetech Laboratories): excitation: 556 / 20 nm (wavelength / bandwidth); emission: 630 / 91 nm; ii) sfGFP (Invitrogen): excitation: 470 / 28 nm; emission: 512 / 23 nm; iii) TagBFP (Evrogen): excitation: 387 / 11 nm; emission: 464 / 23 nm.

[0046] As used herein, the term “orthogonal” refers to non-overlapping, uncorrelated, or independent targets. For example, if two orthogonal nuclease-deficient Cas9 genes fused to different effector domains are present, the sgRNAs encoded by each are considered to have neither crosstalk nor overlap. Not all nuclease-deficient Cas9 genes function the same way, and for this reason, the use of orthogonal nuclease-deficient Cas9 genes fused to different effector domains is possible if appropriate orthogonal sgRNAs are provided.

[0047] As used herein, the term “phenotypic change” or “phenotype” refers to a combination of observable features or traits of an organism, such as its morphology, growth, biochemical or physiological characteristics, phenomenology, behavior, and behavioral outcomes. The phenotype arises from the expression of an organism's genes and the influence of environmental factors and the interaction between the two.

[0048] As used herein, “nucleic acid sequence,” “polynucleotide sequence,” and “nucleotide sequence” refer to oligonucleotides or polynucleotides, and fragments or portions thereof, as well as DNA or RNA of genomic or synthetic origin, which may be single-stranded or double-stranded, and which may represent a sense strand or an antisense strand.

[0049] The term “isolated nucleic acid,” as used herein, refers to any nucleic acid molecule that has been removed from its natural state (e.g., removed from a cell, and in a preferred embodiment, free from other genomic nucleic acids).

[0050] As used herein, the terms "amino acid sequence" and "polypeptide sequence" refer to a series of amino acids that are interconvertible.

[0051] As used herein, the term “part” (as in “a given part of a protein”), when relating to a protein, refers to a fragment of that protein. Unless otherwise specified, fragments may range in size from four amino acid residues to the entire amino acid sequence minus one amino acid.

[0052] When used in relation to a nucleotide sequence, the term "part" refers to a fragment of that nucleotide sequence. Unless otherwise specified, fragments may range in size from 5 nucleotide residues to the entire nucleotide sequence minus 1 nucleic acid residue.

[0053] As used herein, the terms “complementary” or “complementarity” are used in relation to “polynucleotides” and “oligonucleotides” (these are interconvertible terms referring to a series of nucleotides) in relation to base-pairing rules. For example, the sequence “CAGT” is complementary to the sequence “GTCA”. Complementarity can be “partial” or “complete”. “Partial” complementarity is when one or more nucleic acid bases do not match according to base-pairing rules. “Complete” or “complete” complementarity between nucleic acids is when, under base-pairing rules, every possible nucleic acid base matches with another base. The degree of complementarity between nucleic acid strands has a significant effect on the efficiency and strength of hybridization between nucleic acid strands. This is particularly important in amplification reactions and detection methods that rely on binding between nucleic acids.

[0054] As used herein in relation to nucleotide sequences, the terms “homologous” and “homogenetic” refer to the degree of complementarity with other nucleotide sequences. Partial homology or complete homology (i.e., identity) may exist. A nucleotide sequence that is partially complementary to a nucleic acid sequence, i.e., “substantially homologous,” is a nucleotide sequence that at least partially inhibits the hybridization of a fully complementary sequence to the target nucleic acid sequence. The inhibition of hybridization of a fully complementary sequence to a target sequence may be investigated using hybridization assays (Southern or Northern blotting, solution hybridization, and similar methods) under low-strictness conditions. A substantially homologous sequence or probe competes for and inhibits the binding (i.e., hybridization) of a fully homologous sequence to the target sequence under low-strictness conditions. This does not mean that low-strictness conditions allow for nonspecific binding; rather, low-strictness conditions require that the binding of the two sequences to each other be a specific (i.e., selective) interaction. The absence of nonspecific binding can also be tested by using a second target sequence that lacks even partial complementarity (e.g., less than approximately 30% identity); if nonspecific binding is absent, the probe will not hybridize with the second non-complementary target.

[0055] As used herein in relation to amino acid sequences, the terms “homology” and “homological” refer to the degree of primary structural identity between two amino acid sequences. Such degree of identity may apply to a portion of each amino acid sequence or to the entire amino acid sequence. Two or more amino acid sequences that are “substantially homologous” may have at least 75% identity, at least 85% identity, at least 95%, or 100% identity.

[0056] An oligonucleotide sequence that is a "homologous" is defined herein as an oligonucleotide sequence that exhibits more than 50% or equal identity to a sequence when compared to sequences of 100 bp or longer.

[0057] As used herein, the term “hybridization” is used in relation to the pairing of complementary nucleic acids using any method by which a strand of nucleic acid binds to a complementary strand through base pairing to form a hybridization complex. Hybridization and the intensity of hybridization (i.e., the intensity of association between nucleic acids) are influenced by factors such as the degree of complementarity between the nucleic acids, the strictness of the conditions involved, the melting temperature of the hybrid formed, and the G to C ratio within the nucleic acid.

[0058] As used herein, the term “hybridization complex” refers to a complex formed between two nucleic acid sequences due to the formation of hydrogen bonds between complementary G and C bases and between complementary A and T (or U) bases; these hydrogen bonds may be further stabilized by base stacking interactions. The two complementary nucleic acid sequences hydrogen-bond in an antiparallel structure. Hybridization complexes may be formed in solution (e.g., Cot or Rot analysis) or between one nucleic acid sequence present in solution and another nucleic acid sequence immobilized on a solid support (e.g., nylon membranes or nitrocellulose filters used in Southern and Northern blotting, dot blotting, or glass slides used in insight hybridization, including FISH (fluorescent insight hybridization)).

[0059] DNA and RNA molecules are said to have a "5' end" and a "3' end." This is because mononucleotides undergo reactions that produce oligonucleotides in which the 5' phosphate of one mononucleotide pentose ring is bonded in one direction to its adjacent 3' oxygen via a phosphate diester bond. Thus, the end of an oligonucleotide is called the "5' end" unless its 5' phosphate is linked to the 3' oxygen of another mononucleotide pentose ring. The end of an oligonucleotide is called the "3' end" unless its 3' oxygen is linked to the 5' phosphate of another mononucleotide pentose ring. As used herein, nucleic acid sequences can also be said to have 5' and 3' ends, even if they are inside a larger oligonucleotide. In linear or circular DNA or RNA molecules, separate elements are called the "upstream" or 5' element or the "downstream" or 3' element. This terminology reflects the fact that transcription proceeds along the DNA or RNA strand in a 5' to 3' manner. Promoter and enhancer elements that direct the transcription of linked genes are generally located 5' or upstream of the coding region. However, enhancer elements can exert their effect even when located at the 3' of the promoter element and coding region. Transcription termination and polyadenylation signals are located at the 3' or downstream of the coding region.

[0060] The term "transfection" or "transfected" refers to the introduction of foreign DNA or RNA into a cell.

[0061] As used herein, the terms “nucleic acid molecular encoding,” “DNA (or RNA) sequence encoding,” and “DNA (or RNA) encoding” refer to the order or sequence of nucleotides along a nucleic acid chain. The order of these deoxyribonucleotides determines the order of amino acids along a polypeptide (protein) chain. A DNA sequence thus encodes an amino acid sequence.

[0062] In addition to containing introns, the genomic type of a gene may also include sequences located at both the 5' and 3' ends of the sequence present in the RNA transcript. These sequences are called “flanking” sequences or regions (these flanking sequences are located at either the 5' or 3' end of the untranslated sequence present on the mRNA transcript). The 5' flanking region may contain regulatory sequences such as promoters and enhancers that control or influence gene transcription. The 3' flanking region may contain sequences that direct transcription termination, post-transcriptional cleavage, and polyadenylation.

[0063] The terms “patient,” “subject,” and “individual” may be used interchangeably and refer to a human or a non-human animal. “Non-human animal” and “non-human mammal,” as used interchangeably herein, include mammals such as rats, mice, rabbits, sheep, cats, dogs, cattle, pigs, and non-human primates. The term “subject” encompasses any vertebrate, including but not limited to mammals, reptiles, amphibians, and fish. However, advantageously, subject is a mammal such as a human, or another mammal such as a domesticated mammal, e.g., a dog, cat, horse, and its relatives, or a production mammal, e.g., a cattle, sheep, pig, and its relatives. “Patient in need of it” or “subject in need of it” is referred to as a patient diagnosed with or suspected of having a disease or disorder, e.g., diabetes, but not limited to it.

[0064] As used herein, “administer” may mean providing one or more compositions described herein to a patient or subject. By example and without limitation, administration of a composition, e.g., by injection, may be carried out by intravenous (IV), subcutaneous (SC), intradermal (ID), intraperitoneal (IP), or intramuscular (IM) injection. One or more such routes may be used. Parenteral administration may be, for example, by bolus injection or by time-based stepwise perfusion. Alternatively, or simultaneously, administration may be by oral route. Furthermore, administration may also be by surgical deposition of a bolus or cell pellet, or by positioning of a medical device. In embodiments, the compositions of this disclosure may contain engineered cells or host cells expressing a nucleic acid sequence described herein, or a vector containing at least one nucleic acid sequence described herein, in an amount effective to treat or prevent a proliferative disorder. Pharmaceutical compositions may contain a cell population described herein in combination with one or more pharmaceutically or physiologically acceptable carriers, diluents, or excipients. Such compositions may include buffers such as neutral buffered saline, phosphate-buffered saline and the like; carbohydrates such as glucose, mannose, sacrose or dextran, mannitol; proteins; amino acids such as polypeptides or glycine; antioxidants; chelating agents such as EDTA or glutathione; adjuvants (e.g., aluminum hydroxide); and preservatives.

[0065] As used herein, the terms “genetically engineered” and its grammatical equivalent refer to genetic modifications that do not exist in nature. Examples of genetic engineering include the use of gene editing systems such as CRISPR / Cas, TALEN, and / or zinc finger systems to interfere with the expression of one or more gene targets in a cell (e.g., to reduce or eliminate their expression) or to increase their expression in a cell (e.g., by inserting the gene of interest). “Genetically engineered” cells, as used herein, mean cells that have been genetically modified, or cells derived from and / or lineages of genetically engineered cells.

[0066] This disclosure envisions a truncated protein lacking one or more amino acids. In some embodiments, the truncated protein lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58 ,59,60,61,62,63,64,65,66,67,67,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102, Lacking amino acids 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, or 130 (e.g., consecutive amino acids). In some embodiments, the truncated protein lacks 1-130, 1-120, 1-110, 1-100, 1-90, 1-80, 1-70, 1-60, 1-50, 1-40, 1-30, 1-20, 1-15, 1-10, 1-5, 1-40, 1-30, 1-20, 1-15, 1-10, 1-5, 1-3, 100-130, 100-110, 90-110, 80-100, 60-80, 40-60, 20-40, 10-70, 10-50, 10-30, 10-20, 10-15, 5-10, 5-20, or 20-30 amino acids compared to the reference sequence (e.g., SEQ ID NO: 2) or an enumerated portion of the reference sequence (e.g., the truncated protein lacks 90-110 amino acids corresponding to amino acids 505-622 of SEQ ID NO: 2). In some embodiments, the truncated protein lacks any consecutive amino acids in any of the sequences disclosed herein. In some embodiments, the truncated protein lacks any discontinuous amino acids in any of the sequences disclosed herein.

[0067] Compact protein composition In a first embodiment, the disclosure relates to a compact protein comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, wherein the compact protein has an amino acid length between 1175 and 1249. In some embodiments, the compact protein has an amino acid length between 1175 and 1249. In some embodiments, the compact protein has an amino acid length between 1180 and 1249. In some embodiments, the compact protein has an amino acid length between 1190 and 1249. In some embodiments, the compact protein has an amino acid length between 1200 and 1249. In some embodiments, the compact protein has an amino acid length between 1210 and 1249. In some embodiments, the compact protein has an amino acid length between 1220 and 1249. In some embodiments, the compact protein has an amino acid length between 1230 and 1249. In some embodiments, the compact protein has an amino acid length between 1240 and 1249. In some embodiments, the compact protein has an amino acid length between 1180 and 1200. In some embodiments, the compact protein has an amino acid length between 1180 and 1190. In some embodiments, the compact protein has an amino acid length between 1200 and 1249. In some embodiments, the compact protein has an amino acid length between 1200 and 1240. In some embodiments, the compact protein has an amino acid length between 1200 and 1230. In some embodiments, the compact protein has an amino acid length between 1200 and 1220. In some embodiments, the compact protein has an amino acid length between 1200 and 1210. In some embodiments, the compact protein has an amino acid length between 1210 and 1249. In some embodiments, the compact protein has an amino acid length between 1210 and 1240. In some embodiments, the compact protein has an amino acid length between 1210 and 1230. In some embodiments, the compact protein has an amino acid length between 1210 and 1220. In some embodiments, compact proteins have an amino acid length between 1220 and 1249.In some embodiments, the compact protein has an amino acid length between 1220 and 1240. In some embodiments, the compact protein has an amino acid length between 1220 and 1230. In some embodiments, the compact protein has an amino acid length between 1230 and 1249. In some embodiments, the compact protein has an amino acid length between 1230 and 1240. In some embodiments, the compact protein has an amino acid length between 1240 and 1249.

[0068] In certain embodiments of this disclosure, the compact protein is a nucleic acid base editor comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease. In some embodiments, the nucleic acid base editor has an amino acid length between 1175 and 1249. In some embodiments, the nucleic acid base editor has an amino acid length between 1175 and 1249. In some embodiments, the nucleic acid base editor has an amino acid length between 1180 and 1249. In some embodiments, the nucleic acid base editor has an amino acid length between 1190 and 1249. In some embodiments, the nucleic acid base editor has an amino acid length between 1200 and 1249. In some embodiments, the nucleic acid base editor has an amino acid length between 1210 and 1249. In some embodiments, the nucleic acid base editor has an amino acid length between 1220 and 1249. In some embodiments, the nucleic acid base editor has an amino acid length between 1230 and 1249. In some embodiments, the nucleic acid base editor has an amino acid length between 1240 and 1249. In some embodiments, the nucleic acid base editor has an amino acid length between 1180 and 1200. In some embodiments, the nucleic acid base editor has an amino acid length between 1180 and 1190. In some embodiments, the nucleic acid base editor has an amino acid length between 1200 and 1249. In some embodiments, the nucleic acid base editor has an amino acid length between 1200 and 1240. In some embodiments, the nucleic acid base editor has an amino acid length between 1200 and 1230. In some embodiments, the nucleic acid base editor has an amino acid length between 1200 and 1220. In some embodiments, the nucleic acid base editor has an amino acid length between 1200 and 1210. In some embodiments, the nucleic acid base editor has an amino acid length between 1210 and 1249. In some embodiments, the nucleic acid base editor has an amino acid length between 1210 and 1240. In some embodiments, the nucleic acid base editor has an amino acid length between 1210 and 1230. In some embodiments, the nucleic acid base editor has an amino acid length between 1210 and 1220. In some embodiments, the nucleic acid base editor has an amino acid length between 1220 and 1249. In some embodiments, the nucleic acid base editor has an amino acid length between 1220 and 1240.In some embodiments, the nucleic acid base editor has an amino acid length between 1220 and 1230. In some embodiments, the nucleic acid base editor has an amino acid length between 1230 and 1249. In some embodiments, the nucleic acid base editor has an amino acid length between 1230 and 1240. In some embodiments, the nucleic acid base editor has an amino acid length between 1240 and 1249.

[0069] Nucleotide deaminase In some embodiments, the nucleotide deaminase is 145-166 amino acid long. In some embodiments, the nucleotide deaminase is 145-162 amino acid long. In some embodiments, the nucleotide deaminase is 145-158 amino acid long. In some embodiments, the nucleotide deaminase is 145-154 amino acid long. In some embodiments, the nucleotide deaminase is 145-151 amino acid long. In some embodiments, the nucleotide deaminase is 150-166 amino acid long. In some embodiments, the nucleotide deaminase is 150-162 amino acid long. In some embodiments, the nucleotide deaminase is 150-156 amino acid long. In some embodiments, the nucleotide deaminase is 150-153 amino acid long. In some embodiments, the nucleotide deaminase is 152-166 amino acid long. In some embodiments, the nucleotide deaminase is 152-158 amino acid long. In some embodiments, the nucleotide deaminase is 152-154 amino acid long. In some embodiments, the nucleotide deaminase is 157-166 amino acids long. In some embodiments, the nucleotide deaminase is 157-162 amino acids long. In some embodiments, the nucleotide deaminase is 161-166 amino acids long. In some embodiments, the nucleotide deaminase is 161-163 amino acids long.

[0070] In some embodiments, any of the truncated nucleotide deaminases disclosed herein can deaminate adenine. In some embodiments, any of the nucleotide deaminases can deaminate cytosine. In some embodiments, truncated deaminase proteins can deaminate adenine at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% more efficiently than a reference non-truncated deaminase protein (e.g., a deaminase protein containing the amino acid sequence of SEQ ID NO: 1) under the same conditions. In some embodiments, the truncated deaminase protein can deaminate adenine 10-100%, 10-80%, 10-60%, 10-40%, 10-20%, 30-100%, 30-80%, 30-60%, 30-40%, 50-100%, 50-80%, 50-60%, 70-100%, 70-80%, 80-90%, 90-100%, or 85-95% more efficiently than a reference untruncated deaminase protein (e.g., a deaminase protein containing the amino acid sequence of SEQ ID NO: 1) for the same target sequence under the same conditions. In some embodiments, a truncated deaminase protein can deaminate cytosine at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% more efficiently than a reference untruncated deaminase protein (e.g., a deaminase protein containing the amino acid sequence of SEQ ID NO: 1) under the same conditions. In some embodiments, the truncated deaminase protein can deaminate cytosine 10-100%, 10-80%, 10-60%, 10-40%, 10-20%, 30-100%, 30-80%, 30-60%, 30-40%, 50-100%, 50-80%, 50-60%, 70-100%, 70-80%, 80-90%, 90-100%, or 85-95% more efficiently than a reference untruncated deaminase protein (e.g., a deaminase protein containing the amino acid sequence of SEQ ID NO: 1) for the same target sequence under the same conditions.

[0071] In some embodiments, the nucleotide deaminase is adenine deaminase. In some embodiments, the adenine deaminase is TadA deaminase. In some embodiments, the deaminase is either APOBEC1 or APOBEC1 deaminase disclosed in U.S. Patent Application Publication 2017 / 0121693. The aforementioned patent document is incorporated herein by reference in its entirety. In some embodiments, the TadA deaminase is TadA-8e. In some embodiments, the nucleotide deaminase comprises an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 1, or a functional fragment of SEQ ID NO: 1. In some embodiments, the nucleotide deaminase lacks the amino acid corresponding to the first amino acid of SEQ ID NO: 1. In some embodiments, the nucleotide deaminase lacks the amino acids corresponding to the first two amino acids of SEQ ID NO: 1. In some embodiments, the nucleotide deaminase lacks the amino acids corresponding to the first three amino acids of SEQ ID NO: 1. In some embodiments, the nucleotide deaminase lacks the amino acids corresponding to the first four amino acids of SEQ ID NO: 1. In some embodiments, the nucleotide deaminase lacks the amino acids corresponding to the first five amino acids of SEQ ID NO: 1. In some embodiments, the nucleotide deaminase lacks the amino acids corresponding to the first six amino acids of SEQ ID NO: 1. In some embodiments, the nucleotide deaminase lacks the amino acids corresponding to the first seven amino acids of SEQ ID NO: 1. In some embodiments, the nucleotide deaminase lacks the amino acids corresponding to the first eight amino acids of SEQ ID NO: 1. In some embodiments, the nucleotide deaminase lacks the amino acids corresponding to the last 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids of SEQ ID NO: 1. In some embodiments, the nucleotide deaminase lacks the amino acids corresponding to the first six amino acids of SEQ ID NO: 1. SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN also lacks the amino acids corresponding to the last eight amino acids.

[0072] In some embodiments, the nucleotide deaminase is a cytosine deaminase. In some embodiments, the nucleotide deaminase is one of the modified deaminases disclosed in Chen et al., 2022, Nature Biotechnology, 41:663-672. In some embodiments, the nucleotide deaminase contains an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 1, or the functional fragment of SEQ ID NO: 1, except that the amino acid corresponding to the 45th amino acid of SEQ ID NO: 1 is not N. In some embodiments, the nucleotide deaminase contains an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 1, or the functional fragment of SEQ ID NO: 1, wherein the amino acid corresponding to the 45th amino acid of SEQ ID NO: 1 is L.

[0073] Linker between deaminase and nuclease A linker exists between the deaminase and the nuclease. This linker connects the deaminase to the nuclease of the nucleic acid base editor described herein. In some embodiments, the linker between the deaminase and the nuclease is 2 to 10 amino acids long. In some embodiments, the linker between the deaminase and the nuclease is 3 to 7 amino acids long. In some embodiments, the linker between the deaminase and the nuclease is 3 to 5 amino acids long. In some embodiments, the linker between the deaminase and the nuclease is 3 amino acids long. In some embodiments, the linker between the deaminase and the nuclease contains one or more prolines and one or more alanines. In some embodiments, the amino acid sequence of the linker between the deaminase and the nuclease is SEQ ID NO: 11:PAP. In some embodiments, the amino acid sequence of the linker between the deaminase and the nuclease is SEQ ID NO: 12:PAPAP. In some embodiments, the amino acid sequence of the linker between the deaminase and the nuclease is SEQ ID NO: 13: PAPAPAP. In some embodiments, the linker contains glycine. In some embodiments, the linker contains glycine residues 1, 2, 3, 4, 5, 6, 7, 8, or 9. In some embodiments, the linker contains serine. In some embodiments, the linker contains serine residues 1, 2, 3, 4, 5, 6, 7, 8, or 9. In some embodiments, the linker contains one or more proline residues. In some embodiments, the linker contains one or more alanine residues. In some embodiments, the linker contains one or more threonine residues. In some embodiments, the linker contains an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 78 (SGGSSGGSSGSETPGTSESATPESSGGSSGGS). In some embodiments, the linker includes the amino acid sequence of SEQ ID NO: 79 (GGSS), SEQ ID NO: 80 (GGSSGG), or SEQ ID NO: 81 (SGGSSGGS).

[0074] Cas nuclease In some embodiments, the disclosure provides a truncated Cas protein, e.g., a truncated Cas9 protein, compared to a reference Cas protein sequence (e.g., wild-type SaCas9 or SluCas9 protein). In some embodiments, the Cas9 nuclease of the nucleic acid base editor has an amino acid length between 10¹⁰ and 10⁵². In some embodiments, the Cas9 nuclease has an amino acid length between 10¹⁰ and 10⁴². In some embodiments, the Cas9 nuclease has an amino acid length between 10¹⁰ and 10⁵². In some embodiments, the Cas9 nuclease has an amino acid length between 10²⁰ and 10⁵². In some embodiments, the Cas9 nuclease has an amino acid length between 10²⁰ and 10⁴². In some embodiments, the Cas9 nuclease has an amino acid length between 10²⁰ and 10⁵². In some embodiments, the Cas9 nuclease has an amino acid length between 1030 and 1052. In some embodiments, the Cas9 nuclease has an amino acid length between 1030 and 1040. In some embodiments, the Cas9 nuclease has an amino acid length between 1030 and 1035. In some embodiments, the Cas9 nuclease has an amino acid length between 1035 and 1046. In some embodiments, the Cas9 nuclease has an amino acid length between 1040 and 1052. In some embodiments, the Cas9 nuclease has an amino acid length between 1040 and 1046. In some embodiments, the Cas9 nuclease has an amino acid length between 1045 and 1052. In some embodiments, the Cas9 nuclease has an amino acid length between 1045 and 1048.

[0075] In some embodiments, the Cas (e.g., Cas9) nuclease is Cas (e.g., Cas9) nickasase. In some embodiments, the Cas9 nickasase is nSaCas9. In some embodiments, the Cas9 nuclease has the amino acid sequence of Sequence ID No. 2:

[0076] In some embodiments, the Cas protein is the SluCas9 protein. In some embodiments, SluCas9 is the sequence of SEQ ID NO: 65:

[0077] In some embodiments, the Cas protein is one of the engineered Cas proteins disclosed in Schmidt et al., 2021, Nature Communications, "Improved CRISPR genome editing using small highly active and specific engineered RNA-guided nucleases."

[0078] In some embodiments, Cas9 is the sequence of sequence number 66 (referred to herein as sRGN1):

[0079] In some embodiments, Cas9 is the sequence of sequence number 67 (referred to herein as sRGN2):

[0080] In some embodiments, Cas9 is the sequence of sequence number 68 (referred to herein as sRGN3):

[0081] In some embodiments, Cas9 is the sequence of sequence number 69 (referred to herein as sRGN3.1):

[0082] In some embodiments, Cas9 is the sequence of sequence number 70 (referred to herein as sRGN3.2):

[0083] In some embodiments, Cas9 is the sequence of sequence number 71 (referred to herein as sRGN3.3):

[0084] In some embodiments, Cas9 is the sequence of sequence number 72 (referred to herein as sRGN4):

[0085] In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 1-4, 1-6, 77-85, 124-126, 124-129, 124-143, 124-145, 131-134, 135-138, 122-129, 505-622, 687-688, 716-720, 718-719, 718-730, 765, and / or 865-873 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks amino acids corresponding to amino acids 1-4, 1-6, 77-85, 124-126, 124-129, 124-143, 124-145, 131-134, 135-138, 122-129, 505-622, 687-688, 716-720, 718-719, 718-730, 765, and / or 865-873 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 1-4 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 1-6 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 77-85 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 124-126 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 124-129 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 124-143 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 124-145 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 131-134 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 135-138 of SEQ ID NO: 2.In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 122-129 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 505-622 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 687-688 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 716-720 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 718-719 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 718-730 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks the amino acid corresponding to amino acid 765 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 865-873 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks amino acids corresponding to amino acids 718-719 and 765 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks amino acids corresponding to amino acids 718-719 and 687-688 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks amino acids corresponding to amino acids 716-720 and 687-688 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks amino acids corresponding to amino acids 1-4 and 718-730 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks amino acids corresponding to amino acids 1-4, 687-688 and 716-720 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks amino acids corresponding to amino acids 1-4 and 124-145 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks amino acids corresponding to amino acids 1-4, 505-622, and 718-730 of SEQ ID NO: 2.In some embodiments, the Cas9 nuclease is no exception to one or more amino acids from the range of amino acids corresponding to: SEQ ID NO: 2-1 to 2-8, SEQ ID NO: 2-1 to 2-10, SEQ ID NO: 2-81 to 2-94, SEQ ID NO: 2-163 to 2-164, SEQ ID NO: 2-181 to 2-184, SEQ ID NO: 2-244 to 2-424, SEQ ID NO: 2-714 to 2-769, SEQ ID NO: 2-736 to 742, SEQ ID NO: 2-978 to 979, and / or SEQ ID NO: 2-1039 to 1052. In some embodiments, the Cas9 nuclease is not lacking in the range of amino acids corresponding to: SEQ ID NO: 2-1 to 2-8, SEQ ID NO: 2-1 to 2-10, SEQ ID NO: 2-81 to 2-94, SEQ ID NO: 2-163 to 2-164, SEQ ID NO: 2-181 to 2-184, SEQ ID NO: 2-244 to 2-424, SEQ ID NO: 2-714 to 2-769, SEQ ID NO: 2-736 to 2-742, SEQ ID NO: 2-978 to 2-979, and / or SEQ ID NO: 2-1039 to 1052.

[0086] In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 81-94, 123-145, 244-424, 714-769, and / or 717-730 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks amino acids corresponding to amino acids 81-94, 123-145, 244-424, 714-769, and / or 717-730 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 81-94 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 123-145 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 244-424 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 714-769 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 717-730 of SEQ ID NO: 2.

[0087] In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 1-4, 82-85, 122-130, 124-147, 182-185, 720-722, 723-727, 723-733, 731-734, 977-980 and / or 1049-1053 of SEQ ID NO: 65. In some embodiments, the Cas9 nuclease lacks amino acids corresponding to amino acids 1-4, 82-85, 122-130, 124-147, 182-185, 720-722, 723-727, 723-733, 731-734, 977-980 and / or 1049-1053 of SEQ ID NO: 2. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 1-4 of SEQ ID NO: 65. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 122-130 of SEQ ID NO: 65. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 124-147 of SEQ ID NO: 65. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 182-185 of SEQ ID NO: 65. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 720-722 of SEQ ID NO: 65. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 723-733 of SEQ ID NO: 65. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 731-734 of SEQ ID NO: 65. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 977-980 of SEQ ID NO: 65. In some embodiments, the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 1049-1053 of SEQ ID NO: 65.

[0088] In some embodiments, a Cas protein containing an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of the sequences of SEQ ID NO: 2 or 65-72 lacks one or more amino acids from the range of amino acids corresponding to (or homologous to) amino acids 1-4, 1-6, 77-85, 124-126, 124-129, 124-143, 124-145, 131-134, 135-138, 122-129, 505-622, 687-688, 716-720, 718-719, 718-730, 765, and / or 865-873 of SEQ ID NO: 2. In some embodiments, a Cas protein containing an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of the sequences of SEQ ID NO: 2 or 65-72 lacks amino acids corresponding to (or homologous to) amino acids 1-4, 1-6, 77-85, 124-126, 124-129, 124-143, 124-145, 131-134, 135-138, 122-129, 505-622, 687-688, 716-720, 718-719, 718-730, 765, and / or 865-873 of SEQ ID NO: 2. In some embodiments, a Cas protein containing an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of sequence 2, SEQ ID NO: 2, 718-719, and 765, lacks amino acids corresponding to (or homologous to) amino acids 718-719 and 765 of SEQ ID NO: 2. In some embodiments, a Cas protein containing an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of sequence 2, SEQ ID NO: 2, 718-719, and 687-688, lacks amino acids corresponding to (or homologous to) amino acids 718-719 and 687-688 of SEQ ID NO: 2.In some embodiments, a Cas protein containing an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of sequence 2 or 65-72 lacks amino acids corresponding to (or homologous to) amino acids 716-720 and 687-688 of sequence 2. In some embodiments, a Cas protein containing an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of sequence 2 or 65-72 lacks amino acids corresponding to (or homologous to) amino acids 1-4 and 718-730 of sequence 2. In some embodiments, a Cas protein containing an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of sequence 2 or 65-72 lacks amino acids corresponding to (or homologous to) amino acids 1-4, 687-688, and 716-720 of sequence 2. In some embodiments, a Cas protein containing an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of sequence 2 or 65-72 lacks amino acids corresponding to (or homologous to) amino acids 1-4 and 124-145 of sequence 2. In some embodiments, a Cas protein containing an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of sequence 65-72 of SEQ ID NO: 2 lacks amino acids corresponding to (or homologous to) amino acids 1-4, 505-622, and 718-730 of SEQ ID NO: 2.In some embodiments, a Cas protein containing an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of the sequences of SEQ ID NO: 2 or 65-72 is not lacking one or more amino acids from the range of amino acids corresponding to (or homologous to) SEQ ID NO: 2-1-8, 2-1-10, 2-81-94, 2-163-164, 2-181-184, 2-244-424, 2-714-769, 2-736-742, 2-978-979, and / or 2-1039-1052. In some embodiments, a Cas protein containing an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of the sequences of SEQ ID NO: 2, 65-72, is not lacking an amino acid range corresponding to (or homologous to) amino acids: 1-8 of SEQ ID NO: 2, 1-10 of SEQ ID NO: 2, 81-94 of SEQ ID NO: 2, 163-164 of SEQ ID NO: 2, 181-184 of SEQ ID NO: 2, 244-424 of SEQ ID NO: 2, 714-769 of SEQ ID NO: 2, 736-742 of SEQ ID NO: 2, 978-979 of SEQ ID NO: 2, and / or 1039-1052 of SEQ ID NO: 2.

[0089] In some embodiments, any of the truncated Cas proteins disclosed herein can bind to a target DNA sequence. In some embodiments, a truncated Cas protein can bind to a target DNA sequence if it forms a complex with a gRNA containing a sequence complementary to the target DNA sequence (e.g., a spacer sequence). In some embodiments, a truncated Cas protein can bind to a target sequence at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% more efficiently than a reference non-truncated Cas protein (e.g., a Cas protein containing the amino acid sequence of SEQ ID NO: 2) for the same target sequence under the same conditions. In some embodiments, the truncated Cas protein can bind to the target sequence 10-100%, 10-80%, 10-60%, 10-40%, 10-20%, 30-100%, 30-80%, 30-60%, 30-40%, 50-100%, 50-80%, 50-60%, 70-100%, 70-80%, 80-90%, 90-100%, or 85-95% more efficiently than a reference untruncated Cas protein (e.g., a Cas protein containing the amino acid sequence of SEQ ID NO: 2) for the same target sequence under the same conditions. In some embodiments, a truncated Cas protein can bind to a target sequence with a binding affinity that is at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the binding affinity of a reference untruncated Cas protein (e.g., a Cas protein containing the amino acid sequence of SEQ ID NO: 2) for the same target sequence under the same conditions. In some embodiments, a truncated Cas protein can bind to a target sequence with a binding affinity of 10-100%, 10-80%, 10-60%, 10-40%, 10-20%, 30-100%, 30-80%, 30-60%, 30-40%, 50-100%, 50-80%, 50-60%, 70-100%, 70-80%, 80-90%, 90-100%, or 85-95% of the binding affinity of a reference untruncated Cas protein (e.g., a Cas protein containing the amino acid sequence of SEQ ID NO: 2) for the same target sequence under the same conditions.

[0090] In some embodiments, the Cas9 protein may be fused to a heterologous moiety. In some embodiments, the heterologous moiety is a deaminase (e.g., TadA). In some embodiments, the heterologous moiety is a protein useful for gene silencing (e.g., the KRAB domain, Dnmt3A, and / or the Dnmt3L protein domain). In some embodiments, the heterologous moiety is a DNA methyltransferase, e.g., Dnmt1, Dnmt3A, or Dnmt3B. In some embodiments, the heterologous moiety is a factor regulator of DNA methyltransferase [e.g., a catalytically inactive regulator of DNA methyltransferase (e.g., Dnmt3L)]. In some embodiments, the heterologous moiety is a protein useful for reversing gene silencing (e.g., the TET1 DNA demethylase catalytic domain). In some embodiments, the heterologous moiety is a transcriptional repressor. In some embodiments, the heterologous moiety is a transcriptional activator.

[0091] In some embodiments, the compact protein contains a KRAB domain, which is a category of transcriptional repression domain typically 45–75 amino acids long. See, for example, Ecco, G., Imbeault, M., Trono, D., KRAB zinc finger proteins, Development 144, 2017; Lambert et al. The human transcription factors, Cell 172, 2018. In some embodiments, the KRAB domain is the KRAB domain of Kox1. In some embodiments, the KRAB domain contains the sequence of SEQ ID NO: 73. In some embodiments, the KRAB domain contains an amino acid sequence that is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 73.

[0092] In some embodiments, the compact protein comprises the DNA(cytosine-5)-methyltransferase 3A (Dnmt3A) protein or a fragment thereof. In some embodiments, Dnmt3A comprises the sequence of SEQ ID NO: 74. In some embodiments, Dnmt3A comprises an amino acid sequence that is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 74.

[0093] In some embodiments, the compact protein comprises the DNA(cytosine-5)-methyltransferase 3L (Dnmt3L) protein or a fragment thereof. In some embodiments, Dnmt3L comprises the sequence of SEQ ID NO: 75. In some embodiments, Dnmt3L comprises an amino acid sequence that is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 75.

[0094] Nucleic acid base editor activity In some embodiments, the nucleic acid base editor has at least 50%, 60%, 70%, 80%, or 90% of the base editing activity compared to a nucleic acid base editor containing the sequence of SEQ ID NO: 3. In some embodiments, the base editing activity is evaluated using a polynucleotide containing the nucleotide sequence of SEQ ID NO: 4. In some embodiments, the base editing activity is evaluated at positions A5, A8, or A13 of SEQ ID NO: 4.

[0095] In some embodiments, any of the proteins disclosed herein further comprises one or more nuclear localization signals (NLS) and optionally one or more linkers that fuse one or more NLS to any of the proteins disclosed herein (e.g., any of the nucleic acid base editors disclosed herein). In some embodiments, one or more NLS are selected from c-myc NLS (e.g., an NLS containing the amino acid sequence of SEQ ID NO: 5), SV40 NLS (e.g., an NLS containing the amino acid sequence of SEQ ID NO: 7), and / or nucleoplasmin NLS (e.g., an NLS containing the amino acid sequence of SEQ ID NO: 9). In some embodiments, the disclosure provides a nucleic acid base editor comprising a) a first NLS conjugated to the N-terminus of the nucleic acid base editor by optionally a first NLS linker, and b) a second NLS conjugated to the C-terminus of the nucleic acid base editor by optionally a second NLS linker. In some embodiments, the nucleic acid base editor further comprises c) a third NLS conjugated to the second NLS by optionally a third NLS linker.

[0096] In some embodiments, the first NLS comprises the amino acid sequence of SEQ ID NO: PAAKKKKLD, an optional first NLS linker comprises the amino acid sequence of SEQ ID NO: GSVD, the second NLS comprises the amino acid sequence of SEQ ID NO: PKKKRKV, an optional second NLS linker comprises the amino acid sequence of SEQ ID NO: TGGGPGGGAAAGSGS, the third NLS comprises the amino acid sequence of SEQ ID NO: KRPAATKKAGQAKKKK, and an optional third NLS linker comprises the amino acid sequence of SEQ ID NO: GSGS. In some embodiments, one or more NLSs and one or more linkers that fuse one or more NLSs to a nucleic acid base editor have a total of 40 amino acids or less. In some embodiments, one or more NLSs and one or more linkers that fuse one or more NLSs to a nucleic acid base editor have a total of 37 amino acids or less. In some embodiments, one or more NLSs and one or more linkers that fuse one or more NLSs to a nucleic acid base editor have a total of 35 amino acids or less. In some embodiments, one or more NLSs and one or more linkers fusing one or more NLSs to a nucleic acid base editor have a combined length of 32 amino acids or less. In some embodiments, one or more NLSs and one or more linkers fusing one or more NLSs to a nucleic acid base editor have a combined length of 28 amino acids or less. In some embodiments, one or more NLSs and one or more linkers fusing one or more NLSs to a nucleic acid base editor have a combined length of 26 amino acids or less. In some embodiments, one or more NLSs and one or more linkers fusing one or more NLSs to a nucleic acid base editor have a combined length of 10 to 40 amino acids. In some embodiments, one or more NLSs and one or more linkers fusing one or more NLSs to a nucleic acid base editor have a combined length of 11 to 37 amino acids. In some embodiments, one or more NLSs and one or more linkers fusing one or more NLSs to a nucleic acid base editor have a combined length of 13 to 35 amino acids.In some embodiments, one or more NLSs and one or more linkers fusing one or more NLSs to a nucleic acid base editor have a combined length of 15 to 32 amino acids. In some embodiments, one or more NLSs and one or more linkers fusing one or more NLSs to a nucleic acid base editor have a combined length of 16 to 30 amino acids. In some embodiments, one or more NLSs and one or more linkers fusing one or more NLSs to a nucleic acid base editor have a combined length of 18 to 28 amino acids. In some embodiments, one or more NLSs and one or more linkers fusing one or more NLSs to a nucleic acid base editor have a combined length of 20 to 26 amino acids.

[0097] In some embodiments, up to eight amino acids are deleted from the amino terminus of the nucleotide deaminase. In some embodiments, up to eight amino acids are deleted from the carboxy terminus of the nucleotide deaminase. In some embodiments, up to eight amino acids are deleted from the carboxy terminus of the nucleotide deaminase. In some embodiments, up to 22 amino acids are deleted from the REC-lobe domain of the Cas9 nuclease.

[0098] In some embodiments, the linker amino acid sequence is selected from the group of SEQ ID NOs: 11, 12, and 13, and has up to 8 amino acids deleted from the amino terminus of RNA adenine deaminase; up to 8 amino acids deleted from the carboxyl terminus of RNA adenine deaminase; and up to 22 amino acids deleted from the REC-lobe domain of Cas9 nuclease.

[0099] In some embodiments, the amino acid sequence encoding the nucleic acid base editor is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30, SEQ ID NO: 31, SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51, SEQ ID NO: 52, SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, or SEQ ID NO: 56. In some embodiments, the amino acid sequence encoding the nucleic acid base editor is selected from the group SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, and 56.

[0100] In some embodiments, the total activity of the nucleic acid base editor is at least 30% or greater than the total activity of the wild-type nucleic acid base editor. In some embodiments, the total activity of the nucleic acid base editor is at least 40% or greater than the total activity of the wild-type nucleic acid base editor. In some embodiments, the total activity of the nucleic acid base editor is at least 50% or greater than the total activity of the wild-type nucleic acid base editor. In some embodiments, the total activity of the nucleic acid base editor is at least 60% or greater than the total activity of the wild-type nucleic acid base editor. In some embodiments, the total activity of the nucleic acid base editor is at least 70% or greater than the total activity of the wild-type nucleic acid base editor. In some embodiments, the total activity of the nucleic acid base editor is at least 80% or greater than the total activity of the wild-type nucleic acid base editor. In some embodiments, the total activity of the nucleic acid base editor is at least 90% or greater than the total activity of the wild-type nucleic acid base editor. In some embodiments, the base editing activity is compared to the wild-type nucleic acid base editor containing the sequence of Sequence ID No. 3. In some embodiments, base editing activity is evaluated using a polynucleotide containing the nucleotide sequence of SEQ ID NO: 4. In some embodiments, base editing activity is evaluated at positions A5, A8, or A13 of SEQ ID NO: 4:ATGCATTAACTGAAAATGGTCA.

[0101] Some aspects of this disclosure relate to nucleic acids encoding a nucleotide deaminase, a Cas9 nuclease, and a nucleic acid base editor comprising a linker between the deaminase and the Cas9 nuclease, wherein the nucleic acid base editor has an amino acid length between 1175 and 1249. In some embodiments, the nucleic acid further comprises an inverted end sequence (ITR). In some embodiments, the nucleic acid further comprises an sgRNA. In some embodiments, the nucleic acid further comprises a promoter sequence.

[0102] In some embodiments, the total nucleic acid from ITR to ITR, including the ITR, is less than 5kb. In some embodiments, the total nucleic acid from ITR to ITR, including the ITR, is less than 4.9kb. In some embodiments, the total nucleic acid from ITR to ITR, including the ITR, is less than 4.85kb. In some embodiments, the total nucleic acid from ITR to ITR, including the ITR, is less than 4.8kb.

[0103] In some embodiments, the amino acid sequence encoding the nucleic acid base editor is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, SEQ ID NO: 24, SEQ ID NO: 25, SEQ ID NO: 26, SEQ ID NO: 27, SEQ ID NO: 28, SEQ ID NO: 29, SEQ ID NO: 30, SEQ ID NO: 31, SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51, SEQ ID NO: 52, SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, or SEQ ID NO: 56. In some embodiments, the amino acid sequence encoding the nucleic acid base editor is selected from the group SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, and 56.

[0104] vector Some aspects of this disclosure relate to vectors comprising nucleic acids encoding a nucleotide deaminase, a Cas9 nuclease, and a nucleic acid base editor including a linker between the deaminase and the Cas9 nuclease, wherein the nucleic acid base editor has an amino acid length between 1175 and 1249. In some embodiments, the vector is selected from adeno-associated viruses, adenoviruses, lentiviruses, bacteriophages, and virus-like particles. In some embodiments, the vector is an adeno-associated virus.

[0105] Isolated cells Certain embodiments of this disclosure relate to isolated cells comprising a nucleotide deaminase, a Cas9 nuclease, and a nucleic acid base editor comprising a linker between the deaminase and the Cas9 nuclease, wherein the nucleic acid base editor has an amino acid length between 1175 and 1249. In some embodiments, the isolated cells comprise a nucleic acid encoding a nucleic acid base editor comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, wherein the nucleic acid base editor has an amino acid length between 1175 and 1249. In some embodiments, the isolated cells comprise a nucleic acid encoding a nucleic acid base editor comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, wherein the nucleic acid base editor has an amino acid length between 1175 and 1249, and the nucleic acid further comprises an inverted terminal sequence (ITR). In some embodiments, the isolated cells contain a nucleic acid encoding a nucleic acid base editor comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, wherein the nucleic acid base editor has an amino acid length between 1175 and 1249, and the nucleic acid further comprises an ITR and sgRNA. In some embodiments, the isolated cells contain nucleic acids encoding a nucleic acid base editor, which includes a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, the nucleic acid base editor having an amino acid length between 1175 and 1249, and the nucleic acid further contains an ITR, sgRNA, and promoter sequence, and the total nucleic acid from ITR to ITR is less than 5kb.In some embodiments, the isolated cells contain nucleic acids encoding a nucleic acid base editor comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, the nucleic acid base editor having an amino acid length between 1175 and 1249, and the nucleic acid further comprising an ITR, sgRNA, and promoter sequence, with the total nucleic acid from ITR to ITR being less than 4.95 kb. In some embodiments, the isolated cells contain nucleic acids encoding a nucleic acid base editor comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, the nucleic acid base editor having an amino acid length between 1175 and 1249, and the nucleic acid further comprising an ITR, sgRNA, and promoter sequence, with the total nucleic acid from ITR to ITR being less than 4.85 kb.

[0106] In some embodiments, the isolated cells comprise a vector containing a nucleic acid encoding a nucleic acid base editor comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, the nucleic acid base editor having an amino acid length between 1175 and 1249. In some embodiments, the isolated cells comprise a vector containing a nucleic acid encoding a nucleic acid base editor comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, the nucleic acid base editor having an amino acid length between 1175 and 1249, and the nucleic acid further comprises an inverted terminal sequence (ITR). In some embodiments, the isolated cells comprise a vector containing a nucleic acid encoding a nucleic acid base editor comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, the nucleic acid base editor having an amino acid length between 1175 and 1249, and the nucleic acid further comprises an ITR and sgRNA. In some embodiments, the isolated cells comprise a vector containing a nucleic acid encoding a nucleic acid base editor, which includes a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, wherein the nucleic acid base editor has an amino acid length between 1175 and 1249, and the nucleic acid further comprises an ITR, sgRNA, and a promoter sequence, with the total nucleic acid from ITR to ITR being less than 5kb. In some embodiments, the isolated cells contain a vector containing nucleic acids encoding a nucleic acid base editor, which includes a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, the nucleic acid base editor having an amino acid length between 1175 and 1249, and the nucleic acid further includes an ITR, sgRNA, and a promoter sequence, with the total nucleic acid from ITR to ITR being less than 4.95 kb.In some embodiments, the isolated cells comprise a vector containing a nucleic acid encoding a nucleic acid base editor, which includes a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, the nucleic acid base editor having an amino acid length between 1175 and 1249, and the nucleic acid further comprises an ITR, sgRNA, and a promoter sequence, with the total nucleic acid from ITR to ITR being less than 4.9 kb. In some embodiments, the isolated cells comprise a vector containing a nucleic acid encoding a nucleic acid base editor, which includes a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, the nucleic acid base editor having an amino acid length between 1175 and 1249, and the nucleic acid further comprises an ITR, sgRNA, and a promoter sequence, with the total nucleic acid from ITR to ITR being less than 4.85 kb. In some embodiments, the isolated cells contain a vector containing nucleic acids encoding a nucleic acid base editor, which includes a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, the nucleic acid base editor having an amino acid length between 1175 and 1249, and the nucleic acid further includes an ITR, sgRNA, and a promoter sequence, with the total nucleic acid from ITR to ITR being less than 4.8 kb.

[0107] Treatment method The term “treatment” or “to treat” means preventing or delaying the onset of a disease, slowing the rate of progression, and / or restoring the symptoms of a disorder. In some embodiments, any of the proteins disclosed herein (e.g., any of the nucleic acid base editors disclosed herein) may be administered to a subject. In some embodiments, the subject has a disease. In some embodiments, the subject has a liver disease. In some embodiments, the subject has a neurological disorder. [Examples]

[0108] The following embodiments are provided to illustrate the present disclosure. The embodiments are not intended to be limiting in any way.

[0109] Experimental Procedure for Examples Base editor activity was evaluated in HEK293FT cells. Base editor variants and appropriate sgRNAs were delivered into cells via a single plasmid. Base editor expression was regulated by the CMV promoter, and sgRNA expression was regulated by the U6 promoter. Base editors were transfected into HEK293FT cells using Lipofectamine 2000 transfection reagent (ThermoFisher Scientific, 11668019) and incubated at 37°C. After 3 days, cells were lysed, and genomic DNA was extracted using Lucigen QuickExtract DNA extract. Genomic DNA was analyzed at target sites by high-throughput DNA sequencing or Sanger sequencing. The A-to-G conversion percentage was determined. [Example 1]

[0110] Shortened linker between deaminase and Cas9 nuclease The linker between TadA-8e and nSaCas9 was pruned from 16 amino acids to 3 amino acids (SEQ ID NO: 11), 5 amino acids (SEQ ID NO: 12), or 7 amino acids (SEQ ID NO: 13). All three linker substitutions were well tolerable, as shown in Figures 1A, 1B, and Table 1. Nucleic acid base activity was compared to that of the wild-type SaABE nucleotide editor. All three nucleotide editors with reduced linkers exhibited greater nucleotide base activity than 100% of the wild-type nucleotide editor activity. However, all three linker substitutions resulted in a narrower editing window. In some embodiments, a narrowed editing window may be advantageous in certain applications where multiple A-to-G editing is undesirable. [Example 2]

[0111] Amino acid removal from deaminase Next, the ends of TadA-8e were pruned. Deletions of up to eight amino acids at both ends of TadA-8e were well tolerable, as shown in Figures 2A, 2B, and 2C. The shortened end experiments of TadA-8e were then combined with the shortened linker from Example 1. The results are shown in Figures 3A and 3B, where SEQ ID NOs. 36 and 38 exhibited 125% and 65% of the wild-type nucleotide editor activity, respectively. The most efficient combination was the removal of eight amino acids from the N-terminus of TadA-8e and six amino acids from the C-terminus. Such truncation may also be applicable to other Cas9 variants. [Example 3]

[0112] Removal of amino acids from Cas9 nuclease Amino acids were also pruned from the terminals of nSaCas9 (Figures 4A and 4B). Up to four amino acids could be pruned from the N-terminus of nSaCas9. SaABE could not tolerate amino acids being pruned from the C-terminus of nSaCas9. The SaABE editing level was maintained when TadA-terminus truncation, a shortening linker, and nSaCas9 N-terminus truncation were combined.

[0113] A sequence of amino acids was pruned within the REC lobe, the RuvC domain, and the PAM interaction domain of nSaCas9. Certain truncations in the REC lobe and RuvC domain were acceptable, but truncations in the PAM interaction domain were not acceptable in terms of base editing activity. As shown in Figures 5A, 5B, and 5C, a sequence of up to 22 amino acids could be deleted in the REC lobe domain. This truncation was acceptable in combination with other truncations in Examples 1 and 2, and this combination of deletions resulted in an AAV genome size of up to 4.9 kb. Removal of a sequence of up to 13 amino acids was also acceptable in the RuvC domain (Figures 6A, 6B, and 6C). However, when deletions in the RuvC domain were combined with other deletions (in the REC lobe and / or TadA-8e / linker), editing activity was significantly reduced.

[0114] Combining selected deletions in TadA-8e, the linker, the N-terminus of nSaCas9, the REC-lobe, and the RuvC domain resulted in several modified BEs exhibiting activity exceeding 100% of that of wild-type SaABE, SEQ ID NO: 32, SEQ ID NO: 33, and SEQ ID NO: 47 (Figures 9A and 9B). The overall activity after combining deletions in TadA-8e, the linker, the N-terminus of nSaCas9, the REC-lobe, and the RuvC domain is less than 50% of that of the original SaABE. However, it may be possible to restore activity through arginine scanning, which increases the binding energy.

[0115] Table 1 shows the base editor activity of all modified BEs compared to the activity of the reference BE (SEQ ID NO: 3). Table 1 provides representative SEQ ID NOs for the modified BEs and shows the modified sequences compared to the reference BE. The sequences in Table 1 are provided to show a visual representation of the modifications in the sequences used, with TadA in bold text, the linker between the deaminase and nuclease in italics, and the nuclease in normal font.

[0116] [Table 1] TIFF2026516500000002.tif250153 TIFF2026516500000003.tif250153 TIFF2026516500000004.tif250153 TIFF2026516500000005.tif250153 TIFF2026516500000006.tif250153 TIFF2026516500000007.tif250153 TIFF2026516500000008.tif250153 TIFF2026516500000009.tif250153 TIFF2026516500000010.tif250153 TIFF2026516500000011.tif250153 TIFF2026516500000012.tif250153 TIFF2026516500000013.tif250153 TIFF2026516500000014.tif250153 TIFF2026516500000015.tif250153 TIFF2026516500000016.tif250153 TIFF2026516500000017.tif250153 TIFF2026516500000018.tif250153 TIFF2026516500000019.tif250153 TIFF2026516500000020.tif250153 TIFF2026516500000021.tif250153 TIFF2026516500000022.tif250153 TIFF2026516500000023.tif250153 TIFF2026516500000024.tif250153 TIFF2026516500000025.tif36160

[0117] In separate experiments, the nucleic acid base editing activity of modified BEs containing the amino acid sequences of SEQ ID NO: 76 or 77 was tested. Each modified BE was found to edit 9% of the adenine at position 13 of SEQ ID NO: 4 to guanine. The sequences of SEQ ID NOs: 76 and 77 are presented below, with TadA in bold text, the linker between the deaminase and nuclease in italics, and the nuclease in normal font.

[0118] Sequence ID 76 (ABE8e_SaCas9_delta_E506-K623): [ka]

[0119] Sequence ID 77 (ABE8e_SaCas9_delta_K2-Y5_K719-Q731_E506-K623): [ka]

[0120] Array description Table 2 contains descriptions of the sequences included in the electronic sequence listing. The table provides the sequence number, sequence description, and amino acid sequence length. These sequences are included in the aforementioned electronic sequence listing (41816_PCT_SequenceListing; size: 172KB; and creation date: May 6, 2024), which is incorporated in its entirety by reference.

[0121] [Table 2] TIFF2026516500000029.tif249156

Claims

1. A nucleic acid base editor comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, with an amino acid length between 1175 and 1249.

2. 1175-1249, 1180-1249, 1190-1249, 1200-1249, 1210-1249, 1220-1249, 1230-1249, 1240-1249, 1180-1200, 1180-1190, 1200-1249, 1200-1240, 1200-1230, 1200-1220, 1 A nucleic acid base editor according to claim 1, wherein the amino acid length is between 200-1210, 1210-1249, 1210-1240, 1210-1230, 1210-1220, 1220-1249, 1220-1240, 1220-1230, 1230-1249, 1230-1240, or 1240-1249.

3. The nucleic acid base editor according to any one of claims 1 or 2, wherein the linker between the deaminase and the Cas9 nuclease is 2 to 10, 3 to 7, 3 to 5, or 3 amino acid lengths.

4. The nucleic acid base editor according to any one of claims 1 to 3, wherein the nucleotide deaminase has an amino acid length of 145-166, 145-162, 145-158, 145-154, 145-151, 150-166, 150-162, 150-156, 150-153, 152-166, 152-158, 152-154, 157-166, 157-162, 161-166, or 161-163.

5. The nucleic acid base editor according to any one of claims 1 to 4, wherein the Cas9 nuclease has an amino acid length between 1010-1052, 1010-1040, 1010-1025, 1010-1015, 1020-1052, 1020-1040, 1020-1030, 1030-1052, 1030-1040, 1030-1035, 1035-1046, 1040-1052, 1040-1046, 1045-1052, or 1045-1048.

6. The nucleic acid base editor according to any one of claims 1 to 5, wherein the linker between the deaminase and the Cas9 nuclease comprises one or more prolines and one or more alanines.

7. A nucleic acid base editor according to any one of claims 1 to 6, wherein the nucleotide deaminase comprises an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of Sequence ID No. 1, or a functional fragment of Sequence ID No.

1.

8. The nucleic acid base editor according to claim 7, wherein the nucleotide deaminase lacks the amino acids corresponding to the first 1, 2, 3, 4, 5, 6, 7, or 8 amino acids of SEQ ID NO:

1.

9. The nucleic acid base editor according to claim 7 or 8, wherein the nucleotide deaminase lacks the amino acid corresponding to the last 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids of SEQ ID NO:

1.

10. The nucleic acid base editor according to claim 7 or 8, wherein the nucleotide deaminase lacks the amino acids corresponding to the first six amino acids of SEQ ID NO: 1 and lacks the amino acids corresponding to the last eight amino acids of SEQ ID NO:

1.

11. A nucleic acid base editor according to any one of claims 1 to 10, wherein the Cas9 nuclease comprises an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of Sequence ID No. 2, or a functional fragment of Sequence ID No.

2.

12. A nucleic acid base editor according to any one of claims 1 to 12, wherein the Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 1-4, 1-6, 77-85, 124-126, 124-129, 124-143, 124-145, 131-134, 135-138, 122-129, 505-622, 687-688, 716-720, 718-719, 718-730, 765, and / or 865-873 of SEQ ID NO:

2.

13. A nucleic acid base editor according to any one of claims 1 to 12, wherein the Cas9 nuclease lacks amino acids corresponding to amino acids 1-4, 1-6, 77-85, 124-126, 124-129, 124-143, 124-145, 131-134, 135-138, 122-129, 505-622, 687-688, 716-720, 718-719, 718-730, 765, and / or 865-873 of SEQ ID NO:

2.

14. Cas9 nuclease is an amino acid: a) Sequence IDs 718-719 and 765 of Sequence ID 2; b) Sequence ID No. 2, 718-719 and 687-688; c) Sequence IDs 716-720 and 687-688 of Sequence ID No. 2; d) Sequence IDs 21-4 and 718-730; e) Sequence IDs 21-4, 687-688, and 716-720: f) Sequence IDs 21-4 and 124-145; or g) 1-4, 505-622, and 718-730 A nucleic acid base editor according to any one of claims 1 to 13, which lacks the corresponding amino acid.

15. Cas9 nuclease is an amino acid: a) Sequence IDs 2-1 to 2-8, b) Sequence IDs 2-1 to 2-10, c) Sequence ID 2, 81-94, d) Sequence ID 2, lines 163-164, e) Sequence ID 2, lines 181-184, f) Sequence ID 2, 244-424, g) Sequence ID 2, 714-769, h) Sequence ID 2, 736-742, i) Sequence IDs 978-979 of Sequence ID No. 2, and / or j) Sequence ID 2, numbers 1039-1052 A nucleic acid base editor according to any one of claims 1 to 14, wherein one or more amino acids are not missing from the range of amino acids corresponding to [the specified value].

16. Cas9 nuclease is an amino acid: a) Sequence IDs 2-1 to 2-8, b) Sequence IDs 2-1 to 2-10, c) Sequence ID 2, 81-94, d) Sequence ID 2, lines 163-164, e) Sequence ID 2, lines 181-184, f) Sequence ID 2, 244-424, g) Sequence ID 2, 714-769, h) Sequence ID 2, 736-742, i) Sequence IDs 978-979 of Sequence ID No. 2, and / or j) Sequence ID 2, numbers 1039-1052 A nucleic acid base editor according to any one of claims 1 to 15, which does not lack a range of amino acids corresponding to [the specified amino acid].

17. A nucleic acid base editor according to any one of claims 1 to 16, having at least 50%, 60%, 70%, 80%, or 90% of the base editing activity compared to a nucleic acid base editor containing the sequence of Sequence ID No.

3.

18. The nucleic acid base editor according to claim 17, wherein the base editing activity is evaluated using a polynucleotide containing the nucleotide sequence of SEQ ID NO:

4.

19. The nucleic acid base editor according to claim 18, wherein the base editing activity is evaluated at positions A5, A8, or A13 of SEQ ID NO:

4.

20. A nucleic acid base editor according to any one of claims 1 to 19, further comprising one or more nuclear localization signals (NLS) and one or more linkers for optionally fusing one or more NLS to a nucleic acid base editor.

21. The nucleic acid base editor according to claim 20, wherein one or more NLSs are selected from c-myc NLS (e.g., an NLS containing the amino acid sequence of SEQ ID NO: 5), SV40 NLS (e.g., an NLS containing the amino acid sequence of SEQ ID NO: 7), and / or nucleoplasmin NLS (e.g., an NLS containing the amino acid sequence of SEQ ID NO: 9).

22. a) a first NLS conjugated to the N-terminus of the nucleic acid base editor using a first NLS linker as appropriate, and b) a second NLS conjugated to the C-terminus of the nucleic acid base editor using a second NLS linker as appropriate, according to claim 20 or 21.

23. c) The nucleic acid base editor according to claim 22, further comprising a third NLS conjugated to a second NLS using a third NLS linker as appropriate.

24. The nucleic acid base editor according to claim 23, wherein the first NLS comprises the amino acid sequence of SEQ ID NO: 5, any first NLS linker comprises the amino acid sequence of SEQ ID NO: 6, the second NLS comprises the amino acid sequence of SEQ ID NO: 7, any second NLS linker comprises the amino acid sequence of SEQ ID NO: 8, the third NLS comprises the amino acid sequence of SEQ ID NO: 9, and any third NLS linker comprises the amino acid sequence of SEQ ID NO:

10.

25. The nucleic acid base editor according to any one of claims 20 to 24, wherein one or more NLSs and one or more linkers for fusing one or more NLSs to a nucleic acid base editor are collectively 40, 37, 35, 32, 28, or 26 or fewer amino acids.

26. The nucleic acid base editor according to claim 25, wherein one or more NLSs and one or more linkers that fuse one or more NLSs to the nucleic acid base editor have a total of 32 or fewer amino acids.

27. The nucleic acid base editor according to any one of the claims, wherein the amino acid sequence of the linker is selected from the group of SEQ ID NO: 11, SEQ ID NO: 12, and SEQ ID NO:

13.

28. The nucleic acid base editor according to any one of the claims, wherein the amino acid sequence of the linker between the deaminase and the Cas9 nuclease is Sequence ID No.

11.

29. The nucleic acid base editor according to any one of claims 1 to 27, wherein the amino acid sequence of the linker between the deaminase and the Cas9 nuclease is sequence number 12.

30. The nucleic acid base editor according to any one of claims 1 to 27, wherein the amino acid sequence of the linker between the deaminase and the Cas9 nuclease is SEQ ID NO:

13.

31. The amino acid sequences encoding the nucleic acid base editor are: SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, A nucleic acid base editor according to any one of the claims, which is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to 40, SEQ ID NO: 41, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50, SEQ ID NO: 51, SEQ ID NO: 52, SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, or SEQ ID NO:

56.

32. The nucleic acid base editor according to any one of the claims, wherein the amino acid sequence encoding the nucleic acid base editor is selected from the group consisting of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, and 56.

33. A nucleic acid base editor according to any one of the claims, having at least 50% of the base editing activity compared to a nucleic acid base editor containing the sequence of Sequence ID No.

3.

34. A nucleic acid base editor according to any one of the claims, having at least 60% of the base editing activity compared to a nucleic acid base editor containing the sequence of Sequence ID No.

3.

35. A nucleic acid base editor according to any one of the claims, having at least 70% of the base editing activity compared to a nucleic acid base editor containing the sequence of Sequence ID No.

3.

36. A nucleic acid base editor according to any one of the claims, having at least 80% of the base editing activity compared to a nucleic acid base editor containing the sequence of Sequence ID No.

3.

37. A nucleic acid base editor according to any one of the claims, having at least 90% of the base editing activity compared to a nucleic acid base editor containing the sequence of Sequence ID No.

3.

38. A nucleic acid base editor according to any one of claims 33 to 37, wherein base editing activity is evaluated using a polynucleotide containing the nucleotide sequence of SEQ ID NO:

4.

39. The nucleic acid base editor according to claim 38, wherein the base editing activity is evaluated at positions A5, A8, or A13 of sequence number 4.

40. A truncated nucleic acid base editor comprising a nucleotide deaminase, a Cas9 nuclease, and a linker between the deaminase and the Cas9 nuclease, wherein the nucleic acid base editor has an amino acid length between 1175 and 1249. a) A nucleotide deaminase lacking the amino acids corresponding to the first 1, 2, 3, 4, 5, 6, 7, or 8 amino acids of SEQ ID NO: 1 and / or a nucleotide deaminase lacking the amino acids corresponding to the last 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids of SEQ ID NO: 1; b) The Cas9 nuclease lacks one or more amino acids from the range of amino acids corresponding to amino acids 1-4, 1-6, 77-85, 124-126, 124-129, 124-143, 124-145, 131-134, 135-138, 122-129, 505-622, 687-688, 716-720, 718-719, 718-730, 765, and / or 865-873 of SEQ ID NO: c) The linker between the deaminase and the Cas9 nuclease is selected from the group of SEQ ID NOs: 11, 12, and 13; d) The nucleic acid base editor has at least 50%, 60%, 70%, 80%, or 90% of the base editing activity compared to the nucleic acid base editor containing the sequence of Sequence ID No. 3; e) Nucleic acid base editor, i. Further comprising one or more nuclear localization signals (NLS), one or more NLS selected from c-myc NLS (e.g., an NLS containing the amino acid sequence of SEQ ID NO: 5), SV40 NLS (e.g., an NLS containing the amino acid sequence of SEQ ID NO: 7), and / or nucleoplasmin NLS (e.g., an NLS containing the amino acid sequence of SEQ ID NO: 9); ii. One or more NLSs are fused to a nucleic acid base editor via one or more linkers as appropriate; iii. One or more NLSs and one or more linkers fusing one or more NLSs to a nucleic acid base editor have a total of 32 amino acids or less; and f) The amino acid sequence encoding the nucleic acid base editor is selected from the group of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, and 56. A truncation-type nucleic acid base editor.

41. The truncation-type nucleic acid base editor according to claim 40, having substantially the same base editing activity as a nucleic acid base editor containing the sequence of Sequence ID No.

3.

42. A protein containing Cas9 protein, wherein the Cas9 protein has an amino acid length between 1010–1052, 1010–1040, 1010–1025, 1010–1015, 1020–1052, 1020–1040, 1020–1030, 1030–1052, 1030–1040, 1030–1035, 1035–1046, 1040–1052, 1040–1046, 1045–1052, or 1045–1048.

43. A protein comprising the Cas9 protein, wherein the Cas9 protein has an amino acid sequence that is at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of Sequence ID No. 2, or a functional fragment of Sequence ID No.

2.

44. The protein according to claim 42 or 43, wherein the Cas9 protein lacks one or more amino acids from the range of amino acids corresponding to amino acids 1-4, 1-6, 77-85, 124-126, 124-129, 124-143, 124-145, 131-134, 135-138, 122-129, 505-622, 687-688, 716-720, 718-719, 718-730, 765, and / or 865-873 of SEQ ID NO:

2.

45. The protein according to any one of claims 42 to 44, wherein the Cas9 protein lacks amino acids corresponding to amino acids 1-4, 1-6, 77-85, 124-126, 124-129, 124-143, 124-145, 131-134, 135-138, 122-129, 505-622, 687-688, 716-720, 718-719, 718-730, 765, and / or 865-873 of SEQ ID NO:

2.

46. Cas9 protein contains amino acids: a) Sequence ID No. 2, 718-719 and 765, b) Sequence ID No. 2, 718-719 and 687-688, c) Sequence ID No. 2, 716-720 and 687-688, d) Sequence IDs 21-4 and 718-730, e) Sequence IDs 21-4, 687-688, and 716-720, f) Sequence IDs 21-4 and 124-145, or g) Sequence IDs 21-4, 505-622, and 718-730 A protein according to any one of claims 42 to 45, lacking the corresponding amino acid.

47. Cas9 protein contains amino acids: k) Sequence IDs 2-1 to 2-8, l) Sequence IDs 2-1 to 10, m) Sequence ID 2, 81-94, n) Sequence ID 2, lines 163-164, o) Sequence ID 2, lines 181-184, p) Sequence ID 2, 244-424, q) Sequence ID 2, 714-769, r) Sequence ID 2, 736-742, s) Sequence IDs 978-979 of Sequence ID No. 2, and / or t) Sequence ID 2, numbers 1039-1052 The protein according to any one of claims 42 to 46, which does not lack one or more amino acids from the range of amino acids corresponding to [the specified range].

48. Cas9 protein contains amino acids: k) Sequence IDs 2-1 to 2-8, l) Sequence IDs 2-1 to 10, m) Sequence ID 2, 81-94, n) Sequence ID 2, lines 163-164, o) Sequence ID 2, lines 181-184, p) Sequence ID 2, 244-424, q) Sequence ID 2, 714-769, r) Sequence ID 2, 736-742, s) Sequence IDs 978-979 of Sequence ID No. 2, and / or t) Sequence ID 2, numbers 1039-1052 The protein according to any one of claims 42 to 47, which does not lack the corresponding amino acid range.

49. The protein according to any one of claims 42 to 49, wherein the Cas9 protein lacks one or more amino acids from the amino acid ranges corresponding to amino acids 81-94, 123-145, 244-424, 714-769, and / or 717-730.

50. The protein according to claim 49, wherein the Cas9 protein lacks amino acids corresponding to amino acids 81-94, 123-145, 244-424, 714-769, and / or 717-730.

51. The protein according to any one of claims 42 to 50, wherein the Cas9 protein is fused to a heterologous portion.

52. The protein according to claim 51, wherein the heterogeneous portion is a deaminase.

53. The protein according to claim 51, wherein the heterogeneous portion is a protein useful for gene silencing (e.g., a KRAB domain, a Dnmt3A, and / or a Dnmt3L protein domain).

54. The protein according to claim 51, wherein the heterogeneous portion is a DNA methyltransferase, for example, Dnmt1, Dnmt3A, or Dnmt3B.

55. The protein according to claim 51, wherein the heterogeneous portion is a regulator of a DNA methyltransferase factor (for example, a catalytically inactive regulator of DNA methyltransferase (for example, Dnmt3L)).

56. The protein according to claim 51, wherein the heterogeneous portion is a protein useful for reversing gene silencing (e.g., the TET1 DNA demethylase catalytic domain).

57. The protein according to claim 51, wherein the heterogeneous portion is a transcriptional repressor.

58. The protein according to claim 51, wherein the heterogeneous portion is a transcription activator.

59. A nucleic acid encoding a nucleic acid base editor according to any one of the above claims.

60. The nucleic acid according to claim 59, further comprising an inverted terminal sequence (ITR).

61. The nucleic acid according to claim 59 or 60, further comprising sgRNA.

62. The nucleic acid according to any one of claims 59 to 61, further comprising a promoter sequence.

63. The nucleic acid according to any one of claims 59 to 62, wherein the nucleic acid length is less than 5 kb.

64. The nucleic acid according to any one of claims 60 to 62, wherein the nucleic acid from ITR to ITR is less than 5 kb, including ITR.

65. The nucleic acid according to any one of claims 59 to 62, wherein the nucleic acid is less than 4.9 kb.

66. The nucleic acid according to any one of claims 59 to 62, wherein the nucleic acid from ITR to ITR is less than 4.9 kb, including ITR.

67. The nucleic acid according to any one of claims 59 to 62, wherein the nucleic acid is less than 4.85 kb.

68. The nucleic acid according to any one of claims 60 to 62, wherein the nucleic acid from ITR to ITR is less than 4.85 kb, including the ITR.

69. The nucleic acid according to any one of claims 59 to 62, wherein the nucleic acid is less than 4.8 kb.

70. The nucleic acid according to any one of claims 60 to 62, wherein the nucleic acid from ITR to ITR is less than 4.8 kb, including ITR.

71. A vector comprising the nucleic acid according to any one of claims 59 to 70.

72. The vector according to claim 71, wherein the vector is selected from the group consisting of adeno-associated viruses, adenoviruses, lentiviruses, bacteriophages, and virus-like particles.

73. The vector according to claim 71 or 72, wherein the vector is an adeno-associated virus.

74. Isolated cells comprising the composition according to any one of the above claims.

75. A method for editing a genome sequence in a cell, comprising administering to the cell a protein according to any one of claims 1 to 58, a nucleic acid according to any one of claims 59 to 70, or a vector according to any one of claims 71 to 73.

76. The method according to claim 75, further comprising a) administering a gRNA that targets a sequence in a cell, or b) a nucleic acid encoding a gRNA.