Cytosine to guanine base editing factor
Base-editing fusion proteins with nucleic acid-programmable DNA-binding proteins and polymerases enhance the precision and efficiency of C→G nucleotide changes in genomic DNA, addressing the limitations of existing gene editing technologies and providing a therapeutic solution for genetic diseases.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- PRESIDENT & FELLOWS OF HARVARD COLLEGE
- Filing Date
- 2018-03-09
- Publication Date
- 2026-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing gene editing technologies face challenges in achieving precise and efficient C→G nucleotide changes in genomic DNA, which are crucial for treating genetic diseases, as they often result in unwanted mutations and low efficiency.
Development of base-editing fusion proteins, such as those containing a nucleic acid-programmable DNA-binding protein, uracil DNA glycosylase (UDG) domain, and cytidine deaminase, which create a debasement site in nucleic acids and incorporate cytosine opposite to the site using endogenous or engineered polymerases to enhance C→G editing efficiency.
The fusion proteins significantly improve the precision and efficiency of C→G nucleotide changes, offering a promising approach for treating genetic diseases by minimizing unwanted mutations and enhancing therapeutic efficacy.
Smart Images

Figure 0007896846000149 
Figure 0007896846000150 
Figure 0007896846000151
Abstract
Description
[Technical Field]
[0001] Background of the Invention Targeted editing of nucleic acid sequences (e.g., targeted cleavage or targeted introduction of specific modifications into genomic DNA) is an extremely promising approach for studying gene function and also holds promise for providing new therapies for human genetic diseases. Many genetic diseases can, in principle, be treated by causing specific nucleotide changes at specific locations in the genome (e.g., C→G or G→C changes at specific codons in disease-related genes). The development of programmable methods to achieve such precise gene editing represents both a powerful new research tool and a promising new approach to gene-editing-based therapies. [Overview of the Initiative]
[0002] Provided herein are compositions, kits, and methods for modifying polynucleotides (e.g., DNA), such as generating cytosine-to-guanine mutations in polynucleotides. As described in great detail herein, base editing (e.g., C-to-G editing) is achieved by removing a nucleic acid base (e.g., cytosine (C)), thereby creating a debasement site in the nucleic acid sequence. The nucleic acid base opposite the debasement site (e.g., guanine) is then replaced with a different nucleic acid base (e.g., cytosine) by, for example, an endogenous damage-overcoming DNA polymerase. The base editing fusion proteins described herein have the ability to generate specific mutations (e.g., C-to-G mutations) within nucleic acids (e.g., genomic DNA), and can be used, for example, to treat diseases associated with nucleic acid mutations (e.g., C-to-G or G-to-C mutations).
[0003] One example of a C→G base-editing factor is a fusion protein containing a nucleic acid programmable DNA-binding protein (e.g., Cas9 domain), a uracil DNA glycosylase (UDG) domain, and cytidine deaminase. While not wishing to be bound to a specific theory, such base-editing fusion proteins are capable of binding to specific nucleic acid sequences (e.g., via the Cas9 domain), deaminating cytosine within the nucleic acid sequence to uridine, which can then be excised from the nucleic acid molecule by UDG. The nucleic acid base opposite the debasement site can then be replaced by another base (e.g., cytosine) by, for example, an endogenous damage-overcoming polymerase. Typically, base repair machinery (e.g., in cells) replaces the nucleic acid base opposite the debasement site with cytosine, although other bases (e.g., adenine, guanine, or thymine) may replace the nucleic acid base opposite the debasement site. Furthermore, it was found that incorporating damage-transgressing polymerases into base-editing factors can increase the incorporation of cytosine opposite the debasic site. Therefore, base-editing factors were engineered to incorporate various damage-transgressing polymerases to improve base-editing efficiency. Damage-transgressing polymerases that increase the priority of incorporation of C opposite the debasic site improve C→G nucleic acid base editing. It should be understood that other damage-transgressing polymerases that preferentially incorporate non-C nucleic acid bases (e.g., adenine, guanine, and thymine) may be used to generate alternative mutations (e.g., C→A mutations).
[0004] As another example, a base-editing fusion protein may encompass a nucleic acid-programmable DNA-binding protein (e.g., a Cas9 domain) and a base-removal enzyme that removes a nucleic acid base (e.g., cytosine). Rather than deaminating cytosine to uridine and then excising uridine using UDG, as described above, the base-editing factor may encompass a base-removal enzyme that recognizes and removes nucleic acid bases such as cytosine or thymine without performing the first deamination. Thus, a base-editing factor (e.g., a C→G base-editing factor) was manipulated by fusing a nucleic acid-programmable DNA-binding protein (e.g., a Cas9 domain) with a base-removal enzyme that removes cytosine or thymine from a nucleic acid molecule. Furthermore, with the base-editing factors described above, a damage-overcoming polymerase was incorporated into the base-editing factor, increasing the cytosine incorporation opposite to the debase site created by the base-removal enzyme of the base-editing factor. Schematic diagrams outlining exemplary base-editing proteins and base-editing strategies can be seen, for example, in Figures 1-6, 33-36, 40, and 52.
[0005] In some embodiments, the disclosure provides fusion proteins with base-editing capabilities. Exemplary base-editing fusion proteins include: In some embodiments, the fusion protein comprises (i) a nucleic acid-programmable DNA-binding protein (napDNAbp), (ii) a cytidine deaminase domain, and (iii) a uracil-binding protein (UBP). In some embodiments, the fusion protein further comprises (iv) a nucleic acid polymerase domain (NAP). As another example, the fusion protein may comprise (i) a nucleic acid-programmable DNA-binding protein (napDNAbp), (ii) a cytidine deaminase domain, and (iii) a nucleic acid polymerase (NAP) domain. As yet another example, the fusion protein may comprise (i) a nucleic acid-programmable DNA-binding protein (napDNAbp), and (ii) a base-removing enzyme (BEE). In some embodiments, the fusion protein further comprises (iii) a nucleic acid polymerase (NAP) domain. Base-editing factors and methods of using base-editing factors are described in further detail below. [Brief explanation of the drawing]
[0006] [Figure 1] Figure 1 shows a general schematic diagram illustrating C→T and C→G base editing. Certain DNA polymerases (e.g., damage-overcoming polymerases) are known to replace the base opposite the debasidized site with G. One strategy to achieve C→G base editing is to induce the creation of a debasidized site, then replenish or tether such polymerase to replace the G opposite the debasidized site with C.
[0007] [Figure 2] Figure 2 shows a general schematic diagram illustrating the generation of debase sites via base editing and base-specific repair for C→G editing.
[0008] [Figure 3] Figure 3 is a schematic diagram illustrating Scheme 1 from Figure 1, in which a debasement site is formed for C→G base editing. If debasement is efficiently generated, this can increase the total flux of the C→G editing pathway.
[0009] [Figure 4] Figure 4 is a schematic diagram illustrating Approach 1 for C→G base editing, where an increase in the formation of the debase site is utilized. For example, if debase is efficiently generated by using the UDG domain and damage-overcoming polymerase, this can increase the total flux through the C→G editing pathway.
[0010] [Figure 5]Figure 5 shows a schematic diagram illustrating the effect of UdgX on base editing. UdgX, a homologous species of UDG identified as tightly binding to uracil with minimal uracil cleavage activity, increases the amount of C→G editing. 1.) UdgX* is a variant of UDG determined to lack uracil-binding activity via in vitro assay. 2.) UdgX_On is a variant shown to increase uracil cleavage via in vitro assay. 3.) Direct fusion of UDG cleaves uracil.
[0011] [Figure 6] Figure 6 shows a schematic diagram (on the left) illustrating an exemplary C→T base editing factor (e.g., BE3) containing a uracil glycosylase inhibitor (UGI), a Cas9 domain (e.g., nCas9), and cytidine deaminase. On the right is a schematic diagram illustrating a C→G base editing factor containing uracil DNA glucosylase (UDG) (or a variant thereof), a Cas9 domain (e.g., nCas9), and cytidine deaminase.
[0012] [Figure 7] Figure 7 shows the total editing percentage at HEK2 sites in WT Hap1 cells using seven base editing factors (BE3; BE3_UdgX; BE3_UdgX*; BE2_UdgX_On; BE3_UdgX_On; BE2_UDG; and BE3_UDG). Raw editing values are shown in the left panel. The right panel shows a graphical representation of raw editing values, where C→G base editing is graphically indicated by the dot bar (G) becoming a black bar (C) when sequencing is performed on the DNA strand opposite to the strand containing the edited C.
[0013] [Figure 8]Figure 8 shows the total editing percentage at HEK2 sites in WT Hap1 cells by additional C→G base editing factors (BE3; BE3_UdgX; BE3_REV7; and SMUG1, where BE3 and BE3_UdgX are repeated from Figure 4). The upper panel shows the raw editing values. The lower panel shows a graphical representation of the raw editing values, where C→G base editing is graphically indicated by dotted bars (G) becoming black bars (C) when sequencing is performed on the DNA strand opposite to the strand containing the edited C.
[0014] [Figure 9] Figure 9 shows the ratios of editing specificity at the HEK2 site for various C→G base editing factors (BE3;BE3_UdgX;BE3_UdgX*;BE3_REV7;BE2_UDG;BE3_UDG BE2_UdgX_On;BE3_UdgX_On; and SMUG1) in WT Hap1 cells. The upper panel shows the percentage of total edits that resulted in G→A, C, or T, and the ratio of edits. The lower panel is a graphical representation of the specificity ratio values.
[0015] [Figure 10] Figure 10 shows the total editing percentage at RNF2 sites in WT Hap1 cells using seven base editing factors (BE3; BE3_UdgX; BE3_UdgX*; BE2_UdgX_On; BE3_UdgX_On; BE2_UDG; and BE3_UDG). Raw editing values are shown in the left panel. The right panel shows a graphical representation of raw editing values, where C→G base editing is graphically indicated by a dotted bar (G) becoming a black bar (C) when sequencing is performed on the DNA strand opposite to the strand containing the edited C.
[0016] [Figure 11]Figure 11 shows the total editing percentages at the RNF2 site by additional C→G base editing factors (BE3; BE3_UdgX; BE3_REV7; and SMUG1 in WT Hap1 cells, where BE3 and BE3_UdgX are repeated from Figure 7). The upper panel shows the Raw editing values. The lower panel shows the representation of the Raw editing values in a plot, where C→G base editing is depicted pictorially by the dot-patterned bar (G) becoming a solid black bar (C) when sequencing is performed on the opposite DNA strand of the strand containing the edited C.
[0017] [Figure 12] Figure 12 shows the ratio of editing specificities at the RNF2 site by various C→G base editing factors (BE3; BE3_UdgX; BE3_UdgX*; BE3_REV7; BE2_UDG; BE3_UDG BE2_UdgX_On; BE3_UdgX_On; and SMUG1) in WT Hap1 cells. The upper panel shows the percentage of total edits that became G→A, C, or T, and the ratio of edits. The lower panel is a graphical representation of the specificity ratio values.
[0018] [Figure 13] Figure 13 shows the total editing percentages at the FANCF site in WT Hap1 cells using seven base editing factors (BE3; BE3_UdgX; BE3_UdgX*; BE2_UdgX_On; BE3_UdgX_On; BE2_UDG; and BE3_UDG). The Raw editing values are shown in the left panel. The panel on the right shows the graphical representation of the Raw editing values, where C→G base editing is depicted pictorially by the solid black bar (C) becoming a dot-patterned bar (G).
[0019] [Figure 14]Figure 14 shows the total editing percentage at the FANCF site by additional C→G base editing factors (BE3; BE3_UdgX; BE3_REV7; and SMUG1; where BE3 and BE3UdgX are repeated from Figure 10) in WT Hap1 cells. The upper panel shows the raw editing values. The lower panel shows the graphical representation of the raw editing values, where C→G base editing is graphically indicated by the change of a solid black bar (C) to a dotted bar (G).
[0020] [Figure 15] Figure 15 shows the ratios of editing specificity at FANCF sites by various C→G base editing factors (BE3;BE3_UdgX;BE3_UdgX*;BE3_REV7;BE2_UDG;BE3_UDG BE2_UdgX_On;BE3_UdgX_On; and SMUG1) in WT Hap1 cells. The upper panel shows the percentage of total edits that resulted in C→A, G, or T, and the ratio of edits. The lower panel is a graphical representation of the specificity ratio values.
[0021] [Figure 16] Figure 16 shows the total editing percentage at HEK2 sites in UDG- / -Hap1 cells using seven base editing factors (BE3; BE3_UdgX; BE3_UdgX*; BE2_UdgX_On; BE3_UdgX_On; BE2_UDG; and BE3_UDG). Raw editing values are shown in the left panel. The right panel shows a graphical representation of raw editing values, where C→G base editing is graphically indicated by a dotted bar (G) turning into a black bar (C) when sequencing is performed on the DNA strand opposite to the strand containing the edited C.
[0022] [Figure 17]Figure 17 shows the total editing percentage at HEK2 sites by additional C→G base editing factors (BE3;BE3_UdgX;BE3_REV7; and SMUG1, where BE3 and BE3_UdgX are repeated from Figure 13) in UDG- / -Hap1 cells. The upper panel shows the raw editing values. The lower panel shows a graphical representation of the raw editing values, where C→G base editing is graphically indicated by dotted bars (G) becoming black bars (C) when sequencing is performed on the DNA strand opposite to the strand containing the edited C.
[0023] [Figure 18] Figure 18 shows the editing specificity ratios at the HEK2 site in UDG- / -Hap1 cells by various C→G base editing factors (BE3;BE3_UdgX;BE3_UdgX*;BE3_REV7;BE2_UDG;BE3_UDG BE2_UdgX_On; BE3_UdgX_On; and SMUG1). The upper panel shows the percentage of total edits that resulted in G→A, C, or T, and the editing ratio. The lower panel is a graphical representation of the specificity ratio values.
[0024] [Figure 19] Figure 19 shows the total editing percentage using seven base editing factors (BE3;BE3_UdgX;BE3_UdgX*;BE2_UdgX_On;BE3_UdgX_On;BE2_UDG; and BE3_UDG) at the RNF2 site in UDG- / -Hap1 cells. Raw editing values are shown in the left panel. The right panel shows a graphical representation of raw editing values, where C→G base editing is graphically indicated by a dotted bar (G) becoming a black bar (C) when sequencing is performed on the DNA strand opposite to the strand containing the edited C.
[0025] [Figure 20]Figure 20 shows the total editing percentage at RNF2 sites by additional C→G base editing factors (BE3;BE3_UdgX;BE3_REV7; and SMUG1, where BE3 and BE_UdgX are repeated from Figure 16) in UDG- / -Hap1 cells. The upper panel shows the raw editing values. The lower panel shows a graphical representation of the raw editing values, where C→G base editing is graphically indicated by dotted bars (G) becoming black bars (C) when sequencing is performed on the DNA strand opposite to the strand containing the edited C.
[0026] [Figure 21] Figure 21 shows the editing specificity ratios at the RNF2 site in UDG- / -Hap1 cells by various C→G base editing factors (BE3;BE3_UdgX;BE3_UdgX*;BE3_REV7;BE2_UDG;BE3_UDG BE2_UdgX_On;BE3_UdgX_On; and SMUG1). The upper panel shows the percentage of total edits that resulted in G→A, C, or T, and the editing ratio. The lower panel is a graphical representation of the specificity ratio values.
[0027] [Figure 22] Figure 22 shows the total editing percentage at the FANCF site in UDG- / -Hap1 cells using seven base editing factors (BE3; BE3_UdgX; BE3_UdgX*; BE2_UdgX_On; BE3_UdgX_On; BE2_UDG; and BE3_UDG). Raw editing values are shown in the left panel. The panel on the right shows a graphical representation of the raw editing values, where C→G base editing is graphically represented by a black bar (C) becoming a dotted bar (G).
[0028] [Figure 23]Figure 23 shows the total editing percentage at the FANCF site in UDG- / -Hap1 cells by additional C→G base editing factors (BE3; BE3_UdgX; BE3_REV7; and SMUG1, where BE3 and BE3_UdgX are repeated from Figure 19). The upper panel shows the raw editing values. The lower panel shows the graphical representation of the raw editing values, where C→G base editing is graphically indicated by the change of a solid black bar (C) to a dotted bar (G).
[0029] [Figure 24] Figure 24 shows the total editing specificity ratios at the FANCF site in UDG- / -Hap1 cells by various C→G base editing factors (BE3;BE3_UdgX;BE3_UdgX*;BE3_REV7;BE2_UDG;BE3_UDG BE2_UdgX_On;BE3_UdgX_On; and SMUG1). The upper panel shows the percentage of total edits that resulted in C→A, G, or T, and the editing ratio. The lower panel is a graphical representation of the specificity ratio values.
[0030] [Figure 25] Figure 25 shows the total editing percentage at HEK2 sites by various C→G base editing factors (BE3; BE3_UdgX; BE2_UNG; BE3_UNG; BE2UdgX_On; BE3UdgX_On; and SMUG1) in Rev1- / -Hap1 cells. The upper panel shows the raw editing values. The lower panel shows a graphical representation of the raw editing values, where C→G base editing is graphically indicated by dotted bars (G) becoming black bars (C) when sequencing is performed on the DNA strand opposite to the strand containing the edited C.
[0031] [Figure 26]Figure 26 shows the ratios of editing specificity at the HEK2 site by various C→G base editing factors (BE3;BE3_UdgX;BE2_UNG;BE3_UNG;BE2UdgX_On;BE3UdgX_On; and SMUG1) in Rev1- / -Hap1 cells. The upper panel shows the percentage of total edits that resulted in G→A, C, or T, and the ratio of edits. The lower panel is a graphical representation of the specificity ratio values.
[0032] [Figure 27] Figure 27 shows the total editing percentage at the RNF2 site by various C→G base editing factors (BE3;BE3_UdgX;BE2_UNG;BE3_UNG;BE2UdgX_On;BE3UdgX_On; and SMUG1) in Rev1- / -Hap1 cells. The upper panel shows the raw editing values. The lower panel shows a graphical representation of the raw editing values, where C→G base editing is graphically indicated by dotted bars (G) turning into black bars (C) when sequencing is performed on the DNA strand opposite to the strand containing the edited C.
[0033] [Figure 28] Figure 28 shows the ratios of editing specificity at the RNF2 site by various C→G base editing factors (BE3;BE3_UdgX;BE2_UNG;BE3_UNG;BE2UdgX_On;BE3UdgX_On; and SMUG1) in Rev1- / -Hap1 cells. The upper panel shows the percentage of total edits that resulted in G→A, C, or T, and the ratio of edits. The lower panel is a graphical representation of the specificity ratio values.
[0034] [Figure 29]Figure 29 shows the total editing percentage at the FANCF site in Rev1- / -Hap1 cells by various C→G base editing factors (BE3;BE3_UdgX;BE2_UNG;BE3_UNG;BE2UdgX_On;BE3UdgX_On; and SMUG1). The upper panel shows the raw editing values. The lower panel shows the graphical representation of the raw editing values, where C→G base editing is graphically indicated by the change from a solid black bar (C) to a dotted bar (G).
[0035] [Figure 30] Figure 30 shows the ratios of editing specificity at the FANCF site by various C→G base editing factors (BE3;BE3_UdgX;BE2_UNG;BE3_UNG;BE2UdgX_On;BE3UdgX_On; and SMUG1) in Rev1- / -Hap1 cells. The upper panel shows the percentage of total edits that result in C→A, G, or T, and the ratio of edits. The lower panel is a graphical representation of the specificity ratio values.
[0036] [Figure 31] Figure 31 shows a graphical representation of the raw edit values for the total edit percentage at HEK2, RNF2, and FANCF sites using the C→G base editing factors shown.
[0037] [Figure 32] Figure 32 shows a graphical representation of the ratio of specificity with respect to the total edit percentage at the HEK2, RNF2, and FANCF sites.
[0038] [Figure 33] Figure 33 shows a schematic diagram illustrating an approach to increase the incorporation of C to the opposite side of the debase site for C→G base editing. For example, if the preference for incorporation of C to the opposite side of the debase site is increased by using a polymerase (e.g., a damage-overcoming polymerase), the total C→G base editing will also increase.
[0039] [Figure 34]Figure 34 shows a schematic diagram illustrating an approach to increase the incorporation of C to the opposite side of the debase site for C→G base editing. If the preference for incorporation of C to the opposite side of the debase site is increased, for example by incorporating a damage-overcoming polymerase into the base editing factor, the total C→G base editing will also increase.
[0040] [Figure 35] Figure 35 shows a schematic diagram illustrating different polymerases that may be used in the C→G base editing approaches shown in Figures 33 and 34.
[0041] [Figure 36] Figure 36 shows a schematic diagram (on the left) illustrating an exemplary C→T base editing factor (e.g., BE3) containing a uracil glycosylase inhibitor (UGI), a Cas9 domain (e.g., nCas9), and cytidine deaminase. On the right is a schematic diagram illustrating a C→G base editing factor containing a damage-overcoming polymerase, a Cas9 domain (e.g., nCas9), and cytidine deaminase.
[0042] [Figure 37] Figure 37 shows base editing at the HEK2 site in WT cells using base editing factors tethered to REV1, Pol kappa, Pol etha, and Pol iota. C→G editing is graphically represented in the right panel by the dotted bars (G) becoming black bars (C). Tethering to Pol kappa dramatically increases the efficiency of C→G editing. Raw edit values are shown in the left panel.
[0043] [Figure 38]Figure 38 shows base editing at the RNF2 site in WT cells using base editing factors tethered to REV1, Pol kappa, Pol etha, and Pol iota. C→G editing is graphically represented in the right panel by the dotted bars (G) becoming black bars (C). Tethering to Pol kappa dramatically increases the efficiency of C→G editing. Raw edit values are shown in the left panel.
[0044] [Figure 39] Figure 39 shows base editing at the FANCF site in WT cells using base editing factors tethered to REV1, Pol kappa, Pol etha, and Pol iota. This is graphically represented in the right panel by the transformation of black bars (C) into dotted bars (G). Tethering Pol kappa dramatically increases the efficiency of C→G editing. Raw edit values are shown in the left panel.
[0045] [Figure 40] Figure 40 shows a schematic diagram (on the left) illustrating an exemplary C→G base editing factor containing uracil DNA glycosylase (UDG), a damage-overcoming polymerase, a Cas9 domain (e.g., nCas9), and cytidine deaminase. On the right is a schematic diagram illustrating a C→G base editing factor containing a damage-overcoming polymerase, a Cas9 domain (e.g., nCas9), and a base-clearing enzyme (e.g., a UDG variant capable of cleaving C or T residues).
[0046] [Figure 41]Figure 41 shows C→G base editing at HEK2, RNF2, and FANCF sites using constructs that can be tethered to either Pol kappa or Pol iota, using base editing factors described in the left panel of Figure 40 (base editing factors containing uracil DNA glucosylase (UDG), damage-overcoming polymerase, Cas9 domain, and cytidine deaminase). C→G editing is graphically represented for HEK2 and RNF2 by the conversion of dotted bars (G) to black bars (C), and for FANCF by the conversion of black bars (C) to dotted bars (G).
[0047] [Figure 42] Figure 42 shows base editing at the HEK2 site in WT cells using a base editing factor that tethers to one of Pol kappa, Pol etha, Pol iota, or REV1, as shown in the right panel of Figure 40 (damage-overcoming polymerase, Cas9 domain, and T-cleaving base removal enzyme (UDG147)). The amount of C→G is graphically illustrated at specific residues within the HEK2 site. UDG147 is a UDG variant that directly removes T.
[0048] [Figure 43] Figure 43 shows base editing at the RNF2 site in WT cells using a base editing factor that tethers to one of Pol kappa, Pol etha, Pol iota, or REV1, as shown in the right panel of Figure 40 (damage-overcoming polymerase, Cas9 domain, and T-cleaving base-removing enzyme (UDG147)). The amount of C→G is graphically illustrated at specific residues within the HEK2 site. UDG147 is a UDG variant that directly cleaves T.
[0049] [Figure 44]Figure 44 shows base editing at the FANCF site in WT cells using a base editing factor that tethers to either Pol kappa, Pol etha, Pol iota, or REV1, as shown in the right panel of Figure 40 (damage-overcoming polymerase, Cas9 domain, and T-cleaving base removal enzyme (UDG147)). The amount of C→G is graphically illustrated in specific residues within the HEK2 site. UDG147 is a UDG variant that directly cleaves T.
[0050] [Figure 45] Figure 45 shows base editing at the HEK2 site in WT cells using a base editing factor tethered to one of Pol kappa, Pol etha, Pol iota, or REV1, as shown in the right panel of Figure 40 (damage-overcoming polymerase, Cas9 domain, and base-cleaving enzyme (UDG204) that excises C). The amount of C→G is graphically illustrated at specific residues within the HEK2 site. UDG204 is a UDG variant that directly excises C.
[0051] [Figure 46] Figure 46 shows base editing at the RNF2 site in WT cells using a base editing factor that can be tethered to either Pol kappa, Pol etha, Pol iota, or REV1, as shown in the right panel of Figure 40 (a base editing factor containing a damage-overcoming polymerase, Cas9 domain, and a base-cleaving enzyme (UDG204) that excises C). The amount of C→G is graphically illustrated at specific residues within the HEK2 site. UDG204 is a UDG variant that directly excises C.
[0052] [Figure 47]Figure 47 shows base editing at the FANCF site in WT cells using a base editing factor that can be tethered to either Pol kappa, Pol etha, Pol iota, or REV1, as shown in the right panel of Figure 40 (a base editing factor containing a damage-overcoming polymerase, Cas9 domain, and a base-cleaving enzyme (UDG204) that excises C). The amount of C→G is graphically illustrated at specific residues within the HEK2 site. UDG204 is a UDG variant that directly excises C.
[0053] [Figure 48] Figure 48 shows a schematic diagram illustrating the role of MSH2 in base repair, where MSH2 can promote the conversion of uracil (U) to cytosine (C) in DNA.
[0054] [Figure 49] Figure 49 shows base editing at the HEK2 site in MSH2- / - cells using six base editing factors (BE3; BE3_UdgX; BE3_UdgX*; BE2_UdgX_On; BE3_UdgX_On; and BE3_UDG). Raw edit values are shown in the left panel. The panel on the right shows a graphical representation of the raw edit values, where C→G base editing is graphically represented by a dotted bar (G) becoming a solid black bar (C).
[0055] [Figure 50] Figure 50 shows base editing at the RNF2 site in MSH2- / - cells using six base editing factors (BE3;BE3_UdgX;BE3_UdgX*;BE2_UdgX_On;BE3_UdgX_On; and BE3_UDG). Raw edit values are shown in the left panel. The panel on the right shows a graphical representation of the raw edit values, where C→G base editing is graphically represented by a dotted bar (G) becoming a solid black bar (C).
[0056] [Figure 51]Figure 51 shows base editing at the FANCF site in MSH2- / - cells using six base editing factors (BE3;BE3_UdgX;BE3_UdgX*;BE2_UdgX_On;BE3_UdgX_On; and BE3_UNG). Raw edit values are shown in the left panel. The panel on the right shows a graphical representation of the raw edit values, where C→G base editing is graphically represented by a black bar (C) changing to a dotted bar (G).
[0057] [Figure 52] Figure 52 is a schematic diagram illustrating a base editing approach in which a C→G base editing factor containing a UDG (or UDG variant), a Cas9 (e.g., nCas9) domain, and cytidine deaminase is expressed in trans along with a damage-overcoming polymerase.
[0058] [Figure 53] Figure 53 shows base editing at the HEK2 site in HEK293 cells using five base editing factors expressed in trans (BE3;BE3_UdgX;BE3_UdgX*;BE2_UdgX_On; and BE3_UDG) along with various polymerases (Pol kappa, Pol etha, Pol iota, REV1, Pol beta, and Pol delta). C→G base editing is graphically represented by the transformation of a dotted bar (G) into a black bar (C).
[0059] [Figure 54] Figure 54 shows base editing at the RNF2 site in HEK293 cells using five base editing factors expressed in trans (BE3;BE3_UdgX;BE3_UdgX*;BE2_UdgX_On; and BE3_UDG) along with various polymerases (Pol kappa, Pol etha, Pol iota, REV1, Pol beta, and Pol delta). C→G base editing is graphically represented by the transformation of a dotted bar (G) into a black bar (C).
[0060] [Figure 55] Figure 55 shows base editing at the FANCF site in HEK293 cells using five base editing factors expressed in trans (BE3;BE3_UdgX;BE3_UdgX*;BE2_UdgX_On; and BE3_UDG) along with various polymerases (Pol kappa, Pol etha, Pol iota, REV1, Pol beta, and Pol delta). C→G base editing is graphically represented by a black bar (C) being replaced by a dotted bar (G). [Modes for carrying out the invention]
[0061] definition When used herein and in claims, the singular forms "a," "an," and "the" include singular or plural unless the context makes otherwise obvious. Thus, for example, a reference to "an agent" includes both a singular agent and a plural such agent.
[0062] As used herein, the terms “deaminase” or “deaminase domain” refer to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes the hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase domain that catalyzes the hydrolytic deamination of cytosine to uracil. In some embodiments, the terms “deaminase” or “deaminase domain” refer to a naturally occurring deaminase from a living organism, such as a human, chimpanzee, gorilla, monkey, cattle, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from a living organism that does not exist in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring deaminase from an organism.
[0063] The term “base editing factor (BE)” or “nucleic acid base editing factor (NBE)” refers to a drug containing a polypeptide capable of modifying bases (e.g., A, T, C, G, or U) within a nucleic acid sequence (e.g., DNA or RNA). In some embodiments, base editing factors are capable of deaminating bases within nucleic acids. In some embodiments, base editing factors are capable of deaminating bases within DNA molecules. In some embodiments, base editing factors are capable of deaminating cytosine (C) in DNA. In some embodiments, base editing factors are capable of cleaving bases within DNA molecules. In some embodiments, base editing factors are capable of cleaving adenine, guanine, cytosine, thymine, or uracil within nucleic acid (e.g., DNA or RNA) molecules. In some embodiments, base editing factors are proteins containing a nucleic acid programmable DNA-binding protein (napDNAbp) that is fused to cytidine deaminase (e.g., a fusion protein). In some embodiments, the base editing factor is fused to a uracil-binding protein (UBP) such as uracil DNA glucosylase (UDG). In some embodiments, the base editing factor is fused to a nucleic acid polymerase (NAP) domain. In some embodiments, the (NAP) domain is a damage-overcoming DNA polymerase. In some embodiments, the base editing factor comprises napDNAbp, cytidine deaminase, and a UBP (e.g., UDG). In some embodiments, the base editing factor comprises napDNAbp, cytidine deaminase, and a nucleic acid polymerase (e.g., a damage-overcoming DNA polymerase). In some embodiments, the base editing factor comprises napDNAbp, cytidine deaminase, a UBP (e.g., UDG), and a nucleic acid polymerase (e.g., a damage-overcoming DNA polymerase).
[0064] In some embodiments, the napDNAbp of the base editing factor is a Cas9 domain. In some embodiments, the base editing factor comprises a Cas9 protein fused to cytidine deaminase. In some embodiments, the base editing factor comprises Cas9 nickase (nCas9) fused to cytidine deaminase. In some embodiments, the Cas9 nickase comprises the D10A mutation and histidine at residue 840 of SEQ ID NO: 6, or the corresponding mutation in any Cas9 provided herein, such as any one of SEQ ID NOs: 4-26, which gives Cas9 the ability to cleave only one strand of a double-stranded nucleic acid. In some embodiments, the base editing factor comprises nuclease-inactive Cas9 (dCas9) fused to cytidine deaminase. In some embodiments, the dCas9 domain includes the mutations in SEQ ID NO: 6, D10A and H840A, or any corresponding mutation in any Cas9 provided herein, such as any one of SEQ ID NOs: 4-26, which inactivates the nuclease activity of the Cas9 protein.
[0065] The term “linker,” as used herein, refers to a link (e.g., covalently), chemical group, or molecule that connects two molecules or parts, such as a nuclease-inactive Cas9 domain and a nucleic acid editing domain (e.g., cytidine deaminase), or two domains of a fusion protein. In some embodiments, the linker connects the gRNA-binding domain (including the Cas9 nuclease domain) of an RNA-programmable nuclease to the catalytic domain of a nucleic acid editing protein. In some embodiments, the linker connects dCas9 to a nucleic acid editing protein. Typically, the linker is located between or flanked by two groups, molecules, or other parts, and is connected to each other via covalent bonds, thus connecting the two. In some embodiments, the linker is an amino acid or a group of amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5 to 100 amino acid lengths, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acid lengths. Longer or shorter linkers are also intended. In some embodiments, the linker includes the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 102), which may also be referred to as the XTEN linker. In some embodiments, the linker includes the amino acid sequence SGGS (SEQ ID NO: 103).In some embodiments, the linker is (SGGS)n (SEQ ID NO: 103), (GGGS)n (SEQ ID NO: 104), (GGGGS)n (SEQ ID NO: 105), (G)n (SEQ ID NO: 121), (EAAAK)n (SEQ ID NO: 106), (GGS)n (SEQ ID NO: 122), SGSETPGTSESATPES (SEQ ID NO: 102), (XP)n motif (SEQ ID NO: 123), SGGSSGSETPGTSESATPESSGGS (SEQ ID NO: 107), SGGSSGGSSGSETPGTSESATPE This includes SSGGSSGGS (SEQ ID NO: 108), GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS (SEQ ID NO: 109), SGGSGGSGGS (SEQ ID NO: 120), or any combination thereof, where n is an integer independently between 1 and 30, and where X is also any amino acid. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15.
[0066] As used herein, the term “mutation” refers to the substitution of one residue in a sequence (e.g., a nucleic acid sequence or an amino acid sequence) with another residue, or the deletion or insertion of one or more residues in a sequence. Mutations are typically described herein by identifying the original residue in the sequence, followed by identifying the location of the residue, and then identifying the identity of the newly substituted residue. Various methods for producing amino acid substitutions (mutations) provided herein are well known in the art and are provided, for example, in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)).
[0067] As used herein, “uracil-binding protein” or UBP refers to a protein capable of binding to uracil. In some embodiments, the uracil-binding protein is a uracil-modifying enzyme. In some embodiments, the uracil-binding protein is a uracil base-removing enzyme. In some embodiments, the uracil-binding protein is a uracil DNA glucosylase (UDG). In some embodiments, the uracil-binding protein binds to uracil with an affinity that is at least 1%, 2%, 3%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or at least 95% of the affinity that wild-type UDG (e.g., human UDG) has for binding to uracil.
[0068] As used herein, the term “base removal enzyme” or “BEE” refers to a protein capable of removing bases (e.g., A, T, C, G, or U) from nucleic acid molecules (e.g., DNA or RNA). In some embodiments, a BEE is capable of removing cytosine from DNA. In some embodiments, a BEE is capable of removing thymine from DNA. Exemplary BEEs include, without limitation, UDG Tyr147Ala and UDG Asn204Asp, as described by Sang et al., “A Unique Uracil-DNA binding protein of the uracil DNA glycosylase superfamily,” Nucleic Acids Research, Vol. 43, No. 17 2015, the entire contents of which are incorporated herein by reference.
[0069] The term “nucleic acid polymerase” or “NAP” refers to an enzyme that synthesizes nucleic acid molecules (e.g., DNA and RNA) from nucleotides (e.g., deoxyribonucleotides and ribonucleotides). In some aspects, NAP is a DNA polymerase. In some aspects, NAP is a damage-overcoming DNA polymerase. Damage-overcoming DNA polymerases play a role in mutagenesis, for example, by restarting replication forks or filling residual gaps in the genome due to the presence of DNA damage. Exemplary damage-overcoming polymerases include, but are not limited to, Pol beta, Pol lambda, Pol etha, Pol mu, Pol iota, Pol kappa, Pol alpha, Pol delta, Pol gamma, and Pol nu.
[0070] The term “nuclear localization sequence” or “NLS” refers to an amino acid sequence that facilitates the translocation of proteins into the cell nucleus, for example, by nuclear transport. Nuclear localization sequences are known in the art and will be apparent to those skilled in the art. In some embodiments, the NLS is a monopartite NLS. In some embodiments, the NLS is a bipartite NLS. Bipartite NLSs are separated by relatively short spacer sequences (e.g., 2–20 amino acids, 5–15 amino acids, or 8–12 amino acids). For example, the NLS sequence is described in Plank et al.'s international PCT application PCT / EP2000 / 011690, filed November 23, 2000, and published May 31, 2001, as WO / 2001 / 038547, and in Kaethar, KMV et al.'s “Application of bioinformatics-coupled experimental analysis reveals a new transport-competent nuclear localization signal in the nucleoptotein of Influenza A virus strain,” BMC Cell Biol, 2008, 9: 22 (the contents of each are incorporated herein by reference to their disclosures of exemplary nuclear localization sequences). In some embodiments, NLS comprises the amino acid sequence PKKKRKV (SEQ ID NO: 41), MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 42), KRTADGSEFESPKKKRKV (SEQ ID NO: 43), KRGINDRNFWRGENGRKTR (SEQ ID NO: 44), KKTGGPIYRRVDGKWRR (SEQ ID NO: 45), RRELILYDKEEIRRIWR (SEQ ID NO: 46), or AVSRKRKA (SEQ ID NO: 47).
[0071] The term “nucleic acid programmable DNA-binding protein” or “napDNAbp” refers to a protein that binds to a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid, which guides the napDNAbp to a specific nucleic acid sequence. For example, the Cas9 protein can bind to a guide RNA that guides the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, napDNAbp is a Class II microbial CRISPR-Cas effector. In some embodiments, napDNAbp is a Cas9 domain, e.g., nuclease-active Cas9, Cas9 niccas (nCas9), or nuclease-inactive Cas9 (dCas9). Examples of nucleic acid programmable DNA-binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpf1, C2c1, C2c2, C2C3, and Argonaut. However, it should be understood that nucleic acid programmable DNA-binding proteins also include nucleic acid programmable proteins that bind to RNA. For example, napDNAbp may be associated with nucleic acids that induce napDNAbp to bind to RNA. Other nucleic acid programmable DNA-binding proteins are also within the scope of this disclosure, even if they are not specifically mentioned herein.
[0072] The term "Cas9" or "Cas9 domain" refers to an RNA-inducible nuclease containing the Cas9 protein or its fragments (e.g., a protein containing the DNA cleavage domain of the active, inactive, or partially active form of Cas9, and / or the gRNA-binding domain of Cas9). Cas9 nucleases are also sometimes referred to as cason1 nucleases or CRISPR (clustered regularly interspaced short palindromic repeat)-associated nucleases. CRISPR is an adaptive immune system that provides defense against mobile genetic elements (viruses, translocation elements, and mating plasmids). A CRISPR cluster contains a spacer, a sequence complementary to the antecedent mobile element, and target invading nucleic acids. The CRISPR cluster is transcribed and processed to become CRISPR RNA (crRNA). In the type II CRISPR system, the correct processing of pre-crRNA requires trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. tracrRNA acts as a guide for ribonuclease 3-assisted pre-crRNA processing. Cas9 / crRNA / tracrRNA then endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. Target strands not complementary to the crRNA are first endonucleolytically cleaved and then 3'-5' exonucleolytically trimmed. In nature, DNA binding and DNA cleavage typically require proteins and both RNAs. However, a single guide RNA ("sgRNA" or simply "gNRA") can be manipulated to incorporate both crRNA and tracrRNA into a single RNA species.For example, see Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012) (the entire content of which is incorporated herein by reference). Cas9 helps distinguish between self and non-self by recognizing short motifs (PAM or protospacer-adjacent motifs) in CRISPR repeat sequences.Cas9 nuclease sequences and structures are well known to those skilled in the art (e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001);“CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma CM., Gonzales K., Chao See Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I, Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012). The entire contents of each of these are incorporated herein by reference. Cas9 orthologs have been described in a variety of species, including but not limited to S. pyogenes and S. thermophilus. Additional preferred Cas9 nucleases and sequences will be apparent to those skilled in the art based on this disclosure.Such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immune systems” (2013) RNA Biology 10:5, 726-737 (the entire contents of which are incorporated herein by reference). In some embodiments, Cas9 nucleases have an inactive (e.g., inactivating) DNA cleavage domain, i.e., Cas9 is a nickase.
[0073] Nuclease-inactivated Cas9 proteins can interchangeably be referred to as "dCas9" proteins (nuclease-inactive Cas9). Methods for generating Cas9 proteins (or fragments thereof) with inactive DNA cleavage domains are known (see, e.g., Jinek et al., Science. 337:816-821 (2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression" (2013) Cell. 28; 152(5): 1173-83, the entire contents of each of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to contain two subdomains: an HNH nuclease subdomain and a RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can suppress the nuclease activity of Cas9. For example, mutations D10A and H840A completely inactivate the nuclease activity of Cas9 in S. pyogenes (Jinek et al., Science. 337:816-821(2012); Qi et al., Cell. 28; 152(5): 1173-83 (2013)). In some embodiments, proteins containing fragments of Cas9 are provided. For example, in some embodiments, the protein contains one of two Cas9 domains: (1) the gRNA-binding domain of Cas9; or (2) the DNA-cleaving domain of Cas9. In some embodiments, proteins containing Cas9 or its fragments are called "Cas9 variants." Cas9 variants share homology with Cas9 or its fragments. For example, a Cas9 variant is at least approximately 70% identical, at least approximately 80% identical, at least approximately 90% identical, at least approximately 95% identical, at least approximately 96% identical, at least approximately 97% identical, at least approximately 98% identical, at least approximately 99% identical, at least approximately 99.5% identical, or at least approximately 99.9% identical to wild-type Cas9.In some aspects, Cas9 variants may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type Cas9. In some embodiments, a Cas9 variant comprises a fragment of Cas9 (e.g., a gRNA-binding domain or a DNA-cleaving domain) the fragment is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid length of the corresponding wild-type Cas9.
[0074] In some embodiments, the fragment is at least 100 amino acids long. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or 1300 amino acids long. In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_017053.1, SEQ ID NO: 1 (nucleotide); SEQ ID NO: 4 (amino acid)). [ka] [ka] [ka] [ka] [ka]
[0075] In some embodiments, wild-type Cas9 corresponds to or includes SEQ ID NO: 2 (nucleotide) and / or SEQ ID NO: 5 (amino acid). [ka] [ka] [ka] [ka] [ka]
[0076] In some embodiments, wild-type Cas9 corresponds to Cas9 derived from Streptococcus pyogenes (NCBI reference sequence: NC_002737.2, SEQ ID NO: 3 (nucleotide); and uniport reference sequence: Q99ZW2, SEQ ID NO: 6 (amino acid)). [ka] [ka] [ka] [ka] [ka]
[0077] In some embodiments, Cas9 is: Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1);Prevotella intermedia(NCBI Ref: NC_017861.1);Spiroplasma taiwanense(NCBI Ref: NC_021846.1);Streptococcus iniae(NCBI Ref: NC_021314.1);Belliella baltica(NCBI Ref: NC_018010.1);Psychroflexus torquisl(NCBI Ref: NC_018721.1); This refers to Cas9 from Streptococcus thermophilus (NCBI Ref: YP_820832.1), Listeria innocua (NCBI Ref: NP_472073.1), Campylobacter jejuni (NCBI Ref: YP_002344900.1), or Neisseria meningitidis (NCBI Ref: YP_002342100.1), or any other organism.
[0078] In some embodiments, dCas9 corresponds to, or comprises part or all of, a Cas9 amino acid sequence having one or more mutations that inactivate Cas9 nuclease activity. For example, in some embodiments, the dCas9 domain comprises the D10 and H840A mutations of SEQ ID NO: 6 or the corresponding mutations in another Cas9. In some embodiments, dCas9 comprises the amino acids of SEQ ID NO: 7 (D10 and H840A). [ka]
[0079] In some embodiments, the Cas9 domain contains the D10A mutation, while the residue at the corresponding position, such as position 840 in the amino acid sequence provided by SEQ ID NO: 6, or Cas9 described in any amino acid sequence provided by SEQ ID NOs: 4-26, remains histidine. While we do not wish to be constrained by any particular theory, the presence of the H840 catalytic residue maintains the activity of Cas9 to cleave an unedited (e.g., undeaminated) strand containing the opposite T to target A. Restoration of H840 (e.g., from A840 in dCas9) does not result in the cleavage of the A-containing target strand. Such Cas9 variants have the ability to generate single-strand DNA breaks (nicks) at specific locations based on the target sequence defined by the gRNA, leading to the repair of the unedited strand and ultimately resulting in the T-to-C change of the unedited strand.
[0080] In other embodiments, dCas9 variants having mutations other than D10A and H840A are provided, which result in, for example, nuclease-inactivating Cas9 (dCas9). Such mutations include, for example, other amino acid substitutions in D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or RuvC1 subdomain). In some embodiments, variants or homologs of dCas9 (e.g., variants of SEQ ID NOs. 6, 7, 8, 9, or 22) are provided, which are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to SEQ ID NOs. In some embodiments, variants of dCas9 (e.g., variants of SEQ ID NOs. 6, 7, 8, 9, or 22) are provided that have an amino acid sequence that is about 5, 10, 15, 20, 25, 30, 40, 50, 75, 100, or more amino acids shorter or longer than SEQ ID NOs. 7, 8, 9, or 22.
[0081] In some embodiments, the Cas9 fusion proteins provided herein include the full-length amino acid sequence of the Cas9 protein, for example, one of the Cas9 sequences provided herein. In other embodiments, however, the fusion proteins provided herein include only a fragment of the Cas9 sequence, rather than the full-length sequence. For example, in some embodiments, the Cas9 fusion proteins provided herein include a Cas9 fragment that binds to crRNA and tracrRNA or sgRNA, but does not include a functional nuclease domain, for example, that it includes only a shortened version of the nuclease domain or does not include a nuclease domain at all.
[0082] Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and additional suitable sequences of Cas9 domains and fragments will be apparent to those skilled in the art.
[0083] In some embodiments, Cas9 is: Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1);Prevotella intermedia(NCBI Ref: NC_017861.1);Spiroplasma taiwanense(NCBI Ref: NC_021846.1);Streptococcus iniae(NCBI Ref: NC_021314.1);Belliella baltica(NCBI Ref: NC_018010.1);Psychroflexus torquisI(NCBI Ref: This refers to Cas9 from NC_018721.1);Streptococcus thermophilus (NCBI Ref: YP_820832.1);Listeria innocua (NCBI Ref: NP_472073.1);Campylobacter jejuni (NCBI Ref: YP_002344900.1); or Neisseria meningitidis (NCBI Ref: YP_002342100.1).
[0084] It should be understood that additional Cas9 proteins (e.g., nuclease-dead Cas9 (dCas9), Cas9 nickers (nCas9), or nuclease-active Cas9), encompassing these variants and homologs, are within the scope of this disclosure. Exemplary Cas9 proteins include, without limitation, those provided below. In some embodiments, the Cas9 protein is nuclease-dead Cas9 (dCas9). In some embodiments, dCas9 comprises an amino acid sequence (SEQ ID NOs: 7, 8, 9, or 22). In some embodiments, the Cas9 protein is Cas9 nickers (nCas9). In some embodiments, nCas9 comprises an amino acid sequence (SEQ ID NOs: 10, 13, 16, or 21). In some embodiments, the Cas9 protein is nuclease-active Cas9. In some embodiments, the nuclease-active Cas9 comprises an amino acid sequence (SEQ ID NOs: 4, 5, 6, 11, 12, 14, 15, 16, 17, 18, 19, 20, 23, 24, 25, or 26). Exemplary catalyst-inactive Cas9 (dCas9): [ka] [ka] Example Cas9 nickase (nCas9): [ka] Exemplary catalyst-inactive Cas9 (dCas9): [ka]
[0085] As used herein, the term “Cas9 nickase” refers to the Cas9 protein capable of cleaving only one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, Cas9 nickase contains the D10A mutation and has the corresponding mutation in any Cas9 provided in SEQ ID NO: 6 (H840 histidine) or any one of SEQ ID NOs. 4-26, for example, Cas9 nickase may contain the amino acid sequences listed in SEQ ID NOs. 10, 13, 16, or 21. Such Cas9 nickase has an activated HNH nuclease domain and is capable of cleaving the non-target strand of DNA, i.e., the strand bound by gRNA. Furthermore, such Cas9 nickase has an inactivated RuvC nuclease domain and is not capable of cleaving the target strand of DNA, i.e., the strand to which base editing is desired.
[0086] In some embodiments, Cas9 refers to Cas9 derived from archaea (e.g., nanoarchaeons) that constitute the domains and kingdoms of unicellular prokaryotic microorganisms. In some embodiments, Cas9 refers to CasX or CasY, as described, for example, in Burstein et al., “New CRISPR-Cas systems from uncultivated microbes.” Cell Res. 2017 Feb 21. doi: 10.1038 / cr.2017.21 (the entire content is incorporated here by reference). Using genomic analysis and metagenomics, numerous CRISPR-Cas systems have been identified, including Cas9 first reported in the life of the archaeal domain. This diverse Cas9 protein has been found in largely unstudied nanoarchaeons as part of an active CRISPR-Cas system. In bacteria, two previously unknown systems, CRISPR-CasX and CRISPR-CasY, have been discovered, and these are the most compact among those found to date. In some embodiments, Cas9 refers to CasX or a variant of CasX. In some embodiments, Cas9 refers to CasY or a variant of CasY. It should be understood that other RNA-induced DNA-binding proteins may be used as nucleic acid-programmable DNA-binding proteins (napDNAbp), and that this is within the scope of this disclosure.
[0087] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) of any fusion protein provided herein may be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the CasX protein is a nuclease-inactive CasX protein (dCasX), a CasX nickasase (CaxXn), or a nuclease-active CasX. In some embodiments, the napDNAbp is a CasY protein. In some embodiments, the CasY protein is a nuclease-inactive CasY protein (dCasY), a CasY nickasase (CaxYn), or a nuclease-active CasY. In some embodiments, napDNAbp contains an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring CasX or CasY protein. In some embodiments, napDNAbp is a naturally occurring CasX or CasY protein. In some embodiments, napDNAbp contains an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of SEQ ID NOs 27-29. In some embodiments, napDNAbp contains an amino acid sequence of any one of SEQ ID NOs 27-29. It should be understood that CasX and CasY from other bacterial species may be used in accordance with this disclosure. [ka] [ka] [ka] [ka]
[0088] As used herein, the term “effective amount” means the amount of a bioactive substance sufficient to induce a desired biological response. For example, in some embodiments, the effective amount of a nucleic acid base editing factor may mean the amount of the nucleic acid base editing factor sufficient to induce a mutation at a target site to which the nucleic acid base editing factor specifically binds. In some embodiments, the effective amount of a fusion protein provided herein, for example, a fusion protein comprising a nucleic acid programmable DNA-binding protein and a deaminase domain (e.g., an adenosine deaminase domain), may mean the amount of the fusion protein sufficient to induce editing at a target site to which the fusion protein specifically binds and edits. As will be understood by those skilled in the art, the effective amount of a substance, such as a fusion protein, nucleic acid base editing factor, deaminase, hybrid protein, protein dimer, complex of a protein (or protein dimer) with a polynucleotide, or polynucleotide, may vary depending on various factors, such as the desired biological response, e.g., the specific allele, genome, or target site to be edited, the cell or tissue to be targeted, and the substance to be used.
[0089] As used herein, the terms “nucleic acid” and “nucleic acid molecule” refer to compounds comprising nucleic acid bases and acidic moieties, such as nucleosides, nucleotides, or polymers of nucleotides. Typically, polymeric nucleic acids, such as nucleic acid molecules containing three or more nucleotides, are linear molecules in which adjacent nucleotides are linked apart by phosphodiester bonds. In some embodiments, “nucleic acid” refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides). In some embodiments, “nucleic acid” refers to oligonucleotide chains containing three or more individual nucleotide residues. As used herein, the terms “oligonucleotide” and “polynucleotide” may be used interchangeably to refer to polymers of nucleotides (e.g., a chain of at least three nucleotides). In some embodiments, “nucleic acid” encompasses RNA and single- and / or double-stranded DNA. Nucleic acids may exist naturally in the context of genomes, transcripts, mRNA, tRNA, rRNA, siRNA, snRNA, plasmids, cosmids, chromosomes, chromatids, or other naturally occurring nucleic acid molecules. On the other hand, nucleic acid molecules may include molecules not found in nature, such as recombinant DNA or RNA, artificial chromosomes, engineered genomes or their fragments, or synthetic DNA, RNA, DNA / RNA hybrids, or nucleotides or nucleosides not found in nature. Furthermore, the terms “nucleic acid,” “DNA,” “RNA,” and / or similar terms include nucleic acid analogs, such as analogs having something other than a phosphodiester backbone. Nucleic acids may be purified from natural sources, produced and optionally purified using recombinant expression systems, or chemically synthesized. Where appropriate, for example in the case of chemically synthesized molecules, nucleic acids may include nucleoside analogs, such as analogs having chemically modified bases or sugars, and backbone modifications. Nucleic acid sequences are presented in the 5' to 3' direction unless otherwise indicated.In some aspects, nucleic acids are natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyluridine, C5-propynylcytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine); chemically modified bases; biologically modified bases (e.g., methylated bases); intercalated bases; modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioates and 5'-N-phosphoamidite bonds), or comprising them.
[0090] As used herein, the term “proliferative disorder” refers to any disease in which cellular or tissue homeostasis is disrupted in that cells or groups of cells exhibit an abnormally high rate of proliferation. Proliferative disorders include hyperproliferative disorders, such as preneoplastic hyperplasia and neoplastic disorders. Neoplastic disorders are characterized by abnormal cell proliferation and include both benign and malignant neoplasms. Malignant neoplasms are also known as cancer.
[0091] The terms “protein,” “peptide,” and “polypeptide” are used interchangeably herein and refer to polymers of amino acid residues linked together by peptide (amide) bonds. The terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide will be at least three amino acids long. A protein, peptide, or polypeptide may refer to an individual protein or a collection of proteins. One or more amino acids in a protein, peptide, or polypeptide may be modified, for example, by the addition of chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups, conjugations, functionalization, or linkers for other modifications. A protein, peptide, or polypeptide may also be a monomolecule or a multimolecule complex. A protein, peptide, or polypeptide may simply be a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be naturally occurring, recombinant, synthetic, or a combination thereof.
[0092] As used herein, the term “fusion protein” refers to a hybrid polypeptide comprising protein domains from at least two different proteins. One protein may be located at the amino-terminal (N-terminal) or carboxy-terminal (C-terminal) portion of the fusion protein, thus forming an “amino-terminal fusion protein” or a “carboxy-terminal fusion protein,” respectively. The protein may comprise different domains, e.g., a nucleic acid-binding domain of a nucleic acid editing protein (e.g., the gRNA-binding domain of Cas9 that leads to the binding of the protein to a target site) and a nucleic acid-cleaving domain or catalytic domain. In some embodiments, the protein comprises a proteinaceous portion, e.g., an amino acid sequence constituting a nucleic acid-binding domain, and an organic compound, e.g., a compound that can act as a nucleic acid-cleaving agent. In some embodiments, the protein is in complex with or associated with a nucleic acid, e.g., RNA. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced by recombinant protein expression and purification, which is particularly suitable for fusion proteins containing a peptide linker. Methods for recombinant protein expression and purification are well known and include those described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the entire contents of which are incorporated herein by reference.
[0093] The terms “RNA-programmable nuclease” and “RNA-induced nuclease” are used interchangeably herein and refer to a nuclease that forms a complex with (e.g., binds to or associates with) one or more RNAs that are not the target of cleavage. In some embodiments, an RNA-programmable nuclease may be called a nuclease:RNA complex when it is a complex with RNA. Typically, the bound RNA(s) are called guide RNAs (gRNAs). gRNAs may exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA existing as a single RNA molecule may be called a single guide RNA (sgRNA), but “gRNA” is used interchangeably to refer to a guide RNA that exists either as a single molecule or as a complex of two or more molecules. Typically, a gRNA existing as a single RNA species contains two domains: (1) a domain that has homology to the target nucleic acid (e.g., and results in the binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to the tracrRNA provided by Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those containing domain 2) can be found in U.S. Provisional Patent Application USSN61 / 874,682 filed September 6, 2013, entitled "Switchable Cas9 Nucleases And Uses Thereof" and U.S. Provisional Patent Application USSN61 / 874,746 filed September 6, 2013, entitled "Delivery System For Functional Nucleases," the entire contents of which are incorporated herein by reference. In some embodiments, a gRNA may contain two or more domains (1) and (2) and may be referred to as an "extended gRNA."For example, an extended gRNA may bind to, for instance, two or more Cas9 proteins and bind to a target nucleic acid in two or more separate regions as described herein. The gRNA contains a nucleotide sequence complementary to the target, which mediates the binding of the nuclease / RNA complex to the target site and provides sequence specificity for the nuclease:RNA complex.Additionally, RNA-induced RNA-induced polymerization (CRIS PR vaccine) Cas9 strain strain Streptococcus pyogenes and Cas9(Csn1)(Refer to “Complete Genome Sequence of an M1 Strain of Streptococcus pyogenes.” Ferretti JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G, Lyon K, Primeaux C, Sezate S, Suvorov AN, 2005; Kenton S, Lai HS, Lin SP, Qian Y, Jia HG, Ren Q, Zhu H, Song L, White J, Roe BA, McLaughlin RE, Proc RNA and host factor RNase III." Deltcheva E, Chylinski K, Sharma CM, Gonzales K, Chao Y, Pirzada ZA, Eckert MR, Vogel J, Charpentier E, Nature 471 :602–607(2011);および"A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M, Chylinski K, Fonfara I, Hauer M, Doudna JA, Charpentier E. Science 337:816-821(2012) of which the chemical components of this material are prepared in a stable manner.
[0094] RNA-programmable nucleases (e.g., Cas9) perform RNA:DNA hybridization to target DNA cleavage sites; therefore, in principle, these proteins have the ability to target any sequence defined by a guide RNA. Methods using RNA-programmable nucleases, such as Cas9, for site-specific cleavage (e.g., to modify the genome) are well known in this field (e.g., Cong, L. et al., Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas 9. Science 339, 823-826 (2013); Hwang, WY et al., Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic See acids research (2013); Jiang, W. et al., RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature biotechnology 31, 233-239 (2013); the entire contents of each are incorporated herein by reference.
[0095] As used herein, the term "subject" refers to an individual organism, such as an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, goat, cow, cat, or dog. In some embodiments, the subject is a vertebrate, amphibian, reptile, fish, insect, fly, or nematode. In some embodiments, the subject is a research animal. In some embodiments, the subject is genetically engineered, such as a genetically engineered non-human subject. The subject may be of either sex and at any stage of development.
[0096] The term "target site" refers to a sequence in a nucleic acid molecule that is modified by a base editing factor, such as a fusion protein containing cytidine deaminase (e.g., the dCas9-cytidine deaminase fusion protein provided herein).
[0097] The terms “treatment,” “to treat,” and “to treat” refer to a clinical intervention aimed at reversing, reducing, delaying the onset of, or inhibiting the progression of a disease or abnormality or one or more of its symptoms as described herein. In some embodiments, a treatment may be administered after the onset of one or more symptoms and / or after the disease has been diagnosed. In other embodiments, a treatment may be administered in the absence of symptoms, for example, to prevent or delay the onset of symptoms, or to inhibit the onset or progression of the disease. For example, a treatment may be administered to a susceptible individual prior to the onset of symptoms (for example, in light of a history of symptoms and / or genetic or other susceptibility factors). Treatment may also be continued after the symptoms have resolved, for example, to prevent or delay their recurrence.
[0098] As used herein in the context of proteins or nucleic acids, the term “recombinant” means a protein or nucleic acid that does not exist in nature and is a product of human manipulation. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that contains at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations compared to any naturally occurring sequence.
[0099] Detailed description of the invention Nucleic acid programmable DNA-binding protein (napDNAbp) Several aspects of this disclosure provide nucleic acid-programmable DNA-binding proteins that can be used to guide proteins, such as base editing factors, to specific nucleic acid (e.g., DNA or RNA) sequences. Nucleic acid-programmable DNA-binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpf1, C2c1, C2c2, C2C3, and Argonaut. An example of a nucleic acid-programmable DNA-binding protein with distinct PAM specificity from Cas9 is clustered, regularly spaced short batch sequence repeats (Cpf1) derived from Prevotella and Francisella 1. Like Cas9, Cpf1 is also a class 2 CRISPR effector. Cpf1 has been shown to mediate robust DNA interference with features distinct from Cas9. Cpf1 is a monoRNA-inducible endonuclease lacking tracrRNA, utilizing T-rich protospacer flanking motifs (TTN, TTTN, or YTN). Furthermore, Cpf1 cleaves DNA via double-strand breaks of twisted DNA. Among the 16 Cpf1-family proteins, two enzymes derived from Acidaminococcus and Lachnospiraceae have been shown to possess efficient genome editing activity in human cells. The Cpf1 protein is publicly known in this field, for example, previously described in Yamano et al., “Crystal structure of Cpf1 in complex with guide RNA and target DNA.” Cell (165) 2016, pp. 949-962, the entire content of which is incorporated herein by reference.
[0100] A nuclease-inactive Cpf1 (dCpf1) variant, which can be used as a guide nucleotide sequence-programmable DNA-binding protein domain, is also useful in this composition and method. The Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9, but lacks an HNH endonuclease domain, and the N-terminus of Cpf1 does not have the alpha-helix recognition lobe of Cas9. It was shown in Zetsche et al., Cell, 163, 759-771, 2015 (incorporated herein by reference) that the RuvC-like domain of Cpf1 is responsible for cleaving both DNA strands and that inactivation of the RuvC-like domain inactivates the nuclease activity of Cpf1. For example, mutations corresponding to D917A, E1006A, or D1255A in Francisella novicida Cpf1 (SEQ ID NO: 30) inactivate the nuclease activity of Cpf1. In some embodiments, the dCpf1 of this disclosure includes mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A in SEQ ID NO: 30. Any mutation that inactivates the RuvC domain of Cpf1, such as substitutions, deletions, or insertions, may be used in accordance with this disclosure.
[0101] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) of any of the fusion proteins provided herein may be a Cpf1 protein. In some embodiments, the Cpf1 protein is Cpf1 nickase (nCpf1). In some embodiments, the Cpf1 protein is nuclease-inactive Cpf1 (dCpf1). In some embodiments, Cpf1, nCpf1, or dCpf1 contains an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least (at ease) 99.5% identical to any one of SEQ ID NOs. 30-37. In some embodiments, dCpf1 comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least (at ease) 99.5% identical to any one of SEQ ID NOs. 30, and comprises mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A in SEQ ID NOs.
[0102] Wild-type Francisella novicida Cpf1 (SEQ ID NO: 30) (D917, E1006, and D1255 are in bold and underlined)
[0103] [ka] [ka]
[0104] Francisella novicida Cpf1 D917A (Sequence ID 31) (A917, E1006, and D1255 are in bold and underlined)
[0105] [ka] [ka]
[0106] Francisella Novicida Cpf1 E1006A (Sequence ID 32) (D917, A1006, and D1255 are in bold and underlined)
[0107] [ka] [ka]
[0108] Francisella Novicida Cpf1 D1255A (Sequence ID 33) (D917, E1006, and A1255 are in bold and underlined)
[0109] [ka]
[0110] Francisella novicida Cpf1 D917A / E1006A / D1255A (Sequence ID 34) (A917, A1006, and A1255 are bold and underlined)
[0111] [ka] [ka]
[0112] Francisella novicida Cpf1 D917A / D1255A (Sequence ID 35) (A917, A1006, and A1255 are in bold and underlined)
[0113] [ka] [ka]
[0114] Francisella novicida Cpf1 E1006A / D1255A (Sequence ID 36) (D917, A1006, and A1255 are in bold and underlined)"
[0115] [ka] [ka]
[0116] Francisella Novicida Cpf1 D917A / E1006A / D1255A (Sequence ID 37) (A917, A1006, and A1255 are in bold and underlined)
[0117] [ka]
[0118] In some embodiments, nucleic acid programmable DNA-binding proteins (napDNAbp) are nucleic acid programmable DNA-binding proteins that do not require a classical (NGG)PAM sequence. In some embodiments, napDNAbp are Argonaut proteins. An example of such a nucleic acid programmable DNA-binding protein is the Argonaut protein (NgAgo) from Natronobacterium gregoryi. NgAgo is an ssDNA-inducible endonuclease. NgAgo binds to approximately 24 nucleotides of 5' phosphorylated ssDNA (gDNA), guides it to its target site, and will cause a DNA double-strand break at the gDNA site. In contrast to Cas9, the NgAgo-gDNA system does not require a protospacer adjacent motif (PAM). Using nuclease-inactive NgAgo (dNgAgo) can greatly expand the range of bases that can be targeted. The characterization and use of NgAgo are described in Gao et al., Nat Biotechnol., 2016 Jul;34(7):768-73. PubMed PMID: 27136078; Swarts et al., Nature. 507(7491) (2014):258-61; and Swarts et al., Nucleic Acids Res. 43(10) (2015):5120-9, which are incorporated herein by reference, respectively. The sequence of Natronobacterium gregoryi algonaut is provided by Sequence ID No. 38.
[0119] Wild-type Natronobacterium gregoryi Argonaut (SEQ ID NO: 38)
[0120] [ka]
[0121] In some embodiments, napDNAbp is a prokaryotic homolog of Argonaute proteins. Prokaryotic homologs of Argonaute proteins are known, for example, as described in Makarova K., et al., “Prolaryotic homologys of Argonaute proteins are predicted to function as key components of a novel system of defense against mobile genetic elements”, Biol Direct. 2009 Aug 25;4:29. doi: 10.1186 / 1745-6150-4-29, the entire contents of which are incorporated herein by reference. In some embodiments, napDNAbp is the Marinetoga piezophila Argonaute (MpAgo) protein. CRISPR-associated Marinetoga piezophila Argonaute (MpAgo) proteins cleave single-strand target sequences using a 5'-phosphorylation guide. The 5' guide is used by all known Argonaute proteins. The crystal structure of the MpAgo-RNA complex reveals a guide chain binding site containing residues that interfere with 5'-phosphate interactions. This data suggests the evolution of an Argonaut subclass with non-classical specificity for 5'-hydroxylation guides. See, for example, Kaya et al., “A bacterial argonaute with noncanonical guide RNA specificity”, Proc Natl Acad Sci US A. 2016 Apr 12;113(15):4057-62, the entire content of which is incorporated herein by reference. It should be understood that other Argonaut proteins may be used and are within the scope of this disclosure.
[0122] In some embodiments, nucleic acid programmable DNA-binding proteins (napDNAbp) are single effectors of the microbial CRISPR-Cas system. Single effectors of the microbial CRISPR-Cas system include, but are not limited to, Cas9, Cpf1, C2c1, C2c2, and C2c3. Typically, the microbial CRISPR-Cas system is divided into Class 1 and Class 2 systems. Class 1 systems have multi-subunit effector complexes, while Class 2 systems have single protein effectors. For example, Cas9 and Cpf1 are Class 2 effectors. In addition to Cas9 and Cpf1, three distinct Class 2 CRISPR-Cas systems (C2c1, C2c2, and C2c3) are described in Shmakov et al., “Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems”, Mol. Cell, 2015 Nov 5; 60(3): 385-397, the entire contents of which are incorporated herein by reference. The effectors of two systems, C2c1 and C2c3, contain a RuvC-like endonuclease domain related to Cpf1. The third system, C2c2, contains an effector based on two HEPN RNase domains. The production of mature CRISPR RNA is independent of tracrRNA, unlike the production of CRISPR RNA by C2c1. C2c1 relies on both CRISPR RNA and tracrRNA for DNA cleavage. Bacterial C2c2 has been shown to possess unique RNase activity for CRISPR RNA maturation, separate from its active RNA single-strand RNA degradation activity. These RNase functions differ from each other and from the CRISPR RNA processing behavior of CPf1.For example, see East-Seletsky, et al., “Two distinct RNase activities of CRISPR-C2c2 enable guide RNA processing and RNA detection”, Nature, 2016 Oct 13;538(7624):270-273, the entire content of which is incorporated herein by reference. In vitro biochemical analysis of C2c2 in Leptotrichia shahii shows that C2c2 is inducible by a single CRISPR RNA and can be programmed to cleave an ssRNA target that carries a complementary protospacer. Catalytic residues in two conserved HEPN domains mediate the cleavage. Mutations in the catalytic residues produce a catalytically inactive RNA-binding protein. For example, see Abudayyeh et al., “C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector”, Science, 2016 Aug 5; 353(6299), the entire contents of which are incorporated herein by reference.
[0123] The crystal structure of Alicyclobaccillus acidoterrastris C2c1 (AacC2c1) has been reported to be complex with a chimeric single-molecular guide RNA (sgRNA). See, for example, Liu et al., “C2c1-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism”, Mol. Cell, 2017 Jan 19;65(2):310-322, the entire content of which is incorporated herein by reference. The crystal structure has also been reported to be bound to target DNA as a triple complex in Alicyclobaccillus acidoterrastris C2c1. See, for example, Yang et al., “PAM-dependent Target DNA Recognition and Cleavage by C2C1 CRISPR-Cas endonuclease”, Cell, 2016 Dec 15;167(7):1814-1828, the entire content of which is incorporated herein by reference. On both target and non-target DNA strands, the catalytically capable conformation of AacC2c1 is captured and independently located within a single RuvC catalytic pocket, resulting in C2c1-mediated cleavage that occurs as a staggered 7-nucleotide cleavage of the target DNA. Structural comparisons between the C2c1 triple complex and its pre-identified Cas9 and Cpf1 counterparts explain the diversity of mechanisms employed by the CRISPR-Cas9 system.
[0124] In some embodiments, any fusion protein of the nucleic acid programmable DNA-binding protein (napDNAbp) provided herein may be a C2c1, C2c2, or C2c3 protein. In some embodiments, napDNAbp is a C2c1 protein. In some embodiments, napDNAbp is a C2c2 protein. In some embodiments, napDNAbp is a C2c3 protein. In some embodiments, napDNAbp contains an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least (at ease) 99.5% identical to a naturally occurring C2c1, C2c2, or C2c3 protein. In some embodiments, napDNAbp is a naturally occurring C2c1, C2c2, or C2c3 protein. In some embodiments, napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least (at ease) 99.5% identical to either SEQ ID NO: 39 or 40. It should be understood that C2c1, C2c2, or C2c3 from other bacterial species may also be used in accordance with this disclosure. [ka] [ka] [ka] [ka]
[0125] Cas9 domain of nuclear base editing factor In some aspects, the nucleic acid programmable DNA-binding protein (napDNAbp) is the Cas9 domain. Non-limiting, exemplary Cas9 domains are provided herein. The Cas9 domain may be a nuclease-active Cas9 domain, a nuclease-inactive Cas9 domain, or a Cas9 nickasase. In some embodiments, the Cas9 domain is a nuclease-active domain. For example, the Cas9 domain may be a Cas9 domain that cleaves both strands of a double-stranded nucleic acid (e.g., both strands of a double-stranded DNA molecule). In some embodiments, the Cas9 domain comprises any one amino acid sequence presented in Sequence IDs 4-29. In some embodiments, the Cas9 domain includes an amino acid sequence that is identical by at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% to one amino acid sequence presented in any of the Cas9 or SEQ ID NOs. In some embodiments, the Cas9 domain includes an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any one of the amino acid sequences presented in any of the Cas9s provided herein or in SEQ ID NOs: 4-29.In some embodiments, the Cas9 domain includes an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 adjacent amino acid residues that are identical to any Cas9 or any of the amino acid sequences presented in any of the sequences presented in any of the sequences provided herein or in Sequence IDs 4-29.
[0126] In some embodiments, the Cas9 domain is a nuclease-inactive domain (dCas). For example, the dCas9 domain may bind to a double-stranded nucleic acid molecule without cleaving either strand of the double-stranded nucleic acid molecule (e.g., via a gRNA molecule). In some embodiments, the nuclease-inactive dCas9 domain includes a corresponding mutation in any Cas9 provided herein, such as the D10X and H840X mutations of the amino acid sequence presented in SEQ ID NO: 6, or any one of the amino acid sequences provided in SEQ ID NOs: 4-26, where X is any change in amino acid. In some embodiments, the nuclease-inactive dCas9 domain includes a corresponding mutation in any Cas9 provided herein, such as the D10A and H840A mutations of the amino acid sequence presented in SEQ ID NO: 6, or any one of the amino acid sequences provided in SEQ ID NOs: 4-26. As one example, the nuclease-inactivating Cas9 domain includes the amino acid sequence presented in Sequence ID No. 9 (Cloning vector pPlatTET-gRNA2, Accession No. BAV54124).
[0127] [ka] [ka] See, for example, Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression.” Cell. 2013; 152(5):1173-83, the entire content of which is incorporated herein by reference).
[0128] Additional suitable nuclease-inactive dCas9 domains will be apparent to those skilled in the art based on the present disclosure and knowledge of the art, and are within the scope of this disclosure. Such additional exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A variant domains (see Prashant et al., Cas 9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013; 31(9): 833-838, the entire contents of which are incorporated herein by reference). In some embodiments, the dCas9 domain comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the dCas9 domains provided herein. In some embodiments, the Cas9 domain contains an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more mutations compared to any one of the amino acid sequences presented in SEQ ID NOs: 7, 8, 9, or 22.In some embodiments, the Cas9 domain includes an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 adjacent amino acid residues that are identical to any one amino acid sequence presented in SEQ ID NO: 7, 8, 9, or 22.
[0129] In some embodiments, the Cas9 domain is Cas9 nickase. Cas9 nickase may be a Cas9 protein capable of cleaving just one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, Cas9 nickase cleaves a target strand of the double-stranded nucleic acid molecule, meaning that Cas9 nickase cleaves the (complementary) strand which is a base pair of gRNA (e.g., sgRNA) bound to Cas9. In some embodiments, Cas9 nickase comprises the D10A mutation and has histidine at position 840 in SEQ ID NO: 6, or any mutation in Cas9 provided herein, such as any one of SEQ ID NOs: 4-26. For example, Cas9 nickase may comprise the amino acid sequences presented in SEQ ID NOs: 10, 13, 16, or 21. In some embodiments, Cas9 nickase cleaves the non-target, unedited base strand of a double-stranded nucleic acid molecule, meaning that Cas9 nickase cleaves the non-base-paired strand to the gRNA (e.g., sgRNA) bound to Cas9. In some embodiments, Cas9 nickase contains the H840A mutation and has the corresponding mutation in any Cas9 provided herein, such as an aspartic acid residue at position 10 of SEQ ID NO: 6, or one of SEQ ID NOs. 4-26. In some embodiments, Cas9 nickase contains an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the Cas9 nickases provided herein. Additional suitable Cas9 nickases will be apparent from the present disclosure and knowledge in the art, and are within the scope of the present disclosure.
[0130] Cas9 domain with reduced PAM exclusivity Several aspects of the disclosure provide Cas9 domains with different PAM specificities. Typically, Cas9 proteins such as Cas9 from S. pyogenes (SpCas9) require the classical NGGPAM sequence to bind to specific nucleic acid regions, where "N" in "NGG" is adenine (A), thymine (T), guanine (G), or cytosine (C), and G is guanine. This can limit the ability to edit desired bases within the genome. In some embodiments, to edit the fusion proteins provided herein, the bases need to be located in a precise location, for example, within four bases of a region approximately 15 bases upstream of the PAM (e.g., the "deamination window"). See Komor, AC, et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016), the entire contents of which are incorporated herein by reference. In some embodiments, the deamination window is within a 2, 3, 4, 5, 6, 7, 8, 9, or 10-base region. In some embodiments, the deamination window is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 bases upstream of the PAM. Thus, in some embodiments, any fusion protein provided herein may contain a Cas9 domain capable of binding to a nucleotide sequence that does not contain a classical (e.g., NGG) PAM sequence. Cas9 domains that bind to non-classical PAM sequences have been described in the art and will be apparent to those skilled in the art.For example, Cas9 domains that bind to non-classical PAM sequences are described in Kleinstiver, BP, et al., “Engineered CRISPR-Cas9 nucleases with altered PAM specificities” Nature 523, 481-485 (2015); and Kleinstiver, BP, et al., “Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition” Nature Biotechnology 33, 1293-1298 (2015); the entire contents of each are incorporated herein by reference.
[0131] In some embodiments, the Cas9 domain is the Cas9 domain from Staphylococcus aureus (SaCas9). In some embodiments, the SaCas9 domain is nuclease-active SaCas9, nuclease-inactive SaCas9 (SaCas9d), or SaCas9 nickasase (SaCas9n). In some embodiments, SaCas9 comprises the amino acid sequence SEQ ID NO: 12. In some embodiments, SaCas9 comprises the N579X mutation of SEQ ID NO: 12, or the corresponding mutation in any amino acid sequence provided in SEQ ID NO: 13 or 14, where X is any amino acid except N. In some embodiments, SaCas9 comprises the N579A mutation of SEQ ID NO: 12, or the corresponding mutation in any amino acid sequence provided in SEQ ID NO: 13 or 14.
[0132] In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain may bind to nucleic acid sequences having a non-classical PAM. In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain may bind to nucleic acid sequences having an NNGRRT PAM sequence, where N=A, T, C, or G, and R=A or G. In some embodiments, the SaCas9 domain includes one or more of the E781X, N967X, and R1014X mutations in SEQ ID NO: 12, or the corresponding mutation in any amino acid provided in SEQ ID NO: 13 or 14, where X is any amino acid. In some embodiments, the SaCas9 domain includes one or more of the E781K, N967K, and R1014H mutations in SEQ ID NO: 12, or the corresponding mutation in any amino acid provided in SEQ ID NO: 13 or 14. In some embodiments, the SaCas9 domain includes one of the E781K, N967K, or R1014H mutations in SEQ ID NO: 12, or a corresponding mutation in any amino acid provided in SEQ ID NO: 13 or 14.
[0133] In some embodiments, the Cas9 domain of any fusion protein provided herein comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of SEQ ID NOs: 12-14. In some embodiments, the Cas9 domain of any fusion protein provided herein comprises an amino acid sequence that is at least one of SEQ ID NOs: 12-14. In some embodiments, the Cas9 domain of any fusion protein provided herein consists of an amino acid sequence that is at least one of SEQ ID NOs: 12-14.
[0134] Exemplary SaCas9 sequence
[0135] [ka]
[0136] The underlined and bolded residue N579 in sequence number 12 can be mutated (for example to A579) to produce SaCas9 nickase.
[0137] Exemplary SaCas9n sequence
[0138] [ka]
[0139] The mutation at residue A579 in SEQ ID NO: 12, which can produce SaCas9 nickase, is N579 in SEQ ID NO: 13, and is therefore underlined and in bold.
[0140] Exemplary SaKKH Cas9
[0141] [ka] [ka]
[0142] The residue A579 in SEQ ID NO: 14, which can be mutated from residue A579 in SEQ ID NO: 12 to produce SaCas9 nickase, is underlined and in bold. The residues K781, K967, and H1014 in SEQ ID NO: 14, which can be mutated from E781, N967, and R1014 in SEQ ID NO: 12 to produce SaKKH Cas9, are underlined and in italics.
[0143] In some embodiments, the Cas9 domain is a Cas9 domain from Streptococcus Pyogenes (SpCas9). In some embodiments, the SpCas9 domain is nuclease-activated SpCas9, nuclease-inactive SpCas9 (SpCas9d), or SpCas9 nickase (SpCas9n). In some embodiments, SpCas9 comprises the amino acid sequence SEQ ID NO: 15. In some embodiments, SpCas9 comprises a corresponding mutation in any Cas9, such as the D9X mutation in SEQ ID NO: 15, or any amino acid sequence provided in SEQ ID NOs: 4-26, where X is any amino acid except D. In some embodiments, SpCas9 comprises a corresponding mutation in any Cas9 provided herein, such as the D9A mutation in SEQ ID NO: 15, or any amino acid sequence provided in SEQ ID NOs: 4-26. In some embodiments, the SpCas9 domain, SpCas9d domain, or SpCas9n domain may bind to nucleic acid sequences having a non-classical PAM. In some embodiments, the SpCas9 domain, SpCas9d domain, or SpCas9n domain may bind to a nucleic acid sequence having an NGG, NGA, or NGCG PAM sequence. In some embodiments, the SpCas9 domain includes one or more of the D1134X, R1334X, and T1336X mutations of SEQ ID NO: 15, or a corresponding mutation in any Cas9 provided herein, such as any amino acid sequence provided in SEQ ID NOs: 4-26, where X is any amino acid. In some embodiments, the SpCas9 domain includes one or more of the D1134E, R1334Q, and T1336R mutations of SEQ ID NO: 15, or a corresponding mutation in any Cas9 provided herein, such as any amino acid sequence provided in SEQ ID NOs: 4-26. In some embodiments, the SpCas9 domain includes one of the D1134E, one of the R1334Q, and one of the T1336R mutations in SEQ ID NO: 15, or any corresponding mutation in any Cas9 provided herein, such as any amino acid sequence provided in SEQ ID NOs: 4-26.In some embodiments, the SpCas9 domain includes one or more D1134X, R1334X, and T1336X mutations of SEQ ID NO: 15, or any amino acid sequence provided in SEQ ID NOs: 4-26, where X is any amino acid. In some embodiments, the SpCas9 domain includes one or more D1134V, R1334Q, and T1336R mutations of SEQ ID NO: 15, or any amino acid sequence provided in SEQ ID NOs: 4-26, where X is any amino acid. In some embodiments, the SpCas9 domain includes one of the D1134V, one of the R1334Q, and one of the T1336R mutations of SEQ ID NO: 15, or any amino acid sequence provided in SEQ ID NOs: 4-26, where X is any amino acid. In some embodiments, the SpCas9 domain includes one or more mutations in D1134X, G1217X, T1334X, and T1336X of SEQ ID NO: 15, or any amino acid sequence provided in SEQ ID NOs: 4-26, where X is any amino acid. In some embodiments, the SpCas9 domain includes one or more mutations in D1134V, G1217R, T1334Q, and T1336R of SEQ ID NO: 15, or any amino acid sequence provided in SEQ ID NOs: 4-26, where X is any amino acid. In some embodiments, the SpCas9 domain includes one mutation in D1134V, one mutation in G1217R, one mutation in T1334Q, and one mutation in T1336R of SEQ ID NO: 15, or any amino acid sequence provided in SEQ ID NOs: 4-26, where X is any amino acid.
[0144] In some embodiments, the Cas9 domain of any fusion protein provided herein comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of SEQ ID NOs: 15-19. In some embodiments, the Cas9 domain of any fusion protein provided herein comprises an amino acid sequence that is at least one of SEQ ID NOs: 15-19. In some embodiments, the Cas9 domain of any fusion protein provided herein consists of an amino acid sequence that is at least one of SEQ ID NOs: 15-19.
[0145] An example of SpCas9
[0146] [ka] [ka]
[0147] A typical SpCas9n
[0148] [ka] [ka]
[0149] Example SpEQRCas9
[0150] [ka] [ka]
[0151] E1134, Q1334, and R1336 in sequence number 17, which are mutations from D1134, R1334, and T1336 in sequence number 15 and can produce SpEQR Cas9, are underlined and in bold.
[0152] An example of SpVQRCas9
[0153] [ka] [ka]
[0154] V1134, Q1334, and R1336 in sequence number 18, which are mutations from D1134, R1334, and T1336 in sequence number 15 and can produce SpVQR Cas9, are underlined and in bold.
[0155] Example SpVRERCas9
[0156] [ka]
[0157] V1134, R1217, Q1334, and R1336 of sequence number 19, which are mutations from D1134, G1217, R1334, and T1336 of sequence number 15 and can produce SpVRER Cas9, are underlined and in bold.
[0158] High fidelity Cas9 domain Several aspects of this disclosure provide high-fidelity Cas9 domains of nucleic acid base editing factors provided herein. In some embodiments, a high-fidelity Cas9 domain is an engineered Cas9 domain that includes one or more mutations that reduce the electrostatic interaction between the Cas9 domain and the sugar-phosphate backbone of DNA compared to the corresponding wild-type Cas9 domain. While we do not wish to be bound by any particular theory, a high-fidelity Cas9 domain having reduced electrostatic interaction with the sugar-phosphate backbone of DNA has fewer off-target effects. In some embodiments, a Cas9 domain (e.g., a wild-type Cas9 domain) has one or more mutations that reduce the association between the Cas9 domain and the sugar-phosphate backbone of DNA. In some embodiments, the Cas9 domain includes one or more mutations that reduce the association between the Cas9 domain and the sugar-phosphate backbone of DNA by at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, or more.
[0159] In some embodiments, any Cas9 fusion protein provided herein includes one or more N497X, R661X, Q695X, and / or Q926X mutations of the amino acid sequence provided in SEQ ID NO: 6, or a corresponding mutation in any Cas9 provided herein, such as any amino acid sequence provided in SEQ ID NOs: 4-26, where X is any amino acid. In some embodiments, any Cas9 fusion protein provided herein includes one or more N497A, R661A, Q695A, and / or Q926A mutations of the amino acid sequence provided in SEQ ID NO: 6, or a corresponding mutation in any Cas9 provided herein, such as any amino acid sequence provided in SEQ ID NOs: 4-26. In some embodiments, the Cas9 domain includes a corresponding mutation in any Cas9 provided herein, such as the D10A mutation of the amino acid sequence provided in SEQ ID NO: 6, or any amino acid sequence provided in SEQ ID NOs: 4-26. In some embodiments, the Cas9 domain (e.g., any fusion protein provided herein) includes the amino acid sequence presented in SEQ ID NO: 20. In some embodiments, the Cas9 domain contains an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to SEQ ID NO: 20. Cas9 domains with high fidelity are known in the art and will be apparent to those skilled in the art.For example, high-fidelity Cas9 domains are described in Kleinstiver, BP, et al. “High-fidelity CRISPR-Cas 9 nucleases with no detectable genome-wide off-target effects.” Nature 529, 490-495 (2016); and Slaymaker, IM, et al. “Rationally engineered Cas nucleases with improved specificity.” Science 351, 84-88 (2015); the full contents of each are incorporated herein by reference.
[0160] For example, any base editing factor provided herein, such as any C→G base editing factor provided herein, may be converted to a high-fidelity base editing factor by modifying the Cas9 domain as described herein, thereby producing a high-fidelity base editing factor, such as a high-fidelity C→G base editing factor. In some embodiments, the high-fidelity Cas9 domain is a dCas9 domain. In some embodiments, the high-fidelity Cas9 domain is an nCas9 domain.
[0161] The mutation in Sequence ID No. 6 against Cas9 is shown in bold and underlined, indicating a high-fidelity Cas9 domain.
[0162] [ka] [ka]
[0163] This disclosure also provides fragments of napDNAbps, such as shortenings of any napDNAbps provided herein. In some embodiments, a napDNAbps is an N-terminal shortening, where one or more amino acids are missing from the N-terminus of a napDNAbps. In some embodiments, a napDNAbp is missing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids from the N-terminus of a napDNAbps. For example, the N-terminal shortening of napDNAbp may be any napDNAbp provided herein, such as one of the napDNAbp provided in any one of SEQ ID NOs: 4-40. In some embodiments, napDNAbp is a C-terminal shortening, where one or more amino acids are missing from the C-terminus of napDNAbp. In some embodiments, napDNAbp is missing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids from the C-terminus of napDNAbp. For example, the C-terminal shortening of napDNAbp may be any napDNAbp provided herein, such as the NAP provided in any one of SEQ ID NOs: 4-40.
[0164] In some embodiments, any napDNAbps provided herein has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid changes compared to any napDNAbps provided herein, such as any one of the napDNAbps provided in SEQ ID NOs.
[0165] Uracil-binding protein (UBP) A uracil-binding protein, or UBP, refers to a protein capable of binding to uracil. In some embodiments, a uracil-binding protein is a uracil-modifying enzyme. In some embodiments, a uracil-binding protein is a uracil base-removing enzyme. In some embodiments, a uracil-binding protein is a uracil DNA glucosylase (UDG). In some embodiments, a uracil-binding protein binds to uracil with an affinity that is at least 1%, 2%, 3%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or at least 95% of the affinity that wild-type UDG (e.g., human UDG) has for binding to uracil. In some aspects, uracil-binding proteins may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type uracil-binding proteins such as wild-type UDG (e.g., human UDG) that bind to uracil.
[0166] In some embodiments, UBP is a uracil-modifying enzyme. In some embodiments, UBP is a uracil base removal enzyme. In some embodiments, UBP is a uracil DNA glycosylase. In some embodiments, UBP is any uracil-binding protein provided herein. For example, UBP is UDG, UdgX, UdgX * It may be UdgX_On or SMUG1. In some embodiments, UBP comprises an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to a uracil-binding protein, uracil base removal enzyme, or uracil DNA glucosylase (UDG) enzyme. In some embodiments, UBP comprises an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to any uracil-binding protein provided herein, for example, any of the UBP and UBP variants provided below. In some embodiments, UBP comprises an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to any one of SEQ ID NOs: 48-53. In some embodiments, UBP comprises any one amino acid sequence of SEQ ID NOs. 48-53. In some embodiments, the uracil-binding protein has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid changes compared to any UBP provided herein, such as any one of SEQ ID NOs.
[0167] This disclosure also provides fragments of UBP, such as any shortening of UBP provided herein. In some embodiments, UBP is an N-terminal shortening, where one or more amino acids are missing from the N-terminus of UBP. In some embodiments, UBP is missing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids from the N-terminus of UBP. For example, the N-terminal shortening of UBP may be any UBP provided herein, such as one of the UBPs provided in any one of SEQ ID NOs. 48-53. In some embodiments, UBP is a C-terminal shortening, where one or more amino acids are missing from the C-terminus of UBP. In some embodiments, UBP is missing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids from the C-terminus of UBP. For example, the C-terminal shortening of a UBP may be a shortening of the C-terminus of any UBP provided herein, such as any one of the UBPs provided in any one of sequence numbers 48-53.
[0168] Other UBPs will be apparent to those skilled in the art and should be understood to be within the scope of this disclosure. For example, UBPs were previously described in Sang et al., “A Unique uracil-DNA binding protein of the uracil DNA glycosylase superfamily,” Nucleic Acids Research, Vol. 43, No. 17 2015, the entirety of which is incorporated herein by reference. UDG
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
[0171] This disclosure also provides fragments of NAP, such as any shortening of NAP provided herein. In some embodiments, NAP is an N-terminal shortening, where one or more amino acids are missing from the N-terminus of NAP. In some embodiments, NAP is missing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids from the N-terminus of NAP. For example, the N-terminal shortening of NAP may be any NAP provided herein, such as one of the NAPs provided in any one of SEQ ID NOs. 54-64. In some embodiments, NAP is a C-terminal shortening, where one or more amino acids are missing from the C-terminus of NAP. In some embodiments, NAP is missing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids from the C-terminus of NAP. For example, the C-terminal shortening of NAP may be any NAP provided herein, such as one of the NAPs provided in any one of SEQ ID NOs. 54 to 64. Polbeta [ka] [ka] Pol Lambda [ka] Pol Eta [ka] Pol mu
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
Chem.
[0172] Base excision enzyme (BEE) Base removal enzymes, or BEEs, refer to proteins capable of removing bases (e.g., A, T, C, G, or U) from nucleic acid molecules (e.g., DNA or RNA). In some embodiments, BEEs are capable of removing cytosine from DNA. In some embodiments, BEEs are capable of removing thymine from DNA. Exemplary BEEs include, without limitation, UDG Tyr147Ala and UDG Asn204Asp, as described in Sang et al., “A Unique uracil-DNA binding protein of the uracil DNA glycosylase superfamily,” Nucleic Acids Research, Vol. 43, No. 17 2015; the entire contents thereof are incorporated herein by reference.
[0173] In some embodiments, the base recapsulation enzyme (BEE) is a cytosine, thymine, adenine, guanine, or uracil base recapsulation enzyme. In some embodiments, the base recapsulation enzyme (BEE) is a cytosine base recapsulation enzyme. In some embodiments, the BEE is a thymine base recapsulation enzyme. In some embodiments, the base recapsulation enzyme comprises an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to a naturally occurring BEE. In some embodiments, the base recapsulation enzyme comprises an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to any BEE provided herein, for example, UDG(Tyr147Ala) or UDG(Asn204Asp) below. In some embodiments, the base removal enzyme contains an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to any one of sequence numbers 65-66. In some embodiments, the base removal enzyme has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid changes compared to any BEE provided herein, such as any one of SEQ ID NOs.
[0174] The disclosure also provides fragments of BEE, such as any shortening of BEE provided herein. In some embodiments, BEE is an N-terminal shortening, where one or more amino acids are missing from the N-terminus of BEE. In some embodiments, BEE is missing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids from the N-terminus of BEE. For example, the N-terminal shortening of BEE may be any N-terminal shortening of BEE provided herein, such as any one of the BEEs provided in any one of SEQ ID NOs.65-66. In some embodiments, BEE is a C-terminal shortening, where one or more amino acids are missing from the C-terminus of BEE. In some embodiments, BEE is missing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids from the C-terminus of BEE. For example, the C-terminal shortening of BEE may be any C-terminal shortening of BEE provided herein, such as any one of the BEEs provided in any one of SEQ ID NOs. 65 to 66.
[0175] Other BEEs will be obvious to those skilled in the art and should be understood to be within the scope of this disclosure. For example, a BEE is described previously in Sang et al., “A Unique uracil-DNA binding protein of the uracil DNA glycosylase superfamily,” Nucleic Acids Research, Vol. 43, No. 17 2015; its entirety is incorporated herein by reference. UDG(Tyr147Ala)- Mutated residues are indicated by bold and underlined text. [ka] UDG(Asn204Asp)-mutated residues are indicated by bold and underlined text. [ka]
[0176] Deaminase Domain In some aspects, any fusion protein or base editing factor provided herein comprises a cytidine deaminase domain. In some aspects, the cytidine deaminase domain can catalyze a base change from C to U. In some aspects, the cytidine deaminase domain is an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some aspects, the cytidine deaminase domain is APOBEC1 deaminase. In some aspects, the cytidine deaminase domain is APOBEC2 deaminase. In some aspects, the cytidine deaminase domain is APOBEC3 deaminase. In some aspects, the cytidine deaminase domain is APOBEC3A deaminase. In some aspects, the cytidine deaminase domain is APOBEC3B deaminase. In some aspects, the cytidine deaminase domain is APOBEC3C deaminase. In some aspects, the cytidine deaminase domain is APOBEC3D deaminase. In some aspects, the cytidine deaminase domain is APOBEC3E deaminase. In some aspects, the cytidine deaminase domain is APOBEC3F deaminase. In some aspects, the cytidine deaminase domain is APOBEC3G deaminase. In some aspects, the cytidine deaminase domain is APOBEC3H deaminase. In some aspects, the cytidine deaminase domain is APOBEC4 deaminase. In some aspects, the cytidine deaminase domain is activation-inducing deaminase (AID). In some aspects, the cytidine deaminase domain is a vertebrate deaminase. In some aspects, the cytidine deaminase domain is a non-vertebrate deaminase. In some aspects, the cytidine deaminase domain is a human, chimpanzee, gorilla, monkey, cattle, dog, rat, and mouse deaminase. In some embodiments, the cytidine deaminase domain is a starfish deaminase.In some embodiments, the cytidine deaminase domain is a rat deaminase, e.g., rAPOBEC1. In some embodiments, the cytidine deaminase domain is Petromyzon marinus cytidine deaminase 1 (pmCDA1). In some embodiments, the cytidine deaminase domain is human APOBEC3G (SEQ ID NO: 77). In some embodiments, the cytidine deaminase domain is a fragment of human APOBEC3G (SEQ ID NO: 100). In some embodiments, the cytidine deaminase domain is a variant of human APOBEC3G containing the D316R_D317R mutation (SEQ ID NO: 99). In some embodiments, the cytidine deaminase domain is a fragment of human APOBEC3G and contains a mutation corresponding to the D316R_D317R mutation in SEQ ID NO: 77 (SEQ ID NO: 101).
[0177] In some embodiments, the cytidine deaminase domain is at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring cytidine deaminase. In some embodiments, the cytidine deaminase domain is at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any cytidine deaminase provided herein. In some embodiments, the cytidine deaminase domain is at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the deaminase domains of SEQ ID NOs. 67-101. In some embodiments, the nucleic acid editing domain is a nucleic acid editing domain comprising any one amino acid sequence of SEQ ID NOs. 67-101. In some embodiments, the cytidine deaminase domain has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid changes compared to any cytidine deaminase domain provided herein, such as any one of SEQ ID NOs. 67-101.
[0178] This disclosure also provides fragments of the cytidine deaminase domain, such as any shortening of the cytidine deaminase domain provided herein. In some embodiments, the cytidine deaminase domain is N-terminal shortening, where one or more amino acids are missing from the N-terminus of the cytidine deaminase domain. In some embodiments, the cytidine deaminase domain is missing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids from the N-terminus of the cytidine deaminase domain. For example, the N-terminal shortening of a cytidine deaminase domain may be any cytidine deaminase domain provided herein, such as one of the cytidine deaminase domains provided in any one of SEQ ID NOs. 67-101. In some embodiments, the cytidine deaminase domain is a C-terminal shortening, where one or more amino acids are missing from the C-terminus of the cytidine deaminase domain. In some embodiments, a cytidine deaminase domain is delimited by the absence of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acids from the C-terminus of the cytidine deaminase domain. For example, the C-terminal shortening of a cytidine deaminase domain may be any cytidine deaminase domain provided herein, such as any one of the cytidine deaminase domains provided in any one of Sequence IDs 67-101.
[0179] Some exemplary cytidine deaminase domains include, but are not limited to, those provided below. It should be understood that in some embodiments, the active domain of each sequence may be used, for example, a domain without a localization signal (nuclear localization sequence, no nuclear export signal, no cytoplasmic localization signal).
[0180] Human AID: [ka]
[0181] Mouse AI: [ka]
[0182] Dog AID: [ka]
[0183] Bovine AID: [ka]
[0184] Rat AID [ka]
[0185] Mouse APOBEC-3: [ka]
[0186] Rat APOBEC-3: [ka] [ka]
[0187] Rhesus macaque APOBEC-3G: [ka]
[0188] Chimpanzee APOBEC-3G: [ka]
[0189] Green monkey APOBEC-3G: [ka]
[0190] Human APOBEC-3G: [ka] [ka]
[0191] Human APOBEC-3F: [ka]
[0192] Human APOBEC-3B: [ka]
[0193] Rat APOBEC3: [ka]
[0194] Cow APOBEC-3B: [ka]
[0195] Chimpanzee APOBEC-3B: [ka]
[0196] Human APOBEC-3C: [ka]
[0197] Gorilla APOBEC3C [ka]
[0198] Human APOBEC-3A: [ka]
[0199] Rhesus macaque APOBEC-3A: [ka]
[0200] Cow APOBEC-3A: [ka]
[0201] Human APOBEC-3H: [ka]
[0202] Rhesus macaque APOBEC-3H: [ka]
[0203] Human APOBEC-3D: [ka]
[0204] Human APOBEC-1 [ka]
[0205] Mouse APOBEC-1: [ka]
[0206] Rat APOBEC-1: [ka]
[0207] Human APOBEC-2 [ka]
[0208] Mouse APOBEC-2: [ka]
[0209] Rat APOBEC-2: [ka] [ka]
[0210] Bovine APOBEC-2: [ka]
[0211] Petromyzon marinusCDA1(pmCDA1) [ka]
[0212] Human APOBEC3G D316R_D317R [ka]
[0213] Human APOBEC3G chain A [ka]
[0214] Human APOBEC3G chain A D120R_D121R [ka]
[0215] Deaminase domain that regulates the editing window of base editing factors Several aspects of this disclosure are based on the understanding that modifying the catalytic activity of the deaminase domain of any fusion protein provided herein, for example by creating point mutations in the deaminase domain, affects the processability of the fusion protein (e.g., a base-editing factor). For example, a mutation that reduces, but does not eliminate, the catalytic activity of the deaminase domain in a base-editing fusion protein may make it less likely that the deaminase domain will catalyze the deamination of a residue adjacent to the target residue, thereby narrowing the deamination window. The ability to narrow the deamination window may prevent unwanted deamination of residues adjacent to a particular target residue, thereby reducing or preventing off-target effects.
[0216] In some embodiments, any fusion protein provided herein comprises a deaminase domain (e.g., a cytidine deaminase domain) that reduces catalytic deaminase activity. In some embodiments, any fusion protein provided herein comprises a deaminase domain (e.g., a cytidine deaminase domain) that has reduced catalytic deaminase activity compared to a suitable control. For example, a suitable control may be the deaminase activity of a deaminase before introducing one or more mutations into the deaminase. In other embodiments, a suitable control may be a wild-type deaminase. In some embodiments, a suitable control is a wild-type apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, suitable controls are APOBEC1 deaminase, APOBEC2 deaminase, APOBEC3A deaminase, APOBEC3B deaminase, APOBEC3C deaminase, APOBEC3D deaminase, APOBEC3F deaminase, APOBEC3G deaminase, or APOBEC3H deaminase. In some embodiments, suitable controls are activation-inducible deaminases (AIDs). In some embodiments, suitable controls are cytidine deaminase 1 from Petromyzon marinus (pmCDA1). In some embodiments, the deaminase domain may have at least 1%, at least 5%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% less catalytic deaminase activity compared to a suitable control.
[0217] In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase comprising one or more mutations selected from the group consisting of H121X, H122X, R126X, R126X, R118X, W90X, W90X, and R132X of rAPOBEC1 (SEQ ID NO: 93), or one or more corresponding mutations in another APOBEC deaminase, where X is any amino acid. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase comprising one or more mutations selected from the group consisting of H121R, H122R, R126A, R126E, R118A, W90A, W90Y, and R132E of rAPOBEC1 (SEQ ID NO: 93), or one or more corresponding mutations in another APOBEC deaminase.
[0218] In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase comprising one or more mutations selected from the group consisting of D316X, D317X, R320X, R320X, R313X, W285X, W285X, R326X of hAPOBEC3G (SEQ ID NO: 77), or one or more corresponding mutations in another APOBEC deaminase, where X is any amino acid. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase comprising one or more mutations selected from the group consisting of D316R, D317R, R320A, R320E, R313A, W285A, W285Y, R326E of hAPOBEC3G (SEQ ID NO: 77), or one or more corresponding mutations in another APOBEC deaminase.
[0219] In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the H121R and H122R mutations of rAPOBEC1 (SEQ ID NO: 93), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the R126A mutation of rAPOBEC1 (SEQ ID NO: 93), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the R126E mutation of rAPOBEC1 (SEQ ID NO: 93), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the R118A mutation of rAPOBEC1 (SEQ ID NO: 93), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the W90A mutation of rAPOBEC1 (SEQ ID NO: 93) or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the W90Y mutation of rAPOBEC1 (SEQ ID NO: 93) or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the R132E mutation of rAPOBEC1 (SEQ ID NO: 93) or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the W90Y and R126E mutations of rAPOBEC1 (SEQ ID NO: 93) or one or more corresponding mutations in another APOBEC deaminase.In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the R126E and R132E mutations of rAPOBEC1 (SEQ ID NO: 93), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the W90Y and R132E mutations of rAPOBEC1 (SEQ ID NO: 93), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the W90Y, R126E and R132E mutations of rAPOBEC1 (SEQ ID NO: 93), or one or more corresponding mutations in another APOBEC deaminase.
[0220] In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the D316R and D317R mutations of hAPOBEC3G (SEQ ID NO: 77), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the R320A mutation of hAPOBEC3G (SEQ ID NO: 77), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the R320E mutation of hAPOBEC3G (SEQ ID NO: 77), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the R313A mutation of hAPOBEC3G (SEQ ID NO: 77), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the W285A mutation of hAPOBEC3G (SEQ ID NO: 77) or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the W285Y mutation of hAPOBEC3G (SEQ ID NO: 77) or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the R326E mutation of hAPOBEC3G (SEQ ID NO: 77) or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the W285Y and R320E mutations of hAPOBEC3G (SEQ ID NO: 77) or one or more corresponding mutations in another APOBEC deaminase.In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the R320E and R326E mutations of hAPOBEC3G (SEQ ID NO: 77), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the W285Y and R326E mutations of hAPOBEC3G (SEQ ID NO: 77), or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the W285Y, R320E, and R326E mutations of hAPOBEC3G (SEQ ID NO: 77), or one or more corresponding mutations in another APOBEC deaminase.
[0221] A fusion protein containing a nuclease-programmable DNA-binding protein (napDNAbp), cytidine deaminase, and uracil-binding protein (UBP). Several aspects of this disclosure provide fusion proteins comprising nucleic acid programmable DNA-binding proteins (napDNAbp), cytidine deaminase, and uracil-binding proteins (UBP). In some aspects, any fusion protein provided herein is a base-editing factor. In some aspects, UBP is a uracil-modifying enzyme. In some aspects, UBP is a uracil base-removing enzyme. In some aspects, UBP is a uracil DNA glycosylase. In some aspects, UBP is any uracil-binding protein provided herein. For example, UBP may be UDG, UdgX, UdgX *It may be UdgX_On or SMUG1. In some embodiments, UBP contains an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to a uracil-binding protein, uracil base removal enzyme, or uracil DNA glucosylase (UDG) enzyme. In some embodiments, UBP contains an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to any uracil-binding protein provided herein. For example, UBP may contain an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to any one of SEQ ID NOs. 48-53. In some embodiments, UBP contains one of the amino acid sequences of sequence numbers 48-53.
[0222] In some embodiments, the napDNAbp is a Cas9 domain, a Cpf1 domain, a CasX domain, a CasY domain, a C2cl domain, a C2c2 domain, a C2c3 domain, or an Argonaut domain. In some embodiments, the napDNAbp is any napDNAbp provided herein. In some embodiments, the napDNAbp of any fusion protein provided herein is a Cas9 domain. The Cas9 domain may be any Cas9 domain or Cas9 protein provided herein (e.g., dCas9 or nCas9). In some embodiments, any Cas9 domain or Cas9 protein provided herein (e.g., dCas9 or nCas9) may be fused with any cytidine deaminase domain provided herein. In some embodiments, the fusion protein has the following structure: NH2-[cytidine deaminase]-[napDNAbp]-[UBP]-COOH; NH2-[cytidine deaminase]-[UBP]-[napDNAbp]-COOH; NH2-[UBP]-[cytidine deaminase]-[napDNAbp]-COOH; NH2-[UBP]-[napDNAbp]-[cytidine deaminase]-COOH; NH2-[napDNAbp]-[UBP]-[cytidine deaminase]-COOH; or NH2-[napDNAbp]-[cytidine deaminase]-[UBP]-COOH Includes.
[0223] In some embodiments, a fusion protein comprising cytidine deaminase, napDNAbp (e.g., a Cas9 domain), and UBP does not contain a linker sequence. In some embodiments, the linker is located between the cytidine deaminase domain and the napDNAbp. In some embodiments, the linker is located between the cytidine deaminase domain and the UBP. In some embodiments, the linker is located between the napDNAbp and the UBP. In some embodiments, the "-" used in the general constructs above indicates the presence of any linker. In some embodiments, cytidine deaminase and napDNAbp, cytidine deaminase and UBP, and / or napDNAbp and UBP are fused via any linker provided herein. For example, in some embodiments, cytidine deaminase and napDNAbp, cytidine deaminase and UBP, and / or napDNAbp and UBP are fused via any linker provided below in the section titled "Linker". In some embodiments, cytidine deaminase and napDNAbp, cytidine deaminase and UBP, and / or napDNAbp and UBP are fused via a linker containing between 1 and 200 amino acids.In some aspects, cytidine deaminase and napDNAbp, cytidine deaminase and UBP, and / or napDNAbp and UBP are 1-5, 1-10, 1-20, 1-30, 1-40, 1-50, 1-60, 1-80, 1-100, 1-150, 1-200, 5-10, 5-20, 5-30, 5-40, 5-60, 5-80, 5-100, 5-150, 5-200, 10-20, 10-30, 10-40, 10-50, 10-60, 10-80, 10-100, 10-150, 10-200, 20-30, 20-40, 20-50 They are fused via linkers containing amino acid lengths of 20-60, 20-80, 20-100, 20-150, 20-200, 30-40, 30-50, 30-60, 30-80, 30-100, 30-150, 30-200, 40-50, 40-60, 40-80, 40-100, 40-150, 40-200, 50-60, 50-80, 50-100, 50-150, 50-200, 60-80, 60-100, 60-150, 60-200, 80-100, 80-150, 80-200, 100-150, 100-200, or 150-200. In some embodiments, cytidine deaminase and napDNAbp, cytidine deaminase and UBP, and / or napDNAbp and UBP are fused via a linker having an amino acid length of 4, 16, 24, 32, 91, or 104. In some embodiments, cytidine deaminase and napDNAbp, cytidine deaminase and UBP, and / or napDNAbp and UBP are fused via a linker having an amino acid length of 4, 16, 24, 32, 91, or 104. SGSETPGTSESATPES (Sequence ID 102), SGGS (Sequence ID 103), SGGSSGSETPGTSESATPESSGGS (Sequence ID 107), SGGSSGGSSGSETPGTSESATPESSGGSSGGS (Sequence No. 108), GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS (Sequence ID 109), or SGGSGGSGGS (Sequence No. 120) They are fused via a linker containing the amino acid sequence. In some embodiments, cytidine deaminase and napDNAbp, cytidine deaminase and UBP, and / or napDNAbp and UBP are fused via a linker containing the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 102), which is also referred to as the XTEN linker.
[0224] A fusion protein containing a nuclease-programmable DNA-binding protein (napDNAbp), cytidine deaminase, and a nucleic acid polymerase (NAP) domain. Some aspects of this disclosure provide fusion proteins comprising a nucleic acid programmable DNA-binding protein (napDNAbp), a cytidine deaminase, and a nucleic acid polymerase (NAP) domain. In some aspects, any fusion protein provided herein is a base-editing factor. In some aspects, NAP is a eukaryotic nucleic acid polymerase. In some aspects, NAP is a DNA polymerase. In some aspects, NAP has damage-overcoming polymerase activity. In some aspects, NAP is a damage-overcoming DNA polymerase. In some aspects, NAP is the Rev7, Rev1 complex, polymerase iota, polymerase kappa, or polymerase etha. In some aspects, NAP is the eukaryotic polymerase alpha, beta, gamma, delta, epsilon, gamma, etha, iota, kappa, lambda, mu, or nu. In some embodiments, NAP comprises an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to a nucleic acid polymerase (e.g., damage-overcoming DNA polymerase). In some embodiments, NAP comprises an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to any nucleic acid polymerase sequence provided herein. For example, NAP comprises an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to any one of SEQ ID NOs. 54-64. In some embodiments, NAP comprises an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to any one of SEQ ID NOs. 54-64.
[0225] In some embodiments, the napDNAbp is a Cas9 domain, a Cpf1 domain, a CasX domain, a CasY domain, a C2l domain, a C2c2 domain, a C2c3 domain, or an Argonaut domain. In some embodiments, the napDNAbp is any napDNAbp provided herein. In some embodiments, the napDNAbp of any fusion protein provided herein is a Cas9 domain. The Cas9 domain may be any Cas9 domain or Cas9 protein provided herein (e.g., dCas9 or nCas9). In some embodiments, any Cas9 domain or Cas9 protein provided herein (e.g., dCas9 or nCas9) may be fused to any of the cytidine deaminases provided herein. In some embodiments, the fusion protein has the structure: NH2-[cytidine deaminase]-[napDNAbp]-[NAP]-COOH; NH2-[cytidine deaminase]-[NAP]-[napDNAbp]-COOH; NH2-[NAP]-[cytidine deaminase]-[napDNAbp]-COOH; NH2-[NAP]-[napDNAbp]-[cytidine deaminase]-COOH; NH2-[napDNAbp]-[NAP]-[cytidine deaminase]-COOH; or NH2-[napDNAbp]-[cytidine deaminase]-[NAP]-COOH Includes.
[0226] In some embodiments, a fusion protein comprising cytidine deaminase, napDNAbp (e.g., Cas9 domain), and NAP does not contain a linker sequence. In some embodiments, the linker is located between the cytidine deaminase domain and the napDNAbp. In some embodiments, the linker is located between the cytidine deaminase domain and NAP. In some embodiments, the linker is located between the napDNAbp and NAP. In some embodiments, the "-" used in the general constructs above indicates the presence of any linker. In some embodiments, cytidine deaminase and napDNAbp, cytidine deaminase and NAP, and / or napDNAbp and NAP are fused via any linker provided herein. For example, in some embodiments, cytidine deaminase and napDNAbp, cytidine deaminase and NAP, and / or napDNAbp and NAP are fused via any linker provided below in the section titled "Linker". In some embodiments, cytidine deaminase and napDNAbp, cytidine deaminase and NAP, and / or napDNAbp and NAP are fused via a linker containing amino acids between 1 and 200.In some aspects, cytidine deaminase and napDNAbp, cytidine deaminase and NAP, and / or napDNAbp and NAP are 1-5, 1-10, 1-20, 1-30, 1-40, 1-50, 1-60, 1-80, 1-100, 1-150, 1-200, 5-10, 5-20, 5-30, 5-40, 5-60, 5-80, 5-100, 5-150, 5-200, 10-20, 10-30, 10-40, 10-50, 10-60, 10-80, 10-100, 10-150, 10-200, 20-30, 20-40, 20-50 They are fused via linkers containing amino acid lengths of 20-60, 20-80, 20-100, 20-150, 20-200, 30-40, 30-50, 30-60, 30-80, 30-100, 30-150, 30-200, 40-50, 40-60, 40-80, 40-100, 40-150, 40-200, 50-60, 50-80, 50-100, 50-150, 50-200, 60-80, 60-100, 60-150, 60-200, 80-100, 80-150, 80-200, 100-150, 100-200, or 150-200. In some embodiments, cytidine deaminase and napDNAbp, cytidine deaminase and NAP, and / or napDNAbp and NAP are fused via a linker having a length of 4, 16, 32, or 104 amino acids. In some embodiments, cytidine deaminase and napDNAbp, cytidine deaminase and NAP, and / or napDNAbp and NAP are fused via a linker having a length of 4, 16, 32, or 104 amino acids. GSETPGTSESATPES (SEQ ID NO: 102), SGGS (Sequence ID 103), SGGSSGSETPGTSESATPESSGGS (Sequence ID 107), SGGSSGGSSGSETPGTSESATPESSGGSSGGS (Sequence No. 108), GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS (Sequence ID 109), or They are fused via a linker containing the amino acid sequence SGGSGGSGGS (SEQ ID NO: 120). In some embodiments, cytidine deaminase and napDNAbp, cytidine deaminase and NAP, and / or napDNAbp and NAP are fused via a linker containing the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 102), which may also be referred to as the XTEN linker.
[0227] A fusion protein containing a nuclease-programmable DNA-binding protein (napDNAbp), cytidine deaminase, uracil-binding protein (UBP), and a nucleic acid polymerase (NAP) domain. Some aspects of this disclosure provide fusion proteins comprising a nucleic acid programmable DNA-binding protein (napDNAbp), cytidine deaminase, uracil-binding protein (UBP), and a nucleic acid polymerase (NAP) domain. In some aspects, any fusion protein provided herein is a base-editing factor. In some aspects, NAP is a eukaryotic nucleic acid polymerase. In some aspects, NAP is a DNA polymerase. In some aspects, NAP has damage-overcoming polymerase activity. In some aspects, NAP is a damage-overcoming DNA polymerase. In some aspects, NAP is the Rev7, Rev1 complex, polymerase iota, polymerase kappa, or polymerase etha. In some aspects, NAP is the eukaryotic polymerase alpha, beta, gamma, delta, epsilon, gamma, etha, iota, kappa, lambda, mu, or nu. In some embodiments, NAP comprises an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to a nucleic acid polymerase (e.g., damage-overcoming DNA polymerase). In some embodiments, NAP comprises an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to any nucleic acid polymerase provided herein. For example, NAP may comprise an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to any one of SEQ ID NOs. 54-64. In some embodiments, NAP comprises an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to any one of SEQ ID NOs. 54-64.
[0228] In some embodiments, UBP is a uracil-modifying enzyme. In some embodiments, UBP is a uracil-chlorine-removing enzyme. In some embodiments, UBP is a uracil DNA glycosylase. In some embodiments, UBP is any uracil-binding protein provided herein. For example, UBP is UDG, UdgX, UdgX *It may be UdgX_On or SMUG1. In some embodiments, UBP contains an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to a uracil-binding protein, uracil base removal enzyme, or uracil DNA glucosylase (UDG) enzyme. In some embodiments, UBP contains an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to any uracil-binding protein provided herein. For example, UBP may contain an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to any one of SEQ ID NOs. 48-53. In some embodiments, UBP contains one of the amino acid sequences of sequence numbers 48-53.
[0229] In some embodiments, the napDNAbp is a Cas9 domain, a Cpf1 domain, a CasX domain, a CasY domain, a C2l domain, a C2c2 domain, a C2c3 domain, or an Argonaut domain. In some embodiments, the napDNAbp is any napDNAbp provided herein. In some embodiments, the napDNAbp of any fusion protein provided herein is a Cas9 domain. The Cas9 domain may be any Cas9 domain or Cas9 protein provided herein (e.g., dCas9 or nCas9). In some embodiments, any Cas9 domain or Cas9 protein provided herein (e.g., dCas9 or nCas9) may be fused to any of the cytidine deaminases provided herein. In some embodiments, the fusion protein has the structure: NH2-[NAP]-[cytidine deaminase]-[napDNAbp]-[UBP]-COOH; NH2-[cytidine deaminase]-[NAP]-[napDNAbp]-[UBP]-COOH; NH2-[cytidine deaminase]-[napDNAbp]-[NAP]-[UBP]-COOH; NH2-[cytidine deaminase]-[napDNAbp]-[UBP]-[NAP]-COOH; NH2-[NAP]-[cytidine deaminase]-[UBP]-[napDNAbp]-COOH; NH2-[cytidine deaminase]-[NAP]-[UBP]-[napDNAbp]-COOH; NH2-[cytidine deaminase]-[UBP]-[NAP]-[napDNAbp]-COOH; NH2-[cytidine deaminase]-[UBP]-[napDNAbp]-[NAP]-COOH; NH2-[NAP]-[UBP]-[cytidine deaminase]-[napDNAbp]-COOH; NH2-[UBP]-[NAP]-[cytidine deaminase]-[napDNAbp]-COOH; NH2-[UBP]-[cytidine deaminase]-[NAP]-[napDNAbp]-COOH; NH2-[UBP]-[cytidine deaminase]-[napDNAbp]-[NAP]-COOH; NH2-[NAP]-[UBP]-[napDNAbp]-[cytidine deaminase]-COOH; NH2-[UBP]-[NAP]-[napDNAbp]-[cytidine deaminase]-COOH; NH2-[UBP]-[napDNAbp]-[NAP]-[cytidine deaminase]-COOH; NH2-[UBP]-[napDNAbp]-[cytidine deaminase]-[NAP]-COOH; NH2-[NAP]-[napDNAbp]-[UBP]-[cytidine deaminase]-COOH; NH2-[napDNAbp]-[NAP]-[UBP]-[cytidine deaminase]-COOH; NH2-[napDNAbp]-[UBP]-[NAP]-[cytidine deaminase]-COOH; NH2-[napDNAbp]-[UBP]-[cytidine deaminase]-[NAP]-COOH; NH2-[NAP]-[napDNAbp]-[cytidine deaminase]-[UBP]-COOH; NH2-[napDNAbp]-[NAP]-[cytidine deaminase]-[UBP]-COOH; NH2-[napDNAbp]-[cytidine deaminase]-[NAP]-[UBP]-COOH; or NH2-[napDNAbp]-[cytidine deaminase]-[UBP]-[NAP]-COOH Includes.
[0230] In some embodiments, a fusion protein comprising cytidine deaminase, napDNAbp (e.g., Cas9 domain), UBP, and NAP does not contain a linker sequence. In some embodiments, the linker is located between the cytidine deaminase domain and napDNAbp, NAP, and / or UBP. In some embodiments, the linker is located between napDNAbp and the cytidine deaminase domain, NAP, and / or UBP. In some embodiments, the linker is located between NAP and cytidine deaminase, napDNAbp, and / or UBP. In some embodiments, the linker is located between UBP and cytidine deaminase, napDNAbp, and NAP. In some embodiments, the "-" used in the above general constructs indicates the presence of any linker. In some embodiments, the linker is any linker provided herein, for example, in a section titled "Linker". In some embodiments, the linker comprises amino acids between 1 and 200. In this configuration, the linkers are 1-5, 1-10, 1-20, 1-30, 1-40, 1-50, 1-60, 1-80, 1-100, 1-150, 1-200, 5-10, 5-20, 5-30, 5-40, 5-60, 5-80, 5-100, 5-150, 5-200, 10-20, 10-30, 10-40, 10-50, 10-60, 10-80, 10-100, 10-150, 10-200, 20-30, 20-40, 20-50, 20-60, 20-80, 20-100 , containing amino acid lengths of 20-150, 20-200, 30-40, 30-50, 30-60, 30-80, 30-100, 30-150, 30-200, 40-50, 40-60, 40-80, 40-100, 40-150, 40-200, 50-60, 50-80, 50-100, 50-150, 50-200, 60-80, 60-100, 60-150, 60-200, 80-100, 80-150, 80-200, 100-150, 100-200, or 150-200. Linkers containing 4, 16, 32, or 104 amino acid lengths in some embodiments. In some embodiments, SGSETPGTSESATPES (SEQ ID NO: 102), SGGS (SEQ ID NO: 103), SGGSSGSETPGTSESATPESSGGS (Sequence ID 107), SGGSSGGSSGSETPGTSESATPESSGGSSGGS (Sequence No. 108), GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS (Sequence ID 109), Or SGGSGGSGGS (Sequence ID 120) A linker containing the amino acid sequence. In some embodiments, the linker contains the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 102), which is also sometimes referred to as the XTEN linker.
[0231] A fusion protein containing a nuclease-programmable DNA-binding protein and a base-extraction enzyme (BEE). Some aspects of this disclosure provide fusion proteins comprising nucleic acid programmable DNA-binding proteins (napDNAbp) and base-removal enzymes. In some embodiments, any fusion protein provided herein is a base-editing factor. In some embodiments, the base-removal enzyme (BEE) is a cytosine, thymine, adenine, guanine, or uracil base-removal enzyme. In some embodiments, the base-removal enzyme (BEE) is a cytosine base-removal enzyme. In some embodiments, the BEE is a thymine base-removal enzyme. In some embodiments, the base-removal enzyme comprises an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to a naturally occurring BEE. In some embodiments, the base-removal enzyme comprises an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to either SEQ ID NO: 65 or 66. In some embodiments, the base removal enzyme contains either one amino acid sequence of SEQ ID NO: 65 or 66.
[0232] In some embodiments, the napDNAbp is a Cas9 domain, a Cpf1 domain, a CasX domain, a CasY domain, a C21 domain, a C2c2 domain, a C2c3 domain, or an Argonaut domain. In some embodiments, the napDNAbp is any napDNAbp provided herein. In some embodiments, the napDNAbp of any fusion protein provided herein is a Cas9 domain. The Cas9 domain may be any Cas9 domain or Cas9 protein provided herein (e.g., dCas9 or nCas9). In some embodiments, any Cas9 domain or Cas9 protein provided herein (e.g., dCas9 or nCas9) may be fused to any cytidine deaminase provided herein. In some embodiments, the fusion protein has the structure: NH2-[BEE]-[napDNAbp]-COOH; or NH2-[napDNAbp]-[BEE]-COOH; Includes.
[0233] In some embodiments, the fusion protein further comprises a nucleic acid polymerase (NAP). In some embodiments, the NAP is a eukaryotic nucleic acid polymerase. In some embodiments, NAP is a DNA polymerase. In some embodiments, NAP has damage-overcoming polymerase activity. In some embodiments, NAP is a damage-overcoming DNA polymerase. In some embodiments, NAP is the Rev7, Rev1 complex, polymerase iota, polymerase kappa, or polymerase etha. In some embodiments, NAP is the eukaryotic polymerase alpha, beta, gamma, delta, epsilon, gamma, etha, iota, kappa, lambda, mu, or nu. In some embodiments, NAP contains an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to a nucleic acid polymerase (e.g., a damage-overcoming DNA polymerase). In some embodiments, NAP contains an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to any nucleic acid polymerase provided herein. For example, NAP may contain an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to any one of SEQ ID NOs. 54-64. In some embodiments, NAP contains an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% identical to any one of SEQ ID NOs. 54-64. In some embodiments, the fusion protein has the structure: NH2-[BEE]-[napDNAbp]-[NAP]-COOH; NH2-[BEE]-[NAP]-[napDNAbp]-COOH; NH2-[NAP]-[BEE]-[napDNAbp]-COOH; NH2-[NAP]-[napDNAbp]-[BEE]-COOH; NH2-[napDNAbp]-[NAP]-[BEE]-COOH; or NH2-[napDNAbp]-[BEE]-[NAP]-COOH Includes.
[0234] In some embodiments, a fusion protein comprising napDNAbp (e.g., a Cas9 domain) and BEE does not contain a linker sequence. In some embodiments, a fusion protein comprising napDNAbp (e.g., a Cas9 domain), BEE, and NAP does not contain a linker sequence. In some embodiments, the linker is located between napDNAbp and BEE. In some embodiments, the linker is located between BEE and NAP and / or napDNAbp. In some embodiments, the linker is located between NAP and BEE and / or napDNAbp. In some embodiments, the linker is located between napDNAbp and BEE and / or NAP. In some embodiments, the "-" used in the above general constructs indicates the presence of any linker. In some embodiments, the linker is any linker provided herein, for example, in a section titled "Linker". In some embodiments, the linker comprises amino acids between 1 and 200. In this configuration, the linkers are 1-5, 1-10, 1-20, 1-30, 1-40, 1-50, 1-60, 1-80, 1-100, 1-150, 1-200, 5-10, 5-20, 5-30, 5-40, 5-60, 5-80, 5-100, 5-150, 5-200, 10-20, 10-30, 10-40, 10-50, 10-60, 10-80, 10-100, 10-150, 10-200, 20-30, 20-40, 20-50, 20-60, 20-80, 20-100 , containing amino acid lengths of 20-150, 20-200, 30-40, 30-50, 30-60, 30-80, 30-100, 30-150, 30-200, 40-50, 40-60, 40-80, 40-100, 40-150, 40-200, 50-60, 50-80, 50-100, 50-150, 50-200, 60-80, 60-100, 60-150, 60-200, 80-100, 80-150, 80-200, 100-150, 100-200, or 150-200. Linkers containing 4, 16, 32, or 104 amino acid lengths in some embodiments. SGSETPGTSESATPES (SEQ ID NO: 102), SGGS (SEQ ID NO: 103), SGGSSGSETPGTSESATPESSGGS (Sequence ID 107), SGGSSGGSSGSETPGTSESATPESSGGSSGGS (Sequence No. 108), GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS (Sequence ID 109), or SGGSGGSGGS (Sequence No. 120) A linker containing the amino acid sequence. In some embodiments, the linker contains the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 102), which may also be referred to as the XTEN linker.
[0235] Fusion protein containing nuclear localization sequence (NLS) In some embodiments, the fusion proteins provided herein further comprise one or more nuclear target sequences, e.g., nuclear localization sequences (NLS). In some embodiments, the NLS comprises an amino acid sequence that facilitates the transfer of the protein into the cell nucleus (e.g., by nuclear transport), including the NLS. In some embodiments, any fusion protein provided herein further comprises a nuclear localization sequence (NLS). In some embodiments, the NLS is fused to the N-terminus of the fusion protein. In some embodiments, the NLS is fused to the C-terminus of the fusion protein. In some embodiments, the NLS is fused to the N-terminus of napDNAbp. In some embodiments, the NLS is fused to the C-terminus of napDNAbp. In some embodiments, the NLS is fused to the N-terminus of NAP. In some embodiments, the NLS is fused to the C-terminus of NAP. In some embodiments, the NLS is fused to the N-terminus of cytidine deaminase. In some embodiments, the NLS is fused to the C-terminus of cytidine deaminase. In some embodiments, the NLS is fused to the N-terminus of UBP. In some embodiments, the NLS is fused to the C-terminus of UBP. In some embodiments, the NLS is fused to the N-terminus of the BEE. In some embodiments, the NLS is fused to the C-terminus of the BEE. In some embodiments, the NLS is fused to the fusion protein via one or more linkers. In some embodiments, the NLS is fused to the fusion protein without the use of linkers. In some embodiments, the NLS comprises one amino acid sequence of any of the NLS sequences provided or referenced herein. In some embodiments, the NLS comprises the amino acid sequence shown in SEQ ID NO: 41 or SEQ ID NO: 42. Additional nuclear localization sequences are known in the art and will be apparent to those skilled in the art. For example, the NLS sequence is described in Plank et al., PCT / EP2000 / 011690, which is incorporated herein by reference to the disclosure of exemplary nuclear localization sequences.In some embodiments, NLS comprises the amino acid sequence PKKKRKV (SEQ ID NO: 41), MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 42), KRTADGSEFESPKKKRKV (SEQ ID NO: 43), KRGINDRNFWRGENGRKTR (SEQ ID NO: 44), KKTGGPIYRRVDGKWRR (SEQ ID NO: 45), RRELILYDKEEIRRIWR (SEQ ID NO: 46), or AVSRKRKA (SEQ ID NO: 47).
[0236] Linker In certain embodiments, a linker may be used to link any of the proteins or protein domains described herein. The linker may be as simple as a covalent bond, or it may be a polymer linker whose length is many atoms. In certain embodiments, the linker is a polypeptide or based on amino acids. In other embodiments, the linker is not peptide-like. In certain embodiments, the linker is a covalent bond (e.g., carbon-carbon bond, disulfide bond, carbon-heteroatom bond, etc.). In certain embodiments, the linker is a carbon-nitrogen bond of an amide bond. In certain embodiments, the linker is a cyclic or acyclic, substituted or unsubstituted, branched or unbranched, aliphatic or heteroaliphatic linker. In certain embodiments, the linker is polymeric (e.g., polyethylene, polyethylene glycol, polyamide, polyester, etc.). In certain embodiments, the linker comprises a monomer, dimer, or polymer of an aminoalkanoic acid. In certain embodiments, the linker comprises aminoalkanoic acid (e.g., glycine, ethaneic acid, alanine, beta-alanine, 3-aminopropanoic acid, 4-aminobutyric acid, 5-pentanoic acid, etc.). In certain embodiments, the linker comprises monomers, dimers, or polymers of aminohexanoic acid (Ahx). In certain embodiments, the linker is based on a carbocyclic moiety (e.g., cyclopentane, cyclohexane). In other embodiments, the linker comprises a polyethylene glycol moiety (PEG). In other embodiments, the linker comprises an amino acid. In certain embodiments, the linker comprises a peptide. In certain embodiments, the linker comprises an aryl or heteroaryl moiety. In certain embodiments, the linker is based on a phenyl ring. The linker may include a functional moiety (e.g., thiol, amino) to facilitate the attachment of a nucleophile from the peptide to the linker. Any of the electrophiles may be used as part of the linker. Exemplary electrophiles include, but are not limited to, active esters, active amides, Michael acceptors, alkyl halides, aryl halides, acyl halides, and isothiocyanates.
[0237] In some embodiments, the linker is a single amino acid or a group of amino acids (e.g., a peptide or protein). In some embodiments, the linker is a bond (e.g., a covalent bond), an organic molecule, a group, a polymer, or a chemical site. In some embodiments, the linker is 5 to 100 amino acids long. Examples include amino acid lengths of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-110, 110-120, 120-130, 130-140, 140-150, or 150-200 amino acid lengths. Longer or shorter linkers are also intended. In some embodiments, the linker includes the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 102), which is also sometimes referred to as the XTEN linker. In some embodiments, the linker comprises the amino acid sequence SGGS (SEQ ID NO: 103). In some embodiments, the linker comprises (SGGS) n (SEQ ID NO: 103), (GGGS)n (SEQ ID NO: 104), (GGGGS)n (SEQ ID NO: 105), (G)n (SEQ ID NO: 121), (EAAAK)n (SEQ ID NO: 106), (GGS)n (SEQ ID NO: 122), SGSETPGTSESATPES (SEQ ID NO: 102), SGGSGGSGGS (SEQ ID NO: 120), or (XP) nThe linker comprises the motif (SEQ ID NO: 123), or any combination thereof, where n is an integer independently between 1 and 30, and where X is any amino acid. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. In some embodiments, the linker comprises SGSETPGTSESATPES (SEQ ID NO: 102) and SGGS (SEQ ID NO: 103). In some embodiments, the linker comprises SGGSSGSETPGTSESATPESSGGS (SEQ ID NO: 107). In some embodiments, the linker comprises SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 108). In some embodiments, the linker comprises GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS (SEQ ID NO: 109). In some embodiments, the linker includes SGGSGGSGGS (SEQ ID NO: 120).
[0238] Nucleic acid programmable DNA-binding proteins (napDNAbp) bind to guide nucleic acids. Some aspects of this disclosure provide a complex comprising one of the fusion proteins provided herein and a guide nucleic acid bound to the napDNAbp of the fusion protein. Some aspects of this disclosure provide a complex comprising any fusion protein provided herein and a guide RNA bound to the Cas9 domain of the fusion protein (e.g., dCas9, nuclease-active Cas9, or Cas9 niccas).
[0239] In some embodiments, the guide nucleic acid (e.g., guide RNA) is 15 to 100 nucleotides long and contains a sequence of at least 10 consecutive nucleotides that are complementary to the target sequence. In some embodiments, the guide RNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides long. In some embodiments, the guide RNA comprises a sequence of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides that are complementary to the target sequence. In some embodiments, the target sequence is a DNA sequence. In some embodiments, the target sequence is an RNA sequence. In some embodiments, the target sequence is a sequence in the mammalian genome. In some embodiments, the target sequence is a sequence in the human genome. In some embodiments, the 3' end of the target sequence is immediately adjacent to a classical PAM sequence (NGG). In some embodiments, the guide nucleic acid (e.g., guide RNA) is complementary to a guide nucleic acid associated with a disease or abnormality. In some embodiments, the guide nucleic acid (e.g., guide RNA) is complementary to a sequence associated with a disease or disorder having a mutation related to any disease or abnormality provided herein. In some embodiments, the guide nucleic acid (e.g., guide RNA) is complementary to any gene that causes a disease or disorder provided herein.
[0240] Methods using fusion proteins Some aspects of this disclosure provide methods for using a complex comprising any fusion protein (e.g., a base editing factor) provided herein, or a guide nucleic acid (e.g., gRNA) and a fusion protein (e.g., a base editing factor) provided herein. For example, some aspects of this disclosure provide methods for contacting a DNA or RNA molecule with any of the fusion proteins or base editing factors provided herein and with at least one guide nucleic acid (e.g., guideRNA), wherein the guide nucleic acid (e.g., guideRNA) is about 15 to 100 nucleotides long and comprises at least 10 consecutive nucleotides complementary to the target sequence. In some embodiments, the 3' end of the target sequence is immediately adjacent to a classical spCas9PAM sequence (NGG). In some embodiments, the 3' end of the target sequence is not immediately adjacent to a classical spCas9PAM sequence (NGG). In some embodiments, the 3' end of the target sequence is immediately adjacent to an AGC, GAG, TTT, GTG, or CAA sequence.
[0241] In some embodiments, the target DNA sequence includes a sequence related to a disease or disorder. In some embodiments, the target DNA sequence includes a point mutation associated with a disease or disorder. In some embodiments, the activity of a fusion protein (e.g., napDNAbp, cytidine deaminase, and uracil-binding protein UBP), or a complex, results in the modification of a point mutation. In some embodiments, the target DNA sequence includes a G→C or C→G point mutation associated with a disease or disorder, where deamination and / or excision of the mutated C base results in a sequence not associated with a disease or disorder. In some embodiments, the target DNA sequence codes for a protein, and the point mutation is present in a codon, resulting in a change in the amino acid encoded by the mutated codon compared to the wild-type codon. In some embodiments, deamination of the mutated C results in a change in the amino acid encoded by the mutated codon. In some embodiments, deamination of the mutated C results in a codon encoding a wild-type amino acid. In some embodiments, contact is in vivo in the subject. In some embodiments, the subject has or has been diagnosed with a disease or disorder. In some aspects, the disease or disorder is 22q13.3 deletion syndrome; 2-methyl-3-hydroxybutyrateuria; 3-methylcrotonyl-CoA carboxylase 1 deficiency; 3-methylcrotonyl-CoA carboxylase 2 deficiency; 3-methylglutaconic aciduria type 2; 3-methylglutaconic aciduria type 3; 3-methylglutaconic aciduria type V; 3-oxo-5 alpha-steroid delta-4-dehydrogenase deficiency; 46, XY sex reassignment, type 1; 46, XY true hermaphrodite, SRY-related; 4-hydroxyphenylpyruvate dioxygenase deficiency; abnormal facial morphology; abnormal glycosylation (CDG IIa); Chondroplasia type 2; Color blindness 2; Color blindness 5; Color blindness 6; Color blindness 7; Acquired hemoglobin H disease; Acrosacral dactyly type 1; Acroporestrophic osteogenesis imperfecta with or without hormone resistance 1; Acroporestrophic osteogenesis imperfecta with or without hormone resistance 2; Acropacial dysplasia, Cincinnati type; ACTH resistance; Acute neuropathic Gaucher disease; Adams-Oliver syndrome; Adams-Oliver syndrome 2; Adams-Oliver syndrome 4;Adams-Oliver syndrome 6; adenine phosphoribosyltransferase deficiency; adenylosuccinate lyase deficiency; adolescent nephronoplasia; adrenoleukodystrophy; adult junctional epidermolysis bullosa; adult neuronal ceroid lipofuscinosis; ADULT syndrome; age-related macular degeneration 14; age-related macular degeneration 3; Eicardi-Goutier syndrome 5; Eicardi-Goutier syndrome 6; Alexander disease; alpha-thalassemia; alpha-B crystallin disorder; Alport syndrome, autosomal recessive; Alport syndrome, X-linked recessive; alternating hemiplegia in childhood 2; Alzheimer's disease; Alzheimer's disease, type 1; Alzheimer's disease, type 3; ameloidy, hypomaturity, IIA3; ameloidy, type 1E; Amish fatal microcephaly; AML - acute myeloid leukemia; amyloidogenic transthyretin amyloid -sis; amyotrophic lateral sclerosis (ALS) 16, juvenile; amyotrophic lateral sclerosis 6, autosomal recessive; amyotrophic lateral sclerosis type 1; amyotrophic lateral sclerosis type 10; amyotrophic lateral sclerosis type 2; amyotrophic lateral sclerosis type 9; Andersentawil syndrome; anemia, congenital erythrocytosis disorder, type IV; anemia due to G6PD deficiency, nonspherocytic hemolysis; anemia, sideroblastic, pyridoxine-refractory, autosomal recessive; Angelman syndrome; Hereditary vascular disorders with nephropathy, aneurysms, and muscle spasms; anhidrotic ectodermal dysplasia with immunodeficiency; onychomycosis; Antley-Bixler syndrome with genital abnormalities and impaired steroid synthesis; Antley-Bixler syndrome without genital abnormalities or impaired steroid synthesis; aplastic anemia; apolipoprotein AI deficiency; arginase deficiency; arrhythmia-induced right ventricular cardiomyopathy; arrhythmia-induced right ventricular cardiomyopathy, type 11; Arrhythmia-induced right ventricular cardiomyopathy, type 9; arterial calcification in infancy; tortuosomatic arterial syndrome; congenital distal type 1 polyarthritis; articular arrhythmia, renal dysfunction, cholestatic syndrome; articular arrhythmia, distal, type 5d; Arts syndrome; aspartylglucosamineuria, Finnish type; asphyxiated thoracic dystrophy; ataxia with vitamin E deficiency; telangiectatic ataxia syndrome; telangiectatic ataxia-like disorder; atelosteogenesis type 1; atrial fibrillation; familial atrial fibrillation, 10; atrial septal defect, 4; cone atrophy; ATR-X syndrome; atypical hemolytic uremic syndrome, 1; auditory neuropathy, autosomal recessive, 1;Auricular condyle syndrome 1; Autoimmune disease, multiple systems, infantile onset; Autoimmune lymphoproliferative syndrome, type 1A; Autoimmune lymphoproliferative syndrome, type V; Autosomal dominant nocturnal frontal lobe epilepsy; Autosomal dominant progressive extraocular myopia with mitochondrial DNA deletion 2; Autosomal dominant progressive extraocular myopia with mitochondrial DNA deletion 3; Autosomal dominant progressive extraocular myopia with mitochondrial DNA deletion 4; Autosomal recessive congenital ichthyosis 1; Autosomal recessive congenital ichthyosis 5; Autosomal recessive hypophosphatemic vitamin D refractory rickets; Axenfeldt-Rieger anomaly; Axenfeldt-Rieger syndrome type 1; Axenfeldt-Rieger syndrome type 3; Balatter-Winter syndrome 1; Bade-Beedl syndrome; Bade-Beedl syndrome 10; Bade-Beedl syndrome 12; Bade-Beedl syndrome 2; Ba -de-Beedl syndrome 3; Barde-Beedl syndrome 4; Barde-Beedl syndrome 9; Barter syndrome prenatal type 2; Barter syndrome, type 4b; basal ganglia disorder, biotin responsive; Becker muscular dystrophy; benign familial neonatal seizure 1; benign familial neonatal infantile seizure; benign recurrent intrahepatic cholestasis 2; Bernard-Soulier syndrome, type B; beta-thalassemia; Vietti crystalline corneal-retinal dystrophy; bile acid synthesis disorder, congenital, 2; biotinidase deficiency; hemorrhagic disorder, platelet type, 19; blood type - lutelan inhibitor; Bloom syndrome; Bosley-Sally-Alorain syndrome; Boucher-Neuhauser syndrome; brachydactyly type B2; breast cancer; breast-ovarian cancer, familial 1; breast-ovarian cancer, familial 2; bronchiectasis; Brown-Vialetto-Van Laere syndrome; Brown-Vialetto-Van Laere syndrome; Bullous ichthyotic erythroderma; Burkitt lymphoma; Camptomélic dysplasia; Cap myopathy; Carbohydrate deficiency glycoprotein syndrome type I; Carbohydrate deficiency glycoprotein syndrome type II; Colon cancer; Pancreatic cancer; Cardiac arrhythmia; Fatal infantile cytochrome C oxidase deficiency, cardiomyopathy; Cardiofacial cutaneous syndrome; Cardiofacial cutaneous syndrome; Cardiomyopathy, restrictive; Carney complex, type 1; Carnitine palmitoyltransferase I deficiency; Cataract; Congenital cataract with sensorineural hearing loss, Down syndrome-like facial appearance, short stature and intellectual disability;Catecholamine-mediated polymorphic ventricular tachycardia; core diseases; central precocity; cerebellar ataxia and hypogonadism; pediatric cerebellar ataxia with progressive extraocular palsy; cerebellar ataxia, hearing loss, narcolepsy; cerebral amyloid vascular disease, APP-related; autosomal dominant arterial disease with subcortical infarction and leukoencephalopathy; cerebral spongiform malformation1; cerebral palsy, spastic quadriplegia1; cerebromandibular syndrome; neuronal ceroid lipofuscinosis1; neuronal ceroid lipofuscinosis10; neuronal ceroid lipofuscinosis neurons6; neuronal ceroid lipofuscinosis7; neuronal ceroid lipofuscinosis8; neuropathy Transceroid lipofuscinosis, 13; Neuronal ceroid lipofuscinosis, 2; Ch-Higashi syndrome; Charcot-Marie-Tooth disease; Charcot-Marie-Tooth disease type 1B; Charcot-Marie-Tooth disease type 2B; Charcot-Marie-Tooth disease type 2D; Charcot-Marie-Tooth disease type 2I; Charcot-Marie-Tooth disease type 2K; Charcot-Marie-Tooth disease, axonal and vocal cord paralysis, autosomal recessive; Charcot-Marie-Tooth disease, demyelination, type 1C; Charcot-Marie-Tooth disease, dominant intermediate type. Intermediate) E; Charcot-Marie-Tooth disease, type 2; Charcot-Marie-Tooth disease, type 2A2; Charcot-Marie-Tooth disease, type 4C; Charcot-Marie-Tooth disease, type 4G; Charcot-Marie-Tooth disease, type IA; Charcot-Marie-Tooth disease, type IE; Charcot-Marie-Tooth disease, type IF; Charcot-Marie-Tooth disease, X-linked recessive, type 5; CHARGE-related; childhood syndrome; cholestanol storage disorder; cholesterol monooxygenase (side chain break) deficiency; Chondrodysplasia 1, X-linked recessive; Chop syndrome; chromosome 9q deletion syndrome; chronic granulomatous disease, X-linked; ciliary dyskinesia, primary, 14; ciliary dyskinesia, primary, 19; ciliary dyskinesia, primary, 3; ciliary dyskinesia, primary, 7; cranial dysplasia; Cockayne syndrome type A; Coffin-Lowry syndrome; Cohen syndrome; Cohen's disease; colorectal cancer, hereditary, nonpolyposis, type 1; combined cellular humoral immunodeficiency with granuloma; combined oxidative phospholipidosis 24; combined oxidative phospholipidosis 9; unclassifiable immunodeficiency 7; complement component 9 deficiency;Cone-rod dystrophy 10; cone-rod dystrophy 11; cone-rod dystrophy 3; cone-rod dystrophy 5; cone-rod dystrophy 6; congenital adrenal dysplasia, X-linked; congenital megakaryocytic thrombocytopenia; congenital aniridia; congenital bilateral absence of the vas deferens; congenital cataract, hearing loss, and neurological degeneration; congenital contractile arachnoid disease; congenital folate malabsorption; g Congenital lycosylation disorder type 1K; congenital glycosylation disorder type 1M; congenital glycosylation disorder type 1t; congenital glycosylation disorder type 1u; congenital glycosylation disorder type 2C; congenital systemic lipodystrophy type 1; congenital systemic lipodystrophy type 2; congenital heart disease, multiple types, 1, X-linked; congenital lactase deficiency; congenital long QT syndrome; Brain and eyesCongenital muscular dystrophy / dystroglycanopathy with abnormalities, type A2; Congenital muscular dystrophy / dystroglycanopathy with abnormalities of the brain and eyes, type A7; Congenital muscular dystrophy / dystroglycanopathy with intellectual disability, type B1; Congenital muscular dystrophy / dystroglycanopathy with intellectual disability, type B2; Congenital myopathy with fibrous type imbalance; Congenital myotonia, autosomal dominant type; Congenital myotonia, autosomal recessive type; Congenital constant night blindness, autosomal dominant type 3; Congenital constant nocturnal blindness, type 1A; Congenital constant nocturnal blindness, type 1F; Coproporphilia; Corneal dystrophy, Fuchs' endothelium, type 8; Cornea Epithelial dystrophy; corneal fragility, kerotoglobus, blue sclera, and hypermobility of joints; Cornelia de Lange syndrome1; Cornelia de Lange syndrome4; complex cortical dysplasia with other brain malformations3; cortisone reductase deficiency1; Cowden syndrome2; ectoderm dysplasia1; craniofacial hearing loss syndrome; cranial joint disorder; craniosynostosis; craniosynostosis3; craniosynostosis and dental abnormalities; creatine deficiency, X-linked; Crigler-Nadjar syndrome, type 1; Crouzon syndrome; cryptophthalmos syndrome; undescended testicles, unilateral or bilateral; Cushing's synostosis; Beare and Stevenson's Cutis Gyrata syndrome; cystathioniriosis; cystic fibrosis; cystinosis, non-renal ocular impairment; cytochrome c oxidase deficiency; Danon disease; hearing loss, autosomal dominant 12; hearing loss, autosomal dominant 20; hearing loss, autosomal recessive 1A; hearing loss, autosomal recessive 63; hearing loss, autosomal recessive 8; hearing loss, autosomal recessive 9; acetyl-CoA acetyltransferase deficiency; alpha-mannosidase deficiency; feroxidase deficiency; glycerol kinase deficiency; guanidinoethanolate methyltransferase deficiency; hydroxymethylglutaryl-CoA lyase deficiency; iodide peroxidase deficiency; malonyl-CoA decarboxylase deficiency; UDP-glucose-hexose-1-phosphate uridyltransferase deficiency; speech and language development delay; delta-thalassemia; Dent disease 1; Desbucoy Syndrome; Desmosterol disease; DFNA2 asymptomatic hearing loss; Diabetes mellitus type 2; Diabetes mellitus, insulin-dependent, 20; Digitorenocerebral syndrome; Dilated cardiomyopathy 1FF; Dilated cardiomyopathy 1G; Dilated cardiomyopathy 1S; Dilated cardiomyopathy 1X; Dilated cardiomyopathy 3B; Steroid production disorder due to cytochrome p450 oxidoreductase deficiency; Distal inherited motor neuropathy type 2B; Stomatitis-lymphedema syndrome; Dash syndrome; Duchenne muscular dystrophy; Congenital autosomal dominant diskeratosis; X-linked congenital diskeratosis; Congenital diskeratosis, autosomal dominant 2; Congenital Diskerat syndrome, autosomal recessive; 5; Dystonia 1; Dystonia 27; Dystonia 5, dopa-responsive; Dystonia with or without hyperphenylalaninemia, dopa-responsive, autosomal recessive; Early infant epileptic encephalopathy 13; Early infant epileptic encephalopathy 2; Early infant epileptic encephalopathy 8; Early infant epileptic encephalopathy 9; Early muscular encephalopathy; Ectoderm dysplasia syndactyly syndrome 1; Abstaghamonesia, ectoderm dysplasia, and cleft lip / palate syndrome 3; Ehlers-Danlos syndrome, classical type; Ehlers-Danlos syndrome, hydroxylysine deficiency; Ehlers-Danlos Syndrome, contractile type; Ehlers-Danlos syndrome, type 4; Eichsfeldt type congenital muscular dystrophy; elliptocytosis 3; endometrial cancer; plate acetylcholinesterase deficiency; vestibular aqueduct hypertrophy syndrome; enterokinase deficiency; epidermolysis bullosa simple, Koebner type; epilepsy, nocturnal frontal lobe, type 3; epilepsy, progressive muscular clonus 1A (Unverricht and Lundborg); epilepsy, progressive muscular clonus 2b; epileptic encephalopathy, early infant, 1; epileptic encephalopathy, early infant, 24; epileptic encephalopathy, early infant, 28; epileptic encephalopathy, early infant Child, 31; Epiphyseal dysplasia, Miura type; Idiopathic ataxia type 1; Idiopathic ataxia type 6; Idiopathic pain syndrome, familial, 3; Polycythemia, familial, 2; Polycythemia, familial, 3; Erythrokerat dermatitis with ataxia; Exudative vitreoretinopathy 1; Exudative vitreoretinopathy 5; Fabry disease; Fabry disease, cardiac variant; Factor V and Factor VIII, combined deficiency, 2; Familial amyloid nephropathy with urticaria and hearing loss; Familial breast cancer; Familial cold urticaria; Familial febrile seizures 8; Familial hemiplegic migraine type 3; Familial hypertrophic cardiomyopathy 1; Familial hypertrophic cardiomyopathy 10;Familial hypertrophic cardiomyopathy 11; Familial hypertrophic cardiomyopathy 20; Familial hypertrophic cardiomyopathy 23; Familial hypertrophic cardiomyopathy 4; Familial hypertrophic cardiomyopathy 6; Familial dysplasia, glomerulocystic kidney disease; Familial myasthenia gravis; Familial juvenile gout; Familial Mediterranean fever; Familial platelet disorder with myeloid malignancy; Familial porencephaly; Familial cutaneous porphyria; Familial visceral amyloidosis, ostertag type; Fanconi anemia, complementary group C; Fanconi anemia, complementary group F; Fanconi anemia, complementary group G; Fanconi anemia, complementary group J; Fanconi anemia, complementary group T; Far Barr lipogranulomatosis; Fetal hemoglobin quantity trait locus 1; Fetal hemoglobin quantity trait locus 6; fibrochondrosis; Focal epilepsy with or without intellectual disability and speech impairment; Focal glomerulosclerosis 6; Foveal dysplasia and presenile cataract syndrome; Anterior nasal dysplasia 1; Anterior nasal dysplasia 2; Frontotemporal dementia; Fructose biphosphatase deficiency; Fumarase deficiency; Galactosylceramide beta-galactosidase deficiency; Gallbladder disease 4; Gamstorp-Wohlfart syndrome; Ganglioside sialidase deficiency; Gangliosidosis GM1 Type 3; Gardner syndrome; GATA-1-associated thrombocytopenia with erythropoiesis; Gaucher disease; Gaucher disease type 3C; Gaucher disease, perinatal fatality; Gaucher disease, type 1; generalized epilepsy with febrile seizures, type 1; generalized epilepsy with febrile seizures, type 2; generalized epilepsy with febrile seizures, type 9; Gerstmann-Straussler-Scheinker syndrome; Glanzmann thrombosis; glaucoma 1, open angle, F; congenital glaucoma; global developmental delay; glucocorticoid deficiency 4; glutariculosis, type 1; glycogen storage Disease IIIa; Glycogen storage disorder IV, congenital neuromuscular; Glycogen storage disorder IXB; Glycogen storage heart disease, congenitally lethal; Glycogen storage disorder, type II; Glycogen storage disorder, type IV; Glycogen storage disorder, type V; Glycogen storage disorder, type VI; Glycosylphosphatidylinositol deficiency; Grey platelet syndrome; Gliss cell syndrome type 2; Developmental and intellectual disability, mandibular-facial dysplasia, microcephaly, cleft palate; Growth hormone insensitivity with immunodeficiency; Hemochromatosis type 1; Hemochromatosis type 3; Hemolytic anemia due to hexokinase deficiency;Hemolytic anemia and nonspherocytosis due to glucose phosphate isomerase deficiency; systemic hemosiderosis due to aceruloplasminemia; Hennekamm's lymphangiectasia-lymphedema syndrome; hereditary congenital acrodermatitis; hereditary angioedema type 1; hereditary breast and ovarian cancer syndrome; hereditary carcinomatosis syndrome; hereditary diffuse gastric cancer; hereditary diffuse leukoencephalopathy with spheroids; hereditary factor II deficiency; hereditary factor IX deficiency; hereditary factor VIII deficiency; hereditary factor XI deficiency; hereditary fructothuria; hereditary leiomyomatosis and renal cell carcinoma; hereditary lymphedema type I Tumor; Hereditary neuralgia-associated muscular atrophy; Hereditary nonpolyposis colorectal cancer type 5; Hereditary nonpolyposis colorectal tumor; Hereditary pancreatitis; Hereditary paraganglioma-pheochromocytoma syndrome; Hereditary pyropoichylocytosis; Hereditary sensory neuropathy type 1D; Hereditary sideroblastic anemia; Heterotaxy, visceral, X-linked; Ectopic formation; Hirschsprung's disease gangliocyte tumor; Histiocytic myeloid reticulocyte disease; Holoprosencephaly 11; Holoprosencephaly 2; Holoprosencephaly 3; Holoprosencephaly 4; Homocysteinemia due to MTHFR deficiency; Homocystinuria due to CBS deficiency; Hurler syndrome; Eosinophilic cell carcinoma of the thyroid gland; Hutchinson's disease Guilford syndrome; hypercalciuria, childhood, spontaneous remission; hypercholesterolemia; hyperexcitability syndrome; hereditary hyperexpressia; hyperferritinemia cataract syndrome; hyperlipoproteinemia, type I; hyperlipoproteinemia, type ID; hyperlysinemia; hyperornithinemia-hyperammonemia-homocitrlinuria syndrome; hyperproinsulinemia; severe binocular eccentricity with central facial prominence, myopia, intellectual disability, and bone fragility; hypertrophic cardiomyopathy; hypocalcemia, autosomal dominant type I; hypocalcemia, autosomal dominant type I, Bartter syndrome; chondrodysplasia; hypopigmentation with iron overload Microcytic anemia; hypoglycemia with hepatic glycogen synthase deficiency; hypogonadism with or without olfactory dysfunction13; dehiscence-induced X-linked ectodermal dysplasia; periodic paralysis due to hypokalemia1; hypomagnesemia1, intestinal; hypomagnesemia with ocular involvement5, renal; hypomagnesemia, seizures, and intellectual disability; hypomyelinating leukodystrophy7; hypomyelinating leukodystrophy with or without oligodentinopathy and / or hypogonadotropic hypogonadism8; hypoproteinemia, catabolic; hypothyroidism, congenital, non-thyroid,1;Hypothyroidism, congenital, non-thyroid, 5; Hypothyroidism, congenital, non-thyroid, 6; Hypotrichosis 6; Hypotrichosis-lymphedema-telangiectasia syndrome; I-cell disease; Ichthyosis vulgaris; Idiopathic basal ganglia calcification 5; Immunodeficiency 12; Immunodeficiency 23; Immunodeficiency 24; Immunodeficiency 30; Immunodeficiency 31a; Immunodeficiency 31C; Hyper-IgM1 type immunodeficiency; Inclusion body myopathy 2; Degeneration of the infantile cerebellar retina; Infant GM1 gangliosidosis; Infant hypophosphatasia; Infant nystagmus, X-linked; Insulin-resistant diabetes and melanosis; Intellectual disability; Intermediate-type maple syrup urine disease type 2; Invasive pneumococcal infection, recurrent isolated, 2; Iris-corneal-trabecular hypoplasia; Iron accumulation in the brain; Jackson-Weiss syndrome; Jacob's syndrome Creutzfeldt's disease; Joubert syndrome 23; Juvenile GM>1<gangliosidosis; Juvenile polyposis syndrome; Kabuki makeup syndrome; Kallmann syndrome 3; Kallmann syndrome 4; Kallmann syndrome 5; Kallmann syndrome 6; Keratoconus 1; Kohlschutter syndrome; Kugelberg-Welander disease; Lafora disease; Langer moderate dysplasia; Laron-type isolated somatotropin deficiency; Larsen syndrome, dominant type; Lchad deficiency with maternal acute fatty liver during pregnancy; Leber congenital amaurosis 13; Leber congenital amaurosis 4 Leber congenital monocular blindness9; Leigh disease; Leopard syndrome; Leopard syndrome1; Leopard syndrome2; Leprechaunism; Leli-Weil dyschocartilage; Lesch-Nyhan syndrome; Leukodystrophy, hypomyelination6; Leukoencephalopathy with ataxia; Leukoencephalopathy with brainstem and spinal cord involvement and elevated lactate; Leukoencephalopathy with white matter disappearance; Leydig cell aplasia; Lie-Fraumeni syndrome1; Limb-girdle muscular dystrophy; Limb-girdle muscular dystrophy, type 1B; Limb-girdle muscular dystrophy, type 1C; Limb-girdle muscular dystrophy , Type 1E; Limb-girdle muscular dystrophy, Type 2A; Limb-girdle muscular dystrophy, Type 2B; Limb-girdle muscular dystrophy, Type 2E; Limb-girdle muscular dystrophy, Type 2F; Limb-girdle muscular dystrophy, Type 2L; Limb-girdle muscular dystrophy-dystroglycanopathy, Type C1; Limb-girdle muscular dystrophy-dystroglycanopathy, Type C14; Limb-girdle muscular dystrophy-dystroglycanopathy, Type C2; Limb-girdle muscular dystrophy-dystroglycanopathy, Type C7; Lissencephaly 1; Long QT syndrome 1; Long QT syndrome 13; Long QT syndrome 15;QT prolongation syndrome 2; QT prolongation syndrome 9; QT prolongation syndrome, LQT1 subtype; long-chain 3-hydroxyacyl-CoA dehydrogenase deficiency; Lowe's syndrome; luteinizing hormone resistance, female; lymphoproliferative syndrome 1; lymphoproliferative syndrome 1, X-linked; Lynch syndrome I; Lynch syndrome II; macrothrombocytopenia, familial, Bernard-Surier type; macular dystrophy involving the central pyramidal region; Magid syndrome; malignant neoplasm of the esophagus; malignant neoplasm of the prostate; mandibular dysplasia; maple syrup urine disease; maple syrup urine disease type 1A; maple syrup urine disease type 2; Marfan syndrome Group; Marie Unna's hereditary hypotrichosis 1; Adult-onset type 2 diabetes in young people; Adult-onset type 3 diabetes in young people; Medium-chain acyl-CoA dehydrogenase; Mayer-Gorlin syndrome 5; Melnick-Fraser syndrome; MEN2 phenotype: Unclassified; MEN2 phenotype: Unknown; Menkes curly hair syndrome; Menopausal, spontaneous, age, quantitative trait locus 3; Intellectual disability 30, X-linked; Intellectual disability and microcephaly with pontine and cerebellar hypoplasia; Intellectual disability, autosomal dominant 13; Intellectual disability, autosomal dominant 16; Intellectual disability, autosomal dominant 29; Intellectual disability, autosomal dominant 38; Intellectual disability, autosomal dominant 7; Intellectual disability, Autosomal recessive 34; intellectual disability, autosomal recessive 49; intellectual disability, stereotyped motor disorders, epilepsy, and / or brain malformations; intellectual disability, symptomatic, Claes-Jensen type, X-linked; intellectual disability, X-linked, syndrome 13; intellectual disability, X-linked, syndrome 32; intellectual disability, X-linked, syndrome, Raymond type; intellectual disability, X-linked, syndrome, WU type; intellectual disability-hypotonic phase syndrome, X-linked, 1; merosine-deficiency congenital muscular dystrophy; metachromatic leukodystrophy; metaphysical chondrodysplasia, Schmidt type; methylcobalamin deficiency, cblg type; methylmalonic aciduria mut(0) type; microcephaly and chorioretinopathy, autosomal recessive; microcephaly with or without chorioretinopathy, lymphedema, or intellectual disability; microcytic anemia; penile; microophthalmos syndrome; microophthalmos syndrome; isolated microphthalmia; isolated microphthalmia; solitary microphthalmia, coloboma; microvascular complications of diabetes; mild non-PKU hyperphenylalaninemia; mitochondrial complex I deficiency; mitochondrial complex II deficiency; mitochondrial complex III deficiency; mitochondrial DNA depletion syndrome (encephalomyelopathy type); mitochondrial DNA depletion syndrome;Mitochondrial DNA depletion syndrome 9 (encephalomuscular disorder with methylmalonic aciduria); mitochondrial short-chain enoyl-CoA hydratase deficiency; mitochondrial trifunctional protein deficiency; Miyoshi muscular dystrophy 1; Miyoshi muscular dystrophy 3; Mohr-Tranevier syndrome; anomalous mosaic syndrome; Mowat-Wilson syndrome; mucolipidosis III gamma; mucopolysaccharidosis type VI; mucopolysaccharidosis, MPS-II; mucopolysaccharidosis, MP; S-III-B; Mucopolysaccharidosis, MPS-IS; Mucopolysaccharidosis, MPS-IV-A; Mucopolysaccharidosis, MPS-IV-B; Munke syndrome; Mlibrina syndrome; Multiple congenital anomalies; Multiple endocrine neoplasia, type 1; Multiple endocrine neoplasia, type 2; Multiple endocrine neoplasia, type 2a; Multiple epiphyseal dysplasia 1; Multiple epiphyseal dysplasia 5; Multiple exostosis type 2; Multiple pterygium syndrome Escobar type; Multiple sulfatase deficiencies; Corneal dermal transection; Myasthenia gravis, limb girdle, familial; Congenital myasthenic syndrome associated with acetylcholine receptor deficiency 9; Congenital, associated with acetylcholine receptor deficiency Myasthenic syndrome, 9; congenital myasthenic syndrome with presynaptic and postsynaptic defects; congenital myasthenic syndrome, tubular aggregation 2; congenital slow channel type myasthenic syndrome; myoclonus epilepsy myopathy sensory ataxia; myoclonus, familial cortical; myofibrilar myopathy 1; myokymia 1; muscle disorder with postural muscle atrophy, X-linked; myopathy, actin, congenital, excess thin myofilaments; myopathy, central nucleus; myopathy, distal, 1; myopathy, isolated mitochondria, autosomal dominant; myopathy, weight loss, X-linked, early onset, severe; congenital myotonia Diseases; Nail disorders, non-syndromic congenital, 8; Microphthalmia, 4; Narcolepsy, 7; Native American myopathy; Navajo neurohepatic disorder; Nemaline myopathy, 3; Neonatal hypotension; Neonatal insulin-dependent diabetes mellitus; Neonatal intrahepatic cholestasis due to citrin deficiency; Ovarian tumors; Kidney stones / osteoporosis, hypophosphatemia, 2; Nephronoplasia, 16; Nephronoplasia, 18; Nephrotic syndrome, type 10; Neu-Laxova syndrome, 1; Neurodegeneration with cerebral iron accumulation, 5; Pituitary diabetic diabetes insipidus; Nicolaides-Balyser syndrome; Niemann-Pick disease type C1; Niemann-Pick disease, type A; Niemann-Pick Niemann-Pick disease, type B; Niemann-Pick disease, type c1, juvenile type; Nonaka myopathy; Non-ketotic hyperglycinemia; Noonan syndrome 1; Noonan syndrome 5; Noonan syndrome 7; Noonan syndrome 8; Not provided; Not specified; Oculocutaneous albinism type 3; Oculopharyngeal muscular dystrophy; Ocular dysplasia; Optic atrophy 9; Optic atrophy and cataract, autosomal dominant; Optic hypoplasia and central nervous system abnormalities; Orofacial digital syndrome; Ornithine aminotransferase deficiency; Ornithine carbamoyltransferase deficiency; Orofacial fissure 11; Orofasio digital syndrome 6;Orotic aciduria; Osteogenesis imperfecta type 12; Osteogenesis imperfecta type 13; Osteogenesis imperfecta type III; Osteogenesis imperfecta with normal sclera, dominant type; Osteogenesis imperfecta, recessive perinatal lethality; Osteopetrosis autosomal dominant type 1; Osteopetrosis autosomal recessive type 7; Atopalate digital syndrome, type I; Thyroiditis periostosis syndrome; Pallister-Hall syndrome; Papillon-Leffe syndrome; Paraganglioma 1; Paraganglioma 4; Parathyroid carcinoma; Foramen of parietal disease 2; Parkinson's disease 1; Parkinson's disease 7; Parkinson's disease 9; Paroxysmal nocturnal hemoglobinuria 1; Partial hypoxanthine-guanine phosphoribosyltransferase deficiency; Exfoliative skin syndrome, apical type; Pelger-Hu XABT abnormalities; Pelizaeus-Merzbach disease; Pendred syndrome; Permanent neonatal diabetes mellitus; Peroxisome biosynthesis disorder 6B; Peroxisome biosynthesis disorder 9B; Peutz-Jeggers syndrome; Pfeiffer syndrome; Phenylketonuria; Pheochromocytoma; Phosphoglycerate kinase 1 deficiency; Phosphoribosyl pyrophosphate synthetase hyperactivity; Photosensitive trichothiodystrophy; Pearson syndrome; Palathoracic degeneration; Pit-Hopkins syndrome; Pit-Hopkins-like syndrome 2; Pitto-dependent hyperadrenocorticism; Pituitary hormone deficiency, complex type 1; Pituitary hormone deficiency, complex type 4; Pituitary hormone deficiency, complex type 5; Platelet-type bleeding disorder 16; Agglutinating red blood cell syndrome; Polyarteritis nodosa; Polycystic kidney disease, infantile type; Polyglucosane myopathy 2; Polymicrogyria, bilateral Frontoparietal; polyneuropathy, hearing loss, ataxia, retinitis pigmentosa, cataracts; congenital cerebellar dysplasia, type 1B; congenital cerebellar dysplasia, type 1c; congenital cerebellar dysplasia, type 9; Poletti-Bold Schauser syndrome; preaxial polydactyly; early chromatid separation characteristics; premature ovarian dysfunction; premature ovarian failure; premature ovarian dysfunction; primary autosomal recessive microcephaly; primary autosomal recessive microcephaly; primary autosomal recessive microcephaly Microcephaly 5; Primary autosomal recessive microcephaly 6; Primary ciliary dyskinesia; Primary dilated cardiomyopathy; Primary familial hypertrophic cardiomyopathy; Primary hyperoxaluria, type I; Primary hyperoxaluria, type III; Primary focal cutaneous amyloidosis 1; Primary open-angle glaucoma with early onset 1; Primary pulmonary hypertension; Primary pulmonary hypertension 4; Primrose syndrome; Progressive myositis ossificans; Progressive polydystrophy;Proliferative vascular disorders and hydrocephalus - hydrocephalus syndrome; X-linked properdin deficiency; propionic acidemia; pseudo-Hurler's multiple dystrophy; pseudohypoaldosteronism type 1 autosomal dominant; pseudohypoaldosteronism type 2B; pseudohypoaldosteronism type 2; pseudohypoparathyroidism type 1A; pseudoxanthoma elastica; pseudoxanthoma elastica with multiple coagulation factor deficiencies; pulmonary hypertension associated with hereditary hemorrhagic telangiectasia; pulmonary glands Physiological and / or bone marrow dysfunction, related to telomeres; concentrated dysostosis; pyridoxine-dependent epilepsy; pyruvate dehydrogenase E1-alpha deficiency; radial dysplasia-thrombocytopenia syndrome; Reinne syndrome; rhasopathy; recessive dystrophy epidermolysis bullosa; Reifenstein syndrome; renal carnitine transport disorder; renal cell carcinoma, papillary,; renal dysplasia; renal hypouricemia; renal tubular acidosis with hemolytic anemia; retinal cone disease Strophophy 3A; Retinitis pigmentosa; Retinitis pigmentosa 10; Retinitis pigmentosa 11; Retinitis pigmentosa 14; Retinitis pigmentosa 2; Retinitis pigmentosa 25; Retinitis pigmentosa 33; Retinitis pigmentosa 35; Retinitis pigmentosa 4; Retinitis pigmentosa 43; Retinitis pigmentosa 50; Retinitis pigmentosa 56; Retinitis pigmentosa 73; Retinitis pigmentosa 74; Retinoblastoma; Rett syndrome; Rett syndrome, congenital variant; Rett syndrome, zappella variant; Rhabdoid tumor diathesis syndrome 2; Nodular chondrodysplasia type 1; Lienhoff syndrome; Roberts-SC phobia syndrome; Robinaud syndrome; RRM2B-associated mitochondrial disease; Rubinstein-Teyby syndrome; Sätre-Hötsen syndrome; Scapular muscle disorder, X-linked dominant; Schindler's disease, type 1; Schindler's disease, type 3; Schneider crystalline corneal dystrophy; Seckel syndrome 1; Seizures; Selective dentition 1 Tooth agenesis 1); Senior Loken syndrome 8; Sensory ataxic neuropathy, dysarthria, and ocular paresis; Sesame syndrome; Severe combined immunodeficiency due to adenosine deaminase deficiency; Severe combined immunodeficiency with microcephaly, developmental abnormalities, and susceptibility to ionizing radiation; Severe congenital neutropenia; Severe congenital neutropenia 4, autosomal recessive; Severe muscular clonus epilepsy in infancy; Severe X-linked myotubomyopathy; Short QT syndrome; Short QT syndrome 2; Short stature with nonspecific skeletal abnormalities;Short stature, external auditory canal atresia, mandibular dysplasia, skeletal abnormalities; short stature, idiopathic, autosomal; short stature, idiopathic, X-linked; brachydactyly chest dysplasia with or without polydactyly13; brachydactyly chest dysplasia with polydactyly14; pseudocostal dysplasia with or without polydactyly3; Sprinzen syndrome; Sprinzen-Goldberg syndrome; Schwakman syndrome; sialic acid storage disorder, severe infantile type; sialidosis, type II; sinus dysfunction syndrome2, autosomal dominant; sideroblastic anemia with B-cell immunodeficiency, periodic fever, and developmental delay; sitosterolemia; Sj \ xc3 \ XB6Gren-Larsson syndrome; Smith-Lemle-Oppitz syndrome; Sosby retinal degeneration; Sotos syndrome 1; Sotos syndrome 2; Charlevoix-Saguenay type spastic ataxia; Spastic paraplegia 11, autosomal recessive; Spastic paraplegia 30, autosomal recessive; Spastic paraplegia 4, autosomal dominant; Spastic paraplegia 54, autosomal recessive; Spastic paraplegia 6; Spastic paraplegia 7; Spastic paraplegia 8; Spermatogenesis imperfecta 8; Spherocytosis type 4; Sphingolipid activator protein 1 deficiency; Sphingomyelin / cholesterol lipidosis; Spinal muscle Atrophy, lower limb dominant 2, autosomal dominant; spinal muscular atrophy, type II; spinocerebellar ataxia 14; spinocerebellar ataxia 21; spinocerebellar ataxia 35; spinocerebellar ataxia 38; spinocerebellar ataxia, autosomal recessive 12; spondylocostal dysostosis 2; metaphyseal dysplasia with joint laxity; spondylomelastysplasia, Pakistani type; congenital spondyloesthetic dysplasia; spondylomelastysplasia with pyramidal-rod dystrophy; squamous cell carcinoma of the head and neck; Stargart disease 1; Stargart disease 3; Steele syndrome; Stickler syndrome type 1; Stiff skin syndrome Skin syndrome; stinging-associated vascular disorder, infantile onset; subacute neuropathic Gaucher disease; succinyl-CoA acetacetate transferase deficiency; extracellular elevation of superoxide dismutase; supravalvular aortic stenosis; syndactyly-brachydactyly syndrome; syndactyly type 9; Tangier disease; tarsal-carpal syndrome; Tay-Sachs disease; Tay-Sachs disease, B1 variant; T-cell prelymphocytic leukemia; Temple-Balyser syndrome; abdominal anterior axial brachydactyly syndrome; Tetralogy of Fallot; thoracic aortic aneurysm and aortic dissection; thrombocytopenia 2; thrombocytopenia, X-linked; thrombocytopenia, X-linked, intermittent; thrombosis due to activated protein C resistance;Thrombosis, hereditary, protein C deficiency, autosomal dominant; thrombosis, hereditary, protein C deficiency, autosomal recessive; thyroid cancer, non-medullary, 4; thyroid dysfunction, 1; periodic paralysis of thyroid toxicity; Teets syndrome; selective tooth aplasia, 3; selective tooth aplasia, X-linked, 1; transient neonatal diabetes mellitus, 1; transient neonatal diabetes mellitus, 2; Treacher Collins syndrome, 2; piloetabular dysplasia type 1; triglyceride storage disorder with ichthyosis; triose phosphate isomerase deficiency; triphalangeal thumb; tuberous sclerosis, 1; tuberous sclerosis, 2; tuberous sclerosis syndrome; tyrosinase-negative oculocutaneous albinism; tyrosinase-positive oculocutaneous albinism; tyrosinemia type 2; Ulrich-type congenital muscular dystrophy; unclassified; Unverlicht-Lundborg syndrome; Upshaw-Schulman syndrome; uridine-5; ‘ -Hemolytic anemia due to monophosphate hydrolase deficiency; Usher syndrome, type 1D; Usher syndrome, type 1F; Usher syndrome, type 2A; van der Waade syndrome; diverse porphyria; Vater syndrome with macrocephaly and ventricular hypertrophy; ventricular septal defect; vitamin D-dependent rickets, type 1; vitamin D-dependent rickets, type 2; vitamin K-dependent coagulation factor combined deficiency; oval dystrophy; von Hippel-Lindau syndrome; von Willebrand disease, type 2b; Waardenburg syndrome, type 1; Waardenburg syndrome, type 2E, no neurological involvement; Waardenburg syndrome, type 4A; Waardenburg syndrome, type 4B; Waardenburg syndrome, type 4C; Walker-Warburg congenital muscular dystrophy; Warburg microsyndrome; warts, hypogammaglobulinemia, infections, and myeloid cells Retention; Werdnig-Hoffmann disease; Werner syndrome; Wiacker syndrome; Wiedemann-Steiner syndrome; Winchester syndrome; Wolfram syndrome; Hereditary xerocytosis; Xeroderma pigmentosum, group D; Xeroderma pigmentosum, group G; X-linked gammaglobulinemia; X-linked hereditary motor and sensory neuropathies; X-linked ichthyosis with sterylsulfatase deficiency; X-linked intellectual disability; X-linked intellectual disability; X-linked periventricular ectopic formation; Zimmermann-Laband syndrome; or Zimmermann-Laband syndrome.
[0242] In some embodiments, the target DNA sequence includes a sequence associated with a disease or disorder. In some embodiments, the target DNA sequence includes a point mutation associated with a disease or disorder. In some embodiments, the point mutation associated with a disease or disorder is a gene associated with a disease or disorder. In some embodiments, the gene associated with a disease or disorder is AARS2, AASS, ABCA1, ABCA4, ABCB11, ABCB6, ABCC6, ABCC8, ABCD1, ABCG8, ABHD12, ABHD5, ACADM, ACAT1, ACE, ACO2, ACTA1, ACTB, ACTG1, ACTN2, ACVR1, ACVRL1, ADA, ADAMTS13, ADAR, ADGRG1, ADSL, AFF4, AGA, AGBL1, AGL, AGPAT2, AGRN, AGXT, A IPL1, AKR1D1, ALAD, ALAS2, ALDH3A2, ALDH7A1, ALDOB, ALG1, ALPL, ALS2, ALX3, ALX4, AMPD2, AMT, ANKS6, ANO5, APC, APOA1, APOE, A PP, APRT, AQP2, AR, ARHGEF9, ARID2, ARL6, ARSA, ARSB, ARSE, ARX, ASAH1, ASB10, ASPM, ATF6, ATL1, ATM, ATP13A2, ATP1A3, ATP6V1B 2, ATP7A, ATR, ATRX, AVP, B2M, B3GALT6, BAAT, BARD1, BBS10, BBS12, BBS2, BBS4, BBS9, BCKDHA, BCKDHB, BCS1L, BEST1, BHLHA9, BIC D2, BLM, BMP1, BMP4, BMPR2, BRAF, BRCA1, BRCA2, BRIP1, BTD, BTK, C10orf2, C1GALT1C1, C5orf42, C9, CA1, CACNA1S, CALM2, CANT1, CAPN3, CASK, CASQ2, CASR, CAV3, CBS, CCBE1, CCDC39, CD40LG, CDC6, CDC73, CDH1, CDH23, CDKL5, CDKN2A, CDON, CECR1, CENPJ, CEP1 20, CEP83, CFP, CFTR, CHAT, CHCHD10, CHD7, CHRNA1, CHRNB2, CHRNG, CHST14, CHSY1, CLCN1, CLCN2, CLCN5, CLCNKA, CLDN16, CLDN19,CLIC2、CLN6、CLN8、CNGA3、CNNM2、CNTNAP2、COA5、COL11A1、COL1A1、COL1A2 、COL27A1、COL2A1、COL3A1、COL4A1、COL4A5、COL5A1、COL5A2、COL6A1、COL6 A3、COL7A1、COLQ、COMP、CP、CPOX、CPT1A、CPT2、CR2、CRADD、CREBBP、CRH、CR X, CRYAB, CSF1R, CSTB, CTH, CTLA4, CTNS, CTPS1, CTSC, CTSD, CTSF, CTSK, CUL 3、CXCR4、CYBB、CYP1B1、CYP27A1、CYP27B1、CYP4F22、CYP4V2、CYP7B1、DARS 2、DBT、DCLRE1C、DCX、DDHD2、DES、DGUOK、DHCR24、DHCR7、DKC1、DLG3、DLL4、D MD, DMP1, DNAH11, DNAH5, DNAJB6, DNAJC19, DNM1, DNM2, DNMT1, DOCK6, DOK7, DOLK, DPAGT1, DPM2, DSC2, DSP, DYNC1H1, DYNC2H1, DYRK1A, DYSF, ECEL1, EC HS1、EDA、EDN3、EEF1A2、EFHC1、EFTUD2、EGLN1、EHMT1、EIF2B5、ELN、ELOVL4 、ELOVL5、EMP2、ENPP1、EOGT、ERCC2、ERCC8、ESCO2、ETFDH、EXOSC3、EXOSC8、 EXT2、EYA1、EYS、F12、F2、F5、F8、F9、FAM20C、FANCA、FANCF、FANCG、FAS、FBL N5、FBN1、FBN2、FBP1、FBXL4、FCGR3B、FGF8、FGFR1、FGFR2、FGFR3、FH、FHL1、F KTN、FLCN、FLG、FLNA、FLNB、FLT4、FLVCR2、FOXC1、FOXE1、FOXG1、FOXL2、FRA S1、FRMD7、FTL、FUS、G6PC3、G6PD、GAA、GABRA1、GABRG2、GAD1、GALC、GALNS、G ALT, GAMT, GARS, GATA1, GATA6, GBA, GBA2, GBE1, GCDH, GCH1, GKK, GDAP1, GDI1, GFAP, GGCX, GHR, GJA8, GJB1, GJB2, GK, GLB1, GLI3, GLRA1, GMPB, GNAI3GNAS, GNAT1, GNE, GNPTAB, GNPTG, GPI, GPIHBP1, GPT2, GRIA3, GRIN2A, GRIN2B, GRIP1, GRN, GSC, GUCY2D, GYG1, GYS2, H6PD, HADHB, HBB, HBD, HBG1, HBG2 HCN1、HCN4、HESX1、HEXA、HFE、HFM1、HGSNAT、HINT1、HK1、HMGCL、HNF1A、HNF 1B、HOGA1、HOXA1、HPD、HPGD、HPRT1、HR、HSD17B10、HSPB1、IDS、IDUA、IFT122 、IFT80、IGHMBP2、IKBKG、IL11RA、IL12RB1、IMPDH1、IMPG2、INF2、ING1、INP PL1、INSL3、INSR、IRF6、IRX5、ISPD、ITGA2B、ITGB3、ITK、JAGN1、KCNA1、KCNH 1、KCNH2、KCNJ1、KCNJ10、KCNJ11、KCNJ18、KCNJ2、KCNJ5、KCNK3、KCNQ1、KCN Q2、KCNQ4、KDM5C、KIAA0196、KIAA0586、KIF11、KIF1A、KIF2A、KISS1、KISS1R 、KLF1、KMT2A、KMT2D、KRAS、KRIT1、KRT1、KRT5、KRT6A、LAMA1、LAMA2、LAMB2 、LAMB3、LAMP2、LBR、LCT、LDLR、LIPA、LITAF、LMBR1、LMNA、LPIN2、LPL、LRIT3 、LRP5、LRRC6、LRTOMT、LYST、LYZ、MAD1L1、MAF、MALT1、MAN2B1、MAPK1、MAST L、MATN3、MC2R、MCCC1、MCCC2、MCFD2、MCM8、MCOLN1、MCPH1、MECP2、MEF2C、ME FV、MEN1、MESP2、MET、MFN2、MFSD8、MGAT2、MITF、MKKS、MLH1、MLYCD、MMACHC 、MMP14、MOG、MPL、MPV17、MPZ、MRE11A、MRPL3、MSH2、MSH6、MSR1、MSX1、MT-AT P6, MTHFR, MTM1, MT-ND1, MTR, MUSK, MUT, MYBPC3, MYC, MYH7, MYL2, MYL3, MYO1E, MYOC, NAGA, NAGLU, NARS2, NBEAL2, NBN, NDP, NDUFA1, NDUFA13, NDUFAF3NDUFS8、NEFL、NEU1、NEXN、NFIX、NHEJ1、NHLRC1、NIPA1、NIPBL、NKX2-5、NLR P3、NMNAT1、NNT、NOBOX、NOG、NOL3、NOTCH3、NPC1、NPR2、NR0B1、NR3C2、NR5A1 、NRXN1、NSD1、NSDHL、NT5C3A、NYX、OAT、OCA2、OCRL、OFD1、OPA3、OPCML、OSM R、OTC、OTOF、OTX2、OXCT1、PAFAH1B1、PAH、PAK3、PALB2、PANK2、PAPSS2、PARK 7、PAX2、PAX3、PAX6、PAX9、PCCA、PCCB、PCDH15、PCDH19、PCYT1A、PDE4D、PDE 6A、PDE6B、PDE6C、PDE6H、PDGFB、PDHA1、PET100、PEX10、PEX7、PGK1、PGM1、PG M3、PHGDH、PHKB、PHOX2B、PIEZO1、PIGM、PITPNM3、PITX2、PKHD1、PKP2、PLA2 G6、PLK4、PLOD1、PLP1、PMM2、PMP22、PMS2、PNPLA6、POLG、POLG2、POLR1A、POL R1D、POLR3A、POLR3B、POMT1、POMT2、POR、POU1F1、PPOX、PPT1、PRKACG、PRKA G2、PRKAR1A、PRKCG、PRNP、PROC、PROK2、PROKR2、PRPF31、PRPS1、PRSS56、PSA P、PSEN1、PTEN、PTPN11、PURA、PVRL4、PYGL、PYGM、RAB18、RAB27A、RAB7A、RA D21、RAD51C、RAF1、RAG2、RAX、RAX2、RB1、RBM8A、RDH12、RET、RHO、RIT1、RNF2 16、ROGDI、RP2、RPGR、RPS6KA3、RRM2B、RSPO4、RUNX1、RUNX2、RYR1、RYR2、SA CS、SAMHD1、SBDS、SCN11A、SCN1A、SCN2A、SCN5A、SCN8A、SCNN1B、SDHAF1、SDH B, SDHD, SEMA4A, SEPN1, SERPINF1, SERPING1, SETBP1, SGCB, SGCD, SH2D1A, SH3TC2, SHANK3, SHH, SHOX, SIGMAR1, SIX3, SKI, SLC11A2, SLC17A5, SLC19A3Selected from the group consisting of SLC1A3, SLC22A5, SLC25A13, SLC25A15, SLC25A19, SLC25A22, SLC25A38, SLC25A4, SLC26A4, SLC2A10, SLC2A9, SLC33A1, SLC35C1, SLC39A4, SLC46A1, SLC4A1, SLC52A2, SLC52A3, SLC5A5, SLC6A5, SLC6A8, SLC9A3R1, SMAD2, SMAD4, SMARCA2, SMARCA4, SMN1, SMPD1, SNCA, SNRNP200, SNRPB, SOD1, SOD3, SOX9, SPAST, SPATA5, SPG11, SPG7, SPTB, SRD5A2, SRY, STAC3, STAR, STAT1, STAT3, STAT5B, STK11, STS, STX1B, STXBP1, SUCLG1, SUMF1, TARDBP, TAZ, TBC1D24, TBX1, TBX20, TCF12, TCF4, TECTA, TERC, TERT, TFAP2B, TFR2, TGFB3, TGFBI, TGFBR2, TGIF1, TGM1, TGM5, TGM6, THRA, THRB, TIMM8A, TK2, TMEM173, TMEM240, TMEM98, TMPRSS15, TMPRSS3, TMPRSS6, TNFRSF11A, TNNI3, TNNT1, TOR1A, TP53, TP63, TPI1, TPM1, TPM2, TPM3, TPO, TPP1, TRIM37, TRNT1, TRPM6, TRPS1, TSC1, TSC2, TSHR, TSPAN12, TTPA, TTR, TUBB4A, TULP1, TYMP, TYR, TYRP1, UBE2T, UBE3A, UBIAD1, UMOD, UMPS, UROD, USH2A, USP8, VDR, VHL, VPS13B, VPS33B, VWF, WAS, WDR19, WDR45, WDR62, WDR72, WFS1, WNK4, WNT5A, WRN, WT1, WWOX, ZBTB20, ZC"
[0243] Several embodiments provide methods for using DNA-editing fusion proteins, as provided herein. In some embodiments, the fusion protein is used to introduce a point mutation into a nucleic acid by deaminating a target nucleic acid base, for example, a C residue. In some embodiments, the fusion protein is used to deaminate a target C to U and then removed to create an abasic site previously occupied by the C residue. In some embodiments, the deamination of the target nucleic acid base results in the correction of a gene deletion, for example, a correction of a point mutation leading to loss of function of a gene product. In some embodiments, the methods provided herein are used to introduce an inactivating point mutation into a gene or array encoding a gene product associated with a disease or disorder. For example, in some embodiments, a method for introducing an inactivating point mutation into an oncogene using a DNA-editing fusion protein (for example, in the treatment of proliferative disorders) is provided herein. In some embodiments, the inactivating mutation may generate an immature stop codon in the encoding sequence, which results in the expression of a shorter-than-normal gene product, for example, a shorter-than-normal protein lacking the function of a full-length protein.
[0244] In some embodiments, the objective of the methods provided herein is to restore the function of a dysfunctional gene through genome editing. The nucleic acid base editing proteins provided herein may be useful for in vitro human treatment based on gene editing by correcting disease-related mutations in human cell culture, for example. It will be understood by those skilled in the art that the nucleic acid base editing proteins provided herein, such as a fusion protein comprising a nucleic acid programmable DNA-binding protein (e.g., Cas9) and cytidine deaminase and a uracil-binding protein, may be used to correct either one point of a C→G or C→T mutation. In the first case, deamination of the C→U mutation, followed by excision of U, corrects the mutation, and in the latter case, deamination of C→U, followed by excision of U which is base-paired with the G mutation, corrects the mutation following a round of replication.
[0245] Successful modification of point mutations in disease-related genes and alleles opens up new strategies for gene modification with applications in therapy and basic research. Site-directed single-nucleotide modification systems, such as the disclosed fusion proteins including nucleic acid programmable DNA-binding proteins (napDNAbp), cytidine deaminase, and uracil-binding proteins, also have applications in “reverse” gene therapy, where certain gene functions are deliberately suppressed or lost. In those cases, site-directed residue mutations resulting in inactivating mutations in proteins can be used to cause loss or inhibition of protein function in vitro, ex vivo, or in vivo.
[0246] This disclosure provides methods for treating subjects diagnosed with diseases related to or caused by point mutations that can be corrected by DNA editing fusion proteins provided herein. For example, in some embodiments, methods are provided for administering to subjects having such diseases, such as cancers related to point mutations as described above, an effective amount of a base editing fusion protein that corrects a point mutation (e.g., a C→G or G→C point mutation) or introduces an inactivating mutation into a disease-related gene. In some embodiments, the disease is a proliferative disorder. In some embodiments, the disease is a genetic disorder. In some embodiments, the disease is a neoplasm. In some embodiments, the disease is a metabolic disorder. In some embodiments, the disease is a lysosomal storage disorder. Other diseases that can be treated by correcting point mutations or introducing inactivating mutations into disease-related genes will be known to those skilled in the art. This disclosure is not limited in this respect.
[0247] This disclosure provides a list of genes containing pathogenic G→C or C→G mutations. Such pathogenic G→C or C→G mutations may be modified by using methods and compositions provided herein, for example, by mutating from C to G and / or G to C, thereby restoring the function of the gene.
[0248] In some embodiments, the fusion protein recognizes classical PAMs and can therefore correct pathogenic G→C or C→G mutations in adjacent sequences, e.g., NGG. For example, a Cas9 protein that recognizes classical PAMs contains an amino acid sequence that is at least 80%, 85%, 90%, 95%, 97%, 98%, or 99% identical to the amino acid sequence of Streptococcus pyogenes Cas9 provided by SEQ ID NO: 6 or its fragment containing the RuvC and HNH domains of SEQ ID NO: 6.
[0249] It will be apparent to those skilled in the art that, in order to target any of the fusion proteins containing napDNAbp (e.g., a Cas9 domain) provided herein to a target site, e.g., a site containing a point mutation to be edited, it is typically necessary to co-express the fusion protein with a guide RNA, such as sgRNA. As will be described in more detail elsewhere in this specification, the guide RNA typically comprises a tracrRNA framework that enables Cas9 binding and a guide sequence that confers sequence specificity to the Cas9: nucleic acid editing enzyme / domain fusion protein. In some embodiments, the guide RNA comprises the structure 5'-[guide sequence]-guuuuagagcuagaaauagcaaguuaaaauaaaggcuaguccguuaucaacuugaaaaaguggcaccgagucggugcuuuuu-3' (SEQ ID NO: 119), where the guide sequence comprises a sequence complementary to the target sequence. In some embodiments, the guide sequence comprises a nucleic acid sequence complementary to the target nucleic acid. The guide sequence is typically 20 nucleotides long. Suitable guide RNA sequences for targeting Cas9: nucleic acid editing enzyme / domain fusion proteins to specific genomic target sites will be apparent to those skilled in the art based on this disclosure. Such suitable guide RNA sequences typically include a guide sequence that is complementary to a nucleic acid sequence within 50 nucleotides upstream or downstream of the target nucleotide to be edited.
[0250] Base editing factor efficiency Several aspects of this disclosure are based on the understanding that any of the base-editing factors provided herein can modify specific nucleotide bases without generating a significant proportion of indels. As used herein, “indel” refers to an insertion or deletion of a nucleotide base into a nucleic acid. Such insertions or deletions may result in frameshift mutations within the coding region of a gene. In some embodiments, it is desirable to create base-editing factors that efficiently modify (e.g., mutate or deaminate) specific nucleotides in a nucleic acid without generating a large number of insertions or deletions (i.e., indels) in the nucleic acid. In certain embodiments, any of the base-editing factors provided herein can generate a higher proportion of the intended modification (e.g., point mutation or deamination) compared to indels. In some embodiments, the base-editing factors provided herein can generate an intended point mutation to indel ratio greater than 1:1. In some aspects, the base editing factors provided herein can produce a point mutation to indel ratio that is at least 1.5:1, at least 2:1, at least 2.5:1, at least 3:1, at least 3.5:1, at least 4:1, at least 4.5:1, at least 5:1, at least 5.5:1, at least 6:1, at least 6.5:1, at least 7:1, at least 7.5:1, at least 8:1, at least 10:1, at least 12:1, at least 15:1, at least 20:1, at least 25:1, at least 30:1, at least 40:1, at least 50:1, at least 100:1, at least 200:1, at least 300:1, at least 400:1, at least 500:1, at least 600:1, at least 700:1, at least 800:1, at least 900:1, or at least 1000:1, or more. The intended number of mutations and indels can be determined using any preferred method, such as the method used in the example below. In some embodiments, to calculate the indel frequency, sequencing reads are scanned for exact matches to two 10 bp sequences adjacent on either side of the window in which indels may occur. If no exact matches are found, the reads are excluded from the analysis.If the length of this indel window exactly matches that of the reference sequence, the read is classified as not containing an indel. If the indel window is two or more bases longer or shorter than the reference sequence, the sequencing read is classified as an insertion or deletion, respectively.
[0251] In some embodiments, the base-editing factors provided herein can limit the formation of indels in a given region of nucleic acid. In some embodiments, the region is a nucleotide targeted by the base-editing factor, or a region of 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides within a nucleotide targeted by the base-editing factor. In some embodiments, any of the base-editing factors provided herein can limit the formation of indels in a given region of nucleic acid to less than 1%, less than 1.5%, less than 2%, less than 2.5%, less than 3%, less than 3.5%, less than 4%, less than 4.5%, less than 5%, less than 6%, less than 7%, less than 8%, less than 9%, less than 10%, less than 12%, less than 15%, or less than 20%. The number of indels formed in a given nucleic acid region may depend on the amount of time the nucleic acid (e.g., nucleic acid in the genome of a cell) is exposed to the base-editing factor. In some embodiments, the number or proportion of indels is determined at least 1 hour, at least 2 hours, at least 6 hours, at least 12 hours, at least 24 hours, at least 36 hours, at least 48 hours, at least 3 days, at least 4 days, at least 5 days, at least 7 days, at least 10 days, or at least 14 days after exposure of nucleic acid (e.g., nucleic acid in the genome of a cell) to a base editing factor.
[0252] Several aspects of this disclosure are based on the understanding that any of the base-editing factors provided herein can efficiently generate intended mutations, such as point mutations, in nucleic acids (e.g., nucleic acids in the genome of interest) without generating a significant number of unintended mutations, such as unintended point mutations. In some embodiments, the intended mutation is a mutation generated by a specific base-editing factor bound to a gRNA specifically designed to create an intended mutation. In some embodiments, the intended mutation is a mutation associated with a disease or disorder. In some embodiments, the intended mutation is a cytosine (C) → guanine (G) point mutation associated with a disease or disorder. In some embodiments, the intended mutation is a guanine (G) → cytosine (C) point mutation associated with a disease or disorder. In some embodiments, the intended mutation is a cytosine (C) → guanine (G) point mutation within the coding region of a gene. In some embodiments, the intended mutation is a guanine (G) → cytosine (C) point mutation within the coding region of a gene. In some embodiments, the intended mutation is a point mutation that creates a stop codon, e.g., an immature stop codon, within the coding region of a gene. In some embodiments, the intended mutation is a mutation that erases a stop codon. In some embodiments, the intended mutation is a mutation that alters the splicing of a gene. In some embodiments, the intended mutation is a mutation that alters the regulatory sequence of a gene (e.g., a gene promoter or a gene repressor). In some embodiments, any of the base editing factors provided herein can produce an intended mutation to unintended mutation ratio greater than 1:1 (e.g., intended point mutation:unintended point mutation).In some aspects, any of the base editing factors provided herein can produce an intended mutation to unintended mutation ratio (e.g., intended point mutation:unintended point mutation) of at least 1.5:1, at least 2:l, at least 2.5:1, at least 3:l, at least 3.5:1, at least 4:l, at least 4.5:1, at least 5:1, at least 5.5:1, at least 6:1, at least 6.5:1, at least 7:1, at least 7.5:1, at least 8:1, at least 10:1, at least 12:1, at least 15:1, at least 20:1, at least 25:1, at least 30:1, at least 40:1, at least 50:1, at least 100:1, at least 150:1, at least 200:1, at least 250:1, at least 500:1, or at least 1000:1, or more. It should be understood here that the characteristics of the base editing factors described in the "Base Editing Factor Efficiency" section can be applied to either the fusion protein provided here or the method using the fusion protein.
[0253] Methods for nucleic acid editing Several aspects of this disclosure provide methods for editing nucleic acids. In some embodiments, the methods are for editing nucleic acid bases (e.g., base pairs of a double-stranded DNA sequence) of nucleic acids. In some embodiments, the methods include a) contacting a target region of a nucleic acid (e.g., a double-stranded DNA sequence) with a complex comprising a base editing factor (e.g., a Cas9 domain fused to a cytidine deaminase and uracil-binding protein) and a guide nucleic acid (e.g., gRNA) so that the target region contains a target nucleic acid base pair; b) inducing strand separation in the target region; c) converting a first nucleic acid base of the target nucleic acid base pair on a single strand of the target region to a second nucleic acid base; d) excising the second nucleic acid base to create a debase site; and e) replacing a third nucleic acid base complementary to the first nucleic acid base with a fourth nucleic acid base, which is cytosine (C). In some embodiments, the methods result in less than 20% indel formation in the nucleic acid. It should be understood that in some embodiments, step b is omitted. In some embodiments, the first nucleic acid base is cytosine (C). In some embodiments, the second nucleic acid base is deaminated cytosine or uracil. In some embodiments, the third nucleic acid base is guanine. In some embodiments, the fourth nucleic acid base is cytosine (C). In some embodiments, the fifth nucleic acid base is ligated into the generated debase site in step (d). In some embodiments, the fifth nucleic acid base is guanine (G). In some embodiments, the method results in indel formation of 19%, 18%, 16%, 14%, 12%, 10%, 8%, 6%, 4%, 2%, 1%, 0.5%, less than 0.2%, or less than 0.1%. In some embodiments, at least 5% of the intended base pairs are edited. In some embodiments, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the intended base pairs are edited.
[0254] In some embodiments, the ratio of intended product to unintended product in the target nucleotide is at least 2:1, 5:1, 10:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or 200:1, or more. In some embodiments, the ratio of intended point mutation to indel formation is 1:1, 10:1, 50:1, 100:1, 500:1, or 1000:1, or more. In some embodiments, the single strand to be cleaved (the strand containing the nick) hybridizes to the guide nucleic acid. In some embodiments, the single strand to be cleaved is opposite to the strand containing the first nucleic acid base. In some embodiments, the base editing factor contains a Cas9 domain. In some embodiments, the base editing factor contains nickase activity. In some embodiments, the intended edited base pair is upstream of the PAM site. In some embodiments, the intended edited base pair is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides upstream of the PAM site. In some embodiments, the intended edited base pair is downstream of the PAM site. In some embodiments, the intended edited base pair is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides downstream of the PAM site. In some embodiments, the method does not require a classical (e.g., NGG) PAM site. In some embodiments, the nucleic acid base editing factor includes a linker. In some embodiments, the linker is 1 to 25 amino acids long. In some embodiments, the linker is 5 to 20 amino acids long. In some embodiments, the linker has a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids. In some embodiments, the target region includes a target window, and the target window includes a target nucleic acid base pair. In some embodiments, the target window includes 1 to 10 nucleotides. In some embodiments, the target window has a length of 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 to 2, or 1 nucleotide.In some embodiments, the target window is of length 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides. In some embodiments, the intended edited base pair is located within the target window. In some embodiments, the target window contains the intended edited base pair. In some embodiments, the method is carried out using one of the base editing factors provided herein. In some embodiments, the target window is a deamination window.
[0255] In some embodiments, the Disclosure provides methods for editing nucleotides. In some embodiments, the Disclosure provides methods for editing nucleic acid base pairs in a double-stranded DNA sequence. In some embodiments, the method comprises a) contacting a target region of a double-stranded DNA sequence with a complex comprising a base editing factor and a guide nucleic acid (e.g., gRNA) so that the target region contains a target nucleic acid base pair; b) inducing strand separation in the target region; c) converting a first nucleic acid base of the target nucleic acid base pair on a single strand of the target region to a second nucleic acid base; d) excising the second nucleic acid base to create a debase site; and e) replacing a third nucleic acid base complementary to the first nucleic acid base with a fourth nucleic acid base which is cytosine (C), with an efficiency of at least 5% in producing the intended edited base pair. It should be understood that in some embodiments, step b is omitted. In some embodiments, at least 5% of the intended base pair is edited. In some embodiments, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the intended base pairs are edited. In some embodiments, the method causes indel formation of less than 19%, 18%, 16%, 14%, 12%, 10%, 8%, 6%, 4%, 2%, 1%, 0.5%, 0.2%, or less than 0.1%. In some embodiments, the ratio of intended product to unintended product in the target nucleotide is at least 2:1, 5:1, 10:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or 200:1, or more. In some embodiments, the intended point mutation to indel formation ratio is 1:1, 10:1, 50:1, 100:1, 500:1, or 1000:1, or more. In some embodiments, the single strand to be cleaved hybridizes to a guide nucleic acid. In some embodiments, the nucleic acid base editing factor contains nickase activity. In some embodiments, the intended edited base pair is upstream of the PAM site.In some embodiments, the intended edited base pair is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides upstream of the PAM site. In some embodiments, the intended edited base pair is downstream of the PAM site. In some embodiments, the intended edited base pair is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides downstream of the PAM site. In some embodiments, the method does not require a classical (e.g., NGG) PAM site. In some embodiments, the nucleic acid base editing factor includes a linker. In some embodiments, the linker is 1 to 25 amino acids long. In some embodiments, the linker is 5 to 20 amino acids long. In some embodiments, the linker is 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids long. In some embodiments, the target region includes a target window, and the target window includes a target nucleic acid base pair. In some embodiments, the target window includes 1 to 10 nucleotides. In some embodiments, the target window is 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 to 2, or 1 nucleotide long. In some embodiments, the target window is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides long. In some embodiments, the intended edited base pair resides within the target window. In some embodiments, the target window includes the intended edited base pair. In some embodiments, the nucleic acid base editing factor is one of the base editing factors provided herein.
[0256] Pharmaceutical composition Other aspects of this disclosure relate to pharmaceutical compositions comprising any of the base editing factors, fusion proteins, or fusion protein-gRNA complexes described herein. The term “pharmaceutical composition” as used herein refers to a composition formulated for pharmaceutical use. In some embodiments, the pharmaceutical composition further comprises a pharmaceutically acceptable carrier. In some embodiments, the pharmaceutical composition comprises additional agents (e.g., for specific delivery, extension of half-life, or other therapeutic compounds).
[0257] As used herein, the term “pharmaceutically acceptable carrier” means a pharmaceutically acceptable material, composition, or vehicle, such as a liquid or solid filler, diluent, excipient, or manufacturing aid (e.g., lubricant, magnesium talc, calcium or zinc stearate, stearic acid), or a solvent encapsulation material involved in transporting or delivering a compound from one site of the body (e.g., site of delivery) to another site (e.g., organ, tissue, or part of the body). A pharmaceutically acceptable carrier is “acceptable” in the sense that it is compatible with the other components of the formulation and is not harmful to the target tissue (e.g., physiological compatibility, sterility, physiological pH, etc.). Some examples of materials that can serve as pharmaceutically acceptable carriers include: (1) sugars such as lactose, glucose and sucrose; (2) starches such as corn starch and potato starch; (3) cellulose and its derivatives such as sodium carboxymethylcellulose, methylcellulose, ethylcellulose, microcrystalline cellulose and cellulose acetate; (4) powdered tragacanth; (5) malt; (6) gelatin; (7) lubricants such as magnesium stearate, sodium lauryl sulfate and talc; (8) excipients such as cocoa butter and suppository waxes; (9) oils such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil and soybean oil; (10) glycols such as propylene glycol; (1 1) Polyols such as glycerin, sorbitol, mannitol, and polyethylene glycol (PEG); (12) Esters such as ethyl oleate and ethyl laurate; (13) Agar; (14) Buffers such as magnesium hydroxide and aluminum hydroxide; (15) Alginic acid; (16) Water free of pyrogens; (17) Isotonic saline; (18) Ringer's solution; (19) Ethyl alcohol; (20) pH buffer solution; (21) Polyesters, polycarbonates and / or polyanhydrides; (22) Bulking agents such as polypeptides and amino acids; (23) Serum components such as serum albumin, HDL and LDL; (22) C2-C12 alcohols such as ethanol; and (23) Other non-toxic, suitable substances used in pharmaceutical formulations. Wetting agents, colorants, release agents, coating agents, sweeteners, flavoring agents, fragrances, preservatives and antioxidants may also be present in the formulation.The terms "excipient," "carrier," and "pharmaceutically acceptable carrier" are used interchangeably here.
[0258] In some embodiments, the pharmaceutical compositions are formulated for delivery to a target, for example, for gene editing. Preferred routes of administration of the pharmaceutical compositions described herein include, but are not limited to, topical, subcutaneous, percutaneous, intradermal, intrafocal, intraarticular, intraperitoneal, intrabladderal, transmucosal, gingival, intradental, intracochlear, transtympanic, intraorganic, epidural, subarachnoid, intramuscular, intravenous, intravascular, intraosseous, periophthalmosal, intratumoral, intracerebral, and intraventricular administration.
[0259] In some embodiments, the pharmaceutical compositions described herein are administered topically to the site of the disease (e.g., a tumor site). In some embodiments, the pharmaceutical compositions described herein are administered to the subject by injection, by catheter, by suppository, or by implant, the implant being a porous, non-porous, or gelatinous material, and including membranes such as cialastic membranes or fibers.
[0260] In other embodiments, the pharmaceutical compositions described herein are delivered by a controlled release system. In one embodiment, a pump may be used (see, for example, Langer, 1990, Science 249:1527-1533; Sefton, 1989, CRC Crit. Ref. Biomed. Eng. 14:201; Buchwald et al., 1980, Surgery 88:507; Saudek et al., 1989, N. Engl. J. Med. 321:574). In another embodiment, a polymer material may be used. (See, for example, Medical Applications of Controlled Release (Langer and Wise eds., CRC Press, Boca Raton, Fla., 1974); Controlled Drug Bioavailability, Drug Product Design and Performance (Smolen and Ball eds., Wiley, New York, 1984); Ranger and Peppas, 1983, Macromol. Sci. Rev. Macromol. Chem. 23:61. Also see Levy et al., 1985, Science 228:190; During et al., 1989, Ann. Neurol. 25:351; Howard et al., 1989, J. Neurosurg. 71:105.) Other controlled release systems are discussed, for example, in Langer's work mentioned above.
[0261] In some embodiments, pharmaceutical compositions are formulated according to routine procedures as compositions suitable for intravenous or subcutaneous administration to a subject, e.g., a human. In some embodiments, pharmaceutical compositions for administration by injection are solutions in sterile isotonic aqueous buffer. The pharmaceutical may also optionally include a solubilizer and a local anesthetic, such as a lignocaine, to relieve pain at the injection site. Generally, the components are supplied separately or mixed together as unit dosage forms, such as dry lyophilized powder or anhydrous concentrate in sealed containers, such as ampoules or sachets indicating the amount of the active agent. When the pharmaceutical is administered by infusion, it can be administered using an infusion bottle containing sterile pharmaceutical-grade water or saline. When the pharmaceutical composition is administered by injection, ampoules of sterile water or saline for injection can be provided so that the components can be mixed before administration.
[0262] Pharmaceutical compositions for systemic administration may be liquids, such as sterile saline, Ringer's lactate solution, or Hanks' solution. Furthermore, pharmaceutical compositions may be in solid form and can be redissolved or suspended immediately before use. Lyophilized forms are also intended.
[0263] The pharmaceutical composition can be contained within lipid particles or vesicles, such as liposomes or microcrystals, which are also suitable for parenteral administration. The particles can be in any preferred structure, such as monolayer or multilayer, as long as the composition is contained within them. The compound can be captured in "stabilized plasmid-lipid particles" (SPLPs) containing the fusogenic lipid dioleoylphosphatidylethanolamine (DOPE), low levels (5-10 mol%) of cationic lipids, and stabilized by polyethylene glycol (PEG) coating (Zhang YP et al., Gene Ther. 1999, 6:1438-47). Positively charged lipids such as N-[1-(2,3-dioleoyloxy)propyl]-N,N,N-trimethylammonium methyl sulfate, or "DOTAP," are particularly preferred for such particles and vesicles. The preparation of such lipid particles is well known. See, for example, U.S. Patent Nos. 4,880,635; 4,906,477; 4,911,928; 4,917,951; 4,920,016; and 4,921,757. Each of these is incorporated herein by reference.
[0264] The pharmaceutical compositions described herein can be administered or packaged, for example, as unit doses. As used in reference to the pharmaceutical compositions herein, the term “unit dose” refers to physically separate units suitable as unit doses to a subject, each unit comprising a predetermined amount of active material calculated to produce a desired therapeutic effect in relation to the necessary diluent, i.e., a carrier or vehicle.
[0265] Furthermore, the pharmaceutical composition may be provided as a pharmaceutical kit comprising (a) a container containing the compound of the present invention in lyophilized form (e.g., a fusion protein or base editing factor), and (b) a second container containing a pharmaceutically acceptable diluent for injection (e.g., sterile water). The pharmaceutically acceptable diluent can be used to reconstitute or dilute the lyophilized compound of the present invention. Optionally accompanying such container(s) may be a notice in a form prescribed by a government agency regulating the manufacture, use, or sale of a pharmaceutical or biological product, the notice reflecting approval by the agency for manufacture, use, or sale for human control purposes.
[0266] In another aspect, the product includes materials useful for treating the diseases described herein. In some embodiments, the product includes a container and a label. Suitable containers include, for example, bottles, vials, syringes, and test tubes. Containers can be formed from a variety of materials, such as glass or plastic. In some embodiments, the container holds a composition that is effective for treating the diseases described herein and may have a sterile access port. For example, the container may be an intravenous solution bag or vial with a stopper that can be pierced by a subcutaneous injection needle. The active agent in the composition is the compound of the present invention. In some embodiments, a label on or accompanying the container indicates that the composition is used to treat a selected disease. The product may further include a second container containing a pharmaceutically acceptable buffer, such as phosphate-buffered saline, Ringer's solution, or dextrose solution. It may further include other materials desirable from a commercial and user perspective, including accompanying documentation along with other buffers, diluents, filters, needles, syringes, and instructions for use.
[0267] Kits, vectors, cells Some aspects of this disclosure provide a kit comprising (a) a nucleic acid sequence encoding any fusion protein provided herein, and (b) a nucleic acid construct comprising a heterologous promoter that drives the expression of the sequence in (a). In some embodiments, the kit further comprises an expression construct encoding a guide RNA backbone, wherein the construct comprises a cloning site positioned to enable the cloning of a nucleic acid sequence identical or complementary to the target sequence into the guide RNA backbone.
[0268] Some aspects of this disclosure provide polynucleotides encoding napDNAbp (e.g., Cas9 protein) of fusion proteins provided herein. Some aspects of this disclosure provide vectors comprising such polynucleotides. In some embodiments, the vectors include heterologous promoters that drive the expression of the polynucleotides.
[0269] Some aspects of this disclosure provide any fusion protein provided herein, nucleic acid molecules encoding any fusion protein provided herein, complexes comprising any fusion protein and gRNA provided herein, and / or any vectors provided herein.
[0270] The description of the exemplary aspects of the reporter system above is provided for illustrative purposes only and is not intended to be limiting. Additional reporter systems, such as variations of the exemplary systems described in detail above, are also included in this disclosure.
[0271] example Cytosine (C) → guanine (G) base editing factors and manipulated specific repair via the generation of debase sites. The sequencing data for the HEK2, RNF2, and FANCF sites are given below. The data displayed represents the base edit value for the most edited C in the window. This is C6 for HEK2, C6 for RNF2, and C6 for FANCF. The sequences for three different sites before and after base editing are as follows: [ka] ; [ka] and [ka] For both HEK2 and RNF2, the non-target strand was sequenced (this strand contains G' complementary to the target C'). For FANCF, the target strand was sequenced (this strand contains target C's). Schematic diagrams for C→T base editing (e.g., using the C→T base editing factor BE3) and C→G base editing are shown in Figures 1 and 2. Certain DNA polymerases are known to replace the opposite base with G at the abasted site. One strategy to achieve C→G base editing is to induce the creation of an abasted site, then replenish or tether such polymerase to replace the opposite G at the abasted site with C. This would provide access to all editing factors if C and T are excised based on the predetermined base priority of polymerases and can be repaired using all polymerases.
[0272] Different fusion constructs are summarized below and shown in Table 1. UdgX is an isoform of UDG known to bind tightly to uracil with minimal uracil cleavage activity. *UdX_On is a mutant version of UdgX (Sang et al. NAR, 2015) that was observed to lack uracil excision activity in an in vitro assay by Sang et al. UdX_On is another mutant version of UdgX (Sang et al. NAR, 2015) that was observed to have increased uracil excision activity in the same in vitro assay reported by Sang et al. UDG is an enzyme responsible for excising uracil from DNA by creating abasic sites. Rev7 is a component of the Rev1 / Rev3 / Rev7 complex, known to incorporate C opposite abasic sites. Rev1 is the enzymatic component of the complex described above. Polymerases alpha, beta, gamma, delta, epsilon, gamma, eta, iota, kappa, lambda, mu, and nu are eukaryotic polymerases with different priorities for incorporating bases opposite abasic sites. [Table 1-1] [Table 1-2] The construct used in the example: BE_Full Length - This is a C→T base-editing factor construct containing cytidine deaminase, nCas9, and uracilglycosylase inhibitor (UGI) domains. [ka] [ka] BE3_NoUGI - This construct is the BE3 construct described above, but lacks the UGI domain. [ka] [ka] It is used in the Cas9 niccase sequence-BE3. [ka] [ka] Used in dCas9 sequence-BE2 [ka] BE3_UDG replaces UGI, UdgX variant, polymerase - In the construct below, NLS sequences are identified by underlining and linkers by italics. "[UGI]" shown in the sequences below refers to UDG, UDG variant (e.g., UDG, UdgX * Identify the insertion sites of (R107S), and UdgX_On(H109S)), Rev7, and Smug1 (rather than the UGI of BE3). "[Polymerase]" shown in the sequence below identifies the polymerase (e.g., Pol beta, Pol lambda, Pol etha, Pol mu, Pol iota, Pol kappa, Pol alpha, Pol delta, Pol gamma, and Pol neu) and Rev1. [ka] [ka] [ka] N-terminal UDG (inserting UDG(Tyr147Ala) or UDG(Asn204Asp)) + Cas9 nickase and polymerase at the C-terminus - In the construct below, NLS sequences are identified in underline and linkers in italics. "[UDG variant]" shown in the sequence below identifies the insertion sites of UDG Tyr147Ala and UDG Asn204Asp. "[Polymerase]" shown in the sequence below identifies the insertion sites of polymerases (e.g., Pol beta, Pol lambda, Pol etha, Pol mu, Pol iota, Pol kappa, Pol alpha, Pol delta, Pol gamma, and Pol neu), and Rev1. [ka] [ka]
[0273] Example 1: C→G Approach 1 - Increases debasement site formation If abasidation sites are generated more efficiently, the total flux via the C→G base editing pathway will increase. Schematic representations of the base editing factors used in this approach are shown in Figures 3 and 4. Using UdgX, a homologous species of UDG identified as tightly binding to uracil with minimal uracil cleavage activity, increases the amount of C→G editing. While we do not wish to be bound to a specific theory, UdgX, which is nearly covalently bound to U, mimics damage that triggers damage-overcoming polymerase-type repair. Furthermore, UdgX has a low level of catalytic activity, which, combined with its tight binding, leads to U cleavage and abasidation site formation. Abasidation site formation tolerates off-target products, and the preferential generation of this damage results in more products. This is indicated through different experiments and base editing factors and is illustrated in Figures 5 and 6.
[0274] The results of C→G base editing at the HEK2, RNF2, and FANCF sites in WT cells are shown in Figures 7 - 15 using seven base editing factors (BE3; BE3_UdgX; BE3_UdgX * ; BE2_UdgX_On; BE3_UdgX_On; BE2_UDG; and BE3_UDG). These figures show the results of C→G base editing at the most edited position (C6) at three representative sites with high, medium, and low tolerance to perturbation of the sequence from the standard C→T editing.
[0275] The results of C→G base editing at the HEK2, RNF2, and FANCF sites in UDG- / - cells using various C→G base editing factors (BE3; BEUDG_UdgX; BE2_UNG; BE3_UNG; BE2UdgX_On; BE3UdgX_On; and SMUG1) are shown in Figures 16 - 24.
[0276] The results of C→G base editing at the HEK2, RNF2, and FANCF sites in Rev1- / - cells using various C→G base editing factors (BE3; BE3_UdgX; BE2_UNG; BE3_UNG; BE2UdgX_On; BE3UdgX_On; and SMUG1) are shown in Figures 25 - 30.
[0277] The results of C→G base editing factors at the HEK2, RNF2, and FANCF sites in three respective cell types (WT, UDG- / - and REV1- / - cells) using various C→G base editing factors (BE3; BE3_UdgX; BE2_UNG; BE3_UNG; BE2UdgX_On; BE3UdgX_On; and SMUG1) are summarized in Figures 31 and 32.
[0278] Example 2: Increasing C incorporation to the opposite of the abasic site in the C→G approach An increased preference for C incorporation into the opposite chain at the abasic site should result in an increase in overall C→G base editing. A schematic diagram of this approach and the base editing factors used in this approach are illustrated in Figures 33 and 34. Various polymerases that may be used in this approach for C→G base editing are shown in Figure 35. Briefly, the generation of an abasic site results in the formation of a non-T product from C. Rev1 has dC transferase activity. Eliminating this pathway or altering how abasic injury is repaired should result in a new base editing factor. If this pathway is solely responsible for the formation of this product, Rev1- / - knockout cell lines should lack C→G editing. Fusion of various polymerases should result in repair of the opposite chain based on polymerase preference for the repair of the opposite chain at the abasic site, leading to an increase in C→G base editing. Exemplary base editing factors are illustrated in Figure 36.
[0279] The results of C→G base editing at HEK2, RNF2, and FANCF sites in WT cells using various base editing factors (BE3;BE3_UdgX;BE2_UdgX_On;BE3_UdgX_On;BE2_UDG; and BE3_UDG) are shown in Figures 37-39.
[0280] The stable-state kinetic parameters for single-nucleotide incorporations to the opposite side of the G nucleotide by human polymerases η, ι, κ, and REV1 are given in Table 2 (see Choi et al. J mol Bio. 2010). [Table 2]
[0281] The stable state kinetic parameters for single-nucleotide incorporation into the opposite side of the G nucleotide and the debase site by human polymerases α and δ / PCNA are given in Table 3. [Table 3] [Table 4]
[0282] Example 3: C→G Approach 3 - Increases both debase formation and C incorporation. A schematic diagram of base editing factors for increasing both debase formation and C incorporation for increased C→G base editing is illustrated in Figure 40. The addition of polymerases tethered to constructs, particularly Pol kappa, increases C→G base editing. The results of base editing at HEK2, RNF2, and FANCF sites using constructs tethered by either Pol kappa or Pol iota are shown in Figure 41. The results of base editing at cytosine residues in WT cells, at HEK2, RNF2, and FACF sites, using constructs tethered to additional polymerases are shown in Figures 42-47. UDG147 is an enzyme that directly removes T and increases C→G base editing (Figures 42-44), while UDG204 is an enzyme that directly removes C and increases C→G base editing (Figures 45-47).
[0283] Example 4: C→G Approach 4 - Increase C→G flux by eliminating alternative repair pathways. One way to improve C→G editing is to eliminate or downregulate alternative repair pathways. One example is the protein MSH2 - / - Figure 48 shows that removing the repair pathway can lead to an increase in C→G base editing. MSH2 using various base editing factors (BE3;BE3_UdgX;BE2_UdgX_On;BE3_UdgX_On;BE2_UDG; and BE3_UDG) - / - The results of C→G base editing at HEK2, RNF2, and FANCF sites in cells are shown in Figures 49-51.
[0284] Example 5: C→G Approach 5 - Expression of trans components One approach to identifying synergistic base-editing factor components is to express them together in trans in cells. Once base-editing factor components that induce C→G mutations (e.g., polymerases, uracil-binding proteins, base-removal enzymes, cytidine deaminases, and / or nucleic acid-programmable DNA-binding proteins) are identified, they can be tethered to generate base-editing factors. UDG and UdgX variants fused to and expressed in APOBEC-Cas9 nickase, and TLS polymerase co-overexpressed in trans, result in C→G editing at the RNF2 site. A schematic diagram illustrating the trans expression of the components is shown in Figure 52.
[0285] Figures 53–55 show the results of base editing in HEK2, RNF2, and FANCF in HEK293 cells using five different base editing factors (BE3;BE3_UdgX;BE2_UdgX_On;BE3_UdgX_On;BE2_UDG; and BE3_UDG) expressed in trans using various polymerases (Pol kappa, Pol etha, Pol iota, REV1, Pol beta, and Pol delta). References
number
number
[0286] Example 6: -Cas9 variant array This disclosure provides Cas9 variants, for example, Cas9 proteins from one or more organisms, which may contain one or more mutations (e.g., to create dCas9 or Cas9 nickase). In some embodiments, one or more amino acid residues of the Cas9 protein identified below by asterisks (asterek) may be mutated. In some embodiments, the corresponding mutation in any Cas9 provided herein, such as at the D10 and / or H840 residues of the amino acid sequence provided by SEQ ID NO: 6, or in any one of the amino acid sequences provided by SEQ ID NOs: 4-26, is mutated. In some embodiments, the corresponding mutation in any Cas9 provided herein, such as at the D10 residue of the amino acid sequence provided by SEQ ID NO: 6, or in any one of the amino acid sequences provided by SEQ ID NOs: 4-26, is mutated to any amino acid residue other than D. In some embodiments, the corresponding mutation in any Cas9 provided herein, such as at the D10 residue of the amino acid sequence provided by SEQ ID NO: 6, or in any one of the amino acid sequences provided by SEQ ID NOs: 4-26, is mutated to A. In some embodiments, the H840 residue in the amino acid sequence provided by SEQ ID NO: 6, or the corresponding residue in any Cas9, such as any amino acid sequence provided by SEQ ID NOs: 4-26, is H. In some embodiments, the corresponding mutation in any Cas9, such as the H840 residue in the amino acid sequence provided by SEQ ID NO: 6, or any amino acid sequence provided by SEQ ID NOs: 4-26, is a mutation to any amino acid residue other than H. In some embodiments, the corresponding mutation in any Cas9, such as the H840 residue in the amino acid sequence provided by SEQ ID NO: 6, or any amino acid sequence provided by SEQ ID NOs: 4-26, is a mutation to A. In some embodiments, the corresponding residue in any Cas9, such as the D10 residue in the amino acid sequence provided by SEQ ID NO: 6, or any amino acid sequence provided by SEQ ID NOs: 4-26, is D.
[0287] Cas9 sequences from various species were aligned to determine whether the corresponding homologous amino acid residues at D10 and H840 of SEQ ID NO: 6 were identified in other Cas9 proteins and whether this allowed for the generation of Cas9 variants with corresponding mutations in these homologous amino acid residues. Alignment was performed using the NCBI Constraint-based Multiple Alignment Tool (COBALT, accessible at st-va.ncbi.nlm.nih.gov / tools / cobalt) with the following parameters: Alignment parameters: Gap Penalties -11, -1; End-Gap Penalties -5, -1. CDD parameters: Use RPS BLAST on; Blast E-value 0.003; Find Conserved columns and Recompute on. Query clustering parameters: Use query clusters on; Word Size 4; Max cluster distance 0.8; Alphabet Regular.
[0288] An exemplary alignment of four Cas9 sequences is provided below. The Cas9 sequences in the alignment are: Sequence 1 (S1): SEQ ID NO: 23 | WP_010922251 | gi 499224711 | Type II CRISPR RNA-induced endonuclease Cas9 [Streptococcus pyogenes]; Sequence 2 (S2): SEQ ID NO: 24 | WP_039695303 | gi 746743737 | Type II CRISPR RNA-induced endonuclease Cas9 [Streptococcus gallolyticus]; Sequence 3 (S3): SEQ ID NO: 25 | WP_045635197 | gi 782887988 | Type II CRISPR RNA-induced endonuclease Cas9 [Streptococcus mitis]; Sequence 4 (S4): SEQ ID NO: 26 | 5AXW_A | gi 924443546 | Staphylococcus Aureus Cas9. The HNH domain (bold and underlined) and the RuvC domain (boxed) have been identified for each of the four sequences. The homologous amino acids in the aligned sequences to amino acid residues 10 and 840 in S1 are identified by an asterisk following each amino acid residue. [ka] [ka] [ka]
[0289] Alignment demonstrates that amino acid sequences and amino acid residues homologous to a reference Cas9 amino acid sequence or amino acid residue can be identified from Cas9 sequence variants that include, but are not limited to, Cas9 sequences from different species, by identifying amino acid sequences or residues that align with a reference sequence or reference residue using alignment programs and algorithms known in the art. This disclosure provides Cas9 variants in which one or more amino acid residues identified by asterisks in SEQ ID NOs. 23-26 (e.g., S1, S2, S3, and S4, respectively) are mutated as described herein. The residues D10 and H840 of Cas9 in SEQ ID NO. 6, which correspond to the residues identified by asterisks in SEQ ID NOs. 23-26, are hereafter referred to as “homologous” or “corresponding” residues. Such homologous residues can be identified by sequence alignment, for example, as described above, by identifying sequences or residues that align with a reference sequence or residue. Similarly, the identified mutations in SEQ ID NO: 6, for example, the mutations in the Cas9 sequence corresponding to the mutations at residues 10 and 840 of SEQ ID NO: 6, are referred to here as “homologous” or “corresponding” mutations. For example, for the four aligned sequences above, the mutations corresponding to the D10A mutation in SEQ ID NO: 6 or S1 (SEQ ID NO: 23) are D11A in S2, D10A in S3, and D13A in S4; the mutations corresponding to the H840A mutation in SEQ ID NO: 6 or S1 (SEQ ID NO: 23) are H850A in S2, H842A in S3, and H560A in S4.
[0290] Furthermore, several Cas9 sequences from different species have been aligned using the same algorithm and alignment parameters outlined above. Several Cas9 sequences from different species (SEQ ID NOs. 11-260 of “632 Publication”) have been aligned using the same algorithm and alignment parameters outlined above, and are shown, for example, in Patent Publication WO2017 / 070632 (“632 Publication”), titled “Nucleic Acid Base Editing Factor and its Use,” published on April 27, 2017, and incorporated herein by reference. Amino acid residues homologous to residues of other Cas9 proteins may be identified in this manner and used to introduce corresponding mutations into other Cas9 proteins. Amino acid residues homologous to residues 10 and 840 of SEQ ID NO. 6 were identified in the same manner outlined above. The alignments are provided herein and incorporated by reference. The HNH domain (bold and underlined) and the RuvC domain (boxed) have been identified for each of the four sequences (SEQ ID NOs. 23-26). Single residues corresponding to amino acid residues 10 and 840 of SEQ ID NO. 6 are boxed in SEQ ID NO. 23 in the alignment, allowing for the identification of the corresponding amino acid residues in the aligned sequence.
[0291] Equivalents and ranges, incorporated by reference Those skilled in the art will recognize, or can verify by routine experimentation alone, many equivalents to the particular aspects of the invention described herein. The scope of the invention is not intended to be limited to the foregoing description, but rather as described in the appended claims.
[0292] In a claim, articles such as “a,” “an,” and “the” may mean one or more unless otherwise indicated or evident from the context. A claim or description containing “or” between one or more members of a group is considered satisfactory if one, more than one, or all members of the group are present in, used in, or otherwise related to a given product or process, unless otherwise indicated or evident from the context. The invention includes embodiments in which exactly one member of the group is present in, used in, or otherwise related to a given product or process. The invention also includes embodiments in which more than one, or all members of the group are present in, used in, or otherwise related to a given product or process.
[0293] Furthermore, it should be understood that the invention encompasses all variations, combinations, and rearrangements in which one or more limitations, elements, clauses, descriptive terms, etc., from one or more claims or relevant parts of the description are introduced into another claim. For example, any claim dependent on another claim may be modified to include one or more limitations found in any other claim dependent on the same basic claim. Furthermore, if a claim describes a composition, it should be understood that, unless otherwise indicated or unless it would be obvious to a person skilled in the art that this would result in a contradiction or inconsistency, it includes a method of using the composition for any of the purposes disclosed herein and a method of producing the composition according to any of the manufacturing methods disclosed herein or other methods known in the art.
[0294] If elements are presented as a list, for example in Markush group format, it should be understood that each subgroup of the elements is also disclosed, and any element(s) can be removed from that group. It should also be noted that the term “comprising” is intended to be open, allowing for the inclusion of additional elements or steps. Generally, if an invention or aspect of an invention is referred to as including certain elements, features, steps, etc., it should be understood that a given aspect of the invention or part of the invention consists of or is essentially such elements, features, steps, etc. For brevity, those aspects are not specifically described here in these terms. Therefore, for each aspect of an invention including one or more elements, features, steps, etc., the invention also provides aspects consisting of or being essentially such elements, features, steps, etc.
[0295] Where a range is given, the endpoints are encompassed. Furthermore, in various embodiments of the invention, unless otherwise shown or evident from the context and / or the understanding of those skilled in the art, the values expressed as a range may take any particular value within the given range, down to one-tenth of the lower limit unit of the range, unless the context clearly states otherwise. Unless otherwise shown or evident from the context and / or the understanding of those skilled in the art, the values expressed as a range may take any subrange within a given range, and the endpoints of the subranges may be expressed with the same degree of precision as one-tenth of the lower limit unit of the range.
[0296] In addition, it will be understood that any specific aspect of the present invention may be expressly excluded from any one or more claims. Where a range is given, any value within that range may be expressly excluded from any one or more claims. Any aspect, element, feature, use, or aspect of any composition and / or method of the present invention may be excluded from any one or more claims. For the sake of brevity, not all aspects from which one or more elements, features, uses, or aspects have been excluded are expressly shown herein.
[0297] All publications, patents, and sequence databases listed herein, including those listed above, are incorporated herein in whole by reference, as if each individual publication or patent were specifically and individually incorporated by reference. In case of any conflict, this application shall prevail, including any definitions herein.
Claims
1. (i) Nucleic acid programmable DNA-binding protein (napDNAbp) domain, (ii) Cytidine deaminase domain, and (iii) Uracil-binding protein (UBP) A fusion protein containing, The napDNAbp domain is Cas9 nickase (nCas9) or nuclease-inactive Cas9 (dCas9). The fusion protein binds to the target nucleic acid molecule when the napDNAbp domain associates with a guide nucleic acid capable of specifically binding to the target nucleic acid molecule. UBP can remove uracil bases from target nucleic acid molecules. The aforementioned fusion protein.
2. Uracil-binding protein (UBP) (i) Uracil-modifying enzyme, (ii) Uracil base removal enzyme, or (iii) Uracil DNA glucosylase (UDG) enzyme The fusion protein according to claim 1, or any variant thereof.
3. UBP, (i) a uracil DNA glucosylase (UDG) enzyme having uracil DNA glucosylase (UDG) activity from humans, mice, rats, dogs, monkeys, cattle, rhesus macaques, chimpanzees, gorillas, or lampreys; and / or (ii) a UDG containing an amino acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 48(UDG), or a UDG containing the amino acid sequence of SEQ ID NO: 48(UDG); and / or (iii) SMUG1 containing an amino acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 53 (SMUG1), The fusion protein according to claim 1 or 2.
4. UBP, (a) UdgX containing an amino acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 49(UdgX), or UdgX containing the amino acid sequence of SEQ ID NO: 49(UdgX); (b) Sequence ID 50(UdgX * UdgX containing an amino acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence of ). * , or sequence number 50 (UdgX * UdgX containing the amino acid sequence of ) * is; or (c) UdgX_On containing an amino acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence of SEQ ID NO: 51 (UdgX_On), or UdgX_On containing the amino acid sequence of SEQ ID NO: 51 (UdgX_On). The fusion protein according to claim 1 or 2.
5. The fusion protein has the following structure: NH 2 -[cytidine deaminase domain]-[napDNAbp domain]-[UBP]-COOH, or NH 2 -[napDNAbp domain]-[cytidine deaminase domain]-[UBP]-COOH This includes, where each example of "]-[" includes any linker, A fusion protein according to any one of claims 1 to 4.
6. The cytidine deaminase domain and the napDNAbp domain are fused via a linker, or the cytidine deaminase domain and the napDNAbp domain are fused via a linker containing one of the amino acid sequences of SEQ ID NOs: 102–109; and / or The napDNAbp domain and UBP are fused via a linker, or the napDNAbp domain and UBP are fused via a linker containing one of the amino acid sequences of SEQ ID NOs. 102–109. The fusion protein according to any one of claims 1 to 5.
7. The fusion protein, (a) Nucleic acid polymerase domains of eukaryotes; (b) DNA polymerase domain; (c) NAP domain with damage-overcoming polymerase activity; (d) Damage-overcoming DNA polymerase domain; (e) NAP domains selected from the Rev7, Rev1 complex, polymerase iota, polymerase kappa, or polymerase etha, or from the group of eukaryotic polymerases consisting of alpha, beta, gamma, delta, epsilon, gamma, etha, iota, kappa, lambda, mu, and nu; (f) A NAP domain containing an amino acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to any one of the amino acid sequences of SEQ ID NOs. 54–64; or (g) NAP domain containing one of the amino acid sequences from SEQ ID NOs. 54 to 64 A fusion protein according to any one of claims 1 to 6, further comprising (iv) a nucleic acid polymerase (NAP) domain selected from the above.
8. The fusion protein has the following structure: NH 2 -[cytidine deaminase domain]-[napDNAbp domain]-[UBP]-[NAP domain]-COOH; NH 2 -[cytidine deaminase domain]-[napDNAbp domain]-[NAP domain]-[UBP]-COOH; NH 2 -[Cytidine deaminase domain]-[NAP domain]-[napDNAbp domain]-[UBP]-COOH; or, NH 2 -[NAP domain]-[cytidine deaminase domain]-[napDNAbp domain]-[UBP]-COOH The fusion protein according to claim 7, comprising, where in any case of "]-[", comprising any linker.
9. The napDNAbp domain (i) Is it Cas9 nickase (nCas9)? (ia) an nCas9 having an amino acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to any one of sequence numbers 10, 13, 16, or 21, or (ib) nCas9 containing one of the amino acid sequences of sequence numbers 10, 13, 16, or 21; or (ii) Is it a nuclease-inactive Cas9 (dCas9)? (ii-a) dCas9 having an amino acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to any one of sequence numbers 7, 8, 9, or 22, or (ii-b) dCas9 containing one of the amino acid sequences of sequence numbers 7, 8, 9, or 22, The fusion protein according to any one of claims 1 to 8.
10. The cytidine deaminase domain is either a deaminase from the apolipoprotein B mRNA editing complex (APOBEC) family of deaminases, or The cytidine deaminase domain is selected from the group consisting of APOBEC1 deaminase, APOBEC2 deaminase, APOBEC3A deaminase, APOBEC3B deaminase, APOBEC3C deaminase, APOBEC3D deaminase, APOBEC3F deaminase, APOBEC3G deaminase, and APOBEC3H deaminase. A fusion protein according to any one of claims 1 to 9.
11. The cytidine deaminase domain, (i) Contains an amino acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to any one of the amino acid sequences of SEQ ID NOs. 67–101; (ii) Containing one of the amino acid sequences from sequence numbers 67 to 101; (iii) A rat APOBEC1 (rAPOBEC1) deaminase containing one or more mutations selected from the group consisting of W90Y, R126E, and R132E of Sequence ID No. 93, or one or more corresponding mutations in another APOBEC deaminase; (iv) A human APOBEC1 (hAPOBEC1) deaminase containing one or more mutations selected from the group consisting of W90Y, Q126E, and R132E of SEQ ID NO: 91, or one or more corresponding mutations in another APOBEC deaminase; (v) A human APOBEC3G (hAPOBEC3G) deaminase containing one or more mutations selected from the group consisting of W285Y, R320E, and R326E of SEQ ID NO: 77, or one or more corresponding mutations in another APOBEC deaminase; (vi) an activation-inducing deaminase (AID); or (vii) Cytidine deaminase 1 (pmCDA1) from Petromyzon marinus, The fusion protein according to any one of claims 1 to 10.
12. (i) Nucleic acid programmable DNA-binding protein (napDNAbp) domain, (ii-1) Base removal enzymes (BEEs) capable of removing bases from nucleic acid molecules; or (ii-2) Cytosine (C) or thymine (T) base removal enzyme, and (iii) Nucleic acid polymerase (NAP) domain A fusion protein containing, where The napDNAbp domain is Cas9 nickase (nCas9) or nuclease-inactive Cas9 (dCas9), and The fusion protein binds to the target nucleic acid molecule when the napDNAbp domain associates with a guide nucleic acid capable of specifically binding to the target nucleic acid molecule. The aforementioned fusion protein.
13. The base removal enzyme (BEE) contains an amino acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 48, and contains A at amino acid residue 156 of SEQ ID NO: 48, or Base removal enzyme (BEE) contains the amino acid sequence of SEQ ID NO: 65, The fusion protein according to claim 12.
14. Is the base removal enzyme a thymine (T) base removal enzyme, or The base removal enzyme contains an amino acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 48, and contains D at amino acid residue 213 of SEQ ID NO: 48, or The base removal enzyme contains the amino acid sequence of SEQ ID NO: 66, The fusion protein according to claim 12 or 13.
15. The nucleic acid polymerase (NAP) domain (a) Nucleic acid polymerase domains of eukaryotes; (b) DNA polymerase domain; (c) NAP domain with damage-overcoming polymerase activity; (d) Damage-overcoming DNA polymerase domain; (e) NAP domains selected from the group of eukaryotic polymerases consisting of Rev7, Rev1 complex, polymerase iota, polymerase kappa, or polymerase etha, or alpha, beta, gamma, delta, epsilon, gamma, etha, iota, kappa, lambda, mu, and neopolymerases; (f) A NAP domain containing an amino acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to any one of the amino acid sequences of SEQ ID NOs. 54–64; or (g) NAP domain containing one of the amino acid sequences from SEQ ID NOs. 54 to 64 A fusion protein according to any one of claims 12 to 14, selected from the above.
16. The fusion protein has the following structure: NH 2 -[napDNAbp domain]-[BEE]-[NAP domain]-COOH; NH 2 -[napDNAbp domain]-[NAP domain]-[BEE]-COOH; NH 2 -[NAP domain]-[napDNAbp domain]-[BEE]-COOH; NH 2 -[NAP domain]-[BEE]-[napDNAbp domain]-COOH; NH 2 -[BEE]-[napDNAbp domain]-[NAP domain]-COOH; or NH 2 -[BEE]-[NAP domain]-[napDNAbp domain]-COOH A fusion protein according to any one of claims 12 to 15, comprising, where in any case of "]-[", an arbitrary linker.
17. below: (i) The first domain is a nucleic acid programmable DNA-binding protein (napDNAbp) domain containing an amino acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to any one of the amino acid sequences of SEQ ID NOs: 7-10, 13, 16, 21, and 22; (ii) A cytidine deaminase second domain containing an amino acid sequence that is at least 90%, 95%, 98%, or 99% identical to any one of the amino acid sequences of SEQ ID NOs. 67–101; and (iii) A third domain which is a uracil DNA glycosylase (UDG) containing an amino acid sequence that is at least 90%, 95%, 98%, or 99% identical to any one of the amino acid sequences of SEQ ID NOs. 48-53. A fusion protein comprising, wherein the fusion protein binds to a target nucleic acid molecule when the napDNAbp domain associates with a guide nucleic acid capable of specifically binding to the target nucleic acid molecule.
18. (iv) A fourth domain of nucleic acid polymerase (NAP) containing an amino acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to any one of the amino acid sequences of SEQ ID NOs. 54–64. The fusion protein according to claim 17, further comprising:
19. below: (i) A first amino acid sequence containing one of the amino acid sequences of sequence numbers 7-10, 13, 16, 21, and 22; (ii) A second amino acid sequence containing any one of the amino acid sequences of SEQ ID NOs. 67 to 101; and, (iii) A third amino acid sequence containing one of the amino acid sequences of sequence numbers 48-53 A fusion protein containing [the specified ingredient].
20. (iv) A fourth amino acid sequence containing one of the amino acid sequences of SEQ ID NOs. 54–64 The fusion protein according to claim 19, further comprising:
21. A complex comprising a guide nucleic acid and a fusion protein according to any one of claims 1 to 20.
22. The complex according to claim 21, wherein the guide nucleic acid is (i) guide RNA (gRNA) or (ii) single guide RNA (sgRNA).
23. The complex according to claim 21 or 22, wherein the guide nucleic acid comprises a sequence of at least 10 consecutive nucleotides having a length of 15 to 100 nucleotides and being complementary to the target nucleic acid molecule.
24. A pharmaceutical composition comprising a fusion protein according to any one of claims 1 to 20 or a complex according to any one of claims 21 to 23.
25. The pharmaceutical composition according to claim 24, further comprising a pharmaceutically acceptable excipient.
26. A method comprising contacting a nucleic acid molecule with a fusion protein according to any one of claims 1 to 20, or a complex according to any one of claims 21 to 23, or a pharmaceutical composition according to claim 24 or 25, The method, which is carried out in vitro, ex vivo, or in vivo in a non-human animal.
27. The method according to claim 26, wherein the nucleic acid molecule is a DNA molecule.
28. The method according to claim 26 or 27, wherein the nucleic acid molecule is a DNA molecule of a genome.
29. A method for editing base pairs in a double-stranded DNA sequence, wherein the method is as follows: Contacting a double-stranded DNA sequence with a complex comprising a fusion protein and guide nucleic acid as described in any one of claims 1 to 20, or with a pharmaceutical composition as described in claim 24 or 25, wherein the double-stranded DNA sequence includes a target region containing a target base pair; This induces chain separation in the target region; This allows for the cleavage of cytosine or thymine within a single chain of the target region; This involves replacing cytosine or thymine with a second base that is different from both cytosine and thymine; and This results in edited base pairs in the double-stranded DNA sequence. Includes, The method, which is carried out in vitro, ex vivo, or in vivo in a non-human animal.
30. A polynucleotide encoding the fusion protein according to any one of claims 1 to 20.
31. A vector comprising the polynucleotide described in claim 30.
32. below: (i) The fusion protein according to any one of claims 1 to 20, (ii) The composite according to any one of claims 21 to 23, (iii) A nucleic acid molecule encoding the fusion protein described in any one of claims 1 to 20, or (iv) The vector according to claim 31 A cell that includes, The cells present in vitro or ex vivo.
33. A fusion protein according to any one of claims 1 to 20, a complex according to any one of claims 21 to 23, a pharmaceutical composition according to claim 24 or 25, a polynucleotide according to claim 30, a vector according to claim 31, or a cell according to claim 32, for use in treating a subject having or suspected to have a disease or disorder.
34. A method for editing base pairs in a double-stranded DNA sequence, wherein the method is as follows: Contacting a double-stranded DNA sequence with a complex comprising a fusion protein and a guide nucleic acid according to any one of claims 1 to 20, wherein the double-stranded DNA sequence includes a target region containing a target base pair; This induces chain separation in the target region; This involves converting the first base of a target base pair in a single strand of the target region to the second base; This process involves the removal of a second base from a double-stranded DNA sequence, creating a debasement site; This allows cutting only one strand of the target region; and This involves inserting a fifth base complementary to the fourth base into the debasement site, thereby generating the intended edited base pair; and This results in edited base pairs in the double-stranded DNA sequence. Includes, Here, the third base opposite the debasement site is replaced by the fourth base. The method, which is carried out in vitro, ex vivo, or in vivo in a non-human animal.
35. The method according to claim 34, which causes indel formation between less than 20% and less than 1%.
36. The efficiency of generating the intended edited base pairs is (i) at least 5%, or (ii) The method according to claim 34 or 35, wherein the percentage is between 10% and 50%.
37. The ratio of unintended edited base pairs to intended edited base pairs is between 2:1 and 10:1; and / or The method according to any one of claims 34 to 36, wherein the ratio of indel formation to the intended edited base pair is 2:1 to 200:
1.
38. The intended edited base pair is upstream of the PAM site; or The method according to any one of claims 34 to 37, wherein the intended edited base pair is downstream of the PAM site.
39. The intended edited base pair is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides upstream of the PAM site; or The method according to claim 38, wherein the intended edited base pair is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides downstream of the PAM site.
40. In vitro or ex vivo use of the fusion protein and guide nucleic acid bound to the fusion protein according to any one of claims 1 to 20 for editing base pairs of a DNA sequence.
41. The use according to claim 40, wherein the guide nucleic acid comprises a sequence of at least 10 consecutive nucleotides having a length of 15 to 100 nucleotides and being complementary to the target region of the DNA sequence.
42. The use according to claim 40 or 41, wherein the target base pair comprises a point mutation containing a cytosine (C) base associated with a disease or disorder.