Compositions and methods for editing mutations to allow transcription or expression
Programmable nucleobase editors are used to edit the SBDS gene in SDS, introducing functional mutations to address the molecular defects of Shwachman-Diamond syndrome, offering a potential cure by enhancing gene expression and improving patient survival.
Patent Information
- Application Number
- JP2022513443
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-08-29
- Filing Date
- 2020-08-28
- Publication Date
- 2025-12-08
- Estimated Expiration
- 2040-08-28
AI Technical Summary
There is currently no cure for Shwachman-Diamond syndrome (SDS), a rare autosomal recessive disorder characterized by exocrine pancreatic insufficiency, impaired hematopoiesis, and leukemia predisposition, with patients typically surviving only to approximately 35 years of age due to complications.
The use of programmable nucleobase editors to edit genes associated with SDS by splicing mutations that introduce functional gene products, such as altering stop codons, splice acceptor or donor sites, or modifying splice sites, using base editors with programmable DNA-binding domains and deaminase domains to target specific mutations in the SBDS gene.
This approach enables the production of functional SBDS gene products, potentially addressing the underlying molecular defects in SDS, thereby improving patient outcomes and survival rates.
Smart Images

Figure 0007781742000055 
Figure 0007781742000056 
Figure 0007781742000057
Abstract
Description
[Background technology]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is an international PCT application claiming priority to and benefit of U.S. Provisional Patent Application No. 62 / 893,638, filed August 29, 2019, the contents of which are incorporated herein by reference in their entirety.
[0002] Shwachman-Diamond syndrome (SDS) is a rare autosomal recessive multisystem disorder characterized by exocrine pancreatic insufficiency, impaired hematopoiesis, and leukemia predisposition. Patients with SDS present with bone marrow failure. Other clinical manifestations include skeletal, immune, hepatic, and cardiac disorders. Approximately 90% of patients with SDS have biallelic mutations in the evolutionarily conserved Shwachman-Bodian-Diamond syndrome (SBDS) gene located on chromosome 7. The SBDS protein plays a role in ribosome biogenesis and mitotic spindle stabilization, although its precise molecular function remains unknown. Currently, there is no cure for SDS, and patients with this disorder typically experience repeated hospitalizations due to complications and only survive to approximately 35 years of age on average. Therefore, improved methods and therapeutic agents for treating SDS are urgently needed. Summary of the Invention
[0003] As described below, the present invention features products, compositions, and methods for editing genes associated with Shwachman-Diamond Syndrome (SDS) using programmable nucleobase editors so that the genes are spliced to produce functional gene products.
[0004] In one aspect, a method of editing a polynucleotide to enable transcription is provided, the method comprising contacting the polynucleotide with a base editor in complex with one or more guide polynucleotides, wherein the base editor comprises a polynucleotide-programmable DNA-binding domain and a deaminase domain, and the one or more guide polynucleotides target the base editor to effect an alteration that introduces a mutation that enables transcription. In one embodiment, the mutation that enables transcription is a mutation that alters a stop codon, a mutation that introduces a splice acceptor or splice donor site, or a mutation that modifies the splice acceptor or splice donor site.
[0005] In one aspect, a method of editing an SBDS polynucleotide comprising a mutation associated with Shwakman-Diamond Syndrome (SDS) is provided, the method comprising contacting the SBDS polynucleotide with a base editor in complex with one or more guide polynucleotides, the base editor comprising a polynucleotide-programmable DNA-binding domain and a deaminase domain, wherein the one or more guide polynucleotides target the base editor to alter the mutation associated with Shwakman-Diamond Syndrome (SDS). In one embodiment of the method or its embodiments, the mutation associated with Shwakman-Diamond Syndrome (SDS) results from gene conversion. In one embodiment of the method or its embodiments, the mutation associated with Shwakman-Diamond Syndrome (SDS) induces a stop codon or alters splicing of the gene. In one embodiment of the method or its embodiments, the mutation associated with Shwakman-Diamond Syndrome (SDS) encodes an SBDS polypeptide having a truncation.
[0006] In one embodiment of any of the above-mentioned methods and embodiments thereof, the deaminase is a cytidine deaminase or an adenosine deaminase. In one embodiment, the deaminase is an adenosine deaminase. In several embodiments, the adenosine deaminase is selected from ABE8 or ABE8 variants listed herein, such as in Table 7A or Table 7B. In another embodiment of the above-mentioned methods and embodiments thereof, the deaminase is a cytidine deaminase. In one embodiment, the cytosine deaminase is selected from one or more of BE4, rAPOBEC1, PpAPOBEC1, PpAPOBEC1 containing an H122A substitution, AmAPOBEC1, SsAPOBEC2, RrA3F, RrA3F containing an F130L substitution, a variant of BE4 in which APOBEC-1 is replaced with the sequence of rAPOBEC1, a variant of BE4 in which APOBEC-1 is replaced with the sequence of AmAPOBEC1, a variant of BE4 in which APOBEC-1 is replaced with the sequence of SsAPOBEC2, a variant of BE4 in which APOBEC-1 is replaced with the sequence of PpAPOBEC1, or a variant of BE4 in which APOBEC-1 is replaced with the sequence of PpAPOBEC1 containing an H122A substitution. In one embodiment, the PpAPOBEC1 containing the H122A substitution, or the variant of BE4 in which APOBEC-1 is replaced with the sequence of PpAPOBEC1 containing the H122A substitution, further comprises one or more amino acid mutations selected from R33A, W90F, K34A, R52A, H121A, or Y120F. In some embodiments of the above-described methods and embodiments thereof, the two or more guide polynucleotides target base editors to alter two or more mutations associated with Shwachman-Diamond Syndrome (SDS).
[0007] In another aspect, a method for editing an SBDS polynucleotide containing a mutation associated with Shwachman-Diamond Syndrome (SDS) is provided, the method comprising contacting the SBDS polynucleotide with an adenosine base editor (ABE) in complex with one or more guide polynucleotides, wherein the base editor comprises a polynucleotide-programmable DNA-binding domain and a deaminase domain, and the one or more guide polynucleotides target the base editor to effect an A·T to G·C change at 183-184TA>CT Rs113993991, generating a missense mutation. In one embodiment, the one or more guide polynucleotides target one of the following sequences: TGTAAATGTTTCCTAAGGTC or AATGTTCCTAAGGTCAGGT. In one embodiment, the one or more sgRNAs comprise one of the following sequences: UGUAAAUGUUUCCUAAGGUC or AAUGUUUCCUAAGGUCAGGU. In one embodiment, the ABE has a PAM specificity of 5'-NGC-3' or 5'-NGG-3'.
[0008] In another aspect, there is provided a method of editing an SBDS polynucleotide comprising a mutation associated with Shwakman-Diamond Syndrome (SDS), the method comprising contacting the SBDS polynucleotide with a cytidine base editor in complex with one or more guide polynucleotides, wherein the cytidine base editor (CBE) comprises a polynucleotide-programmable DNA-binding domain and a cytidine deaminase domain, and the one or more guide polynucleotides target the base editor to effect a C·G to T·A change at rs113993993 258+2T>C. In one embodiment, the CBE has specificity for a 5'-NGC-3' PAM or a PAM comprising a 5'-NGC-3'. In one embodiment, the guide polynucleotide targets a polynucleotide target sequence selected from GTAAGCAGGCGGGTAACAGCTGC, AGCAGGCGGGTAACAGCTGCAGC, GCGGGTAACAGCTGCAGCATAGC, GTAAGCAGGCGGGTAACAGC, AGCAGGCGGGTAACAGCTGC, GCGGGTAACAGCTGCAGCAT, GCAGGCGGGTAACAGCTGC, CAGGCGGGTAACAGCTGC, AGGCGGGTAACAGCTGC, or AAGCAGGCGGGTAACAGCTGC. In one embodiment, the sgRNA comprises one of the following sequences: GUAAGCAGGCGGGUAACAGC; AGCAGGCGGGUAACAGCUGC; GCGGGUAACAGCUGCAGCA; GCAGGCGGGUAACAGCUGC, CAGGCGGGUAACAGCUGC, AGGCGGGUAACAGCUGC, or AAGCAGGCGGGUAACAGCUGC.
[0009] In other embodiments of any of the above-described methods and their embodiments, the contacting is within a cell, and the cell is a eukaryotic cell, a mammalian cell, or a human cell. In one embodiment, the cell is in vivo or ex vivo. In one embodiment of any of the above-described methods and their embodiments, the base editor introduces a missense mutation, inserts a new splice acceptor or splice donor site, and / or corrects a splice acceptor or splice donor site that contains a mutation. In one embodiment of any of the above-described methods and their embodiments, the polynucleotide-programmable DNA-binding domain is a Cas9 selected from Streptococcus pyrogenes Cas9 (SpCas9), Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), Streptococcus canis Cas9 (ScCas9), or a variant thereof. In one embodiment, the polynucleotide-programmable DNA-binding domain is wild-type or modified Streptococcus pyrogenes Cas9 (SpCas9), or a variant thereof. In one embodiment, the polynucleotide-programmable DNA-binding domain is a modified SpCas9 or SpCas9 variant. In one embodiment, the polynucleotide-programmable DNA-binding domain comprises a modified SpCas9 or SpCas9 variant with altered protospacer adjacent motif (PAM) specificity. In one embodiment, the SpCas9 has specificity for the PAM nucleic acid sequence 5'-NGC-3' or 5'-NGG-3'. In one embodiment, the SpCas9 is a modified SpCas9 or SpCas9 variant with specificity for a PAM nucleic acid sequence comprising the PAM nucleic acid sequence 5'-NGC-3' or 5'-NGC-3'. In one embodiment, the modified SpCas9 or SpCas9 variant comprises an amino acid sequence listed in Table 1. In one embodiment, the modified SpCas9 is spCas9-MQKFRAER. In one embodiment, the modified SpCas9 or SpCas9 variant comprises a combination of amino acid substitutions shown in Figures 3A-3C or 10.In one embodiment, the modified SpCas9 or SpCas9 variant comprises a combination of amino acid sequence substitutions selected from the following: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R(224 SpCas9);D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R(225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q(227 Cas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237 SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242 SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R, D1332, R1335N, and T1337 (244 SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E, and T1337 (245 SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q, and T1337R (259 SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A, and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGCRd1 SpCas9); or D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E, and T1337R(267 (NGCRd2SpCas9)
[0010] In other embodiments of any of the above-described methods and their embodiments, the polynucleotide-programmable DNA-binding domain is a nuclease-inactive or nickase variant. In one embodiment, the nickase variant comprises the amino acid substitution D10A or its corresponding amino acid substitution. In one embodiment, the deaminase domain is capable of deaminating adenosine or cytosine in deoxyribonucleic acid (DNA). In one embodiment, the adenosine deaminase or cytidine deaminase is a modified adenosine deaminase or cytidine deaminase that does not occur in nature. In one embodiment, the adenosine deaminase is TadA deaminase. In one embodiment, the TadA deaminase is TadA*7.10, TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23, or TadA*8.24. In one embodiment, TadA*7.10 contains one or more of the following changes: Y147T, Y147R, Q154S, Y123H, V82S, T166R, Q154R.
[0011] In one embodiment, TadA*7.10 comprises a combination of alterations selected from the group consisting of: Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y147T+Q154R; Y147T+Q154S; V82S+Q154S; and Y123H+Y147R+Q154R+I76Y.
[0012] In another embodiment of any of the aforementioned methods and embodiments thereof, the one or more guide RNAs comprise a CRISPR RNA (crRNA) and a transcoding small RNA (tracrRNA), wherein the crRNA comprises a nucleic acid sequence complementary to an SBDS nucleic acid sequence comprising an SDS-associated modification. In another embodiment of any of the aforementioned methods and embodiments thereof, the base editor is in a complex with a single guide RNA (sgRNA) comprising a nucleic acid sequence complementary to an SBDS nucleic acid sequence comprising an SDS-associated modification.
[0013] In another aspect, a cell is provided that is produced by introducing into the cell or its precursor a base editor comprising a polynucleotide-programmable DNA-binding domain and a deaminase domain, a polynucleotide encoding the base editor, and one or more guide polynucleotides that target the base editor to cause an alteration associated with aberrant splicing. In one embodiment, the cell or its precursor is an embryonic stem cell, an induced pluripotent stem cell, or a hematopoietic stem cell. In one embodiment, the cell expresses an SBDS protein. In one embodiment, the cell is derived from a subject suffering from Shwachman-Diamond syndrome (SDS). In one embodiment, the cell is a mammalian cell or a human cell. In one embodiment of the cell, the mutation or alteration results from a gene conversion containing a stop codon and / or mutation resulting from aberrant splicing. In one embodiment, the cell is selected for a gene conversion associated with SDS. In one embodiment, the polynucleotide-programmable DNA-binding domain is wild-type or modified Streptococcus pyrogenes Cas9 (SpCas9), or a variant thereof. In one embodiment, the polynucleotide programmable DNA-binding domain comprises a wild-type SpCas9 or a modified SpCas9 variant with altered protospacer adjacent motif (PAM) specificity. In one embodiment, the modified SpCas9 has specificity for the nucleic acid sequence 5'-NGC-3' or a PAM nucleic acid sequence comprising 5'-NGC-3'. In one embodiment, the modified SpCas9 is a Cas9 variant listed in Table 1. In one embodiment, the modified SpCas9 is spCas9-MQKFRAER. In one embodiment of the cell, the modified SpCas9 is an SpCas9 variant comprising a combination of amino acid substitutions shown in Figures 3A-3C or 10. In one embodiment of the cell, the SpCas9 variant comprises an amino acid sequence / substitution combination selected from the following: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R(224 SpCas9);D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R(225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q(227 Cas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237 SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242 SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R, D1332, R1335N, and T1337 (244 SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E, and T1337 (245 SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q, and T1337R (259 SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A, and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGCRd1 SpCas9); or D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E, and T1337R (267 (NGC Rd2 SpCas9). In an embodiment of the cell, the programmable polynucleotide binding domain is a nuclease-inactive variant or a nickase variant. In one embodiment, the nickase variant comprises the amino acid substitution D10A or its corresponding amino acid substitution. In one embodiment of the cell, the deaminase domain is a cytidine deaminase domain capable of deaminating cytidine in deoxyribonucleic acid (DNA) or an adenosine deaminase domain capable of deaminating adenosine in DNA. In one embodiment, the adenosine deaminase or cytidine deaminase is a modified adenosine deaminase or cytidine deaminase that does not occur in nature. In another embodiment of the cell, the adenosine deaminase is TadA deaminase. In one embodiment, the TadA deaminase is TadA*7.10, TadA*8.1, TadA*8.2, or TadA*8.3. , TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14 , TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23, or TadA*8.24. In one embodiment, TadA*7.10 comprises one or more of the following alterations: Y147T, Y147R, Q154S, Y123H, V82S, T166R, Q154R. In one embodiment, TadA*7.10 comprises a combination of alterations selected from the group consisting of: Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y147T+Q154R; Y147T+Q154S; V82S+Q154S.In another embodiment of the above cell, the cytosine deaminase is selected from one or more of BE4; rAPOBEC1, PpAPOBEC1, PpAPOBEC1 containing an H122A substitution, AmAPOBEC1, SsAPOBEC2, RrA3F, RrA3F containing an F130L substitution, a variant of BE4 in which APOBEC-1 is replaced with the sequence of rAPOBEC1, a variant of BE4 in which APOBEC-1 is replaced with the sequence of AmAPOBEC1, a variant of BE4 in which APOBEC-1 is replaced with the sequence of SsAPOBEC2, a variant of BE4 in which APOBEC-1 is replaced with the sequence of PpAPOBEC1, or a variant of BE4 in which APOBEC-1 is replaced with the sequence of PpAPOBEC1 containing an H122A substitution. In one embodiment, the PpAPOBEC1 containing the H122A substitution, or the variant of BE4 in which APOBEC-1 is replaced with the sequence of PpAPOBEC1 containing the H122A substitution, further comprises one or more amino acid mutations selected from R33A, W90F, K34A, R52A, H121A, or Y120F. In another embodiment of the cell, the one or more guide RNAs comprise a CRISPR RNA (crRNA) and a transcoding small RNA (tracrRNA), wherein the crRNA comprises a nucleic acid sequence complementary to an SBDS nucleic acid sequence comprising the SDS-associated alteration. In one embodiment of the cell, the base editor and one or more guide polynucleotides form a complex within the cell. In one embodiment, the base editor is in a complex with a single guide RNA (sgRNA) comprising a nucleic acid sequence complementary to an SBDS nucleic acid sequence comprising the SDS-associated gene conversion.
[0014] In another aspect, there is provided a method of treating Shwakman-Diamond Syndrome (SDS) or a disease associated with aberrant splicing in a subject in need thereof, the method comprising administering to the subject cells according to the above-described aspect and described embodiments thereof. In one embodiment of the method, the cells are autologous, allogeneic, or xenogeneic to the subject.
[0015] In another aspect, there is provided an isolated cell or cell population grown or expanded from a cell according to the above aspect and described embodiments thereof.
[0016] In another aspect, a method of treating Shwachman-Diamond syndrome (SDS) in a subject is provided, the method comprising administering to a subject in need thereof a base editor comprising a polynucleotide-programmable DNA-binding domain and a deaminase domain, or a polynucleotide encoding the base editor; and one or more guide polynucleotides that target the base editor to alter a mutation associated with SDS.
[0017] In another aspect, there is provided a method for treating a genetic disease associated with aberrant splicing in a subject, the method comprising administering to a subject in need thereof a base editor comprising a polynucleotide-programmable DNA-binding domain and a deaminase domain, or a polynucleotide encoding the base editor; and one or more guide polynucleotides that target the base editor to alter pathogenic mutations that alter splicing.
[0018] In one embodiment of the above-described method of treating Shwachman-Diamond Syndrome (SDS) in a subject or the above-described method of treating a genetic disease associated with aberrant splicing in a subject, the subject is a mammal or a human. In one embodiment, the above-described method includes delivering a base editor or a polynucleotide encoding the base editor and one or more guide polynucleotides to a cell of the subject. In one embodiment, the cell expresses a truncated polypeptide. In one embodiment of the above-described method, the modification converts a TAA stop to TGG in the SBDS polynucleotide. In another embodiment of the method, the modification changes K62X in the SBDS polypeptide associated with SDS. In another embodiment of the method, the SDS-associated gene conversion results in expression of a truncated SBDS polypeptide. In another embodiment of the method, the base editor modification replaces lysine (K) with tryptophan (W) at amino acid position 62. In another embodiment of the method, the polynucleotide-programmable DNA-binding domain comprises a modified Streptococcus pyrogenes Cas9 (SpCas9) or a variant thereof. In another embodiment of the method, the polynucleotide-programmable DNA-binding domain comprises a modified SpCas9 with altered protospacer adjacent motif (PAM) specificity. In one embodiment, the modified SpCas9 has specificity for the PAM nucleic acid sequence 5'-NGC-3' or a PAM nucleic acid sequence comprising 5'-NGC-3'. In one embodiment, the modified SpCas9 is a Cas9 variant listed in Table 1. In one embodiment, the modified SpCas9 is spCas9-MQKFRAER. In another embodiment of these methods, the modified SpCas9 is an SpCas9 variant comprising a combination of amino acid substitutions shown in Figures 3A-3C or 10. In one embodiment, the SpCas9 variant comprises a combination of amino acid sequence substitutions selected from the following: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R(224 SpCas9);D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R(225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q(227 Cas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237 SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242 SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R, D1332, R1335N, and T1337 (244 SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E, and T1337 (245 SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q, and T1337R (259 SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A, and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGCRd1 SpCas9); or D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E, and T1337R (267 (NGC Rd2 SpCas9). In other embodiments of the above methods and their embodiments, the polynucleotide-programmable DNA-binding domain is a nuclease-inactive variant. In one embodiment of the above methods, the polynucleotide-programmable DNA-binding domain is a nickase variant. In one embodiment, the nickase variant comprises the amino acid substitution D10A or its corresponding amino acid substitution. In one embodiment of the above methods, the deaminase domain is capable of deaminating adenosine or cytidine in deoxyribonucleic acid (DNA). In one embodiment, the deaminase domain is a non-naturally occurring modified adenosine deaminase or cytidine deaminase. In one embodiment, the adenosine deaminase is TadA deaminase. In one embodiment, the TadA deaminase is selected from TadA*7.10, TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23, TadA*8.24, TadA*8.25, TadA*8.26, TadA*8.27, Ta TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23, or TadA*8.24. In one embodiment, TadA*7.10 is or TadA*7.10 comprises one or more of the following alterations: Y147T, Y147R, Q154S, Y123H, V82S, T166R, Q154R; or TadA*7.10 comprises a combination of alterations selected from the group consisting of: Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y147T+Q154R; Y147T+Q154S; V82S+Q154S; and Y123H+Y147R+Q154R+I76Y.In another embodiment of the above-described methods and embodiments thereof, the deaminase domain is a cytidine deaminase selected from one or more of BE4, rAPOBEC1, PpAPOBEC1, PpAPOBEC1 containing an H122A substitution, AmAPOBEC1, SsAPOBEC2, RrA3F, RrA3F containing an F130L substitution, a variant of BE4 in which APOBEC-1 is replaced with the sequence of rAPOBEC1, a variant of BE4 in which APOBEC-1 is replaced with the sequence of AmAPOBEC1, a variant of BE4 in which APOBEC-1 is replaced with the sequence of SsAPOBEC2, a variant of BE4 in which APOBEC-1 is replaced with the sequence of PpAPOBEC1, or a variant of BE4 in which APOBEC-1 is replaced with the sequence of PpAPOBEC1 containing an H122A substitution. In one embodiment, the PpAPOBEC1 containing the H122A substitution, or the variant of BE4 in which APOBEC-1 is replaced with the sequence of PpAPOBEC1 containing the H122A substitution, further comprises one or more amino acid mutations selected from R33A, W90F, K34A, R52A, H121A, or Y120F. In an embodiment of the above-described method and embodiments thereof, the base editor targets the SNP rs113993993 258+2T>C in the SBDS polynucleotide sequence to restore correct splicing. In one embodiment of the above method, the one or more guide polynucleotides comprise a CRISPR RNA (crRNA) and a transcoding small RNA (tracrRNA), wherein the crRNA comprises a nucleic acid sequence complementary to the SBDS nucleic acid sequence comprising the gene conversion. In one embodiment, the base editor is in complex with a single guide RNA (sgRNA) comprising a nucleic acid sequence complementary to the SBDS nucleic acid sequence comprising the gene conversion associated with the SDS.
[0019] In another aspect, a method of producing a cell or a precursor thereof is provided, the method comprising: (a) In induced pluripotent stem cells containing gene conversions associated with Shwachman-Diamond syndrome (SDS), a base editor or a polynucleotide encoding a base editor comprising a polynucleotide-programmable nucleotide-binding domain and a cytidine deaminase domain or an adenosine deaminase domain; and One or more guide polynucleotides that target the base editor to alter the mutation associated with SDS. to introduce; and (b) differentiating the induced pluripotent stem cells or progenitors into a desired cell type. In one embodiment of this method, the mutation is a gene alteration associated with SDS. In one embodiment, the cell or progenitor is obtained from a subject suffering from SDS. In one embodiment, the cell or progenitor is a mammalian cell or a human cell. In another embodiment of this method, the polynucleotide-programmable DNA-binding domain comprises Streptococcus pyrogenes Cas9 (SpCas9), a modified Streptococcus pyrogenes Cas9 (SpCas9), or a variant thereof. In another embodiment, the polynucleotide-programmable DNA-binding domain comprises a modified SpCas9 with altered protospacer adjacent motif (PAM) specificity. In one embodiment of this method, the SpCas9 has specificity for the nucleic acid sequence 5'-NGG-3', and the modified SpCas9 has specificity for the nucleic acid sequence 5'-NGC-3' or a PAM nucleic acid sequence comprising 5'-NGC-3'. In one embodiment of this method, the modified SpCas9 is a Cas9 variant listed in Table 1, or the modified SpCas9 is spCas9-MQKFRAER. In another embodiment of this method, the modified SpCas9 is an SpCas9 variant comprising a combination of amino acid substitutions shown in Figures 3A-3C or 10. In one embodiment of this method, the SpCas9 variant comprises a combination of amino acid sequence substitutions selected from the following: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R(224 SpCas9);D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R(225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q(227 Cas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237 SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242 SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R, D1332, R1335N, and T1337 (244 SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E, and T1337 (245 SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q, and T1337R (259 SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A, and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGCRd1 SpCas9); or D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E, and T1337R (NGC Rd2 SpCas9). In one embodiment of the method, the polynucleotide programmable DNA binding domain is a nuclease-inactive or nickase variant. In one embodiment, the nickase variant comprises the amino acid substitution D10A or its corresponding amino acid substitution. In one embodiment of the method, the adenosine deaminase domain is capable of deaminating adenosines in deoxyribonucleic acid (DNA), and the cytidine deaminase is capable of deaminating cytosines in deoxyribonucleic acid (DNA). In one embodiment, the adenosine deaminase is a modified adenosine deaminase that does not occur in nature. In one embodiment, the adenosine deaminase is selected from the group consisting of TadA*7.10, TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8. In another embodiment of the method, the deaminase domain is selected from BE4, rAPOBEC1, PpAPOBEC1, PpAPOBEC1 containing a H122A substitution, AmAPOBEC1, SsAPOBEC2, RrA3F, RrA3F containing a F130L substitution, APOBEC-1 rAPOBECl, RrA3F containing a F130L substitution, The cytidine deaminase is selected from one or more of a variant of BE4 in which APOBEC-1 has been replaced with the sequence of APOBEC1, a variant of BE4 in which APOBEC-1 has been replaced with the sequence of AmAPOBEC1, a variant of BE4 in which APOBEC-1 has been replaced with the sequence of SsAPOBEC2, a variant of BE4 in which APOBEC-1 has been replaced with the sequence of PpAPOBEC1, or a variant of BE4 in which APOBEC-1 has been replaced with the sequence of PpAPOBEC1 containing an H122A substitution.In one embodiment, the PpAPOBEC1 containing the H122A substitution, or the variant of BE4 in which APOBEC-1 is replaced with the sequence of PpAPOBEC1 containing the H122A substitution, further comprises one or more amino acid mutations selected from R33A, W90F, K34A, R52A, H121A, or Y120F. In one embodiment of the method, the one or more guide polynucleotides comprise a CRISPR RNA (crRNA) and a transcoding small RNA (tracrRNA), wherein the crRNA comprises a nucleic acid sequence complementary to an SBDS nucleic acid sequence comprising the SDS-associated gene conversion. In one embodiment of the method, the base editor and the one or more guide polynucleotides form a complex in a cell. In one embodiment of the method, the base editor is in a complex with a single guide RNA (sgRNA) comprising a nucleic acid sequence complementary to an SBDS nucleic acid sequence comprising the SDS-associated gene conversion.
[0020] In another embodiment, a guide RNA is provided that comprises a 5' to 3' nucleic acid sequence selected from one or more of the following, or a 1, 2, 3, 4, or 5 nucleotide 5' truncation fragment thereof: GUAAGCAGGCGGGUAACAGC; AGCAGGCGGGUAACAGCUGC; GCGGGUAACAGCUGCAGCAU; UGUAAAUGUUUCCUAAGGUC; AAUGUUUCCUAAGGUCAGGU, GCAGGCGGGUAACAGCUGC, CAGGCGGGUAACAGCUGC, AGGCGGGUAACAGCUGC, and AAGCAGGCGGGUAACAGCUGC.
[0021] In another aspect, a base editor system for editing pathogenic mutations in an SBDS gene is provided, the base editor system comprising: (a)(i) a polynucleotide-programmable DNA-binding domain; (ii) a deaminase domain capable of deaminating a polynucleotide or its complementary nucleobase present in the SBDS gene conversion; base editors, including: (b) a guide polynucleotide associated with a polynucleotide-programmable DNA-binding domain, wherein the guide polynucleotide directs the base editor to a target polynucleotide sequence, at least a portion of which is located in the SBDS gene, the SBDS pseudogene, or the reverse complement thereof. Includes; Deamination of the polynucleotide or its complementary nucleobases allows transcription of the SBDS gene.
[0022] In another aspect, a base editor system for editing a mutation in a gene that results in aberrant splicing is provided, the base editor system comprising: (a)(i) a polynucleotide-programmable DNA-binding domain; (ii) a deaminase domain capable of deaminating the mutation or its complementary nucleobase resulting in aberrant splicing; and (b) a guide polynucleotide associated with a polynucleotide-programmable DNA-binding domain, wherein the guide polynucleotide directs the base editor to a target polynucleotide sequence, at least a portion of which is located in the gene, or its reverse complement. Includes; Mutation or deamination of the complementary nucleobase allows transcription.
[0023] In another aspect, a method is provided for editing a pathogenic mutation in a gene that results in aberrant splicing, the method comprising: a target nucleotide sequence, at least a portion of which is located in the gene or its reverse complement, (i) a polynucleotide-programmable DNA-binding domain in association with a guide polynucleotide that directs a base editor to a target polynucleotide sequence, at least a portion of which is located in the gene or its reverse complement; (ii) a deaminase domain capable of deaminating the pathogenic mutation or its complementary nucleobase that results in aberrant splicing; contacting a base editor comprising: Editing a pathogenic mutation by deaminating the pathogenic mutation or its complementary nucleobase when the base editor is directed against the target nucleotide sequence. Including, Deamination of the pathogenic mutation or its complementary nucleobase results in conversion of the pathogenic mutation to a sequence that allows splicing, thereby correcting the pathogenic mutation.
[0024] In another aspect, a method of editing a pathogenic mutation in an SBDS gene is provided, the method comprising: a target nucleotide sequence, at least a portion of which is located in the gene or its reverse complement, (i) a polynucleotide-programmable DNA-binding domain in association with a guide polynucleotide that directs a base editor to a target polynucleotide sequence, at least a portion of which is located in the gene or its reverse complement; (ii) a deaminase domain capable of deaminating the pathogenic mutation or its complementary nucleobase; contacting a base editor comprising: and editing the pathogenic mutation by directing a base editor at the target nucleotide sequence and deaminating the pathogenic mutation or its complementary nucleobase. Including, Deaminating the pathogenic mutation or its complementary nucleobase allows splicing, thereby editing the pathogenic mutation in the SBDS gene. In one embodiment of the above-described method of editing a pathogenic mutation, the pathogenic mutation in SBDS results in gene conversion. In one embodiment, the pathogenic mutation introduces a stop codon or alters splicing of the gene. In one embodiment, the pathogenic mutation encodes a polypeptide with a truncation. In one embodiment, the base editor introduces a missense mutation, inserts a new splice acceptor or splice donor site, or modifies a splice acceptor or splice donor site containing a mutation. In one embodiment, the base editor modifies a splice donor SNP site containing the rs113993993 C→T mutation in the SBDS gene.
[0025] In another aspect, a method is provided for treating SDS in a subject by editing a pathogenic mutation in the SBDS gene, the method comprising: administering to a subject in need thereof a base editor, or a polynucleotide encoding a base editor, wherein the base editor: (i) a polynucleotide-programmable DNA-binding domain; (ii) a deaminase domain capable of deaminating the nucleobase in the pathogenic mutation or its complementary nucleobase; administering; administering a guide polynucleotide to a subject, wherein the guide polynucleotide directs a base editor to a target nucleotide sequence located at least in part in the gene or its reverse complement; and and editing the pathogenic mutation by deaminating the pathogenic mutation or its complementary nucleobase in the SBDS gene when the base editor is directed against the target nucleotide sequence. Including, Deamination of the pathogenic mutation or its complementary nucleobase allows transcription or corrects the pathogenic mutation.
[0026] In another aspect, a method is provided for producing a cell, tissue, or organ for treating SBDS in a subject in need thereof by correcting a pathogenic mutation in the SBDS gene of the cell, tissue, or organ, the method comprising: Cells, tissues, or organs (i) a polynucleotide programmable DNA binding domain; (ii) a deaminase domain capable of deaminating the pathogenic mutation or its complementary nucleobase; contacting a base editor comprising: contacting the cell, tissue, or organ with a guide polynucleotide, wherein the guide polynucleotide directs a base editor to a target nucleotide sequence located at least in part in the gene or its reverse complement; and directing a base editor to target a target nucleotide sequence and editing said mutation by deaminating the pathogenic mutation or its complementary nucleobase. Including, Deamination of the pathogenic mutation or its complementary nucleobase allows splicing, thereby producing a cell, tissue, or organ for treating SDS. In one embodiment, the mutation results from gene conversion. In another embodiment, the mutation associated with Shwachman-Diamond Syndrome induces a stop codon in the gene or alters splicing. In another embodiment, the mutation associated with Shwachman-Diamond Syndrome (SDS) encodes a truncated SBDS polypeptide. In another embodiment, the base editor introduces a missense mutation, inserts a new splice acceptor or splice donor site, or modifies a splice acceptor or splice donor site containing a mutation. In another embodiment, the method includes administering the cell, tissue, or organ to a subject. In one embodiment of the method, the cell, tissue, or organ is autologous, allogeneic, or xenogeneic to the subject. In another embodiment of the method, the deaminase domain is a cytidine deaminase domain or an adenosine deaminase domain. In one embodiment, the adenosine deaminase domain can deaminate adenine in deoxyribonucleic acid (DNA), and the cytidine deaminase can deaminate cytosine in DNA.
[0027] In one embodiment of the above-described base editor system, editing method, or method of treatment, and any of its embodiments, the guide polynucleotide comprises ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). In one embodiment of the above-described base editor system, editing method, or method of treatment, and any of its embodiments, the guide polynucleotide comprises a CRISPR RNA (crRNA) sequence, a trans-activating CRISPR RNA (tracrRNA) sequence, or a combination thereof, wherein the crRNA comprises a nucleic acid sequence complementary to an SBDS nucleic acid sequence containing an SDS-associated modification. In one embodiment of the above-described base editor system, editing method, or method of treatment, and any of its embodiments, the base editor system or method further comprises a second guide polynucleotide. In one embodiment, the second guide polynucleotide comprises ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). In another embodiment, the second guide polynucleotide comprises a CRISPR RNA (crRNA) sequence, a trans-activating CRISPR RNA (tracrRNA) sequence, or a combination thereof. In one embodiment of the above-described base editor system, or method of editing, or method of treatment, and any of its embodiments, the polynucleotide-programmable DNA-binding domain lacks nuclease activity or is a nickase. In one embodiment of the above-described base editor system, or method of editing, or method of treatment, and any of its embodiments, the polynucleotide-programmable DNA-binding domain comprises a Cas9 domain. In one embodiment, the Cas9 domain comprises nuclease-free Cas9 (dCas9), Cas9 nickase (nCas9), or Cas9 with nuclease activity. In one embodiment, the Cas9 domain comprises a Cas9 nickase. In one embodiment of the above-described base editor system, or method of editing, or method of treatment, and any of its embodiments, the polynucleotide-programmable DNA-binding domain is an engineered or modified polynucleotide-programmable DNA-binding domain.In one embodiment of the base editor system, or editing method, or method of treatment, and any of its embodiments, the editing results in less than 20% indel formation, less than 15% indel formation, less than 10% indel formation, less than 5% indel formation, less than 4% indel formation, less than 3% indel formation, less than 2% indel formation, less than 1% indel formation, less than 0.5% indel formation, or less than 0.1% indel formation. In one embodiment of the base editor system, or editing method, or method of treatment, and any of its embodiments, the editing does not result in a translocation. In one embodiment of the base editor system, or editing method, or method of treatment, and any of its embodiments, the base editor corrects a splice donor SNP site comprising the rs113993993 C→T mutation in the SBDS gene.
[0028] In another aspect, there is provided a method of treating Shwakman-Diamond Syndrome (SDS) in a subject in need thereof, the method comprising administering to the subject a cell of the above-described aspect and described embodiments thereof.
[0029] In one embodiment of any of the above-described methods and embodiments thereof, the above-described cells and embodiments thereof, or the above-described base editor systems and embodiments thereof, or the above-described methods for editing, treating, producing cells or tissues, etc. and embodiments thereof, the base editor and / or its components are encoded by mRNA. In another embodiment of the above-described methods and embodiments thereof, the above-described cells and embodiments thereof, or the above-described base editor systems and embodiments thereof, or the above-described methods for editing, treating, producing cells or tissues, etc. and embodiments thereof, the base editor system or method of any of claims 126-157, wherein the base editor is in a complex with a single guide RNA (sgRNA) comprising a nucleic acid sequence complementary to an SBDS nucleic acid sequence. In one embodiment, the sgRNA comprises a nucleic acid sequence comprising at least 10 contiguous nucleotides complementary to an SBDS nucleic acid sequence. In another embodiment, the sgRNA comprises a nucleic acid sequence comprising 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 contiguous nucleotides complementary to an SBDS nucleic acid sequence. In another embodiment, the sgRNA comprises a nucleic acid sequence comprising 18, 19, or 20 contiguous nucleotides complementary to an SBDS nucleic acid sequence.
[0030] In another aspect, a composition is provided comprising a base editor bound to a guide RNA, wherein the guide RNA comprises a nucleic acid sequence complementary to the SBDS gene associated with Shwachman-Diamond Syndrome (SDS). In one embodiment, the base editor comprises an adenosine deaminase or a cytidine deaminase. In one embodiment, the adenosine deaminase is capable of deaminating adenosine in deoxyribonucleic acid (DNA). In one embodiment, the adenosine deaminase is a TadA deaminase selected from one or more of TadA*7.10, TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23, or TadA*8.24. In one embodiment, the cytidine deaminase is capable of deaminating cytidine in deoxyribonucleic acid (DNA). In another embodiment, the cytidine deaminase is APOBEC, A3F, or a derivative thereof. In one embodiment of the composition, the base editor is (i) containing Cas9 nickase; (ii) containing nuclease-inactive Cas9; (iii) SpCas9 variants comprising a combination of amino acid substitutions shown in Figures 3A-3C or Figure 10; (iv) SpCas9 variants comprising a combination of amino acid sequence substitutions selected from: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R(224 SpCas9);D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R(225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q(227 Cas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237 SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242 SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R, D1332, R1335N, and T1337 (244 SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E, and T1337 (245 SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q, and T1337R (259 SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A, and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGCRd1 SpCas9); or D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E, and T1337R(267 (NGC Rd2 SpCas9). (v) does not contain a UGI domain; and / or (vi) A cytidine deaminase selected from BE4, rAPOBEC1, PpAPOBEC1, PpAPOBEC1 containing an H122A substitution, AmAPOBEC1, SsAPOBEC2, RrA3F, RrA3F containing an F130L substitution, a variant of BE4 in which APOBEC-1 is replaced with the sequence of rAPOBEC1, a variant of BE4 in which APOBEC-1 is replaced with the sequence of AmAPOBEC1, a variant of BE4 in which APOBEC-1 is replaced with the sequence of SsAPOBEC2, a variant of BE4 in which APOBEC-1 is replaced with the sequence of PpAPOBEC1, or a variant of BE4 in which APOBEC-1 is replaced with the sequence of PpAPOBEC1 containing an H122A substitution. In one embodiment of the composition, in (vi), the PpAPOBEC1 containing an H122A substitution, or the variant of BE4 in which APOBEC-1 is replaced with the sequence of PpAPOBEC1 containing an H122A substitution, further comprises one or more amino acid mutations selected from R33A, W90F, K34A, R52A, H121A, or Y120F. In one embodiment, the composition further comprises a pharmaceutically acceptable excipient, diluent, or carrier.
[0031] In another aspect, a pharmaceutical composition for the treatment of Shwakman-Diamond Syndrome (SDS) is provided, the pharmaceutical composition comprising the composition of the above aspects and embodiments, and including a pharmaceutically acceptable excipient, diluent, or carrier. In one embodiment of the pharmaceutical composition, the gRNA and base editor are formulated together or separately. In one embodiment of the pharmaceutical composition, the gRNA comprises a 5' to 3' nucleic acid sequence selected from one or more of the following, or a 5' truncated fragment thereof of 1, 2, 3, 4, or 5 nucleotides: GCGGGUAACAGCUGCAGCAU;UGUAAAUGUUUCCUAAGGUC;AAUGUUUCCUAAGGUCAGGU, GCAGGCGGGUAACAGCUGC, CAGGCGGGUAACAGCUGC, AGGCGGGUAACAGCUGC, and AAGCAGGCGGGUAACAGCUGC. In one embodiment, the pharmaceutical composition further comprises a vector suitable for expression in a mammalian cell, wherein the vector comprises a polynucleotide encoding the base editor. In one embodiment of the pharmaceutical composition, the polynucleotide encoding the base editor is mRNA. In one embodiment of the pharmaceutical composition, the vector is a viral vector. In one embodiment, the viral vector is a retroviral vector, an adenoviral vector, a lentiviral vector, a herpesvirus vector, or an adeno-associated viral vector (AAV). In one embodiment, the pharmaceutical composition further comprises a ribonucleopartide suitable for expression in a mammalian cell.
[0032]
[0013] In one aspect, a pharmaceutical composition is provided comprising: (i) a nucleic acid encoding a base editor; and (ii) a guide RNA of the above aspect, such as a guide RNA comprising a 5' to 3' nucleic acid sequence selected from one or more of the following, or a 5' truncated fragment of the nucleic acid sequence by 1, 2, 3, 4, or 5 nucleotides, selected from: GUAAGCAGGCGGGUAACAGC;AGCAGGCGGGUAACAGCUGC;GCGGGUAACAGCUGCAGCAU;UGUAAAUGUUUCCUAAGGUC;AAUGUUUCCUAAGGUCAGGU, GCAGGCGGGUAACAGCUGC, CAGGCGGGUAACAGCUGC, AGGCGGGUAACAGCUGC, and AAGCAGGCGGGUAACAGCUGC. In one embodiment of the pharmaceutical composition of the above aspect or any of its embodiments, the pharmaceutical composition further comprises a lipid.
[0033] In one aspect, a method of treating Shwakman-Diamond Syndrome (SDS) is provided, the method comprising administering to a subject in need thereof a pharmaceutical composition of any of the above aspects and embodiments thereof.
[0034] In one aspect, there is provided a use of the pharmaceutical composition of any of the above aspects and embodiments thereof in treating Shwachman-Diamond Syndrome (SDS) in a subject. In one embodiment of the use, the subject is a human.
[0035] definition The following definitions supplement those in the art and are directed to this application, and are not to be construed as being in any way related or unrelated to any matter, for example, any co-owned patent or application. Although any methods and materials similar or equivalent to those described herein can be used in practice to test the present disclosure, the preferred materials and methods are described herein. Therefore, the terminology used herein is for the purpose of describing specific embodiments only and is not intended to be limiting.
[0036] Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by those skilled in the art to which this specification pertains. The following references provide those skilled in the art with general definitions of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them unless otherwise specified.
[0037] As used herein, the use of the singular includes the plural unless specifically stated otherwise. Note that, as used herein, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. As used herein, the use of "or" means "and / or" unless stated otherwise. Furthermore, the use of the term "including" as well as other forms such as "include," "includes," and "included" is not limiting.
[0038] As used in the specification and claims, the words "comprising" (and any form of comprising, such as "comprise" and "comprises"), "having" (and any form of having, such as "have" and "has"), "including" (and any form of including, such as "includes" and "include"), or "containing" (and any form of containing, such as "contains" and "contain") are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. It is contemplated that any embodiment discussed herein can be implemented with respect to any method or composition of the disclosure, and vice versa. Furthermore, the compositions of the disclosure can be used to practice the methods of the disclosure.
[0039] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which range will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within 1 or more standard deviations per the practice of the art. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, specifically with respect to biological systems or processes, the term can mean within an order of magnitude, e.g., within 5-fold, within 2-fold, of a value. When specific values are described in this application or claims, unless otherwise stated, the term "about" should be considered to mean within an acceptable error range for the specific value.
[0040] References herein to "some embodiments," "embodiments," "one embodiment," or "other embodiments" mean that certain particular features, structures, or characteristics described in connection with the embodiment are included in at least some embodiments of the present disclosure, but not necessarily all embodiments.
[0041] "Adenosine deaminase" refers to a polypeptide or fragment thereof that can catalyze the hydrolytic deamination of adenine or adenosine. In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine to inosine or deoxyadenosine to deoxyinosine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminase (e.g., engineered adenosine deaminase, evolved adenosine deaminase) provided herein can be derived from any organism, such as a bacterium.
[0042] In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase derived from an organism. In some embodiments, the deaminase or deaminase domain is not naturally occurring. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring deaminase. In some embodiments, the adenosine deaminase is derived from bacteria such as Escherichia coli, Staphylococcus aureus, Salmonella typhimurium, Shewanella putrefaciens, Haemophilus influenzae, or Caulobacter crescentus. In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is E. coli TadA (ecTadA) deaminase or a fragment thereof.
[0043] In some embodiments, the adenosine deaminase comprises an alteration in the following sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (Also known as TadA*7.10).
[0044] In some embodiments, TadA*7.10 comprises an alteration at amino acid 82 or 166. In specific embodiments, variants of the above reference sequences comprise one or more of the following alterations: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and Q154R. The alteration Y123H refers to the alteration H123Y in TadA*7.10 that reverts to Y123H TadA (wild type). In other embodiments, the variant of the TadA*7.10 sequence comprises a combination of changes selected from the group consisting of: Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y147T+Q154R; Y147T+Q154S; V82S+Q154S; and Y123H+Y147R+Q154R+I76Y.
[0045] In other embodiments, the invention provides adenosine deaminase variants that include deletions, e.g., TadA*8 that include C-terminal deletions beginning at residues 149, 150, 151, 152, 153, 154, 155, 156, or 157. In other embodiments, the adenosine deaminase variant is a TadA monomer (e.g., TadA*8) that includes one or more of the following alterations: Y147T, Y147R, Q154S, Y123H, V82S, T166R, Q154R. In other embodiments, the adenosine deaminase variant is a monomer comprising the following alterations: Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y147T+Q154R; Y147T+Q154S; V82S+Q154S; and Y123H+Y147R+Q154R+I76Y. In yet other embodiments, the adenosine deaminase variant is a homodimer comprising two adenosine deaminase domains, each domain having one or more of the following alterations: Y147T, Y147R, Q154S, Y123H, V82S, T166R, Q154R. In other embodiments, the adenosine deaminase variant is a heterodimer comprising a wild-type adenosine deaminase domain or a TadA*7.10 domain and an adenosine deaminase variant domain of TadA*7.10 (e.g., TadA*8) that contains one or more of the following alterations: Y147T, Y147R, Q154S, Y123H, V82S, T166R, Q154R. In other embodiments, the adenosine deaminase variant is a heterodimer comprising a TadA*7.10 domain and an adenosine deaminase variant of TadA*7.10 (e.g., TadA*8) comprising the following modifications: Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y147T+Q154R; Y147T+Q154S; V82S+Q154S; and Y123H+Y147R+Q154R+I76Y.
[0046] In one embodiment, the adenosine deaminase is TadA*8 comprising or consisting essentially of the following sequence, or a fragment thereof, having adenosine deaminase activity: MSEVEFSHEYWMRHALTLAKRADEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD.
[0047] In some embodiments, TadA*8 is truncated. In some embodiments, the truncated TadA*8 lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues compared to full-length TadA*8. In some embodiments, the truncated TadA*8 lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues compared to full-length TadA*8. In some embodiments, the adenosine deaminase variant is full-length TadA*8.
[0048] In specific embodiments, the adenosine deaminase heterodimer comprises a TadA*8 domain and an adenosine deaminase domain selected from one of the following: Staphylococcus aureus (S. aureus) TadA: MGSHMTNDIYFMTLAIEEAKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAH AEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGS LMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFKNLRANKKSTN Bacillus subtilis (B. subtilis) TadA: MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEML VIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMN LLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLSE Salmonella Typhimurium (S. Typhimurium) TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEG WNRPIGRHDPTHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIG RVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFRMRRQEIK ALKKADRAEGAGPAV Shewanella putrefaciens (S. putrefaciens) TadA: MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEI LCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGT VVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE Haemophilus influenzae F3031 (H. influenzae) TadA: MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTΑΗ AEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMMCAGAILHSRIKRLVFGASDYK TGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRREEKKIEKALLKSLSDK Caulobacter crescentus (C. crescentus) TadA: MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAH DPTAHAEIAAMRAAAAKLGNYRLTDLTLVVTLEPCAMCAGAISHARIGRVVFGADD PKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAKI Geobacter sulfreducens (G. sulfreducens) TadA: MSSLKKTPIRDDAYWMGKAIREAAKAAAARDEVPIGAVIVRDGAVIGRGHNLREGSN DPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERVVFGCYDP KGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRKKAKATPALF IDERKVPPEP TadA*7.10 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD.
[0049] "Administering" is referred to herein as giving one or more compositions described herein to a patient or subject. By way of example and not limitation, administration, e.g., injection, of a composition can be performed by intravenous (iv) injection, subcutaneous (sc) injection, intradermal (id) injection, intraperitoneal (ip) injection, or intramuscular (im) injection. One or more such routes can be employed. Parenteral administration can be performed, for example, by bolus injection or by gradual perfusion over time. Alternatively, or concurrently, administration can be performed orally.
[0050] By "agent" is meant any small molecule compound, antibody, nucleic acid molecule, or polypeptide, or fragment thereof.
[0051] "Alteration" means a change (increase or decrease) in the sequence, expression level, or activity of a gene or polypeptide as detected by standard art known methods, such as those described herein. As used herein, alteration includes a 10% change in expression level, a 25% change, a 40% change, and a 50% or greater change in expression level.
[0052] "Ameliorate" means to decrease, suppress, attenuate, diminish, arrest, or stabilize the development or progression of a disease.
[0053] "Analog" refers to a molecule that is not identical but has similar functional or structural properties. For example, a polypeptide analog retains the biological activity of the corresponding naturally occurring polypeptide and contains certain biochemical modifications that enhance the function of the analog relative to the naturally occurring polypeptide. Such biochemical modifications may increase the analog's protease resistance, membrane permeability, or half-life, for example, without altering ligand binding. Analogs may contain unnatural amino acids.
[0054] "Base editor (BE)" or "nucleobase editor (NBE)" refers to an agent that binds to a polynucleotide and has nucleobase-modifying activity. In various embodiments, the base editor comprises a nucleobase-modifying polypeptide (e.g., a deaminase) and a polynucleotide-programmable nucleotide-binding domain that associates with a guide polynucleotide (e.g., a guide RNA). In various embodiments, the agent is a biomolecular complex that includes a protein domain with base-editing activity, i.e., a domain that can modify a base (e.g., A, T, C, G, or U) in a nucleic acid molecule (e.g., DNA). In some embodiments, the polynucleotide-programmable DNA-binding domain is fused or linked to a deaminase domain. In one embodiment, the agent is a fusion protein that includes one or more domains with base-editing activity. In another embodiment, the protein domain with base-editing activity is linked to a guide RNA (e.g., via an RNA-binding motif on the guide RNA and an RNA-binding domain fused to the deaminase). In some embodiments, the domain with base-editing activity is capable of deaminating a base in a nucleic acid molecule. In some embodiments, the base editor is capable of deaminating one or more bases in a DNA molecule. In some embodiments, the base editor is capable of deaminating cytosine (C) or adenosine (A) in DNA. In some embodiments, the base editor is capable of deaminating cytosine (C) and adenosine (A) in DNA. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenosine base editor (ABE). In some embodiments, the base editor is an adenosine base editor (ABE) and a cytidine base editor (CBE). In some embodiments, the base editor is a nuclease-inactive Cas9 fused to adenosine deaminase (dCas9). In some embodiments, the Cas9 is a circularly permuted Cas9 (e.g., spCas9 or saCas9).Circularly permuted Cas9s are known in the art and are described, for example, in Oakes et al., Cell 176, 254-267, 2019. In some embodiments, the base editor is fused to a base excision repair inhibitor, such as a UGI domain or a dISN domain. In some embodiments, the fusion protein comprises a Cas9 nickase fused to a deaminase and a base excision repair inhibitor, such as a UGI or dISN domain. In other embodiments, the base editor is an abasic base editor.
[0055] In some embodiments, the adenosine deaminase is evolved from TadA. In some embodiments, the polynucleotide programmable DNA binding domain is a CRISPR-associated (e.g., Cas or Cpfl) enzyme. In some embodiments, the base editor is a catalytically inactive Cas9 (dCas9) fused to a deaminase domain. In some embodiments, the base editor is a Cas9 nickase (nCas9) fused to a deaminase domain. In some embodiments, the base editor is fused to an inhibitor of base excision repair (BER). In some embodiments, the inhibitor of base excision repair is a uracil-DNA glycosylase inhibitor (UGI). In some embodiments, the inhibitor of base excision repair is an inhibitor of inosine base excision repair. Details of base editors are described in International PCT Application Nos. PCT / 2017 / 045381 (WO2018 / 027078) and PCT / US2016 / 058344 (WO2017 / 070632), both of which are incorporated by reference herein in their entireties.See also Komor, AC, et al., "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage," Nature 533, 420-424 (2016); Gaudelli, NM, et al., "Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage," Nature 551, 464-471 (2017); Komor, AC, et al., "Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity," Science Advances 3:eaao4774 (2017), and Rees, HA, et al., "Base editing: precision chemistry on the genome and transcriptome of living cells," Nat Rev Genet. 2018 Dec;19(12):770-788, the entire contents of which are incorporated herein by reference. See also doi: 10.1038 / s41576-018-0059-1.
[0056] In some embodiments, a base editor (e.g., ABE8) is generated by cloning an adenosine deaminase variant (e.g., TadA*8) into a scaffold comprising a circularly permuted Cas9 (e.g., spCAS9) and a bipartite nucleotidic sequence. Circularly permuted Cas9s are known in the art and are described, for example, in Oakes et al., Cell 176, 254-267, 2019. Exemplary circularly permuted sequences are set forth below, with bolded sequences indicating Cas9-derived sequences, italicized sequences indicating linker sequences, and underlined sequences indicating bipartite nucleotidic sequences.
[0057] CP5 (MSP "NGC = Pam variant containing mutation; regular Cas9 prefers NGG"; PID = protein interaction domain and "D10A" nickase): JPEG0007781742000001.jpg117169
[0058] In some embodiments, ABE8 is selected from a base editor from Table 7 below. In some embodiments, ABE8 contains an adenosine deaminase that is evolved from TadA. In some embodiments, the adenosine deaminase variant of ABE8 is a TadA*8 variant as set forth in Table 7 below. In some embodiments, the adenosine deaminase variant is TadA*7.10 that includes one or more changes selected from the group consisting of Y147T, Y147R, Q154S, Y123H, V82S, T166R, Q154R. In various embodiments, ABE8 comprises TadA*7.10 with alterations selected from the group consisting of: Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y147T+Q154R; Y147T+Q154S; V82S+Q154S; and Y123H+Y147R+Q154R+I76Y. In some embodiments, ABE8 is a monomeric construct.
[0059] In some embodiments, the ABE8 is a heterodimeric construct. In some embodiments, the ABE8 base editor comprises the following sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD
[0060] For example, the adenine base editor ABE used in the base editing compositions, systems, and methods described herein has the nucleic acid sequence (8877 base pairs) described below (Addgene, Watertown, MA.; Gaudelli NM, et al., Nature. 2017 Nov 23;551(7681):464-471. doi: 10.1038 / nature24644; Koblan LW, et al., Nat Biotechnol. 2018 Oct;36(9):843-846. doi: 10.1038 / nbt.4172). Polynucleotide sequences having at least 95% or more identity to the ABE nucleic acid sequence are also encompassed. ATATGCCAAGTACGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTATGCCCAGTACAT GACCTTATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATGGTGATGCGG TTTTGGCAGTACATCAATGGGCGTGGATAGCGGTTTGACTCACGGGATTTCCAAGTCTCCACCCCATTG ACGTCAATGGGAGTTTGTTTTGGCACCAAAATCAACGGGACTTTCCAAAATGTCGTAACAACTCCGCCCC ATTGACGCAAATGGGCGGTAGGCGTGTACGGTGGGAGGTCTATATAAGCAGAGCTGGTTTAGTGAACCGT CAGATCCGCTAGAGATCCGCGGCCGCTAATACGACTCACTATAGGGAGAGCCGCCACCATGAAACGGACA GCCGACGGAAGCGAGTTCGAGTCACCAAAGAAGAAGCGGAAAGTCTCTGAAGTCGAGTTTAGCCACGAGT ATTGGATGAGGCACGCACTGACCCTGGCAAAGCGAGCATGGGATGAAAGAGAAGTCCCCGTGGGCGCCGT GCTGGTGCACAACAATAGAGTGATCGGAGAGGGATGGAACAGGCCAATCGGCCGCCACGACCCTACCGCA CACGCAGAGATCATGGCACTGAGGCAGGGAGGCCTGGTCATGCAGAATTACCGCCTGATCGATGCCACCC TGTATGTGACACTGGAGCCATGCGTGATGTGCGCAGGAGCAATGATCCACAGCAGGATCGGAAGAGTGGT GTTCGGAGCACGGGACGCCAAGACCGGCGCAGCAGGCTCCCTGATGGATGTGCTGCACCACCCCGGCATG AACCACCGGGTGGAGATCACAGAGGGAATCCTGGCAGACGAGTGCGCCGCCCTGCTGAGCGATTTCTTTA GAATGCGGAGACAGGAGATCAAGGCCCAGAAGAAGGCACAGAGCTCCACCGACTCTGGAGGATCTAGCGG AGGATCCTCTGGAAGCGAGACACCAGGCACAAGCGAGTCCGCCACACCAGAGAGCTCCGGCGGCTCCTCC GGAGGATCCTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGCCAAGAGGG CACGCGATGAGAGGGAGGTGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGCGAGGGCTG GAACAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATGGCCCTGAGACAGGGCGGCCTG GTCATGCAGAACTACAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCG GCGCCATGATCCACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACGCAAAAACCGGCGCCGCAGG CTCCCTGATGGACGTGCTGCACTACCCCGGCATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCA GATGAATGTGCCGCCCTGCTGTGCTATTTCTTTCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAAGG CCCAGAGCTCCACCGACTCCGGAGGATCTAGCGGAGGCTCCTCTGGCTCTGAGACACCTGGCACAAGCGA GAGCGCAACACCTGAAAGCAGCGGGGGCAGCAGCGGGGGGTCAGACAAGAAGTACAGCATCGGCCTGGCC ATCGGCACCAACTCTGTGGGCTGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGG TGCTGGGCAACACCGACCGGCACAGCATCAAGAAGAACCTGATCGGAGCCCTGCTGTTCGACAGCGGCGA AACAGCCGAGGCCACCCGGCTGAAGAGAACCGCCAGAAGAAGATACACCAGACGGAAGAACCGGATCTGC TATCTGCAAGAGATCTTCAGCAACGAGATGGCCAAGGTGGACGACAGCTTCTTCCACAGACTGGAAGAGT CCTTCCTGGTGGAAGAGGATAAGAAGCACGAGCGGCACCCCATCTTCGGCAACATCGTGGACGAGGTGGC CTACCACGAGAAGTACCCCACCATCTACCACCTGAGAAAGAAACTGGTGGACAGCACCGACAAGGCCGAC CTGCGGCTGATCTATCTGGCCCTGGCCCACATGATCAAGTTCCGGGGCCACTTCCTGATCGAGGGCGACC TGAACCCCGACAACAGCGACGTGGACAAGCTGTTCATCCAGCTGGTGCAGACCTACAACCAGCTGTTCGA GGAAAACCCCATCAACGCCAGCGGCGTGGACGCCAAGGCCATCCTGTCTGCCAGACTGAGCAAGAGCAGA CGGCTGGAAAATCTGATCGCCCAGCTGCCCGGCGGAGAAAGAATGGCCTGTTCGGAAACCTGATTGCCC TGAGCCTGGGCCTGACCCCAACTTCAAGAGCAACTTCGACCTGGCCGAGGATGCCAAACTGCAGCTGAG CAAGGACACCTACGACGACGACGACCTGGACAACCTGCTGGCCCAGATCGGCGACCAGTACGGCCGACCTGTTTT CTGGCCGCCAAGAACCTGCTCGACGCCATCCTGCTGAGCGACATCCTGAGAGTGAACACCGAGATCACCA AGGCCCCCCTGAGCGCCTCTATGATCAAGAGATACGACGAGCACCACCAGGACCTGACCCTGCTGAAAGC TCTCGTGCGGCAGCAGCTGCCTGAGAAGTACAAAGAGATTTTCTTCGACCAGAGCAAGAACGGCTACGCC GGCTACATTGACGGCGGAGCCAGCCAGGAAGAGTTCTACAAGTTCATCAAGCCCATCCTGGAAAAGATGG ACGGCACCGAGGAACTGCTCGTGAAGCTGAACAGAGGACCTGCTGCGGAAGCAGCGGACCTTCGACAA CGGCAGCATCCCCCACCAGATCCACCTGGGAGAGCTGCACGCCATTCTGCGGCGGCAGGAAGATTTTTAC CCATTCCTGAAGGACAACCGGGAAAGATCGAGAAGATCCTGACCTTCCGCATCCCCTACTACGTGGGCC CTCTGGCCAGGGGAAACAGCAGATTCGCCTGGATCCAGAAAAGGCGAGGAAACCATCACCCCCTGGAA CTTCGAGGAAGTGGTGGACAAGGGCGCTTCCGCCCAGAGCTTCATCGAGCGGATGACCAACTTCGATAAG AACCTGCCCAACGAGAAGGTGCTGCCCAAGCACAGCCTGCTGTACGAGTACTTCACCGTGTATAACGAGC TGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAAGCCCGCCTTCCTGAGCGGCGAGCAGAAAAAGGC CATCGTGGACCTGCTGTTCAAGACCAACCGGAAAGTGACCGTGAAGCAGCTGAAAGAGGACTACTTCAAG AAAATCGAGTGCTTCGACTCCGTGGAAATCTCCGGCGTGGAAGATCGGTTCAACGCCTCCCTGGGCACAT ACCACGATCTGCTGAAAATTATCAAGGACAAGGACTTCCTGGACAATGAGGAAAACGAGGACATTCTGGA AGATATCGTGCTGACCCTGACACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTGAAAACCTATGCC CACCTGTTCGACGACAAAGTGATGAAGCAGCTGAAGCGGCGGAGATACACCGGCTGGGGCAGGCTGAGCC GGAAGCTGATCAACGGCATCCGGGACAAGCAGTCCGGCAAGACAATCCTGGATTTCCTGAAGTCCGACGG CTTCGCCAACAGAAACTTCATGCAGCTGATCCACGACGACAGCCTGACCTTTAAAGAGGACATCCAGAAA GCCCAGGTGTCCGGCCAGGGCGATAGCCTGCACGAGCACATTGCCAATCTGGCCGGCAGCCCCGCCATTA AGAAGGGCATCCTGCAGACAGTGAAGGTGGTGGACGAGCTCGTGAAAGTGATGGGCCGGCACAAGCCCGA GAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCACCCAGAAGGGACAGAAGAACAGCCGCGAGAGA ATGAAGCGGATCGAAGAGGGCATCAAAGAGCTGGGCAGCCAGATCCTGAAAGAACACCCCGTGGAAAACA CCCAGCTGCAGAACGAGAAGCTGTACCTGTACTACCTGCAGAATGGGCGGGATATGTACGTGGACCAGGA ACTGGACATCAACCGGCTGTCCGACTACGATGTGGACCATATCGTGCCTCAGAGCTTTCTGAAGGACGAC TCCATCGACAACAAGGTGCTGACCAGAAGCGACAAGAACCGGGGCAAGAGCGACAACGTGCCCTCCGAAG AGGTCGTGAAGAAGATGAAGAACTACTGGCGGCAGCTGCTGAACGCCAAGCTGATTACCCAGAGAAAGTT CGACAATCTGACCAAGGCCGAGAGAGGCGGCCTGAGCGAACTGGATAAGGCCGGCTTCATCAAGAGACAG CTGGTGGAAACCCGGCAGATCACAAAGCACGTGGCACAGATCCTGGACTCCCGGATGAACACTAAGTACG ACGAGAATGACAAGCTGATCCGGGAAGTGAAAGTGATCACCCTGAAGTCCAAGCTGGTGTCCGATTTCCG GAAGGATTTCCAGTTTTACAAAGTGCGCGAGATCAACAACTACCACCACGCCCACGACGCCTACCTGAAC GCCGTCGTGGGAACCGCCCTGATCAAAAAGTACCCTAAGCTGGAAAGCGAGTTCGTGTACGGCGACTACA AGGTGTACGACGTGCGGAAGATGATCGCCAAGAGCGAGCAGGAAATCGGCAAGGCTACCGCCAAGTACTT CTTCTACAGCAACATCATGAACTTTTTCAAGACCGAGATTACCCTGGCCAACGGCGAGATCCGGAAGCGG CCTCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGTGGGATAAGGGCCGGGATTTTGCCACCGTGC GGAAAGTGCTGAGCATGCCCCAAGTGAATATCGTGAAAAAGACCGAGGTGCAGACAGGCGGCTTCAGCAA AGAGTCTATCCTGCCCAAGAGGAACAGCGATAAGCTGATCGCCAGAAAGAAGGACTGGGACCCTAAGAAG TACGGCGGCTTCGACAGCCCCACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAAGGGCAAGT CCAAGAAACTGAAGAGTGTGAAAGAGCTGCTGGGGATCACCATCATGGAAAGAAGCAGCTTCGAGAAGAA TCCCATCGACTTTCTGGAAGCCAAGGGCTACAAAGAAGTGAAAAAGGACCTGATCATCAAGCTGCCTAAG TACTCCCTGTTCGAGCTGGAAAACGGCCGGAAGAGAATGCTGGCCTCTGCCGGCGAACTGCAGAAGGGAA ACGAACTGGCCCTGCCCTCCAAATATGTGAACTTCCTGTACCTGGCCAGCCACTATGAGAAGCTGAAGGG CTCCCCCGAGGATAATGAGCAGAAACAGCTGTTTGTGGAACAGCACAAGCACTACCTGGACGAGATCATC GAGCAGATCAGCGAGTTCTCCAAGAGAGTGATCCTGGCCGACGCTAATCTGGACAAAGTGCTGTCCGCCT ACAACAAGCACCGGGATAAGCCCATCAGAGAGCAGGCCGAGAATATCATCCACCTGTTTACCCTGACCAA TCTGGGAGCCCCTGCCGCCTTCAAGTACTTTGACACCACCATCGACCGGAAGAGGTACACCAGCACCAAA GAGGTGCTGGACGCCACCCTGATCCACCAGAGCATCACCGGCCTGTACGAGACACGGATCGACCTGTCTC AGCTGGGAGGTGACTCTGGCGGCTCAAAAAGAACCGCCGACGGCAGCGAATTCGAGCCCAAGAAGAAGAG GAAAGTCTAACCGGTCATCATCACCATCACCATTGAGTTTAAACCCGCTGATCAGCCTCGACTGTGCCTT CTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCAC TGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGT GGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCT CTATGGCTTCTGAGGCGGAAAGAACCAGCTGGGGCTCGATACCGTCGACCTCTAGCTAGAGCTTGGCGTA ATCATGGTCATAGCTGTTTCCTGTGTGAAATTGTTATCCGCTCACAATTCCACACAACATACGAGCCGGA AGCATAAAGTGTAAAGCCTAGGGTGCCTAATGAGTGAGCTAACTCACATTAATTGCGTTGCGCTCACTGC CCGCTTTCCAGTCGGGAAACCTGTCGTGCCAGCTGCATTAATGAATCGGCCAACGCGCGGGGAGAGGCGG TTTGCGTATTGGGCGCTCTTCCGCTTCCTCGCTCACTGACTCGCTGCGCTCGGTCGTTCGGCTGCGGCGA GCGGTATCAGCTCACTCAAAGGCGGTAATACGGTTATCCACAGAATCAGGGGATAACGCAGGAAAGAACA TGTGAGCAAAAGGCCAGCAAAAGGCCAGGAACCGTAAAAAGGCCGCGTTGCTGGCGTTTTTCCATAGGCT CCGCCCCCCTGACGAGCATCACAAAAATCGACGCTCAAGTCAGAGGTGGCGAAACCCGACAGGACTATAA AGATACCAGGCGTTTCCCCCTGGAAGCTCCCTCGTGCGCTCTCCTGTTCCGACCCTGCCGCTTACCGGAT ACCTGTCCGCCTTTCTCCCTTCGGGAAGCGTGGCGCTTTCTCATAGCTCACGCTGTAGGTATCTCAGTTC GGTGTAGGTCGTTCGCTCCAAGCTGGGCTGTGTGCACGAACCCCCCGTTCAGCCCGACCGCTGCGCCTTA TCCGGTAACTATCGTCTTGAGTCCAACCCGGTAAGACACGACTTATCGCCACTGGCAGCAGCCACTGGTA ACAGGATTAGCAGAGCGAGGTATGTAGGCGGTGCTACAGAGTTCTTGAAGTGGTGGCCTAACTACGGCTA CACTAGAAGAACAGTATTTGGTATCTGCGCTCTGCTGAAGCCAGTTACCTTCGGAAAAAGAGTTGGTAGC TCTTGATCCGGCAAACAAACCACCGCTGGTAGCGGTGGTTTTTTTGTTTGCAAGCAGCAGATTACGCGCA GAAAAAAAGGATCTCAAGAAGATCCTTTGATCTTTTCTACGGGGTCTGACACTCAGTGGAACGAAAACTC ACGTTAAGGGATTTTGGTCATGAGATTATCAAAAAGGATCTTCACCTAGATCCTTTTAAATTAAAAATGA AGTTTTAAATCAATCTAAAGTATATATGAGTAAACTTGGTCTGACAGTTACCAATGCTTAATCAGTGAGG CACCTATCTCAGCGATCTGTCTATTTCGTTCATCCATAGTTGCCTGACTCCCCGTCGTGTAGATAACTAC GATACGGGAGGGCTTACCATCTGGCCCCAGTGCTGCAATGATACCGCGAGACCCACGCTCACCGGCTCCA GATTTATCAGCAATAAACCAGCCAGCCGGAAGGGCCGAGCGCAGAAGTGGTCCTGCAACTTTATCCGCCT CCATCCAGTCTATTAATTGTTGCCGGGAAGCTAGAGTAAGTAGTTCGCCAGTTAATAGTTTGCGCAACGT TGTTGCCATTGCTACAGGCATCGTGGTGTCACGCTCGTCGTTTGGTATGGCTTCATTCAGCTCCGGTTCC CAACGATCAAGGCGAGTTACATGATCCCCCATGTTGTGCAAAAAAGCGGTTAGCTCCTTCGGTCCTCCGA TCGTTGTCAGAAGTAAGTTGGCCGCAGTGTTATCACTCATGGTTATGGCAGCACTGCATAATTCTCTTAC TGTCATGCCATCCGTAAGATGCTTTTCTGTGACTGGTGAGTACTCAACCAAGTCATTCTGAGAATAGTGT ATGCGGCGACCGAGTTGCTCTTGCCCGGCGTCAATACGGGATAATACCGCGCCACATAGCAGAACTTTAA AAGTGCTCATCATTGGAAAACGTTCTTCGGGGCGAAAACTCTCAAGGATCTTACCGCTGTTGAGATCCAG TTCGATGTAACCCACTCGTGCACCCAACTGATCTTCAGCATCTTTTACTTTCACCAGCGTTTCTGGGTGA GCAAAAACAGGAAGGCAAAATGCCGCAAAAAAGGGAATAAGGGCGACACGGAAATGTTGAATACTCATAC TCTTCCTTTTTCAATATTATTGAAGCATTTATCAGGGTTATTGTCTCATGAGCGGATACATATTTGAATG TATTTAGAAAAATAAACAAATAGGGGTTCCGCGCACATTTCCCCGAAAAGTGCCACCTGACGTCGACGGA TCGGGAGATCGATCTCCCGATCCCCTAGGGTCGACTCTCAGTACAATCTGCTCTGATGCCGCATAGTTAA GCCAGTATCTGCTCCCTGCTTGTGTGTTGGAGGTCGCTGAGTAGTGCGCGAGCAAAATTTAAGCTACAAC AAGGCAAGGCTTGACCGACAATTGCATGAAGAATCTGCTTAGGGTTAGGCGTTTTGCGCTGCTTCGCGAT GTACGGGCCAGATATACGCGTTGACATTGATTATTGACTAGTTATTAATAGTAATCAATTACGGGGTCAT TAGTTCATAGCCCATATATGGAGTTCCGCGTTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCC CAACGACCCCCGCCCATTGACGTCAATAATGACGTATGTTCCCATAGTAACGCCAATAGGGACTTTCCAT TGACGTCAATGGGTGGAGTATTTACGGTAAACTGCCCACTTGGCAGTACATCAAGTGTATC
[0061] By way of example, a cytidine base editor (CBE) used in the base editing compositions, systems, and methods described herein has the following nucleic acid sequence (8877 base pairs), as described below (Addgene, Watertown, MA.; Komor AC, et al., 2017, Sci Adv., 30;3(8):eaao4774. doi: 10.1126 / sciadv.aao4774). Polynucleotide sequences having at least 95% or greater identity to the BE4 nucleic acid sequence are also encompassed. 1 ATATGCCAAG TACGCCCCCT ATTGACGTCA ATGACGGTAA ATGGCCCGCC TGGCATTATG 61 CCCAGTACAT GACCTTATGG GACTTTCCTA CTTGGCAGTA CATCTACGTA TTAGTCATCG 121 CTATTACCAT GGTGATGCGG TTTTGGCAGT ACATCAATGG GCGTGGATAG CGGTTTGACT 181 CACGGGGATT TCCAAGTCTC CACCCCATTG ACGTCAATGG GAGTTTGTTT TGGCACCAAA 241 ATCAACGGGA CTTTCCAAAA TGTCGTAACA ACTCCGCCCC ATTGACGCAA ATGGGCGGTA 301 GGCGTGTACG GTGGGAGGTC TATATAAGCA GAGCTGGTTT AGTGAACCGT CAGATCCGCT 361 AGAGATCCGC GGCCGCTAAT ACGACTCACT ATAGGGAGAG CCGCCACCAT GAGCTCAGAG 421 ACTGGCCCAG TGGCTGTGGA CCCCACATTG AGACGGCGGA TCGAGCCCCA TGAGTTTGAG 481 GTATTCTTCG ATCCGAGAGA GCTCCGCAAG GAGACCTGCC TGCTTTACGA AATTAATTGG 541 GGGGGCCGGC ACTCCATTTG GCGACATACA TCACAGAACA CTAACAAGCA CGTCGAAGTC 601 AACTTCATCG AGAAGTTCAC GACAGAAAGA TATTTCTGTC CGAACACAAG GTGCAGCATT 661 ACCTGGTTTC TCAGCTGGAG CCCATGCGGC GAATGTAGTA GGGCCATCAC TGAATTCCTG 721 TCAAGGTATC CCCACGTCAC TCTGTTTATT TACATCGCAA GGCTGTACCA CCACGCTGAC 781 CCCCGCAATC GACAAGGCCT GCGGGATTTG ATCTCTTCAG GTGTGACTAT CCAAATTATG 841 ACTGAGCAGG AGTCAGGATA CTGCTGGAGA AACTTTGTGA ATTATAGCCC GAGTAATGAA 901 GCCCACTGGC CTAGGTATCC CCATCTGTGG GTACGACTGT ACGTTCTTGA ACTGTACTGC 961 ATCATACTGG GCCTGCCTCC TTGTCTCAAC ATTCTGAGAA GGAAGCAGCC ACAGCTGACA 1021 TTCTTTACCA TCGCTCTTCA GTCTTGTCAT TACCAGCGAC TGCCCCCACA CATTCTCTGG 1081 GCCACCGGGT TGAAATCTGG TGGTTCTTCT GGTGGTTCTA GCGGCAGCGA GACTCCCGGG 1141 ACCTCAGAGT CCGCCACACC CGAAAGTTCT GGTGGTTCTT CTGGTGGTTC TGATAAAAAG 1201 TATTCTATTG GTTTAGCCAT CGGCACTAAT TCCGTTGGAT GGGCTGTCAT AACCGATGAA 1261 TACAAAGTAC CTTCAAAGAA ATTTAAGGTG TTGGGGAACA CAGACCGTCA TTCGATTAAA 1321 AAGAATCTTA TCGGTGCCCT CCTATTCGAT AGTGGCGAAA CGGCAGAGGC GACTCGCCTG 1381 AAACGAACCG CTCGGAGAAG GTATACACGT CGCAAGAACC GAATATGTTA CTTACAAGAA 1441 ATTTTTAGCA ATGAGATGGC CAAAGTTGAC GATTCTTTCT TTCACCGTTT GGAAGAGTCC 1501 TTCCTTGTCG AAGAGGACAA GAAACATGAA CGGCACCCCA TCTTTGGAAA CATAGTAGAT 1561 GAGGTGGCAT ATCATGAAAA GTACCCAACG ATTTATCACC TCAGAAAAAA GCTAGTTGAC 1621 TCAACTGATA AAGCGGACCT GAGGTTAATC TACTTGGCTC TTGCCCATAT GATAAAGTTC 1681 CGTGGGCACT TTCTCATTGA GGGTGATCTA AATCCGGACA ACTCGGATGT CGACAAACTG 1741 TTCATCCAGT TAGTACAAAC CTATAATCAG TTGTTTGAAG AGAACCCTAT AAATGCAAGT 1801 GGCGTGGATG CGAAGGCTAT TCTTAGCGCC CGCCTCTCTA AATCCCGACG GCTAGAAAAAC 1861 CTGATCGCAC AATTACCCGG AGAGAAGAAA AATGGGTTGT TCGGTAACCT TATAGCGCTC 1921 TCACTAGGCC TGACACCAAA TTTTAAGTCG AACTTCGACT TAGCTGAAGA TGCCAAATTG 1981 CAGCTTAGTA AGGACACGTA CGATGACGAT CTCGACAATC TACTGGCACA AATTGGAGAT 2041 CAGTATGCGG ACTTATTTTT GGCTGCCAAA AACCTTAGCG ATGCAATCCT CCTATCTGAC 2101 ATACTGAGAG TTAATACTGA GATTACCAAG GCGCCGTTAT CCGCTTCAAT GATCAAAAGG 2161 TACGATGAAC ATCACCAAGA CTTGACACTT CTCAAGGCCC TAGTCCGTCA GCAACTGCCT 2221 GAGAAATATA AGGAAATATT CTTTGATCAG TCGAAAAACG GGTACGCAGG TTATATTGAC 2281 GGCGGAGCGA GTCAAGAGGA ATTCTACAAG TTTATCAAAC CCATATTAGA GAAGATGGAT 2341 GGGACGGAAG AGTTGCTTGT AAAACTCAAT CGGAAGATC TACTGCGAAA GCAGCGGACT 2401 TTCGACAACG GTAGCATTCC ACATCAAATC CACTTAGGCG AATTGCATGC TATACTTAGA 2461 AGGCAGGAGG ATTTTTATCC GTTCCTCAAA GACAATCGTG AAAAGATTGA GAAAATCCTA 2521 ACCTTTCGCA TACCTTACTA TGTGGGACCC CTGGCCCGAG GGAACTCTCG GTTCGCATGG 2581 REPEAT AGTCCREACT AACGATTACT CCATGGAATTT TTREPEAT TGTCREADAA 2641 GGTGCGTCAG CTCAATCGTT CATCGAGAGG ATGACCAACT TTGACAAA TTTACCGAAC 2701 GAAAAGTAT TGCCTAXASSOCIATTACTT CHAPTER TGACTACTACATCHACTC 2761 ACGAAAGTTA AGTATGTCAC TGAGGGCATG CGTAAACCCG CCTTTCTAAG CGGAGAACAG 2821 AAGAAAGCAA TAGTAGATCT GTTATTCAAG ACCAACCGCA AAGTGACAGT TAAGCAATTG 2881 AAAGAGGACT ACTTTAAGAA AATTGAATGC TTCGATTCTG TCGAGATCTC CGGGGTAGAA 2941 GATCGATTTA ATGCGTCACT TGGTACGTAT CATGACCCTCC WINDOW INTERFACE 3001 GACTTCCTGG ATTACK REPEAT ATCTCTGG ATTACKGTT GACTCTTACC 3061 CTCTTTGAAG ATCGGGAAAT GATTGAGGAA AGACTAAAAAA CATACGCTCA CCTGTTCGAC 3121 GATAAGGTTA TGAAACAGTT AAAGAGGCGT CGCTATACGG GCTGGGGACG ATTGTCGCGG 3181 AAACTTATCA ACGGGATAAG AGACAAGCAA AGTGGTAAAA CTATTCTCGA TTTTCTAAAG 3241 AGCGACGGCT TCGCCAATAG GAACTTTATG CAGCTGATCC ATGATGACTC TTTAACCTTC 3301 AAAGAGGATA TACAAAAGGC ACAGGTTTCC GGACAAGGGG ACTCATTGCA CGAACATATT 3361 GCGAATCTTG CTGGTTCGCC AGCCATCAAA AAGGGCATAC TCCAGACAGT CAAAGTAGTG 3421 GATGAGCTAG TTAAGGTCAT GGGACGTCAC AAACCGGAAA ACATTGTAAT CGAGATGGCA 3481 CGCGAAAATC AAACGACTCA GAAGGGGCAA AAAAAACAGTC GAGAGCGGAT GAAGAGAATA 3541 GAAGAGGGTA TTAAAGAACT GGGCAGCCAG ATCTTAAAGG AGCATCCTGT GGAAAATACC 3601 CAATTGCAGA ACGAGAAACT TTACCTCTAT TACCTACAAA ATGGAAGGGA CATGTATGTT 3661 GATCAGGAAC TGGACATAAA CCGTTTATCT GATTACGACG TCGATCACAT TGTACCCCAA 3721 TCCTTTTTGA AGGACGATTC AATCGACAAT AAAGTGCTTA CACGCTCGGA TAAGAACCGA 3781 GGGAAAAGTG ACAATGTTCC AAGCGAGGAA GTCGTAAAGA AAATGAGAA CTATTGGCGG 3841 CAGCTCCTAA ATGCGAAACT GATAACGCAA AGAAAGTTCG ATAACTTAAC TAAAGCTGAG 3901 AGGGGTGGCT TGTCTGAACT TGACAAGGCC GGATTTATTA AACGTCAGCT CGTGGAAACC 3961 CGCCAAATCA CAAAGCATGT TGCACAGATA CTAGATTCCC GAATGAATAC GAAATACGAC 4021 GAGAACGATA AGCTGATTCG GGAAGTCAAA GTAATCACTT TAAAGTCAAA ATTGGTGTCG 4081 GACTTCAGAA AGGATTTTCA ATTCTATAAA GTTAGGGAGA TAATAACTA CCACCATGCG 4141 CACGACGCTT ATCTTAATGC CGTCGTAGGG ACCGCACTCA TTAAGAAATA CCCGAAGCTA 4201 GAAAGTGAGT TTGTGTATGG TGATTACAAA GTTTATGACG TCCGTAAGAT GATCGCGAAA 4261 AGCGAACAGG AGATAGGCAA GGCTACAGCC AAATACTTCT TTTATTCTAA CATTATGAAT 4321 TTCTTTAAGA CGGAAATCAC TCTGGCAAAC GGAGAGATAC GCAAACGACC TTTAATTGAA 4381 ACCAATGGGG AGACAGGTGA AATCGTATGG GATAAGGGCC GGGACTTCGC GACGGTGAGA 4441 AAAGTTTTGT CCATGCCCCA AGTCAACATA GTAAAGAAAA CTGAGGTGCA GACCGGAGGG 4501 TTTTCAAAGG AATCGATTCT TCCAAAAAGG AATAGTGATA AGCTCATCGC TCGTAAAAAG 4561 GACTGGGACC CGAAAAAGTA CGGTGGCTTC GATAGCCCTA CAGTTGCCTA TTCTGTCCTA 4621 GTAGTGGCAA AAGTTGAGAA GGGAAAATCC AAGAAACTGA AGTCAGTCAA AGAATTATTG 4681 GGGATAACGA TTATGGAGCG CTCGTCTTTT GAAAAGAACC CCATCGACTT CCTTGAGGCG 4741 AAAGGTTACA AGGATAAA AAAGGATCTC ATTACK TACTTACA TAGTCTGTTT 4801 GAGTTAGAAA ATGGCCGAAA ACGGATGTTG GCTAGCGCCG GAGAGCTTCA AAAGGGGAAC 4861 GAACTCGCAC TACCGTCTAA ATACGTGAAT TTCCTGTATT TAGCGTCCCA TTACGAGAAG 4921 TTGAAAGGTT CACCTGAAGA TAACGAACAG AAGCAACTTT TTGTTGAGCA GCACAAACAT 4981 TATCTCGACG AAATCATAGA GCAAATTTCG GAATTCAGTA AGAGAGTCAT CCTAGCTGAT 5041 GCCAATCTGG ACAAAGTATT AAGCGCATAC AACAAGCACA GGGATAAACC CATACGTGAG 5101 CAGGCGGAAA ATATTATCCA TTTGTTTACT CTTACCAACC TCGGCGCTCC AGCCGCATTC 5161 AAGTATTTTG ACACAACGAT AGATCGCAAA CGATACACTT CTACCAAGGA GGTGCTAGAC 5221 GCGACACTGA TTCACCAATC CATCACGGGA TTATATGAAA CTCGGATAGA TTTGTCACAG 5281 CTTGGGGGTG ACTCTGGTGG TTCTGGAGGA TCTGGTGGTT CTACTAATCT GTCAGATATT 5341 ATTGAAAAGG AGACCGGTAA GCAACTGGTT ATCCAGGAAT CCATCCTCAT GCTCCCAGAG 5401 GAGGTGGAAG AAGTCATTGG GAACAAGCCG GAAAGCGATA TACTCGTGCA CACCGCCTAC 5461 GACGAGAGCA CCGACGAGAA TGTCATGCTT CTGACTAGCG ACGCCCCTGA ATACAAGCCT 5521 TGGGCTCTGG TCATACAGGA TAGCAACGGT GAGAACAAGA TTAAGATGCT CTCTGGTGGT 5581 TCTGGAGGAT CTGGTGGTTC TACTAATCTG TCAGATATTA TTGAAAAGGA GACCGGTAAG 5641 CAACTGGTTA TCCAGGAATC CATCCTCATG CTCCCAGAGG AGGTGGAAGA AGTCATTGGG 5701 AACAAGCCGG AAAGCGATAT ACTCGTGCAC ACCGCCTACG ACGAGAGCAC CGACGAGAAT 5761 GTCATGCTTC TGACTAGGCGA CGCCCCTGAA TACAAGCCTT GGGCTCTGGT CATACAGGAT 5821 AGCAACGGTG AGAACAAGAT TAAGATGCTC TCTGGTGGTT CTCCCAAGAA GAAGAGGAAA 5881 GTCTAACCGG TCATCATCAC CATCACCATT GAGTTTAAAC CCGCTGATCA GCCTCGACTG 5941 TGCCTTCTAG TTGCCAGCCA TCTGTTGTTT GCCCCTCCCC CGTGCCTTCC TTGACCCTGG 6001 AAGGTGCCAC TCCCACTGTC CTTTCCTAAT AAAATGAGGA AATTGCATCG CATTGTCTGA 6061 GTAGGTGTCA TTCTATTCTG GGGGGTGGGG TGGGGCAGGA CAGCAAGGGG GAGGATTGGG 6121 AAGACAATAG CAGGCATGCT GGGGATGCGG TGGGCTCTAT GGCTTCTGAG GCGGAAAGAA 6181 CCAGCTGGGG CTCGATACCG TCGACCTCTA GCTAGAGCTT GGCGTAATCA TGGTCATAGC 6241 TGTTTCCTGT GTGAAATTGT TATCCGCTCA CAATTCCACA CAACATACGA GCCGGAAGCA 6301 TAAAGTGTAA AGCCTAGGGT GCCTAATGAG TGAGCTAACT CACATTAATT GCGTTGCGCT 6361 CACTGCCCGC TTTCCAGTCG GGAAACCTGT CGTGCCAGCT GCATTAATGA ATCGGCCAAC 6421 GCGCGGGGAG AGGCGGTTTG CGTATTGGGC GCTCTTCCGC TTCCTCGCTC ACTGACTCGC 6481 TGCGCTCGGT CGTTCGGCTG CGGCGAGCGG TATCAGCTCA CTCAAAGGCG GTAATACGGT 6541 TATCCACAGA ATCAGGGGAT AACGCAGGAA AGAACATGTG AGCAAAAGGC CAGCAAAAGG 6601 CCAGGAACCG TAAAAAGGCC GCGTTGCTGG CGTTTTTCCA TAGGCTCCGC CCCCCTGACG 6661 AGCATCACAA AAATCGACGC TCAAGTCAGA GGTGGCGAAA CCCGACAGGA CTATAAAGAT 6721 ACCAGGCGTT TCCCCCTGGA AGCTCCCTCG TGCGCTCTCC TGTTCCGACC CTGCCGCTTA 6781 CCGGATACCT GTCCGCCTTT CTCCCTTCGG GAAGCGTGGC GCTTTCTCAT AGCTCACGCT 6841 GTAGGTATCT CAGTTCGGTG TAGGTCGTTC GCTCCAAGCT GGGCTGTGTG CACGAACCCC 6901 CCGTTCAGCC CGACCGCTGC GCCTTATCCG GTAACTATCG TCTTGAGTCC AACCCGGTAA 6961 GACACGACTT ATCGCCACTG GCAGCAGCCA CTGGTAACAG GATTAGCAGA GCGAGGTATG 7021 TAGGCGGTGC TACAGAGTTC TTGAAGTGGT GGCCTAACTA CGGCTACACT AGAAGAACAG 7081 TATTTGGTAT CTGCGCTCTG CTGAAGCCAG TTACCTTCGG AAAAAGAGTT GGTAGCTCTT 7141 GATCCGGCAA ACAAACCACC GCTGGTAGCG GTGGTTTTTT TGTTTGCAAG CAGCAGATTA 7201 CGCGCAGAAA AAAAGGATCT CAAGAAGATC CTTTGATCTT TTCTACGGGG TCTGACGCTC 7261 AGTGGAACGA AAACTCACGT TAAGGGATTT TGGTCATGAG ATTATCAAAA AGGATCTTCA 7321 CCTAGATCCT TTTAAATTAA AAATGAAGTT TTAAATCAAT CTAAAGTATA TATGAGTAAA 7381 CTTGGTCTGA CAGTTACCAA TGCTTAATCA GTGAGGCACC TATCTCAGCG ATCTGTCTAT 7441 TTCGTTCATC CATAGTTGCC TGACTCCCCG TCGTGTAGAT AACTACGATA CGGGAGGGCT 7501 TACCATCTGG CCCCAGTGCT GCAATGATAC CGCGAGACCC ACGCTCACCG GCTCCAGATT 7561 TATCAGCAAT AAACCAGCCA GCCGGAAGGG CCGAGCGCAG AAGTGGTCCT GCAACTTTAT 7621 CCGCCTCCAT CCAGTCTATT AATTGTTGCC GGGAAGCTAG AGTAAGTAGT TCGCCAGTTA 7681 ATAGTTTGCG CAACGTTGTT GCCATTGCTA CAGGCATCGT GGTGTCACGC TCGTCGTTTG 7741 GTATGGCTTC ATTCAGCTCC GGTTCCCAAC GATCAAGGCG AGTTACATGA TCCCCCATGT 7801 TGTGCAAAAA AGCGGTTAGC TCCTTCGGTC CTCCGATCGT TGTCAGAAGT AAGTTGGCCG 7861 CAGTGTTATC ACTCATGGTT ATGGCAGCAC TGCATAATTC TCTTACTGTC ATGCCATCCG 7921 TAAGATGCTT TTCTGTGACT GGTGAGTACT CAACCAAGTC ATTCTGAGAA TAGTGTATGC 7981 GGCGACCGAG TTGCTCTTGC CCGGCGTCAA TACGGGATAA TACCGCGCCA CATAGCAGAA 8041 CTTTAAAAGT GCTCATCATT GGAAAACGTT CTTCGGGGCG AAAACTCTCA AGGATCTTAC 8101 CGCTGTTGAG ATCCAGTTCG ATGTAACCCA CTCGTGCACC CAACTGATCT TCAGCATCTT 8161 TTACTTTCAC CAGCGTTTCT GGGTGAGCAA AAACAGGAAG GCAAAATGCC GCAAAAAAGG 8221 GAATAAGGGC GACACGGAAA TGTTGAATAC TCATACTCTT CCTTTTTCAA TATTATTGAA 8281 GCATTTATCA GGGTTATTGT CTCATGAGCG GATACATATT TGAATGTATT TAGAAAAATA 8341 AACAAATAGG GGTTCCGCGC ACATTTCCCC GAAAAGTGCC ACCTGACGTC GACGGATCGG 8401 GAGATCGATC TCCCGATCCC CTAGGGTCGA CTCTCAGTAC AATCTGCTCT GATGCCGCAT 8461 AGTTAAGCCA GTATCTGCTC CCTGCTTGTG TGTTGGAGGT CGCTGAGTAG TGCGCGAGCA 8521 AAATTTAAGC TACAACAAGG CAAGGCTTGA CCGACAATTG CATGAAGAAT CTGCTTAGGG 8581 TTAGGCGTTT TGCGCTGCTT CGCGATGTAC GGGCCAGATA TACGCGTTGA CATTGATTAT 8641 TGACTAGTTA TTAATAGTAA TCAATTACGG GGTCATTAGT TCATAGCCCA TATATGGAGT 8701 TCCGCGTTAC ATAACTTACG GTAAATGGCC CGCCTGGCTG ACCGCCCAAC GACCCCCGCC 8761 CATTGACGTC AATAATGACG TATGTTCCCA TAGTAACGCC AATAGGGACT TTCCATTGAC 8821 GTCAATGGGT GGAGTATTTA CGGTAAACTG CCCACTTGGC AGTACATCAA GTGTATC
[0062] In some embodiments, the cytidine base editor is BE4 having a nucleic acid sequence selected from one of the following: Native BE4 nucleic acid sequence: BE4 Codon Optimized 1 Nucleic Acid Sequence: BE4 Codon Optimized 2 Nucleic Acid Sequence:
[0063] "Base editing activity" refers to acting to chemically modify a base within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is, for example, a cytidine deaminase activity that converts a target C·G to T·A. In another embodiment, the base editing activity is, for example, an adenosine or adenine deaminase activity that converts an A·T to G·C. In another embodiment, the base editing activity is, for example, a cytidine deaminase activity that converts a target C·G to T·A, or, for example, an adenosine or adenine deaminase activity that converts an A·T to G·C.
[0064] The term "base editor system" refers to a system for editing nucleobases of a target nucleotide sequence. In various embodiments, the base editor (BE) system comprises: (1) a polynucleotide-programmable nucleotide-binding domain, a deaminase domain for deaminating nucleobases in the target nucleotide sequence, and a cytidine deaminase domain; and (2) one or more guide polynucleotides (e.g., guide RNAs) linked to the polynucleotide-programmable nucleotide-binding domain. In various embodiments, the base editor (BE) system comprises a nucleobase editor domain selected from adenosine deaminase or cytidine deaminase, and a domain with nucleic acid sequence-specific binding activity. In some embodiments, the base editor system comprises: (1) a base editor (BE) comprising a polynucleotide-programmable DNA-binding domain and a deaminase domain for deaminating one or more nucleobases in the target nucleotide sequence; and (2) one or more guide RNAs linked to the polynucleotide-programmable DNA-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable DNA-binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE) or a cytidine base editor (CBE).
[0065] The term "Cas9" or "Cas9 domain" refers to an RNA-guided nuclease that includes the Cas9 protein or a fragment thereof (e.g., a protein containing an active, inactive, or partially active DNA-cleavage domain of Cas9 and / or the gRNA-binding domain of Cas9). Cas9 nucleases are also sometimes referred to as casnl nucleases or CRISPR (clustered regularly interspaced short palindromic repeats)-associated nucleases. An exemplary Cas9 is Streptococcus pyrogenes Cas9 (spCas9), the amino acid sequence of which is shown below: JPEG0007781742000002.jpg114169 (single underline: HNH domain; double underline: RuvC domain)
[0066] The term "conservative amino acid substitution" or "conservative mutation" refers to the substitution of one amino acid for another amino acid with common properties. A functional method for defining common properties between individual amino acids is to analyze the normalized frequency of amino acid changes between corresponding proteins of homologous organisms (Schulz, GE and Schirmer, RH, Principles of Protein Structure, Springer-Verlag, New York (1979)). According to such analysis, amino acid groups can be defined, in which amino acids within a group preferentially exchange with each other and therefore most closely resemble each other in terms of their effect on the overall protein structure (Schulz, GE and Schirmer, RH, supra). Non-limiting examples of conservative mutations include amino acid substitutions of amino acids, such as arginine for lysine, and vice versa, to maintain a positive charge; aspartic acid for glutamic acid, and vice versa, to maintain a negative charge; threonine for serine, to maintain a free -OH; and asparagine for glutamine, to maintain a free NH2.
[0067] The terms "coding sequence" or "protein-coding sequence," as used interchangeably herein, refer to a segment of a polynucleotide that encodes a protein. This region or sequence is bounded near the 5' end by a start codon and near the 3'3' end by a stop codon. Stop codons useful for the base editors described herein include: Glutamine CAG → TAG stop codon CAA → TAA Arginine CGA → TGA Tryptophan TGG → TGA TGG → TAG TGG→TAA A coding sequence can also be referred to as an open reading frame.
[0068] "Cytidine deaminase" refers to a polypeptide or fragment thereof capable of catalyzing a deamination reaction that converts an amino group to a carbonyl group. In one embodiment, the cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine. PmCDA1 (sea lamprey cytosine deaminase 1, "PmCDA1") from the sea lamprey (Petromyzon marinus), AID (activation-induced cytidine deaminase; AICDA) from mammals or different species of mammals (e.g., humans, pigs, cows, horses, monkeys, etc.), and non-mammals such as alligators, and APOBEC are exemplary cytidine deaminases.
[0069] As used herein, the term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes the hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. In some embodiments, the deaminase or deaminase domain is a cytosine deaminase that catalyzes the hydrolytic deamination of cytosine to uracil. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or adenine (A) to inosine (I). In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases (e.g., engineered adenosine deaminases, evolved adenosine deaminases) provided herein can be derived from any organism, such as bacteria. In some embodiments, the adenosine deaminase is derived from bacteria such as Escherichia coli, Staphylococcus aureus, Salmonella typhimurium, Shewanella putrefaciens, Haemophilus influenzae, or Caulobacter crescentus. In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism such as a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain is not naturally occurring.For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical to a naturally occurring deaminase.
[0070] "Direct" refers to determining the presence, absence, or amount of an analyte being detected. In one embodiment, a sequence alteration in a polynucleotide or polypeptide is detected. In another embodiment, the presence of an indel is detected.
[0071] By "detectable label" is meant a composition that, when attached to a molecule of interest, renders the latter detectable via spectroscopic, photochemical, biochemical, immunochemical, or chemical means. For example, useful labels include radioisotopes, magnetic beads, metallic beads, colloidal particles, fluorescent dyes, electron-dense reagents, enzymes (e.g., as commonly used in enzyme-linked immunosorbent assays (ELISAs)), biotin, digoxigenin, or haptens.
[0072] "Disease" refers to any condition or disorder that damages or interferes with the normal function of a cell, tissue, or organ. In a specific embodiment, a disease amenable to treatment with the compositions of the present invention is associated with abnormal splicing. In one specific embodiment, the disease is Shwachman-Diamond Syndrome (SDS).
[0073] By "disease associated with aberrant splicing" is meant any condition or disorder associated with disrupted transcription due to alterations in the gene sequence that affect splicing, such as alterations within the splice acceptor site or splice donor site.
[0074] An "effective amount" refers to the amount of an agent or active compound, e.g., a base editor as described herein, required to ameliorate disease symptoms compared to an untreated patient or disease-free individual, i.e., a healthy individual, or the amount of an agent or active compound sufficient to elicit a desired biological response. The effective amount of an active compound used in practicing the present invention for therapeutic treatment of a disease will vary depending on the mode of administration and the age, weight, and general health of the subject. Ultimately, the appropriate amount and administration regimen will be determined by the attending physician or veterinarian. Such an amount is referred to as an "effective" amount. In one embodiment, an effective amount is the amount of a base editor of the present invention sufficient to introduce an alteration into a gene of interest in a cell (e.g., a cell in vitro or in vivo). In one embodiment, an effective amount is the amount of a base editor required to achieve a therapeutic effect. Such a therapeutic effect need not be sufficient to alter pathogenic genes in all cells of a subject, tissue, or organ, but may be sufficient to alter pathogenic genes in about 1%, 5%, 10%, 25%, 50%, 75%, or more of the cells present in the subject, tissue, or organ. In one embodiment, the effective amount is sufficient to ameliorate one or more symptoms of the disease.
[0075] In some embodiments, an effective amount of an agent or composition comprising a nucleobase editor comprising an nCas9 domain and a deaminase domain (e.g., adenosine deaminase, cytidine deaminase), which may be in the form of a fusion protein as described herein, or a nucleobase editor comprising an nCas9 domain and a deaminase domain (e.g., adenosine deaminase, cytidine deaminase), refers to an amount sufficient to induce editing of a target site specifically bound and edited by the nucleobase editor described herein. As will be understood by one of skill in the art, the effective amount of an agent, e.g., a fusion protein, can vary depending on various factors, e.g., the desired biological response, e.g., the specific allele, genome, or target site to be edited, the cell or tissue being targeted, and / or the agent being used.
[0076] In some embodiments, an effective amount of an agent, which may be in the form of a fusion protein, e.g., a fusion protein comprising an nCas9 domain and a deaminase domain, may refer to the amount of agent, e.g., fusion protein, sufficient to induce editing of a target site specifically bound and edited by the fusion protein. As will be understood by those skilled in the art, the effective amount of an agent, e.g., a fusion protein, a nuclease, a hybrid protein, a protein dimer, a protein complex (or protein dimer), and a polynucleotide, or a polynucleotide, may vary depending on various factors, e.g., the desired biological response, e.g., the specific allele, genome, or target site to be edited, the cell or tissue being targeted, and / or the agent being used.
[0077] By "fragment" is meant a portion of a polypeptide or nucleic acid molecule, which portion contains at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the entire length of the reference nucleic acid molecule or polypeptide. Fragments can contain 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids.
[0078] "Guide RNA" or "gRNA" refers to a polynucleotide that is specific for a target sequence and can form a complex with a polynucleotide-programmable nucleotide-binding domain protein (e.g., Cas9 or Cpfl). In one embodiment, the guide polynucleotide is a guide RNA (gRNA). A gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule is sometimes referred to as a single guide RNA (sgRNA), although "gRNA" is used interchangeably to refer to a guide RNA that exists as a single molecule or as a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species contains two domains: (1) a domain that shares homology with the target nucleic acid (and, for example, directly binds to the target of the Cas9 complex); and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and has a stem-loop structure. For example, in some embodiments, domain (2) corresponds to or is homologous to the tracrRNA described in Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those containing domain 2) can be found in U.S. Patent Application Publication No. 20160208288, entitled "Switchable Cas9 Nucleases and Uses Thereof," and U.S. Patent No. 9,737,604, entitled "Delivery System For Functional Nucleases," the entire contents of each of which are incorporated herein by reference in their entirety. In some embodiments, the gRNA comprises two or more domains (1) and (2), and may also be referred to as an "extended gRNA." An extended gRNA, as described herein, binds two or more Cas9 proteins and binds a target nucleic acid at two or more distinct regions. The gRNA contains a nucleotide sequence complementary to the target site, which mediates binding of the nuclease / RNA complex to the target site and provides sequence specificity for the nuclease:RNA complex.
[0079] "Hybridization" refers to hydrogen bonding between complementary nucleobases, which may be Watson-Crick, Hoogsteen, or reversed Hoogsteen hydrogen bonds. For example, adenine and thymine are complementary nucleobases that pair through the formation of hydrogen bonds.
[0080] By "increase" is meant a positive change of at least 10%, 25%, 50%, 75%, or 100%.
[0081] The terms "inhibitor of base repair," "base repair inhibitor," "IBR," or grammatical equivalents thereof, refer to a protein capable of inhibiting the activity of a nucleic acid repair enzyme, e.g., a base excision repair enzyme. In some embodiments, the IBR is an inhibitor of inosine base excision repair. Exemplary inhibitors of base repair include inhibitors of APE1, EndoIII, EndoIV, EndoV, EndoVIII, Fpg, hOGGl, hNEILl, T7Endol, T4PDG, UDG, hSMUGl, and hAAG. In some embodiments, the base repair inhibitor is an inhibitor of EndoV or hAAG. In some embodiments, the IBR is an inhibitor of EndoV or hAAG. In some embodiments, the IBR is catalytically inactive EndoV or catalytically inactive hAAG. In some embodiments, the base repair inhibitor is catalytically inactive EndoV or catalytically inactive hAAG. In some embodiments, the base repair inhibitor is a uracil glycosylase inhibitor (UGI). UGI refers to a protein capable of inhibiting uracil DNA glycosylase base excision repair enzymes. In some embodiments, the UGI domain comprises wild-type UGI or a fragment of wild-type UGI. In some embodiments, the UGI proteins described herein include fragments of UGI and proteins homologous to UGI or UGI fragments. In some embodiments, the base repair inhibitor is an inhibitor of inosine base excision repair. In some embodiments, the base repair inhibitor is a "catalytically inactive inosine-specific nuclease" or "inactive inosine-specific nuclease." Without wishing to be bound by theory, catalytically inactive inosine glycosylases (e.g., alkyladenine glycosylases (AAG)) can bind to inosine but cannot create an abasic site or remove the inosine, thereby sterically blocking the newly formed inosine moiety from DNA damage / repair mechanisms. In some embodiments, catalytically inactive inosine-specific nucleases can allow inosine to bind within nucleic acids but do not cleave nucleic acids.Non-limiting exemplary catalytically inactive inosine-specific nucleases include catalytically inactive alkyladenosine glycosylase (AAG nuclease), e.g., from humans, and catalytically inactive endonuclease V (EndoV nuclease), e.g., from E. coli. In some embodiments, the catalytically inactive AAG nuclease comprises an E125Q mutation or a corresponding mutation in another AAG nuclease.
[0082] An "intein" is a fragment of a protein that can excise itself and join the remaining fragment (an extein) via a peptide bond in a process known as protein splicing. Inteins are also called "protein introns." The process of an intein excising itself and joining the remainder of the protein is referred to herein as "protein splicing" or "intein-mediated protein splicing." In some embodiments, the inteins of a precursor protein (the intein-containing protein prior to intein-mediated protein splicing) are derived from two genes. Such inteins are referred to herein as split inteins (e.g., split intein-N and split intein-C). For example, in cyanobacteria, DnaE, the catalytic subunit of DNA polymerase III, is encoded by two separate genes, dnaE-n and dnaE-c. The intein encoded by the dnaE-n gene is sometimes referred to herein as "intein-N." The intein encoded by the dnaE-c gene is sometimes referred to herein as "intein-C."
[0083] Other intein systems may also be used. For example, synthetic inteins based on the dnaE intein, the intein pair Cfa-N (e.g., split intein-N) and Cfa-C (e.g., split intein-C), have been described (e.g., Stevens et al., J Am Chem Soc. 2016 Feb. 24; 138(7):2162-5, incorporated herein by reference). Non-limiting examples of intein pairs that can be used in accordance with the present disclosure include the Cfa DnaE intein, Ssp GyrB intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB intein, and Cne Prp8 intein (as described in U.S. Patent No. 8,394,604, incorporated herein by reference).
[0084] Exemplary nucleotide and amino acid sequences of inteins are shown. DnaE intein-N DNA:TGCCTGTCATACGAACCGAGATACTGACAGTAGAATATGGCCTTCTGCCAATCGGGAAGATTGTGGAGAAACGGATAGAATGCACAGTTTACTCTGTCGATAACAATGGTAACATTTATACTCAGCCAGTTGCCCAGTGGCACGACCGGG GAGAGCAGGAAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAGGGCCACTAAGGACCACAAATTTATGACAGTCGATGGCCAGATGCTGCCTATAGACGAAATCTTTGAGCGAGAGTTGGACCTCATGCGAGTTGACAACCTTCCTAAT DnaE Intein-N Protein: CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDR GEQEVFEYCLEDGSLIRATKDHKFMTVDGQMLPIDEIFERELDLMRVDNL PN DnaE intein-C DNA: ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGA TATTGGAGTCGAAAGAGATCACAACTTTGCTCTGAAGAACGGATTCATAG CTTCTAAT Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN Cfa-N DNA: TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGCCTATTGGAAAGATTGTCGAAGAGAGAATTGAATGCACAGTATATACTGTAGACAAGAATGGTTTCGTTTACACACAGCCCATTGCTCAATGGCACAATCGCGGCGAACAAGAAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACGAGCAACTAAAGATCATAAATTCATGACCACTGACGGGCAGATGTTGCCAATAGATGAGATATTCGAGCGGGGCTTGGATCTCAAACAAGTGGATGGATTGCCA Cfa-N protein: CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQMLPIDEIFERGLDLKQVDGLP Cfa-C DNA: ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAGGAAAGTAAAGATAATATCTCGAAAAAGTCTTGGTACCCAAAATGTCTATGATATTGGAGTGGAGAAAGATCACAACTTCCTTCTCAAGAACGGTCTCGTAGCCAGCAAC Cfa-C protein: MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLVASN
[0085] Intein-N and intein-C may be fused to the N-terminal portion of split Cas9 and the C-terminal portion of split Cas9, respectively, for joining the N-terminal portion of split Cas9 and the C-terminal portion of split Cas9. For example, in some embodiments, intein-N is fused to the C-terminus of the N-terminal portion of split Cas9, i.e., forming the structure N-[N-terminal portion of split Cas9]-[intein-N]-C. In some embodiments, intein-C is fused to the N-terminus of the C-terminal portion of split Cas9, i.e., forming the structure N-[intein-C]-[C-terminal portion of split Cas9]-C. The mechanism of intein-mediated protein splicing for joining intein-fused proteins (e.g., split Cas9) is known in the art and is described, for example, in Shah et al., Chem Sci. 2014; 5(1):446-461, which is incorporated herein by reference. Methods for designing and using inteins are known in the art and are described, for example, in WO2014004336, WO2017132580, U.S. Patent Application Publication No. 20150344549, and U.S. Patent Application Publication No. 20180127780, all of which are incorporated herein by reference in their entireties.
[0086] The terms "isolated," "purified," or "biologically pure" refer to varying degrees of freedom from components normally associated with a material when found in its native state. "Isolated" refers to a degree of separation from the original source or environment. "Purified" refers to a degree of separation greater than isolation. A "purified" or "biologically pure" protein is free from other substances to the extent that any impurities do not materially affect the biological properties of the protein or produce other adverse effects. In other words, the nucleic acids or peptides of the invention are purified when they are substantially free from cellular material, viral material, or culture medium when produced by recombinant DNA techniques, or chemical precursors or other chemicals when chemically synthesized. Purity and homogeneity are typically determined using analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high-performance liquid chromatography. The term "purified" can refer to a nucleic acid or protein giving rise to essentially one band in an electrophoretic gel. For proteins that may be subject to modifications, such as phosphorylation or glycosylation, different modifications may result in different isolated proteins that can be purified separately.
[0087] An "isolated polynucleotide" refers to a nucleic acid (e.g., DNA) that is free of the genes that flank it in the naturally occurring genome of the organism from which the nucleic acid molecule of the invention is derived. Thus, the term includes recombinant DNA that is incorporated, for example, within a vector, an autonomously replicating plasmid or virus, or within the genomic DNA of a prokaryote or eukaryote, or that exists as a separate molecule independent of other sequences (e.g., cDNA or genomic DNA fragments generated by PCR or restriction endonuclease digestion). Furthermore, the term includes RNA molecules transcribed from DNA molecules, as well as recombinant DNA that is part of a hybrid gene that encodes additional polypeptide sequences.
[0088] By "isolated polypeptide" is meant a polypeptide of the invention that has been separated from components with which it is naturally associated. Typically, a polypeptide is isolated when it is at least 60%, by weight, free from the proteins and naturally-occurring organic molecules with which it is naturally associated. Preferably, the preparation is at least 75%, more preferably at least 90%, and most preferably at least 99%, by weight, a polypeptide of the invention. Isolated polypeptides of the invention may be obtained, for example, by extraction from a natural source, by expression of a recombinant nucleic acid encoding such a polypeptide, or by chemical protein synthesis. Purity can be measured by any appropriate method, for example, by column chromatography, polyacrylamide gel electrophoresis, or by HPLC analysis.
[0089] The term "linker," as used herein, can refer to a covalent linker (e.g., a covalent bond), a non-covalent linker, a chemical group, or a molecule that links two molecules or moieties, e.g., two components of a protein complex or ribonucleocomplex, or two domains of a fusion protein, e.g., a polynucleotide-programmable DNA-binding domain (e.g., dCas9) and a deaminase domain (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase). A linker can join different components or different parts of a component of a base editor system. For example, in some embodiments, a linker can join the guide polynucleotide-binding domain of a polynucleotide-programmable nucleotide-binding domain and the catalytic domain of a deaminase. In some embodiments, a linker can join a CRISPR polypeptide and a deaminase. In some embodiments, a linker can join Cas9 and a deaminase. In some embodiments, a linker can join a dCas In some embodiments, a linker can join nCas9 and a deaminase. In some embodiments, a linker can join nCas9 and a deaminase. In some embodiments, a linker can join a guide polynucleotide and a deaminase. In some embodiments, a linker can join a deamination component of a base editor system and a polynucleotide-programmable nucleotide-binding component. In some embodiments, a linker can join an RNA-binding portion of a deamination component of a base editor system and a polynucleotide-programmable nucleotide-binding component. In some embodiments, a linker can join an RNA-binding portion of a deamination component of a base editor system and an RNA-binding portion of a polynucleotide-programmable nucleotide-binding component. A linker can be located between or flank two groups, molecules, or other moieties, connected to each other via a covalent or non-covalent bond, thereby connecting the two. In some embodiments, the linker can be an organic molecule, group, polymer, or chemical moiety.In some embodiments, the linker can be a polynucleotide. In some embodiments, the linker can be a DNA linker. In some embodiments, the linker can be an RNA linker. In some embodiments, the linker can comprise an aptamer capable of binding to a ligand. In some embodiments, the ligand can be a carbohydrate, peptide, protein, or nucleic acid. In some embodiments, the linker can comprise an aptamer that can be derived from a riboswitch. The riboswitch from which the aptamer is derived can be selected from a theophylline riboswitch, a thiamine pyrophosphate (TPP) riboswitch, an adenosine cobalamin (AdoCbl) riboswitch, an S-adenosylmethionine (SAM) riboswitch, an SAH riboswitch, a flavin mononucleotide (FMN) riboswitch, a tetrahydrofolate riboswitch, a lysine riboswitch, a glycine riboswitch, a purine riboswitch, a GlmS riboswitch, or a prequosin 1 (PreQ1) riboswitch. In some embodiments, the linker may comprise an aptamer bound to a polypeptide or protein domain, such as a polypeptide ligand. In some embodiments, the polypeptide ligand may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the polypeptide ligand may be part of a base editor system component. For example, a nucleic acid base editing component may comprise a deaminase domain and an RNA recognition motif.
[0090] In some embodiments, the linker can be one amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker can be about 5-100 amino acids in length, e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, or 90-100 amino acids in length. In some embodiments, the linker can be about 100-150, 150-200, 200-250, 250-300, 300-350, 350-400, 400-450, or 450-500 amino acids in length. Longer or shorter linkers are also contemplated.
[0091] In some embodiments, a linker connects the gRNA binding domain of an RNA-programmable nuclease, including a Cas9 nuclease domain, to the catalytic domain of a nucleic acid editing protein (e.g., cytidine or adenosine deaminase). In some embodiments, a linker connects dCas9 to a nucleic acid editing protein. For example, a linker is disposed between or flanked by two groups, molecules, or other moieties, and is connected to each other via a covalent bond, thereby connecting the two. In some embodiments, the linker is one amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is between 5 and 200 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140, 150, 160, 175, 180, 190, or 200 amino acids in length.
[0092] In some embodiments, the base editor domain is fused via a linker comprising the amino acid sequence SGGSSGSETPGTSESATPESSGGS, SGGSSGGSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS. In some embodiments, the base editor domain is fused via a linker comprising the amino acid sequence SGSETPGTSESATPES, sometimes referred to as an XTEN linker. In some embodiments, the linker is 24 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In some embodiments, the linker is 40 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS. In some embodiments, the linker is 64 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGSSGGS. In some embodiments, the linker is 92 amino acids in length. In some embodiments, the linker comprises the amino acid sequence PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS.
[0093] By "marker" is meant any protein or polynucleotide that contains an alteration in expression level or activity that is associated with a disease or disorder.
[0094] The term "mutation" as used herein refers to the substitution of a residue within a sequence, for example, a nucleic acid sequence or an amino acid sequence, the substitution of a residue with another residue, or the deletion or insertion of one or more residues within a sequence. In some embodiments, an insertion is a gene conversion that replaces all or part of the wild-type sequence. Mutations are typically described herein by identifying the original residue, followed by its position within the sequence, and the identity of the newly substituted residue. Various methods for making the amino acid substitutions (mutations) described herein are well known in the art and are described, for example, by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)).
[0095] In some embodiments, the base editors disclosed herein can efficiently generate "mutations of interest," e.g., point mutations, etc., in a nucleic acid (e.g., a nucleic acid in a genome of a subject), without generating a significant number of unintended mutations, e.g., unintended point mutations, etc. In some embodiments, a mutation of interest is a mutation generated by a specific base editor (e.g., a cytidine base editor or an adenosine base editor) attached to a guide polynucleotide (e.g., a gRNA) specifically designed to generate the mutation of interest.
[0096] Generally, mutations made or identified in a sequence (e.g., an amino acid sequence described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain the mutation. Those skilled in the art will readily understand how to determine the position of amino acid and nucleic acid sequence mutations relative to a reference sequence.
[0097] The term "non-conservative mutation" includes amino acid substitutions between different groups, such as tryptophan to lysine or serine to phenylalanine. In this case, it is preferred that the non-conservative amino acid substitution does not interfere with or inhibit the biological activity of the functional variant. The non-conservative amino acid substitution can enhance the biological activity of the functional variant, such that the biological activity of the functional variant is increased compared to the wild-type protein.
[0098] The terms "nuclear localization sequence," "nuclear localization signal," or "NLS" refer to an amino acid sequence that facilitates the import of a protein into a cell nucleus. Nuclear localization sequences are known in the art and are described, for example, in Plank et al., International PCT Application PCT / EP2000 / 011690, filed November 23, 2000, and WO / 2001 / 038547, published May 31, 2001, the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, as described, for example, in Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. Optimized sequences useful in the methods of the present invention are shown in Figures 8A-8E (Koblan et al., supra). In some embodiments, the NLS comprises the amino acid sequence KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.
[0099] The terms "nucleobase," "nitrogenous base," or "base," as used interchangeably herein, refer to nitrogen-containing biological compounds that form the building blocks of nucleosides and, further, nucleotides. The ability of nucleobases to base pair and stack one on top of another directly leads to the formation of long-chain helical structures such as ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). The five nucleobases, adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), are referred to as major or canonical. Adenine and guanine are derived from purines, while cytosine, uracil, and thymine are derived from pyrimidines. DNA and RNA can also contain other (non-major) bases that are modified. Non-limiting exemplary modified nucleobases include hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydromethylcytosine. Hypoxanthine and xanthine can both be created through deamination (replacement of an amine group with a carbonyl group) in the presence of mutagens. Hypoxanthine can be modified from adenine. Xanthine can be modified from guanine. Uracil can result from the deamination of cytosine. A "nucleoside" consists of a nucleobase and a five-carbon sugar (either ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxyuridine, and deoxycytidine. Examples of nucleosides containing modified nucleobases include inosine (I), xanthosine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), and pseudouridine (Ψ). A "nucleotide" consists of a nucleobase, a five-carbon sugar (either ribose or deoxyribonucleotide), and at least one phosphate group.
[0100] The terms "nucleic acid" and "nucleic acid molecule," as used herein, refer to a compound comprising a nucleobase and an acidic moiety, e.g., a nucleoside, a nucleotide, or a polymer of nucleotides. Typically, polymeric nucleic acids, e.g., nucleic acid molecules comprising three or more nucleotides, are linear molecules in which adjacent nucleotides are linked to each other via phosphodiester bonds. In some embodiments, "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides). In some embodiments, "nucleic acid" refers to an oligonucleotide chain comprising three or more individual nucleotide residues. As used herein, the terms "oligonucleotide" and "polynucleotide" can be used interchangeably to refer to a polymer of nucleotides (e.g., a stretch of at least three nucleotides). In some embodiments, "nucleic acid" encompasses RNA and single- and / or double-stranded DNA. Nucleic acids may occur naturally, for example, in the context of a genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, chromatid, or other naturally occurring nucleic acid molecule. Alternatively, a nucleic acid molecule may be a non-natural molecule, e.g., recombinant DNA or RNA, an artificial chromosome, an engineered genome, or a fragment thereof, or synthetic DNA, RNA, or DNA / RNA hybrid, or may contain non-natural nucleotides or nucleosides. Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms include nucleic acid analogs, e.g., analogs having other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems and optionally purified, chemically synthesized, etc. Where appropriate, e.g., in the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, e.g., analogs having chemically modified bases or sugars, and backbone modifications. Nucleic acid sequences are presented in a 5' to 3' direction unless otherwise indicated.In some embodiments, nucleic acids are selected from naturally occurring nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methyl cytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine; chemically modified bases; biologically modified bases (e.g., methylated bases); intercalating bases; modified sugars (2'-e.g., fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages).
[0101] The term "nucleic acid programmable DNA binding protein" or "napDNAbp" may be used interchangeably with "polynucleotide programmable nucleotide binding domain" to refer to a protein associated with a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid or guide polynucleotide (e.g., gRNA), that guides the napDNAbp to a specific nucleic acid sequence. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable DNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable RNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a Cas9 protein. The Cas9 protein may be associated with a guide RNA that guides the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the napDNAbp is a Cas9 domain, such as a nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Non-limiting examples of nucleic acid programmable DNA binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i.Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csn3, Csn4, Csn5, Csn6, Csn7, Csn8, Csn9, Csn10, Csn11, Csn12, Csn13, Csn14, Csn15, Csn16, Csn17, Csn18, Csn19, Csn20, Csn21, Csn22, Csn23, Csn24, Csn25, Csn26, Csn27, Csn28, Csn29, Csn210, Csn211, Csn212, Csn213, Csn22, Csn23, Csn24, Csn25, Csn26, Csn27, Csn28, Csn29, Csn214, Csn215, Csn216, Csn217, Csn22, Csn23, Csn24, Csn25, Csn26, Csn27, Csn Examples of such an effector include sm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V Cas effector proteins, type VI Cas effector proteins, CARF, DinG, homologs thereof, or modified or engineered versions thereof. Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, even though they may not be specifically listed herein. See, for example, Makarova et al. "Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?" CRISPR J. 2018 Oct;1:325-336. doi: 10.1089 / crispr.2018.0033; Yan et al., "Functionally diverse type V CRISPR-Cas systems" Science. 2019 Jan 4;363(6422):88-91. doi: 10.1126 / science.aav7271, the entire contents of each of which are incorporated herein by reference.
[0102] The term "nucleobase editing domain" or "nucleobase editing protein," as used herein, refers to a protein or enzyme that can catalyze the modification of nucleobases in RNA or DNA, such as the deamination of cytosine (or cytidine) to uracil (or uridine) or thymine (or thymidine), and adenine (or adenosine) to hypoxanthine (or inosine), as well as the addition and insertion of non-templated nucleotides. In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., adenine deaminase or adenosine deaminase; or cytidine deaminase or cytosine deaminase). In some embodiments, the nucleobase editing domain is one or more deaminase domains (e.g., adenine deaminase or adenosine deaminase, and cytidine or cytosine deaminase). In some embodiments, the nucleobase editing domain can be a naturally occurring nucleobase editing domain. In some embodiments, the nucleobase-editing domain is an engineered or evolved nucleobase-editing domain from a naturally occurring nucleobase-editing domain. The nucleobase-editing domain can be from any organism, such as a bacterium, human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse.
[0103] As used herein, "obtaining," as in "obtaining an agent," includes synthesizing, isolating, extracting, purchasing, or otherwise acquiring the agent.
[0104] As used herein, "patient" or "subject" refers to a mammalian subject or individual diagnosed with, at risk for, or suspected of developing a disease or disorder. In some embodiments, a subject with a mutation in the gene encoding SDSP is identified as having or at risk for developing Shwachman-Diamond Syndrome (SDS). In some embodiments, the term "patient" refers to a mammalian subject who has a higher-than-average likelihood of developing a disease or disorder. Exemplary patients include humans, non-human primates, cats, dogs, pigs, cows, felines, horses, camels, llamas, goats, sheep, rodents (e.g., mice, rabbits, rats, gerbils, or guinea pigs), and other mammals that can benefit from the therapies disclosed herein. Exemplary human patients can be male and / or female.
[0105] A "patient in need thereof" or "subject in need thereof" is referred to herein as a patient who has been diagnosed with, is at risk of, is predetermined to have, or is suspected of having a disease or disorder, such as SDS.
[0106] The terms "pathogenic mutation," "pathogenic variant," "disease-causing mutation," "disease-causing variant," "deleterious mutation," or "predisposing mutation" refer to a genetic alteration or mutation that increases an individual's susceptibility or predisposition to a particular disease or disorder. In some embodiments, the pathogenic mutation comprises an alteration within a splice acceptor site or a splice donor site in a polynucleotide encoding an SBDS protein. In some embodiments, the pathogenic mutation alters the splicing of the polynucleotide encoding the SBDS protein, resulting in, for example, protein truncation or other negative effect on the expression or activity of the SBDS protein.
[0107] The terms "protein," "peptide," "polypeptide," and their grammatical equivalents are used interchangeably herein and refer to a polymer of amino acid residues linked together by peptide (amide) bonds. The term refers to a protein, peptide, or polypeptide of any size, structure, or function. Typically, a protein, peptide, or polypeptide is at least three amino acids in length. A protein, peptide, or polypeptide can refer to an individual protein or a collection of proteins. One or more amino acids in a protein, peptide, or polypeptide can be modified with chemical entities for conjugation, functionalization, or other modification, such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups, linkers, etc. A protein, peptide, or polypeptide can be a single molecule or a multi-molecular complex. A protein, peptide, or polypeptide can be only a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide can be natural, recombinant, synthetic, or any combination thereof. As used herein, the term "fusion protein" refers to a hybrid polypeptide containing protein domains from at least two different proteins. One protein can be located at the amino-terminal (N-terminal) portion or the carboxy-terminal (C-terminal) portion of the fusion protein, thus forming an amino-terminal fusion protein or a carboxy-terminal fusion protein, respectively. The protein can include different domains, such as a nucleic acid-binding domain of a nucleic acid editing protein (e.g., the gRNA-binding domain of Cas9, which directs binding of the protein to a target site) and a nucleic acid cleavage domain or catalytic domain. In some embodiments, the protein includes a proteinaceous portion, e.g., an amino acid sequence constituting the nucleic acid-binding domain, and an organic compound, e.g., a compound that can act as a nucleic acid cleavage agent. In some embodiments, the protein is complexed with or associated with a nucleic acid, e.g., RNA or DNA.Any protein described herein can be produced by any method known in the art. For example, the protein described herein can be produced through recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known, including those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the entire contents of which are incorporated herein by reference.
[0108] The polypeptides and proteins disclosed herein (including functional portions and functional variants thereof) can contain synthetic amino acids in place of one or more naturally occurring amino acids. Such synthetic amino acids are known in the art and include, for example, aminocyclohexanecarboxylic acid, norleucine, α-amino n-decanoic acid, homoserine, S-acetylaminomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine, β-phenylserine, β-hydroxyphenylalanine, phenylglycine, α-naphthylalanine, cyclohexylalanine, cyclohexylglycine, indoline- 2-carboxylic acid, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, aminomalonic acid monoamide, N'-benzyl-N'-methyl-lysine, N',N'-dibenzyl-lysine, 6-hydroxylysine, ornithine, α-aminocyclopentanecarboxylic acid, α-aminocyclohexanecarboxylic acid, α-aminocycloheptanecarboxylic acid, α-(2-amino-2-norbornane)-carboxylic acid, α,γ-diaminobutyric acid, α,β-diaminopropionic acid, homophenylalanine, and α-tert-butylglycine. Polypeptides and proteins may be associated with post-translational modifications of one or more amino acids of the polypeptide construct. Non-limiting examples of post-translational modifications include phosphorylation, acylation, including acetylation and formylation, glycosylation (including N-linked and O-linked), amidation, hydroxylation, alkylation, including methylation and ethylation, ubiquitination, addition of pyrrolidone carboxylic acid, formation of disulfide bridges, sulfation, myristoylation, palmitoylation, isoprenylation, farnesylation, geranylation, glypiation, lipoylation, and iodination.
[0109] The term "recombinant" as used herein in the context of a protein or nucleic acid refers to a protein or nucleic acid that does not occur in nature but is the product of human engineering. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that contains at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations compared to any naturally occurring sequence.
[0110] By "reduce" is meant a negative alteration of at least 10%, 25%, 50%, 75%, or 100%.
[0111] "Reference" refers to a standard or control condition. In one embodiment, the reference is a wild-type or healthy cell. For example, a wild-type or healthy cell can be derived from or obtained from a healthy and / or disease-free subject. In a specific embodiment, a wild-type or healthy cell is a cell that expresses a wild-type SBDS protein (i.e., an SBDS protein that is the product of a wild-type SBDS gene that exhibits wild-type splicing). In other embodiments, and without limitation, the reference is an untreated cell that is not subjected to the test condition or is subjected to a placebo or normal saline, culture medium, buffer, and / or a control vector that does not contain the polynucleotide of interest.
[0112] A "reference sequence" is a defined sequence used as a basis for sequence comparison. A reference sequence can be a subset or the entirety of a specified sequence, for example, a segment of a full-length cDNA or gene sequence, or the complete cDNA or gene sequence. For polypeptides, the length of a reference polypeptide sequence will generally be at least about 16 amino acids, at least about 20 amino acids, at least about 25 amino acids, at least about 35 amino acids, at least about 50 amino acids, at least about 100 amino acids, or at least about 100 amino acids. For nucleic acids, the length of a reference nucleic acid sequence will generally be at least about 50 nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, at least about 100 nucleotides, or at least about 300 nucleotides, or any integer between or about 50 nucleotides. In some embodiments, the reference sequence is the wild-type sequence of a protein of interest. In other embodiments, the reference sequence is a polynucleotide sequence encoding the wild-type protein.
[0113] The terms "RNA-programmable nuclease" and "RNA-guided nuclease" are used in conjunction with one or more RNAs that are not targets for cleavage (e.g., bind to or associate with such RNAs). In some embodiments, an RNA-programmable nuclease in a complex with an RNA may be referred to as a nuclease:RNA complex. Typically, the bound RNA is referred to as a guide RNA (gRNA). In some embodiments, the RNA-programmable nuclease is a (CRISPR-associated system) Cas9 endonuclease, such as Cas9 (Csnl) from Streptococcus pyrogenes (e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes." Ferretti JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chylinski K., Sharma CM., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607 (2011).
[0114] "Shwachman-Bodian-Diamond Syndrome (SBDS) protein" refers to a polypeptide or fragment thereof having at least about 85% amino acid sequence identity with NCBI Accession Number NP_057122.2 and having the biological activity of SBDS. In various embodiments, the biological activity of SBDS refers to a role in RNA processing, ribosome biogenesis, or binding to an antibody that specifically binds to the SBDS protein.
[0115] An example of the amino acid sequence of an SBDS protein is shown below: MSIFTPTNQI RLTNVAVVRM KRAGKRFEIA CYKNKVVGWR SGVEKDLDEV LQTHSVFVNV SKGQVAKKED LISAFGTDDQ TEICKQILTK GEVQVSDKER HTQLEQMFRD IATIVADKCV NPETKRPYTV ILIERAMKDI HYSVKTNKST KQQALEVIKQ LKEKMKIERA HMRLRFILPV NEGKKLKEKL KPLIKVIESE DYGQQLEIVC LIDPGCFREI DELIKKETKG KGSLEVLNLK DVEEGDEKFE.
[0116] In a specific embodiment, the SBDS protein comprises a protein truncation.
[0117] "Shwachman-Bodian-Diamond Syndrome (SBDS) polynucleotide" means a nucleic acid sequence encoding an SBDS protein. An example of an SBDS polynucleotide sequence is shown in NM_016038.2, reproduced below. The open reading frame (ORF) of the SBDS polynucleotide extends from nucleotide 185 to 937 (underlined). GTAAGTAAGC CTGCCAGACA CACTGTGACG GCTGCCTGAA GCTAGTGAGT CGCGGCGCCG CGCACTGGTG GTTGGGTCAG TGCCGCGCGC CGATCGGTCG TTACCGCGAG GCGCTGGTGG CCTTCAGGCT GGACGGCAGCCGCT CGTCTGCAGCG ATGTCG ATCTTCACCC CCACCAACCA GATCCGCCTA ACCAATGTGG CCGTGGTACG GATGAAGCGT GCCGGGAAGC GCTTCGAAAT CGCCTGCTAC AAAAAAAAAGG TCGTCGGCTG GCGGAGCGGC GTGGAAAAAG ACCTCGATGA AGTTCTGCAG ACCCACTCAG TGTTTGTAAA TGTTTCTAAA GGTCAGGTTG CCAAAAAGGA AGATCTCATC AGTGCGTTTG GAACAGATGA CCAAACTGAA ATCTGTAAGC AGATTTTGAC TAAAGGAGAA GTTCAAGTAT CAGATAAAGA AAGACACACA CAACTGGAGC AGATGTTTAG GGACATTGCA ACTATTGTGG CAGACAAATG TGTGAATCCT GAAACAAAGA GACCATACAC CGTGATCCTT ATTGAGAGAG CCATGAAGGA CATCCACTAT TCGGTGAAAA CCAACAAGAG TACAAAACAG CAGGCTTTGG AAGTGATAAA GCAGTTAAAA GAGAAAATGA AGATAGAACG TGCTCACATG AGGCTTCGGT TCATCCTTCC AGTCAATGAA GGCAAGAAGC TGAAAGAAAA GCTCAAGCCA CTGATCAAGG TCATAGAAAG TGAAGATTAT GGCCAACAGT TAGAAATCGT ATGTCTGATT GACCCGGGCT GCTTCCGAGA AATTGATGAG CTAATAAAAA AGGAAACTAA AGGCAAAGGT TCTTTGGAAG TACTCAATCT GAAAGATGTA GAAGAAGGAG ATGAGAAATT TGAATGA CAC CCATCAATCT CTTCACCTCT AAAACACTAA AGTGTTTCCG TTTCCGACGG CACTGTTTCA TGTCTGTGGT CTGCCAAATA CTTGCTTAAA CTATTTGACA TTTTCTATCT TTGTGTTAAC AGTGGACACA GCAAGGCTTT CCTAACATTATAG T TATGTAATAATGA GGGTCTAAAT CCTAAAGCAA AATTGAAACT CCAAGATGCA AAGTCCAGAG TGGCATTTTG CTACTCTGTC TCATGCCTTG ATAGCTTTCC AAAATGAAG TTACTTGAGG CAGCTCTTGT GGGTGAAAAG TTATTTGTAC AGTAAGGTAGTA GATCAGATGATT TTCCTAAAAA AGAAAACATA TGATGCTTCA TTTCTACTTA ATGGAACTTG TGTTCTGAGG GTCATTATGG TATCGTAATG TAAAGCTTGG ATGATGTTCC TGATTATCTG AGAAACAGAT ATAGAAAAAT TGTGCCTGAC TTGACATTCTTCA TTGGTTAAAA AATAAAAGTC ACTTATTTCT AATTCTTAAA GTTTATATA TATATTTATA TAGCTAAAAT TGTATGTAAT HELPAACC ACTCTTATGT TTATT
[0118] In some embodiments, the Shwachman-Bodian-Diamond Syndrome (SBDS) polynucleotide comprises a polynucleotide derived from an SBDS pseudogene. In some embodiments, the SBDS polynucleotide comprises a mutation resulting from a gene conversion associated with SDS (e.g., a 258+2T>C and / or a 183-184TA>CT mutation), alone or in combination with other alterations in the SBDS pseudogene.
[0119] "Shwachman-Bodian-Diamond Syndrome (SBDS) pseudogene" means a nucleic acid sequence having at least about 85% nucleic acid sequence identity to an SBDS polynucleotide. In one embodiment, exemplary pseudogenes include the following fragments thereof: >NR_024109.1 Human (Homo sapiens) SBDS pseudogene 1 (SBDSP1), transcript variant 4, non-coding RNA CCTTTTTGGGCGTGGAAAGATGGCGGTAAAAGCCACAATGCGCAGGCGTCATCGCTCACTTCTCCCCTCC CGGCTTCTGCTCCACCTGACGCCTGCGCAGTAAGTAAGCCTGCCAGACACGCTGTGGCGGCTGCCTGAAG CTAGTGAGTCGCGGCGCCGCGCACTTGTGGTTGGGTCAGTGCCGCGCGCCGCTCGGTCGTTACCGCGAGG CGCTGGTGGCCTTCAGGCTGGACGGCGCGGGTCAGCCCTGGTTTGCCGGCTTCTGGGTCTTTGAACAGCC GCGATGTCGATCTTCACCCCCACCAACCAGATCCGCCTAACCAATGTGGCCGTGGTACGGATGAAGCGCG CCAGGAAGCGCTTCGAAATCGCCTGCTACAGAAACAAGGTCGTCGGCTGGCGGAGCGGCTTATTTTGACT AAAGGAGAAGTTCAAGTATCAGATAAAGACACACACAACTGGAGCAGATGTTTAGGGACATTGCAATTAT TGTGGCAGACAAATGTGTGACTCCTGAAACAAAGAGACCATACACCGTGATCCTTATTGAGAGAGCCATG AAGGACATCCACTATTTGGTGAAAACCAACAGGAGTACAAAACAGCAGGCTTTGGAAGTGATAAAGCAGT TAAAAGAGAAAATGAAGATAGAACGTGCTCACATGAGGCTTCAGTTCATCCTTCCAGTGAATGAAGGCAA GAAGCTGAAAGAAAAGCTCAAGCCACTGATCAAGGTCATAGAAAGTAAAGATTATGGCCAACAGTTAGAA ATCGTAAGAGTCAAATATTTTCTTTGCTTCATGTTACCTAAATATTGTATTCTCTAGTAATAAATTTGTA GCAAACATTCAAAAAAAAAAAAAAAAAAAA >NR_024110.1 Homo sapiens SBDS pseudogene 1 (SBDSP1), transcript variant 1, non-coding RNA CCTTTTTGGGCGTGGAAAGATGGCGGTAAAAGCCACAATGCGCAGGCGTCATCGCTCACTTCTCCCCTCC CGGCTTCTGCTCCACCTGACGCCTGCGCAGTAAGTAAGCCTGCCAGACACGCTGTGGCGGCTGCCTGAAG CTAGTGAGTCGCGGCGCCGCGCACTTGTGGTTGGGTCAGTGCCGCGCGCCGCTCGGTCGTTACCGCGAGG CGCTGGTGGCCTTCAGGCTGGACGGCGCGGGTCAGCCCTGGTTTGCCGGCTTCTGGGTCTTTGAACAGCC GCGATGTCGATCTTCACCCCCACCAACCAGATCCGCCTAACCAATGTGGCCGTGGTACGGATGAAGCGCG CCAGGAAGCGCTTCGAAATCGCCTGCTACAGAAACAAGGTCGTCGGCTGGCGGAGCGGCTTGGAAAAAGA CCTTGATGAAGTTCTGCAGACCCACTCAGTGTTTGTAAATGTTTCCTAAGGTCAGGTTGCCAAGAAGGAA GATCTCATCAGTGCGTTTGGAACAGATGACCAAACTGAAATCTATTTTGACTAAAGGAGAAGTTCAAGTA TCAGATAAAGACACACACAACTGGAGCAGATGTTTAGGGACATTGCAATTATTGTGGCAGACAAATGTGT GACTCCTGAAACAAAGAGACCATACACCGTGATCCTTATTGAGAGAGCCATGAAGGACATCCACTATTTG GTGAAAACCAACAGGAGTACAAAACAGCAGGCTTTGGAAGTGATAAAGCAGTTAAAAGAGAAAATGAAGA TAGAACGTGCTCACATGAGGCTTCAGTTCATCCTTCCAGTGAATGAAGGCAAGAAGCTGAAAGAAAAGCT CAAGCCACTGATCAAGGTCATAGAAAGTAAAGATTATGGCCAACAGTTAGAAATCGTATGTCTGATTGAC CTGGGCTGCTTCCGAGAAATTGATGAGCTAATAAAAAAGGAAACCAAAGGCAAAGGTTCTTTGGAAGTACTCAATCTGAAAGATTTGAAGAAGGAGATGAGAAATTTGAATGACACCCATCAGTCTCTTCACCTCTAAAA CACTAAAGTGTTTTCGTTTCCAACAGCACTGTTTCATGTCTGTGGTCTGCCAAATACTTGCTCAAACTAT TTGACATTTTCTATCTTTGTGTTAACAGTGGACACAGCAAGGCTTTCCTACATAAGTATAATAATGTGGG AATGATTTGGTTTTAATTATAAACTGGGGTCTAAATCCTAAAGCAAAATTGAAACTCCAGGATGCAAAAT CCAGAGTGGCATTTTGCTACTCTGTCTCATGCCTTGATAGCTTTCCAAAATGAAAGTTACTTGAGGCAGC TCTTGTGGGTGAAAAGTTTTTTGTACAGTAGAGTAAGATTATTAGGGGTATGTCTATACGACAAAAGGGG GGTCTTTCCTAAAAAAGAAAACATGATGCTTCATTTCTACTTAATGGAACTTGTGTTCTGAGGGTCATTA TGGTATCGTAATATAAAGCTTGGATGATGTTCCTGATTATCTGAGAAACAGATATAGAAAAATTGTGTCG GACTTAAATAATTTTCGTTGAACATGCTGCCATAACTTAGATTATTCTTGGTTAAAAAATAAAAGTCACT TATTTCTAATTCTTAAAGTTTATAATATATATTAATATAGCTAAAATTGTATGTAATCAATAAAACCACT CTTATGTTTATTAAACTATGGCTTGTGTTTCTAGACAAAAAAAAAAAAAAAAAA >NR_024111.1 Homo sapiens SBDS pseudogene 1 (SBDSP1), transcript variant 2, non-coding RNA CCTTTTTGGGCGTGGAAAGATGGCGGTAAAAGCCACAATGCGCAGGCGTCATCGCTCACTTCTCCCCTCC CGGCTTCTGCTCCACCTGACGCCTGCGCAGTAAGTAAGCCTGCCAGACACGCTGTGGCGGCTGCCTGAAG CTAGTGAGTCGCGGCGCCGCGCACTTGTGGTTGGGTCAGTGCCGCGCGCCGCTCGGTCGTTACCGCGAGG CGCTGGTGGCCTTCAGGCTGGACGGCGCGGGTCAGCCCTGGTTTGCCGGCTTCTGGGTCTTTGAACAGCC GCGATGTCGATCTTCACCCCCACCAACCAGATCCGCCTAACCAATGTGGCCGTGGTACGGATGAAGCGCG CCAGGAAGCGCTTCGAAATCGCCTGCTACAGAAACAAGGTCGTCGGCTGGCGGAGCGGCTTATTTTGACT AAAGGAGAAGTTCAAGTATCAGATAAAGACACACACAACTGGAGCAGATGTTTAGGGACATTGCAATTAT TGTGGCAGACAAATGTGTGACTCCTGAAACAAAGAGACCATACACCGTGATCCTTATTGAGAGAGCCATG AAGGACATCCACTATTTGGTGAAAACCAACAGGAGTACAAAACAGCAGGCTTTGGAAGTGATAAAGCAGT TAAAAGAGAAAATGAAGATAGAACGTGCTCACATGAGGCTTCAGTTCATCCTTCCAGTGAATGAAGGCAA GAAGCTGAAAGAAAAGCTCAAGCCACTGATCAAGGTCATAGAAAGTAAAGATTATGGCCAACAGTTAGAA ATCGTATGTCTGATTGACCTGGGCTGCTTCCGAGAAATTGATGAGCTAATAAAAAAGGAAACCAAAGGCA AAGGTTCTTTGGAAGTACTCAATCTGAAAGATTTGAAGAAGGAGATGAGAAATTTGAATGACACCCATCA GTCTCTTCACCTCTAAAACACTAAAGTGTTTTCGTTTCCAACAGCACTGTTTCATGTCTGTGGTCTGCCAAATACTTGCTCAAACTATTTGACATTTTCTATCTTTGTGTTAACAGTGGACACAGCAAGGCTTTCCTACA TAAGTATAATAATGTGGGAATGATTTGGTTTTAATTATAAACTGGGGTCTAAATCCTAAAGCAAAATTGA AACTCCAGGATGCAAAATCCAGAGTGGCATTTTGCTACTCTGTCTCATGCCTTGATAGCTTTCCAAAATG AAAGTTACTTGAGGCAGCTCTTGTGGGTGAAAAGTTTTTTGTACAGTAGAGTAAGATTATTAGGGGTATG TCTATACGACAAAAGGGGGGTCTTTCCTAAAAAAGAAAACATGATGCTTCATTTCTACTTAATGGAACTT GTGTTCTGAGGGTCATTATGGTATCGTAATATAAAGCTTGGATGATGTTCCTGATTATCTGAGAAACAGA TATAGAAAAATTGTGTCGGACTTAAATAATTTTCGTTGAACATGCTGCCATAACTTAGATTATTCTTGGT TAAAAAATAAAAGTCACTTATTTCTAATTCTTAAAGTTTATAATATATATTAATATAGCTAAAATTGTAT GTAATCAATAAAACCACTCTTATGTTTATTAAACTATGGCTTGTGTTTCTAGACAAAAAAAAAAAAAAAA AA >NR_001588.2 Homo sapiens SBDS pseudogene 1 (SBDSP1), transcript variant 3, non-coding RNA CCTTTTTGGGCGTGGAAAGATGGCGGTAAAAGCCACAATGCGCAGGCGTCATCGCTCACTTCTCCCCTCC CGGCTTCTGCTCCACCTGACGCCTGCGCAGTAAGTAAGCCTGCCAGACACGCTGTGGCGGCTGCCTGAAG CTAGTGAGTCGCGGCGCCGCGCACTTGTGGTTGGGTCAGTGCCGCGCGCCGCTCGGTCGTTACCGCGAGG CGCTGGTGGCCTTCAGGCTGGACGGCGCGGGTCAGCCCTGGTTTGCCGGCTTCTGGGTCTTTGAACAGCC GCGATGTCGATCTTCACCCCCACCAACCAGATCCGCCTAACCAATGTGGCCGTGGTACGGATGAAGCGCG CCAGGAAGCGCTTCGAAATCGCCTGCTACAGAAACAAGGTCGTCGGCTGGCGGAGCGGCTTGGAAAAAGA CCTTGATGAAGTTCTGCAGACCCACTCAGTGTTTGTAAATGTTTCCTAAGGTCAGGTTGCCAAGAAGGAA GATCTCATCAGTGCGTTTGGAACAGATGACCAAACTGAAATCTATTTTGACTAAAGGAGAAGTTCAAGTA TCAGATAAAGACACACACAACTGGAGCAGATGTTTAGGGACATTGCAATTATTGTGGCAGACAAATGTGT GACTCCTGAAACAAAGAGACCATACACCGTGATCCTTATTGAGAGAGCCATGAAGGACATCCACTATTTG GTGAAAACCAACAGGAGTACAAAACAGCAGGCTTTGGAAGTGATAAAGCAGTTAAAAGAGAAAATGAAGA TAGAACGTGCTCACATGAGGCTTCAGTTCATCCTTCCAGTGAATGAAGGCAAGAAGCTGAAAGAAAAGCT CAAGCCACTGATCAAGGTCATAGAAAGTAAAGATTATGGCCAACAGTTAGAAATCGTAAGAGTCAAATAT TTTCTTTGCTTCATGTTACCTAAATATTGTATTCTCTAGTAATAAATTTGTAGCAAACATTCAAAAAAAA AAAAAAAAAAAA
[0120] The term "single nucleotide polymorphism (SNP)" refers to a single nucleotide variation occurring at a specific position in the genome, with each variation present in a population to some appreciable extent (e.g., >1%). For example, at a specific base position in the human genome, a C nucleotide is likely to appear in most individuals, while an A occupies this position in a minority of individuals. This means that a SNP is present at this specific position, and the two possible nucleotide variations, C or A, are said to be alleles at this position. SNPs underlie differences in susceptibility to disease. The severity of disease and the way our bodies respond to treatment are also manifestations of genetic variation. SNPs can range within the coding region of a gene, within the non-coding region of a gene, or in intergenic regions (regions between genes). In some embodiments, SNPs within a coding sequence do not necessarily change the amino acid sequence of the protein produced due to the degeneracy of the genetic code. SNPs in coding regions are of two types: synonymous SNPs and non-synonymous SNPs. Synonymous SNPs do not affect the protein sequence, whereas nonsynonymous SNPs change the amino acid sequence of a protein. There are two types of nonsynonymous SNPs: missense and nonsense. SNPs not in protein-coding regions can also affect gene splicing, transcription factor binding, messenger RNA degradation, or the sequence of non-coding RNA. Gene expression affected by this type of SNP is called eSNP (expressed SNP) and can be upstream or downstream from the gene. Single nucleotide variants (SNVs) are single nucleotide variations with no frequency restriction and can occur somatically. Somatic single nucleotide variations can also be single nucleotide changes.
[0121] By "specifically binds" is meant a nucleic acid molecule, polypeptide, or complex thereof (e.g., a nucleic acid programmable DNA binding protein and a guide nucleic acid), compound, or molecule that recognizes and binds to a polypeptide and / or nucleic acid molecule of the invention, but does not substantially recognize and bind to other molecules in a sample, e.g., a biological sample.
[0122] Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to an endogenous nucleic acid sequence, but will typically exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence will typically be able to hybridize to at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to an endogenous nucleic acid sequence, but will typically exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence will typically be able to hybridize to at least one strand of a double-stranded nucleic acid molecule. "Hybridize" refers to the pairing between complementary polynucleotide sequences (e.g., genes described herein) or portions thereof to form a double-stranded molecule under various conditions of stringency (see, e.g., Wahl, GM and SL Berger (1987) Methods Enzymol. 152:399; Kimmel, AR (1987) Methods Enzymol. 152:507).
[0123] For example, stringent salt concentrations will typically not exceed about 750 mM NaCl and 75 mM trisodium citrate, preferably about 500 mM NaCl and 50 mM trisodium citrate, and even more preferably about 250 mM NaCl and 25 mM trisodium citrate. Low stringency hybridization can be achieved in the absence of organic solvents, such as formamide, while high stringency hybridization can be achieved in the presence of at least about 35% formamide, more preferably at least about 50% formamide. Stringent temperature conditions will typically include a temperature of at least about 30°C, more preferably at least about 37°C, and most preferably at least about 42°C. Various additional parameters, such as hybridization time, concentration of detergent, e.g., sodium dodecyl sulfate (SDS), and the inclusion or exclusion of carrier DNA, are well known in the art. Various levels of stringency can be achieved by combining these various conditions as needed. In a preferred embodiment, hybridization occurs at 30°C in 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS. In a more preferred embodiment, hybridization occurs at 37°C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / mL denatured salmon sperm DNA (ssDNA). In a most preferred embodiment, hybridization occurs at 42°C in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 μg / mL ssDNA. Useful variations on these conditions will be readily apparent to those skilled in the art.
[0124] In most applications, the washing steps following hybridization will also vary in stringency. Wash stringency conditions can be defined by salt concentration and temperature. As noted above, washing stringency can be increased by decreasing salt concentration or increasing temperature. For example, stringent salt concentrations for washing steps will preferably not exceed about 30 mM NaCl and 3 mM trisodium citrate, and most preferably not exceed about 15 mM NaCl and 1.5 mM trisodium citrate. Stringent temperature conditions for washing steps will usually include temperatures of at least about 25°C, more preferably at least about 42°C, and even more preferably at least about 68°C. In one embodiment, the washing step will occur at 25°C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In another embodiment, the wash steps will occur in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS at 42° C. In a more preferred embodiment, the wash steps will occur in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS at 68° C. Additional variations on these conditions will be readily apparent to one of skill in the art.Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.
[0125] "Split" means to divide into two or more pieces.
[0126] "Split Cas9 protein" or "split Cas9" refers to a Cas9 protein provided as an N-terminal fragment and a C-terminal fragment encoded by two separate nucleotide sequences. The polypeptides corresponding to the N-terminal and C-terminal portions of the Cas9 protein may be separated to form a "reconstituted" Cas9 protein. In specific embodiments, the Cas9 protein is split into two fragments within a disordered region of the protein, e.g., as described in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014, or Jiang et al. (2016) Science 351: 867-871. PDB file: 5F9R, each of which is incorporated herein by reference. In some embodiments, the protein is split into two fragments at any C, T, A, or S within the region of SpCas9, between approximately amino acids A292 to G364, F445 to K483, or E565 to T637, or at the corresponding positions in any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or other napDNAbp. In some embodiments, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, this process of splitting the protein into two fragments is referred to as splitting the protein.
[0127] In other embodiments, the N-terminal portion of the Cas9 protein comprises amino acids 1-573 or 1-637 of S. pyrogenes Cas9 wild-type (SpCas9) (NCBI Reference Sequence: NC_002737.2, Uniprot Reference Sequence: Q99ZW2), and the C-terminal portion of the Cas9 protein comprises amino acids 574-1368 or 638-1368 of SpCas9 wild-type.
[0128] The C-terminal portion of a split Cas9 can be joined to the N-terminal portion of a split Cas9 to form a complete Cas9 protein. In some embodiments, the C-terminal portion of the Cas9 protein begins where the N-terminal portion of the Cas9 protein ends. Thus, in some embodiments, the C-terminal portion of the split Cas9 comprises amino acids (551-651) to 1368 of spCas9. "(551-651) to 1368" means beginning with an amino acid between amino acids 551 and 651 (inclusive) and ending with amino acid 1368.For example, the C-terminal portion of split Cas9 is amino acids 551-1368, 552-1368, 553-1368, 554-1368, 555-1368, 556-1368, 557-1368, 558-1368, 559-1368, 560-1368, 561-1368, 562-1368, 563-1368, 564-1368, 565-1368, 566-1368, 567-1368, 568-1368, 569-1368, 570-1368, 571-1368, 572-1368, 573-1368, 574-1368, 575-1368, 576-1368, 577-1368, 578-1368, 579-1368, 580-1368, 581-1368, 582-1368, 583-1368, 584-1368, 585-1368, 586-1368, 587-1368, 588-1368, 589-1368, 590-1368, 591-1368, 592-1368, 593-1368, 594-1368, 595-1368, 596-1368, 597-1368, 598-1368, 599-1368, 599-1 74~1368, 575~1368, 576~1368, 577~1368, 578~1368, 579~1368, 580~1368, 581~1368, 582~1368, 583~1368, 584~1368, 585~1368, 586~1368, 587~1368, 588~1368, 589~1368, 590~1368, 591~1368, 592~1368, 593~1368, 594~1368, 595~1368, 596~1368, 597~1368, 598~1368, 599~1368, 600~1368 8, 601~1368, 602~1368, 603~1368, 604~1368, 605~1368, 606~1368, 607~1368, 608~1368, 609~1368, 610~1368, 611~1368, 612~1368, 613~1368, 614~1368, 615~1368, 616~1368, 617~1368, 618~1368, 619~1368, 620~1368, 621~1368, 622~1368, 623~1368, 624~1368, 625~1368, 626~1368, 627~ 1368, 628-1368, 629-1368, 630-1368, 631-1368, 632-1368, 633-1368, 634-1368, 635-1368, 636-1368, 637-1368, 638-1368, 639-1368, 640-1368, 641-1368, 642-1368, 643-1368, 644-1368, 645-1368, 646-1368, 647-1368, 648-1368, 649-1368, 650-1368, or 651-1368.In some embodiments, the C-terminal portion of the split Cas9 protein comprises part of amino acids 574-1368 or 638-1368 of SpCas9.
[0129] "Subject" means a mammal, including, but not limited to, a human or a non-human mammal, such as a non-human primate (monkey), cow, horse, dog, sheep, or cat. In some embodiments, the subject described herein comprises a pathogenic mutation in an SDS polynucleotide sequence encoding an SBDS protein that identifies the subject as having or predisposed to developing SDS.
[0130] By "substantially identical" is meant that a polypeptide or nucleic acid molecule exhibits at least 50% identity with a reference amino acid sequence (e.g., any one of the amino acid sequences described herein) or nucleic acid sequence (e.g., any one of the nucleic acid sequences described herein). In one embodiment, such a sequence is at least 60%, 65%, 70%, 75%, 80%, or 85%, 90%, 95%, or even 99% identical at the amino acid or nucleic acid level to the sequence used for comparison.
[0131] Sequence identity is typically measured using sequence analysis software (e.g., the Genetics Computer Group's Sequence Analysis Software Package, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, WI 53705, BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. One exemplary approach to determining the degree of identity is to use a sequence analysis to identify closely related sequences. -3 and e -100 The BLAST program may be used, with probability scores between .
[0132] COBALT is used, for example, with the following parameters: a) Alignment parameters: Gap penalty -11, -1 and End-Gap penalty -5, -1; b) CDD parameters: Use RPS BLAST; Blast E-value 0.003; Find and recalculate the saved columns; and c) Query clustering parameters: Use query cluster; word size 4; max cluster distance 0.8; alphabet regular. Using EMBOSS Needle, for example, with the following parameters: a) Matrix: BLOSUM62; b) GAP OPEN: 10; c) GAP EXTEND: 0.5; d)OUTPUT FORMAT:pair; e)END GAP PENALTY:false; f) END GAP OPEN: 10; and g)END GAP EXTEND:0.5.
[0133] The term "target site" refers to a sequence within a nucleic acid molecule that is deaminated by a deaminase or a fusion protein that includes a deaminase (e.g., a dCas9-adenosine deaminase fusion protein or a base editor disclosed herein).
[0134] Because RNA-programmable nucleases (e.g., Cas9) use RNA:DNA hybridization to target DNA cleavage sites, these proteins can, in principle, target any sequence specified by a guide RNA. Methods for using RNA-programmable nucleases, such as Cas9, for site-specific cleavage (e.g., to modify genomes) are known in the art (e.g., Cong, L. et al., Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, WY et al., Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas, the entire contents of each of which are incorporated herein by reference). see Nucleic acids research (2013); Jiang, W. et ah RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature biotechnology 31, 233-239 (2013)).
[0135] As used herein, the terms "treat," "treating," "treatment," and the like refer to reducing, diminishing, attenuating, diminishing, alleviating, or ameliorating a disease or disorder and / or its associated symptoms, or achieving a desired pharmacological and / or physiological effect. It is understood, but not limited to, that treating a disorder or condition does not require complete elimination of the disorder, condition, or its associated symptoms. In some embodiments, the effect is therapeutic, i.e., without limitation, the effect partially or completely reduces, diminishes, prevents, attenuates, ameliorates, diminishes, or cures the intensity of the disease and / or adverse symptoms resulting from the disease. In some embodiments, the effect is prophylactic, i.e., the effect protects against or prevents the occurrence and recurrence of the disease or condition. To this end, the methods disclosed herein comprise administering a therapeutically effective amount of a composition as described herein. In one embodiment, the invention provides treatment for SDS.
[0136] "Uracil glycosylase inhibitor" or "UGI" refers to an agent that inhibits the uracil excision repair system. In one embodiment, the agent is a protein or fragment thereof that binds to a host's uracil DNA glycosylase and prevents the removal of uracil residues from DNA. In one embodiment, UGI is a protein or fragment or domain thereof that can inhibit the uracil DNA glycosylase base excision repair enzyme. In some embodiments, the UGI domain comprises wild-type UGI or a modified version thereof. In some embodiments, the UGI domain comprises a fragment of an exemplary amino acid sequence defined below. In some embodiments, the UGI fragment comprises an amino acid sequence that comprises at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the exemplary UGI sequence shown below. In some embodiments, the UGI comprises an amino acid sequence homologous to an exemplary UGI amino acid sequence as defined below, or a fragment thereof. In some embodiments, the UGI or portion thereof is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, or 100% identical to wild-type UGI or a UGI sequence or portion thereof as defined below. Exemplary UGIs include amino acid sequences such as: >splP14739IUNGI_BPPB2 Uracil DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLT SD APE YKPW ALVIQDS NGENKIKML.
[0137] Ranges provided herein are understood to be shorthand for all values within the range. For example, a range of 1 to 50 is understood to include any number, combination of numbers, or subrange from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.
[0138] The recitation of a listing of a chemical group in any definition of a variable herein includes that definition of the variable as any single group or combination of the listed groups. The recitation of an embodiment of a variable or aspect herein includes that embodiment as any single embodiment or in combination with any other embodiment or portion thereof.
[0139] Any composition or method described herein can be combined with any one or more of the other compositions and methods described herein.
[0140] The description and examples herein provide detailed descriptions of embodiments of the present disclosure. It will be understood that the present disclosure is not limited to the specific embodiments described herein and may therefore vary. Those skilled in the art will recognize that there are numerous variations and modifications of the present disclosure that may fall within its scope.
[0141] All terms are intended to be understood as understood by one of ordinary skill in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0142] The practice of some embodiments disclosed herein employs, unless otherwise indicated, conventional techniques in immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA that are within the skill of the art. See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (FM Ausubel, et al. eds.); the series Methods in Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (MJ MacPherson, BD Hames, and GR Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (RI Freshney, ed. (2010)).
[0143] Although various features of the present disclosure may be described in the context of a single embodiment, the features may also be presented separately or in any suitable combination. Conversely, although the present disclosure may, for clarity, be described herein in the context of separate embodiments, the present disclosure may also be implemented in a single embodiment. The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.
[0144] The features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages herein will be obtained by reference to the following detailed description, which sets forth illustrative embodiments, in which the principles of the present disclosure are utilized, in view of the accompanying drawings as described herein below. [Brief explanation of the drawings]
[0145] [Figure 1A] Figures 1A and 1B show mutations within SBDS that cause SDS. Figure 1A shows a map of SBDS (coding regions in light shading, noncoding regions in dark shading) and a sequence alignment of the exon 2 region of SBDS with the SBDS protein, indicating gene-specific (gray; top) and pseudogene-specific (gray; bottom) sequences. Compared to SBDS, exon 2 of SBDSP, resulting from the conversion event, contains sequence changes predicted to result in protein truncation (underlined). These include an in-frame stop codon at position 184 and a T→C change at 250+10 (corresponding to the invariant T of the donor splice site at 258+2 in SBDS), the latter change resulting in the use of an alternative donor splice site at 250+1 (the position of the invariant splice site is boxed). Figure 1B shows sequence reads for cloned segments derived from the exon 2 region of SBDS, highlighting sequence changes in individuals with SBDS resulting from gene conversion events between SBDS and its pseudogene. Three converted alleles are shown. These include 183-184TA → CT, 258 + 2T → C, and an extended conversion mutation, 183-184TA → CT + 201A → G + 258 + 2T → C. In each case, informative flanking positions, including 141 and 258 + 124, were unconverted (green). [Figure 1B] Same as above. [Figure 2A]Figures 2A-2D are schematic diagrams illustrating strategies for restoring transcription in SBDS genes containing one or more pathogenic mutations. Figure 2A illustrates a strategy for introducing a mutation that eliminates the stop codon and results in expression of an SBDS protein containing an alternative amino acid (e.g., Trp(W)) at amino acid position 62 (e.g., (K62X)). Figures 2B and 2D illustrate strategies for correcting the splice site at nucleotide position 258 (target SNP rs113993993 C→T). Figure 2C illustrates the splice donor position where the canonical splice donor is restored to correct the SNP mutation. [Figure 2B] Same as above. [Figure 2C] Same as above. [Figure 2D] Same as above. [Figure 3-1]Figures 3A-3C present tables showing amino acid positions where substitutions occur in modified Cas9 proteins, e.g., modified SpCas9, to produce Cas9 variants with specificity for an altered PAM 5'-NGC-3' or a 5'-NGC-3' containing a PAM, and plasmid constructs encoding the SpCas9 variant sequences. Cytidine base editors (CBEs) comprising at least one cytidine deaminase and at least one Cas9 variant as described are used to correct mutations in the SBDS gene associated with SDS, as described in Example 3. Figure 3A presents the amino acid positions in the Cas9 protein that were altered from wild-type to produce Cas9 variants (designated by numbers in the left column) that were capable of binding to the NGC PAM. These Cas9 variants are components of the CBEs evaluated in the base editing studies described herein. Figure 3B presents a subset of Cas9 variants that produced particularly good and robust on-target editing with limited bystander effects in the studies. Also shown in Figure 3B is a schematic representation of the Cas9 protein domains and their location within the Cas9 protein sequence. Figure 3C illustrates the plasmid vector components encoding SpCas9 variants with specificity for the modified PAM 5'-NGC-3' as described herein, and the sequence mutations therein. [Figure 3-2] Same as above. [Figure 3-3] Same as above. [Figure 3-4] Same as above. [Figure 3B] Same as above. [Figure 3C-1] Same as above. [Figure 3C-2] Same as above. [Figure 4] FIG. 4 depicts a graph comparing the relative mutation rates of base editing achieved by CBEs with different cytidine deaminases indicated on the abscissa. [Figure 5]5 is a table showing the guide RNAs (gRNAs) used with the CBEs evaluated in the studies described herein. In embodiments, the gRNA sequences were components of the plasmid constructs used in the base editing studies described in the Examples. [Figure 6A] Figures 6A-6C show the percentage of editing (e.g., on-target editing) versus the percentage of bystander editing achieved by NGC CBE variants as described herein and 19-mer and 20-mer gRNAs, such as G88 and G44. In the right-hand graph of Figure 6A and Figure 6B, "PV226" and "PV230" refer to the plasmids used in the study. The PV226 plasmid contains a polynucleotide encoding Cas9 variant #226, the sequence of which is shown in Figures 3A-3C, and the PV230 plasmid contains a polynucleotide encoding Cas9 variant #230, the sequence of which is shown in Figures 3A-3C. The percentage of editing achieved by another NGC CBE containing a different Cas9 variant, the sequence of which is shown in Figures 3A-3C, and the 20-mer gRNA, G44, is shown in Figure 6C. [Figure 6B] Same as above. [Figure 6C] Same as above. [Figure 7A] Figures 7A and 7B show graphs of percent editing by NGC CBEs containing cytidine deaminases and Cas9 variants shown in Table 13 used in conjunction with a 19-mer gRNA (G88) and a 20-mer gRNA (G44) as described in Example 4 herein. [Figure 7B] Same as above. [Figure 8A]Figures 8A-8J show graphs of the percent base editing (on-target editing and bystander editing) achieved by NGC CBEs containing different cytidine deaminases and Cas9 (e.g., SpCas9) variant polypeptides with specific combinations of mutations within the Cas9 amino acid sequence as presented in Figures 3A-3C or Table 13, in conjunction with either a 19mer or 20mer gRNA, as assessed in a cell-based (HEK293) assay correcting a splice site SNP within the SBDS polynucleotide sequence. Figure 8A shows the contrast between on-target editing and bystander editing exerted by an NGC CBE containing Cas9 variant 225 and PpAPOBEC1, and by NGC CBEs 454 and 459 (Table 13), containing PpAPOBEC1 and Cas9 variants 226 and 244, respectively (Figures 3A-3C), when used with a 19mer (guide 88) gRNA. Figure 8B shows the percentage of on-target versus bystander editing achieved by NGC CBEs 454 and 459 containing Cas9 variant 225 and PpAPOBEC1, respectively, with a 20-mer (Guide44) gRNA, and PpAPOBEC1 and Cas9 variants 226 and 244 (Table 13). Figures 8C and 8D show the percentage of on-target and bystander base editing achieved by NGC CBEs containing AmAPOBEC1 cytidine deaminase and Cas9 variants 225, 226, and 244 (Figures 3A-3C), with either a 19-mer (Guide88) or a 20-mer (Guide44) gRNA. Figures 8E and 8F show the percentage of on-target and bystander base edits of NGC CBEs containing PmCDA1 cytidine deaminase and Cas9 variants 225, 453, and 458 (Table 13), together with either a 19-mer (Guide88) or a 20-mer (Guide44) gRNA.Figures 8G and 8H show the percentage of on-target and bystander base edits for NGC CBEs containing RRA3F cytidine deaminase and Cas9 variants 225, 455, and 460 (Table 13) with either a 19-mer (Guide88) or 20-mer (Guide44) gRNA. Figures 8I and 8J show the percentage of on-target and bystander base edits for NGC CBEs containing SsAPOBEC2 cytidine deaminase and Cas9 variants 225, 456, and 461 (Table 13) with either a 19-mer (Guide88) or 20-mer (Guide44) gRNA. In Figures 8A-8J, Cas9 variant 225 (or PV225) is alternatively referred to as "Beam Shuffle." [Figure 8B] Same as above. [Figure 8C] Same as above. [Figure 8D] Same as above. [Figure 8E] Same as above. [Figure 8F] Same as above. [Figure 8G] Same as above. [Figure 8H] Same as above. [Figure 8I] Same as above. [Figure 8J] Same as above. [Figure 9A] Figures 9A-9D show graphs and dot plots of percent editing by NGC CBEs containing PpAPOBEC1 cytidine deaminase polypeptide sequences containing various mutations as described in Example 4, such as the H122A mutation alone and in combination with the amino acid mutations R33A, W90F, K34A, R52A, H121A, and Y120F, using either a 19-mer gRNA (Figure 9A) or a 20-mer gRNA (Figure 9B). The percentage of on-target versus bystander editing was assessed in an in vitro cell-based assay. Figures 9C and 9D present the data from Figures 9A and 9B, respectively, in dot plot format. [Figure 9B] Same as above. [Figure 9C] Same as above. [Figure 9D] Same as above. [Figure 10-1] Figure 10 presents a table illustrating the mutations and combinations of mutations made in the SpCas9 protein to create SpCas9 variants with the indicated combinations of mutations, including "NRCH" mutations as described in S. Miller et al., April 2020, "Continuous evolution of SpCas9 variants compatible with non-G PAMs," Nature Biotechnology, 38(4):471-481 (published online February 10, 2020, doi: 10.1038 / s41587-020-0412-8), the contents of which are incorporated herein by reference in their entirety. Combinations of NRCH mutations (amino acid substitutions) were included in several different SpCas9 variants to determine which would produce SpCas9 variant components of the NGC CBE for use in correcting a splice site SNP in the SBDS gene associated with SDS with enhanced on-target versus bystander base editing (Example 6). In Figure 10, darker shaded amino acids reflect amino acid substitutions in the Cas9 (SpCas9) amino acid sequence compared to the sequence of the wild-type, unmutated Cas9 (SpCas9) protein, and lighter shaded amino acids reflect amino acid residues in the wild-type, unmutated Cas9 (SpCas9) protein. [Figure 10-2] Same as above. [Figure 11A]11A and 11B show graphs illustrating the percent editing by an NGC CBE containing a cytidine deaminase (e.g., PpAPOBEC1) and an SpCas9 variant containing one or more NRCH mutations as defined in FIG. 10 and Example 5, used in conjunction with either a 19-mer or a 20-mer gRNA, in a cell-based assay evaluating the on-target and bystander editing efficiencies of these CBEs to correct a splice site SNP in the SBDS gene associated with SDS. NGC CBEs 468 and 469 (FIG. 10) showed high levels of on-target versus off-target base editing when used in conjunction with either a 19-mer or a 20-mer gRNA. [Figure 11B] Same as above. [Figure 12A] 12A-12C show graphs illustrating the results of in vitro cell-based assays performed to assess the base editing efficiency of NGC CBEs encoded by the mRNAs described in Example 6 in conjunction with gRNAs of different lengths (17mer, 18mer, 19mer, 20mer, or 21mer), and the percentage of on-target edits versus bystander edits. As observed, mRNA342 with 18mer and 20mer gRNAs has the fewest C-to-A or C-to-G transitions compared to mRNA340 or mRNA341. [Figure 12B] Same as above. [Figure 12C] Same as above. DETAILED DESCRIPTION OF THE INVENTION
[0146] The present invention features compositions and methods that use programmable nucleobase editors to edit pathogenic genetic mutations that cause aberrant splicing within a gene, allowing transcription and achieving a therapeutic effect. In some embodiments, editing comprises converting a stop codon to a codon that allows transcription. In some embodiments, editing comprises providing or correcting a splice acceptor site or splice donor site, or providing an alternative splice acceptor site or splice donor site. In some embodiments, one or more mutations resulting from aberrant splicing are corrected.
[0147] The present invention is based, at least in part, on a strategy to edit pathogenic mutations (e.g., mutations resulting from gene conversion) in genes associated with Shwachman-Diamond syndrome (SDS) using adenosine base editors or cytidine base editors (ABEs, CBEs). Accordingly, the present invention provides base editor systems comprising ABEs or CBEs that are useful for treating or preventing SDS.
[0148] Shwachman-Diamond Syndrome (SDS) Shwachman-Diamond syndrome (SDS) is an autosomal recessive disorder. Approximately 90% of patients who meet the clinical diagnostic criteria for SDS have a mutation in the Shwachman-Bodian-Diamond syndrome (SBDS) gene. The carrier frequency of this mutation is estimated to be approximately 1 in 110. This highly conserved gene contains five exons, encompassing 7.9 kb, and maps to the centromeric region of chromosome 7, 7q11. The SBDS gene encodes a novel 250-amino acid protein lacking homology to known protein functional domains. The adjacent pseudogene, SBDSP, shares 97% homology with SBDS but contains deletions and nucleotide changes that prevent the production of a functional protein. Approximately 75% of SDS patients have mutations resulting from gene conversion events involving this pseudogene. Gene conversion occurs when recombination occurs between homologous (paralogous) sequences located at different genomic loci. The existence of the SBDS pseudogene (also called SBDSP) may result from past gene duplication. SBDS mRNA and protein are widely expressed throughout human tissues at both the mRNA and protein levels. The early truncating SBDS mutation 183TA>CT is common among patients with SDS, but no patients homozygous for this mutation have been identified, suggesting that complete loss of SBDS expression may be lethal in human patients.
[0149] Common sequence alterations associated with SDS include a TA→CT dinucleotide change at positions 183–184 or an 8-bp deletion at the end of exon 2. Analysis of SBDS genomic sequences confirmed the presence of the 183–184 TA→CT alteration and identified the 258+2T→C alteration in individuals with SDS who express the deleted transcript. The 258+2T→C alteration is predicted to disrupt a donor splice site in intron 2, and the 8-bp deletion is consistent with the use of an upstream cryptic splice donor site at positions 251–252. The 183–184 TA→CT dinucleotide change introduces an in-frame stop codon (K62X), and the 258+2T→C and resulting 8-bp deletion cause a frameshift, resulting in premature truncation of the encoded protein (84Cfs3).
[0150] The present invention provides compositions and methods that allow transcription of polynucleotides with one or more alterations (e.g., gene conversions) that result in aberrant splicing, thereby resulting in expression of a functional SBDS protein (e.g., a protein with sufficient activity to mitigate the effects of SBDS gene conversion). In specific embodiments, the present invention provides for introducing an alteration within the SBDS gene, comprising 183-184TA→CT, that converts the TAA stop to TGG, encoding Trp, to allow transcription. In other embodiments, the present invention introduces an alteration in the polynucleotide sequence that introduces a splice donor or effector site that allows splicing of a polynucleotide encoding a biologically active protein. In some embodiments, the present invention corrects a site within exon 2 of the SBDS gene (e.g., by editing a cytosine at nucleotide position 1495, as shown in Figure 2B).
[0151] Nucleobase Editor Disclosed herein is a base editor or nucleobase editor for editing, modifying, or altering a target nucleotide sequence of a polynucleotide. Described herein is a nucleobase editor or base editor comprising a polynucleotide-programmable nucleotide-binding domain and a nucleobase-editing domain (e.g., adenosine deaminase, cytidine deaminase). When linked to a bound guide polynucleotide (e.g., gRNA), the polynucleotide-programmable nucleotide-binding domain specifically binds to the target polynucleotide sequence (i.e., via complementary base pairing between the bases of the bound guide nucleic acid and the bases of the target polynucleotide sequence), thereby localizing the base editor to target the nucleic acid sequence targeted for editing. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid.
[0152] Polynucleotide Programmable Nucleotide Binding Domains It should be understood that polynucleotide programmable nucleotide binding domain can also comprise nucleic acid programmable protein that binds RNA.For example, polynucleotide programmable nucleotide binding domain can be associated with nucleic acid that guides polynucleotide programmable nucleotide binding domain to RNA.Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, although they are not specifically listed in the present disclosure.
[0153] The polynucleotide-programmable nucleotide-binding domain of a base editor may itself comprise one or more domains. For example, the polynucleotide-programmable nucleotide-binding domain may comprise one or more nuclease domains. In some embodiments, the nuclease domain of a polynucleotide-programmable nucleotide-binding domain may comprise an endonuclease or exonuclease. As used herein, the term "exonuclease" refers to a protein or polypeptide capable of digesting nucleic acids (e.g., RNA or DNA) from free ends, and the term "endonuclease" refers to a protein or polypeptide capable of catalyzing (e.g., cleaving) an internal region of a nucleic acid (e.g., DNA or RNA). In some embodiments, an endonuclease can cleave one strand of a double-stranded nucleic acid. In some embodiments, an endonuclease can cleave both strands of a double-stranded nucleic acid molecule. In some embodiments, the polynucleotide-programmable nucleotide-binding domain may be a deoxyribonuclease. In some embodiments, the polynucleotide-programmable nucleotide-binding domain may be a ribonuclease.
[0154] In some embodiments, the nuclease domain of a polynucleotide-programmable nucleotide-binding domain can cleave zero, one, or two strands of a target polynucleotide. In some cases, the polynucleotide-programmable nucleotide-binding domain can include a nickase domain. As used herein, the term "nickase" refers to a polynucleotide-programmable nucleotide-binding domain that includes a nuclease domain that can cleave only one of the two strands in a double-stranded nucleic acid molecule (e.g., DNA). In some embodiments, a nickase can be derived from a fully catalytically active (e.g., native) form of a polynucleotide-programmable nucleotide-binding domain by introducing one or more mutations into the active polynucleotide-programmable nucleotide-binding domain. For example, if the polynucleotide-programmable nucleotide-binding domain includes a nickase domain derived from Cas9, the Cas9-derived nickase domain can include a D10A mutation and a histidine at position 840. In such a case, residue H840 retains catalytic activity, thereby cleaving one strand of a nucleic acid duplex. In another example, a Cas9-derived nickase domain can include an H840A mutation, while the amino acid residue at position 10 remains D. In some embodiments, a nickase can be derived from a fully catalytically active (e.g., native) form of a polynucleotide-programmable nucleotide-binding domain by removing all or part of a nuclease domain that is not required for nickase activity. For example, if a polynucleotide-programmable nucleotide-binding domain includes a nickase domain from Cas9, the Cas9-derived nickase domain can include a deletion of all or part of the RuvC domain or the HNH domain.
[0155] The amino acid sequence of an exemplary catalytically active Cas9 is as follows:
[0156] Thus, a base editor comprising a polynucleotide-programmable nucleotide-binding domain comprising a nickase domain can generate a single-stranded DNA break (nick) in a specific polynucleotide target sequence (e.g., determined by the complementary sequence of a bound guide nucleic acid). In some embodiments, the strand of a nucleic acid duplex target polynucleotide sequence cleaved by a base editor comprising a nickase domain (e.g., a Cas9-derived nickase domain) is the strand not edited by the base editor (i.e., the strand cleaved by the base editor is the opposite strand containing the base to be edited). In other embodiments, a base editor comprising a nickase domain (e.g., a Cas9-derived nickase domain) can cleave the strand of a DNA molecule that is targeted for editing. In such cases, the non-targeted strand is not cleaved.
[0157] Also provided herein are base editors comprising a catalytically inactive (i.e., unable to cleave a target polynucleotide sequence) polynucleotide-programmable nucleotide binding domain. The terms "catalytically inactive" and "nuclease-inactive" are used interchangeably herein to refer to a polynucleotide-programmable nucleotide binding domain having one or more mutations and / or deletions that result in an inability to cleave a strand of nucleic acid. In some embodiments, a catalytically inactive polynucleotide-programmable nucleotide binding domain base editor may lack nuclease activity as a result of specific point mutations within one or more nuclease domains. For example, in the case of a base editor comprising a Cas9 domain, Cas9 may contain both the D10A and H840A mutations. Such mutations inactivate both nuclease domains, thereby resulting in a lack of nuclease activity. In other embodiments, a catalytically inactive polynucleotide-programmable nucleotide binding domain may contain one or more deletions of all or part of a catalytic domain (e.g., the RuvC1 and / or HNH domain). In yet another embodiment, the catalytically inactive polynucleotide programmable nucleotide binding domain comprises a point mutation (eg, D10A or H840A) as well as a deletion of all or part of the nuclease domain.
[0158] Also contemplated herein are mutations that can generate catalytically inactive polynucleotide-programmable nucleotide-binding domains from previously functional versions of the polynucleotide-programmable nucleotide-binding domain. For example, in the case of catalytically inactive Cas9 ("dCas9"), variants are provided that have mutations other than D10A and H840A that result in nuclease-inactivated Cas9. Such mutations include, for example, other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions within the HNH nuclease subdomain and / or RuvC1 subdomain). Additional suitable nuclease-inactive dCas9 domains will be apparent to those skilled in the art based on this disclosure and knowledge in the art and are within the scope of this disclosure. Such additional exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, the D10A / H840A mutant domain, the D10A / D839A / H840A mutant domain, and the D10A / D839A / H840A / N863A mutant domain (see, e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013; 31(9): 833-838, the entire contents of which are incorporated herein by reference).
[0159] Non-limiting examples of polynucleotide-programmable nucleotide binding domains that can be incorporated into base editors include CRISPR protein-derived domains, restriction nucleases, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases (ZFNs). In some cases, the base editor comprises a polynucleotide-programmable nucleotide binding domain comprising a natural or modified protein or a portion thereof that can bind to a nucleic acid sequence via a bound guide nucleic acid during CRISPR (i.e., clustered regularly interspaced short palindromic repeats)-mediated modification of the nucleic acid. Such proteins are referred to herein as "CRISPR proteins." Thus, disclosed herein are base editors that comprise a polynucleotide-programmable nucleotide binding domain comprising all or a portion of a CRISPR protein (i.e., a base editor comprising all or a portion of a CRISPR protein domain, also referred to as the "CRISPR protein-derived domain" of the base editor). The CRISPR protein-derived domain incorporated into the base editor can be modified compared to the wild-type or natural version of the CRISPR protein. For example, as described below, a domain derived from a CRISPR protein can contain one or more mutations, insertions, deletions, rearrangements, and / or recombinations compared to a wild-type or naturally occurring version of the CRISPR protein.
[0160] CRISPR is an adaptive immune system that protects against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain a spacer, a sequence complementary to the preceding mobile element, and a target inlet nucleic acid. The CRISPR cluster is processed into transcribed CRISPR RNA (crRNA). In type II CRISPR systems, correct processing of the crRNA precursor requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3-assisted processing of the crRNA precursor. Subsequently, the Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand that is not complementary to the crRNA is first endonucleolytically cleaved and then exonucleolytically trimmed 3'-5'. In nature, DNA binding and cleavage typically require both protein and RNA. However, single guide RNA ("sgRNA" or simply "gNRA") can be engineered to incorporate both aspects of crRNA and tracrRNA into a single RNA species. For example, see Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif (PAM or protospacer adjacent motif) within the CRISPR repeat sequence to help distinguish self from non-self.
[0161] In some embodiments, the methods described herein can utilize engineered Cas proteins. Guide RNAs (gRNAs) are short synthetic RNAs composed of a scaffold sequence required for Cas binding and a user-defined spacer of approximately 20 nucleotides that defines the genomic target to be modified. Thus, one skilled in the art can alter the genome or polynucleotide target of a Cas protein by altering the target sequence present in the gRNA. The specificity of a Cas protein is determined in part by how specific the gRNA targeting sequence is for a genomic polynucleotide target sequence relative to the rest of the genome.
[0162] In some embodiments, the gRNA scaffold sequence is: GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU.
[0163] In one embodiment, the RNA scaffold comprises a stem loop. In one embodiment, the RNA scaffold comprises the following nucleic acid sequence: GUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG.
[0164] In one embodiment, the RNA scaffold comprises the following nucleic acid sequence: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU.
[0165] In one embodiment, the sgRNA scaffold polynucleotide sequence for S. pyrogenes is as follows: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC.
[0166] In one embodiment, the S. aureus sgRNA scaffold polynucleotide sequence is as follows: GUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGCGAGA.
[0167] In one embodiment, the sgRNA scaffold for BhCas12b has the following polynucleotide sequence: GUUCUGTCUUUUGGUCAGGACAACCGUCUAGCUAUAAGUGCUGCAGGGUGUGAGAAACUCCUAUUGCUGGACGAUGUCUCUUACGAGGCAUUAGCAC.
[0168] In one embodiment, the sgRNA scaffold of BvCas12b has the following polynucleotide sequence: GACCUAUAGGGUCAAUGAAUCUGUGCGUGUGCCAUAAGUAAUUAAAAAUUACCCACCACAGGAGCACCUGAAAACAGGUGCUUGGCAC.
[0169] In some embodiments, the domain derived from a CRISPR protein incorporated into the base editor is an endonuclease (e.g., a deoxyribonuclease or ribonuclease) capable of binding to a target polynucleotide when linked to a bound guide nucleic acid. In some embodiments, the domain derived from a CRISPR protein incorporated into the base editor is a nickase capable of binding to a target polynucleotide when linked to a bound guide nucleic acid. In some embodiments, the domain derived from a CRISPR protein incorporated into the base editor is a catalytically inactive domain capable of binding to a target polynucleotide when linked to a bound guide nucleic acid. In some embodiments, the target polynucleotide bound by the CRISPR protein-derived domain of the base editor is DNA. In some embodiments, the target polynucleotide bound by the CRISPR protein-derived domain of the base editor is RNA.
[0170] Cas proteins that can be used herein include class 1 and class 2 Cas proteins. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, and Csb1. , Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i, CARF, DinG, homologs thereof, or modified versions thereof. Unmodified CRISPR enzymes can have DNA cleavage activity, such as Cas9, which has two functional endonuclease domains: RuvC and HNH.CRISPR enzymes can guide one or both cleavages of target sequences, for example, within the target sequence and / or within the complementary part of the target sequence.For example, CRISPR enzymes can guide one or both cleavages within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500 or more base pairs from the first or last nucleotide of the target sequence.
[0171] A vector encoding a CRISPR enzyme can be used that is mutated relative to the corresponding wild-type enzyme so that the mutant CRISPR enzyme loses the ability to cleave one or both strands of a target polynucleotide containing a target sequence. Cas9 can refer to a polypeptide that has at least 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% or at least about 50%, about 60%, about 70%, about 80%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100% sequence identity and / or sequence homology to a wild-type exemplary Cas9 polypeptide (e.g., Cas9 from S. pyrogenes). Cas9 can refer to a polypeptide with up to 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or homology to a wild-type exemplary Cas9 polypeptide (e.g., from S. pyrogenes). Cas9 can refer to wild-type or modified Cas9 proteins, which can include amino acid alterations, such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof.
[0172] In some embodiments, the base editor CRISPR protein-derived domains are derived from Corynebacterium ulcerans (NCBI Reference Nos.: NC_015683.1, NC_017317.1); Corynebacterium diphtheriae (NCBI Reference Nos.: NC_016782.1, NC_016786.1); Spiroplasma silphidicola (NCBI Reference No.: NC_021284.1); Prevotella intermedia (NCBI Reference No.: NC_017861.1); Spiroplasma taiwanense (NCBI Reference No.: NC_021846.1); Streptococcus iniae (NCBI Reference No.: NC_021314.1). 1); Belliera baltica (NCBI Reference Number: NC_018010.1); Cycloflexus torchis (NCBI Reference Number: NC_018721.1); Streptococcus thermophilus (NCBI Reference Number: YP_820832.1); Listeria innocua (NCBI Reference Number: NP_472073.1); Campylobacter jejuni (NCBI Reference Number: YP_002344900.1); Neisseria meningitidis (NCBI Reference Number: YP_002342100.1), Streptococcus pyrogenes, or Staphylococcus aureus.
[0173] Nucleobase editor Cas9 domain The sequence and structure of Cas9 nuclease are well known to those of skill in the art (see, e.g., "Complete genome sequence of an M1 strain of Streptococcus pyogenes," Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III," Deltcheva E., Chylinski K., Sharma CM, ... See Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607(2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012)). Cas9 orthologs have been described in various species, including, but not limited to, S. pyrogenes and S. thermophilus.Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on the present disclosure, including Cas9 sequences and loci from the organisms disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737, the entire contents of which are incorporated herein by reference.
[0174] In some aspects, the nucleic acid programmable DNA binding protein (napDNAbp) is a Cas9 domain. Non-limiting exemplary Cas9 domains are provided herein. The Cas9 domain may be a nuclease-active Cas9 domain, a nuclease-inactive Cas9 domain, or a Cas9 nickase. In some embodiments, the Cas9 domain is a nuclease-active domain. For example, the Cas9 domain may be a Cas9 domain that cleaves both strands of a double-stranded nucleic acid (e.g., both strands of a double-stranded DNA molecule). In some embodiments, the Cas9 domain comprises any one of the amino acid sequences defined herein. In some embodiments, the Cas9 domain comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the amino acid sequences defined herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any one of the amino acid sequences set forth herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 contiguous amino acid residues identical to any one of the amino acid sequences set forth herein.
[0175] In some embodiments, proteins comprising a fragment of Cas9 are provided. For example, in some embodiments, the protein comprises one of the following two Cas9 domains: (1) the gRNA binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, proteins comprising Cas9 or a fragment thereof are referred to as "Cas9 variants." Cas9 variants share homology with Cas9 or a fragment thereof. For example, the Cas9 variant is at least about 70% homologous to wild-type Cas9, at least about 80% homologous, at least about 90% homologous, at least about 95% homologous, at least about 96% homologous, at least about 97% homologous, at least about 98% homologous, at least about 99% homologous, at least about 99.5% homologous, or at least about 99.9% homologous. In some embodiments, the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type Cas9. In some embodiments, the Cas9 variant has a fragment of Cas9 (e.g., the gRNA binding domain or the DNA cleavage domain) such that the fragment is at least about 70% homologous, at least about 80% homologous, at least about 90% homologous, at least about 95% homologous, at least about 96% homologous, at least about 97% homologous, at least about 98% homologous, at least about 99% homologous, at least about 99.5% homologous, or at least about 99.9% homologous to the corresponding fragment of wild-type Cas9.In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% amino acids in length to the corresponding wild-type Cas9. In some embodiments, the fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids in length.
[0176] In some embodiments, a Cas9 fusion protein as provided herein comprises the full-length amino acid sequence of a Cas9 protein, e.g., one of the Cas9 sequences provided herein. However, in other embodiments, a fusion protein as provided herein does not comprise the full-length Cas9 sequence, but only one or more fragments thereof. Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and additional suitable sequences of Cas9 domains and fragments will be apparent to those skilled in the art.
[0177] The Cas9 protein may be associated with a guide RNA, which guides the Cas9 protein to a specific DNA sequence that is complementary to the guide RNA. In some embodiments, the polynucleotide programmable nucleotide binding domain is a Cas9 domain, such as a nuclease-active Cas9, a Cas9 nickase (nCas9), or a nuclease-inactive Cas9 (dCas9). Examples of nucleic acid programmable DNA binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpf1, Cas12b / C2C1, and Cas12c / C2C3.
[0178] In some embodiments, the wild-type Cas9 corresponds to Cas9 from Streptococcus pyrogenes (NCBI Reference Sequence: NC_017053.1, nucleotide and amino acid sequence as below). JPEG0007781742000003.jpg117168 (single underline: HNH domain; double underline: RuvC domain)
[0179] In some embodiments, the wild-type Cas9 corresponds to or comprises the following nucleotide and / or amino acid sequence: JPEG0007781742000004.jpg113168 (single underline: HNH domain; double underline: RuvC domain)
[0180] In some embodiments, the wild-type Cas9 corresponds to Cas9 from Streptococcus pyrogenes (NCBI Reference Sequence: NC_002737.2 (nucleotide sequence as shown below); and Uniprot Reference Sequence: Q99ZW2 (amino acid sequence as shown below): JPEG0007781742000005.jpg115168 (single underline: HNH domain; double underline: RuvC domain)
[0181] In some embodiments, Cas9 is used in Corynebacterium ulcerans (NCBI Reference Numbers: NC_015683.1, NC_017317.1); Corynebacterium diphtheriae (NCBI Reference Numbers: NC_016782.1, NC_016786.1); Spiroplasma silphidicola (NCBI Reference Number: NC_021284.1); Prevotella intermedia (NCBI Reference Number: NC_017861.1); Spiroplasma taiwanense (NCBI Reference Number: NC_021846.1); Streptococcus iniae (NCBI Reference Number: NC_021314.1) refers to Cas9 derived from Bacillus baltica (NCBI Reference Number: NC_018010.1); Cycloflexus torchis (NCBI Reference Number: NC_018721.1); Streptococcus thermophilus (NCBI Reference Number: YP_820832.1), Listeria innocua (NCBI Reference Number: NP_472073.1), Campylobacter jejuni (NCBI Reference Number: YP_002344900.1), or Neisseria meningitidis (NCBI Reference Number: YP_002342100.1), or to Cas9 derived from any other organism.
[0182] It should be appreciated that additional Cas9 proteins (e.g., nuclease-free Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9), including variants and homologs thereof, are within the scope of this disclosure. Exemplary Cas9 proteins include, but are not limited to, those shown below. In some embodiments, the Cas9 protein is nuclease-free Cas9 (dCas9). In some embodiments, the Cas9 protein is Cas9 nickase (nCas9). In some embodiments, the Cas9 protein is nuclease-active Cas9.
[0183] In some embodiments, the Cas9 domain is a nuclease-inactive Cas9 domain (dCas9). For example, the dCas9 domain can bind to a double-stranded nucleic acid molecule (e.g., via a gRNA molecule) without cleaving either of the double-stranded nucleic acid molecules. In some embodiments, the nuclease-inactive dCas9 domain comprises a D10X mutation and an H840X mutation in the amino acid sequence defined herein, or a corresponding mutation in any of the amino acid sequences shown herein, where X is any amino acid change. In some embodiments, the nuclease-inactive dCas9 domain comprises a D10A mutation and an H840A mutation in the amino acid sequence defined herein, or a corresponding mutation in any of the amino acid sequences shown herein. In one example, the nuclease-inactive Cas9 domain comprises a defined amino acid sequence in the cloning vector pPlatTET-gRNA2 (accession number BAV54124).
[0184] (See, e.g., Qi et al., "Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression." Cell. 2013; 152(5):1173-83, the entire contents of which are incorporated herein by reference.)
[0185] In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is a nickase, also referred to as a "nCas9" protein ("nickase" Cas9). A nuclease-inactivated Cas9 protein may also be referred to interchangeably as a "dCas9" protein (nuclease-inactive" Cas9) or catalytically inactive Cas9. Methods for generating Cas9 proteins (or fragments thereof) with inactive DNA cleavage domains are known (see, e.g., Jinek et al., Science. 337:816-821 (2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression" (2013) Cell. 28;152(5):1173-83, the entire contents of each of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to contain two subdomains: the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can block the nuclease activity of Cas9. For example, mutations D10A and H840A completely inactivate the nuclease activity of S. pyrogenes Cas9 (Jinek et al., Science. 337:816-821(2012); Qi et al., Cell. 28;152(5):1173-83 (2013)).
[0186] In some embodiments, the dCas9 domain comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the dCas9 domains set forth herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more or more mutations compared to any one of the amino acid sequences set forth herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 contiguous amino acid residues identical to any one of the amino acid sequences set forth herein.
[0187] In some embodiments, the dCas9 corresponds to or comprises a Cas9 amino acid sequence with one or more mutations that partially or completely inactivate the Cas9 nuclease activity. For example, in some embodiments, the dCas9 domain comprises D10A and H840A mutations, or corresponding mutations in another Cas9.
[0188] In some embodiments, the dCas9 comprises the amino acid sequence of dCas9 (D10A and H840A): JPEG0007781742000006.jpg116168 (single underline: HNH domain; double underline: RuvC domain)
[0189] In some embodiments, the Cas9 domain comprises a D10A mutation and residue 840 remains a histidine in the amino acid sequence shown above, or at any corresponding position in any of the amino acid sequences shown herein.
[0190] In other embodiments, dCas9 variants with mutations other than D10A and H840A are provided, which, for example, result in nuclease-inactivated Cas9 (dCas9). Such mutations include, for example, other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions within the HNH nuclease subdomain and / or RuvC1 subdomain). In some embodiments, variants or homologs of dCas9 are provided that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical. In some embodiments, variants of dCas9 are provided that have shorter or longer amino acid sequences of about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids, or more.
[0191] In some embodiments, the Cas9 domain is a Cas9 nickase. The Cas9 nickase can be a Cas9 protein that can cleave only one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 nickase cleaves the target strand of the double-stranded nucleic acid molecule, meaning that the Cas9 nickase cleaves the strand that is base-paired (complementary to) the gRNA (e.g., sgRNA) that binds to the Cas9. In some embodiments, the Cas9 nickase comprises a D10A mutation and a histidine at position 840. In some embodiments, the Cas9 nickase cleaves the non-target, non-base-edited strand of the double-stranded nucleic acid molecule, meaning that the Cas9 nickase cleaves the strand that is not base-paired to the gRNA (e.g., sgRNA) that binds to the Cas9. In some embodiments, the Cas9 nickase comprises a H840A mutation and an aspartic acid residue at position 10, or a corresponding mutation. In some embodiments, the Cas9 nickase comprises an amino acid sequence at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the Cas9 nickases provided herein. Additional suitable Cas9 nickases will be apparent to those of skill in the art based on this disclosure and knowledge in the art, and are within the scope of this disclosure.
[0192] The amino acid sequence of an exemplary catalytic Cas9 (nCas9) is as follows:
[0193] In some embodiments, Cas9 refers to Cas9 from archaea (e.g., nanoarchaea), which constitute the domain and kingdom of unicellular prokaryotic microorganisms. In some embodiments, the programmable nucleotide-binding protein may be a CasX or CasY protein, as described, for example, in Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res. 2017 Feb 21. doi: 10.1038 / cr.2017.21, the entire contents of which are incorporated herein by reference. Using genome sequencing metagenomics, multiple CRISPR-Cas systems have been identified, including the first reported Cas9 in the Archaea domain of life. This diverse Cas9 protein was found as part of an active CRISPR-Cas system in the little-studied nanoarchaea. In bacteria, two previously unknown systems, CRISPR-CasX and CRISPR-CasY, have been discovered, which are the most compact systems discovered to date. In some embodiments, in the base editor systems described herein, Cas9 is replaced by CasX or a variant of CasX. In some embodiments, in the base editor systems described herein, Cas9 is replaced by CasY or a variant of CasY. It should be understood that other RNA-guided DNA-binding proteins may be used as nucleic acid programmable DNA-binding proteins (napDNAbp) and are within the scope of the present disclosure.
[0194] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) of any fusion protein provided herein may be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the napDNAbp is a CasY protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or comfortably 99.5% identical to a naturally occurring CasX or CasY protein. In some embodiments, the programmable nucleotide binding protein is a naturally occurring CasX or CasY protein. In some embodiments, the programmable nucleotide-binding protein comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or comfortably 99.5% identical to any CasX or CasY protein described herein. It should be recognized that CasX and CasY from other bacterial species may also be used in accordance with the present disclosure.
[0195] The amino acid sequence of an exemplary CasX ((uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) tr|F0NN87|F0NN87_SULIH CRISPR-associated CasX protein OS = Sulfolobus islandicus (strain HVE10 / 4) GN = SiH_0402 PE=4 SV=1) is as follows: MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKCEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAAAGWVLTRLGKAKV SEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAVGQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG.
[0196] The amino acid sequence of an exemplary CasX (>tr|F0NH53|F0NH53_SULIR CRISPR-associated protein, CasX OS = Sulfolobus islandicus (strain REY15A) GN = SiRe_0771 PE = 4 SV = 1) is as follows: MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKCEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAAAGWVLTRLGKAKV SEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAVGQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG.
[0197] Deltaproteobacteria CasX MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMTDDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAILQVYWQEFKDDHVGLMCKFAQPASKKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFVYKLEQVSEKGKAYTNYFGRCNVAEHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESTHPVKPLAQIAGNRYASGPVGKALSDACMGTIASFLSKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDfAYNEVIARVRMWVNLLWQKLKLSRDDAKPLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENPKKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEMDEKEFYACEIQLQ KWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFTDGTDIKKSGKWQGLLYGGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKIGRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQAAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAKLAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKELSAELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSYKSGKQPFVGAWQAFYKRRLKEVWKPNA
[0198] The amino acid sequence of an exemplary CasY ((ncbi.nlm.nih.gov / protein / APG80656.1) >APG80656.1 CRISPR-associated protein CasY [uncultured Parkobacteria group bacteria]) is as follows:
[0199] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is the single effector of the microbial CRISPR-Cas system. The single effector of the microbial CRISPR-Cas system includes, but is not limited to, Cas9, Cpf1, Cas12b / C2c1, and Cas12c / C2c3. Typically, the microbial CRISPR-Cas system is divided into class 1 and class 2 systems. Class 1 systems have a multi-subunit effector complex, while class 2 systems have a single protein effector. For example, Cas9 and Cpf1 are class 2 effectors. In addition to Cas9 and Cpf1, three distinct Class 2 CRISPR-Cas systems (Cas12b / C2c1, and Cas12c / C2c3) have been described by Shmakov et al., "Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems," Mol. Cell, 2015 Nov. 5; 60(3): 385-397, the entire contents of which are incorporated herein by reference. The effectors of two systems, Cas12b / C2c1 and Cas12c / C2c3, contain a RuvC-like endonuclease domain related to Cpf1. The third system contains an effector with two putative HEPN RNase domains. Unlike Cas12b / C2c1-mediated CRISPR RNA, mature CRISPR RNA production is tracrRNA-dependent. Cas12b / C2c1 depends on both CRISPR RNA and tracrRNA for DNA cleavage.
[0200] The crystal structure of Alicyclobacillus acidoterrestris Cas12b / C2c1 (AacC2c1) has been reported in complex with a chimeric single-stranded guide RNA (sgRNA). See, e.g., Liu et al., "C2c1-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism," Mol. Cell, 2017 Jan. 19; 65(2):310-322, the entire contents of which are incorporated herein by reference. This crystal structure was reported for Alicyclobacillus acidoterrestris C2c1 bound to target DNA as a tertiary complex. See, e.g., Yang et al., "PAM-dependent Target DNA Recognition and Cleavage by C2C1 CRISPR-Cas endonuclease," Cell, 2016 Dec. 15; 167(7):1814-1828, the entire contents of which are incorporated herein by reference. Catalytically competent conformations of AacC2c1 with both the target and non-target DNA strands are independently positioned and trapped within a single RuvC catalytic pocket, resulting in a staggered seven-nucleotide cut in the target DNA after Cas12b / C2c1-mediated cleavage. Structural comparison of the Cas12b / C2c1 tertiary complex with previously identified Cas9 and Cpf1 counterparts demonstrates the diversity of the mechanisms employed by the CRISPR-Cas9 system.
[0201] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) of any fusion protein provided herein may be a Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a Cas12b / C2c1 protein. In some embodiments, the napDNAbp is a Cas12c / C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or comfortably 99.5% identical to a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or comfortably 99.5% identical to any one of the napDNAbp sequences set forth herein. It should be recognized that Cas12b / C2c1 or Cas12c / C2c3 from other bacterial species may also be used in accordance with the present disclosure.
[0202] The amino acid sequence of Cas12b / C2c1 ((uniprot.org / uniprot / T0D7A2#2)sp|T0D7A2|C2C1_ALIAG CRISPR-associated endonuclease C2c1 OS=Alicyclobacillus acidoterrestris (strain ATCC49025 / DSM3922 / CIP106132 / NCIMB13137 / GD3B) GN=c2c1 PE=1 SV=1) is as follows:
[0203] BhCas12b (Bacillus hisashii) NCBI reference sequence: WP_095142515
[0204] In some embodiments, Cas12b is BvCas12B, which is a variant of BhCas12b and contains the following changes relative to BhCas12B: S893R, K846R, and E837G. BvCas12b (Bacillus V3-13) NCBI reference sequence: WP_101661451.1
[0205] The Cas9 nuclease has two functional endonuclease domains: RuvC and HNH. Cas9 undergoes a conformational change upon target binding, which positions the nuclease domains to cleave opposite strands of the target DNA. The end result of Cas9-mediated DNA cleavage is a double-strand break (DSB) within the target DNA (approximately 3–4 nucleotides upstream of the PAM sequence). This resulting DSB is then repaired by one of two general repair pathways: (1) the efficient but error-prone non-homologous end joining (NHEJ) pathway, or (2) the less efficient but high-fidelity homology-directed repair (HDR) pathway.
[0206] The "efficiency" of non-homologous end joining (NHEJ) and / or homology-directed repair (HDR) can be calculated by any convenient method. For example, in some cases, efficiency can be expressed in terms of the percentage of successful HDR. For example, a surveyor nuclease assay can be used to generate cleavage products, and the ratio of the products to the substrate can be used to calculate the percentage. For example, a surveyor nuclease enzyme can be used, which directly cleaves the DNA containing the newly incorporated restriction sequence as a result of successful HDR. A larger number of cleaved substrates indicates a larger percentage of HDR (larger HDR efficiency). As an illustrative example, the rate (percentage) of HDR can be calculated using the following formula: [(cleavage product) / (substrate+cleavage product)] (e.g., (b+c) / (a+b+c), where "a" is the band intensity of the DNA substrate, and "b" and "c" are the cleavage products).
[0207] In some cases, efficiency can be expressed in terms of the percentage of successful NHEJ. For example, T7 nuclease assay I can be used to generate cleavage products, and the ratio of product to substrate can be used to calculate percentage NHEJ. T7 endonuclease I cleaves mismatched heteroduplex DNA resulting from hybridization of wild-type and mutant DNA strands (NHEJ generates small random insertions or deletions (indels) at the site of the original cleavage). More cleavage indicates a higher percentage of NHEJ (higher NHEJ efficiency). As an illustrative example, the rate (percentage) of NHEJ can be calculated using the following formula: (1-(1-(b+c) / (a+b+c)) 1 / 2 ) × 100, where "a" is the band intensity of the DNA substrate, and "b" and "c" are the cleavage products (Ran et al., Cell. 2013 Sep. 12; 154(6):1380-9; and Ran et al., Nat Protoc. 2013 Nov.; 8(11): 2281-2308).
[0208] The NHEJ repair pathway is the most active repair mechanism and frequently induces small nucleotide insertions or deletions (indels) at DSB sites. The random nature of NHEJ-mediated DSB repair has significant practical implications, as a population of cells expressing Cas9 and a gRNA or guide polynucleotide can result in a range of different mutations. In most cases, NHEJ generates small indels in the target DNA, resulting in amino acid deletions, insertions, or frameshift mutations that result in premature stop codons within the open reading frame (ORF) of the target gene. The ideal end result is a loss-of-function mutation within the target gene.
[0209] While NHEJ-mediated DSB repair often interrupts a gene's open reading frame, homology-directed repair (HDR) can be used to generate specific nucleotide changes, ranging from single nucleotide changes to large insertions such as the addition of fluorophores or tags. To utilize HDR for gene editing, a DNA repair template containing the desired sequence can be delivered into the cell type of interest using a gRNA and Cas9 or Cas9 nickase. The repair template can contain the desired edit sequence as well as additional homologous sequences immediately upstream and downstream of the target (also referred to as left and right homology arms). The length of each homology arm can depend on the size of the change being introduced, with larger insertions requiring longer homology arms. The repair template can be a single-stranded oligonucleotide, a double-stranded oligonucleotide, or a double-stranded DNA plasmid. The efficiency of HDR is generally low (<10% for modified alleles), even in cells expressing Cas9, gRNA, and an exogenous repair template. Because HDR occurs during the S and G2 phases of the cell cycle, the efficiency of HDR can be increased by synchronizing cells. Chemical or genetic inhibition of genes involved in NHEJ can also increase the efficiency of HDR.
[0210] In some embodiments, the Cas9 is a modified Cas9. A given gRNA target sequence may have additional sites of partial homology throughout the genome. These sites are called off-targets and need to be considered when designing the gRNA. In addition to optimizing the gRNA design, the specificity of CRISPR can also be increased through modifications to Cas9. Cas9 generates double-strand breaks (DSBs) through the combined activity of two nuclease domains, RuvC and HNH. Cas9 nickase, a D10A mutant of SpCas9, retains one nuclease domain and generates DNA nicks instead of DSBs. This nickase system can also be combined with HDR-mediated gene editing for specific gene editing.
[0211] In some cases, the Cas9 is a variant Cas9 protein. A variant Cas9 polypeptide has an amino acid sequence that differs by one amino acid (e.g., has a deletion, insertion, substitution, or fusion) compared to the amino acid sequence of a wild-type Cas9 protein. In some examples, the variant Cas9 polypeptide has an amino acid change (e.g., a deletion, insertion, or substitution) that reduces the nuclease activity of the Cas9 polypeptide. For example, in some cases, the variant Cas9 polypeptide has less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nuclease activity of the corresponding wild-type Cas9 protein. In some cases, the variant Cas9 protein does not have substantial nuclease activity. When the subject Cas9 protein is a variant Cas9 protein that does not have substantial nuclease activity, it may be referred to as "dCas9."
[0212] In some cases, the variant Cas9 protein has reduced nuclease activity, e.g., the variant Cas9 protein exhibits less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% of the endonuclease activity of a wild-type Cas9 protein, e.g., a wild-type Cas9 protein.
[0213] In some cases, the variant Cas9 protein can cleave the complementary strand of the guide target sequence, but has a reduced ability to cleave the non-complementary strand of the double-stranded guide target sequence. For example, the variant Cas9 protein has a mutation (amino acid substitution) that reduces the function of the RuvC domain. As a non-limiting example, in some embodiments, the variant Cas9 protein has D10A (aspartic acid to alanine at amino acid position 10), and therefore can cleave the complementary strand of the double-stranded guide target sequence, but has a reduced ability to cleave the non-complementary strand of the double-stranded guide target sequence (as a result, when the variant Cas9 protein cleaves the double-stranded target nucleic acid, it generates a single-strand break (SSB) instead of a double-strand break (DSB)) (see, for example, Jinek et al., Science. 2012 Aug. 17; 337(6096):816-21).
[0214] In some cases, a variant Cas9 protein can cleave the non-complementary strand of a double-stranded guide target sequence, but has a reduced ability to cleave the complementary strand of the guide target sequence. For example, the variant Cas9 protein can have a mutation (amino acid substitution) that reduces the function of the HNH domain (RuvC / HNH / RuvC domain motif). As a non-limiting example, in some embodiments, the variant Cas9 protein has an H840A (histidine to alanine at amino acid position 840) mutation, and therefore can cleave the non-complementary strand of a guide target sequence, but has a reduced ability to cleave the complementary strand of a guide target sequence (as a result, when the variant Cas9 protein cleaves a double-stranded guide target sequence, an SSB is generated instead of a DSB). Such a Cas9 protein has a reduced ability to cleave a guide target sequence (e.g., a single-stranded guide target sequence), but retains the ability to bind to the guide target sequence (e.g., a single-stranded guide target sequence).
[0215] In some cases, the variant Cas9 protein has a reduced ability to cleave both the complementary and non-complementary strands of double-stranded target DNA. As a non-limiting example, in some cases, the variant Cas9 protein carries both the D10A and H840A mutations, and therefore the polypeptide has a reduced ability to cleave both the complementary and non-complementary strands of double-stranded target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA).
[0216] As another non-limiting example, in some cases, a variant Cas9 protein carries the W476A and W1126A mutations, such that the polypeptide has a reduced ability to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA).
[0217] As another non-limiting example, in some cases, a variant Cas9 protein carries P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, such that the polypeptide has a reduced ability to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA).
[0218] As another non-limiting example, in some cases, a variant Cas9 protein carries H840A, W476A, and W1126A mutations, such that the polypeptide has a reduced ability to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some cases, a variant Cas9 protein carries H840A, D10A, W476A, and W1126A mutations, such that the polypeptide has a reduced ability to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, a variant Cas9 has the catalytically active His residue restored to position 840 of the Cas9 HNH domain (A840H).
[0219] As another non-limiting example, in some cases, a variant Cas9 protein carries H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, and therefore, the polypeptide has a reduced ability to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some cases, a variant Cas9 protein carries D10A, H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, and therefore, the polypeptide has a reduced ability to cleave target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA). In some cases, when a variant Cas9 protein carries the W476A and W1126A mutations, or when a variant Cas9 protein carries the P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, the variant Cas9 protein does not efficiently bind to the PAM sequence. Therefore, in some such cases, when such a variant Cas9 protein is used in a binding method, the method does not require a PAM sequence. In other words, in some cases, when such a variant Cas9 protein is used in a binding method, the method may include a guide RNA, but the method can be performed in the absence of a PAM sequence (and therefore, binding specificity is provided by the targeting segment of the guide RNA). Other residues can be mutated to achieve the above effect (i.e., inactivate one or other nuclease portion). By way of non-limiting example, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 can be altered (i.e., substituted). Mutations other than alanine substitutions are also suitable.
[0220] In some embodiments, a variant Cas9 protein with reduced catalytic activity (e.g., when the Cas9 protein has a D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 mutation, e.g., D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A), can still bind to target DNA in a site-specific manner (because it is still guided to the target DNA sequence by the guide RNA), so long as the variant Cas9 protein retains the ability to interact with the guide RNA.
[0221] In some embodiments, the variant Cas protein can be spCas9, spCas9-VRQR, spCas9-VRER, xCas9(sp), saCas9, saCas9-KKH, SpCas9-MQKFRAER, spCas9-MQKSER, spCas9-LRKIQK, or spCas9-LRVSQL.
[0222] In one specific embodiment, a modified SpCas9 is used that contains the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (SpCas9-MQKFRAER) and has altered specificity for PAM5'-NGC-3'.
[0223] An alternative to S. pyrogenes Cas9 is an RNA-guided endonuclease from the Cpf1 family that exhibits cleavage activity in mammalian cells. Prevotella and Francisella 1 CRISPR (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. Cpf1 is an RNA-guided endonuclease of the class II CRISPR / Cas system. This adaptive immune mechanism was discovered in Prevotella and Francisella bacteria. The Cpf1 gene is associated with the CRISPR locus, which encodes an endonuclease that uses guide RNA to locate and cleave viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9, overcoming some of the limitations of the CRISPR / Cas9 system. Unlike Cas9 nuclease, Cpf1-mediated DNA cleavage results in a double-strand break with a short 3' overhang. The staggered cleavage pattern of Cpf1 may open up the possibility of directed gene transfer similar to traditional restriction enzyme cloning, thereby increasing the efficiency of gene editing. Like the above-mentioned Cas9 variants and orthologs, Cpf1 can also expand the number of sites that can be targeted by CRISPR to AT-rich regions or AT-rich genomes that are poor in NGG PAM sites preferred by SpCas9. The Cpf1 locus contains a mixed alpha / beta domain, RuvC-I and the subsequent helical region, RuvC-II, and a zinc finger-like domain. The Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9. Furthermore, Cpf1 does not have an HNH endonuclease domain, and the N-terminus of Cpf1 does not have the alpha-helix recognition lobe of Cas9. The structure of the Cpf1 CRISPR-Cas domain indicates that Cpf1 is functionally unique and falls into the category of a class 2, type V CRISPR system. The Cpf1 locus encodes Cas1, Cas2, and Cas4 proteins that are more similar to type I and type III CRISPR systems than to type II systems.Functional Cpf1 does not require transactivating CRISPR RNA (tracrRNA) and therefore only CRISPR (crRNA) is required. This is beneficial for genome editing because Cpf1 is not only smaller than Cas9 but also has a smaller sgRNA molecule (approximately half the number of nucleotides as Cas9). The Cpf1-crRNA complex cleaves target DNA or RNA by recognizing the protospacer-adjacent motif 5'-YTN-3', in contrast to the G-rich PAM targeted by Cas9. After recognizing the PAM, Cpf1 introduces a sticky-end-like DNA double-strand break with a 4- or 5-nucleotide overhang.
[0224] Some aspects of the present disclosure provide a nucleic acid-programmable DNA-binding protein domain and a deaminase domain. Some aspects of the present disclosure provide a fusion protein comprising a domain that functions as a nucleic acid-programmable DNA-binding protein, which can be used to guide proteins such as base editors to specific nucleic acid (e.g., DNA or RNA) sequences. In specific embodiments, the fusion protein comprises a nucleic acid-programmable DNA-binding protein domain and a deaminase domain. DNA-binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i. An example of a programmable polynucleotide-binding protein with a PAM specificity different from Cas9 is clustered regularly interspaced short palindromic repeats (Cpfl) from Prevotella and Francisella 1. Similar to Cas9, Cpf1 is also a class 2 CRISPR effector. Cpf1 has been shown to mediate robust DNA interference through characteristics distinct from Cas9. Cpf1 is a single RNA-guided endonuclease lacking tracrRNA and utilizing a T-rich protospacer adjacent motif (TTN, TTTN, or YTN). Furthermore, Cpf1 cleaves DNA through staggered DNA double-strand breaks. In addition to the 16 Cpf1 family proteins, two enzymes from Acidaminococcus and Lachnospiraceae have been shown to have efficient genome editing activity in human cells. Cpf1 proteins are known in the art and have been previously described, for example, in Yamano et al., "Crystal structure of Cpf1 in complex with guide RNA and target DNA," Cell (165) 2016, pp. 949-962, the entire contents of which are incorporated herein by reference.
[0225] Also useful in the compositions and methods herein is a nuclease-inactive Cpf1 (dCpf1) variant, which can be used as a guide nucleotide sequence programmable polynucleotide binding protein domain.Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9, but does not have the HNH endonuclease domain, and the N-terminus of Cpf1 does not have the alpha-helix recognition lobe of Cas9.Zetsche et al., Cell, 163, 759-771, 2015 (incorporated herein by reference) showed that the RuvC-like domain of Cpf1 is responsible for the cleavage of both DNA strands, and inactivation of the RuvC-like domain inactivates the nuclease activity of Cpf1. For example, mutations corresponding to D917A, E1006A, or D1255A in Francisella novicida Cpf1 inactivate Cpf1 nuclease activity. In some embodiments, a dCpf1 of the present disclosure contains mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A. It will be understood that any mutation that inactivates the RuvC domain of Cpf1, such as a substitution mutation, deletion, or insertion, may be used in accordance with the present disclosure.
[0226] In some embodiments, the nucleic acid-programmable nucleotide-binding protein of any fusion protein provided herein may be a Cpf1 protein. In some embodiments, the Cpf1 protein is a Cpf1 nickase (nCpf1). In some embodiments, the Cpf1 protein is a nuclease-inactive Cpf1 (dCpf1). In some embodiments, the Cpf1, nCpf1, or dCpf1 comprises an amino acid sequence at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a Cpf1 sequence disclosed herein. In some embodiments, the dCpfl comprises an amino acid sequence at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or comfortably 99.5% identical to a Cpfl sequence disclosed herein, and includes corresponding mutations within D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A. It should be recognized that Cpfl from other bacterial species may also be used in accordance with the present disclosure.
[0227] The amino acid sequence of wild-type Francisella novicida Cpf1 is below, with D917, E1006, and D1255 in bold and underlined. MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKL ALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDKAIKENKGEGYKKIVYKKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI DRGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF E DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDY KNFGDKAAKGKWTIASFFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA D ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN.
[0228] The amino acid sequence of Francisella novicida Cpf1 D917A is as follows (A917, E1006, and D1255 are in bold and underlined): MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKL ALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDKAIKENKGEGYKKIVYKKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI ARGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF E DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDY KNFGDKAAKGKWTIASFFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA D ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN.
[0229] The amino acid sequence of Francisella novicida Cpf1 E1006A is as follows (D917, A1006, and D1255 are in bold and underlined): MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKL ALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDKAIKENKGEGYKKIVYKKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI DRGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF A DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDY KNFGDKAAKGKWTIASFFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA D ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN.
[0230] The amino acid sequence of Francisella novicida Cpf1 D1255A is as follows: (The mutation positions of D917, E1006, and A1255 are bolded and underlined.) MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKL ALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDKAIKENKGEGYKKIVYKKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI DRGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF E DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDY KNFGDKAAKGKWTIASFFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA A ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN.
[0231] The amino acid sequence of Francisella novicida Cpf1 D917A / E1006A is as follows (A917, A1006, and D1255 are in bold and underlined): MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKL ALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDKAIKENKGEGYKKIVYKKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI ARGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF A DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDY KNFGDKAAKGKWTIASFFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA D ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN.
[0232] The amino acid sequence of Francisella novicida Cpf1 D917A / D1255A is as follows (A917, E1006, and A1255 are in bold and underlined): MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKL ALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDKAIKENKGEGYKKIVYKKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI ARGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF E DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDY KNFGDKAAKGKWTIASFFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA A ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN.
[0233] The amino acid sequence of Francisella novicida Cpf1 E1006A / D1255A is as follows (D917, A1006, and A1255 are in bold and underlined): MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKL ALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDKAIKENKGEGYKKIVYKKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI DRGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF A DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDY KNFGDKAAKGKWTIASFFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA A ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN.
[0234] The amino acid sequence of Francisella novicida Cpf1 D917 / E1006A / D1255A is as follows (A917, A1006, and A1255 are in bold and underlined): MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKL ALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDKAIKENKGEGYKKIVYKKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFFHCPITINFKSSGANKFNDEINLLLKEKANDVHILSI ARGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAIEKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF A DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVFKDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLDKGYFEFSFDY KNFGDKAAKGKWTIASFFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESDKKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA A ANGAYHIGLKGLMLLGRIKNNQEGKKLNLVIKNEEYFEFVQNRNN.
[0235] In some embodiments, one of the Cas9 domains present in the fusion protein may be replaced with a guide nucleotide sequence-programmable DNA-binding protein domain without the need for a PAM sequence.
[0236] In some embodiments, the Cas domain is a Cas9 domain from Staphylococcus aureus (SaCas9). In some embodiments, the SaCas9 domain is a nuclease-active SaCas9, a nuclease-inactive SaCas9 (SaCas9d), or a SaCas9 nickase (SaCas9n). In some embodiments, the SaCas9 domain comprises an N579A mutation or a corresponding mutation in any of the amino acid sequences set forth herein.
[0237] In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to a nucleic acid sequence having a non-canonical PAM. In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to a nucleic acid sequence having an NNGRRT or NNGRRT PAM sequence. In some embodiments, the SaCas9 domain comprises one or more of the following mutations: E781X, N967X, R1014X, or a corresponding mutation in any amino acid sequence set forth herein, where X is any amino acid. In some embodiments, the SaCas9 domain comprises one or more of the following mutations: E781K, N967K, and R1014H, or a corresponding mutation in any amino acid sequence set forth herein. In some embodiments, the SaCas9 domain comprises the following mutations: E781K, N967K, or R1014H, or a corresponding mutation in any amino acid sequence set forth herein.
[0238] The amino acid sequence of an exemplary SaCas9 is as follows: MKRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHSELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKDGEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKEWYEMLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFK QKKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAINLILDELWHTNDNQIAIFNRKLVPKKVDLSQQK EIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVSFDNSFNNKVLVKQEE NSKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDFINRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFI FKEWKKLDKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLIVNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGD In this sequence, the underlined and bolded residue N579 may be mutated (e.g., to A579) to produce a SaCas9 nickase.
[0239] The amino acid sequence of an exemplary SaCas9n is as follows: KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHSELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKDGEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYEGPGEGSPFGWKDIKEWYEMLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQ KKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAINLILDELWHTNDNQIAIFNRKLVPKKVDLSQQK EIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVSFDNSFNNKVLVKQEE ASKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDFINRNLVDTRYATRGLMNLLRSYFRVNLDVKVKSINGGFTSFLRRKWKFKKERNKGY KHHAEDALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLIVNNLNGLYDKDNDKLK KLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIKYYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIK KENYYEVNSKCYEEAKKLKKISNQAEFIASFYNNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKKYSTDILGNLYEVKSKKHPQIIKKG.
[0240] In this sequence, residue A579, which can be mutated from N579 to produce SaCas9 nickase, is underlined and shown in bold.
[0241] The amino acid sequence of an exemplary SaKKH Cas9 is as follows: JPEG0007781742000007.jpg89168
[0242] Residue A579, shown above, can be mutated from N579 to produce SaCas9 nickase and is underlined and in bold. Residues K781, K967, and H1014, shown above, can be mutated from E781, N967, and R1014 to produce SaKKH Cas9 and are underlined and italicized.
[0243] High-fidelity Cas9 domain Some embodiments of the present disclosure provide high-fidelity Cas9 domains. In some embodiments, high-fidelity Cas9 domains are engineered Cas9 domains that contain one or more mutations that reduce the electrostatic interaction between the Cas9 domain and the sugar phosphate backbone of DNA compared to the corresponding wild-type Cas9 domain. High-fidelity Cas9 domains that reduce the electrostatic interaction with the sugar phosphate backbone of DNA may have fewer off-target effects. In some embodiments, the Cas9 domain (e.g., a wild-type Cas9 domain) contains one or more mutations that reduce the association of the Cas9 domain with the sugar phosphate backbone of DNA. In some embodiments, the Cas9 domain contains one or more mutations that reduce the association of the Cas9 domain with the sugar phosphate backbone of DNA by at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, or at least 70%.
[0244] In some embodiments, any Cas9 fusion protein described herein comprises one or more of N497X, R661X, Q695X, and / or Q926X mutations, or corresponding mutations in any amino acid sequence described herein, where X is any amino acid. In some embodiments, any Cas9 fusion protein described herein comprises one or more of N497A, R661A, Q695A, and / or Q926A mutations, or corresponding mutations in any amino acid sequence described herein. In some embodiments, the Cas9 domain comprises a D10A mutation, or corresponding mutations in any amino acid sequence described herein. High-fidelity Cas9 domains are known in the art and will be apparent to those skilled in the art. For example, high-fidelity Cas9 domains are described in Kleinstiver, BP, et al. "High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects." Nature 529, 490-495 (2016); and Slaymaker, IM, et al. "Rationally engineered Cas9 nucleases with improved specificity." Science 351, 84-88 (2015), the entire contents of each of which are incorporated herein by reference.
[0245] In some embodiments, the modified Cas9 is a high-fidelity Cas9 enzyme. In some embodiments, the high-fidelity Cas9 enzyme is SpCas9(K855A), eSpCas9(1.1), SpCas9-HF1, or an ultra-high-fidelity Cas9 variant (HypaCas9). The modified Cas9, eSpCas9(1.1), contains an alanine substitution that weakens the interaction between the HNH / RuvC groove and the non-target DNA strand, preventing strand separation and cleavage at off-target sites. Similarly, SpCas9-HF1 reduces off-target editing due to an alanine substitution that disrupts the interaction between Cas9 and the DNA phosphate backbone. HypaCas9 contains mutations in the REC3 domain (SpCas9 N692A / M694A / Q695A / H698A) that enhance Cas9 proofreading and target discrimination. All three high-fidelity enzymes generate fewer off-target edits than wild-type Cas9.
[0246] An exemplary high-fidelity Cas9 is shown below. Cas9 domain mutations that confer increased fidelity relative to Cas9 are shown in bold and underlined. MDKKYSIGL AIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLNLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKIKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMT A FDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAHLFDDKVMKQLKRRRYTGWG A LSRKLINGIRDKQSGKTILDFLKSDGFANRNFM A LIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETR AITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKY FFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVA KVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLF VEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD.
[0247] Guide polynucleotide In one embodiment, the guide polynucleotide is a guide RNA. The RNA / Cas complex can help "guide" the Cas protein to the target DNA. Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand not complementary to the crRNA is first endonucleolytically cleaved and then 3'-5' exonucleolytically trimmed. In nature, DNA binding and cleavage typically require both a protein and RNA. However, single guide RNAs ("sgRNAs" or simply "gRNAs") can be engineered to incorporate aspects of both crRNA and tracrRNA into a single RNA species. See, for example, Jinek M. et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif (PAM or protospacer adjacent motif) within the CRISPR repeat sequence to help distinguish self from non-self. The sequence and structure of Cas9 nuclease are well known to those skilled in the art (see, for example, "Complete genome sequence of an M1 strain of Streptococcus pyogenes," Ferretti, JJ et al., Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III," Deltcheva E. et al., Nature 471:602-607 (2011); and "Programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity," Jinek M. et al., Science 337:816-821 (2012)), the entire contents of each of which are incorporated herein by reference).Cas9 orthologs have been described in various species, including, but not limited to, S. pyrogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on this disclosure, including Cas9 sequences and loci from the organisms disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737, the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA-cleavage domain, i.e., Cas9 is a nickase.
[0248] In some embodiments, the guide polynucleotide is at least one single guide RNA ("sgRNA" or "gNRA"). In some embodiments, the guide polynucleotide is at least one tracrRNA. In some embodiments, the guide polynucleotide does not require a PAM sequence to guide the polynucleotide programmable DNA-binding domain (e.g., Cas9 or Cpf1) to the target nucleotide sequence.
[0249] The polynucleotide-programmable nucleotide-binding domain (e.g., a CRISPR-derived domain) of the base editor disclosed herein can recognize a target polynucleotide sequence by associating with a guide polynucleotide. The guide polynucleotide (e.g., a gRNA) is typically single-stranded and can be programmed to bind to a polynucleotide target sequence in a site-specific manner (i.e., via complementary base pairing), thereby directing a base editor linked to the guide nucleic acid to the target sequence. The guide polynucleotide can be DNA. The guide polynucleotide can be RNA. In some cases, the guide polynucleotide includes a natural nucleotide (e.g., adenosine). In some cases, the guide polynucleotide includes an unnatural (or non-natural) nucleotide (e.g., a peptide nucleic acid or a nucleotide analog). In some cases, the targeted region of the guide nucleic acid sequence can be at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. The targeted region of the guide nucleic acid can be between 10 and 30 nucleotides in length, or between 15 and 25 nucleotides in length, or between 15 and 20 nucleotides in length.
[0250] In some embodiments, a guide polynucleotide comprises two or more individual polynucleotides, which may interact with each other, for example, through complementary base pairing (e.g., a dual guide polynucleotide). For example, a guide polynucleotide may comprise a CRISPR RNA (crRNA) and a trans-activating CRISPR RNA (tracrRNA). For example, a guide polynucleotide may comprise one or more trans-activating CRISPR RNAs (tracrRNAs).
[0251] In Type II CRISPR systems, targeting of nucleic acids by CRISPR proteins (e.g., Cas9) typically requires complementary base pairing between a first RNA molecule (crRNA) that contains a sequence that recognizes the target sequence and a second RNA molecule (trRNA) that contains repeats that form a scaffold region that stabilizes the guide RNA-CRISPR protein complex. Such dual guide RNA systems can be employed as guide polynucleotides that target base editors disclosed herein to target polynucleotide sequences.
[0252] In some embodiments, the base editors provided herein utilize a single guide polynucleotide (e.g., a gRNA). In some embodiments, the base editors provided herein utilize a dual guide polynucleotide (e.g., a duplex gRNA). In some embodiments, the base editors provided herein utilize one or more guide polynucleotides (e.g., a multi-partite gRNA). In some embodiments, a single guide polynucleotide is utilized for the different base editors described herein. For example, a single guide polynucleotide may be utilized for a cytidine base editor and an adenosine base editor.
[0253] In other embodiments, a guide polynucleotide can include both a polynucleotide targeting portion of a nucleic acid and a scaffold portion of a nucleic acid in a single molecule (i.e., a single-molecule guide nucleic acid). For example, a single-molecule guide polynucleotide can be a single guide RNA (sgRNA or gRNA). As used herein, the term guide polynucleotide sequence contemplates any single-molecule, bipartite, or multipartite nucleic acid that can interact with and direct a base editor to a target polynucleotide sequence.
[0254] Typically, a guide polynucleotide (e.g., a crRNA / trRNA complex or a gRNA) comprises a "polynucleotide targeting segment" that includes a sequence capable of recognizing and binding to a target polynucleotide sequence, and a "protein-binding segment" that stabilizes the guide polynucleotide within the polynucleotide-programmable nucleotide-binding domain component of a base editor. In some embodiments, the polynucleotide targeting segment of a guide polynucleotide recognizes and binds to a DNA polynucleotide, thereby facilitating base editing within the DNA. In other cases, the polynucleotide targeting segment of a guide polynucleotide recognizes and binds to an RNA polynucleotide, thereby facilitating base editing within the RNA. As used herein, "segment" refers to a section or region of a molecule, e.g., a contiguous stretch of nucleotides within a guide polynucleotide. A segment can also refer to a region / section of a complex; therefore, a segment can comprise a region of more than one molecule. For example, if a guide polynucleotide comprises multiple nucleic acid molecules, the protein-binding segment can comprise all or part of multiple separate molecules, e.g., hybridized along regions of complementarity. In some embodiments, the protein-binding segment of a DNA-targeting RNA comprising two separate molecules can comprise (i) 40-75 base pairs of a first RNA molecule that is 100 base pairs in length and (ii) 10-25 base pairs of a second RNA molecule that is 50 base pairs in length. The definition of "segment," unless otherwise specifically defined in a particular context, is not limited to a particular number of total base pairs, is not limited to any particular number of base pairs from a given RNA molecule, is not limited to a particular number of separate molecules within a complex, and can include regions of an RNA molecule of any overall length, which may include regions having complementarity with other molecules.
[0255] A guide RNA or guide polynucleotide may comprise two or more RNAs, such as a CRISPR RNA (crRNA) and a transactivating crRNA (tracrRNA). A guide RNA or guide polynucleotide may sometimes comprise a single-stranded RNA formed by fusing portions (e.g., functional portions) of a crRNA and a tracrRNA, or a single guide RNA (sgRNA). A guide RNA or guide polynucleotide may also be a double-stranded RNA comprising a crRNA and a tracrRNA. Furthermore, a crRNA can hybridize with a target DNA.
[0256] As discussed above, guide RNA or guide polynucleotide can be an expression product.For example, the DNA encoding guide RNA can be a vector comprising a sequence encoding guide RNA.Guide RNA or guide polynucleotide can be transferred into cells by transfecting cells with isolated guide RNA or plasmid DNA comprising a sequence encoding guide RNA and a promoter.Guide RNA or guide polynucleotide can also be transferred into cells by other methods, such as virus-mediated gene delivery.
[0257] Guide RNA or guide polynucleotide can be isolated.For example, guide RNA can be transfected into cells or living organisms in the form of isolated RNA.Guide RNA can be prepared by in vitro transcription using any in vitro transcription system known in the art.Guide RNA can be transfected into cells in the form of isolated RNA, not in the form of a plasmid containing the coding sequence of guide RNA.
[0258] A guide RNA or guide polynucleotide may comprise three regions: a first region at the 5' end that may be complementary to a target site in a chromosomal sequence, a second intermediate region that may form a stem-loop structure, and a third 3' region that may be single-stranded. The first region of each guide RNA may also be different so that each guide RNA guides the fusion protein to a specific target site. Furthermore, the second and third regions of each guide RNA may be identical in all guide RNAs.
[0259] The first region of the guide RNA or guide polynucleotide can be complementary to the sequence at the target site in the chromosomal sequence, so that the first region of the guide RNA can base pair with the target site. In some cases, the first region of the guide RNA can comprise 10 to 25 nucleotides or about 10 to about 25 nucleotides (i.e., 10 nucleotides to nucleotides; or about 10 nucleotides to about 25 nucleotides; or 10 nucleotides to about 25 nucleotides; or about 10 nucleotides to 25 nucleotides) or more. For example, the region of base pairing between the first region of the guide RNA and the target site in the chromosomal sequence can be 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25, or about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 22, about 23, about 24, about 25, or more nucleotides in length. Sometimes, the first region of the guide RNA can be 19, 20, or 21 nucleotides in length, or about 19, about 20, or about 21 nucleotides in length.
[0260] The guide RNA or guide polynucleotide may comprise a second region that forms a secondary structure. For example, the secondary structure formed by the guide RNA may comprise a stem (or hairpin) and a loop. The length of the loop and stem may vary. For example, the loop may be 3 to 10 or about 3 to 10 nucleotides in length, and the stem may be 6 to 20 or about 6 to about 20 base pairs in length. The stem may comprise one or more bulges of 1 to 10 nucleotides or about 10 nucleotides. The total length of the second region may be 16 to 60 nucleotides or about 16 to about 60 nucleotides in length. For example, the loop may be 4 nucleotides or about 4 nucleotides in length, and the stem may be 12 base pairs or about 12 base pairs.
[0261] The guide RNA or guide polynucleotide can also include a third region at the 3' end, which can be essentially single-stranded. For example, the third region may not be complementary to any chromosomal sequence in the target cell, or may not be complementary to the rest of the guide RNA. Furthermore, the length of the third region can vary. The third region can be more than 4 nucleotides or more than about 4 nucleotides in length. For example, the length of the third region can range from 5 to 60 nucleotides or about 5 to 60 nucleotides in length.
[0262] Guide RNA or guide polynucleotide can target any exon or intron of gene target.In some cases, guide can target exon 1 or 2 of gene, and in other cases, guide can target exon 3 or 4 of gene.Composition can contain multiple guide RNAs that all target the same exon, or in some cases, multiple guide RNAs that can target different exons.Exons and introns of gene can be targeted.
[0263] The guide RNA or guide polynucleotide can target a nucleic acid sequence of 20 nucleotides or about 20 nucleotides. The target nucleic acid can be less than 20 nucleotides or less than about 20 nucleotides. The target nucleic acid can be at least 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, or at least about 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, or any number between 1 and 100 nucleotides in length. The target nucleic acid can be up to 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, or up to about 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, or any number of nucleotides in between 1 and 100 in length. The target nucleic acid sequence can be 20 bases or about 20 bases immediately 5' to the first nucleotide of the AM. A guide RNA can target the nucleic acid sequence. The target nucleic acid can be at least 1-10, 1-20, 1-30, 1-40, 1-50, 1-60, 1-70, 1-80, 1-90, or 1-100 nucleotides, or at least about 1-10, about 1-20, about 1-30, about 1-40, about 1-50, about 1-60, about 1-70, about 1-80, about 1-90, or about 1-100 nucleotides.
[0264] A guide polynucleotide, e.g., a guide RNA, can refer to a nucleic acid that can hybridize to another nucleic acid, e.g., a target nucleic acid or a protospacer in the genome of a cell. A guide polynucleotide can be RNA. A guide polynucleotide can be DNA. A guide polynucleotide can be programmed or designed to site-specifically bind to a nucleic acid sequence. A guide polynucleotide can include a polynucleotide strand and be referred to as a single guide polynucleotide. A guide polynucleotide can include two polynucleotide strands and be referred to as a double guide polynucleotide. A guide RNA can be introduced into a cell or embryo as an RNA molecule. For example, the RNA molecule can be transcribed in vitro and / or chemically synthesized. The RNA can be transcribed from a synthetic DNA molecule, e.g., a gBlocks® gene fragment. The guide RNA can then be introduced into a cell or embryo as an RNA molecule. A guide RNA can also be introduced into a cell or embryo in the form of a non-RNA nucleic acid molecule, e.g., a DNA molecule. For example, DNA encoding the guide RNA can be operably linked to a promoter regulatory sequence for expression of the guide RNA in a cell or embryo of interest. The RNA coding sequence can be operably linked to a promoter sequence recognized by RNA polymerase III (PolIII). Plasmid vectors that can be used to express guide RNA include, but are not limited to, px330 vector and px333 vector. In some cases, a plasmid vector (e.g., px333 vector) can contain at least two guide RNA-encoding DNA sequences.
[0265] Methods for selecting, designing, and validating guide polynucleotides, such as guide RNAs, and target sequences are described herein and are known to those skilled in the art. For example, to minimize the disruptive effects of potential substrates of the deaminase domain (e.g., AID domain) in the nucleic acid base editor system, the number of residues that may be unintentionally targeted for deamination (e.g., off-target C residues that may potentially exist on ssDNA in the target nucleic acid locus) may be minimized. Furthermore, software tools can be used to optimize the gRNA corresponding to the target nucleic acid sequence, for example, to minimize the total off-target activity throughout the genome. For example, for each possible targeting domain selection using S. pyrogenes Cas9, all off-target sequences (preceding the selected PAM, e.g., NAG or NGG) containing up to a certain number (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) of mismatched base pairs may be identified throughout the genome. The first region of the gRNA that is complementary to the target site can be identified, and all first regions (e.g., crRNA) can be ranked according to their total predicted off-target score. That is, the top-ranked targeting domain represents the one that is likely to have the greatest on-target activity and the least off-target activity. Targeting gRNA candidates can be functionally evaluated by using methods known in the art and / or as defined herein.
[0266] As a non-limiting example, target DNA hybridizing sequences in the crRNA of guide RNAs for use with Cas9 may be identified using a DNA sequence search algorithm. gRNA design may be performed using custom gRNA design software based on the publicly available tool cas-offinder, such as that described in Bae S., Park J., & Kim J.-S. Cas-OFFinder: A fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases. Bioinformatics 30, 1473-1475 (2014). This software calculates and then scores guides for their genome-wide off-target propensity. Typically, matches ranging from perfect matches to 7 mismatches are considered for guides ranging in length from 17 to 24. Once off-target sites are determined computationally, a total score is calculated for each guide and summarized in a tabular output using a web interface. In addition to identifying potential target sites adjacent to the PAM sequence, the software also identifies all PAM-flanking sequences that differ from the selected target site by 1, 2, 3, or more nucleotides. A target nucleic acid sequence, such as a genomic DNA sequence for a target gene, can be obtained and screened for repetitive elements using publicly available tools, such as the RepeatMasker program. RepeatMasker searches the input DNA sequence for repetitive elements and low-complexity regions. The output is a detailed annotation of the repeats present in a given query sequence.
[0267] After identification, the first region of the guide RNA, e.g., crRNA, is ranked based on their distance to the target site, their orthogonality, and the presence of a 5' nucleotide to closely match the relevant PAM sequence (e.g., a 5' G based on the identification of a close match in the human genome containing a relevant PAM, e.g., the NGG PAM of S. pyrogenes, or the NNGRRT or NNGRRV PAM of S. aureus). As used herein, orthogonality refers to the number of sequences in the human genome that contain a minimal number of mismatches with the target sequence. "High level of orthogonality" or "good orthogonality" may refer, for example, to a 20-mer targeting domain that has no identical sequences to the intended target in the human genome and no sequences containing one or two mismatches in the target sequence. A targeting domain with good orthogonality may be selected to minimize off-target DNA cleavage.
[0268] In some embodiments, a reporter system may be used to detect base editing activity and test candidate guide polynucleotides. In some embodiments, the reporter system may include a reporter gene-based assay, in which base editing activity results in expression of a reporter gene. For example, the reporter system includes a reporter gene containing an inactivated start codon, e.g., a mutation from 3'-TAC-5' to 3'-CAC-5' on the template strand. Upon successful deamination of the target C, the corresponding mRNA is transcribed as 5'-AUG-3' instead of 5'-GUG-3', allowing translation of the reporter gene. Suitable reporter genes will be apparent to those skilled in the art. Non-limiting examples of reporter genes include genes encoding green fluorescent protein (GFP), red fluorescent protein (RFP), luciferase, secreted alkaline phosphatase (SEAP), or other genes whose expression can be detected and will be apparent to those skilled in the art. Using a reporter system, for example, numerous different gRNAs can be tested to determine which residues in the target DNA sequence each deaminase targets. sgRNAs targeting the non-template strand can also be tested to evaluate the off-target effects of specific base-editing proteins, such as Cas9 deaminase fusion proteins. In some embodiments, such gRNAs can be designed so that mutant start codons do not base-pair with the gRNA. The guide polynucleotide can include standard ribonucleotides, modified ribonucleotides (e.g., pseudouridine), ribonucleotide isomers, and / or ribonucleotide analogs. In some embodiments, the guide polynucleotide can include at least one detectable label. The detectable label can be a fluorophore (e.g., FAM, TMR, Cy3, Cy5, Texas Red, Oregon Green, Alexa Fluors, Halo Tag, or an appropriate fluorescent dye), a detection tag (e.g., biotin, digoxigenin, etc.), a quantum dot, or a gold particle.
[0269] The guide polynucleotide can be chemically synthesized, enzymatically synthesized, or a combination thereof. For example, the guide RNA can be synthesized using standard phosphoramidite-based solid-phase synthesis. Alternatively, the guide RNA can be synthesized in vitro by operably linking DNA encoding the guide RNA to a promoter regulatory sequence recognized by a phage RNA polymerase. Examples of suitable phage promoter sequences include T7, T3, SP6 promoter sequences, or variations thereof. In embodiments where the guide RNA comprises two separate molecules (e.g., crRNA and tracrRNA), the crRNA can be chemically synthesized and the tracrRNA can be enzymatically synthesized.
[0270] In some embodiments, the base editor system may include multiple guide polynucleotides, e.g., gRNAs. For example, the gRNAs may target one or more target loci included in the base editor system (e.g., at least one gRNA, at least two gRNAs, at least five gRNAs, at least 10 gRNAs, at least 20 gRNAs, at least 30 gRNAs, at least 50 gRNAs). The multiple gRNA sequences can be arranged in tandem, preferably separated by direct repeats.
[0271] The DNA sequence encoding the guide RNA or guide polynucleotide can also be part of a vector. Furthermore, the vector can include additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), selection marker sequences (e.g., GFP or antibiotic resistance genes, such as puromycin), replication origins, etc. The DNA molecule encoding the guide RNA (gRNA) can also be linear. The DNA molecule encoding the guide RNA (gRNA) or guide polynucleotide can also be circular.
[0272] In some embodiments, one or more components of the base editor system may be encoded by a DNA sequence. Such DNA sequences may be introduced together or separately into an expression system, e.g., into a cell. For example, DNA sequences encoding a polynucleotide programmable nucleotide binding domain and a guide RNA may be introduced into a cell, with each DNA sequence being part of a separate molecule (e.g., one vector containing the polynucleotide programmable nucleotide binding domain coding sequence and a second vector containing the guide RNA coding sequence), or both being part of the same molecule (e.g., one vector containing coding (and regulatory) sequences for both the polynucleotide programmable nucleotide binding domain and the guide RNA).
[0273] The guide polynucleotide may contain one or more modifications to confer new or enhanced characteristics to the nucleic acid. The guide polynucleotide may contain a nucleic acid affinity tag. The guide polynucleotide may contain synthetic nucleotides, synthetic nucleotide analogs, nucleotide derivatives, and / or modified nucleotides.
[0274] In some cases, gRNA or guide polynucleotide may contain modifications.Modifications can be made anywhere in gRNA or guide polynucleotide.One or more modifications can be made in a single gRNA or guide polynucleotide.GRNA or guide polynucleotide can be subjected to quality control after modification.In some cases, quality control can include PAGE, HPLC, MS, or a combination thereof.
[0275] The modification of the gRNA or guide polynucleotide can be a substitution, insertion, deletion, chemical modification, physical modification, stabilization, purification, or any combination thereof.
[0276] The gRNA or guide polynucleotide may also be modified with any of the following amino acids: 5' adenylate, 5' guanosine-triphosphate cap, 5' N7-methylguanosine-triphosphate cap, 5' triphosphate cap, 3' phosphate, 3' thiophosphate, 5' phosphate, 5' thiophosphate, Cis-Syn thymidine dimer, trimer, C12 spacer, C3 spacer, C6 spacer, d spacer, PC spacer, r spacer, spacer 18, spacer 9, 3'-3' modified, 5'-5' modified, abasic, acridine, azobenzene, biotin, biotin BB, biotin TEG, cholesteryl TEG, desthiobiotin TEG, DNPTEG, DNP-X, DOTA, dT-biotin, dual biotin, PC biotin, psoralen C2, psoralen C6, TINA, 3' DABCYL, black hole quencher 1 , Black Hole Quencher 2, DABCYLSE, dT-DABCYL, IRDyeQC-1, QSY-21, QSY-35, QSY-7, QSY-9, a carboxyl linker, a thiol linker, a 2'-deoxyribonucleoside analog purine, a 2'-deoxyribonucleoside analog pyrimidine, a ribonucleoside analog, a 2'-O-methylribonucleoside analog, a sugar-modified analog, a wobble / universal base, a fluorescent dye label, 2'-fluoroRNA, 2'-O-methylRNA, a methylphosphonate, phosphodiester DNA, phosphodiester RNA, phosphorothioate DNA, phosphorothioate RNA, UNA, pseudouridine-5'-triphosphate, 5'-methylcytidine-5'-triphosphate, or any combination thereof.
[0277] In some cases, modification is permanent.In other cases, modification is temporary.In some cases, multiple modifications are made to gRNA or guide polynucleotide.The modification of gRNA or guide polynucleotide can change the physicochemical properties of nucleotide, such as its conformation, polarity, hydrophobicity, chemical reactivity, base pairing interaction, or any combination thereof.
[0278] The PAM sequence can be any PAM sequence known in the art. Suitable PAM sequences include, but are not limited to, NGG, NGA, NGC, NGN, NGT, NGCG, NGAG, NGAN, NGNG, NGCN, NGCG, NGTN, NNGRRT, NNNRRT, NNGRR(N), TTTV, TYCV, TYCV, TATV, NNNNGATT, NNAGAAW, or NAAAAC. Y is a pyrimidine, N is any nucleotide base, and W is A or T.
[0279] Modifications can also be phosphorothioate substitutions. In some cases, native phosphodiester bonds can be susceptible to rapid degradation by cellular nucleases, and internucleotide linkage modifications using phosphorothioate (PS) bond substitutions can make them more stable to hydrolysis by cellular degradation. Modifications can increase stability in gRNAs or guide polynucleotides. Modifications can also enhance biological activity. In some cases, phosphorothioate-enhanced RNA gRNAs can inhibit RNase A, RNase T1, bovine serum nuclease, or any combination thereof. These properties enable the use of PS-RNA gRNAs for applications where exposure to nucleases is likely in vivo or in vitro. For example, phosphorothioate (PS) bonds can be introduced within the last 3–5 nucleotides of the 5′- or ′-end of the gRNA to inhibit exonucleolytic degradation. In some cases, phosphorothioate bonds can be added throughout the gRNA to reduce attack by endonucleases.
[0280] Protospacer adjacent motif The term "protospacer adjacent motif (PAM)" or PAM-like motif refers to a 2-6 base pair DNA sequence that immediately follows the DNA sequence targeted by the Cas9 nuclease in the CRISPR bacterial adaptive immune system. In some embodiments, the PAM can be a 5' PAM (i.e., located upstream of the 5' end of the protospacer). In other embodiments, the PAM can be a 3' PAM (i.e., located downstream of the 5' end of the protospacer).
[0281] The PAM sequence is essential for target binding, but the exact sequence depends on the type of Cas protein.
[0282] The base editors provided herein may contain a domain derived from a CRISPR protein that can bind a nucleotide sequence containing a canonical or non-canonical protospacer adjacent motif (PAM) sequence. A PAM site is a nucleotide sequence adjacent to a target polynucleotide sequence. Some embodiments of the present disclosure provide base editors that include all or part of a CRISPR protein with different PAM specificities. For example, Cas9 proteins, such as Cas9 from S. pyrogenes (spCas9), typically require the canonical NGG PAM sequence to bind to a specific nucleic acid region, where the "N" in "NGG" is adenine (A), thymine (T), guanine (G), or cytosine (C), and G is guanine. The PAM can be CRISPR protein-specific and can vary between different base editors that contain domains derived from different CRISPR proteins. The PAM can be 5' or 3' of the target sequence. The PAM can be upstream or downstream of the target sequence. A PAM can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides in length. Often, a PAM is between 2 and 6 nucleotides in length. Several PAM variants are listed in Table 1 below. [Table 1]
[0283] In some embodiments, the PAM is NGC. In some embodiments, the NGC PAM is recognized by Cas9 variants. In some embodiments, the NGC PAM variant comprises one or more amino acid substitutions selected from D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (collectively referred to as "MQKFRAER").
[0284] In some embodiments, the PAM is NGT. In some embodiments, the NGT PAM is a variant. In some embodiments, the NGT PAM variant is generated via targeted mutations at one or more of residues 1335, 1337, 1135, 1136, 1218, and / or 1219. In some embodiments, the NGT PAM variant is generated via targeted mutations at one or more of residues 1219, 1335, 1337, 1218. In some embodiments, the NGT PAM variant is generated via targeted mutations at one or more of residues 1135, 1136, 1218, 1219, and 1335. In some embodiments, the NGT PAM variant is selected from the set of targeted mutations set forth in Tables 2 and 3 below. [Table 2] [Table 3]
[0285] In some embodiments, the NGT PAM variant is selected from variants 5, 7, 28, 31, or 36 of Table 2 and Table 3. In some embodiments, the variant has enhanced NGT PAM recognition.
[0286] In some embodiments, the NGT PAM variant has a mutation at residues 1219, 1335, 1337, and / or 1218. In some embodiments, the NGT PAM variant is selected from the variants shown in Table 4 below with a mutation to enhance recognition. [Table 4]
[0287] In some embodiments, the NGT PAM is selected from the variants shown in Table 5 below. [Table 5-1]
[0288] In some embodiments, the Cas9 domain is a Cas9 domain derived from Streptococcus pyrogenes (SpCas9). In some embodiments, the SpCas9 domain is a nuclease-active SpCas9, a nuclease-inactive SpCas9 (SpCas9d), or a SpCas9 nickase (SpCas9n). In some embodiments, the SpCas9 comprises a D9X mutation, or a corresponding mutation in any amino acid sequence provided herein, where X is any amino acid except D. In some embodiments, the SpCas9 comprises a D9A mutation, or a corresponding mutation in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain, SpCas9d domain, or SpCas9n domain is capable of binding to a nucleic acid sequence having a non-canonical PAM. In some embodiments, the SpCas9 domain, SpCas9d domain, or SpCas9n domain can bind to a nucleic acid sequence having an NGG, NGA, or NGCG PAM sequence.
[0289] In some embodiments, the SpCas9 domain comprises one or more of D1135X, R1335X, and T1336X mutations, or corresponding mutations in any amino acid sequence provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of D1135E, R1335Q, and T1336R mutations, or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain comprises D1135E, R1335Q, and T1336R mutations, or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain comprises one or more of D1135X, R1335X, and T1336X mutations, or corresponding mutations in any amino acid sequence provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of D1135V, R1335Q, and T1336R mutations, or corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 domain...
Claims
1. 1. An in vitro or ex vivo method of editing a polynucleotide to introduce or modify a splice acceptor or splice donor site, comprising contacting the polynucleotide with a base editor in complex with one or more guide polynucleotides, wherein the base editor: i) a Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5'-NGC-3', and comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising one of the following combinations of amino acid sequence substitutions relative to SEQ ID NO:69: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); or D1135C, S1136W, G1218N, E1219W, A1322R, and R1335N (244 SpCas9) an SpCas9 domain comprising: ii) a cytidine deaminase domain, A) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 (PpAPOBEC1); a) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 and containing a substitution corresponding to the H122A mutation in the amino acid sequence of SEQ ID NO: 130; or C) an amino acid sequence having at least 90% identity to SEQ ID NO: 127 (AmAPOBEC1) a cytidine deaminase domain comprising wherein the one or more guide polynucleotides target the base editor to result in an alteration that introduces a splice acceptor or a splice donor site or modifies a splice acceptor or a splice donor site, wherein the method does not include editing the genome of a human embryo.
2. 1. An in vitro or ex vivo method of editing an SBDS polynucleotide comprising a mutation associated with Shwachman-Diamond Syndrome (SDS), comprising contacting the SBDS polynucleotide with a base editor in complex with one or more guide polynucleotides, wherein the base editor comprises: i) a Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5'-NGC-3', and comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising one of the following combinations of amino acid sequence substitutions relative to SEQ ID NO:69: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); or D1135C, S1136W, G1218N, E1219W, A1322R, and R1335N (244 SpCas9) an SpCas9 domain comprising: ii) a cytidine deaminase domain, A) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 (PpAPOBEC1); a) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 and containing a substitution corresponding to the H122A mutation in the amino acid sequence of SEQ ID NO: 130; or C) an amino acid sequence having at least 90% identity to SEQ ID NO: 127 (AmAPOBEC1) a cytidine deaminase domain comprising wherein the one or more guide polynucleotides target the base editor to alter a mutation associated with Shwakman-Diamond Syndrome (SDS), and wherein the method does not include editing the genome of a human embryo.
3. 3. The method of claim 1 or 2, wherein the base editor is a variant of BE4 in which APOBEC-1 is substituted for the cytidine deaminase domain.
4. 3. The method of Claim 2, wherein the one or more guide polynucleotides target the base editor to result in a C.G to T.A change at rs113993993 258+2T>C.
5. 5. The method of claim 4, wherein the guide polynucleotide targets a polynucleotide target sequence selected from GTAAGCAGGCGGGTAACAGCTGC, AGCAGGCGGGTAACAGCTGCAGC, GCGGGTAACAGCTGCAGCATAGC, GTAAGCAGGCGGGTAACAGC, AGCAGGCGGGTAACAGCTGC, GCGGGTAACAGCTGCAGCAT, GCAGGCGGGTAACAGCTGC, CAGGCGGGTAACAGCTGC, AGGCGGGTAACAGCTGC, or AAGCAGGCGGGTAACAGCTGC.
6. The method of any one of claims 1 to 5, wherein the contacting is intracellular.
7. 3. The method of claim 2, wherein the mutation associated with Shwachman-Diamond Syndrome (SDS) encodes a truncated SBDS polypeptide.
8. the SpCas9 domain is a nuclease-inactive or nickase variant, or The SpCas9 domain is a nickase variant comprising the amino acid substitution D10A. The method according to any one of claims 1 to 7.
9. i) a Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5'-NGC-3', and comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising one of the following combinations of amino acid sequence substitutions relative to SEQ ID NO:69: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); or D1135C, S1136W, G1218N, E1219W, A1322R, and R1335N (244 SpCas9) an SpCas9 domain comprising: ii) a cytidine deaminase domain, A) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 (PpAPOBEC1); a) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 and containing a substitution corresponding to the H122A mutation in the amino acid sequence of SEQ ID NO: 130; or C) an amino acid sequence having at least 90% identity to SEQ ID NO: 127 (AmAPOBEC1) a cytidine deaminase domain comprising a base editor, or a polynucleotide encoding said base editor, comprising: one or more guide polynucleotides that target the base editor to effect a change associated with aberrant splicing; 10. A cell comprising: a human embryonic stem cell;
10. the cells are induced pluripotent stem cells or hematopoietic stem cells; the cells express an SBDS protein; the cells are derived from a subject suffering from Shwachman-Diamond Syndrome (SDS); and / or the cell is a non-human mammalian cell or a human cell; The cell of claim 9.
11. 11. The cell of claim 9 or 10, wherein the SpCas9 domain is a nuclease-inactive variant, a nickase variant, or a nickase variant comprising the amino acid substitution D10A.
12. the cytidine deaminase domain is a non-naturally occurring modified cytidine deaminase; or the base editor is a variant of BE4 in which APOBEC-1 is replaced with the cytidine deaminase domain; The cell according to any one of claims 9 to 11.
13. 13. A composition for use in a method for treating Shwakman-Diamond Syndrome (SDS) or a disease associated with abnormal splicing in a subject in need thereof, the composition comprising a cell according to any one of claims 9 to 12.
14. 1. A composition for use in a method of treating Shwakman-Diamond Syndrome (SDS) in a subject, comprising: i) a Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5'-NGC-3', and comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising one of the following combinations of amino acid sequence substitutions relative to SEQ ID NO:69: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); or D1135C, S1136W, G1218N, E1219W, A1322R, and R1335N (244 SpCas9) an SpCas9 domain comprising: ii) a cytidine deaminase domain, A) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 (PpAPOBEC1); a) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 and containing a substitution corresponding to the H122A mutation in the amino acid sequence of SEQ ID NO: 130; or C) an amino acid sequence having at least 90% identity to SEQ ID NO: 127 (AmAPOBEC1) a cytidine deaminase domain comprising a base editor, or a polynucleotide encoding said base editor, comprising: one or more guide polynucleotides that target the base editor to effect mutational alterations associated with SDS; A composition comprising:
15. 1. A composition for use in a method of treating a genetic disease associated with aberrant splicing in a subject, comprising: i) a Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5'-NGC-3', and comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising one of the following combinations of amino acid sequence substitutions relative to SEQ ID NO:69: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); or D1135C, S1136W, G1218N, E1219W, A1322R, and R1335N (244 SpCas9) an SpCas9 domain comprising: ii) a cytidine deaminase domain, A) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 (PpAPOBEC1); a) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 and containing a substitution corresponding to the H122A mutation in the amino acid sequence of SEQ ID NO: 130; or C) an amino acid sequence having at least 90% identity to SEQ ID NO: 127 (AmAPOBEC1) a cytidine deaminase domain comprising a base editor, or a polynucleotide encoding said base editor, comprising: one or more guide polynucleotides that target the base editor to modify a pathogenic mutation that alters splicing; A composition comprising:
16. 16. The composition of claim 14 or 15, wherein the SpCas9 domain is a nuclease-inactive variant, a nickase variant, or a nickase variant containing the amino acid substitution D10A.
17. 17. The composition of any one of claims 14 to 16, wherein the deaminase domain is a modified cytidine deaminase that does not occur in nature.
18. the base editor is a variant of BE4 in which APOBEC-1 is replaced with the cytidine deaminase domain; and / or the base editor targets SNP rs113993993 258+2T>C in the SBDS polynucleotide sequence to restore correct splicing; 18. The composition of claim 17.
19. 1. An in vitro or ex vivo method for producing non-human mammalian cells or human cells or precursors thereof, comprising: (a) in induced pluripotent stem cells containing a gene conversion associated with Shwachman-Diamond syndrome (SDS), i) a Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5'-NGC-3', and comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising one of the following combinations of amino acid sequence substitutions relative to SEQ ID NO:69: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); or D1135C, S1136W, G1218N, E1219W, A1322R, and R1335N (244 SpCas9) an SpCas9 domain comprising: ii) a cytidine deaminase domain, A) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 (PpAPOBEC1); a) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 and containing a substitution corresponding to the H122A mutation in the amino acid sequence of SEQ ID NO: 130; or C) an amino acid sequence having at least 90% identity to SEQ ID NO: 127 (AmAPOBEC1) a cytidine deaminase domain comprising a base editor or a polynucleotide encoding the base editor, comprising: One or more guide polynucleotides that target base editors to alter mutations associated with SDS to introduce; and (b) differentiating the induced pluripotent stem cells or progenitors into a desired cell type. wherein the method does not include editing the genome of a human embryo.
20. 20. The method of Claim 19, wherein the SpCas9 domain is a nuclease-inactive variant, a nickase variant, or a nickase variant comprising the amino acid substitution D10A.
21. 21. The method of claim 19 or 20, wherein the base editor is a variant of BE4 in which APOBEC-1 is substituted for the cytidine deaminase domain.
22. 1. An in vitro or ex vivo method for editing a pathogenic mutation in a gene that results in aberrant splicing, comprising: a target nucleotide sequence, at least a portion of which is located in the gene or its reverse complement, (i) a Streptococcus pyogenes Cas9 (SpCas9) domain in association with a guide polynucleotide that directs the base editor to target a target polynucleotide sequence, at least a portion of which is located in the gene or its reverse complement, the Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5′-NGC-3′, and comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising one of the following combinations of amino acid sequence substitutions relative to SEQ ID NO:69: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); or D1135C, S1136W, G1218N, E1219W, A1322R, and R1335N (244 SpCas9) an SpCas9 domain comprising: (ii) a cytidine deaminase domain capable of deaminating a pathogenic mutation or its complementary nucleobase that results in aberrant splicing, wherein the cytidine deaminase domain comprises: A) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 (PpAPOBEC1); a) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 and containing a substitution corresponding to the H122A mutation in the amino acid sequence of SEQ ID NO: 130; or C) an amino acid sequence having at least 90% identity to SEQ ID NO: 127 (AmAPOBEC1) a cytidine deaminase domain comprising contacting a base editor comprising: and editing the pathogenic mutation by directing a base editor at the target nucleotide sequence and deaminating the pathogenic mutation or its complementary nucleobase. Including, deaminating the pathogenic mutation or its complementary nucleobase, thereby converting the pathogenic mutation into a sequence that permits splicing, thereby correcting the pathogenic mutation, wherein the method does not include editing the genome of a human embryo.
23. the pathogenic variant is in the SBDS gene; the pathogenic variant is in the SBDS gene and alters the splicing of said gene; the pathogenic variant is in the SBDS gene and encodes a truncated polypeptide; the base editor inserts a new splice acceptor or splice donor site or modifies a splice acceptor or splice donor site that contains a mutation; and / or the base editor corrects a splice donor SNP site containing the rs113993993 C→T mutation in the SBDS gene; 23. The method of claim 22.
24. 1. An in vitro or ex vivo method of producing a cell, tissue, or organ for treating SDS in a subject in need thereof by correcting a pathogenic mutation in the SDS gene of the cell, tissue, or organ, comprising: the cells, tissues, or organs, (i) a Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5'-NGC-3', and comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising one of the following combinations of amino acid sequence substitutions relative to SEQ ID NO:69: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); or D1135C, S1136W, G1218N, E1219W, A1322R, and R1335N (244 SpCas9) an SpCas9 domain comprising: (ii) a cytidine deaminase domain capable of deaminating a pathogenic mutation or its complementary nucleobase, wherein the cytidine deaminase domain comprises: A) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 (PpAPOBEC1); a) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 and containing a substitution corresponding to the H122A mutation in the amino acid sequence of SEQ ID NO: 130; or C) an amino acid sequence having at least 90% identity to SEQ ID NO: 127 (AmAPOBEC1) a cytidine deaminase domain comprising contacting a base editor comprising: contacting the cell, tissue, or organ with a guide polynucleotide, wherein the guide polynucleotide directs a base editor to a target nucleotide sequence located at least in part in the gene or its reverse complement; and editing the pathogenic mutation or its complementary nucleobase by deaminating the mutation when a base editor is directed against the target nucleotide sequence. Including, deaminating the pathogenic mutation or its complementary nucleobase to permit splicing, thereby producing the cell, tissue, or organ for treating SDS, wherein the method does not include editing the genome of a human embryo.
25. the mutation associated with Shwachman-Diamond Syndrome alters the splicing of the gene; or The mutation associated with Shwachman-Diamond Syndrome (SDS) encodes a truncated SBDS polypeptide.
25. The method of claim 24.
26. 25. The method of Claim 24, wherein the base editor inserts a new splice acceptor or splice donor site, or modifies a splice acceptor or splice donor site that contains a mutation.
27. 1. A composition comprising a base editor bound to a guide RNA, wherein the guide RNA comprises a nucleic acid sequence complementary to an SBDS gene associated with Shwachman-Diamond Syndrome (SDS), and wherein the base editor i) a Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5'-NGC-3', and comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising one of the following combinations of amino acid sequence substitutions relative to SEQ ID NO:69: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); or D1135C, S1136W, G1218N, E1219W, A1322R, and R1335N (244 SpCas9) an SpCas9 domain comprising: ii) a cytidine deaminase domain, A) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 (PpAPOBEC1); a) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 and containing a substitution corresponding to the H122A mutation in the amino acid sequence of SEQ ID NO: 130; or C) an amino acid sequence having at least 90% identity to SEQ ID NO: 127 (AmAPOBEC1) a cytidine deaminase domain comprising A composition comprising:
28. the base editor: (i) comprising a Cas9 nickase; (ii) comprises a nuclease-inactive Cas9; (v) does not contain a UGI domain; and / or (vi) APOBEC-1 is a variant of BE4 in which the cytidine deaminase domain is replaced; 28. The composition of claim 27.
29. 29. The composition of claim 27 or 28, further comprising a pharmaceutically acceptable excipient, diluent, or carrier.
30. 30. A pharmaceutical composition for the treatment of Shwakman-Diamond Syndrome (SDS), comprising the composition of claim 29.
31. 31. The pharmaceutical composition of Claim 30, wherein the guide RNA and the base editor are formulated together or separately.
32. 32. The pharmaceutical composition of claim 30 or 31, wherein the guide RNA comprises a 5' to 3' nucleic acid sequence selected from one or more of the following: GUAAGCAGGCGGGUAACAGC; AGCAGGCGGGUAACAGCUGC; GCGGGUAACAGCUGCAGCAU; UGUAAAUGUUUCCUAAGGUC; AAUGUUUCCUAAGGUCAGGU, GCAGGCGGGUAACAGCUGC, CAGGCGGGUAACAGCUGC, AGGCGGGUAACAGCUGC, and AAGCAGGGCGGGUAACAGCUGC.
33. further comprising a vector suitable for expression in mammalian cells, the vector comprises a polynucleotide encoding the base editor; or the vector comprises an mRNA polynucleotide encoding the base editor; or the vector is a viral vector, or The vector is a retroviral vector, an adenoviral vector, a lentiviral vector, a herpesvirus vector, or an adeno-associated viral vector (AAV). The pharmaceutical composition according to any one of claims 30 to 32.
34. 1. A pharmaceutical composition comprising: (i) a nucleic acid encoding a base editor; and (ii) a guide RNA comprising a 5′ to 3′ nucleic acid sequence chosen from one or more of GUAAGCAGGCGGGUAACAGC; AGCAGGCGGGUAACAGCUGC; GCGGGUAACAGCUGCAGCAU; UGUAAAUGUUUCCUAAGGUC; AAUGUUUCCUAAGGUCAGGU, GCAGGCGGGUAACAGCUGC, CAGGCGGGUAACAGCUGC, AGGCGGGUAACAGCUGC, and AAGCAGGCGGGUAACAGCUGC, or a 1, 2, 3, 4, or 5 nucleotide 5′ truncation fragment thereof, wherein the base editor is i) a Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5'-NGC-3', and comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising one of the following combinations of amino acid sequence substitutions relative to SEQ ID NO:69: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); or D1135C, S1136W, G1218N, E1219W, A1322R, and R1335N (244 SpCas9) an SpCas9 domain comprising: ii) a cytidine deaminase domain, A) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 (PpAPOBEC1); a) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 and containing a substitution corresponding to the H122A mutation in the amino acid sequence of SEQ ID NO: 130; or C) an amino acid sequence having at least 90% identity to SEQ ID NO: 127 (AmAPOBEC1) a cytidine deaminase domain comprising A pharmaceutical composition comprising:
35. 35. The pharmaceutical composition of claim 34, further comprising a lipid.
Citation Information
Patent Citations
Diagnosis of shwachman-diamond syndrome
US20060110734A1
Cited By
Compositions and methods for editing mutation to permit transcription or expression
JP2025183226A