Compositions and methods for editing mutations to allow transcription or expression
By editing SDS-related genes using a programmable nucleobase editor, the problem of difficulty in effectively treating SDS in the prior art is solved, and gene functional correction is achieved, potentially improving patients' health status and lifespan.
Patent Information
- Application Number
- CN202411970155.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-29
- Filing Date
- 2020-08-28
- Publication Date
- 2025-05-02
AI Technical Summary
The prior art is difficult to effectively treat SDS, a rare multisystem disease, leading to insufficient secretion of the external pancreatic glands, impaired hematopoietic function and susceptibility to leukemia.
Use a programmable nucleobase editor to edit genes related to SDS, subjecting them to splicing and producing functional gene products. A specific method includes contacting a polynucleotide with a base editor that includes a programmable DNA binding domain and a deaminase domain of the polynucleotide and implementing a mutation that allows transcription by wizarding polynucleotide targeting.
By editing genes, functional correction of SDS-related genes is achieved, potentially improving the clinical manifestations of patients and extending the average life span.
Smart Images

Figure BDA0005219323980000281 
Figure BDA0005219323980000331 
Figure BDA0005219323980000341
Abstract
Description
[0001] This invention is a divisional application of the PCT patent application entered into China with Chinese patent application number 202080076243.7, invention name is “Compositions and methods for editing mutations to allow transcription or expression”, and international application date is August 28, 2020.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application is an International PCT application, claiming priority to and the benefit of U.S. Provisional Application No. 62 / 893,638, filed on August 29, 2019, the contents of which are incorporated herein by reference in their entirety. Background Art
[0004] Shwachman Diamond Syndrome (SDS) is a rare autosomal recessive multisystem disorder characterized by pancreatic insufficiency, impaired hematopoiesis, and susceptibility to leukemia. Patients with SDS exhibit bone marrow failure. Other clinical features include skeletal, immune, liver, and heart disorders. Approximately 90% of patients with clinical features of SDS have biallelic mutations in the evolutionarily conserved Shwachman-Bodian-Diamond Syndrome (SBDS) gene located on chromosome 7. The SDBS protein plays a role in ribosome biogenesis and mitotic spindle stabilization, but its exact molecular function is unclear. Currently, there is no cure for SDS, and patients with the disease are often repeatedly hospitalized due to complications, with an average life expectancy of only approximately 35 years. Therefore, there is an urgent need for improved methods and therapeutics for the treatment of SDS. Summary of the Invention
[0005] As described below, the invention features products, compositions, and methods for using programmable nucleobase editors to edit a gene associated with Schwann-Diesel syndrome (SDS) such that the gene undergoes splicing and produces a functional gene product.
[0006] On the one hand, a method of editing a polynucleotide to allow transcription is provided, wherein the method comprises contacting the polynucleotide with a base editor, the base editor being complexed with one or more guide polynucleotides, wherein the base editor comprises a polynucleotide programmable DNA binding domain and a deaminase domain, and wherein one or more of the guide polynucleotides targets the base editor to effect a change that introduces a mutation that allows transcription. In one embodiment, the mutation that allows transcription is a mutation that changes a stop codon, a mutation that introduces a splice acceptor or splice donor site, or a mutation that corrects a splice acceptor or splice donor site.
[0007] On the one hand, a method for editing an SBDS polynucleotide comprising a mutation associated with Shu-Diehl syndrome (SDS) is provided, wherein the method includes contacting the SBDS polynucleotide with a base editor, the base editor being compounded with one or more guide polynucleotides, wherein the base editor comprises a polynucleotide programmable DNA binding domain and a deaminase domain, and wherein one or more of the guide polynucleotides targets the base editor to achieve changes in mutations associated with Shu-Diehl syndrome (SDS). In one embodiment of the method or its embodiments, the mutation associated with Shu-Diehl syndrome (SDS) results in gene conversion. In one embodiment of the method or its embodiments, the mutation associated with Shu-Diehl syndrome introduces a stop codon or changes the splicing of the gene. In one embodiment of the method or its embodiments, the mutation associated with Shu-Diehl syndrome (SDS) encodes a truncated SBDS polypeptide.
[0008] In one embodiment of any of the above methods and embodiments thereof, the deaminase is a cytidine deaminase or an adenosine deaminase. In one embodiment, the deaminase is an adenosine deaminase. In an embodiment, the adenosine deaminase is selected from ABE8 or an ABE8 variant as listed in Table 7A or Table 7B and the like herein. In another embodiment of the above methods and embodiments thereof, the deaminase is a cytidine deaminase. In one embodiment, the cytosine deaminase is selected from one or more of the following: BE4; rAPOBEC1; PpAPOBEC1; PpAPOBEC1 containing an H122A substitution; AmAPOBEC1; SsAPOBEC2; RrA3F; RrA3F containing an F130L substitution; a BE4 variant in which APOBEC-1 is replaced with a rAPOBEC1 sequence; a BE4 variant in which APOBEC-1 is replaced with an AmAPOBEC1 sequence; a BE4 variant in which APOBEC-1 is replaced with an SsAPOBEC2 sequence; a BE4 variant in which APOBEC-1 is replaced with a PpAPOBEC1 sequence; or a BE4 variant in which APOBEC-1 is replaced with a PpAPOBEC1 sequence containing an H122A substitution. In one embodiment, the PpAPOBEC1 containing the H122A substitution, or the BE4 variant in which APOBEC-1 is replaced with the PpAPOBEC1 sequence containing the H122A substitution, further comprises one or more amino acid mutations selected from R33A, W90F, K34A, R52A, H121A or Y120F. In embodiments of the above methods and embodiments thereof, two or more guide polynucleotides target base editors to effectuate the alteration of two or more mutations associated with Schwann-Diesel syndrome (SDS).
[0009] In another aspect, a method for editing an SBDS polynucleotide comprising a mutation associated with Scholl-Diesel syndrome (SDS) is provided, wherein the method comprises contacting the SBDS polynucleotide with an adenosine base editor (ABE) complexed with one or more guide polynucleotides, wherein the base editor comprises a polynucleotide programmable DNA binding domain and a deaminase domain, and wherein one or more of the guide polynucleotides targets the base editor to effect a change from A·T to G·C at 183-184TA>CT Rs113993991 to generate a missense mutation. In one embodiment, the one or more guide polynucleotides target one of the following sequences: TGTAAATGTTTCCTAAGGTC or AATGTTTCCTAAGGTCAGGT. In one embodiment, the one or more sgRNAs comprise one of the following sequences: UGUAAAUGUUUCCUAAGGUC or AAUGUUUCCUAAGGUCAGGU. In one embodiment, the ABE has a 5'-NGC-3' or 5'-NGG-3' PAM specificity.
[0010] In another aspect, a method for editing an SBDS polynucleotide comprising a mutation associated with Scholl-Diesel syndrome (SDS) is provided, wherein the method comprises contacting the SBDS polynucleotide with a cytidine base editor complexed with one or more guide polynucleotides, wherein the cytidine base editor (CBE) comprises a polynucleotide-programmable DNA-binding domain and a cytidine deaminase domain, and wherein one or more of the guide polynucleotides targets the base editor to effect a C·G to T·A change of rs113993993 258+2T>C. In one embodiment, the CBE has 5'-NGC-3' PAM specificity or is specific for a PAM comprising 5'-NGC-3'. In one embodiment, the guide polynucleotide targets a polynucleotide target sequence selected from the group consisting of GTAAGCAGGCGGGTAACAGCTGC, AGCAGGCGGGTAACAGCTGCAGC, GCGGGTAACAGCTGCAGCATAGC, GTAAGCAGGCGGGTAACAGC, AGCAGGCGGGTAACAGCTGC, GCGGGTAACAGCTGCAGCAT, GCAGGCGGGTAACAGCTGC, CAGGCGGGTAACAGCTGC, AGGCGGGTAACAGCTGC, or AAGCAGGCGGGTAACAGCTGC. In one embodiment, the sgRNA comprises one of the following sequences: GUAAGCAGGCGGGUAACAGC, AGCAGGCGGGUAACAGCUGC, GCGGGUAACAGCUGCAGCA, GCAGGCGGGUAACAGCUGC, CAGGCGGGUAACAGCUGC, AGGCGGGUAACAGCUGC, or AAGCAGGCGGGUAACAGCUGC.
[0011] In other embodiments of the above methods and embodiments thereof, the binding is carried out in a cell, wherein the cell is a eukaryotic cell, a mammalian cell or a human cell. In one embodiment, the cell is in vivo or in vitro. In one embodiment of any of the above methods and embodiments thereof, the base editor introduces a missense mutation, inserts a new splicing acceptor or splicing donor site, and / or corrects a splicing acceptor or splicing donor site comprising a mutation. In one embodiment of any of the above methods and embodiments thereof, the polynucleotide programmable DNA binding domain is a Cas9 selected from Streptococcus pyogenes Cas9 (SpCas9), Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), Streptococcus canis Cas9 (ScCas9) or a variant thereof. In one embodiment, the polynucleotide programmable DNA binding domain is a wild-type or modified Streptococcus pyogenes Cas9 (SpCas9) or a variant thereof. In one embodiment, the polynucleotide programmable DNA binding domain is a modified SpCas9 or SpCas9 variant. In one embodiment, the polynucleotide programmable DNA binding domain comprises a modified SpCas9 or SpCas9 variant with an altered protospacer sequence adjacent motif (PAM) specificity. In one embodiment, SpCas9 has specificity for the PAM nucleic acid sequence 5'-NGC-3' or 5'-NGG-3'. In one embodiment, SpCas9 is a modified SpCas9 or SpCas9 variant that has specificity for the PAM nucleic acid sequence 5'-NGC-3' or a PAM nucleic acid sequence comprising 5'-NGC-3'. In one embodiment, the modified SpCas9 or SpCas9 variant comprises an amino acid sequence listed in Table 1. In one embodiment, the modified SpCas9 is spCas9-MQKFRAER. In one embodiment, the modified SpCas9 or SpCas9 variant comprises Figures 3A to 3C or Figure 10 In one embodiment,
[0012] The modified SpCas9 or SpCas9 variant comprises a combination of amino acid sequence substitutions selected from the group consisting of:
[0013] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R (224SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225SpCas9);
[0014] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226SpCas9);
[0015] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q (227Cas9);
[0016] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230SpCas9);
[0017] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R , D1332, R1335N and T1337 (244SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E and T1337 (245SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q and T1337R (259SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGC Rd1 SpCas9); or
[0018] D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E, and T1337R (NGCRd2 SpCas9).
[0019] In other embodiments of the above methods and embodiments thereof, the polynucleotide programmable DNA binding domain is an inactive nuclease or nickase variant. In one embodiment, the nickase variant comprises the amino acid substitution D10A or its corresponding amino acid substitution. In one embodiment, the deaminase domain is capable of deaminating adenosine or cytosine in deoxyribonucleic acid (DNA). In one embodiment, the adenosine deaminase or cytidine deaminase is a modified adenosine deaminase or cytidine deaminase that does not exist in nature. In one embodiment, the adenosine deaminase is a TadA deaminase. In one embodiment, the TadA deaminase is TadA*7.10, TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, T adA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23 or TadA*8.24. In one embodiment, TadA*7.10 comprises one or more of the following changes: Y147T, Y147R, Q154S, Y123H, V82S, T166R, Q154R. In one embodiment, TadA*7.10 comprises a combination of changes selected from the group consisting of: Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y147T+Q154R; Y147T+Q154S; V82S+Q154S; and Y123H+Y147R+Q154R+I76Y.
[0020] In another embodiment of any of the above methods and embodiments thereof, the one or more guide RNAs comprise CRISPR RNA (crRNA) and a trans-encoded small RNA (tracrRNA), wherein the crRNA comprises a nucleic acid sequence complementary to a SBDS nucleic acid sequence comprising an SDS-associated change. In another embodiment of any of the above methods and embodiments thereof, the base editor is complexed with a single guide RNA (sgRNA) comprising a nucleic acid sequence complementary to a SBDS nucleic acid sequence comprising an SDS-associated change.
[0021] On the other hand, a cell is provided, which is produced by introducing the following items into the cell or its precursor: a base editor; a polynucleotide encoding the base editor, wherein the base editor comprises a polynucleotide programmable DNA binding domain and a deaminase domain; and one or more guide polynucleotides that target the base editor to achieve changes associated with abnormal splicing. In one embodiment, the cell or its precursor is an induced pluripotent stem cell or a hematopoietic stem cell. In one embodiment, the cell expresses SBDS protein. In one embodiment, the cell is from a subject with Shu-Dieter syndrome (SDS). In one embodiment, the cell is a mammalian cell or a human cell. In one embodiment of the cell, the mutation or change results in a gene conversion comprising a stop codon and / or a mutation causing abnormal splicing. In one embodiment, the cell is selected for gene conversion associated with SDS. In one embodiment, the polynucleotide programmable DNA binding domain is a wild-type or modified Streptococcus pyogenes Cas9 (SpCas9) or a variant thereof. In one embodiment, the polynucleotide programmable DNA binding domain comprises wild-type SpCas9 or a modified SpCas9 with an altered protospacer sequence adjacent motif (PAM) specificity. In one embodiment, the modified SpCas9 is specific for the nucleic acid sequence 5'-NGC-3' or a PAM nucleic acid sequence comprising 5'-NGC-3'. In one embodiment, the modified SpCas9 is a Cas9 variant listed in Table 1. In one embodiment, the modified SpCas9 is spCas9-MQKFRAER. In one embodiment of the cell, the modified SpCas9 is a Figures 3A to 3C or Figure 10 In one embodiment of the cell, the SpCas9 variant comprises a combination of amino acid sequence substitutions selected from the group consisting of:
[0022] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R (224SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225SpCas9);
[0023] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226SpCas9);
[0024] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q (227Cas9);
[0025] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230SpCas9);
[0026] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R , D1332, R1335N and T1337 (244SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E and T1337 (245SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q and T1337R (259SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGC Rd1 SpCas9); or
[0027] D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E and T1337R (267 (NGCRd2 SpCas9). In one embodiment of the cell, the programmable polynucleotide binding domain is an inactive nuclease variant or a nickase variant. In one embodiment, the nickase variant comprises the amino acid substitution D10A or its corresponding amino acid substitution. In one embodiment of the cell, the deaminase domain is a cytidine deaminase domain capable of deaminating cytidine in deoxyribonucleic acid (DNA) or an adenosine deaminase domain capable of deaminating adenosine in DNA. In one embodiment, adenosine deaminase domain is a cytidine deaminase domain capable of deaminating cytidine in deoxyribonucleic acid (DNA). The adenosine deaminase or cytidine deaminase is a modified adenosine deaminase or cytidine deaminase that does not exist in nature. In another embodiment of the cell, the adenosine deaminase is a TadA deaminase. In one embodiment, the TadA deaminase is TadA*7.10, TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA* In one embodiment, TadA*7.10 comprises one of the following changes: In one embodiment, TadA*7.10 comprises a combination of changes selected from the group consisting of: Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y147T+Q154R; Y147T+Q154S; V82S+Q154S.In another embodiment of the cell, the cytosine deaminase is selected from one or more of the following: BE4; rAPOBEC1; PpAPOBEC1; PpAPOBEC1 containing an H122A substitution; AmAPOBEC1; SsAPOBEC2; RrA3F; RrA3F containing an F130L substitution; a BE4 variant in which APOBEC-1 is replaced with a rAPOBEC1 sequence; a BE4 variant in which APOBEC-1 is replaced with an AmAPOBEC1 sequence; a BE4 variant in which APOBEC-1 is replaced with an SsAPOBEC2 sequence; a BE4 variant in which APOBEC-1 is replaced with a PpAPOBEC1 sequence; or a BE4 variant in which APOBEC-1 is replaced with a PpAPOBEC1 sequence containing an H122A substitution. In one embodiment, the PpAPOBEC1 containing the H122A substitution, or the BE4 variant in which APOBEC-1 is replaced with a PpAPOBEC1 sequence containing the H122A substitution, further comprises one or more amino acid mutations selected from R33A, W90F, K34A, R52A, H121A or Y120F. In another embodiment of the cell, the one or more guide RNAs comprise CRISPR RNA (crRNA) and a trans-encoded small RNA (tracrRNA), wherein the crRNA comprises a nucleic acid sequence complementary to an SBDS nucleic acid sequence comprising an SDS-associated change. In one embodiment of the cell, the base editor and the one or more guide polynucleotides form a complex within the cell. In one embodiment, the base editor is complexed with a single guide RNA (sgRNA) comprising a nucleic acid sequence complementary to an SBDS nucleic acid sequence comprising an SDS-associated gene change.
[0028] In another aspect, a method of treating Shu-Dieter syndrome (SDS) or a disease associated with abnormal splicing in a subject in need thereof is provided, wherein the method comprises administering to the subject a cell according to the above aspects and embodiments thereof. In one embodiment, the cell is autologous, allogeneic, or xenogeneic to the subject.
[0029] In another aspect, there is provided an isolated cell or cell population propagated or expanded from a cell according to the above aspects and embodiments thereof.
[0030] In another aspect, a method of treating Shu-Dieter syndrome (SDS) in a subject is provided, wherein the method comprises administering to a subject in need thereof: a base editor or a polynucleotide encoding the base editor, wherein the base editor comprises a polynucleotide-programmable DNA-binding domain and a deaminase domain; and one or more guide polynucleotides that target the base editor to effectuate alteration of a mutation associated with SDS.
[0031] In another aspect, a method for treating a genetic disease associated with abnormal splicing in a subject is provided, wherein the method comprises administering to a subject in need thereof: a base editor or a polynucleotide encoding the base editor, wherein the base editor comprises a polynucleotide programmable DNA binding domain and a deaminase domain; and one or more guide polynucleotides that target the base editor to effectuate an alteration of a pathogenic mutation that alters splicing.
[0032] In one embodiment of the above-mentioned method for treating Shu-Dieter syndrome (SDS) in a subject or in the above-mentioned method for treating a disease associated with abnormal splicing in a subject, the subject is a mammal or a human. In one embodiment, the above-mentioned method includes delivering a base editor or a polynucleotide encoding the base editor and one or more guide polynucleotides to the cells of the subject. In one embodiment, the cell expresses a truncated polypeptide. In one embodiment of the above-mentioned method, the change converts the TAA terminator in the SBDS polynucleotide to TGG. In another embodiment of the method, the change causes a change in the K62X in the SBDS polypeptide associated with SDS. In one embodiment of the method, gene conversion associated with SDS results in the expression of a truncated SBDS polypeptide. In another embodiment of the method, the base editor correction replaces the lysine (K) at amino acid position 62 with tryptophan (W). In another embodiment of the method, the polynucleotide programmable DNA binding domain comprises a modified Streptococcus pyogenes Cas9 (SpCas9) or a variant thereof. In another embodiment of the method, the polynucleotide programmable DNA binding domain comprises a modified SpCas9 having an altered protospacer sequence adjacent to the motif (PAM) specificity. In one embodiment, the modified SpCas9 is specific for the PAM nucleic acid sequence 5'-NGC-3' or a PAM nucleic acid sequence comprising 5'-NGC-3'. In one embodiment, the modified SpCas9 is a Cas9 variant listed in Table 1. In one embodiment, the modified SpCas9 is spCas9-MQKFRAER. In another embodiment of these methods, the modified SpCas9 is a Cas9 variant comprising Figures 3A to 3C or Figure 10 In one embodiment, the SpCas9 variant comprises a combination of amino acid sequence substitutions selected from the group consisting of:
[0033] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R (224SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225SpCas9);
[0034] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226SpCas9);
[0035] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q (227Cas9);
[0036] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230SpCas9);
[0037] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R , D1332, R1335N and T1337 (244SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E and T1337 (245SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q and T1337R (259SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGC Rd1 SpCas9); or
[0038] D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E and T1337R (267 (NGCRd2 SpCas9). In other embodiments of the above methods and embodiments thereof, the polynucleotide programmable DNA binding domain is an inactive nuclease variant. In one embodiment of the above methods, the polynucleotide programmable DNA binding domain is a nickase variant. In one embodiment, the nickase variant comprises the amino acid substitution D10A or its corresponding amino acid substitution. In one embodiment of the above methods, wherein the deaminase domain is capable of deaminating adenosine or cytidine in deoxyribonucleic acid (DNA). In one embodiment In one embodiment, the deaminase domain is a modified adenosine deaminase or cytidine deaminase that does not exist in nature. In one embodiment, the adenosine deaminase is a TadA deaminase. In one embodiment, the TadA deaminase is TadA*7.10, TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA In one embodiment, TadA*7.10 comprises one or more of the following changes: Y147T, Y147 R, Q154S, Y123H, V82S, T166R, Q154R; or wherein the TadA*7.10 comprises a combination of changes selected from the group consisting of: Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y147T+Q154R; Y147T+Q154S; V82S+Q154S; and Y123H+Y147R+Q154R+I76Y.In another embodiment of the above method and embodiments thereof, the deaminase domain is a cytidine deaminase selected from one or more of the following: BE4; rAPOBEC1; PpAPOBEC1; PpAPOBEC1 containing an H122A substitution; AmAPOBEC1; SsAPOBEC2; RrA3F; RrA3F containing an F130L substitution; a BE4 variant in which APOBEC-1 is replaced with a rAPOBEC1 sequence; a BE4 variant in which APOBEC-1 is replaced with an AmAPOBEC1 sequence; a BE4 variant in which APOBEC-1 is replaced with an SsAPOBEC2 sequence; a BE4 variant in which APOBEC-1 is replaced with a PpAPOBEC1 sequence; or a BE4 variant in which APOBEC-1 is replaced with a PpAPOBEC1 sequence containing an H122A substitution. In one embodiment, PpAPOBEC1 containing an H122A substitution, or a BE4 variant in which APOBEC-1 is replaced with a PpAPOBEC1 sequence containing an H122A substitution, further comprises one or more amino acid mutations selected from R33A, W90F, K34A, R52A, H121A, or Y120F. In one embodiment of the above method and its embodiment, the base editor targets SNP rs113993993 258+2T>C in the SBDS polynucleotide sequence to restore correct splicing. In one embodiment of the above method, the one or more guide polynucleotides comprise CRISPR RNA (crRNA) and a trans-encoded small RNA (tracrRNA), wherein the crRNA comprises a nucleic acid sequence complementary to the SBDS nucleic acid sequence of the gene conversion. In one embodiment, the base editor is complexed with a single guide RNA (sgRNA), which comprises a nucleic acid sequence complementary to the SBDS nucleic acid sequence comprising the SDS-related gene conversion.
[0039] In another aspect, another method of producing a cell or a precursor thereof, wherein the method comprises:
[0040] (a) Introduction of SDS-associated genes into induced pluripotent stem cells
[0041] A base editor or a polynucleotide encoding the base editor, wherein the base editor comprises a polynucleotide programmable nucleotide binding domain and a cytidine deaminase domain or an adenosine deaminase domain; and
[0042] one or more guide polynucleotides, wherein the one or more guide polynucleotides target a base editor to effect a change in an SDS-associated mutation; and
[0043] (b) Induced pluripotent stem cells or precursors are differentiated into the desired cell type. In one embodiment of the method, the mutation is a gene conversion associated with SDS. In one embodiment of the method, the cell or precursor is obtained from a subject suffering from SDS. In one embodiment, the cell or precursor is a mammalian cell or a human cell. In another embodiment of the method, the polynucleotide programmable DNA binding domain comprises Streptococcus pyogenes Cas9 (SpCas9), modified Streptococcus pyogenes Cas9 (SpCas9) or a variant thereof. In another embodiment, the polynucleotide programmable DNA binding domain comprises a modified SpCas9 with a modified protospacer sequence adjacent to a motif (PAM) specificity. In one embodiment of the method, SpCas9 is specific for the nucleic acid sequence 5'-NGG-3', while the modified SpCas9 is specific for the nucleic acid sequence 5'-NGC-3' or a PAM nucleic acid sequence comprising 5'-NGC-3'. In one embodiment of the method, the modified SpCas9 is a Cas9 variant listed in Table 1, or wherein the modified SpCas9 is spCas9-MQKFRAER. In another embodiment of the method, the modified SpCas9 is a Figures 3A to 3C or Figure 10 In one embodiment of the method, the SpCas9 variant comprises a combination of amino acid sequence substitutions selected from the group consisting of:
[0044] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R (224SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225SpCas9);
[0045] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226SpCas9);
[0046] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q (227Cas9);
[0047] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230SpCas9);
[0048] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R , D1332, R1335N and T1337 (244SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E and T1337 (245SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q and T1337R (259SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGC Rd1 SpCas9); or
[0049] D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E and T1337R (267 (NGCRd2 SpCas9). In one embodiment of the method, the polynucleotide programmable DNA binding domain is a nuclease-inactive or nickase variant. In one embodiment, the nickase variant comprises the amino acid substitution D10A or its corresponding amino acid substitution. In one embodiment of the method, the adenosine deaminase domain is capable of deaminating adenosine in deoxyribonucleic acid (DNA), and the cytidine deaminase domain is capable of deaminating cytosine in deoxyribonucleic acid (DNA). In one embodiment, the adenosine deaminase is a modified adenosine deaminase that does not exist in nature. In one embodiment, the adenosine deaminase is a TadA selected from the following: Deaminase: TadA*7.10, TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA* 8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA* 8.22, TadA*8.23, or TadA*8.24. In another embodiment of this method, the deaminase domain is a cytidine deaminase selected from one or more of the following: BE4; rAPOBEC1; PpAPOBEC1; PpAPOBEC1 containing an H122A substitution; AmAPOBEC1; SsAPOBEC2; RrA3F; RrA3F containing an F130L substitution; a BE4 variant in which APOBEC-1 is replaced with a rAPOBEC1 sequence; a BE4 variant in which APOBEC-1 is replaced with an AmAPOBEC1 sequence; a BE4 variant in which A POBEC-1 is replaced with an SsAPOBEC2 sequence; a BE4 variant in which APOBEC-1 is replaced with a PpAPOBEC1 sequence; or a BE4 variant in which APOBEC-1 is replaced with a PpAPOBEC1 sequence containing an H122A substitution. In one embodiment, the PpAPOBEC1 sequence containing an H122A substitution, or the BE4 variant in which APOBEC-1 is replaced with a PpAPOBEC1 sequence containing an H122A substitution, further comprises one or more amino acid mutations selected from R33A, W90F, K34A, R52A, H121A, or Y120F.In one embodiment of the method, one or more guide polynucleotides comprise CRISPR RNA (crRNA) and trans-encoded small RNA (tracrRNA), wherein the crRNA comprises a nucleic acid sequence complementary to the SBDS nucleic acid sequence of the SDS-related gene conversion. In one embodiment of the method, the base editor and the one or more guide polynucleotides form a complex in the cell. In one embodiment of the method, the base editor is compounded with a single guide RNA (sgRNA), and the sgRNA comprises a nucleic acid sequence complementary to the SBDS nucleic acid sequence comprising the SDS-related gene conversion.
[0050] In another aspect, a guide RNA is provided, comprising a nucleic acid sequence from 5' to 3', or a 1, 2, 3, 4 or 5 nucleotide 5' truncated fragment thereof, selected from one or more of the following: GUAAGCAGGCGGGUAACAGC, AGCAGGCGGGUAACAGCUGC, GCGGGUAACAGCUGCAGCAU, UGUAAAUGUUUCCUAAGGUC, AAUGUUUCCUAAGGUCAGGU, GCAGGCGGGUAACAGCUGC, CAGGCGGGUAACAGCUGC, AGGCGGGUAACAGCUGC and AAGCAGGCGGGUAACAGCUGC.
[0051] In another aspect, a base editor system for editing a pathogenic mutation in an SBDS gene is provided, wherein the base editor system comprises:
[0052] (a) a base editor comprising:
[0053] (i) a polynucleotide programmable DNA binding domain, and
[0054] (ii) capable of converting the polynucleotide present in the SBDS gene or its complementary nuclease
[0055] a deaminase domain for base deamination; and
[0056] (b) a guide polynucleotide that cooperates with a polynucleotide programmable DNA binding domain, wherein the guide polynucleotide targets the base editor to a target polynucleotide sequence, at least a portion of which is located in an SBDS gene, an SBDS pseudogene, or a reverse complement thereof;
[0057] Deamination of the polynucleotide or its complementary nucleobase enables transcription of the SBDS gene.
[0058] In another aspect, a base editor system for editing a mutation that causes abnormal splicing in a gene is provided, wherein the base editor system comprises:
[0059] (a) a base editor comprising:
[0060] (i) a polynucleotide programmable DNA binding domain, and
[0061] (ii) a deamination domain capable of deaminating the mutation causing aberrant splicing or its complementary nucleobase; and
[0062] (b) a guide polynucleotide that cooperates with a polynucleotide programmable DNA binding domain, wherein the guide polynucleotide targets the base editor to a target polynucleotide sequence, at least a portion of which is located in the gene or its reverse complement sequence;
[0063] Deamination of the mutation or its complementary nucleobase permits transcription.
[0064] In another aspect, a method is provided for editing a pathogenic mutation in a gene that causes aberrant splicing, wherein the method comprises:
[0065] A target nucleotide sequence is contacted with a base editor, at least a portion of which is located in the gene or its reverse complement, the base editor comprising:
[0066] (i) a polynucleotide programmable DNA binding domain that cooperates with a guide polynucleotide that targets the base editor to a target polynucleotide sequence, at least a portion of which is located in the gene or its reverse complement, and
[0067] (ii) a deaminase domain capable of deaminating the pathogenic mutation or its complementary nucleobase that causes aberrant splicing; and
[0068] editing a pathogenic mutation by deaminating the pathogenic mutation or its complementary nucleobase upon targeting the base editor to a target nucleotide sequence,
[0069] Deamination of the pathogenic mutation or its complementary nucleobase results in conversion of the pathogenic mutation to a sequence that permits splicing, thereby correcting the pathogenic mutation.
[0070] In another aspect, a method for editing a pathogenic mutation in an SBDS gene is provided, wherein the method comprises:
[0071] A target nucleotide sequence is contacted with a base editor, at least a portion of which is located in the gene or its reverse complement, the base editor comprising:
[0072] (i) a polynucleotide programmable DNA binding domain that cooperates with a guide polynucleotide that targets the base editor to a target polynucleotide sequence, at least a portion of which is located in the gene or its reverse complement, and
[0073] (ii) a deaminase domain capable of deaminating the pathogenic mutation or its complementary nucleobase; and
[0074] editing a pathogenic mutation by deaminating the pathogenic mutation or its complementary nucleobase upon targeting the base editor to a target nucleotide sequence,
[0075] Wherein deamination of the pathogenic mutation or its complementary nucleobase allows splicing, thereby editing the pathogenic mutation in the SBDS gene. In one embodiment of the above-mentioned method for editing pathogenic mutations, the pathogenic gene in SBDS causes gene conversion. In one embodiment, the pathogenic mutation introduces a stop codon or changes the splicing of the gene. In one embodiment, the pathogenic mutation encodes a polypeptide with truncation. In one embodiment, the base editor introduces a missense mutation, inserts a new splice acceptor or splice donor site, or corrects a splice acceptor or splice donor site containing a mutation. In one embodiment, the base editor corrects a splice donor SNP site in the SBDS gene containing a mutation in rs113993993C→T.
[0076] In another aspect, a method of treating SDS by editing a pathogenic mutation in an SBDS gene is provided, wherein the method comprises:
[0077] administering a clip editor or a polynucleotide encoding the clip editor to a subject in need thereof, wherein the base editor comprises:
[0078] (i) a polynucleotide programmable DNA binding domain, and
[0079] (ii) a deaminase domain capable of deaminating the nucleobase within the pathogenic mutation or its complementary nucleobase; and
[0080] administering a guide polynucleotide to the subject, wherein the guide polynucleotide targets the base editor to a target nucleotide sequence, at least a portion of which is located in the gene or its reverse complement; and
[0081] The method edits the pathogenic mutation in the SBDS gene by deaminating the pathogenic mutation or its complementary nucleobase upon targeting the base editor to the target nucleotide sequence.
[0082] wherein deamination of the pathogenic mutation or its complementary nucleobase allows transcription or correction of the pathogenic mutation.
[0083] In another aspect, a method is provided for producing a cell, tissue, or organ for treating SDS in a subject in need thereof by correcting a pathogenic mutation in the SBDS gene of the cell, tissue, or organ, wherein the method comprises:
[0084] contacting the cell, tissue, or organ with a base editor comprising:
[0085] (i) a polynucleotide programmable DNA binding domain, and
[0086] (ii) a deaminase domain capable of deaminating the pathogenic mutation or its complementary nucleobase; and
[0087] contacting the cell, tissue or organ with a guide polynucleotide, wherein the guide polynucleotide targets the base editor to a target nucleotide sequence, at least a portion of which is located in the gene or its reverse complement; and
[0088] The pathogenic mutation is edited by deaminating the mutation or its complementary nucleobase upon targeting the base editor to the target nucleotide sequence.
[0089] Wherein the pathogenic mutation or its complementary nucleobase deamination allows splicing, thereby producing the cell, tissue or organ for treating SDS. In one embodiment, the mutation causes gene conversion. In another embodiment, the mutation associated with Shu-Dieter syndrome introduces a stop codon or changes the splicing of the gene. In another embodiment of the method, the mutation associated with Shu-Dieter syndrome (SDS) encodes a truncated SBDS polypeptide. In another embodiment, the base editor introduces a missense mutation, inserts a new splice acceptor or splice donor site, or corrects a splice acceptor or splice donor site containing a mutation. In another embodiment, the method further includes administering cells, tissues or organs to the subject. In one embodiment, the cell, tissue or organ is autologous, allogeneic or xenogeneic for the subject. In another embodiment of the method, the deaminase domain is a cytidine deaminase domain or an adenosine deaminase domain. In one embodiment, the adenosine deaminase domain is capable of deaminating adenine in deoxyribonucleic acid (DNA), and the cytidine deaminase is capable of deaminating cytosine in DNA.
[0090] In one embodiment of any of the above-mentioned base editor systems, or editing methods or treatment methods, or embodiments thereof, the guide polynucleotide comprises ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). In one embodiment of any of the above-mentioned base editor systems, or editing methods or treatment methods, or embodiments thereof, the guide polynucleotide comprises CRISPR RNA (crRNA), a reverse-activated CRISPR RNA (tracrRNA) sequence, or a combination thereof, wherein the crRNA comprises a nucleic acid sequence complementary to an SBDS nucleic acid sequence comprising an SDS-related change. In one embodiment of any of the above-mentioned base editor systems, or editing methods or treatment methods, or embodiments thereof, the base editor system or method further comprises a second guide polynucleotide. In one embodiment, the second guide polynucleotide comprises ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). In another embodiment, the second guide polynucleotide comprises a CRISPR RNA (crRNA) sequence, a reverse-activated CRISPR RNA (tracrRNA) sequence, or a combination thereof. In one embodiment of any of the above-mentioned base editor systems, or editing methods or treatment methods, or embodiments thereof, the polynucleotide programmable DNA binding domain is nuclease-dead or a nickase. In one embodiment of any of the above-mentioned base editor systems, or editing methods or treatment methods, or embodiments thereof, the polynucleotide programmable DNA binding domain comprises a Cas9 domain. In one embodiment, the Cas9 domain comprises a nuclease-dead Cas9 (dCas9), a Cas9 nickase (nCas9), or a Cas9 with an inactive nuclease. In some embodiments, the Cas9 domain comprises a Cas9 nickase. In one embodiment of any of the above-mentioned base editor systems, or editing methods or treatment methods, or embodiments thereof, the polynucleotide programmable DNA binding domain is an engineered or modified polynucleotide programmable DNA binding domain. In one embodiment of any of the above-mentioned base editor systems, or editing methods, or treatment methods, or embodiments thereof, editing results in less than 20% indel formation, less than 15% indel formation, less than 10% indel formation, less than 5% indel formation, less than 4% indel formation, less than 3% indel formation, less than 2% indel formation, less than 1% indel formation, less than 0.5% indel formation, or less than 0.1% indel formation. In one embodiment of any of the above base editor systems, or editing methods or treatment methods, or embodiments thereof, the editing does not result in a translocation. In one embodiment of any of the above base editor systems, or editing methods or treatment methods, or embodiments thereof, the base editor corrects a splice donor SNP site comprising a mutation in rs113993993C→T in the SBDS gene.
[0091] In another aspect, a method of treating Schwann-Diesel syndrome (SDS) in a subject in need thereof is provided, wherein the method comprises administering to the subject a cell of the above aspects and embodiments thereof.
[0092] In one embodiment of any of the above methods or embodiments thereof, the above cells or embodiments thereof, or the above base editor systems and embodiments thereof, or the above methods of editing, treating, producing cells, tissues, etc. and embodiments thereof, the base editor and / or its components are encoded by mRNA. In any of the above methods or embodiments thereof, the above cells or embodiments thereof, or the above base editor systems and embodiments thereof, or the above methods of editing, treating, producing cells, tissues, etc., the base editor is complexed with a single guide RNA (sgRNA) comprising a nucleic acid sequence complementary to an SBDS nucleic acid sequence. In one embodiment, the sgRNA comprises a nucleic acid sequence comprising at least 10 consecutive nucleotides complementary to an SBDS nucleic acid sequence. In another embodiment, the sgRNA comprises a nucleic acid sequence comprising 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39 or 40 consecutive nucleotides complementary to an SBDS nucleic acid sequence. In another embodiment, the sgRNA comprises a nucleic acid sequence comprising at least 18, 19, or 20 contiguous nucleotides that are complementary to an SBDS nucleic acid sequence.
[0093] On the other hand, a composition is provided, wherein the composition comprises a base editor bound to a guide RNA, wherein the guide RNA comprises a nucleic acid sequence that is complementary to a SBDS gene associated with Shu-Dai syndrome (SDS). In one embodiment, the base editor comprises an adenosine deaminase or a cytidine deaminase. In one embodiment, the adenosine deaminase is capable of deaminating adenine in deoxyribonucleic acid (DNA). In one embodiment, the adenosine deaminase is a TadA deaminase selected from one or more of the following: TadA*7.10, TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23, or TadA*8.24. In one embodiment, the cytidine deaminase is capable of deaminating cytidine in deoxyribonucleic acid (DNA). In another embodiment, the cytidine deaminase is APOBEC, A3F, or a derivative thereof. In one embodiment of the composition, the base editor
[0094] (i) comprising a Cas9 nickase;
[0095] (ii) Cas9 containing an inactive nuclease;
[0096] (iii) comprising a SpCas9 variant comprising Figures 3A to 3C or Figure 10 Combinations of amino acid substitutions shown in ;
[0097] (iv) comprising a SpCas9 variant comprising a combination of amino acid sequence substitutions selected from the group consisting of:
[0098] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R (224SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225SpCas9);
[0099] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226SpCas9);
[0100] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q (227Cas9);
[0101] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230SpCas9);
[0102] D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R , D1332, R1335N and T1337 (244SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E and T1337 (245SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q and T1337R (259SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGC Rd1 SpCas9); or
[0103] D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E, and T1337R (NGCRd2 SpCas9).
[0104] (v) does not contain a UGI domain; and / or
[0105] (vi) a cytidine deaminase comprising one or more of the following: BE4; rAPOBEC1; PpAPOBEC1; PpAPOBEC1 containing an H122A substitution; AmAPOBEC1; SsAPOBEC2; RrA3F; RrA3F containing an F130L substitution; a BE4 variant in which APOBEC-1 is replaced with a rAPOBEC1 sequence; a BE4 variant in which APOBEC-1 is replaced with an AmAPOBEC1 sequence; a BE4 variant in which APOBEC-1 is replaced with an SsAPOBEC2 sequence; a BE4 variant in which APOBEC-1 is replaced with a PpAPOBEC1 sequence; or a BE4 variant in which APOBEC-1 is replaced with a PpAPOBEC1 sequence containing an H122A substitution. In one embodiment of the composition, in (vi), the PpAPOBEC1 sequence comprising the H122A substitution, or the BE4 variant in which APOBEC-1 is replaced with the PpAPOBEC1 sequence comprising the H122A substitution, further comprises one or more amino acid mutations selected from R33A, W90F, K34A, R52A, H121A, or Y120F. In one embodiment, the composition further comprises a pharmaceutically acceptable excipient, diluent, or carrier.
[0106] On the other hand, a pharmaceutical composition for treating Shu-Dieter syndrome (SDS) is provided, wherein the pharmaceutical composition comprises the composition of the above aspects and embodiments, and comprises a pharmaceutically acceptable excipient, diluent or carrier. In one embodiment of the pharmaceutical composition, the gRNA and the base editor are formulated together or separately. In one embodiment of the pharmaceutical composition, the gRNA comprises a nucleic acid sequence from 5' to 3', or a 1, 2, 3, 4 or 5 nucleotide 5' truncated fragment thereof, selected from one or more of the following: GUAAGCAGGCGGGUAACAGC, AGCAGGCGGGUAACAGCUGC, GCGGGUAACAGCUGCAGCAU, UGUAAAUGUUUCCUAAGGUC, AAUGUUUCCUAAGGUCAGGU, GCAGGCGGGUAACAGCUGC, CAGGCGGGUAACAGCUGC, AGGCGGGUAACAGCUGC and AAGCAGGCGGGUAACAGCUGC. In one embodiment, the pharmaceutical composition further comprises a vector suitable for expression in mammalian cells, wherein the vector comprises a polynucleotide encoding the base editor. In one embodiment of the pharmaceutical composition, the polynucleotide encoding the base editor is mRNA. In one embodiment of the pharmaceutical composition, the vector is a viral vector. In one embodiment, the viral vector is a retroviral vector, an adenoviral vector, a lentiviral vector, a herpes virus vector or an adeno-associated virus vector (AAV). In one embodiment, the pharmaceutical composition further comprises ribonucleic acid particles suitable for expression in mammalian cells.
[0107] In one aspect, a pharmaceutical composition is provided, wherein the pharmaceutical composition comprises (i) a nucleic acid encoding a base editor, and (ii) a guide RNA of the above aspects, such as the following guide RNA, which comprises a nucleic acid sequence from 5' to 3', or a 1, 2, 3, 4 or 5 nucleotide 5' truncated fragment thereof, selected from one or more of the following:
[0108] GUAAGCAGGCGGGUAACAGC, AGCAGGCGGGUAACAGCUGC, GCGGGUAACAGCUGCAGCAU, UGUAAAUGUUUCCUAAGGUC, AAUGUUUCCUAAGGUCAGGU, GCAGGCGGGUAACAGCUGC, CAGGCGGGUAACAGCUGC, AGGCGGGUAACAGCUGC, and AAGCAGGCGGGUAACAGCUGC. In one embodiment of the pharmaceutical composition of any one of the above aspects or embodiments thereof, the pharmaceutical composition further comprises a lipid.
[0109] In one aspect, a method of treating Schwann-Diesel syndrome (SDS) is provided, comprising administering to a subject in need thereof a pharmaceutical composition according to any one of the above aspects and embodiments thereof.
[0110] In one aspect, a pharmaceutical composition according to any one of the above aspects and embodiments thereof is provided for use in treating Schwann-Diesel syndrome (SDS) in a subject. In one embodiment of the use, the subject is a human.
[0111] definition
[0112] The following definitions are supplemental to those in the art and are specific to this application and do not apply to any related or unrelated cases, such as any duplicate granted patents or applications. Although any methods and materials similar or equivalent to those described herein can be used to practice or test the present disclosure, preferred materials and methods are as described herein. Accordingly, the terms used herein are for the purpose of describing specific embodiments only and are not intended to be limiting.
[0113] Unless otherwise defined, all scientific and technological terms used herein have the meanings commonly understood by those skilled in the art to which the present invention belongs. The following references provide a general definition of many of the terms used in the present invention to the technician: "Microbiology and Molecular Biology Dictionary (Second Edition)" (Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed.1994)); "Cambridge Dictionary of Science and Technology" (the Cambridge Dictionary of Science and Technology (Walker ed., 1988)); "Hereditary Professional Dictionary (Fifth Edition)" (The Glossary of Genetics, 5th Ed., R.Rieger et al. (eds.), Springer Verlag (1991)); and "Harper Collins Biology Dictionary" (Hale & Marham, The Harper Collins Dictionary of Biology (1991)). As used herein, unless explicitly specified, the following terms have the meanings described below.
[0114] In this application, unless specifically indicated otherwise, the use of the singular includes the plural. It must be noted that, as used in this specification, the singular forms "a," "an," and "the" include plural referents unless the context clearly indicates otherwise. In this application, the use of "or" means "and / or" unless otherwise specified. Furthermore, the use of the term "include" and other forms, such as verb forms, active or passive forms, is not limiting.
[0115] As used in this specification and claims, the words "comprising" (and any form of comprising, such as singular and plural forms), "having" (and any form of having, such as singular and plural forms), "including" (and any form of including, such as singular and plural forms), or "containing" (and any form of containing, such as singular and plural forms) are inclusive or open-ended and do not exclude additional, uncited elements or method steps. It is contemplated that any embodiment discussed in this specification can be implemented with respect to any method or composition of the present disclosure, and vice versa. In addition, the compositions of the present disclosure can be used to implement the methods of the present disclosure.
[0116] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within 1 or more than 2 standard deviations, depending on the practice in the art. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or organisms, the term can mean within an order of magnitude, e.g., within 5-fold, within 2-fold, of a value. When specific values are described in the application and claims, unless otherwise specified, it should be assumed that the term "about" means within an acceptable error range for the specific value.
[0117] References in the specification to "some embodiments," "an embodiment," "one embodiment," or "other embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least some embodiments of the present disclosure, but not necessarily all embodiments.
[0118] "Adenosine deaminase" means a polypeptide or fragment thereof that can catalyze the hydrolytic deamination of adenine or adenosine. In some embodiments, the deaminase or deaminase domain catalyzes the hydrolytic deamination of adenosine to inosine or deoxyadenosine to deoxyinosine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminase enzymes provided herein (e.g., engineered adenosine deaminase enzymes, evolved adenosine deaminase enzymes) can be from any organism such as bacteria.
[0119] In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism. In some embodiments, the adenosine deaminase or deaminase domain does not exist in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% identical to a naturally occurring deaminase. In some embodiments, the adenosine deaminase is from bacteria such as Escherichia coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, or C. crescentus. In some embodiments, the adenosine deaminase is a TadA deaminase. In some embodiments, the TadA deaminase is an Escherichia coli TadA (ecTadA) deaminase or a fragment thereof.
[0120] In some embodiments, the adenosine deaminase comprises an alteration in the following sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (also known as TadA*7.10).
[0121] In some embodiments, TadA*7.10 comprises an alteration at amino acid 82 or 166. In specific embodiments, variants of the above sequences comprise one or more of the following alterations: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and Q154R. The alteration Y123H refers to the reversal of H123Y in TadA*7.10 to Y123H TadA (wt). In other embodiments, the variant of the TadA*7.10 sequence comprises a combination of changes selected from the group consisting of: Y147R+Q154R+Y123H, Y147R+Q154R+I76Y, Y147R+Q154R+T166R, Y147T+Q154R, Y147T+Q154S, V82S+Q154S and Y123H+Y147R+Q154R+I76Y.
[0122] In other embodiments, the present invention provides adenosine deaminase enzymes comprising deletions, e.g., TadA*8 comprising a C-terminal deletion beginning at residues 149, 150, 151, 152, 153, 154, 155, 156, or 157. In other embodiments, the adenosine deaminase variant is a TadA monomer (e.g., TadA*8) comprising one or more of the following alterations: Y147T, Y147R, Q154S, Y123H, V82S, T166R, Q154R. In other embodiments, the adenosine deaminase variant is a monomer comprising one or more of the following alterations: Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y147T+Q154R; Y147T+Q154S; V82S+Q154S; and Y123H+Y147R+Q154R+I76Y. In yet other embodiments, the adenosine deaminase variant is a homodimer comprising two adenosine deaminase domains, each of which comprises one or more of the following alterations: Y147T, Y147R, Q154S, Y123H, V82S, T166R, Q154R. In other embodiments, the adenosine deaminase variant is a heterodimer comprising a wild-type adenosine deaminase domain or a TadA*7.10 domain and an adenosine deaminase variant domain (e.g., TadA*8) comprising one or more of the following changes: Y147T, Y147R, Q154S, Y123H, V82S, T166R, Q154R. In other embodiments, the adenosine deaminase variant is a heterodimer comprising a TadA*7.10 domain and an adenosine deaminase variant of TadA*7.10 (e.g., TadA*8) comprising the following changes: Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y147T+Q154R; Y147T+Q154S; V82S+Q154S; and Y123H+Y147R+Q154R+I76Y.
[0123] In one embodiment, the adenosine deaminase comprises TadA*8 comprising or consisting essentially of the following sequence, or a fragment thereof having adenosine deaminase activity:
[0124] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD.
[0125] In some embodiments, TadA*8 is truncated. In some embodiments, the truncated TadA*8 lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to full-length TadA*8. In some embodiments, the truncated TadA*8 lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to full-length TadA*8. In some embodiments, the adenosine deaminase is full-length TadA*8.
[0126] In certain embodiments, the adenosine deaminase heterodimer comprises a TadA*8 domain and an adenosine deaminase domain selected from:
[0127] Staphylococcus aureus (S. aureus) TadA:
[0128] MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN
[0129] Bacillus subtilis (B. subtilis) TadA:
[0130] MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLSE
[0131] Salmonella typhimurium (S. typhimurium) TadA:
[0132] MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFRMRRQEIKALKKADRAEGAGPAV
[0133] Shewanella putrefaciens (S. putrefaciens) TadA:
[0134] MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE
[0135] Haemophilus influenzae F3031 (H. influenzae) TadA:
[0136] MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTΑΗAEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRREEKKIEKALLKSLSDK
[0137] Caulobacter crescentus (C. crescentus) TadA:
[0138] MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLTDLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAKI
[0139] Geobacter sulfurreducens (G.sulfurreducens) TadA:
[0140] MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRKKAKATPALF IDERKVPPEPTadA*7.10
[0141] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD.
[0142] As used herein, "administering" refers to providing one or more compositions described herein to a patient or subject. For example, but not limited to, administration (e.g., injection) of a composition can be performed by intravenous (iv) injection, subcutaneous (sc) injection, intradermal (id) injection, intraperitoneal (ip) injection, or intramuscular (im) injection. One or more such routes can be used. Parenteral administration can be performed, for example, by depot injection or by gradual infusion over time. Alternatively or concurrently, administration can be performed by oral route.
[0143] "Agent" means any small molecule chemical compound, antibody, nucleic acid molecule or polypeptide, or fragments thereof.
[0144] "Alteration" means a change (increase or decrease) in the sequence, expression level, or activity of a gene or polypeptide as detected by standard methods known in the art, such as those described herein. As used herein, an alteration includes a 10% change in expression level, a 25% change in expression level, a 40% change, a 50% change, or greater.
[0145] "Relieve" means to reduce, suppress, attenuate, eliminate, arrest, or stabilize the development or progression of a disease.
[0146] "Analog" means a molecule that is not identical but has similar functional or structural characteristics. For example, a polypeptide analog retains the biological activity of the corresponding naturally occurring polypeptide while possessing certain biochemical modifications that enhance the analog's function relative to the naturally occurring polypeptide. Such biochemical modifications may increase the analog's protease resistance, membrane permeability, or half-life without altering, for example, ligand binding. Analogs may include unnatural amino acids.
[0147] "Base editor (BE)" or "nucleobase editor (NBE)" means an agent that binds to a polynucleotide and has nucleobase modification activity. In various embodiments, the base editor comprises a nucleobase modification polypeptide (e.g., a deaminase) and a polynucleotide programmable nucleotide binding domain that act in concert with a guide polynucleotide (e.g., a guide RNA). In various embodiments, the agent is a biomolecule complex comprising a protein domain with base editing activity, i.e., a domain capable of modifying a base (e.g., A, T, C, G, or U) in a nucleic acid molecule (e.g., DNA). In some embodiments, the polynucleotide programmable DNA binding domain is fused or linked to a deaminase domain. In one embodiment, the agent is a fusion protein comprising one or more domains with base editing activity. In another embodiment, the protein domain with base editing activity is linked to a guide RNA (e.g., via an RNA binding motif on the guide RNA and an RNA binding domain fused to a deaminase). In some embodiments, the domain with base editing activity is capable of deaminating bases in nucleic acid molecules. In some embodiments, the base editor is capable of deaminating one or more bases in a DNA molecule. In some embodiments, the base editor is capable of deaminating cytosine (C) or adenosine (A) in DNA. In some embodiments, the base editor is capable of deaminating cytosine (C) and adenosine (A) in DNA. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenosine base editor (ABE). In some embodiments, the base editor is an adenosine base editor (ABE) and a cytidine base editor (CBE). In some embodiments, the base editor is a Cas9 (dCas9) fused to an inactive nuclease of adenosine deaminase. In some embodiments, Cas9 is a circular permutant Cas9 (e.g., spCas9 or saCas9). Circularly arranged Cas9 is known in the art and described in, for example, Oakes et al., Cell 176, 254–267, 2019. In some embodiments, the base editor is fused to an inhibitor of base excision repair, e.g., a UGI domain or a dISN domain. In some embodiments, the fusion protein comprises a Cas9 nickase fused to a deaminase and an inhibitor of base excision repair, such as a UGI or dISN domain. In other embodiments, the base editor is a base-free base editor.
[0148] In some embodiments, adenosine deaminase evolves from TadA. In some embodiments, the polynucleotide programmable DNA binding domain is a CRISPR-related (e.g., Cas or Cpf1) enzyme. In some embodiments, the base editor is a catalytically dead Cas9 (dCas9) fused to a deaminase domain. In some embodiments, the base editor is a Cas9 nickase (nCas9) fused to a deaminase domain. In some embodiments, the base editor is fused to an inhibitor of base excision repair (BER). In some embodiments, the inhibitor of base excision repair is a uracil DNA glycosylase inhibitor (UGI). In some embodiments, the inhibitor of base excision repair is an inosine base excision repair inhibitor. The details of the base editor are described in International PCT Application Nos. PCT / 2017 / 045381 (WO2018 / 027078) and PCT / US2016 / 058344 (WO2017 / 070632), each of which is incorporated herein by reference in its entirety. See also Komor, AC, et al., "Programmable base editing of a target base in genomic DNA without double-strandedDNA cleavage" Nature 533, 420-424 (2016); Gaudelli, NM, et al., "Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage" Nature 551, 464-471 (2017); Komor, AC, et al., "Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T: A base editors with higher efficiency and product purity" Science Advances 3:eaao4774 (2017); and Rees, HA, et al., "Base editing: precision chemistry on the genome and transcriptome of living cells." Nat Rev Genet. 2018 Dec; 19(12): 770-788. doi: 10.1038 / s41576-018-0059-1, the entire contents of which are incorporated herein by reference.
[0149] In some embodiments, a base editor (e.g., ABE8) is generated by cloning an adenosine deaminase variant (e.g., TadA*8) into a scaffold comprising a circularly arranged Cas9 (e.g., spCAS9) and a bipartite nuclear localization sequence. Circularly arranged Cas9 is known in the art and is described in, for example, Oakes et al., Cell 176, 254–267, 2019. Exemplary circularly arranged sequences are described in detail below, where bold sequences represent sequences derived from Cas9, italic sequences represent linker sequences, and underlined sequences represent bipartite nuclear localization sequences.
[0150] CP5 (where MSP "NGC = variant with mutation of conventional Cas9 Pam such as NGG" PID = protein interaction domain, and "D10A" nickase):
[0151]
[0152] In some embodiments, ABE8 is selected from a base editor from Table 7 below. In some embodiments, ABE8 contains an adenosine deaminase evolved from TadA. In some embodiments, the adenosine deaminase variant of ABE8 is a TadA*8 variant as described in Table 7 below. In some embodiments, the adenosine deaminase is TadA*7.10 comprising one or more of the changes selected from the group consisting of: Y147T, Y147R, Q154S, Y123H, V82S, T166R, Q154R. In various embodiments, ABE8 comprises TadA*7.10 with an alteration selected from the group consisting of: Y147R+Q154R+Y123H, Y147R+Q154R+I76Y, Y147R+Q154R+T166R, Y147T+Q154R; Y147T+Q154S, V82S+Q154S, and Y123H+Y147R+Q154R+I76Y. In some embodiments, ABE8 is a monomeric construct.
[0153] In some embodiments, ABE8 is a heterodimer construct. In some embodiments, the ABE8 base editor comprises the sequence:
[0154] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD
[0155] For example, the adenosine base editor ABE to be used in the base editing compositions, systems, and methods described herein has a nucleic acid sequence (8877 base pairs) as described below (Addgene, Watertown, MA.; Gaudelli NM, et al., Nature. 2017 Nov 23; 551(7681): 464-471. doi: 10.1038 / nature24644; Koblan LW, et al., Nat Biotechnol. 2018 Oct; 36(9): 843-846. doi: 10.1038 / nbt.4172.). Polynucleotide sequences having at least 95% or higher identity to the ABE nucleic acid sequence are also encompassed.
[0156]
[0157] For example, the cytidine base editor (CBE) used in the base editing compositions, systems, and methods described herein has the following nucleic acid sequence (8877 base pairs) provided below (Addgene, Watertown, MA; Komor AC, et al., 2017, Sci Adv., 30; 3(8):eaao4774.doi:10.1126 / sciadv.aao4774). Polynucleotide sequences having at least 95% or greater identity to the BE4 nucleic acid sequence are also encompassed.
[0158]
[0159]
[0160]
[0161]
[0162]
[0163] In some embodiments, the cytidine base editor is BE4 having a nucleic acid sequence selected from one of the following:
[0164] Original BE4 nucleic acid sequence:
[0165]
[0166] BE4 codon optimized 1 nucleic acid sequence:
[0167]
[0168] BE4 codon optimized 2 nucleic acid sequence:
[0169]
[0170] "Base editing activity" means acting to chemically change a base within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is cytidine deaminase activity, for example, converting the target C·G to T·A. In another embodiment, the base editing activity is adenosine or adenine deaminase activity, for example, converting the target A·T to G·C. In another embodiment, the base editing activity is cytidine deaminase activity, for example, converting the target C·G to T·A, and adenosine or adenine deaminase activity, for example, converting the target A·T to G·C.
[0171] The term "base editor system" refers to a system for editing the nucleobases of a target nucleotide sequence. In various embodiments, the base editor (BE) system comprises (1) a polynucleotide programmable nucleotide binding domain, a deaminase domain and a cytidine deaminase domain for deaminating the nucleobases in the target nucleotide sequence; and (2) one or more guide polynucleotides (e.g., guide RNAs) that act in concert with the polynucleotide programmable nucleotide binding domain. In various embodiments, the base editor (BE) system comprises a nucleobase editor domain selected from adenosine deaminase and cytidine deaminase, and a domain with nucleic acid sequence-specific binding activity. In some embodiments, the base editor system comprises (1) a base editor (BE) comprising a polynucleotide programmable DNA binding domain and a deaminase domain for deaminating one or more nucleobases in the target nucleotide sequence; and (2) one or more guide RNAs that act in concert with the polynucleotide programmable DNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable DNA binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE) or a cytidine base editor (CBE).
[0172] The term "Cas9" or "Cas9 domain" refers to an RNA-guided nuclease or fragment thereof that comprises a Cas9 protein (e.g., a protein comprising an active, inactive, or partially active DNA cleavage domain of Cas9 and / or a gRNA binding domain of Cas9). Cas9 nucleases are sometimes referred to as casn1 nucleases or CRISPR (clustered regularly interspaced short palindromic repeats)-related nucleases. An exemplary Cas9 is Streptococcus pyogenes Cas9 (spCas9), the amino acid sequence of which is provided below:
[0173]
[0174] (Single underline: HNH domain; double underline: RuvC domain)
[0175] The term "conservative amino acid substitution" or "conservative mutation" refers to the replacement of one amino acid with another amino acid having a common property. A functional approach to defining the common properties between individual amino acids is to analyze the normalized frequency of amino acid changes between corresponding proteins of homologous organisms (Schulz, GE and Schirmer, RH, Principles of Protein Structure, Springer-Verlag, New York (1979)). Based on such analysis, multiple groups of amino acids can be defined in which the amino acids within the group are preferentially exchanged with each other and are therefore most similar to each other in terms of their impact on the overall protein structure (Schulz, GE and Schirmer, RH, supra). Non-limiting examples of conservative mutations include amino acid replacements of amino acids, for example, lysine replaces arginine, and vice versa, so that a positive charge can be maintained; glutamic acid replaces aspartic acid, and vice versa, so that a negative charge can be maintained; serine replaces threonine, so that free –OH can be maintained; and glutamine replaces asparagine, so that free –NH2 can be maintained.
[0176] The terms "coding sequence" or "protein coding sequence," used interchangeably herein, refer to a polynucleotide segment that encodes a protein. This region or sequence has a start codon near the 5' end and a stop codon near the 3' end. Stop codons that can be used with the base editors described herein include the following:
[0177] Glutamine CAG→TAG stop codon
[0178] CAA→TAA
[0179] Arginine CGA→TGA
[0180] Tryptophan TGG→TGA
[0181] TGG→TAG
[0182] TGG→TAA
[0183] A coding sequence may also be referred to as an open reading frame.
[0184] "Cytidine deaminase" means a polypeptide or fragment thereof that can catalyze a deamination reaction that converts an amino group to a carbonyl group. In one embodiment, the cytidine deaminase converts cytosine to uracil or converts 5-methylcytosine to thymine. PmCDA1 (sea lamprey cytosine deaminase 1, "PmCDA1") from sea lamprey (Petromyzon marinus), AID (activation-induced cytidine deaminase; AICDA) and APOBEC from mammals, or mammals of different species (e.g., humans, pigs, cattle, horses, monkeys, etc.), as well as non-mammals such as alligators are exemplary cytidine deaminases.
[0185] As used herein, the term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase, which catalyzes the hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. In some embodiments, the deaminase or deaminase domain is a cytosine deaminase, which catalyzes the hydrolytic deamination of cytosine to uracil. In some embodiments, the deaminase is an adenosine deaminase, which catalyzes the hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase is an adenosine deaminase, which catalyzes the hydrolytic deamination of adenosine or adenine (A) to inosine (I). In some embodiments, the deaminase or deaminase domain is an adenosine deaminase, which catalyzes the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, adenosine deaminase catalyzes the hydrolytic deamination of adenosine in deoxyribonucleic acid (DNA). Adenosine deaminase provided herein (e.g., engineered adenosine deaminase, evolved adenosine deaminase) can be from any organism such as bacteria. In some embodiments, adenosine deaminase is from bacteria, such as Escherichia coli (E. coli), Staphylococcus aureus (S. Aureus), Salmonella typhi (S. typhi), Shewanella putrefaciens (S. putrefaciens), Haemophilus influenzae (H. influenzae) or C. crescentus (C. crescentus). In some embodiments, adenosine deaminase is TadA deaminase. In some embodiments, deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism such as human, chimpanzee, gorilla, monkey, cattle, dog, rat or mouse. In some embodiments, adenosine deaminase or deaminase domain does not exist in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical to a naturally occurring deaminase.
[0186] "Detection" refers to identifying the presence, absence, or amount of an analyte to be detected. In one embodiment, sequence changes in a polynucleotide or polypeptide are detected. In another embodiment, the presence of an indel is detected.
[0187] "Detectable label" means a composition that is linked to a molecule of interest so that the molecule can be detected by spectroscopic, photochemical, biochemical, immunochemical, or chemical means. For example, useful labels include radioactive isotopes, magnetic beads, metallic beads, colloidal particles, fluorescent dyes, electron-dense reagents, enzymes (e.g., as commonly used in enzyme-linked immunosorbent assays (ELISAs)), biotin, digoxigenin, or haptens.
[0188] "Disease" means any condition or disorder that damages or interferes with the normal function of a cell, tissue, or organ. In certain embodiments, the disease amenable to treatment with the compositions of the present invention is associated with abnormal splicing. In a specific embodiment, the disease is Schroeder-Diesel syndrome (SDS).
[0189] By "disease associated with aberrant splicing" is meant a condition or disorder associated with disruption of transcription resulting from alterations in the gene sequence that affect splicing, such as alterations in splice acceptor or splice donor sites.
[0190] “Effective amount” means the amount of an agent or active compound, such as a base editor, described herein, which is required to alleviate the symptoms of a disease relative to an untreated patient or a disease-free individual (ie, a healthy individual); or the amount of an agent or active compound, which is sufficient to cause a desired biological response. The effective amount of the active compound used to practice the present invention for therapeutic treatment of a disease varies with the mode of administration and the age, weight, and general health of the subject. Ultimately, the attending physician or veterinarian will determine the appropriate amount and dosage regimen. This amount is referred to as an “effective” amount. In one embodiment, the effective amount is the amount of a base editor of the present invention that is sufficient to introduce a change into a gene of interest in a cell (eg, a cell in vitro or a cell in vivo). In one embodiment, the effective amount is the amount of the base editor required to achieve a therapeutic effect. Such therapeutic effects do not have to be sufficient to change the pathogenic genes in all cells of a subject, tissue, or organ, but deliberately only change the pathogenic genes in about 1%, 5%, 10%, 25%, 50%, 75% or more cells present in the subject, tissue, or organ. In one embodiment, the effective amount is sufficient to alleviate one or more symptoms of the disease.
[0191] In some embodiments, an effective amount of a base editor comprising an nCas9 domain and a deaminase domain (e.g., adenosine deaminase, cytidine deaminase) that can be in the form of a fusion protein provided herein, or an agent or composition comprising the base editor comprising an nCas9 domain and a deaminase domain (e.g., adenosine deaminase, cytidine deaminase) is sufficient to induce editing of a target site specifically bound and edited by the base editor described herein. As will be appreciated by those skilled in the art, the amount of the agent (e.g., fusion protein) can vary depending on various factors, for example, depending on the desired biological response, for example, depending on the specific allele, genome, or target site to be edited, depending on the cell or tissue to be targeted, and / or depending on the agent used.
[0192] In some embodiments, an effective amount of an agent (e.g., a fusion protein comprising an nCas9 domain and a deaminase domain) that can be present in the form of a fusion protein can refer to an amount of the agent (e.g., a fusion protein) that is sufficient to induce editing of a target site that is specifically bound and edited by the fusion protein. As will be appreciated by those skilled in the art, the amount of the agent (e.g., a fusion protein, a nuclease, a hybrid protein, a protein dimer, a complex of a protein (or protein dimer) and a polynucleotide, or a polynucleotide) can vary depending on various factors, for example, depending on the desired biological response, for example, depending on the specific allele, genome, or target site to be edited, depending on the cell or tissue to be targeted, and / or depending on the agent used.
[0193] "Fragment" means a portion of a polypeptide or nucleic acid molecule. This portion contains at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the entire length of a reference nucleic acid molecule or polypeptide. A fragment can contain 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids.
[0194] "Guide RNA" or "gRNA" means a polynucleotide that is specific for a target sequence and can form a complex with a polynucleotide programmable nucleotide binding domain protein (e.g., Cas9 or Cpf1). In one embodiment, the guide polynucleotide is a guide RNA (gRNA). The gRNA can exist as a complex of two or more RNAs, or as a single RNA molecule. A gRNA that exists as a single RNA molecule can be referred to as a single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to a guide RNA that exists as a single molecule or as a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species contains two domains: (1) a domain that shares homology with the target nucleic acid (e.g., and guides the binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence called tracrRNA and contains a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to the tracrRNA provided in Jinek et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those including domain 2) can be found in US20160208288 entitled “Switchable Cas9 Nucleases and Uses Thereof” and US 9,737,604 entitled “Delivery System For Functional Nucleases”, the entire contents of each of which are incorporated herein by reference in their entirety. In some embodiments, a gRNA comprises two or more domains (1) and (2) and can be an “extended gRNA”. The extended gRNA will bind to two or more Cas9 proteins and bind to the target nucleic acid at two or more different regions, as described herein. The gRNA comprises a nucleotide sequence complementary to the target site that mediates the binding of the nuclease / RNA complex to the target site, providing sequence specificity of the nuclease:RNA complex.
[0195] "Hybridization" means hydrogen bonding between complementary nucleobases, which can be Watson-Crick, Hoogsteen or reversed Hoogsteen hydrogen bonding. For example, adenine and thymine are complementary nucleobases that pair by forming hydrogen bonds.
[0196] By "increase" is meant a positive change of at least 10%, 25%, 50%, 75% or 100%.
[0197] The terms "inhibitor of base repair," "base repair inhibitor," "IBR," or their grammatical equivalents refer to a protein that is capable of inhibiting the activity of a nucleic acid repair enzyme (e.g., a base excision repair enzyme). In some embodiments, the IBR is an inhibitor of inosine base excision repair. Exemplary excision repair inhibitors include inhibitors of APE1, Endo III, Endo IV, Endo V, Endo VIII, Fpg, hOGG1, hNEIL1, T7 Endol, T4PDG, UDG, hSMUG1, and hAAG. In some embodiments, the base repair inhibitor is an inhibitor of Endo V or hAAG. In some embodiments, the IBR is an inhibitor of Endo V or hAAG. In some embodiments, the IBR is catalytically inactive EndoV or catalytically inactive hAAG. In some embodiments, the base repair inhibitor is catalytically inactive EndoV or catalytically inactive hAAG. In some embodiments, the base repair inhibitor is a uracil glycosylase inhibitor (UGI). UGI refers to a protein that can inhibit uracil-DNA glycosidase base excision repair enzymes. In some embodiments, the UGI domain comprises a wild-type UGI or a fragment of wild-type UGI. In some embodiments, the UGI protein provided herein includes a fragment of UGI and a protein homologous to UGI or a fragment of UGI. In some embodiments, the base repair inhibitor is an inhibitor of inosine base excision repair. In some embodiments, the base repair inhibitor is a "catalytically inactive inosine-specific nuclease" or a "dead inosine-specific nuclease." Without being bound by a particular theory, a catalytically inactive inosine glycosidase (e.g., alkyladenine glycosidase (AAG)) can bind to inosine but cannot create a baseless site or remove the inosine, thereby spatially blocking the newly formed inosine portion from the DNA damage / repair mechanism. In some embodiments, a catalytically inactive inosine-specific nuclease may be able to bind to inosine in a nucleic acid but not cut the nucleic acid. Non-limiting exemplary catalytically inactive inosine-specific nucleases include catalytically inactive alkyl adenosine glycosidase (AAG nuclease), e.g., from humans; and catalytically inactive endonuclease V (EndoV nuclease), e.g., from Escherichia coli. In some embodiments, the catalytically inactive AAG nuclease comprises an E125Q mutation or a corresponding mutation in another AAG nuclease.
[0198] An "intein" is a fragment of a protein that is able to excise itself and join the remaining fragment (an extein) with peptide bonds in a process called protein splicing. Inteins are also called "protein introns." The process by which an intein excises itself and joins the remaining part of a protein is referred to herein as "protein splicing" or "intein-mediated protein splicing." In some embodiments, the inteins of a precursor protein (a protein that contains an intein before intein-mediated protein splicing) come from two genes. Herein, such inteins are referred to as split inteins (e.g., split intein-N and split intein-C). For example, in cyanobacteria, DnaE (i.e., the catalytic subunit of DNA polymerase III) is encoded by two separate genes (i.e., dnaE-n and dnaE-c). Herein, the intein encoded by the dnaE-n gene may be referred to as "intein-N." Herein, the intein encoded by the dnaE-c gene may be referred to as "intein-C."
[0199] Other intein systems may also be used. For example, synthetic inteins based on the dnaE intein, namely the Cfa-N (e.g., split intein-N) and Cfa-C (e.g., split intein-C) intein pairs, have been described (e.g., in Stevens et al., J Am Chem Soc. 2016 Feb. 24; 138(7): 2162-5, which is incorporated herein by reference). Non-limiting examples of intein pairs that can be used in accordance with the present disclosure include: Cfa DnaE intein, Ssp GyrB intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB intein, and Cne Prp8 intein (e.g., as described in U.S. Pat. No. 8,394,604, which is incorporated herein by reference).
[0200] Nucleotide and amino acid sequences of exemplary inteins are provided.
[0201] DnaE intein-N DNA:
[0202] TGCCTGTCATACGAAACCGAGATACTGACAGTAGAATATGGCCTTCTGCCAATCGGGAAGATTGTGGAGAAACGGATAGAATGCACAGTTTACTCTGTCGATAACAATGGTAACATTTATACTCAGCCAGTTGCCCAGTGGCACGACCGGGGAGAGCAGGAAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAGGGCCACTAAGGACCACAAATTTATGACAGTCGATGGCCAGATGCTGCCTATAGACGAAATCTTTGAGCGAGAGTTGGACCTCATGCGAGTTGACAACCTTCCTAAT
[0203] DnaE intein-N protein: CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDRGEQEVFEYCLEDGSLIRATKDHKFMTVDGQMLPIDEIFERELDLMRVDNL PN
[0204] DnaE intein-C DNA: ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGATATTGGAGTCGAAAGAGATCACAACTTTGCTCTGAAGAACGGATTCATAG CTTCTAAT
[0205] Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN
[0206] Cfa-N DNA:
[0207] TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGCCTATTGGAAAGATTGTCGAAGAGAGAATTGAATGCACAGTATATACTGTAGACAAGAATGGTTTCGTTTACACACAGCCCATTGCTCAATGGCACAATCGCGGCGAACAAGAAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACGAGCAACTAAAGATCATAAATTCATGACCACTGACGGGCAGATGTTGCCAATAGATGAGATATTCGAGCGGGGCTTGGATCTCAAACAAGTGGATGGATTGCCA
[0208] Cfa-N protein:
[0209] CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQMLPIDEIFERGLDLKQVDGLP Cfa-C DNA:
[0210] ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAGGAAAGTAAAGATAATATCTCGAAAAAGTCTTGGTACCCAAAATGTCTATGATATTGGAGTGGAGAAAGATCACAACTTCCTTCTCAAGAACGGTCTCGTAGCCAGCAAC
[0211] Cfa-C protein: MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLVASN
[0212] Intein-N and intein-C can be fused to the N-terminal portion of the split Cas9 and the C-terminal portion of the split Cas9, respectively, for joining the N-terminal portion of the split Cas9 to the C-terminal portion of the split Cas9. For example, in some embodiments, intein-N is fused to the C-terminus of the N-terminal portion of the split Cas9, i.e., to form a structure of N-[N-terminal portion of the split Cas9]-[intein-N]--C. In some embodiments, intein-C is fused to the N-terminus of the C-terminal portion of the split Cas9, i.e., to form a structure of N--[intein-C]-[C-terminal portion of the split Cas9]--C. Intein-mediated protein splicing mechanisms for joining proteins to which intein peptides are fused (e.g., split Cas9) are known in the art, for example, as described in Shah et al., ChemSci. 2014; 5(1):446-461, which is incorporated herein by reference. Methods for designing and using inteins are known in the art and are described by, for example, WO2014004336, WO2017132580, US20150344549, and US20180127780, each of which is herein incorporated by reference in its entirety.
[0213] The terms "isolated," "purified," or "biologically pure" refer to a material that is free to varying degrees from components that normally accompany it (as found in its native state). "Isolated" means a degree of separation from the original source or surrounding material. "Purified" means a degree of separation greater than isolation. A "purified" or "biologically pure" protein is sufficiently free of other materials such that any impurities do not affect the biological properties of the protein at the material level or cause other negative consequences. In other words, when produced by recombinant DNA technology, a nucleic acid or peptide of the invention is purified if it is substantially free of cellular material, viral material, or culture medium; or when chemically synthesized, it is purified if it is free of chemical precursors or other chemicals. Purity and homogeneity are typically determined using analytical chemistry techniques such as polyacrylamide gel electrophoresis or high performance liquid chromatography. The term "purified" can mean that a nucleic acid or protein produces essentially one band in an electrophoretic gel. For proteins that can be modified, such as phosphorylation or glycosylation, different modifications can produce different isolated proteins that can be purified independently.
[0214] "Isolated polynucleotide" means a nucleic acid (e.g., DNA) that is free of the genes that flank it in the native genome of the organism from which the nucleic acid molecule of the invention is derived. The term thus includes, for example, recombinant DNA that is incorporated into a vector, an autonomously replicating plasmid or virus, or into the genomic DNA of a prokaryotic or eukaryotic organism; or that exists as a separate molecule independent of other sequences (e.g., cDNA or genomic or cDNA fragments generated by PCR or restriction endonuclease digestion). In addition, the term includes RNA molecules transcribed from a DNA molecule, as well as recombinant DNA that is part of a hybrid gene encoding an additional polypeptide sequence.
[0215] "Isolated polypeptide" means a polypeptide of the present invention that has been separated from its naturally associated components. Typically, a polypeptide is free of at least 60% by weight of the proteins and natural organic molecules with which it is naturally associated. Preferably, the preparation is at least 75% by weight, more preferably at least 90% by weight, and most preferably at least 99% by weight of the polypeptide of the present invention. An isolated polypeptide of the present invention can be obtained, for example, by extraction from a natural source, by expressing a recombinant nucleic acid encoding the polypeptide, or by chemically synthesizing the protein. Purity can be measured by any suitable method, such as column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.
[0216] As used herein, the term "linker" may refer to a covalent linker (e.g., a covalent bond), a non-covalent linker, or a chemical group that links two molecules or two parts, for example, two components of a protein complex or a ribonucleic acid complex or two domains of a fusion protein, for example, a polynucleotide programmable DNA binding domain (e.g., dCas9) and a deaminase domain (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase). The linker can engage different components of a base editor system or different parts of multiple components. For example, in some embodiments, the linker can engage the guide polynucleotide binding domain of a polynucleotide programmable nucleotide binding domain and the catalytic domain of a deaminase. In some embodiments, the linker can engage a CRISPR polypeptide and a deaminase. In some embodiments, the linker can engage Cas9 and a deaminase. In some embodiments, the linker can engage dCas9 and a deaminase. In some embodiments, the linker can engage nCas9 and a deaminase. In some embodiments, the linker can engage a guide polynucleotide and a deaminase. In some embodiments, the linker can join the deamination component of the base editor system with the polynucleotide programmable nucleotide binding component. In some embodiments, the linker can join the RNA binding portion of the deamination component of the base editor system with the polynucleotide programmable nucleotide binding component. In some embodiments, the linker can join the RNA binding portion of the deamination component of the base editor system with the RNA binding portion of the polynucleotide programmable nucleotide binding component. The linker can be located between two groups, molecules, or other moieties or the linker can be flanked by two groups, molecules, or other moieties, and can be linked to each other via covalent bonds or non-covalent interactions, thereby linking the two. In some embodiments, the linker can be an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker can be a polynucleotide. In some embodiments, the linker can be a DNA linker. In some embodiments, the linker can be an RNA linker. In some embodiments, the linker can comprise an aptamer capable of binding to a ligand. In some embodiments, the ligand can be a carbohydrate, peptide, protein, or nucleic acid. In some embodiments, the linker can comprise an aptamer, which can be derived from a riboswitch. The riboswitch from which the aptamer is derived can be selected from the group consisting of theophylline riboswitch, thiamine pyrophosphate (TPP) riboswitch, adenosylcobalamin (AdoCbl) riboswitch, S-adenosylmethionine (SAM) riboswitch, SAH riboswitch, flavin mononucleotide (FMN) riboswitch, tetrahydrofolate riboswitch, lysine riboswitch, glycine riboswitch, purine riboswitch, GlmS riboswitch, or precursor of Q 1 (PreQ1) riboswitch. In some embodiments, the linker can comprise an aptamer that binds to a polypeptide or protein domain, such as a polypeptide ligand.In some embodiments, the polypeptide ligand can be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the polypeptide ligand can be part of a base editor system component. For example, a nucleobase editing component can include a deaminase domain and an RNA recognition motif.
[0217] In some embodiments, the connexon can be an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the connexon can be a length of about 5 to 100 amino acids, for example, about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 90 or 90 to 100 amino acids in length. In some embodiments, the connexon can be a length of about 100 to 150, 150 to 200, 200 to 250, 250 to 300, 300 to 350, 350 to 400, 400 to 450 or 450 to 500 amino acids in length. Longer or shorter connexons are also contemplated.
[0218] In some embodiments, the linker engages the gRNA binding domain of an RNA programmable nuclease (including a Cas9 nuclease domain) and the catalytic domain of a nucleic acid editing protein (e.g., a cytidine or adenosine deaminase). In some embodiments, the linker engages dCas9 and a nucleic acid editing protein. For example, the linker is located between two groups, molecules, or other moieties or the linker flanks two groups, molecules, or other moieties and is covalently linked to each other, thereby linking the two. In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5 to 200 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140, 150, 160, 175, 180, 190, or 200 amino acids in length.
[0219] In some embodiments, the domain of the base editor is fused via a linker comprising SGGSSGSETPGTSESATPESSGGS,
[0220] SGGSSGGSSGSETPGTSESATPESSGGSSGGS or
[0221] In some embodiments, the domain of the base editor is fused via a linker comprising the amino acid sequence of SGSETPGTSESATPES, which may also be referred to as an XTEN linker. In some embodiments, the linker is 24 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In some embodiments, the linker is 40 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS. In some embodiments, the linker is 64 amino acids in length. In some embodiments, the linker comprises the amino acid sequence
[0222] SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGS SGGS. In some embodiments, the linker is 92 amino acids in length. In some embodiments, the linker comprises the amino acid sequence
[0223] PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS.
[0224] "Marker" means any protein or polynucleotide having a change in expression level or activity that is associated with a disease or disorder.
[0225] As used herein, the term "mutation" refers to the replacement of a residue within a sequence (e.g., a nucleic acid or amino acid sequence) with another residue, or the deletion or insertion of one or more residues within the sequence. In some embodiments, insertion is a genetic conversion that replaces all or part of a wild-type sequence. Herein, a mutation is typically described as: the identity of the original residue, followed by the position of the residue in the sequence, followed by the identity of the newly replaced residue. Various methods for making amino acid substitutions (mutations) provided herein are well known in the art and are provided by, for example, Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)).
[0226] In some embodiments, the base editors disclosed herein can effectively generate "desired mutations" such as point mutations in nucleic acids (e.g., nucleic acids within the genome of a subject) without generating a large number of undesired mutations such as undesired point mutations. In some embodiments, the desired mutation is a mutation generated by a specific base editor (e.g., a cytidine base editor or an adenosine base editor) that is specifically designed to generate the desired mutation and is bound to a guide polynucleotide (e.g., gRNA).
[0227] Typically, mutations made or identified in a sequence (e.g., an amino acid sequence described herein) are numbered relative to a reference (or wild-type) sequence (i.e., a sequence not containing the mutation). Those skilled in the art will readily understand how to determine the position of mutations in amino acid and nucleic acid sequences relative to a reference sequence.
[0228] The term "non-conservative mutation" includes amino acid substitutions between different groups, such as lysine for tryptophan or phenylalanine for serine. In this case, it is preferred that the non-conservative amino acid substitution does not interfere with or inhibit the biological activity of the functional variant. Non-conservative amino acid substitutions can enhance the biological activity of the functional variant, such that the biological activity of the functional variant is increased compared to the wild-type protein.
[0229] The term "nuclear localization sequence," "nuclear localization signal," or "NLS" refers to an amino acid sequence that facilitates the entry of a protein into the cell nucleus. Nuclear localization sequences are known in the art and are described, for example, in International PCT Application PCT / EP2000 / 011690 by Plank et al., filed November 23, 2000, and in WO / 2001 / 038547, published May 31, 2001, the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS described, for example, by Koblan et al., Nature Biotech. 2018 doi: 10.1038 / nbt.4172. Optimized sequences useful in the methods of the present invention are described in Figures 8A to 8E In some embodiments, the NLS comprises the amino acid sequence KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.
[0230] The terms "nucleobase," "nitrogenous base," or "base," used interchangeably herein, refer to nitrogenous biological compounds that form nucleosides, which in turn serve as components of nucleotides. The ability of nucleobases to form base pairs and to stack one on top of another directly leads to long helical structures, such as ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). Five nucleobases, adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), are referred to as primary or standard nucleobases. Adenine and guanine are derived from purines, while cytosine, uracil, and thymine are derived from pyrimidines. DNA and RNA may also contain other (non-primary) bases that have been modified. Non-limiting exemplary modified nucleobases may include hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydroxymethylcytosine. Both hypoxanthine and xanthine can be produced by deamination (amino group replaced by carbonyl group) in the presence of a mutagen. Hypoxanthine can be modified from adenine. Xanthine can be modified from guanine. Uracil can come from the deamination reaction of cytosine. "Nucleoside" consists of a nucleobase and a five-carbon sugar (ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxyuridine and deoxycytidine. Examples of nucleosides with modified nucleobases include inosine (I), xanthine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C) and pseudouridine (Ψ). "Nucleotide" consists of a nucleobase, a five-carbon sugar (ribose or deoxyribose) and at least one phosphate group.
[0231] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to a compound comprising a core base and an acidic moiety, for example, a polymer of nucleosides, nucleotides or nucleotides. Typically, polymeric nucleic acids, such as nucleic acid molecules comprising three or more nucleotides, are linear molecules in which adjacent nucleotides are linked to each other via phosphodiester linkages. In some embodiments, "nucleic acid" refers to a single nucleic acid residue (e.g., nucleotides and / or nucleosides). In some embodiments, "nucleic acid" refers to an oligonucleotide chain comprising three or more single nucleotide residues. As used herein, the terms "oligonucleotide" and "polynucleotide" are used interchangeably and refer to a polymer of nucleotides (e.g., a string of at least three nucleotides). In some embodiments, "nucleic acid" encompasses RNA and single-stranded and / or double-stranded DNA. Nucleic acids can be naturally occurring, for example, in the context of transcripts, mRNA, tRNA, rRNA, siRNA, snRNA, plasmids, cosmids, chromosomes, chromatids or other naturally occurring nucleic acid molecule genomes. On the other hand, nucleic acid molecules can be non-naturally occurring molecules, for example, recombinant DNA or RNA, bearded chromosome, engineered genome or its fragment, or synthetic DNA, RNA, DNA / RNA hybrid, or include non-naturally occurring nucleotides or nucleosides. In addition, the terms "nucleic acid", "DNA", "RNA" and / or similar terms include nucleic acid analogs, for example, with analogs except the phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using a recombinant expression system and optionally purified, chemically synthesized, etc. If applicable, for example, in the case of chemically synthesized molecules, nucleic acid can include nucleoside analogs such as analogs and backbone modifications with chemically modified bases or sugars. Unless otherwise indicated, nucleotide sequences are presented in 5' to 3' directions. In some embodiments, the nucleic acid is or comprises natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-mercaptothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methyl cytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanosine, and 2-mercaptocytidine); chemically modified bases; biologically modified bases (e.g., methylated bases); intercalated bases; modified sugars (2'-e.g., fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages).
[0232] The term "nucleic acid programmable DNA binding protein" or "napDNAbp" can be used interchangeably with "polynucleotide programmable nucleotide binding domain" and refers to a protein that associates with a nucleic acid (e.g., DNA or RNA) such as a guide nucleic acid or a guide polynucleotide (e.g., gRNA), which guides the napDNAbp to a specific nucleic acid sequence. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable DNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable RNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a Cas9 protein. The Cas9 protein can associate with a guide RNA that guides the Cas9 protein to a specific DNA sequence that is complementary to the guide RNA. In some embodiments, the napDNAbp is a Cas9 domain, for example, a nuclease-active Cas9, a Cas9 nickase (nCas9), or a Cas9 (dCas9). Non-limiting examples of nucleic acid-programmable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i.Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpf1, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2 , Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, a type II Cas effector protein, a type V Cas effector protein, a type VI Cas effector protein, CARF, DinG, a homolog thereof, or a modified or engineered version thereof. Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, but they may not be specifically listed in the present disclosure. See, for example, Makarova et al. "Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?" CRISPR J. 2018 Oct; 1: 325-336. doi: 10.1089 / crispr.2018.0033; Yan et al., "Functionally diverse type VCRISPR-Cas systems" Science. 2019 Jan 4; 363(6422): 88-91. doi: 10.1126 / science.aav7271, the entire contents of each of which are incorporated herein by reference.
[0233] As used herein, the term "nucleobase editing domain" or "nucleobase editing protein" refers to a protein or enzyme that catalyzes nucleobase modifications in RNA or DNA, such as cytosine (or cytidine) to uracil (or uridine) or thymine (or thymidine), and adenine (or adenosine) to hypoxanthine (or inosine) deamination, as well as non-templated nucleotide additions and insertions. In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., an adenine deaminase or adenosine deaminase; or a cytidine deaminase or a cytosine deaminase). In some embodiments, the nucleobase editing domain is more than one deaminase domain (e.g., an adenine deaminase or adenosine deaminase and a cytidine deaminase or a cytosine deaminase). In some embodiments, the nucleobase editing domain may be a naturally occurring nucleobase editing domain. In some embodiments, the nucleobase editing domain may be an engineered or evolved nucleobase editing domain from a naturally occurring nucleobase editing domain. The nucleobase editing domain can be from any organism, such as bacteria, human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse.
[0234] As used herein, "obtaining" in "obtaining a reagent" includes synthesizing, isolating, deriving, purchasing or other methods of obtaining the reagent.
[0235] As used herein, "patient" or "subject" refers to a mammalian subject or individual diagnosed as having or susceptible to or prone to developing a disease or disorder or at risk of having or developing a disease or disorder. In some embodiments, a subject with a mutation in a gene encoding SDSP is identified as having Shu-Dieter syndrome (SDS) or is at risk of developing SDS. In some embodiments, the term "patient" refers to a mammalian subject with a higher than average likelihood of developing a disease or disorder. Exemplary patients can be humans, non-human primates, cats, dogs, pigs, cows, cats, horses, camels, llamas, goats, sheep, rodents (e.g., mice, rabbits, rats, gerbils, or guinea pigs), and other mammals that may benefit from the therapies disclosed herein. Exemplary human patients can be male and / or female.
[0236] As used herein, a "patient in need thereof" or "subject in need thereof" refers to a patient diagnosed with, expected to have, or susceptible to a disease or disorder such as SDS, or at risk of having SDS.
[0237] The terms "pathogenic mutation," "pathogenic variant," "disease-causing mutation," "disease-causing variant," "deleterious mutation," or "predisposing mutation" refer to a genetic alteration or mutation that increases an individual's susceptibility or predisposition to a disease or condition. In some embodiments, the pathogenic mutation comprises an alteration within a splice acceptor or indirect donor within a polynucleotide encoding an SBDS protein. In some embodiments, the pathogenic mutation alters splicing of a polynucleotide encoding an SBDS protein, resulting in, for example, protein truncation or otherwise negatively impacting SBDS protein expression or activity.
[0238] The terms "protein," "peptide," "polypeptide," and their grammatical equivalents are used interchangeably herein and refer to a polymer of amino acid residues linked together by peptide (amide) bonds. The term refers to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide will be at least three amino acids in length. A protein, peptide, or polypeptide may refer to a single protein or a collection of proteins. One or more amino acids in a protein, peptide, or polypeptide may be modified, for example, by the addition of chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups; linkers for conjugation; functionalization; or other modifications. A protein, peptide, or polypeptide may also be a single molecule or may be a multimolecular complex. A protein, peptide, or polypeptide may be merely a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be naturally occurring, recombinant, or synthetic, or any combination thereof. As used herein, the term "fusion protein" refers to a hybrid polypeptide comprising protein domains from at least two different proteins. A kind of protein can be positioned at the amino terminal (N-terminal) part of the fusion protein or at the carboxyl terminal (C-terminal) protein, thus forming an amino terminal fusion protein or a carboxyl terminal fusion protein respectively. Protein can include different domains, for example, nucleic acid binding domain (for example, the gRNA binding domain of Cas9, which guides the combination of protein to the target site) and nucleic acid cleavage domain, or the catalytic domain of nucleic acid editing protein. In some embodiments, protein includes protein moiety (for example, the amino acid sequence of nucleic acid binding domain) and organic compound (for example, a compound that can serve as a nucleic acid cleavage agent). In some embodiments, protein and nucleic acid (for example, RNA or DNA) are compounded or associated. Any protein provided herein can be produced by any method known in the art. For example, the protein provided herein can be produced via recombinant protein expression and purification, which is particularly suitable for fusion proteins comprising peptide linkers. Methods for recombinant protein expression and purification are well known and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the entire contents of which are incorporated herein by reference.
[0239] The polypeptides and proteins disclosed herein (including functional portions and functional variants thereof) may comprise synthetic amino acids in place of one or more naturally occurring amino acids. Such synthetic amino acids are known in the art and include, for example, aminocyclohexanecarboxylic acid, norleucine, α-amino n-decanoic acid, homoserine, S-acetylaminomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine, β-phenylserine, β-hydroxyphenylalanine, phenylglycine, α-naphthylalanine, cyclohexylalanine, cyclohexylglycine, indo ...
[00145] The invention also includes the following: -2-aminobutyric acid, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, aminomalonic acid monoamide, N'-benzyl-N'-methyl-lysine, N',N'-dibenzyl-lysine, 6-hydroxylysine, ornithine, α-aminocyclopentanecarboxylic acid, α-aminocyclohexanecarboxylic acid, α-aminocycloheptanecarboxylic acid, α-(2-amino-2-norleucine)-carboxylic acid, α,γ-diaminobutyric acid, α,β-diaminopropionic acid, homoalanine, and α-tert-butylglycine. Polypeptides and proteins may be associated with post-translational modifications of one or more amino acids of the polypeptide construct. Non-limiting examples of post-translational modifications include phosphorylation, acylation (including acetylation and formylation), glycosylation (including N-linked and O-linked), amidation, hydroxylation, alkylation (including methylation and ethylation), ubiquitination, addition of pyrrolidonecarboxylic acid, formation of disulfide bridges, sulfation, myristoylation, palmitoylation, prenylation, farnesylation, geranylation, glycosylphosphatidylinositolation, lipidation, and iodination.
[0240] As used herein, the term "recombinant" in the context of a protein or nucleic acid refers to a protein or nucleic acid that does not exist in nature but is the product of human engineering. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that comprises one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations compared to any naturally occurring sequence.
[0241] "Decrease" means a change in the opposite direction of at least 10%, 25%, 50%, 75% or 100%.
[0242] "Reference" means a standard or control condition. In one embodiment, the reference may be a wild-type or healthy cell. For example, the wild-type or healthy cell may be derived from or obtained from a healthy and / or disease-free subject. In a specific embodiment, the wild-type or healthy cell is a cell expressing a wild-type SBDS protein (i.e., an SBDS protein that is a wild-type SBDS gene product exhibiting wild-type splicing). In other embodiments and without limitation, the reference is an untreated cell that has not been subjected to the test condition or has been subjected to a placebo or saline, vehicle, buffer, and / or a control vector that does not carry the polynucleotide of interest.
[0243] A "reference sequence" is a defined sequence used as a basis for sequence alignment. A reference sequence can be a subset or the entirety of a particular sequence, for example, a segment of a full-length cDNA or gene sequence, or a complete cDNA or gene sequence. For polypeptides, the length of the reference polypeptide sequence will typically be at least about 16 amino acids, at least about 20 amino acids, at least about 25 amino acids, about 35 amino acids, about 50 amino acids, or about 100 amino acids. For nucleic acids, the length of the reference nucleic acid sequence will typically be at least about 50 nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides, or about 300 nucleotides, or any integer number of nucleotides therein. In some embodiments, the reference sequence is the wild-type sequence of a protein of interest. In other embodiments, the reference sequence is a polynucleotide sequence encoding a wild-type protein.
[0244] The terms "RNA programmable nuclease" and "RNA-guided nuclease" are used in conjunction with (e.g., bound to or associated with) one or more RNAs that are not cleavage targets. In some embodiments, when an RNA programmable nuclease is complexed with RNA, it can be referred to as a nuclease:RNA complex. Typically, the associated RNA is referred to as a guide RNA (gRNA). In some embodiments, the RNA programmable nuclease is a (CRISPR-associated system) Cas9 endonuclease, e.g., Cas9 from Streptococcus pyogenes (Csn1) (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin R.E., Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chylinski K., Sharma CM., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607(2011)).
[0245] "Shwachman Bodian Diamond Syndrome (SBDS) protein" means a polypeptide or fragment thereof that has at least about 85% amino acid sequence identity to NCBI Accession No. NP_057122.2 and that has SBDS biological activity. In various embodiments, the SBDS biological activity refers to a role in RNA processing, ribosome generation, or binding to an antibody that specifically binds to the SBDS protein.
[0246] The amino acid sequence of an exemplary SBDS protein is provided below:
[0247]
[0248] In certain embodiments, the SBDS protein comprises protein truncations.
[0249] "Shwachman Bodian Diamond Syndrome (SBDS) polynucleotide" means a nucleic acid sequence encoding an SBDS protein. An exemplary SBDS polynucleotide sequence is provided at NM_016038.2, which is reproduced below. The SBDS polynucleotide open reading frame (ORF) extends from nucleotides 185 to 937 (shown underlined).
[0250]
[0251]
[0252] In some embodiments, the Shwachman Bodian Diamond syndrome (SBDS) polynucleotide comprises a polynucleotide derived from an SBDS pseudogene. In some embodiments, the SBDS polynucleotide comprises a mutation resulting from a genetic transition associated with SDS (e.g., 258+2T>C and / or 183-184TA>CT mutations), alone or in combination with other alterations present in the SBDS pseudogene.
[0253] "Shwachman Bodian Diamond Syndrome (SBDS) pseudogene" means a nucleic acid sequence that has at least about 85% nucleic acid sequence identity to an SBDS polynucleotide. In one embodiment, exemplary pseudogenes include the following and fragments thereof:
[0254] >NR_024109.1 Homo sapiens SBDS pseudogene 1 (SBDSP1), transcript variant 4, noncoding RNA
[0255] CCTTTTTGGGCGTGGAAAGATGGCGGTAAAAGCCACAATGCGCAGGCGTCATCGCTCACTTCTCCCCTCCCGGCTTCTGCTCCACCTGACGCCTGCGCAGTAAGTAAGCCTGCCAGACACGCTGTGGCGGCTGCCTGAAGCTAGTGAGTCGCGGCGCCGCGCACTTGTGGTTGGGTCAGTGCCGCGCGCCGCTCGGTCGTTACCGCGAGGCGCTGGTGGCCTTCAGGCTGGACGGCGCGGGTCAGCCCTGGTTTGCCGGCTTCTGGGTCTTTGAACAGCCGCGATGTCGATCTTCACCCCCACCAACCAGATCCGCCTAACCAATGTGGCCGTGGTACGGATGAAGCGCGCCAGGAAGCGCTTCGAAATCGCCTGCTACAGAAACAAGGTCGTCGGCTGGCGGAGCGGCTTATTTTGACTAAAGGAGAAGTTCAAGTATCAGATAAAGACACACACAACTGGAGCAGATGTTTAGGGACATTGCAATTATTGTGGCAGACAAATGTGTGACTCCTGAAACAAAGAGACCATACACCGTGATCCTTATTGAGAGAGCCATGAAGGACATCCACTATTTGGTGAAAACCAACAGGAGTACAAAACAGCAGGCTTTGGAAGTGATAAAGCAGTTAAAAGAGAAAATGAAGATAGAACGTGCTCACATGAGGCTTCAGTTCATCCTTCCAGTGAATGAAGGCAAGAAGCTGAAAGAAAAGCTCAAGCCACTGATCAAGGTCATAGAAAGTAAAGATTATGGCCAACAGTTAGAAATCGTAAGAGTCAAATATTTTCTTTGCTTCATGTTACCTAAATATTGTATTCTCTAGTAATAAATTTGTAGCAAACATTCAAAAAAAAAAAAAAAAAAAA
[0256] >NR_024110.1 Homo sapiens SBDS pseudogene 1 (SBDSP1), transcript variant 1, non-coding RNA
[0257]
[0258] >NR_024111.1 Homo sapiens SBDS pseudogene 1 (SBDSP1), transcript variant 2, noncoding RNA
[0259]
[0260] >NR_001588.2 Homo sapiens SBDS pseudogene 1 (SBDSP1), transcript variant 3, noncoding RNA
[0261] CCTTTTTGGGCGTGGAAAGATGGCGGTAAAAGCCACAATGCGCAGGCGTCATCGCTCACTTCTCCCCTCCCGGCTTCTGCTCCACCTGACGCCTGCGCAGTAAGTAAGCCTGCCAGACACGCTGTGGCGGCTGCCTGAAGCTAGTGAGTCGCGGCGCCGCGCACTTGTGGTTGGGTCAGTGCCGCGCGCCGCTCGGTCGTTACCGCGAGGCGCTGGTGGCCTTCAGGCTGGACGGCGCGGGTCAGCCCTGGTTTGCCGGCTTCTGGGTCTTTGAACAGCCGCGATGTCGATCTTCACCCCCACCAACCAGATCCGCCTAACCAATGTGGCCGTGGTACGGATGAAGCGCGCCAGGAAGCGCTTCGAAATCGCCTGCTACAGAAACAAGGTCGTCGGCTGGCGGAGCGGCTTGGAAAAAGACCTTGATGAAGTTCTGCAGACCCACTCAGTGTTTGTAAATGTTTCCTAAGGTCAGGTTGCCAAGAAGGAAGATCTCATCAGTGCGTTTGGAACAGATGACCAAACTGAAATCTATTTTGACTAAAGGAGAAGTTCAAGTATCAGATAAAGACACACACAACTGGAGCAGATGTTTAGGGACATTGCAATTATTGTGGCAGACAAATGTGTGACTCCTGAAACAAAGAGACCATACACCGTGATCCTTATTGAGAGAGCCATGAAGGACATCCACTATTTGGTGAAAACCAACAGGAGTACAAAACAGCAGGCTTTGGAAGTGATAAAGCAGTTAAAAGAGAAAATGAAGATAGAACGTGCTCACATGAGGCTTCAGTTCATCCTTCCAGTGAATGAAGGCAAGAAGCTGAAAGAAAAGCTCAAGCCACTGATCAAGGTCATAGAAAGTAAAGATTATGGCCAACAGTTAGAAATCGTAAGAGTCAAATATTTTCTTTGCTTCATGTTACCTAAATATTGTATTCTCTAGTAATAAATTTGTAGCAAACATTCAAAAAAAAAAAAAAAAAAAA
[0262] The term "single nucleotide polymorphism (SNP)" is a variation in a single nucleotide that occurs at a specific position in the genome, with each variation present in the population to an appropriate degree (e.g., >1%). For example, at a specific base position in the human genome, a C nucleotide may occur in most individuals, but in a minority of individuals, this position is occupied by an A. This means that a SNP exists at this specific position, and the two possible nucleotide variations, C or A, are referred to as alleles at this position. SNPs are the basis for differences in disease susceptibility. The severity of the disease and our body's response to treatment are also manifestations of genetic variation. SNPs can fall within the coding region of a gene, within the non-coding region of a gene, or within the intergenic region (the region between genes). In one embodiment, due to the degeneracy of the genetic code, SNPs within a coding sequence do not necessarily change the amino acid sequence of the protein produced. There are two types of SNPs in a coding region: synonymous and non-synonymous SNPs. Synonymous SNPs do not affect the protein sequence, while non-synonymous SNPs change the amino acid sequence of the protein. There are two types of non-synonymous SNPs: missense and nonsense. SNPs that are not in protein coding regions may still affect gene splicing, transcription factor binding, messenger RNA degradation, or the sequence of non-coding RNA. The gene expression affected by this type of SNP is called eSNP (expression SNP) and can be upstream or downstream of the gene. Single nucleotide variants (SNVs) are variations in single nucleotides without any significant frequency and can occur in somatic cells. Somatic single nucleotide variations can also be referred to as single nucleotide changes.
[0263] "Specifically binds" means that a nucleic acid molecule, polypeptide, or a complex thereof (e.g., a nucleic acid programmable DNA binding protein or a guide nucleic acid), compound, or molecule recognizes and binds to a polypeptide and / or nucleic acid molecule of the invention, but does not substantially recognize and bind to other molecules in a sample (e.g., a biological sample).
[0264] Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to an endogenous nucleic acid sequence, but typically will exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence will typically hybridize to at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to an endogenous nucleic acid sequence, but typically will exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence will typically hybridize to at least one strand of a double-stranded nucleic acid molecule. "Hybridization" refers to the pairing between complementary polynucleotide sequences (e.g., a gene described herein) or fragments thereof under various stringency conditions to form a double-stranded molecule. (See, e.g., Wahl, G. M. and S. L. Berger (1987) Methods Enzymol. 152:399; Kimmel, A. R. (1987) Methods Enzymol. 152:507).
[0265] For example, stringent salt concentrations will generally be less than about 750mM NaCl and 75mM trisodium citrate, preferably less than 500mM NaCl and 50mM trisodium citrate, and more preferably about 250mM NaCl and 25mM trisodium citrate. Low stringency hybridization can be obtained in the absence of an organic solvent such as formamide, while high stringency hybridization can be obtained in the presence of at least about 35% formamide and more preferably at least about 50% formamide. Stringent temperature conditions will generally include a temperature of at least about 30°C, more preferably at least about 37°C, and most preferably at least about 42°C. The inclusion or exclusion of different other parameters such as hybridization time, surfactants such as sodium dodecyl sulfate (SDS), and carrier DNA is well known to those skilled in the art. By combining these conditions as needed, various levels of stringency can be achieved. In a preferred embodiment, hybridization will occur at 30°C in 750mM NaCl, 75mM trisodium citrate, and 1% SDS. In a more preferred embodiment, hybridization will occur in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA) at 37° C. In a most preferred embodiment, hybridization will occur in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 μg / ml ssDNA at 42° C. Useful variations on these conditions will be apparent to those skilled in the art.
[0266] For most applications, the wash step after hybridization will also change the stringency. Wash stringency conditions can be defined by salt concentration and by temperature. As above, wash stringency can be increased by reducing the salt concentration or increasing the temperature. For example, the stringent salt concentration used for the wash step will preferably be less than about 30mM NaCl and 3mM trisodium citrate, and most preferably about 15mM NaCl and 1.5mM trisodium citrate. The stringent temperature conditions used for the wash step will generally include a temperature of at least about 25°C, more preferably at least about 42°C, and even more preferably at least about 68°C. In one embodiment, the wash step will occur at 25°C in 30mM NaCl, 3mM trisodium citrate, and 0.1% SDS. In another embodiment, the wash step will occur at 42°C in 15mM NaCl, 1.5mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the wash step will occur at 68°C in 15mM NaCl, 1.5mM trisodium citrate, and 0.1% SDS. Other variations of these conditions will be apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.
[0267] "Split" means divided into two or more segments.
[0268] "Split Cas9 protein" or "split Cas9" refers to a Cas9 protein that is provided as an N-terminal fragment and a C-terminal fragment encoded by two independent nucleotide sequences. The polypeptides corresponding to the N-terminal portion and the C-terminal portion of the Cas9 protein can be spliced to form a "reconstructed" Cas9 protein. In a specific embodiment, the Cas9 protein is divided into two fragments located within the disordered region of the protein, for example, as described in Nishimasu et al., Cell, Volume 156, Issue 5, pp.935-949, 2014 or as described in Jiang et al. (2016) Science 351: 867-871.PDB file: 5F9R, each of which is incorporated herein by reference. In some embodiments, the protein is divided into two fragments at any C, T, A, or S located in the region between amino acids A292 to G364, between F445 to K483, or between E565 to T637 of SpCas9, or at corresponding positions in any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or other napDNAbp. In some embodiments, the protein is divided into two fragments located at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of dividing a protein into two fragments is referred to as "splitting" the protein.
[0269] In other embodiments, the N-terminal portion of the Cas9 protein comprises amino acids 1 to 573 or 1 to 637 of Streptococcus pyogenes Cas9 wild type (SpCas9) (NCBI reference sequence: NC_002737.2, Uniprot reference sequence: Q99ZW2), and the C-terminal portion of the Cas9 protein comprises amino acids 574 to 1368 or 638 to 1368 of the SpCas9 wild type.
[0270] The C-terminal portion of the split Cas9 can be joined to the N-terminal portion of the split Cas9 to form a complete Cas9 protein. In some embodiments, the C-terminal portion of the Cas9 protein starts where the N-terminal portion of the Cas9 protein ends. Therefore, in some embodiments, the C-terminal portion of the split Cas9 comprises the portion of amino acids (551 to 651) to 1368 of spCas9. "(551 to 651) to 1368" means starting at the amino acids between amino acids 551 to 651 (inclusive) and ending at amino acid 1368.For example, the C-terminal portion of the split Cas9 can include amino acids 551 to 1368, 552 to 1368, 553 to 1368, 554 to 1368, 555 to 1368, 556 to 1368, 557 to 1368, 558 to 1368, 559 to 1368, 560 to 1368, 561 to 1368, 562 to 1368, 563 to 1368, 564 to 1368, 565 to 1368, 566 to 1368, 567 to 1368, 568 to 1368, 569 to 1368, 570 to 1368, 571 to 1368, 572 to 1368, 573 to 1368, 74 to 1368, 575 to 1368, 576 to 1368, 577 to 1368, 578 to 1368, 579 to 1368, 580 to 1368, 581 to 1368, 582 to 1368, 583 to 1368, 584 to 1368, 585 to 1368, 586 to 1368, 587 to 1368, 588 to 1368, 589 to 1368, 590 to 1368, 591 to 1368, 592 to 1368, 593 to 1368, 594 to 1368, 595 to 1368, 596 to 1368, 597 to 1368, 598 to 1368, 599 to 1368, 60 0 to 1368, 601 to 1368, 602 to 1368, 603 to 1368, 604 to 1368, 605 to 1368, 606 to 1368, 607 to 1368, 608 to 1368, 609 to 1368, 610 to 1368, 611 to 1368, 612 to 1368, 613 to 1368, 614 to 1368, 615 to 1368, 616 to 1368, 617 to 1368, 618 to 1368, 619 to 1368, 620 to 1368, 621 to 1368, 622 to 1368, 623 to 1368, 624 to 1368, 625 to 1368, 626 to 1368, 627 to 1368, 628 to 1368, 629 to 1368, 630 to 1368, 631 to 1368, 632 to 1368, 633 to 1368, 634 to 1368, 635 to 1368, 636 to 1368, 637 to 1368, 638 to 1368, 639 to 1368, 640 to 1368, 641 to 1368, 642 to 1368, 643 to 1368, 644 to 1368, 645 to 1368, 646 to 1368, 647 to 1368, 648 to 1368, 649 to 1368, 650 to 1368, or 651 to 1368. In some embodiments, the C-terminal portion of the split Cas9 protein comprises a portion of amino acids 574 to 1368 or 638 to 1368 of SpCas9.
[0271] "Subject" means a mammal, including but not limited to humans and non-human mammals, such as non-human primates (monkeys), cows, horses, dogs, sheep, or cats. In some embodiments, the subjects described herein include a pathogenic mutation in an SDS polynucleotide sequence encoding an SBDS protein, which identifies the subject as having SDS or having a predisposition to develop SDS.
[0272] "Substantially identical" means that a polypeptide or nucleic acid molecule exhibits at least 50% identity to a reference amino acid sequence (e.g., any of the amino acid sequences described herein) or nucleic acid molecule (e.g., any of the nucleic acid sequences described herein). In one embodiment, the sequence is at least 60%, 65%, 70%, 75%, 80% or 85%, 90%, 95% or even 99% identical at the amino acid level or nucleic acid level to the nucleic acid being compared.
[0273] Sequence identity is typically measured using sequence analysis software (e.g., Sequence Analysis Software Package, BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs from the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705). Such software matches identical or similar sequences by assigning degrees of homology to different substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within each group of a series: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. In an exemplary method for determining the degree of identity, the BLAST program can be used, e -3 and e -100 Probability scores between indicate closely related sequences.
[0274] Use COBALT, for example, with the following parameters:
[0275] a) Alignment parameters: gap penalty -11, -1, and end gap penalty -5, -1,
[0276] b) CDD parameters: Use RPS BLAST on; Blast E value 0.003; Find Conserved columns and Recompute on, and
[0277] c) Query clustering parameters: Enable query clusters; word length 4; maximum inter-cluster distance 0.8; AlphabetRegular.
[0278] Use EMBOSS Needle, for example, with the following parameters:
[0279] a) Matrix: BLOSUM62;
[0280] b) GAP OPEN: 10;
[0281] c) GAP EXTEND: 0.5;
[0282] d)OUTPUT FORMAT; yes;
[0283] e)END GAP PENALTY: FALSE;
[0284] f) END GAP OPEN: 10; and
[0285] g)END GAP EXTEND:0.5.
[0286] The term "target" refers to a sequence within a nucleic acid molecule that is deaminated by a deaminase or a fusion protein comprising a deaminase (e.g., a dCas9-adenosine deaminase fusion protein or base editor disclosed herein).
[0287] Because RNA-programmable nucleobases (e.g., Cas9) use RNA:DNA hybridization to target DNA cleavage sites, these proteins can in principle be targeted to any sequence specific for the guide RNA. Methods for site-specific cleavage (e.g., to modify genomes) using RNA-programmable nucleobases such as Cas9 are known in the art (see, e.g., Cong, L. et al., Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, W. et al., Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, J. et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic acids research (2013); Jiang, W. et ah RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature biotechnology 31, 233-239 (2013); the entire contents of each of which are incorporated herein by reference).
[0288] As used herein, the terms "treat", "treat" and the like refer to alleviating, reducing, eliminating, reducing, slowing down or alleviating a disease or disorder and / or symptoms associated therewith or obtaining a desired pharmacological and / or physiological effect. It should be understood that, although not excluded, treating a lesion or condition does not require the complete elimination of the lesion, condition or symptoms associated therewith. In some embodiments, the effect is therapeutic, that is, without limitation, the effect partially or completely alleviates, eliminates, abolishes, weakens, slows down, reduces the intensity of the disease and / or negative symptoms attributable to the disease, or cures the disease and / or negative symptoms attributable to the disease. In some embodiments, the effect is preventive, that is, the effect protects or prevents the occurrence or recurrence of the disease or condition. To this end, the methods of the present disclosure comprise administering a therapeutically effective amount of a composition as described herein. In one embodiment, the present invention provides treatment for SDS.
[0289] "Uracil glycosylase inhibitor" or "UGI" means an agent of the known uracil excision repair system. In one embodiment, the agent is a protein or fragment thereof that binds to the host uracil-DNA glycosidase and prevents the removal of uracil residues from DNA. In one embodiment, UGI is a protein, fragment or domain thereof that is capable of inhibiting the uracil-DNA glycosidase base excision repair enzyme. In some embodiments, the UGI domain comprises wild-type UGI or a modified version thereof. In some embodiments, the UGI domain comprises a fragment of the exemplary amino acid sequence described in detail below. In some embodiments, the UGI fragment comprises an amino acid sequence that comprises at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% of the exemplary UGI sequence provided below. In some embodiments, UGI comprises an amino acid sequence that is homologous to the exemplary UGI amino acid sequence described in detail below or a fragment thereof. In one embodiment, the UGI or portion thereof is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, or 100% identical to a wild-type UGI or UGI sequence described in detail below, or a fragment thereof. An exemplary UGI comprises the following amino acid sequence:
[0290] >splP14739IUNGI_BPPB2 uracil-DNA glycosylase inhibitor
[0291] MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLT SD APEYKPW ALVIQDS NGENKIKML.
[0292] 47, 48, 49, or 50.
[0293] The recitation of a series of chemical groups in any variable definition herein includes definitions of that variable as any single group or combination of the listed groups. The recitation of an embodiment for a variable or aspect herein includes that embodiment as any single embodiment or in combination with other embodiments or portions thereof.
[0294] Any composition or method provided herein can be used in combination with one or more of any other compositions and methods provided herein.
[0295] The description and examples herein illustrate embodiments of the present disclosure in detail. It should be understood that the present disclosure is not limited to the specific embodiments described herein and is therefore variable. Those skilled in the art will recognize that there are a large number of variations and modifications to the present disclosure that are encompassed within the scope of the present disclosure.
[0296] All terms are intended to be understood as understood by those skilled in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this disclosure belongs.
[0297] Unless otherwise indicated, the practice of some embodiments disclosed herein employs conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genetics, and recombinant DNA, which are within the skill of the art. See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, Fourth Edition (2012); Current Protocols in Molecular Biology series (FM Ausubel, et al. eds.); Methods In Enzymology series (Academic Press, Inc.); PCR 2: A Practical Approach (MJ MacPherson, BD Hames and GR Taylor eds. (1995)); Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual; and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, Sixth Edition (RI Freshney, ed. (2010).
[0298] Although various features of the present disclosure may be described in the context of a single embodiment, these features may also be provided separately or in any combination. Conversely, although the invention may be described in the context of separate embodiments for clarity, the invention may also be implemented in a single embodiment. The section headings used herein are for organizational purposes only and are not to be construed as limitations on the subject matter described.
[0299] The features of the present disclosure are particularly described in the appended claims.A better understanding of the features and advantages of the present disclosure may be obtained by referring to the following detailed description of exemplary embodiments (in which the principles of the present disclosure are used) and the accompanying drawings described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0300] Figure 1A and Figure 1B The mutations in SBDS that cause SDS are shown. Figure 1AA map of SBDS is provided (coding regions are indicated by light shading, noncoding regions by dark shading) and a sequence alignment of the exon 2 region of SBDS with the SBDS protein, with gene-specific sequences (grey; top) and pseudogene-specific sequences (grey; bottom) indicated. Compared to SBDS, exon 2 in SBDSP, resulting from the conversion event, contains sequence changes that are predicted to result in protein truncation (underlined). These sequence changes include an in-frame stop codon at position 184 and a T→C change at position 250+10 (corresponding to the invariant T in the donor splice site at position 258+2 in SBDS), resulting in the use of an alternative donor splice site at position 250+1 (the invariant splice site position is boxed). Figure 1B Sequence reads from cloned segments of exon 2 from SBDS are shown, indicating sequence changes in individuals with SDS that arise from a gene conversion event between SBDS and its pseudogene; three conversion alleles are shown. These alleles include 183–184TA→CT, 258+2T→C, and an extended conversion mutation, 183–184TA→CT+201A→G+258+2T→C. In each case, informative flanking positions (including 141 and 258+124) are conversions (green).
[0301] Figures 2A to 2D is a schematic diagram illustrating a strategy for restoring transcription in an SBDS gene containing one or more pathogenic mutations. Figure 2A Shown is a strategy for introducing mutations that eliminate the stop codon and provide for expression of SBDS proteins comprising an alternative amino acid (eg, Trp(W)) at amino acid position 62 (eg, (K62X)). Figure 2B and 2D The strategy for correcting the splice site at nucleotide position 258 (target SNP rs113993993C→T) is shown. Figure 2C The splice donor position is shown where the standard splice donor is restored to correct the SNP mutation.
[0302] Figures 3A to 3C Presented is a table showing amino acid positions where substitutions occur in Cas9 proteins (e.g., modified Cas9, such as modified SpCas9), resulting in Cas9 variants specific for altered PAM 5'-NGC-3' or PAM containing 5'-NGC-3', and plasmid constructs encoding SpCas9 variant sequences. A cytidine base editor (CBE) comprising at least one cytidine deaminase and at least one Cas9 variant as described is used to correct mutations in the SBDS gene associated with SDS, as described in Example 3. Figure 3APresents the amino acid positions in the Cas9 protein that were altered from the wild type to generate Cas9 variants (indicated by numbers in the left column) that are capable of binding to the NGC PAM. These Cas9 variants are components of the CBEs evaluated in the base editing studies described herein; Figure 3B Presents a subset of Cas9 variants that provide particularly good high on-target editing and limited bystander effects in research. Figure 3B Also shown is a schematic diagram of the Cas9 protein domains and their positioning in the Cas9 protein sequence. Figure 3C Shown are plasmid vector components encoding SpCas9 variants specific for altered PAM 5'-NGC-3', as described herein.
[0303] Figure 4 Shown are graphs comparing the relative mutation rates of base editing achieved by CBEs containing different cisidine deaminases (as indicated on the abscissa).
[0304] Figure 5 is a table showing guide RNAs (gRNAs) used with CBEs evaluated in the studies described herein. In embodiments, the gRNA sequences are components of the plasmid constructs used in the base editing studies described in the Examples.
[0305] Figures 6A to 6C The percentage of editing (e.g., on-target editing) versus the percentage of bystander editing achieved by NGC CBE variants and 19mer and 20mer gRNAs (e.g., G88 and G44) as described herein is shown. Figure 6A In the right picture and in Figure 6B In the text, “PV226” and “PV230” refer to the plasmids used in the study. The PV226 plasmid contains the polynucleotide for editing Cas9 variant #226, the sequence of which is shown in Figures 3A to 3C The PV230 plasmid contains a polynucleotide encoding Cas9 variant #230, the sequence of which is shown in Figures 3A to 3C The Cas9 variants (whose sequences are Figures 3A to 3C The editing percentages exhibited by other NGC CBEs (described in ) and 20mer gRNA G44 are shown in Figure 6C middle.
[0306] Figure 7A and Figure 7B Graph showing the percentage of editing achieved by NGC CBEs comprising a cytidine deaminase and the Cas9 variants shown in Table 13, used in conjunction with a 19mer gRNA (G88) and a 20mer gRNA (G44), as described in Example 4 herein.
[0307] Figures 8A to 8JIt is shown that by including different cisidine deaminases and having the following in the Cas9 amino acid sequence: Figures 3A to 3C Graph of the percentage of base editing (on-target and bystander editing) achieved by NGC CBEs of Cas9 (e.g., SpCas9) in combination with 19mer or 20mer gRNAs for the mutation-specific combinations presented in Table 13, as assessed in a cell-based (HEK293) assay to correct splice site SNPs in SBDS polynucleotide sequences. Figure 8A The results show that the NGC CBEs containing Cas9 variant 225 and PpAPOBEC1 and the NGC CBEs containing PpAPOBEC1 and Cas9 variants 226 and 244 ( Figures 3A to 3C ) NGC CBE 454 and 459 (Table 13) in combination with 19mer (guide 88) gRNA showed on-target and bystander editing percentages. Figure 8B Shown are the percentages of on-target and bystander editing exhibited by NGC CBEs containing Cas9 variant 225 and PpAPOBEC1, and by NGC CBEs 454 and 459 containing PpAPOBEC1 and Cas9 variants 226 and 244, respectively (Table 13), in combination with a 20mer (guide 44) gRNA. Figure 8C and 8D The results show that the enzyme contains AmAPOBEC1 cytidine deaminase and Cas9 variants 225, 226 and 244 ( Figures 3A to 3C ) of NGC CBEs with 19mer (guide 88) or 20mer (guide 44) gRNAs. Figure 8E and 8F Shown are the percentages of on-target and bystander base editing achieved by NGC CBEs containing PmCDA1 cytidine deaminase and Cas9 variants 225, 453, and 458 (Table 13) with either 19mer (guide 88) or 20mer (guide 44) gRNAs. Figure 8G and 8H Shown are the percentages of on-target and bystander base editing achieved by NGCCBEs containing RRA3F cytidine deaminase and Cas9 variants 225, 455, and 460 (Table 13) with either 19mer (guide 88) or 20mer (guide 44) gRNAs. Figure 8I and 8J The percentages of on-target and bystander base editing achieved by NGC CBEs containing SsAPOBEC2 cytidine deaminase and Cas9 variants 225, 456, and 461 (Table 13) with either 19mer (guide 88) or 20mer (guide 44) gRNAs are shown. Figures 8A to 8JIn the literature, Cas9 variant 225 (or PV225) is also alternatively referred to as "Beam shuffle".
[0308] 9A to 9D The results show that the NGC CBE comprising the PpAPOBEC1 cytidine deaminase polypeptide sequence (the polypeptide sequence contains various mutations as described in Example 4, such as the H122A mutation alone and in combination with the amino acid mutations R33A, W90F, K34A, R52A, H121A and Y120F) and the 19mer gRNA ( Figure 9A ) or 20mer gRNA ( Figure 9B ) Graph and dot plot of the percentage of editing achieved together. The percentage of on-target versus bystander editing was assessed in a cell-based in vitro assay. Figure 9C and 9D Presented in dot blot format Figure 9A and 9B data.
[0309] Figure 10 A table is presented describing mutations and combinations of mutations that were made in the SpCas9 protein to create SpCas9 variants having the indicated combinations of mutations, including the "NRCH" mutation as described in S. Miller et al., April, 2020, "Continuous evolution of SpCas9 variants compatible with non-G PAMs," Nature Biotechnology, 38(4):471-481 (published online 2020 Feb 10. doi:10.1038 / s41587-020-0412-8), the contents of which are incorporated herein by reference in their entirety. Combinations of NRCH mutations (amino acid substitutions) were included in several different SpCas9 variants to determine which SpCas9 variant components would result in an NGC CBE for use in correcting a splice site SNP of the SBDS gene associated with SDS with high on-target and bystander base editing. (Example 6). Figure 10 In the figure, the darker shaded amino acids reflect amino acid substitutions in the Cas9 (SpCas9) amino acid sequence compared to the wild-type, unmutated Cas9 (SpCas9) protein. The lighter shaded amino acids reflect amino acid residues in the wild-type, unmutated Cas9 (SpCas9) protein.
[0310] Figure 11A and Figure 11B A diagram is shown illustrating the ability of a cytidine deaminase (e.g., PpAPOBEC1) to be expressed by a Cas9 variant comprising a cytidine deaminase (e.g., PpAPOBEC1) and a SpCas9 variant (e.g., Figure 10 The on-target and bystander editing efficiencies of NGC CBEs (468 and 469) for correction of splice site SNPs in the SBDS gene associated with SDS were evaluated in a cell-based assay to compare the percent editing achieved when used in conjunction with 19mer or 20mer gRNAs. Figure 10 ) showed high levels of on-target and off-target base editing.
[0311] 12A to 12C A graph is shown showing the results of a cell-based in vitro assay performed to evaluate the base editing efficiency and percentage of on-target and bystander editing achieved by NGC CBEs encoded by mRNAs with gRNAs of varying lengths (17mer, 18mer, 19mer, 20mer, or 21mer) as described in Example 6. As observed, mRNA 342 had the fewest C to A or C to G conversions when used with 18mer and 20mer gRNAs compared to mRNA 340 or mRNA 341. DETAILED DESCRIPTION
[0312] The present invention features compositions and methods that use programmable nucleobase editors to edit pathogenic gene mutations that cause abnormal splicing in genes to allow transcription and achieve therapeutic effects. In some embodiments, editing includes converting a stop codon to a codon that allows transcription. In some embodiments, editing includes providing and correcting a splice acceptor or splice donor site, or providing an alternative splice acceptor or splice donor site. In some embodiments, more than one mutation that causes abnormal splicing is corrected.
[0313] The present invention is based, at least in part, on a strategy to use adenosine or cytidine base editors (ABEs, CBEs) to edit pathogenic mutations associated with SDS (e.g., mutations due to gene conversion) in genes. Accordingly, the present invention provides base editor systems comprising ABEs or CBEs that can be used to treat or prevent SDS.
[0314] Schwart-Diesel syndrome (SDS)
[0315] Shwachman-Diamond syndrome (SDS) is an autosomal recessive disorder. Approximately 90% of patients who meet the clinical diagnostic criteria for SDS harbor a mutation in the Shwachman-Bodian-Diamond syndrome (SBDS) gene. The carrier frequency of this mutation has been estimated to be approximately 1 in 110. This highly conserved gene has five exons encompassing 7.9 kb and maps to the centromeric region of chromosome 7, 7q11. The SDBS gene encodes a novel 250-amino acid protein that lacks homology to functional domains of known proteins. The adjacent pseudogene, SBDSP, shares 97% homology with SBDS but contains deletions and nucleotide changes that prevent the production of a functional protein. Approximately 75% of patients with SDS harbor mutations that result in a gene conversion event using this pseudogene. Gene conversion occurs when recombination occurs between homologous sequences present at different genomic loci (paralogous sequences). The existence of the SBDS pseudogene (also known as SBDSP) likely results from a previous gene duplication. SBDS mRNA and protein are ubiquitously expressed throughout human tissues at both the mRNA and protein levels. Although the early truncating SBDS mutation 183TA>CT is common in SDS patients, no patients have been found who are homozygous for this mutation, suggesting that complete loss of SBDS expression can be lethal in humans.
[0316] Common sequence changes associated with SDS include a TA→CT dinucleotide change at positions 183 to 184 or an 8-bp deletion at the end of exon 2. Analysis of the SBDS genomic sequence confirmed the presence of the 183 to 184 TA→CT change, and a 258+2T→C change was found in individuals with SDS expressing deleted transcripts. The mutation 258+2T→C is predicted to interrupt the donor splice site of intron 2, and the 8-bp deletion is consistent with the use of an upstream cryptic splice donor site at positions 251 to 252. The dinucleotide change 183 to 184 TA→CT introduces an in-frame stop codon (K62X), and the 258+2T→C and the resulting 8-bp deletion cause premature truncation of the encoded protein by a frameshift (84Cfs3).
[0317] The present invention provides compositions and methods that allow transcription of polynucleotides having one or more alterations (e.g., gene conversion) that result in aberrant splicing, thereby providing expression of a functional SBDS protein (a protein having activity sufficient to improve the effects of SBDS gene conversion). In a specific embodiment, the present invention provides for the introduction of an alteration comprising 183 to 184 TA→CT into the SBDS gene, which converts the TAA terminator to TGG encoding Trp and allows transcription. In other embodiments, the present invention introduces alterations in the polynucleotide sequence that introduce splice donor or effector sites that allow splicing of a polynucleotide encoding a biologically active protein. In some embodiments, the present invention corrects a site in exon 2 of the SBDS gene (e.g., by editing the cytosine at nucleotide position 1495, such as Figure 2B ).
[0318] Nucleobase editors
[0319] Disclosed herein is a base editor or a core base editor for editing, modifying or changing a target nucleotide sequence in a polynucleotide. A core base editor or a base editor is described herein, comprising a polynucleotide programmable nucleotide binding domain and a core base editing domain (e.g., adenosine deaminase, cytidine deaminase). When synergistically acting with a bound guide polynucleotide (e.g., gRNA), the polynucleotide programmable nucleotide binding domain can specifically bind to a target polynucleotide sequence (i.e., via complementary base pairing between the base of the bound guide nucleic acid and the base of the target polynucleotide sequence), and thus the base editor is positioned to the desired target nucleic acid sequence to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid.
[0320] Polynucleotide programmable nucleotide binding domain
[0321] It will be appreciated that a polynucleotide programmable nucleotide binding domain may also include a nucleic acid programmable protein that binds RNA. For example, a polynucleotide programmable nucleotide binding domain may be associated with a nucleic acid that guides the polynucleotide programmable nucleotide binding domain to the RNA. Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, but they are not specifically listed in this disclosure.
[0322] The polynucleotide programmable nucleotide binding domain of the base editor itself may include one or more domains. For example, the polynucleotide programmable nucleotide binding domain may include one or more nuclease domains. In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide binding domain may include an endonuclease or an exonuclease. Herein, the term "exonuclease" refers to a protein or polypeptide capable of digesting a nucleic acid (e.g., RNA or DNA) from a free end, and the term "endonuclease" refers to a protein or polypeptide capable of catalyzing (e.g., cutting) an internal region in a nucleic acid (e.g., DNA or RNA). In some embodiments, an endonuclease can cut a single strand of a double-stranded nucleic acid. In some embodiments, an endonuclease can cut two chains of a double-stranded nucleic acid molecule. In some embodiments, the polynucleotide programmable nucleotide binding domain may be a deoxyribonuclease. In some embodiments, the polynucleotide programmable nucleotide binding domain may be a ribonuclease.
[0323] In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide binding domain can nick zero, one or two chains of the target polynucleotide.In some cases, the polynucleotide programmable nucleotide binding domain may include a nickase domain.Herein, "nickase" refers to a polynucleotide programmable nucleotide binding domain comprising a nuclease domain that can cut only one of the two chains of a double-stranded nucleic acid molecule (e.g., DNA).In some embodiments, a nickase can be derived from a polynucleotide programmable nucleotide binding domain of complete catalytic activity (e.g., natural) form by introducing one or more mutations into an active polynucleotide programmable nucleotide binding domain.For example, if the polynucleotide programmable nucleotide binding domain includes a nickase derived from Cas9, the nickase domain derived from Cas9 may include a D10A mutation and a histidine at position 840.In these cases, residue H840 maintains catalytic activity and therefore can cut a chain of a nucleic acid double helix.In another example, the nickase domain derived from Cas9 may include an H840A mutation, while the amino acid residue at position 10 remains D. In some embodiments, the nickase can be derived from a fully catalytically active (e.g., native) form of a polynucleotide programmable nucleotide binding domain by removing all or a portion of a nuclease that is not required for nickase activity. For example, if the polynucleotide programmable nucleotide binding domain comprises a nickase derived from Cas9, the nickase domain derived from Cas9 can comprise a deletion of all or a portion of a RuvC domain or a HNH domain.
[0324] The amino acid sequence of an exemplary catalytically active Cas9 is as follows:
[0325]
[0326] Therefore, a base editor comprising a polynucleotide programmable nucleotide binding domain with a nickase domain is capable of generating single-stranded DNA breaks (nicks) at a specific polynucleotide target sequence (e.g., determined by the complementary sequence of the combined guide nucleic acid). In some embodiments, the chain of the nucleic acid double helix target polynucleotide sequence cut by a base editor comprising a nickase domain (e.g., a nickase domain derived from Cas9) is a chain that is not edited by a base editor (i.e., the chain cut by the base editor is opposite to the chain comprising the base to be edited). In other embodiments, a base editor comprising a nickase domain (e.g., a nickase domain derived from Cas9) can cut the chain of a DNA molecule to be targeted for editing. In these cases, the non-targeting chain is not cut.
[0327] Also provided herein are base editors comprising a catalytically dead polynucleotide programmable nucleotide binding domain (i.e., unable to cut the target polynucleotide sequence).Herein, the terms “catalytically dead” and “nuclease dead” are used interchangeably to refer to a polynucleotide programmable nucleotide binding domain having one or more mutations and / or deletions that cause it to lose the ability to cut a nucleic acid chain. In some embodiments, the catalytically dead polynucleotide programmable nucleotide binding domain base editor may lack nuclease activity, which is the result of a specific point mutation in one or more nuclease domains. For example, in the case where the base editor comprises a Cas9 domain, Cas9 may comprise both a D10A mutation and an H840A mutation. Such mutations inactivate the two nuclease domains, resulting in the loss of nuclease activity. In other embodiments, the catalytically dead polynucleotide programmable nucleotide binding domain may comprise one or more deletions of all or part of a dialogue domain (e.g., RuvC1 and / or HNH domain). In further embodiments, the catalytically dead polynucleotide programmable nucleotide binding domain comprises a point mutation (eg, D10A or H840A) and a deletion of all or a portion of the nuclease domain.
[0328] It is also contemplated herein that mutations of catalytically dead polynucleotide programmable nucleotide binding domains can be generated from previously functional versions of the polynucleotide programmable nucleotide binding domains. For example, in the case of catalytically dead Cas9 ("dCas9"), variants with mutations other than D10A and H840A are provided, which result in Cas9 with inactive nucleases. For example, such mutations include other amino acid replacements at D10 and H840, or other replacements within the nuclease domain of Cas9 (e.g., replacements in the HNH nuclease subdomain and / or the RuvC1 subdomain). Based on the present disclosure and knowledge in the art, other suitable dCas9 domains of inactive nucleases will be apparent to those skilled in the art and will be within the scope of the present disclosure. Such other exemplary suitable inactive nuclease Cas9 domains include, but are not limited to, D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013; 31(9):833-838, the entire contents of which are incorporated herein by reference).
[0329] Non-limiting examples of polynucleotide programmable nucleotide binding domains that can be incorporated into base editors include domains derived from CRISPR proteins, restriction nucleases, giant nucleases, TAL nucleases (TALENs), and zinc finger nucleases (ZFNs). In some cases, the base editor comprises a polynucleotide programmable nucleotide binding domain having a natural or modified protein or portion thereof, which, via a bound guide polynucleotide, is capable of binding to a nucleic acid sequence during CRISPR (i.e., clustered regularly interspaced short palindromic repeats)-mediated nucleic acid modification. Herein, the protein is referred to as “CRISPR protein.” Accordingly, disclosed herein is a base editor comprising a polynucleotide programmable nucleotide binding domain having all or part of a CRISPR protein (i.e., a base editor comprising all or part of a CRISPR protein as a domain (also referred to as a domain derived from a CRISPR protein)). The domain derived from a CRISPR protein incorporated into the base editor can be modified compared to a wild-type or natural version of the CRISPR protein. For example, as described below, a domain derived from a CRISPR protein can comprise one or more mutations, insertions, deletions, rearrangements, and / or recombination relative to a wild-type or naturally occurring version of the CRISPR protein.
[0330] CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). The CRISPR cluster contains a spacer sequence, a sequence complementary to the precursor mobile element, and a target invading nucleic acid. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In the type II CRISPR system, the correct processing of pre-crRNA requires trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein. TracrRNA acts as a guide for ribonuclease 3-assisted pre-crRNA processing. Subsequently, Cas9 / crRNA / tracrRNA endonucleases cut linear or circular dsDNA targets complementary to the spacer sequence. The target chain that is not complementary to crRNA is first cut endonucleases and then trimmed exonucleases by 3'-5'. In nature, DNA binding and cutting generally require protein and two types of RNA. However, single guide RNA ("sgRNA" or "gRNA" for short) can be engineered to incorporate multiple crRNAs and tracrRNAs into a single RNA species. See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif (PAM or protospacer adjacent motif) in the CRISPR repeat sequence to help distinguish self from non-self.
[0331] In some embodiments, the methods described herein may utilize engineered Cas proteins. A guide RNA (gRNA) is a short synthetic RNA consisting of a scaffold sequence necessary for Cas binding and a user-defined spacer sequence of ~20 nucleotides that defines the genomic target to be modified. Thus, a skilled artisan can alter the genomic or polynucleotide target of a Cas protein by changing the target sequence present in the gRNA. The specificity of a Cas protein is determined in part by how specific the gRNA targeting sequence is for the genomic polynucleotide target sequence versus the rest of the gene.
[0332] In some embodiments, the gRNA scaffold sequence is as follows: GUUUUAGAGC UAGAAAUAGC AAGUUAAAAUAAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU.
[0333] In one embodiment, the RNA scaffold comprises a stem-loop. In one embodiment, the RNA scaffold comprises a nucleic acid sequence:
[0334] GUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG.
[0335] In one embodiment, the RNA scaffold comprises the nucleic acid sequence:
[0336] GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCCGGUGCUUUU.
[0337] In one embodiment, the Streptococcus pyogenes sgRNA scaffold polynucleotide sequence is as follows:
[0338] GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC.
[0339] In one embodiment, the Staphylococcus aureus sgRNA scaffold polynucleotide sequence is as follows:
[0340] GUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGCGAGA.
[0341] In one embodiment, the BhCas12b sgRNA scaffold has the following polynucleotide sequence:
[0342] GUUCUGTCUUUUGGUCAGGACAACCGUCUAGCUAUAAGUGCUGCAGGGUGUGAGAAACUCCUAUUGCUGGACGAUGUCUCUUACGAGGCAUUAGCAC.
[0343] In one embodiment, the BvCas12b sgRNA scaffold has the following polynucleotide sequence:
[0344] GACCUAUAGGGUCAAUGAAUCUGUGCGUGUGCCAUAAGUAAUUAAAAAUUACCCACCACAGGAGCACCUGAAAACAGGUGCUUGGCAC.
[0345] In some embodiments, the domain derived from a CRISPR protein incorporated into a base editor is an endonuclease (e.g., a deoxyribonuclease or a ribonuclease) that is capable of binding to a target polynucleotide when acting in concert with a bound guide nucleic acid. In some embodiments, the domain derived from a CRISPR protein incorporated into a base editor is a nickase that is capable of binding to a target polynucleotide when acting in concert with a bound guide nucleic acid. In some embodiments, the domain derived from a CRISPR protein incorporated into a base editor is a catalytically dead domain that is capable of binding to a target polynucleotide when acting in concert with a bound guide nucleic acid. In some embodiments, the target polynucleotide bound by the domain derived from a CRISPR protein of a base editor is DNA. In some embodiments, the target polynucleotide bound by the domain derived from a CRISPR protein of a base editor is RNA.
[0346] Cas proteins that can be used herein include class 1 and class 2 Cas proteins. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cm Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i, CARF, DinG, homologs thereof, or modified versions thereof. Unmodified CRISPR enzymes can have DNA cleavage activity, such as Cas9, which has two endonuclease domains: RuvC and HNH. The CRISPR enzyme can direct cleavage of one or both strands of a target sequence, such as within the target sequence and / or within the complement of the target sequence. For example, the CRISPR enzyme can direct cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence.
[0347] The following vectors can be used, which encode a CRISPR enzyme that is mutated relative to the corresponding wild-type enzyme such that the mutated CRISPR enzyme lacks the ability to cleave one or both strands of a target polynucleotide containing a target sequence. Cas9 can refer to a polypeptide having at least or at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity and / or sequence homology to a wild-type exemplary Cas9 polypeptide (e.g., Cas9 from Streptococcus pyogenes). Cas9 can refer to a polypeptide having at most or at most about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity and / or sequence homology to a wild-type exemplary Cas9 polypeptide (e.g., from Streptococcus pyogenes). Cas9 can refer to a wild-type or modified form of a Cas9 protein that can comprise amino acid changes such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof.
[0348] In some embodiments, the CRISPR protein-derived domain of the base editor may include all or a portion of Cas9 from Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma (NCBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_021315.1); Ref: NC_018010.1); Psychroflexus torquis (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref: YP_820832.1); Listeria innocua (NCBI Ref: NP_472073.1); Campylobacter jejuni (NCBI Ref: YP_002344900.1); Neisseria meningitidis (NCBI Ref: YP_002342100.1); Streptococcus pyogenes or Staphylococcus aureus.
[0349] Cas9 domains of nucleobase editors
[0350] The Cas9 nuclease sequence and structure are well known to those skilled in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski J., et al., “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” K., Sharma C. M., Gonzales K., Chao Y., Pirzada Z. A., Eckert M. R., Vogel J., Charpentier E., Nature 471:602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease inadaptive bacterial immunity." Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J. A., Charpentier E. Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including S. pyogenes and S. thermophilus.Other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci as disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference.
[0351] In some aspects, the nucleic acid programmable DNA binding domain (napDNAbp) is a Cas9 domain. Non-limiting exemplary Cas9 domains are provided herein. The Cas9 domain can be a Cas9 domain of nuclease activity, a Cas9 domain of inactive nuclease, or a Cas9 nickase. In some embodiments, the Cas9 domain is a domain of nuclease activity. For example, the Cas9 domain can be a Cas9 domain that cuts off two chains of a double-helical nucleic acid (for example, two chains of a double-helical DNA molecule). In some embodiments, the Cas9 domain includes any one of the amino acid sequences described in detail herein. In some embodiments, the Cas9 domain includes an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% identical to any one of the amino acid sequences described in detail herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the amino acid sequences detailed herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical contiguous amino acid residues compared to any of the amino acid sequences detailed herein.
[0352] In some embodiments, proteins comprising fragments of Cas9 are provided. For example, in some embodiments, the protein comprises one or two Cas9 domains: (1) the gRNA binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, proteins comprising Cas9 or a fragment thereof are referred to as "Cas9 variants." Cas9 variants share homology with Cas9 or a fragment thereof. For example, a Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to wild-type Cas9. In some embodiments, the Cas9 variant can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid changes compared to wild-type Cas9. In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain) such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of the corresponding wild-type Cas9. In some embodiments, the fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids in length.
[0353] In some embodiments, the Cas9 fusion protein as provided herein comprises the full-length amino acid sequence of a Cas9 protein (e.g., one of the Cas9 sequences provided herein). However, in other embodiments, the fusion protein provided herein does not comprise the full-length Cas9 sequence, but only comprises one or more fragments thereof. Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and other suitable sequences of Cas9 domains and fragments will be apparent to those skilled in the art.
[0354] The Cas9 protein can be associated with a guide RNA that guides the Cas9 protein to a specific DNA sequence that is complementary to the guide RNA. In some embodiments, the polynucleotide programmable nucleotide binding domain is a Cas9 domain, for example, a nuclease-active Cas9, a Cas9 nickase (nCas9), or a nuclease-inactive Cas9 (dCas9). Examples of nucleic acid programmable DNA binding domains include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpf1, Cas12b / C2C1, and Cas12c / C2C3.
[0355] In some embodiments, the wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI Reference Sequence: NC_017053.1, nucleotide and amino acid sequences below).
[0356]
[0357]
[0358]
[0359] (Single underline: HNH domain; double underline: RuvC domain)
[0360] In some embodiments, the wild-type Cas9 corresponds to or comprises the following nucleotide and / or amino acid sequence:
[0361]
[0362]
[0363]
[0364] (Single underline: HNH domain; double underline: RuvC domain).
[0365] In some embodiments, the wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes pyogenes (NCBI Reference Sequence: NC_002737.2 (nucleotide sequence below); and Uniprot Reference Sequence: Q99ZW2 (amino acid sequence below):
[0366]
[0367]
[0368] (Single underline: HNH domain; double underline: RuvC domain)
[0369] In some embodiments, Cas9 refers to Cas9 from Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1), Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1), Spiroplasma yrphidicola (NCBI Ref: NC_021284.1), Prevotella intermedia (NCBI Ref: NC_017861.1), Spiroplasma (NCBI Ref: NC_021846.1), Streptococcus iniae (NCBI Ref: NC_021314.1), Belliella baltica (NCBI Ref: NC_018010.1), Psychroflexus torquis (NCBI Ref: NC_018721.1), Streptococcus thermophilus (NCBI Ref: YP_820832.1), Listeria innocua (NCBI Ref: NP_472073.1), Campylobacter jejuni (NCBI Ref: P_002344900.1), or Neisseria meningitidis (NCBI Ref: YP_002342100.1); or refers to Cas9 from any other organism.
[0370] It will be appreciated that other Cas9 proteins (e.g., nuclease-dead Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9), including variants and homologs thereof, are within the scope of the present disclosure. Exemplary Cas9 proteins include, but are not limited to, those provided below. In some embodiments, the Cas9 protein is a nuclease-dead Cas9 (dCas9). In some embodiments, the Cas9 protein is a Cas9 nickase (nCas9). In some embodiments, the Cas9 protein is a nuclease-active Cas9.
[0371] In some embodiments, the Cas9 domain is a Cas9 domain (dCas9) of an inactive nuclease. For example, the dCas9 domain can be bound to a double-helical nucleic acid molecule (for example, via a gRNA molecule) without cutting any chain of the double-helical nucleic acid molecule. In some embodiments, the dCas9 domain of the inactive nuclease comprises a D10X mutation and an H840X mutation of the amino acid sequence described in detail herein, or a corresponding mutation in any amino acid sequence provided herein, wherein X is any amino acid change. In some embodiments, the dCas9 domain of the inactive nuclease comprises a D10A mutation and an H840A mutation of the amino acid sequence described in detail herein, or a corresponding mutation in any amino acid sequence provided herein. As an example, the Cas9 domain of the inactive nuclease comprises the amino acid sequence described in detail in the cloning vector pPlatTET-gRNA2 (accession number BAV54124).
[0372]
[0373] (See, e.g., Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression.” Cell. 2013; 152(5): 1173-83, the entire contents of which are incorporated herein by reference).
[0374] In some embodiments, the Cas9 nuclease has an inactivated (e.g., inactivated) DNA cleavage domain, in other words, Cas9 is a nickase, referred to as a "nCas9" protein (for "nickase" Cas9). Cas9 proteins with inactive nucleases can be interchangeably referred to as "dCas9" proteins (for nuclease "dead" Cas9) or catalytically inactive Cas9. It is known to generate Cas9 proteins (or fragments thereof) with inactive DNA cleavage domains (see, e.g., Jinek et al., Science. 337: 816-821 (2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression" (2013) Cell. 28; 152 (5): 1173-83, the entire contents of each of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to include two subdomains: the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations in these subdomains can silence the nuclease activity of Cas9. For example, mutations D10A and H840A completely inactivate the nuclease activity of Streptococcus pyogenes Cas9 (Jinek et al., Science. 337: 816-821 (2012); Qi et al., Cell. 28; 152(5): 1173-83 (2013)).
[0375] In some embodiments, the dCas9 domain comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the dCas9 domains provided herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the amino acid sequences detailed herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical contiguous amino acid residues compared to any of the amino acid sequences detailed herein.
[0376] In some embodiments, the dCas9 corresponds to or comprises, in part or in whole, a Cas9 amino acid sequence having one or more mutations that inactivate the Cas9 nuclease activity. For example, in some embodiments, the dCas9 domain comprises D10A and H840A mutations or corresponding mutations in another Cas9.
[0377] In some embodiments, the dCas9 comprises the amino acid sequence of dCas9(D10A and H840A):
[0378]
[0379] (Single underline: HNH domain; double underline: RuvC domain).
[0380] In some embodiments, the Cas9 domain comprises a D10A mutation, while the residue at position 840 in the amino acid sequence provided above, or the corresponding position in any of the amino acid sequences provided herein, remains a histidine.
[0381] In other embodiments, dCas9 variants are provided having mutations other than D10A and H840A, which result in nuclease-inactive Cas9 (dCas9). For example, such mutations include other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). In some embodiments, variants or homologs of dCas9 are provided that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical. In some embodiments, dCas9 variants are provided having amino acid sequences that are shorter or longer, differing by about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids, or more.
[0382] In some embodiments, the Cas9 domain is a Cas9 nickase. The Cas9 nickase can be a Cas9 protein that can cut only one chain of a double-helical nucleic acid molecule (e.g., a double-helical DNA molecule). In some embodiments, the Cas9 nickase cuts the target chain of the double-helical nucleic acid molecule, meaning that the Cas9 nickase cuts the chain that is base-paired with the gRNA (e.g., sgRNA) bound to Cas9. In some embodiments, the Cas9 nickase comprises a D10A mutation and has histidine at position 840. In some embodiments, the Cas9 nickase cuts the non-target, non-base editing chain of the double-helical nucleic acid molecule, meaning that the Cas9 nickase cuts the chain that is not base-paired with the gRNA (e.g., sgRNA) bound to Cas9. In some embodiments, the Cas9 nickase comprises an H840A mutation and has an aspartic acid residue at position 10, or a corresponding mutation. In some embodiments, the Cas9 nickase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the Cas9 nickases provided herein. Based on this disclosure and knowledge in the art, other suitable Cas9 nickases will be apparent to those skilled in the art and are within the scope of this disclosure.
[0383] The amino acid sequence of an exemplary catalytic Cas9 nickase (nCas9) is as follows:
[0384]
[0385] In some embodiments, Cas9 refers to Cas9 from Archaea (e.g., Nanoarchaea), which constitute the domain and kingdom of unicellular prokaryotic microorganisms. In some embodiments, the programmable nucleotide binding protein can be asX or CasY protein, which has been described in, for example, Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res.2017Feb 21.doi:10.1038 / cr.2017.21, the entire contents of which are incorporated herein by reference. Using genome-resolved metagenomics, a large number of CRISPR-Cas systems have been identified, including Cas9 reported for the first time in the archaeal life domain. This different Cas9 protein was discovered as an active CRISPR-Cas system in rarely studied Nanoarchaea. Two previously unknown systems, CRISPR-CasX and CRISPR-CasY, have been found in bacteria, which are one of the most compact systems discovered so far. In some embodiments, in the base editor system described herein, Cas9 is replaced by CasX or a variant of CasX. In some embodiments, in the base editor systems described herein, Cas9 is replaced by CasY or a variant of CasY. It should be understood that other RNA-guided DNA binding proteins can be used as nucleic acid programmable DNA binding proteins (napDNAbp) and are within the scope of the present disclosure.
[0386] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) of any fusion protein provided herein can be CasX or CasY protein. In some embodiments, napDNAbp is CasX protein. In some embodiments, napDNAbp is CasY protein. In some embodiments, napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring CasX or CasY protein. In some embodiments, the programmable nucleotide binding protein is a naturally occurring CasX or CasY protein. In some embodiments, the programmable nucleotide binding protein comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a CasX or CasY protein described herein. It will be appreciated that CasX and CasY from other bacterial species can also be used according to the present disclosure.
[0387] An exemplary CasX ((uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53)tr|F0NN87|F0NN87_SULIHCRISPR-associated Casx protein OS=Sulfolobus islandicus (strain HVE10 / 4) GN=SiH_0402 PE=4SV=1) amino acid sequence is as follows:
[0388] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLE VEPHYLIIAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAVGQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG.
[0389] An exemplary CasX (>tr|F0NH53|F0NH53_SULIR CRISPR-associated protein, Casx OS=Sulfolobus icelandica (strain REY15A) GN=SiRe_0771 PE=4SV=1) amino acid sequence is as follows:
[0390] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAAAGWVLTRLGKAKV SEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAVGQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG.
[0391] DeltaProteobacteria CasX
[0392] MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMTDDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAILQVYWQEFKDDHVGLMCKFAQPASKKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFVYKLEQVSEKGKAYTNYFGRCNVAEHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESTHPVKPLAQIAGNRYASGPVGKALSDACMGTIASFLSKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDfAYNEVIARVRMWVNLLWQKLKLSRDDAKPLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENPKKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEMDEKEFYACEIQLQ KWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFTDGTDIKKSGKWQGLLYGGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKIGRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQAAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAKLAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKELSAELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSYKSGKQPFVGAWQAFYKRRLKEVWKPNA
[0393] An exemplary CasY ((ncbi.nlm.nih.gov / protein / APG80656.1)>APG80656.1 CRISPR-associated protein CasY [uncultured Parcubacteria group bacteria]) amino acid sequence is as follows:
[0394]
[0395] In some embodiments, nucleic acid programmable DNA binding protein (napDNAbp) is a single effector of a microbial CRISPR-Cas system. Single effectors of microbial CRISPR-Cas systems include, but are not limited to, Cas9, Cpf1, Cas12b / C2c1, and Cas12c / C2c3. Typically, microbial CRISPR-Cas systems are divided into Class 1 and Class 2 systems. Class 1 systems have multi-subunit effector complexes, while Class 2 systems have single protein effectors. For example, Cas9 and Cpf1 are Class 2 effectors. In addition to Cas9 and Cpf1, three different Class 2 CRISPR-Cas systems (Cas12b / C2c1 and Cas12c / C2c3) have been described by Shmakov et al., "Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems", Mol. Cell, 2015 Nov. 5; 60(3): 385-397, the entire contents of which are incorporated herein by reference. Two of the effectors in this system, Cas12b / C2c1 and Cas12c / C2c3, contain a RuvC-like endonuclease domain similar to that of Cpf1. The third system contains an effector with two well-characterized HEPN RNase domains. Unlike CRISPR RNA production by Cas12b / C2c1, production of mature CRISPR RNA is tracrRNA-dependent. Cas12b / C2c1 relies on both CRISPR RNA and tracrRNA for DNA cleavage.
[0396] The crystal structure of Alicyclobaccillus acidoterrastris Cas12b / C2c1 (AacC2c1) has been reported in a complex with a chimeric single-molecule guide RNA (sgRNA). See, for example, Liu et al., "C2c1-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism", Mol. Cell, 2017 Jan. 19; 65(2): 310-322, the entire contents of which are incorporated herein by reference. The crystal structure is also reported as a ternary complex in Alicyclobaccillus acidoterrastris C2c1 bound to a target DNA. See, for example, Yang et al., "PAM-dependent Target DNA Recognition and Cleavage by C2C1 CRISPR-Cas endonuclease", Cell, 2016 Dec. 15; 167(7): 1814-1828, the entire contents of which are incorporated herein by reference. A catalytically competent conformation of AacC2c1, in which the target and non-target DNA strands have been independently captured and positioned within a single RuvC catalytic pocket, and Cas12b / C2c1-mediated cleavage results in staggered heptanucleotide breaks of the target DNA. Structural comparisons between the Cas12b / C2c1 ternary complex and previously identified Cas9 and Cpf1 counterparts demonstrate the diversity of mechanisms used by the CRISPR-Cas9 system.
[0397] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) of any fusion protein provided herein can be Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, napDNAbp is Cas12b / C2c1 protein. In some embodiments, napDNAbp is Cas12c / C2c3 protein. In some embodiments, napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, napDNAbp is naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the napDNAbp sequences provided herein. It will be appreciated that Cas12b / C2c1 and Cas12c / C2c3 from other bacterial species can also be used in accordance with the present disclosure.
[0398] Cas12b / C2c1((uniprot.org / uniprot / T0D7A2#2)sp|T0D7A2|
[0399] C2C1_ALIAG CRISPR-associated endonuclease C2c1 OS = Alicyclobacillus acidoterrestris (strain ATCC49025 / DSM 3922 / CIP 106132 / NCIMB 13137 / GD3B) GN = c2c1 PE = 1 SV = 1) amino acid sequence is as follows:
[0400]
[0401] BhCas12b (Bacillus hisashii) NCBI reference sequence: WP_095142515
[0402]
[0403] In some embodiments, Cas12b is BvCas12B, which is a variant of BhCas12b and comprises the following changes relative to BhCas12B: S893R, K846R, and E837G.
[0404] BvCas12b (Bacillus sp. V3-13) NCBI reference sequence: WP_101661451.1
[0405]
[0406] The Cas9 nuclease has two functional endonuclease domains: RuvC and HNH. Cas9 undergoes a conformational change upon target binding, positioning the nuclease domains to cleave the opposite strand of the target DNA. The end result of Cas9-mediated DNA cleavage is a double-strand break (DSB) within the target DNA (~3-4 nucleotides upstream of the PAM sequence). The resulting DSB is then repaired via one of two general repair pathways: (1) the efficient but error-prone non-homologous end joining (NHEJ) pathway; or (2) the less efficient but high-fidelity homology-directed repair (HDR) pathway.
[0407] The "efficiency" of non-homologous end joining (NHEJ) and / or homology-guided repair (HDR) can be calculated by any convenient method. For example, in some cases, efficiency can be expressed as the percentage of successful HDR. For example, surveyor nuclease determination can be used to generate cleavage products, and the ratio of product to substrate can be used to calculate the percentage. For example, surveyor nuclease can be used, which directly cuts the DNA containing the newly integrated restriction sequence, and the DNA is the result of successful HDR. The more substrates that are cut, the higher the percentage of HDR (the higher the efficiency of HDR). As an illustrative example, the score (percentage) of HDR can be calculated using the following equation: [(cut product) / (substrate plus cut product)] (for example, (b+c) / (a+b+c), wherein "a" is the band intensity of the DNA substrate, and "b" and "c" are cut products).
[0408] In some cases, efficiency can be expressed as a percentage of successful NHEJ. For example, a T7 endonuclease I assay can be used to generate cleavage products, and the ratio of product to substrate can be used to calculate the NHEJ percentage. T7 endonuclease I cuts mismatched duplex DNA hybridized from wild-type and mutant DNA chains (NHEJ generates small random insertions or deletions (indels) at the site of the original break). The more cuts, the higher the percentage of NHEJ (the higher the efficiency of NHEJ). As an illustrative example, the fraction (percentage) of NHEJ can be calculated using the following equation: (1-(1-(b+c) / (a+b+c)) 1 / 2 )×100, where “a” is the band intensity of the DNA substrate, and “b” and “c” are the cleavage products (Ran et al., Cell. 2013 Sep. 12; 154(6): 1380-9; and Ran et al., Nat Protoc. 2013 Nov.; 8(11): 2281–2308).
[0409] The NHEJ repair pathway is the most active repair pathway and frequently causes small nucleotide insertions or deletions (indels) at DSB sites. The randomness of NHEJ-mediated DSB repair has important practical significance because cell populations expressing Cas9 and gRNA or guide polynucleotides can result in mutations of different configurations. In most cases, NHEJ produces small indels in the target DNA, resulting in amino acid deletions, insertions, or frameshift mutations, leading to premature stop codons within the open reading frame (ORF) of the targeted gene. The ideal end result is a loss-of-function mutation within the targeted gene.
[0410] Although NHEJ-mediated DSB repair often disrupts the gene's open reading frame, homology-directed repair (HDR) can be used to generate specific nucleotide changes ranging from single nucleotide changes to large insertions, such as the addition of fluorophores or tags.
[0411] In order to utilize HDR for gene editing, gRNA and Cas9 or Cas9 nickase can be used to deliver a DNA repair template containing the desired sequence to the cell type of interest. The repair template may contain the desired editor and homologous sequences (called left and right homology arms) located immediately upstream and downstream of the target. The length of each homology arm may depend on the size of the variation introduced, and the larger the insertion, the longer the required homology arm. The repair template may be a single-stranded oligonucleotide, a double-stranded oligonucleotide, or a double-stranded DNA plasmid. Even in cells expressing Cas9, gRNA, and an exogenous repair template, the efficiency of HDR is generally low (<10% of modified alleles). The efficiency of HDR can be enhanced by synchronizing the cells because HDR occurs in the S and G2 phases of the cell cycle. Chemically or genetically inhibiting genes involved in NHEJ can also increase the frequency of HDR.
[0412] In some embodiments, Cas9 is a modified Cas9. A given gRNA targeting sequence may have additional sites throughout the genome where partial homology exists. These sites are called off-target and should be considered when gRNA is involved. In addition to optimizing gRNA design, CRISPR specificity can also be increased by modifying Cas9. Cas9 produces double-strand breaks (DSBs) through the combined activity of two nuclease domains, RuvC and HNH. The D10A mutant of Cas9 nickase, SpCas9, maintains one nuclease domain and produces DNA nicks instead of DSBs. The nickase system can also be combined with HDR-mediated editing for specific gene editing.
[0413] In some cases, Cas9 is a variant Cas9 protein. The variant Cas9 polypeptide has an amino acid sequence that differs by one amino acid (e.g., with deletion, insertion, substitution, fusion) from the amino acid sequence of the wild-type Cas9 protein. In some instances, the variant Cas9 polypeptide has an amino acid change (e.g., deletion, insertion, or substitution) that reduces the nuclease activity of the Cas9 polypeptide. For example, in some instances, the variant Cas9 polypeptide has less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nuclease activity of the corresponding wild-type Cas9 protein. In some cases, the variant Cas9 protein has substantially no nuclease activity. When the Cas9 protein under test is a variant Cas9 protein that has substantially no nuclease activity, it may be referred to as "dCas9".
[0414] In some cases, the variant Cas9 protein has reduced nuclease activity. For example, the variant Cas9 protein exhibits less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% of the endonuclease activity of a wild-type Cas9 protein (e.g., a wild-type Cas9 protein).
[0415] In some cases, the variant Cas9 protein can cut the complementary strand of the guide target nucleic acid, but the ability to cut the non-complementary strand of the double-stranded guide target sequence is reduced. For example, the variant Cas9 protein may have a mutation (amino acid substitution) that reduces the function of the RuvC domain. As a non-limiting example, in some embodiments, the variant Cas9 protein has D10A (aspartic acid instead of alanine at amino acid position 10), and therefore can cut the complementary strand of the double-stranded guide target sequence but the ability to cut the non-complementary strand of the double-stranded guide target sequence is reduced (thus when the variant Cas9 protein cuts the double-stranded target nucleic acid, it causes a single-strand break (SSB) rather than a double-strand break (DSB)) (see, e.g., Jinek et al., Science. 2012 Aug. 17; 337 (6096): 816-21).
[0416] In some cases, the variant Cas9 protein can cut the non-complementary strand of a double-stranded guide target nucleic acid, but has a reduced ability to cut the complementary strand of the guide target sequence. For example, the variant Cas9 protein may have a mutation (amino acid substitution) that reduces the function of the HNH domain (RuvC / HNH / RuvC domain motif). As a non-limiting example, in some embodiments, the variant Cas9 protein has an H840A (histidine instead of alanine at amino acid position 840) mutation, and thus can cut the non-complementary strand of the guide target sequence but has a reduced ability to cut the complementary strand of the guide target sequence (thus, when the variant Cas9 protein cuts the double-stranded guide target sequence, it results in an SSB rather than a DSB). This Cas9 protein has a reduced ability to cut a guide target sequence (e.g., a single-stranded guide target sequence), but retains the ability to bind to a guide target sequence (e.g., a single-stranded guide target sequence).
[0417] In some cases, the variant Cas9 protein has a reduced ability to cut both the complementary and non-complementary strands of a double-stranded target DNA. As a non-limiting example, in some cases, the variant Cas9 protein carries both the D10A mutation and the H840A mutation, which reduces the ability of the polypeptide to cut both the complementary and non-complementary strands of a double-stranded target DNA. This Cas9 protein has a reduced ability to cut target DNA (e.g., single-stranded target DNA) but retains the ability to bind to target DNA (e.g., single-stranded target DNA).
[0418] As another non-limiting example, in some cases, the variant Cas9 protein carries W476A and W1126A mutations, such that the ability of the polypeptide to cut target DNA is reduced. Such Cas9 protein has a reduced ability to cut target DNA (e.g., single-stranded target DNA) but retains the ability to bind to target DNA (e.g., single-stranded target DNA).
[0419] As another non-limiting example, in some cases, the variant Cas9 protein carries P475A, W476A, N477A, D1125A, W1126A and D1127A mutations, such that the ability of the polypeptide to cut target DNA is reduced. Such Cas9 protein has a reduced ability to cut target DNA (e.g., single-stranded target DNA) but retains the ability to bind to target DNA (e.g., single-stranded target DNA).
[0420] As another non-limiting example, in some cases, the variant Cas9 protein carries H840A, W476A and W1126A mutations, so that the ability of the polypeptide to cut the target DNA is reduced. The ability of this Cas9 protein to cut the target DNA (for example, single-stranded target DNA) is reduced, but the ability to bind the target DNA (for example, single-stranded target DNA) is maintained. As another non-limiting example, in some cases, the variant Cas9 protein carries H840A, D10A, W476A and W1126A mutations, so that the ability of the polypeptide to cut the target DNA is reduced. The ability of this Cas9 protein to cut the target DNA (for example, single-stranded target DNA) is reduced, but the ability to bind the target DNA (for example, single-stranded target DNA) is maintained. In an embodiment, the variant Cas9 has a catalytic His residue (A840H) recovered at position 840 in the Cas9 HNH domain.
[0421] As another non-limiting example, in some cases, the variant Cas9 protein carries H840A, P475A, W476A, N477A, D1125A, W1126A and D1127A mutations, so that the ability of the polypeptide to cut the target DNA is reduced. The ability of this Cas9 protein to cut the target DNA (e.g., single-stranded target DNA) is reduced, but the ability to bind to the target DNA (e.g., single-stranded target DNA) is maintained. As another non-limiting example, in some cases, the variant Cas9 protein carries D10A, H840A, P475A, W476A, N477A, D1125A, W1126A and D1127A mutations, so that the ability of the polypeptide to cut the target DNA is reduced. The ability of this Cas9 protein to cut the target DNA (e.g., single-stranded target DNA) is reduced, but the ability to bind to the target DNA (e.g., single-stranded target DNA) is maintained. In some cases, when the variant Cas9 protein carries W476A and W1126A mutations or when the variant Cas9 protein carries P475A, W476A, N477A, D1125A, W1126A and D1127A mutations, the variant Cas9 protein does not effectively bind to the PAM sequence. Therefore, in some such cases, when such variant Cas9 protein is used in a binding method, the method does not require a PAM sequence. In other words, in some cases, when such variant Cas9 protein is used in a binding method, the method may include a guide RNA, but the method can be performed in the absence of a PAM sequence (and the binding specificity is therefore provided by the targeting segment of the guide RNA). Other residues can be mutated to achieve the above effects (i.e., inactivating one or other nuclease protein). As non-limiting examples, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986 and / or A987 may be altered (i.e., substituted). Furthermore, mutations other than alanine substitutions are suitable.
[0422] In some embodiments, the variant Cas9 protein has reduced catalytic activity (i.e., when the Cas9 protein has D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986 and / or A987 mutations, for example, D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A and / or D986A), the variant Cas9 protein can still bind to the target DNA in a site-specific manner (because it is still guided to the target DNA sequence by the guide RNA) as long as it retains the ability to interact with the guide RNA.
[0423] In some embodiments, the variant Cas protein can be spCas9, spCas9-VRQR, spCas9-VRER, xCas9(sp), saCas9, saCas9-KKH, SpCas9-MQKFRAER, spCas9-MQKSER, spCas9-LRKIQK, or spCas9-LRVSQL.
[0424] In a specific embodiment, a SpCas9 comprising amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R and having specificity for an altered PAM 5'-NGC-3' (SpCas9-MQKFRAER) is used.
[0425] Alternatives to S. pyogenes Cas9 can include RNA-guided endonucleases from the Cpf1 family, which exhibit cleavage activity in mammalian cells. CRISPR from the genera Prevotella and Francisella 1 (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. Cpf1 is an RNA-guided endonuclease of the class II CRISPR / Cas system. This acquired immunity mechanism is found in bacteria of the genera Prevotella and Francisella. The Cpf1 gene is associated with the CRISPR locus and encodes an endonuclease that uses guide RNA to find and cut viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9, overcoming some limitations of the CRISPR / Cas9 system. Unlike the Cas9 nuclease, the result of Cpf1-mediated DNA cleavage is a double-strand break with a short 3' overhang. The staggered cleavage pattern of Cpf1 can create the potential for targeted gene transfer, similar to traditional restriction enzyme cloning, which can increase the efficiency of gene editing. Similar to the aforementioned Cas9 variants and homologs, Cpf1 can also expand the number of sites that can be targeted by CRISPR to AT-rich regions or AT-rich genomic sites that lack the NGG PAM site preferred by SpCas9. The Cpf1 locus contains a mixed α / β domain, RuvC-I, followed by a helical region, RuvC-II, and zinc finger-like domains. The Cpf1 protein possesses a RuvC-like endonuclease domain similar to the RuvC domain of Cas9. In addition, Cpf1 lacks the HNH endonuclease domain, and the N-terminus of Cpf1 lacks the α-helical recognition lobe of Cas9. The Cpf1 CRISPR-Cas domain architecture reveals that Cpf1 is functionally unique and is classified as a Class 2, Type V CRISPR system. The Cpf1 locus encodes Cas1, Cas2, and Cas4 proteins, which are more similar to type I and type III systems than type II systems. Functional Cpf1 does not require reverse activation of CRISPR RNA (tracrRNA), so only CRISPR (crRNA) is required. This is conducive to genome editing because TCpf1 is not only smaller than Cas9, but it also has a smaller sgRNA molecule (the number of nucleotides is about half the number of nucleotides in Cas9). In contrast to the G-rich PAM targeted by Cas9, the Cpf1-crRNA complex cuts the target DNA or RNA by identifying the protospacer sequence adjacent to the motif 5'-YTN-3'. After identifying the PAM, Cpf1 introduces a sticky-end-like DNA double-strand break of 4 or 5 nucleotide protrusions.
[0426] Some aspects of the present disclosure provide nucleic acid programmable DNA binding protein domain and deaminase domain. Some aspects of the present disclosure provide a fusion protein comprising a domain acting as a nucleic acid programmable DNA binding protein, which can be used to guide proteins such as base editors to specific nucleic acid (e.g., DNA or RNA) sequences. In a specific embodiment, the fusion protein comprises a nucleic acid programmable DNA binding protein domain and a deaminase domain. DNA binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpf1, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i. An example of a programmable polynucleotide binding protein with a PAM specificity different from Cas9 is a clustered regularly spaced short palindromic repeat sequence from Prevotella and Francisella 1 (Cpf1). Similar to Cas9, Cpf1 is also a Class 2 CRISPR effector. It has been shown that Cpf1 mediates robust DNA interference with features different from Cas9. Cpf1 is a single RNA-guided endonuclease lacking tracrRNA, and it utilizes a T-rich protospacer sequence adjacent to a motif (TTN, TTN, or YTN). In addition, Cpf1 cuts DNA via staggered DNA double-strand breaks. Among the 16 Cpf1 family proteins, two enzymes from the genus Acidaminococcus and the family Lachnospiraceae were shown to have effective genome editing activity in human cells. Cpf1 protein is known in the art and has been previously described, for example, Yamano et al., "Crystal structure of Cpf1in complex with guideRNA and target DNA." Cell (165) 2016, p.949-962; its entire contents are incorporated herein by reference.
[0427] Also useful in the compositions and methods of the present invention are Cpf1 (dCpf1) variants of inactive nucleases, which can be used as guide nucleotide sequence programmable polynucleotide binding protein domains. The Cpf1 protein has a RuvC-like nuclease domain similar to the RuvC domain of Cas9 but without an HNH nuclease domain, and the N-terminus of Cpf1 does not have the α-helical recognition lobe of Cas9. Zetsche et al., Cell, 163, 759-771, 2015 (which is incorporated herein by reference) shows that the RuvC-like domain of Cpf1 is responsible for cutting the two DNA chains, and inactivation of the RuvC-like domain inactivates the Cpf1 nuclease activity. For example, mutations corresponding to D917A, E1006A, or D1255A in Francisella novicida Cpf1 inactivate the Cpf1 nuclease activity. In some embodiments, dCpf1 of the present disclosure comprises a mutation corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A. It will be appreciated that any mutation that inactivates the RuvC domain of Cpf1, e.g., a substitution mutation, deletion, or insertion, can be used in accordance with the present disclosure.
[0428] In some embodiments, the nucleic acid programmable nucleotide binding protein of any fusion protein provided herein can be a Cpf1 protein. In some embodiments, the Cpf1 protein is a Cpf1 nickase (nCpf1). In some embodiments, the Cpf1 protein is a nuclease-inactive Cpf1 (dCpf1). In some embodiments, Cpf1, nCpf1, or dCpf1 comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a Cpf1 sequence disclosed herein. In some embodiments, dCpf1 comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the Cpf1 disclosed herein, and comprises a mutation corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A. It will be appreciated that Cpf1 from other bacterial species can also be used according to the present disclosure.
[0429] The remaining amino acid sequence of Francisella novicida Cpf1 is as follows: D917, E1006, and D1255 are bolded and underlined.
[0430]
[0431] The amino acid sequence of Francisella novicida Cpf1 D917A is shown below (A917, E1006, and D1255 are bold and underlined).
[0432]
[0433] The amino acid sequence of Francisella novicida Cpf1 E1006A is shown below (D917, A1006, and D1255 are bold and underlined).
[0434]
[0435]
[0436] The amino acid sequence of Francisella novicida Cpf1 D1255A is shown below. (The D917, E1006, and A1255 mutation positions are bolded and underlined).
[0437]
[0438] The amino acid sequence of Francisella novicida Cpf1 D917A / E1006A is shown below (A917, A1006, and D1255 are bold and underlined).
[0439]
[0440]
[0441] The amino acid sequence of Francisella novicida Cpf1 D917A / D1255A is shown below (A917, E1006, and A1255 are bold and underlined).
[0442]
[0443] The amino acid sequence of Francisella novicida Cpf1E1006A / D1255A is shown below (D917, A1006, and A1255 are bold and underlined).
[0444]
[0445] The amino acid sequence of Francisella novicida Cpf1 D917A / E1006A / D1255A is shown below (A917, A1006, and A1255 are bold and underlined).
[0446]
[0447]
[0448] In some embodiments, one of the Cas9 domains present in the fusion protein can be replaced with a guide nucleotide sequence-programmable DNA-binding protein domain that does not require a PAM sequence.
[0449] In some embodiments, the Cas domain is a Cas9 domain from Staphylococcus aureus (SaCas9). In some embodiments, the SaCas9 domain is a nuclease-active SaCas9, a nuclease-inactive SaCas9 (SaCas9d), or a SaCas9 nickase (SaCas9n). In some embodiments, the SaCas9 domain comprises an N579A mutation, or a corresponding mutation in any of the amino acid sequences provided herein.
[0450] In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to a nucleic acid sequence having a non-standard PAM. In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to a nucleic acid sequence having a NNGRRT or NNGRRT PAM sequence. In some embodiments, the SaCas9 domain comprises one or more of E781X, N967X, and R1014X mutations, or corresponding mutations in any amino acid sequence provided herein, wherein X is any amino acid. In some embodiments, the SaCas9 domain comprises one or more of E781K, N967K, and R1014H mutations, or one or more corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SaCas9 domain comprises E781K, N967K, or R1014H mutations, or corresponding mutations in any amino acid sequence provided herein.
[0451] The amino acid sequence of an exemplary SaCas9 is as follows;
[0452]
[0453]
[0454] In this sequence, the underlined and bolded residue N579 can be mutated (e.g., to A579) to obtain the SaCas9 nickase.
[0455] The amino acid sequence of an exemplary SaCas9n is as follows:
[0456]
[0457] In this sequence, residue A579, which can be mutated from N579 to obtain the SaCas9 nickase, is underlined and bolded.
[0458] The amino acid sequence of an exemplary SaKKH Cas9 is as follows;
[0459]
[0460]
[0461] The above residue A579, which can be mutated from N579 to obtain the SaCas9 nickase, is underlined and bolded. The above residues K781, K967 and H1014, which can be mutated from E781, N967 and R1014 to obtain SaKKH Cas9, are underlined and shown in italics.
[0462] High-fidelity Cas9 domain
[0463] Some aspects of the present disclosure provide high-fidelity Cas9 domains. In some embodiments, the high-fidelity Cas9 domain is an engineered Cas9 domain comprising one or more mutations, which, compared to the corresponding wild-type Cas9 domain, reduces the electrostatic interaction between the Cas9 domain and the sugar-phosphate backbone of the DNA. The high-fidelity Cas9 domain with reduced electrostatic interaction with the sugar-phosphate backbone of the DNA may have a lower off-target effect. In some embodiments, the Cas9 domain (e.g., a wild-type Cas9 domain) comprises one or more mutations that reduce the association between the Cas9 domain and the sugar-phosphate backbone of the DNA. In some embodiments, the Cas9 domain comprises one or more mutations that reduce the association of the Cas9 domain with the sugar-phosphate backbone of DNA by at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, or at least 70%.
[0464] In some embodiments, any Cas9 fusion protein provided herein comprises one or more of N497X, R661X, Q695X and / or Q926X mutations, or corresponding mutations in any amino acid sequence provided herein, wherein X is any amino acid. In some embodiments, any Cas9 fusion protein provided herein comprises one or more of N497A, R661A, Q695A and / or Q926A mutations, or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the Cas9 domain comprises a D10A mutation, or a corresponding mutation in any amino acid sequence provided herein. Cas9 domains with high fidelity are known in the art and will be apparent to those skilled in the art. For example, Cas9 domains with high fidelity have been described in Kleinstiver, BP, et al. “High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects.” Nature 529, 490-495 (2016) and Slaymaker, IM, et al. “Rationally engineered Cas9 nucleases with improved specificity.” Science 351, 84-88 (2015), the entire contents of each of which are incorporated herein by reference.
[0465] In some embodiments, the modified Cas9 is a high-fidelity Cas9 enzyme. In some embodiments, the high-fidelity Cas9 enzyme is SpCas9 (K855A), eSpCas9 (1.1), SpCas9-HF1, or an ultra-precise Cas9 variant (HypaCas9). The modified Cas9eSpCas9 (1.1) contains an alanine substitution that weakens the interaction between the HNH / RuvC groove and the non-target DNA chain, preventing chain separation and cutting at the off-target site. Similarly, SpCas9-HF1 reduces off-target labeling by alanine substitution, which interrupts the interaction of Cas9 with the DNA phosphate backbone. HypaCas9 contains mutations (SpCas9 N692A / M694A / Q695A / H698A) in the REC3 domain that increase Cas9 proofreading and target recognition. All three high-fidelity enzymes produce off-target editing lower than wild-type Cas9.
[0466] Exemplary high-fidelity Cas9s are provided below.
[0467] The high-fidelity Cas9 domain relative to Cas9 is shown in bold and underlined
[0468]
[0469] Guide polynucleotides
[0470] In one embodiment, the guide polynucleotide is a guide RNA. The RNA / Cas complex can help "guide" the Cas protein to the target DNA. Cas9 / crRNA / tracrRNA endonucleolytically cuts linear or circular dsDNA targets complementary to the spacer sequence. The target chain that is not complementary to the crRNA is first cut endonucleolytically and then trimmed exonucleolytically by 3'-5'. In nature, DNA binding and cutting usually require proteins and two RNAs. However, single guide RNA ("sgRNA" or abbreviated as "gNRA") can be engineered to incorporate multiple crRNAs and tracrRNAs into a single RNA substance. See, for example, Jinek M. et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes short motifs (PAM or protospacer sequence adjacent motifs) in CRISPR repeat sequences to help distinguish self from non-self. The Cas9 nuclease sequence and structure are well known to those skilled in the art (see, for example, "Complete genome sequence of an M1 strain of Streptococcus pyogenes." Ferretti, JJ et al., Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E. et al., Nature 471:602-607 (2011); and "Programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M. et al, Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including Streptococcus pyogenes and Streptococcus thermophilus.Based on the present disclosure, other suitable Cas9 nucleases and sequences may be apparent to those skilled in the art, and such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci, as disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5,726-737; the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease has an inactivated (e.g., inactivated) DNA cleavage domain, in other words, Cas9 is a nickase.
[0471] In some embodiments, the guide polynucleotide is at least one single guide RNA ("sgRNA" or "gRNA"). In some embodiments, the guide polynucleotide is at least one tracrRNA. In some embodiments, the guide polynucleotide does not require a PAM sequence to guide a polynucleotide programmable DNA binding domain (e.g., Cas9 or Cpf1) to a target nucleotide sequence.
[0472] The polynucleotide programmable nucleotide binding domains of the base editor disclosed herein (e.g., domains derived from CRISPR) can identify target polynucleotide sequences by associating with guide polynucleotides. Guide polynucleotides (e.g., gRNA) are typically single-stranded and programmable to bind to polynucleotide target sequences site-specifically (i.e., via complementary base pairing), thereby guiding the base editor that cooperates with the guide polynucleotides to the target sequence. The guide polynucleotides can be DNA. The guide polynucleotides can be RNA. In some cases, the guide polynucleotides include natural nucleotides (e.g., adenosine). In some cases, the guide polynucleotides include non-natural (or non-natural) nucleotides (e.g., peptide nucleic acids or nucleotide analogs). In some cases, the targeting region of the guide nucleic acid sequence can be at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides in length. The targeting region of the guide nucleic acid can be between 10 and 30 nucleotides in length, or between 15 and 25 nucleotides in length, or between 15 and 20 nucleotides in length.
[0473] In some embodiments, the guide polynucleotide comprises two or more individual polynucleotides that can interact with each other via, for example, complementary base pairing (e.g., a dual guide polynucleotide). For example, a guide polynucleotide can comprise a CRISPR RNA (crRNA) and a trans-activating CRISPR RNA (tracrRNA). For example, a guide polynucleotide can comprise one or more trans-activating CRISPR RNAs (tracrRNAs).
[0474] In a type II CRISPR system, targeting nucleic acids by a CRISPR protein (e.g., Cas9) typically requires complementary base pairing between a first RNA molecule (crRNA) containing a sequence that recognizes the target sequence and a second RNA molecule (trRNA) containing a repeat sequence of a scaffold region that forms a stabilized guide RNA-CRISPR protein complex. Such dual-guide RNA systems can be used as guide polynucleotides to guide the base editors disclosed herein to the target polynucleotide sequence.
[0475] In some embodiments, the base editors provided herein utilize a single guide polynucleotide (e.g., gRNA). In some embodiments, the base editors provided herein utilize a dual guide polynucleotide (e.g., two gRNAs). In some embodiments, the base editors provided herein utilize one or more guide polynucleotides (e.g., multiple gRNAs). In some embodiments, a single guide polynucleotide is used for different base editors described herein. For example, a single guide polynucleotide can be used for a cytidine base editor and an adenosine base editor.
[0476] In other embodiments, the guide polynucleotide can include both a polynucleotide that targets a portion of a nucleic acid and a scaffold portion of the nucleic acid in one molecule (i.e., a single-molecule guide nucleic acid). For example, a single-molecule guide polynucleotide can be a single guide RNA (sgRNA or gRNA). As used herein, the term guide polynucleotide sequence refers to any single-molecule, bi-molecule, or multi-molecule nucleic acid that is capable of interacting with a base editor and guiding the base editor to a target polynucleotide sequence.
[0477] Typically, a guide polynucleotide (e.g., a crRNA / trRNA complex or gRNA) comprises a: a "polynucleotide targeting segment" that includes a sequence that can recognize and bind to a target polynucleotide sequence; and a "protein binding segment" that stabilizes the guide polynucleotide within the polynucleotide programmable nucleotide binding domain component of the base editor. In some embodiments, the polynucleotide targeting segment of the guide polynucleotide recognizes and binds to a DNA polynucleotide, thereby facilitating the editing of bases in the DNA. In other cases, the polynucleotide targeting segment of the guide polynucleotide recognizes and binds to an RNA polynucleotide, thereby facilitating the editing of bases in the RNA. As used herein, a "segment" refers to a segment or region of a molecule, for example, a continuous stretch of nucleotides in a guide polynucleotide. A segment may also refer to a region / segment of a complex, such that a segment may include a region of more than one molecule. For example, if the guide polynucleotide comprises multiple nucleic acid molecules, the protein binding segment may include, for example, all or a portion of multiple independent molecules that hybridize along a region of complementarity. In some embodiments, a protein-binding segment of a DNA-targeting RNA comprising two separate molecules comprises (i) base pairs 40 to 75 of a first RNA that is 100 base pairs in length and (ii) base pairs 10 to 25 of a second RNA molecule that is 50 base pairs in length. Unless specifically defined otherwise in a particular context, the definition of "segment" is not limited to a specific number of total base pairs, is not limited to a specific number of base pairs from any given RNA molecule, is not limited to a specific number of separate molecules within a complex, and can include regions of RNA molecules that can be of any total length and can include regions of complementarity to other molecules.
[0478] The guide RNA or guide polynucleotide may comprise two or more RNAs, for example, CRISPR RNA (crRNA) and reverse activating crRNA (tracrRNA). Sometimes, the guide RNA or guide polynucleotide may comprise single-stranded RNA, or a single guide RNA (sgRNA) formed by fusing a portion of crRNA (e.g., a functional portion) to tracrRNA. The guide RNA or guide polynucleotide may also be a double RNA comprising crRNA and tracrRNA. In addition, crRNA can hybridize with target DNA.
[0479] As described above, the guide RNA or guide polynucleotide can be an expression product. For example, the DNA encoding the guide RNA can be a vector comprising a sequence encoding the guide RNA. The guide RNA or guide polynucleotide can be transferred into the cell by transfecting the cell with an isolated guide RNA or a plasmid DNA comprising a sequence encoding the guide RNA and a promoter. The guide RNA or guide polynucleotide can also be transferred into the cell by other means, such as using viral-mediated gene delivery.
[0480] The guide RNA or guide polynucleotide can be isolated. For example, the guide RNA can be transferred into a cell or organism as an isolated RNA. The guide RNA can be prepared by in vitro transcription using any in vitro transcription system known in the art. The guide RNA can be transferred into a cell as an isolated RNA rather than as a plasmid containing the coding sequence of the guide RNA.
[0481] The guide RNA or guide polynucleotide can contain three regions: a first region at the 5' end that can be complementary to the target site in the chromosomal sequence; a second internal region that can form a stem-loop structure; and a third 3' region that can be single-stranded. The first region of each guide RNA can also be different, so that each guide RNA guides the fusion protein to a specific target site. Furthermore, the second and third regions of each guide RNA can be the same in all guide RNAs.
[0482] The first region of the guide RNA or guide polynucleotide can be complementary to the sequence at the target site in the chromosome sequence, so that the first region of the guide RNA can be base paired with the target site. In some cases, the guide RNA may comprise about 10 nucleotides to 25 nucleotides (i.e., 10 nucleotides to 25 nucleotides; or about 10 nucleotides to about 25 nucleotides; or 10 nucleotides to about 25 nucleotides; or about 10 nucleotides to 25 nucleotides) or more. For example, the base pairing region between the first region of the guide RNA and the target site in the chromosome sequence can be (approximately) 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25 or more nucleotides in length. Sometimes, the first region of the guide RNA can be a length of (approximately) 19, 20 or 21 nucleotides.
[0483] The guide RNA or guide polynucleotide may also include a second region that forms a secondary structure. For example, the secondary structure formed by the guide RNA may include a stem (or hairpin) and a loop. The length of the loop and the stem can be variable. For example, the loop can be within a length range of about 3 to 10 nucleotides, and the stem can be within a length range of about 6 to 20 base pairs. The stem can include one or more protrusions of 1 to 10 or about 10 nucleotides. The overall length of the second region can be within a length range of (approximately) 16 to 60 nucleotides. For example, the loop can be a length of (approximately) 4 nucleotides, and the stem can be (approximately) 12 base pairs.
[0484] The guide RNA or guide polynucleotide may also include a third region at the 3' end, which may be substantially single-stranded. For example, the third region may not be complementary to any chromosomal sequence in the cell of interest and may not be complementary to the remainder of the guide RNA. Furthermore, the length of the third region may be variable. The third region may be greater than (approximately) 4 nucleotides in length. For example, the length of the third region may be in the range of (approximately) 5 to 60 nucleotides in length.
[0485] A guide RNA or guide polynucleotide can target any exon or intron of a gene target. In some cases, a guide can target exon 1 or 2 of a gene, and in other cases, a guide can target exon 3 or 4 of a gene. A composition can include multiple guide RNAs that all target the same exon, or in some cases, multiple guide RNAs can target different exons. Exons and introns of a gene can be targeted.
[0486] The guide RNA or guide polynucleotide can target a nucleic acid sequence of (about) 20 nucleotides. The target nucleic acid can be less than (about) 20 nucleotides. The target nucleic acid can be at least or at least about 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, or any one between 1 and 100 nucleotides in length. The target nucleic acid can be at most or at most about 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, or any one between 1 and 100 nucleotides in length. The target nucleic acid sequence can be the (about) 20 bases immediately 5' to the first nucleotide of the PAM. The guide RNA can target a nucleic acid sequence. The target nucleic acid can be at least or at least about 1 to 10, 1 to 20, 1 to 30, 1 to 40, 1 to 50, 1 to 60, 1 to 70, 1 to 80, 1 to 90, or 1 to 100 nucleotides.
[0487] A guide polynucleotide, e.g., a guide RNA, can refer to a nucleic acid that can hybridize to another nucleic acid (e.g., a target nucleic acid or a protospacer sequence in the genome of a cell). A guide polynucleotide can be RNA. A guide polynucleotide can be DNA. A guide polynucleotide can be programmed or designed to site-specifically bind to a sequence of a nucleic acid. A guide polynucleotide can comprise one polynucleotide chain and can be referred to as a single guide polynucleotide. A guide polynucleotide can comprise two polynucleotide chains and can be referred to as a dual guide polynucleotide. A guide RNA can be introduced into a cell as an RNA molecule. For example, an RNA molecule can be in vitro transcribed and / or can be chemically synthesized. The RNA can be derived from a synthetic DNA molecule (e.g., Gene fragment) transcription. The guide RNA can then be introduced into the cell as an RNA molecule. The guide RNA can also be introduced into the cell in the form of a non-RNA nucleic acid molecule (e.g., a DNA molecule). For example, the DNA encoding the guide RNA can be operably linked to a promoter control sequence to express the guide RNA in the cell of interest. The RNA coding sequence can be operably linked to a promoter sequence recognized by RNA polymerase III (Pol III). Plasmid vectors that can be used to express guide RNA include, but are not limited to, px330 vectors and px333 vectors. In some cases, a plasmid vector (e.g., px333 vector) may contain at least two DNA sequences encoding the guide RNA.
[0488] Methods for selecting, directing and verifying guide polynucleotides (e.g., guides) and targeting sequences are described herein and are known to those skilled in the art. For example, in order to minimize the impact of potential substrate confusion of the deaminase domain (e.g., AID domain) in the nucleobase editor system, the number of residues that accidentally become deamination targets (e.g., off-target C residues that potentially reside on ssDNA in the target nucleic acid locus) can be minimized. In addition, software tools can be used to optimize the gRNA corresponding to the target nucleic acid sequence, for example, to minimize the total off-target activity across the genome. For example, for each possible targeting structure and selection using Streptococcus pyogenes Cas9, all off-target sequences (previously selected PAM, e.g., NAG or NGG) can be identified across the genome, and the genome contains a maximum of a certain number (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10) of mismatched base pairs. The first region of the gRNA complementary to the target site can be identified, and all first regions (e.g., crRNA) can be ranked according to their total predicted off-target scores; the top-ranked targeting domains represent those that appear to have the highest on-target activity and the lowest off-target activity. Candidate targeting gRNA can be functionally evaluated using methods known in the art and / or as described herein.
[0489] As a non-limiting example, the target DNA hybridization sequence in the crRNA of the guide RNA used in combination with Cas9 can be identified using a DNA sequence search algorithm. Custom gRNA design software can be used to design gRNA based on the public tool cas-offinder, as described in Bae S., Park J., & Kim J.-S.Cas-OFFinder:A fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guide dendonucleases.Bioinformatics 30,1473-1475 (2014). This software scores these guides after calculating the whole genome off-target tendency of the guide. Typically, for guides with a length in the range of 17 to 24, a perfect match to a match within the range of 7 mismatches is considered. Once the off-target site is determined by calculation, the total score of each guide is calculated and summarized in a table output using a web interface. In addition to identifying the potential target site adjacent to the PAM sequence, the software also identifies all PAM adjacent sequences, which are different from the selected target site by 1, 2, 3 or more than 3 nucleotides. The genomic DNA sequence for the target nucleic acid sequence (e.g., target gene) can be obtained, and publicly available tools (e.g., RepeatMasker program) can be used to screen for repetitive elements. RepeatMasker searches the DNA sequence input for repetitive elements and low complexity regions. The output is a detailed annotation of the repetitive sequences present in a given query sequence.
[0490] After identification, the first regions of the guide RNA (e.g., crRNA) can be ranked based on their distance to the target site, their orthogonality, and the presence of a 5' nucleotide for a close match to a relevant PAM sequence (e.g., based on a closely matching 5'G in the human genome containing a relevant PAM (e.g., NGG PAM for S. pyogenes, NNGRRT or NNGRRV PAM for S. aureus). As used herein, orthogonality refers to the number of sequences in the human genome that contain the lowest number of mismatches to the target sequence. "High level of orthogonality" or "good orthogonality" can refer to, for example, a 20-mer targeting domain that has neither identical sequences in the human genome other than the desired target nor any sequences that contain mismatches in one or both target sequences. Targeting domains with good orthogonality can be selected to minimize off-target DNA cleavage.
[0491] In some embodiments, a reporter gene system can be used to detect base editing activity and test candidate guide polynucleotides. In some embodiments, the reporter gene system may include a reporter gene-based assay, wherein the base editing activity results in the expression of a reporter gene. For example, a reporter gene system may include a reporter gene that includes a deactivated start codon, for example, a mutation on the template chain from 3'-TAC-5' to 3'-CAC-5'. After the target C is successfully deaminated, the corresponding mRNA will be transcribed as 5'-AUG-3' rather than 5'-GUG-3', enabling translation of the reporter gene. Suitable reporter genes will be apparent to those skilled in the art. Non-limiting examples of reporter genes include genes encoding: green fluorescent protein (GFP), red fluorescent protein (RFP), luciferase, secreted alkaline phosphatase (SEAP), or any other gene whose expression is detectable and apparent to those skilled in the art. The reporter gene system can be used to test a variety of different gRNAs, for example, to determine which residue(s) of the target DNA sequence each deaminase will target. In order to evaluate specific base editing proteins (e.g., Cas9 deaminase fusion proteins), sgRNAs targeting non-template chains can also be tested. In some embodiments, such gRNAs can be designed so that the mutated start codon does not base pair with the gRNA. The guide polynucleotides may comprise standard ribonucleotides, modified ribonucleotides (e.g., pseudouridine), ribonucleotide isomers and / or ribonucleotide analogs. In some embodiments, the guide polynucleotides may comprise at least one detectable label. The detectable label can be a fluorophore (e.g., FAM, TMR, Cy3, Cy5, Texas Red, Oregon Green, Alexa Fluors, Halo tags or suitable fluorescent dyes), detection tags (e.g., biotin, digoxin, etc.), quantum dots or gold particles.
[0492] The guide polynucleotide can be chemically synthesized, enzymatically synthesized, or a combination thereof. For example, the guide RNA can be synthesized using a standard solid phase synthesis method based on phosphoramidite. Alternatively, the guide RNA can be synthesized in vitro by operably linking the DNA encoding the guide RNA to a promoter control sequence recognized by a bacteriophage or RNA polymerase. Examples of suitable phage promoter sequences include T7, T3, SP6 promoter sequences, or variants thereof. In embodiments where the guide RNA comprises two independent molecules (e.g., crRNA and tracrRNA), the crRNA can be chemically synthesized and the tracrRNA can be enzymatically synthesized.
[0493] In some embodiments, the base editor system may include multiple guide polynucleotides, for example, gRNAs. For example, the gRNAs may be targeted to one or more target loci (e.g., at least 1 gRNA, at least 2 gRNAs, at least 5 gRNAs, at least 10 gRNAs, at least 20 gRNAs, at least 30 gRNAs, at least 50 gRNAs) included in the base editor system. Multiple gRNA sequences can be arranged in series and preferably separated by direct repeat sequences.
[0494] The DNA encoding the guide RNA or guide polynucleotide can also be part of the vector. Furthermore, the vector can contain additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), selectable marker sequences (e.g., GFP or antibiotic resistance genes such as puromycin), replication origins, etc. The DNA molecule encoding the guide RNA (gRNA) can also be linear. The DNA molecule encoding the guide RNA (gRNA) or guide polynucleotide can also be circular.
[0495] In some embodiments, one or more components of the base editor system can be encoded by a DNA sequence. Such DNA sequences can be introduced into an expression system (e.g., a cell) together or independently. For example, a DNA sequence encoding a polynucleotide programmable nucleotide binding domain and a guide RNA can be introduced into a cell, and each DNA sequence can be part of an independent molecule (e.g., one vector contains a polynucleotide programmable nucleotide binding domain encoding sequence, and a second vector contains a guide RNA encoding sequence) or both can be part of the same molecule (e.g., a vector contains coding (and regulatory) sequences for polynucleotide programmable nucleotide binding domains and guide RNAs).
[0496] Guide polynucleotides can include one or more modifications to provide nucleic acids with new or enhanced characteristics. Guide polynucleotides can include nucleic acid affinity tags. Guide polynucleotides can include synthetic nucleotides, synthetic nucleotide analogs, nucleotide derivatives, and / or modified nucleotides.
[0497] In some cases, the gRNA or guide polynucleotide may comprise modifications. Modifications may be made at any position of the gRNA or guide polynucleotide. More than one modification may be made to a single gRNA or guide polynucleotide. The gRNA or guide polynucleotide may undergo quality control after modification. In some cases, quality control may include PAGE, HPLC, MS, or any combination thereof.
[0498] Modifications of the gRNA or guide polynucleotide can be substitutions, insertions, deletions, chemical modifications, physical modifications, stabilization, purification, or any combination thereof.
[0499] The gRNA or guide polynucleotide can also be modified by the following: 5' adenylation, 5' guanosine triphosphate blocking, 5' N7-methylguanosine triphosphate blocking, 5' triphosphate blocking, 3' phosphate, 3' phosphorothioate, 5' phosphate, 5' phosphorothioate and, Cis-Syn thymidine dimer, trimer, C12 spacer, C3 spacer, C6 spacer, dSpacer, PC spacer, rSpacer, Spacer 18, Spacer 9, 3'-3' modification, 5'-5' modification, abasic, acridine, azobenzene, biotin, biotin BB, biotin TEG, cholesteryl TEG, desthiobiotin TEG, DNP TEG, DNP-X, DOTA, dT-biotin, bibiotin, PC biotin, psoralen C2, psoralen C6, TINA, 3'DABCYL, black hole quencher 1, black hole quencher 2, DABCYL SE, dT-DABCYL, IRDye QC-1, QSY-21, QSY-35, QSY-7, QSY-9, carboxyl linker, thiol linker, 2'-deoxyribonucleoside analog purine, 2'-deoxyribonucleoside analog pyrimidine, ribonucleoside analog, 2'-O-methyl ribonucleoside analog, sugar modification analog, wobble / universal base, fluorescent dye label, 2'-fluoro RNA, 2'-O-methyl RNA, methylphosphonate, phosphodiester DNA, phosphodiester RNA, phosphorothioate DNA, phosphorothioate RNA, UNA, guanosine-5'-triphosphate, 5'-methylcytidine-5'-triphosphate, or any combination thereof.
[0500] In some cases, the modification is permanent. In some cases, the modification is temporary. In some cases, multiple modifications are made to the gRNA or guide polynucleotide. The gRNA or guide polynucleotide modifications can alter the physiochemical properties of the nucleotide, such as its conformation, polarity, hydrophobicity, chemical reactivity, base pairing interactions, or any combination thereof.
[0501] The PAM sequence can be any PAM sequence known in the art. Suitable PAM sequences include, but are not limited to, NGG, NGA, NGC, NGN, NGT, NGCG, NGAG, NGAN, NGNG, NGCN, NGCG, NGTN, NNGRRT, NNNRRT, NNGRR(N), TTTV, TYCV, TYCV, TATV, NNNNGATT, NNAGAAW, or NAAAC. Y is a pyrimidine; N is any nucleotide base; and W is A or T.
[0502] The modification can also be a phosphorothioate substitute. In some cases, the natural phosphodiester bond may be easily rapidly degraded by cellular nucleases; and, the modification of the internucleotide linkage using a phosphorothioate (PS) bond substitute may be more stable and less susceptible to hydrolysis by cellular degradation. The modification can increase stability in the gRNA or guide polynucleotide. The modification can also enhance biological activity. In some cases, the phosphorothioate-enhanced RNA gRNA can inhibit RNase A, RNaseT1, calf serum nuclease, or any combination thereof. These properties allow the use of PS-RNA gRNA to be used for specific applications, in which exposure to nucleases is a high-probability event in vivo or in vitro. For example, a phosphorothioate (PS) bond can be introduced between the last 3 to 5 nucleotides at the 5'- or ''-end of the gRNA, which can inhibit exonuclease degradation. In some cases, phosphorothioate bonds can be added to the entire gRNA to reduce attack by exonucleases.
[0503] protospacer adjacent motif
[0504] The term "protospacer adjacent motif (PAM)" or "PAM-like motif" refers to a 2 to 6 indirect DNA sequence that immediately follows the DNA sequence targeted by the Cas9 nuclease in the CRISPR bacterial adaptive immune system. In some embodiments, the PAM can be a 5' PAM (i.e., located upstream of the 5' end of the protospacer sequence). In other embodiments, the PAM can be a 3' PAM (i.e., located downstream of the 5' end of the protospacer sequence).
[0505] The PAM sequence is crucial for target binding, but the exact sequence depends on the type of Cas protein.
[0506] The base editors provided herein may include a domain derived from a CRISPR protein that is capable of binding to a nucleotide sequence containing a standard or non-standard protospacer sequence adjacent to a motif (PAM) sequence. A PAM site is a nucleotide sequence close to a target polynucleotide sequence. Some aspects of the present disclosure provide base editors that include all or part of a CRISPR protein with different PAM specificities. For example, Cas9 proteins, such as Cas9 (spCas9) from Streptococcus pyogenes, typically require a standard NGG PAM sequence to bind to a specific nucleic acid region, where the “N” in “NGG” is adenine (A), thymine (T), guanine (G) or cytosine (C), and G is guanine. PAM can be CRISPR protein-specific and can differ between different base editors comprising different domains derived from CRISPR. PAM can be 5' or 3' of the target sequence. PAM can be upstream or downstream of the target sequence. PAM can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotides in length. PAMs are often between 2 and 6 nucleotides in length. Several PAM variants are described in Table 1 below.
[0507] Table 1. Cas9 proteins and corresponding PAM sequences
[0508] Variants PAM spCas9 NGG spCas9-VRQR NGA spCas9-VRER NGCG SpCas9-MQKFRAER NGC xCas9(sp) NGN saCas9 NNGRRT saCas9-KKH NNNRRT spCas9-MQKSER NGCG spCas9-MQKSER NGCN spCas9-LRKIQK NGTN spCas9-LRVSQK NGTN spCas9-LRVSQL NGTN SpyMacCas9 NAA Cpf1 5'(TTTV)
[0509] In some embodiments, the PAM is NGC. In some embodiments, the NGC PAM is recognized by a Cas9 variant. In some embodiments, the NGC PAM variant comprises one or more amino acid substitutions selected from D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (collectively referred to as "MQKFRAER").
[0510] In some embodiments, the PAM is NGT. In some embodiments, the NGT PAM is a variant. In some embodiments, the NGT PAM variant is generated by targeted mutation at one or more residues 1335, 1337, 1135, 1136, 1218, and / or 1219. In some embodiments, the NGT PAM variant is generated by targeted mutation at one or more residues 1219, 1335, 1337, 1218. In some embodiments, the NGT PAM variant is generated by targeted mutation at one or more residues 1135, 1136, 1218, 1219, and 1335. In some embodiments, the NGT PAM variant is selected from the group of targeted mutations provided in Tables 2 and 3 below.
[0511] Table 2: NGT PAM variant mutations at residues 1219, 1335, 1337, 1218
[0512]
[0513]
[0514] Table 3: NGT PAM variant mutations at residues 1135, 1136, 1218, 1219 and 1335
[0515]
[0516]
[0517] In some embodiments, the NGT PAM variant is selected from variants 5, 7, 28, 31, or 36 in Tables 2 and 3. In some embodiments, the variant has improved NGT PAM recognition.
[0518] In some embodiments, the NGT PAM variant has a mutation at residues 1219, 1335, 1337, and / or 1218. In some embodiments, the NGT PAM variant having a mutation for improved recognition is selected from the variants provided in Table 4 below.
[0519] Table 4: NGT PAM variant mutations at residues 1219, 1335, 1337 and 1218
[0520] Variants E1219V R1335Q T1337 G1218 1 F V T 2 F V R 3 F V Q 4 F V L 5 F V T R 6 F V R R 7 F V Q R 8 F V L R
[0521] In some embodiments, the NGT PAM is selected from the variants provided in Table 5 below.
[0522] Table 5. NGT PAM variants
[0523] NGTN variant D1135 S1136 G1218 E1219 A1322R R1335 T1337 Variant 1 LRKIQK L R K I - Q K Variant 2 LRSVQK L R S V - Q K Variant 3 LRSVQL L R S V - Q L Variant 4 LRKIRQK L R K I R Q K Variant 5 LRSVRQK L R S V R Q K Variant 6 LRSVRQL L R S V R Q L
[0524] In some embodiments, the Cas9 domain is a Cas9 domain (SpCas9) from Streptococcus pyogenes. In some embodiments, the SpCas9 domain is a nuclease-active SpCas9, an inactive nuclease-free SpCas9 (SpCas9d) or a SpCas9 nickase (SpCas9n). In some embodiments, SpCas9 comprises a D9X mutation or a corresponding mutation in any amino acid sequence provided herein, wherein X is any amino acid except D. In some embodiments, the SpCas9 domain comprises a D9A mutation, or a corresponding mutation in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain, SpCas9d domain or SpCas9n domain can be bound to a nucleic acid sequence with a non-standard PAM. In some embodiments, the SpCas9 domain, SpCas9d domain or SpCas9n domain can be bound to a nucleic acid sequence with an NGG, NGA or NGCG PAM sequence.
[0525] In some embodiments, the SpCas9 domain comprises one or more of D1135X, R1335X, and T1336X mutations, or corresponding mutations in any amino acid sequence provided herein, wherein X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of D1135E, R1335Q, and T1336R mutations, or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain comprises one or more of D1135E, R1335Q, and T1336R mutations, or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain comprises one or more of D1135X, R1335X, and T1336X mutations, or corresponding mutations in any amino acid sequence provided herein, wherein X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of D1135V, R1335Q, and T1336R mutations, or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain comprises D1135V, R1335Q and T1336R mutations, or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain comprises one or more of D1135X, G1217X, R1335X and T1336X mutations, or corresponding mutations in any amino acid sequence provided herein, wherein X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of D1135V, G1217R, R1335Q and T1336R mutations, or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain comprises D1135V, G1217R, R1335Q and T1336R mutations, or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain comprises Figures 3A to 3C and Figure 10 One or more of the amino acid substitutions described herein.
[0526] In some embodiments, the Cas9 domain of any fusion protein provided herein comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% identical to the Cas9 polypeptide described herein. In some embodiments, the Cas9 domain of any fusion protein provided herein comprises the amino acid sequence of any Cas9 polypeptide described herein. In some embodiments, the Cas9 domain of any fusion protein provided herein consists of the amino acid sequence of any Cas9 polypeptide described herein.
[0527] In some embodiments, a PAM recognized by a CRISPR protein-based domain of a base editor disclosed herein can be provided on a separate oligonucleotide to an insert encoding the base editor (e.g., an AAV insert) to provide it to a cell. In such embodiments, the PAM provided on the separate oligonucleotide can allow cleavage of a target sequence that cannot otherwise be cleaved because there is no adjacent PAM on the same polynucleotide as the target sequence.
[0528] In one embodiment, Streptococcus pyogenes Cas9 (SpCas9) can be used as a CRISPR nuclease for genome engineering. However, others can be used. In some embodiments, different nucleases can be used to target certain genomic targets. In some embodiments, synthetic variants derived from SpCas9 with non-NGG PAM sequences can be used. In addition, Cas9 orthologs from various species have been identified, and these "non-SpCas9" can be combined with a variety of PAM sequences that can also be used in the present disclosure. For example, the relatively large size of SpCas9 (approximately 4kb coding sequence) can result in plasmids carrying SpCas9 cDNA that cannot be effectively expressed in cells. In contrast, the coding sequence of Staphylococcus aureus Cas9 (SaCas9) is approximately 1 kilobase shorter than SpCas9, which may allow it to be effectively expressed in cells. Similar to SpCas9, SaCas9 endonuclease is able to modify target genes in mammalian cells in vitro and is able to modify target genes in mouse cells in vivo. In some embodiments, Cas proteins can target different PAM sequences. In some embodiments, for example, the target gene may be adjacent to the Cas9 PAM, i.e., 5'-NGG. In some embodiments, for example, the target gene may be adjacent to the Cas9 PAM, i.e., 5'-NGC, or a Cas9 PAM comprising 5'-NGC. In other embodiments, other Cas9 orthologs may have different PAM requirements. For example, other PAMs such as those of Streptococcus thermophilus (5'-NNAGAA for CRISPR1 and 5'-NGGNG for CRISPR3) and those of Neisseria meningiditis (5'-NNNNGATT) may also be found adjacent to the target gene.
[0529] In some embodiments, for the Streptococcus pyogenes system, the target gene sequence can be before (i.e., 5') the 5'-NGG PAM, and the 20-nt guide RNA sequence can be base-paired with the opposite strand to direct the Cas9 cleavage adjacent to the PAM. In some embodiments, the adjacent cut can be at (approximately) 3 base pairs upstream of the PAM. In some embodiments, the adjacent cut can be at (approximately) 10 base pairs upstream of the PAM. In some embodiments, the adjacent cut can be at (approximately) 0 to 20 base pairs upstream of the PAM. For example, the adjacent cut can be at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 base pairs immediately upstream of the PAM. The adjacent cut can also be at 1 to 30 base pairs downstream of the PAM. The sequence of an exemplary SpCas9 protein capable of binding to a PAM is as follows:
[0530] An exemplary PAM-binding amino acid sequence of SpCas9 is as follows:
[0531]
[0532] An exemplary PAM-binding amino acid sequence of SpCas9n is as follows:
[0533]
[0534] An exemplary PAM-binding amino acid sequence of SpEQR Cas9 is as follows:
[0535]
[0536] In this sequence, residues E1135, Q1335 and R1337, which can be mutated from D1135, R1335, and T1337 to obtain SpEQR Cas9, is underlined and in bold show.
[0537] The amino acid sequence of an exemplary PAM binding SpVQR Cas9 is as follows:
[0538] In this sequence, residues V1135, Q1335 and R1336, which can be mutated from D1135, R1335, and T1336 to obtain SpVQR Cas9, is underlined and in bold show.
[0539] The amino acid sequence of an exemplary PAM-binding SpVRER Cas9 is as follows:
[0540]
[0541] In some embodiments, the Cas9 domain is a recombinant Cas9 domain. In some embodiments, the recombinant Cas9 domain is a SpyMacCas9 domain. In some embodiments, the SpyMacCas9 domain is a nuclease-active SpyMacCas9, a nuclease-inactive SpyMacCas9 (SpyMacCas9d), or a SpyMacCas9 nickase (SpyMacCas9n). In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to a nucleic acid sequence with a non-standard PAM. In some embodiments, the SpyMacCas9 domain, SpCas9d domain, or SpCas9n domain can bind to a nucleic acid sequence with an NAA PAM sequence.
[0542] Exemplary SpyMacCas9
[0543]
[0544] In some cases, variant Cas9 protein carries H840A, P475A, W476A, N477A, D1125A, W1126A and D1218A mutations, so that the ability of the polypeptide to cut target DNA or RNA is reduced. The ability of this Cas9 protein to cut target DNA (e.g., single-stranded target DNA) is reduced, but the ability to bind target DNA (e.g., single-stranded target DNA) is maintained. As another non-limiting example, in some cases, variant Cas9 protein carries D10A, H840A, P475A, W476A, N477A, D1125A, W1126A and D1218A mutations, so that the ability of the polypeptide to cut target DNA is reduced. The ability of this Cas9 protein to cut target DNA (e.g., single-stranded target DNA) is reduced, but the ability to bind target DNA (e.g., single-stranded target DNA) is maintained. In some cases, when the variant Cas9 protein carries W476A and W1126A mutations or when the variant Cas9 protein carries P475A, W476A, N477A, D1125A, W1126A and D1218A mutations, the variant Cas9 protein does not effectively bind to the PAM sequence. Therefore, in some such cases, when such variant Cas9 protein is used in a binding method, the method does not require a PAM sequence. In other words, in some cases, when such variant Cas9 protein is used in a binding method, the method may include a guide RNA, but the method can be performed in the absence of a PAM sequence (and the binding specificity is therefore provided by the targeting segment of the guide RNA). Other residues can be mutated to achieve the above effects (i.e., inactivating one or other nuclease protein). As non-limiting examples, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986 and / or A987 may be altered (i.e., substituted). Furthermore, mutations other than alanine substitutions are suitable.
[0545] In some embodiments, the domain derived from the CRISPR protein of the base editor may include all or part of a Cas9 protein with a standard PAM sequence (NGG). In other embodiments, the domain derived from Cas9 of the base editor may adopt a non-standard PAM sequence. Such sequences have been described in the art and will be apparent to those skilled in the art. For example, the Cas9 domain incorporating non-standard PAM sequences has been described in Kleinstiver, BP, et al., “Engineered CRISPR-Cas9 nucleases with altered PAM specificities” Nature 523, 481-485 (2015) and Kleinstiver, BP, et al., “Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition” Nature Biotechnology 33, 1293-1298 (2015), the entire contents of each of which are incorporated herein by reference.
[0546] Fusion proteins containing Cas9 domains and cytidine deaminase and / or adenosine deaminase
[0547] Some aspects of the present disclosure provide fusion proteins comprising Cas9 domains or other nucleic acid programmable DNA binding proteins and one or more adenosine deaminase domains, cytidine deaminase domains and / or DNA glycosidase domains. It should be understood that the Cas9 domain can be any Cas9 domain or Cas9 protein (e.g., dCas9 or nCas9) provided herein. In one embodiment, the Cas9 domain is an SpCas9 domain or an SpCas9 variant domain as described herein. In some embodiments, any Cas9 domain or Cas9 protein (e.g., dCas9 or nCas9) provided herein can be fused with any cytidine deaminase and adenosine deaminase provided herein. The mechanism of the base editor disclosed herein can be arranged in any order.
[0548] For example, but not limitation, in some embodiments, the fusion protein comprises the structure:
[0549] NH2-[cytidine deaminase]-[Cas9 domain]-[adenosine deaminase]-COOH;
[0550] NH2-[adenosine deaminase]-[Cas9 domain]-[cytidine deaminase]-COOH;
[0551] NH2-[adenosine deaminase]-[cytidine deaminase]-[Cas9 domain]-COOH;
[0552] NH2-[cytidine deaminase]-[adenosine deaminase]-[Cas9 domain]-COOH;
[0553] NH2-[Cas9 domain]-[adenosine deaminase]-[cytidine deaminase]-COOH; or NH2-[Cas9 domain]-[cytidine deaminase]-[adenosine deaminase]-COOH.
[0554] In some embodiments, the adenosine deaminase of the fusion protein comprises TadA*8 and a cytidine deaminase. In some embodiments, TadA*8 is TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23, or TadA*8.24.
[0555] Exemplary fusion protein structures include the following:
[0556] NH2-[adenosine deaminase]-[Cas9]-[cytidine deaminase]-COOH;
[0557] NH2-[cytidine deaminase]-[Cas9]-[adenosine deaminase]-COOH;
[0558] NH2-[TadA*8]-[Cas9]-[cytidine deaminase]-COOH; or
[0559] NH2-[Cytidine deaminase]-[Cas9]-[TadA*8]-COOH.
[0560] In some embodiments, the fusion protein comprising cytidine deaminase, base-free editor and adenosine deaminase and napDNAbp (e.g., Cas9 domain) does not include a linker sequence. In some embodiments, a linker is present between cytidine deaminase and adenosine deaminase domain and napDNAbp. In some embodiments, the "-" used in the above-mentioned general framework represents the presence of an optional linker. In some embodiments, cytidine deaminase and adenosine deaminase are fused to napDNAbp via any linker provided herein. For example, in some embodiments, cytidine deaminase and adenosine deaminase are fused to napDNAbp via any linker provided in the sections hereinafter entitled "Linker".
[0561] In some embodiments, the general architecture of exemplary Cas9 or Cas12 fusion proteins having cytidine deaminase, adenosine deaminase, and Cas9 or Cas12 domains comprises any of the following structures, wherein NLS is a nuclear localization sequence (e.g., any NLS provided herein), NH2 is the N-terminus of the fusion protein, and COOH is the C-terminus of the fusion protein.
[0562] NH2-NLS-[cytidine deaminase]-[Cas9 domain]-[adenosine deaminase]-COOH;
[0563] NH2-NLS-[adenosine deaminase]-[Cas9 domain]-[cytidine deaminase]-COOH;
[0564] NH2-NLS-[adenosine deaminase][cytidine deaminase]-[Cas9 domain]-COOH;
[0565] NH2-NLS-[cytidine deaminase]-[adenosine deaminase]-[Cas9 domain]-COOH;
[0566] NH2-NLS-[Cas9 domain]-[adenosine deaminase]-[cytidine deaminase]-COOH;
[0567] NH2-NLS-[Cas9 domain]-[cytidine deaminase]-[adenosine deaminase]-COOH;
[0568] NH2-[cytidine deaminase]-[Cas9 domain]-[adenosine deaminase]-NLS-COOH;
[0569] NH2-[adenosine deaminase]-[Cas9 domain]-[cytidine deaminase]-NL2-COOH;
[0570] NH2-[adenosine deaminase][cytidine deaminase]-[Cas9 domain]-NLS-COOH;
[0571] NH2-[cytidine deaminase]-[adenosine deaminase]-[Cas9 domain]-NLS-COOH;
[0572] NH2-[Cas9 domain]-[adenosine deaminase]-[cytidine deaminase]-NLS-COOH; or
[0573] NH2-[Cas9 domain]-[Cytidine deaminase]-[Adenosine deaminase]-NLS-COOH.
[0574] In some embodiments, the NLS is present within the linker, or the NLS is located on the linker flank, for example, as described herein. In some embodiments, the N-terminus or C-terminus of the NLS is a bipartite NLS. The bipartite NLS comprises two clusters of basic amino acids separated by a relatively short spacer sequence (thus bipartite - 2 parts, while monopartite NLS is not). The NLS of nucleoplasmin, KR[PAATKKAGQA]KKKK, is the prototype of the ubiquitous bipartite signal: a cluster of two basic amino acids separated by a spacer sequence of about 10 amino acids. The sequence of an exemplary bipartite NLS is as follows: PKKKRKVEGADKRTADGSEFESPKKKRKV.
[0575] In some embodiments, the fusion protein comprising cytidine deaminase, adenosine deaminase, Cas9 domain and NLS does not comprise a linker sequence. In some embodiments, a linker sequence is present between one or more of the domain or protein (e.g., cytidine deaminase, adenosine deaminase, Cas9 domain or NLS).
[0576] Should be understood that fusion protein of the present disclosure can comprise one or more additional features.For example, in some embodiments, fusion protein can comprise inhibitor, cytoplasm localization sequence, export sequence such as nuclear export sequence or other c localization sequence, and can be used for the sequence tag of solubilization, purification or detection of fusion protein.Suitable protein tag provided herein includes but is not limited to, biotin carboxylase carrier protein (BCCP) label, myc-label, calmodulin-label, FLAG-label, hemagglutinin (HA)-label, polyhistidine label (also referred to as histidine label or His-label), maltose binding protein (MBP)-label, nus-label, glutathione-S-transferase (GST)-label, green fluorescent protein (GFP)-label, thioredoxin-label, S-label, Softag (for example, Softag 1, Softag 3), strep-label, biotin ligase label, F1AsH label, V5 label and SBP-label.Other suitable sequences will be apparent to those skilled in the art.In some embodiments, fusion protein comprises one or more His tags.
[0577] Exemplary, but non-limiting, fusion proteins are described in International PCT Application Nos. PCT / 2017 / 044935 and PCT / US2020 / 016288, each of which is incorporated herein by reference in its entirety.
[0578] Fusion protein containing nuclear localization sequence F (NLS)
[0579] In some embodiments, the fusion protein provided herein further comprises one or more (e.g., 2, 3, 4, 5) nuclear targeting sequences, e.g., a nuclear localization sequence (NLS). In one embodiment, a bipartite NLS is used. In some embodiments, the NLS comprises an amino acid sequence that promotes entry of the protein comprising the NLS into the nucleus (e.g., by nuclear transport). In some embodiments, any fusion protein provided herein further comprises a nuclear localization sequence (NLS). In some embodiments, the NLS is fused to the N-terminus of the fusion protein. In some embodiments, the NLS is fused to the C-terminus of the fusion protein. In some embodiments, the NLS is fused to the N-terminus of the Cas9 domain. In some embodiments, the NLS is fused to the C-terminus of the nCas9 domain or the dCas9 domain. In some embodiments, the NLS is fused to the N-terminus of the deaminase. In some embodiments, the NLS is fused to the C-terminus of the deaminase. In some embodiments, the NLS is fused to the fusion protein via one or more linkers. In some embodiments, the NLS is fused to the fusion protein without a linker. In some embodiments, the NLS comprises the amino acid sequence of any one of the NLS sequences provided or cited herein. Other nuclear localization sequences are known in the art and will be apparent to the skilled artisan. For example, NLS sequences are described in PCT / EP2000 / 011690 to Plank et al., the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences. In some embodiments, the NLS comprises the amino acid sequence PKKKRKVEGADKRTADGSEFESPKKKRKV, KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRKPKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC. In some embodiments, the NLS is present within a linker, or the NLS flanks a linker, e.g., a linker as described herein. In some embodiments, the N-terminus or C-terminus of the NLS is a bipartite NLS. Bipartite NLSs consist of two clusters of basic amino acids separated by a relatively short spacer sequence (hence bipartite - 2 parts, whereas monopartite NLSs do not). The NLS of nucleoplasmin, KR[PAATKKAGQA]KKKK, is the prototype of the ubiquitous bipartite signal: two clusters of basic amino acids separated by a spacer sequence of approximately 10 amino acids. The sequence of an exemplary bipartite NLS is as follows: PKKKRKVEGADKRTADGSEFES PKKKRKV.
[0580] In some embodiments, the fusion protein of the present invention does not include a linker sequence. In some embodiments, there is a linker between one or more of the domain or protein. In some embodiments, the general architecture of the exemplary Cas9 fusion protein with adenosine deaminase or cytidine deaminase and Cas9 domain includes any of the following structures, wherein NLS is a nuclear localization sequence (e.g., any NLS provided herein), NH2 is the N-terminus of the fusion protein, and COOH is the C-terminus of the fusion protein.
[0581] NH2-NLS-[adenosine deaminase]-[Cas9 domain]-COOH;
[0582] NH2-NLS[Cas9 domain]-[adenosine deaminase]-COOH;
[0583] NH2-[adenosine deaminase]-[Cas9 domain]-NLS-COOH;
[0584] NH2-[Cas9 domain]-[adenosine deaminase]-NLS-COOH;
[0585] NH2-NLS-[cytidine deaminase]-[Cas9 domain]-COOH;
[0586] NH2-NLS-[Cas9 domain]-[cytidine deaminase]-COOH;
[0587] NH2-[Cytidine deaminase]-[Cas9 domain]-NLS-COOH; or
[0588] NH2-[Cas9 domain]-[Cytidine deaminase]-NLS-COOH.
[0589] Should be understood that fusion protein of the present disclosure can comprise one or more additional features.For example, in some embodiments, fusion protein can comprise inhibitor, cytoplasm localization sequence, export sequence such as nuclear export sequence or other c localization sequence, and can be used for the sequence tag of solubilization, purification or detection of fusion protein.Suitable protein tag provided herein includes but is not limited to, biotin carboxylase carrier protein (BCCP) label, myc-label, calmodulin-label, FLAG-label, hemagglutinin (HA)-label, polyhistidine label (also referred to as histidine label or His-label), maltose binding protein (MBP)-label, nus-label, glutathione-S-transferase (GST)-label, green fluorescent protein (GFP)-label, thioredoxin-label, S-label, Softag (for example, Softag 1, Softag 3), strep-label, biotin ligase label, F1AsH label, V5 label and SBP-label.Other suitable sequences will be apparent to those skilled in the art.In some embodiments, fusion protein comprises one or more His tags.
[0590] Vectors encoding CRISPR enzymes comprising one or more nuclear localization sequences (NLSs) can be used. For example, (approximately) 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 NLSs can be used. The CRISPR enzyme can comprise an NLS at or near the amino acid at the amino terminus, about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 NLSs at or near the carboxy terminus, or any combination of these (e.g., one or more NLSs at the amino acid and one or more NLSs at the carboxy terminus). When more than one NLS is present, each can be selected independently of the other, such that a single NLS can be present in more than one copy and / or combined with one or more other NLSs in one or more copies.
[0591] The CRISPR enzyme used in the method can comprise about 6 NLSs. An NLS is considered to be near the N-terminus or C-terminus when the amino acid closest to the NLS is within about 50 amino acids (e.g., within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, or 50 amino acids) of the polypeptide chain along the N-terminus to the C-terminus.
[0592] Fusion proteins with internal insertions
[0593] Provided herein is a fusion protein comprising a heterologous polypeptide fused to a nucleic acid programmable nucleic acid binding domain (e.g., napDNAbp). A heterologous polypeptide can be a polypeptide not found in a natural or wild-type napDNAbp polypeptide sequence. A heterologous polypeptide can be fused to napDNAbp at the C-terminus of napDNAbp, at the N-terminus of napDNAbp, or inserted into an internal position of napDNAbp. In some embodiments, a heterologous polypeptide is inserted into an internal position of napDNAbp.
[0594] In some embodiments, the hospital polypeptide is a deaminase or a functional fragment thereof. For example, the fusion polypeptide may comprise a deaminase flanking the N-terminal fragment and the C-terminal fragment of the Cas9 or Cas12 (e.g., Cas12b / C2c1) polypeptide. The deaminase in the fusion protein may be an adenosine deaminase. In some embodiments, the adenosine deaminase is TadA (e.g., TadA7.10 or TadA*8). In some embodiments, TadA is TadA*8. The TadA sequence as described herein (e.g., TadA7.10 or TadA*8) is a deaminase suitable for use in the above-mentioned fusion protein.
[0595] The deaminase can be a circularly arranged deaminase. For example, the deaminase can be a circularly arranged adenosine deaminase. In some embodiments, the deaminase is a circularly arranged TadA, which is circularly arranged at amino acid residue 116 as numbered in the TadA reference sequence. In some embodiments, the deaminase is a circularly arranged TadA, which is circularly arranged at amino acid residue 136 as numbered in the TadA reference sequence. In some embodiments, the deaminase is a circularly arranged TadA, which is circularly arranged at amino acid residue 65 as numbered in the TadA reference sequence.
[0596] The fusion protein may comprise more than one deaminase. The fusion protein may comprise, for example, 1, 2, 3, 4, 5, or more deaminases. In some embodiments, the fusion protein comprises one deaminase. In some embodiments, the fusion protein comprises two deaminases. The two or more deaminases in the fusion protein may be adenosine deaminase, cytidine deaminase, or a combination thereof. The two or more deaminases may be homodimers. The two or more deaminases may be heterodimers. The two or more deaminases may be inserted in tandem into the apDNAbp. In some embodiments, the two or more deaminases may not be inserted in tandem into the apDNAbp.
[0597] In some embodiments, the napDNAbp in the fusion protein is a Cas9 polypeptide or a fragment thereof. The Cas9 polypeptide may be a variant Cas9 polypeptide. In some embodiments, the Cas9 polypeptide is a Cas9 nickase (nCas9) polypeptide or a fragment thereof. In some embodiments, the Cas9 polypeptide is a nuclease-dead Cas9 (dCas9) polypeptide or a fragment thereof. The Cas9 polypeptide in the fusion protein may be a full-length Cas9 polypeptide. In some cases, the Cas9 polypeptide in the fusion protein may not be a full-length Cas9 polypeptide. The Cas9 polypeptide may be truncated at, for example, the N-terminus or C-terminus relative to a naturally occurring Cas9 protein. The Cas9 polypeptide may be a circularly arranged Cas9 protein. The Cas9 polypeptide may be a fragment, portion, or domain of a Cas9 polypeptide that is still capable of binding to a target nucleotide and a guide nucleic acid sequence.
[0598] In some embodiments, the Cas9 polypeptide is Streptococcus pyogenes Cas9 (SpCas9), Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), or a fragment or variant thereof.
[0599] The Cas9 polypeptide of the fusion protein can comprise an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring Cas9 polypeptide.
[0600] The Cas9 polypeptide of the fusion protein may comprise an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the Cas9 amino acid sequence described in detail below (hereinafter referred to as the "Cas9 reference sequence"):
[0601]
[0602] (Single underline: HNH domain; double underline: RuvC domain).
[0603] Fusion proteins comprising heterologous catalytic domains flanking the N-terminus and C-terminus of Cas9 polypeptides can also be used for base editing in the methods described herein. Fusion proteins comprising Cas9 and one or more deaminase domains (e.g., adenosine deaminase) or comprising adenosine deaminase domains flanking Cas9 sequences can also be used for high specificity and efficient base editing of target sequences. In one embodiment, the chimeric Cas9 fusion protein contains a heterologous catalytic domain (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) inserted into the Cas9 polypeptide. In some embodiments, the fusion protein comprises an adenosine deaminase domain and a cytidine deaminase domain inserted into Cas9. In some embodiments, adenosine deaminase is fused into Cas9, and cytidine deaminase is fused to the C-terminus. In some embodiments, adenosine deaminase is fused into Cas9, and cytidine deaminase is fused to the N-terminus. In some embodiments, cytidine deaminase is fused into Cas9, and adenosine deaminase is fused to the C-terminus. In some embodiments, a cytidine deaminase is fused into Cas9 and an adenosine deaminase is fused to the N-terminus.
[0604] An exemplary structure of a fusion protein with adenosine deaminase and cytidine deaminase and Cas9 is provided below:
[0605] NH2-[Cas9(adenosine deaminase)]-[cytidine deaminase]-COOH;
[0606] NH2-[cytidine deaminase]-[Cas9(adenosine deaminase)]-COOH;
[0607] NH2-[Cas9(cytidine deaminase)]-[adenosine deaminase]-COOH; or
[0608] NH2-[Adenosine deaminase]-[Cas9(cytidine deaminase)]-COOH.
[0609] In some embodiments the "-" used in the above general scheme indicates the presence of an optional linker.
[0610] In various embodiments, the catalytic domain has DNA modification activity (e.g., deaminase activity), such as adenosine deaminase activity. In some embodiments, the adenosine deaminase is TadA (e.g., TadA7.10). In some embodiments, TadA is TadA*8. In some embodiments, TadA*8 is fused into Cas9, and the cytidine deaminase is fused to the C-terminus. In some embodiments, TadA*8 is fused into Cas9, and the cytidine deaminase is fused to the N-terminus. In some embodiments, the cytidine deaminase is fused into Cas9, and TadA*8 is fused to the C-terminus. In some embodiments, the cytidine deaminase is fused into Cas9, and TadA*8 is fused to the N-terminus. Exemplary structures of fusion proteins with TadA*8, cytidine deaminase, and Cas9 are provided below:
[0611] NH2-[Cas9(TadA*8)]-[cytidine deaminase]-COOH;
[0612] NH2-[Cytidine deaminase]-[Cas9(TadA*8)]-COOH;
[0613] NH2-[Cas9(TadA*8)]-[TadA*8]-COOH; or
[0614] NH2-[TadA*8]-[Cas9(TadA*8)]-COOH.
[0615] In some embodiments the "-" used in the above general scheme indicates the presence of an optional linker.
[0616] Heterologous polypeptides (e.g., deaminases) can be inserted into napDNAbp (e.g., Cas9 or Cas12 (e.g., Cas12b / C2c1)) at appropriate positions, for example, so that napDNAbp maintains its ability to bind to target polynucleotides and guide nucleic acids. Deaminases (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted into napDNAbp without destroying the function of the deaminase (e.g., base editing activity) or the function of the napDNAbp (e.g., the ability to bind to target nucleic acids and guide nucleic acids). Deaminases (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted into napDNAbp, and the insertion position is at a disordered region as shown in crystallographic studies or at a region comprising a high temperature factor or a B factor. Regions of poor order, disordered regions, or regions with messy structures (e.g., regions and loops exposed to solvents) of proteins can be used for insertion without destroying structure or function. Deaminases (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted into a flexible loop region or a solvent-exposed region within napDNAbp. In some embodiments, deaminases (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) are inserted into a flexible loop of a Cas9 or Cas12b / C2c1 polypeptide.
[0617] In some embodiments, the insertion position of the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is determined by B-factor analysis of the crystal structure of the Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted into a region of the Cas9 polypeptide that contains a B-factor higher than the average value (e.g., compared to the total protein or a protein domain comprising a disordered region, a higher B-factor). The B-factor or temperature factor can represent the fluctuation of an atom from its average position (e.g., as a result of temperature-dependent atomic vibrations or static disorder in a crystal lattice). The high B-factor of a skeleton atom (e.g., higher than the average B-factor) may be a sign of a region with relatively high local mobility. Such a region can be used to insert a deaminase without destroying structure or function. The deaminase (e.g., an adenosine deaminase, a cytidine deaminase, or both an adenosine deaminase and a cytidine deaminase) can be inserted at a position having a residue with a Cα atom whose B-factor for that atom is 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, or more than 200% greater than the average B-factor of the total protein. A deaminase (e.g., an adenosine deaminase, a cytidine deaminase, or both an adenosine deaminase and a cytidine deaminase) can be inserted at a position having a residue having a Cα atom whose B-factor for that atom is 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, or more than 200% higher than the average B-factor of the Cas9 protein domain comprising that residue. Cas9 polypeptide positions comprising higher-than-average B-factors can include, for example, amino acid residues 768, 792, 1052, 1015, 1022, 1026, 1029, 1067, 1040, 1054, 1068, 1246, 1247, and 1248, as numbered in the Cas9 reference sequence above. Regions of the Cas9 polypeptide comprising an above-average number of B factors can include, for example, amino acid residues 792 to 872, 792 to 906, and 2 to 791, as numbered in the Cas9 reference sequence above.
[0618] A heterologous polypeptide (e.g., a deaminase) can be inserted into napDNAbp at an amino acid residue selected from the group consisting of: 768, 791, 792, 1015, 1016, 1022, 1023, 1026, 1029, 1040, 1052, 1054, 1067, 1068, 1069, 1246, 1247, and 1248, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the heterologous polypeptide is inserted between the following amino acid positions: 768 to 769, 791 to 792, 792 to 793, 1015 to 1016, 1022 to 1023, 1026 to 1027, 1029 to 1030, 1040 to 1041, 1052 to 1053, 1054 to 1055, 1067 to 1068, 1068 to 1069, 1247 to 1248, or 1248 to 1249, as numbered in the above-mentioned Cas reference sequence or their corresponding amino acid positions. In some embodiments, the heterologous polypeptide is inserted between the following amino acid positions: 769 to 770, 792 to 793, 793 to 794, 1016 to 1017, 1023 to 1024, 1027 to 1028, 1030 to 1031, 1041 to 1042, 1053 to 1054, 1055 to 1056, 1068 to 1069, 1069 to 1070, 1248 to 1249, or 1249 to 1250, as numbered in the above-mentioned Cas reference sequence or their corresponding amino acid positions. In some embodiments, the heterologous polypeptide replaces amino acid residues selected from the group consisting of: 768, 791, 792, 1015, 1016, 1022, 1023, 1026, 1029, 1040, 1052, 1054, 1067, 1068, 1069, 1246, 1247, and 1248, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide. It should be understood that references to the insertion positions of the above-mentioned Cas9 reference sequences are for illustrative purposes. The insertions discussed herein are not limited to the Cas9 polypeptide sequences of the above-mentioned Cas9 reference sequences, but include insertions at corresponding positions in variant Cas9 polypeptides (e.g., Cas9 nickase (nCas9), nuclease-dead Cas9 (dCas9), Cas9 variants lacking a nuclease domain, truncated Cas9, or Cas9 domains lacking a partial or complete HNH domain).
[0619] A heterologous polypeptide (e.g., a deaminase) can be inserted into the napDNAbp at an amino acid residue selected from the group consisting of: 768, 792, 1022, 1026, 1040, 1068, and 1247, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the heterologous polypeptide is inserted between the following amino acid positions: 768 to 769, 792 to 793, 1022 to 1023, 1026 to 1027, 1029 to 1030, 1040 to 1041, 1068 to 1069, or 1247 to 1248, as numbered in the Cas9 reference sequence above, or the corresponding amino acid positions. In some embodiments, the heterologous polypeptide is inserted between the following amino acid positions: 769 to 770, 793 to 794, 1023 to 1024, 1027 to 1028, 1030 to 1031, 1041 to 1042, 1069 to 1070, or 1248 to 1249, as numbered in the Cas9 reference sequence above, or the corresponding amino acid positions thereof. In some embodiments, the heterologous polypeptide replaces an amino acid residue selected from the group consisting of: 768, 792, 1022, 1026, 1040, 1068, and 1247, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide.
[0620] A heterologous polypeptide (e.g., a deaminase) can be inserted into napDNAbp at an amino acid residue as described herein, or at the corresponding amino acid residue of another Cas9 polypeptide. In one embodiment, a heterologous polypeptide (e.g., a deaminase) can be inserted into napDNAbp at an amino acid residue selected from the group consisting of: 1002, 1003, 1025, 1052 to 1056, 1242 to 1247, 1061 to 1077, 943 to 947, 686 to 691, 569 to 578, 530 to 539, and 1060 to 1077, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide. A deaminase (e.g., an adenosine deaminase, a cytidine deaminase, or an adenosine deaminase and a cytidine deaminase) can be inserted at the N-terminus or C-terminus of the residue or replace the residue. In some embodiments, a deaminase (eg, an adenosine deaminase, a cytidine deaminase, or an adenosine deaminase and a cytidine deaminase) is inserted at the C-terminus of the residue.
[0621] In some embodiments, an adenosine deaminase (e.g., TadA) is inserted at an amino acid residue selected from the group consisting of: 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, an adenosine deaminase (e.g., TadA) is inserted at residues 792 to 872, 792 to 906, or 2 to 791, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the adenosine deaminase is inserted at the N-terminus of amino acids selected from the group consisting of: 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the adenosine deaminase is inserted at the C-terminus of amino acids selected from the group consisting of: 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the adenosine deaminase is inserted to replace an amino acid selected from the group consisting of 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide.
[0622] In some embodiments, the CBE (e.g., APOBEC1) is inserted at an amino acid residue selected from the group consisting of: 1016, 1023, 1029, 1040, 1069, and 1247, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the ABE is inserted at the N-terminus of an amino acid selected from the group consisting of: 1016, 1023, 1029, 1040, 1069, and 1247, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the ABE is inserted at the C-terminus of an amino acid selected from the group consisting of: 1016, 1023, 1029, 1040, 1069, and 1247, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the ABE is inserted to replace an amino acid selected from the group consisting of 1016, 1023, 1029, 1040, 1069, and 1247, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide.
[0623] In some embodiments, a deaminase (e.g., an adenosine deaminase, a cytidine deaminase, or an adenosine deaminase and a cytidine deaminase) is inserted at amino acid residue 768, as numbered in the above-mentioned Cas9 reference sequence, or at the corresponding ...
Claims
1. A use of a base editor complexed with one or more guide polynucleotides for the preparation of a drug for editing a polynucleotide to allow transcription, comprising contacting the polynucleotide with the drug, The base editor comprises a polynucleotide programmable DNA binding domain and a deaminase domain, and wherein one or more of the guide polynucleotides targets the base editor to effectuate an alteration that introduces a mutation that permits transcription.
2. A use of a base editor complexed with one or more guide polynucleotides for the preparation of a medicament for editing an SBDS polynucleotide comprising a mutation associated with Schwann-Diesel syndrome, comprising contacting the SBDS polynucleotide with the medicament, The base editor comprises a polynucleotide programmable DNA binding domain and a deaminase domain, and wherein one or more of the guide polynucleotides targets the base editor to effectuate alteration of a mutation associated with Schwann-Dieter syndrome.
3. Use of a base editor complexed with one or more guide polynucleotides for the preparation of a medicament for editing an SBDS polynucleotide comprising a mutation associated with Schwann-Diesel syndrome, comprising contacting the SBDS polynucleotide with the medicament, The base editor comprises a polynucleotide programmable DNA binding domain and a deaminase domain, and One or more of the guide polynucleotides targets the base editor to achieve a change from A·T to G·C of 183-184TA>CTRs113993991 to generate a missense mutation.
4. A use of a cytidine base editor complexed with one or more guide polynucleotides for the preparation of a medicament for editing an SBDS polynucleotide comprising a mutation associated with Schwann-Diesel syndrome, comprising contacting the SBDS polynucleotide with the medicament, wherein the cytidine base editor comprises a polynucleotide programmable DNA binding domain and a cytidine deaminase domain, and wherein one or more of the guide polynucleotides targets the base editor to effect a C·G to T·A change of rs113993993 258+2T>C.
5. A cell produced by introducing a base editor or a polynucleotide encoding the base editor into the cell or a precursor thereof, wherein the base editor comprises a polynucleotide programmable DNA binding domain, a deaminase domain, and one or more guide polynucleotides; and Wherein the one or more guide polynucleotides target the base editor to effectuate changes associated with aberrant splicing.
6. Use of the cell according to claim 5 for the preparation of a medicament for treating Schwann-Diesel syndrome or a disease associated with abnormal splicing in a subject in need thereof, wherein the medicament is administered to the subject.
7. Use of a base editor or a polynucleotide encoding the base editor for the preparation of a medicament for treating Schwann-Diesel syndrome in a subject, wherein the drug is administered to the subject, wherein the base editor comprises a polynucleotide programmable DNA binding domain, a deaminase domain, and one or more guide polynucleotides; and Wherein the one or more guide polynucleotides target the base editor to effectuate the alteration of a mutation associated with Schwann-Diesel syndrome.
8. Use of a base editor or a polynucleotide encoding the base editor in the preparation of a medicament for treating a genetic disease associated with abnormal splicing in a subject in need thereof, wherein the drug is administered to the subject, wherein the base editor comprises a polynucleotide programmable DNA binding domain, a deaminase domain, and one or more guide polynucleotides; and Wherein the one or more guide polynucleotides target the base editor to effectuate alteration of a pathogenic mutation that alters splicing.
9. A method for producing a cell or a precursor thereof in vitro, the method comprising: (a) introducing a base editor or a polynucleotide encoding the base editor into an induced pluripotent stem cell comprising a gene transformation associated with Schwann-Diesel syndrome, wherein the base editor comprises a polynucleotide programmable nucleotide binding domain and a cytidine deaminase domain or an adenosine deaminase domain and one or more guide polynucleotides, and wherein the one or more guide polynucleotides target the base editor to effect a change in a mutation associated with Schwann-Dieter syndrome; and (b) differentiating the induced pluripotent stem cells or precursors into desired cell types.
10. A guide RNA comprising a nucleic acid sequence from 5′ to 3′ or a 5′ truncated fragment of 1, 2, 3, 4 or 5 nucleotides thereof, wherein the nucleic acid sequence is selected from one or more of the following: GUAAGCAGGCGGGUAACAGC, AGCAGGCGGGUAACAGCUGC, GCGGGUAACAGCUGCAGCAU, UGUAAAUGUUUCCUAAGGUC, AAUGUUUCCUAAGGUCAGGU, GCAGGCGGGUAACAGCUGC, CAGGCGGGUAACAGCUGC, AGGCGGGUAACAGCUGC or AAGCAGGCGGGUAACAGCUGC.
Citation Information
Patent Citations
Towel rings
CN3315821D
cell phone
CN3329834D
Split inteins, conjugates and uses thereof
US20150344549A1
Switchable cas9 nucleases and uses thereof
US20160208288A1
AAV delivery of nucleobase editors
US20180127780A1