Compositions and methods for treating hemeopathy
By using the improved adenosine deaminase base editor ABE8, SNPs associated with sickle cell disease were targeted and modified, solving the problem that existing technologies could not effectively correct SCD mutations, and achieving efficient genetic repair and disease improvement.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-02-13
- Publication Date
- 2026-04-07
AI Technical Summary
Current treatments cannot effectively edit and correct the genetic mutations associated with sickle cell disease (SCD), which lead to intermittent microvascular occlusion and tissue ischemia/reperfusion injury, affecting life expectancy and quality of life.
Using an improved adenosine deaminase base editor (ABE8), by binding to guide RNA, we target and modify sickle cell disease-associated single nucleotide polymorphisms (SNPs) to convert A·T mutations to G·C, thereby correcting harmful mutations in the HBB gene, such as changing valine to alanine at amino acid position 6.
This technology enables efficient editing of SCD mutations, improving treatment outcomes, extending patients' life expectancy, and reducing the frequency and severity of disease symptoms.
Smart Images

Figure CN120005859B_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese Patent Application No. 202080028554.6, filed on February 13, 2020, entitled “Compositions and Methods of Treating Hemeopathies.”
[0002] Cross Reference to Related Applications
[0003] This application is an international PCT application which claims priority to and the benefit of U.S. Provisional Application Nos. 62 / 805,271, filed February 13, 2019; 62 / 805,277, filed February 13, 2019; 62 / 852,224, filed May 23, 2019; 62 / 852,228, filed May 23, 2019; 62 / 931,722, filed November 6, 2019; 62 / 931,747, filed November 6, 2019; 62 / 941,569, filed November 27, 2019; and 62 / 966,526, filed January 27, 2020, the contents of which are incorporated by reference in their entirety.
[0004] CITED REFERENCES
[0005] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. The contents of all publications, patents, and patent applications mentioned in this specification are hereby incorporated by reference in their entirety. TECHNICAL FIELD
[0006] The present disclosure provides compositions and methods for editing and correcting deleterious mutations associated with sickle cell disease (SCD). In particular, the present disclosure provides a modified adenosine deaminase base editor with improved degree of base editing to correct SCD mutations compared to a reference adenosine deaminase base editor. BACKGROUND
[0007] Sickle cell disease (SCD) is a disorder that affects hemoglobin, the molecule in red blood cells that carries oxygen to cells throughout the body. People with this disorder have atypical hemoglobin molecules, which distort red blood cells into a sickle or crescent shape. Clinical manifestations of sickle cell disease (SCD) are due to intermittent episodes of microvascular occlusion, leading to ischemia / reperfusion injury and chronic hemolysis. Vaso-occlusive events are associated with ischemia / reperfusion injured tissues, causing pain and acute or chronic injury affecting any organ system. Frequently affected are the skeleton / bone marrow, spleen, liver, brain, lungs, kidneys, and joints.
[0008] SCD is a genetic disorder characterized by the presence of at least one heme S-pair gene (HbS; p.Glu6Val in HBB) and a second pathogenic variant of HBB, resulting in abnormal heme polymerization. HbS / S (homozygous p.Glu6Val in HBB) accounts for 60% to 70% of SCD cases in the United States. The life expectancy for men and women with SCD is only 42 and 48 years, respectively. Current treatments focus on managing the symptoms. There is a pressing need for a method to edit the genetic mutations that cause SCD and other heme disorders. Summary of the Invention
[0009] As described below, the present invention is characterized by compositions and methods for editing harmful mutations associated with sickle cell disease (SCD). In a particular embodiment, the present invention provides for correcting SCD mutations using a modified adenosine deaminase base editor called “ABE8” with an unprecedented level of efficacy (e.g., >60 to 70%).
[0010] In one embodiment, the invention is characterized by a method for editing a β-globin polynucleotide containing a sickle cell disease-associated single nucleotide polymorphism (SNP), the method comprising contacting the β-globin polynucleotide with one or more guide RNAs and a fusion protein comprising a polynucleotide programmable DNA-binding domain and at least one base editor domain, the base editor domain being an adenosine deaminase variant containing a modification of amino acid positions 82 and / or 166 in MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD, wherein the guide RNA targets the base editor to perform the modification of the sickle cell disease-related SNP.
[0011] In another embodiment, the invention is characterized by a method for editing a β-globin (HBB) polynucleotide containing a sickle cell disease-associated single nucleotide polymorphism (SNP), the method comprising contacting the β-globin polynucleotide with one or more guide RNAs and a fusion protein containing a polynucleotide-programmable DNA-binding domain and at least one base editor domain containing an adenosine deaminase variant, the polynucleotide-programmable DNA-binding domain comprising the following sequence:
[0012]
[0013] The bolded sequences indicate sequences derived from Cas9, the italicized sequences represent connecting subsequences, and the underlined sequences represent binary kernel localization sequences.
[0014] The adenosine deaminase variant contains modifications at amino acid positions 82 and / or 166 in MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD.
[0015] In another embodiment, the present invention features a base editing system comprising a fusion protein of any of the foregoing embodiments or other embodiments described herein, and a guide RNA comprising a nucleic acid sequence selected from CUUCUCCACAGGAGUCAGAU; ACUUCUCCACAGGAGUCAGAU; and GACUUCUCCACAGGAGUCAGAU. In one embodiment, the gRNA further comprises the nucleic acid sequence GUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG. In another embodiment, the gRNA comprises a nucleic acid sequence selected from...
[0016] CUUCUCCACAGGAGUCAGAUGUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG;
[0017] ACUUCUCCACAGGAGUCAGAUGUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG;and
[0018] The nucleic acid sequence of GACUUCUCCACAGGAGUCAGAUGUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG.
[0019] In another embodiment, the invention is characterized by a cell prepared by introducing the following into the cell or its progenitor cell: a base editor, a polynucleotide encoding the base editor, wherein the base editor comprises a polynucleotide programmable DNA-binding domain and an adenosine deaminase domain as described in any of the embodiments herein; and one or more guide polynucleotides that target the base editor to perform A·T to G·C modifications of sickle cell disease-related SNPs. In one embodiment, the prepared cell is a hematopoietic stem cell, hematopoietic progenitor cell, proerythrocyte, erythroblast, reticulocyte, or erythrocyte. In another embodiment, the cell or its progenitor cell is a hematopoietic stem cell, hematopoietic progenitor cell, proerythrocyte, or erythroblast. In another embodiment, the hematopoietic stem cell is CD34. + Cells. In another embodiment, the cells are derived from an individual suffering from sickle cell disease. In yet another embodiment, the cells are mammalian cells or human cells.
[0020] In another embodiment, the invention is characterized by a method of treating sickle cell disease in an individual, comprising administering to the individual cells according to any of the prior embodiments or any other embodiments of the invention detailed herein. In one embodiment, the cells are the individual's own cells. In another embodiment, the cells are the individual's allogeneic cells.
[0021] In another embodiment, the present invention provides a single cell or a cell population that has been propagated or expanded from said cells in any of the foregoing embodiments or any other embodiments of the present invention detailed herein.
[0022] In another embodiment, the present invention provides a method for manufacturing erythrocytes or their progenitor cells, comprising introducing a base editor or a polynucleotide encoding a base editor into erythrocyte progenitor cells containing sickle cell disease-associated SNPs, wherein the base editor comprises a polynucleotide programmable nucleotide-binding domain and an adenosine deaminase variant domain as described in any of the foregoing embodiments; and one or more guide polynucleotides wherein the one or more guide polynucleotides target the base editor to perform A·T to G·C modifications of the sickle cell disease-associated SNPs; and differentiation of the erythrocyte progenitor cells into erythrocytes. In one embodiment, the method involves differentiating the erythrocyte progenitor cells into one or more hematopoietic stem cells, hematopoietic progenitor cells, proerythrocytes, erythroblasts, reticulocytes, or erythrocytes. In one embodiment, the erythrocyte progenitor cells involved in the method are CD34 cells. + Cells. In another embodiment, the erythrocyte progenitor cells are derived from an individual suffering from sickle cell disease. In another embodiment, the erythrocyte progenitor cells are mammalian cells or human cells. In another embodiment, the A.T to G.C modifications of the sickle cell disease-related SNP change valine to alanine in the HBB polypeptide. In another embodiment, the sickle cell disease-related SNP causes expression of the HBB polypeptide having valine at amino acid position 6. In another embodiment, the sickle cell disease-related SNP replaces glutamic acid with valine. In another embodiment, cells are selected to perform the A.T to G.C modifications of the sickle cell disease-related SNP. In another embodiment, the polynucleotide programmable DNA-binding domain comprises modified Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1Cas9 (St1Cas9), modified Streptococcus pyogenes Cas9 (SpCas9), or variants thereof.
[0023] In various embodiments of the present invention described herein, including any of the above-described embodiments or any other embodiments, the adenosine deaminase variant includes modifications at amino acid positions 82 and 166. In various embodiments of the present invention described herein, including any of the above-described embodiments or any other embodiments, the adenosine deaminase variant includes a V82S modification. In various embodiments of the present invention described herein, including any of the above-described embodiments or any other embodiments, the adenosine deaminase variant includes a T166R modification. In various embodiments of the present invention described herein, including any of the above-described embodiments or any other embodiments, the adenosine deaminase variant includes modifications of both V82S and T166R. In various embodiments of the present invention described herein, the adenosine deaminase variant further includes one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, and Q154R. In various embodiments of any or any other embodiment of the present invention described herein, the adenosine deaminase variant comprises a subset selected from Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y12 3H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; or a modified combination of I76Y+V82S+Y123H+Y147R+Q154R. In embodiments described above, the adenosine deaminase variant comprises Y147R+Q154R+Y123H. In embodiments described above, the adenosine deaminase variant comprises Y147R+Q154R+I76Y. In the embodiments described above, the adenosine deaminase variant comprises Y147R+Q154R+T166R. In the embodiments described above, the adenosine deaminase variant comprises Y147T+Q154R. In the embodiments described above, the adenosine deaminase variant comprises Y147T+Q154S. In the embodiments described above, the adenosine deaminase variant comprises Y147R+Q154S. In the embodiments described above, the adenosine deaminase variant comprises V82S+Q154S. In the embodiments described above, the adenosine deaminase variant comprises V82S+Y147R. In the embodiments described above, the adenosine deaminase variant comprises V82S+Q154R. In the embodiments described above, the adenosine deaminase variant comprises V82S+Y123H.In the embodiments described above, the adenosine deaminase variant comprises I76Y+V82S. In the embodiments described above, the adenosine deaminase variant comprises V82S+Y123H+Y147T. In the embodiments described above, the adenosine deaminase variant comprises V82S+Y123H+Y147R. In the embodiments described above, the adenosine deaminase variant comprises V82S+Y123H+Q154R. In the embodiments described above, the adenosine deaminase variant comprises Y123H+Y147R+Q154R+I76Y. In the embodiments described above, the adenosine deaminase variant comprises V82S+Y123H+Y147R+Q154R. In the embodiments described above, the adenosine deaminase variant comprises I76Y+V82S+Y123H+Y147R+Q154R. In other embodiments of the above examples, the adenosine deaminase variant comprises a C-terminal deletion starting from residues selected from the group consisting of 149, 150, 151, 152, 153, 154, 155, 156, and 157.
[0024] In various embodiments of the invention described herein, whether in vivo or in vitro, the cells are used. In various embodiments of the invention described herein, whether in vivo or in vitro, the A·T to G·C modification of the sickle cell disease-related SNP alters the valine content of the HBB polypeptide to alanine. In various embodiments of the invention described herein, whether in vivo or in vitro, the sickle cell disease-related SNP causes the expression of the HBB polypeptide having a valine content at amino acid position 6. In various embodiments of the invention described herein, whether in vivo or in vitro, the sickle cell disease-related SNP replaces glutamate with valine. In various embodiments of the invention described herein, whether in vivo or in vitro, the A·T to G·C modification of the sickle cell disease-related SNP causes the expression of the HBB polypeptide having alanine content at amino acid position 6. In various embodiments of the invention described herein, whether in vivo or in vitro, the A·T to G·C modification of the sickle cell disease-related SNP replaces glutamate with alanine.
[0025] In various embodiments of the present invention described herein, the polynucleotide programmable DNA binding domain is a modified Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1Cas9 (St1Cas9), a modified Streptococcus pyogenes Cas9 (SpCas9), or a variant thereof. In various embodiments of the present invention described herein, the polynucleotide programmable DNA binding domain comprises a variant of SpCas9 having modified protospacer adjacent motif (PAM) specificity or specificity for non-G PAMs. In various embodiments of the present invention described herein, the modified PAM has specificity for the nucleic acid sequence 5'-NGC-3'. In various embodiments of the present invention described herein, the modified SpCas9 comprises amino acid substitutions of D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R, or their corresponding amino acid substitutions. In various embodiments of the present invention described herein, the polynucleotide programmable DNA-binding domain is a nuclease-inactivating or nicking enzyme variant. In various embodiments of the present invention described herein, the nicking enzyme variant comprises amino acid substitution of D10A, or its corresponding amino acid substitution. In various embodiments of the present invention described herein, the base editor further comprises a zinc finger domain. In various embodiments of the present invention described herein, the zinc finger domain comprises recognizing helical sequences RNEHLEV, QSTTLKR, and RTEHLAR, or recognizing helical sequences RGEHLRQ, QSGTLKR, and RNDKLVP. In various embodiments of the present invention described herein, in any of the above-described embodiments or any other embodiments, the zinc finger domain is one or more zf1ra or zf1rb. In various embodiments of the present invention described herein, the adenosine deaminase domain can deaminate adenine on deoxyribonucleic acid (DNA).In various embodiments of the invention described herein, including any of the above-described embodiments or any other embodiments, the one or more guide RNAs comprise CRISPR RNA (crRNA) and trans-encoded small RNA (tracrRNA), wherein the crRNA comprises a nucleic acid sequence complementary to the HBB nucleic acid sequence containing the sickle cell disease-associated SNP. In various embodiments of the invention described herein, including any of the above-described embodiments or any other embodiments, the base editor is complexed with a single guide RNA (sgRNA), the nucleic acid sequence of which is complementary to the HBB nucleic acid sequence containing the sickle cell disease-associated SNP. In various embodiments of the invention described herein, A·T to G·C modifications of the sickle cell disease-associated SNP change valine in the HBB polypeptide to alanine. In another embodiment, the sickle cell disease-associated SNP causes expression of the HBB polypeptide having valine at amino acid position 6. In another embodiment, the sickle cell disease-associated SNP replaces glutamate with valine. In another embodiment, the A·T to G·C modification of the sickle cell disease-related SNP results in the expression of an HBB polypeptide having an alanine residue at amino acid position 6. In another embodiment, the A·T to G·C modification of the sickle cell disease-related SNP substitutes glutamic acid for alanine. In another embodiment, cells are selected to perform the A·T to G·C modification of the sickle cell disease-related SNP. In another embodiment, the polynucleotide programmable DNA-binding domain is a modified Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1Cas9 (St1Cas9), a modified Streptococcus pyogenes Cas9 (SpCas9), or a variant thereof.
[0026] In one embodiment, a method of treating an individual with sickle cell disease (SCD) is provided, wherein the method includes administering to the individual a fusion protein comprising an adenosine deaminase variant inserted into a Cas9 or Cas12 polypeptide, or a polynucleotide encoding the fusion protein thereon; and one or more guide polynucleotides that target the fusion protein to perform an A·T to G·C modification of an SCD-related single nucleotide polymorphism (SNP), thereby treating the individual's SCD.
[0027] In another embodiment, a method for treating an individual with sickle cell disease (SCD) is provided, wherein the method includes administering to the individual adenosine base editor 8 (ABE8) or a polynucleotide encoding the base editor, wherein ABE8 comprises an adenosine deaminase variant inserted into a Cas9 or Cas12 polypeptide; and one or more guide polynucleotides that target ABE8 to perform A·T to G·C modifications of SCD-related SNPs, thereby treating the individual's SCD.
[0028] In one embodiment of the method described above, ABE8 is selected from: ABE8.1-m, ABE8.2-m, ABE8.3-m, ABE8.4-m, ABE8.5-m, ABE8.6-m, ABE8.7-m, ABE8.8-m, ABE8.9-m, ABE8.10-m, ABE8.11-m, ABE8.12-m, ABE8.13-m, ABE8.14-m, ABE8.15-m, ABE8.16-m, ABE8.17-m, ABE8.18-m, ABE8.19-m, ABE8.20-m, ABE8.21-m, ABE8.22-m, ABE8.23- m, ABE8.24-m, ABE8.1-d, ABE8.2-d, ABE8.3-d, ABE8.4-d, ABE8.5-d, ABE8.6-d, ABE8.7-d, ABE8.8-d, ABE8.9-d, ABE8.10-d, ABE8.11-d, ABE8.12-d, ABE8.13-d, ABE8.14-d, ABE8.15-d, ABE8.16-d, ABE8.17-d, ABE8.18-d, ABE8.19-d, ABE8.20-d, ABE8.21-d, ABE8.22-d, ABE8.23-d, or ABE8.24-d. In one embodiment of the method described above, the adenosine deaminase variant comprises the following amino acid sequence:
[0029] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD, and wherein the amino acid sequence contains at least one modification. In one embodiment, the adenosine deaminase variant contains a modification at amino acid positions 82 and / or 166. In one embodiment, at least one modification comprises: V82S, T166R, Y147T, Y147R, Q154S, Y123H, and / or Q154R.
[0030] In one embodiment of the method described above, the adenosine deaminase variant comprises one of the following modified combinations: Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H +Y147R;V82S+Y123H+Q154R;Y147R+Q154R+Y123H;Y147R+Q154R+I76Y;Y147R+Q154R+T166R;Y123H+Y147R+Q154R+I76Y;V82S+Y123H+Y147R+Q154R;andI76Y+V82S+Y123H+Y147R+Q154R. In one embodiment of the method described above, the adenosine deaminase variant is TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23, or TadA*8.24. In one embodiment, the adenosine deaminase variant comprises a C-terminal deletion starting from residues selected from the group consisting of 149, 150, 151, 152, 153, 154, 155, 156, and 157. In one embodiment, the adenosine deaminase variant is an adenosine deaminase monomer comprising a TadA*8 adenosine deaminase variant domain. In one embodiment, the adenosine deaminase variant is an adenosine deaminase heterodimer comprising a wild-type adenosine deaminase domain and a TadA*8 adenosine deaminase variant domain. In one embodiment, the adenosine deaminase variant is an adenosine deaminase heterodimer comprising a TadA domain and a TadA*8 adenosine deaminase variant domain.
[0031] In one embodiment of the method described above, the SCD-related SNP is located in the β-globin (HBB) gene. In one embodiment of the method described above, the SNP causes expression of an HBB polypeptide containing valine at amino acid position 6. In one embodiment of the method described above, the SNP replaces glutamic acid with valine. In one embodiment of the method described above, modifications A.T to G.C. of the SNP change valine to alanine in the HBB polypeptide. In one embodiment of the method described above, modifications A.T to G.C. of the SNP cause expression of an HBB polypeptide containing alanine at amino acid position 6. In one embodiment of the method described above, modifications A.T to G.C. of the SNP replace glutamic acid with alanine.
[0032] In one embodiment of the method described above, the adenosine deaminase variant is inserted into a flexible loop, α-helix region, unstructured portion, or solvent-accessible portion of the Cas9 or Cas12 polypeptide. In one embodiment of the method described above, the adenosine deaminase variant is side-attached to both the N-terminal and C-terminal fragments of the Cas9 or Cas12 polypeptide. In one embodiment of the method described above, the fusion protein or ABE8 comprises the structural formula NH2-[N-terminal fragment of Cas9 or Cas12 polypeptide]-[adenosine deaminase variant]-[C-terminal fragment of Cas9 or Cas12 polypeptide]-COOH, wherein the “]-[” in each example are linkers as desired. In one embodiment, the C-terminus of the N-terminal fragment or the N-terminus of the C-terminal fragment comprises a portion of a flexible loop of the Cas9 or Cas12 polypeptide. In one embodiment, when the adenosine deaminase variant deaminates a target nucleobase, the flexible loop comprises an amino acid adjacent to the target nucleobase.
[0033] In one embodiment of the method described above, the method further includes administering a guide nucleic acid sequence to an individual to deaminate the target nucleobases of SCD-associated SNPs. In one embodiment, the deamination of the SNP target nucleobases results in the replacement of the target nucleobases with non-wild-type nucleobases, and the deamination of the target nucleobases alleviates the symptoms of sickle cell disease. In one embodiment, the deamination of sickle cell disease-associated SNPs results in alanine replacing glutamic acid.
[0034] In one embodiment of the method described above, the target nucleobase is 1-20 nucleotides from the PAM sequence in the target polynucleotide sequence. In one embodiment, the target nucleobase is 2-12 nucleotides upstream of the PAM sequence. In one embodiment of the method described above, the N-terminal or C-terminal fragment of the Cas9 or Cas12 polypeptide is the binding target polynucleotide sequence. In some embodiments, the N-terminal or C-terminal fragment contains a RuvC domain; the N-terminal or C-terminal fragment contains an HNH domain; neither the N-terminal nor the C-terminal fragment contains an HNH domain; or neither the N-terminal nor the C-terminal fragment contains a RuvC domain. In one embodiment, the Cas9 or Cas12 polypeptide comprises a partial or complete deletion of one or more structured domains, wherein the deaminase is intercalated at the partial or complete deletion site of the Cas9 or Cas12 polypeptide. In some embodiments, the deletion is within the RuvC domain; the deletion is within the HNH domain; or the deletion is a bridge between the RuvC domain and the C-terminal domain.
[0035] In one embodiment of the method described above, the fusion protein or ABE8 comprises a Cas9 polypeptide. In one embodiment, the Cas9 polypeptide is Streptococcus pyogenes Cas9 (SpCas9), Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1Cas9 (St1Cas9), or a variant thereof. In one embodiment, the Cas9 polypeptide comprises the following amino acid sequence (Cas9 reference sequence):
[0036]
[0037]
[0038] (Single underline: HNH domain; Double underline: RuvC domain; (Cas9 reference sequence), or its corresponding region. In some embodiments, the Cas9 peptide contains deleted amino acids 1017-1069 (as numbered in the Cas9 peptide reference sequence), or their corresponding amino acids; the Cas9 peptide contains deleted amino acids 792-872 (as numbered in the Cas9 peptide reference sequence), or their corresponding amino acids; or the Cas9 peptide contains deleted amino acids 792-906 (as numbered in the Cas9 peptide reference sequence).) Or their corresponding amino acids. In one embodiment of the method described above, the adenosine deaminase variant is inserted into a flexible loop of the Cas9 polypeptide. In one embodiment, the flexible loop comprises a region consisting of a group of amino acids selected from positions 530-537, 569-579, 686-691, 768-793, 943-947, 1002-1040, 1052-1077, 1232-1248, and 1298-1300 (which are numbered according to the Cas9 reference sequence), or their corresponding amino acid positions.
[0039] In one embodiment of the method described above, the deaminase variant is inserted between the following amino acid positions: 768-769, 791-792, 792-793, 1015-1016, 1022-1023, 1026-1027, 1029-1030, 1040-1041, 1052-1053, 1054-1055, 1067-1068, 1068-1069, 1247-1248, or 1248-1249 (which are numbered according to the Cas9 reference sequence), or their corresponding amino acid positions. In one embodiment of the method described above, the deaminase variant is inserted between the following amino acid positions: 768-769, 792-793, 1022-1023, 1026-1027, 1040-1041, 1068-1069, or 1247-1248 (which are numbered according to the Cas9 reference sequence), or their corresponding amino acid positions. In one embodiment of the method described above, the deaminase variant is inserted between the following amino acid positions: 1016-1017, 1023-1024, 1029-1030, 1040-1041, 1069-1070, or 1247-1248 (which are numbered according to the Cas9 reference sequence), or their corresponding amino acid positions. In one embodiment of the method described above, the adenosine deaminase variant is inserted into the locus within the Cas9 polypeptide indicated in Table 14A. In one embodiment, the N-terminal fragment comprises amino acid residues 1-529, 538-568, 580-685, 692-942, 948-1001, 1026-1051, 1078-1231, and / or 1248-1297 of the Cas9 reference sequence, or their corresponding residues. In one embodiment, the C-terminal fragment comprises amino acid residues 1301-1368, 1248-1297, 1078-1231, 1026-1051, 948-1001, 692-942, 580-685, and / or 538-568 of the Cas9 reference sequence, or their corresponding residues.
[0040] In one embodiment of the method described above, the Cas9 peptide is a modified Cas9 and is specific to altered PAM or non-G PAM. In one embodiment of the method described above, the Cas9 peptide is a nicking enzyme or wherein the Cas9 peptide is a nuclease inactivated. In one embodiment of the method described above, the Cas9 peptide is a modified SpCas9 peptide. In one embodiment, the modified SpCas9 peptide includes amino acid substitutions of D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (SpCas9-MQKFRAER) and is specific to altered PAM 5'-NGC-3'.
[0041] In another embodiment of the method described above, the fusion protein or ABE8 comprises a Cas12 polypeptide. In one embodiment, the adenosine deaminase variant is inserted into the Cas12 polypeptide. In one embodiment, the Cas12 polypeptide is Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, or Cas12i. In one embodiment, the adenosine deaminase variant is inserted between the following amino acid positions: a) between corresponding amino acid residues of BhCas12b: 153-154, 255-256, 306-307, 980-981, 1019-1020, 534-535, 604-605, or 344-345, or Cas12a, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, or Cas12i; b) between BvCas12b: 147 and 148, 248 and 249, 299 and 30. Between corresponding amino acid residues of Cas12a, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, or Cas12i; or c) between corresponding amino acid residues of AaCas12b, such as 157 and 158, 258 and 259, 310 and 311, 1008 and 1009, or 1044 and 1045, or between corresponding amino acid residues of Cas12a, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, or Cas12i. In one embodiment, the adenosine deaminase variant is inserted into the Cas12 polypeptide indicated in Table 14B. In one embodiment, the Cas12 polypeptide is Cas12b. In one embodiment, the Cas12 polypeptide comprises a BhCas12b domain, a BvCas12b domain, or an AACas12b domain.
[0042] In one embodiment of the method described above, the guide RNA comprises CRISPR RNA (crRNA) and trans-activating crRNA (tracrRNA). In one embodiment of the method described above, the individual is a mammal or a human.
[0043] In another embodiment, a pharmaceutical composition is provided comprising a base editing system containing any of the fusion proteins described in the methods, examples, and embodiments above, and a pharmaceutically acceptable carrier, solvent, or excipient. In one embodiment, the pharmaceutical composition further comprises a guide RNA comprising a sequence selected from CUUCUCCACAGGAGUCAGAU; ACUUCUCCACAGGAGUCAGAU; and GACUUCUCCACAGGAGUCAGAU. In one embodiment, the gRNA further comprises a nucleic acid sequence belonging to the group consisting of the nucleic acid sequences GUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG. In one embodiment, the gRNA comprises a compound selected from CUUCUCCACAGGAGUCAGAUGUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUAUGGCAGGGUG.
[0044] ACUUCUCCACAGGAGUCAGAUGUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG;and
[0045] The nucleic acid sequence of GACUUCUCCACAGGAGUCAGAUGUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG.
[0046] In one embodiment, a pharmaceutical composition is provided comprising a base editor or a polynucleotide encoding a base editor, wherein the base editor comprises any polynucleotide programmable DNA-binding domain and adenosine deaminase domain as described in the methods, examples, and embodiments above; and one or more guide polynucleotides that target the base editor to perform A·T to G·C modifications of sickle cell disease-related SNPs; and a pharmaceutically acceptable carrier, solvent, or excipient.
[0047] In another embodiment, a pharmaceutical composition is provided comprising cells as described in the embodiments and implementation schemes above, and a pharmaceutically acceptable carrier, solvent, or excipient.
[0048] In another embodiment, a kit is provided comprising a base editing system containing any of the fusion proteins described in the methods, embodiments, and implementation schemes above. In one embodiment, the kit further comprises a guide RNA comprising a nucleic acid sequence selected from the group consisting of CUUCUCCACAGGAGUCAGAU; ACUUCUCCACAGGAGUCAGAU; and GACUUCUCCACAGGAGUCAGAU.
[0049] In another embodiment, a kit is provided comprising a base editor or a polynucleotide encoding a base editor, wherein the base editor comprises any polynucleotide programmable DNA-binding domain and adenosine deaminase domain as described in the methods, embodiments and implementations above; and one or more guide polynucleotides that target the base editor to perform A·T to G·C modifications of sickle cell disease-related SNPs.
[0050] In another embodiment, a kit is provided comprising any of the cells described in the embodiments and implementations above. In one embodiment of the kit, the kit further includes a packaging insert with an instruction manual.
[0051] In one embodiment, this document provides a base editor system comprising a polynucleotide programmable DNA-binding domain and at least one base editor domain comprising an adenosine deaminase variant containing a modification at amino acid position 82 or 166 of the following sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD; and a guide RNA wherein the guide RNA targets the base editor to perform modification of α-1 antitrypsin deficiency-related SNPs. In some embodiments, the adenosine deaminase variant comprises a V82S modification and / or a T166R modification. In some embodiments, the adenosine deaminase variant further comprises one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, and Q154R. In some embodiments, the base editor domain comprises an adenosine deaminase heterodimer comprising a wild-type adenosine deaminase domain and an adenosine deaminase variant. In some embodiments, the adenosine deaminase variant is a truncated TadA8, which loses 1, 2, 3, 4, 5, 66, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to the full-length TadA8. In some embodiments, the adenosine deaminase variant is a truncated TadA8, which has lost 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to the full-length TadA8. In some embodiments, the polynucleotide-programmable DNA-binding domain is a modified Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1Cas9 (St1Cas9), a modified Streptococcus pyogenes Cas9 (SpCas9), or a variant thereof. In some embodiments, the polynucleotide-programmable DNA-binding domain is an SpCas9 variant with modified protospacer adjacent motif (PAM) specificity or specificity to non-G PAMs. In some embodiments, the polynucleotide-programmable DNA-binding domain is a nuclease-inactivating Cas9. In some implementations, the polynucleotide programmable DNA binding domain is a Cas9 nickase.
[0052] In one embodiment, this document provides a base editor system comprising one or more guide RNAs and a fusion protein, the latter comprising a multinucleotide programmable DNA-binding domain containing the following sequences:
[0053]
[0054] In one embodiment, a cell is provided comprising any of the base editor systems described above. In some embodiments, the cell is a human cell or a mammalian cell. In some embodiments, the cell is in vitro, in vivo, or in vitro.
[0055] The descriptions and examples herein are illustrative of embodiments of the invention. It is generally understood that the invention is not limited to the specific embodiments described herein and may vary. Those skilled in the art will understand that many variations and modifications are applicable within the scope of this invention.
[0056] This invention provides a composition and method for editing and treating mutations associated with sickle cell disease (SCD). The isolation or manufacture of the compositions and articles as defined herein is related to the examples provided below. Other features and advantages of this invention will be apparent from the detailed description and claims. Unless otherwise stated, the operation of some embodiments disclosed herein employs known immunological, biochemical, chemical, molecular biological, microbiological, cell biological, genomic, and recombinant DNA techniques, etc., which are within the scope of the art. See, for example: Sambrook and Green, 4th edition, Molecular Cloning: A Laboratory Manual (2012); a series of Current Protocols in Molecular Biology (edited by F.M.A. Susubel et al.); a series of Methods in Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (edited by M.J. MacPherson, B.D. Hames, and G.G. Taylor (1995)); Harlow and Lane (edited by 1988), Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th edition (edited by R.R. Freshney (2010)).
[0057] The chapter titles provided in this article are for organizational purposes only and are not intended to limit the topics discussed.
[0058] While various features of the invention may be described in a single embodiment, such features may also be provided separately or in any suitable combination. Conversely, although the invention is described in separate embodiments herein for illustrative purposes, the invention may also be practiced in a single embodiment. Section headings used herein are for organizational purposes only and are not intended to limit the scope of the invention.
[0059] The features of this invention are specifically described in the appended claims. The features and advantages of this invention will be further understood by referring to the embodiments employing the principles of the invention as illustrated in the following detailed description and by taking into account the accompanying drawings.
[0060] definition
[0061] The following definitions are supplementary to those in the relevant art and are relevant to the present application, and are not intended to implicate any related or unrelated cases, such as any jointly owned patents or applications. While any methods and materials similar to or equivalent to those described herein may be used to experiment with the invention, preferred materials and methods will be described herein. Therefore, the terminology used herein is for illustrative purposes only and is not intended to be limiting.
[0062] Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by one of ordinary skill in the art to which this invention pertains. The following references provide general definitions for many of the terms used in this invention that are understood by one of ordinary skill in the art: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd edition, 1994); The Cambridge Dictionary of Science and Technology (Walker, ed., 1988); The Glossary of Genetics, 5th edition, R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991).
[0063] In this application, the singular includes the plural unless otherwise expressly stated. It should be noted that the singular forms “a,” “an,” and “the” used in this specification include the plural related items unless otherwise clearly stated. In this application, the word “or” means “and / or” unless otherwise stated, and is generally understood to be inclusive. Furthermore, the use of the term “including” and other forms such as “include,” “includes,” and “included” is not limited.
[0064] The terms “comprising” (and any comprising type, such as “comprise” and “comprises”), “having” (and any having type, such as “have” and “has”), “including” (and any including type, such as “includes” and “include”), or “containing” (and any containing type, such as “contains” and “contain”) as used in this specification and claims are inclusive or open-ended and do not exclude additional, unextracted elements or method steps. Any embodiment discussed in this specification is contemplated to be carried out using any method or composition of the present invention, and vice versa. Furthermore, the compositions of the present invention can be used to achieve the methods of the present invention.
[0065] The terms "about" or "approximately" mean that a particular value determined by a person skilled in the art is within an acceptable margin of error, with a portion depending on the method of measurement or determination, i.e., limited by the measurement system. For example, depending on the operation of the relevant art, "about" may mean within a standard deviation of 1 or more. Alternatively, "about" may mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of the specified value. Or, particularly for biological systems or processes, the term may mean within an order of magnitude, such as within 5 times or 2 times the specified value. When a particular value is specified in an application or claim, unless otherwise stated, it should be assumed that the term "about" means within an acceptable margin of error for that particular value.
[0066] The ranges provided in this article are generally understood to be shorthand notations of all values within the range. For example, the range 1 to 50 is generally understood to include any number, combination of numbers, or subrange of numbers in the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.
[0067] The terms “some embodiments,” “an embodiment,” “an embodiment,” or “other embodiments” used in this specification refer to specific features, structures, or characteristics described in relation to an embodiment that are included in at least some embodiments of the invention, but not necessarily in all embodiments of the invention.
[0068] "Adenosine deaminase" refers to a polypeptide or fragment thereof that can catalyze the hydrolytic deamination of adenine or adenosine. In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine to form inosine or from deoxyadenosine to deoxyinosine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., engineered adenosine deaminases, evolved adenosine deaminases) may be derived from any organism, such as bacteria. In some embodiments, the adenosine deaminase contains modifications to the following sequences:
[0069] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD
[0070] (Also known as TadA*7.10).
[0071] In some embodiments, TadA*7.10 includes at least one modification. In some embodiments, TadA*7.10 includes modifications to amino acids 82 and / or 166. In specific embodiments, variants of the above-mentioned sequence include one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. In other embodiments, variants of the TadA7.10 sequence include those selected from Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V Modified combinations of 82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R.
[0072] In other embodiments, the present invention provides a variant of adenosine deaminase comprising a deletion, such as TadA*8, which includes a C-terminal deletion beginning at residues 149, 150, 151, 152, 153, 154, 155, 156, or 157. In other embodiments, the adenosine deaminase variant is a TadA (e.g., TadA*8) monomer comprising one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. In other embodiments, the adenosine deaminase variant comprises a subset selected from Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+ The individual units of the modified combinations of Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R.
[0073] In other embodiments, the adenosine deaminase variant is a homodimer comprising two adenosine deaminase domains (e.g., TadA*8), each having one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. In other embodiments, the adenosine deaminase variant is a homodimer comprising two adenosine deaminase domains (e.g., TadA*8), each comprising a domain selected from Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y1 Modified combinations of the groups consisting of 23H+Y147R;V82S+Y123H+Q154R;Y147R+Q154R+Y123H;Y147R+Q154R+I76Y;Y147R+Q154R+T166R;Y123H+Y147R+Q154R+I76Y;V82S+Y123H+Y147R+Q154R; andI76Y+V82S+Y123H+Y147R+Q154R.
[0074] In other embodiments, the adenosine deaminase variant is a heterodimer comprising a wild-type TadA adenosine deaminase domain and an adenosine deaminase variant domain (e.g., TadA*8), comprising one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. In other embodiments, the adenosine deaminase variant is a heterodimer comprising a wild-type TadA adenosine deaminase domain and an adenosine deaminase variant domain (e.g., TadA*8), comprising a domain selected from Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y1 47T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and modified combinations of I76Y+V82S+Y123H+Y147R+Q154R.
[0075] In other embodiments, the adenosine deaminase variant is a heterodimer comprising a TadA*7.10 domain and an adenosine deaminase variant domain (e.g., TadA*8), which comprises one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. In other embodiments, the adenosine deaminase variant is a heterodimer comprising a TadA*7.10 domain and an adenosine deaminase variant domain (e.g., TadA*8), comprising the following modified combinations: Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123 H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q1 54R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; or I76Y+V82S+Y123H+Y147R+Q154R.
[0076] In one embodiment, the adenosine deaminase is TadA*8, which comprises or substantially consists of the following sequence or fragments thereof having adenosine deaminase activity:
[0077] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD.
[0078] In some embodiments, TadA*8 is truncated. In some embodiments, the truncated TadA*8 loses 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to the full-length TadA*8. In some embodiments, the truncated TadA*8 loses 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to the full-length TadA*8. In some embodiments, the adenosine deaminase variant is the full-length TadA*8.
[0079] In a particular embodiment, the adenosine deaminase heterodimer comprises a TadA*8 domain and one of the following adenosine deaminase domains:
[0080] Staphylococcus aureus (S. aureus) TadA:
[0081] MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN
[0082] Bacillus subtilis (B. subtilis) TadA:
[0083] MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLSE
[0084] Salmonella typhimurium (S. typhimurium) TadA:
[0085] MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFRMRRQEIKALKKADRAEGAGPAV
[0086] Shewanella putrefaciens (S. putrefaciens) TadA:
[0087] MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE
[0088] Haemophilus influenzae F3031 (H. influenzae) TadA:
[0089] MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTΑΗAEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRREEKKIEKALLKSLSDK
[0090] Caulobacter crescentus (C. crescentus) TadA:
[0091] MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLTDLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAKI
[0092] Geobacter sulfurreducens (G. sulfurreducens) TadA:
[0093] MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRKKAKATPALFIDERKVPPEP
[0094] TadA*7.10
[0095] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD
[0096] "Adenosine deaminase base editor 8 (ABE8) polypeptide" means a base editor (BE) as defined and / or described herein, which contains an adenosine deaminase variant containing modifications at amino acid positions 82 and / or 166 of the following reference sequence:
[0097] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD. In some implementations, ABE8 contains additional modifications relative to the reference sequence.
[0098] "Adenosine deaminase base editor 8 (ABE8) polynucleotide" refers to the polynucleotide (polynucleotide sequence) that encodes the ABE8 polypeptide.
[0099] "Administration" in this document means providing one or more of the compositions described herein to a patient or individual. Examples, but not limited to, include: intravenous (iv), subcutaneous (sc), intradermal (id), intraperitoneal (ip), or intramuscular (im) administration of injectable compositions. One or more of these routes may be used. Non-enteral administration methods may include, for example, rapid bolus injection or prolonged intravenous infusion. Oral administration may also be used.
[0100] "Agent" refers to any small molecule chemical compound, antibody, nucleic acid molecule, or polypeptide, or fragment thereof.
[0101] "Alteration" refers to changing (e.g., increasing or decreasing) the structure, expression level, or activity of a gene or polypeptide, detected using methods known by standard technique, such as those described herein. The modifications used herein include altering the sequence of polynucleotides or polypeptides or changing the expression level, such as by 25%, 40%, 50%, or more.
[0102] "Alteration" means to reduce, suppress, weaken, eliminate, halt, or stabilize the development or evolution of a disease.
[0103] "Analog" refers to molecules that are different but have similar functions or structural features. For example, polynucleotide or polypeptide analogs retain the biological activity of the corresponding natural polynucleotide or polypeptide, but have certain modifications that enhance the function of the analog compared to naturally occurring polynucleotides or polypeptides. These modifications can improve the analog's affinity for DNA, efficacy, specificity, protease or nuclease resistance, membrane permeability, and / or half-life, without altering, for example, ligand binding. Analogs may include non-natural nucleotides or amino acids.
[0104] "Base editor (BE)" or "nucleobase editor (NBE)" refers to an agent that binds to polynucleotides and has nucleobase modification activity. In various embodiments, the base editor comprises a nucleobase-modifying polypeptide (e.g., a deaminase) and a nucleic acid-programmable nucleotide-binding domain linked to a guide polynucleotide (e.g., guide RNA). In various embodiments, the agent is a biomolecular complex comprising a protein domain with base-editing activity, i.e., a domain capable of modifying bases (e.g., A, T, C, G, or U) within a nucleic acid molecule (e.g., DNA). In some embodiments, the polynucleotide-programmable DNA-binding domain is fused to or linked to a deaminase domain. In one embodiment, the agent is a fusion protein comprising a domain with base-editing activity. In another embodiment, the protein domain with base-editing activity is linked to guide RNA (e.g., via an RNA-binding motif on the guide RNA and an RNA-binding domain fused to a deaminase). In some embodiments, the domain with base-editing activity can deaminate bases within a nucleic acid molecule. In some embodiments, the base editor can deaminate one or more bases within the DNA molecule. In some embodiments, the base editor can deaminate adenosine (A) within the DNA. In some embodiments, the base editor is an adenosine base editor (ABE).
[0105] In some implementations, the base editor is generated (e.g., ABE8) by selecting an adenosine deaminase variant (e.g., TadA*8) into a backbone comprising a circularly arranged Cas9 (e.g., spCAS9 or saCAS9) and a bipartite nucleus localization sequence. Circularly arranged Cas9 is known in the art and is described, for example, in Oakes et al., Cell 176, 254–267, 2019. In the following examples of circular arrangements, bold sequences indicate sequences derived from Cas9, italic sequences represent linker sequences, and underlined sequences represent bipartite nucleus localization sequences.
[0106] CP5 (possessing the MSP "NGC=Pam variant, with mutational patterns such as NGG" PID=protein interaction domain and "D10A" cleavage enzyme):
[0107]
[0108]
[0109] In some embodiments, ABE8 is a base editor selected from Tables 6 to 9, 13, or 14 below. In some embodiments, ABE8 comprises an adenosine deaminase variant evolved from TadA. In some embodiments, the adenosine deaminase variant of ABE8 is the TadA*8 variant described in Tables 7, 9, 13, or 14 below. In some embodiments, the adenosine deaminase variant is the TadA*7.10 variant (e.g., TadA*8) which comprises one or more modifications selected from the group consisting of Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. In various implementations, ABE8 includes the TadA*7.10 variant (e.g., TadA*8) and is selected from Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y Modified combinations of the group consisting of 147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R. In some embodiments, ABE8 is a monomeric construct. In some embodiments, ABE8 is a heterodimeric construct. In some embodiments, ABE8 comprises the sequence:
[0110] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD.
[0111] In some embodiments, the polynucleotide programmable DNA-binding domain is a CRISPR-related enzyme (e.g., Cas or Cpf1). In some embodiments, the base editor is a catalytically inactivated Cas9 (dCas9) fused with a deaminase domain. In some embodiments, the base editor is a Cas9 nickase (nCas9) fused with a deaminase domain. Detailed descriptions of the base editor can be found in international PCT applications PCT / 2017 / 045381 (WO 2018 / 027078) and PCT / US2016 / 058344 (WO 2017 / 070632), the contents of which are incorporated herein by reference in their entirety. See also Komor, AC et al. "Programmable base editing of a target base in genomic DNA without double-stranded DNA cleavage" Nature 533, 420-424 (2016); Gaudelli, NM et al. "Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage" Nature 551,464-471(2017); "Improved base excisionrepair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A baseeditors with higher efficiency and product purity" Science Advances 3:eaao4774(2017) by Komor, AC et al., and "Base editing: precision chemistry on the genome and transcriptome of living" by Rees, HA et al. cells." Nat Rev Genet.2018Dec;19(12):770-788.doi:10.1038 / s41576-018-0059-1, the full content of which has been incorporated herein by reference.
[0112] For example, the adenine base editor (ABE) used in the base editing compositions, systems, and methods described herein has the following nucleic acid sequence (8877 base pairs) provided below (Addgene, Watertown, MA.; Gaudelli NM et al. Nature. 2017 Nov 23; 551(7681):464-471. doi:10.1038 / nature24644; Koblan LW et al. Nat Biotechnol. 2018 Oct; 36(9):843-846. doi:10.1038 / nbt.4172.). It also includes polynucleotide sequences with at least 95% or higher homology to the ABE nucleic acid sequence.
[0113]
[0114] "Base editing activity" refers to the chemical alteration of bases within a polynucleotide. In one embodiment, a first base is converted to a second base. In another embodiment, base editing activity is adenosine or adenine deaminase activity, such as converting A·T to G·C. In some embodiments, base editing activity is analyzed using editing efficacy. Base editing efficacy can be measured using any suitable method, such as Sanger sequencing or next-generation sequencing. In some embodiments, base editing efficacy is measured as the percentage of total sequenced reads with nuclear base transitions performed by a base editor, such as the percentage of total sequenced reads with target AT base pairs converted to GC base pairs. In some embodiments, when base editing is performed in a cell population, base editing efficacy is measured as the percentage of total cells with nuclear base transitions performed by a base editor.
[0115] The term "base editor system" refers to a system for editing the nucleobases of a target nucleotide sequence. In various embodiments, the base editor system comprises (1) a polynucleotide-programmable nucleotide-binding domain (e.g., Cas9); (2) a deaminase domain (e.g., adenosine deaminase) for deaminating the nucleobases; and (3) one or more guide polynucleotides (e.g., guide RNA). In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable DNA-binding domain. In some embodiments, the base editor is an adenine or adenosine base editor (ABE). In some embodiments, the base editor system is ABE8.
[0116] In some embodiments, the base editor system may contain more than one base editing component. For example, the base editor system may include more than one deaminase. In some embodiments, the base editor system may include one or more adenosine deaminases. In some embodiments, a single guide polynucleotide may be used to target different deaminases to the target nucleic acid sequence. In some embodiments, a pair of guide polynucleotides may be used to target different deaminases to the target nucleic acid sequence.
[0117] The deaminase domain and polynucleotide programmable nucleotide-binding component of the base editor system can be linked to each other using any combination of covalent, non-covalent, or associative methods and their interactions. For example, in some embodiments, the deaminase domain can target a target nucleotide sequence using the polynucleotide programmable nucleotide-binding domain. In some embodiments, the polynucleotide programmable nucleotide-binding domain can be fused to or linked to the deaminase domain. In some embodiments, the polynucleotide programmable nucleotide-binding domain can target a target nucleotide sequence by non-covalent interaction with or associating with the deaminase domain. For example, in some embodiments, the deaminase domain may include additional heterologous portions or domains that can interact, associate, or form complexes with additional heterologous portions or domains that are part of the polynucleotide programmable nucleotide-binding domain. In some embodiments, the additional heterologous portions can bind to, interact with, associate with, or form complexes with peptides. In some embodiments, the additional heterologous portions can bind to, interact with, associate with, or form complexes with polynucleotides. In some embodiments, the additional heterologous portion may bind to a guide polynucleotide. In some embodiments, the additional heterologous portion may bind to a peptide linker. In some embodiments, the additional heterologous portion may bind to a polynucleotide linker. The additional heterologous portion may be a protein domain. In some embodiments, the additional heterologous portion may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMuCom coat protein domain, a sterile α motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif.
[0118] The base editor system may further include guide polynucleotide components. It should be understood that components of the base editor system may be linked to each other through covalent bonds, non-covalent interactions, or any combination of such linkages and interactions. In some embodiments, the deaminase domain can target a target nucleotide sequence via the guide polynucleotide. For example, in some embodiments, the deaminase domain may include additional heterologous portions or domains (e.g., polynucleotide-binding domains, such as RNA or DNA-binding proteins) that can interact, link, or form a complex with a portion or segment of the guide polynucleotide (e.g., a polynucleotide motif). In some embodiments, the additional heterologous portions or domains (e.g., polynucleotide-binding domains, such as RNA or DNA-binding proteins) can be fused to or linked to the deaminase domain. In some embodiments, the additional heterologous portions may bind to, interact with, link to, or form a complex with a polypeptide. In some embodiments, the additional heterologous portions may bind to, interact with, link to, or form a complex with a polynucleotide. In some embodiments, the additional heterologous portion may bind to a guide polynucleotide. In some embodiments, the additional heterologous portion may bind to a peptide linker. In some embodiments, the additional heterologous portion may bind to a polynucleotide linker. The additional heterologous portion may be a protein domain. In some embodiments, the additional heterologous portion may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMuCom coat protein domain, a sterile α motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif.
[0119] In some embodiments, the base editor system may further include an inhibitor of the base excision repair (BER) component. It should be understood that the components of the base editor system may be linked to each other through covalent bonds, non-covalent interactions, or any combination of such linkages and interactions. The BER component inhibitor may include a BER inhibitor. In some embodiments, the BER inhibitor may be a uracil DNA glycosidase inhibitor (UGI). In some embodiments, the BER inhibitor may be an inosine BER inhibitor. In some embodiments, the BER inhibitor may target the target nucleotide sequence via a polynucleotide-programmable nucleotide-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain may be fused to or linked to the BER inhibitor. In some embodiments, the polynucleotide-programmable nucleotide-binding domain may be fused to or linked to a deaminase domain and the BER inhibitor. In some embodiments, the polynucleotide-programmable nucleotide-binding domain may allow the BER inhibitor to target the target nucleotide sequence through non-covalent interactions or linkage with the BER inhibitor. For example, in some implementations, the inhibitor component of BER may include additional heterologous portions or domains that can interact with, associate with, or form a complex with an additional heterologous portion or domain that is part of a polynucleotide programmable nucleotide binding domain.
[0120] In some embodiments, the BER inhibitor can target the target nucleotide sequence via a guide polynucleotide. For example, in some embodiments, the BER inhibitor may include an additional heterologous portion or domain (e.g., a polynucleotide-binding domain, such as an RNA or DNA-binding protein) that can interact, link, or form a complex with a portion or segment (e.g., a polynucleotide motif) of the guide polynucleotide. In some embodiments, the additional heterologous portion or domain of the guide polynucleotide (e.g., a polynucleotide-binding domain, such as an RNA or DNA-binding protein) may be fused to or linked to the BER inhibitor. In some embodiments, the additional heterologous portion may be able to bind, interact with, link, or form a complex with the polynucleotide. In some embodiments, the additional heterologous portion may be able to bind to the guide polynucleotide. In some embodiments, the additional heterologous portion may be able to bind to a peptide linker. In some embodiments, the additional heterologous portion may be able to bind to a polynucleotide linker. The additional heterologous portion may be a protein domain. In some embodiments, the additional heterologous portion may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMuCom coat protein domain, a sterile α motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif.
[0121] "β-globin (HBB) protein" means a polypeptide or fragment thereof that has at least about 95% amino acid sequence identity with NCBI accession number NP_000509. In a particular embodiment, the β-globin protein includes one or more modifications relative to the following reference sequence. In one particular embodiment, the sickle cell disease-associated β-globin protein includes an E6V (also known as E7V) mutation. Examples of β-globin amino acid sequences are provided below.
[0122]
[0123] "HBB polynucleotide" refers to a nucleic acid molecule that encodes β-globin protein or a fragment thereof. Sequences of HBB polynucleotide examples can be obtained from NCBI accession number NM_000518, and are provided below:
[0124]
[0125]
[0126] The term "Cas9" or "Cas9 domain" refers to an RNA-guided nuclease containing the Cas9 protein or fragments thereof (e.g., proteins containing the active, inactive, or partially active DNA cleavage domain of Cas9, and / or the gRNA-binding domain of Cas9). Cas9 nucleases are sometimes also called Casn1 nucleases or CRISPR (clustered regularly interspaced short palindromic repeat)-related nucleases. CRISPR is an integral part of the acquired immune system, providing protection against mobile genetic elements (viruses, transposable elements, and conjugating plasmids). A CRISPR cluster contains a spacer, a sequence complementary to the aforementioned mobile elements, and a target invasive nucleic acid. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, the modification of pre-crRNA requires trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. tracrRNA acts as a guide for ribonuclease 3-assisted pre-crRNA processing. Subsequently, Cas9 / crRNA / tracrRNA cleaves the linear or circular dsDNA target complementary to the spacer via endonuclease cleavage. The target strand not complementary to the crRNA is first cleaved via endonuclease cleavage, followed by 3′–5′ trimming via exonuclease cleavage. In nature, DNA binding and cleavage typically require both proteins and two RNAs. However, a single guide RNA (“sgRNA” or simply “gNRA”) can be engineered to incorporate both crRNA and tracrRNA embodiments into a single RNA species. See, for example: Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816–821 (2012), the full contents of which are incorporated herein by reference. Cas9 identifies short motifs (PAM or protospacer adjacent motifs) in CRISPR repeat sequences to help distinguish between self and non-self motifs.The Cas9 nuclease sequence and structure are well known to those skilled in the art (see, for example: "Complete genome sequence of an M1 strain of Streptococcus pyogenes." Ferretti et al., JJ, McShan W.M., Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov A.N., Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chylinski K., Sharma C.M., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012), the full contents of which are incorporated herein by reference. Cas9 is a homologous gene that has been described in various species, including (but not limited to): Streptococcus pyogenes and Streptococcus thermophilus.Other suitable Cas9 nucleases and sequences will be known to those skilled in the art based on this invention, and such Cas9 nucleases and sequences include Cas9 sequences of organisms and loci revealed in “The tracrRNA and Cas9families of type IICRISPR-Cas immunitysystems” (2013) RNA Biology 10:5,726-737 by Chylinski, Rhun, and Charpentier; the full contents of which are incorporated herein by reference.
[0127] One example of Cas9 is Streptococcus pyogenes Cas9 (spCas9), whose amino acid sequence is provided below:
[0128]
[0129]
[0130] (Single bottom line: HNH domain; Double bottom line: RuvC domain)
[0131] Nuclease-inactivated Cas9 proteins can be exchanged for “dCas9” proteins (nuclease-“inactivated” Cas9) or catalytically inactivated Cas9. Methods for generating Cas9 proteins (or fragments thereof) with inactivated DNA cleavage domains are known (see, for example: Jinek et al., Science. 337:816-821 (2012); Qi et al., “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression” (2013) Cell. 28; 152(5):1173-83, the full contents of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to consist of two subdomains: the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains silence the nuclease activity of Cas9. For example, mutations in D10A and H840A completely inactivate the nuclease activity of Cas9 from *S. pyogenes* (Jinek et al., Science. 337:816-821 (2012); Qi et al., Cell. 28; 152(5):1173-83 (2013)). In some embodiments, the Cas9 nuclease has an inactivating (e.g., inactive) DNA cleavage domain, i.e., Cas9 is a nicking enzyme, referred to as the "nCas9" protein (in contrast to the "nicking enzyme" Cas9).
[0132] In some embodiments, a protein comprising a fragment of Cas9 is provided. For example, in some embodiments, the protein comprises one of two Cas9 domains: (1) a gRNA-binding domain of Cas9; or (2) a DNA-cleaving domain of Cas9. In some embodiments, a protein comprising Cas9 or a fragment thereof is referred to as a "Cas9 variant". Cas9 variants and Cas9 or a fragment thereof share common homology. For example, Cas9 variants are at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% homology with wild-type Cas9. In some implementations, the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid alterations compared to the wild-type Cas9. In some embodiments, the Cas9 variant contains a fragment of Cas9 (e.g., a gRNA-binding domain or a DNA-cleavage domain), and therefore the fragment is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid length of wild-type Cas9.
[0133] In some embodiments, the fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids in length. In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_017053.1, nucleotide and amino acid sequences are as follows):
[0134]
[0135]
[0136] (Single bottom line: HNH domain; Double bottom line: RuvC domain)
[0137] In some implementations, wild-type Cas9 corresponds to or contains the following nucleotide and / or amino acid sequences:
[0138]
[0139]
[0140]
[0141] (Single bottom line: HNH domain; Double bottom line: RuvC domain)
[0142] In some implementations, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_002737.2 (nucleotide sequence below); and Uniprot reference sequence: Q99ZW2 (amino acid sequence below):
[0143]
[0144]
[0145] (SEQ ID NO:1. Single bottom line: HNH domain; Double bottom line: RuvC domain).
[0146] In some implementations, Cas9 refers to Cas9 derived from the following: Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense, China (NCBI Ref: NC_021846.1); and Streptococcus iniae (NCBI Ref: NC_021846.1). Ref:NC_021314.1); Burkholderia baltica (NCBI Ref:NC_018010.1); Psychroflexus torquis I (NCBI Ref:NC_018721.1); Streptococcus thermophilus (NCBI Ref:YP_820832.1), Listeria innocua (NCBI Ref:NP_472073.1), Campylobacter jejuni (NCBI Ref:YP_002344900.1), or Neisseria meningitidis (NCBI Ref:YP_002342100.1) or Cas9 from any other organism.
[0147] In some embodiments, Cas9 is Neisseria meningitidis Cas9 (NmeCas9) or a variant thereof. In some embodiments, NmeCas9 is specific to NNNNGAYW PAM, where Y is C or T and W is A or T. In some embodiments, NmeCas9 is specific to NNNNGYTT PAM, where Y is C or T. In some embodiments, NmeCas9 is specific to NNNNGGTCT PAM. In some embodiments, NmeCas9 is Nme1 Cas9. In some embodiments, NmeCas9 is specific to NNNNGATT PAM, NNNCCTA PAM, NNNCCTC PAM, NNNCCTT PAM, NNNCCTG PAM, NNNCCGT PAM, NNNCCGGPAM, NNNCCCA PAM, NNNCCCT PAM, NNNCCCC PAM, NNNNCCAT PAM, NNNCCAG PAM, NNNNCCAT PAM, or NNNNGATT PAM. In some implementations, Nme1Cas9 is specific for NNNGATT PAM, NNNCCTA PAM, NNNCCTC PAM, NNNCCTT PAM, or NNNNCCTGPAM. In some implementations, NmeCas9 is specific for CAA PAM, CAAA PAM, or CCA PAM. In some implementations, NmeCas9 is Nme2Cas9. In some implementations, NmeCas9 is specific for NNNCC(N4CC)PAM, where N is any one of A, G, C, or T. In some implementations, NmeCas9 is specific for NNNCCGT PAM, NNNCCGGPAM, NNNCCCA PAM, NNNCCCT PAM, NNNCCCC PAM, NNNNCCAT PAM, NNNCCAG PAM, NNNNCCAT PAM, or NNNNGATT PAM. In some implementations, NmeCas9 is Nme3Cas9. In some implementations, NmeCas9 is specific to NNNNCAAA PAM, NNNNCC PAM, or NNNNCNNN PAM. In other implementations, the PAM-interaction domains for Nme1, Nme2, or Nme3 are N4GAT, N4CC, and N4CAAA, respectively.Other NmeCas9 characteristics and PAM sequence descriptions can be found in Edraki et al., A Compact, High-Accuracy Cas9 with aDinucleotide PAM for In Vivo Genome Editing, Mol.Cell. (2019) 73(4):714-726, which has been incorporated into this paper in full.
[0148] One example of a Neisseria meningitidis Cas9 protein, Nme1Cas9 (NCBI reference: WP_002235162.1; type II CRISPR RNA-guided endonuclease Cas9), has the following amino acid sequence:
[0149]
[0150]
[0151] Another example of the Neisseria meningitidis Cas9 protein, Nme2Cas9 (NCBI reference: WP_002230835; type II CRISPR RNA-guided endonuclease Cas9), has the following amino acid sequence:
[0152]
[0153] In some embodiments, dCas9 corresponds to or comprises part or all of the Cas9 amino acid sequence having one or more mutations that inactivate Cas9 nuclease activity. For example, in some embodiments, the dCas9 domain contains the D10A and H840A mutations or corresponding mutations in another Cas9. In some embodiments, dCas9 comprises the amino acid sequence of dCas9 (D10A and H840A):
[0154]
[0155] (Single bottom line: HNH domain; Double bottom line: RuvC domain).
[0156] In some implementations, the Cas9 domain contains the D10A mutation, while the residue at position 840 remains in the amino acid sequence provided above or at the corresponding position in any amino acid sequence provided herein, retaining a histidine residue.
[0157] In other embodiments, dCas9 variants with mutations other than D10A and H840A are provided, for example, Cas9 that causes nuclease inactivation (dCas9). Such mutations include, for example, other amino acid substitutions in D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). In some embodiments, variants or homologs of dCas9 are provided, such as at least about 70% homology, at least about 80% homology, at least about 90% homology, at least about 95% homology, at least about 98% homology, at least about 99% homology, at least about 99.5% homology, or at least about 99.9% homology. In some implementations, the provided dCas9 variants have shorter or longer amino acid sequences, differing by approximately 5 amino acids, approximately 10 amino acids, approximately 15 amino acids, approximately 20 amino acids, approximately 25 amino acids, approximately 30 amino acids, approximately 40 amino acids, approximately 50 amino acids, approximately 75 amino acids, or approximately 100 or more amino acids.
[0158] In some embodiments, the Cas9 fusion protein provided herein comprises the full-length amino acid sequence of the Cas9 protein, such as one of the Cas9 sequences provided herein. However, in other embodiments, the fusion protein provided herein does not comprise the full-length Cas9 sequence, but only one or more fragments thereof. Examples of suitable Cas9 domains and Cas9 fragments are provided herein, and those skilled in the art will understand additional suitable sequences for Cas9 domains and fragments.
[0159] It should be understood that additional Cas9 proteins (e.g., nuclease-inactivated Cas9 (dCas9), Cas9 cleavage enzyme (nCas9), or nuclease-active Cas9), including their variants and homologs, are within the scope of this invention. Examples of Cas9 proteins include (but are not limited to) those provided below. In some embodiments, the Cas9 protein is nuclease-inactivated Cas9 (dCas9). In some embodiments, the Cas9 protein is Cas9 cleavage enzyme (nCas9). In some embodiments, the Cas9 protein is nuclease-active Cas9.
[0160] Examples of catalytically deactivated Cas9 (dCas9):
[0161]
[0162] Examples of catalytic Cas9 nickases (nCas9):
[0163]
[0164] Examples of catalytically active Cas9:
[0165]
[0166] In some implementations, Cas9 refers to Cas9 from archaea (e.g., nanoarchaea), which belong to the domain and kingdom Unicellular Prokaryotes. In other implementations, Cas9 refers to CasX or CasY, as described, for example, in Burstein et al., “New CRISPR-Cas Systems from uncultivated microbes.” Cell Res. 2017 Feb 21. doi:10.1038 / cr.2017.21, the full contents of which are incorporated herein by reference. Many CRISPR-Cas systems have been identified using genome-resolved metagenomics, including the first reported Cas9 in the archaea domain of life. This diverse Cas9 protein has appeared in the little-studied nanoarchaea, which is part of an active CRISPR-Cas system. In bacteria, two previously unknown systems, CRISPR-CasX and CRISPR-CasY, have been discovered, and they are among the most robust systems discovered to date. In some embodiments, Cas9 refers to CasX, or a variant of CasX. In other embodiments, Cas9 refers to CasY, or a variant of CasY. It should be understood that other RNA-guided DNA-binding proteins can be used as nucleic acid programmable DNA-binding proteins (napDNAbp), and are within the scope of this invention.
[0167] In some embodiments, the Cas9 is a Cas9 variant specific to the altered PAM sequence. In some embodiments, additional Cas9 variants and PAM sequences are described in Miller et al., Continuous evolution of SpCas9 variants compatible with non-G PAMs. Nat Biotechnol (2020), doi.org / 10.1038 / s41587-020-0412-8, the full contents of which are incorporated herein by reference. In some embodiments, the Cas9 variant does not require a specific PAM. In some embodiments, the Cas9 variant, for example, the SpCas9 variant, is specific to NRNH PAM, where R is A or G and H is A, C, or T. In some embodiments, the SpCas9 variant is specific to PAM sequences AAA, TAA, CAA, GAA, TAT, GAT, or CAC. In some implementations, the SpCas9 variant contains amino acid substitutions at positions 1114, 1134, 1135, 1137, 1139, 1151, 1180, 1188, 1211, 1218, 1219, 1221, 1249, 1256, 1264, 1290, 1318, 1317, 1320, 1321, 1323, 1332, 1333, 1335, 1337, or 1339 (relative to the above reference sequence numbers), or at their corresponding positions.
[0168] (Single bottom line: HNH domain; Double bottom line: RuvC domain).
[0169] In some embodiments, the SpCas9 variant contains amino acid substitutions at positions 1114, 1135, 1218, 1219, 1221, 1249, 1320, 1321, 1323, 1332, 1333, 1335, or 1337 (relative to the reference sequence numbers above), or at their corresponding positions. In other embodiments, the SpCas9 variant contains amino acid substitutions at positions 1114, 1134, 1135, 1137, 1139, 1151, 1180, 1188, 1211, 1219, 1221, 1256, 1264, 1290, 1318, 1317, 1320, 1323, or 1333 (relative to the reference sequence numbers above), or at their corresponding positions. In some embodiments, the SpCas9 variant includes amino acid substitutions at positions 1114, 1131, 1135, 1150, 1156, 1180, 1191, 1218, 1219, 1221, 1227, 1249, 1253, 1286, 1293, 1320, 1321, 1332, 1335, and 1339 (relative to the reference sequence numbers above), or at their corresponding positions. In some embodiments, the SpCas9 variant includes amino acid substitutions at positions 1114, 1127, 1135, 1180, 1207, 1219, 1234, 1286, 1301, 1332, 1335, 1337, 1338, and 1349 (relative to the reference sequence numbers above). Examples of amino acid substitutions and PAM specificity for SpCas9 variants are shown in Tables AD and below. Figure 49 .
[0170] Table A
[0171]
[0172] Table B
[0173]
[0174] Table C
[0175]
[0176] Table D
[0177]
[0178]
[0179] In certain embodiments, the napDNAbp suitable for the method of the present invention comprises a circular arrangement known in the art and described, for example, Oakes et al., Cell 176, 254–267, 2019. An example of a circular arrangement is shown below, wherein the bold sequence indicates a sequence derived from Cas9, the italic sequence represents a linker sequence, and the underlined sequence represents a binuclear localization sequence.
[0180] CP5 (possessing the MSP "NGC=Pam variant, with mutational patterns such as NGG" PID=protein interaction domain and "D10A" cleavage enzyme):
[0181]
[0182]
[0183] Non-limiting examples of multinucleotide programmable nucleotide-binding domains that can be incorporated into a base editor include CRISPR protein-derived domains, restriction nucleases, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases (ZFNs).
[0184] In some embodiments, the nucleic acid-programmable DNA-binding protein (napDNAbp) of any fusion protein provided herein may be a CasX or CasY protein. In some embodiments, napDNAbp is a CasX protein. In some embodiments, napDNAbp is a CasY protein. In some embodiments, napDNAbp contains an amino acid sequence that is identical to that of a native CasX or CasY protein in at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%. In some embodiments, napDNAbp is a native CasX or CasY protein. In some embodiments, the amino acid sequence contained in napDNAbp is identical to any CasX or CasY protein described herein in at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%. It should be understood that Cas12b / C2c1, CasX, and CasY from other bacterial species may also be used according to the present invention.
[0185] Cas12b / C2c1(uniprot.org / uniprot / T0D7A2#2)
[0186] sp|T0D7A2|C2C1_ALIAG CRISPR-related endonuclease C2c1 OS=Alicyclobacillus acidoterrestris (strain ATCC 49025 / DSM 3922 / CIP 106132 / NCIMB13137 / GD3B) GN=c2c1 PE=1 SV=1
[0187]
[0188] CasX(uniprot.org / uniprot / F0NN87;uniprot.org / uniprot / F0NH53)
[0189] >tr|F0NN87|F0NN87_SULIH CRISPR-related CaSX protein OS = Sulfolobus islandicus (strain HVE10 / 4) GN = SiH_0402PE = 4SV = 1
[0190] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLE VEPHYLIIAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAVGQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG
[0191] >tr|F0NH53|F0NH53_SULIR CRISPR-related protein, Casx OS = Sulfolobus islandicus (strain REY15A) GN = SiRe_0771 PE = 4 SV = 1
[0192] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTSNIILPLSGNDKNPWTETLKCYNFPTTTVALSEVFKNFSQVKECEEVSAPSFVKPFEYKFGRSPGMVERTRRVKLEVEPHYLIMAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAVGQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG
[0193] δ- proteobacteria (Deltaproteobacteria) CasX
[0194] MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMTDDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAILQVYWQEFKDDHVGLMCKFAQPASKKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFVYKLEQVSEKGKAYTNYFGRCNVAEHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESTHPVKPLAQIAGNRYASGPVGKALSDACMGTIASFLSKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDAYNEVIARVRMWVNLLWQKLKLSRDDAKPLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENPKKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEMDEKEFYACEIQLQK WYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFTDGTDIKKSGKWQGLLYGGGAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKIGRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQAAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAKLAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKELSAELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSYKSGKQPFVGAWQAFYKRRLKEVWKPNA
[0195] CasY(ncbi.nlm.nih.gov / protein / APG80656.1)
[0196] >APG80656.1 CRISPR-related protein CaSY [uncultured Parcubacteria]
[0197]
[0198] The term "conserved amino acid substitution" or "conserved mutation" refers to the replacement of one amino acid with another amino acid that shares a common property. One functional way to define the common property between individual amino acids is to analyze the normalized frequency of amino acid changes between corresponding proteins in homologous organisms (Schulz, GE, and Schirmer, RH, Principles of Protein Structure, Springer-Verlag, New York (1979)). Based on such analyses, groups of amino acids can be defined, within which amino acids preferentially exchange with each other, thus having the most similar effects on the overall protein structure (Schulz, GE, and Schirmer, RH, as mentioned above). Non-restrictive examples of conserved mutations include amino acid substitutions, such as: lysine replacing arginine and vice versa, thus maintaining a positive charge; glutamic acid replacing aspartic acid and vice versa, thus maintaining a negative charge; serine replacing threonine, thus maintaining free –OH; and glutamic acid replacing aspartic acid, thus maintaining free –NH2.
[0199] As used interchangeably in this document, the terms "coding sequence" or "protein-coding sequence" refer to a segment of polynucleotides that encodes a protein. The region or sequence is defined by a start codon near the 5' end and by a stop codon near the 3' end. A coding sequence may also be referred to as an open reading frame.
[0200] As used herein, the terms "deaminase" or "deaminase domain" refer to proteins or enzymes that catalyze deamination reactions. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenine to form hypoxanthine. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or adenine (A) to form inosine (I). In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or deoxyadenosine to form inosine or deoxyinosine, respectively. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., engineered adenosine deaminases, evolved adenosine deaminases) can be derived from any organism, such as bacteria. In some implementations, adenosine deaminase is derived from bacteria such as Escherichia coli, Staphylococcus aureus, Salmonella typhimurium, Shewanella putrefaciens, Haemophilus influenzae, or Caulobacter crescentus.
[0201] In some embodiments, the adenosine deaminase is a TadA deaminase. In some embodiments, the TadA deaminase is a TadA variant. In some embodiments, the TadA variant is TadA*8. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism such as a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain is not naturally occurring. For example, in some embodiments, the deaminase or deaminase domain is consistent with naturally occurring deaminase at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9%. For example, the description of the deaminase domain in international PCT applications PCT / 2017 / 045381 (WO 2018 / 027078) and PCT / US2016 / 058344 (WO 2017 / 070632) is incorporated herein by reference in its entirety.See also Komor, AC et al. "Programmable base editing of a target base in genomic DNA without double-stranded DNA cleavage" Nature 533, 420-424 (2016); Gaudelli, NM et al. "Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage" Nature 551,464-471(2017); "Improved base excisionrepair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A baseeditors with higher efficiency and product purity" Science Advances 3:eaao4774(2017)) by Komor, AC et al.; and "Base editing: precision chemistry on the genome and transcriptome of living" by Rees, HA et al. cells." Nat Rev Genet.2018Dec;19(12):770-788.doi:10.1038 / s41576-018-0059-1, and its complete contents are incorporated herein by reference.
[0202] "Detection" refers to determining the presence, absence, or quantity of an analyte to be detected. In one embodiment, sequence modifications in polynucleotides or peptides are detected. In another embodiment, the presence of insertions / deletions (indels) is detected.
[0203] "Detectable label" means that when the composition is linked to the molecule of interest, the molecule can be detected using microscopy, photochemistry, biochemistry, immunochemistry, or chemical methods. For example, suitable labels include radioactive isotopes, magnetic beads, metal beads, colloidal particles, fluorescent dyes, high electron density reagents, enzymes (e.g., those commonly used in ELISA), biotin, digoxigenin, or haptens.
[0204] "Disease" means any condition or disorder that impairs or interferes with the normal function of cells, tissues, or organs. In one implementation, the disease is SCD. In another implementation, the disease is β-thalassemia.
[0205] As used herein, the term "effective amount" refers to the amount of a bioactive agent sufficient to elicit the desired biological response. The effective amount of the active compounds (groups) used in operating this invention to treat disease will vary depending on the administration method, the individual's age, weight, and general health. The appropriate dosage and duration of treatment will ultimately be determined by the participating physician or veterinarian. This amount is referred to as the "effective" amount. In a particular embodiment, the effective amount is the amount at which the base editor system of this invention (e.g., a fusion protein comprising a programmable DNA-binding protein, a nucleobase editor, and gRNA) is sufficient to modify SCD mutations in cells to achieve medical efficacy (e.g., reducing or controlling an individual's SCD or its symptoms). Such medical efficacy does not necessarily require sufficient modification of SCD in all cells of a tissue or organ, but only about 1%, 5%, 10%, 25%, 50%, 75%, or more of the cells in the tissue or organ. In one embodiment, the effective amount is sufficient to alleviate one or more SCD symptoms, including anemia and ischemia.
[0206] A “fragment” refers to a portion of a polypeptide or nucleic acid molecule. This portion contains at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the total length of a reference nucleic acid molecule or polypeptide. A fragment may contain 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids.
[0207] "Guide RNA" or "gRNA" refers to a polynucleotide that can specifically target a target sequence and form a complex with a polynucleotide programmable nucleotide binding domain protein (e.g., Cas9 or Cpf1). In one embodiment, the guide polynucleotide is guide RNA (gRNA). gRNA can be a complex of two or more RNAs or a single RNA molecule. A gRNA in the form of a single RNA molecule may be called a single guide RNA (sgRNA), but "gRNA" can be used interchangeably with guide RNA in the form of a single molecule or a complex of two or more molecules. Typically, gRNA in the form of a single RNA species contains two domains: (1) a domain that shares common homology with the target nucleic acid (e.g., and governs the binding of the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and contains a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to the tracrRNA described in Jinek et al., Science 337:816-821 (2012), the full disclosure of which is incorporated herein by reference. Other examples of gRNAs (e.g., those including domain 2) can be found in U.S. Provisional Patent Application USSN 61 / 874,682, filed September 6, 2013, entitled “Switchable Cas9 Nucleases and Uses Thereof”; and U.S. Provisional Patent Application USSN 61 / 874,746, filed September 6, 2013, entitled “Delivery System For Functional Nucleases”, the full disclosures of which are incorporated herein by reference. In some embodiments, the gRNA comprises two or more domains (1) and (2) and may be referred to as an “extended gRNA”. An extended gRNA will bind two or more Cas9 proteins and target nucleic acids in two or more separate regions described herein. gRNA contains a complementary nucleotide sequence to the target site, mediating the binding of the nuclease / RNA complex to the target site and providing sequence specificity for the nuclease:RNA complex. Those skilled in the art will understand that RNA polynucleotide sequences, such as gRNA sequences, including the nucleotide uracil (U), are pyrimidine derivatives, not thymine (T) nucleotides included in DNA polynucleotide sequences. In RNA, uracil is base-paired with adenine and substitutes for thymine during DNA transcription.
[0208] “Hb G-Makassar” or “Makassar” refers to the human β-heme variant, the human heme (Hb) G-Makassar variant or mutation (HB Makassar variant), which is an asymptomatic natural variant (E6A) of heme. Hb G-Makassar was first discovered in Indonesia (Mohamad, AS et al., 2018, Hematol. Rep., 10(3):7210(doi:10.4081 / hr.2018.7210). Hb G-Makassar exhibits slow mobility during electrophoresis. The β-heme variant of Makassar has anatomical abnormalities at the β-6 or A3 position, where glutamyl residues are typically replaced by alanyl residues. The replacement of a single amino acid, β-6 glutamyl, with valine in the gene encoding the β-globin subunit will cause sickle cell disease. Routine processes, such as isoelectric focusing, heme electrophoresis using cation exchange, high performance liquid chromatography (HPLC), and cellulose acetate electrophoresis, cannot separate Hb G-Makassar and HbS globin types because they have been found to have the same properties when analyzed using these methods. The results cannot correctly distinguish Hb G-Makassar and HbS are often confused with each other by those skilled in the art, leading to misdiagnosis of sickle cell disease (SCD).
[0209] "Hybridization" refers to hydrogen bonding, which can be a Watson-Crick, Hoogsteen, or reverse Hoogsteen hydrogen bond between complementary nucleobases. For example, adenine and thymine are complementary nucleobases that pair through hydrogen bonding.
[0210] The term "base repair inhibitor" or "IBR" refers to a protein that can inhibit the activity of nucleic acid repair enzymes, such as base excision repair (BER) enzymes. In some embodiments, the IBR is an inhibitor of inosine base excision repair. Examples of such base repair inhibitors include inhibitors of APE1, Endo III, Endo IV, Endo V, Endo VIII, Fpg, hOGGl, hNEILl, T7Endol, T4PDG, UDG, hSMUGl, and hAAG. In some embodiments, the IBR is an inhibitor of Endo V or hAAG. In some embodiments, the IBR is a catalytically inactivating EndoV or catalytically inactivating hAAG. In some embodiments, the base repair inhibitor is an inhibitor of Endo V or hAAG. In some embodiments, the base repair inhibitor is a catalytically inactivating EndoV or catalytically inactivating hAAG.
[0211] In some embodiments, the base repair inhibitor is a uracil glycosidase inhibitor (UGI). UGI refers to a protein that can inhibit uracil-DNA glycosidase base excision repair enzymes. In some embodiments, the UGI domain contains wild-type UGI or a fragment of wild-type UGI. In some embodiments, the UGI protein provided herein includes fragments of UGI and proteins homologous to UGI or UGI fragments. In some embodiments, the base repair inhibitor is an inhibitor of inosine base excision repair. In some embodiments, the base repair inhibitor is a “catalytically inactivating inosine-specific nuclease” or a “inactivating inosine-specific nuclease.” Without being limited by any specific theory, a catalytically inactivating inosine glycosidase (e.g., alkyl adenine glycosidase (AAG)) can bind inosine but does not create a base-free site or exclude inosine, thereby stereoblocking newly formed inosine moieties from DNA damage / repair mechanisms. In some embodiments, a catalytically inactivating inosine-specific nuclease can bind inosine in nucleic acids but does not cleave the nucleic acids. Non-limiting examples of catalytically inactivating inosine-specific nucleases include catalytically inactivating alkyl adenosine glycosidases (AAG nucleases), such as those from humans, and catalytically inactivating endonuclease V (EndoV nucleases), such as those from *Escherichia coli*. In some embodiments, the catalytically inactivating AAG nuclease contains an E125Q mutation or a corresponding mutation in another AAG nuclease.
[0212] "Improvement" means a positive change of at least 10%, 25%, 50%, 75%, or 100%.
[0213] "Inteins" are protein segments that can self-cleave and bind to other segments (exopeptides) via peptide bonds during a process called protein splicing. Inteins are also called "protein introns." The process by which inteins self-cleave and bind to other protein segments is referred to herein as "protein splicing" or "intein-mediated protein splicing." In some embodiments, the precursor protein inteins (the protein containing inteins preceding intein-mediated protein splicing) are derived from two genes. These inteins are referred to herein as split inteins (e.g., split intein-N and split intein-C). For example, in cyanobacteria, DnaE (the catalytic subunit of DNA polymerase III) is encoded by two separate genes, dnaE-n and dnaE-c. The intein encoded by the dnaE-n gene is referred to herein as "intein-N." The intein encoded by the dnaE-c gene is referred to herein as "intein-C."
[0214] Other intein systems may also be used. For example, synthetic inteins based on dnaE inteins: paired inteins of Cfa-N (e.g., cleaved intein-N) and Cfa-C (e.g., cleaved intein-C) have been described (e.g., Stevens et al., J Am Chem Soc. 2016 Feb. 24; 138(7):2162-5, which is incorporated herein by reference). Non-limiting examples of paired inteins that may be used according to the invention include: Cfa DnaE inteins, Ssp GyrB inteins, Ssp DnaX inteins, Ter DnaE3 inteins, Ter ThyX inteins, Rma DnaB inteins, and Cne Prp8 inteins (e.g., described in U.S. Patent No. 8,394,604, which is incorporated herein by reference).
[0215] Examples of nucleotide and amino acid sequences containing peptides are provided.
[0216] DnaE intima-N DNA:
[0217] TGCCTGTCATACGAAACCGAGATACTGACAGTAGAATATGGCCTTTGCCAATCGGGAAGATTGTGGAGAAACGGATAGAATGCACAGTTTACTCTGTCGATAACAATGGTAACATTTATACTCAGCCAGTTGCCCAGTGGCACGACCGGGGA GAGCAGGAAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAGGGCCACTAAGGACCACAAATTTATGACAGTCGATGGCCAGATGCTGCCTATAGACGAAATCTTTGAGCGAGAGTTGGACCTCATGCGAGTTGACAACCTTCCTAAT
[0218] DnaE contains peptide-N protein:
[0219] CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDRGEQEVFEYCLEDGSLIRATKDHKFMTVDGQMLPIDEIFERELDLMRVDNLPN
[0220] DnaE intipeptide-c DNA:
[0221] ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGATATTGGAGTCGAAAGAGATCACAACTTTGCTCTGAAGAACGGATTCATAGCTTCTAAT
[0222] Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN
[0223] Cfa-N DNA:
[0224] TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGCCTATTGGAAAGATTGTCGAAGAGAGAATTGAATGCACAGTATATACTGTAGACAAGAATGGTTTCGTTTACACACAGCCCATTGCTCAATGGCACAATCGCGGCGAACAAGAAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACGAGCAACTAAAGATCATAAATTCATGACCACTGACGGGCAGATGTTGCCAATAGATGAGATATTCGAGCGGGGCTTGGATCTCAAACAAGTGGATGGATTGCCA
[0225] Cfa-N protein:
[0226] CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQMLPIDEIFERGLDLKQVDGLP
[0227] Cfa-C DNA:
[0228] ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAGGAAAGTAAAGATAATATCTCGAAAAAGTCTTGGTACCCAAAATGTCTATGATATTGGAGTGGAGAAAGATCACAACTTCCTTCTCAAGAACGGTCTCGTAGCCAGCAAC
[0229] Cfa-C protein:
[0230] MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKN GLVASN
[0231] Inteptide-N and inteptide-C may be fused to the N-terminal portion and C-terminal portion of Cas9, respectively, for conjugation of the N-terminal portion and C-terminal portion of Cas9. For example, in some embodiments, inteptide-N is fused to the C-terminus of the N-terminal portion of Cas9, i.e., forming the structure N--[N-terminal portion of Cas9]-[inteptide-N]--C. In some embodiments, inteptide-C is fused to the N-terminus of the C-terminal portion of Cas9, i.e., forming the structure N-[inteptide-C]--[C-terminal portion of Cas9]-C. The mechanism of inteptide-mediated protein splicing of the protein to be fused with the inteptide (e.g., Cas9) is known in the art, for example, described in Shah et al., Chem Sci. 2014; 5(1):446-461, which is incorporated herein by reference. Methods for designing and using peptides are known in the art and described in, for example, WO2014004336, WO2017132580, US20150344549, and US20180127780, the contents of which are incorporated herein by reference in their entirety.
[0232] The terms "isolated," "purified," or "biopure" refer to the exclusion of various degrees of material that is normally naturally present in a component. "Isolated" represents the degree of separation from its original source or environment. "Purified" represents a degree of separation beyond that of isolated. A "purified" or "biopure" protein is one in which other material has been completely excluded so that any impurities cannot materially affect the protein's biological properties or cause other adverse consequences. That is, when the nucleic acid or peptide of this invention is produced by recombinant DNA technology, it is purified if there is substantially no cellular material, viral material, or culture medium, or if it is chemically synthesized, it is free of chemical precursors or other chemical substances. Purity and homology are typically determined using analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high-performance liquid chromatography. The term "purified" can mean that the nucleic acid or protein produces essentially a single band on an electrophoretic gel. Proteins that can be modified, for example, by phosphorylation or glycosylation, and different modifications produce different isolated proteins, can be separated and purified.
[0233] "Single polynucleotide" means that the nucleic acid (e.g., DNA) for which the gene is side-linked has been excluded from the natural genome of the organism from which the nucleic acid molecule of this invention is derived. Therefore, the term includes, for example, recombinant DNA incorporated into a vector; incorporated into an autonomously replicating plasmid or virus; or incorporated into the genome DNA of a prokaryote or eukaryote; or other molecules in which the sequences are separate from each other (e.g., cDNA or genome or cDNA fragments produced by PCR or restriction endonuclease digestion). Furthermore, the term includes RNA molecules transcribed from DNA molecules, and recombinant DNA that forms part of a hybrid gene encoding an additional polypeptide sequence.
[0234] "Isolated polypeptide" means the polypeptide of the present invention that has been isolated from its naturally associated components. Generally, a polypeptide is considered isolated when at least 60% by weight of naturally associated proteins and natural organic molecules have been excluded. Preferably, the polypeptide comprises at least 75%, more preferably at least 90%, and most preferably at least 99% by weight of the polypeptide of the present invention. The isolated polypeptide of the present invention can be obtained, for example, by extraction from natural sources, by expression of recombinant nucleic acids encoding such polypeptides, or by chemical synthesis of proteins. Purity can be determined by any suitable method, such as column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.
[0235] As used herein, the term "linker" can refer to a covalent linker (e.g., a covalent bond), a nonvalent linker, a chemical group, or a molecule that connects two molecules or parts thereof, such as two components of a protein complex or ribonucleoprotein complex, or two domains of a fusion protein, such as, for example, a polynucleotide-programmable DNA-binding domain (e.g., dCas9) and a deaminase domain (e.g., adenosine deaminase, or adenosine deaminase and cytidine deaminase, as described in PCT / US19 / 44935). Linkers can conjugate different components or different parts of a base editor system. For example, in some embodiments, a linker can conjugate a guide polynucleotide-binding domain of a polynucleotide-programmable nucleotide-binding domain and a catalytic domain of a deaminase. In some embodiments, a linker can conjugate a CRISPR peptide and a deaminase. In some embodiments, a linker can conjugate Cas9 and a deaminase. In some embodiments, a linker can conjugate dCas9 and a deaminase. In some embodiments, a linker can conjugate nCas9 and a deaminase. In some embodiments, the linker may conjugate a guide polynucleotide and a deaminase. In some embodiments, the linker may conjugate a deamination component of a base editor system and a programmable nucleotide-binding component of a polynucleotide. In some embodiments, the linker may conjugate an RNA-binding portion of a deamination component of a base editor system and a programmable nucleotide-binding component of a polynucleotide. In some embodiments, the linker may conjugate an RNA-binding portion of a deamination component of a base editor system and an RNA-binding portion of a programmable nucleotide-binding component of a polynucleotide. The linker may be located between or side-attached to two groups, molecules, or other parts, linked to each other via covalent bonds or non-covalent interactions, thus linking the two. In some embodiments, the linker may be an organic molecule, group, polymer, or chemical part. In some embodiments, the linker may be a polynucleotide. In some embodiments, the linker may be a DNA linker. In some embodiments, the linker may be an RNA linker. In some embodiments, the linker may include an aptamer capable of binding a ligand. In some embodiments, the ligand may be a carbohydrate, peptide, protein, or nucleic acid. In some embodiments, the linker may comprise an aptamer derived from a riboswitch. The riboswitch from which the aptamer is derived may be selected from theophylline riboswitch, thiamine pyrophosphate (TPP) riboswitch, adenosylcobalamin (AdoCbl) riboswitch, S-adenosylmethionine (SAM) riboswitch, SAH riboswitch, flavin mononucleotide (FMN) riboswitch, tetrahydrofolate riboswitch, lysine riboswitch, glycine riboswitch, purine riboswitch, GlmS riboswitch, or pre-queosine 1 (PreQ1) riboswitch. In some embodiments, the linker may comprise an aptamer that binds to a polypeptide or protein domain (e.g., a polypeptide ligand).In some embodiments, the peptide ligand may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMuCom coat protein domain, a sterile α motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the peptide ligand may be part of a base editor system component. For example, the nucleobase editing component may include a deaminase domain and an RNA recognition motif.
[0236] In some embodiments, the linker may be one amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker may be about 5-100 amino acids in length, for example: about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, or 90-100 amino acids in length. In some embodiments, the linker may be about 100-150, 150-200, 200-250, 250-300, 300-350, 350-400, 400-450, or 450-500 amino acids in length. Longer or shorter linkers are also considered.
[0237] In some embodiments, the linker conjugates the gRNA-binding domain of an RNA-programmable nuclease (including the Cas9 nuclease domain) and the catalytic domain of a nucleic acid-editing protein (e.g., adenosine deaminase). In some embodiments, the linker conjugates dCas9 and the nucleic acid-editing protein. For example, the linker is located between or side-attached to two groups, molecules, or other parts and covalently linked to each other, thus linking the two. In some embodiments, the linker is one or more amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical part. In some implementations, the linker is 5-200 amino acids in length, for example: 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140, 150, 160, 175, 180, 190, or 200 amino acids in length. Longer or shorter linkers are also considered.
[0238] In some implementations, the domain of the nucleobase editor is fused using a linker containing the following amino acid sequence: SGGSSGSETPGTSESATPESSGGS, SGGSSGGSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSP TSTEEGTSTE PSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS.
[0239] In some embodiments, the nucleobase editor domain is fused using a linker containing the amino acid sequence SGSETPGTSESATPES, which may also be referred to as the XTEN linker. In some embodiments, the linker contains the amino acid sequence SGGS. In some embodiments, the linker contains (SGGS). n (GGGS) n (GGGGS) n (G) n (EAAAK) n (GGS) n , SGSETPGTSESATPES, or (XP) n Motifs, or any combination thereof, wherein n is an integer independently between 1 and 30, and X is any amino acid. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15.
[0240] In some embodiments, the linker is 24 amino acids long. In some embodiments, the linker contains the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In some embodiments, the linker is 40 amino acids long. In some embodiments, the linker contains the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS. In some embodiments, the linker is 64 amino acids long. In some embodiments, the linker contains the amino acid sequence...
[0241] SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGSSGGS. In some embodiments, the linker is 92 amino acids in length. In some embodiments, the linker contains the amino acid sequence PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS.
[0242] "Marker" refers to any protein or polynucleotide that has modifications in its expression level or activity in relation to a disease or ailment.
[0243] As used herein, the term "mutation" refers to the substitution of a residue within a sequence of, for example, nucleic acids or amino acids, by another residue, or the deletion or insertion of one or more residues within the sequence. When describing mutations herein, the original residue is typically indicated first, followed by its position within the sequence, and then the newly substituted residue. Various methods for producing amino acid substitutions (mutations) provided herein are known in the art and are provided, for example, in Green and Sambrook's *Molecular Cloning: A Laboratory Manual* (4th edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)). In some embodiments, the base editor of the present invention can efficiently produce "intended mutations," such as point mutations, in nucleic acids (e.g., nucleic acids within an individual genome) without producing a significant number of unintended mutations, such as unintended point mutations. In some embodiments, the intended mutation is a mutation produced by a specific base editor (e.g., an adenosine base editor) that binds to a guide polynucleotide (e.g., gRNA) and is specifically designed to produce the intended mutation.
[0244] Typically, mutations created or identified in a sequence (e.g., the amino acid sequence described herein) are numbered based on a reference (or wild-type) sequence, i.e., a sequence that does not contain the mutation. Those skilled in the art readily understand how to determine the location of mutations in amino acid and nucleic acid sequences relative to a reference sequence.
[0245] The term "non-conservative mutation" refers to amino acid substitutions between different groups, such as lysine replacing tryptophan, or phenylalanine replacing serine, etc. In this case, non-conservative amino acid substitutions are preferably those that do not interfere with or inhibit the biological activity of the functional variant. Non-conservative amino acid substitutions can enhance the biological activity of the functional variant, making the functional variant more active than the wild-type protein.
[0246] The terms “nuclear localization sequence,” “nuclear localization signal,” or “NLS” refer to an amino acid sequence that facilitates protein importation into the cell nucleus. Nuclear localization sequences are known in the art and are described, for example, in Plank et al., International PCT application PCT / EP2000 / 011690, filed November 23, 2000, and published as WO / 2001 / 038547 on May 31, 2001, the disclosure of such examples of nuclear localization sequences being incorporated herein by reference in their entirety. In other embodiments, the NLS is, for example, an optimized NLS described by Koblan et al. in Nature Biotech. 2018, doi:10.1038 / nbt.4172. In some implementations, the NLS contains the amino acid sequence KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.
[0247] As used herein, the terms “nucleic acid” and “nucleic acid molecule” refer to compounds comprising nucleobases and acidic moieties (e.g., nucleosides, nucleotides, or nucleotide polymers). Typically, polymeric nucleic acids (e.g., nucleic acid molecules comprising three or more nucleotides) are linear molecules in which adjacent nucleotides are linked together by phosphodiester bonds. In some embodiments, “nucleic acid” refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides). In some embodiments, “nucleic acid” refers to an oligonucleotide chain comprising three or more individual nucleotide residues. The terms “oligonucleotide” and “polynucleotide” as used herein may be used interchangeably to refer to polymers of nucleotides (e.g., a single nucleotide chain containing at least three nucleotides). In some embodiments, “nucleic acid” includes RNA and single-stranded and / or double-stranded DNA. Nucleic acids may be naturally occurring in the context of, for example, genomic bodies, transcripts, mRNA, tRNA, rRNA, siRNA, snRNA, plasmids, mucinous plasmids, chromosomes, chromatids, or other naturally occurring nucleic acid molecules. On the other hand, nucleic acid molecules can be non-naturally occurring molecules, such as recombinant DNA or RNA, artificial chromosomes, engineered genomes, fragments thereof, or synthetic DNA, RNA, DNA / RNA hybrids, or including non-naturally occurring nucleotides or nucleosides. Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms include nucleic acid analogs, such as analogs having a non-phosphodiester backbone. Nucleic acids can be purified from natural sources, manufactured using recombinant expression systems, and, as needed, purified, chemically synthesized, etc. Where appropriate, for example, in chemically synthesized molecules, nucleic acids may contain nucleoside analogs, such as analogs having chemically modified bases or sugars, and backbone-modified analogs. Nucleic acid sequences are oriented from 5′ to 3′ unless otherwise specified. In some implementations, the nucleic acid is or contains natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyluridine, C5-propynylcytidine, C5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynylur ...chlorouridine, C5-chlorouridine, C5-propynylcytidine, C5-chlorouridine, C5-methylcytidine, 2-aminoadenosine, C5-chlorouridine, C5-chlorouridine, C5-chlorouridine, C5-chlorouridine, C5-chlorouridine, C5-chlorouridine, C5-chlorouridine, C5-chlorouridine, C5-chloro -aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-sideoxyadenosine, 8-sideoxyguanosine, O(6)-methylguanine, and 2-thiocytidine); chemically modified bases; biologically modified bases (e.g., methylated bases); bases embedded in the middle; modified sugars (2′-e.g., fluororibose, ribose, 2′-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., thiophosphate and 5′-N-iminophosphate chains).
[0248] The terms "nucleic acid programmable DNA-binding protein" or "napDNAbp" are used interchangeably with "polynucleotide programmable nucleotide-binding domain" and refer to a protein associated with a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid or guide polynucleotide (e.g., gRNA) that guides napDNAbp to a specific nucleic acid sequence. In some embodiments, the polynucleotide programmable nucleotide-binding domain is a polynucleotide programmable DNA-binding domain. In some embodiments, the polynucleotide programmable nucleotide-binding domain is a polynucleotide programmable RNA-binding domain. In some embodiments, the polynucleotide programmable nucleotide-binding domain is a Cas9 protein. The Cas9 protein may be associated with a guide RNA that guides the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, napDNAbp is a Cas9 domain, such as an active nuclease Cas9, a Cas9 nickase (nCas9), or an inactive nuclease Cas9 (dCas9). Non-limiting examples of nucleic acid programmable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i.Non-restricted examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, C sm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V Cas effector proteins, type VI Cas effector proteins, CARF, DinG, their homologs, or their modified or engineered forms; other nucleic acid programmable DNA-binding proteins, although not explicitly listed in this invention, are also within the scope of this invention. See, for example: Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct; 1:325-336. doi:10.1089 / crispr.2018.0033; Yan et al., “Functionally diverse type VCRISPR-Cas systems” Science. 2019 Jan 4; 363(6422):88-91. doi:10.1126 / science.aav7271, the full contents of which are incorporated herein by reference.
[0249] In this document, the terms “nucleobase,” “nitrogenous base,” or “base” are used interchangeably to refer to nitrogenous biological compounds that form nucleosides, which in turn become components of nucleotides. The ability of nucleobases to form base pairs and stack one on top of another directly forms long-chain helical structures, such as ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). The five nucleobases – adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U) – are called primitive or canonical. Adenine and guanine are derived from purines, while cytosine, uracil, and thymine are derived from pyrimidines. DNA and RNA may also contain other modified (non-primitive) bases. Non-limiting examples of modified nucleobases include hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydroxymethylcytosine. Hypoxanthine and xanthine can be produced by existing mutagens, both through deamination (replacing the amino group with a carbonyl group). Hypoxanthine can be modified from adenine. Xanthine can be modified from guanine. Uracil can be derived from the deamination of cytosine. A "nucleoside" is composed of a nucleobase and a pentose sugar (ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxyuridine, and deoxycytidine. Examples of nucleosides with modified nucleobases include inosine (I), xanthine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), and pseudouridine (Ψ). A "nucleotide" is composed of a nucleobase, a pentose sugar (ribose or deoxyribose), and at least one phosphate ester group.
[0250] As used herein, the terms "nucleobase editing domain" or "nucleobase editing protein" refer to proteins or enzymes that catalyze nucleobase modifications in RNA or DNA, such as the deamination of cytosine (or cytidine) to uracil (or uridine) or thymine (or thymidine), and the deamination of adenine (or adenosine) to hypoxanthine (or inosine), as well as the addition and insertion of non-template nucleotides. In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., adenine deaminase or adenosine deaminase; or cytidine deaminase or cytosine deaminase). In some embodiments, the nucleobase editing domain is more than one deaminase domain (e.g., adenine deaminase, or adenosine deaminase, and cytidine or cytosine deaminase, e.g., as described in PCT / US19 / 44935). In some embodiments, the nucleobase editing domain may be a naturally occurring nucleobase editing domain. In some implementations, the nucleobase editing domain may be an engineered or evolved nucleobase editing domain derived from a naturally occurring nucleobase editing domain. The nucleobase editing domain can originate from any organism, such as bacteria, humans, chimpanzees, gorillas, monkeys, cows, dogs, rats, or mice.
[0251] In this article, “acquisition” as used in the context of “acquisition agent” includes the methods of synthesizing, purchasing, generating, preparing, or otherwise acquiring the agent.
[0252] As used herein, "patient" or "individual" refers to a mammalian individual diagnosed with, suffering from, at risk of suffering from, developing, susceptible to, or suspected of suffering from or developing a disease or disorder. In some implementations, the term "patient" refers to a mammalian individual with a higher-than-average likelihood of developing a disease or disorder. Examples of patients may include humans, non-human primates, cats, dogs, pigs, cattle, horses, camels, alpacas, goats, sheep, rodents (e.g., mice, rabbits, rats, or guinea pigs), and other mammals that may benefit from the therapies disclosed herein. Human patient examples may be male and / or female.
[0253] "Patients in need" or "individuals in need" in this article refers to patients who have been diagnosed with, are at risk of or have, have previously been diagnosed with or are suspected of having a disease or illness.
[0254] The terms "pathogenic mutation," "pathogenic variant," "disease-causing mutation," "disease-causing variant," "harmful mutation," or "genetic mutation" refer to genetic modifications or mutations that increase an individual's susceptibility or predisposition to a particular disease or ailment. In some implementations, a pathogenic mutation involves the substitution of at least one wild-type amino acid for at least one pathogenic amino acid in a protein encoded by a gene.
[0255] The terms “protein,” “peptide,” “polypeptide,” and their grammatical equivalents are used interchangeably herein and refer to polymers in which amino acid residues are linked together by peptide (amide) bonds. These terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, proteins, peptides, or polypeptides are at least three amino acids in length. Proteins, peptides, or polypeptides refer to individual proteins or collections of proteins. One or more amino acids in a protein, peptide, or polypeptide may be modified, for example, by adding chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups, linkers for conjugation, functionalization, or other modifications, etc. Proteins, peptides, or polypeptides may be single molecules or multi-molecule complexes. Proteins, peptides, or polypeptides may be fragments of naturally occurring proteins or peptides. Proteins, peptides, or polypeptides may be naturally occurring, recombinant, or synthetic, or any combination thereof. The term “fusion protein” as used herein refers to a hybrid polypeptide containing protein domains from at least two different proteins. One of the proteins may be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxyl-terminal (C-terminal) portion, thus forming an amino-terminal fusion protein or a carboxyl-terminal fusion protein, respectively. The protein may contain various domains, such as nucleic acid-binding domains (e.g., the gRNA-binding domain of Cas9, which directs the protein to the target site) and nucleic acid cleavage domains, or catalytic domains of nucleic acid-editing proteins. In some embodiments, the protein contains a proteinaceous portion, such as the amino acid sequence constituting the nucleic acid-binding domain, and an organic compound, such as a compound that can act as a nucleic acid cleavage agent. In some embodiments, the protein is a complex containing or linked to nucleic acids (e.g., RNA or DNA). Any protein provided herein can be manufactured using any method known in the art. For example, the proteins provided herein can be produced by recombinant protein expression and purification, and are particularly suitable for fusion proteins containing peptide linkers. Methods for the expression and purification of recombinant proteins are well known, including their description in Green and Sambrook’s Molecular Cloning: A Laboratory Manual (4th edition, ColdSpring Harbor Laboratory Press, ColdSpring Harbor, NY (2012)), the full contents of which are incorporated herein by reference.
[0256] The polypeptides and proteins (including their functional portions and functional variants) disclosed herein may be modified to include synthetic amino acids replacing one or more naturally occurring amino acids. Such synthetic amino acids are known in the art and include, for example: aminocyclohexanecarboxylic acid, leucine, α-aminodecanoic acid, high-carbon serine, S-acetylaminomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine, β-phenylserine, β-hydroxyphenylalanine, phenylglycine, α-naphthylalanine, cyclohexylalanine, cyclohexylglycine, indoline- 2-Carboxylic acid, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, aminomalonic acid monoacylamine, N'-phenylmethyl-N'-methyl-lysine, N',N'-diphenylmethyl-lysine, 6-hydroxylysine, ornithine, α-aminocyclopentanecarboxylic acid, α-aminocyclohexanecarboxylic acid, α-aminocycloheptanecarboxylic acid, α-(2-amino-2-norbornene)carboxylic acid, α,γ-diaminobutyric acid, α,β-diaminopropionic acid, high-carbon phenylalanine, and α-tertiary butylglycine. Peptides and proteins may be associated with post-translational modifications of one or more amino acids of the peptide construct. Non-limiting examples of post-translational modifications include phosphorylation, acylation (including acetylation and formylation), glycosylation (including N-linking and O-linking), acylation, hydroxylation, alkylation (including methylation and ethylation), ubiquitination, addition of pyrrolidone carboxylic acid, formation of disulfide bridges, sulfation, cardamomylation, palmitoylation, isopreneation, farnesylation, geraniylation, glycosylphosphatidylinositylation, thioctic acidation, and iodination.
[0257] In the context of proteins or nucleic acids, the term "recombinant" as used herein refers to a protein or nucleic acid that does not occur naturally but is a human-engineered product. For example, in some embodiments, the recombinant protein or nucleic acid molecule contains an amino acid or nucleotide sequence that includes at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations relative to any naturally occurring sequence.
[0258] "Reduction" means a negative modification of at least 10%, 25%, 50%, 75%, or 100%.
[0259] "Reference" means title or control condition. In one embodiment, the reference is wild-type or healthy cells. In other embodiments, and without limitation, the reference is untreated cells not subjected to the test conditions, or cells treated with placebo or saline, medium, buffer, and / or a control vector without the polynucleotide of interest.
[0260] A "reference sequence" is a designated sequence used as the basis for sequence alignment. The reference sequence can be a subset or the entirety of a specific sequence; for example, a segment of a full-length cDNA or gene sequence, or the complete cDNA or gene sequence. For peptides, the reference peptide sequence is typically at least about 16 amino acids, at least about 20 amino acids, more preferably at least about 25 amino acids, and even more preferably about 35 amino acids, about 50 amino acids, or about 100 amino acids in length. For nucleic acids, the reference nucleic acid sequence is typically at least about 50 nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, and about 100 nucleotides, or about 300 nucleotides, or any integer between or near that number. In some embodiments, the reference sequence is the wild-type sequence of the protein of interest. In other embodiments, the reference sequence is a polynucleotide sequence encoding a wild-type protein.
[0261] The terms “RNA-programmable nuclease” and “RNA-guided nuclease” are used with one or more RNAs that are not cleavage targets (e.g., binding or associating). In some embodiments, the RNA-programmable nuclease may be referred to as a nuclease:RNA complex when complexed with RNA. Typically, the bound RNA is called guide RNA (gRNA). gRNA may be a complex of two or more RNAs or a single RNA molecule. gRNA in the form of a single RNA molecule may be called a single-guide RNA (sgRNA), but “gRNA” may be used interchangeably with guide RNA in the form of a single molecule or a complex of two or more molecules. Typically, gRNA in the form of a single RNA species contains two domains: (1) a domain that shares common homology with the target nucleic acid (e.g., and predominantly binds the Cas9 complex to the target); and (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and contains a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to the tracrRNA described in Science 337:816-821 (2012) by Jinek et al., the full contents of which are incorporated herein by reference. Other examples of gRNAs (e.g., those including domain 2) can be found in U.S. Provisional Patent Application No. 6 September 2013, USSN 61 / 874,682, entitled “Switchable Cas9 Nucleases and Uses Thereof”, and U.S. Provisional Patent Application No. 6 September 2013, USSN 61 / 874,746, entitled “Delivery System For Functional Nucleases”, the full contents of which are incorporated herein by reference. In some embodiments, the gRNA comprises two or more domains (1) and (2), and may be referred to as “extended gRNA”. For example, an extended gRNA will, for example, bind two or more Cas9 proteins and bind target nucleic acids in two or more separate regions, as described herein. gRNA contains a nucleotide sequence that complements the target site, mediating the binding of the nuclease / RNA complex to the target site and providing sequence specificity for the nuclease:RNA complex.
[0262] In some implementations, the RNA-programmable nuclease is a Cas9 endonuclease (a CRISPR-related system), such as Cas9 (Csnl) from Streptococcus pyogenes (see, for example, "Completegenome sequence of an Ml strain of Streptococcus pyogenes", Ferretti JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III", Deltcheva E.,Chylinski K.,SharmaCM.,Gonzales K.,Chao Y.,Pirzada ZA,Eckert MR,Vogel J.,Charpentier E.,Nature 471:602-607(2011).
[0263] Because RNA-programmable nucleases (e.g., Cas9) use RNA:DNA hybridization to target DNA cleavage sites, these proteins can, in principle, be targeted by any sequence specified by the guide RNA. The use of RNA-programmable nucleases, such as Cas9, for site-specific cleavage (e.g., to modify the genome) is known in the art (see, for example: Cong, L. et al., Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, WY et al., Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic acids). research (2013); Jiang, W. et al., RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature biotechnology 31, 233-239 (2013); the complete contents of these articles are incorporated herein by reference.
[0264] The term "single nucleotide polymorphism (SNP)" refers to a single nucleotide change occurring at a specific location in the genome, where each change occurs at a significant level (e.g., >1%) in the population. For example, at a specific base position in the human genome, the C nucleotide may be present in most individuals, but in a minority of individuals, that position is occupied by A. This indicates the presence of an SNP at that specific location, and the two possible nucleotide changes, C or A, are called the pairwise genes at that location. SNPs represent fundamental differences in disease susceptibility. The severity of disease and our body's response to treatment are also characteristics of genetic changes. SNPs can fall within gene coding regions, non-gene coding regions, or intergenetic regions (regions between genes). In some implementations, due to genetic code degeneracy, SNPs within the coding sequence do not necessarily alter the amino acid sequence of the protein produced. There are two types of SNPs within coding regions: synonymous and non-synonymous SNPs. Synonymous SNPs do not affect the protein sequence, while non-synonymous SNPs alter the amino acid sequence of the protein. There are two types of non-synonymous SNPs: missense and nonsense. SNPs not located in protein-coding regions can still affect gene splicing, transcription factor binding, messenger RNA degradation, or the sequence of non-coding RNA. Gene expression affected by such SNPs is called an eSNP (expressed SNP) and can occur upstream or downstream of the gene. Single nucleotide variants (SNVs) are changes in a single nucleotide, with no frequency restrictions and can originate from body cells. Single nucleotide changes in the body can also be called single nucleotide modifications.
[0265] "Specific binding" means that a nucleic acid molecule, polypeptide, or complex thereof (e.g., a nucleic acid programmable DNA binding domain and a guide nucleic acid), compound, or molecule can recognize and bind to the polypeptide and / or nucleic acid molecules of the present invention, but cannot substantially recognize and bind to other molecules in a sample (e.g., a biological sample).
[0266] Nucleic acid molecules suitable for the methods of this invention include any nucleic acid molecule encoding the polypeptide or fragment thereof of this invention. Such nucleic acid molecules do not need to be 100% identical to the intrinsic nucleic acid sequence, but generally have substantial identity. Polynucleotides having “substantial identity” with the intrinsic sequence can generally hybridize with at least one strand of a double-stranded nucleic acid molecule. “Hybridization” means the pairing of complementary polynucleotide sequences (e.g., genes described herein) or portions thereof, under various stringent conditions to form a double-stranded molecule (see, for example: Wahl, GM & SLBerger (1987) Methods Enzymol. 152:399; Kimmel, AR (1987) Methods Enzymol. 152:507).
[0267] For example, the harsh salt concentration is typically below about 750 mM NaCl and 75 mM trisodium citrate, preferably below about 500 mM NaCl and 50 mM trisodium citrate, and even more preferably below about 250 mM NaCl and 25 mM trisodium citrate. Low-hardness hybridization can be achieved in the absence of organic solvents (e.g., formamide), while high-hardness hybridization can be achieved in the presence of at least about 35% formamide, and more preferably at least about 50% formamide. Harsh temperature conditions typically include temperatures of at least about 30°C, more preferably at least about 37°C, and most preferably at least about 42°C. Other parameters, such as hybridization time, detergent (e.g., sodium dodecyl sulfate (SDS)) concentration, and the inclusion or exclusion of vector DNA, are well known to those skilled in the art. Various degrees of harshness can be achieved by combining these different conditions as needed. In one embodiment, the hybridization is performed at 30°C in 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS. In another embodiment, the hybridization is performed at 37°C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA). In yet another embodiment, the hybridization is performed at 42°C in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 μg / ml ssDNA. Variations in these conditions are readily understood by those skilled in the art.
[0268] For most applications, the washing step following the hybridization process varies in severity. The severity of the washing conditions can be defined by salt concentration and temperature. As mentioned above, the severity of the washing can be increased by decreasing the salt concentration or increasing the temperature. For example, the preferred salt concentration for the washing step is below about 30 mM NaCl and 3 mM trisodium citrate, and most preferably below about 15 mM NaCl and 1.5 mM trisodium citrate. The preferred temperature conditions for the washing step generally include at least about 25°C, more preferably at least about 42°C, and even more preferably at least about 68°C. In one embodiment, the washing step is carried out at 25°C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the washing step is carried out at 42°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In a preferred embodiment, the washing step is performed at 68°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Other variations of these conditions are readily apparent to those skilled in the art. Hybridization temperatures are well known to those skilled in the art and described, for example, by: Benton and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, ColdSpring Harbor Laboratory Press, New York.
[0269] "Splitting" means dividing into two or more segments.
[0270] "Split Cas9 protein" or "split Cas9" refers to the fact that the Cas9 protein is provided as an N-terminal fragment and a C-terminal fragment encoded by two separate nucleotide sequences. The polypeptides corresponding to the N-terminal and C-terminal portions of the Cas9 protein can be spliced to "reassemble" the Cas9 protein. In certain embodiments, the Cas9 protein is split into two fragments within the disordered region of the protein, for example, as described in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014; or in Jiang et al. (2016) Science 351:867-871. PDB file:5F9R, the contents of which are incorporated herein by reference. In some embodiments, the protein is split into two fragments at any C, T, A, or S position within the region of approximately amino acids A292-G364, F445-K483, or E565-T637 of SpCas9, or at the corresponding position in any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or other nap DNAbp. In some embodiments, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of splitting the protein into two fragments is referred to as “splitting” the protein.
[0271] In other embodiments, the N-terminal portion of the Cas9 protein contains amino acids 1-573 or 1-637 of wild-type Cas9 (SpCas9) of Streptococcus pyogenes (NCBI reference sequence: NC_002737.2, Uniprot reference sequence: Q99ZW2), or their corresponding positions / mutations, and the C-terminal portion of the Cas9 protein contains amino acids 574-1368 or 638-1368 of wild-type SpCas9.
[0272] The C-terminal and N-terminal portions of Cas9 are combined to form the complete Cas9 protein. In some embodiments, the C-terminal portion of the Cas9 protein begins at the end of the N-terminal portion. Therefore, in some embodiments, the C-terminal portion of Cas9 includes the (551-651)-1368 amino acid moiety of spCas9. "(551-651)-1368" means starting with amino acids 551-651 (inclusive) and ending with amino acid 1368.For example, the C-terminal portion of Cas9 can contain any part of the amino acids in spCas9: 551-1368, 552-1368, 553-1368, 554-1368, 555-1368, 556-1368, 557-1368, 558-1368, 559-1368, 560-1368, 561-1368, 562-1368, 563-1368, 564-1368, 565-1368, 566-1368, 567-1368, 568-1368, 569-1368, 570-1368, 571-1368, 572-1368, 573- 1368, 574-1368, 575-1368, 576-1368, 577-1368, 578-1368, 579-1368, 580-1368, 581-1368, 582-1368, 583-1368, 584-1368, 585-1368, 586-1368, 587-1368, 588-1368, 589-1368, 590-1368, 591-1368, 592-1368, 593-1368, 594-1368, 595-1368, 596-1368, 597-1368, 598-1368, 599-13 68, 600-1368, 601-1368, 602-1368, 603-1368, 604-1368, 605-1368, 606-1368, 607-1368, 608-1368, 609-1368, 610-1368, 611-1368, 612-1368, 613-1368, 614-1368, 615-1368, 616-1368, 617-1368, 618-1368, 619-1368, 620-1368, 621-1368, 622-1368, 623-1368, 624-1368, 625-1368 626-1368, 627-1368, 628-1368, 629-1368, 630-1368, 631-1368, 632-1368, 633-1368, 634-1368, 635-1368, 636-1368, 637-1368, 638-1368, 639-1368, 640-1368, 641-1368, 642-1368, 643-1368, 644-1368, 645-1368, 646-1368, 647-1368, 648-1368, 649-1368, 650-1368, or 651-1368. In some implementations, the C-terminal portion of the Cas9 protein contains a subset of amino acids 574-1368 or 638-1368 of SpCas9.
[0273] "Individual" means mammal, including (but not limited to): human or non-human mammals, such as: cows, horses, dogs, sheep, or cats. Individuals include livestock, domesticated animals raised to provide labor and consumer goods, such as: food, including (but not limited to): cows, goats, chickens, horses, pigs, rabbits, and sheep.
[0274] "Substantially identical" means that the polypeptide or nucleic acid molecule and a reference amino acid sequence (e.g., any of the amino acid sequences described herein) or nucleic acid sequence (e.g., any of the nucleic acid sequences described herein) have at least 50% similarity. In one embodiment, such sequences and the sequences used for alignment have at least 60%, 80%, 85%, 90%, 95%, or even 99% similarity at the amino acid level or nucleic acid level.
[0275] Sequence identity is typically measured using sequence analysis software (e.g., Sequence Analysis Software Package of the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705; BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs). This software matches identical or similar sequences by specifying and varying degrees of homology with different substitutions, deletions, and / or other modifications. Conserved substitutions typically include substitutions from the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamylamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. An example method for determining identity is the BLAST program, which is used by e -3 to e -100 The probability score between sequences is used to indicate high correlation. For example, when using COBALT, the following parameters are used:
[0276] a) Ranking parameters: Penalty for empty spaces -11, -1, and penalty for empty spaces at the end -5, -1.
[0277] b) CDD parameters: RPS BLAST enabled; Blast E-value 0.003; search for conserved amino acid positions and enable repeat calculations, and
[0278] c) Query cluster parameters: Enable query cluster; character size 4; maximum cluster distance 0.8; general letters.
[0279] Use EMBOSS needles, for example, according to the following parameters:
[0280] a) Matrix: BLOSUM62;
[0281] b) Gap Open: 10;
[0282] c) Gap Extension: 0.5;
[0283] d) Output format: a pair;
[0284] e) Penalty for open play at the end: None;
[0285] f) End gap open: 10; and
[0286] g) End gap extension: 0.5.
[0287] The term "target site" refers to a sequence within a nucleic acid molecule modified by a nucleobase editor. In one embodiment, the target site is deaminated by a deaminase or a fusion protein containing a deaminase (e.g., adenine deaminase).
[0288] As used herein, the terms "treat," "treatment," and similar terms refer to reducing or alleviating a disease, ailment, and / or symptom associated with it, or achieving a desired pharmacological and / or physiological effect. It is generally understood that, while not excluded, treating a disease or ailment does not require the complete elimination of the disease, ailment, and / or symptom associated with it. In some embodiments, the effect is medical, meaning, without limitation, the effect is a partial or complete reduction, elimination, abolition, mitigation, alleviation, or reduction of the severity of the disease and / or symptoms attributed to it, or a cure. In some embodiments, the effect is preventative, meaning the effect protects against or prevents the occurrence or recurrence of the disease or ailment. Therefore, the method of the present invention includes administering a medically effective amount of the composition described herein. In some embodiments, the disease or ailment is sickle cell disease (SCD) or β-thalassemia.
[0289] "Uracil glycosidase inhibitor" or "UGI" refers to an agent that inhibits the uracil-excision repair system. In one embodiment, the agent is a protein or fragment thereof that binds to host uracil-DNA glycosidase and prevents uracil residues from being released from DNA. In one embodiment, UGI is a protein, fragment thereof, or domain that can inhibit the uracil-DNA glycosidase base excision repair enzyme. In some embodiments, the UGI domain comprises wild-type UGI or a modified form thereof. In some embodiments, the UGI domain comprises a fragment of the amino acid sequence listed below. In some embodiments, the UGI fragment comprises an amino acid sequence comprising at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the UGI sequence listed below. In some embodiments, the UGI comprises an amino acid sequence homologous to the UGI amino acid sequence or fragment thereof listed below. In some implementations, the UGI, or a portion thereof, is identical to the wild-type UGI or UGI sequence shown below, or a portion thereof, by at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, or 100%. An example of a UGI contains the following amino acid sequence:
[0290] >splP14739IUNGI_BPPB2 Uracil-DNA glycosidase inhibitor
[0291] MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAY DESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML.
[0292] The term "vector" refers to a method of introducing a nucleic acid sequence into a cell, resulting in transgenic cells. Vectors include plasmids, transposons, bacteriophages, viruses, liposomes, and free genes. An "expression vector" is a nucleic acid sequence containing a nucleotide sequence intended to be expressed in recipient cells. Expression vectors may include additional nucleic acid sequences that promote and / or facilitate the expression of the introduced sequence, such as initiation, termination, enhancer, promoter, and secretion sequences.
[0293] Any composition or method provided herein may be combined with one or more other compositions and methods provided herein.
[0294] DNA editing has become a viable method for modifying diseases at the genetic level by correcting pathogenic mutations. Until recently, all DNA editing platforms have the capability to induce DNA double-strand breaks (DSBs) at specified genomic sites, and the product outcome is determined in a semi-random manner, depending on the endogenous DNA repair pathway, resulting in complex populations of genetic products. Although precise, user-defined repair outcomes can be achieved through the homologous gene-guided repair (HDR) pathway, many challenges remain in medically relevant cell types that prevent the use of HDR for high-efficiency repair. In fact, this pathway is less efficient than the competing and error-prone non-homologous end-joining pathway. Furthermore, HDR is strictly limited to the G1 and S phases of the cell cycle, preventing precise repair of DSBs in post-mitotic cells. Therefore, it has been proven difficult or impossible to efficiently modify genomic sequences in these populations in a user-defined, programmable manner. Attached Figure Description
[0295] Figures 1A-1C Describe the plasmid. Figure 1A This is an expression vector for encoding the TadA7.10-dCas9 base editor. Figure 1B This is a plasmid containing a nucleic acid molecule encoding proteins that confer resistance to chloramphenicol (CamR) and spectinomycin (SpectR). The plasmid also contains a kanamycin resistance gene that has been deprecated by two point mutations. Figure 1C The plasmid contains a nucleic acid molecule encoding proteins that confer chloramphenicol resistance (CamR) and zicromycin resistance (SpectR). The plasmid also contains a kanamycin resistance gene that has been deactivated by three point mutations.
[0296] Figure 2 Image of a bacterial community transduced using the expression vector depicted in Figure 1 is presented, including the defective kanamycin resistance gene. The vector containing the ABE7.10 variant was generated using error-prone PCR. For kanamycin resistance, bacterial cells expressing these "evolved" ABE7.10 variants were selected using progressively increasing concentrations of kanamycin. Bacteria expressing the ABE7.10 variant and possessing adenosine deaminase activity were able to correct the mutation introduced into the kanamycin resistance gene, thereby restoring kanamycin resistance. Kanamycin-resistant cells were selected for further analysis.
[0297] Figure 3A and 3B This study describes the editing of the regulatory region of the heme subunit γ (HGB1) locus, which is a medically relevant site for upregulating fetal heme. Figure 3A The attached figure shows a portion of the regulatory region of the HGB1 gene. Figure 3BThe efficacy and specificity of the adenosine deaminase variants listed in Table 15 were quantified. Editing was performed on the heme subunit γ1 (HGB1) locus in HEK293T cells, a medically relevant site for upregulating fetal heme. The figure above shows the nucleotide residues in the target region of the HGB1 gene regulatory sequence. A5, A8, A9, and A11 represent edited adenosine residues in HGB1.
[0298] Figure 4 This illustrates the relative effectiveness of adenosine base editors containing dCas9 that can recognize atypical PAM sequences. The top figure shows the coding sequence for the heme subunit. The bottom figure demonstrates the effectiveness of base editors using adenosine deaminase variants with guide RNAs of different lengths.
[0299] Figure 5 The accompanying figure illustrates the efficacy and specificity of ABE8. It quantifies the percentage of editing in expected and unexpected target nucleotides (bystanders).
[0300] Figure 6 The accompanying figure illustrates the efficacy and specificity of ABE8. It quantifies the percentage of editing in expected and unexpected target nucleotides (bystanders).
[0301] Figures 7A-7C Graphs and bar charts depicting the transition from A·T to G·C and the phenotypic results in primary cells. Figure 7A Diagrams of the globin genes located on chromosome 11 in embryos, fetuses, and adults are presented, and the HBG1 / 2HPFH site is indicated, where a single-base editor introduces double-helix editing. Figure 7B This diagram illustrates the DNA editing efficiency in CD34+ cells. It shows the A·T to G·C conversion at the -198HBG1 / 2 promoter site in ABE-treated CD34+ cells from two separate donors. NGS analysis was performed at 48 and 144 hours post-treatment. The -198HBG1 / 2 target sequence is as follows: A7 is in bold and underlined with a double line. Draw percentages from A·T to G·C on A7. Figure 7C The attached figure reflects the percentage of γ-globin / α-globin expression in erythrocytes derived from ABE-edited cells. Figure 7C Show the percentage of γ-globin formed, as a ratio relative to α-globin. Figure 7B and 7C The values are from two different donors who underwent ABE treatment and erythrocyte differentiation. For example... Figure 7B The observed ABE8 editing efficacy at the -198HBG1 / 2 promoter target site was 2-3 times higher than at an earlier time point (48hr). Figure 7CObservations showed that ABE8 editing in CD34+ cells increased γ-globin formation in differentiated erythrocytes by approximately 1.4-fold. For example, the ABE8 13-d base editor resulted in 55% γ-globin / α-globin expression.
[0302] Figure 8A and 8B Depicting the A·T to G·C transition of CD34+ cells at the -198 promoter site upstream of HBG1 / 2 after ABE8 treatment. Figure 8A To depict the editing frequencies of A to G in CD34+ cells from two donors, 48 and 144 h after editor treatment, donor 2 was a heterozygote for sickle cell disease. Figure 8B A diagram representing the total ordered segment distribution containing only A7 edits or merged (A7+A8) edits.
[0303] Figure 9 A heatmap depicting the insertion / deletion (INDEL) frequencies at the -198 site of the γ-globin promoter in CD34+ cells after ABE8 treatment. Frequencies from two donors are shown at 48h and 144h time points. At the HBG1 / 2-198 promoter target site described in this paper, complete A·T to G·C conversion creates 10nt poly-G strips. An increased insertion / deletion (INDEL) frequency was observed at this site because the handling of such homopolymers often increases the proportion of PCR- and sequencing-induced errors.
[0304] Figure 10 Present ultra-high performance liquid chromatography (UHPLC) UV-Vis trajectory (220 nm) and integral of globin chain stage of untreated differentiated CD34+ cells (donor 1).
[0305] Figure 11 UHPLC UV-Vis trajectory (220 nm) and integral of globin chain stage of differentiated CD34+ cells (donor 1) treated with ABE7.10-m were plotted.
[0306] Figure 12 UHPLC UV-Vis trajectory (220 nm) and integral of globin chain stage of differentiated CD34+ cells (donor 1) treated with ABE7.10-d were plotted.
[0307] Figure 13 UHPLC UV-Vis trajectory (220 nm) and integral of globin chain stage of differentiated CD34+ cells (donor 1) treated with ABE8.8-m were plotted.
[0308] Figure 14UHPLC UV-Vis trajectory (220 nm) and integral of globin chain stage of differentiated CD34+ cells (donor 1) treated with ABE8.8-d were plotted.
[0309] Figure 15 UHPLC UV-Vis trajectory (220 nm) and integral of globin chain stage of differentiated CD34+ cells (donor 1) treated with ABE8.13-m were plotted.
[0310] Figure 16 UHPLC UV-Vis trajectory (220 nm) and integral of globin chain stage of differentiated CD34+ cells (donor 1) treated with ABE8.13-d were plotted.
[0311] Figure 17 UHPLC UV-Vis trajectory (220 nm) and integral of globin chain stage of differentiated CD34+ cells (donor 1) treated with ABE8.17-m were plotted.
[0312] Figure 18 UHPLC UV-Vis trajectory (220 nm) and integral of globin chain stage of differentiated CD34+ cells (donor 1) treated with ABE8.17-d were plotted.
[0313] Figure 19 UHPLC UV-Vis trajectory (220 nm) and integral of globin chain stage of differentiated CD34+ cells (donor 1) treated with ABE8.20-m were plotted.
[0314] Figure 20 UHPLC UV-Vis trajectory (220 nm) and integral of globin chain stage of differentiated CD34+ cells (donor 1) treated with ABE8.20-d were plotted.
[0315] Figure 21 UHPLC UV-Vis trajectory (220 nm) and integral of globin chain phase were plotted for untreated differentiated CD34+ cells (donor 2). Note: Donor 2 is a heterozygote for sickle cell disease.
[0316] Figure 22 UHPLC UV-Vis trajectory (220 nm) and integral of the globin chain stage were plotted for differentiated CD34+ cells (donor 2) treated with ABE7.10-m. Note: Donor 2 is a heterozygote for sickle cell disease.
[0317] Figure 23UHPLC UV-Vis trajectory (220 nm) and integral of the globin chain stage were plotted for differentiated CD34+ cells (donor 2) treated with ABE7.10-d. Note: Donor 2 is a heterozygote for sickle cell disease.
[0318] Figure 24 UHPLC UV-Vis trajectory (220 nm) and integral of the globin chain stage were plotted for differentiated CD34+ cells (donor 2) treated with ABE8.8-m. Note: Donor 2 is a heterozygote for sickle cell disease.
[0319] Figure 25 UHPLC UV-Vis trajectory (220 nm) and integral of the globin chain stage were plotted for differentiated CD34+ cells (donor 2) treated with ABE8.8-d. Note: Donor 2 is a heterozygote for sickle cell disease.
[0320] Figure 26 UHPLC UV-Vis trajectory (220 nm) and integral of the globin chain stage were plotted for differentiated CD34+ cells (donor 2) treated with ABE8.13-m. Note: Donor 2 is a heterozygote for sickle cell disease.
[0321] Figure 27 UHPLC UV-Vis trajectory (220 nm) and integral of the globin chain stage were plotted for differentiated CD34+ cells (donor 2) treated with ABE8.13-d. Note: Donor 2 is a heterozygote for sickle cell disease.
[0322] Figure 28 UHPLC UV-Vis trajectory (220 nm) and integral of globin chain stage of differentiated CD34+ cells (donor 1) treated with ABE8.17-m were plotted.
[0323] Figure 29 UHPLC UV-Vis trajectory (220 nm) and integral of the globin chain stage were plotted for differentiated CD34+ cells (donor 2) treated with ABE8.17-d. Note: Donor 2 is a heterozygote for sickle cell disease.
[0324] Figure 30A and 30B UHPLC UV-Vis trajectory (220 nm) and integral of globin chain stage of differentiated CD34+ cells treated with ABE8 were plotted. Figure 30A UHPLC UV-Vis trajectory (220 nm) and integral of globin chain phase of differentiated CD34+ cells (donor 2) treated with ABE8.20-m were plotted. Note: Donor 2 is a heterozygote for sickle cell disease. Figure 30BUHPLC UV-Vis trajectory (220 nm) and integral of the globin chain stage were plotted for differentiated CD34+ cells (donor 2) treated with ABE8.20-d. Note: Donor 2 is a heterozygote for sickle cell disease.
[0325] Figures 31A-31E The study described editing using ABE8.8 at two independent sites, achieving 90% editing before enucleation on day 11 post-erythrocyte differentiation and approximately 60% γ-globin on day 18 post-erythrocyte differentiation, exceeding α-globin or total β-family globin. Figure 31A A graph depicting the mean ABE of 8.8 edited from two healthy donors in two independent experiments. Editing efficacy was measured using primer measurements that distinguish between HBG1 and HBG2. Figure 31B A graph depicting the average values of one healthy donor across two independent experiments. Editing efficiency was achieved using primer measurements that simultaneously identify HBG1 and HBG2. Figure 31C A diagram illustrating the editing of ABE8.8 in donors with heterozygous E6V mutations. Figure 31D and 31E A diagram depicting the increase of γ-globin in ABE8.8 edited cells.
[0326] Figure 32A and 32B The percentage of edits corrected for sickle cell mutations using ABE variants is depicted. Figure 32A A diagram depicting different editor variants with approximately 70% edits in fibroblasts from SCD patients. Figure 32B A diagram illustrating CD34 cells from healthy donors edited by a lead ABE variant, which targets the synonymous mutation A13 at an adjacent proline residue within the editing window, acting as a surrogate for editing SCD mutations. The ABE8 variant exhibits an average editing frequency of approximately 40% in the surrogate A13.
[0327] Figure 33A and 33B RNA amplicon sequencing was used to detect A-formation-I editing in RNAs associated with ABE treatment. Individual data points are presented, with error bars representing the standard deviation (sd) of n = 3 independent biological replicate analyses performed on different dates. Figure 33A A diagram illustrating the A-formation-I editing frequency in the target RNA amplicon of the core ABE8 construct compared to the ABE7 and Cas9(D10A) nickase control groups. Figure 33B A diagram depicting the A-formation-I editing frequency of ABE8 with mutations that have been reported to improve off-target RNA editing in target RNA amplicon.
[0328] Figure 34A and 34BPresenting diagrams and UPHLC chromatographic traces related to editing in SCD CD34+ cells. CD34+ cells from SCD patients were transfected using electroporation with ABE8.8 mRNA and sgRNA (HBG1 / 2, 50 nM). Edited cells differentiated into erythrocytes in vitro. The editing rate at the HBG1 / 2 promoter was measured using next-generation genome sequencing (NGS). Figure 34A As shown, 16.5% of the bases were edited by the ABE8.8 base editor at 48 hours post-differentiation and 89.2% were edited at 14 days post-differentiation. Figure 34B The analysis was presented by bystanders 48 hours and 14 days after the differentiation.
[0329] Figures 35A-35D Present the UPHLC chromatographic trajectory of the globin stage and its dependence on... Figure 34A and 34B The diagram illustrates the functions of HbF upregulation and HbS downregulation in edited SCD CD34+ cells. Edited SCD CD34+ cells differentiate into erythrocytes, and the globin stage is analyzed on day 18 post-differentiation. Figure 35A The presented trajectory shows the globin phase of erythrocytes differentiated from SCDCD34+ cells that have never been edited. Figure 35B The presented trajectory shows the globin stage of erythrocytes differentiated from edited SCD CD34+ cells. Figure 35C The results showed that 63.2% of the γ-globin phase was detected in erythrocytes differentiated from edited SCDCD34+ cells compared to unedited cells. Figure 35D The results showed that, compared to unedited cells, the S-globin content differentiated from edited SCD CD34+ cells decreased from 86% to 32.9%. Upregulation of fetal heme is a potentially beneficial treatment for SCD and beta-thalassemia.
[0330] Figures 36A-36C Display and generate ribbon structures, target sequences, and diagrams related to variants of the ABE editor used for editing atypical Cas9 NGG PAM sequences. Design an ABE base editor comprising modified SpCas9 (including MQKFRAER amino acid substitutions and specificity for the altered PAM 5'-NGC-3' described herein). Figure 36A It can target the sickle-shaped paired gene (“target A”) within the ABE editing window, such as Figure 36B As shown, this provides the ability to directly edit this location at the target site, which is typically inaccessible using conventional spCas9. Figure 36CA diagram illustrating the base editing activity of a variant editor incorporating MQKFRAER amino acid substitutions is presented. This editor can identify target sites and convert nucleobase A to nucleobase T (A·T) to achieve the desired correction Val->Ala. The variants are plotted on the x-axis: "Pro→Pro" represents the leftmost bar; "Val→Ala" represents the middle bar; and "Ser→Pro" represents the rightmost bar.
[0331] Figure 37 A diagram, target site sequence, and table are presented related to the generation of additional adenosine deaminase variants, excluding the linker to TadA and placing it closer to the Cas9 complex. These variants demonstrate enhanced efficacy in editing target sites of sickle-pair genes in a model cell line (HEK293T). The terms "ISLAY" or "IBE" refer to base editors in which TadA adenosine deaminase has been inserted into the Cas9 sequence, such as: ISLAY1 V1015, ISLAY2 I1022, ISLAY3 I1029, ISLAY4 E1040, ISLAY5 E1058, ISLAY6 G1347, ISLAY7 E1054, ISLAY8 E1026, and ISLAY9 Q768, as shown in Table 14A below. On the right side of the figure, the target site, PAM site, and corresponding amino acid sequence in the nucleic acid sequence are shown. In the table, “Cp5” (MSP552) refers to ABE8 in the backbone, which includes a cyclic arrangement of Cas9 and has the following amino acid sequence, which is described below.
[0332]
[0333] The 20nt guide sgRNA (1000ng), spCas9-MQKFRAER, specific for NGC PAM, was used in experiments and applied in triplicate to transfect HEK293T cells (2x10). 5 (cells / well).
[0334] Figure 38 and 39 A diagram showing different ISLAY variants of adenosine deaminase demonstrates enhanced editing of target sites (e.g., Figure 37 (As shown). The diagram shown in the middle figure is a comparison with other ABE editors (ABE7.10) that have connectors with the TadA structure domain.
[0335] Figure 40 The percentage of base editing achieved in CD34+ cells expressing the SCD target site is shown, and tables display the edited nucleic acid and amino acid changes. CD34+ cells from heterozygous sickle-shaped patients were treated with the ABE editor, and the editing of the target site (9G) was measured, i.e., the conversion of nucleobase A to nucleobase T to achieve the desired correction Val>Ala. At 96 hours post-electroporation, the variant ABE editor achieved over 50% editing of the sickle-shaped cell pair gene in CD34+ cells. This level was maintained until cells differentiated into erythrocytes (IVD) in vitro, as over 60% editing was observed in differentiated erythrocytes (heterozygotes of the sickle-shaped trait) at 12 days post-IVD. The graph evaluating the Editor_nM mRNA_[sgRNA]:[mRNA]_Timepoint uses 21nt gRNA.
[0336] Figure 41A and 41B This paper presents the high-performance liquid chromatography (UHPLC) trajectory and LC-MS results related to the detection of unique β-globin species in differentiated erythrocytes from heterozygous HbS (β-globin in sickle cells) cells. Prior to these assays and analyses, operators of the relevant techniques typically could not successfully distinguish and isolate the HbG Makassar globin variant from the HbS sickle globin variant using conventional methods. This paper develops a UHPLC method and uses it to distinguish these two distinct globin variants in cells (e.g., CD34+ cells) from SCD patients that have been edited using the ABE8 editor described herein. After editing, CD34+ cells from heterozygous HbSS samples can be analyzed using UHPLC to detect, based on molecular weight, the different β-globin (Hb) variants corresponding to those with Val→Ala substitution. Figure 41AEditing peaks analyzed by liquid chromatography-mass spectrometry (LC-MS) show a charged envelope indicating a novel, independent β-globin variant (Makassar variant). Figure 41B ).
[0337] Figure 42 A table of base editor and sgRNA sequences for base-edited SCD samples containing HbS globin variants is presented to achieve correction of the HbG Makassar globin variant. The ABE8 mutation was introduced into the lead editor candidate, and sgRNAs of different lengths (21nt, 20nt, 19nt prototypical spacers) were analyzed to detect whether they improved on-target editing while reducing potentially harmful 1G editing (Ser10Pro conversion). The “A” nucleotides, indicated by bold / italic / underline, indicate sickle substitutions. Lowercase letters in the sgRNA / prototypical spacer sequences indicate 2′-O-methylated nucleobases. The lowercase “s” in the sgRNA / prototypical spacer sequences indicates phosphate thioesters.
[0338] Figure 43A and 43B Presenting 9G target sites (or 9G and other sites) in CD34+ cells (heterozygous sickle cell phenotypic samples) 48 hours after electroporation using different ABE editors. Figure 43A ) or differentiated erythrocytes in vitro (heterozygous sickle-shaped trait sample) after 7 days of differentiation ( Figure 43B A bar graph showing the total percentage of edits. While additional mutations did not significantly improve mid-target editing, four editors demonstrated comparable mid-target editing efficacy. A 20nt sgRNA length achieved a lower unwanted 1G bystander editing. In these graphs, editor_sgRNA nt or editor_100nM mRNA_μM sgRNA (20nt) were evaluated. Editing was maintained at approximately 80% throughout in vitro erythrocyte differentiation.
[0339] Figure 44A and 44B Show bar graphs and tables displaying the edited nucleic acid sequences and corresponding amino acid sequence transformations related to total base editing at position 9G of HbS in homozygous SCD (HbSS) samples. Cells were obtained from whole blood (non-mobilized) samples from SCD (HbSS) patients and were base-edited using an ABE variant base editor. Figure 44ACD34+ cells (~200,000 cells, homozygous SCD samples) were electroporated using a 50 nM ABE variant editor (MSP619 (ISLAY5)) at a ratio of 100:1 (2 μg mRNA, 4.1 μg sgRNA (21 nt)). The ABE variant base editor achieved approximately 65% editing at position 9G on day 7 post-electroporation and approximately 60% editing at position 9G on day 14 post-electroporation. Figure 44B CD34+ cells (~200,000 cells, homozygous SCD samples) were electroporated using a 30 nM ABE variant editor (MSP616(ISLAY2)) at a ratio of 200:1 (1.3 μg mRNA, 4.95 μg sgRNA (21 nt)). The ABE variant base editor achieved at least approximately 50% editing at the 9G site on erythrocytes on days 7 and 14 post-electroporation.
[0340] Figure 45 The UHPLC chromatographic trajectory after UHPLC analysis is presented. In homozygous HbSS cells obtained from SCD patient samples, after base editing using the ABE variant base editor, it clearly shows the separation and differentiation of HbS and HbGMakassar variants.
[0341] Figure 46A and 46B Present the UHPLC chromatographic trajectory and LC-MS results related to the detection of independent β-globin species in edited heterozygous HbS (β-globin in sickle cells) differentiated erythrocytes. For example... Figure 41A and 41B The description indicates that UHPLC was used to distinguish these two different globin variants. In the edited heterozygous HbSS samples, the different β-globin (Hb) variants corresponding to each other with Val→Ala substitution could be detected based on molecular weight. Figure 46A The edited peaks of the LC-MS trajectory indicate the charged mantle of the novel β-globin variant. Figure 46B ).
[0342] Figure 47Present the UHPLC chromatographic trajectory and LC-MS results of HbSS (SCD) samples with and without base editing (“HbSS – Edited”) or without base editing (“HbSS – Unedited”). As shown in the UHPLC chromatograms in the top and middle figures, based on the difference in dissolution time in UHPLC, the HbG Makassar globin variant separated from the HbS (SCD) globin type at 9.81 minutes (10.03 minutes). Other globin types are easily distinguishable. The LC-MS figure below shows that the Makassar HbG and HbS globin types have different and distinguishable characteristics. Similarly... Figure 41A , 41B The results presented in 45, 46A, and 46B, obtained from UHPLC and LC-MS analyses of SCD (HbSS) erythrocyte samples edited by the ABE variant base editor described in this paper, clearly distinguish and separate the HbG Makassar variant and the HbS (SCD) globin variant in the samples. Therefore, they provide a useful tool for identifying true SCD (HbS) patients and reducing or preventing misdiagnosis of patients who present with the HbG Makassar globin variant SCD (HbSS).
[0343] Figures 48A-48C A bar graph representing the relative area under the peak from UHPLC chromatographic data is presented. The area under the peak was used to quantify the total variation in the content of different β-globin variants in homozygous SCD samples that had undergone base editing using the ABE variant of this invention (base editor MSP619, 50 nM mRNA, 5000 nM sgRNA (21 nt)). The results presented suggest a positive correlation between the degree of HbS variant globin conversion and asymptomatic HbG-Makassar globin.
[0344] Figure 49 A table is provided to depict all possible PAMs accessible within the NRNN PAM region. Only Cas9 variants requiring identification of three or fewer specified nucleotides in their PAMs are listed. Non-G PAM variants include SpCas9-NRRH, SpCas9-NRTH, and SpCas9-NRCH. (Miller, SM et al., Continuous evolution of SpCas9 variants compatible with non-G PAMs, Nat. Biotechnol. (2020), ( / / doi.org / 10.1038 / s41587-020-0412-8), which are incorporated herein by reference in full. Detailed Implementation
[0345] As described below, the present invention is characterized by compositions and methods for modifying sickle cell disease (SCD)-related mutations. In some embodiments, the editing corrects harmful mutations such that the edited polynucleotide is distinguishable from a wild-type reference polynucleotide sequence. In another embodiment, the editing corrects harmful mutations such that the edited polynucleotide contains benign mutations.
[0346] HBB gene editing
[0347] As described herein, the compositions and methods of the present invention are suitable and advantageous for the treatment of sickle cell disease (SCD), caused by a Glu→Val mutation in the sixth amino acid of the β-globin protein encoded by the HBB gene. Despite significant advancements in gene editing, it remains impossible to precisely correct the diseased HBB gene back to Val→Glu, and this cannot currently be achieved using CRISPR / Cas nucleases or CRISPR / Cas base editing methods.
[0348] Genome editing using CRISPR / Cas nuclease methods to replace affected nucleotides in the HBB gene requires cleavage of the genome DNA. However, cleavage of the genome DNA increases the risk of base insertions / deletions, potentially leading to unintended and undesirable results, including premature formation of stop codons, alteration of codon reading frames, and so on. Furthermore, double-strand breaks at the β-globin locus have the potential to fundamentally modify the locus through recombination events. The β-globin locus contains a cluster of globin genes with sequence identity to another -5'-ε-; Gγ-; Aγ-; δ-; and β-globin-3'. Due to the structure of the β-globin locus, recombination repair of double-strand breaks within the locus can potentially result in the loss of sequences interspersed between globin genes, such as between δ- and β-globin genes.
[0349] Unintended modifications to gene loci also carry the risk of causing thalassemia. CRISPR / Cas base editing methods can ensure precise modifications at the nucleobase stage. However, precise correction of Val→Glu(G...)... T G→G AG) A T·A to A·T inversion editor is required, which is currently unknown. Furthermore, the specificity of CRISPR / Cas base editing is partly due to the limited window of editable nucleotides created by the R-loops formed when CRISPR / Cas binds to DNA. Therefore, CRISPR / Cas targeting must occur at or near the sickle cell site to enable base editing, and optimal editing within the window may require additional sequences. One requirement for CRISPR / Cas targeting is the presence of a protospacer adjacent motif (PAM) flanking the targeted site. For example, many base editors are based on SpCas9, which requires an NGGPAM. Even if it is hypothesized that T·A can be inverted to A·T, this SpCas9 base editor lacks an NGG PAM that would place the target "A" in the desired position. Although many new CRISPR / Cas proteins have been discovered or developed that expand the collection of available PAMs, the requirement for PAMs remains a limiting factor in the ability of CRISPR / Cas base editors to target specific nucleotides at any location within the genome.
[0350] This invention is at least in part based on several findings described herein that address the aforementioned challenges and provide a gene editing method for treating sickle cell anemia. In one embodiment, the invention is partly based on the ability to replace amino acid position 6 with alanine, which causes sickle cell disease, thereby producing an Hb variant (HbMakassar) that does not produce the sickle cell phenotype. While precise correction is impossible without a T·A to A·T inversion base editor (G... T G→G A G). However, the experiments conducted in this paper have found that using the A·T to form G·C base editor (ABE) can produce Val→Ala(G). T G→G C G) substitution (i.e., Hb Makassar variant). This is partly achieved through the development of novel base editors and novel base editing strategies presented in this paper. For example, a novel ABE base editor (i.e., with an adenosine deaminase domain) utilizes side sequences (e.g., PAM sequences; zinc finger binding sequences) to achieve optimal base editing at sickle cell target sites.
[0351] Therefore, the present invention includes compositions and methods for base editing thymidine (T) to cytidine (C) at the sixth amino acid codon of the sickle cell variant of β-globin protein (sickle HbS; E6V), thereby replacing valine (V6A) at this amino acid position with alanine. Replacing valine at position 6 of HbS with alanine produces a β-globin protein variant without the sickle cell phenotype (e.g., without the possibility of polymerization as in the pathogenic variant HbS). Therefore, the compositions and methods of the present invention are suitable for treating sickle cell disease (SCD).
[0352] Nucleotide base editor
[0353] This document discloses a base editor or nucleobase editor for editing, modifying, or altering target nucleotide sequences of polynucleotides (e.g., HBB polynucleotides). This document describes a nucleobase editor or base editor comprising a polynucleotide programmable nucleotide-binding domain and a nucleobase editing domain (e.g., adenosine deaminase). The polynucleotide programmable nucleotide-binding domain, when conjugated with a bound guide polynucleotide (e.g., gRNA), can specifically bind to the target polynucleotide sequence (i.e., via complementary base pairing between the bases of the bound guide nucleic acid and the target polynucleotide sequence), thereby allowing the base editor to lock onto the target nucleic acid sequence to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid.
[0354] Multinucleotide programmable nucleotide binding domain
[0355] It should be understood that the polynucleotide programmable nucleotide binding domain may also include a nucleic acid programmable protein that binds to RNA. For example, the polynucleotide programmable nucleotide binding domain may be associated with a nucleic acid that guides the polynucleotide programmable nucleotide binding domain to RNA. Other nucleic acid programmable DNA-binding proteins are also within the scope of this invention, but are not explicitly listed herein.
[0356] The polynucleotide-programmable nucleotide-binding domain of the base editor may itself contain one or more domains. For example, the polynucleotide-programmable nucleotide-binding domain may contain one or more nuclease domains. In some embodiments, the nuclease domain of the polynucleotide-programmable nucleotide-binding domain may contain an endonuclease or an exonuclease. The term "exonuclease" refers to a protein or polypeptide that can digest nucleic acids (e.g., RNA or DNA) from a free terminal, and the term "endonuclease" refers to a protein or polypeptide that can catalyze (e.g., cleave) an internal region of a nucleic acid (e.g., DNA or RNA). In some embodiments, the endonuclease can cleave a single strand of a double-stranded nucleic acid. In some embodiments, the endonuclease can cleave both strands of a double-stranded nucleic acid molecule. In some embodiments, the polynucleotide-programmable nucleotide-binding domain may be a deoxyribonuclease. In some embodiments, the polynucleotide-programmable nucleotide-binding domain may be a ribonuclease.
[0357] In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide binding domain can cleave the zero, one, or both strands of the target polynucleotide. In some embodiments, the polynucleotide programmable nucleotide binding domain may include a nicking enzyme domain. The term "nicking enzyme" herein refers to a polynucleotide programmable nucleotide binding domain containing a nuclease domain that can cleave only one strand of a double-stranded nucleic acid molecule (e.g., DNA). In some embodiments, the nicking enzyme can be introduced into the active polynucleotide programmable nucleotide binding domain through one or more mutations, resulting in the complete catalytic activity (e.g., native) type derived from the polynucleotide programmable nucleotide binding domain. For example, if the polynucleotide programmable nucleotide binding domain includes a Cas9-derived nicking enzyme domain, the Cas9-derived nicking enzyme domain may include a D10A mutation and a histidine residue at position 840. In these examples, residue H840 retains catalytic activity and thereby cleaves the single strand of the nucleic acid double helix. In another example, the Cas9-derived nicking enzyme domain may include an H840A mutation, while the amino acid residue at position 10 remains D. In some embodiments, the nicking enzyme can be derived from a fully catalytically active (e.g., native) type derived from a polynucleotide programmable nucleotide-binding domain by removing all or part of the nuclease domains that are not required for nicking enzyme activity. For example, if the polynucleotide programmable nucleotide-binding domain includes a nicking enzyme domain derived from Cas9, the Cas9-derived nicking enzyme domain may include a domain lacking all or part of the RuvC domain or the HNH domain.
[0358] The amino acid sequences of catalytically active Cas9 examples are as follows:
[0359]
[0360] A base editor containing a polynucleotide programmable nucleotide-binding domain with a nicking enzyme domain can therefore create single-strand DNA breaks (nicks) on specific polynucleotide target sequences (e.g., determined by the complementary sequence of the bound guide nucleic acid). In some embodiments, one strand of the nucleic acid double helix target polynucleotide sequence cleaved by a base editor containing a nicking enzyme domain (e.g., a Cas9-derived nicking enzyme domain) is the strand that will not be edited by the base editor (i.e., the strand cleaved by the base editor is the opposite of the strand containing the base to be edited). In other embodiments, a base editor containing a nicking enzyme domain (e.g., a Cas9-derived nicking enzyme domain) can cleave the targeted strand of the DNA molecule. In this case, the untargeted strand will not be cleaved.
[0361] This document also provides a base editor comprising a catalytically inactivated (i.e., incapable of cleaving a target polynucleotide sequence) polynucleotide programmable nucleotide-binding domain. The terms "catalytically inactivated" and "nuclease inactivated" are used interchangeably herein, referring to a polynucleotide programmable nucleotide-binding domain having one or more mutations and / or deletions that render it incapable of cleaving one strand of a nucleic acid. In some embodiments, the catalytically inactivated polynucleotide programmable nucleotide-binding domain base editor lacks nuclease activity due to specific point mutations in one or more nuclease domains. For example, in a base editor comprising a Cas9 domain, Cas9 may contain both a D10A mutation and an H840A mutation. These mutations inactivate both nuclease domains, resulting in loss of nuclease activity. In other embodiments, the catalytically inactivated polynucleotide programmable nucleotide-binding domain may comprise one or more deletions of all or part of the catalytic domain (e.g., the RuvC1 and / or HNH domains). In other embodiments, the catalytically inactivated polynucleotide programmable nucleotide-binding domain comprises point mutations (e.g., D10A or H840A) and deletions of all or part of the nuclease domain.
[0362] This document also considers mutations that can generate catalytically inactivating polynucleotide programmable nucleotide-binding domains from previously defined functional polynucleotide programmable nucleotide-binding domains. For example, using catalytically inactivating Cas9 (“dCas9”) as an example, variants with mutations other than D10A and H840A are provided, resulting in nuclease inactivation of the Cas9 nuclease. Such mutations include, for example, substitutions of other amino acids in D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). Other suitable nuclease-inactivating dCas9 domains are those skilled in the art and are within the scope of this invention. Other suitable examples of nuclease-inactivated Cas9 domains include (but are not limited to): D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, for example, Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013; 31(9):833-838, the full contents of which are incorporated herein by reference).
[0363] Non-limiting examples of multinucleotide programmable nucleotide-binding domains that can be incorporated into a base editor include CRISPR protein-derived domains, restriction nucleases, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases (ZFNs). In some implementations, the base editor includes a polynucleotide-programmable nucleotide-binding domain comprising a natural or modified protein or a portion thereof, which, through a bound guide nucleic acid, can bind to the nucleic acid sequence during CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) mediated modification of the nucleic acid. Such proteins are referred to herein as “CRISPR proteins.” Therefore, this document discloses a base editor comprising a polynucleotide-programmable nucleotide-binding domain comprising all or a portion of a CRISPR protein (i.e., the base editor comprises all or a portion of a CRISPR protein as a domain, also referred to as a “CRISPR protein-derived domain” of the base editor). The CRISPR protein-derived domain incorporated in the base editor may be modified relative to a wild-type or natural CRISPR protein. For example, the CRISPR protein-derived domain described below may contain one or more mutations, insertions, deletions, rearrangements, and / or recombinations relative to a wild-type or natural CRISPR protein.
[0364] CRISPR acts as an acquired immune system, providing protection against mobile genetic agents (viruses, transposons, and conjugating plasmids). A CRISPR cluster contains a spacer, a sequence complementary to the aforementioned mobile agents, and a target invasive nucleic acid. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In the type II CRISPR system, the modification of pre-crRNA requires a trans-encoded small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. tracrRNA acts as a guide for ribonuclease 3 in assisting pre-crRNA processing. Subsequently, Cas9 / crRNA / tracrRNA cleaves the linear or circular dsDN target complementary to the spacer via endonucleolysis. The target strand not complementary to crRNA is first cleaved via endonucleolysis, followed by exonucleolysis to trim the 3′–5′ segments. In nature, DNA binding and cleavage typically require proteins and two types of RNA. However, single guide RNAs (“sgRNAs” or “gNRAs”) can be engineered to incorporate both crRNA and tracrRNA embodiments into a single RNA species. See, for example: Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the full contents of which are incorporated herein by reference. Cas9 identifies short motifs (PAMs or protospacer adjacent motifs) in CRISPR repeat sequences to help distinguish between self and non-self motifs.
[0365] In some embodiments, the methods described herein can utilize engineered Cas proteins. Guide RNA (gRNA) is a short synthetic RNA consisting of a backbone sequence necessary for Cas-binding and a user-defined spacer of approximately 20 nucleotides, which specifies the target genome to be modified. Therefore, those skilled in the art can modify the specific genome target of the Cas protein, which is partly determined by the specificity of the gRNA targeting sequence to that target genome compared to other genomes.
[0366] In some implementations, the gRNA backbone sequence is as follows: GUUUUAGAGC UAGAAAUAGC AAGUUAAAAUAAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU.
[0367] In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is an endonuclease (e.g., deoxyribonuclease or ribonuclease) that, when linked to a bound guide nucleic acid, can bind the target polynucleotide. In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is a nicking enzyme that, when linked to a bound guide nucleic acid, can bind the target polynucleotide. In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is a catalytic inactivation domain that, when linked to a bound guide nucleic acid, can bind the target polynucleotide. In some embodiments, the target polynucleotide bound by the CRISPR protein-derived domain of the base editor is DNA. In some embodiments, the target polynucleotide bound by the CRISPR protein-derived domain of the base editor is RNA.
[0368] The Cas proteins that can be used in this paper include those of classes 1 and 2. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, and Cmr4. Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i, CARF, DinG, their homologs, or their modifications. Unmodified CRISPR enzymes can possess DNA cleavage activity, such as Cas9, which has two functional endonuclease domains: RuvC and HNH. CRISPR enzymes can predominantly cleave target sequences, such as one or both strands within the target sequence and / or within its complement. For example, a CRISPR enzyme can predominantly cleave one or both strands of a target sequence starting from the first or last nucleotide, approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs.
[0369] A vector encoding a CRISPR enzyme that has been mutated relative to the corresponding wild-type enzyme can be used, so that the mutated CRISPR enzyme lacks the ability to cleave one or both strands of the target polynucleotide containing the target sequence. Cas9 can refer to peptides that have at least or at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology with wild-type Cas9 peptide examples (e.g., Cas9 from Streptococcus pyogenes). Cas9 can refer to and wild-type Cas9 polypeptide instances (e.g., from *Streptococcus pyogenes*) having up to or up to about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology. Cas9 can also refer to wild-type or modified Cas9 proteins, which may include amino acid alterations such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof.
[0370] In some implementations, the CRISPR protein-derived domain of the base editor may include all or part of Cas9 from the following: Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense, China (NCBI Ref: NC_021846.1); Streptococcus iniae (NCBI... Ref:NC_021314.1); Bellella baltica (NCBI Ref:NC_018010.1); Psychroflexus torquis (NCBI Ref:NC_018721.1); Streptococcus thermophilus (NCBI Ref:YP_820832.1); Listeria innocua (NCBI Ref:NP_472073.1); Campylobacter jejuni (NCBI Ref:YP_002344900.1); Neisseria meningitidis (NCBI Ref:YP_002342100.1), Streptococcus pyogenes, or Staphylococcus aureus.
[0371] Cas9 domain of the nucleobase editor
[0372] The Cas9 nuclease sequence and structure are well known to those skilled in the art (see, for example: "Completegenome sequence of an Ml strain of Streptococcus pyogenes." Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chylinski). K., Sharma C.M., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012), the full contents of which are incorporated herein by reference. Cas9 is a homologous gene that has been described in various species, including (but not limited to): Streptococcus pyogenes and Streptococcus thermophilus.Other suitable Cas9 nucleases and sequences are known to those skilled in the art based on this invention, and such Cas9 nucleases and sequences include Cas9 sequences of organisms and loci disclosed by Chylinski, Rhun, and Charpentier in “The tracrRNA and Cas9families of type IICRISPR-Cas immunitysystems” (2013) RNA Biology 10:5,726-737, the full contents of which are incorporated herein by reference.
[0373] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) is a Cas9 domain. This document provides non-limiting examples of Cas9 domains. The Cas9 domain may be a nuclease-active Cas9 domain, a nuclease-inactivating Cas9 domain (dCas9), or a Cas9 cleavage enzyme (nCas9). In some embodiments, the Cas9 domain is a nuclease-active domain. For example, the Cas9 domain may be a Cas9 domain capable of cleaving both strands of a double-stranded nucleic acid (e.g., both strands of a double-stranded DNA molecule). In some embodiments, the Cas9 domain contains any of the amino acid sequences shown herein. In some embodiments, the Cas9 domain contains an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the amino acid sequences shown herein. In some implementations, the amino acid sequence contained in the Cas9 domain has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the amino acid sequences shown herein. In some embodiments, the Cas9 domain contains an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical consecutive amino acid residues compared to any of the amino acid sequences shown herein.
[0374] In some embodiments, a protein containing a fragment of Cas9 is provided. For example, in some embodiments, the protein contains one of two Cas9 domains: (1) a gRNA-binding domain of Cas9; or (2) a DNA-cleaving domain of Cas9. In some embodiments, a protein containing Cas9 or a fragment thereof is referred to as a "Cas9 variant". Cas9 variants and Cas9 or fragments thereof share common homology. For example, Cas9 variants and wild-type Cas9 are at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9%. In some implementations, the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type Cas9. In some embodiments, the Cas9 variant contains a fragment of Cas9 (e.g., a gRNA-binding domain or a DNA-cleavage domain), and therefore the fragment is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to its wild-type Cas9 counterpart. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of the wild-type Cas9. In some embodiments, the fragment is at least 100 amino acids long. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids in length.
[0375] In some embodiments, the Cas9 fusion protein provided herein contains the full-length amino acid sequence of the Cas9 protein, such as one of the Cas9 sequences provided herein. However, in other embodiments, the fusion protein provided herein does not contain the full-length Cas9 sequence, but only one or more fragments thereof. Examples of suitable Cas9 domains and Cas9 fragments are provided herein, and those skilled in the art will generally understand the sequences of other suitable Cas9 domains and fragments.
[0376] The Cas9 protein can bind to a guide RNA, which guides the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the polynucleotide programmable nucleotide-binding domain is a Cas9 domain, such as: active Cas9, Cas9 nickase (nCas9), or inactive Cas9 (dCas9). Examples of nucleic acid programmable DNA-binding proteins include (but are not limited to): Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpf1, Cas12b / C2C1, and Cas12c / C2C3.
[0377] In some implementations, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_017053.1, nucleotide and amino acid sequences are as follows):
[0378]
[0379]
[0380] (Single bottom line: HNH domain; Double bottom line: RuvC domain)
[0381] In some implementations, wild-type Cas9 corresponds to or contains the following nucleotide and / or amino acid sequences:
[0382]
[0383]
[0384]
[0385] (Single bottom line: HNH domain; Double bottom line: RuvC domain)
[0386] In some implementations, wild-type Cas9 corresponds to Cas9 of Streptococcus pyogenes (NCBI reference sequence: NC_002737.2 (nucleotide sequence below); and Uniprot reference sequence: Q99ZW2 (amino acid sequence below):
[0387]
[0388]
[0389]
[0390] (Single bottom line: HNH domain; Double bottom line: RuvC domain)
[0391] In some implementations, Cas9 refers to Cas9 derived from the following: Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense, China (NCBI Ref: NC_021846.1); and Streptococcus iniae (NCBI Ref: NC_021846.1). Ref:NC_021314.1); Belliella baltica (NCBI Ref:NC_018010.1); Psychroflexus torquis I (NCBI Ref:NC_018721.1); Streptococcus thermophilus (NCBI Ref:YP_820832.1), Listeria innocua (NCBI Ref:NP_472073.1), Campylobacter jejuni (NCBI Ref:YP_002344900.1), or Neisseria meningitidis (NCBI Ref:YP_002342100.1) or Cas9 from any other organism.
[0392] It should be understood that additional Cas9 proteins (e.g., nuclease-inactivated Cas9 (dCas9), Cas9 cleavage enzyme (nCas9), or nuclease-active Cas9), including their variants and homologs, are within the scope of this invention. Examples of Cas9 proteins include (but are not limited to): those provided below. In some embodiments, the Cas9 protein is nuclease-inactivated Cas9 (dCas9). In some embodiments, the Cas9 protein is Cas9 cleavage enzyme (nCas9). In some embodiments, the Cas9 protein is nuclease-active Cas9.
[0393] In some embodiments, the Cas9 domain is a nuclease-inactivating Cas9 domain (dCas9). For example, the dCas9 domain may bind to a double-stranded nucleic acid molecule (e.g., via a gRNA molecule) without cleaving either strand of the double-stranded nucleic acid molecule. In some embodiments, the nuclease-inactivating dCas9 domain contains the D10X and H840X mutations of the amino acid sequence shown herein, or the corresponding mutations of any amino acid sequence provided herein, where X represents any amino acid change. In some embodiments, the nuclease-inactivating dCas9 domain contains the D10A and H840A mutations of the amino acid sequence shown herein, or the corresponding mutations of any amino acid sequence provided herein. In one example, the nuclease-inactivating Cas9 domain contains the amino acid sequence shown in the selection vector pPlatTET-gRNA2 (accession number BAV54124).
[0394]
[0395] (See, for example, Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression.” Cell. 2013; 152(5):1173-83, the full text of which is incorporated herein by reference).
[0396] Other suitable nuclease-inactivating dCas9 domains are those skilled in the art and are within the scope of this invention. Examples of such other suitable nuclease-inactivating Cas9 domains include (but are not limited to): D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, for example: Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013; 31(9):833-838, the full contents of which are incorporated herein by reference).
[0397] In some implementations, the Cas9 nuclease has an inactivating (e.g., inactivated) DNA cleavage domain, meaning Cas9 is a nicking enzyme, referred to as the “nCas9” protein (in contrast to the “nicking enzyme” Cas9). The nuclease-inactivated Cas9 protein can be exchanged with the “dCas9” protein (in contrast to the nuclease-“inactivated” Cas9) or the catalytically inactivated Cas9. Methods for generating Cas9 proteins (or fragments thereof) with an inactivating DNA cleavage domain are known (see, for example, Jinek et al., Science. 337:816-821 (2012); Qi et al., “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression” (2013) Cell. 28; 152(5):1173-83, etc., which are incorporated herein by reference in full). For example, the DNA cleavage domain of Cas9 is known to include two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the complementary strand of the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, mutations in D10A and H840A completely inactivate the nuclease activity of Cas9 from Streptococcus pyogenes (Jinek et al., Science. 337:816-821 (2012); Qi et al., Cell. 28; 152(5):1173-83 (2013)).
[0398] In some embodiments, the dCas9 domain contains an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the dCas9 domains provided herein. In some implementations, the amino acid sequence contained in the Cas9 domain has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the amino acid sequences shown herein. In some embodiments, the Cas9 domain contains an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical consecutive amino acid residues compared to any of the amino acid sequences shown herein.
[0399] In some embodiments, dCas9 corresponds to or contains a portion or all of the Cas9 amino acid sequence having one or more mutations that inactivate Cas9 nuclease activity. For example, in some embodiments, the dCas9 domain contains D10A and H840A mutations or corresponding mutations in another Cas9.
[0400] In some implementations, dCas9 contains the amino acid sequence of dCas9 (D10A and H840A):
[0401]
[0402]
[0403] (Single bottom line: HNH domain; Double bottom line: RuvC domain).
[0404] In some implementations, the Cas9 domain contains the D10A mutation, while the residue at position 840 retains a histidine at the corresponding position in the amino acid sequence provided above, or in any amino acid sequence provided herein.
[0405] In other embodiments, dCas9 variants with mutations other than D10A and H840A are provided, for example, Cas9 that causes nuclease inactivation (dCas9). Such mutations include, for example, other amino acid substitutions in D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). In some embodiments, the provided dCas9 variants or homologs are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical. In some implementations, the provided dCas9 variants have shorter or longer amino acid sequences, differing by approximately 5 amino acids, approximately 10 amino acids, approximately 15 amino acids, approximately 20 amino acids, approximately 25 amino acids, approximately 30 amino acids, approximately 40 amino acids, approximately 50 amino acids, approximately 75 amino acids, or approximately 100 or more amino acids.
[0406] In some embodiments, the Cas9 domain is a Cas9 cleavage enzyme. The Cas9 cleavage enzyme may be a Cas9 protein capable of cleaving only one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 cleavage enzyme cleaves the target strand of the double-stranded nucleic acid molecule, meaning the strand cleaved by the Cas9 cleavage enzyme is base-paired (complementary) to gRNA (e.g., sgRNA) already bound to Cas9. In some embodiments, the Cas9 cleavage enzyme contains a D10A mutation and has a histidine residue at position 840. In some embodiments, the Cas9 cleavage enzyme cleaves the non-target, non-base-editing strand of the double-stranded nucleic acid molecule, meaning the strand cleaved by the Cas9 cleavage enzyme is not base-paired to gRNA (e.g., sgRNA) already bound to Cas9. In some embodiments, the Cas9 cleavage enzyme contains an H840A mutation and has an aspartic acid residue at position 10, or a corresponding mutation. In some embodiments, the Cas9 nickase comprises an amino acid sequence that is identical to any of the Cas9 nickases provided herein in at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%. Other suitable Cas9 nickases are those skilled in the art based on the present invention and are within the scope of this invention.
[0407] The amino acid sequence of an example of a catalytic Cas9 nickase (nCas9) is as follows:
[0408]
[0409] In some implementations, Cas9 refers to Cas9 from archaea (e.g., nanoarchaea), which belong to the domain and kingdom of single-celled prokaryotes. In some implementations, the programmable nucleotide-binding protein may be a CasX or CasY protein, as described, for example, in Burstein et al., “New CRISPR-Cas systems from uncultivated microbes.” Cell Res. 2017 Feb 21. doi:10.1038 / cr.2017.21, the full contents of which are incorporated herein by reference. Many CRISPR-Cas systems are identified using genome-resolved metagenomics, including Cas9, which was first reported in the archaea domain of life. This diverse Cas9 protein has appeared in the little-studied nanoarchaea as part of an active CRISPR-Cas system. Two previously unknown systems, CRISPR-CasX and CRISPR-CasY, have been discovered in bacteria, and they are among the most robust systems discovered to date. In some embodiments, in the base editor system described herein, Cas9 is substituted with CasX, or a variant of CasX. In some embodiments, in the base editor system described herein, Cas9 is substituted with CasY, or a variant of CasY. It should be understood that other RNA-guided DNA-binding proteins can be used as nucleic acid programmable DNA-binding proteins (napDNAbp), and are within the scope of this invention.
[0410] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) of any fusion protein provided herein may be a CasX or CasY protein. In some embodiments, napDNAbp is a CasX protein. In some embodiments, napDNAbp is a CasY protein. In some embodiments, napDNAbp contains an amino acid sequence that is identical to that of a native CasX or CasY protein in at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%. In some embodiments, the programmable nucleotide-binding protein is a native CasX or CasY protein. In some embodiments, the programmable nucleotide-binding protein comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any CasX or CasY protein described herein. It should be understood that CasX and CasY from other bacterial species may also be used according to the invention.
[0411] An example of the CasX ((uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53)tr|F0NN87|F0NN87_SULIHCRISPR-associatedCasx protein OS=Sulfolobus islandicus (strain HVE10 / 4)GN=SiH_0402PE=4SV=1) amino acid sequence is as follows:
[0412] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAVGQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGE。
[0413] An exemplary CasX (>tr|F0NH53|F0NH53_SULIR CRISPR - associated protein, Casx OS = Sulfolobus islandicus (strain REY15A) GN = SiRe_0771 PE = 4 SV = 1) amino acid sequence is as follows:
[0414] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAVGQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGE。
[0415] Deltaproteobacteria CasX
[0416] MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMTDDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAILQVYWQEFKDDHVGLMCKFAQPASKKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFVYKLEQVSEKGKAYTNYFGRCNVAEHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESTHPVKPLAQIAGNRYASGPVGKALSDACMGTIASFLSKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDfAYNEVIARVRMWVNLLWQKLKLSRDDAKPLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENPKKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEMDEKEFYACEIQLQ KWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFTDGTDIKKSGKWQGLLYGGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKIGRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQAAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAKLAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKELSAELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSYKSGKQPFVGAWQAFYKRRLKEVWKPNA
[0417] An example of the amino acid sequence of CasY ((ncbi.nlm.nih.gov / protein / APG80656.1)>APG80656.1 CRISPR-associated protein CasY [uncultured Parcubacteria bacteria]) is as follows:
[0418]
[0419] The Cas9 nuclease has two functional endonuclease domains: RuvC and HNH. When Cas9 binds to a target, it undergoes a conformational change, allowing the nuclease domains to cleave the opposite strand of the target DNA. The final result of Cas9-mediated DNA cleavage is a double-strand break (DSB) within the target DNA (3-4 nucleotides upstream of the PAM sequence). The resulting DSB is then repaired by one of two common repair pathways: (1) the efficient but error-prone non-homologous end joining (NHEJ) pathway; or (2) the less efficient but more faithful homologous gene-guided repair (HDR) pathway.
[0420] The “efficiency” of non-homologous end joining (NHEJ) and / or homologous gene guided repair (HDR) can be calculated using any suitable method. For example, in some implementations, efficiency is expressed as a percentage of successful HDR. For example, surveyor nuclease analysis can be used to generate cleavage products, and the percentage can be calculated using the ratio of products to plasmids. For example, surveyor nuclease analysis can be used to directly cleave DNA containing newly integrated restriction enzyme sequences resulting from successful HDR. The more plasmids are cleaved, the higher the percentage of HDR (the higher the efficiency of HDR). In one illustrative example, the percentage of HDR can be calculated using the following formula: [(cleavage products) / (platinum plus cleavage products)] (e.g., (b+c) / (a+b+c), where “a” is the band density of the DNA plasmid, and “b” and “c” are the cleavage products).
[0421] In some implementations, efficacy can be expressed as a percentage of successful NHEJs. For example, T7 endonuclease I assay can be used to generate lysis products, and the percentage of NHEJs can be calculated using the product-to-substrate ratio. T7 endonuclease I cleaves mismatched double-stranded DNA resulting from hybridization of wild-type and mutant DNA strands (NHEJs produce small random insertions or deletions at the original break sites). More cleavages indicate a higher percentage of NHEJs (higher NHEJ efficacy). In one example, the percentage of NHEJs can be calculated using the following formula: (1 - (1 - (b + c) / (a + b + c))) 1 / 2 )×100, where “a” is the density of DNA plasmid bands, and “b” and “c” are cleavage products (Ran et al., Cell. 2013 Sep. 12; 154(6):1380-9; and Ran et al., Nat Protoc. 2013 Nov.; 8(11):2281–2308).
[0422] The NHEJ repair pathway is the most active repair mechanism, frequently inducing small nucleotide insertions or deletions at the DSB site. The randomness of NHEJ-mediated DSB repair has significant practical implications because cell populations expressing Cas9 and gRNA or guide polynucleotides can cause divergent mutations. In most cases, NHEJ produces small insertions or deletions in the target DNA, resulting in amino acid deletions, insertions, or frameshift mutations, leading to premature stop codon formation within the target gene's open reading frame (ORF). The ideal end result is a loss-of-function mutation within the target gene.
[0423] Although NHEJ-mediated DSB repair often disrupts open reading frames of genes, homology-guided repair (HDR) can be used to generate specific nucleotide alterations, ranging from single nucleotide changes to large blocks, such as the addition of fluorescent groups or tags. To utilize HDR for gene editing, a DNA repair template containing the desired sequence is delivered to a cell type equipped with gRNA(s) and Cas9 or Cas9 nickase. The repair template may contain the desired edit and additional homologous sequences immediately upstream and downstream of the target (called the left and right homologous arms). The length of each homologous arm depends on the size of the alteration to be introduced; larger blocks require longer homologous arms. The repair template can be a single-stranded oligonucleotide, a double-stranded oligonucleotide, or a double-stranded DNA plasmid. Even in cells expressing Cas9, gRNA, and extrinsic repair templates, the potency of HDR is generally low (<10% of modified paired genes). HDR potency can be enhanced by cell synchronization, as HDR occurs during the S and G2 phases of the cell cycle. Chemical or genetic repressor genes involved in NHEJ can also increase HDR frequency.
[0424] In some implementations, Cas9 is a modified Cas9. The designated gRNA target sequence may have additional sites throughout the genome with partial homology. These sites are called off-target sites and must be considered when designing the gRNA. In addition to optimizing gRNA design, CRISPR specificity can be improved by modifying Cas9. Cas9 generates double-strand breaks (DSBs) by combining the activities of two nuclease domains, RuvC and HNH. The Cas9 nickase (a D10A mutant of SpCas9) retains one nuclease domain and generates DNA nicks instead of DSBs. The nickase system can also be combined with HDR-mediated gene editing for specific gene editing.
[0425] In some embodiments, Cas9 is a variant Cas9 protein. The amino acid sequence of the variant Cas9 polypeptide differs from that of the wild-type Cas9 protein by one amino acid (e.g., deletion, insertion, substitution, or fusion). In some examples, the amino acid alterations (e.g., deletion, insertion, or substitution) of the variant Cas9 polypeptide reduce the nuclease activity of the Cas9 polypeptide. For example, in some examples, the variant Cas9 polypeptide has less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nuclease activity of the corresponding wild-type Cas9 protein. In some embodiments, the variant Cas9 protein has no substantial nuclease activity. When an individual Cas9 protein is a variant Cas9 protein without substantial nuclease activity, it may be referred to as "dCas9".
[0426] In some implementations, the variant Cas9 protein has reduced nuclease activity. For example, the variant Cas9 protein has less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% of the endonuclease activity of the wild-type Cas9 protein (e.g., the wild-type Cas9 protein).
[0427] In some embodiments, the variant Cas9 protein can cleave the complementary strand of the target sequence, but its ability to cleave the non-complementary strand of the double-stranded target sequence is reduced. For example, the variant Cas9 protein may have a mutation (amino acid substitution) that reduces the function of the RuvC domain. As a non-limiting example, in some embodiments, the variant Cas9 protein has D10A (position 10 aspartic acid becomes alanine amino acid), thus it can cleave the complementary strand of the double-stranded target sequence, but its ability to cleave the non-complementary strand of the double-stranded target sequence is reduced (therefore, when the variant Cas9 protein cleaves the double-stranded target nucleic acid, it causes a single-strand break (SSB) instead of a double-strand break (DSB)) (see, for example: Jinek et al., Science. 2012 Aug. 17; 337(6096):816-21).
[0428] In some embodiments, the variant Cas9 protein can cleave the non-complementary strand of the double-stranded target sequence, but its ability to cleave the complementary strand of the target sequence is reduced. For example, the variant Cas9 protein may have a mutation (amino acid substitution) that reduces the function of the HNH domain (RuvC / HNH / RuvC domain motif). As a non-limiting example, in some embodiments, the variant Cas9 protein has an H840A mutation (histidine at position 840 becomes alanine amino acid), thus enabling it to cleave the non-complementary strand of the target sequence, but reducing its ability to cleave the complementary strand of the target sequence (therefore, when the variant Cas9 protein cleaves the double-stranded target sequence, it produces an SSB instead of a DSB). Such Cas9 proteins have reduced ability to cleave target sequences (e.g., single-stranded target sequences), but retain the ability to bind to target sequences (e.g., single-stranded target sequences).
[0429] In some embodiments, the ability of the variant Cas9 protein to cleave both the complementary and non-complementary strands of the double-stranded target DNA is reduced. As a non-limiting example, in some embodiments, the variant Cas9 protein carries both D10A and H840A mutations, thus reducing the ability of the polypeptide to cleave both the complementary and non-complementary strands of the double-stranded target DNA. While the ability of these Cas9 proteins to cleave target DNA (e.g., single-stranded target DNA) is reduced, their ability to bind to target DNA (e.g., single-stranded target DNA) is retained.
[0430] As another non-limiting example, in some embodiments, the variant Cas9 protein carries mutations of W476A and W1126A, thus reducing the ability of the polypeptide to cleave target DNA. While the ability of these Cas9 proteins to cleave target DNA (e.g., single-stranded target DNA) is reduced, their ability to bind to target DNA (e.g., single-stranded target DNA) is retained.
[0431] As another non-limiting example, in some embodiments, the variant Cas9 protein carries mutations of P475A, W476A, N477A, D1125A, W1126A, and D1127A, thus reducing the peptide's ability to cleave target DNA. These Cas9 proteins have reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA).
[0432] As another non-limiting example, in some embodiments, the variant Cas9 protein carries mutations of H840A, W476A, and W1126A, thus reducing the peptide's ability to cleave target DNA. These Cas9 proteins have reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some embodiments, the variant Cas9 protein carries mutations of H840A, D10A, W476A, and W1126A, thus reducing the peptide's ability to cleave target DNA. These Cas9 proteins have reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, the variant Cas9 has restored the catalytic His residue (A840H) at position 840 of the Cas9HNH domain.
[0433] As another non-limiting example, in some embodiments, the variant Cas9 protein carries mutations of H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A, thus reducing the peptide's ability to cleave target DNA. While these Cas9 proteins have reduced ability to cleave target DNA (e.g., single-stranded target DNA), they retain the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some embodiments, the variant Cas9 protein carries mutations of D10A, H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A, thus reducing the peptide's ability to cleave target DNA. While these Cas9 proteins have reduced ability to cleave target DNA (e.g., single-stranded target DNA), they retain the ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, when the variant Cas9 protein carries mutations of W476A and W1126A, or when the variant Cas9 protein carries mutations of P475A, W476A, N477A, D1125A, W1126A, and D1127A, the variant Cas9 protein will not bind effectively to the PAM sequence. Therefore, in some examples, when using these variant Cas9 proteins in a binding method, the method does not require the PAM sequence. In other words, in some embodiments, when using these variant Cas9 proteins in a binding method, the method may include a guide RNA, but the method can be performed without the PAM sequence (therefore, binding specificity is provided by the target segment of the guide RNA). Other residues can be mutated to achieve the above effect (i.e., to partially inactivate one or more nucleases). As non-restrictive examples, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 may be altered (i.e., substituted). In addition, mutations other than alanine substitution are also acceptable.
[0434] In some implementations, the variant Cas9 protein has reduced catalytic activity (e.g., when the Cas9 protein has mutations such as D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987, for example: D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A). The variant Cas9 protein will still bind to the target DNA in a site-specific manner (because it is still guided to the target DNA sequence via guide RNA), as long as it retains the ability to interact with guide RNA.
[0435] In some implementations, the variant Cas protein may be spCas9, spCas9-VRQR, spCas9-VRER, xCas9(sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK, or spCas9-LRVSQL.
[0436] In some embodiments, a modified SpCas9 is used, comprising amino acid substitutions for D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (SpCas9-MQKFRAER) and is specific to the altered PAM 5'-NGC-3'.
[0437] Alternatives to Cas9 in *Streptococcus pyogenes* can include RNA-guided endonucleases from the Cpf1 family, which exhibit cleavage activity in mammalian cells. CRISPR from *Prevotella* and *Francisella* 1 (CRISPR / Cpf1) is a DNA-editing technology similar to the CRISPR / Cas9 system. Cpf1 is a class II CRISPR / Cas system RNA-guided endonuclease. This acquired immune mechanism is present in *Prevotella* and *Francisella* bacteria. The Cpf1 gene, linked to the CRISPR locus, encodes an endonuclease that uses guide RNA to locate and cleave viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9, overcoming some limitations of the CRISPR / Cas9 system. Unlike Cas9 nucleases, Cpf1-mediated DNA cleavage results in a double-strand break with a short 3′ overhang. The staggered cleavage pattern of Cpf1 opens up possibilities for directional gene transport, similar to traditional restriction enzyme selection, which can improve gene editing efficiency. Like the aforementioned Cas9 variants and direct homologs, Cpf1 can also expand the number of CRISPR-targeted sites to AT-rich regions or AT-rich gene bodies lacking the NGG PAM sites favored by SpCas9. The Cpf1 locus contains a mixed α / β domain, RuvC-I followed by a helical region, RuvC-II, and a zinc finger-like domain. The Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9. Furthermore, Cpf1 lacks an HNH endonuclease domain, and its N-terminus lacks the α-helical recognition leaflet of Cas9. The Cpf1 CRISPR-Cas domain structure reveals that Cpf1 possesses unique functionality, classifying it as a type 2 V CRISPR system. The Cas1, Cas2, and Cas4 proteins encoded by the Cpf1 locus are more similar to those of types I and III compared to type II systems. Functional Cpf1 does not require trans-activation of CRISPR RNA (tracrRNA), and therefore only requires CRISPR (crRNA). This is advantageous for genome editing because Cpf1 is not only smaller than Cas9, but also has a smaller sgRNA molecule (approximately half the number of nucleotides in Cas9). Unlike G-enriched PAMs targeted by Cas9, the Cpf1-crRNA complex cleaves target DNA or RNA by discriminating against the adjacent 5'-YTN-3' motif. After PAM discrimination, Cpf1 introduces a sticky-terminal-like DNA double-strand break containing a 4 or 5-nucleotide overhang.
[0438] The Cas12 domain of the nucleobase editor
[0439] Microbial CRISPR-Cas systems are typically classified into Class I and Class II systems. Class I systems possess multiple subunit effector complexes, while Class II systems possess a single protein effector. For example, Cas9 and Cpf1 are Class II effectors, although they differ in morphology (type II and type V, respectively). Besides Cpf1, Class II type V CRISPR-Cas systems also include Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i. See, for example: Shmakov et al., “Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems”, Mol. Cell, 2015 Nov. 5; 60(3):385-397; Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?”, CRISPR Journal, 2018, 1(5):325-336; and Yan et al., “Functionally Diverse Type VCRISPR-Cas Systems”, Science, 2019 Jan. 4; 363:88-91; the full contents of these articles are incorporated herein by reference. Type V Cas proteins contain RuvC (or RuvC-like) endonuclease domains. Although the production of mature CRISPR RNA (crRNA) usually does not depend on tracrRNA, for example, Cas12b / C2c1 requires tracrRNA to produce crRNA. Cas12b / C2c1 relies on both crRNA and tracrRNA to cleave DNA.
[0440] The nucleic acid programmable DNA-binding proteins considered in this invention include Cas proteins (Cas12 proteins) classified as type 2V. Non-limiting examples of Cas type 2V proteins include Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i, their homologs, or modifications thereof. The Cas12 protein used herein may also be referred to as Cas12 nuclease, Cas12 domain, or Cas12 protein domain. In some embodiments, the Cas12 protein of this invention comprises an amino acid sequence interspersed with an internal fusion protein domain (e.g., a deaminase domain).
[0441] In some embodiments, the Cas12 domain is a nuclease-inactivating Cas12 domain or a Cas12 cleavage enzyme. In some embodiments, the Cas12 domain is a nuclease-active domain. For example, the Cas12 domain may be a Cas12 domain that cleaves one strand of a double-stranded nucleic acid (e.g., a double-stranded DNA molecule). In some embodiments, the Cas12 domain contains any of the amino acid sequences shown herein. In some embodiments, the Cas12 domain contains an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the amino acid sequences shown herein. In some implementations, the Cas12 domain contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the amino acid sequences shown herein. In some embodiments, the Cas12 domain contains an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical consecutive amino acid residues compared to any of the amino acid sequences shown herein.
[0442] In some embodiments, a protein containing a fragment of Cas12 is provided. For example, in some embodiments, the protein contains one or two Cas12 domains: (1) a gRNA-binding domain of Cas12; or (2) a DNA-cleaving domain of Cas12. In some embodiments, a protein containing Cas12 or a fragment thereof is therefore referred to as a "Cas12 variant". Cas12 variants and Cas12 or a fragment thereof share common homology. For example, Cas12 variants are at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% homology with wild-type Cas12. In some implementations, the Cas12 variants may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type Cas12. In some embodiments, the Cas12 variant contains a fragment of Cas12 (e.g., a gRNA-binding domain or a DNA cleavage domain), and therefore the fragment is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to the corresponding fragment of wild-type Cas12. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the corresponding fragment of wild-type Cas12. In some embodiments, the fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids in length.
[0443] In some embodiments, Cas12 corresponds to or includes a portion or all of the Cas12 amino acid sequence having one or more mutations that modify the Cas12 nuclease activity. Such mutations include, for example, amino acid substitutions within the RuvC nuclease domain of Cas12. In some embodiments, the provided Cas12 variants or homologs are at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to wild-type Cas12. In some embodiments, the provided Cas12 variants have shorter or longer amino acid sequences, differing by about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, or about 100 or more amino acids.
[0444] In some embodiments, the Cas12 fusion protein provided herein comprises the full-length amino acid sequence of the Cas12 protein, such as one of the Cas12 sequences provided herein. However, in other embodiments, the fusion protein provided herein does not comprise the full-length Cas12 sequence, but only one or more fragments thereof. The examples of suitable amino acid sequences of Cas12 domains provided herein, as well as other suitable Cas12 domain sequences and fragments, are as commonly understood by those skilled in the art.
[0445] Typically, type 2 V Cas proteins possess a single functional RuvC endonuclease domain (see, for example, Chen et al., “CRISPR-Cas12a target binding unleashes indiscriminate single-stranded DNase activity.” Science 360:436-439 (2018)). In some cases, the Cas12 protein is a variant of the Cas12b protein (see Strecker et al., Nature Communications, 2019, 10(1): Art. No.: 212). In one embodiment, the amino acid sequence of the variant Cas12 peptide differs from that of the wild-type Cas12 protein by 1, 2, 3, 4, 5 or more amino acids (e.g., deletion, insertion, substitution, fusion). In some cases, amino acid alterations (e.g., deletion, insertion, or substitution) of the variant Cas12 peptide reduce the activity of the Cas12 peptide. For example, in some cases, the variant Cas12 is a Cas12b polypeptide with less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the cleavage enzyme activity of the corresponding wild-type Cas12b protein. In some cases, the variant Cas12b protein has no substantial cleavage enzyme activity.
[0446] In some cases, the variant Cas12b protein has reduced cleavage enzyme activity. For example, the variant Cas12b protein has less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% of the cleavage enzyme activity of the wild-type Cas12b protein.
[0447] In some implementations, the Cas12 protein comprises an RNA-guided endonuclease from the Cas12a / Cpf1 family, which exhibits activity in mammalian cells. CRISPR from Prevotella and Francisella 1 (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. Cpf1 is a class II CRISPR / Cas system RNA-guided endonuclease. This acquired immune mechanism is present in Prevotella and Francisella bacteria. The Cpf1 gene is linked to a CRISPR locus and encodes an endonuclease that uses guide RNA to locate and cleave viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9, overcoming some limitations of the CRISPR / Cas9 system. Unlike Cas9 nucleases, Cpf1-mediated DNA cleavage results in a double-strand break with a short 3′ overhang. The staggered cleavage pattern of Cpf1 opens up possibilities for directional gene transport, similar to traditional restriction enzyme selection, which can improve gene editing efficiency. Like the aforementioned Cas9 variants and direct homologs, Cpf1 can also expand the number of sites that allow CRISPR to target AT-rich regions or AT-rich gene bodies lacking the NGG PAM sites favored by SpCas9. The Cpf1 locus contains a mixed α / β domain followed by RuvC-I, RuvC-II, and a zinc finger-like domain. The Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9. Furthermore, unlike Cas9, Cpf1 lacks an HNH endonuclease domain, and the N-terminus of Cpf1 lacks the α-helical recognition leaflet of Cas9. The Cpf1 CRISPR-Cas domain structure reveals that Cpf1 possesses unique functionality, classifying it as a type 2 V CRISPR system. The Cas1, Cas2, and Cas4 proteins encoded by the Cpf1 locus are more similar to those of types I and III compared to type II systems. Functional Cpf1 does not require trans-activation of CRISPR RNA (tracrRNA), and therefore only requires CRISPR (crRNA). This is advantageous for genome editing because Cpf1 is not only smaller than Cas9, but also has a smaller sgRNA molecule (approximately half the number of nucleotides in Cas9). Unlike Cas9, which targets G-enriched PAMs, the Cpf1-crRNA complex cleaves target DNA or RNA by discriminating between adjacent motifs 5'-YTN-3' or 5'-TTTN-3'. After PAM discrimination, Cpf1 introduces a sticky-terminal-like DNA double-strand break containing a 4 or 5-nucleotide overhang.
[0448] In some embodiments of the invention, a vector encoding a CRISPR enzyme that has been mutated on the corresponding wild-type enzyme may be used, thus the mutated CRISPR enzyme lacks the ability to cleave one or both strands of the target polynucleotide containing the target sequence. Cas12 may refer to and exemplify wild-type Cas12 polypeptides (e.g., Cas12 from Bacillus hisashii) that are at least or at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identical and / or sequence homologous peptides. Cas12 can refer to, and exemplify, wild-type Cas12 polypeptides (e.g., polypeptides from Bacillus hisashii (BhCas12b), Bacillus sp. V3-13 (BvCas12b), and Alicyclobacillus acidiphilus (AaCas12b)) having at most or at most about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology. Cas12 can also refer to wild-type or modified forms of Cas12 proteins, which may include amino acid alterations such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof.
[0449] Nucleic acid programmable DNA-binding proteins
[0450] Some embodiments of the invention provide fusion proteins comprising domains that can function as nucleic acid-programmable DNA-binding proteins, which can be used to guide proteins, such as base editors, to specific nucleic acid (e.g., DNA or RNA) sequences. In particular embodiments, the fusion protein comprises a nucleic acid-programmable DNA-binding protein domain and a deaminase domain. Non-limiting examples of nucleic acid-programmable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i. Non-restrictive examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, and Csc2. Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector protein, type V Cas effector protein, type VI Cas effector protein, CARF, DinG, their homologs, or their modified or engineered forms. Other nucleic acid-programmable DNA-binding proteins are also within the scope of this invention, but they may not be explicitly listed in this disclosure.See, for example: Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct; 1:325-336. doi:10.1089 / crispr.2018.0033; Yan et al., “Functionally diverse type V CRISPR-Cas systems” Science. 2019 Jan 4; 363(6422):88-91. doi:10.1126 / science.aav7271, etc., which have been incorporated into this paper in full.
[0451] One example of a nucleic acid-programmable DNA-binding protein with PAM specificity distinct from Cas9 is the clustered regularly interspaced short palindromic repeat from *Prevotella* and *Francisella* 1 (Cpf1). Similar to Cas9, Cpf1 is also a class 2 CRISPR effector. Cpf1 has been shown to mediate reliable DNA interference, distinct from Cas9. Cpf1 is a single RNA-guided endonuclease lacking tracrRNA, utilizing T-enriched proto-intercalating adjacent motifs (TTN, TTTN, or YTN). Furthermore, Cpf1 cleaves DNA via staggered double-strand breaks. Of the 16 Cpf1-f family proteins, two enzymes from the genera *Acidaminococcus* and *Lachnospiraceae* have shown effective gene editing activity in human cells. Cpf1 protein is known in the art and has been described in the past, for example: Yamano et al., “Crystal structure of Cpf1 in complex with guide RNA and target DNA.” Cell (165) 2016, pp. 949-962; the full content of which is incorporated herein by reference.
[0452] Suitable for use in this composition and method is a nuclease-inactivating Cpf1 (dCpf1) variant with a guide nucleotide sequence programmable DNA-binding protein domain. The Cpf1 protein has a RuvC-like endonuclease domain similar to that of Cas9, but lacks an HNH endonuclease domain, and the N-terminus of Cpf1 lacks the α-helical recognition leaf of Cas9. As illustrated in Zetsche et al., Cell, 163, 759-771, 2015 (the contents of which are incorporated herein by reference), the RuvC-like domain of Cpf1 is responsible for cleaving both DNA strands, and inactivation of the RuvC-like domain inactivates Cpf1 nuclease activity. For example, mutations corresponding to D917A, E1006A, or D1255A in *Francisella novicida* Cpf1 inactivate Cpf1 nuclease activity. In some embodiments, the dCpf1 of the present invention includes mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A. It is understood that any mutation that inactivates the RuvC domain of Cpf1, such as substitution, deletion, or insertion, may be used according to the present invention.
[0453] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) of any fusion protein provided herein may be a Cpf1 protein. In some embodiments, the Cpf1 protein is a Cpf1 cleavage enzyme (nCpf1). In some embodiments, the Cpf1 protein is a nuclease-inactivated Cpf1 (dCpf1). In some embodiments, Cpf1, nCpf1, or dCpf1 comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the Cpf1 sequence disclosed herein. In some embodiments, dCpf1 comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the Cpf1 sequence disclosed herein, and contains mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A. It should be understood that Cpf1 from other bacterial species can also be used according to the present invention.
[0454] Wild-type *Francisella novicida* Cpf1 (D917, E1006, and D1255 are in bold and underlined):
[0455]
[0456]
[0457] Francisella novicida Cpf1 D917A (A917, E1006, and D1255 are in bold and underlined):
[0458]
[0459] Francisella novicida Cpf1 E1006A (D917, A1006, and D1255 are in bold and underlined):
[0460]
[0461]
[0462] Francisella novicida Cpf1 D1255A (D917, E1006, and A1255 are in bold and underlined)
[0463]
[0464]
[0465] Francisella novicida Cpf1 D917A / E1006A (A917, A1006, and D1255 are in bold and underlined):
[0466]
[0467]
[0468] Francisella novicida Cpf1 D917A / D1255A (A917, E1006, and A1255 are in bold and underlined):
[0469]
[0470] Francisella novicida Cpf1 E1006A / D1255A (D917, A1006, and A1255 are in bold and underlined):
[0471]
[0472] Francisella novicida Cpf1D917A / E1006A / D1255A (A917, A1006, and A1255 are in bold and underlined):
[0473]
[0474]
[0475] In some implementations, a Cas9 domain present in the fusion protein can be replaced by a guide nucleotide sequence programmable DNA-binding protein domain that is not required by the PAM sequence.
[0476] In some embodiments, the Cas9 domain is the Cas9 domain from Staphylococcus aureus (SaCas9). In some embodiments, the SaCas9 domain is an active nuclease SaCas9, an inactive nuclease SaCas9 (SaCas9d), or a SaCas9 nickase (SaCas9n). In some embodiments, SaCas9 contains the N579A mutation, or a corresponding mutation in any amino acid sequence provided herein.
[0477] In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain may bind to a nucleic acid sequence having an atypical PAM. In some embodiments, the SaCas9 domain, SaCas9d domain, or the SaCas9n domain may bind to a nucleic acid sequence having an NNGRRT or NNGRRT PAM sequence. In some embodiments, the SaCas9 domain contains one or more E781X, N967X, and R1014X mutations, or corresponding mutations in any amino acid sequence provided herein, where X is any amino acid. In some embodiments, the SaCas9 domain contains one or more E781K, N967K, and R1014H mutations, or one or more corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SaCas9 domain contains E781K, N967K, or R1014H mutations, or corresponding mutations in any amino acid sequence provided herein.
[0478] SaCas9 sequence example:
[0479]
[0480] The residue N579, indicated by underline and bold above, can be mutated (e.g., to form A579) to produce the SaCas9 nickase.
[0481] SaCas9n sequence example:
[0482]
[0483]
[0484] The residue A579 that can generate the SaCas9 nickase from the N579 mutation is indicated by underline and bold.
[0485] SaKKH Cas9 example:
[0486]
[0487]
[0488] The residue A579 that can generate the SaCas9 nickase from the N579 mutation is shown in underlined and bold. The residues K781, K967, and H1014 that can generate SaKKH Cas9 from the E781, N967, and R1014 mutations are shown in underlined and italicized.
[0489] In some implementations, the napDNAbp is arranged in a circular pattern. In the following sequences, plain text represents the adenosine deaminase sequence, bold sequences indicate sequences derived from Cas9, italic sequences represent linker sequences, and underlined sequences represent binuclear localization sequences.
[0490] CP5 (containing MSP "NGC" PID and "D10A" nickase):
[0491]
[0492]
[0493] In some implementations, the nucleic acid programmable DNA-binding protein (napDNAbp) is a single effector of the microbial CRISPR-Cas system. Single effectors of microbial CRISPR-Cas systems include (but are not limited to): Cas9, Cpf1, Cas12b / C2c1, and Cas12c / C2c3. Typically, microbial CRISPR-Cas systems are divided into Class 1 and Class 2 systems. Class 1 systems have multiple subunit effector complexes, while Class 2 systems have a single protein effector. For example, Cas9 and Cpf1 are Class 2 effectors. Besides Cas9 and Cpf1, Shmakov et al. described three different Class 2 CRISPR-Cas systems (Cas12b / C2c1 and Cas12c / C2c3) in “Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems”, Mol. Cell, 2015 Nov. 5; 60(3):385-397, the full content of which is incorporated herein by reference. The effectors of two of these systems, Cas12b / C2c1 and Cas12c / C2c3, contain RuvC-like endonuclease domains associated with Cpf1. The third system contains effectors with two predicted HEPN RNase domains. The production of mature CRISPR RNA is independent of tracrRNA, unlike CRISPR RNA produced by Cas12b / C2c1. Cas12b / C2c1 relies on both CRISPR RNA and tracrRNA to cleave DNA.
[0494] The crystal structure of *Alicyclobaccillus acidoterrastris* Cas12b / C2c1 (AacC2c1) has been described in a complex with a chimeric single-molecule guide RNA (sgRNA). See, for example, Liu et al., “C2c1-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism”, Mol. Cell, 2017 Jan. 19; 65(2):310-322, the full contents of which are incorporated herein by reference. The crystal structure described is also reported in *Alicyclobaccillus acidoterrastris* C2c1 bound to target DNA, as a ternary complex. See, for example, Yang et al., “PAM-dependent Target DNA Recognition and Cleavage by C2C1 CRISPR-Cas endonuclease”, Cell, 2016 Dec. 15; 167(7): 1814-1828, the full content of which is incorporated herein by reference. The catalytic component configuration of AacC2c1, both containing target and non-target DNA strands, was independently captured and placed within a single RuvC catalytic pocket. Cas12b / C2c1-mediated cleavage resulted in a staggered seven-nucleotide breakage of the target DNA. Comparison of the structures of the Cas12b / C2c1 ternary complex with previously identified Cas9 and Cpf1 counterparts confirmed the diversity of mechanisms employed in the CRISPR-Cas9 system.
[0495] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) of any fusion protein provided herein may be a Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a Cas12b / C2c1 protein. In some embodiments, the napDNAbp is a Cas12c / C2c3 protein. In some embodiments, the napDNAbp contains an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to that of the native Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a native Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp contains an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any napDNAbp sequence provided herein. It should be understood that Cas12b / C2c1 or Cas12c / C2c3 from other bacterial species may also be used according to the present invention.
[0496] Cas12b / C2c1((uniprot.org / uniprot / T0D7A2#2)sp|T0D7A2|
[0497] C2C1_ALIAG CRISPR-related endonuclease C2c1 OS=Alicyclobacillus acido-terrestris (strain ATCC 49025 / DSM 3922 / CIP 106132 / NCIMB13137 / GD3B) GN=c2c1 PE=1 SV=1) amino acid sequence is as follows:
[0498]
[0499] BhCas12b (Bacillus hisashii) NCBI reference sequence: WP_095142515
[0500]
[0501] In some embodiments, Cas12b is BvCas12B. In some embodiments, Cas12b contains amino acid substitutions S893R, K846R, and E837G, which are the amino acid sequence example numbers of BvCas12b provided below.
[0502] BvCas12b (Bacillus sp. V3-13) NCBI reference sequence: WP_101661451.1
[0503]
[0504] Guide polynucleotides
[0505] In one implementation, the guide polynucleotide is a guide RNA. The RNA / Cas complex assists in “guided” the Cas protein to the target DNA. Cas9 / crRNA / tracrRNA cleaves the linear or circular dsDN target complementary to the spacer via endonucleolysis. The target strand not complementary to the crRNA is first cleaved via endonucleolysis, followed by 3′–5′ trimming via exonucleolysis. In nature, DNA binding and cleavage typically require both proteins and two RNAs. However, a single guide RNA (“sgRNA” or simply “gNRA”) can be engineered to incorporate both crRNA and tracrRNA embodiments into a single RNA species. See, for example, Jinek M. et al., Science 337:816–821 (2012), the full contents of which are incorporated herein by reference. Cas9 recognizes short motifs (PAMs or protospacer adjacent motifs) in CRISPR repeat sequences to help distinguish between self and non-self motifs. The Cas9 nuclease sequence and structure are well known to those skilled in the art (see, for example: “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti, JJ et al., Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E. et al., Nature 471:602-607 (2011); and “Programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M. et al., Science 337:816-821 (2012), the full contents of which are incorporated herein by reference). Cas9 is a homologous gene that has been shown in various species, including (but not limited to): Streptococcus pyogenes and Streptococcus thermophilus.Other suitable Cas9 nucleases and sequences are known to those skilled in the art based on this invention, and such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type IICRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the full contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease has an inactivating (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is a nicking enzyme.
[0506] In some embodiments, the guide polynucleotide is at least one single guide RNA (“sgRNA” or “gNRA”). In some embodiments, the guide polynucleotide is at least one tracrRNA. In some embodiments, the guide polynucleotide does not require a PAM sequence to guide the polynucleotide programmable DNA-binding domain (e.g., Cas9 or Cpf1) to the target nucleotide sequence.
[0507] The multinucleotide programmable nucleotide-binding domains (e.g., CRISPR-derived domains) of the base editor disclosed herein can recognize target polynucleotide sequences by being linked to a guide polynucleotide. The guide polynucleotide (e.g., gRNA) is typically single-stranded and can be programmed to bind to the target sequence of the polynucleotide at a site specificity (i.e., via complementary base pairing), thereby directing the base editor, already linked to the guide nucleic acid, to the target sequence. The guide polynucleotide can be DNA. The guide polynucleotide can be RNA. In some embodiments, the guide polynucleotide comprises a natural nucleotide (e.g., adenosine). In some embodiments, the guide polynucleotide comprises a non-natural (or unnatural) nucleotide (e.g., peptide nucleic acid or nucleotide analogue). In some embodiments, the length of the target region of the guide nucleic acid sequence can be at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. The target region of a guide nucleic acid can be between 10 and 30 nucleotides in length, or between 15 and 25 nucleotides in length, or between 15 and 20 nucleotides in length.
[0508] In some embodiments, the guide polynucleotide comprises two or more individual polynucleotides that can interact with each other via, for example, complementary base pairing (e.g., dual guide polynucleotides). For example, the guide polynucleotide may comprise CRISPR RNA (crRNA) and trans-activating CRISPR RNA (tracrRNA). Alternatively, the guide polynucleotide may comprise one or more trans-activating CRISPR RNAs (tracrRNA).
[0509] In type II CRISPR systems, the targeting of nucleic acids by CRISPR proteins (e.g., Cas9) typically requires complementary base pairing between a first RNA molecule (crRNA) (containing a sequence that recognizes the target sequence) and a second RNA molecule (trRNA) (containing a repetitive sequence that forms a backbone region to stabilize the guide RNA-CRISPR protein complex). Such dual guide RNA systems can be used as guide polynucleotides to direct the base editor to the target polynucleotide sequence as described in this paper.
[0510] In some embodiments, the base editor provided herein utilizes a single guide polynucleotide (e.g., gRNA). In some embodiments, the base editor provided herein utilizes dual guide polynucleotides (e.g., dual gRNAs). In some embodiments, the base editor provided herein utilizes one or more guide polynucleotides (e.g., multiple gRNAs). In some embodiments, a single guide polynucleotide is used for different base editors described herein. For example, a single guide polynucleotide can be used for an adenosine base editor, or an adenosine base editor and a cytidine base editor, as described in PCT / US19 / 44935.
[0511] In other embodiments, the guide polynucleotide may contain both the multinucleotide targeting portion and the backbone portion of the nucleic acid in a single molecule (i.e., a single-molecule guide nucleic acid). For example, a single-molecule guide polynucleotide may be a single guide RNA (sgRNA or gRNA). The term guide polynucleotide sequence used herein refers to any single, dual, or multiple-molecule nucleic acid that can interact with a base editor and be directed to the target polynucleotide sequence.
[0512] Typically, guide polynucleotides (e.g., crRNA / trRNA complexes or gRNA) comprise a "polynucleotide-targeting segment," which includes a sequence capable of recognizing and binding to a target polynucleotide sequence, and a "protein-binding segment," which is a guide polynucleotide within a polynucleotide programmable nucleotide-binding domain that stabilizes the base editor. In some embodiments, the polynucleotide-targeting segment of the guide polynucleotide recognizes and binds to DNA polynucleotides, thereby facilitating base editing in DNA. In other examples, the polynucleotide-targeting segment of the guide polynucleotide binds to RNA polynucleotides, thereby facilitating base editing in RNA. As used herein, a "segment" refers to a segment or region of a molecule, such as a continuous nucleotide segment in a guide polynucleotide. A segment can also refer to a region / segment of a complex, and therefore, the segment may comprise a region of more than one molecule. For example, if the guide polynucleotide comprises multiple nucleic acid molecules, the protein-binding segment may comprise all or part of the multiple separate molecules, for example, hybridizing along complementary regions. In some embodiments, a protein-binding segment of DNA-targeting RNA comprising two separate molecules may comprise (i) 40-75 base pairs of a first RNA molecule of 100 base pairs in length; and (ii) 10-25 base pairs of a second RNA molecule of 50 base pairs in length. Unless otherwise expressly defined in specific contexts, the definition of a “segment” does not limit the specific total number of base pairs, any specific number of base pairs from a specified RNA molecule, or the specific number of separate molecules within the complex, and may include a region of any total length within the RNA molecule, and may include a region complementary to other molecules.
[0513] Guide RNA or guide polynucleotides may contain two or more RNAs, such as CRISPR RNA (crRNA) and trans-activating crRNA (tracrRNA). Guide RNA or guide polynucleotides may sometimes contain single-stranded RNA, or a single guide RNA (sgRNA) formed by the fusion of crRNA and a portion of tracrRNA (e.g., the functional portion). Guide RNA or guide polynucleotides may also be dual RNAs containing crRNA and tracrRNA. Furthermore, crRNA can hybridize with target DNA.
[0514] As discussed above, guide RNA or guide polynucleotides can be expression products. For example, the DNA encoding guide RNA can be a vector containing the sequence encoding guide RNA. Guide RNA or guide polynucleotides can be transferred into cells by transfecting cells with isolated guide RNA or plasmid DNA containing the sequence encoding guide RNA and a promoter. Guide RNA or guide polynucleotides can also be transferred into cells using other methods, such as virus-mediated gene delivery.
[0515] Guide RNA or guide polynucleotides can be isolated. For example, guide RNA can be transfected into cells or organisms in isolated RNA form. Guide RNA can be prepared using any in vitro transcription system known in the art. Guide RNA can be transferred into cells in isolated RNA form, rather than in plasmid form containing a sequence encoding guide RNA.
[0516] Guide RNAs or guide polynucleotides may contain three regions: a first region at the 5' end, which can be complementary to target sites in the chromosomal sequence; a second inner region, which can form a stem-loop structure; and a third 3' region, which can be single-stranded. The first region of each guide RNA can also be different, thus each guide RNA guides the fusion protein to a specific target site. In addition, the second and third regions of all guide RNAs can be the same.
[0517] The first region of the guide RNA or guide polynucleotide is complementary to the sequence of the target site in the chromosomal sequence, thus allowing the first region of the guide RNA to pair with the bases of the target site. In some embodiments, the first region of the guide RNA may contain 10 or about 10 to 25 nucleotides (i.e., 10 to 25 nucleotides; or about 10 to about 25 nucleotides; or 10 to about 25 nucleotides; or about 10 to 25 nucleotides) or more. For example, the length of the base-pairing region between the first region of the guide RNA and the target site in the chromosomal sequence may be about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25, or more nucleotides. Sometimes, the length of the first region of the guide RNA may be about 19, 20, or 21 nucleotides.
[0518] Guide RNA or guide polynucleotides may also contain a second region that forms a secondary structure. For example, a secondary structure formed by guide RNA may contain a stem (or hairpin) and a loop. The lengths of the loop and the stem can vary. For example, the length of the loop can range from 3 or about 3 to 10 nucleotides, and the length of the stem can range from 6 or about 6 to 20 base pairs. The stem may contain one or more protrusions containing 1 to 10 or about 10 nucleotides. The total length of the second region can range from 16 or about 16 to 60 nucleotides. For example, the length of the loop may be 4 or about 4 nucleotides, and the length of the stem may be 12 or about 12 base pairs.
[0519] Guide RNA or guide polynucleotides may also be contained in a third region at the 3' end, which is essentially single-stranded. For example, the third region is sometimes not complementary to any chromosomal sequence in the cell of interest, and sometimes not complementary to the rest of the guide RNA. Furthermore, the length of the third region can vary. The length of the third region can be greater than or equal to about 4 nucleotides. For example, the length of the third region can range from 5 or about 5 to 60 nucleotides.
[0520] Guide RNA or guide polynucleotides can target any exon or intron of a gene target. In some embodiments, the guide sequence can target exon 1 or 2 of the gene; in other examples, the guide sequence can target exon 3 or 4 of the gene. The composition may contain multiple guide RNAs that all target the same exon, or in some embodiments, multiple guide RNAs that target different exons. It can target both exons and introns of a gene.
[0521] Guide RNA or guide polynucleotides can target nucleic acid sequences of 20 or about 20 nucleotides. The target nucleic acid can be fewer than or about 20 nucleotides. The length of the target nucleic acid can be at least or at least about 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, or any number of nucleotides between 1 and 100. The length of the target nucleic acid can be at most or at most about 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, or any number of nucleotides between 1 and 100. The target nucleic acid sequence can be or about 20 bases immediately following the first 5' of the PAM nucleotide. Guide RNA can target nucleic acid sequences. The target nucleic acid may be at least or at least about 1-10, 1-20, 1-30, 1-40, 1-50, 1-60, 1-70, 1-80, 1-90, or 1-100 nucleotides.
[0522] Guide polynucleotides, such as guide RNA, can refer to nucleic acids that can hybridize with another nucleic acid, such as target nucleic acids or prototype spacers in the cellular genome. Guide polynucleotides can be RNA or DNA. Guide polynucleotides can be programmed or designed to bind to nucleic acids in a sequence-site specific manner. Guide polynucleotides can contain a polynucleotide chain and can be called single guide polynucleotides. Guide polynucleotides can contain two polynucleotide chains and can be called double guide polynucleotides. Guide RNA can be introduced into cells as RNA molecules. For example, RNA molecules can be transcribed in vitro and / or can be chemically synthesized. RNA can be transcribed from synthetic DNA molecules, for example: Gene fragments. The guide RNA is then introduced into the cell as an RNA molecule. The guide RNA can also be introduced into the cell as a non-RNA nucleic acid molecule, such as a DNA molecule. For example, the DNA encoding the guide RNA is operably linked to a promoter control sequence to express the guide RNA in the cells of interest. The RNA coding sequence is operably linked to a promoter sequence recognized by RNA polymerase III (Pol III). Plasmid vectors that can be used to express the guide RNA include (but are not limited to): the px330 vector and the px333 vector. In some embodiments, the plasmid vector (e.g., the px333 vector) may contain at least two DNA sequences encoding the guide RNA.
[0523] Methods for selecting, designing, and validating guide polynucleotides, such as guide RNAs and target sequences, are described herein and are well known to those skilled in the art. For example, to minimize the potential promiscuity in the deaminase domains (e.g., the AID domain) of a nucleobase editor system, the number of residues that may be unintentionally targeted and deaminated (e.g., off-target C residues that may remain on ssDNA within the target nucleic acid locus) can be minimized. Furthermore, software tools can be used to optimize gRNAs corresponding to the target nucleic acid sequence, for example, to minimize overall off-target activity throughout the genome. For example, when using *Streptococcus pyogenes* Cas9 to select potential target domains, all off-target sequences (the aforementioned selected PAMs, such as NAG or NGG) containing at most a certain number (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) mismatched base pairs can be identified throughout the genome. The first region of gRNAs complementary to the target site can be identified, and all first regions (e.g., crRNAs) can be sorted according to their predicted off-target scores; the topmost target domain represents the region with the greatest potential on-target activity and the least off-target activity. The function of the target gRNA candidates can be analyzed using methods known in the art and / or those described herein.
[0524] As a non-limiting example, the target DNA hybridization sequence in the crRNA of the guide RNA used by Cas9 can be determined using a DNA sequence search algorithm. gRNA design can be performed using the publicly available tool cas-offinder custom gRNA design software described in: Bae S., Park J., & Kim J.-S. Cas-OFFinder: A fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases. Bioinformatics 30, 1473-1475 (2014). After calculating the off-target probabilities of the isogenome sequences, the software calculates a score for the guide sequences. Typically, for guide sequences ranging in length from 17 to 24, a pairing range from perfect match to 7 mismatches is considered. Once the off-target sites are determined by computer calculation, a total score is calculated for each guide sequence, and the results are summarized and output in a table format using a web interface. In addition to identifying potential target sites adjacent to PAM sequences, the software can also identify all PAM adjacent sequences that differ from a selected target site by 1, 2, 3, or more than 3 nucleotides. The genomic DNA sequence of the target nucleic acid sequence, such as the target gene, can be obtained, and repetitive elements can be screened using publicly available tools, such as the RepeatMasker program. RepeatMasker searches for repetitive elements and regions of low complexity in the input DNA sequence. The output provides detailed annotations of the repetitive sequences present in the specified query sequence.
[0525] Following discrimination, the first region of the guide RNA (e.g., crRNA) can be hierarchically ordered based on its distance from the target site, its direct homology, and the presence of 5' nucleotides to achieve a close match with the associated PAM sequence (e.g., 5'G determined based on close match in human genomes containing associated PAMs such as NGG PAM of *Streptococcus pyogenes*, NNGRRT or NNGRRV PAM of *S. aureus*). In this context, orthogonality refers to the minimum number of sequences in the human genome that mismatch with the target sequence. "High orthogonality" or "good orthogonality" can, for example, refer to a 20 nt target domain in the human genome that has no identical sequences to the intended target and contains no sequences in the target sequence that contain one or two mismatches. Target domains with good orthogonality can be selected to minimize off-target DNA fragmentation.
[0526] In some embodiments, a reporter system may be used to detect base editing activity and test guide polynucleotide candidates. In some embodiments, the reporter system may include an assay based on the reporter gene, wherein base editing activity causes the expression of the reporter gene. For example, the reporter system may include a reporter gene containing a deactivated start code, such as a mutation from 3'-TAC-5' to 3'-CAC-5' on the template strand. When the target C is successfully deaminated, the corresponding mRNA will be transcribed into 5'-AUG-3' instead of 5'-GUG-3', allowing the reporter gene to be translated. Suitable reporter genes are those commonly understood by those skilled in the art. Non-limiting examples of reporter genes include genes encoding green fluorescent protein (GFP), red fluorescent protein (RFP), luciferase, secreted alkaline phosphatase (SEAP), or any other gene whose expression can be detected and is commonly understood by those skilled in the art. The reporter subsystem can be used to test many different gRNAs, for example, to determine which residue(s) are associated with the target DNA sequence of each deaminase. sgRNAs targeting non-template strands can also be used for testing to analyze the off-target effects of specific base-editing proteins, such as Cas9 deaminase fusion proteins. In some embodiments, these gRNAs can be engineered so that the start codon of the mutation does not pair with the gRNA bases. The guide polynucleotide may comprise standard ribonucleotides, modified ribonucleotides (e.g., pseudouridine), ribonucleotide isomers, and / or ribonucleotide analogs. In some embodiments, the guide polynucleotide may comprise at least one detectable tag. The detectable tag may be a fluorophore (e.g., FAM, TMR, Cy3, Cy5, Texas Red, Oregon Green, Alexa Fluors, Halo tags, or suitable fluorescent dyes), a detection tag (e.g., biotin, digoxigenin, and analogs), quantum dots, or gold particles.
[0527] Guide polynucleotides can be chemically synthesized, enzymatically synthesized, or a combination thereof. For example, guide RNA can be synthesized using a standard solid-phase synthesis method based on iminophosphodiester (IPE). Alternatively, guide RNA can be synthesized in vitro by operatively linking DNA encoding the guide RNA to a promoter control sequence recognized by a phage RNA polymerase. Examples of suitable phage promoter sequences include the T7, T3, and SP6 promoter sequences, or variations thereof. In embodiments where the guide RNA comprises two separate molecules (e.g., crRNA and tracrRNA), crRNA can be chemically synthesized, and tracrRNA can be enzymatically synthesized.
[0528] In some implementations, the base editor system may comprise multiple guide polynucleotides, such as gRNAs. For example, the gRNAs may target one or more target loci contained in the base editor system (e.g., at least 1 gRNA, at least 2 gRNAs, at least 5 gRNAs, at least 10 gRNAs, at least 20 gRNAs, at least 30 gRNAs, at least 50 gRNAs). The multiple gRNA sequences may be tandemly arranged, and preferably separated by a forward repeat sequence.
[0529] The DNA sequence encoding guide RNA or guide polynucleotides can also be part of the vector. Furthermore, the vector may contain additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), selectable marker sequences (e.g., GFP or antibiotic resistance genes such as puromycin), origin of replication, and analogues. The DNA molecule encoding guide RNA can also be linear. The DNA molecule encoding guide RNA or guide polynucleotides can also be circular.
[0530] In some implementations, components of one or more base editor systems may be encoded by DNA sequences. These DNA sequences may be introduced together or separately into the expression system, for example, in a cell. For example, a DNA sequence encoding a polynucleotide programmable nucleotide binding domain and a guide RNA may be introduced into the cell, and each DNA sequence may be part of a separate molecule (e.g., a vector containing a sequence encoding the polynucleotide programmable nucleotide binding domain and a second vector containing a sequence encoding the guide RNA) or simultaneously part of the same molecule (e.g., a vector containing both the coding (and regulatory) sequences of the polynucleotide programmable nucleotide binding domain and the guide RNA).
[0531] Guided polynucleotides may contain one or more modifications to provide nucleic acids with novel or enhanced features. Guided polynucleotides may contain nucleic acid affinity tags. Guided polynucleotides may contain synthetic nucleotides, synthetic nucleotide analogs, nucleotide derivatives, and / or modified nucleotides.
[0532] In some embodiments, the gRNA or guide polynucleotide may contain modifications. Modifications can be made at any position on the gRNA or guide polynucleotide. More than one modification can be made on a single gRNA or guide polynucleotide. The gRNA or guide polynucleotide can be quality controlled after modification. In some embodiments, quality control includes PAGE, HPLC, MS, or any combination thereof.
[0533] Modifications to gRNA or guide polynucleotides can include substitution, insertion, deletion, chemical modification, physical modification, stabilization, purification, or any combination thereof.
[0534] gRNA or guide polynucleotides can also be modified as follows: 5' adenosine, 5' guanosine-triphosphate capping, 5' N7-methylguanosine-triphosphate capping, 5' triphosphate capping, 3' phosphorylation, 3' thiophosphorylation, 5' phosphorylation, 5' thiophosphorylation, Cis-Syn thymidine dimer, trimer, C12 spacer, C3 spacer, C6 spacer, d spacer, PC spacer, r spacer, spacer 18, spacer 9, 3'-3' modification, 5'-5' modification, no base, acridine, azobenzene, biotin, biotin BB, biotin TEG, cholesterol-based TEG, dethiobiotin TEG, DNP TEG, DNP-X, DOTA, dT-biotin, double biotin, PC biotin, psoralen C2, psoralen C6, TINA, 3' DABCYL, fluorescence quencher (black hole) quencher) 1. Fluorescent quencher 2. DABCYL SE, dT-DABCYL, IRDye QC-1, QSY-21, QSY-35, QSY-7, QSY-9, carboxyl linker, thiol linker, 2'-deoxyribonucleoside analog purine, 2'-deoxyribonucleoside analog pyrimidine, ribonucleoside analog, 2'-O-methylribonucleoside analog, sugar-modified analog, swing / universal base, fluorescent dye tag, 2'-fluoroRNA, 2'-O-methylRNA, methylphosphonate, phosphodiester DNA, phosphodiester RNA, thiophosphate DNA, thiophosphate RNA, UNA, pseudouridine-5'-triphosphate, 5'-methylcytidine-5'-triphosphate, or any combination thereof.
[0535] In some implementations, the modification is permanent. In other examples, the modification is transient. In some implementations, multiple modifications are made to the gRNA or guide polynucleotide. Modifications to the gRNA or guide polynucleotide can alter the physicochemical properties of the nucleotide, such as its conformation, polarity, hydrophobicity, chemical reactivity, base pairing interactions, or any combination thereof.
[0536] The PAM sequence can be any PAM sequence known in the relevant art. Suitable PAM sequences include (but are not limited to): NGG, NGA, NGC, NGN, NGT, NGCG, NGAG, NGAN, NGNG, NGCN, NGCG, NGTN, NNGRRT, NNNRRT, NNGRR(N), TTTV, TYCV, TYCV, TATV, NNNNGATT, NNAGAAW, or NAAMC. Y is pyrimidine; N is any nucleotide base; W is A or T.
[0537] Modifications can also involve phosphate thioester substitution. In some embodiments, the native phosphodiester bond is subject to rapid degradation by cellular nucleases; and the use of phosphate thioester (PS) bond substitution to modify the internucleotide linkage can provide greater stability against hydrolytic degradation by cells. Modifications can improve the stability of gRNA or guide polynucleotides. Modifications can also enhance biological activity. In some embodiments, phosphate thioester-enhanced RNA gRNA can inhibit RNase A, RNase T1, calf serum nuclease, or any combination thereof. These properties allow PS-RNA gRNA to be used in applications where it is highly likely to be exposed to nucleases in vivo or in vitro. For example, a phosphate thioester (PS) bond can be introduced between the last 3-5 nucleotides of the 5' or 3' end of the gRNA, which can inhibit exonuclease degradation. In some embodiments, a phosphate thioester bond can be added to the entire gRNA, reducing attack by endonucleases.
[0538] Originally spaced adjacent motifs
[0539] The term "proto-spacer adjacent motif (PAM)" or PAM-like motif refers to a 2-6 base pair DNA sequence immediately following the DNA sequence targeted by the Cas9 nuclease in the CRISPR bacterial acquired immune system. In some embodiments, the PAM may be a 5' PAM (i.e., located upstream of the 5' end of the proto-spacer). In other embodiments, the PAM may be a 3' PAM (i.e., located downstream of the 5' end of the proto-spacer).
[0540] The PAM sequence is essential for target binding, but the exact sequence depends on the Cas protein morphology.
[0541] The base editors provided herein may contain CRISPR protein-derived domains that can bind nucleotide sequences containing typical or atypical protospacer adjacent motif (PAM) sequences. The PAM site is a nucleotide sequence adjacent to the target polynucleotide sequence. Some embodiments of the invention provide base editors containing all or part of a CRISPR protein with different PAM specificities. For example, typically, Cas9 proteins, such as Cas9 (spCas9) from *Streptococcus pyogenes*, require a typical NGG PAM sequence to bind specific nucleic acid regions, where "N" in "NGG" stands for adenine (A), thymine (T), guanine (G), or cytosine (C), and "G" stands for guanine. The PAM can be CRISPR protein-specific and can vary between different base editors containing different CRISPR protein-derived domains. The PAM can be the 5' or 3' of the target sequence. The PAM can be upstream or downstream of the target sequence. The length of the PAM can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides. PAMs are typically 2-6 nucleotides in length. Several PAM variants are illustrated in Table 1 below.
[0542] Table 1. Cas9 protein and corresponding PAM sequence
[0543] Variant PAM spCas9 NGG spCas9-VRQR NGA spCas9-VRER NGCG xCas9(sp) NGN saCas9 NNGRRT saCas9-KKH NNNRRT spCas9-MQKSER NGCG spCas9-MQKSER NGCN spCas9-LRKIQK NGTN spCas9-LRVSQK NGTN spCas9-LRVSQL NGTN spCas9-MQKFRAER NGC Cpf1 5’(TTTV) SpyMac 5’-NAA-3’
[0544] In some implementations, PAM is NGC. In some implementations, NGC PAM is recognized by Cas9 variants. In some implementations, NGC PAM variants include one or more amino acid substitutions selected from D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (collectively referred to as "MQKFRAER").
[0545] In some embodiments, PAM is NGT. In some embodiments, NGT PAM is recognized by a Cas9 variant. In some embodiments, an NGT PAM variant is generated through a targeted mutation at one or more residues 1335, 1337, 1135, 1136, 1218, and / or 1219. In some embodiments, an NGT PAM variant is created through a targeted mutation at one or more residues 1219, 1335, 1337, and 1218. In some embodiments, an NGT PAM variant is created through a targeted mutation at one or more residues 1135, 1136, 1218, 1219, and 1335. In some embodiments, the NGT PAM variant is selected from a set of targeted mutations shown in Tables 2 and 3 below.
[0546] Table 2: NGT PAM variant mutations at residues 1219, 1335, 1337, and 1218
[0547] Variant E1219V R1335Q T1337 G1218 1 F V T 2 F V R 3 F V Q 4 F V L 5 F V T R 6 F V R R 7 F V Q R 8 F V L R 9 L L T 10 L L R 11 L L Q 12 L L L 13 F I T 14 F I R 15 F I Q 16 F I L 17 F G C 18 H L N 19 F G C A 20 H L N V 21 L A W 22 L A F 23 L A Y 24 I A W 25 I A F 26 I A Y
[0548] Table 3: NGT PAM variant mutations at residues 1135, 1136, 1218, 1219, and 1335
[0549] Variant D1135L S1136R G1218S E1219V R1335Q 27 G 28 V 29 I 30 A 31 W 32 H 33 K 34 K 35 R 36 Q 37 T 38 N 39 I 40 A 41 N 42 Q 43 G 44 L 45 S 46 T 47 L 48 I 49 V 50 N 51 S 52 T 53 F 54 Y 55 N1286Q I1331F
[0550] In some implementations, the NGT PAM variants are variants 5, 7, 28, 31, or 36 selected from Tables 2 and 3. In some implementations, the variants have improved NGT PAM identification.
[0551] In some embodiments, the NGT PAM variant has mutations at residues 1219, 1335, 1337, and / or 1218. In some embodiments, an NGT PAM variant with a modified identification mutation is selected from the variants provided in Table 4 below.
[0552] Table 4: NGT PAM variant mutations at residues 1219, 1335, 1337, and 1218
[0553] Variant E1219V R1335Q T1337 G1218 1 F V T 2 F V R 3 F V Q 4 F V L 5 F V T R 6 F V R R 7 F V Q R 8 F V L R
[0554] In some implementations, a base editor specific to NGT PAM can be generated by the information provided in Table 5A below.
[0555] Table 5A. NGT PAM variants
[0556]
[0557] In some implementations, the NGTN variant is variant 1. In some implementations, the NGTN variant is variant 2. In some implementations, the NGTN variant is variant 3. In some implementations, the NGTN variant is variant 4. In some implementations, the NGTN variant is variant 5. In some implementations, the NGTN variant is variant 6.
[0558] In some embodiments, the Cas9 domain is the Cas9 domain (SpCas9) derived from *Streptococcus pyogenes*. In some embodiments, the SpCas9 domain is an active nuclease SpCas9, an inactive nuclease SpCas9 (SpCas9d), or a SpCas9 nickase (SpCas9n). In some embodiments, SpCas9 contains a D10X mutation, or a corresponding mutation in any amino acid sequence provided herein, where X is any amino acid other than D. In some embodiments, SpCas9 contains a D10A mutation, or a corresponding mutation in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain, SpCas9d domain, or SpCas9n domain may bind to a nucleic acid sequence having an atypical PAM. In some embodiments, the SpCas9 domain, SpCas9d domain, or SpCas9n domain may bind to a nucleic acid sequence having an NGG, NGA, or NGCG PAM sequence.
[0559] In some embodiments, the Cas9 domain is the Cas9 domain derived from *Streptococcus pyogenes* (SpCas9). In some embodiments, the SpCas9 domain is an active nuclease SpCas9, an inactive nuclease SpCas9 (SpCas9d), or a SpCas9 nickase (SpCas9n). In some embodiments, SpCas9 contains a D9X mutation, or a corresponding mutation in any amino acid sequence provided herein, where X is any amino acid other than D. In some embodiments, SpCas9 contains a D9A mutation, or a corresponding mutation in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain, SpCas9d domain, or SpCas9n domain may bind to a nucleic acid sequence having an atypical PAM. In some embodiments, the SpCas9 domain, SpCas9d domain, or SpCas9n domain may bind to a nucleic acid sequence having an NGG, NGA, or NGCG PAM sequence. In some embodiments, the SpCas9 domain includes one or more D1135X, R1335X, and T1337X mutations, or corresponding mutations in any amino acid sequence provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain includes one or more D1135E, R1335Q, and T1337R mutations, or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain includes one or more D1135X, R1335X, and T1337X mutations, or corresponding mutations in any amino acid sequence provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain includes one or more D1135V, R1335Q, and T1337R mutations, or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain includes the D1135V, R1335Q, and T1337R mutations, or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain includes one or more D1135X, G1218X, R1335X, and T1337X mutations, or corresponding mutations in any amino acid sequence provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain includes one or more D1135V, G1218R, R1335Q, and T1337R mutations, or corresponding mutations in any amino acid sequence provided herein.In some implementations, the SpCas9 domain contains mutations such as D1135V, G1218R, R1335Q, and T1337R, or corresponding mutations in any amino acid sequence provided herein.
[0560] In some embodiments, the Cas9 variant is a Cas9 variant specific to the altered PAM sequence. In some embodiments, other Cas9 variants and PAM sequences are described in Miller et al., Continuous evolution of SpCas9 variants compatible with non-G PAMs. Nat Biotechnol (2020). https: / / doi.org / 10.1038 / s41587-020-0412-8, the full contents of which are incorporated herein by reference. In some embodiments, the Cas9 variant does not require specific PAM. In some embodiments, the Cas9 variant, for example, the SpCas9 variant, is specific to NRNH PAM, where R is A or G and H is A, C, or T. In some embodiments, the SpCas9 variant is specific to PAM sequences AAA, TAA, CAA, GAA, TAT, GAT, or CAC. In some embodiments, the SpCas9 variant contains amino acid substitutions at positions 1114, 1134, 1135, 1137, 1139, 1151, 1180, 1188, 1211, 1218, 1219, 1221, 1249, 1256, 1264, 1290, 1318, 1317, 1320, 1321, 1323, 1332, 1333, 1335, 1337, or 1339, or their corresponding positions, according to SEQ ID NO: 1. In some embodiments, the SpCas9 variant contains amino acid substitutions at positions 1114, 1135, 1218, 1219, 1221, 1249, 1320, 1321, 1323, 1332, 1333, 1335, or 1337, or their corresponding positions, according to SEQ ID NO: 1. In some embodiments, the SpCas9 variant contains amino acid substitutions at positions 1114, 1134, 1135, 1137, 1139, 1151, 1180, 1188, 1211, 1219, 1221, 1256, 1264, 1290, 1318, 1317, 1320, 1323, 1333 or their corresponding positions in SEQ ID NO: 1. In other embodiments, the SpCas9 variant contains amino acid substitutions at positions 1114, 1131, 1...
Claims
1. A base editor system comprising a guide polynucleotide and a fusion protein, said fusion protein comprising an adenosine deaminase domain and a CRISPR / Cas protein domain in sequence from N-terminus to C-terminus, said adenosine deaminase domain being a mutant of SEQ ID NO: 2 or a mutant of the amino acid sequence 2 to 167 of SEQ ID NO: 2, or said adenosine deaminase domain being a heterodimer formed by sequentially linking and fusing wild-type adenosine deaminase of SEQ ID NO: 101 and a mutant of the amino acid sequence 2 to 167 of SEQ ID NO: 2, and said mutant having one of the following amino acid modifications as per SEQ ID NO: 2: a) Y147T; b) Y147R; c) Y147R, Q154R and Y123H; d) I76Y, Y147R, and Q154R; e) Y147R, Q154R and T166R; f) Y147T and Q154R; g) Y147T and Q154S; or h) I76Y, Y123H, Y147R and Q154R, in, The guide polynucleotide dominates the fusion protein to cause nucleobase deamination of the hemoglobin subunit γ1 and / or 2 promoters.
2. The base editor system according to claim 1, wherein, The mutant of SEQ ID NO: 2 or the mutant of the amino acid sequence from position 2 to 167 of SEQ ID NO: 2 in the adenosine deaminase domain contains arginine (R) at position 147 of the amino acid sequence described in SEQ ID NO:
2.
3. The base editor system according to claim 1, wherein, The mutant of SEQ ID NO: 2 in the adenosine deaminase domain or the mutant of the amino acid sequence from position 2 to 167 of SEQ ID NO: 2 contains modifications of Y147R, Q154R, and Y123H.
4. A base editor system comprising a guide polynucleotide and a fusion protein, said fusion protein comprising an adenosine deaminase domain and a CRISPR / Cas protein domain in sequence from N-terminus to C-terminus, said adenosine deaminase domain being a mutant of SEQ ID NO: 2 or a mutant of the amino acid sequence 2 to 167 of SEQ ID NO: 2, or said adenosine deaminase domain being a heterodimer formed by sequentially fusing wild-type adenosine deaminase of SEQ ID NO: 101 and a mutant of the amino acid sequence 2 to 167 of SEQ ID NO: 2, and said mutant having one of the following amino acid modifications as per SEQ ID NO: 2: a) Y147T; b) Y147R; c) Y147R, Q154R and Y123H; d) I76Y, Y147R, and Q154R; e) Y147R, Q154R and T166R; f) Y147T and Q154R; g) Y147T and Q154S; or h) I76Y, Y123H, Y147R and Q154R, in, The guide polynucleotide governs the fusion protein to introduce a modification from A•T to G•C at position 198 of the promoter of hemoglobin subunit γ1 and / or 2.
5. The base editor system according to claim 4, wherein, The CRISPR / Cas protein domain includes the Cas9 domain.
6. The base editor system according to claim 5, wherein, The Cas9 domain contains either inactivated Cas9 or nicking enzyme Cas9.
7. The base editor system according to claim 6, wherein, The Cas9 domain can bind to programmable DNA and is selected from Streptococcus pyogenes (S. pyogenes). Streptococcus pyogenes Cas9, Staphylococcus aureus ( Staphylococcus aureus Cas9, Streptococcus thermophilus ( Streptococcus thermophilus 1. Cas9 and Neisseria meningitidis ( Neisseria meningitidis The group consisting of Cas9.
8. The base editor system according to claim 7, wherein, The meningococcal Cas9 mentioned is meningococcal 2 Cas9.
9. The base editor system according to claim 8, wherein, The Cas9 domain contains the amino acid sequence of SEQ ID NO:
1.
10. The base editor system according to claim 4, wherein, The CRISPR / Cas protein domain contains the amino acid sequence of SEQ ID NO:
3.
11. The base editor system according to claim 4, wherein, The guide polynucleotide comprises a spacer subsequence group consisting of nucleotide sequences selected from SEQ ID NO: 144, SEQ ID NO: 145, SEQ ID NO: 146, SEQ ID NO: 147, SEQ ID NO: 148, SEQ ID NO: 149, SEQ ID NO: 150, SEQ ID NO: 151, SEQ ID NO: 152, SEQ ID NO: 153, SEQ ID NO: 154, SEQ ID NO: 155, SEQ ID NO: 156, SEQ ID NO: 157, SEQ ID NO: 158, SEQ ID NO: 159, SEQ ID NO: 160, SEQ ID NO: 161, SEQ ID NO: 162, SEQ ID NO: 163, SEQ ID NO: 164, SEQ ID NO: 165, SEQ ID NO: 166 and SEQ ID NO:
167.
12. The base editor system according to claim 11, wherein, The guide polynucleotide contains a 2'-O-methyl or thiophosphate modification.
13. The base editor system according to claim 11, wherein, The guide polynucleotide includes a backbone containing the nucleotide sequence of SEQ ID NO:
78.
14. A base editor system comprising guide RNA and mRNA, The mRNA encodes a base editor as shown in SEQ ID NO: 107, and The guide RNA contains the following nucleic acid sequence: mCsmUsmUsGACCAAUAGCCUUGACAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCCGGUGCUsmUsmUsmU (SEQ ID NO: 176), in, "mC" represents 2'-O-methylcytidine, "mU" represents 2'-O-methyluridine, and "s" indicates the position of the thiophosphate.
15. A base editor system comprising a fusion protein or a polynucleotide encoding said fusion protein, said fusion protein comprising, in order from N-terminus to C-terminus, an adenosine deaminase domain and a CRISPR / Cas protein domain, wherein, The adenosine deaminase domain is a mutant of SEQ ID NO: 2 or a mutant of the amino acid sequence from position 2 to 167 of SEQ ID NO: 2, or the adenosine deaminase domain is a heterodimer formed by sequentially connecting and fusing the wild-type adenosine deaminase of SEQ ID NO: 101 and a mutant of the amino acid sequence from position 2 to 167 of SEQ ID NO: 2, wherein the mutant contains histidine (H) at position 123 of the amino acid sequence described in SEQ ID NO: 2, and the mutant has one of the following amino acid modifications described in SEQ ID NO: 2: a) Y147T; b) Y147R; c) Y147R, Q154R and Y123H; d) I76Y, Y147R, and Q154R; e) Y147R, Q154R and T166R; f) Y147T and Q154R; g) Y147T and Q154S; or h) I76Y, Y123H, Y147R and Q154R; and The base editor system also includes one or more guide polynucleotides that direct the fusion protein to cause A•T to G•C modifications of the target nucleobases in the hemoglobin subunit γ1 and / or 2 promoter regions.
16. The base editor system according to claim 15, wherein, The mutant of SEQ ID NO: 2 in the adenosine deaminase domain or the mutant of the amino acid sequence from position 2 to 167 of SEQ ID NO: 2 contains modifications of Y147R, Q154R, and Y123H.
17. The base editor system according to claim 15, wherein, The CRISPR / Cas protein domain is programmable for DNA binding and is selected from the group consisting of Streptococcus pyogenes Cas9, Staphylococcus aureus Cas9, Streptococcus thermophilus Cas9, and Neisseria meningitidis Cas9.
18. The base editor system according to claim 15, wherein, The guide polynucleotide comprises a spacer subsequence group consisting of nucleotide sequences selected from SEQ ID NO: 144, SEQ ID NO: 145, SEQ ID NO: 146, SEQ ID NO: 147, SEQ ID NO: 148, SEQ ID NO: 149, SEQ ID NO: 150, SEQ ID NO: 151, SEQ ID NO: 152, SEQ ID NO: 153, SEQ ID NO: 154, SEQ ID NO: 155, SEQ ID NO: 156, SEQ ID NO: 157, SEQ ID NO: 158, SEQ ID NO: 159, SEQ ID NO: 160, SEQ ID NO: 161, SEQ ID NO: 162, SEQ ID NO: 163, SEQ ID NO: 164, SEQ ID NO: 165, SEQ ID NO: 166 and SEQ ID NO:
167.
19. The base editor system according to claim 18, wherein, The guide polynucleotide comprises a backbone containing the nucleotide sequence of SEQ ID NO:
78.
20. The base editor system according to claim 18, wherein, The base editor is selected from the group consisting of ABE8.8, ABE8.13 and ABE8.
17.
21. The base editor system according to claim 15, wherein, The one or more guide polynucleotides dominate the fusion protein to cause a modification at position 114 of the promoter of the hemoglobin subunit γ1 and / or 2.
22. The base editor system according to claim 15, wherein, The one or more guide polynucleotides dominate the fusion protein to cause a modification at nucleobase 5 or 8 of reference SEQ ID NO:
177.
23. The base editor system according to claim 18, wherein, The guide polynucleotide contains a 2'-O-methyl or thiophosphate modification.
Citation Information
Patent Citations
towel ring
CN3315821D
cell phone
CN3329834D
Improved sun-bonnet for horses
US100000A
Split inteins, conjugates and uses thereof
US20150344549A1
AAV delivery of nucleobase editors
US20180127780A1