Methods of replacing pathogenic amino acids using a programmable base editor system
By using a base editor with a multinucleotide programmable nucleotide binding domain and a deaminase domain, precise editing of target nucleotide sequences is achieved, solving the problems of low efficiency and large side effects in existing technologies, and effectively treating genetic diseases such as sickle cell disease and heme disease.
Patent Information
- Application Number
- CN201980046480.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-17
- Filing Date
- 2019-05-11
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2039-09-16
AI Technical Summary
Existing CRISPR gene editing technology is inefficient in correcting point mutations and often causes random insertions or deletions at the target locus, making it ineffective in treating genetic diseases caused by point mutations.
Using a base editor that includes a multinucleotide programmable nucleotide-binding domain and a deaminase domain, specific nucleobases can be replaced by targeted editing of the target nucleotide sequence. For example, in the sickle cell disease variant of β-globulin protein, the thymidine nucleobase can be edited to a cytidine nucleobase, replacing valine with alanine, to form a normal β-globulin protein variant.
It improves the precision and efficiency of gene editing, reduces the side effects of random insertions or deletions, and effectively treats genetic diseases such as sickle cell disease and heme disease.
Smart Images

Figure CN112534054B_ABST
Abstract
Description
[0001] Related applications
[0002] This application claims the benefits of U.S. Provisional Application No. 62 / 670,521, filed May 11, 2018; U.S. Provisional Application No. 62 / 670,539, filed May 11, 2018; and U.S. Provisional Application No. 62 / 780,890, filed December 17, 2018, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This invention provides compositions and methods for using a base editor, the base editor comprising a polynucleotide programmable nucleotide-binding domain and a nucleobase editing domain for binding to a guide polynucleotide. This invention also provides a base editor system for editing nucleobases of a target nucleotide sequence. Background Technology
[0004] For most known genetic diseases, correction of point mutations at target loci, rather than random gene disruption, is needed to investigate or address the underlying causes of the disease. Current somatic editing techniques using the clustered regularly interspaced short palindromic repeat (CRISPR) system introduce double-stranded DNA breaks at the target locus as the first step in gene correction. In response to these double-stranded DNA breaks, cellular DNA repair processes mostly result in random insertions or deletions (indels) at the DNA cleavage site via non-homologous end joining. Although most genetic diseases are caused by point mutations, current point mutation correction methods are inefficient due to the cellular response to dsDNA breaks and often result in numerous random insertions and deletions (indels) at the target locus. Therefore, a modified form of somatic editing is needed that is more efficient and produces far fewer unwanted products, such as random insertions or deletions (indels) or translocations.
[0005] References are incorporated
[0006] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the same extent that each individual publication, patent, or patent application is expressly and individually indicated to be incorporated by reference. Unless otherwise stated, all publications, patents, and patent applications mentioned in this specification are incorporated herein by reference in their entirety. Summary of the Invention
[0007] This article provides a method for treating a genetic condition in a subject, comprising administering to the subject a base editor or a polynucleotide encoding the base editor, wherein the base editor includes a polynucleotide-programmable nucleotide-binding domain and a deaminase domain; administering to the subject a guide polynucleotide, wherein the guide polynucleotide targets the base editor to a target nucleotide sequence of the subject; and treating the genetic condition by editing the nucleobase by deaminating a nucleobase of the target nucleotide sequence when the base editor is targeted to the target nucleotide sequence, thereby changing the nucleobase to another nucleobase; wherein the genetic condition is caused by a pathogenic amino acid in a protein, and wherein another nucleobase replaces the pathogenic amino acid with a benign amino acid that is different from the wild-type amino acid of the protein.
[0008] This article provides a method for generating cells, tissues, or organs for treating a genetic condition in a subject, wherein the method includes contacting the cell, tissue, or organ with a base editor or a polynucleotide encoding the base editor, wherein the base editor includes a polynucleotide-programmable nucleotide-binding domain and a deaminase domain; contacting the cell, tissue, or organ with a guide polynucleotide, wherein the guide polynucleotide targets the base editor to a target nucleotide sequence of the cell, tissue, or organ; and editing the nucleobase by deaminoing a nucleobase of the target nucleotide sequence after targeting the base editor to the target nucleotide sequence, thereby generating cells, tissues, or organs for treating the genetic condition by changing the nucleobase to another nucleobase; wherein the genetic condition is caused by a pathogenic amino acid in a protein, and wherein another nucleobase replaces the pathogenic amino acid with a benign amino acid that is different from the wild-type amino acid of the protein. In some embodiments, the method further includes delivering the cell, tissue, or organ to the subject. In some embodiments, the cell, tissue, or organ is autologous to the subject. In some implementations, the cell, tissue, or organ is allogeneic to the subject.
[0009] In some embodiments, the nucleotide is located in a gene that is the cause of the genetic condition. In some embodiments, the editing comprises editing a plurality of nucleotides located in the gene, wherein the plurality of nucleotides are not the cause of the genetic condition. In some embodiments, the editing further comprises editing one or more additional nucleotides located in at least one other gene. In some embodiments, the gene and the at least one other gene encode one or more subunits of the protein.
[0010] In some embodiments, the edited nucleobases are in genes listed in Table 3A or 3B, and the editing results in an amino acid alteration in a protein encoded by the gene indicated in Table 3A or 3B. In some embodiments, the genetic condition is ACADM deficiency, sickle cell disease (SCD), heme disease, β-thalassemia, Pendred syndrome, autosomal dominant Parkinson's disease, or α-1 antitrypsin deficiency (A1AD).
[0011] In one embodiment, the invention is characterized by compositions and methods for replacing pathogenic amino acids using a programmable nucleobase editor. Specifically, the invention provides compositions and methods for editing the thymidine (T) nucleobase to a cytidine (C) nucleobase at the codon of the sixth amino acid (Sickle HbS; E6V) of a sickle cell disease variant of the β-globulin protein, thereby replacing valine with alanine (E6A). Replacing valine at position 6 of Sickle HbS with alanine forms a β-globulin protein variant lacking the sickle cell phenotype (e.g., possessing the characteristics of a normal β-globulin protein (HbA; E6) and lacking the possibility of polymerization, as in the case of the pathogenic variant HbS, etc.). Therefore, the compositions and methods of the invention can be used to treat sickle cell disease. In one embodiment, the edited nucleobase is in the HBB gene encoding β-globulin, and this base editing results in a change of amino acid 6 from valine (Val) to alanine (Ala) in the β-globulin (HBB) protein encoded by the HBB gene (β6Val→Ala). In some embodiments, the genetic condition is sickle cell disease or heme disease. In some embodiments, the base editing results in a change of the E6V>E6A amino acid in the β subunit of heme.
[0012] In another embodiment, the present invention provides a method for editing an HBB polynucleotide comprising a single nucleotide polymorphism (SNP) associated with sickle cell disease, wherein the method comprises contacting the HBB polynucleotide with a base editor, the base editor being complexed with one or more guide polynucleotides, wherein the base editor comprises a polynucleotide programmable DNA-binding domain and an adenosine deaminase domain, and wherein the one or more guide polynucleotides target the base editor to achieve an A·T to G·C alteration of the sickle cell disease-associated SNP.
[0013] In another embodiment, the present invention provides a cell generated by introducing a base editor, a polynucleotide encoding the base editor, the base editor comprising a polynucleotide programmable DNA binding domain and an adenosine deaminase domain; and one or more guiding polynucleotides that target the base editor to achieve A·T to G·C alterations in SNPs associated with sickle cell disease.
[0014] In another embodiment, the present invention provides a method for treating sickle cell disease in a subject, comprising administering cells as described herein to a subject in need.
[0015] In another embodiment, the present invention provides an isolated cell or cell population that reproduces or expands from cells of any of the embodiments described herein.
[0016] In another embodiment, the present invention provides a method for treating sickle cell disease in a subject, wherein the method comprises administering a base editor or a polynucleotide encoding the base editor to the subject in need, wherein the base editor comprises a polynucleotide programmable DNA-binding domain and an adenosine deaminase domain; and one or more guide polynucleotides that target the base editor to achieve A·T to G·C alterations in SNPs associated with sickle cell disease.
[0017] In another embodiment, the present invention provides a method for generating red blood cells (erythrocytes) or precursor cells thereof, wherein the method comprises (a) introducing a base editor or a polynucleotide encoding the base editor into a erythrocyte precursor cell containing a sickle cell disease-associated SNP, wherein the base editor comprises a polynucleotide-programmable nucleotide-binding domain and an adenosine deaminase domain, and one or more guide polynucleotides; wherein the one or more guide polynucleotides target the base editor to achieve an A·T to G·C alteration of the sickle cell disease-associated SNP; and (b) differentiating the erythrocyte precursor cell into a erythrocyte.
[0018] In another embodiment, the present invention provides a base editor comprising: (i) a polynucleotide programmable DNA binding domain comprising Streptococcus thermophilus 1Cas9 (St1Cas9) and (ii) an adenosine deaminase domain.
[0019] In another embodiment, the present invention provides a guide RNA (gRNA) comprising a nucleic acid sequence selected from the following: CUUCUCCACAGGAGUCAGAU; ACUUCUCCACAGGAGUCAGAU; and GACUUCUCCACAGGAGUCAGAU.
[0020] In another embodiment, the present invention provides a base editor comprising: (i) a polynucleotide programmable DNA binding domain comprising modified Staphylococcus aureus Cas9 (SaCas9), and (ii) an adenosine deaminase domain.
[0021] In another embodiment, the present invention provides a guide RNA (gRNA) comprising nucleic acid sequences selected from the following: UCCACAGGAGUCAGAUGCAC and UCCACAGGAGUCAGAUGCAC.
[0022] In another embodiment, the present invention provides a guide RNA (gRNA) comprising a nucleic acid sequence selected from the following: UUCUCCACAGGAGUCAGA; CUUCUCCACAGGAGUCAGA; ACUUCUCCACAGGAGUCAGA; GACUUCUCCACAGGAGUCAGA; and AGACUUCUCCACAGGAGUCAGA.
[0023] In one embodiment, the base editing results in an E342K>E342G amino acid alteration in the SERPINA1 gene encoding an α-1 antitrypsin protein. In one embodiment, the genetic disorder is medium-chain acyl-CoA dehydrogenase (ACADM) deficiency. In one embodiment, the base editing results in a K329E>K329G amino acid alteration in the protein encoded by the medium-chain acyl-CoA dehydrogenase (ACADM) gene. In one embodiment, the genetic disorder is a heme disorder. In one embodiment, the base editing results in an E26K>E26G amino acid alteration in the β subunit of heme encoded by the HBB gene. In some embodiments, the genetic disorder is Pendred syndrome. In some embodiments, the base editing results in a T416P>T416F amino acid alteration in SLC26A4; SLC26A4 is a solute carrier family 26 member 4 (PDS) protein encoded by the PDS gene. In some embodiments, the genetic condition is autosomal dominant Parkinson's disease. In some embodiments, the editing results in an alteration of the A30P>A30L amino acids in the α-synuclein (SNCA) protein encoded by the SNCA gene.
[0024] In various embodiments of any of the examples described herein, the A·T to G·C change at the sickle cell disease-associated SNP replaces valine in the HBB polypeptide with alanine. In various embodiments, the sickle cell disease-associated SNP results in the expression of the HBB polypeptide having valine at amino acid position 6. In various embodiments, the sickle cell disease-associated SNP replaces glutamic acid with valine.
[0025] In various embodiments of any of the examples described herein, the contact is in a cell, eukaryotic cell, mammalian cell, or human cell. In various embodiments, the subject is a mammal or human. In various embodiments, the cell is in vivo or ex vivo. In various embodiments, the cell or its precursor cell is an induced pluripotent stem cell hematopoietic stem cell, a common myeloid progenitor cell, a proerythroblast, an erythroblast, a reticulocyte, or a erythrocyte. In various embodiments, the hematopoietic stem cell is CD34. + Cells. In various embodiments, the cells are derived from a subject suffering from sickle cell disease. In various embodiments, the cells are autologous to the subject. In various embodiments, the cells are allogeneic or xenogeneic to the subject. In various embodiments of any of the embodiments described herein, the method comprises delivering the base editor or the polynucleotide encoding the base editor, and one or more guide polynucleotides, to the subject's cells.
[0026] In various embodiments of any of the examples described herein, the polynucleotide-programmable DNA binding domain is a modified Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1Cas9 (St1Cas9), a modified Streptococcus pyogenes Cas9 (SpCas9), or a variant thereof. In various embodiments, the polynucleotide-programmable DNA binding domain comprises a modified SaCas9 having altered protospacer-adjacent motif (PAM) specificity. In various embodiments, the altered PAM comprises the nucleic acid sequence 5'-NNNRRT-3'. In various embodiments, the modified SaCas9 comprises amino acid substitutions of E782K, N968K, and R1015H, or corresponding amino acid substitutions thereof.
[0027] In various embodiments, the polynucleotide programmable DNA binding domain comprises a variant of SpCas9 with altered protospacer adjacent motif (PAM) specificity. In various embodiments, the altered PAM comprises the nucleic acid sequence 5'-NGC-3'.
[0028] In various embodiments, the modified SpCas9 comprises amino acid substitutions of D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R, or corresponding amino acid substitutions thereof. In various embodiments, the polynucleotide programmable DNA-binding domain is a nuclease inactive or nickase variant. In various embodiments, the nickase variant comprises amino acid substitution D10A, or corresponding amino acid substitutions thereof.
[0029] In various embodiments of any of the examples described herein, the base editor further includes a zinc finger domain. In various embodiments, the zinc finger domain includes identifying helical sequences RNEHLEV, QSTTLKR, and RTEHLAR, or identifying helical sequences RGEHLRQ, QSGTLKR, and RNDKLVP. In various cases, the zinc finger domain is one or more of zf1ra or zf1rb.
[0030] In various embodiments of any of the examples described herein, the adenosine deaminase domain deaminates adenine in deoxyribonucleic acid (DNA). In various embodiments, the adenosine deaminase is a modified adenosine deaminase that does not occur in nature. In various embodiments, the adenosine deaminase is a TadA deaminase. In various embodiments, the TadA deaminase is TadA*7.10.
[0031] In various embodiments of any of the examples described herein, the one or more guide RNAs comprise a CRISPR RNA (crRNA) and a trans-coding small RNA (tracrRNA), wherein the crRNA comprises a nucleic acid sequence complementary to an HBB nucleic acid sequence containing a sickle cell disease-associated SNP. In various embodiments, the base editor is complexed with a single guide RNA (sgRNA) comprising a nucleic acid sequence complementary to an HBB nucleic acid sequence containing a sickle cell disease-associated SNP.
[0032] In various embodiments of any of the examples described herein, the St1Cas9 comprises the following amino acid sequence:
[0033]
[0034] In various embodiments, the base editor includes a linker between the polynucleotide programmable DNA binding domain and the adenosine deaminase domain. In various embodiments, the linker comprises the following amino acid sequence: SGGSSGGSSGSETPGTSESATPES. In various embodiments, the base editor includes one or more nuclear localization signals. In various embodiments, the base editor comprises the following amino acid sequence:
[0035]
[0036] In various embodiments of any of the examples described herein, the guide RNA further comprises the following nucleic acid sequence:
[0037] GUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG.
[0038] In various implementations, the guide RNA comprises a nucleic acid sequence selected from the following:
[0039] CUUCUCCACAGGAGUCAGAUGUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG;
[0040] ACUUCUCCACAGGAGUCAGAUGUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG; or
[0041] GACUUCUCCACAGGAGUCAGAUGUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG.
[0042] In various embodiments of any of the examples described herein, the protein-nucleic acid complex comprises a base editor as described herein and a guide RNA as described herein.
[0043] In various embodiments of any of the examples described herein, the modified SaCas9 comprises amino acid substitutions for E782K, N968K, and R1015H, or corresponding amino acid substitutions thereof. In various embodiments, the SaCas9 comprises the following amino acid sequence:
[0044]
[0045] In various embodiments of any of the examples described herein, the base editor comprises the following amino acid sequence:
[0046]
[0047] In various implementations, the base editor comprises the following amino acid sequence:
[0048]
[0049]
[0050] In various embodiments of any of the examples described herein, the guide RNA further comprises the following nucleic acid sequence:
[0051] GUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGCGAGAUUUUU.
[0052] In various implementations, the guide RNA comprises the following nucleic acid sequence:
[0053] UCCACAGGAGUCAGAUGCACGUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGCGAGAUUUUU, or the following
[0054] Nucleic acid sequence:
[0055] CUCCACAGGAGUCAGAUGCACGUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGCGAGAUUUUUU.
[0056] In some embodiments, any of the methods provided herein further includes a second edit of an additional nucleobase. In one embodiment, the additional nucleobase is not the cause of the genetic condition. In another embodiment, the additional nucleobase is the cause of the genetic condition.
[0057] In another embodiment, a method for treating a genetic condition in a subject is provided, the method comprising administering a base editor to the subject in need, wherein the base editor includes a polynucleotide-programmable nucleotide-binding domain and a deaminase domain for binding to a guide polynucleotide; binding the guide polynucleotide to a target nucleotide sequence of the subject's polynucleotide; and editing the nucleotide by deaminating a nucleobase of the target nucleotide sequence after binding the guide polynucleotide to the target nucleotide sequence, thereby treating the genetic condition by changing the nucleotide to another nucleobase; wherein the nucleobase is located in a regulatory element or regulatory region of a gene.
[0058] In another embodiment, a method for producing cells, tissues, or organs for treating a genetic condition in a subject in need includes contacting the cell, tissue, or organ with a base editor, wherein the base editor includes a polynucleotide-programmable nucleotide-binding domain and a deaminase domain for binding to a guide polynucleotide; binding the guide polynucleotide to a target nucleotide sequence of the polynucleotide in the cell, tissue, or organ; and editing the nucleotide by deaminosing a nucleobase of the target nucleotide sequence after binding the guide polynucleotide to the target nucleotide sequence, thereby producing cells, tissues, or organs for treating the genetic condition by changing the nucleotide to another nucleobase; wherein the nucleobase is in a regulatory element of a gene. In some embodiments, the method further includes delivering the cell, tissue, or organ to a subject. In some embodiments, the cell, tissue, or organ is autologous to the subject. In some embodiments, the cell, tissue, or organ is allogeneic to the subject. In some embodiments, the cell, tissue, or organ is xenogeneic to the subject.
[0059] In some embodiments of the above method, the gene is the cause of the genetic condition. In some embodiments, the gene is not the cause of the genetic condition. In some embodiments, the editing results in a change in the amount of transcription of the gene. In some embodiments, the change is an increase in the amount of transcription of the gene. In some embodiments, the change is a decrease in the amount of transcription of the gene. In some embodiments, the editing alters the binding pattern of at least one protein to the regulatory element. In some embodiments, the regulatory element is a promoter, enhancer, repressor, silencer, insulator, start codon, stop codon, Kozak concordant sequence, splice acceptor, splice donor, splice site, 3' untranslated region (UTR), 5' untranslated region (UTR), or intergenic region of the gene. In some embodiments, the editing results in the removal of a splice site. In some embodiments, the editing results in the addition of a splice site. In some embodiments, the editing results in intron inclusion. In some embodiments, the editing results in exon skipping. In some embodiments, the editing results in the removal of a start codon, stop codon, or Kozak concordant sequence. In some implementations, the edit results in the addition of a start codon, a stop codon, or a Kozak common sequence. In some implementations, the edit involves editing multiple nucleobases located in the regulatory elements of the gene.
[0060] In some embodiments of the above method, the editing comprises editing a plurality of nucleobases, wherein at least one nucleobase of the plurality of nucleobases is located in at least one additional regulatory element of at least one additional gene. In some embodiments, the gene and the at least one additional gene encode one or more subunits of at least one protein.
[0061] In some embodiments of the above method, the edit is selected from any of the changes shown in Table 4 herein. In some embodiments, the genetic condition is sickle cell disease (SCD), also known as sickle cell anemia. In some embodiments, the genetic condition is hereditary persistent fetal heme (HPFH) syndrome. In some embodiments, the nucleotide is located at c.-114 to -102 of HBG1 / 2. In some embodiments, the nucleotide is located in the promoter of HBG1 / 2.
[0062] In some embodiments of the above method, the method includes a second edit of at least one additional nucleotide, wherein the additional nucleotide is not located in a regulatory element of the gene. In some embodiments, the additional nucleotide is located in a coding region of the protein.
[0063] In some embodiments of the method described in the above examples, the deaminase domain is an adenosine deaminase domain. In some embodiments, the deaminase domain is a cytidine deaminase domain. In some embodiments, the adenosine deaminase domain can deaminate adenine in deoxyribonucleic acid (DNA). In some embodiments, the guide polynucleotide comprises ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). In some embodiments, the guide polynucleotide comprises a CRISPR RNA (crRNA) sequence, a trans-activated CRISPR RNA (tracrRNA) sequence, or a combination thereof.
[0064] In some embodiments, any of the methods provided herein further comprises a second guide polynucleotide. In some embodiments, the second guide polynucleotide comprises ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). In some embodiments, the second guide polynucleotide comprises a CRISPR RNA (crRNA) sequence, a trans-activated CRISPR RNA (tracrRNA) sequence, or a combination thereof. In some embodiments, the second guide polynucleotide targets the base editor to a second target nucleotide sequence.
[0065] In some embodiments, the polynucleotide-programmable DNA-binding domain comprises a Cas9 domain, a Cpf1 domain, a CasX domain, a CasY domain, a Cas12b / C2c1 domain, or a Cas12c / C2c3 domain. In some embodiments, the polynucleotide-programmable DNA-binding domain is a nuclease-inactivating domain. In some embodiments, the polynucleotide-programmable DNA-binding domain is a nickase. In some embodiments, the polynucleotide-programmable DNA-binding domain comprises a Cas9 domain. In some embodiments, the Cas9 domain comprises a nuclease-inactivated Cas9 (dCas9), a Cas9 nickase (nCas9), or a nuclease-activated Cas9. In some embodiments, the Cas9 domain comprises a Cas9 nickase. In some embodiments, the polynucleotide-programmable DNA-binding domain is an engineered or modified polynucleotide-programmable DNA-binding domain.
[0066] In some embodiments, any of the methods provided herein further includes a second base editor. In some embodiments, the second base editor includes a deaminase domain that is different from the deaminase domain of the other base editors.
[0067] In some embodiments, the base editing results in less than 20% of insertions / deletions. In some embodiments, the base editing results in less than 15% of insertions / deletions. In some embodiments, the base editing results in less than 10% of insertions / deletions. In some embodiments, the base editing results in less than 5% of insertions / deletions. In some embodiments, the base editing results in less than 4% of insertions / deletions. In some embodiments, the base editing results in less than 3% of insertions / deletions. In some embodiments, the base editing results in less than 2% of insertions / deletions. In some embodiments, the base editing results in less than 1% of insertions / deletions. In some embodiments, the base editing results in less than 0.5% of insertions / deletions. In some embodiments, the base editing results in less than 0.1% of insertions / deletions. In some embodiments, the base editing does not cause translocation. Attached Figure Description
[0068] The features of the invention are specifically set forth in the appended claims. The features and advantages of the invention will be further understood by referring to the following illustrative embodiments and their accompanying drawings, in which the principles of the invention are described and utilized:
[0069] Figure 1This diagram compares healthy subjects with patients suffering from antitrypsin deficiency (A1AD). In healthy subjects, the alpha-1 antitrypsin (A1AT) protein protects the lungs from protease damage, and the liver releases alpha-1 antitrypsin into the bloodstream. In patients with alpha-1 antitrypsin deficiency (A1AD), the lack of normal alpha-1 antitrypsin leads to lung tissue damage. The accumulation of abnormal alpha-1 antitrypsin in hepatocytes of the liver leads to cirrhosis.
[0070] Figure 2 This displays typical ranges of serum α-1 antitrypsin (A1AT) levels for different genotypes (normal (MM); atypical carriers of α-1 antitrypsin deficiency (MZ, SZ); and isotype carriers of deficiency (SS, ZZ)). Serum α-1 antitrypsin (AAT) concentrations are expressed in μM on the left "y" axis, which is commonly used in the literature. The right "y" axis shows the approximate conversion of serum AAT concentrations to mg / dL units, as commonly reported in clinical laboratories and various measurement techniques (turbidity assays or radiation immunodiffusion assays).
[0071] Figure 3 This describes the sequence used to correct the target site E342K within the SERPINA1 gene encoding A1AT. Highlights are the atypical spCas9 NGC PAM and the target A nucleotide, editing which will result in the desired E342K correction. Also note the additional off-target A; editing these can create benign paired genes such as E342G or D341G.
[0072] Figure 4A bar graph showing the amount of secreted proteins in the culture supernatant of HEK293T, transiently transfected with plastids encoding variants of different A1AT proteins. A1AT concentrations were determined by ELISA using published methods (Borel et al., 2017, “A-1 Antitrypsin Deficiency: Methods and Protocols,” 10.1007 / 978-1-4939-7163-3). The two most commonly used clinical variants of A1AT (e.g., pathogenic mutations) are E264V (PiS-pair gene) and E342K (PiZ-pair gene). The abundance of the PiS and PiZ proteins produced was lower than that of the wild-type proteins. Either D341G or E342G proteins was produced in amounts similar to the wild-type. Therefore, the adenine base editor and base editing method described herein are used to generate the aforementioned benign paired gene, which restores A1AT secretion in hepatocytes and simultaneously improves hepatotoxicity and increases A1AT circulation to the lungs. In this figure, A1AT: α-1 antitrypsin; A1AD: α-1 antitrypsin deficiency; "Z mutation" is the E342K (PiZ paired gene) mutation; "S mutation" is the E264V (PiS paired gene).
[0073] Figure 5 This is a schematic diagram illustrating a strategy in which DNA deoxyadenosine deaminase evolves from TadA. An *E. coli* library contains a plastid library of the mutant ecTadA (TadA*) gene fused to dCas9 and selectable plastids targeting A·T to G·C mutations to repair antibiotic resistance genes. Mutations from surviving TadA* variants are introduced into ABE structures used for base editing in human cells.
[0074] Figure 6 The table shows the first 8 amino acids of mature heme (Hb), including normal HbA, pathogenic variants Sickle HbS and HbC, and the HbG Makassar variant, which is phenotypically similar to HbA but does not polymerize like HbS. Figure 6 The diagram shows the amino acids encoded at position 6 in each Hb type, as well as the DNA and mRNA sequences, which encode the first 8 amino acids of the Hb protein.
[0075] Figure 7A and 7B The experimental results describe the editing of the nucleobase adenosine (A) in the sequence (CAC) into a guanine nucleoside (G), which is complementary to the codon encoding valine at amino acid position 6 of HbS using various A to G base editors (ABE) that recognize different PAM sequences. Figure 7AThis is a table describing the characteristics of the HBB gRNA and the corresponding ABE test, including the locations of required editing and possible off-target editing. Figure 7B This is a graph showing the results of using ABE to perform base editing at the target site on the sickle cell.
[0076] Figure 8A-8G The experimental results describe the editing of adenosine (A) into guanine nucleoside (G) in the codon (CAC) of valine at amino acid position 6 encoding HbS, using a Staphylococcus aureus Cas9 variant (saKKH) resistant to NNNRRT, which is either alone or fused to a DNA-binding domain that is sequence-specific at the sickle cell target site. Figure 8A A schematic description of the ABE construct is presented, showing the organization of the domains within the polypeptide, including saKKH ABE7.10, saKKH ABE7.10 zf1ra, and saKKH ABE7.10zf1rb. Figure 8B The nucleic acid sequence shown at the target site on the sickle cell, and the target complementary sequence of the guide RNA, as described by the underline (specifying g1 and g4). Figure 8C This is a diagram illustrating the results of binding saKKH ABE7.10, saKKH ABE7.10 zf1ra, and saKKH ABE7.10 zf1rb to the guide RNA g1, which has a 20-nucleotide (nt) nucleic acid sequence that is complementary to the sickle cell target site. Figure 8C The right side shows the nucleic acid sequence at the target site on the sickle cell and the target complementary sequence of the g1 guide RNA. Figure 8D This is a diagram showing the results of binding saKKH ABE7.10, saKKH ABE7.10 zf1ra, and saKKH ABE7.10 zf1rb to the guide RNA g1, which has a 21nt nucleic acid sequence that is complementary to the sickle cell target site. Figure 8E This is a diagram illustrating the results of binding of saKKH ABE7.10, saKKH ABE7.10 zf1ra, and saKKH ABE7.10 zf1rb to the guide RNA g4, which has a 20 nt nucleic acid sequence that is complementary to the sickle cell target site. Figure 8E The right side shows the nucleic acid sequence at the target site on the sickle cell and the target complementary sequence of the g4 guide RNA. Figure 8FThis is a diagram illustrating the results of binding saKKH ABE7.10, saKKH ABE7.10 zf1ra, and saKKH ABE7.10 zf1rb to the guide RNA g4, which has a 21 nt nucleic acid sequence that is complementary to the sickle cell target site. Figure 8G Describes base editing at the control HEK2 site.
[0077] Figures 9A-9E This paper describes the development and evaluation of an adenosine base editor (ABE) with a thermophilic streptococcal Cas9 (St1Cas9) DNA binding domain for base editing at target sites on sickle cells. Figure 9A This shows base editing using ABE St1Cas9 with the typical PAM sequence NNAGAA (TTCTAG; reverse complementary) of St1Cas9. The inset below shows a comparison of ABE St1Cas9, St1Cas9 nuclease, and the percentage of untreated insertions / deletions (Indel%) at the base editing site. Figure 9B This shows base editing using BE St1Cas9 and the typical PAM sequence NNAGAA of St1Cas9. The inset below shows a comparison of ABESt1Cas9, St1Cas9 nucleases, and the percentage of untreated insertions and deletions at the base editing site. Figure 9C This shows base editing using ABE St1Cas9 with the St1Cas9 atypical PAM sequence NNACCA (TGGTNN; reverse complementary). The inset below shows a comparison of ABE St1Cas9, the St1Cas9 nuclease, and the percentage of untreated insertions and deletions at the base editing site. Figure 9D This shows base editing using BE St1Cas9 with the St1Cas9 atypical PAM sequence NNACCA (TGGTNN; reverse complementary). The inset below shows a comparison of ABE St1Cas9, the St1Cas9 nuclease, and the percentage of untreated insertions and deletions at the base editing site. Figure 9E The description describes base editing at the sickle cell target site using ABE St1Cas9 and the St1Cas9 atypical PAM sequence NNACCA. The arrow indicates that the A·T to G·C mutation (Val→Ala) was caused by the ABE-St1Cas9 base editor at the sickle cell target site in Hb.
[0078] Figure 10This describes the percentage of base editing performed at sickle cell target sites using an ABE with an SpCas9 DNA-binding domain, which has been evolved and engineered to accept NGC PAM (ngcABE). In this bar chart, the leftmost bar represents "Pro6Pro"; the middle bar represents "Val7Ala"; and the rightmost bar represents "Ser10Pro".
[0079] Figure 11 This is a schematic diagram representing the promoter region of the HBG1 / 2 gene. Individual purple triangles indicate SNPs and naturally occurring deletions in patients with HPFH. Green arrows (e.g., "BCL11A", "CCAAT", "90BCL11A", and "ZBTB7A") indicate potential transcription binding sites. Clusters of thick, pointed lines (pink) above and below the HBG1 / 2 sequence indicate guide RNAs that can target the regions of interest, such as the target sequence of the gene.
[0080] Figure 12 This figure shows the percentage of targeted base editing in the HBG1 / 2 gene in 293T cells transfected with the indicated gRNA and Cas9 base editor. Base editing efficiency percentages were determined using Miseq. The figure shows the percentage of editing occurring in 293T cells using each type of gRNA; the gene and target sequence are shown in Table 4. "Cs" indicates the location associated with the gRNA, where editing was performed using a CBE that binds to the gRNA. "As" indicates the location associated with the gRNA, where an ABE edits the sequence that binds to the individual gRNA.
[0081] Figure 13 The percentage of base editing performed by each type of gRNA in primary bone marrow CD34+ cells is indicated in Table 4, where the gene and target sequence are shown. CD34+ cells were transfected with the indicated gRNA and base editor. "Cs" indicates the location associated with the gRNA, where editing was performed using a CBE (such as BE4) that binds to the gRNA. "As" indicates the location associated with the gRNA, where ABE edits the target sequence that binds to the individual gRNA. The percentage of base editing at both the HBG1 and HBG2 loci was assessed using Miseq. Detailed Implementation
[0082] As described herein, the present invention is characterized by compositions and methods using a programmable nucleobase editor for replacing pathogenic amino acids. In a particular embodiment, the compositions and methods can be used to treat sickle cell disease caused by a Glu→Val mutation at the sixth amino acid of the β-globulin protein encoded by the HBB gene. Although there have been many advancements in the field of gene editing, precisely correcting the diseased HBB gene to restore Val→Glu remains unsolved and must still be achieved using CRISPR / Cas nucleases or CRISPR / Cas base editing methods.
[0083] Genome editing of the HBB gene using CRISPR / Cas nuclease methods to replace diseased nucleotides requires cutting the genome DNA. However, cutting genome DNA increases the risk of indels, which can lead to unwanted and undesirable consequences, including premature stop codons, altered codon reading frames, and more. Furthermore, double-strand breaks at the β-globulin locus can completely alter the locus via recombination events. The β-globulin locus contains a cluster of globulin genes (-5′-ε-; Gγ-; Aγ-; δ-; and β-globulin-3′), which has a sequence identical to another sequence. Due to the structure of the β-globulin locus, recombination repair of double-strand breaks within it can result in gene loss due to the insertion sequence between globulin genes (e.g., between the δ- and β-globulin genes). Unintended alterations to the locus also carry the risk of inducing thalassemia.
[0084] CRISPR / Cas base editing methods hold great promise because they are capable of producing precise changes at the nuclear base level. However, Val→Glu Precise correction requires a T·A to A·T transversion editor, which is not currently known to exist. Furthermore, the specificity of CRISPR / Cas base editing stems from the limited window of editable nucleotides formed by R-loops after CRISPR / Cas binds to DNA. Therefore, CRISPR / Cas targeting must occur at or near the sickle cell site to enable base editing, and optimal editing within this window may require additional sequence information.
[0085] One requirement for CRISPR / Cas targeting is the presence of pre-spacer adjacent motif (PAM) sequences flanking the target site. For example, many base editors are based on SpCas9, which requires the NGG PAM sequence. Even assuming T·A to A·T transversions are possible, no NGG PAM exists that would place the target "A" at the desired position for this SpCas9 base editor. Although many new CRISPR / Cas proteins have been discovered or developed, expanding the set of available PAMs, PAM requirements remain a limiting factor in the ability to guide a CRISPR / Cas base editor to a specific nucleotide at any location within the genome.
[0086] This invention is at least partly based on several findings described herein that address the aforementioned challenges in providing gene editing methods for treating sickle cell anemia. In one embodiment, the invention is partly based on the substitution of alanine for valine at amino acid position 6 of the Hb protein that causes sickle cell disease, thereby producing an Hb variant (Hb Makassar) that does not produce the sickle cell phenotype. Despite precise correction under a T·A to A·T transversion base editor. It is impossible. The results presented in this paper prove the following finding: Val→Ala The substitution (i.e., the Hb Makassar variant) can be generated using an A·T to G·C base editor (ABE). This is achieved by developing novel base editors and novel base editing strategies, as provided herein. For example, novel ABE base editors (i.e., those with an adenosine deaminase domain) are developed that utilize flanking sequences (e.g., PAM sequences; zinc finger binding sequences) at sickle cell target sites for optimal base editing.
[0087] This document provides and describes compositions and methods for editing the thymidine (T) base to cytidine (C) at the sixth amino acid codon (Sickle HbS; E6V) of a sickle cell disease variant of the β-globulin protein, thereby replacing the valine amino acid residue (V6A) at this amino acid position with an alanine amino acid residue. The substitution of valine at position 6 of HbS with alanine results in a β-globulin protein variant that does not exhibit the sickle cell phenotype (e.g., it does not have the potential to polymerize as in the pathogenic variant HbS). Therefore, the compositions and methods of this invention can be used to treat sickle cell disease.
[0088] This document provides and describes compositions and methods comprising base editors and base editor systems as described herein, for the treatment of diseases or conditions caused by or associated with genes provided in Tables 3A, 3B or 4 herein.
[0089] The following description and examples illustrate embodiments of the present invention in detail. It should be understood that the present invention is not limited to the specific embodiments described herein, and therefore variations are possible. Those skilled in the art will recognize that many variations and modifications exist within the scope of this invention.
[0090] All terms are intended to be understood in the manner of those skilled in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains.
[0091] The chapter titles used in this article are for organizational purposes only and should not be construed as limiting the topics mentioned.
[0092] Although various features of the invention may be described in the text of a single embodiment, these features may also be provided separately or in any suitable combination. Conversely, although the invention may be described in the text of individual embodiments for clarity, the invention may also be practiced in a single embodiment.
[0093] definition
[0094] Unless otherwise defined, all technical and scientific terms used herein have the meaning commonly understood by those skilled in the art to which this invention pertains. The following references provide general definitions of many terms used by those skilled in the art in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd edition, 1994); The Cambridge Dictionary of Science and Technology (Walker, ed., 1988); The Glossary of Genetics, 5th edition, R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991).
[0095] In this application, the use of the singular includes the plural unless otherwise specifically stated. It should be noted that, as used in this specification, the singular forms "a / an" and "the" include the plural references unless otherwise clearly indicated. In this application, the use of "or" means "and / or" unless otherwise stated. Furthermore, the use of the term "including," and other forms such as "include," "includes," and "included," is not restrictive.
[0096] As used in this specification and the claims, the terms "comprising" (and any form of inclusion, such as "comprise" and "comprises"), "having" (and any form of having, such as "have" and "has"), "including" (and any form of inclusion, such as "includes" and "include"), or "containing" (and any form of containing, such as "contains" and "contain") are inclusive or open-ended and do not exclude additional, undescribed elements or method steps. It should be considered that any embodiment discussed in this specification can be implemented with respect to any method or composition of the invention, and vice versa. Furthermore, the compositions of the invention can be used to implement the methods of the invention.
[0097] The term "about / approximately" means within an acceptable range of error for a given value, as determined by a person skilled in the art, which will depend in part on how the value was measured or determined, i.e., the limits of the measurement system. For example, "about" may mean within 1 or more standard deviations according to the practice of the art. Alternatively, "about" may mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Or, particularly with respect to biological systems or methods, the term may mean within an order of magnitude of a value, preferably within 5 times, and more preferably within 2 times. Where a given value is described in the claims of this application and the claims, unless otherwise stated, it should be assumed that the term "about" means within an acceptable range of error for the given value.
[0098] References to "some embodiments," "an embodiment," "an embodiment," or "other embodiments" in this specification mean that a particular feature, structure, or characteristic described in the embodiment is included in at least some embodiments of the invention, but not necessarily in all embodiments.
[0099] "Adenosine deaminase" refers to a polypeptide or fragment thereof that catalyzes the hydrolytic deamination of adenine or adenosine. In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine to inosine or deoxyadenosine to deoxyinosine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., engineered adenosine deaminases, evolved adenosine deaminases) can be derived from any organism, such as bacteria.
[0100] "Administer" as used herein refers to providing a patient or subject with one or more of the products or compositions described herein. By way of example and without limitation, the product or composition may be administered (e.g., by injection) via intravenous (iv), subcutaneous (sc), intradermal (id), intraperitoneal (ip), or intramuscular (im). One or more of these routes may be used. Non-oral administration may be by, for example, rapid bolus injection or by gradual infusion over time. Alternatively or concurrently, administration may be via oral route. Other administration methods are also contemplated, such as, without limitation: intranasal, rectal, intracranial, intravaginal, buccal, thoracic, intradermal, percutaneous, and the like.
[0101] "Reagent" refers to any small molecule chemical compound, antibody, nucleic acid molecule or polypeptide, or fragment thereof.
[0102] "Improvement" means to reduce, suppress, weaken, lower, stop, or stabilize the development or progression of a disease.
[0103] "Change" means a variation (increase or decrease) in the expression level or activity of a gene or polypeptide as detected by standard techniques known by methods such as those described herein. As used herein, a change includes a 10% change in expression level, a 25% change in preferred expression level, a 40% change in even better expression level, and a 50% or greater change in optimal expression level.
[0104] "Analogous" refers to molecules that are not identical but have similar functions or structural features. For example, a polypeptide analog retains the biological activity of the corresponding naturally occurring polypeptide but has certain biochemical modifications that enhance the function of the analog relative to the naturally occurring polypeptide. These biochemical modifications may increase the analog's protease resistance, membrane permeability, or half-life without altering, for example, ligand binding. Analogs may include non-natural amino acids.
[0105] "Base editor (BE)" or "nucleobase editor (NBE)" refers to a reagent that binds to polynucleotides and has nucleobase modification activity. In various embodiments, the base editor comprises a nucleobase-modifying polypeptide (e.g., a deaminase) and a polynucleotide-programmable nucleotide-binding domain that binds to a guide polynucleotide (e.g., guide RNA). In various embodiments, the reagent is a biomolecular complex comprising a protein domain having base-editing activity, i.e., a domain capable of modifying bases (e.g., A, T, C, G, or U) within a nucleic acid molecule (e.g., DNA). In some embodiments, the polynucleotide-programmable DNA-binding domain is fused to or linked to a deaminase domain. In one embodiment, the reagent is a fusion protein comprising a domain having base-editing activity. In another embodiment, the protein domain having base-editing activity is linked to the guide RNA (e.g., via an RNA-binding motif on the guide RNA and an RNA-binding domain fused to a deaminase). In some embodiments, the domain having base-editing activity can deaminate bases within a nucleic acid molecule. In some embodiments, the base editor deaminates bases within a DNA molecule. In some embodiments, the base editor deaminates cytosine (C) or adenosine (A) within the DNA. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenosine base editor (ABE). In some embodiments, an adenosine deaminase evolved from TadA. In some embodiments, the polynucleotide programmable DNA binding domain is a CRISPR-associated enzyme (e.g., Cas or Cpf1). In some embodiments, the base editor is a catalytically inactivating Cas9 (dCas9) fused to a deaminase domain. In some embodiments, the base editor is a Cas9 nickase (nCas9) fused to a deaminase domain. In some embodiments, the base editor is fused to a base excision repair (BER) inhibitor. In some embodiments, the base excision repair inhibitor is a uracil DNA glycosidase inhibitor (UGI). In some embodiments, the base excision repair inhibitor is an inosine base excision repair inhibitor. The detailed description of the base editor is found in international PCT applications PCT / 2017 / 045381 (WO 2018 / 027078) and PCT / US2016 / 058344 (WO 2017 / 070632), the full text of which is incorporated herein by reference.See also Komor, AC et al. "Programmable base editing of atarget base in genomic DNA without double-stranded DNA cleavage" Nature 533, 420-424 (2016); Gaudelli, NM et al. "Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage" Nature 551, 464-471 (2017); "Improved base excisionrepair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A baseeditors with higher efficiency and product purity" by Komor, AC et al. Science Advances3:eaao4774 (2017), and "Base editing: precision chemistry on the genome and transcriptome of living" by Rees, HA et al. cells." Nat Rev Genet. Dec 2018; 19(12):770-788. doi:10.1038 / s41576-018-0059-1, the full text of which is incorporated herein by reference.
[0106] "Cytidine deaminase" refers to a polypeptide or fragment thereof that catalyzes a deamination reaction that converts an amino group to a carbonyl group. In one embodiment, the cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine. PmCDA1, derived from the hagfish (Petromyzon marinus) (hagfish cytosine deaminase 1, "PmCDA1"), AID (activation-inducible cytidine deaminase; AICDA), derived from mammals (e.g., humans, pigs, cattle, horses, monkeys, etc.), and APOBEC are exemplary cytidine deaminases.
[0107] By way of example, the cytidine base editor BE4 has the following nucleic acid sequence. It also contains a polynucleotide sequence having at least 95% or higher identity to the BE4 nucleic acid sequence.
[0108]
[0109] The codon-optimized BE4 nucleic acid sequence is provided below:
[0110]
[0111] Another codon-optimized BE4 nucleic acid sequence (GeneArt, ThermoFisher Scientific) is provided below:
[0112]
[0113] "Base editing activity" refers to the chemical alteration of bases within a polynucleotide through action. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is cytidine deaminase activity, for example, converting a target C·G to T·A. In another embodiment, the base editing activity is adenosine deaminase activity, for example, converting A·T to G·C.
[0114] The term "base editor system" refers to a system for editing the nucleobases of a target nucleotide sequence. In various embodiments, the base editor (BE) system comprises: (1) a polynucleotide-programmable nucleotide-binding domain and a deaminase domain for deamination of the nucleobase; and (2) a guiding polynucleotide (e.g., guiding RNA) that binds to the polynucleotide-programmable nucleotide-binding domain. In some embodiments, the base editor system comprises: (1) a base editor (BE) comprising a polynucleotide-programmable DNA-binding domain and a deaminase domain for deamination of the nucleobase; and (2) guiding RNA that binds to the polynucleotide-programmable DNA-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable DNA-binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE).
[0115] "β-globulin (HBB) protein" means a polypeptide or fragment thereof that has at least about 95% amino acid sequence identity with the amino acid sequence of NCBI accession number NP_000509. In a particular embodiment, the β-globulin protein includes one or more alterations related to the reference sequence below. In a particular embodiment, the β-globulin protein associated with sickle cell disease includes an E6V (also known as E7V) mutation. An exemplary β-globulin amino acid sequence (e.g., a reference sequence) is provided below.
[0116] 1mvhltpeeks avtalwgkvn vdevggealg rllvvypwtq rffesfgdls tpdavmgnpk
[0117] 61vkahgkkvlg afsdglahld nlkgtfatls elhcdklhvd penfrllgnv lvcvlahhfg
[0118] 121keftppvqaa yqkvvagvan alahkyh
[0119] "HBB polynucleotide" refers to a nucleic acid molecule that encodes a β-globulin protein or a fragment thereof. An example HBB polynucleotide sequence (obtainable from NCBI accession number NM_000518) is provided below:
[0120]
[0121] "HBG1 protein" (i.e., Homo sapiens heme subunit γ-1 (HBG1) protein) means a polypeptide or fragment thereof having at least about 95% amino acid sequence identity with the amino acid sequence of NCBI reference sequence number NM_000559.2. In some embodiments, the HBG1 protein may include one or more alterations related to the following amino acid sequence. In a particular embodiment, editing of the regulatory region (e.g., promoter) associated with the HBG1 protein is performed to treat or improve sickle cell disease, as described herein. An exemplary HBG1 amino acid sequence is provided below:
[0122] MGHFTEEDKATITSLWGKVNVEDAGGETLGRLLVVYPWTQRFFDSFGNLSSASAIMGNPKVKAHGKKVLTSLGDATKHLDDLKGTFAQLSELHCDKLHVDPENFKLLGNVLVTVLAIHFGKEFTPEVQASWQKMVTAVASALSSRYH.
[0123] "HBG1 polynucleotide" refers to a nucleic acid molecule that encodes the HBG1 protein or a fragment thereof. The following is an example nucleic acid sequence of an HBG1 polynucleotide:
[0124]
[0125] "HBG2 protein" (i.e., Homo sapiens heme subunit γ-2 (HBG2) protein) means a polypeptide or fragment thereof having at least about 95% amino acid sequence identity with the amino acid sequence of NCBI reference sequence number NM_000184.3. In some embodiments, the HBG2 protein may include one or more changes related to the following amino acid sequence. In a particular embodiment, editing of regulatory regions (e.g., promoters) associated with the HBG2 protein is performed to treat or improve sickle cell disease, as described herein. An exemplary HBG2 amino acid sequence is provided below:
[0126] MGHFTEEDKATITSLWGKVNVEDAGGETLGRLLVVYPWTQRFFDSFGNLSSASAIMGNPKVKAHGKKVLTSLGDAIKHLDDLKGTFAQLSELHCDKLHVDPENFKLLGNVLVTVLAIHFGKEFTPEVQASWQKMVTGVASALSSRYH
[0127] "HBG2 polynucleotide" refers to a nucleic acid molecule that encodes the HBG2 protein or a fragment thereof. An example nucleic acid sequence of an HBG2 polynucleotide is provided below:
[0128]
[0129] "ALAS1 protein" (i.e., Homo sapiens 5′-aminolevulinic acid synthase 1 (ALAS1) protein) means a polypeptide or fragment thereof having at least about 95% amino acid sequence identity with the amino acid sequence of NCBI reference sequence number NM_000688.6. In some embodiments, the ALAS1 protein may include one or more changes related to the following amino acid sequence. In a particular embodiment, editing of the regulatory region (e.g., promoter) associated with the ALAS1 protein is performed to treat or improve sickle cell disease, as described herein. An exemplary ALAS1 amino acid sequence is provided below:
[0130] MESVVRRCPFLSRVPQAFLQKAGKSLLFYAQNCPKMMEVGAKPAPRALSTAAVHYQQIKETPPASEKDKTAKAKVQQTPDGSQQSPDGTQLPSGHPLPATSQGTASKCPFLAAQMNQRGSSVFCKASLELQEDVQEMNAVRKEVAETSAGPSVVSVKTDG GDPSGLLKNFQDIMQKQRPERVSHLLQDNLPKSVSTFQYDRFFEKKIDEKKNDHTYRVFKTVNRRAHIFPMADDYSDSLITKKQVSVWCSNDYLGMSRHPRVCGAVMDTLKQHGAGAGGTRNISGTSKFHVDLERELADLHGKDAALLFSSCFVANDSTL FTLAKMMPGCEIYSDSGNHASMIQGIRNSRVPKYIFRHNDVSHLRELLQRSDPSVPKIVAFETVHSMDGAVCPLEELCDVAHEFGAITFVDEVHAVGLYGARGGGIGDRDGVMPKMDIISGTLGKAFGCVGGYIASTSSLIDTVRSYAAGFIFTTSLPPM LLAGALESVRILKSAEGRVLRRQHQRNVKLMRQMLMDAGLPVVHCPSHIIPVRVADAAKNTEVCDELMSRHNIYVQAINYPTVPRGEELLRIAPTPHHTPQMMNYFLENLLVTWKQVGLELKPHSSAECNFCRRPLHFEVMSEREKSYFSGLSKLVSAQA
[0131] "ALAS1 polynucleotide" refers to a nucleic acid molecule that encodes the ALAS1 protein or a fragment thereof. An example nucleic acid sequence of an ALAS1 polynucleotide is provided below:
[0132]
[0133]
[0134] "BCL11A protein" (i.e., Homo sapiens B-cell CLL / lymphoma 11A (BCL11A) protein) (zinc finger protein) means a polypeptide or fragment thereof having at least about 95% amino acid sequence identity with the amino acid sequence of GenBank accession number ADL_14508.1. In some embodiments, the BCL11A protein may include one or more alterations related to the following amino acid sequences. In a particular embodiment, base editing occurs in a regulatory region (e.g., promoter) associated with the BCL11A protein to treat or improve diseases such as β-thalassemia and sickle cell disease (SCD), for example, by increasing fetal heme production. The BCL11A-encoding gene is highly expressed in several hematopoietic profiles and plays an important role in the transition from γ-globulin expression to β-globulin expression during the transition from fetal to adult erythropoietin. BCL11A may play an important role in inhibiting fetal heme production. It can also be involved in lymphoma pathogenicity; translocation deregulation of BCL11A expression has been found to be associated with B-cell malignancy. An exemplary human BCL11A amino acid sequence is provided below:
[0135] MSRRKQGKPQHLSKREFSPEPLEAILTDDEPDHGPLGAPEGDHDLLTCGQCQMNFPLGDILIFIEHKRKQCNGSLCLEKAVDKPPSPSPIEMKKASNPVEVGIQ VTPEDDDCLSTSSRGICPKQEHIADKLLHWRGLSSPRSAHGALIPTPGMSAEYAPQGICKDEPSSYTCTTCKQPFTSAWFLLQHAQNTHGLRIYLESEHGSPLT PRVGIPSGLGAECPSQPPLHGIHIADNNPFNLLRIPGSVSREASGLAEGRFPPTPPLFSPPPRHHLDPHRIERLGAEEMALATHHPSAFDRVLRLNPMAMEPPAMDFSRRLRELAGNTSSPPLSPGRPSPMQRLLQPFQPGSKPPFLATPPLPPLQSAPPPSQPPVKSKSCEFCGKTFKFQSNLVVHRRSHTGEKPYKCNLCDHACTQA SKLKRHMKTHMHKSSPMTVKSDDGLSTASSPEPGTSDLVGSASSALKSVVAKFKSENDPNLIPENGDEEEEEDDEEEEEEEEEELTESERVDYGFGLSLEAARHHENSSRGAVVGVGDESRALPDVMQGMVLSSMQHFSEAFHQVLGEKHKRGHLAEAEGHRDTCDEDSVAGESDRIDDGTVNGRGCSPGESASGGLSKKLLLGSP SSLSPFSKRIKLEKEFDLPPAAMPNTENVYSQWLAGYAASRQLKDPFLSFGDSRQSPFASSSEHSSENGSLRFSTPPGELDGGISGRSGTGSGGSTPHISGPGP GRPSSKEGRRSDTCEYCGKVFKNCSNLTVHRRSHTGERPYKCELCNYACAQSSKLTRHMKTHGQVGKDVYKCEICKMPFSVYSTLEKHMKKWHSDRVLNNDIKTE
[0136] "BCL11A polynucleotide" refers to a nucleic acid molecule encoding the BCL11A protein or a fragment thereof. An exemplary nucleic acid sequence of the human BCL11A (isomorph 1) polynucleotide (reference sequence number GU324937.1) is provided below:
[0137]
[0138] In some embodiments, the nuclease base editor system may include more than one base editing component. For example, the nuclease base editor system may include more than one deaminase. In some embodiments, the nuclease base editor system may include one or more cytidine deaminases and / or one or more adenosine deaminases. In some embodiments, a single guide polynucleotide may be used to target different deaminases to the target nucleic acid sequence. In some embodiments, a single pair of guide polynucleotides may be used to target different deaminases to the target nucleic acid sequence.
[0139] The nucleobase component of the base editor system and the polynucleotide programmable nucleotide binding component can be covalently or non-covalently associated with each other. For example, in some embodiments, a deaminase domain can be targeted to a target nucleotide sequence via a polynucleotide programmable nucleotide binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain can be fused to or linked to the deaminase domain. In some embodiments, the polynucleotide programmable nucleotide binding domain can target the deaminase domain to a target nucleotide sequence by non-covalently interacting with or associating with the deaminase domain. For example, in some embodiments, the nucleobase editing component (e.g., the deaminase component) may include an additional heterologous portion or domain that can interact with, associate with, or form a complex with an additional heterologous portion or domain, which is a portion of the polynucleotide programmable nucleotide binding domain. In some embodiments, the additional heterologous portion can bind to, interact with, associate with, or form a complex with a polypeptide. In some embodiments, the additional heterologous portion can bind to, interact with, associate with, or form a complex with a polynucleotide. In some embodiments, the additional heterologous portion may bind to a guide polynucleotide. In some embodiments, the additional heterologous portion may bind to a polypeptide linker. In some embodiments, the additional heterologous portion may bind to a polynucleotide linker. The additional heterologous portion may be a protein domain. In some embodiments, the additional heterologous portion may be a K homologous gene (KH) domain, an MS2 sheath protein domain, a PP7 sheath protein domain, an SfMuCom sheath protein domain, a sterile α motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif.
[0140] The base editor system may further include a guide polynucleotide component. It should be understood that the components of the base editor system may be associated with each other via covalent bonds, non-covalent interactions, or any combination of their associated and inter-interactions. In some embodiments, a deaminase domain may target a target nucleotide sequence via a guide polynucleotide. For example, in some embodiments, the nucleobase editing component (e.g., a deaminase component) of the base editor system may include an additional heterologous moiety or domain (e.g., a polynucleotide-binding domain, such as an RNA or DNA-binding protein) that may interact with, associate with, or form a complex with a portion or fragment of the guide polynucleotide (e.g., a polynucleotide motif). In some embodiments, the additional heterologous moiety or domain (e.g., a polynucleotide-binding domain, such as an RNA or DNA-binding protein) may be fused to or linked to the deaminase domain. In some embodiments, the additional heterologous moiety may bind to, interact with, associate with, or form a complex with a polypeptide. In some embodiments, the additional heterologous moiety may bind to, interact with, associate with, or form a complex with a polynucleotide. In some embodiments, the additional heterologous portion may bind to a guide polynucleotide. In some embodiments, the additional heterologous portion may bind to a polypeptide linker. In some embodiments, the additional heterologous portion may bind to a polynucleotide linker. The additional heterologous portion may be a protein domain. In some embodiments, the additional heterologous portion may be a K homologous gene (KH) domain, an MS2 sheath protein domain, a PP7 sheath protein domain, an SfMuCom sheath protein domain, a sterile α motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif.
[0141] In some embodiments, the base editor system may further include an inhibitor of the base excision repair (BER) component. It should be understood that the components of the base editor system may be associated with each other via covalent bonds, non-covalent interactions, or any combination of their association and interaction. The inhibitor of the BER component may include a base excision repair inhibitor. In some embodiments, the base excision repair inhibitor may be a uracil DNA glycosidase inhibitor (UGI). In some embodiments, the base excision repair inhibitor may be an inosine base excision repair inhibitor. In some embodiments, the base excision repair inhibitor may be targeted to the target nucleotide sequence via the polynucleotide-programmable nucleotide binding domain. In some embodiments, the polynucleotide-programmable nucleotide binding domain may be fused to or linked to the base excision repair inhibitor. In some embodiments, the polynucleotide-programmable nucleotide binding domain may be fused to or linked to a deaminase domain and the base excision repair inhibitor. In some embodiments, the polynucleotide-programmable nucleotide binding domain may target the base excision repair inhibitor to the target nucleotide sequence by non-covalently interacting with or associating with the base excision repair inhibitor. For example, in some embodiments, the inhibitor of the base excision repair component may include an additional heterologous portion or domain that can interact with, associate with, or form a complex with the additional heterologous portion or domain, wherein the additional heterologous portion or domain is a portion of a polynucleotide programmable nucleotide binding domain. In some embodiments, the inhibitor of base excision repair can target the target nucleotide sequence by guiding the polynucleotide. For example, in some embodiments, the inhibitor of base excision repair may include an additional heterologous portion or domain (e.g., a polynucleotide binding domain, such as an RNA or DNA-binding protein) that can interact with, associate with, or form a complex with a portion or fragment of the guiding polynucleotide (e.g., a polynucleotide motif). In some embodiments, the additional heterologous portion or domain of the guiding polynucleotide (e.g., a polynucleotide binding domain, such as an RNA or DNA-binding protein) may be fused to or linked to the inhibitor of base excision repair. In some embodiments, the additional heterologous portion may bind to, interact with, associate with, or form a complex with the polynucleotide. In some embodiments, the additional heterologous portion may bind to the guiding polynucleotide. In some embodiments, the additional heterologous portion may bind to a polypeptide linker. In some embodiments, the additional heterologous portion may bind to a polynucleotide linker. The additional heterologous portion may be a protein domain. In some embodiments, the additional heterologous portion may be a K homologous gene (KH) domain, an MS2 sheath protein domain, a PP7 sheath protein domain, an SfMuCom sheath protein domain, a sterile α motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif.
[0142] The term "Cas9" or "Cas9 domain" refers to an RNA-guided nuclease containing the Cas9 protein or a fragment thereof (e.g., a protein containing an activating, inactivating, or partially activated DNA-cutting domain of Cas9, and / or a gRNA-binding domain of Cas9). Cas9 nucleases sometimes also refer to casnl nucleases or CRISPR (clustered regularly spaced short palindromic repeats)-related nucleases. An exemplary Cas9 is Streptococcus pyogenes Cas9, whose amino acid sequence is provided below:
[0143] (Single bottom line: HNH domain; Double bottom line: RuvC domain).
[0144] The term "conserved amino acid substitution" or "conserved mutation" refers to the substitution of one amino acid for another that shares common characteristics. A feasible way to define the common characteristics among individual amino acids is to analyze the normalized frequencies of amino acid changes between corresponding proteins in homologous organisms (Schulz, GE and Schirmer, RH, Principles of Protein Structure, Springer-Verlag, New York (1979)). Based on such analyses, groups of amino acids can be defined where amino acids within a group preferentially exchange with each other and are therefore most similar in their effect on the overall protein structure (Schulz, GE and Schirmer, RH, as above). Unrestricted examples of conserved mutations include amino acid substitutions, such as replacing arginine with lysine and vice versa, thus maintaining a positive charge; replacing aspartic acid with glutamic acid and vice versa, thus maintaining a negative charge; replacing threonine with serine, thus maintaining a free –OH group; and replacing asparagine with glutamic acid, thus maintaining a free –-NH2 group.
[0145] In this document, the terms "coding sequence" or "protein-coding sequence" used interchangeably refer to a segment of polynucleotides that encodes a protein. This region or sequence is delimited by a start codon near the 5' end and by a stop codon near the 3' end. The coding sequence may also be referred to as an open reading frame.
[0146] As used herein, the term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes the hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. In some embodiments, the deaminase or deaminase domain is a cytosine deaminase that catalyzes the hydrolytic deamination of cytosine to uracil. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or adenine (A) to inosine (I). In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., engineered adenosine deaminases, evolved adenosine deaminases) can be derived from any organism, such as bacteria. In some embodiments, the adenosine deaminase is derived from bacteria such as *Escherichia coli*, *Staphylococcus aureus*, *Salmonella typhi*, *Streptococcus putrefaciens*, *Haemophilus influenzae*, or *C. crescentus*. In some embodiments, the adenosine deaminase is a TadA deaminase. In some embodiments, the deaminase or deaminase domain is a variant of a spontaneously occurring deaminase derived from an organism such as a human, chimpanzee, gorilla, monkey, cattle, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain is not spontaneously occurring. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical to spontaneously occurring deaminase. For example, the deaminase domain is described in international PCT applications PCT / 2017 / 045381 (WO 2018 / 027078) and PCT / US2016 / 058344 (WO 2017 / 070632), each of which is incorporated herein by reference in its entirety.See also Komor, AC et al. "Programmable base editing of atarget base in genomic DNA without double-stranded DNA cleavage" Nature 533, 420-424 (2016); Gaudelli, NM et al. "Programmable base editing of A·Tto G·C ingenomic DNA without DNA cleavage" Nature 551, 464-471 (2017); "Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T: Abase editors with higher efficiency and product purity" Science Advances 3:eaao4774 (2017) by Komor, AC et al. and "Base editing: precisionchemistry on the genome and transcriptome of living cells." Nat by Rees, HA et al. Rev Genet. Dec 2018; 19(12):770-788. doi:10.1038 / s41576-018-0059-1, the full text of which is incorporated herein by reference.
[0147] "Detectable tags" refer to a combination of methods that, when linked to a molecule of interest, enable the molecule to be detected by spectroscopic, photochemical, biochemical, immunochemical, or chemical means. Examples of usable tags include radioactive isotopes, magnetic beads, metal beads, colloidal particles, fluorescent dyes, electron-dense reagents, enzymes (e.g., those commonly used in ELISA), biotin, foxglove ligand, or incomplete antigens.
[0148] "Disease" means any symptom or condition that impairs or interferes with the normal function of cells, tissues, or organs. Examples of diseases include: retinitis pigmentosa, Usher syndrome, sickle cell disease, β-thalassemia, hereditary persistent hemoglobinemia of fetus (HPFH), α-1 antitrypsin deficiency (A1AD), hepatic porphyria, medium-chain acyl-CoA dehydrogenase (ACADM) deficiency, lysosomal acid lipase (LAL) deficiency, phenylketonuria, hemochromatosis, Von Gilke's disease, Pompe disease, Gaucher's disease, Hollerithia disease, cystic fibrosis, or chronic pain. In one embodiment, the disease is A1AD. In one embodiment, the disease is sickle cell disease (SCD), also known as "sickle cell anemia".
[0149] "Effective amount" refers to the amount of reagent or activating compound (e.g., a base editor as described herein) required to improve the symptoms of a disease in a subject or patient in need, relative to an untreated patient or disease-free individual (i.e., a healthy individual). The effective amount of the activating compound used to carry out the methods for treating the disease varies depending on the administration method, the subject's age, weight, and overall health. Ultimately, the attending physician or veterinarian will determine the appropriate amount and dosage regimen. This amount is referred to as the "effective" amount. In one embodiment, the effective amount is the amount of the base editor of the present invention, which is sufficient to cause changes in the genes of interest in cells (e.g., cells in vitro or in vivo). In one implementation, the effective amount is the amount of base editor required to achieve a therapeutic effect (e.g., reducing or controlling retinitis pigmentosa, Usher syndrome, sickle cell disease (SCD), β-thalassemia, hereditary persistent fetal hemoglobinemia (HPFH), α-1 antitrypsin deficiency (A1AD), hepatic porphyria, medium-chain acyl-CoA dehydrogenase (ACADM) deficiency, lysosomal acid lipase (LAL) deficiency, phenylketonuria, hemochromatosis, Von Gilke's disease, Pompe disease, Gaucher's disease, Hollerithiasis, cystic fibrosis, or chronic pain). This therapeutic effect does not need to be sufficient to alter the pathogenic gene in all cells of the subject, tissue, or organ, but only to alter the pathogenic gene. The pathogenic gene is present in approximately 1%, 5%, 10%, 25%, 50%, 75%, or higher of the cells in the test subject, tissue, or organ. In one embodiment, the effective amount is sufficient to improve one or more symptoms of a disease (e.g., retinitis pigmentosa, Usher syndrome, sickle cell disease (SCD), β-thalassemia, hereditary persistent fetal hemoglobinemia (HPFH), α-1 antitrypsin deficiency (A1AD), hepatic porphyria, medium-chain acyl-CoA dehydrogenase (ACADM) deficiency, lysosomal acid lipase (LAL) deficiency, phenylketonuria, hemochromatosis, Von Gilke's disease, Pompe disease, Gaucher's disease, Hollerithia disease, cystic fibrosis, or chronic pain).
[0150] "Fragment" refers to a portion of a polypeptide or nucleic acid molecule. This portion preferably contains at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the full length of a reference nucleic acid molecule or polypeptide. A fragment may contain 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids.
[0151] "Hybridization" refers to hydrogen bonds between complementary nucleobases, which can be Watson-Crick, Hoogsteen, or reverse Hoogsteen hydrogen bonds. For example, adenine and thymine are complementary nucleobases that pair up via hydrogen bonds.
[0152] The term "base repair inhibitor" or "IBR" refers to a protein that inhibits the activity of nucleic acid repair enzymes, such as base excision repair enzymes. In some embodiments, the IBR is an inosine base excision repair inhibitor. Exemplary inhibitors of base repair include inhibitors of APE1, Endo III, Endo IV, Endo V, Endo VIII, Fpg, hOGGl, hNEILl, T7 Endol, T4PDG, UDG, hSMUGl, and hAAG. In some embodiments, the IBR is an inhibitor of Endo V or hAAG. In some embodiments, the IBR catalyzes inactive Endo V or catalyzes inactive hAAG. In some embodiments, the base repair inhibitor is an inhibitor of Endo V or hAAG. In some embodiments, the base repair inhibitor catalyzes inactive Endo V or catalyzes inactive hAAG. In some embodiments, the base repair inhibitor is a uracil glycosidase inhibitor (UGI). UGI refers to a protein that inhibits uracil-DNA glycosidase base excision repair enzymes. In some embodiments, the UGI domain comprises wild-type UGI or wild-type UGI fragments. In some embodiments, the UGI protein provided herein comprises UGI fragments and proteins homologous to UGI or UGI fragments. In some embodiments, the base repair inhibitor is an inhibitor of inosine base excision repair. In some embodiments, the base repair inhibitor is a "catalytically inactive inosine-specific nuclease" or an "inactive inosine-specific nuclease." Without wishing to be limited by any particular theory, a catalytically inactive inosine glycosidase (e.g., alkyladenine glycosidase (AAG)) may bind inosine but may not create a debasement site or remove inosine, thereby stereotactically preventing the formation of newly formed inosine moieties from being affected by DNA damage / repair mechanisms. In some embodiments, the catalytically inactive inosine-specific nuclease may bind inosine in nucleic acids but may not cleave the nucleic acids. Non-limiting exemplary catalytically inactive inosine-specific nucleases include catalytically inactive alkyladenine glycosidases (AAG nucleases), for example, from humans, and catalytically inactive endonuclease V (Endo V nucleases), for example, from *Escherichia coli*. In some implementations, the catalytically inactive AAG nuclease contains an E125Q mutation or a corresponding mutation in another AAG nuclease.
[0153] The terms "separation," "purification," or "biopurification" refer to the complete absence of components typically associated with a substance when found in its natural state, or the removal of said components to varying degrees. "Separation" indicates the degree of separation from its original source or surroundings. "Purification" indicates a degree of separation beyond separation. A "purified" or "biopurified" protein is completely free of other substances such that any impurities do not substantially affect the protein's biological properties or cause other adverse results. That is, the nucleic acids or peptides of the present invention are purified if they are substantially free of cellular material, viral material, or culture medium when manufactured using recombinant DNA technology, or substantially free of chemical precursors or other chemicals when chemically synthesized. Purity and homogeneity are typically determined using analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high-performance liquid chromatography. The term "purification" may indicate that the nucleic acid or protein produces essentially one band in the electrophoretic gel. For proteins with acceptable modifications (e.g., phosphorylation or glycosylation), different modifications can result in different isolated proteins that can be purified separately.
[0154] "Separated polynucleotide" means that the gene does not contain nucleic acids (e.g., DNA) flanking it, and that gene is in the naturally occurring genome of an organism derived from the nucleic acid molecule of the present invention. Therefore, this term includes, for example, recombinant DNA incorporated into a vector; incorporated into an autonomously replicating plasmid or virus; or incorporated into the genome DNA of a prokaryote or eukaryote; or existing as a separate molecule independent of other sequences (e.g., cDNA or genome, or a cDNA fragment produced by PCR, or digested by restriction endonucleases). Furthermore, this term includes RNA molecules transcribed from DNA molecules, and recombinant DNA portions of hybrid genes encoding additional polypeptide sequences.
[0155] "Isolated polypeptide" means a polypeptide of the present invention that has been separated from its naturally occurring accompanying components. Generally, a polypeptide is considered isolated when it contains at least 60% by weight of proteins and naturally occurring organic molecules naturally associated with it. Preferably, the preparation is at least 75%, more preferably at least 90%, and most preferably at least 99% (by weight) of the polypeptide of the present invention. For example, the isolated polypeptide of the present invention can be obtained by extraction from a natural source, by expression of a recombinant nucleic acid encoding this polypeptide, or by chemical synthesis of the protein. Purity can be measured by any suitable method, for example, column chromatography, polyacrylamide gel electrophoresis, or by HPLC analysis.
[0156] As used herein, the term "linker" can refer to a covalent linker (e.g., a covalent bond), a non-covalent linker, a chemical group, or a molecule that links two molecules or portions (e.g., two components of a protein complex or ribonucleoside complex, or two fused protein domains, such as a polynucleotide-programmable DNA-binding domain (e.g., dCas9) and a deaminase domain (e.g., adenosine deaminase or cytidine deaminase)). Linkers can conjugate different components of a base editor system, or different portions of a component. For example, in some embodiments, a linker can conjugate a guide polynucleotide-binding domain of a polynucleotide-programmable nucleotide-binding domain and a catalytic domain of a deaminase. In some embodiments, a linker can conjugate a CRISPR peptide and a deaminase. In some embodiments, a linker can conjugate Cas9 and a deaminase. In some embodiments, a linker can conjugate dCas9 and a deaminase. In some embodiments, a linker can conjugate nCas9 and a deaminase. In some embodiments, a linker can conjugate a guide polynucleotide and a deaminase. In some embodiments, the linker may conjugate the deamino group and the polynucleotide programmable nucleotide-binding group of a base editor system. In some embodiments, the linker may conjugate the RNA-binding portion of the deamino group and the polynucleotide programmable nucleotide-binding group of a base editor system. In some embodiments, the linker may conjugate the RNA-binding portion of the deamino group and the RNA-binding portion of the polynucleotide programmable nucleotide-binding group of a base editor system. The linker may be located between or on either side of two groups, molecules, or other parts, and may be connected to each other via covalent or non-covalent interactions, thus connecting the two. In some embodiments, the linker may be an organic molecule, group, polymer, or chemical part. In some embodiments, the linker may be a polynucleotide. In some embodiments, the linker may be a DNA linker. In some embodiments, the linker may be an RNA linker. In some embodiments, the linker may include an aptamer that can bind to a ligand. In some embodiments, the ligand may be a carbohydrate, peptide, protein, or nucleic acid. In some embodiments, the linker may include an aptamer that may be derived from a riboswitch. The aptamer derived from it may be selected from theophylline riboswitches, thiamium pyrophosphate (TPP) riboswitches, adenosylcobalamin (AdoCbl) riboswitches, S-adenosylmethionine (SAM) riboswitches, SAH riboswitches, flavin mononucleotide (FMN) riboswitches, tetrahydrofolate riboswitches, lysine riboswitches, glycine riboswitches, purine riboswitches, GlmS riboswitches, or pre-Q1 riboswitches. In some embodiments, the linker may comprise an aptamer that binds to a polypeptide or protein domain, such as a polypeptide ligand.In some embodiments, the polypeptide ligand may be a K homologous gene (KH) domain, an MS2 sheath protein domain, a PP7 sheath protein domain, an SfMuCom sheath protein domain, a sterile α motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the polypeptide ligand may be part of a base editing system component. For example, the nucleobase editing component may include a deaminase domain and an RNA recognition motif.
[0157] In some embodiments, the linker may be an amino acid or a plurality of amino acids (e.g., a peptide or protein). In some embodiments, the linker is about 5-100 amino acids in length, for example, about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, or 90-100 amino acids in length. In some embodiments, the linker may be about 100-150, 150-200, 200-250, 250-300, 300-350, 350-400, 400-450, or 450-500 amino acids in length. Longer or shorter linkers may also be considered.
[0158] In some embodiments, the linker conjugates the gRNA-binding domain (including the Cas9 nuclease domain) of an RNA-programmable nuclease and the catalytic domain of a nucleic acid-editing protein (e.g., cytidine or adenosine deaminase). In some embodiments, the linker conjugates dCas9 and the nucleic acid-editing protein. For example, the linker is located between or on either side of two groups, molecules, or other parts, and is covalently linked to each other, thus connecting the two. In some embodiments, the linker is an amino acid or a plurality of amino acids (e.g., a peptide or a protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical part. In some embodiments, the linker is 5-200 amino acids in length, for example: 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140, 150, 160, 175, 180, 190, or 200 amino acids. Longer or shorter linkers may also be considered. In some embodiments, the linker comprises the amino acid sequence GSETPGTSESATPES, which may also be referred to as the XTEN linker. In some embodiments, the linker comprises the amino acid sequence SGGS. In some embodiments, the linker comprises: (SGGS)n 、(GGGS) n (GGG GS) n 、(G) n 、(EAAAK) n 、(GGS) n , SGSETPGTSESATPES or (XP) n A motif, or any combination thereof, wherein n is independently an integer between 1 and 30, and wherein X is any amino acid. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. In some embodiments, the linker contains a plurality of proline residues and is 5-21, 5-14, 5-9, or 5-7 amino acids in length, for example: PAPAP, PAPAPA, PAPAPAP, PAPAPAPA, P(AP)4, P(AP)7, P(AP) 10 This proline-rich linker is also known as a "rigid" linker.
[0159] In some embodiments, the base editing domain is fused via a linker comprising the following amino acid sequence: SGGSSGSETPGTSESATPESSGGS, SGGSSGGSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS. In some embodiments, the base editing domain is fused via a linker comprising the following amino acid sequence: SGSETPGTSESATPES, which may also be referred to as an XTEN linker. In some embodiments, the linker is 24 amino acids long. In some embodiments, the linker comprises the following amino acid sequence: SGGSSGGSSGSETPGTSESATPES. In some embodiments, the linker is 40 amino acids long. In some embodiments, the linker comprises the following amino acid sequence: SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS. In some embodiments, the linker is 64 amino acids long. In some embodiments, the linker comprises the following amino acid sequence:
[0160] SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGSSGGS. In some embodiments, the linker is 92 amino acids in length. In some embodiments, the linker contains the following amino acid sequence:
[0161] PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS.
[0162] As used herein, the term "mutation" refers to the substitution of a residue within a sequence (e.g., a nucleic acid or amino acid sequence) with another residue, or the deletion or insertion of one or more residues within a sequence. A mutation is generally described herein by identifying the original residue, followed by the position of the residue within the sequence, and then the identity of the newly substituted residue. Various methods for performing the amino acid substitutions (mutations) described herein are well known in the art and are provided, for example, by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)). In some embodiments, the base editors disclosed herein can efficiently produce "desired mutations," such as point mutations, in nucleic acids (e.g., nucleic acids in the genome of a subject) without producing a large number of unwanted mutations, such as unwanted point mutations. In some embodiments, the desired mutation is a mutation produced by a specific base editor (e.g., a cytidine base editor or an adenosine base editor) that binds to a guide polynucleotide (e.g., gRNA) specifically designed to produce the desired mutation.
[0163] Generally, mutations made or identified in a sequence (e.g., an amino acid sequence as described herein) are associated with a reference (or wild-type) sequence number (i.e., a sequence without the mutation). Those skilled in the art will readily understand how to determine the location of the mutation in the amino acid and nucleic acid sequence associated with the reference sequence.
[0164] The terms "nuclear localization sequence," "nuclear localization signal," or "NLS" refer to an amino acid sequence that facilitates the introduction of a protein into the cell nucleus. Nuclear localization sequences are known in the art and are described, for example, in Plank et al.'s international PCT application PCT / EP2000 / 011690, filed November 23, 2000, and published May 31, 2001 as WO / 2001 / 038547, the contents of which are incorporated herein by reference to disclose exemplary nuclear localization sequences. In other embodiments, the NLS is the optimized NLS described above, for example, Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some implementations, the NLS contains the following amino acid sequence: KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.
[0165] The terms “nucleobase,” “nitrogenous base,” or “base,” used interchangeably herein, refer to nitrogenous biological compounds that form nucleosides, which are then components of nucleotides. The ability of nucleobases to form base pairs and stack on one another directly results in long-chain helical structures, such as ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). The five nucleobases—adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U)—are referred to as primary or typical. Adenine and guanine are derived from purines, while cytosine, uracil, and thymine are derived from pyrimidines. DNA and RNA may also contain other (non-primary) modified bases. Non-limiting exemplary modifications of nucleobases may include hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydromethylcytosine. Hypoxanthine and xanthine can be produced by the presence of a mutagen, both undergoing deamination (the amino group is replaced by a carbonyl group). Hypoxanthine can be modified from adenine. Xanthine is modified from guanine. Uracil can be prepared by deaminating cytosine. A nucleoside consists of a nucleobase and five sugars (ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxyuridine, and deoxycytidine. Examples of nucleosides with modified nucleobases include inosine (I), xanthine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), and pseudouridine (Ψ). A nucleotide consists of a nucleobase, five sugars (ribose or deoxyribose), and at least one phosphate group.
[0166] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to compounds comprising nucleobases and acidic portions, such as nucleosides, nucleotides, or polymers of nucleotides. Typically, polymeric nucleic acids (e.g., nucleic acid molecules comprising three or more nucleotides) are linear molecules in which adjacent nucleotides are linked together by phosphodiester bonds. In some embodiments, "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides). In some embodiments, "nucleic acid" refers to oligonucleotide chains comprising three or more individual nucleotide residues. As used herein, the terms "oligonucleotide," "polynucleotide," and "polynucleic acid" are used interchangeably to refer to polymers of nucleotides (e.g., strings of at least three nucleotides). In some embodiments, "nucleic acid" comprises RNA and single-stranded and / or double-stranded DNA. Nucleic acids can be naturally occurring, for example, in the range of genomic bodies, transcripts, mRNA, tRNA, rRNA, siRNA, snRNA, plastids, cohesive plastids, chromosomes, chromatids, or other naturally occurring nucleic acid molecules. On the other hand, nucleic acid molecules can be non-naturally occurring molecules, such as recombinant DNA or RNA, artificial chromosomes, engineered genomes or fragments thereof, or synthetic DNA, RNA, DNA / RNA hybrids, or molecules containing non-naturally occurring nucleotides or nucleosides. Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms include nucleic acid analogs, such as analogs having a phosphodiester backbone. Nucleic acids can be purified from natural sources, manufactured using recombinant expression systems and purified as appropriate, chemically synthesized, etc. When appropriate, for example, in the case of chemically synthesized molecules, nucleic acids may contain nucleoside analogs, such as analogs having chemically modified bases or sugars, and backbone modifications. Nucleic acid sequences are presented in a 5′ to 3′ orientation unless otherwise specified. In some embodiments, the nucleic acid is a natural nucleoside or contains a natural nucleoside (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); or a nucleoside analogue (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyluridine, C5-propynylcytidine). C5-methylcytidine, 2-aminoadenosine, 7-deadenosine, 7-deadenosine, 8-oxadenosine, 8-oxaguanosine, O(6)-methylguanine and 2-thiocytidine); chemically modified bases; biologically modified bases (e.g., methylated bases); intercalated bases; modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose and hexose); and / or modified phosphate groups (e.g., thiophosphate and 5'-N-phosphorylamine linkage).
[0167] The terms "nucleic acid programmable DNA-binding protein" or "napDNAbp" are used interchangeably with "polynucleotide programmable nucleotide-binding domain" to refer to a protein associated with a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid, which directs the napDNAbp to a specific nucleic acid sequence. For example, a Cas9 protein may be associated with a guide RNA that directs the Cas9 protein to a specific DNA sequence complementary to that guide RNA. In some embodiments, the napDNAbp is a Cas9 domain, such as a nuclease-activated Cas9, a Cas9 nickase (nCas9), or a nuclease-inactive Cas9 (dCas9). In some embodiments, the Cas9 domain comprises any of the amino acid sequences as presented herein. In some embodiments, the Cas9 domain contains at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the same amino acid sequence as any of the amino acid sequences presented herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the amino acid sequences presented herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical consecutive amino acid residues compared to any of the amino acid sequences presented herein.
[0168] Examples of nucleic acid-programmable DNA-binding proteins include (without limitation): Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i. Other nucleic acid-programmable DNA-binding proteins are also within the scope of this invention, although they are not specifically listed herein. See, for example, Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J. Oct 2018; 1:325-336. doi:10.1089 / crispr.2018.0033; Yan et al., “Functionally diverse type V CRISPR-Cassystems” Science. Jan 4, 2019; 363(6422):88-91. doi:10.1126 / science.aav7271, which are incorporated herein by reference in their entirety.
[0169] As used herein, the terms "nucleobase editing domain" or "nucleobase editing protein" refer to a protein or enzyme capable of catalyzing nucleobase modifications in RNA or DNA, such as deamination of cytosine (or cytidine) to uracil (or uridine) or thymine (or thymidine), and deamination of adenine (or adenosine) to hypoxanthine (or inosine), as well as the addition and insertion of non-template nucleotides. In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., cytidine deaminase, cytosine deaminase, adenine deaminase, or adenosine deaminase). In some embodiments, the nucleobase editing domain may be a naturally occurring nucleobase editing domain. In some embodiments, the nucleobase editing domain may be an engineered or evolved nucleobase editing domain derived from a naturally occurring nucleobase editing domain. The nucleobase editing domain may originate from any organism, such as bacteria, humans, chimpanzees, gorillas, monkeys, cattle, dogs, rats, or mice. For example, the description of nucleobase editing proteins is found in international PCT applications PCT / 2017 / 045381 (WO 2018 / 027078) and PCT / US2016 / 058344 (WO 2017 / 070632), the full text of which is incorporated herein by reference. See also Komor, AC et al., “Programmable editing of a target base ingenomic DNA without double-stranded DNA cleavage”, Nature 533, 420-424 (2016); Gaudelli, NM et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage”, Nature 551, 464-471 (2017); and Komor, AC et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity”, Science Advances 3:eaao4774 (2017), all of which are incorporated herein by reference in their entirety.
[0170] As used herein, “obtain” in “obtain a reagent” includes synthesizing, purchasing, isolating or otherwise acquiring the reagent.
[0171] As used herein, "patient" or "subject" refers to a mammalian subject or individual who has been diagnosed with a disease or condition, is at risk of developing or having a disease or condition, or is suspected of developing or having a disease or condition. In some embodiments, the term "patient" refers to a mammalian subject who has a higher-than-average probability of developing a disease or condition. Exemplary patients may be humans, non-human primates, cats, dogs, pigs, cattle, horses, camels, llamas, goats, sheep, rodents (e.g., mice, rabbits, rats, or guinea pigs), and other mammals that may benefit from the treatments disclosed herein. Exemplary human patients may be males and / or females.
[0172] "Patients in need" or "subjects in need" in this article refers to patients or subjects who have been diagnosed with a disease or condition, are at risk of having a disease or condition, or are suspected of having a disease or condition, such as (but not limited to) sickle cell disease (SCD) or α-1 antitrypsin deficiency (A1AD), or diseases or conditions related to genes listed in Tables 3A, 3B or 4 in this article.
[0173] The terms "pathogenic mutation," "pathogenic variant," "disease-causing (or disease-related) mutation," "disease-causing (or disease-related) variant," "adverse mutation," or "susceptibility mutation" refer to a genetic alteration or mutation that increases an individual's susceptibility or predisposition to certain diseases or conditions. In some embodiments, the pathogenic mutation comprises at least one wild-type amino acid substituted with at least one pathogenic amino acid in a protein encoded by a gene.
[0174] The term "non-conservative mutation" refers to an amino acid substitution between different groups, such as replacing tryptophan with lysine, or serine with phenylalanine, etc. In this case, the non-conservative amino acid substitution preferably does not interfere with or inhibit the biological activity of the functional variant. This non-conservative amino acid substitution can enhance the biological activity of the functional variant, resulting in increased biological activity compared to the wild-type protein.
[0175] The terms "protein," "peptide," "polypeptide," and their grammatical equivalents are used interchangeably herein and refer to a polymer of amino acid residues linked together by peptide (amide) bonds. The term refers to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide will be at least 3 amino acids long. A protein, peptide, or polypeptide can refer to a single protein or a collection of proteins. One or more amino acids in a protein, peptide, or polypeptide may be modified, for example by adding chemical entities such as: carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups, linkers for conjugation, functionalization, or other modifications, etc. Proteins, peptides, or polypeptides can also be a single molecule or a multi-molecule complex. A protein, peptide, or polypeptide can be simply a fragment of a naturally occurring protein or peptide. Proteins, peptides, or polypeptides can be naturally occurring, recombinant, or synthetic, or any combination thereof. As used herein, the term "fusion protein" refers to a hybrid polypeptide that contains protein domains from at least two different proteins. A protein may be located at the amino-terminal (N-terminal) portion of a fusion protein or at the carboxyl-terminal (C-terminal) portion, thus forming an amino-terminal fusion protein or a carboxyl-terminal fusion protein, respectively. The protein may contain various domains, such as a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9, which guides the protein to a target site) and a nucleic acid cleavage domain, or a catalytic domain of the nucleic acid used to edit the protein. In some embodiments, the protein comprises a protein moiety, such as an amino acid sequence constituting the nucleic acid binding domain, and an organic compound, such as a compound that can act as a nucleic acid cleavage agent. In some embodiments, the protein is complexed with or associated with a nucleic acid (e.g., RNA or DNA). Any protein provided herein can be manufactured by any method known in the art. For example, the proteins provided herein can be manufactured via recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for the expression and purification of recombinant proteins are also well known, including those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the full text of which is incorporated herein by reference.
[0176] The polypeptides and proteins (including functional parts and their functional variants) disclosed herein may contain synthetic amino acids that substitute for one or more naturally occurring amino acids. Such synthetic amino acids are known in the art and include, for example, aminocyclohexanecarboxylic acid, leucine, α-aminon-decanoic acid, homoserine, S-acetylaminomethylcysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine, β-phenylserine, β-hydroxyphenylalanine, phenylglycine, α-naphthylalanine, cyclohexylalanine, cyclohexylglycine, and indole. The polypeptide and protein contain phyll-2-carboxylic acid, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, aminomalonic acid monoacylamine, N'-benzyl-N'-methyl-lysine, N',N'-dibenzyl-lysine, 6-hydroxylysine, ornithine, α-aminocyclopentanecarboxylic acid, α-aminocyclohexanecarboxylic acid, α-aminocycloheptanecarboxylic acid, α-(2-amino-2-norborneane)-carboxylic acid, α,γ-diaminobutyric acid, α,β-diaminopropionic acid, homophenylalanine, and α-tertiary butylglycine. Post-translational modifications of one or more amino acids in the polypeptide construct may be associated with these polypeptides and proteins. Non-restricted examples of post-translational modifications include: phosphorylation, acylation including acetylation and formylation, glycosylation (including N-linking and O-linking), acylation, hydroxylation, alkylation including methylation and ethylation, ubiquitination, addition of pyrrolidone carboxylic acid, formation of disulfide bonds, sulfation, cardiacidation, palmitate esterification, isopreneation, farnesylation, geraniylation, glycosylphosphatidylinositylation, thioctylation, and iodination.
[0177] As used herein, the term "gene" refers to a polynucleotide, which typically comprises a protein-coding region and a protein-noncoding region. The protein-noncoding region may contain one or more regulatory elements. Non-restricted examples of such regulatory elements include: promoters, enhancers, repressors, silencers, insulators, start codons, stop codons, Kozak concordant sequences, splice acceptors, splice donors, 3' and / or 5' untranslated regions (UTRs), splice sites, or intergenic regions. In some embodiments, the regulatory element is located on a gene that is the cause of a genetic disease or condition. Non-restricted examples of such regulatory elements located on genes that are the cause of a genetic disease or condition include: start codons, stop codons, Kozak concordant sequences, intergenic regions, 3' UTRs, or 5' UTRs, etc. In some embodiments, the regulatory element is not located on a gene that is the cause of a genetic disease or condition. Non-restricted examples of regulatory elements not located on genes that are the cause of a genetic condition include: enhancers, repressors, or insulators, etc.
[0178] The term "polynucleotide-programmable nucleotide binding domain" refers to a protein associated with a nucleic acid (e.g., DNA or RNA), such as a guide polynucleotide (e.g., guide RNA), which directs the polynucleotide-programmable DNA binding domain to a specific nucleic acid sequence. In some embodiments, the polynucleotide-programmable nucleotide binding domain is a polynucleotide-programmable DNA binding domain. In some embodiments, the polynucleotide-programmable nucleotide binding domain is a polynucleotide-programmable RNA binding domain. In some embodiments, the polynucleotide-programmable nucleotide binding domain is a Cas9 protein. The Cas9 protein may be associated with a guide RNA that directs the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the polynucleotide-programmable nucleotide binding domain is a Cas9 domain, such as nuclease-activated Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Unrestricted examples of nucleic acid programmable DNA-binding proteins include: Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i.Unrestricted examples of Cas enzymes include: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa 5. Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas actor proteins, type V Cas actor proteins, type VI Cas actor proteins, CARF, DinG, their homologous genes, or their modified or engineered forms. Other nucleic acid-programmable DNA-binding proteins are also within the scope of this invention, although they are not explicitly listed herein.
[0179] As used herein in the context of proteins or nucleic acids, the term "recombinant" refers to a protein or nucleic acid that is not naturally occurring but is a product of human engineering. For example, in some embodiments, recombinant protein or nucleic acid molecules contain an amino acid or nucleotide sequence that contains at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations compared to any naturally occurring sequence.
[0180] "Reduction" means at least 10%, 25%, 50%, 75%, or 100% negative change.
[0181] "Reference" means standard or control conditions. By way of non-limiting example, after base editing (e.g., benign or regulatory base editing) as described herein, the test results of the activity or function of a gene (and / or its encoded protein product) are compared with the activity or function of a gene (and / or its encoded product) in which no benign or regulatory base editing has occurred, or with the activity or function of a wild-type gene (and / or its encoded product) serving as a reference. In one embodiment, the reference is a wild-type or healthy cell.
[0182] A "reference sequence" is a defined sequence used as the basis for sequence comparison. A reference sequence can be a subset or the entirety of a specific sequence; for example, a fragment of a full-length cDNA or gene sequence, or the complete cDNA or gene sequence. For polypeptides, the reference polypeptide sequence will typically be at least about 16 amino acids in length, preferably at least about 20 amino acids, more preferably at least about 25 amino acids, and even more preferably about 35 amino acids, about 50 amino acids, or about 100 amino acids. For nucleic acids, the reference nucleic acid sequence will typically be at least about 50 nucleotides in length, preferably at least about 60 nucleotides, more preferably at least about 75 nucleotides, and even more preferably about 100 nucleotides or about 300 nucleotides, or any integer around or between these values.
[0183] The terms "RNA-programmable nuclease" and "RNA-guided nuclease" are used with one or more RNAs that are not intended for cleavage (e.g., binding or associating with the target). In some embodiments, when an RNA-programmable nuclease is complexed with RNA, it may be referred to as a nuclease:RNA complex. Typically, the bound RNA is referred to as guide RNA (gRNA). Guide RNA (gRNA) may be a complex of two or more RNAs or exist as a single RNA molecule. gRNA existing as a single RNA molecule may be referred to as single-guide RNA (sgRNA), however, "gRNA" is used interchangeably to refer to guide RNA existing as a single molecule or as a complex of two or more molecules. Typically, gRNA existing as a single RNA species contains two domains: (1) a domain that shares homologous genes with the target nucleic acid (e.g., and guides the Cas9 complex to bind to the target); and (2) a domain that binds the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence known as tracrRNA and contains a stem-loop structure. For example, in some implementations, domain (2) is identical or homologous to the tracrRNA, as provided in Jinek et al., Science 337:816-821 (2012), the entire text of which is incorporated herein by reference. Examples of other gRNAs (e.g., including those of domain 2) can be found in U.S. Provisional Application No. 61 / 874,682, September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof", and U.S. Provisional Application No. 61 / 874,746, September 6, 2013, entitled "Delivery System for Functional Nucleases", each of which is incorporated herein by reference in its entirety. In some implementations, the gRNA comprises two or more domains (1) and (2), and may be referred to as an "expanded gRNA". For example, the expanded gRNA binds two or more Cas9 proteins and target nucleic acids that bind to two or more separate regions, as described herein. The gRNA contains a nucleotide sequence complementary to a target site that mediates the binding of the nuclease / RNA complex to the target site, providing sequence specificity of the nuclease:RNA complex.In some implementations, the RNA-programmable nuclease is a Cas9 intranuclease (CRISPR-related system), for example, Cas9 (Csnl) from Streptococcus pyogenes (see, for example, "Complete genome sequence of an Ml strain of Streptococcus pyogenes." Ferretti JJ, McShan W.M., Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded smallRNA and host factor RNase III." Deltchev E., Chylinski K., SharmaCM., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature471:602-607 (2011).
[0184] The term "single nucleotide polymorphism (SNP)" refers to a variation in a single nucleotide occurring at a specific location in the genome, where each variation is present in a population to a considerable degree (e.g., >1%). For example, at a specific base position in the human genome, the C nucleotide may be present in most individuals, but in a minority of individuals, that position is occupied by A. This means that an SNP is present at that specific location, and the two possible nucleotide variations (C or A) are paired genes at that location. SNPs are the basis for differences in disease susceptibility. Disease severity and how our bodies respond to treatment are also manifestations of gene variation. SNPs can be found in the coding region of a gene, in the non-coding region of a gene, or in intergenetic regions (regions between genes). In some implementations, SNPs within the coding sequence do not necessarily alter the amino acid sequence of the resulting protein due to the degeneracy of the genetic code. SNPs in coding regions have two types: synonymous and non-synonymous SNPs. Synonymous SNPs do not affect the protein sequence, while non-synonymous SNPs alter the amino acid sequence of the protein. Non-synonymous SNPs have two types: missense and nonsense. SNPs not located in protein-coding regions can still affect gene splicing, transcription factor binding, messenger RNA degradation, or non-coding RNA sequences. Gene expression affected by this type of SNP is called an eSNP (expressed SNP) and can occur upstream or downstream of the gene. Single nucleotide variants (SNVs) are variations in a single nucleotide without any frequency restrictions and can occur in somatic cells. Somatic single nucleotide variants (e.g., caused by cancer) are also called single-nucleotide alterations.
[0185] "Specifically binding" means that a nucleic acid molecule, polypeptide, or complex thereof (e.g., a nucleic acid programmable DNA binding domain and a guiding nucleic acid), compound, or molecule recognizes and binds to the polypeptide and / or nucleic acid molecules of the present invention, but does not substantially recognize and bind to other molecules in a sample (e.g., a biological sample).
[0186] The nucleic acid molecules used in the methods of this invention include any nucleic acid molecule encoding the polypeptide or a fragment thereof of the invention. Such nucleic acid molecules do not need to be 100% identical to the endogenous nucleic acid sequence, but will generally exhibit substantial similarity. Polynucleotides having "substantial similarity" to the endogenous sequence can generally hybridize with at least one strand of a double-stranded nucleic acid molecule. "Hybridization" means pairing under various stringent conditions to form a double-stranded molecule between complementary polynucleotide sequences (e.g., the gene described herein) or portions thereof. (See, for example, Wahl, GM and SLBerger (1987) Methods Enzymol. 152:399; Kimmel, AR (1987) Methods Enzymol. 152:507).
[0187] For example, stringent salt concentrations will typically be less than about 750 mM NaCl and 75 mM trisodium citrate, preferably less than about 500 mM NaCl and 50 mM trisodium citrate, and more preferably less than about 250 mM NaCl and 25 mM trisodium citrate. Low-stringency hybridization can be achieved in the absence of organic solvents (e.g., formamide), while high-stringency hybridization can be achieved in the presence of at least about 35% formamide, and more preferably at least about 50% formamide. Stringent temperature conditions will typically include temperatures of at least about 30°C, more preferably at least about 37°C, and most preferably at least about 42°C. Variations in other parameters, such as hybridization time, detergent (e.g., sodium dodecyl sulfate (SDS)) concentration, and the inclusion or exclusion of carrier DNA, are well known to those skilled in the art. Various degrees of stringency are achieved by combining these various desired conditions. In one embodiment, hybridization occurs at 30°C in 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS. In another embodiment, hybridization occurs at 37°C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA). In another embodiment, hybridization occurs at 42°C in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 μg / ml ssDNA. Variations in these conditions will be apparent to those skilled in the art.
[0188] For most applications, the severity of the post-hybridization cleaning step will vary. Cleaning severity conditions can be defined by salt concentration and temperature. As mentioned above, cleaning severity can be increased by decreasing the salt concentration or by increasing the temperature. For example, the stringent salt concentrations used for the cleaning step will be less than about 30 mM NaCl and 3 mM trisodium citrate, and may be less than about 15 mM NaCl and 1.5 mM trisodium citrate. The stringent temperature conditions used for the cleaning step will typically include temperatures of at least about 25°C, more preferably at least about 42°C, and even more preferably at least about 68°C. In a preferred embodiment, the cleaning step is carried out at 25°C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the cleaning step is carried out at 42°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In a preferred embodiment, the washing step is performed at 68°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Other variations under these conditions will be apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in the following: Benton and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.
[0189] "Substantially identical" means that the polypeptide or nucleic acid molecule exhibits at least 50% similarity to a reference amino acid sequence (e.g., any of the amino acid sequences described herein) or nucleic acid sequence (e.g., any of the nucleic acid sequences described herein). Preferably, such a sequence is at least 60%, more preferably 80% or 85%, and even more preferably 90%, 95%, or even 99% identical to the sequence used for comparison at the amino acid level or nucleic acid level.
[0190] Sequence identity is typically measured using sequence analysis software (e.g., the sequence analysis suite software from Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, COBALT, EMBOSS Needle, GAP, or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning homology to various substitutions, deletions, and / or other modifications. Conserved substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamyl; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. In one exemplary method for determining identity, the BLAST program can be used, wherein in e -3 and e -100 The probability scores between sequences indicate closely related sequences. For example, COBALT can be used with the following parameters:
[0191] a) Comparison parameters: gap penalty value -11, -1 and end-gap penalty value -5, -1.
[0192] b) CDD parameters: Enable using RPS BLAST; BLAST E-value 0.003; identify conservative rows and recalculate to enable, and
[0193] c) Query clustering parameters: Enable query clusters; font size 4; maximum cluster distance 0.8; standard letters.
[0194] Use the EMBOSS Needle, for example, with the following parameters:
[0195] a) Matrix: BLOSUM62;
[0196] b) Gap opening: 10;
[0197] c) Gap extension: 0.5;
[0198] d) Output format: in pairs;
[0199] e) End gap penalty value: pseudo-value;
[0200] f) End gap opening: 10; and
[0201] g) End gap extension: 0.5.
[0202] "Subject" means mammal, including (but not limited to): human or non-human mammals such as cows, horses, dogs, sheep or cats.
[0203] The term "target site" refers to a sequence within a nucleic acid molecule modified by a nucleobase editor. In one embodiment, the target site is deaminated by a deaminase or a fusion protein containing a deaminase (e.g., cytidine or adenine deaminase).
[0204] Since RNA-programmable nucleases (e.g., Cas9) use RNA:DNA hybridization to target DNA cleavage sites, the protein can, in principle, be targeted to any sequence specified by the guide RNA. Methods using RNA-programmable nucleases (such as Cas9) for site-specific cleavage (e.g., modification of the genome) are known in the art (see, for example, Cong, L. et al., Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, WY et al., Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic acids). Research (2013); Jiang, W. et al., RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature biotechnology 31, 233-239 (2013), the full text of which is incorporated herein by reference.
[0205] As used herein, the terms "treat / treating / treatment" and similar terms refer to reducing or improving a disease or symptom associated with it and / or its symptoms, or achieving a desired pharmacological and / or physiological effect. It should be understood that, not excluded, treating a disease or symptom does not require complete relief of the associated disease, condition, or symptoms. In some embodiments, the effect is therapeutic, i.e., without limitation, the effect partially or completely reduces, lowers, eliminates, alleviates, slows, or diminishes the intensity of the disease or symptom and / or the adverse symptoms attributable to it, or cures the disease or symptom and / or the adverse symptoms attributable to it. In some embodiments, the effect is preventative, i.e., the effect protects against or prevents the occurrence or recurrence of the disease, symptom, or symptoms. For this purpose, the method of the present invention comprises administering a therapeutically effective amount of the composition, as described herein.
[0206] "Uracil glycosidase inhibitor" refers to an agent that inhibits the uracil-excision repair system. In one embodiment, the agent is a protein or fragment thereof that binds to host uracil-DNA glycosidase and prevents the removal of uracil residues from DNA.
[0207] It should be understood that the ranges provided in this document are abbreviations for all values within the range, including the first and last values, as well as values in between. For example, it should be understood that the range 1 to 50 includes any number, combination of numbers, or subrange of numbers that are groups of the following: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.
[0208] The enumeration of chemical groups in any definition of the variables herein includes the definition of variables as any single group or a combination of the listed groups. The enumeration of embodiments of the variables or examples herein includes embodiments as any single embodiment or in combination with any other embodiment or part thereof.
[0209] Any composition or method provided herein may be combined with one or more other compositions and methods provided herein.
[0210] DNA editing has emerged as a viable means of modifying disease states by correcting pathogenic mutations at the gene level. Until now, all DNA editing platforms have relied on inducing DNA double-strand breaks (DSBs) at specific genomic loci and have depended on endogenous DNA repair pathways to determine product outcomes in a semi-random manner, resulting in complex gene product populations. While precise, user-defined repair outcomes can be achieved via the homologous recombination repair (HDR) pathway, several challenges have hindered the efficient use of HDR in therapeutically relevant cell types. In fact, this pathway is less efficient than the competitive, error-prone non-homologous end-joining pathway. Furthermore, HDR is strictly limited to the G1 and S phases of the cell cycle, preventing precise repair of DSBs in cells undergoing late mitosis. Consequently, it has proven difficult or impossible to efficiently alter genomic sequences in these populations in a user-defined, programmable manner.
[0211] Nucleobase editors
[0212] This document discloses a base editor or nucleobase editor for editing, modifying, or altering a target nucleotide sequence of a polynucleotide. The document describes a nucleobase editor or base editor comprising a polynucleotide programmable nucleotide binding domain and a nucleobase editing domain. The polynucleotide programmable nucleotide binding domain can specifically bind to the target polynucleotide sequence upon binding to a binding-guided polynucleotide (e.g., gRNA) (i.e., via complementary base pairing between bases of the binding-guided nucleic acid and the target polynucleotide sequence), thereby positioning the base editor to the desired editing of the target nucleic acid sequence. In some embodiments, the target polynucleotide sequence comprises single-stranded or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises DNA-RNA hybridization.
[0213] Multinucleotide programmable nucleotide binding domain
[0214] The term "programmable nucleotide binding domain" or "nucleic acid programmable DNA binding protein (napDNAbp)" refers to a protein associated with a nucleic acid (e.g., DNA or RNA), such as a guiding polynucleotide (e.g., guiding RNA), which directs the programmable nucleotide binding domain to a specific nucleic acid sequence. In some embodiments, the programmable nucleotide binding domain is a programmable DNA binding domain. In some embodiments, the programmable nucleotide binding domain is a programmable RNA binding domain. In some embodiments, the programmable nucleotide binding domain is a Cas9 protein. In some embodiments, the programmable nucleotide binding domain is a Cpf1 protein.
[0215] CRISPR is an acquired immune system that provides protection against mobile genetic elements (viruses, translocate elements, and conjugating plasmids). A CRISPR cluster contains a spacer, a sequence complementary to the previously mobile element, and a target invading nucleic acid. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In the type II CRISPR system, the corrective processing of the pre-crRNA requires a trans-coding small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA acts as a guide for ribonuclease 3-assisted processing of the pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endo-cleaves the linear or circular dsDNA target complementary to the spacer. Target strands not complementary to the crRNA are first endo-cleaved, then ex-cleaved at 3'-5'. In nature, DNA binding and cleavage typically require both protein and RNA. However, a single guide RNA ("sgRNA," or simply "gRNA") can be engineered to combine both crRNA and tracrRNA embodiments into a single RNA species. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the full text of which is incorporated herein by reference. Cas9 identifies short motifs (PAM or adjacent motifs of the original spacer sequence) in CRISPR repeat sequences to help distinguish between "self" and "non-self".
[0216] Cas9 domains of nucleobase editors
[0217] The Cas9 nuclease sequence and structure are well known to those skilled in the art (see, for example, “Complete genome sequence of an Ml strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan W.M., Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov A.N., Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma CM., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012), the full text of which is incorporated herein by reference. Cas9 is a homolog that has been described in various species, including (but not limited to): Streptococcus pyogenes and Streptococcus thermophilus. Additional suitable Cas9 nucleases and sequences are readily apparent to those skilled in the art based on this invention, and such Cas9 nucleases and sequences include Cas9 sequences and loci derived from organisms, as disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type IICRISPR-Cas immunitysystems” (2013) RNA Biology 10:5, 726-737; the full text of which is incorporated herein by reference.
[0218] In some implementations, the Cas9 nuclease has an inactive (e.g., deactivated) DNA cleavage domain; that is, the Cas9 is a cleavage enzyme, referred to as an "nCas9" protein (for "cleavage enzyme" Cas9). The nuclease-deactivated Cas9 protein is interchangeably referred to as a "dCas9" protein (for nuclease-inactivated Cas9). Methods for generating Cas9 proteins (or fragments thereof) with an inactive DNA cleavage domain are known (see, for example, Jinek et al., Science. 337:816-821 (2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene"). Expression (2013) Cell. 28; 152(5):1173-83, the full text of which is incorporated herein by reference. For example, it is known that the DNA cleavage domain of Cas9 comprises two subdomains: the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within the subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely deactivate the nuclease activity of Streptococcus pyogenes Cas9 (Jinek et al., Science. 337:816-821 (2012); Qi et al., Cell. 28; 152(5):1173-83 (2013)). In some embodiments, a protein comprising a fragment of Cas9 is provided. For example, in some embodiments, the protein comprises one of two Cas9 domains: (1) a gRNA-binding domain of Cas9; or (2) a DNA cleavage domain of Cas9. In some implementations, proteins containing Cas9 or fragments thereof are referred to as "Cas9 variants." Cas9 variants share homology with Cas9 or fragments thereof. For example, Cas9 variants are at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to wild-type Cas9. In some implementations... In this case, compared to wild-type Cas9, Cas9 variants may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid alterations.In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., a gRNA-binding domain or a DNA-cutting domain) such that, for a corresponding fragment of wild-type Cas9, the fragment is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of the corresponding wild-type Cas9.
[0219] In some embodiments, the fragment is at least 100 amino acids long. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids long.
[0220] In some implementations, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_017053.1, nucleotide and amino acid sequences are as follows):
[0221]
[0222] (Single bottom line: HNH domain; Double bottom line: RuvC domain).
[0223] In some implementations, wild-type Cas9 corresponds to, or contains, the following nucleotide and / or amino acid sequence:
[0224]
[0225] (Single bottom line: HNH domain; Double bottom line: RuvC domain).
[0226] In some implementations, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_002737.2 (nucleotide sequence below); and Uniprot reference sequence: Q99ZW2 (amino acid sequence below).
[0227]
[0228] (Single bottom line: HNH domain; Double bottom line: RuvC domain).
[0229] In some implementations, Cas9 refers to Cas9 derived from the following: *Corynebacterium ulcerans* (NCBI Refs: NC_015683.1, NC_017317.1); *Corynebacterium diphtheria* (NCBI Refs: NC_016782.1, NC_016786.1); *Spiroplasma syrphidicola* (NCBI Ref: NC_021284.1); *Prevotella intermedia* (NCBI Ref: NC_017861.1); *Spiroplasma taiwanense* (NCBI Ref: NC_021846.1); *Streptococcus iniae* (NCBI Ref: NC_021314.1); and *Belliella balolaensis*. baltica)(NCBI Ref: NC_018010.1); Psychroflexustorquis I)(NCBI Ref: NC_018721.1; Streptococcus thermophilus)(NCBI Ref: YP_820832.1), Listeria innocua)(NCBI Ref: NP_472073.1, Campylobacter jejuni)(NCBI Ref: YP_002344900.1 or Neisseria meningitidis)(NCBI Ref: YP_002342100.1, or Cas9 from any other organism.
[0230] In some implementations, the Cas9 domain contains a D10 mutation, while the residue at position 840 retains the histidine in the amino acid sequence provided above, or at the corresponding position in any of the amino acid sequences provided above.
[0231] In some embodiments, dCas9 partially or entirely corresponds to, or contains, a Cas9 amino acid sequence having one or more mutations that deactivate the Cas9 nuclease activity. For example, in some embodiments, the dCas9 domain contains the D10A and H840 mutations or corresponding mutations in another Cas9. In some embodiments, the dCas9 contains the amino acid sequence of dCas9 (D10A and H840A):
[0232] (Single bottom line: HNH domain; Double bottom line: RuvC domain).
[0233] In some implementations, the Cas9 domain contains a D10 mutation, while the residue at position 840 preserves the histidine in the amino acid sequence provided above, or at the corresponding position in any amino acid sequence provided herein.
[0234] In other embodiments, a dCas9 variant is provided having mutations other than D10A and H840A, which, for example, cause nuclease deactivation of Cas9 (dCas9). By way of example, such mutations include other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). In some embodiments, a variant or homolog of dCas9 is provided that is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical. In some embodiments, variants of dCas9 are provided having amino acid sequences of about 5, 10, 15, 20, 25, 30, 40, 50, 75, 100 or more amino acids, whether short or long.
[0235] In some embodiments, such as the Cas9 fusion proteins provided herein, the Cas9 fusion protein comprises the full-length amino acid sequence of a Cas9 protein, for example, one of the Cas9 sequences provided herein. However, in other embodiments, such as the fusion proteins provided herein, the fusion protein does not comprise the full-length Cas9 sequence, but only one or more fragments thereof. Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and additional suitable sequences and fragments of Cas9 domains will be apparent to those skilled in the art.
[0236] The Cas9 protein can be associated with a guide RNA that directs the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the polynucleotide programmable nucleotide binding domain is a Cas9 domain, such as nuclease-activated Cas9, Cas9 nickase (nCas9), or nuclease-inactivated Cas9 (dCas9). Examples of nucleic acid programmable DNA-binding proteins include (without limitation): Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i.
[0237] Nuclease-deactivated Cas9 proteins are interchangeably referred to as "dCas9" proteins (for nuclease-"inactivated" Cas9) or catalytically inactivated Cas9. Methods for generating Cas9 proteins (or fragments thereof) with inactivated DNA cleavage domains are known (see, e.g., Jinek et al., Science. 337:816-821 (2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of..."). GeneExpression (2013) Cell.28; 152(5):1173-83, the full text of which is incorporated herein by reference. For example, it is known that the DNA cleavage domain of Cas9 comprises two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the complementary strand to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely deactivate the nuclease activity of Streptococcus pyogenes Cas9 (Jinek et al., Science.337:816-821 (2012); Qi et al., Cell.28; 152(5):1173-83 (2013)).
[0238] In some embodiments, the Cas9 domain is a Cas9 cleavage enzyme. This Cas9 cleavage enzyme may be a Cas9 protein capable of cleaving only one strand of a duplex nucleic acid molecule (e.g., a duplex DNA molecule). In some embodiments, the Cas9 cleavage enzyme cleaves the target strand of the duplex nucleic acid molecule, meaning that the Cas9 cleavage enzyme cleaves the strand that is base-paired (complementary) to the gRNA (e.g., sgRNA) bound to the Cas9. In some embodiments, the Cas9 cleavage enzyme contains a D10 mutation and has a histidine residue at position 840. In some embodiments, the Cas9 cleavage enzyme cleaves the non-target, non-base-editing strand of the duplex nucleic acid molecule, meaning that the Cas9 cleavage enzyme cleaves the strand that is not base-paired to the gRNA (e.g., sgRNA) bound to the Cas9. In some embodiments, the Cas9 cleavage enzyme contains an H840A mutation and has an aspartic acid residue at position 10, or a corresponding mutation. In some embodiments, the Cas9 nickase comprises at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the same amino acid sequence as any of the Cas9 nickases provided herein. Other suitable Cas9 nickases will be apparent to those skilled in the art based on the present invention and the knowledge of the art, and are within the scope of the present invention.
[0239] In some embodiments, the Cas9 domain is a nuclease-inactive Cas9 domain (dCas9). For example, the dCas9 domain can bind to a dual nucleic acid molecule (e.g., via a gRNA molecule) without cleaving either strand of the dual nucleic acid molecule. In some embodiments, the nuclease-inactive dCas9 domain contains the D10X and H840X mutations of the amino acid sequence proposed herein, or the corresponding mutations in any amino acid sequence provided herein, where X is any amino acid change. In some embodiments, the nuclease-inactive dCas9 domain contains the D10A and H840A mutations of the amino acid sequence proposed herein, or the corresponding mutations in any amino acid sequence provided herein. As an example, the nuclease-inactive Cas9 domain contains the amino acid sequence proposed in the cloning vector pPlatTET-gRNA2 (accession number BAV54124):
[0240]
[0241] It should be understood that additional Cas9 proteins (e.g., nuclease-inactivated Cas9 (dCas9), Cas9 nicking enzyme (nCas9), or nuclease-activated Cas9), including their variants and homologs, are within the scope of this invention. Exemplary Cas9 proteins include (without limitation) those provided below. In some embodiments, the Cas9 protein is nuclease-inactivated Cas9 (dCas9). In some embodiments, the Cas9 protein is Cas9 nicking enzyme (nCas9). In some embodiments, the Cas9 protein is nuclease-activated Cas9.
[0242] Exemplary catalytic non-activated Cas9 (dCas9):
[0243]
[0244] Examples of Cas9 nickases (nCas9) are presented below:
[0245]
[0246] Examples of catalytic activation of Cas9 are presented below:
[0247]
[0248] In some embodiments, Cas9 refers to Cas9 derived from archaea (e.g., hyperthermophilic archaea), which constitute the domains and kingdoms of single-celled prokaryotic microorganisms. In some embodiments, a nucleic acid-programmable DNA-binding protein refers to CasX or CasY, which has been described, for example, in Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res. 2017 Feb 21. doi:10.1038 / cr.2017.21, the full text of which is incorporated herein by reference. Using genomic analysis of the overall genomics, some CRISPR-Cas systems can be identified, including Cas9, which was the first reported in the archaeal domain of life. This divergent Cas9 protein was found in archaea, which are rarely studied, and is part of the activating CRISPR-Cas system. In bacteria, two previously unknown systems, CRISPR-CasX and CRISPR-CasY, have been found, which are among the smallest systems discovered to date. In some embodiments, in the base editor systems described herein, Cas9 is replaced by CasX or CasX variants. In some implementations, Cas9 is replaced by CasY or a CasY variant in the base editor system described herein. It should be understood that other RNA-guided DNA-binding proteins can be used as nucleic acid programmable DNA-binding proteins (napDNAbp) and are within the scope of this invention.
[0249] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) of any fusion protein provided herein may be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the napDNAbp is a CasY protein. In some embodiments, the napDNAbp contains at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the same amino acid sequence as naturally occurring CasX or CasY proteins. In some embodiments, the napDNAbp is a naturally occurring CasX or CasY protein. In some embodiments, the napDNAbp contains an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any CasX or CasY protein described herein. It should be understood that CasX and CasY from other bacterial species may also be used according to the invention.
[0250] The following Cas sequences are provided through examples:
[0251] CasX(uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53)tr|F0NN87|F0NN87_SULIH CRISPR-related Casx protein OS = Sulfolobus islandicus (HVE10 / 4 typing) GN = SiH_0402PE = 4SV = 1:
[0252] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLE VEPHYLIIAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAVGQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG.
[0253] >tr|F0NH53|F0NH53_SULIR CRISPR-related protein, Casx OS = Icelandic Sulfophyllum (REY15A typing) GN = SiRe_0771PE = 4SV = 1:
[0254] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTSNIILPLSGNDKNPWTETLKCYNFPTTTVALSEVFKNFSQVKECEEVSAPSFVKPFEYKFGRSPGMVERTRRVKLEVEPHYLIMAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAVGQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG.
[0255] δ spp.CasX
[0256] 。
[0257] CasY (ncbi.nlm.nih.gov / protein / APG80656.1) > APG80656.1 CRISPR-related protein CasY [uncultured bacteria]:
[0258]
[0259] Cas12b / C2c1(uniprot.org / uniprot / T0D7A2#2)sp|T0D7A2|C2C1_ALIAG CRISPR-related intranuclease C2c1 OS=Bacillus acidus (ATCC 49025 / DSM 3922 / CIP 106132 / NCIMB 13137 / GD3B)GN=c2c1 PE=1SV=1:
[0260]
[0261] In some implementations, one of the Cas9 domains present in the fusion protein can be directed to replace the nucleotide sequence-programmable DNA-binding protein domain, which does not require a PAM sequence.
[0262] In some implementations, the nucleic acid-programmable DNA-binding protein (napDNAbp) is a single-actor in the microbial CRISPR-Cas system. Single-actor microbial CRISPR-Cas systems include (without limitation): Cas9, Cpf1, Cas12b / C2c1, and Cas12c / C2c3. Typically, microbial CRISPR-Cas systems are classified into Class 1 and Class 2 systems. Class 1 systems have multiple single-actor complexes, while Class 2 systems have a single protein actinor. For example, Cas9 and Cpf1 are Class 2 actinors. Besides Cas9 and Cpf1, three different Class 2 CRISPR-Cas systems (Cas12b / C2c1 and Cas12c / C2c3) have been described below: Shmakov et al., “Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems”, Mol. Cell, November 5, 2015; 60(3):385-397, the full text of which is incorporated herein by reference. The controllers of both systems (Cas12b / C2c1 and Cas12c / C2c3) contain a RuvC-like intranuclease domain associated with Cpf1. The third system contains an controller with two predicted HEPN RNase domains. Unlike the production of CRISPR RNA via Cas12b / C2c1, the production of mature CRISPR RNA is tracrRNA-independent. For DNA cleavage, Cas12b / C2c1 depends on both CRISPR RNA and tracrRNA.
[0263] The crystal structure of *Alicyclobaccillus acidoterrestris* Cas12b / C2c1 (AacC2c1) has been reported in a complex containing a chimeric single-molecule guide RNA (sgRNA). See, for example, Liu et al., “C2c1-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism,” *Mol. Cell*, January 19, 2017; 65(2):310-322, the full text of which is incorporated herein by reference. This crystal structure has also been reported in *Alicyclobaccillus acidoterrestris* C2c1, which binds to target DNA as a ternary complex. See, for example, Yang et al., “PAM-dependent Target DNA Recognition and Cleavage by C2C1 CRISPR-Cas endonuclease,” Cell, December 15, 2016; 167(7):1814-1828, the full text of which is incorporated herein by reference. The catalytically competent configuration of AacC2c1 (containing both target and non-target DNA strands) has been independently captured and localized in a single RuvC catalytic bag, where Cas12b / C2c1-mediated cleavage results in staggered seven-nucleotide breaks of the target DNA. Structural comparisons between the Cas12b / C2c1 ternary complex and previously identified Cas9 and Cpf1 control moieties illustrate the diversity of mechanisms employed by the CRISPR-Cas9 system.
[0264] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) of any fusion protein provided herein may be a Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a Cas12b / C2c1 protein. In some embodiments, the napDNAbp is a Cas12c / C2c3 protein. In some embodiments, the napDNAbp contains at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the same amino acid sequence as naturally occurring Cas12b / C2c1 or Cas12c / C2c3 proteins. In some embodiments, the napDNAbp is a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp contains at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the same amino acid sequence as any of the napDNAbp sequences provided herein. It should be understood that Cas12b / C2c1 or Cas12c / C2c3 from other bacterial species may also be used according to the invention.
[0265] The amino acid sequence of Cas12b / C2c1(uniprot.org / uniprot / T0D7A2#2)sp|T0D7A2| / C2C1_ALIAGCRISPR-related intranuclease C2c1 OS=Bacillus acidus (ATCC 49025 / DSM 3922 / CIP106132 / NCIMB 13137 / GD3B)GN=c2c1 PE=1SV=1 is provided below:
[0266]
[0267] BhCas12b (Bacillus hisashii), NCBI reference sequence: WP_095142515, amino acid sequence provided below:
[0268]
[0269] In some implementations, the Cas12b is BvCas12B, which is a variant of BhCas12b and includes the following variations related to BhCas12B: S893R, K846R, and E837G.
[0270] BvCas12b (Bacillus sp.) V3-13, NCBI reference sequence: WP_101661451.1, amino acid sequence provided below:
[0271]
[0272] It should be understood that the polynucleotide programmable nucleotide binding domain may also include a nucleic acid programmable protein that binds to RNA. For example, the polynucleotide programmable nucleotide binding domain may be associated with a nucleic acid that directs the polynucleotide programmable nucleotide binding domain to RNA. Other nucleic acid programmable DNA-binding proteins are also within the scope of this invention, although not explicitly listed herein.
[0273] The Cas proteins used in this article include those of classes 1 and 2. Unrestricted examples of Cas proteins include: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, C sb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i, CARF, DinG, their homologs, or their modified forms. Unmodified CRISPR enzymes can have DNA cleavage activity, such as Cas9, which has two functional intranuclease domains: RuvC and HNH. CRISPR enzymes can guide one or two cuts at a target sequence, such as within the target sequence and / or within complement of the target sequence. For example, CRISPR enzymes can guide one or two cuts within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500 or more base pairs, said base pairs being the first or last nucleotide from the target sequence.
[0274] The vector encoding the corresponding wild-type enzyme mutation of the CRISPR enzyme results in the mutant CRISPR enzyme lacking the ability to cleave one or both strands of the target polynucleotide containing the usable target sequence. Cas9 can refer to a polypeptide having at least or at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology to a wild-type exemplary Cas9 polypeptide (e.g., Cas9 from Streptococcus pyogenes). Cas9 can refer to a polypeptide having at most or at most about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology to a wild-type exemplary Cas9 polypeptide (e.g., from Streptococcus pyogenes). Cas9 can refer to the wild-type or modified form of the Cas9 protein, which may include amino acid changes such as deletion, insertion, substitution, variant, mutation, fusion, chimera or any combination thereof.
[0275] In some embodiments, the methods described herein may use engineered Cas proteins. Guide RNA (gRNA) is a short synthetic RNA consisting of a backbone sequence required for Cas-binding and a user-defined ~20 nucleotide spacer defining the target genosome to be modified. Therefore, it should be understood that, compared to the rest of the genosome, altering the Cas protein-specific target portion is determined by the specificity of the gRNA target sequence for that genosome target.
[0276] The Cas9 nuclease has two functional intracellular nuclease domains: RuvC and HNH. Upon target binding, Cas9 undergoes a second structural change, which positions the nuclease domain to cleave the opposing strands of the target DNA. The final result of Cas9-mediated DNA cleavage is a double-strand break (DSB) within the target DNA (~3-4 nucleotides upstream of the PAM sequence). The resulting DSB is then repaired by one of two general repair pathways: (1) the efficient but error-prone non-homologous end joining (NHEJ) pathway; or (2) the inefficient but highly accurate homologous recombination repair (HDR) pathway.
[0277] The "efficiency" of non-homologous end joining (NHEJ) and / or homologous recombination repair (HDR) can be calculated using any known method. For example, in some cases, efficiency can be expressed as a percentage of successful HDR. For instance, a surveyor nuclease assay can be used to produce cleavage products, and the product-to-matrix ratio can be used to calculate the percentage. For instance, a surveyor nuclease can be used to directly cleave DNA containing newly integrated restriction sequences as a result of successful HDR. More cleavage of the matrix indicates a larger percentage of HDR (more efficient HDR). As an illustrative example, the fraction (percentage) of HDR can be calculated using the following equation: [(cleavage product) / (matrix + cleavage product)] (e.g., (b+c) / (a+b+c), where "a" is the band strength of the DNA matrix, and "b" and "c" are the cleavage products).
[0278] In some cases, efficiency can be expressed as a percentage of successful NHEJs. For example, the NHEJ percentage can be calculated using a T7 intracellular nuclease I assay to produce the cleavage product and the product-to-matrix ratio. T7 intracellular nuclease I cleaves mismatched heteroduplex DNA, which is caused by hybridization of wild-type and mutant DNA strands (NHEJs result in small random insertions or deletions at the original break site). More cleavages indicate a higher NHEJ percentage (greater NHEJ efficiency). As an illustrative example, the NHEJ percentage can be calculated using the following equation: (1 - (1 - (b + c) / (a + b + c))) 1 / 2 )×100, where “a” is the band strength of the DNA matrix, and “b” and “c” are the cleavage products (Ran et al., Cell. 2013 Sep 12; 154(6):1380-9; and Ran et al., Nat Protoc. 2013 Nov; 8(11):2281–2308).
[0279] The NHEJ repair pathway is the most active repair mechanism, and it often induces small nucleotide insertions or deletions (indels) at DSB sites. The randomness of NHEJ-mediated DSB repair has important practical implications because Cas9-expressing cell populations and gRNAs or guide polynucleotides can cause a wide variety of mutations. In most cases, NHEJ induces small insertions or deletions in the target DNA, resulting in amino acid deletions, insertions, or frameshift mutations, leading to premature stop codons within the open reading frame (ORF) of the target gene. The ideal end result is a loss-of-function mutation within the target gene.
[0280] Although NHEJ-mediated DSB repair often disrupts the open reading frame of the gene, homologous recombination repair (HDR) can be used to produce specific nucleotide changes ranging from single nucleotide alterations to large insertions such as the addition of luciferases or tags.
[0281] To utilize HDR for gene editing, a DNA repair template containing the desired sequence can be delivered to a cell type of interest containing gRNA and Cas9 or Cas9 nickase. This repair template may contain the desired edit, as well as additional homologous sequences directly upstream and downstream of the target (called left and right homologous arms). The length of each homologous arm can depend on the size of the change to be introduced, with larger insertions requiring longer homologous arms. The repair template can be a single-stranded oligonucleotide, a double-stranded oligonucleotide, or a double-stranded DNA plasmid. HDR efficiency is typically low (<10% of the modified paired gene), even in cells expressing Cas9, gRNA, and the exogenous repair template. This HDR efficiency can be enhanced by synchronizing the cell, as HDR occurs during the S and G2 phases of the cell cycle. Chemically or genetically repressive genes involved in NHEJ can also increase HDR frequency.
[0282] In some implementations, Cas9 is a modified Cas9. A given RNA target sequence can have additional sites throughout the genome where partial homology exists. These sites are called off-target sites and need to be considered when designing gRNAs. In addition to optimizing gRNA design, CRISPR specificity can also be increased by modifying Cas9. Cas9 produces double-strand breaks (DSBs) through the combined activity of two nuclease domains (RuvC and HNH). The Cas9 nickase, the D10A mutation in SpCas9, retains one nuclease domain and produces a DNA gap instead of a DSB. This nickase system can also be combined with targeted gene editing and HDR-mediated gene editing.
[0283] In some cases, Cas9 is a variant Cas9 protein. A variant Cas9 polypeptide has an amino acid sequence that differs from the wild-type Cas9 protein by one amino acid (e.g., deletion, insertion, substitution, fusion). In some cases, the variant Cas9 polypeptide has amino acid changes (e.g., deletion, insertion, or substitution) that reduce the nuclease activity of the Cas9 polypeptide. For example, in some cases, the variant Cas9 polypeptide has less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nuclease activity of the corresponding wild-type Cas9 protein. In some cases, the variant Cas9 protein does not have significant nuclease activity. When the Cas9 protein is a variant Cas9 protein that does not have significant nuclease activity, it may be referred to as "dCas9".
[0284] In some cases, variant Cas9 proteins exhibit reduced nuclease activity. For example, variant Cas9 proteins exhibit less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% of the nuclease activity of wild-type Cas9 proteins (e.g., wild-type Cas9 proteins).
[0285] In some cases, variant Cas9 proteins can cleave the complementary strand of a double-stranded target sequence, but have reduced ability to cleave the non-complementary strand of the double-stranded target sequence. For example, the variant Cas9 protein may have a mutation (amino acid substitution) that reduces the function of the RuvC domain. As a non-limiting example, in some embodiments, the variant Cas9 protein has D10A (an aspartic acid substitution at amino acid position 10 for alanine) and thus can cleave the complementary strand of a double-stranded target sequence, but has reduced ability to cleave the non-complementary strand of the double-stranded target sequence (therefore, when the variant Cas9 protein cleaves the double-stranded target nucleic acid, it causes a single-stranded break (SSB) instead of a double-stranded break (DSB)) (see, for example, Jinek et al., Science. 2012 Aug 17; 337(6096):816-21).
[0286] In some cases, variant Cas9 proteins can cleave the non-complementary strand of a double-stranded guide target sequence, but have reduced ability to cleave the complementary strand. For example, the variant Cas9 protein may have a mutation (amino acid substitution) that reduces the function of the HNH domain (RuvC / HNH / RuvC domain motif). As a non-limiting example, in some embodiments, the variant Cas9 protein has the H840A mutation (histidine replaced by alanine at amino acid position 840) and thus can cleave the non-complementary strand of the guide target sequence, but has reduced ability to cleave the complementary strand (therefore, when the variant Cas9 protein cleaves the double-stranded guide target sequence, it results in an SSB instead of a DSB). Such Cas9 proteins have reduced ability to cleave guide target sequences (e.g., single guide target sequences), but retain the ability to bind to guide target sequences (e.g., single guide target sequences).
[0287] In some cases, the variant Cas9 protein exhibits reduced ability to cleave both the complementary and non-complementary strands of the double-stranded target DNA. As a non-limiting example, in some cases, the variant Cas9 protein carries mutations in both D10A and H840, resulting in reduced ability of the polypeptide to cleave both the complementary and non-complementary strands of the double-stranded target DNA. This reduces the ability of the Cas9 protein to cleave the target DNA (e.g., single-stranded target DNA) but retains its ability to bind to the target DNA (e.g., single-stranded target DNA).
[0288] As another non-limiting example, in some cases, this variant of the Cas9 protein carries mutations of W476A and W1126A, which reduces the peptide's ability to cleave target DNA. This reduces the ability of the Cas9 protein to cleave target DNA (e.g., single-stranded target DNA), but retains its ability to bind to target DNA (e.g., single-stranded target DNA).
[0289] As another non-limiting example, in some cases, this variant of the Cas9 protein carries mutations of P475A, W476A, N477A, D1125A, W1126A, and D1127A, which reduces the peptide's ability to cleave target DNA. This reduces the ability of the Cas9 protein to cleave target DNA (e.g., single-stranded target DNA), but retains its ability to bind to target DNA (e.g., single-stranded target DNA).
[0290] As another non-limiting example, in some cases, the variant Cas9 protein carries mutations of H840A, W476A, and W1126, resulting in reduced ability of the polypeptide to cleave target DNA. This reduced ability of the Cas9 protein to cleave target DNA (e.g., single-stranded target DNA) retains its ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some cases, the variant Cas9 protein carries mutations of H840A, D10A, W476A, and W1126, resulting in reduced ability of the polypeptide to cleave target DNA. This reduced ability of the Cas9 protein to cleave target DNA (e.g., single-stranded target DNA) retains its ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, this variant Cas9 has restored the catalytic His residue (A840H) at position 840 in the Cas9 HNH domain.
[0291] As another non-limiting example, in some cases, the variant Cas9 protein carries mutations of H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A, resulting in reduced ability of the polypeptide to cleave target DNA. This reduced ability of the Cas9 protein to cleave target DNA (e.g., single-stranded target DNA) retains its ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some cases, the variant Cas9 protein carries mutations of D10A, H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A, resulting in reduced ability of the polypeptide to cleave target DNA. This reduced ability of the Cas9 protein to cleave target DNA (e.g., single-stranded target DNA) retains its ability to bind to target DNA (e.g., single-stranded target DNA). In some cases, when the variant Cas9 protein carries mutations such as W476A and W1126A, or when it carries mutations such as P475A, W476A, N477A, D1125A, W1126A, and D1127A, the variant Cas9 protein does not bind effectively to the PAM sequence. Therefore, in some such cases, when this variant Cas9 protein is used in a binding method, the method does not require the PAM sequence. In other words, in some cases, when this variant Cas9 protein is used in a binding method, the method may include a guide RNA, but the method can be performed without the PAM sequence (and the binding specificity is therefore provided by the target fragment of the guide RNA). Other residues can be mutated to achieve the above effect (i.e., to partially deactivate one or more nucleases). As a non-restrictive example, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 may be altered (i.e., substituted). Furthermore, mutations other than alanine substitution are suitable.
[0292] In some implementations, variant Cas9 proteins with reduced catalytic activity (e.g., when the Cas9 protein has mutations such as D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987, e.g., D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A) can still bind to the target DNA at a specific site (because it is still guided to the target DNA sequence by the guide RNA), as long as it retains the ability to interact with the guide RNA.
[0293] Alternatives to Streptococcus pyogenes Cas9 may include RNA-guided intranucleases derived from the Cpf1 family, which exhibit cleavage activity in mammalian cells. CRISPR (CRISPR / Cpf1) from Prevotella and Francisella 1 is a DNA-editing technology similar to the CRISPR / Cas9 system. Cpf1 is a class II CRISPR / Cas system RNA-guided intranuclease. This acquired immune mechanism is observed in Prevotella and Francisella. Therefore, Cpf1 represents an example of a nucleic acid-programmable DNA-binding protein with PAM specificity different from Cas9. Like Cas9, Cpf1 is also a class II CRISPR agent. Cpf1 has been shown to mediate strong DNA interference with characteristics different from Cas9. Cpf1 is a single RNA-guided intranuclease lacking tracrRNA and utilizes T-enriched protospacer sequence adjacent motifs (TTN, TTTN, or YTN). Furthermore, Cpf1 cleaves DNA via interleaved double-strand breaks. Among the 16 Cpf1-family proteins, two enzymes from *Acidaminococcus* and *Lachnospiraceae* have shown effective genome-editing activity in human cells. The Cpf1 protein is described below, for example, in Yamano et al., “Crystal structure of Cpf1 in complex with guide RNA and target DNA.” Cell (165) 2016, pp. 949-962; the full text of which is incorporated herein by reference.
[0294] The Cpf1 gene is associated with a CRISPR locus and encodes an endonuclease that uses guide RNA to find and cleave viral DNA. Because Cpf1 is a smaller and simpler endonuclease than Cas9, it can overcome some CRISPR / Cas9 system limitations. Unlike Cas9 nucleases, Cpf1-mediated DNA cleavage results in double-stranded breaks with short 3′ overhangs. Cpf1's staggered cleavage pattern opens up the possibility of directional gene transfer, similar to traditional restriction enzyme cloning, which can increase gene editing efficiency. Like the aforementioned Cas9 variants and direct homologs, Cpf1 can also be expanded to target AT-rich regions or AT-rich gene bodies (lacking the SpCas9-preferred NGG PAM site) via CRISPR. The Cpf1 locus contains a mixed α / β domain followed by helical RuvC-I, RuvC-II, and zinc finger-like domains. The Cpf1 protein has a RuvC-like endonuclease domain, similar to the RuvC domain of Cas9. Furthermore, Cpf1 lacks the HNH nuclease domain, and its N-terminus lacks the α-helical recognition leaflet of Cas9. The Cpf1 CRISPR-Cas domain architecture reveals Cpf1 to be functionally unique, classifying it as a type 2 V CRISPR system. The Cpf1 locus encodes Cas1, Cas2, and Cas4 proteins, making it more similar to types I and III than type II systems. Functionally, Cpf1 does not require trans-activation of CRISPR RNA (tracrRNA); therefore, only CRISPR (crRNA) is needed. This is advantageous for genome editing because Cpf1 is not only smaller than Cas9 but also has a smaller sgRNA molecule (approximately half the number of nucleotides in Cas9). The Cpf1-crRN complex cleaves target DNA or RNA by recognizing the 5'-YTN-3' motif adjacent to the protospacer sequence that contrasts with the G-enriched PAM (targeted by Cas9). Upon recognition of PAM, Cpf1 introduces a sticky, end-like DNA double-strand break at a 4 or 5 nucleotide protrusion.
[0295] Nuclease-inactive Cpf1 (dCpf1) variants are also used in the compositions and methods of the present invention, serving as a guide nucleotide sequence-programmable DNA-binding protein domain. This Cpf1 protein possesses a RuvC-like intranuclease domain, similar to the RuvC domain of Cas9, but lacking an HNH intranuclease domain, and the N-terminus of Cpf1 does not possess the α-helical recognition leaf of Cas9. As shown in Zetsche et al., Cell, 163, 759-771, 2015 (incorporated herein by reference), the RuvC-like domain of Cpf1 is responsible for cleaving the DNA strands, and the deactivation of the RuvC-like domain deactivates Cpf1 nuclease activity. For example, mutations corresponding to D917A, E1006A, or D1255A in *Francisella novicida* Cpf1 deactivate Cpf1 nuclease activity. In some embodiments, the dCpf1 of the present invention includes mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A. It should be understood that any mutation (e.g., substitution mutations, deletions, or insertions that deactivate the RuvC domain of Cpf1) may be used according to the present invention.
[0296] In some embodiments, the nucleic acid-programmable DNA-binding protein (napDNAbp) of any fusion protein provided herein may be a Cpf1 protein. In some embodiments, the Cpf1 protein is a Cpf1 cleavage enzyme (nCpf1). In some embodiments, the Cpf1 protein is a nuclease-inactive Cpf1 (dCpf1). In some embodiments, the Cpf1, nCpf1, or dCpf1 comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the Cpf1 sequence disclosed herein. In some embodiments, the dCpf1 comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the Cpf1 sequence disclosed herein, and contains mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A.
[0297] It should be understood that Cpf1 from other bacterial species can also be used according to the present invention. Therefore, the following exemplary Cpf1 sequences from other bacterial species can also be used according to the present invention:
[0298] Wild-type Neofrancium Cpf1 (D917, E1006, and D1255 are in bold and underlined)
[0299]
[0300] Neofrancium Cpf1 D917A (A917, E1006 and D1255 are in bold and underlined)
[0301]
[0302]
[0303] Neofrancium Cpf1 E1006A (D917, A1006 and D1255 are in bold and underlined)
[0304]
[0305]
[0306] Neofrancium Cpf1 D1255A (D917, E1006 and A1255 are in bold and underlined)
[0307]
[0308] Neofrancium Cpf1 D917A / E1006A (A917, A1006, and D1255 are in bold and underlined)
[0309]
[0310]
[0311] Neofrancium Cpf1 D917A / D1255A (A917, E1006, and A1255 are in bold and underlined)
[0312]
[0313]
[0314] Neofrancium Cpf1 E1006A / D1255A (D917, A1006, and A1255 are in bold and underlined)
[0315]
[0316]
[0317] Neofrancium Cpf1 D917A / E1006A / D1255A (A917, A1006, and A1255 are in bold and underlined)
[0318]
[0319] The polynucleotide-programmable nucleotide binding domain of the base editor may itself contain one or more domains. For example, the polynucleotide-programmable nucleotide binding domain may contain one or more nuclease domains. In some embodiments, the nuclease domain of the polynucleotide-programmable nucleotide binding domain may contain an internal nuclease or an external nuclease. The term "external nuclease" refers to a protein or polypeptide capable of digesting nucleic acids (e.g., RNA or DNA) from their free ends, and the term "internal nuclease" refers to a protein or polypeptide capable of catalyzing (e.g., cleaving) an internal region of a nucleic acid (e.g., DNA or RNA). In some embodiments, the internal nuclease can cleave a single strand of a double-stranded nucleic acid. In some embodiments, the internal nuclease can cleave both strands of a double-stranded nucleic acid molecule. In some embodiments, the polynucleotide-programmable nucleotide binding domain may be a deoxyribonuclease. In some embodiments, the polynucleotide-programmable nucleotide binding domain may be a ribonuclease.
[0320] In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide binding domain is capable of cleaving zero, one, or both strands of the target polynucleotide. In some cases, the polynucleotide programmable nucleotide binding domain may include a nicking enzyme domain. The term "nicking enzyme" herein refers to a polynucleotide programmable nucleotide binding domain containing a nuclease domain capable of cleaving only one strand of a duplex nucleic acid molecule (e.g., DNA). In some embodiments, the nicking enzyme can be derived from a fully catalytically activated (e.g., native) form of the polynucleotide programmable nucleotide binding domain by introducing one or more mutations into the activated polynucleotide programmable nucleotide binding domain. For example, in the case where the polynucleotide programmable nucleotide binding domain includes a Cas9-derived nicking enzyme domain, the Cas9-derived nicking enzyme domain may include a D10A mutation and histidine (H) at position 840. In such cases, the residue H840 retains catalytic activity and can thereby cleave a single strand of the nucleic acid duplex. In another example, the Cas9-derived nicking enzyme domain may include an H840A mutation, while the amino acid residue at position 10 retains D. In some embodiments, the nicking enzyme can be derived from a fully catalytically activated (e.g., natural) form of a polynucleotide programmable nucleotide binding domain by removing all or part of a nuclease domain for which nicking enzyme activity is not required. For example, in cases where the polynucleotide programmable nucleotide binding domain includes a Cas9-derived nicking enzyme domain, the Cas9-derived nicking enzyme domain may contain all or part of a RuvC domain or an HNH domain.
[0321] A base editor containing a programmable nucleotide-binding domain of a polynucleotide containing a nicking enzyme domain can thus generate single-strand DNA breaks (gaps) at specific polynucleotide target sequences (e.g., determined by binding to the complementary sequence of a guide nucleic acid). In some embodiments, the strand of the target polynucleotide sequence of the nucleic acid duplex cleaved by a base editor containing a nicking enzyme domain (e.g., a Cas9-derived nicking enzyme domain) is a strand that has not been edited by that base editor (i.e., the strand cleaved by that base editor is opposite to the strand containing the base to be edited). In other embodiments, a base editor containing a nicking enzyme domain (e.g., a Cas9-derived nicking enzyme domain) can cleave strands of a DNA molecule targeted for editing. In this case, the untargeted strand is not cleaved.
[0322] This document also provides base editors comprising a polynucleotide programmable nucleotide binding domain that are catalytically inactivated (i.e., unable to cleave the target polynucleotide sequence). The terms "catalytic inactivation" and "nuclease inactivation" are used interchangeably to refer to a polynucleotide programmable nucleotide binding domain having one or more mutations and / or deletions that prevent it from cleaving strands of nucleic acids. In some embodiments, the catalytically inactivated polynucleotide programmable nucleotide binding domain base editor may lack nuclease activity due to specific point mutations in one or more nuclease domains. For example, in the case of a base editor comprising a Cas9 domain, the Cas9 may contain both a D10A mutation and an H840A mutation. Such mutations deactivate both nuclease domains, resulting in loss of nuclease activity. In other embodiments, the catalytically inactivated polynucleotide programmable nucleotide binding domain may comprise all or part of one or more deletions of the catalytic domain (e.g., the RuvC1 and / or HNH domains). In a further embodiment, the catalytically inactivating polynucleotide programmable nucleotide binding domain includes a point mutation (e.g., D10A or H840A) and a deletion of all or part of the nuclease domain.
[0323] This article also considers mutations that can produce catalytically inactivating polynucleotide programmable nucleotide-binding domains derived from previous functional versions of such domains. For example, in the case of catalytically inactivating Cas9 (“dCas9”), variants with mutations other than D10A and H840A are provided, resulting in nuclease deactivation of Cas9. By way of example, such mutations include other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain).
[0324] Based on the knowledge of this invention and the present technology, additional suitable nuclease-inactive dCas9 domains will be apparent to those skilled in the art and are within the scope of this invention. Such additional exemplary suitable nuclease-inactive Cas9 domains include (but are not limited to): D10A / H840A, D10A / D839A / H840A and D10A / D839A / H840A / N863A mutant domains (see, for example, Prashant et al., CAS9 transcriptional activators for targetspecificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013; 31(9):833-838, the entire text of which is incorporated herein by reference). In some embodiments, the dCas9 domain contains at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the same amino acid sequence as any of the dCas9 domains provided herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having mutations of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more, compared to any of the amino acid sequences presented herein. In some embodiments, the Cas9 domain comprises an amino acid sequence having at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical consecutive amino acid residues, compared to any of the amino acid sequences presented herein.
[0325] Non-restricted examples of polynucleotide-programmable nucleotide-binding domains that can be incorporated into base editors include: CRISPR protein-derived domains, restriction nucleases, giant nucleases, TAL nucleases (TALENs), and zinc finger nucleases (ZFNs). In some cases, base editors include polynucleotide-programmable nucleotide-binding domains containing natural or modified proteins or portions thereof, which can bind to nucleic acid sequences via binding guide nucleic acids during CRISPR (i.e., clustered regularly spaced short palindromic repeats)-mediated modification of nucleic acids. Such proteins are referred to herein as "CRISPR proteins." Therefore, this paper discloses base editors containing polynucleotide-programmable nucleotide-binding domains that contain all or part of a CRISPR protein (i.e., a base editor containing all or part of a CRISPR protein as a domain, also referred to as the "CRISPR protein-derived domain" of the base editor). The CRISPR protein-derived domain incorporated into the base editor can be modified relative to the wild-type or natural version of the CRISPR protein. For example, as described below, the CRISPR protein-derived domain may contain one or more mutations, insertions, deletions, rearrangements, and / or recombinations relative to the wild-type or natural version of the CRISPR protein.
[0326] In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is an endonuclease (e.g., deoxyribonuclease or ribonuclease) that binds to the target polynucleotide when conjugated with a binding guide nucleic acid. In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is a nicking enzyme that binds to the target polynucleotide when conjugated with a binding guide nucleic acid. In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is a catalytic inactivation domain that binds to the target polynucleotide when conjugated with a binding guide nucleic acid. In some embodiments, the target polynucleotide bound by the CRISPR protein-derived domain of the base editor is DNA. In some embodiments, the target polynucleotide bound by the CRISPR protein-derived domain of the base editor is RNA.
[0327] In some implementations, the CRISPR protein-derived domain of the base editor may include all or part of Cas9 from the following: Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheriae (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplastic strain of Hippophae rhamnoides (NCBI Ref: NC_021284.1); Prevotella intermedius (NCBI Ref: NC_017861.1); Spiroplastic strain of Taiwania (NCBI Ref: NC_021846.1); Streptococcus spp. (NCBI Ref: NC_021314.1); Belyella balolaensis (NCBI Ref: NC_018010.1); Curvularia cytogenes (NCBI Ref: NC_018721.1); Streptococcus thermophilus ... Ref: YP_820832.1); harmless Listeria (NCBIRef: NP_472073.1); Campylobacter jejuni (NCBI Ref: YP_002344900.1); Neisseria meningitidis (NCBI Ref: YP_002342100.1), Streptococcus pyogenes or Staphylococcus aureus.
[0328] In some embodiments, the Cas9-derived domain of the base editor is a Cas9 domain (SaCas9) derived from Staphylococcus aureus. In some embodiments, the SaCas9 domain is a nuclease-activated SaCas9, a nuclease-inactivated SaCas9 (SaCas9d), or a SaCas9 nickase (SaCas9n). In some embodiments, the SaCas9 domain contains the N579X mutation. In some embodiments, the SaCas9 domain contains the N579A mutation. In some embodiments, the SaCas9 domain, the SaCas9d domain, or the SaCas9n domain can bind to a nucleic acid sequence having an atypical PAM. In some embodiments, the SaCas9 domain, the SaCas9d domain, or the SaCas9n domain can bind to a nucleic acid sequence having the NNGRRT PAM sequence. In some embodiments, the SaCas9 domain contains one or more of the E781X, N967X, and R1014X mutations.
[0329] In some embodiments, the Cas9 domain is a Cas9 domain derived from Staphylococcus aureus (SaCas9). In some embodiments, the SaCas9 domain is a nuclease-activated SaCas9, a nuclease-inactivated SaCas9 (SaCas9d), or a SaCas9 nickase (SaCas9n). In some embodiments, the SaCas9 contains an N579A mutation, or a corresponding mutation in any of the amino acid sequences provided herein.
[0330] In some embodiments, the SaCas9 domain, the SaCas9d domain, or the SaCas9n domain may bind to a nucleic acid sequence having an atypical PAM. In some embodiments, the SaCas9 domain, the SaCas9d domain, or the SaCas9n domain may bind to a nucleic acid sequence having an NNGRRT or NNNRRT PAM sequence. In some embodiments, the SaCas9 domain contains one or more of the E781X, N967X, and R1014X mutations, or a corresponding mutation in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SaCas9 domain contains one or more of the E781K, N967K, and R1014H mutations, or one or more corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SaCas9 domain contains the E781K, N967K, or R1014H mutation, or a corresponding mutation in any of the amino acid sequences provided herein.
[0331] A base editor may include all or part of a domain derived from a high-precision Cas9. In some embodiments, the high-precision Cas9 domain of the base editor is an engineered Cas9 domain containing one or more mutations that reduce the electrostatic interactions between the Cas9 domain and the sugar-phosphate backbone of the DNA relative to the corresponding wild-type Cas9 domain. A high-precision Cas9 domain with reduced electrostatic interactions with the sugar-phosphate backbone of the DNA may have fewer off-target effects. In some embodiments, the Cas9 domain (e.g., a wild-type Cas9 domain) contains one or more mutations that reduce the association between the Cas9 domain and the sugar-phosphate backbone of the DNA. In some embodiments, the Cas9 domain contains one or more mutations that reduce the association between the Cas9 domain and the sugar-phosphate backbone of the DNA by at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, or more.
[0332] In some implementations, the variant Cas protein may be spCas9, spCas9-VRQR, spCas9-VRER, xCas9(sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK, or spCas9-LRVSQL. An exemplary saCas9 sequence is provided below:
[0333] In the above saCas9 sequence, residue N579 (which is underlined and in bold) can be mutated (e.g., mutated to A579) to produce the SaCas9 nickase.
[0334] An example SaCas9n sequence is provided as follows:
[0335] In the above SaCas9n sequence, residue A579 (which can be mutated from N579 to produce SaCas9 nickase) is underlined and in bold.
[0336] An example sequence of SaKKHCas9 is provided below:
[0337]
[0338] The residues A579 (which may be produced by mutation of N579 to generate SaCas9 nickase) are underlined and in bold. The residues K781, K967 and H1014 (which may be produced by mutation of E781, N967 and R1014 to generate SaKKHCas9) are underlined and in italics.
[0339] In some embodiments, the modified Cas9 is a high-precision Cas9 enzyme. In some embodiments, the high-precision Cas9 enzyme is SpCas9 (K855A), eSpCas9 (1.1), SpCas9-HF1, or a highly precise Cas9 variant (HypaCas9). The modified Cas9 eSpCas9 (1.1) contains an alanine substitution that weakens the interaction between the HNH / RuvC groove and the non-target DNA strand, preventing strand separation and cleavage at off-target sites. Similarly, SpCas9-HF1 reduces off-target editing via an alanine substitution, disrupting the interaction between Cas9 and the DNA phosphate backbone. HypaCas9 contains mutations (SpCas9N692A / M694A / Q695A / H698A) in the REC3 domain that increase Cas9 correction and target identification. All three high-precision enzymes produce less off-target editing than wild-type Cas9. An exemplary high-precision Cas9 is provided below. The high-precision Cas9 domain mutations compared to Cas9 are shown in bold and underlined.
[0340]
[0341] Guided polynucleotides
[0342] As used herein, the term "guide polynucleotide" refers to a polynucleotide that is specific to a target sequence and can form a complex with a polynucleotide programmable nucleotide-binding domain protein (e.g., Cas9 or Cpf1). In one embodiment, the guide polynucleotide is a guide RNA. As used herein, the term "guide RNA (gRNA)" and its grammatical equivalents can refer to RNA that is specific to target DNA and can form a complex with a Cas protein. The RNA / Cas complex assists in "guided" the Cas protein to the target DNA. The Cas9 / crRNA / tracrRNA endonucleotides the linear or circular dsDNA target complementary to the spacer. Target strands not complementary to the crRNA are first endonucleotidely cleaved and then exonucleotidely trimmed 3'-5'. In nature, DNA binding and cleavage typically require both a protein and two RNAs. However, a single guide RNA ("sgRNA" or simply "gRNA") can be engineered to incorporate examples of two crRNAs and tracrRNA into a single RNA species. See, for example, Jinek M. et al., Science 337:816-821 (2012), the full text of which is incorporated herein by reference. Cas9 identifies short motifs (PAM or adjacent motifs of the original spacer sequence) in the CRISPR repeat sequence to help distinguish between "self" and "non-self". The Cas9 nuclease sequence and structure are well known to those skilled in the art (see, for example, “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti, JJ et al., Natl. Acad. Sci. USA 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded smallRNA and host factor RNase III.” Deltchev E. et al., Nature 471:602-607 (2011); and “Programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M. et al., Science 337:816-821 (2012), all of which are incorporated herein by reference in their entirety). Cas9 is a homologous gene described in various species, including (but not limited to) Streptococcus pyogenes and Streptococcus thermophilus.Based on this invention, additional suitable Cas9 nucleases and sequences will be readily apparent to those skilled in the art, and such Cas9 nucleases and sequences include the Cas9 sequences disclosed below from organisms and loci: Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire text of which is incorporated herein by reference. In some embodiments, the Cas9 nuclease has an inactive (e.g., deactivated) DNA cleavage domain, i.e., the Cas9 is a nicking enzyme.
[0343] In some embodiments, the guiding polynucleotide is at least one single guiding RNA ("sgRNA" or "gRNA"). In some embodiments, the guiding polynucleotide is at least one tracrRNA. In some embodiments, the guiding polynucleotide does not require a PAM sequence to guide the polynucleotide-programmable DNA-binding domain (e.g., Cas9 or Cpf1) to the target nucleotide sequence.
[0344] The multinucleotide programmable nucleotide-binding domain (e.g., a CRISPR-derived domain) of the base editor disclosed herein can identify target polynucleotide sequences by associating with a guide polynucleotide. The guide polynucleotide (e.g., gRNA) is typically single-stranded and programmable to bind at a specific site (i.e., via complementary base pairing) to the target sequence of the polynucleotide, thereby guiding the base editor, which is bound to the guide nucleic acid, to the target sequence. The guide polynucleotide may be DNA. The guide polynucleotide may be RNA. In some cases, the guide polynucleotide contains a natural nucleotide (e.g., adenosine). In some cases, the guide polynucleotide contains a non-natural (or unnatural) nucleotide (e.g., peptide nucleic acid or nucleotide analogue). In some cases, the length of the target region of the guide nucleic acid sequence may be at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. The length of the target region of a guiding nucleic acid can be between 10-30 nucleotides, between 15-25 nucleotides, or between 15-20 nucleotides.
[0345] In some embodiments, the guide polynucleotide comprises two or more individual polynucleotides that can interact with each other via, for example, complementary base pairing (e.g., dual guide polynucleotide). For example, the guide polynucleotide may comprise CRISPR RNA (crRNA) and trans-activated CRISPR RNA (tracrRNA). For example, the guide polynucleotide may comprise one or more trans-activated CRISPR RNAs (tracrRNA).
[0346] In type II CRISPR systems, targeting nucleic acids by a CRISPR protein (e.g., Cas9) typically requires complementary base pairing between a first RNA molecule (crRNA) and a second RNA molecule (trRNA). The first RNA molecule (crRNA) contains a sequence that recognizes the target sequence, and the second RNA molecule (trRNA) contains repetitive sequences that form the backbone region that stabilizes the guide RNA-CRISPR protein complex. This dual guide RNA system can be used as a guide polynucleotide to direct the base editor disclosed herein to the target polynucleotide sequence.
[0347] In some embodiments, the base editor provided herein utilizes a single guiding polynucleotide (e.g., gRNA). In some embodiments, the base editor provided herein utilizes dual guiding polynucleotides (e.g., dual gRNA). In some embodiments, the base editor provided herein utilizes one or more guiding polynucleotides (e.g., multiple gRNA). In some embodiments, the single guiding polynucleotide is used for different base editors described herein. For example, a single guiding polynucleotide can be used for both cytidine and adenosine base editors.
[0348] In other embodiments, the guiding polynucleotide may be contained in both the polynucleotide targeting portion and the backbone portion of the nucleic acid in a single molecule (i.e., a single-molecule guiding nucleic acid). For example, a single-molecule guiding polynucleotide may be a single guiding RNA (sgRNA or gRNA). The term guiding polynucleotide sequence in this document refers to any single, double, or multiple-molecule nucleic acid that can interact with a base editor and guide the base editor to the target polynucleotide sequence.
[0349] Typically, a guide polynucleotide (e.g., a crRNA / trRNA complex or gRNA) comprises a "polynucleotide-targeting fragment," which includes a sequence that recognizes and binds to a target polynucleotide sequence, and a "protein-binding fragment," which is the guide polynucleotide within a polynucleotide programmable nucleotide-binding domain that stabilizes the base editor. In some embodiments, the polynucleotide-targeting fragment of the guide polynucleotide recognizes and binds to a DNA polynucleotide, thereby facilitating base editing in the DNA. In other cases, the polynucleotide-targeting fragment of the guide polynucleotide recognizes and binds to an RNA polynucleotide, thereby facilitating base editing in the RNA. Hereinafter, "fragment" refers to a segment or region of a molecule, such as a continuous extension of nucleotides in the guide polynucleotide. A fragment may also refer to a region / segment of a complex, such that a fragment may contain a region of more than one molecule. For example, in cases where the guide polynucleotide comprises multiple nucleic acid molecules, the protein-binding fragment may include all or part of multiple individual molecules, for example, hybridizing along complementary regions. In some embodiments, the protein-binding fragment of DNA-targeted RNA (which comprises two individual molecules) may comprise: (i) 40-75 base pairs of a first RNA molecule of 100 base pairs in length; and (ii) 10-25 base pairs of a second RNA molecule of 50 base pairs in length. The definition of "fragment" (unless otherwise specifically defined in a particular context) is not limited to a specific number of total base pairs, not limited to any specific number of base pairs from a given RNA molecule, not limited to a specific number of individual molecules within a complex, and may include RNA molecule regions having any total length, and may include regions complementary to other molecules.
[0350] Guide RNA or guide polynucleotides may contain two or more RNAs, such as CRISPR RNA (crRNA) and transactivated crRNA (tracrRNA). Guide RNA or guide polynucleotides may sometimes contain single-stranded RNA, or a single guide RNA (sgRNA) formed by the fusion of crRNA and portions of tracrRNA (e.g., the functional portion). Guide RNA or guide polynucleotides may also be dual RNAs containing both crRNA and tracrRNA. Furthermore, crRNA can hybridize with target DNA.
[0351] As discussed above, the guide RNA or guide polynucleotide can be an expression product. For example, the DNA encoding the guide RNA can be a vector containing the sequence encoded by that guide RNA. The guide RNA or guide polynucleotide can be transferred into cells by transfecting cells with isolated guide RNA or plasso DNA containing the sequence encoded by the guide RNA and a promoter. Other methods can also be used to transfer the guide RNA or guide polynucleotide into cells, such as using virus-mediated gene delivery.
[0352] Guide RNA or guide polynucleotides can be isolated. For example, guide RNA can be transfected into cells or organisms in the form of separable RNA. Guide RNA can be prepared by in vitro transcription using any in vitro transcription system known in the art. Guide RNA can be transferred into cells in the form of separable RNA, rather than in the form of plastids containing the coding sequence of the guide RNA.
[0353] Guide RNA or guide polynucleotide may contain three regions: a first region at the 5' end, which is complementary to the target site in the chromosomal sequence; a second inner region, which can form a stem-loop structure; and a third 3' region, which may be single-stranded. The first region of each guide RNA may also be different, so that each guide RNA guides the fusion protein to a specific target site. Furthermore, the second and third regions of each guide RNA may be identical in all guide RNAs.
[0354] The first region of the guide RNA or guide polynucleotide may be complementary to the sequence at the target site in the chromosomal sequence, such that the first region of the guide RNA can pair with the bases at the target site. In some cases, the first region of the guide RNA may contain from or from about 10 nucleotides to 25 nucleotides (i.e., 10 nucleotides to 12 nucleotides; or about 10 nucleotides to about 25 nucleotides; or 10 nucleotides to about 25 nucleotides; or about 10 nucleotides to 25 nucleotides) or more. For example, the base-pairing region between the first region of the guide RNA and the target site in the chromosomal sequence may or may be about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25 or more nucleotides in length. Sometimes, the first region of the guide RNA may or may be about 19, 20 or 21 nucleotides in length.
[0355] Guiding RNA or guiding polynucleotides may also contain a second region that forms a secondary structure. For example, a secondary structure formed by guiding RNA may include a stem (or hairpin) and a loop. The lengths of the loop and the stem can vary. For example, the loop may be 0.5 or in the range of about 3 to 10 nucleotides in length, and the stem may be 0.5 or in the range of about 6 to 20 base pairs in length. The stem may contain one or more protrusions having 1 to 10 or about 10 nucleotides. The total length of the second region may be 0.5 or in the range of about 16 to 60 nucleotides in length. For example, the loop may be 0.5 or about 4 nucleotides in length, and the stem may be 0.5 or about 12 base pairs in length.
[0356] The guide RNA or guide polynucleotide may also contain a third region at its 3' end, which is essentially single-stranded. For example, the third region is sometimes not complementary to any chromosomal sequence in the cell of interest, and sometimes not complementary to the rest of the guide RNA. Furthermore, the length of the third region can vary. The third region can be more than or more than about 4 nucleotides in length. For example, the length of the third region can range from about 5 to about 60 nucleotides.
[0357] Guide RNA or guide polynucleotides can target any exon or intron of a gene. In some cases, it guides exon 1 or 2 of the target gene; in others, it guides exon 3 or 4. The composition may contain multiple guide RNAs, all targeting the same exon, or in some cases, the composition may contain multiple guide RNAs, each targeting a different exon. The exons and introns of the target gene.
[0358] Guide RNA or guide polynucleotides can target nucleic acid sequences of 20 or about 20 nucleotides. The target nucleic acid may be fewer than or less than about 20 nucleotides. The target nucleic acid may be at least or at least about 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, or any nucleotide length between 1 and 100. The target nucleic acid may be at most or at most about 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, or any nucleotide length between 1 and 100. At the 5' of the first nucleotide immediately preceding the PAM, the target nucleic acid sequence may be or may be about 20 bases. Guide RNA can target nucleic acid sequences. The target nucleic acid may be at least or at least about 1-10, 1-20, 1-30, 1-40, 1-50, 1-60, 1-70, 1-80, 1-90 or 1-100 nucleotides.
[0359] A guide polynucleotide (e.g., guide RNA) can refer to a nucleic acid that can hybridize to another nucleic acid (e.g., a target nucleic acid or protospacer in the cellular genome). A guide polynucleotide can be RNA. A guide polynucleotide can be DNA. The guide polynucleotide can be programmed or designed to bind to a nucleic acid sequence at a specific site. A guide polynucleotide can contain a polynucleotide chain and can be called a single guide polynucleotide. A guide polynucleotide can contain two polynucleotide chains and can be called a double guide polynucleotide. Guide RNA can be introduced into cells or embryos as RNA molecules. For example, RNA molecules can be synthesized in vitro by transcription and / or chemical synthesis. RNA can be transcribed from synthetic DNA molecules, for example, Gene fragments. The guide RNA can then be introduced into cells or embryos as an RNA molecule. Alternatively, the guide RNA can be introduced into cells or embryos in the form of a non-RNA nucleic acid molecule (e.g., a DNA molecule). For example, the DNA encoding the guide RNA can be operatively linked to a promoter control sequence for expression of the guide RNA in cells or embryos of interest. The RNA coding sequence can be operatively linked to a promoter sequence recognized by RNA polymerase III (Pol III). Plasmonic vectors that can be used to express guide RNA include (but are not limited to): the px330 vector and the px333 vector. In some cases, the plasmonic vector (e.g., the px333 vector) may contain at least two guide RNA-coding DNA sequences.
[0360] This document describes methods for selecting, designing, and validating guide polynucleotides (e.g., guide RNAs and target sequences), which are known to those skilled in the art. For example, to minimize the potential matrix multifunctionality of deaminase domains (e.g., AID domains) in the nucleobase editor system, the number of residues that may be unintentionally targeted for deamination (e.g., off-target C residues that may potentially remain on ssDNA within the target nucleic acid locus) can be minimized. Furthermore, software tools can be used to optimize gRNAs corresponding to the target nucleic acid sequence, for example, to minimize overall off-target activity across the entire genome. For example, for the various possible target domain selections using Streptococcus pyogenes Cas9, all off-target sequences (previously selected PAMs, e.g., NAG or NGG) can recognize the entire genome containing up to a certain number (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) of mismatched base pairs. The first region of the gRNA complementary to the target site can be identified, and all first regions (e.g., crRNA) can be ranked according to their total estimated off-target score; the highest-ranked target region indicates the target region that is likely to have the greatest on-target and minimal off-target activity. Candidate target gRNAs can be evaluated functionally using methods known in the art and / or as presented herein.
[0361] As a non-restricted example, a DNA sequence search algorithm can be used to identify the target DNA hybridization sequence in the crRNA of the guide RNA used with Cas9. GRNA design can be performed using customized gRNA design software based on the publicly available tool cas-offinder, as described in: Bae S., Park J., & Kim J.-S. Cas-OFFinder: A fast and versatile algorithm that searches for potential off-target sites of Cas9RNA-guided endonucleases. Bioinformatics 30, 1473-1475 (2014). This software scores the guide RNA after calculating its whole genome off-target propensity. Matches are typically considered for guide RNAs ranging from 17 to 24 in length, from perfect matches to 7 mismatches. Once off-target sites are determined by calculation, cumulative scores can be calculated for each guide RNA and a summary can be output in a table using a web interface. In addition to identifying potential target sites adjacent to the PAM sequence, the software also identifies all PAM-adjacent sequences that differ from the selected target site by 1, 2, 3, or more than 3 nucleotides. The genomic DNA sequence of the target nucleic acid sequence (e.g., the target gene) can be obtained, and repetitive elements can be screened using publicly available tools, such as the RepeatMasker program. RepeatMasker searches the input DNA sequence for repetitive elements and regions with low complexity. The output is a detailed annotation of the repetitive sequences present in the given query sequence.
[0362] Following identification, the first region of the guide RNA (e.g., crRNA) can be sorted and graded based on its distance to the target site, its orthogonality, and the presence of a 5' nucleotide closely matching the associated PAM sequence (e.g., based on the recognition of a closely matching 5'G in human genomes containing the associated PAM (e.g., NGG PAM for Streptococcus pyogenes, NNGRRT or NNGRRV PAM for Staphylococcus aureus). As used herein, orthogonality refers to the number of sequences in human genomes containing the minimum number of mismatches to the target sequence. "High orthogonality" or "good orthogonality" may, for example, refer to a 20-mer target region that does not have the same sequence as the desired target in the human genome, nor any sequence containing one or two mismatches in the target sequence. Target regions with good orthogonality can be selected to minimize off-target DNA cleavage.
[0363] In some embodiments, the reporter system can be used to detect base-editing activity and test candidate guide polynucleotides. In some embodiments, the reporter system may include a reporter gene-based assay in which base-editing activity leads to the expression of the reporter gene. For example, the reporter system may include a reporter gene containing a deactivating start codon, for example, a mutation on a template strand from 3'-TAC-5' to 3'-CAC-5'. Upon successful deamination of the target C, the corresponding mRNA will be transcribed to 5'-AUG-3' instead of 5'-GUG-3', thereby enabling the translation of the reporter gene. Suitable reporter genes will be apparent to those skilled in the art. Non-limiting examples of reporter genes include: genes encoding green fluorescent protein (GFP), red fluorescent protein (RFP), luciferase, secretory alkaline phosphatase (SEAP), or any other gene whose expression is detectable and apparent to those skilled in the art. The reporter system can be used to test many different gRNAs, for example, to determine which residues each deaminase will target relative to the target DNA sequence. The off-target effects of sgRNAs targeting non-template strands can also be tested to evaluate specific base-editing proteins (e.g., Cas9 deaminase fusion proteins). In some embodiments, the gRNA can be engineered such that the start codon of the mutation will not pair with the gRNA bases. The guide polynucleotide may comprise a standard ribonucleotide, a modified ribonucleotide (e.g., pseudouridine), a ribonucleotide isomer, and / or a ribonucleotide analog. In some embodiments, the guide polynucleotide may comprise at least one detectable tag. The detectable tag may be a fluorophore (e.g., FAM, TMR, Cy3, Cy5, Texas Red, Oregon Green, Alexa Fluors, Halo tags, or a suitable fluorescent dye), a detection tag (e.g., biotin, foxglove ligand, and the like), a quantum dot, or a gold particle.
[0364] The guide RNA can be synthesized chemically, enzymatically, or in combination. For example, the guide RNA can be synthesized using a standard phosphorylation-solid-phase synthesis method. Alternatively, the guide RNA can be synthesized in vitro by operatively linking DNA encoding the guide RNA to a promoter control sequence recognized by a phage RNA polymerase. Examples of suitable phage promoter sequences include the T7, T3, and SP6 promoter sequences, or variants thereof. In embodiments where the guide RNA comprises two separate molecules (e.g., crRNA and tracrRNA), the crRNA can be synthesized chemically and the tracrRNA can be synthesized enzymatically.
[0365] In some implementations, the base editor system may comprise multiple guide polynucleotides, such as gRNAs. For example, the gRNA may target one or more target loci contained in the base editor system (e.g., at least 1 gRNA, at least 2 gRNAs, at least 5 gRNAs, at least 10 gRNAs, at least 20 gRNAs, at least 30 gRNAs, at least 50 gRNAs). The multiple gRNA sequences may be tandemly configured and preferably separated by forward repeat sequences.
[0366] The DNA sequence encoding guide RNA or guide polynucleotides can also be part of the vector. Furthermore, the vector may contain additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), selectable markers or reporter sequences (e.g., GFP or antibiotic resistance genes such as puromycin), origin of replication, and the like. The DNA molecule encoding guide RNA can also be linear. The DNA molecule encoding guide RNA or guide polynucleotides can also be circular.
[0367] In some implementations, one or more components of the base editor system can be encoded by a DNA sequence. This DNA sequence can be introduced into the expression system (e.g., a cell) together or separately. For example, a DNA sequence encoding a polynucleotide programmable nucleotide binding domain and a guide RNA can be introduced into the cell. The DNA sequences can be parts of different molecules (e.g., a vector containing the coding sequence for the polynucleotide programmable nucleotide binding domain and a second vector containing the coding sequence for the guide RNA) or both can be parts of the same molecule (e.g., a vector containing both the coding (and regulatory) sequences for the polynucleotide programmable nucleotide binding domain and the guide RNA).
[0368] When a DNA sequence encoding an RNA-guided intranuclease and a guide RNA are introduced into a cell, the DNA sequences may be part of an individual molecule (e.g., a vector containing the RNA-guided intranuclease coding sequence and a second vector containing the guide RNA coding sequence) or both may be part of the same molecule (e.g., a vector containing the coding (and regulatory) sequences of both the RNA-guided intranuclease and the guide RNA).
[0369] Leading polynucleotides may contain one or more modifications to provide nucleic acids with novel or enhanced characteristics. Leading polynucleotides may contain nucleic acid affinity tags. Leading polynucleotides may contain synthetic nucleotides, synthetic nucleotide analogs, nucleotide derivatives, and / or modified nucleotides.
[0370] In some cases, the gRNA or guiding polynucleotide may contain modifications. Modifications can be performed at any position on the gRNA or guiding polynucleotide. More than one modification may be performed on a single gRNA or guiding polynucleotide. Quality control of the gRNA or guiding polynucleotide may be performed after modification. In some cases, quality control may include PAGE, HPLC, MS, or any combination thereof.
[0371] Modifications to gRNA or guide polynucleotides may include substitution, insertion, deletion, chemical modification, physical modification, stabilization, purification, or any combination thereof.
[0372] gRNA or guide polynucleotides can also be modified through the following: 5' adenosine, 5' guanosine-triphosphate cap, 5' N7-methylguanosine-triphosphate cap, 5' triphosphate cap, 3' phosphate, 3' thiophosphate, 5' phosphate, 5' thiophosphate, cis-ipsilateral thymidine dimer, trimer, C12 spacer, C3 spacer, C6 spacer, d spacer, PC spacer, r spacer, spacer 18, spacer 9, 3'-3' modification, 5'-5' modification, base removal, acridine, azobenzene, biotin, biotin BB, biotin TEG, cholesterol TEG, desulfurized biotin TEG, DNP TEG, DNP-X, DOTA, dT-biotin, double biotin, PC biotin, psoralen C2, psoralen C6, TINA, 3'DABCYL, black hole quencher 1, black hole quencher 2, DABCYL SE, dT-DABCYL, IRDye QC-1, QSY-21, QSY-35, QSY-7, QSY-9, carboxyl linker, sulfur-containing linker, 2'-deoxyribonucleoside analog purine, 2'-deoxyribonucleoside analog pyrimidine, ribonucleoside analog, 2'-O-methylribonucleoside analog, sugar-modified analog, swing / universal base, fluorescent dye tag, 2'-fluoroRNA, 2'-O-methylRNA, phosphonate methyl ester, phosphodiester DNA, phosphodiester RNA, thiophosphate DNA, thioazophosphate RNA, UNA, pseudouridine-5'-triphosphate, 5'-methylcytidine-5'-triphosphate, or any combination thereof.
[0373] In some cases, the modification is permanent. In others, it is temporary. In some cases, multiple modifications are made to the gRNA or guiding polynucleotide. Modifications to gRNA or guiding polynucleotides can alter the physicochemical properties of nucleotides, such as their conformation, polarity, hydrophobicity, chemical reactivity, base-pairing interactions, or any combination thereof.
[0374] Modification can also involve thiophosphate substitution. In some cases, native phosphodiester bonds are readily and rapidly degraded by cellular nucleases; modifications involving internucleotide linkages using thiophosphate (PS) bonds can provide greater stability against hydrolysis caused by cellular degradation. Modifications can increase stability in gRNA or guide polynucleotides. Modifications can also enhance biological activity. In some cases, thiophosphate-enhanced gRNA can inhibit RNase A, RNase T1, calf serum nuclease, or any combination thereof. These properties make PS-RNA gRNA potentially more readily exposed to nucleases in vivo or in vitro. For example, the introduction of a thiophosphate (PS) bond between the final 3-5 nucleotides at the 5′ or ''- terminus of gRNA can inhibit external nuclease degradation. In some cases, thiophosphate bonds can be added throughout the gRNA to reduce attack by internal nucleases.
[0375] Original spacer sequence adjacent motif
[0376] A "protospacer adjacent motif (PAM)" or PAM-like motif refers to a 2-6 base pair DNA sequence immediately following the DNA sequence targeted by the Cas9 nuclease in the acquired immune system of CRISPR bacteria. In some embodiments, the PAM may be a 5' PAM (i.e., located upstream of the 5' end of the protospacer). In other embodiments, the PAM may be a 3' PAM (i.e., located downstream of the 5' end of the protospacer).
[0377] The protospacer adjacent motif (PAM) or PAM-like motif refers to a 2-6 base pair DNA sequence immediately following the DNA sequence targeted by the Cas9 nuclease in the acquired immune system of CRISPR bacteria. In some embodiments, the PAM may be a 5' PAM (i.e., located upstream of the 5' end of the protospacer). In other embodiments, the PAM may be a 3' PAM (i.e., located downstream of the 5' end of the protospacer). The PAM sequence is crucial for target binding, but the precise sequence depends on the type of Cas protein.
[0378] The base editors provided herein may contain CRISPR protein-derived domains that bind nucleotide sequences containing typical or atypical protospacer adjacent motif (PAM) sequences. The PAM site is a nucleotide sequence adjacent to the target polynucleotide sequence. Some embodiments of the invention provide base editors that contain all or part of a CRISPR protein with different PAM specificities. For example, Cas9 proteins (such as Cas9 (spCas9) from Streptococcus pyogenes) typically require a typical NGGPAM sequence to bind a specific nucleic acid region, where "N" in "NGG" is adenine (A), thymine (T), guanine (G), or cytosine (C), and G is guanine. The PAM may be CRISPR protein-specific and may differ between different base editors containing different CRISPR protein-derived domains. The PAM may be 5' or 3' of the target sequence. The PAM may be upstream or downstream of the target sequence. The PAM may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides in length. Typically, PAM is a length between 2 and 6 nucleotides.
[0379] Cas9 protein sequence
[0380] In some embodiments, the Cas9 domain is a Cas9 domain (SpCas9) derived from Streptococcus pyogenes. In some embodiments, the SpCas9 domain is a nuclease-activated SpCas9, a nuclease-inactivated SpCas9 (SpCas9d), or a SpCas9 cleavage enzyme (SpCas9n). In some embodiments, the SpCas9 contains a D9XD9X mutation, or a corresponding mutation in any amino acid sequence provided herein, where X is an amino acid other than D. In some embodiments, the SpCas9 contains a D9AD9A mutation, or a corresponding mutation in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain, the SpCas9d domain, or the SpCas9n domain may bind to a nucleic acid sequence having an atypical PAM. In some embodiments, the SpCas9 domain, the SpCas9d domain, or the SpCas9n domain may bind to a nucleic acid sequence having an NGG, NGA, or NGCG PAM sequence. In some embodiments, the SpCas9 domain includes one or more of the mutations D1135X, R1335X, and T1337X, or a corresponding mutation in any amino acid sequence provided herein, wherein X is any amino acid. In some embodiments, the SpCas9 domain includes one or more of the mutations D1135E, R1335Q, and T1337R, or a corresponding mutation in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain includes one or more of the mutations D1135X, R1335X, and T1337X, or a corresponding mutation in any amino acid sequence provided herein, wherein X is any amino acid. In some embodiments, the SpCas9 domain contains one or more of the D1135V, R1335Q, and T1337R mutations, or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain contains one or more of the D1135V, R1335Q, and T1337R mutations, or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain contains one or more of the D1135V, G1218X, R1335X, and T1337X mutations, or corresponding mutations in any amino acid sequence provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain contains one or more of the D1135V, G1218R, R1335Q, and T1337R mutations, or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the SpCas9 domain...In some embodiments, the SpCas9 domain contains the D1135V, G1218R, R1335Q, and T1337R mutations, or corresponding mutations in any amino acid sequence provided herein.
[0381] In some embodiments, the Cas9 domain of any fusion protein provided herein comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the Cas9 polypeptide described herein. In some embodiments, the Cas9 domain of any fusion protein provided herein comprises the amino acid sequence of any Cas9 polypeptide described herein. In some embodiments, the Cas9 domain of any fusion protein provided herein consists of the amino acid sequence of any Cas9 polypeptide described herein.
[0382] The following is an example SpCas9 sequence:
[0383]
[0384] The following provides an example SpCas9n sequence:
[0385]
[0386] The following is an example SpEQR Cas9 sequence:
[0387] In the SpEQRCas9 sequence above, residues E1135, Q1335, and R1337 (which can be mutated from D1135, R1335, and T1337 to produce SpEQRCas9) are underlined and shown in bold.
[0388] The following is an example SpVQR Cas9 sequence:
[0389] In the SpVQRCas9 sequence above, residues V1135, Q1335, and R1337 (which can be mutated from D1135, R1335, and T1337 to produce SpVQRCas9) are underlined and shown in bold.
[0390] The following provides an example SpVRER Cas9 sequence:
[0391]
[0392] The following is an example SpVRQR Cas9 sequence:
[0393]
[0394] The residues V1135, R1218, Q1335, and R1337 (which can be mutated from D1135, G1218, R1335, and T1337 to produce SpVRQR Cas9) are underlined and shown in bold.
[0395] In some embodiments, the Cas9 domain is a recombinant Cas9 domain. In some embodiments, the recombinant Cas9 domain is a SpyMacCas9 domain. In some embodiments, the SpyMacCas9 domain is a nuclease-activated SpyMacCas9, a nuclease-inactivated SpyMacCas9 (SpyMacCas9d), or a SpyMacCas9 nickase (SpyMacCas9n). In some embodiments, the SaCas9 domain, the SaCas9d domain, or the SaCas9n domain may bind to a nucleic acid sequence having an atypical PAM. In some embodiments, the SpyMacCas9 domain, the SpCas9d domain, or the SpCas9n domain may bind to a nucleic acid sequence having an NAA PAM sequence.
[0396] Example SpyMacCas9
[0397]
[0398] High-precision Cas9 domain
[0399] Some embodiments of the present invention provide high-precision Cas9 domains. In some embodiments, the high-precision Cas9 domain is an engineered Cas9 domain containing one or more mutations that reduce the electrostatic interactions between the Cas9 domain and the sugar-phosphate backbone of DNA compared to the corresponding wild-type Cas9 domain. Without wishing to be limited by any particular theory, the high-precision Cas9 domain (with reduced electrostatic interactions with the sugar-phosphate backbone of DNA) may have fewer off-target effects. In some embodiments, the Cas9 domain (e.g., a wild-type Cas9 domain) contains one or more mutations that reduce the association between the Cas9 domain and the sugar-phosphate backbone of DNA. In some embodiments, the Cas9 domain contains one or more mutations that reduce the association between the Cas9 domain and the sugar-phosphate backbone of DNA by at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, or at least 70%.
[0400] In some embodiments, any Cas9 fusion protein provided herein contains one or more of the N497X, R661X, Q695X, and / or Q926X mutations, or a corresponding mutation in any amino acid sequence provided herein, where X is any amino acid. In some embodiments, any Cas9 fusion protein provided herein contains one or more of the N497A, R661A, Q695A, and / or Q926A mutations, or a corresponding mutation in any amino acid sequence provided herein. In some embodiments, the Cas9 domain contains the D10A mutation, or a corresponding mutation in any amino acid sequence provided herein. High-precision Cas9 domains are known in the art and will be apparent to those skilled in the art. For example, high-fidelity Cas9 domains have been described in the following: Kleinstiver, BP et al., “High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects.” Nature 529, 490-495 (2016); and Slaymaker, IM et al., “High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects.” Science 351, 84-88 (2015); the full text of each is incorporated herein by reference. In the following high-fidelity Cas9 domains, mutations relative to Cas9 are shown in bold and underlined.
[0401]
[0402] In some cases, variant Cas9 proteins carry mutations such as H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A, which reduces the peptide's ability to cleave target DNA or RNA. This reduced ability of the Cas9 protein to cleave target DNA (e.g., single-stranded target DNA) retains its ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some cases, variant Cas9 proteins carry mutations such as D10A, H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A, which reduces the peptide's ability to cleave target DNA. This reduced ability of the Cas9 protein to cleave target DNA (e.g., single-stranded target DNA) retains its ability to bind to target DNA (e.g., single-stranded target DNA). In some cases, when the variant Cas9 protein carries mutations such as W476A and W1126A, or when it carries mutations such as P475A, W476A, N477A, D1125A, W1126A, and D1127A, the variant Cas9 protein does not effectively bind to the PAM sequence. Therefore, in some cases, when this variant Cas9 protein is used in a binding method, the method does not require the PAM sequence. In other words, in some cases, when this variant Cas9 protein is used in a binding method, the method may include a guide RNA, but the method can be performed without the PAM sequence (and therefore the binding specificity is provided by the target fragment of the guide RNA). Other residues can be mutated to achieve the above effect (i.e., deactivate one nuclease portion or another nuclease portion). As a non-restrictive example, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 may be altered (i.e., substituted). Furthermore, mutations other than alanine substitution are also suitable.
[0403] In some embodiments, the CRISPR protein-derived domain of the base editor may comprise all or part of a Cas9 protein having a typical PAM sequence (NGG). In other embodiments, the Cas9-derived domain of the base editor may use atypical PAM sequences. Such sequences have been described in this art and will be apparent to those skilled in the art. For example, Cas9 domains that bind atypical PAM sequences have been described in: Kleinstiver, BP et al., “Engineered CRISPR-Cas9 nucleases with altered PAM specificities”, Nature, 523, 481-485 (2015); and Kleinstiver, BP et al., “Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition”, Nature Biotechnology, 33, 1293-1298 (2015); each of which is incorporated herein by reference in its entirety.
[0404] In some instances, a PAM recognized by the CRISPR protein-derived domain of the base editor disclosed herein can be provided on a single oligonucleotide to an insertion point (e.g., an AAV insertion point) encoding the base editor. In this case, providing the PAM on a single oligonucleotide makes it possible to cleave a target sequence that would otherwise be inaccessible because no adjacent PAM is present on the same polynucleotide as the target sequence.
[0405] In one embodiment, *Streptococcus pyogenes* Cas9 (SpCas9) can be used as a CRISPR endonuclease for genome engineering. However, others can be used. In some cases, different endonucleases can be used to target certain genome targets. In some cases, synthetic SpCas9-derived variants with non-NGG PAM sequences can be used. Furthermore, other Cas9s from various species that are homologous genes and which are said to bind various PAM sequences can also be used in this invention. For example, the relatively large size of SpCas9 (approximately 4kb coding sequence) can cause plastids carrying the SpCas9 cDNA, which cannot be efficiently expressed in cells. Conversely, the coding sequence for *Staphylococcus aureus* Cas9 (SaCas9) is approximately 1,000 bases shorter than that of SpCas9, which may allow for efficient expression in cells. Similar to SpCas9, the SaCas9 endonuclease can modify target genes in mammalian cells in vitro and in vivo in mice. In some cases, the Cas protein can target different PAM sequences. In some cases, such as the target gene, it may be adjacent to the Cas9 PAM, 5'-NGG. In other cases, other Cas9 homologs may have different PAM requirements. For example, other PAMs, such as the PAMs for Streptococcus thermophilus (5'-NNAGAA for CRISPR1 and 5'-NGGNG for CRISPR3) and the PAM for Neisseria meningitidis (5'-NNNNGATT), may also be found adjacent to the target gene.
[0406] In some implementations, for the *Streptococcus pyogenes* system, the target gene sequence may be preceding (i.e., 5' to) 5'-NGGPAM, and the 20-nt guide RNA sequence may pair with the opposing base pair to mediate Cas9 cleavage adjacent to PAM. In some cases, an adjacent cleavage may be or may be about 3 base pairs upstream of PAM. In some cases, an adjacent cleavage may be or may be about 10 base pairs upstream of PAM. In some cases, an adjacent cleavage may be or may be about 0-20 base pairs upstream of PAM. For example, an adjacent cleavage may be immediately upstream of PAM and be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 base pairs upstream of PAM. An adjacent cleavage can also be 1 to 30 base pairs downstream of PAM.
[0407] Fusion proteins containing nuclear localization sequences (NLS)
[0408] In some embodiments, the fusion protein provided herein further comprises one or more (e.g., 2, 3, 4, 5) nuclear targeting sequences, such as nuclear localization sequences (NLS). In one embodiment, a bimolecular NLS is used. In some embodiments, the NLS comprises an amino acid sequence that facilitates the input of the NLS-containing protein into the cell nucleus (e.g., via nuclear transport). In some embodiments, any fusion protein provided herein further comprises a nuclear localization sequence (NLS). In some embodiments, the NLS is fused to the N-terminus of the fusion protein. In some embodiments, the NLS is fused to the C-terminus of the fusion protein. In some embodiments, the NLS is fused to the N-terminus of the Cas9 domain. In some embodiments, the NLS is fused to the C-terminus of the nCas9 domain or dCas9 domain. In some embodiments, the NLS is fused to the N-terminus of the deaminase. In some embodiments, the NLS is fused to the C-terminus of the deaminase. In some embodiments, the NLS is fused to the fusion protein via one or more linkers. In some embodiments, the NLS is fused to the fusion protein without a linker. In some embodiments, the NLS comprises the amino acid sequence of any of the NLS sequences provided or referenced herein. Additional nuclear localization sequences are known in the art and will be apparent to those skilled in the art. For example, NLS sequences are described in Plank et al. PCT / EP2000 / 011690, the contents of which are disclosed by reference and are incorporated herein by reference. In some embodiments, the NLS comprises the following amino acid sequences: PKKKRKVEGADKRTADGSEFES PKKKRKV, KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRKPKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC. In some embodiments, the NLS is present in a linker or flanked by linkers, such as those described herein. In some embodiments, the N-terminal or C-terminal NLS is a bimolecular NLS. A bimolecular NLS comprises two basic amino acid clusters separated by a relatively short spacer sequence (hence the bimolecular-2 part, unlike monomolecular NLS). The nucleoplasmic NLS (KR[PAATKKAGQA]KKKK) is a prototype of the widely distributed bimolecular message: two basic amino acid clusters separated by a spacer of about 10 amino acids. An exemplary bimolecular NLS sequence is as follows: PKKKRKVEGADKRTADGSEFES PKKKRKV.
[0409] In some embodiments, the fusion protein of the present invention does not contain a linker sequence. In some embodiments, a linker sequence is present between one or more domains or proteins.
[0410] It should be understood that the fusion protein of the present invention may include one or more additional features. For example, in some embodiments, the fusion protein may include: an inhibitor, a cytoplasmic localization sequence, an export sequence (such as a nuclear export sequence) or other localization sequence, and a sequence tag that can be used to dissolve, purify, or detect the fusion protein. Suitable protein tags provided herein include (but are not limited to): biotinylate carrier protein (BCCP) tag, myc-tag, calcitonin-tag, FLAG-tag, hemagglutinin (HA)-tag, multihistidine tag (also known as histidine tag or His-tag), maltose-binding protein (MBP)-tag, nus-tag, glutathione-S-transferase (GST)-tag, green fluorescent protein (GFP)-tag, sulfoxide-reduction protein-tag, S-tag, Softags (e.g., Softags 1, Softags 3), strep-tag, biotin conjugase tag, FLAsH tag, V5 tag, and SBP-tag. Additional suitable sequences will be apparent to those skilled in the art. In some implementations, the fusion protein contains one or more His tags.
[0411] Vectors encoding CRISPR enzymes containing one or more nuclear localization sequences (NLS) can be used. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 NLSs can be used or approximately used. The CRISPR enzyme may contain the NLS at or near the amino terminus, approximately or more than 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 NLSs at or near the carboxyl terminus, or any combination thereof (e.g., one or more NLSs at the amino terminus and one or more NLSs at the carboxyl terminus). When more than one NLS is present, each can be selected independently of the other NLSs, such that a single NLS can be present in more than one copy and / or combined with one or more other NLSs present in one or more copies.
[0412] The CRISPR enzyme used in this method may contain about 6 NLS. An NLS is considered to be close to the N- or C-terminus when the amino acid closest to it is within about 50 amino acids along the polypeptide chain from the N-terminus or C-terminus (e.g., within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, or 50 amino acids).
[0413] In some embodiments, the NLS is present in a linker or flanked by linkers, such as those described herein. In some embodiments, the N-terminal or C-terminal NLS is a bimolecular NLS. A bimolecular NLS comprises two basic amino acid clusters separated by a relatively short spacer sequence (hence the bimolecular-2 part, unlike a monomolecular NLS). The nucleoprotein NLS (KR[PAATKKAGQA]KKKK) is a prototype of the widely distributed bimolecular message: two clusters of basic amino acids separated by a spacer of about 10 amino acids. An exemplary bimolecular NLS sequence is as follows: PKKKRKVEGADKRTADGSEFES PKKKRKV.
[0414] In some embodiments, the NLS is present in a linker or flanked by linkers, such as those described herein. In some embodiments, the N-terminus or C-terminus of the NLS is a bimolecular NLS. A bimolecular NLS comprises two basic amino acid clusters separated by a relatively short spacer sequence (hence the bimolecular-2 part, unlike monomolecular NLS). The nucleoplasmic NLS (KR[PAATKKAGQA]KKKK) is a prototype of the widely distributed bimolecular message: two basic amino acid clusters separated by approximately 10 amino acids. An exemplary bimolecular NLS sequence is as follows: PKKKRKVEGADKRTADGSEFES PKKKRKV.
[0415] The PAM sequence may be any PAM sequence known in the art. Suitable PAM sequences include (but are not limited to): NGG, NGA, NGC, NGN, NGT, NGCG, NGAG, NGAN, NGNG, NGCN, NGCG, NGTN, NNGRRT, NNNRRT, NNGRR(N), TTTV, TYCV, TYCV, TATV, NNNGATT, NNAGAAW, or NAAMC. Y is pyrimidine; N is any nucleotide base; W is A or T.
[0416] Cas9 domain with reduced repulsion
[0417] Typically, Cas9 proteins (such as Cas9 (spCas9) from Streptococcus pyogenes) require a typical NGG PAM sequence to bind to a specific nucleic acid region, where the "N" in "NGG" is adenosine (A), thymidine (T), or cytosine (C), and the "G" is guanosine. This can limit the ability to edit the desired bases within the gene. In some embodiments, the base-editing fusion proteins provided herein may need to be placed at precise locations, such as regions containing the target base upstream of the PAM. See, for example, Komor, AC et al., "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage," Nature 533, 420-424 (2016), the full text of which is incorporated herein by reference. Therefore, in some embodiments, any fusion protein provided herein may contain a Cas9 domain that can bind nucleotide sequences that do not contain a typical (e.g., NGG) PAM sequence. Cas9 domains that bind to atypical PAM sequences have been described in this art and will be apparent to those skilled in the art. For example, Cas9 domains that bind atypical PAM sequences have been described in the following: Kleinstiver, BP et al., “Engineered CRISPR-Cas9 nucleases with altered PAM specificities”, Nature 523, 481-485 (2015); Kleinstiver, BP et al., “Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition”, Nature Biotechnology 33, 1293-1298 (2015); and Nishimasu, H. et al., “Engineered CRISPR-Cas9 nuclease with expanded targeting space”, Science. 2018 Sep 21.
[0418] 361(6408):1259-1262, Chatterjee, P. et al., Minimal PAM specificity of a highly similar SpCas9 ortholog, SciAdv. 2018 Oct 24; 4(10):eaau0766.doi:10.1126 / sciadv.aau0766; all of which are incorporated herein by reference in their entirety. Several PAM variants are described in the following table:
[0419] Table 1. Cas9 protein and corresponding PAM sequence
[0420]
[0421]
[0422] Nucleobase editing domain
[0423] This article describes a base editor comprising a fusion protein including a polynucleotide programmable nucleotide-binding domain and a nucleobase (base) editing domain (e.g., a deaminase domain). The base editor can be programmed to edit one or more bases in the target polynucleotide sequence by interacting with a guide polynucleotide that recognizes the target sequence. Once the target sequence has been recognized, the base editor is anchored to the polynucleotide to be edited, and then the deaminase domain component of the base editor edits a target base.
[0424] In some embodiments, the nucleobase editing domain is a deaminase domain. In some cases, the deaminase domain may be cytosine deaminase or cytidine deaminase. In some embodiments, the terms "cytosine deaminase" and "cytidine deaminase" are used interchangeably. In some cases, the deaminase domain may be adenine deaminase or adenine deaminase. In some embodiments, the terms "adenine deaminase" and "adenine deaminase" are used interchangeably. Detailed descriptions of nucleobase editing proteins are found in International PCT Applications PCT / 2017 / 045381 (WO2018 / 027078) and PCT / US2016 / 058344 (WO2017 / 070632), the full text of which is incorporated herein by reference. See also: Komor, AC et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage”, Nature 533, 420-424 (2016); Gaudelli, NM et al., “Programmable base editing of A·T to G·Cin genomic DNA without DNA cleavage”, Nature 551, 464-471 (2017); and Komor, AC et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity”, Science Advances 3:eaao4774 (2017), all of which are incorporated herein by reference in their entirety.
[0425] C to T editor
[0426] In some embodiments, the base editor disclosed herein comprises a fusion protein containing a cytidine deaminase that deaminates a target cytidine (C) base of a polynucleotide to produce uridine (U), which has the base-pairing properties of thymine. In some embodiments, for example, in the case that the polynucleotide is double-stranded (e.g., DNA), the uridine base can then be substituted with a thymidine base (e.g., via a cellular repair mechanism) to produce a C:G to T:A conversion. In other embodiments, C to U deamination in nucleic acids via the base editor cannot be achieved by U to T substitution.
[0427] The deamination of a target C in a polynucleotide to produce a U is a non-restricted example of base editing, which can be performed by the base editor described herein. In another example, a base editor containing a cytidine deaminase domain can mediate the conversion of a cytosine (C) base to a guanine (G) base. For example, the U of a polynucleotide produced by deaminating cytidine with the cytidine deaminase domain of a base editor can be removed from the polynucleotide by a base excision repair mechanism (e.g., via a uracil DNA glycosidase (UDG) domain), thereby producing a debasement site. The nucleobase relative to the debasement site can then be replaced with another base (such as C) via, for example, by a bypass polymerase (e.g., via a base repair mechanism). Although the nucleobase relative to the debasement site will typically be replaced with C, other substitutions (e.g., A, G, or T) can also occur.
[0428] Therefore, in some embodiments, the base editor described herein includes a deamination domain (e.g., a cytidine deaminase domain) that deaminates a target C in a polynucleotide to U. Further, as described below, the base editor may include additional domains that facilitate the conversion of U (in some embodiments) to T or G caused by deamination. For example, a base editor including a cytidine deaminase domain may further include a uracil glycosidase inhibitor (UGI) domain to mediate the substitution of U with T, thereby completing the C-to-T base editing event. In another example, the base editor may incorporate a bypass polymerase to improve the efficiency of C-to-G base editing, since the bypass polymerase facilitates the incorporation of C opposite the deamination site (i.e., causing G to be incorporated at the deamination site, completing the C-to-G base editing event).
[0429] A base editor incorporating a cytidine deaminase domain can deaminate a target C in any polynucleotide, including DNA, RNA, and DNA-RNA hybrids. Typically, cytidine deaminase catalyzes the C nucleobase located within the single-stranded portion of a polynucleotide. In some embodiments, the entire polynucleotide containing the target C may be single-stranded. For example, a cytidine deaminase incorporated into the base editor can deaminate a target C in a single-stranded RNA polynucleotide. In other embodiments, a base editor incorporating a cytidine deaminase domain can act on a double-stranded polynucleotide, but the target C may be located within a portion of the polynucleotide, which is single-stranded during the deaminating reaction. For example, in an embodiment where the NAGPB domain includes a Cas9 domain, several nucleotides may remain unpaired during the formation of the Cas9-gRNA-target DNA complex, resulting in the formation of a Cas9 "R-loop complex." The unpaired nucleotides can form a single-stranded DNA vesicle, which can serve as a substrate for a single-stranded specific nucleotide deaminase (e.g., cytidine deaminase).
[0430] In some embodiments, the base editor's cytidine deaminase may comprise all or part of the lipoprotein B mRNA editing complex (APOBEC) family of deaminases. APOBEC is an evolutionarily conserved family of cytidine deaminases. Members of this family are C-to-U editing enzymes. The N-terminal domain of APOBEC-like proteins is the catalytic domain, while the C-terminal domain is the pseudocatalytic domain. More specifically, this catalytic domain is a zinc-dependent cytidine deaminase domain and is essential for cytidine deamination. APOBEC family members include: APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D ("APOBEC3E" currently refers to this class), APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, and activation-inducible (cytidine) deaminases. Several modified cytidine deaminases are commercially available, including (but not limited to): SaBE3, SaKKH-BE3, VQR-BE3, EQR-BE3, VRER-BE3, YE1-BE3, EE-BE3, YE2-BE3, and YEE-BE3, which are available from Addgene (plastomers 85169, 85170, 85171, 85172, 85173, 85174, 85175, 85176, 85177). In some embodiments, the deaminase incorporated into the base editor comprises all or part of the APOBEC1 deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or part of the APOBEC2 deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or part of the APOBEC3 deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or part of the APOBEC3A deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or part of APOBEC3B deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or part of APOBEC3C deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or part of APOBEC3D deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or part of APOBEC3E deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or part of APOBEC3F deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or part of APOBEC3G deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or part of APOBEC3H deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or part of APOBEC4 deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or part of an activation-inducible deaminase (AID).
[0431] In some embodiments, the deaminase incorporated into the base editor comprises all or part of cytidine deaminase 1 (CDA1). It should be understood that the base editor may contain a deaminase from any suitable organism (e.g., human or rat). In some embodiments, the deaminase domain of the base editor is derived from human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase domain of the base editor is derived from rats (e.g., rat APOBEC1). In some embodiments, the deaminase domain of the base editor is human APOBEC1. In some embodiments, the deaminase domain of the base editor is pmCDA1.
[0432] The base and amino acid sequences of PmCDA1 and the CDS of human AID are shown below.
[0433] >tr|A5H718|A5H718_PETM Cytosine deaminase OS=Laker OX=7757PE=2SV=1: MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYL RDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKRRSELSIMIQVKILHTTKSPAV
[0434] Nucleic acid sequence: >EF094822.1 PmCDA.21 cytosine deaminase mRNA isolated from hagfish, complete CDS:
[0435] TGACACGACACAGCCGTGTATATGAGGAAGGGTAGCTGGATGGGGGGGGGGGGAATACGTTCAGAGAGGACATTAGCGAGCGTCTTGTTGGTGGCCTTGAGTCTAGACACCTGCAGACATGACCGACGCTGAGTACGTGAGAATCCATGAGAAGTTGGACATCTACACGTTTAAGAAACAGTTTTTCAACA ACAAAAAATCCGTGTCGCATAGATGCTACGTTCCTTTGAATTAAAACGACGGGGTGAACGTAGAGCGTGTTTTGGGGGCTATGCTGTGAATAAACCACAGAGCGGGACAGAACGTGGAATTCACGCCGAAATCTTTAGCATTAGAAAAGTCGAAGAATACCTGCGCGACAACCCCGGACAATTCACGATAA ATTGGTACTCATCCTGGAGTCCTTGTGCAGATTGCGCTGAAAAGATCTTAGAATGGTATAACCAGGAGCTGCGGGGGAACGGCCACACTTTGAAAATCTGGGCTTGCAAACTCTATTACGAGAAAAATGCGAGGAATCAAATTGGGCTGTGGAACCTCAGAGATAACGGGGTTGGGTTGAATGTAATGGTA AGTGAACACTACCAAATGTTGCAGGAAAATATTCATCCAATCGTCGCACAATCAATTGAATGAGAATAGATGGCTTGAGAAGACTTTGAAGCGAGCTGAAAAACGACGGAGCGAGTTGTCCATTATGATTCAGGTAAAAATACTCCACACCACTAAGAGTCCTGCTGTTTAAGAGGCTATGCGGATGGTTTTC
[0436] The amino acid and nucleic acid sequences of the coding sequence (CDS) for human activation-inducible cytidine deaminase (AID) are shown below:
[0437] >tr|Q6QJ80|Q6QJ80_Human Activated-Inducible Cytidine Deaminase OS=Homo sapiens OX=9606GN=AICDA PE=2SV=1
[0438] MDSLMNRRKFLYQFKNVRWAKGRRETYLCYVVKRRDSATSSFLDFGYLRNKNGCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGNPNLSLRIFTARLYFCEDRKAEPEGLRRLHRAGVQIAIMTFKAPV
[0439] Nucleic acid sequence: >NG_011588.1:5001-15681 Homo sapiens activation-induced cytidine deaminase (AICDA), RefSeqGene (LRG_17) on chromosome 12:
[0440]
[0441] Other exemplary deaminases that can be fused to Cas9 according to embodiments of the present invention are provided below. It should be understood that in some embodiments, activation domains of individual sequences may be used, for example, domains without localization information (nuclear localization sequences, without nuclear export information, cytoplasmic localization information).
[0442] Human AID:
[0443]
[0444] (Bottom line: Core location sequence; Double bottom line: Core output message)
[0445] Mouse AID:
[0446]
[0447] (Bottom line: Core location sequence; Double bottom line: Core output message)
[0448] Canine AID:
[0449]
[0450] (Bottom line: Core location sequence; Double bottom line: Core output message) Cow AID:
[0451]
[0452] (Bottom line: Nuclear localization sequence; Double bottom line: Nuclear output message) Rat AID
[0453]
[0454] (Bottom line: nuclear localization sequence; double bottom line: nuclear output message) Mouse APOBEC-3
[0455]
[0456] (Italicized text: Nucleic acid editing domain)
[0457] Rat APOBEC-3:
[0458]
[0459]
[0460] (Italicized: Nucleic acid editing domain)
[0461] Rhesus monkey APOBEC-3G:
[0462]
[0463] (Italics: Nucleic acid editing domain; Underline: Cytoplasmic localization information)
[0464] Chimpanzee APOBEC-3G:
[0465]
[0466] (Italics: Nucleic acid editing domain; Underline: Cytoplasmic localization information)
[0467] Green Monkey APOBEC-3G:
[0468]
[0469] (Italics: Nucleic acid editing domain; Underline: Cytoplasmic localization information)
[0470] Human APOBEC-3G:
[0471]
[0472] (Italics: Nucleic acid editing domain; Underline: Cytoplasmic localization information)
[0473] Human APOBEC-3F:
[0474]
[0475] (Italicized: Nucleic acid editing domain)
[0476] Human APOBEC-3B:
[0477]
[0478] (Italicized: Nucleic acid editing domain)
[0479] Rat APOBEC-3B:
[0480] MQPQGLGPNAGMGPVCLGCSHRRPYSPIRNPLKKLYQQTFYFHFKNVRYAWGRKNNFLCYEVNGMDCALPVPLRQGVFRKQGHIHAELCFIYWFHDKVLRVLSPMEEFKVTWYMSWSPCSKCAEQVARFLAAHRNLSLAIFSSRLYYYLRNPNYQQKLCRLIQEGVHVAAMDLPEFKKCWNKFVDNDGQPFRPWMRLRINFSFYDCKLQEIFSRMNLLREDVFYLQFNNSHRVKPVQNRYYRRKSYLCYQLERANGQEPLKGYLLYKKGEQHVEILFLEKMRSMELSQVRITCYLTWSPCPNCARQLAAFKKDHPDLILRIYTSRLYFWRKKFQKGLCTLWRSGIHVDVMDLPQFADCWTNFVNPQRPFRPWNELEKNSWRIQRRLRRIKESWGL(SEQ ID NO:)
[0481] Bovine APOBEC-3B:
[0482] DGWEVAFRSGTVLKAGVLGVSMTEGWAGSGHPGQGACVWTPGTRNTMNLLREVLFKQQFGNQPRVPAPYYRRKTYLCYQLKQRNDLTLDRGCFRNKKQRHAERFIDKINSLDLNPSQSYKIICYITWSPCPNCANELVNFITRNNHLKLEIFASRLYFHWIKSFKMGLQDLQNAGISVAVMTHTEFEDCWEQFVDNQSRPFQPWDKLEQYSASIRRRLQRILTAPI(SEQ ID NO:)
[0483] Chimpanzee APOBEC-3B:
[0484] MNPQIRNPMEWMYQRTFYYNFENEPILYGRSYTWLCYEVKIRRGHSNLLWDTGVFRGQMYSQPEHHAEMCFLSWFCGNQLSAYKCFQITWFVSWTPCPDCVAKLAKFLAEHPNVTLTISAARL YYYWERDYRRALCRLSQAGARVKIMDDEEFAYCWENFVYNEGQPFMPWYKFDDNYAFLHRTLKEIIRHLMDPDTFTFNFNNDPLVLRRHQTYLCYEVERLDNGTWVLMDQHMGFLCNEAKNLLC GFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGQVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFEYCWDTFVYRQGCPFQPWDGLEEHSQALS GRLRAILQVRASSLCMVPHRPPPPPQSPGPCLPLCSEPPLGSLLPTGRPAPSLPFLLTASFSFPPPASLPPLPSLSLSPGHLPVPSFHSLTSCSIQPPCSSRIRETEGWAVSKEGRDLG(SEQ ID NO:)
[0485] Human APOBEC-3C:
[0486]
[0487] (Italicized: Nucleic acid editing domain)
[0488]
[0489] Human APOBEC-3A:
[0490]
[0491] (Italicized: Nucleic acid editing domain)
[0492] Rhesus monkey APOBEC-3A:
[0493]
[0494] (Italicized: Nucleic acid editing domain)
[0495] Bovine APOBEC-3A:
[0496]
[0497] (Italicized: Nucleic acid editing domain)
[0498] Human APOBEC-3H:
[0499]
[0500] (Italicized: Nucleic acid editing domain)
[0501] Rhesus monkey APOBEC-3H:
[0502] MALLTAKTFSLQFNNKRRVNKPYYPRKALLCYQLTPQNGSTPTRGHLKNKKKDHAEIRFINKIKSMGLDETQCYQVTCYLTWSPCPSCAGELVDFIKAHRHLNLRIF ASRLYYHWRPNYQEGLLLLLCGSQVPVEVMGLPEFTDCWENFVDHKEPPSFNPSEKLEELDKNSQAIKRRLERIKSRSVDVLENGLRSLQLGPVTPSSSIRNSR(SEQ IDNO:)
[0503] Human APOBEC-3D:
[0504]
[0505]
[0506] (Italicized: Nucleic acid editing domain)
[0507] Human APOBEC-1:
[0508] MTSEKGPSTGDPTLRRRIEPWEFDVFYDPRELRKEACLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERDFHPSMSCSITWFLSWSPCWECSQAIREFLSRHPGVTLVIYVARLF WHMDQQNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYPPGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLTFFRLHLQNCHYQTIPPHILLATGLIHPSVAWR(SEQ ID NO:)
[0509] Mouse APOBEC-1:
[0510] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSVWRHTSQNTSNHVEVNFLEKFTTERYFRPNTRCSITWFLSWSPCGECSRAITEFLSRHPYVTLFIYIARLYHHTDQRNRQGLRDLISSGVTIQIMTEQEYCYCWRNFVNYPPSNEAYWPRYPHLWVKLYVLELYCIILGLPPCLKILRRKQPQLTFFTITLQTCHYQRIPPHLLWATGLK(SEQ ID NO:)
[0511] Rat APOBEC-1:
[0512] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK(SEQ ID NO:)
[0513] Human APOBEC-2:
[0514] MAQKEEAAVATEAASQNGEDLENLDDPEKLKELIELPPFEIVTGERLPANFFKFQFRNVEYSSGRNKTFLCYVVEAQGKGGQVQASRGYLEDEHAAAHAEEAFFNTILPAFDPALRYNVTWYVSSSPCAACADRIIKTLSKTKNLRLLILVGRLFMWEEPEIQAALKKLKEAGCKLRIMKPQDFEYVWQNFVEQEEGESKAFQPWEDIQENFLYYEEKLADILK(SEQ ID NO:)
[0515] Mouse APOBEC-2:
[0516] MAQKEEAAEAAAPASQNGDDLENLEDPEKLKELIDLPPFEIVTGVRLPVNFFKFQFRNVEYSSGRNKTFLCYVVEVQSKGGQAQATQGYLEDEHAGAHAEEAFFNTILPAFDPALKYNVTWYVSSSPCAACADRILKTLSKTKNLRLLILVSRLFMWEEPEVQAALKKLKEAGCKLRIMKPQDFEYIWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK(SEQ ID NO:)
[0517] Rat APOBEC-2
[0518] MAQKEEAAEAAAPASQNGDDLENLEDPEKLKELIDLPPFEIVTGVRLPVNFFKFQFRNVEYSSGRNKTFLCYVVEAQSKGGQVQATQGYLEDEHAGAHAEEAFFNTILPAFDPALKYNVTWYVSSSPCAACADRILKTLSKTKNLRLLILVSRLFMWEEPEVQAALKKLKEAGCKLRIMKPQDFEYLWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK(SEQ ID NO:)
[0519] Bovine APOBEC-2
[0520] MAQKEEAAAAAEPASQNGEEVENLEDPEKLKELIELPPFEIVTGERLPAHYFKFQFRNVEYSSGRNKTFLCYVVEAQSKGGQVQASRGYLEDEHATNHAEEAFFNSIMPTFDPALRYMVTWYVSSSPCAACADRIVKTLNKTKNLRLLILVGRLFMWEEPEIQAALRKLKEAGCRLRIMKPQDFEYIWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK(SEQ ID NO:)
[0521] Lamprey CDA1 (pmCDAl)
[0522] MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKRRSELSFMIQVKILHTTKSPAV(SEQ ID NO:)
[0523] Human APOBEC3G D316R D317R
[0524] MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPPLDAKIFRGQ VYSELKYHPEMRFFHWFSKWRKLHRDQEYEVTWYISWSPCTKCTRDMATFLAEDPKVTLTIFVARLYYFWDPDYQEALRSLCQKRDGP Rat MKF NYDEFQHCWSKFVYSQRELFEPWNNLPKYYILLHFMLGEILRHSMDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISK KHVSLCIFTARIYRRQGRCQEGLRTLAEAGAKISF TYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQNQEN
[0525] (SEQ ID NO:)
[0526] Human APOBEC3G Chain A
[0527] MDPPTFTFNFNNEPWWGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYDDQGRCQEGLRTLAEAGAKISF TYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQ(SEQ ID NO:)
[0528] Human APOBEC3G Chain D120R D121R
[0529] MDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYRRQGRCQEGLRTLAEAGAKISFMTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQ(SEQ ID NO:)
[0530] Some embodiments of the present invention are based on the identification of the catalytic activity of the deaminase domain of any fusion protein described herein, for example, by affecting the continuity of the fusion protein (e.g., a base editor) through point mutations in the deaminase domain. For example, mutations that reduce (but do not eliminate) the catalytic activity of the deaminase domain within a base-editing fusion protein make it less likely that the deaminase domain will catalyze deamination of residues adjacent to target residues, thereby narrowing the permissible range of deamination. Narrowing the permissible range of deamination can prevent unwanted deamination of residues adjacent to specific target residues, which can reduce or avoid off-target effects.
[0531] By way of example, in some embodiments, the APOBEC deaminase incorporating a base editor may contain one or more mutations selected from the group consisting of: H121X, H122X, R126X, R126X, R118X, W90X, W90X, and R132X of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase, wherein X is any amino acid. In some embodiments, the APOBEC deaminase incorporating a base editor may contain one or more mutations selected from the group consisting of: H121R, H122R, R126A, R126E, R118A, W90A, W90Y, and R132E of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase.
[0532] In some embodiments, the APOBEC deaminase incorporated into the base editor may contain one or more mutations selected from the group consisting of: D316X, D317X, R320X, R320X, R313X, W285X, W285X, R326X of hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase, wherein X is any amino acid. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing one or more mutations selected from the group consisting of: D316R, D317R, R320A, R320E, R313A, W285A, W285Y, R326E of hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase.
[0533] In some embodiments, the APOBEC deaminase incorporating a base editor may contain the H121R and H122R mutations of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporating a base editor may contain the R126A mutation of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporating a base editor may contain the R126E mutation of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporating a base editor may contain the R118A mutation of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporating a base editor may contain a W90A mutation of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporating a base editor may contain a W90Y mutation of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporating a base editor may contain an R132E mutation of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporating a base editor may contain both W90Y and R126E mutations of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporating a base editor may comprise an APOBEC deaminase containing the R126E and R132E mutations of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporating a base editor may comprise an APOBEC deaminase containing the W90Y and R132E mutations of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporating a base editor may comprise an APOBEC deaminase containing the W90Y, R126E, and R132E mutations of rAPOBEC1, or one or more corresponding mutations in another APOBEC deaminase.
[0534] In some embodiments, the APOBEC deaminase incorporated with a base editor may comprise an APOBEC deaminase containing the D316R and D317R mutations of hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, any fusion protein provided herein comprises an APOBEC deaminase containing the R320 mutation of hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated with a base editor may comprise an APOBEC deaminase containing the R320E mutation of hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporated with a base editor may comprise an APOBEC deaminase containing the R313 mutation of hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporating a base editor may contain a W285 mutation with hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporating a base editor may contain a W285Y mutation with hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporating a base editor may contain an R326E mutation with hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporating a base editor may contain both W285Y and R320E mutations with hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporating a base editor may comprise an APOBEC deaminase containing the R320E and R326E mutations of hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporating a base editor may comprise an APOBEC deaminase containing the W285Y and R326E mutations of hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase. In some embodiments, the APOBEC deaminase incorporating a base editor may comprise an APOBEC deaminase containing the W285Y, R320E, and R326E mutations of hAPOBEC3G, or one or more corresponding mutations in another APOBEC deaminase.
[0535] Some modified cytidine deaminases are commercially available, including (but not limited to): SaBE3, SaKKH-BE3, VQR-BE3, EQR-BE3, VRER-BE3, YE1-BE3, EE-BE3, YE2-BE3, and YEE-BE3, derived from Addgene (plastids 85169, 85170, 85171, 85172, 85173, 85174, 85175, 85176, and 85177).
[0536] Detailed descriptions of C-to-T nucleobase editing proteins can be found in international PCT application number PCT / US2016 / 058344 (WO2017 / 070632) and Komor, AC et al.'s "Programmable editing of a target base ingenomic DNA without double-stranded DNA cleavage" Nature 533, 420-424 (2016), the full text of which is incorporated herein by reference.
[0537] A to G editors
[0538] In some embodiments, the base editor described herein may include a deaminase domain comprising adenosine deaminase. The adenosine deaminase domain of such a base editor facilitates the editing of adenine (A) nucleotides to guanine (G) nucleotides by deaminating A to form inosine (I), exhibiting the base-pairing properties of G. Adenosine deaminase deaminates (i.e., removes the amino group) of adenine residues at deoxyadenosine residues in deoxyribonucleic acid (DNA).
[0539] In some embodiments, the nucleobase editors provided herein can be made by fusing one or more protein domains together to produce a fusion protein. In some embodiments, the fusion proteins provided herein include one or more features that improve the base editing activity (e.g., efficiency, selectivity, and specificity) of the fusion protein. For example, the fusion proteins provided herein may include a Cas9 domain with reduced nuclease activity. In some embodiments, the fusion proteins provided herein may have a Cas9 domain without nuclease activity (dCas9), or a Cas9 domain that cleaves one strand of a double DNA molecule, referred to as Cas9 cleavage enzyme (nCas9). Without being bound by any particular theory, the presence of a catalytic residue (e.g., H840) maintains the activity of the Cas9 to cleave the unedited (e.g., undeaminated) strand containing a T relative to the targeted A. Mutations in the catalytic residues of Cas9 (e.g., D10 to A10) prevent cleavage of the edited strand containing the targeted A residue. This Cas9 variant can generate single-strand DNA breaks (gaps) at specific locations based on a gRNA-defined target sequence to repair unedited strands, ultimately resulting in T-C alterations on those unedited strands. In some embodiments, the A-G base editor further includes an inhibitor of inosine base excision repair, such as a uracil glycosidase inhibitor (UGI) domain or a catalytically inactive inosine-specific nuclease. Without wishing to be limited by any particular theory, the UGI domain or the catalytically inactive inosine-specific nuclease can inhibit or prevent base excision repair of deaminoadenosine residues (e.g., inosine), which may improve the activity or efficiency of the base editor.
[0540] A base editor containing adenosine deaminase can act on any polynucleotide, including DNA, RNA, and DNA-RNA hybrids. In some embodiments, the base editor containing adenosine deaminase can deaminate target A in a polynucleotide containing RNA. For example, the base editor may contain an adenosine deaminase domain that can deaminate target A in RNA polynucleotides and / or DNA-RNA hybrid polynucleotides. In one embodiment, the adenosine deaminase incorporated into the base editor comprises all or part of an adenosine deaminase that acts on RNA (ADAR, e.g., ADAR1 or ADAR2). In another embodiment, the adenosine deaminase incorporated into the base editor comprises all or part of an adenosine deaminase that acts on tRNA (ADAT). A base editor containing an adenosine deaminase domain can also deaminate the A nucleobase of a DNA polynucleotide. In one embodiment, the adenosine deaminase domain of the base editor comprises all or part of an ADAT containing one or more mutations, which allows the ADAT to deaminate target A in DNA. For example, the base editor may comprise all or part of ADAT (EcTadA) from *E. coli*, containing one or more of the following mutations: D108N, A106V, D147Y, E155V, L84F, H123Y, I157F, or a corresponding mutation in another adenosine deaminase. In some embodiments, the TadA deaminase is an *E. coli* TadA (ecTadA) deaminase or a fragment thereof. For example, the truncated ecTadA may lose one or more N-terminal amino acids relative to full-length ecTadA. In some embodiments, the truncated ecTadA may lose 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to full-length ecTadA. In some embodiments, the truncated ecTadA may lose 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to the full-length ecTadA. In some embodiments, the ecTadA deaminase does not contain an N-terminal methionine. In some embodiments, the TadA deaminase is an N-terminal TadA. In a particular embodiment, the TadA is any of the TadA described in PCT / US2017 / 045381, the entire text of which is incorporated herein by reference.
[0541] The adenosine deaminase can be derived from any suitable organism. In some embodiments, the adenosine deaminase is derived from prokaryotes. In some embodiments, the adenosine deaminase is derived from bacteria. In some embodiments, the adenosine deaminase is derived from *Escherichia coli*, *Staphylococcus aureus*, *Salmonella typhi*, *S. putrefaciens*, *Haemophilus influenzae*, *S. lunulatus*, or *Bacillus*. In some embodiments, the adenosine deaminase is derived from *Escherichia coli*. In some embodiments, the adenine deaminase is a naturally occurring adenosine deaminase comprising one or more mutations corresponding to any mutations provided herein (e.g., mutations in ecTadA). Relevant residues in any homologous protein can be identified by, for example, sequence alignment and determination of homologous residues. Mutations in any naturally occurring adenosine deaminase (e.g., homology to ecTadA) can thus be generated, corresponding to any mutation described herein (e.g., any mutation identified in ecTadA).
[0542] TadA (tRNA adenosine deaminase A)
[0543] In a particular implementation, the TadA is any of the TadA described in PCT / US2017 / 045381 (WO 2018 / 027078), the entire text of which is incorporated herein by reference.
[0544] In one embodiment, the fusion protein of the present invention comprises wild-type TadA linked to TadA7.10, which is linked to the Cas9 nickase. In a particular embodiment, the fusion protein comprises a single TadA7.10 domain (e.g., provided as a monomer). In other embodiments, the ABE7.10 editor comprises TadA7.10 and TadA (wt), which can form a heterodimer. The relevant amino acid sequences are as follows:
[0545] (M)SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD, which is called the "TadA reference sequence" or wild-type TadA (TadA(wt)).
[0546] The TadA7.10 amino acid sequence is as follows:
[0547] (M)SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD
[0548] In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the amino acid sequences presented in any adenosine deaminase provided herein. It should be understood that the adenosine deaminases provided herein may include one or more mutations (e.g., any mutations provided herein). The present invention provides any deaminase domain having a certain percentage of consistency plus any mutations or combinations thereof described herein. In some embodiments, the adenosine deaminase comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to a reference sequence or any of the adenosine deaminases provided herein. In some embodiments, the adenosine deaminase comprises an amino acid sequence having at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, or at least 170 identical consecutive amino acid residues compared to any amino acid sequence known in the art or described herein.
[0549] In some embodiments, the TadA deaminase is a whole-cell Escherichia coli TadA deaminase. For example, in some embodiments, the adenosine deaminase comprises the following amino acid sequence:
[0550] MRRAFITGVFFLSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIAQKKAQSSTD.
[0551] It should be understood that additional adenosine deaminases used in this application will be apparent to those skilled in the art and are within the scope of this invention. For example, the adenosine deaminase may be a homolog of an adenosine deaminase acting on tRNA (ADAT). Without limitation, an exemplary ADAT homolog's amino acid sequence includes the following:
[0552] Staphylococcus aureus TadA:
[0553] MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCS GSLMNLLQQS NFNHRAIVDKG VLKE ACS TLLTTFFKNLRANKKS TN
[0554] Bacillus TadA:
[0555] MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEML VIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPKGGC S GTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLSE
[0556] Salmonella typhimurium (S. typhimurium) TadA:
[0557] MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFRMRRQEIKALKKADRAEGAGPAV
[0558] Shewanella putrefaciens TadA:
[0559] MDE YWMQVAMQM AEKAEAAGE VPVGA VLVKDGQQIATGYNLS IS QHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE
[0560] Haemophilus influenzae F3031 TadA:
[0561] MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTΑΗAEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMCAGAILHS RIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRREEKKIEKALLKSLSDK
[0562] Caulobacter crescentus TadA:
[0563] MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIA tag NGPIAAHDPTAHAEIAAMRAAAAKLGNYRLTDLTLVVTLEPCAMCAGAISHARIGRVVFGADD PKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAKI
[0564] Geobacter sulfurreducens TadA:
[0565] MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFFRDLRRRKKAKATPALFIDERKVPPEP.
[0566] Escherichia coli TadA (ecTadA):
[0567] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD.
[0568] In some embodiments, the adenosine deaminase comprises a D108X mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a D108G, D108N, D108V, D108A, or D108Y mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. However, it should be understood that additional deaminases can be similarly compared to identify homologous amino acid residues that can be mutated as provided herein.
[0569] In some embodiments, the adenosine deaminase contains an A106X mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an A106V mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0570] In some embodiments, the adenosine deaminase contains an E155X mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein the presence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an E155D, E155G, or E155V mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0571] In some embodiments, the adenosine deaminase contains a D147X mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein the presence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains a D147Y mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0572] It should be understood that any mutations provided herein (e.g., amino acid sequences based on the TadA reference sequence) can be introduced into other adenosine deaminases, such as Staphylococcus aureus TadA (saTadA), or other adenosine deaminases (e.g., bacterial adenosine deaminases). Those skilled in the art will readily recognize sequences homologous to the mutated residues in the TadA reference sequence. Therefore, any mutations identified relative to the TadA reference sequence can be carried out in other adenosine deaminases having homologous amino acid residues. It should also be understood that any mutations provided herein can be carried out individually or in any combination in TadA or another adenosine deaminase. For example, an adenosine deaminase may contain D108N, A106V, E155V, and / or D147Y mutations relative to the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises the following mutation groups relative to the TadA reference sequence (mutation groups are separated by ";"), or the corresponding mutations in another adenosine deaminase: D108N and A106V; D108N and E155V; D108N and D147Y; A106V and E155V; A106V and D147Y; E155V and D147Y; D108N, A106V and E55V; D108N, A106V and D147Y; D108N, E55V and D147Y; A106V, E55V and D147Y; and D108N, A106V, E55V and D147Y. However, it should be understood that any combination of the corresponding mutations provided herein can be carried out in adenosine deaminases (e.g., ecTadA).
[0573] In some embodiments, the adenosine deaminase comprises one or more of the following mutations relative to the TadA reference sequence: H8X, T17X, L18X, W23X, L34X, W45X, R51X, A56X, E59X, E85X, M94X, I95X, V102X, F104X, A106X, R107X, D108X, K110X, M118X, N127X, A138X, F149X, M151X, R153X, Q154X, I156X, and / or K157X, or one or more corresponding mutations in another adenosine deaminase, wherein the presence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the following mutations relative to the TadA reference sequence: H8Y, T17S, L18E, W23L, L34S, W45L, R51H, A56E or A56S, E59G, E85K or E85G, M94L, 1951, V102A, F104L, A106V, R107C or R107H or R107P, D108G or D108N or D108V or D108A or D108Y, K110I, Ml18K, N127S, A138V, F149Y, M151V, R153C, Q154L, I156D and / or K157R, or one or more corresponding mutations in another adenosine deaminase.
[0574] In some embodiments, the adenosine deaminase comprises one or more of the following mutations relative to the TadA reference sequence: H8X, T17X, L18X, W23X, L34X, W45X, R51X, A56X, E59X, E85X, M94X, I95X, V102X, F104X, A106X, R107X, D108X, K110X, M118X, N127X, A138X, F149X, M151X, R153X, Q154X, I156X, and / or K157X, or one or more corresponding mutations in another adenosine deaminase, wherein the presence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the following mutations relative to the TadA reference sequence: H8Y, T17S, L18E, W23L, L34S, W45L, R51H, A56E or A56S, E59G, E85K or E85G, M94L, 1951, V102A, F104L, A106V, R107C or R107H or R107P, D108G or D108N or D108V or D108A or D108Y, K110I, Ml18K, N127S, A138V, F149Y, M151V, R153C, Q154L, I156D and / or K157R, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the H8X, D108X, and / or N127X mutations relative to the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein X indicates the presence of any amino acid. In some embodiments, the adenosine deaminase comprises one or more of the H8Y, D108N, and / or N127S mutations relative to the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.
[0575] In some embodiments, the adenosine deaminase comprises one or more of the H8X, D108X, and / or N127X mutations relative to the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein X indicates the presence of any amino acid. In some embodiments, the adenosine deaminase comprises one or more of the H8Y, D108N, and / or N127S mutations relative to the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.
[0576] In some embodiments, the adenosine deaminase comprises one or more of the H8X, D108X, and / or N127X mutations relative to the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein X indicates the presence of any amino acid. In some embodiments, the adenosine deaminase comprises one or more of the H8Y, D108N, and / or N127S mutations relative to the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.
[0577] In some embodiments, the adenosine deaminase comprises one or more of the H8X, D108X, and / or N127X mutations relative to the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein X indicates the presence of any amino acid. In some embodiments, the adenosine deaminase comprises one or more of the H8Y, D108N, and / or N127S mutations relative to the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.
[0578] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations relative to the TadA reference sequence, selected from the group consisting of H8X, D108X, N127X, D147X, R152X, and Q154X, or one or more corresponding mutations in another adenosine deaminase, wherein X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations relative to the TadA reference sequence, selected from the group consisting of H8X, M61X, M70X, D108X, N127X, Q154X, E155X, and Q163X, or one or more corresponding mutations in another adenosine deaminase, wherein X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations relative to the TadA reference sequence, selected from the group consisting of H8X, D108X, N127X, E155X, and T166X, or one or more corresponding mutations in another adenosine deaminase, wherein X indicates the presence of any amino acid other than the corresponding amino acid in the reference or wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations relative to the Tad reference sequence, selected from the group consisting of H8X, A106X, and D108X mutations, or mutations in another adenosine deaminase, wherein X indicates the presence of any amino acid other than the corresponding amino acid in the reference or wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations relative to the TadA reference sequence, selected from the group consisting of H8X, R126X, L68X, D108X, N127X, D147X, and E155X, or one or more corresponding mutations in another adenosine deaminase, wherein X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations relative to the TadA reference sequence, selected from the group consisting of H8X, D108X, A109X, N127X, and E155X, or one or more corresponding mutations in another adenosine deaminase, wherein X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase.
[0579] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations relative to the TadA reference sequence, selected from the group consisting of H8Y, D108N, N127S, D147Y, R152C, and Q154H, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations relative to the TadA reference sequence, selected from the group consisting of H8Y, M61I, M70V, D108N, N127S, Q154R, E155G, and Q163H, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase contains one, two, three, four, or five mutations relative to the TadA reference sequence, selected from the group consisting of H8Y, D108N, N127S, E155V, and T166P, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase contains one, two, three, four, five, or six mutations relative to the TadA reference sequence, selected from the group consisting of H8Y, A106T, D108N, N127S, E155D, and K161Q, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations relative to the TadA reference sequence, selected from the group consisting of H8Y, R126W, L68Q, D108N, N127S, D147Y, and E155V, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations relative to the TadA reference sequence, selected from the group consisting of H8Y, D108N, A109T, N127S, and E155G, or one or more corresponding mutations in another adenosine deaminase.
[0580] In some embodiments, the adenosine deaminase contains one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase contains a D108N, D108G, or D108V mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase contains an A106V and D108N mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase contains an R107C and D108N mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase contains H8Y, D108N, N127S, D147Y, and Q154H mutations relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase contains mutations of H8Y, R24W, D108N, N127S, D147Y, and E155V relative to the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase contains mutations of D108N, D147Y, and E155V relative to the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase contains mutations of H8Y, D108N, and N127S relative to the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase contains mutations of A106V, D108N, D147Y, and E155V relative to the TadA reference sequence, or corresponding mutations in another adenosine deaminase.
[0581] In some embodiments, the adenosine deaminase comprises one or more of the following mutations relative to the TadA reference sequence: a, S2X, H8X, I49X, L84X, H123X, N127X, I156X, and / or K160X, or one or more corresponding mutations in another adenosine deaminase, wherein the presence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the following mutations relative to the TadA reference sequence: S2A, H8Y, I49F, L84F, H123Y, N127S, I156F, and / or K160S, or one or more corresponding mutations in another adenosine deaminase.
[0582] In some embodiments, the adenosine deaminase comprises an L84X mutant adenosine deaminase, where X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an L84F mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0583] In some embodiments, the adenosine deaminase contains an H123X mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an H123Y mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0584] In some embodiments, the adenosine deaminase contains an I157X mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an I157F mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0585] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, or seven mutations relative to the TadA reference sequence, selected from the group consisting of L84X, A106X, D108X, H123X, D147X, E155X, and I156X, or one or more corresponding mutations in another adenosine deaminase, wherein X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations relative to the TadA reference sequence, selected from the group consisting of S2X, I49X, A106X, D108X, D147X, and E155X, or one or more corresponding mutations in another adenosine deaminase, wherein X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains one, two, three, four, or five mutations relative to the TadA reference sequence, selected from the group consisting of H8X, A106X, D108X, N127X, and K160X, or one or more corresponding mutations in another adenosine deaminase, wherein X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase.
[0586] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, or seven mutations relative to the TadA reference sequence, selected from the group consisting of L84F, A106V, D108N, H123Y, D147Y, E155V, and I156F, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations relative to the TadA reference sequence, selected from the group consisting of S2A, I49F, A106V, D108N, D147Y, and E155V.
[0587] In some embodiments, the adenosine deaminase contains one, two, three, four, or five mutations relative to the TadA reference sequence, selected from the group consisting of H8Y, A106T, D108N, N127S, and K160S, or one or more corresponding mutations in another adenosine deaminase.
[0588] In some embodiments, the adenosine deaminase comprises one or more of the E25X, R26X, R107X, A142X and / or A143X mutations relative to the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein the presence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the following mutations relative to the TadA reference sequence: E25M, E25D, E25A, E25R, E25V, E25S, E25Y, R26G, R26N, R26Q, R26C, R26L, R26K, R107P, R07K, R107A, R107N, R107W, R107H, R107S, A142N, A142D, A142G, A143D, A143G, A143E, A143L, A143W, A143M, A143S, A143Q, and / or A143R, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase contains one or more mutations described herein corresponding to the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.
[0589] In some embodiments, the adenosine deaminase contains an E25X mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase contains an E25M, E25D, E25A, E25R, E25V, E25S, or E25Y mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0590] In some embodiments, the adenosine deaminase comprises an R26X mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R26G, R26N, R26Q, R26C, R26L, or R26K mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0591] In some embodiments, the adenosine deaminase comprises an R107X mutation relative to the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, wherein X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R107P, R07K, R107A, R10...
Citation Information
Patent Citations
Towel rings
CN3315821D
cell phone
CN3329834D
Adeno-associated virus as eukaryotic expression vector
US4797368A
Dehydrated liposomes
US4880635A
Antineoplastic agent-entrapping liposomes
US4906477A