Methods of substituting pathogenic amino acids using programmable base editor systems
Patent Information
- Application Number
- JP2024183249
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-12-17
- Filing Date
- 2024-10-18
- Publication Date
- 2025-10-14
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] Related Applications This application is a continuation of U.S. Provisional Application No. 62 / 670,521, filed May 11, 2018, and U.S. Provisional Application No. 62 / 670,521, filed May 11, 2018. The benefit of U.S. Provisional Application No. 62 / 670,539, filed December 17, 2018, and U.S. Provisional Application No. 62 / 780,890, filed December 17, 2018. and claims the entire contents of each of which are incorporated herein by reference. Incorporation by Reference All publications, patents, and patent applications mentioned herein are the property of their respective owners. All publications, patents, or patent applications are specifically and individually indicated to be incorporated by reference. Unless otherwise indicated, All publications, patents, and patent applications mentioned herein are hereby incorporated by reference in their entirety. be incorporated into the book.
[0002] For most known genetic disorders, there are several methods to study or address the underlying cause of the disease. , which requires correction of a point mutation at the target locus rather than stochastic disruption of the gene. [Background technology]
[0003] Clustered regularly interspaced short palindromic repeat (CRISPR) system Current genome editing techniques induce double-stranded DNA breaks at target loci as the first step in gene correction. In response to double-stranded DNA breaks, the cell's DNA repair process mostly involves repairing the DNA break. Non-homologous end joining at the break site results in random insertions or deletions (indels). Most genetic disorders result from point mutations, but current approaches to point mutation correction These methods are inefficient and typically result from a cellular response to dsDNA breaks at the target locus. Induce more random insertions and deletions (indels) in the genome, thus making it more efficient and , and unwanted products such as stochastic insertions or deletions (indels) or translocations are much more likely to occur. Fewer and improved forms of genome editing are needed. Summary of the Invention
[0004] Provided herein are methods for treating a genetic disorder in a subject, the methods comprising administering a base administering to a subject an editor, or a polynucleotide encoding a base editor. (Here, the base editor is a programmable nucleotide linkage of polynucleotides. a guide polynucleotide containing a target polypeptide and a deaminase domain; wherein the guide polynucleotide directs the base editor to target nucleotides. Targeting to a base sequence; targeting to a base sequence Upon targeting, the editor deaminates the target nucleobase. By editing the nucleic acid bases of the nucleotide sequence, thereby changing the nucleic acid bases to different nucleic acid bases. and treating a genetic disorder by the addition of a pathogenic component to a protein, wherein the genetic disorder is a pathogenic amino acid, and another nucleic acid base binds the pathogenic amino acid to the wild-type The amino acid is replaced with a different benign amino acid.
[0005] Provided herein are methods for producing cells, tissues, or organs for treating genetic disorders in a subject. and a method of producing the cell, tissue, or organ, comprising: contacting a polynucleotide encoding the base editor (wherein the base editor The nucleotide binding domain and the determinant domain are capable of programming the polynucleotide. a guide polynucleotide; contacting (wherein the guide polynucleotide is a target nucleic acid of the cell, tissue, or organ) targeting a base editor to a nucleotide sequence; By deaminating nucleobases upon targeting base editors to sequences. Editing the nucleic acid base of the target nucleotide sequence, thereby changing the nucleic acid base to another nucleic acid base and producing cells, tissues, or organs for treating genetic disorders by inducing wherein the genetic disorder is caused by a pathogenic amino acid in a protein and is a nucleic acid salt of another nucleic acid. a group that converts the pathogenic amino acid into a benign amino acid that is different from the wild-type amino acid of the protein; In some embodiments, the method comprises administering to a subject a cell, tissue, or organ In some embodiments, the cell, tissue, or organ is autologous to the subject. In some embodiments, the cell, tissue, or organ is autologous to the subject. In some embodiments, the cells, tissues, or organs are allogeneic. It is xenogeneic to
[0006] In some embodiments, the nucleobase is located in a gene that is responsible for a genetic disorder. In some embodiments, the editing comprises editing a plurality of nucleic acid bases located within the gene. In some embodiments, the plurality of nucleic acid bases is not the cause of a genetic disorder. and editing one or more additional nucleic acid bases located in at least one other gene. In some embodiments, the gene and at least one other gene are It encodes one or more subunits of the protein.
[0007] In an embodiment, the nucleobase to be edited is in a gene listed in Table 3A or 3B, The editing may involve the deletion of an amino acid sequence in a protein encoded by a gene shown in Table 3A or 3B. In some embodiments, the genetic disorder is ACADM deficiency, sickle cell disease (SCD), ), hemoglobinopathies, beta-thalassemia, Pendred syndrome, autosomal dominant Parkinson's disease , or alpha-1 antitrypsin deficiency (A1AD).
[0008] In one aspect, the present invention provides a method for the generation of pathogenic amino acids using a programmable nucleobase editor. The present invention features compositions and methods for replacing the sickle portion of the β-globin protein. The codon for the sixth amino acid in the erythrocytosis variant (Sickle HbS; E6V) contains thymidine (T). Base editing of cytidine (C) nucleobases, thereby substituting valine for alanine (E6A) The present invention provides compositions and methods for converting valine to alanine at position 6 of sickle cell hemoglobin S. When substituted, it results in a beta-globin protein variant that lacks the sickle cell phenotype ( For example, it has the properties of normal β-globin protein (HbA; E6) and is a pathogenic variant of HbS. (These compounds are not likely to polymerize as they would be polymerized in the presence of sickle cell disease.) Therefore, the compositions and methods of the present invention are In one embodiment, the edited nucleobases are useful for treating beta (β)-globin. The base editing occurs in the HBB gene, which encodes the β-globulin At amino acid 6 of the HBV (Hepatitis B Blood Group) protein, there is a valine (Val) to alanine (Ala) substitution. In some embodiments, the genetic disorder is sickle cell anemia. In some embodiments, base editing is performed on the β subunit of hemoglobin. This results in an E6V>E6A amino acid change in the unit.
[0009] In another aspect, the present invention provides an HBB polypeptide containing a single nucleotide polymorphism (SNP) associated with sickle cell disease. The present invention provides a method for editing a polynucleotide, the method comprising: contacting a base editor complexed with the guide polynucleotide of The editor contains a polynucleotide programmable DNA binding domain and adenosine and one or more guide polynucleotides comprise a gene encoding a gene encoding a nucleotide deaminase domain related to sickle cell disease. Targeting base editors to modify associated SNPs from A·T to G·C do.
[0010] In another aspect, the invention provides base editors, polynucleotides encoding base editors (The base editor is a programmable DNA binding domain and adenosine deaminase domain.) (including the main) and to result in the A·T to G·C alteration of the SNP associated with sickle cell disease. and one or more guide polynucleotides that target the base editor to the cell or The present invention provides cells produced by introducing the nucleoside analog into a cell line or its precursor.
[0011] In another aspect, the present invention provides a cell according to any of the aspects described herein. and (c) administering to a subject in need thereof a compound selected from the group consisting of benzodiazepines, ... do.
[0012] In another aspect, the present invention provides a method for growing or culturing a cell according to any aspect described herein. The present invention provides isolated or expanded cells or cell populations.
[0013] In another aspect, the present invention provides a method of treating sickle cell disease in a subject, The method comprises combining a base editor, or a polynucleotide encoding a base editor, with a base The base editor generates programmable DNA-binding and adenosine deaminase domains. to produce the A·T to G·C alteration of the SNP associated with sickle cell disease. and one or more guide polynucleotides that target the editor. This includes administering it to elephants.
[0014] In another embodiment, the present invention provides a method for detecting red blood cells (erythrocytes) or The present invention provides a method for producing a progenitor cell thereof, the method comprising: (a) introducing a base editor, or A polynucleotide encoding the base editor is a polynucleotide programming adenosine deaminase domain), A base editor was used to modify the sickle cell disease-associated SNP from A·T to G·C. and one or more guide polynucleotides targeting SNPs associated with sickle cell disease. (b) inducing the differentiation of the erythroid progenitor cells into erythrocytes. Includes toto.
[0015] In another aspect, the present invention provides: (i) Streptococcus thermophilus 1 Cas 9 (St1Cas9) a polynucleotide comprising a programmable DNA-binding domain, and (ii) an adenosine The present invention provides a base editor comprising a deaminase domain.
[0016] In another aspect, the present invention provides a method for producing a medicament for the treatment of a medicament comprising administering to a patient a therapeutically effective amount ... and GACUUCUCCACAGGAGUCAGAU. .
[0017] In another aspect, the present invention provides: (i) a modified Staphylococcus aureus Cas 9 (SaCas9 ) a polynucleotide comprising a programmable DNA-binding domain, and (ii) an adenovirus.
[0010] Provided are base editors comprising a syndeaminase domain.
[0018] In another aspect, the present invention provides a method for producing a compound of formula (I) from UCCACAGGAGUCAGAUGCAC and UCCACAGGAGUCAGAUGCAC. The present invention provides a guide RNA (gRNA) comprising a nucleic acid sequence selected from the group consisting of:
[0019] In another aspect, the present invention provides a method for the preparation of a nucleotide sequence of UUCUCCACAGGAGUCAGA; CUUCUCCACAGGAGUCAGA; ACUUCUCCA Selected from CAGGAGUCAGA; GACUUCUCCACAGGAGUCAGA; and AGACUUCUCCACAGGAGUCAGA A guide RNA (gRNA) comprising a nucleic acid sequence is provided.
[0020] According to one embodiment, the base editing results in the α-1 antithrombin encoded by the SERPINA 1 gene. The resulting amino acid change is E342K>E342G in the psin protein. The genetic disorder is medium-chain acyl-CoA dehydrogenase (ACADM) deficiency. As a result of base editing, the protein encoded by the medium-chain acyl-CoA dehydrogenase (ACADM) gene A K329E>K329G amino acid change occurs in the protein. According to one embodiment, the genetic disorder In one embodiment, the base editing is performed on a gene encoded by the HBB gene. This results in an E26K>E26G amino acid change in the beta subunit of hemoglobin. In embodiments, the genetic disorder is Pendred Syndrome. The gene editing is SLC26A4 (encoded by the PDS gene, Solute Carrier Family 26 Member r 4 (PDS) protein). In one embodiment, the genetic disorder is autosomal dominant Parkinson's disease. A30P>A30 in the gene-encoded alpha-synuclein (SNCA) protein resulting in L amino acid changes.
[0021] In various embodiments of any aspect described herein, the SN associated with sickle cell disease The A·T to G·C modification in P changes valine to alanine in the HBB polypeptide. In various embodiments, the SNP associated with sickle cell disease is a valine at amino acid position 6. In various embodiments, the present invention provides a method for the treatment of sickle cell disease, comprising the steps of: The associated SNP substitutes glutamic acid for valine.
[0022] In various embodiments of all aspects described herein, the contacting is with a cell, eukaryotic cell, In various embodiments, the subject The cell is mammalian or human. In various embodiments, the cell is in vivo or ex vivo. In various embodiments, the cells or their precursors are derived from embryonic stem cells, induced pluripotent stem cells, or hematopoietic stem cells, common myeloid progenitor cells, proerythroblasts, erythroblasts, reticulocytes, or erythrocytes In various embodiments, the hematopoietic stem cells express CD34 + In various embodiments, the cells The cells are from a subject with sickle cell disease. In various embodiments, the cells are In various embodiments, the cells are autologous to the subject. Any of the methods described herein may be allogeneic or xenogeneic. In various embodiments of this aspect, the method further comprises: and one or more guide polynucleotides into a cell of interest. This includes:
[0023] In various embodiments of any aspect described herein, the polynucleotide protease The grammable DNA-binding domain is an engineered Staphylococcus aureus Cas 9 (SaCas9) , Streptococcus thermophilus 1 Cas 9 (St1Cas9), modified Streptococcus pyogene In various embodiments, the polynucleotide is a polynucleotide containing a polypeptide, such as poly(Asp) Cas9 (SpCas9), or a variant thereof. The nucleotide-programmable DNA-binding domain is flanked by engineered protospacers In various embodiments, the modified SaCas9 has specificity for the PAM motif. The modified PAM comprises the nucleic acid sequence 5'-NNNRRT-3'. In various embodiments, the modified S aCas9 catalyzes the amino acid substitutions E782K, N968K, and R1015H or their corresponding amino acid substitutions. Includes exchange.
[0024] In various embodiments, the polynucleotide programmable DNA binding domain , a variant of SpCas9 with altered protospacer adjacent motif (PAM) specificity In various embodiments, the modified PAM comprises the nucleic acid sequence 5'-NGC-3'.
[0025] In various embodiments, the modified SpCas9 comprises the amino acid substitutions D1135M, S1136Q, G1218K, , E1219F, A1322R, D1332A, R1335E, and T1337R or their corresponding amino acid substitutions In various embodiments, the polynucleotide-programmable DNA-binding domain The enzyme is a nuclease-inactivated or nickase variant. The nickase variant has the amino acid substitution D10A or a corresponding amino acid substitution. include.
[0026] In various embodiments of any aspect described herein, the base editor is a zinc fluor In various embodiments, the zinc finger domain further comprises: Recognition helix sequences RNEHLEV, QSTTLKR, and RTEHLAR, or recognition helix sequence RGEHL Various zinc finger domains include zf1ra and zf1ra-RQ, QSGTLKR, and RNDKLVP. is one or more of zf1rb.
[0027] In various embodiments of any aspect described herein, adenosine deaminase The Zed domain can deaminate adenine in deoxyribonucleic acid (DNA). In various embodiments, the adenosine deaminase is a modified adenosine deaminase that does not occur in nature. In various embodiments, the adenosine deaminase is TadA deaminase. In various embodiments, the TadA deaminase is TadA*7.10.
[0028] In various embodiments of any aspect described herein, one or more guide RNAs , CRISPR RNA (crRNA) and transcoding small RNA (tracrRNA), where cr The RNA contains a nucleic acid sequence complementary to an HBB nucleic acid sequence containing a SNP associated with sickle cell disease. In certain embodiments, the base editor is a nucleic acid sequence encoding a HBB nucleic acid sequence containing a SNP associated with sickle cell disease. It is complexed with a single guide RNA (sgRNA) that contains a nucleic acid sequence complementary to
[0029] In various embodiments of any aspect described herein, St1Cas9 is The amino acid sequence includes: SDLVLGLAIGIGSVGVGILNKVTGEIIHKNSRIFPAAQAENNLVRRTNRQGRRLARRKKHRRVRLNRLFEESGLITDFTK ISINLNPYQLRVKGLTDELSNEELFIALKNMVKHRGISYLDDASDDGNSSVGDYAQIVKENSKQLETKTPGQIQLERYQT YGQLRGDFTVEKDGKKHRLINVFPTSAYRSEALRILQTQQEFNPQITDEFINRYLEILTGKRKYYHGPGNEKSRTDYGRY RTSGETLDNIFGILIGKCTFYPDEFRAAKASYTAQEFNLLNDLNNLTVPTETKLSKEQKNQIINYVKNEKAMGPAKLFK YIAKLLSCDVADIKGYRIDKSGKAEIHTFEAYRKMKTLETLDIEQMDRETLDKLAYVLTLNTEREGIQEALEHEFADGSF SQKQVDELVQFRKANSSIFGKGWHNFSVKLMMELIPELYETSEEQMTILTRLGKQKTTSSSNKTKYIDEKLLTEEIYNPV VAKSVRQAIKIVNAAIKEYGDFDNIVIEMARETNEDDEKKAIQKIQKANKDEKDAAMLKAANQYNGKAELPHSVFHGHKQ LATKIRLWHQQGERCLYTGKTISIHDLINNSNQFEVDHILPLSITFDDSLANKVLVYATANQEKGQRTPYQALDSMDDAW SFRELKAFVRESKTLSNKKKEYLLTEEDISKFDVRKKFIERNLVDTLYASRVVLNALQEHFRAHKIDTKVSVVRGQFTSQ LRRHWGIEKTRDTYHHHAVDALIIAASSQLNLWKKQKNTLVSYSEDQLLDIETGELISDDEYKESVFKAPYQHFVDTLKS KEFEDSILFSYQVDSKFNRKISDATIYATRQAKVGKDKADETYVLGKIKDIYTQDGYDAFMKIYKKDKSKFLMYRHDPQT FEKVIEPILENYPNKQINDKGKEVPCNPFLKYKEEHGYIRKYSKKGNGPEIKSLKYDSKLGNHIDITPKDSNNKVVLQS VSPWRADVYFNKTTGKYEILGLKYADLQFDKGTGTYKISQEKYNDIKKKEGVDSDSEFKFTLYKNDLLLVKDTETKEQQL FRFLSRTMPKQKHYVELKPYDKQKFEGGEALIKVLGNVANSGQCKKGLGKSNISIYKVRTDVLGNQHIIKNEGDKPKLDF .
[0030] In various embodiments, the base editor is a polynucleotide programmable It contains a linker between the DNA binding domain and the adenosine deaminase domain. In an embodiment, the linker comprises the amino acid sequence: SGGSSGGSSGSETPGTSESATPES. In embodiments, the base editor comprises one or more nuclear localization signals. In the above, the base editor comprises the following amino acid sequence: MPKKKRKVSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNY RLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRR QEIKAQKKAQSSTDSGGSGGGSSGSETPGTSESATPESSGGSSGGSSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLV LNNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAG SLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTDSGGSGGGSSGSETPGTSESATPESDLVL GLAIGIGSVGVGILNKVTGEIIHKNSRIFPAAQAENNLVRRTNRQGRRLARRKKHRRVRLNRLFEESGLITDFTKISINL NPYQLRVKGLTDELSNEELFIALKNMVKHRGISYLDDASDDGNSSVGDYAQIVKENSKQLETKTPGQIQLERYQTYGQLR GDFTVEKDGKKHRLINVFPTSAYRSEALRILQTQQEFNPQITDEFINRYLEILTGKRKYYHGPGNEKSRTDYGRYRTSGE TLDNIFGILIGKCTFYPDEFRAAKASYTAQEFNLLNDLNNLTVPTETKKLSKEQKNQIINYVKNEKAMGPAKLFKYIAKL LSCDVADIKGYRIDKSGKAEIHTFEAYRKMKTLETLDIEQMDRETLDKLAYVLTLNTEREGIQEALEHEFADGSFSQKQV DELVQFRKANSSIFGKGWHNFSVKLMMELIPELYETSEEQMTILTRLGKQKTTSSSNKTKYIDEKLLTEEIYNPVVAKSV RQAIKIVNAAIKEYGDFDNIVIEMARETNEDDEKKAIQKIQKANKDEKDAAMLKAANQYNGKAELPHSVFHGHKQLATKI RLWHQQGERCLYTGKTISIHDLINNSNQFEVDHILPLSITFDDSLANKVLVYATANQEKGQRTPYQALDSMDDAWSFREL KAFVRESKTLSNKKKEYLLTEEDISKFDVRKKFIERNLVDTLYASRVVLNALQEHFRAHKIDTKVSVVRGQFTSQLRRHW GIEKTRDTYHHHAVDALIIAASSQLNLWKKQKNTLVSYSEDQLLDIETGELISDDEYKESVFKAPYQHFVDTLKSKEFED SILFSYQVDSKFNRKISDATIYATRQAKVGKDKADETYVLGKIKDIYTQDGYDAFMKIYKKDKSKFLMYRHDPQTFEKVI EPILENYPNKQINDKGKEVPCNPFLKYKEEHGYIRKYSKKGNGPEIKSLKYYDSKLGNHIDITPKDSNNKVVLQSVSPWR ADVYFNKTTGKYEILGLKYADLQFDKGTGTYKISQEKYNDIKKKEGVDSDSEFKFTLYKNDLLLVKDTETKEQQLFRFLS RTMPKQKHYVELKPYDKQKFEGGEALIKVLGNVANSGQCKKGLGKSNISIYKVRTDVLGNQHIIKNEGDKPKLDFPKKKR KVEGADKRTADGSEFESPKKKRKV.
[0031] In various embodiments of any aspect described herein, the guide RNA is Further comprising the nucleic acid sequence: GUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUA AGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG.
[0032] In various embodiments, the guide RNA is CUUCUCCACAGGAGUCAGAUGUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUC UUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG; ACUUCUCCACAGGAGUCAGAUGUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAU CUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG; or GACUUCUCCACAGGAGUCAGAUGUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAA UCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG The nucleic acid sequence comprises a nucleic acid sequence selected from:
[0033] In various embodiments of any of the embodiments described herein, the protein nucleic acid complex comprises: A base editor according to any embodiment described herein and a method for producing a base according to any embodiment described herein The present invention includes a guide RNA according to any embodiment.
[0034] In various embodiments of any aspect described herein, the modified SaCas9 is The amino acid substitutions include E782K, N968K, and R1015H or their corresponding amino acid substitutions. In various embodiments, SaCas9 comprises the following amino acid sequence: KRNYILGLAIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTHS ELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKD GEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPGFGWKDIKEWYEMLMGHCTYFP EELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKP EFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAIN LILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELARE KNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPR SVSFDNSFNNKVLVKQEENSKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDF INRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKL DKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRKLINDTLYSTRKDDKGNTLI VNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIK YYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAE FIASFYKNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPHIIKTIASKTQSIKKYSTDILGNLYEV KSKKHPQIIKKG.
[0035] In various embodiments of any aspect described herein, the base editor comprises: Contains the amino acid sequence: MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQK KAQSSTDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIG EGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLH YPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTDSGGSSGGSSGSETPGTSESATPESSGGSSSGGSKRN YILGLAIGITSVGYGIIDYETRDVIDAGVRLFKEANVENGEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHSELS GINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKDGEV RGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKWEYEMLMGHCTYFPEEL RSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFT NLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAINLIL DELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELAREKNS KDAQKMINEMQKRNRQTNERIEEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVS FDNSFNNKVLVKQEENSKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDFINR NLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKLDKA KKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRKLINDTLYSTRKDDKGNTLIVNN LNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIKYYG NKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAEFIA SFYKNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPHIIKTIASKTQSIKKYSTDILGNLYEVKSK KHPQIIKKGEGADKRTADGSEFESPKKKRKV.
[0036] In various embodiments, the base editor comprises the following amino acid sequence: MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHEIMALRQGGLVMQNYRLIDATL YVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQK KAQSSTDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIG EGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLH YPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTDSGGSSGGSSGSETPGTSESATPESSGGSSSGGSKRN YILGLAIGITSVGYGIIDYETRDVIDAGVRLFKEANVENGEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHSELS GINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKDGEV RGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKWEYEMLMGHCTYFPEEL RSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFT NLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAINLIL DELWHTNDNQIAIFNRKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIQUINAIIKKYGLPNDIIIELAREKNS KDAQKMINEMQKRNRQTNERIEEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVS FDNSFNNKVLVKQEENSKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDFINR NLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKLDKA KKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRKLINDTLYSTRKDDKGNTLIVNN LNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIKYYG NKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAEFIA SFYKNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPHIIKTIASKTQSIKKYSTDILGNLYEVKSK KHPQIIKKGEGADKRTADGSEFESPKKKRKVSSGNSNANSRGPSFSSGLVPLSLRGSHSRPGERPFQCRICMRNFSRNEH LEVHTRTHTGEKPFQCRICMRNFSQSTTLKRHLRTHTGEKPFQCRICMRNFSRTEHLARHLKTHLRGSSAQ; or
[0037] MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHEIMALRQGGLVMQNYRLIDATL YVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQK KAQSSTDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIG EGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLH YPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTDSGGSSGGSSGSETPGTSESATPESSGGSSSGGSKRN YILGLAIGITSVGYGIIDYETRDVIDAGVRLFKEANVENGEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHSELS GINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKDGEV RGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKWEYEMLMGHCTYFPEEL RSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFT NLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAINLIL DELWHTNDNQIAIFNRKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIQUINAIIKKYGLPNDIIIELAREKNS KDAQKMINEMQKRNRQTNERIEEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVS FDNSFNNKVLVKQEENSKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDFINR NLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKLDKA KKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRKLINDTLYSTRKDDKGNTLIVNN LNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIKYYG NKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAEFIA SFYKNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPHIIKTIASKTQSIKKYSTDILGNLYEVKSK KHPQIIKKGEGADKRTADGSEFESPKKKRKVSSGNSNANSRGPSFSSGLVPLSLRGSHSRPGERPFQCRICMRNFSRGEH LRQHTRTHTGEKPFQCRICMRNFSQSGTLKRHLRTHTGEKPFQCRICMRNFSRNDKLVPHLKTHLRGSSAQ.
[0038] In various embodiments of any aspect described herein, the guide RNA is Column GUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUG GCGAGAUUUUUU Further includes:
[0039] In various embodiments, the guide RNA comprises the nucleic acid sequence UCCACAGGAGUCAGAUGCACGUUUUAGUACUCU GUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGCGAGAUUUUUU, or the nucleic acid sequence CUCCACAGGAGUCAGAUGCACGUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAG Contains GCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGCGAGAUUUUUU.
[0040] In some embodiments, the methods provided herein further comprise the step of: In one embodiment, the additional nucleobase is not the cause of a genetic disorder. In another embodiment, the additional nucleobase is responsible for a genetic disorder.
[0041] In another aspect, a method of treating a genetic disorder in a subject is provided, the method comprising: The method comprises administering to a subject in need thereof a base editor, wherein the base editor comprises a Polynucleotides with programmable nucleotide binding a target nuclease domain and a deaminase domain; Binding a guide polynucleotide to a guide sequence; Deamination of the nucleic acid bases upon binding of the target polynucleotide By editing the nucleic acid bases of the peptide sequence, thereby changing the nucleic acid bases to different nucleic acid bases. and treating a genetic disorder by using the nucleic acid base, wherein the nucleic acid base is a regulatory element of a gene. or in regulatory regions.
[0042] In another embodiment, a method for treating a genetic disorder in a subject in need thereof. A method of producing a cell, tissue, or organ is provided, the method comprising: contacting the organ with a base editor, wherein the base editor is a polynucleotide sequence encoding a Programmable nucleotide-binding domain and deaminase domain as guideposts target nucleoside of a polynucleotide in a cell, tissue, or organ; Binding a guide polynucleotide to the target sequence; Deamination of the nucleobase upon binding to the target nucleotide sequence by editing the nucleic acid bases in the nucleic acid sequence, thereby changing the nucleic acid bases to different nucleic acid bases. and producing cells, tissues, or organs for treating genetic disorders, and the nucleobase is in a regulatory element of a gene. In some embodiments, the method further comprises administering the cells, tissues, or organs to a subject. In some embodiments, the cell, tissue, or organ is autologous to the subject. In some embodiments, the cell, tissue, or organ is allogeneic to the subject. is heterogeneous to the subject.
[0043] In some embodiments of the above-described methods, the gene is responsible for a genetic disorder. In some embodiments, the gene is not the cause of the genetic disorder. In some embodiments, the change is an increase in the amount of transcription of the gene. In some embodiments, the change is a decrease in the amount of transcription of the gene. The editing alters the binding pattern of at least one protein to a regulatory element. In some embodiments, the regulatory element is a promoter, enhancer, repressor, or -, silencer, insulator, start codon, stop codon, Kozak consensus sequence , splice acceptor, splice donor, splice site, 3' untranslated region (UTR) In some embodiments, the result of editing is a 5′ untranslated region (UTR), a 5′ untranslated region (UTR), or an intergenic region. In some embodiments, editing results in the addition of a splice site. In some embodiments, the editing results in intron inclusion. In some embodiments, editing results in exon skipping. Amplification results in the removal of a start codon, a stop codon, or a Kozak consensus sequence. In some embodiments, the edit is to modify a start codon, a stop codon, or a Kozak consensus codon. In some embodiments, the editing results in the addition of a sequence located in a regulatory element of the gene. This involves editing multiple nucleic acid bases that are located at the base of the nucleic acid.
[0044] In some embodiments of the methods described above, the editing comprises editing a plurality of nucleobases. wherein at least one nucleobase of the plurality of nucleobases is linked to at least one additional residue. In some embodiments, the gene is located in at least one additional regulatory element of the gene. and at least one additional gene encoding one or more subunits of at least one protein. Code the knit.
[0045] In some embodiments of the above-described methods, the editing is any of the changes shown in Table 4 herein. In some embodiments, the genetic disorder is selected from one of: In some embodiments, the genetic disorder is sickle cell disease (SCD), which is characterized by a decrease in fetal hemoglobin In some embodiments, the nucleobase is c. -114 to -102 of HBG1 / 2. In certain embodiments, the nucleobase is located in the promoter of HBG1 / 2.
[0046] In some embodiments of the methods described above, the methods further comprise at least one further a second compilation of nucleobases comprising at least one additional nucleobase selected from the group consisting of: In some embodiments, this additional nucleobase is not present within a regulatory element of the gene. It is located in the protein coding region.
[0047] In certain embodiments of the methods of the above aspects, the deaminase domain is an adenosine deaminase. In some embodiments, the deaminase domain is a cytidine deaminase domain. In some embodiments, the adenosine deaminase domain is a deoxyribonuclease domain. It can deaminate adenine in nucleic acid (DNA). Imidazole polynucleotides include ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). In some embodiments, the guide polynucleotide is a CRISPR RNA (crRNA) sequence, a transgene, or a nucleotide sequence. transactivating CRISPR RNA (tracrRNA) sequences, or a combination thereof.
[0048] In some embodiments, the methods provided herein comprise the step of: In some embodiments, the second guide polynucleotide further comprises a ribonucleic acid (RNP) fragment. A) or deoxyribonucleic acid (DNA). In some embodiments, the second guide The polynucleotides include CRISPR RNA (crRNA) sequences, trans-activating CRISPR RNA (tracrRNA) sequences, and In some embodiments, the second guide polypeptide comprises a sequence of The nucleotide is a nucleic acid sequence that targets the base editor to a second target nucleotide sequence. do.
[0049] In some embodiments, the polynucleotide-programmable DNA-binding domain The domains are the Cas9 domain, Cpf1 domain, CasX domain, CasY domain, and Cas12b / C2c1 domain. In some embodiments, the polynucleotide program The DNA-binding domains capable of binding are nuclease dead. In one embodiment, the polynucleotide programmable DNA binding domain comprises a In some embodiments, the polynucleotide is a programmable DNA sequence. In some embodiments, the A-binding domain comprises a Cas9 domain. Contains nuclease-inactive Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9 In some embodiments, the Cas9 domain comprises a Cas9 nickase. In one embodiment, the polynucleotide programmable DNA binding domain is an engineered , or a modified polynucleotide programmable DNA binding domain.
[0050] In some embodiments, any of the methods provided herein further comprises: In some embodiments, the second base editor further comprises a base editor other than the base editor. It contains a deaminase domain distinct from that of the tar.
[0051] In some embodiments, the base editing results in less than 20% indel formation. In some embodiments, the base editing results in less than 15% indel formation. In embodiments, the base editing results in less than 10% indel formation. In some embodiments, the base editing results in less than 5% indel formation. In some embodiments, the base editing results in less than 4% indel formation. The edits result in less than 3% indel formation. In some embodiments, the base edits In some embodiments, the base editing results in less than 2% indel formation. In some embodiments, the base editing results in less than 0.5% indel formation. In some embodiments, the base editing results in less than 0.1% indel formation. In some embodiments, the base editing does not result in a translocation.
[0052] The features of the present disclosure are set forth with particularity in the appended claims. A better understanding of this point can be obtained by reference to the following detailed description. The present disclosure provides an illustrative embodiment, in which the principles of the present disclosure are explained and utilized, as illustrated in the accompanying drawings. The surface is as follows: [Brief explanation of the drawings]
[0053] [Figure 1]Figure 1 shows a schematic diagram comparing a healthy individual with a patient with antitrypsin deficiency (A1AD). In healthy individuals, the α-1 antitrypsin (A1AT) protein protects the lungs from proteases, and the liver releases α-1 antitrypsin into the blood. In patients with α-1 antitrypsin deficiency (A1AD), the absence of normal α-1 antitrypsin leads to lung tissue damage. Accumulation of abnormal α-1 antitrypsin in liver hepatocytes leads to cirrhosis.
[0054] [Figure 2] Figure 2 shows the typical range of serum alpha-1 antitrypsin (A1AT) levels for different genotypes (normal (MM); heterozygous carriers of alpha-1 antitrypsin deficiency (MZ, SZ); homozygous deficiency (SS, ZZ)). Serum alpha-1 antitrypsin (AAT) concentrations are expressed in μM on the left "y" axis, as is common in the literature. The right "y" axis shows the approximate conversion of serum AAT concentrations to mg / dL, as commonly reported in clinical laboratories and by various measurement techniques (turbidimetry and radial immunodiffusion).
[0055] [Figure 3] Figure 3 shows the sequence of the target site for correction of E342K in the SERPINA 1 gene encoding A1AT. Highlighted are the non-canonical spCas9 NGC PAM and the target A nucleobase whose editing results in the desired correction of E342K. Also shown are additional off-target A editing sites that may result in benign alleles, such as E342G or D341G.
[0056] [Figure 4]Figure 4 is a bar graph showing the secreted protein levels in the culture supernatant of HEK293T cells transiently transfected with plasmids encoding different variants of the A1AT protein. A1AT concentrations were measured by ELISA using a previously published method (Borel et al., 2017, "Alpha-1 Antitrypsin Deficiency: Methods and Protocols," 10.1007 / 978-1-4939-7163-3). The most common clinical variants (e.g., pathogenic mutations) of A1AT are E264V (PiS allele) and E342K (PiZ allele). PiS and PiZ proteins are produced in lower amounts than the wild-type protein. Both D341G and E342G proteins are produced at levels similar to those of the wild-type protein. Therefore, using the adenine base editor and base editing method described herein, we created these benign alleles that can restore A1AT secretion from hepatocytes, simultaneously ameliorating hepatotoxicity and increasing A1AT circulation to the lung. In the figure, A1AT: α-1 antitrypsin; A1AD: α1 antitrypsin deficiency; "Z mutation" is E342K (PiZ allele) mutation; "S mutation" is E264V (PiS allele).
[0057] [Figure 5] Figure 5 shows a schematic diagram of the strategy for evolving DNA deoxyadenosine deaminase starting from TadA. The E. coli library harbored a plasmid library of mutant ecTadA (TadA*) genes fused to dCas9 and a selection plasmid requiring a targeted A·T to G·C mutation to repair the antibiotic resistance gene. Mutations from surviving TadA* variants were transferred into an ABE construct for base editing in human cells.
[0058] [Figure 6]Figure 6 presents a table showing the first eight amino acids of mature hemoglobin (Hb), including normal HbA, the pathogenic variants sickle HbS and HbC, and the HbG Makassar variant, which is phenotypically similar to HbA but does not polymerize like HbS. Figure 6 shows the amino acid encoded at amino acid position 6 in each of the Hb types, as well as the DNA and mRNA sequences encoding the first eight amino acids of these Hb proteins.
[0059] [Figure 7] Figures 7A and 7B show experimental results of editing the nucleobase adenosine (A) to guanosine (G) in the sequence complementary to the codon encoding valine (CAC) at amino acid position 6 of HbS using various A-to-G base editors (ABEs) that recognize different PAM sequences. Figure 7A is a table describing the characteristics of the tested HbS gRNAs and corresponding ABEs, including the locations of the desired edits and potential off-target edits. Figure 7B is a graph showing the results of using ABEs for base editing at the sickle cell target site.
[0060] [Figure 8]Figures 8A-8G show the results of experiments using a Staphylococcus aureus Cas9 variant (saKKH) with tolerance to NNNRRT, alone or fused to a DNA-binding domain with sequence specificity at the sickle cell target site, to edit adenosine (A) to guanosine (G) in the codon encoding valine (CAC) at amino acid position 6 of HbS. Figure 8A shows a schematic diagram of the ABE constructs, showing the organization of domains within the polypeptides, including saKKH ABE7.10, saKKH ABE7.10 zf1ra, and saKKH ABE7.10 zf1rb. Figure 8B shows the nucleic acid sequence at the sickle cell target site, as well as the target-complementary sequences (g1 and g4) of the guide RNAs, as indicated by the underline. Figure 8C is a graph showing the results of using saKKH ABE7.10, saKKH ABE7.10 zf1ra, and saKKH ABE7.10 zf1rb in combination with a guide RNA g1 having a 20-nucleotide (nt) nucleic acid sequence complementary to a sickle cell disease target site. On the right side of the graph in Figure 8C is the nucleic acid sequence at the sickle cell target site and the target-complementary sequence of the g1 guide RNA. Figure 8D is a graph showing the results of using saKKH ABE7.10, saKKH ABE7.10 zf1ra, and saKKH ABE7.10 zf1rb in combination with a guide RNA g1 having a 21-nt nucleic acid sequence complementary to a sickle cell disease target site. Figure 8E is a graph showing the results of using saKKH ABE7.10, saKKH ABE7.10 zf1ra, and saKKH ABE7.10 zf1rb in combination with guide RNA g4, which has a 20-nt nucleic acid sequence complementary to the sickle cell disease target site. The right side of the graph in Figure 8E shows the nucleic acid sequence at the sickle cell disease target site and the target-complementary sequence of the g4 guide RNA. Figure 8F is a graph showing the results of using saKKH ABE7.10, saKKH ABE7.10 zf1ra, and saKKH ABE7.10 zf1rb in combination with guide RNA g4, which has a 21-nt nucleic acid sequence complementary to the sickle cell disease target site. Figure 8G shows base editing at a control HEK2 site.
[0061] [Figure 9] Figures 9A-9E show the development and evaluation of an adenosine base editor (ABE) with the Streptococcus thermophilus Cas9 (St1Cas9) DNA-binding domain for base editing at sickle cell disease target sites. Figure 9A shows base editing using ABE St1Cas9 with the St1Cas9 standard PAM sequence, NNAGAA (TTCTAG; reverse complement). The lower inset shows the indel percentage (indel %) at the base-edited site, comparing ABE St1Cas9, St1 Cas9 nuclease, and untreated. Figure 9B shows base editing using ABE St1Cas9 with the St1Cas9 standard PAM sequence, NNAGAA. The lower inset shows the indel percentage at the base-edited site, comparing ABE St1Cas9, St1Cas9 nuclease, and untreated. Figure 9C shows base editing using ABE St1Cas9 with the St1Cas9 non-canonical PAM sequence NNACCA (TGGTNN; reverse complement). The lower inset shows the percentage of indels at the base-edited site compared with ABE St1Cas9, St1Cas9 nuclease, and untreated. Figure 9D shows base editing using ABE St1Cas9 with the St1Cas9 non-canonical PAM sequence NNACCA (TGGTNN; reverse complement). The lower inset shows the percentage of indels at the base-edited site compared with ABE St1Cas9, St1Cas9 nuclease, and untreated. Figure 9E shows base editing using ABE St1Cas9 with the St1Cas9 non-canonical PAM sequence NNACCA at the sickle cell target site. The arrow indicates the A·T to G·C mutation (Val→Ala) induced by the ABE-St1Cas9 base editor at the sickle cell disease target site in Hb.
[0062] [Figure 10]Figure 10 shows the percent base editing at sickle cell disease target sites using an ABE with an SpCas9 DNA-binding domain evolved and engineered to accept an NGC PAM (ngcABE). The leftmost bar in the bar graph represents "Pro6Pro," the middle bar represents "Val7Ala," and the rightmost bar represents "Ser10Pro."
[0063] [Figure 11] Figure 11 is a schematic diagram depicting the promoter region of the HBG1 / 2 gene. Individual purple triangles indicate SNPs and deletions naturally found in HPFH patients, and green arrows, e.g., "BCL11A," "CCAAT," "90 BCL11A," and "ZBTB7A," indicate potential transcription binding sites. The thick, pointed lines (pink) clustered above and below the HBG1 / 2 sequence indicate guide RNAs that can target these regions of interest, e.g., target sequences in the gene.
[0064] [Figure 12] Figure 12 shows the rate of targeted base editing of target sequences in the HBG1 / 2 gene in 293T cells transfected with the indicated gRNAs and Cas9 base editor. The base editing efficiency percentage was determined by Miseq. Shown in the figure is the percentage of editing that occurred in 293T cells using each type of gRNA, the genes and target sequences of which are shown in Table 4. "Cs" indicates the position relative to the gRNA at which editing will occur by the CBE in conjunction with the gRNA. "As" indicates the position relative to the gRNA at which the ABE in conjunction with the respective gRNA will edit the sequence.
[0065] [Figure 13]Figure 13 shows the percentage of editing in primary bone marrow CD34+ cells performed by each type of gRNA, with the gene and target sequences shown in Table 4. CD34+ cells were transfected with the indicated gRNA and base editor. "Cs" indicates the position relative to the gRNA where editing by a CBE, such as BE4, in conjunction with the gRNA will occur. "As" indicates the position relative to the gRNA where an ABE, together with the respective gRNA, will edit the target sequence. The percentage of base editing at both the HBG1 and HBG2 loci was assessed by Miseq. DETAILED DESCRIPTION OF THE INVENTION
[0066] As described herein, the present invention provides a method for the preparation of nucleic acid sequences using a programmable nucleobase editor. Compositions and methods for replacing pathogenic amino acids are featured. In certain aspects, The compositions and methods described include the use of β-globin protein encoded by the HBB gene. Treatment of sickle cell disease caused by a Glu→Val mutation at the sixth amino acid of Despite many advances in the field of gene editing, →Precise correction of the diseased HBB gene to restore Val remains elusive, and CRISPR / C This has not been achieved using either Cas nucleases or CRISPR / Cas base editing approaches. stomach.
[0067] of the HBB gene using a CRISPR / Cas nuclease approach to replace the affected nucleotide. Genome editing requires cutting genomic DNA, which can result in base insertions / deletions ( This increases the risk of generating premature stop codons, codon readout errors, and other indels. This may result in unintended and undesirable consequences, including frame changes. The generation of a double-strand break at the β-globin locus radically disrupts the locus via recombination events. The β-globin locus contains several globulins with sequence identity. A cluster of bin genes (- 5′-ε- ; Gγ- ; Aγ- ; δ- ; and β-globin-3′) This structure of the β-globin locus allows for recombinational repair of double-strand breaks within the locus. The intervening sequences between globin genes, for example, between the delta-globin gene and the beta-globin gene, are May result in gene loss. Unintentional alteration of the gene locus causes thalassemia. It also poses risks.
[0068] CRISPR / Cas base editing approaches have the ability to generate precise changes at the nucleobase level However, the Val→Glu (G T G → G A G) precise corrections currently exist. It requires a T·A to A·T conversion editor, which is not known to do this. The specificity of CRISPR / Cas base editing is due in part to the R-loop formation that occurs when CRISPR / Cas binds to DNA. This is due to the limited window of editable nucleotides that arises from the synthesis of Therefore, CRISPR / Cas targeting must be performed to enable base editing. It must occur at or near the site and be within the window for optimal editing. There may be different sequence requirements.
[0069] One requirement for CRISPR / Cas targeting is that the target site is flanked by promoters. The presence of a rotospacer adjacent motif (PAM) sequence. For example, many base edits involve PAMs. It is based on SpCas9, which requires the sequence NGG. It allows transversion from T·A to A·T. Even if we assume that this is possible, it is difficult to target the target “ There are no NGG PAMs that place "A" in place. Many new PAMs are expanding the collection of available PAMs. Although new CRISPR / Cas proteins have been discovered or generated, the need for PAM remains unclear. It is the limiting factor in the ability to direct the editor to a specific nucleotide at any position in the genome. Continue to do so.
[0070] The present invention relates to the aforementioned methods for providing genome editing approaches for the treatment of sickle cell anemia. The present invention is based, at least in part, on several discoveries described herein that address the above-mentioned problems. In one aspect, the present invention provides a method for the preparation of a compound comprising the amino acid sequence at amino acid position 6 of the Hb protein that causes sickle cell disease. Hb variants (Hb) that substitute valine for alanine and thereby do not result in the sickle cell phenotype b Makassar). T G → G A G) is T·A to A·T The results described herein demonstrate the conversion of A·T to αT, which would not be possible without a base editor. Using the ABE to G C base editor, Val → Ala (G T G → G C G) substitution (i.e., Hb This demonstrates the finding that it is possible to generate a phenotype of a phenotype that is similar to that of the Makassar variant, as provided herein. This has been achieved in part by the development of novel base editors and novel base editing strategies, as For example, for optimal base editing at the sickle cell disease target site, adjacent sequences (e.g., P A novel ABE base editor (i.e., ABE) utilizing AM sequences; zinc finger binding sequences A new antibody (containing a nosine deaminase domain) has been developed.
[0071] The sixth amino acid of the sickle cell variant of β-globin protein (Sickle HbS; E6V) The thymidine (T) in the donor is base edited to a cytidine (C), thereby Compositions and methods for substituting valine amino acid residues with alanine amino acid residues (V6A) Provided and described herein is a substitution of alanine for valine at position 6 of HbS, which results in the development of sickle cell carcinoma. No blood cell phenotype (e.g., no potential for polymerization, as in the case of the pathogenic variant HbS) Therefore, the compositions and methods of the present invention provide , useful in the treatment of sickle cell disease.
[0072] Provided and described herein are the genes provided in Tables 3A, 3B, or 4 herein. Treating diseases or disorders caused by or associated with genes and base editor systems described herein for Compositions and methods.
[0073] The following description and examples further illustrate embodiments of the present disclosure. It should be understood that the present invention is not limited to the particular embodiments described herein, as such may vary. Those skilled in the art will recognize that this disclosure has numerous variations and modifications that are encompassed within its scope. You will recognize it.
[0074] All terms are intended to be understood as would be understood by a person skilled in the art. Unless otherwise defined, all technical and scientific terms used herein are defined by the principles of the present disclosure. The designations have the same meaning as commonly understood by one of ordinary skill in the art to which they pertain.
[0075] The section headings used herein are for organizational purposes only and are not intended to be limiting unless otherwise stated. This document should not be construed as limiting the subject matter covered.
[0076] Various features of the disclosure may be described in the context of a single embodiment, but these features may also be combined in a single embodiment. They may also be provided separately or in any suitable combination. Although for clarity, may be described in the context of separate embodiments, the present disclosure also provides a single embodiment. It can also be implemented in this state.
[0077] [Definition] Unless otherwise defined, all technical and scientific terms used herein are defined by the It has the meaning commonly understood by one of ordinary skill in the art to which the invention pertains. , provides those skilled in the art with general definitions of many of the terms used in this invention: n et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The C ambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary o f Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991).
[0078] In this application, the use of the singular includes the plural unless specifically stated otherwise. As used herein, the singular forms "a," "an," and "the" are used unless the context clearly indicates otherwise. Please note that in this application, "or" includes plural referents unless otherwise indicated. The use of "includes" means "and / or" unless otherwise specified. "including" and "include," "includes," and "incl The use of other forms such as "uded" is non-limiting.
[0079] As used in this specification and claims, the terms "comprising" (and " "comprise" and "comprises" and all its forms), "having" g) (any of its forms, such as "have" and "has"), "include "including" ("include" and "includes" and any of its forms) or "including containing (any of its forms, such as "contains" and "contain") ) is inclusive or open-ended and does not exclude additional, unrecited elements or method steps. Any embodiment discussed herein may be used with respect to any method or composition of the present disclosure. It is believed that the same can be done, and vice versa. The method of the present disclosure can be achieved by
[0080] The terms "about" or "approximately" refer to a range of values as determined by one of ordinary skill in the art. This means that the value is within an acceptable margin of error for a particular value, which indicates how How it is measured or determined depends in part on the limitations of the measurement system. For example, "about" means, according to practice in the art, within 1 or more than 1 standard deviation. Alternatively, "about" can mean up to 20%, up to 10%, up to 5%, or Alternatively, it may refer to a range of up to 1% of the total mass of a biological system or process. This term means values within the same order of magnitude, preferably within 5-fold, more preferably within 2-fold. Where specific values are described in the application and claims, Unless otherwise stated, "about" means within an acceptable margin of error for that particular value. The term should be inferred.
[0081] In the specification, "some embodiments," "an embodiment," "one embodiment," or Reference to "another embodiment" may include any particular features, structures, or features are included in at least some embodiments of the present disclosure, but not necessarily all. This means that the embodiments are not necessarily included.
[0082] "Adenosine deaminase" refers to the enzyme that hydrolyzes the deamination of adenine or adenosine. In one embodiment, the term "de" refers to a polypeptide or fragment thereof that is capable of catalyzing the deactivation of a protein. The aminase or deaminase domain converts adenosine to inosine or deoxyribonucleic acid. Adenosine deamina catalyzes the hydrolytic deamination of adenosine to deoxyinosine. In some embodiments, the adenosine deaminase is a deoxyribonucleic acid (DNA) The compounds provided herein catalyze the hydrolytic deamination of adenine or adenosine in the ribozyme. Adenosine deaminases (e.g., engineered adenosine deaminases, evolved The adenosine deaminase (enzyme-activated adenosine deaminase) can be from any organism, such as a bacterium.
[0083] "Administering" refers to administering one or more of the products or compositions described herein to a patient or subject. By way of example and not limitation, providing Administration of the product or composition, e.g., injection, can be by intravenous (iv) injection, subcutaneous (sc) injection, intradermal ( This can be done by intraperitoneal (ip) injection, intramuscular (im) injection, or intravenous (id) injection. Any such route can be used. Parenteral administration can be by, for example, bolus injection. Alternatively, this can be achieved by gradually perfusing the blood over time. Administration can be by oral route, including but not limited to intranasal, rectal, or intracranial. Other modes of administration are also envisioned, such as vaginal, buccal, breast, intradermal, transdermal, etc.
[0084] An "agent" is any small molecule compound, antibody, nucleic acid molecule, or polypeptide. , or fragments thereof.
[0085] "Ameliorate" means to reduce, inhibit, or attenuate the occurrence or progression of a disease. means to cause, decrease, stop, or stabilize.
[0086] "Alteration" refers to any alteration that can be detected by standard art known methods such as those described herein. Such a change (increase or decrease) in the expression level or activity of a gene or polypeptide is As used herein, modification means a 10% change in expression level, preferably a 25% change. A change of 10% or more, more preferably a change of 40%, and most preferably a change of 50% or more in expression levels.
[0087] "Analog" means a molecule that is not identical but has similar functional or structural characteristics. For example, a polypeptide analog may possess the biological activity of the corresponding naturally occurring polypeptide. specific biochemical properties that enhance the function of the analog compared to the native polypeptide while retaining the Such biochemical modifications may occur, for example, by altering ligand binding. It can increase the protease resistance, membrane permeability, or half-life of the analogue. The analogs may include unnatural amino acids.
[0088] A "base editor (BE)" or "nucleobase editor (NBE)" is a In various embodiments, the term "agent" refers to an agent that binds to a base and has nucleobase modifying activity. The editors include nucleobase-modifying polypeptides (e.g., deaminases) and polynucleotide proteases. The programmable nucleotide binding domain is coupled to a guide polynucleotide (e.g., a guide In various embodiments, the agent comprises a protein having base editing activity. modifying the bases (e.g., A, T, C, G, U) in the quality domain, i.e., nucleic acid molecules (e.g., DNA) In one embodiment, the polynucleotide is a biomolecular complex that includes a domain capable of binding to a target molecule. The programmable DNA binding domain is fused or linked to the deaminase domain. In one embodiment, the agent is a fusion protein comprising a domain with base editing activity. In another embodiment, the protein domain having base editing activity is a guide RNA. (e.g., an RNA-binding motif on a guide RNA and an R fused to a deaminase) In some embodiments, the domain having base editing activity binds to a nucleic acid. In some embodiments, the base editor can deaminate a base within a molecule. In some embodiments, the base editor can deaminate a base within the A molecule. Cytosine (C) or adenosine (A) in NA can be deaminated. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenosine base editor (ABE). In one embodiment, the deaminase is evolved from TadA. The DNA-binding domain capable of binding is a CRISPR-associated (e.g., Cas or Cpf1) enzyme. In the present study, the base editor was a catalytically dead gene fused to a deaminase domain. In some embodiments, the base editor is a deaminase domain of Cas9 (dCas9). In some embodiments, the base editor is a salt fused to an inhibitor of base excision repair (BER). In some embodiments, the inhibitor of base excision repair The inhibitor is a uracil DNA glycosylase inhibitor (UGI). The inhibitors of repair are inosine base excision repair inhibitors. For more information on base editors, see the National Institute of Genetics. International PCT application number PCT / 2017 / 045381 (International Publication No. PCT / US 2018 / 027078) and PCT / US 2016 / 058344 (International Publication No. PCT / US 2016 / 058344). No. 2017 / 070632, each of which is incorporated herein by reference in its entirety. Komor, AC, et al., “Programmable editing of a target base in genome mic DNA without double-stranded DNA cleavage”Nature 533, 420-424 (2016); Gaudel li, NM, et al., “Programmable base editing of A·T to G·C in genomic DNA wit hout DNA cleavage” Nature 551, 464-471 (2017); Komor, AC, et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to- T:A base editors with higher efficiency and product purity”Science Advances 3:e aao4774 (2017), and Rees, HA, et al., “Base editing: precision chemistry o n the genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec;19(12 ):770-788. See also doi: 10.1038 / s41576-018-0059-1 (the entire contents of which are incorporated by reference) (Incorporated herein).
[0089] "Cytidine deaminase" is a enzyme that converts amino groups into carbonyl groups through a deamination reaction. In one embodiment, the term "catalyzed polypeptide" refers to a polypeptide or fragment thereof that is capable of catalyzing , cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine PmCDA1 (Petromyzon marinus cytosine deaminase) from Petromyzon marinus 1, "PmCDA1"), AID (active AID) derived from mammals (e.g., humans, pigs, cows, horses, monkeys, etc.) Activation-induced cytidine deaminase (AICDA), and APOBEC are exemplary cytidine deaminases. It is Ze.
[0090] As an example, the cytidine base editor BE4 has the following nucleic acid sequence: Polynucleotide sequences having at least 95% or more identity to the sequence are also encompassed. ATGagctcagagactggcccagtggctgtggaccccacattgagacggcggatcgagccccatgagtttgaggtattctt cgatccgagagagctccgcaaggagacctgcctgctttacgaaattaattgggggggccggcactccatttggcgacata catcacagaacactaacaagcacgtcgaagtcaacttcatcgagaagttcacgacagaaagatatttctgtccgaacaca aggtgcagcattacctggtttctcagctggagcccatgcggcgaatgtagtagggccatcactgaattcctgtcaaggta tccccacgtcactctgtttatttacatcgcaaggctgtaccaccacgctgacccccgcaatcgacaaggcctgcgggatt tgatctcttcaggtgtgactatccaaattatgactgagcaggagtcaggatactgctggagaaactttgtgaattatagc ccgagtaatgaagcccactggcctaggtatccccatctgtgggtacgactgtacgttcttgaactgtactgcatcatact gggcctgcctccttgtctcaacattctgagaaggaagcagccacagctgacattctttaccatcgctcttcagtcttgtc attackgcgactgcccccacacattctctgggccaccgggttgaaatctggtggttcttctggtggttctagcggcagc gagactcccgggacctcagagtccgccacacccgaaagttctggtggttcttctggtggttctgataaaaagtattctat tggtttagccatcggcactaattccgttggatgggctgtcataaccgatgaatacaaagtaccttcaaagaaatttaagg tgttggggaacacagaccgtcattcgattaaaaagaatcttatcggtgccctcctattcgatagtggcgaaacggcagag gcgactcgcctgaaacgaaccgctcggagaaggtatacacgtcgcaagaaccgaatatgttacttacaagaaatttttag caatgagatggccaaagttgacgattctttctttcaccgtttggaagagtccttccttgtcgaagaggacaagaaacatg aacggcaccccatctttggaaacatagtagatgaggtggcatatcatgaaaagtacccaacgatttatcacctcagaaaa aagctagttgactcaactgataaagcggacctgaggttaatctacttggctcttgcccatatgataaagttccgtgggca ctttctcattgagggtgatctaaatccggacaactcggatgtcgacaaactgttcatccagttagtacaaacctataatc agttgtttgaagagaaccctataaatgcaagtggcgtggatgcgaaggctattcttagcgcccgcctctctaaatcccga cggctagaaaacctgatcgcacaattacccggagagaagaaaaatgggttgttcggtaaccttatagcgctctcactagg cctgacaccaaattttaagtcgaacttcgacttagctgaagatgccaaattgcagcttagtaaggacacgtacgatgacg atctcgacaatctactggcacaaattggagatcagtatgcggacttatttttggctgccaaaaaccttagcgatgcaatc ctcctatctgacatactgagagttaatactgagattaccaaggcgccgttatccgcttcaatgatcaaaaggtacgatga acatcaccaagacttgacacttctcaaggccctagtccgtcagcaactgcctgagaaatataaggaaatattctttgatc agtcgaaaaacgggtacgcaggttatattgacggcggagcgagtcaagaggaattctacaagtttatcaaacccatatta gagaagatggatgggacggaagagttgcttgtaaaactcaatcgcgaagatctactgcgaaagcagcggactttcgacaa cggtagcattccacatcaaatccacttaggcgaattgcatgctatacttagaaggcaggaggatttttatccgttcctca aagacaatcgtgaaaagattgagaaaatcctaacctttcgcataccttactatgtgggacccctggcccgagggaactct cggttcgcatggatgacaagaaagtccgaagaaacgattactccatggaattttgaggaagttgtcgataaaggtgcgtc agctcaatcgttcatcgagaggatgaccaactttgacaagaatttaccgaacgaaaaagtattgcctaagcacagtttac tttacgagtatttcacagtgtacaatgaactcacgaaagttaagtatgtcactgagggcatgcgtaaacccgcctttcta agcggagaacagaagaaagcaatagtagatctgttattcaagaccaaccgcaaagtgacagttaagcaattgaaagagga ctactttaagaaaattgaatgcttcgattctgtcgagatctccggggtagaagatcgatttaatgcgtcacttggtacgt atcatgacctcctaaagataattaaaagaaggacttcctggataacgaagagaatgaagatatcttagaagatatagtg ttgactcttaccctctttgaagatcgggaaatgattgaggaaagactaaaaacatacgctcacctgttcgacgataaggt tatgaaacagttaaagaggcgtcgctatacgggctggggacgattgtcgcggaaacttatcaacgggataagagacaagc aaagtggtaaaactattctcgattttctaaagagcgacggcttcgccaataggaactttatgcagctgatccatgatgac tctttaaccttcaaagaggatatacaaaaggcacaggtttccggacaaggggactcattgcacgaacatattgcgaatct tgctggttcgccagccatcaaaaagggcatactccagacagtcaaagtagtggatgagctagttaaggtcatgggacgtc acaaaccggaaaacattgtaatcgagatggcacgcgaaaatcaaacgactcagaaggggcaaaaaaacagtcgagagcgg atgaagagaatagaagagggtattaaagaactgggcagccagatcttaaaggagcatcctgtggaaaatacccaattgca gaacgagaaactttacctctattacctacaaaatggaagggacatgtatgttgatcaggaactggacataaaccgtttat ctgattacgacgtcgatcacattgtaccccaatcctttttgaaggacgattcaatcgacaataaagtgcttacacgctcg gataagaaccgagggaaaagtgacaatgttccaagcgaggaagtcgtaaagaaaatgaagaactattggcggcagctcct aaatgcgaaactgataacgcaaagaaagttcgataacttaactaaagctgagaggggtggcttgtctgaacttgacaagg ccggatttattaaacgtcagctcgtggaaacccgccaaatcacaaagcatgttgcacagatactagattcccgaatgaat acgaaatacgacgagaacgataagctgattcgggaagtcaaagtaatcactttaaagtcaaaattggtgtcggacttcag aaaggattttcaattctataaagttagggagataaataactaccaccatgcgcacgacgcttatcttaatgccgtcgtag ggaccgcactcattaagaaatacccgaagctagaaagtgagtttgtgtatggtgattacaaagtttatgacgtccgtaag atgatcgcgaaaagcgaacaggagataggcaaggctacagccaaatacttcttttattctaacattatgaatttctttaa gacggaaatcactctggcaaacggagagatacgcaaacgacctttaattgaaaccaatggggagacaggtgaaatcgtat gggataagggccgggacttcgcgacggtgagaaaagttttgtccatgccccaagtcaacatagtaaagaaaactgaggtg cagaccggagggttttcaaaggaatcgattcttccaaaaaggaatagtgataagctcatcgctcgtaaaaaggactggga cccgaaaaagtacggtggcttcgatagccctacagttgcctattctgtcctagtagtggcaaaagttgagaagggaaaat ccaagaaactgaagtcagtcaaagaattattggggataacgattatggagcgctcgtcttttgaaaagaaccccatcgac ttccttgaggcgaaaggttacaaggaagtaaaaaaggatctcataattaaactaccaaagtatagtctgtttgagttaga aaatggccgaaaacggatgttggctagcgccggagagcttcaaaaggggaacgaactcgcactaccgtctaaatacgtga atttcctgtatttagcgtcccattacgagaagttgaaaggttcacctgaagataacgaacagaagcaactttttgttgag cagcacaaacattatctcgacgaaatcatagagcaaatttcggaattcagtaagagagtcatcctagctgatgccaatct ggacaaagtattaagcgcatacaacaagcacagggataaacccatacgtgagcaggcggaaaatattatccatttgttta ctcttaccaacctcggcgctccagccgcattcaagtattttgacacaacgatagatcgcaaacgatacacttctaccaag gaggtgctagacgcgacactgattcaccaatccatcacgggattatatgaaactcggatagatttgtcacagcttggggg tgactctggtggttctggaggatctggtggttctactaatctgtcagatattattgaaaaggagaccggtaagcaactgg ttatccaggaatccatcctcatgctcccagaggaggtggaagaagtcattgggaacaagccggaaagcgatatactcgtg cacaccgcctacgacgagagcaccgacgagaatgtcatgcttctgactagcgacgcccctgaatacaagccttgggctct ggtcatacaggatagcaacggtgagaacaagattaagatgctctctggtggttctggaggatctggtggttctactaatc tgtcagatattattgaaaaggagaccggtaagcaactggttatccaggaatccatcctcatgctcccagaggaggtggaa gaagtcattgggaacaagccggaaagcgatatactcgtgcacaccgcctacgacgagagcaccgacgagaatgtcatgct tctgactagcgacgcccctgaatacaagccttgggctctggtcatacaggatagcaacggtgagaacaagattaagatgc tctctggtggttctaaaaggacggagggatcagagttcgagagtccgaaaaaaaaacgaaggtcgaataa
[0091] BE4 nucleic acid sequence is provided below: atgtcatccgaaaccgggccagtggccgtagacccaacactcaggaggcggatagaaccccatgagtttgaagtgttctt cgaccccagagagctgcgcaaagagacttgcctcctgtatgaaataaattgggggggtcgccattcaatttggaggcaca ctagccagaatactaacaaacacgtggaggtaaattttatcgagaagtttaccaccgaaagatacttttgccccaataca cggtgttcaattacctggtttctgtcatggagtccatgtggagaatgtagtaggcgataactgagttcctgtctcgata tcctcacgtcacgttgtttatatacatcgctcggctttatcaccatgcggacccgcggaacaggcaaggtcttcgggacc tcatatcctctggggtgaccatccagataatgacggagcaagagagcggatactgctggcgaaactttgttaactacagc ccaagcaatgaggcacactggcctagatatccgcatctctgggttcgactgtatgtccttgaactgtactgcataattct gggacttcgccatgcttgaacattctgcggcggaaacaaccacagctgaccttttcacgattgctctccaaagttgtc actaccagcgattgccaccccacatcttgtgggctactggactcaagtctggaggaagttcaggcggaagcagcgggtct gaaacgcccggaacctcagagagcgcaacgcccgaaagctctggagggtcaagtggtggtagtgataagaaatactccat cggcctcgccatcggtacgaattctgtcggttgggccgttatcaccgatgagtacaaggtcccttctaagaaattcaagg ttttgggcaatacagaccgccattctataaaaaaaaacctgatcggcgcccttttgtttgacagtggtgagactgctgaa gcgactcgcctgaagcgaactgccaggaggcggtatacgaggcgaaaaaaccgaatttgttacctccaggagattttctc aaatgaaatggccaaggtagatgatagtttttttcaccgcttggaagaaagttttctcgttgaggaggacaaaaagcacg agaggcacccaatctttggcaacatagtcgatgaggtcgcataccatgagaaatatcctacgatctatcatctccgcaag aagctggtcgatagcacggataaagctgacctccggctgatctaccttgctcttgctcacatgattaaattcaggggcca tttcctgatagaaggagacctcaatcccgacaattctgatgtcgacaaactgtttattcagctcgttcagacctataatc aactctttgaggagaaccccatcaatgcttcaggggtggacgcaaaggccattttgtccgcgcgcttgagtaaatcacga cgcctcgagaatttgatagctcaactgccgggtgagaagaaaaacgggttgtttgggaatctcatagcgttgagtttggg acttacgccaaactttaagtctaactttgatttggccgaagaatgccaaattgcagctgtccaaagatacctatgatgacg acttggataaccttcttgcgcagattggtgaccaatacgcggatctgtttcttgccgcaaaaaatctgtccgacgccata ctcttgtccgatatactgcgcgtcaatactgagataactaaggctcccctcagcgcgtccatgattaaaagatacgatga gcaccaccaagatctcactctgttgaaagccctggtcgccagcagcttccagagaagtataaggagatattttcgacc aatctaaaaacggctatgcgggttacattgacggtggcgcctctcaagaagaattctacaagtttataaagccgatactt gagaaaatggacggtacagaggaattgttggttaagctcaatcgcgaggacttgttgagaaagcagcgcacatttgacaa tggtagtattccacaccagattcatctgggcgagttgcatgccattcttagaagacaagaagatttttatccgtttctga aagataacagagaaaagattgaaaagatacttacctttcgcataccgtattatgtaggtcccctggctagagggaacagt cgcttcgcttggatgactcgaaaatcagaagaaacaataaccccctggaattttgaagaagtggtagataaaaggtgcgag tgcccaatcttttattgagcggatgacaaattttgacaagaatctgcctaacgaaaaggtgcttcccaagcattcccttt tgtatgaatactttacagtatataatgaactgactaaagtgaagtacgttaccgaggggatgcgaaagccagcttttctc agtggcgagcagaaaaaagcaatagttgacctgctgttcaagacgaataggaaggttaccgtcaaacagctcaaagaaga ttactttaaaaagatcgaatgttttgattcagttgagataagcggagtagaggatagatttaacgcaagtcttggaactt atcatgaccttttgaagatcatcaaggataaagattttttggacaacgaggagaatgaagatatcctggaagatatagta cttaccttgacgctttttgaagatcgagagatgatcgaggagcgacttaagacgtacgcacatctctttgacgataaggt tatgaaacaattgaaacgccggcggtatactggctggggcaggctttctcgaaagctgattaatggtatccgcgataagc agtctggaaagacaatccttgactttctgaaaagtgatggatttgcaaatagaaactttatgcagcttatacatgatgac tctttgacgttcaaggaagacatccagaaggcacaggtatccggccaaggggatagcctccatgaacacatagccaacct ggccggctcaccagctattaaaaagggaatattgcaaaccgttaaggttgttgacgaactcgttaaggttatgggccgac acaaaccagagaatatcgtgattgagatggctagggagaatcagaccactcaaaaaggtcagaaaaattctcgcgaaagg atgaagcgaattgaagagggaatcaaagaacttggctctcaaattttgaaagagcacccggtagaaaacactcagctgca gaatgaaaagctgtatctgtattatctgcagaatggtcgagatatgtacgttgatcaggagctggatatcaataggctca gtgactacgatgtcgaccacatcgttcctcaatctttcctgaaagatgactctatcgacaacaaagtgttgacgcgatca gataagaaccggggaaaatccgacaatgtaccctcagaagaagttgtcaagaagatgaaaaactattggagacaattgct gaacgccaagctcataacacaacgcaagttcgataacttgacgaaagccgaaagaggtgggttgtcagaattggacaaag ctggctttattaagcgccaattggtggagacccggcagattacgaaacacgtagcacaaattttggattcacgaatgaat accaaatacgacgaaaacgacaaattgatacgcgaggtgaaagtgattacgcttaagagtaagttggtttccgatttcag gaaggattttcagttttacaaagtaagagaaataaacaactaccaccacgcccatgatgcttacctcaacgcggtagttg gcacagctcttatcaaaaaatatccaaagctggaaagcgagttcgtttacggtgactataaagtatacgacgttcggaag atgatagccaaatcagagcaggaaattgggaaggcaaccgcaaaatacttcttctattcaaacatcatgaacttctttaa gacggagattacgctcgcgaacggcgaaatacgcaagaggcccctcatagagactaacggcgaaaccggggagatcgtat gggacaaaggacgggactttgcgaccgttagaaaagtactttcaatgccacaagtgaatattgttaaaaagacagaagta caaacaggggggttcagtaaggaatccattttgcccaagcggaacagtgataaattgatagcaaggaaaaaagattggga ccctaagaagtacggtggtttcgactctcctaccgttgcatattcagtccttgtagttgcgaaagtggaaaaggggaaaa gtaagaagcttaagagtgttaaagagcttctgggcataaccataatggaacggtctagcttcgagaaaaatccaattgac tttctcgaggctaaaggttacaaggaggtaaaaaaggacctgataattaaactcccaaagtacagtctcttcgagttgga gaatgggaggaagagaatgttggcatctgcaggggagctccaaaaggggaacgagctggctctgccttcaaaatacgtga actttctgtacctggccagccactacgagaaactcaagggttctcctgaggataacgagcagaaacagctgtttgtagag cagcacaagcattacctggacgagataattgagcaaattagtgagttctcaaaaagagtaatccttgcagacgcgaatct ggataaagttctttccgcctataataagcaccgggacaagcctatacgagaacaagccgagaacatcattcacctcttta cccttactaatctgggcgcgccggccgccttcaaatacttcgacaccacgatagacaggaaaaggtatacgagtaccaaa gaagtacttgacgccactctcatccaccagtctataacagggttgtacgaaacgaggatagatttgtcccagctcggcgg cgactcaggagggtcaggcggctccggtggatcaacgaatctttccgacataatcgagaaagaaaccggcaaacagttgg tgatccaagaatcaatcctgatgctgcctgaagaagtagaagaggtgattggcaacaaacctgagtctgacattcttgtc cacaccgcgtatgacgagagcacggacgagaacgttatgcttctcactagcgacgcccctgagtataaaccatgggcgct ggtcatccaagattccaatggggaaaacaagattaagatgcttagtggtgggtctggagggagcggtgggtccacgaacc tcagcgacattattgaaaaagagactggtaaacaacttgtaatacaagagtctattctgatgttgcctgaagaggtggag gaggtgattgggaacaaaccggagtctgatatacttgttcataccgcctatgacgaatctactgatgagaatgtgatgct tttaacgtcagacgctcccgagtacaaaccctgggctctggtgattcaggacagcaatggtgagaataagattaaaatgt tgagtgggggctcaaagcgcacggctgacggtagcgaatttgagagccccaaaaaaaaacgaaaggtcgaataa
[0092] Another codon-optimized BE4 nucleic acid sequence (GeneArt from ThermoFisher Scientific) is provided below: as follows: atgagcagcgagacaggccctgtggctgtggatcctacactgcggagaagaatcgagccccacgagttcgaggtgttctt cgaccccagagagctgcggaaagagacatgcctgctgtacgagatcaactggggcggcagacactctatctggcggcaca caagccagaacaccaacaagcacgtggaagtgaactttatcgagaagtttacgaccgagcggtacttctgccccaacacc agatgcagcatcacctggtttctgagctggtccccttgcggcgagtgcagcagagccatcaccgagtttctgtccagata tccccacgtgaccctgttcatctatatcgcccggctgtaccaccacgccgatcctagaaatagacagggactgcgcgacc tgatcagcagcggagtgaccatccagatcatgaccgagcaagagagcggctactgctggcggaacttcgtgaactacagc cccagcaacgaagcccactggcctagatatcctcacctgtgggtccgactgtacgtgctggaactgtactgcatcatcct gggcctgcctccatgcctgaacatcctgagaagaaagcagcctcagctgaccttcttcacaatcgccctgcagagctgcc actaccagagactgcctccacacatcctgtgggccaccggacttaagagcggaggatctagcggcggctctagcggatct gagacacctggcacaagcgagtctgccacacctgagagtagcggcggatcttctggcggctccgacaagaagtactctat cggactggccatcggcaccaactctgttggatgggccgtgatcaccgacgagtacaaggtgcccagcaagaaattcaagg tgctgggcaacaccgaccggcacagcatcaagaagaatctgatcggcgccctgctgttcgactctggcgaaacagccgaa gccaccagactgaagagaaccgccaggcggagatacacccggcggaagaaccggatctgctacctgcaagagatcttcag caacgagatggccaaggtggacgacagcttcttccacagactggaagagtccttcctggtggaagaggacaagaagcacg agcggcaccccatcttcggcaacatcgtggatgaggtggcctaccacgagaagtaccccaccatctaccacctgagaaag aaactggtggacagcaccgacaaggccgacctgagactgatctacctggctctggcccacatgatcaagttccggggcca ctttctgatcgagggcgatctgaaccccgacaacagcgacgtggacaagctgttcatccagctggtgcagacctacaacc agctgttcgaggaaaaccccatcaacgcctctggcgtggacgccaaggctatcctgtctgccagactgagcaagagcaga aggctggaaaacctgatcgcccagctgcctggcgagaagaagaatggcctgttcggcaacctgattgccctgagcctggg actgacccctaacttcaagagcaacttcgacctggccgaggatgccaaactgcagctgagcaaggacacctacgacgacg acctggacaatctgctggcccagatcggcgatcagtacgccgacttgtttctggccgccaagaacctgtccgacgccatc ctgctgagcgatatcctgagagtgaacaccgagatcacaaaggcccctctgagcgcctctatgatcaagagatacgacga gcaccaccaggatctgaccctgctgaaggccctcgttagacagcagctgccagagaagtacaaagagattttcttcgatc agtccaagaacggctacgccggctacattgatggcggagccagccaagaggaattctacaagttcatcaagcccatcctg gaaaagatggacggcaccgaggaactgctggtcaagctgaacagagaggacctgctgcggaagcagcggaccttcgacaa tggctctatccctcaccagatccacctgggagagctgcacgccattctgcggagacaagaggacttttacccattcctga aggacaaccgggaaaagatcgagaagatcctgaccttcaggatcccctactacgtgggaccactggccagaggcaatagc agattcgcctggatgaccagaaagagcgaggaaaccatcacaccctggaacttcgaggaagtggtggacaagggcgccag cgctcagtccttcatcgagcggatgaccaacttcgataagaacctgcctaacgagaaggtgctgcccaagcactccctgc tgtatgagtacttcaccgtgtacaacgagctgaccaaagtgaaatacgtgaccgagggaatgagaaagcccgcctttctg agcggcgagcagaaaaaggccattgtggatctgctgttcaagaccaaccggaaagtgaccgtgaagcagctgaaagagga ctacttcaagaaaatcgagtgcttcgacagcgtggaaatcagcggcgtggaagatcggttcaatgccagcctgggcacat accacgacctgctgaaaattatcaaggacaaggacttcctggacaacgaagagaacgaggacattctcgaggacatcgtg ctgaccctgacactgtttgaggacagagagatgatcgaggaacggctgaaaacatacgcccacctgttcgacgacaaagt gatgaagcaactgaagcggaggcggtacacaggctggggcagactgtctcggaagctgatcaacggcatccgggataagc agtccggcaagacaatcctggatttcctgaagtccgacggcttcgccaacagaaacttcatgcagctgatccacgacgac agcctgacctttaaagaggacatccagaaagcccaggtgtccggccaaggcgattctctgcacgagcacattgccaacct ggccggatctcccgccattaagaagggcatcctgcagacagtgaaggtggtggacgagcttgtgaaagtgatgggcagac acaagcccgagaacatcgtgatcgaaatggccagagagaaccagaccacacagaagggccagaagaacagccgcgagaga atgaagcggatcgaagagggcatcaaagagctgggcagccagatcctgaaagaacaccccgtggaaaacacccagctgca gaacgagaagctgtacctgtactacctgcagaatggacgggatatgtacgtggaccaagagctggacatcaaccggctga gcgactacgatgtggaccatatcgtgccccagagctttctgaaggacgactccatcgataacaaggtcctgaccagaagc gacaagaaccggggcaagagcgataacgtgccctccgaagaggtggtcaagaagatgaagaactactggcgacagctgct gaacgccaagctgattacccagcggaagttcgataacctgaccaaggccgagagaggcggcctgagcgaacttgataagg ccggcttcattaagcggcagctggtggaaacccggcagatcaccaaacacgtggcacagattctggactcccggatgaac actaagtacgacgagaatgacaagctgatccgggaagtgaaagtcatcaccctgaagtctaagctggtgtccgatttccg gaaggatttccagttctacaaagtgcgggaaatcaacaactaccatcacgcccacgacgcctacctgaatgccgttgttg gaacagccctgatcaagaagtatcccaagctggaaagcgagttcgtgtacggcgactacaaggtgtacgacgtgcggaag atgatcgccaagagcgaacaagagatcggcaaggctaccgccaagtactttttctacagcaacatcatgaactttttcaa gacagagatcaccctggccaacggcgagatccggaaaagacccctgatcgagacaaacggcgaaaccggggagatcgtgt gggataagggcagagattttgccacagtgcggaaagtgctgagcatgccccaagtgaatatcgtgaagaaaaccgaggtg cagacaggcggcttcagcaaagagtctatcctgcctaagcggaacagcgataagctgatcgccagaaagaaggactggga ccctaagaagtacggcggcttcgatagccctaccgtggcctattctgtgctggtggtggccaaagtggaaaagggcaagt ccaaaaagctcaagagcgtgaaagagctgctggggatcaccatcatggaaagaagcagctttgagaagaacccgatcgac tttctggaagccaagggctacaaagaagtcaagaaggacctcatcatcaagctccccaagtacagcctgttcgagctgga aaatggccggaagcggatgctggcctcagcaggcgaactgcagaaaggcaatgaactggccctgcctagcaaatacgtca acttcctgtacctggccagccactatgagaagctgaagggcagccccgaggacaatgagcaaaagcagctgtttgtggaa cagcacaagcactacctggacgagatcatcgagcagatcagcgagttctccaagagagtgatcctggccgacgctaacct ggataaggtgctgtctgcctataacaagcaccgggacaagcctatcagagagcaggccgagaatatcatccacctgttta ccctgaccaacctgggagcccctgccgccttcaagtacttcgacaccaccatcgaccggaagaggtacaccagcaccaaa gaggtgctggacgccacactgatccaccagtctatcaccggcctgtacgaaacccggatcgacctgtctcagctcggcgg cgattctggtggttctggcggaagtggcggatccaccaatctgagcgacatcatcgaaaaagagacaggcaagcagctcg tgatccaagaatccatcctgatgctgcctgaagaggttgaggaagtgatcggcaacaagcctgagtccgacatcctggtg cacaccgcctacgatgagagcaccgatgagaacgtcatgctgctgacaagcgacgcccctgagtacaagccttgggctct cgtgattcaggacagcaatggggagaacaagatcaagatgctgagcggaggtagcggaggcagtggcggaagcacaaacc tgtctgatatcattgaaaaagaaaccgggaagcaactggtcattcaagagtccattctcatgctcccggaagaagtcgag gaagtcattggaaacaaacccgagagcgatattctggtccacacagcctatgacgagtctacagacgaaaacgtgatgct cctgacctctgacgctcccgagtataagccctgggcacttgttatccaggactctaacggggaaaacaaaatcaaaatgt tgtccggcggcagcaagcggacagccgatggatctgagttcgagagccccaagaagaaacggaaggtggagtaa,
[0093] "Base editing activity" refers to the ability to chemically modify bases within a polynucleotide. In one embodiment, the first base is converted to the second base. The base editing activity is a cytidine deaminase activity that converts the target C·G to T·A, for example. In another embodiment, the base editing activity is adenosine deaminase, e.g., converting A·T to G·C. Enzyme activity.
[0094] The term "base editor system" refers to a system for editing nucleic acid bases of a target nucleotide sequence. In various embodiments, the base editor (BE) system comprises: (1) a polynucleotide sequence; A nucleotide-programmable nucleotide-binding domain and a nucleotide-deaminating nucleotide base and (2) a deaminase domain for polynucleotide programmability. The guide polynucleotide (e.g., guide RNA) is associated with a gene-binding domain. In embodiments, the base editor system comprises: (1) a polynucleotide programmable and a deaminase domain for deaminating nucleic acid bases. (1) base editors (BEs), and (2) polynucleotide programmable DNA binding domains. In some embodiments, the polynucleotide programmable The nucleotide-binding domain is a polynucleotide-programmable DNA-binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenine or adenosine base editor (AB E) is.
[0095] "β-globin (HBB) protein" refers to a protein with a sequence similar to that of NCBI accession number NP_000509. It refers to a polypeptide or fragment thereof having at least about 95% amino acid sequence identity. In certain embodiments, the β-globin protein contains one or more modifications relative to the following reference sequence: In one particular embodiment, the beta-globin protein associated with sickle cell disease contains an E6V (also called E7V) mutation. Exemplary β-globin amino acid sequences (e.g., , reference sequences) are provided below. 1 mvhltpeeks avtalwgkvn vdevggealg rllvvypwtq rffesfgdls tpdavmgnpk 61 vkahgkkvlg afsdglahld nlkgtfatls elhcdklhvd penfrllgnv lvcvlahhfg 121 keftppvqaa yqkvvagvan alahkyh
[0096] "HBB polynucleotide" means a nucleic acid molecule encoding a β-globin protein or The sequence of an exemplary HBB polynucleotide is available under NCBI accession number NM_000518. The columns are provided below: 1 acatttgctt ctgacacaac tgtgttcact agcaacctca aacagacacc atggtgcatc 61 tgactcctga ggagaagtct gccgttactg ccctgtgggg caaggtgaac gtggatgaag 121 ttggtggtga ggccctgggc aggctgctgg tggtctaccc ttggacccag aggttctttg 181 agtcctttgg ggatctgtcc actcctgatg ctgttatggg caaccctaag gtgaaggctc 241 atggcaagaa agtgctcggt gcctttagtg atggcctggc tcacctggac aacctcaagg 301 gcacctttgc cacactgagt gagctgcact gtgacaagct gcacgtggat cctgagaact 361 tcaggctcct gggcaacgtg ctggtctgtg tgctggccca tcactttggc aaagaattca 421 ccccaccagt gcaggctgcc tatcagaaag tggtggctgg tgtggctaat gccctggccc 481 acaagtatca ctaagctcgc tttcttgctg tccaatttct attaaaggtt cctttgttcc 541 ctaagtccaa ctactaaact gggggatatt atgaagggcc ttgagcatct ggattctgcc 601 taataaaaaa catttatttt cattgcaa
[0097] "HBG1 protein," i.e., Homo sapiens hemoglobin subunit γ1 (HBG 1) "Protein" refers to a protein that has at least about 95% identical amino acid sequence to the amino acid sequence of NCBI reference sequence number NM_000559.2. In some embodiments, the amino acid sequence of the polypeptide is a polypeptide having an amino acid sequence identity of the polypeptide or a fragment thereof. In some embodiments, the HBG1 protein may include one or more modifications to the following amino acid sequence: In certain embodiments, methods for treating or ameliorating sickle cell disease as described herein are provided. To achieve this, editing is performed on the regulatory regions associated with the HBG1 protein, such as the promoter. An exemplary HBG1 amino acid sequence is provided below: MGHFTEEDKATITSLWGKVNVEDAGGETLGRLLVVYPWTQRFFDSFGNLSSASAIMGNPKVKAHGKKVLTSLGDATKHLD DLKGTFAQLSELHCDKLHVDPENFKLLGNVLVTVLAIHFGKEFTPEVQASWQKMVTAVASALSSRYH.
[0098] "HBG1 polynucleotide" refers to a nucleic acid molecule that encodes an HBG1 protein or a fragment thereof. The nucleic acid sequences of exemplary HBG1 polynucleotides are provided below: 1 acactcgctt ctggaacgtc tgaggttatc aataagctcc tagtccagac gccatgg gtc 61 atttcacaga ggaggacaag gctactatca caagcctgtg gggcaaggtg aatgtgga ag 121 atgctggagg agaaaccctg ggaaggctcc tggttgtcta cccatggacc cagaggttc t 181 ttgacagctt tggcaacctg tcctctgcct ctgccatcat gggcaacccc aaagtcaag g 241 cacatggcaa gaaggtgctg acttccttgg gagatgccac aaagcacctg gatgatctc a 301 agggcacctt tgcccagctg agtgaactgc actgtgacaa gctgcatgtg gatcctgag a 361 acttcaagct cctgggaaat gtgctggtga ccgttttggc aatccatttc ggcaaagaa t 421 tcacccctga ggtgcaggct tcctggcaga agatggtgac tgcagtggcc agtgccctgt 481 cctccagata ccactgagct cactgcccat gattcagagc tttcaaggat aggctttat t 541 ctgcaagcaa tacaaataat aaatctattc tgctgagaga tcac
[0099] "HBG2 protein," i.e., Homo sapiens hemoglobin subunit γ2 (H BG2) protein is defined as having at least one amino acid sequence identical to that of NCBI reference sequence NM_000184.3 It refers to a polypeptide or fragment thereof having about 95% amino acid sequence identity. In this embodiment, the HBG2 protein comprises one or more modifications to the following amino acid sequence: In certain embodiments, the method of treating or treating sickle cell disease as described herein may be used. To improve the expression of HBG2, the regulatory regions associated with the protein, such as the promoter, are edited. An exemplary HBG2 amino acid sequence is provided below: MGHFTEEDKATITSLWGKVNVEDAGGETLGRLLVVYPWTQRFFDSFGNLSSASAIMGNPKVKAHGKKVLTSLGDAIKHLD DLKGTFAQLSELHCDKLHVDPENFKLLGNVLVTVLAIHFGKEFTPEVQASWQKMVTGVASALSSRYH
[0100] "HBG2 polynucleotide" means a nucleic acid molecule that encodes an HBG2 protein or a fragment thereof. The nucleic acid sequences of exemplary HBG2 polynucleotides are provided below: 1 acactcgctt ctggaacgtc tgaggttatc aataagctcc tagtccagac gccatgggtc 61 atttcacaga ggaggacaag gctactatca caagcctgtg gggcaaggtg aatgtggaag 121 atgctggagg agaaaccctg ggaaggctcc tggttgtcta cccatggacc cagaggttct 181 ttgacagctt tggcaacctg tcctctgcct ctgccatcat gggcaacccc aaagtcaagg 241 cacatggcaa gaaggtgctg acttccttgg gagatgccat aaagcacctg gatgatctca 301 agggcacctt tgcccagctg agtgaactgc actgtgacaa gctgcatgtg gatcctgaga 361 acttcaagct cctgggaaat gtgctggtga ccgttttggc aatccatttc ggcaaagaat 421 tcacccctga ggtgcaggct tcctggcaga agatggtgac tggagtggcc agtgccctgt 481 cctccagata ccactgagct cactgcccat gatgcagagc tttcaaggat aggctttatt 541 ctgcaagcaa tcaaataata aatctattct gctaagagat cacaca
[0101] "ALAS1 protein", i.e., Homo sapiens 5'-aminolevulinic acid synthase 1 ( ALAS1) protein refers to a protein having at least about 95% identical amino acid sequence to that of NCBI reference sequence NM_000688.6. % amino acid sequence identity or a fragment thereof. In particular, the ALAS1 protein may contain one or more modifications to the following amino acid sequence: In embodiments, as described herein, Edits are made to regulatory regions associated with the ALAS1 protein, such as promoters. An exemplary ALAS1 amino acid sequence is provided below: MESVVRRCPPFLSRVPQAFLQKAGKSLLFYAQNCPKMMEVGAKPAPRALSTAA VHYQQIKETPPASEKDKTAKAKVQQTPDGSQQSPDGTQLPSGHPLPATSQGTA SKCPFLAAQMNQRGSSVFCKASLELQEDVQEMNAVRKEVAETSAGPSVVSVK TDGGDPSGLLKNFQDIMQKQRPERVSHLLQDNLPKSVSTFQYDRFFEKKIDEKK NDHTYRVFKTVNRRAHIFPMADDYSDSLITKKQVSVWCSNDYLGMSRHPRVCG AVMDTLKQHGAGAGGTRNISGTSKFHVDLERELADLHGKDAALLFSSCFVAND STLFTLAKMMPGCEIYSDSGNHASMIQGIRNSRVPKYIFRHNDVSHLRELLQRSD PSVPKIVAFETVHSMDGAVCPLEELCDVAHEFGAITFVDEVHAVGLYGARGGGI GDRDGVMPKMDIISGTLGKAFGCVGGYIASTSSLIDTVRSYAAGFIFTTSLPPML LAGALESVRILKSAEGRVLRRQHQRNVKLMRQMLMDAGLPVVHCPSHIIPVRV ADAAKNTEVCDELMSRHNIYVQAINYPTVPRGEELLRIAPTPHHTPQMMNYFLE NLLVTWKQVGLELKPHSSAECNFCRRPLHFEVMSEREKSYFSGLSKLVSAQA
[0102] "ALAS1 polynucleotide" refers to a nucleic acid molecule that encodes an ALAS1 protein or a fragment thereof. The nucleic acid sequences of exemplary ALAS1 polynucleotides are provided below: aggctgctcc cggacaaggg caacgagcgt ttcgtttgga cttctcgact tgagtgcccg cctccttcgc cgc cgcctct gcagtcctca gcgcagttat gcccagttct tccccgctgtg gggacacgac cacggaggaa tccttg cttc agggactcgg gaccctgctg gaccccttcc tcgggtttag gggatgtggg gaccaggaga aagtcagga t ccctaagagt cttccctgcc tggatggatg agtggcttct tctccaccta gattctttcc acaggagcca g catacttcc tgaacatgga gagtgttgtt cgccgctgcc cattcttatc ccgagtcccc caggcctttc tgcagaaagc aggcaaatct ctgttgttct atg cccaaaa ctgccccaag atgatggaag ttggggccaa gccagcccct cgggcattgt ccactgcagc agtaca ctac caacagatca aagaaacccc tccggccagt gagaaagaca aaactgctaa ggccaaggtc caacagact c ctgatggatc ccagcagagt ccagatggca cacagcttcc gtctggacac cccttgcctg ccacaagcca g ggcactgca agcaaatgcc ctttcctggc agcacagatg aatcagagag gcagcagtgt cttctgcaaa gcca gtcttg agcttcagga ggatgtgcag gaaatgaatg ccgtgaggaa agaggttgct gaaacctcag caggccc cag tgtggttagt gtgaaaaccg atggagggga tcccagtgga ctgctgaaga acttccagga catcatgcaa aagcaaagac cagaaagagt gtctcatctt cttcaagata acttgccaaa atctgtttcc acttttcagt at gatcgttt ctttgagaaa aaaattgatg agaaaaagaa tgaccacacc tatcgagttt ttaaaactgt gaacc ggcga gcacacatct tccccatggc agatgactat tcagactccc tcatcaccaa aaagcaagtg tcagtctg gt gcagtaatga ctacctagga atgagtcgcc acccacgggt gtgtggggca gttatggaca ctttgaaaca acatggtgct ggggcaggtg gtactagaaa tatttctgga actagtaaat tccatgtgga cttagagcgg gag ctggcag acctccatgg gaaagatgcc gcactcttgt tttcctcgtg cttgtggcc aatgactcaa ccctct tcac cctggctaag atgatgccag gctgtgagat ttactctgat tctgggaacc atgcctccat gatccaaggg attcgaaaca gccgagtgcc aaa gtacatc ttccgccaca atgatgtcag ccacctcaga gaactgctgc aaagatctga cccctcagtc cccaag attg tggcatttga aactgtccat tcaatggatg gggcggtgtg cccactggaa gagctgtgtg atgtggccc a tgagtttgga gcaatcacct tcgtggatga ggtccacgca gtggggcttt atggggctcg aggcggaggg a ttggggatc gggatggagt catgccaaaa atggacatca tttctggaac acttggcaaa gcctttggtt gtgt tggagg gtacatcgcc agcacgagtt ctctgattga caccgtacgg tcctatgctg ctggcttcat cttcacc acc tctctgccac ccatgctgct ggctggagcc ctggagtctg tgcggatcct gaagagcgct gagggacggg tgcttcgccg ccagcaccag cgcaacgtca aactcatgag acagatgcta atggatgccg gcctccctgt tg tccactgc cccagccaca tcatccctgt gcgggttgca gatgctgcta aaaacacaga agtctgtgat gaact aatga gcagacataa catctacgtg caagcaatca attaccctac ggtgccccgg ggagaagagc tcctacgg at tgcccccacc cctcaccaca caccccagat gatgaactac ttccttgaga atctgctagt cacatggaag caagtggggc tggaactgaa gcctcattcc tcagctgagt gcaacttctg caggaggcca ctgcattttg aag tgatgag tgaaagagag aagtcctatt tctcaggctt gagcaagttg gtatctgctc aggcctgagc atgacc tcaa ttatttcact taaccccagg ccattatcat atccagatgg tcttcagagt tgtctttata tgtgaattaa gttatattaa att ttaatct atagtaaaaa catagtcctg gaaataaatt cttgcttaaa tggtg
[0103] "BCL11A" protein, i.e., Homo sapiens B-cell CLL / lymphoma 11A (BCL11A) The protein (zinc finger protein) is identified by the amino acid sequence of GenBank accession number ADL_14508.1. It refers to a polypeptide or fragment thereof having at least about 95% amino acid sequence identity with the sequence. In some embodiments, the BCL11A protein has the amino acid sequence: In certain embodiments, base editing can involve one or more modifications to the BCL11A protein. or in association with the BCL11A protein, occurs in a regulatory region, e.g., a promoter , e.g., in beta-thalassemia and sickle disease by increasing fetal hemoglobin production Treat or ameliorate diseases such as serotonin-dependent leukemia (SCD). γ-globin is highly expressed in the hematopoietic lineage and during the transition from fetal to adult erythropoiesis BCL11A plays a role in the switch from fetal hemoglobin expression to β-globin expression. It may play a role in suppressing globin production and may also be involved in the pathogenesis of lymphoma. Translocations associated with B cell malignancies have been found to deregulate BCL11A expression. An example of the amino acid sequence of human BCL11A is shown below.
[0104] MSRRKQGKPQHLSKREFSPEPLEAILTDDEPDHGPLGAPEGDHDLLTCGQCQMNFPLGDILIFIEHKRKQCNGSLCLEKA VDKPPSPSPIEMKKASNPVEVGIQVTPEDDDCLSTSSRGICPKQEHIADKLLHWRGLSSPRSAHGALIPTPGMSAEYAPQ GICKDEPSSYTCTTCKQPFTSAWFLLQHAQNTHGLRIYLESEHGSPLTPRVGIPSGLGAECPSQPPLHGIHIADNNPFNL LRIPGSVSREASGLAEGRFPPTPPLFSPPPRHHLDPHRIERLGAEEMALATHHPSAFDRVLRLNPMAMEPPAMDFSRRLR ELAGNTSSPPLSPGRPSPMQRLLQPFQPGSKPPFLATPPLPPLQSAPPPSQPPVKSKSCEFCGKTFKFQSNLVVHRRSHT GEKPYKCNLCDHACTQASKLKRHMKTHMHKSSPMTVKSDDGLSTASSPEPGTSDLVGSASSALKSVVAKFKSENDPNLIP ENGDEEEEEDDEEEEEEEEEEEEELTESERVDYGFGLSLEAARHHENSSRGAVVGVGDESRALPDVMQGMVLSSMQHFSE AFHQVLGEKHKRGHLAEAEGHRDTCDEDSVAGESDRIDDGTVNGRGGCSPGESASGGLSKKLLLGSPSSLSPFSKRIKLEK EFDLPPAAMPNTENVYSQWLAGYAASRQLKDPFLSFGDSRQSPFASSSEHSSENGSLRFSTPPGELDGGISGRSGTGSGG STPHISGPGPGRPSSKEGRRSDTCEYCGKVFKNCSNLTVHRRSHTGERPYKCELCNYACAQSSKLTRHMKTHGQVGKDVY KCEICKMPFSVYSTLEKHMKKWHSDRVLNNDIKTE
[0105] A "BCL11A polynucleotide" is a nucleic acid that encodes a BCL11A protein or a fragment thereof. Exemplary human BCL11A (isoform 1) polynucleotides, reference sequence no. The nucleic acid sequence of No. GU 324937.1 is provided below: atgtctcgccgcaagcaaggcaaaccccagcacttaagcaaacgggaattctcgccccgagcctcttgaagccattcttac agatgatgaaccagaccacggcccgttgggagctccagaaggggatcatgacctcctcacctgtgggcagtgccagatga acttcccattgggggacatt cttattttatcgagcacaaacggaaacaatgcaatggcagcctctgcttagaaaaagc tgtggataagccaccttccccttcaccaatcgagatgaaaaaagcatccaatcccgtggag gttggcatccaggtcacg ccagaggatgacgattgtttatcaacgtcatctagaggaatt tgcccccaaacaggaacacatagcagataaacttctgc actggaggggcctctcctcccctcgttctgcacatggagctctaatccccacgcctgggatgagtgcagaatatgccccg cag ggtatttgtaaagatgagcccagcagctacacatgtacaacttgcaaacagccattcacc agtgcatggtttctc ttgcaacacgcacagaacactcatggattaagaatctacttagaaagcgaacacggaagtcccctgaccccgcgggttgg tatcccttcaggactaggtgcagaa tgtccttcccagccacctctccatgggattcatattgcagacaataaccccttt aacctg ctaagaataccaggatcagtatcgagagaggcttccggcctggcagaagggcgctttccacccactccccccc tgtttagtccaccaccgagacatcacttggacccccaccgcatagagcgcctgggggcggaagagatggccctggccacc catcacccgagtgcctttgacagggtgctgcggttgaatccaatggctatggagcctcccgccatggatttctctaggag acttagagagctggcagggaacacgtctagcccaccgctgtccccaggccggcccagccctatgcaaaggttactgcaac cattccagccaggtagcaagccgcccttcctggcgacgccccccctccctcctctgcaatccgcccctcctccctcccag cccccggtcaagtccaagtcatgcgagttctgcggcaagacgttcaaatttcagagcaacctggtggtgcaccggcgcag ccacacgggcgagaagccctacaagtgcaacctgtgcgaccacgcgtgcacccaggccagcaagctgaagcgccacatga agacgcacatgcacaaatcgtcccccatgacggtcaagtccgacgacggtctctccaccgccagctccccggaacccggc accagcgacttggtgggcagcgccagcagcgcgctcaagtccgtggtggccaagttcaagagcgagaacgaccccaacct gatcccggagaacggggacgaggaggaagaggaggacgacgaggaagaggaagaagaggaggaagaggaggaggaggagc tgacggagagcgagagggtggactacggcttcgggctgagcctggaggcggcgcgccaccacgagaacagctcgcggggc gcggtcgtgggcgtgggcgacgagagccgcgccctgcccgacgtcatgcagggcatggtgctcagctccatgcagcactt cagcgaggccttccaccaggtcctgggcgagaagcataagcgcggccacctggccgaggccgagggccacagggacactt gcgacgaagactcggtggccggcgagtcggaccgcatagacgatggcactgttaatggccgcggctgctccccgggcgag tcggcctcggggggcctgtccaaaaagctgctgctgggcagccccagctcgctgagccccttctctaagcgcatcaagct cgagaaggagttcgacctgcccccggccgcgatgcccaacacggagaacgtgtactcgcagtggctcgccggctacgcgg cctccaggcagctcaaagatcccttccttagcttcggagactccagacaatcgccttttgcctcctcgtcggagcactcc tcggagaacgggagcttgcgcttctccacaccgcccggggagctggacggagggatctcggggcgcagcggcacgggaag tggagggagcacgccccatattagtggtccgggcccggggcaggcccagctcaaaagagggcagacgcagcgacacttgtg agtactgtgggaaagtcttcaagaactgtagcaatctcactgtccacaggagaagccacacgggcgaaaggccttataaa tgcgagctgtgcaactatgcctgtgcccagagtagcaagctcaccaggcacatgaaaacgcatggccaggtggggaagga cgtttacaaatgtgaaatttgtaagatgccttttagcgtgtacagtaccctggagaaacacatgaaaaaatggcacagtg atcgagtgttgaataatgatataaaaactgaatag.
[0106] In some embodiments, the nucleobase editor system comprises a plurality of base editing components. For example, the nucleobase editor system can include multiple deaminases. In some embodiments, the nuclease base editing system can include one or more cytogenes. In some embodiments, the enzyme may comprise one or more adenosine deaminases. In embodiments, a single guide polynucleotide is utilized to target different deaminases. In some embodiments, a pair of guide polynucleotides can be targeted to a nucleic acid sequence. The nucleic acids can be used to target different deaminases to target nucleic acid sequences. do.
[0107] Nucleobase components of base editor systems and polynucleotide programmable nucleases The nucleotide binding moieties can be covalently or non-covalently bound to one another. For example, In some embodiments, the deaminase domain is a polynucleotide programming Targeted to a target nucleotide sequence by a nucleotide-binding domain that can be In some embodiments, the polynucleotide has a programmable nucleotide binding domain. In some embodiments, the deaminase domain may be fused or linked to the deaminase domain. The polynucleotide programmable nucleotide binding domain is a deaminase The deaminase domain binds to the ATPase domain by non-covalently interacting with or binding to the ATPase domain. It can be targeted to a target nucleotide sequence. For example, in some embodiments In the method, a nucleobase editing component, e.g., a deaminase component, is used to program polynucleotides. and a further heterologous moiety or domain that is part of the nucleotide-binding domain that can be bound. Additional heterologous moieties or domains that can interact, associate, or form complexes. In some embodiments, the additional heterologous moiety can include a polypeptide. Some examples are capable of binding, interacting, associating, or forming complexes with In embodiments, the additional heterologous moiety binds to, interacts with, or associates with the polynucleotide, or can form a complex. In some embodiments, the additional heterologous moiety In some embodiments, additional The heterologous moiety can be attached to a polypeptide linker. In this case, the additional heterologous moiety can be attached to the polynucleotide linker. The species portion may be a protein domain. In some embodiments, additional heterologous sequences may be present. The seed portion contains the K homology (KH) domain, the MS2 coat protein domain, and the PP7 coat protein. domain, SfMu Com coat protein domain, steryl α motif, telomerase Ku binding binding motif and Ku protein, telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif.
[0108] The base editing system can further comprise a guide polynucleotide component. The components of the assembly system are linked by covalent bonds, non-covalent interactions, or their binding and It should be understood that the molecules may be coupled to one another via any combination of interactions. In embodiments, the deaminase domain is cleaved by the guide polynucleotide to the target nucleoside. For example, in some embodiments, the base edit The nucleobase editing component of the guide system, e.g., the deaminase component, is a portion of the guide polynucleotide. or segments (e.g., polynucleotide motifs) that interact with, bind to, or complex with Further heterologous moieties or domains that can form complexes (e.g., RNA or DNA binding proteins) In some embodiments, the additional Additional heterologous moieties or domains (e.g., polynucleotides such as RNA or DNA binding proteins) The deaminase domain (binding domain) may be fused or linked to the deaminase domain. In the form, the additional heterologous moiety binds to, interacts with, associates with, or In some embodiments, the additional heterologous polypeptide can be complexed with the The seed moiety binds to, interacts with, associates with, or binds to a polynucleotide. In some embodiments, the additional heterologous moiety can form a complex. In some embodiments, the additional heterologous polynucleotide can be linked to a heterologous polynucleotide. The moiety can be attached to a polypeptide linker. In some embodiments, Additional heterologous moieties can be attached to the polynucleotide linker. may be a protein domain. In some embodiments, the additional heterologous moiety The K homology (KH) domain, the MS2 coat protein domain, and the PP7 coat protein domain are SfMu Com coat protein domain, sterile alpha motif, telomerase Ku binding motif Chief and Ku proteins, telomerase Sm7 binding motifs and Sm7 proteins, or It may be an RNA recognition motif.
[0109] In some embodiments, the base editing system inhibits base excision repair (BER) components. The components of the base editing system can further include covalently or non-covalently bound factors. or any combination of these bonds and interactions. It should be understood that inhibitors of BER components may also inhibit base excision repair. In some embodiments, the inhibitor of base excision repair is uracil DNA glycosylase. In some embodiments, the inhibitor of base excision repair can be an inosine In some embodiments, the inhibitor of base excision repair is a polynucleotide. Nucleotide-programmable nucleotide-binding domains allow target nucleotide sequences to be identified. In some embodiments, polynucleotides can be programmable. A suitable nucleotide binding domain can be fused or linked to an inhibitor of base excision repair. In some embodiments, the polynucleotide programmable nucleotide binding domain comprises: It can be fused or linked to a deaminase domain and an inhibitor of base excision repair. In some embodiments, the polynucleotide programmable nucleotide binding domain is , non-covalently interacting with an inhibitor of base excision repair or inhibiting base excision repair By associating with the inhibitor, the inhibitor targets the base excision repair inhibitor to the target nucleotide sequence. For example, in some embodiments, a base excision repair component can be targeted. The inhibitor of and a further heterologous moiety or domain that is capable of interacting with, associating with, or complexing with a further heterologous moiety or domain that is In some embodiments, the inhibitor of base excision repair comprises a heterologous portion or domain comprising The target nucleotide sequence can be targeted by a guide polynucleotide. For example, in some embodiments, the inhibitor of base excision repair is a guide polynucleotide interacts with, associates with, or binds to a portion or segment (e.g., a polynucleotide motif) of Further heterologous moieties or domains that can be complexed (e.g., RNA or DNA binding proteins) In some embodiments, the polypeptide may comprise a polynucleotide binding domain (e.g., a polynucleotide binding domain). Additional heterologous moieties or domains of the guide polynucleotide (e.g., RNA or DNA binding) Protein-like polynucleotide-binding domain) fused to an inhibitor of base excision repair In some embodiments, the additional heterologous moiety may be a polynucleotide. In some embodiments, the antibody can bind to, interact with, associate with, or complex with the antibody. In some cases, an additional heterologous moiety can be attached to the guide polynucleotide. In this embodiment, the additional heterologous moiety can be attached to the polypeptide linker. In some embodiments, the additional heterologous moiety is attached to the polynucleotide linker. The additional heterologous moiety may be a protein domain. In embodiments, the additional heterologous moiety is a K homology (KH) domain, an MS2 coat protein domain, or a nucleotide sequence encoding the nucleotide sequence of interest. PP7 coat protein domain, SfMu Com coat protein domain, sterile alpha motif, telomerase Ku binding motif and Ku protein, telomerase Sm7 binding motif The Sm7 and Sm8 proteins, or RNA recognition motifs.
[0110] The term "Cas9" or "Cas9 domain" refers to a Cas9 protein or a fragment thereof (e.g., Cas9 Cas9 active, inactive, or partially active DNA cleavage domains, and / or gRNA binding domains of Cas9. Cas9 nuclease refers to an RNA-guided nuclease containing a nucleotide sequence (a protein containing a binding domain). , casnl nuclease or CRISPR (clustered regularly interspaced short palindromic An exemplary Cas9 is the nucleotide sequence of ... The Streptococcus pyogenes Cas9 provided:
[0111] JPEG2025026840000002.jpg162167 (single underline: HNH domain; double underline: RuvC domain).
[0112] The term "conservative amino acid substitution" or "conservative variation" refers to a mutation in which an amino acid is replaced by another amino acid that shares a common characteristic. It refers to the substitution of an amino acid with another amino acid that has the same properties. It defines the common properties between individual amino acids. A functional method for determining the normality of amino acid changes between corresponding proteins of homologous organisms is The goal is to analyze the frequency of the results (Schulz, GE and Schirmer, RH, Principles of f Protein Structure, Springer-Verlag, New York (1979). According to such analysis, Amino acids within a group are preferentially exchanged with each other, thus affecting the overall protein structure. Define the group of amino acids that are most similar to each other in their effect on the (Schulz, GE and Schirmer, RH, supra). Non-limiting examples of conservative mutations include: For example, the amino acids arginine to lysine and the like that can maintain a positive charge are used. Reverse; aspartic acid to glutamic acid and vice versa, which can maintain a negative charge; free threonine to serine, which maintains the -OH; and asparagine to asparagine, which maintains the free NH Examples include amino acid substitutions such as glutamine.
[0113] The terms "coding sequence" or "protein-coding sequence" are used interchangeably herein. " refers to a segment of a polynucleotide that encodes a protein. The sequence is bounded by a start codon near the 5' end and a stop codon near the 3' end. The coding sequence is also called an open reading frame.
[0114] As used herein, the terms "deaminase" or "deaminase domain" and "deaminase domain" refer to refers to a protein or enzyme that catalyzes a deamination reaction. The enzyme or deaminase domain binds to uridine or deoxyuridine, respectively. Cytidine deaminase catalyzing the hydrolytic deamination of cytidine or deoxycytidine In some embodiments, the deaminase or deaminase domain is a cytosine deaminase. enzyme, which catalyzes the hydrolytic deamination of cytosine to uracil. In this case, the deaminase catalyzes the hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase is adenosine or Adenosine deamination catalyzes the hydrolytic deamination of adenine (A) to inosine (I). In some embodiments, the deaminase or deaminase domain is an adenosine deaminase. adenosine or deoxyinosine to inosine or deoxyinosine, respectively. catalyzes the hydrolytic deamination of deoxyadenosine. Adenosine deaminase hydrolyzes the adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., genetically engineered Adenosine deaminase (adenosine deaminase, evolved adenosine deaminase) is a nucleotide that binds to the ATP of any living organism, including bacteria. In some embodiments, the adenosine deaminase can be derived from E. coli, S. aureus, , derived from bacteria such as S. typhi, S. putrefaciens, H. influenzae, or C. crescentus In some embodiments, the adenosine deaminase is TadA deaminase. In such cases, the deaminase or deaminase domain may be derived from, for example, a human, a chimpanzee, Natural deaminases from organisms such as gorillas, monkeys, cows, dogs, rats, or mice In some embodiments, the deaminase or deaminase domain is a non- For example, in some embodiments, a deaminase or deaminase The deaminase domain is at least 50%, at least 55%, or at least At least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85% , at least 90%, at least 91%, at least 92%, at least 93%, at least, At least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 9 9%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9 For example, the deaminase domain is described in International PCT Application No. PCT / 2017 / 04538 1 (International Publication No. 2018 / 027078) and PCT / US 2016 / 058344 (International Publication No. 2017 / 070632) each of which is incorporated herein by reference in its entirety. AC, et al., “Programmable editing of a target base in genomic DNA without do “Uble-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, NM, et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage ” Nature 551, 464-471 (2017); Komor, AC, et al., “Improved base excision re pair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao4774 (2017), and Rees, HA, et al., “Base editing: precision chemistry on the genome and tr anscriptome of living cells.” Nat Rev Genet. 2018 Dec;19(12):770-788. doi: 10. See also 1038 / s41576-018-0059-1, the entire contents of which are incorporated herein by reference.
[0115] A "detectable label" means a label that, when attached to a molecule of interest, is detectable by spectroscopic, photochemical, or biochemical means. By "detectable" is meant a composition that renders it detectable through biological, immunochemical, or chemical means. For example, useful labels include radioisotopes, magnetic beads, metal beads, colloidal particles, fluorescent Dyes, electron-dense reagents, enzymes (e.g., those commonly used in ELISA), biotin, zygotes xygenin, or hapten.
[0116] "Disease" means any condition that damages or interferes with the normal function of a cell, tissue, or organ. or disorders. Examples of diseases include retinitis pigmentosa, Usher syndrome, sickle cell disease, β-thalassemia, hereditary persistent fetal hemoglobin (HPFH), α1-antitrypsin deficiency (A1 AD), hepatic porphyria, medium-chain acyl-CoA dehydrogenase (ACADM) deficiency, lysosomal acidosis Lysosomal acid lipase (LAL) deficiency, phenylketonuria, hemochromatosis, Von Gierke's disease, Pompe disease In one embodiment, the diseases include rheumatoid arthritis, Gaucher's disease, Hurler's syndrome, cystic fibrosis, and chronic pain. The disease is A1AD. In one embodiment, the disease is sickle cell disease (SCD), which is also known as "sickle cell disease." It is also called erythrocyte anemia.
[0117] An "effective amount" is an amount that is effective relative to an untreated patient or a disease-free, i.e., healthy, individual. , a drug or activity necessary to ameliorate the symptoms of a disease in a subject or patient in need thereof. Therapeutic Treatment of Disease The effective amount of active compound used to practice the described methods for This will vary depending on the age, weight, and general health of the subject. A veterinarian will determine the appropriate amount and dosage. Such an amount is called an "effective" amount. In this embodiment, an effective amount is a dose of a gene of interest in a cell (e.g., a cell in vitro or in vivo). In one embodiment, the amount of a base editor of the invention is sufficient to introduce an alteration into The amount is determined to achieve a therapeutic effect (e.g., retinitis pigmentosa, Usher syndrome, sickle cell disease ( SCD), β-thalassemia, hereditary persistence of fetal hemoglobin (HPFH), α-1 antitrypsin deficiency (A1AD), hepatic porphyria, medium-chain acyl-CoA dehydrogenase (ACADM) deficiency, Lysosomal acid lipase (LAL) deficiency, phenylketonuria, hemochromatosis, von Gierke disease, Pompe disease, Gaucher disease, Hurler syndrome, cystic fibrosis, or chronic The amount of base editor required to achieve such a therapeutic effect (reduce or control pain). The effect is sufficient to alter the virulence gene in all cells of the subject, tissue, or organ. It need not be about 1%, 5%, 10%, 25%, 50%, 75% of the cells present in the subject, tissue or organ. In one embodiment, the effective amount of the virulence gene is is associated with diseases (e.g., retinitis pigmentosa, Usher syndrome, sickle cell disease (SCD), beta thalassemia, Hereditary Persistence of Fetal Hemoglobin (HPFH), Alpha-1 Antitrypsin Deficiency (A1AD), Hepatic Porphyromonas reticulata (HRE) dyslipidemia, medium-chain acyl-CoA dehydrogenase (ACADM) deficiency, lysosomal acid lipase (LAL) Deficiency, phenylketonuria, hemochromatosis, von Goerke disease, Pompe disease, Goerke disease Improve one or more symptoms of rheumatoid arthritis (Scher's disease, Hurler's syndrome, cystic fibrosis, or chronic pain) is sufficient.
[0118] By "fragment" is meant a portion of a polypeptide or nucleic acid molecule. is at least 10%, 20%, 30%, 40%, 50%, 60% of the entire length of the reference nucleic acid molecule or polypeptide , 70%, 80%, or 90%. The fragments may be 10, 20, 30, 40, 50, 60, 70, 80, 90, or 1 00, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids It may contain an acid.
[0119] "Hybridization" means hydrogen bonding between complementary nucleobases, as defined by Watson- It can be a Crick, Hoogsteen or reversed Hoogsteen hydrogen bond. For example, Adenine and thymine are complementary nucleobases that form hydrogen bonds to form pairs.
[0120] The term "inhibitor of base repair" or "IBR" refers to an inhibitor of base repair by a nucleic acid repair enzyme, such as base excision repair. IBR refers to a protein that can inhibit the activity of an enzyme. It is an inhibitor of base excision repair. Examples of base repair inhibitors include APE1 and Endo III. , Endo IV, Endo V, Endo VIII, Fpg, hOGGl, hNEILl, T7 Endol, T4 PDG, UDG, hSMUGL and inhibitors of hAAG. In certain embodiments, the IBR is an inhibitor of Endo V or hAAG. In one embodiment, the IBR is a catalytically inactive EndoV or a catalytically inactive In one embodiment, the base repair inhibitor is an inhibitor of Endo V or hAAG. In some embodiments, the base repair inhibitor is a catalytically inactive EndoV or a catalytically inactive EndoV. In one embodiment, the base repair inhibitor is uracil glycosylase. UGI is a uracil-DNA glycosylase inhibitor that inhibits the base excision repair enzyme. In some embodiments, the UGI domain refers to a protein that can inhibit In some embodiments, the UGI provided herein comprises a wild-type UGI or a fragment thereof. GI proteins include fragments of UGI and proteins homologous to UGI or UGI fragments. In some embodiments, the base repair inhibitor is an inhibitor of inosine base excision repair. In some embodiments, the base repair inhibitor is a "catalytically inactive inosine-specific nuclease inhibitor." Without being bound by any particular theory, Although it is not desired to be bundled, catalytically inactive inosine glycosylase (e.g., alkyl Adenine glycosylase (AAG) can bind to inosine but not to the abasic site. Inosine cannot be produced or removed, thereby In some embodiments, the catalytic moiety is sterically blocked from DNA damage / repair machinery. Catalytically inactive inosine-specific nucleases are unable to bind to inosine in nucleic acids. but does not cleave nucleic acids. Non-limiting examples include catalytically inactive alkyl adenosine glycosyltransferases (e.g., of human origin). AAG nuclease, and catalytically inactive endonucleases (e.g., from E. coli) In some embodiments, catalytically active nucleases include EndoV nuclease. Inactive AAG nucleases may be due to the E125Q mutation or the corresponding mutation in another AAG nuclease. This includes mutations that
[0121] The terms "isolated," "purified," or "biologically pure" refer to a substance that is in its native state. Substances that have been removed to varying degrees from components normally associated with them when found in the natural state. "Isolated" refers to the degree of separation from the original source or surrounding environment. "Purified" refers to the degree of separation from the original source or surrounding environment. A "purified" or "biologically pure" protein is one that is free from impurities. that the substance does not materially affect the biological properties of the protein or cause other adverse consequences. In other words, the nucleic acid or peptide of the present invention is or, if produced by recombinant DNA technology, cellular material, viral material, or culture medium. If the substance is not naturally contained in the substance, or if it is chemically synthesized, chemical precursors or other chemical substances Purity and homogeneity are typically determined by analytical methods. Chemical techniques, such as polyacrylamide gel electrophoresis or high performance liquid chromatography The term "purified" refers to the degree to which a nucleic acid or protein is purified by electrophoresis. This can mean that the resulting protein essentially produces one band. For proteins that can undergo glycosylation, the different modifications are purified separately. This can result in different isolated proteins that can be isolated.
[0122] An "isolated polynucleotide" is a nucleic acid molecule that does not occur in the naturally occurring genome of the organism from which the nucleic acid molecule of the invention is derived. It means nucleic acid (e.g., DNA) that does not contain the genes adjacent to the gene. The term refers to, for example, those incorporated into vectors; autonomously replicating plasmids or viruses. integrated into the genomic DNA of prokaryotes or eukaryotes; or independently of other sequences another molecule (e.g., cDNA generated by PCR or restriction endonuclease digestion) or genomic or cDNA fragments). RNA molecules transcribed from DNA molecules, as well as hybrids encoding additional polypeptide sequences. It contains recombinant DNA that is part of the hybrid gene.
[0123] An "isolated polypeptide" is a polypeptide of the invention separated from components that naturally accompany it. Typically, a polypeptide is a polypeptide that is a protein with which it is naturally associated. A substance is isolated if it is at least 60% by weight free from substances and naturally occurring organic molecules. Preferably, the preparation comprises at least 75% by weight of soluble fiber, more preferably at least 90% by weight of soluble fiber, most preferably at least 10% by weight of soluble fiber. or at least 99% of the isolated polypeptide of the present invention. can be prepared by, for example, extraction from a natural source, expression of a recombinant nucleic acid encoding such a polypeptide, or the like. or by chemically synthesizing the protein. Purity can be achieved by any suitable method. Suitable methods, such as column chromatography, polyacrylamide gel electrophoresis, or can be measured by HPLC analysis.
[0124] As used herein, the term "linker" refers to a molecule that binds two molecules or moieties (e.g., a Two components of a protein or ribonucleocomplex, or two domains of a fusion protein In, for example, polynucleotides containing programmable DNA binding domains (e.g., dCas9) and deamidation enzyme domains (e.g., adenosine deaminase or cytidine deaminase) refers to a covalent linker (e.g., a covalent bond), a non-covalent linker, a chemical group, or a molecule that The linker can be used to link different components or components of the base editor system. For example, in some embodiments, The linker is a guide pole of the polynucleotide-programmable nucleotide binding domain. The nucleotide binding domain and the catalytic domain of the deaminase can be joined. In some embodiments, the linker connects the CRISPR polypeptide and the deaminase. In some embodiments, a linker can join the Cas9 and the deaminase. In some embodiments, a linker can join the dCas9 and the deaminase. In some embodiments, a linker can join the nCas9 and the deaminase. In embodiments, the linker can connect the guide polynucleotide and the deaminase. In some embodiments, the linker is a deaminating component of a base editor system. and a polynucleotide capable of linking a programmable nucleotide binding component. In some embodiments, the linker is a base editor system a polynucleotide-programmable nucleotide-binding component and a polynucleotide-binding component. In some embodiments, the linker can be used to link the base editor system. RNA-binding moiety of the amino acid component and polynucleotide-programmable nucleotide binding A linker is a molecule that can connect two groups, molecules, or other molecules to the RNA-binding portion of a component. Located between or sandwiched by moieties and formed by covalent or non-covalent interactions and thus the two can be linked together. In some embodiments, the linker may be an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is a DNA linker. In some embodiments, the linker may be an RNA linker. Thus, the linker can comprise an aptamer capable of binding to the ligand. In embodiments, the ligand can be a carbohydrate, peptide, protein, or nucleic acid. In some embodiments, the linker comprises an aptamer derived from a riboswitch. The aptamer is derived from the theophylline riboswitch and the thiamin TPP riboswitch, adenosine cobalamin (AdoCbl) riboswitch, S- Adenosylmethionine (SAM) riboswitch, SAH riboswitch, flavin mononucleotide FMN riboswitch, tetrahydrofolate riboswitch, lysine riboswitch, glycine riboswitch riboswitch, purine riboswitch, GlmS riboswitch, or prequeosin 1 (Pre Q1) A riboswitch. In some embodiments, the linker can be selected from a polypeptide or The antibody may comprise an aptamer bound to a protein domain, such as a polypeptide ligand. In some embodiments, the polypeptide ligand comprises a K homology (KH) domain, an MS2 coat protein domain, PP7 coat protein domain, SfMu Com coat protein domain, sterile α motif, telomerase Ku binding motif, Ku protein, telomerase Sm7 binding motif Chief and Sm7 proteins, or RNA recognition motifs. The peptide ligand can be part of a base editor system component. For example, a nucleobase editing system The component may include a deaminase domain and an RNA recognition motif.
[0125] In some embodiments, the linker is an amino acid or a plurality of amino acids (e.g., a peptide or In some embodiments, the linker may be about 5 to 100 amino acids in length. For example, lengths of approximately 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20-30 , 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, or 90-100 amino acids. In some embodiments, the linker has a length of about 100-150, 150-200, 200-250, 250-300, It can be 300-350, 350-400, 400-450, or 450-500 amino acids. Shorter linkers are also contemplated.
[0126] In some embodiments, the linker comprises an RNA programmable The gRNA-binding domain of a nuclease and a nucleic acid editing protein (e.g., cytidine or adenine) In some embodiments, the linker connects the catalytic domain of dCas9 to the catalytic domain of dCas9. and the nucleic acid editing protein. For example, a linker can be a linker that connects two groups, molecules, or other Located between or flanked by two groups, molecules, or other moieties, are linked to each other via a covalent bond, thus linking the two. The anchor is an amino acid or multiple amino acids (eg, a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5 to 200 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 70, 80, 95, 100, 150, 100, 250, 350, 400, 0, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 7 0, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140, 150 , 160, 175, 180, 190, or 200 amino acids. Longer or shorter linkers are also possible. In some embodiments, the linker is an amino acid sequence, which may be referred to as an XTEN linker. In some embodiments, the linker comprises the amino acid sequence SGGS. In some embodiments, the linker is (SGGS) n , (GGGS) n , (GGGGS) n , (G) n, (EAAAK) n, (GGS) n , SGSETPGTSESATPES, or (XP) n motif, or any combination thereof, wherein n is independently an integer from 1 to 30, and X is any amino acid. wherein n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. In the above, the linker contains multiple proline residues and has a length of 5 to 21, 5 to 14, 5 to 9, 5 to 7 amino acids. amino acids, such as PAPAP, PAPAPA, PAPAPAP, PAPAPAPA, P(AP)4, P(AP)7, P(AP) 10 Such proline-rich linkers are also referred to as "rigid" linkers. Also called.
[0127] In some embodiments, the base editor domain is SGGSSGSETPGTSESATPESSGGS, SGGSSG GSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSP Contains the amino acid sequence AGSPTSTEEGTSTE PSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS In some embodiments, the base editor domain is fused via a linker. The linker contains the amino acid sequence SGSETPGTSESATPES, which may also be referred to as a linker. In some embodiments, the linker is 24 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In one embodiment, the linker has the amino acid sequence SGGSSGGSSSGSETPGTES In some embodiments, the linker is 64 amino acids in length. In some embodiments, the linker has the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS In one embodiment, the linker is 92 amino acids in length and In some embodiments, the linker has the amino acid sequence PGSPAGSPTSTEEGTSESATPESGPG Contains TSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAP GTSTEPSEGSAPGTSESATPESGPGSEPATS.
[0128] As used herein, the term "mutation" refers to a change in a sequence, e.g., a nucleic acid or amino acid sequence. The substitution of a residue in a sequence of amino acids by another residue, or the deletion of one or more residues in the sequence. Mutations, as used herein, are typically made by identifying the original residue and then and identifying the newly substituted residue, Various methods for making amino acid substitutions (mutations) are provided herein. are well known in the art and are described, for example, in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012). In some embodiments, this The disclosed base editors do not generate a significant number of unintended mutations, e.g., unintended point mutations. without creating an "intended mutation" in a nucleic acid (e.g., a nucleic acid in a subject's genome), e.g. Point mutations can be efficiently generated. In some embodiments, the intended The mutations are generated by a guide polynucleotide sequence specifically designed to produce the intended mutation. A specific base editor (e.g., a cytidine base editor or This mutation is caused by an adenosine base editor (ABE).
[0129] Generally, the amino acid sequence of the present invention is made or identified in a sequence (e.g., an amino acid sequence described herein). The mutations detected are numbered relative to the reference (or wild-type) sequence, i.e., the sequence that does not contain the mutation. Those skilled in the art will recognize mutations in amino acid and nucleic acid sequences relative to a reference sequence. It will be easy to understand how to determine the position.
[0130] The term "nuclear localization sequence," "nuclear localization signal," or "NLS" refers to the localization of a protein. The term "nuclear localization sequence" refers to an amino acid sequence that promotes import into the cell nucleus. For example, WO / 2001 / 038547 filed on November 23, 2000 and published on May 31, 2001 is incorporated herein by reference. and Plank et al. in published international PCT application PCT / EP 2000 / 011690, in which No. 6,239,493, which is incorporated herein by reference for its disclosure of exemplary nuclear localization sequences. In embodiments, the NLS may be any of the NLSs described, for example, in Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4 172. In some embodiments, the NLS is Sequence KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR , RKSGKIAAIVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.
[0131] The terms "nucleobase," "nitrogenous base," or "base" are used interchangeably herein. refers to nitrogen-containing biological compounds that form nucleosides, which are nucleotides The ability of nucleobases to form base pairs and stack with each other is directly related to the synthesis of ribonucleic acid ( Adenine (A) is a nucleotide that gives rise to long helical structures such as RNA and deoxyribonucleic acid (DNA). The five nucleic acid bases are cytosine (C), guanine (G), thymine (T), and uracil (U). They are called primary or standard bases. Adenine and guanine are derived from purines, while cytosine, Uracil and thymine are derived from pyrimidines. DNA and RNA contain modified other (non-primary) bases. Non-limiting exemplary modified nucleobases include hypoxanthine, xanthine, , 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydroxybenzoates. Hypoxanthine and xanthine are produced in the presence of mutagens. Both can be produced by deamination (replacement of an amine group with a carbonyl group). Hypoxanthine can be modified from adenine. Xanthine can be modified from guanine. Uracil can be generated by deamination of cytosine. A "nucleoside" is a nucleic acid base and a five-carbon atom. They consist of a sugar (ribose or deoxyribose). Examples of nucleosides include adenosine , guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, These include deoxyguanosine, thymidine, deoxyuridine, and deoxycytidine. Nucleosides with modified nucleobases include inosine (I), xanthosine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), psoriasis Nucleotides consist of a nucleic acid base, a pentose sugar (ribose or is deoxyribose), and at least one phosphate group.
[0132] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to a nucleic acid molecule comprising a nucleobase and Compounds containing an acidic moiety, such as nucleosides, nucleotides, or polynucleotides Typically, a polymeric nucleic acid, e.g., a nucleic acid molecule containing three or more nucleotides is a linear sequence in which adjacent nucleotides are linked to each other via phosphodiester bonds. In some embodiments, a "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or or nucleosides). In some embodiments, a "nucleic acid" refers to three or more individual nucleosides. As used herein, the term "oligonucleotide" refers to an oligonucleotide chain containing nucleotide residues. "Nucleotide," "polynucleotide," and "polynucleic acid" refer to a polymer of nucleotides. can be used interchangeably to refer to a chain (e.g., a chain of at least three nucleotides). In certain embodiments, "nucleic acid" encompasses RNA and single- and / or double-stranded DNA. Nucleic acids include, for example, genomes, transcripts, mRNA, tRNA, rRNA, siRNA, snRNA, plasmids, cos naturally occurring in the context of a sequence, chromosome, chromatid, or other naturally occurring nucleic acid molecule On the other hand, a nucleic acid molecule can be, for example, a non-naturally occurring molecule, recombinant DNA or RNA, Artificial chromosomes, engineered genomes, or fragments thereof, or synthetic DNA, RNA, DNA / RNA hybrids non-naturally occurring nucleotides or nucleosides, Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or the like may be used interchangeably. The term analogous includes nucleic acid analogs, eg, analogs having other than a phosphodiester backbone. Nucleic acids may be purified from natural sources, produced using recombinant expression systems, and optionally purified. , chemically synthesized, etc. In the case of chemically synthesized molecules, the nucleic acid may be In some cases, for example, analogs with chemically modified bases or sugars and backbone modifications may be used. Nucleic acid sequences may contain nucleoside analogs such as 5'-3' nucleotides unless otherwise indicated. In some embodiments, the nucleic acid is a nucleotide sequence containing natural nucleosides (e.g., adenosine, thiamine, thiamin ... adenosine, cytidine, uridine, deoxyadenosine, deoxythymidine, oxyguanosine, and deoxycytidine; nucleoside analogs (e.g., 2-aminoadenosine, adenosine, 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methyl Cytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodo Uridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2- Aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8- Oxoguanine, O(6)-methylguanine, and 2-thiocytidine; Chemically modified bases; Biology modified bases (e.g., methylated bases); inserted bases; modified sugars (e.g., 2'-fluororibonucleotides); 2'-deoxyribose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages). or containing them.
[0133] The term "nucleic acid programmable DNA binding protein" or "napDNAbp" refers to nucleotide-programmable nucleotide-binding domain" and napDNA A nucleic acid (e.g., DNA or RNA) that guides a nucleotide sequence to a specific nucleic acid sequence, such as a guide nucleic acid. For example, the Cas9 protein binds to a specific DNA sequence complementary to a guide RNA. The sequence can be linked to a guide RNA that guides the Cas9 protein. , napDNAbp is a Cas9 domain, e.g., nuclease-active Cas9, Cas9 nickase (nCas9) , or nuclease-inactive Cas9 (dCas9). In some embodiments, Cas9 The domain comprises any one of the amino acid sequences described herein. In the present invention, the Cas9 domain has at least one amino acid sequence similar to any of the amino acid sequences described herein. at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least At least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% , or at least 99%, or at least 99.5% identity. In some embodiments, the Cas9 domain comprises any of the amino acid sequences described herein. Compared to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 2 0, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 4 Amino acid sequences with 0, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations In some embodiments, the Cas9 domain comprises an amino acid sequence described herein. at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, less At least 90, at least 100, at least 150, at least 200, at least 250, at least 30 0, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 and the amino acid sequence has the same consecutive amino acid residues as the amino acid sequence of the present invention.
[0134] Examples of nucleic acid programmable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9). , Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12 Other nucleic acid programmable DNA fragments include, but are not limited to, Cas12i, Cas13i, and Cas12i. Binding proteins are also within the scope of this disclosure, although they are not specifically described in this disclosure. For example, Makarova et al. "Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct;1:325-336. doi: 10. 1089 / crispr.2018.0033; Yan et al., “Functionally diverse type V CRISPR-Cas syst ems” Science. 2019 Jan 4;363(6422):88-91. doi: 10.1126 / science.aav7271 No. 6,299,499, the entire contents of each of which are incorporated herein by reference.
[0135] The terms "nucleobase editing domain" or "nucleobase editing protein" are used herein. When used in the formula, cytosine (or cytidine) to uracil (or uridine) or hypoxia to thymine (or thymidine) and from adenine (or adenosine) Deamination to sansine (or inosine), and non-templated nucleotide addition and insertion Proteins or enzymes that can catalyze nucleobase modifications in RNA or DNA, such as refers to an enzyme. In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., For example, cytidine deaminase, cytosine deaminase, adenine deaminase, or adenine deaminase. In some embodiments, the nucleobase-editing domain is a naturally occurring In some embodiments, the nucleobase-editing domain may be present in The main focus is on nucleobase editing domains engineered or evolved from naturally occurring nucleobase editing domains. The nucleobase editing domain can be a domain found in bacteria, humans, chimpanzees, gorillas, monkeys, and the like. The nucleic acid may be derived from any organism, such as a human, cow, dog, rat, or mouse. The base-editing proteins are disclosed in International PCT Application No. PCT / 2017 / 045381 (International Publication No. WO 2018 / 027078) and and PCT / US 2016 / 058344 (International Publication No. 2017 / 070632), each of which is , the entire contents of which are incorporated herein by reference. Komor, AC, et al., “Program mable editing of a target base in genomic DNA without double-stranded DNA cleava ge” Nature 533, 420-424 (2016); Gaudelli, NM, et al., “Programmable base edi ting of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); and Komor, A.C., et al., “Improved base excision repair inhibition a nd bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher effic See also “Integrity and product purity” Science Advances 3:eaao4774 (2017) (the entire contents of which are omitted) (The entirety of which is incorporated herein by reference).
[0136] As used herein, "obtaining," as in "obtaining a drug," means This includes synthesizing, purchasing, isolating, or otherwise obtaining a drug. nothing.
[0137] As used herein, a "patient" or "subject" refers to a person who has been diagnosed with a disease or disorder. Mammals at risk of developing them or suspected of having or developing them In certain embodiments, the term "patient" refers to a subject or individual suffering from a disease or disorder. refers to a mammalian subject with a higher than average likelihood of developing a disease. Exemplary patients include humans, non-human primates, cats, dogs, pigs, cattle, horses, camels, llamas, goats, sheep, rodents (e.g. mice, rabbits, rats, guinea pigs) and will benefit from the treatments disclosed herein. Exemplary human patients include males and / or females. could be.
[0138] A "patient in need thereof" or a "subject in need thereof" is used herein to refer to a person with a disease or or disorders, including but not limited to sickle cell disease (SCD) or alpha-1 antitrypsin deficiency Alzheimer's disease (A1AD) or associated with a gene listed in Table 3A, 3B, or 4 herein. Have been diagnosed with, are at risk of having, or have a disease or disorder "Disease" refers to a patient or subject suspected of having one of these diseases or disorders.
[0139] "Pathogenic mutation," "pathogenic variant," "disease-causing (or disease-associated) mutation," "disease" the terms "causative (or disease-associated) variant," "deleterious mutation," or "predisposing mutation" is a genetic change or mutation that increases an individual's susceptibility or predisposition to a particular disease or disorder. refers to a mutation. In some embodiments, a pathogenic mutation is a mutation in a protein encoded by a gene. At least one wild-type amino acid in the protein is replaced by at least one pathogenic amino acid. This includes those that have been converted.
[0140] The term "non-conservative mutation" refers to amino acid substitutions between different groups, e.g. , tryptophan to lysine, or serine to phenylalanine, etc. Non-conservative amino acid substitutions may be made without disrupting or inhibiting the biological activity of the functional variant. Non-conservative amino acid substitutions are preferred because they ensure that the biological activity of the functional variant is comparable to that of the wild-type protein. The biological activity of the functional variant can be enhanced so that it is increased compared to the protein. Cut.
[0141] The terms "protein," "peptide," "polypeptide," and their grammatical equivalents are used interchangeably herein and are linked to each other by a peptide (amide) bond. The term refers to a polymer of amino acid residues bound together in a single chain. It typically refers to a protein, peptide, or polypeptide. A protein, peptide, or polypeptide is at least three amino acids in length. A polypeptide can refer to an individual protein or a group of proteins. One or more amino acids in a protein, peptide, or polypeptide may have, for example, a carbohydrate group, a hydrophobic group, xyl group, phosphate group, farnesyl group, isofarnesyl group, fatty acid group, linkage They may be modified by the addition of chemical entities such as anchors, functionalization, or other modifications. Proteins, peptides, or polypeptides may also be single molecules or multi-molecular. The protein, peptide, or polypeptide may be a naturally occurring It may be merely a fragment of a protein or peptide. Peptides may be naturally occurring, recombinant, or synthetic, or any of the foregoing. As used herein, the term "fusion protein" refers to a combination of: Hybrid polypeptides containing protein domains from at least two different proteins One protein is the amino-terminal (N-terminal) portion of the fusion protein or the capsid. can be located at the carboxy-terminus (C-terminus) of a protein, thus The protein forms a terminal or carboxy-terminal fusion protein. a nucleic acid binding domain (e.g., a nucleic acid binding domain that induces binding of a protein to a target site) Cas9 gRNA binding domain and nucleic acid cleavage domain, or nucleic acid editing protein In some embodiments, the protein may comprise a proteinaceous moiety, e.g., a catalytic domain of For example, an amino acid sequence constituting a nucleic acid binding domain and an organic compound, such as a nucleic acid cleaving agent In some embodiments, proteins include compounds that can act as nucleic acids (e.g., RNA). or DNA) are complexed with or associated with nucleic acids. Any protein can be produced by any method known in the art. For example, the proteins provided herein can be prepared via recombinant protein expression and purification. This is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and are described in Green and Samb rook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Labora and those described by the University of California Press, Cold Spring Harbor, NY (2012). the entire contents of which are incorporated herein by reference.
[0142] The polypeptides and proteins disclosed herein (including functional portions thereof and their functional variants) (including riboflavin) contains synthetic amino acids in place of one or more naturally occurring amino acids Such synthetic amino acids are known in the art and include, for example, aminocyclo Hexanecarboxylic acid, norleucine, α-amino n-decanoic acid, homoserine, S-acetylacetone aminomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenyl phenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine phenylalanine, β-phenylserine, β-hydroxyphenylalanine, phenylglycine, α-Naphthylalanine, cyclohexylalanine, cyclohexylglycine, indoline 2-carboxylic acid, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, Aminomalonic acid monoamide, N'-benzyl-N'-methyllysine, N',N'-dibenzyl-lysine, 6- Hydroxylysine, ornithine, α-aminocyclopentanecarboxylic acid, aminocyclohexyl cyclohexanecarboxylic acid, aminocyclohexanecarboxylic acid, α-aminocycloheptanecarboxylic acid carboxylic acid, α-(2-amino-2-norbornane)-carboxylic acid, α,γ-diaminobutyric acid, α,β-diamino These include monopropionic acid, homophenylalanine, and α-tert-butylglycine. Polypeptides and proteins are derived from post-translational modifications of one or more amino acids of the polypeptide construct. Non-limiting examples of post-translational modifications include phosphorylation, acetylation, and the like. and acylation, including formylation, glycosylation (including N-linked and O-linked), amino Derivatization, hydroxylation, alkylation including methylation and ethylation, ubiquitination, pyrolysis, Addition of lidocaine carboxylic acid, formation of disulfide bridges, sulfation, myristoylation, palmitoylation ylation, isoprenylation, farnesylation, geranylation, glypiation, lipoylation and iodination.
[0143] As used herein, the term "gene" typically refers to a protein-coding region. A polynucleotide that includes a non-coding region and a non-protein-coding region. A region may contain one or more regulatory elements. Non-limiting examples of regulatory elements include promoters. ter, enhancer, repressor, silencer, insulator, start codon, stop Codon, Kozak consensus sequence, splice acceptor, splice donor, 3' and / or comprises a 5' untranslated region (UTR), a slice site, or an intergenic region. In genetic engineering, regulatory elements are located in genes that are responsible for genetic diseases or disorders. Non-limiting examples of regulatory elements located in genes that are responsible for a genetic disease or disorder include: , start codon, stop codon, Kozak consensus sequence, intergenic region, 3'UTR, or 5' In some embodiments, the regulatory element is a regulatory element for a genetic disease or Not located in the gene that causes the disorder. Not located in the gene that causes the genetic disease. Non-limiting examples of elements include enhancers, repressors, or insulators. - and others.
[0144] The term "polynucleotide programmable nucleotide binding domain" refers to a polynucleotide A guide polynucleotide that guides a programmable DNA binding domain to a specific nucleic acid sequence A protein attached to a nucleic acid (such as DNA or RNA) such as a nucleic acid protease (e.g., guide RNA). In some embodiments, the polynucleotide-programmable nucleotide binding domain is In one embodiment, the polynucleotide is a programmable DNA binding domain. The polynucleotide programmable nucleotide binding domain In one embodiment, the polynucleotide programmable nucleic acid is a RNA-binding domain. The nucleotide binding domain is the Cas9 protein. The Cas9 protein binds to the guide RNA. It can bind to a guide RNA that guides the Cas9 protein to a specific complementary DNA sequence. In some embodiments, the polynucleotide programmable nucleotide binding domain is s9 domain, e.g., nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease Non-limiting examples of nucleic acid programmable DNA binding proteins include enzyme-inactive Cas9 (dCas9) and Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, and Cas12d / Ca Non-limiting examples of Cas enzymes include sY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i. are Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7 , Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, C as12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Csc1, Csc1, Csm2, Csm3, Csm3, Csm4, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX , Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1 , Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Type II Cas effector proteins, type V Cas effector proteins, type VI Cas effector proteins, CARF, DinG, and Other nucleic acid fragments include homologs of these fragments, or modified or altered versions thereof. Programmable DNA binding proteins are also within the scope of this disclosure, but are not specifically included in this disclosure. It may not be listed.
[0145] The term "recombinant" as used herein with respect to a protein or nucleic acid means that the protein or nucleic acid is It refers to proteins or nucleic acids that do not occur in nature but are the product of human engineering. For example, In some embodiments, the recombinant protein or nucleic acid molecule is any naturally occurring At least one, at least two, at least three, at least four, or at least at least five, at least six, or at least seven amino acid or nucleotide mutations Contains an octide sequence.
[0146] "Decrease" means a negative change of at least 10%, 25%, 50%, 75%, or 100% .
[0147] "Reference" means a standard or control condition. base edits, such as benign or regulatory base edits, Assays of the activity or function of the protein product (or protein products encoded by the protein) can be used to identify benign or regulatory The activity or function of the gene (and / or its encoded product) that did not undergo base editing , or the activity of the wild-type gene (and / or its encoded product) as a reference, or In one embodiment, the reference is a wild-type or healthy cell.
[0148] A "reference sequence" is a defined sequence used as a basis for sequence comparison. It can be a subset or the entirety of a particular sequence; for example, a full-length cDNA or gene sequence. A segment of a gene, or the complete cDNA or gene sequence. For polypeptides, see the reference polypeptide. The length of the peptide sequence is generally at least about 16 amino acids, preferably at least about 20 amino acids. amino acids, more preferably at least about 25 amino acids, even more preferably about 35 amino acids; For nucleic acids, the length of a reference nucleic acid sequence is about 50 amino acids, or about 100 amino acids. Generally, at least about 50 nucleotides, preferably at least about 60 nucleotides, more preferably Preferably, at least about 75 nucleotides, even more preferably at least about 100 nucleotides or It is about 300 nucleotides or any integer thereabout or therebetween.
[0149] The terms "RNA programmable nuclease" and "RNA-guided nuclease" refer to Used in conjunction with (e.g., bound to or associated with) one or more non-target RNAs In some embodiments, the RNA-programmable nuclease, when complexed with RNA, Typically, the bound RNA is called a guide RNA (gRNA). Guide RNA (gRNA) can exist as a complex of two or more RNAs, or as a single RNA. A gRNA can exist as a single RNA molecule. Although sometimes referred to as sgRNA, "gRNA" can be used as a single molecule or as two or more molecules. These are used interchangeably to refer to either a guide RNA that exists as a complex with a In the present study, gRNAs exist as a single RNA species, consisting of two domains: (1) a domain that shares homology with the target nucleic acid; (2) a domain that contains a Cas9 protein (e.g., that directs binding of the Cas9 complex to a target); and In one embodiment, domain (2) comprises a domain that binds to a protein known as tracrRNA. For example, in some embodiments, , domain (2) is tracrRNA provided by Jinek et al., Science 337:816-821 (2012) The gRNA (e.g., a domain Another example of a method for detecting nucleotides containing ... is "Switchable Cas9 Nucleases" filed on September 6, 2013. and U.S. Provisional Patent Application No. 61 / 874,682, issued September 6, 2013, entitled "Polymer-Based Polymeric Materials and Uses Thereof." The application is filed in U.S. Provisional Patent Application No. 2006 / 0100604 entitled "Delivery System for Functional Nucleases." No. 61 / 874,746. In some embodiments, the gRNA comprises two or more domains. The extended gRNA may comprise sequences (1) and (2) and may be referred to as an "extended gRNA." For example, The engineered gRNA binds to two or more Cas9 proteins, as described herein, and The gRNA binds to the target nucleic acid at these different regions. complements the target site, mediating binding and providing sequence specificity for the nuclease:RNA complex In some embodiments, the RNA programmable nuclease comprises a nucleotide sequence (CRIS PR-related systems) Cas9 endonuclease, e.g., Cas9 from Streptococcus pyogenes (Csnl) (e.g., "Complete genome sequence of an Ml strain of Streptococcus Ferretti JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y. , Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW , Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001); CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Del tcheva E., Chylinski K., Sharma CM., Gonzales K., Chao Y., Pirzada ZA, Eckert See MR, Vogel J., Charpentier E., Nature 471:602-607(2011).
[0150] The term "single nucleotide polymorphism (SNP)" refers to a single nucleotide variation that occurs at a specific position in the genome. where each mutation is present in the population to a noticeable extent (e.g., >1%). For example, At certain base positions in the human genome, C nucleotides can occur in most individuals, In a small number of individuals, the position is occupied by A. This means that there is a SNP at this particular position, and C or A, meaning that two nucleotide variations are alleles at this position. SNPs underlie differences in susceptibility to disease, affecting the severity of the disease and the body's response to treatment. SNPs are found in the coding region of genes, non-coding regions of genes, and so on. It may be present in a gene region, or in an intergenic region (the region between genes). In this case, SNPs within the coding sequence may affect the identity of the protein produced due to the degeneracy of the genetic code. SNPs in the coding region are of two types: synonymous SNPs and non-synonymous SNPs. Synonymous SNPs do not affect the protein sequence, whereas non-synonymous SNPs affect the amino acid sequence of a protein. There are two types of nonsynonymous SNPs: missense and nonsense. SNPs not in the coding region affect gene splicing, transcription factor binding, and messenger receptor (MR) activity. These types of SNPs can affect the degradation of RNA or the sequence of non-coding RNA. The gene expression that is affected is called an eSNP (expressed SNP) and can be upstream or downstream of the gene. Single nucleotide variants (SNVs) are single nucleotide variations with unlimited frequency and occur in somatic Somatic single base variations (e.g., caused by cancer) can occur at a single base It may also be called modification.
[0151] "Specifically bind" means to recognize and bind to the polypeptide and / or nucleic acid molecule of the present invention. and bind to other molecules in the sample (e.g., biological sample) but do not substantially recognize or bind to other molecules in the sample. Nucleic acid molecules, polypeptides, or complexes thereof (e.g., nucleic acid programmable DNA) A binding domain and guide nucleic acid), compound, or molecule.
[0152] Nucleic acid molecules useful in the methods of the present invention may encode a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules include any nucleic acid molecule that is 100% identical to the endogenous nucleic acid sequence. Typically, but not necessarily, substantial identity is shown. A polynucleotide having a nucleotide sequence typically hybridizes with at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention can be modified to The present invention also includes any nucleic acid molecule encoding an endogenous nucleotide sequence or a fragment thereof. It need not be 100% identical to the nucleic acid sequence, but typically will show substantial identity. Polynucleotides having "substantial identity" to a double-stranded nucleic acid molecule typically have a small number of double-stranded nucleic acid molecules. "Hybridize" means to hybridize with at least one of the strands of a given molecule. Complementary polynucleotide sequences (e.g., those described herein) are synthesized under various stringency conditions. This means that a double-stranded molecule is formed between the two genes (or a part of them). For example, Wahl, GM and SL Berger (1987) Methods Enzymol. 152:399; Kimmel, A. R (1987) Methods Enzymol. 152:507).
[0153] For example, a stringent salt concentration is typically less than about 750 mM NaCl and 75 mM citrate. trisodium, preferably less than about 500 mM NaCl and 50 mM trisodium citrate, more preferably Preferably, the concentration is less than about 250 mM NaCl and 25 mM trisodium citrate. See hybridization can be obtained in the absence of organic solvents, e.g., formamide. whereas high stringency hybridization requires at least about 35% formaldehyde. It can be obtained in the presence of at least about 50% formamide. Stringent temperature conditions are typically at least about 30°C, more preferably at least about 37°C. 0 C, and most preferably at least about 42°C. The concentration of detergent (e.g., sodium dodecyl sulfate (SDS)) and carrier DNA content Various additional parameters, such as inclusion or exclusion, are well known to those skilled in the art. By combining these various conditions, various levels of stringency can be achieved. In one embodiment, hybridization is achieved at 30° C. in 750 mM NaCl, 7 In another embodiment, the hybridization occurs in 5 mM trisodium citrate and 1% SDS. The solution was incubated at 37°C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formaldehyde. In another embodiment, the DNA fragment is generated in 100 μg / ml denatured salmon sperm DNA (ssDNA). Hybridization was performed at 42°C in 250 mM NaCl, 25 mM trisodium citrate, 1% The reaction takes place in SDS, 50% formamide, and 200 μg / ml ssDNA. The implementation will be readily apparent to one skilled in the art.
[0154] For most applications, the washing steps that follow hybridization are also stringent. Wash stringency conditions are defined by salt concentration and temperature. As mentioned above, wash stringency can be increased by decreasing the salt concentration or increasing the temperature. This can be increased by increasing the thickness of the strip for the cleaning process. The appropriate salt concentration is less than about 30 mM NaCl and 3 mM trisodium citrate, and the appropriate concentration is about 15 mM NaCl. 1.5 mM NaCl and 1.5 mM trisodium citrate. Gentle temperature conditions are typically at least about 25°C, more preferably at least about 42°C, and even more preferably at least about 50°C. More preferably, the cleaning process includes a temperature of at least about 68° C. The procedure is carried out at 25° C. in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the wash step is performed at 42° C. in 15 mM NaCl, 1.5 mM trisodium citrate. In a more preferred embodiment, the washing step is carried out in 0.1% SDS and 0.1% thorium. The reaction is carried out at 68°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Further variations of these conditions will be readily apparent to those skilled in the art. Redimination techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science nce 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience e, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 19 87, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory This is described in the Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.
[0155] "Substantially identical" means that the amino acid sequence of a reference amino acid sequence (e.g., an amino acid sequence described herein) is substantially identical to the amino acid sequence of a reference amino acid sequence (e.g., an amino acid sequence described herein). any one of the sequences) or nucleic acid sequence (e.g., any one of the nucleic acid sequences described herein) It means a polypeptide or nucleic acid molecule that shows at least 50% identity to Or, such sequences may be identical in amino acid or nucleic acid sequence to the sequences used for comparison. At least 60%, more preferably 80% or 85%, more preferably 90%, 95% or 99% But it has the same identity.
[0156] Sequence identity is typically determined using sequence analysis software (e.g., Genetics Computer Group , University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705 Sequence Analysis Software Package, BLAST, BESTFIT, COBALT, EMBOSS The measurement is performed using software such as Needle, GAP, or PILEUP / PRETTYBOX programs. The software can assign degrees of homology to various substitutions, deletions, and / or other modifications. Identical or similar sequences are matched by: Includes substitutions within the group: glycine, alanine; valine, isoleucine, leucine; asparagine Acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine phenylalanine, tyrosine. Exemplary Approaches for Determining the Degree of Identity The BLAST program can be used in -3 and e -100 The probability score between COBALT is used, for example, with the following parameters: a) Alignment parameters: Gap penalties -11, -1 and End-Gap penalties -5, -1 b) CDD parameters: Use RPS BLAST on; Blast E-value 0.003; Find Conserved column s and Recompute on c) Query clustering parameters: Use query clusters on; Word Size 4; M ax cluster distance 0.8; Alphabet Regular. The EMBOSS Needle is used, for example, with the following parameters: a) Matrix: BLOSUM62; b) GAP OPEN: 10; c) GAP EXTEND: 0.5; d) OUTPUT FORMAT: pair; e) END GAP PENALTY: false; f) END GAP OPEN: 10; and g) END GAP EXTEND: 0.5.
[0157] A "subject" includes, but is not limited to, a cow, horse, dog, sheep, or cat. "A mammal" means a mammal, including a human or non-human mammal.
[0158] The term "target site" refers to a sequence within a nucleic acid molecule that is modified by a nucleobase editor. In one embodiment, the target site is a sequence that is targeted by a deaminase or a fusion enzyme containing the same. Deaminated by a ligase protein (e.g., cytidine or adenine deaminase) .
[0159] RNA-programmable nucleases (e.g., Cas9) use RNA:D to target DNA cleavage sites. Since NA hybridization is used, these proteins can, in principle, be used as guide RNAs. Any sequence specified by A can be targeted. For site-specific cleavage, C Methods using RNA-programmable nucleases such as as9 (e.g., to modify genomes) For this purpose, methods for generating multiplexed genomic DNA (e.g., Cong, L. et al., Multiplex gene expression) are known in the art (see, e.g., Cong, L. et al., Multiplex gene expression). ome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013 ); Hwang, WY et al., Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programme d genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., G enome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic acids research (2013); Jiang, W. et al., RNA-guided editing of bacterial genomes using CRISPR-Cas systems. See Nature biotechnology 31, 233-239 (2013). the entire contents of each of which are incorporated herein by reference).
[0160] As used herein, the terms "treat," "treating," and "treatment" refer to "Treatment" and the like also provide relief from a disease or disorder and / or its associated symptoms. to improve or improve the quality of life or to obtain a desired pharmacological and / or physiological effect Treating a disorder or condition means completely eliminating the associated disorder, condition or symptoms. It will be understood that this does not require complete removal (and complete removal is not excluded). In some embodiments, the effect is therapeutic, i.e., is not limited to However, the effect may be partial relief of the disease or disorder and / or adverse symptoms resulting from the disease or disorder. or completely reduces, diminishes, eliminates, alleviates, mitigates, lessens in intensity, or cures. In some cases, the effect is prophylactic, i.e., the effect is to prevent the occurrence or recurrence of a disease, disorder, or condition. To this end, the methods of the present disclosure include the use of a method for protecting against or preventing the development of a leukemia, as described herein. The method includes administering a therapeutically effective amount of the composition.
[0161] "Uracil glycosylase inhibitor" means a factor that inhibits the uracil excision repair system. In one embodiment, the factor binds to the host uracil-DNA glycosylase to form D It is a protein or fragment thereof that prevents the removal of uracil residues from NA.
[0162] Ranges provided herein include all values within the range, inclusive of the first and last values, and It is understood that the range is an abbreviation for the values between them. For example, the range 1 to 50 is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, Any number, combination of numbers, or sub-category from the group consisting of 46, 47, 48, 49, or 50 It is understood to include the area.
[0163] The recitation of a list of chemical groups in any definition of a variable herein means that the listed group The variables herein include definitions of the variables as any single group or combination of The description of an embodiment or aspect may be used interchangeably with any single embodiment or with any other embodiment. This includes any embodiment in combination with any embodiment or portion thereof.
[0164] Any composition or method provided herein may be used in combination with any other composition or method provided herein. The present invention may be combined with one or more of the compositions and methods.
[0165] DNA editing is a promising approach to modify disease states by correcting pathogenic mutations at the genetic level. Until recently, all DNA editing platforms have focused on specific genomes. It functions by inducing DNA double-strand breaks (DSBs) at genomic sites and is dependent on endogenous DNA repair pathways. The product outcome was determined in a semi-probabilistic manner, resulting in a complex population of genetic products. Achieve precise, user-defined repair results via homogeneous directed repair (HDR) pathway However, many challenges hinder efficient repair using HDR in therapeutically relevant cell types. In practice, this pathway is more efficient than the competing, error-prone non-homologous end-joining pathway. Furthermore, HDR is strictly restricted to the G1 and S phases of the cell cycle and does not occur after mitosis. This prevents accurate repair of DSBs in these cells, resulting in highly efficient repair of DSBs in these populations. , making it difficult or impossible to modify genome sequences in a user-defined, programmable manner. It has been revealed that this is the case.
[0166] [Nucleobase Editor] A base editor for editing, modifying or altering a target nucleotide sequence of a polynucleotide. Disclosed herein are nucleic acid base editors or nucleobase editors. Polynucleotide-programmable nucleotide-binding domains and nucleobase-editing domains A polynucleotide program is a nucleic acid base editor or base editor that includes Possible nucleotide binding domains are those that bind to the guide polynucleotide (e.g., gRNA). When the bases of the bound guide nucleic acid and the target polynucleotide sequence are specifically binds to a target polynucleotide sequence (through complementary base pairing between thereby localizing the base editor to the target nucleic acid sequence desired to be edited. In some embodiments, the target polynucleotide sequence can be single-stranded DNA or In some embodiments, the target polynucleotide sequence comprises double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In an embodiment, the target polynucleotide sequence comprises a DNA-RNA hybrid.
[0167] [Polynucleotide-programmable nucleotide-binding domain] The terms "polynucleotide programmable nucleotide binding domain" or "nucleic acid programmable nucleotide binding domain" are used interchangeably. A "programmable DNA-binding protein" refers to a protein that binds a guide polynucleotide (e.g., a guide RNA) to a target protein. Proteins that associate with nucleic acids (e.g., DNA or RNA) such as polynucleotides. Guiding a programmable nucleotide-binding domain to a specific nucleic acid sequence In some embodiments, polynucleotide programmable nucleotide binding The domain is a polynucleotide-programmable DNA binding domain. In this case, the polynucleotide-programmable nucleotide binding domain is In one embodiment, the polynucleotide protease is a programmable RNA binding domain. In some embodiments, the gram-capable nucleotide binding domain is a Cas9 protein. Polynucleotide-programmable nucleotide-binding domain is the Cpf1 protein .
[0168] CRISPR provides defense against mobile genetic elements (viruses, transposable elements, conjugative plasmids) CRISPR clusters are composed of spacers, sequences that are complementary to the preceding mobile element. The CRISPR cluster contains a sequence of target nucleic acids, and the target invader nucleic acid. The CRISPR cluster is transcribed into CRISPR RNA (crRNA). In type II CRISPR systems, correct processing of pre-crRNA is essential for transcription. transcoding small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein tracrRNA guides the processing of pre-crRNA by ribonuclease 3. Cas9 / crRNA / tracrRNA then binds to a linear or circular dsDNA target complementary to the spacer. The target strand that is not complementary to the crRNA is first cleaved by the endonuclease. It is cleaved protease-wise and then exonucleolytically trimmed to 3'-5'. In the case of mitochondrial DNA, both the protein and the RNA are required for DNA binding and cleavage. A single guide RNA ("sgRNA", or simply "gRNA") can be produced. See, for example, Jinek M., Chylinski K., Fonfa See ra I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012) Cas9 is a CRISPR repeat sequence. It recognizes a short motif (PAM or protospacer adjacent motif) in the sequence and identifies it as "self". Helps distinguish "non-self."
[0169] [Cas9 domain, a nucleobase editor] The sequence and structure of Cas9 nuclease are well known to those skilled in the art (see, e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes.” Ferretti et al., J .J., McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS., Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L.,. Natl. Acad. Sci. USA 98:4658-4663(2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma CM., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentier E., Nature 471:602-607(2011); and “A programmable dual -RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylin ski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(20 12), the entire contents of which are incorporated herein by reference. Cas9 orthologs include: It has been described in a variety of species, including but not limited to S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be readily apparent to those skilled in the art based on the present disclosure. Such Cas9 nucleases and sequences are described in Chylinski, Rhun, and C harpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity syst Ca from organisms and loci disclosed in "ems" (2013) RNA Biology 10:5, 726-737 s9 sequence, the entire contents of which are incorporated herein by reference.
[0170] In some embodiments, the Cas9 nuclease comprises an inactive (e.g., inactivated) DNA cleavage domain. Cas9 has a specific structure, i.e., it is called the "nCas9" protein (for "nickase" Cas9). The nuclease-inactivated Cas9 protein is interchangeably referred to as the "dCas9" protein. It can also be called protein (meaning nuclease-dead Cas9). Methods for producing the as9 protein (or fragments thereof) are known (see, e.g., Jinek et al., Science 337:816-821(2012); Qi et al, “Repurposing CRISPR as an RNA-Guided Platform fo r Sequence-Specific Control of Gene Expression” (2013) Cell. 28; 152(5): 1173-8 3, the contents of each of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 contains two subdomains: the HNH nuclease subdomain and the RuvC1 subdomain It is known that the HNH subdomain cleaves the strand complementary to the gRNA, and the RuvCl subdomain Mutations within these subdomains abolish the nuclease activity of Cas9. For example, mutations D10A and H840A can inhibit the nuclease activity of S. pyogenes Cas9. completely inactivates the ATP-dependent ATPase (Jinek et al., Science. 337:816-821(2012); Qi et al., Cell. 28;152(5): 1173-83 (2013)). In some embodiments, proteins comprising fragments of Cas9 are provided. For example, in some embodiments, the protein comprises two Cas9 domains: (1) the gRNA-binding domain of Cas9; (2) the DNA cleavage domain of Cas9. In embodiments, proteins comprising Cas9 or fragments thereof are referred to as "Cas9 variants." Cas9 variants share homology with Cas9 or fragments thereof. For example, Cas9 variants is at least about 70% identical, at least about 80% identical, at least about 90% identical to wild-type Cas9, At least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical 1. At least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical In some embodiments, the Cas9 mutant has 1, 2, 3, 4, or more amino acid sequences compared to wild-type Cas9. 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 2 6, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 4 It may have 6, 47, 48, 49, 50 or more amino acid changes. Cas9 variants are fragments of Cas9 (e.g., gRNA binding domain or DNA cleavage domain). wherein the fragment is at least about 70% identical to the corresponding fragment of wild-type Cas9 and At least about 80% identical, at least about 90% identical, at least about 95% identical, at least at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical In some embodiments, the fragment is at least 30 amino acids in length of the corresponding wild-type Cas9. %, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, Both are 98%, at least 99%, or at least 99.5%.
[0171] In some embodiments, the fragment is at least 100 amino acids in length. , fragments must be at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least is also 1300 amino acids long.
[0172] In some embodiments, the wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCB Reference sequence: NC_17053.1, with the following nucleotide and amino acid sequence: ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGATTATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG GCAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGCAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAATCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTAGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAGAAATGGCTTGTTTGGGAAT CTCATTGCTTTGTCATTGGGATTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATAGTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAGCGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAGGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGCGCCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGGGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGATATTCAAAAAGCACAGGTGTCTGGACAAGGCCATAGTTTA CATGAACAGATTGCTAACTTAGCTGGCAGTCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAATTGTTGATGAACT GGTCAAAGTAATGGGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCCAGA AAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCTGTT GAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTACAAAATGGAAGAGACATGTATGTGGACCAAGAATT AGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCATTAAAGACGATTCAATAGACAATA AGGTACTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAAAAC TATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGGTTT GAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAATTT TGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCTAAA TTAGTTTCTGACTTCCGAAAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGCGTA TCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATAAAG TTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCTAAT ATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGGGGA AACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATATTG TCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTACCAAAAAGAAATTCGGACAAGCTTATTGCT CGTAAAAAGACTGGGATCCAAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGCTAA GGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCTTTG AAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATAAGGAAGTTAAAAGACTTAATCATTAAACTACCTAATAT AGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAATTACAAAAAGGAAATGAGCTGGCTCT GCCAAGCAAATGTGAATTTTTTATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAACAAA AACAATTGTTTGTGGAGCAGCATAAGCATTATTTGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTTATT TTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGAAAA TATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTAAAC GATATACGTCTACAAAAGAAGTTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATTGAT TTGAGTCAGCTAGGAGGTGACTGA. JPEG2025026840000003.jpg164170 (single underline: HNH domain; double underline: RuvC domain).
[0173] In some embodiments, wild-type Cas9 contains the following nucleotides and / or amino acids: corresponding to or including the amino acid sequence: ATGGATAAAAAGTATTCTATTGGTTTAGACATCGGCACTAATTCCGTTGGATGGGCTGTCATAACCGATGAATACAAAGT ACCTTCAAAGAAATTTAAGGTGTTGGGGAACACAGACCGTCATTCGATTAAAAAAGAATCTTATCGGTGCCCTCCTATTCG ATAGTGGCGAAACGGCAGAGGCGACTCGCCTGAAACGAACCGCTCGGAGAAGGTATACACGTCGCAAGAACCGAATATGT TACTTACAAGAAATTTTTAGCAATGAGATGGCCAAAGTTGACGATTCTTTCTTTCACCGTTTGGAAGAGTCCTTCCTTGT CGAAGAGGACAAGAAACATGAACGGCACCCCATCTTTGGAAACATAGTAGATGAGGTGGCATATCATGAAAAGTACCCAA CGATTTATCACCTCAGAAAAAAGCTAGTTGACTCAACTGATAAAGCGGACCTGAGGTTAATCTACTTGGCTCTTGCCCAT ATGATAAAGTTCCGTGGGCACTTTCTCATTGAGGGTGATCTAAATCCGGACAACTCGGATGTCGACAAACTGTTCATCCA GTTAGTACAAACCTATAATCAGTTGTTTGAAGAACCCTATAAATGCAAGTGGCGTGGATGCGAAGGCTATTCTTAGCG CCCGCCTCTCTAAATCCCGACGGCTAGAAAAACCTGATCGCACAATTACCCGGAGAGAAAAAAATGGTTGTTCGGTAAC CTTATAGCGCTCTCACTAGGCCTGACACCAAATTTAAGTCGAACTTCGACTTAGCTGAAGATGCCAAATTGCAGCTTAG TAAGGACACGTACGATGACGATCTCGCAATCTACTGGCACAAATTGGAGATCAGTATGCGGACTTATTTTTGGCTGCCA AAAACCTTAGCGATGCAATCCTCCTATCTGACATACTGAGAGTTAATACTGAGATTACCAAGGCGCCGTTATCCCGCTTCA ATGATCAAAAGGTACGATGAACATCACCAAGACTTGACACTTCTCAAGGCCCTAGTCCGTCAGCAACTGCCTGAGAAATA TAAGGAAATATTCTTTGATCAGTCGAAAACGGGTACGCAGGTTATATTGACGGGCGGAGCGAGTCAAGAGGAATTCTACA AGTTTATCAAACCCATATTAGAGAATGGATGGGACGGAAGAGTTGCTTGTAAAACTCAATCGCGAAGATCTACTGCGA AAGCAGCGGACTTTCGACAACGGTAGCATTCCACATCAAATCCACTTAGGCGAATTGCATGCTATACTTAGAAGGCAGGA GGATTTTTATCCGTTCCTCAAAGACAATCGTGAAAAGATTGAGAAAATCCTAACCTTTCGCATACCTACTATGTGGGAC CCCTGGCCCGAGGGAACTCTCGGTTCGCATGGATGACAAGAAAGTCCGAAGAAACGATTACTCCATGGAATTTTGAGGAA GTTGTCGATAAAGGTGCGTCAGCTCAATCGTTCATCGAGAGGATGACCAACTTTGACAAGAATTTACCGAACGAAAAAGT ATTGCCTAAGCACAGTTTACTTTACGAGTATTTCACAGTGTACAATGAACTCACGAAAGTTAAGTATGTCACTGAGGGCA TGCGTAAACCCGCCTTTCTAAGCGGAGAACAGAAGAAAGCAATAGTAGATCTGTTATTCAAGACCAACCGCAAAGTGACA GTTAAGCAATTGAAAGAGGACTACTTTAAGAAAATTGAATGCTTCGATTCTGTCGAGATCTCCGGGGTAGAAGATCGATT TAATGCGTCACTTGGTACGTATCATGACCTCCTAAAGATAATTAAAGATAAGGACTTCCTGGATAACGAAGAGAATGAAG ATATCTTAGAAGATATAGTGTTGACTCTTACCCTCTTTGAAGATCGGGAAATGATTGAGGAAAGACTAAAAACATACGCT CACCTGTTCGACGATAAGGTTATGAAACAGTTAAAGAGGCGTCGCTATACGGGCTGGGGACGATTGTCGCGGAAACTTAT CAACGGGATAAAGAGACAAGCAAAGTGGTAAAACTATTCTCGATTTTCTAAAGAGCGACGGCTTCGCCATAGGAACTTTA TGCAGCTGATCCATGATGACTCTTTAACCTTCAAAGAGGATATACAAAAGGCACAGGTTTCCGGACAAGGGGACTCATTG CACGAACATATTGCGAATCTTGCTGGTTCGCCAGCCATCAAAAAGGGCATACTCCAGACAGTCAAAGTAGTGGATGAGCT AGTTAAGGTCATGGGACGTCACAAACCGGAAAACATTGTAATCGAGATGGCACGCGAAAATCAAACGACTCAGAAGGGGC AAAAAAACAGTCGAGAGCGGATGAAGAGAATAGAAGAGGGTATTAAAGAACTGGGCAGCCAGATCTTAAAGGAGCATCCT GTGGAAAATACCCAATTGCAGAACGAGAAACTTTACCTCTATTACCTACAAAATGGAAGGGACATGTATGTTGATCAGGA ACTGGACATAAACCGTTTATCTGATTACGACGTCGATCACATTGTACCCCAATCCTTTTTGAAGGACGATTCAATCGACA ATAAAGTGCTTACACGCTCGGATAAGAACCGAGGGAAAAGTGACAATGTTCCAAGCGAGGAAGTCGTAAAGAAAATGAAG AACTATTGGCGGCAGCTCCTAAATGCGAAACTGATAACGCAAAGAAAGTTCGATAACTTAACTAAAGCTGAGAGGGGTGG CTTGTCTGAACTTGACAAGGCCGGATTTATTAAACGTCAGCTCGTGGAAACCCGCCAAATCACAAAGCATGTTGCACAGA TACTAGATTCCCGAATGAATACGAAATACGACGAGAACGATAAGCTGATTCGGGAAGTCAAAGTAATCACTTTAAAGTCA AAATTGGTGTCGGACTTCAGAAAGGATTTTCAATTCTATAAAGTTAGGGAGATAAATAACTACCACCATGCGCACGACGC TTATCTTAATGCCGTCGTAGGGACCGCACTCATTAAGAAATACCCGAAGCTAGAAAGTGAGTTTGTGTATGGTGATTACA AAGTTTATGACGTCCGTAAGATGATCGCGAAAGCGAACAGGAGATAGGCAAGGCTACAGCCAAATACTTCTTTTATTCT AACATTATGAATTTCTTTAAGACGGAAATCACTCTGGCAAACGGAGAGATACGCAAACGACCTTTAATTGAAACCAATGG GGAGACAGGTGAAATCGTATGGGATAAGGGCCGGGACTTCGCGACGGTGAGAAAAGTTTTGTCCATGCCCCAAGTCAACA TAGTAAAGAAAACTGAGGTGCAGACCGGAGGGTTTTCAAAGGAATCGATTCTTCCAAAAAGGAATAGTGATAAGCTCATC GCTCGTAAAAAGGACTGGGACCCGAAAAAGTACGGTGGCTTCGATAGCCCTACAGTTGCCTATTCTGTCCTAGTAGTGGC AAAAGTTGAGAAGGGAAATCCAAGAAACTGAAGTCAGTCAAAGAATTATTGGGGATAACGATTATGGAGCGCTCGTCTT TTGAAAAGACCCCATCGACTTCCTTGAGGCGAAAGGTTACAAGGAAGTAAAAAAGGATCTCATAATTAAACTACCAAAG TATAGTCTGTTTGAGTTAGAAATGGCCGAAAACGGATGTTGGCTAGCGCCGGAGAGCTTCCAAAGGGGAACGAACTCGC ACTACCGTCTAAATACGTGAATTTCCTGTATTTAGCGTCCCATTACGAGAAGTTGAAAGGTTCACCTGAAGATAACGAAC AGAAGCAACTTTTTGTTGAGCAGCACAAACATTATCTCGACGAAATCATAGAGCAAATTTCGGAATTCAGTAAGAGAGTC ATCCTAGCTGATGCCAATCTGGACAAAGTATTAAGCGCATACAACAAGCACAGGGATAAACCCATACGTGAGCAGGCGGA AAATATTATCCATTTGTTTACTCTTACCAACCTCGGCGCTCCAGCCGCATTCAAGTATTTTGACACAACGATAGATCGCA AACGATACACTTCTACCAAGGAGGTGCTAGACGCGACACTGATTCACCAATCCATCACGGGATTATATGAAACTCGGATA GATTTGTCACAGCTTGGGGGTGACGGATCCCCAAGAAGAAGAGGAAAGTCTCGAGCGACTACAAAGACCATGACGGTGA TTATAAAGATCATGACATCGATTACAAGGATGACGATGACAAGGCTGCAGGA JPEG2025026840000004.jpg161168 (single underline: HNH domain; double underline: RuvC domain).
[0174] In some embodiments, the wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCB Reference sequence: NC_2737.2 (nucleotide sequence below). and Uniprot reference sequence: Q99ZW2 (nucleotide sequence below). amino acid sequence below). ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGAATATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG ACAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGTAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAACCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTGGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAAAAATGGCTTATTTGGGAAT CTCATTGCTTTGTCATTGGGTTTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATACTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAACGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAAGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGTACCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGAGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGACATTCAAAAAGCACAAGTGTCTGGACAAGGCGATAGTTTA CATGAACATATTGCAAATTTAGCTGGTAGCCCTGCTATTAAAAAGGTATTTTACAGACTGTAAAAGTTGTTGATGAATT GGTCAAAGTAATGGGGCGGCATAAGCCAGAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCC AGAAAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAAGAGCATCCT GTTGAAAATACTCAAATTGCAAATGAAAAGCTCTATCTCTATTATCTCCAAAATGGAAGAGACATGTATGTGGACCAAGA ATTAGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCCTTAAAGACGATTCAATAGACA ATAAGGTCTTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAA AACTATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGG TTTGAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAA TTTTGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCT AATTAGTTTCTGACTTCCGAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGC GTATCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATA AAGTTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCT AATATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGG GGAAACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATA TTGTCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTTACCAAAAGAAATTCGGACAAGCTTATT GCTCGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGC TAAGGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCT TTGAAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAAAGACTTAATCATTAAACTACCTAAA TATAGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAAAAGGAAATGAGCTGGC TCTGCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAAC AAAAACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTT ATTTTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGA AAATATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTA AACGATATACGTCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATT GATTTGAGTCAGCTAGGAGGTGACTGA JPEG2025026840000005.jpg164169 (Single underline: HNH domain; double underline: RuvC domain).
[0175] In some embodiments, Cas9 is derived from Corynebacterium ulcerans (NCBI Refs: NC_01 5683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016 786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Strept ococcus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1) ; Psychroflexus torquisI (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NC BI Ref: YP_820832.1), Listeria innocua (NCBI Ref: NP_472073.1), Campylobacter je juni (NCBI Ref: YP_002344900.1) or Neisseria meningitidis (NCBI Ref: YP_00 2342100.1), or represents Cas9 from other organisms.
[0176] In some embodiments, the Cas9 domain comprises a D10A mutation, while the the residue at position 840 in the amino acid sequence provided herein, or The residue at the corresponding position in either sequence remains a histidine.
[0177] In some embodiments, the dCas9 contains one or more mutations that inactivate Cas9 nuclease activity. For example, In some embodiments, the dCas9 domain contains the D10A and H840A mutations or another Cas9 In some embodiments, the dCas9 comprises the following dCas9: (D10A and H840A) containing the amino acid sequence: JPEG2025026840000006.jpg163170 (single underline: HNH domain; double underline: RuvC domain).
[0178] In some embodiments, the Cas9 domain comprises a D10A mutation, while the the residue at position 840 in the amino acid sequence provided herein, or The residue at the corresponding position in either sequence remains a histidine.
[0179] In other embodiments, D10A, e.g., resulting in nuclease-inactivated Cas9 (dCas9), and dCas9 variants having mutations other than H840A. For example, other amino acid substitutions at D10 and H840, or the nuclease domain of Cas9 may be used. Other substitutions within the domain (e.g., HNH nuclease subdomain and / or RuvC1 subdomain) In some embodiments, a variant or homolog of dCas9 is , at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least In some embodiments, those having at least about 99.9% identity are provided. About 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids or more, or longer. Variants of dCas9 having amino acid sequences are provided.
[0180] In some embodiments, the Cas9 fusion proteins provided herein comprise a Cas9 protein. The full-length amino acid sequence of the protein, for example, one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein contain a full-length Cas9 sequence. Examples of Suitable Cas9 Domains and Cas9 Fragments Suitable amino acid sequences are provided herein, and further suitable sequences for Cas9 domains and fragments are , will be apparent to those skilled in the art.
[0181] The Cas9 protein guides the protein to a specific DNA sequence complementary to its guide RNA. In one embodiment, the polynucleotide protease is capable of binding to a guide RNA. The activatable nucleotide binding domain can be a Cas9 domain, e.g., a nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Examples of mutable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl , Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas1 2i, but are not limited to these.
[0182] The nuclease-inactivated Cas9 protein is interchangeably referred to as a "dCas9" protein (nuclease-inactivated Cas9 protein). This enzyme may be referred to as "dead" Cas9, or catalytically inactive Cas9. Methods for generating Cas9 proteins (or fragments thereof) having domains are known (e.g., , Jinek et al., Science. 337:816-821(2012); Qi et al., “Repurposing CRISPR as a n RNA-Guided Platform for Sequence-Specific Control of Gene Expression” (2013) See Cell. 28;152(5):1173-83, the contents of each of which are incorporated herein by reference. The DNA cleavage domain of Cas9 consists of two subdomains: the HNH nuclease subdomain and the RuvC1 subdomain. It is known that the HNH subdomain cleaves the strand complementary to the gRNA. The RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains enhance Cas9 activity. For example, mutations D10A and H840A inhibit the nuclease activity of S. pyogenes Cas9. Completely inactivates nuclease activity (Jinek et al., Science. 337:816-821(2012); Qi et al., Cell. 28;152(5):1173-83 (2013)).
[0183] In some embodiments, the Cas9 domain is a Cas9 nickase. Cas9 protein, which can cut only one strand of a nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 nickase can target the double-stranded nucleic acid molecule. This allows the Cas9 nickase to separate the gRNA (e.g., sgRNA) bound to the Cas9 from the nucleotide sequence. It means cleaving paired (complementary) strands. The Cas9 nickase contains a D10A mutation and has a histidine at position 840. In embodiments, the Cas9 nickase cleaves the non-target, non-base-edited strand of the double-stranded nucleic acid molecule; This is because the Cas9 nickase base pairs with the gRNA (e.g., sgRNA) bound to Cas9. In some embodiments, the Cas9 nickase cleaves the H840A end strand. containing a natural mutation with an aspartic acid residue at position 10, or the corresponding mutation. In some embodiments, the Cas9 nickase is any of the Cas9 nickases provided herein. or at least 60%, at least 65%, at least 70%, at least 75%, at least 80% , at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, containing an amino acid sequence that is at least 98%, at least 99%, or at least 99.5% identical to the Additional suitable Cas9 nickases may be identified based on this disclosure and knowledge in the art. It is obvious to one skilled in the art and is within the scope of this disclosure.
[0184] In some embodiments, the Cas9 domain is a nuclease-inactive Cas9 domain (dCas9). For example, the dCas9 domain can cleave both strands of a double-stranded nucleic acid molecule without cleaving either strand. In some embodiments, the nucleic acid molecule can bind to a nucleic acid molecule (e.g., via a gRNA molecule). The nuclease-inactive dCas9 domain is a D10X mutation in the amino acid sequence described herein. and H840X mutations, or any of the amino acid sequences provided herein. and X is any amino acid change. Therefore, the nuclease-inactive dCas9 domain can be constructed using the D10A mutation in the amino acid sequence described herein. and H840A mutations, or their counterparts in any of the amino acid sequences described herein. As an example, a nuclease-inactive Cas9 domain may be used in cloning vectors. The vector pPlatTET-gRNA 2 (accession number BAV54124) contains the following amino acid sequence: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD (Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence-specific c control of gene expression.” Cell. 2013; 152(5):1173-83, the entire contents of which are incorporated herein by reference. , which is incorporated herein by reference.
[0185] Additional Cas9 proteins (e.g., nuclease-dead Cas9 (dCas9), Cas9 nickase (nC Nuclease-active Cas9, or nuclease-active Cas9, including its variants and homologs, is used in the present invention. It is understood that the scope of the invention is within the scope of the invention. Exemplary Cas9 proteins include, but are not limited to: In some embodiments, the Cas9 protein is a nucleic acid sequence that encodes ... In some embodiments, the Cas9 protein is a Cas9 nickase (dCas9). In some embodiments, the Cas9 protein is a nuclease-active Cas9. be.
[0186] Example of catalytically inactive Cas9 (dCas9): MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD
[0187] An example of Cas9 knickase (nCas9) is described below: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD.
[0188] Examples of catalytically active Cas9 are listed below: MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD.
[0189] In some embodiments, Cas9 is expressed in the Archaea (Archaea), which comprise the domain and kingdom of unicellular prokaryotic microorganisms. In some embodiments, the term refers to a Cas9 derived from a nucleic acid-programmable DNA binding Proteins can be synthesized using the methods described, for example, in Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res. 2017 Feb 21. doi: 10.1038 / cr.2017.21 or CasY, the entire contents of which are incorporated herein by reference. Using genomics, many studies have identified Cas9, including the first reported Cas9 in the archaeal domain of life. The divergent Cas9 protein has been identified as a CRISPR-Cas system, but it has not been studied extensively. It was discovered as part of an active CRISPR-Cas system in nanoarchaea. Two previously unknown systems, CRISPR-CasX and CRISPR-CasY, were discovered, which These are among the most compact systems discovered to date. In the base editor systems described herein, Cas9 is a CasX or a variant of CasX. In some embodiments, the base groups described herein are replaced by In the Deter system, Cas9 is replaced by CasY or a variant of CasY. Nucleic acid programmable DNA-binding protein (napDNAbp) is a novel RNA-guided DNA-binding protein. It should be understood that proteins may also be used and are within the scope of the present disclosure.
[0190] In some embodiments, the nucleic acid of any of the fusion proteins provided herein The programmable DNA binding protein (napDNAbp) can be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the napDNAbp is a CasY protein. is at least 85%, at least 90%, or At least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% In some embodiments, the napDNAbp comprises a naturally occurring amino acid sequence having identity to the napDNAbp. In some embodiments, the napDNAbp is a CasX or CasY protein present. At least 85%, at least at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least CasX and CasY from other bacterial species share at least 99.5% amino acid sequence identity. It should also be understood that they may be used in accordance with the present disclosure.
[0191] As an example, the following Cas sequences are provided:
[0192] CasX(uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) tr|F0NN87|F0NN87_SU LIH CRISPR-associated Casx protein OS = Sulfolobus islandicus (strain HVE10 / 4) G N = SiH_0402 PE=4 SV=1: MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSNVKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAV GQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSSSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNL NSDDGKVRDLKLISAYVNGELIRGEG.
[0193] >tr|F0NH53|F0NH53_SULIR CRISPR associated protein, Casx OS = Sulfolobus islandic us (strain REY15A) GN=SiRe_0771 PE=4 SV=1: MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAV GQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLN SDDGKVRDLKLISAYVNGELIRGEG.
[0194] Deltaproteobacteria CasX MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMDTDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAIL QVYWQEFKDDHVGLMCKFAQPASKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFWYKLEQVSEKGKAITNYFGRCNVA EHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESHTPVKPLAQIAGNRYASGPVGKALSDACMGTIASFL SKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDfAYNEVIARVRMWVNLNLWQKLKLSRDDA KPLRLLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENP KKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEM DEKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFT DGTDIKKSGKWQGLLYGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKI GRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQ AAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAK LAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKEL SAELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSY KSGKQPFVGAWQAFYKRRLKEVWKPNA.
[0195] CasY (ncbi.nlm.nih.gov / protein / APG80656.1) >APG80656.1 CRISPR-associated protein CasY [uncultured Parcubacteria group bacterium]: MSKRHPRISGVKGYRLHAQRLEYTGKSGAMRTIKYPLYSSPSGGRTVPREIVSAINDDYVGLYGLSNFDDLYNAEKRNEE KVYSVLDFWYDCVQYGAVFSYTAPGLLKNVAEVRGGSYELTKTLKGSHLYDELQIDKVIKFLNKKEISRANGSLDKLKKD IIDCFKAEYRERHKDQCNKLADDIKNAKKDAGASLGERQKKLFRDFFGISEQSENDKPSFTNPLNLTCCLLPFDTVNNNR NRGEVLFNKLKEYAQKLDKNEGSLEMWEYIGIGNSGTAFSNFLGEGFLGRLRENKITELKKAMMDITDAWRGQEQEEELE KRLRILAALTIKLREPKFDNHWGGYRSDINGKLSSWLQNYINQTVKIKEDLKGHKKDLKKAKEMINRFGESDTKEEAVVS SLLESIEKIVPDDSADDEKPDIPAIAIYRRFLSDGRLTLNRFVQREDVQEALIKERLEAEKKKKPKKRKKKSDAEDEKET IDFKELFPHLAKPLKLVPNFYGDSKRELYKKYKNAAIYTDALWKAVEKIYKSAFSSSLKNSFFDTDFDKDFFIKRLQKIF SVYRRFNTDKWKPIVKNSFAPYCDIVSLAENEVLYKPKQSRSRKSAAIDKNRVRLPSTENIAKAGIALARELSVAGFDWK DLLKKEEHEEYIDLIELHKTALALLLAVTETQLDISALDFVENGTVKDFMKTRDGNLVLEGRFLEMFSQSIVFSELRGLA GLMSRKEFITRSAIQTMNGKQAELLYIPHEFQSAKITTPKEMSRAFLDLAPAEFATSLEPESLSEKSLLKLKQMRYYPHY FGYELTRTGQGIDGGVAENALRLEKSPVKKREIKCKQYKTLGRGQNKIVLYVRSSYYQTQFLEWFLHRPKNVQTDVAVSG SFLIDEKKVKTRWNYDALTVALEPVSGSERVFVSQPFTIFPEKSAEEEGQRYLGIDIGEYGIAYTALEITGDSAKILDQN FISDPQLKTLREEVKGLKLDQRRGTFAMPSTKIARIRESLVHSLRNRIHHLALKHKAKIVYELEVSRFEEGKQKIKKVYA TLKKADVYSEIDADKNLQTTVWGKLAVASEISASYTSQFCGACKKLWRAEMQVDETITTQELIGTVRVIKGGTLIDAIKD FMRPPIFDENDTPFPKYRDFCDKHHISKKMRGNSCLFICPFCRANADADIQASQTIALLRYVKEEKKVEDYFERFRKLKN IKVLGQMKKI
[0196] Cas12b / C2c1 (uniprot.org / uniprot / T0D7A2#2) sp|T0D7A2|C2C1_ALIAG CRISPR-associate d endo-nuclease C2c1 OS = Alicyclobacillus acido-terrestris (strain ATCC 49025 / DSM 3922 / CIP 106132 / NCIMB 13137 / GD3B) GN=c2c1 PE=1SV=1: MAVKSIKVKLRLDDMPEIRAGLWKLHKEVNAGVRYYTEWLSLLRQENLYRRSPNGDGEQECDKTAEECKAELLERLRARQ VENGHRGPAGSDDELLQLARQLYELLVPQAIGAKGDAQQIARKFLSPLADKDAVGGLGIAKAGNKPRWVRMREAGEPGWE EEKEKAETRKSADRTADVLRALADFGLKPLMRVYTDSEMSSVEWKPLRKGQAVRTWDRDMFQQAIERMMSWESWNQRVGQ EYAKLVEQKNRFEQKNFVGQEHLVHLVNQLQQDMKEASPGLESKEQTAHYVTGRALRGSDKVFEKWGKLAPDAPFDLYDA EIKNVQRRNTRRFGSHDLFAKLAEPEYQALWREDASFLTRYAVYNSILRKLNHAKMFATFTLPDATAHPIWTRFDKLGGN LHQYTFLFNEFGERRHAIRFHKLLKVENGVAREVDDVTVPISMSEQLDNLLPRDPNEPIALYFRDYGAEQHFTGEFGGAK IQCRRDQLAHMHRRRGARDVYLNVSVRVQSQSEARGERRPPYAAVFRLVGDNHRAFVHFDKLSDYLAEHPDDGKLGSEGL LSGLRVMSVDLGLRTSASISVFRVARKDELKPNSKGRVPFFFPIKGNDNLVAVHERSQLLKLPGETESKDLRAIREERQR TLRQLRTQLAYLRLLVRCGSEDVGRRERSWAKLIEQPVDAANHMTPDWREAFENELQKLKSLHGICSDKEWMDAVYESVR RVWRHMGKQVRDWRKDVRSGERPKIRGYAKDVVGGNSIEQIEYLERQYKFLKSWSFFGKVSGQVIRAEKGSRFAITLREH IDHAKEDRLKKLADRIIMEALGYVYALDERGKGKWVAKYPPCQLILLEELSEYQFNNDRPPSENNQLMQWSHRGVFQELI NQAQVHDLLVGTMYAAFSSRFDARTGAPGIRCRRVPARCTQEHNPEPFPWWLNKFVVEHTLDACPLRADDLIPTGEGEIF VSPFSAEGDFHQIHADLNAAQNLQQRLWSDFDISQIRLRCDWGEVDGELVLIPRLTGKRTADSYSNKVFYTNTGVTYYE RERGKKRRKVFAQEKLSEEEAELLVEADEAREKSVVLMRDPSGIINRGNWTRQKEFWSMV NQRIEGYLVKQIRSRVPLQ DSACENTGDI.
[0197] In some embodiments, one of the Cas9 domains present in the fusion protein is a PAM. Sequence-free guide nucleotide sequence-programmable DNA-binding protein domains may be substituted with
[0198] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is a microbial CRI The single effector of the microbial CRISPR-Cas system is Cas. 9, Cpf1, Cas12b / C2c1, and Cas12c / C2c3. Microbial CRISPR-Cas systems are divided into class 1 and class 2 systems. Class 1 systems are multisubsystems. Class 1 systems have unit effector complexes, while Class 2 systems have single protein effectors. For example, Cas9 and Cpf1 are class 2 effectors. In addition to Cas9 and Cpf1, three Different class 2 CRISPR-Cas systems (Cas12b / C2c1 and Cas12c / C2c3) were described by Shmakov et al. , “Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Syst ems”, Mol. Cell, 2015 Nov. 5; 60(3): 385-397 (in its entirety) (The contents of which are incorporated herein by reference.) Two systems, Cas12b / C2c1 and Cas12c / C2c3, The vector contains a RuvC-like endonuclease domain related to Cpf1. The third system is a hybrid of two predicted It contains an effector with a HEPN RNase domain that is used to inhibit CRISPR RNA transcription by Cas12b / C2c1. Unlike the production of CRISPR, the production of mature CRISPR RNA is tracrRNA-independent. Transfection depends on both CRISPRRNA and tracrRNA.
[0199] The crystal structure of Alicyclobaccillus acidoterrestris Cas12b / C2c1 (AacC2c1) was revealed to be a chimeric single It has been reported as a complex with a single guide RNA (sgRNA). For example, Liu et al., "C2 c1-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism”, Mol. Cel l, 2017 Jan. 19; 65(2):310-322, the entire contents of which are incorporated herein by reference. In addition, Alicyclobacillus acidoterrestris C2c1 bound to target DNA as a ternary complex For example, Yang et al., "PAM-dependent Target DN A Recognition and Cleavage by C2C1 CRISPR-Cas endonuclease”, Cell, 2016 Dec. 15 167(7):1814-1828, the entire contents of which are incorporated herein by reference. Target DN The catalytically competent conformation of AacC2c1, along with both the A strand and the non-target DNA strand, is are independently trapped within the RuvC catalytic pocket, and Cas12b / C2c1-mediated cleavage occurs in a staggered fashion at the target DNA. The Cas12b / C2c1 ternary complex, previously identified as Cas9 and C2c1, produces a 7-nucleotide cleavage. Structural comparisons between pf1 counterparts demonstrate the diversity of mechanisms used by the CRISPR-Cas9 system. vinegar.
[0200] In some embodiments, the nucleic acid of any of the fusion proteins provided herein The programmable DNA-binding protein (napDNAbp) binds to the Cas12b / C2c1 or Cas12c / C2c3 proteins. In some embodiments, the napDNAbp can be a Cas12b / C2c1 protein. In some embodiments, the napDNAbp is a Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. At least 85%, at least 90%, at least 91%, at least 92%, or less At least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 9 8%, at least 99%, or at least 99.5% identity. In some embodiments, the napDNAbp is a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 tag. In some embodiments, the napDNAbp is a napD protein provided herein. At least 85%, at least 90%, at least 91%, at least At least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97% , amino acid sequences with at least 98%, at least 99%, or at least 99.5% identity Cas12b / C2c1 or Cas12c / C2c3 from other bacterial species may also be used in accordance with the present disclosure. Please understand that this is possible.
[0201] Cas12b / C2c1 (uniprot.org / uniprot / T0D7A2#2) sp|T0D7A2| / C2C1_ALIAG CRISPR-associ ated endo-nuclease C2c1 OS = Alicyclobacillus acido-terrestris (strain ATCC 4902 5 / DSM 3922 / CIP 106132 / NCIMB 13137 / GD3B) GN=c2c1 PE=1 SV=1 The amino acid sequence is as follows: To provide: MAVKSIKVKLRLDDMPEIRAGLWKLHKEVNAGVRYYTEWLSLLRQENLYRRSPNGDGEQECDKTAEECKAELLERLRARQ VENGHRGPAGSDDELLQLARQLYELLVPQAIGAKGDAQQIARKFLSPLADKDAVGGLGIAKAAGNKPRWVRMREAGEPGWE EEKEKAETRKSADRTADVLRALADFGLKPLMRVYTDSEMSSVEWKPLRKGQAVRTWDRDMFQQAIERMMSWESWNQRVGQ EYAKLVEQKNRFEQKNFVGQEHLVHLVNQLQQDMKEASPGLESKEQTAHYVTGRALRGSDKVFEKWGKLAPDAPFDLYDA EIKNVQRRNTRRFGSHDLFAKLAEPEYQALWREDASFLTRYAVYNSILRKLNHAKMFATFTLPDATAHPIWTRFDKLGGN LHQYTFLFNEFGERRHAIRFHKLLKVENGVAREVDDVTVPISMSEQLDNLLPRDPNEPIALYFRDYGAEQHFTGEFGGAK IQCRRDQLAHMHRRRGARDVYLNVSVRVQSQSEARGERRPPYAAVFRLVGDNHRAFVHFDKLSDYLAEHPDDGKLGSEGL LSGLRVMSVDLGLRTSASISVFRVARKDELKPNSKGRVPFFFPIKGNDNLVAVHERSQLLKLPGETESKDLRAIREERQR TLRQLRTQLAYLRLLVRCGSEDVGRRERSWAKLIEQPVDAANHMTPDWREAFENELQKLKSLHGICSDKEWMDAVYESVR RVWRHMGKQVRDWRKDVRSGERPKIRGYAKDVVGGNSIEQIEYLERQYKFLKSWSFFGKVSGQVIRAEKGSRFAITLREH IDHAKEDRLKKLADRIIMEALGYVYALDERGKGKWVAKYPPCQLILLEELSEYQFNNDRPPSENNQLMQWSHRGVFQELI NQAQVHDLLVGTMYAAFSSRFDARTGAPGIRCRRVPARCTQEHNPEPFPWWLNKFVVEHTLDACPLRADDLIPTGEGEIF VSPFSAEEGDFHQIHADLNAAQNLQQRLWSDFDISQIRLRCDWGEVDGELVLIPRLTGKRTADSYSNKVFYTNTGVTYYE RERGKKRRKVFAQEKLSEEAEELLVEADEAREKSVVLMRDPSGIINRGNWTRQKEFWSMVNQRIEGYLVKQIRSRVPLQD SACENTGDI
[0202] BhCas12b (Bacillus hisashii), NCBI Reference Sequence: WP_095142515のaminoacid Provide the following sequence: MAPKKKRKVGIHGVPAAATRSFILKIEPNEEVKKGLWKTHEVLNHGIAYYMNILKLIRQEAIYEHHEQDPKNPKKVSKAE IQAELWDFVLKMQKCNSFTHEVDKDEVFNILRELYEELVPSSVEKKGEANQLSNKFLYPLVDPNSQSGKGTASSGRKPRW YNLKIAGDPSWEEEKKKWEEDKKKDPLAKILGKLAEYGLIPLFIPYTDSNEPIVKEIKKWMEKSRNQSVRRLDKDMFIQAL ERFLSWESWNLKVKEEYEKVEKEYKTLEERIKEDIQALKALEQYEKERQEQLLRDTLNTNEYRLSKRGLRGWREIIQKWL KMDENEPSEKYLEVFKDYQRKHPREAGDYSVYEFLSKKENHFIWRNHPEYPYLYATFCEIDKKKKDAKQQATFTLADPIN HPLWVRFEERSGSNLNKYRILTEQLHTEKLKKKLTVQLDRLIYPTESGGWEEKGKVDIVLLPSRQFYNQIFLDIEEKGKH AFTYKDESIKFPLKTGTLGGARVQFDRDHLRRYPHKVESGNVGRIYFNMTVNIEPTESPVSKSLKIHRDDFPKVVNFKPKE LTEWIKDSKGKKLKSGIESLEIGLRVMSIDLGQRQAAASIFEVVDQKPDIEGKLFFPIKGTELYAVHRASFNIKLPGET LVKSREVLRKAREDNLKLMNQKLNFLRNVLHFQQFEDITEREKRVTKWISRQENSDVPLVYQDELIQIRELMYKPYKDWV AFLKQLHKRLEVEIGKEVKHWRKSLSDGRKGLYGISLKNIDEIDRTRKFLLRWSLRPTEPGEVRRLEPGQRFAIDQLNHL NALKEDRLKKMANTIIMHALGYCYDVRKKKWQAKNPACQIILFEDLSNYNPYEERSRFENSKLMKWSRREIPRQVALQGE IYGLQVGEVGAQFSSRFHAKTGSPGIRCSVVTKEKLQDNRFFKNLQREGRLTLDKIAVLKEGDLYPDKGGEKFISLSKDR KCVTTHADINAAQNLQKRFWTRTHGFYKVYCKAYQVDGQTVYIPESKDQKQKIIEEFGEGYFILKDGVYEWVNAGKLKIK KGSSKQSSSELVDSDILKDSFDLASELKGEKLMLYRDPSGNVFPSDKWMAAGVFFGKLERILISKLTNQYSISTIEDDSS KQSMKRPAATKKAGQAKKKK.
[0203] In some embodiments, the Cas12b is BvCas12B, which is a variant of BhBvCas12B. It contains the following changes relative to BhBvCas12B: S893R, K846R, and E837G.
[0204] BvCas12b (Bacillus sp. V3-13), NCBI Reference Sequence: WP_101661451.1 The acid sequence is provided below: MAIRSIKLKMKTNSGTDSIYLRKALWRTHQLINEGIAYYMNLLTLYRQEAIGDKTKEAYQAELINIIRNQQRNNGSSEEH GSDQEILALLRQLYELIIPSSIGESGDANQLGNKFLYPLVDPNSQSGKGTSNAGRKPRWKRLKEEGNPDWELEKKKDEER KAKDPTVKIFDNLNKYGLLPLFPLFTNIQKDIEWLPLGKRQSVRKWDKDMFIQAIERLLSWESWNRRVADEYKQLKEKTE SYYKEHLTGGEEWIEKIRKFEKERNMELEKNAFAPNDGYFITSRQIRGWDRVYEKWSKLPESASPEELWKVVAEQQNKMS EGFGDPKVFSFLANRENRDIWRGHSERIYHIAAYNGLQKKLSRTKEQATFTLPDAIEHPLWIRYESPGGTNLNLFKLEEK QKKNYYVTLSKIIWPSEEKWIEKENIEIPLAPSIQFNRQIKLKQHVKGKQEISFSDYSSRISLDGVLGGSRIQFNRKYIK NHKELLGEGDIGPVFFNLVVDVAPLQETRNGRLQSPIGKALKVISSDFSKVIDYKPKELMDWMNTGSASNSFGVASLLEG MRVMSIDMGQRTSASVSIFEVVKELPKDQEQKLFYSINDTELFAIHKRSFLLNLPGEVVTKNNKQQRQERRKKRQFVRSQ IRMLANVLRLETKKTPDERKKAIHKLMEIVQSYDSWTASQKEVWEKELNLLTNMAAFNDEIWKESLVELHHRIEPYVGQI VSKWRKGLSEGRKNLAGISMWNIDELEDTRRLLISWSKRSRTPGEANRIETDEPFGSSLLQHIQNVKDDRLKQMANLIIM TALGFKYDKEEKDRYKRWKETYPACQIILFENLNRYLFNLDRSRRENSRLMKWAHRSIPRTVSMQGEMFGLQVGDVRSEY SSRFHAKTGAPGIRCHALTEEDLKAGSNTLKRLIEDGFINESELAYLKKGDIIPSQGGELFVTLSKRYKKDSDNNELTVI HADINAAQNLQKRFWQQNSEVYRVPCQLARMGEDKLYIPKSQTETIKKYFGKGSFVKNNTEQEVYKWEKSEKMKIKTDTT FDLQDLDGFEDISKTIELAQEQQKKYLTMFRDPSGYFFNNETWRPQKEYWSIVNNIIKSCLKKKILSNKVEL
[0205] Polynucleotide programmable nucleotide binding domains also bind to nuclear RNA. It is understood that the present invention may include an acid programmable protein. The polynucleotide-programmable nucleotide binding domain is The nucleotide-binding domain can be linked to a nucleic acid that guides the RNA. Other DNA-binding proteins are also within the scope of this disclosure, although they are not specifically listed in this disclosure. It has not been done.
[0206] The CAs proteins that can be used herein include class 1 and class 2. Non-limiting examples of proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5d, Cas5t, as5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also called Csn1 or Csx12), Cas10, Csy1, C sy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5, Csn1, Csn2, Csm2, Csm3, Csm4, Csm 5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Cssx16, Cx16, Cx, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1 , Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i, CARF, DinG, and their homologs Unmodified CRISPR enzymes, like Cas9, are composed of two It has functional endonuclease domains, RuvC and HNH, and can have DNA cleavage activity. CRISPR enzymes target sequences, such as within the target sequence and / or the complementary strand of the target sequence. For example, CRISPR enzymes can induce cleavage of one or both strands of a target sequence. Approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 25 nucleotides from the first or last nucleotide of the string Induce breaks in one or both strands at 50, 100, 200, 500 base pairs or more. It is possible.
[0207] Loss of ability to cleave one or both strands of a target polynucleotide containing the target sequence To achieve this, vectors were used to encode CRISPR enzymes that were mutated relative to the corresponding wild-type enzyme. Cas9 can be a wild-type exemplary Cas9 polypeptide (e.g., Cas9 from S. pyogenes). Cas9) and at least approximately 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93% %, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology Cas9 can refer to a polypeptide having a wild-type exemplary Cas9 polypeptide ( For example, from S. pyogenes), at most, or at most, approximately, about 50%, 60%, 70% , 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or polypeptides having sequence homology. or deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any of these It can refer to modified forms of the Cas9 protein that may contain amino acid changes, such as combinations.
[0208] In some embodiments, the methods described herein involve recombinantly engineered The guide RNA (gRNA) is required for Cas binding. The required scaffold sequence and a user-defined approximately 20-base spacer that defines the genomic target to be modified. Therefore, the specificity of Cas proteins can be altered by targeting the genome. Partially due to the specificity of the gRNA targeting sequence for the genomic target relative to other parts of the genome It will be understood that the value is determined by the following equation.
[0209] Cas9 nuclease has two functional endonuclease domains, RuvC and HNH. Upon binding to the target DNA, Cas9 undergoes a second conformational change, which leads to the formation of nucleases. Cas9-mediated DNA cleavage is achieved by positioning the cleavage domain on the opposite strand of the target DNA. The end result is a double-strand break (DSB) in the target DNA (approximately 3–4 nucleotides upstream of the PAM sequence). The resulting DSBs are repaired by one of two general repair pathways: (1) efficient (2) the less efficient but more error-prone non-homologous end joining (NHEJ) pathway; or (3) the less efficient but more highly fidel Homologous directed repair (HDR) pathway.
[0210] The "efficiency" of non-homologous end joining (NHEJ) and / or homology-directed repair (HDR) is not known by any simple It can be calculated in a convenient way. For example, in some cases, the efficiency is For example, a test nuclease assay can be used to determine the percentage of cleavage product. The ratio of product to substrate can be used to calculate the percentage. For example, direct cleavage of DNA containing the newly integrated restriction sequence as a result of successful HDR. A nuclease enzyme can be used to measure the amount of substrate cleaved, which increases the rate of HDR. As an illustrative example, the HDR percentage (percentage) The cleavage percentage can be calculated using the following formula: [(cleavage product) / (substrate + cleavage product)] (e.g., (b+c) / (a+b+c) where "a" is the band intensity of the DNA substrate, and "b" and and "c" are cleavage products.
[0211] In some cases, efficiency can be expressed as the success rate of NHEJ. For example, T7 endonuclease Cleavage products were generated using a cleavage enzyme I assay, and the ratio of product to substrate was used to determine the percentage of NHEJ. The conversion can be calculated using wild-type and mutant T7 endonuclease I. High levels of NHEJ (NHEJ generates small random insertions or deletions (indels) at the initial cleavage site) It cleaves the mismatched heteroduplex DNA resulting from hybridization. This indicates a high rate of NHEJ (high efficiency of NHEJ). (Percentage) is calculated using the formula (1-(1-(b+c) / (a+b+c))1 / 2 ) × 100 , where "a" is the band intensity of the DNA substrate, and "b" and "c" are the cleavage products ( Ran et. al., Cell. 2013 Sep. 12; 154(6):1380-9; and Ran et. al., Nat Protoc. 2013 Nov.;8(11):2281-2308).
[0212] The NHEJ repair pathway is the most active repair mechanism, resulting in small nucleotide insertions or deletions at DSB sites. The random nature of NHEJ-mediated DSB repair is due to the lack of Cas9 and gRNA. Alternatively, a population of cells expressing the guide polynucleotide may result in a diverse array of mutations. In most cases, NHEJ results in small indentations in the target DNA, which has important practical implications. This results in premature deletion of the target gene's open reading frame (ORF). Amino acid deletions, insertions, or frameshift mutations resulting in undesired stop codons occur. The ideal end result is a loss-of-function mutation in the target gene.
[0213] NHEJ-mediated DSB repair often disrupts the open reading frame of a gene, but Homologous-directed repair (HDR) is a method for repairing single nucleotide changes, such as the addition of a fluorophore or tag. These can be used to generate specific nucleotide changes ranging from large insertions such as
[0214] To utilize HDR for gene editing, a DNA repair template containing the desired sequence is prepared using gRNA and It can be delivered to the cell type of interest along with Cas9 or Cas9 nickase. The rate is determined by the desired edit and the regions immediately upstream and downstream of the target (called left and right homology arms). The length of each homologous arm determines the magnitude of the change introduced. The repair template may depend on the size of the insert, with larger inserts requiring longer homology arms. a single-stranded oligonucleotide, a double-stranded oligonucleotide, or a double-stranded DNA plasmid The efficiency of HDR was measured in cells expressing Cas9, gRNA, and exogenous repair template. HDR occurs between the S and G2 phases of the cell cycle, so the risk of developing a mutation is generally low (less than 10% correcting alleles). The efficiency of HDR can be increased by synchronizing cells. Chemically or genetically inhibiting offspring can also increase HDR frequency.
[0215] In some embodiments, the Cas9 is a modified Cas9. There may be additional sites of partial homology throughout the genome. These are called off-target sites ("off-target") and must be considered when designing gRNAs. However, in addition to optimizing the gRNA design, modifying Cas9 may also improve CRISPR transcription. It can also enhance specificity. Cas9 is a combination of two nuclease domains, RuvC and HNH. Cas9 Nicker, a D10A mutant of SpCas9, generates double-strand breaks (DSBs) through its cleavage activity. The enzyme possesses a single nuclease domain and generates DNA nicks rather than DSBs. Nickases can also be combined with HDR-mediated gene editing for gene editing.
[0216] In some cases, the Cas9 is a variant Cas9 protein. Variant Cas9 polypeptide differs by a single amino acid compared to the amino acid sequence of the wild-type Cas9 protein (e.g., In some instances, the amino acid sequence may be a sequence of a variant of the amino acid sequence, such as a nucleotide sequence of a nucleotide, a nucleotide sequence of ... The modified Cas9 polypeptide contains an amino acid sequence that reduces the nuclease activity of the Cas9 polypeptide. have alterations (e.g., deletions, insertions, or substitutions). For example, in some instances, The ant-Cas9 polypeptides exhibited less than 50% of the nuclease activity of the corresponding wild-type Cas9 protein. less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1%. In this case, the variant Cas9 protein does not have substantial nuclease activity. If the protein is a variant Cas9 protein that does not have substantial nuclease activity, This may be referred to as "dCas9."
[0217] In some cases, the variant Cas9 protein has reduced nuclease activity. For example, A variant Cas9 protein may be a variant of a wild-type Cas9 protein (e.g., a wild-type Cas9 protein). Less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 1%, or or less than about 0.1%.
[0218] In some cases, the variant Cas9 protein cleaves the complementary strand of the guide target sequence. but have a reduced ability to cleave the non-complementary strand of the double-stranded guide target sequence. Variant Cas9 proteins contain mutations (amino acid substitutions) that reduce the function of the RuvC domain. As a non-limiting example, in some embodiments, the variant The Cas9 protein has D10A (aspartic acid to alanine at amino acid position 10). , thus, can cleave the complementary strand of the double-stranded guide target sequence, but This variant of Cas9 has a reduced ability to cleave the non-complementary strand of the target sequence (hence, When a protein cleaves a double-stranded target nucleic acid, it creates a single-strand break (S) instead of a double-strand break (DSB). (See, e.g., Jinek et al., Science. 2012 Aug. 17; 337(6096):816-21) .
[0219] In some cases, the variant Cas9 protein cleaves the non-complementary strand of the double-stranded guide target sequence. can cleave the complementary strand of the guide target sequence, but with a reduced ability to cleave the complementary strand of the guide target sequence. Variant Cas9 proteins have mutations (amino acid substitutions) that reduce the function of the HNH domain. (RuvC / HNH / RuvC domain motif). In embodiments, the variant Cas9 protein is H840A (His at amino acid position 840). alanine to cysteine) mutation, thus inhibiting the non-complementary strand of the guide target sequence. can cleave the complementary strand of the guide target sequence, but has a reduced ability to cleave the complementary strand of the guide target sequence (Thus, this variant Cas9 protein cleaves the double-stranded guide target sequence.) (When the Cas9 protein is inserted into the guide target sequence, it generates an SSB instead of a DSB.) The ability to cleave the guide target sequence (e.g., single-stranded guide target sequence) is reduced, but the ability to cleave the guide target sequence (e.g., single-stranded guide target sequence) is reduced. The strand retains the ability to bind to the guide target sequence.
[0220] In some cases, the variant Cas9 protein is capable of cleaving complementary and non-complementary strands of a double-stranded target DNA. As a non-limiting example, in some cases, the variant The Cas9 protein has both the D10A and H840A mutations, resulting in a polypeptide These have a reduced ability to cleave both the complementary and non-complementary strands of double-stranded target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA). However, they retain the ability to bind to target DNA (e.g., single-stranded target DNA).
[0221] As another non-limiting example, in some cases, the variant Cas9 protein is 76A and W1126A mutations, such that the polypeptide binds to the target DNA (e.g., single-stranded target have reduced ability to cleave target DNA (e.g., single-stranded target DNA), but have reduced ability to bind to target DNA (e.g., single-stranded target DNA). The power is retained.
[0222] As another non-limiting example, in some cases, the variant Cas9 protein may be P4 75A, W476A, N477A, D1125A, W1126A, and D1127A mutations, resulting in Such Cas9 proteins have a reduced ability to cleave target DNA. Although the target DNA (e.g., single-stranded target DNA) has a reduced ability to cleave the target DNA (e.g., single-stranded target DNA), It retains the ability to bind to DNA.
[0223] As another non-limiting example, in some cases, the variant Cas9 protein is H8 40A, W476A, and W1126A mutations, so that the polypeptide binds to target DNA (e.g., The target DNA (e.g., single-stranded target DNA) has a reduced ability to cleave, but As another non-limiting example, in some cases, The antCas9 protein has H840A, D10A, W476A, and W1126A mutations, resulting in Such Cas9 proteins have a reduced ability to cleave target DNA. has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but In some embodiments, the variant retains the ability to bind to single-stranded target DNA. This Cas9 has a restored catalytic His residue at position 840 of the Cas9 HNH domain (A840H). .
[0224] As another non-limiting example, in some cases, the variant Cas9 protein is H8 40A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, resulting in The polypeptide has reduced ability to cleave target DNA (e.g., single-stranded target DNA), but not target DNA. As another non-limiting example, some In some cases, the variant Cas9 protein is D10A, H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, such that the polypeptide binds to the target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded targets). have reduced ability to cleave target DNA (e.g., single-stranded target DNA), but have reduced ability to bind to target DNA (e.g., single-stranded target DNA). If the variant Cas9 protein has the W476A and W1126A mutations, or variant Cas9 proteins P475A, W476A, N477A, D1125A, W1126A, and D1127 A mutation, the variant Cas9 protein does not bind efficiently to the PAM sequence. Therefore, in such cases, such variant Cas9 proteins are used in methods of conjugation. In other words, in some cases, such a barrier is not required. When a recombinant Cas9 protein is used in the method of attachment, the method may include a guide RNA. The method can be performed in the absence of a PAM sequence (thus, the specificity of binding depends on the guide RNA). (This is brought about by the target segment of A). To achieve the above effect, other residues were varied. Non-limiting examples include: Residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and and / or A987 can be altered (i.e., substituted). is also suitable.
[0225] In some embodiments, a variant Cas9 protein (e.g., Cas9) with reduced catalytic activity is provided. Proteins D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and and / or A987 mutations, e.g., D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A , H983A, A984A, and / or D986A) that interacts with the guide RNA As long as it retains its ability to bind to target DNA in a site-specific manner (by guide RNA), (This is because the target DNA sequence is guided by the
[0226] As an alternative to S. pyogenes Cas9, the Cpf1 family exhibits cleavage activity in mammalian cells. RNA-guided endonucleases derived from Prevotella and Francisella 1 may be mentioned. The current CRISPR (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. 1 is a class II CRISPR / Cas RNA-guided endonuclease. This adaptive immune mechanism is involved in the Pre Cpf1 is found in the bacteria Vogelia and Francisella. Therefore, Cpf1 has a different PAM specificity than Cas9. Similar to Cas9, Cpf1 also has a gene encoding a nucleic acid-programmable DNA-binding protein. Cpf1 is a CRISPR effector of Cas2. Cpf1 mediates robust DNA interference with characteristics distinct from Cas9. Cpf1 is a single RNA-guided endonuclease that lacks tracrRNA. , utilizing a T-rich protospacer adjacent motif (TTN, TTTN, or YTN). 1 cleaves DNA with staggered double-strand breaks. Of the 16 Cpf1 family proteins, Acid Two enzymes from Aminococcus and Lachnospiraceae enable efficient genome editing in human cells The Cpf1 protein has been shown to have activity. al structure of Cpf1 in complex with guide RNA and target DNA.” Cell (165) 2016 , pp. 949-962, the entire contents of which are incorporated herein by reference.
[0227] The Cpf1 gene is associated with the CRISPR locus and is responsible for finding and cleaving viral DNA. It encodes an endonuclease that uses guide RNA. Cpf1 is smaller and simpler than Cas9. Being an endonuclease, Cpf1 overcomes some of the limitations of the CRISPR / Cas9 system. Unlike Cas9 nuclease, Cpf1-mediated DNA cleavage results in a short 3' overhang. The staggered cleavage pattern of Cpf1 is different from that of traditional restriction enzyme cloning. This opens up the possibility of directional gene transfer, similar to gene editing. Similar to the Cas9 variants and orthologues described above, Cpf1 can enhance the efficiency of CRISPR The number of sites that SpCas9 can target was determined by dividing the target region into AT-rich regions lacking the NGG PAM site preferred by SpCas9. The Cpf1 locus contains the α / β mixed domain, RuvC. It contains RuvC-I and the subsequent helical region, RuvC-II, and a zinc finger-like domain. The protein contains a RuvC-like endonuclease domain similar to the RuvC domain of Cas9. Furthermore, Cpf1 does not have the HNH endonuclease domain, and the N-terminus of Cpf1 is located in the α-helix of Cas9. The Cpf1 CRISPR-Cas domain organization makes Cpf1 functionally unique. This indicates that the Cpf1 gene is classified as a Class 2, Type V CRISPR system. The locus encodes Cas1, Cas2, and Cas4 proteins that are more similar to type I and type III than to type II systems. Functional Cpf1 does not require trans-activating CRISPR RNA (tracrRNA); Cpf1 is not only smaller than Cas9, but also smaller than sgR. It has an NA molecule (about half the number of nucleotides of Cas9), which is beneficial for genome editing. In contrast to the G-rich PAM targeted by Cas9, the Cpf1-crRNA complex targets the motif 5'-YTN-3' Cleavage of target DNA or RNA is achieved by identifying the protospacer adjacent to the PAM. , Cpf1 induces sticky-end-like DNA double-strand breaks with overhangs of 4 or 5 nucleotides. Introduce.
[0228] Also, use as a guide nucleotide sequence-programmable DNA-binding protein domain Nuclease-inactive Cpf1 (dCpf1) variants that can be used in the compositions and methods of the invention are also contemplated. The Cpf1 protein is useful in the RuvC-like endonuclease domain, which is similar to the RuvC domain of Cas9. The N-terminus of Cpf1 contains a clease domain but no HNH endonuclease domain. It does not have the α-helical recognition lobe of Cas9. Zetsche et al., Cell, 163, 759-771, 201 5 (incorporated herein by reference) demonstrated that the RuvC-like domain of Cpf1 induces cleavage of both DNA strands. We showed that inactivation of the RuvC-like domain inactivates Cpf1 nuclease activity. For example, the protons corresponding to D917A, E1006A, or D1255A in Francisella novicida Cpf1 The natural mutation inactivates Cpf1 nuclease activity. dCpf1 is D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or contains mutations corresponding to D917A / E1006A / D1255A, which inactivate the RuvC domain of Cpf1 Any mutation, e.g., a substitution mutation, a deletion, or an insertion, can be used in accordance with the present disclosure. I want you to understand that.
[0229] In some embodiments, the nucleic acid of any of the fusion proteins provided herein The programmable DNA binding protein (napDNAbp) may be the Cpf1 protein. In some embodiments, the Cpf1 protein is Cpf1 nickase (nCpf1). In some cases, the Cpf1 protein is nuclease-inactive Cpf1 (dCpf1). In embodiments, Cpf1, nCpf1, or dCpf1 is selected from the group consisting of Cpf1 sequences disclosed herein. At least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least In some embodiments, the amino acid sequence has at least 99%, or at least 99.5%, identity to the amino acid sequence. In some embodiments, dCpf1 has a sequence that is at least 85% identical to the Cpf1 sequence disclosed herein, at least At least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95% , at least 96%, at least 97%, at least 98%, at least 99%, or at least 99 Contains amino acid sequences with 0.5% identity, D917A, E1006A, D1255A, D917A / E1006A, D917 Contains mutations corresponding to A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A.
[0230] It is understood that Cpf1 from other bacterial species may also be used in accordance with the present disclosure. Therefore, the following exemplary Cpf1 sequences from other bacterial species can also be used in accordance with the present disclosure:
[0231] Wild-type Francisella novicida Cpf1 (D917, E1006, and D1255 are shown in bold and underlined). MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYS DVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDIT DIDEALEIIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAE ELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYK MSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLT DLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILA NFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEH FYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKI FDDKAIKENKGEGYKKIVYKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKF IDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISESYIDSVVNQGKLYLFQIYNKDFSAYSKGR PNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIAANKNKDNPKKESVFEYDLIKDKRFTEDKFFF HCPITINFKSSGANKFNDEINLLLKEKANDVHILSI D RGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAI EKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF E DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVF KDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLD KGYFEFSFDYKNFGDKAAKGKWTIASFFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESD KKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA D ANGAYHIGLKGLMLLGRIKNNQEGK KLNLVIKNEEYFEFVQNRNN
[0232] Francisella novicida Cpf1 D917A (A917, E1006, and D1255 are underlined in bold.) MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYS DVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDIT DIDEALEIIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAE ELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYK MSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLT DLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEKELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILA NFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNLLHKLKIFHISQSEDKANILDKDEH FYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKI FDDKAIKENKGGEGYKKIVYKLLPGANKMLPKVFFSASKIFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKF IDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGR PNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPCKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFF HCPITINFKSSGANKFNDEINLLLKEKANDVHILSI A RGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAI EKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF E DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVF KDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLD KGYFEFSFDYKNFGDKAAKGKWTIASFFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESD KKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA D ANGAYHIGLKGLMLLGRIKNNQEGK KLNLVIKNEEYFEFVQNRNN
[0233] Francisella novicida Cpf1 E1006A (D917, A1006, and D1255 are underlined in bold.) MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYS DVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDIT DIDEALEIIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAE ELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYK MSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLT DLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILA NFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNLLHKLKIFHISQSEDKANILDKDEH FYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKI FDDKAIKENKGGEGYKKIVYKLLPGANKMLPKVFFSASKIFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKF IDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGR PNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPCKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFF HCPITINFKSSGANKFNDEINLLLKEKANDVHILSI D RGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAI EKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF A DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVF KDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLD KGYFESFDYKNFGDKAAKKGKWTIASFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESD KKFFAKLTSVLNTILQMRNSKTTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA D ANGAYHIGLKGLMLLGRIKNNQEGK KLNLVIKNEEYFEFVQNRNN
[0234] Francisella novicida Cpf1 D1255A (D917, E1006, and A1255 are underlined in bold.) MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYS DVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDIT DIDEALEIIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAE ELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYK MSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLT DLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILA NFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEH FYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKI FDDKAIKENKGEGYKKIVYKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKF IDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISESYIDSVVNQGKLYLFQIYNKDFSAYSKGR PNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIAANKNKDNPKKESVFEYDLIKDKRFTEDKFFF HCPITINFKSSGANKFNDEINLLLKEKANDVHILSI D RGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAI EKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF E DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVF KDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLD KGYFEFSFDYKNFGDKAAKGKWTIASFFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESD KKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA A ANGAYHIGLKGLMLLGRIKNNQEGK KLNLVIKNEEYFEFVQNRNN
[0235] Francisella novicida Cpf1 D917A / E1006A (A917, A1006, and D1255 are underlined in bold.) MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYS DVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDIT DIDEALEIIKSFKGWTTYFKGFHENRKNVYSSDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAE ELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYK MSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLT DLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEKELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILA NFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNLLHKLKIFHISQSEDKANILDKDEH FYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKI FDDKAIKENKGGEGYKKIVYKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKF IDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGR PNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPCKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFF HCPITINFKSSGANKFNDEINLLLKEKANDVHILSI A RGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAI EKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVFA DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVF KDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLD KGYFEFSFDYKNFGDKAAKGKWTIASFFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESD KKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA D ANGAYHIGLKGLMLLGRIKNNQEGK KLNLVIKNEEYFEFVQNRNN
[0236] Francisella novicida Cpf1 D917A / D1255A (A917, E1006, and A1255 are underlined in bold.) MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYS DVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDIT DIDEALEIIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAE ELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYK MSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLT DLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEKELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILA NFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNLLHKLKIFHISQSEDKANILDKDEH FYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKI FDDKAIKENKGGEGYKKIVYKLLPGANKMLPKVFFSASKIFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKF IDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGR PNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPCKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFF HCPITINFKSSGANKFNDEINLLLKEKANDVHILSI A RGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAI EKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF E DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVF KDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLD KGYFESFDYKNFGDKAAKKGKWTIASFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESD KKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA A ANGAYHIGLKGLMLLGRIKNNQEGK KLNLVIKNEEYFEFVQNRNN
[0237] Francisella novicida Cpf1 E1006A / D1255A (D917, A1006, and A1255 are underlined in bold.) MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYS DVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDIT DIDEALEIIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAE ELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYK MSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLT DLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILA NFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEH FYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKI FDDKAIKENKGEGYKKIVYKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKF IDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISESYIDSVVNQGKLYLFQIYNKDFSAYSKGR PNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIAANKNKDNPKKESVFEYDLIKDKRFTEDKFFF HCPITINFKSSGANKFNDEINLLLKEKANDVHILSI D RGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAI EKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF A DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVF KDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLD KGYFEFSFDYKNFGDKAAKGKWTIASFFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESD KKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA A ANGAYHIGLKGLMLLGRIKNNQEGK KLNLVIKNEEYFEFVQNRNN
[0238] Francisella novicida Cpf1 D917A / E1006A / D1255A (A917, A1006, and A1255 are underlined in bold) show.) MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEEILSSVCISEDLLQNYS DVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDIT DIDEALEIIKSFKGWTTYFKGFHENRKNVYSSDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAE ELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYK MSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLT DLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEKELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILA NFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNLLHKLKIFHISQSEDKANILDKDEH FYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKI FDDKAIKENKGGEGYKKIVYKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKF IDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGR PNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIAANKNKDNPKKESVFEYDLIKDKRFTEDKFFF HCPITINFKSSGANKFNDEINLLLKEKANDVHILSI A RGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAI EKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF A DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVF KDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLD KGYFEFSFDYKNFGDKAAKGKWTIASFFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESD KKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADVNGNFFDSRQAPKNMPQDA A ANGAYHIGLKGLMLLGRIKNNQEGK KLNLVIKNEEYFEFVQNRNN
[0239] The polynucleotide-programmable nucleotide-binding domain of the base editor is The polynucleotide programmable vector itself can contain one or more domains. The nucleotide-binding domain capable of binding to the nuclease may include one or more nuclease domains. In some embodiments, the nucleic acid of a polynucleotide-programmable nucleotide binding domain The nuclease domain can comprise an endonuclease or an exonuclease. As used herein, the term "exonuclease" refers to an enzyme that liberates nucleic acids (e.g., RNA or DNA). The term "endonuclease" refers to a protein or polypeptide that can be digested from its termini. A "clease" is a nucleic acid that can catalyze (e.g., cleave) an internal region of a nucleic acid (e.g., DNA or RNA). In some embodiments, an endonuclease refers to a protein or polypeptide that It is capable of cleaving a single strand of double-stranded nucleic acid. Both strands of a double-stranded nucleic acid molecule can be cleaved. The programmable nucleotide binding domain can be a deoxyribonuclease. In some embodiments, the polynucleotide programmable nucleotide binding domain is a ribonucleotide. It may be a nuclease.
[0240] In one embodiment, the nucleotide of the polynucleotide-programmable nucleotide binding domain The cleavage domain can cleave zero, one, or two strands of a target polynucleotide. In some cases, the polynucleotide-programmable nucleotide binding domain can comprise a nickase domain. As used herein, the term "nickase" means A nuclease capable of cleaving only one strand of a double-stranded nucleic acid molecule (e.g., DNA) refers to a polynucleotide-programmable nucleotide-binding domain containing a nucleotidease domain In some embodiments, the nickase is an active polynucleotide programmable nucleotide. By introducing one or more mutations into the polypeptide binding domain, Derived from a fully catalytically active (e.g., native) form of a grammable nucleotide-binding domain For example, a polynucleotide programmable nucleotide binding domain can be When the nickase domain derived from Cas9 is included, the nickase domain derived from Cas9 is It may contain a D10A mutation and a histidine (H) at position 840. In this case, residue H840 retains catalytic activity, thereby cleaving one strand of a nucleic acid duplex. In another example, the nickase domain from Cas9 can include an H840A mutation. while the amino acid residue at position 10 remains D. In this case, the nickase may contain all or part of the nuclease domain that is not required for nickase activity. By removing a portion of the polynucleotide, the programmable nucleotide binding domain The polynucleotides can be derived from fully catalytically active (e.g., native) forms of the polypeptides. A nucleotide-programmable nucleotide-binding domain derived from Cas9 When the Cas9-derived nickase domain contains a RuvC domain or an HNH domain, The sequence may include a deletion of all or part of the sequence.
[0241] Thus, polynucleotides containing nickase domains can be used as programmable nucleotides. Base editors containing a base-binding domain can be used to bind to specific polynucleotide target sequences (e.g., binding domains). The DNA fragments are then ligated to the target nucleic acid, which then generates a single-stranded DNA break (a nick) at the target site (determined by the complementary sequence of the guide nucleic acid). In some embodiments, a nickase domain (e.g., a nickase domain derived from Cas9) can be used. Nucleic acid double-stranded target polynucleotides cleaved by base editors containing a base editor domain (a cleavage enzyme domain) The strand of the base sequence is the strand that is not edited by the base editor (i.e., (The strand cleaved by the cleavage is the opposite strand to the strand containing the base to be edited.) and base editors containing a nickase domain (e.g., a nickase domain derived from Cas9). can cut the strand of the DNA molecule targeted for editing. The non-target strand is not cleaved.
[0242] Catalytically dead (i.e., unable to cleave target polynucleotide sequence) polynucleotides Base editors comprising nucleotide-programmable nucleotide-binding domains are also described herein. As used herein, the terms "catalytically dead" and "dead nuclease" are used interchangeably. " refers to one or more mutations and / or mutations that result in the inability to cleave a strand of nucleic acid. Polynucleotides with deletions can be used to refer to programmable nucleotide binding domains. In some embodiments, catalytically dead polynucleotide proteases are used interchangeably. The programmable nucleotide-binding domain base editors bind to one or more nuclease domains. Nuclease activity can be lost as a result of specific point mutations in the For example, in the case of base editors containing the Cas9 domain, Cas9 is able to reverse the D10A mutation and the H840A mutation. Such mutations can render both nuclease domains inactive. In another embodiment, the catalytically dead polynuclease is Nucleotide-programmable nucleotide-binding domains can be used in catalytic domains (e.g., RuvC1 and The present invention may include one or more deletions of all or part of the HNH domain (e.g., the HNH domain and / or the HNH domain). In some embodiments, the catalytically dead polynucleotide programmable nucleotide binding domain The inclusion of point mutations (e.g., D10A or H840A) as well as the entire or partial nuclease domain contains a partial deletion.
[0243] Also provided herein are the following polynucleotide-programmable nucleotide binding domains: Catalytically dead polynucleotide programmable nucleic acid from a previously functioning version Mutations that can generate peptide-binding domains are also contemplated. In the case of dead Cas9 ("dCas9"), mutations other than D10A and H840A are present, leading to the nuclease Variants that result in inactive Cas9 are provided. Such mutations include, for example, D10 and other amino acid substitutions at H840 or other substitutions within the nuclease domain of Cas9. (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain) Includes:
[0244] Additional suitable nuclease-inactive dCas9 domains are described in this disclosure and in the art. Such further examples would be apparent to those skilled in the art based on their knowledge and are within the scope of this disclosure. Exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, D10A / H Includes 840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotec 2013; 31(9): 833-838, the entire contents of which are incorporated herein by reference. In some embodiments, the dCas9 domain is a dCas9 domain provided herein. At least 60%, at least 65%, at least 70%, or at least at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least or at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical In some embodiments, the Cas9 domain comprises an amino acid sequence having the properties described herein. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31 , 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 and In some embodiments, Cas9 comprises an amino acid sequence with more than one mutation. A domain has at least 10 amino acid sequences, At least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least at least 200, at least 250, at least 300, at least 350, at least 400, at least At least 500, at least 600, at least 700, at least 800, at least 900, at least 100 0, at least 1100, or at least 1200 identical consecutive amino acid residues Contains the amino acid sequence.
[0245] Polynucleotide programmable nucleotides that can be incorporated into base editors Non-limiting examples of binding domains include domains derived from CRISPR proteins, restriction nucleases, and the like. Enzymes, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases In some cases, base editors are used to manipulate nucleic acids using CRISPR ( Clustered Regularly Interspaced Short Palindromic Repeats (CIR)-mediated modification natural or modified proteins that can bind to a nucleic acid sequence via a binding guide nucleic acid Polynucleotide-programmable nucleotide-binding domains containing proteins or portions thereof Such proteins are referred to herein as "CRISPR proteins." Thus, disclosed herein are polynucleotides containing all or part of a CRISPR protein. Nucleotide-programmable base editors containing nucleotide-binding domains (i.e. , a base editor containing all or part of a CRISPR protein as a domain (which is a base The CRISPR protein-derived domain of the base editor is also called the "CRISPR protein-derived domain." The CRISPR protein-derived domain integrated into the CRISPR protein is similar to that of the wild-type or naturally occurring CRISPR protein. For example, as described below, CRISPR protein-derived DNA can be modified compared to the target protein. The main one is a gene that contains one or more mutations or insertions compared to the wild-type or native CRISPR protein. , deletions, rearrangements and / or recombinations.
[0246] In some embodiments, the base editor comprises a CRISPR protein-derived domain. The main component is capable of binding to a target polynucleotide when combined with a binding guide nucleic acid. Endonucleases (e.g., deoxyribonucleases or ribonucleases) that can In some embodiments, the CRISPR protein incorporated into the base editor The protein-derived domain binds to the target polynucleotide when combined with the binding guide nucleic acid. In some embodiments, the base editor is a nickase that can bind to The CRISPR protein-derived domains integrated into the CRISPR protein bind to the CRISPR protein when combined with the guide nucleic acid. A catalytically dead domain is capable of binding to a target polynucleotide when In embodiments, a target polynucleotide that binds to a CRISPR protein-derived domain of a base editor is The nucleic acid is DNA, and in some embodiments, the nucleic acid comprises a CRISPR protein-derived domain of a base editor. The target polynucleotide that binds to the in is RNA.
[0247] In some embodiments, the CRISPR protein-derived domain of the base editor is ium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria ( NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021 284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (N CBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Belliella ba ltica (NCBI Ref: NC_018010.1); Psychroflexus torquis (NCBI Ref: NC_018721.1); St reptococcus thermophilus (NCBI Ref: YP_820832.1); Listeria innocua (NCBI Ref: NP _472073.1); Campylobacter jejuni (NCBI Ref: YP_002344900.1); Neisseria meningiti dis (NCBI Ref: YP_002342100.1), Streptococcus pyogenes, or Staphylococcus au The Cas9 fragment may comprise all or part of the Cas9 fragment derived from Cas9.
[0248] In some embodiments, the Cas9-derived domain of the base editor is derived from Staphylococcus aureus. In some embodiments, the SaCas9 domain is a nuclease Active SaCas9, nuclease-inactive SaCas9 (SaCas9d), or SaCas9 nickase (SaCas9n) In some embodiments, the SaCas9 domain comprises an N579X mutation. In some embodiments, the SaCas9 domain comprises an N579A mutation. The main, SaCas9d, or SaCas9n domain binds to a nucleic acid sequence with a non-canonical PAM. In some embodiments, a SaCas9 domain, a SaCas9d domain, Alternatively, the SaCas9n domain can bind to a nucleic acid sequence having an NNGRRT PAM sequence. In some embodiments, the SaCas9 domain contains E781X, N967X, and R1014X mutations. Contains one or more of the following:
[0249] In some embodiments, the Cas9 domain is a Cas9 domain from Staphylococcus aureus (SaCa In some embodiments, the SaCas9 domain is a nuclease-active SaCas9, nuclease The enzyme-inactive SaCas9 (SaCas9d), or SaCas9 nickase (SaCas9n). In embodiments, SaCas9 contains an N579A mutation, or an amino acid sequence provided herein. containing the corresponding mutation in either of the sequences.
[0250] In some embodiments, a SaCas9 domain, a SaCas9d domain, or a SaCas9n domain The nucleotides can bind to nucleic acid sequences with non-canonical PAMs, and in some embodiments The SaCas9 domain, SaCas9d domain, or SaCas9n domain is NNGRRT or NNNRRT In some embodiments, SaCa can bind to a nucleic acid sequence having a PAM sequence. The s9 domain may contain one or more of the E781X, N967X, and R1014X mutations, or any of the mutations provided herein. and corresponding mutations in any of the amino acid sequences described above, where X is any amino acid. In some embodiments, the SaCas9 domain is E781K, N967K, and R101. 4H mutations or in any of the amino acid sequences provided herein. In some embodiments, the SaCas9 domain comprises one or more corresponding mutations in the E781K, N967K, or R1014H mutation, or It includes corresponding mutations in either
[0251] The base editor comprises a domain derived from all or part of Cas9, a high-fidelity Cas9. In some embodiments, the high-fidelity Cas9 domain of the base editor can , between the Cas9 domain and the sugar-phosphate backbone of DNA compared to the corresponding wild-type Cas9 domain. An engineered Cas9 domain containing one or more mutations that reduce electrostatic interactions. High-fidelity Cas9 domains with reduced electrostatic interactions with the sugar-phosphate backbone of NAs exhibited fewer In some embodiments, the Cas9 domain may have off-target effects. (e.g., wild-type Cas9 domain) reduces the binding between the Cas9 domain and the sugar-phosphate backbone of DNA. In some embodiments, the Cas9 domain comprises one or more mutations that reduce Ca The binding between the s9 domain and the sugar-phosphate backbone of DNA is reduced by at least 1%, at least 2%, or at least At least 3%, at least 4%, at least 5%, at least 10%, at least 15%, at least 20%, At least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least at least 50%, at least 55%, at least 60%, at least 65%, at least 70% or more It contains one or more mutations that cause
[0252] In some embodiments, the mutant Cas protein is spCas9, spCas9-VRQR, spCas9 -VRER, xCas9 (sp), saCas9, saCas9-KKKH, spCas9-MQKSER, spCas9-LRKIQK, or spC An exemplary saCas9 sequence is provided below: KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHS ELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKD GEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKWEYEMLMGHCTYFP EELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKP EFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAIN LILDELWHTNDNQIAIFNRKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIQUINAIIKKYGLPNDIIIELARE KNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPR SVSFDNSFNNKVLVKQEE N SKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTTKKEYLLEERDINRFSVQKDF INRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKKERNKGYKHHAEDALIIANADFIFKEWKKL DKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLI VNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIK YYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAE FIASFYNNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKKYSTDILGNLYEV KSKKHPQIIKKG. In the saCas9 sequence above, the bolded and underlined residue N579 was mutated (e.g., to A579). This can produce SaCas9 nickase.
[0253] Exemplary SaCas9n sequences are provided below: KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTHS ELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKD GEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPGFGWKDIKEWYEMLMGHCTYFP EELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKP EFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAIN LILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELARE KNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPR SVSFDNSFNNKVLVKQEE A SKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDF INRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKL DKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLI VNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIK YYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAE FIASFYNNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKKYSTDILGNLYEV KSKKHPQIIKKG. In the above saCas9n sequence, N579 can be mutated to generate SaCas9 nickase. The corresponding residue A579 is underlined in bold.
[0254] The sequence of an exemplary SaKKH Cas9 is provided below:
[0255] KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHS ELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKD GEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKWEYEMLMGHCTYFP EELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKP EFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAIN LILDELWHTNDNQIAIFNRKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIQUINAIIKKYGLPNDIIIELARE KNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPR SVSFDNSFNNKVLVKQEE A SKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTTKKEYLLEERDINRFSVQKDF INRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKKERNKGYKHHAEDALIIANADFIFKEWKKL DKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNR K LINDTLYSTRKDDKGNTLI VNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIK YYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAE FIASFY K NDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPP H IIKTIASKTQSIKKYSTDILGNLYEV KSKKHPQIIKKG. Residue A579 can be mutated from N579 to generate SaCas9 nickase. , bold and underlined. Residues K781, K967, and H1014 listed above are the same as E781, N967, and and R1014, resulting in SaKKH Cas9, are underlined and italicized. is shown.
[0256] In some embodiments, the modified Cas9 is a high-fidelity Cas9 enzyme. In this study, the high-fidelity Cas9 enzyme was selected from SpCas9(K855A), eSpCas9(1.1), SpCas9-HF1, or high-fidelity Cas9. The modified Cas9eSpCas9(1.1) is a precision Cas9 variant (HypaCas9) that binds the HNH / RuvC group and non- It contains alanine substitutions that weaken the interaction between the target DNA strand and the target DNA strand, preventing strand splitting at off-target sites. Similarly, SpCas9-HF1 disrupts the interaction of Cas9 with the DNA phosphate backbone, preventing cleavage and fragmentation. HypaCas9 reduces off-target editing through alanine substitutions that alter the proofreading of Cas9. and contains mutations in the REC3 domain that increase target discrimination (SpCas9 N692A / M694A / Q695A / H698 A) All three high-fidelity enzymes cause fewer off-target edits than wild-type Cas9. An example of a high-fidelity Cas9 is shown in Figure 1. High-fidelity Cas9 domain mutations to Cas9 are shown in bold and in bold. and underlined.
[0257] MDKKYSIGL A IGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEE VVDKGASAQSFIERMT A FDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWG A LSRKLINGIRDKQSGKTILDFLKSDGFANRNFMA LIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETR A ITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD
[0258] [Guide polynucleotide] As used herein, the term "guide polynucleotide" refers to a polynucleotide that is specific to a target sequence. and polynucleotide-programmable nucleotide binding domain proteins (e.g., In one embodiment, the term "polynucleotide" refers to a polynucleotide that can form a complex with a gene encoding a gene encoding a gene encoding a gene (e.g., Cas9 or Cpf1). The guide polynucleotide is a guide RNA. As used herein, the term "guide RNA" "(gRNA)" and its grammatical equivalents are intended to mean a gene that can be specific for a target DNA and can act as a Cas protein. The RNA / Cas complex may represent an RNA that can form a complex with the target DNA. Cas9 / crRNA / tracrRNA can help guide the transcription of the target gene. Endonucleolytic cleavage of linear or circular dsDNA targets that are not complementary to the crRNA. The target strand is first cleaved by an endonuclease and then exonucleolytically cleaved in the 3'-5' direction. In nature, DNA binding and cleavage usually requires both a protein and RNA. However, it has been shown that both crRNA and tracrRNA aspects can be incorporated into a single RNA species. Alternatively, a single guide RNA ("sgRNA," or simply "gRNA") can be generated. For example, , Jinek M. et al., Science 337:816-821 (2012), the entire contents of which are incorporated by reference. Cas9 targets short motifs (PAM or PAM) in the CRISPR repeats. protospacer adjacent motif) to help distinguish "self" from "non-self" The sequence and structure of Cas9 nuclease are well known to those skilled in the art (see, for example, "Complementary te genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti, JJ et al., Natl. Acad. Sci. USA 98:4658-4663(2001); “CRISPR RNA maturation by t rans-encoded small RNA and host factor RNase III.” Deltcheva E. et al., Nature 471:602-607(2011); and “Programmable dual-RNA-guided DNA endonuclease in ad See Jinek M. et al., Science 337:816-821 (2012). The entire contents of which are incorporated herein by reference. Cas9 orthologs include, but are not limited to: Although not common, it has been described in various species, including S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on the present disclosure. Such Cas9 nucleases and sequences include those described by Chylinski, Rhun, and Charpenti er, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2 013) Cas9 sequences from the organisms and loci disclosed in RNA Biology 10:5, 726-737 In one embodiment, Cas9 The nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., Cas9 does not It's Kaze.
[0259] In some embodiments, the guide polynucleotide comprises at least one single guide RNA ("s In some embodiments, the guide polynucleotide is at least In some embodiments, the guide polynucleotide is a polynucleotide. Inject a programmable DNA-binding domain (e.g., Cas9 or Cpf1) into the target nucleotide sequence. It does not require a PAM sequence to guide it.
[0260] The polynucleotide programmable nucleotide sequences of the base editors disclosed herein The guide polynucleotide is associated with a domain-binding domain (e.g., a CRISPR-derived domain). The target polynucleotide sequence can be recognized by the guide polynucleotide ( The target sequence of the polynucleotide is targeted by a gene encoding a gene (e.g., a gRNA), which is typically single-stranded and binds to the target sequence in a site-specific manner. can be programmed to bind (i.e., via complementary base pairing) and Thus, the base editor with the guide nucleic acid is guided to the target sequence. The guide polynucleotide may be DNA. The guide polynucleotide may be RNA. The oligonucleotides include natural nucleotides (e.g., adenosine). Polynucleotides may contain non-natural (or unnatural) nucleotides (e.g., peptide nucleic acids or nucleotides). In some cases, the targeting region of the guide nucleic acid sequence may be short in length. At least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 The targeting region of the guide nucleic acid can be between 10 and 30 nucleotides in length, or can be between 15 and 25 nucleotides in length, or between 15 and 20 nucleotides in length.
[0261] In some embodiments, the guide polynucleotide is a polynucleotide that is capable of, e.g., complementary base pairing. It comprises two or more individual polynucleotides that can interact with each other via a duplex. For example, guide polynucleotides include CRISPR RNA (crRNA) and and trans-activating CRISPR RNA (tracrRNA). For example, a guide polynucleotide The reticle can include one or more trans-activating CRISPR RNAs (tracrRNAs).
[0262] In Type II CRISPR systems, targeting of nucleic acids by CRISPR proteins (e.g., Cas9) is typically Specifically, the first RNA molecule (crRNA) contains a sequence that recognizes the target sequence, and the guide RNA-CRISPR tag. A second RNA molecule (trRNA) containing repeat sequences that form a scaffolding region that stabilizes the protein complex Such a dual guide RNA system requires complementary base pairing between the and a guide polynucleotide for directing a base editor to a target polynucleotide sequence, as disclosed in the document. It can be used as a protease inhibitor.
[0263] In some embodiments, the base editors provided herein comprise a single guidepo In some embodiments, the present invention utilizes a nucleic acid sequence (e.g., a gRNA) as provided herein. The base editors used are dual guide polynucleotides (e.g., dual guide polynucleotides). In some embodiments, the base editors provided herein utilize: Some embodiments utilize one or more guide polynucleotides (e.g., multiple gRNAs). In the method, a single guide polynucleotide is selected from the group consisting of a plurality of different base editors described herein. For example, a single guide polynucleotide can be synthesized using a cytidine base editor. and adenosine base editors.
[0264] In other embodiments, the guide polynucleotide comprises a polynucleotide targeting portion of a nucleic acid and Both the nucleic acid and the nucleic acid scaffold portion can be contained in a single molecule (i.e., a single-molecule guide nucleic acid). For example, a single-molecule guide polynucleotide may be a single guide RNA (sgRNA or gRNA). As used herein, the term "guide polynucleotide sequence" refers to a base pair. any base editor that can interact with the base editor and guide the base editor to the target polynucleotide sequence Single, double or multi-molecule nucleic acids of the formula (I) are contemplated.
[0265] Typically, the guide polynucleotide (e.g., crRNA / trRNA complex or gRNA) is A "polynucleotide" refers to a sequence that can recognize and bind to a polynucleotide sequence. and a base editor for polynucleotide programmable nucleotide sequences. A "protein binding segment" that stabilizes the guide polynucleotide within the binding domain component. In some embodiments, the polynucleotide target segment of the guide polynucleotide The ment recognizes and binds to a DNA polynucleotide, thereby editing the bases in the DNA. In other embodiments, the polynucleotide target segment of the guide polynucleotide is The nucleotides recognize and bind to RNA polynucleotides, thereby facilitating the editing of bases in the RNA. Here, a "segment" refers to a portion or region of a molecule, such as a guide polynucleotide. A segment is a stretch of consecutive nucleotides in a nucleic acid. A segment can also represent a region / section, and thus can include regions of more than one molecule. For example, if the guide polynucleotide comprises multiple nucleic acid molecules, the protein binding segments may be The complement may be, for example, all or part of a plurality of separate molecules hybridized along regions of complementarity. In some embodiments, the targeting of the DNA-targeting RNA comprises two separate molecules. The protein-binding segment is: (i) a first RNA molecule that is 100 base pairs in length and spans base pairs 40-75 of the first RNA molecule; and (ii) 10 to 25 base pairs of a second RNA molecule that is 50 base pairs in length. The definition of "base pair" refers to a specific number of base pairs out of the total number of base pairs, unless otherwise defined in a particular context. is not limited to a specific number of base pairs from a given RNA molecule, but to distinct sequences within the complex. It is not limited to a particular number of molecules and can include regions of any full length RNA molecule, and can be used to identify regions of interest relative to other molecules. It may contain regions of complementarity.
[0266] A guide RNA or guide polynucleotide is a combination of two or more RNAs, such as CRISPR RNA (crRNA ) and trans-activating crRNA (tracrRNA). The polynucleotide sometimes comprises a fusion of portions (e.g., functional portions) of the crRNA and tracrRNA. The guide RNA or single-stranded RNA formed by the Alternatively, the guide polynucleotide may be a duplex RNA comprising a crRNA and a tracrRNA. Furthermore, the crRNA can hybridize with the target DNA.
[0267] As mentioned above, the guide RNA or guide polynucleotide can be an expression product. For example, the DNA encoding the guide RNA is a vector containing a sequence encoding the guide RNA. The guide RNA or guide polynucleotide is an isolated guide RNA or guide Transfect cells with plasmid DNA containing the RNA coding sequence and promoter The guide RNA or guide polynucleotide can be introduced into cells by The peptide can also be introduced into cells by other methods, such as using viral-mediated gene delivery.
[0268] The guide RNA or guide polynucleotide can be isolated. For example, the guide RNA The guide RNA can be transfected into a cell or organism in the form of isolated RNA. It can be prepared by in vitro transcription using any in vitro transcription system known in the art. The guide RNA can be delivered in the form of a plasmid containing the guide RNA coding sequence. , can be introduced into cells in the form of isolated RNA.
[0269] A guide RNA or guide polynucleotide may comprise the following three regions: a chromosomal sequence; a first region at the 5' end that may be complementary to a target site in the a second internal region that can be used to generate a fusion protein, and a third 3' region that can be single-stranded. The first region of each guide RNA can be different to guide the gene to a specific target site. Furthermore, the second and third regions of each guide RNA are identical in all guide RNAs. It could be.
[0270] The first region of the guide RNA or guide polynucleotide is It is complementary to a sequence at the target site in the chromosomal sequence so that it can form base pairs with the target site. In some cases, the first region of the guide RNA may be about 10 nucleotides to 25 nucleotides. nucleotides (i.e., from 10 nucleotides to nucleotides; or from about 10 nucleotides to about 25 nucleotides) or from 10 nucleotides to about 25 nucleotides; or about 10 nucleotides For example, the first nucleotide of a guide RNA may be 25 nucleotides or more. The region of base pairing between the region and the target site in the chromosomal sequence is about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25 or more nucleotides The first region of the guide RNA may be at or about the length of the first region of the guide RNA. may be 9, 20, or 21 nucleotides, or about 19, 20, or 21 nucleotides .
[0271] The guide RNA or guide polynucleotide may include a second region that forms a secondary structure. For example, the secondary structure formed by the guide RNA can be a stem (or hairpin ) and loops. The lengths of the loops and stems can vary. For example, For example, the loop length can range from (approximately) 3 to 10 nucleotides, and the stem length can range from (approximately) 6 to 10 nucleotides. The stem may range from 1 to 10 or about 10 nucleotides. The total length of the second region can range from about 16 to 60 nucleotides in length. For example, the loop can be (about) 4 nucleotides in length and the stem can be (about) 12 nucleotides. It may be a base pair.
[0272] The guide RNA or guide polynucleotide may be essentially single-stranded. A third region may also be included. For example, the third region may be any staining of the cells of interest. It may not be complementary to the target sequence or the rest of the guide RNA. Furthermore, the length of the third region can vary. The third region can be approximately 4 nucleotides in length. For example, the length of the third region can range from (about) 5 to 60 nucleotides in length. could be.
[0273] The guide RNA or guide polynucleotide can be inserted into any exon or intron of the gene target. In some cases, the guide targets exon 1 or 2 of the gene. in other cases, the guide targets exon 3 or 4 of the gene. The composition can contain multiple guide RNAs that all target the same exon, or In some cases, the gene can contain multiple guide RNAs that target different exons. Both exons and introns of a gene can be targeted.
[0274] The guide RNA or guide polynucleotide targets a nucleic acid sequence of (approximately) 20 nucleotides. The target nucleic acid may be less than (about) 20 nucleotides in length. At most (approximately) 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, or 1 to 100 The target nucleic acid may be at most about 5, 10, 15, 16, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170 6, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50 nucleotides, or 1 to 100 nucleotides The target nucleic acid sequence can be any length between the nucleotides. The guide RNA may target a nucleic acid sequence. The target nucleic acid may be: At least (approximately) 1-10, 1-20, 1-30, 1-40, 1-50, 1-60, 1-70, 1-80, 1-90, Or it may be 1 to 100 nucleotides.
[0275] The guide polynucleotide, e.g., guide RNA, is a nucleic acid sequence that encodes a nucleic acid sequence that is encoded by another nucleic acid, e.g., a nucleic acid sequence that is encoded by ... It can refer to a nucleic acid that can hybridize to a target nucleic acid or a protospacer. The guide polynucleotide can be RNA. The guide polynucleotide can be DNA. The polynucleotide is programmed or designed to bind site-specifically to a sequence of nucleic acid. The guide polynucleotide may comprise one polynucleotide strand, A guide polynucleotide may be a sequence of two polynucleotides. The guide RNA may comprise a double-stranded polynucleotide, which may be referred to as a dual-guide polynucleotide. For example, RNA molecules can be transcribed in vitro and / or synthesized. RNA can be synthesized chemically from synthetic DNA molecules, e.g., gBlocks® gene fragments. Guide RNAs can be introduced into cells or embryos as RNA molecules. Guide RNAs can also be It can be introduced into a cell or embryo in the form of a non-RNA nucleic acid molecule, such as a DNA molecule. For example, a guide RNA DNA encoding A is provided that encodes a promoter for expression of the guide RNA in the cells or embryo of interest. The RNA coding sequence may be operably linked to a regulatory sequence that encodes RNA polymerase III (Pol III). ) can be operably linked to a promoter sequence recognized by Plasmid vectors that can be used for this purpose are the px330 vector and the px333 vector. In some cases, a plasmid vector (e.g., px333 vector) may be used. The vector can comprise a DNA sequence encoding at least two guide RNAs.
[0276] Select and design guide polynucleotides (e.g., guide RNAs) and targeting sequences Methods for verifying and validating the identity of the gene are described herein and known to those skilled in the art. For example, Potential substrates of deaminase domains (e.g., AID domains) in nucleobase editor systems To minimize the effects of promiscuity, we selected nucleotides that were inadvertently targeted for deamination. residues that may be off-target (e.g., off-target Cs that may potentially be located on ssDNA within the target nucleic acid locus) Furthermore, software tools can be used to minimize the number of target residues. The gRNA corresponding to the nucleic acid sequence can be optimized, e.g., to maximize the total off-target activity across the genome. For example, targeting domains using S. pyogenes Cas9 can minimize targeting activity. For each possible choice, a PAM (preceding the selected PAM, e.g., NAG or NGG) is All containing up to a number (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) of mismatched base pairs Off-target sequences of gR complementary to the target site can be identified throughout the genome. The primary regions of the RNA can be identified and all primary regions (e.g., crRNA) can be analyzed by their total predicted off-target region. Targeting domains can be ranked according to their target score; the top-ranked targeted domains The compound was selected to have the greatest on-target activity and the least off-target activity. Candidate targeting gRNAs can be prepared using methods known in the art and / or described herein. can be functionally evaluated using
[0277] As a non-limiting example, a target DNA fragment in a guide RNA crRNA for use with Cas9 can be used. Hybridizing sequences can be identified using DNA sequence searching algorithms. The total is from Bae S., Park J., & Kim J.-S. Cas-OFFinder: A fast and versatile algorithm. that searches for potential off-target sites of Cas9 RNA-guided endonucleases. B The public tool cas-offinder, as described in ioinformatics 30, 1473-1475 (2014), This can be done using custom gRNA design software based on this software. scores guides after calculating the genome-wide off-target propensity of the guide. Matches typically range from a perfect match to seven mismatches, with lengths ranging from 17 to 24. Once the off-target sites are computationally determined, An aggregate score is calculated for each guide and output in a tabular format using a web interface. In addition to identifying potential target sites adjacent to the PAM sequence, The software identifies all sequences that differ from the selected target site by 1, 2, 3, or more than 3 nucleotides. The PAM flanking sequences are also identified. The repeat elements were then identified using publicly available tools, such as the RepeatMasker program. RepeatMasker can identify repeats from an input DNA sequence. Finds repeats and low-complexity regions. Detailed annotation of repeats present in a given query sequence. The output is:
[0278] After identification, the first region of the guide RNA (e.g., crRNA) is identified based on its distance to the target site, orthogonality and close match to related PAM sequences. Presence of protease inhibitors (e.g., associated PAMs (e.g., NGG PAM for Streptococcus pyogenes, Close matches in the human genome containing NNGRRT or NNGRRV PAM for Staphylococcus aureus The term "genetic variants" can be ranked hierarchically based on the identification of the 5'G. Orthogonality, when used in this context, refers to the ability to generate sequences in the human genome that contain a minimal number of mismatches to the target sequence. "High level of orthogonality" or "good orthogonality" refers to, for example, the number of sequences that are orthogonal to the intended sequence. It has no identical sequences in the human genome other than the target and has one or two mutations in the target sequence. It can also refer to a 20-mer targeting domain that has no sequences containing a match. The targeting domain can be selected to minimize off-target DNA cleavage.
[0279] In some embodiments, detecting base editing activity and identifying candidate guide polynucleotides. A reporter system can be used to test for nucleotides. In the reporter system, base editing activity results in expression of a reporter gene, It can include reporter gene-based assays. For example, the reporter system can be , including an inactive start codon, e.g., a mutation of 3'-TAC-5' to 3'-CAC-5' on the template strand Upon successful deamination of the target C, the corresponding mRNA is converted to 5'-GUG- The reporter gene is transcribed as 5'-AUG-3' in...
Claims
1. 1. An in vitro or ex vivo method of producing cells for treating sickle cell disease or beta thalassemia in a subject in need thereof, said method comprising: contacting the cell with a base editor or an mRNA encoding same and a guide RNA (gRNA), wherein said base editor comprises a polynucleotide-programmable nucleotide-binding domain and an adenosine deaminase domain or a cytidine deaminase domain; Binding the gRNA together with the base editor to a target nucleotide sequence in the cell located in a regulatory region of the HBG1 and / or HBG2 gene; editing the nucleobase of the target nucleotide sequence by deaminating the nucleobase when the gRNA binds to the target nucleotide sequence; thereby producing cells for treating sickle cell disease or beta thalassemia by changing said nucleobase to another nucleobase, wherein said nucleobase is located within c. -114 to -102 of said HBG1 and / or HBG2 gene, and / or said target nucleotide sequence comprises a nucleotide sequence selected from CTTGACCAATAGCCTTGACAAGG (SEQ ID NO:238; gRNA1), GCTATTGGTCAAGGCAAGGC (SEQ ID NO:240; gRNA4), and AAGTTGCCTTGTCAAGGCTATTGGT (SEQ ID NO:245; gRNA45).
2. 2. The method of claim 1, wherein said polynucleotide-programmable nucleotide-binding domain is a Cas9 nickase.
3. 3. The method of Claim 2, wherein said Cas9 nickase comprises a D10A amino acid substitution as numbered according to SEQ ID NO:
47.
4. The adenosine deaminase has the following amino acid sequence: SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (positions 2-167 of SEQ ID NO: 152) 2. The method of claim 1, wherein the TadA deaminase comprises an amino acid sequence that is at least 85% identical to
5. The method of claim 1 , wherein the cell is a mammalian cell.
6. The method of claim 5, wherein the cells are CD34+ cells.
7. 2. The method of claim 1, wherein the gRNA comprises a nucleotide sequence selected from CUUGACCAAUAGCCUUGACA (SEQ ID NO: 176; gRNA1), GCUAUUGGUCAAGGCAAGGC (SEQ ID NO: 177; gRNA4), AAGUUUGCCUUGUCAAGGCU (SEQ ID NO: 223; gRNA45), GCCUUGACAAGGCAAACUUG (SEQ ID NO: 224), UUGACAAGGCAAACUUGACC (SEQ ID NO: 225), CAAGGCUAUUGGUCAAGGCA (SEQ ID NO: 209), CUUGUCAAGGCUAUUGGUCA (SEQ ID NO: 210), GUUUGCCUUGUCAAGGCUAU (SEQ ID NO: 211), UGGUCAAGUUUGCCUUGUCA (SEQ ID NO: 212), CUUGCCUUGACCAAUAGCCU (SEQ ID NO: 217), and UAGCCUUGACAAGGCAAACU (SEQ ID NO: 218).
8. The method of claim 1, wherein the first and last three bases of the gRNA are 2'OMe modified.
9. 2. The method of claim 1, wherein the mRNA encoding the base editor is an N1MePseudoU-modified mRNA.
10. 10. The method of claim 1, wherein the subject comprises a beta-globin (HBB) protein associated with sickle cell disease.
11. 11. The method of claim 10, wherein the sickle cell disease-associated HBB protein has a valine at amino acid position 7 as numbered according to SEQ ID NO:
37.
12. The method of claim 1, wherein said editing alters the binding pattern of at least one protein to the promoter of said HBG1 and / or HBG2 gene.
13. 1. A composition for use in a method of treating sickle cell disease or beta thalassemia in a subject in need thereof, comprising: The composition comprises cells produced by the method of claim 1, The method of treating comprises administering the cells to the subject. composition.
14. 1. An in vitro or ex vivo method of producing cells for treating sickle cell disease or beta thalassemia in a subject in need thereof, said method comprising: contacting the cell with a base editor or an mRNA encoding same and a guide RNA (gRNA), wherein said base editor comprises a polynucleotide-programmable nucleotide-binding domain and an adenosine deaminase domain; Binding the gRNA together with the base editor to a target nucleotide sequence in the cell located in a regulatory region of the HBG1 and / or HBG2 gene; editing the nucleobase of the target nucleotide sequence by deaminating the nucleobase when the gRNA binds to the target nucleotide sequence; thereby producing cells for treating sickle cell disease or beta thalassemia by changing said nucleobase to another nucleobase, wherein said nucleobase is located within c. -114 to -102 of said HBG1 and / or HBG2 gene, and / or said target nucleotide sequence comprises a nucleotide sequence selected from CTTGACCAATAGCCTTGACAAGG (SEQ ID NO:238; gRNA1), GCTATTGGTCAAGGCAAGGC (SEQ ID NO:240; gRNA4), and AAGTTGCCTTGTCAAGGCTATTGGT (SEQ ID NO:245; gRNA45).
15. 15. The method of claim 14, wherein the polynucleotide-programmable nucleotide-binding domain is a Cas9 nickase.
16. The method of claim 14, wherein the first and last three bases of the gRNA are 2'OMe modified.
17. 15. The method of claim 14, wherein the mRNA encoding the base editor is an N1MePseudoU-modified mRNA.
18. 1. An in vitro or ex vivo method of producing cells for treating sickle cell disease or beta thalassemia in a subject in need thereof, said method comprising: contacting the cell with a base editor or an mRNA encoding the same and a guide RNA (gRNA), wherein the base editor comprises a Cas9 nickase and adenosine deaminase domain; Binding the gRNA together with the base editor to a target nucleotide sequence in the cell located in a regulatory region of the HBG1 and / or HBG2 gene; editing the nucleobase of the target nucleotide sequence by deaminating the nucleobase when the gRNA binds to the target nucleotide sequence; thereby producing a cell for treating sickle cell disease or beta thalassemia by changing said nucleobase to another nucleobase, wherein said gRNA comprises the nucleotide sequence CUUGACCAAUAGCCUUGACA (SEQ ID NO: 176).
19. 19. The method of claim 18, wherein the first and last three bases of the gRNA are 2'OMe modified.
20. 19. The method of Claim 18, wherein the mRNA encoding the base editor is an N1MePseudoU-modified mRNA.
21. A guide RNA comprising the nucleic acid sequence CUUGACCAAUAGCCUUGACA (SEQ ID NO: 176).
22. 22. The guide RNA of claim 21, further comprising the nucleic acid sequence GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (SEQ ID NO: 246).
23. 23. The guide RNA of claim 22, wherein the first and last three bases of the guide RNA are phosphorothioate and 2'OMe modified.
24. 1. A protein-nucleic acid complex comprising a base editor, i) a polynucleotide-programmable base editor comprising a nucleotide-binding domain and an adenosine deaminase domain; and ii) A guide RNA according to any one of claims 21 to 23. A protein-nucleic acid complex comprising:
25. 24. A composition for use in producing cells for treating sickle cell disease or beta-thalassemia in a subject in need thereof, the composition comprising the guide RNA of any one of claims 21-23 and an mRNA encoding a base editor, wherein the base editor comprises a Cas9 nickase and adenosine deaminase domain.