Deaminase, fusion protein, nucleic acid, pharmaceutical composition and use thereof
By developing efficient deaminases and fusion proteins, the limitations of existing gene editing tools in terms of editing efficiency and editing windows are solved, and more efficient and flexible gene editing is achieved.
Patent Information
- Application Number
- PCT/CN2024/137235
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-04
- Filing Date
- 2024-12-05
- Publication Date
- 2025-06-12
AI Technical Summary
Existing gene editing tools have limitations in base editing efficiency and editing window width, making it difficult to meet complex gene editing needs.
A new deaminase and fusion protein has been developed, with high deaminase editing efficiency and a wide editing window, enabling specific deaminology of nucleotides in polynucleotide sequences.
More efficient gene editing is achieved, the scope of the editing window is expanded, and the specificity and flexibility of editing is improved.
Smart Images

Figure PCTCN2024137235-FTAPPB-I100001 
Figure PCTCN2024137235-FTAPPB-I100002 
Figure PCTCN2024137235-FTAPPB-I100003
Abstract
Description
Deaminase, fusion protein, nucleic acid, pharmaceutical composition and use thereof Technical Field
[0001] The present application relates to the field of gene editing, and specifically to deaminases, fusion proteins, nucleic acids, drug combinations and their applications. Background Art
[0002] The CRISPR-Cas system is an adaptive immune defense system developed by bacteria and archaea over a long period of evolution, used to combat invading viruses and foreign DNA. In 2016, David Liu's laboratory developed cytosine base editors (CBEs). Currently, discovering more suitable cytosine deaminases for the construction of base editors and expanding the gene editing toolbox has become a key research direction in gene editing. Summary of the Invention
[0003] The present application provides a deaminase, fusion protein, nucleic acid, drug combination, and application thereof. The deaminase provided herein has high deamination editing efficiency and a wide editing window, and the resulting fusion protein also has high editing efficiency and a wide editing window. Some deaminase fusion proteins provided herein have different editing window locations than existing fusion proteins.
[0004] One aspect of the present application provides a deaminase comprising an amino acid sequence having at least 50% sequence identity to any one of SEQ ID NOs: 1-6, 39-164, 297-312.
[0005] In some embodiments of the present application, the deaminase is capable of deaminating at least one nucleotide in a polynucleotide sequence.
[0006] In some embodiments of the present application, the deaminase is capable of deaminating at least one nucleotide in the target polynucleotide.
[0007] In some embodiments of the present application, the deaminase deaminates at least one nucleotide in a polynucleotide sequence.
[0008] In some embodiments of the present application, the deaminase deaminates at least one nucleotide in the target polynucleotide.
[0009] In some embodiments of the present application, the deaminase comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100% sequence identity to the sequence shown in any one of SEQ ID NOs: 1-6, 39-164, 297-312.
[0010] In some embodiments of the present application, the deaminase comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100% sequence identity to the sequence shown in any one of SEQ ID NO: 1, 3-6, 43, 46, 48, 51, 63, 114, 126, 160, 163, 164, 297, 299 and 301.
[0011] In some embodiments of the present application, the deaminase comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100% sequence identity to the sequence shown in any one of SEQ ID NOs: 3, 5, 6, 126, 160, 163, 164, 297, 299 and 301.
[0012] In some embodiments of the present application, the deaminase comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100% sequence identity to the sequence shown in SEQ ID NO:3.
[0013] In some embodiments of the present application, the deaminase comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100% sequence identity to the sequence shown in SEQ ID NO:5.
[0014] In some embodiments of the present application, the deaminase comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100% sequence identity to the sequence shown in SEQ ID NO:6.
[0015] In some embodiments of the present application, the deaminase comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100% sequence identity to the sequence shown in SEQ ID NO: 126.
[0016] In some embodiments of the present application, the deaminase comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100% sequence identity to the sequence shown in SEQ ID NO: 160.
[0017] In some embodiments of the present application, the deaminase comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100% sequence identity to the sequence shown in SEQ ID NO: 163.
[0018] In some embodiments of the present application, the deaminase comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100% sequence identity to the sequence shown in SEQ ID NO: 164.
[0019] In some embodiments of the present application, the deaminase comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100% sequence identity to the sequence shown in SEQ ID NO: 297.
[0020] In some embodiments of the present application, the deaminase comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100% sequence identity to the sequence shown in SEQ ID NO: 299.
[0021] In some embodiments of the present application, the deaminase comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100% sequence identity to the sequence shown in SEQ ID NO: 301.
[0022] In some embodiments of the present application, the deaminase comprises mutations at positions corresponding to amino acid residues 122, 114, 176, 2 and / or 68 of the reference sequence shown in SEQ ID NO: 301.
[0023] In some embodiments of the present application, the deaminase comprises a mutation at the corresponding position of any one, two, three, four or five of the amino acid residues at positions 122, 114, 176, 2 and 68 of the reference sequence shown in SEQ ID NO: 301.
[0024] In some embodiments of the present application, the deaminase comprises a mutation at the corresponding position of any one, two, three, four or five of the amino acid residues at positions 122, 191, 114, 2 and 89 of the reference sequence shown in SEQ ID NO: 301.
[0025] Without limitation, the mutation may be a replacement, insertion, and / or deletion of an amino acid residue; further, the mutation may be a replacement of an amino acid residue. Without limitation, the corresponding positions can be identified by aligning the amino acid sequence of the deaminase with the reference sequence shown in SEQ ID NO: 301. In some embodiments of the present application, the deaminase comprises a mutation at a position corresponding to amino acid residue 122 of the reference sequence shown in SEQ ID NO: 301. In some embodiments of the present application, the deaminase comprises a mutation at a position corresponding to amino acid residue 191 of the reference sequence shown in SEQ ID NO: 301. In some embodiments of the present application, the deaminase comprises a mutation at a position corresponding to amino acid residue 114 of the reference sequence shown in SEQ ID NO: 301. In some embodiments of the present application, the deaminase comprises a mutation at a position corresponding to amino acid residue 2 of the reference sequence shown in SEQ ID NO: 301. In some embodiments of the present application, the deaminase comprises a mutation at a position corresponding to amino acid residue 176 of the reference sequence shown in SEQ ID NO: 301. In some embodiments of the present application, the deaminase comprises a mutation at a position corresponding to amino acid residue 68 of the reference sequence set forth in SEQ ID NO: 301. In some embodiments of the present application, the deaminase comprises a mutation at a position corresponding to amino acid residue 89 of the reference sequence set forth in SEQ ID NO: 301.
[0026] In some embodiments of the present application, the deaminase is mutated to H at the position corresponding to amino acid residue 122 of the reference sequence set forth in SEQ ID NO: 301. In some embodiments of the present application, the deaminase is mutated to R at the position corresponding to amino acid residue 191 of the reference sequence set forth in SEQ ID NO: 301. In some embodiments of the present application, the deaminase is mutated to T at the position corresponding to amino acid residue 114 of the reference sequence set forth in SEQ ID NO: 301. In some embodiments of the present application, the deaminase is mutated to M at the position corresponding to amino acid residue 2 of the reference sequence set forth in SEQ ID NO: 301. In some embodiments of the present application, the deaminase is mutated to D at the position corresponding to amino acid residue 89 of the reference sequence set forth in SEQ ID NO: 301. In some embodiments of the present application, the deaminase is mutated to F at the position corresponding to amino acid residue 176 of the reference sequence set forth in SEQ ID NO: 301. In some embodiments of the present application, the deaminase is mutated to I at the position corresponding to amino acid residue 68 of the reference sequence shown in SEQ ID NO: 301. In some embodiments of the present application, the deaminase is mutated to I at the position corresponding to amino acid residue 68 of the reference sequence shown in SEQ ID NO: 301, and the position corresponding to amino acid residue 114 is mutated to T.
[0027] In some embodiments of the present application, the deaminase comprises an amino acid sequence having 100% sequence identity with any one of SEQ ID NOs: 1-6, 39-164, 297-312.
[0028] In some embodiments of the present application, the deaminase comprises an amino acid sequence having 100% sequence identity with any one of SEQ ID NOs: 1, 3-6, 43, 46, 48, 51, 63, 114, 126, 160, 163, 164, 297, 299 and 301.
[0029] In some embodiments of the present application, the deaminase comprises an amino acid sequence as shown in any one of SEQ ID NOs: 1-6, 39-164, 297-312.
[0030] In some embodiments of the present application, the deaminase comprises an amino acid sequence shown in any one of SEQ ID NOs: 1, 3-6, 43, 46, 48, 51, 63, 114, 126, 160, 163, 164, 297, 299 and 301.
[0031] In some embodiments of the present application, the amino acid sequence of the deaminase is shown in any one of SEQ ID NOs: 1-6, 39-164, 297-312.
[0032] In some embodiments of the present application, the amino acid sequence of the deaminase is shown in any one of SEQ ID NO: 1, 3-6, 43, 46, 48, 51, 63, 114, 126, 160, 163, 164, 297, 299 and 301.
[0033] In some embodiments of the present application, the amino acid sequence of the deaminase is not any one of SEQ ID NOs: 1-5, 54-60.
[0034] In some embodiments of the present application, the deaminase is a Q122H, S191R, I114T, L2M or N89D mutant of deaminase-145 (amino acid sequence as shown in SEQ ID NO: 301).
[0035] In some embodiments of the present application, the deaminase can be used together with a nucleic acid binding polypeptide to deaminize bases in a nucleic acid sequence. For example, the deaminase can be used together with a TALEN or a ZFN to deaminize bases in a DNA sequence.
[0036] In some embodiments of the present application, the deaminase can form a complex with the nucleic acid-binding polypeptide, thereby deaminating the bases in the nucleic acid sequence.
[0037] In some embodiments of the present application, the deaminase is capable of deaminating bases in a target polynucleotide sequence together with a nucleic acid-binding polypeptide and a guide RNA (gRNA).
[0038] In some embodiments of the present application, the deaminase is capable of forming a complex with a nucleic acid-binding polypeptide and a guide RNA (gRNA), wherein the guide RNA guides the complex to bind to a target polynucleotide and deaminize bases in the target polynucleotide sequence.
[0039] The deaminase and the nucleic acid-binding polypeptide in the complex may be two separate molecules, or may be a single molecule formed by linking the deaminase and the nucleic acid-binding polypeptide (eg, forming a fusion protein).
[0040] As a non-limiting example, the deaminase can be used to replace the deaminase domain in the tBE base editing system described in patent documents WO2022206986A1 or WO2023155901A1, thereby obtaining a new base editing system.
[0041] In some embodiments, the deaminase comprises a homologous or heterologous domain. In some embodiments, the deaminase comprises a nucleic acid binding polypeptide. In some embodiments, the deaminase comprises a cellular localization signal; further, the deaminase comprises a nuclear localization signal, a nuclear export signal, a mitochondrial localization signal, or a chloroplast localization signal. In some embodiments, the deaminase comprises a UGI domain.
[0042] One aspect of the present application provides a fusion protein, comprising: the deaminase as described in the present application.
[0043] In some embodiments of the present application, the fusion protein comprises:
[0044] a) a nucleic acid binding polypeptide that binds to a target polynucleotide; and
[0045] b) a deaminase as described herein.
[0046] The fusion protein can be used for targeted editing of nucleic acids (e.g., DNA nucleic acids, RNA nucleic acids) in vitro, in vitro or in vivo, for example, for the production of mutant cells. These mutant cells can be in plants or animals. This fusion protein can also be used to introduce targeted mutations, for example, for correction of genetic defects in in vitro mammalian cells (e.g., cells obtained from a subject, which are subsequently reintroduced into the same or another subject); and for introducing targeted mutations, for example, correcting genetic defects or introducing mutations in disease-related genes in mammalian subjects. This fusion protein can also be used to introduce targeted mutations in plant cells, for example, for introducing beneficial or agriculturally valuable traits or alleles. The targeted editing, for example, deaminates at least one nucleotide in a nucleic acid.
[0047] Optionally, the nucleic acid binding polypeptide is a DNA binding polypeptide.
[0048] Optionally, the nucleic acid binding polypeptide is an RNA binding polypeptide.
[0049] In some embodiments of the present application, the nucleic acid binding polypeptide is a TALEN nuclease, a zinc finger nuclease, a CRISPR-Cas nuclease or a large nuclease.
[0050] In some embodiments of the present application, the nucleic acid binding polypeptide is an RNA-guided nuclease. Alternatively, the DNA binding polypeptide is a CRISPR-associated nuclease, i.e., a Cas enzyme (also known as a Cas protein, CRISPR-Cas nuclease).
[0051] In some embodiments of the present application, the nucleic acid binding polypeptide is selected from Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9, Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12f / CasZ, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csc3, a5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, C sb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas13a, Cas13b, Cas13c, Cas13d, Cas13e, Cas13f, TnpB, IscB, IsrB, and Fancor; or fragments thereof. Non-limiting examples of such fragments include nucleic acid binding domain fragments.
[0052] In some embodiments of the present application, the nucleic acid binding polypeptide is selected from Cas9, Cas12, Cas13, TnpB, IscB, IsrB, Fancor nuclease; or fragments thereof, including but not limited to nucleic acid binding domain fragments.
[0053] In some embodiments of the present application, the Cas9 is selected from SpCas9, SaCas9, Nme2Cas9, Nme3Cas9, CjCas9, NmCas9, FnCas9, PpnCas9, FrCas9, SauCas9, SauriCas9, ScaCas9, St1Cas9, BlatCas9, CdiCas9 and GeoCas9, fragments thereof, and mutants or fragments of mutants thereof.
[0054] In some embodiments of the present application, the nucleic acid binding polypeptide is selected from AsCpf1, enAsCas12a (addgene plasmid #196724), dFnCas12a (addgene plasmid #136379), ErCas12a, LbCas12a D832A, LbCas12a H759A, LbCas12a E795L, FnCas12a3, FnCas12a D917A, AsCas12a R1226A, AsCas12a D908A, AsCas12a E174R / S542R, AsCas12a(S542R / K548V / N552R), PrCas12a, PxCas12a, PcCas12a, PdCas12a, Mb2Cas12a, Mb3Cas12a, MlCas12a, CM aCas12a, CMtCas12a, HkCas12a, Lb5Cas12a, ErCas12a, TsCas12a, FnCpf1, LbCas12a, dLbCpf1, ttHsCas12a, AaCas12b, AaCas12b D570A, AaCas12b Q119F / E475R / E758R, BhCas12b, BvCas12b, BrCas12b, AkCas12b, AmCas12b, BsCas12b, OspCas12c, Cas12c2(addgene plasmid#183072), Cas12c_4(addgene plasmid#183071), Cas12c1(addgene plasmid #120872), CasY.1 (from Katanobacteria), CasY.2 (from Vogelbacteria), CasY.3 (from Vogelbacteria), CasY.4 (from Parcubacteria), CasY.5 (from Komeilibacteria), CasY.6 (from Kerfeldbacteria), PlmCasX, DpbCasX, Un1Cas12f, CnCas12f1, enRhCas12f1, AsCas12f1, SpaCas12f1, Cas12g1 (addgene plasmid #120879), Cas12h of WO2021113522A1 (SEQ ID NO: 1 in the patent), Cas12i1 (addgene plasmid #171670), Cas12i2 (addgene plasmid #188275), Cas12i1 (addgeneplasmid#120882)、Cas12i2(addgene plasmid#120883) and CN111757889B single-stranded Cas12f.4 / Cas12f.5 / Cas1 2f.6 Cas12i Processor, dSiCas12i(D1049A), SiCas12i, Si2Cas12i, WiCa s12i , Wi2Cas12i , Wi3Cas12i , SaCas12i , Sa2Cas12i , Sa3Cas12i , WaCas12i , Wa2Cas12i , xCas12i , hfCas12Max , Cas12i - Max ( addgene plasmid#188276); plasmid#188498); plasmid#181787)、Cas12k-TnsC(addgene plasmid#181789), Cas12l, MmCas12m, MmCas12mΔZF(H549A,C552A), d Cas12m-ΔZF(D485A,H549A,C552A), AcCas12n, dAcCas12n(D240), TnpB Actinomadura_cellulosilytica_strain_DSM_45823、TnpB Actinomadura_namibiensis_strain_DSM_44197、TnpB Actinomadura_umbrina_strain_DSM_43927_$、TnpB Actinoplanes_lobatus_strain_DSM_43150(TnpB-1and TnpB-2) 、TnpB Alicyclobacillus_macrosporagiidus_strain_DSM_17980 、TnpB Haloactinospora_alba_strain_DSM_45015 、TnpBLipingzhangella_halophila_strain_DSM_102030, TnpB Meiothermus_Silvanus_DSM_9946, TnpB QNFX01000004, ISDra2 TnpB (PDB:8H1J), KraIscB-1, AwaIscB, OgeuIscB, GtFz1 (from Guillardia theta), SpuFz1 (from Spizellomyces punctatus), NlovFz2 (from Percolozoa Naegleria lovaniensis) and MmeFz2 (from Mercenaria mercenaria), fragments thereof, and mutants or fragments thereof.
[0055] In some embodiments of the present application, the DNA binding polypeptide is capable of cleaving one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). Alternatively, the DNA binding polypeptide is an RNA-guided nuclease with nickase activity.
[0056] In some embodiments of the present application, the Cas enzyme cuts the target strand of the double-stranded nucleic acid molecule, which means that the Cas enzyme cuts the strand that is base-paired (complementary) to the gRNA (e.g., sgRNA) bound to the Cas enzyme.
[0057] In some embodiments of the present application, the nucleic acid-binding polypeptide is an RNA-guided nuclease with inactivated nuclease activity.
[0058] In some embodiments of the present application, the nucleic acid-binding polypeptide is a variant in which the nuclease activity of the wild-type RNA-guided nuclease is completely inactivated, such as a dCas enzyme, including but not limited to Cas9, Cas12, Cas13, TnpB, IscB, IsrB and Fancor with completely inactivated nuclease activity (which may be referred to as dead Cas9, dead Cas12, dead Cas13, dead TnpB, dead IscB, dead IsrB and dead Fancor); or fragments or mutants thereof, such as SpRY Cas9 variants.
[0059] In some embodiments of the present application, the nucleic acid-binding polypeptide is a variant in which the nuclease activity of the wild-type RNA-guided nuclease is partially inactivated, such as an nCas enzyme, for example, a polypeptide fragment that retains the nuclease activity of cutting a single strand of double-stranded DNA, including but not limited to nickase Cas9 and nickase Cas12.
[0060] In some embodiments of the present application, the fusion protein further comprises one, two, three, four, or more UGIs (uracil glycosylase inhibitors). In some embodiments of the present application, the fusion protein comprises one UGI. In some embodiments of the present application, the fusion protein comprises two UGIs. In some embodiments of the present application, the fusion protein comprises three UGIs. The UGIs can inhibit human UDG activity.
[0061] In some embodiments of the present application, the UGI comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the amino acid sequence of SEQ ID NO: 38. In some embodiments, the UGI comprises a fragment or homolog of the sequence of SEQ ID NO: 38.
[0062] In a preferred embodiment of the present application, the fusion protein comprises one or more NES (nuclear export signal), such as 1, 2, 3, 4 or more NES.
[0063] In a preferred embodiment of the present application, the fusion protein comprises one or more NLS (nuclear localization signals), such as 1, 2, 3, 4 or more NLS.
[0064] Nuclear localization signals are known in the art and typically comprise a stretch of basic amino acids (see, e.g., Lange et al., J. Biol. Chem. (2007) 282: 5101-5105). In certain embodiments, the fusion protein comprises 2, 3, 4, 5, 6 or more nuclear localization signals. One or more nuclear localization signals can be a heterologous NLS. Non-limiting examples of nuclear localization signals that can be used for the currently disclosed RGNs are the nuclear localization signals of SV40 large T antigen, nucleoplasmin, and c-Myc (see, e.g., Ray et al., (2015) Bioconjug Chem 26 (6): 1004-7). The fusion protein may comprise one or more NLS sequences at its N-terminus, C-terminus, or both the N-terminus and the C-terminus. For example, the fusion protein may comprise two NLS sequences in the N-terminal region and four NLS sequences in the C-terminal region.
[0065] Alternatively, the NLS comprises SPKKKRKVEAS (SEQ ID NO: 14), GPKKKRKVAAA (SEQ ID NO: 15), PKKKRKV (SEQ ID NO: 16), KRPAATKKAGQAKKKK (SEQ ID NO: 17), PAAKRVKLD (SEQ ID NO: 18), RQRRNELKRSP (SEQ ID NO: 19), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 20), RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 21), VSRKRPRP (SEQ ID NO: 22), PPKKARED (SEQ ID NO: 23), POPKKKPL (SEQ ID NO: 24), SALIKKKKKMAP (SEQ ID NO: 25), DRLRR (SEQ ID NO: 26). NO:26), PKQKKRK (SEQ ID NO:27), RKLKKKIKKL (SEQ ID NO:28), REKKKFLKRR (SEQ ID NO:29), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO:30), RKCLQAGMNLEARKTKK (SEQ ID NO:31) and PAAKKKKLD (SEQ ID NO:32).
[0066] In some embodiments of the present application, the nucleic acid binding polypeptide, the deaminase, the one or more UGIs and the one or more NLSs are directly connected or connected through a linker.
[0067] In some embodiments of the present application, when the nucleic acid binding polypeptide, the deaminase, the one or more UGIs and the one or more NLSs are connected by linkers, these linkers may be the same or different.
[0068] In some embodiments of the present application, the linker comprises a sequence as shown in any one of SEQ ID NOs: 33-37 and EFE, or a fragment or variant thereof.
[0069] In some embodiments of the present application, the linker comprises amino acids such as SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 33) and / or SGGSGGSGGS (SEQ ID NO: 34). In some embodiments of the present application, the deaminase and the nucleic acid-binding polypeptide are linked by SEQ ID NO: 33; and / or the nucleic acid-binding polypeptide and the UGI are linked by SEQ ID NO: 34.
[0070] In a preferred embodiment of the present application, the fusion protein comprises the following structures in order from the N-terminus to the C-terminus: a deaminase and a nucleic acid-binding polypeptide. Preferably, the fusion protein comprises the following structures in order from the N-terminus to the C-terminus: a deaminase, a linker 1, and a nucleic acid-binding polypeptide.
[0071] More preferably, the fusion protein comprises the following structures from N-terminus to C-terminus: deaminase, linker 1, nucleic acid binding polypeptide, linker 2, and UGI.
[0072] Further preferably, the fusion protein comprises the following structures from N-terminus to C-terminus: deaminase, linker 1, nucleic acid binding polypeptide, linker 2, UGI, and NLS.
[0073] Even more preferably, the fusion protein comprises the following structures from N-terminus to C-terminus: NLS, deaminase, linker 1, nucleic acid binding polypeptide, linker 2, UGI, NLS.
[0074] In a specific embodiment of the present application, the fusion protein comprises the following structures from N-terminus to C-terminus: SV 40NLS, deaminase, linker 1, nucleic acid binding polypeptide, linker 2, UGI, SV 40NLS.
[0075] The numbers in the connector 1 and the connector 2 are only used to distinguish nouns of the same type and do not indicate actual meanings.
[0076] In some embodiments of the present application, the linker 1 and / or linker 2 are optionally selected from the sequences shown in SEQ ID NO: 33-37 and EFE.
[0077] In some embodiments of the present application, the linker 1 and / or linker 2 is SEQ ID NO: 33.
[0078] In some embodiments of the present application, the linker 1 and / or linker 2 is SEQ ID NO: 34.
[0079] In the present application, the SV 40 refers to Simian vacuolating virus 40.
[0080] In some embodiments of the present application, the structure of the fusion protein is deaminase-XTEN linker-dead Cas.
[0081] In some embodiments of the present application, the structure of the fusion protein is deaminase-XTEN linker-dead Cas-UGI.
[0082] In some embodiments of the present application, the structure of the fusion protein is deaminase-XTEN linker-nickase Cas-UGI.
[0083] In some embodiments of the present application, the fusion protein is obtained by replacing the deaminase in YE1-BE3, YE2-BE3, EE-BE3, YEE-BE3 or HF-CBE3 in the prior art with the deaminase of the present application.
[0084] In some embodiments of the present application, the structure of the fusion protein is deaminase-XTEN linker-dead Cas-UGI-UGI.
[0085] In some embodiments of the present application, the structure of the fusion protein is deaminase-XTEN linker-nickase Cas-UGI-UGI.
[0086] In some embodiments of the present application, the structure of the fusion protein is Gam-deaminase-XTEN linker-nickase Cas9-UGI-UGI. The Gam is a Gam protein derived from bacteriophage Mu or an analog or homolog thereof.
[0087] In some embodiments of the present application, the fusion protein is capable of forming a complex with a guide RNA (gRNA), and the guide RNA guides the complex to bind to the target polynucleotide and deaminize the bases in the target polynucleotide sequence.
[0088] Another aspect of the present application provides an isolated nucleic acid encoding the deaminase as described herein or the fusion protein as described herein.
[0089] In some embodiments of the present application, the nucleic acid is codon-optimized, for example, prokaryotic codon optimization for a prokaryotic expression system, or eukaryotic codon optimization for a eukaryotic expression system.
[0090] In some embodiments of the present application, the isolated nucleic acid comprises a sequence as shown in any one of SEQ ID NOs: 165-296, 313-328. In some embodiments of the present application, the isolated nucleic acid sequence is as shown in any one of SEQ ID NOs: 165-296, 313-328.
[0091] Another aspect of the present application provides an expression vector comprising the nucleic acid described in the present application.
[0092] Those skilled in the art can select a suitable expression vector. For example, the expression vector may include but is not limited to a viral expression vector, a bacterial expression vector, a fungal expression vector, an animal expression vector (for example, insects, fruit flies, nematodes, fish, zebrafish, mammals, mice, rats, rabbits, pigs, monkeys, humans, etc.), a plant expression vector, etc.
[0093] In some embodiments of the present application, the expression vector comprises a regulatory sequence that is used to regulate the expression of the deaminase or fusion protein.
[0094] In some embodiments of the present application, the regulatory sequence is optionally selected from: one or more of a promoter, an enhancer, an internal ribosome entry site and a transcription termination signal; the promoter is, for example, a constitutive promoter, an inducible promoter, a broad-spectrum promoter or a tissue-specific promoter, and / or the transcription termination signal is, for example, a polyadenylation signal or a poly-U sequence.
[0095] In some embodiments, the expression vector comprises a pol III promoter (e.g., U6 and H1 promoters), a pol II promoter (e.g., the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with a RSV enhancer), a cytomegalovirus (CMV) promoter (optionally with a CMV enhancer), an SV40 promoter, a dihydrofolate reductase promoter, a β-actin promoter, a phosphoglycerol kinase (PGK) promoter, or an EF1α promoter.
[0096] In some embodiments, the promoter is a constitutive promoter, which is continuously active and not regulated by external signals or molecules. Suitable constitutive promoters include, but are not limited to, CMV, RSV, SV40, EF1α, CAG, and β-actin promoters. In some embodiments, the promoter is an inducible promoter regulated by external signals or molecules (e.g., transcription factors).
[0097] In some embodiments, the promoter is a tissue-specific promoter, which can be used to drive the tissue-specific expression of the deaminase or fusion protein. Suitable muscle-specific promoters include but are not limited to CK8, MHCK7, myoglobin promoter (Mb), desmin (Desmin) promoter, muscle creatine kinase promoter (MCK) and variants thereof, and SPc5-12 synthetic promoters. Suitable immune cell-specific promoters include but are not limited to B29 promoter (B cells), CD14 promoter (monocytes), CD43 promoter (leukocytes and platelets), CD68 (macrophages) and SV40 / CD43 promoter (leukocytes and platelets). Suitable blood cell-specific promoters include but are not limited to CD43 promoter (leukocytes and platelets), CD45 promoter (hematopoietic cells), INF-β (hematopoietic cells), WASP promoter (hematopoietic cells), SV40 / CD43 promoter (leukocytes and platelets), and SV40 / CD45 promoter (hematopoietic cells). Suitable pancreas-specific promoters include, but are not limited to, elastase-1 promoter. Suitable endothelial cell-specific promoters include, but are not limited to, Fit-1 promoter and ICAM-2 promoter. Suitable neuronal tissue / cell-specific promoters include, but are not limited to, GFAP promoter (astroglial cells), SYN1 promoter (neurons), and NSE / RU5' (mature neurons). Suitable kidney-specific promoters include, but are not limited to, NphsI promoter (podocytes). Suitable bone-specific promoters include, but are not limited to, OG-2 promoter (osteoblasts, odontoblasts). Suitable lung-specific promoters include, but are not limited to, SP-B promoter (lung). Suitable liver-specific promoters include, but are not limited to, SV40 / Alb promoter. Suitable heart-specific promoters include, but are not limited to, α-MHC.
[0098] Another aspect of the present application provides a pharmaceutical composition, which includes the deaminase as described herein, the fusion protein as described herein, the nucleic acid as described herein, or the expression vector as described herein, and a pharmaceutically acceptable delivery vehicle.
[0099] In some embodiments of the present invention, the delivery vector is optionally selected from: viral vectors, VLPs (virus like particles), nanoparticles such as lipid nanoparticles (LNP), liposomes, extracellular vesicles (EV) such as exosomes and microvesicles, and a gene gun.
[0100] In some embodiments of the present invention, the viral vector is an adeno-associated virus, an adenovirus, a lentivirus, or a herpes virus.
[0101] In some embodiments of the present invention, the adeno-associated virus is a recombinant adeno-associated virus vector of serotype AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV PHP.B, AAV PHP.B2, AAV PHP.B3, AAV PHP.A, AAV PHP.eB, AAV PHP.eS, AAV2.7m8, AAV8.7m8, AAV ShH10, AAVrh10 or AAVrh74.
[0102] In some embodiments of the present application, the deaminase, fusion protein, nucleic acid or expression vector described herein is packaged in an AAV vector, for example, packaged into an AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV PHP.B, AAV PHP.B2, AAV PHP.B3, AAV PHP.A, AAV PHP.eB, AAV PHP.eS, AAV2.7m8, AAV8.7m8, AAV ShH10, AAVrh10 or AAVrh74 capsid.
[0103] In some embodiments of the present application, the lentiviral vector is pseudotyped with an envelope protein. Preferably, the nucleic acid is linked to an aptamer sequence.
[0104] In some embodiments of the present application, in the virus-like particle, the nucleic acid is linked to a gene encoding a gag protein.
[0105] Another aspect of the present invention provides an expression system comprising an expression vector as described herein, a pharmaceutical composition as described herein, or a nucleic acid as described herein integrated into the genome. The expression system can be a host cell that can express a fusion protein as described herein, which can cooperate with a gRNA to localize the fusion protein to a target polynucleotide, thereby achieving base editing of the target polynucleotide.
[0106] In some embodiments of the present application, the host cell may be a eukaryotic cell and / or a prokaryotic cell. The eukaryotic cell includes fungi, plant cells, animal cells, such as mammalian cells, such as mouse cells and human cells, and may be, for example, mouse brain neuroma cells, human embryonic kidney cells, human cervical cancer cells, human colon cancer cells, or human osteosarcoma cells, and more specifically, may be N2a cells, HEK293FT cells, Hela cells, HCT116 cells, or U2OS cells.
[0107] Another aspect of the present application provides a system for modifying a target polynucleotide, the system comprising:
[0108] a) one or more guide RNAs (gRNAs) or one or more nucleic acids encoding the one or more guide RNAs, the one or more guide RNAs being capable of hybridizing to the target polynucleotide;
[0109] The system further includes any one or more of the following (a)-(f):
[0110] b) a deaminase as described herein, or a nucleic acid encoding the deaminase; and a nucleic acid-binding polypeptide, or a nucleic acid encoding the nucleic acid-binding polypeptide;
[0111] c) a fusion protein as described herein, or a nucleic acid encoding the fusion protein;
[0112] d) the expression vector as described in this application;
[0113] e) the pharmaceutical composition as described herein;
[0114] f) an expression system as described herein.
[0115] The target polynucleotide can also be called target nucleic acid.
[0116] In some embodiments of the present application, the deaminase in b) forms a complex with the nucleic acid-binding polypeptide, thereby deamining the bases in the nucleic acid sequence.
[0117] In some embodiments of the present application, b) and a) together form a complex, thereby deaminating the bases in the target polynucleotide sequence.
[0118] In some embodiments of the present application, the steps b), c), d), e) and / or f) are directed to the target polynucleotide by a) to deaminize the bases in the target polynucleotide sequence.
[0119] Another aspect of the present application provides a method for modifying a target polynucleotide, the method comprising contacting the target polynucleotide with a deaminase as described herein, a fusion protein as described herein, a nucleic acid as described herein, an expression vector as described herein, a pharmaceutical composition as described herein, an expression system as described herein, or a system for modifying a target polynucleotide as described herein.
[0120] In some embodiments of the present application, the expression vector or expression system comprises: a nucleotide sequence encoding the deaminase or fusion protein.
[0121] In some embodiments of the present application, the modification is that at least one nucleotide in the target polynucleotide is deaminated.
[0122] Optionally, the method is for non-therapeutic or diagnostic purposes.
[0123] On the basis of conforming to the common sense in this field, the above-mentioned preferred conditions can be arbitrarily combined to obtain the preferred embodiments of the present application.
[0124] The reagents and raw materials used in this application are commercially available.
[0125] The positive progress of this application is:
[0126] (1) Some deaminase fusion proteins of this application have high editing efficiency;
[0127] (2) The fusion proteins obtained from some deaminases in this application have a wide editing window;
[0128] (3) The fusion proteins obtained from some deaminases in this application have a narrow editing window. BRIEF DESCRIPTION OF THE DRAWINGS
[0129] Figure 1 is a plasmid map of CBE-27-pCDNA3.1.
[0130] Figure 2 is the Sanger analysis results of CBE-27 single-base editing in Example 2.
[0131] Figure 3 is the Sanger analysis results of CBE-29 single-base editing in Example 2.
[0132] Figure 4 is the Sanger analysis results of single-base editing of the NC group in Example 2.
[0133] Figure 5 is the Sanger analysis results of CBE-136 single-base editing in Example 3.
[0134] Figure 6 is the Sanger analysis results of CBE-140 single-base editing in Example 3.
[0135] Figure 7 is the Sanger analysis results of CBE-139 single-base editing in Example 4.
[0136] Figure 8 is the Sanger analysis results of CBE-145 single-base editing in Example 4.
[0137] Figure 9 is the Sanger analysis results of CBE-141 single-base editing in Example 4.
[0138] Figure 10 is the Sanger analysis results of CBE-143 single-base editing in Example 4.
[0139] Figure 11 shows the verification results of the single-base editor containing the deaminase-145 mutant in Example 5.
[0140] Figure 12 shows the verification results of the single-base editor containing the deaminase-145 mutant in Example 6. DETAILED DESCRIPTION
[0141] the term
[0142] Sequence identity
[0143] As used herein, the term "identity" or "percent identity" refers to the matching of sequences between two polypeptides or between two nucleic acids. When a certain position in the two sequences being compared is occupied by the same base or amino acid monomer subunit (e.g., a certain position in each of the two DNA molecules is occupied by adenine, or a certain position in each of the two polypeptides is occupied by lysine), then the molecules are identical at that position. The "percent sequence identity" between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions compared × 100%. For example, if 6 out of 10 positions in two sequences match, then the two sequences have 60% sequence identity. Typically, comparisons are made when two sequences are aligned to produce maximum sequence identity. Such comparisons can be made using published and commercially available alignment algorithms and programs, such as, but not limited to, CLUSTERalΩ, MAFFT, Probcons, T-Coffee, Probalign, BLAST, which can be reasonably selected by one of ordinary skill in the art. Those skilled in the art can determine appropriate parameters for aligning sequences, including, for example, any algorithms needed to achieve better alignment or optimal comparison over the entire length of the sequences being compared, as well as any algorithms needed to achieve better alignment or optimal comparison over a portion of the sequences being compared.
[0144] Deaminase
[0145] The term "deaminase" refers to an enzyme that catalyzes a deamination reaction (i.e., removes an amino group from a nucleotide base). In some embodiments, the deaminase is a cytidine deaminase. The cytidine deaminase catalyzes the deamination of cytidine to produce uridine, or catalyzes the deamination of deoxycytidine to produce deoxyuridine; that is, it catalyzes the deamination of a cytosine group to produce a uracil group. The deaminase can act on DNA and / or RNA.
[0146] The terms deaminase and deaminase polypeptide are used interchangeably.
[0147] In some embodiments, the deaminase may further comprise a homologous or heterologous domain. Non-limiting examples include: the deaminase further comprises a nucleic acid binding polypeptide (e.g., the deaminase comprises a Cas9 nickase domain, the deaminase comprises a Cas12 domain, the deaminase comprises a TnpB domain, or the deaminase comprises a Talen domain), the deaminase further comprises a cellular localization signal (e.g., a nuclear localization signal, a nuclear export signal, a mitochondrial localization signal, or a chloroplast localization signal), or the deaminase further comprises a UGI domain.
[0148] Deamination
[0149] In some embodiments, the deamination refers to catalyzing the deamination of bases in DNA and / or RNA molecules. In some embodiments, the deamination refers to catalyzing the deamination of bases in DNA molecules. In some embodiments, the deamination refers to catalyzing the deamination of bases in RNA molecules. In some embodiments, the deamination refers to catalyzing the deamination of cytidine to produce uridine, or catalyzing the deamination of deoxycytidine to produce deoxyuridine. In some embodiments, the deamination refers to catalyzing the deamination of a cytosine group to produce a uracil group.
[0150] In some embodiments, the base to be edited is upstream of the PAM site. In some embodiments, the base to be edited is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39 or 40 nucleotides upstream of the PAM site. In some embodiments, the base to be edited is downstream of the PAM site. In some embodiments, the base to be edited is the 1st, 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, 10th, 11th, 12th, 13th, 14th, 15th, 16th, 17th, 18th, 19th, 20th, 21st, 22nd, 23rd, 24th, 25th, 26th, 27th, 28th, 29th, 30th, 31st, 32nd, 33rd, 34th, 35th, 36th, 37th, 38th, 39th or 40th nucleotide downstream of the PAM site. The base to be edited is the base at which deamination occurs. The position of the base to be edited is the editing window. In some embodiments, the target sequence includes an editing window. In some embodiments, the editing window includes 1-10 nucleotides. In some embodiments, the length of the editing window is 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, 1-2 or 1 nucleotide. In some embodiments, the editing window is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length.
[0151] Fusion protein
[0152] In the present disclosure, when it is mentioned that "a fusion protein comprises XXXX deaminase" or similar descriptions are used, it is intended to mean that the amino acid sequence of the fusion protein comprises the amino acid sequence constituting the deaminase.
[0153] In the present disclosure, when it is mentioned that "a fusion protein comprises a XXXX nucleic acid binding polypeptide" or similar descriptions are used, it is intended to mean that the amino acid sequence of the fusion protein comprises the amino acid sequence constituting the nucleic acid binding polypeptide.
[0154] Proteins, peptides and polypeptides
[0155] The terms "protein," "albumin," "peptide," and "polypeptide" are used interchangeably herein to refer to a polymer of amino acid residues linked together by peptide (amide) bonds. The terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide is at least three amino acids in length. A protein, peptide, or polypeptide may refer to a single protein or a collection of proteins. One or more amino acids in a protein, peptide, or polypeptide may be modified, for example, by the addition of chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofamersyl groups, fatty acid groups, linkers for conjugation, functionalization, or other modifications, etc. A protein, peptide, or polypeptide may also be a single molecule or may be a multimolecular complex. A protein, peptide, or polypeptide may simply be a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be naturally occurring, recombinant, or synthetic, or any combination thereof.
[0156] Nucleic acid binding polypeptides
[0157] Optionally, the nucleic acid binding polypeptide is a DNA binding polypeptide.
[0158] Optionally, the nucleic acid binding polypeptide is an RNA binding polypeptide.
[0159] Optionally, the nucleic acid binding polypeptide is a TALEN nuclease, a zinc finger nuclease, a CRISPR-Cas nuclease or a large nuclease.
[0160] In some embodiments of the present application, the nucleic acid-binding polypeptide is an RNA-guided nuclease.
[0161] In some embodiments of the present application, the DNA-binding polypeptide is an RNA-guided nuclease.
[0162] In some embodiments of the present application, the RNA-binding polypeptide is an RNA-guided nuclease.
[0163] RNA-guided nucleases
[0164] The term RNA-guided nuclease (RGN) refers to a polypeptide that binds to a specific target polynucleotide in a sequence-specific manner, and the polypeptide is guided to the target polynucleotide by a guide RNA molecule that is complexed with the polypeptide and hybridized to the target polynucleotide. Although RNA-guided nucleases can cut the target sequence when bound, the term RNA-guided nucleases also include nucleases that can bind but not cut the target sequence. The cutting of the target sequence by the RNA-guided nuclease can result in single-stranded or double-stranded breaks. The nucleases that can only cut the single strand of a double-stranded nucleic acid molecule are referred to as nickases in this application.
[0165] Optionally, the DNA binding polypeptide is a CRISPR-associated nuclease, i.e., a Cas enzyme (also known as a Cas protein).
[0166] In some embodiments of the present application, the nucleic acid binding polypeptide is selected from Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9, Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12f / CasZ, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csc3, a5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, C sb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas13a, Cas13b, Cas13c, Cas13d, Cas13e, Cas13f, TnpB, IscB, IsrB, and Fancor; or fragments thereof. Non-limiting examples of such fragments include nucleic acid binding domain fragments.
[0167] In some embodiments of the present application, the nucleic acid binding polypeptide is selected from Cas9, Cas12, Cas13, TnpB, IscB, IsrB, Fancor nuclease; or fragments thereof, including but not limited to nucleic acid binding domain fragments.
[0168] In some embodiments of the present application, the Cas9 is selected from SpCas9, SaCas9, Nme2Cas9, Nme3Cas9, CjCas9, NmCas9, FnCas9, PpnCas9, FrCas9, SauCas9, SauriCas9, ScaCas9, St1Cas9, BlatCas9, CdiCas9 and GeoCas9, fragments thereof, and mutants or fragments of mutants thereof.
[0169] In some embodiments of the present application, the nucleic acid binding polypeptide is selected from AsCpf1, enAsCas12a (addgene plasmid #196724), dFnCas12a (addgene plasmid #136379), ErCas12a, LbCas12a D832A, LbCas12a H759A, LbCas12a E795L, FnCas12a3, FnCas12a D917A, AsCas12a R1226A, AsCas12a D908A, AsCas12a E174R / S542R, AsCas12a(S542R / K548V / N552R), PrCas12a, PxCas12a, PcCas12a, PdCas12a, Mb2Cas12a, Mb3Cas12a, MlCas12 a, CMaCas12a, CMtCas12a, HkCas12a, Lb5Cas12a, ErCas12a, TsCas12a, FnCpf1, LbCas12a, ttHsCas12a, AaCas12b, AaCas12b D570A, AaCas12b Q119F / E475R / E758R, BhCas12b, BvCas12b, BrCas12b, AkCas12b, AmCas12b, BsCas12b, OspCas12c, Cas12c2(addgene plasmid#183072), Cas12c_4(addgene plasmid#183071), Cas12c1(addgene plasmid #120872), CasY.1 (from Katanobacteria), CasY.2 (from Vogelbacteria), CasY.3 (from Vogelbacteria), CasY.4 (from Parcubacteria), CasY.5 (from Komeilibacteria), CasY.6 (from Kerfeldbacteria), PlmCasX, DpbCasX, Un1Cas12f, CnCas12f1, enRhCas12f1, AsCas12f1, SpaCas12f1, Cas12g1 (addgene plasmid #120879), Cas12h of WO2021113522A1 (SEQ ID NO: 1 in the patent), Cas12i1 (addgene plasmid #171670), Cas12i2 (addgene plasmid #188275), Cas12i1 (addgeneplasmid#120882)、Cas12i2(addgene plasmid#120883) and CN111757889B single-stranded Cas12f.4 / Cas12f.5 / Cas1 2f.6 Cas12i Processor, dSiCas12i(D1049A), SiCas12i, Si2Cas12i, WiCa s12i , Wi2Cas12i , Wi3Cas12i , SaCas12i , Sa2Cas12i , Sa3Cas12i , WaCas12i , Wa2Cas12i , xCas12i , hfCas12Max , Cas12i - Max ( addgene plasmid#188276); plasmid#188498); plasmid#181787)、Cas12k-TnsC(addgene plasmid#181789), Cas12l, MmCas12m, MmCas12mΔZF(H549A,C552A), d Cas12m-ΔZF(D485A,H549A,C552A), AcCas12n, dAcCas12n(D240), TnpB Actinomadura_cellulosilytica_strain_DSM_45823、TnpB Actinomadura_namibiensis_strain_DSM_44197、TnpB Actinomadura_umbrina_strain_DSM_43927_$、TnpB Actinoplanes_lobatus_strain_DSM_43150(TnpB-1and TnpB-2) 、TnpB Alicyclobacillus_macrosporagiidus_strain_DSM_17980 、TnpB Haloactinospora_alba_strain_DSM_45015 、TnpBLipingzhangella_halophila_strain_DSM_102030, TnpB Meiothermus_Silvanus_DSM_9946, TnpB QNFX01000004, ISDra2 TnpB (PDB:8H1J), KraIscB-1, AwaIscB, OgeuIscB, GtFz1 (from Guillardia theta), SpuFz1 (from Spizellomyces punctatus), NlovFz2 (from Percolozoa Naegleria lovaniensis) and MmeFz2 (from Mercenaria mercenaria), fragments thereof, and mutants or fragments thereof.
[0170] In some embodiments of the present application, the DNA binding polypeptide is capable of cleaving one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). Alternatively, the DNA binding polypeptide is an RNA-guided nuclease with nickase activity.
[0171] In some embodiments of the present application, the Cas enzyme cuts the target strand of the double-stranded nucleic acid molecule, which means that the Cas enzyme cuts the strand that is base-paired (complementary) to the gRNA (e.g., sgRNA) bound to the Cas enzyme.
[0172] In some embodiments of the present application, the nucleic acid-binding polypeptide is an RNA-guided nuclease with inactivated nuclease activity.
[0173] In some embodiments of the present application, the nucleic acid-binding polypeptide is a variant in which the nuclease activity of the wild-type RNA-guided nuclease is completely inactivated, such as a dCas enzyme, including but not limited to Cas9, Cas12, Cas13, TnpB, IscB, IsrB and Fancor in which the nuclease activity is completely inactivated (which may be referred to as dead Cas9, dead Cas12, dead Cas13, dead TnpB, dead IscB, dead IsrB and dead Fancor).
[0174] In some embodiments of the present application, the nucleic acid-binding polypeptide is a variant in which the nuclease activity of the wild-type RNA-guided nuclease is partially inactivated, such as an nCas enzyme, for example, a polypeptide fragment that retains the nuclease activity of cutting a single strand of double-stranded DNA, including but not limited to nickase Cas9 and nickase Cas12.
[0175] The target polynucleotide binds to the nucleic acid-binding polypeptide provided herein and hybridizes with the guide RNA associated with the nucleic acid-binding polypeptide. If the nucleic acid-binding polypeptide has nuclease activity, the target sequence can then be cleaved by the nucleic acid-binding polypeptide. For example, the nucleic acid-binding polypeptide can optionally cleave either or both strands of the double-stranded DNA, or can bind only to the target polynucleotide without cleaving either strand. Simultaneously or subsequently to the binding or cleavage of the target polynucleotide by the nucleic acid-binding polypeptide, a deaminase deaminates at least one base of the target polynucleotide.
[0176] In some embodiments, the binding is dependent on recognition of a PAM sequence by the RNA-guided nuclease. In some embodiments, the binding is independent of recognition of a PAM sequence by the RNA-guided nuclease. The PAM sequences recognized by many of the nucleases described herein are known.
[0177] The term "cleave" or "cleavage" refers to the hydrolysis of at least one phosphodiester bond within the backbone of a target polynucleotide, which can result in single-strand or double-strand breaks within the target sequence. The currently disclosed nucleic acid-binding polypeptides can act as endonucleases or can be exonucleases (removing consecutive nucleotides from the ends (5' and / or 3' ends) of a polynucleotide) to cleave nucleotides within a polynucleotide. In other embodiments, the disclosed nucleic acid-binding polypeptides can cleave nucleotides of a target sequence at any position within a polynucleotide, thus acting as both endonucleases and exonucleases. Cleavage of a target polynucleotide by the currently disclosed nucleic acid-binding polypeptides can result in staggered breaks or blunt ends.
[0178] Nuclear localization signal, nuclear localization sequence, NLS
[0179] The terms nuclear localization signal, nuclear localization sequence, and NLS are used interchangeably. The fusion protein currently disclosed may include at least one nuclear localization signal (NLS) to enhance delivery of the nucleic acid binding polypeptide to the cell nucleus. Nuclear localization signals are known in the art and typically include a stretch of basic amino acids (see, e.g., Lange et al., J. Biol. Chem. (2007) 282: 5101-5105). In a specific embodiment, the fusion protein includes 2, 3, 4, 5, 6 or more nuclear localization signals. One or more nuclear localization signals can be a heterologous NLS. Non-limiting examples of nuclear localization signals that can be used for the currently disclosed RGN are the nuclear localization signals of SV40 large T antigen, nucleoplasmin, and c-Myc (see, e.g., Ray et al., (2015) Bioconjug Chem 26 (6): 1004-7).
[0180] connector
[0181] The term "joint" as used herein refers to a chemical group or molecule that connects two molecules or parts (e.g., the binding domain and cleavage domain of a nuclease). Typically, a joint is located between or on both sides of two groups, molecules, or other parts, and is covalently bonded to each group, molecule, or other part, thereby connecting the two. In some embodiments, a joint is an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, a joint is an organic molecule, a group, a polymer, or a chemical moiety. In some embodiments, the length of the linker is 5-100 amino acids, for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated.
[0182] In some embodiments, the linker comprises a sequence as set forth in any one of SEQ ID NOs: 33-37 and EFE, a fragment or variant thereof.
[0183] In some embodiments, the linker comprises a (GGGGS)n (SEQ ID NO: 35), (G)n, (EAAAK)n (SEQ ID NO: 36), (XP)n, or (SGG)n motif, or a combination of any of these, wherein n is independently an integer between 1 and 30. In some embodiments, n is independently 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30, or, if more than one linker or more than one linker motif is present, any combination thereof. Other suitable linker motifs and linker configurations will be apparent to those skilled in the art. In some embodiments, suitable linker motifs and configurations include those described in Chen et al., Fusion protein linkers: property, design and functionality (Adv Drug Deliv Rev. 2013; 65(10): 1357-69, the entire contents of which are incorporated herein by reference). Other suitable linker sequences will be apparent to those skilled in the art based on this disclosure.
[0184] UGI
[0185] The fusion protein of the present application may contain a suitable UGI sequence, and additional suitable UGI sequences are known to those skilled in the art and include, for example, those disclosed in Wang et al., 1989 J. Biol. Chem. 264: 1163-1171; Lundquist et al., 1997 J. Biol. Chem. 272: 21408-21419; Ravishankar et al., 1998 Nucleic Acids Res. 26: 4880-4887; and Putnam et al., 1999 J. Mol. Biol. 287: 331-346 (1999), the entire contents of which are incorporated herein by reference.
[0186] Guide RNA, guide RNA, gRNA, guide polynucleotide
[0187] The terms guide RNA, guide RNA, gRNA, and guide polynucleotide are used interchangeably. In some embodiments of the present application, the guide RNA comprises a guide sequence and a backbone sequence. The backbone sequence interacts with the RNA-guided nuclease. The backbone sequence is the sequence that generally remains unchanged in the guide RNA molecule when designing the guide RNA molecule. For example, the backbone sequence may refer to the portion of the guide RNA molecule other than the guide sequence. For example, when the nucleic acid-binding polypeptide is Cas9, Cas12, or Cas13, the guide RNA may be a crRNA or a complex of crRNA and tracrRNA, for example, a single-molecule gRNA formed by a chimeric crRNA and tracrRNA.
[0188] In some embodiments of the present application, the guide sequence hybridizes with the target polynucleotide.Guide sequence refers to a continuous nucleotide sequence in gRNA, which has partial or complete complementarity with the target polynucleotide (specifically refers to partial or complete complementarity with the target sequence in the target polynucleotide [target nucleic acid]), and can be hybridized with the target sequence in the target polynucleotide by base pairing promoted by the nuclease guided by RNA. The full complementarity of the guide sequence of the present invention to the target sequence is not necessary, as long as there is enough complementarity to cause hybridization and promote the formation of a gene editing complex. As used in this application, the term target sequence refers to a short sequence in a target polynucleotide molecule, which can be complementary (completely complementary or partially complementary) to the guide sequence of the gRNA molecule. The gene editing complex is specifically located at the target sequence by the guide sequence sequence and performs the corresponding function at or near this position. The length of the target sequence is often tens of nt (nucleotides), for example, 10-60nt, 15-50nt, 15-40nt, 15-35nt, 20-35nt, 20-30nt, about 10nt, about 20nt, about 30nt, about 40nt, about 50nt, about 60nt.
[0189] In some embodiments of the present application, the guide sequence has at least 50% sequence identity with the target polynucleotide. In some embodiments of the present application, the guide sequence has at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95% or 100% sequence identity with the target polynucleotide.
[0190] The length of the guide sequence is often tens of nt (nucleotides), for example, can be 10-60nt, 15-50nt, 15-40nt, 15-35nt, 20-35nt, 20-30nt, about 10nt, about 20nt, about 30nt, about 40nt, about 50nt, about 60nt.
[0191] In some embodiments of the present application, the guide sequence is located at the 3' end or 5' end of the backbone sequence. In some embodiments of the present application, the guide sequence is located at the 3' end of the backbone sequence. In some embodiments of the present application, the guide sequence is located at the 5' end of the backbone sequence.
[0192] In some embodiments of the present application, the guide RNA comprises an aptamer sequence.
[0193] In some embodiments of the present application, the aptamer sequence is inserted into the loop of the stem-loop structure of the secondary structure of the guide RNA.
[0194] In some embodiments of the present application, the guide RNA comprises modified nucleotides. The modifications include but are not limited to 2'-O-methyl, 2'-O-methyl-3'-phosphorothioate or 2'-O-methyl-3'-thio PACE modifications. In some embodiments of the present application, the guide RNA comprises modified nucleotides selected from deoxyribonucleotides and locked nucleic acids (LNA). In some embodiments, the guide RNA comprises at least one chemically modified nucleotide.
[0195] In some embodiments, the guide RNA is a hybrid RNA-DNA guide, i.e., some RNA nucleotides in the guide RNA are replaced by DNA nucleotides. In some embodiments, the guide RNA is a hybrid RNA-LNA (locked nucleic acid) guide, i.e., some RNA nucleotides in the guide RNA are replaced by LNA nucleotides.
[0196] In some embodiments of the present application, the target RNA is located in the nucleus and / or cytoplasm of a eukaryotic cell.
[0197] Target polynucleotide
[0198] Depending on the context, the terms target polynucleotide, target nucleic acid, and target gene are used interchangeably.
[0199] The deaminase or fusion protein of the present invention can be used to deaminate bases in the target polynucleotide, catalyzing the deamination of cytidine to produce uridine, or catalyzing the deamination of deoxycytidine to produce deoxyuridine. It can be used to achieve base conversion on DNA molecules or to achieve base conversion on RNA molecules. For example, it can be used to edit the exon region, intron region and / or regulatory region (such as promoter region, enhancer region) of a gene; it can be used to edit pre-mRNA molecules, edit mature mRNA molecules, or edit non-coding RNA molecules.
[0200] Patent application publication number CN108513575A discloses some disease-causing mutations, which are incorporated herein by reference. These mutations can be corrected by the deaminase or fusion protein of the present invention, thereby treating the corresponding diseases.
[0201] The deaminase or fusion protein of the present invention can be used to edit the genes in Appendix 1 (normal gene copies or mutant gene copies), or their corresponding pre-mRNA molecules or mature mRNA molecules, thereby treating the corresponding diseases or conditions.
[0202] When referring to an RNA sequence, the "t" in the sequence is used interchangeably with the "u". When referring to a "guide sequence", the "t" in the sequence is used interchangeably with the "u".
[0203] The present invention is further described below by way of examples, but the present invention is not limited to the scope of the examples. In the following examples, the experimental methods without specific conditions are selected according to conventional methods and conditions or according to the product specifications.
[0204] Example
[0205] The inventors obtained deaminase-23 to deaminase-34 through bioinformatics analysis and screening.
[0206] In addition, deaminase-8 to deaminase-22 and deaminase-36 to deaminase-134 were obtained through structure-based de novo design.
[0207] Experimental Example 1 Synthesis of Different Cytosine Deaminase Expression Plasmids
[0208] The expression frame of the cytosine base editor composed of nSpRYCas9, UGI and different cytosine deaminase sequences was designed and synthesized (using the nucleic acid sequence encoding the deaminase shown in the sequence table), and the expression frame of the sgRNA targeting the VEGFA gene was synthesized at the same time. The two frameworks were then inserted into the pCDNA3.1 expression vector to obtain expression vector plasmids containing cytosine base editors of different deaminases. The number n of deaminase-n corresponds to the same number n in the editor CBE-n, and corresponds to the same number n in the vector plasmid CBE-n-pCDNA3.1. For example, deaminase-27 can be constructed to obtain the editor CBE-27, and its expression vector plasmid is CBE-27-pCDNA3.1. The following other embodiments also use the same representation method.
[0209] Amino acid sequence:
[0210] In this example, multiple cytosine deaminase proteins were screened and tested, and the amino acid sequences of some of the deaminases are as follows:
[0211] >Deaminase-27 (SEQ ID NO: 3)
[0212] >Deaminase-29 (SEQ ID NO: 5)
[0213] >Deaminase-48 (SEQ ID NO: 6)
[0214] >Deaminase-136 (SEQ ID NO: 57)
[0215] >Deaminase-140 (SEQ ID NO: 58)
[0216] Nucleic acid sequence
[0217] >CBE-27-pCDNA3.1 plasmid (SEQ ID NO: 10)
[0218] >CBE-29-pCDNA3.1 plasmid (SEQ ID NO: 12)
[0219] Note: The CBE-27-pCDNA3.1 and CBE-29-pCDNA3.1 vector sequences are specially annotated. The lowercase italicized portion encodes the cytosine base editor fusion protein coding sequence, the lowercase underlined portion encodes the deaminase coding sequence, and the uppercase bold underlined portion encodes the guide sequence of the gRNA targeting VEGFA.
[0220] The sequences of expression plasmids of cytosine base editors containing other deaminases are consistent with the sequences of deaminase-27 and deaminase-29 mentioned above, except that the deaminase coding sequences are replaced with their respective specific sequences (selected from the corresponding deaminase coding sequences in the sequence table).
[0221] An exemplary plasmid map of CBE-27-pCDNA3.1 is shown in FIG1 .
[0222] At the same time, the deaminase reported in the literature (Huang J et al. Discovery of deaminase functions by structure-based protein clustering. Cell. 2023 Jul 20; 186(15): 3182-3195.e14. doi: 10.1016 / j.cell.2023.05.041. Epub 2023 Jun 27. PMID: 37379837.) was used as a control. The coding sequence of deaminase-03 (numbered as Sdd5 in the literature) reported in the literature is as follows:
[0223] >CBE-03 (SEQ ID NO: 7)
[0224] The same method as above was used to construct an expression plasmid of the base editor fused with deaminase-03, namely CBE-03-pCDNA3.1, for subsequent control tests.
[0225] Experimental Example 2: Detection of editing efficiency of base editors composed of different cytosine deaminases
[0226] a. Cell plating: Digest 293T cells to 70%-80% confluency and plate them. Seed cells at a rate of 5 x 10^5 cells / well in a 24-well plate.
[0227] b. Transfection: 12-14 hours after plating, perform transfection. Add 1.5 μl of PEI (Yisheng Bio, 40815ES03, Polyethylenimine Linear (PEI) MW25000) to 100 μl of Opti-MEM per well of a 24-well plate. Add 500 ng of each cytosine base editor plasmid from Example 1, mix well, and incubate at room temperature for 20 minutes before adding the cells to 293T cells for transfection. Replace the culture medium with fresh DMEM + 10% FBS overnight and continue culturing. A negative control group (NC) was not transfected with plasmid.
[0228] c. Cell lysis, PCR amplification near the editing region, and Sanger sequencing: After 72 h of culture, cells were washed with PBS and then 100 μl of cell lysis buffer (Viagen, 302-C, Lysis Reagent (Cell) was used to lyse the cells to obtain a lysis solution containing genomic DNA. The region near the genomic DNA target site was amplified, and the PCR product was subjected to Sanger sequencing.
[0229] d. Data Analysis: Submit the sequencing peaks and sgRNA guide sequences to the editing efficiency analysis website (https: / / moriaritylab.shinyapps.io / editr_v10 / ) for single-base editing efficiency analysis to determine the editing efficiency of cytosine at different positions of the target site using different deaminases.
[0230] e. The specific editing efficiency of different deaminases for the mutation of C bases to T bases at different positions of the target site is shown in Table 1 and Figures 2-4.
[0231] Table 1 Editing efficiency of each base editor
[0232] From the data in Table 1, it can be seen that deaminases-12, 15, 17, 20, 24, 27-29, 38, 48, 90, and 102 are active after constructing base editors; and the activities of deaminases-27 and 29 are better than that of deaminase-03.
[0233] CBE-48 has a very wide editing window, which is beneficial for special needs. CBE-102 has a very narrow editing window, which is very specific; in addition, the editing window position is significantly different from other deaminases.
[0234] Experimental Example 3: Detection of editing efficiency of base editors composed of different cytosine deaminases
[0235] The editing efficiency of each editor in Table 2 was tested using the same method as in Example 2, and the results shown were obtained. CBE-136 and CBE-140 had higher editing efficiency (see Figures 5 and 6).
[0236] Table 2 Editing efficiency of each base editor
[0237] Experimental Example 4. Detection of Editing Efficiency of Base Editors Composed of Different Cytosine Deaminases
[0238] The editing efficiency of the editor composed of deaminase-139, deaminase-145, deaminase-141, and deaminase-143 was tested using the same method as in Example 2. The results are shown in Figures 7-10.
[0239] As can be seen from Figure 8, CBE-145 has a very narrow editing window and good specificity.
[0240] Experimental Example 5. Obtaining and verifying the mutant of deaminase-145
[0241] (1) Screening of highly efficient and specific mutants of CBE-145
[0242] The wild-type amino acid sequence of CBE-145 (207aa) is shown below:
[0243] According to the method of the literature (Neugebauer ME, Hsu A, Arbab M, et al. Evolution of an adenine base editor into a small, efficient cytosine base editor with low off-target activity [J]. Nature biotechnology, 2023, 41 (5): 673-685.), phage expressing deaminase-145 was prepared, and the screening system and strain of the literature were used to evolve deaminase-145 by the PANCE method. Finally, the following deaminase-145 mutants were screened out: deaminase-145-P5 (Q122H), deaminase-145-P1 (S191R), deaminase-145-P2 (I114T), deaminase-145-P4 (L2M), and deaminase-145-P6 (N89D).
[0244] The amino acid sequence of mutant deaminase-145-P5 (Q122H) is:
[0245] (2) Construction of vector plasmid for verification
[0246] Verification vectors were synthesized at an outsourcing service company, including the same CBE-145 expression plasmid vector (SEQ ID NO: 330) as described in Examples 1 and 4, as well as expression plasmid vectors containing deaminase-145 mutants (CBE-145-P5, CBE-145-P1, CBE-145-P2, CBE-145-P4, and CBE-145-P6). The sequence of the deaminase-145-P5 vector plasmid CBE-145-P5 is SEQ ID NO: 331. The Cas protein portion of the editor fusion protein used nSPRY; the gRNA targeted the VEGFA target site, and the spacer was GACCCCCTCCACCCCGCCTC (SEQ ID NO: 332).
[0247] (3) Transfection of 293T cells with the vector to be verified
[0248] Different validation vectors and control vectors were transfected into 293T cells.
[0249] Transfection was performed in 24-well plates according to the instructions of Lipofectamine 2000 (Thermo), and detection was performed 72 h later.
[0250] (4) Detection of changes in VEGFA target genes
[0251] 72 hours after transfection, the cell genome was extracted using a genomic DNA extraction kit (Tiangen Biotechnology), and the extracted product was amplified by PCR. The primers are as follows:
[0252] VEGFA-CBE-F3: ACTTGAATCGGGCCGACG (SEQ ID NO: 333),
[0253] VEGFA-CBE-R3: CCTTGGCATGGTGGAGGTAGAG (SEQ ID NO: 334).
[0254] 2X Pro Taq premix Ver.2 (Acori Biotech) was used for amplification, and the product length was 866 bp. The product was sent to GeneWeichi for sequencing.
[0255] The sequencing results were analyzed, and the results of three batches are shown in Figure 11: 1-20 are the nucleotide position numbers of the spacer, the yellow highlight indicates the base C that can be edited by CBE, and the number below is the percentage of editing from C to T (unit: %), among which the 293T group is the unedited negative control, and the inventors consider ≥10% to have an editing effect.
[0256] In three batches of experiments, CBE-145 containing wild-type deaminase edited Cs at positions 10 and 12 (one batch also showed a weak editing effect at position 14), converting them to Ts. CBE-145-P5 (Q122H) edited only at position 10, with a narrower editing window. Furthermore, across all three batches, CBE-145-P5 performed better at position 10 than CBE-145 containing wild-type deaminase. Therefore, deaminase-145-P5 exhibits superior efficiency and specificity to deaminase-145.
[0257] Across the three batches, CBE-145-P2 and CBE-145-P4 showed better editing performance at position 10 than CBE-145 containing the wild-type deaminase. Therefore, the efficiency of deaminase-145-P2 / P4 is superior to that of deaminase-145.
[0258] Experimental Example 6. Verification of more mutants of deaminase-145
[0259] (1) Screening of highly efficient and specific mutants of CBE-145
[0260] According to the literature (Neugebauer ME, Hsu A, Arbab M, et al. Evolution of an adenine base editor into a small, efficient cytosine base editor with low off-target activity [J]. Nature biotechnology, 2023, 41(5): 673-685.) method to prepare phage expressing deaminase-145, and use the screening system and strain of the literature to evolve deaminase-145 by the PANCE method, and finally screen out deaminase-145 mutants: deaminase-145-P5 (Q122H), deaminase-145-P1 (S191R), deaminase-145-P2 (I114T), deaminase-145-P3 (C176F), deaminase-145-P4 (L2M), deaminase-145-P6 (N89D), deaminase-145-P7 (N68I), deaminase-145-P8 (Q78*, that is, the codon of amino acid Q at position 78 is mutated to a stop codon), and deaminase-145-P11 (N68I and I114T mutations).
[0261] (2) Construction of vector plasmid for verification
[0262] Verification vectors were synthesized at an outsourcing service company, including the same CBE-145 expression plasmid vector (SEQ ID NO: 330) as in Examples 1 and 4, as well as expression plasmid vectors containing deaminase-145 mutants (CBE-145-P5, CBE-145-P1, CBE-145-P2, CBE-145-P3, CBE-145-P4, CBE-145-P6, CBE-145-P7, CBE-145-P8, and CBE-145-P11). The sequence of the deaminase-145-P5 vector plasmid CBE-145-P5 is SEQ ID NO: 331. The Cas protein portion of the editor fusion protein used nSPRY; the gRNA targeted the VEGFA target site, and the spacer was GACCCCCTCCACCCCGCCTC (SEQ ID NO: 332).
[0263] (3) Transfection of 293T cells with the vector to be verified
[0264] Different validation vectors and control vectors were transfected into 293T cells.
[0265] Transfection was performed in 24-well plates according to the instructions of Lipofectamine 2000 (Thermo), and detection was performed 72 h later.
[0266] (4) Detection of changes in VEGFA target genes
[0267] 72 hours after transfection, the cell genome was extracted using a genomic DNA extraction kit, and the extracted product was amplified by PCR. The primers are as follows:
[0268] VEGFA-CBE-F3: ACTTGAATCGGGCCGACG (SEQ ID NO: 333),
[0269] VEGFA-CBE-R3: CCTTGGCATGGTGGAGGTAGAG (SEQ ID NO: 334).
[0270] 2X Pro Taq premix Ver.2 (Acori Biotech) was used for amplification, and the product length was 866 bp. The product was sent to GeneWeichi for sequencing.
[0271] The sequencing results were analyzed, and the results of three batches are shown in Figure 12: 1-20 are the nucleotide position numbers of the spacer, and the numbers below the bases are the percentages of C editing to T (unit: %). The 293T group is an unedited negative control, and the inventors consider ≥10% to have an editing effect.
[0272] The results in Figure 12 show that CBE-145-P2, CBE-145-P3, CBE-145-P4, and CBE-145-P5 all have better editing effects at position 10 than CBE-145 containing the wild-type deaminase. Therefore, the efficiency of the deaminase-145-P2 / P3 / P4 / P5 variants is better than that of deaminase-145.
[0273] The deaminase-145-P1, P6, and P8 mutants almost completely lost their activity.
[0274] The deaminase-145-P7 and P11 mutants still showed strong activity. The P11 mutant showed better editing specificity at position 10 than the wild type.
[0275] Appendix 1: Target nucleic acids / target genes and corresponding diseases or conditions (when there are two or more target genes in a specific cell, it means that these two or more target genes are targeted simultaneously):
Claims
1. A deaminase, characterized in that The deaminase comprises an amino acid sequence having at least 50% sequence identity to any one of SEQ ID NOs: 1-6, 39-164, 297-312; Optionally, the deaminase deaminates at least one nucleotide in the polynucleotide; Optionally, the deaminase can deaminize bases in a nucleic acid sequence together with a nucleic acid binding polypeptide; Optionally, the deaminase is capable of deaminating bases in a target polynucleotide sequence together with the nucleic acid binding polypeptide and the guide RNA; Optionally, the deaminase comprises a homologous or heterologous domain; further, the deaminase comprises a nucleic acid binding polypeptide, a cellular localization signal and / or a UGI domain.
2. The deaminase according to claim 1, characterized in that The deaminase comprises an amino acid sequence as shown in any one of SEQ ID NOs: 1-6, 39-164, 297-312; alternatively, the amino acid sequence of the deaminase is not any one of SEQ ID NOs: 1-5, 54-60; Optionally, the deaminase comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100% sequence identity to any one of SEQ ID NOs: 3, 5, 6, 160, 163, 164, 297, 299 and 301; Optionally, the deaminase comprises mutations at positions corresponding to amino acid residues 122, 191, 114, 2, 176, 68 and / or 89 of the reference sequence shown in SEQ ID NO:
301.
3. A fusion protein, characterized in that The fusion protein comprises: the deaminase according to claim 1 or 2.
4. The fusion protein according to claim 3, characterized in that The fusion protein comprises: a) a nucleic acid binding polypeptide that binds to a target polynucleotide; and b) a deaminase according to claim 1 or 2; Optionally, the nucleic acid binding polypeptide is a TALEN nuclease, a zinc finger nuclease, a CRISPR-Cas nuclease or a meganuclease; Optionally, the nucleic acid binding polypeptide is an RNA-guided nuclease; Optionally, the nucleic acid binding polypeptide is a Cas enzyme (CRISPR-Cas nuclease); Optionally, the nucleic acid binding polypeptide is selected from Cas9, Cas12, Cas13, TnpB, IscB, IsrB, Fancor nuclease; or fragments thereof, including but not limited to nucleic acid binding domain fragments; Further optionally, the fusion protein further comprises one or more UGIs; and / or, the fusion protein comprises one or more NLSs; Further optionally, the nucleic acid binding polypeptide, the deaminase, the one or more UGIs and the one or more NLSs are directly connected or connected via a linker.
5. An isolated nucleic acid, characterized in that The nucleic acid encodes the deaminase according to claim 1 or 2 or the fusion protein according to claim 3 or 4.
6. An expression vector, characterized in that: The expression vector comprises the nucleic acid of claim 5.
7. A pharmaceutical composition, characterized in that The pharmaceutical composition comprises the deaminase according to claim 1 or 2, the fusion protein according to claim 3 or 4, the nucleic acid according to claim 5 or the expression vector according to claim 6, and a pharmaceutically acceptable delivery carrier.
8. An expression system, characterized in that The expression system contains the expression vector according to claim 6, the pharmaceutical composition according to claim 7, or the nucleic acid according to claim 5 is integrated into the genome.
9. A system for modifying a target polynucleotide, characterized in that The system comprises: a) one or more guide RNAs or one or more nucleic acids encoding the one or more guide RNAs, the one or more guide RNAs being capable of hybridizing to the target polynucleotide; The system further includes any one or more of the following (a)-(f): b) a deaminase according to claim 1 or 2, or a nucleic acid encoding the deaminase; and a nucleic acid binding polypeptide, or a nucleic acid encoding the nucleic acid binding polypeptide; c) a fusion protein according to claim 3 or 4, or a nucleic acid encoding the fusion protein; d) the expression vector according to claim 6; e) The pharmaceutical composition according to claim 7; f) The expression system according to claim 8.
10. A method for modifying a target polynucleotide, characterized in that: The method comprises contacting the target polynucleotide with the deaminase of claim 1 or 2, the fusion protein of claim 3 or 4, the nucleic acid of claim 5, the expression vector of claim 6, the pharmaceutical composition of claim 7, the expression system of claim 8, or the system for modifying the target polynucleotide of claim 9.
Citation Information
Patent Citations
Nucleobase editors and uses thereof
CN108513575A
Novel CRISPR / Cas12f enzymes and systems
CN111757889B
Compositions comprising a nuclease and uses thereof
WO2021113522A1
Gene therapy for treating beta-hemoglobinopathies
WO2022206986A1
Mutant cytidine deaminases with improved editing precision
WO2023155901A1
Cited By
Optimized NlovFz2-omega RNA editing system
CN119614632A
An optimized NlovFz2-ω RNA editing system
CN119614632B