Fusion proteins with edited bases and their applications
Patent Information
- Application Number
- CN202111463414.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-02
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2041-12-02
AI Technical Summary
而对于ABE而言,虽然也有提高效率的ABEmax、缩窄窗口的ABE7.10F148A,但其它版本的改进进展缓慢
[0048]本发明的编辑腺嘌呤的融合蛋白可以有效提高近PAM区的编辑效率,并拓宽了编辑窗口;本发明的编辑腺嘌呤和胞嘧啶的融合蛋白有效提高了A/C同时编辑效率。
Smart Images

Figure HDA0003390168470000011 
Figure HDA0003390168470000012 
Figure HDA0003390168470000021
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biotechnology, specifically relating to a fusion protein for editing bases and its applications. This invention also relates to a polynucleotide encoding the fusion protein for editing bases, a recombinant expression vector containing the polynucleotide, a base editing system comprising the aforementioned fusion protein, polynucleotide, or recombinant expression vector, and a pharmaceutical composition. Background Technology
[0002] Since 2013, next-generation gene-editing technologies, represented by CRISPR / Cas9, have entered various experiments in the field of biology, changing traditional gene manipulation methods. In April 2016, David Liu's laboratory first reported a CBE single-base editing tool based on the fusion of cytosine deaminase and CRISPR / Cas9, capable of C-to-T or G-to-A editing on DNA. In October 2017, David Liu's laboratory again reported an ABE single-base editing tool based on the fusion of cytosine deaminase and CRISPR / Cas9, capable of A-to-G or T-to-C editing on DNA.
[0003] Single-base gene editing (SBE) technology has been reported for efficient genome mutation or repair, creation of animal disease models, and gene therapy. Various improved versions of CBE exist, such as the efficiency-enhancing CBEmax, the sequence-specific eA3A-BE3, the narrowed window YE1-BE3, the widened window BE-PLUS, and the hyCBE, which improves editing efficiency targeting the PAM region. For ABE, while there are efficiency-enhancing versions like ABEmax and the narrowed window ABE7.10F148A, improvements in other versions have been slow. Regarding ABEs that improve editing efficiency targeting the PAM region, due to poor adenosine deaminase compatibility, ultra-high activity ABEs have not been reported. Furthermore, the bifunctional base editor A&C-BEmax, based on the fusion of cytosine deaminase and adenosine deaminase with CRISPR / Cas9, suffers from low efficiency in simultaneous editing of both bases due to some loss of adenosine deaminase activity, limiting its widespread application.
[0004] Currently, ABE exhibits low editing efficiency near the PAM region. Furthermore, the dual-function base editor A&C-BEmax, based on the fusion of cytosine deaminase and adenosine deaminase with CRISPR / Cas9, achieves a maximum simultaneous editing efficiency of 30% for both A and C bases, which remains relatively low. This limits the widespread application of ABE and A&C-BEmax to some extent.
[0005] Therefore, developing a universal, ultra-high activity ABE and A&C-BEmax is also an urgent need in this field. Summary of the Invention
[0006] The technical problem this invention aims to solve is to overcome the shortcomings of existing technologies, such as the lack of adenine base editors with high near-PAM region editing efficiency and the lack of base editors that can simultaneously edit adenine and cytosine. This invention provides a fusion protein for editing bases and its applications. The fusion protein of this invention not only effectively improves the editing efficiency of the near-PAM region and widens the editing window, but also effectively improves the efficiency of simultaneous A / C editing.
[0007] Based on their previous research strategy, the inventors fused the DNA-binding domain of Rad51 (Rad51 DBD) into an existing adenine base editor. After screening and trying different Rad51 DBD fusion sites and adenine base editors, they found that fusing Rad51 DBD with ABEmax almost completely lost the editing efficiency from A to G; while fusing Rad51DBD with ABE8e, at a specific fusion site, yielded a highly active adenine base editor, effectively improving the editing efficiency near the PAM region and widening the editing window. Furthermore, using a similar strategy to fuse Rad51 DBD into a base editor for editing adenine and cytosine (A&C-BEmax) did not effectively improve the simultaneous A / C editing efficiency. Therefore, the inventors attempted to improve the adenine deaminase (replacing TadA-TadA* with TadA8e), and found that its simultaneous A / C editing efficiency was effectively improved. Further fusing Rad51 DBD yielded a base editor with ultra-high activity for editing adenine and cytosine.
[0008] The present invention solves the above-mentioned technical problems through the following technical solutions:
[0009] The first aspect of the present invention provides a fusion protein for editing adenine, wherein the fusion protein comprises a TadA8e fragment and a Cas9n fragment sequentially from the N-terminus to the C-terminus, the amino acid sequence of the TadA8e fragment is shown in SEQ ID NO:1, and the Cas9n fragment is a Cas9 protein fragment with single-chain cleavage activity.
[0010] The fusion protein further includes a Rad51 fragment, which is adjacent to the Cas9n fragment and located at the N-terminus or C-terminus of the Cas9n fragment; the amino acid sequence of the Rad51 fragment is shown in SEQ ID NO:2.
[0011] In this invention, the Rad51 fragment is the DNA binding domain of Rad51, and Rad51 is a human protein with DNA break repair capabilities.
[0012] In some embodiments of the present invention, the fusion protein further includes a nuclear localization signal fragment located at the N-terminus and / or C-terminus of the fusion protein.
[0013] In some specific embodiments of the present invention, the amino acid sequence of the nuclear localization signal fragment is shown in SEQ ID NO:3.
[0014] In this invention, the nuclear localization signal refers to the amino acid sequence that promotes protein entry into the cell nucleus (e.g., via nuclear transport), which is well known in the art.
[0015] In some embodiments of the present invention, the Cas9n fragment is selected from the group consisting of SpCas9, SaCas9, NmeCas9, StCas9, CjCas9 or mutants thereof.
[0016] In some preferred embodiments of the present invention, the Cas9n fragment is derived from SpCas9.
[0017] In some specific embodiments of the present invention, the amino acid sequence of the Cas9n fragment is shown in SEQ ID NO:4.
[0018] In this invention, the Cas9n fragment is a fragment of the Cas9 protein with nucleic acid cleavage function or a mutant thereof. Preferably, the Cas9 protein may be derived from the Cas9 protein of *Streptococcus pyogenes* (SpCas9), the Cas9 protein of *Streptococcus aureus* (SaCas9), the Cas9 protein of *Neisseria meningitides* (NmeCas9), the Cas9 protein of *Streptococcus thermophilus* (StCas9), or the Cas9 protein of *Campylobacter jejuni* (CjCas9).
[0019] In this invention, TadA8e is an adenine deaminase that can remove the amino group from the adenine molecule.
[0020] In some specific embodiments of the present invention, the nuclear localization signal fragment of the fusion protein is located at the N-terminus and C-terminus of the fusion protein, and the Cas9n fragment is derived from SpCas9.
[0021] A second aspect of the present invention provides a fusion protein for editing adenine and cytosine, wherein the fusion protein comprises, from the N-terminus to the C-terminus, an hAID fragment, a TadA8e fragment, and a Cas9n fragment, wherein the amino acid sequence of the hAID fragment is shown in SEQ ID NO:5, the amino acid sequence of the TadA8e fragment is shown in SEQ ID NO:1, and the Cas9n fragment is a Cas9 protein fragment with single-strand cleavage activity.
[0022] In some embodiments of the present invention, the fusion protein further includes a Rad51 fragment, which is adjacent to the Cas9n fragment and located at the N-terminus of the Cas9n fragment; the amino acid sequence of the Rad51 fragment is shown in SEQ ID NO:2.
[0023] In some preferred embodiments of the present invention, the Cas9n fragment is the Cas9n fragment of the fusion protein as described in the first aspect. That is, the source of the Cas9n fragment is selected from the group consisting of SpCas9, SaCas9, NmeCas9, StCas9, CjCas9, or mutants thereof. Preferably, the source of the Cas9n fragment is SpCas9. More preferably, the amino acid sequence of the Cas9n fragment is as shown in SEQ ID NO:4.
[0024] In some embodiments of the present invention, the fusion protein further includes a UGI fragment and / or a nuclear localization signal fragment.
[0025] In some preferred embodiments of the present invention, the UGI fragment is located at the C-terminus of the Cas9n fragment; the nuclear localization signal fragment is located at the N-terminus and / or C-terminus of the fusion protein.
[0026] In some preferred embodiments of the invention, the UGI fragment comprises two copies, each copy preferably having the amino acid sequence shown in SEQ ID NO:6. The amino acid sequence of the nuclear localization signal fragment is shown in SEQ ID NO:3.
[0027] In this invention, hAID is a cytosine deaminase, which can remove the amino group from the cytosine molecule.
[0028] A third aspect of the present invention provides a polynucleotide that encodes a fusion protein as described in the first or second aspect.
[0029] In this invention, the polynucleotide can be in the form of DNA or RNA. The DNA form includes cDNA, genomic DNA, or artificially synthesized DNA. The DNA can be single-stranded or double-stranded. The DNA can be a coding strand or a non-coding strand.
[0030] It is well known in the art that, due to the degeneracy of codons, a large number of polynucleotide sequences can encode the same fusion protein. Furthermore, different species exhibit certain codon preferences, and the codons of the fusion protein may be optimized according to the needs of expression in different species. These variants are all within the protection scope of the polynucleotides described in this invention.
[0031] A fourth aspect of the present invention provides a recombinant expression vector comprising the polynucleotides as described in the third aspect.
[0032] In this invention, the recombinant expression vector can be a conventional vector that can be stably expressed in host cells, such as bacterial plasmids, bacteriophages, yeast plasmids, plant cell viruses, mammalian cell viruses such as adenoviruses, and retroviruses.
[0033] A fifth aspect of the present invention provides a base editing system comprising: a base editor and sgRNA; the base editor comprising a fusion protein as described in the first or second aspect.
[0034] In this invention, the sgRNA has the function of binding to the Cas9n protein in the base editor and guiding it to the target gene. The sgRNA has a sequence complementary to the target gene at its 5' end, and binds to the target gene through this complementary sequence, thereby guiding the base editor of this invention to the target gene.
[0035] The target gene contains a PAM sequence, which is a sequence that can be recognized by the Cas9n protein.
[0036] A sixth aspect of the present invention provides a pharmaceutical composition comprising a fusion protein as described in the first or second aspect, or a base editing system as described in the fifth aspect.
[0037] A seventh aspect of the present invention provides a non-therapeutic base editing method, the base editing method comprising base editing in target cells by means of a fusion protein as described in the first or second aspect, a polynucleotide as described in the third aspect, a recombinant expression vector as described in the fourth aspect, or a base editing system as described in the fifth aspect.
[0038] In some embodiments of the present invention, the target cell is a prokaryotic cell or a eukaryotic cell.
[0039] In this invention, the prokaryotic cells can be microbial cells known in the art, preferably microbial cells with medical research value and industrial production value.
[0040] In some preferred embodiments of the present invention, the eukaryotic cells are animal cells or plant cells.
[0041] In this invention, the animal cells can be mammalian cells or rodent cells, including human, horse, cow, sheep, rat, rabbit, etc.
[0042] In this invention, the non-therapeutic purpose is, for example, to evaluate the adenine-editing fusion protein, the adenine-cytosine-editing fusion protein, or the corresponding base editing system described herein by detecting editing occurring in target cells in the laboratory. Conversely, base editing can also be used to study the function of target cells.
[0043] In this invention, the base editing method can also be for therapeutic purposes. In this invention, "treatment" refers to treating a disease in a subject, such as a human, including inhibiting the occurrence or development of the disease, alleviating the symptoms of the disease, or curing the disease.
[0044] The ninth aspect of the present invention provides the use of fusion proteins as described in the first or second aspect, polynucleotides as described in the third aspect, recombinant expression vectors as described in the fourth aspect, base editing systems as described in the fifth aspect, or pharmaceutical compositions as described in the sixth aspect in the preparation of base editing reagents and the construction of animal models.
[0045] Based on common knowledge in the field, the above-mentioned preferred conditions can be combined arbitrarily to obtain various preferred embodiments of the present invention.
[0046] The reagents and raw materials used in this invention are all commercially available.
[0047] The positive and progressive effects of this invention are as follows:
[0048] The adenine-editing fusion protein of the present invention can effectively improve the editing efficiency near the PAM region and widen the editing window; the adenine and cytosine-editing fusion protein of the present invention effectively improves the simultaneous A / C editing efficiency. Attached Figure Description
[0049] Figure 1 This is a schematic diagram illustrating the design and construction of ultra-high activity ABE.
[0050] Figure 2 A schematic diagram showing the comparison of editing efficiency of different ultra-high activity ABE constructs at the human endogenous target CCR5-sg1.
[0051] Figure 3 This is a schematic diagram showing the comparison results of ultra-high activity ABE (hyABE8e) and ABE8e at multiple endogenous targets in humans.
[0052] Figure 4 A schematic diagram illustrating the design and construction of the highly active A&C-BEmax.
[0053] Figure 5 A schematic diagram showing the comparison of the base editing efficiency of different highly active A&C-BEmax constructs at the human endogenous target CTLA-sg2.
[0054] Figure 6 A schematic diagram showing the comparison of the efficiency of simultaneous editing of A / C alleles (A to G or C to T) of different highly active A&C-BEmax constructs at the human endogenous target CTLA-sg2 gene.
[0055] Figure 7 This is a schematic diagram showing the comparison of the base editing efficiency of ultra-high activity A&C-BEmax (hyA&C-BEmax), high activity A&C-BEmax (eA&C-BEmax), and A&C-BEmax at multiple endogenous target sites in humans.
[0056] Figure 8 This is a schematic diagram showing the comparison of the simultaneous editing efficiency of A / C alleles at multiple endogenous target sites in humans by ultra-high activity A&C-BEmax (hyA&C-BEmax), high activity A&C-BEmax (eA&C-BEmax), and A&C-BEmax. Detailed Implementation
[0057] The present invention is further illustrated below by way of embodiments, but the invention is not limited to the scope of the embodiments described herein. Experimental methods in the following embodiments that do not specify specific conditions were performed according to conventional methods and conditions, or as selected according to the product instructions.
[0058] For the plasmid construction and transfection methods in the examples, please refer to Xiaohui Zhang et al., Increasing the efficiency and targeting range of cytidine base editors through fusion of asingle-stranded DNA-binding protein domain, Nature Cell Biology 22(6):1-11.
[0059] Example 1
[0060] 1. For example Figure 1As shown, functional domain fusion constructs of the human single-stranded DNA binding protein Rad51 were designed based on the previously reported ABEmax, and functional domain fusion constructs of the human single-stranded DNA binding protein Rad51 were designed based on the recently reported ABE8e (Addgene#138489). Simultaneously, the endogenous target CCR5-sg1p (tgacatcaattattatacatcgg, SEQ ID NO:7) derived from humans was designed and tested. The above plasmids were constructed sequentially according to the manufacturer's instructions.
[0061] 2. The constructed plasmid and target plasmid were co-transfected into HEK293T cells at a ratio of 750 ng: 250 ng. 120 h post-transfection, cellular DNA was extracted, and PCR was performed followed by PCR library construction, which was then sent for high-throughput sequencing. The results were analyzed to identify the fusion construct hyABE (e.g., ...) with the highest editing efficiency. Figure 2 (As shown).
[0062] 3. Using ABE8e as a control, the hyABE fusion construct with the best editing efficiency was combined with more endogenous targets:
[0063] FANCF-Mb: aagttcgctaatcccggaactgg (SEQ ID NO:8),
[0064] MAGEA1-Mb: gctggcagcaagggcggcgctgg (SEQ ID NO:9),
[0065] ABE site25: agtaaacaaagcatagactgagg (SEQ ID NO: 10) and
[0066] PPP1R12C site6:ggggctcaacatcggaagagggg(SEQ ID NO:11)
[0067] Co-transfected HEK293T cells to verify its functional characteristics, the results are as follows: Figure 3 As shown.
[0068] The results showed that by screening different fusions of the single-stranded DNA-binding protein domain Rad51DBD with the known highly active and compatible ABE8e, it was found that Rad51DBD fusion between TadA8e and Cas9n, as well as Rad51DBD fusion at the C-terminus of Cas9n, effectively improved the editing efficiency near PAM. The highest efficiency was observed when Rad51DBD fused between TadA8e and Cas9n. For ease of description, this fusion protein was named hyABE8e. The editing window of hyABE8e was widened from positions 3-8 of ABE8e to positions 3-13.
[0069] Example 2
[0070] 1. For example Figure 4 As shown, functional domain fusion constructs of the human single-stranded DNA binding protein Rad51 were designed based on the previously reported A&C-BEmax. The two adenine deaminases in A&C-BEmax were replaced with adenine deaminases from the recently reported, highly active and compatible ABE8e to construct the fusion construct (eA&C-BEmax). The functional domains of the human single-stranded DNA binding protein Rad51 were further fused to construct the fusion construct (hyA&C-BEmax). Simultaneously, the human endogenous target CTLA-sg2 (tgtgaagagcttcactgagtagg, SEQ ID NO:12) was designed and tested. The above plasmids were constructed sequentially.
[0071] 2. The constructed plasmid and target plasmid were co-transfected into HEK293T cells at a ratio of 750 ng: 250 ng. 120 h post-transfection, cellular DNA was extracted, and PCR was performed followed by PCR library construction, which was then sent for high-throughput sequencing. The results were analyzed to identify the fusion construct with the optimal base editing efficiency, such as... Figure 5 and Figure 6 As shown.
[0072] The results show that replacing TadA-TadA* in A&C-BEmax with TadA8e of ABE8e yields eA&C-BEmax, which further improves the efficiency of simultaneous A / C editing.
[0073] 3. Using A&C-BEmax as a control, the optimal fusion constructs (eA&C-BEmax and hyA&C-BEmax) were compared with more endogenous targets:
[0074] FANCF-sg1: cgaccaatagcattgcagagagg (SEQ ID NO: 13),
[0075] PPP1R12C site3: gagctcactgaacgctggcatgg (SEQ ID NO:14) and
[0076] EGFR-sg12: atacacgtgccgaacgcaccgg (SEQ ID NO:15)
[0077] Co-transfected HEK293T cells to verify its functional characteristics, such as... Figure 7 and Figure 8 As shown.
[0078] The results showed that fusing the Rad51DBD domain of the single-stranded DNA binding protein with eA&C-BEmax yielded hyA&C-BEmax, which further improved the efficiency of simultaneous A / C editing. SEQUENCE LISTING <110> East China Normal University Shanghai Bangyao Biomedical Co., Ltd. <120> Fusion proteins with edited bases and their applications <130> P21017691C <160> 15 <170> PatentIn version 3.5 <210> 1 <211> 166 <212> PRT <213> Artificial Sequence <220> <223> TadA8e fragment <400> 1 Ser Glu Val Glu Phe Ser His Glu Tyr Trp Met Arg His Ala Leu Thr 1 5 10 15 Leu Ala Lys Arg Ala Arg Asp Glu Arg Glu Val Pro Val Gly Ala Val 20 25 30 Leu Val Leu Asn Asn Arg Val Ile Gly Glu Gly Trp Asn Arg Ala Ile 35 40 45 Gly Leu His Asp Pro Thr Ala His Ala Glu Ile Met Ala Leu Arg Gln 50 55 60 Gly Gly Leu Val Met Gln Asn Tyr Arg Leu Ile Asp Ala Thr Leu Tyr 65 70 75 80 Val Thr Phe Glu Pro Cys Val Met Cys Ala Gly Ala Met Ile His Ser 85 90 95 Arg Ile Gly Arg Val Val Phe Gly Val Arg Asn Ser Lys Arg Gly Ala 100 105 110 Ala Gly Ser Leu Met Asn Val Leu Asn Tyr Pro Gly Met Asn His Arg 115 120 125 Val Glu Ile Thr Glu Gly Ile Leu Ala Asp Glu Cys Ala Ala Leu Leu 130 135 140 Cys Asp Phe Tyr Arg Met Pro Arg Gln Val Phe Asn Ala Gln Lys Lys 145 150 155 160 Ala Gln Ser Ser Ile Asn 165 <210> 2 <211> 113 <212> PRT <213> Artificial Sequence <220> <223> Rad51 fragment <400> 2 Ala Met Gln Met Gln Leu Glu Ala Asn Ala Asp Thr Ser Val Glu Glu 1 5 10 15 Glu Ser Phe Gly Pro Gln Pro Ile Ser Arg Leu Glu Gln Cys Gly Ile 20 25 30 Asn Ala Asn Asp Val Lys Lys Leu Glu Glu Ala Gly Phe His Thr Val 35 40 45 Glu Ala Val Ala Tyr Ala Pro Lys Lys Glu Leu Ile Asn Ile Lys Gly 50 55 60 Ile Ser Glu Ala Lys Ala Asp Lys Ile Leu Ala Glu Ala Ala Lys Leu 65 70 75 80 Val Pro Met Gly Phe Thr Thr Ala Thr Glu Phe His Gln Arg Arg Ser 85 90 95 Glu Ile Ile Gln Ile Thr Thr Gly Ser Lys Glu Leu Asp Lys Leu Leu 100 105 110 Gln <210> 3 <211> 18 <212> PRT <213> Artificial Sequence <220> Nuclear localization signal fragment <400> 3 Lys Arg Thr Ala Asp Gly Ser Glu Phe Glu Ser Pro Lys Lys Lys Arg 1 5 10 15 Lys Val <210> 4 <211> 1367 <212> PRT <213> Artificial Sequence <220> <223> Cas9n fragment <400> 4 Asp Lys Lys Tyr Ser Ile Gly Leu Ala Ile Gly Thr Asn Ser Val Gly 1 5 10 15 Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe Lys 20 25 30 Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile Gly 35 40 45 Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu Lys 50 55 60 Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys Tyr 65 70 75 80 Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser Phe 85 90 95 Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys His 100 105 110 Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr His 115 120 125 Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp Ser 130 135 140 Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His Met 145 150 155 160 Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro Asp 165 170 175 Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr Asn 180 185 190 Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala Lys 195 200 205 Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn Leu 210 215 220 Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn Leu 225 230 235 240 Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe Asp 245 250 255 Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp Asp 260 265 270 Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp Leu 275 280 285 Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp Ile 290 295 300 Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser Met 305 310 315 320 Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys Ala 325 330 335 Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe Asp 340 345 350 Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser Gln 355 360 365 Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp Gly 370 375 380 Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg Lys 385 390 395 400 Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu Gly 405 410 415 Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe Leu 420 425 430 Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile Pro 435 440 445 Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp Met 450 455 460 Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu Val 465 470 475 480 Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr Asn 485 490 495 Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser Leu 500 505 510 Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys Tyr 515 520 525 Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln Lys 530 535 540 Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr Val 545 550 555 560 Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp Ser 565 570 575 Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly Thr 580 585 590 Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp Asn 595 600 605 Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr Leu 610 615 620 Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala His 625 630 635 640 Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr Thr 645 650 655 Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp Lys 660 665 670 Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe Ala 675 680 685 Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe Lys 690 695 700 Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu His 705 710 715 720 Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly Ile 725 730 735 Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly Arg 740 745 750 His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln Thr 755 760 765 Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile Glu 770 775 780 Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro Val 785 790 795 800 Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln 805 810 815 Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg Leu 820 825 830 Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys Asp 835 840 845 Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg Gly 850 855 860 Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys Asn 865 870 875 880 Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe 885 890 895 Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp Lys 900 905 910 Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr Lys 915 920 925 His Gln Ile Leu Asp Ser Arg With Asn Thr Lys Tyr Asp Glu 930,935,940 Asn Asp Lys With Arg Glu Val Val Lys With Thr Lys Ser Ser Lys 945 950 955 960 Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg Glu 965,970,975 Ile Asn Asn Tyr His Ala His Asp Ala Tyr Leu Asn Ala Val Val 980,985,990 Gly Thr Ala Leu Ile Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val 995 1000 1005 Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala Lys 1010 1015 1020 Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr 1025 1030 1035 Serving Asn With Asn Phe Phe Lys Thr Glue With Thr Leu Ala Asn 1040 1045 1050 Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr 1055 1060 1065 Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val Arg 1070 1075 1080 Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr Glu 1085 1090 1095 Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg 1100 1105 1110 Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys 1115 1120 1125 Lys Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val Leu 1130 1135 1140 Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser 1145 1150 1155 Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser Phe 1160 1165 1170 Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu 1175 1180 1185 Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu Phe 1190 1195 1200 Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly Glu 1205 1210 1215 Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn 1220 1225 1230 Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser Pro 1235 1240 1245 Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys His 1250 1255 1260 Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys Arg 1265 1270 1275 Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr 1280 1285 1290 Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile 1295 1300 1305 Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe 1310 1315 1320 Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr 1325 1330 1335 Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr Gly 1340 1345 1350 Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 <210> 5 <211> 181 <212> PRT <213> Artificial Sequence <220> <223> hAID <400> 5 Asp Ser Leu Leu Met Asn Arg Arg Lys Phe Leu Tyr Gln Phe Lys Asn 1 5 10 15 Val Arg Trp Ala Lys Gly Arg Arg Glu Thr Tyr Leu Cys Tyr Val Val 20 25 30 Lys Arg Arg Asp Ser Ala Thr Ser Phe Ser Leu Asp Phe Gly Tyr Leu 35 40 45 Arg Asn Lys Asn Gly Cys His Val Glu Leu Phe Leu Arg Tyr Ile 50 55 60 Ser Asp Trp Asp Leu Asp Pro Gly Arg Cys Tyr Arg Val Thr Trp Phe 65 70 75 80 Thr Ser Trp Ser Pro Cys Tyr Asp Cys Ala Arg His Val Ala Asp Phe 85 90 95 Leu Arg Gly Asn Pro Asn Leu Ser Leu Arg Ile Phe Thr Ala Arg Leu 100 105 110 Tyr Phe Cys Glu Asp Arg Lys Ala Glu Pro Glu Gly Leu Arg Arg Leu 115 120 125 His Arg Ala Gly Val Gln Ile Ala Ile Met Thr Phe Lys Asp Tyr Phe 130 135 140 Tyr Cys Trp Asn Thr Phe Val Glu Asn His Glu Arg Thr Phe Lys Ala 145 150 155 160 Trp Glu Gly Leu His Glu Asn Ser Val Arg Leu Ser Arg Gln Leu Arg 165 170 175 Arg Ile Leu Leu Pro 180 <210> 6 <211> 83 <212> PRT <213> Artificial Sequence <220> <223> UGI Fragment <400> 6 Thr Asn Leu Ser Asp Ile Ile Glu Lys Glu Thr Gly Lys Gln Leu Val 1 5 10 15 Ile Gln Glu Ser Ile Leu Met Leu Pro Glu Glu Val Glu Glu Val Ile 20 25 30 Gly Asn Lys Pro Glu Ser Asp Ile Leu Val His Thr Ala Tyr Asp Glu 35 40 45 Ser Thr Asp Glu Asn Val Met Leu Leu Thr Ser Asp Ala Pro Glu Tyr 50 55 60 Lys Pro Trp Ala Leu Val Ile Gln Asp Ser Asn Gly Glu Asn Lys Ile 65 70 75 80 Lys Met Leu <210> 7 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> CCR5-sg1p <400> 7 tgacatcaat fathers cgg <210> 8 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> FANCF-Mb <400> 8 aagttcgcta atcccggaac tgg <210> 9 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> MAGEA1-Mb <400> 9 gctggcagca agggcggcgc tgg <210> 10 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> LIKE site25 <400> 10 agtaaacaaa gcatagactg agg <210> 11 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> PPP1R12C site6 <400> 11 ggggctcaac atcggaagag ggg <210> 12 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> CTLA-sg2 <400> 12 tgtgaagagc ttcactgagt agg 23 <210> 13 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> FANCF-sg1 <400> 13 cgaccaatag cattgcagag agg 23 <210> 14 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> PPP1R12C site3 <400> 14 gagctcactg aacgctggca tgg 23 <210> 15 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> EGFR-sg12 <400> 15 atacaccgtg ccgaacgcac cgg 23
Claims
1. A fusion protein for editing adenine, characterized in that, The fusion protein comprises a TadA8e fragment and a Cas9n fragment sequentially from its N-terminus to its C-terminus. The amino acid sequence of the TadA8e fragment is shown in SEQ ID NO: 1, and the Cas9n fragment is a Cas9 protein fragment with single-chain cleavage activity. The amino acid sequence of the Cas9n fragment is shown in SEQ ID NO:
4. The fusion protein further includes a Rad51 fragment, which is adjacent to the Cas9n fragment and located at the N-terminus or C-terminus of the Cas9n fragment; the amino acid sequence of the Rad51 fragment is shown in SEQ ID NO:
2.
2. The fusion protein as described in claim 1, characterized in that, The fusion protein also includes a nuclear localization signal fragment located at the N-terminus and / or C-terminus of the fusion protein.
3. The fusion protein as described in claim 2, characterized in that, The amino acid sequence of the nuclear localization signal fragment is shown in SEQ ID NO:
3.
4. A fusion protein for editing adenine and cytosine, characterized in that, The fusion protein comprises, from its N-terminus to its C-terminus, an hAID fragment, a TadA8e fragment, a Cas9n fragment, and a UGI fragment. The amino acid sequence of the hAID fragment is shown in SEQ ID NO: 5, the amino acid sequence of the TadA8e fragment is shown in SEQ ID NO: 1, and the Cas9n fragment is a Cas9 protein fragment with single-strand cleavage activity; wherein, the amino acid sequence of the Cas9n fragment is shown in SEQ ID NO:
4. The fusion protein further includes a Rad51 fragment, which is adjacent to the Cas9n fragment and located at the N-terminus of the Cas9n fragment; the amino acid sequence of the Rad51 fragment is shown in SEQ ID NO:
2.
5. The fusion protein as described in claim 4, characterized in that, The fusion protein also includes a nuclear localization signal fragment.
6. The fusion protein as described in claim 5, characterized in that, The nuclear localization signal fragment is located at the N-terminus and / or C-terminus of the fusion protein.
7. The fusion protein as described in claim 5 or 6, characterized in that, The UGI fragment comprises two copies; and / or the amino acid sequence of the nuclear localization signal fragment is as shown in SEQ ID NO:
3.
8. The fusion protein as described in claim 7, characterized in that, Each copy of the UGI fragment has an amino acid sequence as shown in SEQ ID NO:
6.
9. A polynucleotide, characterized in that, The polynucleotide encodes the fusion protein as described in any one of claims 1-8.
10. A recombinant expression vector, characterized in that, The recombinant expression vector comprises the polynucleotide as described in claim 9.
11. A base editing system, characterized in that, The base editing system includes a base editor and sgRNA; the base editor comprises a fusion protein as described in any one of claims 1-8.
12. A pharmaceutical composition, characterized in that, The pharmaceutical composition comprises the fusion protein as described in any one of claims 1-8, or the base editing system as described in claim 11.
13. A base editing method for non-therapeutic purposes, characterized in that, The base editing method includes base editing in target cells using a fusion protein as described in any one of claims 1-8, a polynucleotide as described in claim 9, a recombinant expression vector as described in claim 10, or a base editing system as described in claim 11.
14. The base editing method as described in claim 13, characterized in that, The target cells are prokaryotic cells or eukaryotic cells.
15. The base editing method as described in claim 14, characterized in that, The eukaryotic cells are animal cells or plant cells.
16. The use of the fusion protein of any one of claims 1-8, the polynucleotide of claim 9, the recombinant expression vector of claim 10, the base editing system of claim 11, or the pharmaceutical composition of claim 12 in the preparation of base editing reagents, the construction of animal models, or the cultivation of new plant varieties.
Citation Information
Patent Citations
Efficient plant wide-targeting adenine single-base editor and construction and application thereof
CN114524879A