Highly efficient base editing system, preparation method therefor and use thereof
By inserting editing enzyme domains into nuclease domains and screening deaminases using transposon systems, a highly efficient base editor was developed, solving the problems of insufficient editing efficiency and scope of existing base editors and achieving more efficient gene editing.
Patent Information
- Application Number
- PCT/CN2025/100627
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-17
- Filing Date
- 2025-06-12
- Publication Date
- 2025-12-26
AI Technical Summary
Existing base editors have shortcomings in terms of editing efficiency and scope, making it difficult to effectively expand the editing scope and improve editing efficiency, and they also pose risks of off-target effects and side effects.
By fusing nuclease domains with editing enzyme domains to form fusion proteins with specific insertion sites, and using transposon systems to randomly insert deaminases, novel base editors with higher editing efficiency and different editing windows can be screened out.
It achieves higher editing efficiency and a wider editing range, reduces off-target risks, and enhances the application potential of gene therapy.
Smart Images

Figure PCTCN2025100627-FTAPPB-I100001 
Figure PCTCN2025100627-FTAPPB-I100002 
Figure PCTCN2025100627-FTAPPB-I100003
Abstract
Description
High-efficiency base editing system, preparation method and application thereof TECHNICAL FIELD
[0001] The present application relates to the field of gene therapy. Specifically, it relates to a high-efficiency base editing system, preparation method and application thereof. BACKGROUND
[0002] The genome editing technology CRISPR-Cas9 was listed as the "Science Annual Top Ten Breakthroughs" by the Science magazine in 2013, which opened the curtain of the era of gene editing. The CRISPR-Cas9 technology produces a DNA double strand break (DSB) at the target site through sgRNA guidance, thereby inducing the homology directed repair (HDR) and non-homologous end joining (NHEJ) repair pathways in cells, and then realizing the modification of the genomic DNA such as insertion, knockout and replacement. However, the occurrence of DSB in cells will not only cause many unexpected gene editing such as large fragment loss, inversion, transposition and the like, but also activate the expression of the tumor suppressor gene p53 and promote the enrichment of mutant p53 cells. In addition, CRISPR-Cas9 also has a relatively serious off-target effect, which may cut the DNA double strand at the mispositioned gene site, thereby causing potential Off-target risk, which is also a major factor limiting the clinical application of CRISPR-Cas9 gene editing.
[0003] Single-nucleotide variants (SNVs) is a new generation of gene editing technology based on CRISPR-Cas9, known as the second generation of gene editing technology, which can accurately and efficiently replace the base without causing DNA double-strand break. By combining different deaminases with nuclease Cas protein, cytosine base editor (CBE) and adenine base editor (ABE) can be constructed to induce cytosine (C) to be replaced by uracil (U) and adenine (A) to be replaced by guanine (G). CBE is a fusion protein composed of nickase Cas9 protein (Nickase Cas9(D10A), nCas9), cytosine deaminase and uracil DNA glycosylase inhibitor (UGI). Its principle is that under the guidance of sgRNA, CBE binds to the target site on the genomic DNA. Under the action of nCas9 helicase, through the complementary pairing of sgRNA and the base sequence on the genomic DNA to form "R-Loop", a single-stranded DNA (ssDNA) window is generated. Under the joint action of deaminase and uracil DNA glycosylase inhibitor (UGI), cytosine (C) in a specific range of ssDNA is deaminated to uracil (U), and then (U) is converted to thymine (T) through the repair mechanism in the cell, thereby realizing the replacement of C·G base pair to T·A base pair. Similarly, ABE is mainly composed of nCas9 and artificially directed evolution of adenine deaminase to form a fusion protein. When the fusion protein targets the genomic DNA under the guidance of sgRNA, adenine deaminase can bind to ssDNA to deaminate adenine (A) in a specific range to inosine (I), which will be read as G at the DNA level and replicated, and ultimately realize the direct replacement of A·T base pair to G·C base pair. Whether CBE or ABE, when replacing the target nucleotide, it will not cause DNA double-strand break, thus effectively reducing the off-target risk on the genome, and has great application potential in the field of gene therapy.
[0004] The most widely used mainstream base editor editing window is at +4~+13 positions (the far PAM end is +1 position), and most research directions are focused on how to narrow the editing window and reduce bystander mutations (Rees and Liu, 2018). There is no relevant research report on how to expand the editing range of the base editor and improve the editing efficiency within the editing range. Although the editing window can be adjusted by changing the length of sgRNA, considering the editing efficiency and off-target effect of specific sgRNA, this adjustment method is not always effective. The expression regulation of genes and the splicing regulation element and mechanism of mRNA are not very clear, and sometimes it is difficult to realize the regulation of gene expression and splicing through single site base mutation, especially under the premise that the splicing regulation element and mechanism are not clear, it is often necessary to simultaneously mutate multiple nucleotides within a certain range to realize the regulation of gene expression and splicing, and multiple sgRNAs are required to edit simultaneously through the existing base editor, which not only increases the difficulty and cost of the experiment, but also brings many uncertain side effects, such as off-target. Therefore, how to improve the base editing efficiency while further expanding the base editing range and greatly improving the application space of base editors in splicing regulation has become another important direction of base editor development. On the basis of the existing base editor, although a large number of base editor mutants can be generated through error-prone PCR, saturation mutation, genome rearrangement, rational design and other protein evolution methods, because the functional elements included in the base editor are relatively independent, it is difficult to obtain effective base editors by using these evolution methods.
[0005] Therefore, there is an urgent need in the art to develop a new type of base editor with higher base editing efficiency and wider or narrower editing range. SUMMARY
[0006] The purpose of the present application is to provide a high-efficiency base editor with higher editing efficiency and wider editing range.
[0007] In a first aspect of the present application, a fusion protein is provided, comprising a nuclease domain and an editing enzyme domain; the editing enzyme domain is inserted into an insertion site in the nuclease domain, wherein the nuclease domain in the fusion protein has nuclease function; and the editing enzyme domain in the fusion protein has editing enzyme function.
[0008] The nuclease domain is derived from a Cas protein or a Cas protein mutant.
[0009] In another preferred example, the fusion protein has the structure of formula I: W0-Z1-L1-Y-L2-Z2-W1-W2 (I)
[0010] wherein,
[0011] W0 is absent or a nuclear localization element;
[0012] Z1 is a left element of a nuclease domain;
[0013] Z2 is a right element of a nuclease domain;
[0014] L1 is absent or a linker peptide element;
[0015] L2 is absent or a linker peptide element;
[0016] Y is an editing enzyme domain element;
[0017] W1 is absent or a uracil glycosylase inhibitor (UGI) element;
[0018] W2 is absent or a nuclear localization element;
[0019] each “-” is independently a chemical bond;
[0020] wherein, Z1 and Z2 are respectively a left element and a right element of an editing enzyme domain formed by the insertion site splitting the editing enzyme domain.
[0021] In another preferred embodiment, the Cas protein or Cas protein mutant is derived from nSaCas9(D10A)KKH.
[0022] In another preferred embodiment, the amino acid sequence of the nSaCas9(D10A)KKH is set forth in SEQ ID NO: 41.
[0023] In another preferred embodiment, the nuclease domain comprises an amino acid sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 41.
[0024] In another preferred embodiment, the insertion site of the editing enzyme domain inserted into the nuclease domain is selected from any one of the following: 361st, 408th, 455th, 469th, 471st, 588th, 634th, 719th, 757th, 928th, and 941st positions of the amino acid sequence set forth in SEQ ID NO: 41.
[0025] In another preferred embodiment, the insertion site of the editing enzyme domain into the nuclease domain is selected from the group consisting of: corresponding to position 361, 408, 588, 634, 719, 757 of the amino acid sequence as set forth in SEQ ID NO: 41.
[0026] In another preferred embodiment, the insertion site is selected from the group consisting of: corresponding to the amino acid sequences YQS and EDI, DEL and HTN, RSF and QSI, KKY and LPN, YGL and NDI, NRT and FQY, DIN and FSV, KEW and KLD, FIT and HQI, FVT and KNL, NYY and VNS, respectively, flanking the insertion site.
[0027] In another preferred embodiment, the insertion site is selected from the group consisting of: RuvC-II, RuvC-III.
[0028] In another preferred embodiment, the editing enzyme domain is inserted into the nuclease domain by a transposon.
[0029] In another preferred embodiment, the editing enzyme is selected from at least one or more of deaminase, DNA glycosylase, methylase and acetylase.
[0030] In another preferred embodiment, the deaminase is selected from the group consisting of APOBEC / AID family, TadA and ADAR family deaminase and deaminase mutants thereof.
[0031] In another preferred embodiment, the APOBEC / AID family includes, but is not limited to, APOBEC1, CDA, AID and APOBEC3 (A3A, A3B, A3C, A3D, A3F, A3G, A3H), and homologues or mutants thereof.
[0032] In another preferred embodiment, the amino acid sequence of the editing enzyme comprises any one selected from the group consisting of the amino acid sequences as set forth in SEQ ID NO: 1 to SEQ ID NO: 20.
[0033] In another preferred embodiment, the amino acid sequence of the editing enzyme is selected from the group consisting of:
[0034] the amino acid sequence as set forth in SEQ ID NO: 1, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 13 or SEQ ID NO: 16.
[0035] In another preferred embodiment, the "having nuclease function" means that the fusion protein has ≥ 70% (preferably 70-200%, more preferably 80-150%) of the nuclease activity of the original nuclease domain.
[0036] In another preferred embodiment, the "having an editing enzyme function" means that the fusion protein has ≥70% (preferably 70-200%, more preferably 80-150%) of the editing enzyme activity of the original editing enzyme domain.
[0037] In another preferred embodiment, the editing enzyme domain is connected to the nuclease domain before and after the insertion site, respectively, by a linker.
[0038] In another preferred embodiment, the linker is a flexible linker.
[0039] In another preferred embodiment, the amino acid sequence of the linker is selected from the group consisting of:
[0040] the amino acid sequence as set forth in any one of SEQ ID NO: 25 to SEQ ID NO: 28, or a combination thereof.
[0041] In another preferred embodiment, the amino acid sequence of the linker is as set forth in SEQ ID NO: 25.
[0042] In another preferred embodiment, the fusion protein further comprises a uracil glycosylase inhibitor (UGI) domain.
[0043] In another preferred embodiment, the UGI domain comprises an amino acid sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence as set forth in SEQ ID NO: 42.
[0044] In another preferred embodiment, the fusion protein further comprises a nuclear localization (NLS) domain comprising an amino acid sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the amino acid sequence as set forth in SEQ ID NO: 43.
[0045] In another preferred embodiment, the amino acid sequence of the fusion protein is selected from any one of the amino acid sequences as set forth in SEQ ID NO: 35-37, SEQ ID NO: 44-49.
[0046] In another preferred embodiment, the amino acid sequence of the fusion protein is selected from any one of the amino acid sequences as set forth in SEQ ID NO: 35, EQ ID NO: 37, SEQ ID NO: 48, or SEQ ID NO: 49.
[0047] In another preferred embodiment, the amino acid sequence of the fusion protein is selected from any one of the amino acid sequences as set forth in SEQ ID NO: 35, SEQ ID NO: 48, or SEQ ID NO: 49.
[0048] In another preferred embodiment, the amino acid sequence of the fusion protein is selected from the amino acid sequence as set forth in SEQ ID NO: 35.
[0049] In another preferred embodiment, the amino acid sequence of the fusion protein is selected from the amino acid sequence as set forth in SEQ ID NO: 48.
[0050] In a second aspect of the present application, there is provided a base editing system comprising the fusion protein of the first aspect of the present application or a polynucleotide encoding the same.
[0051] In another preferred embodiment, the base editing system further comprises a guide RNA.
[0052] In another preferred embodiment, the base editing system is a cytosine base editing system.
[0053] In another preferred embodiment, the cytosine base editing system comprises the fusion protein of the first aspect of the present application, wherein the editing enzyme is a cytosine deaminase.
[0054] The guide RNA targets the fusion protein to a target sequence.
[0055] In another preferred embodiment, the sequence of the guide RNA is selected from any one of the nucleotide sequences as set forth in SEQ ID NO: 21, SEQ ID NO: 24, SEQ ID NO: 29-33, or SEQ ID NO: 38.
[0056] In another preferred embodiment, the sequence of the guide RNA is as set forth in SEQ ID NO: 38.
[0057] In a third aspect of the present application, there is provided a polynucleotide, the sequence of which encodes the fusion protein of the first aspect of the present application or the base editing system of the second aspect of the present application.
[0058] In a fourth aspect of the present application, there is provided a vector comprising the polynucleotide of the third aspect of the present application.
[0059] In another preferred embodiment, the vector comprises one or more promoters operably linked to the nucleic acid sequence, an enhancer, a transcription termination signal, a polyadenylation sequence, an origin of replication, a selectable marker, a nucleic acid restriction site, and / or a homologous recombination site.
[0060] In another preferred embodiment, the vector comprises a plasmid, an mRNA, a viral vector.
[0061] In another preferred embodiment, the viral vector is selected from the group consisting of an adeno-associated virus (AAV), an adenovirus, a lentivirus, a viroid, a retrovirus, a herpes virus, an SV40, a poxvirus, or a combination thereof.
[0062] In another preferred embodiment, the vector comprises an expression vector, a shuttle vector, an integration vector.
[0063] In a fifth aspect of the present application, a host cell is provided, wherein the host cell comprises the vector of the fourth aspect of the present application, or the genome of the host cell has integrated the polynucleotide of the third aspect of the present application.
[0064] In another preferred embodiment, the host cell is a eukaryotic cell, such as a yeast cell, a plant cell, or a mammalian cell (including human and non-human mammal).
[0065] In another preferred embodiment, the host cell is a prokaryotic cell, such as E. coli.
[0066] In another preferred embodiment, the yeast cell is selected from one or more sources of yeast from the group consisting of Pichia pastoris, Kluyveromyces, or a combination thereof; preferably, the yeast cell comprises Kluyveromyces, more preferably Kluyveromyces marxianus, and / or Kluyveromyces lactis.
[0067] In another preferred embodiment, the host cell is selected from the group consisting of E. coli, a wheat germ cell, an insect cell, SF9, Hela, HEK293T, CHO, a yeast cell, or a combination thereof.
[0068] In a sixth aspect of the present application, a method for producing the fusion protein of the first aspect of the present application is provided, wherein the method comprises the steps of:
[0069] culturing the host cell of the fifth aspect of the present application under conditions suitable for expression, thereby expressing the fusion protein; and / or, isolating the fusion protein.
[0070] In a seventh aspect of the present application, a gene editing reagent is provided, wherein the gene editing reagent comprises the fusion protein of the first aspect of the present application.
[0071] In another preferred embodiment, the reagent further comprises one or more reagents selected from the group consisting of:
[0072] (a1) an sgRNA, a crRNA, or a vector for producing the sgRNA or crRNA;
[0073] (a2) a template for homology directed repair: a single-stranded nucleotide sequence or a plasmid vector.
[0074] In an eighth aspect of the present application, a kit is provided, comprising the gene editing reagent of the sixth aspect of the present application.
[0075] In another preferred embodiment, the kit further comprises one or more reagents selected from the group consisting of:
[0076] (a1) sgRNA, crRNA, or a vector for producing the sgRNA or crRNA;
[0077] (a2) a template for homology-directed repair: a single-stranded nucleotide sequence or a plasmid vector.
[0078] In another preferred embodiment, the kit further comprises a label or an instruction.
[0079] In a ninth aspect of the present application, a pharmaceutical composition is provided, comprising:
[0080] (a) the fusion protein of the first aspect of the present application, or a gene encoding the same, or an expression vector thereof; and
[0081] (b) a pharmaceutically acceptable carrier.
[0082] In another preferred embodiment, the expression vector comprises a viral vector.
[0083] In another preferred embodiment, the viral vector is selected from the group consisting of an adeno-associated virus (AAV), an adenovirus, a lentivirus, a viroid, a retrovirus, a herpesvirus, an SV40, a poxvirus, or a combination thereof.
[0084] In another preferred embodiment, the vector is selected from the group consisting of a lentivirus, an adenovirus, an adeno-associated virus (AAV), or a combination thereof, preferably the vector is an adeno-associated virus (AAV).
[0085] In another preferred embodiment, the dosage form of the pharmaceutical composition is selected from the group consisting of a lyophilized formulation, a liquid formulation, or a combination thereof.
[0086] In another preferred embodiment, the dosage form of the pharmaceutical composition is an injection dosage form.
[0087] In another preferred embodiment, the pharmaceutical composition further comprises other drugs for gene therapy.
[0088] In a tenth aspect of the present application, a kit is provided, comprising:
[0089] (a1) a first container, and the fusion protein of the first aspect of the present application, or a gene encoding the same, or an expression vector thereof, or a drug containing the fusion protein of the first aspect of the present application, in the first container.
[0090] In another preferred embodiment, the kit further comprises:
[0091] (a2) a second container, and other drugs for gene therapy or containing other drugs for gene therapy in the second container.
[0092] In another preferred embodiment, the first container and the second container are the same or different containers.
[0093] In another preferred embodiment, the drug in the first container is a single preparation containing the fusion protein of the first aspect of the present application.
[0094] In another preferred embodiment, the drug in the second container is a single preparation containing other drugs for gene therapy.
[0095] In another preferred embodiment, the dosage form of the drug is selected from the group consisting of a lyophilized preparation, a liquid preparation, or a combination thereof.
[0096] In another preferred embodiment, the dosage form of the drug is an injection dosage form.
[0097] In the eleventh aspect of the present application, there is provided a use of the fusion protein of the first aspect of the present application, the base editing system of the second aspect of the present application, the polynucleotide of the third aspect of the present application, the vector of the fourth aspect of the present application, or the host cell of the fifth aspect of the present application in the preparation of a medicament for gene editing.
[0098] In the twelfth aspect of the present application, there is provided a use of the fusion protein of the first aspect of the present application, the base editing system of the second aspect of the present application, the polynucleotide of the third aspect of the present application, the vector of the fourth aspect of the present application, or the host cell of the fifth aspect of the present application in the preparation of a reagent or kit for gene editing, for improving the efficiency of gene editing.
[0099] In the thirteenth aspect of the present application, there is provided a method of gene editing, comprising administering to a target cell the fusion protein of the first aspect of the present application, the base editing system of the second aspect of the present application, the polynucleotide of the third aspect of the present application, or the vector of the fourth aspect of the present application.
[0100] In another preferred embodiment, the method of gene editing is carried out in vivo or in vitro.
[0101] In another preferred embodiment, the target cell is a eukaryotic cell.
[0102] In another preferred embodiment, the method of base editing is for non-therapeutic purposes.
[0103] In the fourteenth aspect of the present application, there is provided a method for improving the efficiency of gene editing, comprising the steps of: a) providing a fusion protein of the first aspect of the present application, a base editing system of the second aspect of the present application, a polynucleotide of the third aspect of the present application, a vector of the fourth aspect of the present application, or a host cell of the fifth aspect of the present application; and b) using the fusion protein, the base editing system, the polynucleotide, the vector, or the host cell in the preparation of a reagent or kit for gene editing, for improving the efficiency of gene editing.
[0104] In the presence of the fusion protein of the first aspect of the present application or the gene editing reagent of the seventh aspect of the present application, the cell is subjected to gene editing, thereby improving the efficiency of gene editing.
[0105] In another preferred embodiment, the cell comprises a human or non-human mammalian cell (e.g., a primate or livestock).
[0106] In another preferred embodiment, the cell comprises a cancer cell or a normal cell.
[0107] In another preferred embodiment, the cell is selected from the group consisting of a kidney cell, a liver cell, a neural cell, a cardiac cell, an epithelial cell, a muscle cell, a somatic cell, a bone marrow cell, an endothelial cell, or a combination thereof.
[0108] In another preferred embodiment, the cell is selected from the group consisting of a 293 cell, an A549 cell, an SW626 cell, an HT-3 cell, a PA-1 cell, a K562 cell, an AC16 cell, a WM115 cell, a U87 cell, a pluripotent induced stem cell, or a combination thereof.
[0109] In another preferred embodiment, the cell comprises a HEK293T.
[0110] It should be understood that, within the scope of the present application, each of the technical features of the present application described above and each of the technical features specifically described hereinafter (e.g., in the examples) can be combined with each other to form a new or preferred technical solution. Due to the limited space, they will not be listed one by one here. BRIEF DESCRIPTION OF DRAWINGS
[0111] FIG. 1 shows a schematic diagram of TAM structure in which AID is located close to the N-terminus of nSaCas9 (D10A)KKH.
[0112] FIG. 2 shows a schematic diagram of a site targeted by sgRNA-1 at the junction between exon 50 and intron 50 of human DMD gene.
[0113] FIG. 3 shows a graph of editing efficiency of TAMs constructed by different AID mutants (mAID) on DMD gene, wherein, panel A represents editing efficiency of hAIDx and AID mutants (mAID-1 to mAID-11) numbered 1 to 11; panel B represents editing efficiency of AID mutants (mAID-12 to mAID-19) numbered 12 to 19; C2 represents the 2nd C targeted by sgRNA, i.e., C2, and similar descriptions are also provided for C6, C10, C12 and C13; * represents P<0.05, and ** represents P<0.01 (n=2, two-way ANOVA).
[0114] Figure 4 shows schematic diagram of plasmid structure of prokaryotic transposition screening system, (A) pGAT-Cas9-UGI and (B) pGAT-Kana-sgRNA; wherein, AmpR: ampicillin resistance gene; LacI: lactose operon repressor protein; nSaCas9(D10A)KKH: SaCas9KKH mutant with 10th aspartic acid mutated to alanine; UGI: uracil glycosylase inhibitor; SmR: streptomycin resistance gene; KanR(D208G): inactivated kanamycin resistance gene with D208G being 208th aspartic acid mutated to glycine; sgRNA: sgRNA expression element against kanamycin resistance gene D208G.
[0115] Figure 5 shows schematic diagram of screening process of prokaryotic transposition system, wherein, AmpR: ampicillin resistance gene; LacI: lactose operon repressor protein; nSaCas9(D10A)KKH: SaCas9KKH mutant with 10th aspartic acid mutated to alanine; UGI: uracil glycosylase inhibitor; SmR: streptomycin resistance gene; KanR(D208G): inactivated kanamycin resistance gene with D208G being 208th aspartic acid mutated to glycine; sgRNA: sgRNA expression element against kanamycin resistance gene D208G; AID and its mutants are randomly inserted into pGAT-Cas9-UGI vector by transposase to construct the library of EB-TAM. EB-TAM is then co-transformed into E. coli with pGAT-Kana-sgRNA. After IPTG induction, colonies growing on plates containing kanamycin indicate functional EB-TAMs that have restored the activity of kanamycin gene. Finally, the insertion site of AID is obtained by sequencing.
[0116] Figure 6 shows schematic diagram of deaminase insertion site in EB-TAM.
[0117] Figure 7 shows comparison of different Cas9 domains, wherein, Arg: arginine-rich helix; NUC: nuclease lobe; REC: recognition lobe; L-I: linker I; L-II: linker II; CTD: C-terminal domain.
[0118] FIG. 8 shows the editing efficiency and editing window of different EB-TAMs on target sites Site1-5, wherein (A, B) show the editing efficiency and editing window of each site of Site1; (C, D) show the editing efficiency and editing window of each site of Site2; (E, F) show the editing efficiency and editing window of each site of Site3; (G, H) show the editing efficiency and editing window of each site of Site4; (I, J) show the editing efficiency and editing window of each site of Site5. * represents P<0.05, ** represents P<0.01 (n=2, two-way ANOVA).
[0119] FIG. 9 shows the editing window mode diagram of different EB-TAMs, and the efficiency is the average value of each site edited by different EB-TAMs on Site1-5; wherein light gray represents the average editing efficiency of 6-10; dark gray represents the average editing efficiency of >10.
[0120] FIG. 10 shows the schematic diagram of sgRNA-8 target site.
[0121] FIG. 11 shows the editing efficiency of different EB-TAMs on DMD Exon 53 target sites, wherein C1 represents that sgRNA targets the first C, i.e. C1, and C5 and C8 have similar descriptions; * represents P<0.05, ** represents P<0.01 (n=2, two-way ANOVA).
[0122] FIG. 12 shows the editing efficiency of EB-TAM757-L3 after replacing different linkers (L1, L2) on Exon 53 target sites; wherein C1 represents that sgRNA targets the first C, i.e. C1, and C5 and C8 have similar descriptions.
[0123] FIG. 13 shows the electropherogram of EB-TAM757-L1 induced DMD Exon 53 skipping, wherein band 1 represents band 1 without DMD Exon 53 skipping, and band 2 represents band 2 of DMD Exon 53 skipping.
[0124] FIG. 14 shows the sequencing diagram of EB-TAM757-L1 induced DMD Exon 53 skipping band 2. DETAILED DESCRIPTION
[0125] Through extensive and in-depth research, the present inventors have unexpectedly developed a class of TAM base editors with unique structures. The TAM base editor of the present application is formed by inserting a nucleic acid editing enzyme into a specific site of a nuclease. Experiments have shown that when the nucleic acid editing enzyme is inserted into these specific sites of the nuclease, not only does it not destroy the function of the nuclease, but it also enables the formed editor of the present application to have unique editing windows and other improved properties. On this basis, the present application is completed.
[0126] Specifically, the present inventors screened AID mutants with higher or lower deaminase activity in eukaryotic cells using the TAM base editor structure in which AID is located at the N-terminal of nSaCas9(D10A)KKH. Then, hAIDx was randomly inserted into a recipient plasmid composed of nSaCas9(D10A)KKH and UGI using a transposon system, and new base editors EB-TAM with narrower (such as EB-TAM361-L3) or wider (such as EB-TAM757-L3) editing windows were obtained by screening. Then, the functions of the new base editors EB-TAM were systematically verified in eukaryotic cells, and the performance of EB-TAM was further improved by optimizing the linker, and EB-TAM cytosine base editors with higher editing activity than eTAM and different editing windows were obtained.
[0127] Terms
[0128] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.
[0129] The term "about" can mean a value or composition that is within an acceptable error range for the particular value or composition as determined by one of ordinary skill in the art, varying depending on how the value or composition is measured or determined.
[0130] As used herein, the term "containing" or "including" can be open, semi-closed and closed. In other words, the term also includes "consisting essentially of" or "consisting of".
[0131] Sequence identity (or homology) is determined by comparing two aligned sequences over a window of comparison that can be 50%, 60%, 70%, 80%, 90%, 95%, or 100% of the length of the reference nucleotide sequence or protein, and determining the number of positions at which the same residue occurs in both sequences. Typically, this is expressed as a percentage. Measurement of sequence identity of nucleotide sequences is a method well known to those skilled in the art.
[0132] Base editor
[0133] Any of the base editors provided herein are capable of modifying a particular nucleotide base without producing a significant proportion of insertions / deletions. As used herein, "insertion / deletion" refers to the insertion or deletion of a nucleotide base within a nucleic acid. Such insertions or deletions can result in a frameshift mutation within a gene coding region. In some embodiments, it is desirable to produce an effective modification (e.g., a mutation or deamination) of a particular nucleotide within a nucleic acid without producing a substantial number of insertions or deletions (i.e., indels) in the nucleic acid. In certain embodiments, any of the base editors provided herein are capable of producing a greater proportion of intended modifications (e.g., point mutations or deaminations) relative to insertions / deletions.
[0134] Any of the base editors of the present application are capable of efficiently producing an intended mutation, such as a point mutation, in a nucleic acid (e.g., within a genome) without producing a substantial number of unintended mutations, such as unintended point mutations.
[0135] Base editors have a wide range of applications, such as for treating diseases caused by single point mutations, such as sickle cell anemia. In sickle cell disease, the glutamic acid at the sixth position of the beta-globin is mutated to valine, and while current base editors are not able to mutate the valine back to glutamic acid, they can convert it to alanine, a naturally occurring non-pathogenic variant, restoring normal beta-globin function. Base editors can also be used to disrupt the binding regions of cis-regulatory elements or trans-acting factors on DNA, such as promoters, enhancers, and silencers, to regulate the transcription of related genes, but editing a single base alone often cannot achieve the regulatory effect. The requirements for base editors vary for different application scenarios. For example, for mutation repair at a specific site, the smaller the editing window of the base editor, the better, and it is best to target only the mutation site to prevent the generation of bystander mutations. For regulatory elements that need to be disrupted, the exact regulatory site cannot be determined, and a larger editing window can be needed to effectively disrupt these regulatory elements.
[0136] In recent years, base editors have been reported to be used for splicing regulation. By effectively mutating splicing regulation elements (SREs), the isoforms of the expression products of related genes can be transformed, or the splicing abnormality caused by splicing regulation can be knocked down the expression of genes. RNA splicing is a key step in the process of gene expression. It connects all exon sequences together by removing intron sequences in genes, and processes pre-messenger RNA (pre-mRNA) into mature messenger RNA (mRNA). The mRNA enters the cytoplasmic matrix through the nuclear pore, and becomes the template for protein synthesis. The boundaries of introns in the mRNA precursors of most genes have common conserved sequences, and the conserved sequence pattern is called the GU-AG rule, also known as the Chambon rule. It means that the 5' splice site (5'ss) sequence at the 5' end of the intron is GU, and the 3' splice site (3'ss) sequence at the 3' end is AG. In addition to the conserved GU-AG, other SRE elements near the splicing site also play a further regulatory role in splicing. For example, exonic splicing enhancer (ESE), intronic splicing enhancer (ISE), exonic splicing silencer (ESS), intronic splicing silencer (ISS), and branch point site located 18-40 nt upstream of 3'ss, etc. These cis-acting elements recruit splicing factors to promote or inhibit the recognition of nearby splicing sites. Trans-acting elements include two main families, namely the serine / arginine-rich protein family and the heterogeneous nuclear ribonucleoprotein (hnRNP), which can bind to pre-mRNA to regulate the assembly of the spliceosome and the recognition of the splicing site. By editing these sites, splicing can be enhanced or weakened. However, in mammals, these sites are severely degenerated, so effective mutations of these sites are needed to regulate splicing. However, the current base editor editing range is difficult to effectively mutate these sites for splicing regulation.
[0137] On the basis of existing base editors, although a large number of base editor mutants can be generated by error-prone PCR, saturation mutation, genome rearrangement, rational design and other protein evolution methods, due to the relative independence of each functional element included in the base editor, it is difficult to obtain effective base editors by using these evolution methods. Transposon is a piece of DNA sequence, which is a basic unit that can autonomously displace. This displacement does not depend on the homology of the sequence, and is also called transposable element or jumping gene. Transposon can transfer genes between different sites on the same chromosome or between different chromosomes, causing mutations and gene rearrangements, and increasing the polymorphism of genes. In recent years, it has been reported that SpCas9 can be inserted into multiple sites by using transposition system without affecting its nuclease activity (Liu et al., 2020; Oakes et al., 2016). Among them, ABE screened by inserting adenine deaminase TadA into nSpCas9 can reduce off-target without affecting the editing effect of the target site. Therefore, by using the random insertion characteristics of transposon, deaminase can be randomly inserted into nuclease by transposition system to form a new fusion protein, which can effectively ensure the activity of deaminase and the PAM sequence recognized by nuclease, and also can change the spatial conformation of nuclease, thereby affecting the editing efficiency and editing window. Finally, through a suitable screening system, a new type of base editor that meets the requirements will be screened, which will be an effective screening method.
[0138] Deaminase AID was first discovered in activated B cells, which plays an important role in the body's immune system. When AID binds to the variable region of the immunoglobulin gene, it deaminates cytosine in ssDNA to uracil nucleoside, and then uses the body's damage repair mechanism to produce high-frequency mutations in the variable region to screen antibodies with high affinity. Compared with other cytidine deaminases such as APOBEC, AID only acts on ssDNA and does not cause off-target editing at the RNA level, and its action motif is WRC (W: A / T, R: A / G), which better matches the GU-AG sequence of exons and introns, and is the most suitable deaminase for splicing regulation of DNA cytosine base editor. Targeted AID-mediated Mutagenesis (TAM) is a general term for cytosine base editors using AID as a deaminase, in which AID is connected to the N-terminal or C-terminal of Cas9, both of which show high base editing activity. At present, there have been reports of using AID mutants (hAIDx) connected to the N-terminal of nSaCas9 (D10A) KKH to form Enhanced TAM (eTAM) to edit specific sites of DNA in vitro and in vivo with high efficiency, but its editing window is +1 ~ +12 positions upstream of PAM (Yuan et al., 2018). Since the selection of sgRNA is limited by the PAM sequence, there is often no suitable PAM sequence near the GU-AG sequence, which seriously limits the application of TAM base editor in gene regulation splicing. The present application uses a transposon random insertion system (Embedding system) to randomly insert hAIDx into nSaCas9 (D10A) KKH, and through a screening system, a new type of cytosine base editor EB-TAM with higher base editing efficiency, wider or narrower editing range is successfully screened.
[0139] Fusion protein
[0140] The fusion protein provided herein comprises a nuclease domain and an editing enzyme domain; the editing enzyme domain is inserted into the nuclease domain, and the insertion site of the editing enzyme domain in the nuclease domain is selected from any one of the following: relative to the 361st, 408th, 455th, 469th, 471st, 588th, 634th, 719th, 757th, 928th or 941st position of the amino acid sequence shown as SEQ ID NO: 41.
[0141] A preferred editing enzyme is selected from at least one or more of deaminase, DNA glycosylase, methylase and acetylase.
[0142] In another embodiment, the deaminase is selected from the group consisting of APOBEC / AID family, TadA and ADAR family deaminases and deaminase mutants thereof.
[0143] In another embodiment, the APOBEC / AID family includes, but is not limited to, APOBEC1, CDA, AID and APOBEC3 (A3A, A3B, A3C, A3D, A3F, A3G, A3H), and homologues or mutants thereof.
[0144] A preferred cytosine deaminase has an amino acid sequence selected from the group consisting of:
[0145] any one of the amino acid sequences set forth in SEQ ID NO: 1 to SEQ ID NO: 20.
[0146] In a preferred embodiment, the cytosine deaminase domain is linked to the nuclease domain before and after the insertion site, respectively, by a linker having an amino acid sequence selected from the group consisting of:
[0147] any one or more combinations of the amino acid sequences set forth in SEQ ID NO: 25 to SEQ ID NO: 28.
[0148] The present application also includes active fragments, derivatives and analogs of the fusion proteins described above. As used herein, the terms "fragment", "derivative" and "analog" refer to polypeptides that substantially maintain the function or activity of the fusion proteins of the present application. A polypeptide fragment, derivative or analog of the present application can be (i) a polypeptide having one or several conservative or non-conservative amino acid residue substitutions, preferably conservative amino acid residue substitutions, or (ii) a polypeptide having a substitution group at one or more amino acid residues, or (iii) a polypeptide formed by fusing an antigenic peptide to another compound, such as a compound that prolongs the half-life of the polypeptide, for example, polyethylene glycol, or (iv) a polypeptide formed by fusing an additional amino acid sequence to the polypeptide sequence (a fusion protein formed by fusing a leader sequence, a secretion sequence or a 6His tag sequence). These fragments, derivatives and analogs are within the scope of those skilled in the art according to the teachings herein.
[0149] A preferred active derivative refers to a polypeptide having at most 3, preferably at most 2, more preferably at most 1 amino acid replaced by a similar or a conservative amino acid compared to the amino acid sequence of Formula I. These conservative variant polypeptides are preferably generated by amino acid substitutions according to Table A.
[0150] Table A
[0151] Adeno-associated virus
[0152] Adeno-associated virus (AAV), also known as adeno-associated virus, belongs to the parvovirus family of dependovirus, is a class of simple structure single-stranded DNA defective virus found at present, which needs the participation of helper virus (usually adenovirus) in replication. It encodes cap and rep genes in two terminal inverted repeat sequences (ITRs). ITRs play a decisive role in virus replication and packaging. The cap gene encodes viral capsid protein, and the rep gene is involved in virus replication and integration. AAV can infect a variety of cells.
[0153] In a preferred embodiment of the present application, the vector is a recombinant adeno-associated virus (rAAV). The recombinant adeno-associated virus vector is derived from a non-pathogenic wild-type adeno-associated virus, and has the advantages of stable physical and chemical properties, weak pathogenicity, low integration risk, persistent expression of exogenous genes, and other structural and biological advantages. It has become one of the most widely used vectors in the field of in vivo gene therapy, and has been widely used in gene therapy and vaccine research worldwide. After more than 10 years of research, the biological characteristics of recombinant adeno-associated virus have been well understood, especially its application effect in various cells, tissues and in vivo experiments has accumulated a lot of data. In medical research, recombinant adeno-associated virus has been used for gene therapy research (including in vivo and in vitro experiments) of various diseases; at the same time, as a characteristic gene transfer vector, it is also widely used in gene function research, construction of disease models, preparation of gene knockout mice, etc.
[0154] AAV vectors can be produced using standard methods in the art. Any serotype of adeno-associated virus is suitable. Methods for purifying vectors can be found, for example, in U.S. Patent Nos. 6566118, 6989264, and 6995006, the disclosures of which are incorporated herein by reference in their entireties. The production of hybrid vectors is described, for example, in PCT Application No. PCT / US2005 / 027091, the disclosure of which is incorporated herein by reference in its entirety. The use of vectors derived from AAV for the in vitro and in vivo transfer of genes has been described (see, e.g., International Patent Application Publication Nos. WO 91 / 18088 and WO 93 / 09239; U.S. Patent Nos. 4,797,368, 6,596,535, and 5,139,941, and European Patent No. 0488528, each of which is incorporated herein by reference in its entirety). These patent publications describe various AAV-derived constructs in which the rep and / or cap genes are deleted and replaced with a gene of interest, and the use of these constructs to transfer the gene of interest in vitro (into cultured cells) or in vivo (directly into an organism). Replication-defective recombinant AAVs can be produced by co-transfecting into a cell line infected with a human helper virus (e.g., adenovirus) a plasmid containing a nucleic acid sequence of interest flanked by two AAV inverted terminal repeat (ITR) regions, and a plasmid carrying AAV encapsidation genes (rep and cap genes). The resulting AAV recombinants are then purified by standard techniques.
[0155] In some embodiments, the recombinant vector is encapsidated into a virion (e.g., an AAV virion including, but not limited to, AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV14, AAV15, and AAV16). Accordingly, the present disclosure includes a recombinant virion (recombinant because it comprises a recombinant polynucleotide) containing any of the vectors described herein. Methods of producing such virions are known in the art and are described in U.S. Patent No. 6,596,535.
[0156] Expression vectors and host cells
[0157] The present application also relates to vectors comprising the polynucleotides of the present application, and to host cells which are genetically engineered to incorporate the vectors of the present application or the fusion protein coding sequences of the present application, and to methods of producing the polypeptides of the present application by recombinant techniques.
[0158] The polynucleotide sequences of the present application can be used to express or produce recombinant fusion proteins by conventional recombinant DNA techniques. In general, the following steps are involved:
[0159] (1) transforming or transducing a suitable host cell with a polynucleotide (or variant) encoding the fusion protein of the present application, or with a recombinant expression vector containing the polynucleotide;
[0160] (2) culturing the host cell in a suitable culture medium;
[0161] (3) isolating, purifying the protein from the culture medium or the cell.
[0162] In the present application, the polynucleotide sequence encoding the fusion protein can be inserted into a recombinant expression vector. The term "recombinant expression vector" refers to a plasmid, bacteriophage, yeast plasmid, plant cell virus, mammalian cell virus such as adenovirus, retrovirus, or other vector known in the art. Any plasmid and vector can be used as long as it can replicate and be stable in the host. An important feature of the expression vector is that it usually contains a replication origin, a promoter, a marker gene, and a translation control element.
[0163] Methods known to those skilled in the art can be used to construct an expression vector containing the DNA sequence encoding the fusion protein of the present application and suitable transcription / translation control signals. These methods include in vitro recombinant DNA techniques, DNA synthesis techniques, in vivo recombination techniques, etc. The DNA sequence described can be operably linked to an appropriate promoter in the expression vector to direct mRNA synthesis. Representative examples of such promoters are the lac or trp promoter of E. coli; the PL promoter of lambda phage; eukaryotic promoters including the CMV immediate early promoter, the HSV thymidine kinase promoter, the early and late SV40 promoters, LTRs of retroviruses, and other promoters known to control expression of genes in prokaryotic or eukaryotic cells or viruses. The expression vector also includes a ribosome binding site for translation initiation and a transcription terminator.
[0164] In addition, the expression vector preferably contains one or more selectable marker genes to provide a phenotypic trait for selection of transformed host cells, such as dihydrofolate reductase for eukaryotic cell culture, neomycin resistance in eukaryotic cells, and green fluorescent protein (GFP), or tetracycline or ampicillin resistance for E. coli. The vector containing the appropriate DNA sequence as described above, as well as an appropriate promoter or control sequence, can be used to transform an appropriate host cell to enable it to express the protein.
[0165] The host cell can be a prokaryote (e.g., E. coli), or a lower eukaryote, or a higher eukaryote, such as a yeast cell, a plant cell, or an animal cell (including human and non-human mammals). Representative examples are E. coli, maize embryo cells, insect cells, SF9, Hela, HEK293, CHO, yeast cells, etc. In a preferred embodiment of the present application, a yeast cell (e.g., Pichia, Kluyveromyces, or a combination thereof; preferably, the yeast cell comprises Kluyveromyces, more preferably Kluyveromyces marxianus, and / or Kluyveromyces lactis) is selected as the host cell.
[0166] The polynucleotide of the present application, when expressed in a higher eukaryote, will be enhanced if an enhancer sequence is inserted in the vector. Enhancers are cis-acting elements of DNA, usually about 10 to 300 base pairs in length, which act to increase the transcription of a gene. Examples of enhancers include the SV40 enhancer, which is about 250 base pairs in length and contains two copies of a 14-base pair sequence; the polyoma enhancer, which is about 100 base pairs in length and contains two copies of a 10-base pair sequence; and the adenovirus enhancer, which is about 200 base pairs in length and contains three copies of a 8-base pair sequence.
[0167] The selection of appropriate vectors, promoters, enhancers, and host cells is well within the level of skill in the art.
[0168] The transformation of host cells with recombinant DNA is performed using conventional techniques well known to those of skill in the art. When the host is a prokaryote, e.g., E. coli, the transformation of the host cell can be effected by using CaCl2treatment of the host cell after it has reached the exponential growth phase, using steps well known in the art. Alternatively, MgCl2can be used. If necessary, the transformation can be performed by electroporation. When the host is a eukaryote, the transformation of the host cell can be effected by using a DNA transfection method, such as calcium phosphate co-precipitation, conventional mechanical methods such as microinjection, electroporation, liposome packaging, etc.
[0169] The resulting transformants are cultured in conventional media using conventional techniques, the culture medium being selected depending upon the host cell used. The culture conditions, such as temperature, pH and the like, are those under which the host cell is grown. After the host cell has grown to an appropriate cell density, the selected promoter is induced by the appropriate method (e.g., temperature shift or chemical induction) and the cells are cultured for an additional period.
[0170] The recombinant polypeptides in the above methods can be expressed in the cell, on the cell membrane, or secreted outside the cell. If necessary, the recombinant proteins can be isolated and purified by various separation methods using their physical, chemical and other properties. These methods are well known to those skilled in the art. Examples of these methods include, but are not limited to: conventional renaturation treatment, treatment with a protein precipitant (salting-out method), centrifugation, osmotic lysis, ultra-treatment, ultra-centrifugation, molecular sieve chromatography (gel filtration), adsorption chromatography, ion exchange chromatography, high performance liquid chromatography (HPLC), and other various liquid chromatography techniques, and combinations of these methods.
[0171] Gene therapy
[0172] Gene therapy for genetic diseases refers to the use of genetic engineering techniques to introduce normal genes into patient cells to correct defective genes and cure diseases. The correction can be in situ repair of defective genes or replacement of defective genes with functional normal genes at a certain site in the cell genome to replace defective genes and play a role. Genes are the basic functional units that carry genetic information of organisms, and are a specific sequence located on a chromosome. The introduction of foreign genes into biological cells must rely on certain technical methods or vectors, and the gene transfer method is divided into biological methods, physical methods and chemical methods. Adenovirus vector is one of the most commonly used viral vectors for gene therapy. Gene therapy is mainly used to treat diseases that seriously threaten human health, including, but not limited to: genetic diseases (such as hemophilia, cystic fibrosis, familial hypercholesterolemia, etc.), malignant tumors, cardiovascular diseases, infectious diseases (such as AIDS, rheumatoid arthritis, etc.). Gene therapy is a biomedical high technology that introduces normal genes or genes with therapeutic effects into target cells in the human body through a certain way to correct gene defects or exert therapeutic effects, thereby achieving the purpose of treating diseases. Gene therapy is different from conventional treatment methods: in general, disease treatment targets various symptoms caused by gene abnormalities, while gene therapy targets the root cause of the disease - abnormal genes themselves. Target cells for gene therapy include, but are not limited to, somatic cells, bone marrow cells, liver cells, neural cells, endothelial cells, and muscle cells.
[0173] In the present application, the target gene is efficiently edited (including gene insertion, replacement, etc.) by gene therapy, thereby restoring normal expression of the gene or enhancing expression of the gene, thereby treating related diseases.
[0174] The main advantages of the present application include
[0175] (1) The present application uses the TAM base editor structure of AID located at the N-terminal of nSaCas9 (D10A) KKH to screen AID mutants with higher or lower deaminase activity in eukaryotic cells.
[0176] (2) The present application inserts hAIDx randomly into the acceptor plasmid composed of nSaCas9(D10A)KKH and UGI by transposon system through prokaryotic screening, and obtains new base editors EB-TAM with narrower (such as EB-TAM361-L3) or wider (such as EB-TAM757-L3) editing window through screening.
[0177] (3) The present application verifies the function of the new base editor EB-TAM in eukaryotic cells, and further improves the performance of EB-TAM by optimizing the linker (such as EB-TAM757-L1), and finally obtains EB-TAM cytosine base editors with different editing windows and higher editing activity than hAIDx located at the N-terminal of nSaCas9(D10A)KKH.
[0178] (4) The present application improves the editing efficiency while changing the TAM editing window by constructing EB-TAM. The EB-TAM with larger editing window (such as EB-TAM757-L3 or EB-TAM757-L1) increases its application field in exon skipping, gene expression regulation and other applications that require multiple base sites to be destroyed at the same time to produce effects; and increases the flexibility of its use scenarios.
[0179] The present application will be further described in conjunction with specific examples. It should be understood that these examples are only used to illustrate the present application and not to limit the scope of the present application. The experimental methods in the following examples are not specified in detail, which are usually carried out according to the conventional conditions, such as the conditions described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or according to the conditions recommended by the manufacturer. Unless otherwise specified, percentages and parts are calculated by weight.
[0180] Example 1: Screening of Cytidine Deaminase AID Mutants
[0181] AID, as one of the core elements of TAM, plays a key role in the performance of TAM, which can convert cytosine in ssDNA to uracil nucleoside, and the deamination efficiency of AID directly affects the editing efficiency of TAM. The present application uses human DMD gene as a target through rational design and TAM base editing platform to screen new AID mutants.
[0182] 1.1 Design of different AID mutants and construction of TAM base editors
[0183] The present application designs and synthesizes 20 AID mutants (Mutant AID, mAID) including hAIDx through structural analysis of AID and literature research, as shown in Table 1, and the sequences of hAIDx and mAID mutants are shown in SEQ ID NO. 1-20 in Table 2.
[0184] Table 1
[0185] Table 2
[0186] The schematic diagram of the constructed TAM structure is shown in Figure 1, and the base editing effects of different AID mutants are compared respectively. The sgRNA-1 targeting site is the-14-+6 region of the junction between the 50th exon and the 50th intron of the human DMD gene, and the schematic diagram is shown in Figure 2, and the sgRNA-1 sequence for DMD Exon50 is shown in SEQ ID NO. 21.
[0187] 1.2 Screening of AID mutants
[0188] 1.2.1 Subculture and culture of HEK293T
[0189] Taking the culture of HEK293T cells in a 100mm culture dish as an example, discard the culture medium, slowly add 3mL D-PBS along the wall of the culture dish and shake gently. Discard the D-PBS and add 1mL trypsin containing EDTA, and digest at room temperature for 3 minutes. Terminate the reaction with DMEM complete medium containing 10% FBS, 1% Penicillin-Streptomycin, collect the cells, centrifuge at 1500rpm for 3 minutes, discard the supernatant and resuspend with 4mL complete medium, and count under a microscope. 1x10 6 HEK293T cells are plated in a 100mm culture dish containing 10mL complete medium, cultured in a 37℃ cell incubator for 2 to 3 days, and the next subculture is performed when the cell confluence is 80%-90%. The HEK293T cells used in the following examples are cells with a subculture number of P8-P15.
[0190] 1.2.2 Plasmid transfection of HEK293T cells
[0191] The day before transfection, normally subculture the HEK293T cells once, and plate 1x10 6Cells were seeded at 1.5 x 105cells / 2 mL medium and placed in a CO2incubator. About 24 h later, when the cell density reached about 80%, plasmid transfection was performed using PEI. 125 μL of serum-free medium Opti-MEM was used to dilute 2 μg of the plasmid of interest, and mixed well. 125 μL of serum-free medium Opti-MEM was used to dilute 6 μL of PEI (1 mg / mL), and mixed well. The diluted plasmid and diluted PEI were mixed immediately and mixed well, and incubated at room temperature for 15 minutes. The plasmid-PEI complex was added to the HEK293T cells. After 8 hours of transfection, the culture medium was discarded and 3 mL of DMEM complete medium was added. After 72 hours of transfection, the genomic DNA of the cells was extracted for analysis.
[0192] 1.2.3 Analysis of editing efficiency of TAMs constructed by different AID mutants on DMD gene
[0193] The genomic DNA of the cells was extracted according to the kit instructions of the Blood / Cell / Tissue Genomic DNA Extraction Kit (Tiangen, DP304), and after passing the test, PCR amplification was performed. The system and reaction conditions of PCR are shown in Table 3, and the forward (F1) and reverse (R1) primer sequences for detecting the editing of human DMD Exon50 are shown in SEQ ID NO. 22-23.
[0194] Table 3
[0195] The PCR products were subjected to Sanger sequencing, and the editing efficiency of each TAM was analyzed by EditR1.0.10. The results are shown in FIGS. 3A-3B: at the target position targeting human DMD gene Exon50, mAID-1, mAID-2, mAID-3, mAID-4, mAID-12, mAID-15, mAID-16 and hAIDx showed higher editing efficiency than other AID mutants at C2, C6, C10, C12 and C13 positions; among them, mAID-4 showed higher editing activity than hAIDx at C2, C6 and C13 positions; mAID-3 showed editing efficiency comparable to hAIDx at C2, C6 and C13 positions; and at C13 position, mAID-1, mAID-2, mAID-3, mAID-4, mAID-12 and hAIDx showed comparable editing efficiency. The editing activity of other mAID mutants was not higher than that of hAIDx at each C position. * represents P<0.05, and ** represents P<0.01 (n=2, two-way ANOVA).
[0196] Example 2 Determination of suitable deaminase insertion site of nSaCas9(D10A)KKH in prokaryotic system by transposon
[0197] The cytosine base editor TAM is mainly composed of deaminase AID, nuclease nSaCas9(D10A)KKH and uracil glycosylase inhibitor UGI. The present application screens the effective deaminase insertion site of nSaCas9(D10A)KKH in a prokaryotic system through a transposon system. The effective insertion site refers to the site that can recognize the PAM of the original nSaCas9(D10A)KKH after inserting the deaminase, and can effectively bind the sgRNA to edit the target point. The present application first constructs a skeleton vector pGAT-Cas9-UGI containing only nSaCas9(D10A)KKH and UGI, and then inserts the AID mutant into the skeleton vector through the transposase to obtain a chimeric cytosine base editor (EB-TAM) library, and then introduces it into E. coli together with the screening vector pGAT-Kana-sgRNA with a point mutation inactivated kanamycin resistance. The kanamycin point mutation on pGAT-Kana-sgRNA is inactivated after being changed to D208G, and the inactivated point site can be corrected by the cytosine base editor to restore the kanamycin resistance of E. coli. An sgRNA-2 against the kanamycin point mutation D208G is also cloned on the plasmid, and the sequence is shown as SEQ ID NO. 24. The clones that successfully grow on the kanamycin plate after induction editing show that the introduced EB-TAM is active, and the random insertion site of the AID mutant is determined by Sanger sequencing. The structures of pGAT-Cas9-UGI and pGAT-Kana-sgRNA are shown in Figure 4, wherein Figure 4(A) shows the structure diagram of pGAT-Cas9-UGI, Figure 4(B) shows the structure diagram of pGAT-Kana-sgRNA, and the prokaryotic transposition system screening process is shown in Figure 5.
[0198] 2.1 Construction of AID mutant and nSaCas9(D10A)KKH chimeric protein plasmid library by transposase
[0199] The present application constructs four different linker AID transposon donor plasmids, wherein L1 linker, L2 linker and L3 linker carry hAIDx, and L4 linker carries mutant mAID-13. The sequences of L1-L4 linkers are shown as SEQ ID NO. 25-28. The transposon fragment is recovered by electrophoresis after enzyme digestion, and the qualified product is subjected to transposition reaction. The transposition reaction system is shown in Table 4:
[0200] Table 4
[0201] After transposition, purification was performed according to the instructions of the MicroElute DNA Clean-Up Kit (Omega bio-tek, D6296) kit, and after inspection, the library was constructed by electroporation. The conditions for constructing the library by electroporation were as follows: 1.0 μL Clean sample and 25.0 μL E. cloni 10G ELITE SixPacks (Lucigen, 60052) were mixed, and then the library was constructed by electroporation according to the parameters C = 25 μF, PC = 200 Ω, V = 1.6 kV. After amplification, the plasmid was extracted according to the instructions of the EndoFree Midi Plasmid Kit (TIANGEN, DP118), and after inspection, the next step was performed.
[0202] 2.2 Screening of functional chimeric proteins by prokaryotic resistance
[0203] The transposition library was screened for functional chimeric proteins by prokaryotic resistance. Specifically, the plasmid pGAT-Kana-sgRNA with streptomycin and inactivated kanamycin was co-transferred into BL21 (DE3) competent cells with the pGAT-Cas9-UGI plasmid with ampicillin, and then plated on LB plates with double-antibiotic streptomycin and ampicillin for overnight culture. The next day, all colonies on the double-antibiotic plate were recovered and expanded in streptomycin and ampicillin double-antibiotic LB medium to an OD of about 0.6, and then the expression of pGAT-Cas9-UGI protein was induced by IPTG. After 16 hours of induction, the bacterial solution was plated on a kanamycin-resistant LB plate and cultured overnight to form clones, which were sequenced by Sanger to determine the sequence of the chimeric protein.
[0204] 2.2.1 Transformation of BL21 competent cells
[0205] 100 μL of BL21 (DE3) competent cells (Sangon biotech, B528414) were taken out from the -80°C refrigerator and quickly inserted into an ice box to dissolve them. 1.0 μL of EB-TAM library and 1.0 μL of pGAT-Kana-sgRNA plasmid were added and mixed gently, and placed in ice for 30 minutes. Heat shock at 42°C for 45 seconds, quickly put back into ice for 2 minutes, add 900 μL of sterile medium without antibiotics, mix well. Incubate at 225 rpm, 37°C for 1 hour.
[0206] 2.2.2 Kanamycin resistance screening
[0207] After transformation in 2.2.1, 4.0 mL of LB medium containing ampicillin resistance and streptomycin resistance was added to the bacterial solution, and the culture was continued for 6-8 h. When the OD was measured at about 0.6, 0.1 mM IPTG was added, and the induction was carried out at 16°C, 225 rpm, overnight. After the induction was completed, the bacterial solution diluted at different gradients was plated on LB agar plates containing kanamycin (50 μg / mL), and after overnight culture, the grown clones were subjected to Sanger sequencing, and the insertion site of the mAID protein on nSaCas9(D10A)KKH was analyzed.
[0208] Results: The present application identified 11 EB-TAMs, EB-TAM361-L3, EB-TAM408-L1, EB-TAM455-L3, EB-TAM469-L4, EB-TAM471-L4, EB-TAM588-L3, EB-TAM634-L1, EB-TAM719-L4, EB-TAM757-L3, EB-TAM928-L2 and EB-TAM941-L2, through E. coli monoclonal sequencing. These EB-TAMs can restore the kanamycin resistance of the screening plasmid through base substitution, indicating that the 361, 408, 455, 469, 471, 588, 634, 719, 757, 928 and 941 amino acid sites of nSaCas9(D10A)KKH (SEQ ID NO. 41) can be inserted with other proteins without affecting the PAM recognition and sgRNA binding ability of nSaCas9(D10A)KKH. SaCas9 is a double-leaf structure composed of a REC recognition leaf and a discontinuous NUC nuclease leaf, and the sgRNA (crRNA) is located in the channel between the two leaves targeted by DNA, and the discontinuous RuvC-like domain in NUC performs cleavage on the non-targeted strand DNA but is inactivated in nSaCas9(D10A)KKH, while the HNH structure is responsible for the cleavage of the targeted strand DNA. The insertion site of each EB-TAM screened by the present application is in the domain as shown in Table 5 and FIG. 6, and the domains other than RuvC-1, Arg and L-1 can tolerate the insertion of other proteins.
[0209] Table 5 Note: aa represents amino acid, "upstream 3 aa sequence" refers to the 3 amino acid sequence on the left or N-terminal of the insertion site, and similarly, "downstream 3 aa sequence" refers to the 3 amino acid sequence on the right or C-terminal of the insertion site.
[0210] Compared with SaCas9, SpCas9, FnCas9 and AnaCas9 four Cas9 homologous proteins as shown in Figure 7, it is found that they contain similar domains and have the same overall structure. The present application successfully screened 11 suitable deaminase insertion sites on nSaCas9(D10A)KKH in the prokaryotic system through the transposon system. It can be inferred that the insertion sites screened in nSaCas9(D10A)KKH in the present application can also be inserted into other proteins at the corresponding sites without affecting their activity.
[0211] Example 3 Screening of EB-TAM with different editing windows in eukaryotic system
[0212] Since EB-TAM is mainly used in eukaryotes, it is necessary to verify whether the EB-TAM screened in Example 2 has editing activity in eukaryotes and to screen EB-TAM with different editing ranges. The present application edits the endogenous target sites Site1-5 rich in C in HEK293T, determines the editing window of each EB-TAM by the position of cytosine in sgRNA, and the sequences of sgRNA-3-sgRNA-7 are shown in SEQ ID NO. 29-33.
[0213] 3.1 Plasmid construction: According to the insertion sites screened in prokaryotes, eukaryotic codon-optimized EB-TAM expression vectors are constructed, wherein the sequences of eTAM, EB-TAM361-L3, EB-TAM719-L4 and EB-TAM757-L3 are shown in SEQ ID NO. 34-37, respectively, and the sequences of EB-TAM408-L1, EB-TAM588-L3, EB-TAM634-L1 and EB-TAM455-L3 are shown in SEQ ID NO. 44-47, respectively.
[0214] 3.2 HEK293T passaging and culture, plasmid transfection and editing efficiency analysis, as described in Example 1, the results are shown in Figure 8, wherein Figure 8(A)-8(B) are the editing efficiency and editing window of each site of Site1; Figure 8(C)-8(D) are the editing efficiency and editing window of each site of Site2; Figure 8(E)-8(F) are the editing efficiency and editing window of each site of Site3. Figure 8(G)-8(H) are the editing efficiency and editing window of each site of Site4. Figure 8(I)-8(J) are the editing efficiency and editing window of each site of Site5. EB-TAM757-L3 has comparable editing efficiency with eTAM at Site1, Site2 and Site5, slightly lower than eTAM at Site4, but significantly higher than eTAM at Site3. And EB-TAM757-L3 shows higher editing efficiency than eTAM after C14 of all Sites. EB-TAM361-L3 shows smaller editing window than eTAM at all Sites, although the editing efficiency is reduced. * indicates P<0.05, ** indicates P<0.01 (n=2, two-way ANOVA). The editing window pattern of different EB-TAMs is shown in Figure 9.
[0215] The results shown in Figure 8 and Figure 9 show that EB-TAM361-L3, EB-TAM408-L1, EB-TAM588-L3, EB-TAM634-L1, EB-TAM719-L4 and EB-TAM757-L3 can all play the role of cytosine base editing in eukaryotic cells, indicating that the insertion of proteins at these sites can construct new tool proteins while maintaining the biological activity of nSaCas9(D10A)KKH. Among them, EB-TAM757-L3 shows higher editing activity than eTAM at 14th-17th position of the editing window at 5 different sites, indicating that the editing window is increased. And EB-TAM361-L3 has the farthest 14th position of the editing window at 5 sites and shows a narrower editing window than eTAM at all sites, with only C8 site edited at Site1.
[0216] This embodiment verified 6 EB-TAMs (EB-TAM361-L3, EB-TAM408-L1, EB-TAM588-L3, EB-TAM634-L1, EB-TAM719-L4 and EB-TAM757-L3) with different editing windows in eukaryotic cells, as the nSaCas9(D10A)KKH cuts the 18th position of the target strand of sgRNA, theoretically it is difficult to effectively edit the bases at positions 19 and 20 of the non-target strand with base editors, and EB-TAM757-L3 has an editing window of 1-17, which is currently the largest window among the reported cytosine base editors with nSaCas9(D10A)KKH as the core.
[0217] Example 4: Screening and optimizing EB-TAM in eukaryotic cells
[0218] Duchenne Muscular Dystrophy (DMD) is a fatal degenerative neuromuscular rare disease, and its pathogenesis is that the gene of dystrophin protein (DMD) has abnormal mutations, resulting in the loss of dystrophin protein. DMD has many mutation types, including small deletions, insertions, point mutations, and repeat sequences, or deletion or sequence duplication of one or more exons in a large fragment, resulting in early clinical symptoms and long-term lack of effective treatment. According to statistics, about 80% of DMD patients can be treated by exon skipping therapy, and skipping exon 53 (Exon 53) can treat about 8% of patients.
[0219] This embodiment takes DMD Exon 53 splice site as a target to verify the editing efficiency of EB-TAM again. Then the editing efficiency of EB-TAM757 is further improved by optimizing the linker connecting the deaminase and nSaCas9(D10A)KKH. Finally, EB-TAM757-L1 is confirmed to effectively induce DMD Exon 53 skipping in iPSC-induced skeletal muscle. The sgRNA design for Exon 53 is in the-13 to +7 region of the junction between intron 52 and exon 53 of the DMD gene, as shown in Figure 10, and the sgRNA-8 sequence for Exon 53 is shown in SEQ ID NO. 38.
[0220] 4.1 Editing efficiency analysis of each EB-TAM by DMD Exon 53 target
[0221] HEK293T cells were passaged and cultured, plasmid transfection and editing efficiency analysis were performed as described in Example 1, the sequences of the amplification primers (F2, R2) are shown in SEQ ID NO. 39-40, and the results are shown in Figure 11: eTAM, EB-TAM455-L3, EB-TAM634-L1, EB-TAM757-L3 all have editing ability at the DMD Exon 53 target site, among which EB-TAM455-L3, EB-TAM634-L1, EB-TAM757-L3 show higher editing effect than eTAM at C1 site, EB-TAM634-L1 shows editing effect comparable to eTAM at C5 site, and EB-TAM757-L3 shows higher editing effect at C8 site.
[0222] 4.2 Improve the editing efficiency of EB-TAM by optimizing the linker
[0223] Based on the data from Examples 3 and 4.1, this invention further optimizes EB-TAM757-L3, which has a wider editing range and higher editing efficiency. It has been reported that when the deaminase is at the N-terminus, optimizing the linker connecting the deaminase and nSaCas9(D10A)KKH, altering its length and rigidity, can affect its editing efficiency at the same target site (Komor et al., 2016). To further improve the editing activity of EB-TAM, this invention replaces the L3 linker sequence connecting to the deaminase at amino acid position 757 of nSaCas9(D10A)KKH. The L3 linker of EB-TAM757-L3 is a (GGGGS)3 sequence, which is a flexible structure and the most common linker in fusion proteins. The L2 linker is an 16-amino acid XTEN, which is used in first-generation cytosine base editors using APOBEC1 as the deaminase due to its relatively balanced adjustment of editing efficiency and editing window. This invention adds an (SGGS)2 sequence to the outside of the L2 linker to form the L1 linker. The L1 and L2 linker sequences are shown in SEQ ID NO. 25-26, and the constructed EB-TAM757-L1 and EB-TAM757-L2 sequences are shown in SEQ ID NO. 48-49, respectively. This embodiment further verified the editing efficiency of EB-TAM757 with different linkers in DMD Exon 53. The editing efficiency of EB-TAM757-L3 at the Exon 53 target site after replacing different linkers (L1 or L2) is shown in Figure 12: at C1, C5, and C8 positions, the editing efficiency of EB-TAM757-L1 is higher than that of EB-TAM757-L2 and EB-TAM757-L3, and the editing efficiency of EB-TAM757-L2 is higher than that of EB-TAM757-L3. Replacing the L3 linker of EB-TAM757 with the L1 linker significantly improved its editing efficiency on the same target. EB-TAM757-L1 offers a 1.75 to 3.75-fold improvement in editing efficiency compared to eTAM, with an efficiency of 15% to 20% for the DMD Exon 53 target.
[0224] 4.3 Utilizing EB-TAM757-L1 to efficiently induce Exon 53 jumping in skeletal muscle cells with DMD
[0225] To more closely approximate the results of in vivo experiments, this embodiment uses iPSC-induced skeletal muscle to verify the effect of EB-TAM757-L1 and sgRNA-8 in inducing DMD Exon 53 jumping.
[0226] 4.3.1 iPS cell culture and passage
[0227] Take the 6-well plate passage as an example. Matrigel coated 6-well plates (coating concentration is 0.013 mg / cm2) are placed in the biosafety cabinet for about 1 hour in advance to recover to room temperature; when the confluence of the clone group is 85%, the iPSC well culture medium is aspirated, 2 mL / well of DPBS is added, and it is gently shaken and aspirated. Add 2 mL / well of Nuwacell EDTA passage working solution to completely cover the bottom of the well, and incubate in a 37°C incubator for 8 min. After digestion, gently take the cell culture plate back to the biosafety cabinet to avoid shaking the cells, and tilt and aspirate. Add 2 mL / well of pre-warmed Blebbistatin+ncTarget complete medium in time, and horizontally cross-shake the 6-well plate to make the cells detach from the matrix, and passage at a ratio of 1:12. 0.166 mL of cells are inoculated into a new Matrigel coated 6-well plate, and 2 mL of pre-warmed Blebbistatin+ncTarget complete medium is added to the well.
[0228] 4.3.2 Skeletal muscle cell induction
[0229] The method of skeletal muscle cell induction of iPSC is carried out according to the Skeletal Muscle Differentiation Kit instruction of Genea Biocells company.
[0230] 4.3.3 Skeletal muscle cell plasmid transfection
[0231] Transfection is carried out within the first 24 hours of myoblast induction to differentiate into myotubes. 2 μg of the plasmid of interest and 4 μL P3000 are diluted with 125 μL of serum-free medium Opti-MEM, 6 μL Lipo3000 is diluted with 125 μL of serum-free medium Opti-MEM, and they are mixed well. Immediately mix the Opti-MEM diluted plasmid and Lipo3000 and mix well, and incubate at room temperature for 15 min. Add the plasmid-liposome complex to the cells. After transfection for 24 h, discard the culture medium and add 2 mL of SKM-03 induction medium; after transfection for 7 days, collect the cells for analysis.
[0232] 4.3.4 Analysis of DMD Exon53 splicing
[0233] Total RNA of the cells was extracted according to the kit instruction of E.Z.N.A. HP Total RNA Kit (Omega bio-tek, R6812), and the qualified samples were used for the next step. Reverse transcription PCR was performed according to the kit instruction of PrimeScript RT reagent Kit with gDNA Eraser (TAKARA, RR047A), and the reaction system of reverse transcription of Total RNA is shown in Table 6:
[0234] Table 6
[0235] The related transcription products were amplified according to the kit instruction of TakaRa Ex PremierTM DNA Polymerase Dye plus (TAKARA, RR371A), and the primer sequences of the forward primer (F2) and the reverse primer (R2) for DMD Exon53 skipping / editing detection were as shown in SEQ ID NO. 39-40, and the reaction system is shown in Table 3. The DMD gene splicing of the skeletal muscle cells edited by TAM757-L1 and sgRNA-8 was analyzed, and the electrophoresis result is shown in Figure 13: two bands of band 1 and band 2 were shown, and the sequencing result of the electrophoresis band is shown in Figure 14: band 2 sequencing found that the entire Exon 53 was missing, indicating that EB-TAM757-L1 and sgRNA-8 can effectively induce DMD Exon53 skipping in skeletal muscle cells, and the skipping efficiency is about 50%.
[0236] The above results show that: by optimizing the linker of EB-TAM757-L3 to EB-TAM757-L1, the editing efficiency of EB-TAM757-L1 is successfully improved. And it is confirmed in skeletal muscle cells that EB-TAM757-L1 and sgRNA-8 can effectively induce DMD Exon53 skipping.
[0237] All the documents mentioned in the present application are incorporated herein by reference as if each document were individually incorporated by reference. In addition, it should be understood that various changes and modifications can be made to the present application by those skilled in the art upon reading the above description of the present application, and such equivalent forms are also within the scope of the present application as defined in the appended claims.
Claims
1. A fusion protein, characterized in that, The fusion protein includes a nuclease domain and an editing enzyme domain; the editing enzyme domain is inserted into an insertion site in the nuclease domain, wherein the nuclease domain in the fusion protein has nuclease function; and the editing enzyme domain in the fusion protein has editing enzyme function. The nuclease domain is derived from the Cas protein or a Cas protein mutant.
2. The fusion protein as described in claim 1, characterized in that, The fusion protein described has the structure of Formula I: W0-Z1-L1-Y-L2-Z2-W1-W2 (I) In the formula, W0 is a non-core or core-positioning element; Z1 is the left-hand element of the nuclease domain; Z2 is the right-hand element of the nuclease domain; L1 is either absent or linked to a peptide element; L2 is either absent or linked to a peptide element; Y represents the editing enzyme structural domain element; W1 is an element with or without uracil glycosylation inhibitor (UGI); W2 is a core positioning element; Each "-" represents an independent chemical bond; Z1 and Z2 are the left and right elements formed by the insertion site splitting an editing enzyme domain, respectively.
3. The fusion protein as described in claim 1, characterized in that, The Cas protein or Cas protein mutant is derived from nSaCas9(D10A)KKH.
4. The fusion protein as described in claim 1, characterized in that, The insertion site of the editing enzyme domain into the nuclease domain is selected from any one of the following groups: positions 361, 408, 455, 469, 471, 588, 634, 719, 757, 928, and 941 corresponding to the amino acid sequence shown in SEQ ID NO:
41.
5. The fusion protein as described in claim 1, characterized in that, The editing enzyme is selected from at least one or more of deaminases, DNA glycosylation enzymes, methyltransferases, and acetyltransferases.
6. A base editing system, characterized in that, It contains the fusion protein as described in claim 1 or its encoded polynucleotide.
7. The base editing system as described in claim 6, characterized in that, The base editing system also includes guide RNA.
8. The base editing system as described in claim 7, characterized in that, The sequence of the guide RNA is selected from any one of the nucleotide sequences shown in SEQ ID NO:21, SEQ ID NO:24, SEQ ID NO:29-33 or SEQ ID NO:
38.
9. A polynucleotide, characterized in that, The polynucleotide sequence encodes the fusion protein of claim 1 or the base editing system of claim 6.
10. A fusion protein as described in claim 1, a base editing system as described in claim 6, The use of the polynucleotide as described in claim 9 in the preparation of a drug for gene editing.
Citation Information
Patent Citations
Base editing tool and use thereof
CN111172133A
Exon splicing enhancer related to Duchenne muscular dystrophy, sgRNA, gene editing tool and application
CN112063621A
Induction method of gene point mutation
CN112251464A
Fusion protein, polynucleotide thereof, base editor and application thereof in medicine preparation
CN113717961A
Editing system and method for efficiently and specifically realizing base transversion and application
CN114835821A