A novel gene editing system mediating a from a to c mutation or a t to g mutation and applications thereof
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-30
- Publication Date
- 2026-08-11
AI Technical Summary
[0007]然而,现有技术中没有已报道的酶可直接将基因组DNA中的腺嘌呤(A)催化为胞嘧啶(C),反链即为胸腺嘧啶(T)到鸟嘌呤(G),而需要A到C或者T到G来纠正的人类致病点突变(SNV)占比16%,C到A或者G到T病原性点突变同时也是第二大最常见的致病性SNV,同时此类突变也超出了经典CBE可以覆盖的疾病范围[1]
[0112] (1) By fusing an artificially evolved mouse-derived 3-methyladenine glycosylase (Aag) variant with a monomeric adenosine deaminase Tad-8e or its variant and Cas9n with impaired catalytic activity, a single-base gene editing system containing an ACBE1.0/2.0/2Q fusion protein was constructed, enabling base transversions from A to C (sense strand) or from T to G (antisense strand), with the highest transversion efficiency of A·T to C·G reaching 45.2%.
Smart Images

Figure CN116200382B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biotechnology, specifically relating to a novel gene editing system that mediates A-to-C mutations or T-to-G mutations and its applications. Background Technology
[0002] Human genetic diseases are essentially caused by gene mutations. About 60% of genetic diseases are caused by single base mutations. Traditional methods of correcting these genetic diseases using homologous recombination mediated by genome editing are very inefficient (0.1%-5%). [1,2] Single-base editors, derived from the CRISPR system, are a newly emerging and highly efficient base editing technology in recent years. Their advantages, such as not causing DNA double-strand breaks, not requiring recombination templates, and high-efficiency editing, have shown great promise for applications in basic research and clinical disease treatment.
[0003] Classic base editors are mainly divided into cytosine base editors and adenine base editors, specifically:
[0004] CBE consists of impaired activity of *Streptococcus pyogenes* spCas9n, rat-derived cytosine deaminase rAPOBEC1, and a uracil glycosidase inhibitor. The Cas9 protein recognizes NGG as a PAM and specifically binds to DNA. Subsequently, under the action of deaminase and DNA repair, C·GT·A substitution is achieved within a 20bp range upstream of the NGG (positions 21-23), with the editing window primarily located at positions 4-8. [2] It is expected to correct 14% of pathogenic point mutations in humans;
[0005] ABE, on the other hand, fuses bacterial-derived TadA with spCas9. With the aid of directed evolution and protein engineering, it underwent seven rounds of evolution to finally obtain ABE7.10, an adenine base editor capable of operating on single-stranded DNA. The active editing region is mainly located at positions 4-7. This system induces an average A·TG·C editing efficiency of approximately 53% in human cells, far exceeding the efficiency of homologous recombination-mediated base mutations. Its product purity reaches 99.9%, with extremely low indel (insertion and deletion) occurrence. [3] More importantly, about 47% of pathogenic point mutations in humans are caused by the C·G mutation to T·A, and the adenine base editor holds the promise of correcting nearly half of these pathogenic point mutations. [1] This demonstrates its great potential in modifying mutant bases and treating genetic diseases. Currently, ABE is widely used in the preparation of animal models. [4-8] and gene therapy [9-14].
[0006] Both CBE and ABE can only achieve base switching. In the early stages of CBE development, it was found that knocking out intracellular uracil glycosidase (UNG) or removing cytosine glycosidase inhibitors (UGI) produces C·G-to-G·C and C·G-to-A·T editing byproducts, i.e., C-based transversion occurs. [3,15] Recently, based on the editing byproducts observed in previous CBE experiments, the CGBE series has been developed by fusing different types of UNG, DNA damage repair proteins, or trans-damage polymerases with CBE after removing UGI. [16-19] It holds promise for treating pathogenic point mutations in 11% of G·C to C·G mutations.
[0007] However, no reported enzymes in the current technology can directly catalyze the conversion of adenine (A) to cytosine (C) in genomic DNA, i.e., the reverse strand conversion from thymine (T) to guanine (G). Pathogenic point mutations (SNVs) requiring A-to-C or T-to-G correction account for 16% of human pathogenic SNVs, and C-to-A or G-to-T pathogenic point mutations are also the second most common type of pathogenic SNV. Furthermore, these mutations exceed the disease range covered by classic CBE (Cellular Pathogenic Biology). [1] .
[0008] [References]
[0009] [1]Rees HA, Liu DR. Base editing: Precision chemistry on the genome and transcriptome of living cells. Nat Rev Genet, 2018, 19: 770-788
[0010] [2]Komor AC, Kim YB, Packer MS, et al. Programmable editing of a targetbase in genomic DNA without double-stranded DNA cleavage. Nature, 2016, 533: 420-424
[0011] [3]Gaudelli NM,Komor AC,Rees HA,et al.Programmable base editing of a*t to g*c in genomic DNA without DNA cleavage.Nature,2017,551:464-471
[0012] [4]Liang P,Sun H,Zhang X,et al.Effective and precise adenine baseediting in mouse zygotes.Protein Cell,2018,9:808-813
[0013] [5]Ma Y,Yu L,Zhang X,et al.Highly efficient and precise base editingby engineered dcas9-guide trna adenosine deaminase in rats.Cell Discov,2018,4:39
[0014] [6]Liu Z,Lu Z,Yang G,et al.Efficient generation of mouse models ofhuman diseases via abe-and be-mediated base editing.Nat Commun,2018,9:2338
[0015] [7]Ryu SM,Koo T,Kim K,et al.Adenine base editing in mouse embryos andan adult mouse model of duchenne muscular dystrophy.Nat Biotechnol,2018,36:536-539
[0016] [8]Yang L,Zhang X,Wang L,et al.Increasing targeting scope ofadenosine base editors in mouse and rat embryos through fusion of tadadeaminase with cas9 variants.Protein Cell,2018,9:814-819
[0017] [9]Suh S,Choi EH,Leinonen H,et al.Restoration of visual function inadult mice with an inherited retinal disease via adenine base editing.NatBiomed Eng,2021,5:169-178
[0018]
[10] Rothgangl T,Dennis MK,Lin PJC,et al.In vivo adenine base editingof pcsk9 in macaques reduces ldl cholesterol levels.Nat Biotechnol,2021,
[0019]
[11] Kim Y,Hong SA,Yu J,et al.Adenine base editing and prime editingof chemically derived hepatic progenitors rescue genetic liver disease.CellStem Cell,2021,
[0020]
[12] Newby GA,Yen JS,Woodard KJ,et al.Base editing of haematopoieticstem cells rescues sickle cell disease in mice.Nature,2021,595:295-302
[0021]
[13] Koblan LW,Erdos MR,Wilson C,et al.In vivo base editing rescueshutchinson-gilford progeria syndrome in mice.Nature,2021,589:608-614
[0022]
[14] Musunuru K,Chadwick AC,Mizoguchi T,et al.In vivo crispr baseediting of pcsk9 durably lowers cholesterol in primates.Nature,2021,593:429-434
[0023]
[15] Komor AC,Zhao KT,Packer MS,et al.Improved base excision repairinhibition and bacteriophage mu gam protein yields c:G-to-t:A base editorswith higher efficiency and product purity.Sci Adv,2017,3:eaao4774
[0024]
[16] Koblan LW,Arbab M,Shen MW,et al.Efficient c*g-to-g*c base editorsdeveloped using crispri screens,target-library analysis,and machinelearning.Nat Biotechnol,2021,
[0025]
[17] Chen L,Park JE,Paa P,et al.Programmable c:G to g:C genome editingwith crispr-cas9-directed base excision repair proteins.Nat Commun,2021,12:1384
[0026]
[18] Zhao D,Li J,Li S,et al.Glycosylase base editors enable c-to-a andc-to-g base changes.Nat Biotechnol,2021,39:35-40
[0027]
[19] Kurt IC, Zhou R, Iyer S, et al. Crispr c-to-g base editors for inducing targeted DNA transversions in human cells. Nat Biotechnol, 2021, 39: 41-46. Summary of the Invention
[0028] This invention first relates to a single-base gene editing system, which can induce A-to-C or T-to-G directional mutations in target DNA. Specifically, it can induce A-to-C directional mutations in the sense strand of the target gene or T-to-G directional mutations in the antisense strand of the target gene. Preferably, the target DNA is genomic DNA.
[0029] The single-base gene editing system comprises gRNA, and the main active region of the single-base gene editing system is the 5th to 7th bases at the 5' end of the gRNA;
[0030] The gene editing system further comprises a fusion protein with single-base editing function, wherein the fusion protein with single-base editing function comprises the following functional blocks:
[0031] (1) One or more nuclear localization signaling peptides (bNLS);
[0032] (2) A functional block with single-base editing capability, which consists of the following structure:
[0033] 1) TadA-8e polypeptide,
[0034] 2) A complete Cas9 / Cas9n polypeptide or multiple secondary polypeptide fragments derived from the same Cas9 / Cas9n polypeptide, and
[0035] 3) A variant of HDG4 polypeptide (3-methyladenine glycosidase, Aag);
[0036] And optional 4), linker, used to link individual peptides, peptide fragments or peptide variants;
[0037] The sequence of the TadA-8e polypeptide is shown in SEQ ID NO.2;
[0038] The variant of the HDG4 polypeptide is a functional polypeptide based on the wild-type mouse 3-methyladenine glycosidase protein (Aag) sequence shown in SEQ ID NO.5, further containing any of the following point mutations or combinations of point mutations: E145A / F, Y147A / F / W, D152F / W / H / F / W, R165E / A / F, Y177A / F / W, V178F / E, Y179A / W / F, Y182W / F / A, G183Q / F / H, M184F / W, Y185W / A / F.
[0039] Preferably, the nuclear localization signal peptide is located at the N-terminus and / or C-terminus of the fusion protein with single-base editing function;
[0040] Most preferably, the fusion protein with single-base editing function has a nuclear localization signal peptide at both its N-terminus and C-terminus, and the sequence of the nuclear localization signal peptide is shown in SEQ ID NO.1;
[0041] Preferably, the polypeptide sequence of the linking fragment is as shown in any one of SEQ ID NO.3, 10-21, and the sequences of the linking fragments linking each block may be the same or different;
[0042] On the other hand, the fusion protein with single-base editing function described in this invention has an ACBE1.0 structure, wherein...
[0043] (2) Functional blocks with single-base editing capabilities are composed of the following structures:
[0044] 1) The TadA-8e polypeptide with the sequence shown in SEQ ID NO.2;
[0045] 2) A complete Cas9 polypeptide;
[0046] 3) Variants of the HDG4 peptide;
[0047] And 4), optionally 0-2 linkers for linking the polypeptide or polypeptide variant;
[0048] Preferably, the complete Cas9 polypeptide is spCas9n polypeptide (D10A), which is the 2-1368 polypeptide sequence of the Saccharomyces cerevisiae spCas9n (D10A) polypeptide sequence shown in SEQ ID NO.4;
[0049] The variant of the HDG4 polypeptide is a functional polypeptide based on the wild-type mouse Aag sequence shown in SEQ ID NO.5, further containing any of the following point mutations or combinations of point mutations:
[0050] E145A / F, Y147A / F / W, D152F / W / H / F / W, R165E / A / F, Y177A / F / W, V178F / E, Y179A / W / F, Y182W / F / A, G183Q / F / H, M184F / W, Y185W / A / F.
[0051] Preferably, the HDG4 variant is a functional polypeptide containing any two of the following point mutation combinations based on the wild-type mouse Aag polypeptide sequence shown in SEQ ID NO.5: E145A / F, Y147A / F / W, D152F / W / H / F / W, R165E / A / F, Y177A / F / W, V178F / E, Y179A / W / F, Y182W / F / A, G183Q / F / H, M184F / W, Y185W / A / F;
[0052] More preferably, the HDG4 variant is a functional polypeptide containing any one of the following mutation sites, based on the wild-type mouse Aag polypeptide sequence shown in SEQ ID NO. 5: (1) E145A, (2) E145F, (3) Y147A, (4) Y147F, (5) Y147W, (6) D152F, (7) D152W, (8) D152H, (9) A154F, (10) A154W, (11) R165E, (12) R165A, ( 13)R165F, (14)Y177A, (15)Y177F, (16)Y177W, (17)V178F, (18)V178E, (19)Y179A, (20)Y179W , (21)Y179F, (22)Y182W, (23)Y182F, (24)Y182A, (25)G183Q, (26)G183F, (27)G183H, (28)M184 F. (29)M184W, (30)Y185W, (31)Y185A, (32)Y185F, (33)R165E&M184F, (34)R165E&Y179F, (35) R165E&Y185F, (36)R165E&G183H, (37)R165E&G183Q, (38)R165E&M184W, (39)Y179F&M184F, (40 )Y179F&Y185F, (41)Y179F&M184W, (42)Y179F&G183H, (43)Y179F&G183Q, (44)M184F&Y185F, ( 45)M184W&Y185F, (46)G183H&Y185F, (47)G183Q&Y185F, (48)G183H&M184W, (49)G183Q&M184W;
[0053] Most preferably, the variant of HDG4 is a polypeptide containing the R165E&Y179F double mutation, based on the wild-type mouse Aag polypeptide sequence shown in SEQ ID NO.5;
[0054] Most preferably, the polypeptide sequences of the linked fragments are as shown in SEQ ID NO.3.
[0055] On the other hand, the fusion protein with single-base editing function described in this invention has an ACBE2.0 structure, wherein,
[0056] (2) Functional blocks with single-base editing capabilities are composed of the following structures:
[0057] 1) The first polypeptide fragment derived from spCas9n polypeptide (D10A);
[0058] 2) Variants of the HDG4 peptide;
[0059] 3) The TadA-8e polypeptide with the sequence shown in SEQ ID NO.2;
[0060] 4) A second polypeptide fragment derived from spCas9n polypeptide (D10A);
[0061] And 5), optionally 0-3 linkers for linking the polypeptide, polypeptide fragment or polypeptide variant;
[0062] The first polypeptide fragment derived from spCas9n polypeptide (D10A) and the second polypeptide fragment derived from spCas9n polypeptide (D10A) are from different regions in the Saccharomyces cerevisiae spCas9n (D10A) polypeptide sequence shown in SEQ ID NO.4;
[0063] Preferably, the first polypeptide fragment derived from spCas9n polypeptide (D10A) is polypeptide 2-1248 in the *Saccharomyces cerevisiae* spCas9n (D10A) polypeptide sequence shown in SEQ ID NO.4; the second polypeptide fragment derived from spCas9n polypeptide (D10A) is polypeptide 1249-1368 in the *Saccharomyces cerevisiae* spCas9n (D10A) polypeptide sequence shown in SEQ ID NO.4;
[0064] The variant of the HDG4 polypeptide is a functional polypeptide based on the wild-type mouse Aag sequence shown in SEQ ID NO.5, further containing any of the following point mutations or combinations of point mutations:
[0065] E145A / F, Y147A / F / W, D152F / W / H / F / W, R165E / A / F, Y177A / F / W, V178F / E, Y179A / W / F, Y182W / F / A, G183Q / F / H, M184F / W, Y185W / A / F.
[0066] Preferably, the HDG4 variant is a functional polypeptide containing any two of the following point mutation combinations based on the wild-type mouse Aag polypeptide sequence shown in SEQ ID NO.5: E145A / F, Y147A / F / W, D152F / W / H / F / W, R165E / A / F, Y177A / F / W, V178F / E, Y179A / W / F, Y182W / F / A, G183Q / F / H, M184F / W, Y185W / A / F;
[0067] More preferably, the HDG4 variant is a functional polypeptide containing any one of the following mutation sites, based on the wild-type mouse Aag polypeptide sequence shown in SEQ ID NO. 5: (1) E145A, (2) E145F, (3) Y147A, (4) Y147F, (5) Y147W, (6) D152F, (7) D152W, (8) D152H, (9) A154F, (10) A154W, (11) R165E, (12) R165A, ( 13)R165F, (14)Y177A, (15)Y177F, (16)Y177W, (17)V178F, (18)V178E, (19)Y179A, (20)Y179W , (21)Y179F, (22)Y182W, (23)Y182F, (24)Y182A, (25)G183Q, (26)G183F, (27)G183H, (28)M184 F. (29)M184W, (30)Y185W, (31)Y185A, (32)Y185F, (33)R165E&M184F, (34)R165E&Y179F, (35) R165E&Y185F, (36)R165E&G183H, (37)R165E&G183Q, (38)R165E&M184W, (39)Y179F&M184F, (40 )Y179F&Y185F, (41)Y179F&M184W, (42)Y179F&G183H, (43)Y179F&G183Q, (44)M184F&Y185F, ( 45)M184W&Y185F, (46)G183H&Y185F, (47)G183Q&Y185F, (48)G183H&M184W, (49)G183Q&M184W;
[0068] Most preferably, the variant of HDG4 is a polypeptide containing the R165E&Y179F double mutation, based on the wild-type mouse Aag polypeptide sequence shown in SEQ ID NO.5;
[0069] Most preferably, the polypeptide sequences of the linked fragments are as shown in SEQ ID NO.3.
[0070] Furthermore, the fusion protein with single-base editing function described in this invention has an ACBE2Q structure, wherein,
[0071] (2) The structure of the functional block with single base editing function is as follows: based on the ACBE2.0 structure of the fusion protein, the TadA-8e polypeptide is replaced with the TadA-8e polypeptide (N108Q) variant, and its amino acid sequence is shown in SEQ ID NO.9.
[0072] Furthermore, the ACBE1.0 fusion protein of the present invention is composed of the following blocks arranged sequentially from the N-terminus to the C-terminus:
[0073] (1) A nuclear localization signal polypeptide with a sequence as shown in SEQ ID NO.1;
[0074] (2) A TadA-8e polypeptide with the sequence shown in SEQ ID NO.2;
[0075] (3) A linked polypeptide with a sequence as shown in SEQ ID NO.3;
[0076] (4) A *Saccharomyces cerevisiae* spCas9n(D10A) polypeptide with the sequence shown in SEQ ID NO.4, 2-1368;
[0077] (5) Another linked polypeptide with a sequence as shown in SEQ ID NO.3;
[0078] (6) Variants of the HDG4 peptide;
[0079] (7) Another nuclear localization signal polypeptide with a sequence as shown in SEQ ID NO.1;
[0080] The variant of the HDG4 polypeptide is a polypeptide containing the R165E&Y179F double mutation, based on the wild-type mouse Aag polypeptide sequence shown in SEQ ID NO.5.
[0081] Furthermore, the ACBE2.0 fusion protein of the present invention is composed of the following blocks arranged sequentially from the N-terminus to the C-terminus:
[0082] (1) A nuclear localization signal polypeptide with a sequence as shown in SEQ ID NO.1;
[0083] (2) A first Cas9n polypeptide fragment with a sequence as shown in SEQ ID NO.4, 2-1248;
[0084] (3) A linked polypeptide with a sequence as shown in SEQ ID NO.3;
[0085] (4) A variant of the HDG4 peptide;
[0086] (5) A TadA-8e polypeptide with the sequence shown in SEQ ID NO.2;
[0087] (6) Another linked polypeptide with a sequence as shown in SEQ ID NO.3;
[0088] (7) A second Cas9n polypeptide fragment with a sequence as shown in SEQ ID NO.4, numbers 1249-1368;
[0089] (8) Another nuclear localization signal polypeptide with a sequence as shown in SEQ ID NO.1;
[0090] The variant of the HDG4 polypeptide is a polypeptide containing the R165E&Y179F double mutation, based on the wild-type mouse Aag polypeptide sequence shown in SEQ ID NO.5.
[0091] Furthermore, the ACBE2Q fusion protein of the present invention is composed of the following blocks arranged sequentially from the N-terminus to the C-terminus:
[0092] (1) A nuclear localization signal polypeptide with a sequence as shown in SEQ ID NO.1;
[0093] (2) A first Cas9n polypeptide fragment with a sequence as shown in SEQ ID NO.4, 2-1248;
[0094] (3) A linked polypeptide with a sequence as shown in SEQ ID NO.3;
[0095] (4) A variant of the HDG4 peptide;
[0096] (5) A variant of the TadA-8e polypeptide (N108Q) with the sequence shown in SEQ ID NO.9;
[0097] (6) Another linked polypeptide with a sequence as shown in SEQ ID NO.3;
[0098] (7) A second Cas9n polypeptide fragment with a sequence as shown in SEQ ID NO.4, numbers 1249-1368;
[0099] (8) Another nuclear localization signal polypeptide with a sequence as shown in SEQ ID NO.1;
[0100] The variant of the HDG4 polypeptide is a polypeptide containing the R165E&Y179F double mutation, based on the wild-type mouse Aag polypeptide sequence shown in SEQ ID NO.5.
[0101] When the single-base gene editing system of the present invention edits the target gene,
[0102] The efficiency of directed mutations from A to C in the justice chain is greater than 20%, and the bystander effect is less than 20%.
[0103] Preferably, the efficiency of directed mutations from A to C in the justice chain is greater than 30% and the bystander effect is less than 15%.
[0104] More preferably, the efficiency of directed mutations from A to C in the justice chain is greater than 40% and the bystander effect is less than 10%.
[0105] The present invention also relates to a nucleotide fragment encoding the fusion protein, a vector or host containing the nucleotide fragment.
[0106] This invention also relates to the following applications of the single-base gene editing system:
[0107] (1) Targeted editing of eukaryotic or prokaryotic cells;
[0108] (2) To prepare drugs for treating diseases caused by SNP mutations from C·G to A·T;
[0109] (3) Used for crop genetic breeding;
[0110] (4) Used to prepare animal models.
[0111] The beneficial effects of this invention are as follows:
[0112] (1) By fusing an artificially evolved mouse-derived 3-methyladenine glycosylase (Aag) variant with a monomeric adenosine deaminase Tad-8e or its variant and Cas9n with impaired catalytic activity, a single-base gene editing system containing an ACBE1.0 / 2.0 / 2Q fusion protein was constructed, enabling base transversions from A to C (sense strand) or from T to G (antisense strand), with the highest transversion efficiency of A·T to C·G reaching 45.2%.
[0113] (2) The single-base gene editing system described in this invention has great therapeutic potential for disease-related SNPs with C·G to A·T mutations, and can also be used in the preparation of disease models, crop genetic breeding and other fields.
[0114] [Terminology Explanation]
[0115] In this invention,
[0116] The term "fusion protein" refers to a protein product obtained by linking the coding regions of two or more genes through gene recombination, chemical methods, or other suitable methods, and expressing the recombinant protein under the control of the same regulatory sequence. Unless otherwise specified, the first block polypeptide at the N-terminus of the fusion protein is linked to the N-terminus of the next block (or linker fragment) polypeptide, and so on; therefore, the N-terminus of the polypeptide located in the N-terminal block of the fusion protein is the N-terminus of the fusion protein, and the C-terminus of the polypeptide located in the C-terminal block of the fusion protein is the C-terminus of the fusion protein. Unless otherwise specified, when constructing the fusion protein described in this invention, except for the linker polypeptide, only the polypeptide of the first block (bNLS) located at the N-terminus of the fusion protein has a complete sequence containing the first amino acid (M), while the polypeptide sequences of other functional blocks do not contain the first amino acid (M) shown in the sequence listing.
[0117] The term "Cas9 protein" refers to the CRISPR-Cas system, comprised of clustered, regularly spaced short palindromic repeats (CRISPR) and associated proteins (Cas), which forms the antiphage immune system present in many bacteria and most archaea. Cas9 proteins are class 2 effector proteins within the larger class of Cas proteins. Cas9 proteins typically contain two cleaving domains: the HNH domain and the RuvC domain. The HNH domain cleaves the DNA strand complementary to the crRNA, while the RuvC domain cleaves the non-complementary strand. Mutating Cas9 to remove nuclease activity from either the RuvC or HNH catalytic domain results in Cas9n, which produces only a single-strand cleavage when interacting with the DNA double helix. Typically, a conversion from aspartic to alanine (e.g., D10A as described in this application) in the RuvC domain produces Cas9n, or a conversion from histidine to alanine (e.g., H840A or H839A) in the HNH domain also produces Cas9n.
[0118] The term "gRNA," also known as "guide RNA," is a single-stranded RNA with a 5' end that is complementary to the target DNA. The gRNA uses this complementary sequence to target Cas9 to a specific genomic site. Typically, in gene editing systems, there is a PAM sequence following (or preceding) the gRNA.
[0119] The term "nuclear localization signal (NLS)," also known as "nuclear localization sequence," is usually a short amino acid sequence that interacts with nuclear vectors to enable proteins to be transported into the cell nucleus. It consists of at least 4-8 basic amino acids, such as Pro, Lys, and Arg.
[0120] The term "Aag" refers to mouse-derived 3-methyladenine glycosidase, which has the ability to recognize / remove hypoxanthine (I).
[0121] The term "Tad-8e" refers to ABE (adenine base editor), which is a fusion of adenosine deaminase TadA and Cas9 protein. Artificially evolved TadA catalyzes the deamination of adenine (A) in single-stranded DNA into hypoxanthine (I). TadA-8e is the latest evolved version of TadA.
[0122] For the definitions of other technical terms, refer to common knowledge familiar to those skilled in the art. Attached Figure Description
[0123] Figure 1 Schematic diagram of the protein structure of the ACBE1.0 / 2.0 / 2Q series constructs.
[0124] Figure 2 1. Sequence and structural alignment of mouse-derived Aag, rat-derived APDG, human-derived ANPG, and Bacillus subtilis-derived Aag; 2A. Mouse-derived Aag sequence; 2B. Rat-derived APDG sequence; 2C. Human-derived ANPG sequence; 2D. Bacillus subtilis-derived Aag sequence.
[0125] Figure 3 Comparison of edited products of mouse-derived 3-methyladenine glycosidase mutant (Aag mutant) at 293T;
[0126] 3A, EMX1-sg7 target editing effect;
[0127] 3B. A-to-C editing effects at different sites on the CCR5-sg1 and EMXl-sg7 genes: (See figure) x-axis A2, A5, A7 represents different positions of the A base. Specifically, in this system, the sgRNA binds to a 20 nt target DNA sequence, starting from 5' as the first... In the position sequence, A2 indicates that there is an A base at the 2nd position, A5 indicates that there is an A base at the 5th position, and A7 indicates that there is an A base at the 7th position. (The same applies below.)
[0128] Figure 4 The structural diagram of the vector used in the ACBE process described in this invention is as follows: 4A, AXBE vector; 4B, ABE8e vector; 4C, U6-sgRNA-EF1α-GFP.
[0129] Figure 5A comparison of the editing efficiency of ACBE1.0 series and control ABE8e, AXBE, and ACBE1.0 at the RUNX1-sgRNA1 target site on 293T. In the figure, transversion represents A to C target editing, and transition represents bystander effect non-target editing. The ACBE1.0 system in the figure uses ACBE1.0 (R165E, Y179F combined mutation).
[0130] Figure 6 The graph compares the editing efficiency of A base sites at two targets on 293T using ABE8e, AXBE, ACBE1.0, ACBE2.0, and ACBE2Q. Specifically, the ACBE1.0 system uses ACBE1.0 (R165E & Y179F combined mutation), the ACBE2.0 system uses ACBE2.0-07 protein, and the ACBE2Q system uses ACBE2Q-07 protein. Detailed Implementation
[0131] Materials and methods, and the sequence structure of the basic sequence units of the basic ACBE construct.
[0132] The mice involved in this invention are C57 / BL6 strain mice;
[0133] See AXBE carrier structure Figure 4 The structure of the ABE8e vector (Addgene#138489) is shown below. Figure 4 B, the structure of U6-sgRNA-EF1α-GFP is shown below. Figure 4 C;
[0134] BBSI enzyme (Thermo);
[0135] Transfection reagent (PEI, Polysciences);
[0136] Cell genome extraction kit (Tiangen DP304);
[0137] Hitom reagent kit (Novogene);
[0138] HEK293T cells, cell source (ATCC CRL-3216), culture medium was DMEM (Gibco) medium supplemented with 10% fetal bovine serum (Gibco), culture parameters: all cells were cultured in an incubator at 37℃ and 5% CO2.
[0139] The amino acid sequence of bNLS is shown in SEQ ID NO.1.
[0140] SEQ ID NO.1: MKRTADGSEFESPKKKRKV;
[0141] The amino acid sequence of TadA-8e is shown in SEQ ID NO.2.
[0142] SEQ ID NO.2:
[0143] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN;
[0144] The amino acid sequence of TadA-8e(N108Q) is shown in SEQ ID NO.9.
[0145] SEQ ID NO.9:
[0146] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVR Q SKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN;
[0147] The amino acid sequence of the linker is shown in SEQ ID NO.3, 10-21.
[0148] SEQ ID NO.3:SGGSSGGSSGSETPGTSESATPESSGGSSGGS
[0149] SEQ ID NO.10: APTLADLAVDLAALRPLEHPNPPLQRAAEALL
[0150] SEQ ID NO.11:EAAAKEAAAKAAAK
[0151] SEQ ID NO.12: GGGGSGGGGSGGGGS
[0152] SEQ ID NO.13:LEASPSNGSGT
[0153] SEQ ID NO.14:SGGSPKKRKVGSSGS
[0154] SEQ ID NO.15:SGSETPGTSESATPES
[0155] SEQ ID NO.16: AEAAAAKEAAAKA
[0156] SEQ ID NO.17: SGGSGGSGGS
[0157] SEQ ID NO.18: GGSGGSGGS
[0158] SEQ ID NO.19: GGGGS
[0159] SEQ ID NO.20: SGGS
[0160] SEQ ID NO.21: SGGSKRTADGSEFEPKKKRKVGSG
[0161] The amino acid sequence of spCas9n(D10A) is shown in SEQ ID NO.4.
[0162] SEQ ID NO.4:
[0163]
[0164] The order of the functional blocks in ABE8e is: bNLS+TadA8e+Linker+spCas9n(D10A)+bNLS;
[0165] The functional blocks of AH4(AXBE) are arranged in the following order: bNLS+TadA8e+Linker+spCas9n(D10A)+Linker+Aag(mouse)+bNLS;
[0166] The amino acid sequence of mouse-derived Aag is shown in SEQ ID NO. 5.
[0167] SEQ ID NO.5:
[0168] MPARGGSARPGRGALKPVSVTLLPDTEQPPFLGRARRPGNARAGSLVTGYHEVGQMPAPLSRKIGQKKQRLADSEQQQTPKERLLSTPGLRRSIYFSSPEDHSGRLGPEFFDQPAVTLARAFLGQVLVRRLADGTELRGRIVETEAYLGPEDEAAHSRGGRQTPRNR GMFMKPGTLYVYLIYGMYFCLNVSSQGAGACVLLRALEPLEGLETMRQLRNSLRKSTVGRSLKDRELCSGPSKLCQALAIDKSFDQRDLAQDDAVWLEHGPLESSSPAVVVAAARIGIGHAGEWTQKPLRFYVQGSPWVSVVDRVAEQMDQPQQTACSEGLLIVQK;
[0169] The amino acid sequence of rat-derived APDG is shown in SEQ ID NO. 6.
[0170] SEQ ID NO.6:
[0171] MRGRGGTARLGRGSLKPVSVVLPDTEHPAFPGRTRRPGNARAGSQVTGSREVGQMPAPLSRKIGQKKQQLAQSEQQQTPKERLSSTPGLLRSIYFSSPEDRPARLGPEYFDQPAVTLARAFLGQVLVRRLADGTELRGRIVETEAYLGPEDEAAHSRGGRQTPRN RGMFMKPGTLYVYLIYGMYFCLNVSSQGAGACVLLRALEPLEGLETMRQLRNSLRKSTVGRSLKDRELCNGPSKLCQALAIDKSFDQRDLAQDEAVWLEHGPLESSSPAVVAAARIGIGHAGEWTQKPLRFYVQGSPWVSVVDRVAEQMYQPQQTACSDCSKVK;
[0172] The amino acid sequence of human ANPG is shown in SEQ ID NO.7.
[0173] SEQ ID NO.7:
[0174] MVTPALQMKKPKQFCRRMGQKKQRPARAGQPHSSSDAAQAPAEQPHSSSDAAQAPCPRERCLGPPTTPGPYRSIYFSSPKGHLTRLGLEFFDQPAVPLARAFLGQVLVRRLPNGTELRGRIVETEAYLGPEDEAAHSRGGRQTPRNRGM FMKPGTLYVYIIYGMYFCMNISSQGDGACVLLRALEPLEGLETMRQLRSTLRKGTASRVLKDRELCSGPSKLCQALAINKSFDQRDLAQDEAVWLERGPLEPSEPAVVAAARVGVGHAGEWARKPLRFYVRGSPWVSVVDRVAEQDTQA;
[0175] The amino acid sequence of Aag derived from Bacillus subtilis is shown in SEQ ID NO. 8.
[0176] SEQ ID NO.8:
[0177] MTREKNPLPITFYQKTALELAPSLLGCLLVKETDEGTASGYIVETEAYMGAGDRAAHSFNNRRTKRTEIMFAEAGRVYTYVMHTHTLLNVVAAEEDVPQAVLIRAIEPHEGQLLMEERRPGRSPREWTNGPGKLTKALGVTMNDYGRWITEQPLYIESGYTPEAISTGPRIGIDNSGEARDYPWRFWVTGNRYVSR;
[0178] Example 1: Structural design, expression construction, and functional verification of the fusion protein ACBE1.0
[0179] Based on DNA base excision repair mechanisms, modification of mouse-derived 3-methyladenine glycosidase (Aag) may alter its ability to recognize and excise inosine, thereby affecting A-to-C or T-to-G base transversions. Different species may also have evolved glycosidases (HDGs) that effectively recognize / excise inosine. Therefore, sequence alignment was performed between mouse-derived Aag, rat-derived APDG, human-derived ANPG, and Bacillus subtilis-derived Aag. Figure 2 It is speculated that 12 amino acids may affect substrate recognition of mouse-derived 3-methyladenine glycosidase.
[0180] Typical structures of the ACBE1.0 system, ACBE2.0 system, ACBE2Q system, and control group fusion proteins ABE8e and AXBE are shown below. Figure 1 As shown.
[0181] 1. Plasmid design and construction
[0182] 1.1 Design of Aag mutants and design of the fusion protein construct ACBE1.0, and design of target genes.
[0183] (1) It was hypothesized that 12 amino acids might affect the substrate recognition of mouse-derived 3-methyladenine glycosidase (Aag). By altering the hydrophobicity or polarity of its amino acids, 32 Aag variants were obtained (single mutation site column in Table 1-1). These variants were then fused with Tad-8e and cas9n to design 32 single-mutant fusion protein constructs (ACBE1.0 structure, see...). Figure 1 );
[0184] (2) The single mutants were combined to obtain 17 combinatorial variants (co-mutation sites listed in Table 1-1), and these were fused with Tad-8e and cas9n to design 17 fusion protein constructs (ACBE1.0 system structure).
[0185] (3) At the same time, one endogenous target from humans, EMX1-sg7, and one endogenous target from humans, CCR5-sg1, were designed (Table 2). At the same time, endogenous test targets, RUNX1-sgRNA1, ABL-sg16, and EGFR-sg41, were designed (Table 2).
[0186] 1.2 Cloning and Expression of Fusion Protein Constructs
[0187] Thirty-two Aag monovariate sequences were synthesized according to the sequence mutation information in Table 1-1, and then seamlessly cloned and assembled using AXBE as a vector.
[0188] Seventeen Aag combinatorial variants were synthesized according to the sequence mutation information in Table 1-1, and then seamlessly cloned and assembled using AXBE as a vector.
[0189] Two oligos were synthesized according to Table 2, with CACC added to the positive strand and AAAC added to the reverse strand. These were then ligated to the U6-sgRNA-EF1α-GFP that had been digested with BbsI to construct the sgRNA plasmid.
[0190] 1.3 The constructed plasmids were sequenced by Sanger sequencing to ensure that the sequences were completely correct.
[0191] Table 1-1 Sequence structure design of mouse-derived Aag mutants
[0192]
[0193]
[0194] Table 2. Targets and sequences used
[0195]
[0196] 2. Cell transfection
[0197] 2.1 On day 1, 293T cells were seeded into 24-well plates.
[0198] HEK293T cells were digested at a rate of 2 × 10⁻⁶. 5 Cells / wells are seeded into 96-well plates.
[0199] Note: After cell resuscitation, cells generally need to be passaged twice before they can be used for transfection experiments.
[0200] 2.2 Day 2 Transfection
[0201] (1) Observe the state of cells in each well.
[0202] Note: The cell density should be 70%-90% before transfection and the cells should be in normal condition.
[0203] (2) Plasmid transfection levels are as follows:
[0204] Using the newly constructed ACBE1.0 fusion protein plasmids from step 1 above (U6-sgRNA-EF1α-GFP = 750 ng: 250 ng plasmid), and PEI (3 μL PEI per 1 μg plasmid), co-transfect the HEK293T host with ABE8e (structure see [link]). Figure 1 ), AXBE (structure see) Figure 1 As a control, n = 3 holes / group.
[0205] 3. Genome extraction and preparation of amplicon libraries
[0206] (1) 72 h after transfection, genomic DNA was extracted from cells using a cell genomic DNA extraction kit (DP304);
[0207] (2) Then, following the operating procedure of the Hitom kit, design the corresponding identification primers (Table 3), that is, add the bridging sequence 5'-ggagtgagtacggtgtgc-3' to the 5' end of the forward identification primer and add the bridging sequence 5'-gagttggatgctggatgg-3' to the 5' end of the reverse identification primer to obtain one round of PCR products.
[0208] Table 3. Primers used for target identification
[0209]
[0210] (3) Then, the first round of PCR products were used as templates for a second round of PCR. The products were mixed together, gel-cleaved and purified, and then sent to a third-party sequencing company for sequencing.
[0211] 4. Analysis and statistics of deep sequencing results
[0212] The deep sequencing results were analyzed using the BE-analyzer website, specifically the editing efficiency from A to C, A to T, and A to G, and statistical graphs were generated using GraphPad Prism 9.1.0.
[0213] 5. Results Analysis
[0214] 5.1 Regarding the EMX1 target
[0215] (1) Using ABE8e and AXBE (wild-type mouse Aag fusion TadA-8e and Cas9n) as controls, single mutants such as R165E, Y179F, G183Q, G183H, M184F, M184W, and Y185F can significantly improve the editing efficiency from A to C (editing target is EMX1-sg7);
[0216] (2) These superior mutations were further combined and screened. Most of the combined mutations further improved the editing efficiency from A to C. Among them, the combination of R165E and Y179F had the best effect. ABE8e had an efficiency of 0.13% in producing A to C, AXBE had an efficiency of 14.1% in producing A to C, and the combination of R165E and Y179F had an efficiency of 39% in producing A to C. Compared to AXBE, editing efficiency is improved. 1.8 times ( Figure 3 A, 3B, Table 4);
[0217] 5.2 Regarding the CCR5 target,
[0218] Deep sequencing results showed that ABE8e had an A-to-C conversion efficiency of 0.1%, AXBE had an A-to-C conversion efficiency of 10.3%, and the combination of R165E and Y179F had an A-to-C conversion efficiency of 25.9% (with CCR5-sg1 as the editing target). Compared to AXBE Editing efficiency increased by 1.5 times. ( Figure 3 B, Table 5);
[0219] The results of the A-to-C editing efficiency test of the ACBE1.0 system for different mutants are shown in Tables 4 and 5.
[0220] Table 4. Editing efficiency results of various mutations in the ACBE1.0 system at the target site EMX1-sg7 (unit, %)
[0221]
[0222]
[0223]
[0224] Table 5. Editing efficiency of various combined mutations in the ACBE1.0 system at the target CCR5-sg1 (unit, %)
[0225]
[0226]
[0227] Example 2: Structural design, expression construction, and functional verification of the ACBE2.0 / ACBE2Q fusion protein system
[0228] To further improve the efficiency of A-to-C generation in the ACBE1.0 system and increase the contact range between Aag and the substrate, we will introduce Aag with the R165E & Y179F combined mutation or TadA-8e (with or without the N108Q mutation) into the middle of Cas9, and design according to the positional variation:
[0229] ACBE1.0-Ce01~ACBE1.0-Ce10 series and ACBE1.0-Ce01Q~ACBE1.0-Ce10Q series fusion proteins (Tables 1-2 and 1-3)
[0230] The new system is named ACBE2.0 (without the aforementioned TadA-8e N108Q mutation) or ACBE2Q system (with the aforementioned TadA-8e N108Q mutation), and the human endogenous test target RUNX1-sgRNA1 (Table 2) is designed for testing.
[0231] 1. Plasmid design and construction
[0232] 1.1 Introduce Aag or TadA-8e with R165E and Y179F mutations into the middle of Cas9 (between positions 1248 and 1249 of the Cas9 protein polypeptide sequence).
[0233] (1) Using ACBE1.0 as a template, and based on the changes in the positions of Tad8e and Aag variants, as well as the addition or removal of linkers, the following constructor is designed:
[0234] 1) Ten ACBE1.0-Ce01 to ACBE1.0-Ce10 constructs (ACBE2.0) (Table 1-2),
[0235] 2) The endogenous test target RUNX1-sgRNA1 derived from humans (sequence shown in Table 2),
[0236] 3) Introducing N108Q into 10 Tad-8e molecules yields the ACBE1.0-Ce01Q~ACBE1.0-Ce10Q constructs (ACBE2Q) (Table 1-3).
[0237] 4) ABL-sg16 and EGFR-sg41 test targets (sequences shown in Table 2)
[0238] The construction method is the same as in Section 1.2 of Example 1.
[0239] 1.2 The plasmid constructed in 1.1 was sequenced by Sanger sequencing to ensure it was completely correct.
[0240] Table 1-2, Sequence Structure Design of ACBE1.0-Ce Construct (ACBE2.0)
[0241]
[0242]
[0243] Note: The amino acid sequence of each block in the ACBE1.0-CeN construct in the table is (N-terminus to C-terminus).
[0244] Table 1-3, Sequence Structure Design of ACBE1.0-CeQ Construct (ACBE2Q)
[0245]
[0246]
[0247] Note: The amino acid sequence of each block in the ACBE1.0-CeNQ construct in the table is (N-terminus to C-terminus).
[0248] 2. Cell transfection
[0249] 2.1 On day 1, 293T cells were seeded into 24-well plates.
[0250] HEK293T cells were digested at a rate of 2 × 10⁻⁶. 5 Cells / wells are seeded into 96-well plates.
[0251] Note: After cell resuscitation, cells generally need to be passaged twice before they can be used for transfection experiments.
[0252] 2.2 Day 2 Transfection
[0253] (1) Observe the state of cells in each well.
[0254] Note: The cell density should be 70%-90% before transfection and the cells should be in normal condition.
[0255] (2) Plasmid transfection levels are as follows:
[0256] Using the newly constructed ACBE1.0-Ce or ACBE1.0-CeQ plasmid from step 1 above: U6-sgRNA-EF1α-GFP = 750 ng: 250 ng, with ABE8e, AXBE, and ACBE1.0 as controls, set n = 3 wells / group.
[0257] The transfection method and steps are the same as in Section 1.2.
[0258] 3. Genome extraction and preparation of amplicon libraries
[0259] (1) 72 h after transfection, genomic DNA was extracted from cells using a cell genomic DNA extraction kit (DP304);
[0260] (2) Design the corresponding identification primers according to the Hitom kit operation procedure (see Table 3). That is, add the bridging sequence 5'-ggagtgagtacggtgtgc-3' to the 5' end of the forward identification primer and add the bridging sequence 5'-gagttggatgctggatgg-3' to the 5' end of the reverse identification primer to obtain one round of PCR products.
[0261] (3) Using the first-round PCR product as a template, a second round of PCR was performed. The products were mixed together, gel-cleaved and purified, and then sent to a third-party sequencing company for sequencing.
[0262] 4. Analysis and Statistics of Deep Sequencing Results
[0263] The deep sequencing results were analyzed using the BE-analyzer website, specifically the editing efficiency from A to C, A to T, and A to G; and statistical graphs were generated using GraphPad Prism 9.1.0.
[0264] 5. Results Analysis
[0265] To further examine the working characteristics of ACBE2.0 / ACBE2Q, ABE8e, AXBE, and ACBE1.0 were used as controls again, and two additional targets (corresponding to ABL and EGFR in Table 2) were designed for evaluation.
[0266] 5.1 For the ACBE1.0-Ce (ACBE2.0) series of fusion proteins:
[0267] (1) Compared to ACBE1.0, except for ACBE1.0-Ce03, all other ACBE1.0-Ce series improve the editing efficiency from A to C to varying degrees. Among them, ACBE1.0-Ce07 (ACBE2.0-07) increases the editing efficiency from the initial 23.1% to 45.2%. Figure 5 );
[0268] (2) It was also found that the embedded version (ACBE1.0-Ce series fusion protein) also increases bystander A to G editing (bystander editing refers to off-target editing of adjacent sites on both sides of the target editing site);
[0269] 5.2 For the ACBE1.0-CeQ (ACBE2Q) construct, the results show:
[0270] (1) The ACBE1.0-CeQ series all reduce the editing efficiency of bystanders from A to G to varying degrees. The ACBE1.0-Ce07Q (ACBE2Q-07) almost completely eliminates the generation of bystander A to G, resulting in an editing efficiency of 31% for A to C. Figure 5 ).
[0271] (2) To further evaluate the editing characteristics of ACBE2.0 and 2Q, two more endogenous test targets, ABL-sg16 and EGFR-sg41 (Table 2), were designed and compared with ABE8e, AXBE, and ACBE1.0. The results are shown in (see Table 2). Figure 6 )show:
[0272] 1) At the ABL-sg16 site, ACBE2.0-07 achieved an A-to-C editing efficiency of 39.2% and an A-to-C product purity of 54%. The A-to-C editing at the more distant A11 site was further improved. In contrast, ACBE2Q-07 achieved an A-to-C editing efficiency of 33.9% and an A-to-C product purity of 53%, with almost no bystander A-to-G editing.
[0273] 2) At the EGFR-sg41 site, ACBE2.0-07 achieved an A-to-C editing efficiency of 34.4% and an A-to-C product purity of 50%, while ACBE2Q-07 achieved an A-to-C editing efficiency of 32% and an A-to-C product purity of 48%. Regarding the editing of bystander A3, compared to ACBE1.0 which produced 61% A-to-G, ACBE2Q-07 only produced a very low level of A-to-G at 4.3% (almost no bystander A-to-G editing).
[0274] The editing efficiency of each ACBE1.0-Ce / Q on the target genes is shown in Table 6-8 below.
[0275] Table 6. Editing efficiency results of ACBE1.0-Ce / Q at target RUNX1-sg1 (unit, %)
[0276]
[0277]
[0278]
[0279] Table 7. Editing efficiency results of ABE8e, AXBE, ACBE1.0, and ACBE2.0 at target ABL-sg16 (unit, %)
[0280]
[0281] Table 8. Editing efficiency results of ABE8e, AXBE, ACBE1.0, and ACBE2.0 at the target EGFR-sg41 (unit, %)
[0282]
[0283] Finally, it should be noted that the above embodiments are only used to help those skilled in the art understand the essence of the present invention, and are not intended to limit the scope of protection of the present invention. SEQUENCE LISTING <110> East China Normal University & Shanghai Bangyao Biomedical Co., Ltd. <120> A novel gene editing system mediating A-to-C or T-to-G mutations and its applications <130> BYP2021‑0009 <160> 21 <170> PatentIn version 3.3 <210> 1 <211> 19 <212> PRT <213> Artificial sequence <400> 1 Met Lys Arg Thr Ala Asp Gly Ser Glu Phe Glu Ser Pro Lys Lys Lys 1 5 10 15 Arg Lys Val <210> 2 <211> 167 <212> PRT <213> Artificial sequence <400> 2 Met Ser Glu Val Glu Phe Ser His Glu Tyr Trp Met Arg His Ala Leu 1 5 1o 15 Thr Leu Ala Lys Arg Ala Arg Asp Glu Arg Glu Val Pro Val Gly Ala 20 25 30 Val Leu Val Leu Asn Asn Arg Val Ile Gly Glu Gly Trp Asn Arg Ala 35 40 45 Ile Gly Leu His Asp Pro Thr Ala His Ala Glu Ile Met Ala Leu Arg 50 55 60 Gln Gly Gly Leu Val Met Gln Asn Tyr Arg Leu Ile Asp Ala Thr Leu 65 70 75 80 Tyr Val Thr Phe Glu Pro Cys Val Met Cys Ala Gly Ala Met Ile His 85 90 95 It should be noted that there is a possible error in the original text where "1o" in line should probably be "15". This has been left as is in the translation for the purpose of following the instructions precisely. Ser Arg Ile Gly Arg Val Val Phe Gly Val Arg Asn Ser Lys Arg Gly 100 105 110 Ala Ala Gly Ser Leu Met Asn Val Leu Asn Tyr Pro Gly Met Asn His 115 120 125 Arg Val Glu Ile Thr Glu Gly Ile Leu Ala Asp Glu Cys Ala Ala Leu 130 135 140 Leu Cys Asp Phe Tyr Arg Met Pro Arg Gln Val Phe Asn Ala Gln Lys 145 150 155 160 Lys Ala Gln Ser Ser Ile Asn 165 <210> 3 <211> 32 <212> PRT <213> Artificial Sequence <400> 3 Ser Gly Gly Ser Ser Gly Gly Ser Ser Gly Ser Glu Thr Pro Gly Thr 1 5 10 15 Ser Glu Ser Ala Thr Pro Glu Ser Ser Gly Gly Ser Ser Gly Gly Ser 20 25 30 <210> 4 <211> 1368 <212> PRT <213> Artificial Sequence <400> 4 Met Asp Lys Lys Tyr Ser Ile Gly Leu Ala Ile Gly Thr Asn Ser Val 1 5 10 15 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Asn Leu Ile 35 40 45 Gly Ala Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 Lys Arg Thr Ala Arg Arg Tyr Thr Arg Arg Lys Asn Arg With Cys 65 70 75 80 Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Asp Ser 85 90 95 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 115 120 125 His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Leu Val Asp 130 135 140 Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His 145 150 155 160 Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro 165 170 175 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 180 185 190 Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 195 200 205 Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn 210 215 220 Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 225 230 235 240 Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 245 250 255 Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp 260 265 270 Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp 275 280 285 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp 290 295 300 Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys 325 330 335 Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe 340 345 350 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 355 360 365 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 370 375 380 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 405 410 415 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 420 425 430 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 435 440 445 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu 465 470 475 480 Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr 485,490,495 Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 500 505 510 Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Thr Lys Val Lys 515,520,525 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 530 535 540 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 545 550 555 560 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Ile Glu Cys Phe Asp 565,570,575 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 580,585,590 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595,600,605 Asn Glu Glu Asn Glu Asp With Glu Asp With Val With Thr With Thr 610 615 620 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755 760 765 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770 775 780 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785 790 795 800 Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu 805 810 815 Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg 820 825 830 Leu Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys 835 840 845 Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys 865 870 875 880 Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys 885 890 895 Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 900 905 910 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr 915 920 925 Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys Tyr Asp 930 935 940 Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Leu Lys Ser 945 950 955 960 Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 965 970 975 Glu Ile Asn Asn Tyr His His Ala His Asp Ala Tyr Leu Asn Ala Val 980 985 990 Val Gly Thr Ala Leu Ile Lys Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala 1010 1015 1020 Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1025 1030 1035 Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1040 1045 1050 Asn Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu 1055 1060 1065 Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1070 1075 1080 Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr 1085 1090 1095 Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys 1100 1105 1110 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1115 1120 1125 Light Light Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1130 1135 1140 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1145 1150 1155 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1160 1165 1170 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1175 1180 1185 Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu 1190 1195 1200 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1205 1210 1215 Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val 1220 1225 1230 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1250 1255 1260 His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys 1265 1270 1275 Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala 1280 1285 1290 Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn 1295 1300 1305 Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala 1310 1315 1320 Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser 1325 1330 1335 Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr 1340 1345 1350 Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 <210> 5 <211> 333 <212> PRT <213> Mus musculus <400> 5 Met Pro Ala Arg Gly Gly Ser Ala Arg Pro Gly Arg Gly Ala Leu Lys 1 5 10 15 Pro Val Ser Val Thr Leu Leu Pro Asp Thr Glu Gln Pro Pro Phe Leu 20 25 30 Gly Arg Ala Arg Arg Pro Gly Asn Ala Arg Ala Gly Ser Leu Val Thr 35 40 45 Gly Tyr His Glu Val Gly Gln Met Pro Ala Pro Leu Ser Arg Lys Ile 50 55 60 Gly Gln Lys Lys Gln Arg Leu Ala Asp Ser Glu Gln Gln Gln Thr Pro 65 70 75 80 Lys Glu Arg Leu Leu Ser Thr Pro Gly Leu Arg Arg Ser Ile Tyr Phe 85 90 95 Ser Ser Pro Glu Asp His Ser Gly Arg Leu Gly Pro Glu Phe Phe Asp 100 105 110 Gln Pro Ala Val Thr Leu Ala Arg Ala Phe Leu Gly Gln Val Leu Val 115 120 125 Arg Arg Leu Ala Asp Gly Thr Glu Leu Arg Gly Arg Ile Val Glu Thr 130 135 140 Glu Ala Tyr Leu Gly Pro Glu Asp Glu Ala Ala His Ser Arg Gly Gly 145 150 155 160 Arg Gln Thr Pro Arg Asn Arg Gly Met Phe Met Lys Pro Gly Thr Leu 165 170 175 Tyr Val Tyr Leu Ile Tyr Gly Met Tyr Phe Cys Leu Asn Val Ser Ser 180 185 190 Gln Gly Ala Gly Ala Cys Val Leu Leu Arg Ala Leu Glu Pro Leu Glu 195 200 205 Gly Leu Glu Thr Met Arg Gln Leu Arg Asn Ser Leu Arg Lys Ser Thr 210 215 220 Val Gly Arg Ser Leu Lys Asp Arg Glu Leu Cys Ser Gly Pro Ser Lys 225 230 235 240 Leu Cys Gln Ala Leu Ala Ile Asp Lys Ser Phe Asp Gln Arg Asp Leu 245 250 255 Ala Gln Asp Asp Ala Val Trp Leu Glu His Gly Pro Leu Glu Ser Ser 260 265 270 Ser Pro Ala Val Val Val Ala Ala Ala Arg Ile Gly Ile Gly His Ala 275 280 285 Gly Glu Trp Thr Gln Lys Pro Leu Arg Phe Tyr Val Gln Gly Ser Pro 290 295 300 Trp Val Ser Val Val Asp Arg Val Ala Glu Gln Met Asp Gln Pro Gln 305 310 315 320 Gln Thr Ala Cys Ser Glu Gly Leu Leu Ile Val Gln Lys 325 330 <210> 6 <211> 329 <212> PRT <213> Rattus norvegicus <400> 6 Met Arg Gly Arg Gly Gly Thr Ala Arg Leu Gly Arg Gly Ser Leu Lys 1 5 10 15 Pro Val Ser Val Val Leu Pro Asp Thr Glu His Pro Ala Phe Pro Gly 20 25 30 Arg Thr Arg Arg Pro Gly Asn Ala Arg Ala Gly Ser Gln Val Thr Gly 35 40 45 Ser Arg Glu Val Gly Gln Met Pro Ala Pro Leu Ser Arg Lys Ile Gly 50 55 60 Gln Lys Lys Gln Gln Leu Ala Gln Ser Glu Gln Gln Gln Thr Pro Lys 65 70 75 80 Glu Arg Leu Ser Ser Thr Pro Gly Leu Leu Arg Ser Ile Tyr Phe Ser 85 90 95 Ser Pro Glu Asp Arg Pro Ala Arg Leu Gly Pro Glu Tyr Phe Asp Gln 100 105 110 Pro Ala Val Thr Leu Ala Arg Ala Phe Leu Gly Gln Val Leu Val Arg 115 120 125 Arg Leu Ala Asp Gly Thr Glu Leu Arg Gly Arg Ile Val Glu Thr Glu 130 135 140 Ala Tyr Leu Gly Pro Glu Asp Glu Ala Ala His Ser Arg Gly Gly Arg 145 150 155 160 Gln Thr Pro Arg Asn Arg Gly Met Phe Met Lys Pro Gly Thr Leu Tyr 165 170 175 Val Tyr Leu Ile Tyr Gly Met Tyr Phe Cys Leu Asn Val Ser Ser Gln 180 185 190 Gly Ala Gly Ala Cys Val Leu Leu Arg Ala Leu Glu Pro Leu Glu Gly 195 200 205 Leu Glu Thr Met Arg Gln Leu Arg Asn Ser Leu Arg Lys Ser Thr Val 210 215 220 Gly Arg Ser Leu Lys Asp Arg Glu Leu Cys Asn Gly Pro Ser Lys Leu 225 230 235 240 Cys Gln Ala Leu Ala Ile Asp Lys Ser Phe Asp Gln Arg Asp Leu Ala 245 250 255 Gln Asp Glu Ala Val Trp Leu Glu His Gly Pro Leu Glu Ser Ser Ser 260 265 270 Pro Ala Val Val Ala Ala Ala Arg Ile Gly Ile Gly His Ala Gly Glu 275 280 285 Trp Thr Gln Lys Pro Leu Arg Phe Tyr Val Gln Gly Ser Pro Trp Val 290 295 300 Ser Val Val Asp Arg Val Ala Glu Gln Met Tyr Gln Pro Gln Gln Thr 305 310 315 320 Ala Cys Ser Asp Cys Ser Lys Val Lys 325 <210> 7 <211> 298 <212> PRT <213> Homo sapiens <400> 7 Met Val Thr Pro Ala Leu Gln Met Lys Lys Pro Lys Gln Phe Cys Arg 1 5 10 15 Arg Met Gly Gln Lys Lys Gln Arg Pro Ala Arg Ala Gly Gln Pro His 20 25 30 Ser Ser Ser Asp Ala Ala Gln Ala Pro Ala Glu Gln Pro His Ser Ser 35 40 45 Ser Asp Ala Ala Gln Ala Pro Cys Pro Arg Glu Arg Cys Leu Gly Pro 50 55 60 Pro Thr Thr Pro Gly Pro Tyr Arg Ser Ile Tyr Phe Ser Ser Pro Lys 65 70 75 80 Gly His Leu Thr Arg Leu Gly Leu Glu Phe Phe Asp Gln Pro Ala Val 85 90 95 Pro Leu Ala Arg Ala Phe Leu Gly Gln Val Leu Val Arg Arg Leu Pro 100 105 110 Asn Gly Thr Glu Leu Arg Gly Arg Ile Val Glu Thr Glu Ala Tyr Leu 115 120 125 Gly Pro Glu Asp Glu Ala Ala His Ser Arg Gly Gly Arg Gln Thr Pro 130 135 140 Arg Asn Arg Gly Met Phe Met Lys Pro Gly Thr Leu Tyr Val Tyr Ile 145 150 155 160 Ile Tyr Gly Met Tyr Phe Cys Met Asn Ile Ser Ser Gln Gly Asp Gly 165 170 175 Ala Cys Val Leu Leu Arg Ala Leu Glu Pro Leu Glu Gly Leu Glu Thr 180 185 190 Met Arg Gln Leu Arg Ser Thr Leu Arg Lys Gly Thr Ala Ser Arg Val 195 200 205 Leu Lys Asp Arg Glu Leu Cys Ser Gly Pro Ser Lys Leu Cys Gln Ala 210 215 220 Leu Ala Ile Asn Lys Ser Phe Asp Gln Arg Asp Leu Ala Gln Asp Glu 225 230 235 240 Ala Val Trp Leu Glu Arg Gly Pro Leu Glu Pro Ser Glu Pro Ala Val 245 250 255 Val Ala Ala Ala Arg Val Gly Val Gly His Ala Gly Glu Trp Ala Arg 260 265 270 Lys Pro Leu Arg Phe Tyr Val Arg Gly Ser Pro Trp Val Ser Val Val 275 280 285 Asp Arg Val Ala Glu Gln Asp Thr Gln Ala 290 295 <210> 8 <211> 196 <212> PRT <213> Bacillus subtilis <400> 8 Met Thr Arg Glu Lys Asn Pro Leu Pro Ile Thr Phe Tyr Gln Lys Thr 1 5 10 15 Ala Leu Glu Leu Ala Pro Ser Leu Leu Gly Cys Leu Leu Val Lys Glu 20 25 30 Thr Asp Glu Gly Thr Ala Ser Gly Tyr Ile Val Glu Thr Glu Ala Tyr 35 40 45 Met Gly Ala Gly Asp Arg Ala Ala His Ser Phe Asn Asn Arg Arg Thr 50 55 60 Lys Arg Thr Glu Ile Met Phe Ala Glu Ala Gly Arg Val Tyr Thr Tyr 65 70 75 80 Val Met His Thr His Thr Leu Leu Asn Val Val Ala Ala Glu Glu Asp 85 90 95 Val Pro Gln Ala Val Leu Ile Arg Ala Ile Glu Pro His Glu Gly Gln 100 105 110 Leu Leu Met Glu Glu Arg Arg Pro Gly Arg Ser Pro Arg Glu Trp Thr 115 120 125 Asn Gly Pro Gly Lys Leu Thr Lys Ala Leu Gly Val Thr Met Asn Asp 130 135 140 Tyr Gly Arg Trp Ile Thr Glu Gln Pro Leu Tyr Ile Glu Ser Gly Tyr 145 150 155 160 Thr Pro Glu Ala Ile Ser Thr Gly Pro Arg Ile Gly Ile Asp Asn Ser 165 170 175 Gly Glu Ala Arg Asp Tyr Pro Trp Arg Phe Trp Val Thr Gly Asn Arg 180 185 190 Tyr Val Ser Arg 195 <210> 9 <211> 167 <212> PRT <213> Artificial sequence <400> 9 Met Ser Glu Val Glu Phe Ser His Glu Tyr Trp Met Arg His Ala Leu 1 5 10 15 Thr Leu Ala Lys Arg Ala Arg Asp Glu Arg Glu Val Pro Val Gly Ala 20 25 30 Val Leu Val Leu Asn Asn Arg Val Ile Gly Glu Gly Trp Asn Arg Ala 35 40 45 Ile Gly Leu His Asp Pro Thr Ala His Ala Glu Ile Met Ala Leu Arg 50 55 60 Gln Gly Gly Leu Val Met Gln Asn Tyr Arg Leu Ile Asp Ala Thr Leu 65 70 75 80 Tyr Val Thr Phe Glu Pro Cys Val Met Cys Ala Gly Ala Met Ile His 85 90 95 Ser Arg Ile Gly Arg Val Val Phe Gly Val Arg Gln Ser Lys Arg Gly 100 105 110 Ala Ala Gly Ser Leu Met Asn Val Leu Asn Tyr Pro Gly Met Asn His 115 120 125 Arg Val Glu Ile Thr Glu Gly Ile Leu Ala Asp Glu Cys Ala Ala Leu 130 135 140 Leu Cys Asp Phe Tyr Arg Met Pro Arg Gln Val Phe Asn Ala Gln Lys 145 150 155 160 Lys Ala Gln Ser Ser Ile Asn 165 <210> 10 <211> 32 <212> PRT <213> artificial sequence <400> 10 Ala Pro Thr Leu Ala Asp Leu Ala Val Asp Leu Ala Ala Leu Arg Pro 1 5 10 15 Leu Glu His Pro Asn Pro Pro Leu Gln Arg Ala Ala Glu Ala Leu Leu 20 25 30 <210> 11 <211> 15 <212> PRT <213> Artificial sequence <400> 11 Glu Ala Ala Ala Lys Glu Ala Ala Ala Lys Glu Ala Ala Ala Lys 1 5 10 15 <210> 12 <211> 15 <212> PRT <213> Artificial sequence <400> 12 Gly Gly Gly Gly Ser Gly Gly Gly Gly Ser Gly Gly Gly Gly Ser 1 5 10 15 <210> 13 <211> 16 <212> PRT <213> Artificial sequence <400> 13 Leu Glu Ala Ser Pro Ser Asn Pro Gly Ala Ser Asn Gly Ser Gly Thr 1 5 10 15 <210> 14 <211> 15 <212> PRT <213> Artificial sequence <400> 14 Ser Gly Gly Ser Pro Lys Lys Arg Lys Val Gly Ser Ser Gly Ser 1 5 10 15 <210> 15 <211> 16 <212> PRT <213> Artificial sequence <400> 15 Ser Gly Ser Glu Thr Pro Gly Thr Ser Glu Ser Ala Thr Pro Glu Ser 1 5 10 15 <210> 16 <211> 12 <212> PRT <213> Artificial sequence <400> 16 Ala Glu Ala Ala Ala Lys Glu Ala Ala Ala Lys Ala 1 5 10 <210> 17 <211> 10 <212> PRT <213> Artificial sequence <400> 17 Ser Gly Gly Ser Gly Gly Ser Gly Gly Ser 1 5 10 <210> 18 <211> 9 <212> PRT <213> Artificial sequence <400> 18 Gly Gly Ser Gly Gly Ser Gly Gly Ser 1 5 <210> 19 <211> 5 <212> PRT <213> Artificial sequence <400> 19 Gly Gly Gly Gly Ser 1 5 <210> 20 <211> 4 <212> PRT <213> Artificial sequence <400> 20 Ser Gly Gly Ser 1 <210> twenty one <211> twenty four <212> PRT <213> Artificial sequence <400> twenty one Ser Gly Gly Ser Lys Arg Thr Ala Asp Gly Ser Glu Phe Glu Pro Lys 1 5 10 15 Lys Lys Arg Lys Val Gly Ser Gly 20
Claims
1. A single-base gene editing system, characterized in that, The single-base gene editing system described above can induce A-to-C directed mutations or T-to-G directed mutations in the target DNA; The single-base gene editing system comprises gRNA, and the main active region of the single-base gene editing system is the 5th to 7th bases at the 5' end of the gRNA; The gene editing system further comprises a fusion protein with single-base editing function, wherein the fusion protein with single-base editing function comprises the following functional blocks: (1) One or more nuclear localization signaling peptides; (2) A functional block with single-base editing capability, which consists of the following structure: 1) TadA-8e polypeptide, 2) A complete Cas9 or Cas9n polypeptide, and 3) Variants of the HDG4 peptide; And optional 4), linking fragments, used to link individual peptides, peptide fragments or peptide variants; The sequence of the TadA-8e polypeptide is shown in SEQ ID NO.2; The variant of the HDG4 polypeptide is a further mutation performed on the wild-type mouse 3-methyladenine glycosidase protein shown in SEQ ID NO.5, as follows: R165E and Y179F.
2. The single-base gene editing system according to claim 1, characterized in that, The nuclear localization signal peptide is located at the N-terminus and / or C-terminus of the fusion protein with single-base editing function.
3. The single-base gene editing system according to claim 2, characterized in that, The sequence of the nuclear localization signal peptide is shown in SEQ ID NO.
1.
4. The single-base gene editing system according to claim 1, characterized in that, The polypeptide sequences of the linking fragments are as shown in any one of SEQ ID NO.3, 10-21, and the sequences of the linking fragments linking the various polypeptides may be the same or different.
5. The single-base gene editing system according to any one of claims 1-4, characterized in that, The functional block with single-base editing capability is composed of the following structure: 1) The TadA-8e polypeptide with the sequence shown in SEQ ID NO.2; 2) A complete Cas9 polypeptide; 3) Variants of the HDG4 peptide; And 4), optionally 0-2 linking fragments for linking the polypeptide or polypeptide variant; The complete Cas9 polypeptide is a D10A type spCas9n polypeptide, which is the 2-1368 polypeptide sequence of the Saccharomyces cerevisiae D10A type spCas9n polypeptide sequence shown in SEQ ID NO.4; The HDG4 variant is derived by further mutation of the wild-type mouse Aag peptide shown in SEQ ID NO.5 as follows: R165E&Y179F.
6. The single-base gene editing system according to claim 5, characterized in that, The polypeptide sequences of the linked fragments are all as shown in SEQ ID NO.
3.
7. A single-base gene editing system, characterized in that, The single-base gene editing system described above can induce A-to-C directed mutations or T-to-G directed mutations in the target DNA; The single-base gene editing system comprises gRNA, and the main active region of the single-base gene editing system is the 5th to 7th bases at the 5' end of the gRNA; The gene editing system further comprises a fusion protein with single-base editing function, the fusion protein having the single-base editing function having the following structure: 1) The first polypeptide fragment derived from the D10A type spCas9n polypeptide; 2) Variants of the HDG4 peptide; 3) The TadA-8e polypeptide with the sequence shown in SEQ ID NO.2; 4) A second polypeptide fragment derived from the D10A type spCas9n polypeptide; And 5), optionally 0-3 linking fragments for linking the polypeptide, polypeptide fragment or polypeptide variant; The first polypeptide fragment derived from the D10A type spCas9n polypeptide and the second polypeptide fragment derived from the D10A type spCas9n polypeptide are from different regions in the Saccharomyces cerevisiae D10A type spCas9n polypeptide sequence shown in SEQ ID NO.4; The first polypeptide fragment derived from the D10A type spCas9n polypeptide is polypeptide 2-1248 in the Saccharomyces cerevisiae D10A type spCas9n polypeptide sequence shown in SEQ ID NO.4; The second polypeptide fragment derived from the D10A type spCas9n polypeptide is polypeptide 1249-1368 in the Saccharomyces cerevisiae D10A type spCas9n polypeptide sequence shown in SEQ ID NO.4; The variant of the HDG4 peptide is: a further mutation of the wild-type mouse Aag peptide shown in SEQ ID NO.5, namely R165E&Y179F.
8. The single-base gene editing system according to claim 7, characterized in that, The polypeptide sequences of the linked fragments are all as shown in SEQ ID NO.
3.
9. The single-base gene editing system according to claim 7, characterized in that, In the functional block with single-base editing capability, the TadA-8e polypeptide sequence shown in SEQ ID NO.2 is replaced with a TadA-8e polypeptide variant with the sequence shown in SEQ ID NO.
9.
10. The single-base gene editing system according to claim 5, characterized in that, The fusion protein with single-base editing function consists of the following structures arranged sequentially from the N-terminus to the C-terminus: (1) A nuclear localization signal polypeptide with a sequence as shown in SEQ ID NO.1; (2) A TadA-8e polypeptide with the sequence shown in SEQ ID NO.2; (3) A linked polypeptide with a sequence as shown in SEQ ID NO.3; (4) A *Saccharomyces cerevisiae* D10A type spCas9n polypeptide with a sequence as shown in SEQ ID NO.4, 2-1368; (5) Another linked polypeptide with a sequence as shown in SEQ ID NO.3; (6) Variants of the HDG4 peptide; (7) Another nuclear localization signal polypeptide with a sequence as shown in SEQ ID NO.1; The variant of the HDG4 peptide is: a double mutation of R165E & Y179F on the wild-type mouse Aag peptide shown in SEQ ID NO.
5.
11. The single-base gene editing system according to claim 7, characterized in that, The fusion protein with single-base editing function consists of the following structures arranged sequentially from the N-terminus to the C-terminus: (1) A nuclear localization signal polypeptide with a sequence as shown in SEQ ID NO.1; (2) A first Cas9n polypeptide fragment with a sequence as shown in SEQ ID NO.4, 2-1248; (3) A linked polypeptide with a sequence as shown in SEQ ID NO.3; (4) A variant of the HDG4 peptide; (5) A TadA-8e polypeptide with the sequence shown in SEQ ID NO.2; (6) Another linked polypeptide with a sequence as shown in SEQ ID NO.3; (7) A second Cas9n polypeptide fragment with a sequence as shown in SEQ ID NO.4, numbers 1249-1368; (8) Another nuclear localization signal polypeptide with a sequence as shown in SEQ ID NO.1; The variant of the HDG4 peptide is: a double mutation of R165E & Y179F on the wild-type mouse Aag peptide shown in SEQ ID NO.
5.
12. The single-base gene editing system according to claim 9, characterized in that, The fusion protein with single-base editing function consists of the following structures arranged sequentially from the N-terminus to the C-terminus: (1) A nuclear localization signal polypeptide with a sequence as shown in SEQ ID NO.1; (2) A first Cas9n polypeptide fragment with a sequence as shown in SEQ ID NO.4, 2-1248; (3) A linked polypeptide with a sequence as shown in SEQ ID NO.3; (4) A variant of the HDG4 peptide; (5) A TadA-8e polypeptide variant with the sequence shown in SEQ ID NO.9; (6) Another linked polypeptide with a sequence as shown in SEQ ID NO.3; (7) A second Cas9n polypeptide fragment with a sequence as shown in SEQ ID NO.4, numbers 1249-1368; (8) Another nuclear localization signal polypeptide with a sequence as shown in SEQ ID NO.1; The variant of the HDG4 peptide is: a double mutation of R165E & Y179F on the wild-type mouse Aag peptide shown in SEQ ID NO.
5.
13. The single-base gene editing system according to any one of claims 1-4, characterized in that, When editing the target gene, the efficiency of directed mutations from A to C on the sense strand is greater than 20% and the bystander effect is less than 20%.
14. The single-base gene editing system according to any one of claims 1-4, characterized in that, When editing the target gene, the efficiency of directed mutations from A to C on the sense strand is greater than 30% and the bystander effect is less than 15%.
15. The single-base gene editing system according to any one of claims 1-4, characterized in that, The efficiency of directed mutations from A to C in the justice chain is greater than 40%, and the bystander effect is less than 10%.
16. A nucleotide fragment encoding a fusion protein having single-base editing function as described in any one of claims 1-12.
17. A vector or host containing the nucleotide fragment of claim 16.