Guide editing tools, fusion rnas and uses thereof
By introducing random sequences and fusion proteins into the 3' end of pegRNA, the problems of low efficiency and off-target effects of guided editing tools were solved, achieving efficient and safe gene editing results.
Patent Information
- Application Number
- CN202110361688.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-02
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2041-04-02
AI Technical Summary
Existing guided editing tools are inefficient, limiting their application in gene editing, and are prone to off-target effects.
By introducing a random sequence at the 3' end of the pegRNA, and combining the Csy4 endonuclease and Cas9n in the fusion protein with the M-MLV reverse transcriptase, the complementary base pairing of the pegRNA's first and last bases is avoided, thereby improving editing efficiency and maintaining safety.
It significantly improves the efficiency of guided editing while reducing off-target effects, and has good prospects for industrialization.
Smart Images

Figure BDA0003005848320000051 
Figure BDA0003005848320000061 
Figure BDA0003005848320000121
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biotechnology and relates to a guided editing tool, fused RNA, and their uses. Background Technology
[0002] The CRISPR / Cas9 system has been widely used in genetic manipulation [Cong, L, et al., Science (New York, NY) 339:819-823; Shen, B, et al., Cell Res 23:720-723.], and was awarded the 2020 Nobel Prize in Chemistry for its enormous impact. Base editing (BE) technology based on the CRISPR / Cas9 system can perform single-base level operations on the genome [Gaudelli, NM, et al., Nature 551:464-471; Komor, AC, et al., Nature 533:420-424.]. Compared with the traditional method using the homology-directed repair (HDR) pathway after Cas9 cleavage, it is significantly more efficient and has been verified to be highly efficient and accurate in plants, animals, and human embryos [Zeng, Y, et al., Mol Ther 26:2631-2637; Li, J, et al., Cell Res 29:174-176; Zong, Y, et al., Nat Biotechnol 35:438-440.]. It has also achieved the repair of pathogenic mutations in human embryos [Zeng, Y, et al., Mol Ther 26:2631-2637; Li, J, et al., Cell Res 29:174-176; Zong, Y, et al., Nat Biotechnol 35:438-440.]. [26:2631-2637], gene therapy was performed in mouse disease models, showing strong promise for gene therapy [Koblan, LW, et al., Nature 589:608-614.].
[0003] However, due to significant off-target effects on DNA and RNA in gene editing (BE) [Grunewald, J, et al., Nature 37:1041-1048; Jin, S, et al., Science (New York, NY) 364:292-295.], and because BE can only target C→T and A→G point mutations, its application is clearly limited, necessitating the development of more powerful gene editing tools. Prime editing (PE), a guided editing technology reported at the end of 2019, can target all mutations, including all point mutation types, and allows for precise insertion and deletion. Therefore, it is highly anticipated as a potential replacement for BE as the next-generation point mutation tool [Anzalone, AV, et al., Nature 576:149-157.].
[0004] PE is essentially an extension of ssDNA through point mutation. Its basic principle involves forming a fusion protein between the Moroni mouse leukemia virus reverse transcriptase M-MLV and the H840A mutant Cas9n, and extending the 3' end of the commonly used sgRNA to form PEgRNA (pegRNA). The extended sequence contains the binding primer (PBS) required by the reverse transcriptase and the template (RT template) needed for repair. The reverse transcriptase reverse transcribes the PBS and RT to obtain the repaired DNA, which can then be used for site-directed mutagenesis. This allows for all types of mutations and precise insertion and deletion of sequences, greatly expanding the scope of gene editing [Anzalone, AV, et al., Nature 576:149-157.].
[0005] Guided editing technology, since its initial report in late 2019, has been applied to plants and animals [Liu, Y, et al., Cell Discov 6:27; Lin, Q, et al., Nat Biotechnol 38:582-585.], demonstrating its feasibility in vector editing. However, the efficiency of guided editing has long been low, limiting its application. Therefore, optimizing and improving guided editing is currently a key area of research. Summary of the Invention
[0006] To address the shortcomings of low efficiency in existing guided editing tools, this invention provides a guided editing tool, fused RNA, and its uses.
[0007] Through extensive exploratory research, the inventors discovered that pegRNA exhibits a complementary pairing of start and end bases in its sequence (e.g. Figure 1 As shown, this may lead to a reduction in the effective expression of pegRNA, thereby affecting the active expression of PE. Adding a random sequence to the 3' end of the pegRNA can reduce potential first- and last-tail complementary base pairing, thereby improving PE activity. Furthermore, it does not affect off-target effects while improving PE editing efficiency, ensuring the safety of PE.
[0008] To address the aforementioned technical problems, the first aspect of this invention provides a guided editing tool, comprising:
[0009] (i) A fusion protein comprising at least one gene editor and a nuclease;
[0010] (ii) A fusion RNA comprising a pegRNA and a recognition site of the endonuclease described in (i);
[0011] The fusion protein has reverse transcription function and can bind to the recognition site and cleave the recognition site, thereby introducing a sequence at the 3' end of the pegRNA and preventing the pegRNA from self-circulating.
[0012] In a preferred embodiment, the fusion RNA consists of pegRNA, a Csy4 endonuclease recognition sequence, and a nick sgRNA from the 5' end to the 3' end; preferably, the nucleotide sequence of the Csy4 endonuclease recognition sequence is shown in SEQ ID NO:5.
[0013] In a preferred embodiment, the fusion protein comprises, for example, a Csy4 endonuclease, Cas9n, and a viral reverse transcriptase, such as Moroni mouse leukemia virus reverse transcriptase M-MLV, sequentially from N-terminus to C-terminus. The fusion protein fuses the Csy4 endonuclease to the N-terminus of the guide editor, enabling guided editing at the target site under the guidance of the fusion RNA, thus effectively improving the editing efficiency of PE.
[0014] Preferably, the amino acid sequence of the Csy4 endonuclease is shown in SEQ ID NO:1, the amino acid sequence of the Cas9n is shown in SEQ ID NO:2, and / or the amino acid sequence of the M-MLV is shown in SEQ ID NO:3.
[0015] In the fusion protein provided by the present invention, the amino acid sequence of the Csy4 endonuclease may include: the amino acid sequence shown in SEQ ID NO:1; or an amino acid sequence having more than 80% sequence similarity to SEQ ID NO:1 and having the function of the amino acid sequence defined in SEQ ID NO:1. Specifically, the amino acid sequence refers to a polypeptide fragment that has the function of the polypeptide fragment shown in SEQ ID NO:1, obtained by substituting, deleting, or adding one or more amino acids (specifically, 1-50, 1-30, 1-20, 1-10, 1-5, 1-3, 1, 2, or 3) of the amino acid sequence shown in SEQ ID NO:1, or obtained by adding one or more amino acids (specifically, 1-50, 1-30, 1-20, 1-10, 1-5, 1-3, 1, 2, or 3) to the N-terminus and / or C-terminus, and has the function of the polypeptide fragment shown in SEQ ID NO:1. For example, it can be a Csy4 endonuclease that still has the targeting activity of the Csy4 endonuclease recognition sequence after mutation, or more specifically, it can be an activity that can target RNA under the guidance of a special targeting sequence to form two independent truncated RNA parts. The amino acid sequence described herein may have 80%, 85%, 90%, 93%, 95%, 97%, or 99% or more similarity to SEQ ID NO:1. The Csy4 endonuclease fragment is typically derived from Pseudomonas aeruginosa.
[0016] In the fusion protein provided by this invention, the amino acid sequence of the second Cas9n fragment may include: the amino acid sequence shown in SEQ ID NO:2; or an amino acid sequence having more than 80% sequence similarity to SEQ ID NO:2 and having the function of the defined amino acid sequence. Specifically, the amino acid sequence refers to: a polypeptide fragment obtained by substituting, deleting, or adding one or more (specifically, 1-50, 1-30, 1-20, 1-10, 1-5, 1-3, 1, 2, or 3) amino acids to the amino acid sequence shown in SEQ ID NO:2, or a polypeptide fragment obtained by adding one or more (specifically, 1-50, 1-30, 1-20, 1-10, 1-5, 1-3, 1, 2, or 3) amino acids to the N-terminus and / or C-terminus, and having the function of the polypeptide fragment shown in SEQ ID NO:2. For example, it may be a polypeptide fragment that still has the targeting activity of Cas9n after mutation, or more specifically, it may be an polypeptide fragment that can target RNA under the guidance of a suitable gRNA. The amino acid sequence described herein may have 80%, 85%, 90%, 93%, 95%, 97%, or 99% similarity to SEQ ID NO:2. The Cas9n fragment is typically derived from Streptococcus pyogenes.
[0017] The fusion protein provided by the present invention may include the amino acid sequence of the M-MLV fragment as shown in SEQ ID NO:3; or an amino acid sequence having more than 80% sequence similarity to SEQ ID NO:3 and having the function of the defined amino acid sequence. Specifically, the amino acid sequence referred to herein refers to a polypeptide fragment that has the function of the polypeptide fragment shown in SEQ ID NO:3, obtained by substitution, deletion, or addition of one or more amino acids (specifically, 1-50, 1-30, 1-20, 1-10, 1-5, 1-3, 1, 2, or 3) of the amino acid sequence shown in SEQ ID NO:3, or obtained by adding one or more amino acids (specifically, 1-50, 1-30, 1-20, 1-10, 1-5, 1-3, 1, 2, or 3) to the N-terminus and / or C-terminus, and has reverse transcription activity, more specifically, the function of reverse transcribing single-stranded RNA (ssRNA) into single-stranded DNA (ssDNA) under the guidance of primers. The amino acid sequence in f) may have 80%, 85%, 90%, 93%, 95%, 97%, or 99% or more similarity to SEQ ID NO:3. The M-MLV fragment is typically derived from mice (Mus musculus). The final fusion protein sequence is shown in SEQ ID NO:4.
[0018] In the fusion protein provided by this invention, the substitution, deletion, or addition can be a conserved amino acid substitution. Specifically, "conserved amino acid substitution" refers to the substitution of an amino acid residue by another amino acid residue with a similar side chain. Families of amino acid residues with similar side chains should be known to those skilled in the art, and may include, but are not limited to, basic side chains (e.g., lysine, arginine, histidine), acidic side chains (e.g., aspartic acid, glutamic acid), uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine), nonpolar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan, isoleucine), and aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine). More specifically, conserved amino acid substitutions may include, but are not limited to, the specific cases listed in the table below. The numbers in Table 1 (amino acid similarity matrix) represent the similarity between two amino acids; a number greater than or equal to 0 is considered a conserved amino acid substitution. Table 2 shows exemplary conserved amino acid substitution schemes.
[0019] Table 1
[0020] C G P S A T D E N Q H K R V M I L F Y W W -8 -7 -6 -2 -6 -5 -7 -7 -4 -5 -3 -3 2 -6 -4 -5 -2 0 0 17 Y 0 -5 -5 -3 -3 -3 -4 -4 -2 -4 0 -4 -5 -2 -2 -1 -1 7 10 F -4 -5 -5 -3 -4 -3 -6 -5 -4 -5 -2 -5 -4 -1 0 1 2 9 L -6 -4 -3 -3 -2 -2 -4 -3 -3 -2 -2 -3 -3 2 4 2 6 I -2 -3 -2 -1 -1 0 -2 -2 -2 -2 -2 -2 -2 4 2 5 M -5 -3 -2 -2 -1 -1 -3 -2 0 -1 -2 0 0 2 6 V -2 -1 -1 -1 0 0 -2 -2 -2 -2 -2 -2 -2 4 R -4 -3 0 0 -2 -1 -1 -1 0 1 2 3 6 K -5 -2 -1 0 -1 0 0 0 1 1 0 5 H -3 -2 0 -1 -1 -1 1 1 2 3 6 Q -5 -1 0 -1 0 -1 2 2 1 4 N -4 0 -1 1 0 0 2 1 2 E -5 0 -1 0 0 0 3 4 D -5 1 -1 0 0 0 4 T -2 0 0 1 1 3 A -2 1 1 1 2 S 0 1 1 1 P -3 -1 6 G -3 5 C 12
[0021] Table 2
[0022]
[0023]
[0024] More preferably, the fusion protein further includes a T2A fragment and / or a BPNLS fragment.
[0025] More preferably, the T2A fragment is located between the Csy4 endonuclease and Cas9n, and its amino acid sequence is shown in SEQ ID NO:6; and / or, the BPNLS fragment is located at the C-terminus, and its amino acid sequence is shown in SEQ ID NO:7.
[0026] In a preferred embodiment, the recognition sequence of the Csy4 endonuclease contained in the fusion RNA is a nucleotide sequence as shown in SEQ ID NO:5, or has more than 95% identity with the nucleotide sequence shown in SEQ ID NO:5 and maintains the function of being recognized by the Csy4 endonuclease.
[0027] The fusion RNA provided by this invention includes a DNA sequence for the Csy4 endonuclease-recognized sequence fragment, which may include: a DNA sequence as shown in SEQ ID NO:5; or a DNA sequence having more than 95% sequence similarity to SEQ ID NO:5 and having the function of the defined DNA sequence. Specifically, the DNA sequence refers to a DNA fragment obtained by substituting, deleting, or adding one or more (1, 2, or 3) DNA molecules to the DNA sequence shown in SEQ ID NO:5, or by adding one or more (specifically, 1, 2, or 3) DNA molecules to the 5'-end and / or 3'-end, and having the function of the DNA fragment shown in SEQ ID NO:5. For example, it may be an activity recognized by the Csy4 endonuclease, more specifically, a function of being recognized by the Csy4 endonuclease in its presence and cleaving within the recognition sequence. The DNA sequence may have more than 95% similarity to SEQ ID NO:5.
[0028] In the fusion RNA provided by this invention, the substitution, deletion, or addition can be RNA substitution. Specifically, "RNA substitution" can refer to RNA mutations that do not affect the recognition function of the Csy4 endonuclease.
[0029] Preferably, the amino acid sequence of the fusion protein is as shown in SEQ ID NO:4, or has 90%, 95%, 96%, 97%, 98%, 99% or more of the same amino acid sequence as SEQ ID NO:4, and has the function of the fusion protein as shown in the amino acid sequence of SEQ ID NO:4.
[0030] To address the aforementioned technical problems, a second aspect of the present invention provides a fusion RNA, wherein the fusion RNA comprises, from the 5' end to the 3' end, pegRNA, a Csy4 endonuclease recognition sequence, and nicking sgRNA, respectively.
[0031] Preferably, the fusion RNA contains a Csy4 endonuclease recognition sequence as shown in SEQ ID NO:5, or has 95% identity with the nucleotide sequence shown in SEQ ID NO:5 and maintains the function of being recognized by the Csy4 endonuclease.
[0032] To address the aforementioned technical problems, a third aspect of the present invention provides a fusion protein, wherein the fusion protein comprises, from the N-terminus to the C-terminus, Csy4 endonuclease, Cas9n, and Moroni mouse leukemia virus reverse transcriptase M-MLV.
[0033] Preferably, the amino acid sequence of the fusion protein is as shown in SEQ ID NO:4, or has 90%, 95%, 96%, 97%, 98%, 99% or more of the same amino acid sequence as SEQ ID NO:4, and has the function of the fusion protein as shown in the amino acid sequence of SEQ ID NO:4.
[0034] To address the aforementioned technical problems, a fourth aspect of the present invention provides an isolated nucleic acid, wherein the isolated nucleic acid comprises a first polynucleotide encoding a fusion protein as described in the third aspect of the present invention; and / or a second polynucleotide transcribed as a fusion RNA as described in the second aspect of the present invention.
[0035] To address the aforementioned technical problems, a fifth aspect of the present invention provides a recombinant expression vector comprising the isolated nucleic acid as described in the fourth aspect of the present invention.
[0036] To address the aforementioned technical problems, a sixth aspect of the present invention provides an expression system comprising the recombinant expression vector as described in the fifth aspect of the present invention.
[0037] The expression system can be a host cell, which can express the fusion protein as described above. The fusion protein can cooperate with the fusion RNA, thereby locating the fusion protein to a target region and achieving guided editing of the target region. In another specific embodiment of the present invention, the host cell of the expression system is selected from eukaryotic or prokaryotic cells, preferably from mouse cells or human cells, more preferably from mouse neuroma cells, human embryonic kidney cells, or human cervical cancer cells, human colon cancer cells, or human osteosarcoma cells, and even more preferably from N2a cells, HEK293T cells, HeLa cells, HCT116 cells, or U2OS cells. The fusion RNA and the fusion protein can be expressed in the same host cell or in different host cells, and the host cell can be the target cell.
[0038] Preferably, in the expression system, the first polynucleotide and the second polynucleotide may be located in the same recombinant expression vector or different recombinant expression vectors, such as pCMV, pCAG or Tet-On.
[0039] To address the aforementioned technical problems, the seventh aspect of this invention provides the use of the guided editing tool as described in the first aspect of this invention, the fusion RNA as described in the second aspect of this invention, the fusion protein as described in the third aspect of this invention, the isolated nucleic acid as described in the fourth aspect of this invention, or the expression system as described in the fifth aspect of this invention in eukaryotic gene editing.
[0040] The eukaryote can specifically be a metazoan, including but not limited to humans and mice. The applications can specifically include, but are not limited to, point mutations, fragment insertions, and deletions. These guided edits can be used to edit splice acceptor / donor sites to regulate RNA splicing, or for constructing models (e.g., disease models, cell models, animal models, etc.) or treating human diseases. In one specific embodiment of the invention, the object being edited can be an embryo, cell, etc. In another specific embodiment of the invention, the gene editing is in vitro gene editing.
[0041] Preferably, the use includes base substitution, insertion, or deletion.
[0042] To solve the above-mentioned technical problems, the eighth aspect of the present invention provides a method for preparing a guided editing tool as described in the first aspect of the present invention, characterized in that it includes the following steps: obtaining the fusion protein and fusion RNA respectively using the expression system described in the sixth aspect of the present invention.
[0043] To address the aforementioned technical problems, a ninth aspect of the present invention provides a method for guided editing, characterized in that the method includes gene editing using a guided editing tool as described in the first aspect of the present invention.
[0044] Existing guided editing systems include PE, pegRNA, and nicking sgRNA. Those skilled in the art can select appropriate pegRNA and nicking sgRNA targeting specific sites based on the target editing region of the gene. For example, the sequence of the pegRNA is typically at least partially complementary to the target region, thereby enabling it to cooperate with the PE and locate itself to the target region, achieving guided editing within the target region, including all types of point mutations, such as C·G-to-A·T, G·C-to-C·G, A·T-to-C·G, and T·A-to-A·T. However, this guided editing system is inefficient; the guided editing tool (ePE) provided in the first aspect of this invention overcomes these shortcomings.
[0045] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention.
[0046] Before further describing specific embodiments of the present invention, it should be understood that the scope of protection of the present invention is not limited to the specific embodiments described below; it should also be understood that the terminology used in the embodiments of the present invention is for describing specific embodiments and not for limiting the scope of protection of the present invention; in the present invention, unless otherwise expressly stated in the text, the singular forms "a", "an" and "this" include the plural forms.
[0047] When numerical ranges are given in the embodiments, it should be understood that, unless otherwise stated in the present invention, both endpoints of each numerical range and any value between the two endpoints may be selected. Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art. In addition to the specific methods, apparatus, and materials used in the embodiments, based on the knowledge of the prior art possessed by one of ordinary skill in the art and the description of this invention, any prior art methods, apparatus, and materials similar to or equivalent to those described, apparatus, and materials in the embodiments of this invention may be used to implement the present invention.
[0048] Unless otherwise stated, the experimental methods, detection methods, and preparation methods disclosed in this invention all employ conventional techniques in molecular biology, biochemistry, chromatin structure and analysis, analytical chemistry, cell culture, recombinant DNA technology, and related fields. These techniques have been well described in existing literature; see Sambrook et al., *MOLECULAR CLONING: ALABORATORY MANUAL*, Second edition, Cold Spring Harbor Laboratory Press, 1989 and Third edition, 2001; Ausubel et al., *CURRENT PROTOCOLS IN MOLECULAR BIOLOGY*, John Wiley & Sons, New York, 1987 and periodic updates; *theseries METHODS IN ENZYMOLOGY*, Academic Press, San Diego; Wolffe, *CHROMATINSTRUCTURE AND FUNCTION*, Third edition, Academic Press, San Diego, 1998; *METHODS IN ENZYMOLOGY*, Vol. 304, Chromatin (PM Wassarman and AP Wolffe, eds.), Academic Press, San Diego, 1999; and *METHODS IN MOLECULAR*. BIOLOGY, Vol. 119, Chromatin Protocols (PB Becker, ed.) Humana Press, Totowa, 1999, etc.
[0049] Based on common knowledge in the field, the above-mentioned preferred conditions can be combined arbitrarily to obtain various preferred embodiments of the present invention.
[0050] The reagents and raw materials used in this invention are all commercially available.
[0051] The positive and progressive effects of this invention are as follows:
[0052] This invention provides a novel guided editing tool (ePE) that uses a Csy4 endonuclease embedded in Cas9n. Compared to traditional PE editors, the Csy4 endonuclease leaves a residual sequence after cutting the recognition sequence, preventing complementary base pairing of the pegRNA's own start and end, significantly improving the efficiency of PE editing without producing off-target effects, and showing good prospects for industrialization. Figure 2 ). Attached Figure Description
[0053] Figure 1 In traditional guided editing systems, pegRNAs form a head-to-tail connection.
[0054] Figure 2 This invention provides an improved guided editing system.
[0055] Figure 3 The guided editing system provided by this invention has a significantly higher base substitution efficiency in HEK293 cells than the traditional method.
[0056] Figure 4 The off-target rate of the guided editing system provided by this invention is not significantly different from that of traditional methods.
[0057] Figure 5 The guided editing system provided by this invention has a significantly higher base substitution efficiency in HeLa cells than the traditional method.
[0058] Figure 6 The guided editing system provided by this invention has a significantly higher base substitution efficiency in mouse N2a cells than the traditional method. Detailed Implementation
[0059] The present invention is further illustrated below by way of embodiments, but the invention is not limited to the scope of the embodiments described herein. Experimental methods in the following embodiments that do not specify specific conditions were performed according to conventional methods and conditions, or as selected according to the product instructions.
[0060] Example 1: Construction of fusion proteins in editing tools
[0061] 1. Construction of a Csy4-based guided editing tool
[0062] The Csy4 endonuclease sequence (SEQ ID NO:1) was synthesized by Genscript Biotech Inc., and PCR amplification was performed using a high-fidelity enzyme kit (Vazyme, P501-d2) from Nanjing Novizan Biotechnology Co., Ltd. The forward primer was SEQ ID NO:8: ATGGACCACTACCTCGACATTC, and the reverse primer was SEQ ID NO:9: GAACCAGGGAACGAAACCTCC.
[0063] The amplification system is shown in Table 3 below:
[0064] Table 3
[0065] water Add water to 50μL 2xbuffer 25μL dNTP 1μL Forward primer (10 μM) 2μL Reverse primer (10 μM) 2μL Synthesis of Csy4 endonuclease template 1ng High-fidelity enzymes 1μL
[0066] PCR conditions are shown in Table 4 below:
[0067] Table 4
[0068]
[0069] The PCR amplification products were purified and recovered using the AxyPrep PCR Clean-up Kit (Axygen, AP-PCR-500G) and were ready for use.
[0070] 2. Construction of pCMV-Csy4-NMRT, a next-generation guided editing tool containing the Csy4 endonuclease.
[0071] The Csy4 product obtained in step 1 was used to construct a vector. PCR amplification was performed using a high-fidelity enzyme kit (Vazyme, P501-d2) from Nanjing Novizan Biotechnology Co., Ltd. The forward primer was SEQ ID NO:10 (GTCAGATCCGCTAGAGATCC GCGGCCGCTAATAC GACTCACTATAGGATGGACCACTACCTCGACATT), and the reverse primer was SEQ ID NO:11 (GACGTCACCGCATGTTAACAGACTTCCTCTGCCCTCGAACCAGGGAACGAAACCTCCTT).
[0072] The PE2 vector was amplified. PCR amplification was performed using a high-fidelity enzyme kit (Vazyme, P501-d2) from Nanjing Novizan Biotechnology Co., Ltd. The forward primer was SEQ ID NO:12 (TGTTAACATGCGGTGACGTCGAGGAGAATCCTGGCCCACCAAAGAAGAAGCGGAAAGTC), and the reverse primer was SEQ ID NO:13 (TGCCGGCCCATCACTTTCAC).
[0073] The amplification system is shown in Table 5 below:
[0074] Table 5
[0075]
[0076]
[0077] PCR conditions are shown in Table 6 below:
[0078] Table 6
[0079]
[0080] The PCR amplification products were purified and recovered using the AxyPrep PCR Clean-up Kit (Axygen, AP-PCR-500G) and were ready for use.
[0081] The pCMV-PE2 (Addgene#132775) plasmid was digested with NotI-HF (NEB, R3189S) and SacI-HF (NEB, R3156S) to obtain a linearized sgRNA vector. The digestion system is shown in Table 7 below:
[0082] Table 7
[0083] water Add water to 50μL pCMV-PE2 5μg 10×cutsmart buffer 5μL NotI-HF enzyme 3μL SacI-HF 3μL
[0084] After the above reaction system was prepared, it was incubated at 37℃ for 5 hours. The enzyme digestion products were then extracted using an AxyPrep DNA gel extraction kit (Axygen, AP-GX-250G) to obtain the linearized vector. 100 ng of the linearized vector and the PCR product fragment were recombined using a recombinase kit (Vazyme, C112) from Nanjing Novizan Biotechnology Co., Ltd. The mixture was incubated at 37℃ for 30 minutes and then transformed into a plate. Sanger sequencing yielded the correct pCMV-Csy4-NMRT vector. The ligation system is shown in Table 8 below.
[0085] Table 8
[0086] water Add water to 20μL 5xbuffer 2μL Segment 1 150ng Segment 2 150ng Linearized pCMV-PE2 100ng Recombinase 1μL
[0087] Example 2: Construction of fusion RNA in editing tools
[0088] The fusion RNA used to detect the targeting and editing efficiency of ePE (Enhanced Prime Editing) in eukaryotic cells was site1. Subsequent detection of ePE at six endogenous gene sites in HEK293T cells revealed fusion RNAs at site1, FBN1, RIT1, RNF2, ALDOB, and MSH2. Subsequent detection of ePE at thirteen endogenous gene sites in N2a cells revealed fusion RNAs at Dnmt1, Fgf21, Ifnar1, Trem2, Rnf2, Tyr, Fgf5, Mstn, Cftr, Hoxd13, SITE3, Ar, and SITE4. The sequence of the Csy4 endonuclease recognition site is shown in SEQ ID NO:5. 20nt spacer primers for pegRNA were designed based on the target site sequences. The upstream primer had ACCG added to the 5' end and GTTTC added to the 3' end, while the downstream primer had CTCTGAAAC added to the 5' end. The PBS and RT sequences of pegRNA and the 20nt spacer sequence of the cleavage sgRNA were designed based on the target site sequence. The PBS, RT, Csy4 protein recognition, and cleavage sgRNA spacer sequences were synthesized onto the same pair of oligonucleotide primers. The upstream primer had GTGC added to its 5' end, and the downstream primer had AAAC added to its 5' end. All primers were synthesized.
[0089] Dissolve in sterile water to a final concentration of 100 μM. The oligonucleotide primer for synthesizing scaffold, scaffold-F, is: agagctagaaatagcaagttgaaataaggctagtccgttatcaacttgaaaaagtggcaccgagtcg (SEQ ID NO:14).
[0090] scaffold-R: gcaccgactcggtgccactttttcaagttgataacggactagccttatttcaacttgctatttctag (SEQ ID NO: 15)
[0091] The synthesized primers were annealed, and the annealing system is shown in Table 9 below:
[0092] Table 9
[0093] forward primer 4.5μL reverse primer 4.5μL 10×NEB buffer2 1μL
[0094] The annealing procedure is shown in Table 10 below:
[0095] Table 10
[0096] 95℃ 5min 95-85℃ -2℃ / s 85-25℃ -0.1℃ / s 4℃ ∞
[0097] The annealed scaffold sequence needs to be phosphorylated. The phosphorylation system is shown in Table 11 below:
[0098] Table 11
[0099] water Add water to 25μL scaffold annealed products 6.25μL 10x T4 DNA ligase buffer (NEB) 2.50μL T4 PNK(NEB) 0.50μL
[0100] Using the pGL3-U6-sgRNA-EGFP (Addgene#107721) plasmid as a template, linearized vector fragments were amplified using primers Csy4peg-bone-F (GAGAGGGTCTCAGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATC, SEQ ID NO:16) and Csy4peg-bone-R (CTCTCGGTCTCACGGTGTTTCGTCCTTTCCAC, SEQ ID NO:17). The linearized vector was then recovered by gel extraction using the AxyPrep DNA gel extraction kit (Axygen, AP-GX-250G). The linearized vector was digested with BsaI (NEB, R0535S) to obtain a fusion RNA vector backbone with sticky ends. The digestion system is shown in Table 12 below.
[0101] Table 12
[0102] water Add water to 30μL Linearized carrier 2μg 10×cutsmart buffer 3μL BsaI enzyme 1μL
[0103] The annealing product was ligated to the fusion RNA vector backbone to construct a targeted fusion RNA. The ligation system is shown in Table 13 below:
[0104] Table 13
[0105] water Add water to 10μL Fusion RNA vector backbone 30ng Annealed product 1 1μL Annealed product 2 1μL Phosphorylated scaffold 1μL Solution I 5μL
[0106] The ligation product was then transformed, allowed to recover for 30 min, plated on aminobenzene-resistant LB agar plates, and incubated overnight at 37°C. Single clones were selected for sequencing verification, yielding correctly sequenced fusion RNA.
[0107] Example 3: Application of Guided Editing Tools in Eukaryotic Cells
[0108] The guided editing tool (ePE) of this invention includes the fusion protein constructed in Example 1 and the fusion RNA constructed in Example 2.
[0109] 1. Targeted editing in human HEK293T cells
[0110] After screening for functional ePEs in prokaryotic cells, we further examined the targeting-guided editing efficiency of ePEs in HEK293T cells, as follows:
[0111] HEK293T cells (from ATCC) were resuscitated and cultured in 10 cm culture dishes (Corning, 430167) in DMEM (HyClone, SH30243.01) containing 10% fetal bovine serum (HyClone, SV30087). The culture temperature was 37°C and the carbon dioxide concentration was 5%. After passage, when the cell density reached 80%, the cells were transferred to 24-well plates. Before use, the 24-well plates were coated with a 1:10 dilution of poly-L-lysine solution (Sigma, P4707-50 mL).
[0112] 1) Transfect cells 12-14 hours after seeding, when the cell concentration is approximately 80%. Transfect each well with 900 ng of pCMV-Csy4-NMRT plasmid, mixing it in 50 μL of Opti-MEM (Gibco, 11058021) medium. Use pCMV-PE2 as a positive control, adding 900 ng of pCMV-PE2 to each well.
[0113] 2) In addition, mix 3 μl of Lipofectamine 2000 transfection reagent (Thermo, 11668019) into 50 μl of Opti-MEM medium and let stand for 5 minutes.
[0114] 3) Add the plasmid-containing Opti-MEM to the Lipofectamine 2000-containing Opti-MEM, mix slowly by pipetting, and let stand for 20 minutes.
[0115] 4) Add the above-mixed and allowed-to-stand transfection solution to the cultured cells.
[0116] 5) Replace the medium with DMEM containing 10% FBS 6 hours after transfection.
[0117] 6) 48 hours after transfection, remove the culture medium, wash the cells once with PBS, then digest the cells with TE (Thermo Fisher, R001100), stop the digestion with DMEM containing 10% FBS, centrifuge to collect the cells, and finally resuspend them in culture medium.
[0118] 7) After resuspending, the cells were sorted by FACS (Fluorescence Activated Cell Sorting), and the cells with the highest GFP fluorescence intensity were collected. At least 10,000 cells were collected for each sample.
[0119] One-sixth of the collected cells were directly lysed, and the target site fragments were amplified by PCR. The PCR primer sequences are shown in SEQ ID NO:10. The genomic target site fragments were amplified by PCR using the high-fidelity enzyme kit (Vazyme, p501-d2) from Nanjing Novizan Biotechnology Co., Ltd. The PCR reaction system is shown in Table 14 below:
[0120] Table 14
[0121] water Add to 50 μL 2xbuffer 25μL dNTP 1μL Forward primer (10 μM) 2μL Reverse primer (10 μM) 2μL High-fidelity enzymes 1μL Cell lysate 3-5μL
[0122] The PCR procedure is shown in Table 15 below:
[0123] Table 15
[0124]
[0125] PCR amplification products were purified and recovered using the AxyPrep PCR Clean-up Kit (Axygen, AP-PCR-500G) and subjected to Sanger sequencing and high-throughput sequencing. All samples were sequenced using an Illumina HiSeq X 10 (2×150PE) at the Novogene Institute of Bioinformatics in Beijing, China, with a read depth of approximately 20 million per sample. Reads were mapped to the human reference genome (hg38) using STAR software (version 2.5.1) with annotations from GENCODE v30. After duplicate removal, variants were identified using GATK HaplotypeCaller (version 4.1.2) and filtered by QD (quality per depth). All variants were validated and quantified using bam-readcount with parameters -q 20-b 30. Given edits, a minimum 10-fold increase was required, and these edits were required to support at least 99% of the reads in the wild-type sample as reference alleles. Specific results are shown below. Figure 3 As shown. By Figure 3 It can be seen that ePE can significantly improve the efficiency of guided editing compared to PE (** indicates p<0.05; *** indicates p<0.01).
[0126] 2. Compare the off-target effects of PE and ePE in human cells.
[0127] 30,000 GFP-positive cells (5% PCR content) were collected and lysed. Genomic target site fragments were amplified by PCR using the high-fidelity enzyme kit (Vazyme, p501-d2) from Nanjing Novizan Biotechnology Co., Ltd. The PCR reaction system is shown in Table 16 below:
[0128] Table 16
[0129] water Add to 50 μL 2xbuffer 25μL dNTP 1μL Forward primer (10 μM) 2μL Reverse primer (10 μM) 2μL High-fidelity enzymes 1μL Cell lysate 3-5μL
[0130] The PCR procedure is shown in Table 17 below:
[0131] Table 17
[0132]
[0133]
[0134] PCR amplification products were purified and recovered using the AxyPrep PCR Clean-up kit (Axygen, AP-PCR-500G) and then subjected to high-throughput sequencing. The sequencing results are shown below. Figure 4 As shown, the results indicate that ePE does not produce additional off-target effects.
[0135] 3. ePE-guided editing results in more cell lines
[0136] The experiments described above have shown that ePE is more efficient than PE in guiding editing without affecting off-target effects. To further illustrate the improvement of guiding editing efficiency by ePE, we also conducted further experiments on human HeLa cell lines and mouse N2a cells, as follows:
[0137] 1) HeLa cells and N2a cells (from ATCC) were resuscitated and cultured separately in 10cm culture dishes (Corning, 430167) in DMEM (HyClone, SH30243.01) containing 10% fetal bovine serum (HyClone, SV30087). The culture temperature was 37℃ and the carbon dioxide concentration was 5%. After passage, when the cell density reached 80%, the cells were transferred to 24-well plates. Before use, the 24-well plates were coated with a 1:10 dilution of poly-L-lysine solution (Sigma, P4707-50mL).
[0138] 2) Transfect cells 12-14 hours after seeding, when the cell concentration is approximately 80%. The amount of plasmid transfected per well is 900 ng of pCMV-Csy4-NMRT plasmid and 300 ng of fusion RNA plasmid. Mix the plasmids in 50 μL of Opti-MEM (Gibco, 11058021) medium. Use pCMV-PE2 as a positive control; add 900 ng of pCMV-Csy4-NMRT, 300 ng of pegRNA plasmid, and 100 ng of nicking sgRNA to each well.
[0139] 3) In addition, mix 3 μl of Lipofectamine 2000 transfection reagent (Thermo, 11668019) into 50 μl of Opti-MEM medium and let stand for 5 minutes.
[0140] 4) Add the plasmid-containing Opti-MEM to the Lipofectamine 2000-containing Opti-MEM, mix slowly by pipetting, and let stand for 20 minutes.
[0141] 5) Add the above-mixed and allowed-to-stand transfection solution to the cultured cells.
[0142] 6) Change the medium with DMEM containing 10% FBS 6 hours after transfection. 48 hours after transfection, remove the medium, wash the cells once with PBS, then digest the cells with TE (Thermo Fisher, R001100), stop the digestion with DMEM containing 10% FBS, centrifuge to collect the cells, and finally resuspend them in culture medium.
[0143] 7) After resuspending, the cells were sorted by FACS (Fluorescence Activated Cell Sorting). Since the GFP signal is on the pegRNA plasmid or fusion RNA plasmid, we directly sorted all GFP-positive cells, collecting at least 10,000 cells per sample.
[0144] The collected cells were directly lysed, and the target site fragments were amplified by PCR. The PCR primer sequences are shown in SEQ ID NO: 11. The genomic target site fragments were amplified by PCR using the Novizan high-fidelity enzyme kit (Vazyme, p501-d2). The PCR reaction system is shown in Table 18 below:
[0145] Table 18
[0146] water Add to 50 μL 2xbuffer 25μL dNTP 1μL Forward primer (10 μM) 2μL Reverse primer (10 μM) 2μL High-fidelity enzymes 1μL Cell lysate 3-5μL
[0147] The PCR procedure is shown in Table 19 below:
[0148] Table 19
[0149]
[0150]
[0151] PCR amplification products were purified and recovered using the AxyPrep PCR Clean-up Kit (Axygen, AP-PCR-500G). PCR products with different barcodes were pooled and deep sequenced on an Illumina Hiseq X Ten (2×150PE) platform at the Novogene Institute of Bioinformatics in Beijing, China. Adapter pairs of paired end reads were removed using AdapterRemoval version 2.2.2, and paired end reads of 11 bp or more were merged into a single shared read. All processed reads were then mapped to the target sequence using the BWA-MEM algorithm (BWA v0.7.16). For each site, the mutation rate was calculated using BAM read counts with parameters -q 20-b 30. Insertion / deletion was calculated based on reads containing at least one inserted or deleted nucleotide in the prototype spacer. The insertion / deletion frequency was calculated as the number of reads containing insertions / deletions divided by the total number of mapped reads. Sequencing results are attached. Figure 5 and attached Figure 6 The results showed that ePE significantly improved the targeted editing efficiency at multiple endogenous sites in HeLa cell lines and N2a compared to PE (** indicates p<0.05; *** indicates p<0.01).
[0152] In summary, this invention effectively overcomes the various shortcomings of the prior art and has high industrial applicability. The above embodiments are merely illustrative of the principles and effects of this invention and are not intended to limit the invention. Any person skilled in the art can modify or change the above embodiments without departing from the spirit and scope of this invention. Therefore, all equivalent modifications or changes made by those skilled in the art without departing from the spirit and technical concept disclosed in this invention should still be covered by the claims of this invention. SEQUENCE LISTING <110> ShanghaiTech University <120> A guided editing tool, fusion RNA and its uses <130> P21013278C <160> 17 <170> PatentIn version 3.5 <210> 1 <211> 187 <212> PRT <213> Artificial Sequence <220> <223> Csy4 endonuclease <400> 1 Met Asp His Tyr Leu Asp Ile Arg Leu Arg Pro Asp Pro Glu Phe Pro 1 5 10 15 Pro Ala Gln Leu Met Ser Val Leu Phe Gly Lys Leu His Gln Ala Leu 20 25 30 Val Ala Gln Gly Gly Asp Arg Ile Gly Val Ser Phe Pro Asp Leu Asp 35 40 45 Glu Ser Arg Ser Arg Leu Gly Glu Arg Leu Arg Ile His Ala Ser Ala 50 55 60 Asp Asp Leu Arg Ala Leu Leu Ala Arg Pro Trp Leu Glu Gly Leu Arg 65 70 75 80 Asp His Leu Gln Phe Gly Glu Pro Ala Val Val Pro His Pro Thr Pro 85 90 95 Tyr Arg Gln Val Ser Arg Val Gln Ala Lys Ser Asn Pro Glu Arg Leu 100 105 110 Arg Arg Arg Leu Met Arg Arg His Asp Leu Ser Glu Glu Glu Ala Arg 115 120 125 Lys Arg Ile Pro Asp Thr Val Ala Arg Ala Leu Asp Leu Pro Phe Val 130 135 140 Thr Leu Arg Ser Gln Ser Thr Gly Gln His Phe Arg Leu Phe Ile Arg 145 150 155 160 His Gly Pro Leu Gln Val Thr Ala Glu Glu Gly Gly Phe Thr Cys Tyr 165 170 175 Gly Leu Ser Lys Gly Gly Phe Val Pro Trp Phe 180 185 <210> 2 <211> 1367 <212> PRT <213> Artificial Sequence <220> <223> Cas9n <400> 2 Asp Lys Lys Tyr Ser Ile Gly Leu Asp Ile Gly Thr Asn Ser Val Gly 1 5 10 15 Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe Lys 20 25 30 Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile Gly 35 40 45 Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu Lys 50 55 60 Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys Tyr 65 70 75 80 Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser Phe 85 90 95 Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys His 100 105 110 Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr His 115 120 125 Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp Ser 130 135 140 Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His Met 145 150 155 160 Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro Asp 165 170 175 Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr Asn 180 185 190 Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala Lys 195 200 205 Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn Leu 210 215 220 Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn Leu 225 230 235 240 Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe Asp 245 250 255 Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp Asp 260 265 270 Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp Leu 275 280 285 Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp Ile 290 295 300 Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser Met 305 310 315 320 Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys Ala 325 330 335 Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe Asp 340 345 350 Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser Gln 355 360 365 Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp Gly 370 375 380 Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg Lys 385 390 395 400 Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu Gly 405 410 415 Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe Leu 420 425 430 Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile Pro 435 440 445 Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp Met 450 455 460 Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu Val 465 470 475 480 Val Asp Lys Gly Ser Ala Gln Ser Phe Ile Glu Arg Met Thr Asn 485,490,495 Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser Leu 500 505 510 Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys Tyr 515,520,525 Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln Lys 530 535 540 Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr Val 545 550 555 560 Lys Gln Leu Lys Glu Asp Tyr Phe Lys Ile Glu Cys Phe Asp Ser 565,570,575 Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly Thr 580 585 590 Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp Asn 595 600 605 Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr Leu 610 615 620 Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala His 625 630 635 640 Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr Thr 645 650 655 Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp Lys 660 665 670 Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe Ala 675 680 685 Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe Lys 690 695 700 Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu His 705 710 715 720 Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly Ile 725 730 735 Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly Arg 740 745 750 His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln Thr 755 760 765 Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile Glu 770 775 780 Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro Val 785 790 795 800 Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln 805 810 815 Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg Leu 820 825 830 Ser Asp Tyr Asp Val Asp Ala Ile Val Pro Gln Ser Phe Leu Lys Asp 835 840 845 Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg Gly 850 855 860 Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys Asn 865 870 875 880 Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe 885,890,895 Asp Asn With Thr Lys Ala Glu Arg Gly Gly Ser Glu With Asp Lys 900 905 910 Only Gly Phe With Lys Arg Gln Leu Val Glu Thr Arg Gln With Thr Lys 915,920,925 His Gln Ile Leu Asp Ser Arg With Asn Thr Lys Tyr Asp Glu 930,935,940 Asn Asp Lys With Arg Glu Val Val Lys With Thr Lys Ser Ser Lys 945 950 955 960 Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg Glu 965,970,975 Ile Asn Asn Tyr His Ala His Asp Ala Tyr Leu Asn Ala Val Val 980,985,990 Gly Thr Ala Leu Ile Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val 995 1000 1005 Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala Lys 1010 1015 1020 Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr 1025 1030 1035 Serving Asn With Asn Phe Phe Lys Thr Glue With Thr Leu Ala Asn 1040 1045 1050 Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr 1055 1060 1065 Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val Arg 1070 1075 1080 Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr Glu 1085 1090 1095 Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg 1100 1105 1110 Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys 1115 1120 1125 Lys Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val Leu 1130 1135 1140 Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser 1145 1150 1155 Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser Phe 1160 1165 1170 Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu 1175 1180 1185 Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu Phe 1190 1195 1200 Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly Glu 1205 1210 1215 Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn 1220 1225 1230 Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser Pro 1235 1240 1245 Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys His 1250 1255 1260 Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys Arg 1265 1270 1275 Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr 1280 1285 1290 Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile 1295 1300 1305 Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe 1310 1315 1320 Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr 1325 1330 1335 Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr Gly 1340 1345 1350 Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 <210> 3 <211> 678 <212> PRT <213> Artificial Sequence <220> <223> M‑MLV <400> 3 Ser Thr Leu Asn Ile Glu Asp Glu Tyr Arg Leu His Glu Thr Ser Lys 1 5 10 15 Glu Pro Asp Val Ser Leu Gly Ser Thr Trp Leu Ser Asp Phe Pro Gln 20 25 30 Ala Trp Ala Glu Thr Gly Gly Met Gly Leu Ala Val Arg Gln Ala Pro 35 40 45 Leu Ile Ile Pro Leu Lys Ala Thr Ser Thr Pro Val Ser Ile Lys Gln 50 55 60 Tyr Pro Met Ser Gln Glu Ala Arg Leu Gly Ile Lys Pro His Ile Gln 65 70 75 80 Arg Leu Leu Asp Gln Gly Ile Leu Val Pro Cys Gln Ser Pro Trp Asn 85 90 95 Thr Pro Leu Leu Pro Val Lys Lys Pro Gly Thr Asn Asp Tyr Arg Pro 100 105 110 Val Gln Asp Leu Arg Glu Val Asn Lys Arg Val Glu Asp Ile His Pro 115 120 125 Thr Val Pro Asn Pro Tyr Asn Leu Leu Ser Gly Leu Pro Pro Ser His 130 135 140 Gln Trp Tyr Thr Val Leu Asp Leu Lys Asp Ala Phe Phe Cys Leu Arg 145 150 155 160 Leu His Pro Thr Ser Gln Pro Leu Phe Ala Phe Glu Trp Arg Asp Pro 165 170 175 Glu Met Gly Ile Ser Gly Gln Leu Thr Trp Thr Arg Leu Pro Gln Gly 180 185 190 Phe Lys Asn Ser Pro Thr Leu Phe Asn Glu Ala Leu His Arg Asp Leu 195 200 205 Ala Asp Phe Arg Ile Gln His Pro Asp Leu Ile Leu Leu Gln Tyr Val 210 215 220 Asp Asp Leu Leu Leu Ala Ala Thr Ser Glu Leu Asp Cys Gln Gln Gly 225 230 235 240 Thr Arg Ala Leu Leu Gln Thr Leu Gly Asn Leu Gly Tyr Arg Ala Ser 245 250 255 Ala Lys Lys Ala Gln Ile Cys Gln Lys Gln Val Lys Tyr Leu Gly Tyr 260 265 270 Leu Leu Lys Glu Gly Gln Arg Trp Leu Thr Glu Ala Arg Lys Glu Thr 275 280 285 Val Met Gly Gln Pro Thr Pro Lys Thr Pro Arg Gln Leu Arg Glu Phe 290 295 300 Leu Gly Lys Ala Gly Phe Cys Arg Leu Phe Ile Pro Gly Phe Ala Glu 305 310 315 320 Met Ala Ala Pro Leu Tyr Pro Leu Thr Lys Pro Gly Thr Leu Phe Asn 325 330 335 Trp Gly Pro Asp Gln Gln Lys Ala Tyr Gln Glu Ile Lys Gln Ala Leu 340 345 350 Leu Thr Ala Pro Ala Leu Gly Leu Pro Asp Leu Thr Lys Pro Phe Glu 355 360 365 Leu Phe Val Asp Glu Lys Gln Gly Tyr Ala Lys Gly Val Leu Thr Gln 370 375 380 Lys Leu Gly Pro Trp Arg Arg Pro Val Ala Tyr Leu Ser Lys Lys Leu 385 390 395 400 Asp Pro Val Ala Ala Gly Trp Pro Pro Cys Leu Arg Met Val Ala Ala 405 410 415 Ile Ala Val Leu Thr Lys Asp Ala Gly Lys Leu Thr Met Gly Gln Pro 420 425 430 Leu Val Ile Leu Ala Pro His Ala Val Glu Ala Leu Val Lys Gln Pro 435 440 445 Pro Asp Arg Trp Leu Ser Asn Ala Arg Met Thr His Tyr Gln Ala Leu 450 455 460 Leu Leu Asp Thr Asp Arg Val Gln Phe Gly Pro Val Val Ala Leu Asn 465 470 475 480 Pro Ala Thr Leu Leu Pro Leu Pro Glu Glu Gly Leu Gln His Asn Cys 485 490 495 Leu Asp Ile Leu Ala Glu Ala His Gly Thr Arg Pro Asp Leu Thr Asp 500 505 510 Gln Pro Leu Pro Asp Ala Asp His Thr Trp Tyr Thr Asp Gly Ser Ser 515 520 525 Leu Leu Gln Glu Gly Gln Arg Lys Ala Gly Ala Ala Val Thr Thr Glu 530 535 540 Thr Glu Val Ile Trp Ala Lys Ala Leu Pro Ala Gly Thr Ser Ala Gln 545 550 555 560 Arg Ala Glu Leu Ile Ala Leu Thr Gln Ala Leu Lys Met Ala Glu Gly 565 570 575 Lys Lys Leu Asn Val Tyr Thr Asp Ser Arg Tyr Ala Phe Ala Thr Ala 580 585 590 His Ile His Gly Glu Ile Tyr Arg Arg Arg Gly Trp Leu Thr Ser Glu 595 600 605 Gly Lys Glu Ile Lys Asn Lys Asp Glu Ile Leu Ala Leu Leu Lys Ala 610 615 620 Leu Phe Leu Pro Lys Arg Leu Ser Ile Ile His Cys Pro Gly His Gln 625 630 635 640 Lys Gly His Ser Ala Glu Ala Arg Gly Asn Arg Met Ala Asp Gln Ala 645 650 655 Ala Arg Lys Ala Ala Ile Thr Glu Thr Pro Asp Thr Ser Thr Leu Leu 660 665 670 Ile Glu Asn Ser Ser Pro 675 <210> 4 <211> 2310 <212> PRT <213> Artificial Sequence <220> <223> fusion protein <400> 4 Met Asp His Tyr Leu Asp Ile Arg Leu Arg Pro Asp Pro Glu Phe Pro 1 5 10 15 Pro Ala Gln Leu Met Ser Val Leu Phe Gly Lys Leu His Gln Ala Leu 20 25 30 Val Ala Gln Gly Gly Asp Arg Ile Gly Val Ser Phe Pro Asp Leu Asp 35 40 45 Glu Ser Arg Ser Arg Leu Gly Glu Arg Leu Arg Ile His Ala Ser Ala 50 55 60 Asp Asp Leu Arg Ala Leu Leu Ala Arg Pro Trp Leu Glu Gly Leu Arg 65 70 75 80 Asp His Leu Gln Phe Gly Glu Pro Ala Val Val Pro His Pro Thr Pro 85 90 95 Tyr Arg Gln Val Ser Arg Val Gln Ala Lys Ser Asn Pro Glu Arg Leu 100 105 110 Arg Arg Arg Leu Met Arg Arg His Asp Leu Ser Glu Glu Glu Ala Arg 115 120 125 Lys Arg Ile Pro Asp Thr Val Ala Arg Ala Leu Asp Leu Pro Phe Val 130 135 140 Thr Leu Arg Ser Gln Ser Thr Gly Gln His Phe Arg Leu Phe Ile Arg 145 150 155 160 His Gly Pro Leu Gln Val Thr Ala Glu Glu Gly Gly Phe Thr Cys Tyr 165 170 175 Gly Leu Ser Lys Gly Gly Phe Val Pro Trp Phe Glu Gly Arg Gly Ser 180 185 190 Leu Leu Thr Cys Gly Asp Val Glu Glu Asn Pro Gly Pro Pro Lys Lys 195 200 205 Lys Arg Lys Val Asp Lys Lys Tyr Ser Ile Gly Leu Asp Ile Gly Thr 210 215 220 Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser 225 230 235 240 Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys 245 250 255 Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala 260 265 270 Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn 275 280 285 Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val 290 295 300 Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu 305 310 315 320 Asp Lys Lys His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu 325 330 335 Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys 340 345 350 Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala 355 360 365 Leu Ala His Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp 370 375 380 Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val 385 390 395 400 Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly 405 410 415 Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg 420 425 430 Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu 435 440 445 Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys 450 455 460 Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp 465 470 475 480 Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln 485 490 495 Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu 500 505 510 Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu 515 520 525 Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr 530 535 540 Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu 545 550 555 560 Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly 565 570 575 Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu 580 585 590 Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp 595 600 605 Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln 610 615 620 Ile His Leu Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe 625 630 635 640 Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr 645 650 655 Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg 660 665 670 Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn 675 680 685 Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu 690 695 700 Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro 705 710 715 720 Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr 725 730 735 Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser 740 745 750 Gly Glu Gln Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg 755 760 765 Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu 770 775 780 Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala 785 790 795 800 Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp 805 810 815 Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu 820 825 830 Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys 835 840 845 Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg 850 855 860 Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly 865 870 875 880 Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser 885 890 895 Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser 900 905 910 Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly 915 920 925 Asp Ser Leu His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile 930 935 940 Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys 945 950 955 960 Val Met Gly Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg 965 970 975 Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met 980 985 990 Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys 995 1000 1005 Glu His Pro Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr 1010 1015 1020 Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu 1025 1030 1035 Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp Ala Ile Val 1040 1045 1050 Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys Val Leu 1055 1060 1065 Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro Ser 1070 1075 1080 Glu Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arg Gln Leu Leu 1085 1090 1095 Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe Asp Asn Leu Thr Lys 1100 1105 1110 Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp Lys Ala Gly Phe Ile 1115 1120 1125 Lys His Val Ala 1130 1135 1140 Gln Ile Leu Asp Ser Arg Met Thr Lys Tyr Asp Glu Asp 1145 1150 1155 Lys Leu Ile Arg Glu Val Lys Val Ile Thr Leu Lys Ser Lys Leu 1160 1165 1170 Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg Glu 1175 1180 1185 Ile Asn Asn Tyr His Ala His Asp Ala Tyr Leu Asn Ala Val 1190 1195 1200 Val Gly Thr Ala Leu Ile Lys Tyr Pro Lys Leu Glu Ser Glu 1205 1210 1215 Phe Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile 1220 1225 1230 Only Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe 1235 1240 1245 Phe Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu 1250 1255 1260 Only Asn Gly Glu With Arg Lys Arg Pro Leu With Thr Asn Gly 1265 1270 1275 Glu Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr 1280 1285 1290 Val Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys 1295 1300 1305 Thr Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro 1310 1315 1320 Lys Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp 1325 1330 1335 Pro Lys Lys Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser 1340 1345 1350 Val Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu 1355 1360 1365 Lys Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser 1370 1375 1380 Ser Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr 1385 1390 1395 Lys Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser 1400 1405 1410 Leu Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala 1415 1420 1425 Gly Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr 1430 1435 1440 Val Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly 1445 1450 1455 Ser Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His 1460 1465 1470 Lys His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser 1475 1480 1485 Lys Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser 1490 1495 1500 Ala Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu 1505 1510 1515 Asn Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala 1520 1525 1530 Ala Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr 1535 1540 1545 Ser Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile 1550 1555 1560 Thr Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly 1565 1570 1575 Asp Ser Gly Gly Ser Ser Gly Gly Ser Ser Gly Ser Glu Thr Pro 1580 1585 1590 Gly Thr Ser Glu Ser Ala Thr Pro Glu Ser Ser Gly Gly Ser Ser 1595 1600 1605 Gly Gly Ser Ser Thr Leu Asn Ile Glu Asp Glu Tyr Arg Leu His 1610 1615 1620 Glu Thr Ser Lys Glu Pro Asp Val Ser Leu Gly Ser Thr Trp Leu 1625 1630 1635 Ser Asp Phe Pro Gln Ala Trp Ala Glu Thr Gly Gly Met Gly Leu 1640 1645 1650 Ala Val Arg Gln Ala Pro Leu Ile Ile Pro Leu Lys Ala Thr Ser 1655 1660 1665 Thr Pro Val Ser Ile Lys Gln Tyr Pro Met Ser Gln Glu Ala Arg 1670 1675 1680 Leu Gly Ile Lys Pro His Ile Gln Arg Leu Leu Asp Gln Gly Ile 1685 1690 1695 Leu Val Pro Cys Gln Ser Pro Trp Asn Thr Pro Leu Leu Pro Val 1700 1705 1710 Lys Lys Pro Gly Thr Asn Asp Tyr Arg Pro Val Gln Asp Leu Arg 1715 1720 1725 Glu Val Asn Lys Arg Val Glu Asp Ile His Pro Thr Val Pro Asn 1730 1735 1740 Pro Tyr Asn Leu Leu Ser Gly Leu Pro Pro Ser His Gln Trp Tyr 1745 1750 1755 Thr Val Leu Asp Leu Lys Asp Ala Phe Phe Cys Leu Arg Leu His 1760 1765 1770 Pro Thr Ser Gln Pro Leu Phe Ala Phe Glu Trp Arg Asp Pro Glu 1775 1780 1785 Met Gly Ile Ser Gly Gln Leu Thr Trp Thr Arg Leu Pro Gln Gly 1790 1795 1800 Phe Lys Asn Ser Pro Thr Leu Phe Asn Glu Ala Leu His Arg Asp 1805 1810 1815 Leu Ala Asp Phe Arg Ile Gln His Pro Asp Leu Ile Leu Leu Gln 1820 1825 1830 Tyr Val Asp Asp Leu Leu Leu Ala Ala Thr Ser Glu Leu Asp Cys 1835 1840 1845 Gln Gln Gly Thr Arg Ala Leu Leu Gln Thr Leu Gly Asn Leu Gly 1850 1855 1860 Tyr Arg Ala Ser Ala Lys Lys Ala Gln Ile Cys Gln Lys Gln Val 1865 1870 1875 Lys Tyr Leu Gly Tyr Leu Leu Lys Glu Gly Gln Arg Trp Leu Thr 1880 1885 1890 Glu Ala Arg Lys Glu Thr Val Met Gly Gln Pro Thr Pro Lys Thr 1895 1900 1905 Pro Arg Gln Leu Arg Glu Phe Leu Gly Lys Ala Gly Phe Cys Arg 1910 1915 1920 Leu Phe Ile Pro Gly Phe Ala Glu Met Ala Ala Pro Leu Tyr Pro 1925 1930 1935 Leu Thr Lys Pro Gly Thr Leu Phe Asn Trp Gly Pro Asp Gln Gln 1940 1945 1950 Lys Ala Tyr Gln Glu Ile Lys Gln Ala Leu Leu Thr Ala Pro Ala 1955 1960 1965 Leu Gly Leu Pro Asp Leu Thr Lys Pro Phe Glu Leu Phe Val Asp 1970 1975 1980 Glu Lys Gln Gly Tyr Ala Lys Gly Val Leu Thr Gln Lys Leu Gly 1985 1990 1995 Pro Trp Arg Arg Pro Val Ala Tyr Leu Ser Lys Lys Leu Asp Pro 2000 2005 2010 Val Ala Ala Gly Trp Pro Pro Cys Leu Arg Met Val Ala Ala Ile 2015 2020 2025 Ala Val Leu Thr Lys Asp Ala Gly Lys Leu Thr Met Gly Gln Pro 2030 2035 2040 Leu Val Ile Leu Ala Pro His Ala Val Glu Ala Leu Val Lys Gln 2045 2050 2055 Pro Pro Asp Arg Trp Leu Ser Asn Ala Arg Met Thr His Tyr Gln 2060 2065 2070 Ala Leu Leu Leu Asp Thr Asp Arg Val Gln Phe Gly Pro Val Val 2075 2080 2085 Ala Leu Asn Pro Ala Thr Leu Leu Pro Leu Pro Glu Glu Gly Leu 2090 2095 2100 Gln His Asn Cys Leu Asp Ile Leu Ala Glu Ala His Gly Thr Arg 2105 2110 2115 Pro Asp Leu Thr Asp Gln Pro Leu Pro Asp Ala Asp His Thr Trp 2120 2125 2130 Tyr Thr Asp Gly Ser Ser Leu Leu Gln Glu Gly Gln Arg Lys Ala 2135 2140 2145 Gly Ala Ala Val Thr Thr Glu Thr Glu Val Ile Trp Ala Lys Ala 2150 2155 2160 Leu Pro Ala Gly Thr Ser Ala Gln Arg Ala Glu Leu Ile Ala Leu 2165 2170 2175 Thr Gln Ala Leu Lys Met Ala Glu Gly Lys Lys Leu Asn Val Tyr 2180 2185 2190 Thr Asp Ser Arg Tyr Ala Phe Ala Thr Ala His Ile His Gly Glu 2195 2200 2205 Ile Tyr Arg Arg Arg Gly Trp Leu Thr Ser Glu Gly Lys Glu Ile 2210 2215 2220 Lys Asn Lys Asp Glu Ile Leu Ala Leu Leu Lys Ala Leu Phe Leu 2225 2230 2235 Pro Lys Arg Leu Ser Ile Ile His Cys Pro Gly His Gln Lys Gly 2240 2245 2250 His Ser Ala Glu Ala Arg Gly Asn Arg Met Ala Asp Gln Ala Ala 2255 2260 2265 Arg Lys Ala Ala Ile Thr Glu Thr Pro Asp Thr Ser Thr Leu Leu 2270 2275 2280 Ile Glu Asn Ser Ser Pro Ser Gly Gly Ser Lys Arg Thr Ala Asp 2285 2290 2295 Gly Ser Glu Phe Glu Pro Lys Lys Lys Arg Lys Val 2300 2305 2310 <210> 5 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> Csy4 endonuclease recognition sequence <400> 5 gttcactgcc gtataggcag 20 <210> 6 <211> 18 <212> PRT <213> Artificial Sequence <220> <223> T2A segment <400> 6 Glu Gly Arg Gly Ser Leu Leu Thr Cys Gly Asp Val Glu Glu Asn Pro 1 5 10 15 Gly Pro <210> 7 <211> 17 <212> PRT <213> Artificial Sequence <220> <223> BPNLS fragments <400> 7 Lys Arg Thr Ala Asp Gly Ser Glu Phe Glu Pro Lys Lys Lys Arg Lys 1 5 10 15 Val <210> 8 <211> twenty two <212> DNA <213> Artificial Sequence <220> <223> Csy4 endonuclease forward primer <400> 8 atggaccact acctcgacat tc 22 <210> 9 <211> twenty one <212> DNA <213> Artificial Sequence <220> <223> Csy4 endonuclease reverse primer <400> 9 gaaccaggga acgaaacctc c 21 <210> 10 <211> 68 <212> DNA <213> Artificial Sequence <220> <223> Csy4 product forward primer <400> 10 gtcagatccg ctagagatcc gcggccgcta atacgactca ctataggatg gaccactacc 60 tcgacatt 68 <210> 11 <211> 59 <212> DNA <213> Artificial Sequence <220> <223> Csy4 product reverse primer <400> 11 gacgtcaccg catgttaaca gacttcctct gccctcgaac cagggaacga aacctcctt 59 <210> 12 <211> 59 <212> DNA <213> Artificial Sequence <220> <223> PE2 vector forward primer <400> 12 tgttaacatg cggtgacgtc gaggagaatc ctggcccacc aaagaagaag cggaaagtc 59 <210> 13 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> PE2 vector reverse primer <400> 13 tgccggccca tcactttcac 20 <210> 14 <211> 67 <212> DNA <213> Artificial Sequence <220> <223> scaffold‑F <400> 14 agagctagaa atagcaagtt aaaataaggc tagtccgtta tcaacttgaa aaagtggcac 60 cgagtcg 67 <210> 15 <211> 67 <212> DNA <213> Artificial Sequence <220> <223> scaffold‑R <400> 15 gcaccgactc ggtgccactt tttcaagttg ataacggact agccttattt taacttgcta 60 tttctag 67 <210> 16 <211> 59 <212> DNA <213> Artificial Sequence <220> <223> Csy4peg‑bone‑F <400> 16 gagagggtct cagttttaga gctagaaata gcaagttaaa ataaggctag tccgttatc 59 <210> 17 <211> 32 <212> DNA <213> Artificial Sequence <220> <223> Csy4peg‑bone‑R <400> 17 ctctcggtct cacggtgttt cgtcctttcc and 32
Claims
1. A guided editing tool, characterized in that, It includes: (i) A fusion protein comprising, from N-terminus to C-terminus, a Csy4 endonuclease, a T2A fragment, a Cas9n, a linker, and Moroni mouse leukemia virus reverse transcriptase M-MLV; (ii) A fusion RNA, wherein the fusion RNA comprises pegRNA, a Csy4 endonuclease recognition sequence, and a nick sgRNA from the 5' end to the 3' end, wherein a 20 nt spacer primer for the pegRNA is designed according to the target site sequence, ACCG is added to the 5' end of the upstream primer, GTTTC is added to the 3' end, and CTCTGAAAC is added to the 5' end of the downstream primer, and oligonucleotide primers for scaffold are synthesized: scaffold-F: agagctagaaatagcaagttgaaataaggctagtccgttatcaacttgaaaaagtggcaccgagtcg (SEQ ID NO: 14), scaffold-R: gcaccgactcggtgccactttttcaagttgataacggactagccttatttcaacttgctatttctag (SEQ ID NO: 15). The fusion protein has reverse transcription function and can bind to and cleave the recognition site, thereby introducing a sequence at the 3' end of the pegRNA and preventing the pegRNA from circularizing itself.
2. The guided editing tool as described in claim 1, characterized in that, The nucleotide sequence of the Csy4 endonuclease recognition sequence is shown in SEQ ID NO:
5.
3. The guided editing tool as described in claim 2, characterized in that, The amino acid sequence of the Csy4 endonuclease is shown in SEQ ID NO: 1, the amino acid sequence of the Cas9n is shown in SEQ ID NO: 2, and / or the amino acid sequence of the M-MLV is shown in SEQ ID NO:
3.
4. The guided editing tool as described in claim 3, characterized in that, The fusion protein also includes a BPNLS fragment.
5. The guided editing tool as described in claim 4, characterized in that, The amino acid sequence of the T2A fragment is shown in SEQ ID NO:6, and / or the BPNLS fragment is located at the C-terminus and its amino acid sequence is shown in SEQ ID NO:
7.
6. The guided editing tool as described in any one of claims 1 to 5, characterized in that, The recognition sequence of the Csy4 endonuclease contained in the fusion RNA is the nucleotide sequence shown in SEQ ID NO:
5.
7. The guided editing tool as described in claim 6, characterized in that, The amino acid sequence of the fusion protein is shown in SEQ ID NO:
4.
8. Use of the guided editing tool as described in any one of claims 1 to 7 in eukaryotic gene editing, wherein the use is for non-therapeutic purposes.
9. The use as described in claim 8, characterized in that, The applications include base substitution, insertion, or deletion.