Pilot editing system based on perv reverse transcriptase

By using the PERV-RT editor instead of MMLV-RT, the problem of low efficiency of leader editing in mammalian cells was solved, achieving more efficient gene editing results.

CN118995664BActive Publication Date: 2026-03-20JIANGXI AGRICULTURAL UNIVERSITY +1
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-21
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing lead editing systems are inefficient in mammalian cells, especially the ePPE system using MMLV-RT, which has failed to significantly improve editing efficiency in mammalian cells.

Method used

By replacing MMLV-RT with PERV-NC fused with RNase H structure and carrying multiple mutation sites, a highly efficient PERV-RT editor is formed for pilot editing in mammalian cells.

Benefits of technology

It significantly improves the efficiency of lead editing in mammalian cells, achieving greater precision and efficiency in gene editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present application relates to the field of genetic engineering. Specifically, the present application relates to a PERV reverse transcriptase-based prime editing system and a method for gene editing by the system. More specifically, the present application provides: 1. A prime editing system using PERV-NC fusion RNase H structure deletion and carrying multiple mutation sites of PERV-RT; 2. A pegRNA fused with a truncated Csy4 (tCsy4-pegRNA); 3. One or more small molecule compounds for treating cells and combinations thereof; to improve the efficiency of prime editing in mammals.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of genetic engineering. In particular, the present application relates to a PERV reverse transcriptase-based prime editing system and a method for gene editing by the system. BACKGROUND

[0002] Prime editing (PE) is a new gene editing technology that has emerged in recent years. It can achieve point mutation, insertion, deletion, replacement and other types of editing without generating DNA double-strand breaks, overcoming the main shortcomings of current CRISPR / Cas9 and single-base editing technology, and is the most precise genome editing technology currently available. The PE editing system is composed of a PE editor and a prime editing guide RNA (pegRNA and / or ngRNA). The PE editor is the most important component of prime editing. The PE editor is a fusion of a Cas9 single-nicking enzyme (nCas9 H840A) and a reverse transcriptase (RT). The current PE2 and its improved version PEmax editors, which are widely used, use the reverse transcriptase of the murine Moloney leukemia virus (MMLV-RT) carrying five mutations of D200N, L603W, T330P, T306K and W313F.

[0003] Other types of RT have been used for prime editing by some teams, but their efficiency is lower or significantly lower than that of PE2. Another report shows that the fusion of nucleocapsid protein (NC) and RT helps to improve the efficiency of prime editing. The nucleocapsid protein (NC) of MMLV and the MMLV-RT with deleted RNase H structure (MMLV-RT(ΔRNase H)) are fused to develop an efficiency-enhanced PE editor: ePPE. However, so far, reports have shown that ePPE only significantly improves editing efficiency in plant cells, and does not significantly improve editing efficiency in mammalian cells.

[0004] There is still a need for efficient prime editing systems, particularly those that can perform efficient editing in mammalian cells. SUMMARY

[0005] The present application uses PERV-NC fusion PERV-RT with deleted RNase H structure and carrying multiple mutation sites to replace MMLV-RT in PE2 editor to improve the efficiency of prime editing in mammals. BRIEF DESCRIPTION OF DRAWINGS

[0006] Figure 1 . shows the amino acid sequence alignment of MMLV-RT(ΔRNase H), PERV-AC-RT(ΔRNase H) and PERV-B-RT(ΔRNase H).

[0007] Figure 2 Figure 2. Vector map of PERV reverse transcriptase editors of the present application showing different NCBI sequences.

[0008] Figure 3 Figure 3. Editing efficiency of PE2c-Puro, epPE-B-M3A-Puro, epPE-B-M4A*-Puro editors in BFP reporter cell line.

[0009] Figure 4 Figure 4. Editing efficiency of PE2c-Puro, epPE-M3A-Puro editors in primary PFF.

[0010] Figure 5 Figure 5. Editing efficiency of PE2c-Puro and epPE-AC-M3A-Puro editors in HEK293T, dog fetal fibroblast (DFF) and sheep fetal fibroblast (SFF) cells.

[0011] Figure 6 Figure 6. Vector map of Bama mini pig PERV reverse transcriptase editors of the present application based on different sequences.

[0012] Figure 7 Figure 7. Editing activity of pPE editors of Bama mini pig PERV reverse transcriptase containing 4 point mutations and wild type in BFP reporter cell line.

[0013] Figure 8 Figure 8. Editing efficiency of Bama mini pig PERV reverse transcriptase pPE editors containing 4 point mutations in PFF.

[0014] Figure 9 Figure 9. Editing efficiency of Bama mini pig PERV reverse transcriptase pPE editors containing 4 point mutations in HEK293T, dog fetal fibroblast (DFF), sheep fetal fibroblast (SFF) cells.

[0015] Figure 10 Figure 10. Editing efficiency of Bama mini pig PERV reverse transcriptase epPE editors based on 4 point mutations in PFF.

[0016] Figure 11 Figure 11. Editing efficiency of Bama mini pig PERV reverse transcriptase epPE editors based on 4 point mutations in HEK293T, dog fetal fibroblast (DFF), sheep fetal fibroblast (SFF) cells.

[0017] Figure 12. Shows editing efficiency of Bama mini pig PERV reverse transcriptase pPE editors containing 1-6 point mutations in HEK293T cells co-transfected with a fluorescent reporter gene.

[0018] Figure 13 . Shows editing efficiency of Bama mini pig PERV reverse transcriptase pPE editors based on M5.1 with different point mutations in HEK293T cells.

[0019] Figure 14 . Shows editing efficiency of Bama mini pig PERV reverse transcriptase pPE editors based on 7 point mutations in HEK293T cells.

[0020] Figure 15 . Shows editing efficiency of Bama mini pig PERV reverse transcriptase epPE editors based on 7 point mutations in HEK293T cells.

[0021] Figure 16 . Shows editing efficiency of Bama mini pig PERV reverse transcriptase epPE editors based on 7 point mutations compared to PEmax, PE6c and PE6d.

[0022] Figure 17 . Shows Csy4 and tCsy4-pegRNA improve editing efficiency of PE editors.

[0023] Figure 18 . Shows small molecule compounds improve editing efficiency of PE editors. DETAILED DESCRIPTION

[0024] I. DEFINITIONS

[0025] In the present application, the scientific and technical terms used herein have the meanings commonly understood by one of ordinary skill in the art, unless otherwise indicated. Also, the terms and techniques employed herein of protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, immunology are those commonly used by those skilled in the art and are generally described in the literature in the corresponding field. For example, the standard recombinant DNA and molecular cloning techniques used in the present application are well known to those skilled in the art and are described more fully in Sambrook, J., Fritsch, E.F. and Maniatis, T., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press: Cold Spring Harbor, 1989 (hereinafter "Sambrook"). Also, for better understanding of the present application, the definitions and explanations of the relevant terms are provided below.

[0026] As used herein, the term "and / or" encompasses all combinations of the items connected by the term. For example, "A and / or B" covers A, B, and A and B. For example, "A, B, and / or C" covers A, B, C, A and B, A and C, B and C, and A and B and C.

[0027] The word "comprise" when used in describing the sequence of a protein or nucleic acid means that the protein or nucleic acid can consist of the sequence, or can have additional amino acids or nucleotides at either or both ends of the protein or nucleic acid, but still have the activity described in the present application. Furthermore, it is clear to a person skilled in the art that the methionine encoded by the start codon at the N-terminus of a polypeptide is in some practical cases (e.g. when expressed in a particular expression system) retained, but does not materially affect the function of the polypeptide. Therefore, when describing a specific amino acid sequence of a polypeptide in the specification and claims of the present application, although it can not comprise the methionine encoded by the start codon at the N-terminus, it nevertheless also encompasses sequences comprising the methionine, and correspondingly, the nucleotide sequence encoding it can comprise the start codon; vice versa.

[0028] "Genome" as used herein encompasses not only chromosomal DNA present in the nucleus of a cell, but also organelle DNA present in subcellular components (such as mitochondria, plastids) of a cell.

[0029] As used herein, "organism" includes any organism suitable for genome editing, preferably a eukaryote. Examples of organisms include, but are not limited to, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cows, cats; poultry such as chickens, ducks, geese; plants including monocotyledons and dicotyledons, for example rice, maize, wheat, sorghum, barley, soybean, peanut, Arabidopsis, etc.

[0030] "Genetically modified organism" means an organism that comprises an exogenous polynucleotide or comprises a modified gene or expression regulatory sequence within its genome. For example, the exogenous polynucleotide can be stably integrated into the genome of the organism and inherited through successive generations. The exogenous polynucleotide can be integrated into the genome either alone or as part of a recombinant DNA construct. The modified gene or expression regulatory sequence is one that comprises one or more nucleotide substitutions, deletions and additions in the gene or expression regulatory sequence in the genome of the organism.

[0031] "Exogenous" with respect to a sequence means a sequence from a foreign species, or if from the same species, a sequence that has been significantly altered in composition and / or locus from its natural form by deliberate human intervention.

[0032] "Polynucleotide," "nucleic acid sequence," "nucleotide sequence," or "nucleic acid fragment" are used interchangeably and are single- or double-stranded RNA or DNA polymers, optionally containing synthetic, non-natural or altered nucleotide bases. Nucleotides are referred to by their single letter designation: "A" is either adenosine or deoxyadenosine (corresponding to RNA or DNA, respectively), "C" denotes cytidine or deoxycytidine, "G" denotes guanosine or deoxyguanosine, "U" denotes uridine, "T" denotes deoxythymidine, "R" denotes purine (A or G), "Y" denotes pyrimidine (C or T), "K" denotes G or T, "H" denotes A or C or T, "D" denotes A, T or G, "I" denotes inosine, and "N" denotes any nucleotide. Although nucleotide sequences herein can be represented in DNA sequence (containing T), the corresponding RNA sequence (i.e., with U in place of T) can be readily determined by one of skill in the art when RNA is referred to.

[0033] "Polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of naturally occurring amino acids, as well as to naturally occurring amino acid polymers. The terms "polypeptide," "peptide," "amino acid sequence," and "protein" can also include modified forms, including but not limited to glycosylation, lipid attachment, sulfation, gamma-carboxylation of glutamic acid residues, hydroxylation, and ADP-ribosylation.

[0034] The sequence "identity" has the conventional meaning in the art and can be calculated as the percentage of nucleotide or amino acid residues in the candidate sequence that match the reference sequence when aligned for maximum correspondence over a specified comparison window. The sequence identity can be measured using publicly available computer programs such as BLAST, ALIGN, or Megalign (DNASTAR) software. The sequence identity can be measured over the entire length of the polynucleotide or polypeptide or over a region of that molecule. (See, e.g., Computational Molecular Biology, Lesk, A. M., ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, D. W., ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, A. M., and Griffin, H. G., eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). While there are a number of methods for measuring sequence identity depending on the comparison window used, the term "identity" is well known to one of skill in the art (Carrillo, H. & Lipman, D., SIAM J Applied Math 48: 1073 (1988)).

[0035] As used herein, "amino acid position with reference to SEQ ID NO: x" (where SEQ ID NO: x is a certain specific sequence listed herein) means that the position number of the specific amino acid described is the position number of the corresponding amino acid on SEQ ID NO: x. Correspondence of amino acids in different sequences can be determined according to methods of sequence alignment known in the art. For example, amino acid correspondence can be determined by the online alignment tool of EMBL-EBI (https: / / www.ebi.ac.uk / Tools / psa / ), where two sequences can be aligned using the Needleman-Wunsch algorithm, using default parameters. For example, a proline amino acid at position 360 from the N-terminus of a polypeptide is aligned with the amino acid at position 365 of SEQ ID NO: x in the sequence alignment, then this proline amino acid in the polypeptide can also be described herein as "a proline at position 365 of the polypeptide, the amino acid position with reference to SEQ ID NO: x".

[0036] As used herein, "expression construct" refers to a vector, such as a recombinant vector, suitable for expression of a nucleotide sequence of interest in an organism. "Expression" refers to the production of a functional product. For example, expression of a nucleotide sequence can refer to transcription of the nucleotide sequence (e.g., to produce mRNA or a functional RNA) and / or translation of the RNA into a precursor or mature protein.

[0037] The "expression construct" of the application can be a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, can be a translatable RNA (e.g., mRNA), such as an in vitro transcribed RNA. For application in a mammal, such as a human, the expression construct is preferably a viral vector, such as an adeno-associated virus (AAV), a lentivirus, an adenovirus vector, most preferably an AAV vector.

[0038] The "expression construct" of the application can comprise regulatory sequences and nucleotide sequences of interest of different origin, or regulatory sequences and nucleotide sequences of interest of the same origin but arranged in a manner different from that which normally occurs in nature.

[0039] A "promoter" refers to a nucleic acid fragment that is capable of controlling the transcription of another nucleic acid fragment. In some embodiments of the application, a promoter is a promoter that is capable of controlling the transcription of a gene in a cell, whether or not it is derived from the cell. A promoter can be a constitutive promoter or a tissue-specific promoter or a developmentally-regulated promoter or an inducible promoter.

[0040] Examples of promoters include, but are not limited to, polymerase pol I, pol II, or pol III promoters. Examples of pol I promoters include the chicken RNA pol I promoter. Examples of pol II promoters include, but are not limited to, the cytomegalovirus immediate early (CMV) promoter, the Rous sarcoma virus long terminal repeat (RSV-LTR) promoter, and the simian virus 40 (SV40) immediate early promoter. Examples of pol III promoters include the U6, U3, and H1 promoters. Inducible promoters such as the metallothionein promoter can be used. Other examples of promoters include the T7 phage promoter, the T3 phage promoter, the beta-galactosidase promoter, and the Sp6 phage promoter. When used in plants, the promoter can be the cauliflower mosaic virus 35S promoter, the maize Ubi-1 promoter, the wheat U6 promoter, the rice U3 promoter, the maize U3 promoter, the rice actin promoter.

[0041] "Introducing" a nucleic acid molecule (e.g., a plasmid, a linear nucleic acid fragment, an RNA, etc.) or a protein into an organism means transforming a cell of the organism with the nucleic acid or protein such that the nucleic acid or protein is able to function in the cell. As used herein, "transforming" includes both stable transformation and transient transformation. "Stable transformation" means introducing an exogenous nucleotide sequence into the genome such that the exogenous gene is stably inherited. Once stably transformed, the exogenous nucleic acid sequence is stably integrated into the genome of the organism and any successive generations thereof. "Transient transformation" means introducing a nucleic acid molecule or protein into a cell to perform a function without the exogenous gene being stably inherited. In transient transformation, the exogenous nucleic acid sequence is not integrated into the genome.

[0042] "Trait" means a physiological, morphological, biochemical, or physical characteristic of a cell or organism.

[0043] II. Reverse transcriptases from PERV and variants thereof

[0044] In one aspect, the present application provides a reverse transcriptase from a Porcine Endogenous Retrovirus (PERV) or a functional variant thereof. The reverse transcriptase has reverse transcriptase activity to generate a complementary DNA strand from a single-stranded RNA template, and the functional variant thereof substantially retains or has enhanced reverse transcriptase activity. The reverse transcriptase from a PERV or a functional variant thereof of the present application can be used for prime editing, particularly for performing prime editing in mammalian cells.

[0045] In some embodiments, the reverse transcriptase from a PERV (also referred to herein as a PERV reverse transcriptase) is from a type A, type B, or type C PERV, or a recombinant of different types of PERV, e.g., a type A / C PERV recombinant.

[0046] In some embodiments, the PERV is a PERV isolated from a Bama mini pig. In some embodiments, the PERV is isolated from other breeds of pigs.

[0047] In some embodiments, the PERV reverse transcriptase or functional variant thereof comprises the amino acid sequence of one of SEQ ID NOs: 1-5 or an amino acid sequence having at least 75%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to one of SEQ ID NOs: 1-5.

[0048] In some embodiments, the functional variant of the PERV reverse transcriptase lacks the RNase H domain. The RNase H domain of the PERV reverse transcriptase, for example, comprises the amino acid sequence corresponding to about position 514 to about position 616 of SEQ ID NO: 1.

[0049] In some embodiments, the functional variant of the PERV reverse transcriptase comprises an amino acid substitution at one or more positions selected from position 63, 199, 222, 305, 312, 329, 330, 331, 332, 408, or 602.

[0050] When referring to an amino acid position of a PERV reverse transcriptase or functional variant thereof herein, the amino acid position is with reference to SEQ ID NO: 5.

[0051] In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions at positions 199, 305, and 312. In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions D199N, T305K, and W312F.

[0052] In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions at positions 199, 305, 312, and 602. In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions D199N, T305K, W312F, and L602W.

[0053] In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions at positions 199, 305, 312, 329, 330, 331, and 332. In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions D199N, T305K, W312F, E329P, K330G, G331T, and E332C.

[0054] In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions at positions 199, 305, 312, 329, 330, 331, 332, and 602. In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions D199N, T305K, W312F, E329P, K330G, G331T, E332C, and L602W.

[0055] In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions at positions 199, 222, 305, 312, and 602. In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions D199N, V222A, T305K, W312F, and L602W.

[0056] In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions at positions 63, 199, 222, 305, 312, and 602. In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions Y63R, D199N, V222A, T305K, W312F, and L602W.

[0057] In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions at positions 199, 222, 305, 312, 408, and 602. In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions D199N, V222A, T305K, W312F, C408R, and L602W.

[0058] In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions at positions 63, 199, 222, 305, 312, 408, and 602. In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions Y63R, D199N, V222A, T305K, W312F, C408R, and L602W.

[0059] In some embodiments, the functional variant of the PERV reverse transcriptase comprises an amino acid sequence selected from the group consisting of SEQ ID NO: 6 (corresponding to AC-RT-M0A), SEQ ID NO: 7 (corresponding to AC-RT-M3A), SEQ ID NO: 8 (corresponding to AC-RT-M4A*), SEQ ID NO: 9 (corresponding to B-RT-M0A), SEQ ID NO: 10 (corresponding to B-RT-M3A), SEQ ID NO: 11 (corresponding to B-RT-M4A*), and SEQ ID NO: 12 (corresponding to bmx-A-RT-M4), SEQ ID NO: 13 (corresponding to bmx-B-RT-M4), SEQ ID NO: 14 (corresponding to bmx-C-RT-M4), SEQ ID NO: 15 (corresponding to bmx-A-RT-M3A), SEQ ID NO: 16 (corresponding to bmx-B-RT-M3A), SEQ ID NO: 17 (corresponding to bmx-C-RT-M3A), SEQ ID NO: 40 (corresponding to bmx-C-RT-M5.1), SEQ ID NO: 41 (corresponding to bmx-C-RT-M5.1+Y63R), SEQ ID NO: 42 (corresponding to bmx-C-RT-M5.1+C408R), SEQ ID NO: 43 (corresponding to bmx-C-RT-M7), and SEQ ID NO: 44 (corresponding to bmx-C-RT-M7A).

[0060] In another aspect, the present application provides the use of a reverse transcriptase from PERV or a functional variant thereof according to the present application for prime editing, in particular for prime editing in a mammalian cell.

[0061] As used herein, "mammal" includes, but is not limited to, humans, mice, rats, monkeys, dogs, pigs, sheep, cows, cats, and the like. In some preferred embodiments, the mammal is a pig. In some preferred embodiments, the mammal is a human.

[0062] III. Prime editing system based on reverse transcriptase from PERV

[0063] In one aspect, the present application relates to a prime editing system for targeted modification of a genomic DNA sequence of an organism, comprising:

[0064] i) a CRISPR nuclease such as a CRISPR nickase and / or an expression construct containing a nucleotide sequence encoding said CRISPR nuclease such as a CRISPR nickase, and a PERV reverse transcriptase or a functional variant thereof and / or an expression construct containing a nucleotide sequence encoding said PERV reverse transcriptase or a functional variant thereof; and / or

[0065] ii) at least one pegRNA and / or an expression construct containing a nucleotide sequence encoding said at least one pegRNA.

[0066] In some embodiments, wherein said at least one pegRNA comprises, in the 5’ to 3’ direction, a leader sequence, a scaffold sequence, a reverse transcription template (RTT) sequence, and a primer binding site (PBS) sequence.

[0067] In some embodiments, said CRISPR nuclease such as CRISPR nickase and said PERV reverse transcriptase or functional variant thereof form a fusion protein.

[0068] In one aspect, the present application relates to a prime editing system for targeted modification of genomic DNA sequences of an organism, comprising:

[0069] i) a prime editing fusion protein and / or an expression construct containing a nucleotide sequence encoding said prime editing fusion protein, wherein said prime editing fusion protein comprises a CRISPR nuclease such as CRISPR nickase and a PERV reverse transcriptase or functional variant thereof; and / or

[0070] ii) at least one pegRNA and / or an expression construct containing a nucleotide sequence encoding said at least one pegRNA.

[0071] In some embodiments, wherein said at least one pegRNA comprises, in the 5’ to 3’ direction, a leader sequence, a scaffold sequence, a reverse transcription template (RTT) sequence, and a primer binding site (PBS) sequence.

[0072] In another aspect, the present application also provides a prime editing fusion protein comprising a CRISPR nuclease such as CRISPR nickase and a PERV reverse transcriptase or functional variant thereof. The present application also provides the use of said prime editing fusion protein for prime editing, in particular for prime editing in mammalian cells.

[0073] In some embodiments, said CRISPR nuclease such as CRISPR nickase and said PERV reverse transcriptase or functional variant thereof in the prime editing fusion protein are connected by a linker.

[0074] In some embodiments, said at least one pegRNA is capable of forming a complex with said CRISPR nuclease such as CRISPR nickase, said PERV reverse transcriptase or functional variant thereof and targeting said fusion protein to a target sequence in the genome, resulting in a nick within said target sequence.

[0075] In some embodiments, the at least one pegRNA is capable of forming a complex with the fusion protein and targeting the fusion protein to a target sequence in the genome, resulting in a nick within the target sequence.

[0076] As used herein, a “prime editing system” refers to a combination of components required for reverse transcription-based genome editing of a genome in a cell. The individual components of the system, such as a CRISPR nuclease, e.g., a CRISPR nickase, a PERV reverse transcriptase or functional variant thereof, a prime editing fusion protein, a pegRNA, etc., can each independently exist, or can exist in any combination as a composition.

[0077] As used herein, a “target sequence” refers to a sequence of about 20 nucleotides in length in a genome characterized by a 5’ or 3’ flanking PAM (protospacer adjacent motif) sequence. Generally, a PAM is required for the complex of a CRISPR nuclease or variant thereof and a guide RNA to recognize a target sequence. For example, for Cas9 nucleases and variants thereof, the target sequence is immediately adjacent to a PAM at the 3’ end, e.g., 5’-NGG-3’. Based on the presence of a PAM, one of skill in the art can readily determine target sequences in a genome that can be targeted. Moreover, depending on the location of the PAM, a target sequence can be on either strand of a genomic DNA molecule. For Cas9 or derivatives thereof, e.g., Cas9 nickases, a target sequence is preferably 20 nucleotides.

[0078] The CRISPR nuclease of the present application can be a CRISPR nuclease of different origin. Examples of CRISPR nucleases suitable for the present application include, but are not limited to, Cas9, Cas12a (i.e., Cpf1), Cas12b (i.e., C2c1), Cas12c, Cas12d (i.e., CasY), Cas12e (i.e., CasX), Cas12f (i.e., Cas14), Cas12h, Cas12i, Cas12j (i.e., Cas ), Cas12k, Cas12m, IsrB, IscB, and TnpB.

[0079] In some embodiments, the CRISPR nuclease further comprises a nickase form thereof. A CRISPR nuclease in the form of a nickase can also be referred to herein as a CRISPR nickase, which forms a nick in only one strand of a double-stranded nucleic acid molecule, but does not completely cleave the double-stranded nucleic acid. In some embodiments, the CRISPR nickase (nickase) is capable of forming a nick within a target sequence in a genomic DNA. Nickase forms of different CRISPR nucleases are available to one of skill in the art, e.g., by mutating the nucleic acid cleavage domain of a CRISPR nuclease.

[0080] In some embodiments, the CRISPR nickase is a Cas9 nickase. However, the CRISPR nickase can also be derived from other CRISPR nucleases.

[0081] In some embodiments, the Cas9 nickase is derived from SpCas9 of S. pyogenes, and comprises at least the amino acid substitution H840A relative to wild-type SpCas9. In some embodiments, the Cas9 nickase comprises the amino acid sequence set forth in SEQ ID NO: 18. In some embodiments, the Cas9 nickase in the fusion protein is capable of forming a nick between the -3 position nucleotide of the PAM (the first nucleotide at the 5' end of the PAM sequence is the +1 position) and the -4 position nucleotide of the PAM of a target sequence.

[0082] In some embodiments, the Cas9 nickase is a Cas9 nickase variant capable of recognizing an altered PAM sequence. Many Cas9 nickase variants capable of recognizing an altered PAM sequence are known in the art. For example, the Cas9 nickase is a Cas9 variant recognizing the PAM sequence 5'-NG-3'. These different variants can all be applied in the present application.

[0083] The nick formed by the Cas9 nickase of the present application can result in the formation of a free single strand with a 3' end (3' free single strand) and a free single strand with a 5' end (5' free single strand) of the target sequence.

[0084] In some embodiments, the PERV reverse transcriptase or functional variant thereof in the fusion protein is a PERV reverse transcriptase or functional variant thereof as described herein before.

[0085] In some embodiments, the PERV reverse transcriptase is from a PERV of type A, type B or type C, or from a recombinant of PERVs of different types, e.g. a recombinant of type A / C PERV.

[0086] In some embodiments, the PERV is a PERV isolated from a Bama mini pig. In some embodiments, the PERV is isolated from a pig of another breed.

[0087] In some embodiments, the PERV reverse transcriptase or functional variant thereof comprises the amino acid sequence of one of SEQ ID NOs: 1-5 or an amino acid sequence having at least 75%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even at least 100% sequence identity to one of SEQ ID NOs: 1-5.

[0088] In some embodiments, the functional variant of the PERV reverse transcriptase lacks the RNase H domain. The RNase H domain of the PERV reverse transcriptase, for example, comprises the amino acid sequence corresponding to about position 514 to about position 616 of SEQ ID NO: 5.

[0089] In some embodiments, the functional variant of the PERV reverse transcriptase comprises an amino acid substitution at one or more positions selected from position 199, 305, 312, 329, 330, 331, 332, or 602.

[0090] Reference to an amino acid position of a PERV reverse transcriptase or a functional variant thereof herein is with reference to SEQ ID NO: 1.

[0091] In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions D199N, T305K, and W312F. In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions D199N, T305K, and W312F at positions 199, 305, and 312.

[0092] In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions D199N, T305K, W312F, and L602W. In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions D199N, T305K, W312F, and L602W at positions 199, 305, 312, and 602.

[0093] In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions D199N, T305K, W312F, E329P, K330G, G331T, and E332C. In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions D199N, T305K, W312F, E329P, K330G, G331T, and E332C at positions 199, 305, 312, 329, 330, 331, and 332.

[0094] In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions D199N, T305K, W312F, E329P, K330G, G331T, E332C, and L602W. In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions D199N, T305K, W312F, E329P, K330G, G331T, E332C, and L602W at positions 199, 305, 312, 329, 330, 331, 332, and 602.

[0095] In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions at positions 63, 199, 222, 305, 312, 408, and 602. In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions Y63R, D199N, V222A, T305K, W312F, C408R, and L602W.

[0096] In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions at positions 63, 199, 222, 305, 312, 408, and 602. In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions Y63R, D199N, V222A, T305K, W312F, C408R, and L602W.

[0097] In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions at positions 63, 199, 222, 305, 312, 408, and 602. In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions Y63R, D199N, V222A, T305K, W312F, C408R, and L602W.

[0098] In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions at positions 63, 199, 222, 305, 312, 408, and 602. In some embodiments, the functional variant of the PERV reverse transcriptase comprises amino acid substitutions Y63R, D199N, V222A, T305K, W312F, C408R, and L602W.

[0099] In some embodiments, the functional variant of the PERV reverse transcriptase comprises an amino acid sequence selected from the group consisting of SEQ ID NO: 6 (corresponding to AC-RT-M0A), SEQ ID NO: 7 (corresponding to AC-RT-M3A), SEQ ID NO: 8 (corresponding to AC-RT-M4A*), SEQ ID NO: 9 (corresponding to B-RT-M0A), SEQ ID NO: 10 (corresponding to B-RT-M3A), SEQ ID NO: 11 (corresponding to B-RT-M4A*), and SEQ ID NO: 12 (corresponding to bmx-A-RT-M4), SEQ ID NO: 13 (corresponding to bmx-B-RT-M4), SEQ ID NO: 14 (corresponding to bmx-C-RT-M4), SEQ ID NO: 15 (corresponding to bmx-A-RT-M3A), SEQ ID NO: 16 (corresponding to bmx-B-RT-M3A), SEQ ID NO: 17 (corresponding to bmx-C-RT-M3A), SEQ ID NO: 40 (corresponding to bmx-C-RT-M5.1), SEQ ID NO: 41 (corresponding to bmx-C-RT-M5.1+Y63R), SEQ ID NO: 42 (corresponding to bmx-C-RT-M5.1+C408R), SEQ ID NO: 43 (corresponding to bmx-C-RT-M7), and SEQ ID NO: 44 (corresponding to bmx-C-RT-M7A).

[0100] In some embodiments, in the fusion protein, the reverse transcriptase or functional variant thereof is fused to a nucleocapsid protein (NC) at the N-terminus or C-terminus, directly or via a linker. The nucleocapsid protein (NC) is, for example, from PERV. Preferably, the nucleocapsid protein and the reverse transcriptase are from the same type of PERV.

[0101] In some embodiments, the nucleocapsid protein comprises an amino acid sequence as set forth in one of SEQ ID NO: 19 or SEQ ID NO: 36-39.

[0102] In some embodiments, the PERV reverse transcriptase or functional variant thereof is fused to a nucleocapsid protein (NC) at the N-terminus, directly or via a linker.

[0103] In some embodiments, the PERV reverse transcriptase or functional variant thereof is fused to a nucleocapsid protein (NC) at the C-terminus, directly or via a linker.

[0104] As used herein, a "linker" can be a 1-50 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 20-25, 25-50) or more amino acids, non-functional amino acid sequence without secondary structure above. For example, the linker can be a flexible linker, etc. For example, can be an XTEN linker as set forth in SEQ ID NO: 20.

[0105] In some embodiments, the CRISPR nuclease, e.g., CRISPR nickase, in the fusion protein is located N-terminal to the reverse transcriptase and / or nucleocapsid protein. In some embodiments, the CRISPR nuclease, e.g., CRISPR nickase, in the fusion protein is located C-terminal to the reverse transcriptase and / or nucleocapsid protein.

[0106] In some embodiments of the application, the fusion protein of the application can further comprise a nuclear localization sequence (NLS). In general, the NLS(s) in the fusion protein should be of sufficient strength to drive accumulation of the fusion protein in the nucleus of the cell in an amount that enables its base editing function. In general, the strength of the nuclear localization activity is determined by the number, location, specific NLS(es) used, or a combination of these factors, of the NLS(s) in the fusion protein. However, depending on the location of the DNA to be edited, the lead editing fusion protein of the application can also include one or more other localization sequences, e.g., a cytoplasmic localization sequence, a chloroplast localization sequence, a mitochondrial localization sequence, etc.

[0107] In some embodiments, the fusion protein can also include a selectable marker protein linked directly or through a self-cleaving peptide.

[0108] As used herein "self-cleaving peptide" means a peptide that can achieve self-cleavage within a cell. For example, the self-cleaving peptide can comprise a protease recognition site, such that it is recognized and specifically cleaved by a protease within the cell.

[0109] Alternatively, the self-cleaving peptide can be a 2A polypeptide. 2A polypeptides are a class of short peptides from viruses, which self-cleavage occurs during translation. When two different proteins of interest are expressed in the same reading frame with a 2A polypeptide, the two proteins of interest are generated in almost 1:1 ratio. A commonly used 2A polypeptide can be P2A from porcine techovirus-1, T2A from Thosea asigna virus, E2A from equine rhinitis A virus, and F2A from foot-and-mouth disease virus. Among them, P2A has the highest cleavage efficiency and is therefore preferred. A variety of functional variants of these 2A polypeptides are also known in the art, which can also be used in the present application.

[0110] An exemplary selectable marker protein is, for example, a Puro-R protein. The Puro-R protein can provide puromycin resistance to select positive cells. An exemplary Puro-R protein comprises the amino acid sequence set forth in SEQ ID NO: 35.

[0111] In some embodiments, the fusion protein comprises an amino acid sequence set forth in one of SEQ ID NOs: 22-34 and 45-49.

[0112] The guide sequence (also referred to as seed sequence or spacer sequence) in at least one pegRNA of the present application is arranged to have sufficient sequence identity (preferably 100% identity) with the target sequence, so as to be able to bind to the complementary strand of the target sequence through base pairing, achieving sequence-specific targeting.

[0113] A variety of scaffold sequences suitable for gRNAs for CRISPR nuclease (e.g. Cas9)-based genome editing are known in the art, which can be used in the pegRNAs of the present application. In some embodiments, the scaffold sequence of the gRNA is set forth in SEQ ID NO: 21.

[0114] In some embodiments, the primer binding sequence (PBS) is arranged to be complementary to at least a portion of the target sequence (preferably fully paired to at least a portion of the target sequence), preferably the primer binding sequence is complementary to (preferably fully paired to) at least a portion of a 3' overhang single strand resulting from the nick in the DNA strand in which the target sequence is located, in particular to the nucleotide sequence of the 3' end of the 3' overhang single strand. When the 3' overhang single strand of the strand is bound to the primer binding sequence by base pairing, the 3' overhang single strand can serve as a primer for reverse transcription of a reverse transcription template (RTT) sequence immediately adjacent to the primer binding sequence under the action of a reverse transcriptase or a functional variant thereof in the fusion protein, extending a DNA sequence corresponding to the reverse transcription (RT) template sequence.

[0115] The primer binding sequence depends on the length of the overhang single strand formed in the target sequence by the CRISPR nickase used, however, it should have a minimum length that ensures specific binding. In some embodiments, the primer binding sequence can be 4-20 nucleotides in length, for example 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides in length.

[0116] The RT template sequence described herein can be any sequence. By reverse transcription as described above, its sequence information can be integrated into the DNA strand in which the target sequence is located (i.e. the strand containing the PAM of the target sequence), and then by the DNA repair action of the cell, a DNA double strand containing the sequence information of the RT template sequence is formed. In some embodiments, the RT template sequence contains a desired modification. For example, the desired modification includes substitution, deletion and / or addition of one or more nucleotides. For example, the modification includes one or more substitutions selected from the group consisting of C to T substitution, C to G substitution, C to A substitution, G to T substitution, G to C substitution, G to A substitution, A to T substitution, A to G substitution, A to C substitution, T to C substitution, T to G substitution, T to A substitution; and / or deletion of one or more nucleotides, for example 1 to about 100 or more, for example 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotides; and / or insertion of one or more nucleotides, for example 1 to about 100 or more, for example 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotides.

[0117] In some embodiments, the RT template sequence is configured to correspond to (e.g., be complementary to at least a portion of) a sequence downstream of a target sequence nick, and comprises a desired modification. The desired modification includes one or more nucleotide substitutions, deletions, and / or additions. For example, the modification includes one or more substitutions selected from the group consisting of: a C to T substitution, a C to G substitution, a C to A substitution, a G to T substitution, a G to C substitution, a G to A substitution, an A to T substitution, an A to G substitution, an A to C substitution, a T to C substitution, a T to G substitution, a T to A substitution; and / or a deletion of one or more nucleotides, e.g., 1 to about 100 or more, e.g., 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotides; and / or an insertion of one or more nucleotides, e.g., 1 to about 100 or more, e.g., 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotides.

[0118] In some embodiments, the RT template sequence can be about 1-300 or more nucleotides in length, e.g., 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100, about 125, about 150, about 175, about 200, about 225, about 250, about 275, about 300 nucleotides or more in length. Preferably, the RT template sequence is 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 nucleotides in length.

[0119] In some embodiments, the prime editing system further comprises a nicking gRNA (for making additional nicks) and / or an expression construct containing a nucleotide sequence encoding the nicking gRNA, the nicking gRNA comprising a prime sequence and a scaffold sequence. In some preferred embodiments, the nicking gRNA does not comprise a reverse transcription (RT) template sequence and a primer binding site (PBS) sequence.

[0120] The guide sequence (also referred to as seed sequence or spacer sequence) in the nicking gRNA of the present application is configured to have sufficient sequence identity (preferably 100% identity) to the nicking target sequence in the genome, such that the fusion protein of the present application can be targeted to the nicking target sequence and cause a nick within the nicking target sequence, which is on the opposite strand of the genomic DNA from the target sequence targeted by the pegRNA (the pegRNA target sequence). In some embodiments, the nick formed by the nicking RNA and the nick formed by the pegRNA are separated by about 1 to about 300 or more nucleotides, e.g., 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100, about 125, about 150, about 175, about 200, about 225, about 250, about 275, about 300 nucleotides or more. In some embodiments, the nick formed by the nicking RNA is upstream or downstream of the nick formed by the pegRNA (both upstream and downstream are with reference to the DNA strand where the pegRNA target sequence is located). In some embodiments, the guide sequence in the nicking gRNA has sufficient sequence identity (preferably 100% identity) to the opposite strand (the modified strand) of the pegRNA target sequence after the editing event occurs, such that the nicking gRNA only targets the nicking target sequence that is created after the pegRNA-induced target sequence targeting and modification is completed. In some embodiments, the PAM of the nicking target sequence is located within the complement of the pegRNA target sequence.

[0121] In some embodiments, the prime editing system comprises at least one pair of pegRNAs and / or an expression construct comprising a nucleotide sequence encoding the at least one pair of pegRNAs. In some embodiments, the two pegRNAs in the pair of pegRNAs are configured to target different target sequences on the same strand of the genomic DNA. In some embodiments, the two pegRNAs in the pair of pegRNAs are configured to target target sequences on different strands of the genomic DNA. In some embodiments, the PAM of the target sequence of one pegRNA in the pair of pegRNAs is located on the sense strand, while the PAM of the other pegRNA is located on the anti-sense strand. In some embodiments, the induced nicks by the two pegRNAs are located on the two sides of the site to be modified, respectively. In some embodiments, the induced nick by the pegRNA targeting the sense strand is located upstream (5’ direction) of the site to be modified, and the induced nick by the pegRNA targeting the anti-sense strand is located downstream (3’ direction) of the site to be modified. The upstream or downstream is with reference to the sense strand. In some embodiments, the induced nicks by the two pegRNAs are separated by about 1 to about 300 or more nucleotides, e.g., 1-15 nucleotides.

[0122] In some embodiments, the two pegRNAs in the pair of pegRNAs are configured to introduce the same desired modification. For example, one pegRNA is configured to introduce a substitution of A to G at the sense strand, and the other pegRNA is configured to introduce a substitution of T to C at the corresponding position of the antisense strand. For another example, one pegRNA is configured to introduce a deletion of two nucleotides at the sense strand, and the other pegRNA is configured to introduce a deletion of two nucleotides at the corresponding position of the antisense strand. Other types of modifications can be similarly configured. The pegRNAs that target two different strands respectively can be configured to introduce the same desired modification by designing appropriate RT template sequences.

[0123] In some preferred embodiments, the pegRNAs of the present application are epegRNAs. The construction of epegRNAs can be referred to James W. Nelson et al., Engineered pegRNAs improve prime editing efficiency. Nature Biotechnology volume 40, 402-410 (2022), which is incorporated herein by reference. In some embodiments, the epegRNAs are epegRNAs with 3’-tevopre Q1-8 nt linker modification.

[0124] In some preferred embodiments, the pegRNAs of the present application are Csy4-pegRNAs. The Csy4-pegRNAs comprise a Csy4 recognition sequence, such as the Csy4 recognition sequence set forth in SEQ ID NO: 50, at the 3’ end of the pegRNA.

[0125] In some preferred embodiments, the pegRNAs of the present application are tCsy4-pegRNAs. The tCsy4-pegRNAs comprise a truncated Csy4 recognition sequence, such as the truncated Csy4 recognition sequence set forth in SEQ ID NO: 51, at the 3’ end of the pegRNA.

[0126] In order to obtain efficient expression in different organisms, in some embodiments of the present application, the nucleotide sequence encoding the fusion protein is codon-optimized for the species of the organism whose genome is to be modified.

[0127] Codon optimization refers to methods of modifying nucleic acid sequences to enhance expression in host cells of interest by replacing at least one codon in the natural sequence (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more) with codons that are more frequently or most frequently used in the gene in the host cell, while maintaining the natural amino acid sequence. Different species exhibit specific preferences for certain codons of specific amino acids. Codon preference (differences in codon use between organisms) is often associated with the translation efficiency of messenger RNA (mRNA), which is thought to depend on the nature of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The dominance of selected tRNAs in a cell generally reflects the codons most frequently used for peptide synthesis. Therefore, genes can be tailored to achieve optimal gene expression in a given organism based on codon optimization. Codon utilization tables are readily available, for example, in the Codon Usage Database (“Codon Usage Database”) available at www.kazusa.orjp / codon / , and these tables can be adapted in various ways. See, Nakamura Y. et al., "Codon usage tabulated from the international DNA sequence databases: status for the year 2000. Nucl. Acids Res., 28:292 (2000).

[0128] In another aspect, the present invention also provides a truncated Csy4 recognition sequence element, the nucleotide sequence of which is shown in SEQ ID NO:51. In another aspect, the present invention also provides a nucleic acid molecule, such as an RNA molecule, comprising the truncated Csy4 recognition sequence element. The RNA molecule is, for example, gRNA, preferably pegRNA. The gRNA and / or pegRNA are as defined above. In another aspect, the present invention also provides the use of the truncated Csy4 recognition sequence element in gene editing.

[0129] IV. Methods for Modifying Cellular Genome Sequences

[0130] In another aspect, the present application provides a method of modifying a genomic sequence of a cell, comprising introducing the prime editing system of the present application into at least one of said cells, thereby causing a modification of a genomic sequence of said at least one cell. The modification comprises one or more substitutions, deletions, and / or additions of nucleotides. For example, the modification comprises one or more substitutions selected from the group consisting of: a C to T substitution, a C to G substitution, a C to A substitution, a G to T substitution, a G to C substitution, a G to A substitution, an A to T substitution, an A to G substitution, an A to C substitution, a T to C substitution, a T to G substitution, a T to A substitution; and / or a deletion of one or more nucleotides, for example, 1 to about 100 or more, for example, 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotides; and / or an insertion of one or more nucleotides, for example, 1 to about 100 or more, for example, 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotides. The modification can be located in the target sequence of the pegRNA or located near the target sequence, for example, downstream of the target sequence.

[0131] In some embodiments, the prime editing system is introduced into the cell in the presence of one or more small molecule compounds selected from the group consisting of RS-1, Nocodazole (Noc), and Trichostatin A (TSA), preferably a combination of RS-1 and TSA or Nocodazole, more preferably a combination of RS-1 and Nocodazole.

[0132] In some embodiments, after introducing the prime editing system into the cell, the cell is treated with one or more small molecule compounds selected from the group consisting of RS-1, Nocodazole (Noc), and Trichostatin A (TSA), preferably a combination of RS-1 and TSA or Nocodazole, more preferably a combination of RS-1 and Nocodazole, for example, culturing the cell in a medium comprising the one or more small molecule compounds, preferably a combination of RS-1 and TSA or Nocodazole, more preferably a combination of RS-1 and Nocodazole.

[0133] The concentration of Nocodazole can be about 50 ng / ml to about 500 ng / ml, such as about 100 ng / ml to about 300 ng / ml, preferably about 200 ng / mL. The concentration of RS-1 can be about 5 mM to about 50 mM, such as about 10 mM to about 30 mM, preferably about 20 mM. The concentration of Trichostatin A can be about 1 nM to about 50 nM, such as about 5 nM to about 20 nM, preferably about 10 nM.

[0134] In another aspect, the present application also provides a method of producing a genetically modified cell, comprising introducing the prime editing system of the present application into the cell.

[0135] The genetic modification comprises one or more substitutions, deletions, and / or additions of nucleotides. For example, the genetic modification comprises one or more substitutions selected from the group consisting of C to T substitution, C to G substitution, C to A substitution, G to T substitution, G to C substitution, G to A substitution, A to T substitution, A to G substitution, A to C substitution, T to C substitution, T to G substitution, T to A substitution; and / or a deletion of one or more nucleotides, such as 1 to about 100 or more, such as 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotides; and / or an insertion of one or more nucleotides, such as 1 to about 100 or more, such as 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotides. The modification can be located in the target sequence of the pegRNA or located near the target sequence, such as downstream of the target sequence.

[0136] In some embodiments, the prime editing system is introduced into the cell in the presence of one or more small molecule compounds selected from the group consisting of RS-1, Nocodazole (Noc), and Trichostatin A (TSA), preferably RS-1 and Nocodazole.

[0137] In some embodiments, after the prime editing system is introduced into the cell, the cell is treated with one or more small molecule compounds selected from the group consisting of RS-1, Nocodazole (Noc), and Trichostatin A (TSA), preferably RS-1 and Nocodazole, such as culturing the cell in a medium comprising the one or more small molecule compounds, preferably RS-1 and Nocodazole.

[0138] In another aspect, the present application also provides a genetically modified organism comprising the genetically modified cell produced by the method of the present application or a progeny cell thereof.

[0139] In the present application, the modification can be located at any position of the genome, for example, within a functional gene such as a protein-coding gene, or for example, can be located in a gene expression regulatory region such as a promoter region or an enhancer region, thereby achieving modification of the function of the gene or modification of the gene expression. The modification in the genome sequence of the cell can be detected by T7EI, PCR / RE or sequencing method.

[0140] In the method of the present application, the prime editing system can be introduced into the cell by various methods well known to those skilled in the art.

[0141] Methods that can be used to introduce the prime editing system of the present application into a cell include, but are not limited to, calcium phosphate transfection, protoplast fusion, electroporation, lipofection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus and other viruses), gene gun method, PEG-mediated protoplast transformation, Agrobacterium-mediated transformation.

[0142] The cell that can be subjected to gene editing by the method of the present application can be from, for example, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, cats; poultry such as chickens, ducks, geese; plants, including monocotyledons and dicotyledons, for example, rice, corn, wheat, sorghum, barley, soybean, peanut, Arabidopsis thaliana, etc. In some preferred embodiments, the cell is from a human. In some preferred embodiments, the cell is from a pig.

[0143] In some embodiments, the method of the present application is performed in vitro. For example, the cell is an isolated cell, or a cell in an isolated tissue or organ.

[0144] In other embodiments, the method of the present application can also be performed in vivo. For example, the cell is a cell in an organism, and the system of the present application can be introduced into the cell in vivo by, for example, viral or Agrobacterium-mediated methods.

[0145] In another aspect, the present application provides the use of one or more small molecule compounds selected from RS-1, Nocodazole (Noc) and Trichostatin A (TSA), preferably a combination of RS-1 and TSA or Nocodazole, more preferably a combination of RS-1 and Nocodazole, for prime editing, in particular for prime editing in mammalian cells.

[0146] V. Therapeutic applications

[0147] The present application also encompasses the use of the present guide-editing system in the treatment of a disease.

[0148] Modifying a disease-associated gene by the present guide-editing system can achieve up-regulation, down-regulation, inactivation, activation or mutation correction of the disease-associated gene, thereby achieving the prevention and / or treatment of the disease. For example, the genomic modification described in the present application can be located within the protein coding region of the disease-associated gene, or for example, can be located in the gene expression regulatory region such as the promoter region or enhancer region, thereby achieving the modification of the function of the disease-associated gene or the modification of the expression of the disease-associated gene. Therefore, the modification of the disease-associated gene described herein includes the modification of the disease-associated gene itself (e.g. the protein coding region), and also includes the modification of the expression regulatory region (such as the promoter, enhancer, intron, etc.) thereof.

[0149] A "disease-associated" gene is any gene that produces a transcriptional or translational product at an abnormal level or in an abnormal form in a cell derived from a tissue affected by a disease, as compared to a non-disease control tissue or cell. It can be a gene that is expressed at an abnormally high level in cases where altered expression is associated with the presence and / or progression of the disease; it can be a gene that is expressed at an abnormally low level. A disease-associated gene also refers to a gene having one or more mutations or genetic variations that are in linkage disequilibrium with one or more genes directly responsible for the etiology of the disease. The mutation or genetic variation is, for example, a single nucleotide variation (SNV). The transcriptional or translational product can be known or unknown, and can be at a normal or abnormal level.

[0150] Accordingly, the present application also provides a method of treating a disease in a subject in need thereof, comprising delivering to the subject an effective amount of the present guide-editing system to modify a gene associated with the disease. The present application also provides the use of a guide-editing system in the manufacture of a pharmaceutical composition for treating a disease in a subject in need thereof, wherein the guide-editing system is used to modify a gene associated with the disease. The present application also provides a pharmaceutical composition for treating a disease in a subject in need thereof, comprising a guide-editing system of the present application, and optionally a pharmaceutically acceptable carrier, wherein the guide-editing system is used to modify a gene associated with the disease.

[0151] Preferably, the pharmaceutical composition of the present application further comprises one or more small molecule compounds selected from the group consisting of RS-1, Nocodazole, and Trichostatin A (TSA) (preferably a combination of RS-1 and TSA or Nocodazole, more preferably a combination of RS-1 and Nocodazole). In other preferred embodiments, the pharmaceutical composition of the present application is for use in combination with one or more small molecule compounds selected from the group consisting of RS-1, Nocodazole, and Trichostatin A (TSA) (preferably a combination of RS-1 and TSA or Nocodazole, more preferably a combination of RS-1 and Nocodazole).

[0152] Preferably, the "subject" of the present application is a mammal, such as a human, mouse, rat, monkey, dog, pig, sheep, cow, or cat, preferably a human.

[0153] In some embodiments, the prime editing system described herein is used to introduce a point mutation into a nucleic acid.

[0154] In some embodiments, the prime editing system described herein is used to cause correction of a genetic deficiency, such as in correction of a point mutation that results in loss of function in a gene product. In some embodiments, the genetic deficiency is associated with a disease or disorder, such as a lysosomal storage disease or a metabolic disease, such as Type I diabetes. In some embodiments, the methods provided herein can be used to introduce an inactivating point mutation into a gene or allele that encodes a gene product associated with a disease or disorder.

[0155] In some embodiments, the aim of the regimen described herein is for the treatment of a disease associated with or caused by a point mutation that can be corrected by the prime editing system provided herein. In some embodiments, the disease is a proliferative disease. In some embodiments, the disease is a genetic disease. In some embodiments, the disease is a de novo disease. In some embodiments, the disease is a metabolic disease. In some embodiments, the disease is a lysosomal storage disease.

[0156] In some embodiments, the aim of the regimen described herein is for the treatment of a mitochondrial disease or disorder. As used herein, "mitochondrial disease" relates to a disease caused by abnormal mitochondria, such as mitochondrial gene mutations, enzyme pathways, etc. Examples of diseases include, but are not limited to, neurological disease, loss of motor control, muscle weakness and pain, gastrointestinal disease and difficulty swallowing, poor growth, heart disease, liver disease, diabetes, respiratory complications, seizures, vision / hearing problems, lactic acidosis, developmental delays, and susceptibility to infection.

[0157] Examples of the diseases described in the present application include, but are not limited to, genetic diseases, circulatory system diseases, muscle diseases, brain, central nervous system and immune system diseases, Alzheimer's disease, secretase disorders, amyotrophic lateral sclerosis (ALS), autism, trinucleotide repeat expansion disorders, hearing diseases, gene targeting therapy of non-dividing cells (neurons, muscle), liver and kidney diseases, epithelial cell and lung diseases, cancer, Usher syndrome or retinitis pigmentosa-39, cystic fibrosis, HIV and AIDS, beta thalassemia, sickle cell disease, herpes simplex virus, autism, drug addiction, age-related macular degeneration, schizophrenia. Other diseases treated by correcting point mutations or introducing inactivating mutations into disease-associated genes are known to those skilled in the art, and thus the present disclosure is not limited in this respect. In addition to the diseases exemplarily described in the present application, other related diseases can also be treated using the strategy and prime editing system provided by the present application, which is obvious to those skilled in the art.

[0158] The administration of the prime editing system or pharmaceutical composition of the present application can be adjusted according to the weight and species of the patient or subject. The administration frequency is within the range permitted in medicine or veterinary medicine. It depends on conventional factors including the age, gender, general health status, other conditions of the patient or subject, and the specific disease or symptom to be addressed.

[0159] Six, kit

[0160] The present application also includes a kit for use in the methods of the present application, which comprises the components of the prime editing system of the present application. The kit can also contain reagents for introducing the prime editing system into an organism or organism cell. The kit generally includes a label indicating the intended use and / or method of use of the contents of the kit. The term label includes any written, printed, or recorded material that is provided with or is otherwise associated with the kit. Examples

[0161] Example 1, PERV-RT sequence acquisition and analysis.

[0162] PERV sequence acquisition: PERV-A / C (accession number: AY953542) and PERV-B (accession number: AY099324) sequences were obtained through the NCBI database, the positions and sequences of NC, RT and RNase H in the PERV sequence were analyzed using the conserved domain database, and PERV-AC-NC, PERV-B-NC and PERV-AC-RT (ΔRNase H) and PERV-B-RT (ΔRNase H) were obtained.

[0163] Sequence analysis: Transform the DNA sequences of MMLV-RT (ΔRNase H), PERV-AC-RT (ΔRNase H) and PERV-B-RT (ΔRNase H) into protein sequences, and compare the differences between MMLV-RT (ΔRNase H), PERV-AC-RT (ΔRNase H) and PERV-B-RT (ΔRNase H) using MegAlign (V7.1.0) software. Figure 1 ) The results show that PERV-AC-RT is highly similar to PERV-B-RT, with only 11 amino acid differences; the similarity between PERV-RT and MMLV-RT is only 70%, and there are 153 and 150 amino acid differences between PERV-AC and PERV-B and MMLV-RT, respectively. It is worth noting that among the 4 mutation sites of MMLV-RT in PE2, 3 are homologous to PERV-RT (red box), and one site has no homology, and the continuous 4 bases of this site are not homologous to PERV-RT (black box). Figure 1 Figure 1

[0164] Example 2, Construction of Enhanced PERV-PE Editor (epPE)

[0165] 2.1 Upgrade PE2 editor to obtain PE2c-Puro

[0166] In order to enhance the editing efficiency of the lead editor in mammalian cells, the PE2 editor is first upgraded. First, the PE2 plasmid is purchased from the American Addgene plasmid sharing platform (number: #132775), and the online codon optimization tool (https: / / www.genscript.com.cn / codon-opt.html) is used to specifically optimize the codons in PE2 MMLV-RT according to the pig genome, obtaining the MMLV-RT (condon opt) sequence, and the sequence is synthesized by Nanjing Kingsriver Biotechnology Co., Ltd. and replaced with the MMLV-RT sequence of the PE2 vector to obtain the PE2c vector. In addition, the 2A-Puro sequence is synthesized and inserted into the C-terminal of the RT structure of the PE2c vector, and the obtained vector is named PE2c-Puro. Figure 2

[0167] 2.2 Construction of epPE editor.

[0168] 2.2.1 Construction of epPE-AC / B-m3-Puro

[0169] ​​​The PERV-AC-RT (ΔRNase H) and PERV-B-RT (ΔRNase H) sequences were specifically codon-optimized according to the pig genome using an online codon optimization tool (https: / / www.genscript.com.cn / codon-opt.html), and the corresponding mutations were made for the three mutant sites (M3) homologous to MMLV-RT in PE2 (D200N, T306K, W313F in MMLV-RT, corresponding to D199N, T305K, W312F in PERV-RT), respectively, to obtain PERV-AC-RT-M3Δ and PERV-B-RT-M3Δ sequences. PERV-AC-RT-M3Δ and PERV-B-RT-M3Δ sequences were synthesized by Nanjing Kingsriver Biotechnology Co., Ltd. and replaced MMLV-RT in PE2c-Puro to obtain epPE-AC-M3Δ-Puro and epPE-B-M3Δ-Puro, respectively. Figure 2 ).

[0170] 2.2.2 Construction of epPE-AC / B-m4-Puro

[0171] The 329-332 amino acids of the RT in epPE-AC-M3Δ-Puro and epPE-B-M3Δ-Puro were mutated from EKGE to PGTL (DNA sequence was mutated from GAGAAGGGTGAA / GAGAAGGGAGAG to CCAGGCACCCTG) by Nanjing Kingsriver Biotechnology Co., Ltd. to obtain epPE-AC-M4Δ*-Puro and epPE-B-M4Δ*-Puro. Figure 2 ).

[0172] Example 3, Test the efficiency of different epPE in immortalized BFP fluorescent reporter pig fetal fibroblast (PFF) cells

[0173] 3.1 Construction of pegRNA and ngRNA

[0174] The pegRNA vector backbone pegRNA-GG (No. 132777) and ngRNA vector backbone (No. 107721) were purchased from the American Addgene plasmid sharing platform, and the vector plasmid DNA was extracted using the Takata MidiBest Kit (Takara, China) kit.

[0175] Vector enzyme digestion: 1x CutSmart buffer (Neb, USA), 1 μL of BsaI endonuclease (Neb, USA) and 2 μg of pegRNA-GG or ngRNA empty plasmid DNA in a 30 μL reaction system, placed in a PCR instrument for 2 hours of enzyme digestion at 37℃, and 15 minutes of inactivation at 80℃. The enzyme digestion product was separated by 1.5% agarose gel to obtain the target fragment, and after cutting the gel, Axypre DNA Gel Extraction Kit (Axygen, USA) was used for purification.

[0176] pegRNA-gRNA, pegRNA 3'-extension, sgRNA scaffold, ngRNA forward and reverse primer annealing into double-stranded: 100 μM / L of forward primer (F) and reverse primer (R), double distilled water and 1x T4 Ligation Buffer (Neb, USA) in a 25 μL reaction system; the EP tube containing the reaction system was sealed and placed in a beaker containing boiling water, and slowly cooled at room temperature.

[0177] pegRNA ligation: the ligation system is shown in Table 1. After the system was prepared, it was placed at room temperature for 10 minutes, then transferred to a PCR instrument for 5 minutes at 16℃, 15 minutes at 37℃ (8 cycles), and finally inactivated at 80℃ for 15 minutes.

[0178] Table 1 pegRNA ligation system

[0179]

[0180] ngRNA ligation: 1 μL of annealed ngRNA was subjected to 10 μL system under the action of Quick Ligase (Neb, USA) for 15 minutes at room temperature, and then stored in a refrigerator at -20℃.

[0181] Preparation of high-purity plasmid. 5 μL of ligated pegRNA / ngRNA was transformed into DH5α according to the steps in the instructions 5α competent cells (Qingke Biotechnology, China) were transformed and spread onto solid LB agar plates (each LB solid medium contained 10g peptone, 5g yeast extract, 10g NaCl, 15g agar powder, and 30mg ampicillin) and cultured overnight at 37°C. Six clones of each pegRNA / ngRNA were picked and expanded in 1mL liquid LB medium. 500μL of the expanded culture was sent to Qingke Biotechnology Co., Ltd. for Sanger sequencing (sequencing primers: U6-F: atggactatcatatgcttaccgta). The correctly ligated clones were added to 200mL liquid LB medium and cultured overnight at 37°C on a shaker. High-purity, endotoxin-free plasmid DNA was extracted using the Takata MidiBest Kit (Takara, China).

[0182] 3.2 Cell culture, electroporation, and puromycin screening.

[0183] 3.2.1 Cell Culture

[0184] Immortalized BFP fluorescent reporter PFF cell lines were cultured in high-glucose DMEM (Gibco, USA) medium containing 12% fetal bovine serum ((v / v), Excell, China) and 1% penicillin / streptomycin ((v / v), BI, China). This cell line was constructed in our laboratory by first knocking the SV40 large T antigen into the H11 safe harbor site of primary Bama fungus PFFs using CRISPR-Cas9 technology, then integrating the BFP vector into the Rosa26 safe harbor site of the immortalized PFFs using CRISPR-Cas9 technology, and finally obtaining the BFP fluorescent reporter PFF cell line through continuous passage.

[0185] 3.2.2 Electro-rotation

[0186] When cell confluence reaches 85-90%, digest the cells using 0.125% trypsin (Gibco, USA) and count them using a hemocytometer. Approximately 4 × 10⁶ cells are collected from each electroporation reaction. 5 Cells were centrifuged in 15 mL centrifuge tubes at 200 × g for 5 minutes. After discarding the supernatant, 20 μL of Entranster-E electroporation buffer (Entranster, China), 1 μg PE editor, 0.5 μg pegRNA, and 0.5 μg ngRNA were added. The mixture was then transferred to 16-well electroporation strips (Lonza, Germany) and placed in the electroporation strip holder of a 4D-Nucleofector™ electroporator. PFF cells were electroporated twice consecutively using the CA137 program. After electroporation, the cells were cultured in 6-well cell culture dishes. The culture medium was changed 7 hours after electroporation to remove dead cells caused by electroporation.

[0187] 3.2.2 Puromycin drug screening

[0188] 48 hours after electroporation, 2 μg / mL puromycin (Sigma, Japan) was added for screening for 1 day, followed by changing the culture medium and continuing to culture for 3 days.

[0189] 3.3 Editing Efficiency Test

[0190] After digesting the cells with 0.125% trypsin (Gibco, USA), the cells were collected into 1.5 mL EP tubes, centrifuged at 200×g for 5 minutes, the supernatant was discarded, and the cells were resuspended in 50 μL of 1×PBS buffer. The cells were then detected using a flow cytometer (BD, USA). With wild-type fluorescent reporter cells as the background, GFP-positive cells were circled using a FITC+PE dual-channel method. The GFP ratio is the editing efficiency.

[0191] Flow cytometry analysis showed that the efficiency of the four epPE editors derived from PERV-RT was significantly higher (p < 0.01) than that of the PE2c-Puro editor derived from mouse RT. Furthermore, the efficiency of epPE-AC-M3Δ-Puro was higher than that of epPE-B-M3Δ-Puro (p = 0.007), and the efficiency of epPE-AC-M4Δ*-Puro was higher than that of epPE-B-M4Δ*-Puro (p = 0.012). Figure 3 Example 4: Testing the efficiency of AC and B type PERV-RT epPE in primary PFF cells.

[0192] 4.1 Obtaining primary PFF cells

[0193] After euthanizing the sows at 31 days of gestation, the entire uterus was removed, and each fetus, along with its membranes, was extracted and transferred to DMEM solution containing 5% (v / v) penicillin and streptomycin (v / v; Gibco: 15140-122) and stored at 4°C. In a laminar flow hood, each separated fetus was transferred to a sterile 10cm culture dish. The head, internal organs, limbs, and tail of the fetus were removed using sterile forceps and scissors. The remaining parts of the fetus were cleaned with PBS solution containing 5% (v / v) penicillin and streptomycin (Gibco: 15140-122) and then minced to approximately 1mm within 10-15 minutes. 3In the minced tissue, 5 ml of collagenase working solution was added; 1 ml of sterile pipette tip was cut with scissors, and the minced tissue was transferred to a 15 ml sterile centrifuge tube, which was sealed with parafilm and incubated in a 37°C shaking incubator for 2 hours. Then, the centrifuge tube was centrifuged at 1000 rpm for 5 minutes, and the upper collagenase working solution was removed. The incubated tissue pieces were evenly transferred to two T75 culture dishes and cultured in a sterile incubator at 37°C and 5% CO2. The culture was continued until the next morning (Note: After about 20 hours of starting cell culture, if the color of the cell culture solution turns yellow, new culture solution needs to be replaced). The adhesion and separation of cells from the cell tissue were observed under a microscope. When the cell proliferation and growth reached more than 85%, the construction of the porcine primary PFF cell line was completed, and it was frozen or subcultured.

[0194] 4.2 Primary PFF cell culture, electroporation and puromycin screening

[0195] The cells were cultured in high-glucose DMEM (Gibco, USA) containing 12% (v / v) fetal bovine serum (Excell, China) and 1% (v / v) penicillin-streptomycin (BI, China) at 37°C in a 5% CO2 incubator. When the cell confluence reached 85-90%, the cells were digested with 0.125% trypsin (Gibco, USA) and counted using a hemocytometer. For each electroporation reaction, 4 x 10^5 cells were taken in a 15 mL centrifuge tube, centrifuged at 200 x g for 5 minutes, and the supernatant was removed. Then, 20 μL of Entranster-E electroporation solution (Ingeng, China), 1 μg of PE editor, 0.5 μg of pegRNA, and 0.5 μg of ngRNA were added, mixed well, and the cell suspension was transferred to a 16-well electroporation strip (Lonza, Germany). The cells were placed in the electroporation strip holder of a 4D-NucleofectorTM electroporator and subjected to two consecutive electric shocks using the CA-137 program. The electroporated cells were cultured in a 6-well cell culture dish. The culture medium was replaced 7 hours after electroporation to remove dead cells caused by electric shock. Two μg / mL of puromycin (Sigma, Japan) was added 48 hours after electroporation and screened for 1 day, followed by replacement of the culture medium for 3 days of continuous culture.

[0196] 4.3 Editing efficiency detection

[0197] 4.3.1 Cell DNA acquisition

[0198] The cells were collected in a 200 μL EP tube after digestion with 0.125% trypsin, centrifuged at 750 x g for 5 minutes, and the supernatant was removed using a vacuum aspirator. 12 μL of cell lysis solution (0.45% NP40 lysis solution (Gibco, USA), 6% proteinase K (Sigma, Japan)) was added to resuspend the cells, which were placed in a PCR instrument with the following reaction program: 56°C for 1.5 hours and 85°C for 15 minutes.

[0199] 4.3.2 Target site amplification.

[0200] Two rounds of PCR amplification were performed on the target site using appropriate primers. The first round of PCR amplified the target site, and the second round of PCR labeled the sample with tagged primers. PrimeSTAR Max DNAPolymerase (Takara, Japan) was used for both rounds of PCR. Each reaction contained 8 μL PrimeSTAR Max DNAPolymerase buffer, 1 μL each of forward and reverse primers, 1 μL of DNA template, and 5 μL of ultrapure water. PCR reaction conditions were: 94℃ for 5 minutes; (94℃ for 30 seconds, 68℃ (decreasing by 0.5℃ per cycle) for 30 seconds, 72℃ for 1 minute) for 26 cycles; (94℃ for 30 seconds, 55℃ for 30 seconds, 72℃ for 1 minute) for 14 cycles; 72℃ for 5 minutes.

[0201] 4.3.3 Sample Mixing and Gel Purification. Amplification products with different tags were mixed in groups of 8-12 samples, with 5 μL of each sample. After adding loading buffer to the mixed samples, electrophoresis was performed on a 2% agarose gel at 180V for 20 minutes. The target fragment was then cut from the gel using a UV cutting platform and purified using the Axypre DNAGel Extraction Kit (Axygen, USA) according to the manufacturer's instructions. Finally, the DNA was dissolved in ultrapure water.

[0202] 4.3.4 High-throughput sequencing. 2 μg of purified PCR product was sent to Beijing Novogene Technology Co., Ltd. for library construction and sequencing. First, high-throughput sequencing was performed using… Ultra TM The DNA Library Prep Kit for Illumina was used to add A-tails to the DNA and ligate it with Illumina sequencing adapters. The DNA was then purified using the AMPure XP system and quantified and diluted using Qubit 3.0. Insert fragments in the library were detected using an Agilent 2100, and for libraries with insert sizes matching the expected values, the effective concentration was accurately quantified using Q-PCR. Finally, paired-end sequencing was performed using the PE150 platform of an Illumina NovaSeq 6000 high-throughput sequencer.

[0203] 4.3.5 Editing Efficiency Analysis

[0204] Firstly, the sequencing data was filtered by Fastp (V0.19.7) software, and the sequences with base quality value Qphred≤5 and N content more than 10% of the read base were filtered out to obtain Clean Data. The mixed sample data was split by fastq-multx (V1.31) software, and each sample would get R1.fastq and R2.fastq. The editing efficiency was analyzed by CRISPResso2 (V2.0.43) software, and each sample was provided with ref_seq, pegRNA_spacer_seq, pegRNA_extension_pegRNA_scaffold_seq and nicking_guide_seq parameters, and the default value was used for other parameters. Among them, ref_seq is the reference sequence amplified by PCR, and pegRNA_extension_pegRNA_scaffold_seq sequence is the complete sgRNA scaffold sequence. For the identification of the final results, Prime-edited UNMODIFILED is considered as accurate editing, and other editing types are considered as unnecessary by-products.

[0205] The results show that the average editing efficiency of epPE editor derived from PERV-RT is higher than that of PE2c-Puro editor derived from murine RT at the 6 editing sites tested in primary PFF, and the average efficiency of epPE-AC-M3Δ-Puro is higher than that of epPE-B-M3Δ-Puro Figure 4 ) at all editing sites.

[0206] Example 5, test the efficiency of epPE in other cells

[0207] 5.1 Cell culture

[0208] Canine fetal fibroblasts (DFF) were isolated from a 30-day pregnant dog fetus, and the specific establishment method was consistent with that of PFF; sheep fetal fibroblasts (SFF) were donated by Wei Hongjiang Laboratory of Yunnan Agricultural University, and the cells were isolated from a small-tail Han sheep fetus; 293T cells were purchased from the Chinese Academy of Sciences Typical Culture Collection Committee Cell Library. The three kinds of cells were placed in 12% (v / v) fetal bovine serum (Excell, China) and 1% (v / v) penicillin-streptomycin (BI, China) in high-glucose DMEM (Gibco, USA) culture medium, 37℃, 5% carbon dioxide incubator for culture.

[0209] 5.2 Cell transfection

[0210] DFF and SFF: the method is similar to the transfection method of PFF, the difference is that for DFF and SFF, after using EN-150 program for electric shock, CA-113 program is used for electric shock again.

[0211] 293T cells: Cells were seeded into 24-well plates 1-2 days before transfection, and when the cells reached 50-60% confluence, transfection was performed using Lipofectamine 3000 reagent (Invitrogen, USA). Each transfection reaction included 1 pg PE editor, 0.5 pg pegRNA, 0.5 pg ngRNA, and 1 pL Lipofectamine 3000 reagent.

[0212] 5.3 Puromycin drug screening

[0213] SFF and DFF cells: 3 pg / mL puromycin (Sigma, Japan) was added 48 hours after electroporation for 2 days of selection, followed by medium change for another 2 days of culture.

[0214] 293T cells: 3 pg / mL puromycin (Sigma, Japan) was added 48 hours after transfection for 3 days of selection, followed by medium change for another 2 days of culture.

[0215] 5.4 Editing efficiency detection

[0216] Methods were consistent with the editing efficiency of PFF.

[0217] As shown in Table 1, the efficiency of epPE editors was higher than that of PE2c-Puro vector of mouse-derived MMLV-RT in SFF, DFF, and 293T and other mammalian cells. Figure 5

[0218] Example 6, Application of RT of PERV isolated from Bama mini-pigs in prime editing

[0219] The present inventors also isolated three types of PERV-RT (bmx-PERV-A-RT-M0, bmx-PERV-B-RT-M0, and bmx-PERV-C-RT-M0) from Bama mini-pig DNA by PCR amplification, replaced MMLV-RT in PE2c-Puro vector respectively, obtained pPE editors of A, B, and C PERV-RT respectively, and constructed PE editors carrying different mutation sites for bmx-PERV-RT; M0 represents wild type without mutation; M4 represents containing 4 mutations (D199N, T305K, W312F, and L602W), as shown in Table 2. Figure 6 The PERV-RT isolated from Bama mini-pigs and the corresponding prime editors are prefixed with “bmx-” to distinguish from PERV-RT and the corresponding prime editors from NCBI.

[0220] The editing activity of different editors in different cell lines was tested as described above.

[0221] ​Figure 7 The pPE editor containing bama pig C-type PERV reverse transcriptase showed higher editing activity in BFP reporter cell line, and after introducing mutations, all three PERV-RTs showed higher editing activity.

[0222] Figure 8 The editing efficiency of bama pig PERV reverse transcriptase pPE editor containing 4 point mutations in PFF is shown, and bmx-pPE-C-M4-puro shows higher overall efficiency.

[0223] Figure 9 The editing efficiency of bama pig PERV reverse transcriptase pPE editor containing 4 point mutations in HEK293T, canine fetal fibroblasts, and sheep fetal fibroblasts is shown, and bmx-pPE-C-M4-puro shows higher efficiency at most sites.

[0224] Figure 10 The editing efficiency of bama pig PERV reverse transcriptase epPE editor containing 4 point mutations in PFF is shown, and compared with type A and type B bmx-epPE, type C bmx-epPE editor shows higher overall efficiency, and is significantly higher than bmx-pPE.

[0225] Figure 11 The editing efficiency of type C bmx-epPE editor (bmx-epPE-C-M3Δ-Puro) in HEK293T, canine fetal fibroblasts, and sheep fetal fibroblasts is shown, and the overall editing efficiency of bmx-epPE-C-M3Δ-Puro in the three cell lines is significantly higher than that of bmx-pPE-C-M4-Puro.

[0226] Example 7, the bama pig PERV-RT guide editing system carrying 7 point mutations has higher efficiency

[0227] Based on the known higher editing efficiency of C-type bmx-PERV-RT in the guide editing system, the inventors screened a large number of mutation sites for C-type bmx-PERV-RT in 293T cells.

[0228] 7.1 Selection of mutation sites

[0229] 7.1.1 Mutation sites reported in MMLV-RT guide editing system and homologous to PERV-RT.

[0230] According to the reports of Anzalone et al. Nature (2019), Ni et al. Genome Biology (2023), the sites that are effective in the MMLV-RT prime editing system and have homology with PERV-RT are selected to obtain six potential beneficial sites of D199N, T305K, W312F, E329P, L602W and V222A.

[0231] 7.1.2 Obtain potential beneficial mutation sites by molecular docking.

[0232] Mutating the polar negatively charged amino acids (D, E) and polar uncharged amino acids (S, T, C, N, Q, Y, G) in RT that bind to DNA or RNA into polar positively charged amino acids (R, K) can enhance the interaction of RT with DNA / RNA and thus improve the efficiency of prime editing. For this purpose, the present application uses Phyre2 software to obtain the predicted protein structure of the bmx-PERV-C-RT-M5.1 variant, and downloads the PDB file of the MMLV-RT-DNA-RNA complex: 4KHQ from the PDB number database, uses PyMOL software to align the bmx-PERV-C-RT-M5.1 variant with 4KHQ, finds out the amino acids of bmx-PERV-C-RT-M5.1 that interact with DNA / RNA, and selects all the polar negatively charged amino acids and the polar uncharged amino acids to be mutated into R, a total of 11 mutation sites: Y43R, Y108R, Q112R, D113R, E116R, D149R, D152R, G190R, N118R, N130R, G269R, T305R, Y324R.

[0233] 7.1.3 Obtain potential beneficial sites according to literature reports

[0234] 7.1.3.1 Select mutation sites that have been reported to improve the performance of MMLV-RT reverse transcription but have not been verified in PE systems, and have homology with PERV-RT, a total of 7 potential mutation sites: S86R, Q112R, Q189F, Q220R, T286A, E332Q, I434R.

[0235] 7.1.3.2 A large number of mutation sites were obtained in the phage-assisted evolution system in the study of Doman et al. Cell (2023), but the efficiency of these mutation sites has not been verified in PE systems, and the present application screens out 9 high-frequency and homologous mutation sites with PERV-RT: T127N, P131S, P195S, K232N, E338N, C408R, D423V, M456I, T457I.

[0236] 7.2 293T cell fluorescent reporter co-transfection experiment

[0237] 1 x 105cells were seeded in 24-well plates 17-24 hours before transfection. 5 Cells were seeded in 24-well plates, and when the confluence reached 50-80%, transfection was performed using H4000 reagent (Ingen, China). Each transfection reaction included 200 ng PE editor, 100 ng pegRNA, 100 ng ngRNA, 10 ng fluorescent reporter vector, and 0.5 μL H4000 reagent. Cells were collected 48 hours after transfection, and the editing efficiency was detected by flow cytometry (BD, USA) using wild-type fluorescent reporter cells as background. The GFP-positive cells were circled using the FITC+PE double channel, and the GFP ratio was the editing efficiency.

[0238] 7.3 293T cell endogenous editing site editing experiment

[0239] 1 x 105cells were seeded in 24-well plates 17-24 hours before transfection. 5 Cells were seeded in 24-well plates, and when the confluence reached 50-80%, transfection was performed using H4000 reagent (Ingen, China). Each transfection reaction included 200 ng PE editor, 100 ng pegRNA, 100 ng ngRNA, 10 ng fluorescent reporter vector, and 0.5 μL H4000 reagent. Cells were collected 48 hours after transfection, and the editing efficiency was detected by flow cytometry (BD, USA) using wild-type fluorescent reporter cells as background. The GFP-positive cells were circled using the FITC+PE double channel, and the GFP ratio was the editing efficiency.

[0240] Figure 12 The efficiency of the 6 reported and homologous to PERV-RT mutation sites and their combinations in the MMLV-RT pilot editing system carried by C-type bmx-PERV-RT in 293T cell fluorescent reporter co-transfection experiments is shown. The data show that the PERV-RT variant carrying 5 amino acid mutations (named M5.1) (D199N, T305K, W312F, L602W, V222A) exhibited the best overall efficiency in the three editing forms.

[0241] Figure 13 The efficiency of the pvPE editor based on M5.1 further superimposed with new potential beneficial mutation sites in 293T cells is shown. Figure 13 A shows the efficiency of a total of 27 new potential beneficial sites in 293T cell fluorescent reporter co-transfection experiments using molecular docking and literature reports, of which 8 sites significantly reduced the PE editing efficiency. Figure 13 B shows the efficiency of the remaining 19 mutation sites in 293T endogenous sites, of which Y63R and C408R exhibited varying degrees of editing efficiency improvement.

[0242] Figure 14 The editing efficiency of M5.1, M5.1+Y63R, M5.1+C408R and their combination M5.1+Y63R+C408R (M7) in 293T cells endogenous sites is shown, and the results show that M7 exhibits superimposed editing efficiency at multiple sites, and the efficiency at all sites is significantly better than M5.1.

[0243] Figure 15 The editing efficiency of M7-based epPE editor in 293T cells is shown, and the efficiency of bmx-epPE-C-M7Δ with 7 point mutations is further improved compared with bmx-pPE-C-M7.

[0244] Figure 16 The efficiency of bmx-epPE-C-M7Δ and the currently most efficient PEmax and PE6c, PE6d editors in 293T cells is shown. The results show that the overall efficiency of bmx-epPE-C-M7Δ is higher than PEmax and PE6c, PE6d.

[0245] Example 8, tCsy4-pegRNA has higher editing efficiency

[0246] The 3' end of the conventional pegRNA is a naked RNA, which is easy to be degraded in cells, thereby affecting the efficiency of the pegRNA. In addition, since PBS is usually 10-16 nt at the 3' end of the pegRNA, and complementary to part of the spacer at the 5' end of the pegRNA, their annealing is expected to cause pegRNA circularization, which can hinder editing. It has been reported that by adding a Csy4 recognition sequence (hairpin structure) to the end of the pegRNA, the efficiency of PE in HKE293T, HeLa and Neuro-2a cells is significantly improved. The inventors also designed a truncated Csy4 (tCsy4, SEQ ID NO: 51). The editing efficiency of tCsy4-pegRNA was compared with Csy4-pegRNA and conventional pegRNA (A). The results show that the editing efficiency of Csy4-pegRNA (5 of 7 sites) and tCsy4-pegRNA (6 of 7 sites) is significantly higher than that of conventional pegRNA, and the editing efficiency of tCsy4-pegRNA is similar to or higher than that of Csy4-pegRNA (B, C). In addition, the efficiency of tCsy4-pegRNA in HEK293T, DFF, SFF was verified, and the data show that tCsy4-pegRNA can further improve the efficiency of PE in the three cell lines (D). Figure 17 A). The results show that the editing efficiency of Csy4-pegRNA (5 of 7 sites) and tCsy4-pegRNA (6 of 7 sites) is significantly higher than that of conventional pegRNA, and the editing efficiency of tCsy4-pegRNA is similar to or higher than that of Csy4-pegRNA (B, C). In addition, the efficiency of tCsy4-pegRNA in HEK293T, DFF, SFF was verified, and the data show that tCsy4-pegRNA can further improve the efficiency of PE in the three cell lines (D). Figure 17 A). The results show that the editing efficiency of Csy4-pegRNA (5 of 7 sites) and tCsy4-pegRNA (6 of 7 sites) is significantly higher than that of conventional pegRNA, and the editing efficiency of tCsy4-pegRNA is similar to or higher than that of Csy4-pegRNA (B, C). In addition, the efficiency of tCsy4-pegRNA in HEK293T, DFF, SFF was verified, and the data show that tCsy4-pegRNA can further improve the efficiency of PE in the three cell lines (D). Figure 17 D).

[0247] Example 9: Small molecules can significantly improve the efficiency of lead editing in mammalian cells.

[0248] It is known that some small molecule compounds can significantly improve the efficiency of CRISPR / Cas9 genome editing. Based on this, it is inferred that some small molecules may also improve the editing efficiency of PE (penetrating genome editing). First, nine small molecules proven effective in the CRISPR / Cas9 system were tested in the fluorescent reporter gene PFF cell line: 10 μM Scr7, 2 μM Nu7441, 25 μM L189, 40 μM L755507, 5 μM Yu238259, 0.1 μM Brefeldin A, 10 nM TSA, 20 μM RS-1, 200 ng / mL Nocodazole, and 25 μM L189. After transfection, the appropriate concentrations of the small molecules or combinations thereof were added to the culture medium to treat the cells.

[0249] Data showed that L755507, BrefeldinA, TSA, Rs-1, and Nocodazole significantly (P<0.05 or P<0.01) improved PE efficiency, with TSA, Rs-1, and Nocodazole being more effective, increasing editing efficiency by 34.4%, 32.7%, and 66.9%, respectively. Figure 18 A). Compared with the use of a single molecule, the combination of Rs-1+TSA or Rs-1+Nocodazole significantly improved (P<0.01) editing efficiency. The Rs-1+Nocodazole group had the highest editing efficiency, which was 106.0% higher than that of the DMSO group. Figure 18 B). The effects of Rs-1, Nocodazole, TSA, and combinations of Rs-1+TSA and Rs-1+Nocodazole in primary PFFs were then verified. Data showed that in primary PFFs using Rs-1, TSA, and Nocodazole, the overall editing efficiency for these six sites was increased by 20.8%, 14.3%, and 36.4%, respectively. The Rs-1+Nocodazole combination increased the editing efficiency by 60.54%, while the Rs-1+TSA combination did not show an additive effect. Figure 18 C). Furthermore, the efficiency of the Rs-1+Nocodazole combination in HEK293T, DFF, and SFF cell lines was verified. Data showed that the Rs-1+Nocodazole combination further improved PE efficiency in all three cell lines. Figure 18 D). Information on small molecule compounds involved in this invention.

[0250] RS-1:

[0251] Selleck part number: S8234

[0252] MedChemExpress (MCE) Cat #: HY-19793

[0253] Formula: C 20 H 16 Br2N2O3S

[0254] Structure:

[0255]

[0256] Nocodazole

[0257] Selleck Cat #: S2775

[0258] MedChemExpress (MCE) Cat #: HY-13520

[0259] Formula: C 14 H 11 N3O3S

[0260] Structure:

[0261]

[0262] Trichostatin A (TSA)

[0263] Selleck Cat #: S1045

[0264] MedChemExpress (MCE) Cat #: HY-15144

[0265] Formula: C 17 H 22 N2O3

[0266] Structure:

[0267]

[0268] The sequence information involved in the present application

[0269] SEQ ID NO: 1 PERV-AC-RT-M0

[0270] TLQLDDEYRLYSPQVKPDQDIQSWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSREAREGIWPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSALPPERNWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGTGRTGQLTWTRLPQGFKNSPTIFDEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKKTVVQIPAPTTAKQVREFLGTAGFCRLWIPGFATLAAPLYPLTKEKGEFSWAPEHQKTFDAIKKALLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLLIEETGVRKDLTDIPL TGEVLTWFTDGSSYVVEG KRMAGAAVVDGTHTIWASSLPEGTSAQKAELMALTQALRLAEGKSINIYTDSRYAFATAHVHGAIYKQRGLLTSAGR EIKNKEEILSLLEALHLPKRLAIIHCPGHQKAKDLISRGNQMADRVAKQAAQAVNLLPIIETPKAPEPRRQ

[0271] Note: Underlined is the RNase H domain

[0272] SEQ ID NO: 2 PERV-B-RT-M0

[0273] TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSKEAQEGIRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRSWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGTGRTGQLTWTRLPQGFKNSPTIFDEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRDGQRWLTEARKKTVVQIPAPTTAKQVREFLGTAGFCRLWIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLLIEETGVRKDLTDIPL TGEVLTWFTDGSSYVVEG KRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQALRLAEGKSINIYTDSRYAFATAHVHGAIYKQRGLLTSAGR EIKNKEEILSLLEALHLPKRLAIIHCPGHQKAKDPISRGNQMADRVAKQAAQGVNLLPMIETPKAPEPGRQ

[0274] Note: underlined is RNase H domain

[0275] SEQ ID NO: 3 bmx-PERV-A-RT-M0

[0276] TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSKEAREGIRSHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRSWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFDEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLMEARKKTVVQIPAPTTAKQVREFLGTAGFCRLWIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEEADEPVTHDCHQLLIEETGVRKDLTDIPL TGEVLTWFTDRSSYVVEG KRMAGAAVVDGTHTIWASSLPEGTSAQKAELMALTQALRLAEGKSINIYTDSRYAFATAHVHRAIYKQRGLLTSAGR EIKNKEEILSLLEALHLPKRLAIIHCPGHQKAKDPISRGNQMADRVTKQAAQGVNLLPIIETPKAPEP

[0277] Note: underlined is RNase H domain

[0278] SEQ ID NO: 4 bmx-PERV-B-RT-M0

[0279] TLQLDDEYRLYSPQVKPDQDIQSWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSREAREGIWPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSALPPEQNWYTVLDLKDAFFCLRLHPTSQPLFAFKWRDPGTGRTGQLTWTRLPQGFKNSPTIFDEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKKTVVQIPAPTTAKQVREFLGTAGFCRLWIPGFATLAAPLYLLTKEKGEFSWAPEHQKAFDAIKKSLLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNACMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLLIEETGVRKDLTDIPL TGEVLTWFTDGSSYVVEG KRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQALRLAEGKSINIYTDSRYAFATAHVHGAIYKQRGLLTSAGR EIKNKEEILSLLEALHLPKRLAIIHCPGHQKAKDLISRGNQMADRVAKQAAQAVNLLPIIETPKAPEP

[0280] Note: underlined is RNase H domain

[0281] SEQ ID NO: 5 bmx-PERV-C-RT-M0

[0282] TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGVGLAKQVPPQVIQLKASATPVSVRQYPLSKEAREGIRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRSWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFDEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKRTVVQIPAPTTAKQVREFLGTAGFCRLWIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLLIEETGVRKDLTDIPL TGEVLTWFTDGSSYVVEG KRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQALRLAEGKSINIYTDSRYAFATAHVHGAIYKQRGLLTSAGR EIKNKEEILSLLEALHLPKRLAIIHCPGHQKAKDPISRGNQMADRVAKQAAQGVNLLPIIETPKAPEP

[0283] Note: Underlined is the RNase H domain

[0284] SEQ ID NO: 6 AC-RT-M0A

[0285] TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSKEAQEGIRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRSWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGTGRTGQLTWTRLPQGFKNSPTIFDEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRDGQRWLTEARKKTVVQIPAPTTAKQVREFLGTAGFCRLWIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLLIEETGVRKDLTDIPL

[0286] SEQ ID NO: 7 AC-RT-M3A

[0287] TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSKEAQEGIRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRSWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGTGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRDGQRWLTEARKKTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLLIEETGVRKDLTDIPL

[0288] SEQ ID NO: 8 AC-RT-M4A

[0289] TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSKEAQEGIRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRSWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGTGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRDGQRWLTEARKKTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKPGTLFSWAPEHQKAFDAIKKALLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLLIEETGVRKDLTDIPL

[0290] SEQ ID NO: 9 B-RT-M0A

[0291] TLQLDDEYRLYSPQVKPDQDIQSWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSREAREGIWPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSALPPERNWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGTGRTGQLTWTRLPQGFKNSPTIFDEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKKTVVQIPAPTTAKQVREFLGTAGFCRLWIPGFATLAAPLYPLTKEKGEFSWAPEHQKTFDAIKKALLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLLIEETGVRKDLTDIPL

[0292] SEQ ID NO: 10 B-RT-M3A

[0293] TLQLDDEYRLYSPQVKPDQDIQSWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSREAREGIWPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSALPPERNWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGTGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKKTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKTFDAIKKALLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLLIEETGVRKDLTDIPL

[0294] SEQ ID NO: 11 B-RT-M4A

[0295] TLQLDDEYRLYSPQVKPDQDIQSWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSREAREGIWPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSALPPERNWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGTGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKKTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKPGTLFSWAPEHQKTFDAIKKALLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLLIEETGVRKDLTDIPL

[0296] SEQ ID NO: 12 bmx-A-RT-M4

[0297] TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSKEAREGIRSHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRSWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLMEARKKTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEEADEPVTHDCHQLLIEETGVRKDLTDIPL TGEVLTWFTDRSSYVVEG KRMAGAAVVDGTHTIWASSLPEGTSAQKAELMALTQALRLAEGKSINIYTDSRYAFATAHVHRAI YKQRGWLTSAGR EIKNKEEILSLLEALHLPKRLAIIHCPGHQKAKDPISRGNQMADRVTKQAAQGVNLLPIIETPKAPEP

[0298] Note: Underlined is RNase H domain

[0299] SEQ ID NO: 13 bmx-B-RT-M4

[0300] TLQLDDEYRLYSPQVKPDQDIQSWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSREAREGIWPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSALPPEQNWYTVLDLKDAFFCLRLHPTSQPLFAFKWRDPGTGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKKTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYLLTKEKGEFSWAPEHQKAFDAIKKSLLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNACMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLLIEETGVRKDLTDIPL TGEVLTWFTDGSSYVVEG KRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQALRLAEGKSINIYTDSRYAFATAHVHGAI YKQRGWLTSAGR EIKNKEEILSLLEALHLPKRLAIIHCPGHQKAKDLISRGNQMADRVAKQAAQAVNLLPIIETPKAPEP

[0301] Note: underlined is RNase H domain

[0302] SEQ ID NO: 14 bmx-C-RT-M4

[0303] TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGVGLAKQVPPQVIQLKASATPVSVRQYPLSKEAREGIRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRSWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKRTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLLIEETGVRKDLTDIPL TGEVLTWFTDGSSYVVEG KRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQALRLAEGKSINIYTDSRYAFATAHVHGAI YKQRGWLTSAGR EIKNKEEILSLLEALHLPKRLAIIHCPGHQKAKDPISRGNQMADRVAKQAAQGVNLLPIIETPKAPEP

[0304] Note: Underlined is RNase H domain

[0305] SEQ ID NO: 15 bmx-A-RT-M3Δ TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSKEAREGIRSHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRSWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLMEARKKTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEEADEPVTHDCHQLLIEETGVRKDLTDIPL

[0306] SEQ ID NO: 16 bmx-B-RT-M3Δ

[0307] TLQLDDEYRLYSPQVKPDQDIQSWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSREAREGIWPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSALPPEQNWYTVLDLKDAFFCLRLHPTSQPLFAFKWRDPGTGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKKTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYLLTKEKGEFSWAPEHQKAFDAIKKSLLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNACMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLLIEETGVRKDLTDIPL

[0308] SEQ ID NO: 17 bmx-C-RT-M3A

[0309] TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGVGLAKQVPPQVIQLKASATPVSVRQYPLSKEAREGIRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRSWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHPQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKRTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPALALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTLGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLLIEETGVRKDLTDIPL

[0310] SEQ ID NO: 18 SpCas9(H840A)

[0311]

[0312] SEQ ID NO: 19 PERV-AC-NC

[0313] DFRKIRSGPRQSGNLGNRTPLDKDQCAYCKEKGHWARNCPKKGNKGLKVLALEEDK

[0314] SEQ ID NO: 20 XTEN

[0315] SGGSSGGSSGSETPGTSESATPESSGGSSGGSS

[0316] SEQ ID NO: 21 gRNA scaffold sequence

[0317] GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC

[0318] SEQ ID NO: 22 epPE-AC-M3A

[0319] MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF

[0320] DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD

[0321] EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQL

[0322] FEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQL

[0323] SKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKA

[0324] LVRQQLPEKY KEIFFDQSKN GYAGYIDGGA SQEEFYKFIK PILEKMDGTE ELLVKLNRED LLRKQRTFDN G

[0325] SIPHQIHLGE LHAILRRQED FYPFLKDNRE KIEKILTFRI PYYVGPLARG NSRFAWMTRK SEETITPWNF E

[0326] EVVDKGASAQ SFIERMTNFD KNLPNEKVLP KHSLLYEYFT VYNELTKVKY VTEGMRKPAF LSGEQKKAIVD

[0327] LLFKTNRKVTV KQLKEDYFKK IECFDSVEIS GVEDRFNASL GTYHDLLKII KDKDFLDNEE NEDILEDIVL

[0328] TLTLFEDREM IEERLKTYAH LFDDKVMKQL KRRRYTGWG RLSRKLINGI RDKQSGKTIL DFLKSDGFAN RN

[0329] FMQLIHDDSL TFKEDIQKAQ VSGQGDSLHE HIANLAGSPA IKKGILQTVK VVDELVKVMG RKHPENIVIE M

[0330] ARENQTTQKG QKNSRERMKR IEEGIKELGS QILKEHPVEN TQLQNEKLYL YYLQNGRDMY VDQELDINRL S

[0331] DYDVDAIVPQ SFLKDDSIDN KVLTRSDKNR GKSDNVPSEE VVKKMKNYWR QLLNAKLITQ RKFDNLTKAE R

[0332] GGLSELDKAG FIKRQLVETR QITKHVAQIL DSRMNTKYDE NKLIREVKVI TLKSKLVSDF RKDFQFYKVR

[0333] EINNYHHAHD AYLNAVVGTA LIKKYPKLES EFVYGDYKVY DVRKMIAKSE QEIGKATAKY FFYSNIMNFF K

[0334] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDK

[0335] LIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE

[0336] VKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE

[0337] QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0338] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSD

[0339] FRKIRSGPRQSGNLGNRTPLDKDQCAYCKEKGHWARNCPKKGNKGLKVLALEEDKSGSETPGTSESATPES

[0340] TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSKEAQE

[0341] GIRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQR

[0342] SWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGTGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQH

[0343] PQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRDGQRWLTEARK

[0344] KTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPA

[0345] LALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLT

[0346] LGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQL

[0347] LIEETGVRKDLTDIPLSGGSKRTADGSEFEPKKKRKV

[0348] SEQ ID NO:23 epPE-AC-M4A

[0349] MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF

[0350] DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD

[0351] EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQL

[0352] FEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQL

[0353] SKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKA

[0354] LVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNG

[0355] SIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFE

[0356] EVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVD

[0357] LLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVL

[0358] TLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRN

[0359] FMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEM

[0360] ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLS

[0361] DYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAER

[0362] GGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR

[0363] EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFK

[0364] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDK

[0365] LIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE

[0366] VKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE

[0367] QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0368] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSD

[0369] FRKIRSGPRQSGNLGNRTPLDKDQCAYCKEKGHWARNCPKKGNKGLKVLALEEDKSGSETPGTSESATPES

[0370] TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSKEAQE

[0371] GIRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQR

[0372] SWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGTGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQH

[0373] PQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRDGQRWLTEARK

[0374] KTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKPGTLFSWAPEHQKAFDAIKKALLSAPA

[0375] LALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLT

[0376] LGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQL

[0377] LIEETGVRKDLTDIPLSGGSKRTADGSEFEPKKKRKV

[0378] SEQ ID NO:24 epPE-B-M3Δ

[0379] MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF

[0380] DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD

[0381] EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQL

[0382] FEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQL

[0383] SKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKA

[0384] LVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNG

[0385] SIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFE

[0386] EVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVD

[0387] LLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVL

[0388] TLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRN

[0389] FMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEM

[0390] ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLS

[0391] DYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAER

[0392] GGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR

[0393] EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFK

[0394] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDK

[0395] LIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE

[0396] VKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE

[0397] QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0398] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSD

[0399] FRKIRSGPRQSGNLGNRTPLDKDQCAYCKEKGHWARNCPKKGNKGPKVLALEEDKSGSETPGTSESATPES

[0400] TLQLDDEYRLYSPQVKPDQDIQSWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSREARE

[0401] GIWPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSALPPER

[0402] NWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGTGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQH

[0403] PQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARK

[0404] KTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKTFDAIKKALLSAPA

[0405] LALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLT

[0406] LGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQL

[0407] LIEETGVRKDLTDIPLSGGSKRTADGSEFEPKKKRKV

[0408] SEQ ID NO: 25 epPE-B-M4A

[0409] MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF

[0410] DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD

[0411] EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQL

[0412] FEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQL

[0413] SKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKA

[0414] LVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNG

[0415] SIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFE

[0416] EVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVD

[0417] LLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVL

[0418] TLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRN

[0419] FMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEM

[0420] ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLS

[0421] DYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAER

[0422] GGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR

[0423] EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFK

[0424] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDK

[0425] LIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE

[0426] VKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE

[0427] QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0428] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSD

[0429] FRKIRSGPRQSGNLGNRTPLDKDQCAYCKEKGHWARNCPKKGNKGPKVLALEEDKSGSETPGTSESATPES

[0430] TLQLDDEYRLYSPQVKPDQDIQSWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSREARE

[0431] GIWPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSALPPER

[0432] NWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGTGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQH

[0433] PQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARK

[0434] KTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKPGTLFSWAPEHQKTFDAIKKALLSAPA

[0435] LALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLT

[0436] LGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQL

[0437] LIEETGVRKDLTDIPLSGGSKRTADGSEFEPKKKRKV

[0438] SEQ ID NO: 26 bmx-pPE-A-M0

[0439] MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF

[0440] DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD

[0441] EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQL

[0442] FEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQL

[0443] SKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKA

[0444] LVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNG

[0445] SIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFE

[0446] EVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVD

[0447] LLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVL

[0448] TLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRN

[0449] FMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEM

[0450] ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLS

[0451] DYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAER

[0452] GGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR

[0453] EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFK

[0454] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDK

[0455] LIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE

[0456] VKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE

[0457] QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0458] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSST

[0459] LQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSKEAREG

[0460] IRSHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRS

[0461] WYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFDEALHRDLANFRIQHP

[0462] QVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLMEARKK

[0463] TVVQIPAPTTAKQVREFLGTAGFCRLWIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPAL

[0464] ALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTL

[0465] GQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEEADEPVTHDCHQLL

[0466] IEETGVRKDLTDIPLTGEVLTWFTDRSSYVVEGKRMAGAAVVDGTHTIWASSLPEGTSAQKAELMALTQAL

[0467] RLAEGKSINIYTDSRYAFATAHVHRAIYKQRGLLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQKA

[0468] KDPISRGNQMADRVTKQAAQGVNLLPI IETPKAPEPSGGSKRTADGSEFEPKKKRKV

[0469] SEQ ID NO: 27 bmx-pPE-B-M0

[0470] MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF

[0471] DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD

[0472] EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQL

[0473] FEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQL

[0474] SKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKA

[0475] LVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNG

[0476] SIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFE

[0477] EVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVD

[0478] LLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVL

[0479] TLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRN

[0480] FMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEM

[0481] ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLS

[0482] DYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAER

[0483] GGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR

[0484] EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFK

[0485] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDK

[0486] LIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE

[0487] VKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE

[0488] QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0489] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSST

[0490] LQLDDEYRLYSPQVKPDQDIQSWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSREAREG

[0491] IWPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSALPPEQN

[0492] WYTVLDLKDAFFCLRLHPTSQPLFAFKWRDPGTGRTGQLTWTRLPQGFKNSPTIFDEALHRDLANFRIQHP

[0493] QVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKK

[0494] TVVQIPAPTTAKQVREFLGTAGFCRLWIPGFATLAAPLYLLTKEKGEFSWAPEHQKAFDAIKKSLLSAPAL

[0495] ALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTL

[0496] GQNITVIAPHALENIVRQPPDRWMTNACMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLL

[0497] IEETGVRKDLTDIPLTGEVLTWFTDGSSYVVEGKRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQAL

[0498] RLAEGKSINIYTDSRYAFATAHVHGAIYKQRGLLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQKA

[0499] KDLISRGNQMADRVAKQAAQAVNLLPI IETPKAPEPSGGSKRTADGSEFEPKKKRKV

[0500] SEQ ID NO: 28 bmx-pPE-C-M0

[0501] MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF

[0502] DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD

[0503] EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQL

[0504] FEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQL

[0505] SKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKA

[0506] LVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNG

[0507] SIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFE

[0508] EVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVD

[0509] LLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVL

[0510] TLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRN

[0511] FMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEM

[0512] ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLS

[0513] DYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAER

[0514] GGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR

[0515] EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFK

[0516] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDK

[0517] LIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE

[0518] VKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE

[0519] QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0520] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSST

[0521] LQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGVGLAKQVPPQVIQLKASATPVSVRQYPLSKEAREG

[0522] IRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRS

[0523] WYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFDEALHRDLANFRIQHP

[0524] QVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKR

[0525] TVVQIPAPTTAKQVREFLGTAGFCRLWIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPAL

[0526] ALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTL

[0527] GQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLL

[0528] IEETGVRKDLTDIPLTGEVLTWFTDGSSYVVEGKRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQAL

[0529] RLAEGKSINIYTDSRYAFATAHVHGAIYKQRGLLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQKA

[0530] KDPISRGNQMADRVAKQAAQGVNLLPI IETPKAPEPSGGSKRTADGSEFEPKKKRKV

[0531] SEQ ID NO: 29 bmx-pPE-A-M4

[0532] MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF

[0533] DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD

[0534] EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQL

[0535] FEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQL

[0536] SKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKA

[0537] LVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNG

[0538] SIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFE

[0539] EVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVD

[0540] LLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVL

[0541] TLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRN

[0542] FMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEM

[0543] ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLS

[0544] DYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAER

[0545] GGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR

[0546] EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFK

[0547] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDK

[0548] LIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE

[0549] VKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE

[0550] QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0551] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSST

[0552] LQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSKEAREG

[0553] IRSHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRS

[0554] WYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHP

[0555] QVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLMEARKK

[0556] TVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPAL

[0557] ALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTL

[0558] GQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEEADEPVTHDCHQLL

[0559] IEETGVRKDLTDIPLTGEVLTWFTDRSSYVVEGKRMAGAAVVDGTHTIWASSLPEGTSAQKAELMALTQAL

[0560] RLAEGKSINIYTDSRYAFATAHVHRAIYKQRGWLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQKA

[0561] KDPISRGNQMADRVTKQAAQGVNLLPI IETPKAPEPSGGSKRTADGSEFEPKKKRKV

[0562] SEQ ID NO: 30 bmx-pPE-B-M4

[0563] MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF

[0564] DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD

[0565] EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQL

[0566] FEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQL

[0567] SKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKA

[0568] LVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNG

[0569] SIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFE

[0570] EVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVD

[0571] LLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVL

[0572] TLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRN

[0573] FMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEM

[0574] ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLS

[0575] DYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAER

[0576] GGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR

[0577] EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFK

[0578] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDK

[0579] LIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE

[0580] VKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE

[0581] QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0582] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSST

[0583] LQLDDEYRLYSPQVKPDQDIQSWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSREAREG

[0584] IWPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSALPPEQN

[0585] WYTVLDLKDAFFCLRLHPTSQPLFAFKWRDPGTGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHP

[0586] QVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKK

[0587] TVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYLLTKEKGEFSWAPEHQKAFDAIKKSLLSAPAL

[0588] ALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTL

[0589] GQNITVIAPHALENIVRQPPDRWMTNACMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLL

[0590] IEETGVRKDLTDIPLTGEVLTWFTDGSSYVVEGKRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQAL

[0591] RLAEGKSINIYTDSRYAFATAHVHGAIYKQRGWLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQKA

[0592] KDLISRGNQMADRVAKQAAQAVNLLPI IETPKAPEPSGGSKRTADGSEFEPKKKRKV

[0593] SEQ ID NO:31 bmx-pPE-C-M4

[0594] MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF

[0595] DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD

[0596] EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQL

[0597] FEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQL

[0598] SKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKA

[0599] LVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNG

[0600] SIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFE

[0601] EVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVD

[0602] LLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVL

[0603] TLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRN

[0604] FMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEM

[0605] ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLS

[0606] DYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAER

[0607] GGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR

[0608] EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFK

[0609] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDK

[0610] LIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE

[0611] VKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE

[0612] QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0613] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSST

[0614] LQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGVGLAKQVPPQVIQLKASATPVSVRQYPLSKEAREG

[0615] IRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRS

[0616] WYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHP

[0617] QVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKR

[0618] TVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPAL

[0619] ALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTL

[0620] GQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLL

[0621] IEETGVRKDLTDIPLTGEVLTWFTDGSSYVVEGKRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQAL

[0622] RLAEGKSINIYTDSRYAFATAHVHGAIYKQRGWLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQKA

[0623] KDPISRGNQMADRVAKQAAQGVNLLPI IETPKAPEPSGGSKRTADGSEFEPKKKRKV

[0624] SEQ ID NO:32 bmx-epPE-A-M3Δ

[0625] MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF

[0626] DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD

[0627] EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQL

[0628] FEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQL

[0629] SKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKA

[0630] LVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNG

[0631] SIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFE

[0632] EVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVD

[0633] LLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVL

[0634] TLTLFEDREMIEEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRN

[0635] FMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEM

[0636] ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLS

[0637] DYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAER

[0638] GGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR

[0639] EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFK

[0640] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDK

[0641] LIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE

[0642] VKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE

[0643] QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0644] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSD

[0645] FRKIRSGPRQSGNLGNRTPLDKDQCAYCKEKGHWARDCPKKGNKGLKVLALEEDKSGSETPGTSESATPES

[0646] TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSKEARE

[0647] GIRSHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQR

[0648] SWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQH

[0649] PQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLMEARK

[0650] KTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPA

[0651] LALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLT

[0652] LGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEEADEPVTHDCHQL

[0653] LIEETGVRKDLTDIPLTGEVLTWFTDRSSYVVEGKRMAGAAVVDGTHTIWASSLPEGTSAQKAELMALTQA

[0654] LRLAEGKSINIYTDSRYAFATAHVHRAIYKQRGWLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQK

[0655] AKDPISRGNQMADRVTKQAAQGVNLLPIIETPKAPEPSGGSKRTADGSEFEPKKKRKV

[0656] SEQ ID NO:33 bmx-epPE-B-M3Δ

[0657] MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF

[0658] DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD

[0659] EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQL

[0660] FEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQL

[0661] SKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKA

[0662] LVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNG

[0663] SIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFE

[0664] EVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVD

[0665] LLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDELEDIVL

[0666] TLTLFEDREMIEEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRN

[0667] FMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEM

[0668] ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLS

[0669] DYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAER

[0670] GGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR

[0671] EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFK

[0672] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDK

[0673] LIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE

[0674] VKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE

[0675] QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0676] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSD

[0677] FRKIRSGPRQSGNLGNRTPLDKDQCAYCKKKGHWARNCPKKGNKGPKVLALEEDKSGSETPGTSESATPES

[0678] TLQLDDEYRLYSPQVKPDQDIQSWLEQFPQAWAETAGMGLAKQVPPQVIQLKASATPVSVRQYPLSREARE

[0679] GIWPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLSALPPEQ

[0680] NWYTVLDLKDAFFCLRLHPTSQPLFAFKWRDPGTGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQH

[0681] PQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARK

[0682] KTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYLLTKEKGEFSWAPEHQKAFDAIKKSLLSAPA

[0683] LALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLT

[0684] LGQNITVIAPHALENIVRQPPDRWMTNACMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQL

[0685] LIEETGVRKDLTDIPLTGEVLTWFTDGSSYVVEGKRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQA

[0686] LRLAEGKSINIYTDSRYAFATAHVHGAIYKQRGWLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQK

[0687] AKDLISRGNQMADRVAKQAAQAVNLLPIIETPKAPEPSGGSKRTADGSEFEPKKKRKV

[0688] SEQ ID NO:34 bmx-epPE-C-M3A

[0689] MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF

[0690] DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD

[0691] EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQL

[0692] FEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQL

[0693] SKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKA

[0694] LVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNG

[0695] SIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFE

[0696] EVVDKGASAQ SFIERMTNFD KNLPNEKVLP KHSLLYEYFT VYNELTKVKY VTEGMRKPAF LSGEQKKAIVD

[0697] LLFKTNRKVTVKQLKEDYFK KIECFDSVEI SGVEDRFNAS LGTYHDLLKI IKDKDFLDNE ENEDILEDIVL

[0698] TLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYT GWGRLSRKLIN GIRDKQSGKTILDFLKSDGFANRN

[0699] FMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEM

[0700] ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLS

[0701] DYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAER

[0702] GGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR

[0703] EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFK

[0704] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDK

[0705] LIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE

[0706] VKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE

[0707] QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0708] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSD

[0709] FRKIRSGPRQLGNLGNRTPLDKDQCAYCKEKGHWARDCPKKGNKGLKVLALEEDKSGSETPGTSESATPES

[0710] TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGVGLAKQVPPQVIQLKASATPVSVRQYPLSKEARE

[0711] GIRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQR

[0712] SWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQH

[0713] PQVTLLQYVDDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARK

[0714] RTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPA

[0715] LALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLT

[0716] LGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQL

[0717] LIEETGVRKDLTDIPLTGEVLTWFTDGSSYVVEGKRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQA

[0718] LRLAEGKSINIYTDSRYAFATAHVHGAIYKQRGWLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQK

[0719] AKDPISRGNQMADRVAKQAAQGVNLLPIIETPKAPEPSGGSKRTADGSEFEPKKKRKV

[0720] SEQ ID NO:35 Puro R amino acid sequence

[0721] MTEYKPTVRLATRDDVPRAVRTLAAAFADYPATRHTVDPDRHIERVTELQELFLTRVGLDIGKVWVADDGA

[0722] AVAVWTTPESVEAGAVFAEIGPRMAELSGSRLAAQQQMEGLLAPHRPKEPAWFLATVGVSPDHQGKGLGSA

[0723] VVLPGVEAAERAGVPAFLETSAPRNLPFYERLGFTVTADVEVPEGPRTWCMTRKPGA

[0724] SEQ ID NO:36 PERV-B-NC

[0725] DFRKIRSGPRQSGNLGNRTPLDKDQCAYCKEKGHWARNCPKKGNKGPKVLALEEDK

[0726] SEQ ID NO:37 bmx-PERV-A-NC

[0727] DFRKIRSGPRQSGNLGNRTPLDKDQCAYCKEKGHWARDCPKKGNKGLKVLALEEDK

[0728] SEQ ID NO:38 bmx-PERV-B-NC

[0729] DFRKIRSGPRQSGNLGNRTPLDKDQCAYCKKKGHWARNCPKKGNKGPKVLALEEDK

[0730] SEQ ID NO:39 bmx-PERV-C-NC

[0731] DFRKIRSGPRQLGNLGNRTPLDKDQCAYCKEKGHWARDCPKKGNKGLKVLALEEDK

[0732] SEQ ID NO:40 bmx-C-RT-M5.1

[0733] TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGVGLAKQVPPQVIQLKASATPVSVRQYPLSKEARE

[0734] GIRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQR

[0735] SWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQH

[0736] PQVTLLQYADDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARK

[0737] RTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPA

[0738] LALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLT

[0739] LGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQL

[0740] LIEETGVRKDLTDIPLTGEVLTWFTDGSSYVVEGKRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQA

[0741] LRLAEGKSINIYTDSRYAFATAHVHGAIYKQRGWLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQK

[0742] AKDPISRGNQMADRVAKQAA

[0743] SEQ ID NO:41 bmx-C-RT-M5.1+Y63R

[0744] TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGVGLAKQVPPQVIQLKASATPVSVRQRPLSKEARE

[0745] GIRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQR

[0746] SWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQH

[0747] PQVTLLQYADDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARK

[0748] RTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPA

[0749] LALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLT

[0750] LGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQL

[0751] LIEETGVRKDLTDIPLTGEVLTWFTDGSSYVVEGKRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQA

[0752] LRLAEGKSINIYTDSRYAFATAHVHGAIYKQRGWLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQK

[0753] AKDPISRGNQMADRVAKQAA

[0754] SEQ ID NO:42 bmx-C-RT-M5.1+C408R

[0755] TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGVGLAKQVPPQVIQLKASATPVSVRQYPLSKEARE

[0756] GIRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQR

[0757] SWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQH

[0758] PQVTLLQYADDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARK

[0759] RTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPA

[0760] LALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVRLKAIAAVAILVKDADKLT

[0761] LGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQL

[0762] LIEETGVRKDLTDIPLTGEVLTWFTDGSSYVVEGKRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQA

[0763] LRLAEGKSINIYTDSRYAFATAHVHGAIYKQRGWLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQK

[0764] AKDPISRGNQMADRVAKQAA

[0765] SEQ ID NO:43 bmx-C-RT-M7

[0766] TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGVGLAKQVPPQVIQLKASATPVSVRQRPLSKEARE

[0767] GIRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQR

[0768] SWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQH

[0769] PQVTLLQYADDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARK

[0770] RTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPA

[0771] LALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVRLKAIAAVAILVKDADKLT

[0772] LGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQL

[0773] LIEETGVRKDLTDIPLTGEVLTWFTDGSSYVVEGKRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQA

[0774] LRLAEGKSINIYTDSRYAFATAHVHGAIYKQRGWLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQK

[0775] AKDPISRGNQMADRVAKQAA

[0776] SEQ ID NO:44 bmx-C-RT-M7Δ

[0777] TLQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGVGLAKQVPPQVIQLKASATPVSVRQRPLSKEARE

[0778] GIRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQR

[0779] SWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQH

[0780] PQVTLLQYADDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARK

[0781] RTVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPA

[0782] LALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVRLKAIAAVAILVKDADKLT

[0783] LGQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQL

[0784] LIEETGVRKDLTDIPL

[0785] SEQ ID NO:45 bmx-pPE-C-M5.1

[0786] MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF

[0787] DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD

[0788] EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQL

[0789] FEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQL

[0790] SKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKA

[0791] LVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNG

[0792] SIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFE

[0793] EVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVD

[0794] LLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVL

[0795] TLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRN

[0796] FMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEM

[0797] ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLS

[0798] DYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAER

[0799] GGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR

[0800] EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFK

[0801] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDK

[0802] LIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE

[0803] VKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE

[0804] QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0805] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSST

[0806] LQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGVGLAKQVPPQVIQLKASATPVSVRQYPLSKEAREG

[0807] IRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRS

[0808] WYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHP

[0809] QVTLLQYADDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKR

[0810] TVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPAL

[0811] ALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTL

[0812] GQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLL

[0813] IEETGVRKDLTDIPLTGEVLTWFTDGSSYVVEGKRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQAL

[0814] RLAEGKSINIYTDSRYAFATAHVHGAIYKQRGWLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQKA

[0815] KDPISRGNQMADRVAKQAA

[0816] SEQ ID NO: 46 bmx-pPE-C-M5.1+Y63R

[0817] MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF

[0818] DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD

[0819] EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQL

[0820] FEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQL

[0821] SKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKA

[0822] LVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNG

[0823] SIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFE

[0824] EVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVD

[0825] LLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVL

[0826] TLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRN

[0827] FMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEM

[0828] ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLS

[0829] DYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAER

[0830] GGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR

[0831] EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFK

[0832] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDK

[0833] LIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE

[0834] VKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE

[0835] QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0836] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSST

[0837] LQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGVGLAKQVPPQVIQLKASATPVSVRQRPLSKEAREG

[0838] IRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRS

[0839] WYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHP

[0840] QVTLLQYADDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKR

[0841] TVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPAL

[0842] ALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVCLKAIAAVAILVKDADKLTL

[0843] GQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLL

[0844] IEETGVRKDLTDIPLTGEVLTWFTDGSSYVVEGKRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQAL

[0845] RLAEGKSINIYTDSRYAFATAHVHGAIYKQRGWLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQKA

[0846] KDPISRGNQMADRVAKQAA

[0847] SEQ ID NO: 47 bmx-pPE-C-M5.1+C408R

[0848] MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF

[0849] DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD

[0850] EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQL

[0851] FEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQL

[0852] SKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKA

[0853] LVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNG

[0854] SIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFE

[0855] EVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVD

[0856] LLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVL

[0857] TLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRN

[0858] FMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEM

[0859] ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLS

[0860] DYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAER

[0861] GGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR

[0862] EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFK

[0863] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDK

[0864] LIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE

[0865] VKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE

[0866] QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0867] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSST

[0868] LQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGVGLAKQVPPQVIQLKASATPVSVRQYPLSKEAREG

[0869] IRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRS

[0870] WYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHP

[0871] QVTLLQYADDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKR

[0872] TVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPAL

[0873] ALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVRLKAIAAVAILVKDADKLTL

[0874] GQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLL

[0875] IEETGVRKDLTDIPLTGEVLTWFTDGSSYVVEGKRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQAL

[0876] RLAEGKSINIYTDSRYAFATAHVHGAIYKQRGWLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQKA

[0877] KDPISRGNQMADRVAKQAA

[0878] SEQ ID NO:48 bmx-pPE-C-M7

[0879] MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF

[0880] DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD

[0881] EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQL

[0882] FEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQL

[0883] SKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKA

[0884] LVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNG

[0885] SIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFE

[0886] EVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVD

[0887] LLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVL

[0888] TLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRN

[0889] FMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEM

[0890] ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLS

[0891] DYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAER

[0892] GGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR

[0893] EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFK

[0894] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDK

[0895] LIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE

[0896] VKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE

[0897] QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0898] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSST

[0899] LQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGVGLAKQVPPQVIQLKASATPVSVRQRPLSKEAREG

[0900] IRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRS

[0901] WYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHP

[0902] QVTLLQYADDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKR

[0903] TVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPAL

[0904] ALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVRLKAIAAVAILVKDADKLTL

[0905] GQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLL

[0906] IEETGVRKDLTDIPLTGEVLTWFTDGSSYVVEGKRMAGAAVVDGTRTIWASSLPEGTSAQKAELMALTQAL

[0907] RLAEGKSINIYTDSRYAFATAHVHGAIYKQRGWLTSAGREIKNKEEILSLLEALHLPKRLAIIHCPGHQKA

[0908] KDPISRGNQMADRVAKQAA

[0909] SEQ ID NO: 49 bmx-epPE-C-M7A

[0910] MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLF

[0911] DSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVD

[0912] EVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQL

[0913] FEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQL

[0914] SKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKA

[0915] LVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNG

[0916] SIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFE

[0917] EVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVD

[0918] LLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVL

[0919] TLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRN

[0920] FMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEM

[0921] ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLS

[0922] DYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAER

[0923] GGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR

[0924] EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFK

[0925] TEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDK

[0926] LIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKE

[0927] VKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVE

[0928] QHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTI

[0929] DRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSSGSETPGTSESATPESSGGSSGGSST

[0930] LQLDDEYRLYSPLVKPDQNIQFWLEQFPQAWAETAGVGLAKQVPPQVIQLKASATPVSVRQRPLSKEAREG

[0931] IRPHVQRLIQQGILVPVQSPWNTPLLPVRKPGTNDYRPVQDLREVNKRVQDIHPTVPNPYNLLCALPPQRS

[0932] WYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPGAGRTGQLTWTRLPQGFKNSPTIFNEALHRDLANFRIQHP

[0933] QVTLLQYADDLLLAGATKQDCLEGTKALLLELSDLGYRASAKKAQICRREVTYLGYSLRGGQRWLTEARKR

[0934] TVVQIPAPTTAKQVREFLGKAGFCRLFIPGFATLAAPLYPLTKEKGEFSWAPEHQKAFDAIKKALLSAPAL

[0935] ALPDVTKPFTLYVDERKGVARGVLTQTLGPWRRPVAYLSKKLDPVASGWPVRLKAIAAVAILVKDADKLTL

[0936] GQNITVIAPHALENIVRQPPDRWMTNARMTHYQSLLLTERVTFAPPAALNPATLLPEETDEPVTHDCHQLL

[0937] IEETGVRKDLTDIPL

[0938] SEQ ID NO:50 Csy4

[0939] GTTCACTGCCGTATAGGCAG

[0940] SEQ ID NO:51 tCsy4

[0941] CTGCCGTATAGGCAG

Claims

1. A pilot editing system for targeted modification of genomic DNA sequences of organisms, comprising: i) (a) A CRISPR nuclease or an expression construct containing a nucleotide sequence encoding the CRISPR nuclease, and a PERV reverse transcriptase or a functional variant thereof or an expression construct containing a nucleotide sequence encoding the PERV reverse transcriptase or a functional variant thereof; or (b) a lead-editing fusion protein or an expression construct containing a nucleotide sequence encoding said lead-editing fusion protein, wherein said lead-editing fusion protein comprises a CRISPR nuclease and a PERV reverse transcriptase or a functional variant thereof; and ii) At least one pegRNA or an expression construct containing a nucleotide sequence encoding said at least one pegRNA, The amino acid sequence of the PERV reverse transcriptase or a functional variant thereof is shown in one of SEQ ID NO: 5, 7-8, 10-17 or 40-44. The CRISPR nuclease is a Cas9 nickase.

2. The lead editing system of claim 1, wherein in the fusion protein, the reverse transcriptase or a functional variant thereof is fused directly or via a linker to the nucleocapsid protein at the N-terminus or C-terminus.

3. The lead editing system of claim 2, wherein the nucleocapsid protein is derived from PERV.

4. The lead editing system of claim 3, wherein the nucleocapsid protein and the reverse transcriptase are derived from the same type of PERV.

5. The lead editing system of claim 3, wherein the amino acid sequence of the nucleocapsid protein is as shown in one of SEQ ID NO:19 or 36-39.

6. The lead editing system of claim 2, wherein the PERV reverse transcriptase or a functional variant thereof is fused to the nucleocapsid protein directly or via a linker at the N-terminus.

7. The lead editing system of any one of claims 1-6, wherein the CRISPR nuclease and the PERV reverse transcriptase or a functional variant thereof in the lead editing fusion protein are linked by a adapter.

8. The lead editing system of claim 7, wherein the amino acid sequence of the Cas9 nickase is as shown in SEQ ID NO:

18.

9. The lead editing system of any one of claims 1-6, wherein the CRISPR nuclease in the fusion protein is located at the N-terminus of the reverse transcriptase or a functional variant thereof.

10. The lead editing system of any one of claims 1-6, wherein the fusion protein further comprises a nuclear localization sequence.

11. The lead editing system of any one of claims 1-6, wherein the fusion protein further comprises a selectable marker protein linked directly or via a self-cleaving peptide.

12. The lead editing system of claim 11, wherein the self-cleaving peptide is a 2A polypeptide and the selectable marker protein is the Puro-R protein.

13. The lead editing system of claim 1, wherein the amino acid sequence of the fusion protein is shown in one of SEQ ID NO:22-25, 28-34 or SEQ ID NO:45-49.

14. The lead editing system of any one of claims 1-6, wherein the at least one pegRNA is capable of forming a complex with the fusion protein and targeting the fusion protein to a target sequence in the genome, resulting in a cut within the target sequence.

15. The leader editing system of any one of claims 1-6, wherein the at least one pegRNA comprises a leader sequence, a scaffold sequence, a reverse transcription template sequence, and a primer binding site sequence from the 5' to 3' direction. At least one of the pegRNAs has a leader sequence that is sufficiently identical to the target sequence, enabling it to bind to the complementary strand of the target sequence via base pairing for sequence-specific targeting.

16. The pre-editing system of claim 15, wherein the scaffold sequence is shown in SEQ ID NO:

21.

17. The lead editing system of claim 15, wherein the primer binding sequence is configured to be complementary to at least a portion of the target sequence.

18. The lead editing system of claim 17, wherein the primer binding sequence is complementary to at least a portion of the 3' free single strand of the DNA strand containing the target sequence, resulting from a nick.

19. The lead editing system of claim 18, wherein the primer-binding sequence is complementary to the nucleotide sequence at the 3' end of the 3' free single strand.

20. The pilot editing system of claim 15, wherein, The reverse transcription template sequence is configured to correspond to the sequence downstream of the target sequence nick and contains one or more nucleotide substitutions, deletions, and / or additions.

21. The lead editing system of any one of claims 1-6, further comprising a nick gRNA or an expression construct containing a nucleotide sequence encoding the nick gRNA, the nick gRNA comprising a lead sequence and a scaffold sequence, the lead sequence being configured to have sufficient sequence identity with a target sequence in the genome, thereby enabling the fusion protein to target the target sequence and resulting in a nick within the target sequence, the target sequence of the nick gRNA and the target sequence of the pegRNA being located on opposite strands of the genomic DNA, the nick induced by the nick gRNA and the nick induced by the pegRNA being 1-300 nucleotides apart.

22. The lead editing system of any one of claims 1-6, comprising at least one pair of pegRNAs or an expression construct containing a nucleotide sequence encoding said at least one pair of pegRNAs.

23. The lead editing system of claim 22, wherein the two pegRNAs in the pegRNA pair are configured to target different target sequences on the same strand of genomic DNA, or the two pegRNAs in the pegRNA pair are configured to target different target sequences on different strands of genomic DNA.

24. The lead editing system of claim 22, wherein the two pegRNAs in the pegRNA pair are configured to introduce the same modification, said modification being a substitution, deletion, and / or addition of one or more nucleotides.

25. The lead editing system of any one of claims 1-6, wherein the pegRNA comprises a truncated Csy4 recognition sequence shown in SEQ ID NO:51 at its 3' end.

26. A method for generating genetically modified cells in vitro, comprising introducing a pilot editing system of any one of claims 1-25 into at least one of the cells, thereby resulting in modification of the genome sequence of the at least one cell, wherein the cell is derived from a mammal.

27. The method of claim 26, wherein the lead editing system is introduced into the cells in the presence of one or more small molecule compounds selected from RS-1, nocodazole, and trichostatin A; or After the pilot editing system is introduced into the cells, the cells are treated with one or more small molecule compounds selected from RS-1, nocodazole, and trichostatin A.

28. The method of claim 27, wherein the one or more small molecule compounds are RS-1 and nocodazole.

29. The method of claim 26, wherein the mammal is selected from humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats.

30. Use of the lead editing system of any one of claims 1-25 in the preparation of a kit for generating genetically modified cells, wherein said cells are derived from mammals.

31. The use of claim 30, wherein the mammal is selected from humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats.

32. A functional variant of a reverse transcriptase derived from porcine endogenous retrovirus (PERV), having an amino acid sequence as shown in one of SEQ ID NO: 7-8, 10-17 or 40-44.

Citation Information

Patent Citations

  • Pilot editing system and gene editing method

    CN116396952A

  • Pilot editing system and application thereof

    CN117720672A

  • Guide editing system based on circular RNA

    CN118995701A

  • Method for producing genetically modified cells

    WO2022148955A1

  • Improved prime editors and methods of use

    WO2023015309A2