Improved CG Base Editing System
By using the C-to-G base editing system of cytosine deaminase, nuclease-inactivated CRISPR effector protein and uracil-DNA glycosylase in the plant genome, combined with guide RNA and other components, the problems of low C-to-G base editing efficiency and many Indel mutations in the prior art are solved, and efficient and accurate base editing is achieved.
Patent Information
- Application Number
- CN202210225327.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-09
- Filing Date
- 2022-03-09
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-03-09
AI Technical Summary
The prior art is difficult to achieve efficient and accurate C-to-G base editing in the animal and plant genomes, especially in plants, and the existing system cannot achieve other types of base substitutions except C·G to T·A and A·T to G·C.
The C-to-G base editing system containing cytosine deaminase, nuclease-inactivated CRISPR effector protein and uracil-DNA glycosylase is adopted to combine guide RNA and edit by targeting the cell genome. It can also be equipped with components such as proliferating cell nuclear antigens whose ubiquitin protein binding sites are mutated and mutated AP endonucleases to optimize editing efficiency and accuracy.
Efficient and accurate C-to-G base editing is achieved in the plant genome, reducing the generation of Indel mutations, expanding the scope of base editing, and improving the efficiency and accuracy of the editing system.
Smart Images

Figure BDA0003538988620000151 
Figure HDA0003538988630000011 
Figure HDA0003538988630000012
Abstract
Description
Technical Field
[0001] The present invention relates to the field of genetic engineering. Specifically, the present invention relates to an improved CG base editing system, which can achieve efficient and precise in vivo C to G base editing. Background of the Invention
[0002] In recent years, with the development of genome editing technologies, a large number of base editors have been continuously developed, improved, and applied. These include cytosine base editors (CBEs) composed of Cas9 nickase (nCas9(D10A)) fused with cytosine deaminase and uracil glycosylase inhibitor (UGI), and adenine base editors (ABEs) composed of nCas9 fused with adenosine deaminase. They respectively mediate precise base substitutions between C·G and T·A, and between A·T and G·C at targeted sites in plant and animal genomes (Komor et al., 2016; Gaudelli et al., 2017; Zong et al., 2018; Li et al., 2018). In 2019, Anzolobe et al. constructed a prime editing system using nCas9(H840A) fused with reverse transcriptase (MMLV), which can successfully achieve substitutions between any types of bases in the genome under the guidance of pegRNA. However, the editing efficiency of this system is limited by multiple factors such as the targeted sequence, the length of PBS, and the melting temperature (Tm) (Lin et al., 2020). In addition, after combining the CBE and ABE8e systems (Richter et al., 2020) with the Cas9 variants SpG and SpRY (Walton et al., 2020), efficient base editing between C·G and T·A, and between A·T and G·C at any targeted site can also be achieved, and the efficiency is much higher than that of the PE system. However, these two systems cannot achieve substitutions between other types of bases. Therefore, there is still a need to develop new base editing systems to achieve substitutions between other types of bases in the current development situation of gene editing. Zhao et al. (2020) and Kurt et al. (2020) constructed a CG base editor using nCas9(D10A) fused with APOBEC1 cytosine deaminase and uracil-DNA glycosylase (UDG), and successfully achieved C-to-G transversions in mammalian cells, but at the same time, a large number of by-products such as C-to-T, C-to-A, and Indels were introduced. Moreover, due to the differences in DNA damage repair pathways between plants and animals, plant CG base editing systems are more prone to generating Indel mutations, which also makes it impossible to establish corresponding CG base editing systems in plants so far. Therefore, it is urgent to establish a CG base editing system in plants to expand the scope of single-base editing, and at the same time further optimize the CG base editing system to enable it to mediate C-to-G base transversions more efficiently and precisely in plant and animal genomes. Brief Description of the Drawings
[0003] Figure 1, showing the occurrence of DNA damage involved in cytosine deamination and its potential repair pathways.
[0004] Figure 2 , showing the vector construction maps of A3A-PBE and PCGBE-1 (Plant C to G Base Editing-1).
[0005] Figure 3 , showing the mutation types and efficiencies mediated by A3A-PBE and PCGBE-1 in rice protoplasts.
[0006] Figure 4 , showing the mutation types and efficiencies mediated by A3A-PBE and PCGBE-1 at the OsNRT1.1B, OsPDS, and OsGRF1 target sites in rice calli.
[0007] Figure 5 , showing the mutation types and efficiencies mediated by A3A-PBE and PCGBE-1 at the OsAAT and OsSWEET14 target sites in rice calli.
[0008] Figure 6 , showing the vector construction maps of the optimized PCGBE-2 to 6.
[0009] Figure 7 , showing the mutation types and efficiencies mediated by PCGBE-1 to 6 at the OsNRT1.1B target site in rice calli.
[0010] Figure 8 , showing the mutation types and efficiencies mediated by PCGBE-1 to 6 at the OsPDS target site in rice calli.
[0011] Figure 9 , showing the mutation types and efficiencies mediated by PCGBE-1 to 6 at the OsGRF1 target site in rice calli.
[0012] Figure 10 , showing the differences in C-to-G editing efficiency and editing purity among PCGBE-1 to 6.
[0013] Figure 11 , showing the construction of additional PCGBE binary vectors.
[0014] Figure 12 , showing Figure 11 b and Figure 11 the mutation types and efficiencies mediated by the two PCGBE systems shown in c in regenerated rice plants. DETAILED DESCRIPTION OF THE INVENTION
[0015] I. Definitions
[0016] In the present invention, unless otherwise specified, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Moreover, the protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, and immunology-related terms and laboratory procedures used herein are all widely used terms and conventional procedures in the corresponding fields. For example, the standard recombinant DNA and molecular cloning techniques used in the present invention are well known to those skilled in the art and are more comprehensively described in the following literature: Sambrook, J., Fritsch, E.F., and Maniatis, T., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press: Cold Spring Harbor, 1989 (hereinafter referred to as "Sambrook"). Meanwhile, to better understand the present invention, the definitions and explanations of relevant terms are provided below.
[0017] As used herein, the term "and / or" encompasses all combinations of the items connected by this term, and should be regarded as each combination having been separately listed herein. For example, "A and / or B" encompasses "A", "A and B", and "B". For example, "A, B, and / or C" encompasses "A", "B", "C", "A and B", "A and C", "B and C", and "A and B and C".
[0018] As used herein, "genome" not only encompasses the chromosomal DNA present in the cell nucleus, but also includes organelle DNA present in the subcellular components of the cell (such as mitochondria, plastids).
[0019] As used herein, "organism" includes any organism suitable for genome editing, preferably eukaryotes. Examples of organisms include, but are not limited to, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, cats; poultry such as chickens, ducks, geese; plants including monocotyledonous plants and dicotyledonous plants, such as rice, corn, wheat, sorghum, barley, soybeans, peanuts, Arabidopsis, etc.
[0020] "Genetically modified organism" or "genetically modified cell" means an organism or cell that contains an exogenous polynucleotide or a modified gene or an expression regulatory sequence within its genome. For example, the exogenous polynucleotide can be stably integrated into the genome of the organism or cell and be inherited through successive generations. The exogenous polynucleotide can be integrated into the genome alone or as part of a recombinant DNA construct. The modified gene or expression regulatory sequence means that the sequence in the genome of the organism or cell contains single or multiple deoxynucleotide substitutions, deletions, and additions.
[0021] "Exogenous," in reference to a sequence, means a sequence that is from a foreign species, or, if from the same species, has been substantially modified in composition and / or genomic locus from its native form through deliberate human intervention.
[0022] "Polynucleotide," "nucleic acid sequence," "nucleotide sequence," or "nucleic acid fragment" are used interchangeably and refer to a single-stranded or double-stranded RNA or DNA polymer, optionally containing synthetic, non-natural or altered nucleotide bases. Nucleotides are referred to by their single letter designations: "A" for adenosine or deoxyadenosine (corresponding to RNA or DNA, respectively), "C" for cytidine or deoxycytidine, "G" for guanosine or deoxyguanosine, "U" for uridine, "T" for deoxythymidine, "R" for purine (A or G), "Y" for pyrimidine (C or T), "K" for G or T, "H" for A or C or T, "I" for inosine, and "N" for any nucleotide.
[0023] "Polypeptide," "peptide," and "protein" are used interchangeably in the present invention and refer to a polymer of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally-occurring amino acids, as well as to naturally-occurring amino acid polymers. The terms "polypeptide," "peptide," "amino acid sequence," and "protein" also include modified forms, including but not limited to glycosylation, lipid attachment, sulfation, γ-carboxylation of glutamic acid residues, hydroxylation and ADP-ribosylation.
[0024] The term "identity" of sequences has the meaning recognized in the art, and the percentage of sequence identity between two nucleic acid or polypeptide molecules or regions can be calculated using publicly available techniques. Sequence identity can be measured along the full length of a polynucleotide or polypeptide or along a region of the molecule. (See, e.g., Computational Molecular Biology, Lesk, A.M., ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, D.W., ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, A.M., and Griffin, H.G., eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). Although there are many methods for measuring the identity between two polynucleotides or polypeptides, the term "identity" is well known to those skilled in the art (Carrillo, H. & Lipman, D., SIAM J Applied Math 48:1073 (1988)).
[0025] When the term "comprising" is used herein to describe the sequence of a protein or nucleic acid, the protein or nucleic acid may consist of the sequence, or may have additional amino acids or nucleotides at one or both ends thereof, but still have the activity described in the present invention. In addition, it is well known to those skilled in the art that the methionine encoded by the start codon at the N-terminus of a polypeptide will be retained in some practical cases (e.g., when expressed in a specific expression system), but does not substantially affect the function of the polypeptide. Therefore, in the description of the specific polypeptide amino acid sequence in the specification and claims of the present application, although it may not contain the methionine encoded by the start codon at the N-terminus, at this time the sequence containing this methionine is also covered. Correspondingly, its encoded nucleotide sequence may also contain the start codon; and vice versa.
[0026] In peptides or proteins, suitable conservative amino acid substitutions are known to those skilled in the art and can generally be made without altering the biological activity of the resulting molecule. Typically, those skilled in the art recognize that a single amino acid substitution in a non-essential region of a polypeptide generally does not alter biological activity (see, e.g., Watson et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub.co., p. 224).
[0027] As used herein, the term "CRISPR effector protein" generally refers to nucleases present in naturally occurring CRISPR systems, as well as modified forms thereof, variants thereof, catalytically active fragments thereof, etc. This term encompasses any effector protein based on the CRISPR system that is capable of achieving gene targeting (e.g., gene editing, gene targeting regulation, etc.) intracellularly. The CRISPR effector proteins described herein can, for example, be selected from Cas3, Cas8a, Cas5, Cas8b, Cas8c, Cas10d, Cse1, Cse2, Csy1, Csy2, Csy3, GSU0054, Cas10, Csm2, Cmr5, Cas10, Csx11, Csx10, Csf1, Cas9, Csn2, Cas4, Cpf1, C2c1, C2c3 or C2c2 proteins, or functional variants of these nucleases. Examples of "CRISPR effector proteins" include Cas9 nucleases or variants thereof. The Cas9 nuclease can be a Cas9 nuclease from different species, such as spCas9 from Streptococcus pyogenes or SaCas9 derived from Staphylococcus aureus. "Cas9 nuclease" and "Cas9" are used interchangeably herein and refer to an RNA-guided nuclease comprising a Cas9 protein or a fragment thereof (e.g., a protein comprising the active DNA cleavage domain of Cas9 and / or the gRNA binding domain of Cas9). Cas9 is a component of the CRISPR / Cas (clustered regularly interspaced short palindromic repeats and their associated systems) genome editing system and can target and cleave a DNA target sequence to form a DNA double-strand break (DSB) under the guidance of a guide RNA. Examples of "CRISPR effector proteins" can also include Cpf1 nucleases or variants thereof such as highly specific variants. The Cpf1 nuclease can be a Cpf1 nuclease from different species, such as the Cpf1 nuclease from Francisella novicida U112, Acidaminococcus sp. BV3L6 and Lachnospiraceae bacterium ND2006.
[0028] As used herein, an "expression construct" refers to a vector, such as a recombinant vector, suitable for the expression of a nucleotide sequence of interest in an organism. "Expression" refers to the production of a functional product. For example, the expression of a nucleotide sequence may refer to the transcription of the nucleotide sequence (such as transcription to produce mRNA or functional RNA) and / or the translation of the RNA into a precursor or mature protein.
[0029] The "expression construct" of the present invention may be a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, may be translatable RNA (such as mRNA).
[0030] The "expression construct" of the present invention may comprise regulatory sequences and a nucleotide sequence of interest from different sources, or regulatory sequences and a nucleotide sequence of interest from the same source but arranged in a manner different from that found in nature.
[0031] "Regulatory sequence" and "regulatory element" are used interchangeably and refer to nucleotide sequences located upstream (5' non-coding sequence), within, or downstream (3' non-coding sequence) of a coding sequence and which affect the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences may include, but are not limited to, promoters, translational leader sequences, introns, and polyadenylation recognition sequences.
[0032] A "promoter" refers to a nucleic acid fragment capable of controlling the transcription of another nucleic acid fragment. In some embodiments of the present invention, the promoter is a promoter capable of controlling the transcription of a gene in a cell, whether or not it is derived from the cell. The promoter may be a constitutive promoter or a tissue-specific promoter or a developmentally regulated promoter or an inducible promoter.
[0033] A "constitutive promoter" refers to a promoter that generally causes a gene to be expressed in most cell types under most circumstances. "Tissue-specific promoter" and "tissue-preferred promoter" are used interchangeably and refer to a promoter that is expressed primarily, but not necessarily exclusively, in one tissue or organ and may also be expressed in a particular cell or cell type. A "developmentally regulated promoter" refers to a promoter whose activity is determined by developmental events. An "inducible promoter" selectively expresses an operably linked DNA sequence in response to an endogenous or exogenous stimulus (such as environment, hormone, chemical signal, etc.).
[0034] Examples of promoters include, but are not limited to, polymerase (pol) I, pol II, or pol III promoters. Examples of pol I promoters include the chicken RNA pol I promoter. Examples of pol II promoters include, but are not limited to, the cytomegalovirus immediate early (CMV) promoter, the Rous sarcoma virus long terminal repeat (RSV-LTR) promoter, and the simian virus 40 (SV40) immediate early promoter. Examples of pol III promoters include the U6 and H1 promoters. Inducible promoters such as the metallothionein promoter can be used. Other examples of promoters include the T7 phage promoter, the T3 phage promoter, the β-galactosidase promoter, and the Sp6 phage promoter. When used in plants, the promoter can be the cauliflower mosaic virus 35S promoter, the maize Ubi-1 promoter, the wheat U6 promoter, the rice U3 promoter, the maize U3 promoter, the rice actin promoter.
[0035] As used herein, the term "operably linked" refers to the connection of a regulatory element (such as, but not limited to, a promoter sequence, a transcription termination sequence, etc.) to a nucleic acid sequence (such as a coding sequence or an open reading frame) such that the transcription of the nucleotide sequence is controlled and regulated by the transcription regulatory element. Techniques for operably linking a regulatory element region to a nucleic acid molecule are known in the art.
[0036] "Introducing" a nucleic acid molecule (such as a plasmid, a linear nucleic acid fragment, RNA, etc.) or a protein into an organism means transforming the cells of the organism with the nucleic acid or protein such that the nucleic acid or protein can function in the cell. "Transformation" as used in the present invention includes stable transformation and transient transformation.
[0037] "Stable transformation" refers to the introduction of an exogenous nucleotide sequence into the genome, resulting in the stable inheritance of the exogenous gene. Once stably transformed, the exogenous nucleic acid sequence is stably integrated into the genome of the organism and any successive generations thereof.
[0038] "Transient transformation" refers to the introduction of a nucleic acid molecule or a protein into a cell, which functions without the stable inheritance of the exogenous gene. In transient transformation, the exogenous nucleic acid sequence is not integrated into the genome.
[0039] II. C-to-G Base Editing System
[0040] The present invention provides a C-to-G base editing system for editing a target sequence in the genome of a cell, which comprises:
[0041] A first polypeptide and / or an expression construct comprising a nucleotide sequence encoding the first polypeptide, wherein the first polypeptide comprises i) a cytosine deaminase, a nuclease-inactivated CRISPR effector protein, and a uracil-DNA glycosylase (UDG), or ii) a cytosine deaminase, a nuclease-inactivated CRISPR effector protein, and a Rad18 protein; and
[0042] A guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, the guide RNA being capable of targeting the first polypeptide to a target sequence in the cell genome.
[0043] In some embodiments, the C-to-G base editing system further comprises
[0044] i) a second polypeptide and / or an expression construct comprising a nucleotide sequence encoding the second polypeptide, wherein the second polypeptide comprises a fusion of a proliferating cell nuclear antigen (PCNA) with a ubiquitin protein binding site mutated; P roliferating C ell N uclear A ntigen, PCNA) and a ubiquitin protein;
[0045] ii) a third polypeptide and / or an expression construct comprising a nucleotide sequence encoding the third polypeptide, wherein the third polypeptide comprises a mutated AP endonuclease (APE);
[0046] iii) a fourth polypeptide and / or an expression construct comprising a nucleotide sequence encoding the fourth polypeptide, wherein the fourth polypeptide comprises a Rad18 protein; and / or
[0047] iv) a fifth polypeptide and / or an expression construct comprising a nucleotide sequence encoding the fifth polypeptide, wherein the fifth polypeptide comprises a uracil-DNA glycosylase (UDG).
[0048] In some embodiments, the C-to-G base editing system comprises:
[0049] A first polypeptide and / or an expression construct comprising a nucleotide sequence encoding the first polypeptide, wherein the first polypeptide comprises a cytosine deaminase, a nuclease-inactivated CRISPR effector protein, and a uracil-DNA glycosylase (UDG);
[0050] A guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, the guide RNA being capable of targeting the first polypeptide to a target sequence in the cell genome; and
[0051] A second polypeptide and / or an expression construct comprising a nucleotide sequence encoding the second polypeptide, wherein the second polypeptide comprises a proliferating cell nuclear antigen (PCNA) with a ubiquitin protein binding site mutatedP proliferating C ell N nuclear A A fusion of antigen, PCNA) and ubiquitin protein.
[0052] In some embodiments, the C-to-G base editing system comprises:
[0053] A first polypeptide and / or an expression construct comprising a nucleotide sequence encoding the first polypeptide, wherein the first polypeptide comprises a cytosine deaminase, a nuclease-inactivated CRISPR effector protein, and uracil-DNA glycosylase (UDG);
[0054] A guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, the guide RNA being capable of targeting the first polypeptide to a target sequence in the cell genome; and
[0055] A third polypeptide and / or an expression construct comprising a nucleotide sequence encoding the third polypeptide, wherein the third polypeptide comprises a mutated AP endonuclease (APE).
[0056] In some embodiments, the C-to-G base editing system comprises:
[0057] A first polypeptide and / or an expression construct comprising a nucleotide sequence encoding the first polypeptide, wherein the first polypeptide comprises a cytosine deaminase, a nuclease-inactivated CRISPR effector protein, and uracil-DNA glycosylase (UDG);
[0058] A guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, the guide RNA being capable of targeting the first polypeptide to a target sequence in the cell genome;
[0059] A second polypeptide and / or an expression construct comprising a nucleotide sequence encoding the second polypeptide, wherein the second polypeptide comprises a fusion of proliferating cell nuclear antigen ( P proliferating C ell N nuclear A antigen, PCNA) with a mutated ubiquitin protein binding site and;
[0060] A third polypeptide and / or an expression construct comprising a nucleotide sequence encoding the third polypeptide, wherein the third polypeptide comprises a mutated AP endonuclease (APE).
[0061] In some embodiments, the C-to-G base editing system comprises:
[0062] A first polypeptide and / or an expression construct comprising a nucleotide sequence encoding the first polypeptide, wherein the first polypeptide comprises a cytosine deaminase, a nuclease-inactivated CRISPR effector protein, and a uracil-DNA glycosylase (UDG);
[0063] A guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, the guide RNA being capable of targeting the first polypeptide to a target sequence in the cell genome; and
[0064] A fourth polypeptide and / or an expression construct comprising a nucleotide sequence encoding the fourth polypeptide, wherein the fourth polypeptide comprises a Rad18 protein.
[0065] In some embodiments, the C-to-G base editing system comprises:
[0066] A first polypeptide and / or an expression construct comprising a nucleotide sequence encoding the first polypeptide, wherein the first polypeptide comprises a cytosine deaminase, a nuclease-inactivated CRISPR effector protein, and a Rad18 protein;
[0067] A guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, the guide RNA being capable of targeting the first polypeptide to a target sequence in the cell genome; and
[0068] A fifth polypeptide and / or an expression construct comprising a nucleotide sequence encoding the fifth polypeptide, wherein the fifth polypeptide comprises a uracil-DNA glycosylase (UDG).
[0069] In some embodiments, the expression construct comprising a nucleotide sequence encoding the first polypeptide, the expression construct comprising a nucleotide sequence encoding the second polypeptide, the expression construct comprising a nucleotide sequence encoding the third polypeptide, the expression construct comprising a nucleotide sequence encoding the fourth polypeptide, the expression construct comprising a nucleotide sequence encoding the fifth polypeptide, and / or the expression construct comprising a nucleotide sequence encoding the guide RNA can be different expression constructs, or any two, any three, or all of them can be the same expression construct. In some embodiments, the first polypeptide is isolated, the second polypeptide is isolated, the third polypeptide is isolated, the fourth polypeptide is isolated, the fifth polypeptide is isolated, and / or the guide RNA is isolated.
[0070] In some embodiments, the gene editing system at least comprises an expression construct, which comprises a nucleotide sequence encoding the first polypeptide, a nucleotide sequence encoding a self-cleaving peptide, and a nucleotide sequence encoding the second polypeptide, the third polypeptide, the fourth polypeptide, or the fifth polypeptide, which are ligated in-frame.
[0071] As used herein, "self-cleaving peptide" means a peptide that can achieve self-cleavage intracellularly. For example, the self-cleaving peptide may contain a protease recognition site and thus be recognized and specifically cleaved by proteases intracellularly.
[0072] Alternatively, the self-cleaving peptide may be a 2A polypeptide. 2A polypeptides are a class of short peptides from viruses, and their self-cleavage occurs during translation. When two different polypeptides of interest are expressed in the same reading frame linked by a 2A polypeptide, the two polypeptides of interest are produced almost in a 1:1 ratio. Commonly used 2A polypeptides may be P2A from porcine teschovirus-1, T2A from Thosea asigna virus, E2A from equine rhinitis A virus, and F2A from foot-and-mouth disease virus. Among them, P2A has the highest cleavage efficiency and is thus preferred. Functional variants of many of these 2A polypeptides are also known in the art and can also be used in the present invention.
[0073] As used herein, the "cytosine deaminase" refers to a deaminase that can accept single-stranded DNA as a substrate and can catalyze the deamination of cytidine or deoxycytidine to uracil or deoxyuracil, respectively. Examples of cytosine deaminases include, but are not limited to, for example, APOBEC1 deaminase, activation-induced cytidine deaminase (AID), APOBEC3G, CDA1, human APOBEC3A deaminase. In some embodiments, the cytosine deaminase is human APOBEC3A deaminase, for example, the amino acid sequence thereof is shown in SEQ ID NO:1. In some specific embodiments, the cytosine deaminase comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to SEQ ID NO:1, or having one or more conservative amino acid substitutions relative to SEQ ID NO:1, but substantially retaining the function of the protein shown in SEQ ID NO:1.
[0074] As used in the present invention, "nuclease-inactivated CRISPR effector protein" means that the double-stranded nucleic acid cleavage activity of the CRISPR effector protein is absent, yet it still retains the gRNA-guided DNA targeting ability. CRISPR effector proteins lacking double-stranded nucleic acid cleavage activity also encompass nickases, which make a nick in a double-stranded nucleic acid molecule but do not completely cut the double-stranded nucleic acid. In some preferred embodiments of the present invention, the nuclease-inactivated CRISPR effector protein of the present invention has nickase activity.
[0075] In some embodiments, the nuclease-inactivated CRISPR effector protein is nuclease-inactivated Cas9. The DNA cleavage domain of the Cas9 nuclease is known to comprise two subdomains: the HNH nuclease subdomain and the RuvC subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC subdomain cleaves the non-complementary strand. Mutations in these subdomains can inactivate the nuclease activity of Cas9, forming "nuclease-inactivated Cas9". The nuclease-inactivated Cas9 still retains the gRNA-guided DNA binding ability. Thus, in principle, when fused to another protein, nuclease-inactivated Cas9 can simply target the said another protein to almost any DNA sequence by co-expression with a suitable guide RNA.
[0076] The nuclease-inactivated Cas9 of the present invention can be derived from Cas9 of different species. For example, it can be derived from Streptococcus pyogenes Cas9 (SpCas9), or from Staphylococcus aureus Cas9 (SaCas9). Simultaneous mutation of the HNH nuclease subdomain and the RuvC subdomain of Cas9 (e.g., comprising mutations D10A and H840A) inactivates the nuclease of Cas9, becoming dead nuclease Cas9 (dCas9). Mutating and inactivating one of the subdomains can enable Cas9 to have nickase activity, i.e., obtaining Cas9 nickase (nCas9). For example, nCas9 having only the mutation D10A. Thus, in some embodiments of the present invention, the nuclease-inactivated Cas9 of the present invention comprises the amino acid substitutions D10A and / or H840A relative to wild-type Cas9. In some specific embodiments of the present invention, the nuclease-inactivated Cas9 may further comprise additional mutations. For example, nuclease-inactivated SpCas9 may further comprise EQR, VQR or VRER mutations and SaCas9 may further comprise KKH mutations (Kim et al. Nat. Biotechnol. 35, 371-376.).
[0077] In some specific embodiments of the present invention, the nuclease-inactivated CRISPR effector protein comprises the amino acid sequence shown in SEQ ID NO: 3.
[0078] As used herein, Uracil-DNA Glycosylase (UDG) or Uracil-N-glycosylase (UNG) refers to an enzyme that can recognize U bases and remove the N-glycosidic bond of the bases to form apurinic or apyrimidinic sites. The UDG can have different sources, such as from Escherichia coli. In some specific embodiments, the UDG has the amino acid sequence shown in SEQ ID NO:5. In some specific embodiments, the UDG comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity with SEQ ID NO:5, or has one or more conservative amino acid substitutions relative to SEQ ID NO:5, but substantially retains the function of the protein shown in SEQ ID NO:5.
[0079] In some embodiments of the present invention, the cytosine deaminase in the first polypeptide is fused to the N-terminus of the nuclease-inactivated CRISPR effector protein.
[0080] In some embodiments of the present invention, the cytosine deaminase, the nuclease-inactivated CRISPR effector protein, and / or the UDG in the first polypeptide are directly linked. In some embodiments of the present invention, the cytosine deaminase, the nuclease-inactivated CRISPR effector protein, and / or the Rad18 protein in the first polypeptide are directly linked. In some embodiments of the present invention, the cytosine deaminase, the nuclease-inactivated CRISPR effector protein, and / or the UDG in the first polypeptide are linked by a linker. In some embodiments of the present invention, the cytosine deaminase, the nuclease-inactivated CRISPR effector protein, and / or the Rad18 protein in the first polypeptide are linked by a linker. The linker can be a non-functional amino acid sequence of 1-50 (such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or 20-25, 25-50) or more amino acids without secondary or higher structures. For example, the linker can be a flexible linker, such as the X-TENT linker shown in SEQ ID NO:2.
[0081] In the present invention, the proliferating cell nuclear antigen (PCNA) can be PCNA from various species sources. In some embodiments, the PCNA is derived from the species to be base edited. In some embodiments, the ubiquitin protein binding site of PCNA is mutated to prevent its ubiquitination by the endogenous system in the cell. Identifying and / or mutating the ubiquitin protein binding site of PCNA to prevent its native ubiquitination is within the ability of those skilled in the art. In some embodiments, the PCNA is rice PCNA. Wild-type rice PCNA, for example, contains the amino acid sequence shown in SEQ ID NO: 20. In some embodiments, the amino acid sequence of the PCNA with the mutated ubiquitin protein binding site contains an amino acid substitution K164R relative to wild-type PCNA, and the amino acid position is referenced to SEQ ID NO: 20. In some embodiments, the PCNA with the mutated ubiquitin protein binding site contains the amino acid sequence shown in SEQ ID NO: 9.
[0082] In some embodiments, the second polypeptide contains a single ubiquitin protein (monoubiquitination) fused to the PCNA with the mutated ubiquitin protein binding site. The ubiquitin protein can be from various species sources. In some embodiments, the ubiquitin protein is derived from the species to be base edited. In some embodiments, the ubiquitin protein is rice ubiquitin protein. In some embodiments, the single ubiquitin protein is a truncated ubiquitin protein. In some embodiments, the truncated ubiquitin protein contains only the N-terminal functional domain of the ubiquitin protein. In some embodiments, the truncated ubiquitin protein contains the amino acid sequence shown in SEQ ID NO: 10. In some embodiments, the ubiquitin protein is fused to the C-terminus of the PCNA with the mutated ubiquitin protein binding site.
[0083] In some embodiments, the second polypeptide further contains MCP (MS2 coat protein), for example, the MCP is fused to the N-terminus of the PCNA with the mutated ubiquitin protein binding site. Exemplary MCP contains the amino acid sequence shown in SEQ ID NO: 7. In some embodiments, the MCP contains an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity with SEQ ID NO: 7, or its amino acid sequence has one or more conservative amino acid substitutions relative to SEQ ID NO: 7, but substantially retains the function of the protein shown in SEQ ID NO: 7. Correspondingly, in some embodiments, the guide RNA contains the MS2 sequence.
[0084] In some embodiments of the present invention, the mutant PCNA, the ubiquitin protein, and optionally the MCP in the second polypeptide are directly linked. In some embodiments of the present invention, the mutant PCNA, the ubiquitin protein, and optionally the MCP in the second polypeptide are linked by a linker. The linker can be a non-functional amino acid sequence of 1 to 50 (such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or 20 - 25, 25 - 50) or more amino acids without secondary or higher structures.
[0085] "AP endonuclease", "AP lyase", "AP lyase", and "apurinic / apyrimidinic lyase" are used interchangeably herein and refer to an enzyme capable of recognizing an apurinic or apyrimidinic site on a nucleic acid and cleaving the nucleic acid. The AP endonuclease can be an AP endonuclease from various species sources. In some embodiments, the AP endonuclease is derived from the species to be base edited. In some embodiments, the AP endonuclease is a rice AP endonuclease. In some embodiments, the AP endonuclease is rice APE01g. Wild-type rice APE01g contains the amino acid sequence shown in SEQ ID NO:18. In some embodiments, the AP endonuclease is rice APE12g. Wild-type rice APE12g contains the amino acid sequence shown in SEQ ID NO:19.
[0086] In some embodiments, the AP endonuclease is mutated to be inactivated, for example, to retain substrate-binding activity but lose catalytic activity.
[0087] In some embodiments, the mutant AP lyase is derived from rice APE01g and contains the amino acid substitution D297A relative to wild-type APE01g, with the amino acid position referring to SEQ ID NO:18. In some embodiments, the mutant AP lyase contains the amino acid sequence shown in SEQ ID NO:15.
[0088] In some embodiments, the mutant AP lyase is derived from rice APE12g and contains the amino acid substitution D327A relative to wild-type APE12g, with the amino acid position referring to SEQ ID NO:19. In some embodiments, the mutant AP lyase contains the amino acid sequence shown in SEQ ID NO:16.
[0089] In some embodiments, the mutant AP lyase is derived from rice APE12g and contains the amino acid substitutions D238A and N240V relative to wild-type APE12g, with the amino acid positions being referenced to SEQ ID NO:19. In some embodiments, the mutant AP lyase contains the amino acid sequence shown in SEQ ID NO:11.
[0090] In the present invention, the Rad18 protein can be Rad18 proteins from various species. In some embodiments, the Rad18 protein is a human Rad18 protein. In some embodiments, the Rad18 protein contains the amino acid sequence shown in SEQ ID NO:17. In some embodiments, the Rad18 protein contains an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity with SEQ ID NO:17, or its amino acid sequence has one or more conservative amino acid substitutions relative to SEQ ID NO:17, but substantially retains the function of the Rad18 protein shown in SEQ ID NO:17.
[0091] As used herein, "gRNA" and "guide RNA" are used interchangeably and refer to an RNA molecule that can form a complex with a CRISPR nuclease and can target the complex to a target sequence due to a certain complementarity with the target sequence. For example, in a Cas9-based gene editing system, the gRNA is typically composed of a crRNA and a tracrRNA molecule that partially complement to form a complex, where the crRNA contains a sequence having sufficient complementarity with the target sequence to hybridize with the target sequence and direct the CRISPR complex (Cas9 + crRNA + tracrRNA) to specifically bind to the target sequence. However, it is known in the art that single guide RNAs (sgRNAs) can be designed that contain the characteristics of both crRNA and tracrRNA. In a Cpf1-based genome editing system, the gRNA is typically composed only of a mature crRNA molecule (also referred to as sgRNA), where the crRNA contains a sequence having sufficient identity with the target sequence to hybridize with the complementary sequence of the target sequence and direct the complex (Cpf1 + crRNA) to specifically bind to the target sequence. Designing a suitable gRNA based on the CRISPR nuclease used and the target sequence to be edited is within the capabilities of those skilled in the art. As used herein, a "target sequence" is a sequence that is complementary or identical (depending on the different CRISPR nucleases) to the approximately 20-nucleotide guide sequence contained in the guide RNA. The guide RNA targets the target sequence through base pairing between the target sequence or its complementary strand.
[0092] In some embodiments of the present invention, the editing results in one or more nucleotide substitutions of C to G in the target sequence.
[0093] In some embodiments of the present invention, the polypeptide of the present invention further comprises a nuclear localization sequence (NLS). Generally, one or more NLSs in the polypeptide should have sufficient strength to drive the accumulation of the polypeptide in the nucleus of the cell in an amount that enables its gene editing function. Generally, the strength of the nuclear localization activity is determined by the number, position, one or more specific NLSs used in the polypeptide, or a combination of these factors. In some embodiments of the present invention, the NLS of the polypeptide of the present invention can be located at the N-terminus and / or C-terminus or in the middle. In some embodiments, the polypeptide comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs. When there are more than one NLS, each can be selected independently of the other NLSs. In some embodiments, the NLS comprises the amino acid sequence shown in SEQ ID NO: 4 or 8.
[0094] In addition, according to the DNA position to be edited, the polypeptide of the present invention may further include other localization sequences, such as cytoplasmic localization sequences, chloroplast localization sequences, mitochondrial localization sequences, etc.
[0095] In some embodiments, the first polypeptide comprises the amino acid sequence shown in SEQ ID NO: 12. In some embodiments, the fusion protein comprising the first polypeptide and the second polypeptide fused by a self-cleaving peptide (T2A) comprises the amino acid sequence shown in SEQ ID NO: 13. In some embodiments, the fusion protein comprising the first polypeptide and the third polypeptide fused by a self-cleaving peptide (T2A) comprises the amino acid sequence shown in SEQ ID NO: 14. In some embodiments, the fusion protein comprising the first polypeptide and the third polypeptide fused by a self-cleaving peptide (T2A) comprises the amino acid sequence shown in SEQ ID NO: 21. In some embodiments, the fusion protein comprising the first polypeptide and the third polypeptide fused by a self-cleaving peptide (T2A) comprises the amino acid sequence shown in SEQ ID NO: 22. In some embodiments, the fusion protein comprising the first polypeptide and the fourth polypeptide fused by a self-cleaving peptide (T2A) comprises the amino acid sequence shown in SEQ ID NO: 23.
[0096] In order to achieve effective expression in cells, in some embodiments of the present invention, the nucleotide sequence encoding the polypeptide is codon-optimized for the organism from which the cells to be gene-edited are derived.
[0097] Codon optimization refers to a method of modifying a nucleic acid sequence to enhance expression in a host cell of interest by replacing at least one codon of a native sequence with a codon that is more frequently or most frequently used in the genes of the host cell (e.g., about or more than about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons while maintaining the native amino acid sequence. Different species exhibit specific preferences for certain codons of a particular amino acid. Codon bias (differences in codon usage among organisms) is often associated with the translation efficiency of messenger RNA (mRNA), which is thought to depend on the nature of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The prevalence of selected tRNAs within a cell generally reflects the codons most frequently used for peptide synthesis. Thus, genes can be tailored for optimal gene expression based on codon optimization in a given organism. Codon usage tables are readily available, for example, in the Codon Usage Database accessible at www.kazusa.orjp / codon / , and these tables can be adapted in different ways. See, Nakamura Y. et al., “Codon usage tabulated from the international DNA sequence databases: status for the year 2000. Nucl. Acids Res., 28:292 (2000).
[0098] The organism from which the cells that can be gene-edited by the system of the present invention are preferably eukaryotes, including but not limited to, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, cats; poultry such as chickens, ducks, geese; plants including monocotyledonous plants and dicotyledonous plants, such as rice, corn, wheat, sorghum, barley, soybeans, peanuts, Arabidopsis thaliana, etc. Preferably, the organism is a plant, more preferably rice.
[0099] As used herein, an “editing system” refers to a combination of components required for base editing of the genome in a cell or organism. Each component of the system, such as one or more polypeptides or expression constructs encoding them, one or more guide RNAs or expression constructs encoding them, can exist independently, or can exist in any combination as a composition.
[0100] III. Method for Modifying a Target Sequence in a Cell Genome
[0101] In another aspect, the present invention provides a method for modifying a target sequence in a cell genome, comprising introducing the base editing system of the present invention into the cell.
[0102] In some embodiments, the modification results in one or more nucleotide substitutions of C to G in the target sequence. In some embodiments, the modification does not include insertion and / or deletion mutations.
[0103] In another aspect, the present invention also provides a method for generating a genetically modified cell, comprising introducing the gene editing system of the present invention into the cell.
[0104] In another aspect, the present invention also provides a genetically modified organism, which comprises the genetically modified cell or its progeny cells generated by the method of the present invention.
[0105] In the present invention, the target sequence to be modified can be located at any position in the genome, for example, within a functional gene such as a protein-coding gene, or can be located, for example, in a gene expression regulatory region such as a promoter region or an enhancer region, so as to achieve modification of the gene function or modification of gene expression. The modification in the cell target sequence can be detected by T7EI, PCR / RE or sequencing methods.
[0106] In the method of the present invention, the base editing system can be introduced into the cell by various methods well known to those skilled in the art.
[0107] Methods that can be used to introduce the base editing system of the present invention into the cell include, but are not limited to: calcium phosphate transfection, protoplast fusion, electroporation, liposome transfection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus and other viruses), gene gun method, PEG-mediated protoplast transformation, Agrobacterium tumefaciens-mediated transformation.
[0108] Cells that can be subjected to base editing by the method of the present invention can be from, for example, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, cats; poultry such as chickens, ducks, geese; plants, including monocotyledonous plants and dicotyledonous plants, such as rice, corn, wheat, sorghum, barley, soybeans, peanuts, Arabidopsis thaliana, etc. Preferably, the cell is a plant cell, such as a rice cell.
[0109] In some embodiments, the cell is a cell in proliferation and / or differentiation. In some embodiments, the cell is a meristem cell. In some embodiments, the cell is a callus cell.
[0110] In some embodiments, the method of the present invention is carried out in vitro. For example, the cell is an isolated cell, or a cell in an isolated tissue or organ.
[0111] In some other embodiments, the method of the present invention can also be carried out in vivo. For example, the cells are cells in a living organism, and the system of the present invention can be introduced into the cells in vivo by methods mediated by, for example, viruses or Agrobacterium tumefaciens.
[0112] IV. Applications in Plants
[0113] The base editing system of the present invention and the method for modifying a target sequence in a cell genome are particularly suitable for genetically modifying plants. Preferably, the plants are crop plants, including but not limited to wheat, rice, corn, soybean, sunflower, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava, and potato. More preferably, the plant is rice.
[0114] In another aspect, the present invention provides a method for generating a genetically modified plant, comprising introducing the base editing system of the present invention into at least one of the plants, thereby resulting in one or more C to G substitutions of the target sequence in the genome of the at least one plant.
[0115] In some embodiments, the method further comprises screening the at least one plant for plants having the desired one or more C to G substitutions.
[0116] In the method of the present invention, the base editing system can be introduced into plants by various methods well known to those skilled in the art. Methods that can be used to introduce the base editing system of the present invention into plants include but are not limited to: particle bombardment, PEG-mediated protoplast transformation, Agrobacterium tumefaciens-mediated transformation, plant virus-mediated transformation, pollen tube pathway method, and ovary injection method. Preferably, the base editing system is introduced into plants by transient transformation.
[0117] In the method of the present invention, modification of the target sequence can be achieved simply by introducing or generating the polypeptide and the guide RNA in plant cells, and the modification can be stably inherited without stably transforming the plant with an exogenous polynucleotide encoding the components of the base editing system. This avoids the potential off-target effects of the stably present (continuously produced) base editing composition and also avoids the integration of exogenous nucleotide sequences into the plant genome, thus having higher biosafety.
[0118] In some preferred embodiments, the introduction is carried out in the absence of a selection pressure, thereby avoiding the integration of exogenous nucleotide sequences into the plant genome.
[0119] In some embodiments, the introduction comprises transforming the base editing system of the present invention into isolated plant cells or tissues, and then regenerating the transformed plant cells or tissues into whole plants. Preferably, the regeneration is carried out in the absence of a selection pressure, that is, no selection agent for the selection gene carried on the expression vector is used during the tissue culture process. Not using a selection agent can improve the regeneration efficiency of plants and obtain modified plants without exogenous nucleotide sequences.
[0120] In other embodiments, the base editing system of the present invention can be transformed into specific parts of a whole plant, such as leaves, shoot tips, pollen tubes, young panicles or hypocotyls. This is particularly suitable for the transformation of plants that are difficult to regenerate by tissue culture.
[0121] In some embodiments, the system is introduced into tissues / cells in plant proliferation and / or differentiation. In some embodiments, the system is introduced into meristems. In some embodiments, the system is introduced into callus.
[0122] In some embodiments of the present invention, the in vitro expressed protein and / or in vitro transcribed RNA molecule (for example, the expression construct is an in vitro transcribed RNA molecule) is directly transformed into the plant. The protein and / or RNA molecule can achieve base editing in plant cells and is subsequently degraded by the cells, avoiding the integration of exogenous nucleotide sequences into the plant genome.
[0123] Therefore, in some embodiments, using the method of the present invention for genetic modification and breeding of plants can obtain plants whose genomes have no integration of exogenous polynucleotides, that is, modified plants that are transgene-free.
[0124] In some embodiments of the present invention, the modified target sequence is related to plant traits such as agronomic traits, whereby the one or more C to G substitutions result in the plant having altered (preferably improved) traits relative to the wild-type plant, such as agronomic traits.
[0125] In some embodiments, the method further comprises the step of screening plants having the desired one or more C to G substitutions and / or desired traits such as agronomic traits.
[0126] In some embodiments of the present invention, the method further comprises obtaining progeny of the genetically modified plant. Preferably, the genetically modified plant or its progeny has the desired one or more C to G substitutions and / or desired traits such as agronomic traits.
[0127] On the other hand, the present invention also provides a genetically modified plant or its progeny or a part thereof, wherein the plant is obtained by the method described above in the present invention. In some embodiments, the genetically modified plant or its progeny or a part thereof is non-transgenic. Preferably, the genetically modified plant or its progeny has the desired genetic modification and / or the desired traits such as agronomic traits.
[0128] On the other hand, the present invention also provides a plant breeding method, which includes crossing a genetically modified first plant containing one or more C to G substitutions in the target nucleic acid region obtained by the method described above in the present invention with a second plant that does not contain the one or more nucleotide substitutions, so as to introduce the one or more nucleotide substitutions into the second plant. Preferably, the genetically modified first plant has the desired traits such as agronomic traits.
[0129] V. Kit
[0130] The present invention also includes a kit for the method of the present invention, which kit includes the base editing system of the present invention and instructions for use. The kit generally includes a label indicating the intended use and / or method of use of the contents of the kit. The term label includes any written or recorded material provided on or with the kit or otherwise provided with the kit. Examples
[0131] Materials and Methods
[0132] 1. Vector Construction
[0133] To construct the PCGBE system, UDG of Escherichia coli (accession number: AMB53293.1), rice's own PCNA (LOC_Os02g56130), Ubiquitin (LOC_Os03g13170), rice's own APE01g (LOC_Os01g58690) and APE12g (LOC_Os12g18200), as well as human-derived hRad18 protein (NP_064550.3) were selected. The PCNA sequence was subjected to a K164R point mutation (PCNA(K164R)), the N-terminal functional domain (Ub) of Ubiquitin was intercepted, and the APE12g sequence was subjected to an inactivating point mutation (dAPE, D238A / N240V), and point mutations were performed on APE01g and APE12g to obtain mAPE01g (D297A) and mAPE12g (D327A). All gene fragments were optimized for rice codons and gene synthesized by Nanjing Genscript Biotechnology Co., Ltd.
[0134] APOBEC3A was fused to the N-terminus of Cas9 with an XTEN linker, and UDG was fused to the C-terminus of Cas9, thereby constructing the PCGBE-1 vector (Figure 2 ) or ADU vector ( Figure 11 a).
[0135] Replace nCas9(D10A) in the PCGBE-1 vector with dCas9(D10A,H840A) to obtain PCGBE-2; secondly, use the self-cleaving 2A polypeptide (T2A) to fuse mAPE01g, mAPE12g, and hRad18 proteins to the C-terminus of the PCGBE-1 vector respectively, thereby constructing PCGBE-3 to 5 vectors; swap the positions of UDG and hRad18 proteins in the PCGBE-5 vector to obtain the PCGBE-6 vector. Finally, use the Gibson method to integrate the fusion gene fragment together with the sgRNA expression component into the pHUE411 backbone to construct the binary vectors pH-PCGBE-2 to 6( Figure 6 ), which is used for Agrobacterium-mediated rice genetic transformation.
[0136] In addition, use the self-cleaving 2A polypeptide (T2A) to fuse the MCP-PCNA-Ub fusion protein and dAPE protein to the C-terminus of the APOBEC3A-nCas9-UDG vector respectively, thereby constructing the Figure 11 CGBE shown in b and Figure 11 the CGBE vector shown in c. Finally, use the Gibson method to integrate the fusion gene fragment together with the sgRNA expression component into the pHUE411 backbone to construct a binary vector for Agrobacterium-mediated rice genetic transformation.
[0137] Select 9 endogenous targets from 9 rice genes (OsAAT, OsALS, OsCDC48, OsDEP1, OsGRF1, OsIPA1, OsNRT1.1B, OsPDS, and OsSWEET14) for systematic testing, and the sequences of all targeted sites are shown in Table 1.
[0138] Table 1. sgRNA Target Sites and Primers
[0139]
[0140] The PAM sequence is shown in bold
[0141] 2. Protoplast Isolation and Transformation
[0142] The rice material used for protoplast isolation and transformation in the present invention is Zhonghua 11.
[0143] 2.1 Rice Etiolated Seedling Culture
[0144] The Zhonghua 11 rice seeds were first rinsed with 75% ethanol for 1 minute, then treated with 4% sodium hypochlorite for 30 minutes, and washed with sterile water more than 5 times. They were cultured on M6 medium for 3 - 4 weeks at 26°C in the dark.
[0145] 2.2 Isolation of rice protoplasts
[0146] (1) Cut off the stem tissue of the etiolated seedlings, cut the middle part into filaments of 0.5 - 1 mm with a blade, place them in 0.6 M Mannitol solution and treat in the dark for 10 min, then filter through a sieve, put them into 50 mL enzyme digestion solution (filtered through a 0.45 μm filter membrane), evacuate (pressure about 15 Kpa) for 30 min, take out and place on a shaker (10 rpm) for enzymatic digestion at room temperature for 5 h; (2) Add 30 - 50 mL of W5 to dilute the enzymatic digestion product, filter the enzymatic digestion solution through a 75 μm nylon filter membrane into a round-bottom centrifuge tube (50 mL); (3) Centrifuge at 23°C, 250 g (rcf), increase 3 and decrease 3, for 3 min, discard the supernatant; (4) Gently suspend the cells with 20 mL of W5, repeat step (3); (5) Suspend with an appropriate amount of MMG and wait for transformation.
[0147] 2.3 Transformation of rice protoplasts
[0148] (1) Add 10 μg of each required transformation vector to a 2 mL centrifuge tube respectively. After mixing, use a pipette tip without the tip to aspirate 200 μL of protoplasts, gently flick to mix, add 220 μL of PEG4000 solution, gently flick to mix, and induce transformation at room temperature in the dark for 20 - 30 min; (2) Add 880 μL of W5 and gently invert to mix, centrifuge at 250 g (rcf), increase 3 and decrease 3, for 3 min, discard the supernatant; (3) Add 1 mL of WI solution, gently invert to mix, and culture in the dark at 23°C for 48 h.
[0149] 3. Agrobacterium-mediated genetic transformation of rice
[0150] The binary vector constructed into the target site was delivered into Agrobacterium tumefaciens AGL1 by electrotransformation. Then, the callus of Zhonghua 11 rice was transformed by the method of Agrobacterium infection, and hygromycin was used as a selection marker for screening transgenic positive plants.
[0151] 4. DNA extraction and amplicon sequencing analysis of protoplasts and transgenic plants
[0152] 3.1 DNA extraction of protoplasts and transgenic plants
[0153] Collect the protoplasts into a 2 mL centrifuge tube, and extract the protoplast DNA (~30 μL) using the CTAB method; sample each transgenic clone separately, extract its genomic DNA using the CTAB method, and measure its concentration (~50 ng / μL) using a NanoDrop ultra-micro spectrophotometer, and store at -20°C.
[0154] 3.2 Amplicon sequencing analysis
[0155] (1) Use genomic primers to perform one round of PCR amplification on the DNA template. The primer information for one round is shown in Table 2. The 20 μL amplification system contains 4 μL of 5×Fastpfu buffer, 1.6 μL of dNTPs (2.5 mM), 0.4 μL of Forward primer (10 μM), 0.4 μL of Reverse primer (10 μM), 0.4 μL of FastPfu polymerase (2.5 U / μL), and 2 μL of DNA template (~60 ng). Amplification conditions: pre-denaturation at 95°C for 5 min; denaturation at 95°C for 30 s, annealing at 50 - 64°C for 30 s, extension at 72°C for 30 s, for 35 cycles; full extension at 72°C for 5 min, store at 12°C;
[0156] (2) Dilute the above amplification product by 10 times, and take 1 μL as the template for the second round of PCR amplification. The amplification primers are sequencing primers containing Barcode, as shown in Table 2 for details. The 50 μL amplification system contains 10 μL of 5×Fastpfu buffer, 4 μL of dNTPs (2.5 mM), 1 μL of Forward primer (10 μM), 1 μL of Reverse primer (10 μM), 1 μL of FastPfupolymerase (2.5 U / μL), and 1 μL of DNA template. The amplification conditions are as above, and the number of amplification cycles is 38 cycles.
[0157] (3) Separate the PCR products by 2% agarose gel electrophoresis, and use the AxyPrepTM DNA Gel Extraction kit to recover the target fragments from the gel. The recovered products are quantitatively analyzed using a NanoDrop ultra-micro spectrophotometer; Take 100 ng of the recovered products respectively for mixing, and send them to Beijing Novogene Bioinformatics Technology Co., Ltd. for amplicon library construction and sequencing analysis.
[0158] (4) After the sequencing is completed, split the original data according to the sequencing primers. At the same time, use the sgRNA sequence and its flanking sequences as the reference sequences to systematically compare and analyze different system editing product types and editing efficiencies at different gene targeting sites.
[0159] Example 1: Construction of a precise CG base editing system
[0160] As early as 2016, David Liu's laboratory had established a cytosine base editor (CBE). The system uses nCas9 (D10A) to guide cytosine deaminase to act on the non-complementary chain of the DNA target site, and deaminates cytosine (C) in a specific area into uracil base (U). U is recognized as thymine (T) during mismatch repair (MMR) or DNA replication, ultimately achieving precise single-base replacement of C-to-T. However, during the body's base excision repair (BER) process, U will be recognized and cut by UDG to form an AP site, which in turn forms an incision under the action of AP nuclease (APE), which leads to the occurrence of some indel by-products. However, it was also found that the CBE system produced some low-frequency C-to-G by-products. Therefore, after fusing the CBE system with UGI, the production of various by-products is also greatly reduced ( Figure 1 ).
[0161] As early as the beginning of 2019, the inventors replaced the UGI in the efficient A3A-PBE system with UDG and built the initial PCGBE system (PCGBE-1) ( Figure 2 ) and tested it in rice protoplasts. The results showed that only a very low frequency of C-to-G base substitutions was detected, and the main frequency was C-to-T base editing ( Figure 3 ). They then tried to test it in rice callus and found that PCGBE-1 mediated about 30% of C-to-G transversions in rice callus ( Figure 4 and 5 This also fully demonstrates that the CGBE-mediated C-to-G base transversion process is dependent on the DNA replication process. In addition to mediating C-to-G transversion, the PCGBE-1 system also produces a high proportion of indel byproducts (about 30% or more) ( Figure 4 and 5 ), which is somewhat similar to the results published by Kurt et al. in 2020. The values in the figure represent the C-to-G editing efficiency; the values in the brackets below represent the C-to-G purity.
[0162] The present invention studies the DNA damage occurrence and potential repair mechanism involved in the process from cytosine deamination to C-to-G generation ( Figure 1), from which a key repair pathway has been identified: translesion DNA synthesis (TLS) repair (Zhuang et al., 2008; Qin et al., 2013; Martin and Wood, 2019). Previous studies have shown that TLS specifically inserts a nucleotide, likely dCTP, into the synthesized strand corresponding to the DNA damage site (e.g., the AP site), primarily used to initiate the TLS repair signal, thus opening the way for C-to-G base editing. However, a key protein factor in TLS repair is proliferating cell nuclear antigen (PCNA), which recognizes DNA damage during DNA replication and mediates diverse DNA repair processes. When monoubiquitinated, PCNA promotes the recruitment of the relevant DNA polymerases (Polη and Polζ) during TLS repair, thereby promoting damage-bypass repair. However, when polyubiquitinated, PCNA participates in the damage-avoidance pathway, which uses the intact sister monomer as a template (Zhuang et al., 2008; Qin et al., 2013; Martin and Wood, 2019). Rad6-Rad18 play an important role in responding to DNA damage and promote the monoubiquitination of PCNA, thereby directing DNA damage repair toward the TLS pathway. However, through comparative analysis of animal and plant protein information, the inventors did not identify a Rad18 homologous protein in the rice genome, and speculate that this may be one of the reasons for the low C-to-G efficiency and purity in plants. Therefore, co-expression of human Rad18 protein is likely to direct mutations toward TLS repair, thereby achieving precise C-to-G base transversion within the target sequence.
[0163] In addition, in the absence of exogenous AP endonucleases, only a portion of the AP site will be cleaved, and the remaining AP site will enter the TLS bypass repair process during DNA replication. Once the AP site is generated, it will also induce the expression of endogenous APE in a short period of time to achieve the purpose of base excision repair (BER). Therefore, mutating the cell's own APE protein to lose its catalytic activity and retain only binding activity, and co-expressing this mutant APE protein (mAPE) in the cell, can competitively inhibit endogenous APE expression and protect AP sites. This strategy can also enable mutations to proceed in the direction of TLS repair and even C-to-G base transversion.
[0164] Based on the above results and the TLS repair mechanism, the inventors formed PCGBE-2 to 6 systems ( Figure 6), and systematically tested the OsNRT1.1B, OsPDS, and OsGRF1 targets using rice callus tissue. The results showed that PCGBE-4 and PCGBE-5 significantly improved both C-to-G editing efficiency and editing purity compared to PCGBE-1 ( Figures 7 - 10 ).
[0165] In addition, the lysine at position 164 of PCNA was converted to arginine (K164R) through point mutation, destroying the ubiquitination protein binding site; at the same time, a ubiquitination protein (Ubiquitin) N-terminal domain was fused to construct a monoubiquitinated proliferating cell nuclear antigen fusion protein (PCNA·Ub), which can stably recruit TLS-related polymerases to induce TLS repair. Therefore, by co-expressing the cell's own PCNA·Ub fusion protein, the mutation is likely to proceed towards TLS repair, thereby achieving precise C-to-G base transversion within the target sequence.
[0166] Based on the above results and the TLS repair mechanism, the inventors used the cell's own PCNA and Ubiquitin genes as templates, performed point mutations, truncations, splicing and synthesis, and then constructed a fusion protein to the C-terminus of the PCGBE-1 system to form Figure 11 b PCGBE system. In addition, the inventors also constructed the inactivated rice APE (D238A / N240V) to the C-terminus of the PCGBE-1 system through T2A to form Figure 11 The CGBE system shown in C. Figure 11 The base editing system was introduced into rice callus. The mutation detection results are shown in Figure 12.
[0167] The results showed that the CG base editing system of the present invention not only maintained a high ratio of C-to-G editing, but also greatly reduced the production of byproducts such as indels and C-to-A. This also shows that the improved PCGBE system can achieve efficient and accurate C-to-G base transversion.
[0168] sequence:
[0169] SEQ ID NO:1 APOBEC3A
[0170] MEASPASGPRHLMDPHIFTSNFNNGIGRHKTYLCYEVERLDNGTSVKMDQHRGFLHNQAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFKHCWDTFVDHQGCPFQPWDGLDEHSQALSGRLRAILQNQGN
[0171] SEQ ID NO:2 XTEN linker
[0172] SGSETPGTSESATPES
[0173] SEQ ID NO:3 nCas9(D10A)
[0174]
[0175] SEQ ID NO:4 nucleoplasmin NLS
[0176] KRPAATKKAGQAKKKK
[0177] SEQ ID NO:5 E - coil UDG
[0178] ANELTWHDVLAEEKQQPYFLNTLQTVASERQSGVTIYPPQKDVFNAFRFTELGDVKVVILGQDPYHGPGQAHGLAFSVRPGIAIPPSLLNMYKELENTIPGFTRPNHGYLESWARQGVLLLNTVLTVRAGQAHSHASLGWETFTDKVISLINQHREGVVFLLWGSHAQKKGAIIDKQRHHVLKAPHPSPLSAHRGFFGCNHFVLANQWLEQRGETPIDWMPVLPAESE
[0179] SEQ ID NO:6 T2A linker
[0180] EGRGSLLTCGDVEENPGP
[0181] SEQ ID NO:7 MCP
[0182] ASNFTQFVLVDNGGTGDVTVAPSNFANGIAEWISSNSRSQAYKVTCSVRQSSAQNRKYTIKVEVPKGAWRSYLNMELTIPIFATNSDCELIVKAMQGLLKDGNPIPSAIAANSGIY
[0183] SEQ ID NO:8 SV40 NLS
[0184] PKKKRKV
[0185] SEQ ID NO:9 OsPCNA(K164R)
[0186] MLELRLVQGSLLKKVLEAIRELVTDANFDCSGTGFSLQAMDSSHVALVALLLRSEGFEHYRCDRNLSMGMNLNNMAKMLRCAGNDDIITIKADDGSDTVTFMFESPNQDKIADFEMKLMDIDSEHLGIPDSEYQAIVRMPSSEFSRICKDLSSIGDTVIISVTREGVKFSTAGDIGTANIVCRQNKTVDKPEDATIIEMQEPVSLTFALRYMNSFTKASPLSEQVTISLSSELPVVVEYKIAEMGYIRFYLAPKIEEDEEMKS
[0187] SEQ ID NO:10 truncated Ubiquitin(Ub)
[0188] MQIFVKTLTGKTITLEVESSDTIDNVKAKIQDKEGIPPDQQRLIFAGKQLEDGRTLADYNIQKESTLHLVLRLR
[0189] SEQ ID NO:11 OsAPE(D238A / N240V), dAPE
[0190] MKRFFQPVPKDGSPAKKRPAAAAAASASDSDSLGGDAPAAAACAVGEGDSPPAPREEEPRRFVTWNANSLLLRMKSDWPAFCQFVSRVDPDVICVQEVRMPAAGSKGAPKNPGQLKDDTSSSRDEKQVVLRALSSPPFKDYRVWWSLSDSKYAGTAMIIKKKFEPKKVSFNLDRTSSKHEPDGRVIIAEFESFLLLNTYAPNNGWKEEENSFQRRRKWDKRMLEFVQQVDKPLIWCGALVVSHEEIDVSHPDFFSSAKLNGYIPPNKEDCGQPGFTLSERRRFGNILSQGKLVDAYRYLHKEKDMDCGFSWSGHPIGKYRGKRMRIDYFLVSEKLKDQIVSCDIHGRGIELEGFYGSDHCPVSLELSEEVEAPKPKSSN
[0191] SEQ ID NO:12 PCGBE-1 system exemplary polypeptide
[0192]
[0193] SEQ ID NO:13 Schematic polypeptide of the PCGBE system ( Figure 11 b)
[0194]
[0195] SEQ ID NO:14 PCGBE system-schematic polypeptide( Figure 11 c)
[0196]
[0197] SEQ ID NO:15 mAPE01g(D297A)
[0198] SAIRASSHRLQTRTVALTRTKMSSMAGLGASQHGYPPRSHEPWTKLVHRERLPEWFAYNPKTMRPPPLSHDTKCMKILSWNINGLHDVVTTKGFSARDLAQRENFDVLCLQETHLEEKDVEKFKNLIADYDSYWSCSVSRLGYSGTAVISRVKPISVQYGIGIREHDHEGRVITLEFDGFYLVNAYVPNSGRFLRRLNYRVNNWDPCFSNYVKILEKSKPVIVAGDLNCARQSIDIHNPPAKTKSAGFTIEERESFETNFSSKGLVDTFRKQHPNAVGYTFWGENQRITNKGWRLAYFLASESITDKVHDSYILPDVSFSDHSPIGLVLKL
[0199] SEQ ID NO:16 mAPE12g(D327A)
[0200] KRFFQPVPKDGSPAKKRPAAAAAASASDSDSLGGDAPAAAACAVGEGDSPPAPREEEPRRFVTWNANSLLLRMKSDWPAFCQFVSRVDPDVICVQEVRMPAAGSKGAPKNPGQLKDDTSSSRDEKQVVLRALSSPPFKDYRVWWSLSDSKYAGTAMIIKKKFEPKKVSFNLDRTSSKHEPDGRVIIAEFESFLLLNTYAPNNGWKEEENSFQRRRKWDKRMLEFVQQVDKPLIWCGDLNVSHEEIDVSHPDFFSSAKLNGYIPPNKEDCGQPGFTLSERRRFGNILSQGKLVDAYRYLHKEKDMDCGFSWSGHPIGKYRGKRMRIAYFLVSEKLKDQIVSCDIHGRGIELEGFYGSDHCPVSLELSEEVEAPKPKSSN
[0201] SEQ ID NO:17 hRad18
[0202] DSLAESRWPPGLAVMKTIDDLLRCGICFEYFNIAMIIPQCSHNYCSLCIRKFLSYKTQCPTCCVTVTEPDLKNNRILDELVKSLNFARNHLLQFALESPAKSPASSSSKNLAVKVYTPVASRQSLKQGSRLMDNFLIREMSGSTSELLIKENKSKFSPQKEASPAAKTKETRSVEEIAPDPSEAKRPEPPSTSTLKQVTKVDCPVCGVNIPESHINKHLDSCLSREEKKESLRSSVHKRKPLPKTVYNLLSDRDLKKKLKEHGLSIQGNKQQLIKRHQEFVHMYNAQCDALHPKSAAEIVREIENIEKTRMRLEASKLNESVMVFTKDQTEKEIDEIHSKYRKKHKSEFQLLVDQARKGYKKIAGMSQKTVTITKEDESTEKLSSVCMGQEDNMTSVTNHFSQSKLDSPEELEPDREEDSSSCIDIQEVLSSSESDSCNSSSSDIIRDLLEEEEAWEASHKNDLQDTEISPRQNRRTRAAESAEIEPRNKRNRN
[0203] SEQ ID NO:18 APE01g
[0204] MSAIRASSHRLQTRTVALTRTKMSSMAGLGASQHGYPPRSHEPWTKLVHRERLPEWFAYNPKTMRPPPLSHDTKCMKILSWNINGLHDVVTTKGFSARDLAQRENFDVLCLQETHLEEKDVEKFKNLIADYDSYWSCSVSRLGYSGTAVISRVKPISVQYGIGIREHDHEGRVITLEFDGFYLVNAYVPNSGRFLRRLNYRVNNWDPCFSNYVKILEKSKPVIVAGDLNCARQSIDIHNPPAKTKSAGFTIEERESFETNFSSKGLVDTFRKQHPNAVGYTFWGENQRITNKGWRLDYFLASESITDKVHDSYILPDVSFSDHSPIGLVLKL
[0205] SEQ ID NO:19 APE12g
[0206] MKRFFQPVPKDGSPAKKRPAAAAAASASDSDSLGGDAPAAAACAVGEGDSPPAPREEEPRRFVTWNANSLLLRMKSDWPAFCQFVSRVDPDVICVQEVRMPAAGSKGAPKNPGQLKDDTSSSRDEKQVVLRALSSPPFKDYRVWWSLSDSKYAGTAMIIKKKFEPKKVSFNLDRTSSKHEPDGRVIIAEFESFLLLNTYAPNNGWKEEENSFQRRRKWDKRMLEFVQQVDKPLIWCGDLNVSHEEIDVSHPDFFSSAKLNGYIPPNKEDCGQPGFTLSERRRFGNILSQGKLVDAYRYLHKEKDMDCGFSWSGHPIGKYRGKRMRIDYFLVSEKLKDQIVSCDIHGRGIELEGFYGSDHCPVSLELSEEVEAPKPKSSN
[0207] SEQ ID NO:20 OsPCNA(wt)
[0208] MLELRLVQGSLLKKVLEAIRELVTDANFDCSGTGFSLQAMDSSHVALVALLLRSEGFEHYRCDRNLSMGMNLNNMAKMLRCAGNDDIITIKADDGSDTVTFMFESPNQDKIADFEMKLMDIDSEHLGIPDSEYQAIVRMPSSEFSRICKDLSSIGDTVIISVTKEGVKFSTAGDIGTANIVCRQNKTVDKPEDATIIEMQEPVSLTFALRYMNSFTKASPLSEQVTISLSSELPVVVEYKIAEMGYIRFYLAPKIEEDEEMKS
[0209] SEQ ID NO:21 Exemplary fusion protein of the PCGBE-3 system
[0210]
[0211] SEQ ID NO:22 PCGBE-4 system schematic fusion protein
[0212]
[0213] SEQ ID NO:23 PCGBE-5 system schematic fusion protein
[0214] Sequence Listing <110> Shanghai Blue Cross Medical Science Research Institute <120> Improved CG Base Editing System <130> P2022TC1988 <160> 23 <170> PatentIn version 3.5 <210> 1 <211> 199 <212> PRT <213> Artificial Sequence <220> <223> APOBEC3A <400> 1 Met Glu Ala Ser Pro Ala Ser Gly Pro Arg His Leu Met Asp Pro His 1 5 10 15 Ile Phe Thr Ser Asn Phe Asn Asn Gly Ile Gly Arg His Lys Thr Tyr 20 25 30 Leu Cys Tyr Glu Val Glu Arg Leu Asp Asn Gly Thr Ser Val Lys Met 35 40 45 Asp Gln His Arg Gly Phe Leu His Asn Gln Ala Lys Asn Leu Leu Cys 50 55 60 Gly Phe Tyr Gly Arg His Ala Glu Leu Arg Phe Leu Asp Leu Val Pro 65 70 75 80 Ser Leu Gln Leu Asp Pro Ala Gln Ile Tyr Arg Val Thr Trp Phe Ile 85 90 95 Ser Trp Ser Pro Cys Phe Ser Trp Gly Cys Ala Gly Glu Val Arg Ala 100 105 110 Phe Leu Gln Glu Asn Thr His Val Arg Leu Arg Ile Phe Ala Ala Arg 115 120 125 Ile Tyr Asp Tyr Asp Pro Leu Tyr Lys Glu Ala Leu Gln Met Leu Arg 130 135 140 Asp Ala Gly Ala Gln Val Ser Ile Met Thr Tyr Asp Glu Phe Lys His 145 150 155 160 Cys Trp Asp Thr Phe Val Asp His Gln Gly Cys Pro Phe Gln Pro Trp 165 170 175 Asp Gly Leu Asp Glu His Ser Gln Ala Leu Ser Gly Arg Leu Arg Ala 180 185 190 Ile Leu Gln Asn Gln Gly Asn 195 <210> 2 <211> 16 <212> PRT <213> Artificial Sequence <220> <223> XTEN linker <400> 2 Ser Gly Ser Glu Thr Pro Gly Thr Ser Glu Ser Ala Thr Pro Glu Ser 1 5 10 15 <210> 3 <211> 1367 <212> PRT <213> Artificial Sequence <220> <223> nCas9 (D10A) <400> 3 Asp Lys Lys Tyr Ser Ile Gly Leu Ala Ile Gly Thr Asn Ser Val Gly 1 5 10 15 Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe Lys 20 25 30 Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile Gly 35 40 45 Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu Lys 50 55 60 Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys Tyr 65 70 75 80 Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser Phe 85 90 95 Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys His 100 105 110 Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr His 115 120 125 Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp Ser 130 135 140 Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His Met 145 150 155 160 Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro Asp 165 170 175 Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr Asn 180 185 190 Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala Lys 195 200 205 Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn Leu 210 215 220 Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn Leu 225 230 235 240 Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe Asp 245 250 255 Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp Asp 260 265 270 Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp Leu 275 280 285 Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp Ile 290 295 300 Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser Met 305 310 315 320 Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys Ala 325 330 335 Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe Asp 340 345 350 Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser Gln 355 360 365 Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp Gly 370 375 380 Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg Lys 385 390 395 400 Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu Gly 405 410 415 Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe Leu 420 425 430 Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile Pro 435 440 445 Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp Met 450 455 460 Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu Val 465 470 475 480 Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr Asn 485 490 495 Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser Leu 500 505 510 Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys Tyr 515 520 525 Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln Lys 530 535 540 Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr Val 545 550 555 560 Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp Ser 565 570 575 Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly Thr 580 585 590 Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp Asn 595 600 605 Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr Leu 610 615 620 Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala His 625 630 635 640 Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr Thr 645 650 655 Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp Lys 660 665 670 Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe Ala 675 680 685 Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe Lys 690 695 700 Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu His 705 710 715 720 Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly Ile 725 730 735 Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly Arg 740 745 750 His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln Thr 755 760 765 Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile Glu 770 775 780 Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro Val 785 790 795 800 Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln 805 810 815 Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg Leu 820 825 830 Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys Asp 835 840 845 Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg Gly 850 855 860 Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys Asn 865 870 875 880 Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe 885 890 895 Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp Lys 900 905 910 Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr Lys 915 920 925 His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys Tyr Asp Glu 930 935 940 Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Leu Lys Ser Lys 945 950 955 960 Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg Glu 965 970 975 Ile Asn Asn Tyr His His Ala His Asp Ala Tyr Leu Asn Ala Val Val 980 985 990 Gly Thr Ala Leu Ile Lys Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val 995 1000 1005 Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala Lys 1010 1015 1020 Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr 1025 1030 1035 Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn 1040 1045 1050 Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr 1055 1060 1065 Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val Arg 1070 1075 1080 Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr Glu 1085 1090 1095 Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg 1100 1105 1110 Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys 1115 1120 1125 Lys Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val Leu 1130 1135 1140 Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser 1145 1150 1155 Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser Phe 1160 1165 1170 Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu 1175 1180 1185 Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu Phe 1190 1195 1200 Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly Glu 1205 1210 1215 Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn 1220 1225 1230 Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser Pro 1235 1240 1245 Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys His 1250 1255 1260 Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys Arg 1265 1270 1275 Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr 1280 1285 1290 Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile 1295 1300 1305 Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe 1310 1315 1320 Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr 1325 1330 1335 Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr Gly 1340 1345 1350 Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 <210> 4 <211> 16 <212> PRT <213> Artificial Sequence <220> <223> nucleoplasmin NLS <400> 4 Lys Arg Pro Ala Ala Thr Lys Lys Ala Gly Gln Ala Lys Lys Lys Lys 1 5 10 15 <210> 5 <211> 228 <212> PRT <213> Artificial Sequence <220> <223> E-coil UDG <400> 5 Ala Asn Glu Leu Thr Trp His Asp Val Leu Ala Glu Glu Lys Gln Gln 1 5 10 15 Pro Tyr Phe Leu Asn Thr Leu Gln Thr Val Ala Ser Glu Arg Gln Ser 20 25 30 Gly Val Thr Ile Tyr Pro Pro Gln Lys Asp Val Phe Asn Ala Phe Arg 35 40 45 Phe Thr Glu Leu Gly Asp Val Lys Val Val Ile Leu Gly Gln Asp Pro 50 55 60 Tyr His Gly Pro Gly Gln Ala His Gly Leu Ala Phe Ser Val Arg Pro 65 70 75 80 Gly Ile Ala Ile Pro Pro Ser Leu Leu Asn Met Tyr Lys Glu Leu Glu 85 90 95 Asn Thr Ile Pro Gly Phe Thr Arg Pro Asn His Gly Tyr Leu Glu Ser 100 105 110 Trp Ala Arg Gln Gly Val Leu Leu Leu Asn Thr Val Leu Thr Val Arg 115 120 125 Ala Gly Gln Ala His Ser His Ala Ser Leu Gly Trp Glu Thr Phe Thr 130 135 140 Asp Lys Val Ile Ser Leu Ile Asn Gln His Arg Glu Gly Val Val Phe 145 150 155 160 Leu Leu Trp Gly Ser His Ala Gln Lys Lys Gly Ala Ile Ile Asp Lys 165 170 175 Gln Arg His His Val Leu Lys Ala Pro His Pro Ser Pro Leu Ser Ala 180 185 190 His Arg Gly Phe Phe Gly Cys Asn His Phe Val Leu Ala Asn Gln Trp 195 200 205 Leu Glu Gln Arg Gly Glu Thr Pro Ile Asp Trp Met Pro Val Leu Pro 210 215 220 Ala Glu Ser Glu 225 <210> 6 <211> 18 <212> PRT <213> Artificial Sequence <220> <223> T2A linker <400> 6 Glu Gly Arg Gly Ser Leu Leu Thr Cys Gly Asp Val Glu Glu Asn Pro 1 5 10 15 Gly Pro <210> 7 <211> 116 <212> PRT <213> Artificial Sequence <220> <223> MCP <400> 7 Ala Ser Asn Phe Thr Gln Phe Val Leu Val Asp Asn Gly Gly Thr Gly 1 5 10 15 Asp Val Thr Val Ala Pro Ser Asn Phe Ala Asn Gly Ile Ala Glu Trp 20 25 30 Ile Ser Ser Asn Ser Arg Ser Gln Ala Tyr Lys Val Thr Cys Ser Val 35 40 45 Arg Gln Ser Ser Ala Gln Asn Arg Lys Tyr Thr Ile Lys Val Glu Val 50 55 60 Pro Lys Gly Ala Trp Arg Ser Tyr Leu Asn Met Glu Leu Thr Ile Pro 65 70 75 80 Ile Phe Ala Thr Asn Ser Asp Cys Glu Leu Ile Val Lys Ala Met Gln 85 90 95 Gly Leu Leu Lys Asp Gly Asn Pro Ile Pro Ser Ala Ile Ala Ala Asn 100 105 110 Ser Gly Ile Tyr 115 <210> 8 <211> 7 <212> PRT <213> Artificial Sequence <220> <223> SV40 NLS <400> 8 Pro Lys Lys Lys Arg Lys Val 1 5 <210> 9 <211> 263 <212> PRT <213> Artificial Sequence <220> <223> OsPCNA (K164R) <400> 9 Met Leu Glu Leu Arg Leu Val Gln Gly Ser Leu Leu Lys Lys Val Leu 1 5 10 15 Glu Ala Ile Arg Glu Leu Val Thr Asp Ala Asn Phe Asp Cys Ser Gly 20 25 30 Thr Gly Phe Ser Leu Gln Ala Met Asp Ser Ser His Val Ala Leu Val 35 40 45 Ala Leu Leu Leu Arg Ser Glu Gly Phe Glu His Tyr Arg Cys Asp Arg 50 55 60 Asn Leu Ser Met Gly Met Asn Leu Asn Asn Met Ala Lys Met Leu Arg 65 70 75 80 Cys Ala Gly Asn Asp Asp Ile Ile Thr Ile Lys Ala Asp Asp Gly Ser 85 90 95 Asp Thr Val Thr Phe Met Phe Glu Ser Pro Asn Gln Asp Lys Ile Ala 100 105 110 Asp Phe Glu Met Lys Leu Met Asp Ile Asp Ser Glu His Leu Gly Ile 115 120 125 Pro Asp Ser Glu Tyr Gln Ala Ile Val Arg Met Pro Ser Ser Glu Phe 130 135 140 Ser Arg Ile Cys Lys Asp Leu Ser Ser Ile Gly Asp Thr Val Ile Ile 145 150 155 160 Ser Val Thr Arg Glu Gly Val Lys Phe Ser Thr Ala Gly Asp Ile Gly 165 170 175 Thr Ala Asn Ile Val Cys Arg Gln Asn Lys Thr Val Asp Lys Pro Glu 180 185 190 Asp Ala Thr Ile Ile Glu Met Gln Glu Pro Val Ser Leu Thr Phe Ala 195 200 205 Leu Arg Tyr Met Asn Ser Phe Thr Lys Ala Ser Pro Leu Ser Glu Gln 210 215 220 Val Thr Ile Ser Leu Ser Ser Glu Leu Pro Val Val Val Glu Tyr Lys 225 230 235 240 Ile Ala Glu Met Gly Tyr Ile Arg Phe Tyr Leu Ala Pro Lys Ile Glu 245 250 255 Glu Asp Glu Glu Met Lys Ser 260 <210> 10 <211> 74 <212> PRT <213> Artificial Sequence <220> <223> truncated Ubiquitin (Ub) <400> 10 Met Gln Ile Phe Val Lys Thr Leu Thr Gly Lys Thr Ile Thr Leu Glu 1 5 10 15 Val Glu Ser Ser Asp Thr Ile Asp Asn Val Lys Ala Lys Ile Gln Asp 20 25 30 Lys Glu Gly Ile Pro Pro Asp Gln Gln Arg Leu Ile Phe Ala Gly Lys 35 40 45 Gln Leu Glu Asp Gly Arg Thr Leu Ala Asp Tyr Asn Ile Gln Lys Glu 50 55 60 Ser Thr Leu His Leu Val Leu Arg Leu Arg 65 70 <210> 11 <211> 379 <212> PRT <213> Artificial Sequence <220> <223> OsAPE (D238A / N240V), dAPE <400> 11 Met Lys Arg Phe Phe Gln Pro Val Pro Lys Asp Gly Ser Pro Ala Lys 1 5 10 15 Lys Arg Pro Ala Ala Ala Ala Ala Ala Ser Ala Ser Asp Ser Asp Ser 20 25 30 Leu Gly Gly Asp Ala Pro Ala Ala Ala Ala Cys Ala Val Gly Glu Gly 35 40 45 Asp Ser Pro Pro Ala Pro Arg Glu Glu Glu Pro Arg Arg Phe Val Thr 50 55 60 Trp Asn Ala Asn Ser Leu Leu Leu Arg Met Lys Ser Asp Trp Pro Ala 65 70 75 80 Phe Cys Gln Phe Val Ser Arg Val Asp Pro Asp Val Ile Cys Val Gln 85 90 95 Glu Val Arg Met Pro Ala Ala Gly Ser Lys Gly Ala Pro Lys Asn Pro 100 105 110 Gly Gln Leu Lys Asp Asp Thr Ser Ser Ser Arg Asp Glu Lys Gln Val 115 120 125 Val Leu Arg Ala Leu Ser Ser Pro Pro Phe Lys Asp Tyr Arg Val Trp 130 135 140 Trp Ser Leu Ser Asp Ser Lys Tyr Ala Gly Thr Ala Met Ile Ile Lys 145 150 155 160 Lys Lys Phe Glu Pro Lys Lys Val Ser Phe Asn Leu Asp Arg Thr Ser 165 170 175 Ser Lys His Glu Pro Asp Gly Arg Val Ile Ile Ala Glu Phe Glu Ser 180 185 190 Phe Leu Leu Leu Asn Thr Tyr Ala Pro Asn Asn Gly Trp Lys Glu Glu 195 200 205 Glu Asn Ser Phe Gln Arg Arg Arg Lys Trp Asp Lys Arg Met Leu Glu 210 215 220 Phe Val Gln Gln Val Asp Lys Pro Leu Ile Trp Cys Gly Ala Leu Val 225 230 235 240 Val Ser His Glu Glu Ile Asp Val Ser His Pro Asp Phe Phe Ser Ser 245 250 255 Ala Lys Leu Asn Gly Tyr Ile Pro Pro Asn Lys Glu Asp Cys Gly Gln 260 265 270 Pro Gly Phe Thr Leu Ser Glu Arg Arg Arg Phe Gly Asn Ile Leu Ser 275 280 285 Gln Gly Lys Leu Val Asp Ala Tyr Arg Tyr Leu His Lys Glu Lys Asp 290 295 300 Met Asp Cys Gly Phe Ser Trp Ser Gly His Pro Ile Gly Lys Tyr Arg 305 310 315 320 Gly Lys Arg Met Arg Ile Asp Tyr Phe Leu Val Ser Glu Lys Leu Lys 325 330 335 Asp Gln Ile Val Ser Cys Asp Ile His Gly Arg Gly Ile Glu Leu Glu 340 345 350 Gly Phe Tyr Gly Ser Asp His Cys Pro Val Ser Leu Glu Leu Ser Glu 355 360 365 Glu Val Glu Ala Pro Lys Pro Lys Ser Ser Asn 370 375 <210> 12 <211> 1842 <212> PRT <213> Artificial Sequence <220> <223> Exemplary polypeptide of the PCGBE-1 system <400> 12 Met Glu Ala Ser Pro Ala Ser Gly Pro Arg His Leu Met Asp Pro His 1 5 10 15 Ile Phe Thr Ser Asn Phe Asn Asn Gly Ile Gly Arg His Lys Thr Tyr 20 25 30 Leu Cys Tyr Glu Val Glu Arg Leu Asp Asn Gly Thr Ser Val Lys Met 35 40 45 Asp Gln His Arg Gly Phe Leu His Asn Gln Ala Lys Asn Leu Leu Cys 50 55 60 Gly Phe Tyr Gly Arg His Ala Glu Leu Arg Phe Leu Asp Leu Val Pro 65 70 75 80 Ser Leu Gln Leu Asp Pro Ala Gln Ile Tyr Arg Val Thr Trp Phe Ile 85 90 95 Ser Trp Ser Pro Cys Phe Ser Trp Gly Cys Ala Gly Glu Val Arg Ala 100 105 110 Phe Leu Gln Glu Asn Thr His Val Arg Leu Arg Ile Phe Ala Ala Arg 115 120 125 Ile Tyr Asp Tyr Asp Pro Leu Tyr Lys Glu Ala Leu Gln Met Leu Arg 130 135 140 Asp Ala Gly Ala Gln Val Ser Ile Met Thr Tyr Asp Glu Phe Lys His 145 150 155 160 Cys Trp Asp Thr Phe Val Asp His Gln Gly Cys Pro Phe Gln Pro Trp 165 170 175 Asp Gly Leu Asp Glu His Ser Gln Ala Leu Ser Gly Arg Leu Arg Ala 180 185 190 Ile Leu Gln Asn Gln Gly Asn Ser Gly Ser Glu Thr Pro Gly Thr Ser 195 200 205 Glu Ser Ala Thr Pro Glu Ser Arg Pro Asp Lys Lys Tyr Ser Ile Gly 210 215 220 Leu Ala Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu 225 230 235 240 Tyr Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg 245 250 255 His Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly 260 265 270 Glu Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr 275 280 285 Thr Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn 290 295 300 Glu Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser 305 310 315 320 Phe Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe Gly 325 330 335 Asn Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr 340 345 350 His Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg 355 360 365 Leu Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His Phe 370 375 380 Leu Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu 385 390 395 400 Phe Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro 405 410 415 Ile Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu 420 425 430 Ser Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu 435 440 445 Lys Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu 450 455 460 Thr Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu 465 470 475 480 Gln Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala 485 490 495 Gln Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu 500 505 510 Ser Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile 515 520 525 Thr Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His 530 535 540 His Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro 545 550 555 560 Glu Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala 565 570 575 Gly Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile 580 585 590 Lys Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys 595 600 605 Leu Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly 610 615 620 Ser Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu Arg 625 630 635 640 Arg Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile 645 650 655 Glu Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala 660 665 670 Arg Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr 675 680 685 Ile Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala 690 695 700 Gln Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn 705 710 715 720 Glu Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val 725 730 735 Tyr Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys 740 745 750 Pro Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu Leu 755 760 765 Phe Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr 770 775 780 Phe Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu 785 790 795 800 Asp Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile 805 810 815 Ile Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu 820 825 830 Glu Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile 835 840 845 Glu Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met 850 855 860 Lys Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg 865 870 875 880 Lys Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu 885 890 895 Asp Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu 900 905 910 Ile His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln 915 920 925 Val Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu Ala 930 935 940 Gly Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val 945 950 955 960 Asp Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile Val 965 970 975 Ile Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn 980 985 990 Ser Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly 995 1000 1005 Ser Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln 1010 1015 1020 Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met 1025 1030 1035 Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp 1040 1045 1050 Val Asp His Ile Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile 1055 1060 1065 Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser 1070 1075 1080 Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys Asn Tyr 1085 1090 1095 Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe 1100 1105 1110 Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 1115 1120 1125 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile 1130 1135 1140 Thr Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys 1145 1150 1155 Tyr Asp Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr 1160 1165 1170 Leu Lys Ser Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe 1175 1180 1185 Tyr Lys Val Arg Glu Ile Asn Asn Tyr His His Ala His Asp Ala 1190 1195 1200 Tyr Leu Asn Ala Val Val Gly Thr Ala Leu Ile Lys Lys Tyr Pro 1205 1210 1215 Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp Tyr Lys Val Tyr Asp 1220 1225 1230 Val Arg Lys Met Ile Ala Lys Ser Glu Gln Glu Ile Gly Lys Ala 1235 1240 1245 Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile Met Asn Phe Phe Lys 1250 1255 1260 Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile Arg Lys Arg Pro Leu 1265 1270 1275 Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile Val Trp Asp Lys Gly 1280 1285 1290 Arg Asp Phe Ala Thr Val Arg Lys Val Leu Ser Met Pro Gln Val 1295 1300 1305 Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly Gly Phe Ser Lys 1310 1315 1320 Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp Lys Leu Ile Ala Arg 1325 1330 1335 Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe Asp Ser Pro 1340 1345 1350 Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val Glu Lys Gly 1355 1360 1365 Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu Gly Ile Thr 1370 1375 1380 Ile Met Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile Asp Phe Leu 1385 1390 1395 Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys Asp Leu Ile Ile Lys 1400 1405 1410 Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu Asn Gly Arg Lys Arg 1415 1420 1425 Met Leu Ala Ser Ala Gly Glu Leu Gln Lys Gly Asn Glu Leu Ala 1430 1435 1440 Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr Leu Ala Ser His Tyr 1445 1450 1455 Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn Glu Gln Lys Gln Leu 1460 1465 1470 Phe Val Glu Gln His Lys His Tyr Leu Asp Glu Ile Ile Glu Gln 1475 1480 1485 Ile Ser Glu Phe Ser Lys Arg Val Ile Leu Ala Asp Ala Asn Leu 1490 1495 1500 Asp Lys Val Leu Ser Ala Tyr Asn Lys His Arg Asp Lys Pro Ile 1505 1510 1515 Arg Glu Gln Ala Glu Asn Ile Ile His Leu Phe Thr Leu Thr Asn 1520 1525 1530 Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe Asp Thr Thr Ile Asp 1535 1540 1545 Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val Leu Asp Ala Thr Leu 1550 1555 1560 Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr Arg Ile Asp Leu 1565 1570 1575 Ser Gln Leu Gly Gly Asp Lys Arg Pro Ala Ala Thr Lys Lys Ala 1580 1585 1590 Gly Gln Ala Lys Lys Lys Lys Gly Thr Asp Ser Gly Gly Ser Ala 1595 1600 1605 Asn Glu Leu Thr Trp His Asp Val Leu Ala Glu Glu Lys Gln Gln 1610 1615 1620 Pro Tyr Phe Leu Asn Thr Leu Gln Thr Val Ala Ser Glu Arg Gln 1625 1630 1635 Ser Gly Val Thr Ile Tyr Pro Pro Gln Lys Asp Val Phe Asn Ala 1640 1645 1650 Phe Arg Phe Thr Glu Leu Gly Asp Val Lys Val Val Ile Leu Gly 1655 1660 1665 Gln Asp Pro Tyr His Gly Pro Gly Gln Ala His Gly Leu Ala Phe 1670 1675 1680 Ser Val Arg Pro Gly Ile Ala Ile Pro Pro Ser Leu Leu Asn Met 1685 1690 1695 Tyr Lys Glu Leu Glu Asn Thr Ile Pro Gly Phe Thr Arg Pro Asn 1700 1705 1710 His Gly Tyr Leu Glu Ser Trp Ala Arg Gln Gly Val Leu Leu Leu 1715 1720 1725 Asn Thr Val Leu Thr Val Arg Ala Gly Gln Ala His Ser His Ala 1730 1735 1740 Ser Leu Gly Trp Glu Thr Phe Thr Asp Lys Val Ile Ser Leu Ile 1745 1750 1755 Asn Gln His Arg Glu Gly Val Val Phe Leu Leu Trp Gly Ser His 1760 1765 1770 Ala Gln Lys Lys Gly Ala Ile Ile Asp Lys Gln Arg His His Val 1775 1780 1785 Leu Lys Ala Pro His Pro Ser Pro Leu Ser Ala His Arg Gly Phe 1790 1795 1800 Phe Gly Cys Asn His Phe Val Leu Ala Asn Gln Trp Leu Glu Gln 1805 1810 1815 Arg Gly Glu Thr Pro Ile Asp Trp Met Pro Val Leu Pro Ala Glu 1820 1825 1830 Ser Glu Pro Lys Lys Lys Arg Lys Val 1835 1840 <210> 13 <211> 2358 <212> PRT <213> Artificial Sequence <220> <223> PCGBE system schematic polypeptide( Figure 11 b) <400> 13 Met Glu Ala Ser Pro Ala Ser Gly Pro Arg His Leu Met Asp Pro His 1 5 10 15 Ile Phe Thr Ser Asn Phe Asn Asn Gly Ile Gly Arg His Lys Thr Tyr 20 25 30 Leu Cys Tyr Glu Val Glu Arg Leu Asp Asn Gly Thr Ser Val Lys Met 35 40 45 Asp Gln His Arg Gly Phe Leu His Asn Gln Ala Lys Asn Leu Leu Cys 50 55 60 Gly Phe Tyr Gly Arg His Ala Glu Leu Arg Phe Leu Asp Leu Val Pro 65 70 75 80 Ser Leu Gln Leu Asp Pro Ala Gln Ile Tyr Arg Val Thr Trp Phe Ile 85 90 95 Ser Trp Ser Pro Cys Phe Ser Trp Gly Cys Ala Gly Glu Val Arg Ala 100 105 110 Phe Leu Gln Glu Asn Thr His Val Arg Leu Arg Ile Phe Ala Ala Arg 115 120 125 Ile Tyr Asp Tyr Asp Pro Leu Tyr Lys Glu Ala Leu Gln Met Leu Arg 130 135 140 Asp Ala Gly Ala Gln Val Ser Ile Met Thr Tyr Asp Glu Phe Lys His 145 150 155 160 Cys Trp Asp Thr Phe Val Asp His Gln Gly Cys Pro Phe Gln Pro Trp 165 170 175 Asp Gly Leu Asp Glu His Ser Gln Ala Leu Ser Gly Arg Leu Arg Ala 180 185 190 Ile Leu Gln Asn Gln Gly Asn Ser Gly Ser Glu Thr Pro Gly Thr Ser 195 200 205 Glu Ser Ala Thr Pro Glu Ser Leu Lys Asp Lys Lys Tyr Ser Ile Gly 210 215 220 Leu Ala Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu 225 230 235 240 Tyr Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg 245 250 255 His Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly 260 265 270 Glu Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr 275 280 285 Thr Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn 290 295 300 Glu Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser 305 310 315 320 Phe Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe Gly 325 330 335 Asn Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr 340 345 350 His Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg 355 360 365 Leu Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His Phe 370 375 380 Leu Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu 385 390 395 400 Phe Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro 405 410 415 Ile Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu 420 425 430 Ser Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu 435 440 445 Lys Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu 450 455 460 Thr Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu 465 470 475 480 Gln Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala 485 490 495 Gln Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu 500 505 510 Ser Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile 515 520 525 Thr Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His 530 535 540 His Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro 545 550 555 560 Glu Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala 565 570 575 Gly Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile 580 585 590 Lys Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys 595 600 605 Leu Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly 610 615 620 Ser Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu Arg 625 630 635 640 Arg Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile 645 650 655 Glu Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala 660 665 670 Arg Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr 675 680 685 Ile Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala 690 695 700 Gln Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn 705 710 715 720 Glu Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val 725 730 735 Tyr Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys 740 745 750 Pro Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu Leu 755 760 765 Phe Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr 770 775 780 Phe Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu 785 790 795 800 Asp Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile 805 810 815 Ile Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu 820 825 830 Glu Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile 835 840 845 Glu Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met 850 855 860 Lys Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg 865 870 875 880 Lys Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu 885 890 895 Asp Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu 900 905 910 Ile His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln 915 920 925 Val Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu Ala 930 935 940 Gly Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val 945 950 955 960 Asp Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile Val 965 970 975 Ile Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn 980 985 990 Ser Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly 995 1000 1005 Ser Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln 1010 1015 1020 Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met 1025 1030 1035 Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp 1040 1045 1050 Val Asp His Ile Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile 1055 1060 1065 Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser 1070 1075 1080 Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys Asn Tyr 1085 1090 1095 Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe 1100 1105 1110 Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 1115 1120 1125 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile 1130 1135 1140 Thr Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys 1145 1150 1155 Tyr Asp Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr 1160 1165 1170 Leu Lys Ser Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe 1175 1180 1185 Tyr Lys Val Arg Glu Ile Asn Asn Tyr His His Ala His Asp Ala 1190 1195 1200 Tyr Leu Asn Ala Val Val Gly Thr Ala Leu Ile Lys Lys Tyr Pro 1205 1210 1215 Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp Tyr Lys Val Tyr Asp 1220 1225 1230 Val Arg Lys Met Ile Ala Lys Ser Glu Gln Glu Ile Gly Lys Ala 1235 1240 1245 Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile Met Asn Phe Phe Lys 1250 1255 1260 Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile Arg Lys Arg Pro Leu 1265 1270 1275 Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile Val Trp Asp Lys Gly 1280 1285 1290 Arg Asp Phe Ala Thr Val Arg Lys Val Leu Ser Met Pro Gln Val 1295 1300 1305 Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly Gly Phe Ser Lys 1310 1315 1320 Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp Lys Leu Ile Ala Arg 1325 1330 1335 Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe Asp Ser Pro 1340 1345 1350 Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val Glu Lys Gly 1355 1360 1365 Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu Gly Ile Thr 1370 1375 1380 Ile Met Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile Asp Phe Leu 1385 1390 1395 Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys Asp Leu Ile Ile Lys 1400 1405 1410 Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu Asn Gly Arg Lys Arg 1415 1420 1425 Met Leu Ala Ser Ala Gly Glu Leu Gln Lys Gly Asn Glu Leu Ala 1430 1435 1440 Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr Leu Ala Ser His Tyr 1445 1450 1455 Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn Glu Gln Lys Gln Leu 1460 1465 1470 Phe Val Glu Gln His Lys His Tyr Leu Asp Glu Ile Ile Glu Gln 1475 1480 1485 Ile Ser Glu Phe Ser Lys Arg Val Ile Leu Ala Asp Ala Asn Leu 1490 1495 1500 Asp Lys Val Leu Ser Ala Tyr Asn Lys His Arg Asp Lys Pro Ile 1505 1510 1515 Arg Glu Gln Ala Glu Asn Ile Ile His Leu Phe Thr Leu Thr Asn 1520 1525 1530 Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe Asp Thr Thr Ile Asp 1535 1540 1545 Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val Leu Asp Ala Thr Leu 1550 1555 1560 Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr Arg Ile Asp Leu 1565 1570 1575 Ser Gln Leu Gly Gly Asp Lys Arg Pro Ala Ala Thr Lys Lys Ala 1580 1585 1590 Gly Gln Ala Lys Lys Lys Lys Thr Arg Asp Ser Gly Gly Ser Ala 1595 1600 1605 Asn Glu Leu Thr Trp His Asp Val Leu Ala Glu Glu Lys Gln Gln 1610 1615 1620 Pro Tyr Phe Leu Asn Thr Leu Gln Thr Val Ala Ser Glu Arg Gln 1625 1630 1635 Ser Gly Val Thr Ile Tyr Pro Pro Gln Lys Asp Val Phe Asn Ala 1640 1645 1650 Phe Arg Phe Thr Glu Leu Gly Asp Val Lys Val Val Ile Leu Gly 1655 1660 1665 Gln Asp Pro Tyr His Gly Pro Gly Gln Ala His Gly Leu Ala Phe 1670 1675 1680 Ser Val Arg Pro Gly Ile Ala Ile Pro Pro Ser Leu Leu Asn Met 1685 1690 1695 Tyr Lys Glu Leu Glu Asn Thr Ile Pro Gly Phe Thr Arg Pro Asn 1700 1705 1710 His Gly Tyr Leu Glu Ser Trp Ala Arg Gln Gly Val Leu Leu Leu 1715 1720 1725 Asn Thr Val Leu Thr Val Arg Ala Gly Gln Ala His Ser His Ala 1730 1735 1740 Ser Leu Gly Trp Glu Thr Phe Thr Asp Lys Val Ile Ser Leu Ile 1745 1750 1755 Asn Gln His Arg Glu Gly Val Val Phe Leu Leu Trp Gly Ser His 1760 1765 1770 Ala Gln Lys Lys Gly Ala Ile Ile Asp Lys Gln Arg His His Val 1775 1780 1785 Leu Lys Ala Pro His Pro Ser Pro Leu Ser Ala His Arg Gly Phe 1790 1795 1800 Phe Gly Cys Asn His Phe Val Leu Ala Asn Gln Trp Leu Glu Gln 1805 1810 1815 Arg Gly Glu Thr Pro Ile Asp Trp Met Pro Val Leu Pro Ala Glu 1820 1825 1830 Ser Glu Pro Lys Lys Lys Arg Lys Val Ser Ala Glu Gly Arg Gly 1835 1840 1845 Ser Leu Leu Thr Cys Gly Asp Val Glu Glu Asn Pro Gly Pro Ala 1850 1855 1860 Ser Asn Phe Thr Gln Phe Val Leu Val Asp Asn Gly Gly Thr Gly 1865 1870 1875 Asp Val Thr Val Ala Pro Ser Asn Phe Ala Asn Gly Ile Ala Glu 1880 1885 1890 Trp Ile Ser Ser Asn Ser Arg Ser Gln Ala Tyr Lys Val Thr Cys 1895 1900 1905 Ser Val Arg Gln Ser Ser Ala Gln Asn Arg Lys Tyr Thr Ile Lys 1910 1915 1920 Val Glu Val Pro Lys Gly Ala Trp Arg Ser Tyr Leu Asn Met Glu 1925 1930 1935 Leu Thr Ile Pro Ile Phe Ala Thr Asn Ser Asp Cys Glu Leu Ile 1940 1945 1950 Val Lys Ala Met Gln Gly Leu Leu Lys Asp Gly Asn Pro Ile Pro 1955 1960 1965 Ser Ala Ile Ala Ala Asn Ser Gly Ile Tyr Gly Gly Ser Pro Lys 1970 1975 1980 Lys Lys Arg Lys Val Ser Gly Ser Glu Thr Pro Gly Thr Ser Glu 1985 1990 1995 Ser Ala Thr Pro Glu Ser Pro Arg Met Leu Glu Leu Arg Leu Val 2000 2005 2010 Gln Gly Ser Leu Leu Lys Lys Val Leu Glu Ala Ile Arg Glu Leu 2015 2020 2025 Val Thr Asp Ala Asn Phe Asp Cys Ser Gly Thr Gly Phe Ser Leu 2030 2035 2040 Gln Ala Met Asp Ser Ser His Val Ala Leu Val Ala Leu Leu Leu 2045 2050 2055 Arg Ser Glu Gly Phe Glu His Tyr Arg Cys Asp Arg Asn Leu Ser 2060 2065 2070 Met Gly Met Asn Leu Asn Asn Met Ala Lys Met Leu Arg Cys Ala 2075 2080 2085 Gly Asn Asp Asp Ile Ile Thr Ile Lys Ala Asp Asp Gly Ser Asp 2090 2095 2100 Thr Val Thr Phe Met Phe Glu Ser Pro Asn Gln Asp Lys Ile Ala 2105 2110 2115 Asp Phe Glu Met Lys Leu Met Asp Ile Asp Ser Glu His Leu Gly 2120 2125 2130 Ile Pro Asp Ser Glu Tyr Gln Ala Ile Val Arg Met Pro Ser Ser 2135 2140 2145 Glu Phe Ser Arg Ile Cys Lys Asp Leu Ser Ser Ile Gly Asp Thr 2150 2155 2160 Val Ile Ile Ser Val Thr Arg Glu Gly Val Lys Phe Ser Thr Ala 2165 2170 2175 Gly Asp Ile Gly Thr Ala Asn Ile Val Cys Arg Gln Asn Lys Thr 2180 2185 2190 Val Asp Lys Pro Glu Asp Ala Thr Ile Ile Glu Met Gln Glu Pro 2195 2200 2205 Val Ser Leu Thr Phe Ala Leu Arg Tyr Met Asn Ser Phe Thr Lys 2210 2215 2220 Ala Ser Pro Leu Ser Glu Gln Val Thr Ile Ser Leu Ser Ser Glu 2225 2230 2235 Leu Pro Val Val Val Glu Tyr Lys Ile Ala Glu Met Gly Tyr Ile 2240 2245 2250 Arg Phe Tyr Leu Ala Pro Lys Ile Glu Glu Asp Glu Glu Met Lys 2255 2260 2265 Ser Ser Gly Gly Ser Met Gln Ile Phe Val Lys Thr Leu Thr Gly 2270 2275 2280 Lys Thr Ile Thr Leu Glu Val Glu Ser Ser Asp Thr Ile Asp Asn 2285 2290 2295 Val Lys Ala Lys Ile Gln Asp Lys Glu Gly Ile Pro Pro Asp Gln 2300 2305 2310 Gln Arg Leu Ile Phe Ala Gly Lys Gln Leu Glu Asp Gly Arg Thr 2315 2320 2325 Leu Ala Asp Tyr Asn Ile Gln Lys Glu Ser Thr Leu His Leu Val 2330 2335 2340 Leu Arg Leu Arg Ser Gly Gly Ser Pro Lys Lys Lys Arg Lys Val 2345 2350 2355 <210> 14 <211> 2248 <212> PRT <213> Artificial Sequence <220> <223> PCGBE system schematic polypeptide( Figure 11 c) <400> 14 Met Glu Ala Ser Pro Ala Ser Gly Pro Arg His Leu Met Asp Pro His 1 5 10 15 Ile Phe Thr Ser Asn Phe Asn Asn Gly Ile Gly Arg His Lys Thr Tyr 20 25 30 Leu Cys Tyr Glu Val Glu Arg Leu Asp Asn Gly Thr Ser Val Lys Met 35 40 45 Asp Gln His Arg Gly Phe Leu His Asn Gln Ala Lys Asn Leu Leu Cys 50 55 60 Gly Phe Tyr Gly Arg His Ala Glu Leu Arg Phe Leu Asp Leu Val Pro 65 70 75 80 Ser Leu Gln Leu Asp Pro Ala Gln Ile Tyr Arg Val Thr Trp Phe Ile 85 90 95 Ser Trp Ser Pro Cys Phe Ser Trp Gly Cys Ala Gly Glu Val Arg Ala 100 105 110 Phe Leu Gln Glu Asn Thr His Val Arg Leu Arg Ile Phe Ala Ala Arg 115 120 125 Ile Tyr Asp Tyr Asp Pro Leu Tyr Lys Glu Ala Leu Gln Met Leu Arg 130 135 140 Asp Ala Gly Ala Gln Val Ser Ile Met Thr Tyr Asp Glu Phe Lys His 145 150 155 160 Cys Trp Asp Thr Phe Val Asp His Gln Gly Cys Pro Phe Gln Pro Trp 165 170 175 Asp Gly Leu Asp Glu His Ser Gln Ala Leu Ser Gly Arg Leu Arg Ala 180 185 190 Ile Leu Gln Asn Gln Gly Asn Ser Gly Ser Glu Thr Pro Gly Thr Ser 195 200 205 Glu Ser Ala Thr Pro Glu Ser Leu Lys Asp Lys Lys Tyr Ser Ile Gly 210 215 220 Leu Ala Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu 225 230 235 240 Tyr Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg 245 250 255 His Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly 260 265 270 Glu Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr 275 280 285 Thr Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn 290 295 300 Glu Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser 305 310 315 320 Phe Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe Gly 325 330 335 Asn Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr 340 345 350 His Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg 355 360 365 Leu Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His Phe 370 375 380 Leu Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu 385 390 395 400 Phe Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro 405 410 415 Ile Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu 420 425 430 Ser Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu 435 440 445 Lys Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu 450 455 460 Thr Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu 465 470 475 480 Gln Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala 485 490 495 Gln Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu 500 505 510 Ser Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile 515 520 525 Thr Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His 530 535 540 His Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro 545 550 555 560 Glu Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala 565 570 575 Gly Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile 580 585 590 Lys Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys 595 600 605 Leu Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly 610 615 620 Ser Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu Arg 625 630 635 640 Arg Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile 645 650 655 Glu Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala 660 665 670 Arg Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr 675 680 685 Ile Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala 690 695 700 Gln Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn 705 710 715 720 Glu Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val 725 730 735 Tyr Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys 740 745 750 Pro Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu Leu 755 760 765 Phe Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr 770 775 780 Phe Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu 785 790 795 800 Asp Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile 805 810 815 Ile Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu 820 825 830 Glu Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile 835 840 845 Glu Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met 850 855 860 Lys Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg 865 870 875 880 Lys Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu 885 890 895 Asp Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu 900 905 910 Ile His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln 915 920 925 Val Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu Ala 930 935 940 Gly Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val 945 950 955 960 Asp Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile Val 965 970 975 Ile Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn 980 985 990 Ser Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly 995 1000 1005 Ser Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln 1010 1015 1020 Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met 1025 1030 1035 Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp 1040 1045 1050 Val Asp His Ile Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile 1055 1060 1065 Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser 1070 1075 1080 Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys Asn Tyr 1085 1090 1095 Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe 1100 1105 1110 Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 1115 1120 1125 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile 1130 1135 1140 Thr Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys 1145 1150 1155 Tyr Asp Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr 1160 1165 1170 Leu Lys Ser Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe 1175 1180 1185 Tyr Lys Val Arg Glu Ile Asn Asn Tyr His His Ala His Asp Ala 1190 1195 1200 Tyr Leu Asn Ala Val Val Gly Thr Ala Leu Ile Lys Lys Tyr Pro 1205 1210 1215 Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp Tyr Lys Val Tyr Asp 1220 1225 1230 Val Arg Lys Met Ile Ala Lys Ser Glu Gln Glu Ile Gly Lys Ala 1235 1240 1245 Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile Met Asn Phe Phe Lys 1250 1255 1260 Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile Arg Lys Arg Pro Leu 1265 1270 1275 Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile Val Trp Asp Lys Gly 1280 1285 1290 Arg Asp Phe Ala Thr Val Arg Lys Val Leu Ser Met Pro Gln Val 1295 1300 1305 Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly Gly Phe Ser Lys 1310 1315 1320 Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp Lys Leu Ile Ala Arg 1325 1330 1335 Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe Asp Ser Pro 1340 1345 1350 Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val Glu Lys Gly 1355 1360 1365 Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu Gly Ile Thr 1370 1375 1380 Ile Met Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile Asp Phe Leu 1385 1390 1395 Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys Asp Leu Ile Ile Lys 1400 1405 1410 Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu Asn Gly Arg Lys Arg 1415 1420 1425 Met Leu Ala Ser Ala Gly Glu Leu Gln Lys Gly Asn Glu Leu Ala 1430 1435 1440 Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr Leu Ala Ser His Tyr 1445 1450 1455 Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn Glu Gln Lys Gln Leu 1460 1465 1470 Phe Val Glu Gln His Lys His Tyr Leu Asp Glu Ile Ile Glu Gln 1475 1480 1485 Ile Ser Glu Phe Ser Lys Arg Val Ile Leu Ala Asp Ala Asn Leu 1490 1495 1500 Asp Lys Val Leu Ser Ala Tyr Asn Lys His Arg Asp Lys Pro Ile 1505 1510 1515 Arg Glu Gln Ala Glu Asn Ile Ile His Leu Phe Thr Leu Thr Asn 1520 1525 1530 Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe Asp Thr Thr Ile Asp 1535 1540 1545 Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val Leu Asp Ala Thr Leu 1550 1555 1560 Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr Arg Ile Asp Leu 1565 1570 1575 Ser Gln Leu Gly Gly Asp Lys Arg Pro Ala Ala Thr Lys Lys Ala 1580 1585 1590 Gly Gln Ala Lys Lys Lys Lys Thr Arg Asp Ser Gly Gly Ser Ala 1595 1600 1605 Asn Glu Leu Thr Trp His Asp Val Leu Ala Glu Glu Lys Gln Gln 1610 1615 1620 Pro Tyr Phe Leu Asn Thr Leu Gln Thr Val Ala Ser Glu Arg Gln 1625 1630 1635 Ser Gly Val Thr Ile Tyr Pro Pro Gln Lys Asp Val Phe Asn Ala 1640 1645 1650 Phe Arg Phe Thr Glu Leu Gly Asp Val Lys Val Val Ile Leu Gly 1655 1660 1665 Gln Asp Pro Tyr His Gly Pro Gly Gln Ala His Gly Leu Ala Phe 1670 1675 1680 Ser Val Arg Pro Gly Ile Ala Ile Pro Pro Ser Leu Leu Asn Met 1685 1690 1695 Tyr Lys Glu Leu Glu Asn Thr Ile Pro Gly Phe Thr Arg Pro Asn 1700 1705 1710 His Gly Tyr Leu Glu Ser Trp Ala Arg Gln Gly Val Leu Leu Leu 1715 1720 1725 Asn Thr Val Leu Thr Val Arg Ala Gly Gln Ala His Ser His Ala 1730 1735 1740 Ser Leu Gly Trp Glu Thr Phe Thr Asp Lys Val Ile Ser Leu Ile 1745 1750 1755 Asn Gln His Arg Glu Gly Val Val Phe Leu Leu Trp Gly Ser His 1760 1765 1770 Ala Gln Lys Lys Gly Ala Ile Ile Asp Lys Gln Arg His His Val 1775 1780 1785 Leu Lys Ala Pro His Pro Ser Pro Leu Ser Ala His Arg Gly Phe 1790 1795 1800 Phe Gly Cys Asn His Phe Val Leu Ala Asn Gln Trp Leu Glu Gln 1805 1810 1815 Arg Gly Glu Thr Pro Ile Asp Trp Met Pro Val Leu Pro Ala Glu 1820 1825 1830 Ser Glu Pro Lys Lys Lys Arg Lys Val Ser Ala Glu Gly Arg Gly 1835 1840 1845 Ser Leu Leu Thr Cys Gly Asp Val Glu Glu Asn Pro Gly Pro Met 1850 1855 1860 Lys Arg Phe Phe Gln Pro Val Pro Lys Asp Gly Ser Pro Ala Lys 1865 1870 1875 Lys Arg Pro Ala Ala Ala Ala Ala Ala Ser Ala Ser Asp Ser Asp 1880 1885 1890 Ser Leu Gly Gly Asp Ala Pro Ala Ala Ala Ala Cys Ala Val Gly 1895 1900 1905 Glu Gly Asp Ser Pro Pro Ala Pro Arg Glu Glu Glu Pro Arg Arg 1910 1915 1920 Phe Val Thr Trp Asn Ala Asn Ser Leu Leu Leu Arg Met Lys Ser 1925 1930 1935 Asp Trp Pro Ala Phe Cys Gln Phe Val Ser Arg Val Asp Pro Asp 1940 1945 1950 Val Ile Cys Val Gln Glu Val Arg Met Pro Ala Ala Gly Ser Lys 1955 1960 1965 Gly Ala Pro Lys Asn Pro Gly Gln Leu Lys Asp Asp Thr Ser Ser 1970 1975 1980 Ser Arg Asp Glu Lys Gln Val Val Leu Arg Ala Leu Ser Ser Pro 1985 1990 1995 Pro Phe Lys Asp Tyr Arg Val Trp Trp Ser Leu Ser Asp Ser Lys 2000 2005 2010 Tyr Ala Gly Thr Ala Met Ile Ile Lys Lys Lys Phe Glu Pro Lys 2015 2020 2025 Lys Val Ser Phe Asn Leu Asp Arg Thr Ser Ser Lys His Glu Pro 2030 2035 2040 Asp Gly Arg Val Ile Ile Ala Glu Phe Glu Ser Phe Leu Leu Leu 2045 2050 2055 Asn Thr Tyr Ala Pro Asn Asn Gly Trp Lys Glu Glu Glu Asn Ser 2060 2065 2070 Phe Gln Arg Arg Arg Lys Trp Asp Lys Arg Met Leu Glu Phe Val 2075 2080 2085 Gln Gln Val Asp Lys Pro Leu Ile Trp Cys Gly Ala Leu Val Val 2090 2095 2100 Ser His Glu Glu Ile Asp Val Ser His Pro Asp Phe Phe Ser Ser 2105 2110 2115 Ala Lys Leu Asn Gly Tyr Ile Pro Pro Asn Lys Glu Asp Cys Gly 2120 2125 2130 Gln Pro Gly Phe Thr Leu Ser Glu Arg Arg Arg Phe Gly Asn Ile 2135 2140 2145 Leu Ser Gln Gly Lys Leu Val Asp Ala Tyr Arg Tyr Leu His Lys 2150 2155 2160 Glu Lys Asp Met Asp Cys Gly Phe Ser Trp Ser Gly His Pro Ile 2165 2170 2175 Gly Lys Tyr Arg Gly Lys Arg Met Arg Ile Asp Tyr Phe Leu Val 2180 2185 2190 Ser Glu Lys Leu Lys Asp Gln Ile Val Ser Cys Asp Ile His Gly 2195 2200 2205 Arg Gly Ile Glu Leu Glu Gly Phe Tyr Gly Ser Asp His Cys Pro 2210 2215 2220 Val Ser Leu Glu Leu Ser Glu Glu Val Glu Ala Pro Lys Pro Lys 2225 2230 2235 Ser Ser Asn Pro Lys Lys Lys Arg Lys Val 2240 2245 <210> 15 <211> 331 <212> PRT <213> Artificial Sequence <220> <223> mAPE01g (D297A) <400> 15 Ser Ala Ile Arg Ala Ser Ser His Arg Leu Gln Thr Arg Thr Val Ala 1 5 10 15 Leu Thr Arg Thr Lys Met Ser Ser Met Ala Gly Leu Gly Ala Ser Gln 20 25 30 His Gly Tyr Pro Pro Arg Ser His Glu Pro Trp Thr Lys Leu Val His 35 40 45 Arg Glu Arg Leu Pro Glu Trp Phe Ala Tyr Asn Pro Lys Thr Met Arg 50 55 60 Pro Pro Pro Leu Ser His Asp Thr Lys Cys Met Lys Ile Leu Ser Trp 65 70 75 80 Asn Ile Asn Gly Leu His Asp Val Val Thr Thr Lys Gly Phe Ser Ala 85 90 95 Arg Asp Leu Ala Gln Arg Glu Asn Phe Asp Val Leu Cys Leu Gln Glu 100 105 110 Thr His Leu Glu Glu Lys Asp Val Glu Lys Phe Lys Asn Leu Ile Ala 115 120 125 Asp Tyr Asp Ser Tyr Trp Ser Cys Ser Val Ser Arg Leu Gly Tyr Ser 130 135 140 Gly Thr Ala Val Ile Ser Arg Val Lys Pro Ile Ser Val Gln Tyr Gly 145 150 155 160 Ile Gly Ile Arg Glu His Asp His Glu Gly Arg Val Ile Thr Leu Glu 165 170 175 Phe Asp Gly Phe Tyr Leu Val Asn Ala Tyr Val Pro Asn Ser Gly Arg 180 185 190 Phe Leu Arg Arg Leu Asn Tyr Arg Val Asn Asn Trp Asp Pro Cys Phe 195 200 205 Ser Asn Tyr Val Lys Ile Leu Glu Lys Ser Lys Pro Val Ile Val Ala 210 215 220 Gly Asp Leu Asn Cys Ala Arg Gln Ser Ile Asp Ile His Asn Pro Pro 225 230 235 240 Ala Lys Thr Lys Ser Ala Gly Phe Thr Ile Glu Glu Arg Glu Ser Phe 245 250 255 Glu Thr Asn Phe Ser Ser Lys Gly Leu Val Asp Thr Phe Arg Lys Gln 260 265 270 His Pro Asn Ala Val Gly Tyr Thr Phe Trp Gly Glu Asn Gln Arg Ile 275 280 285 Thr Asn Lys Gly Trp Arg Leu Ala Tyr Phe Leu Ala Ser Glu Ser Ile 290 295 300 Thr Asp Lys Val His Asp Ser Tyr Ile Leu Pro Asp Val Ser Phe Ser 305 310 315 320 Asp His Ser Pro Ile Gly Leu Val Leu Lys Leu 325 330 <210> 16 <211> 378 <212> PRT <213> Artificial Sequence <220> <223> mAPE12g (D327A) <400> 16 Lys Arg Phe Phe Gln Pro Val Pro Lys Asp Gly Ser Pro Ala Lys Lys 1 5 10 15 Arg Pro Ala Ala Ala Ala Ala Ala Ser Ala Ser Asp Ser Asp Ser Leu 20 25 30 Gly Gly Asp Ala Pro Ala Ala Ala Ala Cys Ala Val Gly Glu Gly Asp 35 40 45 Ser Pro Pro Ala Pro Arg Glu Glu Glu Pro Arg Arg Phe Val Thr Trp 50 55 60 Asn Ala Asn Ser Leu Leu Leu Arg Met Lys Ser Asp Trp Pro Ala Phe 65 70 75 80 Cys Gln Phe Val Ser Arg Val Asp Pro Asp Val Ile Cys Val Gln Glu 85 90 95 Val Arg Met Pro Ala Ala Gly Ser Lys Gly Ala Pro Lys Asn Pro Gly 100 105 110 Gln Leu Lys Asp Asp Thr Ser Ser Ser Arg Asp Glu Lys Gln Val Val 115 120 125 Leu Arg Ala Leu Ser Ser Pro Pro Phe Lys Asp Tyr Arg Val Trp Trp 130 135 140 Ser Leu Ser Asp Ser Lys Tyr Ala Gly Thr Ala Met Ile Ile Lys Lys 145 150 155 160 Lys Phe Glu Pro Lys Lys Val Ser Phe Asn Leu Asp Arg Thr Ser Ser 165 170 175 Lys His Glu Pro Asp Gly Arg Val Ile Ile Ala Glu Phe Glu Ser Phe 180 185 190 Leu Leu Leu Asn Thr Tyr Ala Pro Asn Asn Gly Trp Lys Glu Glu Glu 195 200 205 Asn Ser Phe Gln Arg Arg Arg Lys Trp Asp Lys Arg Met Leu Glu Phe 210 215 220 Val Gln Gln Val Asp Lys Pro Leu Ile Trp Cys Gly Asp Leu Asn Val 225 230 235 240 Ser His Glu Glu Ile Asp Val Ser His Pro Asp Phe Phe Ser Ser Ala 245 250 255 Lys Leu Asn Gly Tyr Ile Pro Pro Asn Lys Glu Asp Cys Gly Gln Pro 260 265 270 Gly Phe Thr Leu Ser Glu Arg Arg Arg Phe Gly Asn Ile Leu Ser Gln 275 280 285 Gly Lys Leu Val Asp Ala Tyr Arg Tyr Leu His Lys Glu Lys Asp Met 290 295 300 Asp Cys Gly Phe Ser Trp Ser Gly His Pro Ile Gly Lys Tyr Arg Gly 305 310 315 320 Lys Arg Met Arg Ile Ala Tyr Phe Leu Val Ser Glu Lys Leu Lys Asp 325 330 335 Gln Ile Val Ser Cys Asp Ile His Gly Arg Gly Ile Glu Leu Glu Gly 340 345 350 Phe Tyr Gly Ser Asp His Cys Pro Val Ser Leu Glu Leu Ser Glu Glu 355 360 365 Val Glu Ala Pro Lys Pro Lys Ser Ser Asn 370 375 <210> 17 <211> 494 <212> PRT <213> Artificial Sequence <220> <223> hRad18 <400> 17 Asp Ser Leu Ala Glu Ser Arg Trp Pro Pro Gly Leu Ala Val Met Lys 1 5 10 15 Thr Ile Asp Asp Leu Leu Arg Cys Gly Ile Cys Phe Glu Tyr Phe Asn 20 25 30 Ile Ala Met Ile Ile Pro Gln Cys Ser His Asn Tyr Cys Ser Leu Cys 35 40 45 Ile Arg Lys Phe Leu Ser Tyr Lys Thr Gln Cys Pro Thr Cys Cys Val 50 55 60 Thr Val Thr Glu Pro Asp Leu Lys Asn Asn Arg Ile Leu Asp Glu Leu 65 70 75 80 Val Lys Ser Leu Asn Phe Ala Arg Asn His Leu Leu Gln Phe Ala Leu 85 90 95 Glu Ser Pro Ala Lys Ser Pro Ala Ser Ser Ser Ser Lys Asn Leu Ala 100 105 110 Val Lys Val Tyr Thr Pro Val Ala Ser Arg Gln Ser Leu Lys Gln Gly 115 120 125 Ser Arg Leu Met Asp Asn Phe Leu Ile Arg Glu Met Ser Gly Ser Thr 130 135 140 Ser Glu Leu Leu Ile Lys Glu Asn Lys Ser Lys Phe Ser Pro Gln Lys 145 150 155 160 Glu Ala Ser Pro Ala Ala Lys Thr Lys Glu Thr Arg Ser Val Glu Glu 165 170 175 Ile Ala Pro Asp Pro Ser Glu Ala Lys Arg Pro Glu Pro Pro Ser Thr 180 185 190 Ser Thr Leu Lys Gln Val Thr Lys Val Asp Cys Pro Val Cys Gly Val 195 200 205 Asn Ile Pro Glu Ser His Ile Asn Lys His Leu Asp Ser Cys Leu Ser 210 215 220 Arg Glu Glu Lys Lys Glu Ser Leu Arg Ser Ser Val His Lys Arg Lys 225 230 235 240 Pro Leu Pro Lys Thr Val Tyr Asn Leu Leu Ser Asp Arg Asp Leu Lys 245 250 255 Lys Lys Leu Lys Glu His Gly Leu Ser Ile Gln Gly Asn Lys Gln Gln 260 265 270 Leu Ile Lys Arg His Gln Glu Phe Val His Met Tyr Asn Ala Gln Cys 275 280 285 Asp Ala Leu His Pro Lys Ser Ala Ala Glu Ile Val Arg Glu Ile Glu 290 295 300 Asn Ile Glu Lys Thr Arg Met Arg Leu Glu Ala Ser Lys Leu Asn Glu 305 310 315 320 Ser Val Met Val Phe Thr Lys Asp Gln Thr Glu Lys Glu Ile Asp Glu 325 330 335 Ile His Ser Lys Tyr Arg Lys Lys His Lys Ser Glu Phe Gln Leu Leu 340 345 350 Val Asp Gln Ala Arg Lys Gly Tyr Lys Lys Ile Ala Gly Met Ser Gln 355 360 365 Lys Thr Val Thr Ile Thr Lys Glu Asp Glu Ser Thr Glu Lys Leu Ser 370 375 380 Ser Val Cys Met Gly Gln Glu Asp Asn Met Thr Ser Val Thr Asn His 385 390 395 400 Phe Ser Gln Ser Lys Leu Asp Ser Pro Glu Glu Leu Glu Pro Asp Arg 405 410 415 Glu Glu Asp Ser Ser Ser Cys Ile Asp Ile Gln Glu Val Leu Ser Ser 420 425 430 Ser Glu Ser Asp Ser Cys Asn Ser Ser Ser Ser Asp Ile Ile Arg Asp 435 440 445 Leu Leu Glu Glu Glu Glu Ala Trp Glu Ala Ser His Lys Asn Asp Leu 450 455 460 Gln Asp Thr Glu Ile Ser Pro Arg Gln Asn Arg Arg Thr Arg Ala Ala 465 470 475 480 Glu Ser Ala Glu Ile Glu Pro Arg Asn Lys Arg Asn Arg Asn 485 490 <210> 18 <211> 332 <212> PRT <213> Artificial Sequence <220> <223> APE01g <400> 18 Met Ser Ala Ile Arg Ala Ser Ser His Arg Leu Gln Thr Arg Thr Val 1 5 10 15 Ala Leu Thr Arg Thr Lys Met Ser Ser Met Ala Gly Leu Gly Ala Ser 20 25 30 Gln His Gly Tyr Pro Pro Arg Ser His Glu Pro Trp Thr Lys Leu Val 35 40 45 His Arg Glu Arg Leu Pro Glu Trp Phe Ala Tyr Asn Pro Lys Thr Met 50 55 60 Arg Pro Pro Pro Leu Ser His Asp Thr Lys Cys Met Lys Ile Leu Ser 65 70 75 80 Trp Asn Ile Asn Gly Leu His Asp Val Val Thr Thr Lys Gly Phe Ser 85 90 95 Ala Arg Asp Leu Ala Gln Arg Glu Asn Phe Asp Val Leu Cys Leu Gln 100 105 110 Glu Thr His Leu Glu Glu Lys Asp Val Glu Lys Phe Lys Asn Leu Ile 115 120 125 Ala Asp Tyr Asp Ser Tyr Trp Ser Cys Ser Val Ser Arg Leu Gly Tyr 130 135 140 Ser Gly Thr Ala Val Ile Ser Arg Val Lys Pro Ile Ser Val Gln Tyr 145 150 155 160 Gly Ile Gly Ile Arg Glu His Asp His Glu Gly Arg Val Ile Thr Leu 165 170 175 Glu Phe Asp Gly Phe Tyr Leu Val Asn Ala Tyr Val Pro Asn Ser Gly 180 185 190 Arg Phe Leu Arg Arg Leu Asn Tyr Arg Val Asn Asn Trp Asp Pro Cys 195 200 205 Phe Ser Asn Tyr Val Lys Ile Leu Glu Lys Ser Lys Pro Val Ile Val 210 215 220 Ala Gly Asp Leu Asn Cys Ala Arg Gln Ser Ile Asp Ile His Asn Pro 225 230 235 240 Pro Ala Lys Thr Lys Ser Ala Gly Phe Thr Ile Glu Glu Arg Glu Ser 245 250 255 Phe Glu Thr Asn Phe Ser Ser Lys Gly Leu Val Asp Thr Phe Arg Lys 260 265 270 Gln His Pro Asn Ala Val Gly Tyr Thr Phe Trp Gly Glu Asn Gln Arg 275 280 285 Ile Thr Asn Lys Gly Trp Arg Leu Asp Tyr Phe Leu Ala Ser Glu Ser 290 295 300 Ile Thr Asp Lys Val His Asp Ser Tyr Ile Leu Pro Asp Val Ser Phe 305 310 315 320 Ser Asp His Ser Pro Ile Gly Leu Val Leu Lys Leu 325 330 <210> 19 <211> 379 <212> PRT <213> Artificial Sequence <220> <223> APE12g <400> 19 Met Lys Arg Phe Phe Gln Pro Val Pro Lys Asp Gly Ser Pro Ala Lys 1 5 10 15 Lys Arg Pro Ala Ala Ala Ala Ala Ala Ser Ala Ser Asp Ser Asp Ser 20 25 30 Leu Gly Gly Asp Ala Pro Ala Ala Ala Ala Cys Ala Val Gly Glu Gly 35 40 45 Asp Ser Pro Pro Ala Pro Arg Glu Glu Glu Pro Arg Arg Phe Val Thr 50 55 60 Trp Asn Ala Asn Ser Leu Leu Leu Arg Met Lys Ser Asp Trp Pro Ala 65 70 75 80 Phe Cys Gln Phe Val Ser Arg Val Asp Pro Asp Val Ile Cys Val Gln 85 90 95 Glu Val Arg Met Pro Ala Ala Gly Ser Lys Gly Ala Pro Lys Asn Pro 100 105 110 Gly Gln Leu Lys Asp Asp Thr Ser Ser Ser Arg Asp Glu Lys Gln Val 115 120 125 Val Leu Arg Ala Leu Ser Ser Pro Pro Phe Lys Asp Tyr Arg Val Trp 130 135 140 Trp Ser Leu Ser Asp Ser Lys Tyr Ala Gly Thr Ala Met Ile Ile Lys 145 150 155 160 Lys Lys Phe Glu Pro Lys Lys Val Ser Phe Asn Leu Asp Arg Thr Ser 165 170 175 Ser Lys His Glu Pro Asp Gly Arg Val Ile Ile Ala Glu Phe Glu Ser 180 185 190 Phe Leu Leu Leu Asn Thr Tyr Ala Pro Asn Asn Gly Trp Lys Glu Glu 195 200 205 Glu Asn Ser Phe Gln Arg Arg Arg Lys Trp Asp Lys Arg Met Leu Glu 210 215 220 Phe Val Gln Gln Val Asp Lys Pro Leu Ile Trp Cys Gly Asp Leu Asn 225 230 235 240 Val Ser His Glu Glu Ile Asp Val Ser His Pro Asp Phe Phe Ser Ser 245 250 255 Ala Lys Leu Asn Gly Tyr Ile Pro Pro Asn Lys Glu Asp Cys Gly Gln 260 265 270 Pro Gly Phe Thr Leu Ser Glu Arg Arg Arg Phe Gly Asn Ile Leu Ser 275 280 285 Gln Gly Lys Leu Val Asp Ala Tyr Arg Tyr Leu His Lys Glu Lys Asp 290 295 300 Met Asp Cys Gly Phe Ser Trp Ser Gly His Pro Ile Gly Lys Tyr Arg 305 310 315 320 Gly Lys Arg Met Arg Ile Asp Tyr Phe Leu Val Ser Glu Lys Leu Lys 325 330 335 Asp Gln Ile Val Ser Cys Asp Ile His Gly Arg Gly Ile Glu Leu Glu 340 345 350 Gly Phe Tyr Gly Ser Asp His Cys Pro Val Ser Leu Glu Leu Ser Glu 355 360 365 Glu Val Glu Ala Pro Lys Pro Lys Ser Ser Asn 370 375 <210> 20 <211> 263 <212> PRT <213> Artificial Sequence <220> <223> OsPCNA (wt) <400> 20 Met Leu Glu Leu Arg Leu Val Gln Gly Ser Leu Leu Lys Lys Val Leu 1 5 10 15 Glu Ala Ile Arg Glu Leu Val Thr Asp Ala Asn Phe Asp Cys Ser Gly 20 25 30 Thr Gly Phe Ser Leu Gln Ala Met Asp Ser Ser His Val Ala Leu Val 35 40 45 Ala Leu Leu Leu Arg Ser Glu Gly Phe Glu His Tyr Arg Cys Asp Arg 50 55 60 Asn Leu Ser Met Gly Met Asn Leu Asn Asn Met Ala Lys Met Leu Arg 65 70 75 80 Cys Ala Gly Asn Asp Asp Ile Ile Thr Ile Lys Ala Asp Asp Gly Ser 85 90 95 Asp Thr Val Thr Phe Met Phe Glu Ser Pro Asn Gln Asp Lys Ile Ala 100 105 110 Asp Phe Glu Met Lys Leu Met Asp Ile Asp Ser Glu His Leu Gly Ile 115 120 125 Pro Asp Ser Glu Tyr Gln Ala Ile Val Arg Met Pro Ser Ser Glu Phe 130 135 140 Ser Arg Ile Cys Lys Asp Leu Ser Ser Ile Gly Asp Thr Val Ile Ile 145 150 155 160 Ser Val Thr Lys Glu Gly Val Lys Phe Ser Thr Ala Gly Asp Ile Gly 165 170 175 Thr Ala Asn Ile Val Cys Arg Gln Asn Lys Thr Val Asp Lys Pro Glu 180 185 190 Asp Ala Thr Ile Ile Glu Met Gln Glu Pro Val Ser Leu Thr Phe Ala 195 200 205 Leu Arg Tyr Met Asn Ser Phe Thr Lys Ala Ser Pro Leu Ser Glu Gln 210 215 220 Val Thr Ile Ser Leu Ser Ser Glu Leu Pro Val Val Val Glu Tyr Lys 225 230 235 240 Ile Ala Glu Met Gly Tyr Ile Arg Phe Tyr Leu Ala Pro Lys Ile Glu 245 250 255 Glu Asp Glu Glu Met Lys Ser 260 <210> 21 <211> 2204 <212> PRT <213> Artificial Sequence <220> <223> Exemplary fusion protein of the PCGBE-3 system <400> 21 Met Glu Ala Ser Pro Ala Ser Gly Pro Arg His Leu Met Asp Pro His 1 5 10 15 Ile Phe Thr Ser Asn Phe Asn Asn Gly Ile Gly Arg His Lys Thr Tyr 20 25 30 Leu Cys Tyr Glu Val Glu Arg Leu Asp Asn Gly Thr Ser Val Lys Met 35 40 45 Asp Gln His Arg Gly Phe Leu His Asn Gln Ala Lys Asn Leu Leu Cys 50 55 60 Gly Phe Tyr Gly Arg His Ala Glu Leu Arg Phe Leu Asp Leu Val Pro 65 70 75 80 Ser Leu Gln Leu Asp Pro Ala Gln Ile Tyr Arg Val Thr Trp Phe Ile 85 90 95 Ser Trp Ser Pro Cys Phe Ser Trp Gly Cys Ala Gly Glu Val Arg Ala 100 105 110 Phe Leu Gln Glu Asn Thr His Val Arg Leu Arg Ile Phe Ala Ala Arg 115 120 125 Ile Tyr Asp Tyr Asp Pro Leu Tyr Lys Glu Ala Leu Gln Met Leu Arg 130 135 140 Asp Ala Gly Ala Gln Val Ser Ile Met Thr Tyr Asp Glu Phe Lys His 145 150 155 160 Cys Trp Asp Thr Phe Val Asp His Gln Gly Cys Pro Phe Gln Pro Trp 165 170 175 Asp Gly Leu Asp Glu His Ser Gln Ala Leu Ser Gly Arg Leu Arg Ala 180 185 190 Ile Leu Gln Asn Gln Gly Asn Ser Gly Ser Glu Thr Pro Gly Thr Ser 195 200 205 Glu Ser Ala Thr Pro Glu Ser Arg Pro Asp Lys Lys Tyr Ser Ile Gly 210 215 220 Leu Ala Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu 225 230 235 240 Tyr Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg 245 250 255 His Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly 260 265 270 Glu Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr 275 280 285 Thr Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn 290 295 300 Glu Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser 305 310 315 320 Phe Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe Gly 325 330 335 Asn Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr 340 345 350 His Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg 355 360 365 Leu Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His Phe 370 375 380 Leu Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu 385 390 395 400 Phe Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro 405 410 415 Ile Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu 420 425 430 Ser Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu 435 440 445 Lys Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu 450 455 460 Thr Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu 465 470 475 480 Gln Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala 485 490 495 Gln Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu 500 505 510 Ser Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile 515 520 525 Thr Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His 530 535 540 His Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro 545 550 555 560 Glu Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala 565 570 575 Gly Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile 580 585 590 Lys Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys 595 600 605 Leu Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly 610 615 620 Ser Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu Arg 625 630 635 640 Arg Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile 645 650 655 Glu Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala 660 665 670 Arg Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr 675 680 685 Ile Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala 690 695 700 Gln Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn 705 710 715 720 Glu Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val 725 730 735 Tyr Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys 740 745 750 Pro Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu Leu 755 760 765 Phe Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr 770 775 780 Phe Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu 785 790 795 800 Asp Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile 805 810 815 Ile Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu 820 825 830 Glu Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile 835 840 845 Glu Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met 850 855 860 Lys Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg 865 870 875 880 Lys Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu 885 890 895 Asp Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu 900 905 910 Ile His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln 915 920 925 Val Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu Ala 930 935 940 Gly Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val 945 950 955 960 Asp Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile Val 965 970 975 Ile Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn 980 985 990 Ser Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly 995 1000 1005 Ser Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln 1010 1015 1020 Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met 1025 1030 1035 Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp 1040 1045 1050 Val Asp His Ile Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile 1055 1060 1065 Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser 1070 1075 1080 Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys Asn Tyr 1085 1090 1095 Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe 1100 1105 1110 Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 1115 1120 1125 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile 1130 1135 1140 Thr Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys 1145 1150 1155 Tyr Asp Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr 1160 1165 1170 Leu Lys Ser Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe 1175 1180 1185 Tyr Lys Val Arg Glu Ile Asn Asn Tyr His His Ala His Asp Ala 1190 1195 1200 Tyr Leu Asn Ala Val Val Gly Thr Ala Leu Ile Lys Lys Tyr Pro 1205 1210 1215 Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp Tyr Lys Val Tyr Asp 1220 1225 1230 Val Arg Lys Met Ile Ala Lys Ser Glu Gln Glu Ile Gly Lys Ala 1235 1240 1245 Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile Met Asn Phe Phe Lys 1250 1255 1260 Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile Arg Lys Arg Pro Leu 1265 1270 1275 Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile Val Trp Asp Lys Gly 1280 1285 1290 Arg Asp Phe Ala Thr Val Arg Lys Val Leu Ser Met Pro Gln Val 1295 1300 1305 Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly Gly Phe Ser Lys 1310 1315 1320 Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp Lys Leu Ile Ala Arg 1325 1330 1335 Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe Asp Ser Pro 1340 1345 1350 Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val Glu Lys Gly 1355 1360 1365 Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu Gly Ile Thr 1370 1375 1380 Ile Met Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile Asp Phe Leu 1385 1390 1395 Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys Asp Leu Ile Ile Lys 1400 1405 1410 Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu Asn Gly Arg Lys Arg 1415 1420 1425 Met Leu Ala Ser Ala Gly Glu Leu Gln Lys Gly Asn Glu Leu Ala 1430 1435 1440 Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr Leu Ala Ser His Tyr 1445 1450 1455 Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn Glu Gln Lys Gln Leu 1460 1465 1470 Phe Val Glu Gln His Lys His Tyr Leu Asp Glu Ile Ile Glu Gln 1475 1480 1485 Ile Ser Glu Phe Ser Lys Arg Val Ile Leu Ala Asp Ala Asn Leu 1490 1495 1500 Asp Lys Val Leu Ser Ala Tyr Asn Lys His Arg Asp Lys Pro Ile 1505 1510 1515 Arg Glu Gln Ala Glu Asn Ile Ile His Leu Phe Thr Leu Thr Asn 1520 1525 1530 Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe Asp Thr Thr Ile Asp 1535 1540 1545 Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val Leu Asp Ala Thr Leu 1550 1555 1560 Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr Arg Ile Asp Leu 1565 1570 1575 Ser Gln Leu Gly Gly Asp Lys Arg Pro Ala Ala Thr Lys Lys Ala 1580 1585 1590 Gly Gln Ala Lys Lys Lys Lys Gly Thr Asp Ser Gly Gly Ser Ala 1595 1600 1605 Asn Glu Leu Thr Trp His Asp Val Leu Ala Glu Glu Lys Gln Gln 1610 1615 1620 Pro Tyr Phe Leu Asn Thr Leu Gln Thr Val Ala Ser Glu Arg Gln 1625 1630 1635 Ser Gly Val Thr Ile Tyr Pro Pro Gln Lys Asp Val Phe Asn Ala 1640 1645 1650 Phe Arg Phe Thr Glu Leu Gly Asp Val Lys Val Val Ile Leu Gly 1655 1660 1665 Gln Asp Pro Tyr His Gly Pro Gly Gln Ala His Gly Leu Ala Phe 1670 1675 1680 Ser Val Arg Pro Gly Ile Ala Ile Pro Pro Ser Leu Leu Asn Met 1685 1690 1695 Tyr Lys Glu Leu Glu Asn Thr Ile Pro Gly Phe Thr Arg Pro Asn 1700 1705 1710 His Gly Tyr Leu Glu Ser Trp Ala Arg Gln Gly Val Leu Leu Leu 1715 1720 1725 Asn Thr Val Leu Thr Val Arg Ala Gly Gln Ala His Ser His Ala 1730 1735 1740 Ser Leu Gly Trp Glu Thr Phe Thr Asp Lys Val Ile Ser Leu Ile 1745 1750 1755 Asn Gln His Arg Glu Gly Val Val Phe Leu Leu Trp Gly Ser His 1760 1765 1770 Ala Gln Lys Lys Gly Ala Ile Ile Asp Lys Gln Arg His His Val 1775 1780 1785 Leu Lys Ala Pro His Pro Ser Pro Leu Ser Ala His Arg Gly Phe 1790 1795 1800 Phe Gly Cys Asn His Phe Val Leu Ala Asn Gln Trp Leu Glu Gln 1805 1810 1815 Arg Gly Glu Thr Pro Ile Asp Trp Met Pro Val Leu Pro Ala Glu 1820 1825 1830 Ser Glu Pro Lys Lys Lys Arg Lys Val Leu Lys Glu Gly Arg Gly 1835 1840 1845 Ser Leu Leu Thr Cys Gly Asp Val Glu Glu Asn Pro Gly Pro Ser 1850 1855 1860 Ala Ile Arg Ala Ser Ser His Arg Leu Gln Thr Arg Thr Val Ala 1865 1870 1875 Leu Thr Arg Thr Lys Met Ser Ser Met Ala Gly Leu Gly Ala Ser 1880 1885 1890 Gln His Gly Tyr Pro Pro Arg Ser His Glu Pro Trp Thr Lys Leu 1895 1900 1905 Val His Arg Glu Arg Leu Pro Glu Trp Phe Ala Tyr Asn Pro Lys 1910 1915 1920 Thr Met Arg Pro Pro Pro Leu Ser His Asp Thr Lys Cys Met Lys 1925 1930 1935 Ile Leu Ser Trp Asn Ile Asn Gly Leu His Asp Val Val Thr Thr 1940 1945 1950 Lys Gly Phe Ser Ala Arg Asp Leu Ala Gln Arg Glu Asn Phe Asp 1955 1960 1965 Val Leu Cys Leu Gln Glu Thr His Leu Glu Glu Lys Asp Val Glu 1970 1975 1980 Lys Phe Lys Asn Leu Ile Ala Asp Tyr Asp Ser Tyr Trp Ser Cys 1985 1990 1995 Ser Val Ser Arg Leu Gly Tyr Ser Gly Thr Ala Val Ile Ser Arg 2000 2005 2010 Val Lys Pro Ile Ser Val Gln Tyr Gly Ile Gly Ile Arg Glu His 2015 2020 2025 Asp His Glu Gly Arg Val Ile Thr Leu Glu Phe Asp Gly Phe Tyr 2030 2035 2040 Leu Val Asn Ala Tyr Val Pro Asn Ser Gly Arg Phe Leu Arg Arg 2045 2050 2055 Leu Asn Tyr Arg Val Asn Asn Trp Asp Pro Cys Phe Ser Asn Tyr 2060 2065 2070 Val Lys Ile Leu Glu Lys Ser Lys Pro Val Ile Val Ala Gly Asp 2075 2080 2085 Leu Asn Cys Ala Arg Gln Ser Ile Asp Ile His Asn Pro Pro Ala 2090 2095 2100 Lys Thr Lys Ser Ala Gly Phe Thr Ile Glu Glu Arg Glu Ser Phe 2105 2110 2115 Glu Thr Asn Phe Ser Ser Lys Gly Leu Val Asp Thr Phe Arg Lys 2120 2125 2130 Gln His Pro Asn Ala Val Gly Tyr Thr Phe Trp Gly Glu Asn Gln 2135 2140 2145 Arg Ile Thr Asn Lys Gly Trp Arg Leu Ala Tyr Phe Leu Ala Ser 2150 2155 2160 Glu Ser Ile Thr Asp Lys Val His Asp Ser Tyr Ile Leu Pro Asp 2165 2170 2175 Val Ser Phe Ser Asp His Ser Pro Ile Gly Leu Val Leu Lys Leu 2180 2185 2190 Ser Gly Gly Ser Pro Lys Lys Lys Arg Lys Val 2195 2200 <210> 22 <211> 2251 <212> PRT <213> Artificial Sequence <220> <223> PCGBE-4 System Schematic Fusion Protein <400> 22 Met Glu Ala Ser Pro Ala Ser Gly Pro Arg His Leu Met Asp Pro His 1 5 10 15 Ile Phe Thr Ser Asn Phe Asn Asn Gly Ile Gly Arg His Lys Thr Tyr 20 25 30 Leu Cys Tyr Glu Val Glu Arg Leu Asp Asn Gly Thr Ser Val Lys Met 35 40 45 Asp Gln His Arg Gly Phe Leu His Asn Gln Ala Lys Asn Leu Leu Cys 50 55 60 Gly Phe Tyr Gly Arg His Ala Glu Leu Arg Phe Leu Asp Leu Val Pro 65 70 75 80 Ser Leu Gln Leu Asp Pro Ala Gln Ile Tyr Arg Val Thr Trp Phe Ile 85 90 95 Ser Trp Ser Pro Cys Phe Ser Trp Gly Cys Ala Gly Glu Val Arg Ala 100 105 110 Phe Leu Gln Glu Asn Thr His Val Arg Leu Arg Ile Phe Ala Ala Arg 115 120 125 Ile Tyr Asp Tyr Asp Pro Leu Tyr Lys Glu Ala Leu Gln Met Leu Arg 130 135 140 Asp Ala Gly Ala Gln Val Ser Ile Met Thr Tyr Asp Glu Phe Lys His 145 150 155 160 Cys Trp Asp Thr Phe Val Asp His Gln Gly Cys Pro Phe Gln Pro Trp 165 170 175 Asp Gly Leu Asp Glu His Ser Gln Ala Leu Ser Gly Arg Leu Arg Ala 180 185 190 Ile Leu Gln Asn Gln Gly Asn Ser Gly Ser Glu Thr Pro Gly Thr Ser 195 200 205 Glu Ser Ala Thr Pro Glu Ser Arg Pro Asp Lys Lys Tyr Ser Ile Gly 210 215 220 Leu Ala Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu 225 230 235 240 Tyr Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg 245 250 255 His Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly 260 265 270 Glu Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr 275 280 285 Thr Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn 290 295 300 Glu Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser 305 310 315 320 Phe Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe Gly 325 330 335 Asn Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr 340 345 350 His Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg 355 360 365 Leu Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His Phe 370 375 380 Leu Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu 385 390 395 400 Phe Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro 405 410 415 Ile Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu 420 425 430 Ser Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu 435 440 445 Lys Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu 450 455 460 Thr Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu 465 470 475 480 Gln Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala 485 490 495 Gln Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu 500 505 510 Ser Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile 515 520 525 Thr Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His 530 535 540 His Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro 545 550 555 560 Glu Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala 565 570 575 Gly Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile 580 585 590 Lys Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys 595 600 605 Leu Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly 610 615 620 Ser Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu Arg 625 630 635 640 Arg Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile 645 650 655 Glu Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala 660 665 670 Arg Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr 675 680 685 Ile Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala 690 695 700 Gln Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn 705 710 715 720 Glu Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val 725 730 735 Tyr Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys 740 745 750 Pro Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu Leu 755 760 765 Phe Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr 770 775 780 Phe Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu 785 790 795 800 Asp Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile 805 810 815 Ile Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu 820 825 830 Glu Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile 835 840 845 Glu Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met 850 855 860 Lys Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg 865 870 875 880 Lys Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu 885 890 895 Asp Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu 900 905 910 Ile His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln 915 920 925 Val Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu Ala 930 935 940 Gly Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val 945 950 955 960 Asp Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile Val 965 970 975 Ile Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn 980 985 990 Ser Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly 995 1000 1005 Ser Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln 1010 1015 1020 Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met 1025 1030 1035 Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp 1040 1045 1050 Val Asp His Ile Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile 1055 1060 1065 Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser 1070 1075 1080 Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys Asn Tyr 1085 1090 1095 Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe 1100 1105 1110 Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 1115 1120 1125 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile 1130 1135 1140 Thr Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys 1145 1150 1155 Tyr Asp Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr 1160 1165 1170 Leu Lys Ser Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe 1175 1180 1185 Tyr Lys Val Arg Glu Ile Asn Asn Tyr His His Ala His Asp Ala 1190 1195 1200 Tyr Leu Asn Ala Val Val Gly Thr Ala Leu Ile Lys Lys Tyr Pro 1205 1210 1215 Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp Tyr Lys Val Tyr Asp 1220 1225 1230 Val Arg Lys Met Ile Ala Lys Ser Glu Gln Glu Ile Gly Lys Ala 1235 1240 1245 Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile Met Asn Phe Phe Lys 1250 1255 1260 Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile Arg Lys Arg Pro Leu 1265 1270 1275 Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile Val Trp Asp Lys Gly 1280 1285 1290 Arg Asp Phe Ala Thr Val Arg Lys Val Leu Ser Met Pro Gln Val 1295 1300 1305 Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly Gly Phe Ser Lys 1310 1315 1320 Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp Lys Leu Ile Ala Arg 1325 1330 1335 Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe Asp Ser Pro 1340 1345 1350 Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val Glu Lys Gly 1355 1360 1365 Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu Gly Ile Thr 1370 1375 1380 Ile Met Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile Asp Phe Leu 1385 1390 1395 Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys Asp Leu Ile Ile Lys 1400 1405 1410 Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu Asn Gly Arg Lys Arg 1415 1420 1425 Met Leu Ala Ser Ala Gly Glu Leu Gln Lys Gly Asn Glu Leu Ala 1430 1435 1440 Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr Leu Ala Ser His Tyr 1445 1450 1455 Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn Glu Gln Lys Gln Leu 1460 1465 1470 Phe Val Glu Gln His Lys His Tyr Leu Asp Glu Ile Ile Glu Gln 1475 1480 1485 Ile Ser Glu Phe Ser Lys Arg Val Ile Leu Ala Asp Ala Asn Leu 1490 1495 1500 Asp Lys Val Leu Ser Ala Tyr Asn Lys His Arg Asp Lys Pro Ile 1505 1510 1515 Arg Glu Gln Ala Glu Asn Ile Ile His Leu Phe Thr Leu Thr Asn 1520 1525 1530 Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe Asp Thr Thr Ile Asp 1535 1540 1545 Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val Leu Asp Ala Thr Leu 1550 1555 1560 Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr Arg Ile Asp Leu 1565 1570 1575 Ser Gln Leu Gly Gly Asp Lys Arg Pro Ala Ala Thr Lys Lys Ala 1580 1585 1590 Gly Gln Ala Lys Lys Lys Lys Gly Thr Asp Ser Gly Gly Ser Ala 1595 1600 1605 Asn Glu Leu Thr Trp His Asp Val Leu Ala Glu Glu Lys Gln Gln 1610 1615 1620 Pro Tyr Phe Leu Asn Thr Leu Gln Thr Val Ala Ser Glu Arg Gln 1625 1630 1635 Ser Gly Val Thr Ile Tyr Pro Pro Gln Lys Asp Val Phe Asn Ala 1640 1645 1650 Phe Arg Phe Thr Glu Leu Gly Asp Val Lys Val Val Ile Leu Gly 1655 1660 1665 Gln Asp Pro Tyr His Gly Pro Gly Gln Ala His Gly Leu Ala Phe 1670 1675 1680 Ser Val Arg Pro Gly Ile Ala Ile Pro Pro Ser Leu Leu Asn Met 1685 1690 1695 Tyr Lys Glu Leu Glu Asn Thr Ile Pro Gly Phe Thr Arg Pro Asn 1700 1705 1710 His Gly Tyr Leu Glu Ser Trp Ala Arg Gln Gly Val Leu Leu Leu 1715 1720 1725 Asn Thr Val Leu Thr Val Arg Ala Gly Gln Ala His Ser His Ala 1730 1735 1740 Ser Leu Gly Trp Glu Thr Phe Thr Asp Lys Val Ile Ser Leu Ile 1745 1750 1755 Asn Gln His Arg Glu Gly Val Val Phe Leu Leu Trp Gly Ser His 1760 1765 1770 Ala Gln Lys Lys Gly Ala Ile Ile Asp Lys Gln Arg His His Val 1775 1780 1785 Leu Lys Ala Pro His Pro Ser Pro Leu Ser Ala His Arg Gly Phe 1790 1795 1800 Phe Gly Cys Asn His Phe Val Leu Ala Asn Gln Trp Leu Glu Gln 1805 1810 1815 Arg Gly Glu Thr Pro Ile Asp Trp Met Pro Val Leu Pro Ala Glu 1820 1825 1830 Ser Glu Pro Lys Lys Lys Arg Lys Val Leu Lys Glu Gly Arg Gly 1835 1840 1845 Ser Leu Leu Thr Cys Gly Asp Val Glu Glu Asn Pro Gly Pro Lys 1850 1855 1860 Arg Phe Phe Gln Pro Val Pro Lys Asp Gly Ser Pro Ala Lys Lys 1865 1870 1875 Arg Pro Ala Ala Ala Ala Ala Ala Ser Ala Ser Asp Ser Asp Ser 1880 1885 1890 Leu Gly Gly Asp Ala Pro Ala Ala Ala Ala Cys Ala Val Gly Glu 1895 1900 1905 Gly Asp Ser Pro Pro Ala Pro Arg Glu Glu Glu Pro Arg Arg Phe 1910 1915 1920 Val Thr Trp Asn Ala Asn Ser Leu Leu Leu Arg Met Lys Ser Asp 1925 1930 1935 Trp Pro Ala Phe Cys Gln Phe Val Ser Arg Val Asp Pro Asp Val 1940 1945 1950 Ile Cys Val Gln Glu Val Arg Met Pro Ala Ala Gly Ser Lys Gly 1955 1960 1965 Ala Pro Lys Asn Pro Gly Gln Leu Lys Asp Asp Thr Ser Ser Ser 1970 1975 1980 Arg Asp Glu Lys Gln Val Val Leu Arg Ala Leu Ser Ser Pro Pro 1985 1990 1995 Phe Lys Asp Tyr Arg Val Trp Trp Ser Leu Ser Asp Ser Lys Tyr 2000 2005 2010 Ala Gly Thr Ala Met Ile Ile Lys Lys Lys Phe Glu Pro Lys Lys 2015 2020 2025 Val Ser Phe Asn Leu Asp Arg Thr Ser Ser Lys His Glu Pro Asp 2030 2035 2040 Gly Arg Val Ile Ile Ala Glu Phe Glu Ser Phe Leu Leu Leu Asn 2045 2050 2055 Thr Tyr Ala Pro Asn Asn Gly Trp Lys Glu Glu Glu Asn Ser Phe 2060 2065 2070 Gln Arg Arg Arg Lys Trp Asp Lys Arg Met Leu Glu Phe Val Gln 2075 2080 2085 Gln Val Asp Lys Pro Leu Ile Trp Cys Gly Asp Leu Asn Val Ser 2090 2095 2100 His Glu Glu Ile Asp Val Ser His Pro Asp Phe Phe Ser Ser Ala 2105 2110 2115 Lys Leu Asn Gly Tyr Ile Pro Pro Asn Lys Glu Asp Cys Gly Gln 2120 2125 2130 Pro Gly Phe Thr Leu Ser Glu Arg Arg Arg Phe Gly Asn Ile Leu 2135 2140 2145 Ser Gln Gly Lys Leu Val Asp Ala Tyr Arg Tyr Leu His Lys Glu 2150 2155 2160 Lys Asp Met Asp Cys Gly Phe Ser Trp Ser Gly His Pro Ile Gly 2165 2170 2175 Lys Tyr Arg Gly Lys Arg Met Arg Ile Ala Tyr Phe Leu Val Ser 2180 2185 2190 Glu Lys Leu Lys Asp Gln Ile Val Ser Cys Asp Ile His Gly Arg 2195 2200 2205 Gly Ile Glu Leu Glu Gly Phe Tyr Gly Ser Asp His Cys Pro Val 2210 2215 2220 Ser Leu Glu Leu Ser Glu Glu Val Glu Ala Pro Lys Pro Lys Ser 2225 2230 2235 Ser Asn Ser Gly Gly Ser Pro Lys Lys Lys Arg Lys Val 2240 2245 2250 <210> 23 <211> 2367 <212> PRT <213> Artificial Sequence <220> <223> PCGBE-5 System Schematic Fusion Protein <400> 23 Met Glu Ala Ser Pro Ala Ser Gly Pro Arg His Leu Met Asp Pro His 1 5 10 15 Ile Phe Thr Ser Asn Phe Asn Asn Gly Ile Gly Arg His Lys Thr Tyr 20 25 30 Leu Cys Tyr Glu Val Glu Arg Leu Asp Asn Gly Thr Ser Val Lys Met 35 40 45 Asp Gln His Arg Gly Phe Leu His Asn Gln Ala Lys Asn Leu Leu Cys 50 55 60 Gly Phe Tyr Gly Arg His Ala Glu Leu Arg Phe Leu Asp Leu Val Pro 65 70 75 80 Ser Leu Gln Leu Asp Pro Ala Gln Ile Tyr Arg Val Thr Trp Phe Ile 85 90 95 Ser Trp Ser Pro Cys Phe Ser Trp Gly Cys Ala Gly Glu Val Arg Ala 100 105 110 Phe Leu Gln Glu Asn Thr His Val Arg Leu Arg Ile Phe Ala Ala Arg 115 120 125 Ile Tyr Asp Tyr Asp Pro Leu Tyr Lys Glu Ala Leu Gln Met Leu Arg 130 135 140 Asp Ala Gly Ala Gln Val Ser Ile Met Thr Tyr Asp Glu Phe Lys His 145 150 155 160 Cys Trp Asp Thr Phe Val Asp His Gln Gly Cys Pro Phe Gln Pro Trp 165 170 175 Asp Gly Leu Asp Glu His Ser Gln Ala Leu Ser Gly Arg Leu Arg Ala 180 185 190 Ile Leu Gln Asn Gln Gly Asn Ser Gly Ser Glu Thr Pro Gly Thr Ser 195 200 205 Glu Ser Ala Thr Pro Glu Ser Arg Pro Asp Lys Lys Tyr Ser Ile Gly 210 215 220 Leu Ala Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu 225 230 235 240 Tyr Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg 245 250 255 His Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly 260 265 270 Glu Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr 275 280 285 Thr Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn 290 295 300 Glu Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser 305 310 315 320 Phe Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe Gly 325 330 335 Asn Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr 340 345 350 His Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg 355 360 365 Leu Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His Phe 370 375 380 Leu Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu 385 390 395 400 Phe Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro 405 410 415 Ile Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu 420 425 430 Ser Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu 435 440 445 Lys Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu 450 455 460 Thr Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu 465 470 475 480 Gln Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala 485 490 495 Gln Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu 500 505 510 Ser Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile 515 520 525 Thr Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His 530 535 540 His Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro 545 550 555 560 Glu Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala 565 570 575 Gly Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile 580 585 590 Lys Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys 595 600 605 Leu Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly 610 615 620 Ser Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu Arg 625 630 635 640 Arg Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile 645 650 655 Glu Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala 660 665 670 Arg Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr 675 680 685 Ile Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala 690 695 700 Gln Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn 705 710 715 720 Glu Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val 725 730 735 Tyr Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys 740 745 750 Pro Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu Leu 755 760 765 Phe Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr 770 775 780 Phe Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu 785 790 795 800 Asp Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile 805 810 815 Ile Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu 820 825 830 Glu Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile 835 840 845 Glu Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met 850 855 860 Lys Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg 865 870 875 880 Lys Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu 885 890 895 Asp Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu 900 905 910 Ile His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln 915 920 925 Val Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu Ala 930 935 940 Gly Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val 945 950 955 960 Asp Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile Val 965 970 975 Ile Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn 980 985 990 Ser Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly 995 1000 1005 Ser Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln 1010 1015 1020 Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met 1025 1030 1035 Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp 1040 1045 1050 Val Asp His Ile Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile 1055 1060 1065 Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser 1070 1075 1080 Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys Asn Tyr 1085 1090 1095 Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe 1100 1105 1110 Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 1115 1120 1125 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile 1130 1135 1140 Thr Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys 1145 1150 1155 Tyr Asp Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr 1160 1165 1170 Leu Lys Ser Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe 1175 1180 1185 Tyr Lys Val Arg Glu Ile Asn Asn Tyr His His Ala His Asp Ala 1190 1195 1200 Tyr Leu Asn Ala Val Val Gly Thr Ala Leu Ile Lys Lys Tyr Pro 1205 1210 1215 Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp Tyr Lys Val Tyr Asp 1220 1225 1230 Val Arg Lys Met Ile Ala Lys Ser Glu Gln Glu Ile Gly Lys Ala 1235 1240 1245 Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile Met Asn Phe Phe Lys 1250 1255 1260 Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile Arg Lys Arg Pro Leu 1265 1270 1275 Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile Val Trp Asp Lys Gly 1280 1285 1290 Arg Asp Phe Ala Thr Val Arg Lys Val Leu Ser Met Pro Gln Val 1295 1300 1305 Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly Gly Phe Ser Lys 1310 1315 1320 Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp Lys Leu Ile Ala Arg 1325 1330 1335 Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe Asp Ser Pro 1340 1345 1350 Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val Glu Lys Gly 1355 1360 1365 Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu Gly Ile Thr 1370 1375 1380 Ile Met Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile Asp Phe Leu 1385 1390 1395 Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys Asp Leu Ile Ile Lys 1400 1405 1410 Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu Asn Gly Arg Lys Arg 1415 1420 1425 Met Leu Ala Ser Ala Gly Glu Leu Gln Lys Gly Asn Glu Leu Ala 1430 1435 1440 Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr Leu Ala Ser His Tyr 1445 1450 1455 Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn Glu Gln Lys Gln Leu 1460 1465 1470 Phe Val Glu Gln His Lys His Tyr Leu Asp Glu Ile Ile Glu Gln 1475 1480 1485 Ile Ser Glu Phe Ser Lys Arg Val Ile Leu Ala Asp Ala Asn Leu 1490 1495 1500 Asp Lys Val Leu Ser Ala Tyr Asn Lys His Arg Asp Lys Pro Ile 1505 1510 1515 Arg Glu Gln Ala Glu Asn Ile Ile His Leu Phe Thr Leu Thr Asn 1520 1525 1530 Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe Asp Thr Thr Ile Asp 1535 1540 1545 Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val Leu Asp Ala Thr Leu 1550 1555 1560 Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr Arg Ile Asp Leu 1565 1570 1575 Ser Gln Leu Gly Gly Asp Lys Arg Pro Ala Ala Thr Lys Lys Ala 1580 1585 1590 Gly Gln Ala Lys Lys Lys Lys Gly Thr Asp Ser Gly Gly Ser Ala 1595 1600 1605 Asn Glu Leu Thr Trp His Asp Val Leu Ala Glu Glu Lys Gln Gln 1610 1615 1620 Pro Tyr Phe Leu Asn Thr Leu Gln Thr Val Ala Ser Glu Arg Gln 1625 1630 1635 Ser Gly Val Thr Ile Tyr Pro Pro Gln Lys Asp Val Phe Asn Ala 1640 1645 1650 Phe Arg Phe Thr Glu Leu Gly Asp Val Lys Val Val Ile Leu Gly 1655 1660 1665 Gln Asp Pro Tyr His Gly Pro Gly Gln Ala His Gly Leu Ala Phe 1670 1675 1680 Ser Val Arg Pro Gly Ile Ala Ile Pro Pro Ser Leu Leu Asn Met 1685 1690 1695 Tyr Lys Glu Leu Glu Asn Thr Ile Pro Gly Phe Thr Arg Pro Asn 1700 1705 1710 His Gly Tyr Leu Glu Ser Trp Ala Arg Gln Gly Val Leu Leu Leu 1715 1720 1725 Asn Thr Val Leu Thr Val Arg Ala Gly Gln Ala His Ser His Ala 1730 1735 1740 Ser Leu Gly Trp Glu Thr Phe Thr Asp Lys Val Ile Ser Leu Ile 1745 1750 1755 Asn Gln His Arg Glu Gly Val Val Phe Leu Leu Trp Gly Ser His 1760 1765 1770 Ala Gln Lys Lys Gly Ala Ile Ile Asp Lys Gln Arg His His Val 1775 1780 1785 Leu Lys Ala Pro His Pro Ser Pro Leu Ser Ala His Arg Gly Phe 1790 1795 1800 Phe Gly Cys Asn His Phe Val Leu Ala Asn Gln Trp Leu Glu Gln 1805 1810 1815 Arg Gly Glu Thr Pro Ile Asp Trp Met Pro Val Leu Pro Ala Glu 1820 1825 1830 Ser Glu Pro Lys Lys Lys Arg Lys Val Leu Lys Glu Gly Arg Gly 1835 1840 1845 Ser Leu Leu Thr Cys Gly Asp Val Glu Glu Asn Pro Gly Pro Asp 1850 1855 1860 Ser Leu Ala Glu Ser Arg Trp Pro Pro Gly Leu Ala Val Met Lys 1865 1870 1875 Thr Ile Asp Asp Leu Leu Arg Cys Gly Ile Cys Phe Glu Tyr Phe 1880 1885 1890 Asn Ile Ala Met Ile Ile Pro Gln Cys Ser His Asn Tyr Cys Ser 1895 1900 1905 Leu Cys Ile Arg Lys Phe Leu Ser Tyr Lys Thr Gln Cys Pro Thr 1910 1915 1920 Cys Cys Val Thr Val Thr Glu Pro Asp Leu Lys Asn Asn Arg Ile 1925 1930 1935 Leu Asp Glu Leu Val Lys Ser Leu Asn Phe Ala Arg Asn His Leu 1940 1945 1950 Leu Gln Phe Ala Leu Glu Ser Pro Ala Lys Ser Pro Ala Ser Ser 1955 1960 1965 Ser Ser Lys Asn Leu Ala Val Lys Val Tyr Thr Pro Val Ala Ser 1970 1975 1980 Arg Gln Ser Leu Lys Gln Gly Ser Arg Leu Met Asp Asn Phe Leu 1985 1990 1995 Ile Arg Glu Met Ser Gly Ser Thr Ser Glu Leu Leu Ile Lys Glu 2000 2005 2010 Asn Lys Ser Lys Phe Ser Pro Gln Lys Glu Ala Ser Pro Ala Ala 2015 2020 2025 Lys Thr Lys Glu Thr Arg Ser Val Glu Glu Ile Ala Pro Asp Pro 2030 2035 2040 Ser Glu Ala Lys Arg Pro Glu Pro Pro Ser Thr Ser Thr Leu Lys 2045 2050 2055 Gln Val Thr Lys Val Asp Cys Pro Val Cys Gly Val Asn Ile Pro 2060 2065 2070 Glu Ser His Ile Asn Lys His Leu Asp Ser Cys Leu Ser Arg Glu 2075 2080 2085 Glu Lys Lys Glu Ser Leu Arg Ser Ser Val His Lys Arg Lys Pro 2090 2095 2100 Leu Pro Lys Thr Val Tyr Asn Leu Leu Ser Asp Arg Asp Leu Lys 2105 2110 2115 Lys Lys Leu Lys Glu His Gly Leu Ser Ile Gln Gly Asn Lys Gln 2120 2125 2130 Gln Leu Ile Lys Arg His Gln Glu Phe Val His Met Tyr Asn Ala 2135 2140 2145 Gln Cys Asp Ala Leu His Pro Lys Ser Ala Ala Glu Ile Val Arg 2150 2155 2160 Glu Ile Glu Asn Ile Glu Lys Thr Arg Met Arg Leu Glu Ala Ser 2165 2170 2175 Lys Leu Asn Glu Ser Val Met Val Phe Thr Lys Asp Gln Thr Glu 2180 2185 2190 Lys Glu Ile Asp Glu Ile His Ser Lys Tyr Arg Lys Lys His Lys 2195 2200 2205 Ser Glu Phe Gln Leu Leu Val Asp Gln Ala Arg Lys Gly Tyr Lys 2210 2215 2220 Lys Ile Ala Gly Met Ser Gln Lys Thr Val Thr Ile Thr Lys Glu 2225 2230 2235 Asp Glu Ser Thr Glu Lys Leu Ser Ser Val Cys Met Gly Gln Glu 2240 2245 2250 Asp Asn Met Thr Ser Val Thr Asn His Phe Ser Gln Ser Lys Leu 2255 2260 2265 Asp Ser Pro Glu Glu Leu Glu Pro Asp Arg Glu Glu Asp Ser Ser 2270 2275 2280 Ser Cys Ile Asp Ile Gln Glu Val Leu Ser Ser Ser Glu Ser Asp 2285 2290 2295 Ser Cys Asn Ser Ser Ser Ser Asp Ile Ile Arg Asp Leu Leu Glu 2300 2305 2310 Glu Glu Glu Ala Trp Glu Ala Ser His Lys Asn Asp Leu Gln Asp 2315 2320 2325 Thr Glu Ile Ser Pro Arg Gln Asn Arg Arg Thr Arg Ala Ala Glu 2330 2335 2340 Ser Ala Glu Ile Glu Pro Arg Asn Lys Arg Asn Arg Asn Ser Gly 2345 2350 2355 Gly Ser Pro Lys Lys Lys Arg Lys Val 2360 2365
Claims
1. A C-to-G base editing system for editing a target sequence in the genome of a plant cell, comprising: A first polypeptide and / or an expression construct comprising a nucleotide sequence encoding the first polypeptide, wherein the first polypeptide is in sequence a cytosine deaminase, a nuclease-inactivated CRISPR effector protein, and uracil-DNA glycosylase (UDG), wherein the UDG has the amino acid sequence shown in SEQ ID NO:5, the nuclease-inactivated CRISPR effector protein is nCas9, whose amino acid sequence is as shown in SEQ ID NO:3, and the cytosine deaminase is human APOBEC3A deaminase, whose amino acid sequence is as shown in SEQ ID NO:1; A guide RNA and / or an expression construct comprising a nucleotide sequence encoding the guide RNA, wherein the guide RNA can target the first polypeptide to the target sequence in the genome of the plant cell; and Any one of i)-iii): i) A second polypeptide and / or an expression construct comprising a nucleotide sequence encoding the second polypeptide, wherein the second polypeptide comprises a fusion of proliferating cell nuclear antigen (PCNA) with a mutated ubiquitin protein binding site and ubiquitin protein, and MCP; the MCP is fused to the N-terminus of the PCNA with the mutated ubiquitin protein binding site, wherein the PCNA with the mutated ubiquitin protein binding site has the amino acid sequence shown in SEQ ID NO:9, the ubiquitin protein is a truncated ubiquitin protein, and the truncated ubiquitin protein has the amino acid sequence shown in SEQ ID NO:10; ii) A third polypeptide and / or an expression construct comprising a nucleotide sequence encoding the third polypeptide, wherein the third polypeptide comprises a mutated AP endonuclease (APE), the APE is derived from rice APE12g and contains the amino acid substitution D327A relative to wild-type APE12g, and the mutated APE has the amino acid sequence shown in SEQ ID NO:16; iii) A fourth polypeptide and / or an expression construct comprising a nucleotide sequence encoding the fourth polypeptide, wherein the fourth polypeptide comprises a Rad18 protein, and the Rad18 protein is a human Rad18 protein, whose amino acid sequence is as shown in SEQ ID NO:17; Wherein the second polypeptide, the third polypeptide, or the fourth polypeptide is located at the C-terminus of the first polypeptide.
2. The C-to-G base editing system of claim 1, wherein the MCP has the amino acid sequence shown in SEQ ID NO:
7.
3. The C-to-G base editing system of any one of claims 1 or 2, wherein the expression construct comprises a nucleotide sequence encoding the amino acid sequence of SEQ ID NO:13 or one of SEQ ID NO:22-23.
4. A method for producing a genetically modified plant cell, comprising introducing the gene editing system of any one of claims 1-3 into the plant cell.
5. The method of claim 4, wherein the genetic modification is one or more C-to-G substitutions in the target sequence.
6. The method according to any one of claims 4 or 5, wherein the plant cells include monocotyledonous plants and dicotyledonous plants, including rice, corn, wheat, sorghum, barley, soybean, peanut, and Arabidopsis thaliana.
7. A kit, which comprises the gene editing system according to any one of claims 1-3, and instructions for use.
Citation Information
Patent Citations
Improved gene editing system
CN114008207A
Novel gene editing system for mediating A-to-C mutation or T-to-G mutation and application thereof
CN116200382A
Improved CG base editing system
CN117043345A
Single-base editing system and application
CN117586987A
Improved cytosine to guanine base editors
WO2022261509A1