A novel genome editing system based on c2c9 nuclease and application thereof

By developing an extremely small CRISPR/C2C9 genome editing system, the limitations of large molecular size and PAM sites have been overcome, enabling precise localization and efficient cutting of genomic DNA, suitable for gene editing in both prokaryotic and eukaryotic cells.

CN116200368BActive Publication Date: 2026-06-02SHANGHAI TECH UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI TECH UNIV
Filing Date
2021-11-30
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing CRISPR/Cas9 and CRISPR/Cas12a genome editing systems are limited in their application in gene therapy and other fields due to their large molecular size and PAM site limitations, as well as their low packaging efficiency and limited selection of genomic sites.

Method used

A very small CRISPR/C2C9 genome editing system has been developed, which includes a C2C9 nuclease and guide RNA, recognizes PAM sites as AAN and/or GAN, and can achieve precise gene editing in prokaryotic and eukaryotic cells.

Benefits of technology

It achieves precise localization and efficient cutting of genomic DNA, reduces packaging difficulty, and improves the efficiency and precision of gene editing, making it suitable for gene editing in prokaryotic bacteria and eukaryotic cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0003384763560000011
    Figure HDA0003384763560000011
  • Figure HDA0003384763560000012
    Figure HDA0003384763560000012
  • Figure HDA0003384763560000021
    Figure HDA0003384763560000021
Patent Text Reader

Abstract

The application belongs to the field of biological medicine, and discloses a use of C2C9 nuclease as an RNA-guided endonuclease and a minisize genome editing system, wherein the genome editing system comprises C2C9 nuclease and / or nucleic acid encoding the C2C9 nuclease, and guide RNA or nucleic acid encoding the guide RNA. The minisize genome editing system can edit at least one target sequence in the genome of prokaryotic bacteria and eukaryotic cells. The application also discloses application of the genome editing system and editing method. The target gene can be precisely knocked out or cut off by using the genome editing system or method, the cutting efficiency is high, and the precise editing of the target gene is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biomedicine and relates to a novel genome editing system based on a very small C2C9 nuclease, as well as the genome editing method and application of this system in prokaryotic bacteria and eukaryotic cells. Background Technology

[0002] CRISPR / Cas (Clustered Regularly Interspaced Short Palindromic Repeats / CRISPR-Associated Protein) is a novel gene-editing system developed in recent years. This system mainly consists of a Cas endonuclease and a corresponding guide RNA. The Cas endonuclease specifically binds to the guide RNA, using complementary base pairing to target and cleave specific sites in the genome, causing double-strand breaks in the genomic DNA. Then, endogenous or exogenous DNA repair mechanisms, such as homologous recombination and non-homologous recombination end-joining repair mechanisms, are used to repair the broken genomic DNA, thereby achieving editing of specific sites in the genome during the repair process.

[0003] The CRISPR / Cas system, used for genetic manipulation of cells or species, has wide applications in basic research, biotechnology development, and disease treatment research in life sciences, medicine, and agriculture. Examples include: correcting mutated genes in genetic diseases or cancer; genetically engineering crops to improve their performance; precisely modifying microbial genomes to drive the production of high-value-added compounds; and cutting and killing pathogenic microorganisms for infection treatment. Due to its simplicity and efficiency, the CRISPR / Cas system has been widely used in medicine, biology, and agriculture.

[0004] Currently, the widely used CRISPR / Cas genome editing systems mainly include two types: CRISPR / Cas9 and CRISPR / Cas12a. In both of these systems, the CRISPR effector proteins, the nucleases Cas9 and Cas12a, are large proteins containing over 1000 amino acids. This presents significant challenges during the delivery of CRISPR / Cas9 or CRISPR / Cas12a to cells. For example, in in vivo gene therapy, CRISPR / Cas9 or CRISPR / Cas12a systems often need to be packaged into lentiviral vectors. However, the limited packaging capacity of lentiviral vectors significantly restricts the packaging efficiency of CRISPR / Cas9 and CRISPR / Cas12a systems due to their large molecular size, increasing packaging difficulty and severely limiting their widespread application in gene therapy and other fields. Furthermore, the limited number of protospacer-adjacent motifs (PAM, PAM sequences, or PAM sites) recognized by the CRISPR / Cas system restricts the selectable genomic sites, hindering its in-depth application in medicine and biology. Therefore, the discovery and development of novel, efficient, and low-molecular-weight genome editing systems is particularly urgent. Summary of the Invention

[0005] This invention provides a very small genome editing system (including a CRISPR / C2C9 system) and method. The genome editing system of this invention allows for precise cutting of target sites in prokaryotic bacteria and eukaryotic cells, thereby achieving precise gene knockout or editing. This invention also provides novel RNA-programmable endonucleases. This invention also relates to other components, including guide RNA (gRNA) and / or target sequences, and methods for generating and using them in various applications. Examples of such applications include methods for regulating transcription, and methods for targeting, editing, and / or manipulating genes (such as DNA or RNA) using novel nucleases and other components such as nucleic acids and / or peptides. This invention also relates to recombinant cells and kits containing the elements of the said editing system.

[0006] This invention provides a genome editing system for editing at least one target sequence in the genome, comprising:

[0007] (a) C2C9 nuclease or nucleic acid encoding C2C9 nuclease; and

[0008] (b) Guide RNA (gRNA) or nucleic acid encoding said gRNA.

[0009] The C2C9 nuclease and its matching guide RNA form a complex.

[0010] In a preferred embodiment, the genome editing system is a CRISPR-C2C9 genome editing system.

[0011] In a preferred embodiment, the CRISPR-C2C9 genome editing system comprises:

[0012] (1) Expression construct of C2C9 nuclease;

[0013] (2) Complete guide RNA expression construct;

[0014] The complete guide RNA comprises a guide RNA backbone and an RNA sequence corresponding to a target sequence added to the 3' end of the guide RNA backbone; the target sequence is a nucleic acid fragment of 12-40 bp in length following the PAM sequence, preferably a nucleic acid fragment of 20 bp in length following the PAM sequence. The C2C9 nuclease described in this invention can be a conventional C2C9 nuclease in the art. Preferably, the size of the C2C9 nuclease is preferably no more than 800 amino acids. More preferably, the C2C9 nuclease is derived from any of the following amino acid sequences or has at least 80% sequence homology with any of the following amino acid sequences: SEQ ID NO.3-117. In some embodiments, the C2C9 nuclease is selected from the group consisting of Actinomaduracraniellae C2C9 (AcC2C9), Corynebacterium glutamicum C2C9 (CgC2C9), Rothiadentocariosa C2C9 (RdC2C9), and Micrococcus luteus C2C9 (MiC2C9).

[0015] In a preferred embodiment of the present invention, the C2C9 nuclease is the AcC2C9 nuclease shown in SEQ ID NO.3 of the sequence listing.

[0016] In some embodiments, the nucleic acid encoding the AcC2C9 nuclease is deoxyribonucleic acid (DNA). In some embodiments, the DNA coding sequence of the AcC2C9 nuclease is as shown in SEQ ID NO:121 in the sequence listing or a variant having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence homology with SEQ ID NO:121.

[0017] In some embodiments, the nucleic acid encoding the AcC2C9 nuclease is ribonucleic acid (RNA). In some embodiments, the RNA encoding the C2C9 nuclease is mRNA.

[0018] In some embodiments, the sequence encoding the AcC2C9 nuclease is a coding sequence optimized for human codons, preferably as shown in SEQ ID NO:122 in the sequence listing or a variant having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence homology with SEQ ID NO:122.

[0019] In some embodiments, the C2C9 nuclease or the nucleic acid encoding the C2C9 nuclease is formulated in liposomes or lipid nanoparticles. In some embodiments, the liposomes or lipid nanoparticles also contain the gRNA or the nucleic acid encoding the gRNA.

[0020] In some embodiments, the system includes a C2C9 nuclease pre-complexed with gRNA to form a ribonucleoprotein (RNP) complex that recognizes a PAM sequence on a target gene (such as a target DNA) sequence; that is, precisely locating a complex flanking the DNA sequence, including the C2C9 nuclease and the guide RNA, to recognize a PAM sequence on the target gene.

[0021] The guide RNA backbone sequence of the AcC2C9 nuclease is shown in SEQ ID NO.120 in the sequence listing.

[0022] In some embodiments, the complete guide RNA expression construct comprises an expression DNA sequence of a guide RNA backbone and an expression DNA sequence of an RNA sequence corresponding to a target sequence added to the 3' end of the expression DNA sequence; wherein the target sequence is a 20 bp fragment following the PAM sequence.

[0023] The CRISPR-C2C9 genome editing system preferably identifies PAM sites as AAN and / or GAN.

[0024] In some embodiments, the system of the present invention may further include a donor template comprising a heterologous polynucleotide sequence, wherein the heterologous polynucleotide sequence can be inserted into the target polynucleotide sequence. In some embodiments, the heterologous polynucleotide donor template is physically linked to gRNA.

[0025] In another aspect, the present invention provides a method for targeting, editing, modifying, or manipulating double-stranded DNA at a target locus, the method comprising mixing a C2C9 nuclease, a complete guide RNA, and the double-stranded DNA in the system described above; wherein the complete guide RNA is: 5'-a guide RNA backbone corresponding to the C2C9 nuclease + an RNA sequence corresponding to the target sequence on the double-stranded DNA - 3'. In some embodiments, the method is used to cleave the double-stranded DNA.

[0026] In some embodiments, the present invention provides a method for targeting, editing, modifying, or manipulating double-stranded DNA at target loci in cells as described above. In some embodiments, the method further includes the step of introducing a polynucleotide donor template into the cell, the donor template being a single-stranded or double-stranded polynucleotide donor template.

[0027] In some embodiments of the present invention, DNA is repaired at DSBs via homology-guided repair, non-homologous end joining, or microhomology-mediated end joining.

[0028] As described above, the preferred guide RNA backbone corresponding to the AcC2C9 nuclease is shown in SEQ ID NO. 120 of the sequence listing. A complete guide RNA can be formed by adding a target RNA sequence for a different target sequence to its 3' end.

[0029] The preferred sources of the C2C9 nuclease are as described above, preferably from the group consisting of SEQ ID NO. 3-117, etc.; more preferably, selected from Actinomadura craniellae C2C9 (AcC2C9), Corynebacterium glutamicum C2C9 (CgC2C9), Rothia dentocariosa C2C9 (RdC2C9), and Micrococcus luteus C2C9 (MiC2C9); more preferably as shown in SEQ ID NO. 3 or a variant thereof in the sequence listing. A more preferred coding sequence of the AcC2C9 nuclease optimized for human codons is shown in SEQ ID NO. 122 in the sequence listing. The variant described herein has at least 20% (e.g., 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%) sequence homology with the amino acid residues of SEQ ID NO.3 and retains RNA-guided DNA binding activity and / or double-stranded DNA cleavage activity.

[0030] The cutting method described in this invention can be conventional in the art. For example, in a preferred embodiment of this invention, the cutting method is carried out in a buffer solution of 150 mM NaCl, 10 mM MgCl2, 10 mM Tris-HCl, pH = 7.5, and 1 mM DTT.

[0031] The preferred reaction temperature for the cutting method is 37°C, and the preferred reaction time is 30-60 min.

[0032] In another aspect, the present invention provides a genome editing method comprising introducing a genome editing system as described above into prokaryotic bacteria or eukaryotic cells containing a target sequence for genome editing.

[0033] The prokaryotic bacteria described in this invention include, but are not limited to, prokaryotic microorganisms such as Escherichia coli and Klebsiella pneumoniae.

[0034] The eukaryotic cells described in this invention include, but are not limited to, mammalian cells, yeast, and other eukaryotic cells.

[0035] The “gene editing” in the genome editing methods described above can be conventional in the field, such as gene cutting, gene deletion, gene insertion, point mutation, transcriptional repression, transcriptional activation, base editing and guided editing, etc., but is not limited to these.

[0036] In another aspect, the present invention provides the application of C2C9 nuclease in genome editing.

[0037] The C2C9 nuclease and the gene editing described herein are further defined as described above.

[0038] The nuclease in this genome editing system is no larger than 800 amino acids.

[0039] In another aspect, the present invention provides a modified cell obtained by editing the genome of a cell using any of the systems or methods described above. In one embodiment, the modified cell comprises: (a) a C2C9 nuclease or nucleic acid encoding the C2C9 nuclease as described above; and (b) a guide RNA (gRNA) or nucleic acid encoding the gRNA, wherein the gRNA is capable of directing the C2C9 nuclease or a variant thereof to a target polynucleotide sequence. In some embodiments, the modified cell further comprises a donor template containing a heterologous polynucleotide sequence, wherein the heterologous polynucleotide sequence is capable of being inserted into the target polynucleotide sequence.

[0040] The positive and progressive effects of this invention include:

[0041] CRISPR / C2C9 is a novel genome editing system discovered by the applicant, comprising a C2C9 nuclease and gRNA. This system is characterized by its small size (C2C9 nucleases are generally no more than, or even less than, 800 amino acids) and the absence of complex PAMs (characteristics such as recognizing PAM sites as AAN). Through the complete guidance and localization functions of the guide RNA, the C2C9 nuclease can precisely locate, target, and cleave genomic DNA in vivo, in vitro, or cell-free environments, achieving double-strand breaks. Utilizing the host cell's own or exogenously supplemented repair mechanisms, this system can efficiently and accurately achieve genome editing in vivo, in vitro, or cell-free environments. Attached Figure Description

[0042] Figure 1 This diagram shows the identification process and results of PAM sequences using AcC2C9. The results indicate that AcC2C9 can effectively recognize 5'-AAN type PAM sequences and a small number of 5'-GAN sequences (where N is a degenerate base, representing any one of A, T, C, or G). The horizontal axes -1, -2, -3, -4, -5, and -6 represent the 1st, 2nd, 3rd, 4th, 5th, and 6th DNA bases upstream of the 5' end of the target-DNA segment sequence on the NTS strand (non-target strand), respectively. The letters above these bases represent the preference of the AcC2C9 nuclease for that DNA sequence; the greater the proportion of the letter, the stronger the preference of AcC2C9 for that base.

[0043] Figure 2 This image shows the PCR results of gene knockout in live Klebsiella pneumoniae cells using the CRISPR-AcC2C9 system. The results show that the selected dha gene can be precisely knocked out by the CRISPR-AcC2C9 system.

[0044] Figure 3 This image shows the results of cutting substrate DNA using the AcC2C9 nuclease. The results show that the substrate DNA was precisely cleaved into two DNA fragments by AcC2C9.

[0045] Figure 4 (A) Schematic diagram of the complete guide RNA sequence of AcC2C9. (B) Schematic diagram of the results of substrate DNA cleavage by AcC2C9 nuclease under tracrRNA and crRNA conditions.

[0046] Figure 5 The image shows the results of cleavage of four 58 bp fluorescently labeled DNA substrates using the AcC2C9 nuclease. The results show that the major cleavage site of AcC2C9 on the target strand is 21-22 nt downstream of the 3' end of the PAM site, while the major cleavage site on the non-target strand is 15-16 nt downstream of the 3' end of the PAM site.

[0047] Figure 6 The figures show the results of gene editing in human cells mediated by the AcC2C9 transient expression plasmid. (A) is a TBE-PAGE gel image from the Indel experiment. The AcC2C9 nuclease successfully achieved efficient gene editing of three genes—VEGFA, HEXA, and DNMT1—at five different selected gene loci. (B) is a statistical graph of high-throughput sequencing results after VEGFA gene editing in human cells mediated by the AcC2C9 transient expression plasmid. As shown, the AcC2C9 nuclease can precisely introduce insertion or deletion mutations at target sequences. Detailed Implementation

[0048] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined by the appended claims.

[0049] The terms “C2C9”, “C2C9 nuclease”, “C2C9 polypeptide”, “C2C9 protein”, and “C2C9 protein” are used interchangeably.

[0050] The terms “guide RNA”, “gRNA”, “single gRNA”, and “chimeric gRNA” are used interchangeably.

[0051] It should be noted that the term "a" or "an" entity in this invention refers to one or more of the entity; therefore, the terms "a" (or "an"), "one or more", and "at least one" are used interchangeably herein.

[0052] The terms "homology," "identity," or "similarity" refer to the sequence similarity between two peptides or two nucleic acid molecules. Homology can be determined by comparing corresponding positions in different polypeptide or nucleic acid molecules. When the same position in the sequence of the compared molecules is occupied by the same base or amino acid in different sequences, then the molecules are homologous at that position. The degree of homology between sequences is determined as a function of the number of common matching or homologous positions. An "unrelated" or "non-homologous" sequence should have less than 20% homology with one of the sequences disclosed in this invention.

[0053] A percentage sequence homology (e.g., 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99%) between a polynucleotide or polynucleotide region (or polypeptide or polypeptide region) and another polynucleotide or polynucleotide region (or polypeptide or polypeptide region) means that, at the time of alignment, the two sequences being aligned contain that percentage of identical bases (or amino acids). This alignment and percentage homology or sequence identity can be determined using software programs and methods known in the art, such as those described in Ausubel et al. eds. (2007) Current Protocols in Molecular Biology. Preferably, default parameters should be used during sequence alignment. One alternative alignment procedure is BLAST, using default parameters. Specifically, when using the BLASTN and BLASTP alignment programs, the following default parameters are used: Genetic code = standard; filter = none; strand = both; cutoff = 60; expect = 10; Matrix = BLOSUM62; Descriptions = 50 sequences; sort by = HIGH SCORE; Databases = non-redundant, GenBank+EMBL+DDBJ+PDB+GenBank CDS translations+SwissProtein+SPupdate+PIR. Biologically equivalent polynucleotides are those that possess the above-defined percentage of homology and encode polypeptides with the same or similar biological activities.

[0054] When the polynucleotide is DNA, its sequence consists of the letters representing the four nucleotide bases: adenine (A), cytosine (C), guanine (G), and thymine (T). When the polynucleotide is RNA, its sequence consists of the letters representing the four nucleotide bases: adenine (A), cytosine (C), guanine (G), and uracil (U). Therefore, the term "polynucleotide sequence" is the letter representation of a polynucleotide molecule. This letter representation can be entered into a database in a computer with a central processing unit and used for bioinformatics applications such as functional genomics and homology searches. The term "polymorphism" refers to the coexistence of more than one form of a gene or part thereof; a "polymorphic region of a gene" refers to a gene having different nucleotide expressions (i.e., different nucleotide sequences) at the same location. A polymorphic region of a gene can be a single nucleotide that differs in different alleles.

[0055] In this invention, the terms "polynucleotide" and "oligonucleotide" are used interchangeably, and they refer to polymeric forms of nucleotides of any length, whether deoxyribonucleotides, ribonucleotides, or analogs thereof. Polynucleotides can have any three-dimensional structure and can perform any function, known or unknown. Examples of polynucleotides include, but are not limited to, the following: genes or gene fragments (including probes, primers, EST or SAGE tags), exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, dsRNA, siRNA, miRNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Polynucleotides also include modified nucleotides, such as methylated nucleotides and nucleotide analogs. If modifications are present on the polynucleotide, these modifications can be conferred before or after assembly. Nucleotide sequences can be broken down by non-nucleotide components. Polynucleotides can be further modified after polymerization, for example, by coupling with labeled components. The term refers to both double-stranded and single-stranded polynucleotide molecules. Unless otherwise stated or required, any embodiment of the polynucleotide disclosed in this invention includes its double-stranded form and any one of two complementary single-stranded forms known or predicted to constitute the double-stranded form.

[0056] When applied to polynucleotides, the term "encoding" refers to a polynucleotide "encoding" a polypeptide, meaning that in its natural state or when manipulated by methods known to those skilled in the art, it can be transcribed and / or translated to produce a target polypeptide and / or fragments thereof, or to produce mRNA capable of encoding the target polypeptide and / or fragments thereof. The antisense strand refers to a sequence complementary to the polynucleotide from which the coding sequence can be deduced.

[0057] The term "genomic DNA" refers to the DNA of an organism's genome, including the DNA of bacteria, archaea, fungi, protozoa, viruses, plants, or animals.

[0058] The term "manipulating" DNA includes binding, creating a nick in one strand, or cutting both strands of DNA, or includes modifying or editing DNA or a polypeptide bound to DNA. Manipulating DNA can silence, activate, or regulate the expression of RNA or polypeptides encoded by said DNA (preventing transcription, reducing transcriptional activity, preventing translation, or reducing translation levels), or prevent or enhance polypeptide binding to DNA. Cutting can be performed by a variety of methods, such as enzymatic or chemical hydrolysis of phosphodiester bonds; it can be single-stranded or double-stranded; DNA cutting can result in blunt ends or staggered ends.

[0059] The terms “hybridizable” or “complementary” or “substantially complementary” refer to nucleic acids (such as RNA) containing nucleotide sequences that enable them to non-covalently bind to another nucleic acid in a sequence-specific, antiparallel manner under appropriate in vitro and / or in vivo temperature and solution ionic strength conditions, i.e., to form Watson-Crick base pairs and / or G / U base pairs, “annealing” or “hybridizing”.

[0060] It is understood in the art that the sequence of a polynucleotide does not need to be 100% complementary to the sequence of its target nucleic acid, which it can specifically hybridize to. The polynucleotide may hybridize on one or more segments. The polynucleotide may contain at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence complementarity with the target region within the target nucleic acid sequence it is targeting. In the art, the percentage complementarity between specific nucleic acid sequence segments can be routinely determined using known BLAST and PowerBLAST programs.

[0061] The terms “peptide,” “polypeptide,” and “protein” are used interchangeably in this invention and refer to a polymeric form of amino acids of any length, which may include encoded and non-coded amino acids, chemically or biochemically modified or derived amino acids, and polypeptides having a modified peptide backbone.

[0062] The term "binding domain" refers to a protein domain capable of non-covalently binding to another molecule. Binding domains can bind to, for example, DNA molecules (DNA-binding proteins), RNA molecules (RNA-binding proteins), and / or protein molecules (protein-binding proteins). In the case of a protein domain-binding protein, it can bind to itself (to form homodimers, homotrimers, etc.) and / or it can bind to one or more molecules of one or more different proteins.

[0063] The term "conserved amino acid substitution" refers to the interchangeability of amino acid residues with similar side chains in a protein. Exemplary sets of conserved amino acid substitutions are: valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine.

[0064] The term "DNA sequence encoding" a specific RNA is a DNA nucleic acid sequence transcribed into RNA. DNA polynucleotides can encode RNA (mRNA) that is translated into protein, or they can encode RNA that is not translated into protein (e.g., tRNA, rRNA, or gRNA; also known as "non-coding" RNA or "ncRNA"). A "protein-coding sequence," or sequence that encodes a specific protein or polypeptide, is a nucleic acid sequence that, under the control of appropriate regulatory sequences, is transcribed into mRNA (in the case of DNA) and translated (in the case of mRNA) into a polypeptide in vivo or in vitro.

[0065] The term "vector" or "expression vector" refers to a replicon, such as a plasmid, bacteriophage, virus, or granule, to which another segment of DNA, or "insertion fragment," can be attached in order to enable the attached segment to replicate within the cell.

[0066] The term "expression cassette" contains a DNA coding sequence operatively linked to a promoter. "Operably linked" means operatively linked, with the components in a relationship that allows them to function in their intended manner. The terms "recombinant expression vector" or "DNA construct" are used interchangeably in this invention to refer to a DNA molecule comprising a vector and at least one insert fragment. Recombinant expression vectors are typically produced for the purpose of expressing and / or amplifying the insert fragment or for constructing other recombinant nucleotide sequences.

[0067] When exogenous DNA, such as a recombinant expression vector, has been introduced into a cell, the cell has been genetically modified, transformed, or transfected by that DNA. The presence of exogenous DNA leads to permanent or transient genetic alterations. The transformed DNA may or may not integrate into the cell's genome.

[0068] The term "target DNA" refers to a DNA polynucleotide containing a "target site" or "target sequence." The terms "target site," "target sequence," "target protospacer DNA," or "protospacer-like sequence" are used interchangeably in this invention to refer to a nucleic acid sequence present in the target DNA to which a DNA-targeting segment of gRNA will bind if sufficient conditions for binding are present. The RNA molecule contains a sequence that binds, hybridizes with, or is complementary to the target sequence within the target DNA, thereby targeting the bound polypeptide to a specific location (target sequence) within the target DNA. "Cleavage" refers to the breakage of the covalent backbone of the DNA molecule.

[0069] The terms "nuclease" and "endonuclease" are used interchangeably to refer to enzymes that have catalytic activity for the degradation of endonucleases for the cleavage of polynucleotides. The "cleavage domain," "active domain," or "nuclease domain" of a nuclease refers to a polypeptide sequence or domain within the nuclease that has catalytic activity for DNA cleavage. The cleavage domain may be contained within a single polypeptide chain, or the cleavage activity may arise from the association of two or more polypeptides.

[0070] The term "targeting peptide" or "RNA-binding site-guided peptide" refers to a peptide that binds to RNA and targets a specific DNA sequence.

[0071] The term “guide sequence” or DNA-targeting segment (or “DNA-targeting sequence”) comprises a nucleotide sequence (complementary strand of the target DNA) that is complementary to a specific sequence within the target DNA, referred to in this invention as a “protospacer-like” sequence.

[0072] The term "recombination" refers to the process of exchanging genetic information between two polynucleotides. As used in this invention, "homology-guided repair (HDR)" refers to a specialized form of DNA repair that occurs, for example, during double-strand break repair in cells. This process requires nucleotide sequence homology, uses a "donor" molecule to provide a template for the repair of a "target" molecule (i.e., the molecule undergoing a double-strand break), and results in the transfer of genetic information from the donor to the target. If the donor polynucleotide differs from the target molecule, and part or all of the donor polynucleotide's sequence is incorporated into the target DNA, homology-guided repair can lead to alterations in the target molecule's sequence (e.g., insertions, deletions, mutations).

[0073] The term "non-homologous end joining (NHEJ)" refers to the repair of double-strand breaks in DNA by directly joining the broken ends together without the need for a homologous template. NHEJ often results in the deletion of nucleotide sequences near the double-strand break site.

[0074] The term "treatment" includes preventing the occurrence of disease or symptoms; suppressing disease or symptoms; or alleviating disease.

[0075] The terms “individual,” “subject,” “host,” and “patient” are used interchangeably in this invention and refer to any mammalian subject, particularly a human, to whom a diagnosis, treatment, or therapy is desired.

[0076] This invention provides a very small C2C9 nuclease suitable for CRISPR editing systems; comprising: i) an RNA-binding portion that interacts with gRNA; and an active portion.

[0077] The C2C9 nuclease possesses at least one of the following activities: regulating transcription within a target gene (e.g., target DNA), cleavage activity (ribonuclease and / or endonuclease activity), gene editing activity, etc., and the types of gene editing achieved include, but are not limited to: gene cleavage, gene deletion, gene insertion, point mutation, transcriptional repression, transcriptional activation, and base editing. The C2C9 nuclease can be derived from any biological species.

[0078] By way of non-limiting example, the C2C9 nuclease can be used as a nuclease, for example, in a CRISPR editing system.

[0079] In this invention, C2C9 nuclease variants can also be formed through modification, mutation, DNA shuffling, etc., so that the C2C9 nuclease variants have improved desired characteristics, such as function, activity, kinetics, half-life, etc. The modification may be, for example, the deletion, insertion, or substitution of amino acids, or, for example, the replacement of the "cleavage domain" of the C2C9 nuclease with a homologous or heterologous cleavage domain from different nucleases (e.g., the HNH domain of CRISPR-associated nucleases); the DNA targeting of the C2C9 nuclease can be altered, for example, by any modification method known in the art for DNA binding and / or DNA-modifying proteins, such as methylation, demethylation, acetylation, etc. The DNA shuffling refers to exchanging sequence fragments between DNA sequences of C2C9 nucleases from different sources to produce a chimeric DNA sequence encoding a synthetic protein with RNA-directed endonuclease activity. The modification, mutation, DNA shuffling, etc., can be used alone or in combination.

[0080] The C2C9 nuclease comprises, from the N-terminus to the C-terminus, the following:

[0081] Domain 1, comprising an amino acid sequence as shown in SEQ ID NO.1 or a variant having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, or at least 90% sequence identity with the amino acid sequence;

[0082] Domain 2, which contains an amino acid sequence as shown in SEQ ID NO.2 or a variant having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, or at least 90% sequence identity with the amino acid sequence.

[0083] The C2C9 nuclease variant can be:

[0084] (I) A variant that has at least 20% (e.g., 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%) sequence homology with the amino acid sequence of wild-type C2C9 nuclease and retains the activity of wild-type C2C9 nuclease;

[0085] (II) Based on the C2C9 nuclease or a variant thereof of (I), which further includes other components, such as nuclear localization signal fragments, so that the constructed C2C9 CRISPR system has appropriate activity in cell-free reaction, prokaryotic cell or eukaryotic cell environment;

[0086] (III) Encoding wild-type C2C9 nuclease or codon-optimized variants of the corresponding polynucleotide sequences of variants according to (I) and (II);

[0087] (IV) Based on a variant of the wild-type C2C9 nuclease and any one of (I) through (III), it further comprises:

[0088] (a) One or more modifications or mutations that produce C2C9 with significantly reduced or undetectable nuclease activity; and

[0089] (b) Peptides or domains with other functional activities.

[0090] The appropriate polypeptide in (IV)(b) is linked to the C-terminal domain of C2C9 nuclease and its variants.

[0091] The C2C9 nuclease variants under (IV) can convert any base pair into any other possible base pair through modification or mutation, rather than introducing double-strand breaks into the target DNA sequence.

[0092] The C2C9 nuclease variants under (IV) can be fusions or chimeric polypeptides formed by wild-type C2C9 nucleases or variants of (I) to (III) with heterologous sequences. These heterologous sequences include, but are not limited to, light-induced transcriptional regulators, small molecule / drug response transcriptional regulators, transcription factors, and transcriptional repressor proteins. When forming fusions or chimeric polypeptides, wild-type C2C9 nucleases or variants of (I) to (III) can be complete, partially, or completely defective C2C9 nucleases. For example, a C2C9 nuclease containing a catalytically active endonuclease domain is fused with a Fokl domain to form a chimeric protein; or a C2C9 nuclease that has lost its endonuclease domain activity through modification is fused with a Fokl domain to form a chimeric protein; or, exogenous genetic modifiers, tags, imaging agents, transcriptional regulators, histones, other components regulating gene structure or activity are fused with a C2C9 nuclease to prepare a chimeric protein. In some embodiments, the heterologous sequence can provide a tag for easy tracking or purification (e.g., a fluorescent protein such as green fluorescent protein (GFP), YFP, RFP, CFP, etc.; a His tag; a hemagglutinin (HA) tag; a FLAG tag; a Myc tag, etc.). In some embodiments, the heterologous sequence can provide subcellular localization of the C2C9 nuclease.

[0093] In some embodiments, the C2C9 nuclease variant is a modified form of the C2C9 nuclease. In some embodiments, the modified form of the C2C9 nuclease comprises amino acid changes that reduce the nuclease activity of the naturally occurring C2C9 nuclease. For example, in some embodiments, the modified form of the C2C9 nuclease has less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the corresponding wild-type C2C9 nuclease activity. In some embodiments, the modified form of the C2C9 nuclease does not have significant nuclease activity but retains the ability to interact with gRNA. In some embodiments, the modified form of the C2C9 nuclease has no nuclease activity. In some embodiments, the C2C9 nuclease has a mutant polypeptide with altered or eliminated DNA endonuclease activity, without a substantial reduction or enhancement of its endonuclease activity or DNA binding affinity.

[0094] In some embodiments, the C2C9 can be used in combination with other enzyme components or other ingredients to further develop various potential applications of the C2C9 nuclease. As non-limiting examples of C2C9 nuclease variants under (IV), for example, developing a C2C9 nuclease-based single-base editing system by fusing inactivated C2C9 with a base deaminase; developing a C2C9 nuclease-based Prime editing system by fusing inactivated C2C9 with a reverse transcriptase; developing a C2C9 nuclease-based transcriptional activation system by fusing inactivated C2C9 with a transcription activator; developing a C2C9 nuclease-based epigenetic modification system by fusing inactivated C2C9 with a nucleic acid epigenetic modification enzyme; and developing a C2C9 nuclease-based transcriptional repression system using inactivated C2C9.

[0095] C2C9 nuclease variants may possess the following specific properties, including but not limited to:

[0096] It has enhanced or reduced ability to bind to the target site, or retains the ability to bind to the target site;

[0097] It has enhanced or reduced ribonuclease and / or nuclease activity, or retains ribonuclease and / or nuclease activity;

[0098] It has deaminase activity, which can act on cytosine, guanine or adenine bases, and then replicate through the deamination site and repair in the cell to produce guanine, thymine and guanine respectively.

[0099] It has the activity of regulating the transcription of target DNA, which can either increase or decrease the transcription of target DNA at specific locations in the target DNA;

[0100] It has altered DNA targeting;

[0101] To increase, decrease, or maintain stability;

[0102] It can cleave the complementary strand of the target DNA, but has a reduced ability to cleave the non-complementary strand of the target DNA;

[0103] It can cleave the non-complementary strand of the target DNA, but has a reduced ability to cleave the complementary strand of the target DNA;

[0104] It has the ability to reduce the cutting of both the complementary and non-complementary strands of the target DNA;

[0105] It has enzymatic activity that modifies DNA-related polypeptides (such as histones). The enzymatic activity can be one or more of the following: methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitin activity, ribosylation activity, etc. (These enzymatic activities catalyze covalent modification of proteins; for example, C2C9 nuclease variants modify histones through methylation, acetylation, ubiquitination, phosphorylation, etc., to induce structural changes in histone-related DNA, thereby controlling the structure and properties of DNA).

[0106] In some embodiments, the C2C9 nuclease variant has no cleavage activity. In some embodiments, the C2C9 nuclease variant has single-strand cleavage activity. In some embodiments, the C2C9 nuclease variant has double-strand cleavage activity.

[0107] Enhanced activity or ability refers to an increase of at least 1%, 5%, 10%, 20%, 30%, 40%, or 50% in activity or ability relative to the wild-type C2C9 nuclease.

[0108] Reduced activity and capacity refer to activities or capacities that are less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% relative to wild-type C2C9 nucleases.

[0109] In this invention, the C2C9 nuclease is preferably derived from the group consisting of SEQ ID NO.3-117, etc.

[0110] In some implementations, C2C9 nucleases from different species are advantageous for providing different applications in order to take advantage of the various enzymatic properties of different C2C9 nucleases (e.g., for different PAM sequence preferences), such as increasing or decreasing enzymatic activity; increasing or decreasing cytotoxicity levels; and altering the balance between NHEJ, homology-directed repair, single-strand breaks, double-strand breaks, etc.

[0111] In some embodiments, the C2C9 nuclease is derived from *Actinomadura craniellae*, comprising the amino acid sequence of SEQ ID NO:3 or a variant thereof having at least about 90% (such as at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) sequence homology with SEQ ID NO:3; *Actinomadura craniellae* C2C9 or “AcC2C9” recognizes the PAM sequence AAN. In some embodiments, the C2C9 nuclease is derived from *Corynebacterium glutamicum* (SEQ ID NO:5), and *Corynebacterium glutamicum* C2C9 or “CgC2C9” recognizes the PAM sequence AAN. In some embodiments, the C2C9 nuclease is derived from *Rothia dentocariosa* (SEQ ID NO:20), and *Rothia dentocariosa* C2C9 or “RdC2C9” recognizes the PAM sequence AAG. In some implementations, the C2C9 is from Micrococcus luteus (SEQ ID NO:12), and Micrococcus luteus C2C9 or “MiC2C9” recognizes the PAM sequence AAN. The AAN includes AAA, AAC, AAG, and AAT.

[0112] C2C9 nucleases from different species or different C2C9 nucleases from the same species may require different PAM sequences in the target DNA. Therefore, the PAM sequence requirement for a particular C2C9 nuclease may be different from the aforementioned PAM sequence.

[0113] In some embodiments, these small C2C9s can be used in any of the systems, compositions, kits, and methods described below in this invention.

[0114] In a preferred embodiment, the C2C9 nuclease is derived from Actinomadura craniellae and is AcC2C9, the sequence of which is shown in SEQ ID NO.3 in the sequence listing.

[0115] In some embodiments, the AcC2C9 nuclease variant is:

[0116] (I) A variant having at least 20% (e.g., 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 99%) sequence homology with the amino acid sequence of the wild-type AcC2C9 nuclease of SEQ ID NO.3, and having increased, decreased, or retained activity relative to the wild-type C2C9 nuclease;

[0117] (II) Based on the AcC2C9 nuclease or a variant thereof of (I), which further includes other components, such as nuclear localization signals, to enable the constructed C2C9 CRISPR system to have appropriate activity in cell-free reaction, prokaryotic cell, or eukaryotic cell environments;

[0118] (III) A codon-optimized variant encoding the wild-type AcC2C9 nuclease SEQ ID NO.3 or the corresponding polynucleotide sequence of the variants according to (I) and (II);

[0119] (IV) A variant of wild-type AcC2C9 nuclease according to SEQ ID NO.3 and any one of (I) to (III), further comprising:

[0120] (a) One or more modifications or mutations that produce C2C9 with significantly reduced or undetectable nuclease activity; and

[0121] (b) Peptides with other functional activities.

[0122] The C2C9 nuclease provided by this invention has a small number of amino acids. In a preferred embodiment, when the amino acid sequence of the C2C9 nuclease is as shown in SEQ ID NO:3, it interacts with the PAM sequence AAN. The cleavage site is typically located within one to three base pairs upstream of the PAM sequence. C2C9 nucleases from different species or different C2C9 nucleases from the same species may require different PAM sequences in the target DNA; therefore, the PAM sequence requirement for a specific C2C9 nuclease may differ from the aforementioned PAM sequence. C2C9 can be engineered to target the PAM sequence "AAN" or other suitable targeting PAM sequences.

[0123] In some embodiments, the AcC2C9 nuclease variant has at least 90% (such as at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%) sequence homology with the amino acid sequence shown in SEQ ID NO:3 and retains the activity of the AcC2C9 nuclease.

[0124] Unless otherwise stated, the terms "C2C9" and "C2C9 nuclease" include wild-type C2C9 nuclease and all its variants. Those skilled in the art can determine the type of C2C9 nuclease variant by conventional means, without being limited to those exemplified above.

[0125] This invention also provides a nucleic acid comprising a nucleotide encoding an AcC2C9 nuclease. In one embodiment, this invention provides a codon-optimized polynucleotide sequence encoding one or more functional C2C9 domains, or encoding a polypeptide having the same function as a polypeptide encoded by the original natural nucleotide sequence. Therefore, those skilled in the art can codon-optimize nucleotide sequences or recombinant nucleic acids for use with the C2C9 nuclease described in this invention in a specific target species. Codon optimization can be performed using other methods known in the art, having less than 100% identity with the natural nucleotide sequence SEQ ID NO:121 (e.g., less than 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%). The polynucleotides of this invention are codon-optimized to increase the expression of the encoded C2C9 nuclease in target cells. In some embodiments, the polynucleotides of this invention are codon-optimized to increase expression in human cells. In some embodiments, the polynucleotides of the present invention are codon-optimized to increase expression in *E. coli* cells. In some embodiments, the polynucleotides of the present invention are codon-optimized to increase expression in fungal cells. In some embodiments, the polynucleotides of the present invention are codon-optimized to increase expression in insect cells.

[0126] In one embodiment, the nucleic acid encoding the AcC2C9 nuclease is shown in SEQ ID NO. 121 of the sequence listing. In one embodiment, the present invention provides a codon-optimized polynucleotide sequence of the AcC2C9 nuclease having at least 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.2%, 99.5%, 99.8%, 99.9%, or 100% sequence homology with SEQ ID NO: 121. In a preferred embodiment, a better coding sequence of the AcC2C9 nuclease optimized for human codons is shown in SEQ ID NO. 122 of the sequence listing, which encodes one or more functional C2C9 domains, or encodes a polypeptide having the same function as the polypeptide encoded by the original natural nucleotide sequence. In some embodiments, the polynucleotide encoding the C2C9 nuclease has at least about 90% (such as at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) sequence homology with SEQ ID NO:121-122.

[0127] In some embodiments, the nucleic acid encoding the C2C9 nuclease is DNA. In some embodiments, the nucleic acid encoding the C2C9 nuclease is RNA. In some embodiments, the nucleic acid encoding the C2C9 nuclease is an expression vector, such as a recombinant expression vector. Any suitable expression vector can be used, as long as it is compatible with the host cell, including but not limited to, viral vectors (e.g., vaccinia virus-based viral vectors; poliovirus; adenovirus; adeno-associated virus; SV40; herpes simplex virus; human immunodeficiency virus); retroviral vectors (e.g., murine leukemia virus, spleen necrosis virus, and vectors derived from retroviruses, such as Rous sarcoma virus, Harvey sarcoma virus, avian leukemia virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus), etc.

[0128] In some embodiments, the nucleotide sequence encoding the C2C9 nuclease is operatively linked to a control element, such as a transcriptional control element, or a promoter. In some embodiments, the nucleotide sequence encoding the C2C9 nuclease is operatively linked to an inducible promoter. In some embodiments, the nucleotide sequence encoding the C2C9 nuclease is operatively linked to a constitutive promoter. The transcriptional control element can function in eukaryotic cells, such as mammalian cells, or prokaryotic cells (such as bacterial or archaea cells). In some embodiments, the nucleotide sequence encoding the C2C9 nuclease is operatively linked to multiple control elements that allow expression of the nucleotide sequence encoding the C2C9 nuclease in both prokaryotic and eukaryotic cells. In some embodiments, the polynucleotide sequence encoding the C2C9 nuclease is operatively linked to a suitable nuclear localization signal for expression in a cellular or in vitro environment.

[0129] In this invention, the polynucleotide encoding the C2C9 nuclease can be synthesized artificially, for example, by chemical methods, thereby allowing for easy modification in various ways. The modifications can be any method known in the art. In some embodiments, the polynucleotide encoding the C2C9 nuclease contains one or more modifications, thereby allowing for the easy incorporation of many modifications, such as enhancing transcriptional activity, altering enzyme activity, improving its translational or stability (e.g., increasing its resistance to proteolysis or degradation) or specificity, altering solubility, altering delivery, or reducing the innate immune response in host cells. The modifications can be any method known in the art. In some embodiments, the DNA or RNA encoding the C2C9 nuclease is introduced into the cell through modification to edit any one or more genomic loci. In some embodiments, the nucleic acid sequence encoding the C2C9 nuclease is a modified nucleic acid, such as codon-optimized. The modifications can be single modifications or combinations of modifications.

[0130] In this invention, the nucleic acid containing the polynucleotide encoding the C2C9 nuclease can be a nucleic acid mimic. For example, a peptide nucleic acid, a polynucleotide mimic with excellent hybridization properties.

[0131] In this invention, the C2C9 nuclease, or the polynucleotide encoding the C2C9 nuclease, is applicable to any organism or in vitro environment, including but not limited to bacteria, archaea, fungi, protozoa, plants, or animals. Accordingly, applicable target cells include, but are not limited to, eukaryotic and prokaryotic cells, such as bacterial cells, archaea cells, fungal cells, protozoan cells, plant cells, or animal cells; the eukaryotic cells include mammalian cells and plant cells, and the prokaryotic cells include *Escherichia coli* and *Klebsiella pneumoniae*. Applicable target cells can be any type of cell, including stem cells, somatic cells, etc. The cells can be in vivo or in vitro. In some embodiments, the C2C9 nuclease, or the nucleic acid encoding the C2C9 nuclease, is formulated in liposomes or lipid nanoparticles.

[0132] This invention also provides the application of the C2C9 nuclease in constructing a CRISPR / C2C9 gene editing system.

[0133] The present invention also provides suitable PAM sequences for use in vivo, in vitro cells or in an in vitro environment, wherein the PAM is AAN and / or GAN.

[0134] This invention also provides suitable PAM sequences and suitable guides for use in prokaryotic, eukaryotic, and in vitro environments. The guide RNA (gRNA) provided by this invention directs nucleases such as C2C9 nuclease to specific target sequences within target genes.

[0135] In some embodiments, the gRNA comprises:

[0136] The first segment of a nucleotide sequence complementary to the target sequence in the target gene (also called a "gene-targeting sequence" or "gene-targeting fragment"); and,

[0137] The second fragment that interacts with the C2C9 nuclease (also known as the "protein-binding sequence" or "protein-binding fragment").

[0138] In some embodiments, the gRNA comprises an array of repeating sequence spacers, wherein the spacers contain nucleic acid sequences complementary to target sequences in the gene.

[0139] In some embodiments, the gRNA comprises:

[0140] i. Gene targeting regions (such as DNA-targeting regions) capable of hybridizing with target sequences.

[0141] ii.tracr pairing sequence, and

[0142] iii. tracr RNA sequence;

[0143] The guide RNA is a single strand, which is formed by sequentially linking the DNA-targeting segment (i) with the tracr pair sequence (ii) and the tracr RNA sequence (iii); or, the guide RNA comprises two strands, one of which is formed by linking the gene-targeting segment (i) with the tracr pair sequence (ii), and the other strand is the tracr RNA sequence (iii).

[0144] The gene-targeting region (i) is preferably the RNA sequence corresponding to a nucleic acid fragment of 20 bp in length following the PAM sequence.

[0145] The tracr pair sequence (ii) hybridizes with the tracrRNA sequence (iii) to form a stem-loop structure.

[0146] The tracr pair sequence (ii) and the tracrRNA sequence (iii) can be linked together to form a single guide RNA backbone sequence.

[0147] The RNA sequence (crRNA) obtained by ligating the gene-targeting region (i) hybridized with the target sequence and the tracr pair sequence (ii), together with the tracrRNA sequence (iii), can mediate C2C9 endonuclease activity when present simultaneously as two separate RNA sequences. Alternatively, a complete guide RNA expression construct targeting the target sequence, obtained by ligating the guide RNA backbone sequence and the DNA-targeting region (i) hybridized with the target sequence, can also mediate C2C9 endonuclease activity.

[0148] The gene-targeting region of a gRNA contains a nucleotide sequence complementary to a sequence in the target gene. This gene-targeting sequence interacts with the target gene in a sequence-specific manner through hybridization (i.e., base pairing). The gene-targeting sequence of the gRNA can be modified, for example, through genetic engineering, so that the gRNA hybridizes with any desired sequence within the target gene. The gRNA then guides the bound polypeptide to a specific nucleotide sequence within the target gene via the aforementioned gene-targeting sequence.

[0149] The stem-loop structure forms a protein-binding structure that interacts with the C2C9 nuclease. In some embodiments, the protein-binding structure of the gRNA comprises four stem-loop structures, and the tracr pairing sequence (ii) typically pairs with the tracr RNA sequence (iii) through base complementarity. Modification of the gRNA's base sequence can enhance the activity of the C2C9 nuclease or reduce nonspecific recognition.

[0150] In some implementations, the target gene is a DNA sequence. In other implementations, the target gene is an RNA sequence.

[0151] In some embodiments, the gRNA is sgRNA. In a preferred embodiment, the sequence of the gRNA is shown in SEQ ID NO. 123.

[0152] In some implementations, the target sequence of the target gene can have a length of 12-40 nucleotides, for example, it can be 13-20, 18-25, 22-32, 26-37, 30-38, or 32-40 nucleotides. The percentage of complementarity between the target region (i) of the guide RNA and the target sequence of the target gene can be at least 50% (e.g., at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%).

[0153] In some embodiments, the gRNA also includes a transcription terminator.

[0154] This invention also provides modified gRNAs that can be modified to achieve hybridization with any desired sequence within a target gene; or, by modifying the gRNA to alter its properties, such as enhancing its stability, including but not limited to increasing its resistance to degradation by ribonucleases (RNases) present in the cell, thereby prolonging its half-life in the cell; or, by modifying it to enhance the formation or stability of a CRISPR-C2C9 genome editing complex comprising gRNA and a nuclease (e.g., C2C9 nuclease); or, by modifying it to enhance the specificity of the genome editing complex; or, by modifying it to enhance the initiation site, stability, or kinetics of the interaction between the genome editing complex and the target sequence in the genome; or, by modifying it to reduce the likelihood or extent to which RNA introduced into the cell triggers an innate immune response, etc. In this invention, various properties of the CRISPR-C2C9 system (described below) can be altered by modifying the gRNA, such as enhancing the formation, target activity, specificity, stability, or kinetics of the CRISPR-C2C9 genome editing complex. RNA can be modified using methods known in the art, including but not limited to 2'-fluorine or 2'-amino modifications on the ribose or base residues of pyrimidine or the reverse base at the 3' end of the RNA. In this invention, gRNA can be modified using any one modification or a combination of modifications. In some embodiments, sgRNA introduced into the cell is modified to edit any one or more genomic loci.

[0155] This invention provides nucleic acids comprising a nucleotide sequence encoding gRNA. In some embodiments, the nucleic acid encoding gRNA is an expression vector, such as a recombinant expression vector. Any suitable expression vector can be used, provided it is compatible with the host cell, including but not limited to, viral vectors (e.g., vaccinia virus-based viral vectors; poliovirus; adenovirus; adeno-associated virus; SV40; herpes simplex virus; human immunodeficiency virus); retroviral vectors (e.g., murine leukemia virus, spleen necrosis virus, and vectors derived from retroviruses, such as Rous sarcoma virus, Harvey sarcoma virus, avian leukemia virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus), etc.

[0156] In some implementations, multiple gRNAs are used simultaneously in the same cell to simultaneously regulate transcription at different locations on the same or different target genes. When multiple gRNAs are used simultaneously, they can be present on the same expression vector or on different vectors, and can be expressed simultaneously; when present on the same vector, they can be expressed under the same control elements.

[0157] In some embodiments, the nucleotide sequence encoding gRNA is operatively linked to a control element, such as a transcriptional control element, like a promoter. In some embodiments, the nucleotide sequence encoding gRNA is operatively linked to an inducible promoter. In some embodiments, the nucleotide sequence encoding gRNA is operatively linked to a constitutive promoter. The transcriptional control element can function in eukaryotic cells, such as mammalian cells, or prokaryotic cells (such as bacterial or archaea cells). In some embodiments, the nucleotide sequence encoding gRNA is operatively linked to multiple control elements, allowing expression of the nucleotide sequence encoding gRNA in both prokaryotic and eukaryotic cells.

[0158] In this invention, the gRNA can be synthesized artificially, for example, by chemical methods, thereby allowing for easy modification in various ways. The modification can be any method known in the art, such as using a polyA tail, adding a 5' cap analogue, a 5' or 3' untranslated region (UTR), including a thiophosphorylated 2'-O-methyl nucleotide at the 5' or 3' end, or treating with a phosphatase to remove the 5' terminal phosphate ester, etc.

[0159] In some embodiments, the nucleotide sequence encoding the gRNA contains one or more modifications that can be used, for example, to enhance activity, stability, or specificity, alter delivery, reduce the innate immune response in host cells, or for other enhancements.

[0160] In some embodiments, one or more target moieties or conjugates that enhance the activity, cellular distribution, or cellular uptake of the nucleotide sequence encoding the gRNA are chemically linked to the gRNA. The target moieties or conjugates may include conjugate groups covalently bound to a functional group; conjugate groups include reporter molecules, polyamines, and polyethylene glycol. In some embodiments, groups that enhance pharmacodynamic properties are linked to the gRNA; these groups include groups that improve uptake, enhance resistance to degradation, and / or enhance sequence-specific hybridization with the target nucleic acid.

[0161] In this invention, the nucleic acid containing a polynucleotide encoding gRNA can be a nucleic acid mimic. For example, a peptide nucleic acid, a polynucleotide mimic with excellent hybridization properties.

[0162] In this invention, the gRNA, or the polynucleotide encoding the gRNA, is applicable to any organism or in vitro environment, including but not limited to bacteria, archaea, fungi, protozoa, plants, or animals. Correspondingly, applicable target cells include, but are not limited to, bacterial cells, archaea cells, fungal cells, protozoan cells, plant cells, or animal cells. Applicable target cells can be any type of cell, including stem cells, somatic cells, etc.

[0163] The present invention also provides a recombinant expression vector comprising (i) a nucleotide sequence encoding a gRNA and (ii) a nucleotide sequence encoding a C2C9 nuclease. In one embodiment, the site of transcription regulated within the target gene is determined by the gRNA.

[0164] The present invention also provides a CRISPR-C2C9 system based on a C2C9 nuclease, comprising: (a) a C2C9 nuclease as described above or a nucleic acid comprising a polynucleotide encoding the C2C9 nuclease as described above; and / or, (b) one or more gRNAs as described above or a nucleic acid encoding the gRNA. The gRNA is capable of directing the C2C9 nuclease to a target sequence.

[0165] In some implementations, the C2C9 nuclease is provided directly as a protein; for example, it can be provided by transforming exogenous proteins into protoplasts and / or transforming fungi with nucleic acids. The C2C9 nuclease can be introduced into cells by any suitable method, such as injection.

[0166] In the system described in this invention, the C2C9 nuclease and gRNA can form a complex in the host cell to recognize the PAM sequence on the target gene (e.g., target DNA) sequence; the target sequence of the CRISPR / C2C9 gene editing system is a nucleic acid fragment (e.g., DNA fragment) 20 bp in length following the PAM sequence. In one embodiment, the complex can selectively regulate the transcription of the target DNA in the host cell. The CRISPR / C2C9 gene editing system can cleave the double strand of the target DNA, causing DNA breaks. The major cleavage site on the target strand (the single strand of DNA complementary to the gRNA, also known as the Targeting strand or TS strand) is 21-22 nt downstream of the 3' end of the PAM site, while the major cleavage site on the non-target strand (the single strand of DNA not complementary to the gRNA, also known as the Non-targeting strand or NTS strand) is 15-16 nt downstream of the 3' end of the PAM site.

[0167] In one embodiment, the system comprises a gRNA and a C2C9 nuclease complex.

[0168] In one embodiment, the system comprises a recombinant expression vector. In another embodiment, the system comprises a recombinant expression vector comprising (i) a nucleotide sequence encoding a gRNA, wherein the gRNA comprises: (a) a first segment comprising a nucleotide sequence complementary to a sequence in the target DNA; and (b) a second segment interacting with a C2C9 nuclease; and (ii) a nucleotide sequence encoding a C2C9 nuclease, wherein the C2C9 nuclease comprises: (a) an RNA-binding portion interacting with the gRNA; and (b) an active portion regulating transcription within the target DNA, wherein the site of transcription regulated within the target DNA is determined by the gRNA.

[0169] In some embodiments, the system comprises (or is composed of) a C2C9 nuclease as shown in SEQ ID NO:3. In some embodiments, the system comprises a nucleic acid encoding a C2C9 nuclease having at least about 90% (e.g., at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%) sequence homology to SEQ ID NO:121 or SEQ ID NO:122. In some embodiments, the system comprises a nucleic acid encoding a C2C9 nuclease, the nucleic acid comprising (or is composed of) the sequence of SEQ ID NO:121 or SEQ ID NO:122. In some embodiments, the gRNA in the system is a single gRNA (sgRNA). In some embodiments, the system further comprises one or more additional gRNAs or nucleic acids encoding one or more additional gRNAs.

[0170] In some embodiments, the system contains a suitable promoter and / or a suitable nuclear localization signal for expression in a cellular or in vitro environment; and a polynucleotide sequence encoding the C2C9 nuclease is operatively linked to (i) a suitable promoter for expression in a cellular or in vitro environment; and / or (ii) a suitable nuclear localization signal.

[0171] In some embodiments, the system includes a polynucleotide sequence encoding the gRNA operatively linked to (i) a suitable promoter for expression in a cellular or in vitro environment; and / or (ii) a suitable nuclear localization signal.

[0172] In some embodiments, the present invention provides a system comprising: (a) a C2C9 nuclease comprising the amino acid sequence of SEQ ID NO:3 or a variant thereof having at least 90% sequence homology with SEQ ID NO:3; and (b) a gRNA, wherein the gRNA is capable of directing the C2C9 nuclease or a variant thereof to a target polynucleotide sequence.

[0173] In some embodiments, the present invention provides a system comprising: (a) a C2C9 nuclease or a nucleic acid encoding a C2C9 nuclease, said C2C9 nuclease comprising the amino acid sequence of SEQ ID NO:3 or a variant thereof having at least 90% sequence homology with SEQ ID NO:3; said nucleic acid encoding the C2C9 nuclease comprising a polynucleotide sequence of any one of SEQ ID NO:121 or SEQ ID NO:122 or a variant thereof having at least 90% sequence homology with the polynucleotide sequence of any one of SEQ ID NO:121 or SEQ ID NO:122; and (b) a gRNA or a nucleic acid encoding a gRNA, wherein said gRNA is capable of directing said C2C9 nuclease or a variant thereof to a target polynucleotide sequence. In some embodiments, said gRNA is sgRNA.

[0174] The system described in this invention may contain one gRNA or multiple gRNAs simultaneously. In one embodiment, the system contains multiple gRNAs to simultaneously modify the same target DNA or different sites on different target DNA. In one embodiment, two or more guide RNAs target the same gene, transcript, or locus. In one embodiment, two or more guide RNAs target different unrelated loci. In some embodiments, two or more guide RNAs target different but related loci.

[0175] The components of the system described in this invention can be delivered via carriers. For example, for polynucleotides, methods that can be used include, but are not limited to, nanoparticles, liposomes, ribonucleoproteins, small RNA conjugates, chimeras, and RNA-fusion protein complexes.

[0176] In some embodiments, the C2C9 or gRNA may be constructed on a vector, with the nucleotide sequence encoding C2C9 and / or gRNA operatively linked to a control element, such as a transcriptional control element, like a promoter. The transcriptional control element may function in eukaryotic cells (e.g., mammalian cells) or prokaryotic cells (e.g., bacterial or archaea cells). In some embodiments, the nucleotide sequence encoding gRNA and / or C2C9 nuclease is operatively linked to multiple control elements, allowing expression of nucleotide sequences encoding guide RNA and / or site-directed modification polypeptides in two prokaryotes. In some embodiments, the nucleotide sequences encoding gRNA and / or C2C9 nuclease are constructed on the same vector. In some embodiments, the nucleotide sequences encoding gRNA and / or C2C9 nuclease are constructed on different vectors. In some embodiments, the nucleotide sequence encoding gRNA and / or C2C9 nuclease is operatively linked to a constitutive promoter. In some embodiments, the nucleotide sequence encoding gRNA and / or C2C9 nuclease is operatively linked to an inducible promoter.

[0177] The system described in this invention may further include one or more donor templates. In some embodiments, the donor template contains a donor sequence for inserting a target gene. In some embodiments, the donor template contains a donor cassette with gRNA target sites on one or both sides. In some embodiments, the donor template has gRNA target sites on both sides of the donor cassette, wherein the gRNA target sites of the donor template are the reverse complementary sequences of the genomic gRNA target sites of the gRNA in this system. The donor sequence is typically different from the genomic sequence it replaces; that is, the donor sequence typically contains base changes, insertions, deletions, inversions, or rearrangements relative to the genomic sequence and has sufficient homology to support homology-guided repair. The donor sequence may be single-stranded DNA, single-stranded RNA, double-stranded DNA, or double-stranded RNA. In some embodiments, the donor sequence is a foreign sequence. In some embodiments, the donor sequence is a donor DNA sequence (including donor single-stranded or double-stranded DNA), wherein the target DNA sequence is edited via homology-guided repair. In some implementations, the donor template is physically linked to gRNA in the system. In some implementations, a complete, partially or completely defective C2C9 nuclease, or gRNA, is linked to the donor template to guide one or more specific target gene (e.g., target DNA) sites via one or more gRNA molecules, thereby promoting homologous recombination of exogenous DNA sequences.

[0178] The system described in this invention may further include a dimer FOK1 nuclease, a complete or partially or completely defective C2C9 nuclease, or gRNA linked to the dimer FOK1 nuclease to direct endonuclease cleavage when directed to one or more specific DNA target sites by one or more gRNA molecules.

[0179] The system described in this invention can edit or modify DNA at multiple locations in cells for use in gene therapy, including but not limited to gene therapy for diseases, biological research, and crop resistance improvement or yield enhancement.

[0180] In this invention, the system is applicable to any biological or in vitro environment, including but not limited to bacteria, archaea, fungi, protozoa, plants, or animals. Correspondingly, suitable target cells include, but are not limited to, bacterial cells, archaea cells, fungal cells, protozoan cells, plant cells, or animal cells. Suitable target cells can be any type of cell, including stem cells, somatic cells, etc. The cells can be in vivo or in vitro.

[0181] The present invention also provides a composition comprising one or more of the C2C9 nuclease or polynucleotide encoding it as described above, gRNA or polynucleotide encoding it, recombinant expression vector, and system, and may further include acceptable carriers, media, etc. The acceptable carriers and media include, for example, sterile water or physiological saline, stabilizers, excipients, antioxidants (ascorbic acid, etc.), buffers (phosphate, citric acid, other organic acids, etc.), preservatives, surfactants (PEG, Tween, etc.), chelating agents (EDTA, etc.), binders, etc. Furthermore, it may also contain other low molecular weight polypeptides; proteins such as serum albumin, gelatin, or immunoglobulins; amino acids such as glycine, glutamine, asparagine, arginine, and lysine; sugars or carbohydrates such as polysaccharides and monosaccharides; and sugar alcohols such as mannitol or sorbitol. When preparing aqueous solutions for injection, such as physiological saline, isotonic solutions containing glucose or other adjuvant drugs, such as D-sorbitol, D-mannose, D-mannitol, or sodium chloride, appropriate solubilizers such as alcohols (ethanol, etc.), polyols (propylene glycol, PEG, etc.), and nonionic surfactants (Tween 80, HCO-50) may be used. In some embodiments, the composition comprises gRNA and a buffer for stabilizing the nucleic acid.

[0182] The present invention also provides a kit comprising the system or composition described above. The kit may further comprise one or more additional reagents, such as those selected from: dilution buffers; wash buffers; control reagents, etc. In some embodiments, the kit comprises (a) a C2C9 nuclease or a nucleic acid encoding a C2C9 nuclease as described above; and (b) gRNA or a nucleic acid encoding the gRNA, wherein the gRNA is capable of directing the C2C9 nuclease or a variant thereof to a target polynucleotide sequence. In some embodiments, the kit further comprises a donor template containing a heterologous polynucleotide sequence, wherein the heterologous polynucleotide sequence is capable of being inserted into the target polynucleotide sequence.

[0183] This invention provides the C2C9 nuclease or its encoding polynucleotide, gRNA or its encoding polynucleotide, recombinant expression vector, system, composition and kit described above for any of the following uses in vivo, in vitro cells or cell-free systems, including but not limited to:

[0184] Cutting target genes;

[0185] Manipulating the expression of target genes;

[0186] Genetic modification target genes;

[0187] Genetic modification target gene-related peptides;

[0188] Used for intentional and controlled damage at any desired location on the target gene;

[0189] Used for intentional and controlled repair at any desired location on the target gene;

[0190] Modifying the target gene in ways other than introducing double-strand breaks (C2C9 nuclease has enzymatic activity, which modifies the target gene in ways other than introducing double-strand breaks; the enzymatic activity may be inherent in C2C9 itself, or obtained by, for example, fusing a heterologous polypeptide with enzymatic activity into C2C9 nuclease to form a chimeric C2C9 nuclease, the enzymatic activity including but not limited to methyltransferase activity, deamination activity, dismutase activity, alkylation activity, demethylase activity, DNA repair activity, transposase activity, recombinase activity, DNA damage activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, etc.).

[0191] In some embodiments, the target gene is target DNA. In some embodiments, the target gene is target RNA.

[0192] In this invention, the C2C9 nuclease can be used in combination with other enzyme components or other ingredients to further develop various potential applications of the C2C9 nuclease. As non-limiting examples, for instance, a single-base editing system based on the C2C9 nuclease can be developed by fusing inactivated C2C9 with a base deaminase; a Prime editing system based on the C2C9 nuclease can be developed by fusing inactivated C2C9 with a reverse transcriptase; a transcription activation system based on the C2C9 nuclease can be developed by fusing inactivated C2C9 with a transcription activator; an epigenetic modification system based on the C2C9 nuclease can be developed by fusing inactivated C2C9 with a nucleic acid epigenetic modification enzyme; and a transcriptional repression system based on the C2C9 nuclease can be developed using inactivated C2C9.

[0193] The C2C9 nuclease or its encoding polynucleotide, gRNA or its encoding polynucleotide, recombinant expression vector, system, composition, and kit of this invention can be applied in research, diagnostics, industry (e.g., microbial engineering), drug discovery (e.g., high-throughput screening), target identification, imaging, and therapeutic fields. In research applications, this can include, for example, determining the effects of enhancing the transcription of target nucleic acids on, for example, development, growth, metabolism (e.g., precisely controlling and regulating biosynthetic pathways by controlling the levels of specific enzymes), and gene expression.

[0194] In this invention, the C2C9 nuclease or its encoding polynucleotide, gRNA or its encoding polynucleotide, recombinant expression vector, system, composition, and kit are applicable to any organism, including but not limited to bacteria, archaea, fungi, protozoa, plants, or animals. Correspondingly, applicable target cells include, but are not limited to, bacterial cells, archaea cells, fungal cells, protozoan cells, plant cells, or animal cells (e.g., rodent cells, human cells, non-human primate cells). Applicable target cells can be any type of cell, including stem cells, somatic cells, etc.

[0195] When the C2C9 nuclease, gRNA, system, composition, and method described in this invention are applicable to eukaryotic cells, they can be used, for example, in mammalian cells. In some embodiments, the cells are not human fetal cells or are not derived from human fetal cells. In some embodiments, the cells are not human embryonic cells or are not derived from human embryonic cells. In some embodiments, the invention does not involve the destruction of a human fetus or human embryo. In some embodiments, the subject or individual is not a human fetus or embryo.

[0196] When the C2C9 nuclease or its encoding polynucleotide, gRNA or its encoding polynucleotide, recombinant expression vector, system, composition, and kit described in this invention are applicable to prokaryotic cells, such as *Escherichia coli*, they are used to shut down gene expression in bacterial cells. In some embodiments, they are used to shut down gene expression in bacterial cells. In some embodiments, nucleic acids encoding suitable gRNA and / or suitable C2C9 nuclease are introduced into the chromosome of a target cell, and the nucleic acids are translated and expressed under the control of an inducible promoter; wherein the encoded gRNA and C2C9 nuclease form a complex that cleaves the target DNA at a site of interest. Therefore, in some embodiments, through engineering, the bacterial genome contains a nucleic acid sequence encoding a suitable C2C9 nuclease and / or a suitable gRNA on a plasmid, and the expression of any targeted gene is controlled by inducing the expression of the gRNA and C2C9 nuclease under the control of an inducible promoter. In some embodiments, the C2C9 nuclease possesses enzymatic activity that modifies target DNA in ways other than introducing double-strand breaks. This enzymatic activity may be inherent in C2C9 itself or obtained, for example, by fusing a heterologous polypeptide with enzymatic activity into the C2C9 nuclease to form a chimeric C2C9 nuclease. The enzymatic activity includes, but is not limited to, methyltransferase activity, deamination activity, dismutase activity, alkylation activity, demethylase activity, DNA repair activity, transposase activity, recombinase activity, DNA damage activity, depurination activity, oxidation activity, and pyrimidine dimer formation activity. In this invention, genetic modification of the target gene at any location within the target gene can be controlled by genetically engineering the desired complementary nucleic acid sequence into the gene-targeting region (e.g., DNA-targeting region) of the gRNA. In this invention, genetic modification of the target nucleic acid at any location within the target gene can be controlled by genetically engineering the desired complementary nucleic acid sequence into C2C9. In some embodiments, it is applied to the intentional and controlled damage to DNA at any desired location within the bacterial target DNA. In some implementations, this is used for sequence-specific and controlled repair of DNA at any desired location within the bacterial target DNA.

[0197] The present invention also provides a method for targeting, editing, modifying or manipulating a target gene (such as a target DNA) in a cell or in vivo, in vitro, or in a cell-free system, comprising: introducing the C2C9 nuclease or the polynucleotide encoding it, gRNA or the polynucleotide encoding it, recombinant expression vector, system, composition, etc. as described above into a kit in vivo, in vitro, or in a cell-free system to target, edit, modify or manipulate the target gene.

[0198] In one embodiment, the method includes the following:

[0199] (a) Introducing the C2C9 nuclease or the nucleic acid encoding the C2C9 nuclease into vivo, in vitro cells, or in a cell-free system; and

[0200] (b) Introducing the gRNA (sgRNA) or a nucleic acid (e.g., DNA) suitable for in situ production of such sgRNA; and

[0201] (c) Contacting a cell or target gene with a C2C9 nuclease or a nucleic acid, gRNA (sgRNA) encoding a C2C9 nuclease or a nucleic acid suitable for producing such sgRNA in situ to create one or more cuts, nicks or edits in the target gene; wherein the C2C9 nuclease is directed to the target gene by its processed or unprocessed gRNA.

[0202] In some embodiments, the target gene is target DNA. In some embodiments, the target DNA may be naked DNA in vitro that is not bound to DNA-associated proteins. In some embodiments, the target DNA is chromosomal DNA in in vitro cells. In some embodiments, the target gene is target RNA. In some embodiments, the target DNA is contacted with a targeting complex comprising the C2C9 nuclease and gRNA, wherein the gRNA provides target specificity to the targeting complex by comprising a nucleotide sequence complementary to the target DNA; and the C2C9 nuclease provides site-specific activity. In some embodiments, the targeting complex modifies the target DNA, thereby causing, for example, DNA cleavage, DNA methylation, DNA damage, DNA repair, etc. In some embodiments, the targeting complex modifies a target DNA-associated polypeptide (e.g., histone, DNA-binding protein, etc.), thereby causing, for example, methylation of the target DNA-associated polypeptide-histone, histone acetylation, histone ubiquitination, etc.

[0203] In the method described in this invention, when applied in vivo, the C2C9 nuclease or the polynucleotide encoding it, gRNA or the polynucleotide encoding it, recombinant expression vector, system, composition and / or donor polynucleotide are directly applied to an individual.

[0204] In the method described in this invention, a C2C9 nuclease or a nucleic acid containing a nucleotide sequence encoding a polypeptide of the C2C9 nuclease can be introduced into cells using known methods. Similarly, gRNA or a nucleic acid containing a nucleotide sequence encoding gRNA can be introduced into cells using known methods. Known methods include DEAE-glucan-mediated transfection, liposome-mediated transfection, viral or bacteriophage infection, lipid transfection, transfection, conjugation, protoplast fusion, polyethyleneimine-mediated transfection, electroporation, calcium phosphate precipitation, gene gun, calcium phosphate precipitation, microinjection, and nanoparticle-mediated nucleic acid delivery. For example, plasmids are delivered via electroporation, calcium chloride transfection, microinjection, and lipid transfection. For viral vector delivery, cells are contacted with viral particles containing nucleic acids encoding gRNA and / or a C2C9 nuclease and / or a chimeric C2C9 nuclease and / or donor polynucleotides.

[0205] In some embodiments, in the method described in this invention, a nuclease cuts target DNA in the cell to produce double-strand breaks, which are then repaired by the cell in the following ways: non-homologous end joining (NHEJ) and homology-guided repair.

[0206] In non-homologous end joining, double-strand breaks are repaired by directly joining the broken ends together. During this process, a few base pairs may be inserted or deleted at the cleavage site. Therefore, the cleavage of DNA by the C2C9 nuclease can be used to remove nucleic acid material from a target DNA sequence by cutting the target DNA sequence and allowing the cell to repair the sequence in the absence of exogenously provided donor polynucleotides. Thus, this method can be used to knock out genes or knock genetic material into selected loci in the target DNA.

[0207] In homology-guided repair (NHEJ) of the target DNA sequence, a donor polynucleotide homologous to the cleaved target DNA sequence is used as a template for repairing the cleaved target DNA sequence, resulting in the transfer of genetic information from the donor polynucleotide to the target DNA. Therefore, new nucleic acid material can be inserted / replicated into the site. In some embodiments, the target DNA is contacted with the donor polynucleotide. In some embodiments, the donor polynucleotide is introduced into the cell. Modifications to the target DNA by NHEJ and / or homology-guided repair can lead to, for example, gene correction, gene substitution, transgene insertion, nucleotide deletion, nucleotide insertion, gene damage, gene mutation, sequence substitution, etc.

[0208] In some embodiments, the cells containing the target gene are in vitro. In some embodiments, the cells containing the target gene are in vivo. Suitable nucleic acids containing nucleotide sequences encoding gRNA and / or C2C9 nucleases include expression vectors, wherein the expression vector containing the nucleotide sequences encoding gRNA and / or C2C9 nucleases is a recombinant expression vector. In some embodiments, the method introduces C2C9 nuclease and one or more gRNAs (gRNAs) into cells via the same or different recombinant vectors, contacting the cells with a vector containing a polynucleotide encoding gRNA and / or C2C9 nucleases to allow the vector to be absorbed by the cells.

[0209] The method described above can also be used to edit or modify DNA at multiple locations within a cell.

[0210] The method described above can also be used to regulate the transcription of target DNA.

[0211] In some embodiments, the method provided by the present invention includes linking a complete or partially or completely defective C2C9 nuclease or gRNA portion to a dimerizing FOK1 nuclease to direct endonuclease cleavage upon guidance to one or more specific DNA target sites via one or more crRNA molecules. In other embodiments, the method provides that the C2C9 nuclease linked to the dimerizing FOK1 nuclease is introduced into a cell along with two or more gRNAs that are either RNA or DNA encoded by a single promoter, each gRNA containing an array of repeat spacer regions, wherein the spacer regions contain nucleic acid sequences complementary to a target sequence in the DNA, and the repeat sequences contain stem-loop structures. In some embodiments, the method provides that the present invention includes linking a complete or partially or completely defective C2C9 nuclease gRNA to a donor single-stranded or double-stranded DNA donor template to direct guidance to one or more specific DNA target sites via one or more gRNA molecules, thereby promoting homologous recombination of exogenous DNA sequences.

[0212] The present invention also provides a method for a target gene-related polypeptide, the method comprising contacting a target gene with a target complex comprising the C2C9 nuclease described above or a nucleic acid encoding a C2C9 nuclease, and contacting gRNA or a nucleic acid suitable for in situ production of such sgRNA.

[0213] In some embodiments, the target gene-associated polypeptide is a histone, DNA-binding protein, etc. In some embodiments, the targeting complex modifies the target DNA-associated polypeptide, thereby causing, for example, methylation, acetylation, or ubiquitination of the target DNA-associated polypeptide, such as histone.

[0214] In some embodiments, the target gene is target DNA. In some embodiments, the target DNA is chromosomal DNA in in vitro cells. In some embodiments, the target DNA is chromosomal DNA in in vivo cells, etc. In some embodiments, the target gene is target RNA. In some embodiments, the target DNA is contacted with a targeting complex containing the C2C9 nuclease and gRNA, the gRNA and C2C9 nuclease forming a complex, the gRNA providing target specificity to the targeting complex by containing a nucleotide sequence complementary to the target DNA; the C2C9 nuclease providing site-specific activity. In some embodiments, modified polynucleotides are used in the CRISPR-C2C9 system based on the C2C9 nuclease to modify the DNA or RNA and / or gRNA encoding the C2C9 nuclease introduced into the cell to edit loci of any one or more genomes.

[0215] The present invention also provides a cell comprising a host cell that has been genetically modified with the above-described C2C9 nuclease or a polynucleotide encoding the same, gRNA or a polynucleotide encoding the same, a recombinant expression vector, a system, or a composition.

[0216] In this invention, the effective dosage of gRNA and / or C2C9 nuclease and / or recombinant expression vector and / or donor polynucleotide is conventional to those skilled in the art. It can be determined according to different routes of administration and the characteristics of the disease being treated.

[0217] C2C9 nuclease is a novel nuclease discovered by the applicant, suitable for CRISPR gene editing systems. Guided by a corresponding guide RNA, it can precisely locate target genes and perform genomic DNA cleavage, achieving double-strand breaks. C2C9 nuclease variants can possess different or advantageous characteristics. These C2C9 nucleases and their variants offer previously unavailable opportunities for genome editing. C2C9 nucleases and their variants can exhibit varying activities depending on the system in which they are used. Although C2C9 has a molecular weight only half or smaller than Cas9 and Cas12a, it possesses genome cleavage efficiency close to or similar to Cas9 and Cas12a, enabling highly efficient genome editing in prokaryotic bacteria and eukaryotic cells, thus showing great promise in biotechnology and gene therapy.

[0218] The C2C9 nuclease provided by this invention exhibits advantageous properties compared to reported CRISPR-Cas endonucleases, such as higher activity in prokaryotic, eukaryotic, and / or in vitro environments. In some embodiments, C2C9 combines one or more of the following: small size, high editing activity, and the requirement of only a short PAM sequence. The small size of the RNA-guided nuclease in this invention refers to a nuclease with a length not exceeding approximately 800 amino acids.

[0219] In this invention, the bacteria or prokaryotic bacteria may be Escherichia coli, Klebsiella pneumoniae, Bacteroides ovalis, Campylobacter jejuni, Staphylococcus saprophyticus, Enterococcus faecalis, Bacteroides polymorpha, Bacteroides vulgaris, Bacteroides monomorpha, Lactobacillus casei, Bacteroides fragilis, Acinetobacter rumeni, Fusobacterium nucleatum, Bacteroides johnsonii, Bacteroides thalianae, Lactobacillus rhamnosus, Bacteroides masei, Bacteroides fecalis, Fusobacterium death, and Bifidobacterium breve, etc.

[0220] In this invention, the eukaryotic cells include, but are not limited to, mammalian cells, fungal cells, and other eukaryotic organisms. The fungi include yeasts and Aspergillus, such as *Saccharomyces cerevisiae*, *Hansenula polymorpha*, *Pichia pastoris*, *Kluyveromyces fragilis*, *Kluyveromyces lactis*, as well as *Schizosaccharomyces cerevisiae*, *Candida albicans*, *Candida dulcis*, *Candida glabrata*, *Candida guinea*, *Candida lactis*, *Candida crocea*, *Candida balsamina*, *Candida merlini*, *Candida oleophila*, *Candida parapsilosis*, *Candida tropicalis*, and *Candida utilis*, *Aspergillus fumigatus*, *Aspergillus flavus*, *Aspergillus niger*, *Aspergillus niger*, *Aspergillus lanceolata*, *Aspergillus glaucus*, *Aspergillus niger*, *Aspergillus oryzae*, *Aspergillus pyrophyllus*, and *Aspergillus versicolor*, etc.

[0221] In embodiments of this invention, a novel genome editing method based on a miniature CRISPR / C2C9 nuclease is disclosed. This invention demonstrates that, through the guiding and localizing functions of the guide RNA, C2C9 can precisely cleave genomic DNA, achieving double-strand breaks. Utilizing the host cell's own or exogenous repair mechanisms, this system can efficiently and accurately achieve gene editing within living cells.

[0222] The embodiments of the present invention will be clearly and completely described below with reference to examples. Obviously, the described embodiments are only used to illustrate a part of the embodiments of the present invention and should not be regarded as limiting the scope of the present invention. Where specific conditions are not specified in the embodiments, conventional conditions or conditions recommended by the manufacturer shall be followed. Where the manufacturers of reagents or instruments used are not specified, they shall be regarded as conventional products that can be purchased commercially.

[0223] The sources of the biological materials used in each embodiment are as follows:

[0224] Klebsiella pneumoniae NCTC9633 strain (purchased from ATCC, USA, catalog number: 13883); competent Escherichia coli DH5α strain (purchased from Kangwei Century Biotechnology Co., Ltd., catalog number: CW0808S). LB liquid medium used in each example was purchased from Sangon Biotech (Shanghai) Co., Ltd., catalog number: A507002-0250; LB solid medium was purchased from Sangon Biotech (Shanghai) Co., Ltd., catalog number: A507003-0250; apopramycin was purchased from BioVision, USA, catalog number: B1521-1G; kanamycin was purchased from Shanghai Aladdin Biochemical Technology Co., Ltd., catalog number: K103024-5g.

[0225] Example 1: Identification of AcC2C9 pre-interstitial sequence adjacent motif (PAM) and construction of plasmids p15a-AcC2C9Control and p15a-AcC2C9PAM

[0226] (1.1) The p15a-AcC2C9Control plasmid sequence is SEQ ID NO:124. Its specific construction method is as follows:

[0227] The DNA sequences of the AcC2C9 nuclease and its corresponding guide RNA expression cassette (SEQ ID NO: 125) and the p15a plasmid backbone (SEQ ID NO: 126) were synthesized in their entirety by Sangon Biotech (Shanghai) Co., Ltd. These two DNA fragments were assembled into the p15a-AcC2C9Control plasmid using Gibson assembly technology. The plasmid was transformed into commercially available E. coli DH5α competent cells and screened on LB agar plates containing 75 μg / mL apopramycin. Single clones were picked, expanded, and the plasmid was extracted. After sequencing to identify the plasmid sequence, the p15a-AcC2C9Control plasmid was obtained.

[0228] (1.2) The p15a-AcC2C9PAM plasmid sequence is SEQ ID NO:127, and its specific construction method is as follows:

[0229] The following two primers were synthesized at Sangon Biotech (Shanghai) Co., Ltd.:

[0230] PAMprimerF: 5'-atcgTGTCCTCTTCCTCTTTAGCG-3' (SEQ ID NO: 128)

[0231] PAMprimerR: 5'-ggccCGCTAAAGAGGAAGAGGACA-3' (SEQ ID NO: 129)

[0232] The two primers were annealed using the following reaction mixture: 5 μl 10×T4 DNA ligase Buffer (NEB), 10 μL PAM primer F (10 μM), 10 μL PAM primer R (10 μM), and 25 μL ddH2O. The mixture was heated at 95°C for 5 min, then slowly cooled to room temperature over 1–2 hours to allow the two single-stranded primers to form double-stranded DNA through base pairing. The resulting product was then diluted 20-fold with ddH2O.

[0233] The double-stranded DNA obtained above was inserted into the BsaI site of the p15a-AcC2C9Control plasmid using Golden Gate technology. The specific reaction system was as follows: 1 μL 10×T4 DNA ligase buffer, 1 μL of the phosphorylated double-stranded DNA diluted 20 times above, 20 fmol p15a-AcC2C9Control plasmid, 0.5 μL T4 DNA ligase (400 units / μL), 0.5 μL BsaI-HF (20 units / μL), and finally, an appropriate amount of ddH2O was added to a total volume of 10 μL. Both the buffer and enzyme used in this reaction were manufactured by NEB. The reaction was carried out in a PCR instrument with the following cycle: 37℃ for 2 min; 16℃ for 5 min, for a total of 25 cycles; then 80℃ for 15 min.

[0234] The 10 μL reaction product was transformed into competent Escherichia coli DH5α strain and plated on LB agar plates containing 75 μg / mL apopramycin. After the transformation solution was absorbed by the agar plates, the plates were incubated upside down at 37°C overnight. The transformed strains that grew on the culture medium were transferred and stored, and the plasmid was extracted and sent to Sangon Biotech (Shanghai) Co., Ltd. for sequencing verification, finally yielding the p15a-AcC2C9PAM plasmid.

[0235] Example 2: Construction of the substrate particle pUC19-6N-library for PAM identification

[0236] The pUC19-6N-library plasmid sequence is SEQ ID NO:130. Its specific construction method is as follows:

[0237] Using the commercial pUC19 plasmid as a template, circular polymerase extension cloning was performed using primer 6N-F (SEQ ID NO: 131) with a random 6N sequence and ordinary primer 6N-R (SEQ ID NO: 132). The PCR product was digested with DpnI and transformed into commercial E. coli DH5α competent cells. The cells were plated on LBA plates containing carbenicillin, and more than 100,000 single clones were collected using a plating applicator. After mixing, the plasmid was extracted to obtain the PAM identification substrate plasmid pUC19-6N-library.

[0238] Example 3: PAM identification of AcC2C9

[0239] The p15a-AcC2C9Control plasmid constructed in Example 1 was transformed into commercial E. coli DH5α competent cells and then plated on LB agar plates containing 75 μg / mL apopramycin. The next day, single colonies were picked and cultured in 100 mL of fresh LB liquid medium until OD... 600 After collecting bacteria at 0.5, the cells were washed twice with pre-chilled sterile 10% v / v glycerol on ice, and finally resuspended in 1 mL of sterile 10% v / v glycerol solution to obtain DH5α competent cells carrying the p15a-AcC2C9Control plasmid.

[0240] DH5α competent cells with the p15a-AcC2C9PAM plasmid constructed in Example 1 were prepared using the same process as those for preparing DH5α competent cells with the p15a-AcC2C9Control plasmid described above.

[0241] 100 ng of the pUC19-6N-library plasmid constructed in Example 2 was electroporated into DH5α competent cells carrying the p15a-AcC2C9Control plasmid and plated onto LB agar plates containing 50 μg / mL carbenicillin and 75 μg / mL apopramycin. More than 100,000 single colonies were collected using a plating applicator, mixed, and the plasmid was extracted to obtain the control library.

[0242] 100 ng of the pUC19-6N-library plasmid constructed in Example 2 was electroporated into DH5α competent cells carrying the p15a-AcC2C9PAM plasmid to prepare an AcC2C9 PAM depletion library. The preparation process was the same as that for the control library.

[0243] Using the control library and PAM depletion library as templates, a first round of PCR was performed using primers PAM-1F (SEQ ID NO: 133) and PAM-1R (SEQ ID NO: 134). The products were detected by electrophoresis. Then, 1 μL of the first round PCR product was used as a template for a second round of PCR using primers PAM-2F (SEQ ID NO: 135) and PAM-2R (SEQ ID NO: 136). The products were then detected by electrophoresis. Products were purified using VAHTS DNA CleanBeads (Novozymes Biotechnology Co., Ltd.). High-throughput sequencing was performed using the Illumina Hiseq 2500 platform, yielding 1 Gb of raw data. 6N random sequence information was extracted. The sequence frequencies in the PAM depletion library were compared with those in the control library to create a web logo, thereby obtaining the PAM information preferred by AcC2C9.

[0244] Figure 1 This diagram shows the identification process and results of PAM sequences using AcC2C9. The results indicate that AcC2C9 can effectively recognize 5'-AAN type PAM sequences and a small number of 5'-GAN sequences (where N is a degenerate base, representing any one of A, T, C, or G). The horizontal axes -1, -2, -3, -4, -5, and -6 represent the 1st, 2nd, 3rd, 4th, 5th, and 6th DNA bases upstream of the 5' end of the target-DNA segment sequence on the NTS strand (non-target strand), respectively. The letters above these bases represent the preference of the AcC2C9 nuclease for that DNA sequence; the greater the proportion of the letter, the stronger the preference of AcC2C9 for that base.

[0245] Example 4: Construction of p15a-AcC2C9 and pSGKP-AcC2C9 plasmids for Klebsiella pneumoniae genome editing

[0246] The p15a-AcC2C9 plasmid sequence is SEQ ID NO:137, and its specific construction method is as follows:

[0247] Using the p15a-AcC2C9Control plasmid from Example 1 as a template, the AcC2C9 nuclease gene and the p15a plasmid backbone were amplified respectively.

[0248] Amplification of the 5' primer sequence of the AcC2C9 nuclease gene: AcC2C9F

[0249] tcatctgtgcatatagctatactgatttcgtcagactcaca (SEQ ID NO: 138)

[0250] Amplification of the 3' primer sequence of the AcC2C9 nuclease gene: ACC2C9R

[0251] cgccaaccagccaCCTTATTGACCTGCACCGCTATG (SEQ ID NO: 139)

[0252] Amplifying the 5' primer sequence of the p15a plasmid backbone: p15aF:

[0253] CAGGTCAATAAGGtggctggttggcgtactgtt (SEQ ID NO: 140)

[0254] Amplifying the 3' primer sequence of the p15a plasmid backbone: p15aR:

[0255] atcagtatagctatatgcacagatgaaaacggtgtaaaaaaga (SEQ ID NO: 141)

[0256] The AcC2C9 gene fragment and p15a plasmid backbone were amplified using Phanta Max Master Mix reagent from Novizan. The reaction system consisted of: 25 μL 2x Phanta Max Master Mix, 1.5 μL 5' Primer (10 μM), 1.5 μL 3' Primer (10 μM), 0.5 μL template DNA (100 ng / μL), 1.5 μL DMSO, and 20 μL ddH2O. After preparation, polymerase chain reaction (PCR) was performed using the following cycle: 98℃ for 2 min; followed by 98℃ for 20 s, 55℃ for 20 s, and 72℃ for 3 min, for a total of 30 cycles; and finally, 72℃ for 5 min. The PCR products were recovered using the SanPrep column-based PCR product purification kit from Sangon Biotech (Shanghai) Co., Ltd. The specific purification steps were performed according to the kit's instruction manual.

[0257] The obtained AcC2C9 gene fragment and p15a plasmid backbone were assembled into a single plasmid using Gibson Assembly. The specific reaction system was as follows: 5 μL NEBuilder HiFi DNA Assembly Master Mix (NEB), 20 fmol AcC2C9 gene fragment, 20 fmol p15a plasmid backbone, and an appropriate amount of ddH2O added to a total volume of 10 μL. The reaction was carried out at 50 °C for 1 hour. 10 μL of the reaction product was transformed into competent E. coli DH5α strain and plated on LB agar plates containing 75 μg / mL apopramycin. After the transformation solution was absorbed by the agar plate, the plates were incubated overnight at 37 °C. The transformed strain grown on the agar plate was transferred and stored, and plasmids were extracted using the SanPrep column-based plasmid DNA mini-extraction kit from Sangon Biotech (Shanghai) Co., Ltd. for subsequent experiments. The specific steps for plasmid extraction were performed according to the kit's instruction manual. Meanwhile, the plasmid was sent to Sangon Biotech (Shanghai) Co., Ltd. for sequencing confirmation, and finally the p15a-AcC2C9 plasmid was obtained.

[0258] The pSGKP-AcC2C9 plasmid sequence is SEQ ID NO:142, and its specific construction method is as follows:

[0259] The DNA sequence of the guide RNA expression cassette corresponding to the AcC2C9 nuclease (SEQ ID NO: 143) and the pSGKP plasmid backbone (SEQ ID NO: 144) were synthesized in their entirety by Sangon Biotech (Shanghai) Co., Ltd. These two DNA fragments were assembled into the pSGKP-AcC2C9 plasmid using Gibson assembly technology. The plasmid was transformed into commercially available E. coli DH5α competent cells and screened on LB agar plates containing 50 μg / mL kanamycin. Single clones were picked, expanded, and the plasmid was extracted. After sequencing to identify the plasmid sequence, the pSGKP-AcC2C9 plasmid was obtained.

[0260] Example 5: Construction of pSGKP-AcC2C9-dha plasmid targeting the dha gene of Klebsiella pneumoniae genome.

[0261] The pSGKP-AcC2C9-dha plasmid sequence is SEQ ID NO:145, and its specific construction method is as follows:

[0262] First, a DNA fragment consisting of 20 bases following an AAG sequence (in this example, the AcC2C9-recognized PAM sequence is AAG) is selected from the target gene of Klebsiella pneumoniae. These 20 bases are referred to as the spacer; AAN is not included. Then, this spacer sequence is inserted into the p15a-AcC2C9Control plasmid constructed in Example 1. A key feature of this step is that, to allow the spacer fragment to insert into the BsaI site of the p15a-AcC2C9Control plasmid, atcg needs to be added to the 5' end of the single-stranded DNA sequence. Simultaneously, the complement sequence of the spacer needs to be synthesized, and ggcc needs to be added to the 5' end of the complement sequence. For example, the DNA sequence (SEQ ID NO. 146) of the dha gene spacer selected in this example is: 5'-TGTTTGTTACCAACTGCCTG-3'. The specific sequence design of the two primers is as follows:

[0263] PAM primerF(dha-spF): 5'-atcgTGTTTGTTACCAACTGCCTG-3'(SEQ ID NO.147)

[0264] PAM primerR(dha-spR): 5'-ggccCAGGCAGTTGTAACAAACA-3'(SEQ ID NO.148)

[0265] The two primers were annealed using the following reaction mixture: 5 μl 10×T4 DNA ligase Buffer (NEB), 10 μL PAM primer F (10 μM), 10 μL PAM primer R (10 μM), and 25 μL ddH2O. The mixture was heated at 95°C for 5 min, then slowly cooled to room temperature over 1–2 hours to allow the two single-stranded primers to form double-stranded DNA through base pairing. The resulting product was then diluted 20-fold with ddH2O.

[0266] The double-stranded DNA obtained above was inserted into the BsaI site of the pSGKP-AcC2C9 plasmid constructed in Example 4 using Golden Gate technology. The specific reaction system was as follows: 1 μL 10×T4 DNA ligase buffer, 1 μL of the phosphorylated double-stranded DNA diluted 20 times above, 20 fmol pSGKP-AcC2C9 plasmid, 0.5 μL T4 DNA ligase (400 units / μL), 0.5 μL BsaI-HF (20 units / μL), and finally, an appropriate amount of ddH2O was added to a total volume of 10 μL. The buffer and enzyme used in this reaction were manufactured by NEB. The reaction was carried out in a PCR instrument with the following cycle: 37℃ for 2 min; 16℃ for 5 min, for a total of 25 cycles; then 80℃ for 15 min.

[0267] The 10 μL reaction product was transformed into competent Escherichia coli DH5α strain and plated on LB agar plates containing 75 μg / mL apopramycin. After the transformation solution was absorbed by the agar plates, the plates were incubated upside down at 37°C overnight. The transformed strains that grew on the culture medium were transferred and stored, and plasmids were extracted and sent to Sangon Biotech (Shanghai) Co., Ltd. for sequencing verification, finally yielding the pSGKP-AcC2C9-dha plasmid.

[0268] Example 6: Preparation of Klebsiella pneumoniae electrocompetent cells containing p15a-AcC2C9 plasmid

[0269] Klebsiella pneumoniae strain NCTC9633 (purchased from ATCC, catalog number: 13883) was streaked onto LB agar plates and incubated overnight at 37°C with the plates inverted. A single colony was picked from the plate and inoculated into 3 mL of LB liquid medium, and incubated overnight at 37°C with shaking at 250 rpm. The next day, 1 mL of the bacterial culture was inoculated into 100 mL of fresh LB liquid medium and incubated at 37°C with shaking. When the OD of the bacterial culture... 600When the culture medium concentration reaches 0.3, cool the bacterial culture on ice for ten minutes. Centrifuge at 8000 rpm for 5 minutes at 4°C to collect the bacteria, discard the supernatant, and resuspend the bacterial pellet at the bottom in 20 mL of 10% (v / v) glycerol (autoclaved and pre-cooled on ice). Centrifuge at the same speed, discard the supernatant, and resuspend the bacterial pellet at the bottom in 20 mL of 10% (v / v) glycerol again. Centrifuge again at the same speed, discard the supernatant, and resuspend the bacterial pellet at the bottom in 1 mL of 10% (v / v) glycerol. Aliquot the obtained electrocompetent bacteria into EP tubes, 50 μL per tube. Quickly freeze the aliquoted bacterial culture in liquid nitrogen and then store at -80°C. Note that the resuscitation process should be gentle; use a pipette to gently resuspend the bacteria.

[0270] Take a tube of freshly prepared Klebsiella pneumoniae electrotransfer competent cells, place it on ice for 5-10 minutes, and then add 1 μg of the p15a-AcC2C9 plasmid prepared in Example 4. After mixing thoroughly, transfer it to a 1 mm electrotransfer cuvette (Bio-Rad) and electrolyze it at room temperature using a GenePulserXcell electroporator (Bio-Rad). The electroporation parameters are: 1800 V, 200 Ω, 25 μF. Immediately after electroporation, add 1 mL of LB liquid medium, mix thoroughly, and transfer it to a clean EP tube. Shake at 37°C for 1-2 hours. Spread 100 μL of the bacterial culture onto an LB agar plate containing 75 μg / mL apopramycin. After the bacterial culture is absorbed by the agar plate, incubate in an inverted incubator at 37°C overnight. Pick a single colony and incubate it overnight at 37°C with a shaker at 250 rpm in 3 mL of LB broth containing 75 μg / mL apopramycin. The next day, inoculate 1 mL of the bacterial culture into 100 mL of fresh LB broth containing 75 μg / mL apopramycin and continue shaking in a shaker at 37°C. When the OD of the bacterial culture... 600 When the culture medium concentration reaches 0.2, add 1 mL of 20% w / v arabinose to the bacterial culture and continue shaking for 2 hours, then cool the culture on ice for 10 minutes. Centrifuge at 8000 rpm for 5 minutes at 4°C to collect the bacteria. Discard the supernatant and resuspend the bacterial pellet at the bottom twice with 20 mL of 10% v / v glycerol (autoclaved and pre-cooled on ice). Centrifuge again at the same speed, discard the supernatant, and resuspend the bacterial pellet at the bottom with 1 mL of 10% v / v glycerol. Aliquot the obtained electrocompetent bacteria into EP tubes, 50 μL per tube. Quickly freeze the aliquots in liquid nitrogen and then store them at -80°C. Note that resuscitation should be done gently; use a pipette to gently resuspend the bacteria.

[0271] Example 7: Efficient Gene Deletion in Live Klebsiella pneumoniae Cells Using the AcC2C9 System

[0272] Take a tube containing Klebsiella pneumoniae electrocompetent cells with the p15a-AcC2C9 plasmid prepared in Example 6, place it on ice for 5-10 minutes until thawed, then add 200 ng of the pSGKP-AcC2C9-dha plasmid prepared in Example 5 and a final concentration of 5 μM of dha gene homologous recombination repair template (its sequence is SEQ ID NO: 149), and mix gently. Use a pipette to transfer the mixed bacterial plasmid mixture into a pre-chilled 1 mm electroporation cuvette (Bio-Rad), and incubate on ice for 5 minutes. Wipe off any condensation from the outside of the electroporation cuvette, and perform electroporation using a GenePulserXcell electroporator (Bio-Rad). The electroporation parameters are: 1800 V, 200 Ω, 25 μF. Immediately after electroporation, 1 mL of LB medium was added to wash out the electroporated cells, which were then transferred to sterile EP tubes and incubated at 37°C with a shaker at 250 rpm / min for 1 h. The bacterial culture was then spread onto LB solid medium containing 75 μg / mL apopramycin and 50 μg / mL kanamycin. After the bacterial culture was absorbed, the cells were incubated upside down overnight at 37°C.

[0273] The edited effect was verified by colony PCR on the grown colonies.

[0274] The 5' primer for colony PCR verification of the dha gene is:

[0275] 5'-TCATCATTGGCATCCTGAAA-3'(SEQ ID NO.150)

[0276] The 3' primer for colony PCR verification of the dha gene is:

[0277] 5'-TGGTCAGCCTGCTGATTAAA-3'(SEQ ID NO.151)

[0278] Figure 2 This image shows the PCR results of gene knockout in live Klebsiella pneumoniae cells using the CRISPR-AcC2C9 system. The results show that the selected dha gene can be precisely knocked out by the CRISPR-AcC2C9 system.

[0279] Example 8: Construction of the AcC2C9 protein heterologous expression plasmid pET28a-sumo-AcC2C9

[0280] Using the p15a-AcC2C9Control plasmid from Example 1 as a template, the AcC2C9 nuclease gene sequence was amplified to construct the pET28a-sumo-AcC2C9 plasmid expressing AcC2C9 (its sequence is SEQ ID NO:152).

[0281] Amplification of the AcC2C9 gene 5' primer sequence pET-AcC2C9F (5' Primer): 5'-tCTTGAAGTCCTCTTTCAGGGACCCATGGTGCAGACGGAGATCCTG-3' (SEQ ID NO: 153):

[0282] Amplifying the AcC2C9 gene 3' primer sequence pET-AcC2C9R (3' Primer):

[0283] 5'-gtcgacggagctcgaattcttaagcTTATGGAGGCGTTGCCGTTTCG-3' (SEQ ID NO: 154):

[0284] The AcC2C9 gene fragment was amplified using Phanta Max Master Mix reagent from Novizan. The reaction mixture consisted of: 25 μL 2x Phanta Max Master Mix, 1.5 μL 5' Primer (10 μM), 1.5 μL 3' Primer (10 μM), 0.5 μL p15a-AcC2C9 plasmid prepared in Example 4 (100 ng / μL), 1.5 μL LDMSO, and 20 μL ddH2O. After preparation, polymerase chain reaction (PCR) was performed using the following cycle: 98℃ for 2 min; followed by 98℃ for 20 s, 55℃ for 20 s, and 72℃ for 1 min, for a total of 30 cycles; and finally, 72℃ for 5 min. The PCR products were purified using the SanPrep column-based PCR product purification kit manufactured by Sangon Biotech (Shanghai) Co., Ltd. The specific purification steps were performed according to the kit's instruction manual.

[0285] The pET28a-sumo plasmid backbone (sequence SEQ ID NO: 155) was synthesized using Sangon Biotech (Shanghai) Co., Ltd. The two AcC2C9 gene fragments and the pET28a-sumo plasmid backbone were assembled into the pET28a-sumo-AcC2C9 plasmid using Gibson assembly technology. The plasmid was transformed into commercially available E. coli DH5α competent cells and screened on LB agar plates containing 50 μg / mL kanamycin. Single clones were picked, expanded, and the plasmid was extracted. After sequencing to identify the plasmid sequence, the pET28a-sumo-AcC2C9 plasmid was obtained.

[0286] Example 9: AcC2C9 Achieves Efficient In Vitro Double-Stranded DNA Cutting

[0287] Preparation of cleavage substrates. Using the genome of *Klebsiella pneumoniae* as a template, the cleavage substrates were amplified using 5' FAM-labeled primers. The 5' primer sequence for amplifying the cleavage substrates was: 5'-CCGCATATATCGCAAAACAA-3' (SEQ ID NO: 156), and the 3' primer sequence was: 5'-CTGGCTGTGAACAAAGTGGA-3' (SEQ ID NO: 157). The PCR products were recovered using the SanPrep column-based PCR product purification kit manufactured by Sangon Biotech (Shanghai) Co., Ltd.

[0288] Preparation of AcC2C9 nuclease. The pET28a-sumo-AcC2C9 plasmid constructed in Example 8 was transformed into Escherichia coli expression strain BL21(DE3). The next day, the transformants were transferred into 1 L of LB medium and cultured with shaking at 37°C. When OD 600 When the culture medium reached 0.6, 0.5 mL of 1 M IPTG was added, and the temperature was lowered to 16 °C for overnight incubation. The strain was collected after overnight incubation, sonicated, and purified using a HisTrap Ni-NTA (GE Healthcare) chromatography column. Further purification was then performed using a HiLoad 16 / 600 Superdex 200 pg molecular sieve (GE Healthcare). The purified protein was concentrated using ultrafiltration and stored in a buffer solution of 500 mM NaCl, 10 mM Tris-HCl, pH 7.5, and 1 mM DTT.

[0289] Preparation of complete guide RNA (SEQ ID NO: 158) targeting substrate cleavage. The complete guide RNA was prepared via in vitro transcription. The transcription template DNA was synthesized by Genewiz Biotechnology Co., Ltd., and its sequence is SEQ ID NO: 159. The complete guide RNA was transcribed using the above template via the HiScribe T7 High Yield RNA Synthesis Kit (NEB). Specific steps were performed according to the kit's instruction manual. The prepared complete guide RNA was purified by phenol-chloroform extraction and ethanol precipitation.

[0290] In vitro cleavage experiments were performed under the following conditions: 150 mM NaCl, 10 mM MgCl2, 10 mM Tris-HCl, pH 7.5, and 1 mM DTT. The total reaction volume was 20 μL, which included 30 nM cleavage substrate, 2 μM AcC2C9, and 2 μM intact guide RNA. The reaction times were 0 min, 1 min, 2 min, 5 min, 15 min, and 30 min, respectively, and the reaction temperature was 37 °C. After the reaction, 50 mM EDTA and 2 mg / mL Protein K were added to terminate the reaction. The reaction products were separated by 6% w / v TBE acrylamide gel electrophoresis and imaged using a Bio-Rad GELDOC gel imaging system after excitation with 488 nm blue light.

[0291] Figure 3 This image shows the results of cutting substrate DNA using the AcC2C9 nuclease. The results show that the substrate DNA was precisely cleaved into two DNA fragments by AcC2C9.

[0292] Example 10: Activity Detection of Complete Guide RNA after Separation into tracrRNA and crRNA

[0293] The complete guide RNA for targeted substrate cleavage includes tracrRNA and crRNA (an RNA sequence formed by linking the DNA-target region hybridized with the target sequence to the tracr pair sequence). To verify whether the complete guide RNA retains its biological activity after being split into tracrRNA and crRNA, this invention performed an in vitro double-stranded DNA cleavage experiment.

[0294] Cut the substrate. Same as described in Example 9.

[0295] Preparation of AcC2C9 nuclease. Same as described in Example 9.

[0296] Preparation of tracrRNA (SEQ ID NO: 160). tracrRNA was prepared via in vitro transcription. The transcription template DNA was synthesized by Genewiz Biotechnology Co., Ltd., and its sequence is SEQ ID NO: 161. tracrRNA was transcribed using the above template via the HiScribe T7 High Yield RNA Synthesis Kit (NEB). Specific steps were performed according to the kit's instruction manual. The prepared tracrRNA was purified by phenol-chloroform extraction and ethanol precipitation.

[0297] Preparation of crRNA (SEQ ID NO: 162). crRNA was prepared via in vitro transcription. The transcription template DNA was synthesized by Genewiz Biotechnology Co., Ltd., and its sequence is SEQ ID NO: 163. The transcription and purification steps for crRNA were the same as those for tracrRNA.

[0298] In vitro cleavage experiments were conducted under the conditions of 150 mM NaCl, 10 mM MgCl2, 10 mM Tris-HCl, pH 7.5, and 1 mM DTT. The total reaction volume was 20 μL, including 30 nM cleavage substrate, 2 μM AcC2C9, 2 μM tracrRNA, and 2 μM crRNA. The reaction time was 30 min, and the reaction temperature was 37 °C. After the reaction, 50 mM EDTA and 2 mg / mL Protein K were added to terminate the reaction. The reaction products were separated by 6% w / v TBE acrylamide gel electrophoresis and imaged under 488 nm blue light excitation in a Bio-Rad GELDOC gel imaging system. It can be seen that under both crRNA and tracrRNA conditions, the substrate DNA was precisely cleaved into two DNA fragments, indicating that the intact guide RNA retains biological activity after being separated into tracrRNA and crRNA.

[0299] Figure 4 (A) Schematic diagram of the complete guide RNA sequence of AcC2C9 (B) Schematic diagram of the results of substrate DNA being cleaved by AcC2C9 nuclease under tracrRNA and crRNA conditions.

[0300] Example 11: Identification of the AcC2C9 cleavage site in double-stranded DNA

[0301] Preparation of complete guide RNA targeting short fluorescently labeled synthetic DNA. Same as described in Example 9.

[0302] Preparation of AcC2C9 nuclease. Same as described in Example 9.

[0303] Substrate preparation. The cleavage substrate consisted of four 58 bp fluorescently labeled synthetic DNA sequences. Two complementary primers were synthesized by Sangon Biotech (Shanghai) Co., Ltd., with sequences 5'-GGCTGTGAGAAAGCGGTTCAGGTGAAAGTGAAAACACTGCCCGACGCCCAGTTCGAAG-3' (SEQ ID NO:164) and 5'-CTTCGAACTGGGCGTCGGGCAGTGTTTTCACTTTCACCTGAACCGCTTTCTCACAGCC-3' (SEQ ID NO:165). FAM fluorescent groups were labeled at the 5' or 3' end of each primer, respectively.

[0304] In vitro cleavage experiments were performed under the following conditions: 150 mM NaCl, 10 mM MgCl2, 10 mM Tris-HCl, pH 7.5, and 1 mM DTT. The total reaction volume was 10 μL, including 20 nM cleavage substrate, 2 μM AcC2C9, and 2 μM intact guide RNA. The reaction temperature was 37 °C, and the reaction time was 30 min. After the reaction, 10 μL of 2×formamide loading buffer (Sangon Biotech (Shanghai) Co., Ltd.) was added to terminate the reaction. The reaction products were separated by 20% TBE-Urea-PAGE and imaged after excitation with 488 nm blue light.

[0305] Figure 5 The image shows the results of cleavage of four 58 bp fluorescently labeled DNA substrates using the AcC2C9 nuclease. The results show that the major cleavage site of AcC2C9 on the target strand is 21-22 nt downstream of the 3' end of the PAM site, while the major cleavage site on the non-target strand is 15-16 nt downstream of the 3' end of the PAM site.

[0306] Example 11: Construction of the mammalian cell gene editing plasmid pAcC2C9Hs for AcC2C9

[0307] The AcC2C9 described in this embodiment uses a better AcC2C9 encoding gene optimized for human codons.

[0308] The pAcC2C9Hs plasmid sequence is SEQ ID NO:166, and its specific construction method is as follows:

[0309] The human transient expression plasmid backbone (SEQ ID NO:167), the puromycin resistance gene expression cassette sequence (SEQ ID NO:168), the AcC2C9 coding gene expression cassette sequence optimized for human codons (SEQ ID NO:169), and the corresponding guide RNA expression cassette for AcC2C9 in human cells (SEQ ID NO:123) (SEQ ID NO:170) were all synthesized by Sangon Biotech (Shanghai) Co., Ltd. The four fragments were assembled into the pAcC2C9Hs plasmid using Gibson assembly technology. The plasmid was transformed into commercially available E. coli DH5α competent cells, plated on LBA plates containing carbenicillin, and single colonies were picked for expansion culture. The plasmid was extracted, sequenced, and identified to obtain the pAcC2C9Hs plasmid.

[0310] Example 12: High-efficiency gene editing in human living cells using AcC2C9

[0311] The AcC2C9 nuclease enables highly efficient gene editing in living human cells. The gene editing effect is demonstrated using human embryonic kidney cells (HEK293) as a model.

[0312] A target spacer sequence was inserted into the pAcC2C9Hs plasmid. In this embodiment, five genes on the human genome—VEGFA, EMX1, HEXA, DNMT1, FANCF, and PDCD1—were selected as target sequences, and a 20 bp target sequence DNA was selected from each of these target sequences. Oligonucleotide sequences Guide1-F, Guide1-R, Guide2-F, Guide2-R, Guide3-F, Guide3-R, Guide4-F, Guide4-R, Guide5-F, and Guide5-R (sequences SEQ ID NO: 171-180) were synthesized. After annealing Guide1-F and Guide1-R, Guide2-F and Guide2-R, Guide3-F and Guide3-R, Guide4-F and Guide4-R, and Guide5-F and Guide5-R, respectively, the five target sequence DNAs were inserted into the pAcC2C9Hs plasmid using Golden gate assembly technology to construct the pAcC2C9Hs-guide1, pAcC2C9Hs-guide2, pAcC2C9Hs-guide3, pAcC2C9Hs-guide4, and pAcC2C9Hs-guide5 plasmids (sequences SEQ ID NO: 181-185, respectively).

[0313] Transient expression plasmid-mediated human cell gene editing. Activation of cryopreserved HEK293 cells. Two days later, HEK293 cells were passaged into 24-well plates, with approximately 1.5 × 10⁻⁶ cells per well. 5 After 16-18 hours, 500 ng of pAcC2C9Hs-guide1, pAcC2C9Hs-guide2, pAcC2C9Hs-guide3, pAcC2C9Hs-guide4, and pAcC2C9Hs-guide5 plasmids were transfected into cells with 0.75 μL of lipofectamine 3000 (Invitrogen). Cells were then cultured for another 72 hours. Adherent cells were digested, and genomic DNA was extracted. Gene fragments were amplified by PCR using the corresponding primers (Guide1-IndelF and Guide1-IndelR, Guide2-IndelF and Guide2-IndelR, Guide3-IndelF and Guide3-IndelR, Guide4-IndelF and Guide4-IndelR, Guide5-IndelF and Guide5-IndelR, with sequences SEQ ID NO: 186-195). PCR products were recovered from the gel. The PCR products were annealed using NEBuffer 2 (NEB). Then, T7 endonuclease 1 (NEB) was added, and the mixture was digested at 37°C for 15 min. The reaction was terminated with 50 mM EDTA, and the products were separated by 6% TBE-PAGE and stained with 4S Red dye (Sangon Biotech (Shanghai) Co., Ltd.) for imaging.

[0314] Figure 6 Figure 1 shows the results of gene editing in human cells mediated by AcC2C9 transient expression plasmid. (A) is a TBE-PAGE gel image from the Indel experiment. The AcC2C9 nuclease successfully achieved efficient gene editing of the VEGFA, HEXA, and DNMT1 genes at five different selected gene loci. (B) is a statistical graph of high-throughput sequencing results after VEGFA gene editing in human cells mediated by the AcC2C9 transient expression plasmid. As shown, the AcC2C9 nuclease can precisely introduce insertion or deletion mutations at target sequences.

[0315] In this invention, the sequence information is as follows:

[0316] SEQ ID NO.1

[0317] Sequence length: 76

[0318] Sequence type: Protein

[0319] Sequence source: Actinomadura craniellae

[0320] Sequence name: AcC2C9 protein domain 1

[0321] SEQ ID NO.2

[0322] Sequence length: 67

[0323] Sequence type: Protein

[0324] Sequence source: Actinomadura craniellae

[0325] Sequence name: AcC2C9 protein domain 2

[0326] SEQ ID NO.3

[0327] Sequence length: 506

[0328] Sequence type: Protein

[0329] Sequence source: Actinomadura craniellae

[0330] Sequence name: AcC2C9

[0331] SEQ ID NO.4

[0332] Sequence length: 493

[0333] Sequence type: Protein

[0334] Sequence source: Candidatus Frankia meridionalis

[0335] Sequence name: C2C9 protein sequence 2

[0336] SEQ ID NO.5

[0337] Sequence length: 537

[0338] Sequence type: Protein

[0339] Sequence source: Corynebacterium glutamicum

[0340] Sequence name: C2C9 protein sequence 3

[0341] SEQ ID NO.6

[0342] Sequence length: 537

[0343] Sequence type: Protein

[0344] Sequence source: Dermacoccus sp. UBA1591

[0345] Sequence name: C2C9 protein sequence 4

[0346] SEQ ID NO.7

[0347] Sequence length: 492

[0348] Sequence type: Protein

[0349] Sequence source: Kitasatospora sp.NA04385

[0350] Sequence name: C2C9 protein sequence 5

[0351] SEQ ID NO.8

[0352] Sequence length: 553

[0353] Sequence type: Protein

[0354] Sequence source: Kocuria indica

[0355] Sequence name: C2C9 protein sequence 6

[0356] SEQ ID NO.9

[0357] Sequence length: 539

[0358] Sequence type: Protein

[0359] Sequence source: Kocuria sp.cx-455

[0360] Sequence name: C2C9 protein sequence 7

[0361] SEQ ID NO.10

[0362] Sequence length: 552

[0363] Sequence type: Protein

[0364] Sequence source: Mesorhizobium sp. B2-6-6

[0365] Sequence name: C2C9 protein sequence 8

[0366] SEQ ID NO.11

[0367] Sequence length: 485

[0368] Sequence type: Protein

[0369] Sequence source: Mycobacterium heckeshornense

[0370] Sequence name: C2C9 protein sequence 9

[0371] SEQ ID NO.12

[0372] Sequence length: 569

[0373] Sequence type: Protein

[0374] Sequence source: Micrococcus luteus

[0375] Sequence name: C2C9 protein sequence 10

[0376] SEQ ID NO.13

[0377] Sequence length: 505

[0378] Sequence type: Protein

[0379] Sequence source: Mycobacterium sp.MS1601

[0380] Sequence name: C2C9 protein sequence 11

[0381] SEQ ID NO.14

[0382] Sequence length: 561

[0383] Sequence type: Protein

[0384] Sequence source: Mycolicibacterium goodii

[0385] Sequence name: C2C9 protein sequence 12

[0386] SEQ ID NO.15

[0387] Sequence length: 496

[0388] Sequence type: Protein

[0389] Sequence source: Nocardiopsis metallicus

[0390] Sequence name: C2C9 protein sequence 13

[0391] SEQ ID NO.16

[0392] Sequence length: 499

[0393] Sequence type: Protein

[0394] Sequence source: Nocardiopsis synnemataformans DSM 44143

[0395] Sequence name: C2C9 protein sequence 14

[0396] SEQ ID NO.17

[0397] Sequence length: 560

[0398] Sequence type: Protein

[0399] Sequence source: Planctomycetes bacterium

[0400] Sequence name: C2C9 protein sequence 15

[0401] SEQ ID NO.18

[0402] Sequence length: 532

[0403] Sequence type: Protein

[0404] Sequence source: Propionimicrobium lymphophilum

[0405] Sequence name: C2C9 protein sequence 16

[0406] SEQ ID NO.19

[0407] Sequence length: 579

[0408] Sequence type: Protein

[0409] Sequence source: Pseudactinotalea sp.HY160

[0410] Sequence name: C2C9 protein sequence 17

[0411] SEQ ID NO.20

[0412] Sequence length: 540

[0413] Sequence type: Protein

[0414] Sequence source: Rothia dentocariosa

[0415] Sequence name: C2C9 protein sequence 18

[0416] SEQ ID NO.21

[0417] Sequence length: 549

[0418] Sequence type: Protein

[0419] Sequence source: Rothia dentocariosa

[0420] Sequence name: C2C9 protein sequence 19

[0421] SEQ ID NO.22

[0422] Sequence length: 541

[0423] Sequence type: Protein

[0424] Sequence source: Rothia kristinae

[0425] Sequence name: C2C9 protein sequence 20

[0426] SEQ ID NO.23

[0427] Sequence length: 558

[0428] Sequence type: Protein

[0429] Sequence source: Rothia mucilaginosa

[0430] Sequence name: C2C9 protein sequence 21

[0431] SEQ ID NO.24

[0432] Sequence length: 559

[0433] Sequence type: Protein

[0434] Sequence source: Rothia mucilaginosa

[0435] Sequence name: C2C9 protein sequence 22

[0436] SEQ ID NO.25

[0437] Sequence length: 559

[0438] Sequence type: Protein

[0439] Sequence source: Rothia mucilaginosa

[0440] Sequence name: C2C9 protein sequence 23

[0441] SEQ ID NO.26

[0442] Sequence length: 485

[0443] Sequence type: Protein

[0444] Sequence source: Rothia nasimurium

[0445] Sequence name: C2C9 protein sequence 24

[0446] SEQ ID NO.27

[0447] Sequence length: 562

[0448] Sequence type: Protein

[0449] Sequence source: Rothia sp. HMSC071C12

[0450] Sequence name: C2C9 protein sequence 25

[0451] SEQ ID NO.28

[0452] Sequence length: 494

[0453] Sequence type: Protein

[0454] Sequence source: Saccharomonospora piscinae

[0455] Sequence name: C2C9 protein sequence 26

[0456] SEQ ID NO.29

[0457] Sequence length: 606

[0458] Sequence type: Protein

[0459] Sequence source: Saccharopolyspora rectivirgula DSM 43747

[0460] Sequence name: C2C9 protein sequence 27

[0461] SEQ ID NO.30

[0462] Sequence length: 459

[0463] Sequence type: Protein

[0464] Sequence source: Streptomyces albidus

[0465] Sequence name: C2C9 protein sequence 28

[0466] SEQ ID NO.31

[0467] Sequence length: 517

[0468] Sequence type: Protein

[0469] Sequence source: Streptomyces bauhiniae

[0470] Sequence name: C2C9 protein sequence 29

[0471] SEQ ID NO.32

[0472] Sequence length: 469

[0473] Sequence type: Protein

[0474] Sequence source: Streptomyces filamentosus

[0475] Sequence name: C2C9 protein sequence 30

[0476] SEQ ID NO.33

[0477] Sequence length: 583

[0478] Sequence type: Protein

[0479] Sequence source: Streptomyces griseocarneus

[0480] Sequence name: C2C9 protein sequence 31

[0481] SEQ ID NO.34

[0482] Sequence length: 466

[0483] Sequence type: Protein

[0484] Sequence source: Streptomyces longispororuber

[0485] Sequence name: C2C9 protein sequence 32

[0486] SEQ ID NO.35

[0487] Sequence length: 463

[0488] Sequence type: Protein

[0489] Sequence source: Streptomyces lydicus

[0490] Sequence name: C2C9 protein sequence 33

[0491] SEQ ID NO.36

[0492] Sequence length: 511

[0493] Sequence type: Protein

[0494] Sequence source: Streptomyces netropsis

[0495] Sequence name: C2C9 protein sequence 34

[0496] SEQ ID NO.37

[0497] Sequence length: 519

[0498] Sequence type: Protein

[0499] Sequence source: Streptomyces niveus NCIMB 11891

[0500] Sequence name: C2C9 protein sequence 35

[0501] SEQ ID NO.38

[0502] Sequence length: 515

[0503] Sequence type: Protein

[0504] Sequence source: Streptomyces rimosus subsp.rimosus

[0505] Sequence name: C2C9 protein sequence 36

[0506] SEQ ID NO.39

[0507] Sequence length: 518

[0508] Sequence type: Protein

[0509] Sequence source: Streptomyces sp.CB04723

[0510] Sequence name: C2C9 protein sequence 37

[0511] SEQ ID NO.40

[0512] Sequence length: 532

[0513] Sequence type: Protein

[0514] Sequence source: Streptomyces sp.CB04723

[0515] Sequence name: C2C9 protein sequence 38

[0516] SEQ ID NO.41

[0517] Sequence length: 513

[0518] Sequence type: Protein

[0519] Sequence source: Streptomyces sp.CNH099

[0520] Sequence name: C2C9 protein sequence 39

[0521] SEQ ID NO.42

[0522] Sequence length: 522

[0523] Sequence type: Protein

[0524] Sequence source: Streptomyces sp. CNS606

[0525] Sequence name: C2C9 protein sequence 40

[0526] SEQ ID NO.43

[0527] Sequence length: 519

[0528] Sequence type: Protein

[0529] Sequence source: Streptomyces sp.IB2014 016-6

[0530] Sequence name: C2C9 protein sequence 41

[0531] SEQ ID NO.44

[0532] Sequence length: 518

[0533] Sequence type: Protein

[0534] Sequence source: Streptomyces sp.man185

[0535] Sequence name: C2C9 protein sequence 42

[0536] SEQ ID NO.45

[0537] Sequence length: 506

[0538] Sequence type: Protein

[0539] Sequence source: Streptomyces sp.S4

[0540] Sequence name: C2C9 protein sequence 43

[0541] SEQ ID NO.46

[0542] Sequence length: 477

[0543] Sequence type: Protein

[0544] Sequence source: Streptomyces sp. SID2563

[0545] Sequence name: C2C9 protein sequence 44

[0546] SEQ ID NO.47

[0547] Sequence length: 509

[0548] Sequence type: Protein

[0549] Sequence source: Streptomyces sp. UH6

[0550] Sequence name: C2C9 protein sequence 45

[0551] SEQ ID NO.48

[0552] Sequence length: 468

[0553] Sequence type: Protein

[0554] Sequence source: Streptomyces spororaveus

[0555] Sequence name: C2C9 protein sequence 46

[0556] SEQ ID NO.49

[0557] Sequence length: 462

[0558] Sequence type: Protein

[0559] Sequence source: Streptomyces xanthochromogenes

[0560] Sequence name: C2C9 protein sequence 47

[0561] SEQ ID NO.50

[0562] Sequence length: 530

[0563] Sequence type: Protein

[0564] Sequence source: Streptosporangium violaceochromogenes

[0565] Sequence name: C2C9 protein sequence 48

[0566] SEQ ID NO.51

[0567] Sequence length: 512

[0568] Sequence type: Protein

[0569] Sequence source: metagenomic sequencing

[0570] Sequence name: C2C9 protein sequence 49

[0571] SEQ ID NO.52

[0572] Sequence length: 758

[0573] Sequence type: Protein

[0574] Sequence source: metagenomic sequencing

[0575] Sequence name: C2C9 protein sequence 50

[0576] SEQ ID NO.53

[0577] Sequence length: 669

[0578] Sequence type: Protein

[0579] Sequence source: Cellulosimicrobium cellulans

[0580] Sequence name: C2C9 protein sequence 51

[0581] SEQ ID NO.54

[0582] Sequence length: 531

[0583] Sequence type: Protein

[0584] Sequence source: Rhodococcus hoagii

[0585] Sequence name: C2C9 protein sequence 52

[0586] SEQ ID NO.55

[0587] Sequence length: 527

[0588] Sequence type: Protein

[0589] Sequence source: Rhodococcus sp. OK519

[0590] Sequence name: C2C9 protein sequence 53

[0591] SEQ ID NO.56

[0592] Sequence length: 560

[0593] Sequence type: Protein

[0594] Sequence source: Rothia sp. HMSC071C12

[0595] Sequence name: C2C9 protein sequence 54

[0596] SEQ ID NO.57

[0597] Sequence length: 496

[0598] Sequence type: Protein

[0599] Sequence source: Nocardiopsis alba

[0600] Sequence name: C2C9 protein sequence 55

[0601] SEQ ID NO.58

[0602] Sequence length: 500

[0603] Sequence type: Protein

[0604] Sequence source: metagenomic sequencing

[0605] Sequence name: C2C9 protein sequence 56

[0606] SEQ ID NO.59

[0607] Sequence length: 539

[0608] Sequence type: Protein

[0609] Sequence source: Streptosporangium subroseum

[0610] Sequence name: C2C9 protein sequence 57

[0611] SEQ ID NO.60

[0612] Sequence length: 595

[0613] Sequence type: Protein

[0614] Sequence source: metagenomic sequencing

[0615] Sequence name: C2C9 protein sequence 58

[0616] SEQ ID NO.61

[0617] Sequence length: 631

[0618] Sequence type: Protein

[0619] Sequence source: Streptomyces agglomeratus

[0620] Sequence name: C2C9 protein sequence 59

[0621] SEQ ID NO.62

[0622] Sequence length: 505

[0623] Sequence type: Protein

[0624] Sequence source: Murinocardiopsis flavida

[0625] Sequence name: C2C9 protein sequence 60

[0626] SEQ ID NO.63

[0627] Sequence length: 566

[0628] Sequence type: Protein

[0629] Sequence source: Streptomyces sp. NRRL F-2664

[0630] Sequence name: C2C9 protein sequence 61

[0631] SEQ ID NO.64

[0632] Sequence length: 715

[0633] Sequence type: Protein

[0634] Sequence source: metagenomic sequencing

[0635] Sequence name: C2C9 protein sequence 62

[0636] SEQ ID NO.65

[0637] Sequence length: 529

[0638] Sequence type: Protein

[0639] Sequence source: Cellulomonas marina

[0640] Sequence name: C2C9 protein sequence 63

[0641] SEQ ID NO.66

[0642] Sequence length: 449

[0643] Sequence type: Protein

[0644] Sequence source: metagenomic sequencing

[0645] Sequence name: C2C9 protein sequence 64

[0646] SEQ ID NO.67

[0647] Sequence length: 577

[0648] Sequence type: Protein

[0649] Sequence source: metagenomic sequencing

[0650] Sequence name: C2C9 protein sequence 65

[0651] SEQ ID NO.68

[0652] Sequence length: 582

[0653] Sequence type: Protein

[0654] Sequence source: Miniimonas sp.S16

[0655] Sequence name: C2C9 protein sequence 66

[0656] SEQ ID NO.69

[0657] Sequence length: 509

[0658] Sequence type: Protein

[0659] Sequence source: Streptomyces sp.SPB074

[0660] Sequence name: C2C9 protein sequence 67

[0661] SEQ ID NO.70

[0662] Sequence length: 509

[0663] Sequence type: Protein

[0664] Sequence source: Streptomyces yunnanensis

[0665] Sequence name: C2C9 protein sequence 68

[0666] SEQ ID NO.71

[0667] Sequence length: 760

[0668] Sequence type: Protein

[0669] Sequence source: Mycobacteroides abscessus

[0670] Sequence name: C2C9 protein sequence 69

[0671] SEQ ID NO.72

[0672] Sequence length: 494

[0673] Sequence type: Protein

[0674] Sequence source: metagenomic sequencing

[0675] Sequence name: C2C9 protein sequence 70

[0676] SEQ ID NO.73

[0677] Sequence length: 545

[0678] Sequence type: Protein

[0679] Sequence source: Streptomyces sp.CB02009

[0680] Sequence name: C2C9 protein sequence 71

[0681] SEQ ID NO.74

[0682] Sequence length: 433

[0683] Sequence type: Protein

[0684] Sequence source: metagenomic sequencing

[0685] Sequence name: C2C9 protein sequence 72

[0686] SEQ ID NO.75

[0687] Sequence length: 681

[0688] Sequence type: Protein

[0689] Sequence source: metagenomic sequencing

[0690] Sequence name: C2C9 protein sequence 73

[0691] SEQ ID NO.76

[0692] Sequence length: 528

[0693] Sequence type: Protein

[0694] Sequence source: metagenomic sequencing

[0695] Sequence name: C2C9 protein sequence 74

[0696] SEQ ID NO.77

[0697] Sequence length: 741

[0698] Sequence type: Protein

[0699] Sequence source: metagenomic sequencing

[0700] Sequence name: C2C9 protein sequence 75

[0701] SEQ ID NO.78

[0702] Sequence length: 611

[0703] Sequence type: Protein

[0704] Sequence source: metagenomic sequencing

[0705] Sequence name: C2C9 protein sequence 76

[0706] SEQ ID NO.79

[0707] Sequence length: 594

[0708] Sequence type: Protein

[0709] Sequence source: metagenomic sequencing

[0710] Sequence name: C2C9 protein sequence 77

[0711] SEQ ID NO.80

[0712] Sequence length: 593

[0713] Sequence type: Protein

[0714] Sequence source: Kitasatospora mediocidica

[0715] Sequence name: C2C9 protein sequence 78

[0716] SEQ ID NO.81

[0717] Sequence length: 493

[0718] Sequence type: Protein

[0719] Sequence source: metagenomic sequencing

[0720] Sequence name: C2C9 protein sequence 79

[0721] SEQ ID NO.82

[0722] Sequence length: 612

[0723] Sequence type: Protein

[0724] Sequence source: metagenomic sequencing

[0725] Sequence name: C2C9 protein sequence 80

[0726] SEQ ID NO.83

[0727] Sequence length: 506

[0728] Sequence type: Protein

[0729] Sequence source: metagenomic sequencing

[0730] Sequence name: C2C9 protein sequence 81

[0731] SEQ ID NO.84

[0732] Sequence length: 607

[0733] Sequence type: Protein

[0734] Sequence source: metagenomic sequencing

[0735] Sequence name: C2C9 protein sequence 82

[0736] SEQ ID NO.85

[0737] Sequence length: 572

[0738] Sequence type: Protein

[0739] Sequence source: metagenomic sequencing

[0740] Sequence name: C2C9 protein sequence 83

[0741] SEQ ID NO.86

[0742] Sequence length: 508

[0743] Sequence type: Protein

[0744] Sequence source: metagenomic sequencing

[0745] Sequence name: C2C9 protein sequence 84

[0746] SEQ ID NO.87

[0747] Sequence length: 542

[0748] Sequence type: Protein

[0749] Sequence source: metagenomic sequencing

[0750] Sequence name: C2C9 protein sequence 85

[0751] SEQ ID NO.88

[0752] Sequence length: 564

[0753] Sequence type: Protein

[0754] Sequence source: metagenomic sequencing

[0755] Sequence name: C2C9 protein sequence 86

[0756] SEQ ID NO.89

[0757] Sequence length: 498

[0758] Sequence type: Protein

[0759] Sequence source: metagenomic sequencing

[0760] Sequence name: C2C9 protein sequence 87

[0761] SEQ ID NO.90

[0762] Sequence length: 538

[0763] Sequence type: Protein

[0764] Sequence source: metagenomic sequencing

[0765] Sequence name: C2C9 protein sequence 88

[0766] SEQ ID NO.91

[0767] Sequence length: 659

[0768] Sequence type: Protein

[0769] Sequence source: Streptomyces sp.

[0770] Sequence name: C2C9 protein sequence 89

[0771] SEQ ID NO.92

[0772] Sequence length: 526

[0773] Sequence type: Protein

[0774] Sequence source: Quadrisphaera granulorum

[0775] Sequence name: C2C9 protein sequence 90

[0776] SEQ ID NO.93

[0777] Sequence length: 549

[0778] Sequence type: Protein

[0779] Sequence source: metagenomic sequencing

[0780] Sequence name: C2C9 protein sequence 91

[0781] SEQ ID NO.94

[0782] Sequence length: 533

[0783] Sequence type: Protein

[0784] Sequence source: metagenomic sequencing

[0785] Sequence name: C2C9 protein sequence 92

[0786] SEQ ID NO.95

[0787] Sequence length: 436

[0788] Sequence type: Protein

[0789] Sequence source: metagenomic sequencing

[0790] Sequence name: C2C9 protein sequence 93

[0791] SEQ ID NO.96

[0792] Sequence length: 483

[0793] Sequence type: Protein

[0794] Sequence source: Prauserella rugosa

[0795] Sequence name: C2C9 protein sequence 94

[0796] SEQ ID NO.97

[0797] Sequence length: 578

[0798] Sequence type: Protein

[0799] Sequence source: Rothia sp. HMSC061D12

[0800] Sequence name: C2C9 protein sequence 95

[0801] SEQ ID NO.98

[0802] Sequence length: 519

[0803] Sequence type: Protein

[0804] Sequence source: metagenomic sequencing

[0805] Sequence name: C2C9 protein sequence 96

[0806] SEQ ID NO.99

[0807] Sequence length: 440

[0808] Sequence type: Protein

[0809] Sequence source: metagenomic sequencing

[0810] Sequence name: C2C9 protein sequence 97

[0811] SEQ ID NO.100

[0812] Sequence length: 495

[0813] Sequence type: Protein

[0814] Sequence source: Mycobacterium grossiae

[0815] Sequence name: C2C9 protein sequence 98

[0816] SEQ ID NO.101

[0817] Sequence length: 494

[0818] Sequence type: Protein

[0819] Sequence source: Mycobacterium avium

[0820] Sequence name: C2C9 protein sequence 99

[0821] SEQ ID NO.102

[0822] Sequence length: 548

[0823] Sequence type: Protein

[0824] Sequence source: Corynebacterium maris

[0825] Sequence name: C2C9 protein sequence 100

[0826] SEQ ID NO.103

[0827] Sequence length: 462

[0828] Sequence type: Protein

[0829] Sequence source: Streptomyces sp. NRRL S-378

[0830] Sequence name: C2C9 protein sequence 101

[0831] SEQ ID NO.104

[0832] Sequence length: 492

[0833] Sequence type: Protein

[0834] Sequence source: Kitasatospora cheerisanensis KCTC 2395

[0835] Sequence name: C2C9 protein sequence 102

[0836] SEQ ID NO.105

[0837] Sequence length: 495

[0838] Sequence type: Protein

[0839] Sequence source: Mycobacterium sp.SWH-M3

[0840] Sequence name: C2C9 protein sequence 103

[0841] SEQ ID NO.106

[0842] Sequence length: 524

[0843] Sequence type: Protein

[0844] Sequence source: Streptomyces roseochromogenus

[0845] Sequence name: C2C9 protein sequence 104

[0846] SEQ ID NO.107

[0847] Sequence length: 513

[0848] Sequence type: Protein

[0849] Sequence source: Mycobacterium intracellulare subsp. yongonense 05-1390

[0850] Sequence name: C2C9 protein sequence 105

[0851] SEQ ID NO.108

[0852] Sequence length: 632

[0853] Sequence type: Protein

[0854] Sequence source: metagenomic sequencing

[0855] Sequence name: C2C9 protein sequence 106

[0856] SEQ ID NO.109

[0857] Sequence length: 495

[0858] Sequence type: Protein

[0859] Sequence source: Mycobacterium sp.JS623

[0860] Sequence name: C2C9 protein sequence 107

[0861] SEQ ID NO.110

[0862] Sequence length: 499

[0863] Sequence type: Protein

[0864] Sequence source: Mycolicibacterium fallax

[0865] Sequence name: C2C9 protein sequence 108

[0866] SEQ ID NO.111

[0867] Sequence length: 449

[0868] Sequence type: Protein

[0869] Sequence source: Nocardiopsis sp. JB363

[0870] Sequence name: C2C9 protein sequence 109

[0871] SEQ ID NO.112

[0872] Sequence length: 726

[0873] Sequence type: Protein

[0874] Sequence source: Rothia sp. HMSC069C03

[0875] Sequence name: C2C9 protein sequence 110

[0876] SEQ ID NO.113

[0877] Sequence length: 494

[0878] Sequence type: Protein

[0879] Sequence source: Mycobacterium avium

[0880] Sequence name: C2C9 protein sequence 111

[0881] SEQ ID NO.114

[0882] Sequence length: 522

[0883] Sequence type: Protein

[0884] Sequence source: metagenomic sequencing

[0885] Sequence name: C2C9 protein sequence 112

[0886] SEQ ID NO.115

[0887] Sequence length: 653

[0888] Sequence type: Protein

[0889] Sequence source: metagenomic sequencing

[0890] Sequence name: C2C9 protein sequence 113

[0891] SEQ ID NO.116

[0892] Sequence length: 427

[0893] Sequence type: Protein

[0894] Sequence source: metagenomic sequencing

[0895] Sequence name: C2C9 protein sequence 114

[0896] SEQ ID NO.117

[0897] Sequence length: 474

[0898] Sequence type: Protein

[0899] Sequence source: metagenomic sequencing

[0900] Sequence name: C2C9 protein sequence 115

[0901] SEQ ID NO.118

[0902] Sequence length: 29

[0903] Sequence type: RNA

[0904] Sequence source: Actinomadura craniellae

[0905] Sequence name: tracr pair sequence of AcC2C9

[0906] SEQ ID NO.119

[0907] Sequence length: 250

[0908] Sequence type: RNA

[0909] Sequence source: Actinomadura craniellae

[0910] Sequence name: AcC2C9 tracr RNA sequence

[0911] SEQ ID NO.120

[0912] Sequence length: 279

[0913] Sequence type: RNA

[0914] Sequence source: Actinomadura craniellae

[0915] Sequence name: Guide RNA backbone sequence of AcC2C9

[0916] SEQ ID NO.121

[0917] Sequence length: 1518

[0918] Sequence type: DNA

[0919] Sequence source: Actinomadura craniellae

[0920] Sequence name: AcC2C9 nuclease base sequence

[0921] SEQ ID NO.122

[0922] Sequence length: 1518

[0923] Sequence type: DNA

[0924] Sequence source: Artificial Sequence

[0925] Sequence Name: Optimized Human Codon Sequence of AcC2C9 Nuclease

[0926] SEQ ID NO.123

[0927] Sequence length: 279

[0928] Sequence type: RNA

[0929] Sequence source: Artificial Sequence

[0930] Sequence name: The gRNA sequence of AcC2C9 used in cell experiments in this example.

[0931] SEQ ID NO.124

[0932] Sequence length: 8198

[0933] Sequence type: DNA

[0934] Sequence source: Artificial Sequence

[0935] Sequence name: p15a-AcC2C9Control

[0936] SEQ ID NO.125

[0937] Sequence length: 1952

[0938] Sequence type: DNA

[0939] Sequence source: Artificial Sequence

[0940] Sequence name: AcC2C9-gRNA-expression

[0941] SEQ ID NO.126

[0942] Sequence length: 6296

[0943] Sequence type: DNA

[0944] Sequence source: Artificial Sequence

[0945] Sequence name: p15a-backbone

[0946] SEQ ID NO.127

[0947] Sequence length: 8201

[0948] Sequence type: DNA

[0949] Sequence source: Artificial Sequence

[0950] Sequence name: p15a-AcC2C9PAM

[0951] SEQ ID NO.128

[0952] Sequence length: 24

[0953] Sequence type: DNA

[0954] Sequence source: Artificial Sequence

[0955] Sequence name: PAM primerF

[0956] SEQ ID NO.129

[0957] Sequence length: 24

[0958] Sequence type: DNA

[0959] Sequence source: Artificial Sequence

[0960] Sequence name: PAM primerR

[0961] SEQ ID NO.130

[0962] Sequence length: 2707

[0963] Sequence type: DNA

[0964] Sequence source: Artificial Sequence

[0965] Sequence name: pUC19-6N-library

[0966] SEQ ID NO.131

[0967] Sequence length: 46

[0968] Sequence type: DNA

[0969] Sequence source: Artificial Sequence

[0970] Sequence name: 6N-F

[0971] SEQ ID NO.132

[0972] Sequence length: 30

[0973] Sequence type: DNA

[0974] Sequence source: Artificial Sequence

[0975] Sequence name: 6N-R

[0976] SEQ ID NO.133

[0977] Sequence length: 51

[0978] Sequence type: DNA

[0979] Sequence source: Artificial Sequence

[0980] Sequence name: PAM-1F

[0981] SEQ ID NO.134

[0982] Sequence length: 43

[0983] Sequence type: DNA

[0984] Sequence source: Artificial Sequence

[0985] Sequence name: PAM-1R

[0986] SEQ ID NO.135

[0987] Sequence length: 66

[0988] Sequence type: DNA

[0989] Sequence source: Artificial Sequence

[0990] Sequence name: PAM-2F

[0991] SEQ ID NO.136

[0992] Sequence length: 63

[0993] Sequence type: DNA

[0994] Sequence source: Artificial Sequence

[0995] Sequence name: PAM-2R

[0996] SEQ ID NO.137

[0997] Sequence length: 8037

[0998] Sequence type: DNA

[0999] Sequence source: Artificial Sequence

[1000] Sequence name: p15a-AcC2C9

[1001] SEQ ID NO.138

[1002] Sequence length: 41

[1003] Sequence type: DNA

[1004] Sequence source: Artificial Sequence

[1005] Sequence name: AcC2C9F

[1006] SEQ ID NO.139

[1007] Sequence length: 36

[1008] Sequence type: DNA

[1009] Sequence source: Artificial Sequence

[1010] Sequence name: ACC2C9R

[1011] SEQ ID NO.140

[1012] Sequence length: 33

[1013] Sequence type: DNA

[1014] Sequence source: Artificial Sequence

[1015] Sequence name: p15aF

[1016] SEQ ID NO.141

[1017] Sequence length: 43

[1018] Sequence type: DNA

[1019] Sequence source: Artificial Sequence

[1020] Sequence name: p15aR

[1021] SEQ ID NO.142

[1022] Sequence length: 5040

[1023] Sequence type: DNA

[1024] Sequence source: Artificial Sequence

[1025] Sequence name: pSGKP-AcC2C9

[1026] SEQ ID NO.143

[1027] Sequence length: 496

[1028] Sequence type: DNA

[1029] Sequence source: Artificial Sequence

[1030] Sequence name: J23119-AcgRNA

[1031] SEQ ID NO.144

[1032] Sequence length: 4594

[1033] Sequence type: DNA

[1034] Sequence source: Artificial Sequence

[1035] Sequence name: pSGKP-backbone

[1036] SEQ ID NO.145

[1037] Sequence length: 5043

[1038] Sequence type: DNA

[1039] Sequence source: Artificial Sequence

[1040] Sequence name: pSGKP-AcC2C9-dha

[1041] SEQ ID NO.146

[1042] Sequence length: 20

[1043] Sequence type: DNA

[1044] Sequence source: Artificial Sequence

[1045] Sequence name: dha-spacer

[1046] SEQ ID NO.147

[1047] Sequence length: 24

[1048] Sequence type: DNA

[1049] Sequence source: Artificial Sequence

[1050] Sequence name: dha-spF

[1051] SEQ ID NO.148

[1052] Sequence length: 24

[1053] Sequence type: DNA

[1054] Sequence source: Artificial Sequence

[1055] Sequence name: dha-spR

[1056] SEQ ID NO.149

[1057] Sequence length: 90

[1058] Sequence type: DNA

[1059] Sequence source: Artificial Sequence

[1060] Sequence name: DHA gene homologous recombination repair template

[1061] SEQ ID NO.150

[1062] Sequence length: 20

[1063] Sequence type: DNA

[1064] Sequence source: Artificial Sequence

[1065] Sequence name: 5' primer for colony PCR verification of the dha gene

[1066] SEQ ID NO.151

[1067] Sequence length: 20

[1068] Sequence type: DNA

[1069] Sequence source: Artificial Sequence

[1070] Sequence name: 3' primers for colony PCR verification of the dha gene

[1071] SEQ ID NO.152

[1072] Sequence length: 7178

[1073] Sequence type: DNA

[1074] Sequence source: Artificial Sequence

[1075] Sequence name: pET28a-sumo-AcC2C9

[1076] SEQ ID NO.153

[1077] Sequence length: 46

[1078] Sequence type: DNA

[1079] Sequence source: Artificial Sequence

[1080] Sequence name: pET-AcC2C9F

[1081] SEQ ID NO.154

[1082] Sequence length: 47

[1083] Sequence type: DNA

[1084] Sequence source: Artificial Sequence

[1085] Sequence name: pET-AcC2C9R

[1086] SEQ ID NO.155

[1087] Sequence length: 5657

[1088] Sequence type: DNA

[1089] Sequence source: Artificial Sequence

[1090] Sequence name: pET28a-sumo-backbone

[1091] SEQ ID NO.156

[1092] Sequence length: 20

[1093] Sequence type: DNA

[1094] Sequence source:

[1095] Sequence name: DNA-substrate-F

[1096] SEQ ID NO.157

[1097] Sequence length: 20

[1098] Sequence type: DNA

[1099] Sequence source: Artificial Sequence

[1100] Sequence name: DNA-substrate-R

[1101] SEQ ID NO.158

[1102] Sequence length: 254

[1103] Sequence type: RNA

[1104] Sequence source: Artificial Sequence

[1105] Sequence name: Complete guide RNA for targeting and cleaving the substrate

[1106] SEQ ID NO.159

[1107] Sequence length: 274

[1108] Sequence type: DNA

[1109] Sequence source: Artificial Sequence

[1110] Sequence name: gRNA-template

[1111] SEQ ID NO.160

[1112] Sequence length: 205

[1113] Sequence type: RNA

[1114] Sequence source: Artificial Sequence

[1115] Sequence name: tracrRNA sequence of AcC2C9

[1116] SEQ ID NO.161

[1117] Sequence length: 225

[1118] Sequence type: DNA

[1119] Sequence source: Artificial Sequence

[1120] Sequence name: Transcription template for AcC2C9 tracrRNA

[1121] SEQ ID NO.162

[1122] Sequence length: 49

[1123] Sequence type: RNA

[1124] Sequence source: Artificial Sequence

[1125] Sequence name: AcC2C9 crRNA sequence targeting substrate cleavage

[1126] SEQ ID NO.163

[1127] Sequence length: 69

[1128] Sequence type: DNA

[1129] Sequence source: Artificial Sequence

[1130] Sequence name: Transcription template for AcC2C9 crRNA

[1131] SEQ ID NO.164

[1132] Sequence length: 58

[1133] Sequence type: DNA

[1134] Sequence source: Artificial Sequence

[1135] Sequence name: 58bp-DNA-substrate F

[1136] SEQ ID NO.165

[1137] Sequence length: 58

[1138] Sequence type: DNA

[1139] Sequence source: Artificial Sequence

[1140] Sequence name: 58bp-DNA-substrate R

[1141] SEQ ID NO.166

[1142] Sequence length: 7204

[1143] Sequence type: DNA

[1144] Sequence source: Artificial Sequence

[1145] Sequence name: pAcC2C9Hs

[1146] SEQ ID NO.167

[1147] Sequence length: 2491

[1148] Sequence type: DNA

[1149] Sequence source: Artificial Sequence

[1150] Sequence name: Human transient expression plasmid backbone

[1151] SEQ ID NO.168

[1152] Sequence length: 1566

[1153] Sequence type: DNA

[1154] Sequence source: Artificial Sequence

[1155] Sequence Name: Puromycin Resistance Gene Expression Cassette Sequence

[1156] SEQ ID NO.169

[1157] Sequence length: 2474

[1158] Sequence type: DNA

[1159] Sequence source: Artificial Sequence

[1160] Sequence name: human-AcC2C9

[1161] SEQ ID NO.170

[1162] Sequence length: 673

[1163] Sequence type: DNA

[1164] Sequence source: Artificial Sequence

[1165] Sequence name: human-AcC2C9-gRNA

[1166] SEQ ID NO.171

[1167] Sequence length: 24

[1168] Sequence type: DNA

[1169] Sequence source: Artificial Sequence

[1170] Sequence name: Guide1-F

[1171] SEQ ID NO.172

[1172] Sequence length: 24

[1173] Sequence type: DNA

[1174] Sequence source: Artificial Sequence

[1175] Sequence name: Guide1-R

[1176] SEQ ID NO.173

[1177] Sequence length: 24

[1178] Sequence type: DNA

[1179] Sequence source: Artificial Sequence

[1180] Sequence name: Guide2-F

[1181] SEQ ID NO.174

[1182] Sequence length: 24

[1183] Sequence type: DNA

[1184] Sequence source: Artificial Sequence

[1185] Sequence name: Guide2-R

[1186] SEQ ID NO.175

[1187] Sequence length: 24

[1188] Sequence type: DNA

[1189] Sequence source: Artificial Sequence

[1190] Sequence name: Guide3-F

[1191] SEQ ID NO.176

[1192] Sequence length: 24

[1193] Sequence type: DNA

[1194] Sequence source: Artificial Sequence

[1195] Sequence name: Guide3-R

[1196] SEQ ID NO.177

[1197] Sequence length: 24

[1198] Sequence type: DNA

[1199] Sequence source: Artificial Sequence

[1200] Sequence name: Guide4-F

[1201] SEQ ID NO.178

[1202] Sequence length: 24

[1203] Sequence type: DNA

[1204] Sequence source: Artificial Sequence

[1205] Sequence name: Guide4-R

[1206] SEQ ID NO.179

[1207] Sequence length: 24

[1208] Sequence type: DNA

[1209] Sequence source: Artificial Sequence

[1210] Sequence name: Guide5-F

[1211] SEQ ID NO.180

[1212] Sequence length: 24

[1213] Sequence type: DNA

[1214] Sequence source: Artificial Sequence

[1215] Sequence name: Guide5-R

[1216] SEQ ID NO.181

[1217] Sequence length: 7198

[1218] Sequence type: DNA

[1219] Sequence source: Artificial Sequence

[1220] Sequence name: pAcC2C9Hs-guide1

[1221] SEQ ID NO.182

[1222] Sequence length: 7198

[1223] Sequence type: DNA

[1224] Sequence source: Artificial Sequence

[1225] Sequence name: pAcC2C9Hs-guide2

[1226] SEQ ID NO.183

[1227] Sequence length: 7198

[1228] Sequence type: DNA

[1229] Sequence source: Artificial Sequence

[1230] Sequence name: pAcC2C9Hs-guide3

[1231] SEQ ID NO.184

[1232] Sequence length: 7198

[1233] Sequence type: DNA

[1234] Sequence source: Artificial Sequence

[1235] Sequence name: pAcC2C9Hs-guide4

[1236] SEQ ID NO.185

[1237] Sequence length: 7198

[1238] Sequence type: DNA

[1239] Sequence source: Artificial Sequence

[1240] Sequence name: pAcC2C9Hs-guide5

[1241] SEQ ID NO.186

[1242] Sequence length: 20

[1243] Sequence type: DNA

[1244] Sequence source: Artificial Sequence

[1245] Sequence name: Guide1-IndelF

[1246] SEQ ID NO.187

[1247] Sequence length: 20

[1248] Sequence type: DNA

[1249] Sequence source: Artificial Sequence

[1250] Sequence name: Guide1-IndelR

[1251] SEQ ID NO.188

[1252] Sequence length: 22

[1253] Sequence type: DNA

[1254] Sequence source: Artificial Sequence

[1255] Sequence name: Guide2-IndelF

[1256] SEQ ID NO.189

[1257] Sequence length: 22

[1258] Sequence type: DNA

[1259] Sequence source: Artificial Sequence

[1260] Sequence name: Guide2-IndelR

[1261] SEQ ID NO.190

[1262] Sequence length: 23

[1263] Sequence type: DNA

[1264] Sequence source: Artificial Sequence

[1265] Sequence name: Guide3-IndelF

[1266] SEQ ID NO.191

[1267] Sequence length: 23

[1268] Sequence type: DNA

[1269] Sequence source: Artificial Sequence

[1270] Sequence name: Guide3-IndelR

[1271] SEQ ID NO.192

[1272] Sequence length: 22

[1273] Sequence type: DNA

[1274] Sequence source: Artificial Sequence

[1275] Sequence name: Guide4-IndelF

[1276] SEQ ID NO.193

[1277] Sequence length: 22

[1278] Sequence type: DNA

[1279] Sequence source: Artificial Sequence

[1280] Sequence name: Guide4-IndelR

[1281] SEQ ID NO.194

[1282] Sequence length: 22

[1283] Sequence type: DNA

[1284] Sequence source: Artificial Sequence

[1285] Sequence name: Guide5-IndelF

[1286] SEQ ID NO.195

[1287] Sequence length: 22

[1288] Sequence type: DNA

[1289] Sequence source: Artificial Sequence

[1290] Sequence name: Guide5-IndelR

[1291] The above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them; although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications can still be made to the specific implementation of the present invention or equivalent substitutions can be made to some technical features; without departing from the spirit and scope of the inventive concept, all changes and advantages that those skilled in the art can conceive of should be covered within the scope of the technical solutions claimed in the present invention. sequence list <110> ShanghaiTech University <120> A novel genome editing system based on C2C9 nuclease and its applications <160> 195 <170> SIPOSequenceListing 1.0 <210> 1 <211> 76 <212> PRT <213> Artificial Sequence <400> 1 Leu His Gln Ile Thr Lys Arg Leu Thr Thr Gly Phe Ala Val Val Ala 1 5 10 15 Leu Glu Asp Leu Asn Val Ala Gly Met Thr Arg Ser Ala Arg Gly Thr 20 25 30 Val Ala Ala Pro Gly Lys Asn Val Arg Gln Lys Ala Gly Leu Asn Arg 35 40 45 Val Ile Leu Asp Ser Ala Pro Ala Glu Leu Arg Arg Gln Val Asn Tyr 50 55 60 Lys Ala Thr Trp Tyr Gly Ser Glu Leu Ala Val Ala 65 70 75 <210> 2 <211> 67 <212> PRT <213> Artificial Sequence <400> 2 Ala Glu Leu Arg Arg Gln Val Asn Tyr Lys Ala Thr Trp Tyr Gly Ser 1 5 10 15 Glu Leu Ala Val Ala Asp Arg Trp Phe Pro Ser Ser Lys Thr Cys Ser 20 25 30 Gly Cys Gly Trp Gln Asn Pro His Leu Lys Leu Ser Asp Arg Val Phe 35 40 45 Arg Cys Thr Asp Cys Gly Leu Val Met Asp Arg Asp Met Asn Ala Ala 50 55 60 Arg Asn Ile 65 <210> 3 <211> 506 <212> PRT <213> Artificial Sequence <400> 3 Met Val Gln Thr Glu Ile Leu Lys Ala Phe Arg Phe Ala Leu Asp Pro 1 5 10 15 Thr Ser Val Gln Val Ala Ala Leu Ser Arg His Ala Gly Ala Ala Arg 20 25 30 Trp Ala Phe Asn His Ala Leu Ala Ala Lys Val Gly Ala His Glu Arg 35 40 45 Trp Arg Ala Glu Val Ala Lys Leu Val Gly Asp Gly Val Pro Glu Glu 50 55 60 Gln Ala Arg Arg Gln Val Arg Val Pro Val Pro Met Lys Pro Ala Ile 65 70 75 80 Gln Lys Ala Leu Asn Ala Val Lys Gly Asp Ser Arg Lys Gly Leu Asp 85 90 95 Gly Ala Cys Pro Trp Trp His Glu Val Asn Thr Tyr Ala Phe Gln Ser 100 105 110 Ala Phe Ile Asp Ala Asp Gln Ala Trp Lys Asn Trp Leu Asp Ser Leu 115 120 125 Ser Gly Lys Arg Val Gly Arg Arg Val Gly Tyr Pro Arg Phe Lys Lys 130 135 140 Lys Gly Arg Ala Arg Asp Ser Phe Arg Leu His His Asp Val Lys Lys 145 150 155 160 Pro Ser Ile Arg Leu Ala Gly Tyr Arg Arg Leu Arg Leu Pro Arg Ile 165 170 175 Gly Glu Val Arg Leu His Asp Ser Gly Lys Arg Leu Ala Arg Leu Ile 180 185 190 Asp Arg Gly Asp Ala Val Val Gln Ser Val Thr Val Ser Arg Gly Gly 195 200 205 His Arg Trp Tyr Ala Ser Val Leu Cys Lys Val Thr Val Gln Val Pro 210 215 220 Asp Arg Pro Ser Arg Arg Gln Arg Glu Arg Gly Ala Val Gly Val Asp 225 230 235 240 Leu Gly Val Lys Val Leu Ala Ala Leu Ser Lys Pro Leu Val Val Asp 245 250 255 Asp Pro Ser Ser Ala Leu Val Arg Asn Pro Gln His Leu Arg Gln Ala 260 265 270 Glu Arg Arg Leu Leu Lys Ala Gln Arg Ala Leu Ala Arg Thr Gln Lys 275 280 285 Gly Ser Ala Arg Arg Glu Lys Ala Lys Arg Arg Val Gly Arg Ala His 290 295 300 His Glu Val Ala Val Arg Arg His Ala Ala Leu His Gln Ile Thr Lys 305 310 315 320 Arg Leu Thr Thr Gly Phe Ala Val Val Ala Leu Glu Asp Leu Asn Val 325 330 335 Ala Gly Met Thr Arg Ser Ala Arg Gly Thr Val Ala Ala Pro Gly Lys 340 345 350 Asn Val Arg Gln Lys Ala Gly Leu Asn Arg Val Ile Leu Asp Ser Ala 355 360 365 Pro Ala Glu Leu Arg Arg Gln Val Asn Tyr Lys Ala Thr Trp Tyr Gly 370 375 380 Ser Glu Leu Ala Val Ala Asp Arg Trp Phe Pro Ser Ser Lys Thr Cys 385 390 395 400 Ser Gly Cys Gly Trp Gln Asn Pro His Leu Lys Leu Ser Asp Arg Val 405 410 415 Phe Arg Cys Thr Asp Cys Gly Leu Val Met Asp Arg Asp Met Asn Ala 420 425 430 Ala Arg Asn Ile Glu Arg His Ala Val Leu Ile Asp Arg Asn Val Ala 435 440 445 Cys Asp Arg Arg Glu Thr Leu Asn Ala Arg Gly Ala Ser Ile Arg Pro 450 455 460 Thr Thr Ser Arg Gly Gly Arg His Asp Ala Ser Lys Arg Glu Gly Ser 465 470 475 480 Gly Ala Arg Pro Glu Ser Pro Gln Gln Ser Asp Leu Leu Thr Leu Pro 485 490 495 Thr Leu Asn Asn Glu Thr Ala Thr Pro Pro 500 505 <210> 4 <211> 493 <212> PRT <213> Artificial Sequence <400> 4 Met Ser Gly Ala Ile Gln Arg Ser His Ala Thr Glu Thr Arg Glu Val 1 5 10 15 Leu Arg Ala Tyr Arg Val Thr Leu Asp Pro Thr Pro Asp Gln Glu Ser 20 25 30 Met Leu Ala Gln His Ala Gly Ala Ala Arg Trp Ala Phe Asn His Ala 35 40 45 Leu Ala Ala Lys Val Asp Ala His Glu Arg Trp Lys Thr Ala Val Ala 50 55 60 Glu Leu Val Ala Ala Gly Val Pro Glu Glu Ala Ala Arg Lys Lys Ile 65 70 75 80 Lys Val Leu Val Pro Gly Lys Gln Val Ile Gln Lys Ala Leu Asn Ala 85 90 95 Val Lys Gly Asp Asp Arg Lys Gly Ile Asp Gly Ala Cys Pro Trp Trp 100 105 110 His Thr Val Ser Thr Tyr Ala Phe Gln Ser Ala Phe Leu Asp Ala Asp 115 120 125 Met Ala Trp Lys Ala Trp Leu Asp Ser Leu Thr Gly Arg Arg Lys Gly 130 135 140 Arg Arg Val Gly Tyr Pro Arg Phe Lys Lys Lys Gly Arg Ser Arg Asp 145 150 155 160 Ser Phe Arg Leu His His Asp Val Lys Arg Pro Thr Ile Arg Pro Ala 165 170 175 Thr Thr Arg Arg Leu Leu Leu Pro Arg Ile Gly Glu Val Arg Thr His 180 185 190 Asp Ser Leu Lys Arg Leu Lys Ala Ala Leu Gly Arg Arg Ser Gly Val 195 200 205 Ile Gln Ser Val Thr Ile Ala Arg Gly Gly His Arg Trp Tyr Ala Ser 210 215 220 Val Leu Val Arg Glu Gln Met Leu Met Pro Ala Pro Thr Arg Ala Gln 225 230 235 240 Gln Ala Ala Gly Thr Val Gly Val Asp Leu Gly Val His Ser Met Ala 245 250 255 Ala Leu Ser Thr Gly Glu Ile Ile Pro Asn Pro Arg Val Lys Ala Arg 260 265 270 His Ala Pro Ala Leu Ala Lys Ala Ser Arg Ala Tyr Ala Arg Thr Val 275 280 285 Lys Gly Ser Arg Arg Arg Glu Gln Ala Arg Arg Arg Leu Ala Arg Leu 290 295 300 His His Arg Glu Ala Thr Arg Arg Ser Thr Gly Leu His His Leu Thr 305 310 315 320 Lys Arg Leu Ala Thr Gly Trp Ala Val Val Ala Val Glu Asp Leu Asp 325 330 335 Val Ala Gly Leu Ile Arg Ser Ala Lys Gly Thr Val Asp Lys Pro Gly 340 345 350 Arg Arg Val Arg Gln Lys Ser Gly Leu Asn Arg Ser Ile Leu Asp Val 355 360 365 Ser Pro Gly Glu Val Arg Arg Gln Leu Asp Tyr Lys Thr Gly Trp Tyr 370 375 380 Gly Ser Thr Ile Ala Ile Leu Asp Arg Phe Tyr Pro Ser Ser Lys Thr 385 390 395 400 Cys Ser Gly Cys Gly Gly Arg Lys Pro Asn Leu Thr Leu Ser Glu Arg 405 410 415 Val Phe Ile Cys Thr Ile Cys Gly Leu Ser Met Asp Cys Asp Val Asn 420 425 430 Ala Ala Val Asn Ile Ala Arg Phe Ala Val Ala Ser Asp Arg Gly Glu 435 440 445 Thr Leu Asn Ala Arg Gly Ala Ser Val Arg Pro Ser His Leu Gln Val 450 455 460 Gly Arg Pro Asp Ala Leu Met Arg Glu Asp Pro Pro Glu Gly Gly Pro 465 470 475 480 Pro Gln His Ser Asp Ala Leu Ala Ser Ser Thr Pro Thr 485 490 <210> 5 <211> 537 <212> PRT <213> Artificial Sequence <400> 5 Met Thr Ser Thr Thr Leu Ala Pro Glu Glu Pro Leu Met Val Phe Arg 1 5 10 15 Gly Ala Arg Phe Arg Leu Asp Pro Thr Gly Glu Gln Gln Gly Ile Leu 20 25 30 Ser Gln Gln Ala Gly Ala Ala Arg Val Ala Tyr Asn Met Met Cys Thr 35 40 45 Leu Asn Lys Asp Ile Leu Glu Ala Arg Ser Gln Leu Tyr Ser Thr Leu 50 55 60 Ile Lys Asp Gly Lys Thr Lys Asp Lys Ala Lys Lys Glu Leu Lys Ala 65 70 75 80 Ala Ala Lys Glu Asp Pro Ser Leu Ala Ile Val Trp Ala Arg Asp Phe 85 90 95 Asp Lys Asn Tyr Ile Thr Pro Glu Arg Asn Arg His Lys His Ala Ala 100 105 110 Gln Arg Ile Ala Ala Gly Glu Asn Pro Val Asp Val Trp Asn Pro Asp 115 120 125 Glu Glu Arg Phe Asn Glu Pro Trp Leu His Thr Ala Asn Arg Arg Val 130 135 140 Leu Arg Ser Gly Gln Lys Gln Tyr Glu Gln Ala Leu Asp Asn Phe Phe 145 150 155 160 Lys Ser Gln Asn Gly Ser Arg Ala Gly Gln Lys Met Gly Lys Pro Arg 165 170 175 Phe Lys Thr Lys Ile Arg Ser Thr Asp Ser Phe Thr Ile Asp Ala Val 180 185 190 Asp Val Ser Ser Ser Thr Thr Leu Ile Arg Asp Ile Gly Pro Lys Asp 195 200 205 His Ala Arg Tyr Lys Thr Gly Glu Ala Ser Thr Gly Ile Ile Ala Asp 210 215 220 Tyr Arg His Val Arg Leu Ser His Leu Gly Thr Phe Arg Val Phe Gly 225 230 235 240 Ser Thr Lys Ala Leu Val Arg Gln Leu Asp Arg Gly Gly Arg Ile Lys 245 250 255 Ser Cys Thr Val Ser Arg Ser Ala Asp Arg Trp Tyr Val Ser Phe Leu 260 265 270 Val Glu Leu Pro Ile Glu Ile Ala Arg Ser Thr Pro Thr Lys Lys Gln 275 280 285 Tyr Lys Ala Gly Ala Val Gly Ile Asp Leu Gly Val Lys Ser Leu Ala 290 295 300 Ala Leu Ser Thr Gly Glu Ile Ile Pro Asn Pro Arg Phe Leu Arg Thr 305 310 315 320 Ala Asp Lys Lys Ile Lys Lys Leu Gln Arg Lys Ile Ala Arg Cys Gln 325 330 335 Lys Gly Ser Lys Asn Ser Ile Arg Leu Lys Arg Arg Leu Ala Arg Cys 340 345 350 His His Glu Leu Ala Leu Gln Arg Ala Gly Tyr Leu Asn Glu Leu Thr 355 360 365 Ser Met Leu Ala Ser Ser Phe Ser Ala Ile Ala Leu Glu Asp Leu Asn 370 375 380 Val Ala Gly Met Thr Ser Ser Ala Arg Gly Thr Val Glu Asn Pro Gly 385 390 395 400 Lys Asn Val Lys Gln Lys Ala Gly Leu Asn Arg Ser Ile Leu Asp Ile 405 410 415 Ser Pro Gly Arg Ile Arg Thr Leu Leu Glu Tyr Lys Cys Thr Asp Arg 420 425 430 Gly Val Glu Leu Gln Val Ile Asp Arg Phe Phe Pro Ser Ser Gln Leu 435 440 445 Cys Ser Ser Cys Gly Ser Lys Thr Thr Ile Pro Leu Ala Gln Arg Ile 450 455 460 Tyr His Cys Asp Val Cys Gly Gly Val Ile Asp Arg Asp Val Asn Ala 465 470 475 480 Ala Ile Asn Ile Val Tyr Glu Ala Lys Arg Leu Val Glu Gln Lys Cys 485 490 495 Ser Glu His Ser Ala Pro Glu Gly Ala Glu Asp Lys Arg Pro Trp Ser 500 505 510 Val His Leu Pro Ser Thr Gln Tyr Val Asp Gly Leu Tyr Thr Arg Lys 515 520 525 Arg Gln Gly Pro Ser Gly His Ser Ser 530 535 <210> 6 <211> 537 <212> PRT <213> Artificial Sequence <400> 6 Met Thr Thr Arg Glu Ile Leu Arg Ala Tyr Arg Val Pro Leu Asp Pro 1 5 10 15 Thr Asp Ala Gln Thr Ala Ala Leu Ala Ser His Ala Gly Ala Ser Arg 20 25 30 Ala Ala Phe Asn Trp Ala Leu Gly Ala Lys Val His Ala His Arg Met 35 40 45 Trp Ser Ala Cys Val Ala Asp Leu Thr Tyr Thr Arg Tyr Gly His Leu 50 55 60 Asp Ala Asp Gln Ala Leu Ala Ala Ala Lys Lys Asp Ala Ser Arg Tyr 65 70 75 80 Tyr Arg Ile Pro Thr Ser Gln Thr Asn Glu Lys Ala Phe Asp Arg Asp 85 90 95 Pro Asp Tyr Ala Trp Arg Thr Glu Val Asn Arg Arg Ser Trp Val Ser 100 105 110 Gly Met Arg Gln Ala Asp Thr Ala Trp Gln Asn Trp Leu Asp Ser Leu 115 120 125 Thr Gly Arg Arg Ala Gly Arg Arg Val Gly Tyr Pro Arg Phe Lys Ser 130 135 140 Lys Gly Arg Cys Arg Asp Ser Phe Thr Leu Ala His Asp Val Lys Arg 145 150 155 160 Pro Ser Ile Arg Pro Asp Gly Tyr Arg Arg Leu Thr Leu Pro Lys Lys 165 170 175 Ile Ser Val Thr Gly Ser Ile Arg Leu Lys Gly Asn Ile Arg His Leu 180 185 190 Ala Arg Arg Ile Arg Arg Gly Val Ala Arg Ile Gln Ser Ala Thr Ile 195 200 205 Ser Arg Ala Gly Asn Gly Trp Ser Val Ser Ile Leu Ala Leu Glu Thr 210 215 220 Leu Asp Ile Pro Asp His Pro Thr Pro Arg Gln Gln Ala Ala Gly Ala 225 230 235 240 Val Gly Val Asp Val Gly Val His His Leu Met Ala Phe Ser Asp Ala 245 250 255 Thr Ile Ile Asp Asn Pro Arg His Leu Arg Ala Ala Gln Lys Arg Leu 260 265 270 Thr Arg Ala Gln Arg Ala Leu Ser Arg Ser Lys Trp Arg Leu Pro Asn 275 280 285 Gly Asp Leu Ile Asp Thr Pro Lys Arg Gly Gln Arg Val Thr Pro Thr 290 295 300 Thr Gly Arg Val Lys Ala Arg Ala Arg Leu Ala Arg Glu His Ala Ala 305 310 315 320 Val Ala Gln His Arg Ala Ser Thr Leu His Ala Ile Thr Lys Gln Leu 325 330 335 Ala Thr Ser His Ala Val Val Ala Val Asp Asp Leu Asn Val Ala Val 340 345 350 Met Thr Arg Ser Ala Arg Gly Ser Ile Asp Lys Pro Gly Arg Asn Val 355 360 365 Ala Ala Lys Ala Gly Leu Asn Arg Ser Ile Leu Asp Ala Ser Phe Ala 370 375 380 Glu Met Arg Arg Gln Leu Thr His Lys Thr Ser Trp Tyr Arg Ser Gln 385 390 395 400 Leu Leu Pro Ser Gly Cys Phe Val Pro Thr Ser Arg Thr Cys Ser Thr 405 410 415 Cys Gly Ala Glu Lys Ala Asn Leu Pro Leu Ser Glu Arg Val Tyr His 420 425 430 Cys Glu Asn Cys Ala Thr Val Leu Asp Arg Asp Val Asn Ala Ala Lys 435 440 445 Asn Val Leu Arg Ala Ala Leu Ala Ser His Asp Ala Pro Gly Met Glu 450 455 460 Glu Ser Gln Asn Ala Arg Gly Gly Arg Gly Glu Thr Ser Val Ala Arg 465 470 475 480 Arg Ser Ala Lys Arg Glu Asp Pro Pro Arg Gly Gly Pro Pro Arg Pro 485 490 495 Cys Asn Arg Val Arg Ser Ser Pro Pro Asp Gly Glu Ala Lys Val Arg 500 505 510 Ala Asn Ala Ser Pro Thr Asn Cys Lys Ala Arg Lys His Ser Thr Pro 515 520 525 Gln Gly Ser Ala Pro Arg Thr Arg Arg 530 535 <210> 7 <211> 492 <212> PRT <213> Artificial Sequence <400> 7 Val Ala Gln Val Glu Thr Leu Arg Ala Tyr Arg Phe Ala Leu Asp Pro 1 5 10 15 Thr Ala Ala Gln Leu Ala Asp Leu Asn Arg His Ala Gly Ala Ala Arg 20 25 30 Trp Ala Phe Asn His Ala Leu Gly Val Lys Val Ala Ala His Arg Gln 35 40 45 Trp Arg Ala Gln Val Asp Ala Leu Val Ser Gly Gly Met Pro Glu Ala 50 55 60 Gln Ala Arg Lys Ser Val Lys Val Pro Val Pro Gly Ser Gln Ala Ile 65 70 75 80 Lys Lys Ala Leu Asn Thr Leu Lys Gly Asp Ser Arg Ala Glu Thr Glu 85 90 95 Leu Pro Asp Gly Phe His Gly Leu His Arg Pro Cys Pro Trp Trp His 100 105 110 Glu Val Ser Thr Tyr Ala Phe Gln Ser Ala Phe Ile Asp Ala Asp Arg 115 120 125 Ala Trp Gly Asn Trp Leu Asp Ser Leu Thr Gly Arg Arg Ala Gly Gly 130 135 140 Pro Val Gly Tyr Pro Arg Phe Lys Arg Lys Gly Arg Ala Arg Asp Ser 145 150 155 160 Phe Arg Leu His His Asp Val Lys Lys Pro Thr Ile Arg Leu Asp Gly 165 170 175 Tyr Arg Arg Leu Gln Leu Pro Arg Leu Gly Ser Ile Arg Val His Asp 180 185 190 Ser Gly Lys Arg Leu Ala Arg Leu Val Ala Lys Gly Gln Ala Val Ile 195 200 205 Gln Ser Val Thr Val Ser Arg Gly Gly Asn Arg Trp Tyr Ala Ser Val 210 215 220 Leu Ala Lys Val Val Gln Glu Val Pro Gly Lys Pro Thr Arg Ala Gln 225 230 235 240 Thr Gly Arg Gly Thr Val Gly Val Asp Trp Gly Ile Ser His Leu Ala 245 250 255 Ser Leu Ser Gln Pro Leu Asp Pro Ala Asp Pro Gly Ser Ile His Ile 260 265 270 Ala Asn Pro Arg His Leu Asp Gln Trp Thr Arg Gln Leu Ala Lys Ala 275 280 285 Gln Arg Ala Leu Ser Arg Thr Glu Arg Gly Ser Arg Arg Arg Val Lys 290 295 300 Ala Ala Arg Lys Val Gly Ala Ile Gln His Arg Ile Ala Gln Arg Arg 305 310 315 320 Ala Ser Thr Val His Leu Leu Thr Lys Lys Leu Ala Thr Gly Phe Ala 325 330 335 Thr Val Ala Val Glu Asp Leu Asn Val Arg Gly Met Ser Ala Ser Ala 340 345 350 Arg Gly Thr Val Glu Lys Pro Gly Arg Arg Val Arg Gln Lys Ala Gly 355 360 365 Leu Asn Arg Ala Ile Leu Asp Ala Ala Pro Gly Glu Leu Arg Arg Gln 370 375 380 Leu Thr Tyr Lys Thr Ser Trp Tyr Gly Ser Ala Leu Ala Val Leu Asp 385 390 395 400 Arg Trp His Pro Ser Ser Lys Thr Cys Ser Ser Cys Gly Thr Ala Arg 405 410 415 Pro Lys Leu Thr Leu Ser Glu Arg Val Phe His Cys Thr Thr Cys Gly 420 425 430 Leu Ala Ile Asp Arg Asp His Asn Ala Ala Ile Asn Ile Ala Arg His 435 440 445 Ala Ala Val Pro Leu Val Glu Gly Asp Val Asn Ala Arg Arg Ser Pro 450 455 460 Pro Pro Ala Asp Asn Gly Arg His Gly Arg Ala Ala Arg Gln Lys Arg 465 470 475 480 Glu Gly Pro Pro Pro Gly Gly Pro Pro Arg Arg Glu 485 490 <210> 8 <211> 553 <212> PRT <213> Artificial Sequence <400> 8 Met Ser Glu Thr Thr Ala Ser Ser Asp Val Arg Ala Tyr Lys Phe Arg 1 5 10 15 Leu Asp Pro Asn Lys Ala Gln Val Val Arg Leu Ser Gln Cys Ala Gly 20 25 30 Ala Ala Arg Val Gly Tyr Asn Met Leu Ile Ala His Asn Arg Glu Ile 35 40 45 His Gln Glu Ala Arg Arg Arg Arg Asp Ala Leu Ile Ala Ser Gly Val 50 55 60 Asp Pro Ala Ala Ala Asp Ala Lys Leu Arg Glu Gln Arg Arg Thr Asp 65 70 75 80 Pro Ala Val Arg Met Leu Ser Tyr Gln Gln Phe Ser Thr Thr His Leu 85 90 95 Thr Pro Glu Ile Ala Arg His Arg Ala Ala Ser Asp Ala Ile Lys Ala 100 105 110 Gly Ala Asp Pro Ala Glu Val Trp Ala Asp Glu Arg Tyr Ala Glu Pro 115 120 125 Trp Leu His Thr Val Pro Arg Arg Val Leu Val Ser Gly Leu His Asn 130 135 140 Ala Trp Asp Ala Phe Thr Asn Trp Met Thr Ser Phe Thr Gly Gln Arg 145 150 155 160 Ala Gly Arg Arg Val Gly Tyr Pro Arg Phe Lys Arg Lys Gly Arg Ser 165 170 175 Arg Asp Ser Phe Thr Ile Pro Ala Glu Thr Thr Gly Gly Lys Gly Thr 180 185 190 Arg Tyr His His Ser Asp Pro His Lys Gly Val Ile Gly Asp Tyr Lys 195 200 205 His Leu Arg Leu Gly Pro Leu Gly Thr Ile Arg Thr His Gln His Thr 210 215 220 Arg Arg Leu Val Arg Ala Cys Glu His Gly Gly Arg Ile Thr Ser Phe 225 230 235 240 Thr Val Ser Arg Ala Ala Asp Tyr Trp Tyr Val Ser Val Leu Val Glu 245 250 255 Thr Pro Ala Ala Glu Pro Pro Ala Pro Ser Arg Ala Gln Arg Ser Arg 260 265 270 Gly Ala Val Gly Val Asp Val Gly Val His Ser Leu Ala Ala Leu Ser 275 280 285 Thr Gly Glu Leu Ile Ala Asn Pro Arg His Gly Gln Ala Thr Arg Glu 290 295 300 Lys Leu Thr Arg Ala Gln Arg Ala Phe Ala Arg Thr Glu Lys Gly Ser 305 310 315 320 Arg Gly Arg Gln Arg Ala Ala Val Arg Val Ala Arg Ile Gln His Lys 325 330 335 Leu Ala Leu Gln Arg Ala Thr Thr Leu His Lys Leu Thr Lys Ala Leu 340 345 350 Ala Thr Gly Tyr Glu Thr Val Ala Ile Glu Asp Leu Asn Val Ala Gly 355 360 365 Met Thr Arg Ser Ala Lys Gly Thr Val Glu Asn Pro Gly Lys Asn Val 370 375 380 Ala Ala Lys Ala Gly Leu Asn Arg Ser Ile Leu Asp Ala Ala Leu Gly 385 390 395 400 Glu Leu Arg Arg Gln Leu Asp Tyr Lys Thr Arg Trp His Gly Ser Thr 405 410 415 Val Thr Thr Ile Gly Arg Tyr Asp Pro Ser Ser Lys Ala Cys Ser Ser 420 425 430 Cys Gly Thr Val Lys Pro Lys Leu Pro Leu Ser Val Arg Val Tyr Asp 435 440 445 Cys Asp Gln Cys Gly Met Ser Leu Asp Arg Asp Val Asn Ala Ala Arg 450 455 460 Asn Ile Leu Ala Trp Ala Ser Asn Gly Arg Asp Ser Phe Thr Val Pro 465 470 475 480 Gln Thr Glu Val Gly Ala Asp Gly Arg Gly Gly Pro Thr Pro Pro Ser 485 490 495 Pro Arg Gly Glu Gly Thr Leu Val Arg Asp Gln Arg Ser Val Lys Thr 500 505 510 Gly Arg Arg Arg Arg Pro Ala Thr Ser Ala Glu Gln Ser Ala Gly His 515 520 525 Leu Ala Thr Pro Arg Glu Gln Thr Ala Pro Pro Glu Arg Glu Thr Pro 530 535 540 Ser Ser Gly Ala Ser Leu His Val Arg 545 550 <210> 9 <211> 539 <212> PRT <213> Artificial Sequence <400> 9 Met Ser Gln Asp Lys Lys Asp Glu Thr Val Asp His Arg Ala Phe Lys 1 5 10 15 Phe Arg Leu Asp Pro Thr Asp Ala Gln Leu Ser Leu Leu Ala Gln Ser 20 25 30 Ala Gly Ala Ala Arg Val Gly Tyr Asn Met Leu Leu Gly His Asn Val 35 40 45 Ala Ala Tyr Gln Ala Arg Gln Glu Leu His Ala Ser Leu Leu Asp Ser 50 55 60 Gly His Asn Glu His Asp Ala Lys Ser Ala Val Thr Asp Arg Lys Thr 65 70 75 80 His Asp Pro Ser Leu Gln Thr Leu Ser Tyr Gln Ser Phe Ser Thr His 85 90 95 Tyr Leu Thr Pro Glu Ile Thr Arg His Arg Ile Ala Ser Asp Ala Ile 100 105 110 Lys Ala Gly Ala Asp Pro Ser Lys Val Trp Asn Asp Pro Arg Tyr Glu 115 120 125 Thr Pro Trp Met His Thr Ile Pro Arg Arg Val Phe Ile Ser Gly Leu 130 135 140 Gln His Ala Asn Thr Ala Tyr Thr Asn Trp Phe Ala Ser Met Arg Gly 145 150 155 160 Asp Arg Ala Gly Ala Arg Val Gly Arg Pro Arg Phe Lys Lys Lys Gly 165 170 175 Arg Cys Arg Asp Ser Phe Thr Ile Pro Ala Pro Glu Ala Met Gly Ala 180 185 190 Lys Gly Ala Pro Tyr Lys Arg Gly Glu Pro Arg Ser Gly Val Ile Glu 195 200 205 Asp Tyr Arg His Val Arg Leu Ala Phe Leu Gly Val Leu Arg Thr His 210 215 220 Asn Ser Thr Lys Arg Leu Val Arg Ala Cys Gln Arg Gly Gly Lys Ile 225 230 235 240 Lys Ser Phe Thr Val Ser Arg Asn Ala Asp Arg Trp Tyr Val Ser Phe 245 250 255 Leu Val Ala Ser Pro Ala Arg Ala Ala Val Thr Thr Lys Arg Gln Arg 260 265 270 Ala Asn Gly Ala Val Gly Val Asp Val Gly Val His Ser Leu Ala Ala 275 280 285 Leu Ser Asn Gly Glu Val Phe Ser Asn Pro Arg His Gly Gln Leu Ala 290 295 300 Gln Lys Lys Leu Lys Arg Ala Gln Arg Lys Leu Ala Arg Thr Gln Lys 305 310 315 320 Gly Ser Lys Gly Arg Ala Arg Ala Ala Gln Arg Val Gly Arg Leu Gln 325 330 335 His Leu Val Thr Leu Gln Arg Glu Thr Thr Ala His His Leu Thr Lys 340 345 350 Tyr Leu Ala Thr His Tyr Cys Ala Val Ala Ile Glu Asp Leu Asn Val 355 360 365 Ser Gly Met Thr Arg Ser Ala Lys Gly Thr Met Glu His Pro Gly Lys 370 375 380 Asn Val Lys Ala Lys Ser Gly Leu Asn Arg Ser Ile Leu Asp Ala Ser 385 390 395 400 Phe Gly Arg Ile His Gln Gln Leu Arg Tyr Lys Met Gly Trp Ser Gly 405 410 415 Ala Asp Leu Gln Ile Ile Gly Arg Phe Val Pro Ser Ser Lys Met Cys 420 425 430 Ser Glu Cys Gly Thr Val Lys Ser Lys Leu Pro Leu Ser Glu Arg Val 435 440 445 Tyr Glu Cys Glu His Cys Gly Leu Lys Ile Asp Arg Asp Val Asn Ala 450 455 460 Ala Ile Asn Ile Leu His Ala Gly Ala Gly Ser Ile Thr Pro Ser Asn 465 470 475 480 Ser Pro Leu Thr Arg Gly Ser Leu Asn Gly Arg Gly Asp Gln Thr Cys 485 490 495 Leu Ser Pro Ser Gly Glu Arg Ile Thr Arg Gly Asp Ser Arg Ser Val 500 505 510 Lys Ala Asp Pro Ile Trp Cys Arg Ser Pro Gln Thr Ser Asn Arg Cys 515 520 525 Ser Ser Leu Thr Val Arg Gly Thr Ala Leu Lys 530 535 <210> 10 <211> 552 <212> PRT <213> Artificial Sequence <400> 10 Met Ser Val Ala Thr Asp Ala Pro Thr Gln Arg Leu Arg Ala Phe Lys 1 5 10 15 His Arg Leu Asp Pro Asn Ala Ser Gln Leu Glu Leu Leu Gly Gln Tyr 20 25 30 Ala Gly Ala Ala Arg Val Ala Phe Asn Met Leu Thr Ala His Asn Arg 35 40 45 Ala Ala Leu Gln Ala Gly Trp Asp Arg Arg Arg Gln Leu Ala Glu Ala 50 55 60 Gly Val Pro Ala Glu Glu Leu Val Gly Arg Met Lys Ala Glu Arg Ala 65 70 75 80 Ala Asp Pro Ala Leu Lys Val Ala Gly Tyr Gln Gln Phe Ala Thr Ala 85 90 95 His Leu Thr Pro Met Val Arg Ala His Arg Glu Ala Ala Glu Ala Ile 100 105 110 Ala Ala Gly Ala Asp Pro Gly Gln Val Trp Ala Asp Glu Arg Tyr Ala 115 120 125 Gln Pro Trp Met His Thr Val Pro Arg Arg Val Leu Val Ser Gly Leu 130 135 140 Gln Asn Ala Ala Lys Ala Thr Glu Asn Trp Met Ala Ser Ala Cys Gly 145 150 155 160 Ala Arg Ala Gly Arg Arg Val Gly Leu Pro Arg Phe Lys Lys Lys Gly 165 170 175 Arg Ser Arg Asp Ser Phe Thr Ile Pro Ala Pro Glu Val Ile Gly Ala 180 185 190 Ala Gly Thr Ala Tyr Lys Arg Gly Glu Ala Arg Gly Gly Val Ile Thr 195 200 205 Asp His Arg His Leu Arg Leu Ala Ser Leu Gly Thr Ile Arg Thr Met 210 215 220 Asp Lys Thr Thr Arg Leu Val Arg Ser Cys Arg Arg Gly Ala Val Val 225 230 235 240 Arg Ser Val Thr Ile Ser Gln Ala Gly Gly His Trp Tyr Ala Ser Val 245 250 255 Leu Val Ala Glu Pro Val Thr Leu Arg Arg Gly Pro Ser Arg Arg Gln 260 265 270 Arg Ala Asn Gly Val Val Gly Val Asp Leu Gly Val Lys His Leu Ala 275 280 285 Ala Leu Ser Thr Gly Glu Leu Ile Pro Asn Ala Arg Ala Gly Gln Ala 290 295 300 Gln Ala Arg Arg Leu Ala Arg Leu Gln Arg Ala Leu Ala Arg Thr Glu 305 310 315 320 Arg Gly Ser Arg Arg Arg Glu Arg Val Arg Arg Gln Ile Ser Ala Leu 325 330 335 His His Gln Val Ala Leu Arg Arg Thr Gly Met Leu His Glu Val Ser 340 345 350 Thr Arg Leu Ala Arg Asp Tyr Ala Ile Val Ala Leu Glu Asp Leu Asn 355 360 365 Val Ala Gly Met Thr Gly Ser Ala Ala Gly Thr Ile Glu His Pro Gly 370 375 380 Lys Asn Val Ala Ala Lys Ala Gly Leu Asn Arg Ala Ile Leu Asp Ala 385 390 395 400 Gly Leu Gly Thr Leu Arg Ala Gln Leu His Tyr Lys Thr Ser Trp Ala 405 410 415 Gly Ser Gln Val Lys Met Ile Asp Arg Phe Ala Pro Ser Ser Lys Thr 420 425 430 Cys Ser Arg Cys Gly Ala Val Lys Ala Thr Leu Ser Leu Arg Glu Arg 435 440 445 Val Tyr Glu Cys Glu Val Cys Ala Leu Val Ile Asp Arg Asp Val Asn 450 455 460 Ala Ala Ile Asn Ile Arg Ala Trp Ala Val Gln Glu Gly Thr Gly His 465 470 475 480 Arg Val Glu Leu Ala Gln Gly Asn Gly Glu Ser Gln Asn Gly Arg Arg 485 490 495 Ala Ala Gly Arg Thr Pro Pro Ser Gly Gly Pro Thr Gly Trp Gln Arg 500 505 510 Arg Ser Val Lys Pro Ala Pro Arg Gly Ala Gly Arg Ser Ser Arg Ala 515 520 525 Thr Gly Trp Ser Ser His Thr Pro His Ile Glu Arg Glu His Ala Glu 530 535 540 Ala Asp Val Phe Gly Pro Val Arg 545 550 <210> 11 <211> 485 <212> PRT <213> Artificial Sequence <400> 11 Met Val Arg Ser Val Gly Asp Asp Leu Arg Ala Tyr Arg Phe Ala Leu 1 5 10 15 Asp Leu Arg Pro Ala Gln Leu Arg Ala Val Ala Glu His Ala Gly Ala 20 25 30 Ala Arg Trp Ala Tyr Asn His Ala Leu Gly Val Lys Phe Ala Ala Leu 35 40 45 Lys Gln Arg Gln Thr Val Ile Ala Glu Leu Val Thr Ala Gly Val Asp 50 55 60 Ala Gln Ala Ala Ala Arg Arg Ala Pro Arg Ile Pro Thr Arg Pro Gln 65 70 75 80 Ile Gln Lys Asp Leu Asn Ala Thr Lys Gly Asp Ser Arg Ser Gly Val 85 90 95 Asp Gly Leu Cys Pro Trp Trp Trp Thr Ala Ser Thr Tyr Ala Phe Gln 100 105 110 Ser Ala Met Ile Asp Ala Asp Arg Ala Trp His Asn Trp Met Ser Ser 115 120 125 Val Thr Gly Gln Arg Ala Gly Arg Arg Val Gly Arg Pro Arg Phe Lys 130 135 140 Ser Lys His Arg Cys Arg Asp Ser Phe Arg Ile His His Asn Val Lys 145 150 155 160 Gln Pro Thr Ile Arg Pro Asp Thr Ser Gly Tyr Arg Arg Leu Ile Val 165 170 175 Pro Arg Leu Gly Ser Leu Arg Thr His Asp Ser Thr Lys Arg Leu Arg 180 185 190 Arg Ala Leu Asp Arg Gly Ala Val Ile Gln Ser Val Thr Ile Ser Arg 195 200 205 Gly Gly His Arg Trp Tyr Ala Ser Val Leu Val Lys Thr Pro Asp Ala 210 215 220 Pro Ala Val Gly Pro Ser Arg Ala Gln Arg Lys Ala Gly Ile Val Gly 225 230 235 240 Val Asp Ile Gly Val His Ser Leu Ala Ala Leu Ser Thr Gly Glu Leu 245 250 255 Val Val Asn Pro Arg His Leu Asn Ser Ser Cys Asp Arg Leu Gly Ala 260 265 270 Ala Gln Arg Ala Leu Ala Arg Cys Gln Lys Gly Ser Asn Arg Arg Arg 275 280 285 Arg Ala Ala Arg Leu Leu Gly Arg Arg His His Glu Leu Ala Glu Arg 290 295 300 Arg Ala Thr Ala Leu His Gln Leu Thr Lys Arg Leu Ala Thr Cys Trp 305 310 315 320 Ser Val Val Ala Val Glu Asp Leu His Val Ala Gly Met Ile Arg Ser 325 330 335 Ala Arg Gly Thr Val Asp Lys Pro Gly Arg Asn Val Arg Ala Lys Ser 340 345 350 Ala Leu Asn Arg Ala Ile Leu Asp Val Ala Pro Gly Glu Leu Arg Arg 355 360 365 Gln Leu Ala Tyr Lys Thr Val Trp Tyr Gly Ser Ser Leu Val Val Cys 370 375 380 Asp Arg Phe Tyr Pro Ser Thr Gln Thr Cys Ser Ala Cys Gly Ala Lys 385 390 395 400 Ala Lys Leu Thr Leu Ala Asp Arg Val Phe Arg Cys Thr Ala Cys Gly 405 410 415 Phe Gly Pro Thr Asp Arg Asp Ile Asn Ala Ala Arg Asn Ile Ala Ala 420 425 430 Asn Ala Ala Val Ala Ser Gly Ala Gly Glu Thr Leu Asn Ala Arg Arg 435 440 445 Ala Asp Gln Gln His Leu Pro Gln Val Gly Pro Cys Arg Arg Pro Ala 450 455 460 Met Lys Arg Glu Gly His His Gln Ser Lys Thr Asp Ser Gly His Leu 465 470 475 480 Ser Arg Ala Thr Gly 485 <210> 12 <211> 569 <212> PRT <213> Artificial Sequence <400> 12 Met Pro Leu Asn Gly Trp Thr Met Leu Ala Gln Glu Glu Val Arg Ala 1 5 10 15 Met Thr Thr Thr Thr Gly Thr Asp Leu Ala Pro Arg Leu Arg Ala Phe 20 25 30 Lys His Arg Leu Asp Pro Asn Pro Ala Gln Ala Thr Leu Leu Ala Gln 35 40 45 Tyr Ala Gly Ala Ala Arg Val Ala Tyr Asn Met Leu Ile Ala His Asn 50 55 60 Arg Ala Ala Leu Ala Ala Gly Ala Ala Arg Arg Thr Glu Leu Ala Glu 65 70 75 80 Ser Gly Leu Ala Gly Pro Glu Leu Ala Ala Arg Met Lys Ala Glu Arg 85 90 95 Ala Ala Asp Pro Thr Leu Arg Val Ala Ser Tyr Gln Ser Tyr Ser Thr 100 105 110 Ala His Leu Thr Pro Leu Ile Arg Arg His Arg Glu Ala Ala Ala Ala 115 120 125 Ile Ala Ala Gly Ala Asp Pro Ala Glu Ala Trp Thr Asp Glu Arg Tyr 130 135 140 Ala Glu Pro Trp Met His Thr Val Pro Arg Arg Val Leu Val Ser Gly 145 150 155 160 Leu Gln Asn Ala Ala Lys Ala Thr Glu Asn Trp Met Ala Ser Ala Ser 165 170 175 Gly Thr Arg Ala Gly Ala Arg Val Gly Leu Pro Arg Phe Lys Lys Lys 180 185 190 Gly Arg Ser Arg Asp Ser Phe Thr Ile Pro Ala Pro Glu Val Ile Gly 195 200 205 Ala Ala Gly Thr Pro Tyr Lys Arg Gly Glu His Arg Arg Gly Val Ile 210 215 220 Thr Asp His Arg His Leu Arg Leu Ala Ser Leu Gly Thr Ile Arg Thr 225 230 235 240 Tyr Asp Lys Thr Ser Arg Leu Val Arg Ala Cys Arg Arg Gly Ala Gln 245 250 255 Ile Arg Ser Met Thr Ile Ser Gln Ala Gly Gly Arg Trp Tyr Ala Ser 260 265 270 Ile Leu Val Ala Asp Pro Thr Pro Ile Arg Thr Gly Pro Ser Arg Arg 275 280 285 Gln Arg Ala Asn Gly Ala Val Gly Val Asp Leu Gly Val Lys His Leu 290 295 300 Ala Ala Leu Ser Thr Gly Glu Val Ile Asp Asn Gly Arg Pro Gly Ala 305 310 315 320 Arg Gln Ala Ala Arg Leu Thr Arg Leu Gln Arg Ala Tyr Ala Arg Thr 325 330 335 Gln Pro Gly Ser Asn Arg Arg Glu Arg Val Arg Arg Gln Ile Ala Ala 340 345 350 Leu His His Gly Ile Ala Leu Arg Arg Ala Gly Leu Leu His Gln Val 355 360 365 Ser Thr Arg Leu Ala Met Asp Phe Ala Val Val Ala Leu Glu Asp Leu 370 375 380 Asn Val Ala Gly Met Thr Arg Ser Ala Arg Gly Thr Leu Glu Ala Pro 385 390 395 400 Gly Arg Asn Val Ala Ala Lys Ser Gly Leu Asn Arg Ala Ile Leu Asp 405 410 415 Ala Gly Leu Gly Met Leu Arg Arg Gln Leu Asp Tyr Lys Thr Ser Trp 420 425 430 Ala Gly Ser Gln Val Lys Met Ile Asp Arg Phe Ala Pro Ser Ser Lys 435 440 445 Ala Cys Ser Arg Cys Gly Thr Val Lys Ser Thr Leu Ser Leu Ala Glu 450 455 460 Arg Thr Phe Glu Cys Glu Ala Cys His Leu Val Ile Asp Arg Asp Val 465 470 475 480 Asn Ala Ala Ile Asn Ile Arg Ala Trp Ala Val Gln Glu Glu Arg Arg 485 490 495 Ala Gly Val Glu Leu Ala Arg Gly Arg Arg Glu Ser Arg Asn Gly Arg 500 505 510 Gly Ala Ala Val Ser Gly Pro Pro Ser Gly Gly Ala Ala Gly Gln Gly 515 520 525 Arg Gly Ser Val Lys Pro Ala Pro Gln Gly Val Gly Met Ser Ser Arg 530 535 540 Ala Thr Gly Trp Ser Ser Gln Pro Pro Ser Thr Glu Gly Glu Ser Ala 545 550 555 560 Glu Arg Gly Ala Ser Ala Leu Ala Arg 565 <210> 13 <211> 505 <212> PRT <213> Artificial Sequence <400> 13 Met Gly Glu Gln Leu Arg Ala Tyr Arg Tyr Ala Leu Asp Leu Thr Asp 1 5 10 15 Ala Gln Ala Ser Val Val Ala Gln His Ala Gly Ala Ala Arg Trp Ala 20 25 30 Tyr Asn His Ala Ile Ala Thr Lys Phe Lys Ala Leu Asp Glu Arg Gln 35 40 45 Ala Ala Ile Arg Glu Ala Val Asp Ala Gly Val Asp Pro Ala Val Ala 50 55 60 Ala Asn Leu Ala Pro Lys Ile Pro Ser Lys Phe Asp Ile Gln Lys Ala 65 70 75 80 Leu Asn Lys Val Lys Gly Asp Asp Arg Thr Gly Arg Glu Gly Glu Cys 85 90 95 Pro Trp Trp His Ser Val Ser Thr Phe Ala Phe Gln Ser Ala Phe Ser 100 105 110 Asp Ala Asp Arg Ala Trp Lys Asn Trp Leu Asp Ser Ile Thr Gly Asn 115 120 125 Arg Lys Gly Arg Arg Val Gly Arg Pro Arg Phe Lys Ala Lys Arg Arg 130 135 140 Ser His Asp Ser Phe Arg Ile His His Asp Val Lys Asn Pro Thr Ile 145 150 155 160 Arg Pro Asp Val Gly Cys Tyr Arg Arg Ile Arg Val Pro Arg Leu Gly 165 170 175 Ser Leu Arg Val His Ser Ser Thr Lys Gln Leu Cys Arg Ala Leu Asp 180 185 190 Arg Gly Ala Val Ile Gln Ser Val Thr Ile Ser Arg Gly Gly His Arg 195 200 205 Trp Tyr Ala Ser Ile Leu Thr Lys Gln Glu Ile Val Ser Glu Val Leu 210 215 220 Pro Thr Arg Arg Gln Arg Leu Ala Gly Thr Val Gly Val Asp Leu Gly 225 230 235 240 Val His Thr Leu Ala Ala Leu Ser Thr Gly Glu Thr Val Ser Asn Pro 245 250 255 Arg His Ala Asn Met Ala Arg Thr Arg Leu Thr Lys Ala Gln Gly Ala 260 265 270 Leu Ser Arg Thr Lys Lys Gly Ser Asn Arg Arg Arg Arg Ala Ala Arg 275 280 285 Leu Val Gly Arg Arg His His Glu Val Ala Glu Arg Arg Ala Ser Thr 290 295 300 Leu His Ala Val Thr Lys Ala Leu Ala Thr Arg Trp Ala Thr Val Ala 305 310 315 320 Val Glu Asp Leu Asn Val Ala Gly Met Thr Ala Ser Ala Arg Gly Thr 325 330 335 Val Asp Lys Pro Gly Arg Asn Val Arg Ala Lys Ala Gly Leu Asn Arg 340 345 350 Ala Ile Leu Asp Val Ser Pro Gly Glu Phe Arg Arg Gln Leu Glu Tyr 355 360 365 Lys Thr Ala Trp Tyr Gly Ser Ala Leu Ala Val Ile Asp Arg Phe Tyr 370 375 380 Pro Ser Ser Gln Thr Cys Ser Ala Cys Gly Ala Arg Ala Lys Leu Thr 385 390 395 400 Leu Ala Glu Arg Ile Tyr Arg Cys Ala Ala Cys Gly Phe Thr Ala Asp 405 410 415 Arg Asp Val Asn Ala Ala Ile Asn Ile Ala Ala Gln Ala Ala Val Ala 420 425 430 Pro Gly Lys Gly Glu Thr Ile Asn Ala Arg Arg Ala Val Ser Asp His 435 440 445 Pro Gln Pro Ala Gly Thr Arg Ser Leu Thr Ala Met Lys Arg Glu Gly 450 455 460 His Arg Asp Asn Ala Val Val Thr Pro Thr Arg Gln Arg Val Gly His 465 470 475 480 Pro Ile Arg Cys Gly Met Arg Leu Asp Glu Arg Ile Gly Ala Gly Gln 485 490 495 Glu Ala Ser Ala Pro His Ala Tyr Gly 500 505 <210> 14 <211> 561 <212> PRT <213> Artificial Sequence <400> 14 Met Lys Glu Leu Ala Ala Asp Gly Ile Pro Val Ala Val Thr Cys Arg 1 5 10 15 Val Leu Lys Leu Ala Arg Gln Pro Tyr Tyr Arg Trp Leu Ala Gln Pro 20 25 30 Val Thr Asp Ala Glu Leu Val Glu Ala His Arg Ala Asn Ala Leu Phe 35 40 45 Asp Ala His Arg Asp Asp Pro Glu Phe Gly Tyr Arg Tyr Leu Ala Glu 50 55 60 Glu Ala Arg Glu Ala Gly Glu Pro Met Ala Glu Arg Thr Ala Trp Arg 65 70 75 80 Ile Cys Ser Gln Asn Arg Leu Trp Ser Val Phe Gly Lys Arg Lys Arg 85 90 95 Gly Lys Asn Gly Lys Thr Gly Pro Pro Val His Asp Asp Leu Val Glu 100 105 110 Arg Asp Phe Thr Ala Glu Ala Pro Asn Gln Leu Trp Leu Ala Asp Ile 115 120 125 Thr Glu His Arg Thr Gly Glu Gly Lys Leu Tyr Leu Cys Ala Ile Lys 130 135 140 Asp Ala Phe Ser Asn Arg Ile Val Gly Tyr Ser Ile Asp Ser Arg Met 145 150 155 160 Lys Ser Arg Leu Ala Thr Thr Ala Leu Ser Ser Ala Val Ala Arg Arg 165 170 175 Gly Asp Val Ala Gly Cys Ile Leu His Ser Asp Arg Gly Ser Gln Phe 180 185 190 Arg Ser Arg Arg Phe Val Gln Ala Leu Asn His His Ala Met Val Gly 195 200 205 Ser Met Gly Arg Val Gly Ala Ala Gly Asp Asn Ala Ala Met Glu Ser 210 215 220 Phe Phe Ser Leu Leu Gln Lys Asn Val Leu Asp Arg Arg Arg Trp Asp 225 230 235 240 Thr Arg Glu Gln Leu Arg Ile Ala Ile Val Thr Trp Ile Glu Arg Thr 245 250 255 Tyr His Arg Arg Arg Arg Gln Ala Gly Pro Arg Thr Ile Asp Pro His 260 265 270 Arg Ile Arg Ser Asp Tyr Glu Arg Thr Gly Gln Ser Gly Arg Val Thr 275 280 285 Glu Thr Val Thr Tyr Ser Cys Ser Arg Pro Ala Gly Arg Val Gly Val 290 295 300 Asp Leu Gly Val His His Leu Ala Ala Leu Ser Thr Gly Glu Leu Ile 305 310 315 320 Asp Asn Pro Arg His His Arg Gln Ala Arg Arg Ala Leu Thr Lys Ala 325 330 335 Gln Gln Ala Leu Ser Arg Thr Glu Lys Gly Ser Gln Arg Arg Arg Lys 340 345 350 Ala Cys Ala Thr Leu Ala Arg Arg His His Leu Leu Ala Glu Arg Arg 355 360 365 Ala Thr Thr Val His Gly Ile Thr Lys Arg Leu Ala Thr Gly Trp Ser 370 375 380 Glu Val Ala Ile Glu Asp Leu Asn Val Ala Gly Met Thr Arg Ser Ala 385 390 395 400 Arg Gly Thr Val Glu Ser Pro Gly Val Asn Val Ala Ala Lys Ala Gly 405 410 415 Leu Asn Arg Ser Ile Leu Asp Ala Ala Pro Asp Glu Leu Arg Arg Gln 420 425 430 Leu Thr Tyr Lys Thr Gln Trp Tyr Gly Ser Arg Leu Ala Val Cys Asp 435 440 445 Arg Trp Phe Pro Ser Thr Gln Ala Cys Ser Ala Cys Gly Ala Lys Ala 450 455 460 Lys Leu Thr Leu Ala Asp Arg Val Phe His Cys Pro Ala Cys Gly Phe 465 470 475 480 Gly Pro Ile Asn Arg Asp His Asn Ala Ala Arg Asn Ile Ala Ala His 485 490 495 Ala Val Ile Val Ala Pro Gly Thr Gly Glu Thr Glu Asn Ala Arg Arg 500 505 510 Ala Asp Ala Gly Pro Pro Pro Arg Val Glu Thr Thr Pro Pro Ser Val 515 520 525 Met Lys Arg Glu Asp His Pro Thr Arg Val Ala Thr Ser Ala Glu Gln 530 535 540 Ser Ala Asp His Pro Lys Pro Ala Asp Gln His Thr Leu Pro Leu Val 545 550 555 560 Asn <210> 15 <211> 496 <212> PRT <213> Artificial Sequence <400> 15 Met Ala Gln Val Leu Arg Ala Phe Arg Tyr Ala Leu Asp Pro Ser Pro 1 5 10 15 Pro Gln Leu Glu Thr Leu Gln Arg Cys Ala Gly Asn Ala Arg Leu Ala 20 25 30 Phe Asn Phe Ala Leu Ala Met Lys Lys Asp Ala His Gln Arg Trp Arg 35 40 45 Asp Glu Val Asp Ile Leu Ile Thr Thr Gly Leu Ser Glu Lys Asp Ala 50 55 60 Arg Thr Arg Leu Lys Gly Thr Ser Lys Ile Pro Asn Lys Pro Asp Val 65 70 75 80 Tyr Lys Ala Phe Gln His Leu Arg Gly Asp Ala Ala Gln Gly Ile Asp 85 90 95 Gly Ile Ala Pro Trp His Ala Glu Ile Pro Thr Tyr Val Phe Gln Ser 100 105 110 Ala Phe Gln Asp Ala Asp Arg Ala Trp Lys Asn Trp Leu Asp Ser Tyr 115 120 125 Thr Gly Lys Arg Ala Gly Arg Arg Val Gly Tyr Pro Arg Phe Lys Lys 130 135 140 Lys Phe Arg Ser Arg Asp Ser Phe Arg Leu His His Asp Ala Lys Lys 145 150 155 160 Pro Lys Pro Ala Leu Arg Leu Glu Gly Tyr Arg Arg Leu Arg Leu Gly 165 170 175 Gly Ala Leu Lys Thr Val Arg Leu His Gly Ser Ala Lys Pro Leu His 180 185 190 Arg Leu Val Ser Ser Gly Arg Ala Val Val Gln Ser Val Thr Val Ser 195 200 205 Arg Gly Gly Thr Arg Trp Tyr Ala Ser Val Leu Cys Lys Val Glu Thr 210 215 220 Asp Val Pro Asp Pro Thr His Arg Gln Lys Ala Asn Gly Arg Ile Gly 225 230 235 240 Leu Asp Trp Gly Leu Asn His Leu Ala Ala Leu Ser Thr Pro Leu Asp 245 250 255 Gly His His Leu Val Asp Asn Pro Arg His Leu Arg His Ala Ser Lys 260 265 270 Arg Leu Thr Lys Ala Gln Arg Ala Leu Ser Arg Thr Gln Lys Gly Ser 275 280 285 Ala Arg Arg Arg Arg Ala Ala Ala Arg Val Gly Lys Leu His His Leu 290 295 300 Val Ala Glu Gln Arg Ala Thr Phe Leu His Thr Leu Thr Lys Arg Leu 305 310 315 320 Thr Thr Thr Phe Ala Cys Val Ala Ile Glu Asp Leu Asn Val Ala Gly 325 330 335 Met Thr Arg Ser Ala Arg Gly Thr Val Gln Leu Pro Gly Lys Asn Val 340 345 350 Ser Ser Lys Ala Gly Leu Asn Arg Ser Ile Leu Asp Ala Ala Pro Ala 355 360 365 Glu Leu Arg Arg Gln Leu Glu Tyr Lys Thr Ser Trp Tyr Gly Ser His 370 375 380 Ile Ala Ile Leu Asp Arg Trp Phe Pro Ser Ser Lys Thr Cys Ser Gly 385 390 395 400 Cys Gly Trp Arg Asn Pro Ser Leu Pro Leu Ser Glu Arg Glu Phe Val 405 410 415 Cys Ala Glu Cys Gly Leu Arg Leu Asp Arg Asp Leu Asn Ala Ala Arg 420 425 430 Asn Ile Ala Ala His Ala Glu Val Pro Ala Ser Gly Thr Gly Ala Pro 435 440 445 Gly Arg Gly Glu Ser Val Asn Ala Arg Gly Gly Cys Val Ser Pro Arg 450 455 460 Leu Leu Arg Glu Ser Gly Gln His Pro Ser Lys Arg Glu Asp Ala Ala 465 470 475 480 Pro Ser Gly Pro Ala Pro Pro Arg Arg Ser Asn Pro Pro Thr Phe Pro 485 490 495 <210> 16 <211> 499 <212> PRT <213> Artificial Sequence <400> 16 Met Ala Glu Val Leu Arg Thr Phe Lys Phe Ala Leu Asp Pro Thr Ser 1 5 10 15 Thr Gln Leu Gln Thr Leu Gly Arg Tyr Ser Gly Asn Ala Arg Leu Ala 20 25 30 Tyr Asn Phe Ala Leu Ala Met Lys Lys Glu Ala His Gln Gln Trp Arg 35 40 45 Asp Arg Val Asp Thr Leu Val Ala Thr Gly Val Asp Glu Lys Ala Ala 50 55 60 Arg Asp Arg Leu Lys Gly Thr Val Lys Val Pro Lys Lys Pro Gln Val 65 70 75 80 Tyr Lys Ala Phe Gln Leu Leu Arg Gly Asp Ala Ala Lys Gly Ile Glu 85 90 95 Gly Leu Ala Pro Trp His Ala Glu Ile Pro Thr His Val Phe Gln Ser 100 105 110 Ala Trp Ile Asn Ala Asp Arg Ala Trp Gln Asn Trp Ile Asp Ser Tyr 115 120 125 Arg Gly Thr Arg Ala Gly Arg Arg Val Gly Tyr Pro Lys Phe Lys Lys 130 135 140 Lys Phe Arg Ser Arg Asp Ser Phe Arg Leu His Gln Ser Ser Gly Arg 145 150 155 160 Pro Val Leu Arg Leu Glu Gly Tyr Arg Arg Leu His Leu Ser Gly Gln 165 170 175 Leu Gly Thr Val Arg Leu His Gly Ser Ala Lys Arg Leu His Arg Leu 180 185 190 Val Ser Ser Ser Gln Ala Lys Val Gln Ser Val Thr Ile Ser Arg Gly 195 200 205 Gly Ser Arg Trp Tyr Ala Asn Val Leu Cys Lys Val Asp Ser Asn Leu 210 215 220 Pro Gly Pro Thr Arg Arg Gln Arg Ala Asn Asn Arg Val Gly Leu Asp 225 230 235 240 Trp Gly Val His His Leu Ala Ala Leu Ser Thr Pro Ile Glu Gly Thr 245 250 255 Thr Leu Val Asp Asn Pro Lys His Leu Ala Lys Ala Thr Thr Arg Leu 260 265 270 Arg Lys Ala Gln Arg Ala Leu Ser Arg Thr Gln Lys Gly Ser Lys Arg 275 280 285 Arg His Thr Ala Ala Ser Arg Val Gly Lys Leu His His Gln Val Ala 290 295 300 Glu Gln Arg Ala Thr Phe Leu His Ala Leu Thr Lys Arg Leu Cys Thr 305 310 315 320 Thr Phe Ala Phe Val Ala Val Glu Asp Leu Asn Val Ala Gly Met Thr 325 330 335 His Ser Ala Arg Gly Thr Lys Asp Lys Pro Gly Thr Asn Val Arg Ala 340 345 350 Lys Ala Gly Leu Asn Arg Ser Ile Leu Asp Ala Ala Pro Gly Glu Leu 355 360 365 Arg Arg Gln Leu Glu Tyr Lys Thr Ser Trp Tyr Gly Ser Gln Leu Ala 370 375 380 Leu Cys Asp Arg Trp Phe Pro Ser Ser Lys Thr Cys Ser Ala Cys Gly 385 390 395 400 Trp Gln Asn Pro Ser Leu Ser Leu Ser Glu Arg Val Phe Val Cys Ala 405 410 415 Glu Cys Gly Leu Arg Val Asp Arg Asp Leu Asn Ala Ala Arg Asn Ile 420 425 430 Ala Ala His Ala Glu Val Pro Ser Gly Thr Gly Ala Pro Gly Arg Gly 435 440 445 Glu Ser Val Asn Ala Arg Gly Gly Arg Val Ser Pro Pro Leu Ser Arg 450 455 460 Glu Arg Gly Gln Arg Pro Ser Lys Arg Glu Asp Ala Gly Leu Lys Gly 465 470 475 480 Pro Ala Pro Pro Arg Arg Ser Asn Pro Pro Thr Ser Leu Thr Arg Val 485 490 495 Ser Arg Thr <210> 17 <211> 560 <212> PRT <213> Artificial Sequence <400> 17 Met Leu Ile Ala His Lys Ile Lys Leu Leu Pro Thr Pro Glu Gln Ala 1 5 10 15 Ala Leu Phe Arg Arg Cys Glu Thr Val Ala Lys Ala Thr Tyr Asn Gln 20 25 30 Ala Leu Arg Met Trp Gln Glu Asp Tyr Gln Arg Tyr Leu Arg Glu Met 35 40 45 Val Ala Leu Cys Gly Ser Arg Gly Ile Ser Ala Leu Thr Asp Ser Pro 50 55 60 Arg Cys Leu Trp Pro Arg Phe Lys Lys Arg Ser Asp Leu Glu Lys Glu 65 70 75 80 Leu Ile Ala Ala Gly Val Ala Lys Asp Asp Ile Pro Ala Thr Ser Gly 85 90 95 Phe Arg Trp Ile Ser Lys His Tyr Ala Asn Arg Pro Ala His Trp Glu 100 105 110 Asp Leu Gly Tyr Trp Pro Gly Ala Gly Lys Ile Thr Asn Leu Lys Asp 115 120 125 Ala Ala Asn His Phe Phe Lys Asp Gly Phe Gly Tyr Pro Arg Pro Ile 130 135 140 Lys Glu Asp Asp Val Cys Gly Phe Leu Val Leu Ser Ser Ala Glu Lys 145 150 155 160 Cys Gly Gly Lys Lys Asp Ala Asn Gly Arg His Asn Gly Pro Phe Ala 165 170 175 Arg His Val Asp Arg Pro Val Trp Gln Arg Pro Leu Glu Phe Asn Val 180 185 190 Pro Lys Ile Gly Lys Val Arg Met Ala Glu Pro Leu Arg Phe Glu Lys 195 200 205 Phe Glu Thr Gln Lys Lys Ile Ser Pro Arg Gly Lys Asn Glu Gly Gln 210 215 220 Glu Ile Glu Thr Phe Pro Val Ala Ser Val Thr Ile Ser Arg Asn Lys 225 230 235 240 Ala Gly Asp Trp Phe Ala Ser Ile Leu Cys Ser Val Pro Glu Ile Glu 245 250 255 Pro Asn Ala Pro Val Arg Asp Tyr Lys Val Gly Val Asp Leu Gly Ile 260 265 270 Lys Thr Leu Ala Lys Leu Ser Asp Glu Val Pro Gly Gln Pro Asp Arg 275 280 285 Phe Pro Ala Arg Lys Ala Leu Ser Asn Ala Glu Arg Arg Leu Leu Arg 290 295 300 Leu Glu Arg Lys Met Ser Arg Gln Gln Ala Gln Ala Ala Arg Arg Lys 305 310 315 320 Arg Lys Leu Ser Glu Cys Lys Asn Tyr Gln Lys Thr Lys Arg Leu His 325 330 335 Ala Leu Ala Trp Lys Lys Val Gly His Arg Arg Glu Ser Tyr Ser His 340 345 350 Gln Leu Thr Thr Val Leu Val Arg Asp Tyr Pro Val Ile Gly Ile Glu 355 360 365 Asp Leu Gly Val Lys Gly Met Met Arg Lys Gly Lys Pro Gly Asn Lys 370 375 380 Pro Lys Lys Ala Arg Leu Asn Arg Ala Ile Gly Asp Ala Gly Phe Gly 385 390 395 400 Arg Phe Arg Ile Gln Leu Glu Tyr Lys Ala Lys Trp His Arg Arg Glu 405 410 415 Val His Tyr Thr Pro Gln Phe Glu Pro Thr Ser Lys Val Cys Asn Gln 420 425 430 Cys Gly Trp Lys Asn Glu Lys Leu Lys Leu Ser Asp Arg Gln Trp Arg 435 440 445 Cys Gly Gly Cys Glu Thr Ile His Asp Arg Asp Leu Asn Ala Ala Leu 450 455 460 Asn Ile Cys Arg Lys Ser Val Pro Lys Ser Ser Asp Asp Pro Ser Val 465 470 475 480 Ala Gly Lys Asp Glu Lys Asp Val Lys Ser Gly Ser Ala Lys Ser Arg 485 490 495 Arg Gly Arg Phe Val Ser Arg Lys Ser Glu Gly Asp Ser Ser Ser Pro 500 505 510 Phe Asp Pro Gly Ala Val Asp Thr Glu Gln Ser Val Val Gly Ala Pro 515 520 525 Asn Gly Gln Lys Ala Leu Glu Glu Ala Thr Gly His Arg Arg Asn Asn 530 535 540 Ala Asp Phe Ser Arg Ser Ser Gly Phe Lys Ser Arg Leu Asp Gly Tyr 545 550 555 560 <210> 18 <211> 532 <212> PRT <213> Artificial Sequence <400> 18 Met Ser Tyr Leu Cys Ala Arg Gly Arg Phe Asn Leu Ser Gly Phe Ser 1 5 10 15 Thr Gly Ser Met Gly Val Leu Pro Val Arg Ala Ser Leu Phe Ala Leu 20 25 30 Leu Phe Thr Thr Leu His Gly Ala Arg Ile Ser Ala His Trp Tyr Gly 35 40 45 Met Ser Lys Glu Asn Asp His Thr Val Thr Cys Ile Lys Ile Cys Leu 50 55 60 Glu Pro Asn Lys Ala Gln Arg Ala Gln Phe Ala Ser Phe Ala Gly Ser 65 70 75 80 Ala Arg Trp Ala Tyr Asn Phe Ala Leu Ala Ile Lys Ile Gly Tyr Gln 85 90 95 Lys Arg Trp Phe Glu Ala Arg Lys Gln Phe Ile Glu Ser Gly Leu Asp 100 105 110 Glu Lys Ala Ala Gly Lys Lys Ala Ser Glu Gln Val Gly Arg Met Pro 115 120 125 Asn Tyr Met Ser Ile Ala Thr Asn Glu Trp Thr Gln Leu Arg Asp Glu 130 135 140 Val Cys Pro Trp Tyr Pro Glu Val Pro Arg Arg Val Phe Val Gly Gly 145 150 155 160 Phe Gln Arg Ala Asp Ala Ala Phe Lys Asn Trp Phe Asp Ser Lys Ser 165 170 175 Gly Arg Arg Ser Gly Ala Ala Met Gly Trp Pro Lys Phe Lys Ser Lys 180 185 190 Ser Lys Ser Arg Glu Ser Phe Val Ile Ala Asn Asp Val Gln Pro Ala 195 200 205 Phe Val Ala Asn Leu Asn Arg Tyr Ile Lys Thr Gly Glu Leu Ala Asp 210 215 220 Met Asp Tyr Arg His Ile Lys Val Pro Lys Cys Gly Glu Val Arg Leu 225 230 235 240 Thr Pro Gly Ser Ala Gly Gln Leu Arg Gln Leu Gly Arg Thr Met Leu 245 250 255 Ala Glu Ala Lys Thr Gly Glu Leu Ile Thr Arg Ile Thr Ser Gly Thr 260 265 270 Ile Ser Arg Leu Gly Asp Arg Trp Tyr Val Ser Leu Val Ile Ser Gly 275 280 285 Pro Phe Val Pro Asp Ala Ile Ser Thr Lys Arg Gln Arg Arg Asn Gly 290 295 300 Val Val Gly Val Asp Leu Gly Ser Gly Arg Phe Tyr Ala Thr Thr Ser 305 310 315 320 Glu Gly Leu Ser Ile Ile Asn Pro Lys Phe Val Ser Lys Tyr Glu Gln 325 330 335 Glu Leu Ala Arg Ala Asn Arg Ala Leu Ala Lys Thr Ala Lys Gly Ser 340 345 350 Ala Ala Arg Lys Lys Ala Leu Ala Arg Leu Arg Arg Val His Ala Arg 355 360 365 Ser Ala Leu Ala Arg Asp Gly Phe Ser His Gln Val Ser Ala Trp Leu 370 375 380 Thr Ser Gln Phe Ala Gly Val Ala Val Glu Lys Phe Asp Leu Ala Ser 385 390 395 400 Met Leu Ala Ser Ala Lys Gly Thr Val Glu Lys Pro Gly Lys Asn Val 405 410 415 Asp Val Lys Ala Arg Phe Asn Ala His Leu Ala Asp Val Gly Ile Ala 420 425 430 Ser Thr Ile Asp Lys Leu Leu Tyr Lys Gly Lys Arg Asp Gly Cys Arg 435 440 445 Val Gln Val Val Asn Thr Leu Asp Asn Ser Ser Thr Thr Cys Ala Lys 450 455 460 Cys Gly His Thr Cys Val Cys Gly Pro Glu Gln Lys Thr Phe Thr Cys 465 470 475 480 Pro Asp Cys Gly Tyr Asn Ala Pro Arg Gln Leu Asn Ser Ala Gln Tyr 485 490 495 Ile Arg Gln Leu Ala Thr Val Gly Phe Asp Glu Leu Gly Leu Asp Met 500 505 510 Thr Ala Ser Leu Thr Pro Asp Thr Gly Lys Arg Pro Ile Ala Phe Met 515 520 525 Thr Ser Ala His 530 <210> 19 <211> 579 <212> PRT <213> Artificial Sequence <400> 19 Met Thr Gly Ala Lys Gln Lys Thr Pro Ile Arg Val Val Arg Phe Ser 1 5 10 15 Ile Asp His Ser Ala Leu Thr Pro Ala Gln Val Val Ala Phe Ala Arg 20 25 30 His Ala Gly Ala Ala Arg Gln Thr Trp Asn Trp Ala Leu Gly Arg Trp 35 40 45 Met Asp Trp Arg Asn Asn Thr Lys Phe Tyr Val Asp Tyr Lys Val Phe 50 55 60 Lys Ala Ala Gly Met Gly Pro Gly Leu Ser Thr Asp Asp Leu Ile Gln 65 70 75 80 Val Ile Glu Arg Ala Val Ser Ile Arg Gln Asp Asp Lys Trp Met Asp 85 90 95 Ala Ala Trp Asp Glu Ala Arg Gln Ile His Gly Glu Trp Asp Gln Phe 100 105 110 Gln Lys Ala Ser Thr Leu Gln Ser Leu Tyr Leu Ala Gly Ala Gln Glu 115 120 125 Pro Phe Asp Pro Ser Arg Asp Asp Gly Ile Asn Pro Tyr His Trp Trp 130 135 140 Val Thr Glu Gly Asp Lys Ser Gly Leu Pro Lys Ala Glu Arg His Asn 145 150 155 160 Val Asn Ser Gly Ala Thr Tyr Thr Ala Pro Leu Arg Ala Phe Glu Glu 165 170 175 Ala Val Gly Arg Phe Tyr Lys Leu Pro Gly Lys Lys Gly Thr Pro Lys 180 185 190 Phe Lys Ser Lys His Asp Asp Glu Gln Gly Phe Cys Ile Gln Arg Leu 195 200 205 Thr Glu Thr Gly Leu Ser Pro Trp Arg Ala Ile Glu Gly Gly His Arg 210 215 220 Ile Lys Val Pro Ser Ile Gly Ser Ile Arg Val Val Gln Ser Thr Lys 225 230 235 240 Arg Leu Arg Gln Leu Ile Lys Arg Gly Gly Lys Thr Thr Ser Ala Arg 245 250 255 Phe Thr Arg Arg Gly Gly Lys Trp Phe Val Ser Val Ser Val Ala Phe 260 265 270 Asp Leu Ser Ala Pro Arg Val Gln Arg Pro Ala Arg Leu Ser Arg Arg 275 280 285 Gln Arg Ala Gly Gly Ser Thr Gly Val Asp Leu Gly Val Asn Arg Leu 290 295 300 Ala Thr Leu Ser Ser Gly Asp Gln Phe Pro Asn Arg Arg Leu Leu Arg 305 310 315 320 Lys Ser Met Ala Glu Ile Lys Arg Leu Gln Arg Lys Phe Asp Arg Gln 325 330 335 His Arg Ala Gly Ser Pro Glu Cys Phe Asn Glu Asp Gly Thr His Lys 340 345 350 Lys Arg Cys Arg Trp Gly Arg Glu Asp Gly Pro Ala Met Ser Arg Ser 355 360 365 Ala Gln Thr Thr Lys Arg Gln Leu Arg Arg Ile His Asp Leu Thr Ala 370 375 380 Arg Arg Arg Ala Gly Val Leu His Glu Ile Thr Lys Asp Leu Ala Thr 385 390 395 400 Arg Phe Glu Leu Ile Gly Val Glu Asp Leu Asn Val Ala Gly Met Thr 405 410 415 Ala Lys Ser Lys Pro Lys Pro Asp Pro Asp Arg Pro Gly His Phe Leu 420 425 430 Pro Asn Arg Arg Ala Ala Lys Ala Gly Leu Asn Arg Ala Ile Leu Asp 435 440 445 Val Gly Phe Tyr Glu Phe Lys Arg Gln Leu Gly Tyr Lys Thr Glu Trp 450 455 460 Tyr Gly Ser Thr Met Gln Met Val His Arg Tyr Ala Ala Thr Ser Lys 465 470 475 480 Thr Cys Ser Gly Cys Gly Trp Val Lys Pro Lys Leu Thr Leu Ala Glu 485 490 495 Arg Thr Phe Asn Cys Thr Gln Cys Gly Leu Ala Met Asp Arg Asp His 500 505 510 Asn Ala Ala Val Asn Ile Arg Ala Leu Ala Leu Glu Gly Ala Ala Pro 515 520 525 Met Glu Arg Glu Gln Pro Ala Pro Val Gly Ala Ala Glu Lys Arg His 530 535 540 Arg Asp Pro Val Ser His Arg Arg Arg Pro Lys Ser Leu Ala Pro Cys 545 550 555 560 Glu Ser Thr Arg Pro Val Arg Asp Leu Ser Pro Pro Ala Thr Gln Glu 565 570 575 Glu Thr Ala <210> 20 <211> 540 <212> PRT <213> Artificial Sequence <400> 20 Met Ala Thr Thr Asp Lys Lys Asp Glu Glu Asn Leu Arg Ala Tyr Lys 1 5 10 15 Phe Arg Leu Asp Pro Asn Gln Ala Gln Thr Thr Ala Leu Tyr Gln Ala 20 25 30 Val Gly Ala Ala Arg Tyr Thr Tyr Asn Met Leu Thr Ala Tyr Asn Leu 35 40 45 Glu Val Asn Arg Leu Arg Asp Asp Tyr Trp Lys Arg Arg His Asp Glu 50 55 60 Asp Ile Ser Asp Ala Asp Ile Lys Lys Glu Leu Asn Ala Leu Ala Lys 65 70 75 80 Glu Asp Lys Arg Tyr Lys Gln Leu Asn Tyr Gly Ala Phe Gly Thr Gln 85 90 95 Tyr Leu Thr Pro Glu Lys Lys Arg His Glu Gln Ala Glu His Arg Ile 100 105 110 Glu Asn Gly Glu Asp Pro Ser Val Val Trp Asn Gln Glu Thr Glu Arg 115 120 125 Ser Ala Asn Pro Trp Leu His Thr Ala Asn Gln Arg Val Leu Val Ser 130 135 140 Gly Leu Gln Asn Ala Ser Asp Ala Trp Asp Asn Phe Trp Ala Ser Arg 145 150 155 160 Thr Gly Lys Arg Ala Gly Arg Leu Val Gly Thr Pro Arg Phe Lys Lys 165 170 175 Lys Gly Val Ser Arg Asp Ser Phe Thr Val Pro Ala Pro Glu Lys Met 180 185 190 Gly Ala Tyr Gly Thr Ala Tyr Leu Arg Gly Glu Pro Ala Tyr Lys Gln 195 200 205 Gly Arg Arg Lys Ile Thr Asp Tyr Arg His Val Arg Leu Ser Tyr Leu 210 215 220 Gly Thr Ile Arg Thr Phe Asn Ser Thr Lys Pro Leu Val Lys Ala Val 225 230 235 240 Val Ala Gly Ala Lys Ile Arg Ser Tyr Thr Val Ser Arg Asn Ala Asp 245 250 255 Arg Trp Tyr Val Ser Phe Leu Val Lys Phe Ser Glu Pro Ile Arg Arg 260 265 270 Ser Ala Thr Lys Arg Ala Arg Ala Ala Gly Ser Val Gly Val Asp Leu 275 280 285 Gly Val Lys Tyr Leu Ala Ser Leu Ser Asp Ser Glu Ala Pro Gln Arg 290 295 300 Phe Pro Asn Leu Lys Phe Val Glu Gly Leu Pro Ser Leu Glu Asn Pro 305 310 315 320 Arg Trp Ser Glu Ala Ser Ser Arg Arg Leu His Lys Leu Gln Arg Ala 325 330 335 Leu Ala Arg Ser Gln Lys Gly Ser Asn Arg Arg Ser Arg Leu Val Lys 340 345 350 Gln Ile Ala Arg Leu His His Met Thr Ala Leu Arg Arg Glu Ser Asn 355 360 365 Leu His Gln Leu Thr Lys Lys Leu Ala Thr Glu Tyr Thr Leu Val Gly 370 375 380 Phe Glu Asp Leu Asn Val Ser Gly Met Thr Ala Ser Ala Lys Gly Thr 385 390 395 400 Val Glu Asn Pro Gly Lys Asn Val Ala Gln Lys Ser Gly Leu Asn Arg 405 410 415 Val Val Leu Asp Ala Ala Phe Gly Val Phe Arg Asn Gln Leu Glu Tyr 420 425 430 Lys Ala Val Trp Tyr Gly Ser Ala Phe Glu Lys Val Asp Arg Tyr Phe 435 440 445 Ala Ser Ser Gln Thr Cys Ser Glu Cys Gly Arg Lys Ala Lys Thr Lys 450 455 460 Leu Thr Leu Arg Asp Arg Val Phe Asp Cys Ala Tyr Cys Gly Asn Met 465 470 475 480 Met Asp Arg Asp Leu Asn Ala Ala Val Asn Ile Cys Arg Glu Ala Gln 485 490 495 Arg Leu Phe Asp Glu Lys Leu Ala Ser Glu Asp Arg Glu Ser Leu Asn 500 505 510 Gly Arg Gly Ser Arg Gly Ala Leu Arg Gly Ala Glu Thr Val Glu Ala 515 520 525 Ser Arg Pro Pro Ala Ser His Arg Arg Gly Ser Pro 530 535 540 <210> 21 <211> 549 <212> PRT <213> Artificial Sequence <400> 21 Met Ala Ala Ala Glu Cys Ser Thr Gln Lys Ala His Arg Ala Tyr Lys 1 5 10 15 Phe Arg Leu Ala Pro Asn Gln Lys Gln Leu Thr Ala Leu Tyr Gln Ala 20 25 30 Ala Gly Ala Ala Arg Tyr Ala Tyr Asn Gln Leu Thr Ala Tyr Asn Leu 35 40 45 Asp Ala Leu Arg Arg Arg Gln Gln Tyr Trp Gly Thr Arg Gln Asp Glu 50 55 60 Gly Ala Ser Glu Ser Glu Ile Lys Ala Glu Leu Arg Ala Ile Ala Arg 65 70 75 80 Glu Asp Ala Ser Tyr Arg Leu Leu Gly Tyr Ile Pro Tyr Ala Thr Glu 85 90 95 Tyr Leu Thr Pro Glu Arg Arg Arg His Gln Asp Ala Ala Ala Gln Ile 100 105 110 Ala Ala Gly Ala Gln Ala His Leu Ile Trp Gly Ala Asp Glu His Phe 115 120 125 Glu His Pro Trp Leu His Thr Ala Ser Arg Arg Val Leu Val Ser Gly 130 135 140 Leu Gln Ser Ala Asp Lys Ala Trp Lys Asn Phe Trp Asp Ser Arg Thr 145 150 155 160 Gly Ala Arg Ala Gly Arg Leu Met Gly Thr Pro Arg Phe Lys Lys Lys 165 170 175 Gly Phe Ser Arg Asp Ser Phe Thr Ile Pro Ala Pro Glu Ala Met Gly 180 185 190 Ala Tyr Gly Ala Ala Tyr Leu Arg Gly Glu Pro Ala Tyr Val Arg Gly 195 200 205 Arg Arg Thr Ile Thr Asp Tyr Arg His Val Arg Leu Ser His Leu Gly 210 215 220 Val Leu Arg Thr Tyr Asp Ser Thr Lys Pro Leu Val Lys Ala Val Ala 225 230 235 240 Ala Gly Ala His Ile Lys Ser Tyr Thr Val Ser Arg His Ala Glu Arg 245 250 255 Trp Tyr Val Ser Phe Leu Val Glu Leu Ser Gln Ser Pro Arg Arg Gln 260 265 270 Ala Ser Lys Arg Ala Arg Ala Ala Gly Thr Val Gly Val Asp Leu Gly 275 280 285 Val Arg Tyr Leu Ala Ser Leu Ser Asp Pro Gln Ala Pro Gln Arg Leu 290 295 300 Thr Gln Leu Asp Phe Leu Ala Asp Thr Pro Ser Ile Lys Asn Pro Arg 305 310 315 320 Trp Ala Asp Arg Ala Ala Arg Arg Leu His Arg Leu Gln Arg Ala Leu 325 330 335 Ala Arg Ser Gln Lys Gly Ser Val Arg Arg Ala Arg Ile Arg Gln Gln 340 345 350 Ile Ala Arg Leu His His Lys Thr Ala Leu Arg Arg Glu Ser Asn Leu 355 360 365 His Gln Leu Thr Lys Arg Leu Ala Thr Gly Tyr Thr Leu Ile Gly Val 370 375 380 Glu Asp Leu Asn Val Ala Gly Met Thr Ala Ser Ala Arg Gly Thr Val 385 390 395 400 Glu His Pro Gly Arg Arg Val Ala Gln Lys Ser Gly Leu Asn Arg Val 405 410 415 Val Leu Asp Ala Gly Phe Gly Val Phe Arg Arg Gln Leu Glu Tyr Lys 420 425 430 Ser Thr Trp Tyr Gly Ser Ala Val Gln Glu Ile Asp Arg Tyr Tyr Pro 435 440 445 Ser Ser Gln Ile Cys Ser Gln Cys Gln Arg Lys Ala Lys Thr Lys Leu 450 455 460 Ser Leu Arg Glu Arg Val Phe Glu Cys Ala Leu Cys Gly Tyr Lys Ala 465 470 475 480 Asp Arg Asp Phe Asn Ala Ala Val Asn Ile Ala Val Gln Ala Gln Arg 485 490 495 Val Phe Glu Glu Lys Leu Ala Ser Glu Ser Gly Glu Ser Leu Asn Gly 500 505 510 Arg Gly Gly Gly Arg Thr His Lys Gly Ala Val Val Ile Glu Ala Ser 515 520 525 Arg Pro Pro Ala Gly Tyr Gly Ser Gly Ser Pro Gln Val Ser Asn His 530 535 540 Leu Pro Ile Arg Thr 545 <210> 22 <211> 541 <212> PRT <213> Artificial Sequence <400> 22 Met Thr Ala Thr Ala Ala Ala Glu Thr Val Met Asp Leu Arg Ala Tyr 1 5 10 15 Lys Phe Arg Leu Asp Pro Asn Ile Ala Gln Thr Gln Ala Leu Ser Gln 20 25 30 Ala Ala Gly Ala Ala Arg Val Ala Tyr Asn Met Val Ile Ala His Asn 35 40 45 Arg Ala Ala His Asp Glu Gly Arg Arg Arg Val Glu Glu Leu Thr Ala 50 55 60 Ser Gly Val Asp Pro Glu Glu Ala Ala Ala Arg Val Arg Ala Gln Arg 65 70 75 80 Thr Thr Asp Pro Ala Leu Ala Met Val Phe Thr Lys Met Gly Phe Asn 85 90 95 Ala Ile Leu Thr Ala Glu Arg Arg Arg His Gln Glu Ala Ala Glu Ala 100 105 110 Ile Ala Ala Gly Ala Asp Pro Ala Gln Val Trp Ala Gly Gln Arg Tyr 115 120 125 Glu Arg Pro Trp Leu His Asp Ala Asn Arg Arg Val Met Val Ser Gly 130 135 140 Val Glu His Ala Thr Ala Ala Phe Thr Asn Trp Val Asn Ser Tyr Thr 145 150 155 160 Gly Gln Arg Ala Gly Arg Arg Val Gly Tyr Pro Arg Phe Lys Lys Lys 165 170 175 Thr Val Ser Arg Asp Ser Phe Thr Ile Pro Ala Pro Glu Leu Met Gly 180 185 190 Gly Lys Gly Ala Pro Tyr Lys Arg Gly Glu Ala Arg Lys Gly Leu Ile 195 200 205 Thr Asp Tyr Arg His Leu Arg Leu Gly Phe Leu Gly Thr Ile Arg Thr 210 215 220 His Gln His Thr Arg Arg Leu Val Arg Ala Cys Gln Arg Gly Ala Val 225 230 235 240 Met Arg Ser Phe Thr Val Ser Arg Ala Ala Asp Arg Trp Tyr Val Ser 245 250 255 Ile Leu Val Ala Thr Pro Arg Glu Thr Ala Pro Glu Pro Thr Arg Ala 260 265 270 Gln Arg Ala Arg Gly Gly Val Gly Val Asp Leu Gly Val Asn Asn Leu 275 280 285 Ala Ala Leu Ser Thr Gly Glu Leu Val Asp Asn Pro Arg Tyr Gly Arg 290 295 300 Arg Gly Ala Ala Lys Leu Arg Arg Ala Gln Arg Ala Tyr Ala Arg Thr 305 310 315 320 Lys Lys Gly Ser Gln Gly Arg Ala Arg Ala Ala Arg Arg Val Ala Arg 325 330 335 Leu His His Leu Val Ala Leu Gln Arg Ser Thr Val Val Asn Thr Leu 340 345 350 Thr Lys Arg Leu Ala Thr Gly Tyr Ala Val Val Ala Ile Glu Asp Leu 355 360 365 Asn Val Ser Gly Met Thr Arg Ser Ala Lys Gly Thr Val Glu Asn Pro 370 375 380 Gly Lys Asn Val Ala Ala Lys Ser Gly Leu Asn Arg Ser Val Leu Asp 385 390 395 400 Ala Ala Phe Gly Glu Leu Arg Arg Gln Leu Glu Tyr Lys Thr Ala Trp 405 410 415 Tyr Gly Ser Ser Val Ala Thr Ile Gly Arg Tyr Glu Pro Ser Ser Lys 420 425 430 Ala Cys Ser Ser Cys Gly Ala Val Lys Pro Lys Leu Ser Leu Gly Glu 435 440 445 Arg Val Tyr Asp Cys Asp Gln Cys Gly Met Arg Leu Asp Arg Asp Val 450 455 460 Asn Ala Ala Arg Asn Ile Leu Ala Trp Ala Thr Gly Gly Arg Asp Gln 465 470 475 480 Pro Pro Ala Pro Gly Thr Gly Ala Asp Gly Arg Gly Glu Pro Ala Pro 485 490 495 Pro Ser Pro Pro Gly Glu Gly Pro Pro Ala Pro Val Gln Arg Ser Ala 500 505 510 Lys Thr Gly Tyr Val Pro His Pro Ala Thr Ala Ala Glu Gln Ser Ala 515 520 525 Ala His Pro His Gln Ser Arg Thr Pro Ile Pro Pro Ser 530 535 540 <210> 23 <211> 558 <212> PRT <213> Artificial Sequence <400> 23 Met Ala Glu Lys Thr Gly Thr Asp Ala Gly Thr Met Asn Arg Ala Tyr 1 5 10 15 Lys Phe Arg Leu Asp Pro Asn Gln Ala Gln Lys Ala Glu Leu Met Arg 20 25 30 Cys Val Gly Ala Ala Arg Tyr Thr Tyr Asn Leu Leu Asn Ala Tyr Asn 35 40 45 Leu Gln Ile Leu Arg Asn Glu Gln Glu Tyr Arg Asn Thr Arg Asn Ala 50 55 60 Glu Gly Ala Asp Tyr Glu Thr Ile Asn Gly Glu Ile Arg Lys Leu Arg 65 70 75 80 Lys Lys Asp Pro Ala Tyr Lys Phe Leu Gly His Ala Glu Tyr Glu Lys 85 90 95 Arg Tyr Leu Thr Pro Glu Lys Gln Arg His Glu Ala Ile Ala Gln Ala 100 105 110 Ile Ala Asp Gly Ala Asp Pro Ala Ile Val Trp Pro Glu Thr Glu Arg 115 120 125 Phe Thr Glu Pro Trp Leu His Thr Ile Ala Arg Arg Val Leu Val Ser 130 135 140 Gly Ile Lys Asn Ala Asp Lys Ala Trp Asp Asn Tyr Asn Lys Ser Arg 145 150 155 160 Met Lys Gln Arg Ala Gly Ala Arg Ile Gly Ile Pro Arg Phe Lys Arg 165 170 175 Lys Gly Val Ser Arg Asp Ser Phe Thr Val Pro His Glu Thr Thr Gly 180 185 190 Ala Tyr Gly Ala Tyr Tyr His Lys Lys Asp Pro Glu Tyr Ala Arg Arg 195 200 205 Lys Ala Gln Leu Lys Arg Arg Gly Ile Ser Val Lys Pro Thr Ile Thr 210 215 220 Asp Tyr Arg His Val Arg Leu Ala Ser Leu Gly Val Ile Arg Thr His 225 230 235 240 Asn Thr Thr Arg Pro Leu Val Lys Ala Val Arg Ala Gly Ala Glu Ile 245 250 255 Lys Ser Phe Thr Val Ser Arg Ala Ala Asp Arg Trp Tyr Val Ser Ile 260 265 270 Leu Val Glu Leu Thr Arg Pro Ser Thr Ala Pro Thr Arg Ala Gln Arg 275 280 285 Ala Ala Gly Ala Val Gly Val Asp Leu Gly Val Arg Tyr Leu Ala Ala 290 295 300 Leu Ser Asp Glu Gln Ala Pro Gln Arg Phe Ala Arg Tyr Pro Asp Leu 305 310 315 320 Glu Phe Thr Gly Asp Gly Asp Pro Thr Leu Ala Asn Pro Arg Trp Ala 325 330 335 Arg Thr Ala Glu Lys Arg Leu Val Arg Leu Gln Arg Ala Leu Ala Arg 340 345 350 Ala Gln Lys Gly Ser Asn Arg Arg Ala Arg Ile Val Gln Gln Ile Ala 355 360 365 Arg His His His Leu Val Ala Leu Arg Arg Glu Ser Gly Leu His Gln 370 375 380 Val Ser Lys Arg Leu Thr Thr Gly Tyr Thr Leu Ile Gly Leu Glu Asp 385 390 395 400 Leu Ala Val Ala Gly Met Thr Ala Ser Ala Ala Gly Thr Ile Glu Ala 405 410 415 Pro Gly Lys Asn Val Arg Gln Lys Ala Gly Leu Asn Arg Ser Ile Leu 420 425 430 Asp Ala Ala Leu Ser Ala Leu Arg Arg Gln Leu Glu Tyr Lys Ser Ser 435 440 445 Trp Tyr Gly Ser Gln Val Gln Ile Ile Asp Arg Phe Phe Ala Ser Ser 450 455 460 Gln Thr Cys Ser Ala Cys Gly Ala Arg Ala Lys Thr Lys Leu Gly Leu 465 470 475 480 His Val Arg Val Phe Glu Cys Ala Ala Cys Gly Ala Arg Ile Asp Arg 485 490 495 Asp Val Asn Ala Ala Arg Asn Ile Arg Ala Glu Ala Val Arg Met Tyr 500 505 510 Glu Ala Gln Leu Ala Pro Gly Thr Gly Glu Ser Leu Asn Gly Arg Gly 515 520 525 Ala Thr Asp Ser Asp Ala Ala Gly Ser Val Val Leu Gly Asp Ala Ala 530 535 540 Leu Asp Ala Ser Arg Pro Ala Ala Thr Gly Gly Gly Ser Pro 545 550 555 <210> 24 <211> 559 <212> PRT <213> Artificial Sequence <400> 24 Met Ala Lys Lys Ile Ser Thr Asp Ala Asp Thr Thr Leu Arg Ala Tyr 1 5 10 15 Lys Phe Arg Leu Asp Pro Thr Gln Thr Gln Lys Ile Ala Leu Ala Gln 20 25 30 Cys Ala Gly Ala Ala Arg Tyr Ala Tyr Asn Leu Leu Thr Ala His Asn 35 40 45 Leu Glu Val Ser Arg Val Arg Ala Asp Tyr Trp His Thr Arg Ile Asp 50 55 60 Ala Gly Ala Glu Glu Thr Thr Val Lys Ala Glu Leu Lys Ala Leu Ala 65 70 75 80 Lys Ser Asp Ala Ala Tyr Lys Ser Ile Gly Tyr Ala Ala Tyr Gly Thr 85 90 95 Gln Tyr Leu Thr Pro Glu Ile Asn Arg His Arg Ala Ala Ser Asn Ala 100 105 110 Ile Thr Asp Gly Ala Asp Pro Ala Ala Val Trp Pro Glu Thr Glu Arg 115 120 125 Ser Ser Glu Pro Trp Met His Thr Ile Pro Arg Arg Val Leu Val Ser 130 135 140 Gly Leu Gln Ser Ala Asp Lys Ala Trp Lys Asn Phe Phe Asp Ser Arg 145 150 155 160 Thr Gly Ala Arg Ala Gly Arg Leu Val Gly Thr Pro Arg Phe Lys Arg 165 170 175 Lys Gly Ala Ser Arg Asp Ser Phe Thr Ile Pro Ala Pro Glu Thr Ile 180 185 190 Gly Ala Tyr Gly Thr Ala Tyr Leu Arg Gly Glu Pro Glu Tyr Ala Arg 195 200 205 Arg Lys Ala Gln Leu Lys His Arg Gly Ile Lys Thr Ala Pro Thr Ile 210 215 220 Glu Asp Tyr Arg His Val Arg Leu Ser His Leu Gly Val Ile Arg Thr 225 230 235 240 His Asn Thr Thr Lys Pro Leu Val Lys Ala Val Arg Ala Gly Ala Gln 245 250 255 Ile Arg Ser Tyr Thr Val Ser Arg Ala Ala Asp Arg Trp Tyr Val Ser 260 265 270 Ile Leu Val Glu Leu Thr Arg Pro Ser Ala Thr Pro Thr Arg Ala Gln 275 280 285 Arg Ser Ala Gly Ala Val Gly Val Asp Leu Gly Val Arg Tyr Leu Ala 290 295 300 Ala Leu Ser Asp Glu Gln Ala Pro Gln Arg Phe Ala Arg Phe Pro Ser 305 310 315 320 Leu Glu Phe Thr Gly Asp Gly Ala Pro Thr Leu Ala Asn Pro Arg Trp 325 330 335 Ala Arg Ala Ala Glu Lys Arg Leu Val Arg Leu Gln Arg Ala Leu Ser 340 345 350 Arg Ala Gln Lys Gly Ser Asn Arg Arg Ala Arg Ile Val Gln Gln Ile 355 360 365 Ala Arg His His His Leu Val Ala Leu Arg Arg Glu Ser Gly Leu His 370 375 380 Gln Val Ser Lys Arg Leu Ser Thr Gly Tyr Thr Leu Ile Gly Leu Glu 385 390 395 400 Asp Leu Ala Val Ala Gly Met Thr Ala Ser Ala Ala Gly Thr Ile Glu 405 410 415 Ala Pro Gly Lys Asn Val Arg Gln Lys Ala Gly Leu Asn Arg Ser Ile 420 425 430 Leu Asp Ala Ala Phe Ser Thr Leu Arg Arg Gln Leu Glu Tyr Lys Ala 435 440 445 Ser Trp Tyr Gly Ser Gln Val Gln Ile Ile Asp Arg Phe Phe Ala Ser 450 455 460 Ser Gln Thr Cys Ser Ala Cys Gly Ala Arg Ala Lys Thr Lys Leu Asp 465 470 475 480 Leu Arg Val Arg Val Phe Glu Cys Ala Ala Cys Gly Val Arg Ile Asp 485 490 495 Arg Asp Val Asn Ala Ala Arg Asn Ile Arg Ala Glu Ala Val Arg Met 500 505 510 Tyr Glu Ala Gln Leu Ala Pro Gly Thr Gly Glu Ser Leu Asn Gly Arg 515 520 525 Gly Val Ala Gly Ser Asp Ala Ala Gly Leu Val Val Leu Gly Asp Val 530 535 540 Ala Leu Asp Ala Ser Arg Pro Ala Ala Thr Gly Gly Gly Ser Pro 545 550 555 <210> 25 <211> 559 <212> PRT <213> Artificial Sequence <400> 25 Met Ala Lys Lys Thr Ser Thr Asp Ala Glu Thr Met Asn Arg Ala Tyr 1 5 10 15 Lys Phe Arg Leu Asp Pro Thr Gln Ala Gln Lys Ala Glu Leu Met Arg 20 25 30 Cys Ala Gly Ala Ala Arg Cys Thr Tyr Asn Leu Leu Asn Glu Tyr Asn 35 40 45 Leu Gln Ile Leu Arg Asn Glu Gln Glu Tyr Arg Arg Thr Arg Ala Ala 50 55 60 Glu Gly Ile Asp His Lys Thr Ile Thr Ser Glu Leu Lys Lys Leu Gly 65 70 75 80 Lys Lys Asp Pro Glu Tyr Lys Phe Leu Gly Gln Ala Glu Tyr Glu Lys 85 90 95 Arg Tyr Leu Thr Pro Glu Lys Gln Arg His Glu Ala Ile Ala Gln Ala 100 105 110 Ile Ala Asp Gly Ala Asp Pro Ala Ile Val Trp Pro Glu Thr Glu Arg 115 120 125 Phe Thr Glu Pro Trp Leu His Thr Ile Ala Arg Arg Val Leu Val Ser 130 135 140 Gly Ile Lys Asn Ala Asp Lys Ala Trp Asp Asn Tyr Asn Lys Ser Arg 145 150 155 160 Met Lys Gln Arg Ala Gly Ala Arg Met Gly Val Pro Arg Arg Lys Arg 165 170 175 Lys Gly Val Ser Arg Asp Ser Phe Thr Ile Pro Ala Pro Glu Ala Ile 180 185 190 Gly Ala Tyr Gly Thr Tyr Tyr His Lys Lys Asp Pro Glu Tyr Ala Arg 195 200 205 Arg Lys Ala Gln Leu Lys Arg Arg Gly Ile Asn Ala Ala Pro Thr Ile 210 215 220 Glu Asp Tyr Arg His Val Arg Leu Ser His Leu Gly Val Ile Arg Thr 225 230 235 240 His Asn Thr Thr Lys Pro Leu Val Lys Ala Val Arg Ala Gly Ala Gln 245 250 255 Ile Arg Ser Tyr Thr Val Ser Arg Thr Ala Asp Arg Trp Tyr Val Ser 260 265 270 Ile Leu Val Glu Leu Thr Arg Pro Ser Ala Thr Pro Thr Arg Ala Gln 275 280 285 Arg Ala Ala Gly Ala Val Gly Val Asp Leu Gly Val Arg Tyr Leu Ala 290 295 300 Ala Leu Ser Asp Glu Trp Ala Pro Gln Arg Phe Ala Arg Tyr Pro Ser 305 310 315 320 Leu Glu Phe Thr Gly Asp Gly Ala Pro Ile Leu Ala Asn Pro Cys Trp 325 330 335 Ala Arg Ala Ala Glu Arg Arg Leu Val Arg Leu Gln Arg Ala Leu Ala 340 345 350 Arg Ala Gln Lys Gly Ser Arg Arg Arg Ala Arg Leu Val Gln Gln Ile 355 360 365 Ala Arg His His His Leu Val Ala Leu Arg Arg Glu Ser Gly Leu His 370 375 380 Gln Val Ser Lys Arg Leu Ser Thr Gly Tyr Thr Leu Ile Gly Leu Glu 385 390 395 400 Asp Leu Ala Val Thr Gly Met Thr Ala Ser Ala Ala Gly Thr Val Glu 405 410 415 Ala Pro Gly Lys Asn Val Arg Gln Lys Ala Gly Leu Asn Arg Ser Ile 420 425 430 Leu Asp Ala Ala Phe Ser Thr Leu Arg Arg Gln Leu Glu Tyr Lys Ser 435 440 445 Gly Trp Tyr Gly Ser Gln Val Gln Ile Ile Asp Arg Phe Phe Ala Ser 450 455 460 Ser Gln Thr Cys Ser Ala Cys Gly Ala Arg Ala Lys Thr Lys Leu Asp 465 470 475 480 Leu Arg Val Arg Val Phe Glu Cys Ala Ala Cys Gly Val Arg Ile Asp 485 490 495 Arg Asp Val Asn Ala Ala Arg Asn Ile Arg Ala Glu Ala Val Arg Met 500 505 510 Tyr Glu Ala Gln Leu Ala Pro Gly Thr Gly Glu Ser Leu Asn Gly Arg 515 520 525 Gly Ala Ala Gly Ser Asp Ala Ala Gly Ser Val Val Leu Gly Asp Val 530 535 540 Ala Leu Asp Ala Ser Arg Pro Ala Ala Thr Gly Gly Gly Ser Pro 545 550 555 <210> 26 <211> 485 <212> PRT <213> Artificial Sequence <400> 26 Met Thr Thr Glu Ile Ala His Lys Ala Phe Lys Tyr Arg Leu Lys Pro 1 5 10 15 Thr Gln Thr Gln Glu Asn Ala Phe Arg Gln Ala Val Gly Ala Ala Arg 20 25 30 Val Ala Tyr Asn Met Leu Thr Ala Leu Asn Arg Asp Val Leu Arg Arg 35 40 45 Gly Trp Glu Ile Arg Asn Gln Leu Ile Glu Gln Gly Leu Ser Glu Ala 50 55 60 Glu Ala Ser Ala Gln Met Lys Lys Leu Arg Lys Glu Asp Leu Ser Leu 65 70 75 80 Lys Ile Met Ser Ser Phe Ala Trp Gln Lys Asp Val Leu Thr Pro Glu 85 90 95 Ile Ala Arg His Lys Gln Ala Ala Ala Arg Ile Ala Ala Gly Asp Pro 100 105 110 Ile Glu Ser Val Trp Ser Gly Glu Arg Phe Ala Glu Pro Trp Phe His 115 120 125 Ala Val Pro Arg Arg Ala Phe Val Ser Gly Ala Asp Gln Ala Ala Thr 130 135 140 Ala Ile Ser Asn Trp Met Lys Ser Val Ala Gly Gln Arg Ala Gly Asn 145 150 155 160 Lys Met Gly Leu Pro Arg Phe Lys Lys Ala Gly Arg Ser His Asp Ser 165 170 175 Ile Thr Ile Pro Val Val Asp Gly Ser Ser Pro Ala Gly Gly Tyr Gly 180 185 190 Ala Ser Tyr Lys Arg Gly Glu Pro Arg Lys Gly Phe Ile Gln Asp His 195 200 205 His Arg Leu Arg Leu Ala Met Phe Gly Thr Val Val Thr Tyr Asn Ser 210 215 220 Thr Ala Pro Leu Thr Gln Leu Val Thr Gln Gly Gly Thr Ile Lys Ser 225 230 235 240 Phe Thr Ile Ser Arg Asp Ala Gln Tyr Trp Tyr Val Ser Leu Leu Val 245 250 255 Glu Ala Pro Ile Asn Leu Leu His Arg Pro Gly Pro Thr Lys Ala Gln 260 265 270 Glu Lys Ala Gly Val Ile Gly Val Asp Leu Gly Val Lys Ala Lys Ala 275 280 285 Ala Leu Ser Asn Gly Glu Ile Ile Asn Asn Pro Arg His Leu Gln Leu 290 295 300 Ser Leu Lys Arg Leu Ala Lys Leu Gln Val Lys Leu Ser Arg Thr Gln 305 310 315 320 Lys Gly Ser Lys Asn Arg Glu His Leu Lys Lys Gln Ile Ser Arg Leu 325 330 335 Gln His Lys Ile Ala Leu Gln Arg Ala Ala Ser Asn His Gln Met Thr 340 345 350 Lys Glu Leu Ser Thr Thr Tyr Val Gly Val Gly Ile Glu Asp Leu Asn 355 360 365 Val Ala Gly Met Thr Lys Ser Ala Ser Gly Thr Leu Glu Asn Pro Gly 370 375 380 Lys Asn Val Ala Ala Lys Ser Gly Leu Asn Arg Asn Ile Leu Asp Val 385 390 395 400 Ser Phe Gly Gln Ile Arg Asn Gln Leu Glu Tyr Lys Thr Lys Met Tyr 405 410 415 Gly Ser Ala Leu Thr Val Ile Gly Arg Tyr Glu Pro Thr Ser Lys Arg 420 425 430 Cys Ser Gln Cys Gly Ala Met Lys Ser Asp Leu Ser Leu Lys Glu Arg 435 440 445 Thr Tyr Ile Cys Val Glu Cys Asn Tyr Gln Asp Asp Arg Asp Ile Asn 450 455 460 Ala Ala Lys Asn Ile Arg Asn Leu Ala Glu Lys Glu Leu Gly Met Val 465 470 475 480 Ser Ser Ile Ser Ser 485 <210> 27 <211> 562 <212> PRT <213> Artificial Sequence <400> 27 Met Ala Lys Lys Thr Ser Thr Asp Asp Gly Ala Leu Tyr Arg Ala Tyr 1 5 10 15 Lys Phe Arg Leu Asp Pro Asn Gln Ala Gln Lys Ala Lys Leu Met Gln 20 25 30 Cys Val Gly Ala Ala Arg Tyr Thr Tyr Asn Leu Phe Thr Asp His Asn 35 40 45 Arg Lys Val Ser Arg Ala Cys Ala Glu Tyr Val Asp Thr Arg Ile Ser 50 55 60 Glu Gly Ala Thr Glu Val Ile Ala Lys Ala Glu Leu Asp Asp Ile Lys 65 70 75 80 Glu Glu Asn Pro Ala Tyr Lys Tyr Ile Gly His Phe Glu Tyr Gly Lys 85 90 95 Arg Tyr Leu Thr Pro Glu Arg Gln Arg His Glu Ala Ile Ala Gln Ala 100 105 110 Ile Ala Asp Gly Ala Asp Pro Ala Val Val Trp Pro Glu Thr Glu Arg 115 120 125 Ser Ser Glu Pro Trp Met His Thr Ile Asp Arg Arg Val Leu Val Ser 130 135 140 Gly Ile Lys Asn Ala His Ala Ala Trp Gly Asn Tyr Phe Glu Ser Leu 145 150 155 160 Lys Gly Leu Arg Ala Gly Pro Arg Met Gly Val Pro Arg Phe Lys Arg 165 170 175 Lys Gly Ile Ser His Asp Ser Phe Thr Val Asp Arg Asp Pro Arg Asn 180 185 190 Asp Glu Leu Gly Gly Tyr Gly Thr Pro Tyr Lys Lys Gly Thr Pro Glu 195 200 205 Tyr Ala Arg Arg Thr Ala Gln Leu Lys Arg Arg Gly Ile Lys Thr Ile 210 215 220 Pro Thr Ile Glu Asp Tyr Arg His Val Arg Leu Ala Ser Leu Gly Val 225 230 235 240 Ile Arg Thr His Asn Ser Thr Lys Pro Leu Val Lys Ala Val Arg Ala 245 250 255 Gly Ala Asp Val Lys Ser Phe Thr Val Ser Arg Ala Ala Asp Arg Trp 260 265 270 Tyr Val Ser Ile Leu Val Lys Leu Ser Arg Thr Pro Val Thr Pro Thr 275 280 285 Arg Ala Gln Arg Ala Ala Gly Ala Val Gly Val Asp Leu Gly Val Arg 290 295 300 Tyr Leu Ala Ala Leu Ser Asp Glu Gln Ala Pro Gln Arg Phe Ala Gln 305 310 315 320 Tyr Pro Ser Leu Glu Phe Thr Gly Asp Gly Ser Pro Thr Leu Ala Asn 325 330 335 Pro Arg Trp Ala Arg Ala Ala Glu Lys Arg Leu Val Arg Leu Gln Arg 340 345 350 Ala Leu Ala Arg Ala Gln Lys Gly Ser Asn Arg Arg Ala Arg Ile Val 355 360 365 Gln Lys Ile Ala Arg His His His Leu Val Ala Leu Arg Arg Glu Ser 370 375 380 Gly Leu His Gln Val Ser Lys Arg Leu Ala Thr Gly Tyr Thr Leu Ile 385 390 395 400 Gly Leu Glu Asp Leu Ala Val Ala Gly Met Thr Ala Ser Ala Ala Gly 405 410 415 Thr Ile Glu Ala Pro Gly Lys Asn Val Arg Gln Lys Ala Gly Leu Asn 420 425 430 Arg Ser Ile Leu Asp Ala Ala Phe Ser Thr Leu Arg Arg Gln Leu Glu 435 440 445 Tyr Lys Ala Ser Trp Tyr Gly Ser Gln Val Gln Ile Ile Asp Arg Phe 450 455 460 Phe Ala Ser Ser Gln Thr Cys Ser Ala Cys Gly Ala Arg Val Lys Thr 465 470 475 480 Lys Leu Gly Leu His Val Arg Val Phe Glu Cys Ala Ala Cys Gly Val 485 490 495 Arg Ile Asp Arg Asp Val Asn Ala Ala Arg Asn Ile Arg Ala Glu Ala 500 505 510 Val Arg Met Cys Glu Ala Gln Leu Ala Pro Gly Thr Gly Glu Ser Leu 515 520 525 Asn Gly Arg Gly Val Ala Asp Ser Asn Val Ala Gly Ser Val Val Leu 530 535 540 Gly Asp Ala Ala Leu Asp Ala Ser Arg Pro Ala Ala Thr Gly Gly Gly 545 550 555 560 Ser Pro <210> 28 <211> 494 <212> PRT <213> Artificial Sequence <400> 28 Met Ser Ser Gly Thr Thr Ala Arg Arg Met Arg Ala Tyr Arg Phe Thr 1 5 10 15 Leu Asp Pro Thr Arg Thr Gln Leu Asp Thr Leu Ala Gln His Ala Gly 20 25 30 Ala Ala Arg Trp Ala Tyr Asn His Ala Leu Ala Thr Lys Leu Asp Ala 35 40 45 Leu Arg Arg Arg Gln Leu Gln Ile Asp Glu Leu Val Thr Leu Gly Leu 50 55 60 Thr Glu Gln Gln Ala Arg Arg Asp Ala Thr Val Thr Val Pro Ser Lys 65 70 75 80 Pro Gln Val Gln Lys Arg Trp Asn Gln Leu Lys Gly Asp Thr Thr Arg 85 90 95 Gly Gly Asp Gly Ile Cys Pro Trp Trp Arg Ala Val Ser Thr Tyr Ala 100 105 110 Phe Gln Ser Ala Phe Leu Asp Ala Asp Arg Ala Trp Lys Asn Trp Met 115 120 125 Asp Ser Leu Thr Gly Lys Arg Ala Gly Arg Arg Val Gly Ala Pro Arg 130 135 140 Phe Lys Lys Lys Gly Arg Cys Arg Asp Ser Phe Arg Leu His His Ala 145 150 155 160 Val Asn Asp Pro Thr Ile Arg Leu Glu Gly Tyr Arg Arg Leu Arg Met 165 170 175 Pro Arg Leu Gly Ser Ile Arg Leu His Asp Ser Gly Lys Arg Leu Ala 180 185 190 Arg Ala Leu Ala Arg Gly Gly Arg Val Gln Ser Val Thr Val Ser Arg 195 200 205 Gly Gly His Arg Trp Tyr Ala Ser Val Leu Val Asp Glu Pro Asp His 210 215 220 Thr Pro Gly Arg Glu Thr Gln His Arg Pro Ser Arg Ala Gln His Thr 225 230 235 240 Ala Gly Gly Val Gly Val Asp Val Gly Val His His Leu Ala Ala Leu 245 250 255 Ser Thr Gly Glu Thr Leu Asp Asn Pro Arg His Leu His His Ala His 260 265 270 Thr Arg Leu Val Lys Ala Gln Arg Ala Leu Ala Arg Thr Gln Lys Gly 275 280 285 Ser His Arg Arg Arg Arg Ala Ala Glu Arg Val Gly Arg Leu His His 290 295 300 Gln Leu Ala Glu Arg Arg Ala Ser His Leu His Thr Ile Thr Lys Arg 305 310 315 320 Leu Ala Thr Gln His Ala Leu Val Ala Val Glu Asp Leu Asn Val Gln 325 330 335 Gly Met Thr Arg Thr Ala Arg Gly Thr Leu Thr Gln Pro Gly Arg Asn 340 345 350 Val Arg Ala Lys Ala Gly Leu Asn Arg Ala Ile Leu Asp Ala Ala Pro 355 360 365 Gly Glu Leu Arg Arg Gln Leu Glu Tyr Lys Ala Ser Trp Tyr Gly Ser 370 375 380 Thr Leu Ala Val Cys Asp Arg Trp Ala Pro Thr Ser Lys Thr Cys Ser 385 390 395 400 Thr Cys Gly Thr Val Lys Thr Lys Leu Pro Leu Ser Thr Arg Val Tyr 405 410 415 Arg Cys Asp Thr Cys Gly Met Val Cys Asp Arg Asp Ile Asn Ala Ala 420 425 430 Arg Asn Ile Leu Lys Asp Ala Asp Pro Val Ala Pro Gly Arg Gly Glu 435 440 445 Thr Leu Asn Ala Cys Gly Gly Pro Val Ser Pro Ser Ala Pro His Ala 450 455 460 Pro Thr Ala Arg Pro Ser Glu Ala Gly Arg Pro Gly His Pro Ala Arg 465 470 475 480 Ser Pro Arg Gly Asn Asp Pro Pro Ser Thr Pro Lys Thr Glu 485 490 <210> 29 <211> 606 <212> PRT <213> Artificial Sequence <400> 29 Met Ala Gln Ala Glu Ala Pro Arg Arg Leu Arg Ala Tyr Lys Phe Ala 1 5 10 15 Leu Asp Pro Thr Glu Ala Gln Leu Arg Glu Phe Glu Gln His Ala Gly 20 25 30 Ser Ala Arg Trp Ala Tyr Asn His Ala Asn Ala Ile Leu Ser Arg Tyr 35 40 45 Ser Asp Thr Leu Arg Asn Arg Trp Asn Ala Trp Ile Ala Gln His His 50 55 60 Gly Leu Ser Arg Glu Gln Leu Tyr Ala Leu Pro Asp Arg Glu Arg Thr 65 70 75 80 Ala Ile Gln Ala Ala Ala Arg Ala Ala Val Lys Ala Glu Asn Ala Gln 85 90 95 Leu Ala Ala Glu Leu Arg Ile Ile Asp Asp His Arg Lys Arg Val Thr 100 105 110 His Lys Gly Lys Pro Ser Val Glu Pro Gly Glu Gln Pro Ala Glu Asp 115 120 125 Ala Pro Glu Arg Ala Tyr Gln Leu Trp Arg Glu Arg Val Glu Leu Ala 130 135 140 Arg Leu His Ala Glu Asp Pro Gln Ala Tyr Arg Ala Glu Arg Lys Arg 145 150 155 160 Ile Leu Asp Glu Ile Arg Pro Leu Val Asn Ala Thr Lys Arg Lys Leu 165 170 175 Ile Glu Gln Gly Ala Tyr Arg Pro Thr Ala Met Asp Ile Ser Thr Leu 180 185 190 Trp Arg Glu Ile Arg Asp Leu Pro Pro Asp Glu Gly Gly Ser Pro Trp 195 200 205 Trp Pro Glu Val Ser Ile Tyr Ala Phe Thr Ser Gly Phe Ala His Ala 210 215 220 Glu Thr Ala Trp Lys Asn Tyr Leu Glu Ser Leu Ala Gly Arg Arg Ala 225 230 235 240 Gly Arg Pro Val Gly Lys Pro Arg Phe Lys Lys Lys Arg Arg Ser Arg 245 250 255 Arg Ser Phe Thr Leu Tyr Gly Ser Val Lys Leu Val Thr Tyr Arg Arg 260 265 270 Ile Gln Val Pro Ser Ile Gly Ser Val Arg Leu His Gly Ser Ala Lys 275 280 285 Arg Leu His Arg Ala Leu Glu Arg Arg Gly Gly Ile Ile Lys Ser Ile 290 295 300 Thr Ile Ser Gln Gly Gly His Arg Trp Tyr Ala Ser Val Leu Val Asp 305 310 315 320 Glu Leu Asp Ile Thr Pro Gly Arg Glu Thr Gln Arg Gly Pro Ser Arg 325 330 335 Arg Gln Arg Asp Arg Gly Ala Val Gly Val Asp Leu Gly Val His His 340 345 350 Leu Val Ala Leu Ser Asp Pro Asn Glu Lys Thr Leu Asp Asn Pro Arg 355 360 365 His Leu Arg Lys Ala Arg Lys Arg Leu Leu Lys Ala Gln Arg Ala Met 370 375 380 Ser Arg Arg Arg Gly Pro Asp Lys Arg Thr Gly Gln Glu Pro Ser Arg 385 390 395 400 Arg Trp Val Lys Ala Arg Asn Arg Val Ala Arg Leu His His Glu Leu 405 410 415 Ala Val Arg Arg Ala Gly His Leu His Glu Ile Thr Lys Arg Leu Ala 420 425 430 Thr Ser Tyr Glu Leu Val Ala Ile Glu Asp Leu Asn Val Ala Gly Met 435 440 445 Thr Arg Ser Ala Arg Gly Thr Ile Asp Gln Pro Gly Arg Gly Val Arg 450 455 460 Ala Lys Ala Gly Leu Asn Arg Ser Ile Leu Asp Thr Ser Pro Ala Glu 465 470 475 480 Phe Arg Arg Gln Leu Gln Tyr Lys Ala Ser Trp Tyr Gly Ala Thr Val 485 490 495 Ala Val Ile Asp Arg Trp Ala Pro Thr Ser Arg Thr Cys Ser Ser Cys 500 505 510 Gly Ala Val Lys Ala Lys Leu Ser Leu Ala Glu Arg Thr Phe Phe Cys 515 520 525 Glu His Cys Gly Met Glu Leu Asp Arg Asp Ile Asn Ala Ala Arg Asn 530 535 540 Ile Leu Ala Phe Ala Gln Ser Ala Tyr Pro Gly Glu Gly Lys Ala Leu 545 550 555 560 Asn Ala Cys Gly Gly Ser Val Ser Pro Gly Ser Gln Ser Val Val Gln 565 570 575 Ala Gly Ala Asp Glu Ala Gly Arg Pro Ala Arg Lys Pro Arg Arg Ser 580 585 590 Ser Arg Gly Ser Asp Pro Pro Ala Thr Pro Thr Thr Arg Ala 595 600 605 <210> 30 <211> 459 <212> PRT <213> Artificial Sequence <400> 30 Met Ser Cys Val Thr Lys Thr Gln Val Val Arg Ala Tyr Arg Phe Ala 1 5 10 15 Leu Asp Ala Thr Pro Ser Gln Val Glu Ile Leu Arg Arg Tyr Ala Asn 20 25 30 Ala Ser Arg Ala Ala Phe Asn Phe Ala Leu Gly Met Lys Thr Ala Ser 35 40 45 His Glu Arg Trp Lys Arg Gly Arg Asp Pro Leu Val Ala Ala Asp Met 50 55 60 Asp Met Ala Glu Ala Lys Lys Lys Ala Pro Lys Val Arg Ile Pro Arg 65 70 75 80 Lys Pro Gln Val His Arg His Phe Ile Ala Thr Lys Gly Lys Gly Pro 85 90 95 Ile Gly Pro Leu Arg Gln Gly Glu Glu Arg Arg Thr Pro Tyr Pro Trp 100 105 110 Trp Glu Gly Val Asn Ala Asn Ser Phe Gln Glu Ala Phe Phe Asp Ala 115 120 125 Asp Ala Ala Trp Lys Asn Trp Leu Asp Ser Met Thr Gly Lys Arg Ala 130 135 140 Gly Ala Pro Val Gly Tyr Pro Arg Phe Lys Arg Lys Gly Ser Arg Glu 145 150 155 160 Ser Phe Arg Leu Thr Asn Thr Ala Val Arg Phe Ala Gly Tyr Arg Arg 165 170 175 Leu Trp Ile Gly Gly Gly Gly Gly Gln Ala Ala Phe Thr Val Arg Leu 180 185 190 His Glu Pro Ala Arg Arg Leu Val Arg Leu Leu Asp Ala Gly Gln Ala 195 200 205 Arg Ile Leu Arg Ala Thr Ile Ala Arg Asp Gly His Arg Trp Phe Ala 210 215 220 Ala Val Val Val Glu Glu Ile Ala Val Leu Pro Asn Arg Pro Thr Arg 225 230 235 240 Arg Gln Ala Asn Ala Gly Arg Val Gly Val Asp Leu Gly Val Lys Ser 245 250 255 Ala Ala Ala Phe Ser Asp Pro Val Val Leu Ala His Gly Ala Pro Gly 260 265 270 Val Leu Glu Leu Ala Asn Pro Arg His Val Arg Asn Thr Glu Lys Lys 275 280 285 Leu Ala Arg Ala Gln Arg Val Ala Ser Arg Arg Phe Val Lys Gly Ala 290 295 300 Gln Lys Gln Ser Lys Gly Tyr His Glu Ala Lys Ala Arg Val Ala Lys 305 310 315 320 Leu His Ala Gln Leu Ala Ala Arg Arg Ala Ser Thr Gln His Leu Leu 325 330 335 Thr Arg Arg Leu Val Glu Gln Tyr Ala Glu Val Ala Leu Glu Lys Leu 340 345 350 Met Val Lys Asn Met Thr Arg Ser Ala Lys Gly Ser Ala Glu Asn Pro 355 360 365 Gly Lys Asn Val Arg Gln Lys Ala Gly Leu Asn Arg Ala Ile Leu Asp 370 375 380 Val Gly Phe Gly Glu Ile Arg Arg Gln Ile Glu Tyr Lys Ala Ala Trp 385 390 395 400 Tyr Gly Ser Ser Thr Arg Tyr Val Gly Thr Tyr Phe Pro Ser Ser Lys 405 410 415 Thr Cys Ser Asn Cys His Trp Lys Asn Lys Glu Leu Thr Leu Ala Asp 420 425 430 Arg Val Phe Lys Cys Thr Gln Cys Gly Met Ile Met Asp Arg Asp Ala 435 440 445 Asn Ala Ala Gln Met Ile Lys Lys His Ala Gln 450 455 <210> 31 <211> 517 <212> PRT <213> Artificial Sequence <400> 31 Met Gln Gln Glu Ile Leu Lys Ala Phe Arg Phe Ala Leu Asp Pro Thr 1 5 10 15 Pro Ser Gln Leu Glu Gly Leu Ala Gln His Ala Gly Ala Ala Arg Trp 20 25 30 Ala Phe Asn His Ala Leu Gly Met Lys Val Ala Ala His Gln Gln Trp 35 40 45 Arg Ser Ala Val Gln Glu Leu Val Asp Asn Gly Met Ser Glu Ala Gln 50 55 60 Ala Arg Lys Glu Val Arg Val Pro Val Pro Thr Lys Ala Val Val Gln 65 70 75 80 Lys His Leu Asn Gln Ile Lys Gly Asp Ser Arg Gly Asp Ala Leu Pro 85 90 95 Asp Gly Ser Leu Gly Pro Glu Arg Pro Cys Pro Trp Trp His Glu Val 100 105 110 Asn Thr His Ala Phe Gln Ser Ala Phe Ile Asp Ala Asp Arg Ala Trp 115 120 125 Lys Asn Trp Leu Asp Ser Leu Lys Gly Thr Arg Ala Gly Arg Arg Val 130 135 140 Gly Tyr Pro Arg Phe Lys Lys Lys Gly Arg Ala Arg Glu Ala Phe Arg 145 150 155 160 Leu His His Asn Val Lys Gln Pro Thr Val Arg Leu Glu Ser Tyr Arg 165 170 175 Arg Leu Arg Leu Pro Thr Ile Gly Ser Val Arg Leu His Asp Thr Ala 180 185 190 Lys Arg Met Ser Arg Ser Ile Gly Arg Gly Asp Ala Val Ile Gln Ser 195 200 205 Val Thr Val Ser Arg Ala Gly Gln Arg Trp Tyr Ala Ser Val Leu Cys 210 215 220 Lys Val Asn Ala Asp Leu Pro Asp Gln Pro Thr Arg Arg Gln Trp Asp 225 230 235 240 Arg Gly Arg Val Gly Val Asp Leu Gly Val Lys His Leu Ala Val Leu 245 250 255 Ser Gln Pro Leu Glu Thr Asp Ala Ala Thr Thr Ala Phe Val Ser Asn 260 265 270 Pro Arg His Ile Arg Lys Ala Glu Gln Gln Leu Ala Lys Ala Gln Arg 275 280 285 Ala Leu Ser Arg Thr Gln Lys Gly Ser Ala Arg Arg Ala Lys Ala Arg 290 295 300 Arg Arg Val Gly Arg Val His His Glu Val Ala Val Arg Arg Ser Thr 305 310 315 320 Val Leu His Ala Leu Thr Lys Gln Leu Ala Thr Arg Phe Ala Glu Val 325 330 335 Ala Ile Glu Asp Leu His Val Ser Gly Met Thr Arg Ser Ser Ser Gly 340 345 350 Thr Pro Asp Lys Pro Gly Arg His Val Lys Gln Lys Ala Gly Leu Asn 355 360 365 Arg Ala Ile Leu Asp Thr Ala Pro Gly Glu Leu Arg Arg Gln Leu Thr 370 375 380 Tyr Lys Thr Arg Trp Tyr Gly Ser Thr Leu Ala Val Leu Asp Arg Trp 385 390 395 400 Phe Pro Ser Ser Lys Thr Cys Ser Ala Cys Gly Trp Gln Asn Pro Arg 405 410 415 Leu Thr Leu Ala Asp Arg Thr Phe His Cys Thr Asn Cys Ser Thr Thr 420 425 430 Ile Asp Arg Asp Leu Asn Ala Ala Arg Asn Ile Ala Gln His Ala Val 435 440 445 Leu Ala Asp Gln Leu Leu Ala Pro Gly Arg Gly Glu Thr Gln Asn Ala 450 455 460 Arg Gly Ala Pro Ile Arg Pro Pro Gly Pro Arg Ala Gly Gly Gln Glu 465 470 475 480 Ala Met Lys Arg Glu Asp Thr Ser Pro Ala Arg Pro Val Pro Pro Gln 485 490 495 Arg Ser Asp Pro Leu Thr Leu Phe Thr Leu Asp Asp Thr Gly His Glu 500 505 510 Thr Ala Lys Arg Ala 515 <210> 32 <211> 469 <212> PRT <213> Artificial Sequence <400> 32 Met Ser Glu Thr Val Thr Ile Val Val Lys Glu Ala Leu Asp Pro Thr 1 5 10 15 Pro Ala Gln Val Val Ile Leu Gln Arg Tyr Ala Asp Ala Ser Arg Cys 20 25 30 Ser Phe Asn Tyr Ala Leu Gly Leu Lys Arg Gly Ala Gln Gln Thr Trp 35 40 45 Ser Gln Gly Arg Asp Arg Leu Val Ala Gln Gly Gln Thr Pro Ala Glu 50 55 60 Ala Ala Arg Asn Ala Pro Lys Val Arg Ile Pro Gly Gln Phe Asp Ile 65 70 75 80 Gln Lys Ile Phe Leu Ala Lys Arg Asp Glu Pro Leu Pro Gly Pro Leu 85 90 95 Leu Pro Gly Gln Glu Pro Arg Leu Leu Tyr Pro Trp Trp Lys Gly Val 100 105 110 Asn Ala Ile Val Cys Gln Gln Ala Phe Arg Asp Ala Asp Ala Ala Phe 115 120 125 Ala Asn Trp Lys Ser Ser Ala Arg Arg Lys Gly Gly Pro Val Gly Phe 130 135 140 Pro Arg Phe Lys Arg Arg Gly Arg Cys Arg Asp Ser Phe Arg Met Phe 145 150 155 160 Ser Ile Arg Leu Val Glu Glu Asp Leu Arg His Val Arg Ile Gly Gly 165 170 175 Glu Arg Gly Gly Gln Pro Ala Phe Thr Val Arg Leu His Arg Pro Ala 180 185 190 Arg Arg Leu Ala Arg Leu Leu Ala Glu Gly Gly Glu Thr Lys Ser Val 195 200 205 Thr Val Ala Arg Glu Gly His Arg Trp Phe Val Ala Phe Asn Val Arg 210 215 220 Val Pro Ala Gly Pro Pro Pro Arg Pro Thr Arg Arg Gln Arg Gly Ala 225 230 235 240 Gly Thr Val Gly Val Asp Leu Gly Val Lys Val Phe Ala Ala Thr Ser 245 250 255 Asp Pro Leu Leu Ile Asn Gly Thr Gly Leu Gln Leu Phe Glu Asn Pro 260 265 270 Arg Leu Leu Asp Asn Ala Arg Arg Gln Leu Arg Lys Trp Gln Arg Arg 275 280 285 Met Ala Arg Arg His Val Arg Gly Leu Arg Ala Asp Glu Gln Ser Ser 290 295 300 Gly Trp Lys Glu Ala Arg Asp Gln Val Ala Arg Leu His Ala Leu Val 305 310 315 320 Ala Ala Arg Arg Ser Ser Thr Gln His Leu Leu Thr Lys Arg Ile Val 325 330 335 Thr Gln Tyr Glu His Val Ala Leu Glu Asp Leu Arg Val Lys Asn Met 340 345 350 Thr Ala Ser Ala Arg Gly Thr Ala Gln Ser Pro Gly Arg Asn Val Lys 355 360 365 Ala Lys Ala Gly Leu Asn Arg Ala Ile Leu Asp Val Gly Phe Gly Glu 370 375 380 Ile Arg Arg Gln Ile Glu Tyr Lys Ala Leu Leu Tyr Gly Thr Thr Val 385 390 395 400 Thr Val Val Asp Pro Ala Tyr Thr Ser Gln Thr Cys Asn Arg Cys Gly 405 410 415 His Val Asp Ala Lys Ser Arg Arg Thr Arg Asp Leu Phe Thr Cys Thr 420 425 430 Arg Cys Gly His Ala Thr His Ala Asp Ile Gly Ala Ala Ile Asn Ile 435 440 445 Lys Ala Arg Ala Gln His Ser His Pro Ala Gln Glu Arg Pro Glu Gly 450 455 460 Thr Thr Asp Val Pro 465 <210> 33 <211> 583 <212> PRT <213> Artificial Sequence <400> 33 Met Phe Glu Arg Thr Thr Thr Lys Ala Ala Phe Leu Ser Pro Leu Asp 1 5 10 15 Leu Arg Pro Thr Gln Ala Thr Asp Leu Glu Arg Phe Ala Gly Thr Thr 20 25 30 Arg Trp Ala Phe Asn Trp Ala Asn Ala Leu Leu Glu Ala His His Gln 35 40 45 Ala Tyr Glu Gly Arg Arg Gln Gln Ala Ala Arg His Leu Phe Gly Leu 50 55 60 Gly Pro Glu Gln Leu Asp Glu Leu Arg Val Leu Ala Asn Gly Thr Arg 65 70 75 80 Asp Glu Asn Gly Lys Lys Ala Lys Gly Asp Pro Val Lys Arg Arg Glu 85 90 95 Tyr Glu Ser Ile Gln Lys Ala Thr Lys Lys Ala Val Ser Glu Glu Asn 100 105 110 Lys Ala Leu Gly Ala Glu Met Lys Leu Trp Asp Glu His Arg Ser Leu 115 120 125 Val Val His Lys Gly Arg Pro Leu Leu Thr Pro Gly Asp Glu Pro Ala 130 135 140 Leu Asp Ala Pro Pro Leu Ala His Arg Leu Tyr Ala Arg Arg Val Glu 145 150 155 160 Leu Ala Gly Ile Gln Lys Thr Asp Pro Asp Tyr Tyr Ala Glu Gln Arg 165 170 175 Lys Lys Glu Arg Glu Ala Ile Thr Pro Asn Val Val Ala Met Lys Arg 180 185 190 Asp Leu Met Ala Lys Gly Ala Tyr Phe Pro Ser Glu Tyr Asp Leu Gln 195 200 205 Tyr Ile Trp Arg Thr Val Arg Asp Leu Pro Lys Glu Glu Gly Gly Ser 210 215 220 Pro Trp Trp Pro Glu Cys Pro Thr Ile Leu Phe Tyr Asp Gly Ile Asn 225 230 235 240 Arg Ala Arg Thr Ala Trp Lys Asn Trp Met Asp Ser Ala Ser Gly Ala 245 250 255 Arg Lys Gly Pro Pro Val Gly Met Pro Arg Phe Lys Ser Lys Tyr Lys 260 265 270 Ala Lys Asp Thr Phe Thr Ile Thr Asn Pro Asn Arg Ser Val Ile Lys 275 280 285 Phe Glu Thr Tyr Arg Arg Ile Ala Ile Thr Gly Ile Gly Ser Met Arg 290 295 300 Leu His Arg Gly Ala Lys Leu Leu Ala Arg Arg Ile Ala Ala Gly Gln 305 310 315 320 Ala Glu Ile Thr Ser Ala Thr Ile Ser Arg Ser Gly Thr Ala Trp Tyr 325 330 335 Val Ser Val Leu Cys Thr Val His Thr Thr Ala Arg Thr Ala Pro Ser 340 345 350 Lys Ala Gln Arg Ser Arg Gly Ala Val Gly Val Asp Trp Gly Val Arg 355 360 365 Ala Leu Ala Thr Thr Ser Lys Pro Ile Ala Leu Thr Pro Gly Lys Pro 370 375 380 Ala Ser Arg Thr Val Pro Ala Glu Lys Tyr Gly Ala Ala Met Ser Gln 385 390 395 400 Lys Ile Ala Arg Ala Gln Arg Gln Leu Ala Arg Met Pro Lys Gly Ser 405 410 415 Ser Arg Arg Arg Lys Ala Ala Arg His Val Ala Asp Leu Gln His Leu 420 425 430 Val Ala Gln Arg Arg Ala Ser Ser Val His Gln Leu Ser Lys Ala Leu 435 440 445 Ala Gln Ser Phe Glu Ile Val Ala Ile Glu Gly Leu Asn Val Arg Gly 450 455 460 Met Thr Lys Ser Ala Lys Gly Thr Val Glu Asn Pro Gly Lys Asn Ile 465 470 475 480 Arg Gln Lys Ala Gly Leu Asn Arg Ala Ile Leu Asp Ala Thr Pro Gly 485 490 495 Glu Leu Lys Arg Gln Leu Glu Tyr Lys Thr Lys Lys Tyr Gly Ser Arg 500 505 510 Leu Val Glu Leu Asp Thr Trp Tyr Pro Ser Ser Lys Thr Cys Ser Arg 515 520 525 Cys Gly Trp Val His Pro Lys Leu Lys Leu Ser Met Arg Thr Phe Arg 530 535 540 Cys Gln Gln Cys Gly Leu Val Glu Asp Arg Asp Phe Asn Ala Ala Val 545 550 555 560 Asn Ile Glu Arg Gln Gly Ile Thr His Ile Val Lys Glu Asn Glu Gly 565 570 575 Thr Asp Asp Arg Glu Glu Gly 580 <210> 34 <211> 466 <212> PRT <213> Artificial Sequence <400> 34 Met Asp Glu Ala Leu Ser Pro Arg Gln Val Thr Arg Val Val Arg Leu 1 5 10 15 Pro Leu Asp Pro Thr Pro Val Glu Gln Ala Ile Leu Gln Arg Tyr Ala 20 25 30 Asp Cys Ser Arg Ala Cys Tyr Asn Phe Ala Phe Ser Leu Lys Asp Ala 35 40 45 Ala Gln Arg Arg Trp Ala Ala Glu Arg Asp Arg Leu Ile Ala Asp Gly 50 55 60 Met Gln Glu Lys Ala Ala Arg Lys Ala Ala Thr Ala Arg Phe Ala Val 65 70 75 80 Pro Arg Gln Phe Asp Leu Gln Lys Ile Phe Leu Ala Val Arg Asp Arg 85 90 95 Pro Phe Thr Gly Pro Glu Arg Ala Gly Glu Pro Phe Pro Arg Tyr Arg 100 105 110 Tyr Arg Trp Trp Ala Gly Val Asn Ala Met Val Cys Gln Gln Ala Phe 115 120 125 Arg Asp Ala Asp Arg Ala Trp Ser Asn Trp Leu Ser Arg Ser Arg Glu 130 135 140 Gly His Gly Tyr Pro Arg Pro Lys Gln Arg Gly Arg Cys Lys Asp Ser 145 150 155 160 Phe Arg Leu Pro Gly Val Ser Leu Ala Ala Glu Asp Leu Arg His Ile 165 170 175 Arg Ile Ser Gly Glu Arg Arg Pro Gly Gly Gln Lys Ala Phe Arg Val 180 185 190 Arg Leu His Arg Pro Ala Tyr Arg Leu Ala Arg Leu Leu Gln Gln Gly 195 200 205 Gly Gln Val Lys Met Val Thr Val Ser Arg Ser Gly Pro Arg Trp Phe 210 215 220 Ala Ala Phe Asn Val Arg Leu Pro Ala Pro Pro Ala Val Val Ala Thr 225 230 235 240 Arg Ala Gln Arg Arg Arg Gly Ala Val Gly Val Asp Leu Gly Val Glu 245 250 255 Val Ile Ala Ala Thr Ser Gln Pro Val Leu Leu Gln Gly Ala Val Thr 260 265 270 Gln Leu Val Ala Asn Pro Arg Tyr Leu Ala Asn Ala Arg Arg Ala Leu 275 280 285 Ala Lys Trp Glu Arg Arg Lys Ala Arg Arg Trp Val Lys Gly Leu Pro 290 295 300 Ala Gly Gln Gln Ser Arg Gly Trp His Glu Ala Lys Asp His Val Ala 305 310 315 320 Lys Leu His Ala Leu Ile Ala Ala His Arg Ala Ser Thr Gln His His 325 330 335 Leu Thr Lys Ala Leu Val Thr Gln Phe Ala Gln Val Val Ile Glu Asp 340 345 350 Leu Gln Val Lys Asn Met Ser Lys Ser Ala Lys Gly Thr Leu Glu Asp 355 360 365 Pro Gly Ser Arg Val Arg Gln Lys Ala Gly Leu Asn Arg Ala Leu Leu 370 375 380 Asp Val Gly Phe Gly Glu Ile Arg Arg Gln Leu Ala Tyr Lys Ala Pro 385 390 395 400 Trp Tyr Gly Ser Val Leu Ser Ala Val Asn Pro Ala Tyr Thr Ser Gln 405 410 415 Thr Cys His Arg Cys Gly His Val Asp Ser Lys Ser Arg Arg Thr Arg 420 425 430 Ser Val Phe Glu Cys Thr Ala Cys Gly Thr Gln Val His Ala Asp Ile 435 440 445 Gly Ala Ala His Asn Ile Leu Ala Arg Gly Gln Gln Thr Pro Ala Ala 450 455 460 Leu Phe 465 <210> 35 <211> 463 <212> PRT <213> Artificial Sequence <400> 35 Met Thr Glu Asn Thr Val Thr Val Val Val Arg Glu Ala Leu Asp Pro 1 5 10 15 Thr Pro Glu Gln Arg Leu Ile Leu His Arg Tyr Ala Asp Ala Ser Arg 20 25 30 Cys Ser Phe Asn Phe Ala Phe Gly Ile Lys His Ala Ala Gln Gln Arg 35 40 45 Trp Ser His Gly Arg Asp Leu Leu Leu Ala Gln Gly Leu Thr Ser Ala 50 55 60 Glu Ala Ala Lys Arg Ala Pro Arg Val His Met Pro Ser Gln Pro Asp 65 70 75 80 Ile Gln Arg Val Phe Leu Ala Val Arg Glu Arg Pro Leu Pro Gly Pro 85 90 95 Cys Arg Val Gly Glu Glu Pro His Ser Leu Tyr Pro Trp Trp Lys Gly 100 105 110 Val Asn Ala Ile Val Cys Gln Gln Ala Phe Arg Asp Ala Asp Lys Ala 115 120 125 Phe Ala Asn Trp Lys Ser Ala Ser Arg Arg Gly Ala Gly Pro Val Gly 130 135 140 Tyr Pro Arg Pro Lys Arg Arg Gly Arg Cys Arg Asp Ser Phe Arg Met 145 150 155 160 Phe Ser Val Arg Leu Val Ser Glu Asp Leu Arg His Val Arg Ile Gly 165 170 175 Gly Glu Arg Asp Pro His Gly Gln Arg Ser Phe Val Val Arg Leu His 180 185 190 Arg Pro Ala Arg Arg Leu Ala Arg Leu Leu Ala Arg Gly Gly Val Ala 195 200 205 Lys Ser Val Thr Val Ala Arg Glu Gly His Arg Trp Phe Ala Ala Phe 210 215 220 Asn Val Arg Leu Pro Ala Pro Pro Ala Val Arg Ala Thr Arg Arg Gln 225 230 235 240 Arg Glu Ala Gly Thr Val Gly Val Asp Leu Gly Val Ala Val Phe Ala 245 250 255 Ala Thr Ser Asp Pro Leu Val Ile Gly Asp Asp Lys Val Gln Leu Phe 260 265 270 Ala Asn Pro Arg His Leu Asp Asn Ala Arg Arg Gln Leu Arg Lys Trp 275 280 285 Gln Arg Arg Met Ser Arg Arg His Val Arg Gly Val Pro Ala His Arg 290 295 300 Gln Ser Lys Gly Trp Lys Glu Ala Arg Asp Gln Val Ala Arg Leu Gln 305 310 315 320 Gly Leu Val Ala Ala Arg Arg Ser Ser Thr Gln His Leu Leu Thr Lys 325 330 335 Arg Leu Val Thr Gln Tyr Ala Gln Val Val Leu Glu Asp Leu Arg Val 340 345 350 Arg Asn Met Thr Ser Ser Ala Arg Gly Ser Val Glu Ala Pro Gly Arg 355 360 365 Asn Val Val Ala Lys Ala Gly Leu Asn Arg Ala Ile Leu Asp Val Gly 370 375 380 Phe Gly Glu Ile Arg Arg Gln Ile Glu Tyr Lys Ala Pro Trp His Gly 385 390 395 400 Thr Thr Val Ser Val Val Asn Pro Ala Tyr Thr Ser Gln Thr Cys Asn 405 410 415 Arg Cys Gly His Thr Asp Ala Lys Ser Arg Arg Thr Arg Ser Leu Phe 420 425 430 Val Cys Thr Arg Cys Arg His Thr Thr His Ala Asp Thr Gly Ala Ala 435 440 445 Val Asn Ile Lys Arg Arg Ala Gln Ser Ala Ala Pro Lys Pro Gly 450 455 460 <210> 36 <211> 511 <212> PRT <213> Artificial Sequence <400> 36 Met Gly Arg Ala Glu Val Leu Arg Ala Phe Arg Phe Ala Leu Asp Pro 1 5 10 15 Thr Asp Thr Gln Leu Ala Ser Leu Asn Gln His Ala Gly Ala Ala Arg 20 25 30 Trp Ala Tyr Asn His Ala Ile Gly Val Lys Ser Gly Ala His Ala Arg 35 40 45 Arg Arg Ala Glu Val Glu Ala Leu Val Ala Ala Gly Met Thr Glu Lys 50 55 60 Asp Ala Arg Lys Ala Val Asn Val Pro Ile Pro Met Lys Pro Ala Ile 65 70 75 80 Gln Thr Ser Leu Asn Ala Ile Lys Gly Asp Pro Arg Val Gln Glu Val 85 90 95 Leu Pro Gly Ala Met Gly Pro His Arg Pro Cys Pro Trp Trp Arg Asp 100 105 110 Val Ala Thr His Ala Phe Gln Ser Ala Phe Arg Asp Ala Asp Thr Ala 115 120 125 Phe Lys Asn Trp Met Asp Ser Leu Arg Gly Gln Arg Ala Gly Arg Pro 130 135 140 Val Gly Tyr Pro Arg Phe Lys Arg Lys Gly Arg Ala Arg Asp Ser Phe 145 150 155 160 Arg Leu His His Asp Val Lys Ala Pro Thr Ile Arg Leu Asp Gly Tyr 165 170 175 Arg Arg Leu Leu Leu Pro Arg Ile Gly Ser Val Arg Leu His Glu Ser 180 185 190 Gly Lys Arg Leu Ala Arg Leu Val Gly Arg Gly Glu Ala Val Val Gln 195 200 205 Ser Val Thr Val Ala Arg Gly Gly Ser Arg Trp Tyr Ala Ser Val Leu 210 215 220 Cys Lys Val Trp Gln Glu Leu Pro Glu Gln Pro Thr His Ala Gln Arg 225 230 235 240 Ala Arg Gly Thr Ile Gly Val Asp Leu Gly Val Lys Val Leu Ala Ala 245 250 255 Val Ser Lys Pro Val Ser Met Thr Arg Gly Gly Pro Gln Leu Asp Leu 260 265 270 Val Pro Asn Ala Arg His Gly Lys Ala Ala Arg Arg Arg Leu Thr Arg 275 280 285 Ala Gln Gln Thr Tyr Ala Arg Thr Ala Lys Gly Ser Lys Arg Arg Ala 290 295 300 Lys Ala Ala Arg Arg Ile Gly Arg Ile Gln His Leu Thr Ala Glu Arg 305 310 315 320 Arg Ala Thr Ser Leu His Thr Leu Thr Lys Arg Leu Val Thr Ala Phe 325 330 335 Ala Thr Val Ala Val Glu Asp Leu Asn Val Ala Gly Met Thr Ala Ser 340 345 350 Ala Arg Gly Thr Val Asp Asn Pro Gly Arg Lys Val Arg Gln Lys Ala 355 360 365 Gly Leu Asn Arg Ser Val Leu Asp Ala Ser Phe Gly Glu Leu Arg Arg 370 375 380 Gln Leu Thr Tyr Lys Ala His Trp Tyr Gly Ala Ala Ile Ala Val Cys 385 390 395 400 Gly Arg Trp Val Pro Ser Ser Gln Thr Cys Ser Val Cys Gly Trp Gln 405 410 415 Asn Pro Arg Arg Leu Thr Leu Ala Asp Arg Glu Phe Glu Cys Thr Pro 420 425 430 Cys Gly Leu Thr Met Asp Arg Asp Leu Asn Ala Ala Arg Asn Ile Glu 435 440 445 Arg Leu Ala Gln His Val Ala Ser Gly Arg Glu Glu Thr Glu Asn Ala 450 455 460 Arg Gly Gly Gly Val Ser Pro Pro Val Arg Leu Gly Gly Arg Gln Ser 465 470 475 480 Pro Val Lys Arg Glu Asp Pro Ala His Arg Ala Gly Pro Pro Arg Arg 485 490 495 Ser Asp Pro Pro Ala Thr Leu Thr Ala Arg Arg Lys Arg Ala Ala 500 505 510 <210> 37 <211> 519 <212> PRT <213> Artificial Sequence <400> 37 Met Ala Glu Thr Ala Val Leu Arg Ala Phe Arg Phe Ala Leu Asp Ala 1 5 10 15 Thr Ala Ala Gln Glu Glu Gly Phe Leu Arg His Ala Gly Ala Ser Arg 20 25 30 Trp Ala Phe Asn His Ala Leu Gly Met Lys Val Ala Ala His Arg Gln 35 40 45 Trp Gln Arg Glu Val Lys Ala Leu Val Glu Gly Gly Met Leu Glu Ala 50 55 60 Gln Ala Arg Lys Thr Val Lys Val Pro Val Pro Thr Arg Pro Thr Ile 65 70 75 80 Gln Lys His Leu Asn Arg Ile Lys Gly Asp Ser Arg Ser Pro Asp His 85 90 95 Pro Glu Gly Ala Gln Gly Pro Gln Arg Pro Cys Pro Trp Phe His Glu 100 105 110 Val Ser Thr Tyr Ala Phe Gln Cys Ala Phe Glu Asp Ala Asp Arg Ala 115 120 125 Trp Asp Asn Trp Gln Ala Ser Leu Ser Gly Arg Arg Ala Gly Arg Lys 130 135 140 Val Gly Tyr Pro Arg Phe Lys Lys Lys Gly Arg Thr Lys Asp Ser Phe 145 150 155 160 Arg Ile Cys His Asp Ala Lys Lys Pro Thr Ile Arg Pro Asp Gly Tyr 165 170 175 Arg Arg Leu Arg Ile Pro Val Leu Gly Ser Val Arg Leu His Asp Thr 180 185 190 Ala Lys Pro Leu Ala Arg Leu Val Asp Arg Gly Ala Ser Val Lys Ser 195 200 205 Val Thr Val Ser Arg Ser Gly Ala Arg Trp Tyr Ala Ser Val Leu Val 210 215 220 Ser Val Leu Gln Asp Leu Pro Glu Arg Pro Thr Arg Arg Gln Arg Gln 225 230 235 240 Ala Gly Thr Val Gly Val Asp Phe Gly Val Lys Thr Leu Ala Ala Leu 245 250 255 Ser Ala Pro Val Thr Leu Pro Asn Leu Gly Thr Leu Thr Met Val Pro 260 265 270 Asn Pro Arg His Leu Ala Ser Asp Thr Arg Arg Leu Thr Lys Ala Gln 275 280 285 Gln Thr Leu Ser Arg Thr Thr Lys Gly Ser Ala Arg Arg Arg Lys Ala 290 295 300 Ala Gln Arg Val Gly Ile Val His His Arg Ile Ala Ser Arg Arg Ala 305 310 315 320 Thr Tyr Leu His Thr Leu Thr Lys Gln Leu Ala Thr Gly Tyr Ala Val 325 330 335 Val Ala Val Glu His Leu Asn Val Ala Gly Met Thr Ser Ser Ala Arg 340 345 350 Gly Thr Val Glu Glu Pro Gly Ser Lys Val Arg Gln Lys Ser Gly Leu 355 360 365 Asn Arg Ser Ile Leu Asp Ala Ser Pro Ala Glu Met Arg Arg Gln Leu 370 375 380 Asp Tyr Lys Thr Arg Trp Asn Gly Ser Gln Leu Ala Val Cys Asp Arg 385 390 395 400 Trp Phe Pro Ser Ser Arg Thr Cys Ser Ala Cys Gly Trp Gln Lys Pro 405 410 415 Arg Leu Thr Leu Ala Glu Arg Val Phe Asn Cys Gly Gln Cys Gly Leu 420 425 430 Val Ile Asp Arg Asp Leu Asn Ala Ala Arg Asn Ile Ala Ala His Ala 435 440 445 Val Leu Val Pro His Gly Thr Ala Ala Pro Gly Ser Gly Glu Ala Ser 450 455 460 Asn Ala Arg Gly Ala Ala Thr Arg Pro Ala Thr Pro Arg Gly Gly Arg 465 470 475 480 Gln Ala Ala Leu Lys Arg Glu Asp Thr Gly Pro Pro Arg Pro Val Pro 485 490 495 Pro Gln Arg Ser Asp Pro Leu Thr Leu Phe Thr Leu Asp Val Pro Asp 500 505 510 Gln Gln Thr Ala Lys Arg Pro 515 <210> 38 <211> 515 <212> PRT <213> Artificial Sequence <400> 38 Met Glu Gln Glu Val Leu Lys Ala Phe Arg Phe Ala Leu Asp Pro Arg 1 5 10 15 Pro Ala Gln Val Glu Val Leu Leu Arg His Ala Gly Ala Ala Arg Trp 20 25 30 Ala Phe Asn Leu Ala Leu Gly Met Lys Val Ala Ala His Lys Glu Trp 35 40 45 Arg Ser Gln Val Gln Ala Leu Val Asp Gln Gly Val Pro Glu Ala Glu 50 55 60 Ala Arg Arg Arg Val Lys Val Pro Val Pro Ser Lys Leu Gln Ile Gln 65 70 75 80 Lys His Leu Asn Leu Ile Lys Gly Asp Ser Arg Glu Gly Gly Leu Pro 85 90 95 Glu Gly Val Leu Gly Pro Glu Arg Pro Cys Pro Trp Trp His Glu Val 100 105 110 Ser Thr Tyr Ala Phe Gln Ser Ala Phe Ile Asp Ala Asp Arg Ala Trp 115 120 125 Gln Asn Trp Leu Ala Ser Leu Ala Gly Lys Arg Ala Gly Arg Ala Val 130 135 140 Gly Tyr Pro Arg Phe Lys Lys Lys Gly Arg Cys Arg Asp Ala Phe Arg 145 150 155 160 Leu His His Asp Val Lys Arg Pro Gly Ile Arg Pro Val Gly Tyr Arg 165 170 175 Arg Leu Arg Leu Pro Thr Val Gly Glu Val Arg Leu His Gly Ser Ala 180 185 190 Lys Arg Leu Val Arg Leu Leu Gly Arg Gly Cys Ala Gln Val Gln Ser 195 200 205 Val Thr Val Ser Arg Gly Gly His Arg Trp Tyr Ala Ser Val Leu Cys 210 215 220 Lys Val Val Val Glu Leu Pro Gly Lys Pro Thr Lys Ala Gln Ala Arg 225 230 235 240 Arg Gly Thr Ile Gly Val Asp Leu Gly Val Lys Tyr Leu Ala Ala Leu 245 250 255 Ser Gln Pro Leu Asp Val Lys Val Pro Glu Ser Arg Phe Val Asp Asn 260 265 270 Pro Arg His Leu Val Arg Ala Glu Lys Gln Leu Ala Lys Ala Gln Arg 275 280 285 Ala Leu Ser Arg Thr Gln Lys Gly Ser Ala Arg Arg His Lys Ala Arg 290 295 300 Cys Arg Val Gly Arg Leu His His Glu Val Ala Val Arg Arg Ala Thr 305 310 315 320 Val Leu His Ala Met Thr Lys Gln Leu Ala Thr Arg Phe Ser Thr Val 325 330 335 Ala Val Glu Asp Leu Tyr Val Ala Gly Met Thr Arg Ser Ala Arg Gly 340 345 350 Thr Val Glu Ala Pro Gly Lys Arg Val Arg Gln Lys Ala Gly Leu Asn 355 360 365 Arg Ala Ile Leu Asp Ala Ala Pro Gly Glu Val Arg Arg Gln Leu Ala 370 375 380 Tyr Lys Thr Arg Trp Tyr Gly Ser Lys Leu Ala Val Leu Asp Arg Trp 385 390 395 400 Phe Pro Ser Ser Lys Thr Cys Phe Ala Cys Gly Trp Gln Asn Pro Arg 405 410 415 Leu Thr Leu Ala Asp Arg Thr Phe His Cys Gly Gly Cys Gly Leu Ser 420 425 430 Thr Asp Arg Asp Glu Asn Ala Ala Arg Asn Ile Ala His His Ala Val 435 440 445 Val Ala Asp Asp Ser Pro Val Ala Pro Gly Lys Glu Glu Thr Arg Asn 450 455 460 Ala Arg Gly Ala Pro Asn Leu Ala Gly Leu Lys Ala Arg Arg Gln Gly 465 470 475 480 Ala Ser Lys Arg Glu Gly Thr Gly Pro Pro Gly Thr Val Pro Pro Gln 485 490 495 Arg Ser Asn Pro Leu Thr Leu Pro Asp Pro His Gly Gln Gln Pro Ala 500 505 510 Lys His Pro 515 <210> 39 <211> 518 <212> PRT <213> Artificial Sequence <400> 39 Met Glu Gln Gln Val Leu Lys Ala Phe Lys Phe Ala Leu Asp Pro Thr 1 5 10 15 Ser Ser Gln Val Glu Glu Leu Thr Arg His Ala Gly Ala Ser Arg Trp 20 25 30 Ala Phe Asn His Ala Leu Gly Met Lys Ile Ala Ala His Arg Asp Trp 35 40 45 Arg Thr Gln Val Gln Glu Leu Ile Asp Ala Gly Val Gln Glu Lys Ala 50 55 60 Ala Arg Gln Gln Val Arg Val Pro Ala Pro Thr Lys Pro Thr Ile Gln 65 70 75 80 Lys His Leu Asn Ser Ile Lys Gly Asp Ser Arg Ser Asp Asp Leu Pro 85 90 95 Pro Gly Ala Met Gly Pro Gln Arg Pro Cys Pro Trp Trp His Glu Val 100 105 110 Asn Thr His Ala Phe Gln Ser Ala Phe Ile Asp Ala Asp Thr Ala Trp 115 120 125 Lys Asn Trp Leu Asp Ser Leu Arg Gly Thr Arg Ala Gly Arg Lys Val 130 135 140 Gly Tyr Pro Arg Phe Lys Lys Lys Gly Arg Ser Arg Asp Ser Phe Arg 145 150 155 160 Leu His His Asp Val Asn Lys Pro Gly Ile Arg Leu Ala Thr Tyr Arg 165 170 175 Arg Leu Arg Leu Pro Lys Ile Gly Glu Val Arg Leu His Asp Ser Gly 180 185 190 Lys Arg Leu Ala Arg Leu Ile Gly Arg Gly Asp Ala Val Ile Gln Ser 195 200 205 Val Thr Val Ser Arg Ala Gly His Arg Trp Tyr Ala Ser Val Leu Ala 210 215 220 Lys Val Thr Val Thr Leu Pro Thr Arg Pro Thr Ala Arg Gln Val Ser 225 230 235 240 Ala Gly Lys Val Gly Val Asp Leu Gly Val Lys Ser Leu Ala Val Leu 245 250 255 Ser Arg Pro Leu Val Pro Gly Asp Asp Thr Thr Ala Phe Val Pro Asn 260 265 270 Pro Arg His Leu Arg His Ala Glu Thr Arg Leu Thr Lys Ala Gln Arg 275 280 285 Ala Leu Ser Arg Thr Thr Lys Gly Ser Ala Arg Arg Glu Lys Ala Arg 290 295 300 Arg Arg Val Ala Arg Leu His His Glu Val Ser Val Arg Arg Ala Gly 305 310 315 320 Gln Leu His Ala Leu Thr Lys Leu Leu Thr Thr Thr Phe Ser Glu Val 325 330 335 Ala Val Glu Asp Leu Asn Val Ala Gly Met Thr Arg Ser Ala Arg Gly 340 345 350 Thr Thr Glu Arg Pro Gly Arg Arg Val Arg Gln Lys Ala Gly Leu Asn 355 360 365 Arg Ala Ile Leu Asp Thr Ser Pro Gly Glu Leu Arg Arg Gln Leu Thr 370 375 380 Tyr Lys Ala Ser Trp Tyr Gly Ser Lys Leu Leu Val Leu Asp Arg Trp 385 390 395 400 Tyr Pro Ser Ser Lys Thr Cys Ser Ala Cys Gly Trp Gln Asn Pro Arg 405 410 415 Leu Thr Leu Ala Asp Arg Thr Phe His Cys Gly Asn Cys His Leu Val 420 425 430 Ile Asp Arg Asp Leu Asn Ala Ala Arg Asn Ile Ala Gln Gln Thr His 435 440 445 Gly Ala Pro His His Pro Val Ala Pro Gly Lys Gly Glu Thr Gln Asn 450 455 460 Ala Arg Arg Ala Thr Gly Arg Pro Pro Ser Pro Arg Ala Gly Arg Gln 465 470 475 480 Glu Ala Gln Lys Arg Glu Asp Thr Gly Pro Pro Arg Ser Val Pro Pro 485 490 495 Gln Arg Ser Asp Pro Leu Thr Ser Pro Pro Pro Gly Ser Pro Gly Gln 500 505 510 Glu Ser Ala Lys Ser Leu 515 <210> 40 <211> 532 <212> PRT <213> Artificial Sequence <400> 40 Met Gly Gln Gln Val Leu Lys Ala Phe Lys Phe Ala Leu Asp Pro Thr 1 5 10 15 Pro Gly Gln Thr Glu Glu Leu Thr Arg His Val Gly Ala Ser Arg Trp 20 25 30 Ala Tyr Asn His Ala Leu Gly Met Lys Ile Ala Ala His Arg Glu Trp 35 40 45 Arg Thr Gln Val Gln Ser Leu Val Glu Glu Gly Leu Pro Glu Lys Asp 50 55 60 Ala Arg Gln Arg Val Arg Val Pro Val Pro Thr Lys Pro Thr Ile Gln 65 70 75 80 Lys His Leu Asn Ser Ile Lys Gly Asp Ser Arg Ala Asp Asp Leu Pro 85 90 95 Pro Arg Ala Leu Gly Pro His Arg Pro Cys Pro Trp Trp His Glu Val 100 105 110 Asn Thr Tyr Ala Phe Gln Ser Ala Phe Ile Asp Ala Asp Thr Ala Trp 115 120 125 Lys Asn Trp Leu Asp Ser Leu Arg Gly Ala Arg Ala Gly Arg Lys Val 130 135 140 Gly Tyr Pro Arg Phe Lys Lys Lys Gly Arg Ser Arg Asp Ser Phe Arg 145 150 155 160 Leu His His Asp Val Asn Lys Pro Gly Ile Arg Leu Ala Thr His Arg 165 170 175 Arg Leu Arg Leu Pro Lys Ile Gly Glu Val Arg Leu His Asp Ser Gly 180 185 190 Lys Arg Leu Ala Arg Leu Ile Gly Arg Gly Asp Ala Val Val Gln Ser 195 200 205 Val Thr Val Ser Arg Ala Gly His Arg Trp Tyr Ala Ser Val Leu Ala 210 215 220 Lys Val Thr Val Thr Leu Pro Glu Arg Pro Thr Ala Arg Gln Thr Ser 225 230 235 240 Ala Gly Lys Val Gly Val Asp Leu Gly Val Lys Asn Leu Ala Val Leu 245 250 255 Ser Arg Pro Leu Leu Pro Gly Asp Asp Thr Thr Ala Phe Val Pro Asn 260 265 270 Pro Arg His Leu Arg His Ala Glu Ala Arg Leu Thr Lys Ala Gln Arg 275 280 285 Ala Leu Ser Arg Thr Thr Lys Gly Ser Ala Arg Arg Glu Lys Ala Arg 290 295 300 Arg Arg Val Ala Arg Leu His His Glu Val Ser Val Arg Arg Asp Gly 305 310 315 320 Gln Leu His Ala Leu Thr Lys Leu Leu Thr Thr Ser Phe Ser Glu Val 325 330 335 Ala Val Glu Asp Leu Asn Val Ala Gly Met Thr Arg Ser Ala Arg Gly 340 345 350 Thr Ile Glu Arg Pro Gly Arg Arg Val Arg Gln Lys Ala Gly Leu Asn 355 360 365 Arg Ala Ile Leu Asp Ala Ser Pro Gly Glu Leu Arg Arg Gln Leu Thr 370 375 380 Tyr Lys Ala Ser Trp Tyr Gly Ser Lys Leu Leu Val Leu Asp Arg Trp 385 390 395 400 Tyr Pro Ser Ser Lys Thr Cys Ser Ala Cys Gly Trp Gln Asn Pro Arg 405 410 415 Leu Thr Leu Ala Asn Arg Thr Phe His Cys Ala Asn Cys His Leu Ala 420 425 430 Ile Asp Arg Asp Leu Asn Ala Ala Arg Asn Ile Ala Gln Gln Thr His 435 440 445 Gly Ala Pro His His Pro Val Pro Pro Gly Arg Gly Asp Ala Lys Arg 450 455 460 Pro Pro Ser His Arg Lys Thr Ala Gln Pro Thr Gly Trp Thr Ala Gly 465 470 475 480 Gly Ala Glu Ala Gly Arg His Arg Pro Thr Gln Val Gly Ala Thr Ser 485 490 495 Ala Glu Arg Ser Ala Asp Ile Pro Thr Thr Arg Ala Thr Trp Ser Arg 500 505 510 Lys Gly Lys Ala Gly Leu Thr Ser Gln Ala Arg Lys Cys Pro Ala Arg 515 520 525 Thr Cys Arg Gly 530 <210> 41 <211> 513 <212> PRT <213> Artificial Sequence <400> 41 Met Ser Ser Glu Arg Ala Pro Lys Leu Arg Asn Val Val Thr Gln Gln 1 5 10 15 Ala Tyr Lys Tyr Ala Leu Glu Pro Thr Pro Arg Gln Gln Cys Ala Phe 20 25 30 Ser Ser His Ala Gly Ala Ala Arg Phe Ala Tyr Asn Trp Gly Ile Ala 35 40 45 Arg Val Ala Asp Ser Leu Asp Ala Tyr Ala Glu Gln Lys Ala Ala Gly 50 55 60 Ile Asp Glu Pro Asp Val Lys Phe Pro Gly His Phe Asp Leu Cys Lys 65 70 75 80 Met Trp Thr Ala Trp Lys Asn Thr Ala Glu Trp Thr Asp Arg His Thr 85 90 95 Gly Gln Thr Thr Thr Gly Val Pro Trp Val Ala Ser Asn Phe Val Gly 100 105 110 Thr Tyr Gln Ala Ala Leu Arg Asp Ala Ala Gly Ala Trp Gln Arg Phe 115 120 125 Phe Arg Ala Arg Lys Thr Gly Ala Arg Ala Gly Arg Pro Arg Phe Lys 130 135 140 Lys Arg Gly Arg Ala Arg Asp Ser Phe Gln Leu His Gly Asp Gly Leu 145 150 155 160 Arg Ile Val Asp Ala Lys His Val Asn Leu Pro Lys Ile Gly Thr Val 165 170 175 Lys Thr Phe Glu Ala Thr Arg Lys Leu Ala Arg Arg Leu Ala Lys Gly 180 185 190 Ser Val Pro Cys Pro Thr Cys Arg Ala Thr Gly Lys Ile Thr Asp Ser 195 200 205 Ala Ser Gly Lys Val Lys Lys Cys Ser Asp Cys Lys Ala Ala Gly Ser 210 215 220 Arg Pro Ala Ala Arg Ile Val Arg Gly Thr Val Ala Arg Asp Ser Ala 225 230 235 240 Gly Arg Trp Tyr Leu Ala Leu Thr Val Glu Leu Val Arg Glu Val Arg 245 250 255 Thr Ala Pro Thr Pro Arg Gln Leu Ala Gly Gly Pro Val Gly Val Asp 260 265 270 Phe Gly Val Arg Gln Val Ala Thr Leu Ser Thr Gly Gln Leu Val Asp 275 280 285 Asn Pro Arg His Leu Glu Ser His Leu Arg Arg Val Lys Thr Ala Gln 290 295 300 Gln Ala Leu Ser Arg Cys Pro Pro Gly Ser Arg Arg Arg Ala Lys Ala 305 310 315 320 Gln Gln Arg Leu Gly Arg Leu His Ala Arg Val Arg His Leu Arg Glu 325 330 335 Asn Ser Leu Gln Gln Ala Thr Ser Ala Leu Ile His Gln His Ser Val 340 345 350 Ile Ala Val Glu Gly Trp Asp Val Gln Gln Thr Ala Gln His Ala Ser 355 360 365 Pro Lys Asn Leu Pro Lys Gln Ile Arg Arg Asn Arg Asn Arg Ala Leu 370 375 380 Leu Asp Thr Gly Ile Gly Ala Ala Arg Trp Gln Leu Gln Ser Lys Gly 385 390 395 400 Ala Trp Tyr Gly Thr Thr Val Val Val Thr Asp Arg His Ala Pro Thr 405 410 415 Gly Arg Gln Cys Ser Ala Cys Gly Thr Val Lys Ala Thr Pro Ile Pro 420 425 430 Pro Thr Gln Asp Glu Tyr Arg Cys Pro Ala Cys Gly Thr Ser Leu Asp 435 440 445 Arg Arg Thr Asn Thr Ala Arg Val Leu Ala Ala Val Ala Ala Gln His 450 455 460 His Asp Ala Pro Ser Gly Gly Glu Ser Lys Asn Ala Arg Gly Glu Asn 465 470 475 480 Thr Arg Pro Thr Ala Pro Arg Arg Asn Gly Gln Phe Ser Ala Lys Arg 485 490 495 Glu Pro Arg Ser Arg Pro Pro Gly Arg Gly Gln Thr Gly Thr Pro Gly 500 505 510 Thr <210> 42 <211> 522 <212> PRT <213> Artificial Sequence <400> 42 Met Ser Glu Gly Val Ser Gly Met Ala Glu Val Leu Arg Ala Phe Lys 1 5 10 15 Phe Thr Leu Asp Pro Thr Arg Ala Gln Val Gly Ala Leu Gln Gln His 20 25 30 Ala Gly Ala Ala Arg Trp Ala Phe Asn Trp Ala Leu Gly Glu Lys Val 35 40 45 Ala Ala His Arg Glu Trp Arg Arg Gln Val Gly Ala Leu Leu Ala Glu 50 55 60 Gly Val Ala Glu Glu Gln Ala Arg Lys Gln Val Arg Val Pro Val Pro 65 70 75 80 Thr Lys Pro Thr Ile Gln Lys Arg Leu Asn Ser Phe Lys Gly Asp Ser 85 90 95 Arg Val Gln Asp Leu Pro Asp Gly Val Leu Gly Pro Arg Arg Pro Cys 100 105 110 Pro Trp Trp Trp Glu Val Lys Thr Tyr Cys Phe Gln Ala Ala Met Ala 115 120 125 Asp Ala Asp Thr Ala Trp Lys Asn Trp Leu Ser Ser Leu Thr Gly Ala 130 135 140 Arg Ala Gly Gln Arg Val Gly Tyr Pro Arg Phe Lys Lys Lys Gly Arg 145 150 155 160 Ala Arg Asp Ser Phe Arg Leu His His Asp Val Lys Lys Pro Gly Ile 165 170 175 Arg Leu Ala Gly Tyr Arg Arg Leu Arg Leu Pro Thr Ile Gly Glu Val 180 185 190 Arg Leu His Asp Phe Gly Lys Arg Leu Ala Arg Leu Ile Asp Arg Gly 195 200 205 Arg Ala Val Val Gln Ser Val Thr Val Ala Arg Cys Gly His Arg Trp 210 215 220 Tyr Ala Ser Val Leu Cys Lys Val Asp Gln Ser Val Pro Gln Arg Ser 225 230 235 240 Thr Arg Ala Gln Arg Arg Arg Gly Arg Val Gly Val Asp Leu Gly Val 245 250 255 Lys His Leu Ala Ala Leu Ser Gln Pro Leu His Pro Tyr Asp Arg Ala 260 265 270 Ser Leu Tyr Val Glu Asn Pro Arg His Leu Arg Arg Ala Ala Gln Arg 275 280 285 Leu Ala Lys Ala Gln Arg Ala Leu Ala Arg Thr Gln Lys Gly Ser Lys 290 295 300 Arg Arg Ala Lys Ala Val Arg Arg Val Gly Arg Leu His His Glu Val 305 310 315 320 Ala Val Arg Arg Glu Ser Thr Leu His Gln Leu Thr Lys Arg Leu Ala 325 330 335 Thr Gly Phe Ala Glu Val Ala Val Glu Asp Leu His Val Ala Gly Met 340 345 350 Thr Arg Ser Ala Lys Gly Thr Ile Asp Ala Pro Ser Arg Asn Val Arg 355 360 365 Ala Lys Ala Gly Leu Asn Arg Ser Ile Leu Asp Thr Ala Pro Gly Glu 370 375 380 Leu Arg Arg Gln Leu Thr Tyr Lys Thr Cys Trp Tyr Gly Ser Arg Leu 385 390 395 400 Ala Val Leu Asp Arg Trp Trp Pro Ser Ser Lys Thr Cys Ser Ala Cys 405 410 415 Gly Arg Gln Asn Pro Arg Leu Thr Leu Ala Asp Arg Thr Phe His Cys 420 425 430 Thr Gly Cys Gly Leu Arg Ile Asp Arg Asp Leu Asn Ala Ser Arg Asn 435 440 445 Ile Ala Thr His Ala Ala Leu Ala Asp Thr Ala Pro Pro Val Ala Pro 450 455 460 Asp Arg Gly Glu Thr Gln Asn Ala Arg Arg Ala Gly Thr Arg Pro Thr 465 470 475 480 Gly Pro Arg Ala Gly Arg His Pro Ala Thr Lys Arg Glu Asp Thr Ala 485 490 495 Pro Ala Val Pro Pro Gln Arg Ser Asn Pro Leu Ala Leu Pro Pro Pro 500 505 510 Pro Thr Gly Tyr Glu Gln Val Thr Leu Phe 515 520 <210> 43 <211> 519 <212> PRT <213> Artificial Sequence <400> 43 Met Ala Glu Thr Glu Val Leu Arg Ala Phe Arg Phe Ala Leu Asp Ala 1 5 10 15 Thr Ala Ala Gln Glu Glu Gly Phe Leu Arg His Ala Gly Ala Ser Arg 20 25 30 Trp Ala Phe Asn His Ala Leu Gly Met Lys Val Ala Ala His Arg Gln 35 40 45 Trp Gln Arg Glu Val Lys Ala Leu Val Glu Gly Gly Met Pro Glu Ala 50 55 60 Gln Ala Arg Lys Thr Val Lys Val Pro Val Pro Thr Arg Pro Thr Ile 65 70 75 80 Arg Lys His Leu Asn Arg Ile Lys Gly Asp Ser Arg Ser Pro Asp Leu 85 90 95 Pro Glu Gly Ala Gln Gly Pro Gln Arg Pro Cys Pro Trp Phe His Glu 100 105 110 Val Ser Thr Tyr Ala Phe Gln Ser Ala Phe Glu Asp Ala Asp Arg Ala 115 120 125 Trp Gly Asn Trp Gln Ala Ser Leu Ser Gly Arg Arg Ala Gly Arg Lys 130 135 140 Val Gly Tyr Pro Arg Phe Lys Lys Lys Gly Arg Thr Lys Asp Ser Phe 145 150 155 160 Arg Ile Cys His Asp Ala Lys Lys Pro Thr Ile Arg Pro Asp Gly Tyr 165 170 175 Arg Arg Leu Arg Ile Pro Ala Leu Gly Ser Val Arg Leu His Asp Thr 180 185 190 Ala Lys Pro Leu Ala Arg Leu Met Asp Arg Gly Ala Ile Val Lys Ser 195 200 205 Val Met Val Ser Arg Ser Gly Ala Arg Trp Tyr Ala Ser Val Leu Val 210 215 220 Ser Val Leu Gln Asp Ile Pro Glu Arg Pro Thr Arg Arg Gln Arg Gln 225 230 235 240 Ala Gly Thr Val Gly Val Asp Phe Gly Val Lys Thr Leu Ala Ala Leu 245 250 255 Ser Ala Pro Val Thr Leu Pro Asn Leu Gly Thr Leu Thr Met Val Pro 260 265 270 Asn Pro Arg His Leu Ala Ser Asp Thr Arg Arg Leu Thr Lys Ala Gln 275 280 285 Arg Ala Leu Ser Arg Thr Thr Lys Gly Ser Ala Arg Arg Arg Lys Ala 290 295 300 Ala Arg Arg Val Gly Val Val His His Arg Ile Ala Glu Arg Arg Ala 305 310 315 320 Thr Tyr Leu His Thr Leu Thr Ser Gln Leu Ala Thr Gly Tyr Ala Val 325 330 335 Val Ala Val Glu His Leu Asn Val Ala Gly Met Thr Ser Ser Ala Arg 340 345 350 Gly Thr Val Glu Glu Pro Gly Ser Lys Val Arg Gln Lys Ser Gly Leu 355 360 365 Asn Arg Ser Ile Leu Asp Ala Ser Pro Ala Glu Met Arg Arg Gln Leu 370 375 380 Asp Tyr Lys Thr Arg Trp Asn Gly Ser Gln Leu Ala Val Cys Glu Arg 385 390 395 400 Trp Phe Pro Ser Ser Arg Thr Cys Ser Ala Cys Gly Trp Gln Lys Pro 405 410 415 Arg Leu Thr Leu Ala Glu Arg Val Phe Ala Cys Gly Gln Cys Gly Leu 420 425 430 Val Ile Asp Arg Asp Leu Asn Ala Ala Arg Asn Ile Ala Ala His Ala 435 440 445 Val Leu Val Pro His Gly Thr Ala Ala Pro Gly Ser Gly Glu Ala Ser 450 455 460 Asn Ala Arg Gly Ala Ala Thr Arg Pro Ala Thr Pro Arg Gly Gly Arg 465 470 475 480 Gln Ala Ala Leu Lys Arg Glu Asp Thr Gly Pro Pro Arg Pro Val Pro 485 490 495 Pro Gln Arg Ser Asp Pro Leu Thr Leu Phe Thr Leu Asp Val Pro Asp 500 505 510 Gln Gln Thr Ala Lys Arg Pro 515 <210> 44 <211> 518 <212> PRT <213> Artificial Sequence <400> 44 Met Gln Thr Glu Val Leu Arg Ala Phe Arg Phe Thr Leu Asp Pro Thr 1 5 10 15 Pro Ala Gln Gln Glu Asp Leu Leu Arg His Ala Gly Ala Ala Arg Trp 20 25 30 Ala Phe Asn His Ala Leu Gly Val Lys Ile Ala Ala His Gln Gln Trp 35 40 45 Arg Thr Lys Val Gln Ala Leu Val Asp Ser Gly Val Ala Glu Ser Ile 50 55 60 Ala Arg Lys Gln Val Arg Val Pro Val Pro Met Lys Pro Glu Ile Gln 65 70 75 80 Lys His Leu Asn Arg Ile Lys Gly Asp Ser Gln Lys Gln Pro Trp Pro 85 90 95 Ala Gly Ser Ile Gly Pro Ala Arg Pro Cys Pro Trp Trp Arg Glu Val 100 105 110 Ser Thr Tyr Ala Phe Gln Ser Ala Phe Ile Asp Ala Asp Gln Ala Trp 115 120 125 Lys Asn Trp Leu Asp Ser Leu Ala Gly Arg Arg Ala Gly Arg Lys Val 130 135 140 Gly Tyr Pro Arg Phe Lys Lys Lys Ala Arg Ser Lys Asp Ser Phe Arg 145 150 155 160 Leu His His Thr Val Thr Gln Pro Thr Ile Arg Leu Asp Gly Tyr Arg 165 170 175 Arg Leu Thr Leu Pro Arg Leu Gly Thr Ile Arg Val His Asp Ser Gly 180 185 190 Lys Arg Leu Ala Arg Leu Ile Thr Arg Gly His Ala Val Val Gln Ser 195 200 205 Val Thr Val Ser Arg Ser Ala Asn Arg Trp Tyr Ala Ala Val Leu Val 210 215 220 Lys Val Arg Gln Asn Leu Pro Asp Arg Pro Thr Arg Arg Gln Gln Ser 225 230 235 240 Asn Gly Thr Val Gly Val Asp Leu Gly Val Lys Thr Leu Ala Ala Leu 245 250 255 Thr Gln Pro Val Thr Leu Pro Gly Ala Thr Asp Thr Leu Leu Val Pro 260 265 270 Asn Pro Arg His Leu Ala Ala Asp Thr Arg Arg Leu Thr Arg Ala Gln 275 280 285 Arg Ala Leu Ser Arg Thr Thr Lys Gly Ser Thr Arg Arg Arg Lys Val 290 295 300 Ala Arg Arg Val Ala Gln Leu His His Arg Ile Ala Glu Arg Arg Ala 305 310 315 320 Thr Tyr Leu His Thr Leu Thr Lys Gln Leu Thr Thr Arg Tyr Ala Ala 325 330 335 Val Ala Val Glu Asp Leu Asn Val Ser Gly Met Ser Ala Ser Ala Arg 340 345 350 Gly Thr Val Glu Glu Pro Gly Arg Lys Val Arg Gln Lys Ala Gly Leu 355 360 365 Asn Arg Ala Ile Leu Asp Met Ala Pro Ala Glu Ile Arg Arg Gln Leu 370 375 380 Gly Tyr Lys Thr Arg Trp Asn Gly Ser Arg Leu Ala Val Cys Asp Arg 385 390 395 400 Tyr Tyr Pro Ser Ser Lys Thr Cys Ser Ala Cys Gly Trp Gln Asn Pro 405 410 415 Arg Leu Thr Leu Ala Asp Arg Thr Phe Ile Cys Gly Gln Cys Gly Leu 420 425 430 Thr Ile Asp Arg Asp Leu Asn Ala Ala Arg Asn Ile Ala Ala His Ala 435 440 445 Val Pro Thr Pro Leu His Ser Thr Val Ala Pro Gly Thr Glu Glu Thr 450 455 460 Gln Asn Ala Arg Arg Ala Ala Thr Arg Pro Thr Thr Pro Arg Gly Gly 465 470 475 480 Arg Gln Thr Ala Lys Lys Arg Glu Asp Thr Gly Arg Gln Pro Pro Val 485 490 495 Pro Pro Gln Arg Ser Asp Pro Leu Thr Leu Gln Pro Pro Gln His Gln 500 505 510 Glu Val Ala Lys Gly Pro 515 <210> 45 <211> 506 <212> PRT <213> Artificial Sequence <400> 45 Met Ala Ser Thr Lys Thr Val Val Leu Arg Ala Phe Lys Phe Thr Leu 1 5 10 15 Ala Pro Thr Ala Thr Gln Asp Gln Gln Leu Leu Arg Trp Cys Gly Asn 20 25 30 Ala Arg Leu Ala Phe Asn Tyr Ala Leu Ala Ser Lys Arg Ala Ala His 35 40 45 Thr Glu Trp Arg Ala Gln Val Asp Ala Leu Val Thr Ser Gly Val Lys 50 55 60 Glu Pro Val Ala Arg Lys Arg Val Thr Gly Pro Lys Thr Pro Thr Lys 65 70 75 80 Pro Ala Val Tyr Lys Ala Phe Ile Ala Glu Arg Gly Asp Thr Arg Glu 85 90 95 Gly Leu Asp Gly Val Cys Pro Trp Ala His Glu Ile Asn Thr His Val 100 105 110 Phe Gln Ser Ala Phe Ile Asp Ala Asp Arg Ala Trp Lys Asn Trp Leu 115 120 125 Asp Ser Phe Lys Gly Thr Arg Lys Gly Arg Arg Val Gly Tyr Pro Arg 130 135 140 Phe Lys Lys Arg Gly Arg Ala Arg Asp Ala Phe Arg Leu His His Thr 145 150 155 160 Val Thr Lys Pro Thr Ile Arg Phe Ser Thr His Arg Arg Leu Arg Leu 165 170 175 Pro Thr Phe Gly Glu Val Arg Leu His Asp Ser Ala Arg Thr Leu Val 180 185 190 Arg Gln Ile Asp Arg Gly Thr Ala Val Val Gln Ser Val Thr Val Ser 195 200 205 Arg Ala Gly His Arg Trp Tyr Ala Ser Val Leu Cys Lys Val Glu Met 210 215 220 Asp Leu Pro Ser Gly Pro Thr Arg Ser Gln Gln Ala Ala Gly Thr Val 225 230 235 240 Gly Ile Asp Phe Gly Val Lys Ala Leu Ala Ala Leu Ser Lys Pro Leu 245 250 255 Val Pro Asp Arg Pro Glu Ser Thr Leu Leu Pro Asn Pro Arg His Leu 260 265 270 Ala Lys Ala Ala His Arg Leu Lys Arg Ala Gln Gln Thr Leu Ser Arg 275 280 285 Arg Gln Lys Gly Ser Ala Arg Arg Glu Lys Ala Arg Arg Arg Val Ala 290 295 300 Arg Leu His His Glu Val Ala Val Arg Arg Gln Ser Ala Leu His Gln 305 310 315 320 Ile Thr Lys Arg Leu Thr Thr Arg Phe Ala Thr Ile Ala Val Glu Asp 325 330 335 Leu His Val Ser Gly Met Thr Arg Ser Ala Arg Gly Thr Met Asp Lys 340 345 350 Pro Gly Arg Lys Val Arg Gln Lys Ala Gly Leu Asn Arg Ala Ile Leu 355 360 365 Asp Ala Ser Met Ala Glu Ala Arg Arg Gln Ile Thr Tyr Lys Thr Ser 370 375 380 Trp Tyr Gly Ser Arg Leu Ala Val Leu Asp Arg Trp Trp Pro Ser Ser 385 390 395 400 Lys Thr Cys Ser Ala Cys Gly Trp Gln Asn Pro Ser Leu Thr Leu Ala 405 410 415 Asp Arg Val Phe Glu Cys Ala Gln Cys Gly Leu Thr Leu Asp Arg Asp 420 425 430 Leu Asn Ala Ala Arg Asn Ile Glu Gln His Ala Val Gln Val Ala Ser 435 440 445 Gly Thr Gly Glu Thr Gln Asn Ala Arg Gly Glu Pro Val Arg Leu Pro 450 455 460 Arg Pro Arg Ala Glu Lys Gln Gly Ser Thr Lys Arg Glu Asp Thr Gly 465 470 475 480 Pro Pro Gly Pro Val Pro Pro Arg Arg Ser Asp Pro Pro Thr Pro Pro 485 490 495 Asn Pro Arg Gln Gly Gln Ala Lys Leu Phe 500 505 <210> 46 <211> 477 <212> PRT <213> Artificial Sequence <400> 46 Met Ser Val Ala His Val Ser Leu Leu Asn Met Thr Thr Thr Glu Asn 1 5 10 15 Thr Thr Thr Ile Val Val Arg Glu Pro Leu Asp Pro Thr Pro Glu Gln 20 25 30 Val Leu Val Leu Lys Arg Tyr Ala Asn Ala Ser Arg Ala Ser Phe Asn 35 40 45 Phe Ala Tyr Gly Leu Lys His Glu Ala Gln Gln Arg Trp Val Arg Gly 50 55 60 Arg Asp Arg Leu Leu Ser Lys Gly Leu Gly Arg Glu Glu Ala Asn Arg 65 70 75 80 Arg Ala Pro Lys Val Leu Val Pro Arg Gly Ser Asp Val Gln Arg Ile 85 90 95 Phe Leu Ala Val Arg Glu Gln Pro Leu Ala Gly Pro Leu Arg Glu Gly 100 105 110 Glu Ser Glu His Arg Arg Met Phe Arg Trp Trp Ala Gly Val Asn Ala 115 120 125 Ile Val Cys Gln Gln Ala Phe Arg Asp Ala Asp Thr Ala Phe Ala Asn 130 135 140 Trp Arg Ser Ser Gly Arg Arg Ala Gly Glu Gly Val Gly Tyr Pro Arg 145 150 155 160 Pro Lys Arg Val Gly Arg Cys Arg Asp Ser Phe Arg Met Thr Ser Val 165 170 175 Arg Leu Val Gly Thr Asp Leu Arg His Val Arg Ile Gly Gly Glu Lys 180 185 190 Asp Pro Ala Gly Gln Arg Ala Leu Ile Val Arg Leu His Arg Pro Gly 195 200 205 Arg Arg Leu Ala Arg Ala Met Ala Arg Gly Gly Val Val Lys Met Val 210 215 220 Thr Val Ala Arg Glu Gly Ser Gln Trp Trp Ala Ser Phe Asn Val Arg 225 230 235 240 Ile Val Leu Pro Pro Pro Ala Gln Pro Ser Arg Arg Gln Arg Glu Ala 245 250 255 Gly Thr Val Gly Val Asp Leu Gly Val Ala Val Phe Ala Ala Thr Ser 260 265 270 Glu Pro Val Leu Thr Ala Ala Gly Lys Glu Gln Leu Phe Asp Asn Pro 275 280 285 Arg His Leu Asp Asn Ala Arg Arg Gln Leu His Lys Trp Gln Arg Arg 290 295 300 Met Ala Arg Arg His Val Lys Gly Leu Pro Val His Arg Gln Ser Ala 305 310 315 320 Gly Trp Arg Glu Ala Arg Asp Gln Val Ala His Leu Met Gly Leu Val 325 330 335 Ala Gln Arg Arg Ala Ser Thr Gln His Leu Leu Thr Lys Gln Leu Val 340 345 350 Thr Gln Phe Glu His Val Ala Leu Glu Asp Leu Arg Val Lys Asn Met 355 360 365 Thr Arg Thr Ala Arg Gly Thr Val Glu Ala Pro Gly Arg Asn Val Ala 370 375 380 Ala Lys Ala Gly Leu Asn Arg Ala Ile Leu Asp Val Gly Phe Gly Glu 385 390 395 400 Ile Arg Arg Gln Ile Glu Tyr Lys Ala Lys Trp His Gly Val Thr Val 405 410 415 Thr Ala Val Asn Pro Ala Tyr Thr Ser Gln Thr Cys His Arg Cys Gly 420 425 430 His Val Asp Arg Lys Ser Arg Arg Thr Arg Ser Val Phe Glu Cys Thr 435 440 445 Arg Cys Gly His Val Thr His Ala Asp Ile Gly Ala Ala His Asn Ile 450 455 460 Lys His Arg Ala Leu Ala Pro Asp Ala Asn Asp Ala Arg 465 470 475 <210> 47 <211> 509 <212> PRT <213> Artificial Sequence <400> 47 Met His Thr Glu Val Leu Arg Ala Tyr Arg Phe Ala Leu Asp Pro Thr 1 5 10 15 Arg Gly Gln Leu Glu Asp Leu Ala Arg His Ala Gly Ala Ala Arg Trp 20 25 30 Ala Phe Asn His Ala Leu Ala Ala Lys Val Ala Ala His Lys Glu Trp 35 40 45 Arg Ala Lys Val Ala Ala Leu Val Glu Ala Gly Thr Ala Glu Ala Glu 50 55 60 Ala Arg Lys Gln Val Arg Val Pro Ile Pro Thr Lys Pro Gly Ile Gln 65 70 75 80 Lys Ala Leu Asn Ala Ala Lys Gly Asp Ser Arg Thr Gly Thr Asp Gly 85 90 95 Leu Cys Pro Trp Trp His Glu Val Asn Thr Tyr Cys Phe Gln Ser Ala 100 105 110 Phe Ala Asp Ala Asp Arg Ala Trp Lys Asn Trp Leu Asp Ser Leu Lys 115 120 125 Gly Val Arg Ala Gly Arg Lys Val Gly Tyr Pro Arg Phe Lys Glu Lys 130 135 140 Gly Arg Ala Arg Asp Ser Phe Arg Leu His His Asp Val Lys Lys Pro 145 150 155 160 Gly Ile Arg Met Val Gly Tyr Arg Arg Leu Arg Leu Pro Lys Val Gly 165 170 175 Glu Val Arg Leu His Gly Ser Gly Arg Glu Leu Ala Arg Ala Val Asp 180 185 190 Arg Gly Arg Ala Val Val Gln Ser Val Thr Val Ala Arg Asp Gly His 195 200 205 Arg Trp Tyr Ala Ser Val Leu Cys Ala Val Ala Met Asp Val Arg Glu 210 215 220 Lys Pro Thr Gln Arg Gln Thr Met Arg Gly Thr Val Gly Val Asp Leu 225 230 235 240 Gly Val Lys Tyr Leu Ala Ala Leu Ser Gln Pro Leu Thr Glu Gly Asp 245 250 255 Glu Ala Ser Lys Phe Val Ala Asn Pro Arg His Leu Lys Ala Ala Glu 260 265 270 Lys Arg Leu Val Lys Ala Gln Arg Ala Phe Ser Arg Thr Gln Lys Gly 275 280 285 Ser Gly Arg Arg Asp Lys Ala Arg Arg Arg Val Ala Arg Leu His His 290 295 300 Gln Val Ala Arg Gln Arg Ala Gly Ala Met His Gln Phe Thr Lys Arg 305 310 315 320 Leu Ala Thr Gly Phe Ala Val Val Ala Val Glu Asp Leu Asn Val Ala 325 330 335 Gly Met Thr Arg Ser Ser Arg Gly Thr Val Glu Ala Pro Gly Arg Arg 340 345 350 Val Arg Gln Lys Ala Gly Leu Asn Arg Ala Ile Leu Asp Val Ala Pro 355 360 365 Gly Glu Leu Arg Arg Gln Leu Ala Tyr Lys Thr Ser Trp Tyr Gly Ser 370 375 380 Lys Leu Ala Val Val Asp Arg Trp Phe Pro Ser Ser Lys Thr Cys Ser 385 390 395 400 Asn Cys Gly Trp Gln Asn Pro Ser Leu Thr Leu Ser Asp Arg Thr Phe 405 410 415 His Cys Ser Asn Cys Glu Ser Ala Met Asp Arg Asp Trp Asn Ala Ala 420 425 430 Arg Asn Ile Ala Arg His Ala Val Leu Gly Asp Ser Gln Val Ala Cys 435 440 445 Asp Arg Arg Glu Thr Glu Asn Ala Arg Gly Ala Leu Val Ser Pro Gly 450 455 460 Ala Ser Arg Gly Gly Arg Arg Arg Ala Val Lys Arg Glu Asp Thr Gly 465 470 475 480 Pro Pro Pro Val Pro Pro Gln Arg Ser Asp Pro Leu Ala Leu Pro Thr 485 490 495 Pro Thr Arg Lys Val Thr Gly Ile Gln Ala Ser Leu Phe 500 505 <210> 48 <211> 468 <212> PRT <213> Artificial Sequence <400> 48 Met Thr Thr Thr Ser Thr Thr Glu Thr Leu Arg Ala Tyr Arg Tyr Ala 1 5 10 15 Leu Asp Pro Thr Pro Ala Gln Ile Glu Ile Leu Gln Arg Tyr Ala Thr 20 25 30 Ala Ala Arg Cys Gly Tyr Asn Phe Ala Leu Gly Tyr Met Val Ala Val 35 40 45 His Gln Lys Trp Ala Arg Gly Arg Asp Ala Leu Ile Ala Ala Gly Met 50 55 60 Asp Lys Ala Val Ala Asn Lys Ala Ala Pro Lys Val Lys Val Pro Asn 65 70 75 80 Ala Phe Arg Ala Gln Ala Phe Phe Arg Glu Thr Lys Gly His Pro Phe 85 90 95 Thr Gly Pro Leu Pro Glu Gly Ala Glu Arg Thr Thr Pro Tyr Pro Trp 100 105 110 Trp Glu Gly Val Ser Asn Arg Ala Tyr Tyr Thr Ala Met Glu Asp Ala 115 120 125 Ala Thr Ala Trp Lys Asn Trp Met Asp Ser Ala Ser Gly Arg Arg Ala 130 135 140 Gly Gly Pro Val Gly Tyr Pro Arg Phe Lys Arg Arg Gly Arg Ala Arg 145 150 155 160 Glu Arg Phe Arg Leu Val His Asn Val Lys Lys Pro Glu Ile Arg Phe 165 170 175 Glu Thr Ser Arg Arg Leu Arg Ile Pro Gly Gly Gly Gly Gln Ser Ala 180 185 190 Phe Thr Val Arg Leu His Gln Asp Ala Arg Asp Leu Leu Arg Leu Ile 195 200 205 Ala Asp Gly Arg Ala Val Val Asn Ser Leu Thr Val Ser Arg Asp Gly 210 215 220 His Arg Trp His Ala Ser Val Leu Cys Arg Val Glu Gln His Ile Pro 225 230 235 240 Ser Gly Pro Ser His Arg Gln Ala Ala Ala Gly Arg Ile Gly Ala Asp 245 250 255 Leu Gly Val Lys Ala Leu Ala Ala Leu Ser Asp Pro Leu Thr Leu Thr 260 265 270 Pro Ala Ala Gly Pro Thr Val Leu Ile Pro Asn Pro Arg His Leu Ala 275 280 285 Ala Thr Glu Arg Lys Leu Ala Arg Ala Gln Arg Val Met Ser Arg Arg 290 295 300 Phe Val Arg Gly Ala Ala Gln Gln Ser Lys Gly Tyr Ser Glu Ala Arg 305 310 315 320 Ala Arg Val Ala Lys Leu His Ala Gln Leu Ala Ala Arg Arg Thr Ser 325 330 335 Ala Leu His Leu Ile Ser Lys Arg Leu Val Gln Gln Tyr Ala Glu Ile 340 345 350 Ala Leu Glu Ser Leu Asn Thr Lys Gly Met Thr Ser Ser Ala Lys Gly 355 360 365 Thr Leu Gln Ser Pro Gly Arg Asn Val Arg Gln Lys Ala Gly Leu Asn 370 375 380 Arg Ala Ile Leu Asp Ala Ser Phe Gly Glu Leu Asn Arg Gln Ile Ala 385 390 395 400 Tyr Lys Ala Ala Trp His Gly Ala Thr Leu Ala Arg Val Pro Thr Phe 405 410 415 Phe Pro Ser Ser Lys Thr Cys Ser Ala Cys Gly Trp Ile Asn Thr Glu 420 425 430 Leu Thr Leu Ala Asp Arg Glu Phe Ala Cys Arg Ser Cys Gly Ile Val 435 440 445 Leu Asp Arg Asp Ala Asn Ala Ala Arg Asn Ile Lys Asn His Ala Ile 450 455 460 Pro Val Gln Pro 465 <210> 49 <211> 462 <212> PRT <213> Artificial Sequence <400> 49 Met Thr Asp Thr Ile Ile Lys Ala Phe Arg Tyr Ala Leu Asp Pro Thr 1 5 10 15 Glu Ala Gln Ile Glu Ile Leu His Arg Tyr Ala Thr Ala Ser Arg Cys 20 25 30 Gly Phe Asn Phe Ala Leu Gly Met Lys Thr Thr Val Tyr Asp Arg Trp 35 40 45 Arg Arg Gly Arg Asp Ala Leu Val Ala Ala Gly Met Asp Lys Ala Glu 50 55 60 Ala Asn Lys Lys Ala Pro Lys Val Arg Val Pro Asn Arg Asn Arg Thr 65 70 75 80 Gln Ala Tyr Trp Arg Glu Thr Arg Gly Gln Gly Phe Ile Gly Pro Leu 85 90 95 Arg Glu Gly Gln Glu Pro Arg Ala Pro Phe Ala Trp Trp Glu Gly Val 100 105 110 Asn Asn Arg Ala Tyr Tyr Thr Ala Phe Glu Asp Ala Asp Thr Ala Trp 115 120 125 Lys Asn Trp Leu Asp Ser Leu Ala Gly Arg Arg Pro Ala Met Gly Phe 130 135 140 Pro Lys Phe Lys Arg Arg Gly Ala Ser Arg Glu Ser Phe Arg Ile Val 145 150 155 160 His Ser Leu Lys Asn Pro Asp Ile Arg Phe Asp Gly Pro Arg Arg Leu 165 170 175 Arg Ile Pro Gly Gly Gly Gly Gln Pro Ala Phe Thr Val Arg Leu Leu 180 185 190 Gln Ser Pro Arg Ala Leu Thr Ser Leu Ile Ala Ser Gly Gln Ala Val 195 200 205 Ile Thr Ser Val Thr Val Ser Arg Glu Gly His Arg Trp His Ala Ser 210 215 220 Val Leu Ala Arg Ile Glu Gln Asp Leu Pro Thr Arg Pro Thr Arg Arg 225 230 235 240 Gln Gln Ala Ala Gly Arg Ile Gly Ile Asp Leu Gly Val Lys Thr Ala 245 250 255 Leu Thr Leu Ser Asp Pro Leu Thr Leu His Arg Gly Gln Ala Pro Val 260 265 270 Leu Ala Ile Asp Asn Pro Arg Leu Leu Glu Asn Thr Ala Arg Lys Leu 275 280 285 Ala Arg Ala Gln Arg Val Met Ala Arg Arg His Val Lys Gly Ala Ala 290 295 300 Gln Gln Ser Gln Gly Tyr Leu Glu Ala Lys Ala Arg Val Ala Lys Leu 305 310 315 320 His Ala Leu Leu Ala Ala Arg Arg Ser Thr Ala Gln His Leu Val Thr 325 330 335 Lys Arg Leu Val Glu Gln Tyr Ala Glu Ile Ala Leu Glu Thr Leu Asn 340 345 350 Thr Lys Gly Met Thr Arg Ser Ala Lys Gly Thr Val Asp Lys Pro Gly 355 360 365 Arg Asn Val Arg Gln Lys Ser Gly Leu Asn Arg Ala Leu Leu Asp Val 370 375 380 Gly Phe Ala Glu Ile Asn Arg Gln Ile Glu Tyr Lys Ala Gly Trp Arg 385 390 395 400 Ala Val Thr Ile Ala Arg Val Pro Gly Leu Phe Pro Ser Ser Lys Ser 405 410 415 Cys His Arg Cys Gly Trp Ile His Thr Asn Leu Thr Leu Ala Asp Arg 420 425 430 Glu Phe Arg Cys Asp Ala Cys Gly Met Ile Ile Asp Arg Asp Val Asn 435 440 445 Ala Ala Gln Asn Ile Lys Asn His Ala Thr Pro Asp Arg Pro 450 455 460 <210> 50 <211> 530 <212> PRT <213> Artificial Sequence <400> 50 Met Ala Leu Gln Arg Ile Gln Gln Ala Phe Lys Tyr Ala Leu Asp Pro 1 5 10 15 Thr Pro Ala Gln Ala Arg Met Leu Thr Ser His Ala Gly Ala Ala Arg 20 25 30 Tyr Ala Phe Asn Trp Gly Leu Ala Thr Phe Ala Thr Ala Leu Asp Ala 35 40 45 Tyr Ser Ala Glu Lys Thr Ala Gly Val Lys Lys Pro Val Thr Lys Leu 50 55 60 Pro Gly His Phe Asp Leu Cys Lys Leu Trp Thr Ala His Lys Asp Asn 65 70 75 80 Pro Thr Ser Asp Leu Gly Trp Val Gly Gln Asn Phe Ser Gly Thr Tyr 85 90 95 Gln Ala Ala Leu Arg Asp Ala His Ala Ala Trp Lys Ala Phe Leu Asp 100 105 110 Ser Lys Asn Gly Arg Arg Arg Gly Arg Gln Val Gly Arg Pro Arg Phe 115 120 125 Lys Ser Arg His Arg Thr Thr Ala Ala Phe Gln Thr His Gly Thr Gly 130 135 140 Leu Arg Val Ala Asp His Arg His Ile Asn Leu Pro Lys Ile Gly Pro 145 150 155 160 Val Arg Ser Tyr Glu Lys Thr Lys Lys Leu Arg Arg Leu Leu Thr Arg 165 170 175 Pro Asp Val Thr Cys Thr Thr Cys Asp Gly Thr Lys Glu Ala Pro Ala 180 185 190 Pro Ala Pro Ala Thr Lys Ser Cys Gly Asp Cys Lys Gly Thr Gly His 195 200 205 Ala Pro His Thr Arg Ile Val Arg Gly Asn Ile Thr Arg Thr Pro Ser 210 215 220 Gly Arg Trp His Ile Ser Leu Thr Val Glu Thr His Arg Asp Ile Arg 225 230 235 240 Thr Arg Pro Ser Gln Arg Gln Arg Glu Gly Gly Ile Val Gly Val Asp 245 250 255 Trp Gly Val Arg Asp Leu Ala Thr Leu Ser Thr Gly Glu Val Ile Gly 260 265 270 Asn Pro Arg His Leu Glu Arg Asn Leu Ala Arg Leu Arg Arg Ala Gln 275 280 285 Gln Asp Leu Ala Arg Lys Ala Asp Gly Ser Ala Gly Arg Glu Ala Ala 290 295 300 Arg Leu Arg Val Ala Arg Leu His Gly Arg Val Ala Asn Leu Arg Arg 305 310 315 320 Asp His Leu Glu Lys Val Thr Ser Arg Leu Val His Ser His Thr Arg 325 330 335 Ile Val Val Glu Gly Trp Asp Val Gln His Ala Met Gln His Ser Gly 340 345 350 Asp Gly Ala Pro Lys Trp Val Arg Arg Asp Arg Asn Arg Ala Leu Ser 355 360 365 Asp Thr Gly Ile Gly Ala Ala Arg Trp Met Leu Gly Arg Lys Ala Ala 370 375 380 Trp Tyr Gly Ser Ala Ile Val Glu Thr Gly Pro His Glu Pro Thr Gly 385 390 395 400 Arg Thr Cys Ser Ala Cys Ala Thr Ala Arg Thr Lys Pro Val Ala Pro 405 410 415 Ala Asp Glu Arg Phe Thr Cys Pro Ser Cys Gly Tyr Ser Gly Asp Arg 420 425 430 Arg Val Asn Thr Ala Arg Val Leu Val Arg Leu Ala His Thr Ser Asp 435 440 445 Ala Pro Ser Ser Gly Glu Ser Leu Asn Ala Arg Gly Gly Asp Val Arg 450 455 460 Pro Ala Ala Pro Ser Thr Arg Pro Gly Gly Gln Thr Pro Thr Arg Pro 465 470 475 480 Ser Gly His Ile Ala Val His Ala Leu Gly Lys Asp Glu Glu Gly Arg 485 490 495 Asp Lys Gly Lys Asp Lys Glu Phe Arg Gly Arg Arg Ser Pro Met Lys 500 505 510 Arg Glu Ala Arg Ser Arg Pro Pro Gly Arg Gly Lys Ala Gly Thr Pro 515 520 525 Asp Pro 530 <210> 51 <211> 512 <212> PRT <213> Artificial Sequence <400> 51 Gly Thr Asp Arg Tyr Gly Lys Pro Gln Lys Pro Asp Trp Glu Ala Ala 1 5 10 15 Arg Gln Val Phe Ile Asp Gly Gly Ile Glu Asp Glu Val Ala Asp Thr 20 25 30 Ile Leu Ser Glu Trp Gln Gln Arg Tyr Arg Arg Glu Pro Ser Asp Thr 35 40 45 Lys Pro Gly Asn Asp Pro Leu Val Asp Gly Ile Ala Pro Trp Leu Ala 50 55 60 Glu Val Pro Asn Ala Leu Val Gln Arg Ala Glu Ser Asp Cys Glu Glu 65 70 75 80 Ala Trp Lys Arg Phe Tyr Lys Met Leu Arg Glu Gly Thr Ala Thr Pro 85 90 95 Ser Lys Arg Ser Arg Pro Arg Lys Thr Ala Arg Pro Asp Gly Ser Phe 100 105 110 Asp Tyr Asp Pro Pro Gly Leu Pro Arg Tyr Lys Arg Arg Ser Pro Gly 115 120 125 Arg Gly Ser Phe Tyr Leu Thr Asn Thr Glu Val Arg Leu Val Pro Asn 130 135 140 Ala Ser Arg Arg Ile Arg Leu Gly Gly Lys Ile Gly Asp Val Arg Thr 145 150 155 160 Leu Glu Gly Arg Thr Met Arg Arg Ile Arg Arg Ser Ile Ala Lys Arg 165 170 175 Asp Gly Val Ile Gln Ser Val Thr Val Ser Arg Gly Ala Ser Arg Trp 180 185 190 Tyr Ala Ser Val Leu Val Lys Glu Thr Phe Thr Pro Pro Lys Pro Thr 195 200 205 Arg Arg Gln Leu Glu Ala Gly Arg Val Gly Val Asp Val Gly Val Lys 210 215 220 His Ala Phe Val Val Ala Gly Pro Ser Thr Thr Ala Gly Gly Leu Ile 225 230 235 240 Val Asp Arg Pro Ala Arg Cys Gly Ser Glu Lys Ala Leu Asn Asp Ala 245 250 255 Arg Lys Thr Leu Arg Glu Thr Ala Pro Arg Ser Lys Ala Arg Met Lys 260 265 270 Ala Arg Ile Thr Val Ser Arg Leu Glu Asp Arg Glu Gln Trp Ala Ile 275 280 285 Glu Gln Ala Gly Arg Ala Val Ser Arg Lys Lys Leu Arg Ser Lys Asn 290 295 300 Trp Tyr Lys Ala Val Ala Arg Leu Ala Glu Leu Lys His Arg Gln Ala 305 310 315 320 Val Arg Arg Lys Thr Phe Ile His Glu Ala Thr Lys Arg Leu Ala Thr 325 330 335 Gly Tyr Ala Glu Ile Val Ile Glu Asp Leu Gln Val Lys Gly Met Ser 340 345 350 Ala Ser Ala Lys Gly Ser Ala Glu Ala Pro Gly Arg Arg Val Arg Gln 355 360 365 Lys Ala Gly Leu Asn Arg Arg Ile Leu Ala Ser Ser Phe Gly Glu Phe 370 375 380 Arg Arg Gln Leu Glu Tyr Lys Gln Gln Trp Tyr Gly Ser His Val Leu 385 390 395 400 Thr Ala Asp Arg Trp Tyr Ala Ser Ser Lys Ile Cys Ala His Cys Gly 405 410 415 Ala Thr Lys Ala Lys Leu Pro Leu Ser Ala Arg Gln Tyr Val Cys Asp 420 425 430 Ser Cys Gly Tyr Thr Ala Asp Arg Asp Val Asn Ala Ala Arg Asn Leu 435 440 445 Ala Arg Cys Ala Lys Val Ala Ser Thr Arg Lys Gly Val Ala Ser Glu 450 455 460 Val Gly Glu Thr Val Asn Asp Arg Arg Gly Arg Tyr Gly Arg Arg Arg 465 470 475 480 Ser Thr Val Ser Ala Asp Ala Ala Ala Val Pro Ala Gly Arg Pro Ala 485 490 495 Ile Asp Ala Gly His Arg Gly Gly Ser Asp Ser Ala Thr Tyr Pro Leu 500 505 510 <210> 52 <211> 758 <212> PRT <213> Artificial Sequence <400> 52 Met Ala Ala Met Glu Thr Gln Met Ile Gln Gln Ala Tyr Leu Phe Ala 1 5 10 15 Leu Asp Pro Thr Gln Ala Gln Ala Ala Thr Leu Ala Ser His Ala Gly 20 25 30 Ala Arg Arg Tyr Ala Phe Asn Trp Ala His Ala Met Ile Ala Ala Ala 35 40 45 Ala Asp Ala Arg Gln Ala Gln Lys Asp Ala Gly Leu Glu Pro Asp Ile 50 55 60 Ala Ile Pro Gly Gln Phe Glu Val Gly Pro Ala Trp Thr Arg Trp Arg 65 70 75 80 Asp Ala Ala Val Gly Cys Lys Thr Cys Arg Arg Leu Leu Ala Arg Asp 85 90 95 Pro Ala Val Pro Ala Ser Leu Trp Ala Asp Ser Arg Thr Gly Glu Ile 100 105 110 Thr Cys Asp Pro Gly Ala Leu Gln Arg Leu Gly Leu Asp Tyr Gly Ala 115 120 125 Pro Pro Leu His Asp Pro Ala Asp Val Gly Cys Arg Arg Cys Trp Lys 130 135 140 Val Leu Arg Glu Thr Pro Ala Gly Glu Trp Ala Asp Ser Ser Gly Ser 145 150 155 160 Ala Ala Cys Pro Glu Ala Arg Arg Pro Gly Pro Pro His Glu Pro Ser 165 170 175 Gly Thr Cys Gln Asp Gly Cys Asp Gly Ala Ala Arg Ala Cys Gly Gly 180 185 190 Pro His Glu Pro Ser Ser Asp Phe Leu Ala Trp Thr Gly Asp Val Phe 195 200 205 Ser Gly Thr Ile Gln Ala Ala Gln Arg Asp Ala Asp Val Ala Trp Lys 210 215 220 Lys Phe Leu Ser Gly Lys Ala Arg Arg Pro Arg Phe Lys Lys Arg Gly 225 230 235 240 Lys ...

Claims

1. Use of a C2C9 nuclease or a nucleic acid encoding the C2C9 nuclease in a gene editing system for purposes other than disease diagnosis or treatment, wherein the amino acid sequence of the C2C9 nuclease is shown in SEQ ID NO.

3.

2. A gene editing system, characterized by, The gene editing system includes a C2C9 nuclease and / or a nucleic acid encoding the C2C9 nuclease, as well as a guide RNA or a nucleic acid encoding the guide RNA; the amino acid sequence of the C2C9 nuclease is shown in SEQ ID NO.

3.

3. The gene editing system as described in claim 2, characterized in that, The gene editing system identifies a PAM sequence on a target sequence; and / or, the gene editing system targets a nucleic acid fragment of 12-40 bp in length after the PAM sequence.

4. The gene editing system as described in claim 3, characterized in that, The gene editing system targets a 20bp nucleic acid fragment following the PAM sequence.

5. The gene editing system as described in claim 3, characterized in that, The PAM sequence is AAN, GAN; where N is a degenerate base, representing any base of A, T, C or G.

6. The gene editing system according to any one of claims 2-5, characterized in that, The guide RNA comprises: Gene targeting regions (i) capable of hybridizing with target sequences. tracr pair sequence (ii). tracr RNA sequence (iii); Wherein, the tracr pair sequence (ii) hybridizes with the tracrRNA sequence (iii) to form a stem-loop structure; The guide RNA is a single strand, which is formed by sequentially linking the nucleic acid target segment (i) with the tracr pair sequence (ii) and the tracrRNA sequence (iii); or, the guide RNA comprises two strands, one of which is formed by linking the gene target segment (i) with the tracr pair sequence (ii), and the other strand is the tracrRNA sequence (iii).

7. The gene editing system as described in claim 6, characterized in that, The guide RNA sequence corresponds to the C2C9 nuclease, wherein the amino acid sequence of the C2C9 nuclease is shown in SEQ ID NO. 3, its corresponding tracr pair sequence is shown in SEQ ID NO. 118, the tracr RNA sequence is shown in SEQ ID NO. 119, and the guide RNA backbone sequence obtained by linking them together is shown in SEQ ID NO.

120.

8. A gene editing method for purposes other than disease diagnosis or treatment, characterized in that, The target gene is brought into contact with the gene editing system of any one of claims 2-7 to achieve the editing of the target gene.

9. The gene editing method as described in claim 8, characterized in that, Includes the following steps: i) Introducing the C2C9 nuclease or the nucleic acid encoding the C2C9 nuclease into cells; ii) Introduce the guide RNA or the nucleic acid encoding the guide RNA into the cell; III) The C2C9 nuclease mediates the creation of one or more nicks in the target gene, or targets, edits, modifies or manipulates the target gene.

10. The gene editing method as described in claim 8 or 9, characterized in that, The C2C9 nuclease is guided to the target gene via a guide RNA in processed or unprocessed form.

11. The gene editing method as described in claim 8 or 9, characterized in that, The C2C9 nuclease and guide RNA form a complex that recognizes the PAM sequence on the target gene; and / or, the target sequence of the gene editing system is a nucleic acid fragment of 12-40 bp in length following the PAM sequence.

12. The gene editing method as described in claim 11, characterized in that, The target sequence of the gene editing system is a 20bp nucleic acid fragment following the PAM sequence.

13. The gene editing method as described in claim 9, characterized in that, The method further includes the step of introducing a donor template containing a heterologous polynucleotide sequence into the cell.

14. The application of the gene editing system as described in any one of claims 2-7 or the method as described in any one of claims 8-13 in gene editing of target genes and / or their related peptides in ex vivo cells or cell-free environments for purposes other than disease diagnosis or treatment, characterized in that, The isolated cells include bacterial cells, archaea cells, fungal cells, protist cells, plant cells, and animal cells.

15. The application of the gene editing system as described in any one of claims 2-7 or the method as described in any one of claims 8-13 in gene editing of target genes and / or their related peptides in ex vivo cells or cell-free environments for purposes other than disease diagnosis or treatment, characterized in that, The gene editing methods described are selected from the group consisting of gene cutting, gene deletion, gene insertion, point mutation, transcriptional repression, transcriptional activation, base editing, and guided editing.