SARS-CoV-2 s protein variants and their use in the preparation of universal vaccines
By designing specific mutated variants of the SARS-CoV-2 S protein, recombinant protein vaccines were prepared, solving the problem that existing vaccines are unable to cope with mutant strains of the novel coronavirus, and achieving effective protection and immune response against multiple mutant strains.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-26
- Publication Date
- 2026-03-31
AI Technical Summary
Existing COVID-19 vaccine designs are mostly based on early viral sequences, which are difficult to effectively combat the constantly emerging mutant strains of the novel coronavirus, leading to a decline in vaccine efficacy. There is a need to develop universal vaccines that can target multiple mutant strains.
Design a SARS-CoV-2 S protein variant containing specific mutations in the S1 and S2 subunits and cleavage sites, express it using recombinant protein technology and fuse it with a tag, and use it to prepare a vaccine to induce a broad-spectrum immune response.
This SARS-CoV-2 S protein variant can effectively prevent infection by multiple mutant strains, induce a strong immune response and cross-neutralizing activity, and provide broad-spectrum protection.
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biotechnology, specifically relating to the SARS-CoV-2 S protein variant and its application in the preparation of universal vaccines. Background Technology
[0002] With the widespread transmission of SARS-CoV-2, many viral mutant strains have emerged. Among them, strain B.1.1.7 from the UK, B.1.351 from South Africa, P.1 from Brazil, and B.1.617.2 from India have been identified by the WHO as Variant of Concern (VOC). These mutant strains are relatively more transmissible and have some impact on the effectiveness of specific drugs and vaccines. Currently approved or undergoing clinical trials of COVID-19 vaccines are designed based on early sequences of the SARS-CoV-2 virus. As the COVID-19 pandemic continues to spread and expand, new mutant strains will continuously emerge and replace existing strains. Due to the long development cycle of vaccines, vaccines targeting a specific mutant strain will always lag behind the emergence of that strain. A universal vaccine that is broadly effective against multiple mutant strains may solve this problem.
[0003] SARS-CoV-2 uses its surface membrane protein Spike (also known as the S protein) to enter susceptible cells. The S protein consists of three domains: the S1 domain at the N-terminus, the S2 domain at the apical membrane, and the transmembrane domain. The susceptibility of SARS-CoV-2 to host cells is determined by the receptor-binding domain (RBD) on the S1 domain. Summary of the Invention
[0004] The purpose of this invention is to prevent disease caused by infection with SARS-CoV-2 WA1 / 2020 and / or mutant strains of SARS-CoV-2 WA1 / 2020; the mutant strains of SARS-CoV-2 WA1 / 2020 are SARS-CoV-2 alpha mutants, SARS-CoV-2 beta mutants, SARS-CoV-2 gamma mutants, or SARS-CoV-2 delta mutants.
[0005] The present invention first protects a SARS-CoV-2 S protein variant, relative to the SARS-CoV-2 WA1 / 2020S protein, which contains one or more mutations in its S1 subunit, one or more mutations in its S2 subunit, and one or more mutations at the cleavage sites of its S1 and S2 subunits.
[0006] In some implementations, the S1 subunit mutation of the SARS-CoV-2 S protein variant includes one or more of L18F, D80A, D215G, 242-244del, S305T, K417N, L452R, E484K, N501Y, and D614G.
[0007] In some implementations, mutations in the S2 subunit of the SARS-CoV-2 S protein variant include K986P, A701V, and / or V987P.
[0008] In some implementations, the mutation sites of the S1 and S2 subunit cleavage sites of the SARS-CoV-2 S protein variant include one or more of R682, R683, and R685.
[0009] In some embodiments, the mutation at the R682 site of the SARS-CoV-2 S protein variant is R682G. The mutation at the R683 site is R683S. The mutation at the R685 site is R685S.
[0010] In some embodiments, the amino acid sequence of the SARS-CoV-2 S protein variant is as shown in SEQ ID NO:2 from the N-terminus 1 to 1205, or is a sequence having at least 90%, 95%, 96%, 97%, 98%, 99%, or 99.9% identity with SEQ ID NO:2 from the N-terminus 1 to 1205.
[0011] The fusion protein obtained by attaching a tag to the N-terminus and / or C-terminus of any of the SARS-CoV-2 S protein variants described above is also within the scope of protection of this invention.
[0012] The label can stabilize the structure, aid in purification, or facilitate identification.
[0013] In some implementations, a tag is added for purification and identification (amino acid sequence is SAWSHPQFEKGGGSGGGGSGGSAWSHPQFEKGSDYKDDDDK, nucleotide sequence is TCTGCCTGGAGCCACCCACAGTTCGAGAAGGGCGGCGGCAGCGGCGGCGGCGGCTCCGGCGGCTCTGCATGGTCTCACCC CCAGTTTGAAAAGGGCAGCGACTACAAGGACGACGATGATAAATGA).
[0014] The amino acid sequence of the fusion protein may be as shown in SEQ ID NO:2.
[0015] The present invention also protects a DNA molecule that encodes any of the SARS-CoV-2 S protein variants described above or any of the fusion proteins described above.
[0016] The nucleotide sequence of the DNA molecule is such as SEQ ID NO:1 or SEQ ID NO:1 from position 41 to 3781 starting from the 5' end, or is a sequence having at least 90%, 95%, 96%, 97%, 98%, 99%, or 99.9% identity with SEQ ID NO:1 or SEQ ID NO:1 from position 41 to 3781 starting from the 5' end.
[0017] The present invention also protects a recombinant plasmid obtained by inserting any of the above-described DNA molecules into the multiple cloning site of an expression vector.
[0018] The present invention also protects a recombinant protein, which can be obtained by transfecting mammalian cells with any of the recombinant plasmids described above and culturing the transfected cells. The cells are preferably 293F cells.
[0019] The present invention also protects a composition for preventing SARS-CoV-2 infection, for neutralizing SARS-CoV-2, or for treating diseases caused by SARS-CoV-2 infection, said composition comprising any of the recombinant proteins described above;
[0020] The SARS-CoV-2 strain is preferably SARS-CoV-2 WA1 / 2020 or a mutant strain of SARS-CoV-2 WA1 / 2020;
[0021] The mutant strain of SARS-CoV-2 WA1 / 2020 is either the SARS-CoV-2 alpha mutant, the SARS-CoV-2 beta mutant, the SARS-CoV-2 gamma mutant, or the SARS-CoV-2 delta mutant.
[0022] This invention also protects the use of any of the above-described recombinant proteins or any of the above-described compositions in the preparation of medicaments for preventing SARS-CoV-2 infection, for neutralizing SARS-CoV-2, or for treating diseases caused by SARS-CoV-2 infection;
[0023] The SARS-CoV-2 strain is preferably SARS-CoV-2 WA1 / 2020 or a mutant strain of SARS-CoV-2 WA1 / 2020;
[0024] The mutant strain of SARS-CoV-2 WA1 / 2020 is either the SARS-CoV-2 alpha mutant, the SARS-CoV-2 beta mutant, the SARS-CoV-2 gamma mutant, or the SARS-CoV-2 delta mutant.
[0025] The present invention also protects a vaccine comprising any of the SARS-CoV-2 S protein variants described above, any of the fusion proteins described above, any of the DNA molecules described above, any of the recombinant plasmids described above, any of the recombinant proteins described above, or any of the combinations described above. Preferably, the vaccine is an mRNA vaccine, an adenovirus vector vaccine, or a recombinant protein vaccine.
[0026] The use of any of the above-described SARS-CoV-2 S protein variants, fusion proteins, DNA molecules, recombinant plasmids, recombinant proteins, or compositions in the preparation of vaccines. Preferably, the vaccine is an mRNA vaccine, an adenovirus vector vaccine, or a recombinant protein vaccine.
[0027] Any of the vaccines mentioned above may be SARS-CoV-2 vaccines.
[0028] The SARS-CoV-2 S protein variant can be used to prepare a universal SARS-CoV-2 vaccine, providing a possibility for the prevention and / or treatment of novel coronavirus pneumonia caused by SARS-CoV-2. This invention has significant application value. Detailed Implementation
[0029] The object of this invention is to provide a composition comprising a recombinant protein containing a sequence encoding a protein that is a coding sequence for the extracellular region of the SARS-CoV-2 S protein, including mutations and / or peptides. This invention also includes administering the composition comprising the recombinant protein of this invention to a desired subject to induce effector and memory T-cell and B-cell immune responses in the subject for the treatment and / or prevention of SARS-CoV-2, particularly SARS-CoV-2 infection.
[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. While any methods and materials similar to or equivalent to those described herein may be used in the practice of testing the invention, preferred materials and methods are described herein. The following terminology will be used in describing and claiming protection for this invention.
[0031] As used herein, the terms “comparison” or “reference” are used interchangeably and refer to the value used as a standard for comparison.
[0032] As used herein, the terms “peptide,” “polypeptide,” and “protein” are used interchangeably and refer to compounds composed of amino acid residues covalently linked by peptide bonds. Polypeptides include any peptide or protein comprising two or more amino acids linked together by peptide bonds. The polypeptide sequences of the present invention can be produced by any suitable means, including recombinant production, chemical synthesis, or other synthetic methods. Suitable production techniques are well known to those skilled in the art. Alternatively, peptides can also be synthesized using well-known solid-phase peptide synthesis methods.
[0033] As used herein, a “fusion protein” refers to a protein comprising two or more proteins linked together by peptide bonds or other chemical bonds. Proteins may be directly linked together by peptide bonds or other chemical bonds, or may have one or more amino acids between two or more proteins, referred to herein as a spacer region.
[0034] In this invention, the amino acid sites are numbered according to the amino acid numbering of the wild-type S protein of SARS-CoV-2 WA1 / 2020.
[0035] As used herein, a “mutation” is a change in the DNA sequence and the protein it encodes that results in an alteration from its native state. In this invention, the S protein expressed by the DNA molecule inserted into the multiple cloning site of the pCAGGS vector is a variant of the SARS-CoV-2 S protein, which, relative to its wild-type S protein (i.e., the S protein of SARS-CoV-2 WA1 / 2020), contains one or more mutations in its S1 subunit (e.g., the presence of L18F, D80A, D215G, 242-244del, S305T, K417N, L452R, E484K, etc.). One or more N501Y and D614G mutations), the S protein variant contains one or more mutations in the region approximately 980 to 990 of its S2 subunit (e.g., one or more of K986P, V987P, and A701V mutations), and one or more mutations at its S1 and S2 subunit cleavage sites (e.g., one or more mutations at R682, R683, and R685, where R682 can be mutated to G or other amino acids of similar properties, R683 can be mutated to S or other amino acids of similar properties, and R685 can be mutated to S or other amino acids of similar properties (for information on the chemical properties of amino acids, see, for example, Stryer et al., Biochemistry, 5th edition, 2002, pp. 44-49)), or variant sequences having at least 90%, 95%, 96%, 97%, 98%, 99%, or 99.9% identity with these variants. Specifically, the SARS-CoV-2 S protein variant can be a variant of SEQ ID NO:2, or a variant sequence having at least 90%, 95%, 96%, 97%, 98%, 99%, or 99.9% identity with SEQ ID NO:2.
[0036] As used herein, the term "nucleic acid" refers to polynucleotides, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The term should also be understood to include RNA or DNA analogs prepared from nucleotide analogs as equivalents, and, as applicable to the described embodiments, to include single-stranded polynucleotides (sense or antisense) and double-stranded polynucleotides.
[0037] As used herein, a "vector" is a composition comprising isolated nucleic acids and a substance that can be used to deliver the isolated nucleic acids into the cell. Many vectors are known in the art, including but not limited to linear polynucleotides, polynucleotides associated with ionic or amphiphilic compounds, plasmids, and viruses. In this document, the term "vector" includes autonomously replicating viruses.
[0038] The term "transfected cell" as used in this article (sometimes referred to in the industry as "production cell," "complementary cell," or "host cell") can be any cell capable of proliferating and expressing the desired recombinant protein.
[0039] Transient expression of a protein in a host cell or stable expression of a protein can be achieved by any means known in the art, including transfection.
[0040] The recombinant protein of the present invention can be formulated into a pharmaceutical composition. Such a pharmaceutical composition may be in a form suitable for administration to a subject, or the pharmaceutical composition may further comprise one or more pharmaceutically acceptable carriers, one or more additional ingredients, or a combination thereof.
[0041] Vaccine compositions containing recombinant proteins disclosed herein can be used to induce immunity against encoded antigen proteins. Vaccines can be formulated using standard techniques and, in addition to encoding the desired protein, may include pharmaceutically acceptable media, such as phosphate-buffered saline (PBS) or other buffer solutions, as well as other components, such as antibacterial and antifungal agents, isotonic and absorption-delaying agents, adjuvants, etc. In some embodiments, the vaccine composition is administered in combination with one or more other vaccines.
[0042] As used herein, the term "prevention" means the prior provision of medication, which may be prior to exposure to a pathogen (pre-exposure prophylaxis) or before the development of disease symptoms (post-exposure prophylaxis). The term "treatment" means the administration of medication during illness.
[0043] As used herein, "subject" or "patient" can refer to a human or a non-human mammal. Non-human mammals include, for example, livestock and pets, such as sheep, bovine, swine, canine, feline, and rodent mammals. Preferably, the subject is a human. The vaccine or pharmaceutical composition of the present invention can be administered via intramuscular injection, intravenous injection, intraperitoneal injection, subcutaneous injection, epidermal application, intradermal application, nasal application, rectal application, or oral administration.
[0044] As used herein, the term "effective amount" or "therapeutic effective amount" refers to the amount of the recombinant protein or composition of the present invention that is necessary for the prevention of a particular condition, or for reducing the severity of the condition or at least one of its symptoms or related symptoms and / or improving the condition or at least one of its symptoms or related symptoms.
[0045] The recombinant proteins, compositions, and methods of the present invention may have one or more of the following superior features compared to the prior art, including but not limited to higher productivity, increased transgenic expression, improved immunogenicity and stability, improved antibody neutralizing activity, or a unique serological cross-reactivity profile. In particular, improved immunogenicity and improved antibody neutralizing activity are significant advantages.
[0046] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.
[0047] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.
[0048] In the quantitative experiments in the following examples, three replicate experiments were set up, and the average value of the results was taken.
[0049] The pCAGGS vector is a circular plasmid, and its nucleotide sequence is shown in SEQ ID NO: 5.
[0050] Example 1: Preparation and Identification of Recombinant Protein
[0051] I. Construction of Recombinant Plasmids
[0052] 1. The DNA fragment between the restriction endonucleases EcoRI and XhoI in the pCAGGS vector was replaced with the DNA molecule shown in SEQ ID NO:1 using homologous recombination to obtain the recombinant plasmid pCAGGS-Uni-S-2PM.
[0053] The recombinant plasmid pCAGGS-Uni-S-2PM expresses the recombinant protein pCAGGS-Uni-S-2PM shown in SEQ ID NO:2, which is the SARS-CoV-2 S protein containing mutations of L18F, D80A, D215G, 242-244del, S305T, K417N, L452R, E484K, N501Y, D614G, A701V, K986P, V987P, R682G, R683S and R685S.
[0054] 2. The DNA fragment between the restriction endonucleases EcoRI and XhoI in the pCAGGS vector was replaced with the DNA molecule shown in SEQ ID NO:3 using homologous recombination to obtain the recombinant plasmid pCAGGS-WH-S-2PM.
[0055] The recombinant plasmid pCAGGS-WH-S-2PM expresses the control protein shown in SEQ ID NO:4.
[0056] II. Preparation of recombinant protein and control protein
[0057] Cell culture conditions: constant temperature culture at 125 rpm, 37°C, and 5% CO2.
[0058] 1. Use FreeStyle TM 293F cells were cultured in 293Expression Medium until the cell density reached 2-3 × 10⁻⁶ cells / year. 6 / ml. During passage, the cell ratio is 1:3-4, and the cell expansion rate is approximately 24 hours per generation.
[0059] 2. After completing step 1, take the recombinant plasmid pCAGGS-Uni-S-2PM, use the transfection reagent PEI to prepare the transfection system and transfect 293F cells, and culture for 72 hours.
[0060] 3. After completing step 2, centrifuge and collect the supernatant; use the Twin-Strep-tag carried by the protein and the binding filler StrepTactin-XT to purify the protein and collect the protein solution.
[0061] 4. After completing step 3, use a 30 kDa concentration tube to concentrate the volume to less than 1 ml.
[0062] 5. Further component separation, purification, and buffer replacement (PBS) were performed using molecular sieves to obtain the recombinant protein pCAGGS-Uni-S-2PM.
[0063] Following the steps described above, replace the recombinant plasmid pCAGGS-Uni-S-2PM with the recombinant plasmid pCAGGS-WH-S-2PM, keeping all other steps unchanged, to obtain the control protein.
[0064] III. Identification of Recombinant Proteins
[0065] HumanACE2 protein is a product of Acrobiosystems, catalog number AC2-H52H8; it is the extracellular domain (18-740) of the human ACE2 protein.
[0066] AM122 is a product of Acrobiosystems, catalog number S1N-M122; it is an RBD binding antibody.
[0067] AM121 is a product of Acrobiosystems, catalog number SPD-M121; it is an NTD-binding antibody.
[0068] The recombinant protein is pCAGGS-Uni-S-2PM.
[0069] 1. Identification of protein expression and purity
[0070] Proteins were collected by measuring the elution time and height of recombinant protein components in each tube of the molecular sieve. The OD value of the recombinant protein was detected by spectrophotometry for the peak with higher UV absorption, and the protein concentration was calculated by using the theoretical absorption coefficient. At the same time, representative component tubes were selected for denatured protein gel electrophoresis analysis. Based on the comparison with protein markers, the band located at approximately 120 kDa was identified as the target band.
[0071] The high-purity fractions were combined into a single tube, and the concentration of all recombinant proteins was adjusted to 1 mg / ml using PBS buffer. The isoelectric point and purity of each purified protein were then determined using capillary electrophoresis.
[0072] 2. Identification of activity
[0073] (1) The binding affinity of the recombinant protein to the receptor humanACE2 protein was determined using the SPR method.
[0074] (2) The ACE2-his protein was coupled to the CM5 chip, and the Kon, Koff and affinity of the protein were detected by Biacore machine using different concentrations of recombinant protein.
[0075] (3) The antigenicity of the recombinant protein was identified by ELISA. The recombinant protein was coated with ELISA plates at 100 ng / well, and the binding ability of the recombinant protein was detected by AM121 and AM122, respectively.
[0076] The results showed that recombinant plasmids pCAGGS-Uni-S-2PM and pCAGGS-WH-S-2PM could express the correct-sized Spike protein. Due to the modification of the Folin restriction site, no detached S1 protein was produced simultaneously. Furthermore, the protein yield was high, exhibiting high purity and a relatively symmetrical peak shape on the molecular sieve pattern. Both Spike proteins could bind to the receptor ACE2 protein.
[0077] The recombinant protein pCAGGS-Uni-S-2PM binds well to AM122 but poorly to AM121. The control proteins bind to both AM121 and AM122.
[0078] Example 2: Application of Recombinant Protein
[0079] I. Animal Immunization
[0080] Six-week-old female BALB / c mice were randomly divided into three groups: n=10 each in the first two groups and n=5 each in the last group. The recombinant protein pCAGGS-Uni-S-2PM prepared in Example 1 and the control protein were mixed with CpG adjuvant and administered intramuscularly in 100 μl volumes. A total of three immunizations were performed: 100 μg for the initial immunization; 100 μg for the second immunization three weeks after the initial immunization; and 50 μg for the third immunization two months after the initial immunization. Blood was collected starting two weeks after the initial immunization, and every two weeks (from the cheek) until four weeks after the third immunization.
[0081] Group 1 (G1): The immunizing agent was recombinant protein pCAGGS-Uni-S-2PM.
[0082] Group 2 (G2 group): The immunizing agent was the control protein.
[0083] Group 3 (G3 group): Adjuvant injection control group.
[0084] II. Preparation of SARS-CoV-2 pseudoviruses
[0085] 1. Replace the small DNA fragment between the restriction endonucleases BamHII and EcoRI in the pcDNA3.1(+) vector with the double-stranded DNA molecule shown in SEQ ID NO:6 (the gene encoding the SARS-CoV-2 WA1 / 2020 S protein) to obtain the SARS-CoV-2WA1 / 2020S protein particle.
[0086] The double-stranded DNA molecule shown in SEQ ID NO:6 encodes the SARS-CoV-2 WA1 / 2020S protein shown in SEQ ID NO:11.
[0087] 2. Replace the small DNA fragment between the restriction endonucleases BamHII and EcoRI in the pcDNA3.1(+) vector with the double-stranded DNA molecule shown in SEQ ID NO:7 (the gene encoding the full-length S protein of the SARS-CoV-2 alpha mutant strain) to obtain the SARS-CoV-2 alpha mutant strain S protein particle.
[0088] 3. Replace the small DNA fragment between the restriction endonucleases BamHII and EcoRI in the pcDNA3.1(+) vector with the double-stranded DNA molecule shown in SEQ ID NO:8 (the gene encoding the full-length S protein of the SARS-CoV-2 beta mutant strain) to obtain the SARS-CoV-2 beta mutant strain S protein particle.
[0089] 4. Replace the small DNA fragment between the restriction endonucleases BamHII and EcoRI in the pcDNA3.1(+) vector with the double-stranded DNA molecule shown in SEQ ID NO:9 (the gene encoding the full-length S protein of the SARS-CoV-2 gamma mutant strain) to obtain the SARS-CoV-2 gamma mutant strain S protein particle.
[0090] 5. Replace the small DNA fragment between the restriction endonucleases BamHII and EcoRI in the pcDNA3.1(+) vector with the double-stranded DNA molecule shown in SEQ ID NO:10 (the gene encoding the full-length S protein of the SARS-CoV-2 delta mutant strain) to obtain the SARS-CoV-2 delta mutant strain S protein particle.
[0091] 6. The SARS-CoV-2 WA1 / 2020S protein particle and backbone plasmid pNL4-3R-E-luciferase (described in the following literature: He J, Choe S, Walker R, Di Marzio P, Morgan DO, Landau NR. Human immunodeficiency virus type 1 viral protein R (Vpr) arrests cells in the G2 phase of the cell cycle by inhibiting p34cdc2 activity. J Virol; 69:6705–6711, 1995) were co-transfected into 293T cells. After incubation, an infectious but non-replicating SARS-CoV-2 WA1 / 2020 pseudovirus was obtained, with infectivity similar to that of the live virus.
[0092] The specific steps are as follows: SARS-CoV-2WA1 / 2020S protein particle and backbone plasmid pNL4-3R-E-luciferase were co-transfected into 293T cells, incubated at 37°C, and the cell culture supernatant was collected 48 hours after transfection, which is the viral fluid containing SARS-CoV-2 WA1 / 2020 pseudovirus.
[0093] 7. Following the method in step 6, replace the SARS-CoV-2 WA1 / 2020S protein particle with the SARS-CoV-2 alpha mutant strain S protein particle, keeping all other steps unchanged, to obtain the SARS-CoV-2 alpha mutant strain pseudovirus, whose infectivity is similar to that of the live virus.
[0094] 8. Following the method in step 6, replace the SARS-CoV-2 WA1 / 2020S protein particle with the SARS-CoV-2 beta mutant strain S protein particle, keeping all other steps unchanged, to obtain the SARS-CoV-2 beta mutant strain pseudovirus, whose infectivity is similar to that of the live virus.
[0095] 9. Following the method in step 6, replace the SARS-CoV-2 WA1 / 2020 S protein particle with the SARS-CoV-2 gamma mutant S protein particle, keeping all other steps unchanged, to obtain the SARS-CoV-2 gamma mutant pseudovirus, whose infectivity is similar to that of the live virus.
[0096] 10. Following the method in step 6, replace the SARS-CoV-2 WA1 / 2020 S protein particle with the SARS-CoV-2 delta mutant S protein particle, keeping all other steps unchanged, to obtain the SARS-CoV-2 delta mutant pseudovirus, whose infectivity is similar to that of the live virus.
[0097] 11. The viral titers of SARS-CoV-2 WA1 / 2020 pseudovirus, SARS-CoV-2 alpha mutant pseudovirus, SARS-CoV-2 beta mutant pseudovirus, SARS-CoV-2 gamma mutant pseudovirus and SARS-CoV-2 delta mutant pseudovirus were detected using the p24 quantitative detection ELISA kit (HIV p24 antigen quantitative detection kit, KEY-BIO, 96T).
[0098] A higher absorbance value indicates a higher virus content.
[0099] III. Detection of total antibodies induced by vaccines
[0100] Take the blood samples obtained in step one, separate the serum, and use ELISA to detect total bound IgG. For total IgG detection, use recombinant protein pCAGGS-Uni-S-2PM or control protein to coat the ELISA plate (100 ng / well). The serum is first diluted to 200-fold volume, and then serially diluted 3-fold (200, 600, 1800, 5400, 16200, 48600, 145800 and 437400, a total of 8 dilutions, diluted with PBS buffer at pH 7.2). The secondary antibody is Anti-mouse IgG HRP.
[0101] The results showed that serum immunized with recombinant protein pCAGGS-Uni-S-2PM could bind very well to recombinant protein pCAGGS-Uni-S-2PM or control protein.
[0102] IV. Detection of antibody neutralizing activity in animal serum after vaccination
[0103] The test solution is serum obtained from the blood sample obtained in step one.
[0104] 1. The test solution was diluted 48 times with DMEM medium containing 10% FBS, and then serially diluted 3 times to obtain different serum concentrations (6 dilutions in total: 48, 144, 432, 1296, 3888 and 11664).
[0105] 2. Mix 100 μl of the diluent obtained in step 1 with 50 μl of the viral fluid of SARS-CoV-2 WA1 / 2020 pseudovirus, SARS-CoV-2 alpha mutant pseudovirus, SARS-CoV-2 beta mutant pseudovirus, SARS-CoV-2 gamma mutant pseudovirus, or SARS-CoV-2 delta mutant pseudovirus (virus content of 100 TCID50) prepared in step 2, and incubate at 37°C for 1 h to form the experimental group.
[0106] Mix 100 μl of DMEM medium containing 10% FBS with 50 μl of the following pseudoviruses prepared in step two: SARS-CoV-2 WA1 / 2020 pseudovirus, SARS-CoV-2 alpha mutant pseudovirus, SARS-CoV-2 beta mutant pseudovirus, SARS-CoV-2 gamma mutant pseudovirus, or SARS-CoV-2 delta mutant pseudovirus (virus content 100 TCID50). Incubate at 37°C for 1 h as a blank control group.
[0107] 3. After completing step 2, add 50 μl of Huh7 cell culture medium (containing approximately 2 × 10⁻⁶ cells). 4 (1 Huh7 cells), incubated at 37°C for 48 hours (in practical applications, 48-72 hours is acceptable).
[0108] 4. After completing step 3, add 100 μl of PBS buffer and 50 μl of cell lysis buffer (Bright-Globe). TM The Luciferase Assay System (Promega, E2650) was used to allow the sample to stand for 2 minutes. Then, the luciferase activity was detected using a chemiluminescence analyzer, and the neutralization activity was further calculated.
[0109] Neutralization activity = (fluorescence intensity of blank control group - fluorescence intensity of experimental group) / fluorescence intensity of blank control group × 100%.
[0110] Three replicate wells were set for each treatment, and the results were averaged.
[0111] The serum dilution factor corresponding to a neutralizing activity of 50% is ID50.
[0112] The results showed that serum immunized with recombinant protein pCAGGS-Uni-S-2PM could effectively neutralize SARS-CoV-2 WA1 / 2020 pseudovirus, SARS-CoV-2 alpha mutant pseudovirus, SARS-CoV-2 beta mutant pseudovirus, SARS-CoV-2 gamma mutant pseudovirus, or SARS-CoV-2 delta mutant pseudovirus. The mouse serum immunized with recombinant protein pCAGGS-Uni-S-2PM showed stronger neutralization against the above pseudoviruses than the control protein, indicating that recombinant protein pCAGGS-Uni-S-2PM can induce more and broader-spectrum neutralizing antibodies with cross-neutralizing activity.
[0113] The present invention has been described in detail above. For those skilled in the art, the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although specific embodiments have been given, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein. Some of the essential features can be applied within the scope of the following appended claims. <110> SARS-CoV-2 S protein variants and their application in the preparation of universal vaccines <120> Tsinghua University <160> 11 <170> PatentIn version 3.5 <210> 1 <211> 3787 <212> DNA <213> Artificial sequence <400> 1 gaattatcga tccggaggta ccatggcgag cggccgcgcc atgttcgtgt tcctggtgct 60 gctgcccctg gtgagcagcc aatgcgtgaa cttcaccaca agaacacagc tgccccccgc 120 ctacaccaac agcttcacaa gaggcgtgta ctaccccgac aaggtgttca gaagcagcgt 180 gctgcacagc acccaagacc tgttcctgcc tttcttctcc aacgtgacct ggttccacgc 240 catccacgtg agcggcacca acggcaccaa gagattcgcc aaccccgtgc tgcccttcaa 300 cgacggcgtg tacttcgcta gcaccgagaa gagcaacatc atcagaggct ggatcttcgg 360 caccaccctg gacagcaaga cacagagcct gctgatcgtg aacaacgcca ccaacgtggt 420 gatcaaggtg tgcgagtttc agttctgcaa cgaccccttc ctgggcgtct actaccataa 480 gaacaacaag agctggatgg agagcgagtt cagagtgtac agcagcgcca acaactgcac 540 cttcgagtac gtgagccaac ccttcctgat ggacctggag ggcaagcaag gcaacttcaa 600 gaacctgaga gagttcgtgt tcaagaacat cgacggctac ttcaagatct acagcaagca 660 cacccccatc aacctggtga gaggcctgcc ccaaggcttc agcgccctgg agcccctggt 720 ggacctgccc atcggcatca acatcacaag atttcagacc ctgcacagaa gctacctgac 780 acctggggat agcagctccg gctggaccgc cggcgccgct gcctactacg tgggctacct 840 gcagcctaga accttcctgc tgaagtacaa cgagaacggc accatcacag acgccgtgga 900 ctgccccctg gaccccctga gcgagaccaa gtgcaccctg aagaccttca ccgtggagaa 960 gggcatctat cagacaagca acttcagagt gcagcccacc gagagcatcg tgagattccc 1020 caacatcacc aacctgtgcc ccttcggcga ggtgttcaac gccacaagat tcgctagcgt 1080 gtacgcttgg aaccggaaga gaatcagcaa ctgcgtggcc gactacagcg tgctgtacaa 1140 cagcgctagc ttcagcacct tcaagtgcta cggcgtgagc cccaccaagc tgaacgacct 1200 gtgcttcacc aacgtgtacg ccgacagctt cgtgatcaga ggcgacgagg tgagacagat 1260 cgcccccggg cagaccggca acatcgccga ctacaactac aagctgcccg aggacttcac 1320 cggctgcgtg atcgcctgga acagcaacaa cctggacagc aaggtgggcg gcaactacaa 1380 ctaccggtac agactgttca gaaagagcaa cctgaagccc ttcgagagag acatcagcac 1440 cgagatctac caagccggca gcaccccctg caacggcgtg aagggcttca actgctactt 1500 ccccctgcag agctacggct ttcagcccac ctacggcgtg ggctatcagc cctacagagt 1560 ggtcgtgctg agcttcgagc tgctgcacgc ccccgccacc gtgtgcggcc ccaagaagag 1620 caccaacctg gtgaagaaca agtgcgtcaa tttcaacttc aacggcctga ccggcaccgg 1680 cgtgctgacc gagagcaaca agaagttcct gccctttcag cagttcggca gagacatcgc 1740 cgacaccacc gacgccgtga gagaccctca gaccctggag atcctggaca tcaccccctg 1800 cagcttcggc ggcgtgagcg tgatcacccc cggcaccaac acaagcaacc aagtggccgt 1860 gctgtaccaa ggcgtgaact gcaccgaggt gcccgtggcc atccacgccg atcagctgac 1920 ccccacctgg agagtgtaca gcaccggcag caacgtgttt cagacaagag ccggctgcct 1980 gatcggcgcc gagcacgtga acaacagcta cgagtgcgac atccccatcg gcgccggcat 2040 ctgcgctagc tatcagacac agaccaacag ccctggatct gcctccagcg tggctagcca 2100 aagcatcatc gcctacacca tgagcctggg cgtggagaac agcgtggcct acagcaacaa 2160 tagcatcgcc atccctacca atttcaccat cagcgtgacc accgaaatat taccagtctc 2220 catgaccaag accagcgtgg actgcaccat gtacatctgc ggcgacagca ccgagtgcag 2280 caatctgctg ctgcagtacg gcagcttctg cacccagctg aatagagccc tgaccggcat cgccgtggag caggacaaga atacccagga ggtgttcgcc caggtgaagc agatctacaa gactccgccg atcaaggact tcggcggctt caatttcagc caaatactcc cagatccaag caagcctagc aagaggagct tcatcgagga cctgctgttc aataaggtga ccctggccga cgccggcttc atcaagcagt acggcgactg cctaggtgat attgcggcaa gagacctgat ctgcgcccag aagtttaacg gtttgacagt actacctcct ctgctgaccg acgagatgat 2640 agcacaatat acgtcggcat tgctcgctgg cacgatcaca tcgggctgga ctttcggcgc cggagcagcg ttgcaaatcc ctttcgccat gcagatggcc tacagattca atggcatcgg cgtgacccag aatgtgctgt acgagaatca gaagctgatc gccaatcagt tcaatagcgc catcggcaag atccaggaca gcctgagcag caccgccagc gccctgggca agctgcagga 2880. cgtggtgaat cagaatgccc aggccctgaa taccctggtg aagcagctga gcagcaattt cggcgccatc holds tcaacgatat cctgagcaga ctggaccctc cagaggccga ggtgcaaatt gatcgtctta ttactggcag actgcagagc ctgcagacct acgtgaccca 3060 gcagctgatc agagccgccg agatcagagc cagcgccaat ctggccgcca ccagatgag 3120 cgagtgcgtg ctgggccaga gcaagagagt ggactctgc ggcaagggct accacctgat 3180 gagcttccct cagagcgctc cacatggcgt ggtgttcctg cacgtgacct acgtgcctgc 3240 ccaggagaag aatttcacca ccgcacccgc aatctgccac gacggcagg cccactccc 3300 tagagaggggc gtgttcgtga gcaatggcac ccactggttc gtgacccaga gaaatttcta 3360 cgagcctcag atcatcacca ccgacaatac cttcgtgagc ggcaattgcg acgtggtgat 3420 cgggatagtc ataatactg tctacgaccc tctgcagcct gagctggaca gcttcagga 3480 ggagctggac aagtacttca agaatcacac cagccctgac gtggaccctcg gtgatatttc 3540 gggaatcaat gccagcgtgg tgaatatcca gaaggaatt gatcggctca acgaagtggc 3600 caagaatctg atgagagcc tgatcgacct gcaggagctg ggcaagtacg agcagctctgc 3660 ctggagccac ccacagttcg agaagggcgg cggcagcggc ggcggcggct ccggcggctc 3720 tgcatggtct cacccccagt ttgaaaaggg cagcgactac aaggacgacg atgataaatg 3780 actcgag 3787 <210> 2 <211> 1246 <212> PRT <213> Artificial sequence <400> 2 Met Phe Val Phe Leu Val Leu Leu Pro Leu Val Ser Ser Gln Cys Val 1 5 10 15 Asn Phe Thr Thr Arg Thr Gln Leu Pro Pro Ala Tyr Thr Asn Ser Phe 20 25 30 Thr Arg Gly Val Tyr Tyr Pro Asp Lys Val Phe Arg Ser Ser Val Leu 35 40 45 His Ser Thr Gln Asp Leu Phe Leu Pro Phe Phe Ser Asn Val Thr Trp 50 55 60 Phe His Ala Ile His Val Ser Gly Thr Asn Gly Thr Lys Arg Phe Ala 65 70 75 80 Asn Pro Val Leu Pro Phe Asn Asp Gly Val Tyr Phe Ala Ser Thr Glu 85 90 95 Lys Ser Asn Ile Ile Arg Gly Trp Ile Phe Gly Thr Thr Leu Asp Ser 100 105 110 Lys Thr Gln Ser Leu Leu Ile Val Asn Asn Ala Thr Asn Val Val Ile 115 120 125 Lys Val Cys Glu Phe Gln Phe Cys Asn Asp Pro Phe Leu Gly Val Tyr 130 135 140 Tyr His Lys Asn Asn Lys Ser Trp Met Glu Ser Glu Phe Arg Val Tyr 145 150 155 160 Ser Ser Ala Asn Asn Cys Thr Phe Glu Tyr Val Ser Gln Pro Phe Leu 165 170 175 Met Asp Leu Glu Gly Lys Gln Gly Asn Phe Lys Asn Leu Arg Glu Phe 180 185 190 Val Phe Lys Asn Ile Asp Gly Tyr Phe Lys Ile Tyr Ser Lys His Thr 195 200 205 Pro Ile Asn Leu Val Arg Gly Leu Pro Gln Gly Phe Ser Ala Leu Glu 210 215 220 Pro Leu Val Asp Leu Pro Ile Gly Ile Asn Ile Thr Arg Phe Gln Thr 225 230 235 240 Leu His Arg Ser Tyr Leu Thr Pro Gly Asp Ser Ser Ser Gly Trp Thr 245 250 255 Ala Gly Ala Ala Ala Tyr Tyr Val Gly Tyr Leu Gln Pro Arg Thr Phe 260 265 270 Leu Leu Lys Tyr Asn Glu Asn Gly Thr Ile Thr Asp Ala Val Asp Cys 275 280 285 Ala Leu Asp Pro Leu Ser Glu Thr Lys Cys Thr Leu Lys Thr Phe Thr 290 295 300 Val Glu Lys Gly Ile Tyr Gln Thr Ser Asn Phe Arg Val Gln Pro Thr 305 310 315 320 Glu Ser Ile Val Arg Phe Pro Asn Ile Thr Asn Leu Cys Pro Phe Gly 325 330 335 Glu Val Phe Asn Ala Thr Arg Phe Ala Ser Val Tyr Ala Trp Asn Arg 340 345 350 Lys Arg Ile Ser Asn Cys Val Ala Asp Tyr Ser Val Leu Tyr Asn Ser 355 360 365 Ala Ser Phe Ser Thr Phe Lys Cys Tyr Gly Val Ser Pro Thr Lys Leu 370 375 380 Asn Asp Leu Cys Phe Thr Asn Val Tyr Ala Asp Ser Phe Val Ile Arg 385 390 395 400 Gly Asp Glu Val Arg Gln Ile Ala Pro Gly Gln Thr Gly Asn Ile Ala 405 410 415 Asp Tyr Asn Tyr Lys Leu Pro Asp Asp Phe Thr Gly Cys Val Ile Ala 420 425 430 Trp Asn Ser Asn Asn Leu Asp Ser Lys Val Gly Gly Asn Tyr Asn Tyr 435 440 445 Arg Tyr Arg Leu Phe Arg Lys Ser Asn Leu Lys Pro Phe Glu Arg Asp 450 455 460 Ile Ser Thr Glu Ile Tyr Gln Ala Gly Ser Thr Pro Cys Asn Gly Val 465 470 475 480 Lys Gly Phe Asn Cys Tyr Phe Pro Leu Gln Ser Tyr Gly Phe Gln Pro 485 490 495 Thr Tyr Gly Val Gly Tyr Gln Pro Tyr Arg Val Val Val Leu Ser Phe 500 505 510 Glu Leu Leu His Ala Pro Ala Thr Val Cys Gly Pro Lys Lys Ser Thr 515 520 525 Asn Leu Val Lys Asn Lys Cys Val Asn Phe Asn Phe Asn Gly Leu Thr 530 535 540 Gly Thr Gly Val Leu Thr Glu Ser Asn Lys Lys Phe Leu Pro Phe Gln 545 550 555 560 Gln Phe Gly Arg Asp Ile Ala Asp Thr Thr Asp Ala Val Arg Asp Pro 565 570 575 Gln Thr Leu Glu Ile Leu Asp Ile Thr Pro Cys Ser Phe Gly Gly Val 580 585 590 Ser Val Ile Thr Pro Gly Thr Asn Thr Ser Asn Gln Val Ala Val Leu 595 600 605 Tyr Gln Gly Val Asn Cys Thr Glu Val Pro Val Ala Ile His Ala Asp 610 615 620 Gln Leu Thr Pro Thr Trp Arg Val Tyr Ser Thr Gly Ser Asn Val Phe 625 630 635 640 Gln Thr Arg Ala Gly Cys Leu Ile Gly Ala Glu His Val Asn Asn Ser 645 650 655 Tyr Glu Cys Asp Ile Pro Ile Gly Ala Gly Ile Cys Ala Ser Tyr Gln 660 665 670 Thr Gln Thr Asn Ser Pro Gly Ser Ala Ser Ser Val Ala Ser Gln Ser 675 680 685 Ile Ile Ala Tyr Thr Met Ser Leu Gly Val Glu Asn Ser Val Ala Tyr 690 695 700 Ser Asn Asn Ser Ile Ala Ile Pro Thr Asn Phe Thr Ile Ser Val Thr 705 710 715 720 Thr Glu Ile Leu Pro Val Ser Met Thr Lys Thr Ser Val Asp Cys Thr 725 730 735 Met Tyr Ile Cys Gly Asp Ser Thr Glu Cys Ser Asn Leu Leu Leu Gln 740 745 750 Tyr Gly Ser Phe Cys Thr Gln Leu Asn Arg Ala Leu Thr Gly Ile Ala 755 760 765 Val Glu Gln Asp Lys Asn Thr Gln Glu Val Phe Ala Gln Val Lys Gln 770 775 780 Ile Tyr Lys Thr Pro Pro Ile Lys Asp Phe Gly Gly Phe Asn Phe Ser 785 790 795 800 Gln Ile Leu Pro Asp Pro Ser Lys Pro Ser Lys Arg Ser Phe Ile Glu 805 810 815 Asp Leu Leu Phe Asn Lys Val Thr Leu Ala Asp Ala Gly Phe Ile Lys 820 825 830 Gln Tyr Gly Asp Cys Leu Gly Asp Ile Ala Ala Arg Asp Leu Ile Cys 835 840 845 Ala Gln Lys Phe Asn Gly Leu Thr Val Leu Pro Pro Leu Leu Thr Asp 850 855 860 Glu Met Ile Ala Gln Tyr Thr Ser Ala Leu Leu Ala Gly Thr Ile Thr 865 870 875 880 Ser Gly Trp Thr Phe Gly Ala Gly Ala Ala Leu Gln Ile Pro Phe Ala 885 890 895 Met Gln Met Ala Tyr Arg Phe Asn Gly Ile Gly Val Thr Gln Asn Val 900 905 910 Leu Tyr Glu Asn Gln Lys Leu Ile Ala Asn Gln Phe Asn Ser Ala Ile 915 920 925 Gly Lys Ile Gln Asp Ser Leu Ser Ser Thr Ala Ser Ala Leu Gly Lys 930 935 940 Leu Gln Asp Val Val Asn Gln Asn Ala Gln Ala Leu Asn Thr Leu Val 945 950 955 960 Lys Gln Leu Ser Ser Asn Phe Gly Ala Ile Ser Ser Val Leu Asn Asp 965 970 975 Ile Leu Ser Arg Leu Asp Pro Pro Glu Ala Glu Val Gln Ile Asp Arg 980 985 990 Leu Ile Thr Gly Arg Leu Gln Ser Leu Gln Thr Tyr Val Thr Gln Gln 995 1000 1005 Leu Ile Arg Ala Ala Glu Ile Arg Ala Ser Ala Asn Leu Ala Ala 1010 1015 1020 Thr Lys Met Ser Glu Cys Val Leu Gly Gln Ser Lys Arg Val Asp 1025 1030 1035 Phe Cys Gly Lys Gly Tyr His Leu Met Ser Phe Pro Gln Ser Ala 1040 1045 1050 Pro His Gly Val Val Phe Leu His Val Thr Tyr Val Pro Ala Gln 1055 1060 1065 Glu Lys Asn Phe Thr Thr Ala Pro Ala Ile Cys His Asp Gly Lys 1070 1075 1080 Ala His Phe Pro Arg Glu Gly Val Phe Val Ser Asn Gly Thr His 1085 1090 1095 Trp Phe Val Thr Gln Arg Asn Phe Tyr Glu Pro Gln Ile Ile Thr 1100 1105 1110 Thr Asp Asn Thr Phe Val Ser Gly Asn Cys Asp Val Val Ile Gly 1115 1120 1125 Ile Val Asn Asn Thr Val Tyr Asp Pro Leu Gln Pro Glu Leu Asp 1130 1135 1140 Ser Phe Lys Glu Glu Leu Asp Lys Tyr Phe Lys Asn His Thr Ser 1145 1150 1155 Pro Asp Val Asp Leu Gly Asp Ile Ser Gly Ile Asn Ala Ser Val 1160 1165 1170 Val Asn Ile Gln Lys Glu Ile Asp Arg Leu Asn Glu Val Ala Lys 1175 1180 1185 Asn Leu Asn Glu Ser Leu Ile Asp Leu Gln Glu Leu Gly Lys Tyr 1190 1195 1200 Glu Gln Ser Ala Trp Ser His Pro Gln Phe Glu Lys Gly Gly Gly 1205 1210 1215 Ser Gly Gly Gly Gly Ser Gly Gly Ser Ala Trp Ser His Pro Gln 1220 1225 1230 Phe Glu Lys Gly Ser Asp Tyr Lys Asp Asp Asp Asp Lys 1235 1240 1245 <210> 3 <211> 3796 <212> DNA <213> Artificial sequence <400> 3 gaattatcga tccggaggta ccatggcgag cggccgcgcc atgttcgtgt tcctggtgct 60 gctgcccctg gtctcttctc agtgcgtgaa tctgacaact cggactcagc tgccccctgc 120 ctatacaaat tccttcaccc ggggcgtgta ctatcctgac aaggtgttta gaagctccgt 180 gctgcacagc acacaggatc tgttctgcc attctttcc aacgtgacct ggttccacgc 240 catccacgtg agcggcacca atggcacaaa gcggttcgac aatccagtgc tgccctttaa 300 cgatggcgtg tacttcgcca gcaccgagaa gtccaacatc atcagaggct ggatctttgg 360 caccacactg gactctaaga cacagagcct gctgatcgtg aacaatgcca ccaacgtggt 420 catcaaggtg tgcgagttcc agttttgtaa tgatcccttc ctgggcgtgt actatcacaa 480 gaacaataag tcctggatgg agtctgagtt tagagtgtat tctagcgcca acaattgcac 540 atttgagtac gtgtcccagc ctttcctgat ggacctggag ggcaagcagg gcaatttcaa 600 gaacctgagg gagttcgtgt ttaagaatat cgatggctac ttcaagatct actctaagca 660 cacccctatc aacctggtgc gcgacctgcc acagggcttc agcgccctgg agcctctggt 720 ggatctgcca atcggcatca acatcacccg gtttcagaca ctgctggccc tgcacagaag 780 ctacctgaca cccggcgact cctctagcgg atggaccgca ggagcagcag cctactatgt 840 gggctatctg cagcctcgga ccttcctgct gaagtacaac gagaatggca ccatcacaga 900 cgccgtggat tgcgccctgg atcctctgtc cgagacaaag tgtacactga agtcttttac 960 cgtggagaag ggcatctatc agacatctaa tttcagggtg cagccaaccg agagcatcgt 1020 gcgctttcca aatatcacaa acctgtgccc ctttggcgag gtgttcaacg ccaccagatt 1080 cgccagcgtg tacgcctgga atcggaagag aatcagcaac tgcgtggccg actattccgt 1140 gctgtacaac agcgcctcct tctctacctt taagtgctat ggcgtgtccc ccacaaagct 1200 gaatgacctg tgctttacca acgtgtacgc cgattctttc gtgatcaggg gcgacgaggt 1260 gagacagatc gcaccaggcc agacaggcaa gatcgccgac tacaattata agctgcccga 1320 cgatttcacc ggctgcgtga tcgcctggaa ctccaacaat ctggattcta aagtgggcgg 1380 1440 catcagcaca gagatctacc aggccggctc caccccttgc aatggcgtgg agggctttaa 1500 ctgttatttc cccctgcaga gctacggctt ccagcctaca aacggcgtgg gctatcagcc 1560 ataccgcgtg gtggtgctga gctttgagct gctgcacgca ccagcaacag tgtgcggacc 1620 taagaagtcc accaatctgg tgaagaacaa gtgcgtgaac ttcaacttca acggcctgac 1680 cggcacaggc gtgctgaccg agtccaacaa gaagttcctg ccatttcagc agttcggcag 1740 ggacatcgca gataccacag acgccgtgcg cgacccccag accctggaga tcctggatat 1800 1860 ggtggccgtg ctgtatcagg acgtgaattg taccgaggtg ccagtggcca tccacgccga 1920 tcagctgacc cccacatggc gggtgtactc taccggcagc aacgtgttcc agacaagagc aggatgcctg atcggagcag agcacgtgaa caattcttat gagtgcgaca tcccaatcgg cgccggcatc tgtgccagct accagaccca gacaaactcc ccaggatctg cctcctctgt ggcaagccag tccatcatcg cctataccat gagcctgggc gccgagaatt ccgtggccta 2160. 2220. cgcacaat tccatcgcca tccccacca cttcacaatc tccgtgacca cagagatcct gcccgtgagc atgaccaaga caagcgtgga ctgcacaatg fatherctgtg gcgatagcac cgagtgctcc aacctgctgc tgcagtacgg cagcttttgt acccagctga atagggccct 2340 gacaggcatc gcagtggagc aggataagaa cacagaggag gtgttcgccc aggtgaagca gatctacaag acccccccta tcaaggactt tggcggcttc aacttcagcc agatcctgcc 2460 cgatccttct aagcctagca agcggtcctt tatcgaggac ctgctgttca acaaggtgac cctggccgat gccggcttca tcaagcagta tggcgattgc ctgggcgaca tcgcagcacg 2580 ggacctgatc tgtgcccaga agtttaatgg cctgaccgtg ctgccacccc tgctgacaga 2640 tgagatgatc gcacagtaca catccgccct gctggcaggc accatcacat ctggatggac 2700 cttcggcgca ggagccgccc tgcagatccc ctttgccatg cagatggcct atcgcttcaa 2760 cggcatcggc gtgacccaga atgtgctgta cgagaaccag aagctgatcg ccaatcagtt 2820 taactccgcc atcggcaaga tccaggactc tctgagctcc acagcaagcg ccctgggcaa 2880 gctgcaggat gtggtgaatc agaacgccca ggccctgaat accctggtga agcagctgtc 2940 tagcaacttc ggcgccatct cctctgtgct gaatgatatc ctgagcaggc tggaccctcc 3000 agaggcagag gtgcagatcg acaggctgat cacaggccgc ctgcagtctc tgcagaccta 3060 tgtgacacag cagctgatca gggcagcaga gatcagagca agcgccaatc tggccgccac 3120 caagatgtct gagtgcgtgc tgggccagag caagcgcgtg gacttttgtg gcaagggcta 3180 tcacctgatg tctttcccac agagcgcccc acacggagtg gtgtttctgc acgtgaccta 3240 cgtgcctgcc caggagaaga acttcaccac agccccagcc atctgccacg atggcaaggc 3300 ccactttcca agggagggcg tgttcgtgtc caacggcacc cactggtttg tgacacagcg 3360 caatttctac gagccccaga tcatcaccac agacaatacc ttcgtgagcg gcaactgtga 3420 cgtggtcatc ggcatcgtga acaataccgt gtatgatcct ctgcagccag agctggactc 3480 ctttaaggag gagctggata agtacttcaa gaatcacacc tctcccgacg tggatctggg 3540 cgacatctcc ggcatcaatg cctctgtggt gaacatccag aaggagatcg accggctgaa 3600 cgaggtggcc aagaatctga acgagagcct gatcgatctg caggagctgg gcaagtatga 3660 gcagtctgcc tggagccacc cacagttcga gaagggcggc ggcagcggcg gcggcggctc 3720 cggcggctct gcatggtctc acccccagtt tgaaaagggc agcgactaca aggacgacga 3780 tgataaatga ctcgag 3796 <210> 4 <211> 1249 <212> PRT <213> Artificial sequence <400> 4 Met Phe Val Phe Leu Val Leu Leu Pro Leu Val Ser Ser Gln Cys Val 1 5 10 15 Asn Leu Thr Thr Arg Thr Gln Leu Pro Pro Ala Tyr Thr Asn Ser Phe 20 25 30 Thr Arg Gly Val Tyr Tyr Pro Asp Lys Val Phe Arg Ser Ser Val Leu 35 40 45 His Ser Thr Gln Asp Leu Phe Leu Pro Phe Phe Ser Asn Val Thr Trp 50 55 60 Phe His Ala Ile His Val Ser Gly Thr Asn Gly Thr Lys Arg Phe Asp 65 70 75 80 Asn Pro Val Leu Pro Phe Asn Asp Gly Val Tyr Phe Ala Ser Thr Glu 85 90 95 Lys Ser Asn Ile Ile Arg Gly Trp Ile Phe Gly Thr Thr Leu Asp Ser 100 105 110 Lys Thr Gln Ser Leu Leu Ile Val Asn Asn Ala Thr Asn Val Val Ile 115 120 125 Lys Val Cys Glu Phe Gln Phe Cys Asn Asp Pro Phe Leu Gly Val Tyr 130 135 140 Tyr His Lys Asn Asn Lys Ser Trp Met Glu Ser Glu Phe Arg Val Tyr 145 150 155 160 Ser Ser Ala Asn Asn Cys Thr Phe Glu Tyr Val Ser Gln Pro Phe Leu 165 170 175 Met Asp Leu Glu Gly Lys Gln Gly Asn Phe Lys Asn Leu Arg Glu Phe 180 185 190 Val Phe Lys Asn Ile Asp Gly Tyr Phe Lys Ile Tyr Ser Lys His Thr 195 200 205 Pro Ile Asn Leu Val Arg Asp Leu Pro Gln Gly Phe Ser Ala Leu Glu 210 215 220 Pro Leu Val Asp Leu Pro Ile Gly Ile Asn Ile Thr Arg Phe Gln Thr 225 230 235 240 Leu Leu Ala Leu His Arg Ser Tyr Leu Thr Pro Gly Asp Ser Ser Ser 245 250 255 Gly Trp Thr Ala Gly Ala Ala Ala Tyr Tyr Val Gly Tyr Leu Gln Pro 260 265 270 Arg Thr Phe Leu Leu Lys Tyr Asn Glu Asn Gly Thr Ile Thr Asp Ala 275 280 285 Val Asp Cys Ala Leu Asp Pro Leu Ser Glu Thr Lys Cys Thr Leu Lys 290 295 300 Ser Phe Thr Val Glu Lys Gly Ile Tyr Gln Thr Ser Asn Phe Arg Val 305 310 315 320 Gln Pro Thr Glu Ser Ile Val Arg Phe Pro Asn Ile Thr Asn Leu Cys 325 330 335 Pro Phe Gly Glu Val Phe Asn Ala Thr Arg Phe Ala Ser Val Tyr Ala 340 345 350 Trp Asn Arg Lys Arg Ile Ser Asn Cys Val Ala Asp Tyr Ser Val Leu 355 360 365 Tyr Asn Ser Ala Ser Phe Ser Thr Phe Lys Cys Tyr Gly Val Ser Pro 370 375 380 Thr Lys Leu Asn Asp Leu Cys Phe Thr Asn Val Tyr Ala Asp Ser Phe 385 390 395 400 Val Ile Arg Gly Asp Glu Val Arg Gln Ile Ala Pro Gly Gln Thr Gly 405 410 415 Lys Ile Ala Asp Tyr Asn Tyr Lys Leu Pro Asp Asp Phe Thr Gly Cys 420 425 430 Val Ile Ala Trp Asn Ser Asn Asn Leu Asp Ser Lys Val Gly Gly Asn 435 440 445 Tyr Asn Tyr Leu Tyr Arg Leu Phe Arg Lys Ser Asn Leu Lys Pro Phe 450 455 460 Glu Arg Asp Ile Ser Thr Glu Ile Tyr Gln Ala Gly Ser Thr Pro Cys 465 470 475 480 Asn Gly Val Glu Gly Phe Asn Cys Tyr Phe Pro Leu Gln Ser Tyr Gly 485 490 495 Phe Gln Pro Thr Asn Gly Val Gly Tyr Gln Pro Tyr Arg Val Val Val 500 505 510 Leu Ser Phe Glu Leu Leu His Ala Pro Ala Thr Val Cys Gly Pro Lys 515 520 525 Lys Ser Thr Asn Leu Val Lys Asn Lys Cys Val Asn Phe Asn Phe Asn 530 535 540 Gly Leu Thr Gly Thr Gly Val Leu Thr Glu Ser Asn Lys Lys Phe Leu 545 550 555 560 Pro Phe Gln Gln Phe Gly Arg Asp Ile Ala Asp Thr Thr Asp Ala Val 565 570 575 Arg Asp Pro Gln Thr Leu Glu Ile Leu Asp Ile Thr Pro Cys Ser Phe 580 585 590 Gly Gly Val Ser Val Ile Thr Pro Gly Thr Asn Thr Ser Asn Gln Val 595 600 605 Ala Val Leu Tyr Gln Asp Val Asn Cys Thr Glu Val Pro Val Ala Ile 610 615 620 His Ala Asp Gln Leu Thr Pro Thr Trp Arg Val Tyr Ser Thr Gly Ser 625 630 635 640 Asn Val Phe Gln Thr Arg Ala Gly Cys Leu Ile Gly Ala Glu His Val 645 650 655 Asn Asn Ser Tyr Glu Cys Asp Ile Pro Ile Gly Ala Gly Ile Cys Ala 660 665 670 Ser Tyr Gln Thr Gln Thr Asn Ser Pro Gly Ser Ala Ser Ser Val Ala 675 680 685 Ser Gln Ser Ile Ile Ala Tyr Thr Met Ser Leu Gly Ala Glu Asn Ser 690 695 700 Val Ala Tyr Ser Asn Asn Ser Ile Ala Ile Pro Thr Asn Phe Thr Ile 705 710 715 720 Ser Val Thr Thr Glu Ile Leu Pro Val Ser Met Thr Lys Thr Ser Val 725 730 735 Asp Cys Thr Met Tyr Ile Cys Gly Asp Ser Thr Glu Cys Ser Asn Leu 740 745 750 Leu Leu Gln Tyr Gly Ser Phe Cys Thr Gln Leu Asn Arg Ala Leu Thr 755 760 765 Gly Ile Ala Val Glu Gln Asp Lys Asn Thr Gln Glu Val Phe Ala Gln 770 775 780 Val Lys Gln Ile Tyr Lys Thr Pro Pro Ile Lys Asp Phe Gly Gly Phe 785 790 795 800 Asn Phe Ser Gln Ile Leu Pro Asp Pro Ser Lys Pro Ser Lys Arg Ser 805 810 815 Phe Ile Glu Asp Leu Leu Phe Asn Lys Val Thr Leu Ala Asp Ala Gly 820 825 830 Phe Ile Lys Gln Tyr Gly Asp Cys Leu Gly Asp Ile Ala Ala Arg Asp 835 840 845 Leu Ile Cys Ala Gln Lys Phe Asn Gly Leu Thr Val Leu Pro Pro Leu 850 855 860 Leu Thr Asp Glu Met Ile Ala Gln Tyr Thr Ser Ala Leu Leu Ala Gly 865 870 875 880 Thr Ile Thr Ser Gly Trp Thr Phe Gly Ala Gly Ala Ala Leu Gln Ile 885 890 895 Pro Phe Ala Met Gln Met Ala Tyr Arg Phe Asn Gly Ile Gly Val Thr 900 905 910 Gln Asn Val Leu Tyr Glu Asn Gln Lys Leu Ile Ala Asn Gln Phe Asn 915 920 925 Ser Ala Ile Gly Lys Ile Gln Asp Ser Leu Ser Ser Thr Ala Ser Ala 930 935 940 Leu Gly Lys Leu Gln Asp Val Val Asn Gln Asn Ala Gln Ala Leu Asn 945 950 955 960 Thr Leu Val Lys Gln Leu Ser Ser Asn Phe Gly Ala Ile Ser Ser Val 965 970 975 Leu Asn Asp Ile Leu Ser Arg Leu Asp Pro Pro Glu Ala Glu Val Gln 980 985 990 Ile Asp Arg Leu Ile Thr Gly Arg Leu Gln Ser Leu Gln Thr Tyr Val 995 1000 1005 Thr Gln Gln Leu Ile Arg Ala Ala Glu Ile Arg Ala Ser Ala Asn 1010 1015 1020 Leu Ala Ala Thr Lys Met Ser Glu Cys Val Leu Gly Gln Ser Lys 1025 1030 1035 Arg Val Asp Phe Cys Gly Lys Gly Tyr His Leu Met Ser Phe Pro 1040 1045 1050 Gln Ser Ala Pro His Gly Val Val Phe Leu His Val Thr Tyr Val 1055 1060 1065 Pro Ala Gln Glu Lys Asn Phe Thr Thr Ala Pro Ala Ile Cys His 1070 1075 1080 Asp Gly Lys Ala His Phe Pro Arg Glu Gly Val Phe Val Ser Asn 1085 1090 1095 Gly Thr His Trp Phe Val Thr Gln Arg Asn Phe Tyr Glu Pro Gln 1100 1105 1110 Ile Ile Thr Thr Asp Asn Thr Phe Val Ser Gly Asn Cys Asp Val 1115 1120 1125 Val Ile Gly Ile Val Asn Asn Thr Val Tyr Asp Pro Leu Gln Pro 1130 1135 1140 Glu Leu Asp Ser Phe Lys Glu Glu Leu Asp Lys Tyr Phe Lys Asn 1145 1150 1155 His Thr Ser Pro Asp Val Asp Leu Gly Asp Ile Ser Gly Ile Asn 1160 1165 1170 Ala Ser Val Val Asn Ile Gln Lys Glu Ile Asp Arg Leu Asn Glu 1175 1180 1185 Val Ala Lys Asn Leu Asn Glu Ser Leu Ile Asp Leu Gln Glu Leu 1190 1195 1200 Gly Lys Tyr Glu Gln Ser Ala Trp Ser His Pro Gln Phe Glu Lys 1205 1210 1215 Gly Gly Gly Ser Gly Gly Gly Gly Ser Gly Gly Ser Ala Trp Ser 1220 1225 1230 His Pro Gln Phe Glu Lys Gly Ser Asp Tyr Lys Asp Asp Asp Asp 1235 1240 1245 Lys <210> 5 <211> 4790 <212> DNA <213> Artificial sequence <400> 5 gtcgacattg attattgact agttattaat agtaatcaat tacggggtca ttagttcata 60 gcccatatat ggagttccgc gttacataac ttacggtaaa tggcccgcct ggctgaccgc 120 ccaacgaccc ccgcccattg acgtcaataa tgacgtatgt tcccatagta acgccaatag 180 ggactttcca ttgacgtcaa tgggtggact atttacggta aactgcccac ttggcagtac 240 atcaagtgta tcatatgcca agtacgcccc ctattgacgt caatgacggt aaatggcccg 300 cctggcatta tgcccagtac atgaccttat gggactttcc tacttggcag tacatctacg 360 tattagtcat cgctattacc atgggtcgag gtgagcccca cgttctgctt cactctcccc 420 atctcccccc cctccccacc cccaattttg tatttattta ttttttaatt attttgtgca 480 gcgatgggg cgggggggg gggggcgcgc gccaggcggg gcggggcggg gcgaggggcg 540 gggcggggcg aggcggagag gtgcggcggc agccaatcag agcggcgcgc tccgaaagtt 600 tcctttatg gcgaggcggc ggcggcggcg gccctataaa aagcgaagcg cgcggcgggc 660 gggagtcgct gcgttgcctt cgccccgtgc cccgctccgc gccgcctcgc gccgcccgcc 720 ccggctctga ctgaccgcgt tactcccaca ggtgagcggg cgggacggcc cttctcctcc 780 gggctgtaat tagcgcttgg tttaatgacg gctcgtttct tttctgtggc tgcgtgaaag 840 ccttaaaggg ctccgggagg gccctttgtg cgggggggag cggctcgggg ggtgcgtgcg 900 tgtgtgtgtg cgtggggagc gccgcgtgcg gcccgcgctg cccggcggct gtgagcgctg 960 cgggcgcggc gcggggcttt gtgcgctccg cgtgtgcgcg aggggagcgc ggccgggggc 1020 ggtgccccgc ggtgcggggg ggctgcgagg ggaacaaagg ctgcgtgcgg ggtgtgtgcg 1080 tgggggggtg agcagggggt gtgggcgcgg cggtcgggct gtaacccccc cctgcacccc 1140 cctccccgag ttgctgagca cggcccggct tcgggtgcgg ggctccgtgc ggggcgtggc 1200 gcggggctcg ccgtgccggg cggggggtgg cggcaggtgg gggtgccggg cggggcgggg 1260 ccgcctcggg ccggggaggg ctcgggggag gggcgcggcg gccccggagc gccggcggct 1320 gtcgaggcgc ggcgagccgc agccattgcc ttttatggta atcgtgcgag agggcgcagg 1380 gacttccttt gtcccaaatc tggcggagcc gaaatctggg aggcgccgcc gcaccccctc 1440 tagcgggcgc gggcgaagcg gtgcggcgcc ggcaggaagg aaatgggcgg ggagggcctt 1500 cgtgcgtcgc cgcgccgccg tccccttctc catctccagc ctcggggctg ccgcaggggg 1560 acggctgcct tcggggggga cggggcaggg cggggttcgg cttctggcgt gtgaccggcg 1620 gctctagagc ctctgctaac catgttcatg ccttcttctt tttcctacag ctcctgggca 1680 acgtgctggt tgttgtgctg tctcatcatt ttggcaaaga attcctcgag gaattcactc 1740 ctcaggtgca ggctgcctat cagaaggtgg tggctggtgt ggccaatgcc ctggctcaca 1800 aataccactg agatcttttt ccctctgcca aaaattatgg ggacatcatg aagccccttg 1860 agcatctgac ttctggctaa taaaggaaat ttattttcat tgcaatagtg tgttggaatt 1920 ttttgtgtct ctcactcgga aggacatatg ggagggcaaa tcatttaaaa catcagaatg 1980 agtatttggt ttagagtttg gcaacatatg ccatatgctg gctgccatga acaaaggtgg 2040 ctataaagag gtcatcagta tatgaaacag ccccctgctg tccattcctt attccataga 2100 aaagccttga cttgaggtta gatttttttt atattttgtt ttgtgttatt tttttcttta 2160 acatccctaa aattttcctt acatgtttta ctagccagat ttttcctcct ctcctgacta 2220 ctcccagtca tagctgtccc tcttctctta tgaagatccc tcgacctgca gcccaagctt ggcgtaatca tggtcatagc tgtttcctgt gtgaaattgt tatccgctca caattccaca 2400. 2400. gccggagca taagtgtaa agcctggggt gcctaatgag tgagctaact cacattaatt gcgttgcgct cactgcccgc tttccagtcg ggaaacctgt cgtgccagcg gatccgcatc tcaattagtc agcaaccata gtcccgcccc taactccgcc catcccgccc ctaactccgc ccagttccgc ccattctccg ccccatggct gactaatttt ttttatttat 2580. gcagaggccg aggccgcctc ggcctctgag ctattccaga agtagtgagg aggctttttt 2640 ggaggcctag gcttttgcaa aaagctaact tgtttattgc agcttataat ggttacaaat aaagcaatag catcacaaat ttcacaaata aagcattttt ttcactgcat tctagttgtg gtttgtccaa actcatcaat gtatcttatc atgtctggat ccgctgcatt aatgaatcgg ccaacgcgcg gggagaggcg gtttgcgtat tgggcgctct tccgcttcct cgctcactga 2880. ctcgctgcgc tcggtcgttc ggctgcggcg agcggtatca gctcactcaa agcggtaat 2940. acggttatcc acagaatcag gggataacgc aggaaagaac atgtgagcaa aaggccagca 3000 aaaggccagg aaccgtaaaa aggccgcgtt gctggcgttt ttccataggc tccgcccccc 3060 tgacgagcat cacaaaaaatc gacgctcaag tcagaggtgg cgaaacccga caggactata 3120 aagataccag gcgtttcccc ctggaagctc cctcgtgcgc tctcctgttc cgaccctgcc 3180 gcttaccgga tacctgtccg cctttctccc ttcgggaagc gtggcgcttt ctcaatgctc 3240 acgctgtagg tatctcagtt cggtgtaggt cgttcgctcc aagctggggct gtgtgcacga 3300 accccccgtt cagcccgacc gctgcgcctt atccggtaac tatcgtcttg agtccaaccc 3360 ggtaagacac gacttatcgc cactggcagc agccactggt aacaggatta gcagagcgag 3420 gtatgtaggc ggtgctacag agttcttgaa gtggtggcct aactacggct acactagaag 3480 gacagtattt ggtatctgcg ctctgctgaa gccagttacc ttcggaaaaa gagttggtag 3540 ctcttgatcc ggcaaacaaa ccaccgctgg tagcggtggt ttttttgttt gcaagcagca 3600 gattacgcgc agaaaaaag gatctcaaga agatcctttg atctttcta cggggtctga 3660 cgctcagtgg aacgaaaact cacgttaagg gattttggtc atgagattat caaaaaggat 3720 cttcacctag atccttttaa attaaaaatg aagttttaaa tcaatctaaa gtatatatga 3780 gtaaacttgg tctgacagtt accaatgctt aatcagtgag gcacctatct cagcgatctg 3840 tctatttcgt tcatccatag ttgcctgact ccccgtcgtg tagataacta cgatacggga 3900 gggcttacca tctggcccca gtgctgcaat gataccgcga gacccacgct caccggctcc 3960 agatttatca gcaataaacc agccagccgg aagggccgag cgcagaagtg gtcctgcaac 4020 tttatccgcc tccatccagt ctattaattg ttgccgggaa gctagagtaa gtagttcgcc 4080 agttaatagt ttgcgcaacg ttgttgccat tgctacaggc atcgtggtgt cacgctcgtc 4140 gtttggtatg gcttcattca gctccggttc ccaacgatca aggcgagtta catgatcccc 4200 catgttgtgc aaaaaagcgg ttagctcctt cggtcctccg atcgttgtca gaagtaagtt 4260 ggccgcagtg ttcactca tggttatggc agcactgcat aattctctta ctgtcatgcc 4320 atccgtaaga tgcttttctg tgactggtga gtactcaacc aagtcattct gagaatagtg 4380 tatgcggcga ccgagttgct cttgcccggc gtcaatacgg gataataccg cgccacatag cagaacttta aaagtgctca tcattggaa acgttcttcg gggcgaaaac tctcaaggat cttaccgctg ttgagatcca gttcgatgta acccactcgt gcacccaact gatcttcagc 4560 atcttttact ttcaccagcg tttctgggtg agcaaaaaca ggaaggcaaa atgccgcaaa aaagggaata agggcgacac ggaaatgttg aatactcata ctcttccttt ttcaatatta ttgaagcatt tatcagggtt attgtctcat gagcggatac attttgaat gtatttagaa 4790. cgcgcacatt tccccgaaaa gtgccacctg <210> 6 <211> 3846 <212> DNA <213> Artificial sequence <400> 6 atgttcgtgt tcctggtgct gctgcctctg gtgagcagcc agtgcgtgaa tctgaccacc agaacccagc tgcctcctgc ctacaccat agcttcacca gaggagttta ttatcccgat aaggtgttca gaagtagtgt attacatagt acccaggacc tgttcctacc tttcttcagt 180 aacgtgacct ggttccacgc catccacgtg agcggcacca atggcaccaa gagattcgac 240 aatcctgtgc tgcctttcaa tgacggcgtg tacttcgcca gcaccgagaa gagcaatatc 300 atcagaggct ggatcttcgg caccaccttg gattccaaga ctcagagcct gctgattgta 360 aacaacgcta caaatgtggt gatcaaggtg tgcgagttcc agttctgcaa tgaccctttc 420 ctgggtgttt attatcataa gaacaacaag agctggatgg agagcgagtt ccgcgtatat 480 tcgtcggcta ataattgcac cttcgagtac gtgagccagc ctttcctgat ggacctggag 540 ggcaagcagg gcaatttcaa gaatctgaga gagttcgtgt tcaagaatat cgacggctac 600 ttcaagatct acagcaagca cacacccatt aatctggtga gagacctgcc tcagggcttc 660 agcgccctgg agcctctggt ggacctgcct atcggcatca atatcaccag attccagacc 720 ctgctggccc tgcacagatc atatcttaca ccaggcgatt cgtcaagcgg ttggaccgct 780 ggagctgcgg catattacgt gggctacctg cagcctagaa ccttcctgct gaagtacaat 840 gagaatggta cgataaccga cgcagttgat tgtgccctgg accctctgag cgagaccaag 900 tgcaccctga agagcttcac cgtggagaag ggcatctacc agaccagcaa tttcagagtg 960 cagcctaccg agagcatcgt gagattccct aatatcacca atctgtgccc ttcggcgag 1020 gtgttcaatg ccaccagatt cgccagcgtg tacgcatgga accgcaagcg gataagcaat 1080 tgcgtggccg actacagcgt gctgtacaat agcgccagct tcagcacctt caaatgttat 1140 ggtgttcgc caacaaagct gaatgacctg tgcttcacca atgtgtacgc cgacagcttc 1200 gtgatcagag gcgacgaggt gagacagatc gcgccagggc agaccggcaa gatcgccgac 1260 tacaattaca agctgcctga cgacttcacc ggctgcgtga tcgcgtggaa ctctaacaat 1320 ctagattcga aagttggagg caattacaat tacctgtaca gactgttcag aaagagcaat 1380 ctgaagcctt tcgagagaga catcagcacc gagatctacc aggccggcag cacaccgtgt 1440 aatggcgtgg agggcttcaa ttgctacttc cctctgcaga gctacggctt ccagcctacc 1500 aatggcgtgg gctaccagcc ttacagagtg gtggtgctga gcttcgagct gctgcacgct 1560 cccgctaccg tgtgcggccc taagaagagc accaatctgg tgaagaataa gtgcgtgaat 1620 ttcaatttca atggtctaac tggaacgggc gtgctgaccg agagcaataa gaagtttctt 1680 ccctttcaac aattcggcag agacatcgcc gacaccacag atgctgtaag agaccctcag 1740 accctggaga tcctggacat cactccgtgt agcttcggcg gcgtgagcgt gatcacaccg 1800 ggtaccaata ccagcaatca ggtggccgtg ctgtaccagg acgtgaattg caccgaggtg 1860 cctgtggcca tccacgccga ccagctgact cccacttgga gggtatattc cacgggaagc 1920 aatgtgttcc agaccagagc cggctgcctg atcggcgccg agcacgtgaa taatagctac 1980 gagtgcgaca tccctatcgg cgccggcatc tgcgccagct accagaccca gaccaatagc 2040 cctagaagag ccagaagcgt ggccagccag agcatcatcg cctacaccat gagcctgggc 2100 gccgagaata gcgtggccta cagcaataat agcatcgcca tccctaccaa tttcaccatc 2160 agcgtgacca ccgaaatatt accagtctcc atgaccaaga ccagcgtgga ctgcaccatg 2220 tacatctgcg gcgacagcac cgagtgcagc aatctgctgc tgcagtacgg cagcttctgc 2280 acccagctga atagagccct gaccggcatc gccgtggagc aggacaagaa tacccaggag 2340 gtgttcgccc aggtgaagca gatctacaag actccgccga tcaaggactt cggcggcttc 2400 aatttcagcc aaatactccc agatccaagc aagcctagca agaggagctt catcgaggac 2460 ctgctgttca ataaggtgac cctggccgac gccggcttca tcaagcagta cggcgactgc 2520 ctaggtgata ttgcggcaag agacctgatc tgcgcccaga agtttaacgg tttgacagta 2580 ctacctcctc tgctgaccga cgagatgata gcacaatata cgtcggcatt gctcgctggc 2640 acgatcacat cgggctggac tttcggcgcc ggagcagcgt tgcaaatccc tttcgccatg 2700 cagatggcct acagattcaa tggcatcggc gtgacccaga atgtgctgta cgagaatcag 2760 aagctgatcg ccaatcagtt caatagcgcc atcggcaaga tccaggacag cctgagcagc 2820 accgccagcg ccctgggcaa gctgcaggac gtggtgaatc agaatgccca ggccctgaat 2880 accctggtga agcagctgag cagcaatttc ggcgccatca gtagtgtact caacgatatc 2940 ctgagcagac tggacaaggt ggaggccgag gtgcaaattg atcgtcttat tactggcaga 3000 ctgcagagcc tgcagaccta cgtgacccag cagctgatca gagccgccga gatcagagcc 3060 agcgccaatc tggccgccac caagatgagc gagtgcgtgc tgggccagag caagagagtg 3120 gacttctgcg gcaagggcta ccacctgatg agcttccctc agagcgctcc acatggcgtg 3180 gtgttcctgc acgtgaccta cgtgcctgcc your foot atttcaccac cgcacccgca 3240 atctgccacg acggcaaggc ccacttccct agagagggcg tgttcgtgag caatggcacc 3300 cactggttcg tgacccagag aaatttctac gagcctcaga tcatcaccac cgacaatacc 3360 3420 ctgcagcctg agctggacag cttcaaggag gagctggaca agtacttcaa gaatcacacc 3480 agccctgacg tggacctcgg tgatatttcg ggaatcaatg ccagcgtggt gaatatccag 3540 areaattg atcggctcaa cgaagtggcc aagaatctga atgagagcct gatcgacctg 3600 caggagctgg gcaagtacga gcagtacatc aagtggcctt ggtacatctg gctggggcttc 3660 atcgccggcc tgatcgccat cgtgatggtg accatcatgc tgtgctgcat gacctcctgt 3720 tgttcctgtt tgaaagggtg ttgttcgtgt gggtcctgct gcaagttcga cgaggacgac 3780 agcgagcctg tgctgaaggg cgtgaagctg cactacacct ggagccaccc tcagttcgag 3840 August 3846 <210> 7 <211> 3813 <212> DNA <213> Artificial sequence <400> 7 atgttcgtgt tcctggtgct gctgcctctg gtgagcagcc agtgcgtgaa tctgaccacc 60 agaacccagc tgcctcctgc ctacaccaat agcttcacca gaggagttta ttatcccgat 120 aaggtgttca gaagtagtgt attacatagt acccaggacc tgttcctacc tttcttcagt 180 aacgtgacct ggttccacgc catcagcggc accaatggca ccaagagatt cgacaatcct 240 gtgctgcctt tcaatgacgg cgtgtacttc gccagcaccg agaagagcaa tatcatcaga 300 ggctggatct tcggcaccac cttggattcc aagactcaga gcctgctgat tgtaaacaac 360 gctacaaatg tggtgatcaa ggtgtgcgag ttccagttct gcaatgaccc tttcctgggt 420 gtttatcata agaacaacaa gagctggatg gagagcgagt tccgcgtata ttcgtcggct 480 aataattgca ccttcgagta cgtgagccag cctttcctga tggacctgga gggcaagcag 540 ggcaatttca agaatctgag agagttcgtg ttcaagaata tcgacggcta cttcaagatc 600 tacagcaagc acacacccat taatctggtg agagacctgc ctcagggctt cagcgccctg 660 gagcctctgg tggacctgcc tatcggcatc aatatcacca gattccagac cctgctggcc 720 ctgcacagat catatcttac accaggcgat tcgtcaagcg gttggaccgc tggagctgcg 780 gcatattacg tgggctacct gcagcctaga accttcctgc tgaagtacaa tgagaatggt 840 acgataaccg acgcagttga ttgtgccctg gaccctctga gcgagaccaa gtgcaccctg 900 aagagcttca ccgtggagaa gggcatctac cagaccagca atttcagagt gcagcctacc 960 gagagcatcg tgagattccc taatatcacc aatctgtgcc ctttcggcga ggtgttcaat 1020 gccaccagat tcgccagcgt gtacgcatgg aaccgcaagc ggataagcaa ttgcgtggcc 1080 gactacagcg tgctctacaa tagcgccagc ttcagcacct tcaaatgtta tggtgtttcg 1140 ccaacaaagc tgaatgacct gtgcttcacc aatgtgtacg ccgacagctt cgtgatcaga 1200 ggcgacgagg tgagacagat cgcgccaggg cagaccggca agatcgccga ctacaattac 1260 aagctgcctg acgacttcac cggctgcgtg atcgcgtgga actctaacaa tctagattcg 1320 aaagttggag gcaattacaa ttacctgtac agactgttca gaaagagcaa tctgaagcct 1380 ttcgagagag acatcagcac cgagatctac caggccggca gcacaccgtg taatggcgtg 1440 gagggcttca attgctactt ccctctgcag agctacggct tccagcctac ctatggcgtg 1500 ggctaccagc cttacagagt ggtggtgctg agcttcgagc tgctgcacgc tcccgctacc 1560 gtgtgcggcc ctaagaagag caccaatctg gtgaagaata agtgcgtgaa tttcaatttc 1620 aatggtctaa ctggaacggg cgtgctgacc gagagcaata agaagtttct tccctttcaa 1680 caattcggca gagacatcga cgacaccaca gatgctgtaa gagaccctca gaccctggag 1740 atcctggaca tcactccgtg tagcttcggc ggcgtgagcg tgatcacacc gggtaccaat 1800 accagcaatc aggtggccgt gctgtaccag ggcgtgaatt gcaccgaggt gcctgtggcc 1860 atccacgccg accagctgac tcccacttgg agggtatatt ccacgggaag caatgtgttc 1920 cagaccagag ccggctgcct gatcggcgcc gagcacgtga ataatagcta cgagtgcgac 1980 atccctatcg gcgccggcat ctgcgccagc taccagaccc gccagaagcg tggccagcca gagcatcatc gcctacacca tgagcctggg cgccgagaat agcgtggcct together tagcatcgcc atccctatca atttcaccat together accgaaatat taccagtctc catgaccaag accagcgtgg actgcaccat gtacatctgc ggcgacagca ccgagtgcag caatctgctg ctgcagtacg gcagcttctg cacccagctg 2340. aatagagccc tgaccggcat cgccgtggag caggacaaga atacccagga ggtgttcgcc caggtgaagc agatctacaa gactccgccg atcaaggact tcggcggctt caatttcagc caaatactcc cagatccaag caagcctagc aagaggagct tcatcgagga cctgctgttc aataggtga ccctggccga cgccggcttc atcaagcagt acggcgactg cctaggtgat attgcggcaa gagacctgat ctgcgcccag aagtttaacg gtttgacagt actacctcct ctgctgaccg bite agcacaatat acgtcggcat tgctcgctgg bitecaca tcgggctgga ctttcggcgc cggagcagcg ttgcaaatcc ctttcgccat gcagatggcc 2700. you are atggcatcgg cgtgacccag aatgtgctgt you are gaagctgatc gccaatcagt tcaatagcgc catcggcaag atccaggaca gcctgagcag caccgccagc gccctgggca agctgcagga cgtggtgaat cagaatgccc aggccctgaa taccctggtg aagcagctga gcagcaattt cggcgccatc agtagtgtac tcaacgatat cctggccaga ctggacaagg tggaggccga ggtgcaaatt gatcgtctta ttactggcag actgcagagc ctgcagacct acgtgaccca gcagctgatc agagccgccg agatcagagc cagcgccaat ctggccgcca ccaagatgag cgagtgcgtg ctgggccaga gcaagagagt ggacttctgc 3120. ggcaagggct accacctgat gagcttccct cagagcgctc cacatggcgt ggtgttcctg 3180. 3240. cacgtgacct acgtgcctgc ccaggagaag aatttcacca ccgcacccgc aatctgccac gacggcaagg cccacttccc tagagggc gtgttcgtga gcaatggcac ccactggttc 3300. 3360. gtgacccaga gaaatttcta cgagcctcag atcatcacca cccacaatac cttcgtgagc ggcaattgcg acgtggtgat cgggatagtc fatheractg tctacgaccc tctgcagcct 3420. gagctggaca gcttcaagga ggagctggac aagtacttca agaatcacac cagccctgac 3480 gtggacctcg gtgatatttc gggaatcaat gccagcgtgg tgaatatcca gaaggaaatt 3540 gatcggctca acgaagtggc caagaatctg aatgagagcc tgatcgacct gcaggagctg 3600 ggcaagtacg agcagtacat caagtggcct tggtacatct ggctgggctt catcgccggc 3660 ctgatcgcca tcgtgatggt gaccatcatg ctgtgctgca tgacctcctg ttgttcctgt 3720 ttgaaagggt gttgttcgtg tgggtcctgc tgcaagttcg acgaggacga cagcgagcct 3780 gtgctgaagg gcgtgaagct gcactacacc tga 3813 <210> 8 <211> 3813 <212> DNA <213> Artificial sequence <400> 8 atgttcgtgt tcctggtgct gctgcccctg gtgagcagcc aatgcgtgaa cttcaccaca 60 agaacacagc tgccccccgc ctacaccaac agcttcacaa gaggcgtgta ctaccccgac 120 aaggtgttca gaagcagcgt gctgcacagc acccaagacc tgttcctgcc tttcttctcc 180 aacgtgacct ggttccacgc catccacgtg agcggcacca acggcaccaa gagattcgcc 240 aaccccgtgc tgcccttcaa cgacggcgtg tacttcgcta gcaccgagaa gagcaacatc 300 atcagaggct ggatcttcgg caccaccctg gacagcaaga cacagagcct gctgatcgtg 360 aacaacgcca ccaacgtggt gatcaaggtg tgcgagtttc agttctgcaa cgaccccttc 420 ctgggcgtct actaccataa gaacaacaag agctggatgg agagcgagtt cagagtgtac 480 agcagcgcca acaactgcac cttcgagtac gtgagccaac ccttcctgat ggacctggag 540 ggcaagcaag gcaacttcaa gaacctgaga gagttcgtgt tcaagaacat cgacggctac 600 ttcaagatct acagcaagca cacccccatc aacctggtga gaggcctgcc ccaaggcttc 660 agcgccctgg agcccctggt ggacctgccc atcggcatca acatcacaag atttcagacc 720 ctgcacagaa gctacctgac acctggggat agcagctccg gctggaccgc cggcgccgct 780 gcctactacg tgggctacct gcagcctaga accttcctgc tgaagtacaa cgagaacggc 840 accatcacag acgccgtgga ctgcgccctg gaccccctga gcgagaccaa gtgcaccctg 900 aagaccttca ccgtggagaa gggcatctat cagacaagca acttcagagt gcagcccacc 960 gagagcatcg tgagattccc caacatcacc aacctgtgcc ccttcggcga ggtgttcaac 1020 gccacaagat tcgctagcgt gtacgcttgg aaccgaaga gaatcagcaa ctgcgtggcc 1080 gactacagcg tgctgtacaa cagcgctagc ttcagcacct tcaagtgcta cggcgtgagc 1140 cccaccaagc tgaacgacct gtgcttcacc aacgtgtacg ccgacagctt cgtgatcaga 1200 ggcgacgagg tgagacagat cgcccccggg cagaccggca acatcgccga ctacaactac 1260 aagctgcccg acgacttcac cggctgcgtg atcgcctgga acagcaacaa cctggacagc 1320 aaggtgggcg gcaactacaa ctacctgtac agactgttca gaaagagcaa cctgaagccc 1380 ttcgagagag acatcagcac cgagatctac caagccggca gcaccccctg caacggcgtg 1440 aagggcttca actgctactt ccccctgcag agctacggct ttcagcccac ctacggcgtg 1500 ggctatcagc cctacagagt ggtcgtgctg agcttcagc tgctgcacgc ccccgccacc 1560 gtgtgcggcc ccaagaagag caccaacctg gtgaagaaca agtgcgtcaa tttcaacttc 1620 aacggcctga ccggcaccgg cgtgctgacc gagagcaaca agaagttcct gccctttcag 1680 cagttcggca gagacatcgc cgacaccacc gacgccgtga gagaccctca gaccctggag 1740 atcctggaca tcaccccctg cagcttcggc ggcgtgagcg tgatcacccc cggcaccaac 1800 acaagcaacc aagtggccgt gctgtaccaa ggcgtgaact gcaccgaggt gcccgtggcc 1860 atccacgccg atcagctgac ccccacctgg agagtgtaca gcaccggcag caacgtgttt 1920 cagacaagag ccggctgcct gatcggcgcc gagcacgtga acaacagcta cgagtgcgac 1980 atccccatcg gcgccggcat ctgcgctagc tatcagacac agaccaacag ccctagaaga 2040 gctagaagcg tggctagcca aagcatcatc gcctacacca tgagcctggg cgtggagaac 2100 agcgtggcct acagcaacaa tagcatcgcc atccctacca atttcaccat cagcgtgacc 2160 accgaaatat taccagtctc catgaccaag accagcgtgg actgcaccat gtacatctgc 2220 ggcgacagca ccgagtgcag caatctgctg ctgcagtacg gcagcttctg cacccagctg 2280 aatagagccc tgaccggcat cgccgtggag caggacaaga atacccagga ggtgttcgcc 2340 caggtgaagc agatctacaa gactccgccg atcaaggact tcggcggctt caatttcagc 2400 caaatactcc cagatccaag caagcctagc aagaggagct tcatcgagga cctgctgttc aataggtga ccctggccga cgccggcttc atcaagcagt acggcgactg cctaggtgat attgcggcaa gagacctgat ctgcgcccag aagtttaacg gtttgacagt actacctcct ctgctgaccg bite agcacaatat acgtcggcat tgctcgctgg bitecaca tcgggctgga ctttcggcgc cggagcagcg ttgcaaatcc ctttcgccat gcagatggcc 2700. you are atggcatcgg cgtgacccag aatgtgctgt you are gaagctgatc gccaatcagt tcaatagcgc catcggcaag atccaggaca gcctgagcag caccgccagc gccctgggca agctgcagga cgtggtgaat cagaatgccc aggccctgaa taccctggtg aagcagctga gcagcaattt cggcgccatc cggcgcatc tcaacgatat cctgagcaga ctggacaagg tggaggccga ggtgcaaatt gatcgtctta ttactggcag actgcagagc ctgcagacct acgtgaccca gcagctgatc agagccgccg agatcagagc cagcgccaat ctggccgcca ccaagatgag cgagtgcgtg ctgggccaga gcaagagagt ggacttctgc 3120. ggcaagggct accacctgat gagcttccct cagagcgctc cacatggcgt ggtgttcctg 3180 cacgtgacct acgtgcctgc ccaggagaag aatttcacca ccgcacccgc aatctgccac 3240 gacggcaagg cccactccc tagagaggggc gtgttcgtga gcaatggcac ccactggttc 3300 gtgacccaga gaaatttcta cgagcctcag atcatcacca ccgacaatac cttcgtgagc 3360 ggcaattgcg acgtggtgat cgggatagtc aaatactg tctacgacc tctgcagcct 3420 gagctggaca gcttcagga ggagctggac aagtacttca agaatcacac cagccctgac 3480 gtggacctcg gtgatatttc gggaatcaat gccagcgtgg tgaatatcca gaaggaaatt 3540 gatcggctca acgaagtggc caagaatctg atgagagcc tgatcgacct gcaggagctg 3600 ggcaagtacg agcagtacat caagtggcct tgtacatct ggctggggctt catcgccggc 3660 ctgatcgcca tcgtgt gaccatcatg ctgtgctgca tgactcctg ttgttcctgt 3720 ttgaaagggt gttgttcgtg tgggtcctgc tgcaagttcg acgaggacga cagcgagcct 3780 gtgctgaagg gcgtgaagct gcactacacc is 3813 <210> 9 <211> 3822 <212> DNA <213> Artificial sequence <400> 9 atgttcgtgt tcctggtgct gctgcccctg gtgagcagcc aatgcgtgaa cttcaccaac 60 agaacacagc tgcctagcgc ctacaccaac agcttcacaa gaggcgtgta ctaccccgac 120 aaggtgttca gaagcagcgt gctgcacagc acccaagacc tgttcctgcc cttcttcagc 180 aacgtgacct ggttccacgc catccacgtg agcggcacca acggcaccaa gagattcgac 240 aaccccgtgc tgcccttcaa cgacggcgtg tacttcgcta gcaccgagaa gagcaacatc 300 atcagaggct ggatcttcgg caccaccctg gacagcaaga cacagagcct gctgatcgtg 360 aacaacgcca ccaacgtggt gatcaaggtg tgcgagtttc agttctgcaa ctaccccttc 420 ctgggcgtgt actaccacaa gaacaacaag agctggatgg agagcgagtt cagagtgtac 480 agcagcgcca acaactgcac cttcgagtac gtgagccaac ccttcctgat ggacctggag 540 ggcaagcaag gcaacttcaa gaacctgagc gagttcgtgt tcaagaacat cgacggctac 600 ttcaagatct acagcaagca cacccccatc aacctggtga gagacctgcc ccaaggcttc 660 agcgccctgg agcccctggt ggacctgccc atcggcatca acatcacaag atttcagacc 720 ctgctggccc tgcacagaag ctacctgacc cccggcgaca gcagcagcgg gtggaccgcc 780 ggcgccgctg cctactacgt gggctacctg cagcctagaa ccttcctgct gaagtacaac 840 gagaacggca ccatcaccga cgccgtggac tgcgccctgg accccctgag cgagaccaag 900 tgcaccctga agagcttcac cgtggagaag ggcatctatc agacaagcaa cttcagagtg 960 cagcccaccg agagcatcgt gagattcccc aacatcacca acctgtgccc cttcggcgag 1020 gtgttcaacg ccacaagatt cgctagcgtg tacgcctgga acagaaaaag aatcagcaac 1080 tgcgtggccg actacagcgt gctgtacaac agcgctagct tcagcacctt caagtgctac 1140 ggcgtgagcc ccaccaagct gaacgacctg tgcttcacca acgtgtacgc cgacagcttc 1200 gtgatcagag gcgacgaggt gagacagatc gcccccgggc agaccggcac catcgccgac 1260 tacaactaca agctgcccga cgacttcacc ggctgcgtga tcgcctggaa ttccaacaac 1320 ctggacagca aggtgggggg caactacaac tacctgtaca gactgttcag aaagagcaac 1380 ctgaagccct tcgagagaga catcagcacc gagatctacc aagccggcag caccccctgc 1440 aacggcgtga agggcttcaa ctgctacttc cccctgcaga gctacggctt tcagcccacc 1500 tatggcgtgg gctaccagcc ttacagagtg gtggtgctga gcttcgagct gctccacgct 1560 cccgctaccg tgtgcggccc taagaagc accaatctgg tgaagaataa gtgcgtgaat 1620 ttcaatttca atggtctaac tggaacgggc gtgctgaccg agagcaataa gaagtttctt 1680 ccctttcaac aattcggcag agacatcgcc gacaccacag atgctgtaag agaccctcag 1740 accctggaga tcctggacat cactccgtgt agcttcggcg gcgtgagcgt gatcacaccg 1800 ggtaccaata ccagcaatca ggtggccgtg ctgtaccagg gcgtgaattg caccgaggtg 1860 cctgtggcca tccacgccga ccagctgact cccacttgga gggtatattc cacgggaagc 1920 aatgtgttcc agaccagagc cggctgcctg atcggcgccg agtacgtgaa taatagctac 1980 gagtgcgaca tccctatcgg cgccggcatc tgcgccagct accagaccca gaccaatagc 2040 cctagagaag ccagaagcgt ggccagccag agcatcatcg cctacaccat gagcctgggc 2100 gccgagaata gcgtggccta cagcaataat agcatcgcca tccctaccaa tttcaccatc 2160 agcgtgacca ccgaaatatt accagtctcc atgaccaaga ccagcgtgga ctgcaccatg 2220 tacatctgcg gcgacagcac cgagtgcagc aatctgctgc tgcagtacgg cagcttctgc 2280 acccagctga atagagccct gaccggcatc gccgtggagc aggacaagaa tacccaggag 2340 gtgttcgccc aggtgaagca gatctacaag actccgccga tcaaggactt cggcggcttc 2400 aatttcagcc aaatactccc agatccaagc aagcctagca agaggagctt catcgaggac 2460 ctgctgttca ataaggtgac cctggccgac gccggcttca tcaagcagta cggcgactgc 2520 ctaggtgata ttgcggcaag agacctgatc tgcgcccaga agtgtaacgg tttgacagta 2580 ctacctcctc tgctgaccga cgagatgata gcacaatata cgtcggcatt gctcgctggc 2640 acgatcacat cgggctggac ttcggcgcc ggagcagcgt tgcaaatccc ttcgccatg 2700 cagatggcct acagattcaa tggcatcggc gtgacccaga atgtgctgta cgagaatcag 2760 aagctgatcg ccaatcagtt caatagcgcc atcggcaaga tccaggacag cctgagcagc 2820 accgccagcg ccctgggcaa gctgcaggac gtggtgaatc agaatgccca ggccctgaat 2880 accctggtga agcagctgag cagcaatttc ggcgccatca gtagtgtact caacgatatc 2940 ctgagcagac tggacaaggt ggaggccgag gtgcaaattg atcgtcttat tacggcaga 3000 ctgcagagcc tgcagaccta cgtgacccag cagctgatca gagccgccga gatcagagcc 3060 agcgccaatc tggccgccat caagatgagc gagtgcgtgc tgggccagag caagagagtg 3120 gacttctgcg gcaagggcta ccacctgatg agcttccctc agagcgctcc acatggcgtg 3180 gtgttcctgc acgtgaccta cgtgcctgcc your foot atttcaccac cgcacccgca 3240 atctgccacg acggcaaggc ccacttccct agagagggcg tgttcgtgag caatggcacc 3300 cactggttcg tgacccagag aaatttctac gagcctcaga tcatcaccac cgacaatacc 3360 3420 ctgcagcctg agctggacag cttcaaggag gagctggaca agtacttcaa gaatcacacc 3480 agccctgacg tggacctcgg tgatatttcg ggaatcaatg ccagcttcgt gaatatccag 3540 aaggaaattg atcggctcaa cgaagtggcc aagaatctga atgagagcct gatcgacctg 3600 caggagctgg gcaagtacga gcagtacatc aagtggcctt ggtacatctg gctgggcttc 3660 atcgccggcc tgatcgccat cgtgatggtg accatcatgc tgtgctgcat gacctcctgt 3720 tgttcctgtt tgaaagggtg ttgttcgtgt gggtcctgct gcaagttcga cgaggacgac 3780 agcgagcctg tgctgaaggg cgtgaagctg cactacacct ga 3822 <210> 10 <211> 3816 <212> DNA <213> Artificial sequence <400> 10 atgttcgtgt tcctggtgct gctgcctctg gtgagcagcc agtgcgtgaa tctgcgcacc 60 agaacccagc tgcctcctgc ctacaccaat agcttcacca gaggagttta ttatcccgat 120 aaggtgttca gaagtagtgt attacatagt acccaggacc tgttcctacc tttcttcagt 180 aacgtgacct ggttccacgc catccacgtg agcggcacca atggcaccaa gagattcgac 240 aatcctgtgc tgcctttcaa tgacggcgtg tacttcgcca gcaccgagaa gagcaatatc 300 atcagaggct ggatcttcgg caccaccttg gattccaaga ctcagagcct gctgattgta 360 aacaacgcta caaatgtggt gatcaaggtg tgcgagttcc agttctgcaa tgaccctttc ctggatgttt attachment gaacaacaag agctggatgg agagcggtgt attcgtcg gctaataatt gcaccttcga gtacgtgagc cagccttcc tgatggacct ggagggcaag 540 cagggcaatt tcaagaatct gagagagttc gtgttcaaga atatcgacgg ctacttcaag atctacagca agcacacacc cattaatctg gtgagagacc tgcctcaggg cttcagcgtc ctggagcctc tggtggacct gcctatcggc atcaatatca ccagattcca gaccctgctg gccctgcaca gatcatatct tacaccaggc gattcgtcaa gcggttggac cgctggagct gcggcatatt acgtgggcta cctgcagcct agaccttcc tgctgaagta caatgagaat ggtacgataa ccgacgcagt tgattgtgcc ctggaccctc tgagcgagac caagtgcacc 900 ctgaagagct tcaccgtgga gaagggcatc taccagacca gcaatttcag agtgcagcct accgagagca tcgtgagatt ccctaatatc accaatctgt gccctttcgg cgaggtgttc aatgccacca gattcgccag cgtgtacgca tggaaccgca agcggataag caattgcgtg gccgactaca gcgtgctgta caatagcgcc agcttcagca ccttcaaatg ttatggtgtt 1140 tcgccaacaa agctgaatga cctgtgcttc accaatgtgt acgccgacag cttcgtgatc 1200 agaggcgacg aggtgagaca gatcgcgcca gggcagaccg gcaagatcgc cgactacaat 1260 tacaagctgc ctgacgactt caccggctgc gtgatcgcgt ggaactctaa caatctagat 1320 tcgaaagttg gaggcaatta caattaccgg tacagactgt tcagaaagag caatctgaag 1380 cctttcgaga gagacatcag caccgagatc taccaggccg gcagcaaacc gtgtaatggc 1440 gtggagggct tcaattgcta cttccctctg cagagctacg gcttccagcc taccaatggc 1500 gtgggctacc agccttacag agtggtggtg ctgagcttcg agctgctgca cgctcccgct 1560 accgtgtgcg gccctaagaa gagcaccaat ctggtgaaga ataagtgcgt gaatttcaat 1620 ttcaatggtc taactggaac gggcgtgctg accgagagca ataagaagtt tcttcccttt 1680 caacaattcg gcagagacat cgccgacacc acagatgctg taagagaccc tcagaccctg 1740 gagatcctgg acatcactcc gtgtagcttc ggcggcgtga gcgtgatcac accgggtacc 1800 aataccagca atcaggtggc cgtgctgtac cagggcgtga attgcaccga ggtgcctgtg 1860 gccatccacg ccgaccagct gactcccact tggagggtat attccacggg aagcaatgtg 1920 ttccagacca gagccggctg cctgatcggc gccgagcacg tgaataatag ctacgagtgc 1980 gacatcccta tcggcgccgg catctgcgcc agctaccaga cccagaccaa tagcagaaga 2040 agagccagaa gcgtggccag ccagagcatc atcgcctaca ccatgagcct gggcgccgag 2100 aatagcgtgg cctacagcaa taatagcatc gccatcccta ccaatttcac catcagcgtg 2160 accaccgaaa tattaccagt ctccatgacc aagaccagcg tggactgcac catgtacatc 2220 tgcggcgaca gcaccgagtg cagcaatctg ctgctgcagt acggcagctt ctgcacccag 2280 ctgaatagag ccctgaccgg catcgccgtg gagcaggaca agaataccca ggaggtgttc 2340 gcccaggtga agcagatcta caagactccg ccgatcaagg acttcggcgg cttcaatttc 2400 agccaaatac tcccagatcc aagcaagcct agcaagagga gcttcatcga ggacctgctg 2460 ttcaataagg tgaccctggc cgacgccggc ttcatcaagc agtacggcga ctgcctaggt 2520 gatattgcgg caagagacct gatctgcgcc cagaagttta acggtttgac agtactacct 2580 cctctgctga ccgacgagat gatagcacaa tatacgtcgg cattgctcgc tggcacgatc 2640 acatcgggct ggactttcgg cgccggagca gcgttgcaaa tccctttcgc catgcagatg 2700 gcctacagat tcaatggcat cggcgtgacc cagaatgtgc tgtacgagaa tcagaagctg 2760 atcgccaatc agttcaatag cgccatcggc aagatccagg acagcctgag cagcaccgcc 2820 agcgccctgg gcaagctgca gaatgtggtg aatcagaatg cccaggccct gaataccctg 2880 gtgaagcagc tgagcagcaa tttcggcgcc atcagtagtg tactcaacga tatcctgagc 2940 agactggaca aggtggaggc cgaggtgcaa attgatcgtc ttattactgg cagactgcag 3000 agcctgcaga cctacgtgac ccagcagctg atcagagccg ccgagatcag agccagcgcc 3060 aatctggccg ccaccaagat gagcgagtgc gtgctgggcc agagcaagag agtggacttc 3120 tgcggcaagg gctaccacct gatgagcttc cctcagagcg ctccacatgg cgtggtgttc 3180 ctgcacgtga cctacgtgcc tgcccaggag aagaatttca ccaccgcacc cgcaatctgc 3240 cacgacggca aggcccactt ccctagagag ggcgtgttcg tgagcaatg cacccactgg 3300 ttcgtgaccc agagaaattt ctacgagcct cagatcatca ccaccgacaa taccttcgtg 3360 agcggcaatt gcgacgtggt gatcgggata gtcaataata ctgtctacga ccctctgcag 3420 cctgagctgg acagcttcaa ggaggagctg gacaagtact tcaagaatca caccagccct 3480 gacgtggacc tcggtgatat ttcgggaatc aatgccagcg tggtgaatat ccagaaggaa 3540 attgatcggc tcaacgaagt ggccaagaat ctgaatgaga gcctgatcga cctgcaggag 3600 ctgggcaagt acgagcagta catcaagtgg ccttggtaca tctggctggg cttcatcgcc 3660 ggcctgatcg ccatcgtgat ggtgaccatc atgctgtgct gcatgacctc ctgttgttcc 3720 tgtttgaaag ggtgttgttc gtgtgggtcc tgctgcaagt tcgacgagga cgacagcgag 3780 cctgtgctga agggcgtgaa gctgcactac acctga 3816 <210> 11 <211> 1281 <212> PRT <213> Artificial sequence <400> 11 Met Phe Val Phe Leu Val Leu Leu Pro Leu Val Ser Ser Gln Cys Val 1 5 10 15 Asn Leu Thr Thr Arg Thr Gln Leu Pro Pro Ala Tyr Thr Asn Ser Phe 20 25 30 Thr Arg Gly Val Tyr Tyr Pro Asp Lys Val Phe Arg Ser Ser Val Leu 35 40 45 His Ser Thr Gln Asp Leu Phe Leu Pro Phe Phe Ser Asn Val Thr Trp 50 55 60 Phe His Ala Ile His Val Ser Gly Thr Asn Gly Thr Lys Arg Phe Asp 65 70 75 80 Asn Pro Val Leu Pro Phe Asn Asp Gly Val Tyr Phe Ala Ser Thr Glu 85 90 95 Lys Ser Asn Ile Ile Arg Gly Trp Ile Phe Gly Thr Thr Leu Asp Ser 100 105 110 Lys Thr Gln Ser Leu Leu Ile Val Asn Asn Ala Thr Asn Val Val Ile 115 120 125 Lys Val Cys Glu Phe Gln Phe Cys Asn Asp Pro Phe Leu Gly Val Tyr 130 135 140 Tyr His Lys Asn Asn Lys Ser Trp Met Glu Ser Glu Phe Arg Val Tyr 145 150 155 160 Ser Ser Ala Asn Asn Cys Thr Phe Glu Tyr Val Ser Gln Pro Phe Leu 165 170 175 Met Asp Leu Glu Gly Lys Gln Gly Asn Phe Lys Asn Leu Arg Glu Phe 180 185 190 Val Phe Lys Asn Ile Asp Gly Tyr Phe Lys Ile Tyr Ser Lys His Thr 195 200 205 Pro Ile Asn Leu Val Arg Asp Leu Pro Gln Gly Phe Ser Ala Leu Glu 210 215 220 Pro Leu Val Asp Leu Pro Ile Gly Ile Asn Ile Thr Arg Phe Gln Thr 225 230 235 240 Leu Leu Ala Leu His Arg Ser Tyr Leu Thr Pro Gly Asp Ser Ser Ser 245 250 255 Gly Trp Thr Ala Gly Ala Ala Ala Tyr Tyr Val Gly Tyr Leu Gln Pro 260 265 270 Arg Thr Phe Leu Leu Lys Tyr Asn Glu Asn Gly Thr Ile Thr Asp Ala 275 280 285 Val Asp Cys Ala Leu Asp Pro Leu Ser Glu Thr Lys Cys Thr Leu Lys 290 295 300 Ser Phe Thr Val Glu Lys Gly Ile Tyr Gln Thr Ser Asn Phe Arg Val 305 310 315 320 Gln Pro Thr Glu Ser Ile Val Arg Phe Pro Asn Ile Thr Asn Leu Cys 325 330 335 Pro Phe Gly Glu Val Phe Asn Ala Thr Arg Phe Ala Ser Val Tyr Ala 340 345 350 Trp Asn Arg Lys Arg Ile Asn Cys Val Ala Asp Tyr Ser Val Leu 355 360 365 Tyr Asn Ser Ala Ser Phe Ser Thr Phe Lys Cys Tyr Gly Val Ser Pro 370 375 380 Thr Lys Leu Asn Asp Leu Cys Phe Thr Asn Val Tyr Ala Asp Ser Phe 385 390 395 400 Val Ile Arg Gly Asp Glu Val Arg Gln Ile Ala Pro Gly Gln Thr Gly 405 410 415 Lys Ile Ala Asp Tyr Asn Tyr Lys Leu Pro Asp Asp Phe Thr Gly Cys 420 425 430 Val Ile Ala Trp Asn Ser Asn Asn Leu Asp Ser Lys Val Gly Gly Asn 435 440 445 Tyr Asn Tyr Leu Tyr Arg Leu Phe Arg Lys Ser Asn Leu Lys Pro Phe 450 455 460 Glu Arg Asp Contains Thr Glu and Tyr Gln Only Gly Serves Thr Pro Cys 465 470 475 480 Asn Gly Val Glu Gly Phe Asn Cys Tyr Phe Pro Leu Gln Ser Tyr Gly 485 490 495 Phe Gln Pro Thr Asn Gly Val Gly Tyr Gln Pro Tyr Arg Val Val Val 500 505 510 Leu Ser Phe Glu Leu Leu His Ala Pro Ala Thr Val Cys Gly Pro Lys 515 520 525 Lys Ser Thr Asn Leu Val Lys Asn Lys Cys Val Asn Phe Asn Phe Asn 530 535 540 Gly Leu Thr Gly Thr Gly Val Leu Thr Glu Ser Asn Lys Lys Phe Leu 545 550 555 560 Pro Phe Gln Gln Phe Gly Arg Asp Ile Ala Asp Thr Thr Asp Ala Val 565 570 575 Arg Asp Pro Gln Thr Leu Glu Ile Leu Asp Ile Thr Pro Cys Ser Phe 580 585 590 Gly Gly Val Ser Val Ile Thr Pro Gly Thr Asn Thr Ser Asn Gln Val 595 600 605 Ala Val Leu Tyr Gln Asp Val Asn Cys Thr Glu Val Pro Val Ala Ile 610 615 620 His Ala Asp Gln Leu Thr Pro Thr Trp Arg Val Tyr Ser Thr Gly Ser 625 630 635 640 Asn Val Phe Gln Thr Arg Ala Gly Cys Leu Ile Gly Ala Glu His Val 645 650 655 Asn Asn Ser Tyr Glu Cys Asp Ile Pro Ile Gly Ala Gly Ile Cys Ala 660 665 670 Ser Tyr Gln Thr Gln Thr Asn Ser Pro Arg Arg Ala Arg Ser Val Ala 675 680 685 Ser Gln Ser Ile Ile Ala Tyr Thr Met Ser Leu Gly Ala Glu Asn Ser 690 695 700 Val Ala Tyr Ser Asn Asn Ser Ile Ala Ile Pro Thr Asn Phe Thr Ile 705 710 715 720 Ser Val Thr Thr Glu Ile Leu Pro Val Ser Met Thr Lys Thr Ser Val 725 730 735 Asp Cys Thr Met Tyr Ile Cys Gly Asp Ser Thr Glu Cys Ser Asn Leu 740 745 750 Leu Leu Gln Tyr Gly Ser Phe Cys Thr Gln Leu Asn Arg Ala Leu Thr 755 760 765 Gly Ile Ala Val Glu Gln Asp Lys Asn Thr Gln Glu Val Phe Ala Gln 770 775 780 Val Lys Gln Ile Tyr Lys Thr Pro Pro Ile Lys Asp Phe Gly Gly Phe 785 790 795 800 Asn Phe Ser Gln Ile Leu Pro Asp Pro Ser Lys Pro Ser Lys Arg Ser 805 810 815 Phe Ile Glu Asp Leu Leu Phe Asn Lys Val Thr Leu Ala Asp Ala Gly 820 825 830 Phe Ile Lys Gln Tyr Gly Asp Cys Leu Gly Asp Ile Ala Ala Arg Asp 835 840 845 Leu Ile Cys Ala Gln Lys Phe Asn Gly Leu Thr Val Leu Pro Pro Leu 850 855 860 Leu Thr Asp Glu Met Ile Ala Gln Tyr Thr Ser Ala Leu Leu Ala Gly 865 870 875 880 Thr Ile Thr Ser Gly Trp Thr Phe Gly Ala Gly Ala Ala Leu Gln Ile 885 890 895 Pro Phe Ala Met Gln Met Ala Tyr Arg Phe Asn Gly Ile Gly Val Thr 900 905 910 Gln Asn Val Leu Tyr Glu Asn Gln Lys Leu Ile Ala Asn Gln Phe Asn 915 920 925 Ser Ala Ile Gly Lys Ile Gln Asp Ser Leu Ser Ser Thr Ala Ser Ala 930 935 940 Leu Gly Lys Leu Gln Asp Val Val Asn Gln Asn Ala Gln Ala Leu Asn 945 950 955 960 Thr Leu Val Lys Gln Leu Ser Ser Asn Phe Gly Ala Ile Ser Ser Val 965 970 975 Leu Asn Asp Ile Leu Ser Arg Leu Asp Lys Val Glu Ala Glu Val Gln 980 985 990 Ile Asp Arg Leu Ile Thr Gly Arg Leu Gln Ser Leu Gln Thr Tyr Val 995 1000 1005 Thr Gln Gln Leu Ile Arg Ala Ala Glu Ile Arg Ala Ser Ala Asn 1010 1015 1020 Leu Ala Ala Thr Lys Met Ser Glu Cys Val Leu Gly Gln Ser Lys 1025 1030 1035 Arg Val Asp Phe Cys Gly Lys Gly Tyr His Leu Met Ser Phe Pro 1040 1045 1050 Gln Ser Ala Pro His Gly Val Val Phe Leu His Val Thr Tyr Val 1055 1060 1065 Pro Ala Gln Glu Lys Asn Phe Thr Thr Ala Pro Ala Ile Cys His 1070 1075 1080 Asp Gly Lys Ala His Phe Pro Arg Glu Gly Val Phe Val Ser Asn 1085 1090 1095 Gly Thr His Trp Phe Val Thr Gln Arg Asn Phe Tyr Glu Pro Gln 1100 1105 1110 Ile Ile Thr Thr Asp Asn Thr Phe Val Ser Gly Asn Cys Asp Val 1115 1120 1125 Val Ile Gly Ile Val Asn Asn Thr Val Tyr Asp Pro Leu Gln Pro 1130 1135 1140 Glu Leu Asp Ser Phe Lys Glu Glu Leu Asp Lys Tyr Phe Lys Asn 1145 1150 1155 His Thr Ser Pro Asp Val Asp Leu Gly Asp Ile Ser Gly Ile Asn 1160 1165 1170 Ala Ser Val Val Asn Ile Gln Lys Glu Ile Asp Arg Leu Asn Glu 1175 1180 1185 Val Ala Lys Asn Leu Asn Glu Ser Leu Ile Asp Leu Gln Glu Leu 1190 1195 1200 Gly Lys Tyr Glu Gln Tyr Ile Lys Trp Pro Trp Tyr Ile Trp Leu 1205 1210 1215 Gly Phe Ile Ala Gly Leu Ile Ala Ile Val Met Val Thr Ile Met 1220 1225 1230 Leu Cys Cys Met Thr Ser Cys Cys Ser Cys Leu Lys Gly Cys Cys 1235 1240 1245 Ser Cys Gly Ser Cys Cys Lys Phe Asp Glu Asp Asp Ser Glu Pro 1250 1255 1260 Val Leu Lys Gly Val Lys Leu His Tyr Thr Trp Ser His Pro Gln 1265 1270 1275 Phe Glu Lys 1280
Claims
1. A SARS-CoV-2 S protein variant, the amino acid sequence of which is shown as positions 1-1205 from the N-terminal end in SEQ ID NO:
2.
2. A fusion protein obtained by connecting a tag to the N-terminal end or / and the C-terminal end of the SARS-CoV-2 S protein variant of claim 1.
3. A DNA molecule encoding the SARS-CoV-2 S protein variant of claim 1 or the fusion protein of claim 2.
4. The DNA molecule of claim 3, wherein: The nucleotide sequence of the DNA molecule is shown as positions 41-3781 from the 5’-terminal end in SEQ ID NO: 1 or SEQ ID NO:
1.
5. A recombinant plasmid obtained by inserting the DNA molecule of claim 3 or 4 into a multiple cloning site of an expression vector.
6. A recombinant protein obtained by transfecting mammalian cells with the recombinant plasmid of claim 5 and culturing the transfected cells.
7. The recombinant protein according to claim 6, characterized in that: The cells are 293F cells.
8. A composition for preventing disease caused by SARS-CoV-2 infection, characterized by: The composition comprises the recombinant protein of claim 6 or 7; The SARS-CoV-2 is SARS-CoV-2 WA1 / 2020 and / or a mutant strain of SARS-CoV-2 WA1 / 2020. The mutant strain of SARS-CoV-2 WA1 / 2020 is SARS-CoV-2 Alpha mutant strain, SARS-CoV-2 Beta mutant strain, SARS-CoV-2 Gamma mutant strain or SARS-CoV-2 Delta mutant strain.
9. Use of the recombinant protein of claim 6 or 7 or the composition of claim 8 in the preparation of a medicament for preventing SARS-CoV-2 infection, for neutralizing SARS-CoV-2 or for treating a disease caused by SARS-CoV-2 infection; The SARS-CoV-2 is SARS-CoV-2 WA1 / 2020 and / or a mutant strain of SARS-CoV-2 WA1 / 2020. The mutant strain of SARS-CoV-2 WA1 / 2020 is SARS-CoV-2 Alpha mutant strain, SARS-CoV-2 Beta mutant strain, SARS-CoV-2 Gamma mutant strain or SARS-CoV-2 Delta mutant strain.
10. A vaccine comprising the SARS-CoV-2 S protein variant of claim 1, the fusion protein of claim 2, the DNA molecule of claim 3 or 4, the recombinant plasmid of claim 5, the recombinant protein of claim 6 or 7 or the composition of claim 8.
11. The vaccine of claim 10, characterized in that: The vaccine is an mRNA vaccine, an adenovirus vector vaccine or a recombinant protein vaccine.
12. Use of the SARS-CoV-2 S protein variant of claim 1, the fusion protein of claim 2, the DNA molecule of claim 3 or 4, the recombinant plasmid of claim 5, the recombinant protein of claim 6 or 7 or the composition of claim 8 in the preparation of a vaccine for preventing SARS-CoV-2 infection. SARS-CoV-2 is SARS-CoV-2 WA1 / 2020 and / or a mutant strain of SARS-CoV-2 WA1 / 2020; the mutant strain of SARS-CoV-2 WA1 / 2020 is SARS-CoV-2 Alpha mutant strain, SARS-CoV-2 Beta mutant strain, SARS-CoV-2 Gamma mutant strain or SARS-CoV-2 Delta mutant strain.
13. Use according to claim 12, characterized in that: the vaccine is an mRNA vaccine, an adenovirus vector vaccine or a recombinant protein vaccine.
Citation Information
Patent Citations
SARS-CoV-2 vaccine and preparation method thereof
CN111217917A
SARS-CoV-2 vaccine and application thereof
CN111671890A