Protein container, polynucleotide, vector, expression cassette, cell, method for producing a container, method for pathogen recognition or disease diagnosis, use of a container, and diagnostic kit
By designing multivalent container proteins, the problem of GFP molecules in the prior art being difficult to express at high levels in the cellular system and tolerate the insertion of multiple exogenous polyamino acid sequences is solved, and efficient fluorescence expression and multifunctional application are achieved.
Patent Information
- Application Number
- CN202080058664.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-27
- Filing Date
- 2020-08-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-08-27
AI Technical Summary
It is difficult to develop a green fluorescent protein (GFP) molecule that can express at a high level in a cellular system, tolerate insertion of multiple exogenous polyamino acid sequences and maintain autofluorescence.
By designing a multivalent container protein, multiple exogenous polyamino acid sequences can be displayed simultaneously at more than 4 different sites, while maintaining sufficient fluorescence intensity and high level of expression.
GFP molecules expressed at high levels in the cellular system are realized, able to tolerate the insertion of multiple exogenous polyamino acid sequences while maintaining autofluorescence, and are suitable for research, diagnostic and vaccine composition development.
Smart Images

Figure CN114258399B_ABST
Abstract
Description
Technical Field
[0001] This invention falls within the fields of chemistry, pharmacy, medicine, and biotechnology, and more specifically, within the field of formulations for biomedical purposes. This invention relates to protein receptacles capable of simultaneously receiving multiple exogenous multi-amino acid sequences for expression in various systems and for different uses. This invention relates to polynucleotides capable of producing the aforementioned protein receptacles. This invention also relates to vectors and expression cassettes comprising one or more of the aforementioned polynucleotides. This invention further relates to cells comprising the aforementioned vectors or expression cassettes. This invention further relates to methods for producing said protein receptacles and for pathogen recognition or in vitro disease diagnosis. This invention further relates to the use of said protein receptacles and kits comprising said protein receptacles for diagnostic purposes or as vaccine compositions. Background Technology
[0002] This document contains multiple references, which are shown in parentheses. The information disclosed in these references is included herein to better describe the prior art to which this invention pertains.
[0003] The green fluorescent protein (GFP) produced by the cnidarian jellyfish (Aequorea victoria) emits high fluorescence in the green region of the visible spectrum (Prasher et al., Gene 15, 111(2):229-33, 1992(1]), and is classically used as a marker for gene expression and localization, and is therefore considered a reporter protein. Several uses of the protein were described after the initial observation that the GFP protein could emit fluorescence.
[0004] The use of GFP as a reporter protein has been patented for various applications and in different systems. Patent application US2018298058 uses GFP in a protein production and purification system. Patent application CN108303539 uses GFP as an internal control for cancer detection assays; application CN108220313 proposes a high-throughput GFP fusion and expression method. Application CN108192904 seeks protection for a GFP fusion protein that can insert itself into a biological membrane; application CN107703219 has highlighted the use of GFP in metabolomics studies of mesenchymal cells. Furthermore, patent application US2018016310 proposes a variant of "superfolded" fused GFP. Several patents exist, such as those with mutated GFP to increase fluorescence expression or modify fluorescence wavelength peaks: US6054321, US6096865, US6027881, and US6025485. Therefore, it is clear that since the description of the GFP protein, it has been used and patented for various purposes as a reporter protein.
[0005] Reporter molecules are commonly used in biological systems to monitor gene expression. GFP proteins offer significant innovation by eliminating the need for any substrates or cofactors—that is, unlike most other reporter proteins, they do not require the addition of any additional reagents for visualization. Another advantage of GFP is its ability to exhibit autofluorescence, eliminating the need for fluorescent labels. Fluorescent labels, for various reasons, lack the appropriate sensitivity or specificity for their intended use.
[0006] Given its ability to produce detectable green fluorescence on its own, GFP has been widely used in the study of gene expression and protein localization and is considered one of the most promising reporter proteins in the literature.
[0007] In its use as a reporter protein, the gene encoding GFP can be used in the production of fusion proteins, i.e., the specific gene of interest can be fused to the gene encoding GFP. The fusion gene cassette can be inserted into a living system, allowing for the monitoring of the expression of the fusion gene and the intracellular localization of the protein of interest (Santos-Beneit & Errington, Archives Microbiology, 199(6):875-880, 2017; Belardinelli & Jackson, Tuberculosis (Edinb), 105:13-17, 2017; Wakabayashi et al., International Journal Food Microbiology, 19, 291:144-150, 2018; Cai et al., Viruses, 20; 10(11), 2018).
[0008] In both yeast and mammalian cell systems, GFP proteins are used as frameworks for peptide display or even peptide libraries (Kamb et al., Proc. Natl. Acad Sci. USA, 95:7508-7513, 1998; WO 2004005322). As a framework protein for displaying random peptides, GFP proteins can be used to define the properties of peptide libraries.
[0009] Advances in the development of novel GFP variants seek to improve the properties of the protein to produce new reagents for a wide range of research purposes. Novel forms of GFP have been developed through mutation, containing DNA sequences optimized for increased production in human cell systems; these are known as humanized GFP proteins (Cormack et al., Gene 173, 33-38, 1996; Haas et al., Current Biology 6, 315-324, 1996; Yang et al., Nucleic Acids Research 24, 4592-4593, 1996).
[0010] One of these forms describes the enhanced green fluorescent protein (eGFP) (Heim & Cubitt, Nature, 373, pp. 663-664, 1995).
[0011] GFP, encoded by the gfp10 gene and originating from the cnidarian jellyfish *Echinochloa victoriae*, is a 238-amino acid protein. This protein has the ability to absorb blue light (with a main excitation peak at 395 nm) and emit green light (with a main emission peak at 509 nm) from a chromophore at its center (Morin & Hastings, *Journal Cell Physiology*, 77(3):313-8, 1971; Prasher et al., *Gene 15*, 111(2):229-33, 1992). The chromophore is composed of a hexapeptide, starting at amino acid 64 and derived from the primary amino acid sequence through cyclization and oxidation of serine, tyrosine, and glycine (positions 65, 66, and 67) (Shimomura, 104(2), 1979; Cody et al., *Biochemistry*, 32(5):1212-8, 1993). The light emitted by GFP is independent of the cell biological species expressing it and does not require any kind of substrate, cofactor, or other gene product from *Bioluminescent jellyfish Victoria* (Chalfie et al., Science, 263(5148):802-5, 1994). This property of GFP allows its fluorescence to be detected in living cells other than *Bioluminescent jellyfish Victoria*, because it can be processed in the cell's protein expression system (Ormo et al., Science 273:1392-1395, 1996; Yang et al., Nature Biotech 14:1246-1251, 1996).
[0012] The basic structure of GFP consists of 11 antiparallel folded β-chains that intertwine to form a tertiary structure in the shape of β-barrels. Each chain connects to the next chain via a loop domain that projects onto the upper and lower surfaces of the barrel, interacting with the environment. By convention, each chain and loop can be identified by numbering for better protein description.
[0013] Targeted mutagenesis experiments revealed that certain biochemical properties of the GFP protein are caused by this barrel structure. Therefore, changes in the amino acid composition of the primary structure are responsible for accelerating protein folding, reducing the aggregation of translation products, and increasing protein stability in solution.
[0014] Specific loops migrate into the cavity of the protein barrel, forming α-helices, which are responsible for the protein's fluorescent properties (Crone et al., GFP-Based Biosensors, InTech, 2013). Certain structural modifications can interfere with the ability to emit fluorescence. Mutations in certain amino acids can alter both the intensity and wavelength of fluorescence emission, thus changing the color of the emitted light. Mutations in Tyr66, an internal residue involved in the chromophore, can produce a large number of fluorescent protein variants with altered structures of the chromophore or its surrounding environment. These alterations interfere with the absorption and emission of light at different wavelengths, resulting in a wide range of different emission colors (Heim & Tsien, Current Biology, 6(2):178-82, 1996).
[0015] Changes in pH can also interfere with fluorescence intensity. At physiological pH, GFP exhibits maximum absorption at 395 nm, while absorbing less light at 475 nm. However, increasing the pH to approximately 12.0 causes maximum absorption to occur in the 475 nm range, while GFP exhibits reduced absorption at 395 nm (Ward et al., Photochemistry and Photobiology, 35(6):803-808, 1982).
[0016] The compact structure of the core protein makes GFP highly stable even under adverse conditions such as protease treatment, making it very useful as a reporter protein.
[0017] Different forms of GFP already exist, but there is always an ongoing pursuit of improving proteins by adding new functions or removing certain restrictions. eGFP itself exhibits an enhanced form, giving proteins greater flexibility in the face of modifications to the F64L and S65T amino acids (Heim et al., Nature 373:663-664, 1995; Li et al., Journal Biology Chemistry, 272(45):28545-9, 1997). Even when its protein sequence contains heterologous sequences, this enhancement allows GFP to achieve both its desired three-dimensional shape and its ability to express fluorescence (Pedelacq et al., Nat Biotechnology, 24(1):79-88, 2006).
[0018] GFAbs are protein modifications that accept exogenous sequences in two loop domains. Their development requires several rounds of directed evolution to select for three mutant protein clones that support the insertion of exogenous peptides into the two proximal regions (i.e., Glu-172-Asp-173 and Asp-102-Asp-103). The authors demonstrated using unmutated proteins that the insertion of both exogenous peptides prevents GFP fluorescence production and protein expression on the surface of yeast cells. After a series of mutations and selections, simultaneous insertion into only the two regions is possible, but still results in a significant loss of the intrinsic activity of GFP, sometimes making the production or expression of its insert impossible (Pavoor et al., PNAS 106(29):11895-11900, 2009).
[0019] Mutations involving circular N- and C-terminal arrangements also indicate that eGFP can be manipulated within the coding sequence without affecting the structural aspects of the core protein (Topell et al., FEBS Letters 457(2):283-289, 1999). However, analysis of 20 circular protein variants revealed low tolerance to the insertion of new ends, and in most cases, they lost the ability to form chromophores. This fact suggests that manipulation of protein sequences can strongly interfere with their properties or even their cellular expression.
[0020] Several attempts have been made to simultaneously insert multiple epitopes into the loop region of GFP, with the aim of enabling the protein to be used for target-specific binding responses. However, given the structural sensitivity of GFP and its chromophore, all these efforts have shown limited success.
[0021] Other mutant GFP proteins exhibit modified forms that emit other types of fluorescence spectra. For example, Heim et al. (Proc Natl Acad Sci USA, 91(26):12501-4, 1994) described a mutant protein that emits blue fluorescence by replacing tyrosine with histidine at amino acid 66. Subsequently, Heim et al. (Nature, 373(6516):663-4, 1995) also described a mutant GFP protein that, by replacing serine with threonine at amino acid 65, has a spectrum similar to that obtained from the marine animal *Renilla reniformis*, with an extinction coefficient per monomer that is more than 10 times greater than the wavelength peak of natural GFP from the genus *Aequorea*. Other patent literature describes mutant GFP proteins that exhibit emission spectra other than green, such as blue and red (US 5625048, WO 2004005322).
[0022] In addition, other mutant GFP proteins possess optimized excitation spectra, making them particularly useful in certain argon laser flow cytometry (FACS) devices (US 5804387). Descriptions also exist of mutant GFP proteins modified for better expression in plant cell systems (WO1996027675). Patent document US5968750 proposes humanized GFP adapted for expression in mammalian cells, including humans. Humanized GFP integrates a preferred codon for reading into the human cell gene expression system.
[0023] In the prior art, as can be clearly seen from the patent documents listed above, GFP can contain genes at its 5' or 3' end without interfering with expression, three-dimensional tangling, and fluorescence generation. Furthermore, GFP has been used as a vector for in vivo peptide display or even peptide libraries. In the case of peptide libraries, GFP can assist in the display of random peptides, thereby helping to define the characteristics of the peptide library (Kamb et al., Proc. Natl. Acad. Sci. USA, 95:7508-7513, 1998; WO2004005322).
[0024] For example, Abedi et al. (1998, Nucleic Acids Res. 26:623-300) inserted peptides into the exposed loop region of the GFP protein from *E. victoria*, demonstrating that the GFP molecule retained autofluorescence when expressed in yeast and *Escherichia coli*. The authors further elucidated that the fluorescence of the GFP frame could be used to monitor peptide diversity and the presence or expression of a specific peptide in a given cell. However, the fluorescence rate of the GFP frame molecule is relatively low compared to native GFP. Kamb and Abedi (US 6025485) prepared a library of GFP arrays from enhanced green fluorescent protein (eGFP) to enhance fluorescence intensity.
[0025] Furthermore, Peele et al. (Chem. & Bio. 8:521-534, 2001) used eGFP as a framework to experiment with peptide libraries with different structural biases in mammalian cells. Anderson et al. further enhanced fluorescence intensity by inserting peptides into GFP loops with tetraglycine ligands (US20010003650). Happe et al. described humanized GFP that could be expressed in large quantities in mammalian cell systems, tolerated peptide insertion, and maintained autofluorescence (WO 2004005322).
[0026] However, there is still a need in this field for GFP molecular frameworks that not only exhibit appropriate fluorescence intensity but are also expressed at high levels in cellular systems.
[0027] In GFP molecules, tolerance to peptide display while retaining autofluorescence is variable. Therefore, there is a need in the current technology to develop GFP molecules that can be expressed at high levels, tolerate insertion, and retain autofluorescence.
[0028] In addition to its ability to support gene expression at the ends, GFP allows epitopes to be inserted into the surface loops of molecules exposed to the medium. Numerous attempts have been made to simultaneously insert multiple peptides into the loop regions of GFP, potentially enabling proteins to be used for target-specific binding reactions. However, given the structural sensitivity of GFP and its chromophores, all these efforts have had limited success.
[0029] Pavoor and colleagues worked to develop a protein modified to accept exogenous sequences in two loop domains. This development required several rounds of directed evolution to select three mutant protein clones that supported the insertion of exogenous peptides in the two proximal regions (i.e., Glu-172-Asp-173 and Asp-102-Asp-103). The authors demonstrated using unmutated proteins that the insertion of both exogenous peptides prevented GFP fluorescence production and protein expression on the surface of yeast cells. After a series of mutations and selections, simultaneous insertion in only the two regions became possible, but still resulted in a significant loss of the intrinsic activity of GFP, sometimes rendering the production or expression of its insert impossible (Pavoor et al., PNAS 106(29):11895-11900, 2009).
[0030] Abedi et al. (Nucleic Acids Research 26(2):623-30, 1998) proposed 10 protein sites in the loop region, 8 of which are between β-sheets, for peptide expression. Chimeric proteins can be used in experiments requiring intracellular expression, so fluorescence uninterruptedness will be a limiting factor. In this study, only three chimeric proteins (with insertion sites at amino acids 157-158, 172-173, and 194-195) showed fluorescence (diminished to one-quarter of the original); and only two insertion sites (studied separately) could have peptides without losing fluorescence. The authors further concluded that "it is intriguing how GFP is so sensitive to structural perturbations even in β-sheets."
[0031] Li et al. (Photochemistry and Photobiology, 84(1):111-9, 2008) proposed a study of chimeric proteins (red fluorescent protein-RFP), pointing out that in this protein, six genetically distinct sites are located in three different loops, where sequences of five residues can be inserted without interfering with the protein's ability to fluoresce. However, the authors did not elucidate the simultaneous use of these sites to insert different peptides.
[0032] Patent application WO02090535 proposes fluorescent GFP for non-simultaneous insertion of peptides into five different rings of a protein. The patent application, in its descriptive report, indicates the possibility of simultaneously inserting peptides into more than one ring of a protein, increasing library complexity and allowing protein display on the same surface. However, the patent document does not demonstrate this possibility, as it only proposes a test of peptide insertion one at a time in the five different protein rings. Notably, the document further emphasizes that rings 1 and 5 themselves do not present good insertion sites because peptide insertion at these sites prevents protein expression. Furthermore, other patent documents propose GFP variants for peptide expression in protein rings; however, these studies have not demonstrated the feasibility of simultaneously expressing more than four peptides in different insertion sites within a GFP protein ring without losing any of their essential characteristics (WO02090535, US2003224412, WO200134824). Summary of the Invention
[0033] To address the aforementioned problems, the present invention offers significant advantages, as container proteins can express a large number of different multi-amino acid sequences, characterized by multivalent container proteins, thus expanding their use for vaccine composition purposes, as internal controls for research and technology development, or for disease diagnosis. There remains a practical need in the prior art to develop container proteins that not only exhibit sufficient fluorescence intensity but are also expressed in large quantities in production cell systems, and further, are tolerant to the accompanying display of multiple exogenous multi-amino acid sequences while still exhibiting detectable autofluorescence.
[0034] In one aspect, the present invention relates to a protein container capable of simultaneously displaying multiple exogenous multi-amino acid sequences at more than four different sites on a container protein.
[0035] In another respect, the present invention relates to polynucleotides capable of producing the aforementioned protein containers.
[0036] In another respect, the present invention relates to a carrier comprising the aforementioned polynucleotides.
[0037] In another aspect, the present invention relates to expression cassettes comprising the aforementioned polynucleotides.
[0038] In another aspect, the present invention relates to methods for producing said protein containers and for pathogen identification or in vitro disease diagnosis.
[0039] In another respect, the present invention relates to the use of the protein container for diagnostic purposes or as a vaccine composition.
[0040] In another aspect, the present invention relates to a kit containing the protein container for diagnostic purposes or as a vaccine composition. Attached Figure Description
[0041] Figure 1 – Purification of PlatCruzi protein by affinity chromatography. (A) Using Elution profiles of container proteins on a nickel column in a liquid chromatography system. (B) Analysis by polyacrylamide gel electrophoresis (SDSPAGE) of collected elution buffers 13–26. Elution was performed in ascending order using buffer B. PM – molecular weight marker.
[0042] Figure 2 – The reactivity of serum obtained by ELISA using PlatCruzi with the international standard biological reference for Trypanosoma cruzi provided by the WHO. (A) Pool of patient serum that recognizes the TcI strain, referred to as IS 09 / 188. (B) Pool of patient serum that recognizes the TcII strain, referred to as IS 09 / 186.
[0043] Figure 3 – Determination of antibody titers from serum of patients with chronic Chagas disease using PlatCruzi container protein as an antigen. LACENS provides serum at a concentration of 500 ng / well and serum dilutions of 1:50–1:1000 for ELISAS.
[0044] Figure 4 – The expression of PlatCruzi antigen by ELISA in serum from patients with various diseases. PlatCruzi antigen was used at a concentration of 500 ng / well and the serum was diluted 1:250.
[0045] Figure 5 – Detection of rabies virus-specific epitopes in rabbit anti-RxRabies2 serum. Rabbit antibodies immunized with RxRabies2 were purified by RxRabies2 affinity chromatography and used as primary antibodies. Immunoblotting was used to detect the crude extract (*; column 2) or semi-purified antibodies at two different concentrations (1x and 0.5x). RxRabies2 in extracts (columns 4 and 6, respectively). Negative controls: Rx container protein (column 3) and PlatCruzi (at two concentrations: 1x and 0.5x, in columns 5 and 7, respectively).
[0046] Figure 6 – Analysis of RxHoIgG3 protein by polyacrylamide gel electrophoresis (SDS PAGE). A, Soluble extract of E. coli that does not produce RxHoIgG3. B, Aqueous insoluble fraction of RxHoIgG3-producing bacteria. C, Soluble fraction of RxHoIgG3-producing bacteria. Arrows indicate the location of the RxHoIgG3 protein.
[0047] Figure 7 – Detection of IgM anti-RxOro antibodies by ELISA. C-: Negative control (serum from patients without oropeceutical virus infection); C+: Standard positive serum for oropeceutical virus infection. Patient: Suspected case of oropeceutical virus infection. Oro+: Positive for oropeceutical virus infection (detection of IgM anti-RxOro antibodies). Oro- (negative control): No IgM reaction to RxOro. Protein concentration: 0.288 μg / μL. Cutoff: 0.0613.
[0048] Figure 8 – This represents polyacrylamide gel electrophoresis (SDS-PAGE) of the production of PlatCruzi, RxMayaro_IgG, and RxMayaro_IgM proteins. Columns 1 through 8 represent: 1) molecular weight; 2) total bacterial extract without recombinant protein induction; 3) total bacterial extract after PlatCruzi induction; 4) total bacterial extract after RxMayaro_IgG induction; 5) total bacterial extract after RxMayaro_IgM induction; 6) insoluble bacterial protein after PlatCruzi induction; 7) insoluble bacterial protein after RxMayaro_IgG induction; 8) insoluble bacterial protein after RxMayaro_IgM induction. Arrows indicate the bands representing PlatCruzi (columns 3 and 6), RxMayaro_IgG (columns 4 and 7), and RxMayaro_IgM (columns 5 and 8).
[0049] Figure 9 – Reactivity of sera from Mayaro virus-positive patients (S MAY) and healthy individuals (SN) was measured by ELISA using the RxMayaro_IgG protein. Colorimetric analysis was performed using anti-IgG immunoglobulin conjugated to alkaline phosphatase (cutoff = 0.0210).
[0050] Figure 10 – Reactivity of serum from individuals considered healthy (SN) and positive for Mayaro virus (S MAY) was measured by ELISA using the RxMayaro_IgM protein. Colorimetric analysis was performed using anti-IgM immunoglobulin conjugated to alkaline phosphatase (cutoff = 0.0547).
[0051] Figure 11 – Polyacrylamide gel electrophoresis (SDS-PAGE) showing the yields of insoluble (I) and soluble (S) proteins from PlatCruzi, TxCruzi, RxPtx, TxNeuza, and RxYFIgG. Columns 1 through 10 include: 1) insoluble proteins from bacteria that induce PlatCruzi production; 2) soluble proteins from bacteria that induce PlatCruzi production; 3) insoluble proteins from bacteria that induce TxCruzi production; 4) soluble proteins from bacteria that induce TxCruzi production; 5) insoluble proteins from bacteria that induce RxPtx production; 6) soluble proteins from bacteria that induce RxPtx production; 7) insoluble proteins from bacteria that induce TxNeuza production; 8) soluble proteins from bacteria that induce TxNeuza production; 9) insoluble proteins from bacteria that induce RxYFIgG production; 10) soluble proteins from bacteria that induce RxYFIgG production. The arrows indicate the bands representing PlatCruzi (columns 1 and 2), TxCruzi (columns 3 and 4), RxPtx (columns 5 and 6), TxNeuza (columns 7 and 8), and RxYFIgG (columns 9 and 10).
[0052] Figure 12– A patterned cellulose membrane of SARS-CoV-2 with multiple amino acids reacting with IgM antibodies from serum of Covid-19 positive patients is visualized as spots of various shades of gray in a grid-delineated area in a checkerboard pattern. Each square comprises a reaction spot of a region of the cellulose membrane, in which different polypeptide sequences synthesized in a linear form are covalently bound to the membrane surface. The relationship between the physical locations within the membrane and the sequences of the multiple amino acids is listed in Table 18. The combined multi-amino acid sequence represents the following coding sequences: spike protein SARS-CoV-2 (S1: aa1-1273, A7-K19), protein ORF3a (OF3: aa1-275, K22-N2), membrane glycoprotein (M: aa1-222, N5-O23); ORF6 (OF6: aa1-61, P2-P12); ORF7 protein (OF7: aa1-121, P15-Q13), ORF8 protein (OF8: aa1-121, Q16-R17), nucleocapsid protein (N: aa1-419, R20-V17), envelope protein (E: aa1-75, W1-W13), and ORF10 protein region (OF10: aa1-38, W15-W20). Each multi-amino acid has a length of 15 amino acids and 10 amino acids that are adjacent and continuously overlapping.
[0053] Figure 13 – A patterned cellulose multiamino acid membrane of SARS-CoV-2 reacting with IgG antibodies from serum of Covid-19 patients is visualized as spots of various shades of gray within a grid-defined area. Each square comprises a reaction spot within a region of the cellulose membrane, where different polypeptide sequences synthesized in a linear fashion are covalently bound to the membrane surface. The relationship between the physical locations within the membrane and the sequences of the multiamino acids is listed in Table 19. The combined multi-amino acid sequences represent the following coding sequences: SARS-CoV-2 ORF3a protein (ORF3: aa 1-275, A7-C11), membrane glycoprotein (G: aa 1-61, C14-E8); ORF6 protein (ORF6: aa 1-61, E11-E21); ORF7 protein (OF7: aa1-121, E24-F22), ORF8 protein (ORF8: aa1-121, G1-G23), spike protein (S: aa 1-1273, H1-R13), nucleocapsid protein (N: aa1-419, R16-V1), envelope protein (E: aa 1-75, W1-W13), and ORF10 (ORF10: aa 1-38, W15-W20). Each multi-amino acid has a length of 15 amino acids and 10 amino acids that are adjacent and continuously overlapping.
[0054] Figure 14 – A patterned cellulose membrane of SARS-CoV-2 with multiple amino acids reacting with IgA antibodies from the serum of Covid-19 patients is visualized as spots of various shades of gray within regions delineated in a grid pattern. Each square comprises a reaction spot within a cellulose membrane region, where different polypeptide sequences synthesized in a linear manner are covalently bound to the membrane surface. The relationship between the physical locations within the membrane and the sequences of the multiple amino acids is listed in Table 20. These combined multi-amino acid sequences represent the following coding sequences: SARS-CoV-2 spike protein (S: aa1-1273, A6-K18), ORF3a (ORF3: aa1-275, K21-N1), membrane glycoprotein (M: aa 1-61, N4-O22); ORF6 (ORF6: aa 1-61, P1-P11); ORF7 (ORF7: aa1-121, P14-Q12), ORF8 (ORF8: aa 1-121, Q15-R13), nucleocapsid protein (N: aa1-419, R16-V1), envelope protein (E: aa 1-75, region V4-V16), and ORF10 (ORF10: aa 1-38, V19-V24). Each multi-amino acid has a length of 15 amino acids and 10 amino acids that are adjacent and continuously overlapping.
[0055] Figure 15 – Reactivity of serum from COVID-19 patients with SARS-CoV-2 peptides synthesized on cellulose membranes, as indicated by alkaline phosphatase-labeled anti-human IgM antibody. Figure 12 ). 15A, surface glycoprotein; 15B, ORF 3a; 15C, membrane glycoprotein; 15D, ORF 6; 15E, ORF 7; 15F, ORF 8; 15G, nucleoprotein; 15H, E protein; 15I, ORF 10.
[0056] Figure 16 – Reactivity of serum from COVID-19 patients with SARS-CoV-2 peptides synthesized on cellulose membranes, as indicated by alkaline phosphatase-labeled anti-human IgG antibodies. Figure 13 ). 16A: ORF 3a; 16B: membrane glycoprotein; 16C: ORF 6; 16D: ORF 7; 16E: ORF 8; 16F: nucleoprotein; 16G: E protein; 16H: ORF 10.
[0057] Figure 17 – Reactivity of serum from COVID-19 patients with SARS-CoV-2 peptides synthesized on cellulose membranes, as indicated by alkaline phosphatase-labeled anti-human IgG antibodies. Figure 14). 17A, surface glycoprotein; 17B, ORF 3a; 17C, membrane glycoprotein; 17D, ORF 6; 17E, ORF 7; 17F, ORF 8; 17G, nucleoprotein; 17H, E protein; 17I, ORF 10.
[0058] Figure 18 – ELISA of serum from hospitalized patients (n=36) (Group 3) using branched-chain synthetic peptides (SARS-X1-SARS-X8) and colorimetric analysis with anti-IgM secondary antibody.
[0059] Figure 19 – ELISA of serum from hospitalized patients (n=36) using branched-chain synthetic peptides (SARS-X1-SARS-X8) and colorimetric analysis with anti-IgG secondary antibody.
[0060] Figure 20 – ELISA of serum from hospitalized patients (n=36) using branched-chain synthetic peptides (SARS-X4-SARS-X8) and colorimetric analysis with anti-IgA secondary antibody.
[0061] Figure 21 – ELISA of serum from four patient groups (Group 1: asymptomatic, 2: suspected, 3: hospitalized, and 4: immune protection) using SARS-X3 branched synthetic peptides and color development with anti-IgM secondary antibody.
[0062] Figure 22 – ELISA of serum from four patient groups (Group 1: asymptomatic, 2: suspected, 3: hospitalized, and 4: immune protection) using SARS-X8 branched synthetic peptides and colorimetric analysis with anti-IgG secondary antibody.
[0063] Figure 23 – ELISA of serum from four patient groups (Group 1: asymptomatic, 2: suspected, 3: hospitalized, and 4: immune protected) using the synthetic peptide SARS-X7 and colorimetric assay with anti-IgA secondary antibody.
[0064] Figure 24Electrophoresis of polyacrylamide gel (SDS-PAGE) showed the production of Ag-Covid19, Ag-COVID19 protein with a six-histidine tail, Tx-SARS-IgM, Tx-SARS2-IgG, Tx-SARS2-G / M, Tx-SARS2-IgA, Tx-SARS2-Universal, and Tx-SARS2-G5. Columns 1 to 10 represent: 1) molecular weight; 2) total bacterial extract after induction without recombinant protein; 3) total bacterial extract after induction of Ag-COVID19 production; 4) total bacterial extract after induction of Ag-COVID19 protein with six histidine tails; 5) total bacterial extract after induction of Tx-SARS2-IgM protein; 6) total bacterial extract after induction of Tx-SARS2-IgG protein; 7) total bacterial extract after induction of Tx-SARS2-G / M protein; 8) total bacterial extract after induction of Tx-SARS2-IgA protein; 9) total bacterial extract after induction of Tx-SARS2-Universal protein; and 10) total bacterial extract after induction of Tx-SARS2-G5 protein. The letters representing the molecular weight standards are A) 250 kDa; B) 130 kDa; C) 100 kDa; D) 70 kDa; E) 55 kDa; F) 35 kDa; and G) 25 kDa.
[0065] Figure 25 - Electrophoresis of polyacrylamide gel (SDS-PAGE) showed the purification of Ag-COVID19 protein by affinity chromatography. (Total) Spectrum of total bacterial extract after induction production; (FT) Spectrum of proteins not bound to a nickel column; (200) Elution spectrum of Ag-COVID19 protein after addition of 200 mM imidazole; (75) Elution spectrum of Ag-COVID19 protein after addition of 75 mM imidazole; and (500) Elution spectrum of Ag-COVID19 protein after addition of 500 mM imidazole.
[0066] Figure 26 - Electrophoresis of polyacrylamide gel (SDS-PAGE) showed the purification of Tx-SARS2-G5 protein by affinity chromatography. (Total) Spectrum of total bacterial extract after induction; (FT) Spectrum of proteins not bound to a nickel column; (200) Elution spectrum of Tx-SARS2-G5 protein after addition of 200 mM imidazole; (75) Elution spectrum of Tx-SARS2-G5 protein after addition of 75 mM imidazole; and (500) Elution spectrum of Tx-SARS2-G5 protein after addition of 500 mM imidazole.
[0067] Figure 27– An ELISA was performed using serum from seven groups of patients infected with malaria, dengue fever, or SARS-CoV-2 who were still hospitalized, recovered, suspected, or asymptomatic. As a control, a collection of serum from healthy individuals collected before the epidemic was used. The Ag-COVID19 protein was used in the ELISA, and the antibody was colorimetrically bound to the secondary antibody against anti-human IgG.
[0068] Figure 28 - ELISA testing of serum from six groups of patients with syphilis, malaria, dengue fever, or who were admitted (hospitalized) or suspected of having SARS-CoV2. As a control, a collection of serum from healthy individuals collected before the epidemic was used. The Tx-SARS2-G5 protein was used in the ELISA, and the antibody was colorimetrically bound to the secondary antibody against anti-human IgG.
[0069] Figure 29 – Antibody titration against Ag-COVID19 by ELISA in mice 2 or 4 weeks after immunization with Ag-COVID19 protein.
[0070] Figure 30 – Antibody purification using Ag-COVID19 protein. Detailed Implementation
[0071] Although the invention is susceptible to various embodiments, preferred embodiments are shown in the accompanying drawings and the following detailed discussion. It should be understood that this specification should be considered as an example of the principles of the invention and is not intended to limit the invention as described herein.
[0072] This article will use some abbreviations. The following is a list of the abbreviations:
[0073] Regarding nitrogenous bases:
[0074] C = cytosine; A = adenine; T = thymine; G = guanine
[0075] About amino acids:
[0076] I = Isoleucine; L = Leucine; V = Valine; F = Phenylalanine; M = Methionine; C = Cysteine; A = Alanine; G = Glycine; P = Proline; T = Threonine; S = Serine; Y = Tyrosine; W = Tryptophan; Q = Glutamine; N = Asparagine; H = Histidine; E = Glutamic acid; D = Aspartic acid; K = Lysine; R = Arginine.
[0077] Protein containers
[0078] This invention relates to the production of protein containers and their use in various methods and compositions based on a sequence of a green fluorescent protein referred to herein as GFP. The methods and compositions utilize the protein containers to simultaneously display multiple different or identical exogenous multi-amino acid sequences at more than four different protein sites, and further to exhibit sufficient fluorescence intensity, efficient expression in cellular protein production systems, and use as reagents for research, diagnostics, or in vaccine compositions.
[0079] In a first embodiment, the present invention relates to a stable protein structure supporting the simultaneous insertion of four or more exogenous multi-amino acid sequences at different sites. In another embodiment, the protein container comprises the amino acid sequence shown in SEQ ID NO:1. In another embodiment, the protein container comprises the amino acid sequence shown in SEQ ID NO:3. In another embodiment, the protein container comprises the amino acid sequence shown in SEQ ID NO:77. In another embodiment, the protein container presents insertion sites for exogenous multi-amino acid sequences in a protein ring facing the external environment. In another embodiment, the simultaneous insertion of exogenous multi-amino acid sequences does not interfere with the manufacturing conditions of the container protein. In another embodiment, the protein container simultaneously contains exogenous multi-amino acid sequences for use in vaccine compositions, for diagnostics, or for the development of laboratory reagents. In another embodiment, the exogenous multi-amino acid sequences do not lose their immunogenic properties when simultaneously inserted into the protein ring of the container protein. In another embodiment, the protein container simultaneously comprises the exogenous multi-amino acid sequences SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, and SEQ ID NO:16. In another embodiment, the protein container comprises the amino acid sequence shown in SEQ ID NO:18. In another embodiment, the protein container simultaneously comprises the exogenous multi-amino acid sequences SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, and SEQ ID NO:29. In another embodiment, the protein container comprises the amino acid sequence shown in SEQ ID NO:20. In another embodiment, the protein container simultaneously comprises multiple copies of the exogenous multi-amino acid sequence SEQ ID NO:30. In another embodiment, the protein container comprises the amino acid sequence shown in SEQ ID NO:31. In another embodiment, the protein container simultaneously comprises the exogenous multi-amino acid sequences SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, and SEQ ID NO:40. In another embodiment, the protein container comprises the amino acid sequence shown in SEQ ID NO:.33.In another embodiment, the protein container simultaneously comprises the exogenous multi-amino acid sequences SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, and SEQ ID NO:44. In another embodiment, the protein container comprises the amino acid sequence shown in SEQ ID NO:45. In another embodiment, the protein container simultaneously comprises the exogenous multi-amino acid sequences SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, and SEQ ID NO:50. In another embodiment, the protein container comprises the amino acid sequence shown in SEQ ID NO:51. In another embodiment, the protein container simultaneously comprises the exogenous multi-amino acid sequences SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58, SEQ ID NO:59, SEQ ID NO:60, SEQ ID NO:61, SEQ ID NO:62, SEQ ID NO:95, and SEQ ID NO:96. In another embodiment, the protein container comprises the amino acid sequence shown in SEQ ID NO:64. In another embodiment, the protein container simultaneously comprises the exogenous multi-amino acid sequences SEQ ID NO:65, SEQ ID NO:66, SEQ ID NO:67, SEQ ID NO:68, SEQ ID NO:69, SEQ ID NO:70, SEQ ID NO:71, SEQ ID NO:72, SEQ ID NO:73, SEQ ID NO:74, and SEQ ID NO:97. In another embodiment, the protein container comprises the amino acid sequence shown in SEQ ID NO:75. In another embodiment, the protein container simultaneously comprises the exogenous multi-amino acid sequences SEQ ID NO:79, SEQ ID NO:80, SEQ ID NO:81, SEQ ID NO:82, SEQ ID NO:83, SEQ ID NO:84, SEQ ID NO:85, SEQ ID NO:86, SEQ ID NO:87, and SEQ ID NO:98. In another embodiment, the protein container comprises the amino acid sequence shown in SEQ ID NO:88.In another embodiment, the protein container simultaneously comprises the exogenous multi-amino acid sequences SEQ ID NO:7, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:12, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:92, SEQ ID NO:93, SEQ ID NO:94, and SEQ ID NO:99. In another embodiment, the protein container comprises the amino acid sequence shown in SEQ ID NO:90. In another embodiment, the protein container simultaneously comprises the exogenous polyamino acid sequences defined as in SEQ ID NO: 100, 124, 125, and 126; or simultaneously comprises the exogenous polyamino acid sequences defined as in SEQ ID NO: 101, 127, 128, 129, 130, 131, 132, 133, 134, and 135; or simultaneously comprises the exogenous polyamino acid sequences defined as in SEQ ID NO: 136, 137, 138, 139, 140, 141, and 142; or simultaneously comprises the exogenous polyamino acid sequences defined as in SEQ ID NO: 129, 133, 135, 137, 140, 141, 142, 143, 144, 146, 147, and 148; or simultaneously comprises the exogenous polyamino acid sequences defined as in SEQ ID NO: 100, 124, 125, and 126; or simultaneously comprises the exogenous polyamino acid sequences defined as in SEQ ID NO: 101, 127, 128, 129, 130, 131, 132, 133, 134, 146, 147, and 148; or simultaneously comprises the exogenous polyamino acid sequences defined as in SEQ ID NO: 100, 124, 125, and 126. The protein container comprises the exogenous multi-amino acid sequences defined in SEQ ID NO: 103, 149, 150, 151, 152, 153, 154, and 155; or, simultaneously comprises the exogenous multi-amino acid sequences defined in SEQ ID NO: 104, 156, 157, 158, 159, 160, 161, 162, and 163; or, simultaneously comprises the exogenous multi-amino acid sequences defined in SEQ ID NO: 136, 139, 140, 141, 142, 143, 144, 146, and 147. In another embodiment, the protein container comprises any amino acid sequence shown in SEQ ID NO: 334-341.
[0080] Another embodiment of the invention relates to the efficient expression of the protein container in a cellular system, based on a sequence of GFP carrying one or two, more than two, three or four, or more than ten exogenous multi-amino acid sequences at 10 different sites on the container protein. Specifically, the invention relates to the efficient expression of the protein container simultaneously displaying exogenous multi-amino acid sequences at up to 10 different protein sites. More specifically, the invention relates to the efficient expression of the container carrying the exogenous multi-amino acid sequences inserted into 10 different protein sites without sacrificing its inherent properties, such as autofluorescence.
[0081] This invention also relates to the production and use of protein containers, “platforms,” “Rx,” and “Tx,” and their amino acid sequences (described in SEQ ID NO:1, SEQ ID NO:3, and SEQ ID NO:77, respectively), their nucleotide sequences (described in SEQ ID NO:2, SEQ ID NO:4, and SEQ ID NO:78, respectively), and their amino acid sequences including selected exogenous multi-amino acid sequences (described in SEQ ID NO:18, SEQ ID NO:20, SEQ ID NO:31, SEQ ID NO:33, SEQ ID NO:45, SEQ ID NO:51, SEQ ID NO:64, SEQ ID NO:75, SEQ ID NO:88, and SEQ ID NO:90). The container proteins can also be modified to suit their intended use by inserting accessory elements. Sequences that facilitate purification processes can also be added to the container proteins, such as, but not limited to, multihistidine tails, chitin-binding proteins, maltose-binding proteins, calmodulin-binding proteins, strep-tags, and GST. For example, thiorodixin and other sequences used for stability can be integrated into container proteins. Sequences such as V5, Myc, HA, Spot, and FLAG, which can facilitate any antibody detection process, can also be added to container proteins.
[0082] Furthermore, for the purposes of this invention, sequences that are at least about 85%, more preferably at least about 90%, 95%, 96%, 97%, 98%, or 99% identical to the proteins and multi-amino acid receptors described herein, as determined by known sequence identity evaluation algorithms such as FASTA, BLAST, or Gap, are also included.
[0083] Sequences that function as target cleavage sites (catalytic sites) for proteases can also be added to protein containers, allowing the main protein to be separated from the aforementioned accessory elements, including for purposes strictly adhering to the optimization of protein production or purification, but without contributing to the proposed end use. These sequences containing sites that function as targets for proteases can be inserted anywhere and include, but are not limited to, thrombin, factor Xa, intestinal peptidase, PreScission, and TEV (Kosobokova et al., Biochemistry, 8:187-200, 2015). Another accessory sequence for labeling proteases can be AviTag, which allows for specific biotinylation at a single site during or after protein expression. In this case, labeled proteins can be generated by combining different elements (Wood, Current Opinion in Structural Biology, 26:54-61, 2014).
[0084] The isolated container proteins can be further modified in vitro for different applications.
[0085] Polynucleotides
[0086] In a first embodiment, the present invention relates to a polynucleotide comprising any one of SEQ ID NO: 2, 4, 78, 17, 19, 32, 34, 46, 52, 63, 76, 89, 91, 326-333 and their degenerate sequences, capable of generating polypeptides defined by SEQ ID NO: 1, 3, 77, 18, 20, 31, 33, 45, 51, 64, 75, 88, 90, 334-341, respectively.
[0087] The present invention also provides isolated container proteins generated from DNA molecules by any expression system, the DNA molecules comprising regulatory elements containing nucleotide sequences encoding selected container proteins.
[0088] With respect to the identity or position of one or more amino acid residues, the DNA sequence encoding the container protein differs from the DNA sequence of the naturally occurring GFP form through the deletion, addition, or substitution of amino acids. However, they still retain some or all of the characteristics inherent in the naturally occurring form, such as, but not limited to, fluorescence production, characteristic three-dimensional shape, ability to be expressed in different systems, and ability to accept exogenous peptides.
[0089] The DNA sequence encoding the container protein of the present invention includes: integration of a preferred codon for expression via certain expression systems; insertion of a cleavage site for a restriction enzyme; insertion of an optimized sequence for facilitating the construction of an expression vector; and insertion of an enhancer sequence to include a selected multi-amino acid sequence of the container protein to be introduced. All of these strategies are known in the art.
[0090] Furthermore, the present invention provides genetic elements, for example, those comprising adding a nucleotide sequence encoding a container protein to the sequences described in SEQ ID NO:2, SEQ ID NO:4, and SEQ ID NO:78. Additionally, elements comprising a nucleotide sequence encoding a container protein, wherein DNA encoding a selected exogenous multi-amino acid sequence is added to the sequence. Such genetic elements comprise sequences described in SEQ ID NO:17, SEQ ID NO:19, SEQ ID NO:32, SEQ ID NO:34, SEQ ID NO:46, SEQ ID NO:52, SEQ ID NO:63, SEQ ID NO:76, SEQ ID NO:89, and SEQ ID NO:91.
[0091] The regulatory elements required for container protein expression include a promoter sequence for binding to RNA polymerase and a translation initiation sequence for binding to the ribosome. For example, bacterial expression vectors must include a start codon, a promoter, and a Shine-Dalgarno sequence suitable for the cellular system and translation initiation. Similarly, eukaryotic expression vectors include a promoter, a start codon, a downstream polyadenylation signal, and a stop codon. Such vectors can be commercially available or constructed from known prior art sequences.
[0092] carrier
[0093] In a first embodiment, the present invention relates to a vector comprising polynucleotides as defined above.
[0094] The conversion from one plasmid to another can be achieved by adding or deleting restriction sites to change the nucleotide sequence without modifying the amino acid sequence, which can be accomplished through nucleic acid amplification technology.
[0095] Expression Box
[0096] In a first embodiment, the present invention relates to an expression cassette comprising polynucleotides as defined above.
[0097] Optimized expression in other systems can be achieved by altering the nucleotide sequence to add or delete restriction sites, and by optimizing codons for alignment with preferred codons of the new expression system without changing the final amino acid sequence.
[0098] cell
[0099] In a first embodiment, the present invention relates to cells comprising a vector or expression cassette as defined above.
[0100] The present invention further provides cells containing a nucleotide sequence encoding a container protein, or a container protein with added DNA encoding a selected exogenous multi-amino acid sequence, to function as a container protein expression system. The cells can be bacterial, fungal, yeast, insect, plant, or even animal cells. The DNA sequence encoding a container protein with added DNA encoding a selected exogenous multi-amino acid sequence can be inserted into a virus, which can be used for container protein expression and production as a delivery system, such as baculovirus, adenovirus, adenovirus-associated virus, alphavirus, herpesvirus, poxvirus, retrovirus, lentivirus, but not limited to these.
[0101] Various methods exist for introducing exogenous genetic material into cells, all of which are known in the prior art. For example, exogenous DNA can be introduced into cells via calcium phosphate precipitation. Other techniques can be used in the development of this invention, such as electroporation, liposome transfection, microinjection, retroviral vectors, and other viral vector systems such as adenovirus-associated virus systems.
[0102] This invention provides a living organism comprising a cell containing at least a DNA molecule, said DNA molecule containing regulatory elements for the expression of sequences encoding container proteins. This invention can be used for the production of container proteins in vertebrates, invertebrates, plants, and microorganisms.
[0103] Container protein expression can be performed in, but is not limited to, the following: *Escherichia coli* cells, *Bacillus subtilis*, *Saccharomyces cerevisiae*, *Pichia pastoris*, *Pichia methanolica*, *Candida boidinii*, *Pichia angusta*, mammalian cells such as CHO cells, HEK293 cells, or insect cells such as Sf9 cells. All known prokaryotic or eukaryotic protein expression systems in the prior art can be used to produce container proteins.
[0104] In one development of this invention, a virus or bacteriophage carrying a coding sequence for a container protein can infect specific types of bacteria or eukaryotic cells and provide expression of the container protein in that cellular system. Infection can be readily observed by detecting the expression of the container protein. Similarly, eukaryotic plant or animal cell viruses carrying sequences encoding container proteins can infect specific cell types and result in the expression of the container protein in eukaryotic cellular systems.
[0105] Methods for producing protein containers
[0106] In a first embodiment, the present invention relates to a method for producing protein containers, comprising introducing a polynucleotide as defined above into competent cells of interest; culturing the competent cells; and isolating the protein containers containing selected exogenous polyamino acids. In another embodiment, the protein containers are not affected by the insertion of various exogenous polyamino acid sequences.
[0107] Container proteins can also be produced through a variety of synthetic biological systems. The generation of fully synthetic genes is essentially related to three connection-based synthetic systems that are widely described in the prior art.
[0108] The present invention provides a method for producing container proteins using a protein expression system comprising: introducing a DNA sequence encoding the container protein, plus DNA encoding a selected exogenous multi-amino acid sequence, into competent cells of interest; culturing these cells under conditions favorable to producing container proteins containing the selected exogenous multi-amino acid sequence; and isolating container proteins containing the selected exogenous multi-amino acid sequence.
[0109] This invention further provides techniques for producing container proteins containing selected exogenous multiple amino acids. This invention demonstrates an efficient method for expressing container proteins containing selected exogenous multiple amino acids, which facilitates the mass production of proteins of interest. The method for producing container proteins can be carried out in various cellular systems, such as yeast, plants, plant cells, insect cells, mammalian cells, and transgenic animals. Various systems can be used by integrating a codon-optimized nucleic acid sequence that produces the desired amino acid sequence into a plasmid suitable for a particular cellular system. The plasmid may contain elements that confer numerous properties to the expression system, including but not limited to sequences promoting retention and replication, selectable markers, promoter sequences for transcription, stabilizing sequences for transcribed RNA, and ribosome binding sites.
[0110] Methods for isolating expressed proteins are known in the prior art, and in this respect, container proteins can be readily isolated by any technique. The presence of a multihistidine tail allows for the purification of recombinant proteins after expression in bacterial systems (Hochuli et al., Bio / Technology 6:1321-25; Bornhorst and Falke, MethodsEnzymology 326:245-54).
[0111] This invention further considers the selection and choice of exogenous multi-amino acids. Container proteins can contain different multi-amino acid sequences from various sources, ranging from vertebrates, invertebrates, plants, microorganisms, or viruses, to facilitate their expression, display, or use in different media. Different multi-amino acid sequence selection methods can be used, such as specific selection based on binding affinity to antibodies or other binding proteins, epitope mapping, or other techniques known in the art.
[0112] A multi-amino acid sequence is a sequence of 5 to 30 amino acids that sensitively or specifically represents an organism for any of the purposes described in this application. Multi-amino acid sequences may represent, but are not limited to, examples of: (i) linear B-cell epitopes; (ii) T-cell epitopes; (iii) neutralizing epitopes; (iv) protein regions specific to pathogen or non-pathogen sources; and (v) regions adjacent to the active sites of enzymes that are not typically targets of an immune response. These epitope regions can be identified by a variety of methods, including but not limited to: spot synthesis analysis, random peptide libraries, phage display, software analysis, using X-ray crystallography data, epitope databases, or other methods of the prior art.
[0113] Inserting exogenous multi-amino acid sequences into container proteins at the sites defined above surprisingly does not disrupt or interfere with the genetic properties of the protein. Eight introduction sites for exogenous multi-amino acid sequences have been identified in container proteins (Kiss et al., Nucleic Acids Res 34:e132, 2006; Pavoor et al., Proc Natl Acad Sci USA 106:11895-900, 2009; Abedi et al., Nucleic Acids Res 26:623-30, 1998; Zhong et al., Biomol Eng 21:67-72, 2004).
[0114] These introduction sites can contain one, two, or more different exogenous multi-amino acid sequences in tandem at the same insertion site, greatly expanding the expression of different multi-amino acid sequences.
[0115] Methods for pathogen identification or in vitro disease diagnosis
[0116] In a first embodiment, the present invention relates to a method for pathogen identification or in vitro disease diagnosis, characterized in that the method uses a container protein as defined above. In another embodiment, the method is used to diagnose Chagas disease, rabies, pertussis, yellow fever, oropic virus infection, Mayaro, IgE hypersensitivity, house dust mite (D. pteronyssinus) allergy, or COVID-19.
[0117] Uses of protein containers
[0118] In a first embodiment, the present invention relates to the use of the protein container as a laboratory reagent. In a further embodiment, the present invention relates to the use of the protein container for the production of vaccine compositions for immunization against Chagas disease, rabies, pertussis, yellow fever, oropic virus infection, Mayaro, IgE hypersensitivity reactions, house dust mite allergies, or COVID-19.
[0119] Furthermore, the present invention relates to a system for the co-expression of multiple multi-amino acid sequences using a single GFP-based protein container for various applications, such as as a research reagent, for diagnostic purposes, or for use in vaccine compositions. Specifically, the expression system can serve as a useful research reagent for purifying antibodies by binding epitopes. Additionally, the expression system can also serve as an immunological and / or molecular technique for diagnosing chronic and infectious diseases. Moreover, the expression system can be advantageously used in vaccine compositions containing multiple antigens for immunization in animals and humans.
[0120] Furthermore, one embodiment of the present invention is a method for the co-production of multiple multi-amino acid sequences using a single protein container based on GFP, for various uses, such as as a reagent for research, for diagnostics, or for vaccine compositions.
[0121] This invention further demonstrates certain uses of container proteins. These uses include, but are not limited to, protein containers as: (i) as reporter molecules in cell screening assays, including intracellular assays; (ii) as proteins for displaying random or selected peptide libraries; (iii) as antigen-presenting proteins, as reagents for developing in vitro immunological diagnostic assays, typically for infectious, parasitic, or other immunological diseases; (iv) as antigen-presenting proteins for selecting, capturing, screening, or purifying binding substances such as antibodies; (v) as antigen-presenting proteins for use in vaccine compositions; (vi) as proteins containing antibody sequences for binding antigens; and (vii) as antigen-presenting proteins with passive immunization activity.
[0122] Container proteins can be used as vaccine compositions by specifically having the following: (a) a large number of simultaneous immune response-inducing multi-amino acid sequences, and (b) non-immune response-inducing core proteins.
[0123] Diagnostic kits
[0124] In a first embodiment, the present invention relates to a diagnostic kit comprising a protein container as defined above.
[0125] Finally, the invention is described in detail with reference to the embodiments given below. It should be emphasized that the invention is not limited to these embodiments and includes variations and modifications within the scope that can be developed. It is also worth noting that the authorized use of all biological sequences of the Brazilian genetic heritage is registered with SISGEN, registration number AC53976.
[0126] Example
[0127] Example 1 - Container protein construction
[0128] The amino acid sequences of different instances of green fluorescent protein, eGFP (GenBank: L29345.1; UniProtKB-P42212), Cycle-3 (GenBank: CAH64883.1), SuperFolder (GenBank: AOH95453.1), Split (Cabantous et al. Science Reports 3:2854, 2013), and Superfast (Fisher & DeLisa. PLoS One 3:e2351, 2008), were used to construct the novel protein of this invention. Sequence alignment and comparison were performed using Intaglio software (Purgatory Design, V3.9.4). From these data, certain modifications were made to enable the container protein to achieve the desired properties.
[0129] Modifications are made to create restriction enzyme action sites. The insertion of these sites is designed such that the physicochemical properties of the GFP protein are not altered, and therefore the properties or quality described in this patent application are not affected. Furthermore, by allowing potential use in genetic engineering methods and processes, the insertion of these restriction sites will allow for genetic manipulation of these proteins to integrate various peptides into different regions of the protein, thereby adding further properties to the container protein.
[0130] Manipulating the nucleotide sequence of the GFP protein to introduce or replace nucleotides creates new restriction enzyme sites. This results in two new container proteins, the "Platform" protein and the "Rx" protein.
[0131] Changes in the nucleotide sequence of the eGFP protein lead to the following amino acid alterations in the container protein:
[0132] Platform protein
[0133] Position 16, amino acid I;
[0134] Position 28, amino acid F;
[0135] Position 30, amino acid R;
[0136] Position 39, amino acid I;
[0137] Position 43, amino acid S;
[0138] Position 72, amino acid S;
[0139] Position 99, amino acid Y;
[0140] Position 105, amino acid T;
[0141] Position 111, amino acid E;
[0142] Position 124, amino acid V;
[0143] Position 128, amino acid I;
[0144] Position 145, amino acid F;
[0145] Position 153, amino acid T;
[0146] Position 163, amino acid A;
[0147] Position 166, amino acid T;
[0148] Position 167, amino acid V;
[0149] Position 171, amino acid V;
[0150] Position 205, amino acid T;
[0151] Position 206, amino acid I;
[0152] Position 208, amino acid L.
[0153] "Rx" protein
[0154] Position 16, amino acid V;
[0155] Position 28, amino acid S;
[0156] Position 30, amino acid R;
[0157] Position 39, amino acid I;
[0158] Position 43, amino acid T;
[0159] Position 72, amino acid A;
[0160] Position 99, amino acid S;
[0161] Position 105, amino acid K;
[0162] Position 111, amino acid V;
[0163] Position 124, amino acid V;
[0164] Position 128, amino acid T;
[0165] Position 145, amino acid F;
[0166] Position 153, amino acid T;
[0167] Position 163, amino acid A;
[0168] Position 166, amino acid T;
[0169] Position 167, amino acid V;
[0170] Position 171, amino acid V;
[0171] Position 205, amino acid T;
[0172] Position 206, amino acid V;
[0173] Position 208, amino acid S.
[0174] For all container proteins, optional amino acid substitutions can still be made at the following positions:
[0175] Position 39, amino acid N;
[0176] Position 72, amino acid S;
[0177] Position 99, amino acid S;
[0178] Position 105, amino acid Y or K;
[0179] Position 206, amino acid I;
[0180] Position 208, amino acid L.
[0181] The presence of some mutations can affect the biochemical properties of proteins. The S30R mutation positively affects the helical properties of proteins; the Y145F and I171V mutations prevent the translation of unwanted intermediates; and the A206V or I mutations reduce the likelihood of nascent protein aggregation.
[0182] Other changes made to the container proteins “Platform” and “Rx”, ranging from including new nucleotide codons to generating new restriction sites, are shown in Table 1 (below).
[0183] Table 1
[0184] amino acid sequence changes Substituted nucleotides Introduced nucleotides The restriction enzyme to be used D102_D103insV - GTC AatII G116_D117insT - ACC KpnI L137_G138insK - AAG AfIII D191_P192insG - GGT RsrII E213_K214insL - CTC SacI
[0185] Furthermore, the nucleotide sequences of the container proteins “platform” and “Rx” have two additional restriction sites for NdeI and NheI, located at the 5' amino terminus of the protein, from the insertion of the sequence CATATGGTGGCTAGC (SEQ ID NO:5), and two other restriction sites for EcoRI and XhoI, located at the 3' carboxyl terminus, from the insertion of the sequence GAATTCTAATGACTCGAG (SEQ ID NO:6). In addition, two stop codons at the amino terminus and a multihistidine tail have been integrated into the container protein.
[0186] The amino acid sequence of the “platform” protein is shown in SEQ ID NO:1 and the corresponding nucleotide sequence is described in SEQ ID NO:2.
[0187] The amino acid sequence of the “Rx” protein is shown in SEQ ID NO:3 and the corresponding nucleotide sequence is described in SEQ ID NO:4.
[0188] By generating restriction sites without altering the three-dimensional structure of the original protein, 10 novel insertion sites for exogenous multi-amino acid sequences are permitted to appear in the protein. These novel insertion sites will be referred to herein as positions 1 through 10.
[0189] The positions of the container proteins “platform” and “Rx” 1 to 10 in the nucleotide and amino acid sequences are shown in Table 2 (below).
[0190] Table 2
[0191]
[0192]
[0193] Based on the comparison of amino acid sequences of different GFP proteins, a common amino acid sequence was identified as CGP (Dai et al., Protein Engineering, Design and Selection 20(2):69-79, 2007). Despite exhibiting high stability, this fluorescent protein was improved through directed evolution to exhibit better stability than CGP (Kiss et al., Protein Engineering, Design & Selection 22(5):313-23, 2009). However, due to the presence of three mutations, the enhanced protein is prone to aggregation. Based on its crystal structure analysis, other mutations were also integrated, leading to the elimination of aggregation and the production of a protein called Thermal Green Protein (TGP) (Close et al., Proteins 83(7):1225-37, 2015). When used as a protein container, the sequence is referred to as “Tx”. The amino acid sequence of the “Tx” protein is shown in SEQ ID NO:77, and the sequence of its corresponding nucleotides is described in SEQ ID NO:78.
[0194] The nucleotide and amino acid sequence positions of positions 1 to 13 of the “Tx” container protein are shown in Table 3 (below). Two multi-amino acid sequences can be inserted consecutively at the amino and carboxyl termini of the container protein. Therefore, insertion sites 1a and 1b and 13a and 13b are characterized by their locations at the amino and carboxyl termini, respectively.
[0195] Table 3
[0196] Location in protein containers Position in amino acid sequence 1 GAHASVIKPE 2 NG 3 YE 4 GAPLPFS 5 AFPE 6 EDQ 7 GD 8 NFPPNGPVMQKK 9 DG 10 EGGG 11 KKDVRLPDA 12 DKDYN 13 RYSG
[0197] Example 2 – Structure of PlatCruzi protein
[0198] The genetically engineered "platform" container protein possesses Trypanosoma cruzi epitopes, which we refer to in this paper as the PlatCruzi platform. The gene corresponding to the PlatCruzi protein, referred to in this paper as the PlatCruzi gene, is described in its nucleotide sequence SEQ ID NO:17.
[0199] Considering experimental data on the specificity and sensitivity of diagnostic tests for Chagas disease, multi-amino acid sequences derived from Trypanosoma cruzi were selected from available existing technical literature (Peralta JM et al. J Clin Microbiol 32:971-974, 1994; Houghton RL et al. J Infect Dis 179:1226-1234, 1999; Thomas et al. Clin Exp Immunol 123:465-471, 2001; Rabello et al., 1999; Gruber & Zingales, Exp Parasitology, 76(1):1-12, 1993; Lafaille et al., Molecular Biochemistry Parasitology, 35(2):127-36, 1989). Ten multi-amino acid sequences, referred to in this paper as TcEp1 to TcEp10, were selected for insertion into ten insertion sites in the “platform” protein, as shown in Table 4 (below).
[0200] After selecting the polyamino acid sequence of Trypanosoma cruzi, a synthetic gene corresponding to the PlatCruzi protein was generated through chemical synthesis and gene ligation synthesis, and then inserted into a plasmid for experimental use. The amino acid sequence corresponding to the PlatCruzi gene containing epitopes TcEp1 to TcEp10 is described in SEQ ID NO:18.
[0201] Table 4
[0202]
[0203] Example 3 – Development of PlatCruzi protein
[0204] The synthetic gene was introduced into the pET28a plasmid using existing molecular biology techniques and restriction sites for the enzymes NdeI and Xhol. To identify whether the synthetic gene paired with the sequence designed for PlatCruzi, the plasmid material was analyzed by transforming E. coli DH5α strain and then digesting it with restriction enzymes followed by sequencing.
[0205] The sequencing methods used were enzymatic, dideoxy, or chain termination methods, which are based on the enzymatic synthesis of complementary strands. The growth of the complementary strand is stopped by adding dideoxy nucleotides (Sanger et al., Proceeding National Academy of Science, 74(12):5463-5467, 1977). The method consists of the following steps: sequencing reaction (DNA replication in a thermal cycler for 25 cycles), DNA precipitation with isopropanol / ethanol, double-strand denaturation (95°C, 2 min), and nucleotide sequence reading in an ABI 3730XL automated sequencer (ThermoFischer SCIENTIFIC) (Otto et al., Genetics and Molecular Research 7:861-871, 2008). The obtained sequences were analyzed using the 4Peaks program (Nucleobytes; Mac OS X, 2004). Primers from the pET-28a vector (T7 promoter and T7 terminator) were used for the reaction.
[0206] The plasmid clone containing the correct PlatCruzi sequence was transferred into *E. coli* strain BL21 to produce the PlatCruzi protein. Upon induction with isopropyl β-D-1-thiogalactoside (IPTG), strain BL21 expressed T7 RNA polymerase. The strain was grown overnight in LB medium and then inoculated into the same medium supplemented with kanamycin (30 μg / ml) and cultured on a shaker at 200 rpm until a turbidity density of 0.6–0.8 (600 nm) was achieved. IPTG (qsp 1 mM) was then added to the culture and cultured under the same conditions for another 3 hours.
[0207] The culture was centrifuged and the precipitate was resuspended in urea buffer (100 mM NaH2PO4, 10 mM Tris-base, 8 M urea, pH 8.0). The solution was subjected to chromatography through a HisTrap™ affinity column (1 mL, GE Healthcare Life Sciences), which allows for high-resolution purification of histidine-labeled proteins. The supernatant was applied to a nickel affinity column (HisTrap™, 1 mL, GE Healthcare Life Sciences) at a flow rate of 0.5 mL / min, the column having been previously equilibrated in buffer A (50 mM Tris-HCl, pH 8.0, 100 mM NaCl, and 5 mM imidazole). After binding, the resin was washed with 10 mL of buffer A. The protein was eluted for 45 min at a flow rate of 0.7 mL / min in a 100% gradient of buffer B (50 mM Tris-HCl, pH 8.0, 100 mM NaCl, and 500 mM imidazole). PlatCruzi was then purified at 280 nm (black line) and shown. Figure 1 A. The percentage of imidazole is marked in red. Elute the protein in approximately 19 to 25 ml.
[0208] Recombinant protein samples (1 μg / well) were subjected to SDS-PAGE (Laemmli, Nature 227:680-685, 1970). Stacking gels and running gels were prepared at 4% and 11% acrylamide concentrations, respectively (Table 5, below). Samples were prepared under denaturing conditions in 62.5 mM Tris-HCl buffer, pH 6.8, 2% SDS, 5% β-mercaptoethanol, and 10% glycerol, and boiled at 95°C for 5 minutes (Hames BD, Gel electrophoresis of proteins: a practical approach. 3. ed. Oxford. 1998). After electrophoresis, proteins were detected by staining with Coomassie Brilliant Blue R250 (Bio-Rad, USA). Kaleidoscope™ Prestained Standards markers were used as molecular weight references (Bio-Rad, USA). Figure 1 B, Display usage The PlatCruzi protein was purified by affinity chromatography using a liquid chromatography system and a nickel-agar column.
[0209] Table 5: Volumes and concentrations of reagents used to prepare 4% sample stacking gel and 11% separating gel
[0210]
[0211] Example 4 – Development of ELISA from PlatCruzi
[0212] The performance of PlatCruzi protein was evaluated against a set of reference biological samples and individuals infected with Trypanosoma cruzi. PlatCruzi protein in carbonate / bicarbonate buffer (50 mM, pH 9.6) was added to 96-well ELISA plates at concentrations of 0.1, 0.25, 0.5, and 1.0 μg / well and incubated at 4 °C for 12–18 hours. The wells were washed with saline-phosphate buffer (PBS) containing Tween 20 (PBS-T, 10 mM sodium phosphate-Na3PO4, 150 mM sodium chloride-NaCl, and 0.05% Tween-20, pH 7.4) and then incubated at 37 °C for 2 hours with 1x PBS buffer containing 5% (w / v) dehydrated skim milk.
[0213] The wells were then washed three times with PBS-T buffer and incubated for 1 hour at 37°C with 2, 4, 8, 16, 32, and 64-fold diluted reference biological samples TCl (IS 09 / 188) or TCII (IS 09 / 186) (World Health Organization). After incubation, the wells were washed three times with PBS-T and incubated for 1 hour at 37°C with alkaline phosphatase-labeled human IgG antibody diluted 1:5000. The wells were washed three times again with PBS-T and the substrate p-nitrophenyl phosphate (PNPP, 1 mg / mL, ThermoFischer SCIENTIFIC) was added. After 30 minutes in the dark, the absorbance was measured at 405 nm using an ELISA reader.
[0214] Regardless of the amount of PlatCruzi used, the results showed satisfactory responses across all dilutions of the reference biological samples using both TCI and TCII. Figure 2 A and 2B). The results strongly support the use of PlatCruzi to identify Trypanosoma cruzi infection caused by any of the six DTUs (discrete typing units or different typing units), covering the entire geographic range of circulating Trypanosoma cruzi strains. The same results were observed when using serum from Chagas disease patients with low or high antibody detection titers. Figure 3 ).
[0215] ELISA plates containing 500 ng of PlatCruzi (in 0.3 M urea, pH 8.0) were prepared as described above. After washing three times with PBS-T, the plates were incubated at 37°C for 1 hour with serum from four patients (6C-CE, 9C-CE, 15C-CE, and 12-SE) and four patients (3C-PB, 6C-PB, 16C-PB, and 17C-PB) with low anti-Trekrus antibodies at different dilutions of 1:50, 1:100, 1:250, 1:500, and 1:1000. Subsequently, the wells were washed three times with PBS-T and incubated for 1 hour with alkaline phosphatase-labeled human IgG antibody diluted 1:5000. Wash the wells three times again with PBS-T buffer and add the substrate p-nitrophenyl phosphate (PNPP, 1 mg / mL, ThermoFischer SCIENTIFIC). After 30 minutes, measure the absorbance at 405 nm using an ELISA reader.
[0216] The results showed that for sera with low antibody titers, the read signals at dilutions of 1:50, 1:100, and 1:250 were clearly higher than the threshold reached by the negative control, indicating the potential of the PlatCruzi platform for detecting anti-Trypanosoma cruzi antibodies in the sera of both high and low antibody patients.
[0217] The use of patient serum samples for experimental purposes was approved by the Ethics Committee of Fiocruz, in accordance with the authorization CEP / IOC-CAAE:52892216.8.0000.5248.
[0218] Example 5 –Sensitivity and specificity of PlatCruzi ELISA
[0219] Seventy-one serum samples from patients diagnosed with Trypanosoma cruzi, plus 18 serum samples from patients diagnosed with leishmaniasis (negative for Trypanosoma cruzi), 20 serum samples from patients diagnosed with dengue fever (negative for Trypanosoma cruzi), and 39 negative serum samples (from other infectious diseases and uninfected individuals), were diluted 1:250 and incubated at 37°C for 1 hour in ELISA plates, which were the ELISA plates developed in the above examples containing 500 ng of PlatCruzi (0.3 M urea, pH 8.0). The plates were then washed and labeled with antibodies, and the color development and reading procedures were performed as described above.
[0220] Correlation analysis from receiver operating characteristic (ROC) curves indicates the superior sensitivity and specificity of the PlatCruzi platform. Figure 4 No false negatives were observed for sera previously identified as positive for Trypanosoma cruzi; and no false positives were observed for other sera known to be negative for Trypanosoma cruzi, including sera positive for other infectious diseases. Both sensitivity and specificity indices were 100%.
[0221] Example 6 – Development of RxRabies2 protein
[0222] The performance and ability of the "Rx" protein to express epitopes derived from other microorganisms, including viruses, were tested. Literature indicates that a large number of specific multi-amino acid sequences can be used as targets for neutralizing antibodies. However, small variations in the sequence observed between viral strains can interfere with neutralization. Therefore, a thorough study of the optimal multi-amino acid sequences requires extensive knowledge of viral biology and the epidemiology of their interactions with their hosts.
[0223] Considering experimental data on the specificity and sensitivity for diagnosing diseases caused by rabies virus, rabies virus multi-amino acid sequences were selected from available existing technical literature (Kuzmina et al., J Antivir Antiretrovir 5:2:37-43, 2013; Cai et al., Microbes Infect 12:948-955, 2010).
[0224] Ten multi-amino acid sequences, referred to herein as RaEp1 to RaEp10, were selected for insertion into ten insertion sites in the Rx protein, as described below. The combination of these multi-amino acid sequences with the Rx protein sequence yields the RxRabies2 protein. The gene corresponding to the RxRabies2 protein, referred herein as the RxRabies2 gene, is described in nucleotide sequence SEQ ID NO:19. The amino acid sequence corresponding to the RxRabies2 gene containing the multi-amino acid sequences RaEp1 to RaEp10 is described in SEQ ID NO:20.
[0225] Table 6
[0226] Multi-amino acid sequence Location in proteins primitive epitope protein sequence SEQ ID no. RaEp 1 1 Antigen site 1 CKLKLCGVLGL SEQ ID no.21 RaEp 2 2 Antigen site 1 CKLKLCGCSGL SEQ ID no.22 RaEp 3 3 Antigen site 1 CKLKLCGVPGL SEQ ID no.23 RaEp 4 4 - VDERGLYK SEQ ID no.24 RaEp 5 5 - WVAMQTSN SEQ ID no.25 RaEp 6 6 Antigen site III KSVRTWNEI SEQ ID no.26 RaEp 8 8 g5 antigen site LHDFHSD SEQ ID no.27 RaEp 9 9 g5 antigen site LHDFRSD SEQ ID no.28 RaEp 10 10 g5 antigen site LHDLHSD SEQ ID no.29
[0227] A synthetic gene containing the sequence encoding the Rx protein and the sequence encoding the multi-amino acid sequence described in Table 6 above has been synthesized.
[0228] The synthetic gene was introduced into the pET28a plasmid using existing molecular biology techniques and restriction sites for the enzymes NdeI and Xhol. To identify whether the synthetic gene paired with the sequence designed for RxRabies2, the plasmid material was transformed into E. coli DH5α strain and analyzed by restriction enzyme digestion followed by sequencing, in the same manner described in PlatCruzi.
[0229] The plasmid clone with the correct RxRabies2 sequence was transferred into *E. coli* strain BL21 to produce the RxRabies2 protein. Upon induction with isopropyl β-D-1-thiogalactoside (IPTG), strain BL21 expressed T7 RNA polymerase. Strain BL21 was grown overnight in LB medium and then inoculated into the same medium with kanamycin (30 μg / ml) and shaken at 200 rpm until it reached a turbidity density of 0.6–0.8 (600 nm). IPTG (qsp 1 mM) was then added to the culture and the culture was maintained at 37°C for another 3 hours.
[0230] The culture was centrifuged and the precipitate was resuspended in urea buffer (100 mM NaH2PO4, 10 mM Tris-base, 8 M urea, pH 8.0). The solution was then subjected to chromatography through a HisTrap™ affinity column (1 mL, GE Healthcare Life Sciences) at a flow rate of 0.5 mL / min, which allows for high-resolution purification of histidine-tagged proteins. The column was previously equilibrated in buffer A (50 mM Tris-HCl, pH 8.0, 100 mM NaCl, and 5 mM imidazole). After binding, the resin was washed with 10 mL of buffer A. The protein was eluted for 45 min at a flow rate of 0.7 mL / min in a 100% gradient of buffer B (50 mM Tris-HCl, pH 8.0, 100 mM NaCl, and 500 mM imidazole).
[0231] RxRabies2 production was analyzed in three different culture volumes: 3, 25, and 50 ml. As shown in Table 7 (below), the average expression of RxRabies2 was 123 μg / ml.
[0232] Table 7
[0233]
[0234] Example 7 –RxRabies2 protein as a vaccine composition
[0235] RxRabies2 protein was generated as described in the examples above. 100 μg of the protein was suspended in Freud's incomplete adjuvant (0.5 mL) and intramuscularly injected into the quadriceps muscle of two 6-month-old male New Zealand rabbits. Seven and fourteen days after the initial injection, the rabbits were re-injected with Rx rabies protein (100 μg / 0.5 mL) suspended in PBS.
[0236] Blood was collected from animals 21 days after the first dose of the vaccine composition. Plasma was collected by centrifugation and purified by affinity by binding to the Rx-Rabies2 protein adsorbed onto a nitrocellulose membrane.
[0237] Nitrocellulose membranes containing Rx-Rabies2 protein were prepared as described below to separate Rx-Rabies2 protein from potential contaminants. Following electrophoresis, the protein was transferred to the nitrocellulose membrane using existing immunoblotting techniques.
[0238] The 11% SDS-PAGE gel was prepared as described in the examples above and Table 5;
[0239] 10 μg of Rx-Rabies2 protein was applied to an 11% SDS-PAGE gel (Table 5) and subjected to an electrophoretic current of 100 volts for about 2 hours.
[0240] Protein transfer to nitrocellulose membrane: Proteins were transferred to nitrocellulose membranes using a Trans-Blot Cell (Bio-Rad, USA) with transfer buffer (25 mM Tris base, 192 mM glycine and 20% methanol) at 100 V for 1 hour.
[0241] The presence of recombinant protein was confirmed by staining with Ponceau S red (Ponceau S 0.1%, acetic acid 5%).
[0242] The membrane was sheared to obtain a piece containing only the RxRabies2 protein.
[0243] The membrane was then decolorized in distilled water and placed in TBS (0.1%) for 12 to 18 hours (overnight).
[0244] Incubate overnight with blocking buffer (containing 0.05% (v / v) Tween 20 (TBS-T) and 5% (w / v) skim milk powder in 25 mM Tris-HCl, 125 mM NaCl pH 7.4 (TBS)), then incubate the membrane again in blocking buffer for 1 hour, then wash 3 times with TBS-T for 5 minutes each time, and then wash 3 more times with TBS for 5 minutes each time.
[0245] Serum from immunized rabbits was diluted 1:500 in TBS, and 10 ml was then placed in contact with a nitrocellulose membrane fragment containing RxRabies2 protein for 1 hour with stirring. The solution was then washed three times with TBS-T for 5 minutes each time, followed by three more washes with TBS for 5 minutes each time. The specifically bound antibody was then released by adding 1 ml of 100 mM glycine (pH 3.0). The pH of the solution was raised to 7 by adding 100 μl of 1 M Tris (pH 9.0) to purify the rabbit antibody. Purification of antibodies that specifically bind to the antigen is an important step in using these antibodies for therapy because it allows for a significant reduction in the amount to be administered while minimizing potential adverse reactions.
[0246] Different extracts were used to demonstrate the specific ability of the generated rabbit antibodies to bind to RxRabies2. The following were used:
[0247] Crude extracts of bacteria expressing Rx protein were obtained using the same conditions as described in PlatCruzi.
[0248] RxRAbies2 protein in crude bacterial extract, purified at 1x and 0.5x concentrations;
[0249] Purified PlatCruzi protein at 1x and 0.5x concentrations.
[0250] As described above, potential ligands were subjected to polyacrylamide gel electrophoresis (11% SDS-PAGE, Table 5), then transferred to nitrocellulose membranes and subjected to immunoblotting, the details of which have been described for RxRabies2.
[0251] The membrane was incubated with purified anti-RxRabies2 serum as described above for 1 hour with stirring. Then, it was washed three times with TBS-T for 5 minutes each time, followed by three more washes with TBS for 5 minutes each time. Subsequently, the membrane was incubated with a 1:10,000 diluted peroxidase and anti-rabbit IgG secondary antibody for 1 hour. After incubation with the secondary antibody, it was washed three times with TBS-T for 5 minutes each time and then three times with TBS for 5 minutes each time. This was performed using SigmaFast. TM DAB Peroxidase Substrate Tablet was used for color development.
[0252] The results showed that the rabbit antibodies generated by RxRabies2 inoculation specifically bind to ligands containing rabies virus 2 protein. Figure 5The images show bands read only in lanes 2, 4, and 6, containing crude bacterial extracts of RxRabies2 protein, as well as purified and diluted RxRabies2 protein at 1x and 0.5x concentrations, respectively.
[0253] The specificity of the immune response was also observed through the absence of bands in lanes 3, 5, and 7, including (i) the Rx container protein without epitope introduction and (ii) the Platcruzi platform diluted 1x and 0.5x, respectively. These results confirm that the immune response is limited to the rabies virus epitope, indicating that the Rx protein itself is not immunogenic.
[0254] Example 8 – Development of RxHolgG3 protein
[0255] Mapping studies of the multi-amino acid sequences of equine immunoglobulins (De-Simone et al., Toxicon 78:83-93, 2014; Wagner et al., Journal of Immunology 173:3230-3242, 2004) have identified the multi-amino acid sequence of equine IgG3 that is recognized by human IgG and IgE, which is applicable to laboratory testing for the diagnosis of equine serum allergy.
[0256] The multi-amino acid sequence DVLFTWYVDGTEV (SEQ ID NO:30) was integrated into the Rx protein at positions 1, 5, 6, 8, 9, and 10 to generate the RxHolgG3 protein. The amino acid sequence of the RxHolgG3 protein is described in SEQ ID NO:31. The nucleotide sequence of the RxHolgG3 protein is described in SEQ ID NO:32. A synthetic gene was synthesized containing the sequence encoding the Rx protein and the sequence encoding the multi-amino acid sequence (SEQ ID NO:30) as described above.
[0257] The synthetic gene was introduced into the pET28a plasmid using existing molecular biology techniques and restriction sites for the enzymes NdeI and Xhol. To identify whether the synthetic gene paired with a sequence designed for the RxHoIgG3 protein, the plasmid material was transformed into the DH5α strain of *E. coli* and analyzed by restriction enzyme digestion followed by sequencing, as described in detail in PlatCruzi.
[0258] The plasmid clone with the correct sequence of RxHoIgG3 was transferred into *E. coli* strain BL21 to produce the RxHoIgG3 protein. Upon induction with isopropyl β-D-1-thiogalactoside (IPTG), strain BL21 expressed T7 RNA polymerase. Strain BL21 was grown overnight in LB medium and then inoculated into the same medium with kanamycin (30 μg / ml) and shaken at 200 rpm until it reached a turbidity density of 0.6–0.8 (600 nm). IPTG (qsp 1 mM) was then added to the culture and cultured under the same conditions for another 3 hours.
[0259] The culture was centrifuged and the precipitate was resuspended in urea buffer (100 mM NaH2PO4, 10 mM Tris-base, 8 M urea, pH 8.0). The solution was chromatographyd using a nickel affinity column (HisTrap™, 1 mL, GE Healthcare Life Sciences) at a flow rate of 0.5 mL / min, the column having been previously equilibrated in buffer A (50 mM Tris-HCl, pH 8.0, 100 mM NaCl, and 5 mM imidazole). After binding, the resin was washed with 10 mL of buffer A. The protein was eluted for 45 min in a 100% gradient of buffer B (50 mM Tris-HCl, pH 8.0, 100 mM NaCl, and 500 mM imidazole) at a flow rate of 0.7 mL / min. Production of RxHolgG3 protein was demonstrated in the eluent following 11% SDS-PAGE electrophoresis (Table 5). Results are shown in… Figure 6 Furthermore, it showed that RxHolgG3 was expressed as a recombinant protein. In uninduced soluble bacterial extracts (…), Figure 6 No bands were detected in column 1, and bands were detected in both the insoluble (column 2) and soluble (column 3) sections.
[0260] Example 9 – Development of RxOro protein
[0261] Based on multi-amino acid sequence mapping studies of oropouche viruses (strains Q71MJ4 and Q9J945, Uniprot) (Acrani et al., Journal of General Virology 96:513–523, 2014; Tilston-Lunel et al., Journal of General Virology 96(Pt 7):1636–1650, 2015), and considering their diagnostic potential, we selected multi-amino acid sequences from dot synthesis or peptide microarray techniques available in the prior art. Six multi-amino acid sequences, referred to in this paper as OrEp1 to OrEp7, were selected for insertion into nine insertion sites in the Rx protein, as shown in Table 8 below:
[0262] Table 8
[0263]
[0264] The combination of these multi-amino acid sequences with the Rx protein produces the RxOro protein. The amino acid sequence corresponding to the RxOro gene, which contains the multi-amino acid sequences OrEp1 to OrEp6, is described in SEQ ID no. 33. The gene corresponding to the RxOro protein, referred to herein as the RxOro gene, is described in nucleotide sequence SEQ ID no. 34.
[0265] The synthetic RxOro gene was introduced into the pET28a plasmid using existing molecular biology techniques and restriction sites for the enzymes NdeI and Xhol. To identify whether the synthetic gene paired with the sequence designed for RxOro, the plasmid material was transformed into the DH5α strain of *E. coli* and analyzed by restriction enzyme digestion and subsequent sequencing, as described in PlatCruzi.
[0266] The plasmid containing the correct sequence of the RxOro protein was transferred into *E. coli* strain BL21. Upon induction with isopropyl β-D-1-thiogalactoside (IPTG), strain BL21 expressed T7 RNA polymerase. Strain BL21 was grown overnight in LB medium and then inoculated into the same medium with kanamycin (30 μg / ml) and shaken at 200 rpm until it reached a turbidity density of 0.6–0.8 (600 nm). IPTG (qsp 1 mM) was then added to the culture and cultured under the same conditions for another 3 hours.
[0267] The culture was centrifuged and the precipitate was resuspended in urea buffer (100 mM NaH2PO4, 10 mM Tris-base, 8 M urea, pH 8.0). The solution was chromatographyd using a nickel affinity column (HisTrap™, 1 mL, GE Healthcare Life Sciences) at a flow rate of 0.5 mL / min, the column having been previously equilibrated in buffer A (50 mM Tris-HCl, pH 8.0, 100 mM NaCl, and 5 mM imidazole). After binding, the resin was washed with 10 mL of buffer A. Protein was eluted for 45 min at a flow rate of 0.7 mL / min in a 100% gradient of buffer B (50 mM Tris-HCl, pH 8.0, 100 mM NaCl, and 500 mM imidazole). RxOro production was tested in three different bacterial growth volumes: 3, 25, and 50 mL. The average expression level of RxOro was 203 μg / mL, and the results are shown in Table 9.
[0268] Table 9
[0269]
[0270] Example 10 – Development of ELISA from RxOro
[0271] The performance of RxOro protein was evaluated using serum samples from a cohort of individuals affected by oropecé virus infection. RxOro protein in solution (0.3 M urea, pH 8.0) was added to 96-well ELISA plates at a rate of 500 ng / well and incubated at 4°C for 12–18 hours. Wells were washed with saline-phosphate buffer (PBS) supplemented with Tween 20 (PBS-T, 10 mM sodium phosphate-Na3PO4, 150 mM sodium chloride-NaCl, and 0.05% Tween-20, pH 7.4) and then incubated at 37°C for 2 hours with 1x PBS buffer containing 5% (w / v) skim milk powder.
[0272] Then, the wells were washed three times with PBS-T buffer and incubated for 1 hour at 37°C with 98 serum samples from patients suspected of having oropucella virus infection and 51 serum samples from healthy patients diluted 1:100. After incubation, the wells were washed three times with PBS-T and then incubated for 1 hour at 37°C with alkaline phosphatase-labeled human IgG antibody diluted 1:5000. The wells were washed three times again with PBS-T buffer and p-nitrophenyl phosphate (PNPP, 1 mg / mL, ThermoFischer SCIENTIFIC) was added. After 30 minutes in the dark, the absorbance was measured at 405 nm using an ELISA reader.
[0273] The results indicate the excellent sensitivity and specificity of using RxOro ( Figure 7 The results strongly support the use of RxOro for the detection of oropecie virus infection.
[0274] The use of patient serum samples for experimental purposes was approved by Fiocruz's ethics committee, in accordance with authorization CEP / IOC-CAAE:52892216.8.0000.5248.
[0275] Example 11 Development of RxMayaro_IgG protein
[0276] Based on multi-amino acid sequence mapping studies of Mayaro viruses (strains Q8QZ73 and Q8QZ72, Uniprot) (Espósito et al., Genome Announcement 3:e01372-15, 2015), and considering their diagnostic potential, we selected multi-amino acid sequences from available prior art literature. Four multi-amino acid sequences, referred to in this paper as MGEp1 to MGEp4, were selected for insertion into nine insertion sites in the Rx protein, as shown in Table 10 below:
[0277] Table 10
[0278] Multi-amino acid sequence Location in proteins primitive epitope protein sequence SEQ ID no. MGEp 1 1 nsP2 KLSATDWSAI SEQ ID no.41 MGEp 2 3 Capsid KPKPQPEK SEQ ID no.42 MGEp 3 4 nsP1 KKMTPSDQI SEQ ID no.43 MGEp 4 5 nsP3 VELPWPLETI SEQ ID no.44 MGEp 2 6 Capsid KPKPQPEK SEQ ID no.42 MGEp 1 7 nsP2 KLSATDWSAI SEQ ID no.41 MGEp 4 8 nsP3 VELPWPLETI SEQ ID no.44 MGEp 3 9 nsP1 KKMTPSDQI SEQ ID no.43 MGEp 1 10 nsP2 KLSATDWSAI SEQ ID no.41
[0279] The combination of these multi-amino acid sequences with the sequence of the Rx protein produces the RxMayaro_IgG protein. The amino acid sequence corresponding to the RxMayaro_IgG gene, which contains the multi-amino acid sequences MGEp1 to MGEp4, is described in SEQ ID no. 45. The gene corresponding to the RxMayaro_IgG protein, referred to herein as the RxMayaro_IgG gene, is described in nucleotide sequence SEQ ID no. 46.
[0280] The synthetic RxMayaro_IgG gene was introduced into the pET28a plasmid using molecular biology techniques known in the art, employing restriction sites for the enzymes NdeI and Xhol. To identify whether the synthetic gene paired with the sequence designed for RxMayaro_IgG, the plasmid material was transformed into the DH5α strain of *E. coli* and analyzed by restriction enzyme digestion followed by sequencing, as cited in PlatCruzi.
[0281] The plasmid containing the correct sequence of the RxMayaro_IgG protein was transferred into *E. coli* strain BL21. Upon induction with isopropyl β-D-1-thiogalactoside (IPTG), strain BL21 expressed T7 RNA polymerase. Strain BL21 was grown overnight in LB medium and then inoculated into the same medium with kanamycin (30 μg / ml) and shaken at 200 rpm until it reached a turbidity density of 0.6–0.8 (600 nm). IPTG (qsp 1 mM) was then added to the culture and cultured under the same conditions for another 3 hours.
[0282] The culture was centrifuged and the precipitate was resuspended in urea buffer (100 mM NaH2PO4, 10 mM Tris-base, 8 M urea, pH 8.0). The solution was chromatographyd at a flow rate of 0.5 mL / min on a nickel affinity column (HisTrap™, 1 mL, GE Healthcare Life Sciences), which had been previously equilibrated in buffer A (50 mM Tris-HCl, pH 8.0, 100 mM NaCl, and 5 mM imidazole). After binding, the resin was washed with 10 mL of buffer A. The protein was eluted for 45 min at a flow rate of 0.7 mL / min in a 100% gradient of buffer B (50 mM Tris-HCl, pH 8.0, 100 mM NaCl, and 500 mM imidazole). Production of RxMayaro_IgG protein could be confirmed in the eluent from 11% SDS-PAGE electrophoresis (Table 5). The RxMayaro_IgG protein was also examined by SDS-PAGE to confirm its production and determine its distribution between the soluble and insoluble fractions. Figure 8 As shown, columns 4 and 7 (arrows) show that RxMayaro_IgG is produced as soluble and insoluble forms.
[0283] RxMayaro_IgG production was tested at three different growth volumes: 3, 25, and 50 ml. As shown in Table 11 (below), the expression level of RxMayaro_IgG was 130 μg / ml on average.
[0284] Table 11
[0285]
[0286] Example 12 – Development of an ELISA from RxMayaro_IgG
[0287] The performance of RxMayaro_IgG protein was evaluated in serum samples from a cohort of individuals affected by Mayaro virus infection. RxMayaro_IgG protein in solution (0.3M urea, pH 8.0) was added to 96-well ELISA plates at a rate of 500 ng / well and incubated at 4°C for 12–18 hours. Wells were washed with saline-phosphate buffer (PBS) supplemented with Tween 20 (PBS-T, 10 mM sodium phosphate-Na3PO4, 150 mM sodium chloride-NaCl, and 0.05% Tween-20, pH 7.4), and then incubated at 37°C for 2 hours with 1x PBS buffer containing 5% (w / v) skim milk powder.
[0288] Then, the wells were washed with PBS-T buffer and incubated for 1 hour at 37°C with 6 serum samples from patients suspected of Mayaro virus infection and 29 serum samples from healthy patients diluted 1:100. After incubation, the wells were washed three times with PBS-T and then incubated for 1 hour at 37°C with alkaline phosphatase-labeled human IgG antibody diluted 1:5000. The wells were washed three times again with PBS-T buffer and p-nitrophenyl phosphate (PNPP, 1 mg / mL, ThermoFischer SCIENTIFIC) was added. After 30 minutes in the dark, the absorbance was measured at 405 nm using an ELISA reader.
[0289] The results indicate the excellent sensitivity and specificity of using RxMayaro_IgG. Figure 9 The results strongly support the use of RxMayaro_IgG for the detection of Mayaro virus infection.
[0290] The use of patient serum samples for experimental purposes was approved by Fiocruz's ethics committee, in accordance with authorization CEP / IOC-CAAE:52892216.8.0000.5248.
[0291] Example 13 Development of RxMayaro IgM protein
[0292] Based on the mapping study of multi-amino acid sequences of Mayaro viruses (strains Q8QZ73 and Q8QZ72, Uniprot) (Espósito et al., Genome Announcement 3:e01372-15, 2015), and considering their diagnostic potential, we selected multi-amino acid sequences from the available existing technical literature. Four multi-amino acid sequences, referred to in this paper as MMEp1 to MMEp4, were selected for insertion into nine insertion sites in the Rx protein, as shown in Table 12 (below):
[0293] Table 12
[0294]
[0295] The combination of these multi-amino acid sequences with the sequence of the Rx protein produces the RxMayaro_IgM protein. The amino acid sequence corresponding to the RxMayaro_IgM gene, which contains the multi-amino acid sequences MMEp1 to MMEp4, is described in SEQ ID no. 51. The gene corresponding to the RxMayaro_IgM protein, referred to herein as the RxMayaro_IgM gene, is described in nucleotide sequence SEQ ID no. 52.
[0296] The synthetic RxMayaro_IgM gene was introduced into the pET28a plasmid using molecular biology techniques known in the art, employing restriction sites for the enzymes NdeI and Xhol. To identify whether the synthetic gene paired with the sequence designed for RxMayaro_IgM, the plasmid material was transformed into the DH5α strain of *E. coli* and analyzed by restriction enzyme digestion followed by sequencing, as previously described in PlatCruzi.
[0297] The plasmid containing the correct sequence of the RxMayaro_IgM protein was transferred into *E. coli* strain BL21. Upon induction with isopropyl β-D-1-thiogalactoside (IPTG), strain BL21 expressed T7 RNA polymerase. Strain BL21 was grown overnight in LB medium and then inoculated into the same medium supplemented with kanamycin (30 μg / ml) and cultured on a shaker at 200 rpm until it reached a turbidity density of 0.6–0.8 (600 nm). IPTG (qsp 1 mM) was then added to the culture and cultured under the same conditions for another 3 hours.
[0298] The culture was centrifuged and the precipitate was resuspended in urea buffer (100 mM NaH2PO4, 10 mM Tris-base, 8 M urea, pH 8.0). The solution was chromatographyd at a flow rate of 0.5 mL / min on a nickel affinity column (HisTrap™, 1 mL, GE Healthcare Life Sciences), which had been previously equilibrated in buffer A (50 mM Tris-HCl, pH 8.0, 100 mM NaCl, and 5 mM imidazole). After binding, the resin was washed with 10 mL of buffer A. Proteins were eluted for 45 min at a flow rate of 0.7 mL / min in a 100% gradient of buffer B (50 mM Tris-HCl, pH 8.0, 100 mM NaCl, and 500 mM imidazole).
[0299] The production of RxMayaro_IgM protein can be demonstrated in the elution buffer by SDS-PAGE electrophoresis, and its distribution between soluble and insoluble fractions can be observed. Figure 8 As shown, columns 5 and 8 (arrows) show RxMayaro_IgM produced as soluble and insoluble.
[0300] RxMayaro_IgM production was also tested at three different growth volumes: 3, 25, and 50 ml. As shown in Table 13 (below), the expression level of RxMayaro_IgM was 205 μg / ml on average.
[0301] Table 13
[0302]
[0303] Example 14 – Development of ELISA based on RxMayaro_IgM
[0304] The performance of RxMayaro_IgM protein was evaluated in serum samples from a cohort of individuals affected by Mayaro virus infection. RxMayaro_IgM protein in solution (0.3M urea, pH 8.0) was added to 96-well ELISA plates at a rate of 500 ng / well and incubated at 4°C for 12–18 hours. Wells were washed with saline-phosphate buffer (PBS) supplemented with Tween 20 (PBS-T, 10 mM sodium phosphate-Na3PO4, 150 mM sodium chloride-NaCl, and 0.05% Tween-20, pH 7.4) and then incubated at 37°C for 2 hours with 1x PBS buffer containing 5% (w / v) skim milk powder.
[0305] Then, the wells were washed three times with PBS-T buffer and incubated for 1 hour at 37°C with 6 serum samples from patients suspected of Mayaro virus infection and 29 serum samples from healthy patients diluted 1:100. After incubation, the wells were washed three times with PBS-T and then incubated for 1 hour at 37°C with alkaline phosphatase-labeled human IgM antibody diluted 1:5000. The wells were washed three times again with PBS-T buffer and p-nitrophenyl phosphate (PNPP, 1 mg / mL, ThermoFischer SCIENTIFIC) was added. After 30 minutes in the dark, the absorbance was measured at 405 nm using an ELISA reader.
[0306] The results indicate the excellent sensitivity and specificity of using RxMayaro_IgM. Figure 10 The results strongly support the use of RxMayaro_IgM for the detection of Mayaro virus infection.
[0307] The use of patient serum samples for experimental purposes was approved by Fiocruz's ethics committee, in accordance with authorization CEP / IOC-CAAE:52892216.8.0000.5248.
[0308] Example 15 – Protein RxPtx Development
[0309] Based on a mapping study of the multi-amino acid sequences of the bacterial toxin protein *Bordetella pertussis* (P04977; P04978; P04979; P0A3R5 and P04981: Uniprot), which causes pertussis, multi-amino acid sequences were selected from available prior art literature, taking into account their diagnostic potential. Ten multi-amino acid sequences, referred to herein as PtxEp1 to PtxEp10, were selected for insertion into nine insertion sites in the Rx protein, as shown in Table 14 below. In this embodiment, the two epitopes are positioned at position 1 using a spacer region (SEQ ID NO:95:SYWKGS). Two additional spacer regions (SEQ ID NO:96:EAAKEAAK) are used to insert the two epitopes at position 10. The purpose of introducing these spacer regions is to create inert physical space between consecutive multi-amino acids, thus helping to prevent binding competition between adjacent antibodies.
[0310] Table 14
[0311]
[0312] The combination of these multi-amino acid sequences with the sequence of the Rx protein produces the protein RxPtx. The amino acid sequence corresponding to the RxPtx gene containing epitopes PtxEp1 to PtxEp10 is described in SEQ ID no. 64. The gene corresponding to the RxPtx protein, referred to herein as the RxPtx gene, is described in nucleotide sequence SEQ ID no. 63.
[0313] The synthetic gene RxPtx was introduced into the pET28a plasmid using existing molecular biology methods and restriction sites for the enzymes NdeI and Xhol. To identify whether the synthetic gene paired with the sequence designed for RxPtx, the plasmid material was transformed into the DH5α strain of *E. coli* and analyzed by restriction enzyme digestion and subsequent sequencing, as previously described in PlatCruzi.
[0314] The plasmid containing the correct sequence of the RxPtx protein was transferred into *E. coli* strain BL21. Upon induction with isopropyl β-D-1-thiogalactoside (IPTG), strain BL21 expressed T7 RNA polymerase. Strain BL21 was grown overnight in LB medium and then inoculated into the same medium supplemented with kanamycin (30 μg / ml) and cultured on a shaker at 200 rpm until it reached a turbidity density of 0.6–0.8 (600 nm). IPTG (qsp 1 mM) was then added to the culture and cultured under the same conditions for another 3 hours.
[0315] The culture was centrifuged and the precipitate was resuspended in 2 mL of PBS with CelLytic (0.5x) and incubated at 4°C for 1 hour. After another centrifugation, the supernatant was collected and the precipitate was resuspended in the same volume of 8M urea solution (pH 8.0). Equal volumes were loaded onto an SDS-PAGE gel (11%, Table 5).
[0316] The RxPtx protein was examined by SDS-PAGE electrophoresis to demonstrate its production and determine its distribution between the soluble and insoluble fractions. Figure 11 As shown, columns 5 and 6 (arrows) show that RxPtx is produced as both soluble and insoluble, with a higher proportion in the insoluble portion.
[0317] Example 16 Development of RxYFIgG protein
[0318] A mapping study of the multi-amino acid sequences of yellow fever virus (strain 17DD and sequence from the p03314-Uniprot archive) was conducted, and multi-amino acid sequences were selected from available prior art literature, taking into account their diagnostic potential. Ten multi-amino acid sequences, referred to herein as YFIgGEp1 to YFIgGEp10, were selected for insertion into nine insertion sites in the Rx protein, as shown in Table 15 below. The two epitopes were positioned at position 10 using a spacer region (SEQ ID NO:97:TSYWKGS). The spacer region functions to create physical space between consecutive epitopes, which helps prevent interactions with antibodies.
[0319] Table 15
[0320]
[0321] The combination of these multi-amino acid sequences with the Rx protein produces the RxYFIgG protein. The amino acid sequence corresponding to the RxYFIgG gene containing epitopes YFIgGEp 1 to YFIgGEp 10 is described in SEQ ID no. 75. The gene corresponding to the RxYFIgG protein, referred to herein as the RxYFIgG gene, is described in nucleotide sequence SEQ ID no. 76.
[0322] The synthetic gene RxYFIgG was introduced into the pET28a plasmid using existing molecular biology methods and restriction sites for the enzymes NdeI and Xhol. To identify whether the synthetic gene paired with the sequence designed for RxYFIgG, the plasmid material was transformed into the DH5α strain of *E. coli* and analyzed by restriction enzyme digestion and subsequent sequencing, as previously described in PlatCruzi.
[0323] The plasmid containing the correct sequence of RxYFIgG was transferred into *E. coli* strain BL21 to produce RxYFIgG protein. Upon induction with isopropyl β-D-1-thiogalactoside (IPTG), strain BL21 expressed T7 RNA polymerase. Strain BL21 was grown overnight in LB medium at 37°C, then inoculated into the same medium with kanamycin (30 μg / ml) and shaken at 200 rpm until it reached a turbidity density of 0.6–0.8 (600 nm). IPTG (qsp 1 mM) was then added to the culture and the culture was maintained at 37°C for another 3 hours.
[0324] The culture was centrifuged and the precipitate was resuspended in 2 mL of PBS with CelLytic (0.5x) and incubated at 4°C for 1 hour. After another centrifugation, the supernatant was collected and the precipitate was resuspended in the same volume of 8M urea solution (pH 8.0). Equal volumes were loaded onto an SDS-PAGE gel (11%, Table 5). Protein RxYFIgG was examined by SDS-PAGE electrophoresis to confirm its production and determine its distribution between soluble and insoluble fractions. Figure 11 As shown, columns 9 and 10 (arrows) show that RxYFIgG is produced as both soluble and insoluble fractions, with a higher proportion in the insoluble fraction.
[0325] Example 17 – Development of TxNeuza protein
[0326] The gene-manipulating container protein “Tx” has a T-cell epitope from the house dust mite (Dermatophogoides pteronyssinus), a major cause of respiratory allergy in humans, which we refer to in this paper as the TxNeuza platform. The gene corresponding to the TxNeuza protein, referred to herein as the TxNeuza gene, is described in nucleotide sequence SEQ ID NO:89.
[0327] Based on the mapping study of T cell polyamino acid sequences of house dust mites, and considering their diagnostic potential for allergies caused by house dust mites, polyamino acid sequences were selected from available prior art literature (Hinz et al., Clin Exp Allergy 45:1601-1612, 2015; Oseroff et al., Clin Exp Allergy 47:577-592, 2017). Nine polyamino acid sequences, referred to herein as NeuzaEp1 to NeuzaEp9, were selected for insertion into nine insertion sites in the Tx protein, as shown in Table 16 below. In this embodiment, a spacer region (SEQ ID NO:98:GGSG) was used to position the two epitopes at position 12.
[0328] Table 16
[0329]
[0330] The combination of these epitopes with the sequence of the Tx protein produces the TxNeuza protein. The amino acid sequences corresponding to the TxNeuza gene containing epitopes NeuzaEp1 to NeuzaEp9 are described in SEQ ID NO:88.
[0331] The synthetic gene TxNeuza was introduced into the pET28a plasmid using existing molecular biology methods and restriction sites for the enzymes NdeI and Xhol. To identify whether the synthetic gene paired with the sequence designed for TxNeuza, the plasmid material was transformed into E. coli DH5α strain and analyzed by restriction enzyme digestion and subsequent sequencing, as previously described in PlatCruzi.
[0332] The plasmid containing the correct TxNeuza sequence was transferred into *E. coli* strain BL21 to produce the TxNeuza protein. Upon induction with isopropyl β-D-1-thiogalactoside (IPTG), strain BL21 expressed T7 RNA polymerase. Strain BL21 was grown overnight in LB medium at 37°C, then inoculated into the same medium with kanamycin (30 μg / ml) and shaken at 200 rpm until it reached a turbidity density of 0.6–0.8 (600 nm). IPTG (qsp 1 mM) was then added to the culture, and the culture was maintained at 37°C for another 3 hours.
[0333] The culture was centrifuged and the precipitate was resuspended in 2 mL of PBS with CelLytic (0.5x) and incubated at 4°C for 1 hour. After another centrifugation, the supernatant was collected and the precipitate was resuspended in the same volume of 8M urea solution (pH 8.0). Equal volumes were loaded onto SDS-PAGE gels (11%, Table 5). TxNeuza protein was analyzed by SDS-PAGE to confirm its production and determine its distribution between soluble and insoluble fractions. Figure 11 As shown, columns 7 and 8 (arrows) show that TxNeuza is produced as an insoluble form.
[0334] Example 18 – Development of TxCruzi protein
[0335] Considering the experimental specificity and sensitivity for diagnostic tests for Chagas disease, multi-amino acid sequences of Trypanosoma cruzi were selected from available existing technical literature (Balouz, et al., Clin Vaccine Immunol 22, 304-312, 2015; Alvarez, et al., Infect Immun 69, 7946-7949, 2001; Fernandez-Villegas, et al., JAntimicrob Chemother 71, 2005-2009, 2016; Thomas, et al., Clin Vaccine Immunol 19, 167-173, 2012). Ten multi-amino acid sequences were selected for insertion into ten insertion sites in the “Tx” protein, as shown in Table 17 below. These ten multi-amino acid sequences are referred to herein as TcEp 1, TcEp 3, TcEp 4, TcEp 6, TcEp 8, TcEp 9, TcEp 10, TcEp 11, TcEp 12, and TcEp 13. The two epitopes are positioned at position 12 using a spacer (SEQ ID NO:99:GGASG).
[0336] Table 17
[0337]
[0338] After selecting the multi-amino acid sequence of Trypanosoma cruzi, a synthetic gene corresponding to the TxCruzi protein was generated through chemical synthesis and gene ligation synthesis, and then inserted into a plasmid for experimental use. The nucleotide sequence of the TxCruzi gene, corresponding to epitopes TcEp 1, TcEp 3, TcEp 4, TcEp 6, TcEp 8, TcEp 9, TcEp 10, TcEp 11, TcEp 12, and TcEp 13, is described in SEQ ID no. 91.
[0339] The combination of the sequences of these epitopes with the sequence of the Tx protein produces the TxCruzi protein. The amino acid sequences of the TxCruzi gene, corresponding to epitopes TcEp 1, TcEp 3, TcEp 4, TcEp 6, TcEp 8, TcEp 9, TcEp 10, TcEp 11, TcEp 12, and TcEp 13, are described in SEQ ID NO:90.
[0340] The synthetic gene was introduced into the pET28a plasmid using existing molecular biology methods and restriction sites for the enzymes NdeI and Xhol. To identify whether the synthetic gene paired with the sequence designed for the TxCruzi protein, the plasmid material was transformed into the DH5α strain of *E. coli* and analyzed by restriction enzyme digestion followed by sequencing, as described in detail in PlatCruzi.
[0341] The plasmid containing the correct TxCruzi sequence was transferred into *E. coli* strain BL21 to produce the TxCruzi protein. Upon induction with isopropyl β-D-1-thiogalactoside (IPTG), strain BL21 expressed T7 RNA polymerase. Strain BL21 was grown overnight in LB medium at 37°C, then inoculated into the same medium with kanamycin (30 μg / ml) and shaken at 200 rpm until it reached a turbidity density of 0.6–0.8 (600 nm). IPTG (qsp 1 mM) was then added to the culture, and the culture was maintained at 37°C for another 3 hours.
[0342] The culture was centrifuged and the precipitate was resuspended in 2 mL of PBS with CelLytic (0.5x) and incubated at 4°C for 1 hour. After another centrifugation, the supernatant was collected and the precipitate was resuspended in the same volume of 8M urea solution (pH 8.0). Equal volumes were loaded onto an SDS-PAGE gel (11%, Table 5).
[0343] The TxCruzi protein was examined by SDS-PAGE to demonstrate its production and determine its distribution between the soluble and insoluble fractions. Figure 11 As shown, columns 3 and 4 (arrows) show the production of TxCruzi as both soluble and insoluble. Compared to PlatCruzi (columns 1 and 2), TxCruzi shows an increased proportion of soluble protein produced, which can be attributed to the use of Tx as a container protein.
[0344] Example 19 - Synthesis of a SARS-CoV-2 peptide library on a cellulose membrane and its reactivity with sera from SARS-CoV-2 positive and negative individuals
[0345] Based on the SARS-CoV-2 genome sequence isolated in Wuhan, China and published in the GenBank database (https: / / www.ncbi.nlm.nih.gov / nuccore / MN908947.3?report=genbank), a polypeptide library covering all protein-coding regions of ORF3a, ORF6, ORF7, ORF8, ORF10, N, M, S, and E of the SARS-CoV-2 virus was synthesized, and annotated as follows:
[0346] Four peptides not encoded by the SARS-CoV-2 virus are included in the peptide library list to represent positive controls for reactivity with human serum. In Table 18, A1, V5 (IHLVNNESSEVIVHK, Clostridium tetani precursor peptide), A2, V6 (GYPKDGNAFNNLDR, Clostridium tetani), A3, V7 (KEVPALTAVETGATG, human poliovirus), and A4, V8 (YPYDVPDYAGYPYD, triple hemagglutinin peptide) are used as such controls. In Tables 19 and 20, A1, V4 (IHLVNNESSEVIVHK, Clostridium tetani precursor peptide), A2, V5 (GYPKDGNAFNNLDR, Clostridium tetani), A3, V6 (KEVPALTAVETGATG, human poliovirus), and A4, V8 (YPYDVPDYAGYPYD, trihemagglutinin peptide) were used as controls.
[0347] As negative controls, peptide-free spot reactants were used for A5, A6, K20, K21, N3, N4, O24, P1, P13, P14, Q14, Q15, R15, R16, V3, V4, V9-V24, W14 in Table 18, and A5, K19, K20, N2, N3, O23, O24, P12, P13, Q13, Q14, R15, R15, V2, V3, V17, V18 in Tables 19 and 20.
[0348] The relationships between the synthesized linear polyamino acids are shown in Tables 18, 19, and 20.
[0349] Table 18 – For Figure 12 List of synthetic SARS-CoV-2 polyamino acids mapped from IgM reactive epitopes in patient serum.
[0350]
[0351]
[0352]
[0353] Table 19 - For Figure 13 List of synthetic SARS-CoV-2 polyamino acids mapped from IgG reactive epitopes derived from patient serum.
[0354]
[0355]
[0356]
[0357] Table 20 - For use Figure 14 List of synthetic SARS-CoV-2 polyamino acids mapped from IgA reactive epitopes in patient serum.
[0358]
[0359]
[0360]
[0361] Example 20 Synthesis of a multi-amino acid library of SARS-CoV-2 on a cellulose membrane
[0362] The polypeptide library described in Example 19 was prepared on an Amino-PEG500-UC540 cellulose membrane according to the manufacturer's instructions, using an Auto-Spot Robot ASP-222, and following standard spot synthesis techniques. A multi-amino acid complex with a length of 15 residues and 10 overlapping adjacent residues, covering the full length of the protein, was synthesized.
[0363] After synthesis, the free sites of the membrane were blocked for 90 minutes with 1.5% BSA (bovine serum albumin) prepared in TBS-T buffer (50 mM Tris, NaCl; 136 mM, 2 mM KCl; 0.05%, Tween-20; pH 7.4). The membrane was then incubated with patient serum (n = 3; 1:100, diluted in TBS-T containing 0.75% BSA) and washed four times with TBS-T. Subsequently, the membrane was incubated for 1.5 h with goat IgG anti-IgM (mu, KPL), anti-IgG (H+L chain, Thermo Scientific), or anti-human IgA (α chain specific, Calbiochem) (1:5000, prepared in TBS-T), followed by washing with TBS-T and CBS (sodium citrate buffer containing 50 mM NaCl, pH 7.0). Then, [the following steps were performed]. Chemiluminescent substrate (0.25 mM) and Nitro-Block-II TM Enhancer (Applied Biosystems, USA) is used to complete the reaction.
[0364] Chemiluminescence signals were detected on an Odyssey FC instrument (LI-COR Bioscience) and the signal intensity was quantified using TotalLabTL100 software (v 2009, Nonlinear Dynamics, USA). Data were analyzed using Microsoft Excel, and only spots with a signal intensity (SI) greater than or equal to the highest value obtained from the set of spots in each membrane were included in the multi-amino acid signature. Background signal intensities of each membrane were used as negative controls.
[0365] Example 21 – Synthesis of branched-chain polyamino acids from SARS-CoV-2
[0366] According to the manufacturer's instructions, multiple branched-chain polyamino acids (SARS-X1-SARS-X8) were synthesized using the F-moc solid-phase polyamino acid synthesis strategy on a Schimdzu synthesizer model PSS8. Wang Kcore resin (dual lysine nucleus, K4) (Novabiochem) was used as the solid support for the synthesis of the branched-chain polyamino acids. The first amino acid to be conjugated was located at the C-terminus of the polypeptide sequence, and the last was located at the N-start. After all cycles of the synthesis of the branched-chain polyamino acids were completed, the branched-chain polyamino acids were deprotected from the solid support by treatment with a cleavage cocktail (trifluoroacetic acid, triisopropylsilane, and ethylene glycol) according to a standard procedure used in the prior art for the deprotection of protecting groups on the side chains of synthesized polyamino acids (Guy and Fields, Methods Emzymol 289, 67-83, 1997). For quality control of the synthesis, the individual polyamino acids were analyzed by HPLC and MALDI-TOF.
[0367] The synthetic polyamino acids X1, X2, and X5 of the SARS-CoV-2 S protein are SEQ ID NO: 100, 101, and 104, respectively. The synthetic polyamino acids X3 and X6 of the SARS-CoV-2 N protein are SEQ ID NO: 102 and 105, respectively. The synthetic polyamino acid X4 of the SARS-CoV-2 E protein is SEQ ID NO: 103. The synthetic polyamino acids X7 and X8 of the SARS-CoV-2 protein encoded by the open reading frame (ORF8) between the S and Se genes are SEQ ID NO: 106 and 107, respectively.
[0368] A gene encoding the X8 protein was found in the small open reading frame (ORF) between the S and Se genes. Genes encoding the ORF6, X4, and X5 proteins were found in the composition of the ORF between the M and N genes.
[0369] The synthetic branched polypeptides are listed in Table 21.
[0370] Table 21 - Branched peptides of SARS-CoV-2 protein synthesis
[0371]
[0372] Example 22 – Human serum sample group
[0373] The 134 serum samples were divided into 5 groups. The groups were:
[0374] Group 0: 10 serum samples (serum #1-#10) from healthy individuals obtained before 2016 from the blood bank (HEMORIO);
[0375] Group 1: 26 serum samples (#1-#26) from “asymptomatic SARS patients” as defined by the WHO case definition;
[0376] Group 2: 24 serum samples (#27-#50) from “suspected patients”;
[0377] Group 3: 38 serum samples (#51-#88) from “SARS hospitalized patients (severe illness)” as defined by the WHO case definition;
[0378] Group 4: 36 serum samples (#89-#124) from “patients immune to SARS”.
[0379] Serum samples #1-#26, high body temperature returned to normal, RT-PCR positive for SARS-CoV-2, and SD rapid test (Standard diagnostics Inc.) for anti-SARS-CoV-2 antibody was negative.
[0380] #27-#50: Patients have a positive RT-PCR result for SARS-CoV-2 and a negative SD rapid antibody test for SARS-CoV-2.
[0381] #51-#88: Hospitalized individuals with worsening signs and symptoms of SARS-CoV-2, positive for rapid RT-PCR, and positive for SD for anti-SARS-CoV-2 antibodies.
[0382] #89-#124: An immune-protected individual is defined as a recovered patient, hospitalized or outpatient, who was diagnosed as SARS-CoV-2 positive or not diagnosed as SARS-CoV-2 positive, sometimes not diagnosed by RT-PCR+ but showing characteristic symptoms.
[0383] Example 23 -Identification of SARS-CoV-2-related IgM polyamino acids
[0384] To identify potential polyamino acids that can be specifically recognized by anti-SARS-CoV-2 IgM antibodies, serum samples from infected patients were analyzed using a synthetic array of SARS-CoV-2 dotted peptide libraries including all regions of S, ORF3a, M, ORF6, ORF7, ORF8, N, E, ORF10, and control polyamino acids.
[0385] Peptide arrays were used to detect the potential binding activity of multi-amino acid sequences in human serum from infected individuals. The peptide arrays covered peptide sequences of 15 amino acid lengths with 10 amino acid overlaps between adjacent spots. Such linear multi-amino acid sequences included the following SARS-CoV-2 proteins:
[0386] - Spike protein (S): aa 1-1273 (speck A7-K19),
[0387] -ORF3a protein (ORF3): aa 1-275 (spot K22-N2),
[0388] - Membrane glycoprotein (M): aa 1-222 (spot N5-O23),
[0389] -ORF6 protein (ORF6): aa 1-61 (spots P2-P12),
[0390] -ORF7 protein (ORF7): aa 1-121 (spot P15-Q13),
[0391] -ORF8 protein (ORF8): aa 1-121 (spot Q16-R17),
[0392] - Nucleocapsid protein (N): aa 1-419 (spot R20-V17),
[0393] -Envelope protein (E): aa 1-75 (spots W1-W13)
[0394] -ORF10 protein (ORF10): aa 1-38 (spots W15-W20)
[0395] - Positive control polyamino acids: A1 and V5 (clostridium tetani precursor peptide), A2 and V6 (clostridium tetani precursor peptide), A3 and V7 (human poliovirus peptide), A4 and V8 (trihemagglutinin epitope).
[0396] - Non-reactive spots serve as negative controls.
[0397] Serological immune responses of human anti-IgM antibodies to various SARS-CoV-2 (S, ORF3a, M, ORF6, ORF7, ORF8, N, E, ORF10) synthetic polyamino acids and control polyamino acids (Examples 19 and 20) were analyzed using polyamino acids covalently synthesized on cellulose membranes (spots) and serum pools from patients #55, #60, and #74 (group 3) (n=3). Figure 12 As shown, supplemented by Table 18.
[0398] Figures 15A to 15IThe results show the signal quantification of membrane spots incubated with human serum, developed using goat anti-human IgM secondary antibody.
[0399] When detected by anti-human IgM antibody, human serum has shown significant reactivity with multiple amino acids from different viral proteins, such as... Figures 15A to 15I As shown, this demonstrates that even in the early stages of infection, a large number of different polyamino acids have great potential for disease diagnosis.
[0400] Example 24 –Identification of SARS-related IgG epitopes
[0401] To identify potential epitopes that can be specifically recognized by anti-SARS-CoV-2 IgG antibodies, serum samples from infected patients were analyzed using a synthetic array of SARS-CoV-2 dotted multiamino acid libraries including all regions of S, ORF3a, M, ORF6, ORF7, ORF8, N, E, ORF10, and control multiamino acids.
[0402] Peptide arrays were used to detect the potential binding activity of multi-amino acid sequences in human serum from infected individuals. The peptide arrays covered peptide sequences of 15 amino acid lengths with 10 amino acid overlaps between adjacent spots. Such linear multi-amino acid sequences included the following SARS-CoV-2 proteins:
[0403] -ORF3a protein (OF3a): aa 1-275 (spot A7-C11),
[0404] - Membrane glycoprotein (M): aa 1-222 (spot C14-E8),
[0405] -ORF6 protein (OF6): aa 1-61 (spots E11-E21),
[0406] -ORF7 protein (OF7(:aa 1-121(spot E24-F22),
[0407] -ORF8 protein (OF8): aa 1-121 (spots G1-G23),
[0408] - Spike protein (S): aa 1-1273 (speck H1-R13),
[0409] - Nucleocapsid protein (N): aa 1-419 (spot R16-V1),
[0410] -Envelope protein (E): aa 1-75 (spots W1-W13),
[0411] -ORF10 protein (OF10): aa 1-38 (spots W15-W20),
[0412] - Positive control peptides: A1 and V4 (clostridium tetani precursor peptide), A2 and V5 (clostridium tetani precursor peptide), A3 and V6 (human poliovirus peptide), A4 and V7 (trihemagglutinin epitope).
[0413] - No reactive spots served as negative controls
[0414] Serological immune responses to various SARS-CoV-2 (ORF3a, M, ORF6, ORF7, ORF8, S, N, E, ORF10) synthetic polyamino acids and control polyamino acids were analyzed using peptides covalently synthesized on cellulose membranes (spots) and serum pools from patients #55, #60, and #74 (group 3) (n=3). Figure 13 As shown, supplemented by Table 19.
[0415] Figures 16A to 16H This displays the signal quantification results from membrane spots incubated with human serum, developed using goat anti-human IgG secondary antibody.
[0416] When tested with anti-human IgG antibodies, human serum showed significant reactivity with multiple amino acids from different viral proteins, such as... Figures 16A to 16H As shown, this demonstrates that even in the later stages of infection, a large number of different polyamino acids have great potential for disease diagnosis.
[0417] Example 25 – Identification of SARS-related multi-amino acid IgA
[0418] To identify potential polyamino acids that can be specifically recognized by anti-SARS-CoV-2 IgA antibodies, serum samples from infected patients were analyzed using a synthetic array of SARS-CoV-2 dotted polyamino acid libraries including all regions of S, ORF3a, M, ORF6, ORF7, ORF8, N, E, ORF10, and control polyamino acids.
[0419] Peptide arrays were used to detect the potential binding activity of multi-amino acid sequences in human serum from infected individuals. The peptide arrays covered peptide sequences of 15 amino acid lengths with 10 amino acid overlaps between adjacent spots. Such linear multi-amino acid sequences included the following SARS-CoV-2 proteins:
[0420] - Spike protein (S): aa 1-1273 (speck A6-K18),
[0421] -ORF3a protein (ORF3): aa 1-275 (spot K21-N1),
[0422] - Membrane glycoprotein (M): aa 1-222(N4-O22),
[0423] -ORF6 protein (ORF6): aa 1-61 (P2-P12),
[0424] -ORF7 protein (ORF7): aa 1-121 (P15-Q13),
[0425] -ORF8 protein (ORF8): aa 1-121 (Q16-R17),
[0426] - Nucleocapsid protein (N): aa 1-419 (spot R20-V17),
[0427] -Envelope protein (E): aa 1-75 (spot W1-W-13),
[0428] -ORF10 protein (ORF10): aa 1-38 (spots W15-W20),
[0429] - Positive control peptides: A1 and V4 (Clostridium tetani precursor peptide), A2 and V5 (Clostridium tetani precursor peptide), A3 and V6 (human poliovirus peptide), A4 and V7 (trihemagglutinin epitope).
[0430] - No reactant spots serve as negative controls.
[0431] B-cell immune responses to various SARS-CoV-2 (ORF3a, M, ORF6, ORF7, ORF8, S, N, E, ORF10) synthetic polyamino acids and control polyamino acids were analyzed using peptides covalently synthesized on cellulose membranes (spots) and serum pools from patients #55, #60, and #74 (group 3) (n=3). Figure 14 As shown, supplemented by Table 20.
[0432] Figures 17A to 17I This displays the signal quantification results from membrane spots incubated with human serum, developed using goat IgG anti-human IgA secondary antibody.
[0433] When tested with anti-human IgA antibodies, human serum showed significant reactivity with multiple amino acids of different viral proteins, such as... Figures 17A to 17I As shown in the figure, a large number of different polyamino acids, primarily found in mucous membranes, have great potential for disease diagnosis through these antibodies.
[0434] Example 26 – Enzyme-linked immunosorbent assay (ELISA) for detecting antibodies against SARS-CoV-2
[0435] Enzyme-linked immunosorbent assay (ELISA) was used to screen for anti-SARS-CoV-2 antibodies in patient serum. ELISA was performed by coating 96-well polystyrene plates with 1 μg / well of branched-chain polyamino acids. Experiments were performed in parallel and simultaneously for comparison of results. To minimize potential variability in performance between assays, the reactivity index (RI), defined as the OD450 of the cutoff OD450 minus the OD450 of the target, was used. Original human serum was diluted 100-fold (100X) in 1% PBS / BSA and secondary antibodies (Merck-Sigma) and 8000x biotin-labeled goat anti-human IgG (Merck-Sigma), followed by incubation with HRP-labeled high-sensitivity neutralvidin (Thermo Fisher Scientific). A colorimetric anti-IgA response was developed using alkaline phosphatase-labeled goat anti-human IgA (KPL). Using TMB (3,3',5,5'-tetramethylbenzidine) as a substrate (Thermo Fisher Scientific), an immune response is defined as significantly enhanced when the reactivity index is greater than 1.
[0436] The results showed that the reactivity of branched-chain polyamino acids SARS-X1 to SARS-X8 to anti-SARS-CoV2 antibodies differed. These differences were correlated with the type of human antibody detected (IgM, IgG, or IgA) and the status of patients diagnosed with SARS-CoV2. Figures 18 to 23 This can be seen in the analysis. Observing such differences allows for the design of diagnostic tests that can provide more accurate or reliable information than a simple positive or negative diagnosis of anti-SARS-CoV2 antibodies. Therefore, in addition to detecting IgM, IgG, or IgA antibodies, diagnostic tests can be designed to indicate whether an individual should be hospitalized, even in the absence of symptoms.
[0437] Example 27 – Development of container proteins for SARS-CoV-2
[0438] Genes were manipulated to incorporate the “Tx” container protein with SARS-CoV-2 epitopes. Reactive epitopes were selected from sera from individuals infected with SARS-CoV-2 to construct eight Tx proteins: Ag-COVID19, Ag COVID19(H), Tx-SARS2-IgM, Tx-SARS2-IgG, Tx-SARS2-G / M, Tx-SARS2-IgA, Tx-SARS2-Universal, and Tx-SARS-G5 (non-RBD).
[0439] The genes corresponding to the Ag-COVID19, Ag COVID19(H), Tx-SARS2-IgM, Tx-SARS2-IgG, Tx-SARS2-G / M, Tx-SARS2-IgA, Tx-SARS2-Universal, and Tx-SARS-G5 (non-RBD) proteins, referred to herein as the Ag-COVID19 gene, Ag-COVID19(H) gene, Tx-SARS2-IgM gene, Tx-SARS2-IgG gene, Tx-SARS2-G / M gene, Tx-SARS2-IgA gene, Tx-SARS2-Universal gene, and Tx-SARS-G5 (non-RBD) gene, are described in nucleotide sequences SEQ ID NO:108 to SEQ ID NO:115. The amino acid sequences corresponding to the Ag-COVID19, Ag COVID19(H), Tx-SARS2-IgM, Tx-SARS2-IgG, Tx-SARS2-G / M, Tx-SARS2-IgA, Tx-SARS2-Universal and Tx-SARS-G5 (without RBD) proteins are described in SEQ ID NO 116 to 123, respectively.
[0440] Based on the SARS-CoV-2 epitope sequence mapping study, and taking into account their diagnostic potential as illustrated in Examples 19 to 26, the above eight proteins with multiple amino acids are shown in Tables 22 to 28.
[0441] Table 22 - Ag-COVID19 and Ag COVID19 Protein (H)
[0442]
[0443] Table 23 - Tx-SARS2-IgM
[0444]
[0445] Table 24 - Tx-SARS2-IgG
[0446]
[0447] Table 25 - Tx-SARS2-G / M
[0448]
[0449] Table 26 - Tx-SARS2-IgA
[0450]
[0451] Table 27-Tx-SARS2-Universal
[0452]
[0453] Table 28-Tx-SARS2-G5
[0454]
[0455] Example 28: Expression of a container protein with multiple amino acids from SARS-CoV-2
[0456] Using the pET24 plasmid containing genes encoding the respective proteins and targeting restriction sites of the BamHI and XhoI enzymes, Ag-COVID19, Ag-COVID19(H), Tx-SARS2-IgM, Tx-SARS2-IgG, Tx-SARS2-G / M, Tx-SARS2-IgA, Tx-SARS2-Universal, and Tx-SARS-G5 (non-RBD) proteins were expressed. Plasmids containing genes targeting specific proteins were then transferred to *E. coli* strain BL21 to promote the expression of the eight different proteins listed above.
[0457] The strain was grown overnight in LB medium, then inoculated into the same medium supplemented with kanamycin (30 μg / ml) and incubated on a shaker at 200 rpm until it reached a turbidity density of 0.6–0.8 (600 nm). When induced by isopropyl β-D-1-thiogalactoside (IPTG), strain BL21 expressed T7 RNA polymerase. IPTG (qsp 1 mM) was then added to the culture and the culture was maintained at 37°C for another 3 hours.
[0458] Cultures of each bacterial strain were centrifuged and the precipitates were resuspended in 10% CelLytic™ (Sigma, BR) at pH 8.0 containing 150 mM NaCl and 50 mM Tris. Recombinant protein samples (1 μg / well) were subjected to SDS-PAGE (Laemmli, Nature 227:680-685, 1970). Stacking gels (stacking gels) and separating gels (separating gels) were prepared at 4% and 11% acrylamide concentrations, respectively (Table 5, below). Samples were prepared under denaturing conditions in 62.5 mM Tris-HCl buffer, pH 6.8, 2% SDS, 5% β-mercaptoethanol, and 10% glycerol, and boiled at 95°C for 5 minutes (Hames BD, Gel electrophoresis of proteins: a practical approach. 3. Ed. Oxford. 1998). Following electrophoresis, proteins were detected by staining with Coomassie Brilliant Blue Simply Blue R250 (ThermoFisher, BR). PageRuler Plus Prestained Standards (ThermoFisher, BR) were used as a molecular weight reference. Figure 24 The image shows bands for Ag-COVID19 (column 3), Ag-COVID19(H) (column 4), Tx-SARS2-IgM (column 5), Tx-SARS2-IgG (column 6), Tx-SARS2-G / M (column 7), Tx-SARS2-IgA (column 8), Tx-SARS2-Universal (column 9), and Tx-SARS-G5 (non-RBD) (column 10). Column 1 shows the molecular weight markers: A) 250 kDa; B) 130 kDa; C) 100 kDa; D) 70 kDa; E) 55 kDa; F) 35 kDa and G) 25 kDa. Column 2 shows the total extract of uninduced bacteria.
[0459] Optionally, the culture was centrifuged and the precipitate was resuspended in urea buffer (100 mM NaH2PO4, 10 mM Tris-base, 8 M urea, pH 8.0). The solution was chromatographically analyzed using a nickel affinity column (HisTrap™, 1 mL, GE Healthcare LifeSciences) pre-equilibrated in buffer A (50 mM Tris-HCl, pH 8.0, 100 mM NaCl, and 5 mM imidazole). After binding, the resin was washed with 10 mL of buffer A. Proteins were eluted at a gradient of 0.7 mL / min in buffer B (50 mM Tris-HCl, pH 8.0, 100 mM NaCl) containing 75 mM, 200 mM, and 500 mM imidazole for 45 min. Figure 25 The pattern shown corresponds to the Ag-COVID19 protein with six histidine tails, indicating a elution concentration of 200 mM. Figure 26 The pattern corresponds to the SARS2-G5 protein purified by affinity using 200 mM imidazole (elution with 75 mM reveals contaminants).
[0460] The results showed that Tx container proteins can be readily used to generate novel and diverse proteins. Furthermore, the same expression protocol can be used to express different container proteins with varying levels of amino acids, significantly reducing input, time, and infrastructure costs. The inclusion of a six-histidine tail appears to be a potential facilitator for purification at high purity levels.
[0461] Example 29 – Enzyme-linked immunosorbent assay (ELISA) for detecting anti-SARS-CoV-2 antibodies using Ag-COVID19 and SARS2-G5 proteins.
[0462] Enzyme-linked immunosorbent assay (ELISA) was used to screen for the presence of antibodies against SARS-CoV-2. The performance of Ag-COVID19 and SARS2-G5 proteins was evaluated using serum samples from a cohort of individuals affected by SARS-CoV-2 virus infection.
[0463] By using Ag-COVID19 protein at 4°C ( Figure 27 ) or Tx-SARS2-G5 protein ( Figure 28ELISA was performed by coating 96-well polystyrene plates with a solution (0.3M urea, pH 8.0) at 1 μg / well for 12–18 hours. The wells were washed with saline-phosphate buffer (PBS) containing Tween 20 (PBS-T, 10 mM sodium phosphate-Na3PO4, 150 mM sodium chloride-NaCl, and 0.05% Tween-20, pH 7.4), and then incubated at 37°C for 2 hours with 1x PBS buffer containing 5% (w / v) skim milk powder.
[0464] Then, the wells were washed three times with PBS-T buffer and incubated for 1 hour at 37°C with human serum samples diluted 1:100 in PBS / BSA 1%. After incubation, the wells were washed three times with PBS-T and then incubated for 1 hour at 37°C with biotin-labeled goat anti-human IgG antibody diluted 1:8000 (Merck-Sigma). Subsequently, HRP-labeled high-sensitivity neutral avidin (Thermo Fisher Scientific) was added. The wells were washed three times again with PBS-T buffer and TMB substrate (3.3',5.5' tetramethylbenzidine, Thermo Fisher Scientific) was used. After 30 minutes in the dark, the absorbance was measured at 405 nm using an ELISA reader.
[0465] The results showed that, by demonstrating excellent sensitivity and specificity indices, the Ag-COVID19 protein ( Figure 27 ) and Tx-SARS2-G5 protein ( Figure 28 It has been shown to be advantageous for detecting antibodies against SARS-CoV-2. For example, from... Figure 27 and 28 As can be seen, antibodies to the protein containing multiple amino acids from SARS-CoV-2 were not detected in the serum of individuals collected before the SARS-CoV-2 epidemic, healthy individuals, or individuals affected by diseases such as dengue fever, malaria, and syphilis. In contrast, antibodies to the protein were detected in the serum of individuals diagnosed as positive for SARS-CoV-2 (symptomatic or asymptomatic), hospitalized patients, or recovered patients.
[0466] Example 30 -Ag-COVID19 protein as a vaccine composition
[0467] Ag-COVID19 protein was produced and purified according to the protocol described in this patent application. Three mice were inoculated on days 0, 14, 21, and 28 with 10 μg of Ag-COVID19 protein suspended in 25 μl of PBS in Freund's incomplete adjuvant (25 μl). Animals inoculated with PBS served as negative controls. Blood samples were collected from the animals prior to each inoculation and subjected to ELISA. Plasma was separated from the collected blood by centrifugation and serially diluted for antibody assays. Figure 29 The results showed that excellent antibodies against Ag-COVID19 were produced four weeks after the first injection.
[0468] Example 31 Use of -Ag-COVID19 protein for the purification of anti-SARS-CoV-2 antibodies
[0469] Anti-SARS-CoV-2 antibodies were purified from serum of patients diagnosed with COVID-19 using the principle of antibody affinity. Ag-COVID19 protein was conjugated to Sepharose activated by CNBr (GE Healthcare, USA). TM 4B. Dilute 10 mL of serum sample from a SARS-CoV-2 positive patient in 10 mL of PBS and subject it to Sepharose-Ag-COVID19 for 1 hour. The mixture is then placed on a chromatography column. After the solution passes through the column, 10 mL of PBS is added to the chromatography system, followed by 5 mL of 100 mM sodium citrate buffer at pH 4. Upon recovery from the column, 0.5 mL fractions are collected sequentially and the presence of antibodies is quantified spectrophotometrically at 280 nm. The absorbance of each fraction is converted to protein concentration and plotted as a function of fraction volume.
[0470] The results showed that the Ag-COVID19 protein can be usefully used as an internal control (neican) for affinity purification of antibodies from patients previously infected with SARS-CoV-2. Figure 30 This demonstrates its importance in generating available internal controls for passive immunization in response to the COVID-19 pandemic. sequence list <110> Oswald Cruz Foundation <120> Protein containers, polynucleotides, vectors, expression cassettes, cells, container manufacturing methods, pathogen identification or disease diagnosis methods, container applications, and diagnostic kits. <130> P139205 <150> BR102019017792-6 <151> August 27, 2019 <160> 163 <170> PatentIn Version 3.5 <210> 1 <211> 246 <212> PRT <213> Artificial Sequence <400> 1 Met Val Ala Ser Lys Gly Glu Glu Leu Phe Thr Gly Val Val Pro Ile 1 5 10 15 Leu Ile Glu Leu Asp Gly Asp Val Asn Gly His Lys Phe Phe Val Arg 20 25 30 Gly Glu Gly Glu Gly Asp Ala Thr Ile Gly Lys Leu Ser Leu Lys Phe 35 40 45 Ile Cys Thr Thr Gly Lys Leu Pro Val Pro Trp Pro Thr Leu Val Thr 50 55 60 Thr Leu Thr Tyr Gly Val Gln Cys Phe Ser Arg Tyr Pro Asp His Met 65 70 75 80 Lys Arg His Asp Phe Phe Lys Ser Ala Met Pro Glu Gly Tyr Val Gln 85 90 95 Glu Arg Thr Ile Tyr Phe Lys Asp Val Asp Gly Thr Tyr Lys Thr Arg 100 105 110 Ala Glu Val Lys Phe Glu Gly Thr Asp Thr Leu Val Asn Arg Ile Val 115 120 125 Leu Lys Gly Ile Asp Phe Lys Glu Asp Gly Asn Ile Leu Lys Gly His 130 135 140 Lys Leu Glu Tyr Asn Phe Asn Ser His Asn Val Tyr Ile Thr Ala Asp 145 150 155 160 Lys Gln Lys Asn Gly Ile Lys Ala Asn Phe Thr Val Arg His Asn Val 165 170 175 Glu Asp Gly Ser Val Gln Leu Ala Asp His Tyr Gln Gln Asn Thr Pro 180 185 190 Ile Gly Asp Gly Pro Val Leu Leu Pro Asp Asn His Tyr Leu Ser Thr 195 200 205 Gln Thr Ile Leu Leu Lys Asp Pro Asn Glu Leu Lys Arg Asp His Met 210 215 220 Val Leu Leu Glu Tyr Val Thr Ala Ala Gly Ile Thr Leu Gly Met Asp 225 230 235 240 Glu Leu Tyr Lys Glu Phe 245 <210> 2 <211> 744 <212> DNA <213> Artificial Sequence <400> 2 catatggtgg ctagcaaagg tgaagaactg ttcaccggtg ttgttccgat cctgattgaa 60 ctggacggtg acgttaacgg tcacaaattc tttgttcgtg gtgaaggtga aggtgacgcg 120 accatcggta aactgagtct gaaattcatc tgcaccaccg gtaagcttcc ggttccgtgg 180 ccgaccctgg ttaccaccct gacctacggt gttcagtgct tctctcgtta cccggaccac 240 atgaaacgtc acgacttctt caaatctgcg atgccggaag gttacgttca ggaacgtacc 300 atctacttca aagacgtcga cggtacttac aaaacccgtg cggaagttaa attcgaaggt 360 accgacaccc tggttaaccg tatcgttctg aaaggcattg acttcaaaga agacggtaac 420 atccttaagg gtcacaaact ggaatacaac ttcaactctc acaacgttta catcaccgcg 480 gacaaacaga aaaacggtat caaagcgaac ttcaccgttc gtcacaacgt tgaagacgga 540 tccgttcagc tggcggacca ctaccagcag aacaccccga tcggtgacgg tccggttctg 600 ctgccggaca accactacct gtctacccag accattctgc tgaaagaccc gaacgagctc 660 aaacgtgacc acatggttct gctggaatat gttacggcgg cgggtatcac cctgggtatg 720 gacgaactgt acaaagaatt ctaa 744 <210> 3 <211> 246 <212> PRT <213> Artificial sequence <400> 3 Met Val Ala Ser Lys Gly Glu Glu Leu Phe Thr Gly Val Val Pro Ile 1 5 10 15 Leu Val Glu Leu Asp Gly Asp Val Asn Gly His Lys Phe Ser Val Arg 20 25 30 Gly Glu Gly Glu Gly Asp Ala Thr Ile Gly Lys Leu Thr Leu Lys Phe 35 40 45 Ile Cys Thr Thr Gly Lys Leu Pro Val Pro Trp Pro Thr Leu Val Thr 50 55 60 Thr Leu Thr Tyr Gly Val Gln Cys Phe Ala Arg Tyr Pro Asp His Met 65 70 75 80 Lys Arg His Asp Phe Phe Lys Ser Ala Met Pro Glu Gly Tyr Val Gln 85 90 95 Glu Arg Thr Ile Ser Phe Lys Asp Val Asp Gly Lys Tyr Lys Thr Arg 100 105 110 Ala Val Val Lys Phe Glu Gly Thr Asp Thr Leu Val Asn Arg Ile Val 115 120 125 Leu Lys Gly Thr Asp Phe Lys Glu Asp Gly Asn Ile Leu Lys Gly His 130 135 140 Lys Leu Glu Tyr Asn Phe Asn Ser His Asn Val Tyr Ile Thr Ala Asp 145 150 155 160 Lys Gln Lys Asn Gly Ile Lys Ala Asn Phe Thr Val Arg His Asn Val 165 170 175 Glu Asp Gly Gly Val Gln Leu Ala Asp His Tyr Gln Gln Asn Thr Pro 180 185 190 Ile Gly Asp Gly Pro Val Leu Leu Pro Asp Asn His Tyr Leu Ser Thr 195 200 205 Gln Thr Val Leu Ser Lys Asp Pro Asn Glu Leu Lys Arg Asp His Met 210 215 220 Val Leu Leu Glu Phe Val Thr Ala Ala Gly Ile Thr Leu Gly Met Asp 225 230 235 240 Glu Leu Tyr Lys Glu Phe 245 <210> 4 <211> 753 <212> DNA <213> Artificial sequence <400> 4 catatggtgg ctagcaaggg cgaggagctg ttcaccggcg tcgtccccat cctggtcgag 60 ctggacggcg acgttaacgg ccacaagttc tccgtccggg gcgagggcga gggcgacgcc 120 accatcggca agctgaccct gaagttcatc tgcaccaccg gcaagcttcc cgtcccctgg 180 cccaccctgg tcaccaccct gacctatggc gtccagtgct tcgcccggta tcccgaccac 240 atgaagcggc acgacttctt caagtccgcc atgcccgagg gctatgtcca ggagcggacc 300 atctccttca aggacgtcga tggcaagtat aagacccggg ccgtcgtcaa gttcgagggt 360 accgacaccc tggtcaaccg gatcgtcctg aagggcaccg acttcaagga ggacggcaac 420 atccttaagg gccacaagct ggagtataac ttcaactccc acaacgtcta tatcaccgcg 480 gacaaacaga agaacggcat caaggccaac ttcaccgtcc ggcacaacgt cgaggacgga 540 tccgtccagc tggccgacca ctatcagcag aacaccccca tcggcgacgg tccggtcctg 600 ctgcccgaca accactatct gtccacccag accgtcctgt ccaaggaccc gaacgagctc 660 aaacgggacc acatggtcct gctggagttc gtcaccgccg ccggcatcac cctgggcatg 720 gacgagctgt ataaggaatt ctaatgactc gag 753 <210> 5 <211> 15 <212> DNA <213> Artificial Sequence <400> 5 catatggtgg ctagc 15 <210> 6 <211> 18 <212> DNA <213> Artificial sequence <400> 6 gaattctaat gactcgag 18 <210> 7 <211> 17 <212> PRT <213> Artificial sequence <400> 7 Lys Phe Ala Glu Leu Leu Glu Gln Gln Lys Asn Ala Gln Phe Pro Gly 1 5 10 15 Lys <210> 8 <211> 7 <212> PRT <213> Artificial sequence <400> 8 Lys Ala Ala Ala Ala Pro Ala 1 5 <210> 9 <211> 7 <212> PRT <213> Artificial sequence <400> 9 Lys Ala Ala Ile Ala Pro Ala 1 5 <210> 10 <211> 15 <212> PRT <213> Artificial sequence <400> 10 Gly Asp Lys Pro Ser Pro Phe Gly Gln Ala Ala Ala Ala Asp Lys 1 5 10 15 <210> 11 <211> 9 <212> PRT <213> Artificial sequence <400> 11 Lys Gln Lys Ala Ala Glu Ala Thr Lys 1 5 <210> 12 <211> 10 <212> PRT <213> Artificial sequence <400> 12 Ala Glu Pro Lys Pro Ala Glu Pro Lys Ser 1 5 10 <210> 13 <211> 10 <212> PRT <213> Artificial sequence <400> 13 Ala Glu Pro Lys Ser Ala Glu Pro Lys Pro 1 5 10 <210> 14 <211> 15 <212> PRT <213> Artificial sequence <400> 14 Gly Thr Ser Glu Glu Gly Ser Arg Gly Gly Ser Ser Met Pro Ser 1 5 10 15 <210> 15 <211> 11 <212> PRT <213> Artificial sequence <400> 15 Ser Pro Phe Gly Gln Ala Ala Ala Gly Asp Lys 1 5 10 <210> 16 <211> 9 <212> PRT <213> Artificial sequence <400> 16 Lys Gln Arg Ala Ala Glu Ala Thr Lys 1 5 <210> 17 <211> 1128 <212> DNA <213> Artificial sequence <400> 17 atgaaattcg cggaactgct ggaacagcag aaaaacgcgc agttcccggg taaagctagc 60 aaaggtgaag aactgttcac cggtgttgtt ccgatcctga ttgaactgga cggtgacgtt 120 aacggtcaca aattctttgt tcgtggtgaa ggtgaaggtg acgcgaccat cggtaaactg 180 agtctgaaat tcatctgcac caccggtaag cttaaagcgg cggcggcgcc ggcgaagctt 240 ccggttccgt ggccgaccct ggttaccacc ctgacctacg gtgttcagtg cttctctcgt 300 tacccggacc acatgaaacg tcacgacttc ttcaaatctg cgatgccgga aggttacgtt 360 caggaacgta ccatctactt caaagacgtc aaagcggcga tcgcgccggc ggacgtcgac 420 ggtacttaca aaacccgtgc ggaagttaaa ttcgaaggta ccggtgacaa accgtctccg 480 ttcggtcagg cggcggcggc ggacaaaggt accgacaccc tggttaaccg tatcgttctg 540 aaaggcattg acttcaaaga agacggtaac atccttaagc agaaagcggc ggaagcgacc 600 aaacttaagg gtcacaaact ggaatacaac ttcaactctc acaacgttta catcaccgcg 660 gacgcggaac cgaaaccggc ggaaccgaaa tctaccgcgg acaaacagaa aaacggtatc 720 aaagcgaact tcaccgttcg tcacaacgtt gaagacggat ccgctgaacc gaaatctgcg 780 gaaccgaaac cgggatccgt tcagctggcg gaccactacc agcagaacac cccgatcggt 840 gacggtccgg gcacctctga agaaggttct cgtggtggtt cttctatgcc gtctgacggt 900 ccggttctgc tgccggacaa ccactacctg tctacccaga ccattctgct gaaagacccg 960 aacgagctct ctccgttcgg tcaggcggcg gcgggtgaca aagagctcaa acgtgaccac 1020 atggttctgc tggaatatgt tacggcggcg ggtatcaccc tgggtatgga cgaactgtac 1080 aaagaattca aacagcgtgc ggcggaagcg accaaatgat gactcgag 1128 <210> 18 <211> 372 <212> PRT <213> Artificial Sequence <400> 18 Met Lys Phe Ala Glu Leu Leu Glu Gln Gln Lys Asn Ala Gln Phe Pro 1 5 10 15 Gly Lys Ala Ser Lys Gly Glu Glu Leu Phe Thr Gly Val Val Pro Ile 20 25 30 Leu Ile Glu Leu Asp Gly Asp Val Asn Gly His Lys Phe Phe Val Arg 35 40 45 Gly Glu Gly Glu Gly Asp Ala Thr Ile Gly Lys Leu Ser Leu Lys Phe 50 55 60 Ile Cys Thr Thr Gly Lys Leu Lys Ala Ala Ala Ala Pro Ala Lys Leu 65 70 75 80 Pro Val Pro Trp Pro Thr Leu Val Thr Thr Leu Thr Tyr Gly Val Gln 85 90 95 Cys Phe Ser Arg Tyr Pro Asp His Met Lys Arg His Asp Phe Phe Lys 100 105 110 Ser Ala Met Pro Glu Gly Tyr Val Gln Glu Arg Thr Ile Tyr Phe Lys 115 120 125 Asp Val Lys Ala Ala Ile Ala Pro Ala Asp Val Asp Gly Thr Tyr Lys 130 135 140 Thr Arg Ala Glu Val Lys Phe Glu Gly Thr Gly Asp Lys Pro Ser Pro 145 150 155 160 Phe Gly Gln Ala Ala Ala Ala Asp Lys Gly Thr Asp Thr Leu Val Asn 165 170 175 Arg Ile Val Leu Lys Gly Ile Asp Phe Lys Glu Asp Gly Asn Ile Leu 180 185 190 Lys Gln Lys Ala Ala Glu Ala Thr Lys Leu Lys Gly His Lys Leu Glu 195 200 205 Tyr Asn Phe Asn Ser His Asn Val Tyr Ile Thr Ala Asp Ala Glu Pro 210 215 220 Lys Pro Ala Glu Pro Lys Ser Thr Ala Asp Lys Gln Lys Asn Gly Ile 225 230 235 240 Lys Ala Asn Phe Thr Val Arg His Asn Val Glu Asp Gly Ser Ala Glu 245 250 255 Pro Lys Ser Ala Glu Pro Lys Pro Gly Ser Val Gln Leu Ala Asp His 260 265 270 Tyr Gln Gln Asn Thr Pro Ile Gly Asp Gly Pro Gly Thr Ser Glu Glu 275 280 285 Gly Ser Arg Gly Gly Ser Ser Met Pro Ser Asp Gly Pro Val Leu Leu 290 295 300 Pro Asp Asn His Tyr Leu Ser Thr Gln Thr Ile Leu Leu Lys Asp Pro 305 310 315 320 Asn Glu Leu Ser Pro Phe Gly Gln Ala Ala Ala Gly Asp Lys Glu Leu 325 330 335 Lys Arg Asp His Met Val Leu Leu Glu Tyr Val Thr Ala Ala Gly Ile 340 345 350 Thr Leu Gly Met Asp Glu Leu Tyr Lys Glu Phe Lys Gln Arg Ala Ala 355 360 365 Glu Ala Thr Lys 370 <210> 19 <211> 1029 <212> DNA <213> Artificial sequence <400> 19 atgctgcatg attttcgttc cgatgctagc aaaggtgaag aactgtttac cggtgttgtt 60 ccgattctgg ttgaactgga tggtgatgtt aatggccaca aattttcagt tcgtggtgaa 120 ggcgaaggtg atgcaaccat tggtaaactg accctgaaat ttatctgtac caccggtaag 180 cttccagttc cgtggccgac cctggttacc accctgacct atggtgttca gtgttttgca 240 cgttatccgg atcacatgaa acgccacgat tttttcaaaa gcgcaatgcc ggaaggttat 300 gttcaagaac gtaccattag ctttaaagac gtctgcaaac tgaaactgtg tggtgttctg 360 ggtctggacg tcgatggtaa atacaaaacc cgtgcagttg tgaaatttga gggtacctgt 420 aaactgaaac tgtgcggagt tagcggtctg ggtaccgata ccctggtgaa tcgtattgtt 480 ctgaaaggca ccgattttaa agaagatggc aacattctta agtgggttgc aatgcagacc 540 agcaatctta agggtcataa actggaatac aacttcaaca gccacaacgt gtatattacc 600 gcggatctgc acgattttca tagcgatacc gcggacaaac agaaaaatgg tattaaagcc 660 aattttaccg tgcggcataa tgttgaagat ggatcctgca aactgaaact gtgcggtgtg 720 cctggtctgg gatccgttca gctggcagat cattatcagc agaatacccc gattggtgac 780 ggtccggttg atgaacgtgg tctgtataaa gacggtccgg tgctgctgcc ggataatcat 840 tatctgagca cccagacagt tctgagcaaa gatccgaatg agctcctgca tgatctgcat 900 agtgatgagc tcaaacgtga tcacatggtt ctgctggaat ttgttaccgc agcaggtatt 960 accctgggta tggatgaact gtacaaagaa ttcaaaagtg tgcgcacctg gaacgaaatc 1020 taactcgag 1029 <210> 20 <211> 340 <212> PRT <213> Artificial Sequence <400> 20 Met Leu His Asp Phe Arg Ser Asp Ala Ser Lys Gly Glu Glu Leu Phe 1 5 10 15 Thr Gly Val Val Pro Ile Leu Val Glu Leu Asp Gly Asp Val Asn Gly 20 25 30 His Lys Phe Ser Val Arg Gly Glu Gly Glu Gly Asp Ala Thr Ile Gly 35 40 45 Lys Leu Thr Leu Lys Phe Ile Cys Thr Thr Gly Lys Leu Pro Val Pro 50 55 60 Trp Pro Thr Leu Val Thr Thr Leu Thr Tyr Gly Val Gln Cys Phe Ala 65 70 75 80 Arg Tyr Pro Asp His Met Lys Arg His Asp Phe Phe Lys Ser Ala Met 85 90 95 Pro Glu Gly Tyr Val Gln Glu Arg Thr Ile Ser Phe Lys Asp Val Cys 100 105 110 Lys Leu Lys Leu Cys Gly Val Leu Gly Leu Asp Val Asp Gly Lys Tyr 115 120 125 Lys Thr Arg Ala Val Val Lys Phe Glu Gly Thr Cys Lys Leu Lys Leu 130 135 140 Cys Gly Val Ser Gly Leu Gly Thr Asp Thr Leu Val Asn Arg Ile Val 145 150 155 160 Leu Lys Gly Thr Asp Phe Lys Glu Asp Gly Asn Ile Leu Lys Trp Val 165 170 175 Ala Met Gln Thr Ser Asn Leu Lys Gly His Lys Leu Glu Tyr Asn Phe 180 185 190 Asn Ser His Asn Val Tyr Ile Thr Ala Asp Leu His Asp Phe His Ser 195 200 205 Asp Thr Ala Asp Lys Gln Lys Asn Gly Ile Lys Ala Asn Phe Thr Val 210 215 220 Arg His Asn Val Glu Asp Gly Ser Cys Lys Leu Lys Leu Cys Gly Val 225 230 235 240 Pro Gly Leu Gly Ser Val Gln Leu Ala Asp His Tyr Gln Gln Asn Thr 245 250 255 Pro Ile Gly Asp Gly Pro Val Asp Glu Arg Gly Leu Tyr Lys Asp Gly 260 265 270 Pro Val Leu Leu Pro Asp Asn His Tyr Leu Ser Thr Gln Thr Val Leu 275 280 285 Ser Lys Asp Pro Asn Glu Leu Leu His Asp Leu His Ser Asp Glu Leu 290 295 300 Lys Arg Asp His Met Val Leu Leu Glu Phe Val Thr Ala Ala Gly Ile 305 310 315 320 Thr Leu Gly Met Asp Glu Leu Tyr Lys Glu Phe Lys Ser Val Arg Thr 325 330 335 Trp Asn Glu Ile 340 <210> 21 <211> 11 <212> PRT < / 213> Artificial sequence <400> 21 Cys Lys Leu Lys Leu Cys Gly Val Leu Gly Leu 1 5 10 <210> twenty two <211> 11 <212> PRT <213> Artificial sequence <400> twenty two Cys Lys Leu Lys Leu Cys Gly Cys Ser Gly Leu 1 5 10 <210> twenty three <211> 11 <212> PRT <213> Artificial sequence <400> twenty three Cys Lys Leu Lys Leu Cys Gly Val Pro Gly Leu 1 5 10 <210> twenty four <211> 8 <212> PRT <213> Artificial sequence <400> twenty four Val Asp Glu Arg Gly Leu Tyr Lys 1 5 <210> 25 <211> 8 <212> PRT <213> Artificial sequence <400> 25 Trp Val Ala Met Gln Thr Ser Asn 1 5 <210> 26 <211> 9 <212> PRT <213> Artificial sequence <400> 26 Lys Ser Val Arg Thr Trp Asn Glu Ile 1 5 <210> 27 <211> 7 <212> PRT <213> Artificial sequence <400> 27 Leu His Asp Phe His Ser Asp 1 5 <210> 28 <211> 7 <212> PRT <213> Artificial sequence <400> 28 Leu His Asp Phe Arg Ser Asp 1 5 <210> 29 <211> 7 <212> PRT <213> Artificial sequence <400> 29 Leu His Asp Leu His Ser Asp 1 5 <210> 30 <211> 13 <212> PRT <213> Artificial sequence <400> 30 Asp Val Leu Phe Thr Trp Tyr Val Asp Gly Thr Glu Val 1 5 10 <210> 31 <211> 344 <212> PRT <213> Artificial sequence <400> 31 Met Asp Val Leu Phe Thr Trp Tyr Val Asp Gly Thr Glu Val Ala Ser 1 5 10 15 Lys Gly Glu Glu Leu Phe Thr Gly Val Val Pro Ile Leu Val Glu Leu 20 25 30 Asp Gly Asp Val Asn Gly His Lys Phe Ser Val Arg Gly Glu Gly Glu 35 40 45 Gly Asp Ala Thr Ile Gly Lys Leu Thr Leu Lys Phe Ile Cys Thr Thr 50 55 60 Gly Lys Leu Pro Val Pro Trp Pro Thr Leu Val Thr Thr Leu Thr Tyr 65 70 75 80 Gly Val Gln Cys Phe Ala Arg Tyr Pro Asp His Met Lys Arg His Asp 85 90 95 Phe Phe Lys Ser Ala Met Pro Glu Gly Tyr Val Gln Glu Arg Thr Ile 100 105 110 Ser Phe Lys Asp Val Asp Gly Lys Tyr Lys Thr Arg Ala Val Val Lys 115 120 125 Phe Glu Gly Thr Asp Thr Leu Val Asn Arg Ile Val Leu Lys Gly Thr 130 135 140 Asp Phe Lys Glu Asp Gly Asn Ile Leu Lys Asp Val Leu Phe Thr Trp 145 150 155 160 Tyr Val Asp Gly Thr Glu Val Leu Lys Gly His Lys Leu Glu Tyr Asn 165 170 175 Phe Asn Ser His Asn Val Tyr Ile Thr Ala Asp Val Leu Phe Thr Trp 180 185 190 Tyr Val Asp Gly Thr Glu Val Gly Gly Ser Gly Thr Ala Asp Lys Gln 195 200 205 Lys Asn Gly Ile Lys Ala Asn Phe Thr Val Arg His Asn Val Glu Asp 210 215 220 Gly Ser Val Gln Leu Ala Asp His Tyr Gln Gln Asn Thr Pro Ile Gly 225 230 235 240 Asp Gly Pro Gly Gly Ser Gly Asp Val Leu Phe Thr Trp Tyr Val Asp 245 250 255 Gly Thr Glu Val Asp Gly Pro Val Leu Leu Pro Asp Asn His Tyr Leu 260 265 270 Ser Thr Gln Thr Val Leu Ser Lys Asp Pro Asn Glu Leu Asp Val Leu 275 280 285 Phe Thr Trp Tyr Val Asp Gly Thr Glu Val Gly Gly Ser Gly Glu Leu 290 295 300 Lys Arg Asp His Met Val Leu Leu Glu Phe Val Thr Ala Ala Gly Ile 305 310 315 320 Thr Leu Gly Met Asp Glu Leu Tyr Lys Glu Phe Asp Val Leu Phe Thr 325 330 335 Trp Tyr Val Asp Gly Thr Glu Val 340 <210> 32 <211> 1035 <212> DNA <213> Artificial Sequence <400> 32 atggatgttc tgtttacctg gtatgttgat ggcaccgaag tggctagcaa aggtgaagaa 60 ctgttaccg gtgttgttcc gattctggtt gaactggatg gtgatgttaa tggccacaaaa 120 ttttcagtc gtggtgaagg cgaaggtgat gcaaccattg gtaaactgac cctgaaattt 180 atctgtacca ccggtaagct tccagttccg tggccgaccc tggttaccac cctgacctat 240 ggtgttcagt gttttgcacg ttccggat cacatgaaac gccacgattt tttcaaaagc 300 gcaatgccgg aaggttatgt tcaagaacgt accattagct tcaaagacgt cgatggtaaa 360 tacaaaaccc gtgccgttgt taaatttgaa ggtaccgata ccctggtgaa tcgtattgtt 420 ctgaaaggca ccgattttaa agaagatggc aacattctta aggacgtgct gttcacatgg 480 tatgtggacg gtacagaagt tcttaagggt cacaaactgg aatacaactt taacagccac 540 aacgtgtata ttaccgcgga cgtactgttt acgtggtacg tagacggaac ggaagttggt 600 ggtagcggca ccgcggacaa acagaaaaat ggtattaaag ccaattttac cgtgcggcat 660 aatgttgaag atggatccgt tcagctggca gatcattatc agcagaatac cccgattggt 720 gacggtccgg gtggttcagg tgatgtcctg ttcacttggt acgtcgatgg aaccgaggtg 780 gacggtccgg ttctgctgcc ggataatcat tatctgagca cccagaccgt tctgagcaaa 840 gatccgaatg agctcgatgt actgttcacg tggtatgtgg atgggactga agtgggtggt 900 agtggtgagc tcaaacgtga tcacatggtg ctgctggaat ttgttaccgc agcaggtatt 960 accctgggta tggatgaact gtataaagaa ttcgatgtgc tgtttacttg gtacgtggac 1020 gggactgagg tttaa 1035 <210> 33 <211> 380 <212> PRT <213> Artificial Sequence <400> 33 Met Tyr Ile Glu Lys Asp Asp Ser Asp Ala Leu Lys Ala Leu Phe Ala 1 5 10 15 Ser Lys Gly Glu Glu Leu Phe Thr Gly Val Val Pro Ile Leu Val Glu 20 25 30 Leu Asp Gly Asp Val Asn Gly His Lys Phe Ser Val Arg Gly Glu Gly 35 40 45 Glu Gly Asp Ala Thr Ile Gly Lys Leu Thr Leu Lys Phe Ile Cys Thr 50 55 60 Thr Gly Lys Leu Pro Val Pro Trp Pro Thr Leu Val Thr Thr Leu Thr 65 70 75 80 Tyr Gly Val Gln Cys Phe Ala Arg Tyr Pro Asp His Met Lys Arg His 85 90 95 Asp Phe Phe Lys Ser Ala Met Pro Glu Gly Tyr Val Gln Glu Arg Thr 100 105 110 Ile Ser Phe Lys Asp Val Gly Asn Phe Met Val Leu Ser Val Asp Asp 115 120 125 Val Asp Gly Lys Tyr Lys Thr Arg Ala Val Val Lys Phe Glu Gly Thr 130 135 140 Gly Asn Phe Met Val Leu Ser Val Asp Asp Gly Thr Asp Thr Leu Val 145 150 155 160 Asn Arg Ile Val Leu Lys Gly Thr Asp Phe Lys Glu Asp Gly Asn Ile 165 170 175 Leu Lys Thr Ser Arg Pro Met Val Asp Leu Thr Phe Gly Gly Val Gln 180 185 190 Leu Lys Gly His Lys Leu Glu Tyr Asn Phe Asn Ser His Asn Val Tyr 195 200 205 Ile Thr Ala Asp Lys Thr Ser Arg Pro Met Val Asp Leu Thr Phe Gly 210 215 220 Gly Val Gln Thr Ala Asp Lys Gln Lys Asn Gly Ile Lys Ala Asn Phe 225 230 235 240 Thr Val Arg His Asn Val Glu Asp Gly Ser Ile Phe Asn Asp Val Pro 245 250 255 Gln Arg Thr Thr Ser Thr Phe Asp Pro Gly Ser Val Gln Leu Ala Asp 260 265 270 His Tyr Gln Gln Asn Thr Pro Ile Gly Asp Gly Pro Ile Phe Asn Asp 275 280 285 Val Pro Gln Arg Thr Thr Ser Thr Phe Asp Pro Asp Gly Pro Val Leu 290 295 300 Leu Pro Asp Asn His Tyr Leu Ser Thr Gln Thr Val Leu Ser Lys Asp 305 310 315 320 Pro Asn Glu Leu Tyr Ser Asp Leu Phe Ser Lys Asn Leu Val Thr Glu 325 330 335 Tyr Glu Leu Lys Arg Asp His Met Val Leu Leu Glu Phe Val Thr Ala 340 345 350 Ala Gly Ile Thr Leu Gly Met Asp Glu Leu Tyr Lys Glu Phe Tyr Ile 355 360 365 Glu Lys Asp Asp Ser Asp Ala Leu Lys Ala Leu Phe 370 375 380 <210> 34 <211> 1149 <212> DNA <213> Artificial Sequence <400> 34 atgtacatcg agaaagacga tagcgacgca ctgaaagcac tgttcgctag caaaggtgaa 60 gaactgttta ccggtgttgt tccgattctg gttgaactgg atggtgatgt taatggccac 120 aaattttcag ttcgtggtga aggcgaaggt gatgcaacca ttggtaaact gaccctgaaa 180 tttatctgta ccaccggtaa gcttccagtt ccgtggccga ccctggttac caccctgacc 240 tatggtgttc agtgttttgc acgttatccg gatcacatga aacgccacga ttttttcaaa 300 agcgcaatgc cggaaggtta tgttcaagaa cgtaccatta gctttaaaga cgtcggcaat 360 tttatggttc tgagcgttga tgacgtcgac ggtaaataca aaacccgtgc agttgttaaa 420 tttgaaggta ccggtaactt tatggtgctg tcagttgatg atggtaccga taccctggtg 480 aatcgtattg ttctgaaagg caccgatttt aaagaagatg gcaacattct taagaccagc 540 cgtccgatgg ttgatctgac ctttggtggt gtgcagctta agggtcataa actggaatat 600 aatttcaaca gccacaacgt gtatatcacc gcggacaaaa cctcacgtcc tatggtggac 660 ctgacattcg gtggcgttca gaccgcggac aaacagaaaa atggtattaa agccaatttt 720 accgtgcggc ataatgttga ggacggatcc atttttaacg atgttccgca gcgtaccacc 780 agtacctttg atccgggatc cgttcagctg gcagatcatt atcagcagaa taccccgatt 840 ggtgacggtc cgatctttaa tgatgtgcct cagcgcacaa cctcaacctt cgatccggac 900 ggtccggttc tgctgccgga taatcattat ctgagcaccc agaccgttct gagcaaagat 960 ccgaatgagc tctataggcga cctgtttagc aaaaatctgg ttaccgaata tgagctcaaa 1020 cgtgatcaca tggtgctgct ggaatttgtt accgcagcag gtattaccct gggtatggat 1080 gaactgtata aagaattcta tatcgaaaaa gatgattccg atgccctgaa agccctgttt 1140 taactcgag 1149 <210> 35 <211> 14 <212> PRT <213> Artificial sequence <400> 35 Tyr Ile Glu Lys Asp Asp Ser Asp Ala Leu Lys Ala Leu Phe 1 5 10 <210> 36 <211> 10 <212> PRT <213> Artificial sequence <400> 36 Gly Asn Phe Met Val Leu Ser Val Asp Asp 1 5 10 <210> 37 <211> 15 <212> PRT <213> Artificial sequence <400> 37 Lys Thr Ser Arg Pro Met Val Asp Leu Thr Phe Gly Gly Val Gln 1 5 10 15 <210> 38 <211> 15 <212> PRT <213> Artificial sequence <400> 38 Ile Phe Asn Asp Val Pro Gln Arg Thr Thr Ser Thr Phe Asp Pro 1 5 10 15 <210> 39 <211> 14 <212> PRT <213> Artificial sequence <400> 39 Leu Tyr Ser Asp Leu Phe Ser Lys Asn Leu Val Thr Glu Tyr 1 5 10 <210> 40 <211> 14 <212> PRT <213> Artificial sequence <400> 40 Tyr Ile Glu Lys Asp Asp Ser Asp Ala Leu Lys Ala Leu Phe 1 5 10 <210> 41 <211> 10 <212> PRT <213> Artificial sequence <400> 41 Lys Leu Ser Ala Thr Asp Trp Ser Ala Ile 1 5 10 <210> 42 <211> 8 <212> PRT <213> Artificial sequence <400> 42 Lys Pro Lys Pro Gln Pro Glu Lys 1 5 <210> 43 <211> 9 <212> PRT <213> Artificial sequence <400> 43 Lys Lys Met Thr Pro Ser Asp Gln Ile 1 5 <210> 44 <211> 10 <212> PRT <213> Artificial sequence <400> 44 Val Glu Leu Pro Trp Pro Leu Glu Thr Ile 1 5 10 <210> 45 <211> 340 <212> PRT <213> Artificial sequence <400> 45 Met Lys Leu Ser Ala Thr Asp Trp Ser Ala Ile Lys Gly Glu Glu Leu 1 5 10 15 Phe Thr Gly Val Val Pro Ile Leu Val Glu Leu Asp Gly Asp Val Asn 20 25 30 Gly His Lys Phe Ser Val Arg Gly Glu Gly Glu Gly Asp Ala Thr Ile 35 40 45 Gly Lys Leu Thr Leu Lys Phe Ile Cys Thr Thr Gly Lys Leu Pro Val 50 55 60 Pro Trp Pro Thr Leu Val Thr Thr Leu Thr Tyr Gly Val Gln Cys Phe 65 70 75 80 Ala Arg Tyr Pro Asp His Met Lys Arg His Asp Phe Phe Lys Ser Ala 85 90 95 Met Pro Glu Gly Tyr Val Gln Glu Arg Thr Ile Ser Phe Lys Asp Val 100 105 110 Lys Pro Lys Pro Gln Pro Glu Lys Asp Val Asp Gly Lys Tyr Lys Thr 115 120 125 Arg Ala Val Val Lys Phe Glu Gly Thr Lys Lys Met Thr Pro Ser Asp 130 135 140 Gln Ile Gly Thr Asp Thr Leu Val Asn Arg Ile Val Leu Lys Gly Thr 145 150 155 160 Asp Phe Lys Glu Asp Gly Asn Ile Leu Lys Val Glu Leu Pro Trp Pro 165 170 175 Leu Glu Thr Ile Leu Lys Gly His Lys Leu Glu Tyr Asn Phe Asn Ser 180 185 190 His Asn Val Tyr Ile Thr Ala Asp Lys Pro Lys Pro Gln Pro Glu Lys 195 200 205 Thr Ala Asp Lys Gln Lys Asn Gly Ile Lys Ala Asn Phe Thr Val Arg 210 215 220 His Asn Val Glu Asp Gly Ser Lys Leu Ser Ala Thr Asp Trp Ser Ala 225 230 235 240 Ile Gly Ser Val Gln Leu Ala Asp His Tyr Gln Gln Asn Thr Pro Ile 245 250 255 Gly Asp Gly Pro Val Glu Leu Pro Trp Pro Leu Glu Thr Ile Asp Gly 260 265 270 Pro Val Leu Leu Pro Asp Asn His Tyr Leu Ser Thr Gln Thr Val Leu [[ID=,18]]275 280 285 Ser Lys Asp Pro Asn Glu Leu Lys Lys Met Thr Pro Ser Asp Gln Ile 290 295 300 Glu Leu Lys Arg Asp His Met Val Leu Leu Glu Phe Val Thr Ala Ala 305 310 315 320 Gly Ile Thr Leu Gly Met Asp Glu Leu Tyr Lys Leu Ser Ala Thr Asp 325 330 335 Trp Ser Ala Ile 340 <210> 46 <211> 1023 <212> DNA <213> Artificial sequence <400> 46 atgaaactga gcgcaaccga ttggagcgca attaaaggtg aagaactgtt taccggtgtt 60 gttccgattc tggttgaact ggatggtgat gttaatggcc acaaattttc agttcgtggt 120 gaaggcgaag gtgatgcaac cattggtaaa ctgaccctga aatttatctg taccaccggt 180 aagcttccag ttccgtggcc gaccctggtt accaccctga cctatggtgt tcagtgtttt 240 gcacgttatc cggatcacat gaaacgccac gattttttca aaagcgcaat gccggaaggt 300 tatgttcaag aacgtaccat tagcttcaaa gacgtcaaac cgaaaccgca gccggaaaaa 360 gacgtcgatg gtaaatacaa aacccgtgcc gttgttaaat ttgaaggtac caaaaaaatg 420 accccgagcg atcagattgg taccgatacc ctggtgaatc gtattgttct gaaaggcacc 480 gattttaaag aagatggcaa catccttaag gtggaactgc cgtggcctct ggaaaccatt 540 cttaagggtc ataaactgga atataatttc aacagccaca acgtgtatat caccgcggac 600 aaacctaaac ctcaacctga aaaaaccgcg gacaaacaga aaaatggcat caaagcaaat 660 tttaccgtgc gccataatgt tgaggatgga tccaaactgt cagccaccga ttggtcagca 720 attggatccg ttcagctggc agatcattat cagcagaata ccccgattgg tgacggtccg 780 gttgagctgc cttggccact ggaaacaatt gacggtccgg tgctgctgcc ggataatcat 840 tatctgagca cccagaccgt tctgagcaaa gatccgaatg agctcaaaaa aatgacaccg 900 tcagatcaga tcgagctcaa acgtgatcac atggttctgc tggaatttgt taccgcagca 960 ggtattaccc tgggtatgga tgaactgtat aaactgagtg cgacagactg gtctgcaatc 1020 taa 1023 <210> 47 <211> 9 <212> PRT <213> Artificial sequence <400> 47 His Arg Ile Arg Leu Leu Leu Gln Ser 1 5 <210> 48 <211> 9 <212> PRT <213> Artificial sequence <400> 48 Ser Tyr Arg Thr Gly Ala Glu Arg Val 1 5 <210> 49 <211> 9 <212> PRT <213> Artificial sequence <400> 49 Asn Gly Val Lys Gln Thr Val Asp Val 1 5 <210> 50 <211> 9 <212> PRT <213> Artificial sequence <400> 50 Gln Ser Arg Thr Leu Asp Ser Arg Asp 1 5 <210> 51 <211> 338 <212> PRT <213> Artificial sequence <400> 51 Met His Arg Ile Arg Leu Leu Leu Gln Ser Lys Gly Glu Glu Leu Phe 1 5 10 15 Thr Gly Val Val Pro Ile Leu Val Glu Leu Asp Gly Asp Val Asn Gly 20 25 30 His Lys Phe Ser Val Arg Gly Glu Gly Glu Gly Asp Ala Thr Ile Gly 35 40 45 Lys Leu Thr Leu Lys Phe Ile Cys Thr Thr Gly Lys Leu Pro Val Pro 50 55 60 Trp Pro Thr Leu Val Thr Thr Leu Thr Tyr Gly Val Gln Cys Phe Ala 65 70 75 80 Arg Tyr Pro Asp His Met Lys Arg His Asp Phe Phe Lys Ser Ala Met 85 90 95 Pro Glu Gly Tyr Val Gln Glu Arg Thr Ile Ser Phe Lys Asp Val Ser 100 105 110 Tyr Arg Thr Gly Ala Glu Arg Val Asp Val Asp Gly Lys Tyr Lys Thr 115 120 125 Arg Ala Val Val Lys Phe Glu Gly Thr Asn Gly Val Lys Gln Thr Val 130 135 140 Asp Val Gly Thr Asp Thr Leu Val Asn Arg Ile Val Leu Lys Gly Thr 145 150 155 160 Asp Phe Lys Glu Asp Gly Asn Ile Leu Lys Gln Ser Arg Thr Leu Asp 165 170 175 Ser Arg Asp Leu Lys Gly His Lys Leu Glu Tyr Asn Phe Asn Ser His 180 185 190 Asn Val Tyr Ile Thr Ala Asp Gln Ser Arg Thr Leu Asp Ser Arg Asp 195 200 205 Thr Ala Asp Lys Gln Lys Asn Gly Ile Lys Ala Asn Phe Thr Val Arg 210 215 220 His Asn Val Glu Asp Gly Ser His Arg Ile Arg Leu Leu Leu Gln Ser 225 230 235 240 Gly Ser Val Gln Leu Ala Asp His Tyr Gln Gln Asn Thr Pro Ile Gly 245 250 255 Asp Gly Pro Ser Tyr Arg Thr Gly Ala Glu Arg Val Asp Gly Pro Val 260 265 270 Leu Leu Pro Asp Asn His Tyr Leu Ser Thr Gln Thr Val Leu Ser Lys 275 280 285 Asp Pro Asn Glu Leu Asn Gly Val Lys Gln Thr Val Asp Val Glu Leu 290 295 300 Lys Arg Asp His Met Val Leu Leu Glu Phe Val Thr Ala Ala Gly Ile 305 310 315 320 Thr Leu Gly Met Asp Glu Leu Tyr Lys His Arg Ile Arg Leu Leu Leu 325 330 335 Gln Ser <210> 52 <211> 1017 <212> DNA <213> Artificial sequence <400> 52 atgcatcgta ttcgtctgct gctgcagagc aaaggtgaag aactgtttac cggtgttgtt 60 ccgattctgg ttgaactgga tggtgatgtt aatggccaca aattttcagt tcgtggtgaa 120 ggcgaaggtg atgcaaccat tggtaaactg accctgaaat ttatctgtac caccggtaag 180 cttccagttc cgtggccgac cctggttacc accctgacct atggtgttca gtgttttgca 240 cgttatccgg atcacatgaa acgccacgat tttttcaaaa gcgcaatgcc ggaaggttat 300 gttcaagaac gtaccattag ctttaaagac gtcagctatc gtaccggtgc agaacgtgtt 360 gacgtcgatg gtaaatacaa aacccgtgcc gttgttaaat ttgaaggtac caatggcgtt 420 aaacagaccg ttgatgttgg taccgatacc ctggtgaatc gtattgttct gaaaggcacc 480 gattttaaag aagatggcaa cattcttaag cagagccgta ccctggatag ccgtgatctt 540 aagggtcata aactggaata taatttcaac agccacaacg tgtatattac cgcggaccag 600 tcacgcaccc tggattcacg tgacaccgcg gacaaacaga aaaatggtat taaagccaat 660 tttaccgtgc ggcataatgt tgaagatgga tcccatcgta tccgcctgct gctgcaaagc 720 ggatccgttc agctggcaga tcattatcag cagaataccc cgattggtga cggtccgagt 780 tatcgcacag gtgccgaacg cgtggacggt ccggttctgc tgccggataa tcattatctg 840 agcacccaga ccgttctgag caaagatccg aatgagctca atggtgtgaa acaaacagtg 900 gatgtagagc tcaaacgtga tcacatggtg ctgctggaat ttgttaccgc agcaggtatt 960 accctgggta tggatgaact gtataaacac cgcattcggc tgctgctgca atcataa 1017 <210> 53 <211> 14 <212> PRT <213> Artificial Sequence <400> 53 Pro Tyr Thr Ser Arg Arg Ser Val Ala Ser Ile Val Gly Thr 1 5 10 <210> 54 <211> 10 <212> PRT <213> Artificial sequence <400> 54 Gln Tyr Tyr Asp Tyr Glu Asp Ala Thr Phe 1 5 10 <210> 55 <211> 10 <212> PRT <213> Artificial sequence <400> 55 Gly Pro Lys Gln Leu Thr Phe Glu Gly Lys 1 5 10 <210> 56 <211> 10 <212> PRT <213> Artificial sequence <400> 56 Asp Ala Thr Phe Glu Thr Tyr Ala Leu Thr 1 5 10 <210> 57 <211> 9 <212> PRT <213> Artificial sequence <400> 57 Leu Thr Val Glu Asp Ser Pro Tyr Pro 1 5 <210> 58 <211> 9 <212> PRT <213> Artificial sequence <400> 58 Ala Leu Ala Thr Tyr Gln Ser Glu Tyr 1 5 <210> 59 <211> 14 <212> PRT <213> Artificial sequence <400> 59 Pro Gly Ile Val Ile Pro Pro Lys Ala Leu Phe Thr Gln Gln 1 5 10 <210> 60 <211> 9 <212> PRT <213> Artificial sequence <400> 60 Ala Val Glu Ala Glu Arg Ala Gly Arg 1 5 <210> 61 <211> 9 <212> PRT <213> Artificial sequence <400> 61 Thr Thr Thr Glu Tyr Ser Asn Ala Arg 1 5 <210> 62 <211> 14 <212> PRT <213> Artificial sequence <400> 62 Glu Arg Ala Gly Glu Ala Met Val Leu Val Tyr Tyr Glu Ser 1 5 10 <210> 63 <211> 1101 <212> DNA <213> Artificial sequence <400> 63 atgccctaca cctcacgtcg cagcgtagct tcgattgttg gcacgtctta ttggaaaggg 60 tctcaatact atgactatga agatgcaact tttaaaggag aggagttgtt taccggcgtg 120 gtgccgatcc ttgtggagtt ggatggagac gttaacggtc acaagttttc agttcgcggt 180 gaggcgag gcgacgcgac tattggtaag cttaccttga agttcatctg tactacgggg cctaagcagc ttacgttcga ggggaatta cctgtcccct ggccaacatt ggtcacgaca ttaacctatg gcgtgcagtg ttttgcgcgt taccccgatc acatgaagcg tcacgacttc 360 tttaagtccg cgatgccgga agggtatgtc caggaacgca cgatttcgtt caaagacgcg acctttgaga cttacgcatt aaccggggggc tcggggcgagg cagccgcaaa agaagctgct 480 gctgacggca aatacaaaac tcgtgcggtc gtaaaattcg agggagacac tttagttaat cgtattgtgc tgaaagggac tgacttcaag gaagatgga acattttgac ggttgaggat agcccctatc ctggccataa acttgagtac aacttcaatt cacataatgt ctacattaca gcacttgcaa cgtatcagtc tgagtacgat aagcagaaga atggtatcaa ggctaatttt acggtccgtc ataatgttga ggatggggagc gtgcagttag ctgaccatta tcaacaaaat acgcccattg gggaccccgg catcgtgatt cccccaaaag cattattcac ccagcaagtg cttttaccag acaaccatta cttgagcacc caaacggtgt taagcaagga ccccaatgag 900 aaagcagtgg aggcggaacg cgctggacgt gatcatatgg tacttcttga attcgtaacg 960 gcggccggga tcactttggg aatggacgag ttatataaag agcgtgcggg tgaggccatg 1020 gtgcttgtgt attacgagtc agaagccgcg aaggaggctg ccaagacaac aacggaatac 1080 agtaacgccc gtaaaaagta a 1101 <210> 64 <211> 366 <212> PRT <213> Artificial Sequence <400> 64 Met Pro Tyr Thr Ser Arg Arg Ser Val Ala Ser Ile Val Gly Thr Ser 1 5 10 15 Tyr Trp Lys Gly Ser Gln Tyr Tyr Asp Tyr Glu Asp Ala Thr Phe Lys 20 25 30 Gly Glu Glu Leu Phe Thr Gly Val Val Pro Ile Leu Val Glu Leu Asp 35 40 45 Gly Asp Val Asn Gly His Lys Phe Ser Val Arg Gly Glu Gly Glu Gly 50 55 60 Asp Ala Thr Ile Gly Lys Leu Thr Leu Lys Phe Ile Cys Thr Thr Gly 65 70 75 80 Pro Lys Gln Leu Thr Phe Glu Gly Lys Leu Pro Val Pro Trp Pro Thr 85 90 95 Leu Val Thr Thr Leu Thr Tyr Gly Val Gln Cys Phe Ala Arg Tyr Pro 100 105 110 Asp His Met Lys Arg His Asp Phe Phe Lys Ser Ala Met Pro Glu Gly 115 120 125 Tyr Val Gln Glu Arg Thr Ile Ser Phe Lys Asp Ala Thr Phe Glu Thr 130 135 140 Tyr Ala Leu Thr Gly Gly Ser Gly Glu Ala Ala Ala Lys Glu Ala Ala 145 150 155 160 Ala Asp Gly Lys Tyr Lys Thr Arg Ala Val Val Lys Phe Glu Gly Asp 165 170 175 Thr Leu Val Asn Arg Ile Val Leu Lys Gly Thr Asp Phe Lys Glu Asp 180 185 190 Gly Asn Ile Leu Thr Val Glu Asp Ser Pro Tyr Pro Gly His Lys Leu 195 200 205 Glu Tyr Asn Phe Asn Ser His Asn Val Tyr Ile Thr Ala Leu Ala Thr 210 215 220 Tyr Gln Ser Glu Tyr Asp Lys Gln Lys Asn Gly Ile Lys Ala Asn Phe 225 230 235 240 Thr Val Arg His Asn Val Glu Asp Gly Ser Val Gln Leu Ala Asp His 245 250 255 Tyr Gln Gln Asn Thr Pro Ile Gly Asp Pro Gly Ile Val Ile Pro Pro 260 265 270 Lys Ala Leu Phe Thr Gln Gln Val Leu Leu Pro Asp Asn His Tyr Leu 275 280 285 Ser Thr Gln Thr Val Leu Ser Lys Asp Pro Asn Glu Lys Ala Val Glu 290 295 300 Ala Glu Arg Ala Gly Arg Asp His Met Val Leu Leu Glu Phe Val Thr 305 310 315 320 Ala Ala Gly Ile Thr Leu Gly Met Asp Glu Leu Tyr Lys Glu Arg Ala 325 330 335 Gly Glu Ala Met Val Leu Val Tyr Tyr Glu Ser Glu Ala Ala Lys Glu 340 345 350 Ala Ala Lys Thr Thr Thr Glu Tyr Ser Asn Ala Arg Lys Lys 355 360 365 <210> 65 <211> 14 <212> PRT <213> Artificial Sequence <400> 65 Ser Pro Trp Ser Trp Pro Asp Leu Asp Leu Lys Pro Gly Ala 1 5 10 <210> 66 <211> 14 <212> PRT <213> Artificial sequence <400> 66 Asp Gly Asn Cys Asp Gly Arg Gly Lys Ser Thr Arg Ser Thr 1 5 10 <210> 67 <211> 14 <212> PRT <213> Artificial sequence <400> 67 Val Phe Ser Pro Gly Arg Lys Asn Gly Ser Phe Ile Ile Asp 1 5 10 <210> 68 <211> 14 <212> PRT <213> Artificial sequence <400> 68 His Val Gln Asp Cys Asp Glu Ser Val Leu Thr Arg Leu Glu 1 5 10 <210> 69 <211> 14 <212> PRT <213> Artificial sequence <400> 69 Asp Cys Asp Gly Ser Ile Leu Gly Ala Ala Val Asn Gly Lys 1 5 10 <210> 70 <211> 9 <212> PRT <213> Artificial sequence <400> 70 Phe Thr Thr Arg Val Tyr Met Asp Ala 1 5 <210> 71 <211> 14 <212> PRT <213> Artificial sequence <400> 71 Arg Asp Ser Asp Asp Trp Leu Asn Lys Tyr Ser Tyr Tyr Pro 1 5 10 <210> 72 <211> 14 <212> PRT <213> Artificial sequence <400> 72 Glu Ser Glu Met Phe Met Pro Arg Ser Ile Gly Gly Pro Val 1 5 10 <210> 73 <211> 14 <212> PRT <213> Artificial sequence <400> 73 Ala Glu Ala Glu Met Val Ile His His Gln His Val Gln Asp 1 5 10 <210> 74 <211> 19 <212> PRT <213> Artificial sequence <400> 74 Leu Glu His Glu Met Trp Arg Ser Arg Ala Asp Glu Ile Asn Ala Ile 1 5 10 15 Phe Glu Glu <210> 75 <211> 430 <212> PRT <213> Artificial sequence <400> 75 Met Gly Ser Ser His His His His His Ser Ser Gly Leu Val Pro 1 5 10 15 Arg Gly Ser His Met Ser Pro Trp Ser Trp Pro Asp Leu Asp Leu Lys 20 25 30 Pro Gly Ala Lys Gly Glu Glu Leu Phe Thr Gly Val Val Pro Ile Leu 35 40 45 Val Glu Leu Asp Gly Asp Val Asn Gly His Lys Phe Ser Val Arg Gly 50 55 60 Glu Gly Glu Gly Asp Ala Thr Ile Gly Lys Leu Thr Leu Lys Phe Ile 65 70 75 80 Cys Thr Thr Gly Lys Leu Pro Val Pro Trp Pro Thr Leu Val Thr Thr 85 90 95 Leu Thr Tyr Gly Val Gln Cys Phe Ala Arg Tyr Pro Asp His Met Lys 100 105 110 Arg His Asp Phe Phe Lys Ser Ala Met Pro Glu Gly Tyr Val Gln Glu 115 120 125 Arg Thr Ile Ser Phe Lys Asp Gly Asn Cys Asp Gly Arg Gly Lys Ser 130 135 140 Thr Arg Ser Thr Gly Gly Ser Gly Glu Ala Ala Ala Lys Glu Ala Ala 145 150 155 160 Ala Lys Lys Tyr Lys Thr Arg Ala Val Val Lys Phe Glu Gly Val Phe 165 170 175 Ser Pro Gly Arg Lys Asn Gly Ser Phe Ile Ile Asp Gly Gly Ser Gly 180 185 190 Glu Ala Ala Ala Lys Glu Ala Ala Ala Lys Asp Thr Leu Val Asn Arg 195 200 205 Ile Val Leu Lys Gly Thr Asp Phe Lys Glu Asp Gly Asn Ile Leu Gly 210 215 220 His Lys His Val Gln Asp Cys Asp Glu Ser Val Leu Thr Arg Leu Glu 225 230 235 240 Tyr Asn Phe Asn Ser His Asn Val Tyr Ile Thr Ala Asp Cys Asp Gly 245 250 255 Ser Ile Leu Gly Ala Ala Val Asn Gly Lys Gln Lys Asn Gly Ile Lys 260 265 270 Ala Asn Phe Thr Val Arg His Asn Val Glu Asp Phe Thr Thr Arg Val 275 280 285 Tyr Met Asp Ala Gly Thr Ser Trp Lys Gly Gly Ser Val Gln Leu Ala 290 295 300 Asp His Tyr Gln Gln Asn Thr Pro Ile Gly Asp Arg Asp Ser Asp Asp 305 310 315 320 Trp Leu Asn Lys Tyr Ser Tyr Tyr Pro Val Leu Leu Pro Asp Asn His 325 330 335 Tyr Leu Ser Thr Gln Thr Val Leu Ser Lys Asp Pro Asn Glu Ser Glu 340 345 350 Met Phe Met Pro Arg Ser Ile Gly Gly Pro Val Lys Arg Asp His Met 355 360 365 Val Leu Leu Glu Phe Val Thr Ala Ala Gly Ile Thr Leu Gly Met Asp 370 375 380 Glu Leu Tyr Lys Leu Glu His Glu Met Trp Arg Ser Arg Ala Asp Glu 385 390 395 400 Ile Asn Ala Ile Phe Glu Glu Thr Ser Tyr Trp Lys Gly Ser Ala Glu 405 410 415 Ala Glu Met Val Ile His His Gln His Val Gln Asp Lys Lys 420 425 430 <210> 76 <211> 1293 <212> DNA <213> Artificial sequence <400> 76 atgggcagca gccatcatca tcatcatcac agcagcggcc tggtgccgcg cggcagccat 60 atgtccccat ggagttggcc tgaccttgat ttaaagcccg gtgctaaagg agaagaactg 120 ttcacagggg tcgttcccat cttagtggag ctggacggcg acgtgaacgg ccataaattc 180 agtgtgcgcg gggaaggaga agggacgca acgattggta agttgacgct gaaatttatc 240 tgtacaactg gtaagcttcc agtgccgtgg ccgacgcttg tgacgacact tacttatgga 300 gtccagtgtt ttgcgcgtta tccagatcac atgaacgcc acgactctt taagtccgca 360 atgccggagg gctacgtgca ggaacgcaca atctcgttta aggacgggaa ctgcgatggc 420 cgtgggaaaa gtacacgctc aacgggcgga tcaggcgaag ccgcagcaaa agaggctgcg 480 gccaaaaaat ataaactcg tgccgtagtg aaatttgagg gagtgttttc cccgggccgt 540 aagaacggca gtttatcat tgacggcggc tcaggagag cggctgcaa ggaggccgcc 600 gcgaaggaca ccttagtgaa cccattgtt tgaaggaa cggattttaa agaggacgga 660 aacattttag gatacaagca tgtccaagac tgcgatgaga gtgttcttac gcgtttagaa 720 tataatttca atagtcataa cgtatacatt acggctgatt gtgatggcag tatcttagga 780 gcggctgtta acgggaagca aaagaatgga atcaaagcca actttactgt tcgtcataat 840 gtcgaagact tcaccactcg tgtctacatg gatgccggta ctagctggaa aggcggtcc 900 gttcagcttg cggatcatta ccaacaaaac actccgatcg gggaccgtga ctccgatgat 960 tggctgaaca agtattcata ctaccctgtg ctgctgcccg ataaccatta tttgagcaca 1020 cagaccgtgc tgtctaaaga tcctaatgaa tcagaaatgt ttatgcctcg ttcgatcggg 1080 ggacccgtta aacgcgatca catggtattg ttggaattcg ttacggcagc gggcatcacg 1140 cttggtatgg atgagttgta taaacttgag catgaaatgt ggcgctcgcg tgcagacgaa 1200 atcaatgcaa tttttgaaga gactagctat tggaagggta gtgcggaagc tgaaatggtc 1260 attcaccacc agcatgtcca ggacaaaaaa taa 1293 <210> 77 <211> 245 <212> PRT <213> Artificial Sequence <400> 77 Met Gly Ala His Ala Ser Val Ile Lys Pro Glu Met Lys Ile Lys Leu 1 5 10 15 Arg Met Glu Gly Ala Val Asn Gly His Lys Phe Val Ile Glu Gly Glu 20 25 30 Gly Ile Gly Lys Pro Tyr Glu Gly Thr Gln Thr Leu Asp Leu Thr Val 35 40 45 Glu Glu Gly Ala Pro Leu Pro Phe Ser Tyr Asp Ile Leu Thr Pro Ala 50 55 60 Phe Gln Tyr Gly Asn Arg Ala Phe Thr Lys Tyr Pro Glu Asp Ile Pro 65 70 75 80 Asp Tyr Phe Lys Gln Ala Phe Pro Glu Gly Tyr Ser Trp Glu Arg Ser 85 90 95 Met Thr Tyr Glu Asp Gln Gly Ile Cys Ile Ala Thr Ser Asp Ile Thr 100 105 110 Met Glu Gly Asp Cys Phe Phe Tyr Glu Ile Arg Phe Asp Gly Thr Asn 115 120 125 Phe Pro Pro Asn Gly Pro Val Met Gln Lys Lys Thr Leu Lys Trp Glu 130 135 140 Pro Ser Thr Glu Lys Met Tyr Val Glu Asp Gly Val Leu Lys Gly Asp 145 150 155 160 Val Glu Met Ala Leu Leu Leu Glu Gly Gly Gly His Tyr Arg Cys Asp 165 170 175 Phe Lys Thr Thr Tyr Lys Ala Lys Lys Asp Val Arg Leu Pro Asp Ala 180 185 190 His Glu Val Asp His Arg Ile Glu Ile Leu Ser His Asp Lys Asp Tyr 195 200 205 Asn Lys Val Arg Leu Tyr Glu His Ala Glu Ala Arg Tyr Ser Gly Gly 210 215 220 Gly Ser Gly Gly Gly Ala Ser Gly Lys Pro Ile Pro Asn Pro Leu Leu 225 230 235 240 Gly Leu Asp Ser Thr 245 <210> 78 <211> 738 <212> DNA <213> Artificial sequence <400> 78 atgggtgcgc acgcgtctgt tatcaaaccg gaaatgaaaa tcaaactgcg tatggaaggt 60 gcggttaacg gtcacaaatt cgttatcgaa ggtgaaggta tcggtaaacc gtacgaaggt 120 acccagaccc tggacctgac cgttgaagaa ggtgcgccgc tgccgttctc ttacgacatc 180 ctgaccccgg cgttccagta cggtaaccgt gcgttcacca aatacccgga agacatcccg 240 gactacttca aacaggcgtt cccggaaggt tactcttggg aacgttctat gacctacgaa 300 gaccagggta tctgcatcgc gacctctgac atcaccatgg aaggtgactg cttcttctac 360 gaaatccgtt tcgacggtac caacttcccg ccgaacggtc cggttatgca gaaaaaaacc 420 ctgaaatggg aaccgtctac cgaaaaaatg tacgttgaag acggtgttct gaaaggtgac 480 gttgaaatgg cgctgctgct ggaaggtggt ggtcactacc gttgcgactt caaaaccacc 540 tacaaagcga aaaaagacgt tcgtctgccg gacgcgcacg aagttgacca ccgtatcgaa 600 atcctgtctc acgacaaaga ctacaacaaa gttcgtctgt acgaacacgc ggaagcgcgt 660 tactctggtg gtggttctgg tggtggtgcg tctggtaaac cgatcccgaa cccgctgctg 720 ggtctggact ctacctaa 738 <210> 79 <211> 20 <212> PRT <213> Artificial sequence <400> 79 Asp Leu Arg Gln Met Arg Thr Val Thr Pro Ile Arg Met Gln Gly Gly 1 5 10 15 Cys Gly Ser Cys 20 <210> 80 <211> 20 <212> PRT <213> Artificial sequence <400> 80 Gly Cys Gly Ser Cys Trp Ala Phe Ser Gly Val Ala Ala Thr Glu Ser 1 5 10 15 Ala Tyr Leu Ala 20 <210> 81 <211> 15 <212> PRT <213> Artificial sequence <400> 81 Gln Glu Ser Tyr Tyr Arg Tyr Val Ala Arg Glu Gln Ser Cys Arg 1 5 10 15 <210> 82 <211> 15 <212> PRT <213> Artificial sequence <400> 82 His Ala Val Asn Ile Val Gly Tyr Ser Asn Ala Gln Gly Val Asp 1 5 10 15 <210> 83 <211> 20 <212> PRT <213> Artificial sequence <400> 83 Cys His Gly Ser Glu Pro Cys Ile Ile His Arg Gly Lys Pro Phe Gln 1 5 10 15 Leu Glu Ala Val 20 <210> 84 <211> 20 <212> PRT <213> Artificial sequence <400> 84 Tyr Asp Ile Lys Tyr Thr Trp Asn Val Pro Lys Ile Ala Pro Lys Ser 1 5 10 15 Glu Asn Val Val 20 <210> 85 <211> 15 <212> PRT <213> Artificial sequence <400> 85 Asn Thr Lys Thr Ala Lys Ile Glu Ile Lys Ala Ser Ile Asp Gly 1 5 10 15 <210> 86 <211> 15 <212> PRT <213> Synthetic sequence <400> 86 Gly Val Leu Ala Cys Ala Ile Ala Thr His Ala Lys Ile Arg Asp 1 5 10 15 <210> 87 <211> 20 <212> PRT <213> Synthetic sequence <400> 87 Pro Lys Asp Pro His Lys Phe Tyr Ile Cys Ser Asn Trp Glu Ala Val 1 5 10 15 His Lys Asp Cys 20 <210> 88 <211> 366 <212> PRT <213> Synthetic sequence <400> 88 Met Pro Lys Asp Pro His Lys Phe Tyr Ile Cys Ser Asn Trp Glu Ala 1 5 10 15 Val His Lys Asp Cys Ser Val Ile Lys Pro Glu Met Lys Ile Lys Leu 20 25 30 Arg Met Glu Gly Ala Val His Ala Val Asn Ile Val Gly Tyr Ser Asn 35 40 45 Ala Gln Gly Val Asp His Lys Phe Val Ile Glu Gly Glu Gly Ile Gly 50 55 60 Lys Pro Tyr Glu Gly Thr Gln Thr Leu Asp Leu Thr Val Glu Glu Gly 65 70 75 80 Ala Pro Leu Pro Phe Ser Tyr Asp Ile Leu Thr Pro Ala Phe Gln Tyr 85 90 95 Gly Asn Arg Ala Phe Thr Lys Tyr Pro Glu Asp Ile Pro Asp Tyr Phe 100 105 110 Lys Gln Ala Phe Pro Glu Gly Tyr Ser Trp Glu Arg Ser Met Thr Tyr 115 120 125 Gly Val Leu Ala Cys Ala Ile Ala Thr His Ala Lys Ile Arg Asp Gly 130 135 140 Ile Cys Ile Ala Thr Ser Asp Ile Thr Met Glu Tyr Asp Ile Lys Tyr 145 150 155 160 Thr Trp Asn Val Pro Lys Ile Ala Pro Lys Ser Glu Asn Val Val Cys 165 170 175 Phe Phe Tyr Glu Ile Arg Phe Asp Gly Thr Gln Glu Ser Tyr Tyr Arg 180 185 190 Tyr Val Ala Arg Glu Gln Ser Cys Arg Thr Leu Lys Trp Glu Pro Ser 195 200 205 Thr Glu Lys Met Tyr Val Glu Asp Gly Val Leu Lys Gly Asp Val Glu 210 215 220 Met Ala Leu Leu Leu Cys His Gly Ser Glu Pro Cys Ile Ile His Arg 225 230 235 240 Gly Lys Pro Phe Gln Leu Glu Ala Val His Tyr Arg Cys Asp Phe Lys 245 250 255 Thr Thr Tyr Lys Ala Gly Cys Gly Ser Cys Trp Ala Phe Ser Gly Val 260 265 270 Ala Ala Thr Glu Ser Ala Tyr Leu Ala His Glu Val Asp His Arg Ile 275 280 285 Glu Ile Leu Ser His Asp Lys Asp Tyr Asn Lys Val Arg Leu Tyr Glu 290 295 300 His Ala Glu Ala Arg Tyr Ser Gly Gly Ser Gly Asp Leu Arg Gln Met 305 310 315 320 Arg Thr Val Thr Pro Ile Arg Met Gln Gly Gly Cys Gly Ser Cys Gly 325 330 335 Gly Ser Gly Asn Thr Lys Thr Ala Lys Ile Glu Ile Lys Ala Ser Ile 340 345 350 Asp Gly Leu Asp Ser Thr His His His His His His Lys Lys 355 360 365 <210> 89 <211> 1101 <212> DNA <213> Artificial sequence <400> 89 atgccgaaag atccccacaa gttctatatt tgctcaaact gggaggccgt gcacaaagat 60 tgttctgtta tcaagccaga gatgaaaatc aaactgcgca tggagggggc tgtccatgcg 120 gtaaacattg tgggttactc aaatgcgcag ggagtggacc ataagtttgt catcgaaggc 180 gagggcatcg gtaaacctta tgaaggcact cagactttgg atttaacggt cgaagaagga 240 gcgcctttgc cattttcgta tgatatctta accccggcgt ttcaatacgg aaatcgcgcg 300 ttcacaaagt atcccgaaga catccccgat tacttcaagc aagcttttcc agaaggctac 360 agctgggagc gttcaatgac gtatggagtc cttgcatgtg ctatcgccac acatgctaaa 420 atccgcgatg gtatttgtat tgccacgtca gatatcacaa tggaatacga tatcaagtac 480 acttggaacg tgccaaaaat cgctcccaag tcagaaaacg tggtgtgttt cttctatgag 540 attcgttttg atggtaccca agagtcttat tatcgttatg ttgcacgtga acaaagttgt 600 cgtacgttga agtgggagcc ttctacagaa aaaatgtatg tggaggacgg cgtgcttaaa 660 ggagacgtgg agatggcact gcttttgtgt catgggtccg aaccgtgtat tattcatcgc 720 ggagacgtgg agatggcact gcttttgtgt catgggtccg aaccgtgtat tattcatcgc 720 ggaaaaccct ttcaattgga agcggtccac taccgctgcg attttaagac cacatataaa 780 ggaaaaccct ttcaattgga agcggtccac taccgctgcg attttaagac cacatataaa 780 gctggatgcg gttcctgctg ggcgttctcc ggtgtagctg ccactgaatc ggcttacttg 840 gctggatgcg gttcctgctg ggcgttctcc ggtgtagctg ccactgaatc ggcttacttg 840 gcgcacgaag tcgatcatcg tatcgagatt ttatcgcacg acaaggatta caacaaagtc 900 gcgcacgaag tcgatcatcg tatcgagatt ttatcgcacg acaaggatta caacaaagtc 900 cgcttatatg agcacgcaga agcacgctac tctggtggca gtggtgactt acgtcagatg 960 cgcttatatg agcacgcaga agcacgctac tctggtggca gtggtgactt acgtcagatg 960 cgtacagtga cacctattcg tatgcaagga ggatgtggta gctgtggagg atcaggcaac 1020 cgtacagtga cacctattcg tatgcaagga ggatgtggta gctgtggagg atcaggcaac 1020 accaagacgg cgaaaattga aattaaagca tccattgatg gacttgattc gacgcaccac 1080 accaagacgg cgaaaattga aattaaagca tccattgatg gacttgattc gacgcaccac 1080 caccatcatc acaaaaagta g 1101 caccatcatc acaaaaagta g 1101 <210> 90 <211> 333 <212> PRT <213> Artificial Sequence <400> 90 Met Gly Ala His Ala Ser Val Ile Lys Phe Ala Glu Leu Leu Glu Gln Met Gly Ala His Ala Ser Val Ile Lys Phe Ala Glu Leu Leu Glu Gln 1 5 10 15 1 5 10 15 Gln Lys Asn Ala Gln Phe Pro Gly Lys Pro Glu Met Lys Ile Lys Leu Gln Lys Asn Ala Gln Phe Pro Gly Lys Pro Glu Met Lys Ile Lys Leu 20 25 30 20 25 30 Arg Met Glu Gly Ala Val Asn Gly His Lys Phe Val Ile Glu Gly Glu Arg Met Glu Gly Ala Val Asn Gly His Lys Phe Val Ile Glu Gly Glu 35 40 45 Gly Ile Gly Lys Pro Tyr Glu Gly Thr Gln Thr Leu Asp Leu Thr Val 50 55 60 Glu Glu Asp Ser Ser Ala His Ser Thr Pro Ser Thr Pro Ala Tyr Asp 65 70 75 80 Ile Leu Thr Pro Ala Phe Gln Tyr Gly Asn Arg Ala Phe Thr Lys Tyr 85 90 95 Pro Glu Asp Ile Pro Asp Tyr Phe Lys Gln Ala Phe Pro Glu Gly Tyr 100 105 110 Ser Trp Glu Arg Ser Met Thr Tyr Glu Asp Gln Gly Ile Cys Ile Ala 115 120 125 Thr Ser Asp Ile Thr Met Glu Gly Asp Lys Pro Ser Pro Phe Gly Gln 130 135 140 Ala Ala Ala Ala Asp Lys Cys Phe Phe Tyr Glu Ile Arg Phe Asp Gly 145 150 155 160 Thr Phe Gly Gln Ala Ala Ala Gly Asp Lys Pro Ser Thr Leu Lys Trp 165 170 175 Glu Pro Ser Thr Glu Lys Met Tyr Val Glu Ala Glu Pro Lys Pro Ala 180 185 190 Glu Pro Lys Ser Val Leu Lys Gly Asp Val Glu Met Ala Leu Leu Leu 195 200 205 Thr Ser Ser Thr Pro Pro Ser Gly Thr Glu Asn Lys Pro Ala Thr Gly 210 215 220 His Tyr Arg Cys Asp Phe Lys Thr Thr Tyr Lys Ala Gly Thr Ser Glu 225 230 235 240 Glu Gly Ser Arg Gly Gly Ser Ser Met Pro Ser His Glu Val Asp His 245 250 255 Arg Ile Glu Ile Leu Ser His Ser Pro Phe Gly Gln Ala Ala Ala Gly 260 265 270 Asp Lys Lys Val Arg Leu Tyr Glu His Ala Glu Ala Arg Tyr Ser Gly 275 280 285 Gly Gly Ser Gly Lys Ala Ala Ile Ala Pro Ala Gly Gly Ala Ser Gly 290 295 300 Lys Gln Arg Ala Ala Glu Ala Thr Lys Pro Ile Pro Asn Pro Leu Leu 305 310 315 320 Gly Leu Asp Ser Thr His His His His His His Lys Lys 325 330 <210> 91 <211> 1002 <212> PRT <213> Artificial Sequence <400> 91 Ala Thr Gly Gly Gly Thr Gly Cys Gly Cys Ala Cys Gly Cys Ala Ala 1 5 10 15 Gly Cys Gly Thr Gly Ala Thr Thr Ala Ala Gly Thr Thr Thr Gly Cys 20 25 30 Gly Gly Ala Ala Cys Thr Thr Thr Thr Ala Gly Ala Ala Cys Ala Gly 35 40 45 Cys Ala Gly Ala Ala Ala Ala Ala Cys Gly Cys Ala Cys Ala Gly Thr 50 55 60 Thr Cys Cys Cys Gly Gly Gly Gly Ala Ala Ala Cys Cys Cys Gly Ala 65 70 75 80 Ala Ala Thr Gly Ala Ala Ala Ala Thr Cys Ala Ala Ala Cys Thr Thr 85 90 95 Cys Gly Cys Ala Thr Gly Gly Ala Gly Gly Gly Gly Gly Cys Gly Gly 100 105 110 Thr Ala Ala Ala Cys Gly Gly Cys Cys Ala Cys Ala Ala Gly Thr Thr 115 120 125 Thr Gly Thr Thr Ala Thr Cys Gly Ala Gly Gly Gly Gly Gly Ala Ala 130 135 140 Gly Gly Thr Ala Thr Cys Gly Gly Ala Ala Ala Gly Cys Cys Thr Thr 145 150 155 160 Ala Cys Gly Ala Gly Gly Gly Ala Ala Cys Gly Cys Ala Gly Ala Cys 165 170 175 Thr Thr Thr Gly Gly Ala Cys Thr Thr Ala Ala Cys Ala Gly Thr Ala 180 185 190 Gly Ala Gly Gly Ala Ala Gly Ala Cys Thr Cys Ala Thr Cys Thr Gly 195 200 205 Cys Ala Cys Ala Thr Thr Cys Ala Ala Cys Cys Cys Cys Gly Thr Cys 210 215 220 Thr Ala Cys Thr Cys Cys Gly Gly Cys Ala Thr Ala Thr Gly Ala Thr 225 230 235 240 Ala Thr Thr Thr Thr Ala Ala Cys Thr Cys Cys Thr Gly Cys Gly Thr 245 250 255 Thr Thr Cys Ala Ala Thr Ala Thr Gly Gly Gly Ala Ala Cys Cys Gly 260 265 270 Thr Gly Cys Ala Thr Thr Thr Ala Cys Thr Ala Ala Ala Thr Ala Cys 275 280 285 Cys Cys Ala Gly Ala Gly Gly Ala Cys Ala Thr Cys Cys Cys Thr Gly 290 295 300 Ala Thr Thr Ala Thr Thr Thr Cys Ala Ala Ala Cys Ala Gly Gly Cys 305 310 315 320 Thr Thr Thr Thr Cys Cys Gly Gly Ala Ala Gly Gly Thr Thr Ala Cys 325 330 335 Thr Cys Ala Thr Gly Gly Gly Ala Gly Cys Gly Cys Thr Cys Thr Ala 340 345 350 Thr Gly Ala Cys Ala Thr Ala Thr Gly Ala Ala Gly Ala Thr Cys Ala 355 360 365 Ala Gly Gly Ala Ala Thr Cys Thr Gly Thr Ala Thr Thr Gly Cys Cys 370 375 380 Ala Cys Thr Thr Cys Gly Gly Ala Cys Ala Thr Cys Ala Cys Cys Ala 385 390 395 400 Thr Gly Gly Ala Gly Gly Gly Ala Gly Ala Cys Ala Ala Gly Cys Cys 405 410 415 Gly Ala Gly Cys Cys Cys Ala Thr Thr Thr Gly Gly Thr Cys Ala Ala 420 425 430 Gly Cys Ala Gly Cys Thr Gly Cys Cys Gly Cys Ala Gly Ala Thr Ala 435 440 445 Ala Ala Thr Gly Thr Thr Thr Thr Thr Thr Cys Thr Ala Thr Gly Ala 450 455 460 Ala Ala Thr Thr Cys Gly Thr Thr Thr Cys Gly Ala Cys Gly Gly Gly 465 470 475 480 Ala Cys Ala Thr Thr Thr Gly Gly Ala Cys Ala Ala Gly Cys Gly Gly 485 490 495 Cys Thr Gly Cys Ala Gly Gly Cys Gly Ala Cys Ala Ala Gly Cys Cys 500 505 510 Thr Ala Gly Cys Ala Cys Gly Thr Thr Gly Ala Ala Gly Thr Gly Gly 515 520 525 Gly Ala Gly Cys Cys Ala Ala Gly Cys Ala Cys Cys Gly Ala Ala Ala 530 535 540 Ala Gly Ala Thr Gly Thr Ala Cys Gly Thr Thr Gly Ala Gly Gly Cys 545 550 555 560 Thr Gly Ala Ala Cys Cys Gly Ala Ala Gly Cys Cys Ala Gly Cys Gly 565 570 575 Gly Ala Gly Cys Cys Gly Ala Ala Ala Thr Cys Ala Gly Thr Cys Thr 580 585 590 Thr Ala Ala Ala Gly Gly Gly Gly Gly Ala Thr Gly Thr Thr Gly Ala 595 600 605 Ala Ala Thr Gly Gly Cys Thr Thr Thr Gly Cys Thr Thr Cys Thr Gly 610 615 620 Ala Cys Gly Ala Gly Cys Ala Gly Cys Ala Cys Gly Cys Cys Ala Cys 625 630 635 640 Cys Ala Ala Gly Thr Gly Gly Cys Ala Cys Ala Gly Ala Ala Ala Ala 645 650 655 Cys Ala Ala Ala Cys Cys Cys Gly Cys Cys Ala Cys Ala Gly Gly Ala 660 665 670 Cys Ala Thr Thr Ala Thr Cys Gly Cys Thr Gly Cys Gly Ala Thr Thr 675 680 685 Thr Thr Ala Ala Gly Ala Cys Thr Ala Cys Ala Thr Ala Thr Ala Ala 690 695 700 Gly Gly Cys Thr Gly Gly Thr Ala Cys Cys Thr Cys Thr Gly Ala Gly 705 710 715 720 Gly Ala Gly Gly Gly Gly Thr Cys Thr Cys Gly Cys Gly Gly Ala Gly 725 730 735 Gly Thr Ala Gly Thr Ala Gly Cys Ala Thr Gly Cys Cys Gly Thr Cys 740 745 750 Ala Cys Ala Cys Gly Ala Gly Gly Thr Thr Gly Ala Cys Cys Ala Cys 755 760 765 Cys Gly Cys Ala Thr Thr Gly Ala Gly Ala Thr Cys Thr Thr Ala Thr 770 775 780 Cys Thr Cys Ala Cys Thr Cys Cys Cys Cys Thr Thr Thr Thr Gly Gly 785 790 795 800 Thr Cys Ala Gly Gly Cys Thr Gly Cys Ala Gly Cys Thr Gly Gly Gly 805 810 815 Gly Ala Thr Ala Ala Gly Ala Ala Gly Gly Thr Gly Cys Gly Thr Cys 820 825 830 Thr Thr Thr Ala Thr Gly Ala Gly Cys Ala Cys Gly Cys Gly Gly Ala 835 840 845 Gly Gly Cys Cys Cys Gly Thr Thr Ala Cys Thr Cys Thr Gly Gly Thr 850 855 860 Gly Gly Ala Gly Gly Cys Ala Gly Thr Gly Gly Gly Ala Ala Ala Gly 865 870 875 880 Cys Gly Gly Cys Ala Ala Thr Thr Gly Cys Cys Cys Cys Cys Gly Cys 885 890 895 Ala Gly Gly Cys Gly Gly Cys Gly Cys Gly Thr Cys Ala Gly Gly Gly 900 905 910 Ala Ala Ala Cys Ala Ala Cys Gly Cys Gly Cys Cys Gly Cys Thr Gly 915 920 925 Ala Gly Gly Cys Gly Ala Cys Gly Ala Ala Ala Cys Cys Gly Ala Thr 930 935 940 Cys Cys Cys Gly Ala Ala Cys Cys Cys Gly Cys Thr Gly Thr Thr Gly 945 950 955 960 Gly Gly Ala Cys Thr Thr Gly Ala Cys Ala Gly Thr Ala Cys Cys Cys 965 970 975 Ala Cys Cys Ala Thr Cys Ala Thr Cys Ala Cys Cys Ala Cys Cys Ala 980 985 990 Cys Ala Ala Gly Ala Ala Ala Thr Ala Gly 995 1000 <210> 92 <211> 12 <212> PRT <213> Artificial sequence <400> 92 Asp Ser Ser Ala His Ser Thr Pro Ser Thr Pro Ala 1 5 10 <210> 93 <211> 11 <212> PRT <213> Artificial sequence <400> 93 Phe Gly Gln Ala Ala Ala Gly Asp Lys Pro Ser 1 5 10 <210> 94 <211> 16 <212> PRT <213> Artificial sequence <400> 94 Thr Ser Ser Thr Pro Pro Ser Gly Thr Glu Asn Lys Pro Ala Thr Gly 1 5 10 15 <210> 95 <211> 6 <212> PRT <213> Artificial sequence <400> 95 Ser Tyr Trp Lys Gly Ser 1 5 <210> 96 <211> 8 <212> PRT <213> Artificial sequence <400> 96 Glu Ala Ala Lys Glu Ala Ala Lys 1 5 <210> 97 <211> 7 <212> PRT <213> Artificial sequence <400> 97 Thr Ser Tyr Trp Lys Gly Ser 1 5 <210> 98 <211> 4 <212> PRT <213> Artificial sequence <400> 98 Gly Gly Ser Gly 1 <210> 99 <211> 5 <212> PRT <213> Artificial sequence <400> 99 Gly Gly Ala Ser Gly 1 5 <210> 100 <211> 15 <212> PRT <213> Artificial sequence <400> 100 Tyr Phe Pro Leu Gln Ser Tyr Gly Phe Gln Pro Thr Asn Gly Val 1 5 10 15 <210> 101 <211> 15 <212> PRT <213> Artificial sequence <400> 101 Arg Ser Tyr Thr Pro Gly Asp Ser Ser Ser Gly Trp Thr Ala Gly 1 5 10 15 <210> 102 <211> 15 <212> PRT <213> Artificial sequence <400> 102 Gly Lys Thr Phe Pro Pro Thr Glu Pro Lys Lys Asp Lys Lys Gly 1 5 10 15 <210> 103 <211> 15 <212> PRT <213> Artificial sequence <400> 103 Met Tyr Ser Phe Val Ser Glu Glu Thr Gly Thr Leu Ile Val Asn 1 5 10 15 <210> 104 <211> 15 <212> PRT <213> Artificial sequence <400> 104 Pro Leu Gln Ser Tyr Gly Phe Gln Pro Thr Asn Gly Val Gly Tyr 1 5 10 15 <210> 105 <211> 15 <212> PRT <213> Artificial sequence <400> 105 Gly Gly Met Lys Asp Leu Ser Pro Arg Trp Tyr Phe Gly Gly Gly 1 5 10 15 <210> 106 <211> 15 <212> PRT <213> Artificial sequence <400> 106 Gly Ser Lys Ser Pro Ile Gln Tyr Ile Asp Gly Gly Gly Gly Gly 1 5 10 15 <210> 107 <211> 15 <212> PRT <213> Artificial sequence <400> 107 Tyr Ile Arg Gly Ala Arg Lys Ser Ala Pro Leu Ile Glu Leu Gly 1 5 10 15 <210> 108 <211> 1041 <212> DNA <213> Artificial sequence <400> 108 atggttaatt atctcggagc acacgctttc gagcgcgaca tctcaacgga aatttatcaa 60 gccggttcaa ccatgaagat caagctccgt atggaagggg cggtcaacgg tcacaaattc 120 gtcatcgaag gggaaggcat tggcaagcca tacgaaggga cacagacttt ggacctcact 180 gtggaagaag gttcatctgg tgaagcggcg aaggaagcgg ccaagggaag cacaccatgc 240 aacggcgtgg aaggtttcaa ttgctacttt tacgacatcc ttaccccggc gttccaatat 300 ggtaaccgtg ccttcacgaa gtacccagaa gatatccctg attactttaa gcaggcattt 360 cctgaaggtt attcgtggga gcgcagtatg acctatgagg accaaggtat ttgtatcgcg 420 acgagcgaca ttaccatgga aggggactgt ttcttctacg agattcgctt cgatggaact 480 aattcgaata acctggacag taaagttggt ggcaactaca actatacctt aaaatgggag 540 ccaagtacag aaaagatgta cgttgaagac ggggtgttga agggtgacgt ggaaatggca 600 ttactcttgt ttgagcgcga catttcaacc gagatttacc aagccgggtc gacaggcggg 660 tccggaacaa gctactggaa ggggagtcac tatcgctgtg attttaagac gacctacaaa 720 gctggtagca ctccatgtaa cggagtagag gggttcaact gctactttca tgaagtagat 780 caccgtattg agattttatc acactacttt cctctgcaat cgtatggctt ccaacccacg 840 aatggcgttg gtagtagcgg cgaagccgcc aaggaagcag ccaagaaggt gcgcctgtac 900 gagcacgccg aggcgtactt tcctctccag tcctatggat tccaacctac caatggagtt 960 ggcggctcgg gcggcggtgc ttccgggaat tcgaacaact tagatagcaa ggtgggtggg 1020 aactataatt atctcgagta g 1041 <210> 109 <211> 1059 <212> DNA <213> Artificial Sequence<00…atggttaatt atctcggagc acacgctttc gagcgcgaca tctcaacgga aatttatcaa 60 gccggttcaa ccatgaagat caagctccgt atggaagggg cggtcaacgg tcacaaattc 120 gtcatcgaag gggaaggcat tggcaagcca tacgaagga cacagacttt ggacctcact 180 gtggagaag gttcatctgg tgaagcggcg aaagcgg ccaagggaag cacaccatgc 240 aacggcgtgg aaggtttcaa ttgctacttt tacgacatcc ttaccccggc gttccaatat 300 ggtaaccgtg ccttcacagaa gtacccagaa gatatccctg attactttaa gcaggcattt 360 cctgaaggtt attcgtggga gcgcagtatg acctatgagg accaaggtat ttgtatcgcg 420 acgagcgaca ttaccatgga aggggactgt ttcttctacg agattcgctt cgatggaact 480 aattcgaata acctggagag taaagttggt ggcaactaca actatacctt aaaatgggag 540 600 ttactcttgt ttgagcgcga catttcaacc gagatttacc aagccgggtc gacaggcggg 660 720 gctggtagca ctccatgtaa cggagtagag gggttcaact gctactttca tgaagtagat 780 caccgtattg agattttatc acactacttt cctctgcaat cgtatggctt ccaacccacg 840 aatggcgttg gtagtagcgg cgaagccgcc aaggaagcag ccaagaaggt gcgcctgtac 900 gagcacgccg aggcgtactt tcctctccag tcctatggat tccaacctac caatggagtt 960 ggcggctcgg gcggcggtgc ttccgggaat tcgaacaact tagatagcaa ggtgggtggg 1020 aactataatt atctcgagca ccaccaccac caccactga 1059 <210> 110 <211> 1128 <212> DNA <213> Artificial sequence <400> 110 atggtaaatt atatagggtc aagtggagtt gtcaacccgg tgatgggtgg ctcgggtggt 60 ccagcgccag cgccgatgaa aattaagttg cgtatggaag gggcggttaa tgctccgcgt 120 atcaccttcg gcggtcctag cgatagcacc ggcggtggtt ctggcgaagc cgcgaagacg 180 agctactgga agggttcgca taaattcgtg attgagggtg agggcattgg taaaccgtat 240 gagggtacgc agaccttaga tctgaccgtt gaggaaggtt ctagcggcgt cgttaacccg 300 gttatgtatg atattttgac cccggcgttc caatatggaa accgcgcatt caccaaatac 360 ccggaatata ttcgtgttgg tgcccgtaaa agcgctccgc tgatcgagct gggttattcc 420 tgggagcgct ccatgaccta cggctccctg atcgacctgc aagagttggg taagtacgag 480 caatacatcg aggcagccgc taaagaagcc gcagcgaagg gcggcgcttc tggcatctgc 540 attgcgacca gcgacatcac catggagggc gactgttttt tctatgaaat ccgcttcgat 600 ggcaccggtc cgtttcagca gtttggacgt gatatcgcgg ataccacgga tgcgaccctg 660 aagtgggagc cgtccactga aaaaatgtac gtcgaggacg gtgtgctgaa gggtgacgtt 720 gaaatggcgc tgctccttgg tggctccggc acgtcatact ggaaaggttc tatgtggctg 780 tcttacttca tcgccagctt tcgtctgcat tatagatgcg attttaaaac tacgtacaaa 840 gcacgcagct acaccccggg cgactccagc agcggctgga ccgcacatga agtggatcat 900 cgtattgaga tcttaagtca cattgtggac gaaccgggcg gtgcgagcgg caaggtgcgt 960 ctgtatgagc acgcggaagc gccggcgccg gctccgggtt ttagcgcatt ggaaccgctg 1020 gtagacctgc cgggcggcag cggtggtggt gcgagcggga agaccttccc gcctacagaa 1080 ccgaaaaagg acaaaaaggg tcatcaccac caccaccaca agaagtaa 1128 <210> 111 <211> 1074 <212> DNA <213> Artificial sequence <400> 111 atggtaaatt acatactagg agtttatcac aagaacaaca aaagctggat ggagagcgaa 60 ttccgcgttt atccggcacc ggcgcctatg aaaatcaagc tgcgtatgga aggtgcagtt 120 aatggtcata aattcgtgat tgagggtgaa ggcattggca aaccaggtgg ttcaggtgag 180 gctgcgaagt tcaactgcta ttttccgctg cagagctacg gctttcagcc gacgggtacg 240 cagaccctgg atctgacggt cgaggaaccg ttgcaatcgt acggtttcca accgacctat 300 gacatcctga ccccggcgtt ccagtatggc aaccgcgcgt tcaccaaata cccggaagcc 360 ggtaacggtg gcgatgccgc cctcgctctg ttactgttag atggctacag ctgggaacgt 420 agcatgacgt acggccgttc ctacctgacc ccgggtgaca gcagctccgg tggcgcgtcg 480 ggtggtatct gcattgcgac cagtgatatt acgatggagg gtgactgttt tttttatgag 540 atccgctttg atggtaccgc ggaccagttg accccaactt ggcgtgtgac ccttaagtgg 600 gagccgagca ccgaaaaaat gtatgtagag gatggggtgc tgaagggcga cgttgaaatg 660 gcattgttgt tgtttatcta taataaaatc gtggacgaac cgggtggcag cggtacttct 720 tactggaaag gttcccatta tcgttgcgac ttcaaaacca cctacaaagc aaagaacccg 780 ttgctatacg atgcgaatta ccatgaagtt gaccatcgta ttgaaattct gagccacggc 840 ggctctggag gcgaggctgc taagggtaga cctcaaggtc tgccgaataa caccgcatcc 900 aaagttcgtc tgtacgagca cgcggaggcg ctggcggaga tcctgcaaaa aaacctgatc 960 cgccagggta cagattataa gcattggccg caaattgcgg gcggctccgg tggcggcaag 1020 atcgccgact acaattacaa gctgggccat caccaccacc accacaagaa gtaa 1074 <210> 112 <211> 1119 <212> DNA <213> Synthetic Sequence <400> 112 atggtaaatt atatatttaa ctgttacttc ccgctgcaga gctacggctt ccagccgacc 60 aatggtgtga tgaaaatcaa actgcgtatg gaaggcgcag ttcgtccgca gggtctgccg 120 aataacaccg caagcggcgg ttccggtggc gaggcggcga agcataagtt cgtgattgaa 180 ggtgagggta ttggtaagcc ggaagctgcg aaagaagctg ctaagggctc tggcttcatc 240 tataacaaaa tcgttgatga gccgggtgcg ggtacgcaga ccctggatct gactgtagag 300 gaagcggacc aactgacccc gacctggcgt gtgggctatg acatcttgac tccggcgttt 360 caatatggta atagagcatt caccaaatac ccggagggca agatcgccga ctacaactac 420 aagttgggtg gctacagctg ggaacgctct atgacctacg aagatcaagg tatttgtatt 480 gcgaccagcg atatcaccat ggagggtgac tgcttttttt atgaaatccg ctttgatggt 540 acatacattc gcgtgggtgc gcgtaaaagc gctccgctga ttgagctgac gctgaaatgg 600 gagccgagca ccgaaaaaat gtatgtcgag atcgtggacg aaccgggagg gtccggtggt 660 gttctgaagg gcgacgttga gatggcactg ttgttgaaga acccgttatt gtacgatgca 720 aattacgggg gttcgggcac ctcctattgg aaaggtcatt atcgctgcga tttcaaaacc 780 acctataaag cgaagacctt tccgccgacg gagcctaaaa aagataaaaa gcacgaagtt 840 gaccatcgta ttgagatcct gagccatgaa gcggctaagg gcggtgccag cggccgtagc 900 tacctcactc cgggtgacag ctctagcaaa gttcgtctgt atgagcacgc agaagccccg 960 gcgccagcgc caggctcgtc cggtgtcgtg aacccggtca tggagccgat ttacgacggc 1020 ggctccggcg gtgcgccagg ccagacggga aagatcgccg actataacta caagcttggt 1080 gcgtcaggca aacaccatca ccaccaccac aagaagtaa 1119 <210> 113 <211> 975 <212> DNA <213> Artificial sequence <400> 113 atggtcaatt atatcgcgct cccacaacgc caaaagaagc agcagaccgt gacgctgttg 60 ccggcaccgg cacccatgaa gattaagtta cgtatggaag gagcggtcaa tgggcacaag 120 tttgtcatcg agggcgaagg aatcgggaaa ccttacgaag ggacccaaac attagatctg 180 accgttgaag agggaagcaa gagcccaatc caatacattg attacgatat tcttacgcct 240 gcatttcagt acggcaatcg tgctttcaca aagtaccctg aagacttctt ggagtatcac 300 gacgtgcgcg tagttttaga cttcgggtat tcttgggaac gtagtatgac gtatggtgga 360 gctagcggtg gtatcaatgc aagtgttgtc aacatccagg gaatctgcat cgccacctcg 420 gatatcacta tggaaggcgg gtccggggaa gctgcgaagc aattcgcccc ctcggctagt 480 gccttcttct gcttcttcta cgaaatccgt tttgacggta ccatgtatag cttcgtaagc 540 gaagaaacgg gcacgctgat tgtcaatact ctgaagtggg agccgtctac ggaaaagatg 600 tatgtagagg acggcgtact gaagggcgac gtcgagatgg cgcttctgtt agaaggcggc 660 ggtcattacc gttgcgattt taagacaacg tataaggcac catccggtac ttggttaact 720 tacacaggcc acgaggtcga ccatcgtatc gagatcttat cacatggcgg ttccggcggc 780 gaagcggcca aggggcagtt cgcgcctagt gcatcggctt tcttcaaggt tcgcctctat 840 gaacacgccg aggcaggcgg tgcatctgga gagctggaca aatatggcgg ttcgggcggc 900 ggtatgtaca gtttcgtatc ggaagagaca ggaactttaa ttgtgaatca ccatcaccac 960 caccataaga agtaa 975 <210> 114 <211> 996 <212> DNA <213> Artificial sequence <400> 114 atggtgaatt acatcccgct ccaatcatac ggatttcagc caacgaatgg ggtcgggtac 60 cctgcgccag cacccgcaat gaaaatcaag cttcgtatgg aaggtgcagt taatggacac 120 aagttcgtga tcgagggcga gggcattggc aagccctatg aagggaccca aactttggac 180 ttgacggtag aagagggtat ctatcagacc agtaatttc gtgtctacga catcttgact 240 300 acccaagcat tcgggcgtcg tggccctgaa ggctatagct gggaacgttc tatgacgtac 360 420 aatcaagtgg cggtgggtgg aagcggtgag gcagcaaagt gcttctttta tgagatccgc 480 ttcgatggta ctaacccggt cctgcccttt aacgatggag tttatttcgc ctcaaccaca 540 ttgaaatggg agccctcaac tgaaaagatg tatgttgaag acggcgtttt gaagggtgac 600 gtagagatgg cacttctgct gtacaactac aagcttcccg atgacttcac tggcggcagc 660 ggtacaagtt attggaaggg ttcacattat cgctgtgatt tcaagacaac ctataaagca 720 atgttccatc tggtagactt tcaagttacg attgctgaga tccttcatga ggttgatcac 780 cgtatcgaaa tcctttctca cggcggaagc ggcggagagg cggccaaggg tatgaaggac 840 ttgagccctc gctggtattt caaggtccgc ttatacgagc acgccgaggc cgacgctgcc 900 cttgcgctcc tgttattaga cggcggaagc ggtggtggca tgaaggattt atcccctcgc 960 tggtacttcc atcaccatca ccaccacaag aagtga 996 <210> 115 <211> 1059 <212> DNA <213> Artificial sequence <400> 115 atggtgaatt atatcttagg ggtgtaccac aagaataaca agtcatggat ggagtcggaa 60 tttcgtgtgt acccggcccc ggctcccatg aagatcaagt tacgcatgga aggcgccgtt 120 aacggccaca aattcgtgat cgagggcgaa ggaatcggta agccaggtgg ctcgggtgaa 180 gccgccaagt tcatttacaa taagatcgtt gatgaaccag gaacacaaac tttagatctt 240 accgtagagg agaagaaccc gcttctttac gacgccaatt actatgacat tttaacgccg 300 gccttccaat acggcaatcg cgcttttacg aagtatccgg aagcgggtaa cggcggcgat 360 gcagccttag cattattatt actggacggt tattcgtggg agcgttcaat gacgtatggc 420 cgtagttact tgacccccgg agacagtagc tctggtggtg ccagcggcgg gatctgtatt 480 gcaactagtg atatcaccat ggaaggcgac tgcttcttct atgagattcg tttcgatggg 540 accgccgatc aactgacccc tacctggcgt gtaacattga agtgggagcc aagcactgag 600 aagatgtacg tggaagacgg cgttctgaag ggcgacgtgg agatggcctt gttattattc 660 atctacaata agatcgtgga cgagccgggt ggtagcggta catcctactg gaagggttcg 720 cattaccgtt gtgatttcaa gacaacctat aaggctaaga atccgttatt gtatgacgca 780 aattaccatg aggtggacca ccgtatcgaa atcctttccc atggcggaag cggtggtgag 840 gcggcgaagg gtcgcccgca gggactgccg aataatacgg cgtctaaggt tcgcttatac 900 gaacacgctg aggccctcgc cgagatcctg caaaagaact taatccgcca aggaaccgac 960 tataagcact ggccgcaaat cgccggtggc tccggtggtg gtaagatcgc agactacaac 1020 tataagctcg gtcatcatca tcatcaccat aagaaataa 1059 <210> 116 <211> 346 <212> PRT <213> Artificial sequence <400> 116 Met Val Asn Tyr Leu Gly Ala His Ala Phe Glu Arg Asp Ile Ser Thr 1 5 10 15 Glu Ile Tyr Gln Ala Gly Ser Thr Met Lys Ile Lys Leu Arg Met Glu 20 25 30 Gly Ala Val Asn Gly His Lys Phe Val Ile Glu Gly Glu Gly Ile Gly 35 40 45 Lys Pro Tyr Glu Gly Thr Gln Thr Leu Asp Leu Thr Val Glu Glu Gly 50 55 60 Ser Ser Gly Glu Ala Ala Lys Glu Ala Ala Lys Gly Ser Thr Pro Cys 65 70 75 80 Asn Gly Val Glu Gly Phe Asn Cys Tyr Phe Tyr Asp Ile Leu Thr Pro 85 90 95 Ala Phe Gln Tyr Gly Asn Arg Ala Phe Thr Lys Tyr Pro Glu Asp Ile 100 105 110 Pro Asp Tyr Phe Lys Gln Ala Phe Pro Glu Gly Tyr Ser Trp Glu Arg 115 120 125 Ser Met Thr Tyr Glu Asp Gln Gly Ile Cys Ile Ala Thr Ser Asp Ile 130 135 140 Thr Met Glu Gly Asp Cys Phe Phe Tyr Glu Ile Arg Phe Asp Gly Thr 145 150 155 160 Asn Ser Asn Asn Leu Asp Ser Lys Val Gly Gly Asn Tyr Asn Tyr Thr 165 170 175 Leu Lys Trp Glu Pro Ser Thr Glu Lys Met Tyr Val Glu Asp Gly Val 180 185 190 Leu Lys Gly Asp Val Glu Met Ala Leu Leu Leu Phe Glu Arg Asp Ile 195 200 205 Ser Thr Glu Ile Tyr Gln Ala Gly Ser Thr Gly Gly Ser Gly Thr Ser 210 215 220 Tyr Trp Lys Gly Ser His Tyr Arg Cys Asp Phe Lys Thr Thr Tyr Lys 225 230 235 240 Ala Gly Ser Thr Pro Cys Asn Gly Val Glu Gly Phe Asn Cys Tyr Phe 245 250 255 His Glu Val Asp His Arg Ile Glu Ile Leu Ser His Tyr Phe Pro Leu 260 265 270 Gln Ser Tyr Gly Phe Gln Pro Thr Asn Gly Val Gly Ser Ser Gly Glu 275 280 285 Ala Ala Lys Glu Ala Ala Lys Lys Val Arg Leu Tyr Glu His Ala Glu 290 295 300 Ala Tyr Phe Pro Leu Gln Ser Tyr Gly Phe Gln Pro Thr Asn Gly Val 305 310 315 320 Gly Gly Ser Gly Gly Gly Ala Ser Gly Asn Ser Asn Asn Leu Asp Ser 325 330 335 Lys Val Gly Gly Asn Tyr Asn Tyr Leu Glu 340 345 <210> 117 <211> 352 <212> PRT <213> Artificial sequence <400> 117 Met Val Asn Tyr Leu Gly Ala His Ala Phe Glu Arg Asp Ile Ser Thr 1 5 10 15 Glu Ile Tyr Gln Ala Gly Ser Thr Met Lys Ile Lys Leu Arg Met Glu 20 25 30 Gly Ala Val Asn Gly His Lys Phe Val Ile Glu Gly Glu Gly Ile Gly 35 40 45 Lys Pro Tyr Glu Gly Thr Gln Thr Leu Asp Leu Thr Val Glu Glu Gly 50 55 60 Ser Ser Gly Glu Ala Ala Lys Glu Ala Ala Lys Gly Ser Thr Pro Cys 65 70 75 80 Asn Gly Val Glu Gly Phe Asn Cys Tyr Phe Tyr Asp Ile Leu Thr Pro 85 90 95 Ala Phe Gln Tyr Gly Asn Arg Ala Phe Thr Lys Tyr Pro Glu Asp Ile 100 105 110 Pro Asp Tyr Phe Lys Gln Ala Phe Pro Glu Gly Tyr Ser Trp Glu Arg 115 120 125 Ser Met Thr Tyr Glu Asp Gln Gly Ile Cys Ile Ala Thr Ser Asp Ile 130 135 140 Thr Met Glu Gly Asp Cys Phe Phe Tyr Glu Ile Arg Phe Asp Gly Thr 145 150 155 160 Asn Ser Asn Asn Leu Asp Ser Lys Val Gly Gly Asn Tyr Asn Tyr Thr 165 170 175 Leu Lys Trp Glu Pro Ser Thr Glu Lys Met Tyr Val Glu Asp Gly Val 180 185 190 Leu Lys Gly Asp Val Glu Met Ala Leu Leu Leu Phe Glu Arg Asp Ile 195 200 205 Ser Thr Glu Ile Tyr Gln Ala Gly Ser Thr Gly Gly Ser Gly Thr Ser 210 215 220 Tyr Trp Lys Gly Ser His Tyr Arg Cys Asp Phe Lys Thr Thr Tyr Lys 225 230 235 240 Ala Gly Ser Thr Pro Cys Asn Gly Val Glu Gly Phe Asn Cys Tyr Phe 245 250 255 His Glu Val Asp His Arg Ile Glu Ile Leu Ser His Tyr Phe Pro Leu 260 265 270 Gln Ser Tyr Gly Phe Gln Pro Thr Asn Gly Val Gly Ser Ser Gly Glu 275 280 285 Ala Ala Lys Glu Ala Ala Lys Lys Val Arg Leu Tyr Glu His Ala Glu 290 295 300 Ala Tyr Phe Pro Leu Gln Ser Tyr Gly Phe Gln Pro Thr Asn Gly Val 305 310 315 320 Gly Gly Ser Gly Gly Gly Ala Ser Gly Asn Ser Asn Asn Leu Asp Ser 325 330 335 Lys Val Gly Gly Asn Tyr Asn Tyr Leu Glu His His His His His His 340 345 350 <210> 118 <211> 375 <212> PRT <213> Synthetic sequence <400> 118 Met Val Asn Tyr Ile Gly Ser Ser Gly Val Val Asn Pro Val Met Gly 1 5 10 15 Gly Ser Gly Gly Pro Ala Pro Ala Pro Met Lys Ile Lys Leu Arg Met 20 25 30 Glu Gly Ala Val Asn Ala Pro Arg Ile Thr Phe Gly Gly Pro Ser Asp 35 40 45 Ser Thr Gly Gly Gly Ser Gly Glu Ala Ala Lys Thr Ser Tyr Trp Lys 50 55 60 Gly Ser His Lys Phe Val Ile Glu Gly Glu Gly Ile Gly Lys Pro Tyr 65 70 75 80 Glu Gly Thr Gln Thr Leu Asp Leu Thr Val Glu Glu Gly Ser Ser Gly 85 90 95 Val Val Asn Pro Val Met Tyr Asp Ile Leu Thr Pro Ala Phe Gln Tyr 100 105 110 Gly Asn Arg Ala Phe Thr Lys Tyr Pro Glu Tyr Ile Arg Val Gly Ala 115 120 125 Arg Lys Ser Ala Pro Leu Ile Glu Leu Gly Tyr Ser Trp Glu Arg Ser 130 135 140 Met Thr Tyr Gly Ser Leu Ile Asp Leu Gln Glu Leu Gly Lys Tyr Glu 145 150 155 160 Gln Tyr Ile Glu Ala Ala Ala Lys Glu Ala Ala Ala Lys Gly Gly Ala 165 170 175 Ser Gly Ile Cys Ile Ala Thr Ser Asp Ile Thr Met Glu Gly Asp Cys 180 185 190 Phe Phe Tyr Glu Ile Arg Phe Asp Gly Thr Gly Pro Phe Gln Gln Phe 195 200 205 Gly Arg Asp Ile Ala Asp Thr Thr Asp Ala Thr Leu Lys Trp Glu Pro 210 215 220 Ser Thr Glu Lys Met Tyr Val Glu Asp Gly Val Leu Lys Gly Asp Val 225 230 235 240 Glu Met Ala Leu Leu Leu Gly Gly Ser Gly Thr Ser Tyr Trp Lys Gly 245 250 255 Ser Met Trp Leu Ser Tyr Phe Ile Ala Ser Phe Arg Leu His Tyr Arg 260 265 270 Cys Asp Phe Lys Thr Thr Tyr Lys Ala Arg Ser Tyr Thr Pro Gly Asp 275 280 285 Ser Ser Ser Gly Trp Thr Ala His Glu Val Asp His Arg Ile Glu Ile 290 295 300 Leu Ser His Ile Val Asp Glu Pro Gly Gly Ala Ser Gly Lys Val Arg 305 310 315 320 Leu Tyr Glu His Ala Glu Ala Pro Ala Pro Ala Pro Gly Phe Ser Ala 325 330 335 Leu Glu Pro Leu Val Asp Leu Pro Gly Gly Ser Gly Gly Gly Ala Ser 340 345 350 Gly Lys Thr Phe Pro Pro Thr Glu Pro Lys Lys Asp Lys Lys Gly His 355 360 365 His His His His His Lys Lys 370 375 <210> 119 <211> 357 <212> PRT <213> Artificial Sequence <400> 119 Met Val Asn Tyr Ile Leu Gly Val Tyr His Lys Asn Asn Lys Ser Trp 1 5 10 15 Met Glu Ser Glu Phe Arg Val Tyr Pro Ala Pro Ala Pro Met Lys Ile 20 25 30 Lys Leu Arg Met Glu Gly Ala Val Asn Gly His Lys Phe Val Ile Glu 35 40 45 Gly Glu Gly Ile Gly Lys Pro Gly Gly Ser Gly Glu Ala Ala Lys Phe 50 55 60 Asn Cys Tyr Phe Pro Leu Gln Ser Tyr Gly Phe Gln Pro Thr Gly Thr 65 70 75 80 Gln Thr Leu Asp Leu Thr Val Glu Glu Pro Leu Gln Ser Tyr Gly Phe 85 90 95 Gln Pro Thr Tyr Asp Ile Leu Thr Pro Ala Phe Gln Tyr Gly Asn Arg 100 105 110 Ala Phe Thr Lys Tyr Pro Glu Ala Gly Asn Gly Gly Asp Ala Ala Leu 115 120 125 Ala Leu Leu Leu Leu Asp Gly Tyr Ser Trp Glu Arg Ser Met Thr Tyr 130 135 140 Gly Arg Ser Tyr Leu Thr Pro Gly Asp Ser Ser Ser Gly Gly Ala Ser 145 150 155 160 Gly Gly Ile Cys Ile Ala Thr Ser Asp Ile Thr Met Glu Gly Asp Cys 165 170 175 Phe Phe Tyr Glu Ile Arg Phe Asp Gly Thr Ala Asp Gln Leu Thr Pro 180 185 190 Thr Trp Arg Val Thr Leu Lys Trp Glu Pro Ser Thr Glu Lys Met Tyr 195 200 205 Val Glu Asp Gly Val Leu Lys Gly Asp Val Glu Met Ala Leu Leu Leu 210 215 220 Phe Ile Tyr Asn Lys Ile Val Asp Glu Pro Gly Gly Ser Gly Thr Ser 225 230 235 240 Tyr Trp Lys Gly Ser His Tyr Arg Cys Asp Phe Lys Thr Thr Tyr Lys 245 250 255 Ala Lys Asn Pro Leu Leu Tyr Asp Ala Asn Tyr His Glu Val Asp His 260 265 270 Arg Ile Glu Ile Leu Ser His Gly Gly Ser Gly Gly Glu Ala Ala Lys 275 280 285 Gly Arg Pro Gln Gly Leu Pro Asn Asn Thr Ala Ser Lys Val Arg Leu 290 295 300 Tyr Glu His Ala Glu Ala Leu Ala Glu Ile Leu Gln Lys Asn Leu Ile 305 310 315 320 Arg Gln Gly Thr Asp Tyr Lys His Trp Pro Gln Ile Ala Gly Gly Ser 325 330 335 Gly Gly Gly Lys Ile Ala Asp Tyr Asn Tyr Lys Leu Gly His His His 340 345 350 His His His Lys Lys 355 <210> 120 <211> 372 <212> PRT <213> Artificial Sequence <400> 120 Met Val Asn Tyr Ile Phe Asn Cys Tyr Phe Pro Leu Gln Ser Tyr Gly 1 5 10 15 Phe Gln Pro Thr Asn Gly Val Met Lys Ile Lys Leu Arg Met Glu Gly 20 25 30 Ala Val Arg Pro Gln Gly Leu Pro Asn Asn Thr Ala Ser Gly Gly Ser 35 40 45 Gly Gly Glu Ala Ala Lys His Lys Phe Val Ile Glu Gly Glu Gly Ile 50 55 60 Gly Lys Pro Glu Ala Ala Lys Glu Ala Ala Lys Gly Ser Gly Phe Ile 65 70 75 80 Tyr Asn Lys Ile Val Asp Glu Pro Gly Ala Gly Thr Gln Thr Leu Asp 85 90 95 Leu Thr Val Glu Glu Ala Asp Gln Leu Thr Pro Thr Trp Arg Val Gly 100 105 110 Tyr Asp Ile Leu Thr Pro Ala Phe Gln Tyr Gly Asn Arg Ala Phe Thr 115 120 125 Lys Tyr Pro Glu Gly Lys Ile Ala Asp Tyr Asn Tyr Lys Leu Gly Gly 130 135 140 Tyr Ser Trp Glu Arg Ser Met Thr Tyr Glu Asp Gln Gly Ile Cys Ile 145 150 155 160 Ala Thr Ser Asp Ile Thr Met Glu Gly Asp Cys Phe Phe Tyr Glu Ile 165 170 175 Arg Phe Asp Gly Thr Tyr Ile Arg Val Gly Ala Arg Lys Ser Ala Pro 180 185 190 Leu Ile Glu Leu Thr Leu Lys Trp Glu Pro Ser Thr Glu Lys Met Tyr 195 200 205 Val Glu Ile Val Asp Glu Pro Gly Gly Ser Gly Gly Val Leu Lys Gly 210 215 220 Asp Val Glu Met Ala Leu Leu Leu Lys Asn Pro Leu Leu Tyr Asp Ala 225 230 235 240 Asn Tyr Gly Gly Ser Gly Thr Ser Tyr Trp Lys Gly His Tyr Arg Cys 245 250 255 Asp Phe Lys Thr Thr Tyr Lys Ala Lys Thr Phe Pro Pro Thr Glu Pro 260 265 270 Lys Lys Asp Lys Lys His Glu Val Asp His Arg Ile Glu Ile Leu Ser 275 280 285 His Glu Ala Ala Lys Gly Gly Ala Ser Gly Arg Ser Tyr Leu Thr Pro 290 295 300 Gly Asp Ser Ser Ser Lys Val Arg Leu Tyr Glu His Ala Glu Ala Pro 305 310 315 320 Ala Pro Ala Pro Gly Ser Ser Gly Val Val Asn Pro Val Met Glu Pro 325 330 335 Ile Tyr Asp Gly Gly Ser Gly Gly Ala Pro Gly Gln Thr Gly Lys Ile 340 345 350 Ala Asp Tyr Asn Tyr Lys Leu Gly Ala Ser Gly Lys His His His His 355 360 365 His His Lys Lys 370 <210> 121 <211> 324 <212> PRT <213> Artificial Sequence <400> 121 Met Val Asn Tyr Ile Ala Leu Pro Gln Arg Gln Lys Lys Gln Gln Thr 1 5 10 15 Val Thr Leu Leu Pro Ala Pro Ala Pro Met Lys Ile Lys Leu Arg Met 20 25 30 Glu Gly Ala Val Asn Gly His Lys Phe Val Ile Glu Gly Glu Gly Ile 35 40 45 Gly Lys Pro Tyr Glu Gly Thr Gln Thr Leu Asp Leu Thr Val Glu Glu 50 55 60 Gly Ser Lys Ser Pro Ile Gln Tyr Ile Asp Tyr Asp Ile Leu Thr Pro 65 70 75 80 Ala Phe Gln Tyr Gly Asn Arg Ala Phe Thr Lys Tyr Pro Glu Asp Phe 85 90 95 Leu Glu Tyr His Asp Val Arg Val Val Leu Asp Phe Gly Tyr Ser Trp 100 105 110 Glu Arg Ser Met Thr Tyr Gly Gly Ala Ser Gly Gly Ile Asn Ala Ser 115 120 125 Val Val Asn Ile Gln Gly Ile Cys Ile Ala Thr Ser Asp Ile Thr Met 130 135 140 Glu Gly Gly Ser Gly Glu Ala Ala Lys Gln Phe Ala Pro Ser Ala Ser 145 150 155 160 Ala Phe Phe Cys Phe Phe Tyr Glu Ile Arg Phe Asp Gly Thr Met Tyr 165 170 175 Ser Phe Val Ser Glu Glu Thr Gly Thr Leu Ile Val Asn Thr Leu Lys 180 185 190 Trp Glu Pro Ser Thr Glu Lys Met Tyr Val Glu Asp Gly Val Leu Lys 195 200 205 Gly Asp Val Glu Met Ala Leu Leu Leu Glu Gly Gly Gly His Tyr Arg 210 215 220 Cys Asp Phe Lys Thr Thr Tyr Lys Ala Pro Ser Gly Thr Trp Leu Thr 225 230 235 240 Tyr Thr Gly His Glu Val Asp His Arg Ile Glu Ile Leu Ser His Gly 245 250 255 Gly Ser Gly Gly Glu Ala Ala Lys Gly Gln Phe Ala Pro Ser Ala Ser 260 265 270 Ala Phe Phe Lys Val Arg Leu Tyr Glu His Ala Glu Ala Gly Gly Ala 275 280 285 Ser Gly Glu Leu Asp Lys Tyr Gly Gly Ser Gly Gly Gly Met Tyr Ser 290 295 300 Phe Val Ser Glu Glu Thr Gly Thr Leu Ile Val Asn His His His His 305 310 315 320 His His Lys Lys <210> 122 <211> 331 <212> PRT <213> Synthetic sequence <400> 122 Met Val Asn Tyr Ile Pro Leu Gln Ser Tyr Gly Phe Gln Pro Thr Asn 1 5 10 15 Gly Val Gly Tyr Pro Ala Pro Ala Pro Ala Met Lys Ile Lys Leu Arg 20 25 30 Met Glu Gly Ala Val Asn Gly His Lys Phe Val Ile Glu Gly Glu Gly 35 40 45 Ile Gly Lys Pro Tyr Glu Gly Thr Gln Thr Leu Asp Leu Thr Val Glu 50 55 60 Glu Gly Ile Tyr Gln Thr Ser Asn Phe Arg Val Tyr Asp Ile Leu Thr 65 70 75 80 Pro Ala Phe Gln Tyr Gly Asn Arg Ala Phe Thr Lys Tyr Pro Glu Lys 85 90 95 Ala Tyr Asn Val Thr Gln Ala Phe Gly Arg Arg Gly Pro Glu Gly Tyr 100 105 110 Ser Trp Glu Arg Ser Met Thr Tyr Glu Asp Gln Gly Ile Cys Ile Ala 115 120 125 Thr Ser Asp Ile Thr Met Glu Gly Thr Asn Thr Ser Asn Gln Val Ala 130 135 140 Val Gly Gly Ser Gly Glu Ala Ala Lys Cys Phe Phe Tyr Glu Ile Arg 145 150 155 160 Phe Asp Gly Thr Asn Pro Val Leu Pro Phe Asn Asp Gly Val Tyr Phe 165 170 175 Ala Ser Thr Thr Leu Lys Trp Glu Pro Ser Thr Glu Lys Met Tyr Val 180 185 190 Glu Asp Gly Val Leu Lys Gly Asp Val Glu Met Ala Leu Leu Leu Tyr 195 200 205 Asn Tyr Lys Leu Pro Asp Asp Phe Thr Gly Gly Ser Gly Thr Ser Tyr 210 215 220 Trp Lys Gly Ser His Tyr Arg Cys Asp Phe Lys Thr Thr Tyr Lys Ala 225 230 235 240 Met Phe His Leu Val Asp Phe Gln Val Thr Ile Ala Glu Ile Leu His 245 250 255 Glu Val Asp His Arg Ile Glu Ile Leu Ser His Gly Gly Ser Gly Gly 260 265 270 Glu Ala Ala Lys Gly Met Lys Asp Leu Ser Pro Arg Trp Tyr Phe Lys 275 280 285 Val Arg Leu Tyr Glu His Ala Glu Ala Asp Ala Ala Leu Ala Leu Leu 290 295 300 Leu Leu Asp Gly Gly Ser Gly Gly Gly Met Lys Asp Leu Ser Pro Arg 305 310 315 320 Trp Tyr Phe His His His His His His Lys Lys 325 330 <210> 123 <211> 352 <212> PRT <213> Artificial Sequence <400> 123 Met Val Asn Tyr Ile Leu Gly Val Tyr His Lys Asn Asn Lys Ser Trp 1 5 10 15 Met Glu Ser Glu Phe Arg Val Tyr Pro Ala Pro Ala Pro Met Lys Ile 20 25 30 Lys Leu Arg Met Glu Gly Ala Val Asn Gly His Lys Phe Val Ile Glu<00时,35 40 45 Gly Glu Gly Ile Gly Lys Pro Gly Gly Ser Gly Glu Ala Ala Lys Phe 50 55 60 Ile Tyr Asn Lys Ile Val Asp Glu Pro Gly Thr Gln Thr Leu Asp Leu 65 70 75 80 Thr Val Glu Glu Lys Asn Pro Leu Leu Tyr Asp Ala Asn Tyr Tyr Asp 85 90 95 Ile Leu Thr Pro Ala Phe Gln Tyr Gly Asn Arg Ala Phe Thr Lys Tyr 100 105 110 Pro Glu Ala Gly Asn Gly Gly Asp Ala Ala Leu Ala Leu Leu Leu Leu 115 120 125 Asp Gly Tyr Ser Trp Glu Arg Ser Met Thr Tyr Gly Arg Ser Tyr Leu 130 135 140 Thr Pro Gly Asp Ser Ser Ser Gly Gly Ala Ser Gly Gly Ile Cys Ile 145 150 155 160 Ala Thr Ser Asp Ile Thr Met Glu Gly Asp Cys Phe Phe Tyr Glu Ile 165 170 175 Arg Phe Asp Gly Thr Ala Asp Gln Leu Thr Pro Thr Trp Arg Val Thr 180 185 190 Leu Lys Trp Glu Pro Ser Thr Glu Lys Met Tyr Val Glu Asp Gly Val 195 200 205 Leu Lys Gly Asp Val Glu Met Ala Leu Leu Leu Phe Ile Tyr Asn Lys 210 215 220 Ile Val Asp Glu Pro Gly Gly Ser Gly Thr Ser Tyr Trp Lys Gly Ser 225 230 235 240 His Tyr Arg Cys Asp Phe Lys Thr Thr Tyr Lys Ala Lys Asn Pro Leu 245 250 255 Leu Tyr Asp Ala Asn Tyr His Glu Val Asp His Arg Ile Glu Ile Leu 260 265 270 Ser His Gly Gly Ser Gly Gly Glu Ala Ala Lys Gly Arg Pro Gln Gly 275 280 285 Leu Pro Asn Asn Thr Ala Ser Lys Val Arg Leu Tyr Glu His Ala Glu 290 295 300 Ala Leu Ala Glu Ile Leu Gln Lys Asn Leu Ile Arg Gln Gly Thr Asp 305 310 315 320 Tyr Lys His Trp Pro Gln Ile Ala Gly Gly Ser Gly Gly Gly Lys Ile 325 330 335 Ala Asp Tyr Asn Tyr Lys Leu Gly His His His His His His Lys Lys 340 345 350 <210> 124 <211> 15 <212> PRT <213> Artificial Sequence <400> 124 Phe Glu Arg Asp Ile Ser Thr Glu Ile Tyr Gln Ala Gly Ser Thr 1 5 10 15 <210> 125 <211> 15 <212> PRT <213> Artificial Sequence <400> 125 Gly Ser Thr Pro Cys Asn Gly Val Glu Gly Phe Asn Cys Tyr Phe 1 5 10 15 <210> 126 <211> 15 <212> PRT <213> Artificial sequence <400> 126 Asn Ser Asn Asn Leu Asp Ser Lys Val Gly Gly Asn Tyr Asn Tyr 1 5 10 15 <210> 127 <211> 10 <212> PRT <213> Artificial sequence <400> 127 Gly Ser Ser Gly Val Val Asn Pro Val Met 1 5 10 <210> 128 <211> 16 <212> PRT <213> Artificial sequence <400> 128 Asn Ala Pro Arg Ile Thr Phe Gly Gly Pro Ser Asp Ser Thr Gly Ser 1 5 10 15 <210> 129 <211> 15 <212> PRT <213> Artificial sequence <400> 129 Tyr Ile Arg Val Gly Ala Arg Lys Ser Ala Pro Leu Ile Glu Leu 1 5 10 15 <210> 130 <211> 15 <212> PRT <213> Artificial sequence <400> 130 Ser Leu Ile Asp Leu Gln Glu Leu Gly Lys Tyr Glu Gln Tyr Ile 1 5 10 15 <210> 131 <211> 15 <212> PRT <213> Artificial sequence <400> 131 Pro Phe Gln Gln Phe Gly Arg Asp Ile Ala Asp Thr Thr Asp Ala 1 5 10 15 <210> 132 <211> 12 <212> PRT <213> Artificial sequence <400> 132 Met Trp Leu Ser Tyr Phe Ile Ala Ser Phe Arg Leu 1 5 10 <210> 133 <211> 5 <212> PRT <213> Artificial sequence <400> 133 Ile Val Asp Glu Pro 1 5 <210> 134 <211> 12 <212> PRT <213> Artificial sequence <400> 134 Gly Phe Ser Ala Leu Glu Pro Leu Val Asp Leu Pro 1 5 10 <210> 135 <211> 13 <212> PRT <213> Artificial sequence <400> 135 Lys Thr Phe Pro Pro Thr Glu Pro Lys Lys Asp Lys Lys 1 5 10 <210> 136 <211> 19 <212> PRT <213> Artificial sequence <400> 136 Leu Gly Val Tyr His Lys Asn Asn Lys Ser Trp Met Glu Ser Glu Phe 1 5 10 15 Arg Val Tyr <210> 137 <211> 15 <212> PRT <213> Artificial sequence <400> 137 Phe Asn Cys Tyr Phe Pro Leu Gln Ser Tyr Gly Phe Gln Pro Thr 1 5 10 15 <210> 138 <211> 10 <212> PRT <213> Artificial sequence <400> 138 Pro Leu Gln Ser Tyr Gly Phe Gln Pro Thr 1 5 10 <210> 139 <211> 15 <212> PRT <213> Artificial sequence <400> 139 Ala Gly Asn Gly Gly Asp Ala Ala Leu Ala Leu Leu Leu Leu Asp 1 5 10 15 <210> 140 <211> 11 <212> PRT <213> Artificial sequence <400> 140 Arg Ser Tyr Leu Thr Pro Gly Asp Ser Ser Ser 1 5 10 <210> 141 <211> 10 <212> PRT <213> Artificial sequence <400> 141 Ala Asp Gln Leu Thr Pro Thr Trp Arg Val 1 5 10 <210> 142 <211> 10 <212> PRT <213> Artificial sequence <400> 142 Phe Ile Tyr Asn Lys Ile Val Asp Glu Pro 1 5 10 <210> 143 <211> 10 <212> PRT <213> Artificial sequence <400> 143 Lys Asn Pro Leu Leu Tyr Asp Ala Asn Tyr 1 5 10 <210> 144 <211> 11 <212> PRT <213> Artificial sequence <400> 144 Arg Pro Gln Gly Leu Pro Asn Asn Thr Ala Ser 1 5 10 <210> 145 <211> twenty three <212> PRT <213> Artificial sequence <400> 145 Leu Ala Glu Ile Leu Gln Lys Asn Leu Ile Arg Gln Gly Thr Asp Tyr 1 5 10 15 Lys His Trp Pro Gln Ile Ala 20 <210> 146 <211> 10 <212> PRT <213> Artificial sequence <400> 146 Gly Lys Ile Ala Asp Tyr Asn Tyr Lys Leu 1 5 10 <210> 147 <211> 15 <212> PRT <213> Artificial sequence <400> 147 Gly Ser Ser Gly Val Val Asn Pro Val Met Glu Pro Ile Tyr Asp 1 5 10 15 <210> 148 <211> 15 <212> PRT <213> Artificial sequence <400> 148 Ala Pro Gly Gln Thr Gly Lys Ile Ala Asp Tyr Asn Tyr Lys Leu 1 5 10 15 <210> 149 <211> 15 <212> PRT <213> Artificial sequence <400> 149 Ala Leu Pro Gln Arg Gln Lys Lys Gln Gln Thr Val Thr Leu Leu 1 5 10 15 <210> 150 <211> 10 <212> PRT <213> Artificial sequence <400> 150 Gly Ser Lys Ser Pro Ile Gln Tyr Ile Asp 1 5 10 <210> 151 <211> 14 <212> PRT <213> Artificial sequence <400> 151 Asp Phe Leu Glu Tyr His Asp Val Arg Val Val Leu Asp Phe 1 5 10 <210> 152 <211> 10 <212> PRT <213> Artificial sequence <400> 152 Gly Ile Asn Ala Ser Val Val Asn Ile Gln 1 5 10 <210> 153 <211> 10 <212> PRT <213> Artificial sequence <400> 153 Gln Phe Ala Pro Ser Ala Ser Ala Phe Phe 1 5 10 <210> 154 <211> 10 <212> PRT <213> Artificial sequence <400> 154 Pro Ser Gly Thr Trp Leu Thr Tyr Thr Gly 1 5 10 <210> 155 <211> 5 <212> PRT <213> Artificial sequence <400> 155 Glu Leu Asp Lys Tyr 1 5 <210> 156 <211> 10 <212> PRT <213> Artificial sequence <400> 156 Gly Ile Tyr Gln Thr Ser Asn Phe Arg Val 1 5 10 <210> 157 <211> 15 <212> PRT <213> Artificial sequence <400> 157 Lys Ala Tyr Asn Val Thr Gln Ala Phe Gly Arg Arg Gly Pro Glu 1 5 10 15 <210> 158 <211> 10 <212> PRT <213> Artificial sequence <400> 158 Gly Thr Asn Thr Ser Asn Gln Val Ala Val 1 5 10 <210> 159 <211> 15 <212> PRT <213> Artificial sequence <400> 159 Asn Pro Val Leu Pro Phe Asn Asp Gly Val Tyr Phe Ala Ser Thr 1 5 10 15 <210> 160 <211> 10 <212> PRT <213> Artificial sequence <400> 160 Tyr Asn Tyr Lys Leu Pro Asp Asp Phe Thr 1 5 10 <210> 161 <211> 15 <212> PRT <213> Artificial sequence <400> 161 Met Phe His Leu Val Asp Phe Gln Val Thr Ile Ala Glu Ile Leu 1 5 10 15 <210> 162 <211> 10 <212> PRT <213> Artificial sequence <400> 162 Met Lys Asp Leu Ser Pro Arg Trp Tyr Phe 1 5 10 <210> 163 <211> 10 <212> PRT <213> Artificial sequence <400> 163 Asp Ala Ala Leu Ala Leu Leu Leu Leu Asp 1 5 10
Claims
1. A protein container, characterized in that, It includes a stable protein structure that supports the simultaneous insertion of more than four exogenous multi - amino acid sequences at different sites. Among them, the amino acid sequence of the stable protein structure is shown as SEQ ID NO: 1, and the amino acid sequence of the protein container is shown as SEQ ID NO:
18.
2. A protein container, characterized in that, It includes a stable protein structure that supports the simultaneous insertion of more than four exogenous multi - amino acid sequences at different sites. Among them, the amino acid sequence of the stable protein structure is shown as SEQ ID NO: 3, and the amino acid sequence of the protein container is shown as SEQ ID NO:
20.
3. A protein container, characterized in that, It includes a stable protein structure that supports the simultaneous insertion of more than four exogenous multi - amino acid sequences at different sites. Among them, the amino acid sequence of the stable protein structure is shown as SEQ ID NO: 3, and the amino acid sequence of the protein container is shown as SEQ ID NO:
33.
4. A protein container, characterized in that, It includes a stable protein structure that supports the simultaneous insertion of more than four exogenous multi - amino acid sequences at different sites. Among them, the amino acid sequence of the stable protein structure is shown as SEQ ID NO: 3, and the amino acid sequence of the protein container is shown as SEQ ID NO: 45 or SEQ ID NO:
51.
5. A protein container, characterized in that, It includes a stable protein structure that supports the simultaneous insertion of more than four exogenous multi - amino acid sequences at different sites. Among them, the amino acid sequence of the stable protein structure is shown as SEQ ID NO: 3, and the amino acid sequence of the protein container is shown as SEQ ID NO:
64.
6. The protein container according to any one of claims 1 to 5, characterized in that, It has insertion sites for exogenous multi - amino acid sequences in the protein loops facing the external environment.
7. The protein container according to any one of claims 1 to 5, characterized in that, The simultaneous insertion of exogenous multi - amino acid sequences does not interfere with the production conditions of the container protein.
8. The protein container according to any one of claims 1 to 5, characterized in that, It simultaneously contains exogenous multi - amino acid sequences for the development of vaccine compositions, for diagnosis, or for laboratory reagents.
9. The protein container according to any one of claims 1 to 5, characterized in that, The exogenous multi - amino acid sequences do not lose their immunogenic properties when simultaneously inserted into the protein loops of the protein container.
10. A polynucleotide, characterized in that, It includes any one of SEQ ID NO: 2, 4, 17, 19, 34, 46, 52, 63, and can separately produce the polypeptides defined by SEQ ID NO: 1, 3, 18, 20, 33, 45, 51, 64.
11. A carrier, characterized in that, It includes the polynucleotide as defined in claim 10.
12. An expression cassette, characterized in that, It includes the polynucleotide as defined in claim 10.
13. A cell, characterized in that, It includes the vector as defined in claim 11 or the expression cassette as defined in claim 12.
14. A method for producing a protein container, characterized in that, It introduces the polynucleotide as defined in claim 10 into competent cells of interest; cultures the competent cells and isolates the container protein containing the selected exogenous multi - amino acid.
15. The method for producing a protein container according to claim 14, characterized in that, It is not interfered by the insertion of various exogenous multi - amino acid sequences.
16. Use of the protein container as defined in any one of claims 1 to 9, characterized in that, The protein container is a laboratory reagent.
17. Use of the protein container as defined in claim 1 in the preparation of a medicament for diagnosing Chagas disease.
18. Use of the protein container as defined in claim 2 in the preparation of a medicament for diagnosing rabies.
19. Use of the protein container as defined in claim 3 in the preparation of a medicament for diagnosing an infection by Oropouche virus.
20. Use of the protein container as defined in claim 4 in the preparation of a medicament for diagnosing an infection by Mayaro virus.
21. Use of the protein container as defined in claim 5 in the preparation of a medicament for diagnosing whooping cough.
22. A diagnostic kit, characterized in that, It comprises a protein container as defined in any one of claims 1 to 9.
Citation Information
Patent Citations
Green fluorescent protein fusions with random peptides
US20010003650A1
Structurally biased random peptide libraries based on different scaffolds
US20030224412A1
Fusion proteins of superfolder green fluorescent protein and use thereof
US20180016310A1
A Multi-Functional Peptide Benefiting Expression, Purification, Stabilization and Catalytic efficiency of Recombinant Proteins
US20180298058A1
Modified green fluorescent proteins
US5625048A