Therapeutic papilloma virus mRNA vaccine and application thereof
By designing a fusion protein vaccine containing HPV16-E6, HPV16-E7, HPV18-E6, and HPV18-E7 proteins, and combining it with an immunostimulatory complex and optimized signal peptides and UTR sequences, a liposome-encapsulated mRNA vaccine was prepared. This solved the problem that existing vaccines could not effectively treat HPV-related diseases and achieved a highly efficient tumor suppression effect.
Patent Information
- Application Number
- CN202410951659.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-15
- Publication Date
- 2026-01-16
AI Technical Summary
Existing preventive HPV vaccines are ineffective in preventing cancers caused by HPV infection in people who are ineligible for vaccination, and there is a lack of effective therapeutic vaccines to treat HPV-related diseases.
A fusion protein vaccine was developed, comprising HPV16-E6, HPV16-E7, HPV18-E6 and HPV18-E7 proteins or their immunogenic fragments, and combined with an immunostimulatory complex and optimized signal peptide and UTR sequence. The mRNA vaccine was prepared by encapsulation with liposome nanoparticles to elicit a TH1-type cellular immune response.
Under the same immune conditions, it significantly improved serum antibody levels and achieved highly efficient inhibition of HPV-related tumors, especially in mouse models where it achieved a tumor elimination rate of over 90%.
Smart Images

Figure BDA0004948327390000061 
Figure BDA0004948327390000062 
Figure BDA0004948327390000064
Abstract
Description
Technical Field
[0001] This invention relates to the field of biopharmaceutical technology, and in particular to therapeutic HPV mRNA vaccines and their preparation methods. Background Technology
[0002] Cancer has a variety of causes, among which bacterial and viral infections are known to be important factors inducing cancer development. Globally, the leading sources of cancer infection are Helicobacter pylori (36.3%) and HPV (31.1%), followed by hepatitis B (16.3%) and hepatitis C (7.1%). Other sources include Epstein-Barr virus (EBV), human herpesvirus 8 (HHV-8), and human T-cell lymphotropic virus 1 (HTLV-1). These infection-mediated cancer factors account for approximately 10% of all cancer development.
[0003] Human papillomavirus (HPV) is a small, non-enveloped, round, double-stranded DNA virus, 50-55 nanometers in diameter. Currently, over 300 different HPV genotypes have been identified, of which more than 200 have been confirmed to be harmful to humans. By comparing the nucleotide sequences of the L1ORF of 118 viruses, de Villiers et al. classified HPV into different genera, species, types, and subtypes. Currently, HPV is classified into five genera: α (65 types, including HPV16, 18, 31, 33, etc.), β (53 types, including HPV5, 9, 49, etc.), γ (98 types, including HPV4, 48, 50, etc.), mu (3 types, including HPV1, HPV63, and HPV204), and nu (HPV41). Most of the highly pathogenic strains that cause HPV-related cancers belong to the α genus. Based on their ability to infect epithelial skin cells or the inner lining of tissues, HPV can be classified into cutaneous and mucosal types. Based on their association with cervical cancer or precancerous lesions, HPV can be further divided into low-risk and high-risk types: (1) Low-risk skin types: including HPV1, 2, 3, 4, 7, 10, 12, 15, etc., which are associated with common warts, flat warts, plantar warts, etc.; (2) High-risk skin types: including HPV5, 8, 14, 17, 20, 36, 38, etc., which are associated with verrucous epidermal dysplasia and also with some malignant tumors, including vulvar cancer, penile cancer, anal cancer, prostate cancer, bladder cancer; (3) Low-risk mucosal types such as HPV6, 11, 13, 32, 34, 40, 42, 43, 44, 54, etc., which are associated with infection of the genital, anal, oropharyngeal, and esophageal mucosa; (4) High-risk mucosal types HPV16, 18, 30, 31, 33, 35, 53, 39, which are associated with cervical cancer, rectal cancer, oral cancer, tonsil cancer, etc.
[0004] Currently, HPV-related cancers can be prevented through vaccination, such as the bivalent vaccine Cervarix (GlaxoSmithKline), the quadrivalent vaccine Gardasil (Merck), and the nonavalent vaccine Gardasil 9 (Merck). However, because preventative vaccines are only effective in the early stages, they cannot prevent HPV-related cancers in populations who are unable to receive vaccinations due to poverty or other reasons. Actively developing therapeutic vaccines for HPV-related cancers could make infection more beneficial. Current research indicates that targeting the major oncogenes (E6 and E7) that drive HPV-related carcinogenesis can effectively inhibit the development of HPV-induced cancers.
[0005] The protein encoded by E6 contains 150-160 amino acids and has a molecular weight of approximately 18 kDa. It contains two zinc finger binding domains, and its C-terminal PDZ binding motif can interact with several cellular proteins, the most important of which is the degradation of p53 by E6. E7 is a relatively small phosphoprotein, containing approximately 100 amino acids, and has three conserved regions 1 / 2 / 3 (CR1 / 2 / 3). The C-terminal CR3 region of E7 is conserved and encodes two zinc finger domains responsible for zinc-dependent dimerization and mediating the interaction between E7 and p21 and pRb, regulating cell cycle and apoptosis. Therefore, E6 and E7 play a crucial role in driving cell carcinogenesis; they can induce uncontrolled cell proliferation, angiogenesis, invasion, metastasis, and unrestricted telomerase activity, as well as evasion of apoptosis and growth inhibitory factors, ultimately leading to cell carcinogenesis. Several in vitro and xenograft studies have also shown that HPV-positive cancer cells senescent or undergo apoptosis in the absence of E6 and E7 activity, demonstrating the crucial role of E6 and E7 in HPV-mediated cancer persistence. Therefore, E6 and E7 are effective targets for the treatment of HPV-positive cancers.
[0006] There is still a need in this field for vaccines that can treat human papillomavirus-related diseases. Summary of the Invention
[0007] This invention is partly based on the inventors' findings that by introducing a combination of multiple immunostimulatory complexes, T cell activation and immunogenicity can be further enhanced, TH1-type cellular immune responses can be elicited, and tumor suppression rates can be increased. This invention also includes overall sequence optimization of the 5' and 3' UTRs, signal peptides, and open reading frames (ORFs) encoding antigens.
[0008] In one aspect, a fusion protein is provided that comprises one or more of the following papillomavirus proteins or immunogenic fragments thereof: HPV16-E6, HPV16-E7, HPV18-E6, and HPV18-E7.
[0009] In one embodiment, when the fusion protein comprises multiple papillomavirus proteins or immunogenic fragments thereof, the papillomavirus proteins or immunogenic fragments thereof are directly linked or linked via adapters. The adapter can be any suitable adapter, such as those defined herein.
[0010] In one embodiment, the HPV16-E7 protein comprises the amino acid sequence of SEQ ID NO:1 or has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence of SEQ ID NO:1.
[0011] In one embodiment, the HPV16-E6 protein comprises the amino acid sequence of SEQ ID NO:2 or has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence of SEQ ID NO:2.
[0012] In one embodiment, the HPV18-E7 protein comprises the amino acid sequence of SEQ ID NO:3 or has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence of SEQ ID NO:3.
[0013] In one embodiment, the HPV18-E6 protein comprises the amino acid sequence of SEQ ID NO:4 or has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence of SEQ ID NO:4.
[0014] In one embodiment, the fusion protein further comprises a stimulant, which is one or more of MITD, LAMP, CRT, IL-2, OX40, IL7, HSP70, KDEL, CD28, ICOS, 4-1BBL, IL1, HSP40, and EDA.
[0015] In one embodiment, the fusion protein further comprises a secretory signal peptide, which is an IgG signal peptide, an IL-2 signal peptide, a tPA signal peptide, an Ig kappa signal peptide, or a SigMHC signal peptide.
[0016] In one embodiment, the fusion protein comprises a first HPV protein or an immunogenic fragment thereof, an optional second HPV protein or an immunogenic fragment thereof, a first stimulant, an optional adapter, and an optional second stimulant. In one embodiment, the first HPV protein is HPV16-E6, HPV16-E7, or a fusion thereof, and the second HPV protein is HPV18-E6, HPV18-E7, or a fusion thereof. In another embodiment, the first HPV protein is HPV18-E6, HPV18-E7, or a fusion thereof, and the second HPV protein is HPV16-E6, HPV16-E7, or a fusion thereof. In some embodiments, the fusion may be formed by direct linking of the components or by linking them through any suitable adapter.
[0017] In one embodiment, the LAMP stimulant comprises the amino acid sequence of SEQ ID NO:10 or has an amino acid sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence of SEQ ID NO:10.
[0018] In one embodiment, the MITD stimulant comprises the amino acid sequence of SEQ ID NO:11 or has an amino acid sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of SEQ ID NO:11.
[0019] In one embodiment, the KDEL stimulant comprises the amino acid sequence of SEQ ID NO:12 or has an amino acid sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of SEQ ID NO:12.
[0020] In one embodiment, the OX40 stimulant comprises the amino acid sequence of SEQ ID NO:13 or has an amino acid sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of SEQ ID NO:13.
[0021] In one embodiment, the CD28 stimulant comprises the amino acid sequence of SEQ ID NO:14 or has an amino acid sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of SEQ ID NO:14.
[0022] In one embodiment, the ICOS stimulant comprises the amino acid sequence of SEQ ID NO:15 or has an amino acid sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of SEQ ID NO:15.
[0023] In one embodiment, the 4-1BBL stimulant comprises the amino acid sequence of SEQ ID NO:16 or has an amino acid sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of SEQ ID NO:16.
[0024] In one embodiment, the CRT stimulant comprises the amino acid sequence of SEQ ID NO:17 or has an amino acid sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of SEQ ID NO:17.
[0025] In one embodiment, the EDA stimulant comprises the amino acid sequence of SEQ ID NO:18 or has an amino acid sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of SEQ ID NO:18.
[0026] In one embodiment, the HSP70 stimulant comprises the amino acid sequence of SEQ ID NO:19 or has an amino acid sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence of SEQ ID NO:19.
[0027] In one embodiment, the IL1 stimulant comprises the amino acid sequence of SEQ ID NO:20 or has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence of SEQ ID NO:20.
[0028] In one embodiment, the IL2 stimulant comprises the amino acid sequence of SEQ ID NO:21 or has an amino acid sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of SEQ ID NO:21.
[0029] In one embodiment, the IL7 stimulant comprises the amino acid sequence of SEQ ID NO:22 or has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence of SEQ ID NO:22.
[0030] In one embodiment, the HSP40 stimulant comprises the amino acid sequence of SEQ ID NO:23 or has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence of SEQ ID NO:23.
[0031] In one embodiment, the IgG signal peptide comprises the amino acid sequence of SEQ ID NO:5 or has an amino acid sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of SEQ ID NO:5.
[0032] In one embodiment, the IL-2 signal peptide comprises the amino acid sequence of SEQ ID NO:6 or has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence of SEQ ID NO:6.
[0033] In one embodiment, the tPA signal peptide comprises the amino acid sequence of SEQ ID NO:7 or has an amino acid sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of SEQ ID NO:7.
[0034] In one embodiment, the Ig kappa signal peptide comprises the amino acid sequence of SEQ ID NO:8 or has an amino acid sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of SEQ ID NO:8.
[0035] In one embodiment, the SigMHC signal peptide comprises the amino acid sequence of SEQ ID NO:9 or has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence of SEQ ID NO:9.
[0036] In another aspect, an isolated nucleic acid is provided that encodes the fusion protein described herein.
[0037] In one implementation, the nucleic acid is a single-stranded or double-stranded DNA molecule or RNA molecule.
[0038] In one implementation, the nucleic acid also includes a 5'UTR and a 3'UTR.
[0039] In one embodiment, the RNA molecule is an mRNA molecule comprising a nucleic acid sequence encoding a first HPV protein or an immunogenic fragment thereof, optionally a nucleic acid sequence encoding a second HPV protein or an immunogenic fragment thereof, a nucleic acid sequence encoding a first stimulant, optionally a linker, optionally a nucleic acid sequence encoding a second stimulant, and further comprising a 5' cap and a Poly-A tail. In one embodiment, the first HPV protein is HPV16-E6, HPV16-E7, or a fusion thereof, and the second HPV protein is HPV18-E6, HPV18-E7, or a fusion thereof. In another embodiment, the first HPV protein is HPV18-E6, HPV18-E7, or a fusion thereof, and the second HPV protein is HPV16-E6, HPV16-E7, or a fusion thereof. In some embodiments, the fusion may be formed by direct linking of the components or by linking them through any suitable linker.
[0040] In one embodiment, the mRNA molecule comprises a 5' cap - 5' UTR - nucleic acid sequence encoding a secretion signal peptide - nucleic acid sequence encoding a first HPV protein or an immunogenic fragment thereof - optionally a nucleic acid sequence encoding a second HPV protein or an immunogenic fragment thereof - nucleic acid sequence encoding a first stimulant - optional adapter - optional nucleic acid sequence encoding a second stimulant - 3' UTR - Poly-A tail.
[0041] In one embodiment, the 5' cap is a compound of formula (I), or a pharmaceutically acceptable salt, stereoisomer, tautomer, or isotopic variant thereof:
[0042]
[0043]
[0044] It is a single key or does not exist.
[0045] X1 is selected from O, S, CH2, CH2CH2, CH=CH, CH=CHO, CH2O, OCH2.
[0046] CH2CH2O, OCH2CH2, three-membered cycloalkyl group,
[0047] R1, R2, R3, and R4 are independently halogenated, OH-, unsubstituted, or OC-substituted, respectively. 1-3 Alkyl-substituted OC 1-3 Alkyl, unsubstituted or OC1-3 Alkyl-substituted OC 1-3 alkyl,
[0048] B1 and B2 are independently selected from natural, modified, or non-natural nucleoside bases, respectively.
[0049] In one embodiment, the compound of formula (I) is any one of the following:
[0050]
[0051]
[0052] In one embodiment, the 5'UTR contains a polynucleotide sequence of SEQ ID NO: 90, 92, 94, 96, 98, 100, 102 or 104 or a polynucleotide sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% of SEQ ID NO: 90, 92, 94, 96, 98, 100, 102 or 104.
[0053] In one embodiment, the 3'UTR contains a polynucleotide sequence of SEQ ID NO: 91, 93, 95, 97, 99, 101, 103 or 105 or a polynucleotide sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% of SEQ ID NO: 91, 93, 95, 97, 99, 101, 103 or 105.
[0054] In one embodiment, the 5'UTR and 3'UTR are selected from the following sequences or combinations thereof with variant polynucleotide sequences having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the following sequences: SEQ ID NO: 90 and 91; SEQ ID NO: 92 and 93; SEQ ID NO: 94 and 95; SEQ ID NO: 96 and 97; SEQ ID NO: 98 and 99; SEQ ID NO: 100 and 101; SEQ ID NO: 102 and 103; and SEQ ID NO: 104 and 105.
[0055] In one embodiment, the Poly-A tail contains a polynucleotide sequence of SEQ ID NO:106 or 107.
[0056] In one embodiment, the nucleic acid is an mRNA molecule containing one or more of the nucleic acid sequences of SEQ ID NO:57-89, 108, and 109.
[0057] In one embodiment, the RNA is modified RNA, wherein uracil, cytosine, adenine, or guanine nucleotides contain modifying groups.
[0058] In one embodiment, the modifying group is selected from at least one of pseudouridine, N1-methylpseudouridine, N1-ethylpseudouridine, 5-methylcytosine, 5-methoxycytosine, N1-methylcytosine, 2-thiouridine, 5-methoxyuridine, or N1-methyladenosine, N1-methylguanine, N1-methylguanine, and isoguanine.
[0059] In one embodiment, the mRNA molecule is a modified mRNA, the modification including the conversion of uridine nucleoside to pseudouridine, N1-methylpseudouridine, N1-ethylpseudouridine, 2-thiouridine, or 5-methoxyuridine; and / or, the conversion of cytosine nucleoside to 5-methylcytosine, 5-methoxycytosine, or N1-methylcytosine; and / or, the conversion of adenine nucleoside to N1-methyladenosine; and / or, the conversion of adenine nucleoside to N1-methylguanine, N1-methylguanine, or isoguanine.
[0060] In another aspect, a carrier is provided that contains the nucleic acid described herein.
[0061] In another aspect, a composition is provided comprising one or more nucleic acids described herein.
[0062] In another aspect, compositions are provided that comprise mRNA molecules containing one or more of the nucleic acid sequences of SEQ ID NO:57-89, 108 and 109.
[0063] In one embodiment, the composition further comprises a pharmaceutically acceptable excipient and / or adjuvant.
[0064] In one embodiment, the composition further comprises liposome nanoparticles.
[0065] In one embodiment, the liposome nanoparticles comprise ionizable lipids, DSPC, cholesterol, and DMG-PEG2000 ethanol.
[0066] In one embodiment, the composition is a vaccine. In another embodiment, the vaccine is an mRNA vaccine.
[0067] In another aspect, a method for preparing the nucleic acid described herein is provided, comprising one or more of the following steps:
[0068] (1) Using linear double-stranded DNA containing a promoter sequence as a template, and ATP, GTP, CTP, and N1-Me-pUTP as substrates, RNA polymerase transcribes the DNA sequence downstream of the promoter to synthesize mRNA; and
[0069] (2) The synthesized mRNA was capped using a one-step chemical method with capping analogs CAP m7Gppp(2'OMeA)pG or CAP5m7G(5')vppp(5')(2'OMeA)pG.
[0070] In one implementation, the promoter is the T7 promoter.
[0071] In one implementation, the RNA polymerase is T7 RNA polymerase.
[0072] In another aspect, the use of nucleic acids described in this article as mRNA vaccines is provided.
[0073] In another aspect, the use of the nucleic acids described herein in the preparation of medicaments for the treatment or prevention of HPV infection or HPV-related diseases is provided.
[0074] In one implementation, HPV is one or more of HPV16 and HPV18.
[0075] In one implementation, HPV infection-related disease is one or more of the following skin and mucous membrane lesions: HPV infection-related tumors or cancers, common warts, genital warts, respiratory warts, plantar warts, sublingual warts, perilingual warts, or flat warts.
[0076] In one implementation, the cancer is one or more of the following: cervical cancer, oropharyngeal cancer, vulvar cancer, anal cancer, penile cancer, vaginal cancer, rectal cancer, squamous cell carcinoma, adenocarcinoma, and head and neck cancer.
[0077] In another aspect, methods for treating or preventing HPV infection or HPV-related diseases are provided, including administering the nucleic acid or vaccine described herein to the subject.
[0078] In one implementation, HPV is one or more of HPV16 and HPV18.
[0079] In one implementation, HPV infection-related disease is one or more of the following skin and mucous membrane lesions: HPV infection-related tumors or cancers, common warts, genital warts, respiratory warts, plantar warts, sublingual warts, perilingual warts, or flat warts.
[0080] In one embodiment, the cancer is one or more of cervical cancer, oropharyngeal cancer, vulvar cancer, anal cancer, penile cancer, vaginal cancer, rectal cancer, squamous cell carcinoma, adenocarcinoma, and head and neck cancer.
[0081] The main advantages of this invention are: (1) Selecting the E6 and E7 proteins of human papillomavirus (HPV) as antigens to develop an optimized mRNA vaccine encapsulated by liposome nanoparticles. (2) Providing 33 preferred sequences of amino acid and nucleotide sequences for designing mRNA expressing the E6 and E7 proteins of HPV and the immunostimulatory complex. (3) Under the same immunization conditions, the mRNA vaccine of this invention can effectively express the dual antigen proteins E6 and E7 simultaneously, and after preparing the mRNA vaccine by LNP encapsulation and immunizing mice, it can induce a high level of serum antibody and stimulate the body to produce a TH1-type cellular immune response. (4) Under the same immunization conditions, the mRNA vaccine of this invention can achieve a tumor elimination rate of more than 90% in the TC-1 tumor-bearing mouse model at a dose of 10 μg. (5) Through a large number of screening experiments, the present invention has obtained the optimal combination of elements for mRNA molecules: 5'UTR and 3'UTR (SEQ ID NO: 90 and 91), signal peptide tPA or SigMHC, first stimulant KDEL + second stimulant EDA and CAP5 m7G(5')vppp(5')(2'OMeA)pG. Detailed Implementation
[0082] The following definitions are provided to enable those skilled in the art to understand the invention. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. While any methods and materials similar to or equivalent to those described herein may be used in the practice of testing the invention, preferred materials and methods are described herein. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.
[0083] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. For example, terms used herein include those in Janeway CAJr, Travers P, Walport M, et al., *Immunobiology*, 5th edition, New York: Garland Science (2001), and “A multilingual glossary of biotechnological terms: (IUPAC Recommendations)”, Leuenberger, HGW, Nagel B., and… The definition is given in H. (ed., 1995), Helvetica Chimica Acta, CH-4010 Basel, Switzerland.
[0084] The term "fusion protein" refers to a protein formed by covalently linking two protein motifs that do not exist together under natural conditions. In this document, fusion proteins comprise one or more immunogenic proteins selected from HPV16-E6, HPV16-E7, HPV18-E6, and HPV18-E7, or one or more immunogenic fragments thereof. These HPV16 and HPV18E6 and E7 proteins are directly linked or linked via suitable adapters. The choice of adapter is conventional to those skilled in the art. For example, adapters may comprise, but are not limited to, (GGGGS)n or (G)n, where n is greater than or equal to 1, for example, 2-10. Fusion proteins may also comprise secretion signal peptides and / or stimulants. In some embodiments, the adapter may be a self-cleaving peptide adapter, such as a P2A peptide, an E2A peptide, or an F2A peptide.
[0085] The term "linker" refers to a segment of amino acids that acts as a connector when two or more parts are present, enabling each part to function. When fusion proteins or fusion products are referred to herein, those skilled in the art will understand that the various parts can be linked by suitable linkers. The choice of linker is conventional for those skilled in the art. For example, linkers may contain, but are not limited to, (GGGGS)n or (G)n, where n is greater than or equal to 1, for example, 2-10. In some embodiments, the linker may be a self-cleaving peptide linker, such as a P2A peptide, an E2A peptide, or an F2A peptide.
[0086] The term "secretion signal peptide" refers to a peptide used to guide the translocation of a synthesized fusion protein into the secretory pathway. Signal peptides are generally essential for transmembrane translocation in the secretory pathway and thus universally control the entry of most proteins into the secretory pathway in eukaryotes and prokaryotes. In eukaryotes, the signal peptide of a nascent precursor protein (preprotein) guides ribosomes to the rough endoplasmic reticulum (ER) membrane and initiates the transport of growing peptide chains across this membrane for processing. ER processing produces mature proteins, in which the signal peptide is typically cleaved from the precursor protein by ER-resident signal peptidases of the host cell, or remains unclew and acts as a membrane anchor. Signal peptides can also promote protein targeting to the cell membrane. Secretion signal peptides can be located at the N-terminus of the fusion protein. The type of secretion signal peptide is not particularly limited, as long as it can guide the secretion of the synthesized fusion protein. Secretion signal peptides can include, but are not limited to, IgG signal peptides, IL-2 signal peptides, tPA signal peptides, Ig kappa signal peptides, and SigMHC signal peptides.
[0087] The term "stimulant" herein refers to a protein used to stimulate an immune response in response to the fusion protein. The selection of stimulants is conventional and will be understood by those skilled in the art. Stimulants may include, but are not limited to, MITD, LAMP, CRT, IL-2, OX40, IL7, HSP70, KDEL, CD28, ICOS, 4-1BBL, IL1, HSP40, and EDA. Stimulants may be located at any suitable position on the fusion protein, for example, at the N-terminus or C-terminus of the immunogenic protein or its immunogenic fragment, preferably at the C-terminus. The fusion protein may contain multiple stimulants, such as two, three, or four. When multiple stimulants are included, the types of stimulants may be the same or different. For example, the fusion protein may contain two or more stimulants selected from MITD, LAMP, CRT, IL-2, OX40, IL7, HSP70, KDEL, CD28, ICOS, 4-1BBL, IL1, HSP40, and EDA. The stimulants may be linked by linkers, such as self-cleaving peptides. The family of self-cleaving peptide linkers (referred to as 2A peptides) is described in the art, for example, in Kim, JH et al., (2011) PLoS ONE 6: el8556. 2A peptides can be P2A peptides, E2A peptides, or F2A peptides.
[0088] The term "immunogenic fragment" refers to a fragment of an immunogenic protein that can elicit an immune response in the body. The length of an immunogenic fragment is shorter than the full length of the immunogenic protein. In this document, those skilled in the art can determine the length and sequence of an immunogenic fragment based on one or more immunogenic proteins selected from HPV16-E6, HPV16-E7, HPV18-E6, and HPV18-E7. In some embodiments, the immunogenic fragment may be at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% but less than 100% of the length of any one of HPV16-E6, HPV16-E7, HPV18-E6, and HPV18-E7.
[0089] The terms “nucleic acid,” “polynucleotide,” and “polynucleotide sequence” are used interchangeably to refer to oligomers and polymers of any length that are essentially composed of nucleotides (such as deoxyribonucleotides and / or ribonucleotides). Nucleic acids can contain purine and / or pyrimidine bases and / or other natural (e.g., xanthine, inosine, hypoxanthine), chemically or biochemically modified (e.g., methylated), non-natural, or derived nucleotide bases. The backbone of a nucleic acid can contain sugar and phosphate groups normally present in RNA or DNA, and / or one or more modified or substituted sugars and / or one or more modified or substituted phosphate groups. Modifications to phosphate groups or sugars can be introduced to improve stability, resistance to enzymatic degradation, or some other useful properties. A “nucleic acid” can be, for example, double-stranded, partially double-stranded, or single-stranded. When single-stranded, a nucleic acid can be a sense strand or an antisense strand. A “nucleic acid” can be circular or linear. As used herein, the term “nucleic acid” encompasses DNA and RNA, including genomes, pre-mRNA, mRNA, cDNA, and recombinant or synthetic nucleic acids containing vectors. For the purposes described herein, it should be understood that polynucleotides can be modified by any method available in the art.
[0090] The term "sequence identity" refers to the degree to which two sequences (amino acid sequences) have identical residues at the same positions when aligned. Such calculations are typically performed using computer programs. Exemplary programs for comparing and aligning sequence pairs include ALIGN (Myers and Miller, 1988), FASTA (Pearson and Lipman, 1988; Pearson, 1990), and gapped BLAST (Altschul et al., 1997), BLASTP, BLASTN, or GCG (Devereux et al., 1984). Furthermore, in determining the degree of sequence identity between two amino acid sequences, those skilled in the art may consider so-called "conserved" amino acid substitutions, which can generally be described as amino acid substitutions in which an amino acid residue is replaced by another amino acid residue having a similar chemical structure, having little or no effect on the function, activity, or other biological properties of the polypeptide. Such conserved amino acid substitutions are well known in the art.
[0091] The term “identity” when used in conjunction with nucleic acids or fragments thereof means that, when an optimized alignment is performed with other nucleic acids (or their complementary strands), at least 50%, 60%, 70%, 80%, 90%, more preferably at least about 95%, 96%, 97%, 98%, or 99% of the nucleotide bases have nucleotide sequence identity, as determined by any sequence identity algorithm well known in the art (such as FASTA, BLAST, or GAP) discussed below.
[0092] A read frame (ORF) is a continuous extension of DNA or RNA that begins with a start codon (e.g., methionine (ATG or AUG)) and ends with a stop codon (e.g., TAA, TAG, or TGA, or UAA, UAG, or UGA). ORFs typically encode proteins. In this paper, a read frame may encode a fusion protein.
[0093] The term "5' untranslated region" (UTR) refers to an mRNA region located directly upstream (i.e., 5') of the start codon (i.e., the first codon of the mRNA transcript translated by the ribosome) and that does not encode a protein or peptide. When an RNA transcript is generated, the 5' UTR may contain promoter sequences. These promoter sequences are known in the art. It should be understood that these promoter sequences will not be present in the mRNA described herein.
[0094] The term "3' untranslated region" (UTR) refers to an mRNA region located directly downstream (i.e., 3') of a stop codon (i.e., the codon that transmits the translation termination signal in the mRNA transcript) and that does not encode a protein or peptide.
[0095] The term "Poly-A tail" refers to an mRNA region containing multiple consecutive adenosine monophosphates (ATPs) located downstream of the 3' UR. For example, a Poly-A tail can be located directly downstream of the 3' UR (i.e., 3'). A Poly-A tail can contain 10 to 300 ATPs, such as 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, or 300 ATPs. A Poly-A tail can protect mRNA, for example, from enzymatic degradation in the cytoplasm, and facilitate transcription termination and / or the export of mRNA from the nucleus for translation.
[0096] The term "ionizable lipid" refers to an amphiphilic molecule (e.g., a lipid or lipidoid, such as a synthetic lipid or lipidoid) containing a group (e.g., a head group) that can ionize, for example, dissociate under given conditions (e.g., pH) to produce one or more charged substances. In some embodiments, the ionizable lipid is SM-102.
[0097] The term "pharmaceutically acceptable" means a molecule or composition that, when administered to a recipient, is harmless to the recipient or provides a benefit to the recipient that outweighs any adverse effects. Regarding carriers or excipients used to formulate compositions as disclosed herein, a pharmaceutically acceptable carrier or excipient must be compatible with the other components of the composition and be harmless to the recipient, or provide a benefit to the recipient that outweighs any adverse effects. The term "pharmaceutically acceptable carrier" means a pharmaceutically acceptable material, composition, or medium, such as a liquid or solid filler, diluent, excipient, solvent, medium, encapsulating material, manufacturing aid (e.g., lubricant, magnesium talc, calcium stearate, or zinc or stearic acid), or solvent encapsulating material, which participates in carrying or transporting a pharmaceutical agent from one part of the body to another (e.g., from one organ to another). Each carrier must be "acceptable" in the sense that it is compatible with the other components of the formulation and is not harmful to the patient. Some examples of materials that can serve as pharmaceutically acceptable carriers include: (1) sugars, such as lactose, glucose, and sucrose; (2) starches, such as corn starch and potato starch; (3) cellulose and its derivatives, such as sodium carboxymethyl cellulose, methyl cellulose, ethyl cellulose, microcrystalline cellulose, and cellulose acetate; (4) powdered tragacanth gum; (5) malt; (6) gelatin; (7) excipients, such as cocoa butter and suppository waxes; (8) oils, such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil; (9) glycols, such as propylene glycol; and (10) polyols, such as glycerol and sorbitol. (11) Mannitol and polyethylene glycol (PEG); (12) Esters, such as ethyl oleate and ethyl laurate; (13) Agar; (14) Buffers, such as magnesium hydroxide and aluminum hydroxide; (15) Alginate; (16) Pyrothermic water; (17) Isotonic saline; (18) Ringer's solution; (19) pH buffer solution; (20) Polyesters, polycarbonates and / or polyanhydrides; (21) Fillers, such as peptides and amino acids; (22) Serum components, such as serum albumin, HDL and LDL; (23) C2-C12 alcohols, such as ethanol; and (24) Other non-toxic and compatible substances used in pharmaceutical formulations. In the context of this document, those skilled in the art may select a carrier suitable for acceptance on a pharmaceutically acceptable protein or nucleic acid as described herein.
[0098] Fusion protein
[0099] The fusion protein described herein may comprise one or more of HPV16-E6, HPV16-E7, HPV18-E6, and HPV18-E7. Mutations (e.g., deletions, additions, substitutions, or insertions) of certain amino acid residues may also be present in HPV16-E6, HPV16-E7, HPV18-E6, and / or HPV18-E7 to reduce potential oncogenicity or enhance immunogenicity. Preferably, when the fusion protein comprises multiple papillomavirus proteins or their immunogenic fragments, the individual papillomavirus proteins or their immunogenic fragments are directly linked or linked via adapters. Suitable adapters are known in the art, such as any suitable flexible adapter, such as any adapter described herein. A fusion protein may comprise a first HPV protein or an immunogenic fragment thereof, optionally a second HPV protein or an immunogenic fragment thereof, a first stimulant, an optional adapter, and an optional second stimulant. The first HPV protein may be HPV16-E6, HPV16-E7, or a fusion thereof, and the second HPV protein may be HPV18-E6, HPV18-E7, or a fusion thereof; or the first HPV protein may be HPV18-E6, HPV18-E7, or a fusion thereof, and the second HPV protein may be HPV16-E6, HPV16-E7, or a fusion thereof. The fusion may be formed by direct linking of the components or by linking them through any suitable adapter. The term "optional" indicates that the element may or may not be present. For example, "first HPV protein or an immunogenic fragment thereof - optional second HPV protein or an immunogenic fragment thereof - first stimulant - optional adapter - optional second stimulant" could represent six possible fusion proteins with or without a second HPV protein or its immunogenic fragment, with or without an adapter, and with or without a second stimulant.
[0100] The HPV16-E7 protein may contain the amino acid sequence of SEQ ID NO:1 or have at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence of SEQ ID NO:1. The HPV16-E6 protein may contain the amino acid sequence of SEQ ID NO:2 or have at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence of SEQ ID NO:2. The HPV18-E7 protein may contain the amino acid sequence of SEQ ID NO:3 or have at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence of SEQ ID NO:3. The HPV18-E6 protein may contain the amino acid sequence of SEQ ID NO:4 or have at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence of SEQ ID NO:4.
[0101] HPV16-E7 protein (SEQ ID NO:1)
[0102] HGDTPTLHEYMLDLQPETTDLYCYEQLNDSSEEEDEIDGPAGQAEPDRAHYNIVTFCCKCDSTLRLCVQSTHVDIRTLEDLLMGTLGIVCPICSQKP
[0103] HPV16-E6 protein (SEQ ID NO:2)
[0104] MHQKRTAMFQDPQERPRKLPQLCTELQTTIHDIILECVYCKQQLLRREVYDFAFRDLCIVYRDGNPYAVCDKCLKFYSKISEYRHYCYSLYGTTLEQQYNKPLCDLLIRCINCQKPLCPEEKQRHLDKKQRFHNIRGRWTGRCMSCCRSSRTRRETQL
[0105] HPV18-E7 protein (SEQ ID NO:3)
[0106] MHGPKATLQDIVLHLEPQNEIPVDLLCHEQLSDSEEENDEIDGVNHQHLPARRAEPQRHTMLCMCCKCEARIKLVVESSADDLRAFQQLFLNTLSFVCPWCASQQ
[0107] HPV18-E6 protein (SEQ ID NO:4)
[0108] MARFEDPTRRPYKLPDLCTELNTSLQDIEITCVYCKTVLELTEVFEFAFKDLFVVYRDSIPHAACHKCIDFYSRIRELRHYSDSVYGDTLEKLTNTGLYNLLIRCLRCQKPLNPAEKLRHLNEKRRFHNIAGHYRGQCHSCCNRARQERLQRRRETQV
[0109] As those skilled in the art will recognize, protein fragments, functional protein domains, and homologous proteins are also considered to fall within the scope of the papillomavirus immunogenic proteins of interest. For example, this document provides any protein fragment of a papillomavirus immunogenic protein (HPV16-E6, HPV16-E7, HPV18-E6, and / or HPV18-E7), provided that the fragment is immunogenic and confers a protective immune response against the papillomavirus. Immunogenic proteins may contain 2, 3, 4, 5, 6, 7, 8, 9, 10, or more mutations, except for truncated fragments identical to any of SEQ ID NO:1-4. The length of the fragment can range from about 4, 6, or 8 amino acids to the full-length protein.
[0110] The fusion protein described herein may also contain a stimulant, which is one or more of MITD, LAMP, CRT, IL-2, OX40, IL7, HSP70, KDEL, CD28, ICOS, 4-1BBL, IL1, HSP40, and EDA. The amino acid sequence of the stimulant may contain mutations (e.g., deletions, additions, substitutions, or insertions) of some amino acid residues. These mutations can enhance the stimulatory activity. The stimulant may contain any of the following sequences or have at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the following amino acid sequences.
[0111] LAMP (SEQ ID NO:10)
[0112] LDENSMLIPIAVGGALAGLVLIVLIAYLVGRKRSHAGYQTI
[0113] MITD (SEQ ID NO:11)
[0114] SQSTIPIVGIVAGLAVLAVVVIGAVVATVMCRRKSSGGKGGSYSQAASSDSAQGSDVSLTA
[0115] KDEL(SEQ ID NO:12)
[0116] KDELKDEL
[0117] OX40(SEQ ID NO:13)
[0118] AAILGLGLVLGLLGPLAILL
[0119] CD28(SEQ ID NO:14)
[0120] LVAYDNAVNLSCKYSYNLFSREFRASLHKGLDSAVEVCVVYGNYSQQLQVYSKTGFNCDGKLGNESVTFYLQNLYVNQTDIYFCKIEVMYPPPYLDNEKSNGTIIHVKG
[0121] ICOS(SEQ ID NO:15)
[0122] FWLPIGCAAFVVVCILGCILI
[0123] 4-1BBL(SEQ ID NO:16)
[0124] MFAQLVAQNVLLIDGPLSWYSDPGLAGVSLTGGLSYKEDTKELVVAKAGVYYVFFQLELRRVVAGEGSGSVSLALHLQPLRSAAGAAALALTVDLPPASSEARNSAFGFQGRLLHLSAGQRLGVHLHTEARARHAWQLTQGATVLGLFRV
[0125] CRT(SEQ ID NO:17)
[0126] MLLSVPLLLGLLGLAVAEPAVYFKEQFLDGDGWTSRWIESKHKSDFGKFVLSSGKFYGDEEKDKGLQTSQDARFYALSASFEPFSNKGQTLVVQFTVKHEQNIDCGGGYVKLFPNSLDQTDMHGDSEYNIMFGPDICGPGTKKVHVIFNYKGKNVLINKDIRCKDDEFTHLYTLIVRPDNTYEVKIDNSQVESGSLEDDWDFLPPKKIKDPDASKPEDWDERAKIDDPTDSKPEDWDKPEHIPDPDAKKPEDWDEEMDGEWEPPVIQNPEYKGEWKPRQIDNPDYKGTWIHPEIDNPEYSPDPSIYAYDNFGVLGLDLWQVKSGTIFDNFLITNDEAYAEEFGNETWGVTKAAEKQMKDKQDEEQRLKEEEEDKKRKEEEEAEDKEDDEDKDEDEEDEEDKEEDEEEDVPGQAKDEL
[0127] EDA(SEQ ID NO:18)
[0128] NIDRPKGLAFTDVDVDSIKIAWESPQGQVSRYRVTYSSPEDGIHELFPAPDGEEDTAELQGLRPGSEYTVSVVALHDDMESQPLIGTQST
[0129] HSP70(SEQ ID NO:19)
[0130] MARAVGIDLGTTNSVVSVLEGGDPVVVANSEGSRTTPSIVAFARNGEVLVGQPAKNQAVTNVDRTVRSVKRHMGSDWSIEIDGKKYTAPEISARILMKLKRDAEAYLGEDITDAVITTPAYFNDAQRQATKDAGQIAGLNVLRIVNEPTAAALAYGLDKGEKEQRILVFDLGGGTFDVSLLEIGEGVVEVRATSGDNHLGGDDWDQRVVDWLVDKFKGTSGIDLTKDKMAMQRLREAAEKAKIELSSSQSTSINLPYITVDADKNPLFLDEQLTRAEFQRITQDLLDRTRKPFQSVIADTGISVSEIDHVVL VGGSTRMPAVTDLVKELTGGKEPNKGVNPDEVVAVGAALQAGVLKGEVKDVLLLDVTPLSLGIETKGVMTRLIERNTTIPTKRSETFTTADDNQPSVQIQVYQGEREIAAHNKLLGSFELTGIPPAPRGIPQIEVTFDIDANGIVHVTAKDKGTGKENTIRIQEGSGLSKEDIDRMIKDAEAHAEEDRKRREEADVRNQAETLVYQTEKFVKEQREAEGGSKVPEDTLNKVDAAVAEKAALGGSDISAIKSAMEKLGQESQALGQAIYEAAQAASQATGAAHPGSADDVVDAEVVDDGREAK
[0131] IL1(SEQ ID NO:20)
[0132] LEADKCKEREEKIILVSSANEIDVRPCPLNPNEHKGTITWYKDDSKTPVSTEQASRIHQHKEKLWFVPAKVEDSGHYYCVVRNSSYCLRIKISAKFVENEPNLCYNAQAIFKQKLPVAGDGGLVCPYMEFFKNENNELPKLQWYKDCKPLLLDNIHFSGVKDRLIVMNVAEKHRGNYTCHASYTYLGKQYPITRVIEFITLEENKPTRPVIVSPANETMEVDLGSQIQLICNVTGQLSDIAYWKWNGSVIDEDDPVLGEDYYSVENPANKRRSTLITVLNISEIESRFYKHPFTCFAKNTHGIDAAYIQLIYPVTNFQK
[0133] IL2(SEQ ID NO:21)
[0134] MYRMQLLSCIALSLALVTNSAPTSSSTKKTQLQLEHLLLDLQMILNGINNYKNPKLTRMLTFKFYMPKKATELKHLQCLEEELKPLEEVLNLAQSKNFHLRPRDLISNINVIVLELKGSETTFMCEYADETATIVEFLNRWITFCQSIISTLT
[0135] IL7(SEQ ID NO:22)
[0136] CDIEGKDGKQYESVLMVSIDQLLDSMKEIGSNCLNNEFNFFKRHICDANKEGMFLFRAARKLRQFLKMNSTGDFDLHLLKVSEGTTILLNCTGQVKGRKPAALGEAQPTKSLEENKSLKEQKKLNDLCFLKRLLQEIKTC
[0137] HSP40(SEQ ID NO:23)
[0138] QDFFNGKELNKSINPDEAVAYGAAVQAAILSGDKSENVQDLLLLDVTPLSLGIETAGG
[0139] VMTVLIKRNTTIPTKQTQTFTTYSDNQPGVLIQVYEGERAMTKDNNLLGKFELTGIPPA
[0140] PRGVPQIEVTFDI
[0141] The fusion protein described in this article may also include a secretion signal peptide. The secretion signal peptide may be an IgG signal peptide, an IL-2 signal peptide, a tPA signal peptide, an Ig kappa signal peptide, or a SigMHC signal peptide. The amino acid sequence of the secretion signal peptide may contain mutations (e.g., deletions, additions, substitutions, or insertions) of some amino acid residues, as long as the mutated secretion signal peptide retains its function of guiding the secretion of the fusion protein. The secretion signal peptide may contain any of the following sequences or have at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the following amino acid sequences.
[0142] IgG signal peptide (SEQ ID NO:5): MDWTWRVFCLLAVTPGAHP
[0143] IL-2 signal peptide (SEQ ID NO:6): MYRMQLLSCIALSLALVTNS
[0144] tPA signal peptide (SEQ ID NO:7): MDAMKRGLCCVLLLCGAVFVSP
[0145] Ig kappa signal peptide (SEQ ID NO:8): METDTLLLWVLLLWVPGSTGD
[0146] SigMHC signal peptide (SEQ ID NO:9): MRVTAPRTVLLLLSAALALTETWA nucleic acid
[0147] This article provides an isolated nucleic acid, which is a DNA or RNA molecule, and may be double-stranded, single-stranded, or partially double-stranded. The isolated nucleic acid encodes the fusion protein described herein. The isolated nucleic acid may contain elements that regulate the expression of the fusion protein, such as enhancer, promoter, and / or terminator sequences. These sequences may be modified or unmodified. Elements regulating the expression of the fusion protein may also be absent from the isolated nucleic acid.
[0148] The nucleic acids of particular interest in this invention are mRNA molecules, such as those containing a read frame encoding an immunogenic protein.
[0149] mRNA molecules
[0150] Messenger RNA (mRNA) is any RNA that encodes a protein and can be translated to produce the protein-encoded RNA in vitro, in vivo, in situ, or ex vivo. Unless otherwise stated, the nucleic acid sequences in this application may be described as “T” in a representative DNA sequence, but in the case of sequences representing RNA (e.g., mRNA), “T” will be replaced by “U”. Therefore, any DNA indicated herein by sequence number also discloses a corresponding RNA (e.g., mRNA) sequence complementary to the DNA, wherein each “T” in the DNA sequence is replaced by a “U”.
[0151] mRNA molecules can be synthetic and modified. mRNA can be chemically modified. mRNA molecules can be chemically synthesized or transcribed in vitro. mRNA molecules can be placed on a vector. The vector can be a viral vector, a bacterial vector, or a eukaryotic expression vector. In one embodiment, the vector is a plasmid. In some instances, mRNA molecules can be delivered to cells via transfection, electroporation, or transduction (e.g., adenovirus or lentivirus transduction).
[0152] Chemical modification
[0153] In some embodiments, the nucleic acid (e.g., mRNA) comprises RNA having a read frame encoding an immunogenic protein, wherein the nucleic acid comprises nucleotides and / or nucleosides that may be standard (unmodified) or modified as known in the art. In some embodiments, the nucleotides and nucleosides of the nucleic acid (e.g., mRNA) comprise modified nucleotides or nucleosides. These modified nucleotides and nucleosides may be naturally occurring modified nucleotides and nucleosides or non-naturally occurring modified nucleotides and nucleosides. These modifications may include modifications at the sugar, backbone, or nucleobase portions of the nucleotides and / or nucleosides as recognized in the art.
[0154] In some implementations, the nucleic acid (e.g., mRNA) may comprise standard nucleotides and nucleosides, naturally occurring nucleotides and nucleosides, non-naturally occurring nucleotides and nucleosides, or any combination thereof.
[0155] In some embodiments, nucleic acids (e.g., DNA nucleic acids and RNA nucleic acids, such as mRNA nucleic acids) comprise different types of standard and / or modified nucleotides and nucleosides. In some embodiments, specific regions of the nucleic acid contain one, two, or more (optionally different) types of standard and / or modified nucleotides and nucleosides.
[0156] In some implementations, modified RNA nucleic acids (e.g., modified mRNA nucleic acids) introduced into cells or organisms exhibit reduced degradation in cells or organisms, relative to unmodified nucleic acids containing standard nucleotides and nucleosides.
[0157] In some implementations, modified RNA nucleic acids (e.g., modified mRNA nucleic acids) introduced into cells or organisms may exhibit reduced immunogenicity (e.g., reduced innate response) in cells or organisms, relative to unmodified nucleic acids containing standard nucleotides and nucleosides.
[0158] In some implementations, nucleic acids (e.g., mRNA) comprise non-natural modified nucleotides introduced during or after nucleic acid synthesis to achieve a desired function or property. Modifications can be present at nucleotide linkages, purine or pyrimidine bases, or sugars. Modifications can be introduced chemically or at any other location at the end of the chain or in the chain using polymerases. Any region of the nucleic acid can be chemically modified.
[0159] Nucleic acids (e.g., mRNA) can contain modified nucleosides and nucleotides. A “nucleoside” is a compound containing a sugar molecule (e.g., pentose or ribose) or a derivative thereof combined with an organic base (e.g., a purine or pyrimidine) or a derivative thereof (also referred to herein as a “nucleobase”). A “nucleotide” refers to a nucleoside, including a phosphate ester group. Modified nucleotides can be synthesized by any useful method, such as chemical, enzymatic, or recombinant methods, to include one or more modified or non-natural nucleosides. Nucleic acids can contain one or more linked nucleoside regions. These regions can have variable backbone linkages. The linkage can be a standard phosphodiester linkage, in which case the nucleic acid will contain the nucleotide region.
[0160] Modified nucleotide base pairings encompass not only standard adenosylthymine, adenosyluracil, or guanosine cytosine base pairs, but also base pairs formed between nucleotides and / or modified nucleotides containing non-standard or modified bases. In nucleic acids, for example, those with at least one chemical modification, the arrangement of hydrogen bond donors and acceptors allows hydrogen bonding to occur between non-standard and standard bases or between two complementary non-standard base structures. An example of such non-standard base pairing is the base pairing between the modified nucleotide inosine and adenine, cytosine, or uracil.
[0161] In some embodiments, the nucleic acid (e.g., mRNA) contains uridine at one or more or all uridine sites. In some embodiments, the mRNA is uniformly modified (e.g., completely modified, modified throughout the entire sequence) for a specific modification. In some embodiments, the nucleic acid can be uniformly modified with methylpseuuridine, meaning that all uridine residues in the mRNA sequence are replaced with 1-methylpseuuridine. Similarly, the nucleic acid can be uniformly modified for any type of nucleoside residue present in the sequence by replacing it with modified residues (e.g., the modified residues described above).
[0162] In some embodiments, the RNA is modified RNA, wherein the uracil, cytosine, adenine, or guanine nucleotide contains a modifying group. The modifying group may be selected from at least one of pseudouridine, N1-methylpseudouridine, N1-ethylpseudouridine, 5-methylcytosine, 5-methoxycytosine, N1-methylcytosine, 2-thiouridine, 5-methoxyuridine, or N1-methyladenosine, N1-methylguanine, N1-methylguanine, and isoguanine.
[0163] In some embodiments, the mRNA molecule is a modified mRNA, the modification including the conversion of uracil nucleoside to pseudouridine, N1-methylpseudouridine, N1-ethylpseudouridine, 2-thiouridine, 5-methoxyuridine; and / or, the conversion of cytosine nucleoside to 5-methylcytosine, 5-methoxycytosine, N1-methylcytosine; and / or, the conversion of adenine nucleoside to N1-methyladenosine; and / or, the conversion of adenine nucleoside to N1-methylguanine, N1-methylguanine, isoguanine.
[0164] Typical mRNA sequence
[0165] An mRNA molecule comprising one or more of the nucleic acid sequences of SEQ ID NO:57-89, 108, and 109 is provided. Those skilled in the art will understand that this invention covers variant sequences of SEQ ID NO:57-89, 108, and 109. For example, the reading frame encoding the immunogenic protein in the mRNA molecule can vary. For example, the reading frame of the immunogenic protein can vary considerably based on codon degeneracy. Similarly, those skilled in the art can make appropriate modifications to other elements while retaining their original function.
[0166] SEQ ID NO:57
[0167] ATGAGGGTGACAGCGCCGCGGACAGTGCTGCTGCTGCTGTCCGCGGCGCTGGCCCTCACCGAG
[0168] ACaTGGGCCCACGGGGACACGCCGACCCTGCATGAGTACATGCTGGACTTGCAGCCTGAGACG
[0169] ACCGACCTTTACTGCTACGAGCAGCTCAACGACTCGTCCGAGGAGGAGGACGAGATCGATGGG
[0170] CCCGCCGGCCAGGCCGAGCCCGACCGCGCCCATTACAACATCGTGACGTTCTGTTGCAAGTGCG
[0171] ACAGCACGTTACGACTGTGCGTGCAGAGCACGCACGTGGACATACGAACCCTGGAGGACCTGC
[0172] TGATGGGGACTCTGGGGATCGTGTGTCCAATCTGCAGCCAGAAGCCAAGCCAGTCCACGATCC
[0173] CCATCGTGGGGATCGTGGCTGGCTTGGCTGTTCTGGCTGTTGTTGTGATCGGCGCGGTCGTGGC
[0174] TACGGTCATGTGCCGGCGGAAGtcCTCTGGCGGTAAAGGTGGGTCGTACTCTCAGGCTGCAAGT
[0175] TCAGACAGTGCTCAGGGGTCGGACGTGTCCCTGACGGCCTGA
[0176] SEQ ID NO:58
[0177] ATGGACGCCATGAAGCGGGGGCTGTGCTGCGTGCTCCTGCTTTGTGGCGCCGTGTTTGTGAGCC
[0178] CGCATGGGGACACTCCGACCCTGCATGAGTACATGCTGGACCTCCAGCCAGAGACtACCGACCT
[0179] CTATTGCTATGAGCAGCTGAACGACTCGTCCGAGGAGGAGGATGAGATAGATGGGCCCGCCGG
[0180] TCAGGCCGAGCCTGACCGGGCCCACTATAACATCGTCACCTTCTGCTGCAAGTGCGACAGCACC
[0181] CTCCGGCTGTGCGTGCAGTCCACTCATGTGGACATACGCACGCTGGAGGACCTGCTGATGGGCA
[0182] CACTGGGCATAGTGTGCCCCATTTGCAGCCAGAAGCCCTTGGACGAGAACAGCATGCTCATAC
[0183] CAATAGCGGTCGGTGGTGCTCTGGCTGGACTGGTCCTGATTGTGCTCATCGCATACCTGGTGGG
[0184] TCGGAAGCGTTCCCATGCGGGCTATCAGACCATCTGA
[0185] SEQ ID NO:59
[0186] ATGCGCGTCACGGCCCCCCGGACCGTCCTCCTCCTCCTGTCCGCGGCTCTGGCGCTGACCGAGA
[0187] CATGGGCTCATGGAGACACCCCCACTCTCCATGAGTACATGCTCGATCTGCAGCCAGAGACgAC
[0188] GGACCTGTACTGTTACGAGCAGCTGAACGACTCCTCCGAGGAGGAGGACGAGATTGATGGGCC
[0189] CGCCGGTCAGGCCGAGCCTGACCGGGCCCATTATAACATCGTCACGTTCTGCTGCAAGTGTGAC
[0190] AGTACCCTCCGCCTTTGCGTCCAGTCCACCCACGTCGACATCCGGACCCTCGAGGACCTCaTTAT
[0191] GGGCACCCTCGGCATCGTCTGCCCCATCTGCAGCCAGAAGCCAAGCCAGTCCACGATCCCCATC
[0192] GTGGGGATCGTGGCTGGCTTGGCTGTTCTGGCTGTGGTGGTCATCGGGGCCGTCGTTGCCACTG
[0193] TGATGTGTCGGAGGAAGTCGTCTGGGGGTAAGGGCGGCTCTTACTCCCAGGCGGCTTCCTCCGA
[0194] CAGCGCACAGGGCAGCGACGTGTCCCTGACCGCCGGGTCCGGGGCCACGAACTTCAGCCTGCT
[0195] GAAGCAGGCTGGAGACGTGGAGGAGAACCCCGGACCCATGCTTCTGTCCGTCCCCCTCCTCCTT
[0196] GGGCTGCTGGGCCTCGCAGTGGCTGAGCCAGCCGTCTACTTCAAGGAGCAGTTCCTTGATGGGG
[0197] ACGGCTGGACCAGCCGCTGGATCGAGTCCAAGCATAAGAGTGACTTCGGAAAGTTCGTGCTGT
[0198] CCTCTGGCAAGTTCTATGGGGACGAGGAGAAGGACAAGGGCCTCCAGACCAGTCAGGATGCTC
[0199] GGTTTTACGCACTGTCCGCCAGCTTCGAGCCCTTTAGCAATAAAGGGCAGACGCTGGTGGTGCA
[0200] GTTCACCGTGAAGCACGAGCAGAACATTGACTGCGGTGGAGGCTATGTCAAGCTCTTtCCCAAT
[0201] AGTCTTGACCAGACGGACATGCACGGAGATTCCGAGTACAACATTATGTTTGGACCCGATATCT
[0202] GCGGCCCTGGTACCAAGAAAGTCCATGTCATCTTCAACTACAAGGGCAAGAATGTCCTGATCA
[0203] ACAAGGACATTCGTTGCAAGGACGATGAGTTTACACACCTGTACACACTCATCGTCCGGCCGG
[0204] ACAACACCTACGAGGTCAAGATTGACAACTCGCAGGTGGAGTCCGGCTCCTTGGAAGATGACT
[0205] GGGACTTTCTGCCCCCCAAGAAGATCAAGGACCCGGACGCCTCCAAGCCTGAGGATTGGGATG
[0206] AGAGAGCGAAGATCGATGACCCGACTGACTCTAAGCCAGAGGACTGGGATAAGCCCGAGCAT
[0207] ATCCCAGATCCAGACGCCAAGAAGCCCGAAGATTGGGATGAGGAGATGGACGGGGAGTGGGA
[0208] GCCTCCGGTGATCCAGAACCCCGAGTATAAGGGGGAGTGGAAGCCTCGGCAGATCGACAACCC
[0209] CGATTACAAGGGGACCTGGATCCACCCGGAGATCGACAACCCCGAGTACTCCCCCCGACCCGTC
[0210] CATCTACGCTTATGACAACTTCGGGGTTCTTGGCCTGGACCTCTGGCAGGTCAAGTCGGGCACG
[0211] ATCTTCGACAACTTTCTCATCACCAATGATGAGGCTTACGCGGAGGAGTTCGGGAACGAGACTT
[0212] GGGGGGTTACCAAGGCCGCAGAGAAGCAGATGAAGGATAAGCAGGACGAGGAACAGCGGCTC
[0213] AAGGAGGAGGAGGAGGACAAGAAGCGGAAGGAGGAGGAGGGAGGCCGAGGATAAGGAGGACG
[0214] ACGAGGACAAGGATGAGGACGAGGAGGACGAGGAGGATAAGGAGGAGGACGAGGGAGGA
[0215] CGTTCCGGGGCAGGCCAAGGACGAGCTGTGA
[0216] SEQ ID NO:60
[0217] ATGCGCGTCACGGCCCCCCGGACCGTCCTCCTCCTCCTGTCCGCCGCCCTGGCTCTGACTGAGA
[0218] CTTGGGCACACGGTGACACCCCAACGCTCCATGAGTACATGCTGGACCTCCAGCCCGAAACCA
[0219] CGGACTTGTACTGCTACGAGCAGCTGAACGACAGCAGCGAGGAGGAGGACGAGATTGATGGG
[0220] CCCGCCGGTCAGGCCGAGCCTGACCGGGCCCATTATAACATCGTCACCTTCTGCTGCAAGTGCG
[0221] ACAGCACTCTTCGGCTGTGCGTGCAGTCAACTCACGTGGACATTCGGACGCTGGAGGACCTGCT
[0222] CATGGGGACGTTGGGGATCGTGTGCCCAATCTGCAGTCAGAAGCCGGGCGGCGGCAGTATGCA
[0223] CCAGAAGCGTACTGCCATGTTCCAGGACCCCCAGGAGCGGCCGCGGAAGCTGCCGCAGCTGTG
[0224] CACGGAGCTCCAGACTACTATCCACGATATTATTCTGGAGTGCGTGTACTGCAAGCAGCAGCTT
[0225] CTGCGGCGCGAAGTCTACGACTTCGCCTTCAGGGACTTGTGCATAGTGTATAGGGATGGCAATC
[0226] CCTATGCAGTGTGCGACAAGTGCCTGAAGTTCTACTCCAAGATCTCCGAGTACAGGCACTACTG
[0227] CTACAGCCTGTACGGGACGACCTTGGAGCAGCAGTACAACAAGCCTCTTTGTGACCTCCTCATC
[0228] CGCTGCATCAATTGCCAGAAGCCGCTGTGTCCCGAGGAGAAGCAGCGGCATCTGGACAAGAAG
[0229] CAGCGTTTCCACAACATCCGCGGCCGCTGGACAGGGAGATGTATGTCCTGCTGCCGGTCGTCAA
[0230] GGACACGTCGGGAGACgCAGCTGAGTCAGAGCACGATCCCCATCGTGGGGATCGTGGCCGGCC
[0231] TGGCTGTGCTGGCCGTCGTTGTCATCGGAGCTGTTGTGGCGACAGTGATGTGTCGCCGCAAGAG
[0232] CTCCGGTGGCAAGGGCGGCTCGTACAGCCAGGCCGCCAGCTCTGACTCAGCTCAGGGCTCCGA
[0233] CGTGTCCTTGACGGCCGGCAGCGGGGCTACAAACTTCTCCCTGTTGAAGCAGGCCGGGGATGTT
[0234] GAGGAGAATCCTGGTCCTATGCTCCTCAGCGTCCCCTTGCTCCTCGGCCTCCTCGGACTCGCAG
[0235] TGGCTGAGCCAGCCGTCTACTTCAAGGAGCAGTTCCTTGATGGGGACGGCTGGACCAGCCGCT
[0236] GGATCGAGTCCAAGCATAAGAGCGATTTTGGGAAGTTTGTGCTCTCGTCCGGCAAGTTCTACGG
[0237] CGACGAGGAGAAGGACAAGGGCCTGCAGACTTCCCAGGACGCTCGCTTTTATGCTCTTTCCGCT
[0238] TCTTTCGAGCCCTTCTCCAATAAGGGGCAGACCCTCGTCGTGCAGTTCACCGTCAAGCATGAGC
[0239] AGAACATCGATTGCGGGGGCGGCTACGTGAAGCTGTTCCCCAATTCCCTGGACCAGACAGATA
[0240] TGCACGGGGACTCCGAATATAACATTATGTTCGGCCCCGATATCTGTGGTCCAGGGACCAAGA
[0241] AAGTCCATGTCATCTTCAACTACAAGGGCAAGAATGTCCTGATCAACAAGGACATTCGTTGCAA
[0242] GGACGATGAGTTTACACACCTGTACACACTCATCGTCCGGCCGGACAACACCTACGAGGTCAA
[0243] GATTGACAACTCGCAGGTGGAGTCCGGCTCCTTGGAAGATGACTGGGACTTTCTGCCCCCCAAG
[0244] AAGATCAAGGACCCGGACGCCTCCAAGCCTGAGGATTGGGATGAGAGAGCGAAGATCGATGA
[0245] CCCGACTGACTCTAAGCCAGAGGACTGGGATAAGCCCGAGCATATCCCAGATCCAGACGCCAA
[0246] GAAGCCCGAAGATTGGGATGAGGAGATGGACGGGGAGTGGGAGCCTCCGGTGATCCAGAACC
[0247] CCGAGTATAAGGGGGAGTGGAAGCCTCGGCAGATCGACAACCCCGATTACAAGGGGACCTGG
[0248] ATCCACCCGGAGATCGACAACCCCGAGTACTCCCCCGACCCGTCCATCTACGCTTATGACAACT
[0249] TCGGGGTTCTTGGCCTGGACCTCTGGCAGGTCAAGTCGGGCACGATCTTCGACAACTTTCTCAT
[0250] CACCAATGATGAGGCTTACGCGGAGGAGTTCGGGAACGAGACTTGGGGGGTGACGAAGGCTGC
[0251] TGAGAAGCAGATGAAGGACAAGCAGGACGAGGAGCAGCGCCTTAAGGAGGAGGAGGAGGAC
[0252] AAGAAGCGGAAAGAGGAGGAGGAGGCCGAGGACAAGGAGGACGATGAGGACAAGGACGAGG
[0253] ATGAGGAGGATGAGGAGGACAAAGAGGAGGACGAGGAGGAGGACGTTCCGGGGCAGGCCAA
[0254] GGACGAGCTGTGA
[0255] SEQ ID NO:61
[0256] ATGGACGCCATGAAGCGGGGTCTGTGCTGCGTCCTGCTGCTGTGTGGTGCAGTGTTCGTGTCCC
[0257] CTCATGGGGACACGCCAACACTGCACGAGTACATGCTGGACCTGCAGCCCGAGACgACCGACC
[0258] TCTATTGTTATGAACAGTTGAATGACAGCAGCGAGGAGGAGGACGAGATTGATGGGCCCGCCG
[0259] GTCAGGCCGAGCCTGACCGGGCCCATTATAACATCGTCACCTTCTGCTGCAAGTGTGATTCAAC
[0260] CCTCCGGCTGTGCGTGCAGTCCACTCATGTGGACATACGCACGCTGGAGGACCTGCTCATGGGG
[0261] ACCCTGGGCATCGTGTGTCCGATCTGCTCCCAGAAGCCTGGGGGCGGATCCATGCACCAGAAG
[0262] CGGACCGCGATGTTCCAGGATCCCCAGGAGCGGCCGCGGAAGCTGCCGCAGCTGTGCACGGAG
[0263] CTCCAGACTACTATCCACGATATTATTCTGGAGTGCGTGTACTGCAAGCAGCAGCTTCTGCGGC
[0264] GCGAAGTCTACGACTTCGCCTTCAGGGACTTGTGCATAGTGTATAGGGATGGCAATCCCTATGC
[0265] AGTGTGCGACAAGTGCCTGAAGTTCTACAGCAAGATCTCAGAGTACCGGCATTATTGCTACTCG
[0266] CTGTACGGGACCACCCTGGAGCAGCAGTACAATAAGCCACTCTGCGATCTGCTGATCCGCTGCA
[0267] TCAATTGCCAGAAGCCGCTGTGTCCCGAGGAGAAGCAGCGGCATCTGGACAAGAAGCAGCGGT
[0268] TTCATAACATTAGAGGTCGGTGGACTGGGCGCTGCATGTCCTGCTGTCGCAGCAGCAGGACGC
[0269] GGCGCGAGACgCAGCTCAAGGACGAGCTCAAGGATGAGCTCTGA
[0270] SEQ ID NO:62
[0271] ATGAGGGTGACCGCCCCGCGGACCGTTCTGCTGCTGCTGTCCGCCGCCCTGGCTCTGACTGAGA
[0272] CTTGGGCACACGGTGACACCCCAACGCTCCATGAGTACATGCTGGACCTCCAGCCCGAAACCA
[0273] CGGACTTGTACTGCTACGAGCAGCTGAACGACAGCAGCGAGGAGGAGGACGAGATTGATGGG
[0274] CCCGCCGGTCAGGCCGAGCCTGACCGGGCCCATTATAACATCGTCACCTTCTGCTGCAAGTGCG
[0275] ACAGCACTCTTCGGCTGTGCGTGCAGTCAACTCACGTGGACATTCGGACGCTGGAGGACCTGCT
[0276] CATGGGGACGTTGGGGATCGTGTGCCCAATCTGCAGTCAGAAGCCGGGCGGCGGCAGCATGCA
[0277] CCAGAAGCGGACCGCGATGTTCCAGGACCCCCAGGAGCGGCCGCGGAAGCTGCCGCAGCTGTG
[0278] CACGGAGCTCCAGACTACTATCCACGATATTATTCTGGAGTGCGTGTACTGCAAGCAGCAGCTT
[0279] CTGCGGCGCGAAGTCTACGACTTCGCCTTCAGGGACTTGTGCATAGTGTATAGGGATGGCAATC
[0280] CCTATGCAGTGTGCGACAAGTGCCTGAAGTTCTACTCCAAGATCTCGGAGTATCGCCACTACTG
[0281] CTACTCCTTGTACGGGACGACCCTGGAGCAGCAGTACAACAAGCCCCTCTGCGACCTCCTGATC
[0282] CGCTGCATCAATTGCCAGAAGCCGCTGTGTCCCGAGGAGAAGCAGCGGCATCTGGACAAGAAG
[0283] CAGCGGTTCCACAACATCAGGGGTCGCTGGACGGGGCGTTGTATGAGCTGCTGCAGGTCGTCC
[0284] CGTACAAGGAGGGAGACACAGCTGAGCCAGTCCACGATCCCCATCGTGGGGATCGTGGCTGGC
[0285] TTAGCTGTGTTGGCAGTGGTGGTGATTGGAGCTGTGGTCGCTACTGTGATGTGCAGGCGGAAGT
[0286] CGTCCGGGGGCAAGGGAGGGAGCTACAGCCAGGCAGCTTCATCCGATTCGGCCCAGGGTTCCG
[0287] ACGTCTCCCTGACTGCTGGCAGCGGGGCGACCAACTTCTCCCTGCTGAAGCAGGCAGGGGACG
[0288] TCGAGGAGAACCCTGGGCCGATGTATCGGATGCAGCTGCTGAGCTGTATCGCTCTCTCCCTTGC
[0289] CCTCGTGACGAATTCCGCCCCCACATCCAGTAGCACCAAGAAGACCCAGCTCCAGCTTGAGCA
[0290] CCTCCTCCTGGATCTGCAGATGATCCTGAACGGGATCAACAACTATAAGAATCCGAAGCTGACT
[0291] CGGATGCTTACGTTTAAGTTCTACATGCCTAAGAAGGCAACAGAGCTTAAGCATCTGCAGTGCC
[0292] TGGAGGAGGAGCTCAAGCCTCTGGAGGAGGTGCTGAACCTCGCCCAGAGCAAGAACTTCCACT
[0293] TGCGTCCGCGGGATCTCATCAGCAACATCAACGTGATCGTCTTGGAGCTGAAGGGCTCCGAGA
[0294] CGACCTTCATGTGTGAGTATGCTGATGAGACgGCGACGATAGTGGAGTTCTTGAACCGGTGGAT
[0295] CACCTTCTGTCAGAGTATCATCAGTACTCTGACATGA
[0296] SEQ ID NO:63
[0297] ATGGACGCAATGAAGCGGGGGCTCTGCTGTGTCCTCCTCCTCTGCGGGGCAGTGTTCGTGTCCC
[0298] CTCATGGGGACACGCCAACACTGCACGAGTACATGCTCGACCTGCAGCCCGAGACgACCGACC
[0299] TCTATTGTTATGAACAGTTGAATGACAGCAGCGAGGAGGAGGACGAGATTGATGGGCCCGCCG
[0300] GTCAGGCCGAGCCTGACCGGGCCCATTATAACATCGTCACCTTCTGCTGCAAGTGTGATTCAAC
[0301] CCTCCGGCTGTGCGTGCAGTCCACTCATGTGGACATACGCACGCTGGAGGACCTGCTCATGGGG
[0302] ACCCTGGGCATCGTGTGTCCGATCTGCTCCCAGAAGCCTGGGGGCGGATCCATGCACCAGAAG
[0303] CGGACCGCGATGTTCCAGGATCCCCAGGAGCGGCCGCGGAAGCTGCCGCAGCTGTGCACGGAG
[0304] CTCCAGACTACTATCCACGATATTATTCTGGAGTGCGTGTACTGCAAGCAGCAGCTTCTGCGGC
[0305] GCGAAGTCTACGACTTCGCCTTCAGGGACTTGTGCATAGTGTATAGGGATGGCAATCCCTATGC
[0306] AGTGTGCGACAAGTGCCTGAAGTTCTACAGCAAGATCTCAGAGTACCGGCATTATTGCTACTCG
[0307] CTGTACGGGACCACCCTGGAGCAGCAGTACAATAAGCCACTCTGCGATCTGCTGATCCGCTGCA
[0308] TCAATTGCCAGAAGCCGCTGTGTCCCGAGGAGAAGCAGCGGCATCTGGACAAGAAGCAGCGGT
[0309] TTCATAACATTAGAGGTCGGTGGACTGGGCGCTGCATGTCCTGCTGCAGGAGCAGCAGGACAC
[0310] GGCGGGAAACACAGCTCAAGGATGAGCTCAAGGATGAGCTCGGGTCCGGGGCCACGAACTTCA
[0311] GCCTGCTGAAGCAGGCTGGAGACGTGGAGGAGAACCCCGGACCCAACATTGACCGGCCCAAG
[0312] GGGCTGGCTTTCACCGATGTTGATGTGGACAGCATCAAGATCGCGTGGGAGTCGCCCCAGGGC
[0313] CAGGTCAGCCGTTATCGGGTGACGTACTCGTCACCCGAGGACGGCATCCATGAGCTGTTTCCCG
[0314] CCCCCGACGGGGAGGAGGACACAGCAGAGCTCCAGGGGTTGCGTCCGGGCTCTGAGTACACGG
[0315] TCAGCGTGGTGGCTCTCCATGACGACATGGAGAGCCAGCCGCTGATCGGTACTCAGAGCACCT
[0316] GA
[0317] SEQ ID NO:64
[0318] ATGCGTGTCACTGCCCCCCGGACGGTCTTGCTTCTGCTGTCCGCCGCGTTGGCCTTGACAGAGA
[0319] CCTGGGCACACGGGGACACCCCCACTCTGCACGAGTACATGCTGGACCTTCAGCCGGAGACCA
[0320] CCGATCTCTACTGCTACGAGCAGCTGAACGACAGCAGCGAGGAGGAGGATGAGATCGATGGCC
[0321] CGGCAGGCCAGGCTGAGCCTGATCGCGCACATTACAACATCGTCACCTTCTGCTGCAAGTGCGA
[0322] CTCCACGCTGCGCCTCTGCGTCCAGAGCACCCATGTCGATATCCGCACCCTGGAAGACCTGCTG
[0323] ATGGGCACGCTGGGCATCGTCTGCCCCATCTGCAGCCAGAAGCCAAGCCAGTCCACGATCCCC
[0324] ATCGTGGGGATCGTGGCTGGCTTGGCTGTTCTGGCTGTGGTGGTCATCGGAGCTGTTGTGGCGA
[0325] CAGTGATGTGTCGCCGCAAGAGCTCCGGTGGGAAGGGCGGCTCCTACAGCCAGGCCGCCTCGT
[0326] CCGATTCCGCCCAGGGCTCCGACGTGTCCTTGACCGCTGGGAGTGGCGCCACCAACTTCTCCCT
[0327] CCTCAAGCAGGCTGGCGACGTTGAGGAGAACCCAGGGCCCATGGCGCGGGCCGTGGGCATCGA
[0328] CCTGGGTACTACCAACAGCGTCGTCAGCGTGCTTGAGGGGGGAGATCCCGTGGTGGTCGCCAA
[0329] CTCCGAGGGGTCAAGGACAACTCCGTCCATCGTGGCGTTCGCCAGGAACGGGGAGGTTCTTGT
[0330] CGGCCAGCCTGCCAAGAACCAGGCTGTGACCAATGTGGACCGCACTGTGCGCTCTGTGAAGAG
[0331] GCACATGGGGTCCGATTGGTCAATCGAGATCGACGGTAAGAAGTATACGGCCCCCGAGATAAG
[0332] CGCCCGCATCCTGATGAAGCTCAAGCGGGATGCGGAGGCTTATCTCGGGGAGGACATCACTGA
[0333] CGCAGTCATCACCACCCCGGCTTATTTCAACGATGCGCAGCGCCAGGCCACGAAGGACGCCGG
[0334] TCAGATCGCTGGCCTGAACGTGCTGCGCATCGTGAATGAGCCGACCGCTGCTGCTCTCGCCTAT
[0335] GGCCTGGACAAGGGCGAGAAGGAGCAGCGGATCCTGGTGTTCGACCTGGGTGGCGGCACCTTC
[0336] GACGTCTCCCTCCTCGAGATCGGGGAGGGAGTCGTCGAGGTGCGCGCCACCTCAGGCGATAAC
[0337] CACCTGGGTGGTGATGACTGGGATCAGCGGGTGGTGGACTGGCTGGTGGACAAGTTCAAGGGC
[0338] ACTTCTGGGATCGACCTCACGAAGGACAAGATGGCCATGCAGCGCCTTCGTGAGGCGGCCGAG
[0339] AAGGCCAAGATTGAGCTGTCCAGCAGCCAGTCCACCAGCATCAACCTGCCGTATATTACCGTCG
[0340] ATGCCGACAAGAACCCCCTGTTCCTGGACGAGCAGCTCACGAGGGCGGAGTTCCAGCGGATCA
[0341] CCCAGGATCTGCTGGATCGCACCCGTAAGCCCTTCCAGTCCGTCATCGCCGATACGGGTATTTC
[0342] CGTCAGCGAGATCGACCACGTCGTGCTCGTTGGCGGAAGTACCCGTATGCCGGCGGTGACGGA
[0343] CCTGGTGAAGGAGCTTACGGGTGGGAAGGAGCCCAACAAGGGCGTGAATCCGGACGAGGTGG
[0344] TGGCTGTGGGAGCCGCCCTTCAGGCCGGAGTGCTTAAGGGGGAGGTGAAGGATGTTCTGCTCC
[0345] TGGACGTCACCCCCCTCAGTCTCGGCATCGAGACTAAGGGGGGGGTGATGACCAGGCTGATCG
[0346] AGCGGAACACCACCATCCCCACTAAGCGCTCCGAGACCTTCACCACGGCAGACGATAACCAGC
[0347] CCAGCGTGCAGATCCAGGTCTACCAGGGTGAGAGAGAGATCGCTGCCCACAACAAGCTCCTGG
[0348] GGTCCTTCGAGCTGACGGGGATCCCCCCGGCCCCCCGGGGGATCCCCCAGATCGAAGTGACCTT
[0349] CGACATCGATGCGAATGGTATTGTCCATGTCACTGCCAAGGACAAGGGCACTGGCAAGGAGAA
[0350] TACCATTCGCATCCAGGAGGGCAGCGGTCTCTCTAAGGAGGATATCGACAGGATGATCAAGGA
[0351] CGCAGAGGCGCACGCGGAGGAGGATCGCAAGCGGCGGGAGGAGGCGGATGTGCGCAATCAGG
[0352] CTGAGACCCTGGTCTACCAGACTGAGAAGTTTGTGAAGGAACAGCGTGAGGCAGAGGGGGGGT
[0353] CCAAGGTGCCCGAGGACACTCTGAACAAGGTCGACGCGGCGGTAGCAGAAGCCAAGGCCGCTC
[0354] TGGGGGGCAGTGACATATCCGCCATCAAGTCCGCCATGGAGAAGCTGGGCCAGGAGAGCCAGG
[0355] CCCTCGGGCAGGCCATCTATGAGGCCGCCCAGGCGGCCTCCCAGGCCACCGGGGCGGCCCATC
[0356] CAGGTGGCGAGCCCGGGGGCGCTCATCCTGGCTCAGCTGACGACGTGGTGGACGCCGAGGTGG
[0357] TGGATGATGGCCGGGAGGCCAAGTGA
[0358] SEQ ID NO:65
[0359] ATGCGCGTGACCGCGCCGCGCACCGTGCTGCTGCTGCTGAGCGCGGCGCTGGCGCTGACCGAA
[0360] ACCTGGGCGATGCATGGCCCGAAAGCGACCCTGCAGGATATTGTGCTGCATCTGGAACCGCAG
[0361] AACGAAATTCCGGTGGATCTGCTGTGCCATGAACAGCTGAGCGATAGCGAAGAAGAAAACGAT
[0362] GAAATTGATGGCGTGAACCATCAGCATCTGCCGGCGCGCCGCGCGGAACCGCAGCGCCATACC
[0363] ATGCTGTGCATGTGCTGCAAATGCGAAGCGCGCATTAAACTGGTGGTGGAAAGCAGCGCGGAT
[0364] GATCTGCGCGCGTTTCAGCAGCTGTTTCTGAACACCCTGAGCTTTGTGTGCCCGTGGTGCGCGA
[0365] GCCAGCAGAGCCAGAGCACCATTCCGATTGTGGGCATTGTGGCGGGCCTGGCGGTGCTGGCGG
[0366] TGGTGGTGATTGGCGCGGTGGTGGCGACCGTGATGTGCCGCCGCAAAAGCAGCGGCGGCAAAG
[0367] GCGGCAGCTATAGCCAGGCGGCGAGCAGCGATAGCGCGCAGGGCAGCGATGTGAGCCTGACC
[0368] GCGTGA
[0369] SEQ ID NO:66
[0370] ATGGATGCGATGAAACGCGGCCTGTGCTGCGTGCTGCTGCTGTGCGGCGCGGTGTTTGTGAGCC
[0371] CGATGCATGGCCCGAAAGCGACCCTGCAGGATATTGTGCTGCATCTGGAACCGCAGAACGAAA
[0372] TTCCGGTGGATCTGCTGTGCCATGAACAGCTGAGCGATAGCGAAGAAGAAAACGATGAAATTG
[0373] ATGGCGTGAACCATCAGCATCTGCCGGCGCGCCGCGCGGAACCGCAGCGCCATACCATGCTGT
[0374] GCATGTGCTGCAAATGCGAAGCGCGCATTAAACTGGTGGTGGAAAGCAGCGCGGATGATCTGC
[0375] GCGCGTTTCAGCAGCTGTTTCTGAACACCCTGAGCTTTGTGTGCCCGTGGTGCGCGAGCCAGCA
[0376] GCTGGATGAAAACAGCATGCTGATTCCGATTGCGGTGGGCGGCGCGCTGGCGGGCCTGGTGCT
[0377] GATTGTGCTGATTGCGTATCTGGTGGGCCGCAAACGCAGCCATGCGGGCTATCAGACCATTTGASEQID NO:67
[0378] ATGCGCGTGACCGCGCCGCGCACCGTGCTGCTGCTGCTGAGCGCGGCGCTGGCGCTGACCGAA
[0379] ACCTGGGCGATGCATGGCCCGAAAGCGACCCTGCAGGATATTGTGCTGCATCTGGAACCGCAG
[0380] AACGAAATTCCGGTGGATCTGCTGTGCCATGAACAGCTGAGCGATAGCGAAGAAGAAAACGAT
[0381] GAAATTGATGGCGTGAACCATCAGCATCTGCCGGCGCGCCGCGCGGAACCGCAGCGCCATACC
[0382] ATGCTGTGCATGTGCTGCAAATGCGAAGCGCGCATTAAACTGGTGGTGGAAAGCAGCGCGGAT
[0383] GATCTGCGCGCGTTTCAGCAGCTGTTTCTGAACACCCTGAGCTTTGTGTGCCCGTGGTGCGCGA
[0384] GCCAGCAGAGCCAGAGCACCATTCCGATTGTGGGCATTGTGGCGGGCCTGGCGGTGCTGGCGG
[0385] TGGTGGTGATTGGCGCGGTGGTGGCGACCGTGATGTGCCGCCGCAAAAGCAGCGGCGGCAAAG
[0386] GCGGCAGCTATAGCCAGGCGGCGAGCAGCGATAGCGCGCAGGGCAGCGATGTGAGCCTGACC
[0387] GCGGGCAGCGGCGCGACCAACTTTAGCCTGCTGAAACAGGCGGGCGATGTGGAAGAAAACCC
[0388] GGGCCCGATGCTGCTGAGCGTGCCGCTGCTGCTGGGCCTGCTGGGCCTGGCGGTGGCGGAACC
[0389] GGCGGTGTATTTTAAAGAACAGTTTCTGGATGGCGATGGCTGGACCAGCCGCTGGATTGAAAG
[0390] CAAACATAAAAGCGATTTTGGCAAATTTGTGCTGAGCAGCGGCAAATTTTATGGCGATGAAGA
[0391] AAAAGATAAAGGCCTGCAGACCAGCCAGGATGCGCGCTTTTATGCGCTGAGCGCGAGCTTTGA
[0392] ACCGTTTAGCAACAAAGGCCAGACCCTGGTGGTGCAGTTTACCGTGAAACATGAACAGAACAT
[0393] TGATTGCGGCGGCGGCTATGTGAAACTGTTTCCGAACAGCCTGGATCAGACCGATATGCATGGC
[0394] GATAGCGAATATAACATTATGTTTGGCCCGGATATTTGCGGCCCGGGCACCAAAAAAGTGCAT
[0395] GTGATTTTTAACTATAAAGGCAAAAACGTGCTGATTAACAAAGATATTCGCTGCAAAGATGAT
[0396] GAATTTACCCATCTGTATACCCTGATTGTGCGCCCGGATAACACCTATGAAGTGAAAATTGATA
[0397] ACAGCCAGGTGGAAAGCGGCAGCCTGGAAGATGATTGGGATTTTCTGCCGCCGAAAAAAATTA
[0398] AAGATCCGGATGCGAGCAAACCGGAAGATTGGGATGAACGCGCGAAAATTGATGATCCGACC
[0399] GATAGCAAACCGGAAGATTGGGATAAACCGGAACATATTCCGGATCCGGATGCGAAAAAACC
[0400] GGAAGATTGGGATGAAGAAATGGATGGCGAATGGGAACCGCCGGTGATTCAGAACCCGGAAT
[0401] ATAAAGGCGAATGGAAACCGCGCCAGATTGATAACCCGGATTATAAAGGCACCTGGATTCATC
[0402] CGGAAATTGATAACCCGGAATATAGCCCGGATCCGAGCATTTATGCGTATGATAACTTTGGCGT
[0403] GCTGGGCCTGGATCTGTGGCAGGTGAAAAGCGGCACCATTTTTGATAACTTTCTGATTACCAAC
[0404] GATGAAGCGTATGCGGAAGAATTTGGCAACGAAACCTGGGGCGTGACCAAAGCGGCGGAAAA
[0405] ACAGATGAAAGATAAACAGGATGAAGAACAGCGCCTGAAAGAAGAAGAAGAAGATAAAAAA
[0406] CGCAAAGAAGAAGAAGAAGCGGAAGATAAAGAAGATGATGAAGATAAAGATGAAGATGAAG
[0407] AAGATGAAGAAGATAAAGAAGAAGATGAAGAAGAAGATGTGCCGGGCCAGGCGAAAGATGA
[0408] ACTGTGA
[0409] SEQ ID NO:68
[0410] ATGCGCGTGACCGCGCCGCGCACCGTGCTGCTGCTGCTGAGCGCGGCGCTGGCGCTGACCGAA
[0411] ACCTGGGCGATGCATGGCCCGAAAGCGACCCTGCAGGATATTGTGCTGCATCTGGAACCGCAG
[0412] AACGAAATTCCGGTGGATCTGCTGTGCCATGAACAGCTGAGCGATAGCGAAGAAGAAAACGAT
[0413] GAAATTGATGGCGTGAACCATCAGCATCTGCCGGCGCGCCGCGCGGAACCGCAGCGCCATACC
[0414] ATGCTGTGCATGTGCTGCAAATGCGAAGCGCGCATTAAACTGGTGGTGGAAAGCAGCGCGGAT
[0415] GATCTGCGCGCGTTTCAGCAGCTGTTTCTGAACACCCTGAGCTTTGTGTGCCCGTGGTGCGCGA
[0416] GCCAGCAGGGCGGCGGCAGCATGGCGCGCTTTGAAGATCCGACCCGCCGCCCGTATAAACTGC
[0417] CGGATCTGTGCACCGAACTGAACACCAGCCTGCAGGATATTGAAATTACCTGCGTGTATTGCAA
[0418] AACCGTGCTGGAACTGACCGAAGTGTTTGAATTTGCGTTTAAAGATCTGTTTGTGGTGTATCGC
[0419] GATAGCATTCCGCATGCGGCGTGCCATAAATGCATTGATTTTTATAGCCGCATTCGCGAACTGC
[0420] GCCATTATAGCGATAGCGTGTATGGCGATACCCTGGAAAAACTGACCAACACCGGCCTGTATA
[0421] ACCTGCTGATTCGCTGCCTGCGCTGCCAGAAACCGCTGAACCCGGCGGAAAAACTGCGCCATCT
[0422] GAACGAAAAACGCCGCTTTCATAACATTGCGGGCCATTATCGCGGCCAGTGCCATAGCTGCTGC
[0423] AACCGCGCGCGCCAGGAACGCCTGCAGCGCCGCCGCGAAACCCAGGTGAGCCAGAGCACCATT
[0424] CCGATTGTGGGCATTGTGGCGGGCCTGGCGGTGCTGGCGGTGGTGGTGATTGGCGCGGTGGTG
[0425] GCGACCGTGATGTGCCGCCGCAAAAGCAGCGGCGGCAAAGGCGGCAGCTATAGCCAGGCGGC
[0426] GAGCAGCGATAGCGCGCAGGGCAGCGATGTGAGCCTGACCGCGGGCAGCGGCGCGACCAACT
[0427] TTAGCCTGCTGAAACAGGCGGGCGATGTGGAAGAAAACCCGGGCCCGATGCTGCTGAGCGTGC
[0428] CGCTGCTGCTGGGCCTGCTGGGCCTGGCGGTGGCGGAACCGGCGGTGTATTTTAAAGAACAGTT
[0429] TCTGGATGGCGATGGCTGGACCAGCCGCTGGATTGAAAGCAAACATAAAAGCGATTTTGGCAA
[0430] ATTTGTGCTGAGCAGCGGCAAATTTTATGGCGATGAAGAAAAAGATAAAGGCCTGCAGACCAG
[0431] CCAGGATGCGCGCTTTTATGCGCTGAGCGCGAGCTTTGAACCGTTTAGCAACAAAGGCCAGAC
[0432] CCTGGTGGTGCAGTTTACCGTGAAACATGAACAGAACATTGATTGCGGCGGCGGCTATGTGAA
[0433] ACTGTTTCCGAACAGCCTGGATCAGACCGATATGCATGGCGATAGCGAATATAACATTATGTTT
[0434] GGCCCGGATATTTGCGGCCCGGGCACCAAAAAAAGTGCATGTGATTTTTAACTATAAAGGCAAA
[0435] AACGTGCTGATTAACAAAGATATTCGCTGCAAAGATGATGAATTTACCCATCTGTATACCCTGA
[0436] TTGTGCGCCCGGATAACACCTATGAAGTGAAAATTGATAACAGCCAGGTGGAAAGCGGCAGCC
[0437] TGGAAGATGATTGGGATTTTCTGCCGCCGAAAAAAAATTAAAGATCCGGATGCGAGCAAACCGG
[0438] AAGATTGGGATGAACGCGCGAAAATTGATGATCCGACCGATAGCAAACCGGAAGATTGGGATA
[0439] AACCGGAACATATTCCGGATCCGGATGCGAAAAAACCGGAAGATTGGGATGAAGAAATGGAT
[0440] GGCGAATGGGAACCGCCGGTGATTCAGAACCCGGAATATAAAGGCGAATGGAACCGCGCCA
[0441] GATTGATAACCCGGATTATAAAGGCACCTGGATTCATCCGGAAATTGATAACCCGGAATATAG
[0442] CCCGGATCCGAGCATTTATGCGTATGATAACTTTGGCGTGCTGGGCCTGGATCTGTGGCAGGTG
[0443] AAAAGCGGCACCATTTTTGATAACTTTCTGATTACCAACGATGAAGCGTATGCGGAAGAATTTG
[0444] GCAACGAAACCTGGGGCGTGACCAAAGCGGCGGAAAAACAGATGAAAGATAAACAGGATGAA
[0445] GAACAGCGCCTGAAAGAAGAAGAAGAAGATAAAAAACGCAAAGAAGAAGAAGAAGCGGAAG
[0446] ATAAAGAAGATGATGAAGATAAAGATGAAGATGAAGAAGATGAAGAAGATAAAGAAGAAGA
[0447] TGAAGAAGAAGATGTGCCGGGCCAGGCGAAAGATGAACTGTGA
[0448] SEQ ID NO:69
[0449] ATGGATGCGATGAAACGCGGCCTGTGCTGCGTGCTGCTGCTGTGCGGCGCGGTGTTTGTGAGCC
[0450] CGATGCATGGCCCGAAAGCGACCCTGCAGGATATTGTGCTGCATCTGGAACCGCAGAACGAAA
[0451] TTCCGGTGGATCTGCTGTGCCATGAACAGCTGAGCGATAGCGAAGAAGAAAACGATGAAATTG
[0452] ATGGCGTGAACCATCAGCATCTGCCGGCGCGCCGCGCGGAACCGCAGCGCCATACCATGCTGT
[0453] GCATGTGCTGCAAATGCGAAGCGCGCATTAAACTGGTGGTGGAAAGCAGCGCGGATGATCTGC
[0454] GCGCGTTTCAGCAGCTGTTTCTGAACACCCTGAGCTTTGTGTGCCCGTGGTGCGCGAGCCAGCA
[0455] GGGCGGCGGCAGCATGGCGCGCTTTGAAGATCCGACCCGCCGCCCGTATAAACTGCCGGATCT
[0456] GTGCACCGAACTGAACACCAGCCTGCAGGATATTGAAATTACCTGCGTGTATTGCAAAACCGT
[0457] GCTGGAACTGACCGAAGTGTTTGAATTTGCGTTTAAAGATCTGTTTGTGGTGTATCGCGATAGC
[0458] ATTCCGCATGCGGCGTGCCATAAATGCATTGATTTTTATAGCCGCATTCGCGAACTGCGCCATT
[0459] ATAGCGATAGCGTGTATGGCGATACCCTGGAAAAACTGACCAACACCGGCCTGTATAACCTGC
[0460] TGATTCGCTGCCTGCGCTGCCAGAAACCGCTGAACCCGGCGGAAAAACTGCGCCATCTGAACG
[0461] AAAAACGCCGCTTTCATAACATTGCGGGCCATTATCGCGGCCAGTGCCATAGCTGCTGCAACCG
[0462] CGCGCGCCAGGAACGCCTGCAGCGCCGCCGCGAAACCCAGGTGGGCAGCGGCGCGACCAACTT
[0463] TAGCCTGCTGAAACAGGCGGGCGATGTGGAAGAAAACCCGGGCCCGAACATTGATCGCCCGAA
[0464] AGGCCTGGCGTTTACCGATGTGGATGTGGATAGCATTAAAATTGCGTGGGAAAGCCCGCAGGG
[0465] CCAGGTGAGCCGCTATCGCGTGACCTATAGCAGCCCGGAAGATGGCATTCATGAACTGTTTCCG
[0466] GCGCCGGATGGCGAAGAAGATACCGCGGAACTGCAGGGCCTGCGCCCGGGCAGCGAATATAC
[0467] CGTGAGCGTGGTGGCGCTGCATGATGATATGGAAAGCCAGCCGCTGATTGGCACCCAGAGCAC
[0468] CTAA
[0469] SEQ ID NO:70
[0470] ATGCGCGTGACCGCGCCGCGCACCGTGCTGCTGCTGCTGAGCGCGGCGCTGGCGCTGACCGAA
[0471] ACCTGGGCGATGCATGGCCCGAAAGCGACCCTGCAGGATATTGTGCTGCATCTGGAACCGCAG
[0472] AACGAAATTCCGGTGGATCTGCTGTGCCATGAACAGCTGAGCGATAGCGAAGAAGAAAACGAT
[0473] GAAATTGATGGCGTGAACCATCAGCATCTGCCGGCGCGCCGCGCGGAACCGCAGCGCCATACC
[0474] ATGCTGTGCATGTGCTGCAAATGCGAAGCGCGCATTAAACTGGTGGTGGAAAGCAGCGCGGAT
[0475] GATCTGCGCGCGTTTCAGCAGCTGTTTCTGAACACCCTGAGCTTTGTGTGCCCGTGGTGCGCGA
[0476] GCCAGCAGGGCGGCGGCAGCATGGCGCGCTTTGAAGATCCGACCCGCCGCCCGTATAAACTGC
[0477] CGGATCTGTGCACCGAACTGAACACCAGCCTGCAGGATATTGAAATTACCTGCGTGTATTGCAA
[0478] AACCGTGCTGGAACTGACCGAAGTGTTTGAATTTGCGTTTAAAGATCTGTTTGTGGTGTATCGC
[0479] GATAGCATTCCGCATGCGGCGTGCCATAAATGCATTGATTTTTATAGCCGCATTCGCGAACTGC
[0480] GCCATTATAGCGATAGCGTGTATGGCGATACCCTGGAAAAACTGACCAACACCGGCCTGTATA
[0481] ACCTGCTGATTCGCTGCCTGCGCTGCCAGAAACCGCTGAACCCGGCGGAAAAACTGCGCCATCT
[0482] GAACGAAAAACGCCGCTTTCATAACATTGCGGGCCATTATCGCGGCCAGTGCCATAGCTGCTGC
[0483] AACCGCGCGCGCCAGGAACGCCTGCAGCGCCGCCGCGAAACCCAGGTGAGCCAGAGCACCATT
[0484] CCGATTGTGGGCATTGTGGCGGGCCTGGCGGTGCTGGCGGTGGTGGTGATTGGCGCGGTGGTG
[0485] GCGACCGTGATGTGCCGCCGCAAAAGCAGCGGCGGCAAAGGCGGCAGCTATAGCCAGGCGGC
[0486] GAGCAGCGATAGCGCGCAGGGCAGCGATGTGAGCCTGACCGCGGGCAGCGGCGCGACCAACT
[0487] TTAGCCTGCTGAAACAGGCGGGCGATGTGGAAGAAAACCCGGGCCCGATGTATCGCATGCAGC
[0488] TGCTGAGCTGCATTGCGCTGAGCCTGGCGCTGGTGACCAACAGCGCGCCGACCAGCAGCAGCA
[0489] CCAAAAAAACCCAGCTGCAGCTGGAACATCTGCTGCTGGATCTGCAGATGATTCTGAACGGCA
[0490] TTAACAACTATAAAAACCCGAAACTGACCCGCATGCTGACCTTTAAATTTTATATGCCGAAAAA
[0491] AGCGACCGAACTGAAACATCTGCAGTGCCTGGAAGAAGAACTGAAACCGCTGGAAGAAGTGCT
[0492] GAACCTGGCGCAGAGCAAAAACTTTCATCTGCGCCCGCGCGATCTGATTAGCAACATTAACGT
[0493] GATTGTGCTGGAACTGAAAGGCAGCGAAACCACCTTTATGTGCGAATATGCGGATGAAACCGC
[0494] GACCATTGTGGAATTTCTGAACCGCTGGATTACCTTTTGCCAGAGCATTATTAGCACCCTGACC
[0495] TAA
[0496] SEQ ID NO:71
[0497] ATGGATGCGATGAAACGCGGCCTGTGCTGCGTGCTGCTGCTGTGCGGCGCGGTGTTTGTGAGCC
[0498] CGATGCATGGCCCGAAAGCGACCCTGCAGGATATTGTGCTGCATCTGGAACCGCAGAACGAAA
[0499] TTCCGGTGGATCTGCTGTGCCATGAACAGCTGAGCGATAGCGAAGAAGAAAACGATGAAATTG
[0500] ATGGCGTGAACCATCAGCATCTGCCGGCGCGCCGCGCGGAACCGCAGCGCCATACCATGCTGT
[0501] GCATGTGCTGCAAATGCGAAGCGCGCATTAAACTGGTGGTGGAAAGCAGCGCGGATGATCTGC
[0502] GCGCGTTTCAGCAGCTGTTTCTGAACACCCTGAGCTTTGTGTGCCCGTGGTGCGCGAGCCAGCA
[0503] GGGCGGCGGCAGCATGGCGCGCTTTGAAGATCCGACCCGCCGCCCGTATAAACTGCCGGATCT
[0504] GTGCACCGAACTGAACACCAGCCTGCAGGATATTGAAATTACCTGCGTGTATTGCAAAACCGT
[0505] GCTGGAACTGACCGAAGTGTTTGAATTTGCGTTTAAAGATCTGTTTGTGGTGTATCGCGATAGC
[0506] ATTCCGCATGCGGCGTGCCATAAATGCATTGATTTTTATAGCCGCATTCGCGAACTGCGCCATT
[0507] ATAGCGATAGCGTGTATGGCGATACCCTGGAAAAACTGACCAACACCGGCCTGTATAACCTGC
[0508] TGATTCGCTGCCTGCGCTGCCAGAAACCGCTGAACCCGGCGGAAAAACTGCGCCATCTGAACG
[0509] AAAAACGCCGCTTTCATAACATTGCGGGCCATTATCGCGGCCAGTGCCATAGCTGCTGCAACCG
[0510] CGCGCGCCAGGAACGCCTGCAGCGCCGCCGCGAAACCCAGGTGAAAGATGAACTGAAAGATG
[0511] AACTGGGCAGCGGCGCGACCAACTTTAGCCTGCTGAAACAGGCGGGCGATGTGGAAGAAAACC
[0512] CGGGCCCGAACATTGATCGCCCGAAAGGCCTGGCGTTTACCGATGTGGATGTGGATAGCATTA
[0513] AAATTGCGTGGGAAAGCCCGCAGGGCCAGGTGAGCCGCTATCGCGTGACCTATAGCAGCCCGG
[0514] AAGATGGCATTCATGAACTGTTTCCGGCGCCGGATGGCGAAGAAGATACCGCGGAACTGCAGG
[0515] GCCTGCGCCCGGGCAGCGAATATACCGTGAGCGTGGTGGCGCTGCATGATGATATGGAAAGCC
[0516] AGCCGCTGATTGGCACCCAGAGCACCTAA
[0517] SEQ ID NO:72
[0518] ATGCGCGTGACCGCGCCGCGCACCGTGCTGCTGCTGCTGAGCGCGGCGCTGGCGCTGACCGAA
[0519] ACCTGGGCGATGCATGGCCCGAAAGCGACCCTGCAGGATATTGTGCTGCATCTGGAACCGCAG
[0520] AACGAAATTCCGGTGGATCTGCTGTGCCATGAACAGCTGAGCGATAGCGAAGAAGAAAACGAT
[0521] GAAATTGATGGCGTGAACCATCAGCATCTGCCGGCGCGCCGCGCGGAACCGCAGCGCCATACC
[0522] ATGCTGTGCATGTGCTGCAAATGCGAAGCGCGCATTAAACTGGTGGTGGAAAGCAGCGCGGAT
[0523] GATCTGCGCGCGTTTCAGCAGCTGTTTCTGAACACCCTGAGCTTTGTGTGCCCGTGGTGCGCGA
[0524] GCCAGCAGAGCCAGAGCACCATTCCGATTGTGGGCATTGTGGCGGGCCTGGCGGTGCTGGCGG
[0525] TGGTGGTGATTGGCGCGGTGGTGGCGACCGTGATGTGCCGCCGCAAAAGCAGCGGCGGCAAAG
[0526] GCGGCAGCTATAGCCAGGCGGCGAGCAGCGATAGCGCGCAGGGCAGCGATGTGAGCCTGACC
[0527] GCGGGCAGCGGCGCGACCAACTTTAGCCTGCTGAAACAGGCGGGCGATGTGGAAGAAAACCC
[0528] GGGCCCGATGGCGCGCGCGGTGGGCATTGATCTGGGCACCACCAACAGCGTGGTGAGCGTGCT
[0529] GGAAGGCGGCGATCCGGTGGTGGTGGCGAACAGCGAAGGCAGCCGCACCACCCCGAGCATTGT
[0530] GGCGTTTGCGCGCAACGGCGAAGTGCTGGTGGGCCAGCCGGCGAAAAACCAGGCGGTGACCA
[0531] ACGTGGATCGCACCGTGCGCAGCGTGAAACGCCATATGGGCAGCGATTGGAGCATTGAAATTG
[0532] ATGGCAAAAAATATACCGCGCCGGAAATTAGCGCGCGCATTCTGATGAAACTGAAACGCGATG
[0533] CGGAAGCGTATCTGGGCGAAGATATTACCGATGCGGTGATTACCACCCCGGCGTATTTTAACGA
[0534] TGCGCAGCGCCAGGCGACCAAAGATGCGGGCCAGATTGCGGGCCTGAACGTGCTGCGCATTGT
[0535] GAACGAACCGACCGCGGCGGCGCTGGCGTATGGCCTGGATAAAGGCGAAAAAGAACAGCGCA
[0536] TTCTGGTGTTTGATCTGGGCGGCGGCACCTTTGATGTGAGCCTGCTGGAAATTGGCGAAGGCGT
[0537] GGTGGAAGTGCGCGCGACCAGCGGCGATAACCATCTGGGCGGCGATGATTGGGATCAGCGCGT
[0538] GGTGGATTGGCTGGTGGATAAATTTAAAGGCACCAGCGGCATTGATCTGACCAAAGATAAAAT
[0539] GGCGATGCAGCGCCTGCGCGAAGCGGCGGAAAAAGCGAAAATTGAACTGAGCAGCAGCCAGA
[0540] GCACCAGCATTAACCTGCCGTATATTACCGTGGATGCGGATAAAAACCCGCTGTTTCTGGATGA
[0541] ACAGCTGACCCGCGCGGAATTTCAGCGCATTACCCAGGATCTGCTGGATCGCACCCGCAAACC
[0542] GTTTCAGAGCGTGATTGCGGATACCGGCATTAGCGTGAGCGAAATTGATCATGTGGTGCTGGTG
[0543] GGCGGCAGCACCCGCATGCCGGCGGTGACCGATCTGGTGAAAGAACTGACCGGCGGCAAAGA
[0544] ACCGAACAAAGGCGTGAACCCGGATGAAGTGGTGGCGGTGGGCGCGGCGCTGCAGGCGGGCG
[0545] TGCTGAAAGGCGAAGTGAAAGATGTGCTGCTGCTGGATGTGACCCCGCTGAGCCTGGGCATTG
[0546] AAACCAAAGGCGGCGTGATGACCCGCCTGATTGAACGCAACACCACCATTCCGACCAAACGCA
[0547] GCGAAACCTTTACCACCGCGGATGATAACCAGCCGAGCGTGCAGATTCAGGTGTATCAGGGCG
[0548] AACGCGAAATTGCGGCGCATAACAAACTGCTGGGCAGCTTTGAACTGACCGGCATTCCGCCGG
[0549] CGCCGCGCGGCATTCCGCAGATTGAAGTGACCTTTGATATTGATGCGAACGGCATTGTGCATGT
[0550] GACCGCGAAAGATAAAGGCACCGGCAAAGAAAACACCATTCGCATTCAGGAAGGCAGCGGCC
[0551] TGAGCAAAGAAGATATTGATCGCATGATTAAAGATGCGGAAGCGCATGCGGAAGAAGATCGC
[0552] AAACGCCGCGAAGAAGCGGATGTGCGCAACCAGGCGGAAACCCTGGTGTATCAGACCGAAAA
[0553] ATTTGTGAAAGAACAGCGCGAAGCGGAAGGCGGCAGCAAAGTGCCGGAAGATACCCTGAACA
[0554] AAGTGGATGCGGCGGTGGCGGAAGCGAAAGCGGCGCTGGGCGGCAGCGATATTAGCGCGATT
[0555] AAAAGCGCGATGGAAAAACTGGGCCAGGAAAGCCAGGCGCTGGGCCAGGCGATTTATGAAGC
[0556] GGCGCAGGCGGCGAGCCAGGCGACCGGCGCGGCGCATCCGGGCGGCGAACCGGGCGGCGCGC
[0557] ATCCGGGCAGCGCGGATGATGTGGTGGATGCGGAAGTGGTGGATGATGGCCGCGAAGCGAAASEQ IDNO:73
[0558] ATGGATGCGATGAAACGCGGCCTGTGCTGCGTGCTGCTGCTGTGCGGCGCGGTGTTTGTGAGCC
[0559] CGCATGGCGATACCCCGACCCTGCATGAATATATGCTGGATCTGCAGCCGGAAACCACCGATCT
[0560] GTATTGCTATGAACAGCTGAACGATAGCAGCGAAGAAGAAGATGAAATTGATGGCCCGGCGGG
[0561] CCAGGCGGAACCGGATCGCGCGCATTATAACATTGTGACCTTTTGCTGCAAATGCGATAGCACC
[0562] CTGCGCCTGTGCGTGCAGAGCACCCATGTGGATATTCGCACCCTGGAAGATCTGCTGATGGGCA
[0563] CCCTGGGCATTGTGTGCCCGATTTGCAGCCAGAAACCGGGCGGCGGCGGCAGCGGCGGCGGCG
[0564] GCAGCATGCATGGCCCGAAAGCGACCCTGCAGGATATTGTGCTGCATCTGGAACCGCAGAACG
[0565] AAATTCCGGTGGATCTGCTGTGCCATGAACAGCTGAGCGATAGCGAAGAAGAAAACGATGAAA
[0566] TTGATGGCGTGAACCATCAGCATCTGCCGGCGCGCCGCGCGGAACCGCAGCGCCATACCATGC
[0567] TGTGCATGTGCTGCAAATGCGAAGCGCGCATTAAACTGGTGGTGGAAAGCAGCGCGGATGATC
[0568] TGCGCGCGTTTCAGCAGCTGTTTCTGAACACCCTGAGCTTTGTGTGCCCGTGGTGCGCGAGCCA
[0569] GCAGGGCAGCGGCAAAGATGAACTGAAAGATGAACTGGCGACCAACTTTAGCCTGCTGAAACA
[0570] GGCGGGCGATGTGGAAGAAAACCCGGGCCCGAACATTGATCGCCCGAAAGGCCTGGCGTTTAC
[0571] CGATGTGGATGTGGATAGCATTAAAATTGCGTGGGAAAGCCCGCAGGGCCAGGTGAGCCGCTA
[0572] TCGCGTGACCTATAGCAGCCCGGAAGATGGCATTCATGAACTGTTTCCGGCGCCGGATGGCGA
[0573] AGAAGATACCGCGGAACTGCAGGGCCTGCGCCCGGGCAGCGAATATACCGTGAGCGTGGTGGC
[0574] GCTGCATGATGATATGGAAAGCCAGCCGCTGATTGGCACCCAGAGCACCTAA
[0575] SEQ ID NO:74
[0576] ATGGATGCGATGAAACGCGGCCTGTGCTGCGTGCTGCTGCTGTGCGGCGCGGTGTTTGTGAGCC
[0577] CGATGGCGCGCTTTGAAGATCCGACCCGCCGCCCGTATAAACTGCCGGATCTGTGCACCGAACT
[0578] GAACACCAGCCTGCAGGATATTGAAATTACCTGCGTGTATTGCAAAACCGTGCTGGAACTGAC
[0579] CGAAGTGTTTGAATTTGCGTTTAAAGATCTGTTTGTGGTGTATCGCGATAGCATTCCGCATGCG
[0580] GCGTGCCATAAATGCATTGATTTTTATAGCCGCATTCGCGAACTGCGCCATTATAGCGATAGCG
[0581] TGTATGGCGATACCCTGGAAAAACTGACCAACACCGGCCTGTATAACCTGCTGATTCGCTGCCT
[0582] GCGCTGCCAGAAACCGCTGAACCCGGCGGAAAAACTGCGCCATCTGAACGAAAAACGCCGCTT
[0583] TCATAACATTGCGGGCCATTATCGCGGCCAGTGCCATAGCTGCTGCAACCGCGCGCGCCAGGA
[0584] ACGCCTGCAGCGCCGCCGCGAAACCCAGGTGGGCGGCGGCGGCAGCATGGCGCGCTTTGAAGA
[0585] TCCGACCCGCCGCCCGTATAAACTGCCGGATCTGTGCACCGAACTGAACACCAGCCTGCAGGA
[0586] TATTGAAATTACCTGCGTGTATTGCAAAACCGTGCTGGAACTGACCGAAGTGTTTGAATTTGCG
[0587] TTTAAAGATCTGTTTGTGGTGTATCGCGATAGCATTCCGCATGCGGCGTGCCATAAATGCATTG
[0588] ATTTTTATAGCCGCATTCGCGAACTGCGCCATTATAGCGATAGCGTGTATGGCGATACCCTGGA
[0589] AAAACTGACCAACACCGGCCTGTATAACCTGCTGATTCGCTGCCTGCGCTGCCAGAAACCGCTG
[0590] AACCCGGCGGAAAAACTGCGCCATCTGAACGAAAAACGCCGCTTTCATAACATTGCGGGCCAT
[0591] TATCGCGGCCAGTGCCATAGCTGCTGCAACCGCGCGCGCCAGGAACGCCTGCAGCGCCGCCGC
[0592] GAAACCCAGGTGAAAGATGAACTGAAAGATGAACTGGGCAGCGGCGCGACCAACTTTAGCCTG
[0593] CTGAAACAGGCGGGCGATGTGGAAGAAAACCCGGGCCCGAACATTGATCGCCCGAAAGGCCT
[0594] GGCGTTTACCGATGTGGATGTGGATAGCATTAAAATTGCGTGGGAAAGCCCGCAGGGCCAGGT
[0595] GAGCCGCTATCGCGTGACCTATAGCAGCCCGGAAGATGGCATTCATGAACTGTTTCCGGCGCCG
[0596] GATGGCGAAGAAGATACCGCGGAACTGCAGGGCCTGCGCCCGGGCAGCGAATATACCGTGAG
[0597] CGTGGTGGCGCTGCATGATGATATGGAAAGCCAGCCGCTGATTGGCACCCAGAGCACCTAASEQ IDNO:75
[0598] ATGGATGCGATGAAACGCGGCCTGTGCTGCGTGCTGCTGCTGTGCGGCGCGGTGTTTGTGAGCC
[0599] CGATGCATGGCCCGAAAGCGACCCTGCAGGATATTGTGCTGCATCTGGAACCGCAGAACGAAA
[0600] TTCCGGTGGATCTGCTGTGCCATGAACAGCTGAGCGATAGCGAAGAAGAAAACGATGAAATTG
[0601] ATGGCGTGAACCATCAGCATCTGCCGGCGCGCCGCGCGGAACCGCAGCGCCATACCATGCTGT
[0602] GCATGTGCTGCAAATGCGAAGCGCGCATTAAACTGGTGGTGGAAAGCAGCGCGGATGATCTGC
[0603] GCGCGTTTCAGCAGCTGTTTCTGAACACCCTGAGCTTTGTGTGCCCGTGGTGCGCGAGCCAGCA
[0604] GGGCGGCGGCGGCAGCGGCGGCGGCGGCAGCCATGGCGATACCCCGACCCTGCATGAATATAT
[0605] GCTGGATCTGCAGCCGGAAACCACCGATCTGTATTGCTATGAACAGCTGAACGATAGCAGCGA
[0606] AGAAGAAGATGAAATTGATGGCCCGGCGGGCCAGGCGGAACCGGATCGCGCGCATTAACAT
[0607] TGTGACCTTTTGCTGCAAATGCGATAGCACCCTGCGCCTGTGCGTGCAGAGCACCCATGTGGAT
[0608] ATTCGCACCCTGGAAGATCTGCTGATGGGCACCCTGGGCATTGTGTGCCCGATTTGCAGCCCAGA
[0609] AACCGAAAGATGAACTGAAAGATGAACTGGGCAGCGGCGCGACCAACTTTAGCCTGCTGAAAC
[0610] AGGCGGGCGATGTGGAAGAAAACCCGGGCCCGAACATTTGATCGCCCGAAAGGCCTGGCGTTTA
[0611] CCGATGTGGATGTGGATAGCATTAAATTGCGTGGGAAAGCCCGCAGGGCCAGGTGAGCCGCT
[0612] ATCGCGTGACCTAGCAGCCCGGAAGATGGCATTCATGAACTGTTTCCGGCGCCGGATGGCG
[0613] AAGAAGATACCGCGGAACTGCAGGGCCTGCGCCCGGGCAGCGAATATACCGTGAGCGTGGTGG
[0614] CGCTGCATGATGATATGGAAAGCCAGCCGCTGATTGGCACCCAGAGCACCTAA
[0615] SEQ ID NO:76
[0616] ATGGATGCGATGAAACGCGGCCTGTGCTGCGTGCTGCTGCTGTGCGGCGCGGTGTTTGTGAGCC
[0617] CGATGGCGCGCTTTGAAGATCCGACCCGCCGCCCGTATAAACTGCCGGATCTGTGCACCGAACT
[0618] GAACACCAGCCTGCAGGATATTGAAATTACCTGCGTGTATTGCAAAACCGTGCTGGAACTGAC
[0619] CGAAGTGTTTGAATTTGCGTTTAAAGATCTGTTTGTGGTGTATCGCGATAGCATTCCGCATGCG
[0620] GCGTGCCATAAATGCATTGATTTTTATAGCCGCATTCGCGAACTGCGCCATTATAGCGATAGCG
[0621] TGTATGGCGATACCCTGGAAAAACTGACCAACACCGGCCTGTATAACCTGCTGATTCGCTGCCT
[0622] GCGCTGCCAGAAACCGCTGAACCCGGCGGAAAAACTGCGCCATCTGAACGAAAAACGCCGCTT
[0623] TCATAACATTGCGGGCCATTATCGCGGCCAGTGCCATAGCTGCTGCAACCGCGCGCGCCAGGA
[0624] ACGCCTGCAGCGCCGCCGCGAAACCCAGGTGGGCGGCGGCGGCAGCATGGCGCGCTTTGAAGA
[0625] TCCGACCCGCCGCCCGTATAAACTGCCGGATCTGTGCACCGAACTGAACACCAGCCTGCAGGA
[0626] TATTGAAATTACCTGCGTGTATTGCAAAACCGTGCTGGAACTGACCGAAGTGTTTGAATTTGCG
[0627] TTTAAAGATCTGTTTGTGGTGTATCGCGATAGCATTCCGCATGCGGCGTGCCATAAATGCATTG
[0628] ATTTTTATAGCCGCATTCGCGAACTGCGCCATTATAGCGATAGCGTGTATGGCGATACCCTGGA
[0629] AAAACTGACCAACACCGGCCTGTATAACCTGCTGATTCGCTGCCTGCGCTGCCAGAAACCGCTG
[0630] AACCCGGCGGAAAAACTGCGCCATCTGAACGAAAAACGCCGCTTTCATAACATTGCGGGCCAT
[0631] TATCGCGGCCAGTGCCATAGCTGCTGCAACCGCGCGCGCCAGGAACGCCTGCAGCGCCGCCGC
[0632] GAAACCCAGGTGAAAGATGAACTGAAAGATGAACTGGGCAGCGGCGCGACCAACTTTAGCCTG
[0633] CTGAAACAGGCGGGCGATGTGGAAGAAAACCCGGGCCCGAACATTGATCGCCCGAAAGGCCT
[0634] GGCGTTTACCGATGTGGATGTGGATAGCATTAAAATTGCGTGGGAAAGCCCGCAGGGCCAGGT
[0635] GAGCCGCTATCGCGTGACCTATAGCAGCCCGGAAGATGGCATTCATGAACTGTTTCCGGCGCCG
[0636] GATGGCGAAGAAGATACCGCGGAACTGCAGGGCCTGCGCCCGGGCAGCGAATATACCGTGAG
[0637] CGTGGTGGCGCTGCATGATGATATGGAAAGCCAGCCGCTGATTGGCACCCAGAGCACCTAASEQ IDNO:77
[0638] ATGGATGCGATGAAACGCGGCCTGTGCTGCGTGCTGCTGCTGTGCGGCGCGGTGTTTGTGAGCC
[0639] CGATGGCGCGCTTTGAAGATCCGACCCGCCGCCCGTATAAACTGCCGGATCTGTGCACCGAACT
[0640] GAACACCAGCCTGCAGGATATTGAAATTACCTGCGTGTATTGCAAAACCGTGCTGGAACTGAC
[0641] CGAAGTGTTTGAATTTGCGTTTAAAGATCTGTTTGTGGTGTATCGCGATAGCATTCCGCATGCG
[0642] GCGTGCCATAAATGCATTGATTTTTATAGCCGCATTCGCGAACTGCGCCATTATAGCGATAGCG
[0643] TGTATGGCGATACCCTGGAAAAACTGACCAACACCGGCCTGTATAACCTGCTGATTCGCTGCCT
[0644] GCGCTGCCAGAAACCGCTGAACCCGGCGGAAAAACTGCGCCATCTGAACGAAAAACGCCGCTT
[0645] TCATAACATTGCGGGCCATTATCGCGGCCAGTGCCATAGCTGCTGCAACCGCGCGCGCCAGGA
[0646] ACGCCTGCAGCGCCGCCGCGAAACCCAGGTGGGCGGCGGCGGCAGCATGCATGGCCCGAAAG
[0647] CGACCCTGCAGGATATTGTGCTGCATCTGGAACCGCAGAACGAAATTCCGGTGGATCTGCTGTG
[0648] CCATGAACAGCTGAGCGATAGCGAAGAAGAAAACGATGAAATTGATGGCGTGAACCATCAGC
[0649] ATCTGCCGGCGCGCCGCGCGGAACCGCAGCGCCATACCATGCTGTGCATGTGCTGCAAATGCG
[0650] AAGCGCGCATTAAACTGGTGGTGGAAAGCAGCGCGGATGATCTGCGCGCGTTTCAGCAGCTGT
[0651] TTCTGAACACCCTGAGCTTTGTGTGCCCGTGGTGCGCGAGCCAGCAGAAAGATGAACTGAAAG
[0652] ATGAACTGGGCAGCGGCGCGACCAACTTTAGCCTGCTGAAACAGGCGGGCGATGTGGAAGAAA
[0653] ACCCGGGCCCGAACATTGATCGCCCGAAAGGCCTGGCGTTTACCGATGTGGATGTGGATAGCA
[0654] TTAAAATTGCGTGGGAAAGCCCGCAGGGCCAGGTGAGCCGCTATCGCGTGACCTATAGCAGCC
[0655] CGGAAGATGGCATTCATGAACTGTTTCCGGCGCCGGATGGCGAAGAAGATACCGCGGAACTGC
[0656] AGGGCCTGCGCCCGGGCAGCGAATATACCGTGAGCGTGGTGGCGCTGCATGATGATATGGAAA
[0657] GCCAGCCGCTGATTGGCACCCAGAGCACCTAA
[0658] SEQ ID NO:78
[0659] ATGGATGCGATGAAACGCGGCCTGTGCTGCGTGCTGCTGCTGTGCGGCGCGGTGTTTGTGAGCC
[0660] CGCATGGCGATACCCCGACCCTGCATGAATATATGCTGGATCTGCAGCCGGAAACCACCGATCT
[0661] GTATTGCTATGAACAGCTGAACGATAGCAGCGAAGAAGAAGATGAAATTGATGGCCCGGCGGG
[0662] CCAGGCGGAACCGGATCGCGCGCATTATAACATTGTGACCTTTTGCTGCAAATGCGATAGCACC
[0663] CTGCGCCTGTGCGTGCAGAGCACCCATGTGGATATTCGCACCCTGGAAGATCTGCTGATGGGCA
[0664] CCCTGGGCATTGTGTGCCCGATTTGCAGCCAGAAACCGGGCGGCGGCGGCAGCGGCGGCGGCG
[0665] GCAGCATGGCGCGCTTTGAAGATCCGACCCGCCGCCCGTATAAACTGCCGGATCTGTGCACCG
[0666] AACTGAACACCAGCCTGCAGGATATTGAAATTACCTGCGTGTATTGCAAAACCGTGCTGGAACT
[0667] GACCGAAGTGTTTGAATTTGCGTTTAAAGATCTGTTTGTGGTGTATCGCGATAGCATTCCGCAT
[0668] GCGGCGTGCCATAAATGCATTGATTTTTATAGCCGCATTCGCGAACTGCGCCATTATAGCGATA
[0669] GCGTGTATGGCGATACCCTGGAAAAACTGACCAACACCGGCCTGTATAACCTGCTGATTCGCTG
[0670] CCTGCGCTGCCAGAAACCGCTGAACCCGGCGGAAAAACTGCGCCATCTGAACGAAAAACGCCG
[0671] CTTTCATAACATTGCGGGCCATTATCGCGGCCAGTGCCATAGCTGCTGCAACCGCGCGCGCCAG
[0672] GAACGCCTGCAGCGCCGCCGCGAAACCCAGGTGAAAGATGAACTGAAAGATGAACTGGGCAG
[0673] CGGCGCGACCAACTTTAGCCTGCTGAAACAGGCGGGCGATGTGGAAGAAAACCCGGGCCCGAA
[0674] CATTGATCGCCCGAAAGGCCTGGCGTTTACCGATGTGGATGTGGATAGCATTAAAATTGCGTGG
[0675] GAAAGCCCGCAGGGCCAGGTGAGCCGCTATCGCGTGACCTATAGCAGCCCGGAAGATGGCATT
[0676] CATGAACTGTTTCCGGCGCCGGATGGCGAAGAAGATACCGCGGAACTGCAGGGCCTGCGCCCG
[0677] GGCAGCGAATATACCGTGAGCGTGGTGGCGCTGCATGATGATATGGAAAGCCAGCCGCTGATT
[0678] GGCACCCAGAGCACCTAA
[0679] SEQ ID NO:79
[0680] ATGGATGCGATGAAACGCGGCCTGTGCTGCGTGCTGCTGCTGTGCGGCGCGGTGTTTGTGAGCC
[0681] CGATGCATGGCCCGAAAGCGACCCTGCAGGATATTGTGCTGCATCTGGAACCGCAGAACGAAA
[0682] TTCCGGTGGATCTGCTGTGCCATGAACAGCTGAGCGATAGCGAAGAAGAAAACGATGAAATTG
[0683] ATGGCGTGAACCATCAGCATCTGCCGGCGCGCCGCGCGGAACCGCAGCGCCATACCATGCTGT
[0684] GCATGTGCTGCAAATGCGAAGCGCGCATTAAACTGGTGGTGGAAAGCAGCGCGGATGATCTGC
[0685] GCGCGTTTCAGCAGCTGTTTCTGAACACCCTGAGCTTTGTGTGCCCGTGGTGCGCGAGCCAGCA
[0686] GGGCGGCGGCGGCAGCGGCGGCGGCGGCAGCATGCATCAGAAACGCACCGCGATGTTTCAGG
[0687] ATCCGCAGGAACGCCCGCGCAAACTGCCGCAGCTGTGCACCGAACTGCAGACCACCATTCATG
[0688] ATATTATTCTGGAATGCGTGTATTGCAAACAGCAGCTGCTGCGCCGCGAAGTGTATGATTTTGC
[0689] GTTTCGCGATCTGTGCATTGTGTATCGCGATGGCAACCCGTATGCGGTGTGCGATAAATGCCTG
[0690] AAATTTTATAGCAAAATTAGCGAATATCGCCATTATTGCTATAGCCTGTATGGCACCACCCTGG
[0691] AACAGCAGTATAACAAACCGCTGTGCGATCTGCTGATTCGCTGCATTAACTGCCAGAAACCGCT
[0692] GTGCCCGGAAGAAAAACAGCGCCATCTGGATAAAAAACAGCGCTTTCATAACATTCGCGGCCG
[0693] CTGGACCGGCCGCTGCATGAGCTGCTGCCGCAGCAGCCGCACCCGCCGCGAAACCCAGCTGAA
[0694] AGATGAACTGAAAGATGAACTGGGCAGCGGCGCGACCAACTTTAGCCTGCTGAAACAGGCGGG
[0695] CGATGTGGAAGAAAACCCGGGCCCGAACATTGATCGCCCGAAAGGCCTGGCGTTTACCGATGT
[0696] GGATGTGGATAGCATTAAAATTGCGTGGGAAAGCCCGCAGGGCCAGGTGAGCCGCTATCGCGT
[0697] GACCTATAGCAGCCCGGAAGATGGCATTCATGAACTGTTTCCGGCGCCGGATGGCGAAGAAGA
[0698] TACCGCGGAACTGCAGGGCCTGCGCCCGGGCAGCGAATATACCGTGAGCGTGGTGGCGCTGCA
[0699] TGATGATATGGAAAGCCAGCCGCTGATTGGCACCCAGAGCACCTAA
[0700] SEQ ID NO:80
[0701] ATGGATGCGATGAAACGCGGCCTGTGCTGCGTGCTGCTGCTGTGCGGCGCGGTGTTTGTGAGCC
[0702] CGCATGGCGATACCCCGACCCTGCATGAATATATGCTGGATCTGCAGCCGGAAACCACCGATCT
[0703] GTATTGCTATGAACAGCTGAACGATAGCAGCGAAGAAGAAGATGAAATTGATGGCCCGGCGGG
[0704] CCAGGCGGAACCGGATCGCGCGCATTATAACATTGTGACCTTTTGCTGCAAATGCGATAGCACC
[0705] CTGCGCCTGTGCGTGCAGAGCACCCATGTGGATATTCGCACCCTGGAAGATCTGCTGATGGGCA
[0706] CCCTGGGCATTGTGTGCCCGATTTGCAGCCAGAAACCGGGCGGCGGCAGCATGCATCAGAAAC
[0707] GCACCGCGATGTTTCAGGATCCGCAGGAACGCCCGCGCAAACTGCCGCAGCTGTGCACCGAAC
[0708] TGCAGACCACCATTCATGATATTATTCTGGAATGCGTGTATTGCAAACAGCAGCTGCTGCGCCG
[0709] CGAAGTGTATGATTTTGCGTTTCGCGATCTGTGCATTGTGTATCGCGATGGCAACCCGTATGCG
[0710] GTGTGCGATAAATGCCTGAAATTTTATAGCAAAATTAGCGAATATCGCCATTATTGCTATAGCC
[0711] TGTATGGCACCACCCTGGAACAGCAGTATAACAAACCGCTGTGCGATCTGCTGATTCGCTGCAT
[0712] TAACTGCCAGAAACCGCTGTGCCCGGAAGAAAAACAGCGCCATCTGGATAAAAAACAGCGCTT
[0713] TCATAACATTCGCGGCCGCTGGACCGGCCGCTGCATGAGCTGCTGCCGCAGCAGCCGCACCCG
[0714] CCGCGAAACCCAGCTGGGCGGCGGCGGCAGCGGCGGCGGCGGCAGCGGCGGCGGCGGCAGCA
[0715] TGCATGGCCCGAAAGCGACCCTGCAGGATATTGTGCTGCATCTGGAACCGCAGAACGAAATTC
[0716] CGGTGGATCTGCTGTGCCATGAACAGCTGAGCGATAGCGAAGAAGAAAACGATGAAATTGATG
[0717] GCGTGAACCATCAGCATCTGCCGGCGCGCCGCGCGGAACCGCAGCGCCATACCATGCTGTGCA
[0718] TGTGCTGCAAATGCGAAGCGCGCATTAAACTGGTGGTGGAAAGCAGCGCGGATGATCTGCGCG
[0719] CGTTTCAGCAGCTGTTTCTGAACACCCTGAGCTTTGTGTGCCCGTGGTGCGCGAGCCAGCAGGG
[0720] CGGCGGCAGCATGGCGCGCTTTGAAGATCCGACCCGCCGCCCGTATAAACTGCCGGATCTGTG
[0721] CACCGAACTGAACACCAGCCTGCAGGATATTGAAATTACCTGCGTGTATTGCAAAACCGTGCTG
[0722] GAACTGACCGAAGTGTTTGAATTTGCGTTTAAAGATCTGTTTGTGGTGTATCGCGATAGCATTC
[0723] CGCATGCGGCGTGCCATAAATGCATTGATTTTTATAGCCGCATTCGCGAACTGCGCCATTATAG
[0724] CGATAGCGTGTATGGCGATACCCTGGAAAAACTGACCAACACCGGCCTGTATAACCTGCTGATT
[0725] CGCTGCCTGCGCTGCCAGAAACCGCTGAACCCGGCGGAAAAACTGCGCCATCTGAACGAAAAA
[0726] CGCCGCTTTCATAACATTGCGGGCCATTATCGCGGCCAGTGCCATAGCTGCTGCAACCGCGCGC
[0727] GCCAGGAACGCCTGCAGCGCCGCCGCGAAACCCAGGTGAAAGATGAACTGAAAGATGAACTG
[0728] GGCAGCGGCGCGACCAACTTTAGCCTGCTGAAACAGGCGGGCGATGTGGAAGAAAACCCGGG
[0729] CCCGAACATTGATCGCCCGAAAGGCCTGGCGTTTACCGATGTGGATGTGGATAGCATTAAAATT
[0730] GCGTGGGAAAGCCCGCAGGGCCAGGTGAGCCGCTATCGCGTGACCTATAGCAGCCCGGAAGAT
[0731] GGCATTCATGAACTGTTTCCGGCGCCGGATGGCGAAGAAGATACCGCGGAACTGCAGGGCCTG
[0732] CGCCCGGGCAGCGAATATACCGTGAGCGTGGTGGCGCTGCATGATGATATGGAAAGCCAGCCG
[0733] CTGATTGGCACCCAGAGCACCTAA
[0734] SEQ ID NO:81
[0735] ATGGATGCGATGAAACGCGGCCTGTGCTGCGTGCTGCTGCTGTGCGGCGCGGTGTTTGTGAGCC
[0736] CGATGCATGGCCCGAAAGCGACCCTGCAGGATATTGTGCTGCATCTGGAACCGCAGAACGAAA
[0737] TTCCGGTGGATCTGCTGTGCCATGAACAGCTGAGCGATAGCGAAGAAGAAAACGATGAAATTG
[0738] ATGGCGTGAACCATCAGCATCTGCCGGCGCGCCGCGCGGAACCGCAGCGCCATACCATGCTGT
[0739] GCATGTGCTGCAAATGCGAAGCGCGCATTAAACTGGTGGTGGAAAGCAGCGCGGATGATCTGC
[0740] GCGCGTTTCAGCAGCTGTTTCTGAACACCCTGAGCTTTGTGTGCCCGTGGTGCGCGAGCCAGCA
[0741] GGGCGGCGGCAGCATGGCGCGCTTTGAAGATCCGACCCGCCGCCCGTATAAACTGCCGGATCT
[0742] GTGCACCGAACTGAACACCAGCCTGCAGGATATTGAAATTACCTGCGTGTATTGCAAAACCGT
[0743] GCTGGAACTGACCGAAGTGTTTGAATTTGCGTTTAAAGATCTGTTTGTGGTGTATCGCGATAGC
[0744] ATTCCGCATGCGGCGTGCCATAAATGCATTGATTTTTATAGCCGCATTCGCGAACTGCGCCATT
[0745] ATAGCGATAGCGTGTATGGCGATACCCTGGAAAAACTGACCAACACCGGCCTGTATAACCTGC
[0746] TGATTCGCTGCCTGCGCTGCCAGAAACCGCTGAACCCGGCGGAAAAACTGCGCCATCTGAACG
[0747] AAAAACGCCGCTTTCATAACATTGCGGGCCATTATCGCGGCCAGTGCCATAGCTGCTGCAACCG
[0748] CGCGCGCCAGGAACGCCTGCAGCGCCGCCGCGAAACCCAGGTGGGCGGCGGCGGCAGCGGCG
[0749] GCGGCGGCAGCGGCGGCGGCGGCAGCCATGGCGATACCCCGACCCTGCATGAATATATGCTGG
[0750] ATCTGCAGCCGGAAACCACCGATCTGTATTGCTATGAACAGCTGAACGATAGCAGCGAAGAAG
[0751] AAGATGAAATTGATGGCCCGGCGGGCCAGGCGGAACCGGATCGCGCGCATTATAACATTGTGA
[0752] CCTTTTGCTGCAAATGCGATAGCACCCTGCGCCTGTGCGTGCAGAGCACCCATGTGGATATTCG
[0753] CACCCTGGAAGATCTGCTGATGGGCACCCTGGGCATTGTGTGCCCGATTTGCAGCCAGAAACCG
[0754] GGCGGCGGCAGCATGCATCAGAAACGCACCGCGATGTTTCAGGATCCGCAGGAACGCCCGCGC
[0755] AAACTGCCGCAGCTGTGCACCGAACTGCAGACCACCATTCATGATATTATTCTGGAATGCGTGT
[0756] ATTGCAAACAGCAGCTGCTGCGCCGCGAAGTGTATGATTTTGCGTTTCGCGATCTGTGCATTGT
[0757] GTATCGCGATGGCAACCCGTATGCGGTGTGCGATAAATGCCTGAAATTTTATAGCAAAATTAGC
[0758] GAATATCGCCATTATTGCTATAGCCTGTATGGCACCACCCTGGAACAGCAGTATAACAAACCGC
[0759] TGTGCGATCTGCTGATTCGCTGCATTAACTGCCAGAAACCGCTGTGCCCGGAAGAAAAACAGC
[0760] GCCATCTGGATAAAAAACAGCGCTTTCATAACATTCGCGGCCGCTGGACCGGCCGCTGCATGA
[0761] GCTGCTGCCGCAGCAGCCGCACCCGCCGCGAAACCCAGCTGAAAGATGAACTGAAAGATGAAC
[0762] TGGGCAGCGGCGCGACCAACTTTAGCCTGCTGAAACAGGCGGGCGATGTGGAAGAAAACCCGG
[0763] GCCCGAACATTGATCGCCCGAAAGGCCTGGCGTTTACCGATGTGGATGTGGATAGCATTAAAA
[0764] TTGCGTGGGAAAGCCCGCAGGGCCAGGTGAGCCGCTATCGCGTGACCTATAGCAGCCCGGAAG
[0765] ATGGCATTCATGAACTGTTTCCGGCGCCGGATGGCGAAGAAGATACCGCGGAACTGCAGGGCC
[0766] TGCGCCCGGGCAGCGAATATACCGTGAGCGTGGTGGCGCTGCATGATGATATGGAAAGCCAGC
[0767] CGCTGATTGGCACCCAGAGCACCTAA
[0768] SEQ ID NO:82
[0769] ATGCGCGTGACCGCGCCGCGCACCGTGCTGCTGCTGCTGAGCGCGGCGCTGGCGCTGACCGAA
[0770] ACCTGGGCGATGCATGGCCCGAAAGCGACCCTGCAGGATATTGTGCTGCATCTGGAACCGCAG
[0771] AACGAAATTCCGGTGGATCTGCTGTGCCATGAACAGCTGAGCGATAGCGAAGAAGAAAACGAT
[0772] GAAATTGATGGCGTGAACCATCAGCATCTGCCGGCGCGCCGCGCGGAACCGCAGCGCCATACC
[0773] ATGCTGTGCATGTGCTGCAAATGCGAAGCGCGCATTAAACTGGTGGTGGAAAGCAGCGCGGAT
[0774] GATCTGCGCGCGTTTCAGCAGCTGTTTCTGAACACCCTGAGCTTTGTGTGCCCGTGGTGCGCGA
[0775] GCCAGCAGGGCGGCGGCGGCAGCGGCGGCGGCGGCAGCCATGGCGATACCCCGACCCTGCAT
[0776] GAATATATGCTGGATCTGCAGCCGGAAACCACCGATCTGTATTGCTATGAACAGCTGAACGAT
[0777] AGCAGCGAAGAAGAAGATGAAATTGATGGCCCGGCGGGCCAGGCGGAACCGGATCGCGCGCA
[0778] TTATAACATTGTGACCTTTTGCTGCAAATGCGATAGCACCCTGCGCCTGTGCGTGCAGAGCACC
[0779] CATGTGGATATTCGCACCCTGGAAGATCTGCTGATGGGCACCCTGGGCATTGTGTGCCCGATTT
[0780] GCAGCCAGAAACCGAAAGATGAACTGAAAGATGAACTGGGCAGCGGCGCGACCAACTTTAGC
[0781] CTGCTGAAACAGGCGGGCGATGTGGAAGAAAACCCGGGCCCGATGTATCGCATGCAGCTGCTG
[0782] AGCTGCATTGCGCTGAGCCTGGCGCTGGTGACCAACAGCGCGCCGACCAGCAGCAGCACCAAA
[0783] AAAACCCAGCTGCAGCTGGAACATCTGCTGCTGGATCTGCAGATGATTCTGAACGGCATTAAC
[0784] AACTATAAAAACCCGAAACTGACCCGCATGCTGACCTTTAAATTTTATATGCCGAAAAAAGCG
[0785] ACCGAACTGAAACATCTGCAGTGCCTGGAAGAAGAACTGAAACCGCTGGAAGAAGTGCTGAAC
[0786] CTGGCGCAGAGCAAAAACTTTCATCTGCGCCCGCGCGATCTGATTAGCAACATTAACGTGATTG
[0787] TGCTGGAACTGAAAGGCAGCGAAACCACCTTTATGTGCGAATATGCGGATGAAACCGCGACCA
[0788] TTGTGGAATTTCTGAACCGCTGGATTACCTTTTGCCAGAGCATTATTAGCACCCTGACCTAASEQ IDNO:83
[0789] ATGCGCGTGACCGCGCCGCGCACCGTGCTGCTGCTGCTGAGCGCGGCGCTGGCGCTGACCGAA
[0790] ACCTGGGCGATGGCGCGCTTTGAAGATCCGACCCGCCGCCCGTATAAACTGCCGGATCTGTGCA
[0791] CCGAACTGAACACCAGCCTGCAGGATATTGAAATTACCTGCGTGTATTGCAAAACCGTGCTGG
[0792] AACTGACCGAAGTGTTTGAATTTGCGTTTAAAGATCTGTTTGTGGTGTATCGCGATAGCATTCC
[0793] GCATGCGGCGTGCCATAAATGCATTGATTTTTATAGCCGCATTCGCGAACTGCGCCATTATAGC
[0794] GATAGCGTGTATGGCGATACCCTGGAAAAACTGACCAACACCGGCCTGTATAACCTGCTGATTC
[0795] GCTGCCTGCGCTGCCAGAAACCGCTGAACCCGGCGGAAAAACTGCGCCATCTGAACGAAAAAC
[0796] GCCGCTTTCATAACATTGCGGGCCATTATCGCGGCCAGTGCCATAGCTGCTGCAACCGCGCGCG
[0797] CCAGGAACGCCTGCAGCGCCGCCGCGAAACCCAGGTGGGCGGCGGCGGCAGCATGGCGCGCTT
[0798] TGAAGATCCGACCCGCCGCCCGTATAAACTGCCGGATCTGTGCACCGAACTGAACACCAGCCT
[0799] GCAGGATATTGAAATTACCTGCGTGTATTGCAAAACCGTGCTGGAACTGACCGAAGTGTTTGAA
[0800] TTTGCGTTTAAAGATCTGTTTGTGGTGTATCGCGATAGCATTCCGCATGCGGCGTGCCATAAAT
[0801] GCATTGATTTTTATAGCCGCATTCGCGAACTGCGCCATTATAGCGATAGCGTGTATGGCGATAC
[0802] CCTGGAAAAACTGACCAACACCGGCCTGTATAACCTGCTGATTCGCTGCCTGCGCTGCCAGAAA
[0803] CCGCTGAACCCGGCGGAAAAACTGCGCCATCTGAACGAAAAACGCCGCTTTCATAACATTGCG
[0804] GGCCATTATCGCGGCCAGTGCCATAGCTGCTGCAACCGCGCGCGCCAGGAACGCCTGCAGCGC
[0805] CGCCGCGAAACCCAGGTGAAAGATGAACTGAAAGATGAACTGGGCAGCGGCGCGACCAACTTT
[0806] AGCCTGCTGAAACAGGCGGGCGATGTGGAAGAAAACCCGGGCCCGATGTATCGCATGCAGCTG
[0807] CTGAGCTGCATTGCGCTGAGCCTGGCGCTGGTGACCAACAGCGCGCCGACCAGCAGCAGCACC
[0808] AAAAAAACCCAGCTGCAGCTGGAACATCTGCTGCTGGATCTGCAGATGATTCTGAACGGCATT
[0809] AACAACTATAAAAACCCGAAACTGACCCGCATGCTGACCTTTAAATTTTATATGCCGAAAAAA
[0810] GCGACCGAACTGAAACATCTGCAGTGCCTGGAAGAAGAACTGAAACCGCTGGAAGAAGTGCTG
[0811] AACCTGGCGCAGAGCAAAAACTTTCATCTGCGCCCGCGCGATCTGATTAGCAACATTAACGTG
[0812] ATTGTGCTGGAACTGAAAGGCAGCGAAACCACCTTTATGTGCGAATATGCGGATGAAACCGCG
[0813] ACCATTGTGGAATTTCTGAACCGCTGGATTACCTTTTGCCAGAGCATTATTAGCACCCTGACCT
[0814] AA
[0815] SEQ ID NO:84
[0816] ATGCGCGTGACCGCGCCGCGCACCGTGCTGCTGCTGCTGAGCGCGGCGCTGGCGCTGACCGAA
[0817] ACCTGGGCGCATGGCGATACCCCGACCCTGCATGAATATATGCTGGATCTGCAGCCGGAAACC
[0818] ACCGATCTGTATTGCTATGAACAGCTGAACGATAGCAGCGAAGAAGAAGATGAAATTGATGGC
[0819] CCGGCGGGCCAGGCGGAACCGGATCGCGCGCATTATAACATTGTGACCTTTTGCTGCAAATGC
[0820] GATAGCACCCTGCGCCTGTGCGTGCAGAGCACCCATGTGGATATTCGCACCCTGGAAGATCTGC
[0821] TGATGGGCACCCTGGGCATTGTGTGCCCGATTTGCAGCCAGAAACCGGGCGGCGGCAGCATGC
[0822] ATCAGAAACGCACCGCGATGTTTCAGGATCCGCAGGAACGCCCGCGCAAACTGCCGCAGCTGT
[0823] GCACCGAACTGCAGACCACCATTCATGATATTATTCTGGAATGCGTGTATTGCAAACAGCAGCT
[0824] GCTGCGCCGCGAAGTGTATGATTTTGCGTTTCGCGATCTGTGCATTGTGTATCGCGATGGCAAC
[0825] CCGTATGCGGTGTGCGATAAATGCCTGAAATTTTATAGCAAAATTAGCGAATATCGCCATTATT
[0826] GCTATAGCCTGTATGGCACCACCCTGGAACAGCAGTATAACAAACCGCTGTGCGATCTGCTGAT
[0827] TCGCTGCATTAACTGCCAGAAACCGCTGTGCCCGGAAGAAAAACAGCGCCATCTGGATAAAAA
[0828] ACAGCGCTTTCATAACATTCGCGGCCGCTGGACCGGCCGCTGCATGAGCTGCTGCCGCAGCAGC
[0829] CGCACCCGCCGCGAAACCCAGCTGGGCGGCGGCGGCAGCGGCGGCGGCGGCAGCGGCGGCGG
[0830] CGGCAGCATGCATGGCCCGAAAGCGACCCTGCAGGATATTGTGCTGCATCTGGAACCGCAGAA
[0831] CGAAATTCCGGTGGATCTGCTGTGCCATGAACAGCTGAGCGATAGCGAAGAAGAAAACGATGA
[0832] AATTGATGGCGTGAACCATCAGCATCTGCCGGCGCGCCGCGCGGAACCGCAGCGCCATACCAT
[0833] GCTGTGCATGTGCTGCAAATGCGAAGCGCGCATTAAACTGGTGGTGGAAAGCAGCGCGGATGA
[0834] TCTGCGCGCGTTTCAGCAGCTGTTTCTGAACACCCTGAGCTTTGTGTGCCCGTGGTGCGCGAGC
[0835] CAGCAGGGCGGCGGCAGCATGGCGCGCTTTGAAGATCCGACCCGCCGCCCGTATAAACTGCCG
[0836] GATCTGTGCACCGAACTGAACACCAGCCTGCAGGATATTGAAATTACCTGCGTGTATTGCAAAA
[0837] CCGTGCTGGAACTGACCGAAGTGTTTGAATTTGCGTTTAAAGATCTGTTTGTGGTGTATCGCGA
[0838] TAGCATTCCGCATGCGGCGTGCCATAAATGCATTGATTTTTATAGCCGCATTCGCGAACTGCGC
[0839] CATTATAGCGATAGCGTGTATGGCGATACCCTGGAAAAACTGACCAACACCGGCCTGTATAAC
[0840] CTGCTGATTCGCTGCCTGCGCTGCCAGAAACCGCTGAACCCGGCGGAAAAACTGCGCCATCTG
[0841] AACGAAAAACGCCGCTTTCATAACATTGCGGGCCATTATCGCGGCCAGTGCCATAGCTGCTGCA
[0842] ACCGCGCGCGCCAGGAACGCCTGCAGCGCCGCCGCGAAACCCAGGTGAGCCAGAGCACCATTC
[0843] CGATTGTGGGCATTGTGGCGGGCCTGGCGGTGCTGGCGGTGGTGGTGATTGGCGCGGTGGTGG
[0844] CGACCGTGATGTGCCGCCGCAAAAGCAGCGGCGGCAAAGGCGGCAGCTATAGCCAGGCGGCG
[0845] AGCAGCGATAGCGCGCAGGGCAGCGATGTGAGCCTGACCGCGGGCAGCGGCGCGACCAACTTT
[0846] AGCCTGCTGAAACAGGCGGGCGATGTGGAAGAAAACCCGGGCCCGATGTATCGCATGCAGCTG
[0847] CTGAGCTGCATTGCGCTGAGCCTGGCGCTGGTGACCAACAGCGCGCCGACCAGCAGCAGCACC
[0848] AAAAAAACCCAGCTGCAGCTGGAACATCTGCTGCTGGATCTGCAGATGATTCTGAACGGCATT
[0849] AACAACTATAAAAACCCGAAACTGACCCGCATGCTGACCTTTAAATTTTATATGCCGAAAAAA
[0850] GCGACCGAACTGAAACATCTGCAGTGCCTGGAAGAAGAACTGAAACCGCTGGAAGAAGTGCTG
[0851] AACCTGGCGCAGAGCAAAAACTTTCATCTGCGCCCGCGCGATCTGATTAGCAACATTAACGTG
[0852] ATTGTGCTGGAACTGAAAGGCAGCGAAACCACCTTTATGTGCGAATATGCGGATGAAACCGCG
[0853] ACCATTGTGGAATTTCTGAACCGCTGGATTACCTTTTGCCAGAGCATTATTAGCACCCTGACCT
[0854] AA
[0855] SEQ ID NO:85
[0856] ATGCGCGTTACAGCTCCCCGAACCGTTCTGCTTCTGCTTAGTGCCGCCCTCGCTCTGACAGAGA
[0857] CTTGGGCGATGCATGGTCCCAAAGCCACACTTCAGGATATCGTCCTGCACCTGGAGCCGCAGA
[0858] ACGAGATTCCAGTTGACCTCCTGTGTCACGAACAGCTCTCTGACTCTGAGGAAGAGAATGACG
[0859] AGATCGATGGCGTGAATCACCAGCACCTGCCCGCCAGGAGGGCCGAACCTCAGAGACATACAA
[0860] TGCTGTGTATGTGTTGCAAGTGTGAAGCCAGGATCAAACTCGTTGTGGAGTCAAGTGCAGACG
[0861] ACCTGCGCGCCTTTCAGCAGCTGTTTCTGAACACCCTGAGTTTCGTATGCCCCTGGTGCGCCTCC
[0862] CAGCAGGGCGGAGGCTCAATGGCGAGGTTTGAAGACCCAACTCGACGGCCTTATAAGCTGCCT
[0863] GACCTGTGCACCGAGCTGAACACCAGTCTGCAGGACATCGAGATAACATGCGTGTATTGTAAG
[0864] ACCGTGCTCGAGTTGACAGAGGTGTTCGAGTTCGCCTTTAAGGATCTCTTTGTGGTTTATCGCG
[0865] ACAGCATACCGCACGCCGCCTGTCACAAGTGTATCGACTTTTATTCCCGCATTAGGGAACTGAG
[0866] GCATTACTCTGACTCCGTTTACGGGGATACTCTGGAGAAGCTGACCAACACCGGTCTGTATAAC
[0867] CTCCTGATAAGGTGTTTGCGATGCCAGAAGCCACTCAATCCTGCCGAGAAACTGAGACATCTCA
[0868] ATGAGAAGAGGCGCTTTCATAACATCGCTGGACACTACCGCGGGCAGTGCCACAGCTGCTGTA
[0869] ATAGAGCCAGACAGGAAAGGTTGCAACGGCGGAGAGAGACCCAGGTGGGCGGCGGTGGCAGT
[0870] GGGGGCGGTGGTTCAGGAGGCGGTGGGAGCCACGGCGACACCCCAACACTTCACGAGTATATG
[0871] CTGGACCTCCAGCCCGAAACCACCGACCTGTATTGCTATGAGCAACTGAACGATTCCTCCGAAG
[0872] AAGAGGATGAGATTGATGGCCCCGCGGGTCAGGCTGAACCAGATAGGGCGCACTACAACATCG
[0873] TCACCTTCTGCTGCAAGTGCGACAGCACTCTCAGACTTTGCGTCCAGTCAACACACGTGGACAT
[0874] CAGAACCCTCGAGGACCTGCTGATGGGCACCCTGGGAATCGTGTGTCCAATTTGCAGCCAAAA
[0875] ACCGGGGGGCGGCTCAATGCATCAGAAGCGAACTGCAATGTTCCAGGACCCCCAGGAACGACC
[0876] AAGAAAACTGCCCCAGCTCTGTACCGAGCTCCAGACCACAATCCACGACATCATCCTCGAGTGT
[0877] GTGTACTGTAAACAGCAACTGCTGAGGCGCGAAGTGTATGACTTCGCCTTTCGGGATCTTTGTA
[0878] TCGTCTACCGCGACGGCAACCCTTATGCCGTGTGCGACAAGTGTCTCAAGTTTTACAGTAAAAT
[0879] CTCCGAGTACAGACACTACTGTTATAGCCTGTACGGAACTACTCTGGAGCAGCAATATAACAA
[0880] GCCGCTGTGTGACCTGCTTATTCGCTGCATTAATTGTCAAAAACCGCTGTGCCCAGAGGAGAAA
[0881] CAGCGCCACCTGGACAAAAAGCAGAGGTTTCATAATATTAGAGGCCGATGGACCGGAAGATGT
[0882] ATGAGCTGTTGTCGATCAAGTCGGACCCGGAGGGAAACTCAGCTGTCTCAGTCCACCATCCCTA
[0883] TCGTTGGTATAGTGGCCGGTCTCGCTGTGTTGGCTGTTGTAGTCATCGGAGCCGTTGTGGCTACC
[0884] GTCATGTGCAGAAGGAAATCTAGCGGTGGAAAAGGGGGTAGCTACAGTCAGGCCGCAAGCTCT
[0885] GATAGCGCTCAAGGATCAGATGTTAGCCTTACCGCCGGATCCGGGGCCACAAACTTCAGCCTTT
[0886] TGAAGCAAGCCGGCGACGTGGAAGAGAATCCAGGCCCTATGTACCGCATGCAGCTGCTCTCCT
[0887] GTATCGCACTGAGCCTGGCACTGGTGACTAATTCAGCCCCAACATCCTCTAGTACTAAAAAGAC
[0888] ACAGCTGCAACTGGAGCACCTGTTGCTCGACCTTCAGATGATCCTGAACGGTATCAACAATTAT
[0889] AAGAACCCTAAGTTGACAAGGATGCTGACCTTTAAGTTCTATATGCCCAAGAAGGCAACAGAA
[0890] CTTAAGCACCTGCAGTGCCTGGAAGAAGAACTCAAGCCTCTGGAGGAGGTCCTCAATCTGGCT
[0891] CAGTCTAAGAACTTCCATCTGAGACCCCGAGACCTTATCTCCAATATCAACGTGATAGTTCTTG
[0892] AGCTGAAAGGTTCCGAGACTACATTTATGTGTGAGTACGCGGACGAAACCGCTACAATAGTAG
[0893] AGTTTCTTAATCGGTGGATCACCTTCTGCCAGTCTATCATCAGCACACTGACTTAA
[0894] SEQ ID NO:86
[0895] ATGCGAGTGACTGCACCTAGAACTGTTCTCCTCCTGCTCTCCGCCGCCCTCGCCCTGACCGAGA
[0896] CATGGGCCCACGGAGACACTCCTACATTGCATGAATATATGCTGGACCTTCAACCAGAGACCA
[0897] CTGACCTCTACTGTTACGAACAACTGAATGATTCCTCCGAAGAGGAGGACGAAATTGATGGGC
[0898] CTGCCGGACAAGCCGAGCCTGACAGGGCCCACTATAATATCGTTACGTTCTGTTGTAAATGTGA
[0899] CAGTACATTGCGGCTTTGCGTCCAGTCTACTCACGTTGACATCAGAACTCTGGAGGATCTTCTG
[0900] ATGGGGACTCTGGGGATCGTATGCCCGATTTGCAGTCAGAAACCAGGCGGTGGGTCCATGCAC
[0901] CAGAAAAGAACAGCCATGTTTCAAGACCCTCAGGAGCGGCCACGCAAATTGCCTCAGCTTTGC
[0902] ACGGAATTGCAGACAACCATCCACGACATAATCCTTGAGTGTGTGTATTGCAAACAGCAGCTCT
[0903] TGCGAAGGGAGGTGTATGATTTTGCATTTAGGGATCTGTGCATCGTATATAGAGACGGAAATCC
[0904] GTATGCCGTCTGTGACAAGTGTCTCAAGTTTTATAGCAAGATCAGCGAGTACAGACATTACTGC
[0905] TACTCACTGTACGGCACCACCCTTGAGCAGCAGTACAACAAACCTCTCTGTGACCTGTTGATCC
[0906] GCTGCATCAACTGCCAGAAGCCACTGTGTCCTGAAGAGAAGCAGAGACACCTGGATAAGAAAC
[0907] AGAGGTTTCACAATATCCGCGGGCGATGGACAGGCCGATGCATGAGCTGTTGTCGGAGCTCTA
[0908] GGACCAGGAGGGAAACCCAGCTTGCCGCTATTCTGGGGCTGGGATTGGTTCTGGGTCTGCTGG
[0909] GTCCATTGGCAATCCTGCTGGGGAGTGGCGCTACCAATTTTTCACTGCTCAAACAGGCCGGGGA
[0910] CGTCGAAGAAAATCCTGGCCCTATGTACAGGATGCAGTTGCTTAGCTGCATCGCTCTTTCACTT
[0911] GCATTGGTTACCAATAGCGCCCCCACAAGTTCATCTACAAAAAAAACGCAACTGCAACTTGAG
[0912] CACCTGCTGCTTGATCTCCAAATGATCCTCAACGGGATCAACAATTACAAGAACCCCAAGCTGA
[0913] CCAGAATGCTGACCTTCAAGTTTTACATGCCAAAAAAGGCTACCGAACTCAAACACCTGCAGT
[0914] GCTTGGAAGAGGAGCTGAAGCCCCTGGAAGAAGTTTTGAACCTGGCACAATCCAAGAATTTTC
[0915] ACCTGAGGCCACGGGATTTGATCTCCAATATCAACGTAATCGTGCTCGAGCTTAAGGGTTCTGA
[0916] GACCACATTCATGTGTGAATATGCAGACGAGACGGCTACCATTGTTGAATTTTTGAATCGGTGG
[0917] ATTACCTTTTGCCAGAGCATTATCTCCACTCTGACGTAA
[0918] SEQ ID NO:87
[0919] ATGCGCGTGACCGCCCCTCGGACCGTTCTCCTCCTCCTCTCTGCAGCCCTGGCGTTGACAGAAA
[0920] CTTGGGCCCATGGGGATACGCCAACACTCCATGAGTATATGCTTGATCTCCAACCCGAAACTAC
[0921] TGACCTCTACTGCTACGAACAGCTGAACGATTCTTCAGAAGAGGAGGATGAGATAGATGGACC
[0922] CGCAGGGCAAGCAGAACCAGATCGGGCGCATTATAACATCGTGACTTTCTGCTGTAAATGCGA
[0923] TTCTACTCTCCGCTTGTGTGTGCAAAGCACCCACGTCGATATAAGAACACTGGAAGACCTTTTG
[0924] ATGGGTACTCTGGGCATCGTGTGTCCTATATGCTCCCAGAAGCCCGGAGGCGGTTCCATGCACC
[0925] AGAAGAGGACAGCTATGTTTCAGGACCCGCAGGAAAGGCCAAGAAAATTGCCCCAGCTGTGTA
[0926] CCGAGCTTCAAACTACTATTCATGACATCATACTCGAGTGCGTGTACTGTAAGCAACAGCTGCT
[0927] GCGACGCGAAGTTTACGACTTCGCTTTCAGAGATCTGTGCATCGTCTACCGAGACGGCAATCCC
[0928] TATGCCGTATGTGATAAATGCCTGAAGTTCTACAGCAAAATCTCTGAATATCGGCATTATTGTT
[0929] ACTCACTGTATGGCACCACTCTGGAACAGCAGTATAATAAGCCGCTTTGCGATCTGTTGATACG
[0930] CTGTATTAATTGCCAAAAGCCCCTGTGTCCAGAGGAAAAACAGCGACACCTCGATAAGAAGCA
[0931] GCGGTTTCACAATATACGCGGACGCTGGACAGGGAGATGTATGTCTTGTTGTCGGTCTAGTAGA
[0932] ACCAGGCGCGAGACACAACTGCTGGTGGCATATGATAACGCTGTGAATTTGTCCTGTAAGTACT
[0933] CATACAACCTGTTTTCTAGAGAATTCAGAGCCAGCCTGCATAAGGGCCTGGATTCAGCCGTAGA
[0934] AGTCTGCGTGGTGTATGGTAATTACAGCCAGCAGCTGCAAGTGTATAGCAAGACTGGCTTCAAC
[0935] TGTGATGGCAAGCTGGGCAACGAGAGCGTAACTTTCTATCTGCAGAACCTGTACGTGAATCAG
[0936] ACTGATATCTATTTCTGCAAGATAGAGGTCATGTACCCACCGCCATACCTGGACAATGAAAAGT
[0937] CCAACGGCACCATTATACACGTTAAGGGGGGCAGCGGCGCTACTAACTTCAGTCTGCTCAAGC
[0938] AGGCAGGTGATGTGGAGGAAAACCCCGGGCCGATGTACAGAATGCAGCTGCTTAGCTGTATCG
[0939] CACTGTCTCTGGCCCTTGTGACCAATTCCGCTCCAACTTCTTCATCCACCAAGAAGACTCAGCTG
[0940] CAGCTGGAACACCTGTTTGTTGGATCTGCAGATGATCCTTAATGGCATTAATAATTATAAGAACC
[0941] CAAAGCTGACCCGCATGCTCACCTTTAAGTTCTACATGCCCAAAAAGGCCACAGAGCTGAAAC
[0942] ATCTTCAGTGCCTTGAAGAAGAACTGAAACCCCTGGAAGAGGTGCTCAACCTGGCCCAGTCCA
[0943] AGAATTTCCACCTGCGCCCTCGAGATCTGATCTCTAATATTAACGTTATTGTGCTCGAGCTGAA
[0944] GGGTAGCGAAACTACCTTCATGTGCGAGTACGCGGACGAGACAGCCACCATTGTGGAATTCCT
[0945] CAACCGCTGGATTACGTTCTGCCAGTCCATTATATCTACCCTGACCTAA
[0946] SEQ ID NO:88
[0947] ATGAGAGTGACTGCCCCAAGGACCGTACTGTTGCTTTTGTCAGCCGCCCTGGCATTGACCGAAA
[0948] CGTGGGCCATGCATGGTCCGAAGGCCACTCTGCAGGATATAGTCTTGCACCTTGAGCCTCAGAA
[0949] CGAAATCCCAGTCGACCTGCTGTGCCATGAACAGCTCTCTGACTCCGAGGAGGAGAACGACGA
[0950] GATTGATGGCGTGAACCACCAGCATCTGCCAGCACGCAGAGCTGAACCACAGCGCCATACCAT
[0951] GTTGTGTATGTGTTGTAAGTGCGAGGCTAGAATCAAACTGGTTGTTGAAAGCTCCGCTGATGAT
[0952] CTGAGGGCATTTCAACAGCTGTTTCTGAACACACTGTCCTTCGTGTGCCCATGGTGCGCATCAC
[0953] AGCAGGGCGGGGGTAGCATGGCCAGGTTCGAAGATCCAACTAGACGACCTTACAAGCTTCCAG
[0954] ACCTGTGCACTGAGCTCAACACATCTTTGCAGGACATCGAGATTACTTGTGTGTACTGCAAAAC
[0955] CGTGCTCGAACTGACTGAAGTTTTTGAGTTCGCATTCAAGGACCTGTTCGTCGTGTACCGCGAC
[0956] TCTATCCCTCATGCCGCTTGCCACAAATGTATAGACTTCTATTCTAGGATTAGAGAGCTCAGAC
[0957] ATTACTCCGACAGCGTATACGGCGACACTCTGGAGAAGCTTACGAACACTGGGCTGTATAATCT
[0958] CCTGATCAGGTGCCTGCGGTGCCAAAAGCCCCTGAACCCCGCTGAAAAGCTGAGGCACCTTAA
[0959] CGAGAAACGGCGGTTCCACAACATTGCTGGCCATTACAGGGGGCAGTGTCACTCATGCTGCAA
[0960] CAGGGCCAGGCAAGAACGGCTGCAGAGGCGCAGAGAGACTCAAGTAGCGGCAATCCTCGGTC
[0961] TCGGCTTGGTCCTTGGACTGCTCGGACCACTTGCGATACTGCTCGGATCCGGCGCCACGAATTT
[0962] TAGCCTTCTGAAACAAGCTGGTGACGTCGAAGAGAACCCTGGACCGATGTATCGCATGCAGTT
[0963] GCTCAGCTGTATTGCCCTTTCACTTGCACTGGTAACTAATAGCGCCCCCACAAGTTCCAGTACA
[0964] AAGAAAACACAACTTCAGCTGGAGCACCTCCTGCTGGATCTTCAGATGATCCTCAACGGCATA
[0965] AATAACTACAAGAATCCCAAGCTCACCCGAATGCTCACCTTCAAATTCTACATGCCTAAGAAGG
[0966] CAACAGAGTTGAAGCACCTGCAATGCCTGGAAGAAGAACTGAAACCTTTGGAGGAGGTGTTGA
[0967] ATTTGGCTCAGTCCAAGAACTTTCACCTTCGGCCACGGGATCTGATATCTAATATCAACGTTAT
[0968] CGTGTTGGAACTTAAGGGTTCAGAGACCACATTCATGTGCGAATATGCCGACGAAACCGCAAC
[0969] CATTGTTGAGTTCTTGAACCGCTGGATCACTTTTTGTCAGTCAATCATTAGCACCCTGACATAASEQID NO:89
[0970] ATGCGCGTTACAGCCCCCCGAACCGTGCTGCTTCTGCTCTCTGCTGCACTTGCACTGACAGAAA
[0971] CATGGGCCATGCATGGTCCCAAAGCTACCCTCCAGGACATCGTTCTGCACCTTGAACCCCAGAA
[0972] TGAGATTCCCGTCGACCTGCTGTGCCACGAACAGCTTAGCGACAGTGAAGAGGAAAACGATGA
[0973] GATAGACGGGGTGAACCACCAGCACCTGCCAGCACGGAGGGCAGAGCCACAACGCCATACCA
[0974] TGCTTTGTATGTGTTGCAAGTGTGAAGCAAGAATTAAGCTGGTTGTAGAATCCTCCGCCGACGA
[0975] CCTCAGGGCGTTCCAACAGCTGTTCCTTAATACTCTGAGCTTCGTTTGTCCATGGTGTGCTAGCC
[0976] AGCAAGGAGGAGGTTCCATGGCTCGCTTTGAGGATCCTACCAGGCGGCCTTACAAGCTTCCCG
[0977] ATCTCTGCACGGAACTTAACACCTCCCTGCAGGACATCGAGATTACATGTGTTTACTGCAAGAC
[0978] GGTGCTGGAGTTGACGGAGGTTTTTGAGTTTGCTTTCAAAGATCTGTTTGTGGTTTACAGGGAC
[0979] AGCATCCCTCATGCAGCTTGTCACAAGTGTATAGACTTTTACAGCAGGATACGGGAACTGCGCC
[0980] ACTACAGCGACTCCGTGTATGGCGACACGCTTGAGAAGTTGACCAACACTGGCCTGTATAACCT
[0981] GCTGATTCGGTGCCTTCGATGCCAGAAACCCCTTAACCCGGCTGAGAAACTGCGACATCTCAAC
[0982] GAGAAACGGCGATTTCATAATATTGCTGGCCACTACAGAGGTCAGTGCCATAGCTGTTGCAAC
[0983] AGAGCGAGGCAAGAGCGCCTTCAGCGGAGACGCGAAACACAAGTGCTGGTCGCATACGATAA
[0984] CGCAGTTAACTTGAGCTGCAAGTACAGTTACAACCTCTTTAGTCGCGAGTTTCGGGCCTCCCTG
[0985] CATAAGGGCCTCGATTCCGCCGTGGAGGTGTGCGTCGTTTACGGGAACTACAGCCAGCAGCTG
[0986] CAGGTCTATAGCAAGACTGGATTCAACTGCGATGGCAAACTGGGTAACGAATCAGTGACTTTCT
[0987] ACCTGCAAAATCTCTATGTCAATCAGACCGATATCTATTTTTGCAAAATCGAGGTGATGTACCC
[0988] TCCGCCATATCTTGACAATGAGAAGTCCAACGGCACAATAATACACGTCAAAGGCGGGTCCGG
[0989] AGCAACAAATTTCTCACTCCTGAAGCAGGCAGGTGACGTAGAGGAGAATCCAGGACCAATGTA
[0990] TCGGATGCAACTGTTGAGCTGCATAGCACTGTCCCTCGCTCTTGTTACGAATAGCGCTCCCACT
[0991] AGTTCTTCAACCAAGAAGACCCAACTTCAACTGGAGCATTTGCTGTTGGACCTGCAGATGATTC
[0992] TCAACGGCATCAACAACTACAAAAACCCTAAACTTACCCGCATGCTTACTTTCAAGTTTTATAT
[0993] GCCTAAGAAGGCTACAGAACTGAAGCATCTGCAGTGTCTGGAGGAGGAATTGAAGCCCCTCGA
[0994] GGAAGTGCTCAACCTGGCTCAGTCAAAAAACTTTCACCTGAGACCCCGGGATCTGATATCTAAC
[0995] ATTAACGTCATCGTCCTGGAACTTAAAGGGTCTGAGACCACCTTCATGTGCGAATATGCAGATG
[0996] AGACCGCCACTATTGTAGAGTTCCTGAACCGGTGGATAACATTCTGTCAGAGCATCATTTCAAC
[0997] ACTGACCTAA
[0998] SEQ ID NO:108
[0999] AGGTCCAACACAACATATACAAAACAAACGAATCTCAAGCAATCAAGCATTCTACTTCTATTGC
[1000] AGCAATTTAAATCATTTCTTTTAAAGCAAAAGCAATTTTCTGAAAATTTTCACCATTTACGAAC
[1001] GATAGCCACCATGAGGGTGACCGCCCCGCGGACCGTTCTGCTGCTGCTGTCCGCCGCCCTGGCT
[1002] CTGACTGAGACTTGGGCACACGGTGACACCCCAACGCTCCATGAGTACATGCTGGACCTCCAG
[1003] CCCGAAACCACGGACTTGTACTGCTACGAGCAGCTGAACGACAGCAGCGAGGAGGAGGACGA
[1004] GATTGATGGGCCCGCCGGTCAGGCCGAGCCTGACCGGGCCCATTATAACATCGTCACCTTCTGC
[1005] TGCAAGTGCGACAGCACTCTTCGGCTGTGCGTGCAGTCAACTCACGTGGACATTCGGACGCTGG
[1006] AGGACCTGCTCATGGGGACGTTGGGGATCGTGTGCCCAATCTGCAGTCAGAAGCCGGGCGGCG
[1007] GCAGCATGCACCAGAAGCGGACCGCGATGTTCCAGGACCCCCAGGAGCGGCCGCGGAAGCTGC
[1008] CGCAGCTGTGCACGGAGCTCCAGACTACTATCCACGATATTATTCTGGAGTGCGTGTACTGCAA
[1009] GCAGCAGCTTCTGCGGCGCGAAGTCTACGACTTCGCCTTCAGGGACTTGTGCATAGTGTATAGG
[1010] GATGGCAATCCCTATGCAGTGTGCGACAAGTGCCTGAAGTTCTACTCCAAGATCTCGGAGTATC
[1011] GCCACTACTGCTACTCCTTGTACGGGACGACCCTGGAGCAGCAGTACAACAAGCCCCTCTGCGA
[1012] CCTCCTGATCCGCTGCATCAATTGCCAGAAGCCGCTGTGTCCCGAGGAGAAGCAGCGGCATCTG
[1013] GACAAGAAGCAGCGGTTCCACAACATCAGGGGTCGCTGGACGGGGCGTTGTATGAGCTGCTGC
[1014] AGGTCGTCCCGTACAAGGAGGGAGACACAGCTGAGCCAGTCCACGATCCCCATCGTGGGGATC
[1015] GTGGCTGGCTTAGCTGTGTTGGCAGTGGTGGTGATTGGAGCTGTGGTCGCTACTGTGATGTGCA
[1016] GGCGGAAGTCGTCCGGGGGCAAGGGAGGGAGCTACAGCCAGGCAGCTTCATCCGATTCGGCCC
[1017] AGGGTTCCGACGTCTCCCTGACTGCTGGCAGCGGGGCGACCAACTTCTCCCTGCTGAAGCAGGC
[1018] AGGGGACGTCGAGGAGAACCCTGGGCCGATGTATCGGATGCAGCTGCTGAGCTGTATCGCTCT
[1019] CTCCCTTGCCCTCGTGACGAATTCCGCCCCCACATCCAGTAGCACCAAGAAGACCCAGCTCCAG
[1020] CTTGAGCACCTCCTCCTGGATCTGCAGATGATCCTGAACGGGATCAACAACTATAAGAATCCGA
[1021] AGCTGACTCGGATGCTTACGTTTAAGTTCTACATGCCTAAGAAGGCAACAGAGCTTAAGCATCT
[1022] GCAGTGCCTGGAGGAGGAGCTCAAGCCTCTGGAGGAGGTGCTGAACCTCGCCCAGAGCAAGAA
[1023] CTTCCACTTGCGTCCGCGGGATCTCATCAGCAACATCAACGTGATCGTCTTGGAGCTGAAGGGC
[1024] TCCGAGACGACCTTCATGTGTGAGTATGCTGATGAGACgGCGACGATAGTGGAGTTCTTGAACC
[1025] GGTGGATCACCTTCTGTCAGAGTATCATCAGTACTCTGACATGATAAGCTGCAGAATTCGTCGA
[1026] CGGATCCGATCTGGTACTGCATGCACGCAATGCTAGCTGCCCCTTTCCCGTCCTGGGTACCCCG
[1027] AGTCTCCCCCGACCTCGGGTCCCAGGTATGCTCCCACCTCCACCTGCCCCACTCACCACCTCTGC
[1028] TAGTTCCAGACACCTCCCAAGCACGCAGCAATGCAGCTCAAAACGCTTAGCCTAGCCACACCC
[1029] CCACGGGAAACAGCAGTGATTAACCTTTAGCAATAAACGAAAGTTTAACTAAGCTATACTAAC
[1030] CCCAGGGTTGGTCAATTTCGTGCCAGCCACACCCTCGAGCTAGCaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa
[1031] gcatatgactaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa
[1032] SEQ ID NO:109
[1033] AGGTCCAACACAACATATACAAAACAAACGAATCTCAAGCAATCAAGCATTCTACTTCTATTGC
[1034] AGCAATTTAAATCATTTCTTTTAAAGCAAAAGCAATTTTCTGAAAATTTTCACCATTTACGAAC
[1035] GATAGCCACCATGGACGCAATGAAGCGGGGGCTCTGCTGTGTCCTCCTCCTCTGCGGGGCAGTG
[1036] TTCGTGTCCCCTCATGGGGACACGCCAACACTGCACGAGTACATGCTCGACCTGCAGCCCGAGA
[1037] CgACCGACCTCTATTGTTATGAACAGTTGAATGACAGCAGCGAGGAGGAGGACGAGATTGATG
[1038] GGCCCGCCGGTCAGGCCGAGCCTGACCGGGCCCATTATAACATCGTCACCTTCTGCTGCAAGTG
[1039] TGATTCAACCCTCCGGCTGTGCGTGCAGTCCACTCATGTGGACATACGCACGCTGGAGGACCTG
[1040] CTCATGGGGACCCTGGGCATCGTGTGTCCGATCTGCTCCCAGAAGCCTGGGGGCGGATCCATGC
[1041] ACCAGAAGCGGACCGCGATGTTCCAGGATCCCCAGGAGCGGCCGCGGAAGCTGCCGCAGCTGT
[1042] GCACGGAGCTCCAGACTACTATCCACGATATTATTCTGGAGTGCGTGTACTGCAAGCAGCAGCT
[1043] TCTGCGGCGCGAAGTCTACGACTTCGCCTTCAGGGACTTGTGCATAGTGTATAGGGATGGCAAT
[1044] CCCTATGCAGTGTGCGACAAGTGCCTGAAGTTCTACAGCAAGATCTCAGAGTACCGGCATTATT
[1045] GCTACTCGCTGTACGGGACCACCCTGGAGCAGCAGTACAATAAGCCACTCTGCGATCTGCTGAT
[1046] CCGCTGCATCAATTGCCAGAAGCCGCTGTGTCCCGAGGAGAAGCAGCGGCATCTGGACAAGAA
[1047] GCAGCGGTTTCATAACATTAGAGGTCGGTGGACTGGGCGCTGCATGTCCTGCTGCAGGAGCAG
[1048] CAGGACACGGCGGGAAACACAGCTCAAGGATGAGCTCAAGGATGAGCTCGGGTCCGGGGCCA
[1049] CGAACTTCAGCCTGCTGAAGCAGGCTGGAGACGTGGAGGAGAACCCCGGACCCAACATTGACC
[1050] GGCCCAAGGGGCTGGCTTTCACCGATGTTGATGTGGACAGCATCAAGATCGCGTGGGAGTCGC
[1051] CCCAGGGCCAGGTCAGCCGTTATCGGGTGACGTACTCGTCACCCGAGGACGGCATCCATGAGC
[1052] TGTTTCCCGCCCCCGACGGGGAGGAGGACACAGCAGAGCTCCAGGGGTTGCGTCCGGGCTCTG
[1053] AGTACACGGTCAGCGTGGTGGCTCTCCATGACGACATGGAGAGCCAGCCGCTGATCGGTACTC
[1054] AGAGCACCTGATAAGCTGCAGAATTCGTCGACGGATCCGATCTGGTACTGCATGCACGCAATG
[1055] CTAGCTGCCCCTTTCCCGTCCTGGGTACCCCGAGTCTCCCCCGACCTCGGGTCCCAGGTATGCTC
[1056] CCACCTCCACCTGCCCCACTCACCACCTCTGCTAGTTCCAGACACCTCCCAAGCACGCAGCAAT
[1057] GCAGCTCAAAACGCTTAGCCTAGCCACACCCCCACGGGAAACAGCAGTGATTAACCTTTAGCA
[1058] ATAAACGAAAGTTTAACTAAGCTATACTAACCCCAGGGTTGGTCAATTTCGTGCCAGCCACACC
[1059] CTCGAGCTAGCaaaaaaaaaaaaaaaaaaaaaaaaaaaagcatatgactaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa
[1060] aaaaaaaaaaaaaaaaaaaaaaaaaaaaa
[1061] mRNA molecule preparation methods
[1062] Methods for preparing and purifying mRNA molecules are known and disclosed in the art. mRNA molecules can be prepared using only in vitro transcription (IVT) enzymes. Methods for preparing IVT polynucleotides are known in the art and described in WO 2013 / 151666, WO 2013 / 151668, etc. Purification methods include purifying RNA transcripts including polyA tails by contacting a sample with a surface linked to multiple thymidines or their derivatives and / or multiple uracils or their derivatives (polyT / U) under conditions that allow RNA transcripts to bind to said surface, and eluting the purified RNA transcripts from said surface (WO2014 / 152031); using ion (e.g., anion) exchange chromatography, which allows for the separation of longer RNAs up to 10,000 nucleotides in length via a scalable method (WO 2014 / 144767); and subjecting the modified mRNA sample to DNase treatment (WO 2014 / 152030).
[1063] In some embodiments, a method for preparing mRNA molecules is provided, comprising (1) transcribing a downstream DNA sequence of a promoter using an RNA polymerase to synthesize mRNA, using linear double-stranded DNA containing a promoter sequence as a template and ATP, GTP, CTP, or N1-Me-pUTP as substrates; and (2) capping the synthesized mRNA using a one-step chemical method with a capping analogue CAP m7Gppp(2'OMeA)pG or CAP5 m7G(5')vppp(5')(2'OMeA)pG. In some embodiments, the promoter is a T7 promoter. In some embodiments, the RNA polymerase is a T7 RNA polymerase.
[1064] During mRNA processing, characteristic structural features of mature mRNA, such as the 5' cap and Poly-A tail, are typically added to the transcribed (immature) mRNA.
[1065] During the in vitro synthesis of mRNA molecules, 5' capping of polynucleotides can be performed simultaneously using chemical RNA cap analogs to produce 5' guanosine cap structures: 5'-guanosine cap structures: 3'-O-Me-m7G(5')ppp(5')G[ARCA cap]; G(5')ppp(5')A; G(5')ppp(5')G; m7G(5')ppp(5')A; m7G(5')ppp(5')G. In some embodiments described herein, the following kits are used for capping: CAP m7Gppp(2'OMeA)pG or CAP5m7G(5')vppp(5')(2'OMeA)pG. 5' capping of mRNA can also be performed post-transcriptionally using vaccinia virus capping enzymes to produce Cap 0 structures. The Cap 1 structure can be generated using both vaccinia virus capping enzyme and 2'-O methyltransferase to produce m7G(5')ppp(5')G 2'O-methyl. The Cap 2 structure can be generated from the Cap 1 structure, followed by 2'-O methylation of the 5' penultimate nucleotide using 2'-O methyltransferase. The Cap 3 structure can be generated from the Cap 2 structure, followed by 2'-O methylation of the 5' penultimate nucleotide using 2'-O methyltransferase. The 3' Poly-A tail is typically an extension of the adenine nucleotide added to the 3' end of the transcribed mRNA. In some embodiments, it can include up to approximately 400 adenine nucleotides.
[1066] Elements of mRNA molecules
[1067] In some implementations, in addition to the read frame encoding the immunogenic protein, the mRNA molecule also contains a 5'UTR and a 3'UTR, as well as a 5' cap structure or a 3' Poly-A tail. The 5'UTR and 3'UTR are typically transcribed from genomic DNA and are elements of immature mRNA.
[1068] When mRNA is engineered to encode immunogenic proteins, it may contain one or more of these untranslated regions (UTRs). Wild-type untranslated regions of nucleic acids are transcribed but not translated. In mRNA, the 5' UTR begins at the transcription start site and continues to the start codon, but does not include the start codon; while the 3' UTR begins immediately after the stop codon and continues until the transcription termination signal. UTRs may play a regulatory role in the stability of nucleic acid molecules and translation. A variety of 5' UTR and 3' UTR sequences are known and available in the art. The 5' UTR is the mRNA region 5' directly upstream of the start codon (the first codon of the mRNA transcript translated by ribosomes). The 5' UTR does not encode proteins (it is non-coding). Native 5' UTRs have features that play a role in translation initiation. They possess features such as the Kozak sequence, which are well known to be involved in the ribosome-initiated translation of many genes. It is also known that 5' UTRs form secondary structures involved in elongation factor binding.
[1069] The 3'UTR is the mRNA region directly downstream (3') of a stop codon (the codon that transmits the translation termination signal in the mRNA transcript). The 3'UTR does not encode proteins (it is non-coding). Strains containing adenosine and uridine are known to be embedded in natural or wild-type 3'UTRs. These AU-rich features are particularly prevalent in genes with high turnover rates. Based on their sequence characteristics and functional properties, AU-rich elements (AREs) can be divided into three categories (Chen et al., 1995): Class I AREs contain several scattered copies of the AUUUA motif within the U-rich region. C-Myc and MyoD contain Class I AREs. Class II AREs have two or more overlapping UUAUUUA(U / A)(U / A) nonmers. Molecules containing this type of ARE include GM-CSF and TNF-α. Class III AREs are less clearly defined. These U-rich regions do not contain the AUUUA motif. c-Jun and myogenin are two well-studied examples of this category. Most proteins that bind to AREs are known to disrupt messenger stability, while members of the ELAV family, particularly HuR, have been shown to increase mRNA stability. HuR binds to all three classes of AREs. Engineering a HuR-specific binding site into the 3'UTR of a nucleic acid molecule will result in HuR binding, thereby stabilizing the messenger in vivo. The introduction, removal, or modification of AU-rich elements (AREs) in the 3'UTR can be used to modulate the stability of polynucleotides (e.g., mRNA). When engineering a particular nucleic acid, one or more copies of an ARE can be introduced to make the nucleic acid of this disclosure less stable, thereby reducing translation and reducing the production of the resulting protein. Similarly, AREs can be identified and removed or mutated to increase intracellular stability, thereby increasing the translation and production of the resulting protein. Those skilled in the art will understand that the 5'UTR can be used with any desired 3'UTR sequence.
[1070] A poly-A tail is an mRNA region containing multiple consecutive adenosine monophosphates (ATPs) located downstream of the 3' UTR, for example, directly downstream (i.e., 3'). A poly-A tail can contain 10 to 300 ATPs. In some embodiments, the poly-A tail contains 10 to 400 ATPs (e.g., 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, or 400). Poly-A tails can be used to protect mRNA from enzymatic degradation, for example, in the cytoplasm, and can facilitate transcription termination and / or the export of mRNA from the nucleus for translation.
[1071] In some implementations, the Poly-A tail contains a polynucleotide sequence of SEQ ID NO:106 or 107.
[1072] SEQ ID NO:106
[1073] aaaaaaaaaaaaaaaaaaaaaaaaaaaaagcatatgactaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa
[1074] SEQ ID NO:107
[1075] aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa.
[1076] In some embodiments, the 5'UTR contains a polynucleotide sequence of SEQ ID NO: 90, 92, 94, 96, 98, 100, 102 or 104 or a variant polynucleotide sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% of SEQ ID NO: 90, 92, 94, 96, 98, 100, 102 or 104.
[1077] In some embodiments, the 3'UTR contains a polynucleotide sequence of SEQ ID NO: 91, 93, 95, 97, 99, 101, 103 or 105 or a variant polynucleotide sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% of SEQ ID NO: 91, 93, 95, 97, 99, 101, 103 or 105.
[1078] In some embodiments, the 5'UTR and 3'UTR are selected from combinations of the following sequences or variant polynucleotide sequences: SEQ ID NO: 90 and 91; SEQ ID NO: 92 and 93; SEQ ID NO: 94 and 95; SEQ ID NO: 96 and 97; SEQ ID NO: 98 and 99; SEQ ID NO: 100 and 101; SEQ ID NO: 102 and 103; and SEQ ID NO: 104 and 105.
[1079] SEQ ID NO:90
[1080] AGGTCCAACACAACATATACAAAACAAACGAATCTCAAGCAATCAAGCATTCTACTTCTATTGCAGCAATTTAAATCATTTCTTTTAAAGCAAAAGCAATTTTCTGAAAATTTTCACCATTTACGAACGATAGCCACC
[1081] SEQ ID NO:91
[1082] GCTGCAGAATTCGTCGACGGATCCGATCTGGTACTGCATGCACGCAATGCTAGCTGCCCCTTTCCCGTCCTGGGTACCCCGAGTCTCCCCCGACCTCGGGTCCCAGGTATGCTCCCACCTCCACCTGCCCCACTCACCACCTCTGCTAGTTCCAGACACCTCCCAAGCACGCAGCAATGCAGCTCAAAACGCTTAGCCTAGCCACACCCCCACGGGAAACAGCAGTGATTAACCTTTAGCAATAAACGAAAGTTTAACTAAGCTATACTAACCCCAGGGTTGGTCAATTTCGTGCCAGCCACACCCTCGAGCTAGC
[1083] SEQ ID NO:92
[1084] GAAAGCTTGGACAGATCGCCTGGAGACGCCATCCACGCTGTTTTGACCTCCATAGAAGACACCGGGACCGATCCAGCCTCCGCGGCCGGGAACGGTGCATTGGAACGCGGATTCCCCGTGCCAAGAGTGACTCACCGTCCTTGACACGGGATCCGCCACC
[1085] SEQ ID NO:93
[1086] GGTACCCGGGTGGCATCCCTGTGACCCCTCCCCAGTGCCTCTCCTGGCCCTGGAAGTTGCCACTCCAGTGCCCACCAGCCTTGTCCTAATAAAATTAAGTTGCATCGGGCCCTATTCTATAG
[1087] SEQ ID NO:94
[1088] AGGCACAGACACCAAGGACAGAGACGCTGGCTAGGCCGCCCTCCCCACTGTTACCAAC
[1089] SEQ ID NO:95AGTGTCCAGACCATTGTCTTCCAACCCCAGCTGGCCTCTAGAACACCCACTGGCCAGTCCTAGAGCTCCTGTCCCTACCCACTCTTTGCTACAATAAATGCTGAATGAATCC
[1090] SEQ ID NO:96
[1091] CTTTTTCGCAACGGGTTTGCCGCCAGAACACAGGTGTCGTGAAAACTACCCCTAAAAGCCAAA
[1092] SEQ ID NO:97
[1093] ATATTATCCCTAATACCTGCCACCCCACTCTTAATCAGTGGTGGAAGAACGGTCTCAGAACTGTTTGTTTCAATTGGCCATTTAAGTTTAGTAGTAAAAGACTGGTTAATGATAACAATGCATCGTAAAACCTTCAGAAGGAAAGGAGAATGTTTTGTGGACCACTTTGGTTTTCTTTTTTGCGTGTGGCAGTTTTAAGTTATTAGTTTTTAAAATCAGTACTTTTTAATGGAAACAACT
[1094] SEQ ID NO:98
[1095] AGGATCCCAAGGCCCAACTCCCCGAACCACTCAGGGTCCTGTGGACAGCTCACCTAGCTGCA
[1096] SEQ ID NO:99
[1097] CTGCCCGGGTGGCATCCCTGTGACCCCTCCCCAGTGCCTCTCCTGGCCCTGGAAGTTGCCACTCCAGTGCCCACCAGCCTTGTCCTAATAAAATTAAGTTGCATCA
[1098] SEQ ID NO:100
[1099] AACTCTATATAGGGAGTTCAACTGGTCACCCAGAGCTGTCCTGTGGCCTCTGCAGCTCAGC
[1100] SEQ ID NO:101
[1101] GGGGCCTTCTGACATGAGTCTGGCCTGGCCCCACCTCCTAGTTCCTCATAATAAAGACAGATTGCTTCTTCGCTTCTCACTGAGGGGCCTTCTGACATGAGTCTGGCCTGGCCCCACCTCCCCAGTTTCTCATAATAAAGACAGATTGCTTCTTCACTTGAATCAAGGGACCT
[1102] SEQ ID NO:102
[1103] ACTCTTCTGGTCCCCACAGACTCAGAGAGAACCCACC
[1104] SEQ ID NO:103
[1105] GCTGGAGCCTCGGTGGCCATGCTTCTTGCCCCTTGGGCCTCCCCCCAGCCCCTCCTCCCCTTCCTGCACCCGTACCCCCGTGGTCTTTGAATAAAGTCTGAGTGGGCGGC
[1106] SEQ ID NO:104
[1107] ACAGAGTAAACTTTTGCTGGGCTCCAAGTGACCGCCCATAGTTTATTATAAAGGTGACTGCACCCTGCAGCCACCAGCACTGCCTGGCTCCACGTGCCTCCTGGTCTCAGT
[1108] SEQ ID NO:105
[1109] CAGGACACAGCCTTGGATCAGGACAGAGACTTGGGGGCCATCCTGCCCCTCCAACCCGACATGTGTACCTCAGCTTTTTCCCTCACTTGCATCAATAAAGCTTCTGTGTTTGGAACAGCTAA
[1110] Formulations or compositions
[1111] Formulations or compositions containing nucleic acids (e.g., mRNA molecules) are known in the art and are described, for example, in WO2013 / 090648. For example, compositions or formulations may be, but are not limited to, nanoparticles, poly(lactic-co-glycolic acid) (PLGA) microspheres, lipids, lipid complexes, liposomes, polymers, carbohydrates (including monosaccharides), cationic lipids, fibrin gels, fibrin hydrogels, fibrin glues, fibrin binders, fibrinogen, thrombin, rapidly eliminating lipid nanoparticles (reLNPs), and combinations thereof.
[1112] vaccine
[1113] This document also provides information on nucleic acid vaccines. Nucleic acid vaccines can be mRNA vaccines. In some embodiments, an mRNA vaccine may contain the same mRNA molecule or multiple different mRNA molecules. In some embodiments, the mRNA in an mRNA vaccine may contain the same read frame of an immunogenic protein. In some embodiments, the mRNA in an mRNA vaccine may contain the same read frame of an immunogenic protein, and the mRNA molecule may be the same. In some embodiments, the mRNA in an mRNA vaccine may contain the same read frame of an immunogenic protein, and the mRNA molecules may be different. That is, an mRNA vaccine contains multiple different mRNA molecules that encode the same immunogenic protein. In some embodiments, the mRNA in an mRNA vaccine may contain different read frames of immunogenic proteins. Nucleic acid vaccines may contain mRNA molecules containing any one or more of SEQ ID NO:57-89, 108 and 109 (e.g. 2-35, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 or 35) of a nucleic acid sequence or a variant sequence.
[1114] In some embodiments, two or more different mRNAs can be formulated into the same lipid nanoparticle. In some embodiments, two or more different RNAs encoding antigens can be formulated into separate lipid nanoparticles. The lipid nanoparticles can then be combined and administered as a single vaccine composition (e.g., containing multiple RNAs encoding multiple antigens), or they can be administered alone.
[1115] Vaccine administration
[1116] In some implementations, the compositions, formulations, or vaccines described herein may be administered to a subject (e.g., a mammalian subject, such as a human subject). The dosage is an effective amount to enable the nucleic acid to be translated in vivo to produce an immunogenic protein. An "effective amount" is based at least in part on the target tissue, target cell type, means of administration, physical characteristics of the RNA (e.g., length, nucleotide composition, and / or degree of nucleoside modification), other components of the vaccine, and other determinants such as the subject's age, weight, height, sex, and general health condition. Typically, an effective amount of vaccine provides an induced or enhanced immune response based on the antigen produced in the subject's cells.
[1117] The term "pharmaceutical composition" refers to a combination of an active agent and an inert or active carrier, making the composition particularly suitable for diagnostic or therapeutic use in vivo or in vitro. A "pharmaceutically acceptable carrier" will not cause undesirable physiological effects when administered to or to a subject. A carrier in a pharmaceutical composition must also be "acceptable" in the sense that it is compatible with and can stabilize the active ingredient. One or more solubilizers may be used as drug carriers to deliver the active agent. Examples of pharmaceutically acceptable carriers include, but are not limited to, biocompatible mediators, adjuvants, additives, and diluents to achieve compositions usable as dosage forms. Other examples of carriers include colloidal silica, magnesium stearate, cellulose, and sodium lauryl sulfate. Further suitable drug carriers and diluents, as well as the pharmaceutical necessities for their use, are described in Remington's Pharmaceutical Sciences.
[1118] In some embodiments, the vaccine described herein can be used to treat or prevent HPV infection. The vaccine can be administered prophylactically or therapeutically to healthy individuals as part of an active immunization program, or during the early stages of infection, either in the incubation period or during active infection after the onset of symptoms. In some embodiments, the vaccine can treat subjects already infected with HPV. In some embodiments, the amount of RNA provided to cells, tissues, or subjects can be an amount effective for immunoprophylaxis or treatment.
[1119] Vaccines may be administered in combination with other prophylactic or therapeutic compounds. As a non-limiting example, the prophylactic or therapeutic compound may be an adjuvant or a booster. As used herein, when referring to a prophylactic composition such as a vaccine, the term "booster" refers to an additional administration of the prophylactic (vaccine) composition. In exemplary embodiments, the time interval between the initial administration of the prophylactic composition and the administration of the booster may be, but is not limited to, 1 week, 2 weeks, 3 weeks, 1 month, 2 months, 3 months, 6 months, or 1 year. In some embodiments, the vaccine may be administered intramuscularly, intranasally, or intradermally.
[1120] In some implementations, the mRNA molecules or vaccines described herein are administered at doses of 5 μg, 10 μg, 20 μg, 30 μg, 40 μg, 50 μg, 60 μg, 70 μg, 80 μg, 90 μg, 100 μg, 200 μg, 300 μg, 400 μg, 500 μg, 600 μg, 700 μg, 800 μg, 900 μg, 1 mg, 2 mg, 3 mg, 4 mg, 5 mg, 6 mg, 7 mg, 8 mg, 9 mg, 10 mg, 20 mg, 30 mg, 40 mg, 50 mg, 60 mg, 70 mg, 80 mg, 90 mg, 100 mg, 200 mg, 300 mg, 400 mg, 500 mg, 600 mg, 700 mg, 800 mg, 900 mg, or more. In some implementations, the mRNA molecules or vaccines described herein are in unit dose form, each unit dose may contain 5 μg, 10 μg, 20 μg, 30 μg, 40 μg, 50 μg, 60 μg, 70 μg, 80 μg, 90 μg, 100 μg, 200 μg, 300 μg, 400 μg, 500 μg, 600 μg, 700 μg, 800 μg, 900 μg, 1 mg, 2 mg, 3 mg, 4 mg, 5 mg, 6 mg, 7 mg, 8 mg, 9 mg, 10 mg, 20 mg, 30 mg, 40 mg, 50 mg, 60 mg, 70 mg, 80 mg, 90 mg, 100 mg, 200 mg, 300 mg, 400 mg, 500 mg, 600 mg, 700 mg, 800 mg, 900 mg or more of the mRNA molecules or vaccines described herein.
[1121] Example
[1122] The invention will be more readily understood by referring to the following examples, which are only used to illustrate certain aspects and embodiments of the invention and are not intended to limit the invention.
[1123] Material
[1124] Unless otherwise stated, all reagents used in this embodiment are commercially available or conventional materials.
[1125] method
[1126] 1. Preparation of HPV mRNA transcripts
[1127] Using linear double-stranded DNA containing the T7 promoter sequence as a template, and ATP, GTP, CTP, and N1-Me-pUTP as substrates, the DNA sequence downstream of the promoter was transcribed by T7 RNA polymerase to efficiently synthesize mRNA. A one-step chemical method was then used to cap the mRNA transcript using different cap analogs, CAP m7GpppAmG or CAP5 m7G(5')vppp(5')(2'OMeA)pG. The mRNA concentration, 260 / 280 ratio, integrity, and capping rate were measured.
[1128] 2. Preparation of HPV mRNA-LNP formulations
[1129] Encapsulation was performed using a fixed lipid formulation. SM-102, DSPC, cholesterol, and DMG-PEG2000 were completely dissolved in ethanol at a molar percentage of (50:10:38.5:1.5) to obtain a lipid mixture. The lipids and mRNA were mixed in a nanomedicine preparation system at a volume ratio of 1:3 (lipids:mRNA) in 20mM sodium citrate buffer (pH 4.0) at a flow rate of 12 ml / min. The collected sample solution was diluted 10-fold in Tris-0.1% NaCl buffer, passed through a 100 kDa Pall ultrafiltration tube, centrifuged, and concentrated to remove ethanol. Finally, the solution was adjusted to a suitable concentration with DPBS buffer for use in the subsequent mRNA vaccine formulation preparation.
[1130] Example 1: HPV E6 / E7 mRNA nucleotide sequence design
[1131] This invention utilizes human hosts to codon-optimize the combination sequence of HPV E6 / E7 antigen and immune-stimulating complex, synthesizing plasmids containing corresponding component amino acid sequences (SEQ ID NO. 1-23), non-coding 5'-UTR and 3'-UTR nucleotide sequences (SEQ ID NO. 90-105), protein-coding amino acid sequences (SEQ ID NO. 24-56), and corresponding protein-coding nucleotide sequences (SEQ ID NO: 57-89). The plasmids were constructed using conventional molecular biology methods (synthesized by Genscript Biotech). Sequence descriptions are shown in Table 1.
[1132] Table 1. Description of SEQ ID NO
[1133]
[1134]
[1135]
[1136]
[1137]
[1138] Example 2: Validation of signal peptides: tPA and SigMHC signal peptides
[1139] In this experiment, C57BL / 6 mice were administered a single dose of different vaccines (as shown in Table 2) via intramuscular injection on day 1, with each mouse receiving a dose of 10 μg. Blood was collected on day 14 post-immunization, and blood cells were separated to detect the vaccine-induced cytokine IFNγ. Specifically, blood cells were added to flow cytometry tubes, and an E7 peptide (sequence RAHYNIVTF, synthesized by GenScript) was used as a stimulant. After 2 hours of stimulation, brefeldin A inhibitor was added, and after 5 hours of incubation, cells were collected and stained for surface CD4 and CD8, as well as intracellular factors. Data were detected and analyzed using flow cytometry.
[1140] The results are shown in Table 2. On day 14, all immunized groups with E7 (HPV-16)-mRNA vaccines containing different signal peptides produced a certain level of E7-specific IFNγ. The T-cell immune response to the E7 (HPV-16) mRNA vaccine was increased to 11.55% and 16.54%, respectively, by the signal peptides tPA and SigMHC. After immune activation, CD4+ T lymphocytes differentiated into TH1 and TH2 cells. IFN-γ is a hallmark cytokine of TH1 cells and can induce the upregulation of MHC-I molecules, assisting CTLs in killing tumor cells.
[1141] As shown in Table 2, compared with IgG, IL-2, and Ig Kappa signal peptides, the signal peptides tPA or SigMHC showed a better effect on enhancing the T-cell immune response to E7 (HPV-16) mRNA vaccines. Therefore, in subsequent examples, the signal peptides tPA or SigMHC were selected as the secretion signal peptides for mRNA.
[1142] Table 2: Validation of different signal peptides: IgG, IL-2, tPA, Ig Kappa, and SigMHC signal peptides
[1143]
[1144] Example 3: UTR Verification: 5'-UTR and 3'-UTR
[1145] In this experiment, C57BL / 6 mice were administered a single dose of different vaccines (as shown in Table 3) via intramuscular injection on day 1, with each mouse receiving a dose of 10 μg. Blood was collected on day 14 post-immunization, and blood cells were separated to detect the vaccine-induced cytokine IFNγ. Specifically, blood cells were added to flow cytometry tubes, and an E7 peptide (sequence RAHYNIVTF, synthesized by GenScript) was used as a stimulant. After 2 hours of stimulation, brefeldin A inhibitor was added, and after 5 hours of incubation, cells were collected and stained for surface CD4 and CD8, as well as intracellular factors. Data were detected and analyzed using flow cytometry.
[1146] The results are shown in Table 3. The tPA-E7 (HPV-16)-mRNA vaccine immunization groups with different UTRs showed varying levels of E7-specific IFNγ production on day 14. Groups 1, 4, 5, 7, and 8, with their 5'-UTR and 3'-UTRs respectively, increased the T-cell immune response to the E7 (HPV-16) vaccine to 16.54%, 9.28%, 10.64%, 17.42%, and 13.45%. Therefore, in subsequent examples, the 5'-UTR and 3'-UTR of group 1 were selected as the 5'-UTR and 3'-UTR of the mRNA.
[1147] Table 3: Validation of 5'-UTR and 3'-UTR
[1148]
[1149] Example 4: Validation of the first stimulus: MITD, LAMP, KDEL, OX40 and CD28
[1150] In this experiment, C57BL / 6 mice were administered a single dose of different vaccines (as shown in Table 4) via intramuscular injection on day 1, with each mouse receiving a dose of 10 μg. Blood was collected on day 14 post-immunization, and blood cells were separated to detect the vaccine-induced cytokine IFNγ. Specifically, blood cells were added to flow cytometry tubes, and an E7 peptide (sequence RAHYNIVTF, synthesized by GenScript) was used as a stimulant. After 2 hours of stimulation, brefeldin A inhibitor was added, and after 5 hours of incubation, cells were collected and stained for surface CD4 and CD8, as well as intracellular factors. Data were detected and analyzed using flow cytometry.
[1151] The results are shown in Table 4. The E7-specific IFNγ production on day 14 differed among the E7 (HPV-16)-mRNA vaccine immunization groups with different first stimuli. For the tPA-E7 (HPV-16)-mRNA vaccine design groups with MITD, LAMP, KDEL, OX40, or CD28 stimuli sequences, the T-cell immune response increased from 11.55% in the group without any stimuli to 20.26%, 24.15%, 20.01%, 18.54%, and 24.02%, respectively.
[1152] Meanwhile, for the SigMHC-E7 (HPV-16)-mRNA vaccine design group incorporating stimulant sequences such as MITD, LAMP, KDEL, OX40, or CD28, the T-cell immune response increased from 16.54% without the stimulant to 26.78%, 30.16%, 31.44%, 25.63%, and 26.48%, respectively. KDEL showed the best enhancement of efficacy. Therefore, KDEL was selected as the first stimulant for the mRNA in subsequent examples.
[1153] Table 4: Validation of the first stimulus (MITD, LAMP, KDEL, OX40, and CD28)
[1154]
[1155] Example 5: Validation of the second stimulus: CRT, EDA, HSP70, IL2 and IL7
[1156] In this experiment, C57BL / 6 mice were administered a single dose of different vaccines (as shown in Table 5) via intramuscular injection on day 1, with each mouse receiving a dose of 10 μg. Blood was collected on day 14 post-immunization, and blood cells were separated to detect the vaccine-induced cytokine IFNγ. Specifically, blood cells were added to flow cytometry tubes, and an E7 peptide (sequence RAHYNIVTF, synthesized by GenScript) was used as a stimulant. After 2 hours of stimulation, brefeldin A inhibitor was added, and after 5 hours of incubation, cells were collected and stained for surface CD4 and CD8, as well as intracellular factors. Data were detected and analyzed using flow cytometry.
[1157] The results are shown in Table 5. The E7-specific IFNγ production varied on day 14 in the E7 (HPV-16)-mRNA vaccine immunization groups with different secondary stimulants. The tPA-E7 (HPV-16)-KDEL-mRNA vaccine design group, which incorporated CRT, EDA, HSP70, IL2, or IL7 as secondary stimulants, showed an increase in T-cell immune response from 24.15% (before addition) to 40.02%, 47.03%, 42.67%, 42.41%, and 39.56%, respectively. As shown in Table 5, the combination of the first stimulant KDEL and the second stimulant EDA exhibited the best enhancement effect. Therefore, in subsequent embodiments, the combination of the first stimulant KDEL and the second stimulant EDA was selected as the mRNA stimulant combination.
[1158] Table 5: Validation of the second stimulus (CRT, EDA, HSP70, IL2, and IL7)
[1159]
[1160] Example 6: Verification of mRNA preparation by capping different cap analogs CAP m7GpppAmG or CAP5m7G(5')vppp(5')(2'OMeA)pG
[1161] Using linear double-stranded DNA containing the T7 promoter sequence as a template (synthesized by GenScript), and ATP, GTP, CTP, and N1-Me-pUTP as substrates, the DNA sequence downstream of the promoter was transcribed using T7 RNA polymerase and cap analogs CAP m7GpppAmG or CAP5 m7G(5')vppp(5')(2'OMeA)pG (20 μL system, see Table 6, 37℃, 2.5 h incubation; after incubation, 2 μL LDNase I was added for 15 min to synthesize mRNA). The results are shown in Table 6. The cap analog CAPm7GpppAmG (structural formula is...) ) or CAP5m7G(5')vppp(5')(2'OMeA)pG (structural formula is All of these methods can achieve an mRNA template capping rate of over 95%. Therefore, in subsequent examples, CAP5m7G(5')vppp(5')(2'OMeA)pG was selected for mRNA template preparation.
[1162] Table 6: Verification of mRNA preparation by capping with different capping analogues CAP m7GpppAmG or CAP5m7G(5')vppp(5')(2'OMeA)pG
[1163]
[1164] Example 7: Validation of different tandem designs for HPV16 subtypes E6 and E7 and HPV18 subtypes E6 and E7
[1165] This experiment used C57BL / 6 mice, and TC-1 cells in the logarithmic growth phase (cell count 1.6 x 10⁻⁶) were harvested. 5 Cells / mouse, inoculation volume 80μl + 80μl of matrix gel, lateral dorsal side) were injected subcutaneously. Three days after tumor inoculation, mice with successful tumor inoculation were screened. On day 4, mice were injected intramuscularly with a single dose of the vaccine (as shown in Table 7). The 5'UTR and 3'UTR used SEQ ID NO:90 and SEQ ID NO:91, respectively. The cap analog used was CAP5 m7G(5')vppp(5')(2'OMeA)pG, at a dose of 10μg per mouse. TC-1 tumor volume was continuously monitored after immunization. The long and short diameters of the tumor tissue were measured using calipers. Tumor volume (mm) was recorded. 3 = 1 / 2 x major axis (mm) x minor axis (mm) 2 The tumor inhibition rate was calculated based on the tumor volume of the control and experimental groups. The formula for calculating the inhibition rate was: Tumor inhibition rate (T / C) = (1-T / C) × 100, where T represents the experimental group and C represents the control group. The results are shown in Table 7. The tumor inhibition rate of the tPA-E7(HPV-16)-mRNA vaccine without the first and second stimulants was 30.6% on day 30, while the tumor inhibition rate with the first and second stimulants was 77.4% on day 30. The tumor inhibition rate of the mRNA vaccine with the E7(HPV-16)-E6(HPV-16) double antigen and the E7(HPV-18)-E6(HPV-18) double antigen was 100% on day 30. Simultaneously, the tumor inhibition rates of the mRNA vaccines with the E7(HPV-16)-E6(HPV-16) double antigen and the E7(HPV-18)-E6(HPV-18) double antigen were also 100%.
[1166] Table 7. Tumor suppression rates after immunization with different designs of HPV16 subtypes E6 and E7 and HPV18 subtypes E6 and E7 in tandem with TC-1 tumor-bearing mice
[1167]
[1168] Example 8: Validation of the tumor elimination rate of TC-1 by different doses of mRNA vaccine in mice
[1169] In this experiment, C57BL / 6 mice were used. TC-1 cells in the logarithmic growth phase (1.6 × 10⁵ cells / mouse, 80 μl + 80 μl of matrix gel, on the lateral dorsal side) were subcutaneously injected. Three days after tumor inoculation, mice with successful tumor inoculation were selected. On the fourth day, mice were injected intramuscularly with a single dose of either tPA-HPV16(E7-E6)-KDEL-2A-EDA or SigMHC-HPV16(E7-E6)-MITD-2A-IL2 vaccine (selected according to Example 6, one group with a 100% tumor inhibition rate, as shown in Table 8). Immunization was performed at low, medium, and high doses, 5 μg, 10 μg, and 20 μg per mouse. TC-1 tumor volume was continuously monitored after immunization. The major and minor diameters of the tumor tissue were measured using calipers. Tumor volume (mm³) = 1 / 2 x major diameter (mm) x minor diameter (mm) 2 The tumor inhibition rate was calculated based on the tumor volume of the control and experimental groups. The formula for calculating the inhibition rate was: Tumor inhibition rate (T / C) = (1-T / C) × 100, where T represents the experimental group and C represents the control group. The results are shown in Table 8. The tPA-E7(HPV-16)-E6(HPV-16)-KDEL-2A-CRT vaccine showed good tumor inhibition effects on TC-1 tumors at different doses. The tumor inhibition rate in the 5μg low-dose group was 88.4%, with one mouse experiencing tumor recurrence and the remaining tumors completely eliminated. The tumor inhibition rate and mouse survival rate in the 10μg and 20μg dose groups were both 100%.
[1170] It should be understood that although the invention has been described in conjunction with its detailed description, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the appended claims. Other aspects, advantages, and modifications are within the scope of the appended claims.
[1171]
Claims
1. A fusion protein comprising one or more of the following papillomavirus proteins or immunogenic fragments thereof: HPV16-E6, HPV16-E7, HPV18-E6, and HPV18-E7, wherein: the HPV16-E7 protein comprises the amino acid sequence of SEQ ID NO: 1 or an amino acid sequence that is at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 1; the HPV16-E6 protein comprises the amino acid sequence of SEQ ID NO: 2 or an amino acid sequence that is at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 2; the HPV18-E7 protein comprises the amino acid sequence of SEQ ID NO: 3 or an amino acid sequence that is at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 3; the HPV18-E6 protein comprises the amino acid sequence of SEQ ID NO: 4 or an amino acid sequence that is at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 4; preferably, the fusion protein further comprises a stimulator, which is any one or more of MITD, LAMP, CRT, IL-2, OX40, IL7, HSP70, KDEL, CD28, ICOS, 4-1BBL, IL1, HSP40, and EDA; preferably, the fusion protein further comprises a secretion signal peptide, which is an IgG signal peptide, an IL-2 signal peptide, a tPA signal peptide, an Ig kappa signal peptide, or a SigMHC signal peptide; preferably, when the fusion protein comprises multiple papillomavirus proteins or immunogenic fragments thereof, each papillomavirus protein or immunogenic fragment thereof is directly linked or linked via a linker; preferably, the fusion protein comprises a first HPV protein or immunogenic fragment thereof - optionally a second HPV protein or immunogenic fragment thereof - a first stimulator - optionally a linker - optionally a second stimulator, the first HPV protein being HPV16-E6, HPV16-E7, or a fusion thereof and the second HPV protein being HPV18-E6, HPV18-E7, or a fusion thereof, or the first HPV protein being HPV18-E6, HPV18-E7, or a fusion thereof and the second HPV protein being HPV16-E6, HPV16-E7, or a fusion thereof; the first and second stimulator are each independently any one or more of MITD, LAMP, CRT, IL-2, OX40, IL7, HSP70, KDEL, CD28, ICOS, 4-1BBL, IL1, HSP40, and EDA; Preferably, the LAMP stimulator comprises the amino acid sequence of SEQ ID NO: 10 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 10; the MITD stimulator comprises the amino acid sequence of SEQ ID NO: 11 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 11; the KDEL stimulator comprises the amino acid sequence of SEQ ID NO: 12 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 12; the OX40 stimulator comprises the amino acid sequence of SEQ ID NO: 13 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 13; the CD28 stimulator comprises the amino acid sequence of SEQ ID NO: 14 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 14; the ICOS stimulator comprises the amino acid sequence of SEQ ID NO: 15 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 15; the 4-1BBL stimulator comprises the amino acid sequence of SEQ ID NO: 16 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 16; the CRT stimulator comprises the amino acid sequence of SEQ ID NO: 17 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 17; the EDA stimulator comprises the amino acid sequence of SEQ ID NO: 18 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 18; The HSP70 stimulator comprises the amino acid sequence of SEQ ID NO: 19 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 19; The IL1 stimulator comprises the amino acid sequence of SEQ ID NO: 20 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 20; The IL-2 stimulator comprises the amino acid sequence of SEQ ID NO: 21 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 21; The IL7 stimulator comprises the amino acid sequence of SEQ ID NO: 22 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 22; The HSP40 stimulator comprises the amino acid sequence of SEQ ID NO: 23 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 23; Preferably, wherein the IgG signal peptide comprises the amino acid sequence of SEQ ID NO: 5 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 5; The IL-2 signal peptide comprises the amino acid sequence of SEQ ID NO: 6 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 6; The tPA signal peptide comprises the amino acid sequence of SEQ ID NO: 7 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 7; The Ig kappa signal peptide comprises the amino acid sequence of SEQ ID NO: 8 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 8; The SigMHC signal peptide comprises the amino acid sequence of SEQ ID NO: 9 or an amino acid sequence that has at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO:
9.
2. An isolated nucleic acid encoding the fusion protein according to claim 1 ; Preferably, the nucleic acid is a single- or double-stranded DNA molecule or an RNA molecule; Preferably, the nucleic acid comprises a nucleic acid sequence encoding a first HPV protein or an immunogenic fragment thereof - optionally a nucleic acid sequence encoding a second HPV protein or an immunogenic fragment thereof - a nucleic acid sequence encoding a first stimulator - optionally a linker - optionally a nucleic acid sequence encoding a second stimulator, the first HPV protein being HPV16-E6, HPV16-E7 or a fusion thereof and the second HPV protein being HPV18-E6, HPV18-E7 or a fusion thereof, or the first HPV protein being HPV18-E6, HPV18-E7 or a fusion thereof and the second HPV protein being HPV16-E6, HPV16-E7 or a fusion thereof, the linker being a 2A peptide, preferably a P2A peptide, E2A peptide or F2A peptide; Preferably, the nucleic acid further comprises a 5’UTR and a 3’UTR; Preferably, wherein the 5’UTR comprises a polynucleotide sequence of SEQ ID NO: 90, 92, 94, 96, 98, 100, 102 or 104 or a variant polynucleotide sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% to SEQ ID NO: 90, 92, 94, 96, 98, 100, 102 or 104, the 3’UTR comprises a polynucleotide sequence of SEQ ID NO: 91, 93, 95, 97, 99, 101, 103 or 105 or a variant polynucleotide sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% to SEQ ID NO: 91, 93, 95, 97, 99, 101, 103 or 105; Preferably, the 5’UTR and the 3’UTR are selected from the following sequences or combinations of variant polynucleotide sequences thereof: SEQ ID NO: 90 and 91; SEQ ID NO: 92 and 93; SEQ ID NO: 94 and 95; SEQ ID NO: 96 and 97; SEQ ID NO: 98 and 99; SEQ ID NO: 100 and 101; SEQ ID NO: 102 and 103; and SEQ ID NO: 104 and 105; Preferably, the nucleic acid is an mRNA molecule and further comprises a 5’ cap and a Poly-A tail; Preferably, wherein the mRNA molecule comprises a 5’ cap - a 5’UTR - a nucleic acid sequence encoding a secretion signal peptide - a nucleic acid sequence encoding a first HPV protein or an immunogenic fragment thereof - optionally a nucleic acid sequence encoding a second HPV protein or an immunogenic fragment thereof - a nucleic acid sequence encoding a first stimulator - optionally a linker - optionally a nucleic acid sequence encoding a second stimulator - a 3’UTR - a Poly-A tail, the linker being a 2A peptide, preferably a P2A peptide, E2A peptide or F2A peptide; preferably wherein the 5' UTR is SEQ ID NO: 90, 92, 94, 96, 98, 100, 102, or 104 or a variant polynucleotide sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 90, 92, 94, 96, 98, 100, 102, or 104, 3' UTR is SEQ ID NO: 91, 93, 95, 97, 99, 101, 103, or 105 or a variant polynucleotide sequence having at least 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% to SEQ ID NO: 91, 93, 95, 97, 99, 101, 103, or 105; preferably the secretion signal peptide is tPA or SigMHC; preferably the first stimulator is KDEL and the second stimulator is EDA; preferably wherein the structure of the 5' cap is a compound of Formula (I), or a pharmaceutically acceptable salt, stereoisomer, tautomer, or isotopic variant thereof: wherein, --- is a single bond or is absent, X1is selected from O, S, CH2, CH2CH2, CH=CH, CH=CHO, CH2O, OCH2, CH2CH2O, OCH2CH2, a three-membered ring alkyl, R1, R2, R3and R4are each independently halogen, OH, unsubstituted or by O-C 1-3 alkyl-substituted O-C 1-3 alkyl, unsubstituted or by O-C 1-3 alkyl-substituted O-C 1-3 alkyl, B1and B2are each independently selected from a natural, modified, or non-natural nucleobase, preferably the compound of Formula (I) is any one of: preferably wherein the Poly-A tail comprises a polynucleotide sequence of SEQ ID NO: 106 or 107; preferably the nucleic acid is an mRNA molecule comprising a nucleic acid sequence of any one or more of SEQ ID NOs: 57-89, 108, and 109; preferably wherein the RNA is a modified RNA, wherein a uracil, cytosine, adenine, or guanine nucleotide contains a modification group; preferably wherein the modification group is selected from at least one of pseudouridine, N1-methylpseudouridine, N1-ethylpseudouridine, 5-methylcytosine, 5-methoxy cytosine, N1-methylcytosine, 2-thiouridine, 5-methoxyuridine, or N1-methyladenosine, N1-methylguanine, N1-methylguanine, isoguanine. preferably wherein the mRNA molecule is a modified mRNA, the modifications comprising conversion of uracil nucleosides to pseudouridine, N1-methylpseudouridine, N1-ethylpseudouridine, 2-thiouridine, 5-methoxyuridine; and / or, conversion of cytosine nucleosides to 5-methylcytosine, 5-methoxy cytosine, N1-methylcytosine; and / or, conversion of adenine nucleosides to N1-methyladenosine; and / or, conversion of adenine nucleosides to N1-methylguanine, N1-methylguanine, isoguanine.
3. A vector comprising the nucleic acid according to claim 2.
4. A composition comprising one or more nucleic acids according to claim 2 or comprising the vector of claim 3. Preferably, the composition comprises an mRNA molecule of the nucleic acid sequence of any one or more of SEQ ID NOs: 57-89, 108 and 109; Preferably, the composition further comprises a pharmaceutically acceptable excipient and / or adjuvant; Preferably, the composition further comprises a liposomal nanoparticle; Preferably, wherein the liposomal nanoparticle comprises ionizable lipid, DSPC, cholesterol and DMG-PEG2000 ethanol; Preferably, wherein the composition is a vaccine, preferably an mRNA vaccine.
5. A method for preparing the nucleic acid according to claim 2, comprising one or more of the following steps: (1) transcribing a DNA sequence downstream of a promoter by an RNA polymerase with ATP, GTP, CTP, N1-Me-pUTP as substrates using a linear double-stranded DNA containing the promoter sequence as a template to synthesize mRNA; and (2) capping the synthesized mRNA with a cap analog CAP m7Gppp(2’OMeA)pG or CAP5 m7G(5’)vppp(5’)(2’OMeA)pG using a chemical one-step method; Preferably, wherein the promoter is a T7 promoter, and / or the RNA polymerase is a T7 RNA polymerase.
6. Use of the nucleic acid according to claim 2 as an mRNA vaccine.
7. Use of the nucleic acid according to claim 2 for the manufacture of a medicament for the treatment or prevention of HPV infection or a disease associated with HPV infection; Preferably, wherein the HPV is one or more of HPV16 and HPV18; Preferably, wherein the disease associated with HPV infection is one or more of a tumor or cancer associated with HPV infection, common warts, genital warts, respiratory warts, plantar warts, sublingual warts, perlingual warts, or cutaneous and mucosal lesions of flat warts; Preferably, wherein the cancer is one or more of cervical cancer, oropharyngeal cancer, vulvar cancer, anal cancer, penile cancer, vaginal cancer, rectal cancer, squamous cell carcinoma, adenocarcinoma, and head and neck cancer.
Citation Information
Patent Citations
Ventilated filter and smoke dispersing mouthpiece
IL71156
Modified nucleoside, nucleotide, and nucleic acid compositions
WO2013090648A1
Modified polynucleotides for the production of biologics and proteins associated with human disease
WO2013151666A2
Modified polynucleotides for the production of secreted proteins
WO2013151668A2
Ion exchange purification of mRNA
WO2014144767A1