Production of heterologous polypeptides in microalgae, microalgal extracellular bodies and compositions, and methods of making and uses thereof

Culturing recombinant microalgae cells of Schizochytrium or Thraustochytrium with encoded viral proteins produces viral proteins and extracellular bodies, addressing the need for efficient heterologous polypeptide expression in microalgae, offering rapid and cost-effective production with post-translational processing.

JP2025114812AInactive Publication Date: 2025-08-05SANOFI VACCINE TECH SAS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025081042
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2010-11-12
Filing Date
2025-05-14
Publication Date
2025-08-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

There is a need for efficient and cost-effective methods to express heterologous polypeptides, such as viral proteins, in microalgae for therapeutic applications, leveraging the advantages of microalgae over traditional production systems like mammalian and microbial cells.

Method used

Culturing recombinant microalgae cells, particularly of the genus Schizochytrium or Thraustochytrium, with nucleic acid molecules encoding viral proteins, allowing secretion or accumulation of proteins like HA, NA, F, G, E, gp120, gp41, and matrix proteins, and producing microalgal extracellular bodies with heterologous polypeptides discontinuous from the plasma membrane.

Benefits of technology

This method enables high-yield production of viral proteins with post-translational processing and secretion, utilizing microalgae's rapid growth and low-cost fermentation, providing a safe and efficient alternative for protein production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025114812000005
    Figure 2025114812000005
  • Figure 2025114812000006
    Figure 2025114812000006
  • Figure 2025114812000007
    Figure 2025114812000007
Patent Text Reader

Abstract

To provide a method for producing viral hemagglutinin (HA) protein using microalgae.SOLUTION: Provided is a method for producing a full length influenza HA protein as a vaccine using a recombinant microalgal cell, where the microalgal cell comprises a nucleic acid molecule comprising a polynucleotide sequence encoding a viral HA protein including a HA membrane domain as well as active regulatory control elements, the microalgal cell is a Schizochytrium or a Thraustochytrid, and the recombinant microalgal cell is cultured in a culture medium into which the HA protein is secreted. Also provided is a method for formulating the recombinant viral protein as a vaccine composition, where the protein is an influenza protein.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE INVENTION The present invention relates to recombinant microalgal cells and their use in the production of heterologous polypeptides, methods for the production of heterologous polypeptides in microalgal extracellular bodies, microalgal extracellular bodies comprising heterologous polypeptides, and compositions comprising them. [Background technology]

[0002] Advances in biotechnology and molecular biology have made it possible to produce proteins in microbial, plant, and animal cells, many of which were previously available only by extraction from human and other animal tissues, blood, or urine. Today, commercially available biologics are typically produced in mammalian cells, such as Chinese hamster ovary (CHO) cells, or in microbial cells, such as yeast or Escherichia coli (E. coli) cell lines.

[0003] Protein production by microbial fermentation offers several advantages over existing systems such as plant and animal cell culture. For example, microbial fermentation-based processes can provide: (i) rapid production of high concentrations of protein; (ii) the ability to use sterile, well-controlled production conditions (such as Good Manufacturing Practice (GMP) conditions); (iii) the ability to use simple, chemically defined growth media that allow for simpler fermentations and fewer impurities; (iv) the absence of contaminating human or animal pathogens; and (v) ease of protein recovery (e.g., by isolation from the fermentation medium). Furthermore, fermentation facilities are typically less expensive to build than cell culture facilities.

[0004] Microalgae, such as thraustochytrids of the Labyrinthulomycota phylum, can be grown in standard fermentation equipment with very short culture cycles (e.g., 1–5 days), inexpensive defined media, and minimal, if any, purification. Furthermore, certain microalgae, such as Schizochytrium, have a proven safety history for food applications of both biomass and derived lipids. For example, a DHA-rich triglyceride oil derived from this microorganism has received GRAS (Recognized as Safe) status from the U.S. Food and Drug Administration.

[0005] Microalgae have been shown to be capable of expressing recombinant proteins. For example, U.S. Patent No. 7,001,772 (Patent Document 1) disclosed the first recombinant constructs suitable for transforming thraustochytrids, including members of the genus Schizochytrium. This publication disclosed, among other things, Schizochytrium nucleic acid and amino acid sequences for acetolactate synthase, acetolactate synthase promoter and terminator regions, α-tubulin promoters, promoters from polyketide synthase (PKS) systems, and fatty acid desaturase promoters. U.S. Patent Application Publication Nos. 2006 / 0275904 (Patent Document 2) and 2006 / 0286650 (Patent Document 3), both of which are incorporated herein by reference in their entireties, subsequently disclosed Schizochytrium sequences for actin, elongation factor 1 alpha (ef1α), and glyceraldehyde-3-phosphate dehydrogenase (gapdh) promoters and terminators.

[0006] There is a continuing need for methods for expressing heterologous polypeptides in microalgae and for identifying alternative compositions for therapeutic applications. [Prior art documents] [Patent documents]

[0007] [Patent Document 1] U.S. Patent No. 7,001,772 [Patent Document 2] U.S. Patent Application Publication No. 2006 / 0275904 [Patent Document 3] U.S. Patent Application Publication No. 2006 / 0286650 Summary of the Invention

[0008] The present invention is directed to a method for producing a viral protein selected from the group consisting of hemagglutinin (HA) protein, neuraminidase (NA) protein, fusion (F) protein, glycoprotein (G) protein, envelope (E) protein, 120 kDa glycoprotein (gp120), 41 kDa glycoprotein (gp41), matrix protein, and combinations thereof, comprising culturing recombinant microalgae cells in a medium, wherein the recombinant microalgae cells contain a nucleic acid molecule comprising a polynucleotide sequence encoding the viral protein to produce the viral protein. In some embodiments, the viral protein is secreted. In some embodiments, the viral protein is recovered from the medium. In some embodiments, the viral protein accumulates in the microalgae cells. In some embodiments, the viral protein accumulates in the membrane of the microalgae cells. In some embodiments, the viral protein is an HA protein. In some embodiments, the HA protein is at least 90% identical to SEQ ID NO: 77. In some embodiments, the microalgae cell is capable of post-translational processing of the HA protein to produce HA1 and HA2 polypeptides in the absence of exogenous proteases. In some embodiments, the viral protein is an NA protein. In some embodiments, the NA protein is at least 90% identical to SEQ ID NO: 100. In some embodiments, the viral protein is an F protein. In some embodiments, the F protein is at least 90% identical to SEQ ID NO: 102. In some embodiments, the viral protein is a G protein. In some embodiments, the G protein is at least 90% identical to SEQ ID NO: 103. In some embodiments, the microalgae cell is a member of the order Thraustochytriales. In some embodiments, the microalgae cell is of the genus Schizochytrium or Thraustochytrium. In some embodiments, the polynucleotide sequence encoding the viral protein further comprises an HA membrane domain.In some embodiments, the nucleic acid molecule further comprises a polynucleotide sequence selected from the group consisting of SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 38, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, and combinations thereof.

[0009] The present invention is directed to isolated viral proteins produced by any of the above methods.

[0010] The present invention is directed to a recombinant microalgae cell comprising a nucleic acid molecule comprising a polynucleotide sequence encoding a viral protein selected from the group consisting of a hemagglutinin (HA) protein, a neuraminidase (NA) protein, a fusion (F) protein, a glycoprotein (G) protein, an envelope (E) protein, a 120 kDa glycoprotein (gp120), a 41 kDa glycoprotein (gp41), a matrix protein, and combinations thereof. In some embodiments, the viral protein is an HA protein. In some embodiments, the HA protein is at least 90% identical to SEQ ID NO: 77. In some embodiments, the microalgae cell is capable of post-translational processing of the HA protein to produce HA1 and HA2 polypeptides in the absence of exogenous proteases. In some embodiments, the viral protein is an NA protein. In some embodiments, the NA protein is at least 90% identical to SEQ ID NO: 100. In some embodiments, the viral protein is an F protein. In some embodiments, the F protein is at least 90% identical to SEQ ID NO: 102. In some embodiments, the viral protein is a G protein. In some embodiments, the G protein is at least 90% identical to SEQ ID NO: 103. In some embodiments, the microalgae cell is a member of the order Thraustochytriales. In some embodiments, the microalgae cell is of the genus Schizochytrium or Thraustochytrium. In some embodiments, the polynucleotide sequence encoding the viral protein further comprises an HA membrane domain. In some embodiments, the nucleic acid molecule further comprises a polynucleotide sequence selected from the group consisting of SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 38, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, and combinations thereof.

[0011] The present invention is directed to a method for producing a microalgal extracellular body comprising a heterologous polypeptide, the method comprising: (a) expressing in a microalgal host cell a heterologous polypeptide comprising a membrane domain; and (b) culturing the microalgal host cell under conditions sufficient to produce an extracellular body comprising the heterologous polypeptide that is discontinuous with the plasma membrane of the host cell.

[0012] The present invention is directed to methods for producing a composition comprising a microalgal extracellular body and a heterologous polypeptide, the method comprising: (a) expressing in a microalgal host cell a heterologous polypeptide comprising a membrane domain; and (b) culturing the microalgal host cell under conditions sufficient to produce extracellular bodies comprising the heterologous polypeptide that are discontinuous with the plasma membrane of the host cell, wherein the composition is produced as a culture supernatant comprising the extracellular bodies. In some embodiments, the method further comprises removing the culture supernatant from the composition and resuspending the extracellular bodies in an aqueous liquid carrier. The present invention is directed to compositions produced by the method.

[0013] In some embodiments, the methods of producing microalgal extracellular bodies and heterologous polypeptides, or the methods of producing compositions comprising microalgal extracellular bodies and heterologous polypeptides, comprise a host cell that is a Labyrinthulomycota host cell, hi some embodiments, the host cell is a Schizochytrium or Thraustochytrium host cell.

[0014] The present invention is directed to microalgal extracellular bodies comprising a heterologous polypeptide discontinuous with the plasma membrane of a microalgal cell. In some embodiments, the extracellular bodies are vesicles, micelles, membrane fragments, membrane aggregates, or mixtures thereof. In some embodiments, the extracellular bodies are mixtures of vesicles and membrane fragments. In some embodiments, the extracellular bodies are vesicles. In some embodiments, the heterologous polypeptide comprises a membrane domain. In some embodiments, the heterologous polypeptide is a glycoprotein. In some embodiments, the glycoprotein comprises high-mannose oligosaccharides. In some embodiments, the glycoprotein is substantially free of sialic acid.

[0015] The present invention is directed to a composition comprising any of the above-claimed extracellular bodies and an aqueous liquid carrier. In some embodiments, the aqueous liquid carrier is a culture supernatant. [1] A method for producing a viral protein selected from the group consisting of hemagglutinin (HA) protein, neuraminidase (NA) protein, fusion (F) protein, glycoprotein (G) protein, envelope (E) protein, 120 kDa glycoprotein (gp120), 41 kDa glycoprotein (gp41), matrix protein, and combinations thereof, comprising culturing recombinant microalgae cells in a medium, wherein the recombinant microalgae cells contain a nucleic acid molecule comprising a polynucleotide sequence encoding the viral protein for producing the viral protein. [2] The method described in [1], wherein the viral protein is secreted. [3] The method described in [1] or [2], further comprising the step of recovering viral proteins from the culture medium. [4] The method described in [1], wherein the viral protein accumulates in the microalgae cells. [5] The method according to any one of [1] to [4], wherein the viral protein accumulates in the membrane of the microalgae cell. [6] The method according to any one of [1] to [5], wherein the viral protein is an HA protein. [7] The method described in [6], wherein the HA protein is at least 90% identical to SEQ ID NO: 77. [8] The method of [6] or [7], wherein the microalgae cells are capable of post-translational processing of HA protein to produce HA1 and HA2 fragments in the absence of exogenous proteases. [9] The method according to any one of [1] to [5], wherein the viral protein is an NA protein.

[10] The method described in [9], wherein the NA protein is at least 90% identical to SEQ ID NO: 100.

[11] The method according to any one of [1] to [5], wherein the viral protein is an F protein.

[12] The method of

[11] , wherein the F protein is at least 90% identical to SEQ ID NO: 102.

[13] The method according to any one of [1] to [5], wherein the viral protein is a G protein.

[14] The method of

[13] , wherein the G protein is at least 90% identical to SEQ ID NO: 103.

[15] The method according to any one of [1] to

[14] , wherein the microalgae cells are members of the order Thraustochytriales.

[16] The method according to any one of [1] to

[15] , wherein the microalgae cells are cells of the genus Schizochytrium or Thraustochytrium.

[17] The method described in any one of [1] to

[16] , wherein the polynucleotide sequence encoding the viral protein further comprises an HA membrane domain.

[18] The method of any one of [1] to

[17] , wherein the nucleic acid molecule further comprises a polynucleotide sequence selected from the group consisting of SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 38, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, and combinations thereof.

[19] An isolated viral protein produced by the method described in any one of [1] to

[18] .

[20] A recombinant microalgae cell comprising a nucleic acid molecule comprising a polynucleotide sequence encoding a viral protein selected from the group consisting of a hemagglutinin (HA) protein, a neuraminidase (NA) protein, a fusion (F) protein, a glycoprotein (G) protein, an envelope (E) protein, a 120 kDa glycoprotein (gp120), a 41 kDa glycoprotein (gp41), a matrix protein, and combinations thereof.

[21] A method of expressing a heterologous polypeptide comprising a membrane domain in a microalgae host cell; culturing microalgal host cells under conditions sufficient to produce extracellular bodies comprising the heterologous polypeptide, wherein the extracellular bodies are discontinuous with the plasma membrane of the host cells; 1. A method for producing microalgal extracellular bodies comprising a heterologous polypeptide, comprising:

[22] Expressing a heterologous polypeptide comprising a membrane domain in a microalgae host cell; culturing microalgae host cells under conditions sufficient to produce microalgae extracellular bodies comprising a heterologous polypeptide, wherein the extracellular bodies are discontinuous with the plasma membrane of the host cells; 1. A method for producing a composition comprising a microalgal extracellular body and a heterologous polypeptide, comprising: The method, wherein the composition is produced as a culture supernatant containing extracellular bodies.

[23] The method of

[22] , further comprising the steps of removing the culture supernatant from the composition and resuspending the extracellular bodies in an aqueous liquid carrier.

[24] The method according to any one of

[21] to

[23] , wherein the host cell is a Labyrinthulomycota host cell.

[25] The method according to any one of

[21] to

[24] , wherein the host cell is a host cell of the genus Schizochytrium or Thraustochytrium.

[26] An extracellular body produced by the method described in

[21] .

[27] A composition produced by the method described in

[22] .

[28] A microalgal extracellular body comprising a heterologous polypeptide, the microalgal extracellular body being discontinuous with the plasma membrane of the microalgal cell.

[29] The extracellular body according to

[28] , which is a vesicle, a micelle, a membrane fragment, a membrane aggregate, or a mixture thereof.

[30] An extracellular body according to

[28] or

[29] , which is a mixture of vesicles and membrane fragments.

[31] An extracellular body according to

[28] or

[29] , which is a vesicle.

[32] An extracellular body according to any one of

[28] to

[31] , wherein the heterologous polypeptide comprises a membrane domain.

[33] The extracellular body according to any one of

[28] to

[32] , wherein the heterologous polypeptide is a glycoprotein.

[34] The extracellular body of

[33] , wherein the glycoprotein comprises high-mannose oligosaccharides.

[35] The extracellular body according to

[33] or

[34] , wherein the glycoprotein is substantially free of sialic acid.

[36] A composition comprising the extracellular body according to any one of

[28] to

[35] and an aqueous liquid carrier.

[37] The composition described in

[36] , wherein the aqueous liquid carrier is a culture supernatant. [Brief explanation of the drawings]

[0016] [Figure 1] 1 shows the polynucleotide sequence (SEQ ID NO: 76) encoding the hemagglutinin (HA) protein of influenza A virus (A / Puerto Rico / 8 / 34 / Mount Sinai (H1N1)) codon-optimized for expression in Schizochytrium sp. ATCC 20888. [Figure 2] The plasmid map of pCL0143 is shown. [Figure 3] The procedure used to analyze the CL0143-9 clone is shown. [Figure 4] Figure 4 shows secretion of HA protein by transgenic Schizochytrium CL0143-9 ("E"). Figure 4A: Recombinant HA protein (as indicated by arrows) recovered from low-speed supernatants (i.e., cell-free supernatants ("CFS")) of cultures at various temperatures (25°C, 27°C, 29°C) and pHs (5.5, 6.0, 6.5, 7.0) is shown in an anti-H1N1 immunoblot. Figure 4B: Recombinant HA protein recovered from the 60% sucrose fraction under non-reducing or reducing conditions is shown in a Coomassie-stained gel ("Coomassie") and anti-H1N1 immunoblot ("IB: anti-H1N1"). [Figure 5]Figure 5 shows the hemagglutination activity of recombinant HA protein from transgenic Schizochytrium sp. CL0143-9 ("E"). Figure 5A: Hemagglutination activity in cell-free supernatant ("CFS"). Figure 5B: Hemagglutination activity in soluble ("US") and insoluble ("UP") fractions. "[Protein]" refers to the concentration of protein, which decreases from left to right with increasing dilution of the sample. "-" refers to the negative control lacking HA. "+" refers to the positive control for influenza hemagglutinin. "C" refers to the negative control wild-type strain of Schizochytrium sp. ATCC 20888. "HAU" refers to hemagglutination units based on fold dilutions of the sample from left to right. The "2" refers to a 2-fold dilution of the sample in the first well; subsequent wells from left to right correspond to 2-fold dilutions relative to the previous well, resulting in 2, 4, 8, 16, 32, 64, 128, 256, 512, 1024, 2048, and 4096 fold dilutions from the first to last well from left to right. [Figure 6] Figure 6 shows the expression and hemagglutination activity of HA protein present in the 60% sucrose fraction for transgenic Schizochytrium sp. CL0143-9 ("E"). Figure 6A shows that the recombinant HA protein (as indicated by the arrow) recovered from the 60% sucrose fraction is shown in a Coomassie-stained gel ("Coomassie") and anti-H1N1 immunoblot ("IB: anti-H1N1"). Figure 6B shows the corresponding hemagglutination activity. "-" indicates a negative control lacking HA. "+" indicates a positive control with influenza HA protein. "C" indicates a negative control wild-type strain of Schizochytrium sp. ATCC 20888. "HAU" indicates hemagglutination activity units based on 2-fold dilution of the sample. [Figure 7] Peptide sequence analysis of the recovered recombinant HA protein was performed, identifying a total of 27 peptides (amino acids associated with the peptides are highlighted in bold), which covers 42% of the entire HA protein sequence (SEQ ID NO: 77). The HA1 polypeptide was identified by a total of 17 peptides, and the HA2 polypeptide was identified by a total of 9 peptides. [Figure 8]A Coomassie-stained gel ("Coomassie") and anti-H1N1 immunoblot ("IB: anti-H1N1") are shown illustrating HA protein glycosylation in Schizochytrium. "EndoH" and "PNGase F" refer to enzyme treatment of 60% sucrose fraction of transgenic Schizochytrium CL0143-9 with the respective enzymes. "NT" refers to transgenic Schizochytrium CL0143-9 incubated without enzymes but under the same conditions as EndoH and PNGase F treatment. [Figure 9] ATCC 20888 culture supernatant protein (g / L) over time (hours). [Figure 10] SDS-PAGE of total Schizochytrium sp. ATCC 20888 culture supernatant proteins in lanes 11-15 is shown, where supernatants were collected at five of the six time points shown in Figure 9 between 37 and 68 hours, excluding 52 hours. Bands identified (by mass spectrometry peptide sequencing) as actin and gelsolin are marked with arrows. Lane 11 was loaded with 2.4 μg of total protein; the remaining wells were loaded with 5 μg of total protein. [Figure 11] Negatively stained vesicles from Schizochytrium sp. ATCC 20888 ("C: 20888") and transgenic Schizochytrium sp. CL0143-9 ("E: CL0143-9") are shown. [Figure 12] Anti-H1N1 immunogold-labeled vesicles from Schizochytrium species ATCC 20888 ("C: 20888") and transgenic Schizochytrium sp. CL0143-9 ("E: CL0143-9") are shown. [Figure 13]The following shows putative signal-anchor sequences specific to the genus Schizochytrium based on the use of the SignalP algorithm. See, for example, Bendsten et al., J. Mol. Biol. 340: 783-795(2004); Nielsen, H. and Krogh, A. Proc. Int. Conf. Intell. Syst. Mol. Biol. 6: 122-130(1998); Nielsen, H., et al., Protein Engineering 12: 3-9(1999); Emanuelsson, O. et al., Nature Protocols 2: 953-971(2007). [Figure 14] 1 shows a putative type I membrane protein in Schizochytrium based on a BLAST search of genomic and EST DNA Schizochytrium databases for genes with homology to known type I membrane proteins from other organisms and with a transmembrane domain in the extreme C-terminal region of the protein. The putative transmembrane domain is shown in bold. [Figure 15] The plasmid map of pCL0120 is shown. [Figure 16] 1 shows a codon usage table for the genus Schizochytrium. [Figure 17] The plasmid map of pCL0130 is shown. [Figure 18] The plasmid map of pCL0131 is shown. [Figure 19] The plasmid map of pCL0121 is shown. [Figure 20] The plasmid map of pCL0122 is shown. [Figure 21] 1 shows the polynucleotide sequence (SEQ ID NO: 92) encoding the Piromyces sp. E2 xylose isomerase protein "XylA," optimized for expression in Schizochytrium sp. ATCC 20888, corresponding to GenBank accession number CAB76571. [Figure 22]1 shows the polynucleotide sequence (SEQ ID NO: 93) encoding the Piromyces sp. E2 xylulose kinase protein "XylB," optimized for expression in Schizochytrium sp. ATCC 20888, corresponding to GenBank accession number AJ249910. [Figure 23] The plasmid map of pCL0132 is shown. [Figure 24] The plasmid map of pCL0136 is shown. [Figure 25] Figure 25A: shows the plasmid map of pCL0140, and Figure 25B: shows the plasmid map of pCL0149. [Figure 26] Figure 26A: Shows the polynucleotide sequence (SEQ ID NO: 100) encoding the neuraminidase (NA) protein of influenza A virus (A / Puerto Rico / 8 / 34 / Mount Sinai (H1N1)), optimized for expression in Schizochytrium species ATCC 20888. Figure 26B: Shows the polynucleotide sequence (SEQ ID NO: 101) encoding the NA protein of influenza A virus (A / Puerto Rico / 8 / 34 / Mount Sinai (H1N1)), followed by a V5 tag and a polyhistidine tag, optimized for expression in Schizochytrium species ATCC 20888. [Figure 27] A scheme of the procedure used to analyze the CL0140 and CL0149 clones is shown. [Figure 28]Neuraminidase activity of recombinant NA from transgenic Schizochytrium strains CL0140-16, -17, -20, -21, -22, -23, -24, -26, and -28 is shown. Activity was determined by measuring the fluorescence of 4-methylumbelliferone produced after hydrolysis of 4-methylumbelliferyl-α-DN-acetylneuraminic acid (4-MUNANA) by sialidase (excitation (Exc): 365 nm, emission (Em): 450 nm). Activity is expressed as relative fluorescence units (RFU) per μg of protein in concentrated cell-free supernatants (cCFS, leftmost bar for each clone) and cell-free extracts (CFE, rightmost bar for each clone). A wild-type strain of Schizochytrium sp. ATCC 20888 ("-") and a PCR-negative strain of Schizochytrium transformed with pCL0140 ("27"), grown and prepared identically to the transgenic strains, were used as negative controls. [Figure 29] Figure 29 shows partial purification of recombinant NA protein from transgenic Schizochytrium strain CL0140-26. Neuraminidase activity of various fractions is shown in Figure 29A. "cCFS" refers to concentrated cell-free supernatant. "D" refers to cCFS diluted with wash buffer, "FT" refers to flow-through fraction, "W" refers to wash, "E" refers to eluate, and "cE" refers to concentrated eluate fraction. Coomassie-stained gels ("Coomassie") of 12.5 μL of each fraction are shown in Figure 29B. Arrows indicate the band identified as NA protein. Proteins were separated by SDS-PAGE on a NuPAGE® Novex® 12% Bis-Tris gel with MOPS SDS running buffer. [Figure 30] Peptide sequence analysis for the recovered recombinant NA protein is shown, identifying a total of 9 peptides (highlighted in bold red) that cover 25% of the protein sequence (SEQ ID NO: 100). [Figure 31]Figure 31A shows the neuraminidase activity of transgenic Schizochytrium strains CL0149-10, -11, and -12, as well as the corresponding Coomassie-stained gel ("Coomassie") and anti-V5 immunoblot ("Immunoblot: anti-V5"). Figure 31B shows neuraminidase activity (Exc: 365 nm, Em: 450 nm) determined by measuring the fluorescence of 4-methylumbelliferone produced after hydrolysis of 4-MUNANA by sialidase. Activity is expressed as relative fluorescence units (RFU) per μg of protein in the cell-free supernatant (CFS). A wild-type strain of Schizochytrium sp. ATCC 20888 ("-"), grown and prepared identically to the transgenic strains, was used as a negative control. Figure 31B: Coomassie-stained gel and corresponding anti-V5 immunoblot for 12.5 μL of CFS for three transgenic Schizochytrium sp. CL0149 strains ("10," "11," and "12"). Positope™ antibody control protein was used as a positive control ("+"). A wild-type strain of Schizochytrium sp. ATCC 20888 ("-"), grown and prepared identically to the transgenic strains, was used as a negative control. [Figure 32] Figure 32A shows the enzymatic activity of influenza HA and NA in cell-free supernatants of transgenic Schizochytrium cotransformed with CL0140 and CL0143. Data are presented for clones CL0140-143-1, -3, -7, -13, -14, -15, -16, -17, -18, -19, and -20. Figure 32A shows neuraminidase activity (Exc: 365 nm, Em: 450 nm) determined by measuring the fluorescence of 4-methylumbelliferone produced after hydrolysis of 4-MUNANA by sialidase. Activity is expressed as relative fluorescence units (RFU) in 25 μL of CFS. A wild-type strain of Schizochytrium sp. ATCC 20888 ("-"), grown and prepared identically to the transgenic strains, was used as a negative control. Figure 32B shows hemagglutination activity. "-" indicates a negative control lacking HA. "+" refers to influenza HA positive control. "HAU" refers to hemagglutination units based on fold dilution of sample. DETAILED DESCRIPTION OF THE INVENTION

[0017] Detailed Description of the Invention The present invention is directed to methods for producing heterologous polypeptides in microalgae host cells. The invention also is directed to heterologous polypeptides produced by the methods, to microalgae cells comprising the heterologous polypeptides, and to compositions comprising the heterologous polypeptides. The invention is also directed to the production in microalgae host cells of heterologous polypeptides associated with microalgae extracellular bodies that are discontinuous with the plasma membrane of the host cell. The invention is also directed to the production of microalgae extracellular bodies comprising the heterologous polypeptides, and to the production of compositions comprising the same. The invention further is directed to microalgae extracellular bodies comprising the heterologous polypeptides, compositions, and uses thereof.

[0018] microalgae host cells Microalgae, also known as microalgae, are often found in freshwater and marine systems. Microalgae are unicellular, but can also grow in chains and groups. Individual cells range in size from a few micrometers to hundreds of micrometers. They can grow in aqueous suspension, allowing them efficient access to nutrients and the aqueous environment.

[0019] In some embodiments, the microalgal host cell is a heterokont or a stramenopiles.

[0020] In some embodiments, the microalgae host cell is a member of the phylum Labyrinthulomycota. In some embodiments, the Labyrinthulomycota host cell is a member of the order Thraustochytridiomycota or Labyrinthulales. In accordance with the present invention, the term "thraustochytrid" refers to any member of the order Thraustochytridiomycota, including the family Thraustochytriaceae, and the term "labyrinthulid" refers to any member of the order Labyrinthulales, including the family Labyrinthulaceae. Members of the Labyrinthula family were once considered to be members of the order Thraustochytridiomycota, but due to more recent revisions to the taxonomy of such organisms, Labyrinthulales are now considered to be members of the order Labyrinthulales. Both Labyrinthulales and Thraustochytridiomycota are considered to be members of the phylum Labyrinthulomycota. Taxonomic theorists now generally place both of these groups of microorganisms within the stramenopiles lineage of algae or algae-like protists. The current taxonomic position of thraustochytrids and labyrinthulidae can be summarized as follows: Kingdom: Stramenopile (Chromista) Phylum: Labyrinthulomycota (Heterokontophytes) Class: Labyrinthulomycetes, Labyrinthulae Order: Labyrinthula Family: Labyrinthulidae Order: Thraustochytrids Family: Thraustochytrididae

[0021] For purposes of the present invention, strains described as Thraustochytrids include the following organisms: Order: Thraustochytriales; Family: Thraustochytrididae; Genus: Thraustochytrid (species: arudimentale, aureum, benthicola, globosum, kinnei, motivum, multirudimentale, pachydermum, proliferum, roseum, striatum), Ulkenia (species: Amoeboidea, Kerguelensis, Minuta, Profunda, Radiata, Sailens, Sarkariana, Schizochytrops, Visurgensis, Yorkensis), Schizochytrium (species: Aggregatum, Limnaceum, Mangrovei, Minutum, Octosporum), Japonochytrium (species: Marinem), Aplanochytrium (species: haliotidis, kerguelensis, profunda, stocchinoi), Althornia (species: crouchii), or Elina (species: marisalba, sinorifica). For the purposes of the present invention, the species described within the genus Ulkenia are considered to be members of the genus Thraustochytrids.Aurantiochytrium, Oblongichytrium, Botryochytrium, Parietichytrium, and Sicyoidochytrium are further genera encompassed by the Labyrinthulomycota in the present invention.

[0022] Strains referred to in the present invention as Labyrinthulidae organisms include the following organisms: Order: Labyrinthulidae, Family: Labyrinthulidae, Genus: Labyrinthula (Species: algeriensis, coenocystis, chattonii, macrocystis, macrocystis atlantica, macrocystis macrocystis, marina, minuta, roscoffensis, valkanovii, vitellina, vitellina pacifica, vitellina vitellina, zopfii), Labyrinthuloides (species: haliotidis, yorkensis), Labyrinthomyxa (species: marina), Diplophrys (species: archeri), Pyrrhosorus (species: marinus), Sorodiplophrys (species: stercorea) or Chlamydomyxa (species: Labyrinthuloides, montana) (however, there is currently no agreement on the exact taxonomic status of the genera Pilosorus, Sorodiplophrys or Chlamydomyxa).

[0023] Microalgal cells of the Labyrinthulomycota phylum included the deposited strains PTA-10212, PTA-10213, PTA-10214, PTA-10215, PTA-9695, PTA-9696, PTA-9697, PTA-9698, PTA-10208, PTA-10209, PTA-10210, PTA-10211, and the microorganism deposited as SAM2179 (named "Ulkenia SAM2179" by the depositor), any Thraustochytrid species (U. bisulgensis, U. amoevoida (U. amoeboida), U. sarkariana, U. profunda, U. radiata, U. minuta, and former Ulkenia species such as Ulkenia species BP-5601), and Thraustochytrium striatum, Thraustochytrium aureum, Thraustochytrium roseum, etc.; as well as any Japonochytrium species. Strains of the order Thraustochytridales include, but are not limited to, Thraustochytrium sp. (23B) (ATCC 20891); Thraustochytrium striatum (Schneider) (ATCC 24473); Thraustochytrium aureum (Goldstein) (ATCC 34304); Thraustochytrium roseum (Goldstein) (ATCC 28210); and Japonochytrium sp. (L1) (ATCC 28207). The genus Schizochytrium includes, but is not limited to, Schizochytrium aggregatum, Schizochytrium limacinum, Schizochytrium sp. (S31) (ATCC 20888), Schizochytrium sp. (S8) (ATCC 20889), Schizochytrium sp. (LC-RM) (ATCC 18915), Schizochytrium sp. (SR21), deposited strain ATCC 28209, and deposited Schizochytrium limacinum strain IFO 32693. In some embodiments, the cell is of the genus Schizochytrium or Thraustochytrium.Schizochytrium can replicate both by successive binary fission and by forming sporangia that eventually release zoospores, however, Thraustochytrids replicate only by forming sporangia that then release zoospores.

[0024] In some embodiments, the microalgae host cells are Labyrinthulomycetes (also called Labyrinthulomycetes). Labyrinthulomycetes produce unique structures called "ectoplasmic nets." These structures are branched, tubular projections of the plasma membrane that contribute significantly to increasing the surface area of the plasma membrane. See, e.g., Perkins, Arch. Mikrobiol. 84:95-118 (1972); Perkins, Can. J. Bot. 51:485-491 (1973). The ectoplasmic nets are formed from unique cellular structures called sagenosomes or bothrosomes. The ectoplasmic nets allow Labyrinthulomycete cells to attach to and penetrate surfaces. See, e.g., Coleman and Vestal, Can. J. Microbiol. 33:841-843 (1987) and Porter, Mycologia 84:298-299 (1992), respectively. For example, Schizochytrium species ATCC 20888 has been observed to produce ectoplasmic nets that extend into the agar when grown on solid medium (data not shown). The ectoplasmic nets in such cases appear to function as pseudorhizoids. Furthermore, actin filaments have been found to be abundant within certain ectoplasmic net membrane protrusions. See, e.g., Preston, J. Eukaryot. Microbiol. 52:461-475 (2005). Based on the importance of actin filaments within cytoskeletal structures in other organisms, it is expected that cytoskeletal elements such as actin are involved in the formation and / or integrity of ectoplasmic net membrane protrusions.

[0025] Further organisms that produce pseudorhizoids include those called chytrid fungi, which are taxonomically classified in various groups, including Chytridiomycota or Phycomyces. Exemplary genera include Chytridium, Chytrimyces, Cladochytium, Lacustromyces, Rhizophydium, Rhisophyctidaceae, Rozella, Olpidium, and Lobulomyces.

[0026] In some embodiments, the microalgae host cell comprises a membrane protrusion. In some embodiments, the microalgae host cell comprises a pseudorhizoid. In some embodiments, the microalgae host cell comprises an ectoplasmic net. In some embodiments, the microalgae host cell comprises a sagenosome or a bothrosome.

[0027] In some embodiments, the microalgae host cell is a Thraustochytrid. In some embodiments, the microalgae host cell is a Schizochytrium or Thraustochytrid cell.

[0028] In some embodiments, the microalgal host cell is a Labyrinthulid.

[0029] In some embodiments, the microalgae host cell is a eukaryote capable of processing polypeptides through the conventional secretory pathway, such as members of the Labyrinthulomycota, including Schizochytrium, Thraustochytrium, and other thraustochytrids. For example, it has been recognized that members of the Labyrinthulomycota produce fewer secreted proteins in large quantities than CHO cells, providing an advantage to using Schizochytrium over, for example, CHO cells. Furthermore, unlike E. coli, members of the Labyrinthulomycota, such as Schizochytrium, perform protein glycosylation, such as N-linked glycosylation, which is required for the biological activity of certain proteins. It has been determined that the N-linked glycosylation exhibited by thraustochytrids, such as Schizochytrium, more closely resembles mammalian glycosylation patterns than yeast glycosylation.

[0030] Effective culture conditions for the host cells of the present invention include, but are not limited to, effective media, bioreactors, temperature, pH, and oxygen conditions that allow for protein production and / or recombination. Effective media refers to any medium in which Thraustochytriales cells, e.g., microalgae cells such as Schizochytrium host cells, are typically cultured. Such media typically include aqueous media having assimilable carbon, nitrogen, and phosphate sources, as well as appropriate salts, minerals, metals, and other nutrients, e.g., vitamins. Non-limiting examples of suitable media and culture conditions are disclosed in the Examples section. Non-limiting culture conditions suitable for Thraustochytriales microorganisms are also described in U.S. Patent No. 5,340,742, which is incorporated herein by reference in its entirety. The cells of the present invention can be cultured in conventional fermentation bioreactors, shake flasks, test tubes, microtiter dishes, and Petri dishes. Culture can be performed at temperatures, pH, and oxygen contents suitable for the recombinant cells.

[0031] In some embodiments, the microalgal host cells of the present invention comprise a recombinant vector including a nucleic acid sequence encoding a selectable marker. In some embodiments, the selectable marker allows for selection of transformed microorganisms. Examples of dominant selectable markers include enzymes that degrade compounds with antibiotic or antifungal activity, such as the Sh ble gene from Streptoalloteichus hindustanus, which encodes the "bleomycin-binding protein" represented by SEQ ID NO: 5. Another example of a dominant selectable marker includes a thraustochytrid acetolactate synthase sequence, such as a variant of the polynucleotide sequence of SEQ ID NO: 6. The acetolactate synthase may be modified, mutated, or otherwise selected for resistance to inhibition by sulfonylurea compounds, imidazolinone class inhibitors, and / or pyrimidinyloxybenzoic acid. Representative examples of thraustochytrid acetolactate synthase sequences include, but are not limited to, amino acid sequences such as SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, or amino acid sequences that differ from SEQ ID NO: 7 by an amino acid deletion, insertion, or substitution at one or more of the following positions: 116G, 117A, 192P, 200A, 251K, 358M, 383D, 592V, 595W, or 599F, and polynucleotide sequences such as SEQ ID NO: 11, SEQ ID NO: 12, or SEQ ID NO: 13, as well as sequences having at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to any of the representative sequences. Further examples of selectable markers that may be included in recombinant vectors for transformation of microalgal cells include ZEOCIN™, paromomycin, hygromycin, blasticidin, or any other suitable resistance marker.

[0032] The term "transformation" refers to any method by which an exogenous nucleic acid molecule (i.e., a recombinant nucleic acid molecule) can be inserted into a microbial cell. In microbial systems, the term "transformation" is used to describe a heritable change resulting from the acquisition of an exogenous nucleic acid by a microorganism and is essentially synonymous with the term "transfection." Suitable transformation techniques for introducing an exogenous nucleic acid molecule into a microalgae host cell include, but are not limited to, particle bombardment, electroporation, microinjection, lipofection, adsorption, infection, and protoplast fusion. For example, an exogenous nucleic acid molecule, including a recombinant vector, can be introduced into microalgae during the exponential growth phase or in the stationary phase when the microalgae reach an optical density of 1.5-2 at 600 nm. Microalgae host cells can also be pretreated with an enzyme having protease activity prior to the introduction of a nucleic acid molecule into the host cell by electroporation.

[0033] In some embodiments, host cells can be genetically modified to introduce or delete genes involved in biosynthetic pathways associated with carbohydrate transport and / or synthesis, including those involved in glycosylation. For example, host cells can be modified by deleting endogenous glycosylation genes and / or inserting human or animal glycosylation genes to enable glycosylation patterns that more closely resemble those of humans. Modifications of glycosylation in yeast can be found, for example, in U.S. Patent No. 7,029,872, and U.S. Patent Application Publication Nos. 2004 / 0171826, 2004 / 0230042, 2006 / 0257399, 2006 / 0029604, and 2006 / 0040353. Host cells of the present invention also include cells in which RNA viral elements are utilized to increase or regulate gene expression.

[0034] Expression system Expression systems used for expression of heterologous polypeptides in microalgal host cells include regulatory control elements active in microalgal cells. In some embodiments, the expression system includes regulatory control elements active in Labyrinthulomycota cells. In some embodiments, the expression system includes regulatory control elements active in Thraustochytrids. In some embodiments, the expression system includes regulatory control elements active in Schizochytrium or Thraustochytrids. Many regulatory control elements, including various promoters, are active in several different species. Therefore, regulatory sequences may be utilized in the same cell type as the cell from which they were isolated, or in a cell type different from the cell from which they were isolated. The design and construction of such expression cassettes utilizes standard molecular biology techniques known to those skilled in the art. For example, see Sambrook et al., 2001, Molecular Cloning: A Laboratory Manual, 3 rd Please refer to the edition.

[0035] In some embodiments, the expression system used to produce a heterologous polypeptide in a microalgae cell comprises a regulatory element derived from a Labyrinthulomycota sequence. In some embodiments, the expression system used to produce a heterologous polypeptide in a microalgae cell comprises a regulatory element derived from a non-Labyrinthulomycota sequence, including a sequence derived from an algal sequence other than Labyrinthulomycota. In some embodiments, the expression system comprises a polynucleotide sequence encoding a heterologous polypeptide, the polynucleotide sequence being associated with any promoter sequence, any terminator sequence, and / or any other regulatory sequence functional in a microalgae host cell. Inducible or constitutively active sequences can be used. Suitable regulatory control elements also include any of the regulatory control elements associated with the nucleic acid molecules described herein.

[0036] The present invention is also directed to an expression cassette for expressing a heterologous polypeptide in a microalgal host cell. The present invention is also directed to any of the host cells described above comprising an expression cassette for expressing a heterologous polypeptide in the host cell. In some embodiments, the expression system comprises an expression cassette including genetic elements, such as at least a promoter, a coding sequence, and a terminator region, operably linked to function in the host cell. In some embodiments, the expression cassette comprises at least one of the isolated nucleic acid molecules of the present invention described herein. In some embodiments, all of the genetic elements of the expression cassette are sequences associated with the isolated nucleic acid molecule. In some embodiments, the control sequence is an inducible sequence. In some embodiments, the nucleic acid sequence encoding the heterologous polypeptide is integrated into the genome of the host cell. In some embodiments, the nucleic acid sequence encoding the heterologous polypeptide is stably integrated into the genome of the host cell.

[0037] In some embodiments, the isolated nucleic acid sequence that encodes the heterologous polypeptide to be expressed is operably linked to a promoter sequence and / or terminator sequence, both of which are functional in host cells.The promoter and / or terminator sequence that the isolated nucleic acid sequence that encodes the heterologous polypeptide to be expressed is operably linked to can comprise any promoter and / or terminator sequence, including but not limited to the nucleic acid sequence disclosed herein, the regulatory sequence disclosed in U.S. Patent No. 7,001,772, the regulatory sequence disclosed in U.S. Patent Application Publication No. 2006 / 0275904 and U.S. Patent Application Publication No. 2006 / 0286650, the regulatory sequence disclosed in U.S. Patent Application Publication No. 2010 / 0233760 and International Publication No. 2010 / 107709, or other regulatory sequence that is operably linked to the isolated polynucleotide sequence that encodes the heterologous polypeptide and is functional in the host cell that is transformed. In some embodiments, the nucleic acid sequence encoding the heterologous polypeptide is codon-optimized for the particular microalgae host cell to maximize translation efficiency.

[0038] The present invention also relates to recombinant vectors comprising the expression cassettes of the present invention. Recombinant vectors include, but are not limited to, plasmids, phages, and viruses. In some embodiments, the recombinant vector is a linearized vector. In some embodiments, the recombinant vector is an expression vector. As used herein, the phrase "expression vector" refers to a vector suitable for producing an encoded product (e.g., a protein of interest). In some embodiments, a nucleic acid sequence encoding a product to be produced is inserted into a recombinant vector to produce a recombinant nucleic acid molecule. A nucleic acid sequence encoding a heterologous polypeptide to be produced is inserted into a vector such that the nucleic acid sequence is operably linked to a regulatory sequence in the vector (e.g., a Thraustochytriales promoter) that enables the transcription and translation of the nucleic acid sequence in the recombinant microorganism. In some embodiments, a selectable marker, including any of the selectable markers described herein, allows for the selection of recombinant microorganisms in which a recombinant nucleic acid molecule of the present invention has been successfully produced.

[0039] In some embodiments, heterologous polypeptides produced by the host cells of the invention are produced on a commercial scale, including production of heterologous polypeptides from microorganisms grown in aerated fermentors of sizes ≥ 100 L, ≥ 1,000 L, ≥ 10,000 L, or ≥ 100,000 L. In some embodiments, commercial scale production is performed in aerated fermentors with stirring.

[0040] In some embodiments, heterologous polypeptides produced by a host cell of the invention may accumulate intracellularly or may be secreted from the cell into the medium, e.g., as a soluble heterologous polypeptide.

[0041] In some embodiments, the heterologous polypeptide produced by the present invention is recovered from the cells, from the medium in which the cells are grown, or from the fermentation medium. In some embodiments, the heterologous polypeptide is a secreted heterologous polypeptide that is recovered from the medium as a soluble heterologous polypeptide. In some embodiments, the heterologous polypeptide is a secreted protein that includes a signal peptide.

[0042] In some embodiments, a heterologous polypeptide produced by the present invention comprises a targeting signal that directs its retention in the endoplasmic reticulum, its extracellular signal, or that directs it to another organelle or cellular compartment. In some embodiments, the heterologous polypeptide comprises a signal peptide. In some embodiments, the heterologous polypeptide comprises a Na / Pi-IIb2 transporter signal peptide or a Sec1 transport protein. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 1 or SEQ ID NO: 37. In some embodiments, a heterologous polypeptide comprising a signal peptide having the amino acid sequence of SEQ ID NO: 1 or SEQ ID NO: 37 is secreted into the culture medium. In some embodiments, the signal peptide is cleaved from the protein during the secretion process, resulting in the mature form of the protein.

[0043] In some embodiments, heterologous polypeptides produced by the host cells of the invention are glycosylated. In some embodiments, the glycosylation pattern of heterologous polypeptides produced by the invention more closely resembles mammalian glycosylation patterns than proteins produced in yeast or E. coli. In some embodiments, heterologous polypeptides produced by the microalgal host cells of the invention comprise an N-linked glycosylation pattern. Glycosylated proteins used for therapeutic purposes are less likely to stimulate an anti-glycotype immune response if their glycosylation pattern resembles that found in the target organism. Conversely, glycosylated proteins with linkages or sugars that are not unique to the target organism are more likely to be antigenic. Specific glycoforms can also modulate effector functions. For example, IgG can mediate pro- or anti-inflammatory responses, correlating with the absence or presence of terminal sialic acid in the Fc region glycoforms, respectively (Kaneko et al., Science 313:670-3 (2006)).

[0044] The present invention further relates to methods for producing a recombinant heterologous polypeptide, comprising culturing a recombinant microalgal host cell of the present invention under conditions sufficient to express a polynucleotide sequence encoding the heterologous polypeptide. In some embodiments, the recombinant heterologous polypeptide is secreted from the host cell and recovered from the culture medium. In some embodiments, the heterologous polypeptide secreted from the cell comprises a secretory signal peptide. Depending on the vector and host system used for production, the recombinant heterologous polypeptide of the present invention may remain within the recombinant cell, be secreted into the fermentation medium, be secreted into the space between two cell membranes, or be retained on the outer surface of the cell membrane. As used herein, the phrase "recovering the protein" refers to collecting the fermentation medium containing the protein and does not necessarily imply further steps of separation or purification. The heterologous polypeptide produced by the methods of the present invention can be purified using a variety of standard protein purification techniques, including, but not limited to, affinity chromatography, ion exchange chromatography, filtration, electrophoresis, hydrophobic interaction chromatography, gel filtration chromatography, reversed-phase chromatography, concanavalin A chromatography, chromatofocusing, and differential solubilization. In some embodiments, the heterologous polypeptide produced by the methods of the present invention is isolated in a "substantially pure" form. As used herein, "substantially pure" refers to a purity that allows for the effective use of the heterologous polypeptide as a commercial product. In some embodiments, the recombinant heterologous polypeptide accumulates intracellularly and is recovered from the cell. In some embodiments, the host cell of the method is a Thraustochytrid. In some embodiments, the host cell of the method is a Schizochytrium or Thraustochytrid. In some embodiments, the recombinant heterologous polypeptide is a therapeutic protein, a food enzyme, or an industrial enzyme. In some embodiments, the recombinant microalgae host cell is a Schizochytrium and the recombinant heterologous polypeptide is a therapeutic protein comprising a secretory signal sequence.

[0045] In some embodiments, the recombinant vector of the present invention is a targeted vector. As used herein, the phrase "targeted vector" refers to a vector used to deliver a specific nucleic acid molecule to a recombinant cell, which nucleic acid molecule is used to delete or inactivate an endogenous gene in a host cell (i.e., used for targeted gene disruption or knockout techniques). Such vectors are also known as "knockout" vectors. In some embodiments, a portion of the targeted vector has a nucleic acid sequence that is homologous to the nucleic acid sequence of a target gene in a host cell (i.e., a gene targeted for deletion or inactivation). In some embodiments, the nucleic acid molecule inserted into the vector (i.e., an insert) is homologous to the target gene. In some embodiments, the nucleic acid sequence of the vector insert is designed to bind to the target gene and undergo homologous recombination between the target gene and the insert, thereby deleting, inactivating, or attenuating the endogenous target gene (i.e., by mutating or deleting at least a portion of the endogenous target gene).

[0046] Isolated nucleic acid molecules According to the present invention, an isolated nucleic acid molecule is a nucleic acid molecule that has been removed (i.e., subjected to human manipulation) from its natural environment, which is the genome or chromosome in which the nucleic acid molecule is originally found. Thus, "isolated" does not necessarily reflect the extent to which the nucleic acid molecule has been purified, but rather indicates that the molecule does not contain the entire genome or chromosome in which the nucleic acid molecule is originally found. Isolated nucleic acid molecules can include DNA, RNA (e.g., mRNA), or derivatives of either DNA or RNA (e.g., cDNA). The term "nucleic acid molecule" refers to the physical nucleic acid molecule itself, and the terms "nucleic acid sequence" and "polynucleotide sequence" refer to the sequence of nucleotides in a nucleic acid molecule itself, although these terms are used interchangeably, particularly with respect to nucleic acid molecules, polynucleotide sequences, or nucleic acid sequences capable of encoding heterologous polypeptides. In some embodiments, the isolated nucleic acid molecules of the present invention are produced using recombinant DNA technology (e.g., polymerase chain reaction (PCR) amplification, cloning) or chemical synthesis. Isolated nucleic acid molecules include naturally occurring nucleic acid molecules and their homologs, including, but not limited to, modified nucleic acid molecules in which nucleotides have been inserted, deleted, substituted and / or inverted in a manner that imparts a desired effect on the sequence, function and / or biological activity of the heterologous polypeptide encoded by the modification, as well as naturally occurring allelic variants.

[0047] The nucleic acid sequence complement of a promoter sequence, a terminator sequence, a signal peptide sequence, or any other sequence refers to the nucleic acid sequence of the nucleic acid strand complementary to the strand having the promoter sequence, the terminator sequence, the signal peptide sequence, or any other sequence. Double-stranded DNA is understood to include a single-stranded DNA and its complementary strand having a sequence complementary to the single-stranded DNA. Thus, nucleic acid molecules may be double-stranded or single-stranded, and include nucleic acid molecules that form stable hybrids with the sequences of the present invention and / or with the complements of the sequences of the present invention under "stringent" hybridization conditions. Methods for deducing complementary sequences are known to those skilled in the art.

[0048] The term "polypeptide" includes not only single polypeptide chain molecules, but also multi-polypeptide complexes in which individual constituent polypeptides are linked by covalent or non-covalent means. According to the present invention, an isolated polypeptide is a polypeptide that has been removed from its natural environment (i.e., has been subjected to human manipulation) and can include, for example, purified proteins, purified peptides, partially purified proteins, partially purified peptides, recombinantly produced proteins or peptides, and synthetically produced proteins or peptides.

[0049] As used herein, a recombinant microorganism has a genome that has been modified (i.e., mutated or altered) from its normal (i.e., wild-type or naturally occurring) form using recombinant techniques. Recombinant microorganisms according to the present invention can include microorganisms in which nucleic acid molecules have been inserted, deleted, or modified (i.e., mutated, for example, by inserting, deleting, substituting, and / or inverting nucleotides) such that such modification produces a desired effect in the microorganism. As used herein, a genetic modification that results in reduced gene expression, reduced gene function, or reduced function of a gene product (i.e., the protein encoded by the gene) can be referred to as gene inactivation (complete or partial), deletion, interruption, disruption, or downregulation. For example, a genetic modification in a gene that results in a decrease in the function of the protein encoded by such gene can be the result of a complete deletion of the gene (i.e., the gene is not present in the recombinant microorganism, and therefore the protein is absent from the recombinant microorganism), a mutation in the gene that results in incomplete or no translation of the protein (e.g., the protein is not expressed), or a mutation in the gene that reduces or eliminates the original function of the protein (e.g., a protein is expressed that has reduced or no activity (e.g., enzymatic activity or action)). A genetic modification that results in an increase in gene expression or function can refer to amplification, overproduction, overexpression, activation, enhancement, addition, or upregulation of the gene.

[0050] promoter A promoter is a region of DNA that directs the transcription of an associated coding region.

[0051] In some embodiments, the promoter is derived from a microorganism of the Labyrinthulomycota phylum, hi some embodiments, the promoter is derived from a Thraustochytrid, including, but not limited to, the microorganism deposited as SAM2179 (named "Ulkenia SAM2179" by the depositor), a microorganism of the genera Ulkenia or Thraustochytrid, or Schizochytrium. The genus Schizochytrium includes, but is not limited to, Schizochytrium agregatum, Schizochytrium limacinum, Schizochytrium sp. (S31) (ATCC 20888), Schizochytrium sp. (S8) (ATCC 20889), Schizochytrium sp. (LC-RM) (ATCC 18915), Schizochytrium sp. (SR21), deposited Schizochytrium strain ATCC 28209, and deposited Schizochytrium strain IFO 32693.

[0052] Promoters can have promoter activity at least in thraustochytrids and include full-length promoter sequences and functional fragments thereof, fusion sequences, and homologs of naturally occurring promoters. Promoter homologs differ from naturally occurring promoters in that at least one, two, three, or several nucleotides have been deleted, inserted, inverted, substituted, and / or derivatized. Promoter homologs can retain promoter activity at least in thraustochytrids, but their activity can be increased, decreased, or produced by certain stimuli. Promoters can contain one or more sequence elements that confer developmental and tissue-specific regulatory control or expression.

[0053] In some embodiments, the isolated nucleic acid molecules described herein comprise a PUFA PKS OrfC promoter ("PKS OrfC promoter"; also known as a PFA3 promoter), such as the polynucleotide sequence represented by SEQ ID NO: 3. The PKS OrfC promoter comprises a PKS OrfC promoter homolog that is sufficiently similar to a native PKS OrfC promoter sequence that the nucleic acid sequence of the homolog can hybridize under moderate, high, or very high stringency conditions to the complement of the nucleic acid sequence of the native PKS OrfC promoter, such as SEQ ID NO: 3, or the OrfC promoter of pCL0001 deposited under ATCC Accession No. PTA-9615.

[0054] In some embodiments, an isolated nucleic acid molecule of the invention comprises an EF1 short promoter ("EF1 short" or "EF1-S" promoter) or an EF1 long promoter ("EF1 long" or "EF1-L" promoter), such as, for example, the EF1 short promoter represented by SEQ ID NO: 42 or the EF1 long promoter represented by SEQ ID NO: 43. An EF1 short or EF1 long promoter includes an EF1 short or long promoter homolog that is sufficiently similar to the native EF1 short and / or long promoter sequence, respectively, that the nucleic acid sequence of the homolog is capable of hybridizing under moderate, high, or very high stringency conditions to, for example, the complement of the nucleic acid sequence of the native EF1 short and / or long promoter, such as SEQ ID NO: 42 and / or SEQ ID NO: 43, respectively, or the EF1 long promoter of pAB0018 deposited under ATCC Accession No. PTA-9616.

[0055] In some embodiments, an isolated nucleic acid molecule of the invention comprises a 60S short promoter ("60S short" or "60S-S" promoter) or a 60S long promoter ("60S long" or "60S-L" promoter), such as, for example, a 60S short promoter represented by SEQ ID NO: 44 or a 60S long promoter having a polynucleotide sequence represented by SEQ ID NO: 45. In some embodiments, the 60S short or 60S long promoter comprises a 60S short or 60S long promoter homologue that is sufficiently similar to a native 60S short or 60S long promoter sequence, respectively, that the nucleic acid sequence of the homologue can hybridise under moderate, high, or very high stringency conditions to the complement of the nucleic acid sequence of a native 60S short and / or 60S long promoter, e.g., SEQ ID NO: 44 and / or SEQ ID NO: 45, respectively, or the 60S long promoter of pAB0011 deposited under ATCC Accession No. PTA-9614.

[0056] In some embodiments, the isolated nucleic acid molecule comprises a Sec1 promoter ("Sec1 promoter"), e.g., the polynucleotide sequence represented by SEQ ID NO: 46. In some embodiments, the Sec1 promoter comprises a Sec1 promoter homolog that is sufficiently similar to a native Sec1 promoter sequence that the nucleic acid sequence of the homolog can hybridize under moderate, high, or very high stringency conditions to the complement of the nucleic acid sequence of the native Sec1 promoter, e.g., SEQ ID NO: 46, or the Sec1 promoter of pAB0022 deposited under ATCC Accession No. PTA-9613.

[0057] Terminator A terminator region is a portion of a gene sequence that marks the end of the gene sequence in genomic DNA for transcription.

[0058] In some embodiments, the terminator region is derived from a microorganism of the Labyrinthulomycota phylum. In some embodiments, the terminator region is derived from a Thraustochytrid. In some embodiments, the terminator region is derived from the genus Schizochytrium or Thraustochytrid. The genus Schizochytrium includes, but is not limited to, Schizochytrium agregatum, Schizochytrium limacinum, Schizochytrium sp. (S31) (ATCC 20888), Schizochytrium sp. (S8) (ATCC 20889), Schizochytrium sp. (LC-RM) (ATCC 18915), Schizochytrium sp. (SR21), deposited strain ATCC 28209, and deposited strain IFO 32693. In some embodiments, the terminator region is a heterologous terminator region, such as, for example, a heterologous SV40 terminator region.

[0059] The terminator region can have terminator activity at least in a thraustochytrid, and includes full-length terminator sequences and functional fragments thereof, fusion sequences, and homologs of naturally occurring terminator regions. Terminator homologs differ from naturally occurring terminators in that at least one or several nucleotides have been deleted, inserted, inverted, substituted, and / or derivatized, including, but not limited to, one or several. In some embodiments, terminator homologs can retain terminator activity at least in a thraustochytrid, but can increase, decrease, or cause its activity to increase in response to certain stimuli.

[0060] In some embodiments, the isolated nucleic acid molecule can include the terminator region of the PUFA PKS OrfC gene (the "PKS OrfC terminator region," also known as the PFA3 terminator), such as the polynucleotide sequence represented by SEQ ID NO: 4. The terminator region disclosed in SEQ ID NO: 4 is a native (wild-type) terminator sequence from a microbial member of the Thraustochytrid family, specifically the Schizochytrium PKS OrfC terminator region, referred to as "OrfC terminator element 1." In some embodiments, the PKS OrfC terminator region comprises a PKS OrfC terminator region homolog that is sufficiently similar to a native PUFA PKS OrfC terminator region such that the nucleic acid sequence of the homolog can hybridize under moderate, high, or very high stringency conditions to the complement of the nucleic acid sequence of the native PKS OrfC terminator region, e.g., SEQ ID NO: 4, or the OrfC terminator region of pAB0011 deposited under ATCC Accession No. PTA-9614.

[0061] signal peptide In some embodiments, the isolated nucleic acid molecule comprises a polynucleotide sequence encoding a signal peptide of a secreted protein from a microorganism of the Labyrinthulomycota phylum. In some embodiments, the microorganism is a Thraustochytrid. In some embodiments, the microorganism is a Schizochytrium or Thraustochytrid.

[0062] Signal peptides can have secretion signal activity in thraustochytrids and include full-length peptides and functional fragments thereof, fusion sequences, and homologs of native signal peptides. Homologues of signal peptides differ from native signal peptides in that at least one or several amino acids have been deleted (e.g., a truncated form of the protein, such as a peptide or fragment), inserted, inverted, substituted, and / or derivatized (e.g., by glycosylation, phosphorylation, acetylation, myristoylation, prenylation, palmitation, amidation, and / or addition of glycosylphosphatidylinositol). In some embodiments, signal peptide homologues retain signal activity at least in thraustochytrids, but can increase, decrease, or cause their activity to increase in response to certain stimuli.

[0063] In some embodiments, the isolated nucleic acid molecule comprises a polynucleotide sequence encoding a Na / Pi-IIb2 transporter protein signal peptide. The Na / Pi-IIb2 transporter protein signal peptide can have signal targeting activity for at least the Na / Pi-IIb2 transporter protein in at least a thraustochytrid, and includes full-length peptides and functional fragments thereof, fusion peptides, and homologs of naturally occurring Na / Pi-IIb2 transporter protein signal peptides. In some embodiments, the Na / Pi-IIb2 transporter protein signal peptide has the amino acid sequence represented by SEQ ID NO: 1. In some embodiments, the Na / Pi-IIb2 transporter protein signal peptide has the amino acid sequence represented by SEQ ID NO: 15. In some embodiments, the isolated nucleic acid molecule comprises a polynucleotide sequence encoding an isolated amino acid sequence comprising a functional fragment of SEQ ID NO: 1 or SEQ ID NO: 15 that functions as a signal peptide for at least the Na / Pi-IIb2 transporter protein in at least a thraustochytrid. In some embodiments, the isolated nucleic acid molecule comprises SEQ ID NO:2.

[0064] The present invention is also directed to an isolated polypeptide comprising the amino acid sequence of the Na / Pi-IIb2 transporter signal peptide.

[0065] In some embodiments, an isolated nucleic acid molecule comprises a polynucleotide sequence encoding an alpha-1,6-mannosyltransferase (ALG12) signal peptide. The ALG12 signal peptide can have signal targeting activity for at least the ALG12 protein in at least a thraustochytrid, and includes the full-length peptide and functional fragments thereof, fusion peptides, and homologs of the native ALG12 signal peptide. In some embodiments, the ALG12 signal peptide has the amino acid sequence represented by SEQ ID NO: 59. In some embodiments, the isolated nucleic acid molecule comprises a polynucleotide sequence encoding an isolated amino acid sequence comprising a functional fragment of SEQ ID NO: 59 that functions as a signal peptide for at least ALG12 in at least a thraustochytrid. In some embodiments, the isolated nucleic acid molecule comprises SEQ ID NO: 60.

[0066] The present invention is also directed to an isolated polypeptide comprising the amino acid sequence of the ALG12 signal peptide.

[0067] In some embodiments, an isolated nucleic acid molecule comprises a polynucleotide sequence encoding a binding immunoglobulin protein (BiP) signal peptide. The BiP signal peptide can have signal targeting activity for at least BiP protein in at least a thraustochytrid, and includes the full-length peptide and functional fragments thereof, fusion peptides, and homologs of the native BiP signal peptide. In some embodiments, the BiP signal peptide has the amino acid sequence represented by SEQ ID NO: 61. In some embodiments, the isolated nucleic acid molecule comprises a polynucleotide sequence encoding an isolated amino acid sequence comprising a functional fragment of SEQ ID NO: 61 that functions as a signal peptide for at least BiP protein in at least a thraustochytrid. In some embodiments, the isolated nucleic acid molecule comprises SEQ ID NO: 62.

[0068] The present invention is also directed to an isolated polypeptide comprising the amino acid sequence of the BiP signal peptide.

[0069] In some embodiments, an isolated nucleic acid molecule comprises a polynucleotide sequence encoding an alpha-1,3-glucosidase (GLS2) signal peptide. The GLS2 signal peptide can have signal targeting activity for at least the GLS2 protein in at least a thraustochytrid, and includes the full-length peptide and functional fragments thereof, fusion peptides, and homologs of the native GLS2 signal peptide. In some embodiments, the GLS2 signal peptide has the amino acid sequence represented by SEQ ID NO: 63. In some embodiments, the isolated nucleic acid molecule comprises a polynucleotide sequence encoding an isolated amino acid sequence comprising a functional fragment of SEQ ID NO: 63 that functions as a signal peptide for at least the GLS2 protein in at least a thraustochytrid. In some embodiments, the isolated nucleic acid molecule comprises SEQ ID NO: 64.

[0070] The present invention is also directed to an isolated polypeptide comprising the amino acid sequence of the GLS2 signal peptide.

[0071] In some embodiments, the isolated nucleic acid molecule comprises a polynucleotide sequence encoding an α-1,3-1,6-mannosidase-like signal peptide. The α-1,3-1,6-mannosidase-like signal peptide can have signal targeting activity for at least an α-1,3-1,6-mannosidase-like protein in at least a thraustochytrid, and includes the full-length peptide and functional fragments thereof, fusion peptides, and homologs of naturally occurring α-1,3-1,6-mannosidase-like signal peptides. In some embodiments, the α-1,3-1,6-mannosidase-like signal peptide has the amino acid sequence represented by SEQ ID NO: 65. In some embodiments, the isolated nucleic acid molecule comprises a polynucleotide sequence encoding an isolated amino acid sequence comprising a functional fragment of SEQ ID NO: 65 that functions as a signal peptide for at least an α-1,3-1,6-mannosidase-like protein in at least a thraustochytrid. In some embodiments, the isolated nucleic acid molecule comprises SEQ ID NO:66.

[0072] The present invention is also directed to an isolated polypeptide comprising the amino acid sequence of an α-1,3-1,6-mannosidase-like signal peptide.

[0073] In some embodiments, the isolated nucleic acid molecule comprises a polynucleotide sequence encoding an α-1,3-1,6-mannosidase-like #1 signal peptide. The α-1,3-1,6-mannosidase-like #1 signal peptide can have signal targeting activity for at least the α-1,3-1,6-mannosidase-like #1 protein in at least a thraustochytrid, and includes the full-length peptide and functional fragments thereof, fusion peptides, and homologs of the native α-1,3-1,6-mannosidase-like #1 signal peptide. In some embodiments, the α-1,3-1,6-mannosidase-like #1 signal peptide has the amino acid sequence represented by SEQ ID NO: 67. In some embodiments, the isolated nucleic acid molecule comprises a polynucleotide sequence encoding an isolated amino acid sequence comprising a functional fragment of SEQ ID NO: 67 that functions as a signal peptide for at least the α-1,3-1,6-mannosidase-like #1 protein in at least a thraustochytrid. In some embodiments, the isolated nucleic acid molecule comprises SEQ ID NO:68.

[0074] The present invention is also directed to an isolated polypeptide comprising the amino acid sequence of an α-1,3-1,6-mannosidase-like #1 signal peptide.

[0075] In some embodiments, the isolated nucleic acid molecule comprises a polynucleotide sequence encoding an α-1,2-mannosidase-like signal peptide. The α-1,2-mannosidase-like signal peptide can have signal targeting activity for at least an α-1,2-mannosidase-like protein in at least a thraustochytrid, and includes the full-length peptide and functional fragments thereof, fusion peptides, and homologs of naturally occurring α-1,2-mannosidase-like signal peptides. In some embodiments, the α-1,2-mannosidase-like signal peptide has the amino acid sequence represented by SEQ ID NO: 69. In some embodiments, the isolated nucleic acid molecule comprises a polynucleotide sequence encoding an isolated amino acid sequence comprising a functional fragment of SEQ ID NO: 69 that functions as a signal peptide for at least an α-1,2-mannosidase-like protein in at least a thraustochytrid. In some embodiments, the isolated nucleic acid molecule comprises SEQ ID NO: 70.

[0076] The present invention is also directed to an isolated polypeptide comprising the amino acid sequence of an α-1,2-mannosidase-like signal peptide.

[0077] In some embodiments, the isolated nucleic acid molecule comprises a polynucleotide sequence encoding a β-xylosidase-like signal peptide. The β-xylosidase-like signal peptide can have signal targeting activity for at least a β-xylosidase-like protein, at least in a thraustochytrid, and includes the full-length peptide and functional fragments thereof, fusion peptides, and homologs of naturally occurring β-xylosidase-like signal peptides. In some embodiments, the β-xylosidase-like signal peptide has the amino acid sequence represented by SEQ ID NO: 71. In some embodiments, the isolated nucleic acid molecule comprises a polynucleotide sequence encoding an isolated amino acid sequence comprising a functional fragment of SEQ ID NO: 71 that functions as a signal peptide for at least a β-xylosidase-like protein, at least in a thraustochytrid. In some embodiments, the isolated nucleic acid molecule comprises SEQ ID NO: 72.

[0078] The present invention is also directed to an isolated polypeptide comprising the amino acid sequence of a β-xylosidase-like signal peptide.

[0079] In some embodiments, the isolated nucleic acid molecule comprises a polynucleotide sequence encoding a carotene synthase signal peptide. The carotene synthase signal peptide can have signal targeting activity for at least a carotene synthase protein, at least in a thraustochytrid, and includes the full-length peptide and functional fragments thereof, fusion peptides, and homologs of naturally occurring carotene synthase signal peptides. In some embodiments, the carotene synthase signal peptide has the amino acid sequence represented by SEQ ID NO: 73. In some embodiments, the isolated nucleic acid molecule comprises a polynucleotide sequence encoding an isolated amino acid sequence comprising a functional fragment of SEQ ID NO: 73 that functions as a signal peptide for at least a carotene synthase protein, at least in a thraustochytrid. In some embodiments, the isolated nucleic acid molecule comprises SEQ ID NO: 74.

[0080] The present invention is also directed to an isolated polypeptide comprising the amino acid sequence of a carotene synthase signal peptide.

[0081] In some embodiments, the isolated nucleic acid molecule comprises a polynucleotide sequence encoding a Sec1 protein ("Sec1") signal peptide. The Sec1 signal peptide can have secretion signal activity for at least Sec1 protein, at least in thraustochytrids, and includes the full-length peptide and functional fragments thereof, fusion peptides, and homologs of the native Sec1 signal peptide. In some embodiments, the Sec1 signal peptide is represented by SEQ ID NO: 37. In some embodiments, the isolated nucleic acid molecule comprises a polynucleotide sequence encoding an isolated amino acid sequence comprising a functional fragment of SEQ ID NO: 37 that functions as a signal peptide for at least Sec1 protein, at least in thraustochytrids. In some embodiments, the isolated nucleic acid molecule comprises SEQ ID NO: 38.

[0082] The present invention is also directed to an isolated polypeptide comprising the amino acid sequence of the Sec1 signal peptide.

[0083] In some embodiments, the isolated nucleic acid molecule can comprise a promoter sequence, a terminator sequence, and / or a signal peptide sequence that is at least 90%, 95%, 96%, 97%, 98%, or 99% identical to any of the promoter, terminator, and / or signal peptide sequences described herein.

[0084] In some embodiments, the isolated nucleic acid molecule comprises an OrfC promoter, an EF1 short promoter, an EF1 long promoter, a 60S short promoter, a 60S long promoter, a Sec1 promoter, a PKS OrfC terminator region, a sequence encoding a Na / Pi-IIb2 transporter protein signal peptide, or a sequence encoding a Sec1 transporter protein signal peptide operably linked to the 5' end of a nucleic acid sequence encoding a heterologous polypeptide. Recombinant vectors (including, but not limited to, expression vectors), expression cassettes, and host cells can similarly comprise an OrfC promoter, an EF1 short promoter, an EF1 long promoter, a 60S short promoter, a 60S long promoter, a Sec1 promoter, a PKS OrfC terminator region, a sequence encoding a Na / Pi-IIb2 transporter protein signal peptide, or a sequence encoding a Sec1 transporter protein signal peptide operably linked to the 5' end of a nucleic acid sequence encoding a heterologous polypeptide.

[0085] As used herein, unless otherwise specified, references to percent identity (% identity) refer to assessments of homology performed using: (1) BLAST 2.0 Basic BLAST homology searches using blastp for amino acid searches and blastn for nucleic acid searches with standard default parameters, in which the query sequence is filtered for low complexity regions by default (see, e.g., Altschul, S., et al., Nucleic Acids Res. 25:3389-3402 (1997), incorporated herein by reference in its entirety); (2) BLAST2 alignments using the parameters described below; and / or (3) PSI-BLAST (position-specific iterative BLAST) with standard default parameters. It should be noted that due to some differences in standard parameters between BLAST 2.0 Basic BLAST and BLAST2, even if two particular sequences may be recognized as having significant homology using the BLAST2 program, a search performed in BLAST 2.0 Basic BLAST using one of the sequences as the query sequence may not identify the second sequence as the top match. Furthermore, PSI-BLAST provides an automated, easy-to-use version of "profile" searching, which is a sensitive method for locating sequence homologs. The program first performs a gapped BLAST database search. The PSI-BLAST program also uses information from any significant alignments returned to construct a position-specific score matrix, which replaces the query sequence for the next round of database searching. Therefore, it should be understood that percent identity can be determined using any one of these programs.

[0086] Two particular sequences can be aligned to each other using BLAST2, for example, as described in Tatusova and Madden, FEMS Microbiol. Lett. 174:247-250 (1999), which is incorporated herein by reference in its entirety. BLAST2 sequence alignment is performed in blastp or blastn using the BLAST2.0 algorithm, which performs a gapped BLAST search (BLAST2.0) between the two sequences, allowing for the introduction of gaps (deletions and insertions) in the resulting alignment. In some embodiments, BLAST2 sequence alignment is performed using the following standard default parameters: For blastn, the following BLOSUM62 matrix is used: Reward for Match = 1 Penalty for mismatch = -2 Open gap (5) and extension gap (2) penalties Gap x_DropOff(50) Expectation(10) WordSize(11) Filter(On) For blastp, the following BLOSUM62 matrix is used: Open gap (11) and extension gap (1) penalties Gap x_DropOff(50) Expectation(10) WordSize(3) Filter(On)

[0087] As used herein, hybridization conditions refer to the standard hybridization conditions used by nucleic acid molecules to identify similar nucleic acid molecules.See, for example, Sambrook J. and Russell D. (2001) Molecular cloning: A laboratory manual, 3rd ed. Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, which is incorporated herein by reference in its entirety.In addition, formulas for calculating the appropriate hybridization and washing conditions for achieving hybridization that allows for varying degrees of nucleotide mismatch are disclosed, for example, in Meinkoth et al., Anal. Biochem. 138:267-284 (1984), which is incorporated herein by reference in its entirety.Those skilled in the art can use the formula of Meinkoth et al. to calculate the appropriate hybridization and washing conditions for achieving a specific level of nucleotide mismatch.Such conditions will vary depending on whether DNA:RNA or DNA:DNA hybrids are formed. The calculated melting temperature for DNA:DNA hybrids is 10° C. lower than for DNA:RNA hybrids. In certain embodiments, stringent hybridization conditions for DNA:DNA hybrids are 6× SSC (0.9 M NaCl, 0.1 ... + ) at temperatures between 20°C and 35°C (low stringency), between 28°C and 40°C (more stringent), and between 35°C and 45°C (even more stringent), along with appropriate wash conditions. In certain embodiments, stringent hybridization conditions for DNA:RNA hybrids include hybridization at an ionic strength of 0.05 M NaCl, ... +) at temperatures between 30°C and 45°C, between 38°C and 50°C, and between 45°C and 55°C, along with similarly stringent wash conditions. These values are based on melting temperature calculations for molecules larger than about 100 nucleotides, 0% formamide, and about 40% G+C content. m can also be calculated empirically, as shown in Sambrook et al. In general, wash conditions should be as stringent as possible and appropriate for the chosen hybridization conditions. For example, hybridization conditions can be determined based on the calculated T m Wash conditions can include a combination of salt and temperature conditions that are about 20°C to about 25°C lower than the calculated T for that particular hybrid. m This includes a combination of salt and temperature conditions that are about 12° C. to about 20° C. lower than those of Example 1. One example of hybridization conditions suitable for use with DNA:DNA hybrids includes hybridization in 6×SSC (50% formamide) at 42° C. for 2 to 24 hours, followed by a wash step including one or more washes in 2×SSC at room temperature, followed by additional washes at a higher temperature and lower ionic strength (e.g., at least one wash in 0.1× to 0.5×SSC at 37° C., followed by at least one wash in 0.1× to 0.5×SSC at 68° C.).

[0088] Heterologous Polypeptides As used herein, the term "heterologous" refers to a sequence that is not naturally found in the microalgal host cell. In some embodiments, heterologous polypeptides produced by the recombinant host cells of the invention include, but are not limited to, therapeutic proteins. As used herein, "therapeutic proteins" include proteins that are useful for the treatment or prevention of diseases, conditions, or disorders in animals and humans.

[0089] In certain embodiments, therapeutic proteins include biologically active proteins, such as, but not limited to, enzymes, antibodies, or antigenic proteins.

[0090] In some embodiments, heterologous polypeptides produced by recombinant host cells of the invention include, but are not limited to, industrial enzymes, including, but not limited to, enzymes used in the manufacture, preparation, preservation, nutrient mobilization, or processing of products, including food, pharmaceutical, chemical, mechanical, and other industrial products.

[0091] In some embodiments, the heterologous polypeptide produced by a recombinant host cell of the invention comprises an auxotrophic marker, a dominant selectable marker (such as, for example, an enzyme that reduces antibiotic activity) or another protein involved in transformation selection, a protein that functions as a reporter, an enzyme involved in protein glycosylation, and an enzyme involved in cellular metabolism.

[0092] In some embodiments, the heterologous polypeptide produced by a recombinant host cell of the invention comprises a viral protein selected from the group consisting of H or HA (hemagglutinin) protein, N or NA (neuraminidase) protein, F (fusion) protein, G (glycoprotein) protein, E or env (envelope) protein, gp120 (120 kDa glycoprotein), and gp41 (41 kDa glycoprotein). In some embodiments, the heterologous polypeptide produced by a recombinant host cell of the invention is a viral matrix protein. In some embodiments, the heterologous polypeptide produced by a recombinant host cell of the invention is a viral matrix protein selected from the group consisting of M1, M2 (membrane channel protein), Gag, and combinations thereof. In some embodiments, the HA, NA, F, G, E, gp120, gp41, or matrix protein is derived from a viral source, such as influenza virus or measles virus.

[0093] Influenza is a leading cause of death in humans due to respiratory viruses. Common symptoms include fever, sore throat, shortness of breath, and muscle aches, among others. Influenza viruses are enveloped viruses that bud from the plasma membrane of infected mammalian and avian cells. They are classified as type A, B, or C based on the nucleoprotein and matrix protein antigens present. Influenza A viruses can be further classified into subtypes according to the combination of HA and NA presented by surface glycoproteins. HA is an antigenic glycoprotein responsible for binding of the virus to infected cells. NA removes terminal sialic acid residues from glycan chains on host cells and viral surface proteins, thereby preventing virus aggregation and promoting viral mobility.

[0094] The influenza virus HA protein is a homotrimer with a receptor-binding pocket in the globular head of each monomer, and the influenza virus NA protein is a tetramer with an enzymatic active site in the head of each monomer. Currently, 16 HA (H1-H16) and 9 NA (N1-N9) subtypes are recognized. Each influenza A virus displays one HA and one NA glycoprotein. Generally, each subtype exhibits species specificity; for example, all HA and NA subtypes are known to infect birds, but only subtypes H1, H2, H3, H5, H7, H9, H10, N1, N2, N3, and N7 have been shown to infect humans. Influenza viruses are characterized by the types of HA and NA they carry, e.g., H1N1, H5N1, H1N2, H1N3, H2N2, H3N2, H4N6, H5N2, H5N3, H5N8, H6N1, H7N7, H8N4, H9N2, H10N3, H11N2, H11N9, H12N5, H13N8, H15N8, H16N3, etc. Subtypes are further divided into strains; each genetically distinct viral isolate is typically considered to be a distinct strain, e.g., Influenza A / Puerto Rico / 8 / 34 / Mount Sinai (H1N1) and Influenza A / Vietnam / 1203 / 2004 (H5N1). In certain embodiments of the invention, the HA is derived from an influenza virus, e.g., the HA is from influenza A, influenza B, or an influenza A subtype selected from the group consisting of H1, H2, H3, H4, H5, H6, H7, H8, H9, H10, H11, H12, H13, H14, H15, and H16. In another embodiment, the HA is from an influenza A subtype selected from the group consisting of H1, H2, H3, H5, H6, H7, and H9. In one embodiment, the HA is from influenza subtype H1N1.

[0095] The influenza virus HA protein is translated intracellularly as a single protein, which, after cleavage of the signal peptide, yields (by conceptual translation) an approximately 62 kDa protein designated HA0 (i.e., hemagglutinin precursor protein). For viral activation, the hemagglutinin precursor protein (HA0) must be cleaved by a trypsin-like serine endoprotease at a specific site, usually encoded by a single basic amino acid (usually arginine) between the HA1 and HA2 polypeptides of the protein. In the specific example of the A / Puerto Rico / 8 / 34 strain, this cleavage occurs between the arginine at amino acid number 343 and the glycine at amino acid number 344. After cleavage, the two disulfide-linked protein polypeptides produce the mature form of the protein subunit, a prerequisite for the conformational changes required for fusion and, therefore, viral infectivity.

[0096] In some embodiments, the HA proteins of the invention are cleaved, e.g., the HA0 proteins of the invention are cleaved into HA1 and HA2. In some embodiments, expression of the HA protein in a microalgae host cell, such as Schizochytrium, results in proper cleavage of the HA0 protein into functional HA1 and HA2 polypeptides without the addition of exogenous proteases. Such cleavage of hemagglutinin in an invertebrate expression system without the addition of exogenous proteases has not previously been demonstrated.

[0097] The viral F protein may contain a single transmembrane domain near the C-terminus. The F protein can be divided into two peptides at the furin cleavage site (amino acid 109). The first portion of the protein, designated F2, contains the N-terminal portion of the complete F protein. The remainder of the viral F protein, including the C-terminal portion of the F protein, is designated F1. The F1 and / or F2 regions may be fused individually to heterologous sequences, such as sequences encoding heterologous signal peptides. Vectors containing the F1 and F2 portions of the viral F protein may be expressed individually or in combination. A vector expressing the complete F protein may be coexpressed with a furin enzyme that cleaves the protein at the furin cleavage site. Alternatively, the sequence encoding the furin cleavage site of the F protein may be replaced with a sequence encoding another protease cleavage site that is recognized and cleaved by a different protease. An F protein containing another protease cleavage site may be coexpressed with a corresponding protease that recognizes and cleaves the other protease cleavage site.

[0098] In some embodiments, the HA, NA, F, G, E, gp120, gp41, or matrix protein is a full-length protein, a fragment, variant, derivative, or analog thereof. In some embodiments, the HA, NA, F, G, E, gp120, gp41, or matrix protein is a polypeptide comprising an amino acid sequence at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the known sequence for the respective viral protein, or a polynucleotide encoding a polypeptide comprising that amino acid sequence, wherein the polypeptide can be recognized by an antibody that specifically binds to the known sequence. The HA sequence can be, for example, a full-length HA protein consisting essentially of the extracellular (ECD) domain, transmembrane (TM) domain, and cytoplasmic (CYT) domain; or a fragment of the entire HA protein consisting essentially of the HA1 polypeptide and the HA2 polypeptide, produced, for example, by cleavage of the entire HA; or a fragment of the entire HA protein consisting essentially of the HA1 polypeptide, the HA2 polypeptide, and the TM domain; or a fragment of the entire HA protein consisting essentially of the CYT domain; or a fragment of the entire HA protein consisting essentially of the TM domain; or a fragment of the entire HA protein consisting essentially of the HA1 polypeptide; or a fragment of the entire HA protein consisting essentially of the HA2 polypeptide. The HA sequence can also include an HA1 / HA2 cleavage site. The HA1 / HA2 cleavage site can be located between the HA1 polypeptide and the HA2 polypeptide, but can be arranged in any order relative to the other sequences of the polynucleotide or polypeptide construct. The viral protein can be derived from a pathogenic virus strain.

[0099] In some embodiments, the heterologous polypeptide of the invention is a fusion polypeptide comprising a full-length HA, NA, F, G, E, gp120, gp41, or matrix protein, or a fragment, variant, derivative, or analog thereof.

[0100] In some embodiments, the heterologous polypeptide is a fusion polypeptide comprising an HA0 polypeptide, an HA1 polypeptide, an HA2 polypeptide, a TM domain, fragments thereof, and combinations thereof. In some embodiments, the heterologous polypeptide comprises a combination of two or more of an HA1 polypeptide, an HA2 polypeptide, a TM domain, or fragments thereof from different subtypes or strains of a virus, such as from different subtypes or strains of influenza virus. In some embodiments, the heterologous polypeptide comprises a combination of two or more of an HA1 polypeptide, an HA2 polypeptide, a TM domain, or fragments thereof from different viruses, such as from influenza virus and measles virus.

[0101] Hemagglutination activity can be determined by measuring the agglutination of red blood cells. The agglutination and subsequent precipitation of red blood cells results from hemagglutinins adsorbed onto the surface of red blood cells. During hemagglutination, clusters of red blood cells are formed, visible to the naked eye, as heaps, lumps, and / or clumps. Hemagglutination is caused by the interaction of agglutinogens present on red blood cells with plasma containing agglutinins. Each agglutinogen has a corresponding agglutinin. Hemagglutination reactions are used, for example, to determine antiserum activity or type of virus. Active hemagglutination, which is caused by the direct action of an agent on red blood cells, is distinguished from passive hemagglutination, which is caused by specific antisera against antigens preadsorbed by red blood cells. The amount of hemagglutination activity in a sample can be measured, for example, in hemagglutination units (HAU). Hemagglutination can be caused, for example, by polysaccharides of the pathogenic bacteria of tuberculosis, plague and tularemia, by polysaccharides of Escherichia coli (colon bacillus), and by viruses of white mouse influenza, mumps, pneumonia, swine and equine influenza, smallpox vaccine, yellow fever and other hemagglutination-inducing diseases.

[0102] microalgae extracellular body The present invention also covers microalgal extracellular bodies that are discontinuous with the plasma membrane. "Discontinuous with the plasma membrane" means that the microalgal extracellular body is not connected to the plasma membrane of the host cell. In some embodiments, the extracellular body is a membrane. In some embodiments, the extracellular body is a vesicle, a micelle, a membrane fragment, a membrane aggregate, or a mixture thereof. As used herein, the term "vesicle" refers to a closed structure comprising a lipid bilayer (unit membrane), e.g., a bubble-like structure formed by the plasma membrane. As used herein, the term "membrane aggregate" refers to any collection of membrane structures that become associated into a single mass. A membrane aggregate can be a collection of a single type of membrane structure, such as, but not limited to, a collection of membrane vesicles, or can be a collection of more than a single type of membrane structure, such as, but not limited to, a collection of at least two vesicles, micelles, or membrane fragments. As used herein, the term "membrane fragment" refers to any portion of a membrane that may contain a heterologous polypeptide described herein. In some embodiments, the membrane fragment is a membrane sheet. In some embodiments, the extracellular body is a mixture of vesicles and membrane fragments. In some embodiments, the extracellular body is a vesicle. In some embodiments, the vesicle is a collapsed vesicle. In some embodiments, the vesicle is a virus-like particle. In some embodiments, the extracellular body is an aggregate of biological material comprising native polypeptides and heterologous polypeptides produced by a host cell. In some embodiments, the extracellular body is an aggregate of native polypeptides and heterologous polypeptides. In some embodiments, the extracellular body is an aggregate of heterologous polypeptides.

[0103] In some embodiments, the ectoplasmic net of the microalgae host cells fragments during cultivation of the microalgae host cells, resulting in the formation of microalgal extracellular bodies. In some embodiments, the microalgae extracellular bodies are formed by fragmentation of the ectoplasmic net of the microalgae host cells as a result of hydrodynamic forces in the agitated medium that physically shear the ectoplasmic net membrane protrusions.

[0104] In some embodiments, the microalgal extracellular body is formed by, but not limited to, extrusion of a microalgal membrane, such as extrusion of the plasma membrane, ectoplasmic net, pseudorhizoid, or a combination thereof, wherein the extruded membrane becomes separated from the plasma membrane.

[0105] In some embodiments, the microalgal extracellular bodies are vesicles or micelles with different diameters, membrane fragments with different lengths, or a combination thereof.

[0106] In some embodiments, the extracellular body is 10 nm to 2500 nm, 10 nm to 2000 nm, 10 nm to 1500 nm, 10 nm to 1000 nm, 10 nm to 500 nm, 10 nm to 300 nm, 10 nm to 200 nm, 10 nm to 100 nm, 10 nm to 50 nm, 20 nm to 2500 nm, 20 nm to 2000 nm. nm, 20 nm~1500 nm, 20 nm~1000 nm, 20 nm~500 nm, 20 nm~300 nm, 20 nm~200 nm, 20 nm~100 nm, 50 nm~2500 nm, 50 nm~2000 nm, 50 nm~1500 nm, 50 nm~1000 nm, 50 nm~500 nm, 50 nm~300 nm, 50 nm~200 nm, 50 nm~100 nm, 100 The vesicles are vesicles having a diameter of 100 nm to 2500 nm, 100 nm to 2000 nm, 100 nm to 1500 nm, 100 nm to 1000 nm, 100 nm to 500 nm, 100 nm to 300 nm, 100 nm to 200 nm, 500 nm to 2500 nm, 500 nm to 2000 nm, 500 nm to 1500 nm, 500 nm to 1000 nm, 2000 nm or less, 1500 nm or less, 1000 nm or less, 500 nm or less, 400 nm or less, 300 nm or less, 200 nm or less, 100 nm or less, or 50 nm or less.

[0107] Non-limiting fermentation conditions for producing microalgal exobodies from thraustochytrid host cells are shown below in Table 1.

[0108] (Table 1) Vessel medium TIFF2025114812000001.tif165136

[0109] Typical culture conditions for producing microalgal extracellular bodies include: pH: 5.5-9.5, 6.5-8.0, or 6.3-7.3 Temperature: 15°C~45°C, 18°C~35°C, or 20°C~30°C Dissolved oxygen: 0.1% to 100% saturation, 5% to 50% saturation, or 10% to 30% saturation Glucose control: 5 g / L to 100 g / L, 10 g / L to 40 g / L, or 15 g / L to 35 g / L

[0110] In some embodiments, the microalgal extracellular bodies are produced from host cells of the phylum Labyrinthulomycota. In some embodiments, the microalgal extracellular bodies are produced from host cells of the class Labyrinthulomycota. In some embodiments, the microalgal extracellular bodies are produced from host cells of the family Thraustochytrid. In some embodiments, the microalgal extracellular bodies are produced from the genus Schizochytrium or Thraustochytrid.

[0111] The present invention is also directed to a microalgal extracellular body comprising a heterologous polypeptide that is discontinuous with the plasma membrane of the microalgal host cell.

[0112] In some embodiments, the microalgal extracellular bodies of the present invention also comprise polypeptides associated with the plasma membrane of the microalgal host cell. In some embodiments, the polypeptides associated with the plasma membrane of the microalgal host cell comprise native membrane polypeptides, heterologous polypeptides, and combinations thereof.

[0113] In some embodiments, the heterologous polypeptide is contained within the microalgal extracellular body.

[0114] In some embodiments, the heterologous polypeptide comprises a membrane domain. As used herein, the term "membrane domain" refers to any domain within a polypeptide that targets the polypeptide to a membrane and / or maintains the polypeptide associated with a membrane, including, but not limited to, a transmembrane domain (e.g., a single or multiple transmembrane regions), an integral monotopic domain, a signal anchor sequence, an ER signal sequence, an N-terminal, internal, or C-terminal transport stop signal, a glycosylphosphatidylinositol anchor, and combinations thereof. The membrane domain can be located anywhere in the polypeptide, including the N-terminus, C-terminus, or middle of the polypeptide. The membrane domain can be associated with a membrane by permanent or transient attachment. In some embodiments, the membrane domain can be cleaved from a membrane protein. In some embodiments, the membrane domain is a signal anchor sequence. In some embodiments, the membrane domain is any of the signal anchor sequences shown in Figure 13 or an anchor sequence derived therefrom. In some embodiments, the membrane domain is a viral signal anchor sequence.

[0115] In some embodiments, the heterologous polypeptide is a polypeptide that naturally comprises a membrane domain. In some embodiments, the heterologous polypeptide does not naturally comprise a membrane domain but is recombinantly fused to a membrane domain. In some embodiments, the heterologous polypeptide is a normally soluble protein that is fused to a membrane domain.

[0116] In some embodiments, the membrane domain is a microalgae membrane domain. In some embodiments, the membrane domain is a Labyrinthulomycota membrane domain. In some embodiments, the membrane domain is a Thraustochytrid membrane domain. In some embodiments, the membrane domain is a Schizochytrium or Thraustochytrid membrane domain. In some embodiments, the membrane domain comprises a signal anchor sequence from Schizochytrium α-1,3-mannosyl-β-1,2-GlcNac-transferase-I-like protein #1 (SEQ ID NO: 78), Schizochytrium β-1,2-xylosyltransferase-like protein #1 (SEQ ID NO: 80), Schizochytrium β-1,4-xylosidase-like protein (SEQ ID NO: 82), or Schizochytrium galactosyltransferase-like protein #5 (SEQ ID NO: 84).

[0117] In some embodiments, heterologous polypeptide is membrane protein.As used herein, the term "membrane protein" refers to any protein that is associated with or bound to cell membrane.For example, as described by Chou and Elrod, Proteins: Structure, Function and Genetics 34:137-153(1999), membrane protein can be classified into various general types. 1) Type 1 membrane proteins: These proteins have a single transmembrane domain in the mature protein. The N-terminus is extracellular and the C-terminus is cytoplasmic. The N-terminus of the protein characteristically has a classical signal peptide sequence that directs the protein for import into the ER. These proteins are subdivided into type Ia (containing a cleavable signal sequence) and type Ib (lacking a cleavable signal sequence). Examples of type I membrane proteins include, but are not limited to, influenza HA, insulin receptor, glycophorin, LDL receptor, and viral G proteins. 2) Type II membrane proteins: These single-domain proteins have an extracellular C-terminus and a cytoplasmic N-terminus. The N-terminus may contain a signal-anchor sequence. Examples of this protein class include, but are not limited to, influenza neuraminidase, Golgi galactosyltransferase, Golgi sialyltransferase, sucrase-isomaltase precursor, asialoglycoprotein receptor, and transferrin receptor. 3) Multi-spanning membrane proteins: In type I and type II membrane proteins, the polypeptide crosses the lipid bilayer once, while in multi-spanning membrane proteins, the polypeptide crosses the membrane multiple times. Multi-spanning membrane proteins are also subdivided into type IIIa and type IIIb. Type IIIa proteins have a cleavable signal sequence. Type IIIb proteins have their amino termini exposed to the outer surface of the membrane but do not have a cleavable signal sequence. Type IIIa proteins include, but are not limited to, the M and L peptides of the photoreaction center. Type IIIb proteins include, but are not limited to, cytochrome P450 and E. coli leader peptidase. Further examples of multi-spanning membrane proteins are membrane transporters, such as sugar transporters (glucose, xylose) and ion transporters. 4) Lipid-chain-anchored membrane proteins: These proteins are attached to the membrane bilayer by one or more covalently attached fatty acid chains or other types of fatty chains called prenyl groups. 5) GPI-anchored membrane proteins: These proteins are attached to the membrane by a glycosylphosphatidylinositol (GPI) anchor. 6) Peripheral membrane proteins: These proteins are indirectly bound to the membrane by non-covalent interactions with other membrane proteins.

[0118] In some embodiments, the membrane domain is the membrane domain of an HA protein.

[0119] In some embodiments, the heterologous polypeptide comprises a native signal anchor sequence or native membrane domain from the wild-type polypeptide corresponding to the heterologous polypeptide. In some embodiments, the heterologous polypeptide is fused to a heterologous signal anchor sequence or heterologous membrane domain that is different from the native signal anchor sequence or native membrane domain. In some embodiments, the heterologous polypeptide comprises a heterologous signal anchor sequence or heterologous membrane domain, while the wild-type polypeptide corresponding to the heterologous polypeptide does not comprise any signal anchor sequence or membrane domain. In some embodiments, the heterologous polypeptide comprises a Schizochytrium signal anchor sequence. In some embodiments, the heterologous polypeptide comprises an HA membrane domain. In some embodiments, the heterologous polypeptide is a therapeutic polypeptide.

[0120] In some embodiments, the membrane domain is, or is derived from, any of the type I membrane proteins shown in Figure 14. In some embodiments, the heterologous polypeptide of the invention is a fusion polypeptide comprising a transmembrane region at the C-terminus of any of the membrane proteins shown in Figure 14. In some embodiments, the C-terminal side of the transmembrane region is further modified by substitution with a similar region from a viral protein.

[0121] In some embodiments, the heterologous polypeptide is a glycoprotein. In some embodiments, the heterologous polypeptide has a glycosylation pattern characteristic of expression in a Labyrinthulomycota cell. In some embodiments, the heterologous polypeptide has a glycosylation pattern characteristic of expression in a Thraustochytrid cell. In some embodiments, the heterologous polypeptide expressed in a microalgae host cell is a glycoprotein with a glycosylation pattern that more closely resembles a mammalian glycosylation pattern than proteins produced in yeast or E. coli. In some embodiments, the glycosylation pattern comprises an N-linked glycosylation pattern. In some embodiments, the glycoprotein comprises high mannose oligosaccharides. In some embodiments, the glycoprotein is substantially free of sialic acid. As used herein, the term "substantially free of sialic acid" refers to less than 10%, less than 9%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, less than 2%, or less than 1% sialic acid. In some embodiments, sialic acid is absent from the glycoprotein.

[0122] In some embodiments, the microalgal extracellular bodies of the present invention comprising heterologous polypeptides are produced on a commercial or industrial scale.

[0123] The present invention is also directed to compositions comprising any of the microalgal extracellular bodies of the present invention described herein and an aqueous liquid carrier.

[0124] In some embodiments, the microalgal extracellular bodies of the present invention comprising heterologous polypeptides are recovered from the culture or fermentation medium in which the microalgal host cells are grown. In some embodiments, the microalgal extracellular bodies of the present invention can be isolated in a "substantially pure" form. As used herein, "substantially pure" refers to a purity that allows for the effective use of the microalgal extracellular bodies as a commercial or industrial product.

[0125] The present invention is also directed to a method for producing a microalgal extracellular body comprising a heterologous polypeptide, the method comprising: (a) expressing in a microalgal host cell a heterologous polypeptide comprising a membrane domain; and (b) culturing the host cell under culture conditions sufficient to produce a microalgal extracellular body comprising the heterologous polypeptide that is discontinuous with the plasma membrane of the host cell.

[0126] The present invention is also directed to methods for producing a composition comprising a microalgal extracellular body and a heterologous polypeptide, the methods comprising: (a) expressing in a microalgal host cell a heterologous polypeptide comprising a membrane domain; and (b) culturing the host cell under culture conditions sufficient to produce a microalgal extracellular body comprising the heterologous polypeptide that is discontinuous with the plasma membrane of the host cell, wherein the composition is produced as a culture supernatant comprising the extracellular body. In some embodiments, the method further comprises removing the culture supernatant and resuspending the extracellular body in an aqueous liquid carrier. In some embodiments, the composition is used as a vaccine.

[0127] Microalgal extracellular bodies containing viral polypeptides Viral envelope proteins are membrane proteins that form the outer layer of virus particles. Their synthesis utilizes membrane domains, such as cellular targeting signals, to target the proteins to the plasma membrane. Envelope coat proteins are classified into several major groups, including, but not limited to, H or HA (hemagglutinin) proteins, N or NA (neuraminidase) proteins, F (fusion) proteins, G (glycoprotein) proteins, E or env (envelope) proteins, gp120 (a 120-kDa glycoprotein), and gp41 (a 41-kDa glycoprotein). Structural proteins, often referred to as "matrix" proteins, help stabilize the virus. Matrix proteins include, but are not limited to, M1, M2 (membrane channel proteins), and Gag. Both envelope and matrix proteins can be involved in viral assembly and function. For example, expression of viral envelope coat proteins, alone or in association with viral matrix proteins, can result in the formation of virus-like particles (VLPs).

[0128] Viral vaccines are often produced from inactivated or attenuated preparations of virus cultures corresponding to the disease they aim to prevent, and typically retain genetic material, such as viral genetic material. Typically, viruses are cultured from the same or similar cell types that they infect in the wild. Such cell cultures are often expensive and difficult to scale. To address this issue, certain specific viral protein antigens have instead been expressed by transgenic hosts, which are less expensive to cultivate and more amenable to scalability. However, viral proteins are typically integral membrane proteins present in the viral envelope. Because membrane proteins are very difficult to produce in large quantities, these viral proteins are typically modified to create soluble forms of the proteins. Although these viral envelope proteins are important for establishing host immunity, numerous attempts to express all or part of them in heterologous systems have met with limited success, likely because the proteins must be presented to the immune system in the context of the viral envelope membrane to be fully immunogenic. Therefore, there is a need for new heterologous expression systems, such as those of the present invention, that are scalable and capable of presenting viral antigens free or substantially free of associated viral material, such as viral genetic material other than the desired viral antigen. As used herein, the term "substantially free of associated viral material" means less than 10%, less than 9%, less than 8%, less than 7%, less than 5%, less than 4%, less than 3%, less than 2%, or less than 1% associated viral material.

[0129] In some embodiments, the microalgae extracellular body comprises a heterologous polypeptide that is a viral glycoprotein selected from the group consisting of H or HA (hemagglutinin) protein, N or NA (neuraminidase) protein, F (fusion) protein, G (glycoprotein) protein, E or env (envelope) protein, gp120 (120 kDa glycoprotein), gp41 (41 kDa glycoprotein), and combinations thereof. In some embodiments, the microalgae extracellular body comprises a heterologous polypeptide that is a viral matrix protein. In some embodiments, the microalgae extracellular body comprises a viral matrix protein selected from the group consisting of M1, M2 (membrane channel protein), Gag, and combinations thereof. In some embodiments, the microalgal extracellular body comprises a combination of two or more viral proteins selected from the group consisting of H or HA (hemagglutinin) protein, N or NA (neuraminidase) protein, F (fusion) protein, G (glycoprotein) protein, E or env (envelope) protein, gp120 (120 kDa glycoprotein), gp41 (41 kDa glycoprotein), and viral matrix protein.

[0130] In some embodiments, the microalgal extracellular bodies of the present invention comprise viral glycoproteins that lack sialic acid that may otherwise interfere with protein accumulation or function.

[0131] In some embodiments, the microalgal extracellular bodies are VLPs.

[0132] As used herein, the term "VLP" refers to a particle morphologically similar to an infectious virus that can be formed by spontaneous self-assembly of viral proteins when the viral proteins are overexpressed. VLPs have been produced in yeast, insect, and mammalian cells, and because they mimic the overall structure of a virus particle without containing infectious genetic material, they are believed to be an effective and safe type of subunit vaccine. This type of vaccine delivery system has been successful in stimulating cellular and humoral responses.

[0133] Studies on paramyxoviruses have revealed that when multiple viral proteins are expressed simultaneously, the VLPs produced are very similar in size and density to authentic virus particles. Expression of the matrix protein (M) alone is necessary and sufficient for VLP formation. In paramyxoviruses, expression of HN alone resulted in very low efficiency of VLP formation. Other proteins alone were not sufficient for NDV budding. HN is a type II membrane glycoprotein present as a tetrameric spike on the surface of virus particles and infected cells. See, e.g., Collins PL and Mottet G, J. Virol. 65:2362-2371 (1991); Mirza AM et al., J. Biol. Chem. 268:21425-21431 (1993); and Ng D et al., J. Cell. Biol. 109:3273-3289 (1989). Interaction with the M protein has been implicated in the incorporation of proteins HN and NP into VLPs. See, e.g., Pantua et al., J. Virology 80:11062-11073 (2006).

[0134] Hepatitis B virus (HBV) or human papillomavirus (HPV) VLPs are simple VLPs that are non-enveloped and produced by expressing one or two capsid proteins. More complex non-enveloped VLPs include particles such as those developed for bluetongue disease. In that case, four of the major structural proteins from bluetongue virus (BTV, Reoviridae) were co-expressed in insect cells. VLPs from viruses with lipid envelopes have also been produced (e.g., hepatitis C and influenza A). VLP-like structures, such as self-assembling polypeptide nanoparticles (SAPNs), which can display repeated antigenic epitopes, also exist. These have been used to design potential malaria vaccines. See, for example, Kaba SA et al., J. Immunol. 183(11):7268-7277 (2009).

[0135] VLPs offer a significant advantage in that they have the potential to generate immunity comparable to that of live-attenuated or inactivated viruses, and are believed to be highly immunogenic due to their particulate nature and the dense, repetitive array of surface epitopes they display. For example, it has been hypothesized that B cells specifically recognize particulate antigens with epitope spacing of 50 angstroms to 100 angstroms as foreign. See Bachman et al., Science 262:1448 (1993). VLPs also possess a particle size that is thought to greatly facilitate uptake by dendritic cells and macrophages. Furthermore, particles between 20 nm and 200 nm diffuse freely to lymph nodes, whereas particles between 500 and 2000 nm do not. There are at least two approved VLP vaccines in humans: hepatitis B (HBV) and human papillomavirus (HPV). However, virus-based VLPs, such as baculovirus-based VLPs, often contain large amounts of viral material that require further purification from the VLP.

[0136] In some embodiments, the microalgae extracellular bodies are VLPs comprising viral glycoproteins selected from the group consisting of H or HA (hemagglutinin) protein, N or NA (neuraminidase) protein, F (fusion) protein, G (glycoprotein) protein, E or env (envelope) protein, gp120 (120 kDa glycoprotein), gp41 (41 kDa glycoprotein), and combinations thereof. In some embodiments, the microalgae extracellular bodies are VLPs comprising viral matrix proteins. In some embodiments, the microalgae extracellular bodies are VLPs comprising viral matrix proteins selected from the group consisting of M1, M2 (membrane channel protein), Gag, and combinations thereof. In some embodiments, the microalgal extracellular body is a VLP comprising a combination of two or more viral proteins selected from the group consisting of H or HA (hemagglutinin) protein, N or NA (neuraminidase) protein, F (fusion) protein, G (glycoprotein) protein, E or env (envelope) protein, gp120 (120 kDa glycoprotein), gp41 (41 kDa glycoprotein), and viral matrix protein.

[0137] Methods for using microalgal extracellular bodies In some embodiments, the microalgal extracellular bodies of the present invention are useful as a medium for protein activity or function. In some embodiments, the protein activity or function is associated with a heterologous polypeptide present in or on the extracellular body. In some embodiments, the heterologous polypeptide is a membrane protein. In some embodiments, the protein activity or function is associated with a polypeptide that binds to a membrane protein present in the extracellular body. In some embodiments, the protein is not functional when soluble, but is functional when part of the extracellular body of the present invention. In some embodiments, microalgal extracellular bodies containing sugar transporters (such as xylose, sucrose, or glucose transporters) can be used to deplete trace amounts of sugars from a medium containing a mixture of sugars or other low molecular weight solutes by trapping the sugars in vesicles that can be separated by various methods, including filtration or centrifugation.

[0138] The present invention also includes the use of any of the microalgal extracellular bodies and compositions thereof of the present invention comprising heterologous polypeptides for therapeutic applications in animals or humans ranging from prophylactic treatments against diseases.

[0139] The terms "treat" and "treatment" refer to both therapeutic treatment and prophylactic or preventative measures, where the purpose is to prevent or slow (reduce) an undesirable physiological condition, disease, or disorder, or to obtain a beneficial or desired clinical result. For purposes of this invention, a beneficial or desired clinical result includes, but is not limited to, alleviation or elimination of symptoms or signs associated with a condition, disease, or disorder; reduction in the extent of the condition, disease, or disorder; stabilization of the condition, disease, or disorder (i.e., in which the condition, disease, or disorder has not worsened); delay in the onset or progression of the condition, disease, or disorder; amelioration of the condition, disease, or disorder; remission (whether partial or total, and whether detectable or undetectable) of the condition, disease, or disorder; or improvement or amelioration of the condition, disease, or disorder. Treatment includes eliciting a clinically significant response without undue side effects. Treatment also includes prolonging survival as compared to expected survival if not receiving treatment.

[0140] In some embodiments, any of the microalgal extracellular bodies of the present invention containing heterologous polypeptides are recovered in the culture supernatant for direct use as a vaccine for animals or humans.

[0141] In some embodiments, the microalgal extracellular bodies comprising heterologous polypeptides are purified according to the requirements of the intended use, e.g., administration as a vaccine. For a typical human vaccine application, the low-speed supernatant will undergo initial purification by concentration (e.g., tangential flow filtration followed by ultrafiltration), chromatographic separation (e.g., anion exchange chromatography), size exclusion chromatography, and sterilization (e.g., 0.2 μm filtration). In some embodiments, the vaccines of the present invention are free of potentially allergenic import proteins, such as egg proteins. In some embodiments, vaccines comprising the extracellular bodies of the present invention do not contain any viral material other than the viral polypeptides associated with the extracellular bodies.

[0142] According to the disclosed methods, microalgal extracellular bodies or compositions thereof comprising heterologous polypeptides can be administered, for example, by intramuscular (im), intravenous (iv), subcutaneous (sc), or intrapulmonary routes. Other suitable routes of administration include, but are not limited to, intratracheal, transdermal, intraocular, intranasal, inhalation, intraluminal, intraductal (e.g., to the pancreas), and intraparenchymal (e.g., to any tissue) administration. Transdermal delivery includes, but is not limited to, intradermal (e.g., to the dermis or epidermis), transdermal (e.g., transcutaneous), and transmucosal (e.g., to or through skin or mucosal tissue) administration. Intraluminal administration includes, but is not limited to, administration into the oral, vaginal, rectal, nasal, peritoneal, and intestinal cavities, as well as intrathecal (e.g., to the spinal canal), intravenous (e.g., to the ventricles of the brain or the ventricles), intraatrial (e.g., to the atrium), and intrathecal (e.g., to the subarachnoid space of the brain) administration.

[0143] In some embodiments, the present invention includes a composition comprising a microalgal extracellular body comprising a heterologous polypeptide. In some embodiments, the composition comprises an aqueous liquid carrier. In further embodiments, the aqueous liquid carrier is a culture supernatant. In some embodiments, the compositions of the present invention can be prepared in a conventional pharmaceutically acceptable excipient known in the art, such as, but not limited to, human serum albumin, ion exchangers, alumina, lecithin, buffer substances such as phosphate, glycine, sorbic acid, potassium sorbate, and salts or electrolytes such as protamine sulfate, as well as in a pharmaceutical composition containing a microalgal extracellular body comprising a heterologous polypeptide. st ed. (2005).

[0144] Alternatively, any of the embodiments described herein directed to microalgal extracellular bodies can be directed to chytrid extracellular bodies.

[0145] The most effective method of administration and dosage regimen for the compositions of this invention will depend on the severity and course of the disease, the subject's health and response to treatment, and the judgment of the treating physician. Accordingly, the dosage of the composition should be titrated to the individual subject. That being said, an effective dose of a composition of the present invention may be in the range of 1 mg / kg to 2000 mg / kg, 1 mg / kg to 1500 mg / kg, 1 mg / kg to 1000 mg / kg, 1 mg / kg to 500 mg / kg, 1 mg / kg to 250 mg / kg, 1 mg / kg to 100 mg / kg, 1 mg / kg to 50 mg / kg, 1 mg / kg to 25 mg / kg, 1 mg / kg to 10 mg / kg, 500 mg / kg to 2000 mg / kg, 500 mg / kg to 1500 mg / kg, 500 mg / kg to 1000 mg / kg, 100 mg / kg to 2000 mg / kg, 100 mg / kg to 1500 mg / kg, 100 mg / kg to 1000 mg / kg, or 100 mg / kg to 500 mg / kg.

[0146] Having generally described the invention, a further understanding can be obtained by reference to the examples provided herein, which are for illustrative purposes only and are not intended to be limiting. [Example]

[0147] Example 1 Construction of pCL0143 expression vector The pCL0143 expression vector (Figure 2) was synthesized and its sequence verified by Sanger sequencing by DNA 2.0 (Menlo Park, CA). The pCL0143 vector contains a promoter from the Schizochytrium elongation factor-1 gene (EF1) to drive expression of the HA transgene, the OrfC terminator (also known as the PFA3 terminator) following the HA transgene, and a selection marker cassette that confers resistance to the antibiotic paromomycin.

[0148] SEQ ID NO: 76 (FIG. 1) encodes the HA protein of influenza A virus (A / Puerto Rico / 8 / 34 / Mount Sinai (H1N1)). This protein sequence matches that of GenBank accession number AAM75158. The specific nucleic acid sequence of SEQ ID NO: 76 was codon-optimized and synthesized by DNA 2.0 for expression in Schizochytrium, as guided by the Schizochytrium codon usage table shown in FIG. 16. Using an alternative signal peptide, a construct was also generated in which the signal peptide (first 51 nucleotides) of SEQ ID NO: 76 was removed and replaced with a polynucleotide sequence encoding the Schizochytrium Sec1 signal peptide (SEQ ID NO: 38).

[0149] Example 2 Expression and characterization of HA proteins produced in Schizochytrium Schizochytrium sp. ATCC 20888 was used as a host cell for transformation with vector pCL0143 using a Biolistic™ Particle Bombarder (BioRad, Hercules, CA). Briefly, a culture of Schizochytrium sp. ATCC 20888 was grown in M2B medium consisting of 10 g / L glucose, 0.8 g / L (NH4)2SO4, 5 g / L Na2SO4, 2 g / L MgSO4·7H2O, 0.5 g / L KH2PO4, 0.5 g / L KCl, 0.1 g / L CaCl2·2H2O, 0.1 M MES (pH 6.0), 0.1% PB26 metals, and 0.1% PB26 vitamins (v / v). PB26 vitamins consisted of 50 mg / mL vitamin B12, 100 μg / mL thiamine, and 100 μg / mL calcium pantothenate. PB26 metals were adjusted to pH 4.5 and consisted of 3 g / L FeSO4·7H2O, 1 g / L MnCl2·4H2O, 800 mg / mL ZnSO4·7H2O, 20 mg / mL CoCl2·6H2O, 10 mg / mL Na2MoO4·2H2O, 600 mg / mL CuSO4·5H2O, and 800 mg / mL NiSO4·6H2O. PB26 stock solutions were filter-sterilized separately and added to the broth after autoclaving. Glucose, KH2PO4, and CaCl2·2H2O were each autoclaved separately from the rest of the broth components before mixing to prevent salt precipitation and caramelization of the carbohydrates. All media components were purchased from Sigma Chemical (St. Louis, MO). Cultures of Schizochytrium were grown to logarithmic phase and transformed using a Biolistic™ particle bombarder (BioRad, Hercules, CA). The Biolistic™ transformation procedure was essentially the same as previously described (see Apt et al., J. Cell. Sci. 115(Pt 21):4061-9 (1996) and U.S. Patent No. 7,001,772).Primary transformants were selected on solid M2B medium containing 20 g / L agar (VWR, West Chester, PA), 10 μg / mL sulfometuron methyl (SMM) (Chem Service, Westchester, PA) after 2–6 days of incubation at 27°C.

[0150] gDNA from primary transformants of pCL0143 was extracted, purified, and used as a template for PCR to test for the presence of the transgene.

[0151] Genomic DNA extraction protocol for Schizochytrium Schizochytrium transformants were grown in 50 ml of medium. 25 ml of culture was aseptically pipetted into a 50 ml conical vial and centrifuged at 3000 × g for 4 minutes to form a pellet. The supernatant was removed, and the pellet was stored at -80°C until use. The pellet was resuspended in approximately 4-5 volumes of a solution consisting of 20 mM Tris pH 8, 10 mM EDTA, 50 mM NaCl, 0.5% SDS, and 100 μg / ml proteinase K in a 50 ml conical vial. The pellet was incubated at 50°C with gentle rocking for 1 hour. Once dissolved, 100 μg / ml RNase A was added, and the solution was rocked at 37°C for 10 minutes. Next, 2 volumes of phenol:chloroform:isoamyl alcohol were added, and the solution was rocked at room temperature for 1 hour, followed by centrifugation at 8000 × g for 15 minutes. The supernatant was transferred to a clean test tube. Two volumes of phenol:chloroform:isoamyl alcohol were added again, and the solution was rocked at room temperature for 1 hour. It was then centrifuged at 8000 × g for 15 minutes, and the supernatant was transferred to a clean test tube. An equal volume of chloroform was added to the resulting supernatant, and the solution was rocked at room temperature for 30 minutes. The solution was centrifuged at 8000 × g for 15 minutes, and the supernatant was transferred to a clean test tube. An equal volume of chloroform was added to the resulting supernatant, and the solution was rocked at room temperature for 30 minutes. The solution was centrifuged at 8000 × g for 15 minutes, and the supernatant was transferred to a clean test tube. 0.3 volumes of 3 M NaOAc and 2 volumes of 100% EtOH were added to the supernatant, and the mixture was gently rocked for several minutes. The DNA was spooled onto a sterile glass rod and immersed in 70% EtOH for 1-2 minutes. The DNA was transferred to a 1.7 ml microcentrifuge tube and air-dried for 10 minutes. Up to 0.5 ml of pre-washed EB was added to the DNA and placed at 4°C overnight.

[0152] Cryostocks of transgenic Schizochytrium (transformed with pCL0143) were grown to confluence in M50-20 and then propagated in 50 mL baffled shake flasks for 48 hours (h) at 27°C and 200 rpm in medium containing the following (per liter) unless otherwise specified: TIFF2025114812000002.tif64128

[0153] The volume was brought to 900 mL with deionized HO, and the pH was adjusted to 6.5 before autoclaving for 35 minutes unless otherwise specified. Filter-sterilized glucose (50 g / L), vitamins (2 mL / L), and trace metals (2 mL / L) were then added to the medium, and the volume was adjusted to 1 liter. The vitamin solution contained 0.16 g / L vitamin B12, 9.75 g / L thiamine, and 3.33 g / L calcium pantothenate. The trace metals solution (pH 2.5) contained 1.00 g / L citric acid, 5.15 g / L FeSO4.7H2O, 1.55 g / L MnCl2.4H2O, 1.55 g / L ZnSO4.7H2O, 0.02 g / L CoCl2.6H2O, 0.02 g / L Na2MoO4.2H2O, 1.035 g / L CuSO4.5H2O, and 1.035 g / L NiSO4.6H2O.

[0154] The Schizochytrium cultures were transferred to 50 ml conical tubes and centrifuged for 15 minutes at 3000 x g or 4500 x g. See Figure 3. The supernatant obtained from this centrifugation, termed "cell-free supernatant" (CFS), was used for immunoblot analysis and hemagglutination activity assays.

[0155] The cell-free supernatant (CFS) was further ultracentrifuged at 100,000 × g for 1 hour (see Figure 3). The resulting pellet (insoluble fraction or "UP") containing HA protein was resuspended in PBS, pH 7.4. This suspension was centrifuged (120,000 × g, 18 hours, 4°C) on a discontinuous sucrose density gradient containing 15-60% sucrose solutions (see Figure 3). The 60% sucrose fraction containing HA protein was used for peptide sequence analysis, glycosylation analysis, and electron microscopy analysis.

[0156] Immunoblot analysis Expression of recombinant HA protein from transgenic Schizochytrium CL0143-9 ("E") was verified by immunoblot analysis according to standard immunoblotting procedures. Proteins from cell-free supernatants (CFS) were separated by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) on NuPAGE® Novex® 12% Bis-Tris gels (Invitrogen, Carlsbad, CA) under reducing conditions with MOPS SDS running buffer unless otherwise specified. Proteins were then stained with Coomassie blue (SimplyBlue Safe Stain, Invitrogen, Carlsbad, CA) or transferred to polyvinylidene difluoride membranes and probed for the presence of HA protein with rabbit anti-influenza A / Puerto Rico / 8 / 34 (H1N1) virus antiserum (1:1000 dilution, a gift from Dr. Albert DME Osterhaus; Fouchier RAM et al., J. Virol. 79:2814-2822 (2005)) followed by alkaline phosphatase-conjugated anti-rabbit IgG (Fc) secondary antibody (1:2000 dilution, #S3731, Promega Corporation, Madison, WI). The membranes were then treated with 5-bromo-4-chloro-3-indolyl phosphate / nitroblue tetrazolium solution (BCIP / NBT) according to the manufacturer's instructions (KPL, Gaithersburg, MD). Anti-H1N1 immunoblots against transgenic Schizochytrium CL0143-9 ("E") grown at various pHs (5.5, 6.0, 6.5, and 7.0) and temperatures (25°C, 27°C, and 29°C) are shown in Figure 4A. The negative control ("C") was the wild-type strain of Schizochytrium species ATCC 20888. Recombinant HA protein was detected in the cell-free supernatant at pH 6.5 (Figure 4A), and the hemagglutination activity detected was highest at pH 6.5 and 27°C (Figure 4A).A Coomassie blue-stained gel ("Coomassie") and corresponding anti-H1N1 immunoblot ("IB: anti-H1N1") for CL0143-9 ("E") grown at pH 6.5, 27°C, under non-reducing and reducing conditions is shown in Figure 4B. The negative control ("C") was a wild-type strain of Schizochytrium species ATCC 20888.

[0157] HA activity The activity of HA protein produced in Schizochytrium was assessed by hemagglutination assay. Functional HA protein exhibits hemagglutination activity, which is easily detected by standard hemagglutination assays. Briefly, 50 μL of a two-fold dilution of the low-speed supernatant in PBS was prepared in a 96-well microtiter plate. An equal volume of approximately 1% chicken red blood cell solution (Fitzgerald Industries, Acton, MA) in PBS, pH 7.4, was then added to each well, followed by incubation at room temperature for 30 minutes. The degree of agglutination was then analyzed visually. The hemagglutination activity unit (HAU) is defined as the highest dilution that caused visible hemagglutination in the well.

[0158] Typical activity was found to be approximately 512 HAU in the cell-free supernatant of transgenic Schizochytrium CL0143-9 ("E") (Figure 5A). PBS ("-") or the wild-type strain of Schizochytrium ATCC 20888 ("C"), grown in the same manner as the transgenic strain, were used as negative controls and did not exhibit any hemagglutination activity. Recombinant HA protein from influenza A / Vietnam / 1203 / 2004 (H5N1) (Protein Sciences Corporation, Meriden, CT, diluted 1:1000 in PBS) was used as a positive control ("+").

[0159] Analysis of the soluble and insoluble fractions of the cell-free supernatant of the transgenic Schizochytrium strain CL0143-9 by hemagglutination assay suggested that the HA protein was found predominantly in the insoluble fraction (Fig. 5B). Typical activities were found to be approximately 16 HAU in the soluble fraction ("US") and approximately 256 HAU in the insoluble fraction ("UP").

[0160] The activity levels of HA protein in 2 L cultures showed similar activity to shake flask cultures grown in the same medium at a constant pH of 6.5.

[0161] In another experiment, the native signal peptide of HA was removed and replaced with the Schizochytrium Sec1 signal peptide (SEQ ID NO: 37, encoded by SEQ ID NO: 38). Transgenic Schizochytrium obtained with this alternative construct exhibited hemagglutination activity and recombinant protein distribution similar to that observed with transgenic Schizochytrium containing the pCL0143 construct (data not shown).

[0162] Peptide sequence analysis The insoluble fraction ("UP") obtained from 100,000 × g centrifugation of the cell-free supernatant was further fractionated against a sucrose density gradient. Fractions containing HA protein, as indicated by hemagglutination activity assays (Figure 6B), were separated by SDS-PAGE as described above, stained with Coomassie Blue, or transferred to PVDF and immunoblotted with anti-H1N1 antiserum from rabbit (Figure 6A). Bands corresponding to cross-reactivity in the immunoblot (HA1 and HA2) were excised from the Coomassie Blue-stained gel and subjected to peptide sequence analysis. Briefly, bands of interest were washed and destained in 50% ethanol and 5% acetic acid. The gel pieces were then dehydrated in acetonitrile, dried in a SpeedVac® (Thermo Fisher Scientific, Inc., Waltham, MA), and digested with trypsin by adding 5 μL of 10 ng / μL trypsin in 50 mM ammonium bicarbonate and incubating overnight at room temperature. The formed peptides were extracted from the polyacrylamide gel into two 30 μL aliquots of 50% acetonitrile containing 5% formic acid. These extracts were combined, evaporated to less than 10 μL in a SpeedVac®, and then resuspended in 1% acetic acid to a final volume of approximately 30 μL for LC-MS analysis. The LC-MS system was a Finnigan™ LTQ™ Linear Ion Trap Mass Spectrometer (Thermo Electron Corporation, Waltham, MA). The HPLC column was a self-packed 9 cm × 75 μm Phenomenex Jupiter™ C18 reverse-phase capillary chromatography column (Phenomenex, Torrance, CA). A μL volume of the extract was then injected, and the peptides were eluted from the column with an acetonitrile / 0.1% formic acid gradient at a flow rate of 0.25 μL / min and introduced into the online mass spectrometer source. The microelectrospray ion source was operated at 2.5 kV. The digests were analyzed using selected reaction mixture (SRM) experiments in which a series of m / z ratios are subdivided by the mass spectrometer over the course of the LC experiment.The fragmentation patterns of the peptides of interest were then used to generate chromatograms. Peak areas for each peptide were determined and normalized to an internal standard. The internal standard used in this analysis was a protein with consistent abundance across the samples tested. Final comparisons between the two systems were determined by comparing the normalized peak ratios for each peptide. The collision-induced dissociation spectra were then searched against the NCBI database. HA protein was identified by a total of 27 peptides, covering 42% of the protein sequence. The specific sequenced peptides are highlighted in bold in Figure 7. More specifically, HA1 was identified by a total of 17 peptides, and HA2 by a total of 9 peptides. This is consistent with an HA N-terminal polypeptide being truncated before position 397. The locations of the peptides identified for HA1 and HA2 are shown within the full amino acid sequence of the HA protein. The putative cleavage site within HA is located between amino acids 343 and 344 (shown as R^G). The italicized peptide sequence beginning at amino acid 402 is related to the HA2 polypeptide but is found in peptides identified in HA1, likely due to minor incorporation of the HA2 peptide in the band excised for HA1 (see, e.g., Figure 3 in Wright et al., BMC Genomics 10:61 (2009)).

[0163] Glycosylation analysis The presence of glycans on the HA protein was assessed by enzymatic treatment. A 60% sucrose fraction of transgenic Schizochytrium CL0143-9 was digested with EndoH or PNGase F according to the manufacturer's instructions (New England Biolabs, Ipswich, MA). Proteins were then separated by SDS-PAGE on NuPAGE Novex 12% Bis-Tris gels (Invitrogen, Carlsbad, CA) using MOPS SDS running buffer. Glycan removal was identified by a shift in expected mobility following staining with Coomassie Blue ("Coomassie") or immunoblotting with anti-H1N1 antiserum ("IB: anti-H1N1") (Figure 8). The negative control for enzymatic treatment was transgenic Schizochytrium CL0143-9 incubated without enzyme ("NT" = no treatment). At least five different species could be identified at the HA1 position on the immunoblot, and two different species could be identified at the HA2 position on the immunoblot, consistent with multiple glycosylation sites on HA1 and a single glycosylation site on HA2, as reported in the literature.

[0164] Example 3 Characterization of proteins from Schizochytrium culture supernatants Schizochytrium sp. ATCC 20888 was grown under typical fermentation conditions as described above. Samples of culture supernatant were collected at 4-hour intervals from 20 to 52 hours of culture, with a final collection at 68 hours.

[0165] Total protein in the culture supernatant for each sample was determined by standard Bradford Assay. See Figure 9.

[0166] Proteins were isolated from culture supernatant samples at 37, 40, 44, 48, and 68 hours using the method in Figure 3. An SDS-PAGE gel of the proteins is shown in Figure 10. Lane 11 was loaded with 2.4 μg of total protein, and the remaining lanes were loaded with 5 μg of total protein. The abundant bands identified (by mass spectrometry peptide sequencing) as actin or gelsolin are indicated by arrows in Figure 10.

[0167] Example 4 Negative staining and electron microscopy of culture supernatant material Schizochytrium sp. ATCC 20888 (control) and transgenic Schizochytrium sp. CL0143-9 (experimental) were grown under typical flask conditions as described above. Cultures were transferred to 50 mL conical tubes and centrifuged at 3000 × g or 4500 × g for 15 minutes. The cell-free supernatant was further ultracentrifuged at 100,000 × g for 1 hour, and the resulting pellet was resuspended in PBS, pH 7.4. This suspension was centrifuged on a discontinuous 15% to 60% sucrose gradient (120,000 × g, 18 hours, 4°C), and the 60% fraction was used for negative staining and electron microscopy.

[0168] Electron microscopy of negatively stained control material revealed a mixture of membrane fragments, membrane aggregates, and vesicles (collectively "exobodies") ranging in diameter from several hundred nanometers to <50 nm (see Figure 11). Vesicle shapes ranged from circular to elongated (tubular), and vesicle edges were smooth or irregular. The interiors of the vesicles appeared faintly stained, suggesting the presence of organic material. Larger vesicles had thickened membranes, suggesting that the vesicle edges overlapped during preparation. The membrane aggregates and fragments were highly irregular in shape and size. The membrane material likely originated from the ectoplasmic net, as indicated by its strong correlation with actin in membranes purified by ultracentrifugation.

[0169] Similarly, electron microscopy of negatively stained material from cell-free supernatants of cultures of transgenic Schizochytrium CL0143-9 expressing heterologous proteins suggested that this material was a mixture of membrane fragments, membrane aggregates, and vesicles ranging in diameter from hundreds of nanometers to <50 nm. See Figure 11.

[0170] Immunolocalization was also performed on this material using the H1N1 antiserum and 12 nm gold particles described for immunoblot analysis in Example 22, as described by Perkins et al., J. Virol. 82:7201-7211 (2008). Extracellular membrane bodies isolated from transgenic Schizochytrium CL0143-9 were heavily decorated with gold particles attached to the antiserum (Figure 12), suggesting that the antibody recognizes HA protein present in the extracellular bodies. Minimal background was observed in areas without membrane material. Few or no gold particles were bound to extracellular bodies isolated from control material (Figure 12).

[0171] Example 5 Construction of xylose transporter, xylose isomerase, and xylulose kinase expression vectors Vector pAB0018 (ATCC Accession No. PTA-9616) was digested with HindIII, treated with mung bean nuclease, purified, and then further digested with KpnI to generate four fragments of varying sizes. The 2552 bp fragment was isolated by standard electrophoresis in an agarose gel and purified using a commercially available DNA purification kit. pAB0018 was then subjected to a second digestion with PmeI and KpnI. A 6732 bp fragment was isolated, purified from this digest, and ligated to the 2552 bp fragment. The ligation product was then used to transform a commercially available strain of competent DH5-α E. coli cells (Invitrogen) using the manufacturer's protocol. Plasmids from ampicillin-resistant clones were propagated, purified, and then screened by restriction digestion or PCR to confirm that the expected plasmid structure had been generated by the ligation. One of the verified plasmids was designated pCL0120. Please refer to Figure 15.

[0172] Sequences encoding the Candida intermedia xylose transporter protein GXS1 (GenBank Accession No. AJ875406) and the Arabidopsis thaliana xylose transporter protein At5g17010 (GenBank Accession No. BT015128) were codon-optimized and synthesized (Blue Heron Biotechnology, Bothell, WA) as guided by the Schizochytrium codon usage table shown in Figure 16. SEQ ID NO: 94 is the codon-optimized nucleic acid sequence of GSX1, while SEQ ID NO: 95 is the codon-optimized nucleic acid sequence of At5g17010.

[0173] SEQ ID NO: 94 and SEQ ID NO: 95 were cloned into pCL0120 using the 5' and 3' restriction sites BamHI and NdeI, respectively, for insertion and ligation by standard techniques. Maps of the resulting vectors, pCL0130 and pCL0131, are shown in Figures 17 and 18, respectively.

[0174] Vectors pCL0121 and pCL0122 were created by ligating a 5095 bp fragment released from pCL0120 by digestion with HindIII and KpnI to synthetic selectable marker cassettes designed to confer resistance to either zeocin or paromomycin. These cassettes consisted of an alpha tubulin promoter to drive expression of either the sh ble gene (for zeocin) or the npt gene (for paromomycin). Transcription products of both selectable marker genes were terminated by an SV40 terminator. The complete sequences of vectors pCL0121 and pCL0122 are provided as SEQ ID NO: 90 and SEQ ID NO: 91, respectively. Maps of vectors pCL0121 and pCL0122 are shown in Figures 19 and 20, respectively.

[0175] Sequences encoding Piromyces sp. E2 xylose isomerase (CAB76571) and Piromyces sp. E2 xylulose kinase (AJ249910) were codon-optimized and synthesized (Blue Heron Biotechnology, Bothell, WA) as guided by the Schizochytrium codon usage table shown in Figure 16. "XylA" (SEQ ID NO: 92) is the codon-optimized nucleic acid sequence of CAB76571 (Figure 21), while "XylB" (SEQ ID NO: 93) is the codon-optimized nucleic acid sequence of AJ249910 (Figure 22).

[0176] SEQ ID NO: 92 was cloned into the vector pCL0121 to give a vector designated pCL0132 (Figure 23), and SEQ ID NO: 21 was cloned into the vector pCL0122 by insertion into the BamHI and NdeI sites to give a vector designated pCL0136 (Figure 24).

[0177] Example 6 Expression and characterization of xylose transporter, xylose isomerase, and xylulose kinase proteins produced in Schizochytrium sp. Schizochytrium sp. ATCC 20888 was used as a host cell for transformation with vectors pCL0130, pCL0131, pCL0132 or pCL0136, respectively.

[0178] Electroporation with enzyme pretreatment Cells were grown in 50 mL of M50-20 medium (see U.S. Patent Application Publication No. 2008 / 0022422) on a shaker at 200 rpm at 30°C for 2 days. Cells were diluted 1:100 into M2B medium (see next paragraph) and grown overnight (16-24 hours) until they reached mid-log phase growth (OD 600 The cells were centrifuged in a 50 ml conical tube at 3000 × g for 5 minutes. The supernatant was removed and the cells were resuspended in an appropriate volume of 1 M mannitol, pH 5.5, and incubated at 2 OD for 1 hour. 600A final concentration of 1000 μg was reached. Five mL of cells were dispensed into a 25 mL shake flask and modified with 10 mM CaCl2 (1.0 M stock, filter-sterilized) and 0.25 mg / mL protease XIV (10 mg / mL stock, filter-sterilized; Sigma-Aldrich, St. Louis, MO). The flask was incubated on a shaker at 100 rpm and 30°C for 4 hours. The cells were monitored under a microscope, and the degree of protoplasting, as desired, was determined. The cells were centrifuged at 2500 × g for 5 minutes in a round-bottom tube (i.e., a 14 mL Falcon™ tube, BD Biosciences, San Jose, CA). The supernatant was removed, and the cells were gently resuspended in 5 mL of ice-cold 10% glycerol. The cells were centrifuged again in the round-bottom tube at 2500 × g for 5 minutes. The supernatant was removed, and the cells were gently resuspended in 500 μL of ice-cold 10% glycerol using a large-bore pipette tip. 90 μL of cells were dispensed into a pre-chilled electrocuvette (Gene Pulser® cuvette - 0.2 cm gap, Bio-Rad, Hercules, CA). 1 μg to 5 μg of DNA (in a volume of 10 μL or less) was added to the cuvette, mixed gently with a pipette tip, and placed on ice for 5 minutes. The cells were electroporated at 200 ohms (resistance), 25 μF (capacitance), and 500 V. 0.5 mL of M50-20 medium was quickly added to the cuvette. The cells were then transferred to 4.5 mL of M50-20 medium in a 25 mL shake flask and incubated on a shaker at 100 rpm and 30°C for 2 to 3 hours. The cells were centrifuged at 2500 × g for 5 minutes in a round-bottom tube. The supernatant was removed and the cell pellet was resuspended in 0.5 mL of M50-20 medium. The cells were plated onto an appropriate number (2–5) of M2B plates with appropriate selection (if necessary) and incubated at 30°C.

[0179] M2B medium consisted of 10 g / L glucose, 0.8 g / L (NH4)2SO4, 5 g / L Na2SO4, 2 g / L MgSO4·7H2O, 0.5 g / L KH2PO4, 0.5 g / L KCl, 0.1 g / L CaCl2·2H2O, 0.1 M MES (pH 6.0), 0.1% PB26 metals, and 0.1% PB26 vitamins (v / v), which consisted of 50 mg / mL vitamin B12, 100 μg / mL thiamine, and 100 μg / mL calcium pantothenate. PB26 metals was adjusted to pH 4.5 and consisted of 3 g / L FeSO4·7H2O, 1 g / L MnCl2·4H2O, 800 mg / mL ZnSO4·7H2O, 20 mg / mL CoCl2·6H2O, 10 mg / mL Na2MoO4·2H2O, 600 mg / mL CuSO4·5H2O, and 800 mg / mL NiSO4·6H2O. PB26 stock solutions were filter-sterilized separately and added to the culture medium after autoclaving. Glucose, KH2PO4, and CaCl2·2H2O were each autoclaved separately from the rest of the culture medium components before mixing to prevent salt precipitation and caramelization of the carbohydrates. All medium components were purchased from Sigma Chemical (St. Louis, MO).

[0180] Transformants were selected for growth on solid medium containing the appropriate antibiotic. Twenty to one hundred primary transformants of each vector were replated onto "xylose-SSFM" solid medium, which is identical to SSFM (described below) except that it contains xylose instead of glucose as the sole carbon source; no antibiotic was added. No growth was observed for any clones under these conditions.

[0181] SSFM medium: 50 g / L glucose, 13.6 g / L NaSO, 0.7 g / L KSO, 0.36 g / L KCl, 2.3 g / L MgSO 7H O, 0.1 M MES (pH 6.0), 1.2 g / L (NH) SO, 0.13 g / L sodium glutamate, 0.056 g / L KH PO, and 0.2 g / L CaCl 2H O. Vitamins were added at 1 mL / L from a stock solution consisting of 0.16 g / L vitamin B, 9.7 g / L thiamine, and 3.3 g / L calcium pantothenate. Trace metals were added at 2 mL / L from a stock solution consisting of 1 g / L citric acid, 5.2 g / L FeSO4·7H2O, 1.5 g / L MnCl2·4H2O, 1.5 g / L ZnSO4·7H2O, 0.02 g / L CaCl2·6H2O, 0.02 g / L Na2MoO4·2H2O, 1.0 g / L CuSO4·5H2O, and 1.0 g / L NiSO4·6H2O adjusted to pH 2.5.

[0182] gDNA from primary transformants of pCL0130 and pCL0131 was extracted, purified, and used as a template for PCR to confirm the presence of the transgene.

[0183] Genomic DNA extraction was performed as described in Example 2.

[0184] Alternatively, after RNase A incubation, DNA was further purified using a Qiagen Genomic Tip 500 / G column (Qiagen, Inc USA, Valencia, CA) according to the manufacturer's protocol.

[0185] PCR - The primers used to detect the GXS1 transgene were: TIFF2025114812000003.tif14153. The primers used to detect the At5g17010 transgene were The file was TIFF2025114812000004.tif12144.

[0186] pCL0130, pCL0132, and pCL0136 ("pCL01310 series") were combined together, or pCL0131, pCL0132, and pCL0136 ("pCL0131 series") were combined together and used to cotransform a wild-type strain of Schizochytrium (ATCC 20888). Transformants were plated directly onto solid xylose SSFM medium, and after 3–5 weeks, colonies were selected and further propagated in liquid xylose-SSFM. Several rounds of serial transfers in liquid medium containing xylose improved the growth rate of the transformants. Cotransformants of the pCL0130 series or pCL0131 series were also plated onto solid SSFM medium containing either SMM, zeocin, or paromomycin. All transformants plated on these media were resistant to each of the antibiotics tested, suggesting that the transformants possessed all three of these respective vectors. Schizochytrium transformed with the xylose transporter, xylose isomerase, and xylulose kinase were able to grow in medium containing xylose as the sole carbon source.

[0187] In future experiments, Western blots of both cell-free extracts and cell-free supernatants from shake flask cultures of selected SMM-resistant transformant clones (pCL0130 or pCL0131 transformants alone, or pCL0130-series cotransformants, or pCL0131-series cotransformants) will be performed to show that both transporters are expressed and found in both fractions, indicating that these membrane-associated proteins are associated with extracellular vesicles, similar to the observations made with the other membrane proteins described herein. Alternatively, Western blots will be performed to show the expression of xylose isomerase and xylulose kinase in the cell-free extracts of all clones where their presence is expected. Extracellular bodies, such as vesicles containing xylose transporters, can be used to deplete trace amounts of xylose from media containing mixtures of sugars or other low-molecular-weight solutes by trapping the sugars within vesicles that can be separated by various methods, including filtration or centrifugation.

[0188] Example 7 Construction of pCL0140 and pCL0149 expression vectors The vector pCL0120 was digested with BamHI and NdeI, generating two fragments of 837 base pairs (bp) and 8454 bp in length. The 8454 bp fragment was fractionated by standard electrophoresis techniques in agarose gels, purified using a commercially available DNA purification kit, and ligated to a synthetic sequence (SEQ ID NO: 100 or SEQ ID NO: 101; see Figure 26) that had also been pre-digested with BamHI and NdeI. SEQ ID NO: 100 (Figure 26) encodes the NA protein of influenza A virus (A / Puerto Rico / 8 / 34 / Mount Sinai (H1N1)). The protein sequence matches that of GenBank accession number NP_040981. The specific nucleic acid sequence of SEQ ID NO: 100 was codon-optimized and synthesized by DNA 2.0 for expression in Schizochytrium as guided by the Schizochytrium codon usage table shown in Figure 16. SEQ ID NO: 101 (Figure 26) encodes the same NA protein as SEQ ID NO: 100, but contains a V5 tag sequence and a polyhistidine sequence at the C-terminus of the coding region.

[0189] The ligation products were then used to transform commercially available strains of competent DH5-α E. coli cells (Invitrogen, Carlsbad, CA) using the manufacturer's protocol. These plasmids were then screened by restriction digestion or PCR to confirm that the ligation had produced the expected plasmid structure. The plasmid vectors resulting from this procedure were verified using Sanger sequencing by DNA 2.0 (Menlo Park, CA) and designated pCL0140 (Figure 25A) containing SEQ ID NO: 100 and pCL0149 (Figure 25B) containing SEQ ID NO: 101. The pCL0140 and pCL0149 vectors contain a promoter from the Schizochytrium elongation factor-1 gene (EF1) to drive expression of the NA transgene, the OrfC terminator (also known as the PFA3 terminator) following the NA transgene, and a selectable marker cassette conferring resistance to sulfometuron methyl.

[0190] Example 8 Expression and characterization of NA proteins produced in Schizochytrium Schizochytrium sp. ATCC 20888 was used as a host cell for transformation with vectors pCL0140 and pCL0149 using a Biolistic™ particle bombarder (BioRad, Hercules, CA) as described in Example 2. Transformants were selected for growth on solid medium containing the appropriate antibiotic. gDNA from primary transformants was extracted, purified, and used as a PCR template to test for the presence of the transgene, as previously described (Example 2).

[0191] Cryopreserved stocks of transgenic Schizochytrium (transformed with pCL0140 and pCL0149) were grown to confluence in M50-20 as described in Example 2 and then propagated in 50 mL baffled shake flasks.

[0192] The Schizochytrium culture was transferred to a 50 ml conical tube and centrifuged at 3000 × g for 15 minutes (see Figure 27). The supernatant obtained from this centrifugation was designated the "cell-free supernatant" (CFS). This CFS fraction was concentrated 50-100 times using a Centriprep™ gravity concentrator (Millipore, Billerica, MA) and designated the "concentrated cell-free supernatant" (cCFS). The cell pellet obtained from the centrifugation was washed in water, frozen in liquid nitrogen, and then resuspended in 2x the pellet weight of lysis buffer (50 mM sodium phosphate (pH 7.4), 1 mM EDTA, 5% glycerol, and 1 mM fresh phenylmethylsulfonyl fluoride) and 2x the pellet weight of 0.5 mm glass beads (Sigma, St. Louis, MO). The cell pellet mixture was then lysed by vortexing at maximum speed in a multitube vortexer (VWR, Westchester, PA) for 3 hours at 4°C. The resulting cell lysate was then centrifuged at 5500 × g for 10 minutes at 4°C. The resulting supernatant was retained and centrifuged again at 5500 × g for 10 minutes at 4°C. The resulting supernatant is defined herein as the "cell-free extract" (CFE). Protein concentrations were determined in cCFS and CFE by a standard Bradford assay (Bio-Rad, Hercules, CA). These fractions were used for neuraminidase activity assays and immunoblot analysis.

[0193] Functional influenza NA proteins exhibit neuraminidase activity that can be detected by a standard fluorometric NA activity assay, based on sialidase-mediated hydrolysis of the (4-methylumbelliferyl)-α-DN-acetylneuraminic acid sodium salt (4-MUNANA) substrate (Sigma-Aldrich, St. Louis, MO), which yields free 4-methylumbelliferone with fluorescence emission at 450 nm following excitation at 365 nm. Briefly, CFS, cCFS, or CFE of transgenic Schizochytrium strains were assayed according to the procedure described by Potier et al., Anal. Biochem. 94:287-296 (1979), using 25 μL of CFS and 75 μL of 40 μM 4-MUNANA or 75 μL of ddH2O for the control. The reaction was incubated at 37°C for 30 min, and fluorescence was measured on a FLUOstar Omega multimode microplate reader (BMG LABTECH, Offenburg, Germany).

[0194] Typical activity observed in concentrated cell-free supernatants (cCFS) and cell-free extracts (CFE) from nine transgenic strains of Schizochytrium transformed with CL0140 is presented in Figure 28. A wild-type strain of Schizochytrium species ATCC 20888 ("-") and a PCR-negative strain of Schizochytrium ("27") prepared by transformation with pCL0140 and growth in the same manner as the transgenic strains were used as negative controls. The majority of activity was found in the concentrated cell-free supernatants, suggesting successful expression and secretion of functional influenza neuraminidase into the external environment by Schizochytrium.

[0195] Peptide sequence analysis The transgenic Schizochytrium strain CL0140-26 was used for partial purification of influenza NA protein, and its successful expression and secretion were confirmed by peptide sequence analysis. The purification procedure was adapted from Tarigan et al., JITV 14(1):75-82 (2008), followed by measuring NA activity as described above (Figure 29A). Briefly, the cell-free supernatant of the transgenic strain CL0140-26 was further centrifuged at 100,000 × g for 1 hour at 4°C. The resulting supernatant was concentrated 100-fold using a Centriprep™ gravity concentrator (Millipore, Billerica, MA) (fraction "cCFS" in Figure 29A) and diluted back to the original volume with 0.1 M sodium bicarbonate buffer (pH 9.1) containing 0.1% Triton X-100 (fraction "D" in Figure 29A). This diluted sample was used for affinity chromatography purification. N-(p-aminophenyl)oxamic acid agarose (Sigma-Aldrich, St. Louis, MO) was packed into a PD-10 column (BioRad, Hercules, CA) that had been activated by washing with 6 column volumes (CV) of 0.1 M sodium bicarbonate buffer (pH 9.1) containing 0.1% Triton X-100, followed by 5 CV of 0.05 M sodium acetate buffer, pH 5.5, containing 0.1% Triton X-100. The diluted sample (fraction "D") was loaded onto the column; unbound material was removed by washing the column with 10 CV of 0.15 M sodium acetate buffer containing 0.1% Triton X-100 (fraction "W" in Figure 29A). The bound NAs were eluted from the column with 5 CV of 0.1 M sodium bicarbonate buffer containing 0.1% Triton X-100 and 2 mM CaCl (fraction "E" in Figure 29A). The NA-rich fraction E solution was concentrated to approximately 10% of the original volume using a 10 kDa molecular cutoff spin concentrator, resulting in fraction cE.

[0196] Proteins from each fraction were separated by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) on a NuPAGE® Novex® 12% Bis-Tris gel (Invitrogen, Carlsbad, CA) under reducing conditions using MOPS SDS running buffer. These proteins were then stained with Coomassie Blue (SimplyBlue Safe Stain, Invitrogen, Carlsbad, CA). Protein bands visible in lane "cE" (Figure 29B) were excised from the Coomassie Blue-stained gel, and peptide sequence analysis was performed as described in Example 2. The protein band containing the NA protein (indicated by an arrow in lane "cE") was identified by a total of nine peptides (113 amino acids) covering 25% of the protein sequence. The specific peptides that were sequenced are highlighted in bold in Figure 30.

[0197] Immunoblot analysis Expression of recombinant NA protein from transgenic Schizochytrium CL0149 (clones 10, 11, and 12) was examined by immunoblot analysis according to standard immunoblotting procedures (Figure 31B). Proteins from cell-free supernatants (CFS) were separated by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) on NuPAGE® Novex® 12% Bis-Tris gels (Invitrogen, Carlsbad, CA) under reducing conditions using MOPS SDS running buffer. These proteins were then stained with Coomassie Blue (SimplyBlue Safe Stain, Invitrogen, Carlsbad, CA) or transferred to polyvinylidene difluoride membranes and probed for the presence of NA protein with anti-V5-AP-binding mouse monoclonal antibody (1:1000 dilution, #962-25, Invitrogen, Carlsbad, CA). The membrane was then treated with 5-bromo-4-chloro-3-indolyl-phosphate / nitroblue tetrazolium solution (BCIP / NBT) according to the manufacturer's instructions (KPL, Gaithersburg, MD). Recombinant NA protein was detected in the cell-free supernatant of clone 11 (Figure 31B). The negative control ("-") was a wild-type strain of Schizochytrium sp. ATCC 20888. The positive control ("+") was Positope™ antibody control protein (#R900-50, Invitrogen, Carlsbad, CA). The corresponding neuraminidase activity is presented in Figure 31A.

[0198] Example 9 Co-expression of influenza HA and NA in Schizochytrium Schizochytrium sp. ATCC 20888 was used as a host cell for co-transformation with vectors pCL0140 (FIG. 25A) and pCL0143 (FIG. 2) using a Biolistic™ particle bombarder (BioRad, Hercules, CA) as described in Example 2.

[0199] Cryopreserved stocks of transgenic Schizochytrium (transformed with pCL0140 and pCL0143) were cultured and processed as described in Example 2. Hemagglutination and neuraminidase activities were measured as described in Examples 2 and 7, respectively, and are shown in Figure 32. Transgenic Schizochytrium transformed with pCL0140 and pCL0143 exhibited HA and NA associated activities.

[0200] Example 10 Expression and characterization of extracellular bodies containing parainfluenza F protein produced in Schizochytrium Schizochytrium species ATCC 20888 was used as a host cell for transformation with a vector containing a sequence encoding the F protein (GenBank accession number P06828) of human parainfluenza 3 virus strain NIH 47885. A representative sequence of the F protein is provided as SEQ ID NO: 102. Some cells were transformed with a vector containing a sequence encoding the native signal peptide sequence associated with the F protein. Other cells were transformed with a vector containing a sequence encoding a different signal peptide sequence (e.g., the Schizochytrium signal anchor sequence) fused to the F protein-encoding sequence so that the F protein is expressed with the heterologous signal peptide sequence. Other cells were transformed with a vector containing a sequence encoding a different membrane domain (e.g., the HA membrane domain) fused to the F protein-encoding sequence so that the F protein is expressed with the heterologous membrane domain. The F protein contains a single transmembrane domain near the C-terminus. The F protein can be separated into two peptides at the furin cleavage site (amino acid 109). The first part of the protein, designated F2, contains the N-terminal portion of the complete F protein. The F2 region may be fused separately to a sequence encoding a heterologous signal peptide. The remainder of the viral F protein, including the C-terminal portion of the F protein, is designated F1. The F1 region may be fused separately to a sequence encoding a heterologous signal peptide. Vectors containing the F1 and F2 portions of the viral F protein may be expressed separately or in combination. A vector expressing the complete F protein may be coexpressed with a furin enzyme that cleaves the protein at the furin cleavage site. Alternatively, the sequence encoding the furin cleavage site of the F protein may be replaced with a sequence encoding another protease cleavage site that is recognized and cleaved by a different protease. An F protein containing another protease cleavage site may be coexpressed with a corresponding protease that recognizes and cleaves the other protease cleavage site.

[0201] Transformation is performed, and frozen stocks are grown and propagated according to one of the methods described herein. The Schizochytrium culture is transferred to a 50 mL conical tube and centrifuged at 3000 x g or 4500 x g for 15 minutes to obtain a low-speed supernatant. The low-speed supernatant is further ultracentrifuged at 100,000 x g for 1 hour. See Figure 3. The resulting pellet of insoluble fraction containing the F protein is resuspended in phosphate-buffered saline (PBS) and used for peptide sequence analysis and glycosylation analysis as described in Example 2.

[0202] Expression of the F protein from transgenic Schizochytrium is verified by immunoblot analysis using anti-F antiserum and secondary antibodies at appropriate dilutions according to standard immunoblotting procedures as described in Example 2. The recombinant F protein is detected in the low-speed supernatant and insoluble fraction. Additionally, the recombinant F protein is detected in cell-free extracts from transgenic Schizochytrium expressing the F protein.

[0203] The activity of the F protein produced in Schizochytrium is assessed by an F activity assay. A functional F protein exhibits F activity that is readily detected by a standard F activity assay.

[0204] Electron microscopy is performed to confirm the presence of extracellular bodies using negatively stained material prepared according to Example 4. Immunogold labeling is performed to confirm protein association with the extracellular membrane bodies.

[0205] Example 11 Expression and characterization of extracellular bodies containing vesicular stomatitis (G Vesicular Stomatitus) virus G protein produced in Schizochytrium Schizochytrium species ATCC 20888 is used as a host cell for transformation with a vector containing a sequence encoding the vesicular stomatitis virus G (VSV-G) protein. A representative sequence for the VSV-G protein is provided as SEQ ID NO: 103 (from GenBank accession number M35214). Some cells are transformed with a vector containing a sequence encoding the native signal peptide sequence associated with the VSV-G protein. Other cells are transformed with a vector containing a sequence encoding a different signal peptide sequence (e.g., a Schizochytrium signal anchor sequence) fused to the VSV-G protein-encoding sequence so that the VSV-G protein is expressed with the heterologous signal peptide sequence. Other cells are transformed with a vector containing a sequence encoding a different membrane domain (e.g., an HA membrane domain) fused to the VSV-G protein-encoding sequence so that the VSV-G protein is expressed with the heterologous membrane domain. Transformations are performed, and cryopreserved stocks are grown and propagated according to any of the methods described herein. The Schizochytrium culture is transferred to a 50 mL conical tube and centrifuged at 3000 x g or 4500 x g for 15 minutes to obtain a low-speed supernatant. The low-speed supernatant is further ultracentrifuged at 100,000 x g for 1 hour. See Figure 3. The resulting pellet of insoluble fraction containing the VSV-G protein is resuspended in phosphate-buffered saline (PBS) and used for peptide sequence analysis and glycosylation analysis as described in Example 2.

[0206] Expression of VSV-G protein from transgenic Schizochytrium is verified by immunoblot analysis using anti-VSV-G antiserum and secondary antibodies at appropriate dilutions according to standard immunoblotting procedures as described in Example 2. Recombinant VSV-G protein is detected in the low-speed supernatant and insoluble fraction. Furthermore, recombinant VSV-G protein is detected in cell-free extracts from transgenic Schizochytrium expressing VSV-G protein.

[0207] The activity of the VSV-G protein produced in Schizochytrium is assessed by a VSV-G activity assay. Functional VSV-G protein exhibits VSV-G activity that is readily detected by a standard VSV-G activity assay.

[0208] Electron microscopy is performed to confirm the presence of extracellular bodies using negatively stained material prepared according to Example 4. Immunogold labeling is performed to confirm protein association with the extracellular membrane bodies.

[0209] Example 12 Expression and characterization of extracellular bodies containing eGFP fusion proteins produced in Schizochytrium Transformation of Schizochytrium species ATCC 20888 with a vector containing a polynucleotide sequence encoding eGFP and expression of eGFP in the transformed Schizochytrium has been described. See U.S. Patent Application Publication No. 2010 / 0233760 and WO 2010 / 107709, which are incorporated by reference in their entireties.

[0210] In future experiments, Schizochytrium species ATCC 20888 will be used as a host cell for transformation with vectors containing sequences encoding fusion proteins between a membrane domain, such as a Schizochytrium-derived membrane domain or a viral membrane domain, such as an HA membrane domain, and eGFP. Representative Schizochytrium membrane domains are provided in Figures 13 and 14. Transformations are performed, and cryopreserved stocks are grown and propagated according to any of the methods described herein. The Schizochytrium culture is transferred to a 50 mL conical tube and centrifuged at 3000 x g or 4500 x g for 15 minutes to obtain a low-speed supernatant. The low-speed supernatant is further ultracentrifuged at 100,000 x g for 1 hour. See Figure 3. The resulting pellet of insoluble fraction containing the transgenic Schizochytrium-derived eGFP fusion protein is resuspended in phosphate-buffered saline (PBS) and used for peptide sequence analysis and glycosylation analysis as described in Example 2.

[0211] Expression of the eGFP fusion protein from transgenic Schizochytrium is verified by immunoblot analysis using anti-eGFP fusion protein antiserum and secondary antibody at appropriate dilutions according to standard immunoblotting procedures as described in Example 2. The recombinant eGFP fusion protein is detected in the low-speed supernatant and insoluble fraction. Furthermore, the recombinant eGFP fusion protein is detected in cell-free extracts from transgenic Schizochytrium expressing the eGFP fusion protein.

[0212] The activity of eGFP fusion proteins produced in Schizochytrium is assessed using an eGFP fusion protein activity assay. Functional eGFP fusion proteins exhibit eGFP fusion protein activity that is readily detected by a standard eGFP fusion protein activity assay.

[0213] Electron microscopy is performed to confirm the presence of extracellular bodies using negatively stained material prepared according to Example 4. Immunogold labeling is performed to confirm protein association with the extracellular membrane bodies.

[0214] Example 13 Detection of heterologous polypeptides produced in cultures of thraustochytrids. A culture of thraustochytrid host cells containing at least one heterologous polypeptide is prepared in a fermentor under suitable fermentation conditions. The fermentor is batch-treated with a medium containing, for example, carbon (glucose), nitrogen, phosphorus, salts, trace metals, and vitamins. The fermentor is inoculated with a typical seed culture and then cultured for 70-120 hours, feeding a carbon (e.g., glucose) source. The carbon source is fed and consumed throughout the fermentation. After 72-120 hours, the fermentor is harvested, and the broth is centrifuged to separate the biomass from the supernatant.

[0215] Protein content is determined for the biomass and cell-free supernatant using standard assays such as Bradford or BCA. Proteins are further analyzed by standard SDS-PAGE and Western blotting to determine the expression of heterologous polypeptides in each biomass and cell-free supernatant fraction. Heterologous polypeptides containing membrane domains are shown to be associated with the microalgal extracellular bodies by routine staining procedures (e.g., negative staining and immunogold labeling) followed by electron microscopy.

[0216] Example 14 Preparation of virus-like particles from microalgae cultures One or more viral envelope polypeptides are heterologously expressed in microalgal host cells under the conditions described above, so that the viral polypeptides are localized in the microalgal extracellular bodies produced under culture conditions. When overexpressed using appropriate culture conditions and regulatory control elements, the viral envelope polypeptides in the microalgal extracellular bodies spontaneously self-assemble into particles morphologically similar to infectious viruses.

[0217] Similarly, one or more viral envelope polypeptides and one or more viral matrix polypeptides are heterologously expressed in microalgal host cells under the above conditions, so that the viral polypeptides are localized in the microalgal extracellular bodies produced under the culture conditions. When overexpressed using appropriate culture conditions and regulatory control elements, the viral polypeptides in the microalgal extracellular bodies spontaneously self-assemble into particles morphologically similar to infectious viruses.

[0218] All of the various aspects, embodiments and options described herein can be combined in any and all variations.

[0219] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.

[0220] Array information SEQUENCE LISTING <110> SANOFI VACCINE TECHNOLOGIES, S.A.S. <120> Production of Heterologous Polypeptides in Microalgae, Microalgal Extracellular Bodies, Compositions, and Methods of Making and Uses Thereof <150> US 61 / 413,353 <151> 2010-11-12 <150> US 61 / 290,469 <151> 2009-12-28 <150> US 61 / 290,441 <151> 2009-12-28 <160> 103 <170> PatentIn version 3.5 <210> 1 <211> 35 <212> PRT <213> Schizochytrium <400> 1 Met Ala Asn Ile Met Ala Asn Val Thr Pro Gln Gly Val Ala Lys Gly 1 5 10 15 Phe Gly Leu Phe Val Gly Val Leu Phe Phe Leu Tyr Trp Phe Leu Val 20 25 30 Gly Leu Ala 35 <210> 2 <211> 105 <212> DNA <213> Schizochytrium <400> 2 atggccaaca tcatggccaa cgtcacgccc cagggcgtcg ccaagggctt tggcctcttt 60 gtcggcgtgc tcttctttct ctactggttc cttgtcggcc tcgcc 105 <210> 3 <211> 1999 <212> DNA <213> Schizochytrium <400> 3 ccgcgaatca agaaggtagg cgcgctgcga ggcgcggcgg cggagcggag cgagggagag 60 ggagagggag agagagggag ggagacgtcg ccgcggcggg gcctggcctg gcctggtttg 120 gcttggtcag cgcggccttg tccgagcgtg cagctggagt tgggtggatt catttggatt 180 ttcttttgtt tttgtttttc tctctttccc ggaaagtgtt ggccggtcgg tgttctttgt 240 tttgatttct tcaaaagttt tggtggttgg ttctctctct tggctctctg tcaggcggtc 300 cggtccacgc cccggcctct cctctcctct cctctcctct cctctccgtg cgtatacgta 360 cgtacgtttg tatacgtaca tacatcccgc ccgccgtgcc ggcgagggtt tgctcagcct 420 ggagcaatgc gatgcgatgc gatgcgatgc gacgcgacgc gacgcgagtc actggttcgc 480 gctgtggctg tggcttgctt gcttacttgc tttcgagctc tcccgctttc ttctttcctt 540 ctcacgccac caccaacgaa agaagatcgg ccccggcacg ccgctgagaa gggctggcgg 600 cgatgacggc acgcgcgccc gctgccacgt tggcgctcgc tgctgctgct gctgctgctg 660 ctgctgctgc tgctgctgct gctgctgctt ctgcgcgcag gctttgccac gaggccggcg 720 tgctggccgc tgccgcttcc agtccgcgtg gagagatcga andgagagata aactggatgg 780 attcatcgag ggatgaatga acgatggttg gatgcctttt tccttttca ggtccacagc 840 gggaagcagg agcgcgtgaa tctgccgcca tccgcatacg tctgcatcgc atcgcatcgc 900 atgcacgcat cgctcgccgg gagccacaga cgggcgacag ggcggccagc cagccaggca 960 gccagccagg caggcaccag agggccagag agcgcgcctc acgcacgcgc cgcagtgcgc 1020 gcatcgctcg cagtgcagac cttgattccc cgcgcggatc tccgcgagcc cgaaacgaag 1080 agcgccgtac gggcccatcc tagcgtcgcc tcgcaccgca tcgcatcgca tcgcgttccc 1140 tagagagtag tactcgacga aggcaccatt tccgcgctcc tcttcggcgc gatcgaggcc 1200 cccggcgccg cgacgatcgc ggcggccgcg gcgctggcgg cggccctggc gctcgcgctg 1260 gcggccgccg cggggcgtctg gccctggcgc gcgcggcgc cgcaggagg gcggcagcgg 1320. ctgctcgccg ccagagaagg agcgcgccgg gcccggggag ggacggggag gagaagga 1380 aggcgcgcaa ggcggccccg aagagaga ccctggactt gaacgcgag aagagaaga aggagaga gttgagaga aagagaga aggagaga gttgagaga acgaggagca ggcgcgttcc aaggcgcgtt ctcttccgga ggcgcgttcc agctgcggcg gcggggcggg 1560 ctgcggggcg ggcgcggggcg cgggtgcgggg cagaggggac gcggcgcggcgg aggcggaggg 1620 ggccgagcgg gagcccctgc tgctgcgggg cgcccgggcc gcaggtgtgg cgcgcgcgac 1680 gacggaggcg acgacgccag cggccgcgac gacaaggccg gcggcgtcgg cgggcggaag 1740 gccccgcgcg gagcaggggc gggagcagga caaggcgcag gagcaggagc agggccggga gcggggagcgg gagcgggcgg cggagcccga ggcagaaccc aatcgagatc cagagcgagc 1860. agaggccggc cgcgagcccg agcccgcgcc gcagatcact agtaccgctg cggaatcaca 1920 gcagcag gcagcag gcagcag gcagcag agggagataa 1980s agaaaaagcg gcagagacg 1999 <210> 4 <211> 325 <212> DNA <213> Schizochytrium <400> 4 gatccgaaag tgaaccttgt cctaacccga cagcgaatgg cgggaggggg cgggctaaaa 60 gatcgtatta catagtattt ttcccctact cttgtgtgttt gtcttttttt ttttttgaac 120 gcattcaagc cacttgtctt ggttacttg tttgtttgct tgcttgcttg cttgcttgcc 180 tgcttcttgg tcagacggcc caaaaaaggg aaaaattca ttcatggcac agataagaaa 240 aagaaaaagtt ttgtcgacca ccgtcatcag aaagcaagag aagagaaaca ctcgcgctca 300 cattctcgct cgcgtaagaa tctta 325 <210> 5 <211> 372 <212> DNA <213> Streptoalloteichus hindustanus <400> 5 atggccaagt tgaccagtgc cgtccggtg ctcaccgcgc gcgacgtcgc cggagcggtc 60 gagttctgga ccgaccggct cgggttctcc cgggacttcg tggaggacga cttcgccggt 120 gtggtccggg acgacgtgac cctgttcatc agcgcggtcc aggaccaggt ggtgccggac 180 aacaccctgg cctgggtgtg ggtgcgcggc ctggacgagc tgtacgccga gtggtcggag 240 gtcgtgtcca cgaacttccg ggacgcctcc gggccggcca tgaccgagat cggcgagcag 300 ccgtgggggc gggagttcgc cctgcgcgac ccggccggca actgcgtgca cttcgtggcc 360 gaggagcagg ac 372 <210> 6 <211> 2055 <212> DNA <213> Schizochytrium <400> 6 atgagcgcga cccgcgcggc gacgaggaca gcggcggcgc tgtcctcggc gctgacgacg 60 cctgtaaagc agcagcagca gcagcagctg cgcgtaggcg cggcgtcggc acggctggcg 120 gccgcggcgt tctcgtccgg cacgggcgga gacgcggcca agaaggcggc cgcggcgagg 180 gcgttctcca cgggacgcgg ccccaacgcg acacgcgaga agagctcgct ggccacggtc 240 caggcggcga cggacgatgc gcgcttcgtc ggcctgaccg gcgcccaaat ctttcatgag 300 ctcatgcgcg agcaccaggt ggacaccatc tttggctacc ctggcggcgc cattctgccc 360 gtttttgatg ccatttttga gagtgacgcc ttcaagttca ttctcgctcg ccacgagcag 420 ggcgccggcc acatggccga gggctacgcg cgcgccacgg gcaagcccgg cgttgtcctc 480 gtcacctcgg gccctggagc caccaacacc atcaccccga tcatggatgc ttacatggac 540 ggtacgccgc tgctcgtgtt caccggccag gtgcccacct ctgctgtcgg cacggacgct 600 ttccaggagt gtgacattgt tggcatcagc cgcgcgtgca ccaagtggaa cgtcatggtc 660 aaggacgtga aggagctccc gcgccgcatc aatgaggcct ttgagattgc catgagcggc 720 cgcccgggtc ccgtgctcgt cgatcttcct aaggatgtga ccgccgttga gctcaaggaa 780 atgcccgaca gctcccccca ggttgctgtg cgccagaagc aaaaggtcga gcttttccac 840 aaggagcgca ttggcgctcc tggcacggcc gacttcaagc tcattgccga gatgatcaac 900 cgtgcggagc gacccgtcat ctatgctggc cagggtgtca tgcagagccc gttgaatggc 960 ccggctgtgc tcaaggagtt cgcggagaag gccaacattc ccgtgaccac caccatgcag 1020 ggtctcggcg gctttgacga gcgtagtccc ctctccctca agatgctcgg catgcacggc 1080 tctgcctacg ccaactactc gatgcagaac gccgatctta tcctggcgct cggtgcccgc 1140 tttgatgatc gtgtgacggg ccgcgttgac gcctttgctc cggaggctcg ccgtgccgag 1200 cgcgagggcc gcggtggcat cgttcacttt gagatttccc ccaagaacct ccacaaggtc 1260 gtccagccca ccgtcgcggt cctcggcgac gtggtcgaga acctcgccaa cgtcacgccc 1320 cacgtgcagc gccaggagcg cgagccgtgg tttgcgcaga tcgccgattg gaaggagaag 1380 cacccttttc tgctcgagtc tgttgattcg gacgacaagg ttctcaagcc gcagcaggtc 1440 ctcacggagc ttaacaagca gattctcgag attcaggaga aggacgccga ccaggaggtc 1500 tacatcacca cgggcgtcgg aagccaccag atgcaggcag cgcagttcct tacctggacc 1560 aagccgcgcc agtggatctc ctcgggtggc gccggcacta tgggctacgg ccttccctcg 1620 gccattggcg ccaagattgc caagcccgat gctattgtta ttgacatcga tggtgatgct 1680 tcttattcga tgaccggtat ggaattgatc acagcagccg aattcaaggt tggcgtgaag 1740 attcttcttt tgcagaacaa ctttcagggc atggtcaaga actggcagga tctcttttac 1800 <h2 style=";text-align:left;direction:ltr">gacaagcgct actcgggcac cgccatgttc aacccgcgct tcgacaaggt cgccgatgcg 1860<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> atgcgtgcca agggtctcta ctgcgcgaaa cagtcggagc tcaaggacaa gatcaaggag 1920<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> tttctcgagt acgatgaggg tcccgtcctc ctcgaggttt tcgtggacaa ggacacgctc 1980<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> gtcttgccca tggtccccgc tggctttccg ctccacgaga tggtcctcga gcctcctaag 2040<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> cccaaggacg cctaa 2055<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <210> 7<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <211> 684<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <212> PRT<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <213> Schizochytrium<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <400> 7<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> Met Ser Ala Thr Arg Ala Ala Thr Arg Thr Ala Ala Ala Leu Ser Ser<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 1 5 10 15<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> Ala Leu Thr Thr Pro Val Lys Gln Gln Gln Gln Gln Gln Leu Arg Val<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 20 25 30<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> Gly Ala Ala Ser Ala Arg Leu Ala Ala Ala Ala Phe Ser Ser Gly Thr<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 35 40 45<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> Gly Gly Asp Ala Ala Lys Lys Ala Ala Ala Ala Arg Ala Phe Ser Thr<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 50 55 60<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> Gly Arg Gly Pro Asn Ala Thr Arg Glu Lys Ser Ser Leu Ala Thr Val<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 65 70 75 80<h2 style=";text-align:left;direction:ltr"> Gln Ala Ala Thr Asp Asp Ala Arg Phe Val Gly Leu Thr Gly Ala Gln 85 90 95 Ile Phe His Glu Leu Met Arg Glu His Gln Val Asp Thr Ile Phe Gly 100 105 110 Tyr Pro Gly Gly Ala Ile Leu Pro Val Phe Asp Ala Ile Phe Glu Ser 115 120 125 Asp Ala Phe Lys Phe Ile Leu Ala Arg His Glu Gln Gly Ala Gly His 130 135 140 Met Ala Glu Gly Tyr Ala Arg Ala Thr Gly Lys Pro Gly Val Val Leu 145 150 155 160 Val Thr Ser Gly Pro Gly Ala Thr Asn Thr Ile Thr Pro Ile Met Asp 165 170 175 Ala Tyr Met Asp Gly Thr Pro Leu Leu Val Phe Thr Gly Gln Val Pro 180 185 190 Thr Ser Ala Val Gly Thr Asp Ala Phe Gln Glu Cys Asp Ile Val Gly 195 200 205 Ile Ser Arg Ala Cys Thr Lys Trp Asn Val Met Val Lys Asp Val Lys 210 215 220 Glu Leu Pro Arg Arg Ile Asn Glu Ala Phe Glu Ile Ala Met Ser Gly 225 230 235 240 Arg Pro Gly Pro Val Leu Val Asp Leu Pro Lys Asp Val Thr Ala Val 245 250 255 Glu Leu Lys Glu Met Pro Asp Ser Ser Pro Gln Val Ala Val Arg Gln 260 265 270 Lys Gln Lys Val Glu Leu Phe His Lys Glu Arg Ile Gly Ala Pro Gly 275 280 285 Thr Ala Asp Phe Lys Leu Ile Ala Glu Met Ile Asn Arg Ala Glu Arg 290 295 300 Pro Val Ile Tyr Ala Gly Gln Gly Val Met Gln Ser Pro Leu Asn Gly 305 310 315 320 Pro Ala Val Leu Lys Glu Phe Ala Glu Lys Ala Asn Ile Pro Val Thr 325 330 335 Thr Thr Met Gln Gly Leu Gly Gly Phe Asp Glu Arg Ser Pro Leu Ser 340 345 350 Leu Lys Met Leu Gly Met His Gly Ser Ala Tyr Ala Asn Tyr Ser Met 355 360 365 Gln Asn Ala Asp Leu Ile Leu Ala Leu Gly Ala Arg Phe Asp Asp Arg 370 375 380 Val Thr Gly Arg Val Asp Ala Phe Ala Pro Glu Ala Arg Arg Ala Glu 385 390 395 400 Arg Glu Gly Arg Gly Gly Ile Val His Phe Glu Ile Ser Pro Lys Asn 405 410 415 Leu His Lys Val Val Gln Pro Thr Val Ala Val Leu Gly Asp Val Val 420 425 430 Glu Asn Leu Ala Asn Val Thr Pro His Val Gln Arg Gln Glu Arg Glu 435 440 445 Pro Trp Phe Ala Gln Ile Ala Asp Trp Lys Glu Lys His Pro Phe Leu 450 455 460 Leu Glu Ser Val Asp Ser Asp Asp Lys Val Leu Lys Pro Gln Gln Val 465 470 475 480 Leu Thr Glu Leu Asn Lys Gln Ile Leu Glu Ile Gln Glu Lys Asp Ala 485 490 495 Asp Gln Glu Val Tyr Ile Thr Thr Gly Val Gly Ser His Gln Met Gln 500 505 510 Ala Ala Gln Phe Leu Thr Trp Thr Lys Pro Arg Gln Trp Ile Ser Ser 515 520 525 Gly Gly Ala Gly Thr Met Gly Tyr Gly Leu Pro Ser Ala Ile Gly Ala 530 535 540 Lys Ile Ala Lys Pro Asp Ala Ile Val Ile Asp Ile Asp Gly Asp Ala 545 550 555 560 Ser Tyr Ser Met Thr Gly Met Glu Leu Ile Thr Ala Ala Glu Phe Lys 565 570 575 Val Gly Val Lys Ile Leu Leu Leu Gln Asn Asn Phe Gln Gly Met Val 580 585 590 Lys Asn Trp Gln Asp Leu Phe Tyr Asp Lys Arg Tyr Ser Gly Thr Ala 595 600 605 Met Phe Asn Pro Arg Phe Asp Lys Val Ala Asp Ala Met Arg Ala Lys 610 615 620 Gly Leu Tyr Cys Ala Lys Gln Ser Glu Leu Lys Asp Lys Ile Lys Glu 625 630 635 640 Phe Leu Glu Tyr Asp Glu Gly Pro Val Leu Leu Glu Val Phe Val Asp 645 650 655 Lys Asp Thr Leu Val Leu Pro Met Val Pro Ala Gly Phe Pro Leu His 660 665 670 Glu Met Val Leu Glu Pro Pro Lys Pro Lys Asp Ala 675 680 <210> 8 <211> 684 <212> PRT <213> Artificial Sequence <220> <223> Mutated ALS 1 <400> 8 Met Ser Ala Thr Arg Ala Ala Thr Arg Thr Ala Ala Ala Leu Ser Ser 1 5 10 15 Ala Leu Thr Thr Pro Val Lys Gln Gln Gln Gln Gln Gln Leu Arg Val 20 25 30 Gly Ala Ala Ser Ala Arg Leu Ala Ala Ala Ala Phe Ser Ser Gly Thr 35 40 45 Gly Gly Asp Ala Ala Lys Lys Ala Ala Ala Ala Arg Ala Phe Ser Thr 50 55 60 Gly Arg Gly Pro Asn Ala Thr Arg Glu Lys Ser Ser Leu Ala Thr Val 65 70 75 80 Gln Ala Ala Thr Asp Asp Ala Arg Phe Val Gly Leu Thr Gly Ala Gln 85 90 95 Ile Phe His Glu Leu Met Arg Glu His Gln Val Asp Thr Ile Phe Gly 100 105 110 Tyr Pro Gly Gly Ala Ile Leu Pro Val Phe Asp Ala Ile Phe Glu Ser 115 120 125 Asp Ala Phe Lys Phe Ile Leu Ala Arg His Glu Gln Gly Ala Gly His 130 135 140 Met Ala Glu Gly Tyr Ala Arg Ala Thr Gly Lys Pro Gly Val Val Leu 145 150 155 160 Val Thr Ser Gly Pro Gly Ala Thr Asn Thr Ile Thr Pro Ile Met Asp 165 170 175 Ala Tyr Met Asp Gly Thr Pro Leu Leu Val Phe Thr Gly Gln Val Pro 180 185 190 Thr Ser Ala Val Gly Thr Asp Ala Phe Gln Glu Cys Asp Ile Val Gly 195 200 205 Ile Ser Arg Ala Cys Thr Lys Trp Asn Val Met Val Lys Asp Val Lys 210 215 220 Glu Leu Pro Arg Arg Ile Asn Glu Ala Phe Glu Ile Ala Met Ser Gly 225 230 235 240 Arg Pro Gly Pro Val Leu Val Asp Leu Pro Lys Asp Val Thr Ala Val 245 250 255 Glu Leu Lys Glu Met Pro Asp Ser Ser Pro Gln Val Ala Val Arg Gln 260 265 270 Lys Gln Lys Val Glu Leu Phe His Lys Glu Arg Ile Gly Ala Pro Gly 275 280 285 Thr Ala Asp Phe Lys Leu Ile Ala Glu Met Ile Asn Arg Ala Glu Arg 290 295 300 Pro Val Ile Tyr Ala Gly Gln Gly Val Met Gln Ser Pro Leu Asn Gly 305 310 315 320 Pro Ala Val Leu Lys Glu Phe Ala Glu Lys Ala Asn Ile Pro Val Thr 325 330 335 Thr Thr Met Gln Gly Leu Gly Gly Phe Asp Glu Arg Ser Pro Leu Ser 340 345 350 Leu Lys Met Leu Gly Met His Gly Ser Ala Tyr Ala Asn Tyr Ser Met 355 360 365 Gln Asn Ala Asp Leu Ile Leu Ala Leu Gly Ala Arg Phe Asp Asp Arg 370 375 380 Val Thr Gly Arg Val Asp Ala Phe Ala Pro Glu Ala Arg Arg Ala Glu 385 390 395 400 Arg Glu Gly Arg Gly Gly Ile Val His Phe Glu Ile Ser Pro Lys Asn 405 410 415 Leu His Lys Val Val Gln Pro Thr Val Ala Val Leu Gly Asp Val Val 420 425 430 Glu Asn Leu Ala Asn Val Thr Pro His Val Gln Arg Gln Glu Arg Glu 435 440 445 Pro Trp Phe Ala Gln Ile Ala Asp Trp Lys Glu Lys His Pro Phe Leu 450 455 460 Leu Glu Ser Val Asp Ser Asp Asp Lys Val Leu Lys Pro Gln Gln Val 465 470 475 480 Leu Thr Glu Leu Asn Lys Gln Ile Leu Glu Ile Gln Glu Lys Asp Ala 485 490 495 Asp Gln Glu Val Tyr Ile Thr Thr Gly Val Gly Ser His Gln Met Gln 500 505 510 Ala Ala Gln Phe Leu Thr Trp Thr Lys Pro Arg Gln Trp Ile Ser Ser 515 520 525 Gly Gly Ala Gly Thr Met Gly Tyr Gly Leu Pro Ser Ala Ile Gly Ala 530 535 540 Lys Ile Ala Lys Pro Asp Ala Ile Val Ile Asp Ile Asp Gly Asp Ala 545 550 555 560 Ser Tyr Ser Met Thr Gly Met Glu Leu Ile Thr Ala Ala Glu Phe Lys 565 570 575 Val Gly Val Lys Ile Leu Leu Leu Gln Asn Asn Phe Gln Gly Met Val 580 585 590 Lys Asn Val Gln Asp Leu Phe Tyr Asp Lys Arg Tyr Ser Gly Thr Ala 595 600 605 Met Phe Asn Pro Arg Phe Asp Lys Val Ala Asp Ala Met Arg Ala Lys 610 615 620 Gly Leu Tyr Cys Ala Lys Gln Ser Glu Leu Lys Asp Lys Ile Lys Glu 625 630 635 640 Phe Leu Glu Tyr Asp Glu Gly Pro Val Leu Leu Glu Val Phe Val Asp 645 650 655 Lys Asp Thr Leu Val Leu Pro Met Val Pro Ala Gly Phe Pro Leu His 660 665 670 Glu Met Val Leu Glu Pro Pro Lys Pro Lys Asp Ala 675 680 <210> 9 <211> 684 <212> PRT <213> Artificial Sequence <220> <223> Mutated ALS 2 <400> 9 Met Ser Ala Thr Arg Ala Ala Thr Arg Thr Ala Ala Ala Leu Ser Ser 1 5 10 15 Ala Leu Thr Thr Pro Val Lys Gln Gln Gln Gln Gln Gln Leu Arg Val 20 25 30 Gly Ala Ala Ser Ala Arg Leu Ala Ala Ala Ala Phe Ser Ser Gly Thr 35 40 45 Gly Gly Asp Ala Ala Lys Lys Ala Ala Ala Ala Arg Ala Phe Ser Thr 50 55 60 Gly Arg Gly Pro Asn Ala Thr Arg Glu Lys Ser Ser Leu Ala Thr Val 65 70 75 80 Gln Ala Ala Thr Asp Asp Ala Arg Phe Val Gly Leu Thr Gly Ala Gln 85 90 95 Ile Phe His Glu Leu Met Arg Glu His Gln Val Asp Thr Ile Phe Gly 100 105 110 Tyr Pro Gly Gly Ala Ile Leu Pro Val Phe Asp Ala Ile Phe Glu Ser 115 120 125 Asp Ala Phe Lys Phe Ile Leu Ala Arg His Glu Gln Gly Ala Gly His 130 135 140 Met Ala Glu Gly Tyr Ala Arg Ala Thr Gly Lys Pro Gly Val Val Leu 145 150 155 160 Val Thr Ser Gly Pro Gly Ala Thr Asn Thr Ile Thr Pro Ile Met Asp 165 170 175 Ala Tyr Met Asp Gly Thr Pro Leu Leu Val Phe Thr Gly Gln Val Gln 180 185 190 Thr Ser Ala Val Gly Thr Asp Ala Phe Gln Glu Cys Asp Ile Val Gly 195 200 205 Ile Ser Arg Ala Cys Thr Lys Trp Asn Val Met Val Lys Asp Val Lys 210 215 220 Glu Leu Pro Arg Arg Ile Asn Glu Ala Phe Glu Ile Ala Met Ser Gly 225 230 235 240 Arg Pro Gly Pro Val Leu Val Asp Leu Pro Lys Asp Val Thr Ala Val 245 250 255 Glu Leu Lys Glu Met Pro Asp Ser Ser Pro Gln Val Ala Val Arg Gln 260 265 270 Lys Gln Lys Val Glu Leu Phe His Lys Glu Arg Ile Gly Ala Pro Gly 275 280 285 Thr Ala Asp Phe Lys Leu Ile Ala Glu Met Ile Asn Arg Ala Glu Arg 290 295 300 Pro Val Ile Tyr Ala Gly Gln Gly Val Met Gln Ser Pro Leu Asn Gly 305 310 315 320 Pro Ala Val Leu Lys Glu Phe Ala Glu Lys Ala Asn Ile Pro Val Thr 325 330 335 Thr Thr Met Gln Gly Leu Gly Gly Phe Asp Glu Arg Ser Pro Leu Ser 340 345 350 Leu Lys Met Leu Gly Met His Gly Ser Ala Tyr Ala Asn Tyr Ser Met 355 360 365 Gln Asn Ala Asp Leu Ile Leu Ala Leu Gly Ala Arg Phe Asp Asp Arg 370 375 380 Val Thr Gly Arg Val Asp Ala Phe Ala Pro Glu Ala Arg Arg Ala Glu 385 390 395 400 Arg Glu Gly Arg Gly Gly Ile Val His Phe Glu Ile Ser Pro Lys Asn 405 410 415 Leu His Lys Val Val Gln Pro Thr Val Ala Val Leu Gly Asp Val Val 420 425 430 Glu Asn Leu Ala Asn Val Thr Pro His Val Gln Arg Gln Glu Arg Glu 435 440 445 Pro Trp Phe Ala Gln Ile Ala Asp Trp Lys Glu Lys His Pro Phe Leu 450 455 460 Leu Glu Ser Val Asp Ser Asp Asp Lys Val Leu Lys Pro Gln Gln Val 465 470 475 480 Leu Thr Glu Leu Asn Lys Gln Ile Leu Glu Ile Gln Glu Lys Asp Ala 485 490 495 Asp Gln Glu Val Tyr Ile Thr Thr Gly Val Gly Ser His Gln Met Gln 500 505 510 Ala Ala Gln Phe Leu Thr Trp Thr Lys Pro Arg Gln Trp Ile Ser Ser 515 520 525 Gly Gly Ala Gly Thr Met Gly Tyr Gly Leu Pro Ser Ala Ile Gly Ala 530 535 540 Lys Ile Ala Lys Pro Asp Ala Ile Val Ile Asp Ile Asp Gly Asp Ala 545 550 555 560 Ser Tyr Ser Met Thr Gly Met Glu Leu Ile Thr Ala Ala Glu Phe Lys 565 570 575 Val Gly Val Lys Ile Leu Leu Leu Gln Asn Asn Phe Gln Gly Met Val 580 585 590 Lys Asn Trp Gln Asp Leu Phe Tyr Asp Lys Arg Tyr Ser Gly Thr Ala 595 600 605 Met Phe Asn Pro Arg Phe Asp Lys Val Ala Asp Ala Met Arg Ala Lys 610 615 620 Gly Leu Tyr Cys Ala Lys Gln Ser Glu Leu Lys Asp Lys Ile Lys Glu 625 630 635 640 Phe Leu Glu Tyr Asp Glu Gly Pro Val Leu Leu Glu Val Phe Val Asp 645 650 655 Lys Asp Thr Leu Val Leu Pro Met Val Pro Ala Gly Phe Pro Leu His 660 665 670 Glu Met Val Leu Glu Pro Pro Lys Pro Lys Asp Ala 675 680 <210> 10 <211> 684 <212> PRT <213> Artificial Sequence <220> <223> Mutated ALS 3 <400> 10 Met Ser Ala Thr Arg Ala Ala Thr Arg Thr Ala Ala Ala Leu Ser Ser 1 5 10 15 Ala Leu Thr Thr Pro Val Lys Gln Gln Gln Gln Gln Gln Leu Arg Val 20 25 30 Gly Ala Ala Ser Ala Arg Leu Ala Ala Ala Ala Phe Ser Ser Gly Thr 35 40 45 Gly Gly Asp Ala Ala Lys Lys Ala Ala Ala Ala Arg Ala Phe Ser Thr 50 55 60 Gly Arg Gly Pro Asn Ala Thr Arg Glu Lys Ser Ser Leu Ala Thr Val 65 70 75 80 Gln Ala Ala Thr Asp Asp Ala Arg Phe Val Gly Leu Thr Gly Ala Gln 85 90 95 Ile Phe His Glu Leu Met Arg Glu His Gln Val Asp Thr Ile Phe Gly 100 105 110 Tyr Pro Gly Gly Ala Ile Leu Pro Val Phe Asp Ala Ile Phe Glu Ser 115 120 125 Asp Ala Phe Lys Phe Ile Leu Ala Arg His Glu Gln Gly Ala Gly His 130 135 140 Met Ala Glu Gly Tyr Ala Arg Ala Thr Gly Lys Pro Gly Val Val Leu 145 150 155 160 Val Thr Ser Gly Pro Gly Ala Thr Asn Thr Ile Thr Pro Ile Met Asp 165 170 175 Ala Tyr Met Asp Gly Thr Pro Leu Leu Val Phe Thr Gly Gln Val Gln 180 185 190 Thr Ser Ala Val Gly Thr Asp Ala Phe Gln Glu Cys Asp Ile Val Gly 195 200 205 Ile Ser Arg Ala Cys Thr Lys Trp Asn Val Met Val Lys Asp Val Lys 210 215 220 Glu Leu Pro Arg Arg Ile Asn Glu Ala Phe Glu Ile Ala Met Ser Gly 225 230 235 240 Arg Pro Gly Pro Val Leu Val Asp Leu Pro Lys Asp Val Thr Ala Val 245 250 255 Glu Leu Lys Glu Met Pro Asp Ser Ser Pro Gln Val Ala Val Arg Gln 260 265 270 Lys Gln Lys Val Glu Leu Phe His Lys Glu Arg Ile Gly Ala Pro Gly 275 280 285 Thr Ala Asp Phe Lys Leu Ile Ala Glu Met Ile Asn Arg Ala Glu Arg 290 295 300 Pro Val Ile Tyr Ala Gly Gln Gly Val Met Gln Ser Pro Leu Asn Gly 305 310 315 320 Pro Ala Val Leu Lys Glu Phe Ala Glu Lys Ala Asn Ile Pro Val Thr 325 330 335 Thr Thr Met Gln Gly Leu Gly Gly Phe Asp Glu Arg Ser Pro Leu Ser 340 345 350 Leu Lys Met Leu Gly Met His Gly Ser Ala Tyr Ala Asn Tyr Ser Met 355 360 365 Gln Asn Ala Asp Leu Ile Leu Ala Leu Gly Ala Arg Phe Asp Asp Arg 370 375 380 Val Thr Gly Arg Val Asp Ala Phe Ala Pro Glu Ala Arg Arg Ala Glu 385 390 395 400 Arg Glu Gly Arg Gly Gly Ile Val His Phe Glu Ile Ser Pro Lys Asn 405 410 415 Leu His Lys Val Val Gln Pro Thr Val Ala Val Leu Gly Asp Val Val 420 425 430 Glu Asn Leu Ala Asn Val Thr Pro His Val Gln Arg Gln Glu Arg Glu 435 440 445 Pro Trp Phe Ala Gln Ile Ala Asp Trp Lys Glu Lys His Pro Phe Leu 450 455 460 Leu Glu Ser Val Asp Ser Asp Asp Lys Val Leu Lys Pro Gln Gln Val 465 470 475 480 Leu Thr Glu Leu Asn Lys Gln Ile Leu Glu Ile Gln Glu Lys Asp Ala 485 490 495 Asp Gln Glu Val Tyr Ile Thr Thr Gly Val Gly Ser His Gln Met Gln 500 505 510 Ala Ala Gln Phe Leu Thr Trp Thr Lys Pro Arg Gln Trp Ile Ser Ser 515 520 525 Gly Gly Ala Gly Thr Met Gly Tyr Gly Leu Pro Ser Ala Ile Gly Ala 530 535 540 Lys Ile Ala Lys Pro Asp Ala Ile Val Ile Asp Ile Asp Gly Asp Ala 545 550 555 560 Ser Tyr Ser Met Thr Gly Met Glu Leu Ile Thr Ala Ala Glu Phe Lys 565 570 575 Val Gly Val Lys Ile Leu Leu Leu Gln Asn Asn Phe Gln Gly Met Val 580 585 590 Lys Asn Val Gln Asp Leu Phe Tyr Asp Lys Arg Tyr Ser Gly Thr Ala 595 600 605 Met Phe Asn Pro Arg Phe Asp Lys Val Ala Asp Ala Met Arg Ala Lys 610 615 620 Gly Leu Tyr Cys Ala Lys Gln Ser Glu Leu Lys Asp Lys Ile Lys Glu 625,630,635,640 Phe Leu Glu Tyr Asp Glu Gly Pro Val Leu Leu Glu Val Phe Val Asp 645,650,655 Lys Asp Thr Leu Val Leu Pro Met Val Pro Ala Gly Phe Pro Leu His 660 665 670 Glu Met Val Leu Glu Pro Pro Lys Pro Lys Asp Here 675 680 <210> 11 <211> 2052 <212> DNA <213> Artificial Sequence <220> <223> Mutated ALS 1 <400> 11 atgagcgcga ccccgcggc gacgaggaca gcggcggcgc tgtcctcggc gctgacgacg 60 cctgtaaagc agcagcagca gcagcagctg cgcgtaggcg cggcgcggc acggctggcg 120 gccgcggcgt tctcgtccgg cacgggcgga gacgcggcca agaaggcggc cgcggcgagg 180 gcgttctcca cgggacgcgg ccccaacgcg acacgcgaga agagctcgct ggccacggtc 240 caggcggcga cggacgatgc gcgcttcgtc ggcctgaccg gcgcccaaat ctttcatgag 300 ctcatgcgcg agcaccaggt ggacaccatc tttggctacc ctggcggcgc cattctgccc 360 gtttttgatg ccatttttga gagtgacgcc ttcaagttca ttctcgctcg ccacgagcag 420 ggcgccggcc acatggccga gggctacgcg cgcgccacgg gcaagcccgg cgttgtcctc 480 gtcacctcgg gccctggagc caccaacacc atcaccccga tcatggatgc ttacatggac 540 ggtacgccgc tgctcgtgtt caccggccag gtgcccacct ctgctgtcgg cacggacgct 600 ttccaggagt gtgacattgt tggcatcagc cgcgcgtgca ccaagtggaa cgtcatggtc 660 aaggacgtga aggagctccc gcgccgcatc aatgaggcct ttgagattgc catgagcggc 720 cgcccgggtc ccgtgctcgt cgatcttcct aaggatgtga ccgccgttga gctcaaggaa 780 atgcccgaca gctcccccca ggttgctgtg cgccagaagc aaaaggtcga gcttttccac 840 aaggagcgca ttggcgctcc tggcacggcc gacttcaagc tcattgccga gatgatcaac 900 cgtgcggagc gacccgtcat ctatgctggc cagggtgtca tgcagagccc gttgaatggc 960 ccggctgtgc tcaaggagtt cgcggagaag gccaacattc ccgtgaccac caccatgcag 1020 ggtctcggcg gctttgacga gcgtagtccc ctctccctca agatgctcgg catgcacggc 1080 tctgcctacg ccaactactc gatgcagaac gccgatctta tcctggcgct cggtgcccgc 1140 tttgatgatc gtgtgacggg ccgcgttgac gcctttgctc cggaggctcg ccgtgccgag 1200 cgcgagggcc gcggtggcat cgttcacttt gagatttccc ccaagaacct ccacaaggtc 1260 gtccagccca ccgtcgcggt cctcggcgac gtggtcgaga acctcgccaa cgtcacgccc 1320 cacgtgcagc gccaggagcg cgagccgtgg tttgcgcaga tcgccgattg gaaggagaag 1380 cacccttttc tgctcgagtc tgttgattcg gacgacaagg ttctcaagcc gcagcaggtc 1440 ctcacggagc ttaacaagca gattctcgag attcaggaga aggacgccga ccaggaggtc 1500 tacatcacca cgggcgtcgg aagccaccag atgcaggcag cgcagttcct tacctggacc 1560 aagccgcgcc agtggatctc ctcgggtggc gccggcacta tgggctacgg ccttccctcg 1620 gccattggcg ccaagattgc caagcccgat gctattgtta ttgacatcga tggtgatgct 1680 tcttattcga tgaccggtat ggaattgatc acagcagccg aattcaaggt tggcgtgaag 1740 attcttcttt tgcagaacaa ctttcagggc atggtcaaga acgttcagga tctcttttac 1800 gacaagcgct actcgggcac cgccatgttc aacccgcgct tcgacaaggt cgccgatgcg 1860 atgcgtgcca agggtctcta ctgcgcgaaa cagtcggagc tcaaggacaa gatcaaggag 1920 tttctcgagt acgatgaggg tcccgtcctc ctcgaggttt tcgtggacaa ggacacgctc 1980 gtcttgccca tggtccccgc tggctttccg ctccacgaga tggtcctcga gcctcctaag 2040 cccaaggacg cc 2052 <210> 12 <211> 2052 <212> DNA <213> Artificial Sequence <220> <223> Mutated ALS 2 <400> 12 atgagcgcga cccgcgcggc gacgaggaca gcggcggcgc tgtcctcggc gctgacgacg 60 cctgtaaagc agcagcagca gcagcagctg cgcgtaggcg cggcgtcggc acggctggcg 120 gccgcggcgt tctcgtccgg cacgggcgga gacgcggcca agaaggcggc cgcggcgagg 180 gcgttctcca cgggacgcgg ccccaacgcg acacgcgaga agagctcgct ggccacggtc 240 caggcggcga cggacgatgc gcgcttcgtc ggcctgaccg gcgcccaaat ctttcatgag 300 ctcatgcgcg agcaccaggt ggacaccatc tttggctacc ctggcggcgc cattctgccc 360 gtttttgatg ccatttttga gagtgacgcc ttcaagttca ttctcgctcg ccacgagcag 420 ggcgccggcc acatggccga gggctacgcg cgcgccacgg gcaagcccgg cgttgtcctc 480 gtcacctcgg gccctggagc caccaacacc atcaccccga tcatggatgc ttacatggac 540 ggtacgccgc tgctcgtgtt caccggccag gtgcagacct ctgctgtcgg cacggacgct 600 ttccaggagt gtgacattgt tggcatcagc cgcgcgtgca ccaagtggaa cgtcatggtc 660 aaggacgtga aggagctccc gcgccgcatc aatgaggcct ttgagattgc catgagcggc 720 cgcccgggtc ccgtgctcgt cgatcttcct aaggatgtga ccgccgttga gctcaaggaa 780 atgcccgaca gctcccccca ggttgctgtg cgccagaagc aaaaggtcga gcttttccac 840 aaggagcgca ttggcgctcc tggcacggcc gacttcaagc tcattgccga gatgatcaac 900 cgtgcggagc gacccgtcat ctatgctggc cagggtgtca tgcagagccc gttgaatggc 960 ccggctgtgc tcaaggagtt cgcggagaag gccaacattc ccgtgaccac caccatgcag 1020 ggtctcggcg gctttgacga gcgtagtccc ctctccctca agatgctcgg catgcacggc 1080 tctgcctacg ccaactactc gatgcagaac gccgatctta tcctggcgct cggtgcccgc 1140 tttgatgatc gtgtgacggg ccgcgttgac gcctttgctc cggaggctcg ccgtgccgag 1200 cgcgagggcc gcggtggcat cgttcacttt gagatttccc ccaagaacct ccacaaggtc 1260 gtccagccca ccgtcgcggt cctcggcgac gtggtcgaga acctcgccaa cgtcacgccc 1320 cacgtgcagc gccaggagcg cgagccgtgg tttgcgcaga tcgccgattg gaaggagaag 1380 cacccttttc tgctcgagtc tgttgattcg gacgacaagg ttctcaagcc gcagcaggtc 1440 ctcacggagc ttaacaagca gattctcgag attcaggaga aggacgccga ccaggaggtc 1500 tacatcacca cgggcgtcgg aagccaccag atgcaggcag cgcagttcct tacctggacc 1560 aagccgcgcc agtggatctc ctcgggtggc gccggcacta tgggctacgg ccttccctcg 1620 gccattggcg ccaagattgc caagcccgat gctattgtta ttgacatcga tggtgatgct 1680 tcttattcga tgaccggtat ggaattgatc acagcagccg aattcaaggt tggcgtgaag 1740 attcttcttt tgcagaacaa ctttcagggc atggtcaaga actggcagga tctcttttac 1800 gacaagcgct actcgggcac cgccatgttc aacccgcgct tcgacaaggt cgccgatgcg 1860 atgcgtgcca agggtctcta ctgcgcgaaa cagtcggagc tcaaggacaa gatcaaggag 1920 tttctcgagt acgatgaggg tcccgtcctc ctcgaggttt tcgtggacaa ggacacgctc 1980 gtcttgccca tggtccccgc tggctttccg ctccacgaga tggtcctcga gcctcctaag 2040 cccaaggacg cc 2052 <210> 13 <211> 2055 <212> DNA <213> Artificial Sequence <220> <223> Mutated ALS 3 <400> 13 atgagcgcga cccgcgcggc gacgaggaca gcggcggcgc tgtcctcggc gctgacgacg 60 cctgtaaagc agcagcagca gcagcagctg cgcgtaggcg cggcgtcggc acggctggcg 120 gccgcggcgt tctcgtccgg cacgggcgga gacgcggcca agaaggcggc cgcggcgagg 180 gcgttctcca cgggacgcgg ccccaacgcg acacgcgaga agagctcgct ggccacggtc 240 caggcggcga cggacgatgc gcgcttcgtc ggcctgaccg gcgcccaaat ctttcatgag 300 ctcatgcgcg agcaccaggt ggacaccatc tttggctacc ctggcggcgc cattctgccc 360 gtttttgatg ccatttttga gagtgacgcc ttcaagttca ttctcgctcg ccacgagcag 420 ggcgccggcc acatggccga gggctacgcg cgcgccacgg gcaagcccgg cgttgtcctc 480 gtcacctcgg gccctggagc caccaacacc atcaccccga tcatggatgc ttacatggac 540 ggtacgccgc tgctcgtgtt caccggccag gtgcagacct ctgctgtcgg cacggacgct 600 ttccaggagt gtgacattgt tggcatcagc cgcgcgtgca ccaagtggaa cgtcatggtc 660 aaggacgtga aggagctccc gcgccgcatc aatgaggcct ttgagattgc catgagcggc 720 cgcccgggtc ccgtgctcgt cgatcttcct aaggatgtga ccgccgttga gctcaaggaa 780 atgcccgaca gctcccccca ggttgctgtg cgccagaagc aaaaggtcga gcttttccac 840 aaggagcgca ttggcgctcc tggcacggcc gacttcaagc tcattgccga gatgatcaac 900 cgtgcggagc gacccgtcat ctatgctggc cagggtgtca tgcagagccc gttgaatggc 960 ccggctgtgc tcaaggagtt cgcggagaag gccaacattc ccgtgaccac caccatgcag 1020 ggtctcggcg gctttgacga gcgtagtccc ctctccctca agatgctcgg catgcacggc 1080 tctgcctacg ccaactactc gatgcagaac gccgatctta tcctggcgct cggtgcccgc 1140 tttgatgatc gtgtgacggg ccgcgttgac gcctttgctc cggaggctcg ccgtgccgag 1200 cgcgagggcc gcggtggcat cgttcacttt gagatttccc ccaagaacct ccacaaggtc 1260 gtccagccca ccgtcgcggt cctcggcgac gtggtcgaga acctcgccaa cgtcacgccc 1320 cacgtgcagc gccaggagcg cgagccgtgg tttgcgcaga tcgccgattg gaaggagaag 1380 cacccttttc tgctcgagtc tgttgattcg gacgacaagg ttctcaagcc gcagcaggtc 1440 ctcacggagc ttaacaagca gattctcgag attcaggaga aggacgccga ccaggaggtc 1500 tacatcacca cgggcgtcgg aagccaccag atgcaggcag cgcagttcct tacctggacc 1560 aagccgcgcc agtggatctc ctcgggtggc gccggcacta tgggctacgg ccttccctcg 1620 gccattggcg ccaagattgc caagcccgat gctattgtta ttgacatcga tggtgatgct 1680 tcttattcga tgaccggtat ggaattgatc acagcagccg aattcaaggt tggcgtgaag 1740 attcttcttt tgcagaacaa ctttcagggc atggtcaaga acgttcagga tctcttttac 1800 gacaagcgct actcgggcac cgccatgttc aacccgcgct tcgacaaggt cgccgatgcg 1860 atgcgtgcca agggtctcta ctgcgcgaaa cagtcggagc tcaaggacaa gatcaaggag 1920 tttctcgagt acgatgaggg tcccgtcctc ctcgaggttt tcgtggacaa ggacacgctc 1980 gtcttgccca tggtccccgc tggctttccg ctccacgaga tggtcctcga gcctcctaag 2040 cccaaggacg cctaa 2055 <210> 14 <211> 12 <212> DNA <213> Schizochytrium <400> 14 cacgacgagt tg 12 <210> 15 <211> 53 <212> PRT <213> Schizochytrium <400> 15 Put Ala Donkey Ile Put Ala Donkey Val Thr Pro Gln Gly Val Ala Lys Gly 1 5 10 15 Phe Gly Leu Phe Val Gly Val Leu Phe Phe Leu Tyr Trp Phe Leu Val 20 25 30 Gly Leu Ala Leu Leu Gly Asp Gly Phe Lys Val Ile Ala Gly Asp Ser 35 40 45 Ala Gly Thr Leu Phe 50 <210> 16 <211> 21 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer S4termF <400> 16 gatcccatgg cacgtgctac g 21 <210> 17 <211> 22 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer S4termR <400> 17 ggcaacatgt atgataagat ac 22 <210> 18 <211> 21 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer C2mcsSmaF <400> 18 gatccccggg ttaagcttgg t 21 <210> 19 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer C2mcsSmaR <400> 19 actggggccc gtttaaactc 20 <210> 20 <211> 34 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 5'tubMCS_BglI <400> 20 gactagatct caattttagg ccccccactg accg 34 <210> 21 <211> 34 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 3'SV40MCS_Sal <400> 21 gactgtcgac catgtatgat aagatacatt gatg 34 <210> 22 <211> 28 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 5'ALSproNde3 <400> 22 gactcatatg gcccaggcct actttcac 28 <210> 23 <211> 34 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 3'ALStermBglII <400> 23 gactagatct gggtcaaggc agaagaattc cgcc 34 <210> 24 <211> 97 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer sec.Gfp5'1b <400> 24 tactggttcc ttgtcggcct cgcccttctc ggcgatggct tcaaggtcat cgccggtgac 60 tccgccggta cgctcttcat ggtgagcaag ggcgagg 97 <210> 25 <211> 32 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer sec.Gfp3'Spe <400> 25 gatcggtacc ggtgttcttt gttttgattt ct 32 <210> 26 <211> 105 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer sec.Gfp5'Bam <400> 26 taatggatcc atggccaaca tcatggccaa cgtcacgccc cagggcgtcg ccaagggctt 60 tggcctcttt gtcggcgtgc tcttctttct ctactggttc cttgt 105 <210> 27 <211> 40 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer ss.eGfpHELD3'RV <400> 27 cctgatatct tacaactcgt cgtggttgta cagctcgtcc 40 <210> 28 <211> 105 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer sec.Gfp5'Bam2 <400> 28 taatggatcc atggccaaca tcatggccaa cgtcacgccc cagggcgtcg ccaagggctt 60 tggcctcttt gtcggcgtgc tcttctttct ctactggttc cttgt 105 <210> 29 <211> 27 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer prREZ15 <400> 29 cggtacccgc gaatcaagaa ggtaggc 27 <210> 30 <211> 27 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer prREZ16 <400> 30 cggatcccgt ctctgccgct ttttctt 27 <210> 31 <211> 30 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer prREZ17 <400> 31 cggatccgaa agtgaacctt gtcctaaccc 30 <210> 32 <211> 27 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer prREZ18 <400> 32 ctctagacag atccgcacca tcggccg 27 <210> 33 <211> 32 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 5'eGFP_kpn <400> 33 gactggtacc atggtgaagc aagggcgagg ag 32 <210> 34 <211> 34 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 3'eGFP_xba <400> 34 gacttctaga ttacttgtac agctcgtcca tgcc 34 <210> 35 <211> 32 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 5'ORFCproKpn-2 <400> 35 gatcggtacc ggtgttcttt gttttgattt ct 32 <210> 36 <211> 31 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 3'ORFCproKpn-2 <400> 36 gatcggtacc gtctctgccg ctttttcttt a 31 <210> 37 <211> 20 <212> PRT <213> Schizochytrium <400> 37 Met Lys Phe Ala Thr Ser Val Ala Ile Leu Leu Val Ala Asn Ile Ala 1 5 10 15 Thr Ala Leu Ala 20 <210> 38 <211> 60 <212> DNA <213> Schizochytrium <400> 38 atgaagttcg cgacctcggt cgcaattttg cttgtggcca acatagccac cgccctcgcg 60 <210> 39 <211> 28 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 5'ss-X Bgl long <400> 39 gactagatct atgaagttcg cgacctcg 28 <210> 40 <211> 33 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 3'ritx_kap_bh_Bgl <400> 40 gactagatct tcagcactca ccgcggttaa agg 33 <210> 41 <211> 702 <212> DNA <213> Artificial Sequence <220> <223> Optimized Sec1 <400> 41 atgaagttcg cgacctcggt cgcaattttg cttgtggcca acatagccac cgccctcgcg 60 cagatcgtcc tcagccagtc ccccgccatc ctttccgctt cccccggtga gaaggtgacc 120 atgacctgcc gcgctagctc ctccgtctcg tacatccact ggttccagca gaagcccggc 180 tcgccccca agccctggat ctacgccacc tccaacctcg cctccggtgt tcccgttcgt 240 ttttccggtt ccggttccgg cacctcctac tccctcacca tctcccgcgt cgaggccgag 300 gatgccgcca cctactactg ccagcagtgg accagcaacc cccccacctt cggcggtggt 360 acgaagctcg agattaagcg caccgtcgcc gccccctccg tcttcatttt tccccctcc 420 gatgagcagc tcaagtccgg taccgcctcc gtcgtttgcc tcctcaacaa cttctacccc 480 cgtgaggcca aggtccagtg gaaggtcgac aacgcgcttc agtccggtaa ctcccaggag 540 tccgtcaccg agcaggattc gaaggacagc acctactccc tctcctccac cctcaccctc 600 tccaaggccg actacgagaa gcacaaggtc tacgcctgcg aggtcacgca ccagggtctt 660 tcctccccg tcacgaagtc ctttaaccgc ggtgagtgct ga 702 <210> 42 <211> 853 <212> DNA <213> Schizochytrium <400> 42 ctccatcgat cgtgcggtca aaaagaaagg aagaagaaag gaaaaagaaa ggcgtgcgca 60 cccgagtgcg cgctgagcgc ccgctcgcgg ccccgcggag cctccgcgtt agtccccgcc 120 ccgcgccgcg cagtcccccg ggaggcatcg cgcacctctc gccgccccct cgcgcctcgc 180 cgattccccg cctccccttt tccgcttctt cgccgcctcc gctcgcggcc gcgtcgcccg 240 cgccccgctc cctatctgct ccccaggggg gcactccgca ccttttgcgc ccgctgccgc 300 cgccgcggcc gccccgccgc cctggtttcc cccgcgagcg cggccgcgtc gccgcgcaaa 360 gactcgccgc gtgccgcccc gagcaacggg tggcggcggc gcggcggcgg gcggggcgcg 420 gcggcgcgta ggcggggcta ggcgccggct aggcgaaacg ccgcccccgg gcgccgccgc 480 cgcccgctcc agagcagtcg ccgcgccaga ccgccaacgc agagaccgag accgaggtac 540 gtcgcgcccg agcacgccgc gacgcgcggc agggacgagg agcacgacgc cgcgccgcgc 600 cgcgcggggg gggggaggga gaggcaggac gcgggagcga gcgtgcatgt ttccgcgcga 660 gacgacgccg cgcgcgctgg agaggagata aggcgcttgg atcgcgagag ggccagccag 720 gctggaggcg aaaatgggtg gagaggatag tatcttgcgt gcttggacga ggagactgac 780 gaggaggacg gatacgtcga tgatgatgtg cacagagaag aagcagttcg aaagcgacta 840 ctagcaagca agg 853 <210> 43 <211> 1064 <212> DNA <213> Schizochytrium <400> 43 ctcttatctg cctcgcgccg ttgaccgccg cttgactctt ggcgcttgcc gctcgcatcc 60 tgcctcgctc gcgcaggcgg gcgggcgagt gggtgggtcc gcagccttcc gcgctcgccc 120 gctagctcgc tcgcgccgtg ctgcagccag cagggcagca ccgcacggca ggcaggtccc 180 ggcgcggatc gatcgatcca tcgatccatc gatccatcga tcgtgcggtc aaaaagaaag 240 gaagaagaaa ggaaaaagaa aggcgtgcgc acccgagtgc gcgctgagcg cccgctcgcg 300 gtcccgcgga gcctccgcgt tagtccccgc cccgcgccgc gcagtccccc gggaggcatc 360 gcgcacctct cgccgccccc tcgcgcctcg ccgattcccc gcctcccctt ttccgcttct 420 tcgccgcctc cgctcgcggc cgcgtcgccc gcgccccgct ccctatctgc tccccagggg 480 ggcactccgc accttttgcg ccgctgccg ccgccgcggc ccctggtttc 540 ccccgcgagc gcggccgcgt cgccgcgcaa agactcgccg cgtgccgccc cgagcaacgg 600 gtggcggcgg cgcggcggcg ggcggggcgc ggcggcgcgt aggcggggct aggcgccggc 660 tagcgaaac gccgccccg ggcgccgccg ccgcccgctc cagagcagtc gccgcgccag 720 accgccaacg cagagaccga gaccgaggta cgtcgcgccc gagcacgccg cgacgcgcgg 780 cagggacgag gagcacgacg ccgcgccgcg ccgcgcgggg ggggggaggg agaggcagga 840 cgcgggagcg agcgtgcatg tttccgcgcg agacgacgcc gcgcgcgctg gagaggagat aaggcgcttg gatcgcgaga gggccagcca ggctggaggc gaaaatgggt ggagaggata 960 gtatcttgcg tgcttggacg aggagactga cgaggaggac ggatacgtcg atgatgatgt gcacagagaa gaagcagttc gaaagcgact actagcaagc aagg <210> 44 <211> 837 <212> DNA <213> Schizochytrium <400> 44 cttcgctttc tcaacctatc tggacagcaa tccgccactt gccttgatcc ccttccgcgc 60 ctcaatcact cgctccacgt ccctcttccc cctcctcatc tccgtgcttt ctctgccccc 120 cccccccccg ccgcggcgtg cgcgcgcgtg gcgccgcggc cgcgacacct tccatactat 180 cctcgctccc aaaatgggtt gcgctatagg gcccggctag gcgaaagtct agcaggcact 240 tgcttggcgc agagccgccg cggccgctcg ttgccgcgga tggagaggga gagagagccc 300 gcctcgataa gcagagacag acagtgcgac tgacagacag acagagagac tggcagaccg 360 gaatacctcg aggtgagtgc ggcgcgggcg agcgggcggg agcgggagcg caagagggac 420 ggcgcggcgc ggcggccctg cgcgacgccg cggcgtattc tcgtgcgcag cgccgagcag 480 cgggacgggc ggctggctga tggttgaagc ggggcggggt gaaatgttag atgagatgat 540 catcgacgac ggtccgtgcg tcttggctgg cttggctggc ttggctggcg ggcctgccgt 600 gtttgcgaga aagaggatga ggagagcgac gaggaaggac gagaagactg acgtgtaggg 660 cgcgcgatgg atgatcgatt gattgattga ttgattggtt gattggctgt gtggtcgatg 720 aacgtgtaga ctcagggagc gtggttaaat tgttcttgcg ccagacgcga ggactccacc 780 cccttctttc gcctttacac agcctttttg tgaagcaaca agaaagaaaa agccaag 837 <210> 45 <211> 1020 <212> DNA <213> Schizochytrium <400> 45 ctttttccgc tctgcataat cctaaaagaa agactatacc ctagtcactg tacaaatggg 60 acatttctct cccgagcgat agctaaggat ttttgcttcg tgtgcactgt gtgctctggc 120 cgcgcatcga aagtccagga tcttactgtt tctctttcct ttcctttatt tcctgttctc 180 ttcttcgctt tctcaaccta tctggacagc aatccgccac ttgccttgat ccccttccgc 240 gcctcaatca ctcgctccac gtccctcttc cccctcctca tctccgtgct ttctctcgcc 300 cccccccccc ccgccgcggc gtgcgcgcgc gtggcgccgc ggccgcgaca ccttccatac 360 tatcctcgct cccaaaatgg gttgcgctat agggcccggc taggcgaaag tctagcaggc 420 acttgcttgg cgcagagccg ccgcggccgc tcgttgccgc ggatggagag ggagagagag 480 cccgcctcga taagcagaga cagacagtgc gactgacaga cagacagaga gactggcaga 540 ccggaatacc tcgaggtgag tgcggcgcgg gcgagcgggc gggagcggga gcgcaagagg 600 gacggcgcgg cgcggcggcc ctgcgcgacg ccgcggcgta ttctcgtgcg cagcgccgag 660 cagcgggacg ggcggctggc tgatggttga agcggggcgg ggtgaaatgt tagatgagat 720 gatcatcgac gacggtccgt gcgtcttggc tggcttggct ggcttggctg gcgggcctgc 780 cgtgtttgcg agaaagagga tgaggagagc gacgaggaag gacgagaaga ctgacgtgta 840 gggcgcgcga tggatgatcg attgattgat tgattgattg gttgattggc tgtgtggtcg 900 atgaacgtgt agactcaggg agcgtggtta aattgttctt gcgccagacg cgaggactcc 960 acccccttct ttcgccttta cacagccttt ttgtgaagca acaagaaaga aaaagccaag 1020 <210> 46 <211> 1416 <212> DNA <213> Schizochytrium <400> 46 cccgtccttg acgccttcgc ttccggcgcg gccatcgatt caattcaccc atccgatacg 60 ttccgccccc tcacgtccgt ctgcgcacga cccctgcacg accacgccaa ggccaacgcg 120 ccgctcagct cagcttgtcg acgagtcgca cgtcacatat ctcagatgca ttgcctgcct 180 gcctgcctgc ctgcctgcct gcctgcctgc ctgcctgcct cagcctctct ttgctctctc 240 tgcggcggcc gctgcgacgc gctgtacagg agaatgactc caggaagtgc ggctgggata 300 cgcgctggcg tcggccgtga tgcgcgtgac gggcggcggg cacggccggc acgggttgag 360 cagaggacga agcgaggcga gacgagacag gccaggcgcg gggagcgctc gctgccgtga 420 gcagcagacc agggcgcagg aatgtacttt tcttgcggga gcggagacga ggctgccggc 480 tgctggctgc cggttgctct gcacgcgccg cccgacttgg cgtagcgtgg acgcgcggcg 540 gcggccgccg tctcgtcgcg gtcggctttg ccgtgtatcg acgctgcggg cttgacacgg 600 gatggcggaa gttcagcatc gctgcgatcc ctcgcgccgc agaacgagga gagcgcaggc 660 cggcttcaag tttgaaagga gaggaaggca ggcaaggagc tggaagcttg ccgcggaagg 720 cgcaggcatg cgtcacgtga aaaaaaggga tttcaagagt agtaagtagg tatggtctac 780 aagtccccta ttcttacttc gcggaacgtg ggctgctcgt gcgggcgtcc atcttgtttt 840 tgtttttttt tccgctaggc gcgtgcattg cttgatgagt ctcagcgttc gtctgcagcg 900 agggcaggaa aataagcggc ccgtgccgtc gagcgcacag gacgtgcaag cgccttgcga 960 gcgcagcatc cttgcacggc gagcatagag accgcggccg atggactcca gcgaggaatt 1020 ttcgaccctc tctatcaagc tgcgcttgac agccgggaat ggcagcctga ggagagaggg 1080 gcgaaggaag ggacttggag aaaagaggta aggcaccctc aatcacggcg cgtgaaagcc 1140 agtcatccct cgcaaagaaa agacaaaagc gggttttttg tttcgatggg aaagaatttc 1200 ttagaggaag aagcggcaca cagactcgcg ccatgcagat ttctgcgcag ctcgcgatca 1260 aaccaggaac gtggtcgctg cgcgccacta tcaggggtag cgcacgaata ccaaacgcat 1320 tactagctac gcgcctgtga cccgaggatc gggccacaga cgttgtctct tgccatccca 1380 cgacctggca gcgagaagat cgtccattac tcatcg 1416 <210> 47 <211> 24 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 5'60S-807 <400> 47 tcgatttgcg gatacttgct caca 24 <210> 48 <211> 21 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 3'60S-2821 <400> 48 gacgacctcg cccttggaca c 21 <210> 49 <211> 34 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 5'60Sp-1302-Kpn <400> 49 gactggtacc tttttccgct ctgcataatc ctaa 34 <210> 50 <211> 32 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 3'60Sp-Bam <400> 50 gactggatcc ttggcttttt ctttcttgtt gc 32 <210> 51 <211> 24 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 5'EF1-68 <400> 51 cgccgttgac cgccgcttga ctct 24 <210> 52 <211> 24 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 3'EF1-2312 <400> 52 cggggtagc ctcggggatg gact 24 <210> 53 <211> 34 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 5'EF1-54-Kpn <400> 53 gactggtacc tcttatctgc ctcgcgccgt tgac 34 <210> 54 <211> 37 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 5'EF1-1114-Bam <400> 54 gactggatcc cttgcttgct agtagtcgct ttcgaac 37 <210> 55 <211> 29 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 5'Sec1P-kpn <400> 55 gactggtacc ccgtccttga cgccttcgc 29 <210> 56 <211> 31 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Primer 3'Sec1P-ba <400> 56 gactggatcc gatgagtaat ggacgatctt c 31 <210> 57 <211> 1614 <212> DNA <213> Artificial sequence <220> <223> Secretion signal <400> 57 ggatccatga agttcgcgac ctcggtcgca attttgcttg tggccaacat agccaccgcc 60 ctcgcgtcga tgaccaacga gacctcggac cgccctctcg tgcactttac ccccaacaag 120 ggttggatga acgatcccaa cggcctctgg tacgacgaga aggatgctaa gtggcacctt 180 tactttcagt acaaccctaa cgacaccgtc tggggcaccc cgctcttctg gggccacgcc 240 acctccgacg acctcaccaa ctgggaggac cagcccattg ctatcgcccc caagcgcaac 300 gactcgggag ctttttccgg ttccatggtt gtggactaca acaacacctc cggttttttt 360 aacgacacca ttgacccccg ccagcgctgc gtcgccatct ggacctacaa cacgcccgag 420 agcgaggagc agtacatcag ctacagcctt gatggaggct acacctttac cgagtaccag 480 aagaaccctg tcctcgccgc caactccacc cagttccgcg accctaaggt tttttggtac 540 gagccttccc agaagtggat tatgaccgcc gctaagtcgc aggattacaa gatcgagatc 600 tacagcagcg acgacctcaa gtcctggaag cttgagtccg cctttgccaa cgagggtttt 660 ctcggatacc agtacgagtg ccccggtctc atcgaggtcc ccaccgagca ggacccgtcc 720 aagtcctact gggtcatgtt tatttccatc aaccctggcg cccctgccgg cggcagcttc 780 aaccagtact tcgtcggctc ctttaacggc acgcattttg aggccttcga caaccagtcc 840 cgcgtcgtcg acttcggcaa ggactactac gccctccaga ccttctttaa caccgacccc 900 acctacggca gcgccctcgg tattgcttgg gcctccaact gggagtactc cgctttcgtc 960 cccactaacc cctggcgcag ctcgatgtcc ctcgtccgca agttttcgct taacaccgag 1020 taccaggcca accccgagac cgagcttatt aacctgaagg ccgagcctat tctcaacatc 1080 tccaacgctg gcccctggtc ccgctttgct actaacacta ccctcaccaa ggccaactcc 1140 tacaacgtcg atctctccaa ctccaccggt actcttgagt ttgagctcgt ctacgccgtc 1200 aacaccaccc agaccatctc caagtccgtc ttcgccgacc tctccctctg gttcaagggc 1260 cttgaggacc ccgaggagta cctgcgcatg ggttttgagg tctccgcctc ctccttcttc 1320 ctcgatcgcg gtaactccaa ggttaagttt gtcaaggaga acccctactt tactaaccgt 1380 atgagcgtca acaaccagcc ctttaagtcc gagaacgatc ttagctacta caaggtttac 1440 ggcctcctcg accagaacat tctcgagctc tactttaacg acggagatgt cgtcagcacc 1500 aacacctact ttatgaccac tggaaacgcc ctcggcagcg tgaacatgac caccggagtc 1560 gacaacctct tttacattga caagtttcag gttcgcgagg ttaagtaaca tatg 1614 <210> 58 <211> 11495 <212> DNA <213> Artificial sequence <220> <223> pCL0076 <400> 58 ctctttatctg cctcgcgccg ttgaccgccg cttgactctt ggcgcttgcc gctcgcatcc 60 tgcctcgctc gcgcaggcgg gcgggcgagt gggtgggtcc gcagccttcc gcgctcgccc 120 gctagctcgc tcgcgccgtg ctgcagccag cagggcagca ccgcacggca ggcaggtccc 180 ggcgcggatc gatcgatcca tcgatccatc gatccatcga tcgtgcggtc aaaaagaaag 240 gaagaaa ggaaaaagaa aggcgtgcgc acccgagtgc gcgctgagcg cccgctcgcg 300 gtcccgcgga gcctccgcgt tagtccccgc cccgcgccgc gcagtcccccc gggaggcatc 360 gcgcacctct cgccgcccccc tcgcgcctcg ccgattcccc gcctcccctt ttccgcttct 420 tcgccgcctc cgctcgcggc cgcgtcgccc gcgccccgct ccctatctgc tccccagggg 480 ggcactccgc accttttgcg cccgctgccg ccgccgcggc cgccccgccg ccctggtttc 540 ccccgcgagc gcggccgcgt cgccgcgcaa agactcgccg cgtgccgccc cgagcaacgg 600 gtggcggcgg cgcggcggcg ggcggggcgc ggcggcgcgt aggcggggct aggcgccggc 660 taggcgaaac gccgccccccg gggcgccgccg ccgcccgctc cagagcagtc gccgcgccag 720 accgccaacg cagagaccga gaccgaggta cgtcgcgccc gagcacgccg cgacgcgcgg 780 cagggacgag gagcacgacg ccgcgccgcg ccgcgcgggg ggggggaggg agaggcagga 840 cgcgggagcg agcgtgcatg tttccgcgcg agacgacgcc gcgcgcgctg gagaggagat aaggcgcttg gatcgcgaga gggccagcca ggctggaggc gaaaatgggt ggagaggata 960 gtatcttgcg tgcttggacg aggagactga cgaggaggac ggatacgtcg atgatgatgt gcacagagaa gaagcagttc gaaagcgact actagcaagc aagggatcca tgaagttcgc gacctcggtc gcaattttgc ttgtggccaa catagccacc gccctcgcgt cgatgaccaa cgagacctcg gaccgccctc tcgtgcactt tacccccaac aagggttgga tgaacgatcc caacggcctc tggtacgacg aagaggatgc taagtggcac ctttactttc agtacaaccc 1260 1320. taacgacacc gtctggggca ccccgctctt ctggggccac gccacctccg acgacctcac 1380. caactgggag gaccagccca ttgctatcgc ccccaagcgc aacgactcgg gagctttttc cggttccatg gttgtggact acaacaacac ctccggtttt tttaacgaca ccattgaccc ccgccagcgc tgcctcgcca tctggaccta caacacgccc gagagcgagg agcagtacat 1500 cagctacagc cttgatggag gctacacctt taccgagtac cagaagaacc ctgtcctcgc 1560 cgccaactcc acccagttcc gcgaccctaa ggttttttgg tacgagcctt cccagaagtg 1620 gattatgacc gccgctaagt cgcaggatta caagatcgag atctacagca gcgacgacct 1680 caagtcctgg aagcttgagt ccgcctttgc caacgagggt tttctcggat accagtacga 1740 gtgccccggt ctcatcgagg tccccaccga gcaggacccg tccaagtcct actgggtcat 1800 gtttatttcc atcaaccctg gcgcccctgc cggcggcagc ttcaaccagt acttcgtcgg 1860 ctcctttaac ggcacgcatt ttgaggcctt cgacaaccag tcccgcgtcg tcgacttcgg 1920 caaggactac tacgccctcc agaccttctt taacaccgac cccacctacg gcagcgccct 1980 cggtattgct tgggcctcca actggagga ctccgctttc gtccccacta acccctggcg 2040 cagctcgatg tccctcgtcc gcaagttttc gcttaacacc gagtaccagg ccaaccccga 2100 gaccgagctt attaacctga aggccgagcc tattctcaac atctccaacg ctggcccctg 2160 gtcccgcttt gctactaaca ctaccctcac caaggccaac tcctacaacg tcgatctctc 2220 caactccacc ggtactcttg agtttgagct cgtctacgcc gtcaacacca cccagaccat 2280 ctccaagtcc gtcttcgccg acctctccct ctggttcaag ggccttgagg accccgagga 2340 gtacctgcgc atgggttttg aggtctccgc ctcctccttc ttcctcgatc gcggtaactc 2400 caaggttaag tttgtcaagg agaaccccta ctttactaac cgtatgagcg tcaacaacca 2460 gccctttaag tccgagaacg atcttagcta ctacaaggtt tacggcctcc tcgaccagaa 2520 cattctcgag ctctacttta acgacggaga tgtcgtcagc accaacacct actttatgac 2580 cactggaaac gccctcggca gcgtgaacat gaccaccgga gtcgacaacc tcttttacat 2640 tgacaagttt caggttcgcg aggttaagta acatatgtta tgagagatcc gaaagtgaac 2700 cttgtcctaa cccgacagcg aatggcggga gggggcgggc taaaagatcg tattacatag 2760 tatttttccc ctactctttg tgtttgtctt tttttttttt ttgaacgcat tcaagccact 2820 tgtctgggtt tacttgtttg tttgcttgct tgcttgcttg cttgcctgct tcttggtcag 2880 acggcccaaa aaagggaaaa aattcattca tggcacagat aagaaaaaga aaaagtttgt 2940 cgaccaccgt catcagaaag caagagaaga gaacactcg cgctcacatt ctcgctcgcg 3000 taagaatctt agccacgcat acgaagtaat ttgtccatct ggcgaatctt tacatgagcg 3060 ttttcaagct ggagcgtgag atcatacctt tcttgatcgt aatgttccaa ccttgcatag 3120 gcctcgttgc gatccgctag caatgcgtcg tactcccgtt gcaactgcgc catcgcctca 3180 ttgtgacgtg agttcagatt ctttcgaga ccttcgagcg ctgctaattt cgcctgacgc 3240 tccttctttt gtgcttccat gacacgccgc ttcaccgtgc gttccacttc ttcctcagac 3300 atgcccttgg ctgcctcgac ctgctcggta agcttcgtcg taatctctc gatctcggaa 3360 ttcttcttgc cctccatcca ctcggcacca tacttggcag cctgttcaac acgctcattg 3420 aaaaaactttt cattctcttc cagctccgca acccgcgctc gaagctcatt cacttccgcc 3480 accacggctt cggcatcgag cgccgaatca gtcgccgaac tttccgaaag atacaccacg 3540 gcccctccgc tgctgctgcg cagcgtcatc atcagtcgcg tgttatcttc gcgcagattc 3600 tccacctgct ccgtaagcag cttcacggtg gcctcttgat tctgagggct cacgtcgtgg 3660 attagcgctt gcagctcttg cagctccgtc agcttggaag agctcgtaat catggctttg 3720 cacttgtcca gacgtcgcag agcgttcgag agccgcttcg cgttatctgc catggacgct 3780 tctgcgctcg cggcctccct gacgacagtc tcttgcagtt tcactagatc atgtccaatc 3840 agcttgcggt gcagctctcc aatcacgttc tgcatcttgt ttgtgtgtcc gggccgcgcc 3900 tcgtcttgcg atttgcgaat ttcctcctcg agctcgcgtt cgagctccag ggcgccttta 3960 agtagctcga agtcagccgc cgttagcccc agctccgtcg ccgcgttcag acagtcggtt 4020 agcttgattc gattccgctt ttccatggca agtttaagat cctggcccag ctgcacctcc 4080 tgcgccttgc gcatcatgcg cggttccgcc tggcgcaaaa gcttcgagtc gtatcctgcc 4140 tgccatgcca gcgcaatggc acgcacgagc gacttgagtt gccaactatt catcgccgag 4200 atgagcagca ttttgatctg catgaacacc tcgtcagagt cgtcatcctc tgcctcctcc 4260 agctctgcgg gcgagcgacg ctctccttgc agatgaagcg agggccgcag gcctccgaag 4320 agcacctctt gcgcgagatc ctcctccgtc gtcgccctcc gcaggattgc ggtcgtgtcc 4380 gccatcttgc cgccacagca gcttttgctc gctctgcacc ttcaatttct ggtgccgctg 4440 gtgccgctgg tgccgcttgt gctggtgctg gtgctggtgc tggtgctggt gccttgtgct 4500 ggtgctgcca cagacaccgc cgctcctgct gctgctcttc cggccccctc gccgccgccg 4560 cgagcccccg ccgcgcgccg tgcctgggct ctccgcgctc tccgcgggct cctcggcctc 4620 ggcctcgccg tccgcgacga cgtctgcgcg gccgatggtg cggatctgct ctagagggcc 4680 cttcgaaggt aagcctatcc ctaaccctct cctcggtctc gattctacgc gtaccggtca 4740 tcatcaccat caccattgag tttaaacggg ccccagcacg tgctacgaga tttcgattcc 4800 accgccgcct tctatgaaag gttgggcttc ggaatcgttt tccgggacgc cggctggatg 4860 atcctccagc gcggggatct catgctggag ttcttcgccc accccaactt gtttattgca 4920 gcttataatg gttacaaata aagcaatagc atcacaaatt tcacaaataa agcatttttt 4980 tcactgcatt ctagttgtgg tttgtccaaa ctcatcaatg tatcttatca tacatggtcg 5040 acctgcagga acctgcatta atgaatcggc caacgcgcgg ggagaggcgg tttgcgtatt 5100 gggcgctctt ccgcttcctc gctcactgac tcgctgcgct cggtcgttcg gctgcggcga 5160 gcggtatcag ctcactcaaa ggcggtaata cggttatcca cagaatcagg ggataacgca 5220 ggaaagaaca tgtgagcaaa aggccagcaa aaggccagga accgtaaaaa ggccgcgttg 5280 ctggcgtttt tccataggct ccgcccccct gacgagcatc acaaaaatcg acgctcaagt 5340 cagaggtggc gaaacccgac aggactataa agataccagg cgtttccccc tggaagctcc 5400 ctcgtgcgct ctcctgttcc gaccctgccg cttaccggat acctgtccgc ctttctccct 5460 tcgggaagcg tggcgctttc tcatagctca cgctgtaggt atctcagttc ggtgtaggtc 5520 gttcgctcca agctgggctg tgtgcacgaa ccccccgttc agcccgaccg ctgcgcctta 5580 tccggtaact atcgtcttga gtccaacccg gtaagacacg acttatcgcc actggcagca 5640 gccactggta acaggattag cagagcgagg tatgtaggcg gtgctacaga gttcttgaag 5700 tggtggccta actacggcta cactagaaga acagtatttg gtatctgcgc tctgctgaag 5760 ccagttacct tcggaaaaag agttggtagc tcttgatccg gcaaacaaac caccgctggt agcggtggtt tttttgtttg caagcagcag attack gaaaaaaagg attack gatcctttga tcttttctac ggggtctgac gctcagtgga acgaaaactc acgttaaggg 5940 attttggtca tgagattatc aaaaaggatc ttcacctaga tccttttaaa ttaaaaatga agttttaaat caatctaaag fathers taaacttggt ctgacagtta ccaatgctta atcagtgagg cacctatctc agcgatctgt ctattcgtt catccatagt tgcctgactc 6120. cccgtcgtgt agataactac gatacgggag ggcttaccat ctggccccag tgctgcaatg ataccgcgag acccacgctc accggctcca gatttatcag caataacca gccagccgga agggccgagc gcagaagtgg tcctgcaact ttatccgcct ccatccagtc tattaattgt 6300. tgccgggag ctaggtaag tagttcgcca gttaatagtt tgcgcaacgt tgttgccatt gctacaggca tcgtggtgtc acgctcgtcg tttggtatgg cttcattcag ctccggttcc 6420 caacgatcaa ggcgagttac atgatccccc atgttgtgca aaaaagcggt tagctccttc ggtcctccga tcgttgtcag aagtaagttg gccgcagtgt tatcactcat ggttatggca 6540 gcactgcata attctcttac tgtcatgcca tccgtaagat gcttttctgt gactggtgag 6600 tactcaacca agtcattctg agaatagtgt atgcggcgac cgagttgctc ttgcccggcg 6660 tcaatacggg samaaccgc gccacatagc agaactttaa aagtgctcat cattggaaaa 6720 cgttcttcgg ggcgaaaact ctcaaggatc ttaccgctgt tgagatccag ttcgatgtaa 6780 cccactcgtg cacccaactg atcttcagca tcttttactt tcaccagcgt ttctgggtga 6840 gcaaaaacag gaaggcaaaa tgccgcaaaa aagggaataa gggcgacacg gaaatgttga 6900 atactcatac tcttccttt tcaatattat tgaagcattt atcagggtta ttgtctcatg 6960 agcggataca tatttgaatg tatttagaaa aataaacaaa taggggttcc gcgcacattt 7020 ccccgaaaag tgccacctga cgtctaagaa accattatta tcatgacatt aacctataaa 7080 aataggcgta tcacgaggcc ctttcgtctc gcgcgtttcg gtgatgacgg tgaaaacctc 7140 tgacacatgc agctcccgga gacggtcaca gcttgtctgt aagcggatgc cgggagcaga 7200 caagcccgtc agggcgcgtc agcgggtgtt ggcgggtgtc ggggctggct taactatgcg 7260 gcatcagagc agattgtact gagagtgcac caagctttgc ctcaacgcaa ctaggcccag 7320 gcctactttc actgtgtctt gtcttgcctt tcacaccgac cgagtgtgca caaccgtgtt 7380 ttgcacaaag cgcaagatgc tcactcgact gtgaagcaaa ggttgcgcgc aagcgactgc 7440 gactgcgagg atgaggatga ctggcagcct gttcaaaaac tgaaaatccg cgatgggtca 7500 gctgccattc gcgcatgacg cctgcgagag acaagttaac tcgtgtcact ggcatgtcct 7560 agcatcttta cgcgagcaaa attcaatcgc tttatttttt cagtttcgta accttctcgc 7620 aaccgcgaat cgccgtttca gcctgactaa tctgcagctg cgtggcactg tcagtcagtc 7680 agtcagtcgt gcgcgctgtt ccagcaccga ggtcgcgcgt cgccgcgcct ggaccgctgc 7740 tgctactgct agtggcacgg caggtaggag cttgttgccg gaacaccagc agccgccagt 7800 cgacgccagc caggggaaag tccggcgtcg aagggagagg aaggcggcgt gtgcaaacta 7860 acgttgacca ctggcgcccg ccgacacgag caggaagcag gcagctgcag agcgcagcgc 7920 gcaagtgcag aatgcgcgaa agatccactt gcgcgcggcg gcgcgcact tgcgggcgcg 7980 gcgcggaaca gtgcggaaag gagcggtgca gacggcgcgc agtgacagtg ggcgcaaagc 8040 cgcgcagtaa gcagcggcgg ggaacggtat acgcagtgcc gcgggccgcc gcacacagaa 8100 gtatacgcgg gccgaagtgg ggcgtcgcgc gcgggaagtg cggaatggcg ggcaggaaa 8160 Gaggagacg gaagaggggc gggagagaga gagagaga gtgaaaaaag aaaaagaaaaaaaaa 8220 agaagaaag aagaaagct cggagccacg ccgcggggg aggagaaat gaagcacgg 8280 cacggcaag CAAGCAAG cagacccagc cgaggcagc cgagggagga gcgcgcgcag 8340 gacccgcgcg gcgagcgagc gagcacggcg cgcgagcgag cgagcgagcg agcgcgcg 8400 cgagcaaggc tgctgcgag cgatcgagcg agcgagcggg aaggatgagc gcgacccgcg 8460 cggcgacgag gagagcggcg gcgctgtcct cggcgctgac gacgcctgta aagcagcagc 8520 agcagcagca gctgcgcgta gcggcggcgt cggcacggct gcggccgcg gcgttctcgt 8580 ccggcacgggg cggagacgcg gccagaagg cggccgcggc gaggcggttc tccacgggac 8640 gcggccccaa cgcgacacgc gagaagagct cgctggccac ggtccaggcg gcgacggacg 8700 atgcgcgctt cgtcggcctg accggcgccc aaatctttca tgagctcatg cgcgagcacc 8760 aggtggacac catctttggc taccctggcg gcgccattct gcccgttttt gatgccattt 8820 ttgagagtga cgcgcttcaa gttcattctc gctcgccacg agcagggcgc cggccacatg 8880 gccgagggct acgcgcgcgc cacgggcaag cccggcgttg tcctcgtcac ctcgggccct 8940 ggagccacca acaccatcac cccgatcatg gatgcttaca tggacggtac gccgctgctc 9000 gtgttcaccg gccaggtgca gacctctgct gtcggcacgg acgctttcca ggagtgtgac 9060 attgttggca tcagccgcgc gtgcaccaag tggaacgtca tggtcaagga cgtgaaggag 9120 ctcccgcgcc gcatcaatga ggcctttgag attgccatga gcggccgccc gggtcccgtg 9180 ctcgtcgatc ttcctaagga tgtgaccgcc gttgagctca aggaaatgcc cgacagctcc 9240 ccccaggttg ctgtgcgcca gaagcaaaag gtcgagcttt tccacaagga gcgcattggc 9300 gctcctggca cggccgactt caagctcatt gccgagatga tcaaccgtgc ggagcgaccc 9360 gtcatctatg ctggccaggg tgtcatgcag agcccgttga atggcccggc tgtgctcaag 9420 gagttcgcgg agaaggccaa cattcccgtg accaccacca tgcagggtct cggcggcttt 9480 gacgagcgta gtcccctctc cctcaagatg ctcggcatgc acggctctgc ctacgccaac 9540 tactcgatgc agaacgccga tcttatcctg gcgctcggtg cccgctttga tgatcgtgtg 9600 acgggccgcg ttgacgcctt tgctccggag gctcgccgtg ccgagcgcga gggccgcggt 9660 ggcatcgttc actttgagat ttcccccaag aacctccaca aggtcgtcca gcccaccgtc 9720 gcggtcctcg gcgacgtggt cgagaacctc gccaacgtca cgccccacgt gcagcgccag 9780 gagcgcgagc cgtggtttgc gcagatcgcc gattggaagg agaagcaccc ttttctgctc 9840 gagtctgttg attcggacga caaggttctc aagccgcagc aggtcctcac ggagcttaac 9900 aagcagattc tcgagattca ggagaaggac gccgaccagg aggtctacat caccacgggc 9960 gtcggaagcc accagatgca ggcagcgcag ttccttacct ggaccaagcc gcgccagtgg 10020 atctcctcgg gtggcgccgg cactatgggc tacggccttc cctcggccat tggcgccaag 10080 attgccaagc ccgatgctat tgttattgac atcgatggtg atgcttctta ttcgatgacc 10140 ggtatggaat tgatcacagc agccgaattc aaggttggcg tgaagattct tcttttgcag 10200 aacaactttc agggcatggt caagaacgtt caggatctct tttacgacaa gcgctactcg 10260 ggccaccgcc atgttcaacc cgcgcttcga caaggtcgcc gatgcgatgc gtgccaaggg 10320 tctctactgc gcgaaacagt cggagctcaa ggacaagatc aaggatttc tcgaagtacga 10380 tgagggtccc gtcctctcg aggttttcgt ggacaaggac acgctcgtct tgcccatggt 10440 ccccgctggc tttccgctcc acgagatggt cctcgagcct cctaagccca aggacgccta 10500 agttcttttt tccatggcgg gcgagcgagc gagcgcgcga gcgcgcaagt gcgcaagcgc 10560 cttgccttgc tttgcttcgc ttcgctttgc tttgcttcac acaacctaag tatgaattca 10620 agttttcttg cttgtcggcg atgcctgcct gccaaccagc cagccatccg gccggccgtc 10680 cttgacgcct tcgcttccgg cgcggccatc gattcaattc acccatccga tacgttccgc 10740 cccctcacgt ccgtctgcgc acgacccctg cacgaccacg ccaaggccaa cgcgccgctc 10800 agctcagctt gtcgacgagt cgcacgtcac atatctcaga tgcatttgga ctgtgagtgt 10860 tattatgcca ctagcacgca acgatcttcg gggtcctcgc tcattgcatc cgttcgggcc 10920 ctgcaggcgt ggacgcgagt cgccgccgag acgctgcagc aggccgctcc gacgcgaggg 10980 ctcgagctcg ccgcgcccgc gcgatgtctg cctggcgccg actgatctct ggagcgcaag 11040 gaagacacgg cgacgcgagg aggaccgaag agagacgctg gggtatgcag gatatacccg 11100 gggcgggaca ttcgttccgc atacactccc ccattcgagc ttgctcgtcc ttggcagagc 11160 cgagcgcgaa cggttccgaa cgcggcaagg attttggctc tggtgggtgg actccgatcg 11220 aggcgcaggt tctccgcagg ttctcgcagg ccggcagtgg tcgttagaaa tagggagtgc 11280 cggagtcttg acgcgcctta gctcactctc cgcccacgcg cgcatcgccg ccatgccgcc 11340 gtcccgtctg tcgctgcgct ggccgcgacc ggctgcgcca gagtacgaca gtgggacaga 11400 gctcgaggcg acgcgaatcg ctcgggttgt aagggtttca agggtcgggc gtcgtcgcgt 11460 gccaaagtga aaatagtagg gggggggggg ggtac 11495 <210> 59 <211> 32 <212> PRT <213> Schizochytrium <400> 59 Met Arg Thr Val Arg Gly Pro Gln Thr Ala Ala Leu Ala Ala Leu Leu 1 5 10 15 Ala Leu Ala Ala Thr His Val Ala Val Ser Pro Phe Thr Lys Val Glu 20 25 30 <210> 60 <211> 96 <212> DNA <213> Schizochytrium <400> 60 atgcgcacgg tgagggggcc gcaaacggcg gcactcgccg cccttctggc acttgccgcg 60 acgcacgtgg ctgtgagccc gttcaccaag gtggag 96 <210> 61 <211> 30 <212> PRT <213> Schizochytrium <400> 61 Met Gly Arg Leu Ala Lys Ser Leu Val Leu Leu Thr Ala Val Leu Ala 1 5 10 15 Val Ile Gly Gly Val Arg Ala Glu Glu Asp Lys Ser Glu Ala 20 25 30 <210> 62 <211> 90 <212> DNA <213> Schizochytrium <400> 62 atgggccgcc tcgcgaagtc gcttgtgctg ctgacggccg tgctggccgt gatcggaggc 60 gtccgcgccg aagaggacaa gtccgaggcc 90 <210> 63 <211> 38 <212> PRT <213> Schizochytrium <400> 63 Put Thr Ser Thr Ala Arg Ala Leu Ala Leu Val Arg Ala Leu Val Leu 1 5 10 15 Ala Leu Ala Val Leu Ala Leu Leu Ala Ser Gln Ser Val Ala Val Val Asp 20 25 30 Arg Lys Lys Phe Arg Thr 35 <210> 64 <211> 114 <212> DNA <213> Schizochytrium <400> 64 atgacgtcaa cggcgcgc gctcgcgctc gtgcgtgctt tggtgctcgc tctgctgtc 60 ttggcgctgc tagcgagcca aagcgtggcc gtggaccgca aaaagttcag gacc 114 <210> 65 <211> 34 <212> PRT <213> Schizochytrium <400> 65 Put Leu Arg Leu Lys Pro Leu Leu Leu Leu Phe Leu Cys Ser Leu Ile 1 5 10 15 Ala Ser Pro Val Val Ala Trp Ala Arg Gly Gly Glu Gly Pro Ser Thr 20 25 30 Ser Glu <210> 66 <211> 102 <212> DNA <213> Schizochytrium <400> 66 atgttgcggc tcaagccact tttactcctc ttcctctgct cgttgattgc ttcgcctgtg 60 gttgcctggg caagaggagg agaagggccg tccacgagcg aa 102 <210> 67 <211> 29 <212> PRT <213> Schizochytrium <400> 67 Put Ala Lys Ile Leu Arg Ser Leu Leu Leu Ala Ala Val Leu Val Val 1 5 10 15 Thr Pro Gln Ser Leu Arg Ala His Ser Thr Arg Asp Ala 20 25 <210> 68 <211> 87 <212> DNA <213> Schizochytrium <400> 68 atggccaaga tcttgcgcag tttgctcctg gcggccgtgc tcgtggtgac tcctcaatca 60<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> ctgcgtgctc attcgacgcg ggacgca 87<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <210> 69<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <211> 36<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <212> PRT<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <213> Schizochytrium<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <400> 69<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> Met Val Phe Arg Arg Val Pro Trp His Gly Ala Ala Thr Leu Ala Ala<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 1 5 10 15<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> Leu Val Val Ala Cys Ala Thr Cys Leu Gly Leu Gly Leu Asp Ser Glu<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 20 25 30<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> Glu Ala Thr Tyr<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 35<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <210> 70<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <211> 108<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <212> DNA<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <213> Schizochytrium<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <400> 70<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> atggtgtttc ggcgcgtgcc atggcacggc gcggcgacgc tggcggcctt ggtcgtggcc 60<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> tgcgcgacgt gtttaggcct gggactggac tcggaggagg ccacgtac 108<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <210> 71<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <211> 30<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <212> PRT<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <213> Schizochytrium<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> <400> 71<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> Met Thr Ala Asn Ser Val Lys Ile Ser Ile Val Ala Val Leu Val Ala<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 1 5 10 15<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> Ala Leu Ala Trp Glu Thr Cys Ala Lys Ala Asn Tyr Gln Trp 20 25 30 <210> 72 <211> 90 <212> DNA <213> Schizochytrium <400> 72 atgacagcta actcggtgaa aataagcatc gtggctgtgc tggtcgcggc actggcttgg 60 gaaacatgcg caaaagctaa ctatcagtgg 90 <210> 73 <211> 35 <212> PRT <213> Schizochytrium <400> 73 Put Ala Arg Arg Ala Ser Arg Leu Gly Ala Ala Val Val Val Val Val Leu 1 5 10 15 Val Val Val Ala Ser Ala Cys Cys Trp Gln Ala Ala Ala Asp Val Val 20 25 30 Asp Ala Gln 35 <210> 74 <211> 105 <212> DNA <213> Schizochytrium <400> 74 atggcgcgca gggcgtcgcg cctcggcgcc gcgtcgtcg tcgtcctcgt cgtcgtcgcc 60 tccgcctgct gctggcaagc cgctgcggac gtcgtggacg cgcag 105 <210> 75 <211> 1785 <212> DNA <213> Artificial Sequence <220> <223> Codon optimized nucleic acid sequence <400> 75 atgaagttcg cgacctcggt cgcaattttg cttgtggcca acatagccac cgccctcgcg 60 gcctccccct cgatgcagac ccgtgcctcc gtcgtcattg attacaacgt cgctcctcct 120 aacctctcca ccctcccgaa cggcagcctc tttgagacct ggcgtcctcg cgcccacgtt 180 cttcccccta acggtcagat tggcgatccc tgcctccact acaccgatcc ctcgactggc 240 ctctttcacg tcggctttct ccacgatggc tccggcattt cctccgccac tactgacgac 300 ctcgctacct acaaggatct caaccagggc aaccaggtca tcgtccccgg cggtatcaac 360 gaccctgtcg ctgttttcga cggctccgtc attccttccg gcattaacgg cctccctacc 420 ctcctctaca cctccgtcag ctacctcccc attcactggt ccatccccta cacccgcggt 480 tccgagacgc agagcctggc tgtctccagc gatggtggct ccaactttac taagctcgac 540 cagggccccg ttattcctgg cccccccttt gcctacaacg tcaccgcctt ccgcgacccc 600 tacgtctttc agaaccccac cctcgactcc ctcctccact ccaagaacaa cacctggtac 660 accgtcattt cgggtggcct ccacggcaag ggccccgccc agtttcttta ccgtcagtac 720 gaccccgact ttcagtactg ggagttcctc ggccagtggt ggcacgagcc taccaactcc 780 acctggggca acggcacctg ggccggccgc tgggccttca acttcgagac cggcaacgtc 840 ttttcgcttg acgagtacgg ctacaacccc cacggccaga tcttctccac cattggcacc 900 gagggctccg accagcccgt tgtcccccag ctcacctcca tccacgatat gctttgggtc 960 tccggtaacg tttcgcgcaa cggatcggtt tccttcactc ccaacatggc cggcttcctc 1020 gactggggtt tctcgtccta cgccgccgcg ggtaaggttc ttccttccac gtcgctcccc 1080 tccaccaagt ccggtgcccc cgatcgcttc atttcgtacg tttggctctc cggcgacctc 1140 tttgagcagg ctgagggctt tcctaccaac cagcagaact ggaccggcac cctcctcctc 1200 ccccgtgagc tccgcgtcct ttacatcccc aacgtggttg ataacgccct tgcgcgcgag 1260 tccggcgctt cctggcaggt cgtctcctcc gatagctcgg ccggtactgt ggagctccag 1320 accctcggca tttccatcgc ccgcgagacc aaggccgccc tcctgtccgg cacctcgttc 1380 actgagtccg accgcactct taactcctcc ggcgtcgttc cctttaagcg ttccccctcc 1440 gagaagtttt tcgtcctctc cgcccagctc tccttccccg cctccgcccg cggctcgggc 1500 ctcaagtccg gcttccagat tctttcctcc gagctcgagt ccaccacggt ctactaccag 1560 tttagcaacg agtccatcat cgtcgaccgc agcaacacca gcgccgccgc ccgtactacc 1620 gacggtatcg actcctccgc cgaggccggc aagctccgcc tctttgacgt cctcaacggc 1680 ggcgagcagg ctattgagac cctcgacctt accctcgtcg ttgataactc cgtgctcgag 1740 atttacgcca acggtcgttt cgcgctttcc acctgggttc gctaa 1785 <210> 76 <211> 1682 <212> DNA <213> Artificial Sequence <220> <223> Codon Optimized HA <400> 76 atgaaggcta acctcctcgt tcttctttcc gctctcgctg ctgcggatgc cgacaccatc 60 tgcattggct accacgctaa caacagcacg gacaccgtcg atactgtcct ggagaagaac 120 gttaccgcac ccattcggtc aacctcctgg aggacagcca caacggcaag ctctgccgtc 180 ttaagggcat cgcccccctc cagctcggca agtgcaacat cgccggctgg ctcctcggca 240 acccggagtg cgatccctcc tccccgttcg ctcctggtcg tacattgtgg agactccgaa 300 cagcgagaac ggtatctgct accccggcga ttttatcgac tacgaggagc tccgcgagca 360 gctctcctcc gtgtccagct tgagcgtttc gagatttttc cgaaggagtc ctcgtggccc 420 aaccacaaca ccaacggcgt caccgccgcc tgctcccacg agggcaagtc gagcttttac 480 cgcaacctgc tttggctcac cgagaagggg gttcgtaccc taagctcaag aactcgtacg 540 tcaacaagaa gggcaaggag gtcctcgtcc tctggggcat ccaccatccc ccgaacagca 600 aggagcagca gaacatctac cagaacgaga acgccacgtt tcggtggtca cgtcgaacta 660 caaccgccgc ttcactcctg agatcgccga gcgccccaag gtgcgcgacc aggctggccg 720 catgaactac tactggaccc tccttaagcc cggtgacacg atatctttga ggccaacggc 780 aaccttatcg cgcccatgta cgcgttcgcc ctctcccgcg gctttggtag cggcatcatt 840 accagcaacg ccagcatgca cgagtgcaac acgaagtgcc agaccccgcc ggtgccatca 900 acagcagcct gccttaccag aacatccacc ccgtcaccat cggtgagtgc ccgaagtacg 960 tgcgctcggc caagctccgc atggtcacgg gcctccgcaa cactcttcg atccagcccg 1020 cggcctcttc ggcgccattg ccggtttcat cgagggcggc tggacgggca tgatcgacgg 1080 ctggtacggc taccaccacc agaacgagca gggctccggt tacgccgcgg accagaagtc 1140 1200 attcagttta ccgctgtcgg caaggagttc aacaagctgg agaagcgcat ggagaacctc 1260 aaaagaagg ggacgatggt ttcctggaca tttggaccta caacgccgag ctcctcgtgc 1320 tccttgagaa cgagcgtacc ctcgacttcc acgactccaa cgtcaagaac ctctacgaga 1380 aggtcaagtc gcagctcaga aaacgccaa ggagattggc aacggttgct tcgagtttta 1440 ccacaagtgc gaacgagt gcatggagtc cgtccgcaac ggcacctacg actacccgaa 1500 gtactccgag gagtcgaagc tgaacgcgag aaggtggacg gcgtgaagct ggagtccatg 1560 ggcatctacc agatcctcgc catttactcg acggttgcct cgtcgctcgt cctccttgtc 1620 tccctcggtg cgatttcgtt ctggatgtgc tgaacggcag ccttcagtgc cgcatctgca 1680 tc 1682 <210> 77 <211> 565 <212> PRT <213> Influenza A virus <400> 77 Met Lys Ala Asn Leu Leu Val Leu Leu Ser Ala Leu Ala Ala Ala Asp 1 5 10 15 Ala Asp Thr Ile Cys Ile Gly Tyr His Ala Asn Asn Ser Thr Asp Thr 20 25 30 Val Asp Thr Val Leu Glu Lys Asn Val Thr Val Thr His Ser Val Asn 35 40 45 Leu Leu Glu Asp Ser His Asn Gly Lys Leu Cys Arg Leu Lys Gly Ile 50 55 60 Ala Pro Leu Gln Leu Gly Lys Cys Asn Ile Ala Gly Trp Leu Leu Gly 65 70 75 80 Asn Pro Glu Cys Asp Pro Leu Leu Pro Val Arg Ser Trp Ser Tyr Ile 85 90 95 Val Glu Thr Pro Asn Ser Glu Asn Gly Ile Cys Tyr Pro Gly Asp Phe 100 105 110 Ile Asp Tyr Glu Glu Leu Arg Glu Gln Leu Ser Ser Val Ser Ser Phe 115 120 125 Glu Arg Phe Glu Ile Phe Pro Lys Glu Ser Ser Trp Pro Asn His Asn 130 135 140 Thr Asn Gly Val Thr Ala Ala Cys Ser His Glu Gly Lys Ser Ser Phe 145 150 155 160 Tyr Arg Asn Leu Leu Trp Leu Thr Glu Lys Glu Gly Ser Tyr Pro Lys 165 170 175 Leu Lys Asn Ser Tyr Val Asn Lys Lys Gly Lys Glu Val Leu Val Leu 180 185 190 Trp Gly Ile His His Pro Pro Asn Ser Lys Glu Gln Gln Asn Ile Tyr 195 200 205 Gln Asn Glu Asn Ala Tyr Val Ser Val Val Thr Ser Asn Tyr Asn Arg 210 215 220 Arg Phe Thr Pro Glu Ile Ala Glu Arg Pro Lys Val Arg Asp Gln Ala 225 230 235 240 Gly Arg Met Asn Tyr Tyr Trp Thr Leu Leu Lys Pro Gly Asp Thr Ile 245 250 255 Ile Phe Glu Ala Asn Gly Asn Leu Ile Ala Pro Met Tyr Ala Phe Ala 260 265 270 Leu Ser Arg Gly Phe Gly Ser Gly Ile Ile Thr Ser Asn Ala Ser Met 275 280 285 His Glu Cys Asn Thr Lys Cys Gln Thr Pro Leu Gly Ala Ile Asn Ser 290 295 300 Ser Leu Pro Tyr Gln Asn Ile His Pro Val Thr Ile Gly Glu Cys Pro 305 310 315 320 Lys Tyr Val Arg Ser Ala Lys Leu Arg Met Val Thr Gly Leu Arg Asn 325 330 335 Thr Pro Ser Ile Gln Ser Arg Gly Leu Phe Gly Ala Ile Ala Gly Phe 340 345 350 Ile Glu Gly Gly Trp Thr Gly Met Ile Asp Gly Trp Tyr Gly Tyr His 355 360 365 His Gln Asn Glu Gln Gly Ser Gly Tyr Ala Ala Asp Gln Lys Ser Thr 370 375 380 Gln Asn Ala Ile Asn Gly Ile Thr Asn Lys Val Asn Thr Val Ile Glu 385 390 395 400 Lys Met Asn Ile Gln Phe Thr Ala Val Gly Lys Glu Phe Asn Lys Leu 405 410 415 Glu Lys Arg Met Glu Asn Leu Asn Lys Lys Val Asp Asp Gly Phe Leu 420 425 430 Asp Ile Trp Thr Tyr Asn Ala Glu Leu Leu Val Leu Leu Glu Asn Glu 435 440 445 Arg Thr Leu Asp Phe His Asp Ser Asn Val Lys Asn Leu Tyr Glu Lys 450 455 460 Val Lys Ser Gln Leu Lys Asn Asn Ala Lys Glu Ile Gly Asn Gly Cys 465 470 475 480 Phe Glu Phe Tyr His Lys Cys Asp Asn Glu Cys Met Glu Ser Val Arg 485 490 495 Asn Gly Thr Tyr Asp Tyr Pro Lys Tyr Ser Glu Glu Ser Lys Leu Asn 500 505 510 Arg Glu Lys Val Asp Gly Val Lys Leu Glu Ser Met Gly Ile Tyr Gln 515 520 525 Ile Leu Ala Ile Tyr Ser Thr Val Ala Ser Ser Leu Val Leu Leu Val 530 535 540 Ser Leu Gly Ala Ile Ser Phe Trp Met Cys Ser Asn Gly Ser Leu Gln 545 550 555 560 Cys Arg Ile Cys Ile 565 <210> 78 <211> 51 <212> PRT <213> Schizochytrium <220> <223> GlcNac-transferase-I-like protein <400> 78 Met Arg Gly Pro Gly Met Val Gly Leu Ser Arg Val Asp Arg Glu His 1 5 10 15 Leu Arg Arg Arg Gln Gln Gln Ala Ala Ser Glu Trp Arg Arg Trp Gly 20 25 30 Phe Phe Val Ala Thr Ala Val Val Leu Leu Val Phe Leu Thr Val Tyr 35 40 45 Pro Ass Val 50 <210> 79 <211> 153 <212> DNA <213> Schizochytrium <220> <223> signal anchor sequence <400> 79 atgcgcggcc cgggcatggt cggcctcagc cgcgtggacc gcgagcacct gcggcggcgg 60 cagcagcagg cggcgagcga atggcggcgc tgggggttct tcgcgcgac ggcgtcgtc 120 ctgctcgtct ttctcaccgt atacccgaac gta 153 <210> 80 <211> 66 <212> PRT <213> Schizochytrium <220> <223> beta-1,2- xylosyltransferase-like protein <400> 80 Met Arg Thr Arg Gly Ala Ala Tyr Val Arg Pro Gly Gln His Glu Ala 1 5 10 15 Lys Ala Leu Ser Ser Arg Ser Ser Asp Glu Gly Tyr Thr Thr Val Asn 20 25 30 Val Val Arg Thr Lys Arg Lys Arg Thr Thr Val Ala Ala Leu Val Ala 35 40 45 Ala Ala Leu Leu Val Thr Gly Phe Ile Val Val Val Val Phe Val Val 50 55 60 Val Val 65 <210> 81 <211> 198 <212> DNA <213> Schizochytrium <220> <223> signal anchor sequence <400> 81 atgcgcacgc ggggcgcggc gtacgtgcgg ccgggacagc acgaggcgaa ggcgctctcg 60 tcaaggagca gcgacgaggg atatacgacg gtcaacgttg tcaggaccaa gcgaaagagg 120 accactgtag ccgcgcttgt agccgcggcg ctgctggtga cgggctttat cgtcgtcgtc 180 gtcttcgtcg tcgttgtt 198 <210> 82 <211> 64 <212> PRT <213> Schizochytrium <220> <223> beta-1,4-xylosidase-like protein <400> 82 Met Glu Ala Leu Arg Glu Pro Leu Ala Ala Pro Pro Thr Ser Ala Arg 1 5 10 15 Ser Ser Val Pro Ala Pro Leu Ala Lys Glu Glu Gly Glu Glu Glu Asp 20 25 30 Gly Glu Lys Gly Thr Phe Gly Ala Gly Val Leu Gly Val Val Ala Val 35 40 45 Leu Val Ile Val Val Phe Ala Ile Val Ala Gly Gly Gly Gly Asp Ile 50 55 60 <210> 83 <211> 192 <212> DNA <213> Schizochytrium <220> <223> signal anchor sequence <400> 83 atggaggccc tgcgcgagcc cttggctgcg ccgccaacgt cggcgcgatc gtcggtgcca 60 gcgccgctcg cgaaggagga gggggaggag gaggacgggg aaaaagggac gtttggggcg 120 ggggtcctcg gtgtcgtggc ggtgctcgtc atcgtggtgt ttgcgatcgt ggcgggaggc 180 ggaggcgata tt 192 <210> 84 <211> 73 <212> PRT <213> Schizochytrium <220> <223> galactosyltransferase-like protein <400> 84 Put Leu Ser Val Ala Gln Val Ala Gly Ser Ala His Ser Arg Pro Arg 1 5 10 15 Arg Gly Gly Glu Arg Met Gln Asp Val Leu Ala Leu Glu Glu Ser Ser 20 25 30 Arg Asp Arg Lys Arg Ala Thr Ala Arg Pro Gly Leu Tyr Arg Ala Leu 35 40 45 Ala Ile Leu Gly Leu Pro Leu Ile Val Phe Ile Val Trp Gln Put Thr 50 55 60 Create Create Leu Thr Thr Wing Pro Create Wing 65 70 <210> 85 <211> 219 <212> DNA <213> Schizochytrium <220> <223> signal anchor sequence <400> 85 atgttgagcg tagcacaagt cgcggggtg gcccactcgc ggccgagacg aggtggtgag 60 cggatgcaag acgtgctggc cctggaggaa agcagcagag atcgaaaacg agcaacagca 120 aggcccgggc tatatcgcgc acttgcgatt ctggggctgc cgctcatcgt attcatcgta 180 tggcaaatga ctagctccct cacgactgcc ccgagcgcc 219 <210> 86 <211> 997 <212> PRT <213> Schizochytrium <220> <223> EMC1 <400> 86 Met Gly Thr Thr Thr Ala Arg Met Ala Val Ala Val Leu Ala Ala Ala 1 5 10 15 Val Ser Val Ala His Gly Leu His Glu Asp Gln Ala Gly Val Asn Asp 20 25 30 Trp Thr Val Arg Asn Leu Gly Ala Tyr Ala His Gly Val Phe Leu Asp 35 40 45 Asp Asp Leu Ala Leu Val Ala Thr Thr Gln Ala Thr Val Gly Ala Val 50 55 60 Arg Met Thr Asp Gly Glu Val Val Trp Arg Glu Thr Leu Pro Thr Ala 65 70 75 80 Arg Ser Ala Pro Leu Ala Ser Gln Val Lys His Glu Leu Phe Ala Thr 85 90 95 Ala Ser Ala Asp Ala Cys Val Ile Glu Leu Trp Ala Thr Pro Ser Gly 100 105 110 Asp Val Met Thr Ser Asp Ser Arg Gln Ala Gly Leu Glu Trp Asp Ala 115 120 125 Lys Ile Cys Asp Asn Thr Asp Ala Asp Ala Thr Gly Val Leu Glu Leu 130 135 140 Leu Asp Asn Asp Phe Asn Asn Asp Gly Thr Pro Asp Val Ala Ala Leu 145 150 155 160 Thr Pro Phe Gln Phe Val Ile Leu Asp Gly Val Ser Gly Arg Val Leu 165 170 175 His Glu Val Asp Leu Asp Lys Thr Ile Ala Trp Gln Gly Leu Val Glu 180 185 190 Ala Ala Gly Ser Ala Thr Gly Gly Lys Arg Lys Arg Pro Ser Ile Met 195 200 205 Ala Tyr Gly Val Asp Ile Lys Thr Gly Lys Leu Glu Val Arg Lys Leu 210 215 220 Ala Asn Ser Gly Ala Thr Leu Asp Pro Val Ser Gly Leu Glu Gly Val 225 230 235 240 Ser Ala Asp Glu Ile Thr Val Leu Lys Ser Gly Val Ala Lys Val Gly 245 250 255 Ser Ala Leu Leu Phe Val Arg Lys Glu Ser Gly Ala Leu Val Ala Phe 260 265 270 Asp Cys Val Ala Asn Gln Leu Gln Glu Leu Thr Asn Ala Pro Ser Ile 275 280 285 Lys Gly Ser Val Gln Ser Leu Gly Ser Ala Arg Phe Phe Ala Thr Asp 290 295 300 Ala Gly Val Ile Tyr Ala Val Asp Gly Glu Leu Lys Ile Ala Glu Thr 305 310 315 320 Leu Lys Gly Val Glu Ala Ala Ala Ile Gly Val Ser Gly Ala Ser Val 325 330 335 Ile Ala Ala Val Gln Ser Ser Thr Ala Ser Gly Thr Gly Asp Glu Ala 340 345 350 Gln Cys Gly Pro Ile Ser Arg Val Leu Val Gln Ser Ala Ser Gly Val 355 360 365 Thr Glu Ile Ala Phe Pro Glu Gln Gln Gly Gln Ser Gly Ala Arg Gly 370 375 380 Leu Val Glu Lys Ile Ile Val Gly Asp Ser Ser Thr Gly Thr Arg Ala 385 390 395 400 Ile Phe Val Phe Glu Asp Ala Ser Ala Val Gly Ile Glu Ile Glu Ser 405 410 415 Gly Ala Ser Glu Ala Ser Thr Leu Phe Val Arg Glu Glu Ala Leu Ala 420 425 430 Asn Val Val Glu Ala Val Ala Val Asp Leu Pro Pro Thr Asp Glu Val 435 440 445 Gly Ser Leu Gly Asp Glu Ala Ala His Val Phe Ala His Gly Ser His 450 455 460 Ala Ser Ile Phe Met Phe Arg Leu Lys Asp Gln Val Arg Thr Val Gln 465 470 475 480 Arg Phe Val Gln Ser Leu Phe Gly Ala Ala Thr Gln His Leu Ser Glu 485 490 495 Phe Val Ala Ser Gln Gly Lys Thr Leu Val Gln Ala Ile Arg Gly Glu 500 505 510 Leu Pro Arg Ala Glu Ser Leu Ser Gln Ser Glu Met Phe Ser Phe Gly 515 520 525 Phe Arg Arg Val Leu Val Leu Arg Ser Ala Ser Gly Lys Val Phe Gly 530 535 540 Leu Asn Ser Ala Asp Gly Ser Leu Leu Trp Ala Ala Gln Ser Pro Gly 545 550 555 560 Ser Arg Leu Phe Val Thr Arg Ala Arg Glu Ala Gly Leu Asp His Pro 565 570 575 Ala Glu Val Ala Ile Val Asp Glu Ala His Gly Arg Val Thr Trp Arg 580 585 590 Asn Ala Ile Thr Gly Ala Val Thr Arg Val Glu Asp Ile Asp Thr Pro 595 600 605 Leu Ala Gln Ile Ala Val Leu Pro Gly Asp Ile Phe Pro Ser Thr Ala 610 615 620 Ser Ser Glu Glu Asp Val Ser Pro Ala Ala Val Leu Ile Ala Leu Asp 625 630 635 640 His Ala Gln Arg Val His Ile Leu Pro Ser Ser Arg Thr Glu Ser Val 645 650 655 Leu Gln Leu Glu Asp Leu Leu Arg Ala Leu His Phe Val Val Tyr Ser 660 665 670 Asn Glu Thr Gly Ala Leu Thr Gly Tyr Ala Val Asp Pro Ser Gln Arg 675 680 685 Ala Gly Val Glu Leu Trp Ser Met Ile Val Pro Ala Ser Gln Thr Leu 690 695 700 Leu Ala Val Glu Gly Gln Ser Gly Gly Ala Leu Asn Asn Pro Gly Ile 705 710 715 720 Lys Arg Gly Asp Gly Ala Val Leu Val Lys Phe Val Asp Pro His Leu 725 730 735 Leu Met Val Ala Thr Gln Ser Gly Pro His Leu Gln Val Ser Ile Leu 740 745 750 Asn Gly Ile Ser Gly Arg Val Ile Ser Arg Phe Thr His Lys Lys Ser 755 760 765 Thr Gly Pro Val His Ala Val Leu Ala Asp Asn Thr Val Thr Tyr Ser 770 775 780 Phe Trp Asn Gln Val Lys Ser Arg Gln Glu Val Ser Val Val Gly Leu 785 790 795 800 Phe Glu Gly Glu Ile Gly Pro Arg Glu Leu Asn Met Trp Ser Ser Arg 805 810 815 Pro Asn Met Gly Ser Gly Lys Ala Met Ser Ala Phe Asp Asp Ser Met 820 825 830 Met Pro Asn Val Gln Gln Lys Thr Phe Tyr Thr Glu Arg Ala Ile Ala 835 840 845 Ala Leu Gly Val Thr Lys Thr Arg Phe Gly Ile Ala Asp Arg Arg Val 850 855 860 Leu Ile Gly Thr Ala Asn Gly Ala Val Asn Met Gln Val Pro Gln Ile 865 870 875 880 Leu Ser Pro Arg Arg Pro Val Gly Lys Leu Ser Asp Met Glu Lys Glu 885 890 895 Glu Gly Leu Met Leu Tyr Ala Pro Glu Leu Pro Leu Ile Pro Thr Gln 900 905 910 Thr Ile Thr Tyr Tyr Glu Ser Ile Pro Gln Leu Arg Leu Ile Arg Ser 915 920 925 Phe Ala Thr Arg Leu Glu Ser Thr Ser Leu Val Leu Ala Ala Gly Leu 930 935 940 Asp Ile Phe Tyr Thr Arg Val Met Pro Ser Arg Gly Phe Asp Val Leu 945 950 955 960 Asp Glu Asp Phe Ala Ser Gly Leu Leu Leu Ala Leu Ile Ala Ala Leu 965 970 975 Leu Ala Leu Thr Ile Tyr Leu Ser Lys Ala Val Gly Lys Ser Thr Leu 980 985 990 Asp Glu Thr Trp Lys 995 <210> 87 <211> 731 <212> PRT <213> Schizochytrium <220> <223> Nicastrin-like <400> 87 Met Gly Ala Ala Arg Arg Ser Met Gly Ala Ala Arg Lys Ala Leu Ala 1 5 10 15 Ala Ser Ala Thr Leu Ala Ala Leu Ala Leu Ala Gly Leu Gln Pro Ala 20 25 30 Arg Ala Glu Val Asn Gly Val Asn Ala Met Thr Glu Ala Met Leu Thr 35 40 45 Glu Tyr Ala Ser Leu Pro Cys Val Arg Ser Ile Ala Arg Asp Gly Ala 50 55 60 Val Gly Cys Gly Ser Pro Ser Asp Arg Ser Val Ala Glu Gly Gly Ala 65 70 75 80 Leu Phe Leu Val Glu Ser Val Glu Asp Val Thr Gly Leu Ile Glu Asn 85 90 95 Ala Gln Gly Leu Asp Ala Val Ala Leu Val Val Asp Asp Ala Leu Leu 100 105 110 His Gly Asp Ser Leu Arg Ala Met Gln Asp Leu Ala Lys Lys Ile Arg 115 120 125 Val Thr Ala Val Ile Val Thr Val Glu Glu Asp Gly Ser Pro Gln Glu 130 135 140 Pro Pro Arg Ser Ser Ala Ala Pro Thr Thr Trp Ile Pro Ser Gly Asp 145 150 155 160 Gly Leu Leu Asn Glu Thr Val Ser Phe Val Val Thr Arg Leu Arg Asn 165 170 175 Ala Thr Gln Ser Glu Glu Ile Arg Ala Leu Ala Ala Ser Asn Arg Asp 180 185 190 Arg Gly Tyr Val Asp Ala Val Phe Gln His Ser Ala Arg Tyr Gln Phe 195 200 205 Tyr Leu Gly Lys Glu Thr Ala Thr Ser Leu Ser Cys Leu Ala Ser Gly 210 215 220 Arg Cys Asp Pro Leu Gly Gly Leu Ser Val Trp Ala Ser Ala Gly Pro 225 230 235 240 Val Pro Val Asn Ser Ala Lys Glu Thr Val Leu Leu Thr Ala Asn Leu 245 250 255 Asp Ala Ala Ser Phe Phe His Asp Val Val Pro Ala Arg Asp Thr Thr 260 265 270 Ala Ser Gly Val Ala Ala Val Leu Leu Ala Ala Lys Ala Leu Ala Ser 275 280 285 Val Asp Glu Ser Val Leu Glu Ala Leu Ser Lys Gln Ile Ala Val Ala 290 295 300 Leu Phe Asn Gly Glu Val Trp Ser Arg Ala Gly Ser Arg Arg Phe Val 305 310 315 320 His Asp Val Ala Leu Gly Glu Cys Leu Ser Pro Gln Thr Ala Ser Pro 325 330 335 Tyr Asn Glu Ser Thr Cys Ala Asn Pro Pro Val Tyr Ala Leu Ala Trp 340 345 350 Thr Ser Leu Gly Leu Asp Asn Ile Thr Asp Val Val Ser Val Asn Asn 355 360 365 Val Ala Gly Ser Glu Ser Gly Ala Phe Tyr Val His Thr Ala Ala Gly 370 375 380 Thr Ala Ser Ala Asn Ala Ala Ala Ala Leu Gln Ser Val Ala Ser Ser 385 390 395 400 Ser Thr Asp Val Asp Val Ser Ile Thr Gly Ala Thr Thr Ser Gly Val 405 410 415 Val Pro Pro Ser Pro Leu Asp Ser Phe Leu Ala Ala Glu Met Glu Thr 420 425 430 Asp Val Ser Phe Ser Gly Ala Gly Leu Val Val Ser Gly Phe Asp Ala 435 440 445 Ala Ile Thr Asp Ala Asn Pro Arg Tyr Ser Ser Arg Tyr Asp Arg Arg 450 455 460 Asp Lys Gly Pro Glu Ala Asp Asp Ala Glu Ala Leu Thr Ala Ala Arg 465 470 475 480 Ile Ala Asp Val Ala Thr Leu Leu Ala Arg His Ala Phe Val Gln Ala 485 490 495 Gly Gly Ser Ile Ser Asp Ala Val Asn Phe Val Leu Val Asp Gly Thr 500 505 510 His Ala Ala Glu Leu Trp Asp Cys Leu Thr Lys Asp Phe Ala Cys Thr 515 520 525 Leu Val Ala Asp Val Ile Gly Ala Glu Asp Thr Thr Ala Val Ala Asp 530 535 540 Phe Met Gly Ser Thr Leu Leu Ala Ala Ser Glu Gly Val Ala Gly Gly 545 550 555 560 Ala Pro Asn Phe Phe Ser Gly Ile Tyr Ser Pro Phe Pro Val Glu Asn 565 570 575 Asn Val Met Arg Pro Val Pro Leu Phe Val Arg Asp Tyr Leu Ala Gln 580 585 590 Tyr Gly Arg Asn Ala Ser Leu Ile Glu Lys Val Thr Glu Ser Ala Lys 595 600 605 Tyr Ala Cys Ala Gln Asp Leu Asp Cys Met Val Met Thr Glu Pro Pro 610 615 620 Ala Cys Glu Leu Gly Arg Ser Ala Leu Ala Cys Leu Arg Gly Gly Cys 625 630 635 640 Val Cys Ser Asn Ala Tyr Phe His Asp Ala Val Ser Pro Ala Leu Val 645 650 655 Tyr Glu Asp Gly Ala Phe Ser Val Asp Ala Gln Lys Leu Thr Asp Asp 660 665 670 Asp Gly Leu Trp Thr Glu Pro Arg Trp Ser Asp Gly Thr Leu Thr Leu 675 680 685 Tyr Thr Ser Ala Asn Ser Ala Ser Thr Thr Ile Ala Leu Leu Val Cys 690 695 700 Gly Ile Leu Leu Thr Ile Gly Cys Val Phe Ala Leu Arg Lys Ala Gln 705 710 715 720 Gly Met Leu Asp Asn Thr Lys Tyr Lys Leu Asn 725 730 <210> 88 <211> 232 <212> PRT <213> Schizochytrium <220> <223> Emp24 <400> 88 Met Ala Thr Thr Glu Asn Glu Ala Arg Leu Pro Pro Gly Lys Gln Arg 1 5 10 15 Leu Gly Arg Arg Arg Arg Gly Arg Val Ser Lys Ala Ser Gly Trp Gly 20 25 30 Thr Thr Leu Ala Leu Ala Ala Ala Val Leu Val Phe Ser Val Asp Arg 35 40 45 Ala Ser Gly Val Arg Phe Glu Val Ala Ser Thr Glu Glu Arg Cys Ile 50 55 60 Phe Asp Val Leu Arg Lys Asp Gln Leu Val Thr Gly Glu Phe Glu Val 65 70 75 80 His Ala Asp Gly Asp Asp Val Asn Met Asp Ile His Val Thr Gly Pro 85 90 95 Leu Gly Glu Glu Val Phe Ser Lys Gln Asn Ser Lys Met Ala Lys Phe 100 105 110 Gly Phe Thr Ala Glu Ala Ala Gly Glu His Val Leu Cys Leu Arg Asn 115 120 125 Asn Asp Met Ile Met Arg Glu Val Gln Val Lys Leu Arg Ser Gly Val 130 135 140 Glu Ala Lys Asp Leu Thr Glu Val Val Gln Arg His His Leu Lys Pro 145 150 155 160 Leu Ser Ala Glu Val Ile Arg Ile Gln Glu Thr Ile Arg Asp Val Arg 165 170 175 His Glu Leu Thr Ala Leu Lys Gln Arg Glu Ala Glu Met Arg Asp Met 180 185 190 Asn Glu Ser Ile Asn Thr Arg Val Ser Leu Phe Ser Phe Phe Ser Ile 195 200 205 Ala Val Val Gly Ser Leu Gly Ala Trp Gln Ile Met Tyr Leu Lys Ser 210 215 220 Tyr Phe Gln Arg Lys Lys Leu Ile 225 230 <210> 89 <211> 550 <212> PRT <213> Schizochytrium <220> <223> Calnexin-like <400> 89 Met Arg Thr Thr Phe Val Ala Ala Tyr Ala Ala Val Ala Ala Leu Ala 1 5 10 15 Leu Gly Gln Cys Glu Ala Ile Asn Phe Arg Glu Ser Phe Glu Gly Ala 20 25 30 Asn Val Glu Lys Glu Trp Val Lys Ser Ala Ser Asp Arg Tyr Ala Gly 35 40 45 Ser Glu Trp Ala Phe Asp Thr Ser Lys Asp Thr Gly Asp Val Gly Leu 50 55 60 Gln Thr Val Lys Pro His Lys Phe Tyr Gly Ile Ser Arg Lys Phe Glu 65 70 75 80 Asn Pro Ile Pro Val Gly Asp Gly Glu Lys Pro Phe Val Ala Gln Tyr 85 90 95 Glu Val Lys Phe Thr Glu Gly Val Ser Cys Ser Gly Ala Tyr Leu Lys 100 105 110 Leu Leu Glu Gln Asp Asp Ala Phe Thr Pro Lys Asp Leu Val Glu Ser 115 120 125 Ser Pro Tyr Ser Ile Met Phe Gly Pro Asp Asn Cys Gly Ala Asn Asn 130 135 140 Lys Val His Leu Ile Phe Arg Gln Glu Asn Pro Val Thr Lys Glu Tyr 145 150 155 160 Glu Glu Lys His Met Thr Lys Lys Val Thr Ser Val Arg Asp Arg Thr 165 170 175 Ser His Val Tyr Thr Leu Glu Val His Pro Asp Asn Thr Phe Lys Val 180 185 190 Lys Val Asp Gly Lys Val Glu Ala Glu Gly Ser Leu Thr Asp Asp Glu 195 200 205 Ala Phe Ser Pro Pro Phe Gln Gln Pro Lys Glu Ile Asp Asp Pro Asn 210 215 220 Asp Glu Lys Pro Asp Asp Trp Val Asp Gln Ala Lys Ile Pro Asp Pro 225 230 235 240 Glu Ala Ser Lys Pro Asp Asp Trp Asp Glu Asp Ala Pro Lys Arg Ile 245 250 255 Ala Asp Pro Asp Ala Val Lys Pro Glu Gly Trp Leu Asp Asp Glu Pro 260 265 270 Asp Gln Val Pro Asp Pro Ala Ala Ser Glu Pro Glu Asp Trp Asp Glu 275 280 285 Glu Asp Asp Gly Ile Trp Glu Ala Pro Leu Val Ala Asn Pro Lys Cys 290 295 300 Thr Ala Gly Pro Gly Cys Gly Glu Trp Asn Ala Pro Met Ile Glu Asn 305 310 315 320 Pro Asn Tyr Lys Gly Lys Trp Ser Ala Pro Met Ile Asp Asn Pro Glu 325 330 335 Tyr Lys Gly Val Trp Lys Pro Arg Arg Ile Glu Asn Pro Ala Tyr Phe 340 345 350 Glu Glu Ser Ser Pro Val Thr Thr Ile Lys Pro Ile Gly Ala Val Ala 355 360 365 Ile Glu Ile Leu Ala Asn Asp Lys Gly Ile Arg Phe Asp Asn Ile Ile 370 375 380 Ile Gly Asn Asp Val Lys Glu Ala Ala Glu Phe Ile Asp Lys Glu Phe 385 390 395 400 Leu Ala Lys Gln Ala Asp Glu Lys Ala Lys Val Lys Glu Glu Ala Ala 405 410 415 Gln Ala Ala Gln Asn Ser Arg Trp Glu Glu Tyr Lys Lys Gly Ser Ile 420 425 430 Gln Gly Tyr Val Met Trp Tyr Ala Gly Asp Tyr Ile Asp Tyr Val Met 435 440 445 Glu Leu Tyr Glu Ala Ser Pro Ile Ala Val Gly Val Gly Ala Ala Ala 450 455 460 Ala Gly Leu Ala Val Leu Val Ala Leu Met Val Met Cys Met Ser Gly 465 470 475 480 Ala Pro Glu Glu Tyr Asp Asp Asp Val Ala Leu His Lys Lys Asp Asp 485 490 495 Asp Ala Ala Ala Gly Asp Asp Asp Glu Ala Glu Ala Glu Ala Glu Asn 500 505 510 Asp Ala Ala Asp Glu Asp Glu Asp Glu Glu Asp Asp Asp Asp Glu Glu 515 520 525 Asp Glu Asp Glu Glu Glu Asp Glu Asp Glu Ala Thr Gly Pro Arg Arg 530 535 540 Arg Val Asn Arg Ala Asn 545 550 <210> 90 <211> 6175 <212> DNA <213> Artificial <220> <223> pCL0121 <400> 90 ctctttatctg cctcgcgccg ttgaccgccg cttgactctt ggcgcttgcc gctcgcatcc 60 tgcctcgctc gcgcaggcgg gcgggcgagt gggtgggtcc gcagccttcc gcgctcgccc 120 gctagctcgc tcgcgccgtg ctgcagccag cagggcagca ccgcacggca ggcaggtccc 180 ggcgcggatc gatcgatcca tcgatccatc gatccatcga tcgtgcggtc aaaaagaaag 240 gaagaaa ggaaaaagaa aggcgtgcgc acccgagtgc gcgctgagcg cccgctcgcg 300 gtcccgcgga gcctccgcgt tagtccccgc cccgcgccgc gcagtcccccc gggaggcatc 360 gcgcacctct cgccgcccccc tcgcgcctcg ccgattcccc gcctcccctt ttccgcttct 420 tcgccgcctc cgctcgcggc cgcgtcgccc gcgccccgct ccctatctgc tccccagggg 480 ggcactccgc accttttgcg cccgctgccg ccgccgcggc cgccccgccg ccctggtttc 540 ccccgcgagc gcggccgcgt cgccgcgcaa agactcgccg cgtgccgccc cgagcaacgg 600 gtggcggcgg cgcggcggcg ggcggggcgc ggcggcgcgt aggcggggct aggcgccggc 660 tagcgaaac gccgccccg ggcgccgccg ccgcccgctc cagagcagtc gccgcgccag 720 accgccaacg cagagaccga gaccgaggta cgtcgcgccc gagcacgccg cgacgcgcgg 780 cagggacgag gagcacgacg ccgcgccgcg ccgcgcgggg ggggggaggg agaggcagga 840 cgcgggagcg agcgtgcatg tttccgcgcg agacgacgcc gcgcgcgctg gagaggagat aaggcgcttg gatcgcgaga gggccagcca ggctggaggc gaaaatgggt ggagaggata 960 gtatcttgcg tgcttggacg aggagactga cgaggaggac ggatacgtcg atgatgatgt gcacagagaa gaagcagttc gaaagcgact actagcaagc aagggatcca tgaagttcgc gacctcggtc gcaattttgc ttgtggccaa catagccacc gccctcgcgc agagcgatgg 1140 ctgcaccccc accgaccaga cgatggtgag caagggcgag gagctgttca ccggggtggt gcccatcctg gtcgagctgg acggcgacgt aaacggccac aagttcagcg tgtccggcga 1260 gggcgagggc gatgccacct acggcaagct gaccctgag ttcatctgca ccaccggcaa gctgcccgtg ccctggccca ccctcgtgac caccctgacc tacggcgtgc agtgcttcag 1380 ccgctacccc gaccacatga agcagcacga cttcttcaag tccgccatgc ccgaaggcta 1440 cgtccaggag cgcaccatct tcttcaagga cgacggcaac tacaagaccc gcgccgaggt 1500 gaagttcgag ggcgacaccc tggtgaaccg catcgagctg aagggcatcg acttcaagga 1560 ggacggcaac atcctgggac acaagctgga gtacaactac aacagccaca acgtctatat 1620 catggccgac aagcagaaga acggcatcaa ggtgaacttc aagatccgcc acaacatcga 1680 ggacggcagc gtgcagctcg ccgaccacta ccagcagaac acccccatcg gcgacggccc 1740 cgtgctgctg cccgacaacc actacctgag cacccagtcc gccctgagca aagaccccaa 1800 cgagaagcgc gatcacatgg tcctgctgga gttcgtgacc gccgccggga tcactctcgg 1860 catggacgag ctgtacaagc accaccatca ccaccactaa catatgagtt atgagatccg 1920 aaagtgaacc ttgtcctaac ccgacagcga atggcgggag ggggcgggct aaaagatcgt 1980 attacatagt atttttcccc tactctttgt gtttgtcttt tttttttttt tgaacgcatt 2040 caagccactt gtctgggttt acttgtttgt ttgcttgctt gcttgcttgc ttgcctgctt 2100 cttggtcaga cggcccaaaa aagggaaaaa attcattcat ggcacagata agaaaaagaa 2160 aaagtttgtc gaccaccgtc atcagaaagc aagagaagag aaacactcgc gctcacattc 2220 tcgctcgcgt aagaatctta gccacgcata cgaagtaatt tgtccatctg gcgaatcttt 2280 acatgagcgt tttcaagctg gagcgtgaga tcataccttt cttgatcgta atgttccaac 2340 cttgcatagg cctcgttgcg atccgctagc aatgcgtcgt actcccgttg caactgcgcc 2400 atcgcctcat tgtgacgtga gttcagattc ttctcgagac cttcgagcgc tgctaatttc 2460 gcctgacgct ccttcttttg tgcttccatg acacgccgct tcaccgtgcg ttccacttct 2520 tcctcagaca tgcccttggc tgcctcgacc tgctcggtaa aacgggcccc agcacgtgct 2580 acgagatttc gattccaccg ccgccttcta tgaaaggttg ggcttcggaa tcgttttccg 2640 ggacgccggc tggatgatcc tccagcgcgg ggatctcatg ctggagttct tcgcccaccc 2700 caacttgttt attgcagctt ataatggtta caataaagc aatagcatca caaatttcac 2760 aaataaagca tttttttcac tgcattctag ttgtggtttg tccaaactca tcaatgtatc 2820 ttatcataca tggtcgacct gcaggaacct gcattaatga atcggccaac gcgcggggag 2880 aggcggtttg cgtattgggc gctcttccgc ttcctcgctc actgactcgc tgcgctcggt 2940 cgttcggctg cggcgagcgg tatcagctca ctcaaaggcg gtaatacggt tatccacaga 3000 atcaggggat aacgcaggaa agaacatgtg agcaaaaggc cagcaaaagg ccaggaaccg 3060 taaaaaggcc gcgttgctgg cgtttttcca taggctccgc ccccctgacg agcatcacaa 3120 aaatcgacgc tcaagtcaga ggtggcgaaa cccgacagga ctataaagat accaggcgtt 3180 tccccctgga agctccctcg tgcgctctcc tgttccgacc ctgccgctta ccggatacct 3240 gtccgccttt ctcccttcgg gaagcgtggc gctttctcat agctcacgct gtaggtatct 3300 cagttcggtg taggtcgttc gctccaagct gggctgtgtg cacgaacccc ccgttcagcc 3360 cgaccgctgc gccttatccg gtaactatcg tcttgagtcc aacccggtaa gacacgactt 3420 atcgccactg gcagcagcca ctggtaacag gattagcaga gcgaggtatg taggcggtgc 3480 tacagagttc ttgaagtggt ggcctaacta cggctacact agaagaacag tatttggtat 3540 ctgcgctctg ctgaagccag ttaccttcgg aaaaagagt ggtagctctt gatccggcaa 3600 acaaaccacc gctggtagcg gtggtttttt tgtttgcaag cagcagatta cgcgcagaaa 3660 aaaaggatct caagaagatc ctttgatctt ttctacgggg tctgacgctc agtggaacga 3720 aaactcacgt taagggattt tggtcatgag attatcaaaa aggatcttca cctagatcct 3780 tttaaattaa aaatgaagtt ttaaatcaat ctaaagtata tatgagtaaa cttggtctga 3840 cagttaccaa tgcttaatca gtgaggcacc tatctcagcg atctgtctat ttcgttcatc 3900 catagttgcc tgactccccg tcgtgtagat aactacgata cgggagggct taccatctgg 3960 ccccagtgct gcaatgatac cgcgagaccc acgctcaccg gctccagatt tatcagcaat 4020 aaaccagcca gccggaaggg ccgagcgcag aagtggtcct gcaactttat ccgcctccat 4080 ccagtctatt aattgttgcc gggaagctag agtaagtagt tcgccagtta atagtttgcg 4140 caacgttgtt gccattgcta caggcatcgt ggtgtcacgc tcgtcgtttg gtatggcttc 4200 attcagctcc gttcccaac gatcaggcg agttacatga tcccatgt tgtgcaaaaa 4260 agcggttagc tccttcggtc ctccgatcgt tgtcagaagt aagttggccg cagtgttatc 4320 actcatggtt atggcagcac tgcataattc tcttactgtc atgccatccg tagatgctt 4380 ttctgtgact ggtgagtact caccaagtc attctgagaa tagtgtatgc ggcgaccgag 4440 ttgctcttgc ccggcgtca tacggtaa taccgcgcca catagcagaa ctttaaaagt 4500 gctcatcatt ggaaaacgtt cttcggggcg aaaacctca aggatcttac cgctgttgag 4560 atccagttcg atgtaaccca ctcgtgcacc caacgatct tcagcatctt ttactttcac 4620 cagcgttttct gggtgagcaaaacaggaag gcaaatgcc gcaaaaagg gataaggggc 4680 gabacggaaa tgttgaatac tcatactctt cctttttcaa tattattgaa gcatttatca 4740 gggttattgt ctcatgagcg gatacatatt tgaatgtatt tagaaaata aaaatagg 4800 ggttccgcgc acatttcccc gaaagtgcc acctgacgtc taagaaacca ttattatcat 4860 gataaacc tataaaaata ggcgtatcac gaggccctt cgtctcgcgc gtttcggtga 4920 tgacggtgaa aacctctgac acatgcagct cccggagacg gtcacagctt gtctgtaagc 4980 ggatgccggg agcagacaag cccgtcaggg cgcgtcagcg ggtgttggcg ggtgtcgggg 5040 ctggcttaac tatgcggcat cagagcagat tgtactgaga gtgcaccaag cttccaattt 5100 taggcccccc actgaccgag gtctgtcgat aatccacttt tccattgatt ttccaggttt 5160 cgttaactca tgccactgag caaaacttcg gtctttccta acaaaagctc tcctcacaaa 5220 gcatggcgcg gcaacggacg tgtcctcata ctccactgcc acacaaggtc gataaactaa 5280 gctcctcaca aatagaggag aattccactg acaactgaaa acaatgtatg agagacgatc 5340 accactggag cggcgcggcg gttgggcgcg gaggtcggca gcaaaaacaa gcgactcgcc 5400 gagcaaaccc gaatcagcct tcagacggtc gtgcctaaca acacgccgtt ctaccccgcc 5460 ttcttcgcgc cccttcgcgt ccaagcatcc ttcaagttta tctctctagt tcaacttcaa 5520 gaagaacaac accaccaaca ccatggccaa gttgaccagt gccgttccgg tgctcaccgc 5580 gcgcgacgtc gccggagcgg tcgagttctg gaccgaccgg ctcgggttct cccgggactt 5640 cgtggaggac gacttcgccg gtgtggtccg ggacgacgtg accctgttca tcagcgcggt 5700 ccaggaccag gtggtgccgg acaacaccct ggcctgggtg tgggtgcgcg gcctggacga 5760 gctgtacgcc gagtggtcgg aggtcgtgtc cacgaacttc cgggacgcct ccgggccggc 5820 catgaccgag atcggcgagc agccgtgggg gcgggagttc gccctgcgcg acccggccgg 5880 caactgcgtg cacttcgtgg ccgaggagca ggactgacac gtgctacgag atttcgattc 5940 caccgccgcc ttctatgaaa ggttgggctt cggaatcgtt ttccgggacg ccggctggat 6000 gatcctccag cgcggggatc tcatgctgga gttcttcgcc caccccaact tgtttattgc 6060 agcttataat ggttacaaat aaagcaatag catcacaaat ttcacaaata aagcattttt 6120 ttcactgcat tctagttgtg gtttgtccaa actcatcaat gtatcttatc ggtac 6175 <210> 91 <211> 6611 <212> DNA <213> Artificial <220> <223> pCL0122 <400> 91 ctcttatctg cctcgcgccg ttgaccgccg cttgactctt ggcgcttgcc gctcgcatcc 60 tgcctcgctc gcgcaggcgg gcgggcgagt gggtgggtcc gcagccttcc gcgctcgccc 120 gctagctcgc tcgcgccgtg ctgcagccag cagggcagca ccgcacggca ggcaggtccc 180 ggcgcggatc gatcgatcca tcgatccatc gatccatcga tcgtgcggtc aaaaagaaag 240 gaagaaa ggaaaaagaa aggcgtgcgc acccgagtgc gcgctgagcg cccgctcgcg 300 gtcccgcgga gcctccgcgt tagtccccgc cccgcgccgc gcagtcccccc gggaggcatc 360 gcgcacctct cgccgcccccc tcgcgcctcg ccgattcccc gcctcccctt ttccgcttct 420 tcgccgcctc cgctcgcggc cgcgtcgccc gcgccccgct ccctatctgc tccccagggg 480 ggcactccgc accttttgcg cccgctgccg ccgccgcggc cgccccgccg ccctggtttc 540 ccccgcgagc gcggccgcgt cgccgcgcaa agactcgccg cgtgccgccc cgagcaacgg 600 gtggcggcgg cgcggcggcg ggcggggcgc ggcggcgcgt aggcggggct aggcgccggc 660 taggcgaaac gccgccccccg gggcgccgccg ccgcccgctc cagagcagtc gccgcgccag 720 accgccaacg cagagaccga gaccgaggta cgtcgcgccc gagcacgccg cgacgcgcgg 780 cagggacgag gagcacgacg ccgcgccgcg ccgcgcgggg ggggggaggg agaggcagga 840 cgcgggagcg agcgtgcatg tttccgcgcg agacgacgcc gcgcgcgctg gagaggagat aaggcgcttg gatcgcgaga gggccagcca ggctggaggc gaaaatgggt ggagaggata 960 gtatcttgcg tgcttggacg aggagactga cgaggaggac ggatacgtcg atgatgatgt gcacagagaa gaagcagttc gaaagcgact actagcaagc aagggatcca tgaagttcgc gacctcggtc gcaattttgc ttgtggccaa catagccacc gccctcgcgc agagcgatgg 1140 ctgcaccccc accgaccaga cgatggtgag caagggcgag gagctgttca ccggggtggt gcccatcctg gtcgagctgg acggcgacgt aaacggccac aagttcagcg tgtccggcga 1260 gggcgagggc gatgccacct acggcaagct gaccctgag ttcatctgca ccaccggcaa gctgcccgtg ccctggccca ccctcgtgac caccctgacc tacggcgtgc agtgcttcag 1380 ccgctacccc gaccacatga cttcttcaag tccgccatgc ccgaaggcta cgtccaggag cgcaccatct tcttcaagga cgacggcaac tacaagaccc gcgccgaggt gaagttcgag ggcgacaccc tggtgaaccg catcgagctg aagggcatcg acttcaagga 1560 ggacggcaac atcctgggac acaagctgga gtacaactac aacagccaca acgtctatat 1620 catggccgac aagcagaaga acggcatcaa ggtgaacttc aagatccgcc acaacatcga 1680 ggacggcagc gtgcagctcg ccgaccacta ccagcagaac acccccatcg gcgacggccc 1740 cgtgctgctg cccgacaacc actacctgag cacccagtcc gccctgagca aagaccccaa 1800 cgagaagcgc gatcacatgg tcctgctgga gttcgtgacc gccgccggga tcactctcgg 1860 catggacgag ctgtacaagc accaccatca ccaccactaa catatgagtt atgagatccg 1920 aaagtgaacc ttgtcctaac ccgacagcga atggcgggag ggggcggct aaaagatcgt 1980 attacatagt atttttcccc tactctttgt gtttgtcttt tttttttttt tgaacgcatt 2040 caagccactt gtctgggttt acttgtttgt ttgcttgctt gcttgcttgc ttgcctgctt 2100 cttggtcaga cggcccaaaa aagggaaaaaa attcattcat ggcacagata agaaaaagaa 2160 aaagtttgtc gaccaccgtc atcagaaagc aagagaagag aaacactcgc gctcacattc 2220 tcgctcgcgt aagaatctta gccacgcata cgaagtaatt tgtccatctg gcgaatcttt 2280 acatgagcgt tttcaagctg gagcgtgaga tcataccttt cttgatcgta atgttccaac 2340 cttgcatagg cctcgttgcg atccgctagc aatgcgtcgt actcccgttg caactgcgcc 2400 atcgcctcat tgtgacgtga gttcagattc ttctcgagac cttcgagcgc tgctaatttc 2460 gcctgacgct ccttcttttg tgcttccatg acacgccgct tcaccgtgcg ttccacttct 2520 tcctcagaca tgcccttggc tgcctcgacc tgctcggtaa aacgggcccc agcacgtgct 2580 acgagatttc gattccaccg ccgccttcta tgaaaggttg ggcttcggaa tcgttttccg 2640 ggacgccggc tggatgatcc tccagcgcgg ggatctcatg ctggagttct tcgcccaccc 2700 caacttgttt attgcagctt ataatggtta caataaagc aatagcatca caaatttcac 2760 aaataaagca tttttttcac tgcattctag ttgtggtttg tccaaactca tcaatgtatc 2820 ttatcataca tggtcgacct gcaggaacct gcattaatga atcggccaac gcgcggggag 2880 aggcggtttg cgtattgggc gctcttccgc ttcctcgctc actgactcgc tgcgctcggt 2940 cgtcggctg cggcgagcgg tatcagctca ctcaaaggcg gtaatacggt tatccacaga 3000 atcaggggat aacgcaggaa agaacatgtg agcaaaaggc cagcaaaagg ccaggaaccg 3060 taaaaaggcc gcgttgctgg cgtttttcca taggctccgc ccccctgacg agcatcacaa 3120 aaatcgacgc tcaagtcaga ggtggcgaaa cccgacagga ctataaagat accaggcgtt 3180 tccccctgga agctccctcg tgcgctctcc tgttccgacc ctgccgctta cggatacct 3240 gtccgccttt ctcccttcgg gaagcgtggc gctttctcat agctcacgct gtaggtatct 3300 cagttcggtg tagtcgttc gctccaagct gggctgtgtg cacgaacccc ccgttcagcc 3360 cgaccgctgc gccttatccg gtaactatcg tcttgagtcc aacccggtaa gacacgactt 3420 atcgccactg gcagcagcca ctggtaacag gattagcaga gcgaggtatg taggcggtgc 3480 tacagagttc ttgaagtggt ggcctaacta cggctacact agaagaacag tatttggtat 3540 ctgcgctctg ctgaagccag ttaccttcgg aaaaagagt ggtagctctt gatccggcaa 3600 acaaaccacc gctggtagcg gtggtttttt tgtttgcaag cagcagatta cgcgcagaaa 3660 aaaaggatct caagaagatc ctttgatctt ttctacgggg tctgacgctc agtggaacga 3720 aaactcacgt taagggattt tggtcatgag attatcaaaa aggatcttca cctagatcct 3780 tttaaattaa aaatgaagtt ttaaatcaat ctaaagtata tatgagtaaa cttggtctga 3840 cagttaccaa tgcttaatca gtgaggcacc tatctcagcg atctgtctat ttcgttcatc 3900 catagttgcc tgactccccg tcgtgtagat aactacgata cgggaggct taccatctgg 3960 ccccagtgct gcaatgatac cgcgagaccc acgctcaccg gctccagatt tatcagcaat 4020 aaaccagcca gccggaaggg ccgagcgcag aagtggtcct gcaactttat ccgcctccat 4080 ccagtctatt aattgttgcc gggaagctag agtaagtagt tcgccagtta atagtttgcg 4140 caacgttgtt gccattgcta caggcatcgt ggtgtcacgc tcgtcgtttg gtatggcttc 4200 attcagctcc ggttcccaac gatcaaggcg agttacatga tcccccatgt tgtgcaaaaa 4260 agcggttagc tccttcggtc ctccgatcgt tgtcagaagt aagttggccg cagtgttatc 4320 actcatggtt atggcagcac tgcataattc tcttactgtc atgccatccg taagatgctt 4380 ttctgtgact ggtgagtact caaccaagtc attctgagaa tagtgtatgc ggcgaccgag 4440 ttgctcttgc ccggcgtcaa tacgggataa taccgcgcca catagcagaa ctttaaaagt 4500 gctcatcatt ggaaaacgtt cttcggggcg aaaactctca aggatcttac cgctgttgag 4560 atccagttcg atgtaaccca ctcgtgcacc caactgatct tcagcatctt ttactttcac 4620 cagcgtttct gggtgagcaa aaacaggaag gcaaaatgcc gcaaaaaagg gaataagggc 4680 gacacggaaa tgttgaatac tcatactctt cctttttcaa tattattgaa gcatttatca 4740 gggttattgt ctcatgagcg gatacatatt tgaatgtatt tagaaaaata aaaatagg 4800 ggttccgcgc acatttcccc gaaaagtgcc acctgacgtc taagaaacca ttattatcat 4860 gacattaacc tataaaaata ggcgtatcac gaggcccttt cgtctcgcgc gtttcggtga 4920 tgacggtgaa aacctctgac acatgcagct cccggagacg gtcacagctt gtctgtaagc 4980 ggatgccggg agcagacaag cccgtcaggg cgcgtcagcg ggtgttggcg ggtgtcgggg 5040 ctggcttaac tatgcggcat cagagcagat tgtactgaga gtgcaccaag cttccaattt 5100 taggcccccc actgaccgag gtctgtcgat aatccacttt tccattgatt ttccaggttt 5160 cgttaactca tgccactgag caaaacttcg gtctttccta acaaaagctc tcctcacaaa 5220 gcatggcgcg gcaacggacg tgtcctcata ctccactgcc acacaaggtc gataaactaa 5280 gctcctcaca aatagaggag aattccactg acaactgaaa acaatgtatg agagacgatc 5340 accactggag cggcgcggcg gttgggcgcg gaggtcggca gcaaaaacaa gcgactcgcc 5400 gagcaaaccc gaatcagcct tcagacggtc gtgcctaaca acacgccgtt ctaccccgcc 5460 ttcttcgcgc cccttcgcgt ccaagcatcc ttcaagttta tctctctagt tcaacttcaa 5520 gaagaacaac accaccaaca ccatgattga acaagatgga ttgcacgcag gttctccggc 5580 cgcttgggtg gagaggctat tcggctatga ctgggcacaa cagacaatcg gctgctctga 5640 tgccgccgtg ttccggctgt cagcgcaggg gcgcccggtt ctttttgtca agaccgacct 5700 gtccggtgcc ctgaatgaac tgcaggacga ggcagcgcgg ctatcgtggc tggccacgac 5760 gggcgttcct tgcgcagctg tgctcgacgt tgtcactgaa gcgggaaggg actggctgct 5820 attgggcgaa gtgccggggc aggatctcct gtcatctcac cttgctcctg ccgagaaagt 5880 atccatcatg gctgatgcaa tgcggcggct gcatacgctt gatccggcta cctgcccatt 5940 cgaccaccaa gcgaaacatc gcatcgagcg agcacgtact cggatggaag ccggtcttgt 6000 cgatcaggat gatctggacg aagagcatca ggggctcgcg ccagccgaac tgttcgccag 6060 gctcaaggcg cgcatgcccg acggcgatga tctcgtcgtg acccatggcg atgcctgctt 6120 gccgaatatc atggtggaaa atggccgctt ttctggattc atcgactgtg gccggctggg 6180 tgtggcggac cgctatcagg acatagcgtt ggctacccgt gatattgctg aagagcttgg 6240 cggcgaatgg gctgaccgct tcctcgtgct ttacggtatc gccgctcccg attcgcagcg 6300 catcgccttc tatcgccttc ttgacgagtt cttctgacac gtgctacgag atttcgattc 6360 caccgccgcc ttctatgaaa ggttgggctt cggaatcgtt ttccgggacg ccggctggat 6420 gatcctccag cgcggggatc tcatgctgga gttcttcgcc caccccaact tgtttattgc 6480 agcttataat ggttacaaat aaagcaatag catcacaaat ttcacaaata aagcattttt 6540 ttcactgcat tctagttgtg gtttgtccaa actcatcaat gtatcttatc atgtctgaat 6600 tcccggggta c 6611 <210> 92 <211> 1314 <212> DNA <213> Artificial <220> <223> Codon optimized Isomerase <400> 92 atggctaagg agtacttccc ccagatccag aagattaagt tcgagggtaa ggacagcaag 60 aacccgctcg cctttcatta ctacgacgcc gagaaggagg tgatgggcaa gaagatgaag 120 gactggcttc gctttgctat ggcttggtgg cacactctct gcgctgaggg cgcggaccag 180 tttggcggcg gtacgaagag ctttccgtgg aacgagggca ctgacgctat tgagattgct 240 aagcagaagg ttgacgctgg tttcgagatt atgcagaagc tcggtattcc gtactactgc 300 tttcacgatg tcgacctcgt ttccgagggc aactcgatcg aggagtacga gtcgaacctc 360 aaggctgtgg ttgcctacct caaggagaag cagaaggaga ccggaatcaa gctcctctgg 420 agcaccgcca acgttttcgg ccacaagcgc tacatgaacg gcgcctccac caaccctgac 480 ttcgatgttg ttgcccgcgc tattgtccag attaagaacg ccatcgacgc tggtatcgag 540 ctcggagccg agaactacgt tttttggggc ggacgcgagg gttacatgtc cctcctcaac 600 accgaccaga agcgtgagaa ggagcacatg gccactatgc ttaccatggc ccgcgactac 660 gcccgcagca agggttttaa gggtactttt ctcattgagc cgaagcccat ggagccgacc 720 aagcaccagt acgacgtcga caccgagacc gccattggct tccttaaggc ccacaacctt 780 gacaaggatt ttaaggtgaa catcgaggtt aaccacgcta cgcttgccgg ccacaccttt 840 gagcatgagc tcgcctgcgc tgttgacgcc ggaatgcttg gttccattga cgccaaccgc 900 ggcgactacc agaacggctg ggacaccgac cagtttccga ttgaccagta cgagctcgtc 960 caggcctgga tggagatcat ccgtggtgga ggctttgtta ccggtggtac gaacttcgac 1020 gccaagacgc gccgtaacag cacggacctc gaggacatca tcattgctca tgtgtcgggc 1080 atggacgcca tggctcgcgc ccttgagaac gctgctaagc tcctccagga gagcccctac 1140 acgaagatga agaaggagcg ctacgcgtcg tttgacagcg gaatcggtaa ggacttcgag 1200 gatggcaagc tcaccctgga gcaggtgtac gagtacggta agaagaacgg cgagccgaag cagaccagcg gcaagcagga gctctacgag gccattgtcg ccatgtacca gtag <210> 93 <211> 1485 <212> DNA <213> Artificial <220> <223> Codon optimized Kinase <400> 93 atgaagaccg tcgccggcat cgatcttgga acccagtcca tgaaggttgt catttacgac tacgagaaga aggagatcat cgagtccgcc tcgtgcccta tggagctcat tagcgagtcg 120 180. gacggaaccc gcgagcagac gactgagtgg tttgacaagg gtctcgaggt gtgctttgga aagctctccg ctgataacaa gaagaccatt gaggcgattg gcatctccgg ccagctccac ggcttcgtcc ctctcgatgc gaacggaag gcgctctaca acatcaagct ctggtgcgac accgccactg tggaggagtg caagatcatt actgacgccg ccggcggcga caaggctgtc atcgacgcgc tcggcaacct catgctcacc ggattcaccg ccccgaagat tctctggctc 420 aagcgcaaca agcccgaggc ctttgctac ctcaagtaca ttatgctgcc ccacgattac ctcaactgga agctgactgg agactacgtc atggagtacg gcgacgcctc cggcaccgcc 540 ctttttgatt cgaagaaccg ctgctggtcg aagaagattt gcgacattat tgatcctaag 600 ctgctcgacc ttctccctaa gctcattgag ccctcggccc ccgccggtaa ggtcaacgac 660 gaggccgcca aggcgtacgg cattcccgcc ggaatccccg tttccgctgg cggcggtgat 720 aacatgatgg gtgcggtcgg tactggcacc gtcgctgacg gattcctcac gatgagcatg 780 ggcacctccg gaactcttta cggctactcg gacaagccta tttccgaccc ggctaacggc 840 ctcagcggct tctgcagctc cacgggcggc tggcttcccc tcctttgcac catgaactgc 900 accgtcgcca ccgagttcgt ccgcaacctt tttcagatgg atatcaagga gctgaacgtc 960 gaggctgcta agtccccctg cggcagcgag ggcgttcttg tcattccttt cttcaacggc 1020 gagcgcaccc cgaacctccc caacggccgc gcctcgatta ccggcctcac ctccgcgaac 1080 acgtcccgcg ccaacatcgc tcgcgcctcc tttgagtcgg ccgtctttgc catgcgcggt 1140 ggcctcgatg cgtttcgtaa gctcggattc cagcccaagg agattcgcct catcggcggt 1200 ggttcgaagt ccgacctctg gcgccagatc gctgctgaca ttatgaacct tcccatccgt 1260 gtcccccttc tcgaggaggc cgccgccctc ggcggagctg tccaggccct ttggtgcctt 1320 aagaaccagt ccggtaagtg cgacatcgtc gagctttgca aggagcatat caagattgac 1380 gagtccaaga acgccaaccc gattgccgag aacgtcgccg tgtacgataa ggcctacgat 1440 gagtactgca aggtcgttaa cacgctcagc cctctgtacg cctaa 1485 <210> 94 <211> 1569 <212> DNA <213> Artificial <220> <223> Codon optimized transporter <400> 94 atgggcctcg aggataaccg catggttaag cgctttgtca acgtgggcga gaagaaggcc 60 ggtagcaccg ccatggccat cattgttggc ctcttcgcgg cctcgggcgg cgtcctcttc 120 ggctacgaca ccggcactat ctcgggcgtc atgactatgg actacgttct cgcccgctac 180 ccctccaaca agcactcctt caccgctgac gagtcgtcgc tcatcgtttc cattctttcg 240 gtcggcacct tcttcggcgc cctctgcgcc ccgttcctca acgataccct cggccgccgc 300 tggtgcctca tcctcagcgc cctcattgtc tttaacatcg gcgccatcct ccaggtcatt 360 tccaccgcca tccccctgct ctgcgcgggc cgcgttatcg ccggtttcgg tgtcggcctc 420 atttccgcca ccatcccgct ctaccagtcc gagactgctc cgaagtggat tcgcggcgcc 480 atcgtttcct gctaccagtg ggccatcact atcggacttt tcctcgcttc ctgcgtcaac 540 aagggcaccg agcacatgac caactccggt tcgtaccgta ttcctctggc catccagtgc 600 ctctggggcc tcatccttgg tattggcatg attttcctcc ctgagacccc ccgcttctgg 660 atttcgaagg gcaaccagga gaaggccgcc gagtccctcg cccgtctccg caagctcccc 720 atcgaccatc ctgatagcct tgaggagctt cgcgatatta ctgccgccta cgagttcgag 780 accgtctacg gtaagtccag ctggtcccag gtcttttccc acaagaacca tcagctcaag 840 cgcctcttta ccggcgttgc cattcaggcc tttcagcagc tcaccggagt taactttatc 900 ttttactacg gcaccacctt ttttaagcgc gccggagtca acggattcac catcagcctt 960 gccaccaaca tcgttaacgt cggcagcact attcccggca ttcttctcat ggaggtcctc 1020 ggccgccgca acatgctcat gggcggtgcc accggcatgt cgctgtcgca gcttatcgtc 1080 gccattgtcg gagttgccac gtcggagaac aacaagtcga gccagtcggt cctcgtcgct 1140 ttctcgtgca ttttatcgc ttttttgcc gccacctggg gtccctgcgc ctgggtcgtc 1200 gtcggcgagc tctttcccct tcgcactcgc gctaagtcg ttccctctg caccgcgtcc 1260 aactggctct ggaactgggg cattgcttac gccaccccct acatggtcga cgaggataag 1320 ggtaacctcg gcagcaacgt ttttttatt tggggaggct tcaacctcgc ttgcgtcttt 1380 ttcgcgtggt acttcattta cgagaccaag ggcctttccc tcgagcaggt tgatgagctc 1440 tacgagcatg ttcgaaggc gtggaagtcc aagggttttg tcccgtccaa gcactccttt 1500 cgcgagcagg tcgaccagca gatggactcc aagaccgagg ccattatgag cgaggaggcg 1560 tcggtttaa 1569 <210> 95 <211> 1512 <212> DNA <213> Artificial <220> <223> Codon optimized transporter <400> 95 atggccctcg accctgagca gcagcagccc atttcctccg tgtcgcgcga gtttggtaag 60 tcgtccggtg agatctcccc cgagcgtgag cctctcatta aggagaacca cgtccccgag 120 aactactccg ttgttgccgc catcctcccc ttcctcttcc cggccctggg tggcctcctt 180 tacggttacg agattggcgc tacgtcgtgc gctacgattt cccttcagtc cccctccctc 240 tccggcatct cctggtacaa cctctcctcc gtcgatgttg gcctcgtcac ttccggttcc 300 ctctacggtg ctctgtttgg ctccattgtt gccttcacca ttgccgacgt tattggccgt 360 cgcaaggagc ttatcctcgc tgctctcctc tacctcgtcg gtgccctcgt taccgctctc 420 gcccctacgt actccgttct catcatcggc cgtgtcattt acggtgtttc cgtcggtctt 480 gccatgcatg ctgcccctat gtacatcgcg gagaccgccc cgtcccccat ccgcggccag 540 ctcgtttccc tcaaggagtt tttcatcgtt ctcggtatgg tcggcggata cggcattggt 600 tccctcaccg tcaacgtcca ctccggttgg cgctacatgt acgctacctc cgttcccctc 660 gctgtgatca tgggcattgg catgtggtgg cttcctgcct ccccccgttg gctcctcctc 720 cgcgtcattc agggtaaggg taacgttgag aaccagcgcg aggctgccat taagtccctc 780 tgctgcctcc gtggtcctgc cttcgtcgac tcggccgccg agcaggtcaa cgagattctc 840 gccgagctta ccttcgttgg cgaggataag gaggtcacct tcggcgagct cttccaggga 900 aagtgcctca aggccctcat tatcggcggc ggccttgttc tctttcagca gatcaccggt 960 cagccttcgg tcctctacta cgccccctcg atcctccaga ctgcgggctt ctccgccgcc 1020 ggcgatgcta cccgcgtttc cattcttctc ggcctcctca agctcattat gaccggtgtc 1080 gccgtcgtcg ttatcgatcg tctcggccgt cgccctctcc tcctcggcgg agtcggtggt 1140 atggttgttt cgctctttct ccttggctcg tactaccttt tcttcagcgc ttcccccgtc 1200 gtcgccgttg tcgccctcct tctctacgtg ggttgctacc agctctcctt tggccccatt 1260 ggctggctta tgatttccga gatttttccc ctcaagctcc gtggtcgcgg actctccctt 1320 gccgtgcttg tcaactttgg tgccaacgcc ctcgtcacct ttgccttttc ccctctcaag 1380 gagctcctcg gcgccggcat cctgttttgc ggctttggcg ttatctgcgt tctctccctt 1440 gtttttatct tttttatcgt cccggagact aagggcctca cgctcgagga gatcgaggcg 1500 aagtgcctct aa 1512 <210> 96 <211> 19 <212> DNA <213> Artificial <220> <223> Primer 5' CL0130 <400> 96 cctcgggcgg cgtcctctt 19 <210> 97 <211> 20 <212> DNA <213> Artificial <220> <223> Primer 3' CL0130 <400> 97 ggcggccttc tcctggttgc 20 <210> 98 <211> 24 <212> DNA <213> Artificial <220> <223> Primer 5' CL0131 <400> 98 ctactccgtt gttgccgcca tcct 24 <210> 99 <211> 22 <212> DNA <213> Artificial <220> <223> Primer 3' CL0131 <400> 99 ccgccgacca taccgagaac ga 22 <210> 100 <211> 1362 <212> DNA <213> Artificial <220> <223> Codon Optimized NA <400> 100 atgaacccca accagaagat tactactatc ggtagcattt gcctcgtcgt tggacttatc tcccttattc ttcagattgg taacattatc tccatttgga tctcgcatag cattcagacc ggctcccaga accacaccgg catttgcaac cagaacatta ttacttacaa gaactccact tgggtcaagg acactactag cgttattctt accggtaact cgtcgctttg ccctattcgc 240 ggctgggcta tttacagcaa ggacaactcg atccgcatcg gtagcaaggg cgacgttttt gtcatccgtg agccttttat ttcctgcagc cacctcgagt gccgtacttt ttttctgact 360 cagggcgctc tcctcaacga tagcattcc aacggcactg tcaaggatcg cagcccctac 420 cgcgccctta tgtcctgccc tgtcggcgag gctcccagcc cctacaactc ccgttttgag 480 tccgttgcct ggtccgccag cgcctgccac gacggatgg gatggctcac tattggtatt 540 tccggccctg fathercggcgc tgtcgccgtc cttagtaca acggcattat caccgagacc 600 atcaagtcct ggcgtaagaa gatcctccgc acccaggagt ccgagtgcgc ctgcgtcaac 660 ggcagctgct tcacgattat gaccgacggc ccctccgacg gcctcgcttc ctacaagatt 720 tttaagattg agaagggtaa ggtcacgaag tccatcgagc ttaacgcccc gaactcccac tacgaggagt gctcctgcta ccctgacact ggcaaggtga tgtgcgtctg ccgcgataac 840 tggcatggct ccaaccgccc ctgggttagc ttcgatcaga accttgacta ccagattgga tacatttgct ccggtgtttt tggcgacaac ccgcgccccg aggatggac tggttcgtgc 960 ggtcctgttt acgttgacgg cgccaacggc gttaagggtt tttcctaccg ttacggtac 1020 ggagtctgga tcggccgcac caagtcgcac agctcgcgcc acggatttga gatgatctgg gaccccaacg gatggactga gaccgattcc aagtttagcg ttcgccagga tgtcgttgct atgaccgatt ggtcgggata ctccggttcc tttgtgcagc accctgagct caccggcctt gactgcatgc gcccttgctt ttgggtcgag ctcattcgcg gtcgccctaa ggagaagact 1260. atttggacct ccgccagcag catttccttt tgcggcgtta actccgacac cgtcgactgg 1320 tcgtggcccg atggcgccga gcttcccttt tccattgata ag 1362. <210> 101 <211> 1431 <212> DNA <213> Artificial <220> <223> Codon Optimized NA with V5 Tag and a Polyhistidine Tag <400> 101 at...

Claims

1. 1. A method for the production of a viral neuraminidase (NA) protein, comprising the steps of: providing a recombinant microalgae cell containing a nucleic acid molecule comprising a polynucleotide sequence encoding a viral NA protein comprising an NA membrane domain and an active regulatory control element; Cultivating the recombinant microalgae cells in a culture medium such that the NA protein is secreted into the culture medium; and Recovering the secreted recombinant viral NA protein from the culture medium.

2. The method of claim 1, wherein the recombinant viral NA protein is a full-length viral NA protein.

3. The method of claim 1, wherein the recombinant viral NA protein is an influenza NA protein.

4. The method of claim 1, wherein the recombinant viral NA protein is a full-length influenza NA protein.

5. The method of claim 1, wherein the recombinant viral NA protein is at least 90% identical to the amino acid sequence encoded by SEQ ID NO:

100.

6. The method of claim 1, wherein the recombinant viral NA protein has an amino acid sequence encoded by SEQ ID NO:

100.

7. 10. The method of claim 1, wherein the microalgae cells are members of the order Thraustochytriales.

8. 2. The method of claim 1, wherein the microalgae cells are Schizochytrium or Thraustochytrium.

9. 1. A method of making a vaccine composition, comprising: Providing a recombinant viral NA protein produced by the method of any one of claims 1 to 8; and Formulating the recombinant viral NA protein as a vaccine composition for administration by at least one route selected from intramuscular, intravenous, subcutaneous, intrapulmonary, intratracheal, transdermal, intraocular, intranasal, inhalation, intraluminal, intraductal, and intraparenchymal administration.

Citation Information

Patent Citations

  • Method for introducing a gene into Labyrinthulomycota

    US20060275904A1

  • Vector capable for transformation of labyrinthulomycota

    US20060286650A1

  • Product and process for transformation of Thraustochytriales microorganisms

    US7001772B2