Recombinant microalgae capable of producing peptides, polypeptides or proteins of collagen, elastin and their derivatives in the chloroplasts of microalgae and related methods
By introducing nucleic acid sequences encoding recombinant proteins and polypeptides into the chloroplast genome of microalgae and utilizing homologous recombination and endogenous disulfide bond formation mechanisms, the stability and purity problems of amino acid repeating unit proteins and polypeptides in microalgae were solved, achieving efficient expression and simplified purification.
Patent Information
- Application Number
- CN202180017364.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-28
- Filing Date
- 2021-02-26
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2041-02-26
AI Technical Summary
The expression of existing recombinant proteins and peptides in microalgae has stability and purity problems, especially the expression of collagen and elastin and their derivatives containing repeating amino acid units is difficult, resulting in low yield and poor activity.
By introducing nucleic acid sequences encoding recombinant proteins, polypeptides or peptides into the chloroplast genome of microalgae, the nucleic acid sequences are integrated into the chloroplast genome using homologous recombination technology, combined with the endogenous disulfide bond formation mechanism to improve stability and activity, and identified and purified through expression vectors and selection markers.
The invention realizes efficient accumulation and stable expression of recombinant proteins, polypeptides or peptides of amino acid repeating units in microalgae chloroplasts, improves solubility and activity, simplifies purification steps and reduces production costs.
Smart Images

Figure BDA0003818011110000081 
Figure BDA0003818011110000211 
Figure BDA0003818011110000221
Abstract
Description
Technical Field
[0001] The present invention relates to recombinant microalgae comprising a nucleic acid sequence encoding a recombinant protein, polypeptide, or peptide comprising repeating amino acid units, selected from the group consisting of collagen, elastin, and derivatives thereof, wherein the nucleic acid sequence is located in the chloroplast genome of the microalgae. The present invention also relates to a method for producing a recombinant protein, polypeptide, or peptide comprising repeating amino acid units, selected from the group consisting of collagen, elastin, and derivatives thereof, in the chloroplasts of microalgae, wherein the method comprises transforming the chloroplast genome of the microalgae with the nucleic acid sequence encoding the recombinant protein, polypeptide, or peptide. Background Art
[0002] In recent years, the demand for recombinant proteins has been growing due to their high-value applications in a wide range of industries such as personal care, cosmetics, healthcare, tissue engineering, biomaterials, agriculture, and the paper industry. Numerous commercial pharmaceutical proteins produced in various recombinant systems have been introduced, for example, insulin, human growth hormone, erythropoietin, and interferon.
[0003] Additionally, these various industrial applications require large quantities of proteins and peptides.
[0004] The most current industrial expression systems include E. coli, yeast (S. cerevisiae and Pichia pastoris) and mammalian cell lines. Emerging technologies are insect cell culture, plants and microalgae.
[0005] However, the expression of recombinant peptides and proteins is still limited because a lot of effort is required to obtain the desired peptides and proteins with native conformation, in large quantities and at high purity. Even current bacterial systems such as E. coli have limitations in expressing recombinant peptides / polypeptides / proteins. In fact, the formation of insoluble aggregates (or inclusion bodies) occurs due to the lack of mature mechanisms for post-translational modifications, such as disulfide bond formation or glycosylation. This leads to poor solubility of the protein of interest and / or lack of protein activity.
[0006] In recent years, interest in microalgae as an alternative platform for recombinant protein production has grown.
[0007] Recombinant algae offer several advantages over other recombinant protein production platforms. Microalgae are photosynthetic, single-celled microorganisms that require few nutrients for growth. They are capable of photoautotrophic, mixotrophic, or heterotrophic growth. The cost of protein production in algae is much lower than other production systems that rely on photoauxotrophic growth. Proteins purified from algae, like those purified from plants, should be free of toxins and viral agents that may be present in preparations derived from bacterial or mammalian cell cultures. In fact, some microalgae species have GRAS (generally recognized as safe) status granted by the FDA, such as the microalgae Chlorella vulgaris, Chlorella protothecoides S106, Dunaliella bardawil, Chlamydomonas reinhardtii, and the cyanobacterium Arthrospira plantesis.
[0008] As in transgenic plants, algae have been engineered to express recombinant genes from both the nuclear and chloroplast genomes.
[0009] Furthermore, the recombinant synthesis of peptides and polypeptides comprising repeating units of a specific amino acid sequence is difficult because DNA sequences encoding the peptide or polypeptide frequently undergo genetic recombination, leading to genetic instability and often resulting in the production of proteins that are smaller than the native protein.
[0010] Recombinant production of relatively small peptides can also be challenging because they can self-assemble or be subject to proteolytic degradation.
[0011] Moreover, algae represent a robust industrial chassis with competitive production costs compared to plant expression systems, and can be produced on an industrial scale under reproducible, sterile, and well-controlled production conditions in photobioreactors and fermenters, or in disposable wave bioreactors (wave bags). Furthermore, they can secrete recombinant proteins extracellularly and, therefore, in the culture medium, simplifying subsequent purification steps. Algae are not seasonal and do not use arable land.
[0012] The development of chloroplast transformation in algae to produce proteins of interest is newer than in plants and is in need of improvement. Indeed, recombinant protein yields typically range between 0.5% and 5% of the total soluble protein in Chlamydomonas reinhardtii chloroplasts, which remains low compared to established microbial platforms.
[0013] Additionally, some mammalian proteins are not easily expressed (Rasala et al., 2010).
[0014] The inventors of the present invention have surprisingly been able to produce recombinant proteins, polypeptides or peptides comprising several amino acid repeating units selected from the group consisting of collagen, elastin and their derivatives in the chloroplasts of microalgae.
[0015] In fact, homologous recombination between similar or identical sequences is very efficient in the chloroplast genome, and transgenes with repetitive sequences are very unstable.
[0016] By using the described methods, the formation of endogenous disulfide bonds necessary for the increase of protein, peptide and polypeptide stability and activity, as well as protein, peptide and polypeptide accumulation, is allowed. Summary of the Invention
[0017] Therefore, the present invention relates to recombinant microalgae comprising a nucleic acid sequence encoding a recombinant protein, polypeptide or peptide comprising amino acid repeating units, said protein, polypeptide or peptide being selected from collagen, elastin and their derivatives, and said nucleic acid sequence being located in the chloroplast genome of the microalgae.
[0018] The present invention also relates to the use of the recombinant algae for producing recombinant proteins, polypeptides or peptides comprising repeating amino acid units, wherein the proteins, polypeptides or peptides are selected from collagen, elastin and their derivatives.
[0019] The present invention further relates to a method for producing a recombinant protein, polypeptide or peptide comprising repeating amino acid units in the chloroplasts of microalgae, wherein the recombinant protein, polypeptide or peptide is selected from the group consisting of collagen, elastin and their derivatives, wherein the method comprises transforming the chloroplast genome of the microalgae with a nucleic acid sequence encoding the recombinant protein, polypeptide or peptide.
[0020] In particular, the method comprises:
[0021] (i) providing a nucleic acid sequence encoding the recombinant protein, polypeptide or peptide;
[0022] (ii) introducing the nucleic acid sequence according to (i) into an expression vector capable of expressing the nucleic acid sequence in a microalgae host cell; and
[0023] (iii) Transforming the chloroplast genome of a microalgae host cell with the expression vector.
[0024] In particular, the method further comprises:
[0025] (iv) identifying the transformed microalgae host cell;
[0026] (v) characterizing a microalgal host cell for producing a recombinant protein, polypeptide, or peptide expressed by the nucleic acid sequence;
[0027] (vi) extracting the recombinant protein, polypeptide or peptide; and optionally
[0028] (vii) Purification of recombinant proteins, polypeptides or peptides.
[0029] More particularly, the method according to the invention allows increasing the accumulation and / or stability and / or solubility and / or folding and / or activity of recombinant peptides, polypeptides or proteins comprising repeating amino acid units in the chloroplasts of microalgae.
[0030] In particular, the method further comprises a further step (viii) of step (vii) wherein the polypeptide is cleaved to allow release of the peptide unit.
[0031] The cleavage may be performed by any method known to those skilled in the art, such as using a suitable intracellular protease.
[0032] In one embodiment, the recombinant proteins, polypeptides or peptides obtained by the method according to the invention are chemically modified at their N- or C-terminus, for example by adding a palmitoyl group, a hydroxyl group, an alkoyl chain (i.e. an alkyl chain comprising a hydroxyl group) or a biotin group.
[0033] In the art and in the context of the present invention, "recombinant peptide / polypeptide / protein" means an exogenous peptide / polypeptide / protein expressed by a recombinant gene (or recombinant nucleic acid sequence), i.e., an exogenous gene (or exogenous nucleic acid sequence) from a different species (heterologous) or from the same species (homologous).
[0034] "Recombinant microalgae" means microalgae that comprises a nucleic acid sequence encoding a recombinant protein, polypeptide or peptide. In the context of the present invention, recombinant microalgae are transformed, as described in further detail below.
[0035] "Peptide", "polypeptide" and "protein" have the meanings commonly understood by those skilled in the art to which the present invention belongs. In particular, peptides, polypeptides and proteins are polymers of amino acids linked by peptide (amide) bonds.
[0036] More particularly, the protein according to the invention has a unique and stable three-dimensional structure and comprises more than 50 amino acids, such as proteins of 54, 60, 66, 72, 75, 78, 84, 90, 96, 100, 102, 108, 114, 120, 150, 180, 200, 300, 350 or more amino acids; the peptides according to the invention are short oligopeptides, for example peptides of 2 to 10 amino acids, such as peptides of 4, 5, 6, 7, 8, 9 or 10 amino acids; a polypeptide according to the classical meaning may comprise a polypeptide of 11 to 50 amino acids, such as a polypeptide of 11, 12, 15, 18, 20, 24, 25, 30, 35, 36, 40, 42, 45, 48 or 50 amino acids; but in the context of the present invention, a polypeptide is a repetition of n units of identical or different amino acid sequences, or a repetition of n units of identical or different peptides, n being 2 to 400, in particular 2 to 100.
[0037] In the context of the present invention, the recombinant protein, polypeptide or peptide according to the present invention is a recombinant protein, polypeptide or peptide selected from collagen, elastin and their derivatives comprising repeating units of amino acids.
[0038] Elastin is the major structural protein of the extracellular matrix. It is present in the connective tissues of all vertebrates and provides tissue elasticity. Elastin is first synthesized as a soluble monomeric precursor, tropoelastin, which is subsequently assembled into mature elastin, a stable polymeric structure.
[0039] This protein is well known in the art. The amino acid sequence of elastin and tropoelastin contains short, repetitive amino acid motifs and many hydrophobic residues. During the aging process, after exposure to UVB radiation, or in pathological processes, elastin is degraded into short peptides, called "elastin peptides," which act as signal peptides that promote, for example, cell proliferation. These elastin peptides are part of the matrix peptides (matricins peptide).
[0040] Elastin peptides are believed to be components found in natural elastin and have a short sequence of amino acids.
[0041] Examples of elastin peptides according to the invention are the pentapeptides: KGGVG (SEQ ID N° 1), VGGVG (SEQ ID N° 2), GVGVP (SEQ ID N° 3), VPGXG (X is V, I or K) (SEQ ID N° 4°); the hexapeptide: VGVAPG (SEQ ID N° 5); the heptapeptide: LGAGGAG (SEQ ID N6); or the nonapeptide: LGAGGAGVL (SEQ ID N° 7).
[0042] In the context of the present invention, recombinant peptides, polypeptides or proteins comprising repeating units of elastin include repeating units of elastin peptides (same or different), and in particular those of SEQ ID N° 1 to 7 (same or different). "Derivatives" of elastin proteins / polypeptides / peptides encompass elastin-like proteins / polypeptides / peptides and proteins / polypeptides / peptides in which the amino acid sequence of native elastin proteins / polypeptides / peptides is mutated or contains one or more amino acids at the N- or C-terminus. The supplementary amino acids may be any amino acids. Derivatives of elastin proteins / polypeptides / peptides also encompass derivatives of elastin-like proteins / polypeptides / peptides. "Derivatives" of elastin-like proteins / polypeptides / peptides encompass proteins / polypeptides / peptides in which the amino acid sequence of native elastin proteins / polypeptides / peptides is mutated or contains one or more amino acids at the N- or C-terminus.
[0043] Therefore, the peptides, polypeptides and elastin peptides and their derivatives according to the present invention also encompass elastin-like peptides, polypeptides, peptides and their derivatives.
[0044] Elastin-like proteins, polypeptides, or peptides (ELPs) are synthetic molecules that Mainly includes Peptides or derivatives containing multiple repeating units.
[0045] "Mutated" peptide, polypeptide or protein means that the nucleic acid or amino acid sequence of the mutant peptide, polypeptide or protein contains one or more mutations. These mutations include deletions, substitutions, insertions and / or cleavages of one or more nucleic acids or amino acids.
[0046] There are many variants of elastin-like polypeptides, including repeat units of different elastin peptides, such as those of SEQ ID Nos. 1 to 7.
[0047] The elastin-like polypeptide or protein described in the present invention may thus, for example, comprise n-fold repeats of the pentapeptide: (KGGVG) n 、(VGGVG) n 、(GVGVP) n 、(VPGXG) n ; Hexapeptide: (VGVAPG) n ; Heptapeptide: (LGAGGAG) n ; or nonapeptide: (LGAGGAGVL) n or derivatives thereof. The number n may be, for example, 2 to 200, preferably 2 to 100.
[0048] As an example, the elastin peptide described in the present invention is the hexapeptide of SEQ ID N° 5 (VGVAPG) or a derivative thereof.
[0049] The elastin-like polypeptide or protein described in the present invention may therefore, for example, comprise n-fold repeats of the hexapeptide: (VGVAPG)n or a derivative thereof. The number n may, for example, be 2 to 200, preferably 2 to 100.
[0050] In particular, the peptide derivative of SEQ ID No. 5 contains the VGVAPG sequence with one or more amino acids at its N- and / or C-terminus.
[0051] In particular, the supplementary amino acid in the peptide derivative of SEQ ID N° 5 is aspartic acid or glutamic acid.
[0052] More particularly, the derivative may be SEQ ID N° 8 (VGVAPGD) or SEQ ID N° 9 (VGVAPGE).
[0053] Still particularly, the elastin-like polypeptides described in the present invention include repetitions of the hexapeptide of SEQ ID No. 5, more particularly, four times this hexapeptide (designated in the present invention as ELP4 and having SEQ ID No. 81: VGVAPGVGVAPGVGVAPGVGVAPG). Another example of the elastin-like polypeptide derivatives described in the present invention includes repetitions of the hexapeptide of SEQ ID No. 9, more particularly, four times this hexapeptide (designated in the present invention as ELPE4 and having SEQ ID No. 80: VGVAPGEVGVAPGEVGVAPGEVGVAPGE).
[0054] An interesting feature of ELPs is that they can self-aggregate in response to increasing temperature or changes in pH and ionic strength. This feature can facilitate the purification of ELPs after production in host cells.
[0055] Collagen is a superfamily of structurally related proteins that constitute the fundamental building blocks of connective tissue and participate in many biological functions in animals. These proteins exhibit a characteristic triple-helical tertiary structure resulting from the association of three polypeptide chains comprising the repeating amino acid sequence Gly-XY (or GXY), where amino acids X and Y are typically proline or 4-hydroxyproline, which participate in triple helix formation.
[0056] In humans, there are at least 27 different types of collagen found in different tissues (such as bones, skeleton, skin, tendons, blood vessels, eyes, etc....).
[0057] In the context of the present invention, a recombinant peptide, polypeptide or protein of collagen comprises repeating units of the collagen motif, ie repeating units of the sequence GXY.
[0058] “Derivatives” of collagen proteins / polypeptides / peptides encompass collagen-like proteins / polypeptides / peptides and proteins / polypeptides / peptides in which the amino acid sequence of native collagen proteins / polypeptides / peptides is mutated or contains one or more amino acids at its N- or C-terminus. Such supplementary amino acids may be any amino acids. “Derivatives” of collagen proteins / polypeptides / peptides also encompass derivatives of collagen-like proteins / polypeptides / peptides. “Derivatives” of collagen-like proteins / polypeptides / peptides encompass proteins / polypeptides / peptides in which the amino acid sequence of native collagen-like proteins / polypeptides / peptides is mutated or contains one or more amino acids at its N- or C-terminus. Such supplementary amino acids may be any amino acids.
[0059] Collagen-like proteins, peptides, or polypeptides Mainly includes The collagen motif of repeated units, that is, the sequence of repeated units GXY (and the peptides, polypeptides or proteins of collagen Contains only those repeating units ).
[0060] These repeating units are called "collagen-like domains" and are capable of forming triple helices (as collagens), such as C-type lectins (collectins) that are involved in host defense mechanisms.
[0061] Furthermore, screening of genomic databases for genes encoding collagen-like sequences containing repetitive GXY motifs has identified genes in bacterial and bacteriophage genomes. However, these organisms appear to lack proline hydroxylases. Recent studies have shown that two recently identified Streptococcal collagen-like proteins, Scl1 and Scl2, used as models, are able to form stable triple helices without hydroxylation of proline residues.
[0062] In particular, the collagen-like protein, peptide or polypeptide in the present invention is in particular and / or is derived from (one or more amino acids of) the CCMP2712 protein (referred to as GtCLP) of microalgae, for example, Guillardia theta.
[0063] In one embodiment, the derivative according to the invention comprises an amino acid sequence that is at least 80% identical to the amino acid sequence of the recombinant peptide, polypeptide or protein according to the invention.
[0064] By "an amino acid sequence that is at least 80% identical" is specifically intended an amino acid sequence that is 81, 82, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99% identical. By an amino acid sequence that is at least 95% "identical" to a query amino acid sequence of the invention, it is intended that the amino acid sequence of the peptide, polypeptide or protein is identical to the query sequence, except that the amino acid sequence may include up to five amino acid alterations per 100 amino acids of the query amino acid sequence. In other words, up to 5% (5 out of 100) of the amino acid residues in the sequence may be inserted, deleted or substituted with another amino acid in order to obtain an amino acid sequence that is at least 95% identical to a query amino acid sequence.
[0065] In the framework of the present application, the percentage of identity is calculated using a global comparison (i.e., comparing two sequences within their entire length). The method for comparing the identity of two or more sequences is well known in the art. For example, the "needle" program can be used, which uses the Needleman-Wunsch global alignment algorithm (Needleman and Wunsch (1970) J.Mol.Biol.48:443-453) to find their optimal comparison (including room) when considering the entire length of two sequences. For example, the needle program can be obtained on the ebi.ac.uk World Wide Web website. Preferably, EMBOSS::needle (global) program (wherein "room opens" parameter equals 10.0, "room expands" parameter equals 0.5) and Blosum62 matrix are used to calculate the percentage of identity according to the present invention.
[0066] An amino acid sequence that is "at least 80%, 85%, 90%, 95% or 99% identical" to a reference sequence may include mutations, such as deletions, insertions and / or substitutions, compared to the reference sequence. In the case of substitutions, an amino acid sequence that is at least 80%, 85%, 90%, 95% or 99% identical to a reference sequence may correspond to a homologous sequence derived from another species other than the reference sequence. In another preferred embodiment, the substitutions preferably correspond to conservative substitutions as shown in the following table.
[0067]
[0068] According to the present invention, the nucleic acid sequence encoding the recombinant protein, polypeptide or peptide is a nucleic acid sequence encoding the above protein, peptide or polypeptide.
[0069] In particular, as proteins, nucleic acid sequences encoding elastin proteins, elastin-like proteins, collagen proteins or collagen-like proteins come into consideration, in particular collagen-like proteins come into consideration.
[0070] For example, the following nucleic acid sequence encoding a collagen-like protein, 3f-tv-Gtclp-tv-ha (SEQ ID No. 10) can be cited.
[0071] In particular, as peptides, nucleic acid sequences encoding collagen-like peptides, collagen peptides, elastin peptides and elastin-like peptides are considered; more particularly, nucleic acid sequences encoding collagen-like peptides or elastin-like peptides.
[0072] For example, the following nucleic acid sequences can be cited: SEQ ID N°11 (GTAGGTGTAGCTCCTGGT), SEQ ID N°12 (GTTGGTGTTGCTCCTGGA), SEQ ID N°13 (GTAGGTGTTGCTCCAGGT) and SEQ ID N°14 (GTGGGTGTAGCTCCTGGT), all of which encode the aforementioned elastin peptide VGVAPG. Other nucleic acid sequences can be cited as examples, such as SEQ ID N°15 (GTAGGTGTAGCTCCTGGTGAA), SEQ ID N°16 (GTTGGTGTTGCTCCTGGAGAA), SEQ ID N°17 (GTAGGTGTGTGTCCAGGTGAA) and SEQ ID N°18 (GTGGGTGTAGCTCCTGGTGAA), all of which encode the aforementioned elastin peptide derivative VGVAPGE.
[0073] In particular, with respect to polypeptides, nucleic acid sequences encoding collagen-like polypeptides, collagen polypeptides, elastin polypeptides, and elastin-like polypeptides are contemplated; more specifically, nucleic acid sequences encoding collagen polypeptides or elastin-like polypeptides.
[0074] For example, the following nucleic acid sequences can be cited: ha-sp-3f-Gtccld-3ha (SEQ ID N°19), ha-sp-3f-Gtccld (SEQ ID N°20) and ha-sp-3f-Gtcld (SEQ ID N°21), elp4 (SEQID N°22) and elpe4 (SEQID N°23).
[0075] Still particularly, the nucleic acid sequence encoding the derivative according to the present invention includes a nucleic acid sequence at least 80% identical to the nucleic acid sequence encoding the recombinant peptide, polypeptide or protein according to the present invention." Nucleic acid sequence at least 80% identical" particularly means a nucleic acid sequence of 81, 82, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99% identical. For example, a nucleic acid sequence 95% "identical" to a query sequence of the present invention is intended to mean that the sequence of the polynucleotide is identical to the query sequence, except that the sequence can include up to five nucleotide changes in 100 nucleotides of each query sequence. In other words, in order to obtain a polynucleotide with a sequence at least 95% identical to the query sequence, up to 5% (5 in 100) of the sequence can be inserted, deleted or replaced with another nucleotide. In other words, sequences should be compared over their entire length (i.e., by preparing a global comparison). For example, a first polynucleotide of 100nt (nucleotides) compared within a second polynucleotide of 200nt is 50% identical to said second polynucleotide. For example, the needle program can be used, which uses the Needleman-Wunsch global alignment algorithm (Needleman and Wunsch (1970), Ageneral methodapplicable to the search for similarities in the amino acidsequence of two proteins, J.Mol.Biol.48:443-453) to find the best alignment of two sequences (including gaps) over the entire length of the two sequences. Preferably, the identity percentage according to the present invention is calculated using the needle program (wherein the "gap opening" parameter is equal to 10.0, the "gap extension" parameter is equal to 0.5) and the Blosum 62 matrix. For example, the needle program is available on the ebi.ac.uk World Wide Web site.
[0076] In one embodiment, the nucleic acid sequence encoding a protein, polypeptide or peptide comprising amino acid repeating units according to the present invention is codon-optimized for expression in the chloroplast genome of a microalgae host cell.
[0077] As mentioned above, the nucleic acid sequence according to the present invention is introduced into an expression vector capable of expressing the nucleic acid sequence.
[0078] "Introducing" means cloning the nucleic acid sequence encoding the recombinant protein / polypeptide / peptide into an expression vector using methods well known to the skilled person and in a manner that results in the expression of the nucleic acid sequence.
[0079] "Expression vector" or "transformation vector" or "recombinant DNA construct" or similar terms are defined herein as DNA sequences required for the transcription of recombinant genes and the translation of their mRNA in microalgal host cells. An "expression vector" comprises one or more expression cassettes for expressing recombinant genes (one or more genes encoding a protein, peptide or polypeptide of interest and typically a selectable marker). In the case of chloroplast genome transformation, the expression vector also contains homologous recombination regions for integration of the expression cassette into the chloroplast genome.
[0080] In the context of the present invention, the expression vector may in particular be a circular molecule having a plasmid backbone containing two homologous recombination regions and flanking the expression cassette, or a linearized molecule corresponding to an expression vector linearized by enzymatic digestion or to a PCR fragment containing only the expression cassette flanked by two homologous recombination regions.
[0081] In particular, the expression vector of the invention comprises at least one expression cassette and is, for example, the vector pCO86, pCO96, pCO26, pCO28, pLA01, pLA02, pAL03 or pAL04.
[0082] An "expression cassette" contains a coding sequence operably fused to one or more regulatory elements or regulatory sequences (eg, fused to a promoter and / or 5'UTR at its 5' end and / or to a 3'UTR at its 3' end).
[0083] A "coding sequence" is the portion of a gene and its corresponding transcribed mRNA that is translated into a recombinant protein / polypeptide / peptide. A coding sequence includes, for example, a translation initiation control sequence and a stop codon. In some embodiments, an expression cassette may contain a polycistronic sequence comprising more than one coding sequence encoding several proteins under the control of only one promoter / 5'UTR and 3'UTR.
[0084] The expression cassette is flanked by left (LHRR) and right (RHRR) endogenous sequences identical to the sequences surrounding the targeted integration site in the chloroplast genome. These left (LHRR) and right (RHRR) homology regions allow for integration of the expression cassette following homologous recombination exchange between the homologous regions.
[0085] Homologous recombination is the ability of complementary DNA sequences to align and exchange homologous regions. A transgenic DNA ("donor") containing a sequence homologous to the targeted genomic sequence ("template") is introduced into an organism and then recombined into the genome at the site of the corresponding genomic homologous sequence.
[0086] By its very nature, homologous recombination is a precise gene targeting event; therefore, most transgenic lines generated with the same targeting sequence will be essentially identical in phenotype, requiring far fewer transformation events to be screened.
[0087] In the case of chloroplast genome transformation of microalgae, integration of the expression cassette within the chloroplast genome occurs following homologous recombination between endogenous homologous sequences of the expression vector, where the genomic sequences are identical or similar to sequences surrounding the targeted integration site in the chloroplast genome.
[0088] In the context of the present invention, different integration sites can be used between genes rbcL and atpA, or psaB and trnG, or atpB and 16SrDNA, or psaA exon 3 and trnE, or trnE and psbH, or psbN and psbT, or psbB and trnD.
[0089] In some embodiments, to enhance its accumulation, the recombinant protein or polypeptide or peptide can be fused to an endogenous protein, for example, to the large subunit of ribulose bisphosphate carboxylase (Rubisco LSU). In this case, after homologous recombination of the transformation vector into the chloroplast genome, the promoter and 5'UTR will be those of the endogenous rbcL gene.
[0090] Depending on the processing system chosen, the protein, peptide or polypeptide of interest will be further separated from RBCL either in vivo (using self-cleaving peptides) or in vitro (by site-specific proteolysis).
[0091] In one embodiment, the coding sequence of the expression cassette according to the present invention further comprises a nucleic acid sequence encoding an epitope tag, particularly a Flag epitope tag, more particularly a Flag epitope tag repeated three times (3xFlag tag), in order to identify and / or purify the recombinant protein, polypeptide or peptide.
[0092] In particular, the epitope tag sequence is located at the N-terminus of the protein, peptide or polypeptide. More particularly, another epitope tag sequence may be placed alone or in addition to one at the N-terminus at the C-terminus of the protein, peptide or polypeptide according to the invention in order to monitor the release of the peptide / polypeptide / protein of interest in order to follow its cleavage, for example by an endoprotease.
[0093] Examples of epitope tag sequences are Flag tag (SEQ ID N°24: DYKDDDDK), 3xFlag tag (SEQ ID N°25: DYKDDDDKDYKDDDDKDYKDDDDK), HA tag (SEQ ID N°26: YPYDVPDYA), 3xHA tag (SEQ ID N°27: YPYDVPDYAYPYDVPDYAYPYDVPDYA), His tag (SEQ ID N°28: HHHHHH), which are described in the experimental part of the present invention.
[0094] In one embodiment, the coding sequence in the expression cassette comprises a nucleic acid sequence that encodes not only the recombinant protein, polypeptide or peptide but also an amino acid sequence that allows the production of the recombinant protein, polypeptide or peptide in a specific cellular compartment.
[0095] As used herein, "promoter" refers to a nucleic acid control sequence that directs transcription of a nucleic acid.
[0096] The "5'UTR" or 5' untranslated region (also called leader sequence or leader RNA) is the region of the mRNA directly upstream of the start codon.
[0097] The "3'UTR" or 3' untranslated region is the portion of the messenger RNA (mRNA) that immediately follows the translation stop codon.
[0098] The 5'UTR and 3'UTR are required for transcript (mRNA) stability and translation initiation.
[0099] For microalgae chloroplast expression, promoters, 5'UTRs and 3'UTRs that can be used in the context of the present invention are, for example: the promoters and 5'UTRs of the genes psbD, psbA, psaA, atpA and atpB; the 16S rRNA promoter (Prrn) promoter fused to the 5'UTR; the psbA 3'UTR; the atpA 3'UTR; or the rbcL 3'UTR.
[0100] A 5'UTR from an exogenous source, such as the 5'UTR of gene 10L from bacteriophage T7, can also be fused downstream of a microalgae promoter. In particular, the nucleic acid sequence is operably linked to the Chlamydomonas reinhardtii 16S rRNA promoter (Prrn) at its 5'-end.
[0101] Stable expression and translation of the nucleic acid sequence according to the invention can be controlled, for example, by the promoter and 5'UTR from psbD and the atpA 3'UTR.
[0102] Furthermore, in one embodiment of the recombinant microalgae or method according to the present invention, the nucleic acid sequence encoding the recombinant protein, polypeptide or peptide is operably linked to at least one regulatory sequence selected from the group consisting of: psbD promoter and 5'UTR (SEQ ID N°29); or 16S rRNA promoter (Prrn) promoter fused to atpA 5'UTR (SEQ ID N°30); psaA promoter and 5'UTR; atpA promoter and 5'UTR; 3'UTR from atpA (SEQ ID N°31) and rbcL (SEQ ID N°32).
[0103] In one embodiment, a promoterless gene encoding a protein, peptide or polypeptide of interest can be integrated after the homologous recombination region within the chloroplast genome and just downstream of the native promoter.
[0104] As mentioned above, the chloroplast genome of the microalgae host cell is transformed with the expression vector. Genetic transformation of microalgae host cells, and more particularly the chloroplast genome of microalgae, by the expression vector according to the present invention can be performed according to any suitable technique well known to those skilled in the art, including but not limited to biolistic methods (Boynton et al., 1988; Goldschmidt-Clermont, 1991), electroporation (Fromm et al., (1985) Proc. Natl. Acad. Sci. (USA) 82:5824-5828; see Maruyama et al., (2004) Biotechnology Techniques 8:821-826), glass bead transformation (Purton et al., revue), protoplasts treated with CaCl2 and polyethylene glycol (PEG) (see Kim et al., (2002) Mar. Biotechnol. 4:63-73) or microinjection.
[0105] Specifically, the transformation uses a helium gun bombardment technique of gold microparticles complexed with transforming DNA.
[0106] To identify microalgae transformants, selectable marker genes can be used. For example, the aadA gene, which encodes aminoglycoside 3″-adenylyltransferase and confers resistance to spectinomycin and streptomycin in the case of Chlamydomonas reinhardtii chloroplast transformation, can be mentioned. In another embodiment, the selectable marker gene can be the aphA-6Ab gene of Acinetobacter baumannii, which encodes type VI 3′-aminoglycoside phosphotransferase and confers kanamycin resistance.
[0107] Therefore, chloroplast genome engineering can be performed using selectable marker genes that confer antibiotic resistance or using the rescue of photosynthetic mutants.
[0108] In particular, in one embodiment, the expression vector for chloroplast genome transformation comprises two expression cassettes comprising nucleic acid sequences encoding the recombinant protein, polypeptide or peptide or the selectable marker gene according to the present invention.
[0109] More particularly, the expression vector for chloroplast genome transformation comprises an expression cassette comprising a nucleic acid sequence encoding a recombinant protein, polypeptide or peptide according to the present invention and an expression cassette comprising the aadA gene encoding aminoglycoside 3″-adenylyl transferase.
[0110] In another embodiment of the present invention, corresponding to the case of rescuing a photosynthetic mutant that is sensitive to light, the expression vector includes a wild-type RHRR region that is deleted in the mutant. After homologous recombination, the deleted region is restored in the genome of the photosynthetic mutant, and the photosynthetic mutant is then able to grow under light.
[0111] In particular, the coding sequence of the expression cassette according to the present invention also includes a nucleic acid sequence encoding a signal peptide. "Signal peptide" (SP) in the present invention means an amino acid sequence located at the N-terminus of a newly synthesized recombinant protein, polypeptide or peptide. The signal peptide should allow the protein to translocate within the lumen of the chloroplast thylakoid rather than within the chloroplast stroma. The signal peptide is cleaved after translocation across the thylakoid membrane.
[0112] For example, such a signal peptide sequence is selected from proteins known to be translocated within the lumen of the thylakoid using the twin-arginine protein translocation (Tat) pathway or the Sec pathway.
[0113] For example, the signal peptide can be derived from an algal protein localized in the thylakoid lumen, such as the signal peptides from the Chlamydomonas reinhardtii 16 and 23 kDa subunits of the oxygen-evolving complex of photosystem II, or the Chlamydomonas reinhardtii Rieske subunit of the b6f complex, or the α subunit of cryptophytic phycoerythrin (e.g., from Cryptophyta cyanobacteria).
[0114] In particular, a signal peptide can be extracted from the sequence of the Escherichia coli TorA gene encoding trimethylamine-N-oxide reductase 1 (UniProt No. P33225) (SP; SEQ ID No. 33: NNNDLFQAASRRRFLAQLGGLTVAGMLGPSLLTPRRATAAQA; the nucleic acid sequence encoding SP is SEQ ID No. 34). This signal peptide utilizes the Tat system. After the protein passes through the thylakoid membrane, this amino acid sequence is cleaved from the protein.
[0115] Other signal peptides may be used which do not leave additional amino acids at the N-terminus of the recombinant protein, such as in particular signal peptides from algae, and in particular from Chlamydomonas reinhardtii.
[0116] In one embodiment, the recombinant protein or polypeptide or peptide may be produced as a fusion protein.
[0117] Therefore, the present invention also relates to the recombinant microalgae or the method according to the invention, wherein the nucleic acid sequence encoding the recombinant protein, polypeptide or peptide is operably fused at its 5' or 3' end to the nucleic acid sequence encoding the vector.
[0118] Fusion partners or vectors have been developed in recombinant protein production to increase accumulation yield and / or solubility and / or folding and / or facilitate protein purification. Fusion partners of different sizes (or molecular weights) have been used in various production systems to enhance protein solubility and accumulation (maltose binding protein (MBP), glutathione-S-transferase (GST), thioredoxin, GB1, N-utilization substance A (NusA), ubiquitin, small ubiquitin-like modifier (SUMO), Fh8) and to facilitate detection and purification (for example, but not limited to, MBP, GST, and small epitope tag peptides such as c-myc tags, polyhistidine tags (His Tag), Flag tags, HA tags. Another type of fusion tag for purification is a stimulus-responsive tag (or environmentally responsive polypeptide), which allows the fusion protein to precipitate when the stimulus is adjusted as a modification of temperature or solution ionic strength.
[0119] In particular, the carrier according to the invention is aprotinin.
[0120] The vector is fused to the recombinant protein, polypeptide or peptide to form a fusion protein.
[0121] "Aprotinin" means alkaline trypsin inhibitor (BPTI), a small single-chain protein cross-linked by three disulfide bridges, comprising 58 amino acid residues, with a molecular weight of 6.5 kDa and an isoelectric point of 10.9.
[0122] The protein is well known to those skilled in the art and is commercially available. For example, it can be produced in recombinant systems such as plants (in the cytoplasm by nuclear transformation (Pogue et al., 2010) or in the thylakoid lumen by chloroplast transformation (Tissot et al., 2008).
[0123] Its molecular formula is C 284 H 432 N 84 O 79 S7, and its molar mass is 6511.51 g / mol.
[0124] The amino acid sequence of aprotinin from Bos Taurus (cattle) is RPDFC LEPPY TGPCK ARIIR YFYNAKAGLC QTFVY GGCRA KRNNF KSAED CMRTC GGA (SEQ ID No. 35). The nucleic acid sequence encoding this amino acid sequence is SEQ ID No. 36.
[0125] In the context of the present invention, the term "aprotinin" also encompasses chimeric aprotinins and mutant aprotinins.
[0126] "Chimeric aprotinin" means that aprotinin is linked at its N-terminus and / or at its C-terminus to an epitope tag peptide and / or a signal peptide and / or a protease recognition cleavage site.
[0127] The chimeric aprotinin may be, for example, the protein designated HA-APRO (SEQ ID N° 37), which comprises
[0128] Aprotinin fused to an HA epitope tag at its N-terminus; or 3F-APRO (SEQ ID N° 39), which comprises aprotinin fused to a 3xFlag epitope tag (3F) at its N-terminus.
[0129] Other examples of chimeric aprotinins may be proteins designated HA-SP-3F-FX-APRO (SEQ ID N°41 and 42), which comprise aprotinin fused at its N-terminus to an amino acid sequence consisting of an HA epitope tag (HA), followed by a signal peptide (SP), a 3xFlag epitope tag (3F), and a cleavage site for Factor Xa (FX; IEGR); or aprotinin designated HA-SP-3F-APRO (SEQ ID N°43 and 44), which comprise aprotinin fused at its N-terminus to an HA epitope tag, followed by a signal peptide SP and a 3xFlag epitope tag (3F); or chimeric aprotinin designated HA-SP-APRO (SEQ ID N°45 and 46), which comprise aprotinin fused at its N-terminus to an HA epitope tag, followed by a signal peptide SP.
[0130] "Mutated" aprotinin means that the nucleic acid or amino acid sequence of a "mutated" aprotinin contains one or more mutations in the nucleic acid or amino acid sequence of aprotinin or chimeric aprotinin. These mutations include deletions, substitutions, insertions and / or cleavages of one or more nucleic acids or amino acids.
[0131] The signal peptide was as described previously.
[0132] Other signal peptides that do not leave additional amino acids at the N-terminus of the recombinant protein can be used. If the signal peptide is cleaved after translocation to the lumen of the chloroplast thylakoid (or across the thylakoid membrane), two other chimeric aprotinins 3F-APRO or 3F-FX-APRO (SEQ ID N° 47) can be produced in vivo.
[0133] In particular, according to the present invention, fusion partners are used to improve the accumulation and / or stability of recombinant peptides, polypeptides and proteins.
[0134] Still particularly, and as described above, the fusion protein further comprises a cleavage site recognized by a specific protease.
[0135] Cleavage sites recognized by specific proteases are well known to those skilled in the art. They are used to separate aprotinin from the recombinant protein, polypeptide or peptide of interest and should be removed if the carrier may interfere with the activity or structure of the protein, polypeptide or peptide and thus interfere with its use.
[0136] In particular, the cleavage site is an endoprotease and / or intracellular protease recognition sequence (or a protease cleavage site or a protease recognition site). More particularly, the sequence of the cleavage site is located between two coding sequences (one of aprotinin and one of the recombinant protein of interest, polypeptide or peptide according to the present invention).
[0137] Cleavage of the fusion protein can be performed in vivo (in recombinant host cells prior to extraction, or on the skin when applying the cosmetic peptide) or in vitro by addition of a protease after extraction and purification.
[0138] Non-limiting examples of proteases are Factor Xa (FX), tobacco edge virus protease (TEV), enterokinase (EK), SUMO protease, thrombin, human rhinovirus 3C protease (HRV 3C), intracellular protease Arg-C, intracellular protease Asp-C, intracellular protease Asp-N, intracellular protease Lys-C, intracellular protease Glu-C, proteinase K, IgA-protease, trypsin, chymotrypsin, and thermolysin.
[0139] Self-cleaving peptides can also be used, such as the interferon system (Yang et al., 2003), the viral 2A system (Rasala et al., 2012), or the preferredoxin site from Chlamydomonas (Muto et al., 2009).
[0140] In one embodiment, a linker can be placed between aprotinin and the protease cleavage site. Linkers can be divided into three types: flexible, rigid, and cleavable. The usual function of a linker is to fuse the two partners of a fusion protein (e.g., flexible linkers or rigid linkers) or to release them under specific conditions (cleavable linkers) or to provide other functions of the protein in drug design, such as improving their biological activity or their targeted delivery.
[0141] In one embodiment of the invention, the linker may also make the protease cleavage site more accessible to the enzyme, if desired.
[0142] In one embodiment, the flexible linker contains small non-polar (e.g., Gly) or polar (e.g., Ser or Thr) amino acids. Examples of such linkers are given in Chen et al., 2013.
[0143] The flexible linker according to the present invention may be LG (SEQ ID N° 49: RSGGGGSGGGGSGS) or LGM (SEQ ID N° 50: RSGGGGSSGGGGGGSSRS).
[0144] When a fusion protein with a carrier is involved, step (vii) of the method according to the invention is a step of purifying the fusion protein.
[0145] In that case, the method optionally comprises a step (viii) in which the fusion protein is cleaved.
[0146] The cleavage can be performed by any method known to those skilled in the art, such as using a suitable protease to release the recombinant peptide, polypeptide or protein.
[0147] Said step (viii) is optionally followed by a purification step (ix) of the recombinant protein, polypeptide or peptide.
[0148] In particular, the method further comprises a step (viii') between step (viii) and step (ix), wherein the polypeptide is cleaved to allow release of the peptide unit.
[0149] The cleavage may be performed by any method known to the person skilled in the art, such as using a suitable intracellular protease.
[0150] Characterization of microalgal host cells producing recombinant proteins, polypeptides or peptides can be performed by techniques known to those skilled in the art, such as by PCR screening for antibiotic-resistant transformants or Western blot analysis of total protein extracts.
[0151] Extraction of total protein can be performed using well-known techniques (centrifugation, lysis, sonication, etc.).
[0152] Identification of fusion proteins or recombinant proteins can be performed by Western blotting using specific antibodies.
[0153] Purification can be performed using well-known techniques. In one embodiment, it comprises affinity chromatography and / or a step of separating the peptide, polypeptide or protein according to the invention from the carrier (e.g. by enterokinase protease digestion) and / or size exclusion chromatography.
[0154] In one embodiment, the affinity chromatography step can be replaced by ion exchange chromatography, which is less expensive for large-scale purification.
[0155] According to the present invention, "microalgae" are eukaryotic microbial organisms containing chloroplasts or plastids and optionally capable of photosynthesis, or prokaryotic microbial organisms capable of photosynthesis (cyanobacteria).
[0156] In particular, the microalgae are selected from the group consisting of Chlorophyta (green algae), Rhodophyta (red algae), Stramenopiles (heterokonts), Xanthophyceae (yellow-green algae), Glaucocystophyceae (gray algae), Chlorarachniophyceae (chlorarachniophytes), Euglenida (euglena), Haptophyceae (coccolithophytes), Chrysophyceae (chrysophytes), phyceae (chrysophyceae), Cryptophyta (cryptoalgae), Dinophyceae (flagellate algae), Haptophyceae (coccolithophores), Bacillariophyta (diatoms), Eustigmatophyceae (eustigmatophytes), Raphidophyceae (raphidophytes), Scenedesmaceae, Phaeophyceae (brown algae).
[0157] More specifically, the microalgae are selected from the group consisting of Chlamydomonas, Chlorella, Dunaliella, Haematococcus, diatoms, Scenedesmus, Tetraselmis, Ostreococcus, Porphyridium and Nannochloropsis.
[0158] Even more particularly, the microalgae is chosen from the group consisting of Chlamydomonas, more particularly Chlamydomonas reinhardtii, even more particularly Chlamydomonas reinhardtii 137c or a defective strain of Chlamydomonas reinhardtii CW15.
[0159] In particular, the microalgae are cultured under classical conditions known to those skilled in the art. For example, Chlamydomonas reinhardtii is grown in TAP (Tris Acetate Phosphate) medium to mid-logarithmic phase (approximately 1-2 x 10 6 cells / ml), and / or at a temperature between 23°C and 25°C (ideally 25°C), and / or at constant light (70-150 μE / m 2 The experimental section explains the culture conditions.
[0160] All embodiments mentioned in the context of the present invention can be combined.
[0161] The present invention will be further explained by the following figures and examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0162] Figure 1 :Codon usage in the chloroplast genome of Chlamydomonas reinhardtii
[0163] Figure 2 : Schematic representation of the chloroplast transformation vectors used to produce the G. theta CCMP2712 protein (GtCLP), native and chimeric collagen-like domains (GtCLD) of G. theta CCMP2712.
[0164] Figure 3Western blot analysis of independent algal transformants CW-CO86 and CW-CO96 expressing the gene for the C. cyanobacteria CCMP2712 protein (A) and a chimeric collagen-like domain of C. cyanobacteria CCMP2712 (B) from the algal chloroplast genome using a monoclonal anti-Flag M2 antibody. 100 μg of each total soluble protein sample extracted from wild-type (WT) CW15 and transformants CW-CO86 and CW-CO96, lysed with SDS buffer, was separated on a 15% SDS polyacrylamide gel. MW: Molecular weight marker. Arrows indicate the location of the recombinant proteins.
[0165] Figure 4 Western blot analysis of the 137c-CO96-4 transformant using monoclonal anti-Flag M2 (A) or anti-HA (B) antibodies. 50 μg of total soluble protein from the 137c-CO96 transformant extracted by sonication was digested with enterokinase (EK) (+EK; A: lane 4; B: lane 5) or not (A: lanes 2 and 3; B: lanes 3 and 4). 50 μg of total soluble protein from wild-type (WT) 137c was lysed and extracted using SDS buffer. MW: molecular weight marker. Arrows indicate the position of recombinant proteins with or without the Flag tag.
[0166] Figure 5 : Western blot analysis of different elution fractions from anti-HA affinity chromatography of protein extracts from 137c-CO96-4 using a monoclonal anti-HA antibody. Protein samples (50 μg of protein extracted by sonication of wild-type (WT) 137c or 25 μg of elution or wash fractions) were loaded on a 15% SDS polyacrylamide gel. MW: molecular weight standard. Load: total soluble protein extracted by sonication before incubation with anti-HA resin. FT: flow through. EA: elution fraction. W: wash fraction. Arrows indicate the position of purified recombinant protein.
[0167] Figure 6 : Schematic representation of chloroplast transformation vectors for elastin polypeptide and peptide production.
[0168] Figure 7 Western blot analysis of algal cells transformed with pLA01 using monoclonal anti-Flag (A) or anti-HA (B) antibodies. 50 μg of total soluble protein samples extracted from wild-type (WT) CW15 cells and independent CW-LA01 transformants were lysed with SDS buffer and separated on a 15% SDS polyacrylamide gel. MW: molecular weight marker. Arrows indicate the location of recombinant proteins. Example
[0169] Example 1
[0170] Materials and methods
[0171] All oligonucleotides and synthetic genes were purchased from Eurofins. All enzymes were purchased from New England Biolabs, Promega, Invitrogen, and Sigma Aldrich / Merck. All plasmids were constructed on the pBluescript II backbone.
[0172] Algal strains and growth conditions
[0173] The two algae strains used were Chlamydomonas reinhardtii wild type (137c; mt+) and the cell wall-deficient strain CW15 (CC-400; mt+), obtained from the Chlamydomonas Resource Center, University of Minnesota).
[0174] Prior to transformation, all strains were grown at a temperature between 23°C and 25°C (ideally 25°C) in a room with constant light (70-150 μE / m 2 The cells were grown on a rotary shaker in TAP (Tris acetate phosphate) medium in the presence of 4% paraformaldehyde (Pb) and 2% paraformaldehyde (Pb) to mid-logarithmic phase (approximately 1-2 × 10 6 cells / mL density).
[0175] Transformants were grown under the same conditions and in the same medium containing 100 μg / mL spectinomycin or 100 μg / mL kanamycin, depending on the selectable marker gene present in the transformation vector.
[0176] Growth kinetics were also followed by measuring the optical density at 750 nm using a spectrophotometer.
[0177] Algae transformation
[0178] Chlamydomonas reinhardtii cells were transformed using the helium gun bombardment technique using gold microprojectiles complexed with transforming DNA as described in Boynton et al., 1988. Briefly, Chlamydomonas reinhardtii cells were grown to mid-logarithmic growth in TAP medium, harvested by gentle centrifugation, and then resuspended in TAP medium to 1.10 μg / mL. 8 The final concentration of cells / mL was 400 μL. Depending on the selectable marker gene present in the transformation vector, 300 μL of this cell suspension was plated on TAP agar medium supplemented with 100 μg / mL spectinomycin or 100 μg / mL kanamycin. As described by the manufacturer, the plate was bombarded with gold particles (S550d; Seashell Technology) coated with the transformation vector. The plate was then placed at 25°C under standard light conditions to allow selection and formation of transformed colonies.
[0179] Total DNA extraction and PCR screening of positive transformants
[0180] Total DNA was extracted from single colonies (approximately 1 mm in diameter) of wild-type and / or antibiotic-resistant transformant Chlamydomonas strains using the chelating resin Chelex 100 (Biorad).
[0181] From the isolated colonies, a few cells corresponding to a diameter of approximately 0.5 mm were picked with a sharp tip and resuspended in 20 μL of HO. 200 μL of ethanol was added and incubated at room temperature for 1 minute. 200 μL of 5% Chelex was added and vortexed. After incubation at 100°C for 8 minutes, the mixture was cooled and centrifuged at 13,000 rpm for 5 minutes. Finally, the supernatant was collected.
[0182] Following transformation, algal colonies growing on restrictive solid medium plates are expected to have the antibiotic resistance gene and other transgenes incorporated into their genome.
[0183] To identify stable integration of the recombinant gene into the algal genome, antibiotic-resistant transformants were screened by polymerase chain reaction (PCR or PCR amplification) in a thermal cycler using 1 μL of previously extracted total DNA as a template, two synthetic and specific oligonucleotides (primers), and Taq polymerase (GoTaq, Promega). PCR amplification cycles followed the manufacturer's recommended guidelines. PCR reactions were subjected to gel electrophoresis to examine the PCR fragment of interest.
[0184] Protein extraction and western blot analysis
[0185] Chlamydomonas cells were collected by centrifugation (50 mL, 1-2.10 6 Cells were collected at 4% 4% PBS (100 mM EDTA) and the cell pellet was resuspended in lysis buffer (50 mM Tris-HCl pH 6.8, 2% SDS, 10 mM EDTA). As described in further detail below, in some embodiments of the Examples, the lysis buffer does not contain 10 mM EDTA. After 30 minutes at room temperature, cell debris was removed by centrifugation at 13,000 rpm, and the supernatant containing total soluble protein was collected.
[0186] Total soluble protein was extracted under native conditions in various buffers, depending on the analytical procedure. The cell pellet was resuspended in a buffer containing 50 mM Tris-HCl (pH 6.8 or 8) or 20 mM Tris-HCl (pH 6.8 or 8). The algal cell suspension was kept on ice during the sonication step, using a cell disruptor, a sonicator FB505500W (Sonic / FisherBrand), and a microtip probe set to 20% power for 5 minutes. Following sonication, cell debris was removed by centrifugation at 13,000 rpm for 30 minutes.
[0187] Total soluble protein present in the supernatant was quantified using the Pierce BCA protein assay kit according to the supplier's instructions (Thermofisher).
[0188] Total soluble protein samples (50 or 100 μg or another amount mentioned further in the examples depending on the experiment) were separated in 12 or 15% Tris-glycine SDS-PAGE prepared according to Laemmli (1970).
[0189] For experiments performed under reducing conditions, samples were prepared in Laemmli sample loading buffer with 50 mM DTT (or more, depending on the fusion protein) or 5% β-mercaptoethanol and further denatured at 95° C. for 5 min before loading. SDS PAGE experiments were performed using a Protein Gel tank from BioRad.
[0190] After separation, transfer the cells using standard transfer buffer and Turbo TM The sample was blotted onto a nitrocellulose membrane (GE Healthcare) using a transfer system. To visualize the transferred proteins, the nitrocellulose membrane was stained with Ponceau S dye. The membrane was further blocked with Tris-buffered saline Tween buffer (TBS-T) (50 mM Tris-HCl pH 7.5, 150 mM NaCl, 0.1% Tween-20) containing 5% bovine serum albumin (BSA). After gentle shaking at room temperature for one hour, the membrane was incubated overnight at 4°C with TTBS buffer containing mouse primary antibodies (see Table 1).
[0191] Table 1: Primary Antibodies
[0192]
[0193]
[0194] After washing three times with TBS-T-BSA buffer, the membrane was incubated with TBS-T-BSA buffer containing secondary antibody (anti-mouse IgG (H+L), HRP conjugate; Promega) at room temperature for one hour. After washing four times with TTBS buffer and once with TBS buffer, the membrane was incubated in enhanced chemiluminescence (ECL) substrate (Clarity Max ECL substrate; Biorad). ChemiDoc was used. TM ECL signals were visualized using an XRS+ system (Biorad).
[0195] Protein purification
[0196] After centrifugation, the algal cell pellet was resuspended in different buffers, depending on the protein and the subsequent steps to be performed. If the next step was anti-FLAG M2 affinity chromatography, the buffer contained 50 mM Tris-HCl pH 8, 500 mM NaCl, and 0.1% Tween 20. If the next step was anti-HA affinity chromatography, the buffer contained 20 mM Tris-HCl pH 8. Approximately, 10 mL of buffer was used per gram of wet algal cells, depending on the transformant strain. The resuspended cells were sonicated under the same conditions as previously described.
[0197] Affinity chromatography
[0198] All recombinant proteins were tagged at their N-termini with a Flag tag epitope that bound specifically to anti-Flag M2 affinity gel (Sigma / Merck). This resin contained a mouse monoclonal antibody covalently linked to agarose. M2 antibody.
[0199] All steps of the experiment were performed as described by the manufacturer. Briefly, total soluble protein samples were filtered using cellulose acetate 0.45 μm filters and stained with the antibody prepared as recommended by the manufacturer. M2 affinity gel is mixed and balanced in binding buffer (50mM Tris-HCl pH8, 500mM NaCl, 0.1% Tween 20). Roughly, 1mL of resin is used for every 4 to 8g of wet algae cells, depending on the transformant. The binding of the recombinant fusion protein is carried out at 4°C for 4 hours or overnight, gently and continuously mixing by inversion. After incubation, the soluble protein mixture incubated with the resin is loaded onto an empty Bio-rad Econo-pac column by gravity or collected by centrifugation, and washed several times with 40 column volumes of TBS and 20 column volumes of TBS. The protein of interest is eluted from the resin using 100mM glycine pH 3.5, 500mM NaCl, and neutralized to a final concentration of 50mM with Tris-HCl pH 8.
[0200] Some recombinant proteins were tagged with an HA epitope tag that specifically binds to an anti-HA agarose resin (Pierce / ThermoScientific). All steps of the experiment were performed as described by the manufacturer. Briefly, a filtered sample of total soluble protein was mixed with an anti-HA agarose resin prepared as recommended by the manufacturer and balanced in TBS and incubated overnight at 4°C with gentle continuous inversion mixing or on a shaking platform. After incubation, the resin was precipitated at 12,000g for 5 to 10 seconds (repeated 3 times). The supernatant was retained for further analysis. The precipitated resin was washed several times with 10 bed volumes of TBST. After incubating the resin with 10 bed volumes of 1 mg / ml Pierce HA peptide at 30°C for 15 minutes, the protein of interest was eluted from the resin. The resin was precipitated by centrifugation (at 12,000g for 5 to 10 seconds). The supernatant containing the fusion protein was collected. This elution step was repeated another 3 times.
[0201] Each elution fraction of affinity chromatography was further analyzed by SDS-PAGE and western blotting.
[0202] According to further steps, as described by the manufacturer, the eluted fractions containing the protein of interest are dialyzed against the buffer used in the protease digestion in Slide-A-Lyzer dialysis cassettes (3.5 kDa MWCO, Thermo Scientific). The dialyzed samples are concentrated using Vivaspin 6 (3 kDa MWCO, GE Healthcare).
[0203] Separation of the protein of interest from the carrier
[0204] Separation of the protein of interest from the carrier is performed by protease digestion, in particular in the present invention by enterokinase (light chain) or tobacco etch virus (TEV) protease from New England BioLabs (NEB).
[0205] Enzymatic digestion was performed according to the manufacturer's recommendations.
[0206] For example, for enterokinase light chain digestion, the reaction was combined with 25 μg of the protein of interest and 1 μL of enterokinase light chain in 20 μL of buffer (20 mM Tris-HCl pH 8.0, 50 mM NaCl, 2 mM CaCl2) and incubated at 25°C for 16 h.
[0207] For example, for TEV digestion, the manufacturer recommends combining 15 μg of protein substrate with 5 μL of TEV protease reaction buffer (10X) to create a total reaction volume of 50 μL. After adding 1 μL of TEV protease, the reaction is incubated at 30°C for 1 hour or at 4°C overnight.
[0208] For example, for Factor Xa digestion, the manufacturer recommends digesting 50 μg of fusion protein with 1 μg of FXa in a volume of 50 μL for 6 hours at 23° C. The reaction buffer includes 20 mM Tris-HCl pH 8.0, 100 mM NaCl, and 2 mM CaCl 2 .
[0209] Cleavage of peptides by intracellular proteases
[0210] The selection of the intracellular protease for cutting the polypeptide of interest is based on the amino acid sequence of the polypeptide. The intracellular protease can be, for example, intracellular protease Glu-C, intracellular protease Arg-C, intracellular protease Asp-C, intracellular protease Asp-N or intracellular protease Lys-C.
[0211] Enzyme digestion was performed according to the manufacturer's recommendations. For example, for intracellular protease Glu-C digestion (from NEB), the manufacturer recommends digesting 1 μg of substrate protein with 50 ng of intracellular protease Glu-C at 37°C for 16 h. The reaction buffer included 50 mM Tris-HCl pH 8.0 and 0.5 mM GluC-GluC.
[0212] Size Exclusion Chromatography (SEC)
[0213] The purified and digested fusion protein was subjected to size exclusion chromatography using an AKTA Pure system (GE Healthcare) to separate the protein of interest from the vector.
[0214] First, a Superdex S30 Plus G10 / 300GL column (GE Healthcare) and a HiLoad 26 / 600 Superdex 30 preparative column were calibrated using two standards: aprotinin (bovine lung; 6.5 kDa) and glycine (75 Da) diluted in 2X PBS buffer (or appropriate buffer for further steps).
[0215] After the washing step in water, a Superdex S30 G10 / 300GL chromatographic column is added and balanced in running buffer (2XPBS, pH 7.4; or 1X PBS, pH 7.4; or the appropriate buffer for further steps), and 200 to 500 μL samples are flowed through the chromatographic column at a speed of 0.5mL / min. The elution of protein is detected by measuring the optical absorbance at 280, 224, and 214nm. 0.5mL fractions are collected and analyzed by SDS-PAGE, followed by western blotting or staining with Coomassie blue dye.
[0216] After the washing step in water, a HiLoad 26 / 600 Superdex 30 preparative column was balanced in running buffer (2X PBS, pH 7.4; or 1X PBS, pH 7.4; or an appropriate buffer for further steps), and a sample (4 to 30 mL) was flowed through the column at a speed of 2.6 mL / min. Protein elution was detected by measuring the optical absorbance at 280, 224, and 214 nm. 4 mL fractions were collected and analyzed by SDS-PAGE followed by Western blotting.
[0217] In some embodiments, the eluted fractions of interest are pooled and evaporated using a SpeedVac (Eppendorf). The peptides or polypeptides or proteins present in these evaporated samples are subjected to Edman degradation to confirm the N-terminal amino acid sequence of the protein of interest.
[0218] Example 2
[0219] Collagen-like proteins, collagen-like domains and their roles in chloroplasts transformed with the chloroplast genome of Chlamydomonas reinhardtii Production of collagen-like peptides
[0220] Construction of transformation vector for expression of collagen-like proteins
[0221] To generate novel collagen-like proteins (CLPs) and / or collagen-like domains, we screened databases for collagen-like genes encoding collagen-like domains from various sources. One of the sequences discovered was the CCMP2712 protein from the microalga G. theta. The amino acid sequence of the CCMP2712 protein from G. theta (designated GtCLP, SEQ ID N° 51) contains a collagen-like domain and was extracted from GenBank Accession No. XM_005827950.
[0222] Nothing has been described regarding the ability of the proteins encoded by the identified genes to form collagen-like triple helical structures.
[0223] In Chlamydomonas reinhardtii, codon usage has been shown to play a significant role in protein accumulation (Franklin et al., 2002; Mayfield and Schultz, 2004).
[0224] The nucleic acid sequence encoding CCMP2712 protein was designed and optimized to improve its expression in Chlamydomonas reinhardtii host cells.
[0225] Methods for altering nucleic acid sequences for improved expression in host cells are known in the art, particularly in algal cells, in particular in Chlamydomonas reinhardtii.
[0226] The codon usage database was found at http: / / www.kazusa.or.jp / codon / (see Codon usage of Chlamydomonas reinhardtii chloroplast genome; Figure 1 ).
[0227] To improve expression of genes of interest in C. reinhardtii chloroplasts, codons that are not commonly used in their native sequences are replaced with codons encoding the same or similar amino acid residues that are more commonly used in the C. reinhardtii chloroplast codon bias. In addition, other codons are replaced to avoid multiple or extended codon repeats, restriction enzyme sites, or a higher probability of secondary structures that could reduce or interfere with expression efficiency.
[0228] In order to check and fulfill all the above mentioned criteria, the amino acid sequence of the protein of interest was also optimized by the software GENEius from Eurofins using the appropriate codon usage of the Chlamydomonas reinhardtii chloroplast.
[0229] After its codon optimization, the gene Gtclp encoding the native CCMP2712 protein from C. cyanobacteria (referred to as "recombinant GtCLP" or 3F-TV-GtCLP-TV-HA) was designed to be operably fused at its 5' end to a codon-optimized nucleic acid sequence encoding an amino acid sequence containing a 3xFlag epitope tag (SEQ. DYKDDDDKDYKDDDDKDYKDDDDK; SEQ ID N°25) followed by a recognition site for TEV protease (SEQENLYFQG; SEQ ID N°52). At its 3' end, the optimized gene Gtclp was operably fused to an optimized nucleic acid sequence encoding a TEV protease recognition site followed by an HA epitope tag (SEQ ID N°26). The recombinant GtCLP produced in vivo in Chlamydomonas reinhardtii chloroplasts is also referred to as 3F-TV-GtCLP-TV-HA. The two recognition sites for TEV protease allow for the removal of the 3xFlag and HA tags by in vitro protease digestion.
[0230] The resulting fusion gene 3f-tv-Gtclp-tv-ha (SEQ ID N° 10) encoding the recombinant GtCLP designated 3F-TV-GtCLP-TV-HA (SEQ ID N° 55) was synthesized and cloned into the vector pEX-A258 by Eurofins Genomics, generating the vector pAL70.
[0231] After PCR amplification from the vector pAL70 using primers O5'SCL70 (SEQ ID N°56) and O3'SCL70 (SEQ ID N°57), the 1317 bp PCR fragment FPCR-SCL70 (SEQ ID N°58) was cloned into the expression cassette of the gene of interest (goi) present in the chloroplast transformation vector PLE56 linearized with NcoI and SalI using Gibson assembly from New England Biolabs (as per the manufacturer's recommendations) to form the vector pCO86 ( Figure 2 ).
[0232] The chloroplast expression vector pLE56 contains two expression cassettes for expressing genes encoding a selectable marker (gos) and a recombinant protein of interest (goi). The selection cassette contains the selectable marker aadA gene, which encodes an aminoglycoside 3″-adenylyltransferase and confers resistance to spectinomycin and streptomycin. This gene is operably linked at its 5′ end to the C. reinhardtii 16S rRNA promoter (Prrn) fused to the 5′ UTR of atpA and at its 3′ end to the 3′ UTR of the C. reinhardtii rbcL gene. In the second cassette, stable expression of the recombinant goi is controlled by the promoter and the 5′ UTR from C. reinhardtii psbD and the 3′ UTR from C. reinhardtii atpA.
[0233] The two expression cassettes are flanked by left (LHRR) and right (RHRR) endogenous homologous recombination sequences that are identical to the sequences surrounding the targeted integration site of the Chlamydomonas reinhardtii chloroplast genome. The insertion site within the chloroplast genome is typically chosen, for example, to avoid disrupting essential genes or interfering with expression of the polycistronic unit. In a preferred embodiment, the chloroplast transformation vector of the present invention allows for targeted integration of a transgene into the chloroplast genome of Chlamydomonas reinhardtii between the 5S rDNA and psbA genes (and is derived, for example, from sequence GenBank Accession No. NC005352).
[0234] Construction of a chloroplast transformation vector for expression of a chimeric collagen-like domain from the CLP of the cyanobacterium CCMP2712.
[0235] The collagen-like domain from the Cryptophyte CCMP2712 protein (SEQ ID NO 59; designated GtCLD) was engineered to be fused to an amino acid sequence containing a Cys knot and CR4 repeats (SEQ. GPCCGPPGPPGPPGPP, SEQ ID NO 60) at its N-terminus. The Cys knot and CR4 repeats are known in the art to increase the conformational rigidity of a sequence. At its C-terminus, GtCLD was fused to the CR4 repeats, followed by the Cys knot sequence and the folded fiber protein of T4 bacteriophage (SEQ. GYIPEAPRDGQAYVRKDGEWVLLSTFL, SEQ ID NO 61). The resulting recombinant protein was designated Chimeric Collagen-Like Domain from Cryptophyte or GtCCLD (SEQ ID NO 62). The GtCCLD protein was also designed to be fused with a 3xHA epitope tag (referred to as 3HA; SEQ ID N° 63: YPYDVPDYAYPYDVPDYAYPYDVPDYA) at its C-terminus.
[0236] The synthetic codon-optimized gene Gtccld-3ha (SEQ ID N°84) was synthesized by Eurofins Genomics and cloned into the vector pEX-A258 to form pAL81. The gene Gtccld-3ha was subcloned into the goi expression cassette of the chloroplast transformation vector pLE63, which had been previously linearized by digestion with BamHI and PmeI, to generate pCO96. More precisely, Gtccld-3ha was subcloned downstream of the nucleic acid sequence ha-sp-3F (HA-SP-3F) encoding the HA epitope tag (HA) linked to a signal peptide (SP) followed by a 3xFlag epitope tag ( Figure 2 ).
[0237] Then, in the chloroplasts of Chlamydomonas reinhardtii, recombinant GtCCLD-3HA (SEQ ID No. 64) was produced, fused at its N-terminus to the amino acid sequence HA-SP-3F and designated HA-SP-3F-GtCCLD-3HA (SEQ ID No. 83). This recombinant protein is encoded by the nucleic acid sequence designated ha-sp-3f-Gtccld-3ha (SEQ ID No. 19). The chloroplast expression vector pLE63 contains the same selectable marker expression cassette as pLE56 ( Figure 2 Stable expression of goi is controlled by a promoter and the 5' UTR of psbD from Chlamydomonas reinhardtii and the 3' UTR of atpA from Chlamydomonas reinhardtii. In this expression cassette, goi is fused to the nucleic acid sequence ha-sp-3F at its 5' end. pLE63 allows for the same targeted integration of recombinant genes into the Chlamydomonas reinhardtii chloroplast genome as pLE56 and pCO86.
[0238] To remove the nucleic acid sequence encoding the 3xHA tag, a PCR fragment Gtccld-2 (SEQ ID N° 65) containing the nuclear sequence encoding the recombinant chimeric collagen-like domain from Cyanobacterium cyanobacteria GtCCLD was amplified by PCR from pAL81 and cloned into pLE63 linearized with BamHI and PmeI between the psbD promoter / 5'UTR and atpA 3'UTR by the Gibson assembly method to form pCO26 ( Figure 2 As with pCO96, the gene ccld-2 was subcloned downstream of the nucleic acid sequence ha-sp-3F encoding the amino acid sequence HA-SP-3F ( Figure 2 ).
[0239] Then, in the chloroplasts of Chlamydomonas reinhardtii, a recombinant chimeric collagen-like domain from C. cyanobacteria GtCCLD was produced, fused to the amino acid sequence HA-SP-3F at its N-terminus and called HA-SP-3F-GtCCLD (SEQ ID N°85) encoded by the nucleic acid sequence ha-sp-3f-Gtccld (SEQ ID N°20).
[0240] To produce only collagen-like domains from Cryptobium cyanobacteria without Cys knot and CR4 repeats, a PCR fragment containing the nuclear sequence of the gene Gtcld (SEQ ID N° 66) was amplified by PCR from pAL81 and cloned by the Gibson assembly method between the psbD promoter / 5'UTR and atpA 3'UTR into pLE63 linearized with BamHI and PmeI to form pCO28 ( Figure 2 As in pCO96 and pCO26, the gene Gtcld was subcloned downstream of the nucleic acid sequence ha-sp-3F encoding the amino acid sequence HA-SP-3F ( Figure 2 ).
[0241] Then, in the chloroplasts of Chlamydomonas reinhardtii, a recombinant collagen-like domain from C. cyanobacteria GtCLD was produced, fused to the amino acid sequence HA-SP-3F at its N-terminus and called HA-SP-3F-GtCLD (SEQ ID N°86) encoded by the nucleic acid sequence ha-sp-3f-Gtcld (SEQ ID N°21).
[0242] After they are produced in algal chloroplasts, the signal peptide (SP) will be cleaved from the recombinant proteins HA-SP-3F-GtCCLD-3HA, HA-SP-3F-GtCCLD, and HA-SP-3F-GtCLD in vivo during translocation into the thylakoids. Thus, the other three proteins 3F-GtCCLD-3HA, 3F-GtCCLD, and 3F-GtCLD can be produced.
[0243] The 3xFlag tag will be cleaved by in vitro enterokinase digestion of these recombinant proteins.
[0244] Transformation of algae
[0245] The transformation vectors pCO86, pCO96, pCO26 and pCO28 were bombarded in Chlamydomonas reinhardtii cells (137c and CW15) as described in Example 1 .
[0246] To identify stable integration of the recombinant gene encoding the fusion protein into the chloroplast algal genome, colonies were screened for spectinomycin resistance by PCR analysis. For positive PCR screening of the fusion protein gene in CO96, CO26, and CO28 transformants, primers O5'ASTatpA2 (SEQ ID NO: 67) and O3'SUTRpsbD (SEQ ID NO: 68) annealing to the atpA 3'UTR and psbD 5'UTR, respectively, were used. For CO86 transformants, two additional primers were used: O5'SCL70 (SEQ ID NO: 69) and O5'PpsbDC1a2 (SEQ ID NO: 70), annealing to the gene of interest and the psbD promoter, respectively.
[0247] Analysis and Results
[0248] Western blot analysis was performed under reducing conditions on total soluble protein samples extracted from several transformants obtained after transformation with pCO86, pCO96, pCO26 and pCO28 (from strains 137c and CW15).
[0249] The results showed that four recombinant proteins were produced and detected by anti-Flag antibody and / or anti-HA antibody.
[0250] Figure 2 Shown are the results of Western blotting of total soluble protein extracts from CW-CO86 and CW-CO96 transformants producing the recombinant proteins 3F-TV-GtCLP-TV-HA and HA-SP-3F-GtCCLD-3HA.
[0251] To release the recombinant protein of interest from the N-terminal epitope tag fused or not to the signal peptide SP, total soluble protein samples extracted from one clone of the CO96, CO26 and CO28 transformants were digested in vitro by enterokinase digestion as described in Example 1. Figure 4 Shown is a Western blot of enterokinase digested total soluble protein samples extracted from 137c-CO96-4 transformants. The in vivo produced recombinant protein HA-SP-3F-GtCCLD-3HA was cleaved to form GtCCLD-3HA.
[0252] In the case of CO86 transformants producing the recombinant protein 3F-TV-GtCLP-TV-HA, total soluble protein samples were digested in vitro by enterokinase to produce protein TV-GtCLP-TV-HA or by TEV protease to produce protein GtCLP.
[0253] 137c-CO96-4 cells were generated from approximately 900 mL of culture. Algal cells (approximately 3 g) were resuspended and sonicated as described in Example 1. 14.6 mL of total soluble protein extract was obtained.
[0254] Recombinant proteins from 6.5 mL of protein extract of 137c-CO96-4 were purified by affinity chromatography using anti-HA resin (150 μl). The eluted fractions were analyzed by Western blot analysis. As an example, Figure 5 The results shown in reveal the effectiveness of affinity chromatography purification of the chimeric collagen-like domain.
[0255] Western blot analysis performed under non-reducing conditions (without DDT and boiling or without protein samples before loading on polyacrylamide gels) showed that the chimeric collagen-like domain produced in the CO96 transformant strain formed multimeric structures of high apparent molecular weight in vivo, in contrast to the apparent molecular weight of the protein under reducing conditions.
[0256] Example 3
[0257] Chloroplast transformation in Chlamydomonas reinhardtii chloroplasts using aprotinin as a vector for fusion proteins Producing elastin-like peptides or polypeptides or derivatives
[0258] Construction of transformation vectors (pLA01, pLA02, pAL03, and pAL04)
[0259] In the chloroplast transformation vector ELP4, an elastin-like polypeptide consisting of repeated VGVAPG hexapeptides (SEQ ID No. 5), more specifically fourfold repeats of this hexapeptide (SEQ ID No. 81: VGVAPGVGVAPGVGVAPGVGVAPG), is expressed in a fusion protein fused to the C-terminus of the chimeric aprotinin HA-SP-3F-FX-APRO (SEQ ID Nos. 41 and 42). This fusion partner contains aprotinin fused to its N-terminus with an amino acid sequence consisting of an HA epitope tag (HA) followed by a signal peptide (SP), a 3xFlag epitope tag (3F), and a cleavage site for Factor Xa (FX; SEQ ID No. 71: IEGR). A Flag epitope tag sequence (SEQ ID No. 24: DYKDDDDK), representing a cleavage site for enterokinase, is inserted between the chimeric aprotinin and ELP4 to allow for site-specific proteolysis of the fusion protein in vitro by enterokinase.
[0260] After the fusion protein HA-SP-3F-FX-APRO-F-ELP4 (SEQ ID N°87) is produced in algal chloroplasts, the N-terminal fragment HA-SP will be cleaved during the translocation of the protein to the thylakoids, and the following recombinant protein 3F-FX-APRO-F-ELP4 will be produced in vivo.
[0261] As explained in Example 2, in Chlamydomonas reinhardtii, codon usage in nucleic acid sequences encoding proteins of interest has been shown to play a significant role in protein accumulation.
[0262] Nucleic acid sequences encoding aprotinin were designed and optimized to improve their expression in Chlamydomonas reinhardtii host cells as described in Example 2. Following optimization, the gene encoding aprotinin (APRO) was operably fused at its 5' end to a codon-optimized nucleic acid sequence encoding an HA epitope tag (HA), followed by a signal peptide, a 3xFlag epitope tag (3F), and a cleavage site recognized by Factor Xa protease (FX), to form the chimeric aprotinin gene ha-sp-3f-fx-apro (SEQ ID N° 42).
[0263] Using the same method described in Example 1 and the codon usage of the Chlamydomonas reinhardtii chloroplast genome, the nucleic acid sequence encoding ELP4 was first codon-optimized. The resulting sequence was used to design two overlapping oligomers, O5'Gibs-ELP4 (SEQ ID NO 72) and O3'Gibs-ELP4 (SEQ ID NO 73), which were used as primers and templates to amplify the 194 bp fragment FGibs-ELP4 by PCR. This amplified DNA was cloned into the chloroplast transformation vector pAU76 using Gibson Assembly Master Mix from New England Biolabs (as recommended by the manufacturer) and linearized with PmeI to form the vector pLA00.
[0264] The expression vector pAU76 used for chloroplast genome transformation contains two expression cassettes for the genes encoding a selectable marker (same as the previous transformation vector) and the chimeric aprotinin HA-SP-3F-FX-APRO. pAU76 allows the same targeted integration of recombinant genes into the Chlamydomonas reinhardtii chloroplast genome as pLE63.
[0265] The nucleic acid sequence encoding ELP4 was amplified from pLA00 by PCR using primers 05'Gibs01BE (SEQ ID N°74) and 03'Gibs01BE (SEQ ID N°75). The 359 pb PCR fragment FPCR-AP-FELP4 (SEQ ID N°76) was cloned into pLE63 linearized with BamHI and PmeI using Gibson Assembly Master Mix to form the vector pLA01. Transformation of the vector pLA01 allowed the production of the fusion protein HA-SP-3F-FX-APRO-F-ELP4, which contains ELP4 (SEQ ID N°76) linked to the chimeric aprotinin HA-SP-3F-FX-APRO at its N-terminus and subsequently to the 1X Flag tag. Figure 6 ).
[0266] The chloroplast transformation vector pLA02 was obtained by cloning into pLE63 (linearized with BamHI and PmeI) using the Gibson assembly method. The PCR fragment FPCR-FELP4-HA (SEQ ID No. 77) (359 pb) was amplified from pLA00 using primers 05'Gibs02BE (SEQ ID No. 78) and 03'Gibs02BE (SEQ ID No. 79). The transformation vector pLA02 allows the production of the fusion protein HA-SP-3F-1F-ELP4, which contains ELP4 (SEQ ID No. 77) linked at its N-terminus to the chimeric sequence HA-SP-3F followed by a 1X Flag tag. Figure 6 ).
[0267] In the case of algal chloroplasts transformed with pLA01 or pLA02, two other different proteins, 3F-FX-APRO-1F-ELP4 or 3F-1F-ELP4 (SEQ ID N° 88), can be produced in vivo if the signal peptide SP is cleaved after translocation of the fusion protein into the thylakoid lumen.
[0268] In both types of transfer vectors, the release of ELP4 from the fusion protein can be performed in vitro by enterokinase digestion, which cleaves the protein sequence just after the second lysine amino acid in the motif sequence DYKDDDDK (SEQ ID N° 24) present in the 1X Flag tag upstream of ELP4.
[0269] The elastin-like polypeptide called ELPE4 consists of repeated GVAPGE (SEQ ID No. 9), a derivative of the peptide VGVAPG, more specifically a 4-fold repeat of this peptide (SEQ ID No. 80, VGVAPGEVGVAPGEVGVAPGEVGVAPGE). ELPE4 is also expressed in a fusion protein in a chloroplast transformation vector, where it is fused to the C-terminus of the chimeric aprotinin HA-SP-3F-FX-APRO.
[0270] For cleavage from the vector in vitro by ELPE4, the flexible linker LGM (SEQ ID N°50: RSGGGGSSGGGGGGSSRS) was added followed by TEV protease (TV; SEQ ID N°52: ENLYFQG) or enterokinase (EK; SEQ ID N°38: DDDDK).
[0271] Two types of fusion proteins have been produced from two different chloroplast expression vectors: HA-SP-3F-FX-APRO-LGM-TV-ELPE4 (SEQ ID N°89) or HA-SP-3F-FX-APRO-LGM-EK-ELPE4 (SEQ ID N°90).
[0272] The nucleic acid sequences encoding LGM-TV-ELPE4 or LGM-EK-ELPE4 were also codon-optimized using the same method described in Example 1 and the codon usage of the Chlamydomonas reinhardtii chloroplast genome. After codon optimization, different synthetic genes, lgm-tv-elpe4 (SEQ ID N°40) and lgm-ek-elpe4 (SEQ ID N°48), were synthesized by Eurofins. These optimized genes were cloned into an expression cassette (SEQ N°82) in the chloroplast transformation vector pAU76 linearized with PmeI by the Gibson assembly method downstream of the gene encoding the vector, generating pLA03 and pLA04, respectively.
[0273] Transformation of algae
[0274] The transformation vectors pAL01, pLA02, pLA03 and pLA04 were bombarded in Chlamydomonas reinhardtii cells (137c and CW15) as described in Example 1.
[0275] To identify recombinant genes encoding fusion proteins stably integrated into the chloroplast algal genome, spectinomycin-resistant colonies were screened by PCR analysis using primers O5'ASTatpA2 (SEQ ID N°67) and O3'SUTRpsbD (SEQ ID N°54) annealing in atpA 3'UTR and psbD 5'UTR, respectively.
[0276] Analysis and Results
[0277] Western blot analysis of total soluble proteins extracted from different independent strains of LA01, LA03 or LA04 transformants using anti-Flag antibody revealed the production of the fusion proteins HA-SP-3F-FX-APRO-F-ELP4 (SEQ ID N°87), HA-SP-3F-FX-APRO-LGM-TV-ELPE4 (SEQ ID N°89) and HA-SP-3F-FX-APRO-LGM-EK-ELPE4 (SEQ ID N°91) in the chloroplasts of C. reinhardtii.
[0278] like Figure 7 As shown in Figure 2, Western blot analysis showed that the ELP4 polypeptide fused to 3F-FX-APR was produced very well in LA01 transformants using an anti-Flag antibody. In addition, the HA epitope tag and signal peptide appeared to be cleaved in all LA01 transformants, as Western blots of the same total soluble protein extracts showed that the primary anti-HA antibody did not recognize any of the fusion proteins ( Figure 7 ).
[0279] In the LA02 transformant, no recombinant protein was detected ( Figure 7 ), indicating that fusion of ELP4 to 3F-FX-APRO (or aprotinin as a general approach) allows accumulation of ELP4.
[0280] Biomass of one transformant, CW-LA01, was generated. The cell pellet was resuspended in sonication buffer.
[0281] The fusion protein was purified by anti-Flag M2 affinity chromatography. The eluted fractions containing the fusion protein were identified by Western blot analysis, dialyzed, and concentrated. Enterokinase protease digestion was performed, followed by size exclusion chromatography (HiLoad 26 / 00 Superdex 30) to allow purification of the ELP4 polypeptide.
[0282] The same method was used for the purification of ELPE4. The fusion protein was purified by affinity chromatography. The eluted fractions containing the fusion protein were identified by Western blot analysis, dialyzed, and concentrated. Depending on the transformant, enterokinase or TEV protease digestion was performed, followed by size exclusion chromatography (HiLoad 26 / 00 Superdex 30) to allow purification of the ELPE4 polypeptide.
[0283] In the case of the LA03 transformant, and after TEV protease digestion, the released polypeptide is GVGVAPGEVGVAPGEVGVAPGEVGVAPGE (SEQ ID N° 53).
[0284] To cleave the polypeptide ELPE4 into the peptide VGVAPGE by intracellular proteases, the SEC elution fraction was evaporated and dialyzed using dialysis tubing with a 1 kDa cutoff to remove salts and exchange the buffer.
[0285] After digestion with Glu-C_ intracellular protease of the dialyzed samples, the released peptides were purified by size exclusion chromatography as described in Example 1. Sequence Listing <110> Algae and Cell Company <120> Recombinant microalgae capable of producing peptides, polypeptides or proteins of collagen, elastin and their derivatives in the chloroplasts of microalgae and related methods <130> BEX21L0054 <150> EP20305210.5 <151> 2020-02-28 <160> 91 <170> PatentIn Version 3.5 <210> 1 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Elastin Pentapeptide <400> 1 Lys Gly Gly Val Gly 1 5 <210> 2 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Elastin Pentapeptide <400> 2 Val Gly Gly Val Gly 1 5 <210> 3 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Elastin Pentapeptide <400> 3 Gly Val Gly Val Pro 1 5 <210> 4 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Elastin Pentapeptide <220> <221> MISC_FEATURE <222> (4)..(4) <223> X is V, I, or K <400> 4 Val Pro Gly Xaa Gly 1 5 <210> 5 <211> 6 <212> PRT <213> Artificial sequence <220> <223> Elastin Hexapeptide <400> 5 Val Gly Val Ala Pro Gly 1 5 <210> 6 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Elastin heptapeptide <400> 6 Leu Gly Ala Gly Gly Ala Gly 1 5 <210> 7 <211> 9 <212> PRT <213> Artificial sequence <220> <223> Elastin Nonapeptide <400> 7 Leu Gly Ala Gly Gly Ala Gly Val Leu 1 5 <210> 8 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Elastin derivatives <400> 8 Val Gly Val Ala Pro Gly Asp 1 5 <210> 9 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Elastin derivatives <400> 9 Val Gly Val Ala Pro Gly Glu 1 5 <210> 10 <211> 1251 <212> DNA <213> Artificial sequence <220> <223> 3f-tv-Gtclp-tv-ha <400> 10 atggcagatt ataaagacga tgatgataaa gattacaaag acgatgacga caaagactat 60 aaagatgatg acgataaaga aaacctttat tttcaaggta gtcgtcgtcg tcaaggtgct 120 gctgccgcag ccgctggtgc agcaggatta atcttacttg tgttatactttc attagtatct 180 ttacgtgctc atgaacgcaa accaacatct ttagttgaag gttcacgcac tacaatgatg 240 ccaacatgga aagtaactgg tacattaaat ccagttgaag aaacagagaa agaagtagaa 300 gaagaaccat cttcacctcc accaccagca gtacaagcta caccagatgt acctcaagag 360 caagaagccg aaccaagtca gtcagaatgg gttccaccac ctaaaccttg ggaaccatta 420 gctgaaaaaa atattggtcg tttagaacgt actaatcgtg aagtagttcg tggtgatatg 480 gaaaatgcaa acatgataaa aaaattacgt gctgctatgc gtttataaa gaaacaattt 540 gcagctaaaa ttgtgaattt agaacaagct tatgatcgtc gtttagctcg tgaaagtcgt 600 attttacacc aacgtattaa tacagtagct atgcaacctg gtccaacagg catgtcaggt 660 cctcaaggtt taccaggtcc tactggtcca gcaggccaaa atggttctcc aggctcacca 720 ggagcaatgg gtccacaagg accaatgggt ccaagaggtt atcgtggttt acgtggtcca 780 cctggtgatc aaggtcgtcc tggtgataca ggtagacctg gtgctcctgg tattccaggt 840 caaattggtc aacgtggacc tattggtcca atgggtgttc aaggccctga aggtccacgt 900 ggtaatcctg gattaccagg tccaccaggt gcaccaggtg ttccaggtat tggtatcaa 960 ggaccagcag gtcctcctgg tccaccagga cgctttcctg ctacagaaca ttgtagctat 1020 gttactggtg aatgtgtaat taacggcaat gacattttcc atgctatgcg tgaaactggc 1080 gatgtaagtt gtccaagaaa ctactatgtt aaaggtgttg attacatcaa atgcaataat 1140 ggtgctgaat accacaatac tttacagtta cgtttaacat gttgtttatt aggtttaaac 1200 gagaatcttt atttccaagg ttatccttat gatgttccag attatgctta a 1251 <210> 11 <211> 18 <212> DNA <213> Artificial sequence <220> <223> Nuclear sequence encoding the elastin peptide VGVAPG <400> 11 gtaggtgtag ctcctggt 18 <210> 12 <211> 18 <212> DNA <213> Artificial sequence <220> <223> Nuclear sequence encoding the elastin peptide VGVAPG <400> 12 gttggtgttg ctcctgga 18 <210> 13 <211> 18 <212> DNA <213> Artificial sequence <220> <223> Nuclear sequence encoding the elastin peptide VGVAPG <400> 13 gtaggtgttg ctccaggt 18 <210> 14 <211> 18 <212> DNA <213> Artificial sequence <220> <223> Nuclear sequence encoding the elastin peptide VGVAPG <400> 14 gtgggtgtag ctcctggt 18 <210> 15 <211> twenty one <212> DNA <213> Artificial sequence <220> <223> Nuclear sequence encoding the elastin peptide derivative VGVAPGE <400> 15 gtaggtgtag ctcctggtga a 21 <210> 16 <211> twenty one <212> DNA <213> Artificial sequence <220> <223> Nuclear sequence encoding the elastin peptide derivative VGVAPGE <400> 16 gttggtgttg ctcctggaga a 21 <210> 17 <211> twenty one <212> DNA <213> Artificial sequence <220> <223> Nuclear sequence encoding the elastin peptide derivative VGVAPGE <400> 17 gtaggtgttg ctccaggtga a 21 <210> 18 <211> twenty one <212> DNA <213> Artificial sequence <220> <223> Nuclear sequence encoding the elastin peptide derivative VGVAPGE <400> 18 gtgggtgtag ctcctggtga a 21 <210> 19 <211> 852 <212> DNA <213> Artificial sequence <220> <223> ha-sp-3f-Gtccld-3ha <400> 19 atggcatatc cttatgatgt accagattat gccaataata acgacttatt tcaagcatca 60 cgtcgtcgtt tcttagctca attaggaggt ttaacagttg ctggtatgtt aggtccatct 120 ttacttactc cacgtagagc tacagcagct caagctgatt acaaagatga tgacgataaa 180 gactataaag atgatgatga caaagattac aaagacgatg atgataaagg atccggtcca 240 tgttgtggac cacctggtcc tccaggtcca ccaggaccac ctggaccaac aggtatgagt 300 ggaccacaag gtttaccagg tcctactggt cctgctggtc aaaacggttc accaggttct 360 ccaggtgcaa tgggtccaca aggtcctatg ggaccacgtg gttatcgtgg tttacgtggt 420 ccaccaggtg atcaaggtcg tcctggtgat acaggtcgtc caggtgctcc aggtattcca 480 ggtcagattg gtcaacgtgg tccaattggt cctatgggtg ttcaaggacc agaaggtcca 540 agaggtaatc caggccttcc aggaccacca ggtgcaccag gtgtacctgg tattggcatt 600 caaggtcctg ctggcccacc tggtccacct ggtcgttttg gtcctcctgg ccctccaggc 660 ccaccaggtc ctcctggtcc atgttgcggc tatatcccag aagctccacg tgatggtcaa 720 gcttatgttc gcaaagatgg tgaatgggtg ttattatcaa cattcttata cccttatgac 780 gtaccagatt acgcatatcc ttatgatgtt ccagattatg cctatccata cgatgtacct 840 gactatgctt aa 852 <210> 20 <211> 771 <212> DNA <213> Artificial sequence <220> <223> ha-sp-3f-Gtccld <400> 20 atggcatatc cttatgatgt accagattat gccaataata acgacttatt tcaagcatca 60 cgtcgtcgtt tcttagctca attaggaggt ttaacagttg ctggtatgtt aggtccatct 120 ttacttactc cacgtagagc tacagcagct caagctgatt acaaagatga tgacgataaa 180ccaccaggtg atcaaggtcg tcctggtgat acaggtcgtc caggtgctcc aggtattcca 480 ggtcagattg gtcaacgtgg tccaattggt cctatgggtg ttcaaggacc agaaggtcca 540 agaggtaatc caggccttcc aggaccacca ggtgcaccag gtgtacctgg tattggcatt 600 caaggtcctg ctggcccacc tggtccacct ggtcgttttg gtcctcctgg ccctccaggc 660 ccaccaggtc ctcctggtcc atgttgcggc tatatcccag aagctccacg tgatggtcaa 720 gcttatgttc gcaaagatgg tgaatgggtg ttattatcaa cattcttata a 771 <210> 21 <211> 594 <212> DNA <213> Artificial sequence <220> <223> ha-sp-3f-Gtcld <400> 21 atggcatatc cttatgatgt accagattat gccaataata acgacttatt tcaagcatca 60 cgtcgtcgtt tcttagctca attaggaggt ttaacagttg ctggtatgtt aggtccatct 120 ttacttactc cacgtagagc tacagcagct caagctgatt acaaagatga tgacgataaa 180 [[ID=3】2]]gactataaag atgatgatga caaagattac aaagacgatg atgataaagg atccggacca 240 acaggtatga gtggaccaca aggtttacca ggtcctactg gtcctgctgg tcaaaacggt 300 tcaccaggtt ctccaggtgc aatgggtcca caaggtccta tgggaccacg tggttatcgt 360 ggtttacgtg gtccaccagg tgatcaaggt cgtcctggtg atacaggtcg tccaggtgct 420 ccaggtattc caggtcagat tggtcaacgt ggtccaattg gtcctatggg tgttcaagga 480 ccagaaggtc caagaggtaa tccaggcctt ccaggaccac caggtgcacc aggtgtacct 540 ggtattggca ttcaaggtcc tgctggccca cctggtccac ctggtcgttt ttaa 594 <210> 22 <211> 72 <212> DNA <213> Artificial sequence <220> <223> elp4 <400> 22 gtaggtgtag ctcctggtgt tggtgttgct cctggagtag gtgttgctcc aggtgtgggt 60 gtagctcctg gt 72 <210> 23 <211> 84 <212> DNA <213> Artificial sequence <220> <223> elpe4 <400> 23 gtaggtgtag ctcctggtga agttggtgtt gctcctggag aagtaggtgt tgctccaggt 60 gaagtgggtg tagctcctgg tgaa 84 <210> twenty four <211> 8 <212> PRT <213> Artificial sequence <220> <223> Flag Tags <400> twenty four Asp Tyr Lys Asp Asp Asp Asp Lys 1 5 <210> 25 <211> twenty four <212> PRT <213> Artificial sequence <220> <223> 3X Flag Tags <400> 25 Asp Tyr Lys Asp Asp Asp Asp Lys Asp Tyr Lys Asp Asp Asp Asp Lys 1 5 10 15 Asp Tyr Lys Asp Asp Asp Asp Lys 20 <210> 26 <211> 9 <212> PRT <213> Artificial sequence <220> <223> HA tag <400> 26 Tyr Pro Tyr Asp Val Pro Asp Tyr Ala 1 5 <210> 27 <211> 27 <212> PRT <213> Artificial sequence <220> <223> 3xHA tag <400> 27 Tyr Pro Tyr Asp Val Pro Asp Tyr Ala Tyr Pro Tyr Asp Val Pro Asp 1 5 10 15 Tyr Ala Tyr Pro Tyr Asp Val Pro Asp Tyr Ala 20 25 <210> 28 <211> 6 <212> PRT <213> Artificial sequence <220> <223> His label <400> 28 His His His His His His 1 5 <210> 29 <211> 424 <212> DNA <213> Artificial sequence <220> <223> PpsbD-5'UTR <400> 29 atgaaattaa atggatattt ggtacattta attccacaaa aatgtccaat acttaaaata 60 caaaattaaa agtattagtt gtaaacttga ctaacatttt aaattttaaa ttttttccta 120 attatatatt ttacttgcaa aatttataaa aattttatgc atttttat cataataata 180 aaacctttat tcatggttta taatataata attgtgatga ctatgcacaa agcagttcta 240 gtcccatata tataactata tataacccgt ttaaagattt atttaaaaaat atgtgtgtaa 300 aaaatgctta tttttaattt tattttat aagttataat attaaataca caatgattaa 360 aattaaataa taataaattt aacgtaacga tgagttgttt ttttattttg gagatacacg 420 cacc 424 <210> 30 <211> 1025 <212> DNA <213> Artificial sequence <220> <223> Prrn-5'UTRatpA <400> 30 tttttaatta agtaggaact cggtatatgc tcttttgggg tcttattagc tagtattagt 60 taactaacaa aagatcaata ttttagtttg ttttatatat tttattactt aagtagtaag 120 gatttgcatt tagcaatctt aaatacttaa gtaataatct ataaataaaa tatattttcg 180 ctttaaaact tataaaaatt atttgctcgt tataagccta aaaaaacgta ggatctctac 240 gagatattac attgtttttt tctttaattg gctttaatat tactttgtat atataaacca 300 aagtacttgt taatagttat taaattatat taactataca gtacaaagaa attttttgct 360 aaaaaaagta tgttaacatt aaaaattttt gtttatacag actagtcaga ctcggggggc 420 aggcaacaaa tttatttatt gtcccgtaag gggaagggga aaacaattat tattttactg 480 cggagcagct tgttattgaa attttattaa aaaaaaaata aaaatttgac aaaaaaaaat 540 aaaaaagtta aattaaaaac actgggaatg ttctacatca taaaaatcaa aagggtttaa 600 aatcccgaca aaatttaaac tttaaagagt ggcgcctacc ttttttttaa tttgcatgat 660 tttaatgctt atgctatctt ttttatttag tccataaaac ctttaaagga ccttttctta 720 tgggatattt atattttcct aacaaagcaa tcggcgtcat aaactttagt tgcttacgac 780 gcctgtggac gtccccccct tccccttacg ggcaagtaaa cttagggatt ttaatgcaat 840 aaataaattt gtcctcttcg ggcaaatgaa ttttagtatt taaatatgac aagggtgaac 900 cattactttt gttaacaagt gatcttacca ctcactattt ttgttgaatt ttaaacttat 960 ttaaaattct cgagaaagat tttaaaaata aactttttta atcttttatt tattttttct 1020 ttttt 1025 <210> 31 <211> 400 <212> DNA <213> Artificial sequence <220> <223> 3'UTRatpA <400> 31 tttttaatta agtaggaact cggtatatgc tcttttgggg tcttattagc tagtattagt 60 taactaacaa aagatcaata ttttagtttg ttttatatat tttattactt aagtagtaag 120 gatttgcatt tagcaatctt aaatacttaa gtaataatct ataaataaaa tatattttcg 180 ctttaaaact tataaaaatt atttgctcgt tataagccta aaaaaacgta ggatctctac 240 gagatattac attgtttttt tctttaattg gctttaatat tactttgtat atataaacca 300 aagtacttgt taatagttat taaattatat taactataca gtacaaagaa attttttgct 360 aaaaaaagta tgttaacatt aaaaattttt gtttatacag 400 <210> 32 <211> 435 <212> DNA <213> Artificial sequence <220> <223> 3'UTRrbcL <400> 32 aagcttgtac tcaagctcgt aacgaaggtc gtgaccttgc tcgtgaaggt ggcgacgtaa 60 ttcgttcagc ttgtaaatgg tctccagaac ttgctgctgc atgtgaagtt tggaaagaaa 120 ttaaattcga atttgatact attgacaaac tttaattttt atttttcatg atgtttatgt 180 gaatagcata aacatcgttt ttatttttta tggtgtttag gttaaatacc taaacatcat 240 tttacatttt taaaattaag ttctaaagtt atcttttgtt taaatttgcc tgtgctttat 300 aaattacgat gtgccagaaa aataaatct tagcttttta ttatagaatt tatctttatg fathers fathers aaaagaata gtaacatact aaagcggatg taactcaatc ggtagtgc gatcc 435 <210> 33 <211> 41 <212> PRT <213> The snowstorm <220> <223> SP <400> 33 Asn Asn Asn Asp There Is Nothing Like You 1 5 10 15 Leu Gly Gly Leu Thr Val Ala Gly Met Leu Gly Pro Ser Leu Leu Thr 20 25 30 Pro Arg Arg God Thr God God Gln God 35 40 <210> 34 <211> 123 <212> DNA <213> The snowstorm <220> <223> Buy the SP bubble <400> 34 grandfather acttatttca agcatcacgt cgtcgtttct tagctcaatt aggaggttta acagttgctg gtatgttagg tccatcttta cttactccac gtaggctac agcagctcaa gct 123 <210> 35 <211> 58 <212> PRT <213> Cattle (Bos taurus) <400> 35 Arg Pro Asp Phe Cys Leu Glu Pro Pro Tyr Thr Gly Pro Cys Lys Ala 1 5 10 15 Arg Ile Ile Arg Tyr Phe Tyr Asn Ala Lys Ala Gly Leu Cys Gln Thr 20 25 30 Phe Val Tyr Gly Gly Cys Arg Ala Lys Arg Asn Asn Phe Lys Ser Ala 35 40 45 Glu Asp Cys Met Arg Thr Cys Gly Gly Ala 50 55 <210> 36 <211> 174 <212> DNA <21x> Cattle (Bos taurus) <400> 36 cggcctgact tctgcctaga gcctccatat acgggtccct gcaaggccag aattatcaga 60 tacttctaca acgccaaggc tgggctctgc cagacctttg tatatggcgg ctgcagagct 120 aaaagaaaca atttcaagag cgcagaggac tgcatgagga cctgtggtgg tgct 174 <210> 37 <211> 67 <212> PRT <213> Artificial sequence <220> <223> HA-APRO <400> 37 Tyr Pro Tyr Asp Val Pro Asp Tyr Ala Arg Pro Asp Phe Cys Leu Glu 1 5 10 15 Pro Pro Tyr Thr Gly Pro Cys Lys Ala Arg Ile Ile Arg Tyr Phe Tyr 20 25 30 Asn Ala Lys Ala Gly Leu Cys Gln Thr Phe Val Tyr Gly Gly Cys Arg 35 40 45 Ala Lys Arg Asn Asn Phe Lys Ser Ala Glu Asp Cys Met Arg Thr Cys 50 55 60 Gly Gly Ala 65 <210> 38 <211> 5 <212> PRT <213> Artificial sequence <220> <223> EK <400> 38 Asp Asp Asp Asp Lys 1 5 <210> 39 <211> 85 <212> PRT <213> Artificial sequence <220> <223> 3F-APRO <400> 39 Ala Gln Ala Asp Tyr Lys Asp Asp Asp Asp Lys Asp Tyr Lys Asp Asp 1 5 10 15 Asp Asp Lys Asp Tyr Lys Asp Asp Asp Asp Lys Arg Pro Asp Phe Cys[[ID=�6]] 20 25 30 Leu Glu Pro Pro Tyr Thr Gly Pro Cys Lys Ala Arg Ile Ile Arg Tyr 35 40 45 Phe Tyr Asn Ala Lys Ala Gly Leu Cys Gln Thr Phe Val Tyr Gly Gly 50 55 60 Cys Arg Ala Lys Arg Asn Asn Phe Lys Ser Ala Glu Asp Cys Met Arg 65 70 75 80 Thr Cys Gly Gly Ala 85 <210> 40 <211> 234 <212> DNA <213> Artificial sequence <220> <223> lgm-tv-elpe4 <400> 40 cagctgaaga ttgtatgcgt acttgtggtg gtgctagatc tggtggaggc ggttcaagtg 60 gtggaggagg tggcggatct tcaagatctg aaaatttata ttttcaaggt gtaggtgtag 120 ctcctggtga agttggtgtt gctcctggag aagtaggtgt tgctccaggt gaagtgggtg 180 tagctcctgg tgaataagtt taaaccaggt gacctgcaga gctagctact gcag 234 <210> 41 <211> 140 <212> PRT <213> Artificial sequence <220> <223> HA-SP-3F-FX-APRO <400> 41 Met Ala Tyr Pro Tyr Asp Val Pro Asp Tyr Ala Asn Asn Asn Asp Leu 1 5 10 15 Phe Gln Ala Ser Arg Arg Arg Phe Leu Ala Gln Leu Gly Gly Leu Thr 20 25 30 Val Ala Gly Met Leu Gly Pro Ser Leu Leu Thr Pro Arg Arg Ala Thr 35 40 45 Ala Ala Gln Ala Asp Tyr Lys Asp Asp Asp Asp Lys Asp Tyr Lys Asp 50 55 60 Asp Asp Asp Lys Asp Tyr Lys Asp Asp Asp Asp Lys Gly Ser Ile Glu 65 70 75 80 Gly Arg Arg Pro Asp Phe Cys Leu Glu Pro Pro Tyr Thr Gly Pro Cys 85 90 95 Lys Ala Arg Ile Ile Arg Tyr Phe Tyr Asn Ala Lys Ala Gly Leu Cys 100 105 110 Gln Thr Phe Val Tyr Gly Gly Cys Arg Ala Lys Arg Asn Asn Phe Lys 115 120 125 Ser Ala Glu Asp Cys Met Arg Thr Cys Gly Gly Ala 130 135 140 <210> 42 <211> 423 <212> DNA <213> Synthetic sequence <220> <223> Nuclear sequence encoding HA-SP-3F-FX-APRO <400> 42 atggcatatc cttatgatgt accagattat gccaataata acgacttatt tcaagcatca 60 cgtcgtcgtt tcttagctca attaggaggt ttaacagttg ctggtatgtt aggtccatct 120 ttacttactc cacgtagagc tacagcagct caagctgatt acaaagatga tgacgataaa 180 gactataaag atgatgatga caaagattac aaagacgatg atgataaagg atccattgaa 240 ggtcgtagac cagacttttg cttagaacca ccatatacag gtccatgtaa agctcgtatc 300 attcgctatt tctacaatgc aaaagcagga ttatgtcaaa catttgtgta tggaggttgt 360 cgtgctaaac gtaacaactt taaatcagct gaagattgta tgcgtacttg tggtggtgct 420 taa 423 <210> 43 <211> 134 <212> PRT <213> Artificial Sequence3] <220> <223> HA-SP-3F-APRO <400> 43 Met Ala Tyr Pro Tyr Asp Val Pro Asp Tyr Ala Asn Asn Asn Asp Leu 1 5 10 15 Phe Gln Ala Ser Arg Arg Arg Phe Leu Ala Gln Leu Gly Gly Leu Thr 20 25 30 Val Ala Gly Met Leu Gly Pro Ser Leu Leu Thr Pro Arg Arg Ala Thr 35 40 45 Ala Ala Gln Ala Asp Tyr Lys Asp Asp Asp Asp Lys Asp Tyr Lys Asp 50 55 60 Asp Asp Asp Lys Asp Tyr Lys Asp Asp Asp Asp Lys Arg Pro Asp Phe 65 70 75 80 Cys Leu Glu Pro Pro Tyr Thr Gly Pro Cys Lys Ala Arg Ile Ile Arg 85 90 95 Tyr Phe Tyr Asn Ala Lys Ala Gly Leu Cys Gln Thr Phe Val Tyr Gly 100 105 110 Gly Cys Arg Ala Lys Arg Asn Asn Phe Lys Ser Ala Glu Asp Cys Met 115 120 125 Arg Thr Cys Gly Gly Ala 130 <210> 44 <211> 349 <212> DNA <213> Artificial sequence <220> <223> Nuclear sequence encoding HA-SP-3F-APRO <400> 44 atcacgtcgt cgtttcttag ctcaattagg aggtttaaca gttgctggta tgttaggtcc 60 atctttactt actccacgta gagctacagc agctcaagct gattacaaag atgatgacga 120 taaagactat aaagatgatg atgacaaaga ttacaaagac gatgatgata aaagaccaga 240. cttttgctta gaaccaccat atacaggtcc atgtaaagct cgtatcattc gctatttcta caatgcaaaa gcaggattat gtcaaacatt tgtgtatgga ggttgtcgtg ctaaacgtaa caactttaaa tcagctgaag attgtatgcg tacttgtggt ggtgcttaa <210> 45 <211> 108 <212> PRT <213> The snowstorm <220> <223> AND-SP-APRO <400> 45 Thyr Pro Tyr Asp Val Pro Asp Tyr Ala Asn Asn Asn Asp Leu Phe Gln 1 5 10 15 Click Download to save Ala Ser Arg Arg Arg Phe Leu Ala Gln Leu Gly Gly Leu Thr Val Ala mp3 youtube com 20 25 30 Gly Met Leu Gly Pro Ser Leu Leu Thr Pro Arg Arg Free Mp3 Download 35 40 45 Gln Ala Arg Pro Asp Phe Cys Leu Glu Pro Tyr Thr Gly Pro Cys 50 55 60 Lys Ala Arg Ile Ile Arg Tyr Phe Tyr Asn Ala Lys Ala Gly Leu Cys 65 70 75 80 Gln Thr Phe Val Tyr Gly Gly Cys Arg Ala Lys Arg Asn Asn Phe Lys 85 90 95 Ser Ala Glu Asp Cys Met Arg Thr Cys Gly Gly Ala 100 105 <210> 46 <211> 324 <212> DNA <213> Artificial sequence <220> <223> Nuclear sequence encoding HA-SP-APRO <400> 46 tatccttatg atgtaccaga ttatgccaat aataacgact tatttcaagc atcacgtcgt 60 cgtttcttag ctcaattagg aggtttaaca gttgctggta tgttaggtcc atctttactt 120 actccacgta gagctacagc agctcaagct agaccagact tttgcttaga accaccatat 180 acaggtccat gtaaagctcg tatcattcgc tatttctaca atgcaaaagc aggattatgt 240 caaacatttg tgtatggagg ttgtcgtgct aaacgtaaca actttaaatc agctgaagat 300 tgtatgcgta cttgtggtgg tgct 324 <210> 47 <211> 91 <212> PRT <213> Artificial sequence <220> <223> 3F-FX-APRO <400> 47 Ala Gln Ala Asp Tyr Lys Asp Asp Asp Asp Lys Asp Tyr Lys Asp Asp 1 5 10 15 Asp Asp Lys Asp Tyr Lys Asp Asp Asp Asp Lys Gly Ser Ile Glu Gly 20 25 30 Arg Arg Pro Asp Phe Cys Leu Glu Pro Pro Tyr Thr Gly Pro Cys Lys 35 40 45 Ala Arg Ile Ile Arg Tyr Phe Tyr Asn Ala Lys Ala Gly Leu Cys Gln 50 55 60 Thr Phe Val Tyr Gly Gly Cys Arg Ala Lys Arg Asn Asn Phe Lys Ser 65 70 75 80 Ala Glu Asp Cys Met Arg Thr Cys Gly Gly Ala 85 90 <210> 48 <211> 231 <212> DNA <213> Artificial sequence <220> <223> lgm-ek-elpe4 <400> 48 cagctgaaga ttgtatgcgt acttgtggtg gtgctagatc tggtggaggc ggttcaagtg 60 gtggaggagg tggcggatct tcaagatctg acgatgacga caagggtgta ggtgtagctc 120 ctggtgaagt tggtgttgct cctggagaag taggtgttgc tccaggtgaa gtgggtgtag 180 ctcctggtga ataagtttaa accaggtgac ctgcagagct agctactgca g 231 <210> 49 <211> 14 <212> PRT <213> Artificial Sequence <220> <223> LG <400> 49 Arg Ser Gly Gly Gly Gly Ser Gly Gly Gly Gly Ser Gly Ser 1 5 10 <210> 50 <211> 18 <212> PRT <213> Artificial Sequence <220> <223> LGM <400> 50 Arg Ser Gly Gly Gly Gly Ser Ser Gly Gly Gly Gly Gly Gly Ser Ser 1 5 10 15 Arg Ser <210> 51 <211> 365 <212> PRT <213> Artificial Sequence <220> <223> CCMP2712 <400> 51 Met Ser Arg Arg Arg Gln Gly Ala Ala Ala Ala Ala Ala Gly Ala Ala[[ID=�5]] 1 5 10 15 Gly Leu Ile Leu Leu Val Leu Leu Ser Leu Val Ser Leu Arg Ala His 20 25 30 Glu Arg Lys Pro Thr Ser Leu Val Glu Gly Ser Arg Thr Thr Met Met 35 40 45 Pro Thr Trp Lys Val Thr Gly Thr Leu Asn Pro Val Glu Glu Thr Glu 50 55 60 Lys Glu Val Glu Glu Glu Pro Ser Ser Pro Pro Pro Pro Ala Val Gln 65 70 75 80 Ala Thr Pro Asp Val Pro Gln Glu Gln Glu Ala Glu Pro Ser Gln Ser 85 90 95 Glu Trp Val Pro Pro Pro Lys Pro Trp Glu Pro Leu Ala Glu Lys Asn 100 105 110 Ile Gly Arg Leu Glu Arg Thr Asn Arg Glu Val Val Arg Gly Asp Met 115 120 125 Glu Asn Ala Asn Met Ile Lys Lys Leu Arg Ala Ala Met Arg Leu Leu 130 135 140 Lys Lys Gln Phe Ala Ala Lys Ile Val Asn Leu Glu Gln Ala Tyr Asp 145 150 155 160 Arg Arg Leu Ala Arg Glu Ser Arg Ile Leu His Gln Arg Ile Asn Thr 165 170 175 Val Ala Met Gln Pro Gly Pro Thr Gly Met Ser Gly Pro Gln Gly Leu 180 185 190 Pro Gly Pro Thr Gly Pro Ala Gly Gln Asn Gly Ser Pro Gly Ser Pro 195 200 205 Gly Ala Met Gly Pro Gln Gly Pro Met Gly Pro Arg Gly Tyr Arg Gly 210 215 220 Leu Arg Gly Pro Pro Gly Asp Gln Gly Arg Pro Gly Asp Thr Gly Arg 225 230 235 240 Pro Gly Ala Pro Gly Ile Pro Gly Gln Ile Gly Gln Arg Gly Pro Ile 245 250 255 Gly Pro Met Gly Val Gln Gly Pro Glu Gly Pro Arg Gly Asn Pro Gly 260 265 270 Leu Pro Gly Pro Pro Gly Ala Pro Gly Val Pro Gly Ile Gly Ile Gln 275 280 285 Gly Pro Ala Gly Pro Pro Gly Pro Pro Gly Arg Phe Pro Ala Thr Glu 290 295 300 His Cys Ser Tyr Val Thr Gly Glu Cys Val Ile Asn Gly Asn Asp Ile 305 310 315 320 Phe His Ala Met Arg Glu Thr Gly Asp Val Ser Cys Pro Arg Asn Tyr 325 330 335 Tyr Val Lys Gly Val Asp Tyr Ile Lys Cys Asn Asn Gly Ala Glu Tyr 340 345 350 His Asn Thr Leu Gln Leu Arg Leu Thr Cys Cys Leu Leu 355 360 365 <210> 52 <211> 7 <212> PRT <213> Artificial sequence <220> <223> TEV protease <400> 52 Glu Asn Leu Tyr Phe Gln Gly 1 5 <210> 53 <211> 29 <212> PRT <213> Artificial sequence <220> <223> The polypeptide of Example 3 <400> 53 Gly Val Gly Val Ala Pro Gly Glu Val Gly Val Ala Pro Gly Glu Val 1 5 10 15 Gly Val Ala Pro Gly Glu Val Gly Val Ala Pro Gly Glu 20 25 <210> 54 <211> 34 <212> DNA <213> Artificial sequence <220> <223> O3'SUTRpsbD <400> 54 cgatgagttg tttttttatttggagatac acgc 34 <210> 55 <211> 416 <212> PRT <213> Artificial sequence <220> <223> 3F-TV-GtCLP-TV-HA <400> 55 Met Ala Asp Tyr Lys Asp Asp Asp Asp Lys Asp Tyr Lys Asp Asp Asp 1 5 10 15 Asp Lys Asp Tyr Lys Asp Asp Asp Asp Lys Glu Asn Leu Tyr Phe Gln 20 25 30 Gly Ser Arg Arg Arg Gln Gly Ala Ala Ala Ala Ala Ala Gly Ala Ala 35 40 45 Gly Leu Ile Leu Leu Val Leu Leu Ser Leu Val Ser Leu Arg Ala His 50 55 60 Glu Arg Lys Pro Thr Ser Leu Val Glu Gly Ser Arg Thr Thr Met Met 65 70 75 80 Pro Thr Trp Lys Val Thr Gly Thr Leu Asn Pro Val Glu Glu Thr Glu 85 90 95 Lys Glu Val Glu Glu Glu Pro Ser Ser Pro Pro Pro Pro Ala Val Gln 100 105 110 Ala Thr Pro Asp Val Pro Gln Glu Gln Glu Ala Glu Pro Ser Gln Ser 115 120 125 Glu Trp Val Pro Pro Pro Lys Pro Trp Glu Pro Leu Ala Glu Lys Asn 130 135 140 Ile Gly Arg Leu Glu Arg Thr Asn Arg Glu Val Val Arg Gly Asp Met 145 150 155 160 Glu Asn Ala Asn Met Ile Lys Lys Leu Arg Ala Ala Met Arg Leu Leu 165 170 175 Lys Lys Gln Phe Ala Ala Lys Ile Val Asn Leu Glu Gln Ala Tyr Asp 180 185 190 Arg Arg Leu Ala Arg Glu Ser Arg Ile Leu His Gln Arg Ile Asn Thr 195 200 205 Val Ala Met Gln Pro Gly Pro Thr Gly Met Ser Gly Pro Gln Gly Leu 210 215 220 Pro Gly Pro Thr Gly Pro Ala Gly Gln Asn Gly Ser Pro Gly Ser Pro 225 230 235 240 Gly Ala Met Gly Pro Gln Gly Pro Met Gly Pro Arg Gly Tyr Arg Gly 245 250 255 Leu Arg Gly Pro Pro Gly Asp Gln Gly Arg Pro Gly Asp Thr Gly Arg 260 265 270 Pro Gly Ala Pro Gly Ile Pro Gly Gln Ile Gly Gln Arg Gly Pro Ile 275 280 285 Gly Pro Met Gly Val Gln Gly Pro Glu Gly Pro Arg Gly Asn Pro Gly 290 295 300 Leu Pro Gly Pro Pro Gly Ala Pro Gly Val Pro Gly Ile Gly Ile Gln 305 310 315 320 Gly Pro Ala Gly Pro Pro Gly Pro Pro Gly Arg Phe Pro Ala Thr Glu 325 330 335 His Cys Ser Tyr Val Thr Gly Glu Cys Val Ile Asn Gly Asn Asp Ile 340 345 350 Phe His Ala Met Arg Glu Thr Gly Asp Val Ser Cys Pro Arg Asn Tyr 355 360 365 Tyr Val Lys Gly Val Asp Tyr Ile Lys Cys Asn Asn Gly Ala Glu Tyr 370 375 380 His Asn Thr Leu Gln Leu Arg Leu Thr Cys Cys Leu Leu Gly Leu Asn 385 390 395 400 Glu Asn Leu Tyr Phe Gln Gly Tyr Pro Tyr Asp Val Pro Asp Tyr Ala 405 410 415 <210> 56 <211> 62 <212> DNA <213> Artificial sequence <220> <223> O5'SCL70 <400> 56 ctgcagtagc tagctctgca ggtcacctgt taagcataat ctggaacatc ataaggataa 60 cc 62 <210> 57 <211> 62 <212> DNA <213> Artificial sequence <220> <223> O3'SCL70 <400> 57 cgatgagttg tttttttatt ttggagatac acgcaccatg gcagattata aagacgatga 60 tg 62 <210> 58 <211> 1317 <212> DNA <213> Artificial Sequence <220> <223> FPCR-SCL70 <400> 58 ctgcagtagc tagctctgca ggtcacctgt taagcataat ctggaacatc ataaggataa 60 ccttggaaat aaagattctc gtttaaacct aataaacaac atgttaaacg taactgtaaa 120 gtattgtggt attcagcacc attattgcat ttgatgtaat caacaccttt aacatagtag 180 tttcttggac aacttacatc gccagtttca cgcatagcat ggaaaatgtc attgccgtta 240 attacacatt caccagtaac atagctacaa tgttctgtag caggaaagcg tcctggtgga 300 ccaggaggac ctgctggtcc ttgaatacca atacctggaa cacctggtgc acctggtgga 360 cctggtaatc caggattacc acgtggacct tcagggcctt gaacacccat tggaccaata 420 ggtccacgtt gaccaatttg acctggaata ccaggagcac caggtctacc tgtatcacca 480 ggacgacctt gatcaccagg tggaccacgt aaaccacgat aacctcttgg acccattggt 540 ccttgtggac ccattgctcc tggtgagcct ggagaaccat tttggcctgc tggaccagta 600 ggacctggta aaccttgagg acctgacatg cctgttggac caggttgcat agctactgta 660 ttaatacgtt ggtgtaaaat acgactttca cgagctaaac gacgatcata agcttgttct 720 aaattcacaa ttttagctgc aaattgtttc tttaataaac gcatagcagc acgtaatttt 780 tttatcatgt ttgcattttc catatcacca cgaactactt cacgattagt acgttctaaa 840 cgaccaatat ttttttcagc taatggttcc caaggtttag gtggtggaac ccattctgac 900 tgacttggtt cggcttcttg ctcttgaggt acatctggtg tagcttgtac tgctggtggt 960 ggaggtgaag atggttcttc ttctacttct ttctctgttt cttcaactgg atttaatgta 1020 ccagttactt tccatgttgg catcattgta gtgcgtgaac cttcaactaa agatgttggt 1080 ttgcgttcat gagcacgtaa agatactaat gaaagtaaca caagtaagat taatcctgct 1140 gcaccagcgg ctgcggcagc agcaccttga cgacgacgac taccttgaaa ataaaggttt 1200 tctttatcgt catcatcttt atagtctttg tcgtcatcgt ctttgtaatc tttatcatca 1260 tcgtctttat aatctgccat ggtgcgtgta tctccaaaat aaaaaaacaa ctcatcg 1317 <210> 59 <211> 119 <212> PRT <213> Synthetic Sequence <220> <223> Collagen-like domain from Chroomonas CCMP2712 <400> 59 Gly Pro Thr Gly Met Ser Gly Pro Gln Gly Leu Pro Gly Pro Thr Gly 1 5 10 15 Pro Ala Gly Gln Asn Gly Ser Pro Gly Ser Pro Gly Ala Met Gly Pro 20 25 30 Gln Gly Pro Met Gly Pro Arg Gly Tyr Arg Gly Leu Arg Gly Pro Pro 35 40 45 Gly Asp Gln Gly Arg Pro Gly Asp Thr Gly Arg Pro Gly Ala Pro Gly 50 55 60 Ile Pro Gly Gln Ile Gly Gln Arg Gly Pro Ile Gly Pro Met Gly Val 65 70 75 80 Gln Gly Pro Glu Gly Pro Arg Gly Asn Pro Gly Leu Pro Gly Pro Pro 85 90 95 Gly Ala Pro Gly Val Pro Gly Ile Gly Ile Gln Gly Pro Ala Gly Pro 100 105 110 Pro Gly Pro Pro Gly Arg Phe 115 <210> 60 <211> 16 <212> PRT <213> Artificial sequence <220> <223> Cys knot and CR4 repeats <400> 60 Gly Pro Cys Cys Gly Pro Pro Gly Pro Pro Gly Pro Pro Gly Pro Pro 1 5 10 15 <210> 61 <211> 27 <212> PRT <213> Artificial sequence <220> <223> Folded fiber protein of T4 bacteriophage <400> 61 Gly Tyr Ile Pro Glu Ala Pro Arg Asp Gly Gln Ala Tyr Val Arg Lys 1 5 10 15 Asp Gly Glu Trp Val Leu Leu Ser Thr Phe Leu 20 25 <210> 62 <211> 180 <212> PRT <213> Artificial sequence <220> <223> GtCCLD <400> 62 Gly Ser Gly Pro Cys Cys Gly Pro Pro Gly Pro Pro Gly Pro Pro Gly 1 5 10 15 Pro Pro Gly Pro Thr Gly Met Ser Gly Pro Gln Gly Leu Pro Gly Pro 20 25 30 Thr Gly Pro Ala Gly Gln Asn Gly Ser Pro Gly Ser Pro Gly Ala Met 35 40 45 Gly Pro Gln Gly Pro Met Gly Pro Arg Gly Tyr Arg Gly Leu Arg Gly 50 55 60 Pro Pro Gly Asp Gln Gly Arg Pro Gly Asp Thr Gly Arg Pro Gly Ala 65 70 75 80 Pro Gly Ile Pro Gly Gln Ile Gly Gln Arg Gly Pro Ile Gly Pro Met 85 90 95 Gly Val Gln Gly Pro Glu Gly Pro Arg Gly Asn Pro Gly Leu Pro Gly 100 105 110 Pro Pro Gly Ala Pro Gly Val Pro Gly Ile Gly Ile Gln Gly Pro Ala 115 120 125 Gly Pro Pro Gly Pro Pro Gly Arg Phe Gly Pro Pro Gly Pro Pro Gly 130 135 140 Pro Pro Gly Pro Pro Gly Pro Cys Cys Gly Tyr Ile Pro Glu Ala Pro 145 150 155 160 Arg Asp Gly Gln Ala Tyr Val Arg Lys Asp Gly Glu Trp Val Leu Leu 165 170 175 Ser Thr Phe Leu 180 <210> 63 <211> 27 <212> PRT <213> Artificial Sequence <220> <223> 3HA <400> 63 Tyr Pro Tyr Asp Val Pro Asp Tyr Ala Tyr Pro Tyr Asp Val Pro Asp 1 5 10 15 Tyr Ala Tyr Pro Tyr Asp Val Pro Asp Tyr Ala 20 25 <210> 64 <211> 207 <212> PRT <213> Artificial sequence <220> <223> GtCCLD-3HA <400> 64 Gly Ser Gly Pro Cys Cys Gly Pro Pro Gly Pro Pro Gly Pro Pro Gly 1 5 10 15 Pro Pro Gly Pro Thr Gly Met Ser Gly Pro Gln Gly Leu Pro Gly Pro 20 25 30 Thr Gly Pro Ala Gly Gln Asn Gly Ser Pro Gly Ser Pro Gly Ala Met 35 40 45 Gly Pro Gln Gly Pro Met Gly Pro Arg Gly Tyr Arg Gly Leu Arg Gly 50 55 60 Pro Pro Gly Asp Gln Gly Arg Pro Gly Asp Thr Gly Arg Pro Gly Ala 65 70 75 80 Pro Gly Ile Pro Gly Gln Ile Gly Gln Arg Gly Pro Ile Gly Pro Met 85 90 95 Gly Val Gln Gly Pro Glu Gly Pro Arg Gly Asn Pro Gly Leu Pro Gly 100 105 110 Pro Pro Gly Ala Pro Gly Val Pro Gly Ile Gly Ile Gln Gly Pro Ala 115 120 125 Gly Pro Pro Gly Pro Pro Gly Arg Phe Gly Pro Pro Gly Pro Pro Gly 130 135 140 Pro Pro Gly Pro Pro Gly Pro Cys Cys Gly Tyryl Pro Glu Ala Pro 145 150 155 160 Arg Asp Gly Gln Ala Tyr Val Arg Lys Asp Gly Glu Trp Val Leu Leu 165 170 175 Ser Thr Phe Leu Tyr Pro Tyr Asp Val Pro Asp Tyr Ala Tyr Pro Tyr 180 185 190 Asp Val Pro Asp Tyr Ala Tyr Pro Tyr Asp Val Pro Asp Tyr Ala 195 200 205 <210> 65 <211> 617 <212> DNA <213> The snowstorm <220> <223> PCR product Gtccld-2 <400> 65 gatgatgatg acaaagatta caaagacgat gatgataag gatccggtcc atgttgtgga ccacctggtc ctccaggtcc accaggacca cctggacca caggtatgag tggaccaca 180. ggtttaccag gtcctactgg tcctgctggt caaaacggtt caccaggttc tccaggtgca atgggtccac aaggtcctat gggaccacgt ggttatcgtg gtttacgtgg tccaccaggt 240 gatcaaggtc gtcctggtga tacaggtcgt ccaggtgctc caggtattcc aggtcagatt ggtcaacgtg gtccaattgg tcctatgggt gttcaaggac cagaaggtcc aagaggtaat ccaggccttc caggaccacc aggtgcacca ggtgtacctg gtattggcat tcaaggtcct 420 gctggcccac ctggtccacc tggtcgtttt ggtcctcctg gccctccagg cccaccaggt 480 cctcctggtc catgttgcgg ctatatccca gaagctccac gtgatggtca agcttatgtt 540 600. cgcaaagatg gtgaatgggt gttattatca acattcttat aagtttaaac gtcgacctgc agagctagct actgcag 617 <210> 66 <211> 440 <212> DNA <213> The snowstorm <220> <223> PCR Probe Gtcld <400> 66 gatgatgatg acaaagatta caaagacgat gatgataaag gatccggacc aacaggtatg 120. agtggaccac aaggtttacc aggtcctact ggtcctgctg gtcaaaacgg ttcaccaggt tctccaggtg caatgggtcc acaaggtcct atgggaccac gtggttatcg tggtttacgt 180 ggtccaccag gtgatcaagg tcgtcctggt gatacaggtc gtccaggtgc tccaggtatt 240 ccaggtcaga ttggtcaacg tggtccaatt ggtcctatgg gtgttcaagg accagaaggt 300 ccaagaggta atccaggcct tccaggacca ccaggtgcac caggtgtacc tggtattggc 360 attcaaggtc ctgctggccc acctggtcca cctggtcgtt tttaagttta aacgtcgacc 420 tgcagagcta gctactgcag 440 <210> 67 <211> 35 <212> DNA <213> Artificial sequence <220> <223> O5'ASTatpA2 <400> 67 cctacttaat taaaaactgc agtagctagc tctgc 35 <210> 68 <211> 34 <212> DNA <213> Artificial sequence <220> <223> O3'SUTRpsbD <400> 68 cgatgagttg tttttttatt ttggagatac acgc 34 <210> 69 <211> 62 <212> DNA <21 <400> 69 ctgcagtagc tagctctgca ggtcacctgt taagcataat ctggaacatc ataaggataa 60 cc 62 <210> 70 <211> 33 <212> DNA <213> Artificial sequence <220> <223> O5'PpsbDCla2 <400> 70 catcgatgat gaaattaaat ggatatttgg tac 33 <210> 71 <211> 4 <212> PRT <213> Artificial sequence <220> <223> FX <400> 71 Ile Glu Gly Arg 1 <210> 72 <211> 116 <212> DNA <213> Artificial sequence <220> <223> O5'Gibs-ELP4 <400> 72 ctgcagtagc tagctctgca ggtcgacgtt taaacttaag cgtaatctgg aacatcatat 60 ggataaccag gagctacacc cacacctgga gcaacaccta ctccaggagc aacacc 116 <210> 73 <211> 110 <212> DNA <213> Artificial sequence <220> <223> O3'Gibs-ELP4 <400> 73 gctgaagatt gtatgcgtac ttgtggtggt gctgattaca aagacgatga tgacaaagta 60 ggtgtagctc ctggtgttgg tgttgctcct ggagtaggtg ttgctccagg 110 <210> 74 <211> 62 <212> DNA <213> Artificial sequence <220> <223> O5'Gibs01BE <400> 74 ctgcagtagc tagctctgca ggtcgacgtt taaacttaac caggagctac acccacacct 60 gg 62 <210> 75 <211> 70 <2aatgcaaaag caggattatg tcaaacattt gtgtatggag gttgtcgtgc taaacgtaac 180 aactttaaat cagctgaaga ttgtatgcgt acttgtggtg gtgctgatta caaagacgat 240 gatgacaaag taggtgtagc tcctggtgtt ggtgttgctc ctggagtagg tgttgctcca 300 ggtgtgggtg tagctcctgg ttaagtttaa acgtcgacct gcagagctag ctactgcag 359 <210> 77 <211> 207 <212> DNA <213> Artificial sequence gg 62 <210> 79 <211> 75 <212> DNA <213> Artificial sequence <220> <223> O3'Gibs02BE <400> 79 agatgatgat gacaaagatt acaaagacga tgatgataaa ggatccgatt acaaagacga 60 tgatgacaaa gtagg 75 <210> 80 <211> 28 <212> PRT <213> Artificial sequence <220> <223> ELPE4 <400> 80 Val Gly Val Ala Pro Gly Glu Val Gly Val Ala Pro Gly Glu Val Gly 1 5 10 15 Val Ala Pro Gly Glu Val Gly Val Ala Pro Gly Glu 20 25[[ID=4<213> Artificial Sequence <220> <223> Expression cassette of Example 3 <400> 82 atgaaattaa atggatattt ggtacattta attccacaaa aatgtccaat acttaaaata 60 caaaattaaa agtattagtt gtaaacttga ctaacatttt aaattttaaa ttttttccta 120 attatatatt ttacttgcaa aatttataaa aattttatgc atttttatat cataataata 180 aaacctttat tcatggttta taatataata attgtgatga ctatgcacaa agcagttcta 240 gtcccatata tataactata tataacccgt ttaaagattt atttaaaaat atgtgtgtaa 300 aaaatgctta tttttaattt tattttatat aagttataat attaaataca caatgattaa 360 aattaaataa taataaattt aacgtaacga tgagttgttt ttttattttg gagatacacg 420 caccatggca tatccttatg atgtaccaga ttatgccaat aataacgact tatttcaagc 480 atcacgtcgt cgtttcttag ctcaattagg aggtttaaca gttgctggta tgttaggtcc 540 atctttactt actccacgta gagctacagc agctcaagct gattacaaag atgatgacga 600 taaagactat aaagatgatg atgacaaaga ttacaaagac gatgatgata aaggatccat 660 tgaaggtcgt agaccagact tttgcttaga accaccatat acaggtccat gtaaagctcg 720 tatcattcgc tatttctaca atgcaaaagc aggattatgt caaacatttg tgtatggagg 780 ttgtcgtgct aaacgtaaca actttaaatc agctgaagat tgtatgcgta cttgtggtgg 840 tgcttaagtt taaacgtcga cctgcagagc tagctactgc agtttttaat taagtaggaa 900 ctcggtatat gctcttttgg ggtcttatta gctagtatta gttaactaac aaaagatcaa 960 tattttagtt tgttttatat attttattac ttaagtagta aggatttgca tttagcaatc 1020 ttaaatactt aagtaataat ctataataa aatatatttt cgctttaaaa cttataaaaa 1080 ttatttgctc gttataagcc taaaaaaacg taggatctct acgagatatt acattgtttt 1140 tttctttaat tggctttaat attactttgt atatataaac caaagtactt gttaatagtt 1200 attaaattat attaactata footcaaag aaattttttg ctaaaaaaaag tatgttaaca 1260 ttaaaaattt ttgtttatac ag 1282 <210> 83 <211> 283 <212> PRT <213> Artificial sequence <220> <223> HA-SP-3F-GtCCLD-3HA <400> 83 Met Ala Tyr Pro Tyr Asp Val Pro Asp Tyr Ala Asn Asn Asn Asp Leu 1 5 10 15 Phe Gln Ala Ser Arg Arg Arg Phe Leu Ala Gln Leu Gly Gly Leu Thr 20 25 30 Val Ala Gly Met Leu Gly Pro Ser Leu Leu Thr Pro Arg Arg Ala Thr 35 40 45 Ala Ala Gln Ala Asp Tyr Lys Asp Asp Asp Asp Lys Asp Tyr Lys Asp 50 55 60 Asp Asp Asp Lys Asp Tyr Lys Asp Asp Asp Asp Lys Gly Ser Gly Pro 65 70 75 80 Cys Cys Gly Pro Pro Gly Pro Pro Gly Pro Pro Gly Pro Pro Gly Pro 85 90 95 Thr Gly Met Ser Gly Pro Gln Gly Leu Pro Gly Pro Thr Gly Pro Ala 100 105 110 Gly Gln Asn Gly Ser Pro Gly Ser Pro Gly Ala Met Gly Pro Gln Gly 115 120 125 Pro Met Gly Pro Arg Gly Tyr Arg Gly Leu Arg Gly Pro Pro Gly Asp 130 135 140 Gln Gly Arg Pro Gly Asp Thr Gly Arg Pro Gly Ala Pro Gly Ile Pro 145 150 155 160 Gly Gln Ile Gly Gln Arg Gly Pro Ile Gly Pro Met Gly Val Gln Gly 165 170 175 Pro Glu Gly Pro Arg Gly Asn Pro Gly Leu Pro Gly Pro Pro Gly Ala 180 185 190 Pro Gly Val Pro Gly Ile Gly Ile Gln Gly Pro Ala Gly Pro Pro Gly 195 200 205 Pro Pro Gly Arg Phe Gly Pro Pro Gly Pro Pro Gly Pro Pro Gly Pro 210 215 220 Pro Gly Pro Cys Cys Gly Tyr Ile Pro Glu Ala Pro Arg Asp Gly Gln 225 230 235 240 Ala Tyr Val Arg Lys Asp Gly Glu Trp Val Leu Leu Ser Thr Phe Leu 245 250 255 Tyr Pro Tyr Asp Val Pro Asp Tyr Ala Tyr Pro Tyr Asp Val Pro Asp 260 265 270 Tyr Ala Tyr Pro Tyr Asp Val Pro Asp Tyr Ala 275 280 <210> 84 <211> 652 <212> DNA <213> Artificial sequence <220> <223> Synthetic codon-optimized gene Gtccld-3ha <400> 84 gtccatggct ggatccggtc catgttgtgg accacctggt cctccaggtc caccaggacc 60 gtccatggct ggatccggtc catgttgtgg accacctggt cctccaggtc caccaggacc 60 acctggacca acaggtatga gtggaccaca aggtttacca ggtcctactg gtcctgctgg 120 acctggacca acaggtatga gtggaccaca aggtttacca ggtcctactg gtcctgctgg 120 tcaaaacggt tcaccaggtt ctccaggtgc aatgggtcca caaggtccta tgggaccacg 180 tcaaaacggt tcaccaggtt ctccaggtgc aatgggtcca caaggtccta tgggaccacg 180 tggttatcgt ggtttacgtg gtccaccagg tgatcaaggt cgtcctggtg atacaggtcg 240 tggttatcgt ggtttacgtg gtccaccagg tgatcaaggt cgtcctggtg atacaggtcg 240 tccaggtgct ccaggtattc caggtcagat tggtcaacgt ggtccaattg gtcctatggg 300 tccaggtgct ccaggtattc caggtcagat tggtcaacgt ggtccaattg gtcctatggg 300 tgttcaagga ccagaaggtc caagaggtaa tccaggcctt ccaggaccac caggtgcacc 360 tgttcaagga ccagaaggtc caagaggtaa tccaggcctt ccaggaccac caggtgcacc 360 aggtgtacct ggtattggca ttcaaggtcc tgctggccca cctggtccac ctggtcgttt 420 aggtgtacct ggtattggca ttcaaggtcc tgctggccca cctggtccac ctggtcgttt 420 tggtcctcct ggccctccag gcccaccagg tcctcctggt ccatgttgcg gctatatccc 480 tggtcctcct ggccctccag gcccaccagg tcctcctggt ccatgttgcg gctatatccc 480 agaagctcca cgtgatggtc aagcttatgt tcgcaaagat ggtgaatggg tgttattatc 540 agaagctcca cgtgatggtc aagcttatgt tcgcaaagat ggtgaatggg tgttattatc 540 aacattctta tacccttatg acgtaccaga ttacgcatat ccttatgatg ttccagatta 600 aacattctta tacccttatg acgtaccaga ttacgcatat ccttatgatg ttccagatta 600 tgcctatcca tacgatgtac ctgactatgc ttaagtttaa acgtcgacct gc 652 tgcctatcca tacgatgtac ctgactatgc ttaagtttaa acgtcgacct gc 652 <210> 85 <210> 85 <211> 256 <211> 256 <212> PRT <212> PRT <213> 人工序列 <213> Artificial Sequence <220> <220> <223> HA-SP-3F-GtCCLD <400> 85 Met Ala Tyr Pro Tyr Asp Val Pro Asp Tyr Ala Asn Asn Asn Asp Leu 1 5 10 15 Phe Gln Ala Ser Arg Arg Arg Phe Leu Ala Gln Leu Gly Gly Leu Thr 20 25 30 Val Ala Gly Met Leu Gly Pro Ser Leu Leu Thr Pro Arg Arg Ala Thr 35 40 45 Ala Ala Gln Ala Asp Tyr Lys Asp Asp Asp Asp Lys Asp Tyr Lys Asp 50 55 60 Asp Asp Asp Lys Asp Tyr Lys Asp Asp Asp Asp Lys Gly Ser Gly Pro 65 70 75 80 Cys Cys Gly Pro Pro Gly Pro Pro Gly Pro Pro Gly Pro Pro Gly Pro 85 90 95 Thr Gly Met Ser Gly Pro Gln Gly Leu Pro Gly Pro Thr Gly Pro Ala 100 105 110 Gly Gln Asn Gly Ser Pro Gly Ser Pro Gly Ala Met Gly Pro Gln Gly 115 120 125 Pro Met Gly Pro Arg Gly Tyr Arg Gly Leu Arg Gly Pro Pro Gly Asp 130 135 140 Gln Gly Arg Pro Gly Asp Thr Gly Arg Pro Gly Ala Pro Gly Ile Pro 145 150 155 160 Gly Gln Ile Gly Gln Arg Gly Pro Ile Gly Pro Met Gly Val Gln Gly 165 170 175 Pro Glu Gly Pro Arg Gly Asn Pro Gly Leu Pro Gly Pro Pro Gly Ala 180 185 190 Pro Gly Val Pro Gly Ile Gly Ile Gln Gly Pro Ala Gly Pro Pro Gly 195 200 205 Pro Pro Gly Arg Phe Gly Pro Pro Gly Pro Pro Gly Pro Pro Gly Pro 210 215 220 Pro Gly Pro Cys Cys Gly Tyr Ile Pro Glu Ala Pro Arg Asp Gly Gln 225 230 235 240 Ala Tyr Val Arg Lys Asp Gly Glu Trp Val Leu Leu Ser Thr Phe Leu 245 250 255 <210> 86 <211> 197 <212> PRT <213> Artificial Sequence<000188�><220> <223> HA-SP-3F-GtCLD <4OO> 86 Met Ala Tyr Pro Tyr Asp Val Pro Asp Tyr Ala Asn Asn Asn Asp Leu 1 5 10 15 Phe Gln Ala Ser Arg Arg Arg Phe Leu Ala Gln Leu Gly Gly Leu Thr 20 25 30 Val Ala Gly Met Leu Gly Pro Ser Leu Leu Thr Pro Arg Arg Ala Thr 35 40 45 Ala Ala Gln Ala Asp Tyr Lys Asp Asp Asp Asp Lys Asp Tyr Lys Asp 50 55 60 Asp Asp Asp Lys Asp Tyr Lys Asp Asp Asp Asp Lys Gly Ser Gly Pro 65 70 75 80 Thr Gly Met Ser Gly Pro Gln Gly Leu Pro Gly Pro Thr Gly Pro Ala 85 90 95 Gly Gln Asn Gly Ser Pro Gly Ser Pro Gly Ala Met Gly Pro Gln Gly 100 105 110 Pro Met Gly Pro Arg Gly Tyr Arg Gly Leu Arg Gly Pro Pro Gly Asp 115 120 125 Gln Gly Arg Pro Gly Asp Thr Gly Arg Pro Gly Ala Pro Gly Ile Pro 130 135 140 Gly Gln Ile Gly Gln Arg Gly Pro Ile Gly Pro Met Gly Val Gln Gly 145 150 155 160 Pro Glu Gly Pro Arg Gly Asn Pro Gly Leu Pro Gly Pro Pro Gly Ala 165 170 175 Pro Gly Val Pro Gly Ile Gly Ile Gln Gly Pro Ala Gly Pro Pro Gly 180 185 190 Pro Pro Gly Arg Phe 195 <210> 87 <211> 172 <212> PRT <213> Artificial Sequence <220> <223> HA - SP - 3F - FX - APRO - F - ELP4 <400> 87 Met Ala Tyr Pro Tyr Asp Val Pro Asp Tyr Ala Asn Asn Asn Asp Leu 1 5 10 15 Phe Gln Ala Ser Arg Arg Arg Phe Leu Ala Gln Leu Gly Gly Leu Thr 20 25 30 Val Ala Gly Met Leu Gly Pro Ser Leu Leu Thr Pro Arg Arg Ala Thr 35 40 45 Ala Ala Gln Ala Asp Tyr Lys Asp Asp Asp Asp Lys Asp Tyr Lys Asp 50 55 60 Asp Asp Asp Lys Asp Tyr Lys Asp Asp Asp Asp Lys Gly Ser Ile Glu 65 70 75 80 Gly Arg Arg Pro Asp Phe Cys Leu Glu Pro Pro Tyr Thr Gly Pro Cys 85 90 95 Lys Ala Arg Ile Ile Arg Tyr Phe Tyr Asn Ala Lys Ala Gly Leu Cys 100 105 110 Gln Thr Phe Val Tyr Gly Gly Cys Arg Ala Lys Arg Asn Asn Phe Lys 115 120 125 Ser Ala Glu Asp Cys Met Arg Thr Cys Gly Gly Ala Asp Tyr Lys Asp 130 135 140 Asp Asp Asp Lys Val Gly Val Ala Pro Gly Val Gly Val Ala Pro Gly 145 150 155 160 Val Gly Val Ala Pro Gly Val Gly Val Ala Pro Gly 165 170 <210> 88 <211> 67 <212> PRT <213> Synthetic Sequence <220> <223> 3F-1F-ELP4 <400> 88 Asp Tyr Lys Asp Asp Asp Asp Lys Asp Tyr Lys Asp Asp Asp Asp Lys 1 5 10 15 Asp Tyr Lys Asp Asp Asp Asp Lys Gly Ser Asp Tyr Lys Asp Asp Asp 20 25 30 Asp Lys Val Gly Val Ala Pro Gly Val Gly Val Ala Pro Gly Val Gly 35 40 45 Val Ala Pro Gly Val Gly Val Ala Pro Gly Tyr Pro Tyr Asp Val Pro 50 55 60 Asp Tyr Ala 65 <210> 89 <211> 193 <212> PRT <213> Artificial sequence <220> <223> HA-SP-3F-FX-APRO-LGM-TV-ELPE4 <400> 89 Met Ala Tyr Pro Tyr Asp Val Pro Asp Tyr Ala Asn Asn Asn Asp Leu 1 5 10 15 Phe Gln Ala Ser Arg Arg Arg Phe Leu Ala Gln Leu Gly Gly Leu Thr 20 25 30 Val Ala Gly Met Leu Gly Pro Ser Leu Leu Thr Pro Arg Arg Ala Thr [[ID=2I]]35 40 45 Ala Ala Gln Ala Asp Tyr Lys Asp Asp Asp Asp Lys Asp Tyr Lys Asp 50 55 60 Asp Asp Asp Lys Asp Tyr Lys Asp Asp Asp Asp Lys Gly Ser Ile Glu 65 70 75 80 Gly Arg Arg Pro Asp Phe Cys Leu Glu Pro Pro Tyr Thr Gly Pro Cys 85 90 95 Lys Ala Arg Ile Ile Arg Tyr Phe Tyr Asn Ala Lys Ala Gly Leu Cys 100 105 110 Gln Thr Phe Val Tyr Gly Gly Cys Arg Ala Lys Arg Asn Asn Phe Lys 115 120 125 Ser Ala Glu Asp Cys Met Arg Thr Cys Gly Gly Ala Arg Ser Gly Gly It should be noted that in the original text, there is a potential error in the line "[[ID=2I]]35 40 四十五 ", where "四十五" should be "45" for consistency. The above translation is based on the corrected content. 130 135 140 Gly Gly Ser Ser Gly Gly Gly Gly Gly Gly Ser Ser Arg Ser Glu Asn 145 150 155 160 Leu Tyr Phe Gln Gly Val Gly Val Ala Pro Gly Glu Val Gly Val Ala 165 170 175 Pro Gly Glu Val Gly Val Ala Pro Gly Glu Val Gly Val Ala Pro Gly 180 185 190 Glu <210> 90 <211> 191 <212> PRT <213> Artificial Sequence <220> <223> HA-SP-3F-FX-APRO-LGM-EK-ELPE4 <400> 90 Met Ala Tyr Pro Tyr Asp Val Pro Asp Tyr Ala Asn Asn Asn Asp Leu 1 5 10 15 Phe Gln Ala Ser Arg Arg Arg Phe Leu Ala Gln Leu Gly Gly Leu Thr 20 25 30 Val Ala Gly Met Leu Gly Pro Ser Leu Leu Thr Pro Arg Arg Ala Thr 35 40 45 Ala Ala Gln Ala Asp Tyr Lys Asp Asp Asp Asp Lys Asp Tyr Lys Asp 50 55 60 Asp Asp Asp Lys Asp Tyr Lys Asp Asp Asp Asp Lys Gly Ser Ile Glu 65 70 75 80 Gly Arg Arg Pro Asp Phe Cys Leu Glu Pro Pro Tyr Thr Gly Pro Cys 85 90 95 Lys Ala Arg Ile Ile Arg Tyr Phe Tyr Asn Ala Lys Ala Gly Leu Cys 100 105 110 Gln Thr Phe Val Tyr Gly Gly Cys Arg Ala Lys Arg Asn Asn Phe Lys 115 120 125 Ser Ala Glu Asp Cys Met Arg Thr Cys Gly Gly Ala Arg Ser Gly Gly 130 135 140 Gly Gly Ser Ser Gly Gly Gly Gly Gly Gly Ser Ser Arg Ser Asp Asp 145 150 155 160 Asp Asp Lys Val Gly Val Ala Pro Gly Glu Val Gly Val Ala Pro Gly 165 170 175 Glu Val Gly Val Ala Pro Gly Glu Val Gly Val Ala Pro Gly Glu 180 185 190 <210> 91 <211> 191 <212> PRT <213> Artificial Sequence <220> <223> HA-SP-3F-FX-APRO-LGM-EK-ELPE4 <400> 91 Met Ala Tyr Pro Tyr Asp Val Pro Asp Tyr Ala Asn Asn Asn Asp Leu 1 5 10 15 Phe Gln Ala Ser Arg Arg Arg Phe Leu Ala Gln Leu Gly Gly Leu Thr 20 25 30 Val Ala Gly Met Leu Gly Pro Ser Leu Leu Thr Pro Arg Arg Ala Thr 35 40 45 Ala Ala Gln Ala Asp Tyr Lys Asp Asp Asp Asp Lys Asp Tyr Lys Asp 50 55 60 Asp Asp Asp Lys Asp Tyr Lys Asp Asp Asp Asp Lys Gly Ser Ile Glu 65 70 75 80 Gly Arg Arg Pro Asp Phe Cys Leu Glu Pro Pro Tyr Thr Gly Pro Cys 85 90 95 Lys Ala Arg Ile Ile Arg Tyr Phe Tyr Asn Ala Lys Ala Gly Leu Cys 100 105 110 Gln Thr Phe Val Tyr Gly Gly Cys Arg Ala Lys Arg Asn Asn Phe Lys 115 120 125 Ser Ala Glu Asp Cys Met Arg Thr Cys Gly Gly Ala Arg Ser Gly Gly 130 135 140 Gly Gly Ser Ser Gly Gly Gly Gly Gly Gly Ser Ser Arg Ser Asp Asp 145 150 155 160 Asp Asp Lys Val Gly Val Ala Pro Gly Glu Val Gly Val Ala Pro Gly 165 170 175 Glu Val Gly Val Ala Pro Gly Glu Val Gly Val Ala Pro Gly Glu 180 185 190
Claims
1. A recombinant microalgae comprising a nucleic acid sequence encoding a recombinant protein, polypeptide, or peptide comprising repeating amino acid units, wherein the protein, polypeptide, or peptide is selected from the group consisting of SEQ ID NO: 55, SEQ ID NO: 83, SEQ ID NO: 85, SEQ ID NO: 86, SEQ ID NO: 87, SEQ ID NO: 89, and SEQ ID NO: 90; And the nucleic acid sequence is located in the chloroplast genome of the microalgae, wherein the microalgae is Chlamydomonas.
2. A method for producing a recombinant protein, polypeptide or peptide comprising repeating amino acid units in the chloroplasts of Chlamydomonas microalgae, wherein the protein, polypeptide or peptide is selected from the group consisting of SEQ ID NO: 55, SEQ ID NO: 83, SEQ ID NO: 85, SEQ ID NO: 86, SEQ ID NO: 87, SEQ ID NO: 89 and SEQ ID NO: 90; The method comprises transforming the chloroplast genome of Chlamydomonas microalgae with a nucleic acid sequence encoding the recombinant protein, polypeptide or peptide.
3. The method according to claim 2, comprising: (i) providing a nucleic acid sequence encoding the recombinant protein, polypeptide or peptide; (ii) introducing the nucleic acid sequence according to (i) into an expression vector capable of expressing said nucleic acid sequence; and (iii) Transforming the chloroplast genome of a Chlamydomonas microalgae host cell with the expression vector.
4. The method according to claim 3, further comprising: (iv) identifying the transformed Chlamydomonas microalgae host cell; (v) characterizing a Chlamydomonas microalgae host cell for producing a recombinant protein, polypeptide, or peptide expressed by the nucleic acid sequence; as well as (vi) Extracting recombinant proteins, polypeptides or peptides.
5. The method according to claim 4, further comprising: (vii) purifying the recombinant protein, polypeptide or peptide.
6. The method according to any one of claims 3 to 5, wherein the expression vector further comprises at least one expression cassette comprising a nucleic acid sequence encoding the recombinant protein, polypeptide or peptide.
7. The recombinant Chlamydomonas microalgae according to claim 1 or the method according to any one of claims 2 to 5, wherein the nucleic acid sequence encoding the protein, polypeptide or peptide is codon-optimized for expression in the chloroplast genome of the Chlamydomonas microalgae host cell.
8. The recombinant Chlamydomonas microalgae of claim 1 or the method of any one of claims 3 to 5, wherein the nucleic acid sequence encoding the recombinant protein, polypeptide or peptide is operably fused at its 5' or 3' end to a nucleic acid sequence encoding a vector.
9. The recombinant Chlamydomonas microalgae of claim 1 or the method of any one of claims 3 to 5, wherein the nucleic acid sequence encoding the recombinant protein, polypeptide or peptide is operably linked to at least one regulatory sequence selected from the group consisting of a psbD promoter and 5′UTR or a 16S rRNA promoter fused to an atpA 5′UTR, a psaA promoter and 5′UTR, an atpA promoter and 5′UTR, and atpA and rbcL 3′UTRs.
10. The method of claim 6, wherein the at least one expression cassette further comprises a nucleic acid sequence encoding an epitope tag peptide operably fused at its 5' or 3' end to the nucleic acid sequence encoding the recombinant protein, polypeptide or peptide.
11. The method of claim 6, wherein the at least one expression cassette further comprises a nucleic acid sequence encoding a signal peptide.
12. The method of claim 6, wherein the at least one expression cassette further comprises a nucleic acid sequence encoding an amino acid sequence that allows production of the recombinant protein, polypeptide, or peptide in a specific cellular compartment.
13. Use of the recombinant Chlamydomonas microalgae according to claims 1 and 7-9 or the recombinant Chlamydomonas microalgae produced according to the method of any one of claims 2-12 for producing a recombinant protein, polypeptide or peptide comprising amino acid repeating units selected from the group consisting of SEQ ID NO: 55, SEQ ID NO: 83, SEQ ID NO: 85, SEQ ID NO: 86, SEQ ID NO: 87, SEQ ID NO: 89 and SEQ ID NO: 90.