Methods and compositions for protein synthesis and secretion

CN117794941BActive Publication Date: 2026-08-11HELENA CORP
View PDF 18 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-29
Publication Date
2026-08-11

Smart Images

  • Figure BDA0004686044970000241
    Figure BDA0004686044970000241
  • Figure BDA0004686044970000251
    Figure BDA0004686044970000251
  • Figure BDA0004686044970000271
    Figure BDA0004686044970000271
Patent Text Reader

Abstract

In some aspects, this document discloses synthetic secretion signal peptides. It also discloses nucleic acid molecules encoding such signal peptides, in some cases operatively linked to protein-coding sequences, and cells comprising such nucleic acid molecules. Furthermore, it discloses methods for secreting polypeptides, including expressing the disclosed signal peptide linked to said polypeptide in a cell. Some aspects include proteins produced by such methods (e.g., human milk proteins), and compositions comprising such proteins.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing of related patent applications

[0002] This application claims priority and interest in U.S. Provisional Application No. 63 / 227,820, filed July 30, 2021, and U.S. Provisional Application No. 63 / 273,858, filed October 21, 2021, the entire contents of which are incorporated herein by reference.

[0003] sequence list

[0004] This application contains a sequence list submitted in XML format, the entire contents of which are incorporated herein by reference. The XML copy was created on July 76, 2022, named HELA_P0005WO_Sequence_Listing.xml, and is 61,471 bytes in size.

[0005] background Technical Field

[0006] Various aspects of this invention relate to at least the fields of microbiology, genetics, and biotechnology. Background Technology

[0007] Yeast is an ideal host for producing recombinant proteins due to its rapid growth, ability to achieve high cell densities, growth on defined minimum media, high protein yields, and post-translational modifications in eukaryotes. The yeast most relevant to protein production is *Pichia pastoris* (also known as *Komagataella pastoris* or *Komagataella phaffii*), because of the extensive availability of genomic information and molecular tools for genome manipulation. These enable the use of *Pichia pastoris* for the production of GRAS (Generally Recognized As Safe) components according to FDA standards.

[0008] For various biotechnological applications, it is generally preferred to produce proteins that are secreted into the growth medium for easy recovery. Pichia pastoris can secrete active recombinant proteins while maintaining low levels of endogenous protein secretion.

[0009] In eukaryotes, secreted proteins are first transported from the cytoplasm to the lumen of the endoplasmic reticulum (ER). This transport into the ER can occur both post-translational (i.e., after polypeptide chain synthesis) and during co-translation (i.e., during the translation of mRNA into its amino acid sequence). Post-translational transport requires molecular chaperones in the cytoplasm to maintain the polypeptide chain in a loose conformation, as well as the role of the ER-resident molecular chaperone Kar2, which acts as a molecular ratchet. Therefore, this process can be hindered by partially folded domains and / or cytoplasmic aggregation. Thus, for biotechnological applications, it is necessary to facilitate co-translational transport. Once in the ER, proteins are glycosylated, their disulfide bonds are isomerized, and they fold back to their native state. Successfully folded proteins are then transported to the Golgi complex, where further glycosylation occurs, and they are packaged into secretory granules fused to the cell membrane, releasing the protein into the extracellular environment.

[0010] The targeting of proteins to the secretion pathway is mediated by secretory peptides. The most widely used peptide in *Pichia pastoris* is the precursor peptide of mating factor α from *Saccharomyces cerevisiae*. It consists of two distinct regions: ii) the pre-region of the first 19 amino acids, which facilitates post-translational transport and is cleaved upon entry into the endoplasmic reticulum; and 2) the pro-segment of 70 amino acids, which serves as the output signal from the endoplasmic reticulum to the Golgi apparatus and is cleaved at the KR site, a dibasic amino acid cleavage site, in the Golgi apparatus.

[0011] There is a need for synthetic secretory signal peptides to achieve higher extracellular protein production. Summary of the Invention

[0012] Various aspects of this disclosure address certain needs by providing novel secretory signal peptides that can efficiently enhance the extracellular production of proteins, including mammalian proteins such as human milk proteins. Certain aspects of this disclosure are (at least in part) based on the development of signal peptides produced by in-frame fusion of a pre-secreting peptide from *P. pastoris* (from i) the α subunit (Ost1) of an oligosaccharide transferase complex from the ER cavity or ii) the GPI-anchored protein Pst1 with (i) a mating factor from *Saccharomyces cerevisiae* or ii) the pro-region of *P. pastoris* Epx1. Therefore, what is described herein is the isolated nucleic acid encoding such a secretory signal peptide, in some cases linked to a recombinant protein (such as human milk protein), and cells comprising such nucleic acids, as well as methods for producing and collecting recombinant proteins from such cells.

[0013] In some embodiments, the isolated nucleic acid encoding the polypeptide described herein comprises a sequence having at least 90% sequence identity with SEQ ID NO: 1, 2, 3, or 4. In some embodiments, the sequence comprises SEQ ID NO: 1, 2, 3, or 4. In some embodiments, the polypeptide further comprises a sequence of a mammalian protein. In some embodiments, the mammalian protein is a human milk protein. In some embodiments, the human milk protein is secretory IgA (sIgA), xanthine dehydrogenase, lactoferrin, lactoperoxidase, butyrophilin, lactadherin, adiponectin, β-casein, κ-casein, leptin, lysozyme, or α-lactalbumin. In some embodiments, the human milk protein is human lactoferrin.

[0014] In some embodiments, the sequence has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO:1. In some embodiments, the sequence includes SEQ ID NO:1. In some embodiments, the isolated nucleic acid includes a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with SEQ ID NO:41. In some embodiments, the nucleic acid sequence includes SEQ ID NO:41. In some embodiments, the polypeptide includes SEQ ID NO:5. In some embodiments, the isolated nucleic acid comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with SEQ ID NO:46. In some embodiments, the nucleic acid sequence comprises SEQ ID NO:46.

[0015] In some embodiments, the sequence has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO:2. In some embodiments, the sequence includes SEQ ID NO:2. In some embodiments, the isolated nucleic acid includes a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with SEQ ID NO:42. In some embodiments, the nucleic acid sequence includes SEQ ID NO:42. In some embodiments, the polypeptide includes SEQ ID NO:6. In some embodiments, the isolated nucleic acid comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with SEQ ID NO:47. In some embodiments, the nucleic acid sequence comprises SEQ ID NO:47.

[0016] In some embodiments, the sequence has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO:3. In some embodiments, the sequence includes SEQ ID NO:3. In some embodiments, the isolated nucleic acid includes a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with SEQ ID NO:43. In some embodiments, the nucleic acid sequence includes SEQ ID NO:43. In some embodiments, the polypeptide includes SEQ ID NO:7. In some embodiments, the isolated nucleic acid comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with SEQ ID NO:48. In some embodiments, the nucleic acid sequence comprises SEQ ID NO:48.

[0017] In some embodiments, the sequence has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO:4. In some embodiments, the sequence includes SEQ ID NO:4. In some embodiments, the isolated nucleic acid includes a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with SEQ ID NO:44. In some embodiments, the nucleic acid sequence includes SEQ ID NO:44. In some embodiments, the polypeptide includes SEQ ID NO:8. In some embodiments, the isolated nucleic acid comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with SEQ ID NO:49. In some embodiments, the nucleic acid sequence comprises SEQ ID NO:49.

[0018] In some embodiments, vectors comprising the nucleic acids disclosed herein (e.g., isolated nucleic acids or their sequences or portions thereof) are also disclosed herein.

[0019] In some aspects, engineered eukaryotic cells, including the nucleic acids disclosed herein, are also disclosed. In some embodiments, the cells are fungal cells. In some embodiments, the fungal cells are *Arxula*, *Aspecium*, *Aurantiochytrium*, *Candida*, *Claviceps*, *Cryptococcus*, *Cunninghamella*, *Geotrichum*, *Hansenula*, *Kluyveromyces*, *Kodamaea*, *Komagataella*, *Leucosporidiella*, and *Oleoma*. Cells of the genera *Lipomyces*, *Mortierella*, *Ogataea*, *Pichia*, *Prototheca*, *Rhizopus*, *Rhodosporidium*, *Rhodotorula*, *Saccharomyces*, *Schizosaccharomyces*, *Tremella*, *Trichosporon*, *Wickerhamomyces*, or *Yarrowia*. In some embodiments, the cells are yeast cells. In some embodiments, the yeast cells are *Komagataella* cells. In some embodiments, the yeast cells are *Komagataella phaffii*, *Komagataella pastoris*, or *Komagataella pseudopastoris* cells. In some aspects, nucleic acids are integrated into the cell's genome. In some respects, nucleic acids are not integrated into the cell's genome.

[0020] In some aspects, methods for producing secreted proteins are also disclosed, the methods comprising growing engineered eukaryotic cells of the present disclosure under conditions sufficient to secrete polypeptides from cells. In some embodiments, the method further comprises collecting the secreted proteins. In some aspects, the secreted proteins are human milk proteins. In some embodiments, the human milk protein is secretory IgA (sIgA), xanthine dehydrogenase, lactoferrin, lactoperoxidase, lactolipoprotein, lactobacillus, adiponectin, β-casein, κ-casein, leptin, lysozyme, or α-lactalbumin. In some embodiments, the human milk protein is human lactoferrin. In some embodiments, the human milk protein comprises one or more human-like N-glycans. In some embodiments, the method further comprises generating a mixture comprising one or more components of human milk protein and infant formula.

[0021] In some aspects, this document also discloses engineered yeast cells comprising nucleic acids encoding polypeptides, the polypeptides comprising sequences having at least 90% sequence identity with SEQ ID NO:1, 2, 3, or 4. In some embodiments, the sequences comprise SEQ ID NO:1, 2, 3, or 4. In some embodiments, the sequences comprise SEQ ID NO:1. In some embodiments, the sequences comprise SEQ ID NO:2. In some embodiments, the sequences comprise SEQ ID NO:3. In some embodiments, the sequences comprise SEQ ID NO:4. In some embodiments, the polypeptides further comprise sequences of mammalian proteins. In some embodiments, the mammalian protein is a human milk protein. In some embodiments, the human milk protein is secretory IgA (sIgA), xanthine dehydrogenase, lactoferrin, lactoperoxidase, lactolipoprotein, lactobacin, adiponectin, β-casein, κ-casein, leptin, lysozyme, or α-lactalbumin. In some embodiments, the human milk protein is human lactoferrin.

[0022] In some aspects, this document describes engineered yeast cells comprising: (a) a first nucleic acid encoding a polypeptide, said polypeptide comprising: (i) a sequence having at least 90% sequence identity with SEQ ID NO: 1, 2, 3 or 4, and (ii) a sequence of a human milk protein; and (b) a second nucleic acid encoding an α-1,2-mannosidase (Man-I) protein, wherein the cell does not express a functional OCH1 protein. In some embodiments, the sequence of (i) comprises SEQ ID NO: 1, 2, 3 or 4. In some embodiments, the human milk protein is secretory IgA (sIgA), xanthine dehydrogenase, lactoferrin, lactoperoxidase, lactolipoprotein, lactobacin, adiponectin, β-casein, κ-casein, leptin, lysozyme, or α-lactalbumin. In some embodiments, the human milk protein is human lactoferrin. In some embodiments, the human milk protein is human α-lactalbumin. In some embodiments, the Man-I protein is fused to an HDEL C-terminal tag. In some embodiments, the cell also includes a third nucleic acid encoding one or more of the following proteins: (a) N-acetylglucosamine transferase-I (GnT-I) protein; (b) α-1,3 / 6-mannosidase (Man-II) protein; (c) β-1,2-acetylglucosamine transferase (GnT-II) protein; and (d) β-1,4-galactosyltransferase (GalT) protein. In some embodiments, the yeast cell is a *Cotyledon* cell. In some embodiments, the yeast cell is *Cotyledon phafofrum*, *Cotyledon pastoris*, or *Cotyledon pseudopasteurella* cell. In some aspects, the nucleic acid is integrated into the cell's genome. In some aspects, the nucleic acid is not integrated into the cell's genome.

[0023] It should be considered that any embodiment discussed in this specification can be implemented using any method or composition of the disclosed embodiments, and vice versa. Furthermore, the compositions of the embodiments disclosed herein can be used to implement the methods of these embodiments.

[0024] Other objects, features, and advantages of the embodiments disclosed herein will become apparent from the following detailed description. However, it should be understood that while the detailed description and specific examples indicate particular embodiments, they are given by way of example only, as various changes and modifications within the spirit and scope of the embodiments disclosed herein will be apparent to those skilled in the art based on this detailed description. Attached Figure Description

[0025] The following figures form part of this specification and are included to further illustrate certain aspects of this disclosure. This can be better understood by referring to one or more of these figures in conjunction with the detailed description of the specific embodiments presented herein.

[0026] Figure 1 These are images of the protein blot from the supernatant. Lane 1 contained the protein standard Genscript, M00624 (Thermo Fisher Scientific, Waltham, MA, USA). Lane 2 contained lactoferrin from human milk, Sigma Aldrich, SRP6519 (Sigma Aldrich, St. Louis, MO, USA). Lane 3 contained the control (pre-pro-MFα from Saccharomyces cerevisiae). Lane 4 contained the negative control, i.e., the supernatant of unconverted yeast cells. Lanes 5-6 contained the supernatant from yeast cells converted with SP2-lactoferrin. Lanes 7-8 contained the supernatant from yeast cells converted with SP3-lactoferrin. Lanes 9-10 contained yeast cells converted with SP1-lactoferrin.

[0027] Figure 2 This is a bar chart showing protein expression levels. Extracellular proteins are quantified using ELISA. Detailed Implementation

[0028] This document describes the generation of novel synthetic secretion signal peptides. Cells engineered to express one or more exogenous proteins (e.g., human milk proteins) including such signal peptides (e.g., fungal cells, such as yeast cells) are also disclosed. As disclosed herein, in-frame fusion of a “pre-region” sequence from *Pichia pastoris* Ost1 or Pst1 with a “pro-region” sequence from *Saccharomyces cerevisiae* mating factor α or *Pichia pastoris* Epx1 can promote increased extracellular protein production compared to previously used signal peptides. The disclosed signal peptides include, for example, peptides comprising SEQ ID NO: 1, 2, 3, or 4, and peptides comprising 1, 2, 3, 4, or 5 amino acid substitutions (or more) relative to SEQ ID NO: 1, 2, 3, or 4. As described herein, in-frame fusion of these hybrid signal peptides with the N-terminus of mammalian proteins (e.g., human milk proteins such as lactoferrin or α-lactalbumin) promotes efficient protein secretion.

[0029] I. Definition

[0030] The term "biologically active moiety" refers to an amino acid sequence that is less than the full-length amino acid sequence but exhibits at least one activity of the full-length sequence. For example, the biologically active moiety of an enzyme can refer to one or more domains (i.e., catalytic domains) of an enzyme that have catalytic activity. In some aspects, the biologically active moiety of an enzyme is a part of the enzyme that includes the enzyme's catalytic domain. The biologically active moiety of a protein includes peptides or polypeptides that comprise an amino acid sequence sufficiently identical to or derived from the amino acid sequence of the protein, comprising fewer amino acids than the full-length protein, and exhibiting at least one activity of the protein (e.g., enzymatic activity, functional activity, etc.).

[0031] The term "exogenous" refers to any substance introduced into or already introduced into a cell. "Exogenous nucleic acid" is a nucleic acid that has entered or is already inside the cell through the cell membrane. "Exogenous nucleic acid sequence" is the nucleic acid sequence of an exogenous nucleic acid. Exogenous nucleic acids may contain nucleotide sequences present in the cell's natural genome and / or nucleotide sequences not previously present in the cell's genome. Exogenous nucleic acids include exogenous genes. An "exogenous gene" is a nucleic acid that encodes the expression of RNA and / or proteins, which has been introduced into the cell (through, for example, transformation / transfection), also known as a "transgenic gene." Cells containing exogenous nucleic acids are called recombinant cells, into which additional exogenous genes can be introduced. Exogenous genes can originate from the same or different species as the transformed cell. Therefore, exogenous genes can include natural genes that occupy a different location or are subject to different control in the cell's genome relative to their endogenous copies. Multiple copies of an exogenous gene can exist in the cell. Exogenous genes can be maintained in the cell as intercalations within the genome (nucleus, mitochondria, or plastids) or as free molecules.

[0032] "In operable linkage" (or "operably linked") refers to a functional connection between two nucleic acid sequences, such as a control sequence (usually a promoter) and a linking sequence (usually a protein-coding sequence, also called a coding sequence). If a promoter can mediate the transcription of a gene, it is in an operable linkage with the gene.

[0033] The term "natural" refers to the composition of cells or parent cells prior to a transformation event. A "natural gene" (also called an "endogenous gene") is a nucleotide sequence encoding a protein that is not introduced into the cell through a transformation event. A "natural protein" (also called an "endogenous protein") is an amino acid sequence encoded by a natural gene.

[0034] “Recombination” refers to cells, nucleic acids, proteins, or vectors modified by the introduction of exogenous nucleic acids or by alteration of native nucleic acids. The resulting cells, nucleic acids, proteins, or vectors are considered recombinant, and their progeny, offspring, duplication, or replication are also considered recombinant. Thus, for example, recombinant cells may express genes not found in the natural (non-recombinant) form of the cell, or express natural genes in a manner different from how non-recombinant cells express the same gene. Recombinant cells may include, but are not limited to, recombinant nucleic acids encoding gene products or repressive elements, such as mutations, knockouts, antisenses, interfering RNA (RNAi), or dsRNA that reduce the level of active gene products in the cell. “Recombinant nucleic acids” are derived from nucleic acids initially formed in vitro, typically manipulated by, for example, polymerases, ligases, exonucleases, and endonucleases, or in forms uncommon in nature. Once a recombinant nucleic acid is prepared and introduced into a host cell or organism, it can replicate using the host cell's in vivo cellular mechanisms; however, once such a nucleic acid is produced by a recombinant method, it is considered recombinant (for the purposes of this disclosure) even though it subsequently replicates within the cell. Furthermore, recombinant nucleic acids refer to nucleotide sequences that include both endogenous and exogenous nucleotide sequences; therefore, endogenous genes that have undergone recombination with exogenous promoters are recombinant nucleic acids. "Recombinant proteins" are proteins prepared using recombinant technology (i.e., by expressing recombinant nucleic acids).

[0035] "Transformation" refers to the transfer of nucleic acids into a host organism or the host organism's genome. A host organism (and its progeny) containing the transformed nucleic acid fragment is referred to as a "recombinant," "transgenic," or "transformed" organism. Therefore, the isolated polynucleotides of this disclosure can be incorporated into recombinant constructs (typically DNA constructs) that can be introduced into host cells and replicate within them. Such constructs can be vectors comprising replication systems and sequences capable of transcription and translation of sequences encoding polypeptides in a given host cell. Typically, expression vectors include, for example, one or more clonal genes under transcriptional control of 5' and 3' regulatory sequences, as well as selectable markers. Such vectors may also contain promoter regulatory regions (e.g., regulatory regions controlling inducible or constitutive, environmental or developmental, or position-specific expression), transcription start sites, ribosome binding sites, transcription termination sites, and / or polyadenylation signals. Alternatively, cells can be transformed with a single genetic element (e.g., a promoter), which may achieve genetically stable inheritance upon integration into the host organism's genome (e.g., via homologous recombination).

[0036] The term "transformed cell" refers to a cell that has undergone transformation. Therefore, a transformed cell includes both the parent's genome and heritable genetic modifications. Implementation schemes include the offspring and descendants of such transformed cells.

[0037] The term "vector" refers to the way nucleic acids can replicate and / or transfer between organisms, cells, or cellular components. Vectors include plasmids, linear DNA fragments, viruses, bacteriophages, proviruses, phage particles, transposons, and artificial chromosomes, which may or may not be able to replicate autonomously or integrate into the host cell's chromosome.

[0038] The terms “individual,” “subject,” and “patient” are used interchangeably and can refer to humans or non-humans.

[0039] Throughout this application, the term "about" is used to indicate that numerical values ​​include inherent error variations in the measurement or quantitative method.

[0040] When used with the term “including”, the expression “a” or “an” may mean “one”, but it is also consistent with the meanings “one or more”, “at least one” and “one or more”.

[0041] The phrase “and / or” means “and” or “or”. For example, A, B and / or C includes: A alone, B alone, C alone, a combination of A and B, a combination of A and C, a combination of B and C, or a combination of A, B and C. In other words, “and / or” is an inclusive “or”.

[0042] The expressions “comprising” (and any form of inclusion, such as “comprise” and “comprises”), “having” (and any form of having, such as “have” and “has”), “including” (and any form of inclusion, such as “includes” and “include”), or “containing” (and any form of inclusion, such as “contains” and “contain”) are inclusive or open-ended and do not exclude additional, unreferenced elements or method steps.

[0043] The composition and its method of use may “comprise” any ingredient or step disclosed throughout the specification, “consistently constitute” or “made up of”. A composition or method “consistently constitutes any disclosed ingredient or step” limits the scope of the claims to the specified materials or steps that do not substantially affect the essential and novel features of the claimed embodiment.

[0044] II. Proteins and Nucleic Acids

[0045] As used herein, “protein” or “peptide” refers to a molecule comprising at least five amino acid residues. The term “wild-type” as used herein refers to the endogenous form of a molecule naturally occurring in an organism. In some embodiments, the wild-type form of a protein or peptide is used; however, in many embodiments of this disclosure, a modified protein or peptide is used. The terms above are used interchangeably. “Modified protein” or “modified peptide” or “variant” refers to a protein or peptide whose chemical structure, particularly its amino acid sequence, has been altered relative to the wild-type protein or peptide. In some embodiments, the modified / variant protein or peptide has at least one modified activity or function (recognizing that proteins or peptides may have multiple activities or functions). It is particularly considered that a modified / variant protein or peptide may be altered in one activity or function but retain the wild-type activity or function in other respects.

[0046] When proteins are specifically referred to herein, they generally mean native (wild-type) or recombinant (modified) proteins, or optionally proteins in which any signal sequence has been removed. The protein can be isolated directly from its naturally occurring organism, produced by recombinant DNA / exogenous expression methods, or produced by solid-phase peptide synthesis (SPPS) or other in vitro methods. In a specific embodiment, there is an isolated nucleic acid fragment and a recombinant vector incorporated into the nucleic acid sequence encoding the polypeptide. The term "recombinant" can be used with the polypeptide or with the name of a specific polypeptide, which generally refers to a polypeptide produced from a nucleic acid molecule already manipulated in vitro or from a nucleic acid molecule that is a replication product of such a molecule.

[0047] In some embodiments, the size of the protein or polypeptide (wild-type or modified) may include, but is not limited to, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43. 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92 93, 94, 95, 96, 97, 98, 99, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 525, 550, 575, 600, 625, 6 50, 675, 700, 725, 750, 775, 800, 825, 850, 875, 900, 925, 950, 975, 1000, 1100, 1200, 1300, 1400, 1500, 1750, 2000, 2250, 2500 or more amino acid residues, or any range thereof, or derivatives of the corresponding amino sequences described or cited herein. It should be taken into consideration that peptides can be mutated by truncation to make them shorter than their corresponding wild-type forms, and they can also be altered by fusion or conjugation with heterologous protein or peptide sequences having specific functions (e.g., for targeting or localization, for enhancing immunogenicity, for purification purposes, etc.). As used herein, the term “domain” refers to any distinct functional or structural unit of a protein or peptide and generally refers to an amino acid sequence having a structure or function recognizable to those skilled in the art.

[0048] The term "polynucleotide" refers to a recombinant nucleic acid molecule or one isolated from the total genomic nucleic acid. The term "polynucleotide" includes oligonucleotides (nucleic acids of 100 residues or less in length), recombinant vectors, including, for example, plasmids, granules, bacteriophages, viruses, etc. In some respects, polynucleotides include regulatory sequences that are substantially isolated from their naturally occurring gene or protein-coding sequences. Polynucleotides can be single-stranded (coding or antisense) or double-stranded, and can be RNA, DNA (genomic, cDNA, or synthetic), analogues thereof, or combinations thereof. Additional coding or non-coding sequences may, but do not necessarily, be present in the polynucleotide.

[0049] In this regard, the terms "gene," "polynucleotide," or "nucleic acid" are used to refer to nucleic acids (including any sequence required for proper transcription, post-translational modification, or localization) that encode proteins, polypeptides, or peptides. As those skilled in the art will understand, the term includes genomic sequences, expression cassettes, cDNA sequences, and smaller engineered nucleic acid fragments that express or can be engineered to express proteins, polypeptides, domains, peptides, fusion proteins, and mutants. Nucleic acids encoding all or part of a polypeptide may comprise a continuous nucleic acid sequence encoding all or part of such a polypeptide. It should also be considered that a particular polypeptide can be encoded by containing nucleic acids with slightly different variations in nucleic acid sequence but still encoding the same or substantially similar proteins.

[0050] In some embodiments, there are polynucleotide variants that are substantially identical to the sequences disclosed herein; those polynucleotide variants, compared to the polynucleotide sequences of this document provided using the methods described herein (e.g., BLAST analysis using standard parameters), comprise at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity, including all values ​​and ranges therebetween. In some aspects, the isolated polynucleotide will comprise a nucleotide sequence encoding a polypeptide that, over its full length, has at least 90% and in some cases 95% or more identity with the amino acid sequence described herein; or a nucleotide sequence complementary to the isolated polynucleotide.

[0051] Regardless of the length of the coding sequence itself, nucleic acid fragments can be combined with other nucleic acid sequences, such as promoters, polyadenylation signals, additional restriction enzyme sites, multiple cloning sites, and other coding fragments, allowing for significant variations in their total length. The nucleic acid can be of any length. Its length can be, for example, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 75, 100, 125, 175, 200, 250, 300, 350, 400, 450, 500, 750, 1000, 1500, 3000, 5000 nucleotides or more, and / or can contain one or more additional sequences (e.g., regulatory sequences), and / or be part of a larger nucleic acid (e.g., a vector). Therefore, it should be understood that nucleic acid fragments of virtually any length can be used, with their full length preferably limited by the ease of preparation and use in the intended recombinant nucleic acid protocol. In some cases, nucleic acid sequences can encode peptide sequences with additional heterologous coding sequences, for example, to allow for peptide purification, transport, secretion, post-translational modification, or to allow for therapeutic benefits (e.g., targeting or efficacy). As discussed above, tags or other heterologous peptides can be added to the modified peptide coding sequence, where "heterologous" means a peptide different from the modified peptide.

[0052] The polypeptide, protein, or polynucleotide encoding such polypeptide or protein in this disclosure may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 (or any extrapolable range thereof) or more variant amino acids or nucleic acid substitutions of SEQ ID NO: 1-49, or at least 60. %, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% (or any derivable range thereof) are similar, identical, or homologous.Having at least or at most 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 8 1, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 20 3, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 300, 400, 500, 550, 1000 consecutive amino acids or nucleic acids or more, or any deducible range thereof.

[0053] In some embodiments, the protein or polypeptide may include SEQ ID. NO: Amino acids 1-14 or 34-40, 1 to 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 7 9, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 1 45, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266267、268、269、270、271、272、273、274、275、276、277、278、279、280、281、282、283、284、285、286、287、288、289、290、291、292、293、294、295、296、297、298、299、300、301、302、303、304、305、306、307、308、309、310、311、312、313、314、315、316、317、318、319、320、321、322、323、324、325、326、327、328、329、330、331、332、333、334、335、336、337、338、339、340、341、342、343、344、345、346、347、348、349、350、351、352、353、354、355、356、357、358、359、360、361、362、363、364、365、366、367、368、369、370、371、372、373、374、375、376、377、378、379、380、381、382、383、384、385、386、387、388、389、390、391、392、393、394、395、396、397、398、399、400、401、402、403、404、405、406、407、408、409、410、411、412、413、414、415、416、417、418、419、420、421、422、423、424、425、426、427、428、429、430、431、432、433、434、435、436、437、438、439、440、441、442、443、444、445、446、447、448、449、450、451、452、453、454、455、456、457、458、459、460、461、462、463、464、465、466、467、468、469、470、471、472、473、474、475、476、477、478、479、480、481、482、483、484、485、486、487、488、489、490、491、492、493、494、495、496、497、498、499、500、501、502、503、504、505、506、507、508、509、510、511、512、513、514、515、516、517, 518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 557, 558, 559, 560, 561, 562, 563 564, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 61 0, 611, 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 6 57, 658, 659, 660, 661, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699, or 700 (or any derivable range thereof).

[0054] In some implementations, the protein, polypeptide, or nucleic acid may include SEQ ID. NO:1-49 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 8 2, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 14 7, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 2 08, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268269、270、271、272、273、274、275、276、277、278、279、280、281、282、283、284、285、286、287、288、289、290、291、292、293、294、295、296、297、298、299、300、301、302、303、304、305、306、307、308、309、310、311、312、313、314、315、316、317、318、319、320、321、322、323、324、325、326、327、328、329、330、331、332、333、334、335、336、337、338、339、340、341、342、343、344、345、346、347、348、349、350、351、352、353、354、355、356、357、358、359、360、361、362、363、364、365、366、367、368、369、370、371、372、373、374、375、376、377、378、379、380、381、382、383、384、385、386、387、388、389、390、391、392、393、394、395、396、397、398、399、400、401、402、403、404、405、406、407、408、409、410、411、412、413、414、415、416、417、418、419、420、421、422、423、424、425、426、427、428、429、430、431、432、433、434、435、436、437、438、439、440、441、442、443、444、445、446、447、448、449、450、451、452、453、454、455、456、457、458、459、460、461、462、463、464、465、466、467、468、469、470、471、472、473、474、475、476、477、478、479、480、481、482、483、484、485、486、487、488、489、490、491、492、493、494、495、496、497、498、499、500、501、502、503、504、505、506、507、508、509、510、511、512、513、514、515、516、517、518、519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 557, 558, 559, 560, 561, 562, 563, 564, 56 5, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610, 611, 6 12, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658 659, 660, 661, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699, or 700 (or any deducible range thereof) consecutive amino acids.

[0055] In some implementations, the polypeptide, protein, or nucleic acid may include at least, at most, or exactly SEQ ID. NO:1-49, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 14 6, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266267、268、269、270、271、272、273、274、275、276、277、278、279、280、281、282、283、284、285、286、287、288、289、290、291、292、293、294、295、296、297、298、299、300、301、302、303、304、305、306、307、308、309、310、311、312、313、314、315、316、317、318、319、320、321、322、323、324、325、326、327、328、329、330、331、332、333、334、335、336、337、338、339、340、341、342、343、344、345、346、347、348、349、350、351、352、353、354、355、356、357、358、359、360、361、362、363、364、365、366、367、368、369、370、371、372、373、374、375、376、377、378、379、380、381、382、383、384、385、386、387、388、389、390、391、392、393、394、395、396、397、398、399、400、401、402、403、404、405、406、407、408、409、410、411、412、413、414、415、416、417、418、419、420、421、422、423、424、425、426、427、428、429、430、431、432、433、434、435、436、437、438、439、440、441、442、443、444、445、446、447、448、449、450、451、452、453、454、455、456、457、458、459、460、461、462、463、464、465、466、467、468、469、470、471、472、473、474、475、476、477、478、479、480、481、482、483、484、485、486、487、488、489、490、491、492、493、494、495、496、497、498、499、500、501、502、503、504、505、506、507、508、509、510、511、512、513、514、515、516、517, 518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 557, 558, 559, 560, 561, 562, 563, 5 64, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610, 611 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658 659, 660, 661, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 697, 698, 699, or 700 (or any deducible range thereof) consecutive amino acids, which, together with SEQ ID NO, constitute SEQ ID NO. NO:1-49 has at least, at most, or exactly 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% (or any derivable range therein) similarity, identity, or common origin.

[0056] In some respects, a nucleic acid molecule or polypeptide begins with SEQ ID. Positions of any sequence from NO:1-49: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267268、269、270、271、272、273、274、275、276、277、278、279、280、281、282、283、284、285、286、287、288、289、290、291、292、293、294、295、296、297、298、299、300、301、302、303、304、305、306、307、308、309、310、311、312、313、314、315、316、317、318、319、320、321、322、323、324、325、326、327、328、329、330、331、332、333、334、335、336、337、338、339、340、341、342、343、344、345、346、347、348、349、350、351、352、353、354、355、356、357、358、359、360、361、362、363、364、365、366、367、368、369、370、371、372、373、374、375、376、377、378、379、380、381、382、383、384、385、386、387、388、389、390、391、392、393、394、395、396、397、398、399、400、401、402、403、404、405、406、407、408、409、410、411、412、413、414、415、416、417、418、419、420、421、422、423、424、425、426、427、428、429、430、431、432、433、434、435、436、437、438、439、440、441、442、443、444、445、446、447、448、449、450、451、452、453、454、455、456、457、458、459、460、461、462、463、464、465、466、467、468、469、470、471、472、473、474、475、476、477、478、479、480、481、482、483、484、485、486、487、488、489、490、491、492、493、494、495、496、497、498、499、500、501、502、503、504、505、506、507、508、509、510、511、512、513、514、515、516、517、518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 557, 558, 559, 560, 561, 562, 563, 56 4, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610, 6 11, 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658, 659, 660, 661, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699, or 700, and including at least, at most, or exactly SEQ. Any sequence of ID NO: 1-49, including the numbers 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82.83、84、85、86、87、88、89、90、91、92、93、94、95、96、97、98、99、100、101、102、103、104、105、106、107、108、109、110、111、112、113、114、115、116、117、118、119、120、121、122、123、124、125、126、127、128、129、130、131、132、133、134、135、136、137、138、139、140、141、142、143、144、145、146、147、148、149、150、151、152、153、154、155、156、157、158、159、160、161、162、163、164、165、166、167、168、169、170、171、172、173、174、175、176、177、178、179、180、181、182、183、184、185、186、187、188、189、190、191、192、193、194、195、196、197、198、199、200、201、202、203、204、205、206、207、208、209、210、211、212、213、214、215、216、217、218、219、220、221、222、223、224、225、226、227、228、229、230、231、232、233、234、235、236、237、238、239、240、241、242、243、244、245、246、247、248、249、250、251、252、253、254、255、256、257、258、259、260、261、262、263、264、265、266、267、268、269、270、271、272、273、274、275、276、277、278、279、280、281、282、283、284、285、286、287、288、289、290、291、292、293、294、295、296、297、298、299、300、301、302、303、304、305、306、307、308、309、310、311、312、313、314、315、316、317、318、319、320、321、322、323、324、325、326、327、328、329、330、331、332、333、334、335、336、337、338、339、340、341、342、343、344、345、346、347、348、349、350、351、352、353、354、355、356、357、358、359、360、361、362、363、364、365、366、367、368、369、370、371、372、373、374、375、376、377、378、379、380、381、382、383、384、385、386、387、388、389、390、391、392、393、394、395、396、397、398、399、400、401、402、403、404、405、406、407、408、409、410、411、412、413、414、415、416、417、418、419、420、421、422、423、424、425、426、427、428、429、430、431、432、433、434、435、436、437、438、439、440、441、442、443、444、445、446、447、448、449、450、451、452、453、454、455、456、457、458、459、460、461、462、463、464、465、466、467、468、469、470、471、472、473、474、475、476、477、478、479、480、481、482、483、484、485、486、487、488、489、490、491、492、493、494、495、496、497、498、499、500、501、502、503、504、505、506、507、508、509、510、511、512、513、514、515、516、517、518、519、520、521、522、523、524、525、526、527、528、529、530、531、532、533、534、535、536、537、538、539、540、541、542、543、544、545、546、547、548、549、550、551、552、553、554、555、556、557、558、559、560、561、562、563、564、565、566、567、568、569、570、571、572、573、574、575、576、577、578、579、580、581、582、583、584、585、586、587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610, 611, 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658, 659, 660, 661, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699, or 700 (or any deducible range thereof) consecutive amino acids or nucleotides.

[0057] Nucleotides, as well as protein, polypeptide, and peptide sequences of various genes, have been previously published and can be found in recognized computerized databases. Two commonly used databases are the National Center for Biotechnology Information's Genbank and GenPept databases (ncbi.nlm.nih.gov / on the World Wide Web) and the Universal Protein Resource (UniProt; uniprot.org on the World Wide Web). The coding regions of these genes can be amplified and / or expressed using the techniques disclosed herein or techniques known to those skilled in the art.

[0058] It should be considered that the compositions of this disclosure contain about 0.001 mg to about 10 mg of total polypeptides, peptides and / or proteins per milliliter. The concentration of protein in the composition may be about, at least about or at most about 0.001, 0.010, 0.050, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.5, 2.0, 2.5, 3.0, 3.5, 4.0, 4.5, 5.0, 5.5, 6.0, 6.5, 7.0, 7.5, 8.0, 8.5, 9.0, 9.5, 10.0 mg / ml or more (or any of the derived ranges thereof).

[0059] In the case of proteins with catalytic activity (e.g., enzymes), the Enzyme Classification (EC) nomenclature can be used to describe such proteins. EC classifications of various enzymes have been previously published and can be found in recognized databases, such as the ENZYME database (Bairoch A. The ENZYME database in 2000. Nucleic Acids Res. 2000 Jan 1; 28(1):304-5. doi:10.1093 / nar / 28.1.304; the entire contents of which are incorporated herein by reference).

[0060] A. Signal peptide

[0061] This disclosure relates to synthetic signal peptides, and to polynucleotides and nucleic acids encoding such signal peptides. Cells comprising such signal peptides are also disclosed, as well as methods for utilizing cells to produce and secrete proteins (e.g., mammalian proteins, such as human milk proteins). As used herein, a “signal peptide” (or “signal peptide sequence”) describes any peptide that, when present at the N-terminus of a newly synthesized polypeptide, can direct the polypeptide across or into the cell membrane (e.g., plasma membrane, endoplasmic reticulum membrane, etc.). In some aspects, the signal peptides of this disclosure can direct the polypeptide into the cell via a secretory pathway and subsequently secrete the polypeptide (referred to herein as a “secreting signal peptide”).

[0062] As described herein, aspects of this disclosure relate to synthetic signal peptides, including:

[0063] (a) From the following pre-region sequences:

[0064] (i) Pichia pastoris Ost1; or

[0065] (ii) Pichia pastoris Pst1; and

[0066] (b) From the following original-region sequences:

[0067] (i) Saccharomyces cerevisiae mating factor α (MFα); or

[0068] (ii) Pichia pastoris Epx1.

[0069] Some of the signal peptides disclosed in this disclosure are described in Table 1 below.

[0070] Table 1 - Signal Peptides

[0071]

[0072]

[0073] In some respects, a polypeptide comprising the signal peptide of this disclosure is disclosed. A nucleic acid encoding such a polypeptide is also disclosed. Further disclosed are cells expressing the polypeptide comprising the signal peptide of this disclosure.

[0074] In some aspects, the polypeptide of this disclosure includes SEQ ID NO:1. In some embodiments, the polypeptide of this disclosure includes a sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with SEQ ID NO:1. In some aspects, the polypeptide of this disclosure includes a sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions (or more) relative to SEQ ID NO:1.

[0075] In some aspects, the polypeptide of this disclosure includes SEQ ID NO:2. In some embodiments, the polypeptide of this disclosure includes a sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with SEQ ID NO:2. In some aspects, the polypeptide of this disclosure includes a sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions (or more) relative to SEQ ID NO:2.

[0076] In some aspects, the polypeptide of this disclosure includes SEQ ID NO:3. In some embodiments, the polypeptide of this disclosure includes a sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with SEQ ID NO:3. In some aspects, the polypeptide of this disclosure includes a sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions (or more) relative to SEQ ID NO:3.

[0077] In some aspects, the polypeptide of this disclosure includes SEQ ID NO:4. In some embodiments, the polypeptide of this disclosure includes a sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with SEQ ID NO:4. In some aspects, the polypeptide of this disclosure includes a sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions (or more) relative to SEQ ID NO:4.

[0078] Any one or more signal peptides disclosed herein may be excluded from certain implementations.

[0079] B. Secretory proteins

[0080] This disclosure includes aspects of secretory proteins (also known as "secreted proteins"), compositions including secretory proteins, methods of expressing secretory proteins, and methods of using them. As used herein, "secretory protein" describes any protein secreted extracellularly. In some cases, the secretory proteins of this disclosure are proteins present in human secretions (e.g., colostrum, milk, tears, semen, vaginal fluid, saliva, or other secretions). In some aspects, the secretory proteins of this disclosure are human milk proteins. In some aspects, the secretory proteins of this disclosure are not human milk proteins.

[0081] 1. Human milk protein

[0082] This disclosure includes human milk proteins, compositions comprising human milk proteins (e.g., infant formula compositions), methods of producing human milk proteins, and methods of using them. In some aspects, the disclosure discloses cells expressing human milk proteins linked to a signal peptide (e.g., including SEQ ID NO: 1, 2, 3, or 4) of this disclosure. As used herein, “human milk protein” describes any protein present in human breast milk. Human milk proteins include proteins derived from (e.g., isolated from) human breast milk, as well as any protein having the amino acid sequence of proteins present in human breast milk, produced by other methods (e.g., recombinant expression, chemical synthesis, etc.). Various human milk proteins are recognized in the art and are considered herein. Human milk proteins considered herein include, but are not limited to, secretory IgA (sIgA), human serum albumin, xanthine dehydrogenase, lactoferrin, lactoperoxidase, lactolipoprotein, lactobacin, adiponectin, β-casein, κ-casein, leptin, lysozyme, and α-lactalbumin. In some embodiments, the human milk protein of this disclosure is human whey protein. In some embodiments, the human milk protein of this disclosure is a recombinant human milk protein (e.g., produced by non-mammalian cells such as yeast cells).

[0083] Certain aspects of this disclosure relate to human milk proteins having “human-like” glycans. Human-like glycans (also referred to as “human-like glycan structures”) describe glycans having structures present in human glycoproteins. These glycans include, for example, hybrid N-glycans, complex N-glycans, bi-antennary, tri-antennary, and tetra-antennary N-glycans, as well as glycans comprising sialic acid, galactose, N-acetylgalactosamine, or fucose. Human-like glycans include those having a Man3GlcNAc2 core structure. Therefore, the human milk proteins of this disclosure include those proteins having one or more human-like glycans (e.g., hybrid N-glycans, complex N-glycans, bi-antennary N-glycans, tri-antennary N-glycans, tetra-antennary N-glycans, and combinations thereof).

[0084] Therefore, in some embodiments, recombinant human milk proteins comprising one or more human-like glycans (e.g., recombinant human lactoferrin) are disclosed. Such recombinant proteins include, for example, recombinant proteins produced by engineered mammals, fungi, yeast, bacteria, or other cells (including engineered cells described elsewhere herein). In some aspects, such recombinant proteins have a glycan pattern different from that of the corresponding native human milk proteins. For example, in some embodiments, recombinant human lactoferrin comprising one or more human-like glycans is disclosed, wherein the lactoferrin has a glycan pattern different from that of any naturally occurring human lactoferrin (e.g., human lactoferrin in human breast milk).

[0085] a. Lactoferrin

[0086] This disclosure relates to lactoferrin and compositions comprising lactoferrin, including infant formula compositions. In some aspects, it discloses cells expressing human lactoferrin linked to a signal peptide (e.g., including SEQ ID NO: 1, 2, 3, or 4). Lactoferrin (also known as “lactoferrin”) is a whey protein found in exocrine fluids such as breast milk and encoded by the LTF gene. Without wishing to be bound by theory, lactoferrin is understood to possess antibacterial and anti-inflammatory properties. Certain aspects of this disclosure relate to human lactoferrin (UniProtKB / Swiss-Prot accession number P02788), including its subtypes. The complete sequence of human lactoferrin (including the signal peptide) is provided in the form of SEQ ID NO: 34. The sequence of mature human lactoferrin (after cleavage of the signal peptide) is provided in the form of SEQ ID NO: 9.

[0087] Table 2 - Human lactoferrin sequence

[0088]

[0089]

[0090]

[0091]

[0092]

[0093]

[0094]

[0095]

[0096]

[0097]

[0098]

[0099]

[0100]

[0101] In some aspects, the human lactoferrin of this disclosure is recombinant human lactoferrin (rhLactoferrin). In some aspects, the recombinant human lactoferrin of this disclosure is obtained from mammals, fungi, yeast, bacteria, or other cells. In some aspects, the recombinant human lactoferrin of this disclosure is not obtained from mammalian cells. In some aspects, the recombinant human lactoferrin of this disclosure is obtained from fungal cells. Fungal cells may be, for example, cells of *Arxula*, *Aspergillus*, *Schizochytrium*, *Candida*, *Clavicipitaceae*, *Cryptococcus*, *Hypericum*, *Geotrichum*, *Hansenula*, *Kluyveromyces*, *Kodakosomal*, *Codonopsis*, *Leucobacter*, *Leucobacterium*, *Oligospora*, *Pichia pastoris*, *Protocol*, *Rhizopus*, *Rhodotorula*, *Saccharomyces*, *Schizosaccharomyces*, *Tremella*, *Trichosporium*, *Wickhamia*, or *Yarlocha*. In some aspects, the fungal cells are yeast cells. In some respects, the yeast cells are cells of the genus *Cotyledon* (e.g., *Cotyledon phaffeoides*, *Cotyledon pastoris*, *Cotyledon pseudopasteuroides*). Additional cells suitable for recombinant protein production are recognized in the art and are considered herein. In some respects, the recombinant human lactoferrin of this disclosure is obtained from bacterial cells. In other respects, the human lactoferrin of this disclosure is isolated from natural sources.

[0102] Specific aspects of this disclosure relate to human lactoferrin having at least one hybrid or complex N-glycan. In some aspects, human lactoferrin comprises a glycan containing one or more of sialic acid, galactose, N-acetylgalactosamine, or fucose. In some aspects, human lactoferrin comprises biantennary, triantennary, or tetraantennary N-glycans. As disclosed herein, human lactoferrin having one or more hybrid, complex, biantennary, triantennary, or tetraantennary N-glycans can be used, for example, in infant formula or other nutritional compositions or supplements.

[0103] b. α-lactalbumin

[0104] This disclosure relates to alpha-lactalbumin and compositions comprising alpha-lactalbumin, including infant formula compositions. In some aspects, it discloses cells expressing human alpha-lactalbumin linked to a signal peptide (e.g., including SEQ ID NO: 1, 2, 3, or 4). Alpha-lactalbumin (also known as α-lactalbumin) is a whey protein found in breast milk and encoded by the LALBA gene. Certain aspects of this disclosure relate to human alpha-lactalbumin (UniProtKB / Swiss-Prot accession number P00709), including its subtypes. The complete sequence of human alpha-lactalbumin (including the signal peptide) is provided in the form of SEQ ID NO: 36. The sequence of mature human alpha-lactalbumin (after cleavage of the signal peptide) is provided in the form of SEQ ID NO: 35.

[0105] Table 3 - Human lactoferrin sequence

[0106]

[0107]

[0108] In some aspects, the human α-lactalbumin of this disclosure is recombinant human α-lactalbumin. In some aspects, the recombinant human α-lactalbumin of this disclosure is obtained from mammals, fungi, yeast, bacteria, or other cells. In some aspects, the recombinant human α-lactalbumin of this disclosure is not obtained from mammalian cells. In some aspects, the recombinant human α-lactalbumin of this disclosure is obtained from yeast cells. The yeast cells may be, for example, cells of *Arxula*, *Aspergillus*, *Schizochytrium*, *Candida*, *Cryptococcus*, *Cryptococcus*, *Hypericum*, *Gastromycium*, *Hansenula*, *Kluyveromyces*, *Kodakosomalia*, *Codonopsis*, *Leucobacter*, *Lactobacillus*, *Oligococcus*, *Pichia pastoris*, *Protocol*, *Rhizopus*, *Rhodotorula*, *Saccharomyces*, *Schizophys*, *Tremella*, *Trichosporium*, *Wickhamia*, or *Yarlosporium*. In some respects, the yeast cell is a cell of the genus *Cotyledon* (e.g., *Cotyledon phaf.*, *Cotyledon pastoris*, *Cotyledon pseudo-Cotyledon pastoris*). Additional yeast cells suitable for recombinant protein production are recognized in the art and are considered herein. In some respects, the recombinant human α-lactalbumin of this disclosure is obtained from bacterial cells. In other respects, the human α-lactalbumin of this disclosure is isolated from a natural source.

[0109] Specific aspects of this disclosure relate to human alpha-lactalbumin having at least one hybrid or complex N-glycan. In some aspects, human alpha-lactalbumin includes a glycan containing one or more of sialic acid, galactose, N-acetylgalactosamine, or fucose. In some aspects, human lactoferrin includes biantennary, triantennary, or tetraantennary N-glycans. As disclosed herein, human alpha-lactalbumin having one or more hybrid, complex, biantennary, triantennary, or tetraantennary N-glycans can be used, for example, in infant formula or other nutritional compositions or supplements.

[0110] c. Additional human milk protein

[0111] Additional human milk proteins considered in the compositions (e.g., infant formula compositions) and methods of this disclosure include, but are not limited to, secretory IgA (sIgA), human serum albumin, xanthine dehydrogenase, lactoperoxidase, lactolipoprotein, lactobacin, adiponectin, β-casein, κ-casein, leptin, osteopontin, bile salt-stimulated lipase (BSSL), and lysozyme. Any one or more of these human milk proteins may be included in the compositions of this disclosure (e.g., infant formula). In some embodiments, any one or more of these human milk proteins may be excluded.

[0112] CN-acetylglucosamine transferase

[0113] This disclosure relates to N-acetylglucosamine transferase proteins. As used herein, “N-acetylglucosamine transferase protein” describes any polypeptide having N-acetylglucosamine transferase activity. N-acetylglucosamine transferase describes an enzyme that catalyzes the transfer of a monosaccharide from a specific glyconucleotide donor to a specific hydroxyl position of the monosaccharide in a growing glycan chain, via one of two possible anomeric linkages (α or β).

[0114] N-acetylglucosamine transferase proteins can be derived from any suitable organism. In some respects, N-acetylglucosamine transferase proteins are eukaryotic N-acetylglucosamine transferase proteins. In other respects, N-acetylglucosamine transferase proteins are mammalian N-acetylglucosamine transferase proteins.

[0115] 1. N-acetylglucosamine transferase I

[0116] In some embodiments, the N-acetylglucosamine transferase protein is N-acetylglucosamine transferase I protein (EC 2.4.1.101). The systematic name for this enzyme class is α-1,3-mannosyl-glycoprotein beta-1,2-N-acetylglucosaminyltransferase. Other names include: GnT-I, N-acetylglucosamine transferase I, and Uridine diphosphoacetylglucosamine-alpha-1,3-mannosylglycoprotein beta-1,2-N-acetylglucosaminyltransferase. In some embodiments, the N-acetylglucosamine transferase I protein of this disclosure is human (Homo Sapiens) GnT-I; however, N-acetylglucosamine transferase I proteins from any eukaryotic organism can be used as part of the methods and compositions of this disclosure.

[0117] 2. β-1,2-N-acetylglucosamine transferase

[0118] In some embodiments, the N-acetylglucosamine transferase protein is β-1,2-N-acetylglucosamine transferase protein (EC 2.4.1.143). The systematic name for this enzyme class is α-1,6-mannosyl-glycoprotein 2-β-N-acetylglucosaminyltransferase. Other names include GnT-II, N-acetylglucosamine transferase II, and Uridine diphosphoacetylglucosamine-alpha-1,6-mannosylglycoprotein beta-1-2-N-acetylglucosaminyltransferase. In some embodiments, the β-1,2-N-acetylglucosamine transferase protein of this disclosure is brown rat (Rattus norvegicus) GnT-II; however, β-1,2-N-acetylglucosamine transferase proteins from any eukaryotic organism may be used as part of the methods and compositions of this disclosure.

[0119] D. α-1,3 / 6-Mannosidase

[0120] This disclosure relates to α-1,3 / 6-mannosidase protein (EC 3.2.114). As used herein, “α-1,3 / 6-mannosidase protein” (or alpha-1,3 / 6-mannosidase protein) describes any polypeptide possessing α-1,3 / 6-mannosidase activity. α-1,3 / 6-mannosidase describes an enzyme that catalyzes the removal of two mannose residues from an N-glycan. The systematic name for this class of enzymes is Mannosyl-oligosaccharide 1,3-1,6-alpha-mannosidase. Other names include Man-II and Mannosidase II. α-1,3 / 6-mannosidase protein can originate from any suitable organism. In some embodiments, the α-1,3 / 6-mannosidase protein is a eukaryotic α-1,3 / 6-mannosidase protein. In some embodiments, the α-1,3 / 6-mannosidase protein is from Drosophila melanogaster Man-II; however, α-1,3 / 6-mannosidase proteins from any eukaryotic organism can be used as part of the methods and compositions of this disclosure.

[0121] E. α-1,2-mannosidase

[0122] This disclosure relates to α-1,2-mannosidase protein (EC 3.2.1.130). As used herein, "α-1,2-mannosidase protein" (or alpha-1,2-mannosidase protein) describes any polypeptide possessing α-1,2-mannosidase activity. The systematic name for this enzyme class is glycoprotein endo-alpha-1,2-mannosidase. Other names include endo-α-D-mannosidase and Man-I. In some embodiments, the α-1,2-mannosidase protein is fungal Man-I. In some embodiments, Man-I is *Trichoderma reesei* Man-I.

[0123] F. β-1,4-galactosyltransferase

[0124] This disclosure relates to β-1,4-galactosyltransferase protein (EC 2.4.1.38). As used herein, “β-1,4-galactosyltransferase protein” describes any polypeptide possessing β-1,4-galactosyltransferase activity. The systematic name for this enzyme class is β-N-acetylglucosaminylglycopeptide beta-1,4-galactosyltransferase. Other names include: glycoprotein 4-beta-galactosyltransferase, UDP-galactose-glycoprotein galactosyltransferase, and GalT. In some embodiments, the β-1,4-galactosyltransferase protein is mammalian GalT. In some embodiments, GalT is human GalT.

[0125] Glycosylated proteins

[0126] This disclosure relates to methods and compositions for producing glycosylated proteins (also known as "glycoproteins") having a glycosylation pattern similar to that of glycoproteins produced by human cells. In some embodiments, the glycoproteins of this disclosure are N-linked glycoproteins. N-linked glycoproteins contain an N-acetylglucosamine residue linked to the amide nitrogen of an asparagine residue in the protein. The main sugars present on the glycoprotein are glucose, galactose, mannose, fucose, N-acetylgalactosamine (GalNAc), N-acetylglucosamine (GlcNAc), and sialic acid, such as N-acetyl-neuraminic acid (NANA). Glycosyl processing occurs in the endoplasmic reticulum lumen via co-translation and continues in the Golgi apparatus to form N-linked glycoproteins.

[0127] H. Protein Targeting

[0128] Certain aspects of this disclosure include cells expressing one or more proteins by nucleic acid molecules, wherein the proteins are targeted to desired subcellular locations (e.g., organelles, such as the Golgi apparatus). In some cases, proteins are targeted to subcellular locations by forming fusion proteins comprising a portion of a protein (e.g., the catalytic domain of an enzyme) and a cell-targeting signal peptide (e.g., a heterologous signal peptide, such as a signal peptide comprising SEQ ID NO: 1, 2, 3, or 4), which is typically not linked to or bound to the portion of the protein. Fusion proteins may be encoded by polynucleotides encoding a cell-targeting signal peptide that are linked within the same translation reading frame (“frame”) to a nucleic acid fragment encoding a protein (e.g., an enzyme) or a fragment of its catalytic activity.

[0129] The targeting signal peptide component of the fusion construct or protein can be derived from membrane-bound proteins, recovery signals, type II membrane proteins, type I membrane proteins, transmembrane nucleotide sugar transporters, mannosidases, sialic acid transferases, glucosidases, mannosyltransferases, and phosphomannosyltransferases of the endoplasmic reticulum or Golgi apparatus. In some respects, the targeting signal peptide is a Golgi localization tag. Examples of Golgi localization tags include, but are not limited to, transmembrane domains from *Saccharomyces cerevisiae* Kre2p, *Saccharomyces cerevisiae* Mnn2p, *Saccharomyces cerevisiae* Mnn9, *Phaeodactyces phalloides* Bmt2, *Phaeodactyces phalloides* Bmt3, or *Phaeodactyces phalloides* Ktr2.

[0130] III. Sequence

[0131] Some of the example peptide and nucleic acid sequences conceived in this article are shown in Table 4 below.

[0132] Table 4

[0133]

[0134]

[0135]

[0136]

[0137]

[0138]

[0139]

[0140]

[0141]

[0142]

[0143]

[0144]

[0145]

[0146]

[0147]

[0148] IV. Genetic Engineering

[0149] According to this disclosure, vectors for transforming microorganisms (e.g., fungal cells, yeast cells) can be prepared using known techniques familiar to those skilled in the art. Vectors typically contain one or more genes, each encoding the expression of a desired product (gene product), and are operatively linked to one or more control sequences that regulate gene expression in recombinant cells or target the gene product to a specific location.

[0150] Exogenous nucleic acid sequences (including, for example, nucleic acid sequences encoding fusion proteins, nucleic acid sequences encoding wild-type or mutant proteins) can be introduced into many different host cells. As further described herein, nucleic acid sequences configured to promote genetic mutations in genes can also be introduced into a variety of host cells. Suitable host cells are microbial hosts widely found in the fungi family. Examples of suitable host strains include, but are not limited to, fungal or yeast species, such as *Arxula*, *Aspergillus*, *Schizochytrium*, *Candida*, *Clavicipitaceae*, *Cryptococcus*, *Hypericum*, *Hansenula*, *Kluyveromyces*, *Codonopsis*, *Leucobacter*, *Oligosaccharomyces*, *Pichia pastoris*, *Protocol*, *Rhizopus*, *Rhodotorula*, *Saccharomyces*, *Schizophytes*, *Tremella*, *Trichosporium*, and *Yarlium*. In some embodiments, the host cell of this disclosure is *Codonopsis* cells. In some embodiments, the host cell of this disclosure is *Codonopsis phafuncsis*. In some embodiments, the host cell of this disclosure is *Pasteurella pastoris*. In some embodiments, the host cell of this disclosure is *Pseudopasteurella pastoris*.

[0151] Microbial expression systems and expression vectors are well known to those skilled in the art. Any such expression vector can be used to introduce immediate gene and nucleic acid sequences into an organism. Nucleic acid sequences can be introduced into suitable microorganisms through transformation techniques. For example, nucleic acid sequences can be cloned in suitable plasmids, and parental cells can be transformed with the resulting plasmids. There are no particular limitations on plasmids, as long as they make the desired nucleic acid sequence heritable for the offspring of the microorganism.

[0152] Vectors or cassettes used to transform suitable host cells are well-established in the art. Typically, a vector or cassette contains a gene, a sequence (including a promoter) that directs the transcription and translation of the associated gene, optional markers, and a sequence that allows autonomous replication or chromosomal integration. Suitable vectors include a 5' region of the gene containing a promoter and other transcription initiation control elements, and a 3' region of the DNA fragment that controls transcription termination.

[0153] The promoter, cDNA, and 3'UTR, as well as other elements of the vector, can be generated using cloning techniques from fragments isolated from natural sources (Green & Sambrook, Molecular Cloning: A Laboratory Manual, (4th ed., 2012); US Pat. No. 4, 683, 202; cited herein). Alternatively, elements can be generated synthetically using known methods (Gene 164:49-53 (1995)).

[0154] A. Carrier and carrier components

[0155] According to this disclosure, vectors for transforming microorganisms (e.g., yeast cells) can be prepared using known techniques familiar to those skilled in the art. Vectors typically contain one or more genes, each encoding the expression of a desired product (gene product), and are operatively linked to one or more control sequences (e.g., promoter sequences, signal peptide sequences) that regulate gene expression or target the gene product to a specific location in recombinant cells.

[0156] 1. Control sequence

[0157] Control sequences are nucleic acid sequences that regulate the expression of coding sequences or direct gene products to specific locations inside or outside the cell. Control sequences regulating expression include, for example, promoters that regulate transcription of coding sequences and terminators that terminate transcription of coding sequences. Another control sequence is the 3' untranslated sequence located at the end of a coding sequence encoding a polyadenylation signal. Control sequences directing gene products to specific locations include sequences encoding signal peptides, which direct the proteins they link to to specific locations inside or outside the cell.

[0158] Therefore, example vectors designed for gene expression in microorganisms contain coding sequences of the desired gene product (e.g., selectable markers, enzymes, fusion proteins, etc.) operatively linked to an active promoter in yeast. Alternatively, if the vector does not contain a promoter operatively linked to the target coding sequence, the coding sequence can be transformed into the cell such that it is operatively linked to an endogenous promoter at the point of vector integration. Example promoters conceived herein include, but are not limited to, the AOX1, GAP, TEF1, TPI1, DAS1, DAS2, CAT1, and FMD promoters.

[0159] The promoter used to express a gene can be a naturally occurring promoter linked to the gene, or it can be a different promoter.

[0160] Promoters can typically be characterized as constitutive or inducible. Constitutive promoters are generally active or functional to drive expression at the same level consistently (or at some point in the cell's life cycle). In contrast, inducible promoters are active (or become inactive) only in response to stimuli, or are significantly upregulated or downregulated. Both types of promoters can be used in the methods disclosed. Useful inducible promoters include those that mediate the transcription of operable linker genes in response to stimuli such as exogenously provided small molecules, temperature (hot or cold), nitrogen deficiency in the culture medium, etc. A suitable promoter can activate the transcription of substantially silent genes or upregulate the transcription of operable linker genes with low levels of transcription.

[0161] Including the termination control sequence is optional. The termination region can be a natural region of the transcription start region (promoter), a natural region of the target DNA sequence, or can be obtained from other sources (see, for example, Chen & Orozco, Nucleic Acids Research 16:8411 (1988)).

[0162] In some cases, the full nucleotide sequence of the promoter is not necessary to drive transcription; shorter sequences than the full nucleotide sequence of the promoter can drive the transcription of operable genes. The smallest part of the promoter (called the core promoter) includes the transcription start site, the RNA polymerase binding site, and the transcription factor binding site.

[0163] By introducing a promoter and a target into a nucleic acid molecule (e.g., a vector), the promoter and target can be linked. The vector can then be introduced into a cell to express the promoter and target. In one embodiment, the promoter is integrated into the cell's genome by introducing the promoter into the cell's DNA and linking the promoter and target (e.g., through homologous recombination).

[0164] B. Gene and codon optimization

[0165] Typically, a gene includes a promoter, a coding sequence, and a termination control sequence. When assembled using recombinant DNA technology, the gene can be called an expression cassette, and its flanking parts can be restriction sites to facilitate insertion into a vector used to introduce the recombinant gene into the host cell. The flanking parts of the expression cassette can be DNA sequences from the genome or other nucleic acid targets to facilitate stable integration of the expression cassette into the genome via homologous recombination. Alternatively, the vector and its expression cassette can remain in an unintegrated state (e.g., in a free form), in which case the vector typically includes an origin of replication to ensure the replication of the vector DNA.

[0166] The common genes present on vectors are protein-coding genes, the expression of which allows differentiation between recombinant cells containing the protein and cells that do not express it. Such genes and their corresponding gene products are called selectable markers or selection markers. Any of a variety of selectable markers can be used in transgenic constructs that can be used to transform organisms covered in the disclosed embodiments.

[0167] To achieve optimal expression of recombinant proteins, it may be beneficial to use a coding sequence as described below, which produces mRNA with codons best suited for use by the cells to be transformed. Therefore, proper expression of the transgene may require that the codon usage of the transgene matches the specific codon preference of the organism in which it is expressed. The exact mechanisms behind this effect are numerous, but include an appropriate balance between the available pool of aminoacylated tRNAs and the proteins synthesized in the cell, and more efficient translation of the transgene messenger RNA (mRNA) when this requirement is met. When codon usage in the transgene is not optimized, the available tRNA pool may be insufficient to allow for efficient translation of the transgene mRNA, leading to ribosome arrest and termination, and potentially causing instability of the transgene mRNA.

[0168] The coding sequences of this disclosure can be codon-optimized for a specific host cell by replacing one or more rare codons with one or more codons that are more frequently present in the host cell. Rare codons in the host cell describe codons that make up less than 5%, less than 10%, or less than 20% of the coding sequence in the host cell. Rare codons can be identified using methods known to those skilled in the art.

[0169] Aspects of this disclosure include transforming microorganisms using nucleic acid sequences containing genes encoding proteins. This gene may be native to the cell or derived from a different species. The gene may originate from a different species but be modified (e.g., codon optimization) to achieve optimal expression in the microorganism. In some embodiments, the gene is heritable to the offspring of the transformed cells. In some embodiments, the gene is heritable because it resides on a plasmid. In some embodiments, the gene is heritable because it is integrated into the genome of the transformed cells.

[0170] Other aspects of this disclosure may include transforming microorganisms using nucleic acid sequences configured to produce mutations in the genes of the microorganisms. For example, aspects of this disclosure may include transforming microorganisms using nucleic acid sequences comprising upstream and downstream sequences of a gene (e.g., the OCH1 gene) to facilitate reduced gene expression or gene deletion through homologous recombination. Various methods for producing mutations (including deletion or knockout mutations, and mutations that reduce gene expression) in microbial genes are recognized in the art and are considered herein. Microorganisms with gene deletion or knockout mutations do not produce a functional copy of the protein. For example, recombinant yeast cells of this disclosure may include the deletion of the endogenous OCH1 gene, such that the recombinant yeast cells do not express the endogenous, functional OCH1 protein. Microorganisms with reduced gene or protein expression produce a functional copy of the protein, but in a reduced number compared to wild-type (i.e., non-recombinant or non-genetically modified) microorganisms of the same species. Methods for reducing protein expression are recognized in the art, including, for example, replacing an endogenous promoter and / or modifying one or more regulatory elements.

[0171] C. Transformation

[0172] Cells can be transformed using any suitable technique, including gene gun, electroporation, glass bead transformation, and silicon carbide whisker transformation. The embodiments disclosed herein can employ any convenient technique for introducing transgenes into microorganisms.

[0173] Vectors for microbial transformation can be prepared using known techniques familiar to those skilled in the art. In one embodiment, an exemplary vector for expressing a gene in a microorganism is designed to contain a gene encoding an enzyme, operatively linked to an active promoter in the microorganism. Alternatively, if the vector does not contain a promoter operatively linked to the target gene, the gene can be transformed into a cell such that it is operatively linked to a native promoter at the point of vector integration. The vector may also contain a second gene encoding a protein. Optionally, one or both genes are followed by a 3' untranslated sequence containing a polyadenylation signal. Expression cassettes encoding both genes may be physically linked in the vector or located on separate vectors. Co-transformation of microorganisms can also be used, where different vector molecules are used simultaneously to transform cells (Protist 155:381-93 (2004)). Transformed cells may be selectively chosen based on their ability to grow in the presence of antibiotics or other optional markers, provided that cells lacking resistance cassettes will not grow.

[0174] D. Genetically engineered cells

[0175] Various aspects of this disclosure include genetically engineered cells (also referred to as "engineered cells" or "recombinant cells") and methods for preparing and using such cells. In some embodiments, recombinant cells comprising one or more exogenous nucleic acid sequences are disclosed. Methods for generating such recombinant cells are also disclosed, including introducing one or more exogenous nucleic acid sequences into a host cell. Further described are methods for collecting one or more products (e.g., mammalian proteins) from such recombinant cells, including culturing the cells and collecting the products.

[0176] In some embodiments, the recombinant cells are prokaryotic cells, such as bacterial cells. In some embodiments, the recombinant cells are eukaryotic cells, such as mammalian cells, yeast cells, filamentous fungal cells, protist cells, algal cells, avian cells, plant cells, or insect cells. In some embodiments, the cells are yeast cells. Those skilled in the art will recognize that many forms of filamentous fungi produce yeast-like growth, and the definition of yeast herein includes such cells. The recombinant cells of this disclosure can be selected from algae, bacteria, molds, fungi, plants, and yeast. In some embodiments, the recombinant cells of this disclosure are bacterial cells (e.g., Escherichia coli), fungal cells, or yeast cells.

[0177] In some embodiments, the recombinant cells of this disclosure are recombinant fungal cells. The recombinant fungal cells can be any suitable fungal cells recognized in the art. In some aspects, the fungal cells are Arxula, Aspergillus, Schizochytrium, Candida, Clavicipitaceae, Cryptococcus, Clavicipitium spp., Gastrodia, Hansenula, Kluyveromyces, Kodakosomal, Komajiri, Echinosporium, Lepidosporium, Morphozoa, Mucor, Augerium, Pichia pastoris, Protocellus, Rhizopus, Rhodotorula, Rhodotorula, Yeast, Schizochytozoa, Tremella, Trichosporium, Wickham's yeast, or Yarrow's yeast cells. In some implementations, the fungal cells are *Arxula adeninivorans*, *Aspergillus niger*, *Aspergillus orzyae*, *Aspergillus terreus*, *Aurantiochytrium limacinum*, *Candida utilis*, *Claviceps purpurea*, *Cryptococcus albidus*, *Cryptococcus curvatus*, *Cryptococcus ramirezgomezianus*, *Cryptococcus terreus*, *Cryptococcus wieringae*, *Cunninghamella echinulata*, *Cunninghamella japonica*, *Geotrichumfermentans*, and *Hansenula polymorpha*. Kluyveromyces lactis, Komagataella phaffii, Komagataella pastoris, Komagataella pseudopastoris, Kluyveromyces marxianus, Kodamaea ohmeri, Leucosporidiella creatinivora, Lipomyces lipofer, Lipomyces starkeyi, Lipomyces tetrasporus, MortierellaThe following yeasts are listed: *Isabellina*, *Mortierella alpina*, *Ogataea polymorpha*, *Pichia ciferrii*, *Pichia guilliermondii*, *Pichia pastoris*, *Pichia stipites*, *Prototheca zopfii*, *Rhizopus arrhizus*, *Rhodosporidium babjevae*, *Rhodosporidium toruloides*, *Rhodosporidium paludigenum*, *Rhodotorula glutinis*, *Rhodotorula mucilaginosa*, *Saccharomyces pombe*, *Tremella enchepala*, and *Trichosporon*. The yeasts include *Cutaneum*, *Trichosporon fermentans*, *Wickerhamomyces ciferrii*, or *Yarrowia lipolytica*.

[0178] In some embodiments, the fungal cell is a yeast cell. In some embodiments, the yeast cell is a cell of the genus *Kluyveromyces*. In some embodiments, the yeast cell is *Kluyveromyces phaffii*, *Kluyveromyces pastoris*, or *Pseudo-Kluyveromyces pastoris*. In a specific embodiment, the yeast cell is *Kluyveromyces phaffii*.

[0179] In some embodiments, the engineered cells of this disclosure are yeast cells that include one or more modifications to improve the production of N-glycans (including human-like N-glycans). Examples of such cells and modifications are described, for example, in U.S. Patent 9,617,550, the entire contents of which are incorporated herein by reference.

[0180] E. Gene Editing System

[0181] Some embodiments of this disclosure relate to the use of gene editing technologies to produce gene knockouts or other mutations in the genes of a cell population. Various methods and systems for gene editing are known in the art, including, for example, zinc finger nuclease (ZFN)-based gene editing, transcription activator-like effector nuclease (TALEN)-based gene editing, and CRISPR / Cas-based gene editing. Various methods and systems for gene editing are recognized in the art and are conceived herein. In some embodiments, the methods of this disclosure include CRISPR / Cas-based gene editing, which includes the use of components of a CRISPR system, such as guide RNA (gRNA) and Cas nucleases. In some embodiments, the methods of this disclosure do not include CRISPR / Cas-based gene editing (e.g., including ZFN-based, TALEN-based, or any other gene editing methods or systems).

[0182] The term “CRISPR system” generally refers to transcripts and other elements involved in expressing or directing the activity of CRISPR-related (“Cas”) genes, including sequences encoding Cas genes, tracr (trans-activation CRISPR) sequences (e.g., tracrRNA or active tracrRNA), tracr pairing sequences (in the context of an endogenous CRISPR system, these include “direct repeat sequences” and tracrRNA-processed partial direct repeat sequences), guide sequences (also known as “spacer regions” in the context of an endogenous CRISPR system), and / or other sequences and transcripts derived from CRISPR sites.

[0183] CRISPR / Cas nucleases or CRISPR / Cas nuclease systems may include a non-coding RNA molecule (guide RNA) that specifically binds to DNA and a Cas protein (e.g., Cas9) with nuclease function (e.g., two nuclease domains). One or more elements of a CRISPR system may be derived from type I, type II, or type III CRISPR systems, for example, from specific organisms that include endogenous CRISPR systems, such as Streptococcus pyogenes.

[0184] In some respects, Cas nucleases and gRNAs (including fusions of crRNAs specifically targeting target sequences and immobilized tracrRNAs) are introduced into cells. Cas nucleases and gRNAs can be introduced indirectly into cells by introducing one or more nucleic acids encoding Cas nucleases and / or gRNAs (e.g., vectors). Cas nucleases and gRNAs can also be introduced directly into cells by introducing Cas nuclease proteins and gRNA molecules. Typically, the target site at the 5' end of the gRNA uses complementary base pairing to target the Cas nuclease at a target site, such as a gene. The target site can be selected based on its position at the 5' of its protospacer adjacent motif (PAM) sequence, such as the typical NGG or NAG. In this regard, the gRNA can be targeted to the desired sequence by modifying the first 20, 19, 18, 17, 16, 15, 14, 14, 12, 11, or 10 nucleotides of the guide RNA to correspond to the target DNA sequence. Typically, CRISPR systems are characterized by elements that facilitate the formation of the CRISPR complex at the target sequence site. The "target sequence" typically refers to a sequence designed to be complementary to the guide sequence, where hybridization between the target and guide sequences promotes the formation of the CRISPR complex. Perfect complementarity is not necessarily required if sufficient complementarity is sufficient to induce hybridization and promote CRISPR complex formation.

[0185] As discussed herein, the CRISPR system can induce double-strand breaks (DSBs) at the target site, leading to disruption. In other embodiments, a Cas9 variant (considered a “nicking enzyme”) is used to nick the target site in a single strand. Paired nicking enzymes can be used, for example to improve specificity, with each nicking enzyme guided by a pair of different gRNA targeting sequences, such that a 5' overhang is introduced when the nick is introduced simultaneously. In other embodiments, a catalytically inactive Cas9 is fused with a heterologous effector domain (such as a transcriptional repressor or activator) to influence gene expression.

[0186] The target sequence can include any polynucleotide, such as DNA or RNA polynucleotides. The target sequence can be located in the cell nucleus or cytoplasm, for example, within a cell's organelles. Typically, the sequence or template that can be used for recombination into a target site that includes the target sequence is called an "edit template," "edit polynucleotide," or "edit sequence." In some respects, exogenous template polynucleotides can be called edit templates. In some respects, recombination is homologous recombination.

[0187] Typically, in the context of an endogenous CRISPR system, the formation of a CRISPR complex (including a guide sequence that hybridizes to a target sequence and complexes with one or more Cas proteins) results in the cleavage of one or both strands within or near the target sequence (e.g., within the range of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50 or more base pairs). This tracr sequence may include all or part of a wild-type tracr sequence, or may consist of all or part of a wild-type tracr sequence (e.g., about or greater than about 20, 26, 32, 45, 48, 54, 63, 67, 85 or more nucleotides of a wild-type tracr sequence), which may also constitute part of the CRISPR complex, for example, by hybridizing along at least a portion of the tracr sequence with all or part of a tracr pairing sequence, said tracr pairing sequence being operatively linked to the guide sequence. The tracr sequence and its pairing sequence have sufficient complementarity to hybridize and participate in the formation of the CRISPR complex. For example, when the alignment is optimal, the length of the pairing sequence along the tracr is at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence complementarity.

[0188] One or more vectors that drive the expression of one or more elements of the CRISPR system can be introduced into cells, such that the expression of the CRISPR system elements guides the formation of a CRISPR complex at one or more target sites. Components can also be delivered to cells as proteins and / or RNA. For example, the Cas enzyme, the guide sequence linked to the tracr pairing sequence, and the tracr sequence can all be operatively linked to separate regulatory elements on different vectors. Alternatively, two or more elements expressed by the same or different regulatory elements can be combined in a single vector to provide any components of the CRISPR system not included in the first vector via one or more additional vectors. The vector may include one or more insertion sites, such as restriction endonuclease recognition sequences (also known as “cloning sites”). In some embodiments, one or more insertion sites are located upstream and / or downstream of one or more sequence elements of one or more vectors. When using multiple different guide sequences, a single expression construct can be used to target CRISPR activity to multiple different, corresponding target sequences within the cell.

[0189] The vector may include a regulatory element operatively linked to an enzyme-coding sequence encoding a Cas protein (also known as a "Cas nuclease"). Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Cas12a (Cpf1), Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, and C... Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csfl, Csx1, Csx1, Csfl, Csf2, Csf3, Csf4, their homologues or modified forms thereof. These enzymes are known; for example, the amino acid sequence of the *Streptococcus pyogenes* Cas9 protein can be found in the SwissProt database accession number Q99ZW2.

[0190] The Cas nuclease can be Cas9 (e.g., from Streptococcus pyogenes or Streptococcus pneumoniae). The Cas nuclease can be Cas12a. The Cas nuclease can direct the cleavage of one or both strands at the location of the target sequence, such as within the target sequence and / or within its complementary sequence. The vector can encode a Cas nuclease mutated relative to the corresponding wild-type enzyme, such that the mutated Cas nuclease lacks the ability to cleave one or both strands of the target polynucleotide containing the target sequence. In some embodiments, the Cas9 cleavage enzyme can be used in combination with a guide sequence, for example, two guide sequences that target the sense and antisense strands of the DNA target, respectively. This combination allows cleavage to occur on both strands and can be used to induce NHEJ or HDR.

[0191] In some implementations, the enzyme coding sequence encoding the CRISPR enzyme is codon-optimized for expression in specific cells (e.g., yeast cells).

[0192] The guide sequence is typically any polynucleotide sequence that is sufficiently complementary to the target polynucleotide sequence to hybridize with the target sequence and guide the CRISPR complex to bind sequence-specifically to the target sequence. In some embodiments, when optimal alignment is performed using a suitable alignment algorithm, the complementarity between the guide sequence and its corresponding target sequence is 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or higher.

[0193] The optimal alignment can be determined by using any suitable sequence alignment algorithm, non-limiting embodiments of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transform (e.g., Burrows-Wheeler Aligner), Clustal W, Clustal X, BLAST, Novoalign (Novocraft Technologies, ELAND (Illumina, San Diego, Calif.), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net).

[0194] Cas nucleases can be part of a fusion protein comprising one or more heterologous protein domains. Cas nuclease fusion proteins can include any additional protein sequences, as well as optionally, a linker sequence between any two domains. Examples of protein domains that can be fused with Cas nucleases include, but are not limited to, epitope tags, reporter gene sequences, and protein domains having one or more of the following activities: methyltransferase activity, demethyltransferase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, and nucleic acid binding activity. Non-limiting examples of epitope tags include histidine (His) tags, V5 tags, FLAG tags, influenza hemagglutinin (HA) tags, Myc tags, VSV-G tags, and thioredoxin (Trx) tags. Examples of reporter genes include, but are not limited to, glutathione-5-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and autofluorescent proteins, including blue fluorescent protein (BFP). Cas nucleases can be fused with gene sequences encoding proteins or protein fragments that bind to DNA molecules or other cellular molecules, including, but not limited to, maltose-binding protein (MBP), S-tags, Lex A DNA-binding domain (DBD) fusions, GAL4A DNA-binding domain fusions, and herpes simplex virus (HSV) BP16 protein fusions. Other domains that can form a fusion protein comprising a Cas nuclease are described in US20110059502 and are incorporated herein by reference.

[0195] Example

[0196] The following embodiments are included to demonstrate certain embodiments disclosed herein. Those skilled in the art will understand that the techniques disclosed in the following embodiments represent techniques discovered by the inventors that work well in the practice of the disclosed embodiments, and thus can be considered to constitute certain patterns for their practice. However, based on this disclosure, those skilled in the art will understand that many changes can be made to the specific embodiments disclosed without departing from the spirit and scope of the disclosed embodiments, and similar or related results can still be obtained.

[0197] Example 1 - Novel signal peptides increase extracellular protein levels

[0198] To determine the effect of the novel signal peptide on extracellular protein levels, DNA encoding SEQ ID NO:1 (“SP1”), SEQ ID NO:2 (“SP2”), and SEQ ID NO:4 (“SP4”) was cloned within the 5' end frame of the DNA encoding the target protein (POI) (i.e., Pichia pastoris codon-optimized human lactoferrin), replacing the pro-pro-MFα derived from Saccharomyces cerevisiae. This is the most widely used signal peptide in yeast and was used as a control. Single copies of the resulting and control sequences were integrated into the AOX1 locus via double crossover. Multiple colonies were cultured in each of the 96-well plates.

[0199] To confirm the presence of the target protein, Western blotting was performed using the supernatant. For example... Figure 1 As shown, when a single copy of human lactoferrin is integrated, secretion is driven by the widely used pre-pro-MFα of *Saccharomyces cerevisiae*, and the protein is not detected in the supernatant. Conversely, when secretion is driven by SEQ ID NO:1 (“SP1”), SEQ ID NO:2 (“SP2”), and SEQ ID NO:3 (“SP3”), extracellular proteins are detected.

[0200] To assess the extent of improved secretion, extracellular proteins were quantified using ELISA. For example... Figure 2 As shown, compared with the control group (pro-pro-MFα), the newly engineered signals of SEQ ID NO:1 (“SP1”), SEQ ID NO:2 (“SP2”) and SEQ ID NO:3 (“SP3”) increased extracellular protein levels by 2.38-fold, 2.41-fold and 2.20-fold, respectively.

[0201] Materials and Methods

[0202] Vector and strain construction. Oligonucleotides and gBlocks were ordered from Integrated DNA Technologies (San Diego, California, USA), as shown in Table 5. HiFi DNA Assembly Premix The DNA polymerase and E. coli DH5α cells were obtained from New England Biolabs. All polymerase chain reaction (PCR) amplified sequences were confirmed by Genewiz sequencing.

[0203]

[0204] Linear dsDNA transformation for integration was performed using the methods described in the following literature (Madden, Tolstorukov, & Cregg (2014) Fungi, Volume 1, Fungal Biology). The Easy DNA kit from Invitrogen (ThermoFisher, Applied Biosystems, PrepSEQ) was used. TMTM Total yeast genomic DNA was extracted using the 1-2-3 Nucleic Acid Extraction Kit (catalog number: 4452222). The resulting plasmids are summarized in Table 6.

[0205]

[0206] The leader peptide sequences of the endogenous proteins Ost1 and Pst1 in Pichia pastoris were determined using SignalP-5.0 bioinformatics software, which is publicly available from the Center Biological Sequence Analysis (CBS). The original region of Epx1 is described in the following literature (Heiss et al. (2015) Microbiology, 161(7)).

[0207] Genscript synthesized plasmid P1 containing the gene encoding human lactoferrin (excluding its naturally secreted peptide), which was fused within the pre-pro-lead peptide frame of *Saccharomyces cerevisiae* mating factor-α. The human lactoferrin gene was codon-optimized for expression in *Pichia pastoris*.

[0208] To construct plasmid P2 containing the signal sequence SP1 (SEQ ID NO:1), the Ost1 leader sequence was amplified using primers PMR1 (SEQ ID NO:16) and PMR2 (SEQ ID NO:17) with gBLOCK1 as a template. Polymerase chain reaction (PCR) of plasmid P1 using primers PMR3 (SEQ ID NO:18) and PMR4 (SEQ ID NO:19) yielded a backbone containing human lactoferrin, yeast HIS4 auxotrophic marker, and E. coli antibiotic resistance and origin of replication. Following the manufacturer's instructions, [the process was carried out using...]. The two fragments obtained from the assembly of HiFi DNA assembly premix solution.

[0209] To generate plasmid P3 containing the signal sequence SP2 (SEQ ID NO:2), primers PMR5 (SEQ ID NO:20) and PMR6 (SEQ ID NO:21) were used.

[0210] gBLOCK1 (SEQ ID NO:15) was used as a template for amplification. Primers were used...

[0211] PCR of P1 plasmids from PMR7 (SEQ ID NO:22) and PMR8 (SEQ ID NO:23) yielded a backbone containing human lactoferrin, yeast HIS4 auxotrophic markers, and E. coli antibiotic resistance and origin of replication. Following the manufacturer's instructions, [the following was performed] using... The two fragments obtained from the assembly of HiFi DNA assembly premix solution.

[0212] To generate plasmid P4 containing the signal sequence SP3 (SEQ ID NO:3), amplification was performed using primers PMR9 (SEQ ID NO:24) and PMR10 (SEQ ID NO:25) with gBLOCK1 as a template. PCR of plasmid P1, performed using primers PMR11 (SEQ ID NO:26) and PMR12 (SEQ ID NO:27), yielded a backbone containing human lactoferrin, yeast HIS4 auxotrophic markers, and E. coli antibiotic resistance and origin of replication. Following the manufacturer's instructions, [the following steps were performed]. The two fragments obtained from the assembly of HiFi DNA assembly premix solution.

[0213] To generate plasmid P5 containing the signal sequence SP4 (SEQ ID NO:4), primers PMR13 (SEQ ID NO:28) and PMR14 (SEQ ID NO:29) were used.

[0214] gBLOCK1 (SEQ ID NO:15) was used as a template for amplification. Primers were used...

[0215] PCR of P1 plasmids PMR15 (SEQ ID NO:30) and PMR16 (SEQ ID NO:31) yielded a backbone containing human lactoferrin, yeast HIS4 auxotrophic markers, and E. coli antibiotic resistance and origin of replication. Following the manufacturer's instructions, [the following was performed] using... The two fragments obtained from the assembly of HiFi DNA assembly premix solution.

[0216] Transform the assembly mixture into *E. coli* DH5α cells according to the manufacturer's instructions and inoculate them into Luria broth (LB) agar plates containing 100 μg / mL ampicillin. Positive clones were selected by colony polymerase chain reaction (PCR) and inoculated overnight in 5 mL of liquid Luria broth supplemented with 100 μg / mL ampicillin. The GeneJET plasmid miniprep kit was used. The plasmid (catalog number K0502) was isolated from E. coli cells. Correct assembly was confirmed by Sanger DNA sequencing.

[0217] Using Q5 high-fidelity DNA polymerase, primers PMR17 (SEQ ID NO:32) and PMR18 (SEQ ID NO:33) and plasmids P1, P2, P3, P4, or P5 as templates, linear dsDNA fragments for integration into yeast were obtained. Electrotransformation of competent Pichia pastoris cells was performed according to the following literature: Madden, Tolstorukov, & Cregg (2014) Fungi, Volume 1, Fungal Biology. Cells were spread on MD plates (1.34% yeast nitrogen basal, 4 x 10⁻⁶ m² / cm²). -5 % Biotin, 2% Dextrose, 20% Agar), which allows for the selection of his4 + Cells were incubated at 30°C for 72 hours. Single yeast colonies (approximately 10-20) were then re-stripened onto MD plates and allowed to grow at 30°C for 24 hours. P1-transformed cells were used as a control to evaluate the higher efficiency of SP1 (SEQ ID NO:1), SP2 (SEQ ID NO:2), SP3 (SEQ ID NO:3), and SP4 (SEQ ID NO:5) in the secretion of the target protein (POI).

[0218] Single colonies from restretched plates were inoculated into 96-well plates using 600 μl of 2% YPD (2% dextrose, 2% peptone, 1% yeast extract). Cells were grown at 1,000 rpm and 30°C for 48 h. 50 μl of the resulting cell suspension was transferred to 550 μl of BMG (100 mM potassium phosphate buffer (pH 6.0), 1.34% yeast nitrogen basal solution, 4 x 10⁻⁶ oz.) supplemented with 0.5% CAS amino acids. -5 The cells were incubated with 1% biotin and 1% glycerol at 1,000 rpm and 30°C for 48 hours. Cells were then pelleted by centrifugation at 4,500 x g for 5 minutes and resuspended in 1% BMM (100 mM potassium phosphate buffer (pH = 6.0), 1.34% yeast nitrogen basal solution, 4 x 10⁻⁶ ppm). -5The proteins were induced for 72 hours at 1,000 rpm and 20°C in a solution of 1% biotin and 1% methanol. The proteins secreted into the extracellular culture medium were then analyzed by SDS-PAGE, ELISA, and Western blotting.

[0219] ***

[0220] In view of this disclosure, all methods disclosed and claimed herein can be made and performed without excessive experimentation. While the compositions and methods disclosed herein have been described according to certain embodiments, it will be apparent to those skilled in the art that variations may be imposed on the methods and steps or the order of steps described herein without departing from the concept, spirit, and scope of the disclosed embodiments. More specifically, certain chemically and physiologically relevant reagents may obviously substitute for the reagents described herein while achieving the same or similar results. All such similar substitutions and modifications that will be apparent to those skilled in the art are considered to be within the spirit, scope, and concept of the embodiments disclosed herein as defined by the appended claims.

[0221] References

[0222] The following references are specifically incorporated herein by reference to provide exemplary procedures or other details that supplement those described herein.

[0223] Bernauer et al.,Komagataella phaffii as emerging model organism infundamental research.Front.Microbiol.(January 11,2021).

[0224] Besada-Lombana&Da Silva(2019)Engineering the early secretory pathway for increased protein secretion in Saccharomyces cerevisiae.MetabolicEngineering,55,142-151(September 2019).

[0225] Dalvie et al.(2020)“Host-informed expression of CRISPR guide RNA forgenomic engineering in Komagataella phaffii.”ACS Synth.Biol.,9(1),26-35(December 11,2019).

[0226] Duran&Kahve(2017)The use of lactoferrin in food industry.AcademicJournal of Science,07(02),89-94.

[0227] Heiss et al.(2015)Multi-step processing of the secretion leader ofthe extracellular protein Epx1 in Pichia pastoris and implications forprotein localization.Microbiology,161(7)(July 1,2015).

[0228] Madden,Tolstorukov,&Cregg,Book Chapter:Electroporation of Pichiapastoris.Genetic Transformation Systems 87in Fungi,Volume 1,FungalBiology.M.A.van den Berg and K.Maruthachalam(eds.)(2014).

[0229] Nicholl,An Introduction to Genetic Engineering.2nd edition(Cambridge:Cambridge University Press,2002),Glossary.

[0230] Recombinant Protein Production in Yeast,Brigitte Gasser&DiethardMattanovich(eds.)(Springer,2019).

[0231] US Patent No. 4,977,137 (Nicols et al.)

[0232] US Patent No. 5,571,691 (Conneely et al.)

[0233] US Patent No. 7,335,512 (Callewaert et al.)

[0234] US Patent No. 7,344,867 (Connolly)

[0235] US Patent No. 7,749,960 (Vidal et al.)

[0236] U.S. Patent No. 7,524,815 (Vidal et al.)

[0237] US Patent No. 7,914,822 (Medo)

[0238] US Patent No. 8,440,456 (Callewaert et al.)

[0239] US Patent No. 8,871,445 (Cong et al.)

[0240] US Patent No. 8,802,650 (Buck et al.)

[0241] US Patent No. 8,821,878 (Medo et al.)

[0242] U.S. Patent No. 8,927,027 (Fournell et al.)

[0243] U.S. Patent No. 7,449,308 (Gerngross et al.)

[0244] US Patent Publication No. 2012 / 0142580 (Nutten et al.)

Claims

1. An isolated nucleic acid encoding a polypeptide, said polypeptide comprising the sequence of SEQ ID NO:

3.

2. The isolated nucleic acid according to claim 1, wherein the polypeptide further comprises a sequence of a mammalian protein.

3. The isolated nucleic acid according to claim 2, wherein the mammalian protein is a human milk protein.

4. The isolated nucleic acid according to claim 3, wherein the human milk protein is secretory IgA (sIgA), xanthine dehydrogenase, lactoferrin, lactoperoxidase, lactolipoprotein, lactobacin, adiponectin, β-casein, κ-casein, leptin, lysozyme, or α-lactalbumin.

5. The isolated nucleic acid according to claim 4, wherein the human milk protein is human lactoferrin.

6. The isolated nucleic acid according to any one of claims 1-5, wherein the isolated nucleic acid comprises the nucleic acid sequence of SEQ ID NO:

43.

7. The isolated nucleic acid according to any one of claims 1-5, wherein the isolated nucleic acid is codon-optimized.

8. The isolated nucleic acid according to claim 7, wherein the isolated nucleic acid comprises the nucleic acid sequence of SEQ ID NO:

48.

9. A vector comprising the nucleic acid according to any one of claims 1-8.

10. An engineered Pichia pastoris ( Pichia pastoris Cells, including the nucleic acid of any one of claims 1-8 or the vector of claim 9.

11. A method for producing a secreted protein, the method comprising growing Pichia pastoris cells of claim 10 under conditions sufficient to secrete a polypeptide from the cell.

12. The method of claim 11, wherein the method includes collecting the secreted protein.

13. The method of claim 11, wherein the secreted protein is a human milk protein.

14. The method according to claim 13, wherein the human milk protein is secretory IgA (sIgA), xanthine dehydrogenase, lactoferrin, lactoperoxidase, lactolipoprotein, lactobacin, adiponectin, β-casein, κ-casein, leptin, lysozyme, or α-lactalbumin.

15. The method according to claim 13 or 14, wherein the human milk protein comprises one or more human-like N-glycans.

16. The method according to any one of claims 11-14, further comprising generating a mixture comprising one or more components including human milk proteins and infant formula.

17. The engineered Pichia pastoris cells according to claim 10, wherein the nucleic acid encodes a mammalian polypeptide.

18. The engineered Pichia pastoris cells according to claim 10, wherein the polypeptide encoded by the isolated nucleic acid comprises SEQ ID NO:

3.

19. The engineered Pichia pastoris cells according to claim 17, wherein the polypeptide is a mammalian protein.

20. The engineered Pichia pastoris cells according to claim 19, wherein the mammalian protein is human milk protein.

21. The engineered Pichia pastoris cells according to claim 20, wherein the human milk protein is secretory IgA (sIgA), xanthine dehydrogenase, lactoferrin, lactoperoxidase, lactolipoprotein, lactobacin, adiponectin, β-casein, κ-casein, leptin, lysozyme, or α-lactalbumin.

22. The engineered Pichia pastoris cells according to claim 20 or 21, wherein the human milk protein is human lactoferrin.

Citation Information

Patent Citations

  • Multiple domain proteins

    US20110059502A1

  • Opioid receptors stimulating compounds (thymoquinone, nigella sativa) and food allergy

    US20120142580A1

  • Lactoferrin as a dietary ingredient promoting the growth of the gastrointestinal tract

    US4977137A

  • Production of recombinant lactoferrin and lactoferrin polypeptides using CDNA sequences in various organisms

    US5571691A

  • Marker for measuring liver cirrhosis

    US7335512B2