Methods and compositions for protein synthesis and secretion
Novel synthetic secretory signal peptides in Pichia pastoris improve co-translational translocation and secretion of mammalian proteins, addressing inefficiencies in existing methods and enhancing extracellular production.
Patent Information
- Application Number
- JP2024506169
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-10-29
- Filing Date
- 2022-07-29
- Publication Date
- 2025-10-20
- Estimated Expiration
- 2042-07-29
AI Technical Summary
Existing methods for protein production in Pichia pastoris face challenges in promoting co-translational translocation and efficient secretion of recombinant proteins, particularly mammalian proteins like human milk proteins, due to issues with partially folded domains and aggregation in the cytoplasm, leading to suboptimal extracellular production.
Development of novel synthetic secretory signal peptides derived from in-frame fusions of P. pastoris pre-secretory peptides with Saccharomyces cerevisiae mating factor proregions, enhancing extracellular production of proteins such as human milk proteins by improving translocation efficiency and secretion.
The novel signal peptides significantly enhance the extracellular production of mammalian proteins, including human milk proteins, by promoting efficient co-translational translocation and secretion, outperforming traditional signal peptides.
Smart Images

Figure 0007756787000026 
Figure 0007756787000027 
Figure 0007756787000001
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 227,820, filed July 30, 2021, and U.S. Provisional Application No. 63 / 273,858, filed October 21, 2021, which are incorporated by reference in their entirety into this specification.
[0002] Array List
[0002] This application contains a sequence listing that has been submitted in XML format, which is hereby incorporated by reference in its entirety. Said XML copy, created on July 7, 2022, is named HELA_P0005WO_Sequence_Listing.xml and is 61,471 bytes in size. [Background technology]
[0003] I.Technical field Aspects of the present invention relate to at least the fields of microbiology, genetics and biotechnology.
[0004] II. Background Yeast is a desirable host for recombinant protein production due to its rapid growth, ability to reach high cell densities, ability to grow on defined minimal media, high protein yields, and ability to perform eukaryotic post-translational modifications. The most suitable yeast for protein production is Pichia pastoris (Komagataella pastoris, Komagataella phaffii) due to the widespread availability of genome information and molecular tools for genome manipulation. These have enabled the use of Pichia pastoris for the production of GRAS ingredients under FDA standards.
[0005] For a variety of biotechnological applications, it is often preferred to produce proteins that are secreted into the growth medium for ease of recovery. Pichia pastoris is capable of secreting active recombinant proteins while maintaining low levels of secretion of endogenous proteins.
[0006] In eukaryotes, secreted proteins are first targeted from the cytoplasm to the lumen of the endoplasmic reticulum (ER) by translocation. Translocation into the ER can occur either post-translationally (i.e., once the polypeptide chain has been synthesized) or co-translationally (i.e., while the mRNA is being translated into its amino acid sequence). Post-translational translocation requires the action of the ER-resident chaperone Kar2, which acts as a chaperone and molecular ratchet to maintain the polypeptide chain in a loose conformation in the cytoplasm. Consequently, this process can be hindered by partially folded domains and / or aggregation in the cytoplasm. Therefore, for biotechnological applications, it is desirable to promote co-translational translocation. Once in the ER, proteins are glycosylated, their disulfide bonds are isomerized, and they fold to their native state. Successfully folded proteins then translocate to the Golgi complex, where they undergo further glycosylation before being packaged into secretory granules, which fuse with the plasma membrane and release the protein into the extracellular environment.
[0007]
[0007] Targeting of proteins to the secretory pathway is mediated by secretory peptides. The most widely used in Pichia pastoris is the leader peptide of mating factor alpha from Saccharomyces cerevisiae. It consists of two distinct regions: ii) a preregion of the first 19 amino acids that promotes posttranslational translocation and is cleaved upon ER entry; and ii) a 70-amino acid prosegment that functions as an export signal from the ER to the Golgi apparatus, where it is cleaved at a dibasic amino acid cleavage site, KR.
[0008]
[0008] There is a need for synthetic secretory signal peptides that lead to higher extracellular production of proteins. Summary of the Invention
[0009] Aspects of the present disclosure address a particular need by providing novel secretory signal peptides that are effective in improving the extracellular production of proteins, including mammalian proteins such as human milk proteins. Certain aspects of the present disclosure are based, at least in part, on the development of signal peptides generated from the in-frame fusion of 1) a P. pastoris pre-secretory peptide from either i) the alpha subunit of the ER lumen oligosaccharyltransferase complex (Ost1) or ii) the GPI-anchored protein Pst1 with 2) the proregion of either i) the Saccharomyces cerevisiae mating factor or ii) the proregion of P. pastoris Epx1. Accordingly, described herein are isolated nucleic acids encoding such secretory signal peptides, in some cases linked to recombinant proteins such as human milk proteins, as well as cells containing such nucleic acids and methods for producing and recovering recombinant proteins from such cells.
[0010]
[0010] Described herein, in certain embodiments, is an isolated nucleic acid encoding a polypeptide comprising a sequence having at least 90% sequence identity to SEQ ID NO:1, 2, 3, or 4. In certain embodiments, the sequence comprises SEQ ID NO:1, 2, 3, or 4. In certain embodiments, the polypeptide further comprises the sequence of a mammalian protein. In certain embodiments, the mammalian protein is a human milk protein. In certain embodiments, the human milk protein is secretory IgA (sIgA), xanthine dehydrogenase, lactoferrin, lactoperoxidase, butyrophilin, lactadherin, adiponectin, β-casein, κ-casein, leptin, lysozyme, or α-lactalbumin. In certain embodiments, the human milk protein is human lactoferrin.
[0011] In some embodiments, the sequence has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO:1. In some embodiments, the sequence comprises SEQ ID NO:1. In some embodiments, the isolated nucleic acid comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO:41. In some embodiments, the nucleic acid sequence comprises SEQ ID NO:41. In some embodiments, the polypeptide comprises SEQ ID NO:5. In some embodiments, the isolated nucleic acid comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 46. In some embodiments, the nucleic acid sequence comprises SEQ ID NO:46.
[0012] In some embodiments, the sequence has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO:2. In some embodiments, the sequence comprises SEQ ID NO:2. In some embodiments, the isolated nucleic acid comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO:42. In some embodiments, the nucleic acid sequence comprises SEQ ID NO:42. In some embodiments, the polypeptide comprises SEQ ID NO:6. In some embodiments, the isolated nucleic acid comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 47. In some embodiments, the nucleic acid sequence comprises SEQ ID NO: 47.
[0013] In some embodiments, the sequence has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO:3. In some embodiments, the sequence comprises SEQ ID NO:3. In some embodiments, the isolated nucleic acid comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO:43. In some embodiments, the nucleic acid sequence comprises SEQ ID NO:43. In some embodiments, the polypeptide comprises SEQ ID NO:7. In some embodiments, the isolated nucleic acid comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 48. In some embodiments, the nucleic acid sequence comprises SEQ ID NO:48.
[0014] In some embodiments, the sequence has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO:4. In some embodiments, the sequence comprises SEQ ID NO:4. In some embodiments, the isolated nucleic acid comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO:44. In some embodiments, the nucleic acid sequence comprises SEQ ID NO:44. In some embodiments, the polypeptide comprises SEQ ID NO:8. In some embodiments, the isolated nucleic acid comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 49. In some embodiments, the nucleic acid sequence comprises SEQ ID NO: 49.
[0015] Also disclosed herein, in certain aspects, are vectors that include a nucleic acid disclosed herein (eg, an isolated nucleic acid or a sequence or portion thereof).
[0016]
[0016] Further disclosed in certain aspects are engineered eukaryotic cells comprising the nucleic acids disclosed herein. In certain embodiments, the cell is a fungal cell. In certain embodiments, the fungal cell is a fungal cell selected from the group consisting of Arxula, Aspergillus, Aurantiochytrium, Candida, Claviceps, Cryptococcus, Cunninghamella, Geotrichum, Hansenula, Kluyveromyces, Kodamaea, Komagataella, Leucosporidiella, Lipomyces, and the like. The cell is a cell of the genus Lipomyces, Mortierella, Ogataea, Pichia, Prototheca, Rhizopus, Rhodosporidium, Rhodotorula, Saccharomyces, Schizosaccharomyces, Tremella, Trichosporon, Wickerhamomyces, or Yarrowia. In some embodiments, the cell is a yeast cell. In some embodiments, the yeast cell is a Komagataella cell. In some embodiments, the yeast cell is a Komagataella puffii, Komagataella pastoris, or Komagataella pseudopastoris cell. In some aspects, the nucleic acid is integrated into the genome of the cell. In some aspects, the nucleic acid is not integrated into the genome of the cell.
[0017] Also disclosed in certain aspects are methods for producing secreted proteins, the methods comprising growing an engineered eukaryotic cell of the present disclosure under conditions sufficient to secrete the polypeptide from the cell. In certain embodiments, the methods further comprise recovering the secreted protein. In certain aspects, the secreted protein is a human milk protein. In certain embodiments, the human milk protein is secretory IgA (sIgA), xanthine dehydrogenase, lactoferrin, lactoperoxidase, butyrophilin, lactadherin, adiponectin, β-casein, κ-casein, leptin, lysozyme, or α-lactalbumin. In certain embodiments, the human milk protein is human lactoferrin. In certain embodiments, the human milk protein comprises one or more human-like N-glycans. In certain embodiments, the methods further comprise producing a mixture comprising the human milk protein and one or more components of infant formula.
[0018]
[0018] Further, in certain aspects, disclosed herein is an engineered yeast cell comprising a nucleic acid encoding a polypeptide comprising a sequence having at least 90% sequence identity to SEQ ID NO:1, 2, 3, or 4. In certain embodiments, the sequence comprises SEQ ID NO:1. In certain embodiments, the sequence comprises SEQ ID NO:2. In certain embodiments, the sequence comprises SEQ ID NO:3. In certain embodiments, the sequence comprises SEQ ID NO:4. In certain embodiments, the polypeptide further comprises the sequence of a mammalian protein. In certain embodiments, the mammalian protein is a human milk protein. In certain embodiments, the human milk protein is secretory IgA (sIgA), xanthine dehydrogenase, lactoferrin, lactoperoxidase, butyrophilin, lactadherin, adiponectin, β-casein, κ-casein, leptin, lysozyme, or α-lactalbumin. In some embodiments, the human milk protein is human lactoferrin.
[0019]
[0019] In one aspect, described herein is an engineered yeast cell comprising: (a) a first nucleic acid encoding a polypeptide comprising (i) a sequence having at least 90% sequence identity to SEQ ID NO: 1, 2, 3, or 4, and ii) the sequence of a human milk protein; and (b) a second nucleic acid encoding an alpha-1,2-mannosidase (Man-I) protein, wherein the cell does not express a functional OCH1 protein. In one aspect, the sequence of (i) comprises SEQ ID NO: 1, 2, 3, or 4. In one embodiment, the human milk protein is secretory IgA (sIgA), xanthine dehydrogenase, lactoferrin, lactoperoxidase, butyrophilin, lactadherin, adiponectin, β-casein, κ-casein, leptin, lysozyme, or α-lactalbumin. In one embodiment, the human milk protein is human lactoferrin. In some embodiments, the human milk protein is human α-lactalbumin. In some embodiments, the Man-I protein is fused to an HDEL C-terminal tag. In some embodiments, the cell further comprises a third nucleic acid encoding one or more of: (a) an N-acetylglucosaminyltransferase I (GnT-I) protein; (b) an α-1,3 / 6-mannosidase (Man-II) protein; (c) a β-1,2-acetylglucosaminyltransferase (GnT-II) protein; and (d) a β-1,4-galactosyltransferase (GalT) protein. In some embodiments, the yeast cell is a Komagataella cell. In some embodiments, the yeast cell is a Komagataella puffii, Komagataella pastoris, or Komagataella pseudopastoris cell. In some aspects, the nucleic acid is integrated into the genome of the cell. In some aspects, the nucleic acid is not integrated into the genome of the cell.
[0020]
[0020] It is contemplated that any embodiment discussed herein can be implemented with respect to any method or composition of the disclosed embodiments, and vice versa. Furthermore, compositions of the embodiments disclosed herein can be used to achieve the methods of those embodiments.
[0021]
[0021] Other objects, features, and advantages of the present embodiments disclosed herein will become apparent from the following detailed description. However, it should be understood that the detailed description and specific examples, while indicating particular embodiments, are given by way of illustration only, since various changes and modifications within the spirit and scope of the embodiments disclosed herein will become apparent to those skilled in the art from this detailed description.
[0022] The following drawings form part of the present specification and are included to further demonstrate certain aspects of the present disclosure, which may be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein. [Brief explanation of the drawings]
[0023] [Figure 1]
[0023] Figure 1 is an image of a Western blot of supernatants. Lane 1 is loaded with a protein standard, Genscript, M00624 (ThermoFisher Scientific, Waltham, Massachusetts, USA). Lane 2 is loaded with human milk lactoferrin, Sigma Aldrich, SRP6519 (Sigma Aldrich, St. Louis, Missouri, USA). Lane 3 is loaded with a control (Saccharomyces cerevisiae pre-pro-MFα). Lane 4 is loaded with a negative control (supernatant from untransformed yeast cells). Lanes 5-6 are loaded with supernatant from SP2-lactoferrin transformed yeast cells. Lanes 7-8 are loaded with supernatant from SP3-lactoferrin transformed yeast cells. Lanes 9-10 are loaded with SP1-lactoferrin transformed yeast cells.
[0024] [Figure 2]
[0024] Figure 2 is a bar graph showing protein expression levels. Quantification of extracellular proteins was performed by ELISA. DETAILED DESCRIPTION OF THE INVENTION
[0025]
[0025] Described herein is the generation of novel synthetic secretory signal peptides. Also disclosed are cells (e.g., fungal cells such as yeast cells) engineered to express one or more exogenous proteins (e.g., human milk proteins) containing such signal peptides. As disclosed herein, in-frame fusions of the "preregion" sequence from P. pastoris Ost1 or Pst1 and the "proregion" sequence from Saccharomyces cerevisiae mating factor α or P. pastoris Epx1 can promote increased extracellular protein production compared to previously used signal peptides. Disclosed signal peptides include, for example, peptides containing SEQ ID NO:1, 2, 3, or 4, as well as peptides containing 1, 2, 3, 4, or 5 amino acid substitutions (or more) relative to SEQ ID NO:1, 2, 3, or 4. As described herein, in-frame fusion of these hybrid signal peptides to the N-terminus of mammalian proteins (e.g., human milk proteins such as lactoferrin or α-lactalbumin) promote highly efficient protein secretion.
[0026] I. Definition The term "biologically active portion" refers to an amino acid sequence that is less than the full-length amino acid sequence but that exhibits at least one activity of the full-length sequence. For example, a biologically active portion of an enzyme can refer to one or more domains of the enzyme that have the catalytic activity of the enzyme (i.e., the catalytic domain). In one aspect, a biologically active portion of an enzyme is a portion of the enzyme that includes the catalytic domain of the enzyme. A biologically active portion of a protein includes a peptide or polypeptide that includes an amino acid sequence that is sufficiently identical to or derived from the amino acid sequence of the protein, which contains fewer amino acids than the full-length protein, and which exhibits at least one activity of the protein (e.g., enzymatic activity, functional activity, etc.).
[0027] The term "exogenous" refers to anything that is introduced into or has been introduced into a cell. An "exogenous nucleic acid" is a nucleic acid that enters or enters a cell through a cell membrane. An "exogenous nucleic acid sequence" is the nucleic acid sequence of an exogenous nucleic acid. An exogenous nucleic acid can contain a nucleotide sequence present in the cell's native genome and / or a nucleotide sequence not previously present in the cell's genome. Exogenous nucleic acids include exogenous genes. An "exogenous gene" is a nucleic acid that encodes the expression of RNA and / or protein introduced into a cell (e.g., by transformation / transfection) and is also called a "transgene." Cells containing exogenous nucleic acids may be called recombinant cells, into which additional exogenous gene(s) can be introduced. Exogenous genes can be from the same or a different species as the cell being transformed. Thus, an exogenous gene can include a native gene that occupies a different location in the cell's genome or is under different control relative to the endogenous copy of that gene. An exogenous gene can exist in more than one copy in a cell. The exogenous gene may be maintained in the cell either as an insertion in the genome (nuclear, mitochondrial, or plastid) or as an episomal molecule.
[0028] "Operable linkage" (or "operably linked") refers to a functional linkage between two nucleic acid sequences, such as a regulatory sequence (typically a promoter) and a linked sequence (typically a sequence that encodes a protein, also called a coding sequence). A promoter is operably linked to a gene if it is capable of mediating transcription of the gene.
[0029] The term "native" refers to the composition of a cell or parent cell prior to a transformation event. A "native gene" (also called an "endogenous gene") refers to a nucleotide sequence that encodes a protein that has not been introduced into a cell by a transformation event. A "native protein" (also called an "endogenous protein") refers to an amino acid sequence encoded by a native gene.
[0030] "Recombinant" refers to a cell, nucleic acid, protein, or vector that has been altered by the introduction of an exogenous nucleic acid or the alteration of a naturally occurring nucleic acid. The resulting cell, nucleic acid, protein, or vector is considered a recombinant, and their progeny, descendants, duplicates, or copies are also considered recombinant. Thus, for example, a recombinant cell can express genes not found in the native (non-recombinant) form of the cell, or express native genes differently from those same genes expressed by a non-recombinant cell. Recombinant cells can include, but are not limited to, recombinant nucleic acids encoding gene products or inhibitory elements that reduce the level of an active gene product in a cell, such as mutations, knockouts, antisense, interfering RNA (RNAi), or dsRNA. "Recombinant nucleic acids" generally are derived from nucleic acids originally formed in vitro by manipulating nucleic acids, for example, with polymerases, ligases, exonucleases, and endonucleases, in a form not normally found in nature otherwise. Once a recombinant nucleic acid is generated and introduced into a host cell or organism, it can be replicated using the host cell's in vivo cellular machinery; however, such a nucleic acid, once produced recombinantly, although subsequently replicated within the cell, is still considered recombinant for purposes of this disclosure. In addition, recombinant nucleic acid refers to a nucleotide sequence that includes endogenous and exogenous nucleotide sequences; thus, an endogenous gene that has been recombined with an exogenous promoter is a recombinant nucleic acid. A "recombinant protein" is a protein made using recombinant techniques, i.e., by expression of a recombinant nucleic acid.
[0031] "Transformation" refers to the transfer of nucleic acid into a host organism or the genome of a host organism. Host organisms (and their progeny) containing the transformed nucleic acid fragments are referred to as "recombinant," "transgenic," or "transformed" organisms. Thus, the isolated polynucleotides of the present disclosure can be incorporated into recombinant constructs, typically DNA constructs, capable of introduction into and replication in a host cell. Such constructs can be vectors containing a replication system and sequences capable of transcription and translation of a polypeptide-encoding sequence in a given host cell. Typically, expression vectors contain one or more cloned genes under the transcriptional control of, for example, 5' and 3' regulatory sequences and a selectable marker. Such vectors may also contain a promoter regulatory region (e.g., a regulatory region controlling inducible or constitutive, environmentally or developmentally regulated, or site-specific expression), a transcription initiation start site, a ribosome binding site, a transcription termination site, and / or a polyadenylation signal. Alternatively, cells can be transformed with a single genetic element, such as a promoter, which, upon integration into the genome of the host organism by homologous recombination or the like, can result in genetically stable inheritance.
[0032] The term "transformed cell" refers to a cell that has undergone transformation. Thus, a transformed cell contains the parental genome and the inheritable genetic alteration. Embodiments include the progeny and descendants of such transformed cells.
[0033] The term "vector" refers to a means by which nucleic acids can be propagated and / or transferred between organisms, cells, or cellular components. Vectors include plasmids, linear DNA fragments, viruses, bacteriophages, proviruses, phagemids, transposons, artificial chromosomes, and the like, whether or not they are capable of autonomous replication or of integrating into a host cell chromosome.
[0034]
[0034] "Individual," "subject," and "patient" are used interchangeably and can refer to a human or a non-human.
[0035]
[0035] Throughout this application, the term "about" is used to indicate that a value includes the variation of error inherent in the measurement or quantification method.
[0036]
[0036] When used in conjunction with the term "comprising," the use of the words "a" or "an" can mean "one," but is also consistent with the meaning of "one or more," "at least one," and "one or more than one."
[0037] The phrase "and / or" means "and" or "or." Reciting A, B, and / or C includes: A only, B only, C only, a combination of A and B, a combination of A and C, a combination of B and C, or a combination of A, B, and C. In other words, "and / or" operates as inclusive.
[0038]
[0038] The words "comprising" (and all forms of comprising, e.g., "comprise" and "comprises"), "having" (and all forms of having, e.g., "have" and "has"), "including" (and all forms of including, e.g., "includes" and "include") or "containing" (and all forms of containing, e.g., "contains" and "contain") are inclusive or open-ended and do not exclude additional, unrecited elements or method steps.
[0039] Compositions and methods for their use can "comprise," "consist essentially of," or "consist" of any of the components or steps disclosed throughout this specification. Compositions and methods "consisting essentially of" any of the disclosed components or steps limit the scope of the claim to the specified materials or steps that do not materially affect the basic and novel characteristics of the claimed embodiment.
[0040] II. Proteins and Nucleic Acids As used herein, "protein" or "polypeptide" refers to a molecule comprising at least five amino acid residues. As used herein, the term "wild-type" refers to the endogenous version of a molecule that occurs naturally in an organism. In some embodiments, a wild-type version of a protein or polypeptide is employed, while in many embodiments of the present disclosure, a modified protein or polypeptide is employed. The above terms may be used interchangeably. A "modified protein" or "modified polypeptide" or "variant" refers to a protein or polypeptide whose chemical structure, particularly its amino acid sequence, has been altered relative to the wild-type protein or polypeptide. In some embodiments, a modified / variant protein or polypeptide has at least one modified activity or function (recognizing that a protein or polypeptide can have multiple activities or functions). It is specifically contemplated that a modified / variant protein or polypeptide can be altered with respect to one activity or function yet retain wild-type activities or functions in other respects.
[0041]
[0041] When a protein is specifically referred to herein, it generally refers to a native (wild-type) or recombinant (modified) protein, or a protein from which any signal sequence may be removed. A protein can be isolated directly from an organism in which it occurs naturally, produced by recombinant DNA / exogenous expression methods, or produced by solid-phase peptide synthesis (SPPS) or other in vitro methods. In certain embodiments, there are isolated nucleic acid segments and recombinant vectors incorporating nucleic acid sequences encoding polypeptides. The term "recombinant" can be used in conjunction with the name of a polypeptide or a specific polypeptide, and generally refers to a polypeptide produced from a nucleic acid molecule manipulated in vitro or a product of replication of such a molecule.
[0042] In certain embodiments, the size of the protein or polypeptide (wild type or modified) is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 1 5, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97 , 98, 99, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 525, 550, 575, 600, 625, 650, 675, 700, 725, 750, 775, 80 A domain may include, but is not limited to, 0, 825, 850, 875, 900, 925, 950, 975, 1000, 1100, 1200, 1300, 1400, 1500, 1750, 2000, 2250, 2500 amino acid residues or more, and any range derivable therein, or derivatives of the corresponding amino acid sequences described or referenced herein. Polypeptides can be mutated by truncation to make them shorter than their corresponding wild-type forms, and they can also be altered by fusing or conjugating heterologous protein or polypeptide sequences having a particular function (e.g., for targeting or localization, for enhanced immunogenicity, for purification purposes, etc.). As used herein, the term "domain" refers to any discrete functional or structural unit of a protein or polypeptide, and generally refers to a sequence of amino acids having a structure or function recognizable to one of skill in the art.
[0043] The term "polynucleotide" refers to a nucleic acid molecule that is either recombinant or isolated from total genomic nucleic acid. Included in the term "polynucleotide" are oligonucleotides (nucleic acids 100 residues or less in length), recombinant vectors, including plasmids, cosmids, phages, viruses, and the like. Polynucleotides, in certain aspects, contain regulatory sequences isolated substantially away from their naturally occurring gene or protein-coding sequences. Polynucleotides can be single-stranded (coding or antisense) or double-stranded, and can be RNA, DNA (genomic, cDNA, or synthetic), analogs thereof, or combinations thereof. Additional coding or non-coding sequences can be, but need not be, present within a polynucleotide.
[0044] In this regard, the terms "gene," "polynucleotide," or "nucleic acid" are used to refer to a nucleic acid that encodes a protein, polypeptide, or peptide (including any sequences required for proper transcription, post-translational modification, or localization). As will be understood by those skilled in the art, this term encompasses genomic sequences, expression cassettes, cDNA sequences, and smaller engineered nucleic acid segments that express, or can be adapted to express, proteins, polypeptides, domains, peptides, fusion proteins, and variants. A nucleic acid that encodes all or a portion of a polypeptide may contain a contiguous nucleic acid sequence that encodes all or a portion of such a polypeptide. It is also contemplated that a particular polypeptide may be encoded by a nucleic acid that contains variations having slightly different nucleic acid sequences, but still encode the same or substantially similar protein.
[0045] In certain embodiments, polynucleotide variants having substantial identity to the sequences disclosed herein exist, including at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or more sequence identity, including all values and ranges therebetween, compared to the polynucleotide sequences provided herein using the methods described herein (e.g., BLAST analysis using standard parameters). In certain aspects, the isolated polynucleotide will comprise a nucleotide sequence encoding a polypeptide having at least 90%, and in some cases 95% or more identity over the entire length of the sequence, to the amino acid sequences described herein; or a nucleotide sequence complementary to said isolated polynucleotide.
[0046]
[0046] Regardless of the length of the coding sequence itself, nucleic acid segments can be combined with other nucleic acid sequences, such as promoters, polyadenylation signals, additional restriction enzyme sites, multiple cloning sites, other coding segments, etc., and therefore their overall length can vary considerably. Nucleic acids can be of any length. They can be, for example, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 75, 100, 125, 175, 200, 250, 300, 350, 400, 450, 500, 750, 1000, 1500, 3000, 5000, or longer nucleotides in length, and / or can contain one or more additional sequences, such as regulatory sequences, and / or can be part of a larger nucleic acid, such as a vector. Thus, it is contemplated that nucleic acid fragments of almost any length can be employed, with the total length preferably being limited by the ease of preparation and use in the intended recombinant nucleic acid protocol. In some cases, the nucleic acid sequence may encode a polypeptide sequence with additional heterologous coding sequences, for example, to allow for purification, transport, secretion, post-translational modification of the polypeptide, or for therapeutic benefits such as targeting or efficacy. As discussed above, tags or other heterologous polypeptides can be added to the modified polypeptide-encoding sequence, where "heterologous" refers to a polypeptide that is not the same as the modified polypeptide.
[0047] A polypeptide, protein or polynucleotide encoding such a polypeptide or protein of the present disclosure can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 (or any derivable range therein) or more variant amino acid or nucleic acid substitutions, or can include at least or up to the sequence of SEQ ID NO: NO: 1-49, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 1 2, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116 6, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 21 61, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205,206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 300, 400, 500, 550, 1000 or more consecutive amino acids or nucleic acids, or derivatizable therein The nucleic acid sequence may be at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% (or any derivable range therein) similar, identical, or homologous to any suitable nucleic acid sequence.
[0048] In one embodiment, the protein or polypeptide has SEQ ID NO: 1-14 or 34-40 amino acids 1-2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121 7, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262,263、264、265、266、267、268、269、270、271、272、273、274、275、276、277、278、279、280、281、282、283、284、285、286、287、288、289、290、291、292、293、294、295、296、297、298、299、300、301、302、303、304、305、306、307、308、309、310、311、312、313、314、315、316、317、318、319、320、321、322、323、324、325、326、327、328、329、330、331、332、333、334、335、336、337、338、339、340、341、342、343、344、345、346、347、348、349、350、351、352、353、354、355、356、357、358、359、360、361、362、363、364、365、366、367、368、369、370、371、372、373、374、375、376、377、378、379、380、381、382、383、384、385、386、387、388、389、390、391、392、393、394、395、396、397、398、399、400、401、402、403、404、405、406、407、408、409、410、411、412、413、414、415、416、417、418、419、420、421、422、423、424、425、426、427、428、429、430、431、432、433、434、435、436、437、438、439、440、441、442、443、444、445、446、447、448、449、450、451、452、453、454、455、456、457、458、459、460、461、462、463、464、465、466、467、468、469、470、471、472、473、474、475、476、477、478、479、480、481、482、483、484、485、486、487、488、489、490、491、492、493、494、495、496、497、498、499、500、501、502、503、504、505、506、507、508、509、510、511、512、513, 514, 515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 557, 558, 559, 560, 561 , 562, 563, 564, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610 , 611, 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658, 659 , 660, 661, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699, or 700 (or any derivable range therein).
[0049] In one embodiment, the protein, polypeptide or nucleic acid has a sequence ID NO: 1 to 49, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 5, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 24 05, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264,265、266、267、268、269、270、271、272、273、274、275、276、277、278、279、280、281、282、283、284、285、286、287、288、289、290、291、292、293、294、295、296、297、298、299、300、301、302、303、304、305、306、307、308、309、310、311、312、313、314、315、316、317、318、319、320、321、322、323、324、325、326、327、328、329、330、331、332、333、334、335、336、337、338、339、340、341、342、343、344、345、346、347、348、349、350、351、352、353、354、355、356、357、358、359、360、361、362、363、364、365、366、367、368、369、370、371、372、373、374、375、376、377、378、379、380、381、382、383、384、385、386、387、388、389、390、391、392、393、394、395、396、397、398、399、400、401、402、403、404、405、406、407、408、409、410、411、412、413、414、415、416、417、418、419、420、421、422、423、424、425、426、427、428、429、430、431、432、433、434、435、436、437、438、439、440、441、442、443、444、445、446、447、448、449、450、451、452、453、454、455、456、457、458、459、460、461、462、463、464、465、466、467、468、469、470、471、472、473、474、475、476、477、478、479、480、481、482、483、484、485、486、487、488、489、490、491、492、493、494、495、496、497、498、499、500、501、502、503、504、505、506、507、508、509、510、511、512、513、514、515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 557, 558, 559, 560, 561, 562, 563 , 564, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610, 611, 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658, 659, 660, 661, 662, 663, 664, 665 2, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658, 659, 660, 661, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699, 700, 701, 702, 703, 704, 705, 706, 707, 708, 709, 710, 711, 712, 713, 71 The amino acid sequence may comprise 61, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699, or 700 contiguous amino acids (or any range derivable therein).
[0050] In certain embodiments, a polypeptide, protein, or nucleic acid has at least, at most, or exactly the same sequence as SEQ ID NO: NO: 1 to 49: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79 , 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185 3, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 24 02, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260,261、262、263、264、265、266、267、268、269、270、271、272、273、274、275、276、277、278、279、280、281、282、283、284、285、286、287、288、289、290、291、292、293、294、295、296、297、298、299、300、301、302、303、304、305、306、307、308、309、310、311、312、313、314、315、316、317、318、319、320、321、322、323、324、325、326、327、328、329、330、331、332、333、334、335、336、337、338、339、340、341、342、343、344、345、346、347、348、349、350、351、352、353、354、355、356、357、358、359、360、361、362、363、364、365、366、367、368、369、370、371、372、373、374、375、376、377、378、379、380、381、382、383、384、385、386、387、388、389、390、391、392、393、394、395、396、397、398、399、400、401、402、403、404、405、406、407、408、409、410、411、412、413、414、415、416、417、418、419、420、421、422、423、424、425、426、427、428、429、430、431、432、433、434、435、436、437、438、439、440、441、442、443、444、445、446、447、448、449、450、451、452、453、454、455、456、457、458、459、460、461、462、463、464、465、466、467、468、469、470、471、472、473、474、475、476、477、478、479、480、481、482、483、484、485、486、487、488、489、490、491、492、493、494、495、496、497、498、499、500、501、502、503、504、505、506、507、508、509、510、511, 512, 513, 514, 515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610, 611, 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658, 659, 660, 661, 662, 11, 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658, 659, 660, 661, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699, 700, 701, 702, 703, 704, 705, 706, 707, 708, 709, 710, 711, 712, 7 61, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699 or 700 (or any range derivable therein) contiguous amino acids, which is At least, at most, or exactly 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%,99% or 100% (or any range derivable therein) similar, identical, or homologous.
[0051] In one aspect, SEQ ID NO: Any of positions 1 to 49 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267,268、269、270、271、272、273、274、275、276、277、278、279、280、281、282、283、284、285、286、287、288、289、290、291、292、293、294、295、296、297、298、299、300、301、302、303、304、305、306、307、308、309、310、311、312、313、314、315、316、317、318、319、320、321、322、323、324、325、326、327、328、329、330、331、332、333、334、335、336、337、338、339、340、341、342、343、344、345、346、347、348、349、350、351、352、353、354、355、356、357、358、359、360、361、362、363、364、365、366、367、368、369、370、371、372、373、374、375、376、377、378、379、380、381、382、383、384、385、386、387、388、389、390、391、392、393、394、395、396、397、398、399、400、401、402、403、404、405、406、407、408、409、410、411、412、413、414、415、416、417、418、419、420、421、422、423、424、425、426、427、428、429、430、431、432、433、434、435、436、437、438、439、440、441、442、443、444、445、446、447、448、449、450、451、452、453、454、455、456、457、458、459、460、461、462、463、464、465、466、467、468、469、470、471、472、473、474、475、476、477、478、479、480、481、482、483、484、485、486、487、488、489、490、491、492、493、494、495、496、497、498、499、500、501、502、503、504、505、506、507、508、509、510、511、512、513、514、515、516、517、518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554, 555, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565 5, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610, 611, 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658, 659, 66 starting at 0, 661, 662, 663, 664, 665, 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699 or 700 and continuing at least, at most, or exactly at SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78,79、80、81、82、83、84、85、86、87、88、89、90、91、92、93、94、95、96、97、98、99、100、101、102、103、104、105、106、107、108、109、110、111、112、113、114、115、116、117、118、119、120、121、122、123、124、125、126、127、128、129、130、131、132、133、134、135、136、137、138、139、140、141、142、143、144、145、146、147、148、149、150、151、152、153、154、155、156、157、158、159、160、161、162、163、164、165、166、167、168、169、170、171、172、173、174、175、176、177、178、179、180、181、182、183、184、185、186、187、188、189、190、191、192、193、194、195、196、197、198、199、200、201、202、203、204、205、206、207、208、209、210、211、212、213、214、215、216、217、218、219、220、221、222、223、224、225、226、227、228、229、230、231、232、233、234、235、236、237、238、239、240、241、242、243、244、245、246、247、248、249、250、251、252、253、254、255、256、257、258、259、260、261、262、263、264、265、266、267、268、269、270、271、272、273、274、275、276、277、278、279、280、281、282、283、284、285、286、287、288、289、290、291、292、293、294、295、296、297、298、299、300、301、302、303、304、305、306、307、308、309、310、311、312、313、314、315、316、317、318、319、320、321、322、323、324、325、326、327、328、329、330、331、332、333、334、335、336、337、338、339、340、341、342、343、344、345、346、347、348、349、350、351、352、353、354、355、356、357、358、359、360、361、362、363、364、365、366、367、368、369、370、371、372、373、374、375、376、377、378、379、380、381、382、383、384、385、386、387、388、389、390、391、392、393、394、395、396、397、398、399、400、401、402、403、404、405、406、407、408、409、410、411、412、413、414、415、416、417、418、419、420、421、422、423、424、425、426、427、428、429、430、431、432、433、434、435、436、437、438、439、440、441、442、443、444、445、446、447、448、449、450、451、452、453、454、455、456、457、458、459、460、461、462、463、464、465、466、467、468、469、470、471、472、473、474、475、476、477、478、479、480、481、482、483、484、485、486、487、488、489、490、491、492、493、494、495、496、497、498、499、500、501、502、503、504、505、506、507、508、509、510、511、512、513、514、515、516、517、518、519、520、521、522、523、524、525、526、527、528、529、530、531、532、533、534、535、536、537、538、539、540、541、542、543、544、545、546、547、548、549、550、551、552、553、554、555、556、557、558、559、560、561、562、563、564、565、566、567、568、569、570、571、572、573、574、575、576、577、578、579、580、581、582、583、584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, 604, 605, 606, 607, 608, 609, 610, 611, 612, 613, 614, 615, 616, 617, 618, 619, 620, 621, 622, 623, 624, 625, 626, 627, 628, 629, 630, 631, 632, 633, 634, 635, 636, 637, 638, 639, 640, 641, 642, 643, 644, 645, 646, 647, 648, 649, 650, 651, 652, 653, 654, 655, 656, 657, 658, 659, 660, 661, 662, 663, 664, 665 , 666, 667, 668, 669, 670, 671, 672, 673, 674, 675, 676, 677, 678, 679, 680, 681, 682, 683, 684, 685, 686, 687, 688, 689, 690, 691, 692, 693, 694, 695, 696, 697, 698, 699 or 700 (or any range derivable therein) consecutive amino acids or nucleotides.
[0052] Nucleotide and protein, polypeptide and peptide sequences for various genes have been previously disclosed and can be found in recognized computerized databases. Two commonly used databases are the National Center for Biotechnology Information's Genbank and GenPept databases (ncbi.nlm.nih.gov / on the World Wide Web) and the Universal Protein Resource (UniProt; uniprot.org on the World Wide Web). The coding regions for these genes can be amplified and / or expressed using the techniques disclosed herein or as would be known to those skilled in the art.
[0053] It is contemplated that the compositions of the present disclosure have about 0.001 mg to about 10 mg of total polypeptide, peptide, and / or protein per ml. The concentration of protein in the composition can be about (at least about or at most about) 0.001, 0.010, 0.050, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.5, 2.0, 2.5, 3.0, 3.5, 4.0, 4.5, 5.0, 5.5, 6.0, 6.5, 7.0, 7.5, 8.0, 8.5, 9.0, 9.5, 10.0 mg / ml or more (or any range derivable therein).
[0054]
[0054] In the case of proteins with catalytic activity (e.g., enzymes), such proteins can be described using the Enzyme Classification (EC) nomenclature. The EC classification of various enzymes has been previously disclosed and can be found in recognized databases, such as the ENZYME database (Bairoch A. The ENZYME database in 2000. Nucleic Acids Res. 2000 Jan 1;28(1):304-5. doi: 10.1093 / nar / 28.1.304; incorporated herein by reference in its entirety).
[0055] A. Signal Peptide Aspects of the present disclosure are directed to synthetic signal peptides and polynucleotides and nucleic acids encoding such signal peptides. Also disclosed are cells containing such signal peptides and methods of using the cells in the production and secretion of proteins (e.g., mammalian proteins such as human milk proteins). As used herein, "signal peptide" (or "signal peptide sequence") describes any peptide that, when present at the N-terminus of a newly synthesized polypeptide, is capable of directing the polypeptide across or into a cell's cellular membrane (e.g., the plasma membrane, the endoplasmic reticulum membrane, etc.). In certain aspects, signal peptides of the present disclosure are capable of directing a polypeptide into the cell's secretory pathway and subsequent secretion of the polypeptide (referred to herein as "secretory signal peptides").
[0056] As described herein, aspects of the present disclosure relate to a synthetic signal peptide comprising: (a) Preregion sequence from: (i) P. pastoris Ost1; or (ii) P. pastoris Pst1; and (b) Pro-region sequence from (i) Saccharomyces cerevisiae mating factor α (MFα); or (ii) P. pastoris Epx1.
[0057]
[0057] Particular signal peptides of the present disclosure are listed in Table 1 below.
[0058] [Table 1-1]
[0059] [Table 1-2]
[0058] In one aspect, polypeptides comprising the signal peptides of the present disclosure are disclosed. Nucleic acids encoding such polypeptides are also disclosed. Additionally, cells expressing polypeptides comprising the signal peptides of the present disclosure are also disclosed.
[0060] In certain aspects, a polypeptide of the present disclosure comprises SEQ ID NO: 1. In certain embodiments, a polypeptide of the present disclosure comprises a sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 1. In certain aspects, a polypeptide of the present disclosure comprises a sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions (or more) relative to SEQ ID NO: 1.
[0061] In certain aspects, a polypeptide of the present disclosure comprises SEQ ID NO:2. In certain embodiments, a polypeptide of the present disclosure comprises a sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO:2. In certain aspects, a polypeptide of the present disclosure comprises a sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions (or more) relative to SEQ ID NO:2.
[0062] In certain aspects, a polypeptide of the present disclosure comprises SEQ ID NO:3. In certain embodiments, a polypeptide of the present disclosure comprises a sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO:3. In certain aspects, a polypeptide of the present disclosure comprises a sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions (or more) relative to SEQ ID NO:3.
[0063] In certain aspects, a polypeptide of the present disclosure comprises SEQ ID NO:4. In certain embodiments, a polypeptide of the present disclosure comprises a sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO:4. In certain aspects, a polypeptide of the present disclosure comprises a sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions (or more) relative to SEQ ID NO:4.
[0064] Any one or more of the signal peptides disclosed herein may be excluded from certain embodiments.
[0065] B. Secreted proteins Aspects of the present disclosure include secreted proteins (also referred to as "secreted proteins"), as well as compositions comprising secreted proteins, methods of expressing secreted proteins, and methods of their use. As used herein, "secreted protein" describes any protein that is secreted outside of a cell. In certain cases, secreted proteins of the present disclosure are proteins present in human secretions, such as colostrum, milk, tears, semen, vaginal fluid, saliva, or other secretions. In some aspects, secreted proteins of the present disclosure are human milk proteins. In some aspects, secreted proteins of the present disclosure are not human milk proteins.
[0066] 1. Human milk protein Aspects of the present disclosure include human milk proteins, as well as compositions comprising the human milk proteins (e.g., infant formula compositions), methods of producing the human milk proteins, and methods of use thereof. In one aspect, cells are disclosed that express a human milk protein (e.g., comprising SEQ ID NO: 1, 2, 3, or 4) linked to a signal peptide of the disclosure. As used herein, "human milk protein" describes any protein present in human breast milk. Human milk proteins include proteins derived from (e.g., isolated from) human breast milk, as well as any protein produced by other means (e.g., recombinant expression, chemical synthesis, etc.) that has the amino acid sequence of a protein present in human breast milk. A variety of human milk proteins are recognized in the art and are contemplated herein. The human milk proteins contemplated herein include, but are not limited to, secretory IgA (sIgA), human serum albumin, xanthine dehydrogenase, lactoferrin, lactoperoxidase, butyrophilin, lactadherin, adiponectin, β-casein, κ-casein, leptin, lysozyme, and α-lactalbumin.In some embodiments, the human milk proteins of the present disclosure are human whey proteins.In some embodiments, the human milk proteins of the present disclosure are recombinant human milk proteins (e.g., produced by non-mammalian cells such as yeast cells).
[0067] Certain aspects of the present disclosure are directed to human milk proteins having "human-like" glycans. Human-like glycans (also referred to as "human-like glycan structures") describe glycans having structures present on human glycoproteins. Such glycans include, for example, hybrid N-glycans, complex N-glycans, biantennary, triantennary, and tetraantennary N-glycans, as well as glycans containing sialic acid, galactose, N-acetylgalactosamine, or fucose. Human-like glycans include those having a Man3GlcNAc2 core structure. Thus, the human milk proteins of the present disclosure include those having one or more human-like glycans, such as hybrid N-glycans, complex N-glycans, biantennary N-glycans, triantennary N-glycans, tetraantennary N-glycans, and combinations thereof.
[0068]
[0067] Accordingly, in some embodiments, recombinant human milk proteins (e.g., recombinant human lactoferrin) comprising one or more human-like glycans are disclosed. Such recombinant proteins include recombinant proteins produced by engineered mammalian, fungal, yeast, bacterial, or other cells, including, for example, the engineered cells described elsewhere herein. In certain aspects, such recombinant proteins have a glycan pattern that differs from the glycan pattern of the corresponding native human milk protein. For example, in some embodiments, recombinant human lactoferrin comprising one or more human-like glycans is disclosed, wherein the lactoferrin has a glycan pattern that differs from the glycan pattern of any naturally occurring human lactoferrin (e.g., human lactoferrin in human breast milk).
[0069] Lactoferrin Aspects of the present disclosure are directed to lactoferrin and compositions comprising lactoferrin, including infant formula compositions. In one aspect, cells expressing human lactoferrin linked to a signal peptide of the present disclosure (e.g., comprising SEQ ID NO:1, 2, 3, or 4) are disclosed. Lactoferrin (also called "lactotransferrin") is a whey protein found in exocrine fluids, such as breast milk, and is encoded by the LTF gene. Without wishing to be bound by theory, lactoferrin is understood to have antimicrobial and anti-inflammatory properties. Certain aspects of the present disclosure are directed to human lactoferrin (UniProtKB / Swiss-Prot Accession No. P02788), including its isoforms. The full sequence of human lactoferrin, including the signal peptide, is provided as SEQ ID NO:34. The sequence of mature human lactoferrin after cleavage of the signal peptide is provided as SEQ ID NO:9.
[0070] [Table 2-1]
[0071] [Table 2-2]
[0072] [Table 2-3]
[0073] [Table 2-4]
[0074] [Table 2-5]
[0075] [Table 2-6]
[0076]
Table 2-7
[0077]
Table 2-8
[0078]
Table 2-9
[0079]
Table 2-10
[0080]
Table 2-11
[0069] In some aspects, the human lactoferrin of the present disclosure is recombinant human lactoferrin (rhLactoferrin). In some aspects, the recombinant human lactoferrin of the present disclosure is obtained from mammalian, fungal, yeast, bacterial, or other cells. In some aspects, the recombinant human lactoferrin of the present disclosure is not obtained from mammalian cells. In certain aspects, the recombinant human lactoferrin of the present disclosure is obtained from fungal cells. The fungal cell can be, for example, a cell of the genus Arcthura, Aspergillus, Aurantiochytrium, Candida, Cryptococcus, Cryptococcus, Geotrichum, Hansenula, Kluyveromyces, Kodamaea, Komagataella, Leucosporidia, Lipomyces, Mortierella, Ogataea, Pichia, Prototheca, Rhizopus, Rhodosporidium, Rhodotorula, Saccharomyces, Schizosaccharomyces, Tremella, Trichosporon, Wickerhamomyces, or Yarrowia. In one aspect, the fungal cell is a yeast cell. In one aspect, the yeast cell is a Komagataella cell (Komagataella puffii, Komagataella pastoris, Komagataella pseudopastoris). Additional cells suitable for recombinant protein production are recognized in the art and are contemplated herein. In one aspect, the recombinant human lactoferrin of the present disclosure is obtained from a bacterial cell. In another aspect, the human lactoferrin of the present disclosure is isolated from a natural source.
[0081]
[0070] Certain aspects of the present disclosure are directed to human lactoferrin having at least one hybrid or complex N-glycan. In some aspects, the human lactoferrin comprises glycans containing one or more of sialic acid, galactose, N-acetylgalactosamine, or fucose. In some aspects, the human lactoferrin comprises biantennary, triantennary, or tetraantennary N-glycans. As disclosed herein, human lactoferrin having one or more hybrid, complex, biantennary, triantennary, or tetraantennary N-glycans can be useful, for example, in infant formula or other nutritional compositions or supplements.
[0082] B. Alpha-lactalbumin (α-lactalbumin) Aspects of the present disclosure are directed to alpha-lactalbumin and compositions comprising alpha-lactalbumin, including infant formula compositions. In one aspect, cells expressing human alpha-lactalbumin linked to a signal peptide of the present disclosure (e.g., comprising SEQ ID NO: 1, 2, 3, or 4) are disclosed. Alpha-lactalbumin (also referred to as "α-lactalbumin") is a whey protein found in breast milk and is encoded by the LALBA gene. Certain aspects of the present disclosure are directed to human α-lactalbumin (UniProtKB / Swiss-Prot Accession No. P00709), including its isoforms. The full sequence of human α-lactalbumin, including the signal peptide, is provided as SEQ ID NO: 36. The sequence of mature human α-lactalbumin after cleavage of the signal peptide is provided as SEQ ID NO: 35.
[0083] [Table 3] In some aspects, the human alpha-lactalbumin of the present disclosure is recombinant human alpha-lactalbumin. In some aspects, the recombinant human alpha-lactalbumin of the present disclosure is obtained from mammalian, fungal, yeast, bacterial, or other cells. In some aspects, the recombinant human alpha-lactalbumin of the present disclosure is not obtained from mammalian cells. In certain aspects, the recombinant human alpha-lactalbumin of the present disclosure is obtained from yeast cells. The yeast cell can be, for example, a cell of the genus Arcthura, Aspergillus, Aurantiochytrium, Candida, Cryptococcus, Cryptococcus, Geotrichum, Hansenula, Kluyveromyces, Kodamaea, Komagataella, Leucosporidium, Lipomyces, Mortierella, Ogataea, Pichia, Prototheca, Rhizopus, Rhodosporidium, Rhodotorula, Saccharomyces, Schizosaccharomyces, Tremella, Trichosporon, Wickerhamomyces, or Yarrowia. In one aspect, the yeast cell is a Komagataella cell (e.g., Komagataella paphii, Komagataella pastoris, Komagataella pseudopastoris). Additional yeast cells suitable for recombinant protein production are recognized in the art and are contemplated herein. In one aspect, the recombinant human alpha-lactalbumin of the present disclosure is obtained from bacterial cells. In another aspect, the human alpha-lactalbumin of the present disclosure is isolated from natural sources.
[0084]
[0073] Certain aspects of the present disclosure are directed to human α-lactalbumin having at least one hybrid or complex N-glycan.In some aspects, human α-lactalbumin comprises glycans comprising one or more of sialic acid, galactose, N-acetylgalactosamine or fucose.In some aspects, human lactoferrin comprises biantennary, triantennary or tetraantennary N-glycans.As disclosed herein, human α-lactalbumin having one or more hybrid, complex, biantennary, triantennary or tetraantennary N-glycans can be useful, for example, in infant formula or other nutritional compositions or supplements.
[0085] c. Additional human milk proteins Additional human milk proteins contemplated in the compositions (e.g., infant formula compositions) and methods of the present disclosure include, but are not limited to, secretory IgA (sIgA), human serum albumin, xanthine dehydrogenase, lactoperoxidase, butyrophilin, lactadherin, adiponectin, β-casein, κ-casein, leptin, osteopontin, bile salt-stimulated lipase (BSSL), and lysozyme. Any one or more of these human milk proteins may be included in the compositions (e.g., infant formulas) of the present disclosure. Any one or more of these human milk proteins may be excluded in certain embodiments.
[0086] CN-acetylglucosaminyltransferase Aspects of the present disclosure relate to N-acetylglucosaminyltransferase proteins. As used herein, "N-acetylglucosaminyltransferase protein" describes any polypeptide having N-acetylglucosaminyltransferase activity. N-acetylglucosaminyltransferase describes an enzyme that catalyzes the transfer of a monosaccharide from a specific sugar nucleotide donor to a specific hydroxyl position of the monosaccharide in a growing glycan chain in one of two possible anomeric linkages (either α or β).
[0087] The N-acetylglucosaminyltransferase protein can be an N-acetylglucosaminyltransferase protein from any suitable organism. In one aspect, the N-acetylglucosaminyltransferase protein is a eukaryotic N-acetylglucosaminyltransferase protein. In one aspect, the N-acetylglucosaminyltransferase protein is a mammalian N-acetylglucosaminyltransferase protein.
[0088] 1. N-acetylglucosaminyltransferase I In some embodiments, the N-acetylglucosaminyltransferase protein is an N-acetylglucosaminyltransferase I protein (EC 2.4.1.101). The systematic name for this enzyme class is alpha-1,3-mannosyl-glycoprotein beta-1,2-N-acetylglucosaminyltransferase. Other names include: GnT-I, N-acetylglucosaminyltransferase I, and uridine diphosphoacetylglucosamine-alpha-1,3-mannosylglycoprotein beta-1,2-N-acetylglucosaminyltransferase. In certain embodiments, the N-acetylglucosaminyltransferase I protein of the present disclosure is Homo sapiens GnT-I, although N-acetylglucosaminyltransferase I proteins from any eukaryotic organism can be used as part of the methods and compositions of the present disclosure.
[0089] 2. β-1,2-N-acetylglucosaminyltransferase In certain embodiments, the N-acetylglucosaminyltransferase protein is a β-1,2-N-acetylglucosaminyltransferase protein (EC 2.4.1.143). The systematic name for this enzyme class is alpha-1,6-mannosyl-glycoprotein 2-beta-N-acetylglucosaminyltransferase. Other names include: GnT-II, N-acetylglucosaminyltransferase II, and uridine diphosphoacetylglucosamine-alpha-1,6-mannosylglycoprotein beta-1-2-N-acetylglucosaminyltransferase. In a specific embodiment, the β-1,2-N-acetylglucosaminyltransferase protein of the present disclosure is Rattus norvegicus GnT-II, although β-1,2-N-acetylglucosaminyltransferase proteins from any eukaryotic organism can be used as part of the methods and compositions of the present disclosure.
[0090] D. Alpha-1,3 / 6-mannosidase (α-1,3 / 6-mannosidase) Aspects of the present disclosure relate to α-1,3 / 6-mannosidase proteins (EC 3.2.114). As used herein, "α-1,3 / 6-mannosidase protein" (or "alpha-1,3 / 6-mannosidase protein") describes any polypeptide having α-1,3 / 6-mannosidase activity. α-1,3 / 6-mannosidase describes an enzyme that catalyzes the removal of two mannosyl residues from an N-glycan. The systematic name for this enzyme class is mannosyloligosaccharide 1,3-1,6-alpha-mannosidase. Other names include: Man-II and mannosidase II. α-1,3 / 6-mannosidase proteins can be derived from any suitable organism. In some embodiments, the α-1,3 / 6-mannosidase protein is a eukaryotic α-1,3 / 6-mannosidase protein. In certain embodiments, the α-1,3 / 6-mannosidase protein is Drosophila Man-II, although α-1,3 / 6-mannosidase proteins from any eukaryotic organism can be used as part of the methods and compositions of the present disclosure.
[0091] E. Alpha-1,2-mannosidase (α-1,2-mannosidase) Aspects of the present disclosure relate to α-1,2-mannosidase proteins (EC 3.2.1.130). As used herein, "α-1,2-mannosidase protein" (or "alpha-1,2-mannosidase protein") describes any polypeptide having α-1,2-mannosidase activity. The systematic name for this enzyme class is glycoprotein endo-alpha-1,2-mannosidase. Other names include: endo-alpha-D-mannosidase and Man-I. In some embodiments, the α-1,2-mannosidase protein is fungal Man-I. In particular embodiments, the Man-I is Trichoderma reesei Man-I.
[0092] F. Beta-1,4-galactosyltransferase (β-1,4-galactosyltransferase) Aspects of the present disclosure relate to β-1,4-galactosyltransferase proteins (EC 2.4.1.38). As used herein, "β-1,4-galactosyltransferase protein" (or "beta-1,4-galactosyltransferase protein") describes any polypeptide having β-1,4-galactosyltransferase activity. The systematic name for this enzyme class is beta-N-acetylglucosaminylglycopeptide beta-1,4-galactosyltransferase. Other names include: glycoprotein 4-beta-galactosyltransferase, UDP-galactose-glycoprotein galactosyltransferase, and GalT. In some embodiments, the β-1,4-galactosyltransferase protein is a mammalian GalT. In certain embodiments, the GalT is a Homo sapiens GalT.
[0093] G. Glycosylated Proteins Aspects of the present disclosure are directed to methods and compositions for producing glycosylated proteins (also referred to as "glycoproteins") having glycosylation patterns similar to those of glycoproteins produced by human cells. In certain embodiments, the glycoproteins of the present disclosure are N-linked glycoproteins. N-linked glycoproteins contain N-acetylglucosamine residues attached to the amide nitrogen of asparagine residues in the protein. The predominant sugars found on glycoproteins are glucose, galactose, mannose, fucose, N-acetylgalactosamine (GalNAc), N-acetylglucosamine (GlcNAc), and sialic acid, e.g., N-acetylneuraminic acid (NANA). Processing of sugar groups occurs cotranslationally in the lumen of the ER and, for N-linked glycoproteins, continues in the Golgi apparatus.
[0094] H. Protein Targeting Certain aspects of the present disclosure include cells that express one or more proteins from nucleic acid molecules, where the proteins are targeted to a desired subcellular location (e.g., an organelle such as the Golgi apparatus). In some cases, the protein is targeted to a subcellular location by forming a fusion protein that includes a portion of the protein (e.g., the catalytic domain of an enzyme) and a cellular targeting signal peptide, such as a heterologous signal peptide that is not normally ligated to or associated with the portion of the protein (e.g., a signal peptide comprising SEQ ID NO:1, 2, 3, or 4). A fusion protein can be encoded by ligating a polynucleotide encoding the cellular targeting signal peptide in the same translational reading frame ("in frame") to a nucleic acid fragment encoding the protein (e.g., an enzyme) or a catalytically active fragment thereof.
[0095] The targeting signal peptide component of the fusion construct or protein can be derived from ER or Golgi membrane-associated proteins, retrieval signals, type II membrane proteins, type I membrane proteins, membrane-spanning nucleotide sugar transporters, mannosidases, sialyltransferases, glucosidases, mannosyltransferases, and phosphomannosyltransferases. In one aspect, the targeting signal peptide is a Golgi apparatus localization tag. Examples of Golgi apparatus localization tags include, but are not limited to, transmembrane domains from Saccharomyces cerevisiae Kre2p, Saccharomyces cerevisiae Mnn2p, Saccharomyces cerevisiae Mnn9, Komagataella puffii Bmt2, Komagataella puffii Bmt3, or Komagataella puffii Ktr2.
[0096] III. Sequence
[0085] Specific examples of polypeptide and nucleic acid sequences contemplated herein are shown in Table 4 below.
[0097] [Table 4-1]
[0098] [Table 4-2]
[0099] [Table 4-3]
[0100] [Table 4-4]
[0101] [Table 4-5]
[0102] [Table 4-6]
[0103] [Table 4-7]
[0104] [Table 4-8] IV. Genetic Engineering
[0086] Vectors for transforming microorganisms (e.g., fungal cells, yeast cells) according to the present disclosure can be prepared by known techniques well known to those skilled in the art in light of the disclosure herein. Vectors typically contain one or more genes, where each gene codes for expression of a desired product (gene product) and is operably linked to one or more control sequences that regulate gene expression or target the gene product to a specific location in the recombinant cell.
[0105] For example, exogenous nucleic acid sequences, including nucleic acid sequences encoding fusion proteins, nucleic acid sequences encoding wild-type or mutant proteins, can be introduced into many different host cells. Nucleic acid sequences configured to promote genetic mutation in genes can also be introduced into various host cells, as further described herein. Suitable host cells are microbial hosts that can be widely found in the fungal family. Examples of suitable host strains include, but are not limited to, fungal or yeast species, such as Arctura, Aspergillus, Aurantiochytrium, Candida, Bacillus, Cryptococcus, Cuscuta, Hansenula, Kluyveromyces, Komagataella, Leucosporidium, Lipomyces, Mortierella, Ogataea, Pichia, Prototheca, Rhizopus, Rhodosporidium, Rhodotorula, Saccharomyces, Schizosaccharomyces, Tremella, Trichosporon, and Yarrowia. In some embodiments, the host cell of the present disclosure is a Komagataella cell. In some embodiments, the host cell of the present disclosure is Komagataella puffii. In some embodiments, the host cell of the present disclosure is Komagataella pastoris. The host cell of the present disclosure is Komagataella pseudopastoris.
[0106] Microbial expression systems and expression vectors are well known to those skilled in the art. Any such expression vector can be used to introduce instant genes and nucleic acid sequences into organisms. The nucleic acid sequence can be introduced into a suitable microorganism by transformation techniques. For example, the nucleic acid sequence can be cloned into a suitable plasmid, and the parent cell can be transformed with the resulting plasmid. The plasmid is not particularly limited, as long as it allows the desired nucleic acid sequence to be inherited by the progeny of the microorganism.
[0107]
[0089] Vectors or cassettes useful for transforming suitable host cells are recognized in the art. Typically, vectors or cassettes contain genes, sequences that direct the transcription and translation of related genes, including promoters, selectable markers, and sequences that allow autonomous replication or chromosomal integration. Suitable vectors include the 5' region of the gene with promoters and other transcription initiation controls, and the 3' region of the DNA fragment that controls transcription termination.
[0108]
[0090] The promoter, cDNA, and 3'UTR, as well as other elements of the vector, can be produced by cloning techniques using fragments isolated from natural sources (Green & Sambrook, Molecular Cloning: A Laboratory Manual, (4th ed., 2012); U.S. Pat. No. 4,683,202; incorporated by reference). Alternatively, the elements can be produced synthetically using known methods (Gene 164:49-53 (1995)).
[0109] A. Vectors and Vector Components
[0091] Vectors for transforming microorganisms (e.g., yeast cells) according to the present disclosure can be prepared by known techniques well known to those skilled in the art in view of the disclosure herein. Vectors typically contain one or more genes, where each gene encodes the expression of a desired product (gene product) and is operably linked to one or more control sequences (e.g., promoter sequences, signal peptide sequences) that regulate gene expression or target the gene product to a specific location in the recombinant cell.
[0110] 1. Control arrays
[0092] A control sequence is a nucleic acid sequence that regulates the expression of a coding sequence or directs a gene product to a specific location within or outside a cell. Control sequences that regulate expression include, for example, promoters, which control transcription of a coding sequence, and terminators, which terminate transcription of a coding sequence. Another control sequence is a 3' untranslated sequence located at the end of a coding sequence that encodes a polyadenylation signal. Control sequences that direct a gene product to a specific location include those that encode signal peptides, which direct proteins to which they are bound to a specific location within or outside the cell.
[0111] Thus, an example of a vector design for expressing a gene in a microorganism contains a coding sequence for a desired gene product (e.g., a selectable marker, an enzyme, a fusion protein, etc.) operably linked to a promoter active in yeast. Alternatively, if the vector does not contain a promoter operably linked to the coding sequence of interest, the coding sequence can be transformed into the cell so that it is operably linked to an endogenous promoter at the time of vector integration. Examples of promoters contemplated herein include, but are not limited to, the AOX1, GAP, TEF1, TPI1, DAS1, DAS2, CAT1, and FMD promoters.
[0112]
[0094] The promoter used to express a gene can be the promoter naturally linked to that gene or a different promoter.
[0113] Promoters are generally characterized as constitutive or inducible. Constitutive promoters are generally active or function to drive expression at the same level at all times (or at specific times in the cell's life cycle). Inducible promoters, conversely, are active (or inactivated) or are significantly up- or down-regulated only in response to a stimulus. Both types of promoters are applicable to the disclosed methods. Useful inducible promoters include those that mediate transcription of an operably linked gene in response to a stimulus such as an exogenously provided small molecule, temperature (heat or cold), or nitrogen deficiency in the culture medium. Suitable promoters can activate transcription of an essentially inactive gene or up-regulate transcription of an operably linked gene that is transcribed at low levels.
[0114] The insertion of a termination region regulatory sequence is optional. The termination region may be native to the transcription initiation region (promoter), may be native to the DNA sequence of interest, or may be available from another source (see, e.g., Chen & Orozco, Nucleic Acids Research 16:8411 (1988)).
[0115] In some cases, the entire nucleotide sequence of a promoter is not required to drive transcription, and a sequence shorter than the entire nucleotide sequence of a promoter can drive transcription of an operably linked gene. The smallest portion of a promoter is called the core promoter and includes the transcription initiation site, the binding site for RNA polymerase, and the binding site for transcription factors.
[0116] The promoter can be linked to the target by introducing the promoter and the target into a nucleic acid molecule, for example, a vector.The vector is introduced into cell, thereby expressing the promoter and the target.In one embodiment, the promoter is linked to the target by, for example, introducing the promoter into the DNA of cell via homologous recombination, thereby integrating the promoter into the genome of cell.
[0117] B. Gene and Codon Optimization
[0099] Typically, a gene comprises a promoter, a coding sequence, and a termination control sequence. When assembled by recombinant DNA technology, the gene can be called an expression cassette, and can be flanked by restriction sites for convenient insertion into a vector used to introduce the recombinant gene into a host cell. The expression cassette can be flanked by DNA sequences from a genome or other nucleic acid target to facilitate stable integration of the expression cassette into the genome by homologous recombination. Alternatively, the vector and its expression cassette can remain non-integrated (e.g., episomal), in which case the vector typically comprises an origin of replication capable of providing replication of the vector DNA.
[0118]
[0100] A common gene present on a vector is a protein-encoding gene, the expression of which allows recombinant cells containing the protein to be differentiated from cells that do not express the protein. Such a gene and its corresponding gene product are called selectable markers or selection markers. Any of a wide variety of selectable markers can be utilized in transgene constructs useful for transforming organisms covered in the disclosed embodiments.
[0119] For optimal expression of recombinant proteins, it may be beneficial to utilize a coding sequence that produces mRNA with codons optimally used by the host cell to be transformed. Thus, proper expression of a transgene may require that the codon usage of the transgene matches the specific codon bias of the organism in which the transgene is expressed. The exact mechanisms underlying this effect are numerous, but include a proper balance between the available aminoacylated tRNA pool and the proteins synthesized in the cell, coupled with more efficient translation of transgenic messenger RNA (mRNA) when this requirement is met. If the codon usage in the transgene is not optimized, the available tRNA pool may not be sufficient to enable efficient translation of the transgenic mRNA, resulting in ribosome stalling and termination, and possible instability of the transgenic mRNA.
[0120] The coding sequences of the present disclosure can be codon-optimized for a particular host cell by replacing one or more rare codons with one or more codons that are more frequently found in the host cell. A rare codon in a host cell describes a codon that is found in less than 5%, less than 10%, or less than 20% of the coding sequences in the host cell. Rare codons can be identified using methods known to those skilled in the art.
[0121] Aspects of the present disclosure include the transformation of a microorganism with a nucleic acid sequence comprising a gene encoding a protein. The gene can be native to the cell or from a different species. The gene can be from a different species but modified (e.g., codon-optimized) for optimal expression in the microorganism. In certain embodiments, the gene is heritable to progeny of the transformed cell. In some embodiments, the gene is heritable because it is present on a plasmid. In certain embodiments, the gene is heritable because it is integrated into the genome of the transformed cell.
[0122]
[0104] Further aspects of the present disclosure may include transforming a microorganism with a nucleic acid sequence configured to generate a mutation in a gene of the microorganism. For example, aspects of the present disclosure may include transforming a microorganism with a nucleic acid sequence comprising upstream and downstream sequences of a gene (e.g., the OCH1 gene), thereby facilitating the reduction or deletion of gene expression via homologous recombination. Various methods for generating mutations in a gene of a microorganism (including deletion or knockout mutations, as well as mutations that reduce gene expression) are recognized in the art and are contemplated herein. A microorganism with a gene deletion or knockout mutation does not produce a functional copy of the protein. For example, a recombinant yeast cell of the present disclosure may include a deletion of the endogenous OCH1 gene such that the recombinant yeast cell does not express the endogenous functional OCH1 protein. A microorganism with reduced expression of a gene or protein produces a functional copy of the protein, but in a reduced amount compared to a wild-type (i.e., non-recombinant or non-genetically modified) microorganism of the same species. Methods for reducing protein expression are art-recognized and include, for example, replacing the endogenous promoter and / or modifying one or more regulatory elements.
[0123] C. Transformation
[0105] Cells can be transformed by any suitable technique, including, for example, gene gun, electroporation, glass bead transformation, and silicon carbide whisker transformation. Any convenient technique for introducing transgenes into microorganisms can be employed in the embodiments disclosed herein.
[0124]
[0106] Vectors for microbial transformation can be prepared by known techniques well known to those skilled in the art. In one embodiment, an exemplary vector design for expressing a gene in a microorganism contains a gene encoding an enzyme operably linked to a promoter active in the microorganism. Alternatively, if the vector does not contain a promoter operably linked to the gene of interest, the gene can be transformed into a cell so that it is operably linked to its native promoter at the time of vector integration. The vector can also contain a second gene encoding a protein. Optionally, one or both gene(s) are followed by a 3' untranslated sequence containing a polyadenylation signal. Expression cassettes encoding the two genes can be physically linked within the vector or on separate vectors. Co-transformation of microorganisms can also be used, in which separate vector molecules are used simultaneously to transform cells (Protist 155:381-93 (2004)). Transformed cells can optionally be selected based on their ability to grow in the presence of an antibiotic or other selectable marker under conditions in which cells lacking the resistance cassette would not grow.
[0125] D. Genetically engineered cells Aspects of the present disclosure include genetically engineered cells (also called "engineered cells" or "recombinant cells") and methods for making and using such cells. In certain embodiments, recombinant cells comprising one or more exogenous nucleic acid sequences are disclosed. Methods for producing such recombinant cells comprising introducing one or more exogenous nucleic acid sequences into a host cell are also disclosed. Further described are methods for harvesting one or more products (e.g., mammalian proteins) from such recombinant cells comprising culturing the cells and harvesting the products.
[0126] In some embodiments, the recombinant cell is a prokaryotic cell, such as a bacterial cell. In some embodiments, the recombinant cell is a eukaryotic cell, such as a mammalian cell, a yeast cell, a filamentous fungal cell, a protist cell, an algae cell, an avian cell, a plant cell, or an insect cell. In some embodiments, the cell is a yeast cell. Those skilled in the art will recognize that many forms of filamentous fungi provide yeast-like growth, and the definition of yeast herein encompasses such cells. The recombinant cell of the present disclosure can be selected from the group consisting of algae, bacteria, mold, fungi, plants, and yeast. In some embodiments, the recombinant cell of the present disclosure is a bacterial cell (e.g., E. coli), a fungal cell, or a yeast cell.
[0127] In some embodiments, the recombinant cell of the present disclosure is a recombinant fungal cell. The recombinant fungal cell can be any suitable fungal cell recognized in the art. In some aspects, the fungal cell is a cell of the genus Arcthura, Aspergillus, Aurantiochytrium, Candida, Cryptococcus, Cryptococcus, Geotrichum, Hansenula, Kluyveromyces, Kodamaea, Komagataella, Leucosporidia, Lipomyces, Mortierella, Ogataea, Pichia, Prototheca, Rhizopus, Rhodosporidium, Rhodotorula, Saccharomyces, Schizosaccharomyces, Tremella, Trichosporon, Wickerhamomyces, or Yarrowia. In some embodiments, the fungal cell is selected from the group consisting of Arxula adeninivorans, Aspergillus niger, Aspergillus orzyae, Aspergillus terreus, Aurantiochytrium limacinum, Candida utilis, Claviceps purpurea, Cryptococcus albidus, Cryptococcus curvatus, Cryptococcus ramirezgomezianus, Cryptococcus terreus, and the like. terreus, Cryptococcus wieringae, Cunninghamella echinulata, Cunninghamella japonica, Geotrichum fermentans, Hansenula polymorpha, Kluyveromyces lactislactis, Komagataella paphii, Komagataella pastoris, Komagataella pseudopastoris, Kluyveromyces marxianus, Kodamaea ohmeri, Leucosporidiella creatinivora, Lipomyces lipofer, Lipomyces starkeyi, Lipomyces tetrasporus, Mortierella isabellina, Mortierella alpina, Ogataea polymorpha, Pichia ciferi ciferrii, Pichia guilliermondii, Pichia pastoris, Pichia stipites, Prototheca zopfii, Rhizopus arrhizus, Rhodosporidium babjevae, Rhodosporidium toruloides, Rhodosporidium paludigenum, Rhodotorula glutinis, Rhodotorula mucilaginosa, Saccharomyces cerevisiae, Schizosaccharomyces cerevisiae, Tremella enchepala, Trichosporon cutaneum cutaneum, Trichosporon fermentans, Wickerhamomyces ciferrii or Yarrowia lipolytica.
[0128]
[0110] In one aspect, the fungal cell is a yeast cell. In one embodiment, the yeast cell is a Komagataella cell. In one embodiment, the yeast cell is Kluyveromyces puffii, Komagataella pastoris, or Komagataella pseudopastoris. In a particular embodiment, the yeast cell is Kluyveromyces puffii.
[0129] In some embodiments, the engineered cells of the present disclosure are yeast cells that contain one or more modifications to improve production of N-glycans, including human-like N-glycans. Examples of such cells and modifications are described, for example, in U.S. Patent No. 9,617,550, which is incorporated herein by reference in its entirety.
[0130] E. Gene Editing Systems Certain aspects of the present disclosure are directed to the use of gene editing techniques to generate knockouts or other mutations in genes in a population of cells.Various methods and systems for gene editing are known in the art, including, for example, zinc finger nuclease (ZFN)-based gene editing, transcription activator-like effector nuclease (TALEN)-based gene editing, and CRISPR / Cas-based gene editing.Various methods and systems for gene editing are recognized in the art and are contemplated herein.In some embodiments, the method of the present disclosure includes CRISPR / Cas-based gene editing, which includes the use of components of the CRISPR system, such as guide RNA (gRNA) and Cas nuclease.In some embodiments, the method of the present disclosure does not include CRISPR / Cas-based gene editing (including, for example, ZFN-based, TALEN-based, or any other gene editing method or system).
[0131]
[0113] Generally, a "CRISPR system" refers collectively to transcripts and other elements involved in the expression of or directing the activity of CRISPR-associated ("Cas") genes, including sequences encoding Cas genes, tracr (trans-activating CRISPR) sequences (e.g., tracrRNA or active partial tracrRNA), tracr-mate sequences (including "direct repeats" and partial direct repeats in the context of an endogenous CRISPR system processed by tracrRNA), guide sequences (also referred to as "spacers" in the context of an endogenous CRISPR system), and / or other sequences and transcripts derived from a CRISPR locus.
[0132]
[0114] A CRISPR / Cas nuclease or CRISPR / Cas nuclease system can include a non-coding RNA molecule (guide) RNA that binds to DNA in a sequence-specific manner, and a Cas protein (e.g., Cas9) that has nuclease functionality (e.g., two nuclease domains). One or more elements of the CRISPR system can be derived from a type I, type II, or type III CRISPR system, and can be derived from a particular organism that contains an endogenous CRISPR system, such as Streptococcus pyogenes.
[0133] In one aspect, a Cas nuclease and gRNA (comprising a fusion of a target sequence-specific crRNA and a fixed tracrRNA) are introduced into a cell. The Cas nuclease and gRNA can be introduced into a cell indirectly through the introduction of one or more nucleic acids (e.g., vectors) encoding the Cas nuclease and / or gRNA. The Cas nuclease and gRNA can be introduced into a cell directly by the introduction of a Cas nuclease protein and a gRNA molecule. Generally, a target site at the 5' end of the gRNA targets the Cas nuclease to the target site, e.g., a gene, using complementary base pairing. The target site can be selected based on its immediate 5' position of a protospacer adjacent motif (PAM) sequence, e.g., typically NGG or NAG. In this respect, gRNA can be targeted to desired sequence by modifying the first 20, 19, 18, 17, 16, 15, 14, 14, 12, 11 or 10 nucleotides of guide RNA to correspond to target DNA sequence.Generally, CRISPR system is characterized by the element that promotes the formation of CRISPR complex at the site of target sequence.Typically, " target sequence " generally refers to the sequence that guide sequence is designed to have complementarity with, where hybridization between target sequence and guide sequence promotes the formation of CRISPR complex.Perfect complementarity is not necessarily required, provided that there is sufficient complementarity to cause hybridization and promote the formation of CRISPR complex.
[0134] CRISPR systems can induce double-strand breaks (DSBs) at target sites, followed by disruption as discussed herein. In other embodiments, Cas9 variants considered "nickases" are used to nick a single strand at the target site. Paired nickases can be used to improve the specificity mediated by pairs of distinct gRNA targeting sequences, for example, so that a 5' overhang is simultaneously introduced upon nick introduction. In other embodiments, catalytically inactive Cas9 is fused to a heterologous effector domain, such as a transcriptional repressor or activator, to affect gene expression.
[0135]
[0117] The target sequence can comprise any polynucleotide, such as a DNA or RNA polynucleotide. The target sequence can be located in the nucleus or cytoplasm of a cell, such as within a cellular organelle. Generally, a sequence or template that can be used for recombination into a target locus that contains a target sequence is referred to as an "editing template" or "editing polynucleotide" or "editing sequence." In one aspect, an exogenous template polynucleotide can be referred to as an editing template. In one aspect, the recombination is homologous recombination.
[0136] Typically, in the context of an endogenous CRISPR system, formation of a CRISPR complex (comprising a guide sequence hybridized to a target sequence and complexed with one or more Cas proteins) results in cleavage of one or both strands in or near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs from the target sequence). A tracr sequence that may comprise or consist of all or a portion of a wild-type tracr sequence (e.g., about 20, 26, 32, 45, 48, 54, 63, 67, 85, or more nucleotides of the wild-type tracr sequence) can also form part of a CRISPR complex, for example, by hybridization along at least a portion of the tracr sequence to all or a portion of a tracr mate sequence operably linked to the guide sequence. The tracr sequence has sufficient complementarity to the tracr mate sequence to hybridize and participate in the formation of a CRISPR complex, such as at least 50%, 60%, 70%, 80%, 90%, 95% or 99% sequence complementarity along the length of the tracr mate sequence when optimally aligned.
[0137] One or more vectors driving the expression of one or more elements of the CRISPR system can be introduced into a cell so that expression of the CRISPR system elements directs the formation of a CRISPR complex at one or more target sites. Components can also be delivered to a cell as protein and / or RNA. For example, a Cas enzyme, a guide sequence linked to a tracr-mate sequence, and a tracr sequence can each be operably linked to separate regulatory elements on separate vectors. Alternatively, two or more elements expressed from the same or different regulatory elements can be combined in a single vector, with one or more additional vectors providing any components of the CRISPR system not included in the first vector. A vector can contain one or more insertion sites, such as restriction endonuclease recognition sequences (also called "cloning sites"). In some embodiments, one or more insertion sites are located upstream and / or downstream of one or more sequence elements of one or more vectors. When multiple different guide sequences are used, a single expression construct can be used to target CRISPR activity to multiple different corresponding target sequences within a cell.
[0138] The vector may include regulatory elements operably linked to an enzyme coding sequence encoding a Cas protein (also called a "Cas nuclease"). Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Cas12a (Cpf1), Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, These enzymes include Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csfl, Csf2, Csf3, Csf4, their homologs, or modified versions thereof. These enzymes are known; for example, the amino acid sequence of the Streptococcus pyogenes Cas9 protein can be found in the SwissProt database under accession number Q99ZW2.
[0139] The Cas nuclease can be Cas9 (e.g., from Streptococcus pyogenes or Streptococcus pneumoniae). The Cas nuclease can be Cas12a. The Cas nuclease can direct cleavage of one or both strands at the location of a target sequence, such as within the target sequence and / or within the complementary sequence of the target sequence. The vector can encode a Cas nuclease that is mutated relative to the corresponding wild-type enzyme, such that the mutated Cas nuclease lacks the ability to cleave one or both strands of a target polynucleotide containing the target sequence. In some embodiments, the Cas9 nickase can be used in combination with guide sequence(s), for example, two guide sequences that target the sense and antisense strands, respectively, of a DNA target. This combination allows both strands to be nicked and used to induce NHEJ or HDR.
[0140]
[0122] In some embodiments, the enzyme coding sequence encoding the CRISPR enzyme is codon-optimized for expression in a particular cell, such as a yeast cell.
[0141] Generally, guide sequence is any polynucleotide sequence that has sufficient complementarity with target polynucleotide sequence to hybridize with target sequence and direct the sequence-specific binding of CRISPR complex to target sequence.In some embodiments, the degree of complementarity between guide sequence and its corresponding target sequence is 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or more when optimally aligned using suitable alignment algorithm.
[0142] Optimal alignment can be determined using any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler Transform (e.g., Burrows Wheeler Aligner), Clustal W, Clustal X, BLAST, Novoalign (Novocraft Technologies, ELAND (Illumina, San Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net).
[0143] Cas nucleases can be part of fusion proteins containing one or more heterologous protein domains. Cas nuclease fusion proteins can include any additional protein sequences and, optionally, linker sequences between any two domains. Examples of protein domains that can be fused to Cas nucleases include, but are not limited to, epitope tags, reporter gene sequences, and protein domains with one or more of the following activities: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, and nucleic acid binding activity. Non-limiting examples of epitope tags include histidine (His) tags, V5 tags, FLAG tags, influenza hemagglutinin (HA) tags, Myc tags, VSV-G tags, and thioredoxin (Trx) tags. Examples of reporter genes include, but are not limited to, glutathione-5-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, luciferase, green fluorescent protein (GFP), autofluorescent proteins including HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and blue fluorescent protein (BFP). Cas nucleases can be fused to gene sequences encoding proteins or protein fragments that bind to DNA molecules or other cellular molecules, including, but not limited to, maltose binding protein (MBP), S-tag, Lex A DNA binding domain (DBD) fusions, GAL4A DNA binding domain fusions, and herpes simplex virus (HSV) BP16 protein fusions. Additional domains that can form part of fusion proteins containing Cas nucleases are described in U.S. Patent No. 20110059502, incorporated herein by reference. [Example]
[0144]
[0126] The following examples are included to demonstrate specific embodiments disclosed herein. It should be understood by those skilled in the art that the techniques disclosed in the following examples represent techniques discovered by the inventors to function well in implementing the disclosed embodiments, and therefore can be considered to constitute specific modes for their implementation. However, those skilled in the art should understand, in light of the present disclosure, that many changes can be made in the specific embodiments disclosed and still obtain the same or similar results, without departing from the spirit and scope of the embodiments disclosed herein.
[0145] Example 1 - Novel signal peptides increase extracellular protein levels To determine the effect of the novel signal peptides on extracellular protein levels, DNA encoding SEQ ID NO:1 ("SP1"), SEQ ID NO:2 ("SP2"), and SEQ ID NO:4 ("SP4") was cloned in frame at the 5' end of DNA encoding the protein of interest (POI), i.e., Pichia pastoris codon-optimized human lactoferrin, resulting in the replacement of pre-pro-MFα from Saccharomyces cerevisiae. This is the most widely used signal peptide in yeast and serves as a control. Single copies of the resulting sequences and controls were integrated into the AOX1 locus via double crossover. Multiple colonies from each transformation plate were grown in 96-deep-well plates.
[0146] To confirm the presence of the protein of interest, Western blots were performed on the supernatant. As shown in Figure 1, when a single copy of human lactoferrin is integrated and secretion is driven by the widely used preproMFα from Saccharomyces cerevisiae, no protein is detected in the supernatant. Conversely, when secretion is driven by SEQ ID NO:1 ("SP1"), SEQ ID NO:2 ("SP2"), and SEQ ID NO:3 ("SP3"), extracellular protein is detected.
[0147] To assess the magnitude of the secretion enhancement, extracellular protein quantification was performed by ELISA. As can be seen in Figure 2, the novel engineered signals enhanced extracellular protein levels by 2.38-fold, 2.41-fold, and 2.20-fold over the control (preproMFα) for SEQ ID NO:1 ("SP1"), SEQ ID NO:2 ("SP2"), and SEQ ID NO:3 ("SP3"), respectively.
[0148] Materials and Methods Vector and strain construction. Oligonucleotides and gBlocks were ordered from Integrated DNA Technologies (San Diego, CA, USA) and are listed in Table 5. NEBuilder® HiFi DNA Assembly Master Mix, OneTaq® Quickload® DNA polymerase, and E. coli DH5α cells were from New England Biolabs. All polymerase chain reaction (PCR) amplified sequences were confirmed by sequencing at Genewiz.
[0149] [Table 5] Transformation of linear dsDNA for integration was performed using the method described by Madden, Tolstorukov, & Cregg (2014) Fungi, Volume 1, Fungal Biology. Total yeast genomic DNA extraction was performed using the Easy DNA kit from Invitrogen (ThermoFisher, Applied Biosystems™, PrepSEQ™ 1-2-3 Nucleic Acid Extraction Kit, Catalog number: 4452222). The resulting plasmids are summarized in Table 6.
[0150] [Table 6] The leader peptide sequences from the Pichia pastoris endogenous proteins Ost1 and Pst1 were determined using publicly available SignalP-5.0 bioinformatics software from the Center for Biological Sequence Analysis (CBS). The pro-region of Epx1 was described by Heiss et al. (2015) Microbiology, 161(7).
[0151] Plasmid P1 containing a gene encoding human lactoferrin lacking its native secretory peptide fused in-frame with the prepro-leader peptide of mating factor alpha from Saccharomyces cerevisiae was synthesized by Genscript. The human lactoferrin gene was codon-optimized for expression in Pichia pastoris.
[0152] To generate plasmid P2 containing the signal sequence SP1 (SEQ ID NO:1), primers PMR1 (SEQ ID NO:16) and PMR2 (SEQ ID NO:17) were used to amplify the Ost1 leader sequence using gBLOCK1 as a template. A backbone containing human lactoferrin, a yeast HIS4 auxotrophic marker, and an E. coli antibiotic resistance and replication origin was obtained by polymerase chain reaction (PCR) of the P1 plasmid using primers PMR3 (SEQ ID NO:18) and PMR4 (SEQ ID NO:19). The two resulting fragments were assembled using NEBuilder® HiFi DNA Assembly Master Mix according to the manufacturer's instructions.
[0153] To generate plasmid P3 containing the signal sequence SP2 (SEQ ID NO:2), primers PMR5 (SEQ ID NO:20) and PMR6 (SEQ ID NO:21) were used for amplification, and gBLOCK1 (SEQ ID NO:15) was used as a template. A backbone containing human lactoferrin, the yeast HIS4 auxotrophic marker, and the E. coli antibiotic resistance and replication origin was obtained by PCR of the P1 plasmid using primers PMR7 (SEQ ID NO:22) and PMR8 (SEQ ID NO:23). The two resulting fragments were assembled using NEBuilder® HiFi DNA Assembly Master Mix according to the manufacturer's instructions.
[0154] To generate plasmid P4, which contains the signal sequence SP3 (SEQ ID NO:3), primers PMR9 (SEQ ID NO:24) and PMR10 (SEQ ID NO:25) were used for amplification using gBLOCK1 as a template. A backbone containing human lactoferrin, the yeast HIS4 auxotrophic marker, and E. coli antibiotic resistance and replication origin was obtained by PCR of the P1 plasmid using primers PMR11 (SEQ ID NO:26) and PMR12 (SEQ ID NO:27). The two resulting fragments were assembled using NEBuilder® HiFi DNA Assembly Master Mix according to the manufacturer's instructions.
[0155] To generate plasmid P5, which contains the signal sequence SP4 (SEQ ID NO:4), primers PMR13 (SEQ ID NO:28) and PMR14 (SEQ ID NO:29) were used for amplification using gBLOCK1 (SEQ ID NO:15) as a template. A backbone containing human lactoferrin, the yeast HIS4 auxotrophic marker, and E. coli antibiotic resistance and replication origin was obtained by PCR of the P1 plasmid using primers PMR15 (SEQ ID NO:30) and PMR16 (SEQ ID NO:31). The resulting two fragments were assembled using NEBuilder® HiFi DNA Assembly Master Mix according to the manufacturer's instructions.
[0156] The assembly mixture was transformed into E. coli DH5α cells as instructed by the manufacturer and plated on Luria Broth (LB) agar plates containing 100 μg / mL ampicillin. Positive clones were selected by colony polymerase chain reaction (PCR) and inoculated overnight into 5 mL of liquid Luria Broth medium supplemented with 100 μg / mL ampicillin. Plasmids from E. coli cells were isolated using a GeneJET Plasmid Miniprep Kit (ThermoFisher®, Cat. No. K502). Proper assembly was confirmed by Sanger DNA sequencing.
[0157] Primers PMR17 (SEQ ID NO:32) and PMR18 (SEQ ID NO:33) and plasmids P1, P2, P3, P4, or P5 were used as templates to generate linear dsDNA fragments for integration into yeast using Q5 high-fidelity DNA polymerase. Electrocompetent Pichia pastoris cells were transformed as described by Madden, Tolstorukov, & Cregg (2014) Fungi, Volume 1, Fungal Biology. Cells were plated on MD plates (1.34% yeast nitrogen base, 4 × 10 -5% biotin, 2% glucose, 20% agar), which + Cells were allowed to select and incubated at 30°C for 72 hours. Individual yeast colonies (~10-20) were then restreaked on MD plates and grown at 30°C for 24 hours. Cells transformed with P1 were used as a control to evaluate the higher efficiency of SP1 (SEQ ID NO:1), SP2 (SEQ ID NO:2), SP3 (SEQ ID NO:3), and SP4 (SEQ ID NO:5) in secreting the protein of interest (POI).
[0158] Individual colonies from the restreaked plates were inoculated into 96-deep-well plates using 600 μl of 2% YPD (2% dextrose, 2% peptone, 1% yeast extract). Cells were grown at 1,000 rpm and 30°C for 48 hours. Fifty microliters of the resulting cell suspension was added to 550 μl of BMG (100 mM potassium phosphate buffer (pH = 6.0), 1.34% yeast nitrogen base, 4 × 10) supplemented with 0.5% cas-amino acids. -5 % biotin, 1% glycerol) and incubated at 1,000 rpm and 30°C for 48 hours. The cells were then pelleted by centrifugation at 4,500 × g for 5 minutes and transferred to 1% BMM (100 mM potassium phosphate buffer (pH = 6.0), 1.34% yeast nitrogen base, 4 × 10 -5 The cells were resuspended in 1% PBS (1% biotin, 1% methanol). Proteins secreted into the extracellular medium were then analyzed by SDS-PAGE, ELISA, and Western blot.
[0159]
[0142] All of the methods disclosed and claimed herein can be made and executed without undue experimentation in light of the present disclosure. While the compositions and methods disclosed herein have been described with reference to specific embodiments, it will be apparent to those skilled in the art that variations can be applied to the methods described herein, and in the steps or in the order of the steps of the methods, without departing from the concept, spirit, and scope of the disclosed embodiments. More specifically, it will be understood that certain chemically and physiologically related agents can be substituted for the agents described herein while the same or similar results can be achieved. All such similar substitutes and modifications apparent to those skilled in the art are deemed to be within the spirit, scope, and concept of the embodiments disclosed herein, as defined by the appended claims.
[0160] This specification includes the disclosure of the following inventions: [Item 1] An isolated nucleic acid encoding a polypeptide comprising a sequence having at least 90% sequence identity to SEQ ID NO: 1, 2, 3 or 4. [Item 2] An isolated nucleic acid according to Item 1, wherein the sequence comprises SEQ ID NO: 1, 2, 3, or 4. [Item 3] An isolated nucleic acid according to item 1 or 2, wherein the polypeptide further comprises a sequence of a mammalian protein. [Item 4] The isolated nucleic acid according to Item 3, wherein the mammalian protein is a human milk protein. [Item 5] The isolated nucleic acid according to Item 4, wherein the human milk protein is secretory IgA (sIgA), xanthine dehydrogenase, lactoferrin, lactoperoxidase, butyrophilin, lactadherin, adiponectin, β-casein, κ-casein, leptin, lysozyme, or α-lactalbumin. [Item 6] The isolated nucleic acid according to Item 5, wherein the human milk protein is human lactoferrin. [Item 7] An isolated nucleic acid according to any one of Items 1 to 6, wherein the sequence has at least 90% sequence identity to SEQ ID NO:1. [Item 8] An isolated nucleic acid according to any one of Items 1 to 6, wherein the sequence comprises SEQ ID NO:1. [Item 9] An isolated nucleic acid according to Item 8, wherein the isolated nucleic acid comprises a nucleic acid sequence having at least 80% identity to SEQ ID NO:41. [Item 10] An isolated nucleic acid according to Item 9, wherein the nucleic acid sequence comprises SEQ ID NO:41. [Item 11] An isolated nucleic acid according to Item 8, wherein the polypeptide comprises SEQ ID NO:5. [Item 12] An isolated nucleic acid according to Item 11, wherein the isolated nucleic acid comprises a nucleic acid sequence having at least 80% identity to SEQ ID NO:46. [Item 13] An isolated nucleic acid according to Item 12, wherein the nucleic acid sequence comprises SEQ ID NO:46. [Item 14] An isolated nucleic acid according to any one of Items 1 to 6, wherein the sequence has at least 90% sequence identity to SEQ ID NO:2. [Item 15] An isolated nucleic acid according to any one of Items 1 to 6, wherein the sequence comprises SEQ ID NO:2. [Item 16] An isolated nucleic acid according to Item 15, wherein the isolated nucleic acid comprises a nucleic acid sequence having at least 80% identity to SEQ ID NO:42. [Item 17] An isolated nucleic acid according to Item 16, wherein the nucleic acid sequence comprises SEQ ID NO:42. [Item 18] An isolated nucleic acid according to Item 15, wherein the polypeptide comprises SEQ ID NO:6. [Item 19] An isolated nucleic acid according to Item 18, wherein the isolated nucleic acid comprises a nucleic acid sequence having at least 80% identity to SEQ ID NO:47. [Item 20] An isolated nucleic acid according to Item 19, wherein the nucleic acid sequence comprises SEQ ID NO:47. [Item 21] An isolated nucleic acid according to any one of Items 1 to 6, wherein the sequence has at least 90% sequence identity to SEQ ID NO:3. [Item 22] An isolated nucleic acid according to any one of Items 1 to 6, wherein the sequence comprises SEQ ID NO:3. [Item 23] An isolated nucleic acid according to Item 22, wherein the isolated nucleic acid comprises a nucleic acid sequence having at least 80% identity to SEQ ID NO:43. [Item 24] An isolated nucleic acid according to Item 23, wherein the nucleic acid sequence comprises SEQ ID NO:43. [Item 25] An isolated nucleic acid according to Item 22, wherein the polypeptide comprises SEQ ID NO:7. [Item 26] An isolated nucleic acid according to Item 25, wherein the isolated nucleic acid comprises a nucleic acid sequence having at least 80% identity to SEQ ID NO:48. [Item 27] An isolated nucleic acid according to Item 26, wherein the nucleic acid sequence comprises SEQ ID NO:48. [Item 28] An isolated nucleic acid according to any one of Items 1 to 6, wherein the sequence has at least 90% sequence identity to SEQ ID NO:4. [Item 29] An isolated nucleic acid according to any one of Items 1 to 6, wherein the sequence comprises SEQ ID NO:4. [Item 30] An isolated nucleic acid according to Item 29, wherein the isolated nucleic acid comprises a nucleic acid sequence having at least 80% identity to SEQ ID NO:44. [Item 31] An isolated nucleic acid according to Item 30, wherein the nucleic acid sequence comprises SEQ ID NO:44. [Item 32] An isolated nucleic acid according to Item 32, wherein the polypeptide comprises SEQ ID NO:8. [Item 33] An isolated nucleic acid according to Item 32, wherein the isolated nucleic acid comprises a nucleic acid sequence having at least 80% identity to SEQ ID NO:49. [Item 34] An isolated nucleic acid according to Item 33, wherein the nucleic acid sequence comprises SEQ ID NO:49. [Item 35] A vector comprising the nucleic acid according to any one of Items 1 to 34. [Item 36] An engineered eukaryotic cell comprising the nucleic acid according to any one of Items 1 to 34 or the vector according to Item 35. [Item 37] The engineered eukaryotic cell according to Item 36, wherein the cell is a fungal cell. [Item 38] The engineered eukaryotic cell according to Item 37, wherein the fungal cell is a cell of the genus Arcthura, Aspergillus, Aurantiochytrium, Candida, Cryptococcus, Cryptococcus, Geotrichum, Hansenula, Kluyveromyces, Kodamaea, Komagataella, Leucosporidia, Lipomyces, Mortierella, Ogataea, Pichia, Prototheca, Rhizopus, Rhodosporidium, Rhodotorula, Saccharomyces, Schizosaccharomyces, Tremella, Trichosporon, Wickerhamomyces, or Yarrowia. [Item 39] The engineered eukaryotic cell according to Item 38, wherein the cell is a yeast cell. [Item 40] The engineered eukaryotic cell according to Item 39, wherein the yeast cell is a cell of the genus Komagataella. [Item 41] The engineered eukaryotic cell according to Item 40, wherein the yeast cell is a Komagataella puffii, Komagataella pastoris or Komagataella pseudopastoris cell. [Item 42] An engineered eukaryotic cell according to any one of Items 36 to 41, wherein the nucleic acid is integrated into the genome of the cell. [Item 43] An engineered eukaryotic cell according to any one of Items 36 to 41, wherein the nucleic acid is not integrated into the genome of the cell. [Item 44] A method for producing a secreted protein, the method comprising growing a cell according to any one of Items 36 to 43 under conditions sufficient to secrete the polypeptide from the cell. [Item 45] The method described in Item 44, further comprising collecting the secreted protein. [Item 46] The method described in Item 44 or 45, wherein the secreted protein is a human milk protein. [Item 47] The method described in Item 46, wherein the human milk protein is secretory IgA (sIgA), xanthine dehydrogenase, lactoferrin, lactoperoxidase, butyrophilin, lactadherin, adiponectin, β-casein, κ-casein, leptin, lysozyme, or α-lactalbumin. [Item 48] The method according to any one of Items 44 to 47, wherein the human milk protein comprises one or more human-like N-glycans. [Item 49] The method according to any one of Items 44 to 48, further comprising producing a mixture comprising the human milk protein and one or more components of infant formula. [Item 50] An engineered yeast cell comprising a nucleic acid encoding a polypeptide comprising a sequence having at least 90% sequence identity to SEQ ID NO: 1, 2, 3 or 4. [Item 51] An engineered yeast cell according to Item 50, wherein the sequence comprises SEQ ID NO: 1, 2, 3 or 4. [Item 52] An engineered yeast cell according to Item 51, wherein the sequence comprises SEQ ID NO:3. [Item 53] The engineered yeast cell according to any one of Items 50 to 52, wherein the polypeptide further comprises a sequence of a mammalian protein. [Item 54] The engineered yeast cell according to Item 53, wherein the mammalian protein is a human milk protein. [Item 55] The engineered yeast cell according to Item 54, wherein the human milk protein is secretory IgA (sIgA), xanthine dehydrogenase, lactoferrin, lactoperoxidase, butyrophilin, lactadherin, adiponectin, β-casein, κ-casein, leptin, lysozyme, or α-lactalbumin. [Item 56] The engineered yeast cell according to Item 55, wherein the human milk protein is human lactoferrin. [Item 57] An engineered yeast cell comprising: (a) a first nucleic acid encoding a polypeptide comprising: (i) a sequence having at least 90% sequence identity to SEQ ID NO: 1, 2, 3 or 4; and (ii) the sequence of a human milk protein; and (b) an engineered yeast cell comprising a second nucleic acid encoding an α-1,2-mannosidase (Man-I) protein, wherein the cell does not express a functional OCH1 protein. [Item 58] An engineered yeast cell according to Item 57, wherein the sequence of (i) comprises SEQ ID NO: 1, 2, 3 or 4. [Item 59] The engineered yeast cell according to Item 57 or 58, wherein the human milk protein is secretory IgA (sIgA), xanthine dehydrogenase, lactoferrin, lactoperoxidase, butyrophilin, lactadherin, adiponectin, β-casein, κ-casein, leptin, lysozyme, or α-lactalbumin. [Item 60] The engineered yeast cell according to Item 57, wherein the human milk protein is human lactoferrin. [Item 61] The engineered yeast cell according to Item 57, wherein the human milk protein is human alpha-lactalbumin. [Item 62] The engineered yeast cell according to any one of Items 57 to 61, wherein the Man-I protein is fused to an HDEL C-terminal tag. [Item 63] The engineered yeast cell according to any one of Items 57 to 62, further comprising: (a) N-acetylglucosaminyltransferase I (GnT-I) protein; (b) α-1,3 / 6-mannosidase (Man-II) protein; (c) β-1,2-acetylglucosaminyltransferase (GnT-II) protein; and (d) β-1,4-galactosyltransferase (GalT) protein; The engineered yeast cell further comprising a third nucleic acid encoding one or more of: [Item 64] Infant formula containing human glycoproteins with human-like N-linked glycosylation. [Item 65] The infant formula according to Item 64, wherein the human glycoprotein is a human milk protein. [Item 66] The infant formula according to Item 65, wherein the human milk protein is secretory IgA ( sIgA), xanthine dehydrogenase, lactoferrin, lactoperoxidase, Infant formula that is butyrophilin, lactadherin, adiponectin, beta-casein, kappa-casein, leptin, lysozyme, or alpha-lactalbumin. [Item 67] The infant formula according to Item 66, wherein the human milk protein is human lactoferrin. [Item 68] An infant formula according to any one of Items 64 to 67, wherein the lactoferrin has a glycan pattern that is different from the glycan pattern of any human lactoferrin naturally present in human breast milk. [Item 69] The infant formula according to any one of Items 64 to 68, wherein the human glycoprotein is produced by the method according to any one of Items 44 to 49. References The following references, to the extent that they provide exemplary procedural or other details supplementary to those set forth herein, are specifically incorporated herein by reference.
[0161] [ka] U.S. Patent No. 4,977,137 (Nichols et al.) U.S. Patent No. 5,571,691 (Conneely et al.) U.S. Patent No. 7,335,512 (Callewaert et al.) U.S. Patent No. 7,344,867 (Connolly) U.S. Patent No. 7,749,960 (Vidal et al.) U.S. Patent No. 7,524,815 (Vidal et al.) U.S. Patent No. 7,914,822 (Medo) U.S. Patent No. 8,440,456 (Callewaert et al.) U.S. Patent No. 8,871,445 (Cong et al.) U.S. Patent No. 8,802,650 (Buck et al.) U.S. Patent No. 8,821,878 (Medo et al.) U.S. Patent No. 8,927,027 (Fournell et al.) U.S. Patent No. 7,449,308 (Gerngross et al.) U.S. Patent Application Publication No. 2012 / 0142580 (Nutten et al.)
Claims
1. An isolated nucleic acid encoding a polypeptide comprising a signal peptide comprising a sequence having at least 90% sequence identity to the sequence of SEQ ID NO: 1, 2, 3 or 4.
2. 2. The isolated nucleic acid of claim 1, wherein the polypeptide comprises a signal peptide comprising the sequence of SEQ ID NO: 1, 2, 3, or 4.
3. 2. The isolated nucleic acid of claim 1, wherein the polypeptide further comprises a sequence of a mammalian protein.
4. 4. The isolated nucleic acid of claim 3, wherein the mammalian protein is a human milk protein.
5. 5. The isolated nucleic acid of claim 4, wherein the human milk protein is secretory IgA (sIgA), xanthine dehydrogenase, lactoferrin, lactoperoxidase, butyrophilin, lactadherin, adiponectin, β-casein, κ-casein, leptin, lysozyme, or α-lactalbumin.
6. 6. The isolated nucleic acid according to any one of claims 1 to 5, wherein the polypeptide comprises a signal peptide comprising the sequence of SEQ ID NO:
3.
7. 2. The isolated nucleic acid of claim 1, wherein the isolated nucleic acid comprises a nucleic acid sequence having at least 80% identity to the sequence of SEQ ID NO: 41, 42, 43 or 44.
8. 2. The isolated nucleic acid of claim 1, wherein the nucleic acid sequence comprises the sequence of SEQ ID NO: 41, 42, 43 or 44.
9. 2. The isolated nucleic acid of claim 1, wherein the polypeptide comprises the sequence of SEQ ID NO: 5, 6, 7 or 8.
10. 2. The isolated nucleic acid of claim 1, wherein the isolated nucleic acid comprises a nucleic acid sequence having at least 80% identity to the sequence of SEQ ID NO: 46, 47, 48 or 49.
11. 2. The isolated nucleic acid of claim 1, wherein the nucleic acid sequence comprises the sequence of SEQ ID NO: 46, 47, 48 or 49.
12. A vector comprising the nucleic acid according to any one of claims 1 to 5.
13. An engineered Pichia pastoris cell comprising a nucleic acid according to any one of claims 1 to 5.
14. 14. A method for producing a secreted polypeptide, comprising growing the Pichia pastoris cell of claim 13 under conditions sufficient to secrete the polypeptide from the cell.
15. 15. The method of claim 14, wherein the secreted polypeptide is a human milk protein.
16. 16. The method of claim 15, wherein the human milk protein comprises one or more human-like N-glycans.
17. 15. The method of claim 14, further comprising collecting the secreted polypeptide, wherein the secreted polypeptide is used in the manufacture of an infant formula.
18. 14. The engineered Pichia pastoris cell of claim 13, wherein the nucleic acid encodes a mammalian polypeptide.
19. 14. The engineered Pichia pastoris cell of claim 13, wherein the sequence comprises the sequence of SEQ ID NO: 1, 2, 3 or 4.
20. 19. The engineered Pichia pastoris cell of claim 18, wherein the mammalian polypeptide is a human protein.
21. 21. The engineered Pichia pastoris cell of claim 20, wherein the human protein is secretory IgA (sIgA), xanthine dehydrogenase, lactoferrin, lactoperoxidase, butyrophilin, lactadherin, adiponectin, β-casein, κ-casein, leptin, lysozyme, or α-lactalbumin.
22. 14. The engineered Pichia pastoris cell of claim 13, which is Pichia pastoris, Komagataella pastoris, or Komagataella puffii.
23. 15. The method of claim 14, wherein the Pichia pastoris cell is Pichia pastoris, Komagataella pastoris, or Komagataella puffii.
Citation Information
Patent Citations
Bacillus circulans chitoanase as well as preparation method and application thereof
CN107586768A
Expression sequence
JP2015533286A
Compositions and methods for producing high secreted yields of recombinant proteins
JP2020509751A
Nucleic acids of pichia pastoris and use thereof for recombinant production of proteins
US20110021378A1
Signal sequence for protein expression in pichia pastoris
US20160168198A1