Compositions and methods for producing heterologous globins in filamentous fungal cells

Recombinant filamentous fungal strains with optimized expression cassettes and fermentation processes address the need for cost-effective large-scale globin protein production, achieving high yields for plant-based meat substitutes.

JP2026505371APending Publication Date: 2026-02-13DANISCO US INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025545990
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-02-08
Filing Date
2024-01-26
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

There is an ongoing need for enhanced globin protein expression systems suitable for cost-effective large-scale production of globin proteins to meet the increasing demand for plant-based meat substitutes.

Method used

The development of recombinant filamentous fungal strains with enhanced globin protein productivity phenotypes, utilizing expression cassettes encoding globin proteins, and industrial-scale fermentation processes to produce heterologous globin proteins, such as leghemoglobin, myoglobin, and hemoglobin, in fermentation broths.

Benefits of technology

The method achieves high yields of globin proteins, with strains like T. reesei producing up to 1 gram per liter of fermentation broth, suitable for use in food materials and flavor/aroma modifiers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026505371000001_ABST
    Figure 2026505371000001_ABST
Patent Text Reader

Abstract

The present disclosure generally relates to methods and compositions for producing heterologous globin proteins of interest in recombinant filamentous fungal cells. Accordingly, certain embodiments are directed to compositions and methods for the production of globin proteins, recombinant filamentous fungal strains comprising an enhanced globin protein productivity phenotype, polynucleotides (e.g., expression constructs) encoding one or more globin proteins, industrial-scale fermentation, globin protein recovery processes, and the like.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates generally to the fields of biology, genetics, molecular biology, filamentous fungi, food proteins, industrial protein production, and the like. Certain embodiments relate to methods and compositions for producing heterologous globin proteins in filamentous fungal strains. As described herein, the recombinant fungal strains of the present disclosure are particularly well suited for growth in submerged culture for large-scale production of heterologous globin proteins.

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 483,859, filed February 8, 2024, which is incorporated herein by reference in its entirety.

[0003] Sequence Listing Reference The contents of the electronic submission of the Sequence Listing text file named "NB42149-WO-PCT_SequenceListing.xml" was created on January 23, 2024, is 227KB in size, and is incorporated herein by reference in its entirety. [Background technology]

[0004] As will be appreciated by those skilled in the art, over the past 10 to 15 years, there has been an ongoing shift to develop and produce plant-based meat substitutes with taste and / or aroma profiles similar to those of animal-based meat. In particular, as described in PCT Publication WO 2013 / 010042, livestock farming has a significant negative impact on the environment, with an estimated 30% of the Earth's land surface devoted to animal farming and livestock accounting for over 20% of all terrestrial animal biomass. As further described in WO 2013 / 010042, due to large-scale livestock farming, such practices account for over 18% of net greenhouse gas emissions, suggesting that livestock farming may be the largest source of water pollution and the world's greatest threat to biodiversity. For example, it has been estimated that if the human population could shift from a meat-containing diet to a diet free of animal products (vegetarian diet), over 26% of the Earth's land surface would be freed up for other uses, significantly reducing water and energy consumption.

[0005] In certain aspects, the WO 2013 / 010042 publication speculates that one or more plant-based (meat) proteins may be isolated and purified from genetically modified organisms (e.g., genetically modified bacteria or yeast cells), where the one or more isolated and purified plant proteins include hemoglobin, myoglobin, leghemoglobin, non-symbiotic hemoglobin, etc. For example, soybean-derived leghemoglobin protein is an important food additive that imparts meat-like flavor and color to meat analogs. In particular, the WO 2013 / 010042 specification and experimental examples section describe the construction of muscle replicas (compositions), muscle tissue analogs, adipose tissue analogs, and connective tissue analogs, each constructed from one or more plant proteins (hemoglobin, myoglobin, leghemoglobin, etc.) isolated and purified from one or more natural plant sources, i.e., as opposed to being expressed and recovered from genetically modified organisms.

[0006] PCT Publication No. WO 2014 / 110532 describes methods and compositions for modifying the flavor and aroma profile of consumable foods using so-called "plant-based meat substitutes" that have properties similar to animal-based meat compositions, where the plant-based meat substitutes contain one or more highly conjugated heterocycles (i.e., heme prosthetic groups) complexed with one or more flavor precursors (e.g., sugars, oils, FFAs, amino acids, nucleosides, vitamins, etc.) and iron complexes. U.S. Patent Publication No. U.S. Patent Application Publication No. 2014 / 0161958 describes a meat substitute product comprising vegetable protein blended with starch, hydrocolloids, and oil from a plant source. U.S. Patent Publication No. U.S. Patent Application Publication No. 2021 / 0289813 describes a meat substitute that includes two or more plant protein sources, or a meat substitute that includes one or more plant protein sources and fruit, fruit powder, or chia seed extract, or a hypoallergenic meat substitute that is optionally soy-free and optionally free of other allergenic ingredients.

[0007] In other aspects, attempts have been described to produce certain globin proteins (e.g., leghemoglobin, cyanoglobin) in recombinant microbial host cells (e.g., E. coli, cyanobacteria, yeast). In one particular aspect, PCT Publication No. WO 2016 / 183163 generally describes methods for constructing engineered methylotrophic yeast cells (P. patoris) for expression of recombinant proteins, where the engineered yeast cells co-express the entire heme biosynthetic pathway from a methanol-inducible promoter. For example, WO 2016 / 183163 teaches the use of P. patoris strains that overexpress the transcriptional activator Mxr1 under the control of the alcohol oxidase 1 (AOX1) promoter element to increase expression / co-expression of recombinant proteins and the heme biosynthetic pathway.

[0008] However, as reviewed in Krainer et al. (2015), insufficient heme incorporation is considered a central obstacle in the recombinant production of active hemoproteins. In particular, Krainer et al. (2015) investigated the effects of both pathway engineering and medium supplementation to optimize the recombinant production of the hemoprotein horseradish peroxidase (HRP) in the yeast P. patoris. As summarized by Krainer et al. (2015), in contrast to other studies, (a) co-overexpression of genes of the endogenous heme biosynthetic pathway in P. patoris did not improve recombinant production of active heme (HRP) enzyme, (b) medium supplementation with the commonly used precursor 5-aminolevulinic acid (ALA) did not affect the yield of active heme (HRP) enzyme, but (c) medium supplementation with hemin increased the yield of active heme (HRP) enzyme, leading to the conclusion that the yield of active peroxidase enzyme from P. patoris can be easily enhanced by supplementing the culture medium with hemin.

[0009] PCT Publication No. WO 2019 / 079135 generally describes non-animal derived meat-like materials / ingredients obtained from genetically modified cyanobacteria containing polynucleotides encoding heterologous globin proteins (e.g., leghemoglobin, cyanoglobin). PCT Publication No. WO 2023 / 278968 describes non-heme iron-binding protein pigment compositions for meat substitutes that impart a pink and / or red color to the meat substitute composition. Summary of the Invention [Problem to be solved by the invention]

[0010] Despite the current knowledge related to globin proteins, their production and / or uses and applications, there remains an ongoing unmet need in the art. In particular, due to the increasing demand for plant-based meat substitutes, plant-based (meat) proteins, non-animal derived meat-like products, etc., there continues to be an ongoing unmet need in the art for enhanced globin (protein) expression systems suitable for cost-effective large-scale production of globin proteins. [Means for solving the problem]

[0011] As described later herein, the present disclosure addresses ongoing unmet needs in the art related to the production of heterologous globin proteins. More particularly, as shown and described herein, certain embodiments of the present disclosure provide, inter alia, novel methods and compositions for the production of globin proteins, recombinant filamentous fungal strains with enhanced globin protein productivity phenotypes, polynucleotides (e.g., expression cassettes) encoding globin proteins, polynucleotide (linker DNA) sequences encoding protein / amino acid cleavage sites, industrial-scale fermentation processes, and the like. Accordingly, certain embodiments are directed to recombinant filamentous fungal cells capable of producing heterologous globin proteins for use in, inter alia, food materials, food ingredients, flavor modifiers, aroma modifiers, and the like.

[0012] In certain embodiments, the present disclosure relates to recombinant filamentous fungal cells that express a heterologous globin protein. In related embodiments, the present disclosure provides recombinant filamentous fungal cells that, when fermented under suitable conditions, express and secrete the heterologous globin protein into the fermentation broth. In one or more other embodiments, the recombinant filamentous fungal cells comprise an introduced expression cassette encoding the globin protein.

[0013] In certain embodiments, the expression cassette comprises an upstream promoter (pro) sequence operably linked to a downstream nucleic acid (sig-seq) encoding a protein (signal) secretion sequence operably linked to a downstream (3') nucleic acid (globin CDS) encoding at least a globin protein (e.g., 5'-[pro]-[sig-seq]-[globin CDS]). In other related embodiments, the nucleic acid encoding the globin protein (globin CDS) can comprise an upstream nucleic acid encoding an N-terminal protein fusion (N-fusion) and / or an upstream nucleic acid encoding an N-terminal protein cleavage site (N-linker) and / or a downstream nucleic acid encoding a C-terminal protein fusion (C-fusion) and / or a downstream nucleic acid encoding a C-terminal protein cleavage site (C-linker) and / or an upstream nucleic acid encoding an N-terminal protein fusion (N-fusion) and / or a downstream nucleic acid encoding a C-terminal protein fusion (C-fusion), and / or a combination thereof.

[0014] In certain other embodiments, one or more expression cassettes are integrated into the genome of the cells. In other embodiments, the recombinant filamentous fungal cells comprise one or more introduced expression cassettes encoding one or more protease inhibitor proteins. In other embodiments, the recombinant fungal cells are fermented for at least about 96 hours to about 300 hours to produce globin proteins. In related embodiments, the recombinant fungal cells are fermented for at least about 96 hours to about 300 hours and produce at least 0.1, 0.2, 0.3, 0.4, or 0.5 grams of globin protein per liter of fermentation broth (g / L). In certain other embodiments, the recombinant fungal cells are fermented for about 180-190 hours and produce at least 1.0 gram of globin protein per liter of fermentation broth (g / L). In related embodiments, the globin proteins produced are selected from the group consisting of leghemoglobin, myoglobin, hemoglobin, cyanoglobin, and non-symbiotic hemoglobin.

[0015] In other embodiments, the present disclosure provides methods for producing a heterologous globin protein in a filamentous fungal cell. In certain embodiments, the method includes, but is not limited to, introducing into the filamentous fungal cell an expression cassette encoding a globin protein (the cassette comprising at least an upstream promoter (pro) sequence operably linked to a downstream nucleic acid (sig-seq) encoding a protein (signal) secretion sequence operably linked to a downstream (3') nucleic acid encoding the globin protein (globin CDS)) and fermenting the engineered cell under conditions suitable for production of the globin protein. In certain embodiments of the method, the filamentous fungal cell is selected from the group consisting of an Acremonium sp. cell, an Aspergillus sp. cell, an Emericella sp. cell, a Fusarium sp. cell, a Humicola sp. cell, a Mucor sp. cell, a Myceliophthora sp. cell, a Neurospora sp. cell, a Penicillium sp. cell, a Scytalidium sp. cell, a Talaromyces sp. cell, a Thielavia sp. cell, a Tolypocladium sp. cell, and a Trichoderma sp. cell. In other embodiments of the method, one or more expression cassettes encoding globin proteins are integrated into the genome of the cell. In other embodiments of the method, the fungal cell comprises introduced expression cassettes encoding at least two globin proteins and / or comprises at least two introduced expression cassettes encoding at least two globin proteins. In certain preferred embodiments of the method, the recombinant fungal cell comprises an introduced expression cassette encoding a protease inhibitor.In still other embodiments of the method, the recombinant fungal cells are fermented for at least about 96 hours to about 300 hours and produce at least 0.1, 0.2, 0.3, 0.4, or 0.5 grams of globin protein per liter of fermentation broth (g / L). In certain other embodiments, the fungal cells are fermented for about 180 to 190 hours and produce at least 1 gram of globin protein per liter of fermentation broth (g / L). In still other embodiments of the method, the expressed globin is secreted and recovered from the fermentation broth, and the recovered globin protein is optionally purified. Thus, in certain other embodiments of the method, the secreted globin protein is selected from the group consisting of leghemoglobin, myoglobin, hemoglobin, cyanoglobin, and non-symbiotic hemoglobin. [Brief explanation of the drawings]

[0016] [Figure 1] 1 shows the amino acid and codon-optimized DNA sequences encoding exemplary globin proteins. In particular, FIG. 1 shows the amino acid sequences of natural soybean leghemoglobin (SEQ ID NO: 1) encoded by the DNA of SEQ ID NO: 2, natural bovine myoglobin (SEQ ID NO: 3) encoded by the DNA of SEQ ID NO: 4, and natural kidney bean leghemoglobin (SEQ ID NO: 18) encoded by the DNA of SEQ ID NO: 19.

[0017] [Figure 2] Shown are the native BASI protein (SEQ ID NO: 5), the codon-optimized DNA sequence encoding the native BASI protein (SEQ ID NO: 6), the Cbh1 core domain protein containing a 17 amino acid signal peptide at the N-terminus (SEQ ID NO: 7), and the full-length Cbh1 protein containing a 17 amino acid signal peptide at the N-terminus (SEQ ID NO: 9).

[0018] [Figure 3]Figure 3 shows centrifuged culture broth collected at 24-hour intervals during fed-batch fermentation of T. reesei strain BFZ28, which expresses / produces soybean leghemoglobin by secretion. As shown in Figure 3, fermentation broth samples were taken at 21, 48, 67, 91, 120, 143, 167, and 188 hours from the start of the fermentation run.

[0019] [Figure 4] 1 shows the total secreted protein titer from a fermentation run of strain BFZ28 (see, e.g., FIG. 3). The titer represents the total soluble protein present in the culture supernatant, including the CBH1 core domain protein, soybean leghemoglobin protein, and other background proteins secreted by the T. reesei host strain.

[0020] [Figure 5] 5 shows an SDS-PAGE analysis for a BFZ28 strain fermentation run; molecular weight markers (kDa) are shown to the left of the gel, and the CBH1 core protein, leghemoglobin protein / BASI protease inhibitor are shown with labels to the right of the gel. More specifically, as shown in FIG. 5, the CBH1 core protein has an approximate molecular weight of about 49 kDa, the leghemoglobin protein has an approximate Mw of about 15.5 kDa, and the BASI protease inhibitor has an approximate Mw of about 20 kDa.

[0021] [Figure 6] HPLC analysis of a 188-hour supernatant sample from a fermentation run of strain BFZ28 is shown. Notably, the heme group of leghemoglobin protein was detected at a wavelength of 410 nm (heme prosthetic group), which co-migrated with a protein peak detected at 280 nm with a retention time of 6.4 minutes. The inset (Figure 6) shows a spectral scan of leghemoglobin protein, with a peak at a retention time of approximately 6.4 minutes (410 nm wavelength), confirming the peak absorbance at 410 nm of the heme-containing leghemoglobin protein.

[0022] [Figure 7] Figure 1 shows the overall secreted protein titers from fermentation runs of strains BFZ28 (Pcbh1-CBH1core-KEX2-LegGm1b), BGJ74 (Pcbh1-LegGm1b, Pcbh2-BASI), BGJ75 (Pcbh1-Pv1b, Pcbh2-BASI), and BGJ76 (Pcbh1-LegGm1b.Pep1ss, Pcbh2-BASI). The titers represent the overall soluble protein present in the culture supernatant, including the CBH1 core domain protein (BFZ28), BASI proteins (BGJ74, BGJ75, and BGJ76), soybean leghemoglobin protein (BFZ28, BGJ74, BGJ75, and BGJ76), and other background proteins secreted by the T. reesei host strain.

[0023] [Figure 8] SDS-PAGE analysis of fermentation runs of strains BGJ74, BGJ75, and BGJ76 is shown. As shown in Figure 8, the red arrow indicates the BASI and leghemoglobin bands that co-migrate at approximately 20 kDa.

[0024] [Figure 9] Chromatograms of HPLC analysis of 188-hour supernatant samples from the 188-hour end of fermentation runs of strains BGJ74, BGJ75, and BGJ76 are shown. As shown in Figure 9, a protein peak is detected at 280 nm. The heme prosthetic group of the leghemoglobin protein is detected at 410 nm in Figure 9, which co-migrated with the protein peak detected at 280 nm with a retention time of 6.4 minutes.

[0025] [Figure 10]

[0033] Figure 1 shows an SDS-PAGE gel analysis of a fermentation run of strain BHX46. The first sample lane (labeled "C") contains equine myoglobin. Subsequent lanes contain culture supernatant samples collected at 43, 72, 94, 114, 137, 161, and 186 hours. DETAILED DESCRIPTION OF THE INVENTION

[0026] A brief description of biological sequences SEQ ID NO: 1 is the amino acid sequence of native Glycine max (soybean) leghemoglobin protein (NCBI catalog: NP_001235248).

[0027] SEQ ID NO:2 is a nucleic acid (DNA) sequence encoding the native leghemoglobin protein of SEQ ID NO:1, which is codon-optimized for expression in T. reesei fungal cells.

[0028] SEQ ID NO: 3 is the amino acid sequence of native Bos taurus (cow) myoglobin protein (NCBI Catalog: NP_776306.1; GI: 27806939).

[0029] SEQ ID NO:4 is the DNA sequence encoding the native myoglobin protein of SEQ ID NO:3, which is codon-optimized for expression in T. reesei fungal cells.

[0030] SEQ ID NO: 5 is the amino acid sequence of the native barley amylase subtilisin inhibitor (BASI) protein (NCBI Catalog: 1210227A).

[0031] SEQ ID NO: 6 is the DNA sequence encoding the native BASI protein of SEQ ID NO: 5, which is codon-optimized for expression in T. reesei fungal cells.

[0032] SEQ ID NO: 7 is the amino acid sequence of the Cbh1 core domain protein.

[0033] SEQ ID NO: 8 is the DNA sequence encoding the Cbh1 core domain.

[0034] SEQ ID NO: 9 is the amino acid sequence of the native (full length) Cbh1 protein (NCBI catalog: A0A024RXP8).

[0035] SEQ ID NO: 10 is the DNA sequence encoding the full-length Cbh1 protein.

[0036] SEQ ID NO: 11 is the cbh1 promoter region (DNA) sequence of the T. reesei chb1 gene.

[0037] SEQ ID NO: 12 is the cbh2 promoter region (DNA) sequence of the T. reesei chb2 gene.

[0038] SEQ ID NO: 13 is the DNA sequence encoding the Cbh1 signal peptide sequence.

[0039] SEQ ID NO: 14 is the DNA sequence encoding the signal peptide of the Pep1 protein.

[0040] SEQ ID NO: 15 is the terminator (DNA) sequence of the cbh1 gene.

[0041] SEQ ID NO: 16 is the T. reesei pyr2 gene (marker) encoding orotate phosphoribosyltransferase.

[0042] SEQ ID NO: 17 is the DNA sequence encoding the A. nidulans acetamidase.

[0043] SEQ ID NO: 18 is the DNA sequence encoding the Kex2 protease cleavage site.

[0044] SEQ ID NO: 19 is the amino acid sequence of native kidney bean (Phaseolus vulgaris) leghemoglobin (NCBI catalog: AAA33767.1).

[0045] SEQ ID NO:20 is the DNA sequence encoding the native leghemoglobin protein of SEQ ID NO:19, which is codon-optimized for expression in T. reesei fungal cells.

[0046] SEQ ID NO: 21 is the sequence of a synthetic DNA primer designated OT4268.

[0047] SEQ ID NO: 22 is the sequence of a synthetic DNA primer designated OT4269.

[0048] SEQ ID NO: 23 is the sequence of a synthetic DNA primer designated OT4337.

[0049] SEQ ID NO: 24 is the sequence of a synthetic DNA primer designated OT4338.

[0050] SEQ ID NO: 25 is the sequence of a synthetic DNA primer designated OT4333.

[0051] SEQ ID NO: 26 is the sequence of a synthetic DNA primer designated OT4334.

[0052] SEQ ID NO: 27 is an integration cassette designated HG1.

[0053] SEQ ID NO: 28 is an integration cassette designated HG2.

[0054] SEQ ID NO: 29 is an integration cassette designated HG3.

[0055] SEQ ID NO: 30 is an integration cassette designated HG4.

[0056] SEQ ID NO: 31 is an integration cassette designated HG5.

[0057] SEQ ID NO: 32 is an integration cassette designated HG6.

[0058] SEQ ID NO: 33 is an integration cassette designated HG7.

[0059] SEQ ID NO: 34 is an integration cassette designated HG8.

[0060] SEQ ID NO: 35 is an integration cassette designated HG9.

[0061] SEQ ID NO: 36 is a synthetic cassette designated pCHL853.

[0062] SEQ ID NO: 37 is a synthetic cassette designated pCHL856.

[0063] SEQ ID NO: 38 is a synthetic cassette designated pLH1088.

[0064] SEQ ID NO: 39 is a synthetic cassette designated pLH1104.

[0065] SEQ ID NO: 40 is a synthetic cassette designated pLH1105.

[0066] SEQ ID NO: 41 is a synthetic cassette designated pLH1106.

[0067] SEQ ID NO: 42 is a synthetic cassette designated pLH1107.

[0068] SEQ ID NO: 43 is a synthetic cassette designated pLH1108.

[0069] SEQ ID NO: 44 is a synthetic cassette designated pLH1109.

[0070] SEQ ID NO: 45 is a synthetic single guide RNA designated "sgRNA-TrC144F."

[0071] SEQ ID NO: 46 is a synthetic DNA sequence containing the TrpC transcription terminator sequence.

[0072] SEQ ID NO: 47 is a synthetic cassette designated pCHL852.

[0073] SEQ ID NO: 48 is an integration cassette designated HG10.

[0074] SEQ ID NO:49 is a second codon-optimized bovine myoglobin gene (Mb1c) that encodes the same bovine myoglobin protein of SEQ ID NO:3.

[0075] SEQ ID NO: 50 is a synthetic DNA containing the T. reesei Egl1 terminator (Tegl1).

[0076] SEQ ID NO: 51 is a tandem-copy expression vector designated "pLHX143."

[0077] Detailed Description As described herein, certain embodiments of the present disclosure provide, inter alia, compositions and methods for the production of globin proteins, recombinant filamentous fungal strains comprising enhanced globin protein productivity phenotypes, polynucleotides (e.g., expression constructs) encoding one or more globin proteins, industrial-scale fermentation and recovery processes for globin proteins, and the like. Accordingly, in one or more embodiments or aspects, the present disclosure relates to recombinant (modified) filamentous fungal cells capable of producing heterologous globin proteins. In certain embodiments or aspects, the recombinant filamentous fungal cells comprise introduced expression cassettes encoding one or more heterologous globin proteins of interest. In certain embodiments, the one or more cassettes comprise an upstream (5′) promoter (pro) region sequence operably linked to a downstream nucleic acid (sig-seq) encoding a protein secretion (signal) sequence, which is operably linked to a downstream (3′) nucleic acid (globin CDS) encoding the globin protein. In other one or more embodiments or aspects, the recombinant filamentous fungal cells comprising the one or more introduced cassettes are fermented under conditions suitable for the production of globin proteins. In certain other embodiments, the secreted globin protein is recovered from the end of fermentation (EOF) broth. More particularly, as shown and described later herein, the recombinant fungal strains of the present disclosure are particularly well suited for growth in submerged culture for large-scale production of heterologous globin proteins.

[0078] I. Definition Before describing the strains and methods in detail, the following terms are defined for clarity. Terms not defined should be accorded their usual meaning as used in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the compositions and methods belong.

[0079] All publications and patents cited herein are hereby incorporated by reference.

[0080] Where a range of values ​​is presented, unless the context clearly dictates otherwise, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limits of that range, and any other stated or intervening value within that stated range, is encompassed within the compositions and methods. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the compositions and methods of the invention, subject to any specifically excluded limit in the stated range. When the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the compositions and methods.

[0081] Certain ranges are described herein by numerical values ​​preceded by the term "about." The term "about" is used herein to provide literal support for the exact number it precedes and a number that is close to or approximately the number preceded by the term. In determining whether a number is close to or approximately a specifically recited number, the unrecited near or approximately number may be a number that, in the context in which it is presented, provides a substantial equivalent number to the specifically recited number. For example, in connection with numerical values, the term "about" does not necessarily mean the exact number of the number unless the term is clearly defined otherwise in the context. - 10%~ + As another example, the phrase "a pH value of about 6" refers to a pH value of 5.4 to 6.6, unless the pH value is specifically defined otherwise.

[0082] The headings provided herein are not limitations of the various aspects or embodiments of the present compositions and methods, which can be had by reference to the specification as a whole, and accordingly, the terms defined immediately below are more fully defined by reference to the specification as a whole.

[0083] In accordance with this detailed description, the following abbreviations and definitions apply. Note that the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to an "enzyme" includes a plurality of such enzymes, reference to a "dosage" includes a reference to one or more dosages and equivalents thereof known to those skilled in the art, and so forth.

[0084] It is further noted that the claims may be drafted to exclude any optional element. Accordingly, this statement is intended to serve as a prelude to the use of exclusive terminology such as "solely," "only," "excluding," and "not including" or the use of a "negative" limitation in connection with the recitation of claim elements.

[0085] It is further noted that the term "comprises," as used herein, means "including but not limited to" the component following the term "comprises." The component following the term "comprises" is required or essential, but a composition including that component may further include other non-essential or optional components.

[0086] It should also be noted that the term "consisting of," as used herein, means to include and be limited to the component(s) following the term "consisting of." Thus, the component(s) following the term "consisting of" are required or essential, and other components are not present in the composition.

[0087] It will be apparent to those skilled in the art upon reading this disclosure that each of the individual embodiments depicted and illustrated herein has individual components and features that can be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the compositions and methods described herein. Any described method can be carried out in the order of events recited or in any other order that is logically possible.

[0088] As used herein, the terms "wild-type" and "native" are used interchangeably to refer to a gene, protein, fungal cell, or strain as found in nature.

[0089] As used herein, the terms "recombinant" or "non-naturally occurring" refer to an organism, microorganism, cell, nucleic acid molecule, or vector that has at least one engineered genetic change or that has been modified by the introduction of a heterologous nucleic acid molecule, or to a cell (e.g., a microbial cell) that has been altered so that expression of a heterologous or endogenous nucleic acid molecule or gene can be controlled. Recombinant also refers to a cell that is derived from, or is the progeny of, a non-naturally occurring cell that has one or more such changes. Genetic changes include, for example, a change that introduces an expressible nucleic acid molecule that encodes a protein, an addition, deletion, or substitution of a nucleic acid molecule, or other functional alteration of the cell's genetic material. For example, a recombinant cell may express a gene or other nucleic acid molecule that is not found in the same or homologous form in a native (wild-type) cell, or may provide an altered expression pattern of an endogenous gene, such that it is overexpressed, underexpressed, minimally expressed, or not expressed at all.

[0090] "Recombination," "recombining," or producing a "recombinant" nucleic acid generally refers to the assembly of two or more nucleic acid fragments, which assembly results in a chimeric gene.

[0091] As used herein, the term "gene" is synonymous with the term "allele," which refers to a nucleic acid that encodes and directs the expression of a protein or RNA. Because vegetative propagation forms of filamentous fungi are generally haploid, a single copy of a particular gene (i.e., a single allele) is sufficient to confer a particular phenotype.

[0092] As used herein, the term "gene" refers to a segment of DNA involved in producing a polypeptide (protein) chain, which may or may not include regions preceding and following the coding region (e.g., 5' untranslated (5' UTR) or "leader" sequence, 3' UTR or "trailer" sequence, promoter sequence, and terminator sequence), and intervening sequences (introns) between individual coding segments (exons). For example, a gene (DNA) sequence of interest (GOI) may encode a globin protein of interest, a structural protein, a commercially important industrial protein or peptide, such as an enzyme (e.g., protease, mannanase, xylanase, amylase, glucoamylase, cellulase, oxidase, phytase, lipase), etc. A gene of interest can be a naturally occurring gene, a mutated (modified) gene, or a synthetic gene.

[0093] As used herein, the term "promoter" refers to a nucleic acid sequence that functions to direct transcription of a downstream gene or its open reading frame (ORF). A promoter will generally be appropriate for the host cell (e.g., a filamentous fungal cell) in which the target gene is being expressed. A promoter, along with other transcriptional and translational regulatory nucleic acid sequences (also referred to as "control sequences"), are necessary to express a given gene. Generally, transcriptional and translational regulatory sequences include, but are not limited to, promoter and terminator sequences, including a core promoter and enhancer or activator or repressor sequences, and transcriptional and translational start and stop sequences. In certain embodiments, the promoter is an inducible promoter or a constitutive promoter. In certain embodiments, the inducible promoter is an inducible cellulase gene promoter.

[0094] As used herein, the term "promoter activity" refers to the ability of a nucleic acid to direct transcription of a downstream (3') polynucleotide in a host cell. To test promoter activity, a (promoter) nucleic acid may be operably linked to a downstream polynucleotide to generate a recombinant nucleic acid. The recombinant nucleic acid may be introduced into a cell to assess transcription of the polynucleotide. In certain cases, the polynucleotide may encode a protein, and transcription of the polynucleotide can be assessed by assessing production of the protein in the cell.

[0095] As used herein, the term "operably linked" refers to a functional linkage between two or more nucleic acid sequences. Thus, a nucleic acid sequence is operably linked when it functionally relates to another nucleic acid sequence. For example, a promoter or terminator sequence is operably linked to a coding sequence if it affects the transcription of the coding sequence; a ribosome binding site is operably linked to a coding sequence if it is positioned so as to facilitate translation; and a nucleic acid sequence encoding a secretory leader (i.e., signal peptide) is operably linked to a nucleic acid sequence encoding a polypeptide (e.g., ORF) if it is expressed as a preprotein that participates in the secretion of the polypeptide. Generally, "operably linked" means that the DNA (nucleic acid) sequences being linked are contiguous, and, in the case of a secretory leader, contiguous and in reading phase. However, enhancers need not be contiguous. Linking (i.e., operably linking) two or more nucleic acid sequences can be accomplished using any number of methods known to those skilled in the art.

[0096] As used herein, a "functional gene" is a gene for which cellular components can be used to produce an active gene product, typically a protein. In contrast, a "non-functional gene" is one for which cellular components cannot be used to produce an active gene product (i.e., a functional protein), or for which the ability of cellular components to be used to produce an active gene product (i.e., a functional protein) is reduced.

[0097] As used herein, a "functional protein" is a protein that has a function or activity, such as an enzymatic function / activity, a binding function / activity (e.g., DNA binding), and a surfactant property, and that has not been mutated, truncated, or otherwise modified to eliminate or reduce that function / activity.

[0098] As used herein, the terms "modified filamentous fungal cell," "mutant or variant filamentous fungal cell," "recombinant fungal cell," "modified filamentous fungal strain," and the like may be used interchangeably and refer to a filamentous fungal cell derived from (obtained from) a control or parent filamentous fungal cell belonging to the subdivision Pezizomycotina. For example, an "modified" filamentous fungal cell may be derived from (obtained from) a control or parent filamentous fungal cell, where the modified cell contains at least one genetic modification not found in the control or parent cell.

[0099] As used herein, the term "ascomycota fungal cell" refers to any organism in the phylum Ascomycota in the kingdom Fungi. Examples of ascomycota fungal cells include, but are not limited to, filamentous fungi in the subphylum Pezizomycotina, such as Trichoderma sp., Aspergillus sp., Myceliophthora sp., and Penicillium sp.

[0100] As used herein, the term "filamentous fungi" refers to all filamentous forms of Eumycota and Oomycota. For example, filamentous fungi include, but are not limited to, species of Acremonium, Aspergillus, Emericella, Fusarium, Humicola, Mucor, Myceliophthora, Neurospora, Penicillium, Scytalidium, Talaromyces, Thielavia, Tolypocladium, or Trichoderma. In some embodiments, the filamentous fungus can be Aspergillus aculeatus, Aspergillus awamori, Aspergillus foetidus, Aspergillus japonicus, Aspergillus nidulans, Aspergillus niger, or Aspergillus oryzae.

[0101] In some embodiments, the filamentous fungus is a Fusarium species, such as Fusarium bactridioides, Fusarium cerealis, Fusarium crookwellense, Fusarium culmorum, Fusarium graminearum, Fusarium graminum, Fusarium heterosporum, Fusarium negundi, Fusarium oxysporum, Fusarium reticulatum, Fusarium roseum, or the like. These include Fusarium roseum, Fusarium sambucinum, Fusarium sarcochroum, Fusarium sporotrichioides, Fusarium sulphureum, Fusarium torulosum, Fusarium trichothecioides, and Fusarium venenatum. In other embodiments, the filamentous fungus is Humicola insolens, Humicola lanuginosa, Mucor miehei, Myceliophthora thermophila, Neurospora crassa, Scytalidium thermophilum, Thielavia terrestris, or the like.In certain other embodiments, the filamentous fungus is Trichoderma harzianum, Trichoderma koningii, Trichoderma longibrachiatum, Trichoderma reesei, Trichoderma viride, or the like.

[0102] As used herein, exemplary parent Trichoderma reesei strains include, but are not limited to, T. reesei strain QM6a (ATCC Accession No. 13631), T. reesei strain RL-P37 (NRRL Accession No. 15709), and T. reesei strain RUT-C30 (ATCC Accession No. 56765); exemplary parent Aspergillus niger strains include, but are not limited to, A. niger strain ATCC Accession No. 1015; exemplary parent Aspergillus oryzae oryzae strains include, but are not limited to, A. oryzae strain RIB40 (ATCC Accession No. 42149); exemplary parent Myceliophthora thermophila strains include, but are not limited to, the M. thermophila strain designated as ATCC Accession No. 42464. For example, Trichoderma strains Rut-RUT-C30 and RL-P37 are mutagenized (cellulase-overproducing) derivatives of Trichoderma natural isolate QM6a (Sheir-Neiss and Montenecourt, 1984), with strain NG14 being the immediate common ancestor. In certain aspects, suitable Trichoderma strains are derived / obtainable from T. reesei strains that contain a deletion of the T. reesei pyr2 gene (Δpyr2), generally as described by Sheir-Neiss and Montenecourt (1984) and WO 2011 / 153449, the entire contents of which are specifically incorporated herein by reference. In certain embodiments, in one or more embodiments, the T. reesei cells / strains of the present disclosure are obtained / derived from T. reesei strain RL-P37.Thus, in certain one or more embodiments, T. reesei cells derived from strain RL-P37 are abbreviated herein as "Tr1 cells," "Tr1 strain," "Tr1 parent" cells, "Tr1 control" cells, etc.

[0103] As used herein, the terms "lignocellulolytic enzymes," "cellulase enzymes," and / or "cellulases" are used interchangeably and include glycoside hydrolase (GH) enzymes, such as cellobiohydrolases, xylanases, endoglucanases, and β-glucosidases, that hydrolyze the β-(1,4)-linked glycosidic bonds of cellulose (hemi-cellulose) to produce glucose.

[0104] In certain embodiments, cellobiohydrolases include enzymes classified under the Enzyme Commission number (EC 3.2.1.91), endoglucanases include enzymes classified under EC 3.2.1.4, endo-β-1,4-xylanases include enzymes classified under EC 3.2.1.8, β-xylosidases include enzymes classified under EC 3.2.1.37, and β-glucosidases include enzymes classified under EC 3.2.1.21.

[0105] As used herein, "endoglucanase" proteins may be abbreviated as "EG," "cellobiohydrolase" proteins may be abbreviated as "CBH," "β-glucosidase" proteins may be abbreviated as "BG," and "xylanase" proteins may be abbreviated as "XYL." Thus, as used herein, genes (gene CDS or ORFs) encoding EG proteins may be abbreviated as "eg," genes (or ORFs) encoding CBH proteins may be abbreviated as "cbh," genes (or ORFs) encoding BG proteins may be abbreviated as "bg," and genes (or ORFs) encoding XYL proteins may be abbreviated as "xyl."

[0106] As used herein, a "globin" or "globin protein" is a metalloprotein that contains a "porphyrin prosthetic group." In particular, globin proteins incorporate a series of α-helical segments known as a globin fold that accommodate / bind the porphyrin prosthetic group. Globin proteins include, but are not limited to, "leghemoglobin," "myoglobin," and "hemoglobin." In certain aspects, the porphyrin prosthetic group confers functionality to the globin protein, which may include oxygen transport or transport, oxygen reduction, electron transfer, and other processes.

[0107] As used herein, the term "porphyrin" has the same meaning as understood in the art; porphyrins are a group of heterocyclic macrocyclic organic compounds composed of four modified pyrrole subunits interconnected at their α-carbon atoms via methine bridges. In particular, porphyrins (ring structures) are often described as highly conjugated aromatic rings that strongly absorb electromagnetic radiation in the visible region of the spectrum. For example, porphyrins bind to metal ions in the N4 pocket by simultaneous replacement of two NH protons, and metal ions typically bind to two + or 3 + In particular, when there is no metal ion (or atom) bound to the nitrogen in the ring (structural center), the compounds are called "free porphyrins," while when they are bound to a metal ion (or atom) in the ring (structural center), they are called "bound porphyrins." Examples of porphyrins with iron atoms bound include, but are not limited to, myoglobin and hemoglobin. In particular, one of the most well-known families of porphyrin complexes is heme.

[0108] As used herein, the term "hemin" refers to ferric iron (Fe) with a coordinated chloride ligand. 3+ ) ion (protoporphyrin IX).

[0109] As used herein, phrases such as "retaining globin protein function or activity" and "comprising globin protein function or activity" mean that heme (porphyrin) is bound to the heme (porphyrin)-binding pocket of a globin protein in a pentacoordinated state for some globins (e.g., leghemoglobin) or a hexacoordinated state for other globins (e.g., cyanoglobin). Thus, in certain embodiments or aspects of the present disclosure, functional globin protein may be assayed and detected according to the bound heme (porphyrin) prosthetic group, which can be measured / detected based on the UV-Vis absorbance of the porphyrin molecule. For example, the bound heme described in this example has a signature UV-Vis absorption peak at 410 nm.

[0110] As used herein, a "native soybean leghemoglobin protein" comprises at least about 90% to 100% identity to the native Glycine max (soybean) leghemoglobin protein of SEQ ID NO: 1. In certain embodiments, the native soybean leghemoglobin protein comprises at least about 90% to 100% identity to the native soybean leghemoglobin protein of SEQ ID NO: 1 and retains (includes) a function or activity of native leghemoglobin.

[0111] As used herein, a "gene coding sequence (CDS)" encoding a native soybean leghemoglobin protein encodes a leghemoglobin protein that comprises at least about 40% to 100% identity to the native protein of SEQ ID NO: 1. In one or more certain embodiments, the gene CDS encoding a native soybean leghemoglobin protein comprises at least about 90% to 100% identity to SEQ ID NO: 1 and retains the function or activity of native leghemoglobin. In certain embodiments, the wild-type (WT) soybean leghemoglobin C2 gene CDS of the present disclosure is abbreviated as "LegGm1b," and the WT LegGm1b gene CDS has been codon-optimized for expression in T. reesei cells as set forth in SEQ ID NO: 2.

[0112] As used herein, a "native kidney bean leghemoglobin protein" comprises at least about 90% to 100% identity to the native Phaseolus vulgaris (kidney bean) leghemoglobin protein of SEQ ID NO: 19. In certain embodiments, the native kidney bean leghemoglobin protein comprises at least about 90% to 100% identity to the native kidney bean leghemoglobin protein of SEQ ID NO: 19 and comprises a function or activity of native leghemoglobin. In certain embodiments, the WT kidney bean leghemoglobin gene CDS of the present disclosure is abbreviated as "LegPv," and the WT LegPv gene CDS has been codon-optimized for expression in T. reesei cells as set forth in SEQ ID NO: 20. In other embodiments, the gene CDS encoding the native kidney bean leghemoglobin protein encodes a leghemoglobin protein that comprises at least about 40% to 100% identity to the native protein of SEQ ID NO:19.

[0113] As used herein, a "native bovine (Bos taurus) myoglobin protein" comprises at least about 80% to 100% identity to the native myoglobin protein of SEQ ID NO: 3. In certain embodiments, the native myoglobin protein comprises at least about 80% to 100% identity to the native myoglobin protein of SEQ ID NO: 3 and retains (inclusive) a function or activity of native myoglobin. In one or more certain embodiments, the genetic CDS encoding the native myoglobin protein encodes a myoglobin protein comprising at least about 80% to 100% identity to the wild-type protein of SEQ ID NO: 3. In one or more certain embodiments, the genetic CDS encoding the native myoglobin protein comprises at least about 80% to 100% identity to the native myoglobin protein of SEQ ID NO: 3 and retains (inclusive) a function or activity of native myoglobin. In certain aspects, the WT myoglobin gene CDS is abbreviated as "Mb1b," and the WT Mb1b gene CDS has been codon-optimized for expression in T. reesei cells as set forth in SEQ ID NO: 4. In certain other embodiments, the gene CDS encoding a native bovine myoglobin protein encodes a myoglobin protein comprising at least about 40%-100% identity to the native protein of SEQ ID NO: 3. In certain other embodiments, the gene CDS encoding a native bovine myoglobin protein encodes a myoglobin protein comprising at least about 40%-100% identity to the native protein of SEQ ID NO: 3.

[0114] As used herein, phrases such as "protease inhibitor protein," "protein protease inhibitor," or "protease inhibitor" may be used interchangeably, and a protease inhibitor protein may be any peptide or protein that reversibly inhibits a protease of interest. According to the MEROPS database (www.ebi.ac.uk / merops / ), a repository of proteases / peptidases and proteins that inhibit these protein / peptide hydrolases, protease inhibitors can be classified into 38 clans and subdivided into 78 families. In certain embodiments, protease inhibitors include, but are not limited to, trypsin inhibitor proteins and subtilisin inhibitor proteins. Examples are trypsin inhibitors and subtilisin inhibitors known in the art, such as those generally described in Laskowski and Kato (1980), Strickler et al. (1992), PCT Publication No. WO 1992 / 03529, etc. In certain embodiments, exemplary protease inhibitors are family IV trypsin inhibitors and family III, VI and VII subtilisin inhibitors.Examples include naturally occurring Streptomyces subtilisin inhibitors (SSIs) and functional variants thereof, naturally occurring Streptomyces antifibrinolytic plasminostreptin inhibitors and functional variants thereof, naturally occurring barley subtilisin inhibitors (BASIs) and functional variants thereof, naturally occurring potato subtilisin inhibitors and functional variants thereof, naturally occurring tomato subtilisin inhibitors and functional variants thereof, naturally occurring eglin C inhibitors and functional variants thereof, naturally occurring Vicia faba These include, but are not limited to, a series of protease inhibitors such as (faba) subtilisin inhibitors and functional variants thereof, natural leupeptin inhibitors and functional variants thereof, natural soybean trypsin inhibitors and functional variants thereof, natural pea aspartic protease inhibitors and functional variants thereof, Bowman-Birk inhibitors (BBI) and functional variants thereof, natural serpin inhibitors and functional variants thereof, natural phytocystatin inhibitors and functional variants thereof, natural Kunitz-type inhibitors (KTI) and functional variants thereof, bifunctional α-amylase-trypsin inhibitors and functional variants thereof, mustard-type inhibitors, potato metallocarboxypeptidase inhibitors and functional variants thereof, natural pumpkin inhibitors and functional variants thereof, and natural cyclotide inhibitors and functional variants thereof.

[0115] As used herein, a "native barley amylase subtilisin (protease) inhibitor-encoding gene" protein comprises at least about 90% to 100% identity to the wild-type BASI gene set forth in SEQ ID NO: 6. In certain aspects, the wild-type barley amylase subtilisin inhibitor gene is abbreviated as "BASI" (italics). In certain one or more embodiments or aspects of the present disclosure, the wild-type BASI gene is codon-optimized for expression in a Trichoderma strain.

[0116] As used herein, a "native barley amylase subtilisin (protease) inhibitor" protein (abbreviated herein as "BASI" protein) comprises protease inhibitor activity and at least about 90% to 100% identity to the native BASI protein of SEQ ID NO: 6. For example, one or more variant BASI proteins may be derived from the native BASI protein (SEQ ID NO: 6), and the native or variant BASI proteins are particularly suitable for reducing / alleviating certain undesirable protease activities as described herein.

[0117] As used herein, the modified T. reesei strain designated "BFZ28" is derived from the Tr1 parent strain and contains an introduced expression cassette (Pcbh1-CBH1core-KEX2-LegGm1b) encoding a heterologous (soybean) leghemoglobin protein.

[0118] As used herein, the modified T. reesei strain designated "BGJ74" is derived from the Tr1 parent strain and contains an introduced cassette encoding a heterologous (soybean) leghemoglobin protein (Pcbh1-LegGm1b) and an introduced cassette encoding a protease inhibitor (Pcbh2-BASI).

[0119] As used herein, the modified T. reesei strain designated "BGJ75" is derived from the Tr1 parent strain and contains an introduced cassette encoding a heterologous (phase bean) leghemoglobin protein (Pcbh1-Pv1b) and an introduced cassette encoding a protease inhibitor (Pcbh2-BASI).

[0120] As used herein, the modified T. reesei strain designated "BGJ76" is derived from the Tr1 parent strain and contains an introduced cassette encoding a heterologous (soybean) leghemoglobin protein (Pcbh1-LegGm1b.Pep1ss) and an introduced cassette encoding a protease inhibitor (Pcbh2-BASI).

[0121] As described in Table 1 in the Examples section below, various plasmids (vectors) have been designed, constructed, and screened for secreted expression of globin proteins in filamentous fungal strains. In particular, Table 1 provides the names (column 1) and descriptions (column 2) of the various vectors constructed, containing relevant genetic elements such as heterologous promoter regions (column 3), signal (secretion) sequences (column 4), N-terminal fusion (protein) sequences (column 5), gene CDS descriptions (column 6), and C-terminal fusion (protein) sequences (column 7) as described and exemplified herein. Table 1 also includes the names of the 5' PCR primer (column 9) and 3' PCR primer (column 10) used herein; the primer DNA sequences are listed in the Sequence Listing.

[0122] As used herein, chromosomal integration cassettes for the expression of leghemoglobin include cassette "HG1" (SEQ ID NO: 27; pI1-Pcbh1-LegGm1b), cassette "HG2" (SEQ ID NO: 28; pI1-Pcbh1-LegGm1b-Pep1ss), cassette "HG3" (SEQ ID NO: 29; pI1-Pcbh1-Cbh1core-LegGm1b), cassette "HG4" (SEQ ID NO: 30; pI1-Pcbh1-Cbh1FL-LegGm1b), cassette "HG5" (SEQ ID NO: 31; pI1-Pcbh1-Cbh1-LegGm1b-His6), and cassette "HG9" (SEQ ID NO: 35; pI1c-Pcbh1-BASI-LegGm1b).

[0123] As used herein, chromosomal integration cassettes for the expression of myoglobin include cassette "HG6" (SEQ ID NO: 32; pI1-Pcbh1-Cbh1FL-Mb1b) and cassette "HG7" (SEQ ID NO: 33; pI1-Pcbh1-Cbh1FL-Mb1b-His6).

[0124] As used herein, the chromosomal integration cassette for expression of the protease inhibitor (BASI) is designated cassette "HG8" (SEQ ID NO: 34; pKS923-Pcbh2-BASI).

[0125] As used herein, plasmids designated "pCHL852" and "pCHL853" contain codon-optimized genes encoding kidney bean leghemoglobin and soybean leghemoglobin, respectively. More specifically, plasmid pCHL852 (SEQ ID NO: 47) contains an upstream (5') cbh1 promoter sequence operably linked to a downstream DNA sequence encoding a Cbh1 signal sequence operably linked to a downstream gene CDS (LegPv1b) encoding the leghemoglobin protein, and plasmid pCHL853 (SEQ ID NO: 36) contains an upstream (5') cbh1 promoter sequence operably linked to a downstream DNA sequence encoding a Cbh1 (protein) signal sequence operably linked to a downstream gene CDS (LegGm1b) encoding the leghemoglobin protein.

[0126] As used herein, the plasmid designated "pCHL856" contains a codon-optimized gene encoding soybean leghemoglobin. More specifically, plasmid pCHL856 (SEQ ID NO: 37) contains an upstream (5') cbh1 promoter sequence (SEQ ID NO: 11) operably linked to a downstream DNA sequence encoding a Pep1 (protein) signal sequence (SEQ ID NO: 14), which is operably linked to a downstream gene CDS (LegGm1b) encoding the leghemoglobin protein (SEQ ID NO: 1).

[0127] As used herein, the plasmid designated "pLH1088" contains a codon-optimized gene encoding soybean leghemoglobin. More specifically, plasmid pLH1088 (SEQ ID NO: 38) contains an upstream (5') cbh1 promoter sequence (SEQ ID NO: 11) operably linked to a downstream DNA sequence encoding a Cbh1 signal sequence (SEQ ID NO: 13), which is operably linked to a downstream DNA sequence encoding a Cbh1 protein core domain (abbreviated as "Cbh1 core"; SEQ ID NO: 7), which is operably linked to a downstream gene CDS (LegGm1b) encoding the leghemoglobin protein (SEQ ID NO: 1).

[0128] As used herein, the plasmid designated "pLH1104" contains a codon-optimized gene encoding soybean leghemoglobin. More specifically, plasmid pLH1104 (SEQ ID NO: 39) contains an upstream (5') cbh1 promoter sequence (SEQ ID NO: 11) operably linked to a downstream DNA sequence encoding a Cbh1 signal sequence (SEQ ID NO: 13), which is operably linked to a downstream DNA sequence encoding a full-length Cbh1 protein (abbreviated as "Cbh1 FL"; SEQ ID NO: 9), which is operably linked to a downstream gene CDS (LegGm1b) that encodes the leghemoglobin protein (SEQ ID NO: 1).

[0129] As used herein, the plasmid designated "pLH1105" contains a codon-optimized gene encoding soybean leghemoglobin. More specifically, plasmid pLH1105 (SEQ ID NO: 40) contains an upstream (5') cbh1 promoter sequence (SEQ ID NO: 11) operably linked to a downstream DNA sequence encoding a Cbh1 signal sequence (SEQ ID NO: 13) operably linked to a downstream DNA sequence encoding a Cbh1 FL protein (SEQ ID NO: 9) operably linked to a downstream gene CDS (LegGm1b) encoding a leghemoglobin protein (SEQ ID NO: 1) operably linked to a downstream 6-histidine (6-His) tag.

[0130] As used herein, the plasmid designated "pLH1106" contains a codon-optimized gene encoding bovine myoglobin. More specifically, plasmid pLH1106 (SEQ ID NO:41) contains an upstream (5') cbh1 promoter sequence (SEQ ID NO:11) operably linked to a downstream DNA sequence encoding a Cbh1 signal sequence (SEQ ID NO:13), which is operably linked to a downstream DNA sequence encoding a Cbh1 FL protein (SEQ ID NO:9), which is operably linked to a downstream gene CDS (Mb1b) encoding a myoglobin protein (SEQ ID NO:3).

[0131] As used herein, the plasmid designated "pLH1107" contains a codon-optimized gene encoding bovine myoglobin. More specifically, plasmid pLH1107 (SEQ ID NO:42) contains an upstream (5') cbh1 promoter sequence (SEQ ID NO:11) operably linked to a downstream DNA sequence encoding a Cbh1 signal sequence (SEQ ID NO:13) operably linked to a downstream DNA sequence encoding a Cbh1 FL protein (SEQ ID NO:9) operably linked to a downstream gene CDS (Mb1b) encoding a myoglobin protein (SEQ ID NO:3) operably linked to a downstream 6-histidine tag.

[0132] As used herein, the plasmid designated "pLH1108" contains a gene encoding the barley amylase subtilisin inhibitor (BASI) protein. More specifically, plasmid pLH1108 (SEQ ID NO:43) contains an upstream (5') cbh2 promoter sequence (SEQ ID NO:12) operably linked to a downstream DNA sequence encoding the Pep1 signal sequence (SEQ ID NO:14), which is operably linked to a downstream DNA sequence (BASI) encoding the BASI protein (SEQ ID NO:5).

[0133] As used herein, the plasmid designated "pLH1109" contains a codon-optimized gene encoding the BASI protein. More specifically, plasmid pLH1109 (SEQ ID NO:44) contains an upstream (5') cbh1 promoter sequence (SEQ ID NO:11) operably linked to a downstream DNA sequence encoding the Pep1 signal sequence (SEQ ID NO:14), which is operably linked to a downstream DNA sequence (BASI) encoding the BASI protein (SEQ ID NO:5), which is operably linked to a downstream gene CDS (LegGm1b) encoding the leghemoglobin protein (SEQ ID NO:1).

[0134] As used herein, the terms "polypeptide" and "protein" (and / or their respective plurals) are used interchangeably to refer to polymers of any length comprising amino acid residues linked by peptide bonds. Conventional one-letter or three-letter codes for amino acid residues are used herein. The polymers may be linear or branched, may comprise modified amino acids, and may be interrupted by non-amino acids. The terms also encompass amino acid polymers that are modified, either naturally or by intervention, such as, for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification, such as, for example, conjugation with a labeling component. Also included within this definition are, for example, polypeptides containing one or more analogs of an amino acid (including, for example, unnatural amino acids), as well as other modifications known in the art.

[0135] As further described in Sections II-III and the Examples below, in certain embodiments, one or more globin proteins of interest may be designed and constructed as fusion proteins. For example, in certain non-limiting aspects, a globin fusion protein may comprise an N-terminal fusion of one or more amino acid residues (an "N-fusion") and / or a C-terminal fusion of one or more amino acid residues (a "C-fusion"). For example, certain globin fusion proteins having N-terminal and / or C-terminal fusions have been designed, constructed, and described herein as generally set forth in Tables 1-2 of the Examples.

[0136] As used herein, the term "derived polypeptide / protein" refers to a protein that is derived from or can be obtained from a protein by the addition of one or more amino acids to either or both of the N-terminus and C-terminus, the substitution of one or more amino acids at one or several different sites in the amino acid sequence, the deletion of one or more amino acids at one or both termini of the protein or at one or more sites in the amino acid sequence, and / or the insertion of one or more amino acids at one or more sites in the amino acid sequence. Preparation of a protein derivative can be achieved by modifying a DNA sequence encoding the native protein, transforming the DNA sequence into a suitable host, and expressing the modified DNA sequence to form the derived protein.

[0137] Related (and derived) proteins include "variant proteins." Variant proteins differ from a reference / parent protein (e.g., a wild-type protein) by substitution, deletion, and / or insertion of a small number of amino acid residues. The number of different amino acid residues between a variant protein and a parent protein can be one or more, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, or more amino acid residues. A variant protein can share at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%, or more, amino acid sequence identity with the reference protein. Variant proteins may also differ from the reference protein in selected motifs, domains, epitopes, conserved regions, and the like.

[0138] As used herein, the term "analogous sequence" refers to a sequence within a protein that provides a similar function, tertiary structure, and / or conserved residues to a protein of interest (i.e., typically the original protein of interest). For example, in epitope regions containing an α-helical or β-sheet structure, the substituted amino acids in the analogous sequence preferably maintain the same specific structure. The term also refers to nucleotide and amino acid sequences. In some embodiments, analogous sequences are developed such that the amino acid replacement results in a variant enzyme that exhibits similar or improved function. In some embodiments, the tertiary structure and / or conserved amino acid residues within the protein of interest are located in or near the segment or fragment of interest. Thus, if the segment or fragment of interest contains, for example, an α-helical or β-sheet structure, the substituted amino acids preferably maintain that specific structure.

[0139] As used herein, the term "homologous protein" refers to a protein that has a similar activity and / or structure to a reference protein. It is not intended that homologs are necessarily evolutionarily related. Thus, the term is intended to encompass identical, similar, or corresponding proteins (i.e., with respect to structure and function) obtained from different organisms. In some embodiments, it is desirable to identify homologs that have similar quaternary, tertiary, and / or primary structure to the reference protein.

[0140] The degree of homology between sequences may be determined using any suitable method known in the art (e.g., programs such as GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics software package (Genetics Computer Group, Madison, Wis.)). As shown and described in Section II below, other globin gene / protein homologs may be identified by reference to one or more exemplary globin proteins that are well suited for production in one or more modified filamentous fungal strains of the present disclosure.

[0141] For example, PILEUP is a useful program for determining sequence homology levels. PILEUP creates a multiple sequence alignment from a group of related sequences using progressive pairwise alignments. It can also plot a tree showing the clustering relationships used to create the alignment. PILEUP uses a simplified version of the progressive alignment method of Feng and Doolittle (1987). Useful PILEUP parameters include a default gap weight of 3.00, a default gap length weight of 0.10, and weighted end gaps. Another example of a useful algorithm is the BLAST algorithm. One particularly useful BLAST program is the WU-BLAST-2 program. The parameters "W," "T," and "X" determine the sensitivity and speed of the alignment. The BLAST program uses defaults of word length (W) of 11, BLOSUM62 scoring matrix alignment (B) of 50, expectation (E) of 10, M'5, N'-4, and comparison of both strands.

[0142] As used herein, the phrases "substantially similar" and "substantially identical," in the context of at least two nucleic acids or polypeptides, generally mean that the polynucleotides or polypeptides contain sequences having at least about 40% to 100% sequence identity. Thus, in one or more embodiments, a substantially similar or substantially identical nucleic acid or polypeptide of the present disclosure contains at least about 40%, 50%, 60%, 70%, 80%, 90%, or 100% identity to one or more sequences described herein. In certain related embodiments, one or more nucleic acid sequences and / or one or more protein sequences of the present disclosure contain at least about 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 50%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, %, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity. Sequence identity can be determined using known programs such as BLAST, ALIGN, and CLUSTAL using standard parameters. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information. Databases can also be searched using FASTA. One indication that two polypeptides are substantially identical is that the first polypeptide and the second polypeptide are immunologically cross-reactive.Usually, polypeptides that differ by conservative amino acid substitution are immunologically cross-reactive.Therefore, a polypeptide is substantially identical to a second polypeptide, for example, when the two peptides are only different by conservative substitution.Another indication that two nucleic acid sequences are substantially identical is that the two molecules hybridize to each other under stringent conditions (eg, within moderate to high stringency).

[0143] As used herein, "nucleic acid" refers to nucleotide or polynucleotide sequences, and fragments or portions thereof, as well as DNA, cDNA, and RNA of genomic or synthetic origin, which may be double-stranded or single-stranded, whether representing the sense or antisense strand.

[0144] As used herein, the term "expression" refers to the transcription and stable accumulation of sense (mRNA) or antisense RNA derived from a nucleic acid molecule of the present disclosure. Expression may also refer to the translation of mRNA into a polypeptide. Thus, the term "expression" includes any step involved in producing a polypeptide, including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification, and secretion.

[0145] As used herein, the terms "modification" and "genetic modification" are used interchangeably and include, but are not limited to: (a) the introduction, substitution, or removal of one or more nucleotides within a gene, or the introduction, substitution, or removal of one or more nucleotides within a regulatory element required for the transcription or translation of a gene; (b) gene disruption; (c) gene conversion; (d) gene deletion; (e) gene downregulation (e.g., antisense RNA, siRNA, miRNA, etc.); (f) directed mutagenesis (including, but not limited to, CRISPR / Cas9-mediated mutagenesis); and / or (g) random mutagenesis of any one or more genes disclosed herein.

[0146] As used herein, "the introduction, substitution, or removal of one or more nucleotides in a gene encoding a protein" refers to such genetic modification including the coding sequences (i.e., exons) and non-coding intervening (intron) sequences of the gene.

[0147] As used herein, "gene disruption," "gene disruption," "gene inactivation," and "gene inactivation" are used interchangeably and broadly refer to any genetic modification that substantially disrupts / inactivates a target gene. Exemplary gene disruption methods include, but are not limited to, complete or partial elimination of any portion of a gene, including a polypeptide coding sequence (CDS), promoter, enhancer, or other regulatory element, or mutagenesis of the gene, where mutagenesis encompasses substitutions, insertions, deletions, inversions, and any combinations and variations thereof that disrupt / inactivate the target gene and substantially reduce or prevent expression / production of a functional gene product. In certain embodiments of the present disclosure, such gene disruption prevents a host cell from expressing / producing the encoded lov gene product.

[0148] In other embodiments, a protein of interest (e.g., a globin POI) expressed / produced by a fungal cell of the present disclosure may be detected, measured, assayed, etc. by protein quantification methods, gene transcription methods, mRNA translation methods, etc., including, but not limited to, protein migration / mobility (SDS-PAGE), mass spectrometry, HPLC, size exclusion, ultracentrifugation sedimentation velocity analysis, transcriptomics, proteomics, fluorescent tags, epitope tags, fluorescent protein (e.g., GFP, RFP) chimeras / hybrids, etc.

[0149] As used herein, functionally and / or structurally similar proteins are considered to be "related proteins." Such related proteins may be from organisms of different genera and / or species, or from different classes of organisms (e.g., bacteria and fungi). Related proteins also encompass homologs and / or orthologs, as determined by primary sequence analysis, by secondary or tertiary structure analysis, or by immunological cross-reactivity.

[0150] The term "promoter," as used herein, refers to a nucleic acid sequence capable of controlling the expression of a coding sequence (CDS) or functional RNA. Generally, the coding sequence (CDS) is located downstream (3') of the promoter (pro) sequence. Promoters may be derived entirely from a native gene, be composed of different elements from various naturally occurring promoters, or contain synthetic nucleic acid segments. Those skilled in the art will understand that different promoters can direct the expression of a gene in different cell types, at different developmental stages, or in response to different environmental or physiological conditions. Promoters that most frequently cause gene expression in most cell types are generally referred to as "constitutive promoters." Furthermore, it is recognized that because the exact boundaries of regulatory sequences in most cases have not been completely defined, DNA fragments of different lengths may have identical promoter activity.

[0151] As defined herein, the term "introducing," when used in phrases such as "introducing at least one polynucleotide open reading frame (ORF), or gene thereof, or vector thereof, into a fungal cell, includes methods known in the art for introducing polynucleotides into cells, including, but not limited to, protoplast fusion, natural or artificial transformation (e.g., calcium chloride, electroporation), transduction, and transfection.

[0152] As used herein, "transformed" or "transformation" means that a cell has been transformed by the use of recombinant DNA techniques. Transformation typically occurs by inserting one or more nucleotide sequences (e.g., polynucleotides, ORFs, or genes) into the cell. The inserted nucleotide sequences may be heterologous nucleotide sequences (i.e., sequences that do not naturally occur in the cell to be transformed).

[0153] As used herein, "transformation" refers to the introduction of foreign DNA into a host cell such that the DNA is maintained as a chromosomal integrant or a self-replicating extrachromosomal vector. As used herein, "transforming DNA," "transforming sequence," and "DNA construct" refer to DNA used to introduce a sequence into a host cell. The DNA may be generated in vitro by PCR or any other suitable technique. In some embodiments, the transforming DNA includes the incoming sequence, while in other embodiments, the transforming DNA further includes the incoming sequence flanked by homology boxes. In yet other embodiments, the transforming DNA includes other non-homologous sequences (i.e., stuffer sequences or flanking sequences) added to the ends. The ends can be closed to form a closed circle of the transforming DNA, for example, for insertion into a vector.

[0154] As used herein, "incoming sequence" refers to a DNA sequence that is introduced into a fungal cell chromosome. In some embodiments, the incoming sequence is part of a DNA construct. In other embodiments, the incoming sequence encodes one or more proteins of interest. In some embodiments, the incoming sequence comprises a sequence that may or may not already be present in the genome of the cell to be transformed (i.e., it may be a homologous or heterologous sequence). In some embodiments, the incoming sequence encodes one or more proteins of interest, genes, and / or mutant or modified genes. In alternative embodiments, the incoming sequence encodes a functional wild-type gene or operon, a functional mutant gene or operon, or a non-functional gene or operon. In some embodiments, the incoming sequence is a non-functional sequence that is inserted into a gene to disrupt the function of the gene. In another embodiment, the incoming sequence comprises a selectable marker. In a further embodiment, the incoming sequence comprises two homology boxes.

[0155] As used herein, a "homology box" refers to a nucleic acid sequence that is homologous to a sequence within a fungal cell chromosome. More specifically, a homology box is an upstream or downstream region that shares about 80-100% sequence identity, about 90-100% sequence identity, or about 95-100% sequence identity with the immediately adjacent coding region of a gene or portion of a gene to be deleted, disrupted, inactivated, downregulated, etc., according to the present invention. These sequences direct the location of integration of the DNA construct within the fungal cell chromosome and direct which portion of the fungal cell chromosome will be replaced by the incoming sequence. While not intended to limit the present disclosure, a homology box can include from about 1 base pair (bp) to 200 kilobases (kb). Preferably, a homology box includes from about 1 bp to 10.0 kb; 1 bp to 5.0 kb; 1 bp to 2.5 kb; 1 bp to 1.0 kb; and 0.25 kb to 2.5 kb. The homology box can also include approximately 10.0 kb, 5.0 kb, 2.5 kb, 2.0 kb, 1.5 kb, 1.0 kb, 0.5 kb, 0.25 kb, and 0.1 kb. In some embodiments, the 5' and 3' ends of the selectable marker are flanked by homology boxes, wherein the homology boxes comprise nucleic acid sequences that immediately flank the coding region of the gene.

[0156] As used herein, the term "nucleotide sequence encoding a selectable marker" refers to a nucleotide sequence expressible in a host cell, where expression of the selectable marker confers on cells containing the expressed gene the ability to grow in the presence of a corresponding selection agent or in the absence of an essential nutrient.

[0157] As used herein, the terms "selectable marker" and "selection marker" refer to a nucleic acid (e.g., a gene) that can be expressed in a host cell, facilitating the selection of hosts containing a vector. Examples of such selectable markers include, but are not limited to, antimicrobial agents. Thus, the term "selectable marker" refers to a gene that indicates that a host cell has taken up incoming DNA of interest or that some other reaction has occurred. Typically, a selectable marker is a gene that confers antimicrobial resistance or a metabolic advantage to the host cell, allowing cells containing foreign DNA to be distinguished from cells that have not received any foreign sequences during transformation.

[0158] As defined herein, a host cell "genome," a fungal cell "genome," or a filamentous fungal cell "genome" includes chromosomal genes and extrachromosomal genes.

[0159] As used herein, the terms "plasmid," "vector," and "cassette" refer to extrachromosomal elements that often carry genes that are not typically part of the cell's central metabolism and are usually in the form of circular double-stranded DNA molecules. Such elements can be linear or circular, single- or double-stranded, self-replicating sequences of DNA or RNA, genome-integrating sequences, phage, or nucleotide sequences from any source in which multiple nucleotide sequences have been joined or recombined into a unique structure that can introduce into a cell a promoter fragment and DNA sequence for a selected gene product, along with appropriate 3' untranslated sequences.

[0160] As used herein, the term "vector" refers to any nucleic acid that can replicate (propagate) within a cell and carry a new gene or DNA segment (e.g., an "incoming sequence") into the cell. Thus, the term refers to a nucleic acid construct designed for transport between various host cells. Vectors include viruses, bacteriophages, proviruses, plasmids, phagemids, transposons, and artificial chromosomes, such as YACs (yeast artificial chromosomes), BACs (bacterial artificial chromosomes), and PLACs (plant artificial chromosomes), which are either "episomal" (i.e., autonomously replicating) or can be integrated into a host cell chromosome.

[0161] As used herein, "transformation cassette" refers to a particular vector that contains a gene and has elements in addition to that gene that facilitate transformation of a particular host cell.

[0162] As used herein, "expression vector" refers to a vector capable of incorporating and expressing heterologous DNA in a cell. Many prokaryotic and eukaryotic expression vectors are commercially available and known to those skilled in the art. The selection of an appropriate expression vector is within the knowledge of one skilled in the art.

[0163] As used herein, the term "expression cassette" refers to a nucleic acid construct, produced recombinantly or synthetically, with a series of designated nucleic acid elements that allow for transcription of a specific nucleic acid in a target cell (e.g., a vector or vector element as described above). Recombinant expression cassettes can be incorporated into a plasmid, chromosome, mitochondrial DNA, plastid DNA, virus, or nucleic acid fragment. Typically, the recombinant expression cassette portion of an expression vector includes, among other sequences, a nucleic acid sequence to be transcribed and a promoter. In some embodiments, the DNA construct also includes a series of specified nucleic acid elements that allow for transcription of a specific nucleic acid in a target cell. In certain embodiments, the DNA construct of the present disclosure includes a selectable marker and an inactivated chromosomal segment or gene segment or DNA segment, as defined herein.

[0164] As used herein, a "targeting vector" is a vector that contains a polynucleotide sequence homologous to a region in a chromosome of a host cell into which the targeting vector is transformed and that can drive homologous recombination at that region. For example, a targeting vector is used to introduce a genetic modification into a chromosome of a host cell via homologous recombination. In some embodiments, the targeting vector contains other non-homologous sequences (i.e., stuffer sequences or flanking sequences), for example, added to the ends. The ends can be closed so that the targeting vector forms a closed circle, for example, for insertion into a vector.

[0165] As used herein, the terms "purified," "isolated," or "enriched" refer to a biomolecule (e.g., a polypeptide or polynucleotide) that has been altered from its native state by separation from some or all of the naturally occurring components with which it is naturally associated. Such isolation or purification can be accomplished by art-recognized separation techniques, such as ion exchange chromatography, affinity chromatography, hydrophobic separation, dialysis, protease treatment, heat treatment, ammonium sulfate precipitation or other protein salting-out, crystallization, centrifugation, size exclusion chromatography, filtration, microfiltration, gel electrophoresis, or gradient separation, to remove whole cells, cell debris, impurities, extraneous proteins, or enzymes that are not desired in the final composition. Further components that impart additional benefits, such as activators, anti-inhibitors, desired ions, pH-adjusting compounds, or other enzymes or chemicals, can then be added to the purified or isolated biomolecule composition.

[0166] As used herein, a "protein preparation" is any material, typically a generally aqueous solution, that contains one or more proteins.

[0167] As used herein, the terms "broth," "culture broth," "fermentation broth," and / or "whole fermentation broth" may be used interchangeably and refer to a preparation produced by cell fermentation that does not undergo processing steps after fermentation is complete. For example, whole fermentation broth is typically produced when a microbial culture is grown to saturation and incubated under carbon-limited conditions that allow protein synthesis (e.g., expression of proteins by the host cells; and optionally, secretion of proteins into the cell culture medium). Typically, whole fermentation broth is unfractionated and includes spent cell culture medium, metabolic products, extracellular polypeptides, and microbial cells.

[0168] As used herein, the phrase "treated broth" refers to broth that has been conditioned by making changes to the chemical composition and / or physical properties of the broth. Broth "conditioning" can include one or more treatments such as cell lysis, pH modification, heating, cooling, addition of chemicals (e.g., calcium, salts, flocculants, reducing agents, enzyme activators, enzyme inhibitors, and / or surfactants), mixing, and / or holding the broth for a set period (e.g., 0.5 to 200 hours) without further processing.

[0169] As used herein, the "cell lysis" process includes any cell lysis technique known in the art, including, but not limited to, enzymatic treatment (e.g., lysozyme, proteinase K treatment), chemical means (e.g., ionic liquids), physical means (e.g., French press, ultrasound), simply maintaining the culture without feed, etc.

[0170] The terms "recovered," "recovered," and "recovering," as used herein, refer to at least partial separation of the protein from one or more components of the microbial broth and / or from one or more solvents (e.g., water or ethanol) in the broth.

[0171] In certain embodiments, the broth in which host cells are fermented for the production of globin proteins, with or without broth treatment, is clarified. As used herein, "clarified" broth refers to a broth that has been subjected to at least one clarification process to remove cellular debris and / or other insoluble components. Clarification processes, as understood in the art, include, but are not limited to, centrifugation techniques, cross-flow membrane filtration techniques, solid / liquid filtration techniques, and the like.

[0172] "Cell debris" refers to cell walls and other insoluble components that are released or formed after disruption of the cell membrane (eg, after undergoing a cell lysis process).

[0173] In certain embodiments, separation of the solvent, as understood in the art, includes, but is not limited to, ultrafiltration, evaporation, spray drying, freeze drying, etc. The resulting solution is referred to as a "clarified broth concentrate," "UF concentrate," or "ultrafiltration concentrate."

[0174] As used herein, the term "cell mass" refers to the cellular components (including intact and lysed cells) present in a liquid (submerged) culture. Cell mass can be expressed as dry cell weight (DCW) or wet cell weight (WCW).

[0175] II. Recombinant filamentous fungal cells producing heterologous globins As briefly described above, despite current knowledge regarding globin proteins (e.g., leghemoglobin, cyanoglobin, hemoglobin, myoglobin, leghemoglobin, etc.), the isolation and / or production of globin proteins, and their various uses and applications, there remain ongoing unmet needs in the art. For example, due to the growing demand for so-called "non-animal" derived meat substitutes, certain legume (plant) derived globins (i.e., leghemoglobin) have been proposed for use. In one particular example, the isolation of leghemoglobin for use as a meat substitute has been described, wherein leghemoglobin protein was isolated and purified from a plant (legume) source (PCT Publication No. WO 2013 / 010042). In other cases, expression of leghemoglobin protein in yeast cells (e.g., S. cerevisiae; P. pastoris) has been attempted, generally requiring extensive genetic modification of the host cell to co-express the entire heme biosynthetic pathway and leghemoglobin protein (PCT Publication No. WO 2016 / 183163).

[0176] As described in Krainer et al. (2015), insufficient heme incorporation is considered a central obstacle in the recombinant production of active hemoproteins. In particular, Krainer et al. (2015) investigated the effects of both pathway engineering and medium supplementation to optimize the recombinant production of the hemoprotein horseradish peroxidase (HRP) in the methylotrophic yeast P. patoris. As summarized by Krainer et al. (2015), in contrast to studies with other host organisms, (a) co-overexpression of genes from the endogenous heme biosynthetic pathway in P. patoris did not improve recombinant production of active heme (HRP) enzymes, (b) medium supplementation with the commonly used precursor 5-aminolevulinic acid (ALA) did not affect the yield of active heme (HRP) enzymes, but (c) medium supplementation with hemin increased the yield of active heme (HRP) enzymes, leading them to conclude that the yield of active peroxidase enzymes from P. patoris can be easily enhanced by supplementing the culture medium with hemin.

[0177] More recently, Shao et al. (2022) described a P. pastoris strain capable of high-yield secretory production of functional leghemoglobin, developed through gene dosage optimization and heme pathway enhancement. In particular, the heme biosynthesis pathway was engineered by increasing the copy number of heterologous leghemoglobin and enhancing the native heme biosynthetic pathway to address challenges in heme depletion and leghemoglobin secretion. This P. pastoris strain engineering strategy increased leghemoglobin secretion without the need to supplement the medium with expensive precursors (e.g., hemin; Shao et al., 2022).

[0178] In other cases, non-animal-derived meat-like materials / ingredients obtained from genetically modified cyanobacteria (blue-green algae) containing polynucleotides encoding heterologous heme-containing proteins (e.g., leghemoglobin, cyanoglobin) have been described (PCT Publication WO 2019 / 079135). As generally described in WO 2019 / 079135, certain genetic modifications of the cyanobacterial host strain were required to improve the levels of globin proteins, such as genetic modifications to reduce (lower) the levels of heme oxygenase in the host strain, genetic modifications to knock down the gene encoding magnesium chelatase (MgCh) in the host strain, genetic modifications to overexpress ferrochelatase (FeCh) in the host strain, and genetic modifications to knock down the levels of gun4 (an MgCh activator) in the host strain.

[0179] In one particular case, the heme biosynthetic pathway of Aspergillus niger strains was evaluated in the context of producing ligninolytic peroxidases (i.e., heme-containing class II peroxidases; Franken et al., 2011). As summarized by Franken et al. (2011), cofactor availability and incorporation have been shown to be limiting factors in the production of fungal peroxidases in A. niger, which require heme as a cofactor. Notably, Franken et al. (2011) noted that peroxidase production can be increased by supplementing the fermentation medium with hemoglobin or hemin, but the mechanism behind heme incorporation is poorly understood, making this approach too expensive to be suitable for industrial purposes.

[0180] Thus, as generally described above, it is apparent that ongoing uncertainties and inconsistencies remain in the art, particularly with respect to identifying optimal host cell genetic modifications necessary for enhanced production of globin proteins, the need or requirement for media (broth) supplements (e.g., hemin, ALA) for enhanced production of globin proteins, and the like. For example, as demonstrated in certain yeast strains (e.g., S. cerevisiae; P. pastoris), the need or requirement for upregulation (e.g., overexpression) and / or downregulation of one or more heme pathway biosynthetic genes varies depending on the strain of yeast selected and the particular globin protein expressed thereby. Similarly, in other aspects, the need or requirement for media supplementation can vary significantly depending on the strain of yeast being fermented for the globin protein expressed thereby. As reviewed in Franken et al. (2011), the use of Aspergillus sp. fungal strains for the production of ligninolytic peroxidases requires media supplementation (e.g., hemin), suggesting that a better understanding of the heme biosynthetic pathway and its regulation in A. niger is needed for further strain improvement, most preferably without the need for media supplementation.

[0181] Therefore, it remains generally unknown whether filamentous fungal strains such as Aspergillus can produce (heme-containing) globin proteins at sufficiently high levels desired in the art. In particular, the design, construction, identification, culture / fermentation, etc. of genetically engineered (recombinant) host organisms and methods thereof that contain enhanced globin-producing phenotypes (e.g., protein titer, specific productivity, protein yield, volumetric productivity, carbon conversion efficiency, etc.) are important economic factors in the cost of globin protein production, especially under large-scale (industrial) fermentation conditions.

[0182] Based on the above, Applicant has surprisingly observed that recombinant filamentous fungal cells are particularly suitable for the large-scale production of heterologous globin proteins. More specifically, as described in the Examples section below, Applicant has designed, constructed, evaluated, etc., recombinant polynucleotides (e.g., expression cassettes) encoding heterologous globin proteins, and the cassettes were introduced into filamentous fungal cells for the expression and secretion of the globin proteins. In particular, as described in Example 1, Applicant has constructed vectors for the expression of soybean leghemoglobin (SEQ ID NO: 1) as a secreted protein (Example 1A), for the expression of kidney bean leghemoglobin (SEQ ID NO: 19) as a secreted protein (Example 1B), and for the expression of soybean leghemoglobin (SEQ ID NO: 1) as a secreted fusion protein (Example 1C). As presented and described in Example 2, Applicant has further constructed a vector for the expression of a heterologous bovine myoglobin protein (SEQ ID NO: 3) in filamentous fungal cells. For example, various nucleic acid (DNA) sequences and genetic elements used for the construction of one or more modified (recombinant) filamentous fungal strains of the present disclosure are set forth in the Examples section below in Table 1. Similarly, certain chromosomal integration vectors for expression of leghemoglobin, myoglobin, and protease inhibitor protein (BASI; SEQ ID NO: 5) are set forth in Table 2 of the Examples section.

[0183] Applicant has further designed and constructed recombinant filamentous fungal cells through targeted integration of one or more expression cassettes, as described in Example 3. In particular, leghemoglobin- and myoglobin-expressing filamentous fungal strains were generated by Cas9-induced targeted genomic integration (Example 3). Additionally, to assess possible proteolysis of secreted globin proteins (e.g., hydrolytic degradation via natural proteases secreted by the host strain), an expression cassette encoding an exemplary secreted protease inhibitor protein (BASI) was integrated into certain strains for coexpression of a globin protein (e.g., leghemoglobin) and a protease inhibitor (e.g., BASI).

[0184] Fed-batch fermentation was carried out in a two liter (2 L) bioreactor for the filamentous fungal strain BFZ28 (Pcbh1-CBH1core-KEX2-LegGm1b) as generally shown and described in Example 4, with ten milliliter (10 mL) whole broth samples taken every 24 hours and frozen at -20°C. More specifically, as shown in Figure 3, red (dark) color formation in the supernatant indicates secretion of globin protein into the culture supernatant. Notably, total protein secretion of BFZ28 is shown in Figure 4, with a total protein concentration of approximately 50 grams per liter (50 g / L) after approximately 188 hours of fermentation. As shown in Figure 5, supernatant from strain BFZ28 was analyzed by SDS-PAGE; the upper protein band at approximately 50 kDa is the Cbh1 core protein, and the leghemoglobin band (approximately 16 kDa) co-migrates with the BASI protease inhibitor at approximately 20 kDa, or is present as a minor band at approximately 17 kDa. Additionally, to verify / confirm leghemoglobin expression, fermentation supernatant was analyzed by HPLC analysis of 188-hour supernatant samples from BFZ28 fermentation runs. As shown in Figure 6, the porphyrin (heme) prosthetic group of the leghemoglobin protein was detected at a wavelength of 410 nm, which co-migrated with the protein peak detected at 280 nm with a retention time of 6.4 minutes. The inset in Figure 6 shows a spectral scan of the leghemoglobin peak at a retention time of 6.4 minutes, confirming the peak absorbance of leghemoglobin protein at 410 nm.

[0185] As briefly described above, to assess the role of proteolysis on secreted globin proteins, an expression cassette encoding an exemplary secreted protease inhibitor (BASI) was incorporated into certain strains for co-expression of the globin protein and the (BASI) protease inhibitor. More specifically, engineered strains "BGJ74" (Pcbh1-LegGm1b, Pcbh2-BASI), "BGJ75" (Pcbh1-Pv1b, Pcbh2-BASI), and "BGJ76" (Pcbh1-LegGm1b.Pep1ss, Pcbh2-BASI) were constructed and fermented as generally described in Example 5. For example, as shown in FIG. 7, total soluble protein secreted in a fermentation run was plotted against effective fermentation time (EFT, hours), along with data from a fermentation run of BFZ28 (Example 4). Fermentation supernatants were analyzed by SDS-PAGE, as shown in Figure 8, and the presence of (heme-containing) leghemoglobin was confirmed via HPLC, as shown in Figure 9. As described in Example 5, the total protein secretion titers at the end of the 188-hour fermentation run are shown in Table 4, and leghemoglobin titers (g / L) were calculated based on the protein peak area at 280 nm. For example, recombinant filamentous fungal strains expressing soybean leghemoglobin (BFZ28, BGJ74, BGJ76) and kidney bean leghemoglobin (BGJ75) were able to secrete leghemoglobin at high titers (g / L) during fermentation, as shown in Table 4.

[0186] Accordingly, certain embodiments of the present disclosure relate to recombinant (modified) filamentous fungal cells capable of producing heterologous globin proteins for use in food materials, food ingredients, flavor modifiers, aroma modifiers, and the like. In certain embodiments, the globin protein is selected from the group consisting of leghemoglobin, myoglobin, hemoglobin, cyanoglobin, and non-symbiotic hemoglobin. In certain other embodiments, the globin protein comprises the group consisting of leghemoglobin, hemoglobin, non-symbiotic hemoglobin, myoglobin, neuroglobin, cytoglobin, protoglobin, truncated 2 / 2 globin, HbN, cyanoglobin, HbO, Glb3, and Hell's gate globin, bacterial hemoglobin, and ciliary myoglobin. For example, as is generally known in the art, members of the globin-like superfamily include a wide variety of all-helical proteins that bind porphyrins and play a variety of roles in all three kingdoms of life, including as sensors or transporters of oxygen. The globin-like superfamily includes the M / myoglobin-like, S / sensor globin, and T / truncated globin (TrHb) families, and the phycobiliproteins (PBPs).

[0187] Thus, in certain other one or more embodiments, the present disclosure provides methods for identifying and obtaining suitable globin (DNA / protein) sequences for expression in one or more modified filamentous fungal cells described herein. In certain aspects, one or more suitable leghemoglobin sequences can be identified and obtained from various plant sources, such as, but not limited to, various legume species and varieties thereof, including soybean, broad bean, lima bean, cowpea, pea, yellow pea, lupin, sand pea, chickpea, peanut, alfalfa, Vicia faba, clover, bush clover, and spotted bean. In certain embodiments, the leghemoglobin protein comprises at least about 90% identity to a leghemoglobin protein derived from a plant selected from the group consisting of soybean, broad bean, lima bean, cowpea, pea, common bean, yellow pea, lupin, sand bean, chickpea, peanut, alfalfa, Vicia faba, clover, bush clover, and spotted bean. In certain other embodiments, the leghemoglobin protein comprises at least about 90% identity to soybean leghemoglobin protein SEQ ID NO: 1. In certain other embodiments, the leghemoglobin protein comprises at least about 90% identity to common bean leghemoglobin protein SEQ ID NO: 19. In certain other embodiments, the myoglobin protein comprises at least about 90% identity to SEQ ID NO: 3. In one or more other embodiments, the recombinant filamentous fungal cells of the present disclosure comprise an introduced polynucleotide (expression cassette) encoding a protease inhibitor protein.

[0188] For example, the native soybean leghemoglobin protein (SEQ ID NO: 1; Figure 1) contains 145 amino acid residues, with amino acid positions 8-111 containing a globin-like superfamily domain. In addition, the native soybean leghemoglobin protein (Figure 1) contains a class 1-2 non-symbiotic hemoglobin domain at residue positions 4-144 of SEQ ID NO: 1. The leghemoglobin protein heme-binding site includes amino acid positions L44, F45, S46, F47, K58, H62, K65, L66, F67, L69, V70, A88, L89, I92, H93, K96, I98, Q102, F103, Y134, L137, A138, and I141 (Figure 1; SEQ ID NO: 1, shown as bolded residues).

[0189] Thus, in certain embodiments, the one or more leghemoglobin protein sequences are selected from the group consisting of Jequirity Bean (Abrus precatorius) (Inventory No. XP_027365674.1), Astragalus canadensis (Astragalus canadensis) (Inventory No. QAX32739.1), Astragalus sinicus (Astragalus sinicus) (Inventory No. ABB13622.1), Pigeonpea (Cajanus cajan) (Inventory No. XP_020222796.1), Sea Buckthorn (Canavalia lineata) (Inventory No. P42511.1), Chickpea (Cicer arietinum) (Inventory No. XP_004490880.1), Galega (Galega orientalis) (Inventory No. QAX32752.1), Wild Soybean (Glycine soja) (Inventory No. XP_004490880.1), and the like. soja (Inventory No. KAG4983849.1), Ural licorice (Glycyrrhiza uralensis) (Inventory No. QAX32708.1), Lotus japonicus (Inventory No. AFK42883.1), Medicago sativa (Inventory No. AAA32657.1), Medicago truncatula (Inventory No. XP_003616494.1), Mucuna pruriens (Inventory No. RDX62803.1), Sainfoin (Onobrychis viciifolia) (Inventory No. QAX32720.1), Ononis spinosa (Inventory No. QAX32757.1), Phaseolus vulgaris (Phaseolus vulgaris (Inventory No. XP_007144265.1), pea (Pisum sativum) (Inventory No. XP_050917485.1), winged bean (Psophocarpus tetragonolobus) (Inventory No. P27199.1), sesbania (Sesbania rostrata) (Inventory No. P14848.2), red clover (Trifolium pratense) (Inventory No. XP_045786945.1), desert clover (Trifolium subterraneum) (Inventory No. GAU42435.1), faba bean (Vicia faba) (Inventory No.P93849.3), azuki bean (Vigna angularis) (Inventory No. XP_017409386.1), mung bean (Vigna radiata var. radiata) (Inventory No. XP_014513332.1), bamboo bean (Vigna umbellate) (Inventory No. XP_047166207.1), and cowpea (Vigna unguiculata) (Inventory No. NP_001363034.1). In certain embodiments, the leghemoglobin protein will comprise an absorbance spectrum and appearance nearly identical to that of myoglobin protein derived from animal muscle.

[0190] As shown in FIG. 1, the native bovine myoglobin protein (SEQ ID NO: 3) contains 154 amino acid residues, with amino acid positions 7 to 112 containing the globin-like superfamily domain. Thus, in certain embodiments, the one or more myoglobin protein sequences are derived from a mammal, such as a giant panda (Ailuropoda melanoleuca) (Inventory No. XP_002925619.1), a bowhead whale (Balaena mysticetus) (Inventory No. R9RZK8.1), a minke whale (Balaenoptera acutorostrata) scamonia (Inventory No. XP_007165766.1), a blue whale (Balaenoptera musculus) (Inventory No. XP_036722853.1), a yak (Bos mutus) (Inventory No. MXQ80090.1), a cow (Bos taurus) (Inventory No. NP_776306.1), a water buffalo (Bubalus bubalis) (Inventory No. XP_006074486.1), Bactrian camel (Camelus ferus) (Inventory No. XP_006183931.1), goat (Capra hircus) (Inventory No. XP_005680660.1), southern white rhinoceros (Ceratotherium simum simum) (Inventory No. XP_004418145.1), red deer (Cervus elaphus) (Inventory No. P02191.2), spotted hyena (Crocuta) (Inventory No. KAF0881531.1), beluga whale (Delphinapterus leucas) (Inventory No. XP_022455612.1), spiny dolphin (Delphinus capensis (Inventory No. AMN15044.1), Alaskan sea otter (Enhydra lutris kenyoni) (Inventory No. XP_022374654.1), donkey (Equus asinus) (Inventory No. XP_014688187.1), horse (Equus caballus) (Inventory No. NP_001157488.1), long-finned pilot whale (Globicephala melas) (Inventory No. XP_030711253.1), gorilla (Gorilla gorilla) (Inventory No. XP_018874109.1), wolverine (Gulo gulo) (Inventory No. VCW91177.1), pygmy hippopotamus (Hexaprotodon liberiensis) (Inventory No. AGM75766.1), striped hyena (Hyaena hyaena) (Inventory No. XP_039086409.1), northern bottlenose whale (Hyperoodon ampullatus) (Inventory No. AGM75769.1), Pacific ruffed dolphin (Indopacetus pacificus) (Inventory No. Q0KIY9.3), Amazon river dolphin (Inia geoffrensis) (Inventory No. P02181.2), Pacific white-sided dolphin (Lagenorhynchus obliquidens) (Inventory No. XP_026964057.1), ring-tailed lemur (Lemur catta) (Inventory No. XP_045408951.1), Weasel lemur (Lepilemur mustelinush) (Inventory No. P02169.2), Siberian river dolphin (Lipotes vexillifer) (Inventory No. XP_007456317.1), Canadian river otter (Lontra canadensis) (Inventory No. XP_032692617.1), Eurasian river otter (Lutra lutra) (Inventory No. P11343.3), Eared pangolin (Manis pentadactyla) (Inventory No. XP_036747296.1), Humpback whale (Megaptera novaeangliae) (Inventory No. P02178.2), European badger (Meles meles (Inventory No. XP_045868724.1), Habbo's beaked whale (Mesoplodon carlhubbsi) (Inventory No. P02183.2), grey mouse lemur (Microcebus murinus) (Inventory No. XP_012642551.1), narwhal (Monodon monoceros) (Inventory No. XP_029061405.1), finless porpoise (Neophocaena asiaeorientalis) (Inventory No. XP_024599230.1), killer whale (Orcinus orca) (Inventory No. XP_004286254.1), European rabbit (Oryctolagus cuniculus) (Inventory No. XP_008255312.3), sheep (Ovis aries) (Inventory No. NP_001072126.1), bonobo (Pan paniscus) (Inventory No. XP_008973239.1), Cape warthog (Phacochoerus africanus) (Inventory No. XP_047642382.1), harbor seal (Phoca vitulina) (Inventory No. XP_032272566.1), sperm whale (Physeter catodon) (Inventory No. 5YCG_A), raccoon (Procyon lotor) (Inventory No. AGM75763.1), Cockerell's sifaka (Propithecus coquereli) (Inventory No. XP_012500179.1), spotted dolphin (Stenella attenuate) (Inventory No. Q0KIY6.3), meerkat (Suricata The animal may be derived from organisms including, but not limited to, wild boar (Sus suricatta) (Inventory No. XP_029811662.1), wild boar (Sus scrofa) (Inventory No. NP_999401.1), tree shrew (Tupaia chinensis) (Inventory No. XP_006170382.1), bottlenose dolphin (Tursiops truncatus) (Inventory No. AMN15042.1), polar bear (Ursus maritimus) (Inventory No. NP_001288305.1), and beaked whale (Ziphius cavirostris) (Inventory No. P02182.2).

[0191] As described in Section III below, in one or more other embodiments, the recombinant filamentous fungal cell of the present disclosure comprises an introduced polynucleotide (expression cassette) encoding a protease inhibitor protein. In certain one or more embodiments, the expression cassette comprises or consists of a nucleic acid (DNA) sequence encoding a fusion protein, such as an upstream (5') DNA fusion sequence (N-terminal fusion) and / or a downstream (3') DNA fusion sequence (C-terminal fusion). In other embodiments, one or more cassettes are integrated into the genome of the cell. In certain other embodiments, the recombinant filamentous fungal cell comprises at least two introduced expression cassettes encoding the same or different globin proteins.

[0192] In certain other embodiments, the recombinant filamentous fungal cell comprises a deletion or disruption of one or more endogenous genes encoding one or more secreted proteases, including, but not limited to, subtilisin-like serine proteases, aspartic acid proteases, trypsin-like serine proteases, glutamic acid proteases, and aminopeptidases.

[0193] In yet other embodiments, the recombinant filamentous fungal cells of the present disclosure comprise a deletion of one or more highly expressed endogenous genes. For example, in the case of T. reesei cells, highly expressed endogenous genes include, but are not limited to, one or more secreted lignocellulolytic enzymes (e.g., cellobiohydrolase, xylanase, endoglucanase, β-glucosidase). Thus, in certain embodiments, the filamentous fungal cells of the present disclosure are genetically modified to be deficient in the production of one or more highly expressed lignocellulolytic enzymes. For example, in T. reesei cells, suitable genetic modifications include rendering the cells deficient in the production of one or more endogenous genes selected from the cellobiohydrolase 1 (cbh1) gene, the cellobiohydrolase 2 (cbh2) gene, the endoglucanase 1 (egl1) gene, and the endoglucanase 2 (egl2) gene.

[0194] III. Polynucleotide Constructs and Molecular Biology As noted above, certain embodiments of the present disclosure relate to engineered filamentous fungal strains comprising enhanced globin protein production phenotypes (e.g., protein titer, specific productivity, protein yield, volumetric productivity, carbon conversion efficiency, etc.). Accordingly, certain embodiments of the present disclosure relate to, among other things, molecular biology, genetic modifications, polynucleotides, genes, gene coding sequences (CDS), ORFs, vectors, expression cassettes, fusion proteins, protein linker sequences, cleavable protein linker sequences, and the like. In certain embodiments, the present disclosure provides recombinant nucleic acids (polynucleotides) comprising a gene or gene CDS encoding a globin protein. In particular, certain embodiments provide polynucleotide constructs (e.g., expression cassettes) encoding globin proteins for expression and secretion of the globin protein into a culture medium / fermentation broth. In certain aspects, one or more expression cassettes for secretion of a globin protein may generally be represented by one or more schematic diagrams.

[0195] For example, an expression cassette encoding a secreted globin protein may be depicted schematically as 5'-[pro]-[sig-seq]-[globin CDS]-3'; the expression cassette comprises (in the 5' to 3' direction) a promoter (pro) region sequence operably linked to a nucleic acid encoding a protein (signal) secretion sequence (sig-seq), which is operably linked to a nucleic acid encoding a globin protein (globin CDS).

[0196] In certain embodiments, the promoter (pro) region sequence is a strong promoter functional in the host filamentous fungal strain. By way of example, such strong promoter (pro) region sequences functional in Trichoderma sp. filamentous fungal cells include, but are not limited to, the T. reesei CBH1 promoter, CBH2 promoter, Xyn3 promoter, Gla1 promoter, Egl2 promoter, etc. In certain embodiments, a strong promoter may be referred to as a promoter that overexpresses a globin protein.

[0197] While certain protein (signal) secretion sequences are exemplified herein (e.g., Cbh1 secretion sequence (SEQ ID NO: 13), Pep1 secretion sequence (SEQ ID NO: 14), one of skill in the art may screen, identify, and select other suitable protein signal / secretion sequences that are functional in filamentous fungal cells. For example, in certain aspects, globin protein secretion in one or more filamentous fungal cells of the present disclosure may be identified using one or more signal (secretion) peptide sequences from highly secreted filamentous fungal proteins known in the art. In certain other embodiments, globin protein secretion in one or more filamentous fungal cells may be identified using one or more signal (secretion) peptide sequences from Talaromyces sp. β-mannanase secretion sequence, Talaromyces sp. glucoamylase secretion sequence, Trichoderma sp. Cbh2 secretion sequence, Trichoderma sp. glucoamylase secretion sequence, Humicola sp. The secretory / signal sequence may be selected and identified by reference to one or more naturally occurring protein secretory / signal sequences and functional variants thereof, including, but not limited to, the Cel45 secretory sequence, the Neurospora sp. chitin synthase secretory sequence, the Aspergillus sp. alpha-galactosidase (GlaA) secretory sequence, the Aspergillus sp. PepN secretory sequence, the Trichoderma harzianum aspartyl protease (PapA) secretory sequence, the Myceliophthora sp. IMI 387099 xylanase secretory sequence, and the like.

[0198] As further shown in the Examples, one or more expression cassettes encoding secreted globin proteins can comprise or consist of nucleic acid (DNA) sequences encoding fusion proteins, protein linker sequences, cleavable (protein) linker sequences, etc. For example, an expression cassette encoding a secreted globin protein having an N-terminal fusion (N-fusion) may be depicted schematically as 5'-[pro]-[sig-seq]-[N-fusion]-[globin CDS]-3'; the cassette comprises (in the 5' to 3' direction) a promoter (pro) region sequence operably linked to a nucleic acid encoding a protein (signal) secretion sequence (sig-seq) operably linked to a nucleic acid encoding an N-terminal fusion protein (N-fusion), which is operably linked to a nucleic acid encoding a globin protein (globin CDS). Similarly, an expression cassette encoding a secreted globin protein having a C-terminal fusion (C-fusion) may be shown schematically as 5'-[pro]-[sig-seq]-[globin CDS]-[C-fusion]-3'; the expression cassette comprises (in the 5' to 3' direction) a promoter (pro) region sequence operably linked to a nucleic acid encoding a protein (signal) secretion sequence (sig-seq), which is operably linked to a nucleic acid encoding a globin protein (globin CDS), which is operably linked to a nucleic acid encoding the C-terminal fusion protein (C-fusion).

[0199] Thus, in certain other embodiments, an expression cassette encoding a secreted globin protein having an N-terminal fusion (N-fusion) and a C-terminal fusion (C-fusion) may be depicted schematically as 5'-[pro]-[sig-seq]-[N-fusion]-[globin CDS]-[C-fusion]-3'; the expression cassette comprises (in the 5' to 3' direction) a promoter (pro) region sequence operably linked to a nucleic acid encoding a protein (signal) secretion sequence (sig-seq) operably linked to a nucleic acid encoding an N-terminal fusion protein (N-fusion) operably linked to a nucleic acid encoding a globin protein (globin CDS) operably linked to a nucleic acid encoding the C-terminal fusion protein (C-fusion).

[0200] For example, as shown in Tables 1-2 (Example 2), the T. reesei Cbh1 protein is an exemplary N-terminal fusion protein, in which a nucleic acid encoding a leghemoglobin protein (leghemoglobin CDS) is located upstream (5') and includes a nucleic acid encoding a Cbh1 protein (Cbh1) operably linked to the nucleic acid encoding the leghemoglobin protein (e.g., 5'-[pro]-[sig-seq]-[Cbh1]-[leghemoglobin CDS]-3'). In other embodiments, a 6-histidine amino acid (6-His) tag is an exemplary C-terminal fusion protein (peptide), and a nucleic acid encoding a leghemoglobin protein (leghemoglobin CDS) is located downstream (3') and includes a nucleic acid encoding a 6-His peptide (6-His) operably linked to the nucleic acid encoding the leghemoglobin protein (e.g., 5'-[pro]-[sig-seq]-[leghemoglobin CDS]-[6-His]-3'). In certain other embodiments, a Cbh1 protein and a 6-His peptide are exemplary N- and C-terminal proteins fused to a leghemoglobin protein (e.g., 5'-[pro]-[sig-seq]-[Cbh1]-[leghemoglobin CDS]-[6-His]-3').

[0201] In certain other embodiments, one or more cassettes comprise one or more upstream (N-linker) and / or downstream (C-linker) nucleic acids encoding one or more (protein / peptide / amino acid) linker amino acid sequences. In certain embodiments, the protein / peptide / amino acid linker sequence is a cleavable sequence (e.g., a KEX2 cleavable site). As described in Example 4, strain BFZ28 contains an introduced cassette encoding a Cbh1-lectin fusion protein (Pcbh1-CBH1core-KEX2-LegGm1b), with the nucleic acid encoding the cleavable KEX2 linker (KEX2) positioned between the nucleic acid encoding the Cbh1 protein (Cbh1) and the nucleic acid encoding the leghemoglobin protein (leghemoglobin CDS) (e.g., 5'-[pro]-[sig-seq]-[Cbh1]-[KEX2]-[leghemoglobin CDS]-3'). For example, during protein secretion in fungal cells, certain proteins are cleaved by KEX2, a member of the KEX2 or "kexin" family of serine peptidases (EC 3.4.21.61). As described in U.S. Patent Publication Nos. U.S. Patent Application Publication No. 2014 / 0024067 and U.S. Patent No. 8,936,917 (each incorporated herein by reference in its entirety), KEX2 is a highly specific, calcium-dependent endopeptidase that cleaves peptide bonds immediately C-terminal to a pair of basic amino acids ("KEX2 site") in a protein substrate (e.g., globin) during secretion of that (globin) protein. For example, a KEX2 (cleavable) site can be included in the construction of an expression cassette encoding one or more globin proteins, as generally described in U.S. Patent Publication No. U.S. Patent Application Publication No. 2014 / 0024067. Similarly, U.S. Patent No. 8,936,917 describes a modified KEX2 cleavage site with a presequence (VAVE) that improves cleavage efficiency at the post-presequence KEX2 site. Although the KEX2 cleavage site is exemplified, other protease cleavage sites functional in filamentous fungal cells can be used for cleavage of the peptide linker between the protein fusion partner and the globin protein.Examples of additional protease-cleavable linkers include, but are not limited to, STE13 described in La Maquer et al. (2019) and the self-cleaving 2A peptide described in Subramanian et al. (2017).

[0202] In other embodiments, the one or more cassettes encoding globin proteins are operably linked and include a terminator region sequence (term) located at the 3' end.

[0203] In certain embodiments, polynucleotides of the present disclosure may comprise one or more selectable markers. Selectable markers for use in filamentous fungi include, but are not limited to, alsl, amdS, hygR, pyr2, pyr4, pyrG, sucA, bleomycin resistance markers, blasticidin resistance markers, pyrithiamine resistance markers, chlorimuron ethyl resistance markers, neomycin resistance markers, adenine pathway genes, tryptophan pathway genes, thymidine kinase markers, and the like. In certain embodiments, the selectable marker is pyr2, the construction and use of which are generally described in PCT Publication No. WO 2011 / 153449.

[0204] To transform the fungal host cells of the present disclosure, standard techniques for transforming filamentous fungi and culturing fungi, which are well known to those skilled in the art, are used. Thus, introduction of a DNA construct or vector into a fungal host cell can be achieved by transformation, electroporation, nuclear microinjection, transduction, transfection (e.g., lipofection-mediated and DEAE-dextrin-mediated transfection), incubation with calcium phosphate DNA precipitates, high-velocity gun with DNA-coated microprojectiles, gene gun or biolistic transformation, protoplast fusion, and other techniques. General transformation techniques are known in the art.

[0205] In many cases, transformation of Trichoderma sp. fungal cells typically takes place within 10 5 ~10 7 / mL, specifically, about 2x10 6 Permeabilized protoplasts or cells at a density of 100 μL / mL are used. A volume of 100 μL of these protoplasts or cells in an appropriate solution (e.g., 1.2 M sorbitol and 50 mM CaCl2) is mixed with the desired DNA. Generally, a high concentration of polyethylene glycol (PEG) is added to the uptake solution. Additives such as dimethyl sulfoxide, heparin, spermidine, potassium chloride, etc. may also be added to the uptake solution to facilitate transformation. Similar procedures are available for other fungal host cells (see, e.g., U.S. Pat. Nos. 6,022,725 and 6,268,328, both of which are incorporated by reference).

[0206] Thus, the methods and compositions of the present disclosure generally rely on conventional techniques in the field of recombinant genetics. For example, in certain embodiments, a gene encoding a heterologous globin protein of interest is introduced into a filamentous fungal (host) cell. In certain embodiments, the gene (or gene CDS) is cloned into an intermediate vector before being transformed into a filamentous fungal (host) cell for replication and / or expression. These intermediate vectors may be, for example, prokaryotic vectors such as plasmids or shuttle vectors. In certain embodiments, the expression of the gene or gene CDS encoding the globin protein is under the control of a heterologous promoter, which may be a heterologous constitutive promoter or a heterologous inducible promoter, particularly a strong promoter capable of overexpressing the globin protein.

[0207] Expression vectors usually contain a transcription unit or "expression cassette" that contains all additional elements required for expression of a heterologous sequence. For example, a typical expression cassette contains an upstream (5') promoter operably linked to a nucleic acid sequence encoding a protein of interest and may further include a nucleic acid sequence encoding a protein (signal) secretion sequence, nucleic acid sequences required for efficient polyadenylation of the transcript, a ribosome binding site, and a translation termination sequence. Additional elements of the cassette may include an enhancer and, if genomic DNA is used as the structural gene, an intron with functional splice donor and acceptor sites.

[0208] In addition to a promoter sequence, the expression cassette may also contain a transcription termination region downstream of the structural gene to provide efficient termination. The termination region may be obtained from the same gene as the promoter sequence or from a different gene. While any fungal terminator is likely to be functional in the present invention, preferred terminators include those derived from the Trichoderma cbhI gene, the Aspergillus nidulans trpC gene, and the Aspergillus awamori or Aspergillus niger glucoamylase gene.

[0209] The particular expression vector used to transport genetic information into cells is not particularly critical. Any conventional vector used for expression in eukaryotic or prokaryotic cells can be used. Standard bacterial expression vectors include bacteriophage λ and M13, as well as plasmids such as pBR322-based plasmids, pSKF, pET23D, and fusion expression systems such as MBP, GST, and LacZ. Epitope tags, such as c-myc, can also be added to recombinant proteins to provide convenient isolation methods.

[0210] Elements that can be included in the expression vector can also be a replicon, a gene encoding antibiotic resistance to allow for selection of bacteria harboring the recombinant plasmid, or a unique restriction site in a non-essential region of the plasmid to allow for the insertion of heterologous sequences. The particular antibiotic resistance gene selected is not critical, as any of the many resistance genes known in the art can be suitable. Prokaryotic sequences are preferably selected so as not to interfere with DNA replication or integration in the filamentous fungal host.

[0211] The transformation methods of the present invention may result in stable integration of all or part of the transformation vector into the filamentous fungal genome. However, transformation resulting in the maintenance of a self-replicating extrachromosomal transformation vector is also contemplated. Since many standard transfection methods can be used to generate filamentous fungal cell lines expressing large amounts of heterologous proteins, any of the known procedures for introducing foreign nucleotide sequences into fungal host cells may be used. These include calcium phosphate transfection, polybrene, protoplast fusion, electroporation, biolistics, liposomes, microinjection, plasma vectors, viral vectors, and any of the other known methods for introducing cloned genomic DNA, cDNA, synthetic DNA, or other foreign genetic material into host cells. Agrobacterium-mediated transfection methods, such as those described in U.S. Pat. No. 6,255,115, are also useful.

[0212] After the expression vector is introduced into the cells, the transformed cells are cultured under conditions favorable for gene expression. Large batches of transformed cells can be cultured as described herein. Finally, the protein product is recovered from the culture using standard techniques. Thus, the present disclosure provides for enhanced expression and production of a desired protein of interest, as described herein.

[0213] In certain one or more embodiments or aspects of the present disclosure, the filamentous fungal cell (strain) may comprise one or more genetic modifications, including, but not limited to, (a) introduction, substitution, or removal of one or more nucleotides within a gene (gene CDS, or ORF thereof), or introduction, substitution, or removal of one or more nucleotides within a regulatory element required for transcription or translation of a gene (gene CDS or ORF); (b) gene disruption; (c) gene conversion; (d) gene deletion; (e) gene downregulation; (f) directed mutagenesis; and / or (g) random mutagenesis of a gene (gene CDS or ORF thereof).

[0214] As generally indicated herein above and described below, by reference to one or more nucleic acid and / or protein sequences disclosed herein, one skilled in the art can easily perform one or more genetic modifications to construct recombinant / modified / variant filamentous fungal strains. For example, gene deletion techniques allow for the partial or complete removal of a gene, thereby eliminating or reducing protein expression / production and / or eliminating or reducing expression / production of the encoded protein. In such methods, gene deletion can be achieved by homologous recombination using an integration plasmid / vector constructed to contain adjacent 5' and 3' regions flanking the gene. The flanking 5' and 3' regions can be introduced into a filamentous fungal cell, for example, by an integration plasmid / vector associated with a selectable marker that allows the plasmid to be integrated into the cell.

[0215] In other embodiments, the engineered strain of filamentous fungi comprises a genetic modification that disrupts or inactivates a gene of interest. Typical methods of gene disruption / inactivation include disrupting any portion of the gene, including the polypeptide coding sequence (CDS), promoter, enhancer, or other regulatory elements, and the disruption may be substitution, insertion, deletion, inversion, or combinations and variations thereof. Non-limiting examples of gene disruption techniques include inserting (integrating) an integration plasmid containing a nucleic acid fragment homologous to the gene of interest into one or more genes of the present disclosure, which creates an overlap between the homologous region and the integration (insertion) vector DNA. In certain other non-limiting examples, gene disruption techniques include inserting an integration plasmid containing a nucleic acid fragment homologous to the gene of interest into the gene of interest, creating an overlap between the region of homology and the overlapping region of the integration (insertion) vector DNA, where the inserted vector DNA, for example, separates the gene's promoter from the protein-coding region or interrupts (disrupts) the gene's coding or non-coding sequence, resulting in a phenotype of enhanced protein production. The disruption construct may also be a selectable marker gene (e.g., pyr2) accompanied by 5' and 3' regions homologous to the gene of interest. The selectable marker allows for identification of transformants containing the disrupted gene. Thus, in certain embodiments, gene disruption involves modification of a gene's regulatory elements, such as the promoter, ribosome binding site (RBS), untranslated region (UTR), or codon change.

[0216] In other embodiments, engineered strains of filamentous fungi are constructed (i.e., genetically modified) by introducing, substituting, or removing one or more nucleotides within genes or regulatory elements required for their transcription or translation. For example, nucleotides can be inserted or removed to introduce a premature stop codon, remove a start codon, or cause a frameshift in the open reading frame (ORF). Such modifications can be achieved by site-directed mutagenesis or PCR-mediated mutagenesis, according to methods known in the art.

[0217] In other embodiments, engineered strains of filamentous fungi are constructed by a gene conversion process. For example, in gene conversion methods, a nucleic acid sequence corresponding to a target gene is mutated in vitro to generate a defective nucleic acid sequence, which is then transformed into a parent cell to produce mutant cells containing the defective gene. The defective nucleic acid sequence replaces the endogenous gene through homologous recombination. It may be desirable for the defective gene or gene fragment to also encode a marker that can be used to select for transformants containing the defective gene. For example, the defective gene can be introduced into a non-replicating or temperature-sensitive plasmid associated with a selectable marker. Selection for integration of the plasmid is affected by selection for the marker under conditions that do not permit plasmid replication. Selection for a second recombination event resulting in gene replacement is affected by examining colonies for loss of the selectable marker and acquisition of the mutated gene.

[0218] In other embodiments, modified strains of filamentous fungi are constructed using established antisense (gene silencing) techniques, using a nucleotide sequence complementary to the nucleic acid sequence of a gene of interest. More specifically, gene expression by a filamentous fungal strain may be reduced (downregulated) or eliminated by introducing a nucleotide sequence complementary to the nucleic acid sequence of the gene, which is transcribed intracellularly and can hybridize to mRNA produced intracellularly. Thus, under conditions in which the complementary antisense nucleotide sequence can hybridize to the mRNA, the amount of translated protein is reduced or eliminated. Such antisense methods include, but are not limited to, RNA interference (RNAi), small interfering RNA (siRNA), microRNA (miRNA), and antisense oligonucleotides, all of which are well known to those skilled in the art.

[0219] In other embodiments, modified strains of filamentous bacteria are constructed by random or directed mutagenesis using methods well known in the art, including, but not limited to, chemical mutagenesis and transposition. Genetic modification can be performed by subjecting parent cells to mutagenesis and selecting for mutant cells in which gene expression is reduced or eliminated. Mutagenesis can be directed or random, for example, by using suitable physical or chemical mutagens, by using suitable oligonucleotides, or by subjecting DNA sequences to PCR-generated mutagenesis. Furthermore, mutagenesis can be performed by using any combination of these mutagenesis methods. Examples of physical or chemical mutagens suitable for purposes of the present invention include ultraviolet (UV) irradiation, hydroxylamine, N-methyl-N'-nitro-N-nitrosoguanidine (MNNG), N-methyl-N'-nitrosoguanidine (NTG), O-methylhydroxylamine, nitrous acid, ethyl methanesulfonate (EMS), sodium bisulfite, formic acid, and nucleotide analogs. When such agents are used, mutagenesis is typically carried out by incubating the parent cells to be mutagenized in the presence of the mutagenizing agent of choice under appropriate conditions and selecting for mutant cells that exhibit reduced or no expression of the gene.

[0220] In certain other embodiments, modified strains of filamentous fungi are constructed by site-specific gene editing techniques. For example, in certain embodiments, variant strains of filamentous fungi are constructed (i.e., genetically modified) by using transcription activator-like endonucleases (TALENs), zinc finger endonucleases (ZFNs), homing (mega) endonucleases, etc. More specifically, the part of the gene to be modified (e.g., coding region, non-coding region, leader sequence, propeptide sequence, signal sequence, transcription terminator, transcription activator, or other regulatory elements required for expressing the coding region) is subjected to genetic modification by ZFN gene editing, TALEN gene editing, homing (mega) endonucleases, etc., and these modification methods are well known and available to those skilled in the art.

[0221] In certain other embodiments, modified strains of filamentous fungi are constructed by CRISPR / Cas9 editing. More specifically, compositions and methods for fungal genome modification using the CRISPR / Cas9 system have been described and are well known in the art (see, for example, PCT Publication Nos. WO 2016 / 100571, WO 2016 / 100568, WO 2016 / 100272, WO 2016 / 100562, etc.). Thus, genes of interest can be disrupted, deleted, mutated, or otherwise genetically modified by nucleic acid-guided endonucleases that find their target DNA by binding either guide RNA (e.g., Cas9) or guide DNA (e.g., NgAgo), which recruits the endonucleases to target sequences on the DNA, and the endonucleases can generate single- or double-strand breaks in the DNA. This targeted DNA break provides a substrate for DNA repair, which, combined with the provided editing template, can disrupt or delete the gene. For example, a gene encoding a nucleic acid-guided endonuclease (e.g., Cas9 from S. pyogenes or a codon-optimized gene encoding a Cas9 nuclease) is operably linked to a promoter active in filamentous fungal cells and a terminator active in filamentous fungal cells, thereby creating a filamentous fungal Cas9 expression cassette. Similarly, one or more target sites unique to a gene of interest are easily identified by those skilled in the art. For example, to construct a DNA construct encoding a gRNA directed to a target site within a gene of interest, a variable targeting domain (VT) would contain the 5' (PAM) protospacer adjacent motif (TGG) target site nucleotides, which are fused to DNA encoding the Cas9 endonuclease recognition domain (CER) for S. pyogenes Cas9. DNA encoding the gRNA is generated by combining DNA encoding the VT domain and DNA encoding the CER domain.Thus, a filamentous fungal expression cassette for a gRNA is created by operably linking DNA encoding the gRNA to a promoter active in a filamentous fungal cell and a terminator active in a filamentous fungal cell.

[0222] In certain embodiments, the DNA break induced by endonuclease is repaired / replaced by the incoming sequence.For example, a nucleotide editing template is provided so that the DNA repair mechanism of the cell can use the editing template to precisely repair the DNA break generated by the above-mentioned Cas9 expression cassette and gRNA expression cassette.For example, about 500bp of the 5' target gene can be fused to about 500bp of the 3' target gene to generate an editing template, and this template is used by the mechanism of the filamentous fungus host to repair the DNA break generated by RGEN (RNA-guided endonuclease).

[0223] The Cas9 expression cassette, gRNA expression cassette, and editing template can be co-delivered into filamentous fungal cells using a number of different methods (e.g., protoplast fusion, electroporation, natural competence, or induced competence). Transformed cells are selected by PCR by amplifying the target locus with forward and reverse primers. These primers can amplify the wild-type locus or the modified locus edited by RGEN. These fragments are then sequenced using sequencing primers to identify edited colonies.

[0224] Another method for genetically modifying a gene of interest is by altering the expression level of the gene of interest. For example, nuclease-deficient variants of such nucleotide-guided endonucleases (e.g., Cas9 D10A, N863A or Cas9 D10A, H840A) can be used to enhance or attenuate the transcription of target genes, thereby regulating the expression level of the gene. These Cas9 variants are inactive due to all nuclease domains present in the protein sequence, but retain RNA-guided DNA binding activity (i.e., these Cas9 variants cannot cleave either strand of DNA when bound to their cognate target site). Thus, nuclease-deficient proteins (i.e., Cas9 variants) can be expressed as filamentous fungal expression cassettes, and when combined with a filamentous fungal gRNA expression cassette, the resulting Cas9 variant protein is directed to a specific target sequence within the cell. The binding of Cas9 (variant) protein to specific gene target sites can reduce the amount of gene product produced by blocking the binding or movement of the transcription machinery on cellular DNA.Therefore, any of the genes disclosed herein can be targeted to reduce gene expression using this method.Gene silencing can be monitored in cells containing nuclease-deficient Cas9 expression cassettes and gRNA expression cassettes by using methods such as RNA sequencing.

[0225] IV. Fermentation and Recovery of Globin Protein As briefly described in the preceding sections, the present strains and methods are used for the production of commercially important proteins in submerged culture of filamentous fungi. For example, in certain embodiments, recombinant (modified) filamentous fungal cells containing the introduced expression cassette, when fermented under conditions suitable for the production of globin protein, produce at least about 30 grams of total protein per liter of broth (g / L). In one or more other embodiments, modified filamentous fungal cells containing the introduced expression cassette, when fermented under conditions suitable for the production of globin protein, produce at least about 0.5 grams of bin protein per liter of broth (g / L). Thus, in certain embodiments, total protein titer and globin protein titer can be defined as the amount of total protein per volume (g / L) and the amount of globin protein per volume (g / L), respectively. For example, protein titer can be measured by methods known in the art (e.g., ELISA, HPLC, Bradford assay, LC / MS, etc.).

[0226] In certain other embodiments, modified filamentous fungal cells containing the introduced cassette may be described by their volumetric productivity, defined as the amount of protein (g) produced during fermentation per nominal bioreactor volume (L) per total fermentation time (h). For example, volumetric productivity can be measured by methods known in the art (e.g., ELISA, HPLC, Bradford assay, LC / MS, etc.).

[0227] In certain other embodiments, modified filamentous fungal cells containing the introduced cassette may be described according to total protein yield, which is defined as the amount of protein (g) produced per gram of carbohydrate fed, compared to the (unmodified) parent strain. Thus, as used herein, total protein yield (g / g) may be calculated using the following formula: Yf=Tp / Tc where "Yf" is the total protein yield (g / g), "Tp" is the total protein produced during the fermentation (g), and "Tc" is the total carbohydrate (g) fed during the fermentation (bioreactor) run.

[0228] Total protein yield can also be described as carbon conversion efficiency / carbon yield, e.g., as the percentage (%) of supplied carbon that is incorporated into total protein. Thus, in certain embodiments, modified filamentous fungal cells containing an introduced cassette can be described according to their carbon conversion efficiency (e.g., an increase in the percentage (%) of supplied carbon that is incorporated into total protein).

[0229] In certain other embodiments, modified filamentous fungal cells containing the introduced cassette may be described according to the specific productivity (Qp) of globin protein. For example, determining the specific productivity (Qp) is a suitable method for assessing the rate of globin protein production, and Qp may be determined using the following formula: "Qp = gP / gDCW·hr" where "gP" is the grams of protein produced in the tank, "gDCW" is the grams of dry cell weight (DCW) in the tank, and "hr" is the fermentation time (hours) from the time of inoculation, which includes the production time and growth time.

[0230] In certain embodiments, the present disclosure provides, inter alia, compositions and methods for producing globin proteins, comprising fermenting filamentous fungal cells containing one or more introduced globin protein cassettes, wherein the fungal cells express and secrete (i.e., produce) the globin proteins. Generally, fermentation methods well known in the art are used to ferment the fungal cells. In some embodiments, the fungal cells are grown under batch or continuous fermentation conditions.

[0231] Classical batch fermentation is a closed system in which the composition of the medium is set at the beginning of the fermentation and remains unchanged during the fermentation. At the beginning of the fermentation, the medium is inoculated with the desired organism. This method allows fermentation to occur without adding any components to the system. Typically, batch fermentation is considered "batch" with respect to the addition of a carbon source, and factors such as pH and oxygen concentration are often controlled. The metabolite and biomass composition of a batch system changes constantly until the fermentation is stopped. Within a batch culture, cells progress through a static lag phase, a high-growth logarithmic phase, and finally a stationary phase where growth rate slows or stops. If untreated, cells in the stationary phase eventually die. Generally, cells in the logarithmic phase are responsible for the majority of product production.

[0232] A suitable variation of the standard batch system is the "fed-batch fermentation" system. In this variation of the typical batch system, substrate is added gradually as the fermentation progresses. Fed-batch systems are useful when catabolite repression is likely to inhibit cellular metabolism and when a limited amount of substrate is desired in the medium. Measurement of the actual substrate concentration in a fed-batch system is difficult and therefore is estimated based on changes in measurable factors such as pH, dissolved oxygen, and the partial pressure of waste gases such as CO2. Batch and fed-batch fermentation are common and well known in the art.

[0233] Continuous fermentation is an open system in which a defined fermentation medium is continuously added to a bioreactor and an equal amount of conditioned medium is simultaneously removed for processing. Continuous fermentation generally maintains the culture at a constant high density, with cells primarily in logarithmic growth phase. Continuous fermentation allows for the adjustment of one or more factors that affect cell growth and / or product concentration. For example, in one embodiment, a limiting nutrient, such as a carbon or nitrogen source, can be maintained at a fixed ratio while all other parameters are adjusted. In other systems, multiple factors affecting growth can be continuously varied while the cell concentration, as measured by medium turbidity, remains constant. Continuous systems attempt to maintain steady-state growth conditions. Therefore, cell loss due to medium removal must be balanced against the cell growth rate during fermentation. Methods for adjusting nutrients and growth factors in continuous fermentation processes, as well as techniques for maximizing product formation rates, are well known in the art of industrial microbiology.

[0234] To ensure proper microbial growth, maximize assimilation of carbon and energy sources by the cells in the microbial conversion process, and achieve maximum cell yield at maximum cell density in the fermentation medium, it is essential to supply appropriate amounts of inorganic nutrients in the proper proportions in addition to the carbon and energy sources, oxygen, assimilable nitrogen, and microbial inoculant.

[0235] The composition of the aqueous mineral medium can vary over a wide range, depending in part on the microorganism and substrate utilized, as is known in the art. In addition to nitrogen, the mineral medium will contain appropriate amounts of phosphorus, magnesium, calcium, potassium, sulfur, and sodium in suitable soluble, absorbable, ionic, complex forms, and preferably also certain trace elements such as copper, manganese, molybdenum, zinc, iron, boron, and iodine, again in suitable soluble, assimilable forms, all of which are known in the art.

[0236] Fermentation reactions are aerobic processes in which the necessary molecular oxygen is supplied by a molecular oxygen-containing gas, such as air, oxygen-enriched air, or even substantially pure molecular oxygen, provided to maintain the contents of the fermentor at an appropriate oxygen partial pressure effective to support vigorous growth of the microbial species.

[0237] Microorganisms also require an assimilable nitrogen source. The assimilable nitrogen source can be any nitrogen-containing compound or a compound capable of releasing nitrogen in a form suitable for metabolic utilization by the microorganism. While various organic nitrogen source compounds, such as protein hydrolysates, can be used, typically inexpensive nitrogen-containing compounds such as ammonia, ammonium hydroxide, urea, and various ammonium salts, e.g., ammonium phosphate, ammonium sulfate, ammonium pyrophosphate, ammonium chloride, or various other ammonium compounds, are available. Ammonia gas itself is convenient for large-scale operations and can be used by bubbling an appropriate amount through the aqueous fermentation product (fermentation medium). At the same time, such ammonia can also be used to assist in pH control.

[0238] The pH range of the aqueous microbial fermentation product (fermentation mixture) should be in the exemplary range of about 2.0 to 8.0. For filamentous fungi, the pH is typically within the range of about 2.5 to 8.0; for T. reesei, the pH is typically within the range of about 3.0 to 7.0. The preferred pH range for a microorganism will depend to some extent on the medium and the specific microorganism utilized and will vary somewhat with changes in the medium, as can be readily determined by one of ordinary skill in the art. In certain embodiments, the fermentation is carried out in a manner that allows for control of the carbon-containing substrate as the limiting factor, thereby achieving good conversion of the carbon-containing substrate to the cells and avoiding contamination of the cells with significant amounts of unconverted substrate. The latter is not an issue with water-soluble substrates, as any traces remaining can be easily washed away. However, this can be problematic with water-insoluble substrates, necessitating additional product processing steps, such as appropriate washing steps.

[0239] As noted above, the time to reach this level is not critical and may vary depending on the particular microorganism and fermentation process being performed, however, methods for determining the carbon source concentration in a fermentation medium and whether the desired carbon source level has been achieved are well known in the art.

[0240] Fermentation can be carried out as a batch or continuous operation, with fed-batch operation being preferred for ease of control, production of uniform amounts of product, and the most economical use of all equipment.

[0241] If necessary, some or all of the carbon and energy source materials and / or some of the assimilable nitrogen source, such as ammonia, can be added to the aqueous mineral medium before it is fed to the fermenter.

[0242] Each of the streams introduced into the reactor is preferably controlled at a predetermined rate or according to needs that can be determined by monitoring the concentrations of carbon and energy substrates, pH, dissolved oxygen, oxygen or carbon dioxide in the fermentor off-gas, cell density as measured by dry cell weight, light transmittance, etc. The feed rates of the various materials can be varied to obtain the fastest possible cell growth rate and the highest possible yield of microbial cells relative to the substrate feed, consistent with efficient utilization of the carbon and energy sources.

[0243] In a batch or, preferably, fed-batch operation, all equipment, reactors, or fermentation means, tanks or vessels, piping, and associated circulation or cooling devices are first sterilized, typically by using steam, for example, at about 121° C. for at least about 15 minutes. The sterilized reactor is then inoculated with a culture of the selected microorganism in the presence of all necessary nutrients, including oxygen, and a carbon-containing substrate. The type of fermentor utilized is not critical.

[0244] Harvesting and purification of globin proteins from the fermentation broth can be carried out by procedures known to those skilled in the art. For example, as described above, the recombinant fungal strains of the present disclosure can be constructed to secrete one or more globin proteins into the fermentation broth, simplifying the globin protein recovery process (e.g., not requiring cell lysis), thereby reducing the cost of globin protein production. The fermentation broth will generally contain cells, various suspended solids and other biomass contaminants, and cellular debris, including the desired globin protein, which are removed from the fermentation broth by means known in the art.

[0245] In certain other embodiments or aspects, the purified globin (protein) preparation may be derived from or recovered from a harvested and collected fermentation broth.

[0246] As used herein, the terms "purified," "isolated," or "enriched" with respect to globin (protein) mean that the globin has been converted from a less pure state by separating it from some or all of the contaminants with which it is associated, including, but not limited to, microbial cells, metabolites, solvents, chemicals, color, aggregates, processing aids, inhibitors, fermentation media, cellular debris, nucleic acids, proteins other than the target leghemoglobin, host cell proteins, cross-contaminants from production equipment, and the like.

[0247] Thus, in the context of "purified globin," as used herein, purification may be performed by any art-recognized separation technique, including, but not limited to, ion exchange chromatography, affinity chromatography, hydrophobic separation, dialysis, protease treatment, heat treatment, ammonium sulfate precipitation or other protein salting-out, crystallization, centrifugation, size exclusion chromatography, filtration, microfiltration, gel electrophoresis, or gradient separation to remove whole cells, cell debris, impurities, extraneous proteins, or enzymes undesired in the final composition.

[0248] Further components that impart additional benefits, such as activators, anti-inhibitors, desired ions, pH-adjusting compounds, or other enzymes or chemicals, can then be added to the purified or isolated globin composition.

[0249] As used herein, globin "purity" is a relative term and is not meant to be limiting when used in phrases such as "the recovered globin is more pure, the same purity, or less pure than before the recovery process." For example, the relative "purity" of globin (protein) before and after the recovery process can be determined using methods known in the art, including, but not limited to, common quantification methods (e.g., Bradford, UV-Vis, activity assays), electrophoretic analysis (SDS-PAGE), analytical HPLC-mass spectrometry, hydrophobic interaction chromatography, etc.

[0250] Non-limiting examples of obtaining relative purity of globin include SDS-PAGE analysis and / or the A of non-globin (impurities) relative to globin. 280 These include, but are not limited to, ratios. For example, relative globin purity via SDS-PAGE can be determined by the visual abundance of globin (protein) bands compared to non-globin protein (undesired contaminant; impurity) bands present in the preparation. Alternatively, the relative purity of globin can be determined by the ratio of non-globin to globin (A 280 ) ratio.

[0251] For example, A 280 The ratio is a measure of the amount of absorbance at 280 nm contributed by non-globin impurities (e.g., non-globin A) relative to one unit of absorbance at 280 nm contributed by globin in a protein preparation. 280 / Globin A 280 ), a lower number means higher purity. More specifically, globin (A 280) To determine the concentration, the concentration can be measured by HPLC using purified globin as a standard. The HPLC concentration is determined using an absorbance of 1 mg / mL globin at 280 nm = 1.00. 280 The non-globin (A 280 ) method to determine the concentration is to zero with MilliQ water and, if necessary, A with MilliQ water. 280 Non-globin A can be measured using a 1 cm glass cuvette diluted to <1 280 The concentration was determined by measuring the globin A 280 It is calculated by subtracting

[0252] Thus, in one or more embodiments or aspects of the present disclosure, globin protein is recovered from a fermentation broth of filamentous fungal cells fermented under conditions suitable for the production of globin protein. In certain embodiments, the filamentous fungal host cells are constructed for secretion of globin protein into the fermentation broth, and the globin protein is recovered from the end of fermentation (EOF) broth. In certain other embodiments, such as when the filamentous fungal host cells are constructed for intracellular globin expression, the EOF fermentation broth is subjected to a cell lysis process, and the globin protein is recovered from the lysed cell broth.

[0253] Thus, in certain embodiments or aspects, the recovered globin is of higher purity after undergoing one or more recovery processes described herein. For example, the fermentation broth (e.g., whole broth at the end of fermentation) may be subjected to one or more protein recovery processes, including, but not limited to, a broth conditioning process, a broth clarification process, a protein concentration and / or protein purification process (e.g., protein concentration, filtration, precipitation, crystallization, crystal separation, crystal precipitate dissolution process, etc.), a buffer exchange process, a sterile filtration process, etc. In certain other aspects, the fermentation broth is subjected to a broth treatment (broth conditioning) process to improve subsequent broth handling characteristics. In certain other embodiments, the fermentation broth is subjected to a cell lysis process before recovering the globin, such as when a filamentous fungal host cell is constructed for intracellular globin expression. Cell lysis processes include, but are not limited to, enzymatic treatment (e.g., lysozyme, proteinase K treatment), chemical means (e.g., ionic liquids), physical means (e.g., French press, ultrasound), simply maintaining the culture without feed, and the like.

[0254] Thus, the disclosed methods / processes are not meant to be limiting, as one of skill in the art can readily adapt or modify one or more of the compositions and / or methods disclosed herein for the recovery of particular globin proteins, and / or combinations thereof, as described herein. In certain aspects, the fermentation broth obtained by fermenting filamentous fungal cells that express and secrete globin proteins can be processed by harvesting, clarifying, and concentrating the broth as generally described herein.

[0255] V. Illustrative Embodiments Non-limiting embodiments of the present disclosure include, but are not limited to, the following:

[0256] 1. Recombinant filamentous fungal cells expressing heterologous globin proteins.

[0257] 2. Recombinant filamentous fungal cells that express and secrete heterologous globin proteins when fermented under suitable conditions.

[0258] 3. The recombinant cell of embodiment 1 or embodiment 2, wherein the recombinant cell is selected from the group consisting of an Acremonium sp. cell, an Aspergillus sp. cell, an Emericella sp. cell, a Fusarium sp. cell, a Humicola sp. cell, a Mucor sp. cell, a Myceliophthora sp. cell, a Neurospora sp. cell, a Penicillium sp. cell, a Scytalidium sp. cell, a Talaromyces sp. cell, a Thielavia sp. cell, a Tolypocladium sp. cell, and a Trichoderma sp. cell.

[0259] 4. The recombinant cell of any one of embodiments 1-3, comprising an introduced expression cassette encoding a globin protein, the cassette comprising an upstream (5') promoter (pro) sequence operably linked to a downstream nucleic acid (sig-seq) encoding a protein (signal) secretion sequence operably linked to a downstream (3') nucleic acid encoding the globin protein (globin CDS).

[0260] 5. The recombinant cell of embodiment 4, wherein the nucleic acid encoding the globin protein (globin CDS) comprises an upstream (5') nucleic acid encoding an N-terminal protein fusion (N-fusion).

[0261] 6. The recombinant cell of embodiment 4, wherein the nucleic acid encoding the globin protein (globin CDS) comprises an upstream (5') nucleic acid encoding an N-terminal protein cleavage site (N-linker).

[0262] 7. The recombinant cell of embodiment 4, wherein the nucleic acid encoding the globin protein (globin CDS) comprises a downstream (3') nucleic acid encoding a C-terminal protein fusion (C-fusion).

[0263] 8. The recombinant cell of embodiment 4, wherein the nucleic acid encoding the globin protein (globin CDS) comprises a downstream (3') nucleic acid encoding a C-terminal protein cleavage site (C-linker).

[0264] 9. The recombinant cell of embodiment 4, wherein the nucleic acid encoding the globin protein (globin CDS) comprises an upstream (5') nucleic acid encoding an N-terminal protein fusion (N-fusion) and a downstream (3') nucleic acid encoding a C-terminal protein fusion (C-fusion).

[0265] 10. The recombinant cell of embodiment 4, wherein the nucleic acid encoding the globin protein (globin CDS) comprises an upstream (5') nucleic acid encoding an N-terminal protein cleavage site (N-linker) and a downstream (3') nucleic acid encoding a C-terminal protein cleavage site (C-linker).

[0266] 11. The recombinant cell of any one of embodiments 4 to 10, wherein the nucleic acid encoding the globin protein (globin CDS) comprises a combination of upstream nucleic acids encoding an N-terminal protein fusion and an N-terminal protein cleavage site, in any order, and / or a combination of downstream nucleic acids encoding a C-terminal protein fusion and a C-terminal protein cleavage site, in any order.

[0267] 12. The recombinant cell of any one of embodiments 4 to 11, wherein the cassette further comprises a terminator sequence operably linked to and located at the 3' end of the cassette.

[0268] 13. The recombinant cell of any one of embodiments 4 to 12, wherein the cassette is integrated into the genome of the cell.

[0269] 14. The recombinant cell of any one of embodiments 4 to 13, comprising at least two introduced polynucleotides encoding the same or different globin proteins and / or at least two introduced expression cassettes encoding the same or different globin proteins.

[0270] 15. The recombinant cell of any one of embodiments 1 to 14, further comprising an introduced expression cassette encoding a protease inhibitor.

[0271] 16. The recombinant cell of embodiment 15, wherein the protease inhibitor is selected from the group consisting of naturally occurring barley amylase subtilisin inhibitor (BASI) proteins and functional variants thereof, naturally occurring soybean trypsin inhibitor (STI) proteins and functional variants thereof, naturally occurring Bowman-Birk inhibitor (BBI) proteins and functional variants thereof, naturally occurring Kunitz inhibitor (KTI) proteins and functional variants thereof, and naturally occurring potato metallocarboxypeptidase inhibitor proteins and functional variants thereof.

[0272] 17. The recombinant cell of any one of embodiments 1-16, further comprising a deletion of one or more endogenous genes encoding one or more secreted proteases.

[0273] 18. The recombinant cell of embodiment 17, wherein the one or more proteases are selected from the group consisting of subtilisin-like serine proteases, aspartic acid proteases, trypsin-like serine proteases, glutamic acid proteases, and aminopeptidases.

[0274] 19. The recombinant cell of any one of embodiments 1-18, which is fermented for at least about 96 hours to about 300 hours.

[0275] 20. The recombinant cell of embodiment 19, wherein the cell produces at least 0.1 grams of globin protein per liter of fermentation broth (g / L).

[0276] 21. The recombinant cell of embodiment 19, wherein the cell produces at least 0.2 to 0.5 grams of globin protein per liter of fermentation broth (g / L).

[0277] 22. The recombinant cell of embodiment 19, which is fermented for about 180 to 190 hours.

[0278] 23. The recombinant cell of embodiment 22, wherein the cell produces at least 1 gram of globin protein per liter of fermentation broth (g / L).

[0279] 24. The recombinant cell of embodiment 1 or embodiment 2, wherein the globin protein is selected from the group consisting of leghemoglobin, myoglobin, hemoglobin, cyanoglobin, and non-symbiotic hemoglobin.

[0280] 25. The recombinant cell of embodiment 24, wherein the leghemoglobin protein comprises at least about 60% to 100% sequence identity to a native soybean leghemoglobin protein (SEQ ID NO: 1) or at least about 60% to 100% sequence identity to a native kidney bean leghemoglobin protein (SEQ ID NO: 19).

[0281] 26. The recombinant cell of embodiment 24, wherein the myoglobin protein comprises at least about 60% to 100% sequence identity to native bovine myoglobin protein (SEQ ID NO: 3).

[0282] 27. The recombinant cell of embodiment 24, wherein the globin protein comprises a globin superfamily domain and a functional porphyrin (heme) binding site.

[0283] 28. A method for producing a heterologous globin protein in a filamentous fungal cell, comprising introducing into the filamentous fungal cell an expression cassette encoding the globin protein, the cassette comprising an upstream (5') promoter (pro) sequence operably linked to a downstream nucleic acid (sig-seq) encoding a protein (signal) secretion sequence operably linked to a downstream (3') nucleic acid encoding the globin protein (globin CDS), and fermenting the modified cell under conditions suitable for production of the globin protein.

[0284] 29. The filamentous fungal cell is selected from the group consisting of an Acremonium sp. cell, an Aspergillus sp. cell, an Emericella sp. cell, a Fusarium sp. cell, a Humicola sp. cell, a Mucor sp. cell, a Myceliophthora sp. cell, a Neurospora sp. cell, a Penicillium sp. cell, a Scytalidium sp. cell, a Talaromyces sp., a Thielavia sp. cell, a Tolypocladium sp. cell, and a Trichoderma sp. cell. 29. The method of embodiment 28, wherein the cell is selected from the group consisting of: (Illegible) cells.

[0285] 30. The method of embodiment 28, wherein the nucleic acid encoding the globin protein (globin CDS) comprises an upstream (5') nucleic acid encoding an N-terminal protein fusion (N-fusion).

[0286] 31. The method of embodiment 28, wherein the nucleic acid encoding the globin protein (globin CDS) comprises an upstream (5') nucleic acid encoding an N-terminal protein cleavage site (N-linker).

[0287] 32. The method of embodiment 28, wherein the nucleic acid encoding the globin protein (globin CDS) comprises a downstream (3') nucleic acid encoding a C-terminal protein fusion (C-fusion).

[0288] 33. The method of embodiment 28, wherein the nucleic acid encoding the globin protein (globin CDS) comprises a downstream (3') nucleic acid encoding a C-terminal protein cleavage site (C-linker).

[0289] 34. The method of embodiment 28, wherein the nucleic acid encoding the globin protein (globin CDS) comprises an upstream (5') nucleic acid encoding an N-terminal protein fusion (N-fusion) and a downstream (3') nucleic acid encoding a C-terminal protein fusion (C-fusion).

[0290] 35. The method of embodiment 28, wherein the nucleic acid encoding the globin protein (globin CDS) comprises an upstream (5') nucleic acid encoding an N-terminal protein cleavage site (N-linker) and a downstream (3') nucleic acid encoding a C-terminal protein cleavage site (C-linker).

[0291] 36. The method of any one of embodiments 28-35, wherein the nucleic acid encoding the globin protein (globin CDS) comprises a combination of upstream nucleic acids encoding an N-terminal protein fusion and an N-terminal protein cleavage site, in any order, and / or a combination of downstream nucleic acids encoding a C-terminal protein fusion and a C-terminal protein cleavage site, in any order.

[0292] 37. The method of any one of embodiments 28-36, wherein the cassette further comprises a terminator sequence operably linked and located at the 3' end of the cassette.

[0293] 38. The method of any one of embodiments 28-37, wherein the cassette is integrated into the genome of the cell.

[0294] 39. The method of any one of embodiments 28-38, comprising at least two introduced polynucleotides encoding the same or different globin proteins and / or comprising at least two introduced expression cassettes encoding the same or different globin proteins.

[0295] 40. The method of any one of embodiments 28-39, further comprising an introduced expression cassette encoding a protease inhibitor.

[0296] 41. The method of embodiment 40, wherein the protease inhibitor is selected from the group consisting of naturally occurring barley amylase subtilisin inhibitor (BASI) proteins and functional variants thereof, naturally occurring soybean trypsin inhibitor (STI) proteins and functional variants thereof, naturally occurring Bowman-Birk inhibitor (BBI) proteins and functional variants thereof, naturally occurring Kunitz inhibitor (KTI) proteins and functional variants thereof, and naturally occurring potato metallocarboxypeptidase inhibitor proteins and functional variants thereof.

[0297] 42. The method of any one of embodiments 28-41, further comprising a deletion of one or more endogenous genes encoding one or more secreted proteases.

[0298] 43. The method of embodiment 42, wherein the one or more proteases are selected from the group consisting of subtilisin-like serine proteases, aspartic acid proteases, trypsin-like serine proteases, glutamic acid proteases, and aminopeptidases.

[0299] 44. The method of any one of embodiments 28-43, wherein the cells are fermented for at least about 96 hours to about 300 hours.

[0300] 45. The method of embodiment 44, wherein the cells produce at least 0.1 grams of globin protein per liter of fermentation broth (g / L).

[0301] 46. ​​The method of embodiment 44, wherein the cells produce at least 0.2 to 0.5 grams of globin protein per liter of fermentation broth (g / L).

[0302] 47. The method of embodiment 44, wherein the fermentation is carried out for about 180 to 190 hours.

[0303] 48. The method of embodiment 47, wherein the cells produce at least 1 gram of globin protein per liter of fermentation broth (g / L).

[0304] 49. The method of embodiment 28, wherein fermenting the modified cells under conditions suitable for producing globin protein does not require or include a hemin medium supplement and / or does not require a 5-aminolevulinic acid (ALA) medium supplement.

[0305] 50. The method of embodiment 28, wherein the modified cells do not require or include heme biosynthetic pathway engineering for the production of globin proteins.

[0306] 51. The method of any one of embodiments 28-50, wherein the expressed globin is secreted and recovered from the fermentation broth.

[0307] 52. The method of embodiment 51, wherein the recovered globin protein is purified.

[0308] 53. The method of embodiment 28, wherein the globin protein is selected from the group consisting of leghemoglobin, myoglobin, hemoglobin, cyanoglobin, and non-symbiotic hemoglobin.

[0309] 54. The method of embodiment 53, wherein the leghemoglobin protein comprises at least about 60% to 100% sequence identity to a native soybean leghemoglobin protein (SEQ ID NO: 1) or at least about 60% to 100% sequence identity to a native kidney bean leghemoglobin protein (SEQ ID NO: 19).

[0310] 55. The method of embodiment 53, wherein the myoglobin protein comprises at least about 60% to 100% sequence identity to native bovine myoglobin protein (SEQ ID NO: 3). [Example]

[0311] Certain aspects of the present invention can be further understood in light of the following examples, which should not be construed as limiting. Modifications to materials and methods will be apparent to those skilled in the art. Standard recombinant DNA and molecular cloning techniques used herein are well known in the art (Ausubel et al., 1987; Sambrook et al., 1989).

[0312] Example 1 Expression and secretion of heterologous leghemoglobin proteins As briefly described above, the applicants of the present disclosure have contemplated, designed, and constructed recombinant (modified) filamentous fungal cells capable of producing heterologous globin proteins. In one or more specific embodiments, a recombinant polynucleotide (e.g., an expression cassette) encoding one or more heterologous globin proteins is introduced into a filamentous fungal cell of the present disclosure. For example, in certain embodiments, the expression cassette encoding a secreted globin protein includes (in the 5' to 3' direction), a promoter (pro) region sequence operably linked to a nucleic acid (sig-seq) encoding a protein (signal) secretion sequence operably linked to a nucleic acid (globin CDS) encoding the globin protein.

[0313] A. Construction of a Vector for Expression of Soybean Leghemoglobin as a Secreted Protein One strategy for improving the expression of heterologous globin proteins in filamentous fungal strains is to express the gene of interest (GOI) with a signal sequence and a strong promoter, such as the T. reesei cellobiohydrolase I gene promoter (Pcbh1). In particular, for expression of the soybean leghemoglobin C2 gene (NP_001235248.2, GI:1229203762), the gene was codon-optimized for expression in T. reesei (SEQ ID NO:2). For example, a synthetic DNA sequence containing a codon-optimized leghemoglobin C2 gene (designated "LegGm1b") containing the native Cbh1 signal sequence (SEQ ID NO:13) or the native Pep1 signal sequence (SEQ ID NO:14) was synthesized using GeneArt Seamless Cloning and Assembly Enzyme Mix (Thermo Fisher Scientific, Carlsbad, CA) (Twist Biosciences, San Francisco, CA).

[0314] The backbone of the expression vector was amplified from a plasmid containing the following features: 1 kb upstream (5') flanking homology sequence suitable for integration into the genomic locus of a T. reesei strain, the cbh1 promoter (Pcbh1) sequence (SEQ ID NO: 11), the Cbh1 protein secretion sequence (SEQ ID NO: 13) or the Pep1 protein secretion sequence (SEQ ID NO: 14), the cbh1 terminator (Tcbh1) region sequence (SEQ ID NO: 15), the T. reesei pyr2 gene marker sequence (SEQ ID NO: 16) for transformation in T. reesei, 1 kb downstream (3') flanking homology suitable for integration into the genomic locus of a T. reesei strain, and bacterial vector sequences for selection and maintenance of the plasmid in E. coli. The synthetic DNA encoding leghemoglobin (LegGm1b) contains 25 bp of 5'-flanking sequence overlapping the cbh1 promoter (Pcbh1) region and 25 bp of 3'-flanking sequence overlapping the cbh1 terminator (Tcbh1) region. This vector construction produced leghemoglobin expression vectors designated pCHL853 (pI1-Pcbh1-LegGm1b; SEQ ID NO: 36) and pCHL856 (pI1-Pcbh1-LegGm1b-pep1ss; SEQ ID NO: 37), as shown in Table 1 below.

[0315] B. Construction of a Vector for Expression of Bean Leghemoglobin as a Secreted Protein Common bean leghemoglobin (SEQ ID NO: 19, NCBI catalog: AAA33767.1) is a second exemplary globin protein, and Applicant constructed and screened a recombinant filamentous fungal strain for the expression and secretion of common bean leghemoglobin protein. More specifically, a synthetic DNA sequence (LegPv1b; SEQ ID NO: 19) encoding common bean leghemoglobin is codon-optimized for expression in T. reesei cells. For example, common bean leghemoglobin (LegPv1b) was synthesized (Twist Biosciences), and an expression construct was constructed similarly to that for LegGm1b into pCHL852.

[0316] C. Construction of a Vector for Expression of Soybean Leghemoglobin as a Secreted Fusion Protein Another strategy for improving heterologous globin protein expression / production in filamentous fungal strains is to express a globin gene of interest (GOI) with an N-terminal fusion protein, typically a highly expressed native filamentous fungal protein such as the T. reesei lignocellulolytic enzymes Cbh1, Cbh2, Eg1, Eg2, or glucoamylase. The globin GOI encoding the globin protein can be linked (fused) to the gene for the highly expressed native filamentous fungal protein (e.g., Cbh1) via a linker peptide sequence that can be cleaved by a native protease (e.g., Kex2 site; SEQ ID NO: 18; Goller et al. 1998) before secretion. The globin protein can also be linked to other suitable domains of highly expressed filamentous fungal proteins, such as the Cbh1 core domain (i.e., Cbh1 without the C-terminal cellulose-binding domain; SEQ ID NO: 7).

[0317] For expression of soybean leghemoglobin and bovine myoglobin, the Cbh1 full-length, mature protein (Cbh1 FL; SEQ ID NO: 9) or the Cbh1 core protein domain (Cbh1 core; SEQ ID NO: 7) can be used as an exemplary N-terminal fusion partner (see Table 1). The leghemoglobin C2 gene was codon-optimized for expression in T. reesei. A synthetic DNA sequence (LegGm1b; SEQ ID NO: 2) containing the codon-optimized leghemoglobin C2 gene, including upstream (5') and downstream (3') flanking sequences for construction of leghemoglobin fusion protein expression vectors, was synthesized using GeneArt Seamless Cloning and Assembly Enzyme Mix (Thermo Fisher Scientific, Carlsbad, CA) (Twist Biosciences, San Francisco, CA). More specifically, the backbone of the expression vector was amplified from a plasmid containing the following features: T. reesei strain single guide RNA sgRNA-TrC144F (SEQ ID NO: 45; 5'-GCUUUCGCCUUACUUCUGCAGGG-3') (Synthego, Redwood City, MD). a 1 kb upstream (5') flanking homology sequence suitable for integration into a preselected genomic locus "TrC114F" located at a targeting site in T. reesei (T. City, CA); a cbh1 promoter (Pcbh1) region sequence (SEQ ID NO:11); a DNA sequence encoding a Cbh1 secretion signal (SEQ ID NO:13); a DNA sequence encoding the Cbh1 protein core domain (Cbh1 core; (SEQ ID NO:7); a DNA sequence encoding a Kex2 protease cleavage site (SEQ ID NO:18); and a cbh1 terminator (Tcbh1) region sequence (SEQ ID NO:15). A T. reesei pyr2 gene (marker) cassette as a selectable marker for transformation in T. reesei, 1 kb of 3' flanking homology sequence downstream of the "TrC144F" locus, and a bacterial vector sequence for selection and maintenance of the plasmid in E. coli.The synthetic DNA sequence (LegGm1b; SEQ ID NO:2) encoding the leghemoglobin protein (SEQ ID NO:1) contains 25 base pairs (bp) of upstream (5') flanking sequence overlapping the Cbh1core-KEX2 region and 25 bp of downstream (3') flanking sequence overlapping the cbh1 terminator (Tcbh1) region. This vector construction generated a leghemoglobin expression vector designated "pLH1088" (SEQ ID NO:38).

[0318] A second leghemoglobin expression vector was constructed for expression of a fusion leghemoglobin protein with the complete (full-length, mature) Cbh1 protein. This vector contains the following features: 1 kilobase (kb) of upstream (5') flanking homology sequence at the "TrC114F" locus as described above, the cbh1 promoter region sequence (SEQ ID NO: 11), the full-length Cbh1 sequence (Cbh1 FL) including the Cbh1 signal sequence, and the C-terminal cellulose-binding domain, the leghemoglobin gene CDS (LegGm1b) followed by the KEX2 sequence (SEQ ID NO: 18), the cbh1 terminator region (SEQ ID NO: 15), and 1 kb of downstream (3') flanking homology sequence at the genomic locus "TrC114F." The remainder of the expression vector contains the T. reesei pyr2 gene and bacterial vector sequences for plasmid selection and maintenance in E. coli. The resulting expression vector was designated "pLH1104" (SEQ ID NO: 39).

[0319] A similar vector was constructed with the same sequence as pLH1104, except that leghemoglobin was expressed with both an N-terminal fusion of Cbh1 and a C-terminal fusion of a 6-His tag; this vector was designated "pLH1105" (SEQ ID NO: 40).

[0320] Example 2 Expression and secretion of heterologous bovine myoglobin proteins The bovine myoglobin gene (Mb; NP_776306.1, GI:27806939) was codon-optimized for expression in T. reesei (Mb1b, SEQ ID NO:4), and a synthetic DNA fragment containing Mb1b was obtained from Twist Biosciences. An Mb1b expression vector was constructed for expression of a fusion protein with the full-length Cbh1 (Cbh1 FL) protein. This vector contains the following features: 1 kb of 5' flanking homology sequence at the targeted genomic locus, a Pcbh1 promoter, the full-length Cbh1 sequence (Cbh1 FL), a synthetic Mb1b coding sequence followed by a KEX2 sequence, a Tcbh1 terminator, a pyr2 selection marker, and 1 kb of 3' flanking homology sequence at the targeted genomic locus. The remainder of the expression vector contains bacterial vector sequences for selection and maintenance of the plasmid in E. coli. The full-length Cbh1-myoglobin (fusion) protein expression vector (Cbh1FL-Mb1b) was designated "pLH1106" (pl1-Pcbh1-Cbh1FL-Mb1b), and the full-length Cbh1-myoglobin-His6 (fusion) protein expression vector was designated "pLH1107" (pl1-Pcbh1-Cbh1FL-Mb1b-His6). [Table 1] [Table 2]

[0321] Example 3 CAS9-directed targeted integration of a leghemoglobin expression cassette into the T. reesei genome In this example, leghemoglobin- and myoglobin-expressing T. reesei strains were generated by Cas9-guided targeted integration into the genome via homologous recombination (HR) or nonhomologous end joining (NHEJ) mechanisms. The Cas9-ribonucleoprotein complex (Cas9RNP) is composed of the Cas9 protein and a single-stranded guide RNA (sgRNA). Upon protoplast transformation of the Cas9RNP complex and a targeted integration cassette into T. reesei cells, the Cas9RNP enters the nucleus via a nuclear localization signal (NLS) at the C-terminus of the Cas9 protein. The guide RNA then guides the Cas9RNP to the targeted genomic locus, creating a double-strand break, which is subsequently repaired by either the integration cassette containing homologous sequences to both ends of the break site or various DNA repair mechanisms, such as NHEJ (nonhomologous end joining). For targeting via a homologous recombination approach, integration cassettes typically contain 50-1000 bp of sequence homologous to both the 5' and 3' ends of the integration locus. For integration cassettes constructed for leghemoglobin or myoglobin, 1 kb of 5' and 3' homologous sequence is included to improve the efficiency of chromosomal integration and homology-based recombination at the desired locus.

[0322] For the construction of leghemoglobin and myoglobin expression strains, linear expression cassettes were amplified by PCR using primers OT4268 and OT4269 (see Table 3 below) to generate DNA fragments containing 5' and 3' 1 kb flanking sequences for Cas9-targeted chromosomal integration at the chromosomal locus. [Table 3]

[0323] PCR reactions were prepared in 6x PCR strips, each containing eight PCR tubes with 50 μL reactions per tube in a total volume of 1.2 mL. The 5' PCR primer (OT4268) and 3' PCR primer (OT4269) were each added to a final concentration of 0.5 μM with 0.5 μg / mL of template DNA plasmid (Table 2). PCR reactions were performed using NEB-NEXT PCR Master Mix (New England Biolabs, MA) under the following conditions: 98°C, 30 seconds; 35 cycles (98°C, 10 seconds; 70°C, 30 seconds; 72°C, 4 minutes); 72°C, 4 minutes. The final PCR products (HG1-HG7, Table 2) were digested with DpnI enzyme (New England Biolabs) to remove the plasmid template DNA. The reaction mixture was purified using Zymo DNA Clean and Concentrator according to the manufacturer's protocol (Zymo Research, Irvine, CA) and dissolved in the elution buffer provided in the kit to a final concentration of 0.5–1.0 μg / μL.

[0324] The Cas9RNP complex used for targeted chromosomal integration of the above expression cassette was generated as follows: sgRNA targeting the locus in the T. reesei genome (TrC114F, without the protospacer adjacent motif (PAM) sequence "GGG") was obtained from Synthego (South San Francisco, CA) with the RNA sequence of SEQ ID NO: 45 (5'-GCUUUCGCCUUACUUCUGCA-3') and dissolved to 100 μM in TE buffer (10 mM Tris, 1 mM EDTA, pH 8.0). The Cas9RNP assembly reaction contains 12 μM Cas9 protein (New England Biolabs), 12 μM sgRNA-EclipseA in 1x NEB 3.1 buffer (New England Biolabs). The reaction is incubated at room temperature for 10 minutes and stored on ice until protoplast transformation is performed. Protoplast transformation of the T. reesei host was performed according to standard procedures using 5 μL of Cas9RNP complex, 4 μg of the PCR product of the linear integration cassette, and 300 μL of T. reesei protoplasts (10 per mL). 8 ) were combined. Transformation reactions were plated onto Vogel agar plates (for selection of the pyr2 marker) and incubated at 32°C for 5 days. Single colonies were picked from the Vogel agar plates and transferred onto fresh Vogel agar plates and incubated at 32°C for 3 days.

[0325] Colonies were then inoculated into 1 mL of NREL medium (pH 6.5), grown in microtiter plates at 28 °C for 4 days, and selected for leghemoglobin or myoglobin expression by adding 1xHalt protease inhibitor cocktail (Thermo Fisher Scientific). SDS-PAGE analysis was performed to screen for transformants that produced full-length Cbh1 protein (54 kDa), Cbh1 core domain protein (50 kDa), leghemoglobin protein (16 kDa), myoglobin protein (17 kDa), or a 6-histidine-tagged fusion protein. Integration of the linear expression cassette at the desired locus was screened and verified by colony PCR amplification from T. reesei transformants using primers OT4333 and OT4334.

[0326] To assess the reduction (alleviation) of proteolysis of secreted leghemoglobin and myoglobin proteins by natural proteases secreted by T. reesei, an expression cassette for barley amylase subtilisin (protease) inhibitor (BASI, SEQ ID NO: 6) was integrated into certain strains for coexpression of leghemoglobin or myoglobin with BASI. The BASI expression cassette contains the cbh2 promoter region (Pcbh2; SEQ ID NO: 12), a DNA sequence encoding the Pep1 signal sequence (SEQ ID NO: 14), the BASI gene CDS (SEQ ID NO: 6), and the TrpC terminator (TTrpC) region (SEQ ID NO: 46). Using pLH1108 as a template, the Pcbh2-BASI cassette was PCR amplified with primers OT4337 and OT4338 using NEB-NEXT Master Mix (98°C, 30 s; 35 cycles of (98°C, 10 s; 70°C, 30 s; 72°C, 4'); 72°C, 4') to yield a 4 kb product (HG8).

[0327] To construct a T. reesei strain coexpressing leghemoglobin and BASI, linear integration cassettes HG3 derived from PCR amplification of pLH1088 (PCR primers OT4268 and OT4269) and HG8 derived from pLH1108 amplification (PCR primers OT4337 and OT4338) were cotransformed with Cas9RNP-sgRNA-TrC114F into T. reesei protoplasts. The HG3 cassette integrated into the locus as confirmed by colony PCR using primers OT4333 and OT4334, whereas the HG8 cassette, which does not contain 5' or 3' homologous sequences to the T. reesei genome, randomly integrated into the genome. Transformants were used to inoculate 1 mL wells of 24-well microtiter plates containing NREL medium (pH 6.5, 1xHalt protease inhibitor cocktail) at 28°C for 4 days, and leghemoglobin (or myoglobin) and BASI expression were analyzed by SDS-PAGE.

[0328] Example 4 Expression and production of CBH1-leghemoglobin (N-terminus) fusion protein by secretion A fed-batch fermentation run was performed in a 2 L bioreactor for T. reesei strain BFZ28 (Pcbh1-CBH1core-KEX2-LegGm1b) as generally described in PCT Publication No. WO 2004 / 035070 (incorporated herein by reference). Specifically, cells were initially grown in minimal medium containing 75 g / L glucose until glucose was depleted. The production phase was initiated by adding 30 g / L VEG-Pro as a nutrient supplement, along with a combined supply of glucose and sophorose at pH 6.5 for induction of protein expression under the control of the cbh1 promoter (Pcbh1). For example, 10 mL whole broth samples were collected every 24 hours and frozen at -20°C, resulting in a total fermentation time of approximately 180 hours. The collected whole broth samples were thawed and centrifuged. The formation of red (dark) color in the supernatant indicates the secretion of globin protein into the culture supernatant, as shown in Figure 3. Similarly, the supernatant was analyzed by SDS-PAGE as shown in Figure 5, and the upper protein band at approximately 55 kDa is the Cbh1 core protein, and the leghemoglobin protein (approximately 16 kDa) co-migrates with the BASI protease inhibitor at approximately 20 kDa, or a minor band is present at approximately 17 kDa.

[0329] To verify leghemoglobin expression, fermentation supernatants were further analyzed by HPLC analysis of 188-hour supernatant samples from BFZ28 strain fermentation runs (Figure 6). The heme prosthetic group of the leghemoglobin protein was detected at 410 nm, which co-migrated with a protein peak detected at 280 nm with a retention time of 6.4 min. The inset in Figure 6 shows a spectral scan of the leghemoglobin peak at a retention time of 6.4 min, confirming the peak absorbance of the leghemoglobin protein at 410 nm.

[0330] Example 5 Co-expression and production of secreted leghemoglobin and secreted protease inhibitors Three fed-batch fermentation runs were performed in 2 L bioreactors for three T. reesei strains: (1) BGJ74 (Pcbh1-LegGm1b, Pcbh2-BASI) for direct expression of soybean leghemoglobin under the cbh1 promoter, (2) BGJ75 (Pcbh1-Pv1b, Pcbh2-BASI) for direct expression of common bean leghemoglobin under the cbh1 promoter, and (3) BGJ76 (Pcbh1-LegGm1b.Pep1ss, Pcbh2-BASI) for direct expression of soybean leghemoglobin under the cbh1 promoter.

[0331] Cells were initially grown in minimal medium containing 75 g / L glucose until glucose was depleted. The production phase was initiated with a combined glucose and sophorose feed at pH 7 containing 30 g / L VEG-Pro for induction of protein expression under the control of the cbh1 promoter (Pcbh1). Ten mL whole broth samples were collected every 24 hours and frozen at -20°C, resulting in a total fermentation time of 188 hours. For example, as shown in Figure 7, the total soluble protein secreted in the fermentation run is plotted against effective fermentation time (EFT, hours) for the three fermentation runs (strains BGJ74, BGJ75, and BGJ76) described above, along with data from the previously performed BFZ28 fermentation (BFZ28; Pcbh1-CBH1core-KEX2-LegGm1b) described in Example 4. The fermentation supernatant was analyzed by SDS-PAGE as shown in Figure 8, and a major protein band was detected at 20 kDa containing the BASI protease inhibitor and leghemoglobin protein (approximately 16 kDa) that co-migrated with the BASI protein. The presence of heme-containing leghemoglobin was further confirmed by HPLC.

[0332] To estimate leghemoglobin expression in these fermentation runs, supernatant samples 188 hours from the end of the fermentation runs were analyzed by HPLC (Figure 9). The heme prosthetic group of the leghemoglobin protein was detected at 410 nm, which co-migrated with a protein peak detected at 280 nm with a retention time of 6.4 minutes (Figure 9). The total protein secretion titers at the end of the 188-hour fermentation runs are shown below in Table 4, and leghemoglobin levels (g / L) were calculated based on the protein peak area at 280 nm. [Table 4]

[0333] Example 6 Secreted myoglobin production in double-copy expressing strains A. Construction of a double-copy expression vector for secretion of bovine myoglobin To improve myoglobin expression, we constructed a strain containing two different codon-optimized sequences of the bovine myoglobin gene. Specifically, a tandem copy expression vector (pLHX143) was constructed containing two expression cassettes for bovine myoglobin. The first cassette contains the Pcbh1 promoter (SEQ ID NO: 11), the Pep1 signal sequence (Pep1ss, SEQ ID NO: 14), the first codon-optimized bovine myoglobin gene (Mb1b, SEQ ID NO: 4), and the terminator sequence of the endoglucanase 1 gene (Tegl1, SEQ ID NO: 50). This sequence is immediately followed by a second cassette containing the Pcbh2 promoter (SEQ ID NO:12), the Cbh1 signal sequence (Cbh1ss, SEQ ID NO:13), a second codon-optimized bovine myoglobin gene (Mb1c, SEQ ID NO:49) encoding the same bovine myoglobin protein (SEQ ID NO:3), and the terminator sequence of the Cbh1 gene (Tcbh1, SEQ ID NO:15). This vector also contains the following features: 1 kb of 5' flanking homology sequence at the targeted genomic locus and 1 kb of 3' flanking homology sequence at the targeted genomic locus. The remainder of the expression vector contains bacterial vector sequences for plasmid selection and maintenance in E. coli. This dual-copy myoglobin expression vector was designated "pLHX143" (pI1-Pcbh1-Mb1b-Pcbh2-Mb1c, SEQ ID NO:51).

[0334] A linear double-copy myoglobin expression cassette was amplified by PCR using primers OT4268 and OT4269 (see, e.g., Table 3) to generate a DNA fragment containing 1 kb of 5' and 3' flanking sequences for Cas9-targeted chromosomal integration at the chromosomal locus. The myoglobin cassette PCR fragment and the Pcbh2-BASI cassette PCR fragment were integrated into the T. reesei genome as described in Example 3. Colonies were screened for myoglobin expression by inoculating 1 mL of NREL medium (pH 6.5) and growing in microtiter plates at 28°C for 4 days without the addition of protease inhibitors. SDS-PAGE was performed to analyze the secreted protein. One of the best-expressing strains (BHX46) was selected for a 1-liter fermentation run, and samples were collected at 43, 72, 94, 114, 137, 161, and 186 hours. Culture supernatants from these samples were analyzed by SDS-PAGE (Figure 10). As shown in Figure 10, myoglobin protein was detected as a 17 kDa band, and BASI protease inhibitor was detected at 20 kDa. Integration of the linear double-copy myoglobin expression cassette at the desired locus was verified by colony PCR amplification and genome sequencing. References PCT Publication Number: International Publication No. 1992 / 03529 PCT Publication Number: International Publication No. 2004 / 035070 PCT Publication Number: International Publication No. 2011 / 153449 PCT Publication Number: International Publication No. 2016 / 100272 PCT Publication Number: International Publication No. 2016 / 100562 PCT Publication Number: International Publication No. 2016 / 100568 PCT Publication Number: International Publication No. 2016 / 100571 PCT Publication Number: International Publication No. 2021 / 092356 PCT Publication Number: International Publication No. 2023 / 278968 U.S. Patent No. 6,255,115 U.S. Patent No. 6,268,328 U.S. Patent No. 7,713,725 U.S. Patent No. 8,936,917 U.S. Patent Application Publication No. 2014 / 0024067 U.S. Patent Application Publication No. 2014 / 0024067 U.S. Patent Application Publication No. 2014 / 0161958 U.S. Patent Application Publication No. 2021 / 0289813 U.S. Patent No. 6,022,725 Cao et al., Science, 9: 991-1001, 2000. Goller et al., “Role of endoproteolytic dibasic proprotein processing in maturation of secretory proteins in Trichoderma reesei”, Appl Environ Microbiol., 64(9):3202-3208, 1998. Krainer et al., “Optimizing cofactor availability for the production of recombinant heme peroxidase in Pichia pastoris”, Microbial Cell Factories, 14:4, 2015. Laskowski and Kato, “Protein Inhibitors of Proteinases”, Ann.Rev.Bioch., 49:593-626, 1980. Le Marquer et al., “Identification of new signaling peptides through a genome-wide survey of 250 fungal secretomes”, BMC Genomics 20:64 2019. Shao et al., “High-level secretory production of leghemoglobin in Pichia pastoris through enhanced globin expression and heme biosynthesis”, Bioresource Technology, Volume 363, 2022. Strickler et al., “Two Novel Streptomyces Protein Protease Inhibitors.Purification, Activity, Cloning and Expression”, The Journal of Biological Chemistry, Vol.267, No.5, pages 3236-3241, 1992. Subramanian et al., “A versatile 2A peptide-based bicistronic protein expressing platform for the industrial cellulase producing fungus, Trichoderma reesei”, Biotechnol for Biofuels 10:34 (2017).

Claims

1. Recombinant filamentous fungal cells expressing heterologous globin proteins.

2. 10. The recombinant cell of claim 1, which expresses and secretes said heterologous globin protein when fermented under suitable conditions.

3. 2. The recombinant cell of claim 1, comprising an introduced expression cassette encoding said globin protein, said cassette comprising an upstream promoter (pro) sequence operably linked to downstream nucleic acid encoding a protein (signal) secretion sequence (sig-seq) operably linked to downstream nucleic acid encoding said globin protein (globin CDS).

4. The recombinant cell of claim 3 , wherein the cassette encoding the globin protein is integrated into the genome of the cell.

5. 2. The recombinant cell of claim 1, comprising an introduced expression cassette encoding at least two globins or comprising at least two introduced expression cassettes encoding at least two globin proteins.

6. 10. The recombinant cell of claim 1, comprising an introduced expression cassette encoding a protease inhibitor protein.

7. 2. The recombinant cell of claim 1, wherein the globin protein expressed is selected from the group consisting of leghemoglobin, myoglobin, hemoglobin, cyanoglobin, and non-symbiotic hemoglobin.

8. 2. The recombinant cell of claim 1, wherein the expressed globin protein comprises a globin superfamily domain and a functional porphyrin (heme) binding site.

9. 10. The recombinant cell of claim 1, wherein the cell is fermented for at least about 96 hours to about 300 hours and produces at least 0.1 grams of globin protein per liter of fermentation broth (g / L).

10. 10. The recombinant cell of claim 1, wherein the cell is fermented for about 180 to about 190 hours and produces at least 1 gram of globin protein per liter of fermentation broth (g / L).

11. 1. A method for producing a heterologous globin protein in a filamentous fungal cell, comprising: (a) introducing into the filamentous fungal cell an expression cassette encoding a globin protein, the cassette comprising an upstream (5') promoter (pro) sequence operably linked to downstream nucleic acid (sig-seq) encoding a protein (signal) secretion sequence operably linked to downstream (3') nucleic acid (globin CDS) encoding the globin protein; and (b) fermenting the modified cells under conditions suitable for the production of said globin protein.

12. The method of claim 11 , wherein the cassette encoding the globin protein is integrated into the genome of the cell.

13. 12. The method of claim 11, wherein the cell comprises an introduced expression cassette encoding at least two globins or comprises at least two introduced expression cassettes encoding at least two globin proteins.

14. The method of claim 11 , wherein the cell comprises an introduced expression cassette encoding a protease inhibitor protein.

15. 12. The method of claim 11, wherein the cells are fermented for at least about 96 hours to about 300 hours to produce at least 0.1 grams of globin protein per liter of fermentation broth (g / L).

16. 12. The method of claim 11, wherein the cells are fermented for about 180 hours to about 190 hours, and the cells produce at least 1 gram of globin protein per liter of fermentation broth (g / L).

17. 12. The method of claim 11, wherein fermenting the modified cells under conditions suitable for production of the globin protein does not require or include a hemin media supplement and / or does not require or include a 5-aminolevulinic acid (ALA) media supplement.

18. 12. The method of claim 11, wherein the secreted globin is recovered from the fermentation broth, and optionally the recovered globin is purified.

19. 12. The method of claim 11, wherein the globin protein is selected from the group consisting of leghemoglobin, myoglobin, hemoglobin, cyanoglobin, and non-symbiotic hemoglobin.

20. 20. The method of claim 19, wherein the leghemoglobin protein comprises at least about 80% identity to the native soybean leghemoglobin protein of SEQ ID NO: 1 or at least about 80% identity to the native kidney bean leghemoglobin protein of SEQ ID NO:

19.

21. 20. The method of claim 19, wherein the myoglobin protein comprises at least 80% identity to the native bovine myoglobin protein of SEQ ID NO:

3.

22. 20. The method of claim 19, wherein the globin protein comprises a globin superfamily domain and a functional porphyrin (heme) binding site.