Genetically engineered yeast strains producing malic acid
Genetically engineered S. pombe strains efficiently produce malic acid by expressing MDH and PYC enzymes with specific promoters and gene modifications, achieving high yields of malic acid production.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ARCHER DANIELS MIDLAND CO
- Filing Date
- 2025-11-07
- Publication Date
- 2026-05-15
AI Technical Summary
There is a need for microorganisms to more efficiently produce C4 dicarboxylic acids, such as malic acid, as traditional petrochemical processes are costly and environmentally unfriendly, and existing yeast fermentation methods require improvements.
Genetically engineered Schizosaccharomyces pombe strains are developed to produce malic acid by expressing exogenous malate dehydrogenase (MDH) and optionally pyruvate carboxylase (PYC) enzymes, with specific promoters and gene modifications to enhance production, including mitochondrial localization of fumarase and inactivation of the mitochondrial citrate transporter.
The engineered S. pombe strains can produce malic acid at high yields, reaching up to 12.2 g/L within 48 hours, predominantly as the 4-carbon diacid, overcoming the inefficiencies of traditional methods.
Smart Images

Figure IMGF000018_0001_TABLE 
Figure IMGF000020_0001_TABLE 
Figure IMGF000021_0001_TABLE
Abstract
Description
[0001] GENETICALLY ENGINEERED YEAST STRAINS PRODUCING MALIC ACID
[0002] REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
[0003] The instant application contains a Sequence Listing, which has been submitted electronically in.xml format. The contents of the electronic sequence listing (BIO_0076_W001_SL.xml; Size 637,102 bytes, and Date of Creation: November 4, 2025) is incorporated by reference herein in its entirety.
[0004] INCORPORATION BY REFERENCE
[0005] All publications, patents, and patent applications cited herein are incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. In the event of a conflict between a term herein and a term in an incorporated reference, the term herein controls.
[0006] BACKGROUND
[0007] Microorganisms have been designed to convert inexpensive raw materials into commercially important products (e.g., biofuels, specialty chemicals, and pharmaceutical compositions). Typically, the microorganisms are engineered, among other things, to efficiently produce and secrete the desired products into the culture medium. Such a manufacturing process is considered advantageous at least for exerting minimized feedback inhibition or toxicity, avoiding product degradation, and simplifying downstream processing.
[0008] Four-carbon (C4) dicarboxylic acids (e.g., succinic, malic, and fumaric acid), also called 4-carbon diacid, are widely used in the food, pharmaceutical, and polymers industries. Traditionally, these acids have been produced through petrochemical processes, which are costly and environmentally unfriendly. Recently, the production of succinic and malic acids by yeast fermentation has been developed and commercialized (e.g., Reverdia and Bio Amber).
[0009] Nevertheless, there continues to be a need for developing microorganisms to more efficiently produce C4 dicarboxylic acids, e.g., malic acid. SUMMARY OF THE INVENTION
[0010] The present disclosure provides genetically engineered Schizosaccharomyces pombe strains capable of fermenting glucose to produce malic acid and the methods of producing malic acid from the S. pombe strains. The S. pombe strains are genetically engineered to produce an exogenous malate dehydrogenase (MDH) enzyme, and optionally produce an exogenous pyruvate carboxylase (PYC) enzyme or an exogenous phosphoenolpyruvate carboxylase (PPC) enzyme. Additionally, the S. pombe strains may be further engineered to produce malic acid as the predominant C4 dicarboxylic acid, which can be achieved by (i) having the 5. pombe strains produce a fumarase selectively located in mitochondria and / or (ii) inactivating the mitochondrial citrate transporter in the S. pombe strains.
[0011] In an aspect, provided is a genetically engineered Schizosaccharomyces pombe strain comprising a nucleic acid encoding an exogenous malate dehydrogenase (MDH) enzyme that is operably linked to a first promoter to express the MDH enzyme in the S. pombe strain, wherein the S'. pombe strain is capable of fermenting glucose to produce malic acid. The MDH enzyme may comprise an amino acid sequence with at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to an amino acid sequence selected from SEQ ID NOs: 29-142. Alternatively, the MDH enzyme may comprise an amino acid sequence selected from SEQ ID NOs: 111, 119-121, 124, and 130. The MDH enzyme may comprise an amino acid sequence selected from SEQ ID NOs 121, 124, and 130. The first promoter may be an ACT! promoter (PACTI), a PYK1 promoter (PPYKI), or a ZYM1 promoter (PZYMI). The ACT1 promoter (PACTI) may comprise a nucleotide sequence of SEQ ID NO: 24 or an operably functional portion thereof. The PYK1 promoter (PPYKI) may comprise a nucleotide sequence of SEQ ID NO: 25 or an operably functional portion thereof. The ZYM1 promoter (PZYMI) may comprise a nucleotide sequence of SEQ ID NO: 26 or an operably functional portion thereof.
[0012] The nucleic acid encoding the MDH enzyme may be integrated at a locus of either the inactivated ADH1 gene or the inactivated ADH4 gene. Alternatively, the S. pombe strain may comprise two copies of the nucleic acid encoding the MDH enzy me.
[0013] In another aspect, the 5. pombe strain may further compnse (i) a nucleic acid encoding an exogenous pyruvate carboxylase (PYC) enzyme that is operably linked to a second promoter to express the PYC enzy me in the S. pombe strain, or (ii) a nucleic acid encoding an exogenous phosphoenolpyruvate carboxylase (PPC) enzyme that is operably linked to a third promoter to express the PPC enzyme in the S. pombe strain. The PYC enzy me may comprise an amino acid sequence with at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to an amino acid sequence selected from SEQ ID NOs: 143-228. The PPC enzyme may comprise an amino acid sequence with at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to an amino acid sequence selected from SEQ ID NOs: 229-371. Alternatively, the PYC enzyme may compnse an amino acid sequence selected from SEQ ID NOs: 154, 156, 166, 168, and 176. The PPC enzyme may comprise an amino acid sequence selected from SEQ ID NOs: 245, 253, 259, 266, 271, 275, 283, 289, 293, 311, 320, 337, 346, and 348-350. The PYC enzyme may comprise an amino acid sequence selected from SEQ ID NOs: 154, 156. and 166. The PPC enzyme may comprise an amino acid sequence selected from SEQ ID NOs: 245, 253, 266, 271, 275, 283, 311, and 348-350. Each of the second promoter and the third promoter may be an ACT1 promoter (PACTI), a PYK1 promoter (PPYKI), or a ZYM1 promoter (PZYMI). Each of the nucleic acid encoding the PYC enzyme and the nucleic acid encoding the PPC enzyme may be integrated at a locus of either the inactivated ADH 1 gene or the inactivated ADH4 gene. The S. pombe strain may further comprise two copies of the nucleic acid encoding the PYC enzyme or two copies of the nucleic acid encoding the PPC enz me.
[0014] In a further aspect, the S. pombe strain may further comprises one or more features selected from:
[0015] (i) an inactivated alcohol dehydrogenase 1 (ADH1) gene;
[0016] (ii) an inactivated alcohol dehydrogenase 4 (ADH4) gene;
[0017] (iii) an inactivated glycerol-3-phosphate dehydrogenase 1 (GPD1) gene;
[0018] (iv) an inactivated pyruvate decarboxylase 1 (PDC201) gene; and
[0019] (v) an inactivated malic enzyme (MAE2) gene.
[0020] The S. pombe strain may produce a fumarase selectively located in mitochondria. The fumarase may comprise a mitochondrial localization sequence (MLS) from a succinate-CoA ligase beta subunit or a mitochondrial DNA polymerase gamma subunit. For example, the fumarase may comprise an amino acid sequence with at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 388 or 390.
[0021] In yet another aspect, the S. pombe strain may comprise an inactivated mitochondrial citrate transporter (YHM2) gene.
[0022] The provided genetically engineered Schizosaccharomyces pombe strain may comprise:
[0023] (i) a nucleic acid encoding an exogenous malate dehydrogenase (MDH) enzyme that is operably linked to a first promoter to express the MDH enzyme in the S'. pombe strain;
[0024] (ii) a nucleic acid encoding an exogenous pyruvate carboxylase (PYC) enzyme that is operably linked to a second promoter to express the PY C enzyme in the S. pombe strain or a nucleic acid encoding an exogenous phosphoenolpyruvate carboxylase (PPC) enzyme that is operably linked to a third promoter to express the PPC enzyme in the S. pombe strain;
[0025] (iii) an inactivated alcohol dehydrogenase 1 (ADH1) gene; (iv) an inactivated alcohol dehydrogenase 4 (ADH4) gene; (v) an inactivated glycerol-3 -phosphate dehydrogenase 1 (GPD1) gene; (vi) an inactivated pyruvate decarboxylase 1 (PDC201) gene; and (vii) an inactivated malic enz me (MAE2) gene,
[0026] wherein the S. pombe strain is capable of fermenting glucose to produce malic acid. The S. pombe strain may produce a fumarase selectively located in mitochondria or comprise an inactivated mitochondrial citrate transporter (YHM2) gene.
[0027] A culture of the S. pombe strain may be capable of producing at least 12.2 g / L malic acid after 48 hours culture in a baffled shake flask incubated at 33 °C. A culture of the S’. pombe strain after 48 hours culture in a baffled shake flask incubated at 33 °C may be capable of producing a 4-carbon diacid cocktail comprising about 90% or more malic acid.
[0028] In a further aspect, the S'. pombe strain selected from the group consisting of:
[0029] S'. pombe Sk41 deposited as NRRL Y-68447;
[0030] S. pombe Sk75 deposited as NRRL Y-68448;
[0031] S'. pombe Sk94 deposited as NRRL Y-68449;
[0032] S', pombe Sk220 deposited as NRRL Y-68450; S'. pombe Sk223 deposited as NRRL Y -68451;
[0033] S. pombe Sk468 deposited as NRRL Y-68452;
[0034] S. pombe Sk471 deposited as NRRL Y-68453;
[0035] S. pombe Sk513 deposited as NRRL Y-68463;
[0036] S'. pombe Sk528 deposited as NRRL Y-68454; and
[0037] S. pombe Sk530 deposited as NRRL Y-68455.
[0038] In yet another aspect, also provided is a method of producing malic acid comprising the steps of:
[0039] culturing a cell of the S’, pombe strain of any one of claims 1 to 25 in the presence of a carbon source under conditions suitable for producing at least 12.2 g / L malic in 48 hours, and
[0040] isolating malic acid from the culture.
[0041] The carbon source used for culturing the S’, pombe strain may be selected from glucose, sucrose, fructose, a glucose oligomer, raffinose, glycerol, starch, and any combination thereof.
[0042] Further provided is (i) an isolated nucleic acid encoding an MDH enzy me having an amino acid sequence selected from SEQ ID NOs: 29-142, wherein the nucleic acid sequence is operably linked to an ACT1 promoter (PACTI), a PYK1 promoter (PPYKI), or a ZYM1 promoter (PZYMI); (ii) an isolated nucleic acid encoding a PYC enzyme having an amino acid sequence selected from SEQ ID NOs: 143-228, wherein the nucleic acid sequence is operably linked to an ACT1 promoter (PACTI), a PYK1 promoter (PPYKI), or a ZYM1 promoter (PZYMI); and (iii) an isolated nucleic acid encoding a PPC enzyme having an amino acid sequence selected from SEQ ID NOs: 229-371, wherein the nucleic acid sequence is operably linked to an ACT1 promoter (PACTI), a PYK1 promoter (P YKI), or a ZYM1 promoter (PZYMI).
[0043] The processes disclosed herein may be used to produce malic acid from S. pombe with a high yield or produce a 4-carbon diacid cocktail acid from S. pombe predominantly comprising malic acid.
[0044] BRIEF DESCRIPTION OF THE FIGURES FIG. 1 illustrates a comparison among strains Sk5, Sk41, Sk75, and Sk94 in malic acid production. The culture conditions for the strains are described in Example 2. The data presented suggest that increasing copy number of heterologous genes PPC17 and MDH50 increased the biological production of malic acid.
[0045] FIG. 2 illustrates a phylogenetic tree of MDH enzy me library as described in Example 3.
[0046] FIG. 3 illustrates in vitro malate dehydrogenase (MDH) activities of various enzyme variants as described in Example 3.
[0047] FIG. 4 illustrates in vivo malate dehydrogenase (MDH) activities of various enzyme variants as described in Example 3.
[0048] FIG. 5 illustrates a phylogenetic tree of the PYC enzyme library as described in Example 3.
[0049] FIG. 6 illustrates in vivo pyruvate carboxylase (PYC) activities of various enzyme variants as described in Example 3.
[0050] FIG. 7 illustrates a phylogenetic tree of the PPC enzyme library as described in Example 3.
[0051] FIG. 8 illustrates in vitro phosphoenolpyruvate carboxylase (PPC) activities of various enzyme variants in the absence of Acetyl-CoA as described in Example 3.
[0052] FIG. 9 illustrates in vitro phosphoenolpyruvate carboxylase (PPC) activities of various enzyme variants in the presence of Acetyl-CoA as described in Example 3 FIG. 10 illustrates in vitro phosphoenolpyruvate carboxylase (PPC) activities of selected enzy me variants in the presence of malate as described in Example 3.
[0053] FIG. 11 illustrates malic acid production from selected strains expressing various PPCs as described in Example 3.
[0054] FIG. 12 illustrates C4 dicarboxylic acid production from various strains as described in Example 4.
[0055] DETAILED DESCRIPTION
[0056] The following are abbreviations used in the present disclosure:
[0057] ADH alcohol dehydrogenase
[0058] PDC pyruvate decarboxylase
[0059] GPD glycerol-3-phosphate dehydrogenase
[0060] MAE2 malic enzy me
[0061] MDH malate dehydrogenase
[0062] PYC pyruvate carboxylase PPC phosphoenolpyruvate carboxylase
[0063] 2,3-BDO 2,3-Butanedio
[0064] 5-FOA 5-fluoroorotic acid
[0065] EMM Edinburgh Minimal Medium
[0066] CRISPR clustered regularly interspaced short palindromic repeats LDH lactate dehydrogenase
[0067] CFUs colony forming units
[0068] CSL com steep liquor
[0069] OAA oxaloacetate
[0070] ATP adenosine triphosphate
[0071] HPLC high-performance liquid chromatography
[0072] MSA multiple sequence alignment
[0073] TCA pathway tricarboxylic acid cycle, citric acid cycle, or the Krebs cycle
[0074] MLS mitochondrial localization sequence
[0075] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs. See, e g., Singleton et al., DICTIONARY OF MICROBIOLOGY AND MOLECULAR BIOLOGY 2nd ed., J. Wiley & Sons (New York, N. Y. 1994); Sambrook et al., MOLECULAR CLONING, A LABORATORY MANUAL, Cold Spring Harbor Press (Cold Spring Harbor, N. Y. 1989).
[0076] The use of any examples, or exemplary language (e.g. "such as”) provided for certain embodiments, herein is intended to better illuminate the present disclosure and does not pose a limitation on the scope of the present disclosure otherwise claimed. This disclosure is not limited to the particular methodology, protocols, reagents, etc., described herein and as such can vary. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the present disclosure. The terminology used herein is to describe particular embodiments and is not intended to limit the scope of the present disclosure, which is defined solely bv the claims.
[0077] As used in the specification, embodiments, and claims, the singular forms "a.” “an” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a cell” includes a plurality of cells, including mixtures thereof. Similarly, use of “a compound” in a method as described herein contemplates using one or more compounds of this invention for such method unless the context clearly dictates otherwise.
[0078] As used herein, the term “comprising” is intended to mean that the compositions and methods include the recited elements, but not excluding others. “Consisting essentially of’ when used herein to define compositions and methods, shall mean excluding other elements of any essential significance to the combination. Thus, a method consisting essentially of the recited steps as defined herein would not exclude additional steps that do not alter the method from yielding the claimed result. “Consisting of’ shall mean in the context of a method excluding additional steps and limiting the claim only to the recited steps. Embodiments defined by each of these transition terms are within the scope of this invention.
[0079] The term “about” in relation to a reference numerical value, and its grammatical equivalents as used herein, can include the reference numerical value itself and a range of values plus or minus 10% from that reference numerical value. For example, the term “about 10” includes 10 and any amount from and including 9 to 11. In some cases, the term “about” in relation to a reference numerical value can also include a range of values plus or minus 10%, 9%, 8%, 7%, 6%. 5%, 4%, 3%, 2%, or 1% from that reference numerical value. In some embodiments, “about” in connection with a number or range measured by a particular method indicates that the given numerical value includes values determined by the variability of that method.
[0080] As used herein, any concentration range, percentage range, ratio range, or integer range is to be understood to include the value of any integer within the recited range and, when appropriate, fractions thereof (such as one-tenth or one-hundredth of an integer), unless otherwise indicated. In addition, all ranges are intended to expressly include the boundaries of the range individually. For clarity, the range 3-6 is intended to include individually 3, 4, 5. and 6 as well as any fraction (e.g., about 0.1 fraction) within that range.
[0081] Polypeptides and nucleic acids have a form of directionality when discussed in certain orientations. A nucleic acid is discussed in a 5‘ to 3’ direction, or sense direction, which relates for example, to when a nucleic acid is translated into a polypeptide. A polypeptide is described in an amino-terminal (N-terminal or N-terminus) to carboxy -terminal (C-terminal or C-terminus) orientation.
[0082] The term “heterologous” or “exogenous” when used with respect to a nucleic acid (DNA or RNA) or protein refers to a nucleic acid or protein that does not occur naturally as part of the organism, cell, genome or DNA or RNA sequence in which it is present, or that is found in a cell or location or locations in the genome or DNA or RNA sequence that differ from that in which it is found in nature. Heterologous or exogenous nucleic acids or proteins are not endogenous to the cell into which it is introduced but have been obtained from another cell or synthetically or recombinantly produced.
[0083] The term ‘'gene”, as used herein, refers to a nucleic acid sequence containing a template for a nucleic acid polymerase, in eukaryotes, RNA polymerase II. Genes are transcribed into mRNAs that are then translated into proteins.
[0084] The term "nucleic acid” as used herein, includes reference to a deoxyribonucleotide or ribonucleotide polymer, i.e. a polynucleotide, in either single- or double-stranded form, and unless otherwise limited, encompasses known analogs having the essential nature of natural nucleotides in that they hybridize to single-stranded nucleic acids in a manner similar to naturally occurring nucleotides (e.g.. peptide nucleic acids). A polynucleotide can be full-length or a subsequence of a native or heterologous structural or regulatory gene. Unless otherwise indicated, the term includes reference to the specified sequence as well as the complementary' sequence thereof.
[0085] The terms "polypeptide”, "peptide”, and "protein” are used interchangeably herein to refer to a polymer of amino acid residues. The terms apply to amino acid polymers containing one or more artificial chemical analogs, each of which corresponds to a naturally occurring amino acid residue. The essential nature of such analogs of naturally occurring amino acids is that, when incorporated into a protein, that protein is specifically reactive to antibodies elicited to the same protein but consisting entirely of naturally occurring amino acids. The terms “polypeptide”, “peptide”, and '‘protein” are also inclusive of modifications including, but not limited to, glycosylation, lipid attachment, sulfation, gamma-carboxylation of glutamic acid residues, hydroxylation, and ADP-ribosylation. A polypeptide can be a fragment of an immature or mature protein, for example, encoded by a gene.
[0086] The term “enzyme” as used herein refers to a protein capable of catalyzing a (bio)chemical reaction in a cell.
[0087] Polynucleotide or polypeptide sequences can be compared by performing a sequence alignment, which may be gapped or ungapped. In an ungapped alignment, two or more sequences are compared as “contiguous” sequences, i.e., one sequence is aligned with the other sequence and each amino acid or nucleotide in one sequence is directly compared with the corresponding amino acid or nucleotide in the other sequence, one residue at a time. In an ungapped alignment, in an otherwise identical pair of sequences, one insertion or deletion may cause the other nucleotide or amino acid residues to be put out of alignment, thus resulting in a potentially non-optimal global alignment. In a gapped alignment, sequences are compared “non-contiguously”, and insertions and deletions (collectively “gaps”) may be inserted to optimally align the sequences.
[0088] The “percentage identity” between polypeptide sequences can be calculated using commercially available algorithms that compare a reference sequence with a query sequence. In some embodiments, polypeptides are 70%, at least 70%, 75%, at least 75%, 80%, at least 80%, 85%, at least 85%, 90%, at least 90%, 92%, at least 92%, 95%, at least 95%, 97%, at least 97%, 98%, at least 98%, 99%, or at least 99% or 100% identical to a reference polypeptide, or a fragment thereof (e.g., as measured by BLASTP or CLUSTAL, or other alignment software) using default parameters. Similarly, nucleic acids can also be described with reference to a starting nucleic acid, e.g., they can be 50%, at least 50%, 60%, at least 60%, 70%, at least 70%, 75%, at least 75%, 80%, at least 80%, 85%, at least 85%, 90%, at least 90%, 95%, at least 95%, 97%, at least 97%, 98%, at least 98%, 99%, at least 99%, or 100% identical to a reference nucleic acid or a fragment thereof (e.g., as measured by BLASTN or CLUSTAL, or other alignment software using default parameters). When one molecule is said to have a certain percentage of sequence identity with a larger molecule, it means that when the two molecules are optimally aligned, the percentage of residues in the smaller molecule finds a match residue in the larger molecule in accordance with the order by which the two molecules are optimally aligned, and the “%” (percent) identity is calculated in accord with the length of the smaller molecule.
[0089] Methods to determine identity and similarity' are codified in publicly available algorithms or software, and can include but are not limited to: BLAST, FASTA. T-COFFEE, or M-COFFEE. In some methods, a scaled similarity score matrix or equivalent can be used to assign a score to each pairwise comparison based on chemical similarity or evolutionary distance. An example of such a matrix commonly used is the BLOSUM62 matrix — the default matrix for the BLAST suite of programs. There are also alternative computational methods used to determine identity or similarity (e.g., INFERNAL or R-COFFEE) which also consider aspects of sequence relatedness (e.g., covariance models, secondary structure, or tertiary' structure) in addition to the primary sequence (Eddy & Durbin. Nucleic Acids Research (1994); DOI: 10.1093 / nar / 22.1 1.2079) (Rivas et al., Bioinformatics (2020); DOI: 10.1093 / bioinformatics / btaa080) (Nawrocki & Eddy; Bioinformatics (2013); DOI: 10.1093 / bioinformatics / btt509).
[0090] In some embodiments, the engineered S. pombe strain may produce an MDH enzyme having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to an MDH having an amino acid sequence selected SEQ ID NOs: 29-142. In other embodiments, the MDH enzyme may have at least 90%, 91%. 92%. 93%. 94%. 95%, 96%, 97%, 98%, or 99% sequence identity to an MDH having an amino acid sequence selected from SEQ ID NOs: 111, 119-121, 124, and 130.
[0091] In some embodiments, the engineered S. pombe strain may produce a PY C enzyme having at least 80%, 81%, 82%, 83%. 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%. 96%. 97%. 98%. or 99% sequence identity to a PYC having an amino acid sequence selected SEQ ID NOs: 143-288. In other embodiments, the PYC enzyme may have at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to a PYC having an amino acid sequence selected from SEQ ID NOs: 154, 156, 166, 168. and 176.
[0092] In some embodiments, the engineered S. pombe strain may produce a PPC enzyme having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to a PPC having an amino acid sequence selected SEQ ID NOs: 229-371. In other embodiments, the PPC enzyme may have at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to a PPC having an amino acid sequence selected from SEQ ID NOs: 245, 253, 259, 266, 271, 275, 283, 289, 293, 311, 320, 337, 346, and 348-350.
[0093] Usually, a nucleotide sequence encoding an enzyme is operably linked to a promoter that causes sufficient expression of the corresponding nucleotide sequence in the eukaryotic cell (e.g., a Schizosaccharomyces pombe cell) according to the present disclosure to confer to the cell the ability’ to produce an intended product.
[0094] As used herein, the term ‘‘operably linked” refers to a juxtaposition wherein the components described are in a relationship permitting them to function in their intended manner. A control sequence “operably linked” to a coding sequence is ligated in such a way that expression of the coding sequence is achieved under conditions compatible with the control sequences. For example, a promoter is said to be operably linked to a gene, open reading frame, or coding sequence, if the linkage or connection allows or effects transcription of said gene. In a further example, a 5’ and a 3’ gene, cistron, open reading frame, or coding sequence are said to be operably linked in a polycistronic expression unit, if the linkage or connection allows or effects translation of at least the 3’ gene. For example, DNA sequences, such as, e.g.. a promoter and an open reading frame, are said to be operably linked if the nature of the linkage between the sequences does not (1) result in the introduction of a frame-shift mutation, (2) interfere with the ability' of the promoter to direct the transcription of the open reading frame, or (3) interfere with the ability of the open reading frame to be transcribed by the promoter region sequence.
[0095] As used herein, the term '‘promoter” refers to a nucleic acid fragment that functions to control the transcription of one or more genes, located upstream (e.g., 5’ to) with respect to the direction of transcription of the transcription initiation site of the gene, and is structurally identified by the presence of a binding site for DNA-dependent RNA polymerase, transcription initiation sites and any other DNA sequences known to one of skilled in the art. A “constitutive” promoter is a promoter that is active under most environmental and developmental conditions. An “inducible” promoter is a promoter that is active under environmental or developmental regulation.
[0096] A promoter that could be used to achieve expression of a nucleotide sequence encoding an enzyme (e.g., the MDH, PYC, and / or PPC described herein) may be not native to the nucleotide sequence coding for the enzy me to be expressed, i.e. a promoter that is heterologous to the nucleotide sequence (coding sequence) to which it is operably linked. Alternatively, the promoter is homologous, i.e. endogenous to the host cell.
[0097] The promoters that are useful to express the gene encoding the exogenous MDH, PYC, or PPC enzymes in the S. pombe strains of this disclosure are not particularly limited. Representative promoters may include an ACT1 promoter (PACTI) having a nucleotide sequence of SEQ ID NO: 24 or an operably functional portion thereof, a PYK1 promoter (PPYKI) comprises a nucleotide sequence of SEQ ID NO: 25 or an operably functional portion thereof, and a ZYM1 promoter (PZYMI) comprises a nucleotide sequence of SEQ ID NO: 26 or an operably functional portion thereof. The “operably function portion” as described herein refers to a nucleotide sequence containing one or more substitutions, insertions, and / or deletions of the reference nucleotide sequence, while maintaining substantially similar (e.g., 80% or more) promoter activity7as the reference nucleotide sequence. Cells of the engineered A’, pombe strains of this disclosure can be cultivated to produce malic acid, either in the free acid form or in salt form (or both), or a metabolization product of malic acid. The recombinant cell can be cultured in a medium that includes at least one carbon source that can be fermented by the cell. Examples of suitable carbon sources for use in culturing include, but are not limited to, twelve carbon sugars such as sucrose, hexose sugars such as glucose or fructose, glycan, starch, or other glucose polymers, glucose oligomers such as maltose, maltotriose and isomaltotriose, panose, and fructose oligomers, and pentose sugars such as xylose, xylan, other oligomers of xylose, or arabinose. In some embodiments, the carbon source may be selected from glucose, sucrose, fructose, a glucose oligomer, raffinose, glycerol, starch, and any combination thereof.
[0098] The culture medium typically may contain, in addition to the carbon source, nutrients as required by the engineered S. pombe cells, including a source of nitrogen (such as amino acids, proteins, inorganic nitrogen sources such as ammonia or ammonium salts, and the like), and various vitamins, minerals and the like. In some embodiments, the cells of the engineered S. pombe strain can be cultured in a chemically defined medium.
[0099] Other cultivation conditions, such as temperature, cell density, selection of substrate(s), selection of nutrients, and the like ca be generally selected and optimized to provide an economical process. Temperatures during each of the growth phases and the production phase(s) may range from above the freezing temperature of the medium to about 50 °C. depending on the ability of the strain to tolerate elevated temperatures. A typical culture temperature, particularly during the production phase, is about 30 °C to 45 °C.
[0100] During cultivation, aeration and agitation conditions may be selected to produce a desired oxygen uptake rate. The cultivation may be conducted aerobically, microaerobically, or anaerobically, depending on pathway requirements. In some embodiments, the cultivation conditions may be adjusted to produce an oxygen uptake rate in the range of about 2-25 mmol / L / hour, in the range of about 5-20 mmol / L / hour, or in the range of about 8-15 mmol / L / hour.
[0101] The culturing process may be divided up into phases. For example, the cell culture process may be divided into a cultivation phase, a production phase, and a recovery phase. The pH may be allowed to range freely during cultivation or may be buffered if necessary' to prevent the pH from falling below or rising above predetermined levels. For example, the medium may be buffered to prevent the pH of the solution from falling below around 2.0 or above about 8.0 during cultivation. In some of these embodiments, the medium may be buffered to prevent the pH of the solution from falling below about 3.0 or rising above about 7.0. Alternatively, in some of these embodiments, the medium may be buffered to prevent the pH of the solution from falling below about 3.0 or rising above about 4.5. Suitable buffering agents include basic materials that neutralize the acid as it is formed, and include, for example, calcium hydroxide, calcium carbonate, sodium hydroxide, potassium hydroxide, potassium carbonate, sodium carbonate, ammonium carbonate, ammonia, ammonium hydroxide, or any combination thereof.
[0102] In a buffered fermentation, acidic fermentation products (e.g., malic acid) are neutralized to the corresponding salt as they are formed. Recovery of the acidic product therefore may involve regenerating the free acid. Acidic product is typically achieved by removing the cells and acidulating the fermentation broth with a strong acid such as sulfuric acid. A salt by-product is formed (gypsum in the case where a calcium salt is the neutralizing agent and sulfuric acid is the acidulating agent), which is separated from the broth.
[0103] The following examples are exemplary only with additional modifications to the methods and products obtained thereby in accordance with the teachings herein and known in the art suitable for use in modifying the examples. The Examples should not be viewed as limiting to the methods, compounds, and compositions described herein.
[0104] EXAMPLES
[0105] Example 1: Construction of yeast strains producing malic acid
[0106] The construction of Schizosaccharomyces pombe strains that can ferment dextrose to malic acid started from strain Sp285, which is an engineered S. pombe strain producing lactic acid. Strain Sp285 was described in PCT / US2022 / 073915 (entitled “A GENETICALLY ENGINEERED YEAST PRODUCING LACTIC ACID) filed on July 20, 2022, and published on January' 26, 2023, as WO 2023 / 004336, which is incorporated by reference in its entirety for all purposes. Strain Sp285 has been deposited in the United States Department of Agriculture strain depository as NRRL B-6805. As described, Strain Sp285 has the engineered genotype of pdc201A:: PACTi-LcLDH ura4\:: BC4241 adhl \:: P ACTI -LcLDH adh4A:: BC59 gpdlA:: URA4. SEQ ID NO: 16 is the coding nucleotide sequence of ADH1, while SEQ ID NO: 8 is the amino acid sequence of ADH1. SEQ ID NO: 17 is the coding nucleotide sequence of ADH4, while SEQ ID NO: 9 is the amino acid sequence of ADH4. SEQ ID NO: 20 is the nucleotide sequence of PDC201, while SEQ ID NO: 12 is the amino acid sequence of PDC201.
[0107] Strain Sp285 has been performance-evaluated at commercial production scales and routinely achieves commercial KPIs at 7500L fermentations. Strain Sp285 makes 97.5% pure lactic acid with all other metabolites ethanol, glycerol, acetate, pyruvate, and 2,3-BDO comprising 2.5% of the final fermentation broth. The lactic acid strain was engineered to eliminate these by-products through targeted gene deletion, heterologous gene expression, attenuated enzyme activity, and adapted laboratory evolution as described herein.
[0108] For this disclosure, the prefix “Sk” is used for all engineered strains. Since strain Sp285 is the parent for all of the engineered strains described herein, strain Sp285 was given the alternate name of strain Sk2.
[0109] Strain Sk2.5 was built from strain Sk2 by replacing the URA4 selection marker at the GPD1 chromosomal locus with the non-coding molecular barcode (BC4010) sequence: 5’-CGCGTGGGAAGTATTATATCTCGACCATTTCTTAGCAGTG-3’ (SEQ ID NO: 6). SEQ ID NO: 18 is the coding nucleotide sequence of GPD1, while SEQ ID NO: 10 is the amino acid sequence of GPD1. This gene replacement was mediated by CRISPR-Mad7 genome editing tools using the PAM-Protospacer sequence for URA4, which is 5 -TTTGTGATATGAGCCCAAGAAGCAA-3’ (SEQ ID NO: 5). Transformants were recovered on EMM + uracil + 5 -fluorooro tic acid (5-FOA) media to select for strains that lost the URA4 gene.
[0110] Next, strain Sk3 was built from strain Sk2.5 by deleting the MAE2 / SPCC794.12c gene (SEQ ID NO: 22: coding nucleotide sequence; SEQ ID NO: 14: amino acid sequence). It was rationalized that MAE2 must be deleted in a malic acid-producing yeast strain because MAE2 encodes a malic enzyme (a.k.a. malate dehydrogenase (oxaloacetate decarboxylating) enzyme) that converts malic acid to pyruvic acid. A functional MAE2 gene in a malic acid-producing yeast strain would result in a futile cycle of malic acid production by the synthetic pathway and malic acid consumption via Mae2p. The deletion of the MAE2 / SPCC794.12c gene was achieved via homology directed repair (HDR) with a URA4 selection marker (SEQ ID NO: 21: coding nucleotide sequence; SEQ ID NO: 13: amino acid sequence). Transformants were recovered using EMM - uracil to select for uracil prototrophic strains that had gained the URA4 gene via transformation. Deletion of MAE2 was confirmed by diagnostic PCR using standard PCR procedures.
[0111] Strain Sk4 was built from Sk3 by replacing the URA4 selection marker wi th a noncoding DNA molecular barcode sequence (BC4298) sequence: 5’-ACTCCATCAAATCACTAGGGACGTACAAATACATCTCCGG-3' (SEQ ID NO: 7) by CRISPR-Mad7 genome editing. The PAM-Protospacer sequence for URA4, which is 5’-TTTGTGATATGAGCCCAAGAAGCAA-3’ (SEQ ID NO: 5) was used to target and replace the URA4 marker. The genotype of strain Sk4 was confirmed by diagnostic PCR using standard procedures. Strain Sk4 was confirmed to be a uracil auxotroph and a 5-FOA resistant strain.
[0112] Furthermore, strain Sk5 was built from strain Sk4 by deleting one copy of the heterologous Lactobacillus cerevisiae L-lactate dehydrogenase gene (also referred to as LDH58) (SEQ ID NO: 19: coding nucleotide sequence; SEQ ID NO: 11: amino acid sequence). Strain Sk2 contains two copies of L. cerevisiae LDH, which were introduced as descnbed in PCT / US2022 / 073915 (entitled “A GENETICALLY ENGINEERED YEAST PRODUCING LACTIC ACID ) filed on July 20, 2022, and published on January' 26, 2023, as WO 2023 / 004336, which is incorporated herein by reference in its entirety for all purposes. LDH58 was replaced in strain Sk4 with a URA4 selection marker at the PDC201 locus using CRJSPR-Mad7. The PAM-Protospacer for LDH58 is 5’-TTTGATGGCTAAATCCGCTGAAACC-3’ (SEQ ID NO: 1). Strain Sk5 was confirmed by diagnostic PCR using standard procedures. Strain Sk5 is further confirmed as a uracil prototroph and a 5-FOA-sensitive strain. Since the LDH58 was replaced with a URA4 selection marker at PDC201, strain Sk5 contains only one copy of LDH58 at the ADH1 locus. To fully eliminate the production of lactic acid and to test whether these modified cells could produce malic acid, the second copy of LDH58 was deleted by replacing it with heterologous malic acid pathway genes, phosphoenolpyruvate carboxylase (PPC) and malate dehydrogenase (MDH).
[0113] Strain Sk41 was constructed from strain Sk5 by deleting LDH58 using the PAM-Protospacer sequence 5’-TTTGATGGCTAAATCCGCTGAAACC-3’ (SEQ ID NO: 1) in a one-step CRISPR-based replacement using a repair DNA template that contained a sequence encoding a PPC and a sequence encoding MDH. Both the lactic acid and malic acid metabolic pathways are NAD(H) redox balanced with glycolysis. For every glucose molecule metabolized, two molecules of NADH are generated by glyceraldehyde 3-phosphate dehydrogenase (GPD); the two NADH molecules are then used by LDH or MDH to regenerate NAD+. Under microaerobic conditions, such as the inside of a yeast colony on an agar plate, only cells with a functional malic acid pathway will be NAD(H) redox balanced with glycolysis. That means that colony-forming units (CFUs) will only be recovered if the cell was transformed with a functional PPC and MDH. Given this scientific reasoning, LDH58 was targeted for replacement with CRISPR-Mad7 using a repair DNA template that contained different versions of PPC DNA expression cassettes and MDH expression cassettes. One transformant that resulted in a large CFU contained the PPC 17 (sequence derived from Chlamydomonas reinhardlii) gene, encoding the amino acid sequence of SEQ ID NO: 245, and MDH50 (sequence derived from Gibberella moniliformis)' gene, encoding the amino acid sequences of and SEQ ID NO: 78. This strain was named strain Sk41. Production of malic acid by Sk41 was confirmed in shake-flask fermentations (see below).
[0114] Strain Sk75 contains a second copy of the PPC17-MDH50 expression cassette at the ADH1 chromosomal locus. Sk75 was built from Sk41 by deleting SPAPB1 Al 1.03 using the PAM-Protospacer sequence 5’-TTTGAGAACAATTGGGCCATCCCAA-3’ (SEQ ID NO: 2) with a second copy of PPC17-MDH50 that contains homology to the SPAPB1A11.03 chromosomal locus. SPAPB1A11.03 encodes an FMN-dependent alpha-hydroxy acid dehydrogenase; its deletion has no apparent effect on malic acid production in cells.
[0115] Strain Sk94 was built from Sk75 by deleting PDC202 using the PAM-Protospacer sequence 5’-TTTGATCACCGAGACTGGAACTGCA-3’ (SEQ ID NO: 3) with a repair DNA molecule that contained a third copy of MDH. SEQ ID NO: 23 is the coding nucleotide sequence of PDC202, while SEQ ID NO: 15 is the amino acid sequence of PDC202.
[0116] Table 1 summarizes the genotypes of the “Sk’‘ strains described and constructed herein. Table 1.
[0117]
[0118] Example 2: Characterization of malic acid production by various strains
[0119] Cells of Sk5, Sk41. Sk74, and Sk94 strains were grown separately in a baffled shake flask containing 5 g / L Roquette com steep liquor (CSL); 25 g / L yeast extract (Neogen, Lansing MI); and 30 g / L dextrose. The culture volume was 100 rnL in a 500 mL baffled shake-flask and incubated at 33 °C. The growth phase shake-flask was incubated in an Infers HT shaker at 350 rpm for 24 hours. Seed cultures were cultivated for 24 hours at which point the cells were harvested by centrifugation, washed with sterile water and resuspended in the fermentation / main media to an OD of 1.0. The production flask conditions are: 15 g / L Roquette CSL; 50 g / L dextrose; 1.0 mg / L biotin; and 50 g / L calcium carbonate. The initial OD600 of the production shake-flask was 1.0. The culture volume was 10 mL in a 25 mL unbaffled shake-flask, which was incubated at 33 °C, at 250 rpm for 48 hours. Supernatant malic acid concentrations were measured with high-performance liquid chromatography (HPLC) and illustrated in FIG. 1. The data in FIG. 1 suggest that increasing the copy number of heterologous genes PPC17 and MDH50 increased the biological production of malic acid in the strains tested.
[0120] Example 3, Screening and characterizing various enzyme variants for malic acid production in engineered S. pombe cells.
[0121] To further optimize the malic acid production in the strains, a library approach was employed to identify a diverse set of enzy mes that can perform the required bioconversions within an S. pombe cell to synthesize malic acid. For this, bioinformatics was first used to identify a diverse set of enzyme variants for PYC (EC 6.4.1.1), PPC (EC 4.1.1.31), and MDH (EC 1.1.1.37). Second, each enzyme was evaluated in vitro for its purported activity. In vitro activity data were used to qualify each enzyme for expression in S. pombe cells. Third, the top-ranking enzy mes were expressed in S. pombe cells and quantified for malic acid production. Enzymes that result in relatively high malic acid production in an S. pombe cell are the most likely enzymes that will be present in the commercial production strain.
[0122] The Malate Dehydrogenase (MDH) (EC 1.1.1.37) Variants
[0123] A malate dehydrogenase (MDH) gene library was designed to identify other MDH variants that may have improved activity over MDH50. To construct the phylogenetic tree, multiple sequence alignment (MSA) was performed using Clustal Omega 1.2.2. A phylogenetic tree was then calculated, submitting these aligned sequences to the Geneious Tree Builder (Geneious Prime® 2020.0.4). The tree builder uses the Jukes Cantor Genetic distance model to build a neighbor-joining tree with no outgroup and no bootstrap calculations. Bootstrapping was not necessary' for this application, because the intent was not to represent true evolutionary' relatedness but rather as a simple gauge for genetic diversity. A phylogenetic tree of the MDH enzyme library is shown in FIG. 2. Table 2 provides all the MDH variants tested.
[0124] Table 2.
[0125]
[0126]
[0127]
[0128]
[0129]
[0130] The enzyme kinetics characterization of MDH50 showed it had a Kcat of ~1000 s-1at pH 7.2 at 33 °C, and a Km of 0.32 mM. Additionally, MDH50 demonstrated substrate-based inhibition at higher concentrations of oxaloacetate.
[0131] To obtain variants with improved enzyme activities, five MDHs (MDH 51-55) were selected from the BRENDA database that had been previously reported to have Kcat > 2500 s'1at pH < 9.0 and an assay temperature of < 37 °C. Also, 3 MDH mutants (and their WT sequences) were selected (MDH56-60) that previously had been reported to have no substrate inhibition. An additional 15 variants (MDH61-75) were selected from the GotEnzymes database for the reaction (ECI.1.1.37) using oxaloacetate (C00036) as substrate. These enzymes were predicted to have a Kcat of > 10,000 s'1based on the deep-leaming prediction tool, DLKcat. To identify other variants that are similar to MDH50, a PSI-BLAST search was performed using the MDH50 as query sequence against the NR_clustered database and all enzy mes having sequence identities of greater than 85% were selected (MDH76-107). Variant CCF47222.1 was identified to contain ambiguous residues, so it was replaced with another variant from Colletotrichum sp. (EFQ28643.1), following 5 rounds of SIB-BLAST. To expand the sequence space beyond the sequences deposited in the NCBI protein database, seven MDH variants were selected (MDH108-114). These seven variants were previously reported to be generated using generative machine learning methods (ProteinGAN) (Repecka, D. et al.. Expanding functional protein sequence spaces using generative adversarial networks. Nat Mach Intell 3, 324-333 (2021)).
[0132] The MDH enzyme library was ordered from Twist Bioscience (South San Francisco, CA) with N-terminal His-tag constructs in a pET28a vector and tested in in vitro enzyme assays. Plasmids were transformed into chemically competent BL21-DE3 cells and expressed in TB-autoinduction media for 24 hours at 30 °C. The overexpressed proteins were affinity purified using Ni-NTA-filled PhyTip columns and a Hamilton Star liquid handling system. The eluates were desalted using size-exclusion resin and their protein recovery was analyzed by Bradford protein assay. Malate dehydrogenase activity was determined by using a 384-well plate-based method that monitored the decrease in NADH at 340 nm in the presence of 0.5 mM oxaloacetate (OAA). The malate dehydrogenase activities are summarized in FIG. 3.
[0133] Based on the data in FIG. 3, a subset of the MDH enzyme variants identified as high-performing in the biochemical evaluations were selected for expression in A pombe cells to rank the in vivo activity of those MDH enzymes directly in the production strain. The S. pombe genome contains a cytosolic pyruvate carboxylase PYRl / SPBC17G9.11c, so there is cytoplasmic oxaloacetate available to the heterologous MDHs. Nineteen MDH variants were selected for expression in S. pombe cells using the constitutive promoter of the ACTl / SPBC32H8.12c gene (SEQ ID NO: 24). Synthetic DNA for each MDH variant was codon optimized for expression in S'. pombe. Nineteen (19) 2* P CTI-MDH50 strains were made by a 1-step duplex CRISPR-Mad7 replacement of LDH58 at the PDC201 and ADH1 loci in Sk3 cells using the PAM-protospacer sequence 5’-TTTGATGGCTAAATCCGCTGAAACC-3’ (SEQ ID NO: 1). Sk3 was selected because it is mae2A, so these cells are unable to directly convert malate to pyruvate. MDH50 has been shown in shake-flask experiments to produce malic acid (FIG. 1) so it was included for reflecting a baseline activity, whereas Sk3 is the negative control (no MDH). Shake-flask fermentations were conducted as described in Example 2 and evaluated for malic acid production after 48 hours. All 19 strains produced malic acid, while MDH50 was ranked 16th out of 19 (FIG. 4). A small amount of malic acid was observed from the no MDH control strain because S. pombe contains pyruvate decarboxylase (PYR1 / SPBC17G9.11 c) and malate dehydrogenase MDHl / SPCC306.08c.
[0134] In this 2x MDH shake-flask experiment, titer was used to rank MDH performance. MDH93 (SEQ ID NO: 121), MDH96 (SEQ ID NO: 124), MDH102 (SEQ ID NO: 130), MDH92 (SEQ ID NO: 120), MDH83 (SEQ ID NO: 111), and MDH91 (SEQ ID NO: 119) produced the most malic acid in these 2* MDH strains, which means that these MDH variants are the most active enzy mes in S. pombe cells. With a cutoff of greater than 5 standard deviations above MDH50, thirteen MDH enzymes had better activity than MDH50 in S. pombe cells. The top three MDHs: MDH93 (Colletotrichum fructicoldy, MDH96 (Verticillium dcihliay. and MDH102 (C olletotrichum graminicola) are homologs of MDH50, thereby illustrating strong malate dehydrogenase activity from this clade. These MDH-only data demonstrate that by adding heterologous MDH to a redox-imbalanced S', pombe cell that has heterologous LDH replaced by MDH and lacks functional endogenous ADH1 and ADH4 genes, such variants can generate some metabolic pull from pyruvate to malate and that metabolic pull is scientifically sufficient to rank the in vivo activity of heterologous MDH enzymes. However, without a highly active carboxylase enzyme (e.g., heterologous PPC / PYC), the oxaloacetate pools are too low in these variant cells to achieve high titers. This suggests that the addition of heterologous PPC / PYC would remove the malic acid pathway bottleneck converting pyruvate to oxaloacetate.
[0135] The Pyruvate Carboxylase (PYC) (EC 6.4.1.1) Variants
[0136] The conversion of 3-carbon pyruvate to 4-carbon oxaloacetate (OAA) is necessary to produce high titers of malic acid from dextrose. Pyclp is a large ATP-requiring, homo-tetrameric protein that requires biotin and carbonate to carboxylate pyruvate to form oxaloacetate. Control of Pyclp activity is complex and dependent on multiple intracellular metabolites. There are reports in the literature on the allosteric regulation of Pyclp by acetyl-CoA, adenosine triphosphate (ATP), and pyruvate. There are reports of allosteric inhibition of Pyclp by malate and / or oxaloacetate. The S pombe genome contains one PYC enzyme encoded by the gene PYRl / SPBC17G9.11c, which explains why strains that only contain heterologous MDH are capable of making malic acid (albeit low titers). Gene libraries of heterologous pyruvate carboxylase (PYC) enzymes were designed and integrated into S. pombe genome, and the variants obtained were tested for activity in vivo.
[0137] Bioinformatics was used to design a diverse library containing 86 PYC enzymes.
[0138] A phylogenetic tree of the PYC enzy me library is shown in FIG. 5. Table 3 provides all the PYC variants tested.
[0139] Table 3.
[0140]
[0141]
[0142]
[0143]
[0144]
[0145] To avoid Pyclp variants that are allosterically regulated or product-inhibited, a set of PYC variants was first selected from fungal organisms that natively produce large amounts of C3-C4 organic acids. Aspergillus niger is a natural producer of C4 diacids, and it encodes a PYC (AnPYC) which is reported to not require acetyl-CoA for activation. S. pombe has relatively low acetyl-CoA pools that could be highly dynamic based on the fermentation phase. Thus, it is desirable to utilize PYC enzymes that can function independently of cellular acetyl-CoA concentrations. AnPYC (UNIPROT accession ID: PYC_ASPNG) was used as the seed sequence to identify homologs from public protein databases. The AnPYC sequence was used to perform a BLAST search of the Genbank Non-Redundant (NR) Clustered Protein Database. Multiple Sequence Alignment (MSA) was performed using Clustal Omega 1.2.2 on 2500 homologous sequences retrieved by the search. Proteins that were more than 3 standard deviations greater or smaller in length than the average sequence length and proteins containing non-standard amino acids were removed. The remaining sequences were split into 2 groups: 1) the 460 closest hits (mostly fungal); and 2) the rest (mostly bacterial). The bacterial set of PYCs was reduced to 1000 randomly chosen sequences. The final library' was selected from the 60 most diverse fungal sequences and the 30 most diverse bacterial using a Similarity Reduction Algorithm. Also included in the PYC library was Rhizopus oryzae PYC (Ro PYC), which is also reported to not require acetyl-CoA for activation; and a mutant RoPY C that contains an arginine to proline substitution at amino acid 485 (RoPYC_R485P) previously reported to improve fumaric acid production (SEQ ID NO: 228). Fumaric acid is produced by the enzymatic dehydration of malic acid by fumarase (also known as fumarate dehydratase). Organisms that can make high titers of fumaric acid must make high titers of malic acid, therefore PYC enzymes selected from organisms that can produce fumaric acid should also be able to produce malic acid.
[0146] To ensure that the library members are localized to the cytoplasm, an artificially intelligent deep-leaming model of protein localization in eukaryotic cells was used.
[0147] DeepLoc 2.0 is a Deep Neural Network classifier for determining the localization of eukaryotic proteins which was developed using 13,858 proteins of known localization from the UNIPROT database. It predicts the localization of unknown proteins to 1 of 10 potential cellular sub compartments including mitochondria / chloroplast and cytoplasm (Almagro Armenteros JJ. Sonderby CK, Sonderby SK, Nielsen H, Winther O. '‘DeepLoc: prediction of protein subcellular localization using deep learning,” Bioinformatics. 2017 Nov 1; 33(21): 3387-3395). The PYC proteins were submitted to DeepLoc 2.0 to determine their predicted localization and none were predicted to localize to a compartment other than the cytoplasm. This set of 92 proteins is the library delivered for PYC.
[0148] Strain Skl43 (2x PACTI-MDH50) was the recipient strain selected for receiving and systematically evaluating PYC enzymes. In this experiment, the heterologous PYCs were all expressed using the same ZYM1 / SPAC22H10.13 promoter (SEQ ID NO: 26) (see SPAC22H10.13 at PomBase (https: / / www.pombase.org / gene / SPAC22H10.13)). The ZYM1 promoter was identified as a constitutive promoter. The PYC sequences were codon optimized for expression in S. pombe cells, and one-step CRIPSR-Mad7 genome editing was used to integrate the PYC expression cassettes at GORI / SPACUNK4.10 locus using the PAM-Protospacer sequence 5’-TTTGGGCGGTATTGGTAAGACCATG-3’ (SEQ ID NO: 4), resulting in the genotypes gorl:: PzYMi-PYC. A non-PYC Green Fluorescent Protein (GFP) was integrated to make the Sk269 (gorl:: PZYMI-GFP) negative control strain. Eighty-five PYC-producing. pombe strains were built and confirmed by diagnostic PCR using standard PCR methods.
[0149] The PYCs were ranked for their activity based on malic acid titer in a 96-well plate using a 2-phase fermentation. The results are summarized in FIG. 6. The two fermentation phases are Seed, which promotes biomass production; and Main, which promotes malic acid production.
[0150] Seed fermentation was performed in square well, V-bottom 96 well plates with a culture volume of 200 µL, and incubation was performed at 33 °C for 24 hours. The Main phase of fermentation was conducted in a round well, U-bottom 96 well plate, 400 µL culture volume, and incubation was performed at 33 °C for 48 hours. These production conditions model a more microaerobic (fermentative) condition, promoting product formation, because only those cells that pass carbon through malate dehydrogenase (MDH) would achieve NAD(H) redox balance with glycolysis. The S. pombe PYC (PYR1 / SPBC17G9.11c) was not deleted in these strains, so the “no PYC control” produces some malic acid via the S. pombe PYC. Therefore, the negative control strain Sk265 (GFP) produced 2.4g / L malic acid. As shown in FIG. 6, eighty-six strains made more malic acid than Sk265. The highest 3 PYCs were: PYC77 (Spizellomyces palustris) (SEQ ID NO: 166); 2) PYC65 (Hesseltinella vesiculosa) (SEQ ID NO: 154); and 3) PYC67 (Lichtheimia hyalospora) (SEQ ID NO: 156), PYC89 (Alkalibacillus haloalkaliphilus) (SEQ ID NO: 176), and PYC80 (Gaertneriomyces semiglobifer) (SEQ ID NO: 168), which all made more than 1.5x malic acid as Sk265. These data show diverse PYC enzymes are functional in 5. pombe cells and can generate sufficient oxaloacetate pools and pathway flux for producing malic acid from dextrose.
[0151] Phosphoenolpyruvate Carboxylase (PPC) (EC 4.1.1.31) Variants
[0152] The conversion of 3-carbon phosphoenolpyruvate (PEP) to 4-carbon oxaloacetate (OAA) is an alternate pathway to produce malic acid and other C4 diacids from dextrose. Phosphoenolpyruvate carboxylase (PPC) (EC 4.1.1.31) catalyzes the irreversible conversion of pyruvate to oxaloacetate (OAA) in the presence of bicarbonate (HCO₃⁻). To test the feasibility of this pathway, libraries of heterologous PPC enzymes were designed and inserted in the S. pombe genome for heterologous expression, and the variants obtained were tested for in vitro and in vivo activity7.
[0153] Bioinformatics was used to identify a sequence library consisting of 143 phosphoenolpyruvate carboxylase (PPC) enzymes. Due to the presence of multiple groups of evolutionarily diverse PPCs found in nature, a two-pronged approach to library7construction was used. Two seed sequences of known PPC enzymes: 1) Sorghum bicolor (UNIPROT accession ID: CAPP1_SORBI); and 2) Corynebacterium glutamicum (UNIPROT accession ID: CAPP_CORGL) were used to identify homologues from public protein databases. These PPCs were chosen because they are reported to not be allosterically inhibited by malic acid. These sequences were used to perform BLAST searches of the Genbank Non-Redundant (NR) Clustered Protein Database. Multiple Sequence Alignment (MSA) was performed using Clustal Omega 1.2.2 on 1000 homologous sequences retrieved by each search. Proteins that were more than 3 standard deviations greater or smaller in length than the average as well as those annotated as “partial” or containing non-standard amino acids were removed. To ensure that the library7members are localized to the cytoplasm, DeepLoc 2.0 was used as with the PYC library. The gathered PPC proteins were submitted to DeepLoc 2.0 analysis to determine their predicted localization. Those predicted to localize to a compartment other than the cytoplasm were discarded. Each of the protein sets was reduced using a Similarity Reduction Algorithm until 45 representative proteins for each seed remained. Additionally, twenty-five PPC variants were selected from the database search and tested in vivo for carboxylation activity based upon malic acid production. PPC25 (SEQ ID NO: 253) was identified as an unusual, Archaeal-like PPC. The majority of PPCs that were tested had either plant or bacterial origins whereas PPC25 is from the bacteria Clostridium perfringens. PPC25 (SEQ ID NO: 253) contains an atypical amino acid sequence and is smaller in size, being only 538 AAs long whereas most other PEPCs tested were around 1,000 amino acids long. Performing a BLAST analysis on PPC25 revealed that it is more closely related to Archaeal PPCs than bacterial PPCs. The sequence of PPC25 (Uniprot ID: Q8XLE8) was used to perform a BLAST search of the Genbank Non-Redundant (NR) Clustered Protein Database and 100 homologs were retrieved from the public protein database. Sequences that were less than 400 and greater than 600 amino acids long were removed. Multiple Sequence Alignment (MSA) was performed using Clustal Omega 1.2.2 on 90 homologous sequences along with PPC25 and a phylogenetic tree was generated using the neighbor-joining method and Jukes-Cantor distance model. PPC25 was selected as the variants’ tree roots and branches were ordered by increasing distance. Based on subtree distance, a representative variant was selected from each tree branch and 24 heterologous proteins were selected for in vitro enzyme screening. A phylogenetic tree of the PPC enzyme library is shown in FIG. 7. Table 4 provides all the PPC variants tested.
[0154] Table 4.
[0155]
[0156]
[0157]
[0158]
[0159]
[0160]
[0161] This set of 143 PPC variants were selected for in vitro enzyme screening. Escherichia coli PPC (PPC 141), Escherichia coli PPC K620S mutant (PPC 142), Corynebacterium glutamicum PPC N917G mutant (PPC 143), and Sorghum vulgare PPC S8D mutant (PPC144) (SEQ ID NOs: 342-347) were also selected for enzyme screening. Deeploc2 analysis of the selected enzyme variants confirmed that all enzymes were free of subcellular localization tags and were predicted to be localized in the cytoplasm.
[0162] Full-length sequences of selected PPC variants were ordered from a Twist Bioscience as N-terminal His-tagged constructs and cloned into pET28a vector backbone. These plasmids were transformed into the BL21(DE3) host strain cells and the enzymes were expressed in LB media induced with 0.5 mM IPTG for 22 hours at 20 °C. The collected cell pellets were frozen, lysed, and soluble enzymes were purified using TALON Cobalt IMAC purification and buffer exchanged by gel-filtration to remove imidazole. The purified enzymes were tested for carboxylation activity using a coupled enzyme assay using a porcine malate dehydrogenase enzyme that converts oxaloacetate (produced by carboxylation of PEP) to malate using NADH cofactor. The continuous enzyme assay monitored the conversion of NADH to NAD+ at 340 nm in a 384-well plate reader for 1 hour at 30 °C. PPC enzymes are reported to be allosterically activated by fructose- 1,6-biphosphate and acetyl-CoA, and inhibited by aspartate, malate, and phosphate ion. Enzyme activity7was determined both in the presence and absence of 350 µM Acetyl-CoA in Tris HC1 buffer at pH 7.2.
[0163] The in vitro PPC activities are summarized in FIG. 8 and FIG. 9. As expected, the literature reported E. coli PPC (PPC 141) did not show any activity in the absence of the activator acetyl-CoenzymeA (AcCoA). Eighty-four out of the one hundred and nineteen enzymes showed carboxylase activity7in the presence of 350 µM AcCoA whereas only sixty-seven showed some carboxylase activity in the absence of AcCoA. This observation confirmed that a distinct subgroup of selected PPCs has a different propensity to allosteric activation with AcCoA. The cellular abundance of cytoplasmic acetyl-CoA throughout any part of the malic acid fermentation process is unknown. It thus may be beneficial or necessary7to have one of each type of PPC. Ten enzymes that did not require acetyl-CoA for activation and ten that had improved carboxylation activity in the presence of acetyl-CoA were selected for rational protein engineering to disrupt allosteric malate inhibition. The C-terminal asparagine (N917) in Corynebacterium glutamicum PPC has been reported to be a critical residue for the allosteric interaction of malate with the PPC enzyme. It has been shown this interaction can be disrupted by an N917G mutation to reduce feedback inhibition (Chen, Zhen et al. '‘Deregulation of feedback inhibition of phosphoenolpyruvate carboxylase for improved lysine production in Corynebacterium glutamicum, ” Applied & Environmental Microbiology 80(4): 1388-93, 2014). All twenty7enzy mes selected for initial enzyme activity screening were mutated from asparagine to glycine (PPC 169-188) for enzyme screening. The enzyme assays were then performed for all forty variants in the presence of 5 mM malate to find enzyme variants that were not malate inhibited. Fifteen out of thirty-one PPCs tested retained carboxylation activity7in the presence of malate. The results of malate inhibition are shown in FIG. 10. Tn summary, a diverse library of one hundred and thirty-nine unique phosphoenolpyruvate carboxylase (PPC) library was evaluated using in vitro methods for the carboxylation of phosphoenolpyruvate to oxaloacetate. Experiments were performed with and without the addition acetyl-CoA. which is a known allosteric activator of PPC; and malate, which is known to inhibit PPC activity. Based on the biochemical data, twenty PPC variants were selected for expression in S. pombe cells that contain two copies of MDH50 (Skl43).
[0164] All the S. pombe variants expressed PPCs that were transcribed using the strong, constitutive promoter from ZYM1 / SPAC22H10.13. These strains were evaluated for malic acid production in shake-flasks (as described in Example 2) and ranked based on titer.
[0165] Sixteen strains (PPC72 (SEQ ID NO: 275), PPC80 (SEQ ID NO: 283), PPC68 (SEQ ID NO: 271), PPC60 (SEQ ID NO: 266). PPC 108 (SEQ ID NO: 311), PPC145 (SEQ ID NO: 348), PPC146 (SEQ ID NO: 349), PPC25 (SEQ ID NO: 253), PPC17 (SEQ ID NO: 245), PPC 147 (SEQ ID NO: 350), PPC 143 (SEQ ID NO: 346), PPC 117 (SEQ ID NO: 320), PPC134 (SEQ ID NO: 337), PPC56 (SEQ ID NO: 259), PPC90 (SEQ ID NO: 293), and PPC86 (SEQ ID NO: 289)) produced more malic acid than the control strain Skl43, which produces malic acid via the endogenous pyruvate carboxylase PYR1 and heterologous malate dehydrogenase (MDH50). Ten strains of these sixteen strains (PPC72, PPC80, PPC68, PPC60, PPC108, PPC145, PPC146, PPC25, PPC17, and PPC147), produced more than 1.5x malic acid than the control strain Skl43. The malic acid production results are illustrated in FIG. IE
[0166] Example 4, Characterization of various engineered S. pombe strains for C4 dicarboxylic acid production.
[0167] Malic acid fermentations from dextrose often result in a 4-carbon diacid cocktail comprising malic acid, fumaric acid, and succinic acid. This may be due to metabolic crosstalk between the engineered heterologous reductive TCA (rTCA) pathway and the natural oxidative TCA (TCA) pathway encoded in the host genome. The genetically engineered 5. pombe strains of the above examples were subject to further modifications to optimize malic acid production, e.g., producing a 4-carbon diacid cocktail comprising predominantly malic acid.
[0168] The dehydration of malate to fumarate in the genetically engineered malic acid strains is expected to result in the undesired production of fumaric acid. In S. pombe, the fumarase enzyme (Fumlp) encoded by the gene FUM1 / SPCC18.18c catalyzes the reversible hydration / dehydration of fumarate to malate (SEQ ID NO: 377 is the amino acid sequence; SEQ ID NO: 378 is the coding nucleotide sequence). In S. pombe, the Fumlp can localize to both the mitochondria and cytoplasm. It might be possible to sequester the Fumlp to the mitochondria to minimize the production of fumarate into the culture supernatant. For this, FUM1-MLS swap was performed and characterized. The mitochondrial localization sequence of FUM1 in the genetically engineered S. pombe strains was replaced with that of LSC2 or POG1. LSC2 / SPCC 1620.08 encodes succinate-CoA ligase beta subunit; and Lsc2p is localized exclusively in the mitochondria. POG1 / SPCC24B10.22 encodes the mitochondrial DNA polymerase, gamma subunit; and Poglp is localized exclusively to in the mitochondria. The mitochondrial localization sequences (MLS) for the FUM1, OSM1 (fumarate reductase), POG1, and LSC2 genes (SEQ ID NOs: 380, 382, 384, and 386 are the amino acid sequences; SEQ ID NOs: 381, 383, 385, and 387 are the coding nucleotide sequences) were identified using the DeepLoc 2.0 program available at the Technical University' of Denmark (DTU) Health Tech site (https: / / senices.healthtech.dtu.dk / services / DeepLoc-2.1 / ). To replace the FUM1 localization sequence with those of POG1 and LSC2, the amino acid sequences of FUM1, POG1, and LSC2 were entered into the DeepLoc 2.0 program using the “High-quality' model” and “Long output” settings. The MLS regions were determined to be the region starting with the first methionine residue with a Sorting Signal Importance score above 0.00.
[0169] Sk220 was built from Sk94 by replacing the mitochondrial localization sequence (MLS) region of FUM1 (SEQ ID NO: 380) with that of LSC2 (SEQ ID NO: 386) using the PAM-Protospacer sequence 5 -TTTAGTCGTGCGTTATCCCATTGAA-3’ (SEQ ID NO: 372) within the FUM1 intron (SEQ ID NO: 379). The repair DNA template contained a codon-optimized variant of the LSC2 MLS region (SEQ ID NO: 387) flanked by homology to the FUM1 gene. Sk220 contains a synthetic FUM1 gene fused to the LSC2 localization sequence (FUM1-LSC2) (SEQ ID NO: 388 is the amino acid sequence; SEQ ID NO: 389 is the coding nucleotide sequence).
[0170] Sk471 was built from Sk220 by replacing the mitochondrial localization sequence (MLS) region of OSM (SEQ ID NO: 382) with that of LSC2 (SEQ ID NO: 386) using the PAM-Protospacer sequence 5’ - TTTATACATGGACTTTTAGGCGTCT-3’ (SEQ ID NO: 373) within the OSM1 gene (SEQ ID NO: 395). The repair DNA template contained a codon-optimized variant of the LSC2 MLS region (SEQ ID NO: 387) flanked by homology to the 0SM1 gene. Sk471 contains a synthetic OSM1 gene fused to the LSC2 localization sequence (OSM1-LSC2) (SEQ ID NO: 398 is the amino acid sequence; SEQ ID NO: 399 is the coding nucleotide sequence).
[0171] Sk223 was built from Sk94 by replacing the mitochondrial localization sequence (MLS) region of FUM1 (SEQ ID NO: 380) with that of POG1 (SEQ ID NO: 384) using the PAM-Protospacer sequence 5 -TTTAGTCGTGCGTTATCCCATTGAA-3’ (SEQ ID NO: 372) within the FUM1 intron (SEQ ID NO: 379). The repair DNA template contained a codon-optimized variant of the POG1 MLS region (SEQ ID NO: 385) flanked by homology to the FUM1 gene. Sk223 contains a synthetic FUM1 gene fused to the POG1 localization sequence (FUM1-POG1) (SEQ ID NO: 390 is the amino acid sequence; SEQ ID NO: 391 is the coding nucleotide sequence).
[0172] Sk468 was built from Sk223 by replacing the mitochondrial localization sequence (MLS) region of OSM1 (SEQ ID NO: 382) with that of POG1 (SEQ ID NO: 384) using the PAM-Protospacer sequence 5’ - TTTATACATGGACTTTTAGGCGTCT-3’ (SEQ ID NO: 373) within the OSM1 gene (SEQ ID NO: 395; SEQ ID NO: 394 is the corresponding amino acid sequence). The repair DNA template contained a codon-optimized variant of the POG1 MLS region (SEQ ID NO: 385) flanked by homology to the OSM1 gene. Sk468 contains a synthetic OSM1 gene fused to the POG1 localization sequence (OSM1-POG1) (SEQ ID NO: 396 is the amino acid sequence; SEQ ID NO: 397 is the coding nucleotide sequence).
[0173] Malic acid production by Sk94, Sk220, Sk223, Sk471 and Sk468 was characterized in the shake-flask experiments conducted under similar conditions as Example 2. As shown in FIG. 12, Sk94, containing the wild-type FUM1 gene, produced 6.3% fumaric acid, while Sk220 produced 1.1% fumaric acid and Sk223 produced 1.7% fumaric acid. These data suggest that the FUM1-MLS swap results in about a 75% reduction in fumaric acid. The OSM1-MLS swap did not further improve malic acid yield and did not reduce or eliminate succinic acid production, suggesting that succinic acid w as not produced via the activity' of the cytoplasmic Osmlp.
[0174] As shown in FIG. 12, Sk220 and Sk223 still produced similar amounts of succinic acid (about 10%) as Sk94, suggesting that the mechanism for succinic acid production in these strains is independent from the production of fumaric acid. It was known that the efflux of succinic acid from the mitochondria as a TCA intermediate is independent of the engineered rTC A pathway. Thus, the role of the efflux of succinic acid was next investigated as the source of off-target succinic acid production.
[0175] Although YHM2 / SPBC83.13 has been annotated as a mitochondrial citrate transporter, it was tested for succinic acid transport by loss-of-function mutagenesis in the example. For this, YHM2 was deleted from the genome of each of Sk94 to make Sk513, Sk471 to make Sk530, and Sk468 to make Sk528 strains using CRISPR-Mad7 with the PAM-Protospacer sequence 5’-TTTGGGCTTCTATCAGGGTCTTATT-3’ (SEQ ID NO: 375). In each of the resulting strains, the YHM2 gene was replaced by a non-coding unique molecular barcode sequence (BC2952) 5’- TCCTAGCGAGCGCATGAGTAGAGGTGCCAATCGTCTGCAA-3’ (SEQ ID NO: 376), so that all the YHM2 deletion strains had the genotype yhm2A:: BC2952 (see also Table 1).
[0176] The YHM2 mutant strains were characterized, along with the corresponding parent strains containing a functional Yhm2p transporter, in shake-flask fermentations as described in Example 2. As shown in FIG. 12, all strains containing the YHM2 gene deletion (Sk513, Sk528, and Sk530) was below the limit of detection for succinic acid. These data suggest that YHM2 is necessary’ and sufficient for mitochondrial / cytoplasmic succinic acid transport in 5. pombe cells, and the succinate produced during malic acid fermentations was coming from mitochondrial TCA. Deletion in YHM2 was also performed in the strains containing the Fumlp localized to the mitochondria, resulting in Sk528 and Sk530. Both Sk528 and Sk530 produced > 99% pure malic acid. In contrast, Sk94 having native Fumlp and YHM2 produced 82.4% malic acid, 6.2% fumaric acid and 11.4% succinic acid (FIG. 12).
Claims
What is claimed is:Claim 1. A genetically engineered Schizosaccharomyces pombe strain comprising a nucleic acid encoding an exogenous malate dehydrogenase (MDH) enzyme that is operably linked to a first promoter to express the MDH enzyme in the S. pombe strain, wherein the S. pombe strain is capable of fermenting glucose to produce malic acid.Claim 2. The 5. pombe strain of claim 1, wherein the MDH enzyme comprises an amino acid sequence with at least 90% sequence identity to an amino acid sequence selected from SEQ ID NOs: 29-142.Claim 3. The S. pombe strain of either claim 1 or claim 2, wherein the MDH enzyme comprises an amino acid sequence selected from SEQ ID NOs: 111, 119-121, 124, and 130.Claim 4. The S. pombe strain of any one of claims 1 to 3, wherein the MDH enzyme comprises an amino acid sequence selected from SEQ ID NOs 121, 124, and 130.Claim 5. The S. pombe strain of any one of claims 1 to 4, wherein the first promoter is an ACT! promoter (PACTI), a PYK1 promoter (PPYKI), or a ZYM1 promoter (PZYMI).Claim 6. The S. pombe strain of any one of claims 1 to 5, wherein the nucleic acid encoding the MDH enzyme is integrated at a locus of either the inactivated ADH1 gene or the inactivated ADH4 gene.Claim 7. The S. pombe strain of claim 1, further comprising two copies of the nucleic acid encoding the MDH enzyme.Claim 8. The S. pombe strain of any one of claims 1 to 7, further comprising (i) a nucleic acid encoding an exogenous pyruvate carboxylase (PYC) enzy me that is operably linked to a second promoter to express the PYC enzy me in the S. pombestrain, or (ii) a nucleic acid encoding an exogenous phosphoenolpyruvate carboxylase (PPC) enzy me that is operably linked to a third promoter to express the PPC enzyme in the S. pombe strain.Claim 9. The S. pombe strain of claim 8, wherein the PYC enzyme comprises an amino acid sequence with at least 90% sequence identity to an amino acid sequence selected from SEQ ID NOs: 143-228, and wherein the PPC enzyme comprises an amino acid sequence with at least 90% sequence identity to an amino acid sequence selected from SEQ ID NOs: 229-371.Claim 10. The S. pombe strain of either claim 8 or claim 9, wherein the PYC enzyme comprises an amino acid sequence selected from SEQ ID NOs: 154, 156, 166, 168. and 176, and wherein the PPC enzyme comprises an amino acid sequence selected from SEQ ID NOs: 245, 253, 259, 266, 271, 275, 283, 289, 293, 311, 320, 337, 346, and 348-350.Claim 11. The A pombe strain of any one of claims 8 to 10, wherein the PYC enzyme comprises an amino acid sequence selected from SEQ ID NOs: 154, 156, and 166, and wherein the PPC enzyme comprises an amino acid sequence selected from SEQ ID NOs: 245, 253, 266, 271, 275, 283, 311, and 348-350.Claim 12. The < S’. pombe strain of any one of claims 8 to 11, wherein each of the second promoter and the third promoter is an ACT1 promoter (PACTI), a PYK1 promoter (PPYKI), or aZYMl promoter (PZYMI).Claim 13. The < S’. pombe strain of any one of claims 8 to 12, wherein each of the nucleic acid encoding the PYC enzyme and the nucleic acid encoding the PPC enzyme is integrated at a locus of either the inactivated ADH1 gene or the inactivated ADH4 gene.Claim 14. The 5. pombe strain of claim 8, further comprising two copies of the nucleic acid encoding the PYC enzyme or two copies of the nucleic acid encoding the PPC enzyme.Claim 15. The X pombe strain of claim 5 or claim 12, wherein the ACT1 promoter (PACTI) comprises a nucleotide sequence of SEQ ID NO: 24 or an operably functional portion thereof,wherein the PYK1 promoter (PPYKI) comprises a nucleotide sequence of SEQ ID NO: 25 or an operably functional portion thereof, andwherein the ZYM1 promoter (PZYMI) comprises a nucleotide sequence of SEQ ID NO: 26 or an operably functional portion thereof.Claim 16. The S. pombe strain of any one of claims 1 to 15, further comprising one or more features selected from:(i) an inactivated alcohol dehydrogenase 1 (ADH1) gene;(ii) an inactivated alcohol dehydrogenase 4 (ADH4) gene;(iii) an inactivated glycerol-3-phosphate dehydrogenase 1 (GPD1) gene;(iv) an inactivated pyruvate decarboxylase 1 (PDC201) gene; and(v) an inactivated malic enzy me (MAE2) gene.Claim 17. The S. pombe strain of any one of claims 1 to 16. wherein the S. pombe strain produces a fumarase selectively located in mitochondria.Claim 18. The S. pombe strain of clam 17, wherein the fumarase comprises a mitochondrial localization sequence (MLS) from a succinate-CoA ligase beta subunit or a mitochondrial DNA polymerase gamma subunit.Claim 19. The S. pombe strain of claim 17 or claim 18, wherein the fumarase comprises an amino acid sequence with at least 95% sequence identity to SEQ ID NO: 388 or 390.Claim 20. The S. pombe strain of any one of claims 1 to 19, wherein the S. pombe strain comprises an inactivated mitochondrial citrate transporter (YHM2) gene.Claim 21. A genetically engineered Schizosaccharomyces pombe strain comprising:(i) a nucleic acid encoding an exogenous malate dehydrogenase (MDH) enzyme that is operably linked to a first promoter to express the MDH enzy me in the S. pombe strain;(ii) a nucleic acid encoding an exogenous pyruvate carboxylase (PYC) enzyme that is operably linked to a second promoter to express the PYC enzyme in the Y pombe strain or a nucleic acid encoding an exogenous phosphoenolpyruvate carboxylase (PPC) enzyme that is operably linked to a third promoter to express the PPC enzyme in the S. pombe strain;(hi) an inactivated alcohol dehydrogenase 1 (ADH1) gene;(iv) an inactivated alcohol dehydrogenase 4 (ADH4) gene;(v) an inactivated glycerol-3-phosphate dehydrogenase 1 (GPD1) gene;(vi) an inactivated pyruvate decarboxylase 1 (PDC201) gene; and(vii) an inactivated malic enzyme (MAE2) gene.wherein the S. pombe strain is capable of fermenting glucose to produce malic acid.Claim 22. The S. pombe strain of claim 21, wherein the S. pombe strain produces a fumarase selectively located in mitochondria or comprises an inactivated mitochondrial citrate transporter (YHM2) gene.Claim 23. The S. pombe strain of any one of claims 1 to 22, wherein a culture of the S. pombe strain is capable of producing at least 12.2 g / L malic acid after 48 hours culture in a baffled shake flask incubated at 33 °C.Claim 24. The S. pombe strain of any one of claims 17 to 20 and 22, wherein a culture of the S. pombe strain after 48 hours culture in a baffled shake flask incubated at 33 °C is capable of producing a 4-carbon diacid cocktail comprising about 90% or more malic acid.Claim 25. A Y pombe strain selected from the group consisting of:Y pombe Sk41 deposited as NRRL Y-68447;Y pombe Sk75 deposited as NRRL Y-68448;Y pombe Sk94 deposited as NRRL Y-68449;Y pombe Sk220 deposited as NRRL Y-68450;Y pombe Sk223 deposited as NRRL Y-68451;S. pombe Sk468 deposited as NRRL Y-68452;S. pombe Sk471 deposited as NRRL Y-68453;S', pombe Sk513 deposited as NRRL Y-68463;S. pombe Sk528 deposited as NRRL Y-68454; andS. pombe Sk530 deposited as NRRL Y-68455.Claim 26. A method of producing malic acid comprising the steps of: culturing a cell of the S. pombe strain of any one of claims 1 to 25 in the presence of a carbon source under conditions suitable for producing at least 12.2 g / L malic in 48 hours, andisolating malic acid from the culture.Claim 27. The method of claim 26. wherein the carbon source used for culturing the S', pombe strain is selected from glucose, sucrose, fructose, a glucose oligomer, raffinose, glycerol, starch, and any combination thereof.Claim 28. An isolated nucleic acid encoding an MDH enzyme having an amino acid sequence selected from SEQ ID NOs: 29-142, wherein the nucleic acid sequence is operably linked to an ACT1 promoter (PACTI), a PYK1 promoter (PPYKI), or a ZYM1 promoter (PZYMI).Claim 29. An isolated nucleic acid encoding a PYC enzyme having an amino acid sequence selected from SEQ ID NOs: 143-228, wherein the nucleic acid sequence is operably linked to an ACT1 promoter (PACTI), a PYK1 promoter (PPYKI), or a ZYM1 promoter (PZYMI).Claim 30. An isolated nucleic acid encoding a PPC enzyme having an amino acid sequence selected from SEQ ID NOs: 229-371, wherein the nucleic acid sequence is operably linked to an ACT1 promoter (PACTI), a PYK1 promoter (PPYKI), or a ZYM1 promoter (PZYMI).