Glycoengineering in THERMOTHELOMYCES HETEROTHALLICA
Patent Information
- Application Number
- JP2024544911
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-01-27
- Filing Date
- 2023-01-26
- Publication Date
- 2026-02-03
AI Technical Summary
The prior art is difficult to efficiently produce proteins with human N-glycosylation in fungi and yeast, especially since the glycosylation pattern is different from that in mammals, resulting in low yields and high impurities, affecting the efficacy and immune response.
By genetically engineering the tropical fungus Thermotheromyces heterothallica, deletion of the alg3 gene, expressing ER-targeted mannocidase 1 and glucuronidase 2 subunits, and introducing heterologous GlcNAc transferases 1 and 2, as well as other glycosyltransferases, its glycosylation pathways are adjusted to approach human N-glycosylation patterns.
A high yield of human N-glycosylated proteins has been achieved. More than 90% of the N-glycosylated structure is consistent with human proteins. It is suitable for therapeutic efficacy and immune response, and avoids the impact of cell survival.
Smart Images

Figure 00000042_0000 
Figure 00000042_0001 
Figure 00000042_0002
Abstract
Description
[Technical field]
[0001] The present invention relates to genetically modified Thermothelomyces heterothallica (formerly Myceliophthora thermophila) in which the protein glycosylation pathway has been engineered to produce proteins with N-glycans similar to those of mammalian proteins, particularly human proteins, with minimal disruption of endogenous genes. [Background technology]
[0002] Most therapeutic proteins require glycosylation to ensure proper folding, function, and activity. Glycosylation of therapeutic proteins is also particularly important for their immunogenicity. Therefore, such proteins cannot be produced in standard prokaryotic expression systems that lack the necessary glycosylation machinery. Because glycosylation and other post-translational modifications are essential for therapeutic glycoproteins, most of them are currently produced in mammalian cells. However, fermentation processes based on mammalian cell medium cultures (e.g., CHO, mouse, or human cells) are typically very slow, require expensive nutrients and cofactors (e.g., fetal bovine serum or specific growth factors), often result in low product titers, and are prone to infections that can contaminate the resulting protein product. Therefore, there is a shift towards serum-free expression systems. In particular, yeast and fungi have been developed as alternative protein expression systems.
[0003] Eukaryotic organisms, yeasts, and fungi are capable of post-translational modifications including N- and O-glycosylation, but protein glycosylation in yeasts and fungi differs considerably from that in mammalian cells. To overcome these problems, the possibility of re-engineering the N-glycosylation pathway has been explored, especially in the species most frequently used for the production of heterologous proteins (e.g., S. cerevisiae, Pichia pastoris, Yarrowia lipolytica, Hansenula polymorpha, and Aspergillus and Trichoderma species). However, protein yields still require improvement, and in particular the glycosylation pattern needs to be improved so that a high percentage of the produced proteins have the desired glycoforms, i.e., the glycoforms of mammalian (especially human and companion animal) proteins.
[0004] Parsaie Nasab et al., 2013, Appl Environ Microbiol., 79(3):997-1007, described a synthetic N-glycosylation pathway for producing recombinant proteins with human N-glycans in Saccharomyces cerevisiae. A Δalg3Δalg11 double mutant strain was used, which was further genetically modified to express an artificial flippase, a protozoan oligosaccharyltransferase, and Golgi-targeted human N-acetylglucosaminyltransferase I and II. The results confirmed the presence of the complex human N-glycan structure GlcNAc2Man3GlcNAc2 on a secreted monoclonal antibody recombinantly expressed in the mutant strain. However, heterogeneity of N-linked glycans was observed due to interference with Golgi-localized mannosyltransferase.
[0005] This work by Parsaie Nasab et al. is also described in US Patent Application Publication No. 2011 / 0207214, which discloses cells engineered to express lipid-linked oligosaccharide (LLO) flippase activity that can flip LLOs containing one mannose residue, two mannose residues, and three mannose residues from the cytoplasmic side to the luminal side of intracellular organelles, and is further reviewed along with other related work in De Wachter et al., 2018, Engineering of Yeast Glycoprotein Expression. In: Advances in Biochemical Engineering / Biotechnology. Springer, Berlin, Heidelberg.
[0006] Nos. 7,029,872, 7,326,681, 7,629,163, and 7,981,660 disclose cell lines with genetically engineered glycosylation pathways that allow them to carry out a series of enzymatic reactions that mimic the processing of glycoproteins in humans. Eukaryotic organisms, such as unicellular and multicellular fungi, that normally produce high mannose-containing N-glycans, are engineered to produce N-glycans such as Man5GlcNAc2 or other structures along the human glycosylation pathway.
[0007] U.S. Patent Nos. 7,449,308 and 7,935,513 disclose eukaryotic host cells with modified oligosaccharides that can be further modified by heterologous expression of a series of glycosyltransferases, sugar transporters, and mannosidases to provide host strains for the production of mammalian, e.g., human, therapeutic glycoproteins. The N-glycans produced in the engineered host cells have a Man5GlcNAc2 core structure, which can then be further modified by heterologous expression of one or more enzymes, e.g., glycosyltransferases, sugar transporters, and mannosidases, to produce human-like glycoproteins.
[0008] No. 7,795,002 discloses eukaryotic host cells, such as yeast and filamentous fungi, that produce human-like glycoproteins characterized by terminal β-galactose residues and essentially lacking fucose and sialic acid residues. Additionally, methods are disclosed for catalyzing the transfer of galactose residues from UDP-galactose onto acceptor substrates in recombinant eukaryotic host cells, which can be used as therapeutic glycoproteins.
[0009] No. 8,986,949 discloses genetically engineered strains of non-mammalian eukaryotes expressing catalytically active endomannosidase genes to enhance processing of N-linked glycan structures with the overall goal of obtaining more human-like glycan patterns. Additionally, the cloning and expression of novel human and mouse endomannosidases is disclosed.
[0010] US Patent No. 9,359,628 discloses genetically engineered strains of Pichia capable of producing proteins with smaller glycans. Specifically, the genetically engineered strains can express either or both of α-1,2-mannosidase and glucosidase II. The genetically engineered strains can be further modified to disrupt the OCH1 gene. Methods of producing glycoproteins with smaller glycans using such genetically engineered strains of Pichia are also provided.
[0011] No. 9,695,454 discloses compositions comprising filamentous fungal cells, such as Trichoderma fungal cells, having reduced protease activity and expressing a fucosylation pathway. Additionally, methods are described for producing glycoproteins having fucosylated N-glycans using genetically modified filamentous fungal cells, such as Trichoderma fungal cells, as an expression system.
[0012] Thermothelomyces heterothallica (Th. heterothallica) strain C1 (recently renamed from Myceliophthora thermophila, which was renamed from Chrysosporium lucknowense) is a thermotolerant ascomycete fungus that produces high levels of cellulases, making it attractive for the production of these and other enzymes on a commercial scale.
[0013] For example, U.S. Patent Nos. 8,268,585 and 8,871,493 disclose a transformation system in the field of filamentous fungal hosts for expressing and secreting heterologous proteins or polypeptides. Also disclosed is a process for producing large amounts of polypeptides or proteins in an economical manner. The system includes transformed or transgenic fungal strains of the genus Chrysosporium, more specifically Chrysosporium lucknowense, and mutants or derivatives thereof. Also disclosed are transformants containing Chrysosporium coding sequences and expression control sequences for Chrysosporium genes.
[0014] Wild type C1 was deposited under the Budapest Treaty under number VKM F-3500 D, date of deposit August 29, 1996. High Cellulase (HC) and Low Cellulase (LC) strains have also been deposited, for example, as described in U.S. Patent No. 8,268,585, e.g., strain UV13-6, deposit number VKM F-3632 D, strain NG7C-19, deposit number VKM F-3633 D, strain UV18-25, deposit number VKM F-3631 D.
[0015] Additional improved C1 strains that have been deposited include: (i) HC strain UV18-100f (Δalp1 Δpyr5)—deposit number CBS141147, (ii) HC strain UV18-100f (Δalp1 Δpep4 Δalp2 Δpyr5 Δprt1)—deposit number CBS141143, (iii) LC strain W1L#100I (Δchi1 Δalp1 Δalp2 Δpyr5)—deposit number CBS141153, and (iv) LC strain W1L#100I (Δchi1 Δalp1 Δpyr5)—deposit number CBS141149.
[0016] EP 2505651 discloses an isolated fungus that has been mutated or selected to have low protease activity, which has less than 50% of the protease activity compared to the non-mutated fungus, the fungus being of the genus Chrysosporium, preferably a strain of Chrysosporium lucknowense.
[0017] WO 2021 / 094935 to the applicant of the present invention discloses genetically modified Thermothelomyces heterothallica in which the protein glycosylation pathway has been engineered to produce proteins with N-glycans similar to human N-glycans. WO 2021 / 094935 discloses deletion or disruption of alg3 and alg11 genes, overexpression of flippase, and expression of heterologous GlcNAc transferase 1 (GlcNAc transferase 1, GNT1) and GlcNAc transferase 2 (GNT2).
[0018] There is a need for additional and alternative expression systems for producing recombinant human, pet, and other mammalian proteins that are capable of producing mammalian proteins, particularly glycoproteins having N-glycans of human and pet proteins, in high yields so that the proteins are suitable for therapeutic use in humans, pets, and other mammals. Summary of the Invention
[0019] The present invention provides Thermothelomyces heterothallica genetically modified to produce glycoproteins having mammalian protein N-glycans, particularly human protein N-glycans. The genetic modification of Th. heterothallica of the present invention includes deletion or disruption of the alg3 gene, heterologous expression or overexpression of ER-targeted mannosidase 1 (alpha-1,2-mannosidase) and ER-targeted glucosidase 2 alpha-subunit. The genetic modification of Th. heterothallica may also further include expression of heterologous GlcNAc transferase 1 (GNT1) and GlcNAc transferase 2 (GNT2). In some embodiments, the genetic modification further includes expression of the STT3 subunit of heterologous oligosaccharyltransferase (OST). In additional embodiments, the genetic modification further includes expression of a heterologous galactosyltransferase. In some embodiments, the genetic modification further comprises overexpression of an endogenous flippase or expression of a heterologous flippase.
[0020] The present invention is based in part on the discovery that the genetically modified Th. heterothallica disclosed herein produces glycoproteins in which the desired mammalian / human N-glycans constitute more than 90% of the N-glycans found on the glycoprotein, in some cases more than 95% or even more than 98% of the N-glycans. In addition, when further modified to express a heterologous mammalian glycoprotein (e.g., an antibody), the genetically modified Th. heterothallica disclosed herein produces high levels of heterologous glycoproteins with the desired mammalian / human N-glycans constituting more than 90% of the N-glycans. This is in contrast to previously described expression systems that produce large variations in the resulting N-glycans. Notably, no serious adverse effects on cell viability were observed with any of the modifications made.
[0021] It is further disclosed that expression of ER-targeted mannosidase 1 (alpha-1,2-mannosidase) and ER-targeted glucosidase 2 alpha-subunit does not compromise cell viability or production levels of heterologous glycoproteins. It is now disclosed that, unexpectedly, desired mammalian / human N-glycans on the produced glycoproteins were achieved without the need to interfere with expression of endogenous alpha-1,2-mannosyltransferase (ALG11).
[0022] It is noted that Th. heterothallica, unlike most fungi and yeasts, does not have hypermannosylated N-glycans that may contain up to 50 mannose residues, but rather "oligomannose" N-glycans containing 3-9 mannose residues, as well as hybrid-type N-glycans containing mannose and HexNAc residues, the structures of which have not been fully characterized. Because the structure and synthetic pathway of hybrid N-glycans have not been fully characterized, it was unclear whether such glycans could be eliminated using the genetic modifications described herein. Surprisingly, the genetic modifications according to the present invention are sufficient to result in the essential elimination of these structures, with more than 90% of the N-glycoforms being the desired mammalian / human N-glycans, without the need to reduce the expression of alg11.
[0023] Advantageously, the Th. heterothallica cells of the invention produce high yields of protein: protein levels obtained using the Th. heterothallica cells of the invention are much higher than those obtained using, for example, yeast.
[0024] Thus, the present invention provides an efficient system for producing glycoproteins with desired N-glycans suitable for therapeutic use in humans.
[0025] According to one aspect, the present invention provides a Thermothelomyces heterothallica genetically modified to produce glycoproteins having mammalian N-glycans, the genetic modification comprising: (i) a deletion or disruption of the alg3 gene such that Th. heterothallica does not produce a functional alpha-1,3-mannosyltransferase; (ii) expression of ER-targeted mannosidase 1 (alpha-1,2-mannosidase), and (iii) expression of ER-targeted glucosidase 2 alpha-subunit.
[0026] In some embodiments, the mannosidase 1 is Trichoderma reesei mannosidase 1. In other embodiments, the mannosidase 1 is Th. heterothallica mannosidase 1. In some embodiments, the glucosidase 2 alpha-subunit is selected from the group consisting of Th. heterothallica, T. reesei, and Aspergillus niger glucosidase 2 alpha-subunit. In additional embodiments, the genetic modification further comprises expression of a glucosidase beta-subunit.
[0027] In some embodiments, the ER-targeted Trichoderma reesei mannosidase 1 comprises the amino acid sequence set forth in SEQ ID NO: 2. In some embodiments, the ER-targeted Trichoderma reesei mannosidase 1 is encoded by an exogenous polynucleotide introduced into Th. heterothallica comprising the sequence set forth in SEQ ID NO: 1, or an analog or derivative thereof having at least 90% sequence identity.
[0028] In some embodiments, the ER targeted Th. heterothallica glucosidase 2 alpha-subunit comprises the amino acid set forth in SEQ ID NO: 4. In some embodiments, the ER targeted Th. heterothallica glucosidase 2 alpha-subunit is encoded by an exogenous polynucleotide introduced into Th. heterothallica comprising the sequence set forth in SEQ ID NO: 3, or an analog or derivative thereof having at least 90% sequence identity.
[0029] In some embodiments, the ER-targeted mannosidase 1 (alpha-1,2-mannosidase) and / or the ER-targeted glucosidase 2 alpha-subunit are integrated into the alp3 protease locus in the Th. Heterothallica genome. In certain embodiments, the ER-targeted mannosidase 1 (alpha-1,2-mannosidase) and the ER-targeted glucosidase 2 alpha-subunit are both integrated into the alp3 protease locus in the Th. Heterothallica genome.
[0030] In some embodiments, the genetic modification further comprises expression of heterologous GlcNAc transferase 1 (GNT1) and GlcNAc transferase 2 (GNT2).
[0031] In some embodiments, the heterologous GNT1 and GNT2 according to the present invention are of animal origin.
[0032] In some embodiments, the animal-derived GNT1 according to the invention comprises a heterologous Golgi localization signal.
[0033] In some embodiments, the animal-derived GNT1 according to the invention is human GNT1. In some embodiments, the animal-derived GNT1 according to the invention is human GNT1 that includes a heterologous Golgi localization signal.
[0034] In some embodiments, the animal-derived GNT1 according to the invention is bovine GNT1. In some embodiments, the animal-derived GNT1 according to the invention is bovine GNT1 that includes a heterologous Golgi localization signal.
[0035] In some embodiments, a heterologous Golgi localization signal according to the invention is a Th. heterothallica Golgi localization signal. In some embodiments, the Th. heterothallica Golgi localization signal is from the Th. heterothallica protein KRE2.
[0036] In other embodiments, a heterologous Golgi localization signal according to the invention is a yeast Golgi localization signal, hi some embodiments, the yeast Golgi localization signal is from the yeast protein KRE2.
[0037] In some embodiments, the animal-derived GNT2 according to the present invention is rat GNT2. In other embodiments, the animal-derived GNT2 according to the present invention is human GNT2.
[0038] In some embodiments, the animal-derived GNT1 is human GNT1 and the animal-derived GNT2 is rat GNT2. In some embodiments, the human GNT1 comprises a Th. heterothallica Golgi localization signal. In some embodiments, the Th. heterothallica Golgi localization signal is from the C1 Th. heterothallica protein KRE2. In other embodiments, the human GNT1 comprises a yeast Golgi localization signal.
[0039] In some embodiments, the animal-derived GNT1 is human GNT1 and the animal-derived GNT2 is rat GNT2. In some embodiments, the human GNT1 comprises a Th. heterothallica Golgi localization signal. In some embodiments, the Th. heterothallica Golgi localization signal is from C1 protein KRE2. In other embodiments, the human GNT1 comprises a yeast Golgi localization signal.
[0040] In some embodiments, Th. heterothallica according to the present invention is genetically modified to overexpress the endogenous Th. heterothallica RFT1 flippase.
[0041] In another embodiment, the Th. heterothallica according to the present invention is genetically modified to express a heterologous flippase, wherein the heterologous flippase is yeast FLC2p flippase.
[0042] In some embodiments, the genetic modification according to the present invention further comprises expression of a heterologous oligosaccharyltransferase STT3 subunit (heterologous STT3). In some embodiments, the heterologous STT3 according to the present invention is Leishmania major STT3.
[0043] In some embodiments, the genetic modification according to the present invention further comprises expression of a heterologous galactosyltransferase. In some embodiments, the heterologous galactosyltransferase is an animal-derived galactosyltransferase.
[0044] In some embodiments, the animal-derived galactosyltransferase is a human galactosyltransferase. In some embodiments, the animal-derived galactosyltransferase according to the invention is a human galactosyltransferase that includes a heterologous Golgi localization signal, including, for example, the Th. heterothallica KRE2 Golgi localization signal.
[0045] In additional embodiments, the animal-derived galactosyltransferase is a Xenopus tropicalis galactosyltransferase. In some embodiments, an animal-derived galactosyltransferase according to the invention is a Xenopus tropicalis galactosyltransferase that includes a heterologous Golgi localization signal, including, for example, the S. cerevisiae KRE2 Golgi localization signal.
[0046] In some embodiments, the Th. heterothallica is Th. heterothallica C1.
[0047] In some embodiments, C1 is a strain that has been modified to delete one or more genes encoding endogenous proteases.
[0048] In some embodiments, C1 is a strain that has been modified to delete the gene encoding an endogenous chitinase.
[0049] In some embodiments, C1 is selected from the group consisting of wild-type C1 accession number VKM F-3500 D, UV13-6 accession number VKM F-3632 D, NG7C-19 accession number VKM F-3633 D, UV18-25 accession number VKM F-3631 D, D, W1L#100I (prt-Δalp1Δchi1Δalp2Δpyr5) accession number CBS141153, UV18-100f (prt-Δalp1, Δpyr5) accession number CBS141147, W1L#100I (prt-Δalp1Δchi1Δpyr5) accession number CBS141149, and UV18-100f (prt-Δalp1Δpep4Δalp2Δprt1Δpyr5) accession number CBS141143. Each possibility represents a separate embodiment of the invention. According to certain embodiments, the C1 strain has reduced expression and / or activity of at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14 or more proteases. According to certain exemplary embodiments, the C1 strain has reduced expression or activity of ALP1, ALP2, PEP4, PRT1, SRP1, ALP3, PEP1, and MTP2 (ΔalpΔalp2Δpep4Δprt1Δsrp1Δalp3Δpep1Δmtp2).
[0050] According to some embodiments, Th. heterothallica is capable of producing heterologous mammalian glycoproteins having mammalian / human N-glycans that constitute more than 90%, 95%, or 98% of the N-glycans found on said glycoprotein.
[0051] In some embodiments, Th. heterothallica is further genetically modified to express a heterologous mammalian glycoprotein, hi some embodiments, the heterologous mammalian glycoprotein is an antibody or an antigen-binding fragment thereof.
[0052] According to some embodiments, the heterologous mammalian glycoprotein comprises mammalian / human N-glycans that constitute more than 90% of its total N-glycans. According to some embodiments, the heterologous mammalian glycoprotein comprises mammalian / human N-glycans that constitute more than 92%, 94%, 96%, 98%, or 99% of its total N-glycans.
[0053] According to some embodiments, the heterologous mammalian glycoprotein comprises an amount of mammalian / human N-glycans that is at least 80%, 85%, 90%, 95%, or more of the amount of N-glycans found on the same glycoprotein produced by a mammalian / human cell.
[0054] According to another aspect, the present invention provides a method for generating Th. heterothallica that produces glycoproteins having mammalian N-glycans, comprising: (a) deleting or disrupting the alg3 gene of Th. heterothallica such that Th. heterothallica does not produce a functional alpha-1,3-mannosyltransferase; (b) introducing into Th. heterothallica an exogenous polynucleotide encoding ER-targeted mannosidase 1 (alpha-1,2-mannosidase); (c) introducing into Th. heterothallica an exogenous polynucleotide encoding an ER-targeted glucosidase 2 alpha-subunit.
[0055] In some embodiments, the method further comprises introducing into Th. Heterothallica an exogenous polynucleotide encoding an endogenous flippase to induce overexpression of the endogenous flippase in Th. Heterothallica, or an exogenous polynucleotide encoding a heterologous flippase to induce expression of the heterologous flippase in Th. Heterothallica.
[0056] According to some embodiments, the ER-targeted mannosidase 1 (alpha-1,2-mannosidase) and the ER-targeted glucosidase 2 alpha-subunit are introduced using a single exogenous polynucleotide encoding both enzymes.
[0057] In some embodiments, the mannosidase 1 is Trichoderma reesei mannosidase 1. In some embodiments, the glucosidase 2 alpha-subunit is Th. heterothallica glucosidase 2 alpha-subunit.
[0058] According to a further aspect, the present invention provides a method for producing a glycoprotein having mammalian N-glycans, comprising the steps of: (a) providing a genetically modified Th. heterothallica according to the present invention; (b) culturing Th. heterothallica in a medium under conditions suitable for expressing the glycoprotein; (c) recovering the glycoprotein.
[0059] In some embodiments, the glycoprotein is a heterologous mammalian glycoprotein recombinantly expressed in Th. heterothallica. In some particular embodiments, the glycoprotein is a human protein recombinantly expressed in Th. heterothallica. In other embodiments, the glycoprotein is a companion animal protein recombinantly expressed in Th. heterothallica. In some embodiments, the heterologous mammalian glycoprotein is an antibody or antigen-binding fragment thereof.
[0060] According to some embodiments, the heterologous mammalian glycoprotein comprises mammalian / human N-glycans that constitute more than 90%, 92%, 94%, 96%, 98%, or 99% of its total N-glycans. According to additional embodiments, the heterologous mammalian glycoprotein comprises an amount of mammalian / human N-glycans that is at least 80%, 85%, 90%, 95%, or more of the amount of N-glycans found in the same glycoprotein produced by a mammalian / human cell.
[0061] According to a further aspect, the present invention provides a recombinant glycoprotein produced by the genetically modified Th. heterothallica according to the present invention, wherein the glycoprotein comprises a GlcNAc2Man3GlcNAc2 (G0) glycan.
[0062] According to a further aspect, the present invention provides a recombinant glycoprotein produced by the genetically modified Th. heterothallica according to the present invention, wherein the glycoprotein comprises Gal1GlcNAc2Man3GlcNAc2(G1) glycan, Gal2GlcNAc2Man3GlcNAc2(G2) glycan, or a combination thereof.
[0063] In some embodiments, the recombinant glycoprotein produced by the genetically modified Th. heterothallica according to the present invention is a pharmaceutical grade glycoprotein.
[0064] According to some embodiments, the heterologous mammalian glycoprotein comprises mammalian / human N-glycans that constitute more than 90%, 92%, 94%, 96%, 98%, or 99% of its total N-glycans. According to additional embodiments, the heterologous mammalian glycoprotein comprises an amount of mammalian / human N-glycans that is at least 80%, 85%, 90%, 95%, or more of the amount of N-glycans found in the same glycoprotein produced by a mammalian / human cell.
[0065] It is to be understood that any combination of each of the aspects and embodiments disclosed herein is expressly encompassed within the present disclosure.
[0066] These and further aspects and features of the present invention will become apparent from the following detailed description, examples and claims. [Brief description of the drawings]
[0067] [Figure 1A] Analysis of released N-glycans of Protein A affinity purified Nivolumab from glycoengineered strains. One chromatogram from each transformation is shown as an example. Results for the main N-glycans for all strains described in the text are shown in the corresponding table below the chromatograms. Figure 1A - Strains M4855 and M4856 (M3291-based strains with T. reesei Mns1-HDEL and C1 gls2a-HDEL expression) cultivated in shake flasks. Chromatogram of M4855 and table of main released N-glycans for both M4855 and M4856 are shown. Figure 1B-D - G1 / 2 strains M5129-M5132 (M4855-based strains with expression of G1 / 2 machinery) cultivated in shake flasks. Chromatograms of M5130 (FIG. 1B) and M5132 (FIG. 1C) and a table of the major released N-glycans for M5129-M5132 (FIG. 1D) are shown. [Figure 1B]Analysis of released N-glycans of Protein A affinity purified Nivolumab from glycoengineered strains. One chromatogram from each transformation is shown as an example. Results for the main N-glycans for all strains described in the text are shown in the corresponding table below the chromatograms. Figure 1A - Strains M4855 and M4856 (M3291-based strains with T. reesei Mns1-HDEL and C1 gls2a-HDEL expression) cultivated in shake flasks. Chromatogram of M4855 and table of main released N-glycans for both M4855 and M4856 are shown. Figure 1B-D - G1 / 2 strains M5129-M5132 (M4855-based strains with expression of G1 / 2 machinery) cultivated in shake flasks. Chromatograms of M5130 (FIG. 1B) and M5132 (FIG. 1C) and a table of the major released N-glycans for M5129-M5132 (FIG. 1D) are shown. [Figure 1C] Analysis of released N-glycans of Protein A affinity purified Nivolumab from glycoengineered strains. One chromatogram from each transformation is shown as an example. Results for the main N-glycans for all strains described in the text are shown in the corresponding table below the chromatograms. Figure 1A - Strains M4855 and M4856 (M3291-based strains with T. reesei Mns1-HDEL and C1 gls2a-HDEL expression) cultivated in shake flasks. Chromatogram of M4855 and table of main released N-glycans for both M4855 and M4856 are shown. Figure 1B-D - G1 / 2 strains M5129-M5132 (M4855-based strains with expression of G1 / 2 machinery) cultivated in shake flasks. Chromatograms of M5130 (FIG. 1B) and M5132 (FIG. 1C) and a table of the major released N-glycans for M5129-M5132 (FIG. 1D) are shown. [Figure 1D]Analysis of released N-glycans of Protein A affinity purified Nivolumab from glycoengineered strains. One chromatogram from each transformation is shown as an example. Results for the main N-glycans for all strains described in the text are shown in the corresponding table below the chromatograms. Figure 1A - Strains M4855 and M4856 (M3291-based strains with T. reesei Mns1-HDEL and C1 gls2a-HDEL expression) cultivated in shake flasks. Chromatogram of M4855 and table of main released N-glycans for both M4855 and M4856 are shown. Figure 1B-D - G1 / 2 strains M5129-M5132 (M4855-based strains with expression of G1 / 2 machinery) cultivated in shake flasks. Chromatograms of M5130 (FIG. 1B) and M5132 (FIG. 1C) and a table of the major released N-glycans for M5129-M5132 (FIG. 1D) are shown. [Diagram 2] Comparison of the amount of released N-glycans of nivolumab produced by strains M5129-M5132 (M4855-based strains with expression of the G1 / 2 machinery) cultured in shake flasks. The amount of released N-glycans has been normalized between samples using a constant amount of internal standard added to each sample for analysis. The response % is set to 100% for the reference protein Opdivo. [Figure 3A] Analysis of released N-glycans of Protein A affinity purified nivolumab from fermentation samples of the three glycoengineered strains. Results for the main N-glycans for all strains described in the text are shown in the corresponding table below the chromatograms. Figure 3A - Starting strain M3291 with only alg3 deletion. Chromatogram and table of main released N-glycans for M3291 from fermentation conditions. Figures 3B-3D - G1 / 2 strains M5130 and M5132. Chromatograms and table of main released N-glycans (Figure 3D) for M5130 (Figure 3B) and M5132 (Figure 3C) from fermentation conditions are shown. [Figure 3B]Analysis of released N-glycans of Protein A affinity purified nivolumab from fermentation samples of the three glycoengineered strains. Results for the main N-glycans for all strains described in the text are shown in the corresponding table below the chromatograms. Figure 3A - Starting strain M3291 with only alg3 deletion. Chromatogram and table of main released N-glycans for M3291 from fermentation conditions. Figures 3B-3D - G1 / 2 strains M5130 and M5132. Chromatograms and table of main released N-glycans (Figure 3D) for M5130 (Figure 3B) and M5132 (Figure 3C) from fermentation conditions are shown. [Figure 3C] Analysis of released N-glycans of Protein A affinity purified nivolumab from fermentation samples of the three glycoengineered strains. Results for the main N-glycans for all strains described in the text are shown in the corresponding table below the chromatograms. Figure 3A - Starting strain M3291 with only alg3 deletion. Chromatogram and table of main released N-glycans for M3291 from fermentation conditions. Figures 3B-3D - G1 / 2 strains M5130 and M5132. Chromatograms and table of main released N-glycans (Figure 3D) for M5130 (Figure 3B) and M5132 (Figure 3C) from fermentation conditions are shown. [Figure 3D] Analysis of released N-glycans of Protein A affinity purified nivolumab from fermentation samples of the three glycoengineered strains. Results for the main N-glycans for all strains described in the text are shown in the corresponding table below the chromatograms. Figure 3A - Starting strain M3291 with only alg3 deletion. Chromatogram and table of main released N-glycans for M3291 from fermentation conditions. Figures 3B-3D - G1 / 2 strains M5130 and M5132. Chromatograms and table of main released N-glycans (Figure 3D) for M5130 (Figure 3B) and M5132 (Figure 3C) from fermentation conditions are shown. [Figure 4]Comparison of the amount of released N-glycan of nivolumab produced by strains M5130 and M5132 cultured under fermentation conditions. The amount of released N-glycan is normalized between samples using a constant amount of internal standard added to each sample for analysis. % response is set to 100% for the reference protein Opdivo. [Figure 5A] Analysis of released N-glycans of all secreted proteins from glycoengineered strains. One chromatogram from each transformation is shown as an example. The results for the major N-glycans for all strains described in the text are shown in the table (Figure 5C) and chromatograms (Figures 5A-B). M6589 and M6596 are strains with the G1 / 2 machinery where TrMns1-HDEL is under the ubiquitin-like protein promoter, whereas M6590 and M6597 are strains with the G1 / 2 machinery where TrMns1-HDEL is under the bgl8 promoter. [Figure 5B] Analysis of released N-glycans of all secreted proteins from glycoengineered strains. One chromatogram from each transformation is shown as an example. The results for the major N-glycans for all strains described in the text are shown in the table (Figure 5C) and chromatograms (Figures 5A-B). M6589 and M6596 are strains with the G1 / 2 machinery where TrMns1-HDEL is under the ubiquitin-like protein promoter, whereas M6590 and M6597 are strains with the G1 / 2 machinery where TrMns1-HDEL is under the bgl8 promoter. [Figure 5C] Analysis of released N-glycans of all secreted proteins from glycoengineered strains. One chromatogram from each transformation is shown as an example. The results for the major N-glycans for all strains described in the text are shown in the table (Figure 5C) and chromatograms (Figures 5A-B). M6589 and M6596 are strains with the G1 / 2 machinery where TrMns1-HDEL is under the ubiquitin-like protein promoter, whereas M6590 and M6597 are strains with the G1 / 2 machinery where TrMns1-HDEL is under the bgl8 promoter. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0068] The present invention is directed to the genetic modification of the fungus Thermothelomyces heterothallica, particularly strain C1, to produce glycoproteins bearing mammalian protein N-glycans, particularly human, companion animal, and other mammalian protein N-glycans.
[0069] Glycoproteins produced by genetically modified Th. heterothallica as described herein are suitable for therapeutic use in humans, companion animals such as dogs, cats, and horses, and other mammals.
[0070] Protein glycosylation, the covalent attachment of oligosaccharides to the side chains of newly synthesized polypeptide chains in cells, is an ordered process in eukaryotic cells that involves a series of enzymes that sequentially add and remove sugar moieties. N-glycosylation is the process by which oligosaccharides are attached to asparagine residues, specifically to the side chain of asparagine that occurs in the sequence Asn-Xaa-Ser / Thr, where Xaa represents any amino acid except Pro.
[0071] N-glycosylation begins in the endoplasmic reticulum (ER), where the oligosaccharide Glc3Man9GlcNAc2 is assembled on the lipid carrier dolichol-pyrophosphate and then transferred to selected asparagine residues of polypeptides that have entered the lumen of the ER. The biosynthesis of lipid-linked oligosaccharides requires the activity of several specific glycosyltransferases (e.g., ALG1, ALG2, and ALG3). The biosynthesis begins on the cytoplasmic side of the ER membrane and ends in the lumen, where oligosaccharyltransferases (OSTs) select the NXS / T sequon of the nascent polypeptide and generate an N-glycosidic bond between the asparagine side chain amide and the oligosaccharide. The flipping of lipid-linked oligosaccharides from the outside to the inside of the ER is performed by flippases located in the ER membrane. Following transfer to the nascent polypeptide, the oligosaccharides are typically trimmed by glucosidases and mannosidases and the nascent glycoprotein is then transferred to the Golgi apparatus for further processing.
[0072] The synthesis of dolichol pyrophosphate-linked oligosaccharides is essentially conserved in all known eukaryotes. However, the further processing of oligosaccharides as glycoproteins move along the secretory pathway varies greatly between lower eukaryotes, such as fungi or yeast, and higher eukaryotes, such as animals and plants. Thus, the final composition of the sugar side chains differs between various organisms and is host dependent.
[0073] In microorganisms such as yeast, additional mannose and / or mannosylphosphate sugars are typically added, resulting in "high mannose" N-glycans that can contain up to 30-50 mannose residues.
[0074] In animal cells, including human, pet, and other mammalian cells, nascent glycoproteins are transferred to the Golgi apparatus where mannose residues are removed by Golgi-specific 1,2-mannosidases. Processing continues as the protein progresses through the Golgi by a number of modifying enzymes, including N-acetylglucosamine transferases (GnT I, GnT II, GnT III, GnT IV, GnT V, GnT VI), mannosidase II, and fucosyltransferases, which add and remove specific sugar residues. Finally, the N-glycans are acted upon by galactosyl transferases (GalT) and sialyltransferases (ST) and the completed glycoprotein is released from the Golgi apparatus. The N-glycans of animal glycoproteins may have biantennary, triantennary, or tetraantennary structures and typically contain galactose, fucose, and N-acetylglucosamine. Generally, the terminal residue of the N-glycan consists of sialic acid.
[0075] Th. heterothallica, unlike most fungi and yeasts, does not have hypermannosylated N-glycans, but rather "oligomannose" glycans (Man3-Man 8-9 ), as well as hybrid type glycans (Man3HexNac-Man8HexNac) that contain both Man and HexNAc residues. The exact structures of these hybrid glycans are not fully known. The hybrid glycans have the typical mannose residues, but in addition, an unknown HexNAc attached via a yet to be characterized linkage.
[0076] Because the structures and synthetic pathways of hybrid glycans have not been fully characterized, it was unclear whether such glycans could be eliminated using the genetic modifications described herein. Surprisingly, the genetic modifications according to the present invention resulted in the essential elimination of these structures, with more than 90% and in many cases more than 98% of the N-glycoforms being the desired mammalian / human glycans.
[0077] The present invention relates to GlcNAc2Man3GlcNAc2 ("G0"), GlcNAc2Man3GlcNAc2(Fuc) ("FG0"), Gal 1-2 GlcNAc2Man3GlcNAc2 ("G1" / "G2"), and Gal 1-2 The present invention relates to genetic modification of the N-glycosylation pathway in Th. heterothallica to produce a high percentage of glycoproteins having mammalian N-glycans, particularly human N-glycans, such as GlcNAc2Man3GlcNAc2(Fuc) ("FG1" / "FG2").
[0078] Specifically, in some embodiments, the genetic modification of the N-glycosylation pathway in Th. heterothallica includes: 1. Deletion of the C1alg3 gene (encoding alpha-1,3-mannosyltransferase), 2. Expression of ER-targeted Trichoderma reesei mannosidase 1 (alpha-1,2-mannosidase), 3. Expression of ER-targeted C1 glucosidase 2 alpha-subunit, 4. Expression of heterologous GlcNAc transferase 1 (GNT1) 5. Expression of heterologous GlcNAc transferase 2 (GNT2), and 6. Expression of heterologous galactosyltransferase 1 (GalT1).
[0079] The deletion of alg3 terminates the synthesis of the N-glycan precursor at Man5GlcNAc2 with one or two terminal glucoses. This glycan serves as a substrate for GNT1 and GNT2, which are introduced into Th. heterothallica. Additional genetic modifications may include the introduction of additional enzymes from human, companion animal, and other mammalian glycosylation pathways, such as galactosyltransferases and / or fucosyltransferases.
[0080] The heterologous enzymes described above are expressed with a targeting peptide such that the expressed enzyme is targeted to a specific cellular compartment.
[0081] As used herein, when an enzyme is referred to it includes enzymatically active fragments thereof and enzymatically active variants thereof.
[0082] The present invention is particularly directed to the manipulation of the N-glycosylation pathway of Th. heterothallica. It is noted that O-glycans may be present or may be removed or modified by further genetic modification of Th. heterothallica.
[0083] It is understood that the genetic modifications according to the present invention are such that the genetically modified Th. heterothallica is capable of growing at a sufficient rate suitable for its intended use.
[0084] As used herein, "C1" or "Thermothelomyces heterothallica C1" or "Th. heterothallica C1" all refer to Thermothelomyces heterothallica strain C1. Descriptions of the Thermothelomyces genus and species can be found, for example, in Marin-Felix Y (2015. Mycologica 107(3):619-632) and van den Brink J et al. (2012, Fungal Diversity 52(1):197-207).
[0085] It should be noted that the above mentioned authors (Marin-Felix et al., 2015) proposed a division of the genus Myceliophthora based on differences in optimal growth temperature, conidial morphology, and details of the sexual reproduction cycle. According to the proposed criteria, C1 clearly belongs to the newly established genus Thermothelomyces, which contains the former thermotolerant Myceliophthora species, and not to the genus Myceliophthora, which still includes non-thermostable species. Because C1 can form ascospores with other Thermothelomyces (formerly Myceliophthora) strains with the opposite mating type, C1 is best classified as Th. heterothallica strain C1, not Th. thermophila C1.
[0086] It should also be understood that fungal taxonomy has been constantly changing in the past, so that the current names listed above may have been preceded by various older names other than Myceliophthora thermophila (van Oorschot, 1977. Personia 9(3):403), which are now considered synonyms. For example, Thermothelomyces heterothallica (Marin-Felix et al., 2015. Mycologica, 3:619-63) is synonymous with Corynascus heterotchallica (von Arx et al., 1983), Thielavia heterothallica (von Klopotek, 1976. Archives of Microbiology 107(2), 223-224), Chrysosporium lucknowense and thermophile (von Klopotek, 1974. Archives of Microbiology 98(1), 365-369), and Sporotrichium thermophile (Alpinis 1963. Nova Hedwigia 5:74).
[0087] It should be further clearly understood that the present invention encompasses any strain containing a ribosomal DNA (rDNA) sequence that exhibits 99% or greater homology to SEQ ID NO:22, and all such strains are considered to be of the same species as Thermothelomyces heterothallica.
[0088] SEQ ID NO:22 is 99.98% identical to the rDNA sequence found on chromosome 7 of Th. heterothallica / thermophila (listed as Myceliophtora thermophilica) ATCC 42464 rDNA sequence (ncbi.nlm.nih.gov / nucleotide / CP003008.1).
[0089] Th. heterothallica strain C1 (as Chrysosporium lucknowense strain C1) was deposited according to the Budapest Treaty under number VKM F-3500 D, date of deposit 29 August 1996.
[0090] The above term also encompasses genetically modified substrains derived from wild-type strains, which have been mutated using random or directed techniques, such as UV mutagenesis, or by deleting one or more endogenous genes. For example, C1 strain may refer to a wild-type strain that has been modified to delete one or more endogenous genes encoding endogenous proteases and / or one or more genes encoding endogenous chitinases. For example, C1 strains (substrains) encompassed by the present invention include UV18-25, accession number VKM F-3631 D, strain NG7C-19, accession number VKM F-3633 D; and strain UV13-6, accession number VKM F-3632 D. Additional C1 strains that may be used in accordance with the teachings of the present invention include HC strain UV18-100f accession number CBS141147, HC strain UV18-100f accession number CBS141143, LC strain W1L#100I accession number CBS141153; and LC strain W1L#100I accession number CBS141149.
[0091] The Th. heterothallica fungi in general, and strain C1 in particular, show higher biomass production when grown in favorable conditions compared to yeast strains. The Th. heterothallica fungi can be grown in large-scale three-dimensional (3D) liquid cultures as well as on solid media. Some strains developed by the applicant of the present invention are less sensitive to feedback inhibition by glucose and other fermentable sugars present in the fungal growth medium as carbon sources compared to conventional yeasts and other fungi, and can tolerate high feed rates of carbon sources leading to high yields. Furthermore, some of these strains, when grown in commercial fermenters, have a significantly reduced medium viscosity compared to the high viscosity obtained with wild-type Th. heterothallica fungi not repressed by glucose, or with other filamentous fungi known to be used for protein production. The low viscosity may be due to a morphological change in the strains from long and highly entangled hyphae in the parent strain to short and less entangled hyphae in the developed strains. The low medium viscosity is highly advantageous for large-scale industrial production in fermenters. For example, Th. heterothallica C1 strain UV18-25, accession number VKM F-3631 D, which shows reduced sensitivity to glucose repression, has been grown industrially to produce recombinant enzyme in quantities exceeding 100,000 liters.
[0092] In some embodiments, the C1 strain of the present invention is a strain that has been modified to delete multiple (i.e., at least two) genes encoding endogenous proteases. In some embodiments, the C1 strain is a strain that has been modified to delete at least four genes encoding endogenous proteases. In additional embodiments, the C1 strain is a strain that has been modified to delete at least five genes encoding endogenous proteases. In some particular embodiments, the C1 strain is a strain that has been modified to delete at least six genes encoding endogenous proteases. In additional particular embodiments, the C1 strain is a strain that has been modified to delete at least eight genes encoding endogenous proteases. In additional particular embodiments, the C1 strain is a strain that has been modified to delete at least eight, nine, ten, eleven, twelve, thirteen, fourteen, or more genes encoding endogenous proteases. In certain exemplary embodiments, the C1 strain is a strain that has been modified to delete at least 13 or 14 genes encoding endogenous proteases.
[0093] It should be clearly understood that the teachings of the present invention encompass mutants, derivatives, progeny, clones, and analogs of Th. heterothallica C1 strain, provided that such derivatives, progeny, clones, and analogs, when genetically modified in accordance with the teachings of the present invention, are capable of growing and producing proteins having N-glycans as described herein.
[0094] It should be clearly understood that the term "derivative" with respect to a fungal lineage encompasses any fungal parent lineage with modifications that positively affect product yield, efficiency, or potency, or any trait that improves the fungal derivative as a tool for producing mammalian proteins, particularly heterologous proteins with N-glycans of human, companion animal, and other mammalian proteins, as described herein. As used herein, the term "progeny" refers to an unmodified descendant from a parent fungal lineage, such as a cell from a cell.
[0095] As used herein, "glycan" refers to an oligosaccharide chain that can be linked to a carrier such as an amino acid, peptide, polypeptide, lipid, or reducing end conjugate. The present invention particularly relates to N-linked glycans ("N-glycans") conjugated to a polypeptide N-glycosylation site such as -Asn-Xxx-Ser / Thr- by an N-linkage to the side chain amide nitrogen of an asparagine residue (Asn), where Xxx is any amino acid residue except Pro. The present invention may further relate to glycans as part of the dolichol-phospho-oligosaccharide (Dol-PP-OS) precursor lipid structure, which is the precursor of N-linked glycans in the endoplasmic reticulum of eukaryotic cells. The precursor oligosaccharides are linked by their reducing ends to two phosphate residues on the dolichol lipid.
[0096] Monosaccharides that typically constitute N-glycans found on mammalian glycoproteins include, but are not limited to, N-acetylglucosamine (abbreviated as "GlcNAc"), mannose (abbreviated as "Man"), glucose (abbreviated as "Glc"), galactose (abbreviated as "Gal"), sialic acid (abbreviated as "Neu5Ac"), and fucose (abbreviated as "Fuc").
[0097] The N-glycans share a common pentasaccharide, designated the "core" structure Man3GlcNAc2 (abbreviated as "Man3"). Important target glycan structures of the present invention include N-glycans having one GlcNAc residue on the terminal 1,3 mannose arm of the core structure and one GlcNAc residue on the terminal 1,6 mannose arm of the core structure. Such N-glycans include GlcNAc2Man3GlcNAc2 (referred to as the "G0" glycoform), Gal 1-2GlcNAc2Man3GlcNAc2 (termed "G1" or "G2" glycoforms according to the number of galactose residues), and their core fucosylated glycoforms: GlcNAc2Man3GlcNAc2(Fuc) ("G0F" or "FG0") and Gal 1-2 GlcNAc2Man3GlcNAc2(Fuc) ("G1F" and "G2F", or "FG1" and "FG2").
[0098] The term "alg3 gene" refers to a gene encoding alpha-1,3-mannosyltransferase. The term "alpha-1,3-mannosyltransferase" refers to dolichyl-P-Man:Man5GlcNAc2-PP-dolichol alpha-1,3-mannosyltransferase (EC 2.4.1.258), which is an ER resident enzyme that catalyzes the following reaction: Dolichyl beta-D-mannosyl phosphate + D-Man-alpha-(1->2)-D-Man-alpha-(1->2)-D-Man-alpha-(1->3)-[D-Man-alpha-(1->6)]-D-Man-beta-(1->4)-D-GlcNAc-beta-(1->4)-D-GlcNAc-diphosphodolichol
[0099]
number
[0100] In some specific embodiments, the "alg3 gene" is the gene encoding the C1 alpha-1,3-mannosyltransferase (ortholog of the JGI M. thermophila genome (mycocosm.jgi.doe.gov) accession number 2310419).
[0101] The Th. heterothallica of the present invention is genetically modified by deletion or disruption of the alg3 gene such that the Th. heterothallica does not produce a functional alpha-1,3-mannosyltransferase. The Th. heterothallica of the present invention does not exhibit detectable alpha-1,3-mannosyltransferase activity.
[0102] The term "Mannosidase 1" (alpha-1,2-mannosidase), abbreviated as "MDS1" or "MNS1," catalyzes the following reaction: 3xH2O+N4-(α-D-Man-(1→2)-α-D-Man-(1→2)-α-D-Man-(1→3)-[α-D-Man-(1→3)-[α-D-Man-(1→2)-α-D-Man-(1→6)]-α- D-Man-(1→6)]-β-D-Man-(1→4)-β-D-GlcNAc-(1→4)-α-D-GlcNAc)-L-asparaginyl-[protein](N-glucanmannose isomer 8A1,2,3B1,3)
[0103]
number
[0104] An exemplary accession number for alpha-1,2-mannosidase is Uniprot Q9P8T8.
[0105] The term "glucosidase 2 alpha-subunit" (GLS2-alpha or GLS2α) refers to an enzyme that sequentially cleaves the innermost two alpha-1,3-linked glucose residues from the Glc2Man9GlcNAc2 oligosaccharide precursor of immature glycoproteins.
[0106] The term "flippase" (EC 7.6.2.1) refers to an enzyme that transfers lipid-linked glycan precursors from the cytoplasmic side to the luminal side of the ER during their synthesis in the ER.
[0107] The term "GlcNAc transferase 1", abbreviated as "GNT1" (also "GnTI"), refers to alpha-1,3-mannosyl-glycoprotein 2-beta-N-acetylglucosaminyltransferase (EC 2.4.1.101), which is a Golgi-resident enzyme that transfers a GlcNAc residue from UDP-GlcNAc to the acceptor substrate Man5GlcNAc2 to produce GlcNAcMan5GlcNAc2. In the present invention, synthesis of the N-glycan precursor takes into account the deletion of alg3 and the expression of alpha-1,2-mannosidase and glucosidase 2 alpha-subunits to produce Man3GlcNAc2, and thus the glycan Man3GlcNAc2 serves as a substrate for GNT1 to produce GlcNAcMan3GlcNAc2.
[0108] The term "GlcNAc transferase 2," abbreviated as "GNT2" (also "GnTII"), refers to alpha-1,6-mannosyl-glycoprotein 2-beta-N-acetylglucosaminyltransferase (EC 2.4.1.143), which is a Golgi-resident enzyme that transfers a GlcNAc residue from UDP-GlcNAc to a free terminal mannose residue in GlcNAcMan3GlcNAc2 to produce GlcNAc2Man3GlcNAc2.
[0109] The terms "STT3 subunit of oligosaccharyltransferase", "STT3 protein", or simply "STT3" are used interchangeably herein to refer to the dolichyl-diphosphooligosaccharide-protein glycosyltransferase subunit (EC:2.4.99.18). The STT3 subunit is the catalytic subunit of the oligosaccharyltransferase (OST) complex, which catalyzes the initial transfer of a defined glycan (Glc3Man9GlcNAc2 in eukaryotes) from the lipid carrier dolichol-pyrophosphate to an asparagine residue within an Asn-X-Ser / Thr consensus motif in a nascent polypeptide chain (the first step in protein N-glycosylation). The STT3 protein catalyzes the following reaction: Dolichyl diphosphooligosaccharide + L-asparaginyl-[protein]
[0110]
number
[0111] The term "galactosyltransferase" (EC 2.4.1.38) refers to a Golgi-resident enzyme that transfers β-linked galactosyl residues to terminal N-acetylglucosamine.
[0112] The term "heterologous", when referring to a gene, enzyme, protein, or peptide sequence, such as a subcellular localization signal, is used herein to describe a gene, enzyme, protein, or peptide sequence that is not naturally found or expressed in C1. When referring to a subcellular localization signal, the term also describes a subcellular localization signal that is different from that naturally found in the respective protein.
[0113] The term "endogenous" when referring to a gene, enzyme, protein, or peptide sequence, such as a subcellular localization signal, refers to a gene, enzyme, protein, or peptide sequence that is naturally present in C1.
[0114] The term "exogenous," when referring to a polynucleotide, is used herein to describe a synthetic polynucleotide that is exogenously introduced into C1 via transformation. The exogenous polynucleotide can be introduced into C1 in a stable or transient manner to produce a ribonucleic acid (RNA) molecule and, subsequently, a polypeptide molecule.
[0115] Expression constructs and vectors The terms "expression construct", "DNA construct" or "expression cassette" are used interchangeably herein and refer to an artificially assembled or isolated nucleic acid molecule that contains a nucleic acid sequence encoding a protein of interest and is assembled such that the protein of interest is expressed in a target host cell. An expression construct typically contains appropriate regulatory sequences operably linked to the nucleic acid sequence encoding the protein of interest. An expression construct may further contain a nucleic acid sequence encoding a selection marker.
[0116] The terms "nucleic acid sequence," "nucleotide sequence," and "polynucleotide" are used herein to refer to polymers of deoxyribonucleotides (DNA), ribonucleotides (RNA), and modified forms thereof, either in the form of separate fragments or as components of larger constructs. A nucleic acid sequence can be a coding sequence, i.e., a sequence that codes for an intracellular end product, such as a protein. A nucleic acid sequence can also be a regulatory sequence, e.g., a promoter.
[0117] The terms "peptide," "polypeptide," and "protein" are used herein to refer to a polymer of amino acid residues. The term "peptide" typically refers to an amino acid sequence of 2 to 50 amino acids, while "protein" refers to an amino acid sequence of more than 50 amino acid residues.
[0118] Sequences (such as nucleic acid and amino acid sequences) that are "homologous" to a reference sequence herein refer to the percent identity between the sequences, the percent identity being at least 75%, preferably at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99%. Each possibility represents a separate embodiment of the invention. Homologs of the sequences described herein are encompassed by the invention. Protein homologs are encompassed so long as they maintain the activity of the original protein. Homologous nucleic acid sequences include mutations associated with codon usage and the degeneracy of the genetic code. Sequence identity can be determined using nucleotide / amino acid sequence comparison algorithms as known in the art.
[0119] Nucleic acid sequences encoding the polypeptides of the invention can be optimized for expression. Examples of such sequence modifications include, but are not limited to, altered G / C content to more closely resemble that typically found in Th. heterothallica, commonly referred to as codon optimization, and elimination of codons atypically found in this fungus.
[0120] "Codon optimization" refers to the process of modifying a nucleic acid sequence to enhance expression in a host cell of interest by selecting appropriate DNA nucleotides within a structural gene or fragment thereof that approach the codon usage in the organism of interest, and / or replacing at least one codon of the native sequence (e.g., more than about or about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more codons) with a codon more frequently or most frequently used in the genes of that host cell while maintaining the native amino acid sequence. Different species show a particular bias for certain codons of certain amino acids. Codon bias (differences in codon usage among organisms) often correlates with the efficiency of translation of messenger RNA (mRNA), which is believed to depend, among other things, on the properties of the codon being translated and the availability of certain transfer RNA (tRNA) molecules. The dominance of selected tRNAs in a cell is generally a reflection of the codons most frequently used in protein synthesis. Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Thus, an optimized gene or nucleic acid sequence refers to a gene in which the nucleotide sequence of a native or naturally occurring gene has been modified to utilize statistically preferred or statistically favored codons in the organism. The present invention expressly encompasses polynucleotides encoding the enzymes of interest disclosed herein that are codon optimized for expression in Th. heterothallica.
[0121] The term "regulatory sequences" refers to DNA sequences that control the expression (transcription) of coding sequences, such as promoters and terminators.
[0122] The term "promoter" refers to a regulatory DNA sequence that controls or directs the transcription of another DNA sequence in vivo or in vitro. Usually, promoters are located in the 5' region of the transcribed sequence (i.e., preceding, located upstream). Promoters may be derived in their entirety from natural sources or be composed of different elements derived from different promoters found in nature, or even contain synthetic nucleotide segments. Promoters may be constitutive (i.e., promoter activation is not regulated by an inducing agent, and thus the transcription rate is constant) or inducible (i.e., promoter activation is regulated by an inducing agent). In most cases, the exact boundaries of regulatory sequences have not been completely defined, and in some cases cannot be completely defined, so that several variants of DNA sequences may have identical promoter activity.
[0123] The term "terminator" refers to another regulatory DNA sequence that controls the termination of transcription. The terminator sequence is operably linked to the 3' end of the nucleic acid sequence to be transcribed.
[0124] The terms "Th. heterothallica promoter" and "Th. heterothallica terminator" refer to promoter and terminator sequences suitable for use in Th. heterothallica, i.e. capable of directing gene expression in Th. heterothallica. In some particular embodiments, the C1 promoter and C1 terminator are used, which refer to promoter and terminator sequences capable of directing gene expression in C1.
[0125] According to some embodiments, the Th. heterothallica promoter / terminator is derived from an endogenous gene of Th. heterothallica. According to other embodiments, the Th. heterothallica promoter / terminator is derived from a gene exogenous to Th. heterothallica.
[0126] Suitable constitutive promoters and terminators include those of the C1 glycolysis pathway genes, such as the phosphoglycerate kinase gene (PGK) (Uniprot: G2QLD8, NCBI Reference Sequence: XM_003665967), glyceraldehyde 3-phosphate dehydrogenase (GPD) (Uniprot: G2QPQ8, NCBI Reference Sequence: XM_003666768), phosphofructokinase (PFK) (Uniprot: G2Q605, NCBI Reference Sequence: XM_003659879), or those of the β-glucosidase 1 gene bgl1 (Accession Number: XM_003662656), or the triose phosphate isomerase (TPA) (Uniprot: G2Q605, NCBI Reference Sequence: XM_003659879). isomerase, TPI (Uniprot: G2QBR0, NCBI Reference Sequence: XM_003663200), or actin (ACT) (Uniprot: G2Q7Q5, NCBI Reference Sequence: XM_003662111), or the C1cbh1 promoter (GenBank AX284115) or C1chi1 promoter (GenBank HI550986). Additional promoters that can be used are the Aspergillus nidulans gpdA promoter, and the synthetic promoters described in Rantasalo et al. (2018 NAR 46(18):e111). Synthetic promoters that can be used with the present invention are further described in WO 2017 / 144777. As exemplary terminators, the terminator of the C1 chitinase 1 gene chi1 (GenBank HI550986), the cellobiohydrolase 1cbh1 (GenBank AX284115), or the yeast adh1 terminator can be used.
[0127] Exemplary promoter and terminator sequences, and promoter / terminator pairs, are provided in the Examples section below. In some embodiments, promoter sequences for use with the present invention are selected from the group consisting of promoter-8, bgl8 promoter, promoter-9, promoter-3, promoter-1, and TEF1A promoter, as exemplified below. Each possibility represents a separate embodiment of the present invention.
[0128] The term "operably linked" means that a selected nucleic acid sequence is adjacent to regulatory elements (promoter or terminator) such that the regulatory elements are capable of regulating expression of the selected nucleic acid sequence.
[0129] The terms "localization signal," "localization sequence," "intracellular targeting peptide / signal / sequence," and the like, are used interchangeably herein and refer to short peptide sequences (usually 5-30 amino acids in length) contained within a protein sequence (typically present at one end of the protein, such as the N-terminus) that direct the protein to a specific subcellular localization within the cell. For example, a Golgi localization signal targets a protein to the Golgi apparatus. A "heterologous localization signal," e.g., a "heterologous Golgi localization signal," refers to a localization signal that is not naturally found in the protein. In some embodiments, "heterologous" refers to a localization signal from another organism.
[0130] In some embodiments, the localization signal of the protein expressed in Th. heterothallica according to the present invention is derived from an endogenous gene of Th. heterothallica. For example, in some embodiments, the Golgi localization signal from C1 protein KRE2a (ortholog of the M. thermophila genome (mycocosm.jgi.doe.gov) accession number 2300989) is used.
[0131] In other embodiments, the localization signal of the protein expressed in Th. heterothallica according to the present invention is derived from a gene exogenous to Th. heterothallica (heterologous localization signal). For example, in some embodiments, the animal-derived enzymes expressed in Th. heterothallica according to the present invention are expressed with their own naturally occurring Golgi localization signal. As another example, in some embodiments, a Golgi localization signal from a yeast protein, such as the S. cerevisiae protein KRE2 (GenBank accession number CAA44516), is used.
[0132] According to some embodiments, the protein expressed in Th. heterothallica comprises an ER targeting sequence. In certain embodiments, the ER targeting sequence is the sequence HDEL.
[0133] As used herein, the term "in frame," when referring to one or more nucleic acid sequences, indicates that the sequences are linked such that their correct reading frame is maintained.
[0134] An expression construct according to some embodiments of the invention comprises a Th. heterothallica promoter sequence and a Th. heterothallica terminator sequence operably linked to a nucleic acid sequence encoding an enzyme such as flippase, GNT1, or GNT2. In some particular embodiments, an expression construct of the invention comprises a C1 promoter sequence and a C1 terminator sequence operably linked to a nucleic acid sequence encoding an enzyme such as flippase, GNT1, or GNT2.
[0135] Specific expression constructs may be assembled by a variety of different methods, including conventional molecular biology methods such as polymerase chain reaction (PCR), restriction endonuclease digestion, in vitro and in vivo assembly methods, and gene synthesis methods, or combinations thereof. Exemplary expression constructs and methods for constructing them are provided in the Examples section below.
[0136] alg3 deletion Gene deletion techniques allow for the partial or complete removal of a gene, thereby eliminating its expression. In such methods, gene deletion can be achieved by homologous recombination using a plasmid constructed to contiguously contain the 5' and 3' regions flanking the gene.
[0137] Gene deletion can also be performed by inserting a disruptive nucleic acid construct, also referred to herein as a deletion construct, into a gene. The disruptive construct can simply be a selectable marker gene with 5' and 3' regions of homology to the gene. The selectable marker allows for the identification of transformants containing the disrupted gene. Alternatively or additionally, the disruptive nucleic acid construct can include one or more polynucleotides encoding heterologous proteins to be expressed in the host cell.
[0138] Exemplary deletion constructs for alg3 and procedures for performing the deletions are described, for example, in WO 2021 / 094935. Deletions can be confirmed using PCR with appropriate primers flanking the disruptive construct.
[0139] In some embodiments, the Th. heterothallica of the present invention is genetically modified to express a heterologous flippase. In some particular embodiments, the heterologous flippase is a yeast flippase. In further particular embodiments, the yeast flippase is the S. cerevisiae mutant flippase FLC2p, which is a C-terminally truncated version of the S. cerevisiae ER-localized flippase FLC2.
[0140] In additional embodiments, the Th. heterothallica of the present invention is genetically modified to overexpress endogenous Th. heterothallica RFT1 flippase. In some specific embodiments, Th. heterothallica C1 is genetically modified to overexpress endogenous C1 RFT1 flippase. Overexpression of RFT1 in Th. heterothallica according to the present invention can be achieved by introduction of an exogenous polynucleotide encoding RFT1, comprising a nucleic acid sequence encoding RFT1 operably linked to a regulatory sequence operable in Th. heterothallica. An exemplary nucleotide sequence of rft1 is set forth in SEQ ID NO:7. An exemplary amino acid sequence of rft1 is set forth in SEQ ID NO:8.
[0141] In some exemplary embodiments, Th. heterothallica is genetically modified to delete or disrupt alg3, express mannosidase 1 (alpha-1,2-mannosidase) and ER-targeted C1 glucosidase 2 alpha-subunit, overexpress endogenous Th. heterothallica RFT1 flippase, and further to express animal-derived GNT1 and animal-derived GNT2 containing heterologous Golgi localization signals, e.g., human GNT1 and human GNT2 containing the Golgi localization signal from yeast protein KRE2.
[0142] In some exemplary embodiments, Th. heterothallica has been genetically modified to delete or disrupt alg3, express mannosidase 1 (alpha-1,2-mannosidase) and glucosidase 2 alpha-subunit, express yeast FLC2p flippase, and further to express animal-derived GNT1 and animal-derived GNT2 containing heterologous Golgi localization signals, e.g., human GNT1 containing the Golgi localization signal from Th. heterothallica protein KRE2, and rat GNT2.
[0143] In additional exemplary embodiments, Th. heterothallica is genetically modified by deletion or disruption of alg3, expression of mannosidase 1 (alpha-1,2-mannosidase) and ER-targeted C1 glucosidase 2 alpha-subunit, overexpression of endogenous Th. heterothallica RFT1 flippase, expression of human GNT1 containing the Th. heterothallica KRE2a Golgi localization signal, and rat GNT2, and expression of Leishmania major STT3. In some embodiments, such Th. heterothallica is further genetically modified by expression of human galactosyltransferase or Xenopus tropicalis galactosyltransferase.
[0144] In additional exemplary embodiments, Th. heterothallica is genetically modified by deletion or disruption of alg3, expression of mannosidase 1 (alpha-1,2-mannosidase) and ER-targeted C1 glucosidase 2 alpha-subunit, overexpression of endogenous Th. heterothallica RFT1 flippase, expression of bovine GNT1 containing the Th. heterothallica KRE2 Golgi localization signal, and rat GNT2, and expression of Leishmania major STT3.
[0145] Mannosidase 1 (alpha-1,2-mannosidase) and glucosidase 2 alpha-subunit According to some embodiments, expression constructs for ER-targeted mannosidase 1 (alpha-1,2-mannosidase) and ER-targeted C1 glucosidase 2 alpha-subunit are integrated into the alp3 protease locus of Thermothelomyces heterothallica.
[0146] An exemplary nucleotide sequence of an ER-targeted Trichoderma reesei mannosidase 1 is set forth in SEQ ID NO: 1. An exemplary amino acid sequence of an ER-targeted Trichoderma reesei mannosidase 1 is set forth in SEQ ID NO: 2.
[0147] An exemplary nucleotide sequence of an ER-targeted C1 glucosidase 2 alpha-subunit (gls2a-HDEL) is set forth in SEQ ID NO: 3. An exemplary amino acid sequence of an ER-targeted C1 glucosidase 2 alpha-subunit is set forth in SEQ ID NO: 4.
[0148] GNT1 and GNT2 In some embodiments, the Th. heterothallica of the present invention is genetically modified to express heterologous GNT1 and GNT2. In some embodiments, the heterologous GNT1 and GNT2 are of animal origin. As used herein, "animal origin" includes mammalian origin, including, for example, pet animals such as dogs and cats, and additional mammals such as horses. As exemplified herein below, animal origin includes, for example, rat origin. The term "animal origin" further includes human origin, as further exemplified below.
[0149] Heterologous GNT1 and GNT2 can be expressed in Th. heterothallica according to the present invention by introduction of one or more exogenous polynucleotides encoding GNT1 and GNT2, comprising nucleic acid sequences encoding GNT1 and GNT2 operably linked to regulatory sequences operable in Th. heterothallica. In some embodiments, the nucleic acid sequences encoding GNT1 and GNT2 are included in a single polynucleotide that is introduced into Th. heterothallica. In other embodiments, the nucleic acid sequences encoding GNT1 and GNT2 are included in two different polynucleotides that are introduced into Th. heterothallica.
[0150] In some embodiments, GNT1 is expressed in Th. heterothallica with its own naturally occurring Golgi localization signal. In other embodiments, GNT1 is expressed in Th. heterothallica with a heterologous Golgi localization signal.
[0151] In some embodiments, the heterologous Golgi localization signal is a yeast Golgi localization signal. In some specific embodiments, the heterologous Golgi localization signal is from yeast protein KRE2 alpha-1,2-mannosyltransferase. In some exemplary embodiments, the heterologous Golgi localization signal is from KRE2 of S. cerevisiae.
[0152] In other embodiments, the heterologous Golgi localization signal is from a filamentous fungus. In some embodiments, the heterologous Golgi localization signal is from Th. heterothallica. In some particular embodiments, the heterologous Golgi localization signal is from the C1 homologue of yeast protein KRE2.
[0153] In some embodiments, the GNT1 is human GNT1. In some embodiments, the human GNT1 introduced into Th. heterothallica comprises a heterologous Golgi localization signal. In some embodiments, the GNT1 is human GNT1 comprising a yeast Golgi localization signal. In some particular embodiments, the GNT1 is human GNT1 comprising a Golgi localization signal from the S. cerevisiae protein KRE2. An exemplary nucleotide sequence of KRE2 signal-GNT1 is set forth in SEQ ID NO:5. An exemplary amino acid sequence of KRE2 signal-GNT1 is set forth in SEQ ID NO:6.
[0154] In some embodiments, GNT2 is rat GNT2. GNT2 is typically expressed with its own naturally occurring Golgi localization signal. The amino acid sequence of rat GNT2 is set forth in SEQ ID NO: 16. An exemplary nucleic acid sequence of the polynucleotide for use according to the present invention that encodes rat GNT2 is set forth in SEQ ID NO: 15.
[0155] In other embodiments, the GNT2 is human GNT2. GNT2 is typically expressed with its own naturally occurring Golgi localization signal.
[0156] Exemplary, but non-limiting, combinations of GNT1 and GNT2 according to the present invention include: - human GNT1 with the yeast KRE2 Golgi localization signal, and human GNT2; - human GNT1 with the yeast KRE2 Golgi localization signal, and rat GNT2; -C1 human GNT1 with KRE2a Golgi localization signal, and human GNT2; -C1 human GNT1 with KRE2a Golgi localization signal, and rat GNT2; -C1 bovine GNT1 with the KRE2a Golgi localization signal, and rat GNT2.
[0157] Each combination represents a separate embodiment of the present invention.
[0158] Galactosyltransferase In some embodiments, the Th. Heterothallica of the present invention is genetically modified to express a heterologous galactosyltransferase. In some embodiments, the heterologous galactosyltransferase is of animal origin.
[0159] The galactosyltransferase can be expressed in Th. heterothallica according to the present invention by introduction of an exogenous polynucleotide encoding a galactosyltransferase, comprising a nucleic acid sequence encoding a galactosyltransferase operably linked to a regulatory sequence operable in Th. heterothallica.
[0160] In some embodiments, the galactosyltransferase is expressed in Th. heterothallica with its own naturally occurring Golgi localization signal, hi other embodiments, the galactosyltransferase is expressed in Th. heterothallica with a heterologous Golgi localization signal.
[0161] In some embodiments, the galactosyltransferase is human galactosyltransferase (huGalT1). In some embodiments, the human galactosyltransferase introduced into Th. heterothallica comprises a heterologous Golgi localization signal. In some particular embodiments, the human galactosyltransferase comprises the S. cerevisiae KRE2 Golgi localization signal. The amino acid sequence of a human galactosyltransferase having a Golgi localization signal from the S. cerevisiae protein KRE2 is set forth in SEQ ID NO: 14. An exemplary nucleic acid sequence of a polynucleotide used in accordance with the present invention encoding a human galactosyltransferase having a Golgi localization signal from the S. cerevisiae protein KRE2 is set forth in SEQ ID NO: 13.
[0162] According to other embodiments, the galactosyltransferase is from Xenopus tropicalis (XtGalT1). In some embodiments, the galactosyltransferase from Xenopus tropicalis introduced into Th. heterothallica comprises a heterologous Golgi localization signal. In some particular embodiments, the galactosyltransferase from Xenopus tropicalis comprises a Golgi localization signal from the protein KRE2 of S. cerevisiae.
[0163] STT3 oligosaccharyltransferase In some embodiments, the Th. heterothallica of the present invention is genetically modified to express a heterologous STT3 subunit of oligosaccharyltransferase. In some specific embodiments, the heterologous STT3 is Leishmania major STT3. The amino acid sequence of Leishmania major STT3 is set forth in SEQ ID NO: 12. Leishmania major STT3 can be expressed in Th. heterothallica according to the present invention by introducing an exogenous polynucleotide encoding Leishmania major STT3, comprising a nucleic acid sequence encoding Leishmania major STT3 operably linked to a regulatory sequence operable in Th. heterothallica. An exemplary nucleic acid sequence encoding Leishmania major STT3 for use according to the present invention is set forth in SEQ ID NO: 11.
[0164] Genetically engineered Th. heterothallica Th. heterothallica cells genetically engineered to produce glycoproteins having N-glycans of mammalian proteins (particularly human and pet proteins) according to the present invention are generated by modifying (e.g. deleting) endogenous genes of Th. heterothallica alg3 such that the genes do not produce functional proteins, and expressing exogenous polynucleotides encoding various enzymes.
[0165] It is understood that the genetic modifications of Th. heterothallica disclosed herein do not necessarily require modification of each and every cell of the genetically modified Th. heterothallica, so long as the desired outcome disclosed herein of production of glycoproteins having N-glycans of mammalian proteins, particularly human and companion animal proteins, is achieved.
[0166] In some embodiments, Th. heterothallica is further genetically modified to express a heterologous mammalian glycoprotein, hi some embodiments, the heterologous mammalian glycoprotein is an antibody or an antigen-binding fragment thereof.
[0167] In some exemplary embodiments, Th. heterothallica is genetically modified to express nivolumab or an antigen-binding fragment thereof. In additional exemplary embodiments, Th. heterothallica is genetically modified to express a nivolumab light chain with a Th. heterothallica CBH1 signal sequence and a nivolumab heavy chain with a Th. heterothallica CBH1 signal sequence.
[0168] In some embodiments, the present invention provides Th. heterothallica cells genetically modified as disclosed herein.
[0169] The expression of the exogenous polynucleotide is carried out by introducing an expression construct comprising a nucleic acid encoding the protein to be expressed in C1 into the Th. heterothallica cell, in particular into the nucleus of the Th. heterothallica cell. In particular, the genetic modification according to the present invention refers to the integration of the expression construct into the host genome.
[0170] Introduction of the expression construct into Th. heterothallica cells, ie, transformation of Th. heterothallica, can be carried out by methods for transforming filamentous fungi.
[0171] To facilitate easy selection of transformed cells, a selection marker can be transformed into the Th. heterothallica cells. A "selection marker" refers to a polynucleotide that encodes a gene product that confers a certain type of phenotype not present in untransformed cells, such as antibiotic resistance (resistance marker), the ability to utilize a certain resource (utilization / auxotrophic marker), or the expression of a reporter protein that can be detected, for example, by spectrometry. Auxotrophic markers are typically preferred as a means of selection in the food or pharmaceutical industry. The selection marker can be on a separate polynucleotide cotransformed with the expression construct, or on the same polynucleotide of the expression construct. Following transformation, positive transformants are selected, for example, by culturing the C1 cells on a selective medium according to the selected selection marker. In some cases, a split marker system is used in which the selection marker is split into two plasmids and a functional selection marker is only formed when the two plasmids are cotransformed and joined together via homologous recombination.
[0172] When a synthetic expression system is used, an expression cassette encoding a suitable synthetic transcription factor (sTF) is introduced into a host cell.
[0173] The transformed DNA can be integrated into the Th. heterothallica chromosome through homologous recombination or non-homologous end joining. To facilitate targeted integration into a specific locus in the genome, sequences corresponding to the target locus are incorporated into the same polynucleotide as the expression construct.
[0174] Selected clones are then grown and investigated for production of proteins with the desired N-glycoforms. The genetically modified Th. heterothallica is cultured in a medium under suitable conditions. According to certain embodiments, the fungus is grown at a temperature ranging from about 20° C. to about 45° C. and a medium pH of about 4.0 to about 8.0. The particular medium type may be selected according to the regulatory requirements of the final product. The produced glycoproteins may be isolated and analyzed.
[0175] The expression of GNT1, GNT2, and optionally additional enzymes such as galactosyltransferase in Th. heterothallica can be determined by structural analysis of the N-glycans produced by C1.
[0176] The genetically modified Th. heterothallica according to the present invention produces G0 (Man3GlcNAc2) as the final N-glycan structure or as an intermediate N-glycan structure.
[0177] In some embodiments, Th. heterothallica genetically modified by deletion or disruption of alg3, expression of ER-targeted Trichoderma reesei mannosidase 1 (alpha-1,2-mannosidase) and ER-targeted C1 glucosidase 2 alpha-subunit, optionally expression of a heterologous flippase or overexpression of an endogenous flippase, expression of heterologous GNT1 and GNT2, and optionally expression of a heterologous STT3 oligosaccharyltransferase, produces secreted glycoproteins, where G0 constitutes at least 80% of the N-glycans on the secreted glycoprotein, preferably at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or even at least 95% of the N-glycans on the secreted glycoprotein. Each value represents a separate embodiment of the invention.
[0178] In some embodiments, Th. heterothallica further genetically modified to express a heterologous galactosyltransferase produces secreted glycoproteins, wherein G1 and G2 (the sum of both G1 and G2) constitute at least 75% of the N-glycans on the secreted glycoprotein, preferably at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or even at least 95% of the N-glycans on the secreted glycoprotein, each value representing a separate embodiment of the invention.
[0179] In some embodiments, the genetic modification of Th. heterothallica does not include expression of a heterologous oligosaccharyl transferase (OST).
[0180] The following examples are presented to more fully illustrate certain embodiments of the present invention. However, they should in no way be construed as limiting the broad scope of the present invention. Those skilled in the art can easily devise many variations and modifications of the principles disclosed herein without departing from the scope of the present invention. EXAMPLES
[0181] Example 1 - Generating Nivolumab producing strains in three steps using humanized G1 / 2 glycans The first step in generating a Thermothelomyces heterothallica C1 strain producing human-type galactosylated glycans was the deletion of the alg3 gene from a strain carrying a deletion of eight protease genes and expressing nivolumab (as described in WO 2021 / 094935). After the alg3 deletion, the next C1 glycoengineering step was to integrate ER-targeted Trichoderma reesei mannosidase 1 (alpha-1,2-mannosidase) and ER-targeted C1 glucosidase 2 alpha-subunit into the alp3 protease locus to trim the glycan precursor to DolP-GlcNAc2-Man3, which is suitable for synthesis of human-type glycans. Expression cassettes containing human or bovine GNT1 (GlcNAc transferase I), rat GNT2 (GlcNAc transferase II), C1 flippase RFT1, oligosaccharyltransferase STT3 from Leishmania major, and human GalT1 (galactosyltransferase I) were then integrated into the alp6 protease locus. Expression of these genes results in the addition of GlcNAc residues onto both branches of the GlcNAc2Man3 glycan resulting in the G0 (GlcNAc2Man3GlcNAc2) glycan, followed by the addition of galactose residues onto one or both branches of the G0 glycan resulting in the production of G1 (GlcNAc2Man3GlcNAc2Gal) and G2 (GlcNAc2Man3MlcNAc2Gal2) glycans.
[0182] In the first transformation step, DNA constructs designed to integrate into the alp3 locus (JGI M. thermophila genome database ID 2306020) and simultaneously express T. reesi (Tr) mns1-HDEL (JGI T. reesei genome database ID 45717) and C1 gls2a-HDEL (JGI M. thermophila genome database ID 2125259) were assembled in two parts into two separate plasmids. HDEL is a four amino acid C-terminal ER localization signal. The first plasmid contained the alp3 5' flanking region fragment for integration, an expression cassette for Tr mns1-HDEL with the gene between the C1 bgl8 promoter and terminator (JGI M. thermophila Genome Database ID 115968), a synthetic transcription factor (sTF, 3' for the synthetic promoter in the plasmid), a direct repeat to the C1 cbh1 terminator (JGI M. thermophila Genome Database ID 109566), and the first two-thirds of the amdS marker gene (encoding acetamidase from Aspergillus nidulans). The second plasmid contained the last 2 / 3 of the amdS marker, a direct repeat fragment targeted to the end of the sTF cassette (for amdS marker removal by recombination), an expression cassette for C1 gls2a-HDEL between the synthetic AnSES promoter (Rantasalo A et al. 2018, A universal gene expression system for fungi. Nucl. Acids Research 46(18):e111) and the chi1 terminator (JGI M. thermophila Genome Database ID 50608), and the alp3 3' flanking region fragment for integration. The amdS marker fragments in these two plasmids overlap each other, and this region undergoes homologous recombination in C1 between the plasmids at the same time that the 5' and 3' flanking region fragments recombine with genomic DNA on either side of the alp3 gene. Recombination between the selection marker fragments renders the marker gene functional, allowing transformants to grow under selection.
[0183] The construct containing the alp3 5' flanking region fragment, the expression cassette for Tr mns1-HDEL, sTF, and the first 2 / 3 of the amdS marker is depicted in SEQ ID NO: 17 (pMYT1288). The 5' flanking sequence corresponds to positions 9-1196 of SEQ ID NO: 17. The bgl8 promoter sequence corresponds to positions 1202-2593 of SEQ ID NO: 17. The nucleic acid sequence encoding the genomic Trichoderma reesei mannosidase 1 with an artificial HDEL ER retention signal corresponds to positions 2594-4420 of SEQ ID NO: 17. The nucleic acid sequence encoding the T. reesei MNS1 with an HDEL signal is also depicted as SEQ ID NO: 1 (Tr mns1-HDEL nt) and the amino acid sequence is depicted as SEQ ID NO: 2 (Tr MNS1-HDEL aa). The bgl8 terminator sequence corresponds to positions 4421-4887 of SEQ ID NO: 17. The nucleic acid sequence for the synthetic transcription factor cassette corresponds to positions 4901-6550 of SEQ ID NO: 17. The first 2 / 3 of the amdS marker gene corresponds to positions 7060-9126 of SEQ ID NO: 17. Fragments were generated by PCR, separated in agarose gel, purified and cloned into the backbone vector pRS426 by the Gibson Assembly (NEBuilder® HiFi DNA Assembly Cloning Kit, New England Biolabs) method to obtain the 5' arm expression vector pMYT1288, which was verified by sequencing.
[0184] The construct containing the last 2 / 3 of the amdS marker, the expression cassette for C1 gls2a-HDEL, and the alp3 3' flanking region fragment for integration is described in SEQ ID NO: 18 (pMYT0721). The 2 / 3 of the amdS marker gene corresponds to positions 17-1738 of SEQ ID NO: 18. The synthetic AnSES promoter sequence corresponds to positions 2255-2747 of SEQ ID NO: 18. The sequence encoding C1 glucosidase 2 alpha with an artificial HDEL ER retention signal corresponds to positions 2748-5760 of SEQ ID NO: 18. The nucleic acid sequence encoding C1 GLS2 alpha with HDEL signal is also described as SEQ ID NO: 3 (C1 gls2a-HDEL nt) and the amino acid sequence is described as SEQ ID NO: 4 (C1 GLS2A-HDEL aa). The chi1 terminator sequence corresponds to positions 5761-6406 of SEQ ID NO: 18. The alp3 3' flanking sequence corresponds to positions 6415-7548 of SEQ ID NO: 18. The fragment was generated by PCR, separated in an agarose gel, purified and cloned into the backbone vector pRS426 by the Gibson Assembly (NEBuilder® HiFi DNA Assembly Cloning Kit, New England Biolabs) method to obtain the 3' arm expression vector pMYT0721, which was verified by sequencing.
[0185] For the second transformation step, DNA constructs designed to integrate into the alp6 locus (JGI M. thermophila Genome Database ID 94536) and simultaneously express GNT1 and GNT2, and RFT1, STT3, and GalT1, were assembled in two parts onto two separate plasmids. The first plasmid contained an alp6 5' flanking region fragment for integration, an expression cassette for GNT1 with either the human GNT1 gene or the bovine GNT1 gene fused to the C1 KRE2 Golgi localization signal between the C1 bgl8 promoter and bgl8 terminator (JGI M. thermophila Genome Database ID 2300989), the C1 flippase RFT1 between the promoter and terminator of the ubiquitin-like protein gene (JGI M. thermophila Genome Database ID 2307799), and the first two-thirds of the C1 pyr4 marker gene (JGI M. thermophila Genome Database ID 2311494). The second plasmid contained the last two-thirds of the C1 pyr4 marker gene, a direct repeat fragment targeted to the beginning of the pyr4 cassette (for pyr4 marker removal by recombination), a reverse expression cassette for STT3 from Leishmania major between the promoter and terminator of the C1 hypothetical protein (JGI M. thermophila genome database ID 2315630), a reverse expression cassette for human GalT1 fused to the Saccharomyces cerevisiae KRE2 Golgi localization signal between the promoter of the ubiquitin-like protein gene and the bgl8 terminator, the GNT2 gene from rat between the promoter of translation elongation factor 1A (JGI M. thermophila genome database ID 2298136) and the terminator of the C1 hypothetical protein (JGI M. thermophila genome database ID 114107), and finally, the alp6 3' flanking region fragment for integration.The C1 pyr4 marker fragments in these two plasmids overlap each other and undergo homologous recombination at C1 between the plasmids while the 5' and 3' flanking fragments recombine with genomic DNA on either side of the alp6 gene. Recombination between the selection marker fragments renders the marker gene functional, allowing transformants to grow under selection.
[0186] The first 5' construct containing the alp6 5' flanking region fragment, an expression cassette for human GNT1, C1 RFT1, and the first 2 / 3 of the C1 pyr4 marker gene is set forth as SEQ ID NO:19 (pMYT1451). The 5' flanking sequence corresponds to positions 8-1156 of SEQ ID NO:19. The bgl8 promoter sequence corresponds to positions 1165-2556 of SEQ ID NO:19. The nucleic acid sequence encoding human GNT1 fused to the C1 KRE2 Golgi localization signal corresponds to positions 2557-4042 of SEQ ID NO:19, where positions 2557-2821 encode the C1 KRE2 localization signal and positions 2822-4042 encode human GNT1. The nucleic acid sequence encoding human GNT1 fused to the C1 KRE2 Golgi localization signal is also set forth as SEQ ID NO:5 (C1 kre2-huGNT1 nt). The nucleic acid sequence encoding human GNT1 was codon-optimized for M. thermophila C1 and synthesized by GenScript (USA). In the synthesized sequence, the region of amino acids 1-100 of human GNT1 was replaced by the C1 KRE2 Golgi localization signal (amino acids 1-70). The complete amino acid sequence of C1 KRE2-GNT1 is set forth as SEQ ID NO:6 (C1 KRE2-huGNT1 aa). The bgl8 terminator sequence corresponds to positions 4046-4512 of SEQ ID NO:19. The ubiquitin-like protein promoter sequence corresponds to positions 4513-5532 of SEQ ID NO:19. The C1 flippase rft1 gene corresponds to positions 5533-7449 of SEQ ID NO:19. The nucleic acid sequence encoding C1 RFT1 is also set forth as SEQ ID NO:7 (C1 rft1 nt) and the amino acid sequence is set forth as SEQ ID NO:8 (C1 RFT1 aa). The ubiquitin-like protein terminator sequence corresponds to positions 7458-7953 of SEQ ID NO: 19. The first two-thirds of the C1 pyr4 marker gene corresponds to positions 7969-9748 of SEQ ID NO:19.
[0187] A second 5' construct containing the alp6 5' flanking region fragment, an expression cassette for bovine GNT1, C1 RFT1, and the first 2 / 3 of the C1 pyr4 marker gene is set forth as SEQ ID NO:20 (pMYT1452). The 5' flanking sequence corresponds to positions 8-1156 of SEQ ID NO:20. The bgl8 promoter sequence corresponds to positions 1165-2556 of SEQ ID NO:20. The nucleic acid sequence encoding bovine GNT1 fused to the C1 KRE2 Golgi localization signal corresponds to positions 2557-4051 of SEQ ID NO:20, where positions 2557-2821 encode the C1 KRE2 localization signal and positions 2822-4051 encode bovine GNT1. The nucleic acid sequence encoding bovine GNT1 fused to the C1 KRE2 Golgi localization signal is also set forth as SEQ ID NO:9 (C1 kre2-boGNT1 nt). The nucleic acid sequence encoding bovine GNT1 was codon-optimized for M. thermophila C1 and synthesized by GenScript (USA). In the synthesized sequence, the region of amino acids 1-38 of bovine GNT1 was removed and replaced by the C1 KRE2 Golgi localization signal (amino acids 1-70) during cloning of the expression plasmid. The complete amino acid sequence of C1 KRE2-GNT1 is depicted as SEQ ID NO: 10 (C1 KRE2-boGNT1 aa). The bgl8 terminator sequence corresponds to positions 4052-4518 of SEQ ID NO: 20. The ubiquitin-like protein promoter sequence corresponds to positions 4524-5543 of SEQ ID NO: 20. The C1 flippase RFT1 gene described for pMYT1451 corresponds to positions 5544-7460 of SEQ ID NO: 20. The ubiquitin-like protein terminator sequence corresponds to positions 7469-7964 of SEQ ID NO: 20. The first two-thirds of the C1 pyr4 marker gene corresponds to positions 7980 to 9759 of SEQ ID NO:20.Fragments for the 5' plasmids were generated by PCR, separated in agarose gel, purified and cloned into the backbone vector pRS426 by the Gibson Assembly (NEBuilder® HiFi DNA Assembly Cloning Kit, New England Biolabs) method to obtain the 5' arm expression vectors pMYT1451 and pMYT1452, which were verified by sequencing.
[0188] The construct containing the last 2 / 3 of the C1 pyr4 marker gene, a reverse expression cassette for STT3 from Leishmania major, a reverse expression cassette for human GalT1, a reverse expression cassette for the rat GNT2 gene, and an alp6 3' flanking region fragment for integration is described in SEQ ID NO: 21 (pMYT1453). The 2 / 3 of the C1 pyr4 marker gene corresponds to positions 17-1273 of SEQ ID NO: 21. The hypothetical protein (ID2315630) promoter sequence corresponds to positions 5965-4957 of SEQ ID NO: 21. The sequence coding for STT3 from Leishmania major, codon-optimized for M. thermophila C1 and synthesized by GenScript (USA), corresponds to positions 4956-2383 of SEQ ID NO: 21. The nucleic acid sequence coding for the Leishmania major STT3 gene is also described as SEQ ID NO: 11 (LmSTT3 nt). The complete amino acid sequence of L. major STT3 is set forth as SEQ ID NO: 12 (LmSTT3 aa). The hypothetical protein (ID2315630) terminator sequence corresponds to positions 2374-1795 of SEQ ID NO: 21. The sequence encoding human GalT1 fused to the Sc KRE2 Golgi localization signal corresponds to positions 7704-6439 of SEQ ID NO: 21, where positions 7704-7405 encode the Sc KRE2 localization signal and positions 7404-6439 encode human GalT1. The ubiquitin-like protein promoter sequence corresponds to positions 8724-7705 of SEQ ID NO: 21. The nucleotide sequence encoding human GalT1 fused to the Sc KRE2 Golgi localization signal is also set forth as SEQ ID NO: 13 (Sckre2-huGalT1 nt). The nucleic acid sequence encoding human GalT1 was codon-optimized for M. thermophila C1 and synthesized by GenScript (USA). In the synthesized sequence, the region of amino acids 1-77 of human GalT1 was removed and replaced by Sc KRE2 Golgi localization signal (amino acids 1-100) during cloning into the expression plasmid. The complete amino acid sequence of Sc KRE2-GalT1 is set forth as SEQ ID NO: 14 (ScKRE2-huGalT1 aa).The bgl8 terminator sequence corresponds to positions 6438-5972 of SEQ ID NO:21. The translation elongation factor 1A promoter sequence corresponds to positions 11618-10562 of SEQ ID NO:21. The sequence encoding rat GNT2, codon-optimized for M. thermophila C1 and synthesized by GenScript (USA), corresponds to positions 10561-9233 of SEQ ID NO:21. The nucleic acid sequence encoding rat GNT2 is also set forth as SEQ ID NO:15 (rat GNT2 nt). The complete amino acid sequence of rat GNT2 is set forth as SEQ ID NO:16 (rat GNT2 aa). The hypothetical protein (ID114107) terminator sequence corresponds to positions 9232-8729 of SEQ ID NO:21. The alp6 3' flanking sequence corresponds to positions 11626-12681 of SEQ ID NO:21. Fragments were generated by PCR, separated in agarose gels, purified, and cloned into the backbone vector pRS426 by the Gibson Assembly (NEBuilder® HiFi DNA Assembly Cloning Kit, New England Biolabs) method to yield the 3′ arm expression vector pMYT1453, which was verified by sequencing.
[0189] To obtain a nivolumab producing strain with G1 / 2 N-glycans, the above expression plasmids were transformed sequentially into the pyr4-minus nivolumab producing alg3 deletion strain M3291 (described in WO 2021 / 094935). In each transformation, a pair of one 5' arm expression vector and one 3' arm expression vector (cut out of the expression plasmid backbone with MssI) was used. These pairs were: First round: SEQ ID NO: 17 (pMYT1288) (alp3 5' flanking region-Tr Mns1-HDEL cassette-sTF cassette-2 / 3 amdS) + SEQ ID NO: 18 (pMYT0721) (2 / 3 amdS-C1 gls2a-HDEL cassette-alp3 3' flanking fragment) Second round: SEQ ID NO: 19 (pMYT1451) or SEQ ID NO: 20 (pMYT1452) (alp6 5' flanking region-human or bovine GNT1 cassette-C1 RFT1 cassette-2 / 3 pyr4) + SEQ ID NO: 21 (pMYT1453) (2 / 3 pyr4-LmSTT3 cassette-human GalT1 cassette-rat GNT2 cassette-alp6 3' flanking fragment)
[0190] C1 transformation and transformant selection were performed essentially as described in WO 2021 / 094935. For the first round of transformation, selection was based on a functional amdS gene, and for the second round on a functional pyr4 gene. Transformants were screened by PCR to find clones in which the alp3 deletion site (first round) or the alp6 gene (second round) had been replaced by the construct.
[0191] Correct integration of the plasmid / gene into the transformant genome was verified using specific primers (primenr). The first round transformants with correct integration of the expression construct and additional deletion of the alp3 locus were kept as M4855 and M4856. Strain M4855 was used for the second round of transformation. The second round transformants with correct integration of the construct and deletion of the alp6 gene were kept as M5129 and M5130 (with human GNT1) and M5131 and M5132 (with bovine GNT1), respectively.
[0192] The C1 strains constructed from both transformation rounds were grown in 250 mL shake flasks in 50 mL liquid medium as described for the shake flask culture of parent strain M3291 in WO 2021 / 094935. The medium culture was carried out at 35° C. and approximately 200 RPM for 4 days. The mycelium was removed by centrifugation and the supernatant from the culture was used for Protein A affinity purification of nivolumab using an AKTA Start automated HPLC system (Cytiva) and a prepacked 1 mL column HiTrap MabSelect Sure or HiTrap MabSelect PrismA according to the manufacturer's (Cytiva) instructions. Peak fractions from all samples were subjected to analysis of released N-glycans as described in Example 1 of WO 2021 / 094935. The results are summarized in Figure 1 for all strains. FIG. 1A shows that by expressing both Tr MNS1-HDEL and C1 GLS2A-HDEL, the predominant N-glycan on nivolumab is Man3 (M3) at nearly 90% abundance, while there is little higher mannose structure (GlcNAc2Man4, GlcNAc2Man5) and virtually no Hex6 (GlcNAc2Man5Glc) that is thought to block N-glycan modification beyond G0. FIG. 1B-D show that by using the alp6-targeted expression vector described in this example, very high G1 (GlcNAc2Man3GlcNAc2Gal) and G2 (GlcNAc2Man3GlcNAc2Gal2) N-glycan levels are reached on the target mAb. Conversion to human G1 / 2 N-glycans is 88-98%, i.e., nearly complete conversion to human glycans by this approach. In addition to the very high levels of human-type N-glycans, Figure 2 shows the amount of released N-glycans that was normalized between samples using a constant amount of internal standard that was added to each sample for analysis.The percentage of N-glycans released from the different nivolumab samples was at similar levels to the reference Opdivo (commercially available nivolumab produced in a mammalian system, Bristol-Myers Squibb), suggesting that this approach to humanize glycosylation does not adversely affect N-glycosylation site occupancy.
[0193] C1 strains M3291, M5130, and M5132 were then cultured in ambr250 or 1 L bioreactors using a fed-batch process in medium with yeast extract as the organic nitrogen source and glucose as the carbon source. The medium cultures were performed at 38 °C for 7 days. After the end of the culture, the mycelium was removed by centrifugation at 4000 g for 20 min, phenylmethylsulfonyl fluoride (PMSF) was added at a concentration of 1-2 mM to inhibit protease activity in the resulting medium culture supernatants, and the samples were stored at -80 °C. Nivolumab was purified from the day 7 fermentation samples using essentially the same Protein A affinity purification method as described above for the purification of the shake flask samples. Peak fractions from all samples were subjected to released N-glycan analysis. The results are summarized in Figure 3 for all three strains. FIG. 3A shows that the parental strain M3291 has more than 80% of the desired precursor Man3 (M3), but also more than 3.5% Hex6 (M6), which is believed to prevent N-glycan humanization (conversion to G0 and further modification) and some higher mannose structures (Man4, Man5). FIG. 3B-D shows that the released N-glycans on the target mAb are almost completely humanized, i.e., more than 99% of the detected N-glycans belong to the G0-G2 species. FIG. 4 shows the amount of released N-glycans, which is normalized between samples using a constant amount of internal standard (sialylglycan peptide) added to each sample for analysis. As with shake flasks, the percentage of released N-glycans from the two nivolumab strains with humanized N-glycans appears to be at the same level as the reference Opdivo (commercially available nivolumab produced in a mammalian system), further confirming the absence of adverse effects on N-glycan site occupancy.
[0194] Example 2 - Generating an empty strain in one step using humanized G1 / 2 glycans This second approach combined all three glycomodification steps required to generate a Thermothelomyces heterothallica C1 strain producing human-type galactosylated glycans, as described in Example 1. Briefly, deletion of the alg3 gene from a strain carrying a deletion of the 14 protease genes and the kex2 protease gene under weaker promoters was combined with integration of ER-targeted Trichoderma reesei mannosidase 1 (alpha-1,2-mannosidase), ER-targeted C1 glucosidase 2 alpha-subunit, human GNT1 (GlcNAc transferase I), rat GNT2 (GlcNAc transferase II), oligosaccharyltransferase STT3 from Leishmania major, and human GalT1 (galactosyltransferase I) into the alg3 locus. Expression of these genes results first in the trimming of DolP-GlcNAc2-Man3, a glycan precursor suitable for the synthesis of human-type glycans, followed by the addition of GlcNAc residues onto both branches of the GlcNAc2Man3 glycan resulting in the G0 (GlcNAc2Man3GlcNAc2) glycan, followed by the addition of galactose residues onto one or both branches of the G0 glycan resulting in the production of G1 (GlcNAc2Man3GlcNAc2Gal) and G2 (GlcNAc2Man3MlcNAc2Gal2) glycans.
[0195] DNA constructs designed to integrate into the alg3 locus (JGI M. thermophila genome database ID 2310419) and simultaneously express T. reesei (Tr) mns1-HDEL (JGI T. reesei genome database ID 45717), C1 gls2a-HDEL (JGI M. thermophila genome database ID 2125259), human GNT1, rat GNT2, as well as Leishmania major (Lm) STT3 and human GalT1 were assembled in two parts onto two separate plasmids. HDEL represents a four amino acid C-terminal ER localization signal.
[0196] The first plasmid contained an alg3 5' flanking region fragment for integration, an expression cassette for GNT1 in which the human GNT1 gene is fused to the C1 KRE2 Golgi localization signal (JGI M. thermophila genome database ID 2300989) between the C1 bgl8 promoter and the bgl8 terminator (JGI M. thermophila genome database ID 115968), an expression cassette for human GalT1 in which the Saccharomyces cerevisiae KRE2 Golgi localization signal (JGI M. thermophila genome database ID 2315548) between the promoter of the ubiquitin-like protein gene (JGI M. thermophila genome database ID 2302731) and the terminator of the C1 hypothetical protein (JGI M. thermophila genome database ID 2302731), and a synthetic AnSES promoter (Rantasalo A et al. 2018, A universal gene expression system for fungi.Nucl.Acids Research The plasmid contained an expression cassette for C1 gls2a-HDEL between the nucleotide sequence of the recombinant M. thermophila genome (nucleotide sequence 46(18):e111) and the chi1 terminator (JGI M. thermophila Genome Database ID 50608), an expression cassette for a synthetic transcription factor (sTF, for the synthetic promoter of C1 gls2a-HDEL), and the first two-thirds of the pyr4 marker gene (JGI M. thermophila Genome Database ID 2311494).
[0197] The second plasmid contained the last two-thirds of the C1 pyr4 marker gene, a direct repeat fragment targeted to the beginning of the pyr4 cassette (for pyr4 marker removal by recombination), a reverse expression cassette for STT3 from Leishmania major between the promoter and terminator of the C1 hypothetical protein (JGI M. thermophila Genome Database ID 2315630), a reverse expression cassette for Tr mns1-HDEL, whose gene is between either the C1 bgl8 promoter (JGI M. thermophila Genome Database ID 115968) or the promoter of the ubiquitin-like protein gene (JGI M. thermophila Genome Database ID 2315548) and the ubiquitin-like protein gene terminator (JGI M. thermophila Genome Database ID 2315548), a reverse expression cassette for Tr mns1-HDEL, whose gene is between the promoter of translation elongation factor 1A (JGI M. thermophila Genome Database ID 2298136) and the terminator of the C1 hypothetical protein (JGI The plasmid contained a reverse expression cassette for the GNT2 gene from rat (M. thermophila genome database ID 114107) between the two plasmids, and finally, the alg3 3' flanking region fragment for integration. The C1 pyr4 marker fragments in these two plasmids overlap each other and this region undergoes homologous recombination in C1 between the plasmids at the same time that the 5' and 3' flanking region fragments recombine with genomic DNA on either side of the alg3 gene. Recombination between the selection marker fragments renders the marker gene functional, allowing transformants to grow under selection.
[0198] A construct containing expression cassettes for the alg3 5' flanking region fragment, human GNT1, human GalT1, C1 gls2a-HDEL, sTF, and the first 2 / 3 of the pyr4 marker is set forth in SEQ ID NO:23 (pMYT1974). The alg3 5' flanking sequence corresponds to positions 9-1010 of SEQ ID NO:23. The bgl8 promoter sequence corresponds to positions 1018-2409 of SEQ ID NO:23. The nucleic acid sequence encoding human GNT1 fused to the C1 KRE2 Golgi localization signal corresponds to positions 2410-3898 of SEQ ID NO:23, where positions 2410-2674 encode the C1 KRE2 localization signal and positions 2675-3898 encode human GNT1. The nucleic acid sequence encoding human GNT1 fused to the C1 KRE2 Golgi localization signal is also set forth as SEQ ID NO:5 (C1 kre2-huGNT1 nt). The nucleic acid sequence encoding human GNT1 was codon-optimized for M. thermophila C1 and synthesized by GenScript (USA). In the synthesized sequence, the region of amino acids 1-100 of human GNT1 was replaced by the C1 KRE2 Golgi localization signal (amino acids 1-70). The complete amino acid sequence of C1 KRE2-GNT1 is set forth as SEQ ID NO:6 (C1 KRE2-huGNT1 aa). The bgl8 terminator sequence corresponds to positions 3899-4365 of SEQ ID NO:23. The ubiquitin-like protein promoter sequence corresponds to positions 4374-5393 of SEQ ID NO:23. The sequence encoding human GalT1 fused to the Sc KRE2 Golgi localization signal corresponds to positions 5394-5693 of SEQ ID NO: 23, where positions 5394-5693 encode the Sc KRE2 localization signal and positions 5694-6659 encode human GalT1. The nucleotide sequence encoding human GalT1 fused to the Sc KRE2 Golgi localization signal is also set forth as SEQ ID NO: 13 (Sckre2-huGalT1 nt). The nucleic acid sequence encoding human GalT1 was codon-optimized for M. thermophila C1 and synthesized by GenScript (USA).In the synthesized sequence, the region of amino acids 1-77 of human GalT1 was removed and replaced by the Sc KRE2 Golgi localization signal (amino acids 1-100) during cloning of the expression plasmid. The complete amino acid sequence of Sc KRE2-GalT1 is described as SEQ ID NO: 14 (ScKRE2-huGalT1 aa). The terminator sequence of the hypothetical protein (ID2302731) corresponds to positions 6660-7065 of SEQ ID NO: 23. The synthetic AnSES promoter sequence corresponds to positions 7073-7565 of SEQ ID NO: 23. The sequence coding for C1 glucosidase 2 alpha with an artificial HDEL ER retention signal corresponds to positions 7566-10578 of SEQ ID NO: 23. The nucleic acid sequence coding for C1 GLS2 alpha with an HDEL signal is also described as SEQ ID NO: 3 (C1 gls2a-HDEL nt) and the amino acid sequence is described as SEQ ID NO: 4 (C1 GLS2A-HDEL aa). The chi1 terminator sequence corresponds to positions 10579-11224 of SEQ ID NO:23. The nucleic acid sequence for the synthetic transcription factor cassette corresponds to positions 11225-12874 of SEQ ID NO:23. The first 2 / 3 of the C1 pyr4 marker gene corresponds to positions 12888-14667 of SEQ ID NO:23. Fragments were generated by PCR, separated in agarose gel, purified, and cloned stepwise using yeast recombination methods into the backbone vector pRS426 to yield the final 5' arm expression vector pMYT1974, which was verified by sequencing.
[0199] Two parallel constructs containing the last 2 / 3 of the C1 pyr4 marker gene, direct repeats for marker removal, a reverse expression cassette for STT3 from Leishmania major, a reverse expression cassette for Tr Mns1-HDEL, a reverse expression cassette for the rat GNT2 gene, and an alg3 3' flanking region fragment for integration are described in SEQ ID NO:24 (pMYT1963) and SEQ ID NO:25 (pMYT1964). The last 2 / 3 of the C1 pyr4 marker gene corresponds to positions 17-1273 of SEQ ID NO:24. The direct repeat fragment targeted to the start of the pyr4 cassette corresponds to positions 1290-1786 of SEQ ID NO:24. The hypothetical protein (ID2315630) promoter sequence corresponds to positions 5965-4957 of SEQ ID NO:24. The sequence encoding STT3 from Leishmania major, codon-optimized for T. heterothallica C1 and synthesized by GenScript (USA), corresponds to positions 4956-2383 of SEQ ID NO:24. The nucleic acid sequence encoding the Leishmania major STT3 gene is also set forth as SEQ ID NO:11 (LmSTT3 nt). The complete amino acid sequence of L. major STT3 is set forth as SEQ ID NO:12 (LmSTT3 aa). The hypothetical protein (ID2315630) terminator sequence corresponds to positions 2374-1795 of SEQ ID NO:24. The ubiquitin-like protein promoter sequence corresponds to positions 9316-8297 of SEQ ID NO:24. The nucleic acid sequence encoding genomic Trichoderma reesei mannosidase 1 with an artificial HDEL ER retention signal corresponds to positions 8296-6470 of SEQ ID NO:24. The nucleic acid sequence encoding T. reesei MNS1 with HDEL signal is also set forth as SEQ ID NO:1 (Tr mns1-HDEL nt) and the amino acid sequence is set forth as SEQ ID NO:2 (Tr MNS1-HDEL aa). The ubiquitin-like protein terminator sequence corresponds to positions 6469-5974 of SEQ ID NO:24. The translation elongation factor 1A promoter sequence corresponds to positions 12210-11154 of SEQ ID NO:24.The sequence encoding rat GNT2, codon-optimized for T. heterothallica C1 and synthesized by GenScript (USA), corresponds to positions 11153-9825 of SEQ ID NO:24. The nucleic acid sequence encoding rat GNT2 is also set forth as SEQ ID NO:15 (rat GNT2 nt). The complete amino acid sequence of rat GNT2 is set forth as SEQ ID NO:16 (rat GNT2 aa). The hypothetical protein (ID114107) terminator sequence corresponds to positions 9824-9321 of SEQ ID NO:24. The alg3 3' flanking sequence corresponds to positions 12218-13217 of SEQ ID NO:24. The fragments were generated by PCR, separated in agarose gel, purified, and cloned stepwise using yeast recombination methods into the backbone vector pRS426 to obtain the final 3' arm expression vector pMYT1963, which was verified by sequencing.
[0200] In SEQ ID NO:25 (pMYT1964), the last 2 / 3 of the C1 pyr4 marker gene corresponds to positions 17-1273. The direct repeat fragment targeted to the beginning of the pyr4 cassette corresponds to positions 1290-1786 of SEQ ID NO:25. The hypothetical protein (ID2315630) promoter sequence corresponds to positions 5965-4957 of SEQ ID NO:25. The sequence encoding STT3 from Leishmania major, codon-optimized for T. heterothallica C1 and synthesized by GenScript (USA), corresponds to positions 4956-2383 of SEQ ID NO:25. The nucleic acid sequence encoding the Leishmania major STT3 gene is also set forth as SEQ ID NO:11 (LmSTT3 nt). The complete amino acid sequence of L. major STT3 is set forth as SEQ ID NO:12 (LmSTT3 aa). The hypothetical protein (ID2315630) terminator sequence corresponds to positions 2374-1795 of SEQ ID NO:25. The bgl8 promoter sequence corresponds to positions 9688-8297 of SEQ ID NO:25. The nucleic acid sequence encoding the genomic Trichoderma reesei mannosidase 1 with an artificial HDEL ER retention signal corresponds to positions 8296-6470 of SEQ ID NO:25. The nucleic acid sequence encoding the T. reesei MNS1 with an HDEL signal is also set forth as SEQ ID NO:1 (Tr mns1-HDEL nt) and the amino acid sequence is set forth as SEQ ID NO:2 (Tr MNS1-HDEL aa). The ubiquitin-like protein terminator sequence corresponds to positions 6469-5974 of SEQ ID NO:25. The translation elongation factor 1A promoter sequence corresponds to positions 12584-11528 of SEQ ID NO:25. The sequence encoding rat GNT2, codon-optimized for M. thermophila C1 and synthesized by GenScript (USA), corresponds to positions 11527-10199 of SEQ ID NO:25. The nucleic acid sequence encoding rat GNT2 is also set forth as SEQ ID NO:15 (rat GNT2 nt). The complete amino acid sequence of rat GNT2 is set forth as SEQ ID NO:16 (rat GNT2 aa). The hypothetical protein (ID114107) terminator sequence corresponds to positions 10198-9695 of SEQ ID NO:25.The alg3 3' flanking sequence corresponds to positions 12592-13591 of SEQ ID NO: 25. Fragments were generated by PCR, separated in agarose gel, purified, and cloned stepwise using yeast recombination methods into the backbone vector pRS426 to yield the final 3' arm expression vector pMYT1964, which was verified by sequencing.
[0201] To obtain two different empty strains (not expressing heterologous mammalian glycoproteins) with G1 / 2 N-glycans, the expression plasmids described above were transformed in a single step into a pyr4-minus strain carrying a deletion of the 14 protease genes and the kex2 protease gene under a weaker promoter. In each transformation, pairs of one 5'-arm and one 3'-arm expression vector (cut out of the expression plasmid backbone with MssI) were used. These pairs were SEQ ID NO:23 (pMYT1974; alg3 5' flanking fragment-human GNT1 cassette-human GalT1 cassette-C1 gls2a-HDEL cassette-sTF cassette-2 / 3 pyr4) + SEQ ID NO:24 (pMYT1963) or SEQ ID NO:25 (pMYT1964) (2 / 3 pyr4-DR-LmSTT3 cassette-TrMns1-HDEL cassette-rat GNT2 cassette-alg3 3' flanking fragment). C1 transformation and transformant selection were performed essentially as described in WO2021 / 094935. Transformant selection was based on a functional pyr4 gene. Transformants were screened by PCR to find clones in which the alg3 gene was replaced by the construct. Correct integration of the plasmid / gene into the transformant genome was verified using specific primers. Transformants with correct integration of the construct and deletion of the alg3 gene were saved as M6589 and M6596 (with the ubiquitin-like protein promoter for TrMns1-HDEL) and M6590 and M6597 (with the bgl8 promoter for TrMns1-HDEL), respectively.
[0202] The C1 strain constructed from both transformations was grown in 3.5 mL liquid medium in 24-well deep well plates as described for culturing the strains in WO 2021 / 094935. The medium culture was carried out for 4 days at 35°C, 800 RPM and 80% humidity. The mycelium was removed by centrifugation and the supernatant from the culture was subjected to analysis of released N-glycans as described in Example 1 of WO 2021 / 094935. The results are summarized in Figure 5 for all four strains. Figure 5 shows that by expressing all six genes affecting glycan structure from the alg3 locus, the predominant N-glycans on all secreted proteins are more than 80% of human type (N-glycans belonging to species G0 to G2) with no detectable Hex6 (GlcNAc2Man5Glc) which is thought to block N-glycan modification to G0 and beyond. Also of note is the abundance of the final G2 N-glycan. The amount (%) of human N-glycans on all secreted proteins was found to be consistently less than that observed on any purified target protein, such as nivolumab in Example 1.
[0203] array SEQ ID NO:1 - T. reesei mns1-HDEL nt (coding sequence, i.e. intron removed, 1684 bp) SEQ ID NO:2 - T. reesei MNS1-HDEL aa (527 aa) SEQ ID NO:3 - C1gls2a-HDEL nt (coding sequence, i.e., intron removed, 2961 bp) SEQ ID NO: 4-C1 GLS2A-HDEL aa (986 aa) SEQ ID NO:5 - C1kre2-huGNT1 nt (coding sequence, i.e., intron removed, 1434 bp) SEQ ID NO: 6: C1 KRE2-huGNT1 aa (477 aa) SEQ ID NO: 7-C1 rft1 nt (coding sequence, i.e., intron removed, 1839 bp) SEQ ID NO:8-C1 RFT1 aa (612 aa) SEQ ID NO:9 - C1kre2-boGNT1 nt (coding sequence, i.e. intron removed, 1440 bp) SEQ ID NO:10-C1 KRE2-boGNT1 aa (479 aa) SEQ ID NO:11-LmSTT3 nt (2574 bp) SEQ ID NO: 12: LmSTT3 aa (857 aa) SEQ ID NO:13 - Sckre2-huGalT1 nt (1266 bp) SEQ ID NO:14 - ScKRE2-huGalT1 aa (421 aa) SEQ ID NO:15 - Rat GNT2 nt (1329 bp) SEQ ID NO:16 - Rat GNT2 aa (442 aa) SEQ ID NO:17 - pMYT1288 (14648 bp) SEQ ID NO:18-pMYT0721 (13062 bp) SEQ ID NO:19 - pMYT1451 (15270 bp) SEQ ID NO:20 - pMYT1452 (15281 bp) SEQ ID NO:21 - pMYT1453 (18195 bp) SEQ ID NO:22 - rDNA of Thermothelomyces heterothallica C1 SEQ ID NO:23-pMYT1974 (20189 bp) SEQ ID NO:24-pMYT1963 (18731 bp) SEQ ID NO:25-pMYT1964 (19105 bp)
[0204] The foregoing description of specific embodiments fully reveals the general nature of the invention, so that others, by applying current knowledge, can easily modify and / or adapt to various uses of such specific embodiments without undue experimentation and without departing from the general concept, and therefore such adaptations and modifications should be understood within the meaning and range of equivalents of the disclosed embodiments, and are intended to be so. It should be understood that the expressions or terms used herein are for the purpose of description and not for the purpose of limitation. The means, materials, and steps for carrying out the various disclosed chemical structures and functions can take various alternative forms without departing from the invention.
Claims
1. 1. A Thermothelomyces heterothallica genetically modified to produce glycoproteins having mammalian N-glycans, wherein the genetic modification comprises: (i) a deletion or disruption of the alg3 gene such that Th. heterothallica does not produce a functional alpha-1,3-mannosyltransferase; (ii) expression of ER-targeted mannosidase 1 (alpha-1,2-mannosidase), and (iii) Expression of ER-targeted glucosidase 2 alpha-subunit, including Thermothelomyces heterothallica.
2. 2. The Th. heterothallica of claim 1, wherein the mannosidase 1 is Trichoderma reesei mannosidase.
3. 2. The Th. heterothallica of claim 1, wherein the glucosidase 2 alpha-subunit is a Th. heterothallica glucosidase 2 alpha-subunit.
4. 2. The Th. heterothallica of claim 1, wherein the genetic modification further comprises expression of heterologous GlcNAc transferase 1 (GNT1) and GlcNAc transferase 2 (GNT2).
5. 5. The Th. heterothallica of claim 4, wherein the heterologous GNT1 and GNT2 are of animal origin.
6. 6. The Th. heterothallica according to claim 5, wherein the animal-derived GNT1 is human GNT1.
7. 7. The Th. heterothallica of claim 6, wherein the animal-derived GNT1 is human GNT1 further comprising a Th. heterothallica Golgi localization signal.
8. 6. The Th. heterothallica according to claim 5, wherein the animal-derived GNT1 is bovine GNT1.
9. 9. The Th. heterothallica of claim 8, wherein the animal-derived GNT1 is bovine GNT1 containing a Th. heterothallica Golgi localization signal.
10. 6. The Th. heterothallica according to claim 5, wherein the animal-derived GNT2 is rat GNT2.
11. 6. The Th. heterothallica of claim 5, wherein the animal-derived GNT1 is human or bovine GNT1 containing a Golgi localization signal from the Th. heterothallica protein KRE2, and the animal-derived GNT2 is rat GNT2.
12. 6. The Th. heterothallica of claim 5, wherein the animal-derived GNT1 is bovine GNT1 containing a Golgi localization signal from the Th. heterothallica protein KRE2, and the animal-derived GNT2 is rat GNT2.
13. 2. The Th. heterothallica of claim 1, wherein the genetic modification further comprises overexpression of an endogenous flippase or expression of a heterologous flippase.
14. 2. The Th. heterothallica of claim 1, wherein the genetic modification further comprises expression of the STT3 subunit of a heterologous oligosaccharyltransferase.
15. 2. The Th. heterothallica of claim 1, wherein the genetic modification further comprises expression of a heterologous galactosyltransferase.
16. 2. The Th. heterothallica of claim 1, wherein the Th. heterothallica is Th. heterothallica C1, and the C1 is a strain that has been modified to delete one or more genes encoding endogenous proteases.
17. 17. The Th. heterothallica of any one of claims 1 to 16, further genetically modified to express a heterologous mammalian glycoprotein.
18. 1. A method for generating Th. heterothallica that produces glycoproteins having mammalian N-glycans, comprising: (a) deleting or disrupting the alg3 gene of the Th. heterothallica such that the Th. heterothallica does not produce a functional alpha-1,3-mannosyltransferase; (b) introducing into said Th. heterothallica exogenous polynucleotides encoding ER-targeted mannosidase 1 (alpha-1,2-mannosidase) and ER-targeted glucosidase 2 alpha-subunit.
19. 1. A method for producing glycoproteins having mammalian N-glycans, comprising: (a) providing a genetically modified Th. heterothallica of claim 1; (b) culturing Th. heterothallica in a medium under conditions suitable for expressing the glycoprotein; (c) recovering the glycoprotein.
20. 2. A recombinant glycoprotein produced by the genetically modified Th. heterothallica of claim 1, wherein the glycoprotein is GlcNAc 2 Man 3 GlcNAc 2 (G0) A recombinant glycoprotein containing glycans.