Composition
By introducing oligosaccharide transferases derived from Trypanosoma brevicornu in mammalian cells, the problem of low N-glycan site occupancy efficiency was solved, improving the yield, stability, and functionality of recombinant proteins and achieving batch-to-batch uniform expression.
Patent Information
- Application Number
- CN202480040739.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-22
- Filing Date
- 2024-06-21
- Publication Date
- 2026-01-16
AI Technical Summary
In mammalian cells, the low occupancy rate of N-glycan sites affects the yield, function, and stability of recombinant proteins, leading to batch-to-batch inconsistencies in drug production.
Oligosaccharide transferase (OST) derived from Trypanosoma brevicornuate was introduced into mammalian cells, and through genetic engineering, the occupancy efficiency of N-glycan sites was improved, thereby enhancing the glycosylation profile of recombinant proteins.
It significantly improved the yield, stability, and activity of recombinant proteins, ensuring batch-to-batch uniformity and functional expression of recombinant proteins.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to the use of oligosaccharyltransferase (OST) enzymes obtained from Trypanosoma brucei, T. brucei, to improve the expression of functional and stable recombinant proteins in mammalian expression systems, which includes creating new engineered mammalian cell lines and genetic constructs comprising multiple nucleotides encoding OST enzymes obtained from T. brucei. BACKGROUND
[0002] The addition of glycan moieties to the surface of proteins, known as glycosylation, is one of the most common post-translational modifications (PTMs) and is a natural process that is important for the function and stability of many proteins. This complex process begins in the endoplasmic reticulum, where glycan moieties are added to newly synthesized proteins. This in turn affects the structure of the protein and the biological activity of the protein. Indeed, the inherent structural variations of glycan moieties can modulate the properties of the protein. Therefore, the production of therapeutic glycoproteins in industry has great interest. The production of proteins with added glycan moieties has the potential to improve the yield and stability of therapeutic proteins and also allows for the modulation of activity and new drug discovery.
[0003] Protein glycosylation is a naturally occurring process and in the human body, approximately 50% of proteins are glycosylated, hence the term glycoproteins. In the pharmaceutical industry, two-thirds of regulatory approved therapeutic proteins are glycoproteins (Delobel, Mass Spectrometry of Glycoproteins, pp 1-21, 2021) and they all have at least one glycan moiety (which is a specific sugar molecule) on their surface, but some glycoproteins have tens of glycan moieties attached. These glycan moieties are fundamentally important for the functional performance of the protein - they affect how active the protein is, how long it is stable and how well the protein is folded.
[0004] There are two types of glycosylation: N-glycosylation and O-glycosylation. Over 90% of glycoproteins are modified by N-glycosylation (Helenius, Mol Biol Cell, 5(3); 253-65, 1994). In mammals, N-glycosylation is the covalent attachment of Glc3MAN9GlcNAc2 to an asparagine (Asn or N) residue on the polypeptide chain, which is located within the Asn-X-Ser / Thr motif, where X can be any amino acid except proline (Reilly et al., Nature Reviews Nephrology, 15(346-366), 2019). This can occur co-translationally or post-translationally within the endoplasmic reticulum (ER) of the cell (Canada et al., Cell, 136(2): 272-283, 2010) (Kleizen & Braakman, Curr Opin Cell Biol, 16(4):343-9, 2004).
[0005] The addition of N-glycans to proteins is catalyzed by a multi-subunit enzyme called oligosaccharyltransferase (OST or also referred to as OTase). Mammalian OST is a multi-subunit membrane protein, comprising two distinct catalytic subunits (STT3A and STT3B), and at least six other non-catalytic subunits whose function is poorly understood (Mohanty et al., Biomolecules, 10(4):624, 2020), (Pfeffer et al., Nature Communications, 5 (3072), 2014). These two catalytic subunits have different substrate specificities (Cherepanova & Gilmore, Scientific Reports, 6(20946), 2016), and the entire OST catalyzes the transfer of a preassembled high-mannose oligosaccharide as a whole to the asparagine residue (Kheller & Gilmore, Glycobiology, 16(4):47-62, 2005).
[0006] However, in general, OSTs that perform glycosylation in eukaryotes are inefficient, with about 35% of N-glycan sequences unoccupied (Petrescu et al., Glycobiology, 14(2): 103-14, 2004). This represents a major limitation in the industry, where mammalian cells are commonly used for the bioproduction of glycosylated proteins, including therapeutic proteins. Indeed, N-glycan site occupancy is one of the key quality attributes analyzed by the Food and Drug Administration (FDA) for regulatory approval of therapeutic proteins. This is because N-glycosylation plays an important role in many aspects of protein production, including protein yield, protein function, proper protein folding, and controlling batch-to-batch variation.
[0007] Accordingly, there is a need in the art to improve the production of recombinantly produced proteins in mammalian cell lines. In particular, there is a need to enhance N-glycan site occupancy of recombinant proteins in order to substantially improve the yield, activity, stability, and reproducibility of recombinant proteins in the pharmaceutical production and discovery industry. SUMMARY
[0008] The inventors of the present application have surprisingly found that specific OST enzymes derived from the Trypanosoma brucei parasite can be engineered into mammalian bioproduction cells in order to substantially improve the N-glycan profile of recombinantly produced proteins in mammalian cells. This has the potential to improve the yield, stability, and activity of recombinantly produced proteins.
[0009] In a first aspect of the application, there is provided a mammalian cell comprising at least one nucleic acid sequence encoding at least one oligosaccharyltransferase (OST) protein or functional fragment thereof, wherein the at least one OST protein is a Trypanosoma spp. OST protein, preferably wherein the at least one OST protein is a Trypanosoma brucei OST protein.
[0010] In a second aspect of the application, there is provided a mammalian cell comprising at least one nucleic acid sequence encoding at least one oligosaccharyltransferase (OST) protein or functional fragment thereof, wherein the at least one nucleic acid sequence encoding the at least one OST protein or functional fragment thereof comprises a sequence according to any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22, or a sequence having at least 70% sequence identity thereto. In preferred embodiments, the mammalian cell can comprise a combination of nucleic acids, such as the combinations set out in the DETAILED DESCRIPTION.
[0011] In a third aspect of the application, there is provided a mammalian cell comprising at least one nucleic acid sequence encoding at least one oligosaccharyltransferase (OST) protein or functional fragment thereof, wherein the at least one OST protein comprises an amino acid sequence according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, and / or 20 or a sequence having at least 70% sequence identity thereto. In preferred embodiments, the mammalian cell can comprise a combination of nucleic acids, such as the combinations listed in the detailed description.
[0012] In a fourth aspect of the application, there is provided an isolated nucleic acid molecule comprising a sequence according to any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22 or a sequence having at least 70% identity thereto, preferably wherein the isolated nucleic acid molecule comprises a sequence according to SEQ ID NO: 3, 5, 7, 9, 11, and / or 13.
[0013] In a fifth aspect of the application, there is provided a nucleic acid vector comprising: (i) at least one nucleic acid sequence according to any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22 or a sequence having at least 70% identity thereto; or (ii) at least one isolated nucleic acid molecule according to the fourth aspect of the application, preferably wherein the nucleic acid vector is selected from the group comprising, but not limited to, a plasmid-based expression vector, a bacterial artificial chromosome (BAC) vector, and a viral vector such as an adenoviral vector, an adeno-associated vector (AAV), a retroviral vector, or a lentiviral vector. In preferred embodiments, the vector can comprise a combination of nucleic acids, such as the combinations listed in the detailed description.
[0014] In a sixth aspect of the application, there is provided a method of modifying the glycosylation profile of a recombinant protein, the method comprising contacting the recombinant protein with (i) a mammalian cell according to the first aspect of the application; or (ii) at least one OST protein or functional fragment thereof, wherein the at least one OST protein is a Trypanosoma sp. OST protein, preferably wherein the at least one OST protein is a T. brucei OST protein, more preferably wherein the at least one OST protein comprises an amino acid sequence according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, and / or 20 or a sequence having at least 70% sequence identity thereto; or producing the recombinant protein in a mammalian cell according to the first aspect of the application, preferably wherein the mammalian cell is engineered to express the recombinant protein. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1AAmino acid sequence alignment of OST proteins from T. brucei (TbSTT3 A, B, and C) showing residues 367-408, including the predicted active region 371-408. Positively charged residues are indicated in grey shading, negatively charged residues are indicated in bold and underlined, and neutral charge residues are indicated in bold.
[0016] Figure 1B AlphaFold predicted structure of TbSTT3 A, B, C, with the predicted active region highlighted in dark grey. Key amino acid residues are shown and labeled in black.
[0017] Figure 2 Table showing the three OST enzymes isolated from T. brucei, along with their respective gene identification numbers.
[0018] Figure 3 Table showing confirmation of expression of T. brucei enzymes in mammalian cells, verified via mass spectrometry.
[0019] Figure 4 Demonstration of increased hormone production in the presence of kinetoplastid OST. EPO hormone production was detected 72 h after transient transfection of cells. “Mock” is mock transfected cells (control).
[0020] Figure 5 Demonstration of increased N-glycans in the presence of T. brucei OST. Data shows the results of N-glycan site occupancy in the presence of OST (enzyme A) and mock transfected cells (control). The percentage of N-glycosylated peptides was determined by mass spectrometry.
[0021] Figure 6 Demonstration that transient expression of T. brucei OST enzymes in Expi293 cells (derived from HEK-293 cells) does not affect the overall viability of the cell population. DETAILED DESCRIPTION
[0022] To facilitate an understanding of the present application, a number of terms are defined below. Additional definitions are set forth throughout the detailed description.
[0023] As used herein, the term "mammalian cell" refers to a eukaryotic cell derived from or isolated from a mammalian tissue. In some embodiments, the eukaryotic cell can be modified to proliferate indefinitely, generating an immortalized cell line. In other embodiments, the eukaryotic cell can not be immortalized. In the context of the present invention, the mammalian cell is used as an expression system to express a recombinant protein. Thus, any mammalian cell that fulfills this function is within the scope of the present invention. For example, the mammalian cell can be a Chinese hamster ovary (CHO) cell, a baby hamster kidney (BHK21) cell, a murine myeloma cell (NS0 and Sp2 / 0), or a human embryonic kidney (HEK) cell. Preferably, the mammalian cell of the present invention is a Chinese hamster ovary (CHO) cell and / or an immortalized human embryonic kidney (HEK) cell, such as a HEK293 cell.
[0024] As used herein, the term "exogenous" refers to a nucleic acid sequence (i.e., a heterologous nucleic acid), also referred to as a DNA sequence, that originates from outside the mammalian cell and is subsequently transformed into the mammalian cell. As will be appreciated by those skilled in the art, the nucleic acid molecules of the present invention are exogenous to the mammalian cells of the present invention, as they are derived from the T. brucei OST-encoding nucleic acid. The terms "nucleic acid sequence" and "gene" are used interchangeably herein, although it will be appreciated that "gene" is generally used to refer to a nucleic acid sequence or combination of nucleic acid sequences that is transcribed and translated into a single protein or multiple proteins that form a complex with a prescribed function. In the context of the present invention, the term "exogenous gene" refers to a gene that originates from a different species, in particular a parasite such as T. brucei, and is therefore considered a heterologous gene. The exogenous nucleic acid sequence of the present invention can have a sequence according to any one of SEQ ID NOs: 1 to 7.
[0025] As used herein, the terms "oligosaccharyltransferase", "OST", and "OTase" are used interchangeably and can refer to mono- and multi-subunit glycosyltransferases that catalyze the addition of N-glycans to proteins. In the context of the present invention, an OST derived from or obtained from a parasite is used.
[0026] As used herein, the terms "sequence identity" and "sequence homology" are interchangeable and refer to the number of identical residues within a defined length that goes into a given alignment. To calculate the % sequence identity of any one of the sequences disclosed herein, sequence comparison software can be used, for example, using the default settings on the BLAST software package (V2.10.1).
[0027] As used herein, the term "recombinant protein" refers to a protein that has been encoded by a particular nucleic acid sequence, i.e., a gene that has been cloned in an expression system that supports gene expression and translation of messenger RNA (mRNA) (recombinant DNA). The gene introduced into the expression system can be a heterologous gene (or foreign gene), i.e., the gene is derived or originated from a cell type that is derived from a different organism than the recipient expression system. Methods for heterologous expression of recombinant proteins are well known to those skilled in the art. The expression system of the present application is a mammalian expression system. Preferably, the mammalian expression system is a mammalian cell, such as a CHO or HEK cell.
[0028] As used herein in reference to the methods of the present application, the term "contacting" is intended to encompass any manner of bringing the recombinant protein and the at least one OST protein into sufficient contact to allow glycosylation of the recombinant protein by the at least one OST protein. This can involve a chaperone as desired. It is contemplated that the recombinant protein can be contacted with the at least one OST protein within a mammalian cell or in a cell-free system, such as a cell-free expression system. The recombinant protein and the at least one OST protein can be co-expressed in a mammalian cell or cell-free system, or they can be expressed at different times. For example, the mammalian cell can constitutively express the at least one OST protein, and then recombinant cell expression is performed by the mammalian cell. The at least one OST protein can be expressed and secreted from the mammalian cell or extracted from the mammalian cell, and then contacted with the recombinant protein. For example, the mammalian cell expressing the at least one OST protein can be lysed to allow contact with the recombinant protein. As will be appreciated by those skilled in the art, there are many ways in which the methods of the present application can be performed, but the core concept requires that the at least one OST protein be presented to the recombinant protein in a manner that allows glycosylation of the recombinant protein. Ideally, the glycosylated recombinant protein is maintained in a conformation that maintains its biological activity, as measured in vitro or in vivo by ligand affinity assays, activity assays, stability assays, half-life assays, pharmacokinetic assays, dose studies, cellular uptake assays, and other assays that assess potency, efficacy, and quality. Again, the skilled artisan will know how to do this.
[0029] As used herein, the term "glycosylation" refers to the process of attaching oligosaccharides (glycan moieties) to target molecules such as proteins and lipids. The skilled artisan will readily recognize that this process is critical for both the functional performance and stability of proteins. Oligosaccharides are carbohydrates composed of chains or polymers of monosaccharides (single sugar molecules). Glycosylation is a form of co-translational and post-translational modification of proteins in which carbohydrates are conjugated or covalently attached to proteins. In biology, the process of glycosylation is an enzyme-catalyzed reaction and can involve enzymes such as oligosaccharyltransferases (OSTs) and glycosyltransferases. Within mammalian cells, glycosylation primarily occurs within the rough endoplasmic reticulum, cytoplasm, and nucleus. There are two major types of glycosylation: N-glycosylation (also referred to as "N-linked glycosylation") and O-glycosylation (also referred to as "O-linked glycosylation"), the former being the most prevalent. There is also C-glycosylation (also referred to as "C-linked glycosylation"), which occurs when a glycan is attached to a carbon atom on a tryptophan side chain. In N-glycosylation, glycans are attached to the nitrogen atom of an amino acid (e.g., asparagine and arginine). In O-glycosylation, glycans are attached to the oxygen atom of a hydroxyl group on an amino acid (e.g., serine, threonine, and tyrosine).
[0030] As used herein, the term "active region" refers to a region in at least one OST protein that is proposed or predicted to interact with a target molecule such as a polypeptide. The interaction between the OST and the target molecule can involve glycosylation and / or binding. For example, the interaction can involve OST-mediated glycosylation of a recombinant polypeptide or protein. As used herein, the term "active region" is also used interchangeably with "active site," "recognition site," "consensus," "conserved region," and "binding site."
[0031] Accordingly, in the context of the present application, the addition of "N-glycans" to a protein of interest is of particular interest. The addition of N-glycans refers to the addition of any of three general types of N-glycans: oligomannose, complex, and hybrid, each of which contains the common core Man3GlcNAc2, with any number of branches and / or elongations.
[0032] As used herein, the terms "N-glycan occupancy," "N-glycan site occupancy," "N-glycosylation occupancy," and "N-glycosylation site occupancy" are used interchangeably and refer to the number of potential N-glycan (or N-glycosylation) sites on a protein that are occupied, i.e., the number of sequons that are modified. This is also referred to as the "macroheterogeneity" of N-glycosylation. For example, 50% N-glycan occupancy would reflect a protein in which half of the potential N-glycan sites of the protein are occupied. 100% N-glycan occupancy would reflect a protein in which all of the potential N-glycan sites of the protein are occupied. The present invention is particularly directed to N-glycan occupancy of NXS / T sequons (i.e., the number of NXS / T consensus sites that are occupied by a glycan).
[0033] As used herein, the term "N-glycosylation profile" refers to both the sites of N-glycosylation as well as the types of N-glycans that are present at the sites on the surface of a fully folded protein. Thus, the N-glycosylation profile reflects both the extent of N-glycosylation site occupancy on a fully folded protein as well as the types of individual N-glycans. As will be appreciated by one of skill in the art, in determining the N-glycosylation profile, the extent of N-glycosylation site occupancy as well as the types of N-glycans present will reflect the profile of all proteins in a population of a given protein (i.e., those proteins present in a sample). In particular, in the context of the present invention, the term "N-glycosylation profile" refers to the overall profile of the extent of N-glycosylation site occupancy as well as the types of N-glycans present in a given population of recombinant proteins. Preferably, in the context of the present invention, the improved glycosylation profile consists of an increase in the extent of N-glycosylation site occupancy on a recombinant protein that has a consistent N-glycan portion that is reproducible between production batches of the recombinant protein. Mass spectrometry can be utilized in order to determine the % N-glycosylation of a given protein. Preferably, the OST proteins encompassed by the present invention will increase the % N-glycosylation of a given protein (e.g., a recombinant protein) to at least 90% site occupancy. For example, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% N-glycan site occupancy.
[0034] As used herein, the term "expression vector" or "expression construct" refers to a recombinantly or synthetically generated nucleic acid construct having a defined series of nucleic acid elements that permit the transcription of a particular nucleic acid sequence in a host cell. An expression vector can be a plasmid, a bacterial artificial chromosome, a virus, or a portion of a segment of nucleic acid sequences. Typically, an expression vector includes a nucleic acid sequence linked to a promoter sequence. An expression vector can encode at least one of any of the nucleic acid sequences of SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22.
[0035] To produce the T. brucei OST in mammalian cells, the above expression vectors containing the nucleic acid sequences are used to stably or transiently transfect the mammalian cells. Stable transfection refers to a method of integrating the T. brucei OST gene into the genome of the mammalian cell host that allows for stable expression of the OST gene under the control of a constitutive or inducible promoter. Stable transfection allows for integration into the host genome, thus the gene is replicated and the T. brucei OST gene is expressed long term. Transient transfection refers to any method of introducing the T. brucei OST gene and expressing it under the control of a constitutive or inducible promoter. Transient transfection does not allow for integration into the host genome, thus the gene is not inherited during cell division, thus expression of the T. brucei OST gene is sustained for a limited period of time.
[0036] The use of the alternative (e.g.,“or”) should be understood to mean either one, but not both, of the alternatives. As used herein, the indefinite articles“a” or“an” should be understood to refer to“one or more” of any recited or enumerated component.
[0037] As used herein,“about” means within an acceptable error range of the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, “about” can mean within 1 or more than 1 standard deviation, per the practice in the art. Alternatively, “about” can mean ranges approximately 20% above and below a given value. When a particular value is provided in the application and claims, unless otherwise stated the meaning of“about” should be assumed to be within an acceptable error range of the particular value.
[0038] It is contemplated that the expression systems and methods thereof disclosed herein will have numerous benefits compared to expression systems utilizing native mammalian OSTs. For example, the recombinant products produced in the expression systems defined herein are expected to have improved stability and functionality, enhanced effectiveness, higher yields, and higher batch-to-batch uniformity, all of which are key advantageous features in the bioproduction industry.
[0039] In a first aspect, the present application provides a mammalian cell comprising at least one nucleic acid sequence encoding at least one oligosaccharyltransferase (OST) protein or functional fragment thereof, wherein the at least one OST protein is a Trypanosoma sp. OST protein, preferably wherein the at least one OST protein is a T. brucei OST protein.
[0040] In a second aspect, the present application provides a mammalian cell comprising at least one nucleic acid sequence encoding at least one oligosaccharyltransferase (OST) protein or a functional fragment thereof, wherein the at least one nucleic acid sequence encoding the at least one OST protein comprises a sequence according to any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22, or a sequence having at least 70% identity thereto.
[0041] Accordingly, the present application utilizes genetic engineering to introduce a highly efficient OST enzyme from a unicellular parasitic organism into a mammalian cell. Thus, the inventors believe that improved N-glycan occupancy and sialyation of recombinant proteins can be achieved. The nucleic acid sequence according to the first aspect can have a sequence according to any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22. The sequence encodes a protein from T. brucei, a known causative agent of human African trypanosomiasis. There are three subspecies that cause different types of trypanosomiasis: T. brucei brucei, T. brucei gambiense, and T. brucei rhodiense.
[0042] In any aspect or embodiment of the present application disclosed herein, it is believed that the cell or vector can comprise any combination of one, two, three, or more of the nucleic acid sequences disclosed herein, particularly those combinations set out in the detailed description.
[0043] SEQ ID NO: 1 corresponds to a consensus DNA sequence encoding a fragment of a TbSTT3A protein.
[0044] SEQ ID NO: 21 corresponds to a consensus DNA sequence encoding a fragment of a TbSTT3B protein.
[0045] SEQ ID NO: 22 corresponds to a consensus DNA sequence encoding a fragment of a TbSTT3C protein.
[0046] The highly conserved consensus sequences of SEQ ID NOs: 1, 21, and 22 can contain residues that are critical to the activity of the OST protein. Without being bound by theory, it is believed that SEQ ID NOs: 1, 21, and 22 can encode an active region of a trypanosomal OST protein. This region defined by SEQ ID NOs: 1, 21, and 22 can be referred to herein as an active region, active site, recognition site, binding site, consensus site, or highly conserved site. It is believed that Figure 1AThe underlined residues can be particularly important. Thus, the present application also includes cells and vectors comprising any sequence that is at least 70% identical to any of SEQ ID NOs: 1, 21, and / or 22 as listed herein, also including lysine-369, arginine-369, and / or lysine-381, as numbered according to their position in SEQ ID NOs: 8, 10, and 12.
[0047] SEQ ID NOs: 2, 4, and 6 correspond to DNA sequences encoding the active region sequences of the TbSTT3A, TbSTT3B, and TbSTT3C proteins, respectively.
[0048] SEQ ID NOs: 3, 5, and 7 correspond to DNA sequences encoding the active region sequences of the TbSTT3A, TbSTT3B, and TbSTT3C proteins, respectively, optimized for expression in CHO mammalian cells.
[0049] SEQ ID NOs: 8, 10, and 12 correspond to DNA sequences encoding the full-length sequences of the TbSTT3A, TbSTT3B, and TbSTT3C proteins, respectively.
[0050] SEQ ID NOs: 9, 11, and 7 correspond to DNA sequences encoding the full-length sequences of the TbSTT3A, TbSTT3B, and TbSTT3C proteins, respectively, optimized for expression in CHO mammalian cells. While the SEQ ID NOs: 1-13 and 21-22 disclosed herein represent DNA nucleic acid sequences, the skilled artisan will appreciate that the corresponding RNA nucleic acid sequences will also be compatible with the disclosure of the present application, such that thymine residues can be replaced with uracil residues.
[0051] The active region of the Tb. brucei OST enzyme is predicted to be within the region spanning amino acid residue numbers 371-408 of TbSTT3A, B, and C. The inventors of the present application have identified potential key residues within this highly conserved region and around it that can affect peptide specificity. These residues include lysine-369 or arginine-369 or lysine-381 and arginine-397 or histidine-397. In particular, the key amino acid residues for TbSTT3A are lysine-369, lysine-381, and arginine-397. The key amino acid residues for TbSTT3B are arginine-369, lysine-381, and histidine-397. The key amino acid residues for TbSTT3C are arginine-369, lysine-381, and arginine-397.
[0052] According to a third aspect of the application, there is provided a mammalian cell comprising at least one nucleic acid sequence encoding at least one oligosaccharyltransferase (OST) protein or a functional fragment thereof, wherein the at least one OST protein comprises an amino acid sequence according to any one of SEQ ID NO: 14, 15, 16, 17, 18, 19 and / or 20 or a sequence having at least 70% sequence identity thereto. The OST protein can comprise a full-length protein (i.e. comprising SEQ ID NO: 17, 18 and / or 19), or the OST protein can comprise a highly conserved region (i.e. comprising SEQ ID NO: 14, 15, 16 and / or 20). As the skilled person will readily appreciate, the amino acid sequences as defined in the third aspect of the application can be derived from the nucleic acid sequences as defined in the first and second aspects of the application. However, there are many possible permutations of the nucleic acid codons that can give rise to these amino acid sequences, and this is also appreciated by the skilled person. Thus, there can be other nucleic acid sequences from which the amino acid sequences of the third aspect of the application can be derivable.
[0053] SEQ ID NO: 14, 15 and 16 correspond to the amino acid sequences corresponding to the active regions of the TbSTT3A, TbSTT3B and TbSTT3C proteins, respectively.
[0054] SEQ ID NO: 17, 18 and 19 correspond to the amino acid sequences corresponding to the full-length TbSTT3A, TbSTT3B and TbSTT3C proteins, respectively.
[0055] SEQ ID NO: 20 corresponds to the amino acids conserved in the active regions of TbSTT3A, TbSTT3B and TbSTT3C. SEQ ID NO: 1, 21 and 22 are consensus DNA sequences encoding the amino acids according to SEQ ID NO: 20.
[0056] Jinnelov et al. (2017) previously demonstrated that arginine-397 and histidine-397 in the TbOSTa (TbSTT3A) and TbOSTb (TbSTT3B) respectively are extremely important for determining peptide specificity between the TbOST enzymes. Notably, both of these residues are positively charged. The inventors of the present invention have observed that residues in the L. major and T. cruzi OST enzyme active regions that can be involved in peptide specificity are neutral or negatively charged amino acids (i.e., these key residues are not positively charged). Surprisingly, however, in contrast, the key residues in the TbOST enzyme active region that can be involved in peptide specificity are positively charged. This pattern occurs again twice at other residues mentioned above. Thus, the inventors hypothesize that the unique glycan specificity of the TbOST protein can be influenced by, or be a result of, the presence of positively charged residues at these positions.
[0057] The present application also provides a nucleic acid comprising or having a sequence that is at least 70% identical to any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22. For example, the nucleic acid sequence can be at least 75% identical to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22; at least 80% identical to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22; at least 85% identical to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22; at least 90% identical to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22; at least 91% identical to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22; at least 92% identical to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22; at least 93% identical to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22; at least 94% identical to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22; at least 95% identical to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22; at least 96% identical to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22; at least 97% identical to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22; at least 98% identical to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22; or at least 99% identical to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22.
[0058] The mammalian cell can comprise at least one exogenous nucleic acid comprising or having a sequence according to SEQ ID NO: 1. In preferred embodiments, the exogenous nucleic acid has the sequence of SEQ ID NO: 1, or a sequence having at least 70% identity to said sequence. For example, the exogenous nucleic acid can comprise a sequence having at least 75% identity to SEQ ID NO: 1; at least 80% identity to SEQ ID NO: 1; at least 85% identity to SEQ ID NO: 1; at least 90% identity to SEQ ID NO: 1; at least 91% identity to SEQ ID NO: 1; at least 92% identity to SEQ ID NO: 1; at least 93% identity to SEQ ID NO: 1; at least 94% identity to SEQ ID NO: 1; at least 95% identity to SEQ ID NO: 1; at least 96% identity to SEQ ID NO: 1; at least 97% identity to SEQ ID NO: 1; at least 98% identity to SEQ ID NO: 1; or at least 99% identity to SEQ ID NO: 1. According to the sequence as set forth in SEQ ID NO: 1, the amino acid residues at positions 369, 381 and 397 will be positively charged amino acids, for example amino acids selected from lysine, arginine and histidine, or artificial analogs thereof.
[0059] The mammalian cell can comprise at least one exogenous nucleic acid comprising or having a sequence according to SEQ ID NO: 21. In preferred embodiments, the exogenous nucleic acid has the sequence of SEQ ID NO: 21, or a sequence having at least 70% identity to said sequence. For example, the exogenous nucleic acid can comprise a sequence having at least 75% identity to SEQ ID NO: 21; at least 80% identity to SEQ ID NO: 21; at least 85% identity to SEQ ID NO: 21; at least 90% identity to SEQ ID NO: 21; at least 91% identity to SEQ ID NO: 21; at least 92% identity to SEQ ID NO: 21; at least 93% identity to SEQ ID NO: 21; at least 94% identity to SEQ ID NO: 21; at least 95% identity to SEQ ID NO: 21; at least 96% identity to SEQ ID NO: 21; at least 97% identity to SEQ ID NO: 21; at least 98% identity to SEQ ID NO: 21; or at least 99% identity to SEQ ID NO: 21. According to the sequence as set forth in SEQ ID NO: 21, the amino acid residues at positions 369, 381, and 397 will be positively charged amino acids, for example, amino acids selected from the group consisting of lysine, arginine, and histidine, or artificial analogs thereof.
[0060] The mammalian cell can comprise at least one exogenous nucleic acid comprising or having a sequence according to SEQ ID NO: 22. In preferred embodiments, the exogenous nucleic acid has the sequence of SEQ ID NO: 22, or a sequence having at least 70% identity to said sequence. For example, the exogenous nucleic acid can comprise a sequence having at least 75% identity to SEQ ID NO: 22; at least 80% identity to SEQ ID NO: 22; at least 85% identity to SEQ ID NO: 22; at least 90% identity to SEQ ID NO: 22; at least 91% identity to SEQ ID NO: 22; at least 92% identity to SEQ ID NO: 22; at least 93% identity to SEQ ID NO: 22; at least 94% identity to SEQ ID NO: 22; at least 95% identity to SEQ ID NO: 22; at least 96% identity to SEQ ID NO: 22; at least 97% identity to SEQ ID NO: 22; at least 98% identity to SEQ ID NO: 22; or at least 99% identity to SEQ ID NO: 22. According to the sequence as set forth in SEQ ID NO: 22, the amino acid residues at positions 369, 381 and 397 will be positively charged amino acids, for example, amino acids selected from the group consisting of lysine, arginine and histidine, or artificial analogs thereof.
[0061] In any embodiment of the application, SEQ ID NOs: 1, 21 and 22 are interchangeable.
[0062] The mammalian cell can comprise one or more exogenous nucleic acids encoding different OSTs, wherein the exogenous nucleic acids are selected from the group consisting of any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21 and / or 22. The mammalian cell can comprise a single exogenous nucleic acid sequence, two exogenous nucleic acid sequences or three exogenous nucleic acid sequences. For example, the mammalian cell can comprise an exogenous nucleic acid sequence according to any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21 and / or 22; or any combination thereof. When the mammalian cell comprises two or more exogenous nucleic acids, the exogenous nucleic acids can be expressed separately (i.e., on separate vectors) or together (i.e., on the same vector). The one or more nucleic acid molecules can be expressed as a fusion protein, or they can be expressed as separate proteins. They can be expressed as a fusion protein comprising a linker, for example, a cleavable linker. As will be appreciated by those skilled in the art, the cleavable linker can be a linker that is cleavable by a protease, or it can be a self-cleaving peptide, for example, a viral self-cleaving peptide, for example, P2A.
[0063] In preferred embodiments, the mammalian cell can comprise an exogenous nucleic acid sequence according to SEQ ID NO: 2, 4, or 6; 2 and 4; 2 and 6; 4 and 6; and / or 2, 4, and 6. In another preferred embodiment, the mammalian cell can comprise an exogenous nucleic acid sequence according to SEQ ID NO: 8, 10, or 12; 8 and 10; 8 and 12; 10 and 12; and / or 8, 10, and 12. In another preferred embodiment, the mammalian cell can comprise an exogenous nucleic acid sequence according to SEQ ID NO: 3, 5, or 7; 3 and 5; 3 and 7; 5 and 7; and / or 3, 5, and 7. In another preferred embodiment, the mammalian cell can comprise an exogenous nucleic acid sequence according to SEQ ID NO: 9, 11, or 13; 9 and 11; 9 and 13; 11 and 13; and / or 9, 11, and 13. As the skilled artisan will appreciate, any combination of the DNA sequences disclosed herein can be used for the purposes of the present application, including combinations of optimized nucleic acid sequences with non-optimized nucleic acid sequences. Furthermore, the skilled artisan will readily appreciate that any of the nucleic acid sequences encoding the active regions disclosed herein (optimized or non-optimized) can be combined with any of the nucleic acid sequences encoding the full-length proteins disclosed herein (optimized or non-optimized).
[0064] In preferred embodiments, the exogenous nucleic acid has the sequence of SEQ ID NO: 2, or a sequence that is at least 70% identical to said sequence. For example, the exogenous nucleic acid can comprise a sequence that is at least 75% identical to SEQ ID NO: 2; at least 80% identical to SEQ ID NO: 2; at least 85% identical to SEQ ID NO: 2; at least 90% identical to SEQ ID NO: 2; at least 91% identical to SEQ ID NO: 2; at least 92% identical to SEQ ID NO: 2; at least 93% identical to SEQ ID NO: 2; at least 94% identical to SEQ ID NO: 2; at least 95% identical to SEQ ID NO: 2; at least 96% identical to SEQ ID NO: 2; at least 97% identical to SEQ ID NO: 2; at least 98% identical to SEQ ID NO: 2; or at least 99% identical to SEQ ID NO: 2.
[0065] In preferred embodiments, the exogenous nucleic acid has the sequence of SEQ ID NO: 3, or a sequence that is at least 70% identical to said sequence. For example, the exogenous nucleic acid can comprise a sequence that is at least 75% identical to SEQ ID NO: 3; at least 80% identical to SEQ ID NO: 3; at least 85% identical to SEQ ID NO: 3; at least 90% identical to SEQ ID NO: 3; at least 91% identical to SEQ ID NO: 3; at least 93% identical to SEQ ID NO: 3; at least 93% identical to SEQ ID NO: 3; at least 94% identical to SEQ ID NO: 3; at least 95% identical to SEQ ID NO: 3; at least 96% identical to SEQ ID NO: 3; at least 97% identical to SEQ ID NO: 3; at least 98% identical to SEQ ID NO: 3; or at least 99% identical to SEQ ID NO: 3.
[0066] In preferred embodiments, the exogenous nucleic acid has the sequence of SEQ ID NO: 4, or a sequence that is at least 70% identical to said sequence. For example, the exogenous nucleic acid can comprise a sequence that is at least 75% identical to SEQ ID NO: 4; at least 80% identical to SEQ ID NO: 4; at least 85% identical to SEQ ID NO: 4; at least 90% identical to SEQ ID NO: 4; at least 91% identical to SEQ ID NO: 4; at least 94% identical to SEQ ID NO: 4; at least 94% identical to SEQ ID NO: 4; at least 94% identical to SEQ ID NO: 4; at least 95% identical to SEQ ID NO: 4; at least 96% identical to SEQ ID NO: 4; at least 97% identical to SEQ ID NO: 4; at least 98% identical to SEQ ID NO: 4; or at least 99% identical to SEQ ID NO: 4.
[0067] In preferred embodiments, the exogenous nucleic acid has the sequence of SEQ ID NO: 5, or a sequence that is at least 70% identical to said sequence. For example, the exogenous nucleic acid can comprise a sequence that is at least 75% identical to SEQ ID NO: 5; at least 80% identical to SEQ ID NO: 5; at least 85% identical to SEQ ID NO: 5; at least 90% identical to SEQ ID NO: 5; at least 91% identical to SEQ ID NO: 5; at least 92% identical to SEQ ID NO: 5; at least 93% identical to SEQ ID NO: 5; at least 94% identical to SEQ ID NO: 5; at least 95% identical to SEQ ID NO: 5; at least 96% identical to SEQ ID NO: 5; at least 97% identical to SEQ ID NO: 5; at least 98% identical to SEQ ID NO: 5; or at least 99% identical to SEQ ID NO: 5. This particular exogenous nucleic acid is believed to be particularly successful due to its broad substrate specificity and expression in improving N-glycan occupancy.
[0068] In preferred embodiments, the exogenous nucleic acid has the sequence of SEQ ID NO: 6, or a sequence that is at least 70% identical to said sequence. For example, the exogenous nucleic acid can comprise a sequence that is at least 75% identical to SEQ ID NO: 6; at least 80% identical to SEQ ID NO: 6; at least 85% identical to SEQ ID NO: 6; at least 90% identical to SEQ ID NO: 6; at least 91% identical to SEQ ID NO: 6; at least 92% identical to SEQ ID NO: 6; at least 93% identical to SEQ ID NO: 6; at least 94% identical to SEQ ID NO: 6; at least 95% identical to SEQ ID NO: 6; at least 96% identical to SEQ ID NO: 6; at least 97% identical to SEQ ID NO: 6; at least 98% identical to SEQ ID NO: 6; or at least 99% identical to SEQ ID NO: 6.
[0069] In preferred embodiments, the exogenous nucleic acid has the sequence of SEQ ID NO: 7, or a sequence having at least 70% identity to said sequence. For example, the exogenous nucleic acid can comprise a sequence having at least 75% identity to SEQ ID NO: 7; at least 80% identity to SEQ ID NO: 7; at least 85% identity to SEQ ID NO: 7; at least 90% identity to SEQ ID NO: 7; at least 91% identity to SEQ ID NO: 7; at least 92% identity to SEQ ID NO: 7; at least 93% identity to SEQ ID NO: 7; at least 94% identity to SEQ ID NO: 7; at least 95% identity to SEQ ID NO: 7; at least 96% identity to SEQ ID NO: 7; at least 97% identity to SEQ ID NO: 7; at least 98% identity to SEQ ID NO: 7; or at least 99% identity to SEQ ID NO: 7.
[0070] In preferred embodiments, the exogenous nucleic acid has the sequence of SEQ ID NO: 8, or a sequence having at least 70% identity to said sequence. For example, the exogenous nucleic acid can comprise a sequence having at least 75% identity to SEQ ID NO: 8; at least 80% identity to SEQ ID NO: 8; at least 85% identity to SEQ ID NO: 8; at least 90% identity to SEQ ID NO: 8; at least 91% identity to SEQ ID NO: 8; at least 92% identity to SEQ ID NO: 8; at least 93% identity to SEQ ID NO: 8; at least 94% identity to SEQ ID NO: 8; at least 95% identity to SEQ ID NO: 8; at least 96% identity to SEQ ID NO: 8; at least 97% identity to SEQ ID NO: 8; at least 98% identity to SEQ ID NO: 8; or at least 99% identity to SEQ ID NO: 8.
[0071] In preferred embodiments, the exogenous nucleic acid has the sequence of SEQ ID NO: 9, or a sequence that is at least 70% identical to said sequence. For example, the exogenous nucleic acid can comprise a sequence that is at least 75% identical to SEQ ID NO: 9; at least 80% identical to SEQ ID NO: 9; at least 85% identical to SEQ ID NO: 9; at least 90% identical to SEQ ID NO: 9; at least 91% identical to SEQ ID NO: 9; at least 92% identical to SEQ ID NO: 9; at least 93% identical to SEQ ID NO: 9; at least 94% identical to SEQ ID NO: 9; at least 95% identical to SEQ ID NO: 9; at least 96% identical to SEQ ID NO: 9; at least 97% identical to SEQ ID NO: 9; at least 98% identical to SEQ ID NO: 9; or at least 99% identical to SEQ ID NO: 9.
[0072] In preferred embodiments, the exogenous nucleic acid has the sequence of SEQ ID NO: 10, or a sequence that is at least 70% identical to said sequence. For example, the exogenous nucleic acid can comprise a sequence that is at least 75% identical to SEQ ID NO: 10; at least 80% identical to SEQ ID NO: 10; at least 85% identical to SEQ ID NO: 10; at least 90% identical to SEQ ID NO: 10; at least 91% identical to SEQ ID NO: 10; at least 92% identical to SEQ ID NO: 10; at least 93% identical to SEQ ID NO: 10; at least 94% identical to SEQ ID NO: 10; at least 95% identical to SEQ ID NO: 10; at least 96% identical to SEQ ID NO: 10; at least 97% identical to SEQ ID NO: 10; at least 98% identical to SEQ ID NO: 10; or at least 99% identical to SEQ ID NO: 10.
[0073] In preferred embodiments, the exogenous nucleic acid has the sequence of SEQ ID NO: 11, or a sequence having at least 70% identity to that sequence. For example, the exogenous nucleic acid can comprise a sequence having at least 75% identity to SEQ ID NO: 11; at least 80% identity to SEQ ID NO: 11; at least 85% identity to SEQ ID NO: 11; at least 90% identity to SEQ ID NO: 11; at least 91% identity to SEQ ID NO: 11; at least 92% identity to SEQ ID NO: 11; at least 93% identity to SEQ ID NO: 11; at least 94% identity to SEQ ID NO: 11; at least 95% identity to SEQ ID NO: 11; at least 96% identity to SEQ ID NO: 11; at least 97% identity to SEQ ID NO: 11; at least 98% identity to SEQ ID NO: 11; or at least 99% identity to SEQ ID NO: 11.
[0074] In preferred embodiments, the exogenous nucleic acid has the sequence of SEQ ID NO: 12, or a sequence having at least 70% identity to that sequence. For example, the exogenous nucleic acid can comprise a sequence having at least 75% identity to SEQ ID NO: 12; at least 80% identity to SEQ ID NO: 12; at least 85% identity to SEQ ID NO: 12; at least 90% identity to SEQ ID NO: 12; at least 91% identity to SEQ ID NO: 12; at least 92% identity to SEQ ID NO: 12; at least 93% identity to SEQ ID NO: 12; at least 94% identity to SEQ ID NO: 12; at least 95% identity to SEQ ID NO: 12; at least 96% identity to SEQ ID NO: 12; at least 97% identity to SEQ ID NO: 12; at least 98% identity to SEQ ID NO: 12; or at least 99% identity to SEQ ID NO: 12.
[0075] In preferred embodiments, the exogenous nucleic acid has the sequence of SEQ ID NO: 13, or a sequence having at least 70% identity to said sequence. For example, the exogenous nucleic acid can comprise a sequence having at least 75% identity to SEQ ID NO: 13; at least 80% identity to SEQ ID NO: 13; at least 85% identity to SEQ ID NO: 13; at least 90% identity to SEQ ID NO: 13; at least 91% identity to SEQ ID NO: 13; at least 92% identity to SEQ ID NO: 13; at least 93% identity to SEQ ID NO: 13; at least 94% identity to SEQ ID NO: 13; at least 95% identity to SEQ ID NO: 13; at least 96% identity to SEQ ID NO: 13; at least 97% identity to SEQ ID NO: 13; at least 98% identity to SEQ ID NO: 13; or at least 99% identity to SEQ ID NO: 13.
[0076] The three OST paralogs expressed in T. brucei have different substrate specificities. Analysis of the parasite protein glycosylation revealed that TbSTT3A selectively transfers bi-antennary Man5GlcNAc2glycans, while both TbSTT3B and TbSTT3C transfer tri-antennary Man9GlcNAc2glycans. It has been shown that the substrate specificity arises from the presence of an ALG12-dependent c-branch of the conventional tri-antennary Man9GlcNAc2glycan, which is required for TbSTT3B and TbSTT3C but not TbSTT3A. However, studies have also reported that both TbSTT3A and TbSTT3B can transfer Man5GlcNAc2glycans as well as Man7GlcNAc2glycans to T. brucei VSG proteins. The unique functionality of TbSTT3B with respect to its promiscuous substrate specificity is relevant. Most OTases display different preferences for oligosaccharides, limiting the efficiency of N-glycosylation. Indeed, it is conceivable that the use of an adaptable OTase that can transfer N-glycans to almost any amino acid sequence without bias would increase N-glycosylation in mammalian cells. Thus, it is desirable that in alternative embodiments of the present application, T. brucei OST enzymes are exploited to provide a combination of their unique specificities and functionalities.
[0077] Thus, in alternative preferred embodiments, the mammalian cell can comprise two exogenous nucleic acids which together at least encode the active regions of TbSTT3B and TbSTT3C. For example, the exogenous nucleic acids can comprise a sequence according to SEQ ID NO: 4 or a sequence having at least 70% identity to said sequence, and a further sequence according to SEQ ID NO: 6 or a sequence having at least 70% identity to said sequence. For example, the exogenous nucleic acids can comprise a sequence having at least 75% identity to SEQ ID NO: 4 and / or 6; at least 80% identity to SEQ ID NO: 4 and / or 6; at least 85% identity to SEQ ID NO: 4 and / or 6; at least 90% identity to SEQ ID NO: 4 and / or 6; at least 91 % identity to SEQ ID NO: 4 and / or 6; at least 92% identity to SEQ ID NO: 4 and / or 6; at least 93% identity to SEQ ID NO: 4 and / or 6; at least 94% identity to SEQ ID NO: 4 and / or 6; at least 95% identity to SEQ ID NO: 4 and / or 6; at least 96% identity to SEQ ID NO: 4 and / or 6; at least 97% identity to SEQ ID NO: 4 and / or 6; at least 98% identity to SEQ ID NO: 4 and / or 6; or at least 99% identity to SEQ ID NO: 4 and / or 6.
[0078] In alternative preferred embodiments, the mammalian cell can comprise two exogenous nucleic acids encoding full-length TbSTT3B and TbSTT3C. For example, the exogenous nucleic acids can comprise a sequence according to SEQ ID NO: 10 or a sequence with at least 70% identity to said sequence, and a further sequence according to SEQ ID NO: 12 or a sequence with at least 70% identity to said sequence. For example, the exogenous nucleic acids can comprise a sequence with at least 75% identity to SEQ ID NO: 10 and / or 12; at least 80% identity to SEQ ID NO: 10 and / or 12; at least 85% identity to SEQ ID NO: 10 and / or 12; at least 90% identity to SEQ ID NO: 10 and / or 12; at least 91% identity to SEQ ID NO: 10 and / or 12; at least 92% identity to SEQ ID NO: 10 and / or 12; at least 93% identity to SEQ ID NO: 10 and / or 12; at least 94% identity to SEQ ID NO: 10 and / or 12; at least 95% identity to SEQ ID NO: 10 and / or 12; at least 96% identity to SEQ ID NO: 10 and / or 12; at least 97% identity to SEQ ID NO: 10 and / or 12; at least 98% identity to SEQ ID NO: 10 and / or 12; or at least 99% identity to SEQ ID NO: 10 and / or 12.
[0079] In alternative preferred embodiments, the mammalian cell can comprise at least two exogenous nucleic acids encoding active regions of TbSTT3A and TbSTT3B. For example, the exogenous nucleic acids can comprise a sequence according to SEQ ID NO: 2 or a sequence with at least 70% identity to said sequence, and a further sequence according to SEQ ID NO: 4 or a sequence with at least 70% identity to said sequence. For example, the exogenous nucleic acids can comprise a sequence with at least 75% identity to SEQ ID NO: 2 and / or 4; at least 80% identity to SEQ ID NO: 2 and / or 4; at least 85% identity to SEQ ID NO: 2 and / or 4; at least 90% identity to SEQ ID NO: 2 and / or 4; at least 91% identity to SEQ ID NO: 2 and / or 4; at least 92% identity to SEQ ID NO: 2 and / or 4; at least 93% identity to SEQ ID NO: 2 and / or 4; at least 94% identity to SEQ ID NO: 2 and / or 4; at least 95% identity to SEQ ID NO: 2 and / or 4; at least 96% identity to SEQ ID NO: 2 and / or 4; at least 97% identity to SEQ ID NO: 2 and / or 4; at least 98% identity to SEQ ID NO: 2 and / or 4; or at least 99% identity to SEQ ID NO: 2 and / or 4.
[0080] In alternative preferred embodiments, the mammalian cell can comprise two exogenous nucleic acids encoding full-length TbSTT3A and TbSTT3B. For example, the mammalian cell can comprise two exogenous nucleic acids. The exogenous nucleic acids can comprise a sequence according to SEQ ID NO: 8 or a sequence with at least 70% identity to said sequence, and a further sequence according to SEQ ID NO: 10 or a sequence with at least 70% identity to said sequence. For example, the exogenous nucleic acids can comprise a sequence with at least 75% identity to SEQ ID NO: 8 and / or 10; at least 80% identity to SEQ ID NO: 8 and / or 10; at least 85% identity to SEQ ID NO: 8 and / or 10; at least 90% identity to SEQ ID NO: 8 and / or 10; at least 91% identity to SEQ ID NO: 8 and / or 10; at least 92% identity to SEQ ID NO: 8 and / or 10; at least 93% identity to SEQ ID NO: 8 and / or 10; at least 94% identity to SEQ ID NO: 8 and / or 10; at least 95% identity to SEQ ID NO: 8 and / or 10; at least 96% identity to SEQ ID NO: 8 and / or 10; at least 97% identity to SEQ ID NO: 8 and / or 10; at least 98% identity to SEQ ID NO: 8 and / or 10; or at least 99% identity to SEQ ID NO: 8 and / or 10.
[0081] In alternative preferred embodiments, the mammalian cell can comprise at least two exogenous nucleic acids encoding active regions of TbSTT3A and TbSTT3C. For example, the mammalian cell can comprise two exogenous nucleic acids. The exogenous nucleic acids can comprise a sequence according to SEQ ID NO: 2 or a sequence with at least 70% identity to said sequence, and a further sequence according to SEQ ID NO: 6 or a sequence with at least 70% identity to said sequence. For example, the exogenous nucleic acids can comprise a sequence with at least 75% identity to SEQ ID NO: 2 and / or 6; at least 80% identity to SEQ ID NO: 2 and / or 6; at least 85% identity to SEQ ID NO: 2 and / or 6; at least 90% identity to SEQ ID NO: 2 and / or 6; at least 91% identity to SEQ ID NO: 2 and / or 6; at least 92% identity to SEQ ID NO: 2 and / or 6; at least 93% identity to SEQ ID NO: 2 and / or 6; at least 94% identity to SEQ ID NO: 2 and / or 6; at least 95% identity to SEQ ID NO: 2 and / or 6; at least 96% identity to SEQ ID NO: 2 and / or 6; at least 97% identity to SEQ ID NO: 2 and / or 6; at least 98% identity to SEQ ID NO: 2 and / or 6; or at least 99% identity to SEQ ID NO: 2 and / or 6.
[0082] In alternative preferred embodiments, the mammalian cell can comprise two exogenous nucleic acids encoding full-length TbSTT3A and TbSTT3C. For example, the exogenous nucleic acids can comprise a sequence according to SEQ ID NO: 8 or a sequence with at least 70% identity to said sequence, and a further sequence according to SEQ ID NO: 12 or a sequence with at least 70% identity to said sequence. For example, the exogenous nucleic acids can comprise a sequence with at least 75% identity to SEQ ID NO: 8 and / or 12; at least 80% identity to SEQ ID NO: 8 and / or 12; at least 85% identity to SEQ ID NO: 8 and / or 12; at least 90% identity to SEQ ID NO: 8 and / or 12; at least 91% identity to SEQ ID NO: 8 and / or 12; at least 92% identity to SEQ ID NO: 8 and / or 12; at least 93% identity to SEQ ID NO: 8 and / or 12; at least 94% identity to SEQ ID NO: 8 and / or 12; at least 95% identity to SEQ ID NO: 8 and / or 12; at least 96% identity to SEQ ID NO: 8 and / or 12; at least 97% identity to SEQ ID NO: 8 and / or 12; at least 98% identity to SEQ ID NO: 8 and / or 12; or at least 99% identity to SEQ ID NO: 8 and / or 12.
[0083] In more preferred embodiments, the mammalian cell can comprise at least three exogenous nucleic acids encoding active regions of TbSTT3A, TbSTT3B, and TbSTT3C. For example, the exogenous nucleic acids can comprise a sequence according to SEQ ID NO: 2 or a sequence with at least 70% identity to said sequence, a further sequence according to SEQ ID NO: 4 or a sequence with at least 70% identity to said sequence, and a sequence according to SEQ ID NO: 6 or a sequence with at least 70% identity to said sequence. For example, the exogenous nucleic acids can comprise a sequence with at least 75% identity to SEQ ID NO: 2, 4, and / or 6; at least 80% identity to SEQ ID NO: 2, 4, and / or 6; at least 85% identity to SEQ ID NO: 2, 4, and / or 6; at least 90% identity to SEQ ID NO: 2, 4, and / or 6; at least 91% identity to SEQ ID NO: 2, 4, and / or 6; at least 92% identity to SEQ ID NO: 2, 4, and / or 6; at least 93% identity to SEQ ID NO: 2, 4, and / or 6; at least 94% identity to SEQ ID NO: 2, 4, and / or 6; at least 95% identity to SEQ ID NO: 2, 4, and / or 6; at least 96% identity to SEQ ID NO: 2, 4, and / or 6; at least 97% identity to SEQ ID NO: 2, 4, and / or 6; at least 98% identity to SEQ ID NO: 2, 4, and / or 6; or at least 99% identity to SEQ ID NO: 2, 4, and / or 6.
[0084] In alternative preferred embodiments, the mammalian cell can comprise three exogenous nucleic acids encoding full-length TbSTT3A, TbSTT3B, and TbSTT3C. For example, the exogenous nucleic acids can comprise a sequence according to SEQ ID NO: 8 or a sequence with at least 70% identity to said sequence, a further sequence according to SEQ ID NO: 10 or a sequence with at least 70% identity to said sequence, and a further sequence according to SEQ ID NO: 12 or a sequence with at least 70% identity to said sequence. For example, the exogenous nucleic acids can comprise a sequence with at least 75% identity to SEQ ID NO: 8, 10, and / or 12; at least 80% identity to SEQ ID NO: 8, 10, and / or 12; at least 85% identity to SEQ ID NO: 8, 10, and / or 12; at least 90% identity to SEQ ID NO: 8, 10, and / or 12; at least 91% identity to SEQ ID NO: 8, 10, and / or 12; at least 92% identity to SEQ ID NO: 8, 10, and / or 12; at least 93% identity to SEQ ID NO: 8, 10, and / or 12; at least 94% identity to SEQ ID NO: 8, 10, and / or 12; at least 95% identity to SEQ ID NO: 8, 10, and / or 12; at least 96% identity to SEQ ID NO: 8, 10, and / or 12; at least 97% identity to SEQ ID NO: 8, 10, and / or 12; at least 98% identity to SEQ ID NO: 8, 10, and / or 12; or at least 99% identity to SEQ ID NO: 8, 10, and / or 12.
[0085] A skilled artisan will readily appreciate that any of the above combinations can also be used with optimized versions of the sequences disclosed herein. For example, a mammalian cell can comprise two optimized exogenous nucleic acids that encode at least the active region of any one of TbSTT3A, TbSTT3B, and / or TbSTT3C, and thus comprise or have a sequence according to SEQ ID NO: 5 and / or 7; 3 and / or 5; 3 and / or 7; or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto. Alternatively, a mammalian cell can comprise three optimized exogenous nucleic acids that encode the active region of any one of TbSTT3A, TbSTT3B, and / or TbSTT3C, and thus comprise or have a sequence according to SEQ ID NO: 3, 5, and / or 7; or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto. Similarly, a mammalian cell can comprise two optimized exogenous nucleic acids that encode the full length of any one of TbSTT3A, TbSTT3B, and / or TbSTT3C, and thus comprise or have a sequence according to SEQ ID NO: 11 and / or 13; 9 and / or 11; 9 and / or 13; or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto. Alternatively, a mammalian cell can comprise three optimized exogenous nucleic acids that encode the full length of any one of TbSTT3A, TbSTT3B, and / or TbSTT3C, and thus comprise or have a sequence according to SEQ ID NO: 9, 11, and / or 13; or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto. A skilled artisan will also readily appreciate that any combination of the DNA sequences disclosed herein can be used for the purposes of the present application, including combinations of optimized nucleic acid sequences with non-optimized nucleic acid sequences.
[0086] Furthermore, the skilled artisan will readily appreciate that any of the nucleic acid sequences (optimized or non-optimized) encoding the active regions disclosed herein can be combined with any of the nucleic acid sequences (optimized or non-optimized) encoding the full-length proteins disclosed herein. For example, in the presence of two exogenous nucleic acids, the nucleic acid sequence can comprise any of the sequences according to SEQ ID NOs: 2, 3, 8, and / or 9, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto, and further comprise any of the sequences according to SEQ ID NOs: 4, 5, 10, and / or 11, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto.
[0087] Furthermore, the nucleic acid sequence can comprise any of the sequences according to SEQ ID NOs: 2, 3, 8, and / or 9, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto, and further comprise any of the sequences according to SEQ ID NOs: 6, 7, 12, and / or 13, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto.
[0088] Furthermore, the nucleic acid sequence can comprise any of the sequences according to SEQ ID NOs: 4, 5, 10, and / or 11, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto, and further comprise any of the sequences according to SEQ ID NOs: 6, 7, 12, and / or 13, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto.
[0089] In the presence of three exogenous nucleic acids, the nucleic acid sequence can comprise a sequence according to any one of the sequences of SEQ ID NO: 2, 3, 8 and / or 9, or having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto, and further comprising any one of the sequences according to SEQ ID NO: 4, 5, 10 and / or 11, or having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto, and further comprising any one of the sequences according to SEQ ID NO: 6, 7, 12 and / or 13, or having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto.
[0090] In a fourth aspect, the present application provides an isolated nucleic acid molecule comprising a sequence according to any one of SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21 and / or 22, or a sequence having at least 70% identity thereto, preferably wherein the isolated nucleic acid molecule comprises a sequence according to SEQ ID NO: 3, 5, 7, 9, 11 and / or 13. Thus, in preferred embodiments, the isolated nucleic acid molecule comprises a sequence according to SEQ ID NO: 3, 5, 7, 9, 11 and / or 13 (i.e. a sequence optimised for expression in a mammalian cell, such as a CHO cell). It is considered that the nucleic acids of the fourth aspect of the present application can be successfully transfected into a mammalian cell. In another embodiment, the nucleic acid comprises a sequence according to SEQ ID NO: 2, 4, 6, 8, 10 and / or 12 (i.e. a sequence which has not been optimised for expression in a mammalian cell, such as a CHO cell). Again, the skilled person will readily appreciate that any combination of the nucleic acid sequences disclosed herein can be used for the purposes of the present application, including combinations of optimised nucleic acid sequences with non-optimised nucleic acid sequences. Furthermore, the skilled person will readily appreciate that any of the nucleic acid sequences encoding the active regions disclosed herein (optimised or non-optimised) can be combined with any of the nucleic acid sequences encoding the full-length proteins disclosed herein (optimised or non-optimised).
[0091] Accordingly, in a fifth aspect, the present application provides a vector comprising at least one nucleic acid sequence according to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21 and / or 22, or comprising an isolated nucleic acid molecule of the fourth aspect of the application. A vector deemed suitable can be designed to encode at least one nucleic acid comprising or having a sequence according to SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21 and / or 22, for transfection into a mammalian cell. In one embodiment, the vector comprises at least one nucleic acid sequence according to SEQ ID NO: 3, 5, 7, 9, 11 and / or 13 (i.e. sequences optimised for expression in a mammalian cell, such as a CHO cell). It is considered that a vector of the fourth aspect of the application can be successfully transfected into a mammalian cell. In another embodiment, the vector comprises at least one nucleic acid sequence according to SEQ ID NO: 2, 4, 6, 8, 10 and / or 12 (i.e. sequences that have not been optimised for expression in a mammalian cell, such as a CHO cell). Again, the skilled person will readily appreciate that any combination of the DNA sequences disclosed herein can be used for the purposes of the present application, including combinations of optimised nucleic acid sequences with non-optimised nucleic acid sequences.
[0092] Suitable vectors can include, but are not limited to, plasmid-based expression vectors, bacterial artificial chromosome (BAC) vectors or viral vectors (e.g. adenovirus, retrovirus, lentivirus). Methods for transfection into mammalian cells are routine in the art and readily understood by the skilled person. For example, suitable transfection methods can include, but are not limited to, electroporation, viral transfection, microinjection and chemical transfection (e.g. reagent-based transfection).
[0093] The mammalian cell of the present application can be any mammalian cell into which the OST from a parasite disclosed herein can be successfully transfected and expressed. For example, the mammalian cell can be a Chinese hamster ovary (CHO) cell, a baby hamster kidney (BHK21) cell, a murine myeloma cell (e.g. NS0 and Sp2 / 0) or a human embryonic kidney (e.g. HEK293) cell. In a preferred embodiment, the mammalian cell is a CHO cell or a HEK293 cell.
[0094] Both transient and stable mammalian cell lines can be used in the present application. Preferably, the mammalian cell to be used is a stable mammalian cell. Even more preferably, the mammalian cell to be used is a stable CHO cell. Stable mammalian cells have advantages over transiently transfected mammalian cells in that they allow the genetic modification to be passed on to the progeny of the cell, allowing the cell that has been genetically modified to maintain stable expression of the gene of interest over an extended period of time.
[0095] Disclosed herein are mammalian cells comprising at least one exogenous nucleic acid sequence encoding at least one OST from Trypanosoma brucei. The exogenous nucleic acid sequence can be expressed by the mammalian cell to provide a functional protein. Expression of the nucleic acid sequence in the mammalian cell can be achieved via inducible expression, constitutive expression, stable expression, and transient expression.
[0096] In one embodiment, the mammalian cell can express at least one Trypanosoma brucei OST enzyme as well as one or more recombinant proteins encoded by a further exogenous nucleic acid sequence. The recombinant protein can be a protein of interest to be glycosylated by the Trypanosoma brucei OST enzyme. As the skilled person will appreciate, the recombinant protein can be any protein or peptide comprising an amino acid residue suitable for N-glycosylation. For example, the recombinant protein can be a therapeutic protein or therapeutic peptide, a diagnostic protein or diagnostic peptide, a protein or peptide for research and development purposes, and / or a protein or peptide for supplementation of growth media that requires or benefits from N-glycosylation in a mammalian cell to improve its yield, activity, stability, and / or reproducibility.
[0097] The recombinant protein can be any protein that can be produced using the mammalian expression system disclosed herein. A key advantage of the present invention is that its application is not limited to a small number of recombinant proteins, but can be applied to any recombinant protein of interest. For example, the recombinant protein can be a therapeutic protein. Examples of therapeutic proteins include, but are not limited to, hormones, cytokines, antibodies, enzymes, complement proteins, coagulation factors, functional fragments thereof, or any combination thereof. The antibody can be a monoclonal or polyclonal antibody. The antibody can be derived from the following classes and subclasses: IgG including the IgG1, IgG2, IgG3, and IgG4 subclasses, IgA including the IgA1 and IgA2 subclasses, IgM, IgD, and IgE. The antibody can be a humanized monoclonal antibody. The antibody can be a bispecific antibody.
[0098] As used in relation to an OST protein in any embodiment of the present invention or in relation to a recombinant protein or any other protein mentioned herein, the term "functional fragment thereof" refers to a polypeptide which is derived from a longer polypeptide, e.g. a full-length polypeptide, and which has been truncated in the N-terminal region and / or C-terminal region to generate a fragment of the full-length polypeptide. To be a functional fragment, the fragment must retain at least 40%, preferably at least 50%, more preferably at least 60%, more preferably at least 70%, more preferably at least 80%, even more preferably at least 90%, most preferably at least 95%, even most preferably at least 100% of the activity of the full-length / mature polypeptide.
[0099] In preferred embodiments, the recombinant protein can be a protein selected from the group comprising erythropoietin (EPO), rituximab, butyrylcholinesterase (BuChE), Factor VIII, ENPP1-Fc, Olamkicept (sgp130-Fc), GM-CSF, FSH, eCG, alpha-1-antitrypsin, viral glycoproteins, e.g. SARS-CoV2 spike protein, or any combination thereof.
[0100] In alternative embodiments, the mammalian cell can express at least one T. brucei OST enzyme alone, without a recombinant protein. It is contemplated that such cell can be further modified to express any one or more recombinant proteins of interest that require N-glycosylation by the T. brucei OST enzyme. In addition, the mammalian cell can be further modified to become increasingly suitable for the introduction of further exogenous nucleic acid sequences encoding recombinant proteins.
[0101] In preferred embodiments, the mammalian cell expresses an OST protein to provide a functional / active OST protein. In particular, the mammalian cell can produce at least one OST protein comprising or having an amino acid sequence according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, and / or 20, or a sequence having at least 70% identity to said sequence.
[0102] The mammalian cell can produce at least one OST protein comprising or having an amino acid sequence according to any one of SEQ ID NOs: 14, 15, and / or 16, or a functional fragment thereof, or a sequence having at least 70% identity to said sequence or fragment. Alternatively, the mammalian cell can produce at least one OST protein comprising or having an amino acid sequence according to any one of SEQ ID NOs: 17, 18, and / or 19, or a functional fragment thereof, or a sequence having at least 70% identity to said sequence or fragment.
[0103] The mammalian cell can produce at least one OST protein comprising or having an amino acid sequence according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, and / or 20, or a functional fragment thereof, or any combination of such sequences or fragments. The mammalian cell can comprise a single OST protein, two OST proteins, or three OST proteins.
[0104] Thus, in one embodiment, the mammalian cell can comprise an OST protein. In preferred embodiments, the OST protein has the sequence of SEQ ID NO: 14, or a sequence that is at least 70% identical to said sequence. For example, the OST protein can have a sequence that is at least 75% identical to SEQ ID NO: 14; at least 80% identical to SEQ ID NO: 14; at least 85% identical to SEQ ID NO: 14; at least 90% identical to SEQ ID NO: 14; at least 91% identical to SEQ ID NO: 14; at least 92% identical to SEQ ID NO: 14; at least 93% identical to SEQ ID NO: 14; at least 94% identical to SEQ ID NO: 14; at least 95% identical to SEQ ID NO: 14; at least 96% identical to SEQ ID NO: 14; at least 97% identical to SEQ ID NO: 14; at least 98% identical to SEQ ID NO: 14; or at least 99% identical to SEQ ID NO: 14.
[0105] In preferred embodiments, the OST protein has the sequence of SEQ ID NO: 15, or a sequence that is at least 70% identical to said sequence. For example, the OST protein can have a sequence that is at least 75% identical to SEQ ID NO: 15; at least 80% identical to SEQ ID NO: 15; at least 85% identical to SEQ ID NO: 15; at least 90% identical to SEQ ID NO: 15; at least 91% identical to SEQ ID NO: 15; at least 92% identical to SEQ ID NO: 15; at least 93% identical to SEQ ID NO: 15; at least 94% identical to SEQ ID NO: 15; at least 95% identical to SEQ ID NO: 15; at least 96% identical to SEQ ID NO: 15; at least 97% identical to SEQ ID NO: 15; at least 98% identical to SEQ ID NO: 15; or at least 99% identical to SEQ ID NO: 15.
[0106] In preferred embodiments, the OST protein has the sequence of SEQ ID NO: 16, or a sequence that is at least 70% identical to said sequence. For example, the OST protein can have a sequence that is at least 75% identical to SEQ ID NO: 16; at least 80% identical to SEQ ID NO: 16; at least 85% identical to SEQ ID NO: 16; at least 90% identical to SEQ ID NO: 16; at least 91% identical to SEQ ID NO: 16; at least 92% identical to SEQ ID NO: 16; at least 93% identical to SEQ ID NO: 16; at least 94% identical to SEQ ID NO: 16; at least 95% identical to SEQ ID NO: 16; at least 96% identical to SEQ ID NO: 16; at least 97% identical to SEQ ID NO: 16; at least 98% identical to SEQ ID NO: 16; or at least 99% identical to SEQ ID NO: 16.
[0107] In preferred embodiments, the OST protein has the sequence of SEQ ID NO: 17, or a sequence that is at least 70% identical to said sequence. For example, the OST protein can have a sequence that is at least 75% identical to SEQ ID NO: 17; at least 80% identical to SEQ ID NO: 17; at least 85% identical to SEQ ID NO: 17; at least 90% identical to SEQ ID NO: 17; at least 91% identical to SEQ ID NO: 17; at least 92% identical to SEQ ID NO: 17; at least 93% identical to SEQ ID NO: 17; at least 94% identical to SEQ ID NO: 17; at least 95% identical to SEQ ID NO: 17; at least 96% identical to SEQ ID NO: 17; at least 97% identical to SEQ ID NO: 17; at least 98% identical to SEQ ID NO: 17; or at least 99% identical to SEQ ID NO: 17.
[0108] In preferred embodiments, the OST protein has the sequence of SEQ ID NO: 18, or a sequence that is at least 70% identical to that sequence. For example, the OST protein can have a sequence that is at least 75% identical to SEQ ID NO: 18; at least 80% identical to SEQ ID NO: 18; at least 85% identical to SEQ ID NO: 18; at least 90% identical to SEQ ID NO: 18; at least 91% identical to SEQ ID NO: 18; at least 92% identical to SEQ ID NO: 18; at least 93% identical to SEQ ID NO: 18; at least 94% identical to SEQ ID NO: 18; at least 95% identical to SEQ ID NO: 18; at least 96% identical to SEQ ID NO: 18; at least 97% identical to SEQ ID NO: 18; at least 98% identical to SEQ ID NO: 18; or at least 99% identical to SEQ ID NO: 18.
[0109] In preferred embodiments, the OST protein has the sequence of SEQ ID NO: 19, or a sequence that is at least 70% identical to that sequence. For example, the OST protein can have a sequence that is at least 75% identical to SEQ ID NO: 19; at least 80% identical to SEQ ID NO: 19; at least 85% identical to SEQ ID NO: 19; at least 90% identical to SEQ ID NO: 19; at least 91% identical to SEQ ID NO: 19; at least 92% identical to SEQ ID NO: 19; at least 93% identical to SEQ ID NO: 19; at least 94% identical to SEQ ID NO: 19; at least 95% identical to SEQ ID NO: 19; at least 96% identical to SEQ ID NO: 19; at least 97% identical to SEQ ID NO: 19; at least 98% identical to SEQ ID NO: 19; or at least 99% identical to SEQ ID NO: 19.
[0110] In preferred embodiments, the OST protein has the sequence of SEQ ID NO: 20, or a sequence that is at least 70% identical to said sequence. For example, the OST protein can have a sequence that is at least 75% identical to SEQ ID NO: 20; at least 80% identical to SEQ ID NO: 20; at least 85% identical to SEQ ID NO: 20; at least 90% identical to SEQ ID NO: 20; at least 91% identical to SEQ ID NO: 20; at least 92% identical to SEQ ID NO: 20; at least 93% identical to SEQ ID NO: 20; at least 94% identical to SEQ ID NO: 20; at least 95% identical to SEQ ID NO: 20; at least 96% identical to SEQ ID NO: 20; at least 97% identical to SEQ ID NO: 20; at least 98% identical to SEQ ID NO: 20; or at least 99% identical to SEQ ID NO: 20.
[0111] The mammalian cell can produce at least one OST protein comprising or having an amino acid sequence according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, and / or 20, or any combination thereof. The mammalian cell can comprise a single OST protein, two OST proteins, or three OST proteins.
[0112] Thus, in another embodiment, the mammalian cell can comprise or express two separate OST proteins. For example, the OST protein can have a sequence comprising the active regions of TbSTT3B and TbSTT3C, thus comprising or having a sequence according to SEQ ID NO: 15, or a sequence with at least 70% identity to said sequence, and a further sequence according to SEQ ID NO: 16, or a sequence with at least 70% identity to said sequence. For example, the OST protein can have a sequence with at least 75% identity to SEQ ID NO: 15 and / or 16; at least 80% identity to SEQ ID NO: 15 and / or 16; at least 85% identity to SEQ ID NO: 15 and / or 16; at least 90% identity to SEQ ID NO: 15 and / or 16; at least 91 % identity to SEQ ID NO: 15 and / or 16; at least 92% identity to SEQ ID NO: 15 and / or 16; at least 93% identity to SEQ ID NO: 15 and / or 16; at least 94% identity to SEQ ID NO: 15 and / or 16; at least 95% identity to SEQ ID NO: 15 and / or 16; at least 96% identity to SEQ ID NO: 15 and / or 16; at least 97% identity to SEQ ID NO: 15 and / or 16; at least 98% identity to SEQ ID NO: 15 and / or 16; or at least 99% identity to SEQ ID NO: 15 and / or 16.
[0113] In another embodiment, the OST protein can have a sequence comprising the sequence of full-length TbSTT3B and TbSTT3C, thus comprising or having the sequence according to SEQ ID NO: 18, or a sequence with at least 70% identity to said sequence, and the further sequence according to SEQ ID NO: 19, or a sequence with at least 70% identity to said sequence. For example, the OST protein can have a sequence with at least 75% identity to SEQ ID NO: 18 and / or 19; at least 80% identity to SEQ ID NO: 18 and / or 19; at least 85% identity to SEQ ID NO: 18 and / or 19; at least 90% identity to SEQ ID NO: 18 and / or 19; at least 91 % identity to SEQ ID NO: 18 and / or 19; at least 92% identity to SEQ ID NO: 18 and / or 19; at least 93% identity to SEQ ID NO: 18 and / or 19; at least 94% identity to SEQ ID NO: 18 and / or 19; at least 95% identity to SEQ ID NO: 18 and / or 19; at least 96% identity to SEQ ID NO: 18 and / or 19; at least 97% identity to SEQ ID NO: 18 and / or 19; at least 98% identity to SEQ ID NO: 18 and / or 19; or at least 99% identity to SEQ ID NO: 18 and / or 19.
[0114] In another embodiment, the OST protein can have a sequence comprising the active region of TbSTT3A and TbSTT3B, thus comprising or having a sequence according to SEQ ID NO: 14, or a sequence having at least 70% identity to said sequence, and a further sequence according to SEQ ID NO: 15, or a sequence having at least 70% identity to said sequence. For example, the OST protein can have a sequence having at least 75% identity to SEQ ID NO: 14 and / or 15; at least 80% identity to SEQ ID NO: 14 and / or 15; at least 85% identity to SEQ ID NO: 14 and / or 15; at least 90% identity to SEQ ID NO: 14 and / or 15; at least 91 % identity to SEQ ID NO: 14 and / or 15; at least 92% identity to SEQ ID NO: 14 and / or 15; at least 93% identity to SEQ ID NO: 14 and / or 15; at least 94% identity to SEQ ID NO: 14 and / or 15; at least 95% identity to SEQ ID NO: 14 and / or 15; at least 96% identity to SEQ ID NO: 14 and / or 15; at least 97% identity to SEQ ID NO: 14 and / or 15; at least 98% identity to SEQ ID NO: 14 and / or 15; or at least 99% identity to SEQ ID NO: 14 and / or 15.
[0115] In another embodiment, the OST protein can have a sequence comprising the full-length TbSTT3A and TbSTT3B, thus comprising or having the sequence according to SEQ ID NO: 17, or a sequence having at least 70% identity to said sequence, and the further sequence according to SEQ ID NO: 18, or a sequence having at least 70% identity to said sequence. For example, the OST protein can have a sequence having at least 75% identity to SEQ ID NO: 17 and / or 18; at least 80% identity to SEQ ID NO: 17 and / or 18; at least 85% identity to SEQ ID NO: 17 and / or 18; at least 90% identity to SEQ ID NO: 17 and / or 18; at least 91 % identity to SEQ ID NO: 17 and / or 18; at least 92% identity to SEQ ID NO: 17 and / or 18; at least 93% identity to SEQ ID NO: 17 and / or 18; at least 94% identity to SEQ ID NO: 17 and / or 18; at least 95% identity to SEQ ID NO: 17 and / or 18; at least 96% identity to SEQ ID NO: 17 and / or 18; at least 97% identity to SEQ ID NO: 17 and / or 18; at least 98% identity to SEQ ID NO: 17 and / or 18; or at least 99% identity to SEQ ID NO: 17 and / or 18.
[0116] In another embodiment, the OST protein can have a sequence comprising the active regions of TbSTT3A and TbSTT3C, thus comprising or having a sequence according to SEQ ID NO: 14, or a sequence having at least 70% identity to said sequence, and a further sequence according to SEQ ID NO: 16, or a sequence having at least 70% identity to said sequence. For example, the OST protein can have a sequence having at least 75% identity to SEQ ID NO: 14 and / or 16; at least 80% identity to SEQ ID NO: 14 and / or 16; at least 85% identity to SEQ ID NO: 14 and / or 16; at least 90% identity to SEQ ID NO: 14 and / or 16; at least 91 % identity to SEQ ID NO: 14 and / or 16; at least 92% identity to SEQ ID NO: 14 and / or 16; at least 93% identity to SEQ ID NO: 14 and / or 16; at least 94% identity to SEQ ID NO: 14 and / or 16; at least 95% identity to SEQ ID NO: 14 and / or 16; at least 96% identity to SEQ ID NO: 14 and / or 16; at least 97% identity to SEQ ID NO: 14 and / or 16; at least 98% identity to SEQ ID NO: 14 and / or 16; or at least 99% identity to SEQ ID NO: 14 and / or 16.
[0117] In another embodiment, the OST protein can have a sequence comprising the sequence of full-length TbSTT3A and TbSTT3C, thus comprising or having the sequence according to SEQ ID NO: 17, or a sequence with at least 70% identity to said sequence, and the further sequence according to SEQ ID NO: 19, or a sequence with at least 70% identity to said sequence. For example, the OST protein can have a sequence with at least 75% identity to SEQ ID NO: 17 and / or 19; at least 80% identity to SEQ ID NO: 17 and / or 19; at least 85% identity to SEQ ID NO: 17 and / or 19; at least 90% identity to SEQ ID NO: 17 and / or 19; at least 91 % identity to SEQ ID NO: 17 and / or 19; at least 92% identity to SEQ ID NO: 17 and / or 19; at least 93% identity to SEQ ID NO: 17 and / or 19; at least 94% identity to SEQ ID NO: 17 and / or 19; at least 95% identity to SEQ ID NO: 17 and / or 19; at least 96% identity to SEQ ID NO: 17 and / or 19; at least 97% identity to SEQ ID NO: 17 and / or 19; at least 98% identity to SEQ ID NO: 17 and / or 19; or at least 99% identity to SEQ ID NO: 17 and / or 19.
[0118] For the avoidance of any doubt, it is considered that when a cell comprises nucleic acids encoding more than one OST protein according to the present disclosure, these sequences are not necessarily on the same nucleic acid molecule or construct. Currently, the nucleic acids encoding the OST proteins are on separate nucleic acid molecules or genetic constructs, which are used for transient expression and stable expression of the OST proteins. However, the disclosure herein does not exclude this possibility of having nucleic acids encoding more than one OST protein on the same nucleic acid molecule, which also encompasses co-expression of multiple OST molecules as fusion proteins or proteins separated by self-cleaving peptide sequences, such as viral P2A sequences, or enzymatically cleavable sequences, such as sequences susceptible to protease-mediated degradation or multiple expression cassettes within the same vector.
[0119] The skilled person will readily understand that any of the amino acid sequences comprising the active region can be combined with any of the amino acid sequences corresponding to the full-length proteins.
[0120] The mammalian cell can produce an OST protein comprising or having the amino acid sequence according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, and / or 20, or any combination thereof. The mammalian cell can comprise a single OST protein, two OST proteins, or three OST proteins.
[0121] Thus, in one embodiment, the mammalian cell can comprise three separate OST proteins. For example, the OST protein can have a sequence comprising the proposed polypeptide binding site of TbSTT3A, TbSTT3B and TbSTT3C, thus comprising or having a sequence according to SEQ ID NO: 14 or a sequence having at least 70% identity to said sequence, and a further sequence according to SEQ ID NO: 15 or a sequence having at least 70% identity to said sequence, and a further sequence according to SEQ ID NO: 16 or a sequence having at least 70% identity to said sequence. For example, the OST protein can have a sequence which is at least 75% identical to SEQ ID NO: 14, 15 and / or 16; at least 80% identical to SEQ ID NO: 14, 15 and / or 16; at least 85% identical to SEQ ID NO: 14, 15 and / or 16; at least 90% identical to SEQ ID NO: 14, 15 and / or 16; at least 91% identical to SEQ ID NO: 14, 15 and / or 16; at least 92% identical to SEQ ID NO: 14, 15 and / or 16; at least 93% identical to SEQ ID NO: 14, 15 and / or 16; at least 94% identical to SEQ ID NO: 14, 15 and / or 16; at least 95% identical to SEQ ID NO: 14, 15 and / or 16; at least 96% identical to SEQ ID NO: 14, 15 and / or 16; at least 97% identical to SEQ ID NO: 14, 15 and / or 16; at least 98% identical to SEQ ID NO: 14, 15 and / or 16; or at least 99% identical to SEQ ID NO: 14, 15 and / or 16.
[0122] In another embodiment, the OST protein can have a sequence comprising the sequence of full-length TbSTT3A, TbSTT3B and TbSTT3C, thus comprising or having a sequence according to SEQ ID NO: 17 or a sequence having at least 70% identity to said sequence, and a further sequence according to SEQ ID NO: 18 or a sequence having at least 70% identity to said sequence, and a further sequence according to SEQ ID NO: 19 or a sequence having at least 70% identity to said sequence. For example, the OST protein can have a sequence having at least 75% identity to SEQ ID NO: 17, 18 and / or 19; at least 80% identity to SEQ ID NO: 17, 18 and / or 19; at least 85% identity to SEQ ID NO: 17, 18 and / or 19; at least 90% identity to SEQ ID NO: 17, 18 and / or 19; at least 91% identity to SEQ ID NO: 17, 18 and / or 19; at least 92% identity to SEQ ID NO: 17, 18 and / or 19; at least 93% identity to SEQ ID NO: 17, 18 and / or 19; at least 94% identity to SEQ ID NO: 17, 18 and / or 19; at least 95% identity to SEQ ID NO: 17, 18 and / or 19; at least 96% identity to SEQ ID NO: 17, 18 and / or 19; at least 97% identity to SEQ ID NO: 17, 18 and / or 19; at least 98% identity to SEQ ID NO: 17, 18 and / or 19; or at least 99% identity to SEQ ID NO: 17, 18 and / or 19.
[0123] The skilled person will readily appreciate that any of the amino acid sequences comprising the proposed polypeptide binding sites (highly conserved regions) can be combined with any of the amino acid sequences corresponding to the full-length proteins.
[0124] The native gene sequence from T. brucei encoding the OST enzyme has been codon-optimised for recombinant protein production in CHO cells. As the skilled person will readily appreciate, codon optimisation is a genetic engineering tool which is performed in order to improve gene expression and protein production. Codon optimisation seeks to address the issue of codon bias, which refers to the more frequent use of some synonymous codons over others within a region of a transcribed gene. The frequency of these synonymous codons varies within and between species. Codon optimisation changes the gene sequence to accommodate the codon bias of the host organism without changing the amino acid sequence in order to increase protein production. The GenScript Codon Optimizer Tool was used. Thus, while the present disclosure discloses codon-optimised sequences for producing recombinant proteins in CHO cells, the skilled person will appreciate that this also applies to producing recombinant proteins in other mammalian cells, including HEK cells (e.g. HEK293 cells).
[0125] In accordance with SEQ ID NO: 3, 5, 7, 9, 11 and / or 13 disclosed herein, the inventors of the present application have developed nucleic acid sequences encoding T. brucei OST proteins (in accordance with SEQ ID NO: 14, 15, 16, 17, 18 and / or 19) in which the codons have been optimised for production in mammalian cells.
[0126] In a sixth aspect, the present application provides a method of modifying the glycosylation profile of a recombinant protein, the method comprising contacting the recombinant protein with (i) a mammalian cell according to the first aspect of the present application; or (ii) with an OST protein or functional fragment thereof, wherein the OST protein is a Trypanosoma sp. OST protein, preferably wherein the OST protein is a T. brucei OST protein, more preferably wherein the OST protein comprises an amino acid sequence according to any one of SEQ ID NO: 14, 15, 16, 17, 18, 19 and / or 20; or producing the recombinant protein in a mammalian cell according to the first aspect of the present application, preferably wherein the mammalian cell is engineered to express the recombinant protein.
[0127] TbSTT3 proteins target certain proteins via recognition of a common region on the target by the active region of the TbSTT3. TbSTT3 proteins have polypeptide specificity for certain acceptor substrates, and the active region involved in the peptide acceptor specificity of TbSTT3s falls within a region spanning amino acid positions 371-408 of each of TbSTT3A, TbSTT3B, and TbSTT3C (Figure 1). In particular, arginine-397 in TbSTT3A and TbSTT3C and histidine-397 in TbSTT3B are known to interact with NX(S / T) adjacent amino acids and affect the interaction between the acceptor peptide and the enzyme surface (Jinnelov et al., 2017). These interactions can increase the efficiency of substrate recognition, and thus this region is important for the specific function of TbSTT3 proteins. Both residues arginine-397 / histidine-397 in TbSTT3s are positively charged amino acids. Potentially, without wishing to be bound by theory, the positively charged active region helps define TbSTT3 function by facilitating specific sequence preferences or polypeptide acceptor specificity. Thus, this unique property of TbSTT3 proteins can be particularly useful for a biotechnology platform to improve efficient N-glycosylation of recombinant glycoproteins in eukaryotic expression systems.
[0128] The three STT3 paralogs in T. brucei have different substrate specificities. For example, TbSTT3A has been found to selectively transfer bisantennary Man5GlcNAc2glycans, while both TbSTT3B and TbSTT3C transfer trisantennary Man9GlcNAc2glycans (Jinnelov et al., 2017). Of particular importance is the ability of TbSTT3 proteins to transfer Man9GlcNAc2moieties, which are the preferred substrates of mammalian OST proteins.
[0129] The present invention discloses engineering high efficiency enzymes from a unicellular parasitic organism (T. brucei) into mammalian cell lines to increase N-glycan occupancy of recombinant proteins. Thus, the methods disclosed herein provide methods in which the frequency of N-glycan occupancy on a recombinant protein can be altered. In preferred embodiments, the methods disclosed herein provide methods in which the frequency of N-glycan occupancy on a recombinant protein can be increased. However, sialyation and density of glycans on the surface of recombinant proteins can also be increased. In the context of the present invention, the term “increased” refers to a comparison between the number of potential N-glycosylation sites (i.e., N-glycan occupancy) on a recombinant protein produced in a mammalian cell line having a native / unmodified form of OST and the number of N-glycan occupancy on a recombinant protein produced in a mammalian cell line having an OST obtained from a parasite as disclosed herein.
[0130] As disclosed above, the recombinant protein in relation to the method of modifying the glycosylation profile of a recombinant protein can be any protein that can be produced using the mammalian expression system disclosed herein. For example, the recombinant protein can be a therapeutic protein. Examples of therapeutic proteins include, but are not limited to, hormones, cytokines, antibodies, enzymes, complement proteins, coagulation factors, functional fragments thereof, or any combination thereof. The antibody can be a monoclonal or polyclonal antibody. The antibody can be a humanized antibody. The antibody can be a bispecific antibody. The term “functional fragment” refers to a polypeptide that is derived from a longer polypeptide, such as a full-length polypeptide, and has been truncated in the N-terminal region and / or C-terminal region to generate a fragment of the full-length polypeptide. To be a functional fragment, the fragment must retain at least 40%, preferably at least 50%, more preferably at least 60%, more preferably at least 70%, more preferably at least 80%, even more preferably at least 90%, most preferably at least 95%, even most preferably at least 100% of the activity of the full-length / mature polypeptide.
[0131] In preferred embodiments, the recombinant protein can be a protein selected from the group comprising erythropoietin (EPO), rituximab, butyrylcholinesterase (BuChE), Factor VIII, ENPP1-Fc, Orlacizumab (sgp130-Fc), GM-CSF, FSH, eCG, alpha-1-antitrypsin, viral glycoproteins, e.g. SARS-CoV2 spike protein, or any combination thereof.
[0132] The application is further described with reference to the following non-limiting examples:
[0133] Example
[0134] Expression of T. brucei OST in mammalian cells results in improved yield and N-glycan site occupancy of therapeutic targets
[0135] Example 1: Results
[0136] We sought to investigate the impact of expressing OSTs derived from Trypanosoma brucei in mammalian cells used for the bioproduction of therapeutics Figure 2 ). To this end, a previously validated suspension-adapted CHO cell line engineered to stably express the recombinant hormone erythropoietin (CHO-EPO) was used for the study. Four OST enzymes derived from Trypanosoma brucei were selected. Each OST was cloned into a mammalian expression vector for transient expression. The OST vectors were transfected into the CHO-EPO cell line and expression of the OSTs and EPO was monitored. As a control, the CHO-EPO cell line was mock transfected with an empty mammalian expression vector.
[0137] Expression of T. brucei OST in mammalian cells
[0138] Following transfection, cell number, viability, clumping, and diameter were assessed. To confirm the expression of relevant OSTs, cell pellets from each transiently transfected cell line were collected 72 h post-transfection and sent for proteomics analysis by mass spectrometry (MS) to identify the presence of OSTs. In all transiently transfected cell lines, relevant OSTs, along with stably expressed EPO, were detected by MS. Figure 3 ).
[0139] T. brucei OST increases yield of therapeutic proteins secreted by CHO cells
[0140] To examine the effect of OST expression on the yield of therapeutic proteins produced by cells, the total protein yield in the supernatant was measured after 3 days. This showed an increase in EPO production relative to the simulated transfected cell line after each introduction of OST. Figure 4 For example, in the case of enzyme A (TbSTT3A), we observed a 2.75-fold increase in recombinant EPO production, demonstrating the predicted and expected effect of heterologous OST enzymes on improving protein secretion.
[0141] T. brucei OST increases N-glycan site occupancy of EPO produced in CHO cells
[0142] Next, we determined whether heterologous OST expression altered the N-glycosylation levels of secreted therapeutic proteins. For this purpose, EPO was purified from OST-transfected cells and digested with PNGase F for N-glycan site occupancy analysis. The number of N-glycan sites occupied on EPO-derived peptides increased from 80% to 89%. Figure 5 This demonstrates that OST expression can modify the N-glycosylation state of therapeutic proteins, even when expressed transiently for only a short period of time.
[0143] Concluding remarks
[0144] This study demonstrates that OSTs derived from Trypanosoma brevicornu can indeed be expressed in mammalian cells, and their presence increases the yield and modification of both therapeutic proteins produced in the cells. This research focuses on the production of the EPO hormone, and similar findings are expected regarding the production of other therapeutic proteins in mammalian cells.
[0145] Example 2: Methods
[0146] Transfection
[0147] Cells were transfected according to standard protocol. Briefly, cells were washed in electroporation buffer and a total of 160 ug of plasmid DNA was mixed with 40,000,000 CHO cells to a final volume of 400 uL in electroporation buffer. Electroporation was performed according to manufacturer’s instructions and 400 uL of cells were transferred to a 125 mL shake flask immediately after transfection and placed in a C02 incubator at 37C. After 40 minutes, 30 mL of growth media was added and the shake flask was shaken. Protein concentration in the supernatant was then measured.
[0148] Purification
[0149] For whole proteome analysis, cell pellets were harvested. EPO was purified with the EPO purification kit.
[0150] Mass spectrometry analysis
[0151] Whole proteome
[0152] S-Trap processing of samples
[0153] Cell pellets were lysed in lysis buffer (100 mM TEAB, 5% SDS), sonicated three times to shear DNA, and protein concentration was estimated using micro-BCA assay. An aliquot of 350 pg of protein was processed using the S-Trap mini protocol. Protein disulfide bonds were reduced in the presence of 20 mM DTT and then alkylated in 40 mM IAA. After the sample was applied to the S-Trap mini spin column, 5 washes were performed with S-Trap binding buffer. Peptides were digested with trypsin (1 :40) overnight at 37C and trypsin peptides were pooled, dried and quantified.
[0154] Mass spectrometry analysis
[0155] Peptides (1.5 pg in equivalent) were injected onto a nano-scale C18 reverse phase chromatography system and electrosprayed into an Orbitrap mass spectrometer. Peptides were eluted from the column at a constant flow rate and two blanks were run between each sample to reduce carry-over. The column was kept at a constant temperature of 50C. The MS was operated in DIA mode.
[0156] N-glycan site occupancy studies on purified EPO
[0157] PNGase treatment
[0158] Purified EPO samples were mixed with 4 μΐ of PNGase buffer and dried in a speed-Vac. Subsequently, 20 μΐ of heavy water was added to each sample and the mixture was dried again. The dried PNGase F was resuspended in 10 μΐ of heavy water, vortexed for 3 seconds and the contents added to the sample, another 10 μΐ of heavy water was added to the PNGase F vial, vortexed for 2 seconds and the contents added to the sample. The mixture (20 μΐ) was then incubated at 50°C for 20 minutes.
[0159] Peptide cleavage
[0160] Samples were reduced using DTT and alkylated by the addition of 5 μΐ of iodoacetamide (final concentration 300 mM). The samples and label were then run on a bis-tris 4-12% gradient gel, washed, stained with coomassie stain, protein bands excised and gel pieces washed.
[0161] In-gel cleavage was performed with trypsin at a final concentration of 12 μg / mL (in 20 mM ammonium bicarbonate) and incubated at 30°C on a shaker for 16 hr. Acetonitrile was added (an equal volume to cover the gel pieces) and the peptide mixture extracted by shaking at 30°C for 15 minutes. The resulting supernatant was transferred to a fresh microfuge tube. Peptides were further extracted from the gel pieces by the addition of 5% formic acid followed by 100% acetonitrile. The supernatant was collected and transferred to the first fraction. The gel pieces were finally washed with acetonitrile for 10 minutes and the three pooled fractions dried in a speed-Vac.
[0162] LC-MS / MS analysis of N-glycan site occupancy
[0163] Analysis of peptide reads was performed on a Q-exactive HF mass spectrometer coupled to a Dionex Ultimate 3000 RS. Samples were reconstituted in 1% formic acid and an aliquot of each sample was loaded at 10 μΐ / min onto a trap column equilibrated in 0.1% TFA. Peptides were eluted from the column at a constant flow rate of 300 nl / min and the column maintained at a constant temperature of 50°C.
[0164] The Q-exactive Hf was operated in data dependent positive ionisation mode. Three blanks were run between each sample to reduce carryover and the mass accuracy checked prior to the start of sample analysis.
[0165] Database searching, protein and peptide identification
[0166] Raw MS data was searched against the human uniprot proteome using the MASCOT search engine (Matrix Science, version 2.6). Protein identifications and peptide derivations were exported into Microsoft Excel.
[0167] Example 3: Stable expression of OST enzymes in CHO-S cells
[0168] The OST gene was randomly integrated into the CHO-S genome using bacterial artificial chromosomes (BACs) containing the T. brucei OST nucleic acid sequence, resulting in stable expression of the T. brucei OST enzyme. Briefly, BACs were linearized and transfected into CHO-S cells using an Amaxa nucleofector. Antibiotic selection was performed for 7 days, followed by fluorescence- assisted cell sorting (FACS) to separate live and dead cells. Individual live cells were isolated, after which single cell clones were recovered and expanded. Clones were monitored for cell viability, gene copy number, titer, and glycosylation pattern. The stable cell lines generated to date can be found in Table 1 below.
[0169]
[0170] Table 1: Cell lines for stable expression of OST enzyme.
[0171] Transient expression of OST enzymes in Expi293 cells
[0172] The OST enzyme was transiently expressed in Expi293 cells (derived from the HEK293 cell line) using mammalian expression plasmids containing the T. brucei OST nucleic acid sequence. Briefly, plasmids were transfected into Expi293 cells using Expifectamine reagent from the Expi293 Expression System. Constitutive expression of the OST enzyme occurred for 6 days, during which time the cell viability of the Expi293 cells was monitored. Figure 6 The results indicate that expression of the OST cell line in the mammalian-derived Expi293 cell line did not negatively impact cell viability.
[0173] Sequences forming part of the specification
[0174] SEQ ID NO: 1 - Consensus DNA sequence for the active region of TbSTT3A
[0175]
[0176] SEQ ID NO: 2 - DNA sequence for the active region of TbSTT3A
[0177]
[0178] SEQ ID NO: 3 - DNA sequence of the active region of TbSTT3A (CHO optimized)
[0179]
[0180] SEQ ID NO: 4 - DNA sequence of the active region of TbSTT3B
[0181]
[0182] SEQ ID NO: 5 - DNA sequence of the active region of TbSTT3B (CHO optimized)
[0183]
[0184] SEQ ID NO: 6 - DNA sequence of the active region of TbSTT3C
[0185]
[0186] SEQ ID NO: 7 - DNA sequence of the active region of TbSTT3C (CHO optimized)
[0187]
[0188] SEQ ID NO: 8 - DNA sequence of the full-length TbSTT3A
[0189]
[0190]
[0191]
[0192] SEQ ID NO: 9 - DNA sequence of the full-length TbSTT3A (CHO optimized)
[0193]
[0194]
[0195] SEQ ID NO: 10 - DNA sequence of the full-length TbSTT3B
[0196]
[0197]
[0198] SEQ ID NO: 11 - DNA sequence of full-length TbSTT3B (CHO optimized)
[0199]
[0200]
[0201] SEQ ID NO: 12 - DNA sequence of full-length TbSTT3C
[0202]
[0203]
[0204] SEQ ID NO: 13 - DNA sequence of full-length TbSTT3C (CHO optimized)
[0205]
[0206]
[0207]
[0208] SEQ ID NO: 14 - Amino acid sequence of active region of TbSTT3A
[0209]
[0210] SEQ ID NO: 15 - Amino acid sequence of active region of TbSTT3B
[0211]
[0212] SEQ ID NO: 16 - Amino acid sequence of active region of TbSTT3C
[0213]
[0214] SEQ ID NO: 17 - Amino acid sequence of full-length TbSTT3A
[0215]
[0216]
[0217] SEQ ID NO: 18 - Amino acid sequence of full-length TbSTT3B
[0218]
[0219] SEQ ID NO: 19 - Amino acid sequence of full-length TbSTT3C
[0220]
[0221]
[0222] SEQ ID NO: 20 - Conserved amino acid residues in the active region of TbSTT3A, B, and C
[0223]
[0224] SEQ ID NO: 21 - Consensus DNA sequence of the active region of TbSTT3B
[0225]
[0226] SEQ ID NO: 22 - Consensus DNA sequence of the active region of TbSTT3C
[0227]
[0228]
[0229] References
[0230] Canada et al., Cell, 136(2): 272-283, 2010.
[0231] Cherepanova & Gilmore, Scientific Reports, 6(20946), 2016.
[0232] Delobel, Mass Spectrometry of Glycoproteins, pp 1-21, 2021.
[0233] Izquierdo, Mehlert, & Ferguson, Glycobiology, 22(5): 696-703, 2012.
[0234] Jinnelov et al., J Biol Chem, 292(49):20328-20341, 2017.
[0235] Kheller & Gilmore, Glycobiology, 16(4):47-62, 2005.
[0236] Kleizen & Braakman, Curr Opin Cell Biol, 16(4):343-9, 2004.
[0237] Mohanty et al., Biomolecules, 10(4):624, 2020.
[0238] Petrescu et al., Glycobiology, 14(2):103-14, 2004.
[0239] Pfeffer et al., Nature Communications, 5 (3072), 2014.
[0240] Reilly et al., Nature Reviews Nephrology, 15(346-366), 2019.
Claims
1. A mammalian cell comprising at least one nucleic acid sequence encoding at least one oligosaccharyltransferase (OST) protein or functional fragment thereof, wherein the at least one OST protein is a Trypanosoma sp. OST protein, preferably wherein the at least one OST protein is a T. brucei OST protein.
2. The mammalian cell according to claim 1, wherein the at least one nucleic acid sequence encoding the at least one OST protein or functional fragment thereof comprises a sequence according to any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22, or a sequence having at least 70% sequence identity thereto.
3. A mammalian cell comprising at least one nucleic acid sequence encoding at least one oligosaccharyltransferase (OST) protein or functional fragment thereof, wherein the at least one nucleic acid sequence encoding the at least one OST protein or functional fragment thereof comprises a sequence according to any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21, and / or 22, or a sequence having at least 70% sequence identity thereto.
4. A mammalian cell comprising at least one nucleic acid sequence encoding at least one oligosaccharyltransferase (OST) protein or functional fragment thereof, wherein the at least one OST protein comprises an amino acid sequence according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, and / or 20, or a sequence having at least 70% sequence identity thereto.
5. The mammalian cell according to any one of claims 1 to 4, wherein the at least one nucleic acid sequence encoding the at least one OST protein or functional fragment thereof comprises a sequence according to SEQ ID NOs: 2, 3, 8, and / or 9, or a sequence having at least 70% identity thereto, and wherein the cell further comprises a nucleic acid sequence encoding a further OST protein or functional fragment thereof, and wherein the further nucleic acid sequence comprises a sequence according to SEQ ID NOs: 4, 5, 10, and / or 11, or a sequence having at least 70% identity thereto, and further comprises a sequence according to SEQ ID NOs: 6, 7, 12, and / or 13, or a sequence having at least 70% identity thereto.
6. The mammalian cell according to any preceding claim, wherein the at least one nucleic acid sequence encoding the at least one OST protein or functional fragment thereof comprises a sequence according to any one of SEQ ID NOs: 4, 5, 10, and / or 11, or a sequence having at least 70% identity thereto, and wherein the cell further comprises a nucleic acid sequence encoding a further OST protein or functional fragment thereof, and wherein the further nucleic acid comprises a sequence according to any one of SEQ ID NOs: 6, 7, 12, and / or 13, or a sequence having at least 70% identity thereto.
7. The mammalian cell according to any one of claims 1 to 4, wherein the at least one nucleic acid sequence encoding the at least one OST protein or functional fragment thereof comprises a sequence according to any one of SEQ ID NOs: 2, 3, 8 and / or 9 or a sequence having at least 70% identity thereto, and wherein the cell further comprises a nucleic acid sequence encoding a further OST protein or functional fragment thereof, and wherein the further nucleic acid sequence comprises a sequence according to any one of SEQ ID NOs: 4, 5, 10 and / or 11 or a sequence having at least 70% identity thereto.
8. The mammalian cell according to any one of claims 1 to 4, wherein the at least one nucleic acid sequence encoding the at least one OST protein or functional fragment thereof comprises a sequence according to any one of SEQ ID NOs: 2, 3, 8 and / or 9 or a sequence having at least 70% identity thereto, and wherein the cell further comprises a nucleic acid sequence encoding a further OST protein or functional fragment thereof, and wherein the further nucleic acid sequence comprises a sequence according to any one of SEQ ID NOs: 6, 7, 12 and / or 13 or a sequence having at least 70% identity thereto.
9. An isolated nucleic acid molecule comprising a sequence according to any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21 and / or 22, or a sequence having at least 70% identity thereto, preferably wherein the isolated nucleic acid molecule comprises a sequence according to SEQ ID NOs: 3, 5, 7, 9, 11 and / or 13.
10. A nucleic acid vector comprising at least one nucleic acid sequence according to any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 21 and / or 22, or comprising at least one nucleic acid according to claim 9, preferably wherein the nucleic acid vector is selected from the group comprising, but not limited to, a plasmid-based expression vector, a bacterial artificial chromosome (BAC) vector and a viral vector, such as an adenoviral vector, an adeno-associated vector (AAV), a retroviral vector or a lentiviral vector.
11. The mammalian cell according to any one of claims 1-8, wherein the mammalian cell is a Chinese hamster ovary (CHO) cell, a baby hamster kidney (BHK21) cell, a murine myeloma cell or a human embryonic kidney (HEK293) cell, preferably wherein the mammalian cell is a Chinese hamster ovary (CHO) cell, more preferably wherein the CHO mammalian cell is a stable CHO (CHO-S) mammalian cell.
12. The mammalian cell according to claim 1-3, 5-8 or 11, wherein the OST protein or functional fragment thereof has an amino acid sequence according to SEQ ID NO: 14, 15, 16, 17, 18, 19 and / or 20, or a sequence having at least 70% identity thereto, preferably wherein the OST protein is expressed by the mammalian cell as defined in any one of claims 1-3, 5-8 or 11.
13. The mammalian cell according to claim 11 or 12, wherein the mammalian cell is engineered to express a recombinant protein, preferably wherein the recombinant protein is a therapeutic protein.
14. A method of modifying the glycosylation profile of a recombinant protein, the method comprising contacting the recombinant protein with (i) the mammalian cell according to any one of claims 1-8 or 11-13; or with (ii) an OST protein or functional fragment thereof, wherein the OST protein is a Trypanosoma sp. OST protein, preferably wherein the OST protein is a T. brucei OST protein, more preferably wherein the OST protein comprises an amino acid sequence according to any one of SEQ ID NO: 14, 15, 16, 17, 18, 19 and / or 20; or producing the recombinant protein in the mammalian cell according to any one of claims 1-8 or 11-13, preferably wherein the mammalian cell is engineered to express the recombinant protein.
15. The method according to claim 14, wherein the frequency of N-glycan occupancy on the recombinant protein is altered, preferably wherein the frequency of N-glycan occupancy on the recombinant protein is increased; and / or wherein the recombinant protein is a protein selected from the group comprising a hormone, a cytokine, an antibody, an enzyme, a complement protein, a coagulation factor, a functional fragment thereof or any combination thereof.