Method for recombinant production of novel glycolipids
The production of glycolipids in recombinant host cells using enzymes E1, E2, and E3 addresses low yield and high cost issues, enabling efficient and environmentally friendly production of novel glycolipids for industrial applications.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- UNIV HOHENHEIM
- Filing Date
- 2024-11-12
- Publication Date
- 2026-05-21
AI Technical Summary
Current methods for producing glycolipids face low yields and high recovery and purification costs, and there is a need for non-pathogenic organisms to produce uncharged glycolipids efficiently using renewable raw materials.
A method for producing glycolipids in recombinant host cells by culturing cells with recombinant nucleic acid molecules encoding enzymes E1, E2, and E3 under suitable conditions, resulting in glycolipids such as glycomonolipids and glycodilipids, which can be ionic, amphoteric, or non-ionic.
The method enables high-yield production of novel glycolipids with desirable chemical properties, offering alternatives to conventional surfactants and expanding industrial applications.
Smart Images

Figure IMGF000002_0001 
Figure IMGF000003_0001 
Figure IMGF000003_0002
Abstract
Description
[0001] Method for recombinant production of novel glycolipids Technical field
[0002] The present invention relates to a method for producing glycolipids in recombinant host cells and to a novel glycolipid obtained by said method as well as nucleic acid molecules used in this method.
[0003] Background of the invention
[0004] Microbial biosurfactants have become increasingly attractive as promising ingredients for environmentally and climate-friendly products. The reasons for this are their generally good performance and biodegradability, low toxicity, production from renewable raw materials and benefits for the climate and environment perceived by consumers. Some microbial biosurfactants also display desired performance characteristics such as foaming, low critical micelle concentration (CMC), efficient emulsification and strong reduction of interfacial tension.
[0005] Based on their utility in detergent, paint, cosmetics, textile, agriculture, food and pharmaceutical industries, there is a general demand in the market for such biosurfactants. Biosurfactants may also potentially replace conventional surfactants made from petroleum or products thereof, and thus invariably improve the environmental performance of the resulting formulations.
[0006] Microorganisms such as yeasts, bacteria and some filamentous fungi are capable of producing biosurfactants with different molecular structures and surface activities. Different types of biosurfactants such as glycolipid biosurfactants, peptide biosurfactants and fatty acid biosurfactants are known. Glycolipids are one of the most important classes of biosurfactants, as they possess various bioactivities as well as excellent surface lowering properties. Glycolipids comprise at least one fatty acid chain and a carbohydrate moiety linked by a glycosidic or ester bond. The variety of glycolipids includes rhamnolipids, sophorolipids, mannosylerythritol lipids, cellobiose lipids, trehalolipids, xylolipids, and glucoselipids which differ in the fatty acid chains and / or the carbohydrate moiety. Many currently used methods to produce biosurfactants employ wild-type isolates of various bacteria naturally producing the biosurfactants but face low yields in production processes and high recovery and purification costs. For example, the gram negative enterobacter Rouxiella sp. DSM 100043 was described as a new glycolipid producing species in a study by J. H. Kugler et al. (2015) AMB Express 5(1 ):82. Methods for production of biosurfactants such as sophorolipids in non-pathogenic yeasts are described in WO 2020 / 006194 A1 and WO 2020 / 104582 A1 and the production of rhamnolipids in non-pathogenic organisms using carbon sources is taught in WO 2012 / 013554 A1 and EP 2 949 214 A1. The production and separation of a trehalolipid biosurfactant is described in a study by D. A. White et al. (2013) J Appl Microbiol 115(3):744-55.
[0007] Nevertheless, there is still a need to produce glycolipids efficiently and in high amounts using non-pathogenic organisms and a renewable raw material as an alternative to conventional surfactants made of petroleum. Moreover, since most commercially available biosurfactants are negatively charged, it is desirable to provide uncharged glycolipids having interesting chemical properties which can be advantageously used for new industrial applications.
[0008] Brief summary of the invention
[0009] The present invention is directed to a method for the production of a glycolipid comprising a glucose molecule linked to at least one fatty acid molecule, comprising culturing recombinant host cells under suitable conditions for the production of said glycolipid, wherein the recombinant host cells comprise one or more recombinant nucleic acid molecules encoding at least one enzyme selected from the group consisting of the enzymes E1 , E2 and E3. The present invention further relates to a novel glycolipid obtained by said method and to isolated nucleic acid molecules used in said method.
[0010] Brief description of the drawings
[0011] Figure 1. LC-MS / MS analysis of the purified glycodilipid produced by E. coll pCAT2: The base peak chromatogram in negative ion mode shows an intense signal of the glycolipid at a retention time of 13.26 min. Figure 2. LC-MS / MS analysis of the purified glycomonolipid produced by E. coll pAFP1 : The base peak chromatogram in negative ion mode shows an intense signal of the glycolipid at a retention time of 17.02 min. Figure 3. Time course of fed-batch bioreactor cultivations of E. coli pCAT2 showing glucose concentration (triangle), cell dry weight (cross), and glucoselipid (glycodilipid) concentration (circle).
[0012] Figure 4. Stability test of 50 mg / L glycodilipid biosurfactant isolated from R. badensis DSM 100043Tby measuring emulsification index in olive oil in dependence of temperature (A), salinity (B), and pH (C).
[0013] Detailed description of the invention
[0014] General definitions
[0015] Before the invention is described in detail with respect to some of its preferred embodiments, the following general definitions are provided.
[0016] The present invention as illustratively described in the following may suitably be practiced in the absence of any element or elements, limitation or limitations, not specifically disclosed herein.
[0017] The present invention will be described with respect to particular embodiments and with reference to certain figures, but the invention is not limited thereto but only by the claims.
[0018] Where the term “comprising” is used in the present description and claims, it does not exclude other elements. For the purposes of the present invention, the term “consisting of is considered to be a preferred embodiment of the term “comprising”. If hereinafter a group is defined to comprise at least a certain number of embodiments, this is also to be understood to disclose a group which preferably consists only of these embodiments.
[0019] For the purposes of the present invention, the term “obtained” is considered to be a preferred embodiment of the term “obtainable”. If hereinafter e.g. a compound is defined to be obtainable from a specific source, this is also to be understood to disclose a compound which is obtained from this source.
[0020] Where an indefinite or definite article is used when referring to a singular noun, e.g. “a”, “an” or “the”, this includes a plural of that noun unless something else is specifically stated. The terms “about” or “approximately” in the context of the present invention denote an interval of accuracy that the person skilled in the art will understand to still ensure the technical effect of the feature in question. The term typically indicates deviation from the indicated numerical value of ±10%, and preferably of ±5%.
[0021] Technical terms are used by their common sense. If a specific meaning is conveyed to certain terms, definitions of terms will be given in the following in the context of which the terms are used.
[0022] General definition of glycolipids
[0023] The present disclosure relates to a method for the production of a glycolipid. According to the invention, a glycolipid is a compound comprising a glucose molecule linked to at least one fatty acid molecule. Within the present application the terms glycolipid and glucoselipid are used interchangeably.
[0024] In one embodiment the produced glycolipid can be ionic, meaning it can be positively or negatively charged. In another embodiment, the produced glycolipid can be amphoteric, meaning it can be both positively and negatively charged. In still another embodiment, the produced glycolipid can be non-ionic, meaning it is neither positively nor negatively charged. In a preferred embodiment, the produced glycolipid is non-ionic.
[0025] A glucose molecule is a monosaccharide, that is made of 6 carbon atoms, 12 hydrogen atoms, and 6 oxygen atoms. The positions C1-C6 of the carbon atoms comprised in a glucose molecule are numbered as shown in the following structure:
[0026] <
[0027]
[0028] According to the invention, at least one fatty acid molecule can be linked to at least one of the positions C1-C6 of the carbon atoms comprised in a glucose molecule to form the glycolipid.
[0029] A fatty acid molecule comprises a straight chain of a number of carbon atoms, with hydrogen atoms along the length of the chain and at one end of the chain and a carboxyl group at the other end of the chain. Fatty acid molecules can be differentiated by the number of carbon atoms comprised in the chain. In one embodiment, fatty acid molecules are selected from the group consisting of propanoic acid molecules (C3), butanoic acid molecules (C4), pentanoic acid molecules (C5), hexanoic acid molecules (Ce), heptanoic acid molecules (C7), octanoic acid molecules (Cs), nonanoic acid molecules (Cg), decanoic acid molecules (C10), undecanoic acid molecules (On), dodecanoic acid molecules (C12), tridecanoic acid molecules (C13), tetradecanoic acid molecules (C14), pentadecanoic acid molecules (C15), hexadecanoic acid molecules (Cie), heptadecanoic acid molecules (C17), octadecanoic acid molecules (Cis), nonadecanoic acid molecules (C19) and icosanoic acid molecules ( 20). In one embodiment, the fatty acid molecules linked to the glucose molecule are selected from the group consisting of decanoic acid molecules (C10) and dodecanoic acid molecules (C12).
[0030] In one embodiment, the fatty acid molecule linked to the glucose molecule is a decanoic acid molecule (C10). The decanoic acid molecule in its standard form is a saturated fatty acid molecule consisting of a 10-carbon chain with single bonds between the carbon atoms and a carboxyl group (-COOH) at one end. In one embodiment, the decanoic acid molecule is modified by introducing one or more double bonds, resulting in an unsaturated decanoic acid molecule and / or by adding hydroxyl (-OH) groups to different carbon atoms in the chain, resulting in a hydroxylated decanoic acid molecule. In a preferred embodiment, the decanoic acid molecule is a hydroxylated decanoic acid molecule such as a 2-hydroxydecanoic acid molecule, a 3-hydroxydecanoic acid molecule, a 4-hydroxydecanoic acid molecule or a 5-hydroxydecanoic acid molecule. More preferably, the decanoic acid molecule linked to the glucose molecule is a 3-hydroxydecanoic acid molecule as shown in the following structure:
[0031]
[0032] In one embodiment, the fatty acid molecule linked to the glucose molecule is a dodecanoic acid molecule (C12). The dodecanoic acid molecule in its standard form is a saturated fatty acid molecule consisting of a 12-carbon chain with single bonds between the carbon atoms and a carboxyl group (-COOH) at one end. In one embodiment, the dodecanoic acid molecule is modified by introducing one or more double bonds, resulting in an unsaturated dodecanoic acid molecule and / or by adding hydroxyl (-OH) groups to different carbons in the chain, resulting in a hydroxylated dodecanoic acid molecule. In a preferred embodiment, the dodecanoic acid molecule is a monounsaturated hydroxylated dodecanoic acid molecule such as a 2-hydroxy-7-dodecenoic acid molecule, a 3-hydroxy-5-dodecenoic acid molecule, a 4-hydroxy-2-dodecenoic acid molecule, a 6-hydroxy-4-dodecenoic acid molecule or a 7-hydroxy-9-dodecenoic acid molecule. More preferably, the dodecanoic acid molecule linked to the glucose molecule is a 3-hydroxy-5-dodecenoic acid molecule as shown in the following structure:
[0033] In one embodiment, the produced glycolipid comprises a glucose molecule linked to one fatty acid molecule and no further fatty acid molecules are linked to the glucose molecule. In a preferred embodiment, the produced glycolipid comprises a glucose molecule linked to a dodecanoic acid molecule. More preferred, the produced glycolipid comprises a glucose molecule linked to a 3-hydroxy-5-dodecenoic acid molecule.
[0034] In one embodiment, the produced glycolipid comprises a glucose molecule linked to one fatty acid molecule at position C2 and no further fatty acid molecules are linked to the glucose molecule. Such a glycolipid is herein also designated as glycomonolipid. In a preferred embodiment, the produced glycolipid comprises a glucose molecule
[0035]
[0036] linked to a dodecanoic acid molecule at position C2. More preferably, the produced glycolipid comprises a glucose molecule linked to a 3-hydroxy-5-dodecenoic acid molecule at position C2. Most preferably, the produced glycolipid has the following structure (I):
[0037]
[0038] In another embodiment, the produced glycolipid comprises a glucose molecule linked to two fatty acid molecules and no further fatty acid molecules are linked to the glucose molecule. Such a glycolipid is herein also designated as glycodilipid. In a preferred embodiment, produced glycolipid comprises a glucose molecule linked to a dodecanoic acid molecule and a decanoic acid molecule. More preferably, the produced glycolipid comprises a glucose molecule linked to a 3-hydroxy-5-dodecenoic acid molecule and a 3-hydroxydecanoic acid molecule. In one embodiment, the produced glycolipid comprises a glucose molecule linked to a first fatty acid molecule at position C2 and a second fatty acid molecule at position C3 and no further fatty acid molecules are linked to the glucose molecule. In a preferred embodiment, the produced glycolipid comprises a glucose molecule linked to a dodecanoic acid molecule at position C2 and a decanoic acid molecule at position C3. More preferably, the produced glycolipid comprises a glucose molecule linked to a 3-hydroxy-5-dodecenoic acid molecule at position C2 and a 3-hydroxydecanoic acid molecule at position C3. Most preferably, the produced glycolipid has the following structure (II):
[0039]
[0040] In one embodiment, the recombinant cell produces a glycolipid comprising at least one fatty acid molecule being linked to a glucose molecule by an ester bond. In a preferred embodiment, the ester bond is formed between the carboxyl group of the fatty acid molecules and the hydroxyl group of the glucose molecule. An ester bond is a bond between a hydroxyl group (-OH) and a carboxyl group (-COOH), formed by the elimination of a molecule of water (H2O).
[0041] Recombinant cells
[0042] The method of the invention comprises culturing recombinant host cells under suitable conditions for the production of a glycolipid, wherein the recombinant host cells comprise one or more recombinant nucleic acid molecules encoding at least one enzyme selected from the group consisting of the enzymes E1 , E2 and E3.
[0043] As used herein, the terms “produce” or “production” refer to the synthesis of a molecule. The synthesis of a molecule can comprise one or more chemical reactions that can be catalyzed by one or more enzymes and can take place within a cell or organism or within a cell-free environment comprising the required enzymes. In one embodiment, production means the synthesis of a molecule within a cell and the further isolation of said molecule within a cell-free environment. In a preferred embodiment, said molecule is a glycolipid produced by recombinant host cells. The term “catalyzing” or “catalyze” as used herein when referring to an enzymatic reaction means to cause or accelerate the initiation or the progression of a chemical reaction. Enzymes may use cellular or thermal energy and / or proton or electron donors and acceptors while catalyzing reactions. Catalyzing means reducing the activation energy needed to start a reaction by weakening the chemical bonds, usually by temporarily bonding with the reacting molecules.
[0044] The cells used in the method of the present invention can be prokaryotic or eukaryotic cells, such as mammalian cells (such as, for example, human cells, CHO (Chinese Hamster Ovary) cells, HEK293 (Human Embryonic Kidney) cells or NS0 (mouse myeloma) cells), plant cells or microorganisms such as yeasts, fungi or bacteria, wherein microorganisms are preferred, and bacterial cells are most preferred.
[0045] Suitable bacteria, yeasts or fungi are in particular those bacteria, yeasts or fungi that are deposited in the Deutsche Sammlung von Mikroorganismen und Zellkulturen (German Collection of Microorganisms and Cell Cultures) GmbH (DSMZ), Braunschweig, Germany, as bacterial, yeast or fungal strains.
[0046] According to the invention, the bacterial cells may be selected from the genera Aspergillus, Corynebacterium, Brevibacterium, Bacillus, Acinetobacter, Alcaligenes, Lactobacillus, Paracoccus, Lactococcus, Candida, Pichia, Hansenula, Kluyveromyces, Saccharomyces, Escherichia, Zymomonas, Yarrowia, Methylobacterium, Ralstonia, Pseudomonas, Rhodospirillum, Rhodobacter, Burkholderia, Clostridium and Cupriavidus. More in particular, the bacterial cells may be selected from the group consisting of Aspergillus nidulans, Aspergillus niger, Alcaligenes latus, Bacillus megaterium, Bacillus subtilis, Brevibacterium flavum, Brevibacterium lactofermentum, Burkholderia andropogonis, B. brasilensis, B. caledonica, B. caribensis, B. caryophylili, B. fungorum, B. gladioli, B. glathei, B. glumae, B. graminis, B. hospita, B. kururiensis, B. phenazinium, B. phymatum, B. phytofirmans, B. plantarii, B. sacchari, B. singaporensis, B. sordidicola, B. terricola, B. tropica, B. tuberum, B. ubonensis, B. unamae, B. xenovorans, B. anthina, B. pyrrocinia, B. thailandensis, Candida blankii, Candida rugosa, Corynebacterium glutamicum, Corynebacterium efficiens, Escherichia coli, Hansenula polymorpha, Kluveromyces lactis, Methylobacterium extorquens, Paracoccus versutus, Pseudomonas argentinensis, P. borbori, P. citronellolis, P. flavescens, P. mendocina, P. nitroreducens, P. oleovorans, P. pseudoalcaligenes, P. resinovorans, P. straminea, P. aurantiaca, P. aureofaciens, P. chlororaphis, P. fragi, P. lundensis, P. taetrolens, P. antarctica, P. azotoformans, ’P. blatchfordae’, P. brassicacearum, P. brenneri, P. cedrina, P. corrugata, P. fluorescens, P. gessardii, P. libanensis, P. mandelii, P. marginalis, P. mediterranea, P. meridiana, P. migulae, P. mucidolens, P. orientalis, P. panacis, P. proteolytica, P.rhodesiae, P. synxantha, P. thivervalensis, P. tolaasii, P. veronii, P. denitrificans, P. pertucinogena, P. cremoricolorata, P. fulva, P. monteilii, P. mosselii, P. parafulva, P. putida, P. balearica, P. stutzeri, P. amygdali, P. avellanae, P. caricapapayae, P. cichorii, P. coronafaciens, P. ficuserectae, 'P. helianthi’, P. meliae, P. savastanoi, P. syringae, P. tomato, P. viridiflava, P. abietaniphila, P. acidophila, P. agaric!, P. alcaliphila, P. alkanolytica, P. amyloderamosa, P. asplenii, P. azotifigens, P. cannabina, P. coenobios, P. congelans, P. costantinii, P. cruciviae, P. delhiensis, P. excibis, P. extremorientalis, P. frederiksbergensis, P. fuscovaginae, P. gelidicola, P. grimontii, P. indica, P. jessenii, P. jinjuensis, P. kilonensis, P. knackmussii, P. koreensis, P. Uni, P. lutea, P. moraviensis, P. otitidis, P. pachastrellae, P. palleroniana, P. papaveris, P. pell, P. perolens, P. poae, P. pohangensis, P. psychrophila, P. psychrotolerans, P. rathonis, P. reptilivora, P. resiniphila, P. rhizosphaerae, P. rubescens, P. salomonii, P. segitis, P. septica, P. simiae, P. suis, P. thermotolerans, P. aeruginosa, P.tremae, P.trivialis, P. turbinellae, P.tuticorinensis, P. umsongensis, P. vancouverensis, P. vranovensis, P. xanthomarina, Ralstonia eutropha, Rhodospirillum rubrum, Rhodobacter sphaeroides, and Zymomonas mobile.
[0047] Suitable fungal cells for the production of glycolipids may be selected from the group consisting of Saccharomyces cerevisiae, Yarrowia lipolytica, Candida bombicola and Candida tropicalis.
[0048] Preferably, the cells used in the method of the present invention are Escherichia cells, more preferably they are Escherichia coli (E. coli) cells.
[0049] In one embodiment, the Escherichia coli cells are selected from the group of strains consisting of wild-type, BL21 (DE3), MG1655, BI21 , 60E4, Ita23 or lta36A. In a preferred embodiment, the E. coli cells are from the E. coli BL21 (DE3) strain.
[0050] E. coli BL21 (DE3) is a chemically competent E. coli strain suitable for transformation and protein expression. Cells of said strain carry a chromosomal copy of the T7 RNA polymerase gene under the control of the lacUV5 promoter. Such strains are suitable for production of proteins from target genes cloned in appropriate T7 expression vectors, using isopropyl p-D-1 -thiogalactopyranoside (IPTG) as an inducer. The term “expression” or “gene expression” as used herein refers to the process of synthesis of a gene product, preferably a functional RNA or protein. Gene expression generally comprises DNA transcription, optionally RNA processing and in the case of protein-expressing genes, RNA translation into protein. According to the invention, the expressed protein can be an enzyme, being a specific type of protein that catalyzes a specific reaction.
[0051] Genetic modifications
[0052] “Recombinant host cells” used in the method of the invention are cells that produce glycolipids through genetic engineering.
[0053] For the purposes of the invention, "recombinant" with regard to cells means that said recombinant cells have been genetically modified such that said cells differ from the wild-type of these cells. Said recombinant cells express a recombinant polypeptide characterized by an amino acid sequence which is encoded by a polynucleotide which is introduced by gene technology. Said polynucleotide includes all those constructions brought about by man, by gene technology I recombinant DNA techniques in which either
[0054] (a) the sequence of the polynucleotide or a part thereof, or
[0055] (b) one or more genetic control sequences which are operably linked with the polynucleotide, including but not limited thereto a promoter, or
[0056] (c) both a) and b)
[0057] are not located in their wildtype genetic environment or have been modified.
[0058] “Wild-type" of a cell herein designates a cell, the genome of which is present in a state as is formed naturally by evolution. The term “wild-type” is used herein both for the entire cell as well as for individual genes within the cell. The term "wild-type" therefore in particular does not include those cells or those genes, the gene sequences of which have been modified at least partially by man by means of recombinant methods. According to any aspect of the present invention, the wild-type cells may be incapable of forming detectable amounts of glycolipids and / or have no or no detectable activity of one or more of the enzymes E1 , E2 and E3.
[0059] The terms “introduction of a polynucleotide” or “transformation of a polynucleotide” as referred to herein encompass the transfer of an exogenous polynucleotide into a host cell, irrespective of the method used for the transfer. That is, the term “transformation of a polynucleotide” as used herein is independent of vector, shuttle system, or host cell, and it not only relates to the polynucleotide transfer method of transformation as known in the art (cf. , for example, Sambrook, J. et al. (1989) Molecular Cloning: A Laboratory Manual, 2nd Ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY), but it encompasses any further kind of polynucleotide transfer methods such as, but not limited to, transduction or transfection.
[0060] To express the polypeptides in a recombinant host cell, a recombinant nucleic acid molecule, preferably an expression cassette comprising one or more of the nucleotide sequences as described herein, is used. Typically, the expression cassette comprises three elements: a promoter sequence, an open reading frame and a 3' untranslated region that, in eukaryotes, usually contains a polyadenylation site. Additional regulatory elements may include transcriptional as well as translational enhancers. The promoter may be either constitutive or inducible. A constitutive promoter drives continuous, unregulated expression of a polypeptide, whereas an inducible promoter controls the expression of a polypeptide in response to specific environmental signals or chemicals. Examples for inducible promoters and their inducers include, but are not limited to a lac operon promoter induced by IPTG, Tet-On / Tet-Off systems induced by tetracycline, galactose inducible promoters and heat shock promoters activated by heat or stress conditions. According to the disclosure, the preferred inducible promoter is a lac operon promoter induced by IPTG. An intron sequence may also be added to the 5' untranslated region (UTR) or to the coding sequence to increase the amount of the mature message that accumulates in the cytosol. The expression cassette may be part of a vector which is present separate from the genome of the host cell or may be integrated into the genome of a host cell and replicated together with the genome of its host cell. The expression cassette usually is capable of increasing or decreasing expression of a polypeptide. In the method of the present invention the expression cassette is capable of increasing the expression of at least one enzyme selected from the group consisting of the enzymes E1 , E2 and E3.
[0061] Also disclosed herein is an expression vector comprising one or more of the nucleotide sequences as described herein. The term “vector” as used herein comprises any kind of construct suitable to carry foreign polynucleotide sequences for transfer to another cell, or for stable or transient expression within a given cell. The polynucleotide encoding the polypeptide used in the method of the present invention may be introduced into a vector by means of standard recombinant DNA techniques. Once introduced into the vector, the polynucleotide comprising a coding sequence may be suitable to be introduced (by transformation, transduction, transfection, etc.) into a host cell or host cell organelles. A cloning vector may be chosen which is suitable for expression of the polynucleotide sequence in the host cell or host cell organelles. Suitable vectors for this purpose are available to the person skilled in the art and can be taken, for example, from the brochures of companies including Novagen, Promega, New England Biolabs, Biosearch technologies, Cambio, VWR, Clontech or Gibco BRL. According to the disclosure, the cloning vector belongs to the pET (plasmid for expression by T7 RNA polymerase) system. Preferably, the cloning vector is a pET21 a(+) vector.
[0062] A vector as used herein may provide additional segments fortranscription and translation of a foreign polynucleotide upon transformation into a host cell or host cell organelles. Such additional segments may include regulatory nucleotide sequences, one or more origins of replication that is required for its maintenance and / or replication in a specific cell type, one or more selectable markers, a polyadenylation signal, a suitable site for the insertion of foreign coding sequences such as a multiple cloning site etc. Non-limiting examples of suitable origins of replication include the f1-ori and colE1.
[0063] The term “polypeptide” is defined herein as a polymer consisting of at least ten amino acids. A heterologous, exogenous, foreign or recombinant polypeptide as described herein is a polypeptide which is not native to the host cell.
[0064] According to the disclosure said heterologous, exogenous, foreign or recombinant polypeptide can be a polypeptide native to the host cell in which structural modifications, e.g., deletions, substitutions, and / or insertions, have been made by recombinant DNA techniques to alter the native polypeptide, or a polypeptide native to the host cell whose expression is quantitatively altered or whose expression is directed from a genomic location different from that in the native host cell as a result of manipulation of the DNA ofthe host cell by recombinant DNA techniques, or whose expression is quantitatively altered as a result of manipulation of the regulatory elements of the polynucleotide by recombinant DNA techniques e.g., a stronger promoter; or a polynucleotide native to the host cell, but not integrated within its natural genetic environment as a result of genetic manipulation by recombinant DNA techniques.
[0065] The term “amino acid sequence” herein refers to the specific order in which amino acids are arranged to form a protein or polypeptide. The sequence of amino acids can be defined using different abbreviations for each amino acid. There are 20 standard amino acids commonly found in proteins, and they are represented by their abbreviations, i.e. each amino acid can be denoted by a single letter (one-letter code) or by a three-letter abbreviation (three-letter code).
[0066] The terms “nucleic acid” or “nucleic acid molecule” are used interchangeably herein to refer to a biomolecule composed of nucleotides. The terms “nucleic acid sequence” or “nucleotide sequence” refer to the order of nucleotides in a strand of DNA or RNA. These sequences determine the genetic information that can be transcribed and translated into proteins. The nucleic acid can be synthesized in vitro or can be naturally occurring, i.e. isolated from nature. The nucleic acid molecule can be comprised within a eukaryotic or prokaryotic organism, a eukaryotic or prokaryotic cell, a cell nucleus or a cell organelle, as part of a genome or as an individual molecule. According to the disclosure a synthesized nucleic acid molecule comprising elements not linked in nature is a recombinant nucleic acid molecule and it can be comprised within a plasmid, a vector or an artificial chromosome.
[0067] The terms “sequence identity”, “% sequence identity”, “% identity”, “% identical” or “sequence alignment” are used interchangeably herein and refer to the result of the comparison of a first nucleic acid sequence to a second nucleic acid sequence, or a comparison of a first amino acid sequence to a second amino acid sequence. The sequence identity is calculated as a percentage based on the comparison. The result of this calculation can be described as “percent identical” or “percent ID.” A sequence identity may be determined by a program, which produces an alignment, and calculates identity counting both mismatches at a single position and gaps at a single position as non-identical positions in the final sequence identity calculation. Preferably, the sequence identity is determined over the entire length of the first and second nucleic acid sequence.
[0068] According to this invention, a pairwise global alignment is produced, meaning that two sequences are aligned over their complete length, which is usually produced by using a mathematical approach, called alignment algorithm. According to the invention, the alignment is generated by using the algorithm of Needleman and Wunsch (J. Mol. Biol. (1979) 48, p. 443-453). Preferably, the program “NEEDLE” (The European Molecular Biology Open Software Suite (EMBOSS)) is used for the purposes of the current invention, with using the programs default parameter (polynucleotides: gap open=10.0, gap extend=0.5 and matrix=EDNAFULL; polypeptides: gap open=10.0, gap extend=0.5 and matrix=EBLOSUM62). After aligning two sequences, in a second step, an identity value is determined from the alignment produced. For this purpose, the %-identity is calculated by dividing the number of identical residues by the length of the alignment region which is showing the respective sequence over its complete length multiplied with 100: %-identity = (identical residues / length of the alignment region which is showing the respective sequence over its complete length) *100.
[0069] For calculating the percent identity of two nucleic acid sequences the same applies as for the calculation of percent identity of two amino acid sequences with some specifications. For nucleic acid sequences encoding a protein the pairwise alignment shall be made over the complete length of the coding region of the sequence from start to stop codon excluding introns. Introns present in the other sequence, to which the sequence is compared, shall also be removed for the pairwise alignment. After aligning two sequences, in a second step, an identity value is determined from the alignment produced. Percent identity is calculated by %-identity = (identical residues I length of the alignment region which is showing the sequence from start to stop codon excluding introns over its complete length) *100.
[0070] Enzymes
[0071] In the present invention, the recombinant cell producing glycolipids comprises one or more recombinant nucleic acid molecules for the expression of at least one enzyme selected from the group consisting of the enzymes E1 , E2 and E3. Preferably, the recombinant cell producing glycolipids comprises one or more recombinant nucleic acid molecules for the expression of at least two enzymes selected from the group consisting of the enzymes E1 , E2 and E3. More preferably, the recombinant cell producing glycolipids comprises one or more recombinant nucleic acid molecules for the expression of the enzymes E1 , E2 and E3.
[0072] The term "fragment of a polypeptide" as used herein refers to a portion of an amino acid sequence comprising a deletion of one or more amino acids at the N terminus and / or the C terminus of the polypeptide. Fragments which essentially retain the same enzyme activity as the full-length protein are herein designated as "functional fragments".
[0073] Enzyme E1
[0074] In one embodiment, the recombinant cell comprises a recombinant nucleic acid molecule encoding the enzyme E1. In one embodiment, the enzyme E1 catalyzes the transfer of an acyl group to a substrate. In one embodiment, the amino acid sequence of the enzyme E1 is at least 70%, at least 75 %, at least 78 %, at least 80 %, at least 82 %, at least 84 %, at least 85 %, at least 86 %, at least 87 %, at least 88 %, at least 89 %, at least 90 %, at least 90.5%, at least 91 %, at least 91.5%, at least 92%, at least 92.5%, at least 93%, at least 93.5%, at least 94%, at least 94.5%, at least 95%, at least 95.1 %, at least 95.2%, at least 95.3%, at least 95.4%, at least 95.5%, at least 95.6%, at least 95.7%, at least 95.8%, at least 95.9% at least 96%, at least 96.1%, at least 96.2%, at least 96.3%, at least 96.4%, at least 96.5%, at least 96.6%, at least 96.7%, at least 96.8%, at least 96.9%, at least 97%, at least 97.1 %, at least 97.2%, at least 97.3%, at least 97.4%, at least 97.5%, at least 97.6%, at least 97.7%, at least 97.8%, at least 97.9%, at least 98.0%, at least 98.1 %, at least 98.2%, at least 98.3%, at least 98.4%, at least 98.5%, at least 98.6%, at least 98.7%, at least 98.8%, at least 98.9%, at least 99.0%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% identical to the amino acid sequence according to SEQ ID NO: 1.
[0075] In another embodiment, the enzyme E1 has an amino acid sequence that is encoded by a nucleic acid sequence which at least 70%, at least 75 %, at least 78 %, at least 80 %, at least 82 %, at least 84 %, at least 85 %, at least 86 %, at least 87 %, at least 88 %, at least 89 %, at least 90 %, at least 90.5%, at least 91%, at least 91.5%, at least 92%, at least 92.5%, at least 93%, at least 93.5%, at least 94%, at least 94.5%, at least 95%, at least 95.1 %, at least 95.2%, at least 95.3%, at least 95.4%, at least 95.5%, at least 95.6%, at least 95.7%, at least 95.8%, at least 95.9% at least 96%, at least 96.1 %, at least 96.2%, at least 96.3%, at least 96.4%, at least 96.5%, at least 96.6%, at least 96.7%, at least 96.8%, at least 96.9%, at least 97%, at least 97.1%, at least 97.2%, at least 97.3%, at least 97.4%, at least 97.5%, at least 97.6%, at least 97.7%, at least 97.8%, at least 97.9%, at least 98.0%, at least 98.1%, at least 98.2%, at least 98.3%, at least 98.4%, at least 98.5%, at least 98.6%, at least 98.7%, at least 98.8%, at least 98.9%, at least 99.0%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% identical to the nucleic acid sequence according to SEQ ID NO: 4.
[0076] In one embodiment, the recombinant cell comprises the enzyme E1 having the amino acid sequence according to SEQ ID NO: 1 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 1. In another embodiment, the recombinant cell comprises the enzyme E1 having an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 4 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 4. In another embodiment, the recombinant cell comprises a functional fragment of the enzyme E1 having the amino acid sequence according to SEQ ID NO: 1 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 1.
[0077] Preferably, the functional fragment of the enzyme E1 has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80% identical, at least 81 %, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 98.5 %, at least 99%, or at least 99.5% of the length of the original full length amino acid sequence. Preferably, the functional fragment of the enzyme E1 comprises 100 to 195, 120 to 195, 140 to 195, 160 to 195, 180 to 195, or 190 to 195 consecutive amino acids of the full-length polypeptide. Also preferably, the functional fragment of the enzyme E1 retains at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80% identical, at least 81 %, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 98.5 %, at least 99%, at least 99.5% or at least 100% of the enzyme activity of the original full length amino acid sequence. The functional fragment of the enzyme E1 comprises consecutive amino acids compared to the original full length amino acid sequence, respectively. In said embodiment, the "original full length amino acid sequence" of the enzyme E1 is an amino acid sequence which is at least 70% identical to the amino acid sequence according to SEQ ID NO: 1 and which has the same length as the amino acid sequence according to SEQ ID NO: 1.
[0078] In one embodiment, the fragment of the enzyme E1 does not comprise any internal deletions compared to the amino acid sequence according to SEQ ID NO: 1. However, as discussed above, the fragment of the enzyme E1 comprises a deletion of one or more amino acids at the N terminus and / or the C terminus of the polypeptide. In another embodiment, the recombinant cell comprises the enzyme E1 that is encoded by a fragment of a nucleic acid sequence according to SEQ ID NO: 4 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 4.
[0079] According to the invention, any enzyme E1 comprising an amino acid sequence with less than 100% sequence identity according to SEQ ID NO: 1 has at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80% identical, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 98.5 %, at least 99%, at least 99.5% or at least 100% of the enzyme activity of the enzyme E1 comprising an amino acid sequence with 100% sequence identity according to SEQ ID NO: 1. According to the disclosure, any enzyme E1 encoded by a nucleic acid sequence with less than 100% sequence identity according to SEQ ID NO: 4 has at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80% identical, at least 81 %, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 98.5 %, at least 99%, at least 99.5% or at least 100% of the enzyme activity of the enzyme activity of the enzyme E1 encoded by a nucleic acid sequence with 100% sequence identity according to SEQ ID NO: 4.
[0080] According to the invention, “the enzyme activity of enzyme E1” means that when enzyme E1 is expressed with the enzymes E2 comprising the amino acid sequence according to SEQ ID NO: 2 and E3 comprising the amino acid sequence according to SEQ ID NO: 3 in a recombinant host cell, a glycodilipid according to structure (II) is produced.
[0081] Enzyme E2
[0082] In one embodiment, the recombinant cell comprises a recombinant nucleic acid molecule encoding the enzyme E2. In one embodiment, the enzyme E2 has an acyltransferase activity. An acyltransferase is an enzyme that catalyzes the transfer of an acyl group (a functional group derived from a carboxylic acid, typically R-C=O) from one molecule to another. An acyltransferase can catalyze the transfer of an acyl group from an acyl-CoA molecule to an alcohol and form ester bonds thereby.
[0083] In one embodiment, the amino acid sequence of the enzyme E2 is at least 70%, at least 75 %, at least 78 %, at least 80 %, at least 82 %, at least 84 %, at least 85 %, at least 86 %, at least 87 %, at least 88 %, at least 89 %, at least 90 %, at least 90.5%, at least 91 %, at least 91.5%, at least 92%, at least 92.5%, at least 93%, at least 93.5%, at least 94%, at least 94.5%, at least 95%, at least 95.1%, at least 95.2%, at least 95.3%, at least 95.4%, at least 95.5%, at least 95.6%, at least 95.7%, at least 95.8%, at least 95.9% at least 96%, at least 96.1 %, at least 96.2%, at least 96.3%, at least 96.4%, at least 96.5%, at least 96.6%, at least 96.7%, at least 96.8%, at least 96.9%, at least 97%, at least 97.1 %, at least 97.2%, at least 97.3%, at least 97.4%, at least 97.5%, at least 97.6%, at least 97.7%, at least 97.8%, at least 97.9%, at least 98.0%, at least 98.1%, at least 98.2%, at least 98.3%, at least 98.4%, at least 98.5%, at least 98.6%, at least 98.7%, at least 98.8%, at least 98.9%, at least 99.0%, at least 99.1 %, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% identical to an amino acid sequence according to SEQ ID NO: 2.
[0084] In one embodiment, the enzyme E2 has an amino acid sequence that is encoded by a nucleic acid sequence which is at least 70%, at least 75 %, at least 78 %, at least 80 %, at least 82 %, at least 84 %, at least 85 %, at least 86 %, at least 87 %, at least 88 %, at least 89 %, at least 90 %, at least 90.5%, at least 91%, at least 91.5%, at least 92%, at least 92.5%, at least 93%, at least 93.5%, at least 94%, at least 94.5%, at least 95%, at least 95.1 %, at least 95.2%, at least 95.3%, at least 95.4%, at least 95.5%, at least 95.6%, at least 95.7%, at least 95.8%, at least 95.9% at least 96%, at least 96.1%, at least 96.2%, at least 96.3%, at least 96.4%, at least 96.5%, at least 96.6%, at least 96.7%, at least 96.8%, at least 96.9%, at least 97%, at least 97.1%, at least 97.2%, at least 97.3%, at least 97.4%, at least 97.5%, at least 97.6%, at least 97.7%, at least 97.8%, at least 97.9%, at least 98.0%, at least 98.1 %, at least 98.2%, at least 98.3%, at least 98.4%, at least 98.5%, at least 98.6%, at least 98.7%, at least 98.8%, at least 98.9%, at least 99.0%, at least 99.1 %, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% identical to a nucleic acid sequence according to SEQ ID NO: 5. In one embodiment, the recombinant cell comprises the enzyme E2 having the amino acid sequence according to SEQ ID NO: 2 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 2. In another embodiment, the recombinant cell comprises the enzyme E2 having an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 5 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 5.
[0085] In another embodiment, the recombinant cell comprises a fragment of the enzyme E2 having the amino acid sequence according to SEQ ID NO: 2 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 2.
[0086] Preferably, the functional fragment of the enzyme E2 has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80% identical, at least 81 %, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 98.5 %, at least 99%, or at least 99.5% of the length of the original full length amino acid sequence. Preferably, the functional fragment of the enzyme E2 comprises 140 to 282, 160 to 282, 180 to 282, 200 to 282, 220 to 282, 240 to 282, 260 to 282, or 270 to 282 consecutive amino acids of the full-length polypeptide. Also preferably, the functional fragment of the enzyme E2 retains at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80% identical, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 98.5 %, at least 99%, at least 99.5% or at least 100% of the enzyme activity of the original full length amino acid sequence. The functional fragment of the enzyme E2 comprises consecutive amino acids compared to the original full length amino acid sequence, respectively. In said embodiment, the "original full length amino acid sequence" of the enzyme E2 is an amino acid sequence which is at least 70% identical to the amino acid sequence according to SEQ ID NO: 2 and which has the same length as the amino acid sequence according to SEQ ID NO: 2.
[0087] In one embodiment, the fragment of the enzyme E2 does not comprise any internal deletions compared to the amino acid sequence according to SEQ ID NO: 2. However, as discussed above, the fragment of the enzyme E2 comprises a deletion of one or more amino acids at the N terminus and / or the C terminus of the polypeptide. In another embodiment, the recombinant cell comprises the enzyme E2 that is encoded by a fragment of a nucleic acid sequence according to SEQ ID NO: 5 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 5.
[0088] According to the disclosure, any the enzyme E2 comprising an amino acid sequence with less than 100% sequence identity according to SEQ ID NO: 2 has at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80% identical, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 98.5 %, at least 99%, at least 99.5% or at least 100% of the enzyme activity of the enzyme the enzyme E2 comprising an amino acid sequence with 100% sequence identity according to SEQ ID NO: 2. According to the disclosure, any the enzyme E2 encoded by a nucleic acid sequence with less than 100% sequence identity according to SEQ ID NO: 5 has at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80% identical, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 98.5 %, at least 99%, at least 99.5% or at least 100% of the enzyme activity of the enzyme E2 encoded by a nucleic acid sequence with 100% sequence identity according to SEQ ID NO: 5.
[0089] According to the invention, “the enzyme activity of enzyme E2” means that when enzyme E2 is expressed with the enzymes E1 comprising the amino acid sequence according to SEQ ID NO: 1 and E3 comprising the amino acid sequence according to SEQ ID NO: 3 in a recombinant host cell, a glycodilipid according to structure (II) is produced.
[0090] Enzyme E3
[0091] In one embodiment, the recombinant cell comprises a recombinant nucleic acid molecule encoding the enzyme E3. In one embodiment, the enzyme E3 has a phosphatase activity. A phosphatase is an enzyme that catalyzes the removal of a phosphate group from a molecule through hydrolysis in a process called dephosphorylation. Phosphatases act on different types of molecules, including proteins, lipids, and nucleotides, by removing phosphate groups.
[0092] In one embodiment, the amino acid sequence of the enzyme E3 is at least 70%, at least 75 %, at least 78 %, at least 80 %, at least 82 %, at least 84 %, at least 85 %, at least 86 %, at least 87 %, at least 88 %, at least 89 %, at least 90 %, at least 90.5%, at least 91 %, at least 91.5%, at least 92%, at least 92.5%, at least 93%, at least 93.5%, at least 94%, at least 94.5%, at least 95%, at least 95.1%, at least 95.2%, at least 95.3%, at least 95.4%, at least 95.5%, at least 95.6%, at least 95.7%, at least 95.8%, at least 95.9% at least 96%, at least 96.1 %, at least 96.2%, at least 96.3%, at least 96.4%, at least 96.5%, at least 96.6%, at least 96.7%, at least 96.8%, at least 96.9%, at least 97%, at least 97.1 %, at least 97.2%, at least 97.3%, at least 97.4%, at least 97.5%, at least 97.6%, at least 97.7%, at least 97.8%, at least 97.9%, at least 98.0%, at least 98.1%, at least 98.2%, at least 98.3%, at least 98.4%, at least 98.5%, at least 98.6%, at least 98.7%, at least 98.8%, at least 98.9%, at least 99.0%, at least 99.1 %, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% identical to an amino acid sequence according to SEQ ID NO: 3.
[0093] In one embodiment, the enzyme E3 has an amino acid sequence that is encoded by a nucleic acid sequence which is at least 70%, at least 75 %, at least 78 %, at least 80 %, at least 82 %, at least 84 %, at least 85 %, at least 86 %, at least 87 %, at least 88 %, at least 89 %, at least 90 %, at least 90.5%, at least 91%, at least 91.5%, at least 92%, at least 92.5%, at least 93%, at least 93.5%, at least 94%, at least 94.5%, at least 95%, at least 95.1 %, at least 95.2%, at least 95.3%, at least 95.4%, at least 95.5%, at least 95.6%, at least 95.7%, at least 95.8%, at least 95.9% at least 96%, at least 96.1%, at least 96.2%, at least 96.3%, at least 96.4%, at least 96.5%, at least 96.6%, at least 96.7%, at least 96.8%, at least 96.9%, at least 97%, at least 97.1%, at least 97.2%, at least 97.3%, at least 97.4%, at least 97.5%, at least 97.6%, at least 97.7%, at least 97.8%, at least 97.9%, at least 98.0%, at least 98.1 %, at least 98.2%, at least 98.3%, at least 98.4%, at least 98.5%, at least 98.6%, at least 98.7%, at least 98.8%, at least 98.9%, at least 99.0%, at least 99.1 %, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or 100% identical to a nucleic acid sequence according to SEQ ID NO: 6.
[0094] In one embodiment, the recombinant cell comprises the enzyme E3 having the amino acid sequence according to SEQ ID NO: 3 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 3. In another embodiment, the recombinant cell comprises the enzyme E3 having an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 6 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 6.
[0095] Preferably, the functional fragment of the enzyme E3 has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80% identical, at least 81 %, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 98.5 %, at least 99%, or at least 99.5% of the length of the original full length amino acid sequence. Preferably, the functional fragment of the enzyme E3 comprises 120 to 236, 140 to 236, 160 to 236, 180 to 236, 200 to 236, 220 to 236 or 230 to 236 consecutive amino acids of the full-length polypeptide. Also preferably, the functional fragment of the enzyme E3 retains at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80% identical, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 98.5 %, at least 99%, at least 99.5% or at least 100% of the enzyme activity ofthe original full length amino acid sequence. The functional fragment of the enzyme E3 comprises consecutive amino acids compared to the original full length amino acid sequence, respectively. In said embodiment, the "original full length amino acid sequence" of the enzyme E3is an amino acid sequence which is at least 70% identical to the amino acid sequence according to SEQ ID NO: 3 and which has the same length as the amino acid sequence according to SEQ ID NO: 3.
[0096] In one embodiment, the fragment of the enzyme E3 does not comprise any internal deletions compared to the amino acid sequence according to SEQ ID NO: 3. However, as discussed above, the fragment of the enzyme E3 comprises a deletion of one or more amino acids at the N terminus and / or the C terminus of the polypeptide. In another embodiment, the recombinant cell comprises a fragment of the enzyme E3 having the amino acid sequence according to SEQ ID NO: 3 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 3.
[0097] In another embodiment, the recombinant cell comprises the enzyme E3 that is encoded by a fragment of a nucleic acid sequence according to SEQ ID NO: 6 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 6.
[0098] According to the disclosure, any the enzyme E3 comprising an amino acid sequence with less than 100% sequence identity according to SEQ ID NO: 3 has at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80% identical, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 98.5 %, at least 99%, at least 99.5% or at least 100% of the enzyme activity of the enzyme E3 comprising an amino acid sequence with 100% sequence identity according to SEQ ID NO: 3. According to the disclosure, any the enzyme E3 encoded by a nucleic acid sequence with less than 100% sequence identity according to SEQ ID NO: 6 has at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80% identical, at least 81 %, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 98.5 %, at least 99%, at least 99.5% or at least 100% of the enzyme activity of the enzyme E3 encoded by a nucleic acid sequence with 100% sequence identity according to SEQ ID NO: 6.
[0099] According to the invention, “the enzyme activity of enzyme E3” means that when enzyme E3 is expressed with the enzymes E1 comprising the amino acid sequence according to SEQ ID NO: 1 and E2 comprising the amino acid sequence according to SEQ ID NO: 2 in a recombinant host cell, a glycodilipid according to structure (II) is produced. In one embodiment, the recombinant cell comprises one or more recombinant nucleic acid molecules encoding at least one enzyme selected from the group consisting of the enzymes E1 , E2 and E3.
[0100] In one embodiment, the recombinant cell comprises a recombinant nucleic acid molecule encoding the enzyme E1. In one embodiment, the recombinant cell comprises a recombinant nucleic acid molecule encoding the enzyme E2. In one embodiment, the recombinant cell comprises a recombinant nucleic acid molecule encoding the enzyme E3. In one embodiment, the recombinant cell comprises one or more recombinant nucleic acid molecules encoding the enzyme E1 and the enzyme E2. In one embodiment, the recombinant cell comprises one or more recombinant nucleic acid molecules encoding the enzyme E1 and the enzyme E3. In one embodiment, the recombinant cell comprises one or more recombinant nucleic acid molecules encoding the enzyme E2 and the enzyme E3. In one embodiment, the recombinant cell comprises one or more recombinant nucleic acid molecules encoding the enzymes E1 , E2 and E3.
[0101] The enzymes used in the method of the present invention can also be enzymes which share essentially the same three-dimensional structure than the enzymes E1 , E2 and / or E3. The three-dimensional structure can be determined using models such as ESM (evolutionary scale modeling), LLM (large language model) and ProteinBERT (see Brandes et al. (2022) Bioinformatics 38(8): 2102-2110).
[0102] In one embodiment, the recombinant cell comprises a recombinant nucleic acid molecule encoding a polypeptide having the activity of enzyme E1 , wherein:
[0103] (i) the polypeptide has an amino acid sequence according to SEQ ID NO: 1 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 1 ; or (ii) the polypeptide has an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 4 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 4;
[0104] (iii) the polypeptide is a fragment of (i) or (ii).
[0105] In one embodiment, the recombinant cell comprises a recombinant nucleic acid molecule encoding a polypeptide having the activity of enzyme E2, wherein: (i) the polypeptide has an amino acid sequence according to SEQ ID NO: 2 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 2; or (ii) the polypeptide has an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 5 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 5;
[0106] (Hi) the polypeptide is a fragment of (i) or (ii).
[0107] In one embodiment, the recombinant cell comprises a recombinant nucleic acid molecule encoding a polypeptide having the activity of enzyme E3, wherein:
[0108] (i) the polypeptide has an amino acid sequence according to SEQ ID NO: 3 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 3; or (ii) the polypeptide has an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 6 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 6;
[0109] (Hi) the polypeptide is a fragment of (i) or (ii).
[0110] In one embodiment, the recombinant cell comprises one or more recombinant nucleic acid molecules encoding a) a polypeptide having the activity of enzyme E1 , wherein:
[0111] (i) the polypeptide has an amino acid sequence according to SEQ ID NO: 1 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 1 ; or (ii) the polypeptide has an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 4 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 4;
[0112] (Hi) the polypeptide is a fragment of (i) or (ii); and
[0113] b) a polypeptide having the activity of enzyme E2, wherein:
[0114] (i) the polypeptide has an amino acid sequence according to SEQ ID NO: 2 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 2; or (ii) the polypeptide has an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 5 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 5;
[0115] (Hi) the polypeptide is a fragment of (i) or (ii).
[0116] In one embodiment, the recombinant cell comprises one or more recombinant nucleic acid molecules encoding a) a polypeptide having the activity of enzyme E1 , wherein:
[0117] (i) the polypeptide has an amino acid sequence according to SEQ ID NO: 1 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 1 ; or (ii) the polypeptide has an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 4 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 4;
[0118] (Hi) the polypeptide is a fragment of (i) or (ii); and
[0119] b) a polypeptide having the activity of enzyme E3, wherein:
[0120] (i) the polypeptide has an amino acid sequence according to SEQ ID NO: 3 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 3; or (ii) the polypeptide has an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 6 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 6;
[0121] (Hi) the polypeptide is a fragment of (i) or (ii).
[0122] In one embodiment, the recombinant cell comprises one or more recombinant nucleic acid molecules encoding a) a polypeptide having the activity of enzyme E2, wherein:
[0123] (i) the polypeptide has an amino acid sequence according to SEQ ID NO: 2 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 2; or (ii) the polypeptide has an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 5 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 5;
[0124] (Hi) the polypeptide is a fragment of (i) or (ii); and
[0125] b) a polypeptide having the activity of enzyme E3, wherein:
[0126] (i) the polypeptide has an amino acid sequence according to SEQ ID NO: 3 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 3; or (ii) the polypeptide has an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 6 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 6;
[0127] (Hi) the polypeptide is a fragment of (i) or (ii).
[0128] In one embodiment, the recombinant cell comprises one or more recombinant nucleic acid molecules encoding a) a polypeptide having the activity of enzyme E1 , wherein:
[0129] (i) the polypeptide has an amino acid sequence according to SEQ ID NO: 1 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 1 ; or (ii) the polypeptide has an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 4 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 4;
[0130] (Hi) the polypeptide is a fragment of (i) or (ii); and
[0131] b) a polypeptide having the activity of enzyme E2, wherein:
[0132] (i) the polypeptide has an amino acid sequence according to SEQ ID NO: 2 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 2; or (ii) the polypeptide has an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 5 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 5;
[0133] (Hi) the polypeptide is a fragment of (i) or (ii); and
[0134] c) a polypeptide having the activity of enzyme E3, wherein:
[0135] (i) the polypeptide has an amino acid sequence according to SEQ ID NO: 3 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 3; or (ii) the polypeptide has an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 6 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 6;
[0136] (Hi) the polypeptide is a fragment of (i) or (ii).
[0137] In one embodiment, the recombinant host cells comprise one or more recombinant nucleic acid molecules encoding the enzyme E1 and the enzyme E3 and the glycolipid comprises a glucose molecule linked to one fatty acid. In a preferred embodiment, the recombinant host cells comprise one or more recombinant nucleic acid molecules encoding the enzyme E1 and the enzyme E3 and the glycolipid comprises a glucose molecule linked to a dodecanoic acid molecule. More preferably, the recombinant host cells comprise one or more recombinant nucleic acid molecules encoding the enzyme E1 and the enzyme E3 and the glycolipid comprises a glucose molecule linked to a 3-hydroxy-5-dodecenoic acid molecule.
[0138] In one embodiment, the recombinant host cells comprise one or more recombinant nucleic acid molecules encoding the enzyme E1 and the enzyme E3 and the glycolipid comprises a glucose molecule linked to one fatty acid at position C2. In a preferred embodiment, the recombinant host cells comprise one or more recombinant nucleic acid molecules encoding the enzyme E1 and the enzyme E3 and the glycolipid comprises a glucose molecule linked to a dodecanoic acid molecule at position C2. More preferably, the recombinant host cells comprise one or more recombinant nucleic acid molecules encoding the enzyme E1 and the enzyme E3 and the glycolipid comprises a glucose molecule linked to a 3-hydroxy-5-dodecenoic acid molecule at position C2, which has the following structure (I):
[0139]
[0140] In said preferred embodiment, wherein the produced glycolipid comprises a glucose molecule linked to a 3-hydroxy- 5-dodecenoic acid molecule at position C2, the recombinant host cells comprise one or more recombinant nucleic acid molecules encoding the enzyme E1 having the amino acid sequence according to SEQ ID NO: 1 and / or that is encoded by a nucleic acid sequence according to SEQ ID NO: 4 and the enzyme E3 having the amino acid sequence according to SEQ ID NO: 3 and / or that is encoded by a nucleic acid sequence according to SEQ ID NO: 6.
[0141] In one embodiment, the recombinant host cells comprise one or more recombinant nucleic acid molecules encoding the enzymes E1 , E2 and E3 and the glycolipid comprises a glucose molecule linked to two fatty acids. In a preferred embodiment, the recombinant host cells comprise one or more recombinant nucleic acid molecules encoding the enzymes E1 , E2 and E3 and the glycolipid comprises a glucose molecule linked to a dodecanoic acid molecule and a decanoic acid molecule. More preferably, the recombinant host cells comprise one or more recombinant nucleic acid molecules encoding the enzymes E1 , E2 and E3 and the glycolipid comprises a glucose molecule linked to a 3-hydroxy-5-dodecenoic acid molecule and a 3-hydroxydecanoic acid molecule.
[0142] In one embodiment, the recombinant host cells comprise one or more recombinant nucleic acid molecules encoding the enzymes E1 , E2 and E3 and the glycolipid comprises a glucose molecule linked to a first fatty acid molecule at position C2 and a second fatty acid at position C3. In a preferred embodiment, the recombinant host cells comprise one or more recombinant nucleic acid molecules encoding the enzymes E1 , E2 and E3 and the glycolipid comprises a glucose molecule linked to a dodecanoic acid molecule at position C2 and a decanoic acid molecule at position C3. More preferably, the recombinant host cells comprise one or more recombinant nucleic acid molecules encoding the enzymes E1 , E2 and E3 and the glycolipid comprises a glucose molecule linked to a 3-hydroxy-5-dodecenoic acid molecule at position C2 and a 3-hydroxydecanoic acid molecule at position C3, which has the following structure (II):
[0143]
[0144] In said preferred embodiment, wherein the produced glycolipid comprises a glucose molecule linked to a 3-hydroxy-5-dodecenoic acid molecule at position C2 and a 3-hydroxydecanoic acid molecule at position C3, the recombinant host cells comprises one or more recombinant nucleic acid molecules encoding the enzyme E1 having the amino acid sequence according to SEQ ID NO: 1 and / or that is encoded by a nucleic acid sequence according to SEQ ID NO: 4, the enzyme E2 having the amino acid sequence according to SEQ ID NO: 2 and / or that is encoded by a nucleic acid sequence according to SEQ ID NO: 5 and the enzyme E3 having the amino acid sequence according to SEQ ID NO: 3 and / or that is encoded by a nucleic acid sequence according to SEQ ID NO: 6.
[0145] In one embodiment, the enzymes E1 , E2 and E3 can be encoded by a single recombinant nucleic acid molecule. In another embodiment, a first recombinant nucleic acid molecule encoding the enzyme E1 , a second recombinant nucleic acid molecule encoding the enzyme E2 and a third recombinant nucleic acid molecule encoding the enzyme E3 can be used to express the enzymes E1 , E2 and E3.
[0146] According to the disclosure, the glycolipids produced by the method of the invention are non-ionic. A non-ionic glycolipid is a type of glycolipid that does not carry a net electrical charge (neither positive nor negative) and its non-ionic nature comes from the absence of charged groups.
[0147] Culture methods
[0148] In the method of the present invention, the recombinant host cells are cultured. Culturing of recombinant host cells means herein that said cells are maintained in a cell culture medium under conditions suitable for maintaining viability and supporting growth of said cells. Culturing of recombinant cells can be initiated with one or more preculture steps. The term “preculture” refers to a culture with small volume which is used to grow cells for a main culture with a higher volume. The culture medium used for the preculture may be the same or may be different from the culture medium used for the main culture. The duration of the preculture is dependent on the concentration of cells suspended in the culture medium, which can be estimated by measurement of the optical density (OD) of the culture. The standard wavelength for measuring OD in cell cultures is 600 nm (QD600). According to the disclosure, the culturing of recombinant host cells comprises preferably two precultures. In one aspect of the present invention, the culture medium used in the first preculture differs from the culture medium of the main culture, whereas the culture medium of the second preculture is the same as used for the main culture. In one aspect of the present invention, the second preculture lasts until an OD600 of 0.5 is reached.
[0149] Cell culture systems can differ in regard to the process of adding required components to the cell culture and may include batch cell culture and fed-batch cell culture. The term batch cell culture refers to a system, in which all the necessary components for the cell culture are added at the beginning of the process and culturing is allowed to proceed without any further addition of materials, in particular without any further addition of cell culture medium or components thereof. The term fed-batch cell culture refers to a system, wherein initial components are added at the beginning of the process, and during the process, additional components are intermittently or continuously added to the cell culture without removing or replacing the culture medium. Initial components of a fed-batch cell culture comprise a carbon source such as glucose, salts, trace elements and essential minerals. According to the disclosure, the recombinant cells are preferably cultured in a fed-batch cell culture. Preferably, one additional component is a carbon source, more preferably the carbon source is glucose. Also preferably, one additional component is an essential mineral, more preferably the essential mineral is a magnesium source. Also preferably, one additional component being intermittently or continuously added to the cell culture may be a suitable inducer, which can activate the expression of a gene of interest which is placed under the control of an inducible promoter. Preferably, the additional components of the fed-batch culture comprise IPTG as inducer.
[0150] The skilled person will be aware of suitable cell culture media for supporting the growth of recombinant cells. The culture medium to be used must satisfy in a suitable manner the demands of the respective strains. In one embodiment, a suitable cell culture medium for culturing E. coli cells comprises a carbon source such as glucose and salts that supply nitrogen, phosphorus, and trace metals. In one embodiment, a suitable cell culture medium for culturing E. coli cells additionally comprises amino acids, nucleotide precursors, vitamins, and other metabolites. In one embodiment, a suitable cell culture medium for culturing E. coli cells comprises antibiotics such as ampicillin, kanamycin, or chloramphenicol for selection of transformed cells. According to the invention, the cell culture medium for culturing E. coli cells comprises ampicillin.
[0151] Typically, the cell culture medium comprises at least one carbon source. A carbon source in a cell culture medium is a compound that provides cells with the necessary carbon atoms to support their growth, metabolism, and energy production. In particular, the carbon source may be selected from carbohydrates such as, for example, glucose, sucrose, arabinose, xylose, lactose, fructose, maltose, molasses, starch, cellulose and hemicellulose, vegetable and animal oils and fats such as, for example, soybean oil, safflower oil, peanut oil, hempseed oil, jatropha oil, coconut fat, calabash oil, linseed oil, corn oil, poppyseed oil, evening primrose oil, olive oil, palm kernel oil, palm oil, rapeseed oil, sesame oil, sunflower oil, grapeseed oil, walnut oil, wheat germ oil and coconut oil, fatty acids, such as, for example, caprylic acid, capric acid, lauric acid, myristic acid, palmitic acid, palmitoleic acid, stearic acid, arachidonic acid, behenic acid, oleic acid, linoleic acid, linolenic acid, gamma-linolenic acid and its methyl or ethyl ester as well as fatty acid mixtures, mono-, di- and triglycerides containing any fatty acids mentioned above, alcohols such as, for example, glycerol, ethanol and methanol, hydrocarbons such as methane, carbon-containing gases and gas mixtures, such as CO, CO2, synthesis or flue gas, amino acids such as L-glutamate or L-valine or organic acids such as, for example, acetic acid. In a preferred embodiment, the recombinant host cells are cultured in a medium comprising glucose. In another preferred embodiment, the recombinant host cells are cultured in a medium comprising only one carbon source. More preferably, the recombinant host cells are cultured in a medium comprising glucose as sole carbon source. In a preferred embodiment, the recombinant host cells are cultured in a medium comprising glucose in the fed-batch mode. In another preferred embodiment, the recombinant host cells are cultured in a medium comprising only one carbon source in the fed-batch mode. More preferably, the recombinant host cells are cultured in a medium comprising glucose as sole carbon source in the fed-batch mode.
[0152] Isolation methods
[0153] In the method according to any aspect of the present invention, the glycolipids produced by the recombinant cell can be isolated from the cells and / or the medium after completion of the production process. Methods for isolation of low molecular weight substances from complex compositions are known in the art. In one embodiment, methods such as filtration, extraction, adsorption, chromatography, crystallization and the like may be used for isolation of the glycolipids. In a preferred embodiment, the glycolipid is isolated by using chromatography methods, more preferably by using Medium Pressure Liquid Chromatography (MPLC). For example, in said MPLC method, the cell culture is centrifuged to separate recombinant host cells and supernatant. Then, the supernatant is subjected to acidification using phosphoric acid and subsequently to liquid-liquid extraction using ethyl acetate. The solvent phase is placed in a separation funnel and evaporated to obtain the crude extract. The crude extract is dissolved in DMSO and purified using MPLC. The glycolipid containing fractions are pooled and evaporated to obtain the pure solid-form of the produced glycolipid. Industrial application
[0154] The glycolipids that can be produced according to the method of the present invention can advantageously be employed in cleaning or care agents that are used in housekeeping or industry, in cosmetic, dermatological or pharmaceutical formulations as well as in plant protection formulations, surfactant concentrates and the like. More specifically, they can be used in household and industrial cleaners, laundry detergents, shampoos and body washes, facial cleansers, emulsifiers, foaming agents, drug delivery systems, topical medications, pesticide and herbicide formulations, soil remediation, enhanced oil recovery, oil spill remediation, textile processing and leather tanning.
[0155] The term "care agents" is understood here as meaning a formulation that fulfills the purpose of maintaining an article in its original form, reducing or avoiding the effects of external influences (e.g. time, light, temperature, pressure, pollution, chemical reaction with other reactive compounds coming into contact with the article and the like) and aging, pollution, material fatigue, and / or even for improving desired positive properties of the article.
[0156] While the invention has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive. The invention is not limited to the disclosed embodiments. Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing a claimed invention, from a study of the drawings, the disclosure, and the dependent claims. The detailed description is merely exemplary in nature and is not intended to limit application and uses. The following examples further illustrate the present invention without, however, limiting the scope of the invention thereto. Various changes and modifications can be made by those skilled in the art on the basis of the description of the invention, and such changes and modifications are also included in the present invention.
[0157] 1. Material and Methods
[0158] 1.1 Media and cultivation procedures of E. coli (pCAT2 or pAFP1) cells
[0159] E. coli cells were transformed using either a pCAT2 or a pAFP1 vector. The pCAT2 vector comprises the nucleic acid sequence according to SEQ ID NO: 4 encoding the enzyme E1 , the nucleic acid sequence according to SEQ ID NO: 5 encoding the enzyme E2 and the nucleic acid sequence according to SEQ ID NO: 6 encoding the enzyme E3. For constructing the pCAT2 vector, PCR (polymerase chain reaction) was used to amplify genomic DNA comprising the nucleic acid sequences according to SEQ ID NO: 4, 5 and 6 with primers having the nucleic acid sequences according to SEQ ID NO: 7 and 8, and cloned into a pET21 a(+) vector. The pAFP1 vector comprises the nucleic acid sequence according to SEQ ID NO: 4 encoding the enzyme E1 and the nucleic acid sequence according to SEQ ID NO: 6 encoding the enzyme E3. For constructing the pAFP1 vector, PCR was used to amplify genomic DNA comprising the nucleic acid sequences according to SEQ ID NO: 4 and 6 with primers having the nucleic acid sequences according to SEQ ID NO: 9 and 10, and cloned into a pET21a(+) vector. Recombinant E. coli cells transformed with pCAT2 produce a glycodilipid comprising a glucose molecule linked to a 3-hydroxy-5-dodecenoic acid molecule at position C2 and a 3-hydroxydecanoic acid molecule at position C3. Recombinant E. coli cells transformed with pAFP1 produce a glycomonolipid comprising a glucose molecule linked to a 3-hydroxy-5-dodecenoic acid molecule at position C2.
[0160] For the first preculture, 25 pL glycerol stock of recombinant E. coli transformed with pCAT2 or pAFP1 was added into 25 mL Lysogeny Broth medium (10 g / L NaCI, 10 g / L tryptone and 5 g / L yeast extract) supplemented with ampicillin (100 pg / mL) in a 250 mL baffled conical flask. The first preculture was done for 12 hours in an incubator shaker at 120 rpm and 37 °C.
[0161] The second preculture was done in 200 mL modified Riesenberg’s medium (see Table 1 ) supplemented with ampicillin (100 pg / mL) in 2 L baffled conical flask. Culture broth from the first preculture was used for inoculation of the second preculture with a starting optical density (QD600) of 0.1.
[0162] For the fed-batch bioreactor cultivation, 9.6 L modified Riesenberg’s medium (see Table 1 ) supplemented with 25 g / L glucose as sole carbon source and ampicillin (100 pg / mL) was used at the beginning of the batch phase. Culture broth of the second preculture (-400 mL) was then pumped into a 40 L bioreactor for inoculation with a starting optical density (QD600) of -0.5. As the initial conditions, the bioreactor was set with the following parameters: temperature of 37°C; p©2 of 20%; pH of 7.0; agitation rate of 300 rpm; and gassing rate of 2 NL / min. For pH control, 20% ammonia solution and 4 M H3PO4 are used as base and acid, respectively.
[0163] The batch phase was left to grow overnight (-10 h) and then the fed-batch phase was started by injecting feed medium, comprising 0.1 mM IPTG, glucose as sole carbon source (500 g / L), MgSO4'7H2O as magnesium source (19.7 g / L) and ampicillin (100 pg / mL). The fed-batch phase was finished, after adding 5 L of said feed medium exponentially to the bioreactor.
[0164] Table 1 : Composition of modified Riesenberq’s Medium
[0165]
[0166] 1.2 Glucose quantification
[0167] Glucose quantification in the supernatant of the cell culture was done using an enzymatic assay kit following the protocol of the manufacturer (Enzytec™ Liquid D-Glucose, r-biopharm, Germany).
[0168] 1.3 Glycolipid quantification
[0169] The glycodilipid produced by E. coli pCAT2 was initially precipitated from 2 mL supernatant by the addition of 20 pL 85% v / v H3PO4. The acidification was followed by two-fold liquid-liquid extraction where 2.5 mL ethyl acetate was added to the acidified supernatant and mixed for 5 s. After centrifugation (3000 rpm, 15 min, 4°C, Heraeus Multifuge X3R), 1.5 mL of the upper phase were transferred into a 15 mL falcon tube, resulting in a final volume of 3 mL. To concentrate the glycodilipid, the sample was evaporated at 10 mbar and 40°C for 40 min using the rotavapor (RVC 2-25 CDplus, Martin Christ, Germany). The residue was resolved in 1 mL ethyl acetate for subsequent High-Performance Thin-Layer Chromatography (HPTLC) measurement.
[0170] The glycolipid quantification method using HPTLC was first developed and further validated based on Validation of Analytical Procedures: Methodology (FDA Guidance) with respect to the parameters linearity, limit of detection (LOD), limit of quantification (LOQ), precision, accuracy and repeatability according to Food and Drug Administration (FDA) Guidance (1999) (see also Geissler et al. (2017) J Chromatogr B Analyt Technol Biomed Life Sci 1044-1045: 214-224). The purified glycodilipid (95% purity) was dissolved in ethyl acetate then used as reference material at a concentration of 1 g / L. The calibration curves were developed by applying different volumes (5; 10; 15; 20; 25; 30; 35; 40 pL) of this reference material on the Thin-Layer Chromatography (TLC) plate then the glycodilipid concentration of the samples was calculated by considering 95% purity of the reference material. WinCATS Software 1.4.7 (CAMAG, Muttenz, Switzerland) was used to control all the HPTLC instruments (CAMAG, Muttenz, Switzerland). A sufficient amount of sample material was applied automatically by an Automatic TLC Sampler 4 (ATS 4) on 10 x 20 cm TLC silica gel 60 RP-18 F254S (Merck KGaA, Germany): 8 mm distance from the lower edge and 15 mm distance from the left edge. The application parameters were set as follows: filling speed 15 pL / s, dosage speed 150 nL / s, filling vacuum time 1.0 s, rinsing vacuum time 2.0 s, band length 6.0 mm. Methanol was used as the rinsing solvent. The development was performed using Automatic Developing Chamber (ADC 2) equipped with a 20 cm » 10 cm twin-trough chamber using 10 mL isopropyl acetate / methanol / acetic acid (100: 10: 1 , v / v / v) as the mobile phase until a migration distance of 70 mm was reached. A chamber saturation step for 5 min was attained to reach an equilibrium between the vapor and the solvent. The preconditioning step was done for 1 min by the usage of a filter paper soaked with 25 mL of the mobile phase. A final drying step was then performed for 5 min. Diphenylamine-Aniline-Phosphoric Acid (DPA) was used as derivatization reagent to visualize the glycodilipid on the TLC plate as DPA was found to be sensitive, convenient, and widely used for revealing glycoconjugates on TLC plate (see Buchan et al. (1952) Analyst 77: 401-406, Harris et al. (1954) Chem. Ind. p.
[0171] 249, Svennerholm et al. (1957) Biochim Biophys Acta 24(3):604-11 ). The DPA reagent was prepared by dissolving 2.4 g diphenylamine and 2.4 g aniline in 200 mL methanol then adding 20 mL 85% H3PO4. The developed plate was immersed in the derivatization reagent using the TLC Immersion Device at 3 cm / s speed for 3 s and the reaction was completed by incubating the plate on the TLC Plate Heater 3 at 120°C for 10 min. The derivatized plate was scanned with a scanning speed of 20 mm / s, a data resolution of 100 pm / step and a slit dimension of 3.0 x 0.30 mm at 620 nm by the TLC Scanner 4. For quantification of the glycodilipid, peak areas were evaluated using linear regression based on a calibration curve obtained from the reference material.
[0172] 1.4 Isolation and purification of glycolipids
[0173] The fermentation broth from the bioreactor cultivation was first subjected to centrifugation at 4700 rpm and 4°C for 15 min (Heraeus Multifuge X3R) to separate the supernatant from cell pellets. The cell-free supernatant was acidified with 1 % volume 85% H3PO4 (v / v) and subsequently extracted twice using 1.25 volumes of ethyl acetate (v / v) in a separation funnel. The extraction process was done by manually shaking the separation funnel vigorously for 5 min at room temperature. The organic phase was removed and concentrated using a rotary vacuum evaporator (R-215, Buchi Labortechnik AG, Switzerland) at 215 mbar and 40°C followed by further vacuum evaporation with a rotavapor (RVC 2-25 CDplus, Martin Christ, Germany) at 10 mbar, 1200 rpm, and 40°C for 6 h to gain the final crude extract.
[0174] The crude extract was dissolved in 3 mL dimethyl sulfoxide (DMSO) for purification with Medium Pressure Liquid Chromatography (MPLC; SepacoreX50, Buchi, Flawil, Switzerland) using a prepacked 40-60 pm particle size reverse phase C18 column (FlashPure EcoFlex C18; 40 g; column volume 80 mL, Buchi, Flawil, Switzerland). A 7.5 mL / min acetonitrile (ACN)Zwater as well as a 10 mL / min water / methanol gradient system was used as mobile phase (see Kugler et al. (2015) AMB Express 5(1 ):82). The eluate was collected in 10 mL fractions. All fractions were then analyzed by using HPTLC where silica gel 60 RP-18 F254S plate were derivatized with p-anisaldehyde / sulphuric acid / glacial acetic acid (1 :2: 100 v / v / v) reagent. Fractions containing the glycolipid were pooled and solvents were evaporated by using a rotary vacuum evaporator (R-215, Buchi Labortechnik AG, Switzerland) at 10 mbar and 40°C for structure elucidation.
[0175] 1.5 Nuclear Magnetic Resonance (NMR) analysis
[0176] For the structure elucidation of the glycolipids, 1 D and 2D NMR-spectra were recorded on an Avance HD III 600 MHz spectrometer, equipped with a 5 mm BBO Prodigy cryo-probe (Bruker, Billerica, United States). The sample was dissolved in 600 pl methanol-d4 and transferred to a standard 5 mm NMR tube. 1 H and 13C chemical shifts were referenced to the residual solvent signal at 5H / C 3.35 ppm / 49.0 ppm. 1 H, 13C, HSQC, HMBC, COSY, Heteronuclear Single Quantum Coherence Total Correlation Spectroscopy (HSQCTOCSY), F1 homoband decoupled HSQC, bandselective HSQC with and without decoupling, bandselective HMBC, H2BC (Heteronuclear 2-Bond Correlation) and selective 1 D-TOCSY spectra were recorded using standard Bruker pulse sequences at 298 K. A super long-range HMBC was measured by an in-house modifed Bruker pulse sequence (see Abdel-Mohsen et al. (2013) J Org Chem 78(16):7986-8003, Furihata K et al. (1995) Tetrahedron Lett 36(16):2817-20). Triple Spin Echo Pure Shift Yielded by Chirp Excitation (TSE-PSYCHE) and F1-homodecoupled PSYCHE TOCSY pulse sequence and parameters were obtained from the Manchester NMR methodology group (see Foroozandeh et al. (2014) J Am Chem Soc 136(34): 11867-9, Foroozandeh et al. (2015) Chem Commun 51 (84): 15410-3, Foroozandeh et al. (2014) Angew Chem Int Ed Engl 53(27):6990-2). The recorded NMR spectra were processed with Topspin 4.1.3 (copyright 2021 , Bruker Biospin, Billerica, MA, USA) and SpinWorks 4.2.10 (Copyright 2019, K. Marat, University of Manitoba, CA).
[0177] 1.6 Structure elucidation of glycolipids: Liguid Chromatography-Tandem Mass Spectrometry (LC-MS / MS)
[0178] The LC-MS / MS analysis of glycolipids was performed on a 1290 UHPLC system (Agilent, Waldbronn, Germany) coupled to a Q-Exactive Plus Orbitrap mass spectrometer equipped with a heated electrospray ionization source (HESI, Thermo Fisher Scientific, Bremen, Germany). The glycolipids were separated by a CSH column (2.1 pm x 150 mm, Waters, Eschborn, Germany). The temperature of the column was maintained at 40°C.
[0179] For the separation of the glycodilipid produced by E. coll pCAT2, samples were dissolved in methanol and 0.37 pL of each sample was injected. Mobile phase A was 0.2% formic acid in water, and mobile phase B was 0.2% formic acid in acetonitrile. A constant flow rate of 0.3 mL / min was used and the gradient elution was performed as follows: 45-53% B from 0 to 5 min, 53-59% B from 5 to 10 min, 59-90% B from 10 to 20 min, isocratic at 90% B from 20 to 25 min. Finally, the system was returned to initial conditions from 90% B to 45% B from 25 to 26 min.
[0180] For the separation of the glycomonolipid produced by E. coll pAFP1 , samples were dissolved in methanol and 2 pL of each sample was injected. Mobile phase A was 0.2% formic acid in water, and mobile phase B was 0.2% formic acid in methanol. A constant flow rate of 0.3 mL / min was used and the gradient elution was performed as follows: 10-25% B from 0 to 5 min, 25-45% B from 5 to 10 min, 45-90% B from 10 to 20 min, isocratic at 90% B from 20 to 25 min. Finally, the system was returned to initial conditions from 90% B to 10% B from 25 to 26 min and reequilibrated at 10% B from 26 to 32 min.
[0181] The HESI source was operated in positive and negative ion mode with a spray voltage of 4.0 kV in positive ion mode and 3.5 kV in negative ion mode. The ion transfer capillary temperature was set to 360°C and the sweep gas and auxiliary pressure rates were set to 60 and 20, respectively. The S-lens RF level was set to 50%. The Q-Exactive Plus mass spectrometer was calibrated externally in positive and negative ion mode using the manufacturers calibration solutions (Pierce, Thermo Fisher Scientific, Germany). Mass spectra were acquired within the mass range of 100 to 1400 m / z at a resolution of 70,000 FWHM using an Automatic Gain Control (AGC) target of 3.0 x 10E6 of and a maximum ion injection time of 100 ms. Data-dependent MS / MS spectra in the mass range of 50 to 2000 m / z were generated for the five most abundant precursor ions with a resolution of 17,500 FWHM using an AGC target of 1.0 x 10E6 and 100 ms maximum ion injection time and a normalized collision energy of 19. Xcalibur software version 4.3.73.11 and Compound Discoverer Software version 3.3 (both Thermo Fisher Scientific, San Jose, USA) were used for data acquisition and data analysis. Identification and assignment of the glycolipids were based on the precise m / z value of the precursor ion and manual inspection of the corresponding MS / MS spectra. Structure predictions from NMR were compared to MS / MS spectra using in silico fragmentation prediction provided by the Fragment Ion Search (FISh)ZMass Frontier algorithm in Compound Discoverer.
[0182] 1.7 Emulsification assay
[0183] Emulsification assay was done as indirect method to observe glycolipid production and performance as biosurfactant in bioreactor cultivation. The analyzed glycolipid was a glycodilipid isolated from Rouxiella badensis DSM 100043T, comprising a glucose molecule linked to a 3-hydroxy-5-dodecenoic acid molecule at position C2 and a 3-hydroxydecanoic acid molecule at position C3. Three milliliters of cell-free supernatant samples were mixed with 0.5 mL test oil and vortexed vigorously for 2 min then incubated at 37°C for 1 h. In this assay, both commercial olive oil (R Brandie GmbH, Germany) and low viscosity paraffin oil (Carl Roth, Germany) were used as test oils. The aqueous phase was transferred into cuvette and the absorbance was recorded at 400 nm using UV-visible spectrophotometer (GENESYS 150, Thermo scientific). The blank was prepared similarly with sterile MSM for bioreactor cultivation. At 400 nm, an absorbance of 0.010 units multiplied by any applicable dilution factor is considered equivalent to one unit of emulsification activity per milliliter (EU / mL) (see Patil et al. (2001 ) J Appl Microbiol 91 (2):290-8).
[0184] 1.8 Critical Micelle Concentration (CMC) determination
[0185] The CMC of a glycodilipid isolated from Rouxiella badensis DSM 100043T, comprising a glucose molecule linked to a 3-hydroxy-5-dodecenoic acid molecule at position C2 and a 3-hydroxydecanoic acid molecule at position C3, was determined by measuring the surface tension of air-water surface at 25°C by using a DCAT 11 tensiometer equipped with a Wilhelmy plate (DataPhysics GmbH, Filderstadt, Germany). The glass beaker and Wilhelmy plate were cleaned with purified water and ethanol and then rinsed thoroughly with purified water. The Wilhelmy plate was further heated to a light red glow with a Bunsen burner to remove any contaminants. The stock aqueous glycodilipid solution (800 mg / L, M = 546 g mol1) was titrated manually into purified water in a glass beaker in a series of increasing volumes. After each titration, the sample solution was stirred at 30% stirring rate for 60 s and equilibrated for 10 s before measuring the surface tension. The surface tension of the purified water was 72.049 ± 0.012 mN nr1. The density of the glycodilipid solution was 0.99780 kg rm3as determined by using a DMA 35N density meter (Anton Paar GmbH, Graz, Austria).
[0186] 1.9 Stability test
[0187] The stability of a glycodilipid isolated from Rouxiella badensis DSM 100043T, comprising a glucose molecule linked to a 3-hydroxy-5-dodecenoic acid molecule at position C2 and a 3-hydroxydecanoic acid molecule at position C3, was evaluated under a wide range of temperature, salinity, and pH conditions above its CMC. A modified method from Samykannu et al. ((2017) Appl Biochem Biotechnol 183(1 ):70-90), ) was performed using 2 mL of 50 mg / L glycodilipid dissolved in ultrapure water. To analyze the temperature stability, the glycodilipid solution was incubated at 0, 20, 40, 60, 80, and 100°C for 1 h. To analyze the stability in different salinities, the glycodilipid solution was incubated at 0, 3, 6, 9, 12, and 15% (w / v) NaCI. To examine the pH stability of the glycodilipid biosurfactant, the pH of the glycodilipid solution was adjusted to different pH values ranging from 2 to 12 using 6 N HCI and 6 N NaOH. All samples were then subjected to an emulsification activity test (E24) according to Hamzah et al. ((2020) J Pet Explor Prod Technol. 10(8):3767-77), where 2 mL of olive oil were added into incubated samples, vortexed for 5 min, and again incubated for 24 h. The E24 index was obtained by calculating percentage of the height of the emulsified layer divided by the total height of the liquid column (see Cooper et al. (1987) Appl Environ Microbiol 53(2):224-29).
[0188] 2. Results
[0189] 2.1 LC-MS / MS analysis of the qlycodilipid produced by E. coli pCAT2
[0190] The purified glycodilipid sample produced by E. coli pCAT2 was analyzed by high-resolution mass spectrometry. The base peak chromatogram of the LC-MS / MS analysis in negative ion mode showed an intense signal at 13.26 min (see Figure 1 ).
[0191] 2.2 LC-MS / MS analysis the qlycomonolipid produced by E. coli pAFP1
[0192] The purified glycomonolipid sample produced by E. coli pAFP1 was analyzed by high-resolution mass spectrometry. The base peak chromatogram of the LC-MS / MS analysis in negative ion mode showed an intense signal at 17.02 min (Figure 2).
[0193] 2.3 Fed-batch bioreactor cultivations of E. coli pCAT2
[0194] Fed-batch bioreactor cultivations of E. coli pCAT2 were conducted over the time course of 25 hours (Figure 3). Using quantitative high-performance thin-layer chromatography (HPTLC) measurement, the glycodilipid could be detected and measured. The method for glycolipid quantification by HPTLC has been validated according to US FDA Guidance (1999). The batch phase was conducted for 10 hours, at the end of which glucose was completely consumed. Then the fed-batch phase was started by injecting feed medium comprising IPTG and glucose. The feed medium was pumped into the bioreactor at the same rate as the glucose consumption rate of the E. coli pCAT2 cell culture. The highest glycodilipid titer as well as the highest cell dry weight was measured after 25 hours of cultivation.
[0195] 2.4 Surface properties of various glycolipids
[0196] An isolated glycodilipid, comprising a glucose molecule linked to a 3-hydroxy-5-dodecenoic acid molecule at position C2 and a 3-hydroxydecanoic acid molecule at position C3, produced by R. badensis DSM 100043T shows excellent surface activity property indicated by a very low CMC value compared to other microbial biosurfactants as shown in Table 2. This outstanding property might be explained by the structure of the glycodilipid, where the polar glucose molecule acts as hydrophilic moiety while the two non-polar fatty acid tails act as hydrophobic moiety. Typically, other microbial biosurfactants have only a single fatty acid tail contributing to their hydrophobic moiety, resulting in lower hydrophobicity compared to the glycodilipid produced by R. badensis DSM 100043T.
[0197] Table 2: Comparison of surface tension and CMC values of different microbial biosurfactants
[0198]
[0199]
[0200] 2.5 Stability test of 50 mq / L qlycodilipid biosurfactant from R. badensis DSM 100043TThe stability at different range of temperatures, pH, and salinity of an isolated glycodilipid, comprising a glucose molecule linked to a 3-hydroxy-5-dodecenoic acid molecule at position C2 and a 3-hydroxydecanoic acid molecule at position C3, produced by R. badensis DSM 100043Twas evaluated using an emulsification test (%El24> in olive oil (Figure 4). In the emulsification test, 50 mg / L of pure glycodilipid was used, which is above its CMC value, to maintain emulsification performance of the glycodilipid. In general, the glycodilipid showed good emulsification stability against changes in slightly acidic pH, temperature, and saline concentration as shown in Figure 4.
[0201] Figure 4A shows the effect of different temperatures on the emulsification index of the glycodilipid in olive oil. The glycodilipid demonstrated thermostability, maintaining its stability at temperatures up to 100°C. The thermal stability of the glycodilipid suggests its potential for use across various industries that operate under high temperatures, as well as its suitability for enhanced oil recovery (see Samykannuet et al. (2017) Appl Biochem Biotechnol 183(1 ):70-90).
[0202] The glycodilipid was also found to be stable over a wide range of NaCI concentrations as shown in Figure 4B. The stability of the glycodilipid in high salinity conditions demonstrates its potential for the treatment of oil spills in the marine environment (see Nikolova et al. (2021 ) Front Bioeng Biotechnol 9:626639).
[0203] Regarding pH stability, the results showed that pH plays a significant role in the chemical stability of the glycodilipid as %El24 changed at different pH values as shown in Figure 4C. The glycodilipid was found to be more stable at acidic pH rather that in basic environment where the highest emulsification index was observed at pH 4. This phenomenon contrasts with that observed in rhamnolipids, where acidic pH leads to the precipitation of rhamnolipid, thereby reducing its stability (see Dabaghi et al. (2023) BMC Biotechnol 23: 2). The glycodilipid of R. badensis DSM 100043Tdoes not have an acidic group that could be protonated, meaning that low pH value should not play a major role in the precipitation, and remained stable at low pH levels, making it suitable for cosmetic applications where effectiveness is expected on the slightly acidic pH of human skin.
[0204] The glycodilipid produced by R. badensis DSM 100043Tshowed increased stability and surface properties compared to other microbial biosurfactants having only a single fatty acid tail contributing to their hydrophobic moiety, resulting in lower hydrophobicity. Hence, the advantageous and outstanding properties of said glycodilipid are due to the two non-polar fatty acid tails, a characteristic that also applies for the glycolipid produced by E. coli pCAT2.
Claims
Claims1 . A method for the production of a glycolipid comprising a glucose molecule linked to at least one fatty acid molecule, comprising culturing recombinant host cells under suitable conditions for the production of said glycolipid, wherein the recombinant host cells comprise one or more recombinant nucleic acid molecules encoding at least one enzyme selected from the group consisting of E1 , E2 and E3, wherein:(a) the enzyme E1 comprises:(i) a polypeptide having an amino acid sequence according to SEQ ID NO: 1or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 1 ; or (ii) a polypeptide having an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 4 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 4;(Hi) a fragment of the polypeptide of (i) or (ii);(b) the enzyme E2 comprises:(i) a polypeptide having an amino acid sequence according to SEQ ID NO: 2 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 2; or(ii) a polypeptide having an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 5 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 5;(Hi) a fragment of the polypeptide of (i) or (ii), and(c) the enzyme E3 comprises:(i) a polypeptide having an amino acid sequence according to SEQ ID NO: 3 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 3; or(ii) a polypeptide having an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 6 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 6;(Hi) a fragment of the polypeptide of (i) or (ii).
2. The method according to claim 1 , wherein the recombinant host cells are recombinant Escherichia coli cells.
3. The method according to claim 1 or 2, wherein the recombinant host cells are cultured in a culture medium containing glucose as sole carbon source.
4. The method according to any of claims 1 to 3, wherein the recombinant host cells are cultured in fed-batch mode.
5. The method according to any of claims 1 to 4, further comprising the step of isolating the glycolipid.
6. The method according to any of claims 1 to 5, wherein the recombinant host cells comprise one or more recombinant nucleic acid molecules encoding the three enzymes E1 , E2 and E3 and produce a glycolipid comprising a glucose molecule linked to a first and a second fatty acid molecule.
7. The method according to claim 6, wherein the first fatty acid molecule has 12 carbon atoms (C12) and is a 3-hydroxy-5-dodecenoic acid.
8. The method according to claim 6 or 7, wherein the second fatty acid molecule has 10 carbon atoms (C10) and is a 3-hydroxydecanoic acid.
9. The method according to any of claims 6 to 8, wherein the two fatty acid molecules are linked to two different carbon atoms of the glucose molecule.
10. The method according to any of claims 6 to 9, wherein no fatty acid molecule is linked to C1 of the glucose molecule.11 . The method according to any of claims 8 to 10, wherein the 3-hydroxydecanoic acid is linked to C3 of the glucose molecule.
12. The method according to any of claims 7 to 11 , wherein the 3-hydroxy-5-dodecenoic acid is linked to C2 of the glucose molecule.
13. The method according to any of claims 6 to 12, wherein the fatty acid molecules are linked to the glucose molecule by ester bonds.
14. The method according to claim 13, wherein the ester bonds are formed between the carboxyl group of the fatty acid molecules and the hydroxyl group of the glucose molecule.
15. The method according to any of claims 6 to 14, wherein the glycolipid is non-ionic.
16. The method according to any of claims 6 to 15, having the following structure (II):
17. An isolated glycolipid comprising a glucose molecule linked to one fatty acid molecule, wherein the fatty acid molecule has 12 carbon atoms (C12).
18. The glycolipid according to claim 17, wherein the fatty acid molecule having 12 carbon atoms(Ci2) is a 3-hydroxy-5-dodecenoic acid.
19. The glycolipid according to claim 18, wherein the 3-hydroxy-5-dodecenoic acid is linked to C2 of the glucose molecule.
20. The glycolipid according to any of claims 17 to 19, wherein the fatty acid molecule is linked to the glucose molecule by an ester bond.
21. The glycolipid according to claim 20, wherein the ester bond is formed between the carboxyl group of the fatty acid molecule and the hydroxyl group of the glucose molecule.
22. The glycolipid according to any of claims 17 to 21 , wherein no further fatty acid molecule is linked to the glucose molecule.
23. The glycolipid according to any of claims 17 to 22, wherein the glycolipid is non-ionic.
24. The glycolipid according to any of claims 17 to 23, having the following structure (I):
25. The method according to any of claims 1 to 5, wherein the recombinant host cells comprise one or more recombinant nucleic acid molecules encoding the two enzymes E1 and E3 and the produced glycolipid is the glycolipid of any of claims 17 to 24.
26. An isolated nucleic acid molecule encoding a polypeptide having the activity of enzyme E1 , wherein: (i) the polypeptide has an amino acid sequence according to SEQ ID NO: 1 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 1 ; or(ii) the polypeptide has an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 4 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 4;(Hi) the polypeptide is a fragment of (i) or (ii),wherein the enzyme activity of enzyme E1 means that when enzyme E1 is expressed with the enzymes E2 comprising the amino acid sequence according to SEQ ID NO: 2 and E3 comprising the amino acid sequence according to SEQ ID NO: 3 in a recombinant host cell, a glycodilipid according to structure (II) is produced.
27. An isolated nucleic acid molecule encoding a polypeptide having the activity of enzyme E2, wherein: (i) the polypeptide has an amino acid sequence according to SEQ ID NO: 2 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 2; or(ii) the polypeptide has an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 5 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 5; or(Hi) the polypeptide is a fragment of (i) or (ii),wherein the enzyme activity of enzyme E2 means that when enzyme E2 is expressed with the enzymes E1 comprising the amino acid sequence according to SEQ ID NO: 1 and E3 comprising the amino acid sequence according to SEQ ID NO: 3 in a recombinant host cell, a glycodilipid according to structure (II) is produced.
28. An isolated nucleic acid molecule encoding a polypeptide having the activity of enzyme E3, wherein: (i) the polypeptide has an amino acid sequence according to SEQ ID NO: 3 or an amino acid sequence with at least 70% identity to the amino acid sequence according to SEQ ID NO: 3; or(ii) the polypeptide has an amino acid sequence that is encoded by a nucleic acid sequence according to SEQ ID NO: 6 or a nucleic acid sequence having at least 70% identity to the nucleic acid sequence according to SEQ ID NO: 6; (Hi) the polypeptide is a fragment of (i) or (ii),wherein the enzyme activity of enzyme E3 means that when enzyme E3 is expressed with the enzymes E1 comprising the amino acid sequence according to SEQ ID NO: 1 and E2 comprising the amino acid sequence according to SEQ ID NO: 2 in a recombinant host cell, a glycodilipid according to structure (II) is produced.
29. A recombinant nucleic acid molecule comprising one or more of the isolated nucleic acid molecules according to any one of claims 26 to 28.
30. A recombinant host cell comprising one or more recombinant nucleic acid molecules according to claim 29.31 . An isolated polypeptide encoded by the nucleic acid molecule according to any one of claims 26 to 28.
32. A composition comprising the glycolipid of any one of claims 17 to 23.
33. A household, pharmaceutical, medical, food, agricultural or cosmetic composition comprising the glycolipid of any one of claims 17 to 23.
34. Use of the glycolipid of any one of claims 17 to 23 for environmental remediation or in the petroleum industry.