Novel antibiotic compounds
By screening and identifying new biosynthetic gene clusters in marine actinomycete strains, the new antibiotic compound nidaromycin was successfully discovered and prepared, which solved the problem of antibiotic resistance in the prior art and achieved effective antibacterial effects on multidrug-resistant bacteria.
Patent Information
- Application Number
- CN202380054765.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-05-17
- Filing Date
- 2023-05-16
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to effectively solve the problem of antibiotic resistance of multiple bacterial pathogens, and the discovery of new antimicrobial agents is slowed down.
By screening and identifying new biosynthetic gene clusters in marine actinomycetes strains, gene clusters encoding the new antibiotic compound nidaromycin were found and cloned, and the gene cluster was expressed in a heterologous host to produce nidaromycin.
The new antibiotic compound nidaromycin was successfully discovered and prepared, showing significant antibacterial activity against multidrug-resistant bacteria and was nontoxic to mammalian cells.
Smart Images

Figure BDA0005247991670000021 
Figure BDA0005247991670000071 
Figure BDA0005247991670000081
Abstract
Description
Technical Field
[0001] The present disclosure and the present invention relate to a new antibiotic compound (which we will refer to as nidaromycin), its uses, and its biosynthesis. In particular, a new biosynthetic gene cluster (BGC) has been identified, sequenced, and cloned such that the BGC is introduced into and expressed in a host cell to produce the compound. Accordingly, the present disclosure and the present invention also relate to new genes and nucleic acid molecules encoding the biosynthetic machinery for producing nidaromycin, constructs, vectors, and host cells for expressing the BGC, and methods for producing the compound. Background Art
[0002] Natural products produced by bacteria and fungi are very important because of their potential use as drugs or veterinary products, most notably as antibiotics. Actinomycetes are a class of filamentous Gram-positive bacteria with a high GC content that produce the vast majority of all known antibiotics of microbial origin, approximately half of which are obtained from the genus Streptomyces.
[0003] Genes in actinomycetes for synthesizing secondary metabolites such as antibiotics tend to be organized in clusters, including genes encoding biosynthetic enzymes, transporters, and other proteins involved in their synthesis or regulation. Various gene clusters for synthesizing multiple antibiotics in different organisms have been reported in the literature.
[0004] The increasing antibiotic resistance of multiple bacterial pathogens is a major health concern globally, and at the same time, the rate of discovery of new antimicrobials is continuously slowing down. This has promoted the development of new strategies for antibiotic discovery, one of the new methods being genome mining of actinomycetes to identify novel secondary metabolite gene clusters of unknown function that may potentially encode genes for the biosynthesis of clinically useful antibiotics. For this purpose, there are various strategies, including attempts to induce the expression of cryptic gene clusters in natural hosts by culturing under different conditions, or heterologous expression of gene clusters in alternative hosts. We have adopted the latter approach, combined with bioinformatics and phylogenetic analysis, to identify a new BGC encoding the biosynthesis of a new antibiotic compound. Summary of the Invention
[0005] For this purpose, based on various criteria, including phylogenetic novelty, gene cluster diversity, and the antibacterial activity of the strains, 1200 marine actinomycete strains isolated from the Trondheim Fjord and 576 actinomycete type strains retrieved from public databases were selected and their genomes were analyzed. Based on this, various strains were selected for further study, including the actinomycete strain P08-G05. The genome of this strain was sequenced and bioinformatics analysis was performed to identify and analyze potential new BGCs containing genes encoding the synthesis of antibiotic-like compounds. A new BGC was thus identified, which is referred to herein as P08-G05-cluster 16 (P08-G05-c16), and its DNA sequence is shown in SEQ ID NO.1. This cluster has been cloned and expressed in a heterologous actinomycete host, specifically in the Streptomyces coelicolor M1152ΔmatAB strain, which is a modified derivative of the model strain Streptomyces coelicolor A3(2) / M145 (ATCC BAA-471), and its preparation process is described in the following examples. The P08-G05-c16 gene cluster was cloned into an inducible bacterial artificial chromosome (BAC) vector and transferred into the Streptomyces coelicolor M1152ΔmatAB strain by triparental conjugation to prepare the transconjugant strain M1152ΔmatAB(P08-G05_C16). The transconjugant expresses this BGC and synthesizes compounds; the extracts have been shown to have antibacterial activity. We refer to this new compound as nidaromycin, which has currently been extracted, purified, structurally analyzed, and characterized from the heterologous host, and the results have confirmed its structural novelty, antibiotic activity, and low cytotoxicity.
[0006] Accordingly, in a first aspect, the present invention provides a compound of formula (I):
[0007]
[0008] wherein R 1 is -SO 2 OH, -SO 2 OR or -SO 2 R and R 2 is H, or wherein R 2 is
[0009] -SO 2 OH, -SO 2 OR or -SO 2 R and R 1 is H;
[0010] wherein R is C 1 -C 20 hydrocarbyl; and
[0011] wherein each R 3 is independently selected from H or C 1 -C 20 hydrocarbyl group;
[0012] or a pharmaceutically acceptable salt, solvate, hydrate or ester thereof.
[0013] In a second aspect, the present invention provides a compound of formula I as defined herein for use as a medicament, or in other words, for treatment.
[0014] In particular, the medicament is an antibiotic and the treatment is an antimicrobial treatment, particularly an antibacterial treatment.
[0015] Accordingly, a third aspect provides a compound of formula I as defined herein for use as an antibiotic medicament, or for treating a microbial infection, particularly a bacterial infection.
[0016] A fourth aspect provides the use of a compound of formula I as defined herein in the preparation of a medicament for treating a microbial infection, particularly a bacterial infection.
[0017] A fifth aspect provides a method of treating a microbial infection, particularly a bacterial infection, in a subject, the method comprising administering to the subject an effective amount of a compound of formula I as defined herein.
[0018] The subject can be any human or non-human animal, particularly a mammal. Accordingly, the medical uses herein include human clinical and veterinary uses, as well as uses in animal husbandry and agriculture, including uses in aquaculture and as a plant protection agent against plant pathogens.
[0019] In one embodiment, the bacterial infection is a Gram-positive bacterial infection.
[0020] A sixth aspect provides a pharmaceutical composition comprising a compound of formula I as defined herein, and at least one pharmaceutically acceptable carrier, additive and / or excipient. In one embodiment, the pharmaceutical composition is suitable for parenteral, oral or topical administration.
[0021] In addition to the medical uses outlined above, the compound may also have non-medical (i.e., non-therapeutic) uses, taking advantage of its antimicrobial / antibacterial properties in vitro or ex vivo, such as for decontaminating, disinfecting or sterilizing surfaces and the like.
[0022] Accordingly, a seventh aspect provides the use of a compound of formula I as defined herein as an antimicrobial agent, particularly an antibacterial agent.
[0023] In other words, this aspect also provides a method for controlling bacteria on a surface, the method comprising the step of applying a compound as defined herein to the surface (or bringing the surface into contact with the compound). Controlling bacteria includes inhibiting the growth and / or viability of bacteria. This may also include reducing the number of bacteria (e.g., killing bacteria), and / or reducing or preventing their proliferation (replication).
[0024] The compound can be prepared by biosynthesis in a host that has been modified or engineered to express a BCG containing the biosynthetic genes for its synthesis, or in other words, a BGC, or more specifically a nucleic acid molecule containing a BGC or its constituent genes, has been introduced into the host. In one embodiment, the host is a heterologous host, i.e., a host that does not naturally contain a BGC. However, in another embodiment, the BGC can be introduced into the organism from which it was obtained, i.e., isolate P08 - G05, or more generally a strain that endogenously contains a BGC. Alternatively, the compound can be prepared in an in vitro transcription and translation (IVTT) system.
[0025] Accordingly, the eighth aspect provides a nucleic acid molecule comprising:
[0026] (a) the nucleotide sequence shown in SEQ ID NO.1; or
[0027] (b) a nucleotide sequence complementary to SEQ ID NO.1; or
[0028] (c) a nucleotide sequence degenerate to SEQ ID NO.1; or
[0029] (d) a nucleotide sequence having at least 85% sequence identity with SEQ ID NO.1; or
[0030] (e) a portion of any one of (a) to (d),
[0031] wherein the nucleic acid molecule encodes one or more polypeptides or is complementary to a nucleic acid molecule encoding the one or more polypeptides, or comprises one or more genetic elements or is complementary to a nucleic acid molecule containing one or more genetic elements, and has functional activity in the synthesis of an antibiotic compound.
[0032] The functional activity can be enzyme activity, or transport or transfer activity, or regulatory activity (e.g., regulation of gene expression), or any other activity that contributes to the synthesis or transport of the compound or its components.
[0033] Accordingly, more generally, the nucleic acid molecule can be defined as comprising one or more nucleotide sequences that contribute to the biosynthesis of the compound. In other words, the nucleic acid molecule can comprise one or more nucleotide sequences that constitute or are part of a biosynthetic gene cluster for the synthesis of the compound.
[0034] In particular, the antibiotic compound is a compound of formula I or a derivative thereof as defined herein.
[0035] In particular, in part (e), the portion of the nucleotide sequence comprises a sequence corresponding to a biosynthetic gene or an open reading frame (ORF) encoding a protein involved in the biosynthesis of the compound, or is complementary or degenerate to such a sequence.
[0036] In one embodiment, the nucleic acid molecule comprises the nucleotide sequences (a) to (d) and encodes a polypeptide for the synthesis of an antibiotic compound, in particular for the synthesis of a compound of formula I or a derivative thereof as defined herein. In other words, the nucleic acid molecule comprises nucleotide sequences that together provide the biosynthetic machinery for producing the compound. Thus, in this embodiment, the nucleic acid molecule can be defined as encoding a biosynthetic system for synthesizing the compound, or as comprising a BGC for synthesizing the compound.
[0037] In one embodiment, the nucleic acid molecule comprises a nucleotide sequence for producing the compound in an actinomycete host, specifically a Streptomyces host, more specifically a Streptomyces coelicolor host, especially in a host strain of Streptomyces coelicolor strain A3(2), M145, M1152 or M1152ΔmatAB, as defined or described herein.
[0038] As described in part (e), the nucleic acid molecule can also encode a part of a complete biosynthetic system, such as a single component of a BGC. Individual genes or ORFs in a BGC have been identified and annotated, as described in more detail below (see Table 1). Thus, a portion of the nucleic acid molecule can represent or correspond to, for example, a single gene or ORF encoding a polypeptide involved in the biosynthesis of a compound, or two or more such genes or ORFs. Thus, in one embodiment, the nucleic acid molecule comprises the nucleotide sequence shown in any one or more of SEQ ID NOs: 2-29, or a nucleotide sequence complementary or degenerate thereto, or a nucleotide sequence having at least 85% sequence identity thereto.
[0039] In another embodiment, the nucleic acid molecule comprises a nucleotide sequence encoding the amino acid sequence shown in any one or more of SEQ ID NOs: 30-57, or a nucleotide sequence encoding an amino acid having at least 85% sequence identity thereto.
[0040] The ninth aspect provides a polypeptide encoded by the nucleic acid molecule as defined above.
[0041] The tenth aspect provides a recombinant construct comprising the nucleic acid molecule as defined herein.
[0042] In one embodiment, the recombinant construct comprises one or more other nucleic acid sequences, such as regulatory sequences, expression control sequences, or genetic elements involved in the replication or transfer of the nucleic acid molecule.
[0043] The eleventh aspect provides a vector comprising a nucleic acid molecule or a recombinant construct as defined herein.
[0044] In one embodiment, the vector is a plasmid, cosmid, or artificial chromosome, particularly a bacterial artificial chromosome (BAC).
[0045] The twelfth aspect provides a microbial host cell comprising a nucleic acid molecule, a recombinant construct, or a vector as defined herein. In other words, this aspect provides a modified or engineered microbial host cell into which a nucleic acid molecule, a recombinant construct, or a vector has been introduced. As described above, the host can be a heterologous host. By definition, a heterologous host or a modified host cell does not include the natural producer of the compound or the natural strain that endogenously contains the BGC. However, as also described above, introduction of the nucleic acid molecule, recombinant construct, or vector into the original strain from which the nucleic acid molecule is derived is not excluded. The nucleic acid molecule can be introduced into the host cell in increased copy numbers, i.e., one or more copies can be introduced.
[0046] The host cell can be a production host cell for the production of a compound, or it can be a host cell generated for the purpose of cloning a nucleic acid molecule (e.g., propagating or producing a nucleic acid molecule) or for transferring it to another host cell (i.e., it can be a cloning or transfer host cell).
[0047] The thirteenth aspect provides a method for producing a compound of formula I as defined herein, the method comprising introducing a nucleic acid molecule, a recombinant construct, or a vector as defined herein into a microbial host cell and allowing the nucleic acid molecule to be expressed (particularly allowing the individual gene sequences or ORFs of the nucleic acid molecule to be expressed, i.e., allowing the genes of the BGC to be expressed).
[0048] As further defined, this aspect provides a method for producing a compound of formula I as defined herein, the method comprising introducing a nucleic acid molecule, a recombinant construct, or a vector as defined herein into a microbial host cell and culturing (or growing) the host cell under conditions that express a biosynthetic system for synthesizing the compound.
[0049] More specifically, the conditions are those that allow the compound to be synthesized by the expressed biosynthetic system.
[0050] The production method may further comprise the step of recovering the compound (or in other words, collecting or harvesting the compound). In addition, the method may comprise the steps of separating, isolating, or purifying the compound.
[0051] In one embodiment, the host cell is a bacterial host cell, such as from Actinomycetes. In a specific embodiment, the host cell is an Actinomycete, specifically a Streptomyces host cell, more specifically Streptomyces coelicolor, especially Streptomyces coelicolor strain A3(2), M145, M1152 or M1152ΔmatAB as defined or described herein.
[0052] The fourteenth aspect provides a compound obtainable or able to be obtained by a production method as defined herein. In particular, the compound is obtainable or able to be obtained by a method comprising expressing the nucleic acid molecule in a heterologous Actinomycete host cell, which Actinomycete host cell is specifically a Streptomyces host cell, more specifically Streptomyces coelicolor, especially Streptomyces coelicolor strain A3(2), M145, M1152 or M1152ΔmatAB as defined or described herein.
[0053] The compound obtainable or able to be obtained by this method can be used for any of the uses or methods listed above, or can be comprised in a pharmaceutical composition. In other words, in any of the aspects listed above, the compound can alternatively or additionally be defined as a compound obtainable or able to be obtained by a production method as defined herein.
[0054] Furthermore, the production method herein allows the biosynthetic system used for synthesis to be modified to modify the compound produced. Thus, the nucleic acid molecule encoding the biosynthetic enzyme or the individual nucleotide sequences comprised therein can be modified, or inactivated or deleted, and / or additional enzyme activities can be introduced. Such modifications can alter the biosynthetic pathway and allow the obtaining of modified derivatives of the compound.
[0055] Therefore, the fifteenth aspect provides a method for preparing a nucleic acid molecule encoding a modified biosynthetic system for synthesizing a modified derivative of a compound of formula I as defined herein, the method comprising modifying a nucleic acid molecule as defined herein. The nucleic acid molecule can be modified by introducing, mutating, deleting, substituting or inactivating a sequence encoding one or more activities or proteins encoded by the nucleic acid molecule.
[0056] In one embodiment, one or more nucleotide sequences of SEQ ID NOs. 2-29 are modified, or nucleotide sequences complementary or degenerate to any of SEQ ID NOs. 2-29. Detailed Description
[0057] The disclosure and invention of the present text relate to the novel antibiotic compound nidaromycin, which has the structure shown in Formula I as defined above. As described above, this novel compound was discovered by screening and identifying novel biosynthetic gene clusters in actinomycete strains. The prior bioactivity data of the strains, combined with extensive bioinformatics and phylogenetic analysis work, including genome sequencing, gene annotation, clustering analysis, and manual processing of the analysis results, as detailed in the following examples, enabled us to select certain strains and identify novel BGCs in their genomes. Based on our analysis work, one such cluster, namely Cluster 16, was selected from the isolate identified as strain P08-G05. Based on bioinformatics and clustering analysis, it was hypothesized that this cluster encodes a novel monomycin-like compound. The cluster was cloned and transferred into a heterologous host, specifically Streptomyces coelicolor M1152ΔmatAB, the preparation of which is further detailed below. The resulting transconjugant strain, namely strain M1152ΔmatAB(P08-G05_C16), was cultured and shown to have antibacterial activity against various Gram-positive bacteria, including Staphyloccus aureus and Enterococcus faecium, in cell-free extracts. Additionally, in the toxicity tests described in the following examples, the compound was shown not to exhibit toxicity to mammalian cell lines. The compound was purified and subjected to structure elucidation studies, as detailed in the following examples. This led to the structure elucidation of the novel compound.
[0058] As shown in Formula I, we believe there are two possibilities for the position of the sulfate group in this molecule, denoted as R 1 and R 2 at C4 and C3 of ring moiety D, respectively (see below), and are referred to as nidaromycin D4 and nidaromycin D3, respectively.
[0059]
[0060] In the compound of formula (I),
[0061] R 1 is -SO 2 OH, -SO 2 OR or -SO 2 R and R 2 is H, or
[0062] R 2 is -SO 2 OH, -SO 2 OR or -SO 2 R and R 1 is H;
[0063] wherein R is C1 -C 20 hydrocarbyl group.
[0064] Preferably, R 1 is -SO 2 OH or -SO 2 OR and R 2 is H, or R 2 is -SO 2 OH or -SO 2 OR and R 1 is H. Most preferably, R 1 is -SO 2 OH and R 2 is H, or R 2 is -SO 2 OH and R 1 is H. It should be understood that when the compound is in salt form, the H on any -SO 2 OH group can be replaced by a non-H ion.
[0065] The carboxylic acid groups in the nidaromycin compounds can be modified to their ester forms, as shown by the group -COOR 3 in Formula I, where R 3 is C 1 -C 20 hydrocarbyl group. As is well known in the art, this can be achieved by chemical and enzymatic methods.
[0066] Hydrocarbyl group generally refers to alkyl, alkenyl, alkynyl or aryl group in this article, preferably alkyl and aryl groups, and most preferably alkyl group. C 1 -C 20 Hydrocarbyl group generally refers to C 1 -C 20 alkyl group, C 1 -C 20 alkenyl group, C 1 -C 20 alkynyl group or C 6 -C 20 aryl group. Alkyl, alkenyl and alkynyl groups can be cyclic or acyclic. Alkyl, alkenyl and alkynyl groups can also be straight-chain or branched-chain. C 1 -C 20 Hydrocarbyl group is generally C 1 -C 10 hydrocarbyl group (e.g., C 1 -C 10 alkyl group, alkenyl group, alkynyl group or aryl group), such as C 1 -C 6 hydrocarbyl group (e.g., C 1 -C 6 alkyl group, alkenyl group, alkynyl group or aryl group), and most preferably C 1 -C6 Alkyl. The above definition of hydrocarbyl applies particularly to R and R 3 .
[0067] Each R 3 group may independently be selected from H or C 1 -C 20 hydrocarbyl. It is possible that all R 3 groups are H, or that two R 3 groups are H and the other is C 1 -C 20 hydrocarbyl, or that one R 3 group is H and the other two are C 1 -C 20 hydrocarbyl, or that all three R 3 groups are C 1 -C 20 hydrocarbyl. Preferably, all R 3 groups are H. When the compound is in salt form, H (i.e., R 3 is H) can obviously be replaced by a non-H cation.
[0068] Thus, in one embodiment, the compound has the structure of formula II:
[0069]
[0070] wherein R 1 is -SO 2 OH, -SO 2 OR or -SO 2 R and R 2 is H, or wherein R 2 is -SO 2 OH, -SO 2 OR or -SO 2 R and R 1 is H;
[0071] wherein R is C 1 -C 20 hydrocarbyl,
[0072] preferably wherein R 1 is -SO 2 OH and R 2 is H, or wherein R 2 is -SO 2 OH and R 1 is H;
[0073] or a pharmaceutically acceptable salt, solvate or hydrate thereof.
[0074] In a specific embodiment, the compound has structure I:
[0075]
[0076] or a pharmaceutically acceptable salt, solvate or hydrate thereof.
[0077] In another specific embodiment, the compound has Structure II:
[0078]
[0079] or a pharmaceutically acceptable salt, solvate or hydrate thereof.
[0080] It should be understood that these compounds may be in the form of pharmaceutically acceptable salts, solvates or hydrates.
[0081] The compound may be in the form of a metal salt, such as a lithium, sodium, potassium or calcium salt (in which case, usually one or more R 3 groups are lithium, sodium or potassium).
[0082] Pharmaceutically acceptable salts can also be readily prepared by using the desired acid. The salts can be precipitated from solution and collected by filtration, or can be recovered by evaporation of the solvent. Suitable addition salts are formed from inorganic or organic acids that form non-toxic salts, and examples are hydrochloride, hydrobromide, hydroiodide, sulfate, bisulfate, nitrate, phosphate, hydrogen phosphate, acetate, trifluoroacetate, maleate, malate, fumarate, lactate, tartrate, citrate, formate, gluconate, succinate, pyruvate, oxalate, oxaloacetate, trifluoroacetate, saccharate, benzoate, alkyl or aryl sulfonates (such as mesylate, esylate, benzenesulfonate or p-toluenesulfonate) and isethionate.
[0083] Those skilled in the art of organic chemistry will understand that many organic compounds can form complexes with solvents, in which they react or precipitate or crystallize from the solvent. These complexes are called "solvates". Complexes with water are called "hydrates". Solvates and hydrates of the compounds of the present invention are within the scope of the present invention. Salts of the compounds can form solvates and hydrates, and the present invention also includes all such solvates and hydrates.
[0084] As described above, the compound has been shown to have antibiotic activity. In other words, the compound has antimicrobial activity. More specifically, the compound has antibacterial activity. That is to say, the compound is capable of inhibiting the growth and / or viability of microorganisms, especially bacteria. In one embodiment, the compound has antibacterial activity against Gram-positive bacteria.
[0085] As detailed in the following examples, extracts of the transconjugant strain producing the compound were tested in a bioassay against a panel of strains and their activity was demonstrated. In addition, the compound has been purified and the activity of the purified compound has been confirmed. The minimum inhibitory concentrations (MICs) against Gram-positive indicator organisms have been determined as follows:
[0086] - MIC 70 Staphylococcus aureus ATCC 29213: 0.53 μg / ml
[0087] - MIC 70 Staphylococcus aureus ATCC 43300 (MRSA): 0.53 μg / ml
[0088] - MIC 70 : Enterococcus faecalis CTC 492: 8.45 μg / ml
[0089] - MIC 70 : Enterococcus faecalis CCUG 37832: 2.11 μg / ml
[0090] With the exception of the case of Enterococcus faecalis CCUG 37832, these values are advantageous compared to vancomycin used as a reference compound. In particular, the compound is more potent against both methicillin-resistant and methicillin-susceptible strains of Staphylococcus aureus than vancomycin.
[0091] Thus, in one embodiment, the compound has activity against drug-resistant (or antibiotic-resistant) bacteria, particularly multi-drug resistant (MDR) bacteria. Specifically, the compound has activity against drug-resistant (or antibiotic-resistant) Gram-positive bacteria, particularly multi-drug resistant (MDR) Gram-positive bacteria.
[0092] Specifically, the compound is effective against bacteria resistant to methicillin and / or vancomycin, particularly Gram-positive bacteria.
[0093] More generally, the compound is particularly effective against Gram-positive bacteria resistant to any class of antibiotics, including β-lactam antibiotics (including penicillins and cephalosporins), glycopeptides, (phospho)lipoglycans, macrolides, tetracyclines, sulfonamides, aminoglycosides, carbapenems, and quinolones (including fluoroquinolones) antibiotics. In one embodiment, the class of antibiotics is β-lactams or glycopeptides.
[0094] Specifically, the compound is effective against Staphylococcus and / or Enterococcus. In one embodiment, the compound is effective against Staphylococcus aureus and / or Enterococcus faecalis. In a more specific embodiment, the compound is effective against methicillin-resistant Staphylococcus aureus (MRSA) and / or vancomycin-resistant Enterococcus faecalis.
[0095] However, more generally, the compound can be used against any kind of Gram-positive bacteria, particularly clinically relevant Gram-positive bacteria, including, for example, Micrococcus, Streptococcus, Pneumococcus, Bacillus, Listeria, Clostridium.
[0096] The antibiotic or antimicrobial activity can be readily evaluated according to methods well-known in the art. For example, the broth microdilution assay for determining MIC (used, for example, in the following examples) is widely used and reported, and so are agar plate-based methods (such as disk diffusion assays).
[0097] Notably, the compound has also been shown to have no cytotoxicity or negligible cytotoxicity to mammalian cells, indicating its suitability for clinical use.
[0098] Thus, as described above, the compounds herein have a medical use, particularly as therapeutic antimicrobials or more specifically antibacterial agents, i.e., for treating or preventing microbial or bacterial infections.
[0099] The subject to be treated with the compound can be any subject suffering from or at risk of infection. As described above, the subject is typically a human, but includes veterinary uses, so the subject can be any animal, particularly a vertebrate, such as an animal selected from mammals, birds, amphibians, fish, and reptiles. Thus, the compound can be used both clinically and in the livestock and agricultural environments, including, for example, aquaculture. In one embodiment, the subject is a mammalian. The animal can be a domestic or farm animal or an animal of commercial value, including laboratory animals or animals in zoos or amusement parks. Thus, representative animals include dogs, cats, rabbits, mice, guinea pigs, hamsters, horses, pigs, sheep, goats, cows, chickens, turkeys, guinea fowl, ducks, geese, parrots, budgerigars, pigeons, salmon, trout, cod, haddock, sea bass, and carp. The subject can be regarded as a patient.
[0100] In addition to its use in the animal context, the compound can also be used as an antimicrobial (or antibacterial agent) in the plant context, i.e., as a plant protection agent against plant pathogens. Thus, the compound is generally problematic in the agricultural and horticultural context, or in other words, in the treatment or prevention of plant infections.
[0101] The compound is administered to a subject in an amount effective to treat or prevent an infection. An "effective amount" of the compound is an amount effective to inhibit the growth and / or viability of a microorganism (e.g., a bacterium) and / or to provide a clinical benefit to the subject, such as an amount that provides a measurable or discernible improvement in one or more clinical parameters or symptoms of a clinical condition of the subject (e.g., an infection).
[0102] One of ordinary skill in the art will be able to readily determine the effective amount of the compound based on conventional dose-response protocols and by conveniently using conventional techniques such as those described above for assessing microbial growth inhibition and the like.
[0103] The appropriate dose of the compound will vary depending on the subject and can be determined by a physician or veterinarian based on the weight, age, and sex of the subject, the severity of the condition, and the mode of administration.
[0104] The term "treatment" is used herein broadly to include any therapeutic effect, i.e., any beneficial effect on a condition or on an infection. Thus, "treatment" includes not only eradicating or eliminating an infection, or curing a subject or an infection, but also improving an infection or condition of a subject. Thus, "treatment" includes, for example, an improvement in any symptom or sign of an infection or condition, or an improvement in any clinically acceptable indicator of an infection / condition. Thus, treatment includes, for example, curative and palliative treatment of a pre-existing or diagnosed infection / condition.
[0105] As used herein, the term "prevention" refers to any prophylactic or preventive effect. Thus, "prevention" includes delaying, limiting, reducing, or preventing an infection or its onset, or one or more of its symptoms or manifestations, e.g., relative to a condition or symptom or manifestation prior to prophylactic treatment. Thus, prevention expressly includes absolutely preventing the occurrence or development of an infection or its symptoms or manifestations, as well as any delay, or reduction or limitation in the development or progression of the onset or development of an infection or symptoms or manifestations.
[0106] For such uses, the compound can be formulated into a pharmaceutical composition. The pharmaceutical composition comprises one or more pharmaceutically acceptable carriers, additives, or excipients.
[0107] The pharmaceutical composition can be formulated for administration by any convenient or desired means, such as parenterally, enterally, orally (especially per os), or topically, or by inhalation.
[0108] Conventional Galenic preparations include tablets, pills, powders (e.g., inhalable powders), lozenges, sachets, cachets, elixirs, suspensions, emulsions, solutions, syrups, aerosols (as solids or in liquid media), sprays (e.g., nasal sprays), compositions for nebulizer ointments, soft and hard (e.g., gelatin) capsules, suppositories, sterile injectable solutions, sterile packaged powders, etc.
[0109] Examples of suitable carriers, excipients and diluents are lactose, glucose, sucrose, sorbitol, mannitol, starch, gum arabic, calcium phosphate, inert alginates, tragacanth, gelatin, calcium silicate, microcrystalline cellulose, polyvinylpyrrolidone, cellulose, syrups, water, water / ethanol, water / ethylene glycol, water / polyethylene, hypertonic saline, ethylene glycol, propylene glycol, methylcellulose, methyl hydroxybenzoate, propyl hydroxybenzoate, talc, magnesium stearate, mineral oil or fatty substances such as stearin or suitable mixtures thereof. Notable excipients and diluents are mannitol and hypertonic saline (saline).
[0110] The composition may additionally contain additives such as lubricants, wetting agents, emulsifying agents, suspending agents, preservatives, sweetening agents, flavoring agents, etc. Another therapeutic active agent may be included in the pharmaceutical composition.
[0111] Parenteral dosage forms (e.g., intravenous solutions) should be sterile and free of physiologically unacceptable agents, and should have a low osmolality to minimize irritation or other side effects upon administration, and thus the solution should preferably be isotonic or slightly hypertonic, e.g., hypertonic saline (saline). Suitable vehicles include aqueous carriers commonly used for administering parenteral solutions, such as sodium chloride injection, Ringer's injection, glucose injection, glucose and sodium chloride injection, lactated Ringer's injection and other solutions known in the art. The solution may contain preservatives, antimicrobial agents, buffers and antioxidants, excipients and other additives commonly used in parenteral solutions, which are compatible with the compound and do not interfere with the manufacture, storage or use of the product.
[0112] For topical administration, the compound may be incorporated into creams, ointments, gels, transdermal patches, etc. The compound may also be incorporated into medical dressings, e.g., wound dressings, e.g., woven (e.g., fabric) dressings or non-woven dressings (e.g., gels or dressings having a gel component).
[0113] Additional delivery systems include in-situ drug delivery systems, such as gels in which a solid, semi-solid, amorphous or liquid crystal gel matrix is formed in-situ and may contain the compound. Such matrices can be conveniently designed to control the release of the compound from the matrix, e.g., the release can be delayed and / or sustained over a selected period of time. Such systems can form gels only upon contact with biological tissue or fluid. Typically, the gels are bioadhesive. The pre-gel composition can be targeted to any body site that can retain or is adapted to retain the pre-gel composition by this delivery technique.
[0114] For oral, buccal and dental surface applications, such as for oral care or hygiene purposes, toothpastes, dental gels, dental foams and mouthwashes are specifically mentioned.
[0115] Inhalable compositions can take the form of, for example, inhalable powders, solutions or suspensions. These can include, for example, aerosolizable solutions without propellants.
[0116] The compound can be used in combination with other therapeutic active agents, including other antibiotics or antimicrobials, such as antifungal or antiviral agents. These agents can be used alone, or simultaneously or sequentially or separately in the same composition, e.g., at any desired time intervals.
[0117] Accordingly, the present invention also provides products, such as kits, which contain a compound as defined herein and a second therapeutic active agent, for use alone, sequentially or simultaneously as a combined preparation for treating a subject, particularly for treating an infection in a subject.
[0118] Other possible therapeutic active agents include immunomodulators, such as immunostimulants, such as cytokines or interferons, growth factors, enzymes, mucolytics, analgesics, anti-inflammatory agents, bronchodilators or steroids, etc.
[0119] As described above, the compound and the second therapeutic active agent can be formulated together in the same composition, or in separate formulations, for simultaneous or separate administration, e.g., according to a defined dosage regimen.
[0120] In addition to medical uses, the antimicrobial properties of the compound can also be used in non-clinical settings. Thus, in addition to the above-mentioned plant protection uses, the compound can be used in non-biological environments, such as on non-biological (or in other words inanimate) surfaces or in non-biological locations, for the purpose of disinfection or purification, or for preventing or reducing bacterial colonization.
[0121] Accordingly, the compound can be used as an antibacterial agent against bacteria on any surface. The surface is not limited and includes any surface on which bacteria may appear. Inanimate (or abiotic) surfaces include any surface that may be exposed to microbial contact or contamination. Thus, it particularly includes surfaces on medical devices or machinery (e.g., industrial machinery), or any surface exposed to an aquatic environment (e.g., marine equipment, or a ship or boat or parts or components thereof), or any surface exposed to any part of the environment, such as on pipes or buildings. Such inanimate surfaces exposed to microbial contact or contamination particularly include any part of the following: food or beverage processing, preparation, storage, or distribution machinery or equipment, air conditioning units, industrial machinery (e.g., in chemical or biotech processing plants), storage tanks, medical or surgical equipment, and cell and tissue culture equipment. Any device or equipment used for carrying, transporting, or delivering materials is vulnerable to microbial contamination. Such surfaces will particularly include pipes (the term is used broadly herein to include any conduit or pipeline). Representative inanimate or abiotic surfaces include, but are not limited to, food processing, storage, distribution, or preparation equipment or surfaces, tanks, conveyor belts, floors, drain pipes, coolers, freezers, equipment surfaces, walls, valves, belts, pipes, air conditioning ducts, cooling devices, food or beverage distribution pipelines, heat exchangers, the hull of a ship or any part of a ship structure exposed to water, dental water pipes, oil drilling pipelines, contact lenses, and storage bins.
[0122] As described above, medical or surgical devices or apparatuses represent a particular class of surfaces on which bacterial contamination can form. This can include any kind of pipeline, including catheters (e.g., central venous and urinary catheters), prosthetic devices such as heart valves, artificial joints, dentures, dental crowns, dental caps, and soft tissue implants (e.g., breast, hip, and lip implants). It includes any kind of implantable (or "indwelling") medical device (e.g., stents, intrauterine devices, pacemakers, intubation tubes (e.g., endotracheal tubes or tracheostomy tubes), prosthetic or orthotic devices, pipelines or catheters). An "indwelling" medical device can include a device any part of which is contained within the body, i.e., the device can be indwelling in whole or in part.
[0123] The surface can be made of any material. For example, it can be metal, such as aluminum, steel, stainless steel, chromium, titanium, iron, their alloys, etc. The surface can also be plastic, glass, brick, tile, ceramic, porcelain, wood, vinyl, linoleum, or carpet, their combinations, etc. These surfaces can also be food, such as beef, poultry, pork, vegetables, fruits, fish, shellfish, their combinations, etc.
[0124] The compound can also be incorporated into materials and products for such disinfection purposes.
[0125] For such in vitro applications, the compound can also be used in combination with other agents, such as other antimicrobials, disinfectants, cleaning agents, etc., or it can be incorporated into paints and coatings, etc.
[0126] The compound is biosynthetically prepared by expressing the BGC in a suitable host. For this purpose, nucleic acid molecules are provided herein that contain nucleotide sequences corresponding to the BGC (or, in other words, nucleic acid molecules that contain the constituent nucleotide sequences of the BGC).
[0127] The nucleotide sequence of the BGC is shown in SEQ ID NO.1, which represents the sequence of the BGC cloned from the marine actinomycete isolate PG08-G05. SEQ ID NO.1 has been annotated and shown to contain many genes or ORFs that encode various polypeptides responsible for the activities required for compound synthesis.
[0128] In particular, the antiSMASH software is used in combination with manual analysis and processing to predict the putative functions of the genes. antiSMASH is a software package for the identification, annotation, and analysis of secondary metabolite biosynthetic gene clusters in microbial genome sequences (Medema et al., Nucleic Acids Research, 2011, Vol. 39, Web server issue, W339-W346).
[0129] The BGC encodes the components necessary to produce the compound in the host strain or IVTT system. However, not all of the encoded polypeptides are considered to play a role in biosynthesis, and thus not all ORFs are essential. Various genes / ORFs can encode enzymes that catalyze one or more reactions in the biosynthetic pathway, or proteins that do not have enzymatic activity but are involved in other processes, such as regulation of the synthesis process (e.g., transcription factors), or transport (e.g., transport of the compound or intermediate or substrate compound within the cell or in and out of the cell), or confer resistance (or "immunity") to the synthesized compound. Many sequences encoding transporters have been identified. The BGC contains 28 ORFs as shown in Table 1 below.
[0130] Table 1
[0131]
[0132] The amino acid sequences corresponding to the translations of SEQ ID NOs. 2-29 are shown in SEQ ID NOs. 30-57, respectively.
[0133] As described above, the nucleic acid molecule may comprise a nucleotide sequence corresponding to all or part of a BGC. Thus, it may comprise a nucleotide sequence corresponding to all 28 ORFs in SEQ ID NOs. 2-29 (or a sequence having at least 85% sequence identity therewith), or it may comprise a subgroup thereof. In other words, it may comprise all or part of SEQ ID NO.1 (or a sequence having at least 85% sequence identity therewith). This part may correspond to or represent an individual ORF or gene, for example it may encode a polypeptide involved in the biosynthesis of the compound. Such polypeptides and their coding sequences per se may represent individual useful products and may have other uses in addition to the biosynthesis of the compound.
[0134] In another embodiment, the part may comprise a plurality, or two or more ORFs, but less than the complete complementary sequences of 28 ORFs, i.e., 2-27 ORFs of any one of SEQ ID NOs. 2-29. The part may comprise ORFs that are necessary and sufficient for the biosynthesis of the compound.
[0135] Although the nucleotide sequence encoding a polypeptide is a specific embodiment herein, in another embodiment, the nucleotide sequence may comprise functional genetic elements such as promoters, operators, promoter-operators, enhancers or other regulatory regions.
[0136] The nucleic acid molecule need not comprise the entire cluster, as shown in Table 1, but may comprise a part or parts thereof. This may include one or more genes and / or regulatory sequences, or non-coding or coding functional genetic elements, etc. Generally, for the production of the compound, the nucleic acid molecule will comprise many different genes and / or regulatory molecules, resulting in the synthesis of an antibiotic compound, e.g., nidaromycin or its derivatives. One or more additional nucleotide sequences encoding another activity may also be included in the nucleic acid molecule, e.g., a polypeptide having an activity that modifies the structure of nidaromycin, e.g., to produce a derivative or analogue, e.g., an ester included in Formula I. As further described below, the sequence encoding an enzyme may also be modified to alter the activity of the enzyme, or the sequence may be deleted or inactivated to alter the structure of the synthesized compound, thereby producing a derivative molecule.
[0137] In a specific embodiment, the nucleic acid molecule comprises a nucleotide sequence sufficient to synthesize the compound in a suitable host cell. In other words, the nucleic acid molecule encodes a biosynthetic system or mechanism for synthesizing the compound. In other words, the nucleic acid molecule comprises a BGC for synthesizing the compound.
[0138] The nucleotide sequence contained in the nucleic acid molecule can be defined as a biosynthetic gene or ORF, i.e., a gene or ORF encoding a polypeptide that functions in the biosynthesis of a compound of formula I (or more particularly formula II) or nidaromycin (as shown in structures I and II above) or a derivative or related molecule thereof.
[0139] In one embodiment, a portion of the nucleic acid molecule can be at least 300, such as at least 350, 400, 450, 500, 600, 700, 800, 900, or 1000 bases in length, particularly in the case of a single gene / ORF. However, when the molecule contains a sequence corresponding to all or a major part of BCG, the portion will be considerably larger, such as at least 10,000, 20,000, 25,000, or 30,000 bases.
[0140] The nucleic acid molecule can be an isolated molecule, separated from the components usually found in nature, or it can be a recombinant or synthetic nucleic acid molecule. Generally, since the BGC is cloned from its natural host, the molecule will be an artificial molecule.
[0141] It can be any nucleic acid, but generally it will be DNA.
[0142] As defined above, the nucleic acid molecule can contain a nucleotide sequence that is a variant of the sequences of SEQ ID NOs. 1-29, or a nucleotide sequence encoding a variant (e.g., a functionally equivalent variant) of the amino acid sequences of SEQ ID NOs. 30-57. Such variants can include portions, degenerate sequences, or homologs, which are defined by their % sequence identity to any one or more of SEQ ID NOs. 1-57. The activity of the variant polypeptide or the polypeptide encoded by the variant nucleotide sequence can be as defined above.
[0143] The term "biosynthetic gene or ORF" as used above includes such variant sequences. The variant sequences retain at least one function of the entity from which they are derived, such as encoding a polypeptide having substantially the same properties or activities as the original / source / parental polypeptide, or at least the same general type of properties or functions.
[0144] Generally, the term "gene" includes an ORF encoding a polypeptide and can include regulatory sequences such as a promoter. The term "ORF" refers only to the portion of the gene responsible for encoding the polypeptide.
[0145] As described above, a nucleic acid molecule can comprise a nucleotide sequence selected from any one or more of SEQ ID NO.1 or SEQ ID NOs. 2-29, or a nucleotide sequence that exhibits at least 85% sequence identity with any of the foregoing sequences. More specifically, this can be at least 87%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity, or a sequence complementary or degenerate thereto.
[0146] In addition, a nucleic acid molecule can comprise a nucleotide sequence encoding one or more amino acid sequences selected from SEQ ID NOs. 30-57, or a nucleotide sequence of an amino acid sequence that exhibits at least 85% sequence identity therewith.
[0147] Similarly, a polypeptide herein, i.e., a polypeptide encoded by a nucleic acid molecule as defined and described herein, can comprise all or part of the amino acid sequence shown in any of SEQ ID NOs. 30-57, or an amino acid sequence having at least 85% sequence identity therewith.
[0148] More specifically, in the case of the foregoing amino acid sequences, this can have at least 87%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with the amino acid sequence shown in any of SEQ ID NOs. 30-57.
[0149] Percent sequence identity can be readily determined by commercially available sequence comparison programs that can calculate the percent homology or identity between two or more sequences.
[0150] Percent homology or sequence identity can be calculated by contiguous sequences, i.e., aligning one sequence with another and comparing each amino acid in one sequence directly with the corresponding amino acid in the other sequence, one residue at a time. This is referred to as a "gapless" alignment. Usually, such a gapless alignment is performed only on a relatively small number of residues.
[0151] Although this is a very simple and consistent method, it does not take into account, for example, that in a pair of otherwise identical sequences, an insertion or deletion in the nucleotide sequence can cause a shift in the subsequent codons and thus can cause a substantial reduction in the percent homology when performing a global alignment. Thus, most sequence comparison methods are designed to produce an optimal alignment that takes into account possible insertions and deletions without unduly penalizing the overall homology score. This is achieved by inserting "gaps" in the sequence alignment to attempt to maximize local homology.
[0152] However, these more complex methods assign a "gap penalty" to each gap that appears in the alignment, such that for the same number of identical amino acids, a sequence alignment with as few gaps as possible (reflecting a higher relatedness between the two compared sequences) will receive a higher score than a sequence alignment with many gaps. An "affine gap cost" is commonly used, which charges a relatively high cost for the presence of a gap and a smaller penalty for each subsequent residue in the gap. This is the most commonly used gap scoring system. A high gap penalty will of course result in an optimized alignment with fewer gaps. Most alignment programs allow modification of the gap penalty. However, when using such software for sequence comparison, it is preferred to use the default values. For example, when using the GCG Wisconsin Bestfit software package, the default gap penalty for amino acid sequences is -12 for a gap and -4 for each extension.
[0153] Thus, taking into account the gap penalty, the calculation of the maximum homology / sequence identity percentage first requires generating an optimal alignment. A suitable computer program for performing such an alignment is the GCG Wisconsin Bestfit software package (University of Wisconsin, USA; Devereux et al., (1984) Nucleic Acids Res. Vol. 12: p. 387). Examples of other software that can perform sequence comparisons include, but are not limited to, the BLAST software package (see Ausubel et al., (1999) ibid - Chapter 18), FASTA (Atschul et al., (1990) J. Mol. Biol. pp. 403 - 410), and the GENEWORKS comparison tool suite. Both BLAST and FASTA can be used for offline and online searches (see Ausubel et al., (1999) ibid, pp. 7 - 58 to 7 - 60). However, for some applications, it is preferred to use the GCG Bestfit program. Another tool called BLAST 2Sequences can also be used to compare protein and nucleotide sequences (see FEMS Microbiol. Lett. (1999) Vol. 174: pp. 247 - 250; FEMS Microbiol. Lett. (1999) Vol. 177: pp. 187 - 188).
[0154] Although the final percent homology can be measured by identity, the alignment process itself is generally not based on an all-or-none pairwise comparison. Instead, a scaled similarity scoring matrix is typically used, which assigns a score for each pairwise comparison based on chemical similarity or evolutionary distance. An example of such a commonly used matrix is the BLOSUM62 matrix - the default matrix for the BLAST program suite. The GCG Wisconsin programs generally use the common default values or a custom symbol comparison table if provided (see the user manual for more details). For some applications, it is preferred to use the common default values of the GCG software package, or in the case of other software, a default matrix such as BLOSUM62. Appropriately, the percent identity is determined over the entire reference sequence and / or query sequence.
[0155] Once the software has produced the optimal alignment, the percent homology, preferably the percent sequence identity, can be calculated. The software generally performs this as part of the sequence comparison and generates a numerical result.
[0156] Such variants of a particular sequence can be readily prepared using recombinant DNA techniques such as site-directed mutagenesis or gene replacement or gene editing techniques, homologous recombination, etc. A variety of such methods are known and described in the art.
[0157] As described above, nucleic acid molecules and nucleic acid sequences as defined herein can include any nucleic acid and can be DNA or RNA. They can be single-stranded or double-stranded. Those skilled in the art will understand that due to the degeneracy of the genetic code, many different nucleic acid molecules / nucleotide sequences can encode the same polypeptide. In addition, it should be understood that those skilled in the art can use conventional techniques for nucleotide substitution to reflect the codon usage of any particular host organism in which the polypeptides of the present invention will be expressed, and these nucleotide substitutions do not affect the polypeptide sequence encoded by the nucleic acid molecule / polynucleotide / nucleotide sequence as defined herein.
[0158] Nucleic acid molecules / nucleotide sequences such as DNA nucleic acid molecules / sequences can be produced by recombinant, synthetic, or any means available to those skilled in the art. They can also be cloned by standard techniques.
[0159] Longer nucleic acid molecules / polynucleotides / nucleotide sequences are generally produced using recombinant means, such as by cloning using polymerase chain reaction (PCR) or other cloning techniques.
[0160] The nucleic acid molecules of the present invention can also contain nucleotide sequences encoding selectable markers. Suitable selectable markers are well known in the art and include, but are not limited to, fluorescent proteins such as GFP. Appropriately, the selectable marker can be a fluorescent protein, such as GFP, YFP, RFP, tdTomato, dsRed, or variants thereof.
[0161] The nucleic acid molecule can be provided as part of a nucleic acid construct that comprises the nucleic acid molecule and one or more other nucleotide sequences. These other nucleotide sequences can encode selectable markers or other polypeptides, which can be any other polypeptides that are desired to be introduced into the host cell together with the BGC.
[0162] In another embodiment, the nucleic acid molecule can be provided in the form of a recombinant construct that comprises a nucleic acid molecule that is operably linked to one or more expression control sequences (such as a promoter), optionally operably linked to one or more additional regulatory sequences. Thus, for example, a nucleotide sequence corresponding to a single or selected ORF from a BGC can be provided in a construct having heterologous regulatory sequences for controlling gene expression. A construct comprising a nucleic acid molecule that contains one or more coding sequences and one or more expression control sequences can be referred to herein as an expression construct.
[0163] The nucleic acid molecule or recombinant construct can be contained within a vector, which can be used for cloning, transfer, or expression purposes. Thus, in an embodiment, the vector can be a cloning vector, a transfer vector, or an expression vector. As used herein, the term "vector" refers to any genetic element that can serve as a vehicle for the genetic transfer, expression, or replication of an exogenous nucleic acid sequence in a host strain.
[0164] The vector can exist as a single nucleic acid molecule or as two or more separate nucleic acid molecules. When present in a host strain, the vector can be a single-copy vector or a multi-copy vector.
[0165] A particular vector used herein is an expression vector. In such a vector, one or more genes / coding sequences can be inserted into the vector molecule in an appropriate orientation and in proximity to the expression control elements so as to direct the expression of one or more proteins when the vector molecule is present in a host strain. The expression control elements can be provided by the vector, but conveniently, they are part of the nucleic acid molecule inserted into the vector (i.e., a nucleic acid molecule derived from a BGC as defined and described herein), particularly when the nucleic acid molecule contains a nucleotide sequence corresponding to a complete or substantially complete BGC or a major part thereof. In other words, in one embodiment, the nucleic acid molecule in the vector can contain the control and regulatory sequences for expressing the coding nucleotide sequence, and the vector can simply serve as a vehicle for the molecule for its cloning or introduction into cells, propagation in cells, etc.
[0166] The construction of suitable vectors for the present disclosure and other recombinant or genetic modification techniques are well known in the art (see, e.g., Green and Sambrook, “Molecular Cloning, A Laboratory Manual”, Cold Spring Harbor Laboratory Press (Cold Spring Harbor, N.Y.) (2012), and Ausubel et al., “Short Protocols in Molecular Biology, Current Protocols”, John Wiley and Sons (New Jersey) (2002)).
[0167] Vectors can be plasmids, cosmids, phagemids or other phage vectors, viral vectors, episomes, artificial chromosomes such as bacterial artificial chromosomes (BACs) or P1 artificial chromosomes (PACs), or other polynucleotide constructs.
[0168] Conveniently, the vector is an artificial chromosome, particularly a BAC. This is especially the case when the nucleic acid molecule comprises a sequence corresponding to the whole BGC or a substantial or major part thereof, given the size of the molecule. As described in the following examples, the BGC is cloned into a BAC and, in general, for cloning the nucleic acid molecule for the synthesis of the compound in a host, an artificial chromosome, particularly a BAC, will be used. An exemplary BAC for this purpose is the BAC vector pDualP from Varigen Biosciences, Madison, WI, USA.
[0169] For the cloning and expression of smaller nucleic acid molecules, a variety of plasmids suitable for the selected or desired bacterial host strain are available and known in the art.
[0170] Typically, regulatory control sequences are operably linked to the coding nucleic acid sequence and include constitutive, regulatory and inducible promoters, transcriptional enhancers, transcriptional terminators, etc. well known in the art. The coding nucleic acid sequence can be operably linked to a common expression control sequence or to different expression control sequences. However, as mentioned above, conveniently the native control sequences of the BGC gene are used, particularly when the vector comprises the nucleic acid molecule for the synthesis of the compound.
[0171] Suitable promoter sequences for expression in bacteria, particularly actinomycetes, are known in the art. When using a heterologous promoter, it may be advantageous to use a strong promoter. In particular, strong inducible promoters can be used. For example, this may be the case for the expression of a single gene or ORF.
[0172] The choice of vector will generally depend on the size of the nucleic acid molecule, as well as the compatibility between the vector and the host strain into which the vector is to be introduced. The vector can be a linear or closed circular plasmid. The vector can also be an autonomously replicating vector, i.e., a vector that exists as an extrachromosomal entity and whose replication is independent of chromosomal replication, such as a plasmid, an extrachromosomal element, a minichromosome, or an artificial chromosome. The vector can contain any means to ensure self-replication. Alternatively, the vector can be a vector that integrates into the genome when introduced into the host strain and replicates with the chromosome into which it integrates. Integrative plasmids are known in the art. In addition, a single vector or plasmid, or two or more vectors or plasmids, which together contain the total nucleic acid to be introduced into the genome of the host strain, can be used, or transposons can be used.
[0173] As described above, the vector can contain one or more selectable markers that allow for the easy selection of transformed cells. The selectable marker gene can, for example, encode a detectable product (such as a fluorescent protein), provide resistance to an antibiotic or toxin, complement a nutritional auxotrophic defect, or provide a key nutrient not present in the medium, and / or provide control over chromosomal integration. Examples of bacterial selectable markers are markers that confer antibiotic resistance, such as resistance to ampicillin, kanamycin, chloramphenicol, apramycin, or tetracycline.
[0174] The vector can also contain one or more elements that allow the vector to integrate into the genome of the host strain or to replicate autonomously in the host independently of the genome. For integration into the genome of the host strain, the vector can rely on coding nucleic acid sequences or other elements of the vector to integrate into the genome by homologous or non-homologous recombination. CRISPR-based systems can also be used to achieve integration. For autonomous replication, the vector can also contain an origin of replication that enables the vector to replicate autonomously in the strain under discussion. The origin of replication can be any plasmid replicon that functions in the cell to mediate autonomous replication. The term "origin of replication" or "plasmid replicon" is defined herein as the nucleotide sequence that enables a plasmid or vector to replicate in vivo.
[0175] The vector can be introduced into the host cell by any convenient or desired means, and this depends on the nature of the vector. As used herein, the term "introducing" refers to a method of inserting foreign nucleic acid (e.g., DNA or RNA) into a cell. This includes conjugation and transformation methods, or virtually any method suitable for transferring a nucleic acid molecule or vector into a host cell. A host cell that has been modified by the introduction of a nucleic acid molecule or vector can be referred to as an engineered host cell. Thus, an engineered host cell is distinguished from a natural or wild-type host cell in that a nucleic acid molecule that is not present in the natural or wild-type host cell has been introduced into the engineered host cell.
[0176] For plasmid vectors and the like, transformation methods are known in the art. However, for larger vectors, such as vectors for transferring nucleic acid molecules for the synthesis of compounds, methods for transferring plasmids into host cells by conjugation are typically used. This can include transferring the vector into an intermediate host, i.e., a transfer host, before introducing it into the host for expression, e.g., for the production of a compound. Conveniently, to transfer a vector (e.g., a BAC) for the production of the molecule into a production host, a triparental conjugation method can be used, as known and reported in the art. Thus, by triparental mating of a cloning host containing the vector, a host containing the helper plasmid, and the transfer host, the vector containing the nucleic acid molecule for the synthesis of the compound can be transferred together with the helper plasmid into the transfer host. This method is described in the following examples. Subsequently, the transfer host containing the vector and the helper plasmid is conjugated with the intended production host cell to transfer the vector into the production host cell for the synthesis of the compound.
[0177] As described above, once introduced, the vector can be maintained as a chromosomal integrant or as a self-replicating extrachromosomal vector.
[0178] Methods known in the art can be used to confirm transformation. These methods include, for example, PCR at the integration site (primers in the vector and the host chromosome) or genomic sequencing. Alternatively or additionally, analysis of gene expression levels can be performed, using, for example, Northern blotting or polymerase chain reaction (PCR) amplification of mRNA, or Western blotting for the expressed gene product, or other suitable analytical methods to test the expression of the introduced nucleic acid sequence or its corresponding gene product. In the case where the vector contains a nucleic acid molecule for the synthesis of a compound, the production and detection of the compound will confirm that expression has occurred. Expression levels can be further optimized using methods well known in the art to obtain sufficient expression.
[0179] Which host cell the vector is transferred into will depend on the purpose of the host cell, i.e., whether it is for cloning or transfer, for expressing one or more polypeptides, or for production (synthesis) of a compound.
[0180] Host cells suitable as cloning hosts can be any cells known in the art for this purpose and will depend on the size of the nucleic acid molecule or vector. For example, for cloning smaller nucleic acid molecules containing nucleotide sequences encoding a single polypeptide or a smaller selection thereof, a variety of host cells can be used, including, for example, various strains of Escherichia coli (E. coli). When the nucleic acid molecule is a large molecule, especially for the synthesis of a compound, the host cell needs to be suitable for the propagation of large constructs, and such hosts are also known in the art, including, for example, Escherichia coli strain 10Beta.
[0181] Suitable hosts for transfer by conjugation are also known in the art, including, for example, Escherichia coli ET12567.
[0182] Suitable hosts for expressing a single polypeptide are also known in the art and include many strains of Escherichia coli, Bacillus, and other bacteria, including actinomycetes, especially Streptomyces.
[0183] For the production of a compound, the host cell can be any suitable host cell in which the nucleic acid molecule can be expressed and the compound can be synthesized. The production host is generally a bacterium, conveniently an actinomycete, more specifically a bacterium of the genus Streptomyces.
[0184] The production host cell can be a heterologous host cell, i.e., a host cell that does not naturally contain the BGC or the synthetic compound. However, in an alternative embodiment, the nucleic acid molecule can be introduced into an organism that has cloned the BGC, i.e., isolate P08-G05, or more generally a strain that endogenously contains the BGC.
[0185] A series of different Streptomyces hosts are known in the art and available for use. In one embodiment, the host is Streptomyces coelicolor. Similarly, various strains and isolates of Streptomyces coelicolor are available for use. Strain A(3)2 is a well-known model strain.
[0186] The Streptomyces coelicolor strain M145, also known as Streptomyces violaceoruber (Waksman and Curtis) Pridham, is a prototrophic derivative of strain A(3)2 and is available from the ATCC under the number BAA-471. The Streptomyces coelicolor strain M145 (ATCC BAA-471) is described in Bentley et al., 2002, Nature, Vol. 417, pp. 141-147, and in particular lacks the plasmids SCP1 and SCP2 of the parental A3(2) strain. Derivatives of the Streptomyces coelicolor strain M145 (ATCC BAA-471) have been engineered for the heterologous expression of secondary metabolite gene clusters, as described in Gomez-Escribano and Bibb, Microbial Biotechnology, 2011, Vol. 4, No. 2, pp. 207-215. Any of the strains described in this document can be used. The modified strains have been made to lack all or part of four antibiotic gene clusters (the act, red, cda, and cpk gene clusters) in strain M1146, and further point mutations have been introduced in rpoB and / or rpsL, which encode the RNA polymerase β-subunit and ribosomal protein S12, respectively. It has been shown that each of these mutations can increase the antibiotic production level of Streptomyces without compromising its growth. The mutant alleles were incorporated into a suicide plasmid and the wild-type genes were replaced by homologous recombination, as described in the above document. Any of the strains M1141-M1146 or M1151-M1156 described in this document can be used. Of particular note is the strain M1152 (ΔactΔredΔcpkΔcda rpoB(C1298T)). In addition, as described in the examples below, mutants of the strain M1152 have been generated, in particular in-frame deletion mutants of SCO2963 and SCO2962, in which the matAB locus has been deleted. The strain Streptomyces coelicolor M1152ΔmatAB described in Example 1 below represents a preferred host cell for the production of the compound.
[0187] Once the nucleic acid molecule has been introduced into the host strain for the production of the compound, the modified bacteria are grown or cultured under conditions suitable for the expression of the encoded polypeptide and the synthesis of the compound. Similarly, the processes and conditions for this purpose are known in the art and can be readily achieved according to the techniques and principles well known in the art.
[0188] Growth media suitable for Streptomyces bacteria are known in the art and are described in the examples below, such as the MG-2.5w / NaCl medium.
[0189] Alternatively, compounds can be prepared in an in vitro transcription and translation (IVTT) system (i.e., a cell-free system) according to principles and techniques known in the art. This can include cell extracts of the above-described host cells, including cell extracts from the various Streptomyces coelicolor strains described above, or more generally from Streptomyces or Actinomycetes cells.
[0190] After synthesizing the compound or expressing the desired polypeptide, the compound or polypeptide can be harvested from the culture, or in other words recovered or collected.
[0191] In particular, the compound or polypeptide can be extracted or isolated from bacterial cells by a cell lysis procedure, which is also well known in the art. Thus, a crude extract containing the compound or polypeptide can be obtained.
[0192] In addition, separation and purification procedures known in the art can be used to separate or purify the compound or polypeptide. These methods include, for example, precipitation, chromatography, or filtration methods such as ammonium sulfate precipitation, ion exchange chromatography, reverse phase chromatography, size exclusion chromatography, gel filtration, HPLC methods, etc. Any desired or convenient combination of purification methods can be used. To isolate the compound, solvent and / or acid extraction and preparative chromatography, such as HPLC, can be used. Suitable methods are described in the following examples.
[0193] Thus, the methods herein can include additional steps of purifying the compound or polypeptide.
[0194] In a specific aspect and embodiment herein, the compound is a compound that can be obtained by expressing a nucleic acid molecule comprising SEQ ID NO.1 in Streptomyces coelicolor strain M1152ΔmatAB. More specifically, according to the methods described in the following examples, the compound can be obtained by expressing a nucleic acid molecule comprising SEQ ID NO.1 in Streptomyces coelicolor strain M1152ΔmatAB.
[0195] Thus, the compound can have the structure shown in Formula I above, more specifically including the structure of Formula II or the structure of Formula I or II. However, as described above, by chemically modifying the compound or by modifying one or more coding sequences of the nucleic acid molecule, modifications of the compound can be obtained, i.e., derivatives can be produced. Such modifications can alter one or more enzyme activities, thereby resulting in modifications of the resulting synthesized compound. Modifications of genes in antibiotic gene clusters to modify the resulting antibiotic compounds have been described in the art, for example, described in WO 2001 / 059126 (nystatin) or WO 2009 / 115822 (BE-14106).
[0196] The present invention will now be described in more detail in the following non-limiting examples with reference to the following figures. BRIEF DESCRIPTION OF THE DRAWINGS
[0197] Figure 1 Shows the LC-DAD-isoplot of the extracts of M1152ΔmatAB(P08-G05_C16) (A) and control M1152ΔmatAB (B) cultured in a well plate containing 2.5×MG-2.5 w / NaCl and 0.1% inducer ε-caprolactam. Only two compounds eluting between 13 and 15 minutes were observed in the extract of the transconjugant, namely, Nidaromycin eluting at 14.6 minutes and a derivative eluting at 13.2 minutes.
[0198] Figure 2 Shows the MS spectrum of the compound produced by the transconjugant P08-g05_C16 (corresponding to the main UV peak observed in the transconjugant extract). The compound also forms a Na+ adduct in MS; the red bars show the theoretical isotope distribution of the proposed molecular formula, while the black bars show the measured isotope distribution of the compound;
[0199] Figure 3 Shows the LC-DAD isoplot of the HPLC-purified compound produced by the transconjugant M1152ΔmatAB(P08-G05_C16). In this chromatogram, Nidaromycin elutes as a single peak at approximately 16 minutes, and no other UV-absorbing compounds were observed in the chromatogram.
[0200] Figure 4 Shows the toxicity data of the compound nidaromycin produced by the transconjugant M1152ΔmatAB(P08-G05_C16) at different concentrations from 0.5 μg / ml to 50 μg / ml against cell lines HepG and LLC-PK1 as described in Example 2: (A) viability after 48 hours of exposure, expressed as %, and (B) LDH leakage rate (%) after 48 hours of exposure;
[0201] Figure 5 Shows 15 the MS spectrum of N-labeled nidaromycin (A) and 13 C- and 15 N-labeled nidaromycin (B), showing an increase in molecular weight by 2 Da and 63 Da respectively, indicating that the molecular formula contains 2 N and 61 C.
[0202] Figure 6 Shows the MSMS fragmentation pattern of purified nidaromycin. The precursor mass is M+H+1349.5674
[0203] Figure 7 The predicted structure of the compound produced by the transfer conjugator M1152ΔmatAB(P08-G05_C16) as determined by NMR studies in Example 8 is shown, including the atomic numbering for the specific assignments of atoms in the NMR spectra used in the example.
[0204] Example
[0205] Example 1
[0206] Preparation of Host Strain Streptomyces coelicolor M1152ΔmatAB
[0207] Streptomyces coelicolor strain M1152 described in Gomez-Escribano 2010 (supra) was obtained from the John Innes Centre, Norwich, UK.
[0208] The in-frame deletion mutants of SCO2963 / SCO2962 in Streptomyces coelicolor M1152 were generated as described previously (van Dissel et al., 2015, Microbial Cell Factories, Vol. 14, No. 1, pp. 1-10). Briefly, the upstream region of SCO2963 in the range of -1326 to +43 relative to the start codon and the downstream region of SCO2962 in the range of +2190 to +3610 were amplified from the Streptomyces coelicolor genome by PCR using the primers listed in Table 2. The amplified flanks were cloned into the unstable shuttle vector pWHM3-oriT (Wu et al., 2019, Angewandte Chemie, Vol. 131, No. 9, pp. 2835-2840) using EcoRI and HindIII restriction sites. XbaI site all appears in two amplification zones, for inserting the apramycin resistance box aacC4 flanked by loxP sites between the flanking regions. The completed vector (pMAT1) is introduced into Escherichia coli ET12567+PUZ8002, which allows pMAT1 to be transferred toward Streptomyces coelicolor M1152 by joining. Mutants with Thio- / Apra+ phenotypes are screened out by repeated plate culture, and the matAB loci of these mutants are replaced by aacC4 resistance boxes, and pWHM3 carriers are lost. Unmarked Streptomyces coelicolor M1152 ΔmatAB strains are obtained by introducing the pUWLcre plasmid expressing Cre recombinase, which cuts the loxP sites around the apramycin resistance gene.
[0209] Table 2
[0210] matA_-1326 <![CDATA[AGTC GAATTC CAGCCGGGCGGTGAGATTCC]]> matA_+43 <![CDATA[ACTG TCTAGA CGAGCACTCGTCGGCCGAAC]]> matB_+2190 <![CDATA[AGTC TCTAGA AGGCCGGTCGGATGACCACC]]> matB_+3610 <![CDATA[AGTC AAGCTT CCCTGTTCACTCCCGCAACCG]]>
[0211] Example 2
[0212] Identification of Biosynthetic Gene Cluster (Cluster 16) of Marine Actinomycete Isolate P08-G05
[0213] Source of the marine isolate P08-G05
[0214] The marine isolate P08-G05 was obtained from the SINTEF / NTNU Marine Actinomycete Strain Collection, which was constructed from water samples, sediment samples, and sponge samples taken from the Trondheim fjord. The strain was selected based on a comprehensive assessment of the draft genomes of 1200 isolates from this collection, based on different criteria as described below: phylogenetic novelty, gene cluster diversity, and previously observed bioactivity.
[0215] The frozen glycerol cultures from this collection were streaked on TSA (Tryptic Soy Broth Agar) supplemented with 0.5x artificial seawater (Engelhardt et al. 2010, Applied and Environmental Microbiology 76(15):4969-4976.). The pure isolates were cultured in TSB with artificial seawater to produce mycelia for the working cell bank.
[0216] Illumina sequencing and de novo assembly of the strain P08-G05 genome
[0217] The biomass of strain P08-G05 for genome sequencing was produced in TSB medium supplemented with 50% artificial seawater at 30 °C. The biomass was collected by centrifugation and sent to BaseClear BV for sequencing, where DNA extraction, sequencing, and post-sequencing data processing were performed.
[0218] Paired-end sequence reads were generated using the Illumina HiSeq2500 system. The FASTQ sequence files were generated using the Illumina Casava pipeline version 1.8.3. The initial quality assessment was based on data filtered by Illumina Chastity. Subsequently, reads containing adapters and / or PhiX control signals were removed using the BaseClear in-house filtering scheme. The second quality assessment was based on the remaining reads using the FASTQC quality control tool version 0.10.0. The quality of the FASTQ sequences was improved by trimming off low-quality bases using the "Trim Sequences" option of CLC Genomics Workbench version 8.0.
[0219] For genome assembly and scaffolding, quality-filtered sequence reads were assembled into contig sequences. The analysis was performed using the "de novo assembly" option of CLC Genomics Workbench version 8.0. Misassemblies and nucleotide inconsistencies between Illumina data and contig sequences were corrected with Pilon version 1.11. Contigs were joined and placed into scaffolds or supercontigs. This resulted in an assembly of 7,315,765 bp and 980 scaffolds. The orientation, order, and distance between contigs were estimated using the insert size between paired-end and / or mate-pair reads. The analysis was performed using SSPACE Premium scaffolder version 2.3. The gap regions within scaffolds were (partially) closed automatically using GapFiller version 1.10, taking advantage of the insert size between paired-end and / or mate-pair reads. The resulting draft genome was then used for phylogenetic analysis and genome annotation. The quality of the Illumina-sequenced de novo genome assembly of P08-G05 was evaluated by the checkM software (version 1.07), showing a high integrity of 95.9% and a low contamination of 1.6%.
[0220] PacBio sequencing and hybrid de novo genome assembly
[0221] Cell pellets for PacBio sequencing and direct cloning were generated in 500 ml shake flasks at 30 °C with 200 rpm and 2.5 orbital movement until OD 600 = 5.6. The shake flasks contained 3 g of 3 mm glass beads and 120 ml of TSB medium supplemented with 0.5x artificial seawater. Cell pellets were harvested by centrifugation, stored at -40 °C until transportation, and then shipped on dry ice to BaseClear BV in the Netherlands.
[0222] Long-read PacBio sequencing was performed at BaseClear using a PacBio Sequel instrument, and the obtained data were processed and filtered using the SMRT Link software suite, discarding subreads shorter than 50 bp. This resulted in a number of 622,557 reads with a yield of 2,886,923,195 bp.
[0223] The quality of Illumina HiSeq reads was improved by trimming low-quality bases using BBDuk, which is part of the BBMap suite version 36.77. High-quality reads were assembled into contigs using ABySS version 2.0.2.
[0224] The long reads were mapped to the assembled draft using BLASR version 1.3.1. Based on these alignments, the contigs were joined together and placed into scaffolds. The orientation, order, and distance between contigs were estimated using SSPACE-LongRead version 1.0. The gap regions within the scaffolds were closed using Illumina reads with GapFiller version 1.10 (partial). Finally, the assembly errors and nucleotide inconsistencies between the Illumina reads and the scaffold sequences were corrected using Pilon version 1.21. This resulted in an assembly of 7,840,734 bp with 22 scaffolds.
[0225] Phylogenetic positioning of strain P08-G05
[0226] The phylogenetic position of isolate P08-G05 was determined by performing a comprehensive phylogenetic analysis together with 1200 selected strains from the SINTEF / NTNU Marine Actinomycetes Strain Collection and 576 actinomycete type strains retrieved from public databases, whose genomic sequences were generated in a similar manner as that of P08-G05 (as described above). The analysis was performed using the IQTREE software (IQ-TREE MPI multi-core version 1.6.7.1) with 92 housekeeping genes as references to determine the phylogenetic novelty of strain P08-G05. Strain P08-G05 was classified as other strains in the actinomycete strain collection rather than other type strains, indicating that this strain may represent a new actinomycete species.
[0227] Identification of the biosynthetic gene cluster P08-G05_c16
[0228] Using an in-house Python script, the abundances and diversities of different biosynthetic gene cluster (BGC) categories from 1200 strains in the actinomycete strain collection were evaluated based on a collection of BGC profile hidden Markov models (pHMMs). The obtained matrices containing the pHMM hit counts for the corresponding strains were used to cluster the strains into different populations by implementing the t-distributed stochastic neighbor embedding (t-SNE) algorithm of AFG in the programming language R.
[0229] Strain P08-G05 was clustered together with other strains in cluster 34 (a total of 40 t-SNE clusters). This strain, together with other strains from different t-SNE clusters, was selected into a candidate list of 86 strains for further characterization, including long-read PacBio sequencing.
[0230] Based on the antiSMASH results of the PacBio-sequenced genome of P08-G05, new clusters were identified through manual processing, and these clusters encode (at least) the (core) metabolic mechanisms for the synthesis of the nidaromycin compound. An internal script was used to analyze the resistance genes on the gene clusters. No resistance genes were identified in cluster P08-G06_c16.
[0231] Example 3
[0232] Cloning and Expression of Gene Cluster P08-G06_c16
[0233] Cloning and conjugation of gene cluster P08-G06_c16
[0234] Based on the antiSMASH results, it is hypothesized that gene cluster P08-G06_c16 encodes a new compound similar to monomycin. The molecular formula of monomycin is C 68 H 106 N 5 O 34 P, with a mass of 1567.645683 g / mol.
[0235] Based on the chromosomal DNA of strain P08-G05, cluster P08-G05_c16 was cloned into an inducible bacterial artificial chromosome (BAC) vector (pDualP, proprietary to Varigen Biosciences in Madison, WI, USA), which was purchased from Varigen Biosciences. The construct was obtained from Varigen Bioscience and is suitable for the propagation of large constructs (10Beta) in E. coli strains. The cluster was transferred into Streptomyces coelicolor M1152ΔmatAB prepared according to Example 1 by triparental conjugation and by a procedure similar to the previous method (Jones et al., 2013, PLoS ONE 8(7):e69319.doe:10.1371). In short: the construct containing the BGC was transferred into E. coli ET12567 together with the driver plasmid pR9406 by triparental mating. For this reason, each strain was first cultured overnight on LB agar without selective culture medium. For all three strains, several colonies were scraped using an inoculation loop and streaked together into a small plaque on LB agar containing apramycin, chloramphenicol and ampicillin. As a control, each strain was also streaked separately on the sample selective culture medium. The single colony of ET12567+pR9406+DualP-BGC was grown in a 13ml culture tube containing 5ml LB+ampicillin, chloramphenicol and apramycin until OD was 0.6, after which the culture was precipitated and washed twice with cold LB culture medium. Meanwhile, coelicolor spores were pre-germinated by heat shock at 50°C for 10 minutes and incubated at 30°C for 2 hours-3 hours. Escherichia coli and coelicolor spores were mixed and spread on soy flour mannitol (SFM) agar plates and incubated at 30°C for 18 to 24 hours, then covered with apramycin+nalidixic acid to select for transfer of conjugated Streptomyces colonies. Single colonies were subsequently picked and streaked onto selective SFM plates and expanded to confluent plates for spore harvesting and storage by standard procedures. The new transconjugant strain carrying the nidaromycin gene cluster was abbreviated as M1152ΔmatAB (P08-G05_C16).
[0236] Cultivation of the transfer conjugate M1152ΔmatAB(P08-G05_C16) and expression of the gene cluster P08-G05_c16
[0237] The following is the well plate culture of M1152ΔmatAB(P08 - G05_C16) and M1152ΔmatAB (control): Seed cultures were produced in 250 ml shake flasks with 50 ml of 0.5x tryptone soy broth (TSB) and 1.5 g of 3 mm glass beads (without antibiotics), shaken at 30 °C and 225 rpm for two days until OD600 = 5 - 7. Production was carried out in 24 - well plates (AXYGP - DW10ML24C) using 5254SW medium (Králová et al., 2021, Frontiers in Microbiology, Vol. 12: p. 2131) or MG - 2.5 medium supplemented with 1 g / L NaCl (Doull and Vining, 1990, Applied Microbiology and Biotechnology, Vol. 32: pp. 449 - 454; Martínez - Castro et al., 2013, Applied Microbiology and Biotechnology, Vol. 97: pp. 2139 - 2152). 2.5 ml of medium and 4 × 3 mm glass beads were added to each well, and 1.3% of the seed culture was inoculated. The plates were incubated at 30 °C in a New Brunswick incubator at 800 rpm and 85% humidity for 6 days. The broth was freeze - dried and extracted with DMSO for one hour in a volume equal to the broth volume.
[0238] The cell - free extracts were analyzed using an Agilent LC - DAD - QTOF equipped with a Zorbax Bonus RP 2.1×50 mm, 3.5 μL. 50 mM ammonium acetate [A] and acetonitrile [B] were used as the mobile phase. The gradient was as follows: 5% acetonitrile was used from 0 min - 2 min, and then the acetonitrile concentration was increased to 95% over the next 25 min. The QTOF was operated in both positive and negative ion modes, capillary voltage: 3.5 kV, fragmentation voltage: 150 V, skimmer: 65 V, gas temperature 325 °C, drying gas: 10 L / min, nebulizer: 50. Data were processed using Agilent's Mass Hunter and Mass Profiler Professional software.
[0239] The LC - DAD - isoplot showed that two peaks were observed in the transconjugant extracts but not in the control ( Figure 1 ). The abundance of these two peaks was higher in MG - 2.5 w / 0.5x seawater than in 5254SW medium. The MS data indicated that a cluster of mass peaks was observed at the retention time corresponding to the main UV peak. The three main mass peaks were M + H = 1349.5668, and its adduct was M + Na = 1371.5496 (Figure 2 )。These mass peaks were not found in the extract of the control M1152ΔmatAB. It was inferred that this compound was related to the heterologous expression of the introduced P08-G05_c16 and was named nidaromycin.
[0240] Example 4
[0241] Scale-up Production and Purification of Heterologously Expressed Compounds
[0242] The scale-up production of the bioactive compound was carried out in a 500 ml shake flask with 125 ml of MG-2.5w / NaCl. The medium was inoculated with 3% of the seed culture and incubated at 30 °C with a 200 rpm agitation and a 2.5 cm orbital movement for 6 days.
[0243] The broth was lyophilized and homogenized in a mortar. The substance was extracted with DMSO acidified with trifluoroacetic acid (TFA) to a final concentration of 0.1%. The amount of the organic solvent was 0.4 times the original broth volume. The DMSO extract was fractionated by Agilent preparative HPLC equipped with a Zorbax Bonus RP, 9.4 × 250 mm, 7 μm column (Agilent), a diode array detector (DAD), and a fraction collector. The mobile phase was water containing 20 mM ammonium acetate [A] and acetonitrile [B]. The gradient was 5% [B] during injection, and then increased from 55% to 75% [B] within 10 minutes. The column was washed with 95% [B] for 1 minute before column equilibration with 5% [B]. Acetonitrile in the HPLC fractions was removed by a rotary evaporator, and the aqueous phase was further purified and concentrated using a 500 mg HLB solid-phase extraction column (Waters). The compound was eluted from the SPE column with methanol. Methanol was removed by evaporation at 50 °C using a Speedvac (ThermoFisher). Water was added to the sample, frozen at -80 °C, and lyophilized.
[0244] Figure 3 The DAD chromatogram in confirmed that the purified compound had been obtained. The obtained substance was used for the inhibition and toxicity assays as described in Example 5, to determine the molecular formula of nidaromycin as described in Example 6, and for structure analysis by NMR and determination of the position of the sulfate group as described in Examples 7 and 8.
[0245] Example 5
[0246] Determination of Activities of Crude Extracts and Purified Compounds
[0247] Bioassay of the crude extract
[0248] A cell-free extract was prepared from a culture of strain M1152ΔmatAB (P08-G05_C16) (prepared as described in Example 3) and tested in a bioassay against a panel of strains, namely Enterococcus faecalis CCUG 37832, Micrococcus luteus TO-09 ATCC 9341, Pseudomonas aeruginosa ATCC 15692, and Candida albicans CCUG. The extract showed activity against Enterococcus faecalis CCUG37832.
[0249] Specifically, the extract of the transconjugant inhibited the growth of Enterococcus faecalis CCUG 37832 at 16x dilution (MG-2.5w / NaCl) and 4x dilution (5254SW), while the extract of the control did not inhibit any strain.
[0250] In vitro MIC bioassay
[0251] According to the Clinical and Laboratory Standards Institute protocol, the minimum inhibitory concentration (MIC) against a series of Gram-positive indicator organisms was determined by using a microdilution test in 384-well format. The indicator strains were cultured overnight in TSB medium until OD600 = 0.4, then diluted in TSB medium to OD600 = 0.1 and further diluted 45x in the assay medium (Mueller-Hinton broth, Difco). The inoculated medium was dispensed into the assay plates. Serial 2x dilutions of the isolated compound or vancomycin (reference) diluted in DMSO were added to the inoculated wells to give a final DMSO concentration of 2.7% in each well and a final concentration of the active compound from 0 μg / ml to 540 μg / ml (23 different concentrations). Four parallel values were determined for each compound and concentration. 0.5 mg of the compound produced by strain M1152ΔmatAB (P08-G05_C16) was purified on preparative HPLC and the pure compound was tested in a bioassay against a panel of strains. The indicator organisms were Micrococcus luteus ATCC 9341, Staphylococcus aureus ATCC 29213, Staphylococcus aureus ATCC 43300 (MRSA), Enterococcus faecalis CCUG 37832, Enterococcus faecalis CTC 492. The results are shown in Table 3 below.
[0252] Table 3
[0253]
[0254] In vitro toxicity assay
[0255] The human cell lines HepG2, LLC-PK1, and L929 were used to evaluate the toxicity of nidaromycin on human cell lines. These cell lines were cultured in the following media respectively: RPMI 1640 supplemented with 10% fetal bovine serum (FBS), 2 mM L-glutamine, and 100 U / ml Pen-Strep (HepG2); Medium 199 supplemented with 3% FBS, 2 mM L-glutamine, 100 U / ml Pen-Strep (LLC-PK1); and Dulbecco's Modified Eagle Medium (DMEM) supplemented with 10% FBS, 2 mM L-glutamine, 1 mM sodium pyruvate, and 100 U / ml Pen-Strep (L929). Cells were passaged according to standard protocols using a Tecan EVO robotic workstation and an MCA384 pipetting device with disposable tips (Tecan MCA125 Catalog number 300-5-1-808), and the cell suspension was transferred from a stirred reservoir and inoculated into 384-well plates (Corning assay plates, 3712). The reservoir (flat bottom, 300 mL, Thermo Scientific, 10723363) was equipped with a sterile magnetic stir bar (15×4.5 mm VWR 442-4522) and stirred at 350 rpm. The number of cells in each well was 50,000 (HepG2), 25,000 (LLC-PK1), and 10,000 (L929). After inoculation, the microplates containing the cell suspension were shaken at 1600 rpm and 2.5 mm amplitude (Bioshake) for 20 seconds. The microplates containing the cells were incubated in an atmosphere of 37 °C and 5% CO2. On the day of cell exposure, serial dilutions were prepared in DMSO. The serial dilutions containing the compound were further diluted in cell culture medium and transferred to the assay wells, where a total DMSO concentration of 0.6% was obtained. After exposure, the plates were incubated for an additional 48 hours in a 5% CO 2 atmosphere at 37 °C. Cell viability was measured using the Promega CellTiter-GLO 2.0 viability assay at 24 hours and 48 hours of incubation. The highest concentration tested was 50 μg / ml of Nidaromycin, and no toxic effects were detected at this concentration or lower concentrations.
[0256] Data are shown in Figure 4 .
[0257] Example 6
[0258] Determination of Molecular Formula of Nidaromycin
[0259] MS1 analysis of isotope-labeled broth
[0260] To determine the molecular formula of nidaromycin, strain M1152ΔmatAB (P08-G05_C16) was grown in isotope-labeled medium as follows, in which all carbon sources were 13 C labeled, and all nitrogen sources were 15 N labeling: 3 ml seed culture produced in 0.5×TSB medium was washed once with 10 ml sterile 0.9% NaCl, resuspended in 0.9% NaCl to OD=5, and used to inoculate production culture (0.5% in 20 ml medium). Production was carried out in 250 ml shake flasks containing 1.5 g 3 ml glass beads and 20 ml of the following medium from Silante: 1 g / 100 ml 13C Silex medium powder for E. coli (115204100), E. coli OD2N (110301402), or E. coli OD2CN (110601402). The culture was harvested after six days.
[0261] The freeze-dried isotope-labeled broth was extracted with 2 ml DMSO for 1 hour, and 0.1% trifluoroacetic acid was added per 20 ml original broth volume. The volumetric yield of nidaromycin in these media was very low, but high enough to detect the mass of isotope-labeled nidaromycin. 13 C. 15 N. 13 C and 15 The masses of N-labeled nidaromycin were 1410.7769Da, 1351.5562Da, and 1412.7585Da, respectively, indicating that the molecular formula of nidaromycin contains 61 carbon atoms and 2 nitrogen atoms ( Figure 5 ). Based on MS1 data ( Figure 2 ), the most likely molecular formula is C 61 H 92 N 2 O 29 S, yielding a theoretical monoisotopic mass of 1348.5507 Da.
[0262] MS2 analysis of purified nidaromycin
[0263] The purified nidaromycin was analyzed by LC-MS / MS, and the fragmentation of the precursor mass was m / z = 1349.567. The LC conditions were the same as those described in Example 4, and the MS / MS data were generated in positive mode using a Bruker Impact II QTOF. The MS conditions were: spectral rate: 12 Hz, capillary voltage: 4500 V, end plate offset: 500 V, drying gas: 10 L / min, nebulizer gas: 220, data acquisition control: dynamic MS / MS, collision energy 5 V, with multiCE 20, 50, and 100.
[0264] Molecular weight and MS / MS fragmentation pattern ( Figure 6 ) indicated a molecular formula of C 61 H 92 N 2 O 29 S, and 29 MS / MS fragments could be explained by this molecular formula (data not shown).
[0265] Example 7
[0266] Determination of Structure by NMR
[0267] The structure determination of the compound produced by the transconjugant M1152ΔmatAB (P08-G05_C16) was carried out by Red Glead Discovery AB using 1D and 2D NMR spectra.
[0268] The determined structure is shown in Figure 7 , in which the atom numbers of the atom-specific assignments that have been carried out are also shown. The structure consists of four substituted sugar moieties A–D, a linked 2,3-dihydroxypropanoic acid (E), and a hydrocarbon moiety (F) with the formula C30H45. There are several alternative positions for the proposed sulfate group, and the position shown at position 4 on the uronic acid unit "D" may be valid. The alternative positions are position 3 in unit D and position 4 in unit A, as shown below:
[0269]
[0270] The NMR studies carried out are detailed below.
[0271] Sample Information
[0272] The sample under study was provided to ReadGlead as a solid material. The material was stored at -20 °C upon receipt. Between two measurements, the prepared NMR samples were stored in the dark at 4 °C - 8 °C. The following sample information was provided:
[0273] Sample ID: P08-G05_c16
[0274] Single isotope mass: 1348.5604
[0275] Total molecular formula: C61H92N2O29S
[0276] Amount obtained: 5.13 mg
[0277] Solubility: 10 mg / mL DMSO solution after sonication
[0278] Materials and Methods
[0279] NMR Samples
[0280] PN102-62-01
[0281] The NMR sample was prepared by weighing 2.899 mg of sample P08 - G05_c16 in a screw - cap vial and adding 540 μL of DMSO - d 6 The sample was slowly dissolved and heated to 40 °C for 1 - 2 minutes, then sonicated for 3 × 10 seconds. Visually, the sample still showed traces of finely dispersed undissolved particles but was transferred to a 5 mm NMR tube.
[0282] PN102-62-01B
[0283] The NMR sample was prepared by adding 20 μL of D 2 2O to the NMR tube of the above - mentioned sample PN102 - 62 - 01.
[0284] PN102-62-01C
[0285] The NMR sample was prepared by adding 2 μL of TFA - d to the NMR tube of the above - mentioned sample PN102 - 62 - 01B.
[0286] PN102-62-02
[0287] The NMR sample was prepared by adding 540 μL of CD 3 3OD directly to an Eppendorf tube containing the residue of P08 - G05_c16 (about 2.2 mg). The sample was slowly dissolved and heated to 40 °C for 1 - 2 minutes, then sonicated for 3 × 10 seconds. Visually, the sample still contained a large amount of undissolved material, but the supernatant was transferred to a 5 mm NMR tube.
[0288] Chemicals and Materials
[0289] Equipment :
[0290] Mettler Toledo MT5 Balance
[0291] Bandelin Sonorex Ultrasonic Cleaner, Model RK 31
[0292] Agilent 2 mL Clear Screw Neck Vial, Part Number 5190 - 9062
[0293] Agilent Technologies Screw Cap, 9 mm, with PTFE / Silicone Septum, Part Number 5190 - 9068
[0294] Hilgenberg Standard NMR Tube, 5 mm in diameter, Article Number 2001745
[0295] Hilgenberg Sealing Cap for NMR Tube, 5 mm in diameter, Article Number 9400312
[0296] NMR Solvents and Chemicals
[0297]
[0298] NMR Spectra
[0299] 500 MHz Bruker Avance Neo Spectrometer equipped with a 5 mm iProbe BBF / H / D probe and 500 MHz Varian Inova Spectrometer equipped with a 5 mm 1 H / 13 C / 15 N triple resonance probe were used to perform NMR experiments. Data were recorded at 25 °C or 40 °C. The recorded spectra are listed in Table 4 below.
[0300] Table 4
[0301]
[0302] DMSO-d 6 (2.50 / 39.52 ppm) and CD 3 OD (3.31 / 49.00 ppm) solvent residual signals were used as 1 H and 13 C chemical shift references.
[0303] NMR data were processed and analyzed using MestreNova 12.0.1 (Mestrelab Research S.L.). The plug-in "NMRPredict" in MestReNova 12.0.1 was used and the solvent was set to DMSO-d6 The predictor "Mnovabest" is used to predict chemical shifts.
[0304] Results and Conclusions
[0305] NMR Experiments and Conditions
[0306] For the sample material dissolved in DMSO-d 6 or CD 3 OD, NMR experiments are performed on the obtained material. The NMR data are recorded on a 500 MHz Bruker Avance NMR spectrometer and a 500 MHz Varian Inova spectrometer.
[0307] For the sample material provided in DMSO-d 6 medium-quality 1D and 2D 1 H / 13 C / 1 5 N NMR spectral data are obtained. A large number of low-intensity signals are observed, which may involve structure-related impurities and / or minor conformational isomers of the main substance in the solution. The spectral region of the sugar moiety becomes complex due to signal overlap and signal broadening, such that in a few cases, it is not easy to observe 1 H- 13 C HSQC cross-peaks. Heating the DMSO-d 6 sample to 40 °C does not result in signal sharpening or only very slight signal sharpening, so all the data used in the structure analysis are recorded at 25 °C. However, acidifying the DMSO-d 6 sample with TFA-d causes some signals to sharpen significantly, and the chemical shifts of the sugar and sugar derivatives also change (while the chemical shifts of the hydrocarbon tails are basically unaffected).
[0308] CD 3 The overall appearance of the CD 6 OD spectrum is slightly clearer and sharper than that of the DMSO-d 6 spectrum. However, due to solubility limitations, the signal-to-noise ratio intensity in methanol is too low to obtain useful 2D long-range and through-space NMR data, which are crucial for the structure identification process. In the few cases where the DMSO-d 3 data are ambiguous, the CD
[0309] A total of four different NMR data sets are used for structure analysis. For completeness, except for the hydrocarbon tail "F" with very similar chemical shifts in these two samples, for the DMSO-d 6 sample (PN102-62-01) and DMSO-d 6For the / TFA-d sample (PN102-62-01C), chemical shift assignments of the proposed structure were reported. In addition to the major compound, the sample also appears to contain a large number of unassigned small molecules.
[0310] Elucidation of Chemical Structure and Atom-Specific Assignments
[0311] The proposed structure of P08-G05_c16 ( Figure 7 ) has several structural features in common with the related moenomycin compounds, namely that they all contain a substituted tetrasaccharide linked to a hydrocarbon tail. However, in contrast to moenomycin, P08-G05_c16 lacks the linked phosphodiester, and the hydrocarbon tail contains 30 rather than 25 carbon atoms. The structural evidence for P08-G05_c16 is strong because, with the exception of the amide nitrogen of sugar unit "C" and the carbonyl carbon of the 2,3-dihydroxypropionic acid unit "E", essentially all 1 H / 13 C / 15 N atoms were observed and assigned, and the atomic connectivity is in complete agreement with the 2D data obtained. In addition, a reasonably good degree of agreement was observed between the experimental and predicted chemical shift values.
[0312] The sugar units designated as "A" and "D" are assigned as uronic acids, and the sugar units designated as "B" and "C" are assigned as N-acetyl-glucosamine. Stereospecific assignments are not included in this study. The connectivity of the sugar moieties was determined by correlations between the anomeric proton signals in the HMBC spectrum and the corresponding carbon signals via the glycosidic bond (C-4 of sugars "B" and "C" and C-2 of sugar "D"), and / or NOE correlations between the anomeric proton signals and the corresponding protons in the next sugar moiety. The connectivity between sugar "D" and the linked 2,3-dihydroxypropionic acid "E" was confirmed by the NOE correlation between the anomeric proton of "D" and the methylene proton of "E" (observed only in the acidified sample PN102-62-01C). The connectivity between "E" and the hydrocarbon tail "F" was also established by the NOE observed between both the CH proton and the CH2 proton of "E" and the two closest CH protons of "F" (atom numbers 47 and 48 in Table 6).
[0313] The assigned chemical shifts of P08-G05_c16 together with the corresponding chemical shifts predicted from the proposed chemical structure are given in Tables 5 and 6 below. The overall situation is that the predicted chemical shift values fully support the proposed molecular structure, and only a small deviation from the predicted values was observed in the "D" moiety.
[0314] Based on the sum of the proposed molecular formula, it is assumed that there is a sulfate substituent on one of the sugar oxygens. There are several available positions for this, namely the sugar positions where no OH protons were detected and no other substituents / bonds were determined. Due to C-3 and C-6 of13 The C chemical shifts in part “B” and part “C” are very similar, so it is considered that it is unlikely that either of these positions bears a sulfate group. Therefore, the C-4 position of unit “A”, the C-3 position of unit “D” and the C-4 position of unit “D” are retained as reasonable candidate positions. The position of the sulfate group cannot be determined solely based on NMR data. However, since ring system “D” is more sensitive to pH changes and there are deviations from the predicted chemical shift values, it seems to have a higher structural complexity than ring system “A” and is thus considered a more reasonable choice. It can be expected that O-sulfation will cause a slight downfield shift of the O-sulfated carbon and the proton bound to it, so the C-4 position of unit “D” is tentatively assigned.
[0315] Table 5
[0316] The chemical shift (δ) values of P08-G05_c16 in parts A-D in Figure 8 were predicted using “Mnova Predict” and determined experimentally in DMSO-d 2 after adding D 6 O and TFA-d (NMR sample PN102-62-01C) and in DMSO-d 6 (NMR sample PN102-62-01) and DMSO-d 1 H / 13 C solvent residual signals (2.50 / 39.52 ppm) were reported, and an indirect reference was applied to 15N. The data were recorded at 25 °C. NO = not observed.
[0317]
[0318] No 15th amide nitrogen in sugar “C” was observed from the recorded 1 H- 15 N HSQC data. This is an expected result due to the broadening of the relevant amide proton signal observed in the 1D 1 H spectrum data. Nevertheless, due to the characteristic chemical shifts of the amide proton and the adjacent C-2 carbon, as well as the overall consistency between parts “B” and “C” and the expected total molecular formula, the structural assignment of acetylated amino sugar “C” is very likely. Similarly, no 85th carbonyl carbon in part “E” was observed from the recorded 1 H- 13 C HMBC data, which is expected due to the broadening observed for the adjacent 75th methine. Due to the predicted and experimental 13The C chemical shift values are very close, and this structural motif is known to exist in related moenomycins. Considering the overall molecular formula sum, the proposed "E" substructure remains possible. The high number of quaternary carbon atoms observed in fragment "F" and the splitting of the methylene proton signals explain the cyclic subunit, which is also consistent with the total number of rings and double bonds expected from the proposed molecular formula sum. The structural unit "F" deviates significantly from the structure of moenomycin and has not been evaluated from a biosynthetic perspective.
[0319] Structural parts "D" and "E" showed significantly broadened 1 H signals in DMSO-d6, and the cross-peaks in the 1 H- 13 C HSQC data were broadened to the point of being unrecognizable. After adding TFA-d to the sample, the signals became sharpened, allowing for 13 C chemical shift assignments, but the data also revealed a tendency for these signals to double - the sharpening effect in this region is consistent with the presence of protonated carboxylic acid moieties after acidification. The origin behind these observations has not been explored in this study, and only the main signals were evaluated in the structure elucidation.
[0320] Table 6
[0321] The chemical shift (δ) values of P08 - G05_c16 in the E - F part in Figure 8 were predicted using "Mnova Predict" and determined experimentally in DMSO-d 6 (NMR sample PN102 - 62 - 01). *The shifts reported in DMSO-d 2 after adding D 6 O and TFA-d (NMR sample PN102 - 62 - 01C). The experimental shifts relative to the solvent residual signals of 1 H / 13 C (2.50 / 39.52 ppm) are reported, and an indirect reference is applied to 15 N. The data were recorded at 25 °C. NO = Not observed.
[0322]
[0323] Example 8
[0324] Determination of Position of Sulfate Group in Nidaromycin
[0325] Background:
[0326] ReadGlead has determined the structure of Nidaromycin, the active compound produced by P08-G05_c16. However, there is some uncertainty regarding the position of the sulfate group (SO4- group). Here we use MSMS fragmentation followed by computational simulated fragmentation with the aim of determining the position of the SO4- group in Nidaromycin.
[0327] Based on the proposed molecular formula sum, it is assumed that there is a sulfate substituent on one of the sugar oxygens. There are several available positions for this, namely, sugar positions where no OH proton was detected and no other substituents / bonds were determined. Since the 13C chemical shifts of C-3 and C-6 13 are very similar in part "B" and part "C", it is considered unlikely that either of these positions bears a sulfate. Thus, the C-4 position of unit "A", the C-3 position of unit "D", and the C-4 position of unit "D" are retained as reasonable candidates. The position of the sulfate cannot be determined solely based on NMR data. However, since ring system "D" is more sensitive to changes in pH and since there are deviations from the predicted chemical shift values, it seems to have a higher structural complexity than ring system "A" and is therefore considered a more reasonable choice. It can be expected that O-sulfation would result in a slight downfield shift of the O-sulfated carbon and the proton bound to it, and thus the C-4 position of unit "D" is tentatively assigned.
[0328] These three possible structures are given in SMILES:
[0329] Nidaromycin A4
[0330] C / C(C)=C\CC1CCC(C)( / C=C / C2=CCC(C)C(C)(C\C=C(\C) / C=C / OC(COC3OC(C(=O)O)C(O)C(O)C3OC3OC(CO)C(OC4OC(CO)C(OC5OC(C(=O)O)C(OS(=O)(=O)O)C(O)C5O)C(O)C4 / N=C(\C)O)C(O)C3\N=C( / C)O)C(=O)O)C2=C)C1(C)C
[0331] Nidaromycin D3.
[0332] CC(CC=C1 / C=C / C2(C)C(C)(C)C(CC=C(C)C)CC2)C(C)(C / C=C(\C) / C=C / OC(COC(C(C(C2O)OS(O)(=O)=O)OC(C(C3O)NC(C)=O)OC(CO)C3OC(C(C3O)NC(C)=O)OC(CO)C3OC(C(C(C3O)O)O)OC3C(O)=O)OC2C(O)=O)C(O)=O)C1=C
[0333] Nidaromycin D4:
[0334] CC1CC=C(\C=C\C2(C)CCC(CC=C(C)C)C2(C)C)C(=C)C1(C)C\C=C( / C)\C=C\OC(COC 1OC(C(OS(O)(=O)=O)C(O)C1OC1OC(CO)C(OC2OC(CO)C(OC3OC(C(O)C(O)C3O)C(O)=O)C(O)C2NC(C)=O)C(O)C1NC(C)=O)C(O)=O)C(O)=O
[0335] Materials and methods:
[0336] LC-MS method: The cell-free extract was analyzed using an Agilent LC-DAD system connected to a Bruker Impact II QTOF. LC was performed with 10 mM ammonium acetate buffer [mobile phase A] and 90:10 acetonitrile: water containing 10 mM ammonium acetate [mobile phase B]. The gradient was 5% B for 2 minutes, then 5% - 100% B for 2 - 25 minutes. MS was carried out in the electrospray ionization positive ion mode with the following MS parameters: mass range 100 - 1800, spectral rate: 12 Hz, absolute threshold: 25 counts, fragmentation threshold: 100 counts, capillary voltage: 4500 V, endplate offset: 500 V, drying gas: 10 L / min, nebulizer: 31.9 psi, drying temperature: 220 °C, precursor ion list: 1000 - 1500, data acquisition control: dynamic MSMS or fixed MSMS, collision energy: 5 V, CID: acqCtr+MultiCe, MultiCe20.
[0337] In silico fragmentation. In silico fragmentation was performed using MetFrag (Schymanski et al., 2015, Anal Bioanal Chem, Vol. 407, No. 21: pp. 6237–6255).
[0338] Results:
[0339] Research fragment: Computer simulation of fragmentation strongly suggests that the SO4- group is located at D3 or D4 rather than A4. As shown in Table 7, if the SO4- group were in A4, several fragments should not form. Additionally, it is difficult to distinguish between D3 and D4 because these structures typically produce the same fragments. Moreover, the SO4- group is often lost during fragmentation, and only a few low-abundance fragments contain the SO4- group. However, we observed a fragment (M+H = 398.1976) that can be explained by D3 but not by D4. This indicates that the SO4- group is located at D3, but only one matching fragment supports this.
[0340] To support the QTOF data, data from FT-ICR MSMS fragmentation at negative ionization was examined. Even the FT-ICR fragmentation pattern could not distinguish between positions D3 and D4. However, several fragments could only be explained by the sulfate group in position D3 or D4 and not by position A4 (Table 5).
[0341] Table 7
[0342] The fragments obtained from MSMS fragmentation using a Bruker Impact II QTOF were compared with the computer simulation of fragmentation of the three proposed structures. Several fragments could not be explained by the structure with the SO4- group in the A4 position.
[0343]
[0344] Table 8
[0345] The fragments obtained from MSMS fragmentation using a Bruker FT-ICR were compared with the computer simulation of fragmentation of the three proposed structures. The fragments shown here could not be explained by the structure with the SO4- group in the A4 position.
[0346]
Claims
1. A compound of formula (I): wherein R 1 is -SO 2 OH, -SO 2 OR or -SO 2 R and R 2 is H, or wherein R 2 is -SO 2 OH, -SO 2 OR or -SO 2 R and R 1 is H; wherein R is C 1 -C 20 hydrocarbyl; where each R 3 is independently selected from H or C 1 -C 20 hydrocarbyl; or a pharmaceutically acceptable salt, solvate or hydrate thereof.
2. The compound according to claim 1, wherein the compound has the following structure: wherein R 1 is -SO 2 OH, -SO 2 OR or -SO 2 R and R 2 is H, or wherein R 2 is -SO 2 OH, -SO 2 OR or -SO 2 R and R 1 is H; wherein R is C 1 -C 20 hydrocarbyl; or a pharmaceutically acceptable salt, solvate or hydrate thereof.
3. The compound according to claim 1 or 2, wherein R 1 is -SO 2 OH or -SO 2 OR and R 2 is H, or wherein R 2 is -SO 2 OH or -SO 2 OR and R 1 is H; wherein R is C 1 -C 20 hydrocarbyl; Preferably, R 1 is -SO 2 OH and R 2 is H, or wherein R 2 is -SO 2 OH and R 1 is H.
4. A compound according to any one of claims 1 to 3, wherein the compound has the following structure: or a pharmaceutically acceptable salt, solvate or hydrate thereof.
5. The compound according to any one of claims 1 to 3, wherein the compound has the following structure: or a pharmaceutically acceptable salt, solvate or hydrate thereof.
6. A nucleic acid molecule comprising: (a) the nucleotide sequence shown in SEQ ID NO.1; or (b) a nucleotide sequence complementary to SEQ ID NO.1; or (c) a nucleotide sequence degenerate with SEQ ID NO.1; or (d) a nucleotide sequence having at least 85% sequence identity with SEQ ID NO.1; or (e) a part of any one of (a) to (d); wherein the nucleic acid molecule encodes one or more polypeptides or is complementary to a nucleic acid molecule encoding the one or more polypeptides, or contains one or more genetic elements or is complementary to a nucleic acid molecule containing one or more genetic elements, and has functional activity in the synthesis of antibiotic compounds.
7. The nucleic acid molecule according to claim 6, wherein the compound is defined according to any one of claims 1 to 5.
8. The nucleic acid molecule according to claim 6 or claim 7, wherein the molecule encodes a biosynthetic system for synthesizing the compound.
9. The nucleic acid molecule according to claim 6 or claim 7, wherein the molecule: (i) comprises the nucleotide sequence shown in any one or more of SEQ ID NOs. 2 - 29, or a nucleotide sequence complementary or degenerate to any one or more of SEQ ID NOs. 2 - 29, or a nucleotide sequence having at least 85% sequence identity with any one or more of SEQ ID NOs. 2 - 29; or (ii) comprises a nucleotide sequence encoding the amino acid sequence shown in any one or more of SEQ ID NOs. 30 - 57, or a nucleotide sequence encoding an amino acid having at least 85% sequence identity with any one or more of SEQ ID NOs. 30 - 57.
10. A polypeptide encoded by the nucleic acid molecule defined according to any one of claims 6 to 9.
11. A recombinant construct comprising the nucleic acid molecule defined according to any one of claims 6 to 9.
12. A vector comprising the nucleic acid molecule defined according to any one of claims 6 to 9 or the construct defined according to claim 11.
13. A microbial host cell comprising the nucleic acid molecule, recombinant construct or vector defined according to any one of claims 6 to 9, 11 or 12.
14. The host cell according to claim 13, wherein the host cell is a production host cell for producing the polypeptide or antibiotic compound according to claim 10 and is an actinomycete.
15. The host cell according to claim 13 or claim 14, wherein the host cell is: (i) of the genus Streptomyces; (ii) Streptomyces coelicolor; (iii) Streptomyces coelicolor strain M145 (ATCC BAA - 471); (iv) Streptomyces coelicolor strain M1152, which is a derivative of (ii) and contains the modification ΔactΔredΔcpkΔcdarpoB (C1298T); or (v) Streptomyces coelicolor strain M1152ΔmatAB, which is a derivative of (iii) and further contains a deletion of the locus matAB.
16. A method for producing an antibiotic compound, the method comprising introducing a nucleic acid molecule, construct or vector as defined in any one of claims 8, 9 or 12 into a microbial host cell, expressing the nucleic acid molecule, and synthesizing the compound through the expressed biosynthetic system.
17. The method according to claim 16, wherein the host cell is as defined in claim 15.
18. The method according to claim 16 or claim 17, wherein the method further comprises recovering the compound.
19. The method according to any one of claims 16 to 18, wherein the method further comprises purifying the compound.
20. A compound obtainable or obtained by the method according to any one of claims 16 to 19.
21. The compound according to any one of claims 1 to 5 or claim 20, which is used as a medicine.
22. The compound according to any one of claims 1 to 5 or claim 20, which is used as an antibiotic medicine.
23. The compound according to claim 21 or claim 22, which is used as an antibiotic medicine against Gram - positive bacteria.
24. The compound for use according to claim 23, wherein the bacterium is Staphylococcus aureus or Enterococcus faecium, including their antibiotic - resistant strains.
25. A pharmaceutical composition, the pharmaceutical composition comprising a compound according to any one of claims 1 to 5 or 20, and further comprising at least one carrier, additive and / or excipient.
26. Use of the compound according to any one of claims 1 to 5 or claim 20 as an antibacterial agent.
27. The use according to claim 26, wherein the compound is used as an antibacterial agent for plants.
28. A method for preparing a nucleic acid molecule encoding a modified biosynthetic system for synthesizing a modified derivative of a compound as defined in any one of claims 1 to 5 or 20, the method comprising modifying a nucleic acid molecule as defined in any one of claims 6 to 9, optionally, wherein the nucleic acid molecule is modified by introducing, mutating, deleting, substituting or inactivating a sequence encoding one or more activities or proteins encoded by the nucleic acid molecule.
Citation Information
Patent Citations
Gene cluster encoding a nystatin polyketide synthase and its manipulation and utility
WO2001059126A2
NRPS-PKS gene cluster and its manipulation and utility
WO2009115822A1