Heterologous biosynthesis of noshizolide

By using polynucleotides encoding enzymes from the NA biosynthetic pathway for heterologous expression, the challenges of producing nodulisporic acid are addressed, achieving efficient production of NA and its precursors.

JP7688973B2Active Publication Date: 2025-06-05VICTORIA LINK LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2020518002
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-09-29
Filing Date
2018-09-28
Publication Date
2025-06-05
Estimated Expiration
2038-09-28

AI Technical Summary

Technical Problem

The production of nodulisporic acid (NA) is challenging due to the difficulty in biosynthesizing it from its natural producer, Hypoxylon pulicicidum, and the inefficiency of existing chemical synthesis methods, which have not achieved complete synthesis of NA.

Method used

The development of polynucleotides encoding enzymes involved in the NA biosynthetic pathway, such as NodW, NodR, NodX, NodM, NodB, and others, which can be used for heterologous expression in a permissive host to produce NA and its precursors.

Benefits of technology

This approach enables the production of useful amounts of NA and its precursors, overcoming the limitations of traditional biosynthesis and chemical synthesis methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007688973000025
    Figure 0007688973000025
  • Figure 0007688973000026
    Figure 0007688973000026
  • Figure 0007688973000027
    Figure 0007688973000027
Patent Text Reader

Abstract

Nodulisporic acid (NA) comprises a group of indole diterpenes known for their potent insecticidal activity; however, NA biosynthesis by the natural producer, Hypoxylon pulicicidum (a Nodulisporium species), is extremely difficult to achieve. Identification of genes involved in NA production may enable biosynthetic pathway optimization, providing access to NA for commercial applications. Because it is difficult to obtain useful quantities of NA using published fermentation methods, gene knockout studies are not desirable as a method for confirming gene function. Instead, heterologous expression of H. pulicicidum genes in a more robust host species, such as Penicillium paxilli, provides a method for rapidly identifying the function of genes that play a role in NA biosynthesis. In this study, we identified the functions of four secondary metabolism genes required for the biosynthesis of nodulisporic acid F (NAF) and reconstituted these genes in the genome of P. paxili to enable heterologous expression of NAF in this fungus.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to novel polypeptides that catalyze at least one biochemical reaction leading to the production of nodulisporic acid (NA), polynucleotides encoding such polypeptides, methods for producing such polypeptides and polynucleotides, and methods for producing at least one NA by heterologous expression in a permissive host using such polypeptides and polynucleotides.

Background Art

[0002] Filamentous fungi produce an interesting and diverse repertoire of useful compounds. One member of such a compound class, indole diterpenes (IDTs), is of particular interest due to their broad chemical diversity and associated biological activities, such as anti-MRSA 1 , anti-cancer 2、3 , anti-H1N1 4 , insecticidal 5 and antispasmodic 6 activities. NA (Figure 1) is a particularly bioactive group of quasipaspalin-like IDTs produced by Hypoxylon pulicicidum, which was previously classified as a Nodulisporium species. 7 . Nodulisporic acid A (NAA) 10 is particularly important because it exhibits very strong insecticidal activity against blood-sucking arthropods without showing observable side effects in mammals. 5、8 .

[0003] NA is particularly difficult to biosynthesize from its natural producer, H. pulicicidum. The reported NA biosynthesis method requires growing H. pulicicidum for 21 days in a very nutrient-rich medium in complete darkness. 9Since the biosynthesis of NAA 10 in H. prischidum is difficult, it is difficult to obtain useful amounts of NAA 10 using the published fermentation methods, and the production of commercial amounts of NAA 10 is essentially unachievable. Accordingly, attempts have been made to chemically synthesize NAA 10, and the synthetic mechanisms of nostoc-sporene F (NAF) 5a 10 and nostoc-sporene D 7a 11 have occurred, but the complete synthesis of NAA 10 has not been achieved 12 As a result, there is a need in the art for novel methods of NAA 10 synthesis and / or biosynthesis that provide useful amounts of NAA 10.

Summary of the Invention

Problems to be Solved by the Invention

[0004] It is an object of the present invention to provide a polynucleotide encoding at least one enzyme in the NAA 10 biosynthetic pathway of H. prischidum and / or, using such a vector, to provide a method for producing at least one indole diterpene compound that is NA and / or, in a heterologous host, to provide a method for producing a precursor of NAA 10 and / or, at least, to provide useful alternatives to the public.

[0005] As used herein, when reference is made to a patent specification, other external documents, or other information sources, this is generally for the purpose of providing a background for discussing the features of the present invention. Unless specifically stated otherwise, such reference to external documents is not to be construed as an admission that such documents, or such information sources, are prior art or form part of the common general knowledge in the art in any respect.

Means for Solving the Problems

[0006] In one aspect, the present invention relates to an isolated polypeptide comprising an amino acid sequence selected from the group consisting of NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodM (SEQ ID NO: 12), NodB (SEQ ID NO: 15), NodO (SEQ ID NO: 18), NodJ (SEQ ID NO: 21), NodC (SEQ ID NO: 24), NodY1 (SEQ ID NO: 27), NodD2 (SEQ ID NO: 30), NodD1 (SEQ ID NO: 33), NodY2 (SEQ ID NO: 36), NodZ (SEQ ID NO: 39), NodS (SEQ ID NO: 50), and NodI (SEQ ID NO: 56) or a functional variant or fragment thereof.

[0007] In another aspect, the present invention relates to an isolated polynucleotide encoding a polypeptide comprising an amino acid sequence selected from the group consisting of NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodM (SEQ ID NO: 12), NodB (SEQ ID NO: 15), NodO (SEQ ID NO: 18), NodJ (SEQ ID NO: 21), NodC (SEQ ID NO: 24), NodY1 (SEQ ID NO: 27), NodD2 (SEQ ID NO: 30), NodD1 (SEQ ID NO: 33), NodY2 (SEQ ID NO: 36), NodZ (SEQ ID NO: 39), NodS (SEQ ID NO: 50), and NodI (SEQ ID NO: 56) or a functional variant or fragment thereof.

[0008] In another aspect, the present invention relates to an isolated polynucleotide comprising at least 70% nucleic acid sequence identity to a nucleic acid sequence selected from the group consisting of nodW cDNA (SEQ ID NO: 2), nodW genomic DNA (SEQ ID NO: 1), nodR cDNA (SEQ ID NO: 5), nodR genomic DNA (SEQ ID NO: 4), nodX cDNA (SEQ ID NO: 8), nodX genomic DNA (SEQ ID NO: 7), nodM cDNA (SEQ ID NO: 11), nodM genomic DNA (SEQ ID NO: 10), nodB cDNA (SEQ ID NO: 14), nodB genomic DNA (SEQ ID NO: 13), nodO cDNA (SEQ ID NO: 17), nodO genomic DNA (SEQ ID NO: 16), nodJ cDNA (SEQ ID NO: 20), nodJ genomic DNA (SEQ ID NO: 19), nodC cDNA (SEQ ID NO: 23), nodC genomic DNA (SEQ ID NO: 22), nodY1 cDNA (SEQ ID NO: 26), nodY1 genomic DNA (SEQ ID NO: 25), nodD2 cDNA (SEQ ID NO: 29), nodD2 genomic DNA (SEQ ID NO: 28), nodD1 cDNA (SEQ ID NO: 32), nodD1 genomic DNA (SEQ ID NO: 31), nodY2 cDNA (SEQ ID NO: 35), nodY2 genomic DNA (SEQ ID NO: 34), nodZ cDNA (SEQ ID NO: 38), nodZ genomic DNA (SEQ ID NO: 37), nodS cDNA (SEQ ID NO: 49), nodS genomic DNA (SEQ ID NO: 48), nodI cDNA (SEQ ID NO: 55), and nodI genomic DNA (SEQ ID NO: 54).

[0009] In another aspect, the present invention relates to a transcription unit (TU) comprising at least one isolated polynucleotide described in the present invention. In another aspect, the present invention relates to a vector encoding the isolated polypeptide described in the present invention.

[0010] In another aspect, the present invention relates to a vector comprising the isolated nucleic acid sequence or TU described in the present invention. In another aspect, the present invention relates to an isolated host cell comprising the isolated polypeptide, isolated polynucleotide, TU and / or vector described in the present invention.

[0011] In another aspect, the present invention relates to a method for producing at least one NA comprising the step of heterologously expressing at least one polypeptide, isolated nucleic acid sequence, TU or vector described in the present invention in an isolated host cell.

[0012] In another aspect, the present invention relates to at least one NA produced by the method of the present invention. In another aspect, the present invention relates to an isolated polypeptide from a Hypoxylon species, or a functional fragment or variant thereof, that catalyzes a biochemical reaction in the biosynthetic pathway leading from 3-geranylgeranylindole (GGI) 2 to NAA 10.

[0013] In another aspect, the present invention relates to an isolated polynucleotide encoding at least one polypeptide from a Hypoxylon species, or a functional variant or fragment thereof, that catalyzes a biochemical reaction in the biosynthetic pathway leading from GGI 2 to NAA 10.

[0014] In another aspect, the present invention relates to a method for producing at least one Hypoxylon species polypeptide or a functional variant or fragment thereof comprising the step of heterologously expressing the isolated nucleic acid sequence or vector described in the present invention in an isolated host cell.

[0015] In another aspect, the present invention relates to a method for producing at least one NA comprising the step of heterologously expressing at least one polypeptide that catalyzes a biochemical reaction in the biosynthetic pathway leading from GGI 2 to NAA 10 in an isolated host cell.

[0016] In another aspect, the present invention relates to an isolated host cell that expresses at least one heterologous polypeptide that catalyzes the conversion of a substrate in the biosynthetic pathway leading to the formation of NAA 10 from GGI 2.

[0017] In another aspect, the present invention relates to an isolated host cell that produces at least one polypeptide involved in the biosynthetic pathway leading from GGI 2 to NAA 10 by heterologous expression.

[0018] In another aspect, the present invention relates to a method for producing at least one NA, comprising contacting a recombinant cell transformed with a nucleic acid that results in an increased activity level of a polypeptide selected from the group consisting of NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodM (SEQ ID NO: 12), NodB (SEQ ID NO: 15), NodO (SEQ ID NO: 18), NodJ (SEQ ID NO: 21), NodC (SEQ ID NO: 24), NodY1 (SEQ ID NO: 27), NodD2 (SEQ ID NO: 30), NodD1 (SEQ ID NO: 33), NodY2 (SEQ ID NO: 36), NodZ (SEQ ID NO: 39), NodS (SEQ ID NO: 50), and NodI (SEQ ID NO: 56) or a functional variant or fragment thereof with a carbohydrate containing a substrate such that the substrate is metabolized to at least one NA, compared to the cell before transformation.

[0019] In another aspect, the present invention relates to an isolate of Hypoxylon prischidum comprising at least one heterologous nucleic acid sequence encoding an enzyme in the biosynthetic pathway leading to NAA 10.

[0020] In another aspect, the present invention relates to an isolate of Hypoxylon prischidum that expresses at least two different GGPPS enzymes. In another aspect, the present invention relates to an isolate of Hypoxylon prischidum comprising a genetic modification that leads to an increase in the biosynthesis of NAA 10.

[0021] In another aspect, the present invention relates to a method for producing NAA 10, comprising the step of expressing at least one heterologous nucleic acid sequence in Hypoxylon prischidum, wherein the at least one heterologous nucleic acid sequence encodes an enzyme in the biosynthetic pathway leading to NAA 10.

[0022] The various aspects of the different aspects of the present invention discussed above are also shown in the following detailed description of the present invention, but the present invention is not limited to such description. Other aspects of the present invention may become apparent from the following description, which is provided by way of example only, and from the accompanying figures.

Brief Description of the Drawings

[0023] The present invention is hereby referred to with reference to the images of the accompanying figures.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Mode for Carrying Out the Invention

[0024] Definitions The term "comprising" means, in this specification and the claims, "consisting at least in part of"; that is, when interpreting references in this specification and the claims that include "comprising", in each reference, it is meant that all of the features preceded by this term must be present, but other features may also be present. Related terms, such as "comprise" and "comprised", shall be interpreted in a similar manner.

[0025] The term "consisting essentially of" means, in this specification, materials or steps that are specified, and that do not substantially affect the basic and novel characteristics (singular or plural) of the claimed invention.

[0026] The term "consisting of" means, in this specification, the materials or steps specified in the claimed invention, excluding any element, step, or component not specified in the claims. The terms "recognition site" and "restriction site" are used interchangeably in this specification and mean the same thing. These terms, as used in this specification in relation to restriction enzymes, mean a nucleic acid sequence or polynucleotide sequence that defines a binding site on a polynucleotide of a given restriction enzyme.

[0027] The term "indole diterpene (IDT) compound" or "indole diterpenoid" refers to any compound derived from an indole-containing precursor, preferably indole-3-glycerol phosphate 1b, and geranylgeranyl pyrophosphate (GGPP) 1a.

[0028] In some embodiments, the IDT compound is selected from the group consisting of GGI 2, emindole SB 4a, and NAF 5a. The term "gene construct" refers to a polynucleotide molecule, usually double-stranded DNA, that is conjugated to another polynucleotide molecule. In one non-limiting example, a gene construct is produced by inserting a first polynucleotide molecule into a second polynucleotide molecule by restriction / ligation, such as is known in the art. In some embodiments, the gene construct includes, but is not limited to, a single polynucleotide module, at least two polynucleotide modules, or a number of polynucleotide modules assembled into a single continuous polynucleotide molecule (also referred to herein as a "multi-gene construct").

[0029] The term "gene construct" refers to a polynucleotide molecule, usually double-stranded DNA, that is conjugated to another polynucleotide molecule. In one non-limiting example, a gene construct is produced by inserting a first polynucleotide molecule into a second polynucleotide molecule by restriction / ligation, such as is known in the art. In some embodiments, the gene construct includes, but is not limited to, a single polynucleotide module, at least two polynucleotide modules, or a number of polynucleotide modules assembled into a single continuous polynucleotide molecule (also referred to herein as a "multi-gene construct").

[0030] A gene construct may contain the elements necessary to enable transcription of a polynucleotide molecule and, optionally, the elements necessary to translate the transcript into a polypeptide. The polynucleotide molecule contained in and / or by the gene construct may be derived from the host cell, or may be derived from a different cell or organism, and / or may be a recombinant polynucleotide. Once inside the host cell, the gene construct may integrate into the host chromosomal DNA. The gene construct may be linked to a vector.

[0031] As used herein, the term “transcription unit” (TU) refers to a polynucleotide encoding a single RNA molecule, including all nucleotide sequences necessary for the transcription of a single RNA molecule, including but not limited to a promoter, an RNA coding sequence, and a terminator.

[0032] As used herein, the term “transcription unit module” (TUM) refers to a polynucleotide containing a nucleotide sequence that encodes a single RNA molecule or a portion thereof; or encodes a protein coding sequence (CDS) or a portion thereof; or encodes a regulatory sequence element or a portion thereof that controls the transcription of the RNA molecule; or encodes a regulatory sequence element or a portion thereof that controls the translation of the CDS. Such regulatory sequence elements may include, but are not limited to, a promoter, an untranslated region (UTR), a terminator, a polyadenylation signal, a ribosome binding site, a transcription enhancer, and a translation enhancer.

[0033] As used herein, the term “multigene construct” means a gene construct that is a polynucleotide containing at least two TUs. As used herein, the term “marker” means a nucleic acid sequence in a polynucleotide that encodes a selectable marker or a scoreable marker.

[0034] As used herein, the term "selectable marker" refers to a TU that, when introduced into a cell, confers at least one trait on the cell and enables the selection of the cell based on the presence or absence of that trait. In one embodiment, cells are selected based on their ability to survive under conditions that kill cells that do not contain at least one selectable marker.

[0035] As used herein, the term "scorable marker" refers to a TU that, when introduced into a cell, confers at least one trait on the cell and enables the scoring of the cell based on the presence or absence of that trait. In one embodiment, cells containing the TU are scored by phenotypically identifying the cells from a plurality of cells.

[0036] As used herein, the term "gene element" refers to any polynucleotide sequence that is not a TU or does not form part of a TU. Such polynucleotide sequences may include, but are not limited to, origins of replication of plasmids and viruses, centromeres, telomeres, repetitive sequences, sequences used for homologous recombination, site-specific recombination sequences, and sequences that control DNA transfer between organisms.

[0037] As used herein, the term "vector" refers to any type of polynucleotide molecule that can be used to manipulate genetic material so that it can be amplified, replicated, manipulated, partially replicated, modified, and / or expressed, but is not limited thereto. In some embodiments, a vector may be used to transport a polynucleotide contained therein into a cell or organism.

[0038] As used herein, the term "source vector" refers to a vector capable of cloning a polynucleotide sequence of interest. In some embodiments, the polynucleotide sequence is a TU and a TUM as described herein. In some embodiments, the source vector is selected from the group consisting of plasmids, bacterial artificial chromosomes (BACs), phage artificial chromosomes (PACs), yeast artificial chromosomes (YACs), bacteriophages, phagemids, and cosmids. In some embodiments, a source vector containing a polynucleotide sequence of interest is referred to as an entry clone. In some embodiments, the entry clone may function as a shuttle or destination vector for accepting additional polynucleotide sequences.

[0039] As used herein, the term "shuttle vector" refers to a vector in which a polynucleotide sequence of interest can be cloned and from which these sequences are operable. In some embodiments, the polynucleotide sequence is a TU and a TUM as described herein. In some embodiments, the shuttle vector is selected from the group consisting of plasmids, BACs, PACs, YACs, bacteriophages, phagemids, and cosmids. In some embodiments, a shuttle vector containing a polynucleotide sequence of interest may function as a destination vector for accepting additional polynucleotide sequences.

[0040] As used herein, the term "destination vector" refers to a vector capable of cloning a polynucleotide sequence of interest. In some embodiments, the polynucleotide sequence is a TU and a TUM as described herein. In some embodiments, the destination vector is selected from the group consisting of plasmids, BACs, PACs, YACs, bacteriophages, phagemids, and cosmids. In some embodiments, a destination vector comprising a polynucleotide sequence of interest is an entry clone. In some embodiments, an entry clone may function as a destination vector for accepting additional polynucleotide sequences.

[0041] As used herein, the term "polynucleotide(s)" means a single-stranded or double-stranded deoxyribonucleotide or ribonucleotide polymer of any length, and includes, by way of non-limiting example, coding and non-coding sequences of genes, sense and antisense sequences, exons, introns, genomic DNA, cDNA, pre-mRNA, mRNA, rRNA, siRNA, miRNA, tRNA, ribozymes, recombinant polynucleotides, isolated and purified naturally occurring DNA or RNA sequences, synthetic RNA and DNA sequences, nucleic acid probes, primers, fragments, gene constructs, vectors, and modified polynucleotides. References to nucleic acids, nucleic acid molecules, nucleotide sequences, and polynucleotide sequences are to be understood as being equivalent.

[0042] As used herein, the term "gene" refers to the biological unit of heredity, which is self-replicating and located at a defined position (locus) on a particular chromosome. In one embodiment, the particular chromosome is a eukaryotic or bacterial chromosome. The term, bacterial chromosome, is used interchangeably herein with the term, bacterial genome.

[0043] The term "gene cluster" as used herein refers to a group of genes located in close proximity to each other on the same chromosome, whose products play a coordinated role in specific aspects of primary or secondary metabolism of a cell. In one example, a gene cluster includes a CDC group that is involved in all of a series of biochemical reactions, including a biosynthetic pathway or array whose products produce a given metabolite, particularly a secondary metabolite.

[0044] The term "secondary metabolite" as used herein refers to a compound that is not involved in primary metabolism and thus is different from more widely present macromolecules such as proteins and nucleic acids that constitute the basic mechanisms of life.

[0045] The terms "under conditions where a single enzyme is active" and "under conditions where multiple enzymes are active", and their grammatical variations, when used with respect to enzyme activity, mean that the enzyme performs its expected function; for example, a restriction endonuclease cleaves a nucleic acid at an appropriate restriction site and a DNA ligase covalently links two polynucleotides together.

[0046] The term "endogenous" as used herein refers to a component of a cell, tissue or organism that is derived from or naturally produced within these. "Endogenous" components may be, but are not limited to, any component including polynucleotides, polypeptides including non-ribosomal polypeptides, fatty acids or polyketides.

[0047] The term "exogenous" as used herein refers to any component of a cell, tissue or organism that is not derived from or not naturally produced within these. An exogenous component may be, for example, a polynucleotide sequence introduced into a cell, tissue or organism, or a polypeptide expressed in a cell, tissue or organism from that polynucleotide sequence.

[0048] "Naturally occurring", as used herein with respect to the polynucleotide sequences described in the present invention, refers to the primary polynucleotide sequences found in nature. Synthetic polynucleotide sequences that are identical to wild-type polynucleotide sequences are considered to be naturally occurring sequences for the purposes of this disclosure. What is important with respect to naturally occurring polynucleotide sequences is that the actual sequence of nucleotide bases comprising the polynucleotide is found or known to be found in nature.

[0049] For example, wild-type polynucleotide sequences are naturally occurring polynucleotide sequences, but are not limited thereto. Naturally occurring polynucleotide sequences also refer to variant polynucleotide sequences that are found in nature and that differ from the wild-type. For example, allelic variants and naturally occurring recombinant polynucleotide sequences resulting from hybridization or horizontal gene transfer are included, but are not limited thereto.

[0050] "Non-naturally occurring", as used herein with respect to the polynucleotide sequences described in the present invention, refers to polynucleotide sequences that are not found in nature. Examples of non-naturally occurring polynucleotide sequences include, but are not limited to, artificially produced mutants and variant polynucleotide sequences created by, for example, point mutations, insertions, or deletions. Non-naturally occurring polynucleotide sequences also include chemically evolved sequences. What is important with respect to the non-naturally occurring polynucleotide sequences described in the present invention is that the actual sequence of nucleotide bases comprising the polynucleotide is not found and is not known to be found in nature.

[0051] The term "wild-type", as used herein when referring to a polynucleotide, refers to the naturally occurring; non-mutated form of the polynucleotide. A mutant polynucleotide means a polynucleotide that maintains a mutation known in the art, such as, but not limited to, a point mutation, insertion, deletion, substitution, amplification, or translocation.

[0052] As used herein, the term "wild-type" when used with respect to a polypeptide refers to the naturally occurring, non-mutated form of the polypeptide. A wild-type polypeptide is a polypeptide that can be expressed from a wild-type polynucleotide.

[0053] The term "coding sequence" or "open reading frame" (ORF) refers to the sense strand of a genomic DNA sequence or cDNA sequence that is capable of producing a transcript and / or polypeptide under the control of appropriate control sequences. A CDS is identified by the presence of a 5' translation initiation codon and a 3' translation stop codon. When inserted into a gene construct or expression cassette, a "coding sequence (CDS)" is capable of being expressed when ligated so as to be functional with a promoter sequence and / or other control elements.

[0054] "Ligated so as to be functional" means that the sequence to be expressed is placed under the control of a control element. "Control element" as used herein refers to any nucleic acid sequence element that controls or affects the expression of a polynucleotide insert from a vector, gene construct or expression cassette, and includes promoters, transcriptional control sequences, translational control sequences, origins of replication, tissue-specific control elements, temporal control elements, enhancers, polyadenylation signals, repressors and terminators. A control element may be "homologous" or "heterologous" to the polynucleotide insert to be expressed from a gene construct, expression cassette or vector as described herein. When a gene construct, expression cassette or vector as described herein is present in a cell, a control element may be "endogenous", "exogenous", "naturally occurring" and / or "non-naturally occurring" with respect to the cell.

[0055] The term "untranslated region" refers to the untranslated sequences upstream of the translation start site and downstream of the translation stop site. These sequences are also sometimes referred to as the 5’UTR and 3’UTR, respectively. These regions contain elements necessary for transcription initiation and termination, as well as for the control of translation efficiency.

[0056] A terminator is a sequence that terminates transcription and is found at the 3’untranslated end of a gene downstream of the sequence that is translated. Terminators are important determinants of mRNA stability and in some cases have been found to have a spatial control function.

[0057] The term "promoter" refers to a non-transcribed cis-regulatory element upstream of the coding region that controls the transcription of a polynucleotide sequence. A promoter contains a cis-initiator element that specifies the transcription start site and conserved boxes. In one non-limiting example, a bacterial promoter may include a "Pribnow box" (also known as the -10 region), and other motifs to which transcription factors bind and which promote transcription. A promoter may be homologous or heterologous with respect to the polynucleotide sequence to be expressed. When attempting to express a polynucleotide sequence in a cell, the promoter may be an endogenous or exogenous promoter. A promoter may be a constitutive promoter, an inducible promoter or a controllable promoter, such as are known in the art.

[0058] "Homologous", as used herein with respect to a polynucleotide control element, means a polynucleotide control element that is native and naturally occurring. A homologous polynucleotide control element may be linked to a polynucleotide of interest in a functional manner such that the polynucleotide of interest is expressible from the TU, genetic element or vector described in the present invention.

[0059] As used herein, "homologous" with respect to a polynucleotide or polypeptide in a host organism means that the polynucleotide or polypeptide is an endogenous and naturally occurring polynucleotide or polypeptide within that host organism. A homologous polynucleotide may be linked to homologous or heterologous control elements such that a homologous polypeptide is expressible from a TU, genetic element, or vector containing such a homologous polynucleotide as described herein.

[0060] As used herein, "introduced homologous" with respect to a polynucleotide or polypeptide in a host organism means that the polynucleotide or polypeptide is an endogenous and naturally occurring polynucleotide or polypeptide within that host organism that has been introduced into the organism by experimental techniques. An introduced homologous polynucleotide may be linked to homologous or heterologous control elements such that a homologous polypeptide is expressible from a TU, genetic element, or vector containing such a homologous polynucleotide as described herein.

[0061] As used herein, "heterologous" with respect to a polynucleotide control element means a polynucleotide control element that is not an endogenous and naturally occurring polynucleotide control element. A heterologous polynucleotide control element is not normally associated with a CDS to which it is operably linked. A heterologous control element may be operably linked to a polynucleotide of interest such that the polynucleotide of interest is expressible from a vector, gene construct, or expression cassette of the invention. Such promoters may include promoters normally associated with other genes, ORFs, or coding regions, and / or any other promoter isolated from any other bacterial, viral, eukaryotic, or mammalian cell.

[0062] "Heterologous", as used herein with respect to a polynucleotide or polypeptide in a host organism, means a polynucleotide or polypeptide that is not a polynucleotide or polypeptide that is native and naturally occurring in that host organism. A heterologous polynucleotide may be linked to heterologous or homologous regulatory elements such that a heterologous polypeptide is expressible from a TU, genetic element or vector comprising the heterologous polynucleotide as described herein and is functional with respect to such regulatory elements.

[0063] The terms "heterologously express" and "heterologous expression" mean the expression of a heterologous polypeptide in a host cell. "A biochemical reaction in the biosynthetic pathway leading from GGI 2 to NAA 10" means one of the specific reactions catalyzed by one of the specific enzymes involved in the conversion of the substrate molecule GGI 2 through the following intermediates: mono-expoxidized GGI 3a, emindole SB 4a, NAF 5a, NAE 6a, NAD 7a, NAC 8, NAB 9 to NAA 10, and does not include similar enzymes in the host cell that may have a similar function but do not act on the specific intermediates listed above.

[0064] A "functional variant or fragment thereof" of a polypeptide is a subsequence of the polypeptide that performs a function necessary for the biological activity or binding of that polypeptide and / or provides the three-dimensional structure of the polypeptide. The term may refer to a polypeptide, an aggregate of polypeptides, such as a dimer or other multimer, a fusion polypeptide, a polypeptide fragment, a polypeptide variant, or a functional polypeptide derivative thereof capable of performing the polypeptide activity.

[0065] As used herein, "isolated" describes a polynucleotide or polypeptide sequence that has been removed from its natural cellular environment. An isolated molecule may be obtained by any method or combination of methods known in the art and used, including biochemical, recombinant, and synthetic techniques. A polynucleotide or polypeptide sequence may be prepared by at least one purification step.

[0066] As used herein, "isolated" when used in reference to a cell or host cell describes a cell or host cell that has been obtained from or removed from an organism or its natural environment and is subsequently maintained in a laboratory environment as is known in the art. The term includes the single cell itself, as well as cells or host cells contained in a cell culture, and may include single cells or single host cells.

[0067] The term "isolated host cell" as used herein includes single cells of single-celled fungi, as well as hyphae and mycelia of filamentous fungi, including septate and aseptate forms, with respect to fungal host cells. As used herein, the term "recombinant" refers to a polynucleotide sequence that has been removed from the sequences that surround it in its natural context and / or has been recombined with a sequence that does not exist in its natural context. A "recombinant" polypeptide sequence is produced by translation from a "recombinant" polynucleotide sequence.

[0068] As used herein, the term "variant" refers to a polynucleotide or polypeptide sequence in which one or more nucleotide or amino acid residues have been deleted, substituted, or added, as compared to a specifically identified sequence. A variant may be a naturally occurring allelic variant or a non-naturally occurring variant. A variant may be from the same species or from another species and may include homologs, paralogs, and orthologs. In certain embodiments, variants of a polypeptide useful in the present invention have the same or similar biological activity as the corresponding wild-type molecule; i.e., the parent polypeptide or polynucleotide.

[0069] In certain embodiments, variants of the polypeptides described herein have biological activities that are similar or substantially similar to the corresponding wild-type molecules. In certain embodiments, the similarity is similar activity and / or binding specificity.

[0070] In certain embodiments, variants of the polypeptides described herein have biological activities that are different from the corresponding wild-type molecules. In certain embodiments, the difference is altered activity and / or binding specificity.

[0071] The term "variant" with respect to polynucleotides and polypeptides includes all types of polynucleotides and polypeptides as defined herein. The variant polynucleotide sequences exhibit at least 50%, at least 60%, preferably at least 70%, preferably at least 71%, preferably at least 72%, preferably at least 73%, preferably at least 74%, preferably at least 75%, preferably at least 76%, preferably at least 77%, preferably at least 78%, preferably at least 79%, preferably at least 80%, preferably at least 81%, preferably at least 82%, preferably at least 83%, preferably at least 84%, preferably at least 85%, preferably at least 86%, preferably at least 87%, preferably at least 88%, preferably at least 89%, preferably at least 90%, preferably at least 91%, preferably at least 92%, preferably at least 93%, preferably at least 94%, preferably at least 95%, preferably at least 96%, preferably at least 97%, preferably at least 98%, and preferably at least 99% identity to the sequences of the present invention. The identity is found over a comparison window of at least 8 nucleotides, preferably at least 10 nucleotides, preferably at least 15 nucleotides, preferably at least 20 nucleotides, preferably at least 27 nucleotides, preferably at least 40 nucleotides, preferably at least 50 nucleotides, preferably at least 60 nucleotides, preferably at least 70 nucleotides, preferably at least 80 nucleotides, preferably over the full length of the polynucleotides used or identified according to the methods of the present invention.

[0072] Polynucleotide variants also include those that are likely to retain the functional equivalence of these sequences and that show similarity to one or more of the specifically identified sequences that could not reasonably be expected to occur by random chance.

[0073] Polynucleotide sequence identity and similarity can be readily determined by those skilled in the art. Variant polynucleotides also include polynucleotides that differ from the polynucleotide sequences described herein but, as a result of the degeneracy of the genetic code, encode a polypeptide having an activity similar to that of the polypeptide encoded by the polynucleotides of the present invention. Sequence modifications that do not change the amino acid sequence of the polypeptide are "silent mutations". Except for ATG (methionine) and TGG (tryptophan), other codons for the same amino acid may be changed by techniques recognized in the art, for example, to optimize codon expression in a particular host organism.

[0074] Polynucleotide sequence modifications that result in conservative substitutions of one or several amino acids in the encoded polypeptide sequence without significantly modifying its biological activity are also included in the present invention. Those skilled in the art know methods for making amino acid substitutions that are phenotypically silent (see, for example, Bowie et al., 1990, Science 247, 1306).

[0075] The term "variant" with respect to a polypeptide also includes naturally occurring, recombinant, and synthetically produced polypeptides. Variant polypeptide sequences preferably exhibit at least 35%, preferably at least 40%, preferably at least 50%, preferably at least 60%, preferably at least 70%, preferably at least 71%, preferably at least 72%, preferably at least 73%, preferably at least 74%, preferably at least 75%, preferably at least 76%, preferably at least 77%, preferably at least 78%, preferably at least 79%, preferably at least 80%, preferably at least 81%, preferably at least 82%, preferably at least 83%, preferably at least 84%, preferably at least 85%, preferably at least 86%, preferably at least 87%, preferably at least 88%, preferably at least 89%, preferably at least 90%, preferably at least 91%, preferably at least 92%, preferably at least 93%, preferably at least 94%, preferably at least 95%, preferably at least 96%, preferably at least 97%, preferably at least 98%, and preferably at least 99% identity to the sequences of the present invention. Identity is found over a comparison window of at least 2 amino acid positions, preferably at least 3 amino acid positions, preferably at least 4 amino acid positions, preferably at least 5 amino acid positions, preferably at least 7 amino acid positions, preferably at least 10 amino acid positions, preferably at least 15 amino acid positions, preferably at least 20 amino acid positions, preferably over the full length of the polypeptide being used or identified according to the methods of the present invention.

[0076] Polypeptide variants also include those that are likely to retain the functional equivalence of these sequences and that show similarity to one or more specifically identified sequences that could not reasonably be expected to occur by random chance.

[0077] Polypeptide sequence identity and similarity can be readily determined by those skilled in the art. Variant polypeptides include polypeptides whose amino acid sequences differ from the polypeptides of the present specification by one or more conservative amino acids, or non-conservative substitutions, deletions, additions or insertions that do not affect the biological activity of the peptide.

[0078] Conservative substitutions typically involve the substitution of one amino acid for another with similar properties, such as within the following groups: valine, glycine; glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid; asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine.

[0079] Analysis of evolutionary biological sequences has shown that, at least in part, reflecting the differences between conservative and non-conservative substitutions at the biological level, not all sequence changes are equally likely. For example, certain amino acid substitutions are often likely to occur, while others are very rare. Evolutionary changes or substitutions at amino acid residues can be modeled by a scoring matrix, also called a substitution matrix. Such matrices are used in bioinformatics analysis to identify relationships between sequences and are known to those skilled in the art.

[0080] Other variants include peptides containing modifications that affect peptide stability. Such analogs may, for example, contain one or more non-peptide bonds (substituting for peptide bonds) in the peptide sequence. Also included are analogs containing residues other than the naturally occurring L-amino acids, such as D-amino acids or non-naturally occurring synthetic amino acids, such as beta or gamma amino acids and cyclic analogs.

[0081] Substitutions, deletions, additions or insertions may be made by mutagenesis methods known in the art. Those skilled in the art know methods for making amino acid substitutions that are phenotypically silent. See, for example, Bowie et al., 1990, Science 247, 1306.

[0082] Polypeptides as used herein may also refer to polypeptides that are modified during or after synthesis, for example, by biotinylation, benzylation, glycosylation, phosphorylation, amidation, derivatization with blocking groups / protecting groups, etc. Such modifications may increase the stability or activity of the polypeptide.

[0083] The terms "regulate the expression of", "regulated expression" and "expression regulation" of a polynucleotide or polypeptide are intended to include situations where the genomic cDNA corresponding to the polynucleotide to be expressed according to the present invention is modified and thus leads to the regulated expression of the polynucleotide or polypeptide of the present invention. The modification of genomic DNA may be through other methods known in the art for inducing genetic transformation or mutation. "Regulated expression" may be related to an increase or decrease in the amount of messenger RNA and / or polypeptide produced, and may also result in an increase or decrease in the activity of the polypeptide due to modification of the sequences of the polynucleotide and polypeptide produced.

[0084] The terms "regulate the activity of", "regulated activity" and "activity regulation" of a polynucleotide or polypeptide are intended to include situations where the genomic cDNA corresponding to the polynucleotide to be expressed according to the present invention is modified and thus leads to the regulated expression of the polynucleotide of the present invention, or the regulated expression or activity of the polypeptide. The modification of genomic DNA may be through other methods known in the art for inducing genetic transformation or mutation. "Regulated activity" may be related to an increase or decrease in the amount of messenger RNA and / or polypeptide produced, and may also result in an increase or decrease in the activity of the polypeptide due to modification of the sequences of the polynucleotide and polypeptide produced.

[0085] References to a range of numbers disclosed herein (e.g., 1 to 10) also incorporate references to all relevant numbers within that range (e.g., 1, 1.1, 2, 3, 3.9, 4, 5, 6, 6.5, 7, 8, 9, and 10) and also to any range of rational numbers within that range (e.g., 2 to 8, 1.5 to 5.5, and 3.1 to 4.7), and thus all sub-ranges of all ranges explicitly disclosed herein are explicitly disclosed. These are merely exemplary and are not intended to be limiting, and all possible combinations of numerical values between the recited minimum and maximum values are to be considered as explicitly recited herein in a similar manner.

[0086] Detailed Description Identification of the Biosynthesis of IDT-Paxilline 6b in Penicillium paxilli 6、13-16 Since then, the gene functionality in seven other IDT biosynthetic pathways has been elucidated. 17-22 These IDT pathways share homologous genes encoding the enzymes for the first three steps in IDT biosynthesis (Figure 2): (I) geranylgeranyl pyrophosphate synthase (GGPPS) converts farnesyl pyrophosphate and isopentenyl pyrophosphate to GGPP 1a, (II) geranylgeranyl transferase (GGT) catalyzes the indole condensation of GGPP 1a and indole-3-glycerol phosphate 1b to produce GGI 2, and (III) site-selective flavin adenine dinucleotide (FAD)-dependent epoxidase generates single and / or double epoxidized GGI products 3a / 3b. In a fourth enzymatic step involving IDT cyclization, the pathway diverges into four important branches, yielding mono / dioxygenated antimycotoxin-derived cyclic cores such as emindole SB 4a and paspaline 4b, or mono / dioxygenated mycotoxin-derived cyclic cores such as aflatrem 4c and emindole DB 4d. These cyclic cores are often further modified by decorating enzymes to generate the biological activity diversity seen across IDT.

[0087] NA is a bioactive IDT produced by Hypoxylon pricincium, and NAA 10 is particularly important because it has strong insecticidal activity against blood-sucking arthropods and lacks mammalian toxicity. 5、8 . However, as described herein, the production of NAA 10 by direct synthesis has not been achieved, and the biosynthesis production of this compound in useful amounts, even at small-scale commercial levels, would be difficult if not impossible.

[0088] Accordingly, the present invention generally relates to a series of isolated genes from the fungus, H. pricincium, which, when combined, form a gene cluster that mediates the production of NA, and the use of such gene clusters to direct the heterologous expression of NA in isolated host cells, preferably in isolated fungal cells. Using a recently developed technique for manipulating gene sequences, called the modular iterated DNA assembly system (MIDAS), the inventors have reconstituted the biosynthetic pathway for NAF 5a from H. pricincium in an alternative fungal host, Penicillium paxilli. The MIDAS platform and methods using the MIDAS platform are described herein and in related patent application AU2017903955, which is incorporated herein by reference in its entirety.

[0089] The present inventors analyzed the genomic sequence of H. prischidum and identified a cluster containing 15 predicted coding sequences (CDSs) predicted to encode enzymes required for the biosynthesis of NAA10: nodW cDNA (SEQ ID NO: 2) and genomic DNA (SEQ ID NO: 1), nodR cDNA (SEQ ID NO: 5) and genomic DNA (SEQ ID NO: 4), nodX cDNA (SEQ ID NO: 8) and genomic DNA (SEQ ID NO: 7), nodM cDNA (SEQ ID NO: 11) and genomic DNA (SEQ ID NO: 10), nodB cDNA (SEQ ID NO: 14) and genomic DNA (SEQ ID NO: 13), nodO cDNA (SEQ ID NO: 17) and genomic DNA (SEQ ID NO: 16), nodJ cDNA (SEQ ID NO: 20) and genomic DNA (SEQ ID NO: 19), nodC cDNA (SEQ ID NO: 23) and genomic DNA (SEQ ID NO: 22), nodY1 cDNA (SEQ ID NO: 26) and genomic DNA (SEQ ID NO: 25), nodD2 cDNA (SEQ ID NO: 29) and genomic DNA (SEQ ID NO: 28), nodD1 cDNA (SEQ ID NO: 32) and genomic DNA (SEQ ID NO: 31), nodY2 cDNA (SEQ ID NO: 35) and genomic DNA (SEQ ID NO: 34), nodZ cDNA (SEQ ID NO: 38) and genomic DNA (SEQ ID NO: 37), nodS cDNA (SEQ ID NO: 49) and genomic DNA (SEQ ID NO: 48), nodI cDNA (SEQ ID NO: 55) and genomic DNA (SEQ ID NO: 54).

[0090] The boundaries of this cluster were determined by identifying adjacent genes with high similarity and synthetic machinery compared to equivalent genomic loci in another Hypoxylon strain that does not produce nostrilic acid. The details of these predicted genes in the cluster and their proposed functions are shown in Table 1. Seven of the cluster genes are homologous to those found in other IDT biosynthetic gene clusters (Tables 2 - 6). The protein products of the seven predicted genes homologous to IDT biosynthetic genes from other fungi have at least 35% amino acid identity to their homologs in the PAX cluster of P. paxilli, the JAN cluster of P. janthinellum, and / or the PEN cluster of P. crustosum, and this includes GGT (NodC (SEQ ID NO: 24)), two FAD-dependent oxidases (NodM (SEQ ID NO: 12)) and NodO (SEQ ID NO: 18)), the IDT cyclase (NodB (SEQ ID NO: 15)), two prenyltransferases (NodD2 (SEQ ID NO: 30), and NodD1 (SEQ ID NO: 33)), and one cytochrome P450 oxygenase (NodR (SEQ ID NO: 6)). The other seven putative ORFs were predicted to encode four cytochrome P450 oxygenases (NodW (SEQ ID NO: 3), NodX (SEQ ID NO: 9), NodJ (SEQ ID NO: 21), and NodZ (SEQ ID NO: 39)), a paralogous FAD-dependent oxygenase pair (NodY1 (SEQ ID NO: 27), and NodY2 (SEQ ID NO: 36)), and two gene products (NodS (SEQ ID NO: 50), and NodI (SEQ ID NO: 56)) that may be involved in the biosynthesis of NA of unknown function. The TER gene cluster from Chaunopycnis alba (Tolypocladium album) involved in terpendol biosynthesis 17Similarly, NOD clusters do not appear to contain secondary metabolite-specific GGPPS genes. In particular, the inventors have identified only one GGPPS-encoding gene in the genome of H. prischidum, and the amino acid sequence of its predicted protein product, its exon / intron structure, and its location outside the identified cluster strongly suggest that the gene is involved in a primary metabolic function similar to ggs1 in P. paxilli. 23 .

[0091] To confirm the function of the gene products and directly establish their respective roles in NAA 10 biosynthesis, the inventors constructed a series of plasmids harboring these genes in various combinations, which were then transformed into appropriate P. paxilli hosts (Table 7) for the heterologous production of NAA 10 precursors. Thus, the CDSs of the H. prischidum genes of interest were amplified (see Table 8 for primers) and cloned into the MIDAS level 1 destination vector, pML1 (Table 9). At the MIDAS level 2, the cloned CDSs were placed under the control of heterologous promoter (ProUTR) and transcription terminator (UTRterm) modules to generate full-length TUs (Table 10), which were then used to generate multi-gene plasmids (Table 11). The inventors used the P. paxilli knockout strain repertoire (Table 7) to perform functional complementation and pathway reconstruction to determine the function of the genes in the NOD cluster. After transformation of the P. paxilli host with the multi-gene plasmid, the inventors first determined the chemical phenotypes of the transformants by normal-phase thin-layer chromatography (TLC, results not shown) of the fungal extracts and subsequently by reverse-phase liquid chromatography-mass spectrometry (LC-MS) analysis. The inventors purified the newly expressed metabolites as determined by semi-preparative reverse-phase high-performance liquid chromatography (HPLC) by high-resolution mass spectrometry (HRMS) and subjected the compounds to nuclear magnetic resonance (NMR) spectroscopy ( 1 H, 13 C, and HSQC, HMBC, COSY) for final identification.

[0092] Using this methodology, the inventors identified that nodC (cDNA (SEQ ID NO: 23) and genomic DNA (SEQ ID NO: 22)) is a functional ortholog of paxC (cDNA (SEQ ID NO: 43) and genomic DNA (SEQ ID NO: 42)), and that NodC (SEQ ID NO: 24) mediates the production of GGI 2, the second step in IDT biosynthesis in H. priscusidum (Figure 3, Trace i.b, Figure 4). NodC (SEQ ID NO: 24) shares 52.3% amino acid sequence identity with PaxC (SEQ ID NO: 44) from P. paxilli (Table 2). The inventors also identified that nodM (cDNA (SEQ ID NO: 11) and genomic DNA (SEQ ID NO: 10)) is a homolog of paxM (cDNA (SEQ ID NO: 46) and genomic DNA (SEQ ID NO: 45)), and that NodM (SEQ ID NO: 12) is a GGI 2 monoepoxidase that catalyzes the production of monoepoxidized GGI 3a (Figure 3, Trace ii.b, Figure 5). NodM (SEQ ID NO: 12) shares 48.6% sequence identity with PaxM (SEQ ID NO: 47) from P. paxilli (Table 3). In particular, the inventors showed that NodM (SEQ ID NO: 12) is different from PaxM (SEQ ID NO: 47), which is a mono- or diepoxidase, and is specifically a monoepoxidase. The inventors further found that NodB (SEQ ID NO: 15) from H. priscusidum acts as an IDT cyclase that cyclizes the monoepoxidized GGI product 3a to form emindole SB 4a (Figure 3, Trace iii.b, Figure 6), and further confirmed that nodB (cDNA (SEQ ID NO: 14) and genomic DNA (SEQ ID NO: 13)) is a functional ortholog of paxB (cDNA (SEQ ID NO: 52) and genomic DNA (SEQ ID NO: 51)) from P. paxilli (Figure 3, Trace iv.b, Figure 7). NodB (SEQ ID NO: 15) from H. priscusidum shares 63% identity with PaxB (SEQ ID NO: 53) from P. paxilli (Table 4).

[0093] The dedicated NAA 10-core is NAF 5a, which is generated by the oxidation of the terminal methyl carbon of emindole SB 4a, C5” (Figure 8B). Thus, the inventors confirmed the identity of this oxidase as a P450 oxygenase encoded by H. precidum nodW (cDNA (SEQ ID NO: 2) and genomic DNA (SEQ ID NO: 1)) by co-expression of nodM (cDNA (SEQ ID NO: 11) and genomic DNA (SEQ ID NO: 10)) and nodW (cDNA (SEQ ID NO: 2) and genomic DNA (SEQ ID NO: 1)) within the paxM (cDNA (SEQ ID NO: 46) and genomic DNA (SEQ ID NO: 45)) deletion background (PN2257) that results in the production of NAF 5a (Figure 3, trace v.b, Figure 9).

[0094] To confirm that only five genes are essential for the production of NAF 5a in P. paxilli and to establish that none of the other P. paxilli IDT genes from the PAX cluster contribute to NAF 5a production in the paxM KO strain (PN2257), the inventors assembled a multi-gene construct containing paxG (genomic DNA (SEQ ID NO: 40)) from P. paxilli, as well as nodC (genomic DNA (SEQ ID NO: 22)), nodM (genomic DNA (SEQ ID NO: 10)), nodB (genomic DNA (SEQ ID NO: 13)), and nodW (genomic DNA (SEQ ID NO: 1)) from H. precidum. This multi-gene construct was transformed into the P. paxilli PAX gene cluster knockout strain (PN2250). As expected, the expression of the five genes indicated that NAF 5a was produced (Figure 10), and it was shown that these five genes are indeed required for the biosynthesis of NFA 5a.

[0095] Based on the research described in this specification, the inventors disclose the use of heterologous expression to identify the first five steps for producing the NAA 10-core compound, NAF 5a. In particular, the inventors have confirmed the functions of four previously unknown genes from H. prischidum: nodC (cDNA (SEQ ID NO: 23) and genomic DNA (SEQ ID NO: 22)), nodM (cDNA (SEQ ID NO: 11) and genomic DNA (SEQ ID NO: 10)), nodB (cDNA (SEQ ID NO: 14) and genomic DNA (SEQ ID NO: 13)), and nodW (cDNA (SEQ ID NO: 2) and genomic DNA (SEQ ID NO: 1)), and have discovered a second filamentous fungal species, H. prischidum, which appears to lack the secondary metabolism GGPPS gene but is still capable of producing IDT. Without wishing to be bound by theory, the inventors believe that H. prischidum relies on its primary metabolism GGPPS to provide GGPP for IDT synthesis. The lack of secondary metabolism GGPPS may explain why H. prischidum produces such a small amount of NA. The low amount of NA produced by H. prischidum is difficult both in terms of elucidating the details of biosynthesis and for the use of the compound or its derivatives. Using efficient gene recombination of MIDAS and heterologous expression in P. paxilli, the inventors have overcome both of these problems. Furthermore, the inventors have demonstrated that P. paxilli has much more favorable growth conditions and is a suitable host for heterologous expression studies, enabling the functions of genes to be confirmed more quickly and easily than would be possible if the inventors had relied on the biosynthetic mechanism of H. prischidum.

[0096] Elucidation of the biosynthetic pathway for the heterologous production of NAF 5a in P. paxilli provides a reasonable expectation of success in fully identifying the gene products from H. prischidum that are involved in the "decoration" steps leading to the production of fully functionalized NAA 10. This reasonable expectation stems from the inventors' identification of the nucleic acid CDSs expected to encode the enzymes necessary for prenylation to form nostryporic acid E 6a and for oxidation to form nostryporic acid D 7a. Each of these nucleic acid sequences was identified by inference from the H. prischidum biosynthetic gene cluster described herein. Collectively, the inventors' studies described herein confirm that heterologous expression of IDT genes in heterologous hosts is a viable method that can be used to produce natural products that would otherwise be difficult to obtain and that can be useful across many industries.

[0097] Polypeptide Accordingly, in one aspect, the present invention relates to an isolated polypeptide comprising an amino acid sequence selected from the group consisting of SEQ ID NO: NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodM (SEQ ID NO: 12), NodB (SEQ ID NO: 15), NodO (SEQ ID NO: 18), NodJ (SEQ ID NO: 21), NodC (SEQ ID NO: 24), NodY1 (SEQ ID NO: 27), NodD2 (SEQ ID NO: 30), NodD1 (SEQ ID NO: 33), NodY2 (SEQ ID NO: 36), NodZ (SEQ ID NO: 39), NodS (SEQ ID NO: 50), and NodI (SEQ ID NO: 56) or a functional variant or fragment thereof.

[0098] Preferably, the functional variant or fragment thereof has at least 70%, preferably at least 75%, preferably at least 80%, preferably at least 85%, preferably at least 90%, preferably at least 95%, preferably at least 99% amino acid sequence identity to an amino acid sequence selected from the group consisting of SEQ ID NO: NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodM (SEQ ID NO: 12), NodB (SEQ ID NO: 15), NodO (SEQ ID NO: 18), NodJ (SEQ ID NO: 21), NodC (SEQ ID NO: 24), NodY1 (SEQ ID NO: 27), NodD2 (SEQ ID NO: 30), NodD1 (SEQ ID NO: 33), NodY2 (SEQ ID NO: 36), NodZ (SEQ ID NO: 39), NodS (SEQ ID NO: 50), and NodI (SEQ ID NO: 56).

[0099] Preferably, the isolated polypeptide comprising NodW (SEQ ID NO: 3) or a functional variant or fragment thereof has oxygenase activity, preferably cytochrome P450 oxygenase activity.

[0100] Preferably, the isolated polypeptide comprising NodR (SEQ ID NO: 6) or a functional variant or fragment thereof has oxygenase activity, preferably cytochrome P450 oxygenase activity.

[0101] Preferably, the isolated polypeptide comprising NodX (SEQ ID NO: 9) or a functional variant or fragment thereof has oxygenase activity, preferably cytochrome P450 oxygenase activity.

[0102] Preferably, the isolated polypeptide comprising NodM (SEQ ID NO: 12) or a functional variant or fragment thereof has oxygenase activity, preferably FAD-dependent oxygenase activity.

[0103] Preferably, the isolated polypeptide comprising NodB (SEQ ID NO: 15) or a functional variant or fragment thereof has cyclase activity, preferably IDT cyclase activity. Preferably, an isolated polypeptide comprising NodO (SEQ ID NO: 18) or a functional variant or fragment thereof has oxidase activity, preferably FAD-dependent oxidase activity.

[0104] Preferably, an isolated polypeptide comprising NodJ (SEQ ID NO: 21) or a functional variant or fragment thereof has oxidase activity, preferably cytochrome P450 oxidase activity.

[0105] Preferably, an isolated polypeptide comprising NodC (SEQ ID NO: 24) or a functional variant or fragment thereof has transferase activity, preferably GGT activity. Preferably, an isolated polypeptide comprising NodY1 (SEQ ID NO: 27) or a functional variant or fragment thereof has oxidase activity, preferably FAD-dependent oxidase activity.

[0106] Preferably, an isolated polypeptide comprising NodD2 (SEQ ID NO: 30) or a functional variant or fragment thereof has transferase activity, preferably prenyltransferase activity.

[0107] Preferably, an isolated polypeptide comprising NodD1 (SEQ ID NO: 33) or a functional variant or fragment thereof has transferase activity, preferably prenyltransferase activity.

[0108] Preferably, an isolated polypeptide comprising NodY2 (SEQ ID NO: 36) or a functional variant or fragment thereof has oxidase activity, preferably FAD-dependent oxidase activity.

[0109] Preferably, an isolated polypeptide comprising NodZ (SEQ ID NO: 39) or a functional variant or fragment thereof has oxidase activity, preferably cytochrome P450 oxidase activity.

[0110] In one aspect, the isolated polypeptide comprises SEQ ID NO: NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodM (SEQ ID NO: 12), NodB (SEQ ID NO: 15), NodO (SEQ ID NO: 18), NodJ (SEQ ID NO: 21), NodC (SEQ ID NO: 24), NodY1 (SEQ ID NO: 27), NodD2 (SEQ ID NO: 30), NodD1 (SEQ ID NO: 33), NodY2 (SEQ ID NO: 36), NodZ (SEQ ID NO: 39), NodS (SEQ ID NO: 50), and NodI (SEQ ID NO: 56) or a functional variant or fragment thereof.

[0111] In one aspect, the isolated polypeptide consists essentially of SEQ ID NO: NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodM (SEQ ID NO: 12), NodB (SEQ ID NO: 15), NodO (SEQ ID NO: 18), NodJ (SEQ ID NO: 21), NodC (SEQ ID NO: 24), NodY1 (SEQ ID NO: 27), NodD2 (SEQ ID NO: 30), NodD1 (SEQ ID NO: 33), NodY2 (SEQ ID NO: 36), NodZ (SEQ ID NO: 39), NodS (SEQ ID NO: 50), and NodI (SEQ ID NO: 56) or a functional variant or fragment thereof.

[0112] In one aspect, the isolated polypeptide consists of SEQ ID NO: NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodM (SEQ ID NO: 12), NodB (SEQ ID NO: 15), NodO (SEQ ID NO: 18), NodJ (SEQ ID NO: 21), NodC (SEQ ID NO: 24), NodY1 (SEQ ID NO: 27), NodD2 (SEQ ID NO: 30), NodD1 (SEQ ID NO: 33), NodY2 (SEQ ID NO: 36), NodZ (SEQ ID NO: 39), NodS (SEQ ID NO: 50), and NodI (SEQ ID NO: 56) or a functional variant or fragment thereof.

[0113] Polynucleotide In another aspect, the present invention relates to an isolated polynucleotide encoding a polypeptide comprising an amino acid sequence selected from the group consisting of SEQ ID NO: NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodM (SEQ ID NO: 12), NodB (SEQ ID NO: 15), NodO (SEQ ID NO: 18), NodJ (SEQ ID NO: 21), NodC (SEQ ID NO: 24), NodY1 (SEQ ID NO: 27), NodD2 (SEQ ID NO: 30), NodD1 (SEQ ID NO: 33), NodY2 (SEQ ID NO: 36), NodZ (SEQ ID NO: 39), NodS (SEQ ID NO: 50), and NodI (SEQ ID NO: 56) or a functional variant or fragment thereof.

[0114] Preferably, the functional variant or fragment thereof comprises at least 70%, preferably at least 75%, preferably at least 80%, preferably at least 85%, preferably at least 90%, preferably at least 95%, preferably at least 99% amino acid sequence identity to an amino acid sequence selected from the group consisting of SEQ ID NO: NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodM (SEQ ID NO: 12), NodB (SEQ ID NO: 15), NodO (SEQ ID NO: 18), NodJ (SEQ ID NO: 21), NodC (SEQ ID NO: 24), NodY1 (SEQ ID NO: 27), NodD2 (SEQ ID NO: 30), NodD1 (SEQ ID NO: 33), NodY2 (SEQ ID NO: 36), NodZ (SEQ ID NO: 39), NodS (SEQ ID NO: 50), and NodI (SEQ ID NO: 56) or a functional variant or fragment thereof.

[0115] Preferably, the isolated polynucleotide encodes a polypeptide comprising NodW (SEQ ID NO: 3) or a functional variant or fragment thereof having oxygenase activity, preferably cytochrome P450 oxygenase activity.

[0116] Preferably, the isolated polynucleotide encodes a polypeptide comprising NodR (SEQ ID NO: 6) or a functional variant or fragment thereof having oxygenase activity, preferably cytochrome P450 oxygenase activity.

[0117] Preferably, the isolated polynucleotide encodes a polypeptide comprising NodX (SEQ ID NO: 9) or a functional variant or fragment thereof, which has oxygenase activity, preferably cytochrome P450 oxygenase activity.

[0118] Preferably, the isolated polynucleotide encodes a polypeptide comprising NodM (SEQ ID NO: 12) or a functional variant or fragment thereof, which has oxygenase activity, preferably FAD-dependent oxygenase activity.

[0119] Preferably, the isolated polynucleotide encodes a polypeptide comprising NodB (SEQ ID NO: 15) or a functional variant or fragment thereof, which has cyclase activity, preferably IDT cyclase activity.

[0120] Preferably, the isolated polynucleotide encodes a polypeptide comprising NodO (SEQ ID NO: 18) or a functional variant or fragment thereof, which has oxygenase activity, preferably FAD-dependent oxygenase activity.

[0121] Preferably, the isolated polynucleotide encodes a polypeptide comprising NodJ (SEQ ID NO: 21) or a functional variant or fragment thereof, which has oxygenase activity, preferably cytochrome P450 oxygenase activity.

[0122] Preferably, the isolated polynucleotide encodes a polypeptide comprising NodC (SEQ ID NO: 24) or a functional variant or fragment thereof, which has transferase activity, preferably GGT activity.

[0123] Preferably, the isolated polynucleotide encodes a polypeptide comprising NodY1 (SEQ ID NO: 27) or a functional variant or fragment thereof, which has oxygenase activity, preferably FAD-dependent oxygenase activity.

[0124] Preferably, the isolated polynucleotide encodes a polypeptide comprising NodD2 (SEQ ID NO: 30) or a functional variant or fragment thereof, which has transferase activity, preferably prenyltransferase activity.

[0125] Preferably, the isolated polynucleotide encodes a polypeptide comprising NodD1 (SEQ ID NO: 33) or a functional variant or fragment thereof, which has transferase activity, preferably prenyltransferase activity.

[0126] Preferably, the isolated polynucleotide encodes a polypeptide comprising NodY2 (SEQ ID NO: 36) or a functional variant or fragment thereof, which has oxidase activity, preferably FAD-dependent oxidase activity.

[0127] Preferably, the isolated polynucleotide encodes a polypeptide comprising NodZ (SEQ ID NO: 39) or a functional variant or fragment thereof, which has oxidase activity, preferably cytochrome P450 oxidase activity.

[0128] In one embodiment, the isolated polynucleotide encodes a polypeptide comprising NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodM (SEQ ID NO: 12), NodB (SEQ ID NO: 15), NodO (SEQ ID NO: 18), NodJ (SEQ ID NO: 21), NodC (SEQ ID NO: 24), NodY1 (SEQ ID NO: 27), NodD2 (SEQ ID NO: 30), NodD1 (SEQ ID NO: 33), NodY2 (SEQ ID NO: 36), NodZ (SEQ ID NO: 39), NodS (SEQ ID NO: 50), and NodI (SEQ ID NO: 56) or a functional variant or fragment thereof.

[0129] In one embodiment, the isolated polynucleotide encodes a polypeptide consisting essentially of NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodM (SEQ ID NO: 12), NodB (SEQ ID NO: 15), NodO (SEQ ID NO: 18), NodJ (SEQ ID NO: 21), NodC (SEQ ID NO: 24), NodY1 (SEQ ID NO: 27), NodD2 (SEQ ID NO: 30), NodD1 (SEQ ID NO: 33), NodY2 (SEQ ID NO: 36), NodZ (SEQ ID NO: 39), NodS (SEQ ID NO: 50), and NodI (SEQ ID NO: 56) or a functional variant or fragment thereof.

[0130] In one embodiment, the isolated polynucleotide encodes a polypeptide consisting of NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodM (SEQ ID NO: 12), NodB (SEQ ID NO: 15), NodO (SEQ ID NO: 18), NodJ (SEQ ID NO: 21), NodC (SEQ ID NO: 24), NodY1 (SEQ ID NO: 27), NodD2 (SEQ ID NO: 30), NodD1 (SEQ ID NO: 33), NodY2 (SEQ ID NO: 36), NodZ (SEQ ID NO: 39), NodS (SEQ ID NO: 50), and NodI (SEQ ID NO: 56) or a functional variant or fragment thereof.

[0131] In another aspect, the present invention relates to an isolated polynucleotide comprising at least 70% nucleic acid sequence identity to a nucleic acid sequence selected from the group consisting of nodW cDNA (SEQ ID NO: 2), nodW genomic DNA (SEQ ID NO: 1), nodR cDNA (SEQ ID NO: 5), nodR genomic DNA (SEQ ID NO: 4), nodX cDNA (SEQ ID NO: 8), nodX genomic DNA (SEQ ID NO: 7), nodM cDNA (SEQ ID NO: 11), nodM genomic DNA (SEQ ID NO: 10), nodB cDNA (SEQ ID NO: 14), nodB genomic DNA (SEQ ID NO: 13), nodO cDNA (SEQ ID NO: 17), nodO genomic DNA (SEQ ID NO: 16), nodJ cDNA (SEQ ID NO: 20), nodJ genomic DNA (SEQ ID NO: 19), nodC cDNA (SEQ ID NO: 23), nodC genomic DNA (SEQ ID NO: 22), nodY1 cDNA (SEQ ID NO: 26), nodY1 genomic DNA (SEQ ID NO: 25), nodD2 cDNA (SEQ ID NO: 29), nodD2 genomic DNA (SEQ ID NO: 28), nodD1 cDNA (SEQ ID NO: 32), nodD1 genomic DNA (SEQ ID NO: 31), nodY2 cDNA (SEQ ID NO: 35), nodY2 genomic DNA (SEQ ID NO: 34), nodZ cDNA (SEQ ID NO: 38), nodZ genomic DNA (SEQ ID NO: 37), nodS cDNA (SEQ ID NO: 49), nodS genomic DNA (SEQ ID NO: 48), nodI cDNA (SEQ ID NO: 55), and nodI genomic DNA (SEQ ID NO: 54).

[0132] Preferably, the isolated polynucleotide comprises at least 75%, preferably at least 80%, preferably at least 85%, preferably at least 90%, preferably at least 95%, preferably at least 99% nucleic acid sequence identity to a nucleic acid sequence selected from the group consisting of SEQ ID NO: nodW cDNA (SEQ ID NO: 2), nodW genomic DNA (SEQ ID NO: 1), nodR cDNA (SEQ ID NO: 5), nodR genomic DNA (SEQ ID NO: 4), nodX cDNA (SEQ ID NO: 8), nodX genomic DNA (SEQ ID NO: 7), nodM cDNA (SEQ ID NO: 11), nodM genomic DNA (SEQ ID NO: 10), nodB cDNA (SEQ ID NO: 14), nodB genomic DNA (SEQ ID NO: 13), nodO cDNA (SEQ ID NO: 17), nodO genomic DNA (SEQ ID NO: 16), nodJ cDNA (SEQ ID NO: 20), nodJ genomic DNA (SEQ ID NO: 19), nodC cDNA (SEQ ID NO: 23), nodC genomic DNA (SEQ ID NO: 22), nodY1 cDNA (SEQ ID NO: 26), nodY1 genomic DNA (SEQ ID NO: 25), nodD2 cDNA (SEQ ID NO: 29), nodD2 genomic DNA (SEQ ID NO: 28), nodD1 cDNA (SEQ ID NO: 32), nodD1 genomic DNA (SEQ ID NO: 31), nodY2 cDNA (SEQ ID NO: 35), nodY2 genomic DNA (SEQ ID NO: 34), nodZ cDNA (SEQ ID NO: 38), nodZ genomic DNA (SEQ ID NO: 37), nodS cDNA (SEQ ID NO: 49), nodS genomic DNA (SEQ ID NO: 48), nodI cDNA (SEQ ID NO: 55), and nodI genomic DNA (SEQ ID NO: 54).

[0133] In one embodiment, the isolated polynucleotide comprises a nucleic acid sequence selected from the group consisting of nodW cDNA (SEQ ID NO: 2), nodW genomic DNA (SEQ ID NO: 1), nodR cDNA (SEQ ID NO: 5), nodR genomic DNA (SEQ ID NO: 4), nodX cDNA (SEQ ID NO: 8), nodX genomic DNA (SEQ ID NO: 7), nodM cDNA (SEQ ID NO: 11), nodM genomic DNA (SEQ ID NO: 10), nodB cDNA (SEQ ID NO: 14), nodB genomic DNA (SEQ ID NO: 13), nodO cDNA (SEQ ID NO: 17), nodO genomic DNA (SEQ ID NO: 16), nodJ cDNA (SEQ ID NO: 20), nodJ genomic DNA (SEQ ID NO: 19), nodC cDNA (SEQ ID NO: 23), nodC genomic DNA (SEQ ID NO: 22), nodY1 cDNA (SEQ ID NO: 26), nodY1 genomic DNA (SEQ ID NO: 25), nodD2 cDNA (SEQ ID NO: 29), nodD2 genomic DNA (SEQ ID NO: 28), nodD1 cDNA (SEQ ID NO: 32), nodD1 genomic DNA (SEQ ID NO: 31), nodY2 cDNA (SEQ ID NO: 35), nodY2 genomic DNA (SEQ ID NO: 34), nodZ cDNA (SEQ ID NO: 38), nodZ genomic DNA (SEQ ID NO: 37), nodS cDNA (SEQ ID NO: 49), nodS genomic DNA (SEQ ID NO: 48), nodI cDNA (SEQ ID NO: 55), and nodI genomic DNA (SEQ ID NO: 54).

[0134] In one embodiment, the isolated polynucleotide consists essentially of a nucleic acid sequence selected from the group consisting of nodW cDNA (SEQ ID NO: 2), nodW genomic DNA (SEQ ID NO: 1), nodR cDNA (SEQ ID NO: 5), nodR genomic DNA (SEQ ID NO: 4), nodX cDNA (SEQ ID NO: 8), nodX genomic DNA (SEQ ID NO: 7), nodM cDNA (SEQ ID NO: 11), nodM genomic DNA (SEQ ID NO: 10), nodB cDNA (SEQ ID NO: 14), nodB genomic DNA (SEQ ID NO: 13), nodO cDNA (SEQ ID NO: 17), nodO genomic DNA (SEQ ID NO: 16), nodJ cDNA (SEQ ID NO: 20), nodJ genomic DNA (SEQ ID NO: 19), nodC cDNA (SEQ ID NO: 23), nodC genomic DNA (SEQ ID NO: 22), nodY1 cDNA (SEQ ID NO: 26), nodY1 genomic DNA (SEQ ID NO: 25), nodD2 cDNA (SEQ ID NO: 29), nodD2 genomic DNA (SEQ ID NO: 28), nodD1 cDNA (SEQ ID NO: 32), nodD1 genomic DNA (SEQ ID NO: 31), nodY2 cDNA (SEQ ID NO: 35), nodY2 genomic DNA (SEQ ID NO: 34), nodZ cDNA (SEQ ID NO: 38), nodZ genomic DNA (SEQ ID NO: 37), nodS cDNA (SEQ ID NO: 49), nodS genomic DNA (SEQ ID NO: 48), nodI cDNA (SEQ ID NO: 55), and nodI genomic DNA (SEQ ID NO: 54).

[0135] In one embodiment, the isolated polynucleotide consists of a nucleic acid sequence selected from the group consisting of nodW cDNA (SEQ ID NO: 2), nodW genomic DNA (SEQ ID NO: 1), nodR cDNA (SEQ ID NO: 5), nodR genomic DNA (SEQ ID NO: 4), nodX cDNA (SEQ ID NO: 8), nodX genomic DNA (SEQ ID NO: 7), nodM cDNA (SEQ ID NO: 11), nodM genomic DNA (SEQ ID NO: 10), nodB cDNA (SEQ ID NO: 14), nodB genomic DNA (SEQ ID NO: 13), nodO cDNA (SEQ ID NO: 17), nodO genomic DNA (SEQ ID NO: 16), nodJ cDNA (SEQ ID NO: 20), nodJ genomic DNA (SEQ ID NO: 19), nodC cDNA (SEQ ID NO: 23), nodC genomic DNA (SEQ ID NO: 22), nodY1 cDNA (SEQ ID NO: 26), nodY1 genomic DNA (SEQ ID NO: 25), nodD2 cDNA (SEQ ID NO: 29), nodD2 genomic DNA (SEQ ID NO: 28), nodD1 cDNA (SEQ ID NO: 32), nodD1 genomic DNA (SEQ ID NO: 31), nodY2 cDNA (SEQ ID NO: 35), nodY2 genomic DNA (SEQ ID NO: 34), nodZ cDNA (SEQ ID NO: 38), nodZ genomic DNA (SEQ ID NO: 37), nodS cDNA (SEQ ID NO: 49), nodS genomic DNA (SEQ ID NO: 48), nodI cDNA (SEQ ID NO: 55), and nodI genomic DNA (SEQ ID NO: 54).

[0136] The nucleic acid molecules of the invention or otherwise described herein are preferably isolated. They may be isolated from biological samples using a variety of techniques known to those of ordinary skill in the art. For example, such polynucleotides may be isolated through the use of polymerase chain reaction (PCR) as known in the art. The nucleic acid molecules of the invention may be amplified using primers as defined herein that are derived from the polynucleotide sequences of the invention.

[0137] Additional methods for isolating polynucleotides include the use of all or a portion of the polynucleotides of the invention as hybridization probes. A genomic or cDNA library may be screened using techniques for hybridizing a labeled polynucleotide probe to a polynucleotide immobilized on a solid support such as a nitrocellulose filter or nylon membrane. Similarly, the probe may be coupled to beads and hybridized to the target sequence. Isolation may be achieved using protocols known in the art such as magnetic separation. Selection of appropriately stringent hybridization and wash conditions is considered to be within the skill of the art.

[0138] Polynucleotide fragments may be produced by techniques well known in the art such as restriction endonuclease digestion and oligonucleotide synthesis. In a sample, partial polynucleotide sequences may be used as probes in methods well known in the art to identify the corresponding full-length polynucleotide sequence. Such methods include PCR-based methods, 5' RACE and hybridization-based methods, computer / database-based methods as known in the art. Detectable labels such as radioisotopes, fluorescence, chemiluminescence and bioluminescence labels may be used to facilitate detection. Reverse PCR is also known and used in the art to enable the acquisition of unknown sequences flanking the polynucleotide sequences disclosed herein, starting from primers based on known regions. The method uses several restriction enzymes to generate appropriate fragments in the known region of the gene. The fragments are then circularized by intramolecular ligation and used as a PCR template. Divergent primers are designed from the known region. Standard molecular biology approaches as known in the art may be utilized to physically assemble the full-length clone. Primers and primer pairs that enable amplification of the polynucleotides of the invention also form a further aspect of the invention.

[0139] Variants (including orthologs) may be identified by the methods described. PCR-based methods as are known in the art may be used to identify variant polynucleotides. Typically, the polynucleotide sequences of primers useful for amplifying variants of polynucleotide molecules by PCR may be based on sequences encoding conserved regions of the corresponding amino acid sequences.

[0140] Further methods for identifying variant polynucleotides include the use of all or portions of the polynucleotides specified as hybridization probes for screening genomic or cDNA libraries as described above. Typically, probes based on sequences encoding conserved regions of the corresponding amino acid sequences may be used. The hybridization conditions may also be less stringent than those used when screening for sequences identical to the probe.

[0141] In another aspect, the invention relates to a TU comprising at least one isolated polynucleotide as described herein. In one embodiment, the TU is contained within a vector, preferably an expression vector. In one embodiment, the vector is selected from the group consisting of plasmids, BACs, (PACs), YACs, bacteriophages, phagemids, and cosmids. Preferably, the vector is a plasmid.

[0142] In another aspect, the invention relates to a vector encoding the isolated polypeptide or a functional variant or fragment thereof as described in the invention. In another aspect, the invention relates to a vector comprising the isolated nucleic acid sequence described in the invention.

[0143] In one embodiment, the isolated nucleic acid sequence is contained within a TU. In one embodiment, the vector is selected from the group consisting of plasmids, BACs, PACs, YACs, bacteriophages, phagemids, and cosmids. Preferably, the vector is a plasmid. In one embodiment, the vector is an expression vector.

[0144] The TU containing the polynucleotide of the present invention may incorporate its polynucleotide, or, where applicable, the encoded polypeptide of the present invention, into any suitable vector capable of expression in vitro or in a host cell. Preferably, the vector is an expression vector. Examples of suitable expression vectors include, but are not limited to, plasmid DNA vectors, viral DNA vectors (such as adenovirus and adeno-associated virus), or viral RNA vectors (such as retroviral vectors). In some embodiments, plasmids and / or phage vectors may be selected from the following vectors or variants thereof including pUC18, pU19, Mp18, Mp19, ColE1, PCR1 and pKRC; lambda gt10 and M13 plasmids, such as pBR322, pACYC184, pT127, RP4, p1J101, SV40 and BPV. Also included are vectors such as cosmids, YACs, BAC shuttle vectors, such as pSA3, PAT28 transposons (such as those described in US 5,792,294), etc., although not limited thereto.

[0145] Suitable viral vectors include, but are not limited to, vectors derived from adenovirus (AV), adeno-associated virus (AAV), retroviruses (such as lentivirus (LV), rhabdovirus, murine leukemia virus), herpesvirus, etc. The viral vectors used herein may be appropriately modified by pseudotyping with envelope proteins from other viruses or other surface antigens known and used in the art, or by substitution with different viral capsid proteins.

[0146] In one embodiment, the expression vector comprises at least 1, preferably at least 2, preferably at least 3, preferably at least 4, preferably at least 5, preferably at least 6, preferably at least 7, preferably at least 8, preferably at least 9, preferably at least 10 isolated polynucleotides as described herein.

[0147] In one embodiment, the expression vector comprises at least 1, preferably at least 2, preferably at least 3, preferably at least 4, preferably at least 5, preferably at least 6, preferably at least 7, preferably at least 8, preferably at least 9, preferably at least 10 TUs as described herein.

[0148] In one embodiment, the vector is a component of a cloning system. In one embodiment, the cloning system is useful for creating a gene construct comprising at least one TU.

[0149] In one embodiment, the vector is included in a vector set, and the vector set is part of a cloning system. In one embodiment, the cloning system is useful for creating a gene construct comprising at least one TU.

[0150] In one embodiment, the cloning system is useful for creating a gene construct comprising at least one TU. In one embodiment, the gene construct is a multi-gene construct comprising at least 2 TUs. In one embodiment, the multi-gene construct comprises at least 3, preferably at least 4, preferably at least 5, preferably at least 6, preferably at least 7, preferably at least 8, preferably at least 9, preferably at least 10 TUs.

[0151] The TU described herein may comprise one or more of the polynucleotide sequences disclosed by the present invention and / or polynucleotides encoding the polypeptides disclosed by the present invention. The TU may be constructed to drive the expression of at least one polypeptide involved in the biosynthesis of NAA 10, either in vitro or in vivo. In one embodiment, the TU comprises a polynucleotide of the present invention operably linked to a 5' or 3' untranslated regulatory sequence. The design of a particular TU will depend on a variety of factors, including the host cell in which the polynucleotide is to be expressed operably linked and the desired level of polynucleotide expression.

[0152] Similarly, the selection of various promoters, enhancers and / or other genetic elements for the TU will depend on a variety of factors, including the host cell and expression level discussed above. In one embodiment, the TU comprises a homologous promoter operably linked to a polynucleotide of the present invention. In another embodiment, the expression cassette comprises a heterologous promoter operably linked to a polynucleotide of the present invention. In one embodiment, the homologous or heterologous promoter is an inducible, repressible or controllable promoter. An appropriate promoter may be selected and used under conditions suitable to direct high level expression of the polynucleotide of the present invention. Many such elements are described in the literature and are available through commercial suppliers.

[0153] Although only by way of example, the promoter useful in the expression cassette may be any suitable eukaryotic or prokaryotic promoter. In one embodiment, the eukaryotic promoter may be eukaryotic RNA polymerase I (pol I), RNA polymerase II (pol II), or RNA polymerase III (pol III). In a particular cell type, the expression level of the polynucleotide ligated to be functional will be determined by the presence (or absence) in the vicinity of a particular gene regulatory sequence (such as an enhancer, silencer, etc.). The expression of the polynucleotide of the present invention may be driven using any suitable combination of promoter / enhancer (see the eukaryotic promoter database EPDB).

[0154] Additional promoters useful in the expression cassette include the β-lactamase, alkaline phosphatase, tryptophan, and tac promoter systems, all well known in the art. Yeast promoters include, but are not limited to, 3-phosphoglycerate kinase, enolase, hexokinase, pyruvate decarboxylase, glucokinase, and glyceraldehyde-3-phosphate dehydrogenase.

[0155] Prokaryotic promoters useful in the expression cassette include constitutive promoters (such as the int promoter of bacteriophage λ and the bla promoter of the β-lactamase gene sequence of pBR) and controllable promoters (such as lacZ, recA, and gal) as known in the art. A ribosome binding site upstream of the CDS may also be necessary for expression.

[0156] Enhancers useful in the TU include the SV40 enhancer, cytomegalovirus immediate early promoter enhancer, globin, albumin, insulin, etc. In one embodiment, the TU may be driven by a T3, T7, or SP6 cytoplasmic expression system.

[0157] The selection of a specific promoter / enhancer / cell type combination for protein expression is within the ordinary skill of those in the art of molecular biology (see, for example, Sambrook et al. (1989), which is incorporated herein by reference).

[0158] In another aspect, the present invention relates to an isolated host cell comprising the isolated polypeptide, isolated polynucleotide, TU and / or isolated vector described in the present invention. In one embodiment, the isolated host cell is a prokaryotic or eukaryotic cell. The prokaryote most commonly used as a host cell is a strain of Escherichia coli (E. coli). Other prokaryotic hosts include, but are not limited to, Pseudomonas, Bacillus, Serratia, Klebsiella, Streptomyces, Listeria, Salmonella and Mycobacteria.

[0159] In one embodiment, the eukaryotic cell is an animal cell, a plant cell, a fungal cell or a protozoan cell. In one embodiment, the animal cell is an insect cell or a mammalian cell. In one embodiment, the fungal cell is a single cell of a unicellular fungal host strain. In one embodiment, the fungal cell comprises fungal hyphae or a mycelium of a fungal host strain.

[0160] In one embodiment, the fungal cell, hyphae or mycelium of a fungal host cell line is from the genus Aspergillus, Trichoderma, Neurospora, Fusarium, Mortierella, Chrysosporium, Candida, Geotrichum, Yarrowia, Eremothecium, Trichoplusia, Ashbya, Hansenula, Pichia, Kluveromyces, Schizzosaccharomyces, Monascus, Talaromyces, Cryptonectria, Endothia, Tolypocladium, Hypocrea, Gibberella, Acremonium, Agaricus, Pleurotus, Penicillium, Volvariella, Flammulina, Lentinula, Auricularia, Ganoderma, (Rhizo)mucor, Riopus, or Saccharomyces, preferably from the genus Penicillium, Aspergillus, Saccharomyces, Pichia, Trichoplusia, and Spondoptera. Preferably, the fungal cell is from the genus Saccharomyces. Preferably, the fungal hyphae or mycelium is from the genus Penicillium, preferably P. paxilli.

[0161] In another aspect, the present invention relates to a method of making at least one NA comprising the step of heterologously expressing at least one polypeptide, isolated nucleic acid sequence, TU or vector as described herein in an isolated host cell.

[0162] In one embodiment, the NA is selected from the group of NAs shown in Figure 1. Preferably, the NA is NAF 5a or NAA 10, preferably NAA 10. In one embodiment, the polypeptide is a polypeptide as described herein or a functional variant or fragment thereof.

[0163] Specifically contemplated as aspects within this aspect of the invention are the various aspects set forth herein with respect to heterologous expression (including selection of appropriate control sequences), expression cassettes, genetic elements, TUs, multi-gene constructs, host cells, and vectors, any other aspect of the invention.

[0164] In certain embodiments, heterologous expression of the polypeptide comprises expression from at least one vector as described herein in the mycelium of an isolated fungal host cell or isolated fungal strain as described herein of at least one polynucleotide as described herein or at least one TU encoding at least one polypeptide of the present invention. In one embodiment, the polypeptide is NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), or NodZ (SEQ ID NO: 39), and preferably, the polypeptide is an enzyme that catalyzes the biological conversion of NAB 9 to NAA 10. In one embodiment, the fungal cell or strain is a cell or strain of the genus Penicillium, preferably P. paxilli.

[0165] In one embodiment, the TU is contained in a multi-gene construct comprising polynucleotides encoding at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, and / or at least ten polypeptides as described herein.

[0166] In another aspect, the present invention relates to at least one NA produced by the method of the present invention. In one embodiment, the NA is selected from the group of NAs shown in Figure 1. Preferably, the NA is NAF 5a or NAA 10, preferably NAA 10.

[0167] In another aspect, the present invention relates to an isolated polypeptide or a functional variant or fragment thereof derived from a Hypoxylon species that catalyzes a biochemical reaction in the biosynthetic pathway leading from GGI 2 to NAA 10.

[0168] In another aspect, the present invention relates to an isolated polynucleotide encoding at least one polypeptide derived from a Hypoxylon species that catalyzes a biochemical reaction in the biosynthetic pathway leading from GGI 2 to NAA 10.

[0169] In one embodiment, the isolated polypeptide is an oxygenase, preferably a cytochrome P450 oxygenase or an FAD-dependent oxygenase. Preferably, the cytochrome P450 oxygenase is NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodJ (SEQ ID NO: 21), or NodZ (SEQ ID NO: 39). Preferably, the FAD-dependent oxygenase is NodM (SEQ ID NO: 12), NodO (SEQ ID NO: 18), NodY1 (SEQ ID NO: 27), or NodY2 (SEQ ID NO: 36). In one embodiment, the isolated polypeptide is a transferase, preferably GGT, or a prenyltransferase. Preferably, GGT is NodC (SEQ ID NO: 24). Preferably, the prenyltransferase is NodD1 (SEQ ID NO: 33), or NodD2 (SEQ ID NO: 30). In one embodiment, the isolated polypeptide is an IDT cyclase. Preferably, the IDT cyclase is NodB (SEQ ID NO: 15). In one embodiment, the isolated polypeptide is NodS (SEQ ID NO: 50). In one embodiment, the isolated polypeptide is NodI (SEQ ID NO: 56).

[0170] In one embodiment, the isolated polypeptide catalyzes a biochemical reaction in the biosynthetic pathway leading from GGI 2 to NAF 5a. Preferably, the isolated polypeptide is a GGT, FAD-dependent oxygenase, IDT cyclase, or cytochrome P450 oxygenase. Preferably, the GGT is NodC (SEQ ID NO: 24). Preferably, the FAD-dependent oxygenase is NodM (SEQ ID NO: 12). Preferably, the IDT cyclase is NodB (SEQ ID NO: 15). Preferably, the cytochrome P450 oxygenase is NodW (SEQ ID NO: 3).

[0171] In one embodiment, the isolated polypeptide, or a functional variant or fragment thereof, is encoded by the nucleic acids described in the present invention. In another aspect, the present invention relates to a method for producing at least one Hypoxylon species polypeptide, or a functional variant or fragment thereof, comprising the step of heterologously expressing the isolated nucleic acid sequence, TU, or vector described in the present invention in an isolated host cell.

[0172] In one embodiment, the at least one Hypoxylon species polypeptide is a polypeptide described in the present invention as intended herein with respect to any other aspect of the present invention. In one embodiment, the at least one Hypoxylon species polypeptide is a polypeptide comprising the amino acid sequence of SEQ ID NO: NodW (SEQ ID NO: 3) or a functional variant or fragment thereof. Preferably, the polypeptide consists essentially of or consists of SEQ ID NO: NodW (SEQ ID NO: 3). In one embodiment, the isolated host cell comprises a fungal mycelium of the genus Penicillium, preferably P. paxilli.

[0173] Particularly contemplated in connection with this aspect of the present invention are the various aspects shown with respect to any other aspect of the present invention relating to heterologous expression (including selection of appropriate control sequences), genetic elements, TUs, multi-gene constructs, host cells, and vectors.

[0174] In another aspect, the present invention relates to a method for producing at least one NA, comprising the step of heterologously expressing in an isolated host cell at least one polypeptide that catalyzes a biochemical reaction in a biosynthetic pathway leading from GGI 2 to NAA 10.

[0175] In one embodiment, the at least one polypeptide is an oxygenase, preferably a cytochrome P450 oxygenase or an FAD-dependent oxygenase. Preferably, the cytochrome P450 oxygenase is NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodJ (SEQ ID NO: 21), or NodZ (SEQ ID NO: 39). Preferably, the FAD-dependent oxygenase is NodM (SEQ ID NO: 12), NodO (SEQ ID NO: 18), NodY1 (SEQ ID NO: 27), or NodY2 (SEQ ID NO: 36). In one embodiment, the isolated polypeptide is a transferase, preferably GGT, or a prenyltransferase. Preferably, GGT is NodC (SEQ ID NO: 24). Preferably, the prenyltransferase is NodD1 (SEQ ID NO: 33), or NodD2 (SEQ ID NO: 30). In one embodiment, the isolated polypeptide is an IDT cyclase. Preferably, the IDT cyclase is NodB (SEQ ID NO: 15). In one embodiment, the isolated polypeptide is NodS (SEQ ID NO: 50). In one embodiment, the isolated polypeptide is NodI (SEQ ID NO: 56).

[0176] In one embodiment, the at least one polypeptide catalyzes a biochemical reaction in a biosynthetic pathway leading from GGI 2 to NAF 5a. Preferably, the at least one polypeptide is GGT, an FAD-dependent oxygenase, an IDT cyclase, or a cytochrome P450 oxygenase. Preferably, GGT is NodC (SEQ ID NO: 24). Preferably, the FAD-dependent oxygenase is NodM (SEQ ID NO: 12). Preferably, the IDT cyclase is NodB (SEQ ID NO: 15). Preferably, the cytochrome P450 oxygenase is NodW (SEQ ID NO: 3).

[0177] In one embodiment, at least one polypeptide comprises the amino acid sequence of SEQ ID NO: NodW (SEQ ID NO: 3) or a functional variant or fragment thereof. Preferably, the polypeptide consists essentially of or consists of the sequence of SEQ ID NO: NodW (SEQ ID NO: 3). In one embodiment, the isolated host cell comprises fungal mycelia of the genus Penicillium, preferably P. paxilli.

[0178] Particularly contemplated in this aspect of the invention are the various aspects shown with respect to any other aspect of the invention, relating to heterologous expression (including selection of appropriate control sequences), genetic elements, TUs, multi-gene constructs, host cells, and vectors.

[0179] In another aspect, the invention relates to an isolated host cell that expresses at least one heterologous polypeptide that catalyzes the conversion of a substrate in a biosynthetic pathway leading to the formation of NAA 10. In one embodiment, at least one heterologous polypeptide catalyzes the conversion of a substrate in a biosynthetic pathway leading to the formation of NAF 5a.

[0180] In one embodiment, the substrate is selected from the group consisting of GGPP 1a, indole-3-glycerol phosphate 1b, GGI 2, monoepoxidized GGI 3a, emindole SB 4a, NAF 5a, NAE 6a, NAD 7a, NAC 8, and NAB 9.

[0181] In one embodiment, the conversion is selected from the group consisting of condensation, oxidation, or cyclization. In one embodiment, the substrate to be converted is GGPP 1a and indole-3-glycerol phosphate 1b, and the conversion is a condensation.

[0182] In one embodiment, the substrate to be converted is GGI 2, and the conversion is an oxidation. In one embodiment, the substrate to be converted is monoepoxidized GGI 3a, and the conversion is a cyclization.

[0183] In one embodiment, the substrate to be converted is emindole SB 4a, and the conversion is oxidation. In one embodiment, the substrate to be converted is NAF 5a, and the conversion is condensation.

[0184] In one embodiment, the substrate to be converted is NAE 6a, and the conversion is oxidation. In one embodiment, the substrate to be converted is NAD 7a, and the conversion is oxidation and condensation.

[0185] In one embodiment, the substrate to be converted is NAC 8, and the conversion is oxidation. In one embodiment, the substrate to be converted is NAB 9, and the conversion is oxidation. In another aspect, the present invention relates to an isolated host cell that produces at least one polypeptide involved in the biosynthetic pathway leading to NAA 10 by heterologous expression.

[0186] In one embodiment, the at least one polypeptide catalyzes a biochemical reaction in the biosynthetic pathway leading from GGI 2 to NAF 5a. Preferably, the at least one polypeptide is a GGT, an FAD-dependent oxygenase, an IDT cyclase, or a cytochrome P450 oxygenase. Preferably, the GGT is NodC (SEQ ID NO: 24). Preferably, the FAD-dependent oxygenase is NodM (SEQ ID NO: 12). Preferably, the IDT cyclase is NodB (SEQ ID NO: 15). Preferably, the cytochrome P450 oxygenase is NodW (SEQ ID NO: 3).

[0187] In some embodiments specifically contemplated with respect to this aspect of the invention, the at least one polypeptide is a polypeptide involved in the biosynthetic pathway leading to NAA 10, as defined herein, with respect to any other aspect of the invention.

[0188] In one aspect, at least one polypeptide is a polypeptide of the present invention or a functional variant or fragment thereof. In one aspect, the polypeptide or a functional variant or fragment thereof is encoded by a nucleic acid sequence of the present invention.

[0189] In one aspect, at least one polypeptide is involved in the biosynthetic pathway leading to NAF 5a. In one aspect, at least one polypeptide comprises the amino acid sequence of SEQ ID NO: NodW (SEQ ID NO: 3) or a functional variant or fragment thereof. Preferably, the polypeptide consists essentially of or consists of SEQ ID NO: NodW (SEQ ID NO: 3). In one aspect, the isolated host cell comprises a fungal mycelium of the genus Penicillium, preferably P. paxilli.

[0190] Particularly contemplated in this aspect of the invention are the various aspects shown with respect to heterologous expression (including selection of appropriate control sequences), genetic elements, TUs, multi-gene constructs, host cells, and any other aspect of the invention related to vectors.

[0191] In another aspect, the present invention relates to a method for producing at least one NA, comprising contacting a carbohydrate containing a substrate with a recombinant cell transformed with a nucleic acid that results in an increase in the level or activity of a polypeptide selected from the group consisting of NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodM (SEQ ID NO: 12), NodB (SEQ ID NO: 15), NodO (SEQ ID NO: 18), NodJ (SEQ ID NO: 21), NodC (SEQ ID NO: 24), NodY1 (SEQ ID NO: 27), NodD2 (SEQ ID NO: 30), NodD1 (SEQ ID NO: 33), NodY2 (SEQ ID NO: 36), NodZ (SEQ ID NO: 39), NodS (SEQ ID NO: 50), and NodI (SEQ ID NO: 56) or a functional variant or fragment thereof, as compared to the cell before transformation, such that the substrate is metabolized to at least one NA.

[0192] In one aspect, the nucleic acid encodes at least one polypeptide that catalyzes a biochemical reaction in the biosynthetic pathway leading from GGI 2 to NAF 5a, preferably a biochemical reaction leading from emindole SB 4a to NAF 5a.

[0193] In one aspect, the recombinant host cell is the isolated host cell of the invention as described herein. In one aspect, the carbohydrate is included in the medium. In one aspect, the medium is CDYE or a variant thereof that supports the growth of recombinant cells.

[0194] In one aspect, the nucleic acid encodes at least one polypeptide that is an oxygenase, preferably a cytochrome P450 oxygenase or an FAD-dependent oxygenase. Preferably, the cytochrome P450 oxygenase is NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodJ (SEQ ID NO: 21), or NodZ (SEQ ID NO: 39). Preferably, the FAD-dependent oxygenase is NodM (SEQ ID NO: 12), NodO (SEQ ID NO: 18), NodY1 (SEQ ID NO: 27), or NodY2 (SEQ ID NO: 36). In one aspect, the isolated polypeptide is a transferase, preferably GGT, or a prenyltransferase. Preferably, GGT is NodC (SEQ ID NO: 24). Preferably, the prenyltransferase is NodD1 (SEQ ID NO: 33), or NodD2 (SEQ ID NO: 30). In one aspect, the isolated polypeptide is an IDT cyclase. Preferably, the IDT cyclase is NodB (SEQ ID NO: 15). In one aspect, the isolated polypeptide is NodS (SEQ ID NO: 50). In one aspect, the isolated polypeptide is NodI (SEQ ID NO: 56).

[0195] In one aspect, the nucleic acid encodes at least one GGT, FAD-dependent oxygenase, IDT cyclase, or cytochrome P450 oxygenase. In one aspect, the nucleic acid encodes at least two, preferably at least three, preferably all four of GGT, FAD-dependent oxygenase, IDT cyclase, or cytochrome P450 oxygenase. Preferably, GGT is NodC (SEQ ID NO: 24). Preferably, the FAD-dependent oxygenase is NodM (SEQ ID NO: 12). Preferably, the IDT cyclase is NodB (SEQ ID NO: 15). Preferably, the cytochrome P450 oxygenase is NodW (SEQ ID NO: 3).

[0196] In one aspect, a polypeptide selected from the group consisting of NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodM (SEQ ID NO: 12), NodB (SEQ ID NO: 15), NodO (SEQ ID NO: 18), NodJ (SEQ ID NO: 21), NodC (SEQ ID NO: 24), NodY1 (SEQ ID NO: 27), NodD2 (SEQ ID NO: 30), NodD1 (SEQ ID NO: 33), NodY2 (SEQ ID NO: 36), NodZ (SEQ ID NO: 39), NodS (SEQ ID NO: 50), and NodI (SEQ ID NO: 56) or a functional variant or fragment thereof comprises the amino acid sequence of NodW (SEQ ID NO: 3), NodR (SEQ ID NO: 6), NodX (SEQ ID NO: 9), NodM (SEQ ID NO: 12), NodB (SEQ ID NO: 15), NodO (SEQ ID NO: 18), NodJ (SEQ ID NO: 21), NodC (SEQ ID NO: 24), NodY1 (SEQ ID NO: 27), NodD2 (SEQ ID NO: 30), NodD1 (SEQ ID NO: 33), NodY2 (SEQ ID NO: 36), NodZ (SEQ ID NO: 39), NodS (SEQ ID NO: 50), and NodI (SEQ ID NO: 56) or a functional variant or fragment thereof of the present invention.

[0197] In one aspect, the polypeptide comprises the amino acid sequence of NodW (SEQ ID NO: 3) or a functional variant or fragment thereof. Preferably, the polypeptide consists essentially of or consists of SEQ ID NO: NodW (SEQ ID NO: 3). In one aspect, the isolated host cell comprises fungal mycelia of the genus Penicillium, preferably P. paxilli.

[0198] Particularly contemplated in this aspect of the invention are the various aspects shown with respect to heterologous expression (including selection of appropriate control sequences), genetic elements, TUs, multi-gene constructs, host cells, and any other aspect of the invention related to vectors.

[0199] In one aspect, at least one heterologous or introduced homologous nucleic acid sequence is at least one NAA 10 biosynthetic gene selected from the group consisting of nodW, nodR, nodX, nodM, nodB, nodO, nodJ, nodY1, nodD2, nodD1, nodY2, nodZ, nodS, and nodI as described herein.

[0200] In one aspect, one of two different GGPPS enzymes is produced in H. prischidum by heterologous expression. In one aspect, one of two different GGPPS enzymes is encoded by a second copy of the native H. prischidum that encodes the GGPPS enzyme.

[0201] In another aspect, the invention relates to an isolate of Hypoxylon prischidum comprising a genetic modification that leads to increased biosynthesis of NAA 10. In one aspect, the isolate comprises increased expression of at least one GGPPS enzyme when compared to a control strain of H. prischidum, preferably H. prischidum ATCC 74245.

[0202] In one aspect, the increased expression is an increase in the expression of the native primary GGPPS gene of H. prischidum through modification of the gene regulatory elements. In one aspect, modification of the gene control element includes ligation such that it is functional to another or modified promoter of the native primary GGPPS gene.

[0203] In one aspect, modification of the gene control element includes ligation such that it is functional to a more robust native promoter of the native primary GGPPS gene. In one aspect, the native primary GGPPS gene is an introduced homologous gene.

[0204] In one aspect, the increased expression is the result of heterologous expression of a biosynthetic gene contributing to NAA 10 biosynthesis. In one aspect, the increased expression is due to the expression of a heterologous gene in H. prischidum that has a biochemical function equivalent to the genes identified in the Nod cluster as described herein.

[0205] In one aspect, the increased expression is due to the expression of a heterologous gene in H. prischidum that corrects a limitation in the supply of a substrate compound or biosynthetic intermediate required for NAA 10 biosynthesis.

[0206] In one aspect, the increased expression is due to the heterologous expression in H. prischidum of any gene encoding GGT that catalyzes the condensation of GGPP 1a and indole-3-glycerol phosphate 1b to produce 3-geranylgeranyl indole 2.

[0207] In one aspect, the increased expression is due to the heterologous expression in H. prischidum of any gene encoding GGPPS. In one aspect, the increased expression is due to the heterologous expression in H. prischidum of any gene encoding a FAD-dependent oxidase that produces the monoepoxidized GGI product 3a.

[0208] In one aspect, the increased expression is due to the heterologous expression in H. prischidum of any gene encoding an enzyme that cyclizes the monoepoxidized GGI 3a to produce emindole SB 4a.

[0209] In one embodiment, the increased expression is by heterologous expression in H. prischidum of any gene encoding an oxidase that oxidizes emindole SB 4a to produce nostocyclic acid.

[0210] In one embodiment, the increased expression is by at least one genetic modification that leads to increased expression of an NA biosynthetic gene selected from the group consisting of nodW, nodR, nodX, nodM, nodB, nodO, nodJ, nodY1, nodD2, nodD1, nodY2, nodZ, nodS, and nodI as described herein.

[0211] In another aspect, the present invention relates to a method for producing NAA 10 comprising the step of expressing at least one heterologous nucleic acid sequence in Hypoxylon prischidum, wherein the at least one heterologous nucleic acid sequence encodes an enzyme in the biosynthetic pathway leading to NAA 10.

[0212] Particularly contemplated in connection with this aspect of the invention are all aspects shown with respect to increased expression of NAA 10, and also heterologous expression (including selection of appropriate control sequences), gene elements, TUs, multi-gene constructs, host cells, and vectors as described herein, and also all aspects shown with respect to any other aspect of the invention with respect to an isolate of Hypoxylon prischidum as described herein.

[0213] As used herein, when reference is made to a patent specification, other external document, or other source of information, this is generally for the purpose of providing background for discussing the features of the present invention. Unless specifically stated otherwise, such reference to such external documents is not to be construed as an admission that such documents; or such sources of information, are prior art or form part of the common general knowledge in the art in any respect.

[0214] The present invention will now be illustrated in a manner not limited by the following examples.

Example

[0215] Materials and Methods gDNA Isolation for Genome Sequencing and TUM Amplification Genomic DNA for genome sequencing and TUM amplification by PCR was isolated with modifications, following Byrd et al. 26 from Penicillium paxilli strain ATCC® 26601 TM (PN2013) 24 and Hypoxylon pruinatum strain ATCC® 74245 TM、25 in Milli-Q® water, sterile 2.4% (w / v) Difco TM Potato Dextrose Broth (Becton, Dickinson and Company, Maryland, USA) was prepared in 25 mL aliquots in 125 mL Erlenmeyer flasks and 5x10 6 spores or ~1 cm 2Freshly ground mycelium (for non-sporulating strains) was inoculated. The culture was incubated at 22 °C for 2 - 4 days with shaking (200 rpm). The fermentation broth was filtered through a sterile, fluffed nappy liner, and the mycelium was rinsed three times with sterile water. The mycelium was transferred to a sterile 15 mL centrifuge tube and flash frozen in liquid nitrogen for 24 - 48 hours for lyophilization. 15 - 20 mg of freeze-dried mycelium was placed in a mortar with liquid nitrogen and ground to a powder. The ground mycelium was transferred to a 2 mL tube and resuspended in 1 mL of extraction buffer (150 mM EDTA, 50 mM Tris-HCl, and 1% (w / v) sodium lauroyl sarcosinate). 1.6 mg of proteinase K was added to the tube, and the contents were incubated at 37 °C for 30 minutes. The tube was centrifuged at 13,000 rpm for 10 minutes, and the supernatant was transferred to a fresh 2 mL tube. 500 μL of phenol and 500 μL of chloroform were added to the tube, and the contents were mixed by vortexing and then centrifuged at 13,000 rpm for 10 minutes. The aqueous phase was transferred to a fresh 2 mL tube and washed two more times with 500 μL of phenol and 500 μL of chloroform as described above. The aqueous phase was then transferred to a fresh 2 mL tube and washed with 1 mL of chloroform (vortexing and centrifuging at 13,000 rpm for 10 minutes). The aqueous phase was transferred to a fresh 2 mL tube and mixed with 1 mL of chilled isopropanol. The DNA was precipitated overnight at -20 °C and pelleted by centrifugation at 13,000 rpm for 10 minutes. The supernatant was discarded, and the DNA was resuspended in 1 mL of 1 M NaCl. The tube was incubated at room temperature for 10 minutes and then centrifuged at 13,000 rpm for 10 minutes to pellet the polysaccharides. The supernatant was transferred to a fresh tube and mixed with 1 mL of isopropanol. The tube was incubated at room temperature for 10 minutes, and the DNA was pelleted by centrifugation at 13,000 rpm for 10 minutes. The supernatant was discarded, and 1 mL of chilled 70% ethanol was added to the pellet without resuspension. The tube was centrifuged at 13,000 rpm for 2 minutes, and the supernatant was discarded.The tube was centrifuged at 13,000 rpm for 1 minute, and the remaining 70% ethanol was aspirated with a pipette. The pellet was air-dried at room temperature, resuspended in 50 μL of Milli-Q® water, and stored at -20 °C.

[0216] Overview of MIDAS Design The MIDAS toolkit is based on Golden Gate assembly technology 27 and utilizes the ability of type IIS restriction enzymes to seamlessly ligate multiple DNA fragments together in a single reaction. MIDAS uses three type IIS restriction enzymes, AarI, BsaI, and BsmBI, which generate 4 bp overhangs defined by the user upon cleavage. Through appropriate selection of these user-defined overhangs and through appropriate orientation of the type IIS sites flanking each DNA fragment, multiple fragments can be assembled in an ordered (directional) manner into a recipient plasmid (also referred to as a destination vector) using a one-pot restriction ligation reaction. The recipient plasmid contains a marker gene (typically the lacZα gene for blue / white screening) flanked by two divergently oriented recognition sites for the type IIS enzyme; these elements collectively referred to as the "Golden Gate cloning cassette" are replaced by the insert during the assembly reaction.

[0217] Similar to other recently described Golden Gate-based modular assembly technologies 28-32 the assembly of genes and multi-gene constructs using MIDAS is a hierarchical process. At the first level (MIDAS level 1), functional modules (promoters, CDSs, terminators, tags, etc.) are cloned into a level 1 destination vector (pML1), where they form a library of reusable and sequence-verified parts. Due to the complementary design of the modules and the destination vector, once cloned into pML1, these modules are ensured to be releasable from the vector by digestion with BsaI.

[0218] At the second level (Level 2), a set of the array that is verified to be compatible with the Level 1 modules is released from pML1 and then assembled into a Level 2 destination vector (pML2) using the BsaI-mediated Golden Gate reaction, leading to the generation of a Level 2 plasmid containing a eukaryotic TU. Again, by the design rules, it is ensured that each of the assembled TUs can be released from the pML2 vector by digestion with either AarI or BsmBI (depending on the pML2 vector used to assemble the TU).

[0219] At Level 3, the TUs assembled at Level 2 are released from the pML2 plasmid and then successively assembled together into a Level 3 destination vector (pML3) using either AarI- or BsmBI-mediated Golden Gate reaction to form a functional multi-gene construct, which can then be transformed into a desired expression host.

[0220] Level 1: Module Cloning At Level 1, a functional TUM is generated either as a PCR product or as a synthetic polynucleotide sequence (from a gene synthesis company) and then cloned into a Level 1 destination vector (pML1) by BsmBI-mediated Golden Gate cloning. For cloning into the pML1 vector, each of the amplified TUMs is flanked by two convergent BsmBI sites, BsmBI[CTCG] and [AGAC]BsmBI, and PCR primers are designed to generate sticky ends that are compatible with those of the BsmBI sites present in the pML1 destination vector upon restriction enzyme digestion. Thus, the Golden Gate cloning cassette present in pML1 consists of two divergent BsmBI sites flanking a lacZα scorable marker: 5’-[CTCG]BsmBI-lacZα-BsmBI[AGAC]-3’ (Figure 11A).

[0221] To enable the subsequent (i.e., level 2) assembly of full-length TUs, each TUM is designed such that four module-specific nucleotides (NNNN) flank it at the 5' end and four module-specific nucleotides (NNNN) flank it at the 3' end, and these are included as part of the PCR primer sequences. Due to the complementary design of the module to be amplified and the pML1 vector, when the amplified TUM is cloned into pML1 using the BsmBI-mediated Golden Gate reaction, each TUM becomes flanked by convergent BsaI recognition sites, and when the module is released from pML1 during BsaI-mediated Golden Gate assembly of the subsequent (i.e., level 2) full-length TU, it is ensured that the module-specific nucleotides (NNNN and NNNN) become the BsaI-specific 4bp overhangs. Thus, the overall structure of each module in the PCR product (or synthetic polynucleotide) is of the type: 5'-BsmBI[CTCG]NNNN-TUM-NNNNtg[AGAC]BsmBI-3' (Figure 11B), which, after BsmBI-mediated cloning, becomes 5'-BsaI[NNNN]-TUM-[NNNN]BsaI-3' in pML1 (Figure 11C).

[0222] Since each TUM is defined by its flanking four nucleotides, these module-specific bases effectively form the addressing system of each TUM, and they determine its position and orientation within the assembled TU. The developers of MoClo and GoldenBraid2.0 have already worked in concert to develop a common syntax or standard addressing set (referred to as "fusion sites" in the MoClo system and "barcodes" in GoldenBraid2.0) for a very diverse set of TUMs for plant expression, and this standard is also adopted herein for MIDAS-based assembly of TUs for expression in filamentous fungi (Figure 12). 33 、and for the expression in filamentous fungi, this standard is also adopted herein for MIDAS-based assembly of TUs (Figure 12).

[0223] Thus, for filamentous fungal expression, the ProUTR module (including a promoter, 5’untranslated region (UTR) and ATG start codon) has GGAG as the module-specific 5’ nucleotide and AATG as the module-specific 3’ nucleotide (i.e., 5’-GGAG-ProUTR-A ATG -3’), with the translation start codon underlined. Similarly, the CDS module is flanked by AATG and GCTT (i.e., 5’-A ATG -CDS-GCTT-3’), while the UTRterm module (consisting of the 3’UTR and 3’ non-transcribed region including the polyadenylation signal) has the form 5’-GCTT-UTRterm-CGCT-3’. Considerations regarding the design of PCR primers for amplifying these three types of TUMs are shown in Table 12.

[0224] After BsmBI-mediated assembly of the TUM in pML1, the reaction was transformed into an E. coli strain, e.g., DH5α (or equivalent), and plated onto LB plates supplemented with spectinomycin, IPTG and X-Gal. Plasmids harboring the cloned TUM were identified by screening for white colonies and confirmed by sequencing.

[0225] At MIDAS level 1, it is important to mask or remove all internal recognition sites for AarI, BsaI and BsmBI from the TUM. The process of masking or removing such restricted sites, termed “domestication”, can be achieved by: (i) excluding these sites when ordering the sequence from a gene synthesis company, (ii) performing site-directed mutagenesis, or (iii) using masking oligonucleotides that form triplexes with the target DNA, thereby preventing restriction enzyme cleavage. 30 Type IIS enzymes are used for mutagenesis 34 and for Golden Gate domestication purposes 27、35In the same manner as previously utilized, the inventors adapted the MIDAS module by designing PCR primers (referred to as adaptor primers) that contain a single nucleotide mismatch overlapping with and disrupting the internal IIS-type restriction site. Since the PCR products are designed to be assembled together in MIDAS using the BsmBI-mediated Golden Gate reaction to form a full-length adapted TUM in pML1, it is important that the MIDAS adaptor primers are designed to include the BsmBI restriction site that generates a compatible overhang at the 5'-end.

[0226] Level 2: TU assembly At level 2, a compatible set of level 1 TUMs (e.g., ProUTR, CDS, and UTRterm modules) that have been cloned and sequence-verified were assembled into the pML2 destination vector using the BsaI-mediated Golden Gate reaction, leading to the generation of a level 2 plasmid (pML2 entry clone) containing a complete (i.e., full-length) eukaryotic TU. The module address standard described previously ensures that TU assembly proceeds in an ordered and directional manner, with the 3'-end of one module compatible with the 5'-end of the next module.

[0227] The module-specific bases GGAG located at the 5'-end of the ProUTR module and CGCT located at the 3'-end of the UTRterm module are compatible with the overhangs generated by digestion of the pML2 destination vector with BsaI, and thus these bases define the outermost cloning boundaries for level 2 assembly.

[0228] In MIDAS, there are eight level 2 (pML2) destination vectors capable of assembling TUs, and the selection depends on the desired arrangement of the TUs in the multi-gene plasmid produced at level 3, namely: (i) the desired order of adding TUs to the multi-gene assembly, (ii) the desired direction of assembling the multi-gene plasmid, and (iii) the desired orientation of each TU in the multi-gene plasmid. These features are discussed further below.

[0229] The pML2 vectors are distinguished from each other by the arrangement of specific sequence features that are very important for the operation of MIDAS. These sequence features, collectively referred to as the MIDAS cassette (Figure 13), define the level 2 assembly of the TUs and govern the assembly of the multi-gene constructs produced at level 3.

[0230] Each MIDAS cassette is defined by (i) having a Golden Gate cloning cassette containing adjacent divergent BsaI recognition sites, (ii) different arrangements of the recognition sites for AarI and BsmBI, and (iii) the presence or absence of a lacZα scorable marker. These features are described in more detail.

[0231] In contrast to a normal Golden Gate cloning cassette (typically containing the lacZα gene for blue / white screening), the Golden Gate cloning cassette in all eight pML2 vectors contains a mutant E. coli pheS gene with adjacent divergent BsaI recognition sites (driven by the promoter of the E. coli gene for chloramphenicol acetyltransferase). The Thr of the E. coli pheS gene used herein 251 Ala / Ala 294 Gly double mutant confers high lethality to cells grown on LB medium supplemented with the phenylalanine analog 4-chloro-phenylalanine, 4CP. 36During TU BsaI-mediated level 2 assembly, the mutant pheS gene is removed from the pML2 vector and thus may be used as a negative selection marker.

[0232] The eight pML2 vectors may be divided into two classes, "blue" and "white", according to the presence or absence of the lacZα gene in the MIDAS cassette (see Figure 13). There are four "blue" pML2 vectors (indicated by "B" in the plasmid name) and four "white" pML2 vectors (indicated by "W" in the plasmid name). In the "blue" and "white" vectors, the relative arrangement of the AarI and BsmBI restriction sites in the MIDAS cassette also differs. Thus, in the "blue" vector, the entire MIDAS cassette contains the lacZα gene nested adjacent to the convergent BsmBI site and with a divergent AarI site adjacent within it. In the "white" vector, the enzyme arrangement is switched (the entire MIDAS cassette contains the convergent AarI site with two divergent BsmBI sites nested within it), and there is no lacZα gene. The lacZα chromogenic marker in the pML2 vector is not used for blue / white screening during level 2 Golden Gate assembly of the TU (reserved for level 3 cloning), but it is important to note that the choice of "blue" or "white" vector into which the TU is to be assembled must be made during level 2 assembly of the TU because this will determine the order in which the TUs are added to the multi-gene construct at level 3. Similarly, the AarI and BsmBI sites are not used for level 2 assembly of the TU; instead, they are incorporated into the level 3 assembly of the multi-gene construct. These considerations, including the differences between the (+) and (-) vectors, are further discussed below in the level 3 description.

[0233] The orientation (transcription direction) of each TU can be freely defined by assembling each TU into either the pML2 "forward" vector (indicated by "F" in the plasmid name) or the pML2 "reverse" vector (indicated by "R" in the plasmid name). The pML2 "reverse" vector has a switched BsaI recognition site (for Golden Gate assembly of the TU) compared to the BsaI fusion site in the pML2 "forward" vector. Thus, the pML2 "forward" vector has a Golden Gate cassette based on pheS oriented as 5’-[GGAG]BsaI-pheS→-BsaI[CGCT]-3’, while the pML2 "reverse" vector has the switched BsaI recognition site: 5’-[AGCG]BsaI-←pheS-BsaI[CTCC]-3’; where the arrow indicates the direction of transcription of the mutant pheS gene.

[0234] When Escherichia coli DH5α cells (or equivalents) transformed in the assembly reaction were plated on LB plates supplemented with kanamycin and 4CP, in contrast to the cloned level 1 modules, the pML2 destination vector confers kanamycin resistance and enables efficient counter-selection against the level 1 module backbone, while the mutant pheS gene provides a strong negative selection against any parental pML2 destination plasmid.

[0235] Level 3: Assembly of multi-gene constructs At MIDAS level 3, the TUs assembled in the pML2 plasmid are consecutively loaded (by binary assembly) into a level 3 destination vector (pML3) to form a multi-gene construct.

[0236] The assembly of multi-gene constructs at level 3 depends critically on the relative arrangement of AarI and BsmBI restriction sites in the MIDAS cassettes located in the "blue" and "white" pML2 vectors; the nested and inverted arrangement of these restriction sites in the white vector, as compared to the blue vector, is a defining feature of the MIDAS multi-gene assembly process. In the "blue" vector, the entire MIDAS cassette has adjacent convergent BsmBI sites, and nested within it is the divergent AarI site adjacent to the lacZα gene. In the "white" vector, the enzyme arrangement is reversed (the entire MIDAS cassette has adjacent convergent AarI sites, and nested within it are two divergent BsmBI sites), and there is no lacZα gene. As illustrated in Figure 14, the nesting and inversion of restriction sites in the "blue" and "white" vectors means that TUs assembled in the "white" MIDAS cassette can be inserted into the "blue" MIDAS cassette using an AarI-mediated Golden Gate reaction, and conversely, TUs assembled in the "blue" MIDAS cassette can be cloned into the "white" MIDAS cassette using a BsmBI-mediated Golden Gate reaction. This cloning cycle (i.e., the alternation between "white" and "blue" pML2 entry clones) can be repeated indefinitely.

[0237] The level 3 destination vector, the Golden Gate cloning cassette found in pML3, consists of the lacZα gene flanked by divergent AarI sites: [CATT]AarI-lacZα-AarI[CGTA], thus MIDAS level 3 assembly always starts with an AarI-mediated Golden Gate reaction between pML3, and the TUs assembled within the pML2 "white" destination vector (i.e., the first TU is always added) (Figure 14). The generated plasmid is then used in a BsmBI-mediated Golden Gate reaction with the TUs cloned within the pML2 "blue" destination vector. Additionally, further TUs are added by following this approach of alternating AarI- and BsmBI-mediated Golden Gate reactions, using the pML2 "white" and pML2 "blue" entry clones, respectively. Thus, each plasmid generated by cloning TUs into the multi-gene construct becomes the destination vector for the next cycle of TU addition (Figure 14).

[0238] After each cloning cycle, Escherichia coli DH5α cells (or equivalents) are transformed by the Golden Gate reaction, plated on LB plates supplemented with spectinomycin, IPTG, and X-Gal, and positive clones are identified by blue / white screening. Spectinomycin selects for cells that have taken up the level 3 plasmid and counter selects against any pML2 plasmid backbone. The lacZα chromogenic marker present in the pML2 "blue" vector was not utilized previously during level 2 assembly of the TU, but here it is noted that at the multi-gene assembly level (level 3), it comes to be used for blue / white screening. Thus, for TUs assembled into the multi-gene construct using the AarI-mediated Golden Gate reaction, white colonies are picked for analysis, while for TUs assembled into the multi-gene construct using the BsmBI-mediated Golden Gate reaction, blue colonies are picked for analysis (see Table 13).

[0239] In the simplest arrangement, MIDAS can achieve multi-gene assembly using only two pML2 destination vectors: one "white" vector and one "blue" vector (Figure 15A). The full set of eight pML2 vectors is provided to allow maximum user control over (i) the order in which each TU is added to the growing multi-gene construct, (ii) the desired orientation of each TU (i.e., the direction of transcription), and (iii) the polarity of the assembly, i.e., the direction in which the next TU is loaded into the multi-gene construct.

[0240] First, and as described previously, the order of addition of each TU to the growing multi-gene construct is governed by the choice of "white" or "blue" pML2 destination vector into which the TU is assembled.

[0241] Second, as previously described when discussing the level 2 features, the orientation of the TU (direction of transcription) can be freely defined by the choice of the "forward" or "reverse" pML2 vector in which the TU is assembled. Extending MIDAS to include the option to assemble the TU in either orientation results in a set of vectors that is extended to four pML2 plasmids (see FIGS. 13 and 15B).

[0242] Third, the polarity of the multi-gene assembly (i.e., the direction in which new TUs are added to the growing multi-gene assembly) can also, in this case, be freely defined by assembling the TU in either a "plus" (+) or "minus" (-) polarity pML2 destination vector (FIG. 13). Use of the pML2(+) entry clone for level 3 assembly ensures that the next added TU is added in the same direction as the transcription direction of the Spec R gene in pML3, i.e., the next TU to be assembled in the multi-gene construct is added to the right of the TU added using the pML2(+) entry clone (as illustrated in FIGS. 15A and 15B). In contrast, use of the pML2(-) entry clone for level 3 assembly ensures that the next added TU is added in the direction opposite to the transcription direction of the Spec R gene in pML3, and thus the next TU to be loaded into the multi-gene construct is added to the left of the TU added using the pML2(-) entry clone. However, constructing the multi-gene construct using both polarity entry clones (i.e., both pML2(+) and pML2(-) entry clones) gives MIDAS the ability to switch the direction in which new TUs are added to level 3 assembly, and for the hypothetical assembly shown in FIG. 15C, indicates that all subsequently added TUs will be nested between TU3 and TU2.

[0243] Bacterial and fungal strains The routine growth of Escherichia coli was carried out in LB broth at 37°C. Chemically competent Escherichia coli HST08 Stellar cells (Clontech Laboratories, Inc.) were used for routine transformation and plasmid maintenance. The Penicillium paxilli strains used in this study are shown in Table 7.

[0244] Protocol for MIDAS Level 1 Module Cloning The PCR amplification modules were purified using a spin column protocol and cloned into the MIDAS Level 1 plasmid pML1 by BsmBI-mediated Golden Gate assembly. Typically, 1 - 2 μL (approximately 50 - 200 ng) of pML1 plasmid DNA from a miniprep was mixed with 1 - 2 μL of each purified PCR fragment, 1 μL of BsmBI (20 U / μL), 1 μL of T4 DNA ligase (20 U / μL), and 2 μL of 10xT4 DNA ligase buffer in a total reaction volume of 20 μL. The reaction was incubated at 37°C for 1 - 3 hours and an aliquot (typically 2 - 3 μL) was transformed into 30 μL of Escherichia coli HST08 Stellar competent cells by heat shock. After the recovery period (i.e., addition of 250 μL of SOC medium and incubation at 37°C for 1 hour), an aliquot of the transformation mixture was spread onto LB agar plates supplemented with 50 μg / mL spectinomycin, 1 mM IPTG, and 50 μg / mL X-Gal. The plates were incubated overnight at 37°C and white colonies were selected for analysis.

[0245] Protocol for MIDAS Level 2 TU Assembly Using the modules cloned at level 1, the full-length TU was assembled into the MIDAS level 2 plasmid by BsaI-mediated Golden Gate assembly. Typically, 40 fmol of pML2 plasmid DNA was mixed with 40 fmol each of the plasmid DNA of the level 1 entry clones, 1 μL of BsaI-HF (20 U / μL), 1 μL of T4 DNA ligase (20 U / μL), and 2 μL of 10x T4 DNA ligase buffer in a total reaction volume of 20 μL. The reaction was incubated in a DNA Engine PTC-200 Peltier thermal cycler (Bio-Rad) using the following parameters: 45 cycles of (37 °C for 2 minutes and 16 °C for 5 minutes), then 50 °C for 5 minutes, and 80 °C for 10 minutes. The reaction was transformed as described for level 1 assembly and plated on LB agar plates containing 75 μg / mL kanamycin and 1.25 mM 4CP. After incubation overnight at 37 °C, colonies were picked for analysis.

[0246] Protocol for MIDAS Level 3 Multigene Assembly Using the full-length TUs assembled at level 2, either AarI (for TUs cloned into the pML2 "white" vector) or BsmBI (for TUs cloned into the pML2 "blue" vector) was used, and by performing Golden Gate assembly alternately, multi-gene assembly was generated in the level 3 destination vector. Typically, 40 fmol of the level 3 destination vector plasmid DNA was mixed with 40 fmol of the level 2 entry clone plasmid DNA, 1 μL of BsaI-HF (20 U / μL), 1 μL of T4 DNA ligase (20 U / μL), and 2 μL of 10x T4 DNA ligase buffer in a total reaction volume of 20 μL. The reaction was incubated in a DNA Engine PTC-200 Peltier thermal cycler (Bio-Rad) using the following parameters: 45 cycles of (37 °C for 2 minutes and 16 °C for 5 minutes), then 37 °C for 5 minutes, and 80 °C for 10 minutes. The reaction was transformed as described for level 1 assembly and plated onto LB agar plates supplemented with 50 μg / mL spectinomycin, 1 mM IPTG, and 50 μg / mL X-Gal. The plates were incubated overnight at 37 °C. White colonies were selected for analysis for AarI-mediated assembly reactions, while blue colonies were selected for BsmBI-mediated assembly reactions.

[0247] Media and reagents for fungal research CDYE (Czapex-Dox / yeast extract) medium containing trace elements was prepared with deionized water, and the medium contained 3.34% (w / v) Czapex-Dox (Oxoid Ltd, Hampshire, UK), 0.5% (w / v) yeast extract (Oxoid Ltd, Hampshire, UK), and 0.5% (v / v) trace element solution. For agar plates, Select agar (Invitrogen, California, USA) was added up to 1.5% (w / v).

[0248] A trace element solution was prepared in deionized water, and the solution contained 0.004% (w / v) cobalt(II) chloride hexahydrate (Ajax Finechem, Auckland, New Zealand), 0.005% (w / v) copper(II) sulfate pentahydrate (Schariau, Barcelona, Spain), 0.05% (w / v) iron(II) sulfate heptahydrate (Merck, Darmstadt, Germany), 0.014% (w / v) manganese(II) sulfate tetrahydrate, and 0.05% (w / v) zinc sulfate heptahydrate (Merck, Darmstadt, Germany). The solution was stored with one drop of 12 M hydrochloric acid.

[0249] A regeneration (RG) medium was prepared in deionized water, and the medium contained 2% (w / v) malt extract (Oxoid Ltd, Hampshire, UK), 2% (w / v) D(+)-anhydrous glucose (VWR International BVBA, Leuven, Belgium), 1% (w / v) peptone for fungi (Oxoid Ltd, Hampshire, UK), and 27.6% sucrose (ECP Ltd. Birkenhead, Auckland, New Zealand). Depending on whether the medium was used for plates (1.5% RGA) or overlaid (0.8% RGA), 1.5% or 0.8% (w / v) of Select agar (Invitrogen, California, USA) was added respectively.

[0250] Fungal protocol - protoplast preparation The preparation of fungal protoplasts for transformation was modified and followed Yelton et al. 37 In a 100 mL Erlenmeyer flask, five 25 mL aliquots of CDYE medium containing trace elements were inoculated with 5x10 6 spores and incubated at 28 °C for 28 hours with shaking (200 rpm). From all five flasks, the fermentation broth was filtered through a sterile fluffy liner, and the combined mycelia were washed three times with sterile water and then with OM buffer (10 mM Na 2 HPO 4 and 1.2 M MgSO 4 .7H 2 O, 100 mM NaH 2PO 4 .2H 2 O to pH 5.8) and rinsed once. The weight of the mycelium was measured and resuspended in 10 mL of filter-sterilized lytic enzyme solution per gram of mycelium (prepared by resuspending lytic enzyme from Trichoderma harzianum (Sigma) in OM buffer at 10 mg / mL) and incubated at 30 °C for 16 h with shaking at 80 rpm. The protoplasts were filtered through a sterile fluffy liner into a 250 mL Erlenmeyer flask. An aliquot (5 mL) of the filtered protoplasts was transferred into a 15 mL sterile centrifuge tube and overlaid with 2 mL of ST buffer (0.6 M sorbitol and 0.1 M Tris-HCl, pH 8.0). The tubes were centrifuged at 2600 x g at 4 °C for 15 min. In each tube, the white layer of protoplasts formed between the OM and ST buffers was transferred into a sterile 15 mL centrifuge tube (in 2 mL aliquots), gently washed by pipetting and resuspending in 5 mL of STC buffer (1 M sorbitol, 50 mM Tris-HCl, pH 8.0, and 50 mM CaCl 2 ) and centrifuged at 2600 x g at 4 °C for 5 min. The supernatant was decanted and the protoplasts pelleted from multiple tubes were combined by resuspending in 5 mL aliquots of STC buffer. After repeating the STC buffer wash three times, the protoplasts were pooled into a single 15 mL centrifuge tube. The final protoplast pellet was resuspended in 500 μL of STC buffer and the protoplast concentration was estimated using a hemocytometer. The protoplast stock was diluted to a final concentration of 1.25 x 10 8 protoplasts per mL of STC buffer. An aliquot (100 μL) of the protoplasts was used immediately for fungal transformation and the excess protoplasts were held in 8% PEG solution (80 μL of protoplasts added to 40% (w / v) PEG 4000 in 20 μL of STC buffer) in 1.7 mL microcentrifuge tubes and stored at -80 °C.

[0251] Fungal protocol - Transformation of P. paxilli Vollmer and Yanofsky 38 , as well as Oliver et al 39 -modified fungal transformation was performed in 1.7 mL microcentrifuge tubes containing 100 μL (1.25x10 7 ) of either freshly prepared protoplasts in STC buffer or protoplasts stored in 8% PEG solution (as described above). A solution containing 2 μL of spermidine (50 mM in H 2 O), 5 μL of heparin (5 mg / mL in STC buffer), and 5 μg of plasmid DNA (250 μg / mL) was added to the protoplasts and incubated on ice for 30 minutes, after which 900 μL of 40% PEG solution (40% (w / v) PEG 4000 in STC buffer) was added. The transformation mixture was incubated on ice for an additional 15 - 20 minutes, transferred to a sterile 50 mL tube containing 17.5 mL of 0.8% RGA medium (pre-warmed to 50 °C), mixed by inversion, and 3.5 mL aliquots were dispensed onto 1.5% RGA plates. After incubation overnight at 25 °C, 5 mL of 0.8% RGA (containing sufficient geneticin to achieve a final concentration of 150 μg per mL of solid medium) was overlaid onto each plate. The plates were incubated at 25 °C for an additional 4 days, and spores were picked from individual colonies and streaked onto CDYE agar plates supplemented with 150 μg / mL geneticin. The streaked plates were incubated at 25 °C for an additional 4 days. Spores from individual colonies were resuspended in 50 μL of 0.01% (w / v) triton X-100, and 5 x 5 μL aliquots of the spore suspension were transferred onto fresh CDYE agar plates supplemented with 150 μg / mL geneticin. The sporulation plates were incubated at 25 °C for 4 days, and spore stocks were prepared as follows. Colony plugs from the sporulation plates were suspended in 2 mL of 0.01% (v / v) triton X-100, and 800 μL of the suspended spores were mixed with 200 μL of 50% (w / v) glycerol in 1.7 mL microcentrifuge tubes. Spore stocks were used to inoculate 50 mL of CDYE medium, flash frozen in liquid nitrogen, and stored at -80 °C.

[0252] Indole diterpene production and extraction The fungal transformant was grown at 28 °C for 7 days in a 250 mL Erlenmeyer flask capped with cotton wool in a shaker culture (≥200 rpm) in 50 mL of CDYE medium containing trace elements. The mycelium was isolated from the fermentation broth by filtration through a fluffy liner, transferred to a 50 mL centrifuge tube (Lab Serv® , Thermo Fisher Scientific), and the indole diterpene was extracted by vigorously shaking the mycelium (≥200 rpm) in 2-butanone for ≥45 minutes.

[0253] Thin layer chromatography The 2-butanone supernatant (containing the extracted indole diterpene) was used for thin layer chromatography (TLC) analysis on a solid phase silica gel 60 aluminum plate (Merck). The indole diterpene was chromatographed with 9:1 chloroform:acetonitrile or 8:2 dichloromethane:acetonitrile and visualized with Ehrlich's reagent (1% (w / v) p-dimethylaminobenzaldehyde in 24% (v / v) HCl and 50% ethanol).

[0254] Liquid chromatography - mass spectrometry Samples were prepared for liquid chromatography-mass spectrometry (LC-MS) from the transformants that gave positive results by TLC. Thus, 1 mL samples of the 2-butanone supernatant (containing the extracted indole diterpenes) were transferred to 1.7 mL microcentrifuge tubes and the 2-butanone was evaporated overnight. The contents were resuspended in 100% acetonitrile and filtered through a 0.2 μm membrane into LC-MS vials. The LC-MS samples were chromatographed on a reversed-phase Thermo Scientific Accucore 2.6 μm C18 (50 x 2.1 mm) column attached to an UltiMate® 3000 standard LC system (Dionex, Thermo Fisher Scientific) running at a flow rate of 0.200 mL / min and eluted with an aqueous acetonitrile solution containing 0.01% formic acid using a multi-step gradient method (Table 14). maXis TM Mass spectra were acquired through in-line analysis on an maXis II quadrupole time-of-flight mass spectrometer (Bruker).

[0255] Large-scale indole diterpene purification for NMR analysis Fungal transformants producing high-level novel indole diterpenes were grown in ≧1 liter of CDYE medium containing trace elements as described below in "Indole Diterpene Production and Extraction". The mycelia were pooled into a 1 liter shot bottle containing a stir bar. 2-butanone was added and the indole diterpenes were extracted overnight with stirring (≧700 rpm). The extract was filtered through Celite® 545 (J.T. Baker®), and after dry loading onto silica while rotary evaporating for rough purification by silica column, it was finally purified by semi-preparative HPLC. A 1 mL aliquot of the crude extract was injected onto a semi-preparative reverse-phase Phenomenex 5μm C18(2) 100Å (250x15mm) column attached to an UltiMate® 3000 standard LC system (Dionex, Thermo Fisher Scientific) that was eluted at a flow rate of 8.00 mL / min. A multi-step gradient method was optimized for the purification of different sets of indole diterpenes. The purity of each indole diterpene was evaluated by LC-MS and the structure was identified by NMR.

[0256] NMR NMR samples were prepared in deuterated chloroform. Compounds were analyzed by standard one-dimensional proton and carbon-13 NMR, two-dimensional correlation spectroscopy (COSY), heteronuclear single quantum correlation spectroscopy (HSQC) and heteronuclear multiple bond correlation spectroscopy (HMBC).

[0257] Tables 1 to 14 referred to in this specification are shown below: Table 1. Functional assignment of predicted genes in the putative nosperic acid gene cluster

[0258]

Table 1-1

[0259]

Table 1-2

[0260] Genes in the IDT cluster are named according to A. nidulans naming convention, where genes are given names that include a three-letter prefix in lowercase indicating the species, written in italic font, followed by a single-letter suffix in uppercase indicating gene function (e.g., paxC). Naming of the corresponding protein products follows the same rules, except that the first letter of the prefix is uppercase and the whole name is written in normal (non-italic) font (e.g., PaxC is the protein product of paxC). Thus, nod names were assigned to each H. prischidum gene in the NAA 10 gene cluster. Asterisks ( * ) are shown after H. prischidum genes that share homology with genes found in the known IDT pathway (>35% amino acid identity of the predicted translation products), and letters corresponding to known confirmed genes are shown, except for nodR (e.g., the protein encoded by nodC shares 52.8% amino acid identity with the protein product of paxC). Genes that do not share homology with known IDT genes were assigned letters that are not shared with any of the confirmed IDT genes. In particular, the cluster contains two sets of paralogous genes (>40% amino acid identity with each other), which we distinguished using numbers (i.e., nodD1 / nodD2 and nodY1 / nodY2). The closest matches were identified using BLASTp (protein-protein BLAST) against a non-redundant protein sequence database with an "expect threshold" set to 10 and a "word size" set to 6. The BLOSUM62 scoring matrix was applied with a gap opening penalty of 11 and a gap extension penalty of 1, with conditional compositional score matrix adjustment.

[0261] Table 2. Similarity matrix of geranylgeranyl transferase (the "C" enzyme)

[0262]

Table 2

[0263] The lightly shaded gray regions indicate the % identity scores for amino acid residues, and the darkly shaded gray regions indicate the % similarity scores. Table 3. Similarity Matrix of 3-Geranylgeranyl Indole Epoxidase (the "M" enzyme)

[0264]

Table 3

[0265] The lightly shaded gray regions indicate the % identity scores for amino acid residues, and the darkly shaded gray regions indicate the % similarity scores. Table 4. Similarity Matrix of Indole Diterpene Cyclase (the "B" enzyme)

[0266]

Table 4

[0267] The lightly shaded gray regions indicate the % identity scores for amino acid residues, and the darkly shaded gray regions indicate the % similarity scores. Table 5. Similarity Matrix of Indole Diterpene Prenyltransferase (the "D" and "E" enzymes compared to NodD1 and NodD2)

[0268]

Table 5

[0269] The lightly shaded gray regions indicate the % identity scores for amino acid residues, and the darkly shaded gray regions indicate the % similarity scores. Table 6. Similarity Matrix of Indole Diterpene FAD-Dependent Oxidative Cyclase (the "O" enzyme)

[0270]

Table 6

[0271] The area of the light gray shadow indicates the % identity score for amino acid residues, and the area of the dark gray shadow indicates the % similarity score. Table 7. Table of fungal species used in this study

[0272]

Table 7

[0273] Table 8. PCR primers for amplification of transcription unit modules (TUMs)

[0274]

Table 8-1

[0275]

Table 8-2

[0276]

Table 8-3

[0277]

Table 8-4

[0278]

Table 8-5

[0279] List the forward and reverse PCR primers used for the amplification of TUM (i.e., promoter (ProUTR), coding sequence (CDS), and terminator (UTRterm)). Shade the primers used for amplifying the TUM fragment for the purpose of adaptation (i.e., removal of internal sites of AarI, BsaI, or BsmBI) in light gray. The template for the amplification of the nod CDS was genomic DNA derived from the Hypoxylon prischidulum strain ATCC® 74245 TM was 25 The template for the amplification of the pax gene TUM was genomic DNA derived from the Penicillium paxilli strain ATCC® 26601 TM (PN2013) 24 The PCR products used to produce the trpC ProUTR module, the nptII CDS module (conferring resistance to geneticin), and the trpC UTRterm module were all amplified from the plasmid pII99 41 Emphasize the BsmBI recognition site (cgtctc) and show the overhangs generated after BsmBI cleavage in a lighter gray shade. The 5' (prefix) and 3' (suffix) nucleotide bases adjacent to each TUM and forming the basis of the addressing system for each MIDAS module are shown in light gray and dark gray, respectively

[0280] Table 9. MIDAS Level 1 plasmid library: Assembly of TUMs in pML1

[0281] [Table 9]

[0282] This table shows the MIDAS level 1 TUMs that the inventors used to assemble the MIDAS level 2 TU (Table 10). The 4-base prefixes and suffixes (from 5' to 3') adjacent to each TUM are shown at the top of the table, both of which bind to the TUM to highlight the sequences that create the MIDAS level 2 TU. These 4-base adjacent regions are shown in light gray (forward address) and dark gray (reverse address) in the primer table (Table 8).

[0283] Table 10. Assembly of TUs in the MIDAS level 2 plasmid library: pML2 destination vector

[0284]

Table 10

[0285] This table shows the construction of the MDIAS level 2 TUs used to assemble the MIDAS level 3 multi-gene plasmids for heterologous expression studies. The names of the generated level 2 entry plasmids are shown in the shaded columns in gray. The TUs are indicated by the CDS they contain, and the TU orientation determined by the pML2 destination vector is shown by the arrowheads in the level 2 description (→ for the forward (F) destination vector and ← for the reverse (R) destination vector).

[0286] Table 11. MIDAS level 3 plasmid library: Multi-gene assembly in pML3

[0287]

Table 11-1

[0288]

Table 11-2

[0289] The table shows the level 2 entry clones and level 3 destination vectors used to construct the multi-gene plasmid. The names of the plasmids produced during each cycle of the level 3 assembly are shown in the column with gray shading. The number of level 3 assembly reactions used to generate the level 3 plasmids is indicated by the numbers in the step column. The TUs are annotated with the names of the CDSs they contain. The TU orientation is indicated by an arrow.

[0290] Table 12. General primer design for amplification of ProUTR, CDS, and UTRterm modules to be cloned into pML1

[0291] [Table 12]

[0292] List the general characteristics of the forward and reverse PCR primers used for amplification of TUMs. The BsmBI recognition site is shown in color (cgtctc), and the overhangs generated after BsmBI digestion are shown by gray shading. The 5' and 3' nucleotide-specific bases adjacent to each TUM and forming the basis of the address system for each MIDAS module are shown in light gray and dark gray, respectively.

[0293] Table 13. Construct the level 3 multi-gene assembly by alternately performing Golden Gate cloning reactions using the TUs assembled in the "white" and "blue" pML2 vectors.

[0294] [Table 13]

[0295] The table shows the cloning steps used to produce a hypothetical multi-gene construct containing four TUs. Each column shows the input plasmids (level 2 entry clones and destination plasmids), the type of Golden Gate reaction used for assembly, the product plasmid, and the type of colonies screened.

[0296] Table 14. Multi-step acetonitrile gradient used for LC-MS analysis of fungal extracts

[0297]

Table 14

[0298] Industrial Applicability The present invention has industrial applicability in the production of indole diterpene compounds, particularly NA. References

[0299]

Chemical formula

[0300]

Chemical formula

[0301]

Chemical formula

[0302]

Chemical formula

Claims

1. An isolated polypeptide comprising the amino acid sequence of NodW (SEQ ID NO: 3) and having P450 oxygenase activity.

2. An isolated polynucleotide encoding the polypeptide of Claim 1.

3. An isolated polynucleotide encoding a polypeptide that comprises at least 70% nucleic acid sequence identity to SEQ ID NO: 2, comprises the amino acid sequence of NodW (SEQ ID NO: 3), and has P450 oxygenase activity.

4. A transcription unit (TU) comprising the isolated polynucleotide of Claim 3.

5. A vector encoding the isolated polypeptide of Claim 1.

6. A vector comprising the isolated polynucleotide of Claim 2 or Claim 3, or the TU of Claim 4.

7. An isolated host cell comprising the isolated polypeptide of Claim 1, the isolated polynucleotide of Claim 2 or Claim 3, the TU of Claim 4, and / or the vector of Claim 5 or Claim 6.

8. A method for producing nosylspolonic acid F (NAF) comprising the step of heterologously expressing in an isolated host cell the isolated polypeptide of Claim 1, the isolated polynucleotide of Claim 2 or Claim 3, the TU of Claim 4, and / or the vector of Claim 5 or Claim 6, wherein the isolated host cell is Penicillium paxilli transformed with a nucleic acid sequence that expresses the NodM polypeptide.

9. A method for producing at least one Hypoxylon species polypeptide comprising the step of heterologously expressing in an isolated host cell the isolated polynucleotide of Claim 2 or Claim 3, or the TU of Claim 4, and / or the vector of Claim 5 or Claim 6.

10. A method for producing NAF comprising the step of contacting a recombinant cell transformed with a nucleic acid sequence that comprises at least 70% sequence identity to SEQ ID NO: 2 and encodes a polypeptide that comprises the amino acid sequence of NodW (SEQ ID NO: 3) and has P450 oxygenase activity with a carbohydrate comprising a substrate such that the substrate, emindole SB, is metabolized to nosylspolonic acid F (NAF).

11. An isolate of Hypoxylon pulicicidum comprising at least one heterologous nucleic acid sequence encoding an enzyme in the biosynthetic pathway leading to NAA 10, wherein the heterologous nucleic acid sequence comprises the amino acid sequence of NodW (SEQ ID NO: 3) and encodes a polypeptide having P450 oxygenase activity, said isolate.

12. A method for producing NAA 10 in Hypoxylon pulicicidum, comprising the step of expressing at least one heterologous nucleic acid sequence, wherein the at least one heterologous nucleic acid sequence comprises at least 70% nucleic acid sequence identity to SEQ ID NO: 2 and comprises the amino acid sequence of NodW (SEQ ID NO: 3) in the biosynthetic pathway leading to NAA 10 and encodes a polypeptide having P450 oxygenase activity, said method.

13. The isolated polynucleotide of claim 2, comprising at least 75% nucleic acid sequence identity to SEQ ID NO:

2.

14. The isolated polynucleotide of claim 2, comprising at least 80% nucleic acid sequence identity to SEQ ID NO:

2.

15. The isolated polynucleotide of claim 2, comprising at least 85% nucleic acid sequence identity to SEQ ID NO:

2.

16. The isolated polynucleotide of claim 2, comprising at least 90% nucleic acid sequence identity to SEQ ID NO:

2.

17. The isolated polynucleotide of claim 2, comprising at least 95% nucleic acid sequence identity to SEQ ID NO:

2.

18. The isolated polynucleotide of claim 2, comprising at least 99% nucleic acid sequence identity to SEQ ID NO: 2.