Human therapeutic targets and their modulators
By studying embedded target genes (ETaGs) in fungal biosynthetic gene clusters, we solved the problem of identifying druggable targets in the human proteome and achieved the effective identification and design of human targets with therapeutic significance.
Patent Information
- Application Number
- CN201880072708.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-09-14
- Filing Date
- 2018-09-14
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2038-09-14
AI Technical Summary
Identifying druggable targets in the human proteome is a major challenge. Only about 2% of human proteins have been successfully targeted using current technologies, and most human proteins are difficult to regulate with drugs.
By studying non-biosynthetic genes in fungal biosynthetic gene clusters, especially embedded target genes (ETaGs), which have homology to human genes and are co-regulated with biosynthetic gene clusters, we provide new methods for identifying and designing human targets.
This method provides a new way to identify and characterize human targets of therapeutic interest, improves druggability, and expands the scope of target selection.
Smart Images

Figure BDA0002484093520000601 
Figure BDA0002484093520000611 
Figure BDA0002484093520000621
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Application No. 62 / 558,744, filed September 14, 2017, which is incorporated herein by reference in its entirety. Background Art
[0003] Identifying so-called "druggable" targets in the human proteome has been described as a "major challenge." See, e.g., Dixon et al Curr. Opin. Chem. Biol. 13:549, 2009. As of 2011, reports estimate that only about 2% of human proteins have been successfully targeted by approved drugs, and further, only 10% to 15% of human proteins are even easily targetable (i.e., are "druggable"). See, e.g., Stockwell Sci. Am 305:20, 2011. Summary of the Invention
[0004] There is increasing evidence that some microbial biosynthetic gene clusters sometimes contain genes that do not appear to be involved in the synthesis of the relevant biosynthetic products produced by the enzymes encoded by the cluster (referred to herein as "passenger" genes). In some cases, such passenger genes are described as "self-protective" because they encode proteins that apparently make the host organism resistant to the relevant biosynthetic products. For example, in some cases, genes encoding transporters of biosynthetic products, detoxification enzymes that act on biosynthetic products, or resistant variants of proteins whose activities are targeted by biosynthetic products have been reported. See, for example, Cimermancic et al Cell 158: 412, 2014; Keller Nat. Chem. Biol. 11: 671, 2015. Researchers have proposed that the identification of such genes and their functions can be used to determine the role of the biosynthetic products synthesized by the enzymes of the cluster. See, e.g., Yeh et al. ACS Chem. Biol. 11:2275, 2016; Tang et al. ACS Chem. Biol. 10:2841, 2015; Regueira et al. Appl, Environ. Microbiol. 77:3035, 2011; Kennedy et al., Science 284:1368, 1999; Lowtheret al., Proc. Natl. Acad. Sci. USA 95: 12153, 1998; Abe et al., Mol. Genet. Genomics 268: 130, 2002.
[0005] The present disclosure provides, inter alia, different perspectives on non-biosynthetic genes present in biosynthetic gene clusters as described herein, or in proximity to biosynthetic genes in such clusters, and provides new insights into the potential usefulness of certain such genes in human therapeutics. In some embodiments, the present disclosure provides techniques for utilizing such insights to develop and / or improve human therapeutics.
[0006] The present disclosure provides, inter alia, the insight that certain non-biosynthetic genes present in biosynthetic gene clusters or in proximity to biosynthetic genes in such clusters, and particularly in biosynthetic gene clusters of eukaryotes (e.g., fungi, as compared to bacteria), may represent homologs of human genes that represent targets of therapeutic interest. The present disclosure defines parameters for characterizing such non-biosynthetic genes of interest, referred to herein as "embedded target genes" or "ETaGs." The present disclosure provides, inter alia, techniques for identifying and / or characterizing ETaGs, databases comprising biosynthetic gene clusters and / or ETaG gene sequences (and optionally associated annotations), systems for identifying and / or characterizing human target genes corresponding to ETaGs, and methods for making and / or using such human target genes and / or systems comprising and / or expressing such human target genes.
[0007] The present disclosure provides additional insight: the relationship between ETaG and its associated biosynthetic gene cluster (a biosynthetic gene cluster comprising a biosynthetic gene with the ETaG in the vicinity of the biosynthetic gene) provides information for identifying, designing, and / or characterizing effective modulators of the corresponding human target gene. The present disclosure provides techniques for such identification, design, and / or characterization, and also provides agents that achieve modulation of the associated human target gene, as well as methods for providing and / or using such agents.
[0008] As described above, the present disclosure encompasses the insight that ETaGs can be used as functional homologs (e.g., orthologs) of human targets with medical (e.g., therapeutic) relevance. According to the present disclosure, the sequences of passenger (i.e., non-biosynthetic) genes within a biosynthetic gene cluster of a eukaryote (e.g., fungus) or in a neighboring region relative to a biosynthetic gene in the cluster can be compared with the sequences of human genes. For the compared sequences, nucleic acid sequence similarity, peptide sequence similarity, and / or phylogenetic relationships can be determined (e.g., quantitatively assessed and / or visualized by a phylogenetic tree). Alternatively or in addition, the conservation of known structural and / or protein effector elements can be assessed. In some embodiments, those passenger genes that have relatively high homology to human sequences and / or conserved structural and / or protein effector elements can be prioritized as ETaGs that are as meaningful as human drug targets.
[0009] In some embodiments, the present disclosure provides a method comprising the steps of:
[0010] a set of query nucleic acid sequences, each of which is present in a fungal strain and comprises a biosynthetic gene cluster; and
[0011] An embedded target gene (ETaG) sequence is identified within at least one fungal nucleic acid sequence, wherein the embedded target gene (ETaG) sequence is characterized in that:
[0012] within a vicinity relative to at least one gene in the cluster; and
[0013] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0014] Generally speaking, a biosynthetic gene cluster comprises one or more biosynthetic genes. In some embodiments, a biosynthetic gene cluster comprises one or more biosynthetic genes and one or more non-biosynthetic genes. In some embodiments, non-biosynthetic genes are regulatory, such as transcription factors. In some embodiments, in a biosynthetic gene cluster identified by bioinformatics, non-biosynthetic genes can be hypothetical genes. In some embodiments, the boundary of a biosynthetic gene cluster is defined by bioinformatics methods such as antiSMASH. In some embodiments, biosynthetic genes and non-biosynthetic genes are specified based on bioinformatics. In some embodiments, non-biosynthetic genes may have biosynthetic functions, even if they are identified as non-biosynthetic genes (and / or are labeled as non-biosynthetic genes in this disclosure) by bioinformatics methods.
[0015] In some embodiments, the present disclosure provides a method comprising the steps of:
[0016] a set of query nucleic acid sequences, each of which is present in a fungal strain and comprises a biosynthetic gene cluster; and
[0017] An embedded target gene (ETaG) sequence is identified within at least one fungal nucleic acid sequence, wherein the embedded target gene (ETaG) sequence is characterized in that:
[0018] within a vicinity relative to at least one biosynthetic gene in the cluster; and
[0019] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0020] In some embodiments, the present disclosure encompasses the recognition that ETaGs from eukaryotic fungi can have a higher similarity to mammalian genes than, for example, their counterparts in prokaryotes such as certain bacteria, if any. In some embodiments, fungi contain more therapeutically relevant ETaGs and / or contain more therapeutically relevant ETaGs than organisms that are more evolutionarily distant from humans.
[0021] In some embodiments, the present disclosure provides a method comprising the steps of:
[0022] a set of query nucleic acid sequences, each of which is present in a fungal strain and comprises a biosynthetic gene cluster; and
[0023] An embedded target gene (ETaG) sequence is identified within at least one fungal nucleic acid sequence, wherein the embedded target gene (ETaG) sequence is characterized in that:
[0024] within a vicinity relative to at least one gene in the cluster;
[0025] is homologous to an expressed mammalian nucleic acid sequence; and
[0026] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0027] In some embodiments, the present disclosure provides a method comprising the steps of:
[0028] a set of query nucleic acid sequences, each of which is present in a fungal strain and comprises a biosynthetic gene cluster; and
[0029] An embedded target gene (ETaG) sequence is identified within at least one fungal nucleic acid sequence, wherein the embedded target gene (ETaG) sequence is characterized in that:
[0030] within a vicinity relative to at least one biosynthetic gene in the cluster;
[0031] is homologous to an expressed mammalian nucleic acid sequence; and
[0032] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0033] In some embodiments, the adjacent region is no more than 1 to 100 kb upstream or downstream of a biosynthetic gene in a cluster, e.g., no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, or 100 kb. In some embodiments, the adjacent region is no more than 1 to 100 kb upstream or downstream of a biosynthetic gene in a cluster, e.g., no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 kb. In some embodiments, the ETaG is within a biosynthetic gene cluster. In some embodiments, the adjacent region is between two biosynthetic genes in a biosynthetic gene cluster.
[0034] In some embodiments, the ETaG sequence is homologous to a mammalian nucleic acid sequence. In some embodiments, the mammalian sequence is a human nucleic acid sequence. In some embodiments, the ETaG sequence is homologous to a human nucleic acid sequence. In some embodiments, the ETaG sequence is homologous to an expressed mammalian nucleic acid sequence. In some embodiments, the ETaG sequence is homologous to an expressed human nucleic acid sequence. In some embodiments, mammalian nucleic acids, such as human nucleic acid sequences, are associated with human diseases, disorders, or conditions. In some embodiments, such human nucleic acid sequences are existing targets of therapeutic significance. In some embodiments, such human nucleic acid sequences are new targets of therapeutic significance. In some embodiments, such human nucleic acid sequences are targets that were previously thought to be difficult to target, for example, by small molecules. In some embodiments, the biosynthetic product or its analog produced by the enzyme encoded by the relevant biosynthetic gene cluster is a modulator (e.g., activator, inhibitor, etc.) of the human target.
[0035] In some embodiments, the ETaG sequence is homologous to an expressed mammalian nucleic acid sequence in that its sequence, or a portion thereof, is at least 50%, 60%, 70%, 80%, or 90% identical to the sequence, or a portion thereof, of the expressed mammalian nucleic acid sequence. In some embodiments, the ETaG sequence is homologous to a mammalian nucleic acid sequence in that the mRNA produced by ETaG, or a portion thereof, is homologous to the mRNA, or a portion thereof, of the mammalian nucleic acid sequence. In some embodiments, the homologous portion is at least 50, 100, 150, or 200 base pairs in length. In some embodiments, the homologous portion encodes a protein or a conserved portion of a protein that is conserved from fungi to mammals, such as a protein domain, a collection of residues associated with a function (e.g., interaction with another molecule (e.g., protein, small molecule, etc.), enzymatic activity, etc.), and the like.
[0036] In some embodiments, the ETaG sequence is homologous to a mammalian nucleic acid sequence in that the product encoded by ETaG, or a portion thereof, is homologous to a product encoded by a mammalian nucleic acid sequence, or a portion thereof. In some embodiments, the ETaG sequence is homologous to a mammalian nucleic acid sequence in that the protein encoded by ETaG, or a portion thereof, is homologous to a protein encoded by a mammalian nucleic acid sequence, or a portion thereof. In some embodiments, the ETaG sequence is homologous to a mammalian nucleic acid sequence in that a portion of the protein encoded by ETaG is homologous to a portion of the protein encoded by the mammalian nucleic acid sequence.
[0037] In some embodiments, a portion of a protein is a protein domain. In some embodiments, a protein domain is an enzyme domain. In some embodiments, a protein domain interacts with one or more agents such as small molecules, lipids, carbohydrates, nucleic acids, proteins, and the like.
[0038] In some embodiments, a portion of a protein is a functional domain and / or a structural domain that defines the protein family to which the protein belongs. Amino acid residues within a particular catalytic domain or structural domain that defines a protein family can be selected based on predicted subfamily domain architecture and optionally validated by various assays for use in homology comparisons.
[0039] In some embodiments, a portion of a protein is a collection of continuous or discontinuous key residues that are important for the function of the protein. In some embodiments, the function is enzymatic activity, and a portion of a protein is a collection of residues required for that activity. In some embodiments, the function is enzymatic activity, and a portion of a protein is a collection of residues that interact with a substrate, an intermediate, or a product. In some embodiments, the collection of residues interacts with a substrate. In some embodiments, the collection of residues interacts with an intermediate. In some embodiments, the collection of residues interacts with a product.
[0040] In some embodiments, the function is to interact with one or more agents, such as small molecules, lipids, carbohydrates, nucleic acids, proteins, etc., and a portion of the protein is a set of residues required for the interaction. In some embodiments, the residues in the set are each independently in contact with an interacting agent. For example, in some embodiments, each residue in the set is independently in contact with an interacting small molecule. In some embodiments, the protein is a kinase and the interacting small molecule is or comprises a core base, and the residues in the set are each independently in contact with the core base by, for example, hydrogen bonding, electrostatic force, van der Waals force (van der Waals force), aromatic stacking, etc. In some embodiments, the interacting agent is another macromolecule. In some embodiments, the interacting agent is a nucleic acid. In some embodiments, the residues in the set are those that are in contact with interacting nucleic acids, such as those in transcription factors. In some embodiments, the residues in the set are those that are in contact with interacting proteins.
[0041] In some embodiments, a portion of a protein is or comprises an essential structural element for protein effector recruitment and / or binding, such as based on the tertiary protein structure of a human target.
[0042] Portions of proteins, such as protein domains, groups of residues responsible for a biological function, etc., can be conserved between species, for example, in some embodiments, from fungi to humans, as exemplified in this disclosure.
[0043] In some embodiments, protein homology is measured based on exact identity, such as the same amino acid residue at a given position. In some embodiments, homology is measured based on one or more properties, such as amino acid residues having one or more identical or similar properties (e.g., polarity, non-polarity, hydrophobicity, hydrophilicity, size, acidity, basicity, aromaticity, etc.). Exemplary methods for assessing homology are well known in the art and can be used in accordance with the present disclosure, for example, MUSCLE, TCoffee, ClustalW, etc.
[0044] In some embodiments, the protein encoded by ETaG or a portion thereof (e.g., those described in the present disclosure) is at least 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%, or 100% (where 100% is identical) homologous to a protein encoded by a mammalian nucleic acid sequence, or a portion thereof. In some embodiments, the protein encoded by ETaG or a portion thereof is at least 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%, or 100% homologous to a protein encoded by an expressed mammalian nucleic acid sequence, or a portion thereof.
[0045] In some embodiments, ETaG is co-regulated with at least one biosynthetic gene in a biosynthetic gene cluster. In some embodiments, ETaG is co-regulated with two or more genes in a biosynthetic gene cluster. In some embodiments, ETaG is co-regulated with a biosynthetic gene cluster in that the expression of ETaG is increased or turned on when a biosynthetic product produced by an enzyme encoded by the biosynthetic gene cluster (the biosynthetic product of the biosynthetic gene cluster) is produced. In some embodiments, ETaG is co-regulated with a biosynthetic gene cluster in that the expression of ETaG is increased or turned on when the level of the biosynthetic product of the biosynthetic gene cluster is increased.
[0046] In some embodiments, the ETaG gene sequence optionally has greater than about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99% homology to one or more gene sequences in the same genome. In some embodiments, the ETaG gene sequence optionally has greater than about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99% homology to 2, 3, 4, 5, 6, 7, 8, 9, or more gene sequences in the same genome. In some embodiments, the homology is greater than 10%. In some embodiments, the homology is greater than 20%. In some embodiments, the homology is greater than 30%. In some embodiments, the homology is greater than 40%. In some embodiments, the homology is greater than 50%. In some embodiments, the homology is greater than 60%. In some embodiments, the homology is greater than 70%. In some embodiments, the homology is greater than 80%. In some embodiments, the homology is greater than 90%. Certain examples are provided in the accompanying drawings.
[0047] In some embodiments, the ETaG gene sequence optionally has no more than about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99% identity to any expressed gene sequence in a collection of at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% fungal nucleic acid sequences from different fungal strains and comprising homologous biosynthetic gene clusters. In some embodiments, the ETaG gene sequence optionally has no more than about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99% identity to at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% of a fungal gene sequence from a different fungal strain within the contiguous region relative to the biosynthetic genes in the homologous biosynthetic gene cluster. In some embodiments, the ETaG gene sequence optionally has no more than about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99% identity to at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% of a fungal gene sequence from a different fungal strain within the contiguous region relative to the biosynthetic genes in the homologous biosynthetic gene cluster. In some embodiments, the ETaG gene sequence optionally has no more than about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99% identity with any expressed gene sequence in any fungal nucleic acid sequence in the collection, the fungal nucleic acid sequence being from a different fungal strain and comprising a homologous biosynthetic gene cluster. In some embodiments, the ETaG gene sequence optionally has no more than about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99% identity with any expressed gene sequence from a different fungal strain within the vicinity of the biosynthetic genes in the homologous biosynthetic gene cluster. In some embodiments, no more than about 10% identity. In some embodiments, no more than about 20% identity. In some embodiments, no more than about 30% identity. In some embodiments, no more than about 40% identity. In some embodiments, no more than about 50% identity.In some embodiments, no more than about 60% identity. In some embodiments, no more than about 70% identity. In some embodiments, no more than about 80% identity. In some embodiments, no more than about 90% identity.
[0048] In some embodiments, a human target gene and / or its product is readily regulated by a biosynthetic product of a biosynthetic gene cluster or an analog thereof, wherein the human target gene has its cognate ETaG within the biosynthetic gene cluster or in a neighboring region relative to the biosynthetic genes in the cluster. In some embodiments, a protein encoded by a human target gene is readily regulated by a biosynthetic product of a biosynthetic gene cluster or an analog thereof, wherein the human target gene has its cognate ETaG within the biosynthetic gene cluster or in a neighboring region relative to the biosynthetic genes in the cluster. Thus, in some embodiments, the present disclosure not only provides novel human targets, but also provides methods and agents for regulating such human targets.
[0049] In some embodiments, the present disclosure provides techniques, e.g., methods, databases, systems, etc., for identifying ETaGs and / or their medical relevance, e.g., their therapeutic relevance. In some embodiments, the present disclosure provides databases, optionally with multiple annotations, configured for efficient identification, retrieval, use, etc. of ETaGs, related biosynthetic gene clusters, related biosynthetic products of biosynthetic gene clusters and / or analogs thereof, related homologous mammalian nucleic acid sequences (e.g., human genes), etc. The present disclosure provides, inter alia, databases and / or sequences configured to improve, e.g., computational efficiency and / or accuracy of ETaG identification.
[0050] For example, in some embodiments, the database provided is constructed so as to identify and annotate all biosynthetic gene clusters. The nucleotide sequences of these clusters are then calculated, removed, and databased from the remaining nucleotide sequences in the fungal genome. The resulting biosynthetic gene cluster database is then used for an ETaG search. In particular, when using such a database to identify a hit in an ETaG search, the hit is an ETaG because only the sequence in the biosynthetic cluster (or its adjacent region) is retrieved. Separating the biosynthetic gene cluster sequence from the full genome sequence improves the signal-to-noise ratio and greatly accelerates the ETaG search process. In particular, compared with using the database provided, searching for ETaG in the full fungal genome sequence frequently results in false positives of the hits identified, which are "housekeeping" genes that are located in the genome but not in the biosynthetic gene cluster or its adjacent region. In some embodiments, the hits identified, for example, ETaG, from the technology provided (e.g., method, database, etc.) are not housekeeping genes. In some embodiments, a hit identified from the provided techniques, such as ETaG, is or comprises a sequence that shares homology with a second nucleic acid sequence (e.g., a gene) in the same genome or a portion thereof. The sequence homology of the sequences of the present disclosure can be at least 20%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5%. In some embodiments, the homology is at least 50%; in some embodiments, at least 60%; in some embodiments, at least 70%; in some embodiments, at least 75%; in some embodiments, at least 80%; in some embodiments, at least 85%; in some embodiments, at least 90%; and in some embodiments, at least 95%. A portion of a sequence of the present disclosure may comprise at least 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, or 1000 amino acid residues (for protein sequences) or nucleobases (for nucleic acid sequences). In some embodiments, a portion of a nucleic acid sequence is at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, or 1000 nucleobases in length. In some embodiments, the length is at least 20 nucleobases. In some embodiments, the length is at least 30 nucleobases. In some embodiments, the length is at least 40 nucleobases. In some embodiments, the length is at least 50 nucleobases.In some embodiments, the length is at least 100 nucleobases. In some embodiments, the length is at least 150 nucleobases. In some embodiments, the length is at least 200 nucleobases. In some embodiments, the length is at least 300 nucleobases. In some embodiments, the length is at least 400 nucleobases. In some embodiments, the length is at least 500 nucleobases. In some embodiments, the hit identified from the provided technology, such as ETaG, is or comprises a sequence encoding a product (e.g., a protein) that has a common homology (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5%) with a product encoded by a second nucleic acid sequence (e.g., a gene) in the same genome, or a portion thereof (e.g., a set of key residues of a protein as described in the present disclosure, a protein domain, etc.). As described herein, homology / similarity can be assessed using a variety of techniques understood by those skilled in the art. In some embodiments, the second nucleic acid sequence is or comprises a housekeeping gene. In some embodiments, the second nucleic acid sequence is shared between two or more species. In some embodiments, ETaG is homologous to the second nucleic acid sequence but differs from the second nucleic acid sequence in that ETaG encodes a product (e.g., a protein) that confers resistance to the product (e.g., a small molecule) of its corresponding biosynthetic cluster, while the second nucleic acid sequence does not.
[0051] In some embodiments, the present disclosure provides a system comprising:
[0052] One or more non-transitory machine-readable storage media storing data representing a set of nucleic acid sequences, each of which is present in a fungal strain and comprises a biosynthetic gene cluster.
[0053] In some embodiments, the present disclosure provides a system comprising:
[0054] One or more non-transitory machine-readable storage media storing data representing a set of nucleic acid sequences, each of which is or comprises an ETaG sequence.
[0055] In some embodiments, at least 10, 20, 50, 100, 200, or 500, or at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%, or all of the nucleic acid sequences in the collection contain indexed and / or annotated ETaGs. In some embodiments, the provided system can greatly improve computational efficiency because it is configured to greatly reduce the amount of data to be processed. For example, instead of processing all genomic or biosynthetic gene cluster sequence data for one or more (in some cases, hundreds or thousands or even more) fungal genomes to retrieve ETaGs, the provided system can only retrieve genes that are indexed / labeled as ETaGs, thereby saving the time and cost of processing sequences that are not indexed as ETaGs. Additionally or alternatively, an ETaG can be independently annotated with information such as its associated biosynthetic gene cluster (which comprises the biosynthetic gene to which the ETaG is located in a neighborhood relative to the biosynthetic gene), the structure of the biosynthetic product of the associated biosynthetic gene cluster, and / or the human homolog of the ETaG. In some embodiments, at least 10, 20, 50, 100, 200, or 500, or at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%, or all of the ETaGs in the collection are independently annotated with at least one of: an associated biosynthetic gene cluster and a human homolog of the ETaG. In some embodiments, at least 10, 20, 50, 100, 200, or 500, or at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%, or all of the ETaGs in the collection are independently annotated with at least one of the following: related biosynthetic gene clusters, biosynthetic products of related biosynthetic gene clusters, and human homologs of the ETaGs. In some embodiments, by utilizing ETaG indexes and annotations to structure sequence data, the provided system can provide numerous advantages. For example, in some embodiments, the provided system allows for rapid access to ETaGs with useful relevant information, such as their related biosynthetic gene clusters and human homologs, and vice versa, while maintaining data size and cost.
[0056] In some embodiments, the provided methods and systems can be used to identify and / or characterize human targets because, among other things, the provided methods and systems provide a link between a biosynthetic gene cluster, ETaG, and a human target gene. In some embodiments, the present disclosure provides insights into targets that were previously considered undruggable by providing homologous ETaGs and related biosynthetic gene clusters in fungi. In some embodiments, the present disclosure greatly improves the druggability of targets previously considered undruggable, in some cases essentially converting them into druggable targets, through, for example, their homologous ETaGs in fungi, related biosynthetic gene clusters, biosynthetic products of the related biosynthetic gene clusters (which can be used directly as modulators of the human target, and / or their analogs can be used as modulators of the human target).
[0057] In some embodiments, the present disclosure provides methods for identifying and / or characterizing human targets of a biosynthetic product of a biosynthetic gene cluster or an analog of the product.
[0058] In some embodiments, the present disclosure provides a method comprising:
[0059] Identifying a human homolog of ETaG that is within a proximal region relative to at least one gene in a biosynthetic gene cluster or within a proximal region relative to at least one gene in a second biosynthetic gene cluster that encodes an enzyme that produces a biosynthetic product produced by an enzyme encoded by the biosynthetic gene cluster; and
[0060] Optionally, the effect of the biosynthetic products produced by the enzymes encoded by the biosynthetic gene cluster, or analogs of the products, on the human target is assayed.
[0061] In some embodiments, the present disclosure provides a method comprising:
[0062] identifying a human homolog of ETaG that is within a proximal region relative to at least one biosynthetic gene in a biosynthetic gene cluster or within a proximal region relative to at least one biosynthetic gene in a second biosynthetic gene cluster that encodes an enzyme that produces a biosynthetic product produced by an enzyme encoded by the biosynthetic gene cluster; and
[0063] Optionally, the effect of the biosynthetic products produced by the enzymes encoded by the biosynthetic gene cluster, or analogs of the products, on the human target is assayed.
[0064] In some embodiments, the present disclosure provides a method comprising:
[0065] Identifying a human homolog of ETaG within a contiguous region relative to at least one gene in the biosynthetic gene cluster; and
[0066] Optionally, the effect of the biosynthetic products produced by the enzymes encoded by the biosynthetic gene cluster, or analogs of said products, on the human target is assayed.
[0067] In some embodiments, the present disclosure provides a method comprising:
[0068] identifying a human homolog of ETaG within a proximal region relative to at least one biosynthetic gene in the biosynthetic gene cluster; and
[0069] Optionally, the effect of the biosynthetic products produced by the enzymes encoded by the biosynthetic gene cluster, or analogs of said products, on the human target is assayed.
[0070] In some embodiments, for a biosynthetic gene cluster that does not contain a biosynthetic gene containing ETaG in its vicinity, a mammalian target, such as a human target, of a product of such a biosynthetic gene cluster (and / or its analog) can be identified by the presence of ETaG in its vicinity relative to a biosynthetic gene in a second biosynthetic gene cluster that encodes an enzyme that produces the same biosynthetic product. In some embodiments, the second biosynthetic gene cluster is in a different organism. In some embodiments, the second biosynthetic gene cluster is in a different fungal strain.
[0071] In some embodiments, the present disclosure provides a method for identifying and / or characterizing a human target of a biosynthetic product of a biosynthetic gene cluster or an analog of the product, comprising:
[0072] identifying a human homolog of ETaG that is within a proximal region relative to at least one biosynthetic gene in a second biosynthetic gene cluster that encodes an enzyme that produces the same biosynthetic product produced by the enzyme encoded by the biosynthetic gene cluster; and
[0073] Optionally, the effect of the biosynthetic products produced by the enzymes encoded by the biosynthetic gene cluster, or analogs of the products, on the human target is assayed.
[0074] In some embodiments, the provided technology can be used to assess the interaction of a human target with a compound. In some embodiments, the present disclosure provides a method for assessing the interaction of a human target with a compound, comprising:
[0075] The nucleic acid sequence of the human target or a nucleic acid sequence encoding the human target is compared to a collection of nucleic acid sequences comprising one or more ETaGs.
[0076] In some embodiments, homology to ETaG (at the nucleic acid level or protein level, including portions thereof) relates to the biosynthetic gene cluster associated with the ETaG and its biosynthetic products. In some embodiments, such association between the biosynthetic product and the human target indicates interaction and / or regulation of the human target or the product encoded thereby. In some embodiments, such biosynthetic product interacts with and / or regulates the human target or the product encoded thereby.
[0077] In some embodiments, the provided technology can be used to design and / or provide modulators of human targets because, among other things, the provided technology provides a link between the biosynthetic gene cluster, ETaG, and the human target gene.
[0078] In some embodiments, the disclosure provides compounds that are products of enzymes encoded by a biosynthetic gene cluster, wherein an ETaG is present in proximity to at least one gene in the biosynthetic gene cluster, said ETaG:
[0079] is homologous to a human target or a nucleic acid sequence encoding a human target; and
[0080] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0081] In some embodiments, the provided compounds are products of enzymes encoded by the provided biosynthetic gene clusters. In some embodiments, the provided compounds are analogs of products of enzymes encoded by the provided biosynthetic gene clusters. In some embodiments, the provided biosynthetic gene clusters comprise Figures 5 to 12 and one or more biosynthetic genes shown in one of 20 to 39. In some embodiments, the biosynthetic gene cluster provided is Figures 5 to 12 and one of 20 to 39. In some embodiments, provided compounds are Figures 5 to 12 and the product of an enzyme encoded by a provided biosynthetic gene cluster as shown in one of 20 to 39. In some embodiments, the provided compound is Figures 5 to 12 and the provided biosynthetic gene cluster shown in one of 20 to 39 or comprising Figures 5 to 12 and a product of a biosynthetic gene cluster of one or more biosynthetic genes shown in one of 20 to 39. In some embodiments, a provided compound is Figures 5 to 12 and a product of a provided biosynthetic gene cluster as shown in one of 20 to 39. In some embodiments, a provided compound is a compound comprising Figures 5 to 12 and one or more of the biosynthetic genes shown in any one of 20 to 39. In some embodiments, the provided compound is Figures 5 to 12and an analog of the product of an enzyme encoded by a provided biosynthetic gene cluster as shown in one of 20 to 39. In some embodiments, the provided compound is a compound comprising Figures 5 to 12 and an analog of the product of the provided biosynthetic gene cluster of one or more biosynthetic genes shown in one of 20 to 39. In some embodiments, the provided compound regulates the function of the human target. In some embodiments, the disclosure provides pharmaceutical compositions of the provided compounds. In some embodiments, the disclosure provides pharmaceutical compositions comprising the provided compounds or pharmaceutically acceptable salts thereof. In some embodiments, the disclosure provides pharmaceutical compositions comprising the provided compounds or pharmaceutically acceptable salts thereof, and a pharmaceutically acceptable carrier. In some embodiments, the provided compound in the provided composition is an analog of the product of an enzyme encoded by the biosynthetic gene cluster or a salt thereof. In some embodiments, the provided compound in the provided composition is a non-natural salt of the product of an enzyme encoded by the biosynthetic gene cluster.
[0082] In some embodiments, the present disclosure provides methods for identifying and / or characterizing a modulator of a human target comprising:
[0083] A product or an analog thereof is provided, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster, wherein an ETaG is present in the vicinity of at least one gene in the biosynthetic gene cluster, wherein the ETaG:
[0084] is homologous to a human target or a nucleic acid sequence encoding a human target; and
[0085] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0086] In some embodiments, the present disclosure provides methods for identifying and / or characterizing a modulator of a human target comprising:
[0087] A product or an analog thereof is provided, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster, wherein an ETaG is present in a proximal region relative to at least one biosynthetic gene in the biosynthetic gene cluster, wherein the ETaG:
[0088] is homologous to a human target or a nucleic acid sequence encoding a human target; and
[0089] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0090] In some embodiments, the present disclosure provides methods for modulating a human target comprising:
[0091] A product or an analog thereof is provided, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster, wherein an ETaG is present in the vicinity of at least one gene in the biosynthetic gene cluster, wherein the ETaG:
[0092] is homologous to a human target or a nucleic acid sequence encoding a human target; and
[0093] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0094] In some embodiments, the present disclosure provides methods for modulating a human target comprising:
[0095] A product or an analog thereof is provided, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster, wherein an ETaG is present in a proximal region relative to at least one biosynthetic gene in the biosynthetic gene cluster, wherein the ETaG:
[0096] is homologous to a human target or a nucleic acid sequence encoding a human target; and
[0097] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0098] In some embodiments, the present disclosure provides methods for treating a condition, disorder, or disease associated with a human target, comprising administering a biosynthetic product or an analog thereof to a subject susceptible to or suffering from the condition, disorder, or disease, wherein:
[0099] The biosynthetic product is of a biosynthetic gene cluster, wherein an ETaG is present in the vicinity of at least one gene in the biosynthetic gene cluster, said ETaG:
[0100] is homologous to a human target or a nucleic acid sequence encoding a human target; and
[0101] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0102] In some embodiments, the present disclosure provides methods for treating a condition, disorder, or disease associated with a human target, comprising administering a biosynthetic product or an analog thereof to a subject susceptible to or suffering from the condition, disorder, or disease, wherein:
[0103] The biosynthetic product is of a biosynthetic gene cluster, wherein an ETaG is present in the vicinity of at least one biosynthetic gene in the biosynthetic gene cluster, said ETaG:
[0104] is homologous to a human target or a nucleic acid sequence encoding a human target; and
[0105] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0106] In some embodiments, the human target is a Ras protein. In some embodiments, the human target comprises a RasGEF domain. In some embodiments, the human target comprises a RasGAP domain.
[0107] In some embodiments, ETaG is identified by provided methods.
[0108] In some embodiments, the product (e.g., biosynthetic product) is produced by fungi. In some embodiments, the product is acyclic. In some embodiments, the product is a polyketide. In some embodiments, the product is a terpene compound. In some embodiments, the product is non-ribosomal synthesis.
[0109] In some embodiments, analog is a substance that has one or more specific structural features, elements, components or parts in common with a reference substance. Typically, analog shows significant structural similarity with a reference substance, such as shared core or shared structure, but is also different in some discrete aspects. In some embodiments, analog is a substance that can be produced by a reference substance, such as by carrying out chemical operations on the reference substance. In some embodiments, analog is a substance that can be produced by a synthetic process that is substantially similar to (for example, having multiple steps in common with) the synthetic process that produces the reference substance. In some embodiments, analog is produced by or can be produced by a synthetic process different from the synthetic process that produces the reference substance. In some embodiments, the analog of a substance is a substance that is substituted at one or more locations in its substitutable position.
[0110] In some embodiments, the analog of the product comprises the structural core of the product. In some embodiments, the biosynthetic product is cyclic, e.g., monocyclic, bicyclic, or polycyclic, and the structural core of the product is or comprises a monocyclic, bicyclic, or polycyclic ring system. In some embodiments, the product is or comprises a polypeptide, and the structural core is the backbone of the polypeptide. In some embodiments, the product is or comprises a polyketide, and the structural core is the backbone of the polyketide.
[0111] In some embodiments, an analog is a substituted biosynthetic product.In some embodiments, an analog is or comprises a structural core substituted with one or more substituents as described herein.
[0112] In some embodiments, the present disclosure provides compositions of biosynthetic products of provided biosynthetic gene clusters or analogs thereof, wherein ETaG is present in a region adjacent to at least one gene in the biosynthetic gene cluster. In some embodiments, the provided compositions are pharmaceutical compositions. In some embodiments, the provided pharmaceutical compositions comprise: a pharmaceutically acceptable salt of a biosynthetic product of a provided biosynthetic gene cluster or an analog thereof, wherein ETaG is present in a region adjacent to at least one gene in the biosynthetic gene cluster; and a pharmaceutically acceptable carrier.
[0113] In some embodiments, two events or entities are associated with each other if the presence, level, and / or form of one event or entity is correlated with the presence, level, and / or form of another event or entity. For example, a particular entity (e.g., a polypeptide, genetic signature, metabolite, microorganism, etc.) is considered to be associated with a disease, disorder, or condition if the presence, level, and / or form of the particular entity is correlated with the incidence and / or susceptibility of the particular disease, disorder, or condition (e.g., in a relevant population).
[0114] In some embodiments, the disease is cancer. In some embodiments, the disease is an infectious disease. In some embodiments, the disease is heart disease. In some embodiments, the disease is associated with the levels of lipids, proteins, human metabolites, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0115] Figure 1 : Figure 1 Brefeldin A ETaG identified in Penicillium vulpinum IBT 29486 is shown. The exemplary ETaG identified is a Sec7 guanine-nucleotide-exchange-factor superfamily (pfam01369). Sequence similarity is the similarity of the Sec7 domain calculated using the MUSCLE alignment algorithm.
[0116] Figure 2 : Figure 2 Shown are lovastatin ETaGs identified in Aspergillus terreus ATCC 20542. The exemplary ETaG identified is hydroxymethylglutaryl-coenzyme A reductase (HMG-CoA; pfam00368). Sequence similarity is the similarity of the HMG-CoA domain calculated using the MUSCLE alignment algorithm.
[0117] Figure 3 : Figure 3Fellutamide ETaGs identified in Aspergillus nidulans FGSCA4 are shown. The exemplary ETaG identified is the proteasome 20S β-subunit (pfam00227). Sequence similarity is that of the 20S β-subunit calculated using the MUSCLE alignment algorithm.
[0118] Figure 4 : Figure 4 Shown are cyclosporin ETaGs identified in Tolypocladium inflatum NRRL 8044. The exemplary ETaG identified is a cyclophilin-type peptidyl-prolyl cis-trans isomerase (pfam00160). Sequence similarity is that of the cyclophilin domain calculated using the MUSCLE alignment algorithm.
[0119] Figure 5 : Figure 5 Shown are Ras ETaGs identified in Thermomyces lanuginosus ATCC 200065 (public). The exemplary ETaGs identified are from the Ras family (pfam00071). Sequence similarity is the similarity of the Ras domains calculated using the MUSCLE alignment algorithm. The ETaGs are shown below the scale bar.
[0120] Figure 6 : Figure 6 Shown are Ras ETaGs identified in Talaromyces leycettanus strain CBS 398.68. The exemplary ETaGs identified are from the Ras family (pfam00071). Sequence similarity is the similarity of the Ras domains calculated using the MUSCLE alignment algorithm. The ETaGs are shown below the scale bar.
[0121] Figure 7 : Figure 7 Ras ETaGs identified in Sistotremastrum niveocremeum HHB9708 or Sistotremastrum suecicum HHB10207 (National Forestry Service) are shown. The exemplary ETaGs identified are from the Ras family (pfam00071). Sequence similarity is the similarity of the Ras domains calculated using the MUSCLE alignment algorithm. ETaGs are shown below the scale bar.
[0122] Figure 8 : Figure 8 Shown are Ras ETaGs identified in Agaricus bisporus var. burnettii JB137-S8 (Fungal Genome Stock Center). The exemplary ETaGs identified are from the Ras family (pfam00071). Sequence similarity is the similarity of the Ras domains calculated using the MUSCLE alignment algorithm. The ETaGs are shown below the scale bar.
[0123] Figure 9 : Figure 9 Shown are Ras ETaGs identified in Coprinopsis cinerea okayama 7#130 (Fungal Genome Library). The exemplary ETaGs identified are from the Ras family (pfam00071). Sequence similarity is the similarity of the Ras domains calculated using the MUSCLE alignment algorithm. The ETaGs are shown below the scale bar.
[0124] Figure 10 : Figure 10 Ras ETaGs identified in Colletotrichum higginsianum IMI 349063 (CABI) are shown. The exemplary ETaGs identified are from the Ras family (pfam00071). Sequence similarity is the similarity of the Ras domains calculated using the MUSCLE alignment algorithm. ETaGs are shown below the scale bar.
[0125] Figure 11 : Figure 11 Shown is a RasETaG identified in Gyalolechia flavorubescens KoLRI002931. The exemplary ETaG identified is from the Ras family (pfam00071). Sequence similarity is the similarity of the Ras domain calculated using the MUSCLE alignment algorithm. ETaG is shown below the scale bar.
[0126] Figure 12 : Figure 12 Ras ETaGs identified in Bipolaris maydis ATCC 48331 are shown. The exemplary ETaGs identified are from the Ras family (pfam00071). Sequence similarity is the similarity of the Ras domains calculated using the MUSCLE alignment algorithm. ETaGs are shown below the scale bar.
[0127] Figure 13 : Figure 13 An alignment of the human Ras gene with some of the identified Ras ETaGs is shown. As shown, the human Ras gene shares identical amino acid residues with the indicated ETaGs at many positions of the KRAS nucleotide binding residue.
[0128] Figure 14 : Figure 14 An alignment of the human Ras gene with some of the identified Ras ETaGs is shown. As shown, the human Ras gene shares identical amino acid residues with the indicated ETaGs at many positions within the 4A region of the KRAS residues of BRAF.
[0129] Figure 15 : Figure 15 An alignment of the human Ras gene with some of the identified Ras ETaGs is shown. As shown, the human Ras gene is aligned with the indicated ETaGs at the rasGAP Many positions of KRAS residues within the mitochondria share the same amino acid residues.
[0130] Figure 16 : Figure 16 The alignment of the human Ras gene with some of the identified Ras ETaGs is shown. As shown, the human Ras gene is aligned with the indicated ETaGs in the SOS region. Many positions of KRAS residues within the mitochondria share the same amino acid residues.
[0131] Figure 17 : Figure 17 An exemplary sequence is shown where ETaG is indexed / labeled (dark color).
[0132] Figure 18 : Figure 18 The biosynthetic gene cluster with a Sec7 homolog in Penicillium fowleri IBT 29486 is shown.
[0133] Figure 19 : Figure 19 Sequence alignment of Sec7 is shown. (A) Exemplary Brefeldin A interacting residues. (B) Exemplary sequence alignment.
[0134] Figure 20 : Figure 20 Exemplary biosynthetic gene clusters associated with Ras, for example from Thermomyces lanuginosus ATCC 200065, Aspergillus rambelli, and Aspergillus ochraceoroseus, are shown. Ras homologs are shown in black.
[0135] Figure 21 : Figure 21 Exemplary biosynthetic gene clusters associated with Ras are shown, for example, from Agaricus bisporus variety JB137-S8, Agaricus bisporus H97, Coprinus Okayama, and Hypholoma sublateritum FD-334. Ras homologs are shown in black.
[0136] Figure 22 : Figure 22 Exemplary biosynthetic gene clusters associated with Ras are shown, for example from Sistotremastrum niveocremeum HHB9708 and Sistotremastrum suecicum HHB 10207. Ras homologs are shown in black.
[0137] Figure 23 : Figure 23 An exemplary biosynthetic gene cluster associated with Ras, for example from Talaromyces leycettanus strain CBS 398.68, is shown. Ras homologs are shown in black.
[0138] Figure 24 : Figure 24 An exemplary biosynthetic gene cluster associated with Ras, for example from Thermoascus crustaceus, is shown. Ras homologs are shown in black.
[0139] Figure 25 : Figure 25 Shown is an exemplary biosynthetic gene cluster related to Ras, for example from Helicoverpa maydis ATCC 48331. Ras homologs are shown in black.
[0140] Figure 26 : Figure 26 An exemplary biosynthetic gene cluster (CABI) associated with Ras is shown, for example from Colletotrichum higgins IMI 349063. Ras homologs are shown in black.
[0141] Figure 27 : Figure 27 An exemplary biosynthetic gene cluster associated with Ras, for example from Gyalolechia flavorubescens, is shown. Ras homologs are shown in black.
[0142] Figure 28 : Figure 28Exemplary biosynthetic gene clusters related to RasGEFs, for example from Penicillium chrysogenum Wisconsin 54-1255 and Lecanosticta acicula CBS 871.95, are shown. The RasGEF homologs shown are marked in black.
[0143] Figure 29 : Figure 29 An exemplary biosynthetic gene cluster associated with RasGEF, for example from Magnaporthe oryzae 70-15, is shown. The RasGEF homologs shown are marked in black.
[0144] Figure 30 : Figure 30 An exemplary biosynthetic gene cluster related to RasGEF, for example from Arthroderma gypseum CBS 118893, is shown. The RasGEF homologs shown are marked in black.
[0145] Figure 31 : Figure 31 An exemplary biosynthetic gene cluster related to RasGEF, for example from Endocarpon pusillum strain KoLRI No. LF000583, is shown. The RasGEF homologs shown are marked in black.
[0146] Figure 32 : Figure 32 An exemplary biosynthetic gene cluster associated with RasGEF, for example from Fistulina hepatica ATCC 64428, is shown. The RasGEF homologs shown are marked in black.
[0147] Figure 33 : Figure 33 An exemplary biosynthetic gene cluster associated with RasGEF, for example from Aureobasidium pullulans var. pullulans EXF-150, is shown. The RasGEF homologs shown are marked in black.
[0148] Figure 34 : Figure 34An exemplary biosynthetic gene cluster related to RasGAP, for example from Acremonium furcatum var. pullulan EXF-150, is shown. The RasGAP homologs shown are marked in black.
[0149] Figure 35 : Figure 35 Exemplary biosynthetic gene clusters related to RasGEFs, for example from Purpureocillium lilacinum strain TERIBC 1 and Fusarium sp. JS1030, are shown. The RasGEF homologs shown are marked in black.
[0150] Figure 36 : Figure 36 Exemplary biosynthetic gene clusters related to RasGAP are shown, for example, from Corynespora cassiicola UM591 and Oryza sativa strain SV9610. RasGAP homologs are shown in black.
[0151] Figure 37 : Figure 37 An exemplary biosynthetic gene cluster related to RasGAP, for example from Colletotrichum acutatum strain 1 KC05_01, is shown. RasGAP homologs are shown in black.
[0152] Figure 38 : Figure 38 Exemplary biosynthetic gene clusters related to RasGAP, for example from Hypoxylon sp. E7406B and Diaporthe ampelina isolate DA912, are shown. RasGAP homologs are shown in black.
[0153] Figure 39 : Figure 39 Exemplary biosynthetic gene clusters related to RasGAP are shown, for example, from Talaromyces piceae strain 9-3 and Sporothrix insectorum RCEF 264. RasGAP homologs are shown in black. DETAILED DESCRIPTION
[0154] 1. Definition
[0155] As used herein, unless otherwise indicated, the following definitions shall apply. For the purposes of this disclosure, chemical elements are according to the Periodic Table of the Elements, CAS version, Handbook of Chemistry and Physics, 75 th Ed. In addition, the general principles of organic chemistry are described in "Organic Chemistry", Thomas Sorrell, University Science Books, Sausalito: 1999, and "March's Advanced Organic Chemistry", 5 th Ed., Ed.: Smith, M. B. and March, J., John Wiley & Sons, New York: 2001.
[0156] Aliphatic: As used herein, "aliphatic" means a straight chain (i.e., unbranched) or branched, substituted or unsubstituted hydrocarbon chain that is fully saturated or contains one or more unsaturated units, or a substituted or unsubstituted monocyclic, bicyclic, or polycyclic hydrocarbon ring that is fully saturated or contains one or more unsaturated units, or a combination thereof. Unless otherwise indicated, an aliphatic group contains 1 to 100 aliphatic carbon atoms. In some embodiments, an aliphatic group contains 1 to 20 aliphatic carbon atoms. In other embodiments, an aliphatic group contains 1 to 10 aliphatic carbon atoms. In other embodiments, an aliphatic group contains 1 to 9 aliphatic carbon atoms. In other embodiments, an aliphatic group contains 1 to 8 aliphatic carbon atoms. In other embodiments, an aliphatic group contains 1 to 7 aliphatic carbon atoms. In other embodiments, an aliphatic group contains 1 to 6 aliphatic carbon atoms. In yet other embodiments, the aliphatic group comprises 1 to 5 aliphatic carbon atoms, and in yet other embodiments, the aliphatic group comprises 1, 2, 3, or 4 aliphatic carbon atoms. Suitable aliphatic groups include, but are not limited to, linear or branched, substituted or unsubstituted alkyl, alkenyl, alkynyl, and heterocomplexes thereof.
[0157] Alkyl: The term "alkyl" as used herein is given its ordinary meaning in the art and may include saturated aliphatic groups, including straight chain alkyls, branched chain alkyls, cycloalkyls (alicyclic groups), cycloalkyls substituted with alkyls, and alkyls substituted with cycloalkyls. In some embodiments, an alkyl group has 1 to 100 carbon atoms. In certain embodiments, a straight chain or branched chain alkyl group has about 1 to 20 carbon atoms in its backbone (e.g., C1-C2 for a straight chain). 20; For branched chains, C2-C 20 ), or, about 1 to 10 carbon atoms. In some embodiments, a cycloalkyl ring has about 3 to 10 carbon atoms in its ring structure, wherein such ring is monocyclic, bicyclic, or polycyclic, or has about 5, 6, or 7 carbons in the ring structure. In some embodiments, the alkyl group can be a lower alkyl group, wherein the lower alkyl group contains 1 to 4 carbon atoms (e.g., C1-C4 for a straight chain lower alkyl group).
[0158] Aryl: The term "aryl" as used alone or as part of a larger moiety as in "aralkyl," "aralkoxy," or "aryloxyalkyl" refers to a monocyclic, bicyclic, or polycyclic ring system having a total of five to thirty ring members, wherein at least one ring in the system is aromatic. In some embodiments, an aryl group is a monocyclic, bicyclic, or polycyclic ring system having a total of five to fourteen ring members, wherein at least one ring in the system is aromatic, and wherein each ring in the system contains 3 to 7 ring members. In some embodiments, an aryl group is a biaryl group. The term "aryl" can be used interchangeably with the term "aromatic ring." In certain embodiments of the present disclosure, "aryl" refers to an aromatic ring system including, but not limited to, phenyl, biphenyl, naphthyl, binaphthyl, anthracenyl, and the like, which may have one or more substituents. Also included within the scope of the term "aryl," as used herein, in some embodiments, are groups in which an aromatic ring is fused to one or more non-aromatic rings, such as indanyl, phthalimidyl, naphthimidyl, phenanthridinyl, or tetrahydronaphthyl, among others, where the radical or point of attachment is on the aromatic ring.
[0159] Cycloaliphatic: The term "cycloaliphatic" as used herein refers to a saturated or partially unsaturated aliphatic monocyclic, bicyclic or polycyclic ring system having, for example, 3 to 30 members, wherein the aliphatic ring system is optionally substituted. Cycloaliphatic groups include, but are not limited to, cyclopropyl, cyclobutyl, cyclopentyl, cyclopentenyl, cyclohexyl, cyclohexenyl, cycloheptyl, cycloheptenyl, cyclooctyl, cyclooctenyl, norbornyl, adamantyl and cyclooctadienyl. In some embodiments, the cycloalkyl group has 3 to 6 carbon atoms. The term "cycloaliphatic" may also include aliphatic rings fused to one or more aromatic or non-aromatic rings, such as decahydronaphthyl or tetrahydronaphthyl, wherein the linking group or point of attachment is on the aliphatic ring. In some embodiments, the carbocyclic group is bicyclic. In some embodiments, the carbocyclic group is tricyclic. In some embodiments, the carbocyclic group is polycyclic. In some embodiments, "cycloaliphatic" (or "carbocycle" or "cycloalkyl") refers to a monocyclic C3-C6 hydrocarbon or C8-C6 hydrocarbon that is completely saturated or contains one or more unsaturated units but is not aromatic. 10Bicyclic hydrocarbons, either completely saturated or containing one or more unsaturated units but not aromatic, C9-C 16 Tricyclic hydrocarbons.
[0160] Halogen: The term "halogen" means F, Cl, Br or I.
[0161] Heteroaliphatic: The term "heteroaliphatic" is given its ordinary meaning in the art and refers to an aliphatic group as described herein in which one or more carbon atoms are replaced by one or more heteroatoms (e.g., oxygen, nitrogen, sulfur, silicon, phosphorus, etc.).
[0162] Heteroalkyl: The term "heteroalkyl" is given its ordinary meaning in the art and refers to an alkyl group as described herein in which one or more carbon atoms are replaced by a heteroatom (e.g., oxygen, nitrogen, sulfur, silicon, phosphorus, etc.). Some examples of heteroalkyl groups include, but are not limited to, alkoxy, poly(ethylene glycol)-, alkyl-substituted amino, tetrahydrofuranyl, piperidinyl, morpholinyl, and the like.
[0163] Heteroaryl: The terms "heteroaryl" and "heteroar-" used alone or as part of a larger moiety, such as "heteroaralkyl" or "heteroaralkoxy", refer to a monocyclic, bicyclic, or polycyclic ring system having, for example, 5 to 30 total ring members, wherein at least one ring in the system is aromatic and at least one aromatic ring atom is a heteroatom. In some embodiments, the heteroatom is nitrogen, oxygen, or sulfur. In some embodiments, a heteroaryl group is a group having 5 to 10 ring atoms (i.e., monocyclic, bicyclic, or polycyclic), in some embodiments, 5, 6, 9, or 10 ring atoms. In some embodiments, a heteroaryl group has 6, 10, or 14 pi electrons in common in the ring array; and 1 to 5 heteroatoms in addition to carbon atoms. Heteroaryl groups include, but are not limited to, thienyl, furanyl, pyrrolyl, imidazolyl, pyrazolyl, triazolyl, tetrazolyl, Azolyl, iso Azolyl, In some embodiments, heteroaryl is a heterobiaryl, such as bipyridyl, etc. The terms "heteroaryl" and "heteroaryl-" as used herein also include groups in which a heteroaromatic ring is fused to one or more aromatic rings, cycloaliphatic rings, or heterocyclyl rings, wherein the linking group or point of attachment is on the heteroaromatic ring. Some non-limiting examples include indolyl, isoindolyl, benzothiophenyl, benzofuranyl, dibenzofuranyl, indazolyl, benzimidazolyl, benzothiazolyl, quinolyl, isoquinolyl, cinnolinyl, phthalazinyl, quinazolinyl, quinoxalinyl, 4H-quinolyl, carbazolyl, acridinyl, phenazinyl, phenothiazinyl, phen ... oxazinyl, tetrahydroquinolinyl, tetrahydroisoquinolinyl and pyrido[2,3-b]-1,4- Oxazin-3(4H)-one. A heteroaryl group can be monocyclic, bicyclic, or polycyclic. The term "heteroaryl" is used interchangeably with the terms "heteroaromatic ring," "heteroaryl group," or "heteroaromatic," any of which include optionally substituted rings. The term "heteroaralkyl" refers to an alkyl group substituted with a heteroaryl group, wherein the alkyl and heteroaryl portions are independently optionally substituted.
[0164] Heteroatom: The term "heteroatom" means an atom that is not carbon or hydrogen. In some embodiments, the heteroatom is oxygen, sulfur, nitrogen, phosphorus, boron, or silicon (including any oxidized form of nitrogen, sulfur, phosphorus, or silicon; any basic nitrogen or quaternized form of a substitutable nitrogen of a heterocyclic ring (e.g., N (as in 3,4-dihydro-2H-pyrrolyl), NH (as in pyrrolidinyl), or NR + (e.g., in N-substituted pyrrolidinyl)); etc.). In some embodiments, the heteroatom is boron, nitrogen, oxygen, silicon, sulfur, or phosphorus. In some embodiments, the heteroatom is nitrogen, oxygen, silicon, sulfur, or phosphorus. In some embodiments, the heteroatom is nitrogen, oxygen, sulfur, or phosphorus. In some embodiments, the heteroatom is nitrogen, oxygen, sulfur, or phosphorus. In some embodiments, the heteroatom is nitrogen, oxygen, or sulfur.
[0165] Heterocyclyl: As used herein, the terms "heterocycle," "heterocyclyl," "heterocyclic group," and "heterocyclic ring" are used interchangeably and refer to a monocyclic, bicyclic, or polycyclic moiety (e.g., 3 to 30 members) that is saturated or partially unsaturated and has one or more heteroatom ring atoms. In some embodiments, the heteroatom is boron, nitrogen, oxygen, silicon, sulfur, or phosphorus. In some embodiments, the heteroatom is nitrogen, oxygen, silicon, sulfur, or phosphorus. In some embodiments, the heteroatom is nitrogen, oxygen, sulfur, or phosphorus. In some embodiments, the heteroatom is nitrogen, oxygen, sulfur, or phosphorus. In some embodiments, the heteroatom is nitrogen, oxygen, or sulfur. In some embodiments, the heterocyclyl is a stable 5- to 7-membered monocyclic or 7- to 10-membered bicyclic heterocyclic moiety that is saturated or partially unsaturated and has one or more, preferably 1 to 4, heteroatoms as defined above in addition to carbon atoms. The term "nitrogen," when used to refer to the ring atoms of a heterocycle, includes substituted nitrogens. As an example, in a saturated or partially unsaturated ring having 0 to 3 heteroatoms selected from oxygen, sulfur or nitrogen, nitrogen can be N (as in 3,4-dihydro-2H-pyrrolyl), NH (as in pyrrolidinyl) or + NR (as in N-substituted pyrrolidinyl). The heterocycle may be attached to its side group at any heteroatom or carbon atom that would produce a stable structure, and any ring atom may be optionally substituted. Some examples of such saturated or partially unsaturated heterocyclic groups include, but are not limited to, tetrahydrofuranyl, tetrahydrothiophenyl, pyrrolidinyl, piperidinyl, pyrrolinyl, tetrahydroquinolinyl, tetrahydroisoquinolinyl, decahydroquinolinyl, Oxazolidinyl, piperazinyl, di Alkyl, dioxolanyl, diazepine Oxazolin Thiazepam The terms "heterocycle", "heterocyclyl", "heterocyclyl ring", "heterocyclic group", "heterocyclic moiety" and "heterocyclic group" are used interchangeably herein and also include groups in which the heterocyclyl ring is fused to one or more aromatic rings, heteroaromatic rings or cycloaliphatic rings, such as indolinyl, 3H-indolyl, chromanyl, phenanthridinyl or tetrahydroquinolinyl, wherein the linking group or point of attachment is on the heteroaliphatic ring. The heterocyclyl group can be monocyclic, bicyclic or polycyclic. The term "heterocyclylalkyl" refers to an alkyl group substituted with a heterocyclyl, wherein the alkyl and heterocyclyl moieties are independently optionally substituted.
[0166] Partially unsaturated: As used herein, the term "partially unsaturated" refers to a moiety containing at least one double or triple bond. The term "partially unsaturated" is intended to encompass groups having multiple sites of unsaturation, but is not intended to include aryl or heteroaryl moieties.
[0167] Pharmaceutical composition: As used herein, the term "pharmaceutical composition" refers to an active agent formulated together with one or more pharmaceutically acceptable carriers. In some embodiments, the active agent is present in a unit dosage amount suitable for administration in a therapeutic regimen that shows a statistically significant probability of achieving a predetermined therapeutic effect when administered to a relevant population. In some embodiments, the pharmaceutical composition can be specially formulated for administration in solid or liquid form, including those suitable for the following: oral administration, for example, a drench (aqueous or non-aqueous solution or suspension), a tablet, for example, those targeted for buccal, sublingual and systemic absorption, a pill, a powder, a granule, a paste applied to the tongue; parenteral administration, for example, by subcutaneous, intramuscular, intravenous or epidural injection, such as, for example, a sterile solution or suspension, or a sustained-release formulation; topical application, for example, as a cream, ointment or controlled-release patch or spray, which is applied to the skin, lungs or mouth; intravaginal or intrarectal, for example, as a pessary, cream or foam; sublingual; ophthalmic; transdermal; or nasal, pulmonary and to other mucosal surfaces.
[0168] Pharmaceutically acceptable: As used herein, the phrase "pharmaceutically acceptable" refers to those compounds, materials, compositions and / or dosage forms that are, within the scope of sound medical judgment, suitable for use in contact with the tissues of humans and animals without excessive toxicity, irritation, allergic response, or other problem or complication, commensurate with a reasonable benefit / risk ratio.
[0169] Pharmaceutically acceptable carrier: As used herein, the term "pharmaceutically acceptable carrier" means a pharmaceutically acceptable material, composition, or vehicle involved in delivering or transporting the subject compound from one organ or part of the body to another organ or part of the body, such as a liquid or solid filler, diluent, excipient, or solvent encapsulating material. Each carrier must be "acceptable" in the sense of being compatible with the other ingredients of the formulation and not injurious to the patient. Some examples of materials that can be used as pharmaceutically acceptable carriers include: sugars such as lactose, glucose, and sucrose; starches such as corn starch and potato starch; cellulose and its derivatives such as sodium carboxymethylcellulose, ethylcellulose, and cellulose acetate; powdered tragacanth; malt; gelatin; talc; excipients such as cocoa butter and suppository waxes; oils such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil; glycols such as propylene glycol; polyols such as glycerol, sorbitol, mannitol, and polyethylene glycol; esters such as ethyl oleate and ethyl laurate; agar; buffers such as magnesium hydroxide and aluminum hydroxide; alginic acid; pyrogen-free water; isotonic saline; Ringer's solution; ethanol; pH buffered solutions; polyesters, polycarbonates, and / or polyanhydrides; and other nontoxic, compatible substances used in pharmaceutical formulations.
[0170] Pharmaceutically acceptable salt: As used herein, the term "pharmaceutically acceptable salt" refers to salts of such compounds that are suitable for use in a pharmaceutical setting, i.e., salts that are, within the scope of sound medical judgment, suitable for use in contact with the tissues of humans and lower animals without excessive toxicity, irritation, allergic response, and the like, and are commensurate with a reasonable benefit / risk ratio.
[0171] Pharmaceutically acceptable salts are well known. For example, SM Berge, et al. describe pharmaceutically acceptable salts in detail in J. Pharmaceutical Sciences, 66: 1-19 (1977). In some embodiments, pharmaceutically acceptable salts include, but are not limited to, non-toxic acid addition salts, which are salts of amino groups formed with inorganic acids such as hydrochloric acid, hydrobromic acid, phosphoric acid, sulfuric acid, and perchloric acid, or with organic acids such as acetic acid, maleic acid, tartaric acid, citric acid, succinic acid, or malonic acid, or by using other known methods such as ion exchange. In some embodiments, pharmaceutically acceptable salts include, but are not limited to, adipate, alginate, ascorbate, aspartate, benzenesulfonate, benzoate, bisulfate, borate, butyrate, camphorate, camphorsulfonate, citrate, cyclopentanepropionate, digluconate, dodecylsulfate, ethanesulfonate, formate, fumarate, glucoheptonate, glycerophosphate, gluconate, hemisulfate, heptanoate, hexanoate, hydroiodide, 2-hydroxy-ethanesulfonate, lactobionate, lactate, laurate, lauryl sulfate, malate, maleate, malonate, methanesulfonate, 2-naphthalenesulfonate, nicotinate, nitrate, oleate, oxalate, palmitate, pamoate, pectinate, persulfate, 3-phenylpropionate, phosphate, picrate, pivalate, propionate, stearate, succinate, sulfate, tartrate, thiocyanate, p-toluenesulfonate, undecanoate, valerate, and the like. In some embodiments, pharmaceutically acceptable salts include, but are not limited to, non-toxic base addition salts such as those formed by acidic groups of provided compounds (e.g., phosphate groups of oligonucleotides, phosphorothioate groups of oligonucleotides, etc.) and bases. Representative alkali metal salts or alkaline earth metal salts include salts of sodium, lithium, potassium, calcium, magnesium, and the like. In some embodiments, pharmaceutically acceptable salts are ammonium salts (e.g., -N(R)3 + In some embodiments, the pharmaceutically acceptable salt is a sodium salt. In some embodiments, pharmaceutically acceptable salts include non-toxic ammonium, quaternary ammonium, and amine cations formed using counterions such as halides, hydroxides, carboxylates, sulfates, phosphates, nitrates, alkyls having 1 to 6 carbon atoms, sulfonates, and arylsulfonates, where appropriate.
[0172] Protecting group: the phrase "protecting group" used herein refers to a temporary substituent that protects a potential reactive functional group from undesirable chemical transformations. Some examples of such protecting groups include esters of carboxylic acids, silyl ethers of alcohols, and acetals and ketals of aldehydes and ketones, respectively." Si protecting group " is a protecting group comprising a Si atom, such as Si-trialkyl (e.g., trimethylsilyl, tributylsilyl, tert-butyldimethylsilyl), Si-triaryl, Si-alkyl-diphenyl (e.g., tert-butyldiphenylsilyl) or Si-aryl-dialkyl (e.g., Si-phenyldialkyl). Generally speaking, the Si protecting group is connected to an oxygen atom. The field of protecting group chemistry has been reviewed (Greene, TW; Wuts, PGM Protective Groups mOrganic Synthesis, 5th ed.; John Wiley and Sons: Hoboken, NJ, 2014). Exemplary protecting groups (and associated protected moieties) are described in detail below.
[0173] Protected hydroxyl groups are well known in the art and are included in Protecting Groups in Organic Synthesis, TW Greene and PGM Wuts, 3 rdThose described in detail in edition, John Wiley & Sons, 1999, are incorporated herein by reference in their entirety. Some examples of the hydroxyl group of suitable protection also include, but are not limited to, esters, carbonates, sulfonates, allyl ethers, ethers, silyl ethers, alkyl ethers, aryl alkyl ethers and alkoxyalkyl ethers. Some examples of suitable esters include formates, acetates, propionates, valerates, crotonates and benzoates. Some specific examples of suitable esters include formates, benzoyl formate, chloroacetate, trifluoroacetate, methoxyacetate, triphenylmethoxyacetate, p-chlorophenoxyacetate, 3-phenylpropionic acid ester, 4-oxopentanoate, 4,4-(ethylenedithio) valerate, pivalate (trimethylacetate), crotonates, 4-methoxy-crotonates, benzoates, p-benzylbenzoate, 2,4,6-trimethylbenzoate. Some examples of suitable carbonates include 9-fluorenylmethyl carbonate, ethyl carbonate, 2,2,2-trichloroethyl carbonate, 2-(trimethylsilyl) ethyl carbonate, 2-(phenylsulfonyl) ethyl carbonate, ethylene carbonate, allyl carbonate and p-nitrobenzyl carbonate. Some examples of suitable silyl ethers include trimethylsilyl ether, triethylsilyl ether, tert-butyldimethylsilyl ether, tert-butyldiphenylsilyl ether, triisopropylsilyl ether and other trialkylsilyl ethers. Some examples of suitable alkyl ethers include methyl ether, benzyl ether, p-methoxybenzyl ether, 3,4-dimethoxybenzyl ether, trityl ether, tert-butyl ether and allyl ether or derivatives thereof. Alkoxyalkyl ethers include acetals, such as methoxymethyl ether, methylthiomethyl ether, (2-methoxyethoxy) methyl ether, benzyloxymethyl ether, β-(trimethylsilyl) ethoxymethyl ether and tetrahydropyran-2-yl ether. Some examples of suitable arylalkyl ethers include benzyl ether, p-methoxybenzyl (MPM) ether, 3,4-dimethoxybenzyl ether, o-nitrobenzyl ether, p-nitrobenzyl ether, p-halobenzyl ether, 2,6-dichlorobenzyl ether, p-cyanobenzyl ether, 2-picolyl ether, and 4-picolyl ether.
[0174] Protected amines are well known in the art and include those described in detail in Greene (1999). Suitable mono-protected amines also include, but are not limited to, arylalkylamines, carbamates, allylamines, amides, and the like. Some examples of suitable mono-protected amino moieties include tert-butyloxycarbonylamino (-NHBOC), ethoxycarbonylamino, methoxycarbonylamino, trichloroethoxycarbonylamino, allyloxycarbonylamino (-NHAlloc), benzyloxycarbonylamino (-NHCBZ), allylamino, benzylamino (-NHBn), fluorenylmethylcarbonyl (-NHFmoc), formylamino, acetylamino, chloroacetylamino, dichloroacetylamino, trichloroacetylamino, phenylacetylamino, trifluoroacetylamino, benzylamino, tert-butyldiphenylsilyl, and the like. Suitable di-protected amines include amines substituted with two substituents independently selected from those described above for the mono-protected amines, and also include cyclic imides such as phthalimide, maleimide, succinimide, and the like. Suitable double-protected amines also include pyrrole, 2,2,5,5-tetramethyl-[1,2,5]azadisilolidine, and azide.
[0175] Protected aldehydes are well known in the art and include those described in detail in Greene (1999). Suitable protected aldehydes also include, but are not limited to, acyclic acetals, cyclic acetals, hydrazones, imines, and the like. Some examples of such groups include dimethyl acetal, diethyl acetal, diisopropyl acetal, dibenzyl acetal, bis(2-nitrobenzyl)acetal, 1,3-di ...isopropyl acetal, dibenzyl acetal, bis(2-nitrobenzyl)acetal, 1,3-dimethyl acetal, diisopropyl acetal, dibenzyl acetal, bis(2-nitrobenzyl)acetal, 1,3-dimethyl acetal, diisopropyl Alkanes, 1,3-dioxolanes, semicarbazones, and their derivatives.
[0176] Protected carboxylic acids are well known in the art and include those described in detail in Greene (1999). Suitable protected carboxylic acids also include, but are not limited to, optionally substituted C 1-6 Aliphatic esters, optionally substituted aryl esters, silyl esters, activated esters, amides, hydrazides, and the like. Some examples of such ester groups include methyl, ethyl, propyl, isopropyl, butyl, isobutyl, benzyl, and phenyl esters, each of which is optionally substituted. Additional suitable protected carboxylic acids include Oxazoline and orthoester.
[0177] Protected thiols are well known in the art and include those described in detail in Greene (1999). Suitable protected thiols also include, but are not limited to, disulfides, thioethers, silyl thioethers, thioesters, thiocarbonates, and thiocarbamates, among others. Some examples of such groups include, but are not limited to, alkyl thioethers, benzyl thioethers and substituted benzyl thioethers, triphenylmethyl thioether, and trichloroethoxycarbonyl thioester, among others.
[0178] Substitution: As described herein, the compounds of the present disclosure may include optionally substituted and / or substituted parts. Generally speaking, the term "substituted", whether or not preceded by the term "optionally", means that one or more hydrogens of the designated part are replaced by a suitable substituent. Unless otherwise indicated, an "optionally substituted" group may have a suitable substituent at each substitutable position of the group, and when more than one position in any given structure can be substituted by more than one substituent selected from a specified group, the substituent may be the same or different at each position. The combinations of substituents envisioned by the present disclosure are preferably those that form stable or chemically feasible compounds. The term "stable" as used herein refers to compounds that do not substantially change when subjected to conditions that allow them to be produced, detected, and in certain embodiments, recovered, purified, and used for one or more of the purposes disclosed herein. In some embodiments, some exemplary substituents are described below.
[0179] Suitable monovalent substituents are halogen; -(CH2) 0-4 R o ; -(CH2) 0-4 OR o ;-O(CH2) 0-4 R o ;-O-(CH2) 0-4 C(O)OR o ; -(CH2) 0-4 CH(OR o )2; can be R o Substituted -(CH2) 0-4 Ph; can be R o Substituted -(CH2) 1-4 O(CH2) 0-1 Ph; can be R o Substituted -CH=CHPh; can be R o Substituted -(CH2) 0-4 O(CH2) 0-1 -pyridinyl; -NO2; -CN; -N3; -(CH2) 0- 4N(R o )2;-(CH2)0-4N(R o )C(O)Ro ;-N(R o )C(S)R o ;-(CH2) 0-4 N(R o )C(O)N(R o )2;-N(R o )C(S)N(R o )2;-(CH2) 0-4 N(R o )C(O)OR o ;-N(R o )N(R o )C(O)R o ;-N(R o )N(R o )C(O)N(R o )2;-N(R o )N(R o )C(O)OR o ;-(CH2) 0-4 C(O)R o ;-C(S)R o ;-(CH2) 0-4 C(O)OR o ;-(CH2) 0-4 C(O)SR o ;-(CH2) 0-4 C(O)OSi(R o )3;-(CH2) 0-4 OC(O)R o ;-OC(O)(CH2) 0-4 SR o ;-SC(S)SR o ;-(CH2) 0-4 SC(O)R o ;-(CH2) 0-4 C(O)N(R o )2;-C(S)N(R o )2;-C(S)SR o ;-SC(S)SR o ;-(CH2) 0-4 OC(O)N(R o )2;-C(O)N(OR o )R o ;-C(O)C(O)R o ;-C(O)CH2C(O)R o ;-C(NOR o )R o ;-(CH2) 0-4 SSR o; -(CH2) 0-4 S(O)2R o ; -(CH2) 0-4 S(O)2OR o ; -(CH2) 0-4 OS(O)2R o ;-S(O)2N(R o )2;-(CH2) 0-4 S(O)R o ;-N(R o )S(O)2N(R o )2-N(R o )S(O)2R o ;-N(OR o )R o ;-C(NH)N(R o )2;-Si(R o )3;-OSi(R o )3;-P(R o )2;-P(OR o )2;-OP(R o )2;-OP(OR o )2;-N(R o )P(R o )2;-B(R o )2;-OB(R o )2;-P(O)(R o )2;-OP(O)(R o )2;-N(R o )P(O)(R o )2;-(C 1-4 linear or branched alkylene)ON(R o )2; or -(C 1-4 linear or branched alkylene) C(O)ON(R o )2; where each R o may be substituted as defined below and are independently hydrogen; C 1-20 aliphatic; C having 1 to 5 heteroatoms independently selected from nitrogen, oxygen, sulfur, silicon and phosphorus 1-20 Heteroaliphatic; -CH2-(C 6-14 aryl); -O(CH2) 0-1 (C 6-14 aryl); -CH2-(5- to 14-membered heteroaromatic ring); a 5- to 20-membered monocyclic, bicyclic or polycyclic saturated, partially unsaturated or aromatic ring having 0 to 5 heteroatoms independently selected from nitrogen, oxygen, sulfur, silicon and phosphorus, or, notwithstanding the above limitations, two independent occurrences of R oand its intervening atoms are taken together to form a 5- to 20-membered monocyclic, bicyclic or polycyclic saturated, partially unsaturated or aryl ring having 0 to 5 heteroatoms independently selected from nitrogen, oxygen, sulfur, silicon and phosphorus, which may be substituted as defined below.
[0180] In R o (or through two independent occurrences of R o Suitable monovalent substituents on the ring formed by taking together with its intervening atoms are independently halogen; -(CH2) 0-2 R · ;-(halogenated R · ); -(CH2) 0-2 OH; -(CH2) 0-2 OR · ; -(CH2) 0-2 CH(OR · )2;-O(halogenated R · ); -CN; -N3; -(CH2) 0-2 C(O)R · ; -(CH2) 0-2 C(O)OH; -(CH2) 0-2 C(O)OR · ; -(CH2) 0-2 SR · ; -(CH2) 0-2 SH; -(CH2) 0-2 NH2; -(CH2) 0-2 NHR · ; -(CH2) 0-2 NR · 2;-NO2;-SiR · 3;-OSiR · 3;-C(O)SR · ;-(C 1-4 linear or branched alkylene)C(O)OR · ; or -SSR · , where each R · is unsubstituted or, in the case of "halo" in front, is substituted only with one or more halogens, and is independently selected from C 1-4 Aliphatic; -CH2Ph; -O(CH2) 0-1 Ph; or a 5- to 6-membered saturated, partially unsaturated or aromatic ring having 0 to 4 heteroatoms independently selected from nitrogen, oxygen and sulfur. o Suitable divalent substituents on a saturated carbon atom of include =0 and =S.
[0181] Suitable divalent substituents are the following: =O; =S; =NNR * 2; =NNHC(O)R* ;=NNHC(O)OR * ; =NNHS(O)2R * ; =NR * ; =NOR * ;-O(C(R * 2)) 2-3 O-; or -S(C(R * 2)) 2-3 S-, where each independent occurrence of R * is selected from hydrogen; C which may be substituted as defined below 1-6 aliphatic; or an unsubstituted 5- to 6-membered saturated, partially unsaturated, or aromatic ring having 0 to 4 heteroatoms independently selected from nitrogen, oxygen, and sulfur. Suitable divalent substituents bound to the ortho-substitutable carbon of an "optionally substituted" group include: -O(CR * 2) 2-3 O-, where each independent occurrence of R * is selected from hydrogen; C which may be substituted as defined below 1-6 aliphatic; or an unsubstituted 5- to 6-membered saturated, partially unsaturated, or aryl ring having 0 to 4 heteroatoms independently selected from nitrogen, oxygen, and sulfur.
[0182] In R * Suitable substituents on the aliphatic group of are halogen; -R · ;-(halogenated R · );-OH;-OR · ;-O(halogenated R · );-CN;-C(O)OH;-C(O)OR · ;-NH2;-NHR · ;-NR · 2; or -NO2, where each R · is unsubstituted or, in the case of being preceded by "halo", substituted only by one or more halogens, and is independently C 1-4 Aliphatic; -CH2Ph; -O(CH2) 0-1 Ph; or a 5- to 6-membered saturated, partially unsaturated or aryl ring having 0 to 4 heteroatoms independently selected from nitrogen, oxygen and sulfur.
[0183] In some embodiments, suitable substituents on the substitutable nitrogen are or Each of these are independently hydrogen; C 1-6aliphatic; unsubstituted -OPh; or an unsubstituted 5- to 6-membered saturated, partially unsaturated, or aryl ring having 0 to 4 heteroatoms independently selected from nitrogen, oxygen, and sulfur, or, notwithstanding the above limitations, two independent occurrences of Together with their intervening atoms they form an unsubstituted 3- to 12-membered saturated, partially unsaturated or aromatic monocyclic or bicyclic ring having 0 to 4 heteroatoms independently selected from nitrogen, oxygen and sulfur.
[0184] exist Suitable substituents on the aliphatic group of are independently halogen; -R · ;-(halogenated R · );-OH;-OR · ;-O(halogenated R · );-CN;-C(O)OH;-C(O)OR · ;-NH2;-NHR · ;-NR · 2; or -NO2, where each R · is unsubstituted or, in the case of being preceded by "halo", substituted only with one or more halogens, and is independently C 1-4 Aliphatic; -CH2Ph; -O(CH2) 0-1 Ph; or a 5- to 6-membered saturated, partially unsaturated or aryl ring having 0 to 4 heteroatoms independently selected from nitrogen, oxygen and sulfur.
[0185] Unsaturated: The term "unsaturated" as used herein means that the moiety has one or more units of unsaturation.
[0186] Unless otherwise indicated, salts of provided compounds, such as pharmaceutically acceptable acid or base addition salts, stereoisomeric forms, and tautomeric forms are included.
[0187] 2. Detailed Description of Certain Embodiments
[0188] The present disclosure encompasses, among other things, the recognition that many products produced by enzymes encoded by fungal biosynthetic gene clusters can be used to develop therapeutics directed against human targets to treat a variety of diseases. The present disclosure recognizes that one challenge in using fungal products is identifying their human targets. In some embodiments, the present disclosure provides techniques for efficiently identifying human targets for biosynthetic products produced by enzymes encoded by fungal biosynthetic gene clusters. In some embodiments, the provided techniques identify embedded target genes (ETaGs) in the vicinity of biosynthetic genes within a biosynthetic gene cluster, and optionally also identify human targets for biosynthetic products produced by enzymes encoded by the biosynthetic gene cluster by comparing the ETaG sequences with human nucleic acid sequences, particularly expressed human nucleic acid sequences (including protein-encoding human genes). As will be readily understood by those skilled in the art, once the association between a biosynthetic product from a biosynthetic gene cluster, the ETaG, and the human target is established, it can be used in a variety of methods. For example, one can start with a biosynthetic product produced by an enzyme encoded by a biosynthetic gene cluster, proceed to an ETaG within the vicinity of a biosynthetic gene within the biosynthetic gene cluster, and then proceed to a human target homologous to the ETaG. Once a human target is identified, it can be prioritized (even if it was previously considered undruggable) and the biosynthetic products used to develop modulators of the human target using the many methods available to those skilled in the art, including optionally further optimizing the biosynthetic products for medical use, such as by preparing and testing analogs of the products. One can also start from a human target of therapeutic interest, to an ETaG homologous to the human target, and then to a biosynthetic gene cluster containing the biosynthetic gene, with the ETaG being included relative to the neighborhood of the biosynthetic gene. Once a biosynthetic gene cluster is identified, the biosynthetic products produced by the enzymes encoded by the biosynthetic gene cluster can be characterized and tested to modulate the human target or its products. In accordance with the present disclosure, the biosynthetic products can be used as leads to be optimized using the many methods in the art to provide agents that can be used for many medical purposes, such as therapeutic purposes.
[0189] Without intending to be bound by any theory, in some embodiments, the present disclosure encompasses the recognition that ETaG from eukaryotes and / or products encoded thereby may have a higher similarity to mammalian genes and / or products encoded thereby than, for example, their counterparts in prokaryotes such as bacteria (if any); in some embodiments, eukaryotic ETaG may be more therapeutically relevant. In some embodiments, ETaG in fungi may be particularly useful for developing human therapeutics given the close relationship of fungi to mammals in the phylogenetic tree.
[0190] In some embodiments, the present disclosure provides techniques for identifying and / or characterizing ETaGs, which are non-biosynthetic genes in that they are not necessarily involved in the synthesis of products produced by enzymes encoded by the biosynthetic gene cluster that contains the ETaG or, in some embodiments, the vicinity of the biosynthetic genes relative to the genes of the biosynthetic gene cluster that contain the ETaG (the enzymes encoded by the biosynthetic gene cluster can produce the biosynthetic products in the absence of the ETaG). In some embodiments, the ETaG is not required for the synthesis of products produced by enzymes encoded by the biosynthetic gene cluster that contains the ETaG or, in some embodiments, the vicinity of the biosynthetic genes relative to the genes of the biosynthetic gene cluster that contain the ETaG (the enzymes encoded by the biosynthetic gene cluster can produce the biosynthetic products in the absence of the ETaG). In some embodiments, the ETaG is not involved in the synthesis of a product produced by an enzyme encoded by a biosynthetic gene cluster that contains the ETaG, or relative to a gene of the biosynthetic gene cluster, in some embodiments, a region adjacent to a biosynthetic gene that contains the ETaG (the enzyme encoded by the biosynthetic gene cluster can produce the biosynthetic product in the absence of the ETaG). In some embodiments, the ETaG is homologous to a human gene or comprises a sequence homologous to a human gene, e.g., shares at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% homology with a human protein or sequence (e.g., a functional unit and / or structural unit, e.g., a domain, a functional structural feature (helix, sheet, etc.), etc.).
[0191] In some embodiments, ETaG is co-regulated with at least one biosynthetic gene in a biosynthetic gene cluster. In some embodiments, ETaG is co-regulated with a biosynthetic gene cluster in that its expression is correlated with the production of a product encoded by an enzyme in the biosynthetic gene cluster. In some embodiments, ETaG provides an autoprotective function. In some embodiments, ETaG encodes a transporter for a product produced by an enzyme in the biosynthetic gene cluster. In some embodiments, ETaG encodes a product, such as a protein, that detoxifies a product produced by an enzyme in the biosynthetic gene cluster. In some embodiments, ETaG encodes a resistance variant of a protein whose activity is targeted by a product produced by an enzyme in the biosynthetic gene cluster.
[0192] In some embodiments, the present disclosure provides a method comprising:
[0193] a set of query nucleic acid sequences, each of which is present in a fungal strain and comprises a biosynthetic gene cluster; and
[0194] An embedded target gene (ETaG) sequence is identified within at least one fungal nucleic acid sequence, wherein the embedded target gene (ETaG) sequence is characterized in that:
[0195] is not involved in the synthesis of products produced by enzymes encoded by biosynthetic gene clusters;
[0196] within a vicinity relative to at least one biosynthetic gene in the biosynthetic gene cluster; and
[0197] Optionally co-regulated with at least one biosynthetic gene in a biosynthetic gene cluster.
[0198] In some embodiments, ETaG is homologous to a mammalian nucleic acid sequence. In some embodiments, the present disclosure provides a method comprising:
[0199] a set of query nucleic acid sequences, each of which is present in a fungal strain and comprises a biosynthetic gene cluster; and
[0200] An embedded target gene (ETaG) sequence is identified within at least one fungal nucleic acid sequence, wherein the embedded target gene (ETaG) sequence is characterized in that:
[0201] is not involved in the synthesis of products produced by enzymes encoded by biosynthetic gene clusters;
[0202] within a vicinity relative to at least one biosynthetic gene in a biosynthetic gene cluster;
[0203] is homologous to an expressed mammalian nucleic acid sequence; and
[0204] Optionally co-regulated with at least one biosynthetic gene in a biosynthetic gene cluster.
[0205] Neighborhood
[0206] In some embodiments, ETaG is generally within the vicinity of at least one gene in a biosynthetic gene cluster. In some embodiments, ETaG is within the vicinity of at least one biosynthetic gene in a biosynthetic gene cluster. In some embodiments, the vicinity is no more than 1 to 100 kb upstream or downstream of a gene. In some embodiments, the vicinity is no more than 1 to 50 kb upstream or downstream of a gene. In some embodiments, the vicinity is no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, or 90 kb upstream or downstream of a gene. In some embodiments, the vicinity is no more than 1 kb upstream or downstream of a gene. In some embodiments, the vicinity is no more than 5 kb upstream or downstream of a gene. In some embodiments, the vicinity is no more than 10 kb upstream or downstream of a gene. In some embodiments, the vicinity is no more than 15 kb upstream or downstream of a gene. In some embodiments, the adjacent region is no more than 20 kb upstream or downstream of the gene. In some embodiments, the adjacent region is no more than 25 kb upstream or downstream of the gene. In some embodiments, the adjacent region is no more than 30 kb upstream or downstream of the gene. In some embodiments, the adjacent region is no more than 35 kb upstream or downstream of the gene. In some embodiments, the adjacent region is no more than 40 kb upstream or downstream of the gene. In some embodiments, the adjacent region is no more than 45 kb upstream or downstream of the gene. In some embodiments, the adjacent region is no more than 50 kb upstream or downstream of the gene.
[0207] In some embodiments, the ETaG is within a biosynthetic gene cluster. In some embodiments, the ETaG is not within the region defined by the first and last genes in the biosynthetic gene cluster, but is instead within a region adjacent to the first or last gene in the biosynthetic gene cluster.
[0208] Homology
[0209] In some embodiments, ETaG is homologous to an expressed mammalian nucleic acid sequence. In some embodiments, the mammalian nucleic acid sequence is an expressed mammalian nucleic acid sequence. In some embodiments, the mammalian nucleic acid sequence is a mammalian gene. In some embodiments, the mammalian nucleic acid sequence is an expressed mammalian gene. In some embodiments, the mammalian nucleic acid is a human nucleic acid sequence. In some embodiments, the human nucleic acid sequence is an expressed human nucleic acid sequence. In some embodiments, the human nucleic acid sequence is a human gene. In some embodiments, the human nucleic acid sequence is an expressed human gene. In some embodiments, the human nucleic acid sequence is or its encoded product is an existing target of therapeutic significance. In some embodiments, the human nucleic acid sequence is or its encoded product is a new target of therapeutic significance. In some embodiments, the human nucleic acid sequence is or its encoded product is a target that was previously considered undruggable. In some embodiments, the human nucleic acid sequence is or its encoded product is a target that was previously considered undruggable by small molecules. In some embodiments, the present disclosure provides the unexpected discovery that targets traditionally considered undruggable can be effectively modulated or targeted by small molecules, which are biosynthetic products or analogs of biosynthetic products produced by enzymes encoded by biosynthetic gene clusters, which biosynthetic gene clusters comprise biosynthetic genes that contain ETaG (or portions thereof, or products encoded thereby and / or portions thereof) that are homologous to the target relative to the adjacent regions of the biosynthetic genes.
[0210] In some embodiments, the present disclosure provides a method comprising:
[0211] a set of query nucleic acid sequences, each of which is present in a fungal strain and comprises a biosynthetic gene cluster; and
[0212] An embedded target gene (ETaG) sequence is identified within at least one fungal nucleic acid sequence, wherein the embedded target gene (ETaG) sequence is characterized in that:
[0213] is not involved in the synthesis of products produced by enzymes encoded by biosynthetic gene clusters;
[0214] within a vicinity relative to at least one biosynthetic gene in a biosynthetic gene cluster;
[0215] homologous to an expressed human nucleic acid sequence; and
[0216] Optionally co-regulated with at least one biosynthetic gene in a biosynthetic gene cluster.
[0217] In some embodiments, ETaG shares nucleic acid sequence homology with a nucleic acid sequence. In some embodiments, the ETaG sequence is homologous to another nucleic acid sequence (e.g., an expressed human nucleic acid sequence) in that the ETaG nucleic acid sequence, or a portion thereof, shares similarity at the nucleic acid base sequence level with the other nucleic acid sequence, or a portion thereof. In some embodiments, the sequence of ETaG shares nucleic acid base sequence similarity with the other nucleic acid sequence. In some embodiments, a portion of the sequence of ETaG shares nucleic acid base sequence similarity with a portion of the other nucleic acid sequence.
[0218] In some embodiments, the length of the homologous portion is at least 50, 100, 150, 200, 300, 400, 500, 600, 70, 800, 900, or 1000 base pairs. In some embodiments, the length is at least 50 base pairs. In some embodiments, the length is at least 100 base pairs. In some embodiments, the length is at least 150 base pairs. In some embodiments, the length is at least 200 base pairs. In some embodiments, the length is at least 300 base pairs. In some embodiments, the length is at least 400 base pairs. In some embodiments, the length is at least 500 base pairs.
[0219] In some embodiments, the homologous portion encodes amino acid residues that are certain structural and / or functional units of the encoded protein. For example, in some embodiments, the homologous portion can encode a protein domain that is characteristic of the family of the encoded protein, has enzymatic activity, is responsible for interaction with effectors, etc., as described in the present disclosure.
[0220] Methods for assessing similarity / homology of nucleic acid sequences are well known in the art and can be used in accordance with the present disclosure.
[0221] In some embodiments, ETaG shares homology with a nucleic acid sequence in the product it encodes, such as a protein. In some embodiments, ETaG is homologous to a nucleic acid sequence in that the product encoded by ETaG, or a portion thereof, shares similarity with the product encoded by the nucleic acid sequence, or a portion thereof. In some embodiments, the encoded product is a protein. In some embodiments, the product encoded by ETaG and the nucleic acid sequence shares similarity over their entire length. In some embodiments, the product encoded by ETaG and the nucleic acid sequence shares similarity over some portion.
[0222] In some embodiments, ETaG is homologous to a nucleic acid in that a protein encoded by the ETaG or a portion thereof shares similarity with a protein encoded by the nucleic acid or a portion thereof. The proteins encoded by the ETaG and nucleic acid sequence may share similarity over their entire length or over a portion thereof. In some embodiments, all amino acid residues in the homologous portion are contiguous. In some embodiments, not all amino acid residues in the homologous portion are contiguous.
[0223] In some embodiments, a portion of a protein is a protein domain. In some embodiments, a protein domain forms a structure characteristic of a protein family. In some embodiments, a protein domain performs a characteristic function. For example, in some embodiments, a protein domain has an enzymatic function. In some embodiments, such a function is shared by the protein encoded by ETaG and a protein encoded by a homologous nucleic acid sequence, such as a human gene. In some embodiments, the characteristic function is non-enzymatic. In some embodiments, the characteristic function is an interaction with another entity, such as a small molecule, a nucleic acid, a protein, etc.
[0224] In some embodiments, a portion of a protein is a collection of contiguous or discontinuous amino acid residues that are important for protein function. In some embodiments, the function is enzymatic activity. In some embodiments, a portion of a protein is a collection of residues required for activity. In some embodiments, a portion is a collection of residues that interact with a substrate, intermediate, product, or cofactor. In some embodiments, a portion is a collection of residues that interact with a substrate. In some embodiments, a portion is a collection of residues that interact with an intermediate. In some embodiments, a portion is a collection of residues that interact with a product. In some embodiments, a portion is a collection of residues that interact with a cofactor.
[0225] In some embodiments, the function is an interaction with another entity. In some embodiments, the entity is a small molecule. In some embodiments, the entity is a lipid. In some embodiments, the entity is a carbohydrate. In some embodiments, the entity is a nucleic acid. In some embodiments, the entity is a protein. In some embodiments, a portion is a collection of amino acid residues that come into contact with an interacting agent. For example, Figure 13 The portion (amino acid set) that interacts with the nucleotides of the Ras protein and its homolog ETaG is shown, and Figures 14 to 16 Portions involved in protein-protein interactions are shown.
[0226] In some embodiments, the interaction of an amino acid residue with an interacting entity can be assessed by hydrogen bonding, electrostatic forces, van der Waals forces, aromatic stacking, etc. In some embodiments, the interaction can be assessed by the distance of the amino acid residue from the interacting entity (e.g., as used in some cases). ) to evaluate.
[0227] In some embodiments, the similarity is that the two structures have a Ca backbone rmsd (root mean square deviation) within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, or 50 square angstroms and have the same overall fold or core domain. In some embodiments, the Ca backbone rmsd is within.
[0228] In some embodiments, a portion of a protein is or comprises a structural element that is essential for protein effector recruitment. In some embodiments, such a portion can be selected based on structural and / or activity data of a protein encoded by a nucleic acid sequence homologous to ETaG (e.g., a human gene encoding a protein homologous to ETaG).
[0229] In some embodiments, a portion of a protein comprises at least 2 to 200, 2 to 100, 2 to 50, 2 to 40, 2 to 30, 2 to 20, 2 to 15, 2 to 10, 3 to 200, 3 to 100, 3 to 50, 3 to 40, 3 to 30, 3 to 20, 3 to 15, 3 to 10, 4 to 200, 4 to 100, 4 to 50, 4 to 40, 4 to 30, 4 to 20, 4 to 15, 4 to 10, 5 to 200, 5 to 100, 5 to 50, 5 to 40, 5 to 30, 5 to 20, 5 to 15, or 5 to 10 amino acid residues. In some embodiments, a portion of a protein comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, or 150 amino acid residues. In some embodiments, a portion comprises at least 2 amino acid residues. In some embodiments, a portion comprises at least 3 amino acid residues. In some embodiments, a portion comprises at least 4 amino acid residues. In some embodiments, a portion comprises at least 5 amino acid residues. In some embodiments, a portion comprises at least 6 amino acid residues. In some embodiments, a portion comprises at least 7 amino acid residues. In some embodiments, a portion comprises at least 8 amino acid residues. In some embodiments, a portion comprises at least 9 amino acid residues. In some embodiments, a portion comprises at least 10 amino acid residues. In some embodiments, a portion comprises at least 15 amino acid residues. In some embodiments, a portion comprises at least 20 amino acid residues. In some embodiments, a portion comprises at least 25 amino acid residues. In some embodiments, a portion comprises at least 30 amino acid residues.
[0230] According to the present disclosure, the similarity of nucleic acid sequences and protein sequences can be assessed by a variety of methods, including those known in the art. For example, MUSCLE is used for protein sequences. In some embodiments, similarity is measured based on exact identity, such as the same amino acid residues at a given position. In some embodiments, similarity is measured based on one or more common properties, such as amino acid residues with one or more identical or similar properties (e.g., acidic, basic, aromatic, etc.).
[0231] In some embodiments, ETaG is homologous to a nucleic acid sequence (e.g., an expressed human nucleic acid sequence) in that the similarity between ETaG and the nucleic acid base sequence is no less than the level of similarity based on the nucleic acid sequence or a portion thereof, or a protein encoded by ETaG and the nucleic acid sequence or a portion thereof, as described herein. In some embodiments, ETaG is homologous to a nucleic acid sequence in that the similarity between ETaG and the nucleic acid sequence is no less than the level of similarity based on the nucleic acid sequence or a portion thereof. In some embodiments, ETaG is homologous to a nucleic acid sequence in that the similarity between ETaG and the nucleic acid sequence is no less than the level of similarity based on the nucleic acid sequence or a portion thereof. In some embodiments, the level is at least 10% to 99%. In some embodiments, the level is at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%. In some embodiments, the level is at least 10%. In some embodiments, the level is at least 20%. In some embodiments, the level is at least 30%. In some embodiments, the level is at least 40%. In some embodiments, the level is at least 50%. In some embodiments, the level is at least 60%. In some embodiments, the level is at least 70%. In some embodiments, the level is at least 80%. In some embodiments, the level is at least 90%. In some embodiments, the level is 100%. In some embodiments, the level is less than 100%. In some embodiments, the level does not exceed 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%.
[0232] In some embodiments, ETaG is homologous to a nucleic acid sequence in that the protein encoded by ETaG or a portion thereof has a 3-dimensional structure similar to the 3-dimensional structure of the protein encoded by the nucleic acid sequence. In some embodiments, similarity is assessed by, for example, a Cα backbone rmsd (root mean square deviation) within 1 to 100 square angstroms, for example, 5, 10, 20, 30, 40, 50 square angstroms. In some embodiments, sequences sharing similarity have a Cα backbone rmsd of no more than 10 square angstroms and also have the same overall fold or core domain. In some embodiments, structural similarity is assessed by interaction with another entity, such as a small molecule, nucleic acid, protein, etc. In some embodiments, structural similarity is assessed by small molecule binding. In some embodiments, a protein encoded by an embedded target gene or a portion thereof has a similar 3-dimensional structure to a protein encoded by a nucleic acid sequence in that a small molecule that binds to a protein encoded by an embedded target gene or a portion thereof also binds to a protein encoded by the nucleic acid sequence or a portion thereof. In some embodiments, the Kd for binding is no more than 1 to 100 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100) μM.
[0233] Co-regulation
[0234] In some embodiments, ETaG is co-regulated with at least one biosynthetic gene in a biosynthetic gene cluster, the biosynthetic gene cluster comprising the biosynthetic gene and the region adjacent to the biosynthetic gene comprising the ETaG. In some embodiments, ETaG is co-regulated with a biosynthetic gene cluster, the biosynthetic gene cluster comprising the biosynthetic gene and the region adjacent to the biosynthetic gene comprising the ETaG. In some embodiments, ETaG is co-regulated with a biosynthetic gene cluster in that expression of the ETaG and / or production of a product encoded by the ETaG, such as a protein, is correlated with production of a biosynthetic product produced by an enzyme encoded by the biosynthetic gene cluster. In some embodiments, production of a product encoded by the ETaG, such as a protein, overlaps temporally with production of a biosynthetic product by an enzyme encoded by the biosynthetic gene cluster. In some embodiments, ETaG is co-regulated with a biosynthetic gene cluster in that expression of the ETaG is increased or turned on when a biosynthetic product produced by an enzyme encoded by the biosynthetic gene cluster is produced. In some embodiments, ETaG is co-regulated with a biosynthetic gene cluster in that expression of ETaG is increased or turned on when the level of production of a biosynthetic product produced by an enzyme encoded by the biosynthetic gene cluster is increased.
[0235] In some embodiments, ETaG provides advantages to its host organism, such as a fungus, when producing a biosynthetic product produced by an enzyme encoded by a co-regulated biosynthetic gene cluster. For example, in some embodiments, the protein encoded by ETaG facilitates transport of the biosynthetic product out of the cell in which the product is produced. In some embodiments, the protein encoded by ETaG detoxifies the biosynthetic product so that the biosynthetic product does not harm the organism that produced the biosynthetic product, but rather affects the growth or survival of other organisms.
[0236] In some embodiments, the present disclosure provides various methods for identifying ETaGs. For example, in some embodiments, homologous biosynthetic gene clusters, typically from different fungal strains, are compared, e.g., a collection of biosynthetic gene clusters whose encoded enzymes produce the same biosynthetic product (based on prediction (e.g., sequence-based prediction) and / or identification of the product). Non-biosynthetic genes that are present in only one or a few biosynthetic gene clusters (within the biosynthetic gene cluster or in a region adjacent to the biosynthetic genes in the biosynthetic gene cluster) but not in the majority of the biosynthetic gene clusters in the collection are identified as ETaG candidates, and are optionally further compared to mammalian (e.g., human) nucleic acid sequences to identify homologous mammalian nucleic acid sequences. In some embodiments, such methods can be used to identify ETaGs on a genomic scale, e.g., from sequences of many (e.g., hundreds, thousands, or even more) genomes, as shown in the Examples. The identified ETaGs can be prioritized based on the therapeutic importance of their mammalian homologs, particularly human homologs. In some embodiments, as shown in the accompanying figures, organisms containing ETaGs contain one or more homologous genes of ETaG.
[0237] In some embodiments, ETaG is present in no more than 1%, 5%, or 10% of the biosynthetic gene clusters in a collection. In some embodiments, ETaG is present in no more than 1%, 5%, or 10% of the homologous biosynthetic gene clusters in a collection. In some embodiments, ETaG is present in no more than 1%, 5%, or 10% of the biosynthetic gene clusters in a collection that encode enzymes that produce the same biosynthetic product. In some embodiments, this percentage is less than 1%. In some embodiments, this percentage is less than 5%. In some embodiments, this percentage is less than 10%.
[0238] In some embodiments, the present disclosure provides methods for particularly effective and efficient identification of homologous ETaGs encoding human nucleic acids of therapeutically interesting targets by querying a provided collection of nucleic acid sequences comprising a biosynthetic gene cluster and / or an ETaG within a proximal region relative to a biosynthetic gene in the biosynthetic gene cluster.
[0239] In some embodiments, the present disclosure provides a set of nucleic acid sequences as described herein. In some embodiments, the present disclosure provides a set of nucleic acid sequences, each of which is present in a fungal strain and comprises a biosynthetic gene cluster. In some embodiments, the present disclosure provides a set of nucleic acid sequences, each of which is present in a fungal strain and comprises an ETaG. In some embodiments, the present disclosure provides a set of nucleic acid sequences, each of which is present in a fungal strain and comprises a biosynthetic gene cluster and an ETaG within a region adjacent to a biosynthetic gene in the biosynthetic gene cluster. In some embodiments, the nucleic acid sequences comprising a biosynthetic gene cluster do not comprise sequences other than the regions adjacent to the biosynthetic genes in the biosynthetic gene cluster and the sequence of the biosynthetic gene cluster. In some embodiments, the present disclosure provides a database comprising the provided set of nucleic acid sequences.
[0240] In some embodiments, the biosynthetic gene clusters of the provided technology comprise biosynthetic genes that encode enzymes that may be involved in the synthesis of compounds that share at least one common chemical attribute. In some embodiments, the common chemical attribute is a cyclic core structure. In some embodiments, the common chemical attribute is a macrocyclic core structure. In some embodiments, the common chemical attribute is a shared acyclic backbone. In some embodiments, the common chemical attribute is that the compounds all belong to a certain category, such as non-ribosomal peptides (NPRS), terpenes, isoprenes, alkaloids, etc. In some embodiments, by identifying individual ETaGs of a biosynthetic gene cluster, the present disclosure can distinguish compounds that share a common chemical attribute, even if they may be structurally similar.
[0241] The provided collections can be of different sizes and / or diversity. In some embodiments, it is desirable to have more sequences from more species to increase the number of ETaGs and biosynthetic gene clusters. In some embodiments, the collection comprises at least 100, 200, 300, 400, 500, 1,000, 1,500, 2,000, 3,000, 5,000, 10,000, 20,000, 50,000, 100,000, 500,000, 1,000,000, 1,500,000, or 2,000,000 nucleic acid sequences comprising biosynthetic gene clusters. In some embodiments, the collection comprises at least 100, 200, 300, 400, 500, 1,000, 1,500, 2,000, 3,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 1,500,000, or 2,000,000 biosynthetic gene clusters. In some embodiments, the collection comprises at least 100, 200, 300, 400, 500, 1,000, 1,500, 2,000, 3,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 1,500,000, or 2,000,000 biosynthetic gene clusters associated with ETaG (biosynthetic gene clusters comprising biosynthetic genes relative to the proximity of the biosynthetic genes that comprise ETaG). In some embodiments, the collection comprises at least 100, 200, 300, 400, 500, 1,000, 1,500, 2,000, 3,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 1,500,000, or 2,000,000 ETaGs. In some embodiments, the sequences in the provided collections are from at least 100, 200, 300, 400, 500, 1,000, 1,500, 2,000, 3,000, 5,000, 10,000, 20,000, 50,000, 100,000 genomes from different species, e.g., different fungal species.
[0242] In particular, the provided databases and / or provided collections are constructed to particularly improve the efficiency of, for example, identifying ETaGs, identifying ETaGs associated with a given biosynthetic gene cluster, identifying a biosynthetic gene cluster associated with a given ETAG, identifying ETaGs homologous to a given mammalian nucleic acid sequence (e.g., a human gene), identifying a biosynthetic gene cluster associated with a given mammalian nucleic acid sequence (e.g., a human gene; optionally via an associated ETaG), identifying mammalian nucleic acid sequences (e.g., a human gene) homologous to a given ETaG, mammalian nucleic acid sequences (e.g., a human gene) homologous to a given biosynthetic gene cluster (optionally via an associated ETaG), identifying mammalian nucleic acid sequences (e.g., a human gene) associated with a given product (and / or its analogue) produced by an enzyme encoded by a biosynthetic gene cluster (optionally via an associated ETaG and a biosynthetic gene cluster), identifying a product (and / or its analogue) produced by an enzyme encoded by a biosynthetic gene cluster associated with a given mammalian nucleic acid sequence (e.g., a human gene; optionally via an associated biosynthetic gene cluster and an ETaG), and the like.
[0243] For example, in some embodiments, the ETaGs in the provided collections and / or databases are indexed / tagged for retrieval. For example, Figure 17 (Applicants note that the provided collections and databases may contain hundreds, thousands, or millions of sequences.) Exemplary sequences from the provided collections and / or databases are shown, in which ETaGs are specifically indexed / labeled (dark). In particular, such structural features can greatly improve query efficiency, for example: rather than searching tens, hundreds, or thousands of genomes for ETaGs homologous to a human gene of interest, the provided technology can instead be used to focus the search on indexed / labeled ETaGs (e.g., skipping non-biosynthetic gene cluster sequences and / or non-ETaG sequences (e.g., Figure 17 and the empty arrows in between) to quickly locate hits (e.g., Figure 17 Thus, the time and resources for searching the vast majority of irrelevant genomic information are saved.
[0244] Additionally or alternatively, the provided sequence collections and databases are structured such that an ETaG can be independently annotated with information such as its associated biosynthetic gene cluster (the associated biosynthetic gene cluster of an ETaG is a biosynthetic gene cluster comprising a biosynthetic gene, the ETaG being in a proximal region relative to the biosynthetic gene), products produced by enzymes encoded by the associated biosynthetic gene cluster and their analogs, their homologous mammalian nucleic acid sequences (e.g., human genes), etc. Similarly, a biosynthetic gene cluster can be independently annotated with information such as its associated ETaG (the associated ETaG of a biosynthetic gene cluster is an etg in a proximal region relative to the biosynthetic genes in the biosynthetic gene cluster), biosynthetic products produced by enzymes encoded by the biosynthetic gene cluster and their analogs, their homologous mammalian nucleic acid sequences and the products encoded thereby, etc. By structuring sequence data using indices and annotations, the provided collections and databases can provide numerous advantages. For example, in some embodiments, the provided system allows for rapid access to ETaGs, and vice versa, with available related information such as their related biosynthetic gene clusters and human homologs, while keeping data size and query costs low.
[0245] In some embodiments, at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 2,500, 5,000, or 10,000, or at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%, or all of the ETaGs in the collection are independently annotated. In some embodiments, at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 2,500, 5,000, or 10,000, or at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%, or all of the ETaGs in the collection are independently annotated with their cognate biosynthetic gene clusters and homologous mammalian nucleic acid sequences. In some embodiments, at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 2,500, 5,000, or 10,000, or at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%, or all of the biosynthetic gene clusters in the collection are independently annotated. In some embodiments, at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 2,500, 5,000, or 10,000, or at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%, or all of the biosynthetic gene clusters in the collection are independently annotated with their associated ETaGs.
[0246] In some embodiments, the provided sequence collections and / or databases are contained in a computer-readable medium. In some embodiments, the present disclosure provides systems comprising one or more non-transitory machine-readable storage media that store data representing the provided sequence collections and / or databases. Non-transitory machine-readable storage media suitable for containing the provided data include all forms of non-volatile storage, including, for example, semiconductor storage devices, such as EPROM, EEPROM, and flash storage devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. In particular, the provided systems can be particularly effective because the provided collections and databases have the specific structures described herein.
[0247] In some embodiments, the present disclosure provides computer systems that can perform the provided techniques. In some embodiments, the present disclosure provides computer systems suitable for performing the provided methods. In some embodiments, the present disclosure provides computer systems suitable for querying provided sequence sets. In some embodiments, the present disclosure provides computer systems suitable for querying provided databases. In some embodiments, the present disclosure provides computer systems suitable for accessing provided databases.
[0248] Computer systems that can be used to implement all or part of the provided technology may include various forms of digital computers. Examples of digital computers include, but are not limited to, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, smart televisions, and other suitable computers. Mobile devices can be used to implement all or part of the provided technology. Mobile devices include, but are not limited to, tablet computing devices, personal digital assistants, cellular phones, smartphones, digital cameras, digital glasses, and other portable computing devices. The computing devices, their connections and relationships, and their functions described herein are intended only as examples and are not intended to limit the implementation of the present technology.
[0249] All or part of the techniques described herein, and various modifications thereof, may be implemented at least in part by a computer program product that is executed by or controls the operation of a data processing device (e.g., a programmable processor, a computer, or multiple computers), for example, a computer program tangibly embodied in one or more information carriers, such as one or more tangible machine-readable storage media.
[0250] Computer programs for the provided technology can be written in any form of programming language (including compiled or interpreted languages), and can be deployed in any form (including as stand-alone programs or as modules, parts, subroutines, or other units suitable for a computing environment). A computer program can be deployed to be executed on one computer or on multiple computers at one site or distributed across multiple sites and interconnected by a network.
[0251] Actions such as those associated with implementing the procedures and techniques can be performed by one or more programmable processors executing one or more computer programs to perform the provided techniques. All or part of the processes can be implemented as dedicated logic circuitry, such as an FPGA (field programmable gate array) and / or an ASIC (application-specific integrated circuit).
[0252] Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, a processor will receive instructions and data from a read-only memory area or a random access memory area, or both. Some elements of a computer (including a server) include one or more processors for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include or be effectively coupled to receive data from or transfer data to one or more machine-readable storage media, such as mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or both. Non-transitory machine-readable storage media suitable for containing computer program instructions and data include all forms of non-volatile storage, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0253] Each computing device, such as a tablet computer, may include a hard drive for storing data and computer programs, as well as a processing device (e.g., a microprocessor) and memory (e.g., RAM) for executing computer programs. Each computing device may include an image capture device, such as a still camera or a video camera. The image capture device may be built-in or simply accessible to the computing device.
[0254] Each computing device may include a graphics system that includes a display screen. A display screen, such as an LCD or CRT (cathode ray tube), displays images generated by the computing device's graphics system to a user. As is well known, displaying on a computer display (e.g., a monitor) physically transforms the computer display. For example, if the computer display is LCD-based, the orientation of the liquid crystals can be altered by applying a bias voltage, resulting in a physical transformation that is visually apparent to the user. As another example, if the computer display is a CRT, the state of the fluorescent screen can be altered by an electrical impulse, also in a physical transformation that is visually apparent. Each display screen may be touch-sensitive, allowing a user to input information into the display screen via a virtual keyboard. Some computing devices, such as desktop computers or smartphones, may include a physical QWERTY keyboard and scroll wheel for inputting information into the display screen. Each computing device and the computer programs executed thereon may also be configured to accept voice commands and perform functions in response to such commands.
[0255] In particular, the provided technology (methods, collections, databases, systems, etc.) establishes connections between biosynthetic gene clusters, products produced by enzymes encoded by the biosynthetic gene clusters, ETaG, homologous mammalian nucleic acid sequences of ETaG (e.g., human genes), and the like. Thus, in some embodiments, the provided technology can be particularly useful for identifying and / or characterizing human targets of products produced by enzymes encoded by the biosynthetic gene clusters. The provided technology can also be particularly useful for identifying and developing modulators for human targets. For example, in some embodiments, to develop therapeutics for human targets, ETaG of a human target (or a nucleic acid sequence encoding a human target) can be rapidly identified using the provided technology and information about its associated biosynthetic gene cluster and / or biosynthetic products produced by the enzymes of the biosynthetic gene cluster. The products of the associated biosynthetic gene clusters can be further characterized, and if necessary, analogs thereof can be prepared, characterized, and assayed to develop therapeutics with improved properties. The provided technology can be particularly useful for targeting human targets that were challenging and / or considered undruggable prior to the present disclosure.
[0256] In some embodiments, the present disclosure provides methods for evaluating compounds using the identified ETaGs and products encoded thereby. In some embodiments, the present disclosure provides methods comprising:
[0257] At least one test compound is contacted with a gene product encoded by a target gene embedded in a fungal nucleic acid sequence, said embedded target gene being characterized in that it:
[0258] is not required for or involved in the biosynthesis of the product of the biosynthetic gene cluster;
[0259] within a vicinity relative to at least one biosynthetic gene in the cluster;
[0260] homologous to a mammalian nucleic acid sequence; and
[0261] Optionally co-regulated with at least one biosynthetic gene in the cluster; and
[0262] Sure:
[0263] The level or activity of the gene product is altered in the presence of the test compound compared to in the absence of the test compound; or
[0264] The level or activity of the gene product is comparable to the level or activity observed in the presence of a reference agent that has a known effect on that level or activity.
[0265] In some embodiments, the present disclosure provides methods for identifying and / or characterizing a mammalian (e.g., human) target of a product produced by an enzyme encoded by a biosynthetic gene cluster, or an analog of the product, comprising:
[0266] identifying a human homolog of ETaG within a proximal region relative to at least one biosynthetic gene in the biosynthetic gene cluster or within a proximal region relative to at least one biosynthetic gene in a second biosynthetic gene cluster that encodes an enzyme that produces the same biosynthetic product produced by the enzyme encoded by the biosynthetic gene cluster; and
[0267] Optionally, the effect on the target of a product produced by an enzyme encoded by the biosynthetic gene cluster, or an analog of the product, is determined.
[0268] In some embodiments, the present disclosure provides methods for evaluating compounds using products encoded by mammalian (e.g., human) nucleic acid sequences homologous to ETaG. In some embodiments, the present disclosure provides methods comprising:
[0269] At least one test compound is contacted with a gene product encoded by a mammalian nucleic acid sequence homologous to an embedded target gene, wherein the embedded target gene:
[0270] is not required for or involved in the biosynthesis of the product of the biosynthetic gene cluster;
[0271] within a vicinity relative to at least one biosynthetic gene in the cluster;
[0272] homologous to a mammalian nucleic acid sequence; and
[0273] Optionally co-regulated with at least one biosynthetic gene in the cluster; and
[0274] Sure:
[0275] The level or activity of the gene product is altered in the presence of the test compound compared to in the absence of the test compound; or
[0276] The level or activity of the gene product is comparable to the level or activity observed in the presence of a reference agent that has a known effect on that level or activity.
[0277] In some embodiments, the present disclosure provides methods for identifying and / or characterizing a mammalian (e.g., human) target of a product produced by an enzyme encoded by a biosynthetic gene cluster, or an analog of the product, comprising:
[0278] identifying a human homolog of ETaG within a contiguous region relative to at least one biosynthetic gene in the biosynthetic gene cluster, and
[0279] Optionally, the effect of a product produced by an enzyme encoded by the biosynthetic gene cluster, or an analog of the product, on the target is determined.
[0280] In some embodiments, the methods and systems provided can be used to assess the interaction of a human target with a compound. In some embodiments, the present disclosure provides a method for assessing the interaction of a human target with a compound, comprising:
[0281] The nucleic acid sequence of the human target or a nucleic acid sequence encoding the human target is compared to a collection of nucleic acid sequences comprising one or more ETaGs.
[0282] In some embodiments, a compound produced by an enzyme of a biosynthetic gene cluster interacts with a target encoded by a mammalian (eg, human) nucleic acid sequence homologous to ETaG associated with the biosynthetic gene cluster.
[0283] In some embodiments, the provided technology is particularly useful for designing and / or providing modulators for human targets because, among other things, the provided technology provides a link between the biosynthetic gene cluster, ETaG, and the human target gene.
[0284] In some embodiments, the present disclosure provides methods for identifying and / or characterizing a modulator of a human target comprising:
[0285] A product or an analog thereof is provided, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster, wherein an ETaG is present in a proximal region relative to at least one biosynthetic gene in the biosynthetic gene cluster, wherein the ETaG:
[0286] is homologous to a human target or a nucleic acid sequence encoding a human target; and
[0287] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0288] In some embodiments, the human target is a Ras protein. In some embodiments, the Ras protein is a HRas protein. In some embodiments, the Ras protein is a KRas protein. In some embodiments, the Ras protein is a NRas protein. In some embodiments, the human target is a protein comprising a RasGEF domain. In some embodiments, the protein is KNDC1, PLCE1, RALGDS, RALGPS1, RALGPS2, RAPGEF1, RAPGEF2, RAPGEF3, RAPGEF4, RAPGEF5, RAPGEF6, RAPGEFL1, RASGEF1A, RASGEF1B, RASGEF1C, RASGRF1, RASGRF2, RASGRP1, RASGRP2, RASGRP3, RASGRP4, RGL1, RGL2, RGL3, RGL4 / RGR, SOS1, SOS2, or a human guanine nucleotide exchange factor. In some embodiments, the protein is SOS1. In some embodiments, the protein is a human guanine nucleotide exchange factor. In some embodiments, the human target is a protein comprising a RasGAP domain. In some embodiments, the protein is DAB2IP, GAPVD1, IQGAP1, IQGAP2, IQGAP3, NF1, RASA1, RASA2, RASA3, RASA4, RASAL1, RASAL2, or SYNGAP1. In some embodiments, the protein is protein p120. In some embodiments, the protein is human guanine nucleotide activator.
[0289] In some embodiments, the present disclosure provides a method for identifying and / or characterizing a modulator of human Ras protein, comprising:
[0290] preparing analogs of compounds produced by enzymes encoded by biosynthetic gene clusters;
[0291] wherein an ETaG is present in the vicinity of at least one biosynthetic gene in the biosynthetic gene cluster, said ETaG:
[0292] Homologous to a human Ras protein, a RasGEF domain, or a RasGAP domain, or a nucleic acid sequence encoding a human Ras protein, a RasGEF domain, or a RasGAP domain; and
[0293] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0294] In some embodiments, a protein comprising a RasGEF domain modulates one or more functions of a human Ras protein. In some embodiments, a protein comprising a RasGAP domain modulates one or more functions of a human Ras protein.
[0295] In some embodiments, the present disclosure provides a method for identifying and / or characterizing a modulator of human Ras protein, comprising:
[0296] preparing analogs of compounds produced by enzymes encoded by biosynthetic gene clusters;
[0297] wherein an ETaG is present in the vicinity of at least one biosynthetic gene in the biosynthetic gene cluster, said ETaG:
[0298] Homologous to a human Ras protein or a nucleic acid sequence encoding a human Ras protein; and
[0299] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0300] In some embodiments, the present disclosure provides methods for identifying and / or characterizing modulators of a protein comprising a RasGEF domain, comprising:
[0301] preparing analogs of compounds produced by enzymes encoded by biosynthetic gene clusters;
[0302] wherein an ETaG is present in the vicinity of at least one biosynthetic gene in the biosynthetic gene cluster, said ETaG:
[0303] homologous to a RasGEF domain or a nucleic acid sequence encoding a RasGEF domain; and
[0304] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0305] In some embodiments, the present disclosure provides methods for identifying and / or characterizing modulators of a protein comprising a RasGAP domain, comprising:
[0306] preparing analogs of compounds produced by enzymes encoded by biosynthetic gene clusters;
[0307] wherein an ETaG is present in the vicinity of at least one biosynthetic gene in the biosynthetic gene cluster, said ETaG:
[0308] homologous to a RasGAP domain or a nucleic acid sequence encoding a RasGAP domain; and
[0309] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0310] In some embodiments, the biosynthetic gene cluster is shown in the accompanying figures, for example, along with ETaG, which is homologous to Ras proteins. Figures 5 to 12, and 20 to 27, or a biosynthetic gene cluster comprising one or more of the biosynthetic genes described therein. In some embodiments, the biosynthetic gene cluster is shown in the accompanying drawings, for example, together with ETaG, which is homologous to the Ras protein. Figures 5 to 12 , and an exemplary biosynthetic gene cluster of one of 20 to 27. In some embodiments, the biosynthetic gene cluster is shown in the accompanying drawings, e.g., Figures 28 to 33 , and 35, or a biosynthetic gene cluster comprising one or more of the biosynthetic genes described therein. In some embodiments, the biosynthetic gene cluster is shown in the accompanying drawings, for example, together with ETaG homologous to the RasGEF domain. Figures 28 to 33 , and 35. In some embodiments, the biosynthetic gene cluster is shown in the accompanying drawings, for example, together with ETaG homologous to the RasGEF domain. Figure 34 , and 36 to 39, or a biosynthetic gene cluster comprising one or more of the biosynthetic genes described therein. In some embodiments, the biosynthetic gene cluster is shown in the accompanying drawings, for example, together with ETaG homologous to the RasGEF domain. Figure 34 , and an exemplary biosynthetic gene cluster of one of 36 to 39. Exemplary ETaG sequences are provided in the present disclosure and are particularly useful for locating and identifying biosynthetic gene clusters, biosynthetic genes, and the like.
[0311] In some embodiments, the present disclosure provides methods for modulating a human target comprising:
[0312] A product or an analog thereof is provided, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster, wherein an ETaG is present in a proximal region relative to at least one biosynthetic gene in the biosynthetic gene cluster, wherein the ETaG:
[0313] is homologous to a human target or a nucleic acid sequence encoding a human target; and
[0314] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0315] In some embodiments, the present disclosure provides methods for modulating Ras protein, comprising:
[0316] Provide a product or an analog thereof, said product being obtained by Figures 5 to 12 , and enzymes encoded by one of the biosynthetic gene clusters 20 to 27.
[0317] In some embodiments, the present disclosure provides methods for modulating a RasGEF protein, comprising:
[0318] Provide a product or an analog thereof, said product being obtained by Figures 28 to 33 , and 35, wherein the enzyme is encoded by the biosynthetic gene cluster.
[0319] In some embodiments, the present disclosure provides methods for modulating RasGAP protein, comprising:
[0320] Provide a product or an analog thereof, said product being obtained by Figure 34 , and enzymes encoded by one of the biosynthetic gene clusters 36 to 39 are produced.
[0321] In some embodiments, ETaG is identified by provided methods.
[0322] In some embodiments, the product produced by an enzyme encoded by a biosynthetic gene cluster is a secondary metabolite produced by the biosynthetic gene cluster.
[0323] In some embodiments, the analog of the product comprises the structural core of the product. In some embodiments, the product is cyclic, such as a monocyclic, bicyclic or polycyclic ring. In some embodiments, the structural core of the product is or comprises a monocyclic, bicyclic or polycyclic ring system. In some embodiments, the structural core of the product comprises a ring in the bicyclic or polycyclic ring system of the product.
[0324] In some embodiments, the product is linear and the structural core is its backbone. In some embodiments, the product is or comprises a polypeptide and the structural core is the backbone of the polypeptide. In some embodiments, the product is or comprises a polyketide and the structural core is the backbone of the polyketide.
[0325] In some embodiments, an analog is a product substituted with one or more suitable substituents as described herein. In some embodiments, an analog is a structural core substituted with one or more suitable substituents as described herein.
[0326] In particular, the present disclosure provides the following exemplary embodiments:
[0327] 1. A method comprising the following steps:
[0328] a set of query nucleic acid sequences, each of which is present in a fungal strain and comprises a biosynthetic gene cluster; and
[0329] An embedded target gene (ETaG) sequence is identified within at least one fungal nucleic acid sequence, wherein the embedded target gene (ETaG) sequence is characterized in that:
[0330] is not required for or involved in the biosynthesis of the product of the biosynthetic gene cluster;
[0331] within a vicinity relative to at least one gene in the cluster;
[0332] homologous to a mammalian nucleic acid sequence; and
[0333] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0334] 2. The method of embodiment 1, wherein the ETaG sequence is in a contiguous region relative to at least one biosynthetic gene in the cluster.
[0335] 3. The method of any one of the preceding embodiments, wherein the nucleic acid sequence comprising the biosynthetic gene cluster does not comprise a sequence other than the nucleic acid sequence of the adjacent region relative to the biosynthetic genes in the biosynthetic gene cluster and the nucleic acid sequence of the biosynthetic gene cluster.
[0336] 4. The method of any one of the preceding embodiments, wherein the contiguous region is no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 kb upstream or downstream of a biosynthetic gene in the cluster.
[0337] 5. The method of any one of the preceding embodiments, wherein the contiguous region is no more than 50 kb upstream or downstream of the biosynthetic genes in the cluster.
[0338] 6. The method of any one of the preceding embodiments, wherein the contiguous region is no more than 40 kb upstream or downstream of the biosynthetic genes in the cluster.
[0339] 7. The method of any one of the preceding embodiments, wherein the contiguous region is no more than 30 kb upstream or downstream of the biosynthetic genes in the cluster.
[0340] 8. The method of any one of the preceding embodiments, wherein the contiguous region is no more than 20 kb upstream or downstream of the biosynthetic genes in the cluster.
[0341] 9. The method of any one of the preceding embodiments, wherein the contiguous region is no more than 10 kb upstream or downstream of a biosynthetic gene in the cluster.
[0342] 10. The method of any one of the preceding embodiments, wherein the contiguous region is a region between two biosynthetic genes in a biosynthetic gene cluster.
[0343] 11. The method of any one of the preceding embodiments, wherein the mammalian nucleic acid sequence is an expressed sequence.
[0344] 12. The method of any one of the preceding embodiments, wherein the mammalian nucleic acid sequence is a gene.
[0345] 13. The method of any of the preceding embodiments, wherein the mammalian nucleic acid sequence is a human nucleic acid sequence.
[0346] 14. The method of any of the preceding embodiments, wherein the embedded target gene sequence is homologous to the expressed mammalian nucleic acid sequence in that the base sequence of the embedded target gene sequence, or a portion thereof, is at least 50%, 60%, 70%, 80% or 90% identical to the base sequence of the mammalian nucleic acid sequence, or a portion thereof.
[0347] 15. The method of embodiment 14, wherein the sequence or a portion thereof is at least 50, 100, 150 or 200 base pairs in length.
[0348] 16. The method of any one of embodiments 1 to 13, wherein the embedded target gene sequence is homologous to the expressed mammalian nucleic acid sequence in that the product encoded by the embedded target gene, or a portion thereof, is homologous to the product of the mammalian nucleic acid sequence, or a portion thereof.
[0349] 17. The method of embodiment 16, wherein the product is a protein.
[0350] 18. The method of embodiment 16, wherein the protein encoded by the embedded target gene or a portion thereof is at least 50%, 60%, 70%, 80% or 90% similar to the protein encoded by the mammalian nucleic acid sequence or a portion thereof.
[0351] 19. The method of embodiment 16, wherein the protein encoded by the embedded target gene or a portion thereof has a 3-dimensional structure similar to the 3-dimensional structure of the protein encoded by the mammalian nucleic acid sequence or a portion thereof.
[0352] 20. The method of embodiment 19, wherein the portion of the protein encoded by the embedded target gene has a similar 3-dimensional structure to the portion of the protein encoded by the mammalian nucleic acid sequence.
[0353] 21. The method of any one of embodiments 19 to 20, wherein the similarity is that the structures have a Cα backbone rmsd (root mean square deviation) within 10 square angstroms and have the same overall fold or core domain.
[0354] 22. The method of any one of embodiments 19 to 20, wherein the protein encoded by the embedded target gene or a portion thereof has a similar 3-dimensional structure to the protein encoded by the mammalian nucleic acid sequence, in that a small molecule that binds to the protein encoded by the embedded target gene or a portion thereof also binds to the protein encoded by the mammalian nucleic acid sequence or a portion thereof.
[0355] 23. The method of embodiment 22, wherein the Kd of the small molecule binding to the protein encoded by the embedded target gene and the mammalian nucleic acid sequence or a portion thereof is no more than 100 μM, 50 μM, 10 μM, 5 μM or 1 μM.
[0356] 24. The method of any one of embodiments 22 to 23, wherein the small molecule is produced by a fungus.
[0357] 25. The method of embodiment 24, wherein the small molecule is acyclic.
[0358] 26. The method of embodiment 24, wherein the small molecule is cyclic.
[0359] 27. The method of any one of embodiments 24 to 26, wherein the small molecule is a secondary metabolite molecule produced by a fungus.
[0360] 28. The method of any one of embodiments 24 to 27, wherein the small molecule is non-ribosomally synthesized.
[0361] 29. The method of any one of embodiments 24 to 28, wherein the small molecule is a biosynthetic product of a biosynthetic gene cluster.
[0362] 30. The method of embodiment 16, wherein a portion of the protein encoded by the embedded target gene is at least 50%, 60%, 70%, 80% or 90% similar to a portion of the protein encoded by the expressed mammalian nucleic acid sequence.
[0363] 31. The method of embodiment 30, wherein the portion of the protein is a protein domain.
[0364] 32. The method of any one of embodiments 30 to 31, wherein the portion of the protein is a set of amino acid residues necessary for function.
[0365] 33. The method of embodiment 32, wherein the function is an enzymatic function.
[0366] 34. The method of embodiment 33, wherein the set of amino acid residues contacts a substrate.
[0367] 35. The method of embodiment 33, wherein the set of amino acid residues contacts an intermediate.
[0368] 36. The method of embodiment 33, wherein the set of amino acid residues is contacted with the product.
[0369] 37. The method of embodiment 32, wherein the function is an interaction with another entity.
[0370] 38. The method of embodiment 37, wherein the entity is a small molecule.
[0371] 39. The method of embodiment 37, wherein the entity is a lipid.
[0372] 40. The method of embodiment 37, wherein the entity is a carbohydrate.
[0373] 41. The method of embodiment 37, wherein the entity is a nucleic acid.
[0374] 42. The method of embodiment 37, wherein the entity is a protein.
[0375] 43. The method of any one of embodiments 32 to 42, wherein each of the residues in the set is in the entity Inside.
[0376] 44. The method of any of the preceding embodiments, wherein the embedded target gene is co-regulated with at least one gene in the cluster.
[0377] 45. The method of any of the preceding embodiments, wherein the embedded target gene is absent from 80%, 90%, 95% or 100% of all fungal nucleic acid sequences from different fungal strains and comprising homologous or identical biosynthetic gene clusters in the collection.
[0378] 46. The method of any one of the preceding embodiments, wherein the collection comprises at least 100, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 1,500,000, 2,000,000, or 2,500,000 independent fungal nucleic acid sequences.
[0379] 47. The method of any one of the preceding embodiments, wherein the collection comprises nucleic acid sequences from at least 100, 500, 1,000, 5,000, 10,000, 15,000, 20,000, 22,000, 25,000, or 30,000 independent fungal strains.
[0380] 48. The method of any one of the preceding embodiments, wherein the ETaG sequence is not a housekeeping gene.
[0381] 49. The method of any one of the preceding embodiments, wherein the ETaG sequence is or comprises a sequence having homology to a second nucleic acid sequence or a portion thereof in the same genome.
[0382] 50. The method of any one of the preceding embodiments, wherein the ETaG sequence is or comprises a sequence encoding a product having homology to a product or a portion thereof encoded by a second nucleic acid sequence in the same genome.
[0383] 51. The method of embodiment 49 or 50, wherein the homology is at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5%.
[0384] 52. The method of embodiment 49, wherein the homology is at least 70%.
[0385] 53. The method of embodiment 49, wherein the homology is at least 80%.
[0386] 54. The method of embodiment 49, wherein the homology is at least 90%.
[0387] 55. The method of any one of embodiments 48 to 54, wherein the second nucleic acid sequence is or comprises a housekeeping gene.
[0388] 56. The method of any one of embodiments 48 to 55, wherein the ETaG sequence encodes a product that confers resistance to the product of the biosynthetic gene cluster, and the second nucleic acid sequence does not.
[0389] 57. The method of embodiment 56, wherein the ETaG sequence encodes a protein that provides resistance to the small molecule product of the biosynthetic gene cluster, while the protein encoded by the second nucleic acid sequence does not.
[0390] 58. The method of any of the preceding embodiments, wherein the nucleic acid sequences within the collection comprise biosynthetic gene clusters, the biosynthetic genes of which encode enzymes involved in the synthesis of compounds sharing at least one common chemical attribute.
[0391] 59. The method of any one of the preceding embodiments, wherein the nucleic acid sequences are from a plurality of fungal strains.
[0392] 60. The method of any one of the preceding embodiments, wherein the common chemical attribute is or comprises a ring system.
[0393] 61. The method of any one of the preceding embodiments, wherein the common chemical attribute is or comprises a macrocycle.
[0394] 62. The method of any one of embodiments 52 to 61, wherein the common chemical attribute is or comprises an acyclic skeleton.
[0395] 63. The method of any one of embodiments 52 to 62, wherein the compounds sharing at least one common chemical attribute are polyketides.
[0396] 64. The method of any one of embodiments 52 to 62, wherein the compounds sharing at least one common chemical attribute are non-ribosomal peptides.
[0397] 65. The method of any one of embodiments 52 to 62, wherein the compounds sharing at least one common chemical attribute are alkaloids.
[0398] 66. The method of any one of embodiments 52 to 62, wherein the compounds sharing at least one common chemical attribute are terpenes / isoprenes.
[0399] 67. A method comprising the steps of:
[0400] At least one test compound is contacted with a gene product encoded by an embedded target gene in a fungal nucleic acid sequence, wherein the embedded target gene (ETaG) is characterized in that it:
[0401] is not required for or involved in the biosynthesis of the product of the biosynthetic gene cluster;
[0402] within a vicinity relative to at least one biosynthetic gene in the cluster;
[0403] homologous to a mammalian nucleic acid sequence; and
[0404] Optionally co-regulated with at least one biosynthetic gene in the cluster; and
[0405] Sure:
[0406] The level or activity of the gene product is altered in the presence of the test compound compared to in the absence of the test compound; or
[0407] The level or activity of the gene product is comparable to the level or activity observed in the presence of a reference agent that has a known effect on that level or activity.
[0408] 68. The method of embodiment 67, wherein the ETaG is the ETaG of any one of embodiments 1 to 66.
[0409] 69. The method of embodiment 67 or 68, wherein the mammalian nucleic acid sequence is a human Ras sequence.
[0410] 70. The method of embodiment 69, wherein the mammalian nucleic acid sequence is a KRas, HRas, or NRas sequence.
[0411] 71. The method of embodiment 67 or 68, wherein the mammalian nucleic acid sequence is a sequence encoding a RasGEF domain.
[0412] 72. The method of embodiment 67 or 68, wherein the mammalian nucleic acid sequence is a sequence encoding a RasGAP domain.
[0413] 73. The method of any one of embodiments 66 to 72, wherein the ETaG is Figures 1 to 39 One of the ETaG.
[0414] 74. The method of any one of embodiments 66 to 73, wherein the biosynthetic gene cluster is Figures 1 to 39 A biosynthetic gene cluster in one of the
[0415] 75. The method of any one of embodiments 66 to 74, wherein the test compound is a biosynthetic product of the biosynthetic gene cluster or an analog thereof.
[0416] 76. A method comprising the steps of:
[0417] At least one test compound is contacted with a gene product encoded by an expressed mammalian nucleic acid sequence that is homologous to the embedded target gene sequence of any one of embodiments 1 to 75.
[0418] 77. The method of embodiment 76, wherein the mammalian nucleic acid sequence is a human Ras sequence.
[0419] 78. The method of embodiment 77, wherein the mammalian nucleic acid sequence is a KRas, HRas or NRas sequence.
[0420] 79. The method of embodiment 76 or 77, wherein the mammalian nucleic acid sequence is a sequence encoding a RasGEF domain.
[0421] 80. The method of embodiment 76 or 77, wherein the mammalian nucleic acid sequence is a sequence encoding a RasGAP domain.
[0422] 81. The method of any one of embodiments 76 to 80, wherein the ETaG is Figures 1 to 39 One of the ETaG.
[0423] 82. The method of any one of embodiments 76 to 81, wherein the biosynthetic gene cluster is Figures 1 to 39 A biosynthetic gene cluster in one of the
[0424] 83. The method of any one of embodiments 76 to 82, wherein the test compound is a biosynthetic product of the biosynthetic gene cluster or an analog thereof.
[0425] 84. A method comprising:
[0426] identifying a human homolog of ETaG within a proximal region relative to at least one biosynthetic gene in the biosynthetic gene cluster; and
[0427] Optionally, the effect of a product produced by an enzyme encoded by the biosynthetic gene cluster, or an analog of the product, on the human homolog is determined.
[0428] 85. The method of embodiment 77, wherein the ETaG is the ETaG of any one of embodiments 1 to 66.
[0429] 86. A method for identifying and / or characterizing a modulator of a human target, comprising:
[0430] A product or an analog thereof is provided, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster, wherein an ETaG is present in a proximal region relative to at least one gene in the biosynthetic gene cluster, wherein the ETaG:
[0431] is homologous to the human target or a nucleic acid sequence encoding the human target; and
[0432] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0433] 87. The method of embodiment 86, wherein the ETaG is the ETaG of any one of embodiments 1 to 83.
[0434] 88. The method of embodiment 86, wherein the human target is Ras protein.
[0435] 89. The method of embodiment 88, wherein the human target is KRas, HRas or NRas.
[0436] 90. The method of embodiment 86, wherein the human target comprises a RasGEF domain.
[0437] 91. The method of embodiment 86, wherein the human target comprises a RasGAP domain.
[0438] 92. The method of any one of embodiments 86 to 91, wherein the ETaG is Figures 1 to 39 One of the ETaG.
[0439] 93. The method of any one of embodiments 86 to 92, wherein the biosynthetic gene cluster is Figures 1 to 39 A biosynthetic gene cluster in one of the
[0440] 94. A method for modulating a human target, comprising:
[0441] A product or an analog thereof is provided, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster, wherein an ETaG is present in a proximal region relative to at least one biosynthetic gene in the biosynthetic gene cluster, wherein the ETaG:
[0442] is homologous to the human target or a nucleic acid sequence encoding the human target; and
[0443] Optionally co-regulated with at least one biosynthetic gene in the cluster.
[0444] 95. The method of embodiment 94, wherein the human target is Ras protein.
[0445] 96. The method of embodiment 94, wherein the human target is KRas, HRas or NRas.
[0446] 97. The method of embodiment 94, wherein the human target comprises a RasGEF domain.
[0447] 98. The method of embodiment 94, wherein the human target comprises a RasGAP domain.
[0448] 99. The method of any one of embodiments 94 to 98, wherein the ETaG is Figures 1 to 39 One of the ETaG.
[0449] 100. The method of any one of embodiments 94 to 99, wherein the biosynthetic gene cluster is Figures 1 to 39 A biosynthetic gene cluster in one of the
[0450] 101. The method of embodiment 94, wherein the ETaG is the ETaG of any one of embodiments 1 to 93.
[0451] 102. A database comprising:
[0452] a collection of nucleic acid sequences, each of which is present in a fungal strain and comprises a biosynthetic gene cluster;
[0453] The nucleic acid sequence collection is contained in a computer-readable medium.
[0454] 103. The database of embodiment 102, wherein one or more embedded target genes of any one of embodiments 1 to 101 are indexed.
[0455] 104. A system comprising:
[0456] One or more non-transitory machine-readable storage media storing data representing a set of nucleic acid sequences, each of which is present in a fungal strain and comprises a biosynthetic gene cluster.
[0457] 105. A system comprising:
[0458] One or more non-transitory machine-readable storage media storing data representing a set of nucleic acid sequences, each of which is or comprises an ETaG sequence.
[0459] 106. The system of embodiment 105, wherein one or more embedded target genes of any one of embodiments 1 to 101 are indexed.
[0460] 107. A computer system adapted to perform the method of any one of embodiments 1 to 101.
[0461] 108. A computer system adapted to access the database of any one of embodiments 95 to 103.
[0462] Example
[0463] Some non-limiting examples of the provided technology are described below.
[0464] Example 1: Construction of an exemplary database and its exemplary use
[0465] Approximately 2,000 reported fungal genomes were processed using, for example, antiSMASH to identify potential biosynthetic gene clusters, and approximately 70,000 identified biosynthetic gene clusters were added to the database. The initial database was queried using a human target. For example, a BLAST search was performed against the initial library using the protein sequence of human Sec7 to identify ETaGs. As an alternative or supplement, biosynthetic gene clusters can be compared between them. For example, in one process, non-biosynthetic genes that are present in one or more biosynthetic gene clusters (within the adjacent region relative to at least one biosynthetic gene in the biosynthetic gene cluster) but not in most other homologous biosynthetic gene clusters of the same biosynthetic product are identified as potential ETaGs, and are further confirmed by analyzing whether they have homologous mammalian nucleic acid sequences (e.g., human genes) at the nucleic acid level and / or preferably at the protein level. The identified ETaGs can be indexed / labeled and annotated. The database can be searched by nucleotide sequence (e.g., BLASTN; tBLASTx) or protein sequence (e.g., tBLASTn).
[0466] In some embodiments, the results of a BLAST query against a human target are listed in order of sequence homology strength, indicating all putative hits within the database. The DNA sequences of all hit biosynthetic gene clusters are then checked to verify whether one or more open reading frame (gene) homologs of the target protein are within the predicted range of the biosynthetic gene cluster.
[0467] In some embodiments, GenBank formatted sequence files (*.gbk) for each biosynthetic cluster are assembled and collated, and ETaG protein sequences are obtained therefrom using prediction algorithms (e.g., those including antiSMASH) and / or methods. The protein family (pfam) function of the open reading frame can be predicted by, for example, antiSMASH, and the nucleotide distance between each identified ETaG and its closest biosynthetic enzyme predicted by antiSMASH can be determined. In some embodiments, the closer the predicted ETaG is to the biosynthetic enzyme, the higher the probability that the open reading frame encodes a true ETaG.
[0468] Applicants have successfully identified multiple biosynthetic gene clusters with related ETaGs, including several biosynthetic gene clusters containing authentic ETaGs (biosynthetic gene clusters for cyclosporine, figlutamine, lovastatin, mycophenolic acid, and brefeldin).
[0469] In some embodiments, the present disclosure encompasses the recognition that ETaG can be used as a functional homolog (ortholog) of a putative human target protein. In some embodiments, the protein sequence of a putative ETaG hit is compared to the sequence of a human target ortholog. For example, in a project searching for ETaG of human protein A, n biosynthetic gene clusters containing putative protein A homologs are found, and all n predicted ETaG proteins are aligned with human protein A. In some embodiments, only amino acids within a specific catalytic domain or structural domain (e.g., based on a predictive subfamily domain architecture) that defines the pfam boundary of ETaG / target are used in the alignment analysis. By aligning all ETaGs with the human target protein, ETaG sequences are directly compared to their human counterparts, generating quantitative correlation data (e.g., peptide sequence similarity and / or phylogenetic tree visualization) based on their phylogenetic relationships. Additional analysis can include conservation / similarity of essential structural elements for protein effector recruitment / binding, for example, based on an examination of the tertiary protein structure of the human target. For example, in some embodiments, the aligned sequences are compared to PDB crystal structures corresponding to target protein residues within 4 angstroms of the corresponding conjugated protein. Without wishing to be bound by any theory, in the case where these structural motifs are conserved within fungal ETaG, this may indicate an increased likelihood that metabolites produced by ETaG-related biosynthetic gene clusters are effectors of both fungal and human target proteins, and that the produced metabolites may be drug candidates or leads for drug development against human targets. In some embodiments, the above analysis is used to prioritize ETaG and its related biosynthetic gene clusters, as well as the metabolites produced by the biosynthetic gene clusters, for targeting human targets.
[0470] Example 2: Modulators for the human target - Sec7
[0471] In particular, the present disclosure provides techniques for identifying modulators of human targets.In some embodiments, human sequences are used to query provided databases to identify biosynthetic gene clusters in whose vicinity homologs to the human sequences exist.
[0472] For example, in particular, the present disclosure provides biosynthetic gene clusters whose biosynthetic products can modulate Sec7 function. To identify modulators for the human Sec7 domain, the Sec7 protein sequence is used to query a database, such as the database provided in Example 1. An exemplary Sec7 homolog ETaG was identified in Penicillium foetidum IBT 29486, which has a related biosynthetic gene cluster - the ETaG is in a region adjacent to one of the biosynthetic genes in the biosynthetic gene cluster. See Figure 1 、 Figure 18 and Figure 19In particular, the identified biosynthetic gene cluster shares homology with the biosynthetic gene cluster of brefeldianin A in Eupenicillium brefeldianum and is expected to produce brefeldianin A. Thus, brefeldianin A was identified as a candidate modulator of Sec7 and / or a lead compound for modulators thereof. If desired, this result can be optionally validated according to the present disclosure by expressing the biosynthetic gene cluster of P. vulgaris IBT 29486, isolating and characterizing its products, and then assaying the function of the products against Sec7 using a variety of methods available in the art. Since brefeldianin A has been reported to target the Sec7 domain of human GBF1, this example illustrates that the provided technology can be successfully used to identify modulators of human targets.
[0473] Example 3: ETaG of Lovastatin, Felutamide and Cyclosporine
[0474] The provided technology can be used to identify ETaGs of a variety of entities. For example, as demonstrated herein, the provided technology can be effectively used to identify ETaGs associated with lovastatin, figlutamine, and cyclosporine. Exemplary results are shown in Figures 2 to 4 middle.
[0475] Example 4: Modulators for the human target - Ras
[0476] In particular, the present disclosure provides biosynthetic gene clusters whose biosynthetic products can regulate one or more functions of the following proteins: Ras proteins and / or proteins comprising a RasGEF domain (e.g., KNDC1, PLCE1, RALGDS, RALGPS1, RALGPS2, RAPGEF1, RAPGEF2, RAPGEF3, RAPGEF4, RAPGEF5, RAPGEF6, RAPGEFL1, RASGEF1A, RASGEF1B, RASGEF1C, Proteins comprising RASGRF1, RASGRF2, RASGRP1, RASGRP2, RASGRP3, RASGRP4, RGL1, RGL2, RGL3, RGL4 / RGR, SOS1, SOS2, etc.) and / or RasGAP domains (DAB2IP, GAPVD1, IQGAP1, IQGAP2, IQGAP3, NF1, RASA1, RASA2, RASA3, RASA4, RASAL1, RASAL2, SYNGAP1, etc.). Ras proteins (e.g., HRas, KRas, and NRas) are associated with many human cancers but are notoriously difficult targets for drug discovery. In particular, the present disclosure provides techniques for developing Ras modulators (including Ras inhibitors).
[0477] Human Ras sequences were used to query provided databases, such as the database of Example 1. Eight exemplary ETaGs with varying levels of sequence similarity to human Ras proteins were identified from different strains. Related biosynthetic gene clusters encode enzymes that produce different types of compounds. Figures 5 to 12 and Figures 20 to 27 The protein encoded by the identified ETaG is highly homologous to the human Ras protein. For example, the similarity of nucleotide binding residues can be found in Figure 13 , BRAF interacting residues see Figure 14 rasGAP interacting residues see Figure 15 , and SOS interacting residues see Figure 16 .
[0478] Similarly, biosynthetic gene clusters were identified whose biosynthetic products can regulate RasGEF and RasGAP domains. As demonstrated herein, the exemplary biosynthetic gene clusters identified can contain genes and / or modules involved in the synthesis of various types of moieties / products, e.g., terpenes, PKSs, NRPSs, etc. For example, the biosynthetic gene clusters identified and RasGEF and RasGAP homologs are described in Figures 28 to 39 .
[0479] Exemplary ETaG sequences identified are listed below:
[0480] Figure 5 :Thermomyces lanuginosus, Ras ETaG sequence:
[0481]
[0482] Figure 6 : Talaromyces leycettanus CBS 398.68, Ras ETaG sequence:
[0483]
[0484]
[0485] Figure 7 :Sistotremastrum niveocremeum, Ras ETaG sequence:
[0486]
[0487] Sistotremastrum suecicum, Ras ETaG sequence:
[0488]
[0489]
[0490] Figure 8 : Agaricus bisporus var. beinart JB137-S8, Ras ETaG sequence:
[0491]
[0492]
[0493] Figure 9 Okayama Coprinus cinerea, Ras ETaG sequence:
[0494]
[0495] Figure 10 : Higgins anthracis, Ras ETaG sequence:
[0496]
[0497] Figure 11 :Gyalolechia flavorubescens KoLRI002931, Ras ETaG sequence:
[0498]
[0499] Figure 12 : Helminthosporium maydis ATCC 48331, Ras ETaG sequence:
[0500]
[0501]
[0502] Figure 18 : Penicillium foetidum IBT 29486, Sec7 ETaG sequence:
[0503]
[0504]
[0505]
[0506] Figure 20 :Thermomyces lanuginosus ATCC 200065, Ras ETaG sequence:
[0507]
[0508] Aspergillus rambleli, Ras ETaG sequence:
[0509]
[0510]
[0511] Aspergillus ochraceus, Ras ETaG sequence:
[0512]
[0513]
[0514] Figure 21 : Agaricus bisporus var. beinart JB137-S8, Ras ETaG sequence:
[0515]
[0516] Agaricus bisporus H97, Ras ETaG sequence:
[0517]
[0518]
[0519] Okayama cinerea, Ras ETaG sequence:
[0520]
[0521]
[0522] Subbrick red weeping mushroom FD-334, Ras ETaG sequence:
[0523]
[0524]
[0525] Figure 22 :Sistotremastrum niveocremeum, Ras ETaG sequence:
[0526]
[0527] Sistotremastrum suecicum, Ras ETaG sequence:
[0528]
[0529]
[0530] Figure 23: Talaromyces leycettanus CBS 398.68, Ras ETaG sequence:
[0531]
[0532]
[0533] Figure 24 :Thermoascus fragilis, Ras ETaG sequence:
[0534]
[0535] Figure 25 : Helminthosporium maydis ATCC 48331, Ras ETaG sequence:
[0536]
[0537] Figure 26 : Higgins anthracis IMI 349063, Ras ETaG sequence:
[0538]
[0539] Figure 27 :Gyalolechia flavorubescens, Ras ETaG sequence:
[0540]
[0541] Figure 28 :Pine needle brown spot fungus CBS 871.95, RasGEF ETaG sequence:
[0542]
[0543]
[0544]
[0545] Penicillium chrysogenum Wisconsin 54-1255, RasGEF ETaG sequence:
[0546]
[0547]
[0548] Figure 29 : Rice 70-15, RasGEF ETaG sequence:
[0549]
[0550]
[0551]
[0552] Figure 30 : Arthroderma gypseum CBS 118893, RasGEF ETaG sequence:
[0553]
[0554]
[0555]
[0556] Figure 31 : Endocarpon pusillum strain KoLRI No.LF000583, RasGEF ETaG sequence:
[0557]
[0558]
[0559] Figure 32 : Bologna hepatica ATCC 64428, RasGEF ETaG sequence:
[0560]
[0561]
[0562]
[0563] Figure 33 : Aureobasidium pullulans var. EXF-150, RasGEF ETaG sequence:
[0564]
[0565]
[0566] Figure 34 : Acremonium cladodes, RasGAP ETaG sequence:
[0567]
[0568]
[0569] Figure 35 : Pseudomonas lilacinus strain TERIBC 1, RasGEF ETaG sequence:
[0570]
[0571]
[0572]
[0573]
[0574] Neurospora JS1030, RasGEF ETaG sequence:
[0575]
[0576]
[0577] Figure 36 : Corynespora polytomata UM 591, RasGAP ETaG sequence:
[0578]
[0579]
[0580]
[0581]
[0582] Rice strain SV9610, RasGAP ETaG sequence:
[0583]
[0584]
[0585]
[0586] Figure 37 : Colletotrichum oxysporum strain 1 KC05_01, RasGAP ETaG sequence:
[0587]
[0588]
[0589] Figure 38 :Pyrophaga sp. E7406B, RasGAP ETaG sequence:
[0590]
[0591]
[0592] Grape interstices Separation beads DA912 , RasGAP ETaG sequence:
[0593]
[0594]
[0595]
[0596] Figure 39 : Picea truncatula strain 9-3, RasGAP ETaG sequence:
[0597]
[0598]
[0599] Entomophthora RCEF 264, RasGAP ETaG sequence:
[0600]
[0601]
[0602] Using the identified biosynthetic gene clusters, according to the present disclosure, a variety of methods can be used to identify and characterize compounds produced by the enzymes of these biosynthetic gene clusters (e.g., those described in Clevenger, et al., Nat. Chem. Bio., 13, 895-901 (2017) and references cited therein). Once identified, the compounds can be assayed to assess their ability to regulate human Ras proteins. In addition or in lieu thereof, the compounds can be used as lead compounds to prepare more analogs for, for example, SAR studies, to further improve affinity, potency, selectivity, etc. for regulating Ras activity. It is expected that useful compounds will be developed from the biosynthetic gene clusters associated with the identified ETaG.
[0603] Although multiple embodiments have been described and illustrated herein, a variety of other means and / or structures for performing the functions described in this disclosure and / or obtaining the results and / or one or more advantages described in this disclosure will readily occur to those skilled in the art, and each of such variations and / or modifications is considered to be included. More generally, it will be readily understood by those skilled in the art that all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and that actual parameters, dimensions, materials, and / or configurations will depend on one or more specific applications of the teachings of the present disclosure. Those skilled in the art will recognize or be able to determine many equivalents of the specific embodiments of the disclosure described in this disclosure using only routine experiments. Therefore, it should be understood that the aforementioned embodiments are given only by way of example, and that the technology provided (including those claimed) can be practiced in a manner other than that specifically described and claimed. In addition, if two or more features, systems, products, materials, kits, and / or methods are not mutually contradictory, any combination of such features, systems, products, materials, kits, and / or methods is included within the scope of this disclosure. Sequence Listing <110> LIFEMINE THERAPEUTICS, INC. <120> Human therapeutic targets and their modulators <130> 2012710-0002 <140> PCT / US2018 / 051134 <141> 2018-09-14 <150> 62 / 558,744 <151> 2017-09-14 <160> 49 <170> PatentIn version 3.5 <210> 1 <211> 1044 <212> DNA <213> Thermomyces lanuginosus <400> 1 atgcagccgc ggtgagtgtt ggtcgcgctc cttggcaaag gtcaatacta attggacaca 60 ggcgggaata tcatattgtc gtcctgggag ctggtatgtc gagaagagat tcgccacagc 120 ctatcagtcg atatgtgtcc ctaacaatgt tatacaggag gcgtcgggaa gagctgcttg 180 acaggtatgg acgcgatgga ctgcggcgac aacatgcgac cgatggctca ctaacttatc 240 tcatagctca atttgtacaa aatgtttgga ttgagagtta cgacccgaca attgaagatt 300 cctatcgaaa gcagattgaa gtcgatgtga gttcccgtgg cattgatgcg attataccac 360 ctgcttacga tattctattc gcagggtcga caatgcattc tcgagatgta cgtctctctt 420 cagagctgtc gcggagctat ttcatcttac tgatcaccgt gcagtctgga cacagccgga 480 acagagcaat tcagtacgtc ttaacctccc aactccgatg aaaaggacca tccactaacg 540 atgacgacag ctgcgatgag gtattacacg tcaatgcggc gcacatggcc aatgaagttg 600 acatgactgt ccagggaaat ttacatgaaa caagggcagg gattcctgct agtcttctcg 660 atcaccagca tgtcatcgct gaacgagtta tcggaaatcc gggagcagat cctccgcatc 720 aaggacgatg acaaggtccc tatggtgatc gtcggcaaca agtccgatct cgaggaaaac 780 cgagctgtgc ctcgtagcaa agcgtttgcg ctctcgcaga gctggggcaa cgctccttac 840 tacgaaacat ccgctcgacg gcgagcaaac gtcaacgagg tcttcattga cctgtgccga 900 cagatcatcc ccaaggatct ccaagctaca caggcaaagc aagcggaagc cagacaagtt 960 aagcgagagg cgactcctcg caatgacagg agcaagaagg atagaaaatc cacaaggcgt 1020 cggcatcaat gcgcgattat gtga 1044 <210> 2 <211> 2143 <212> DNA <213> Talaromyces leycettanus <400> 2 atgctgggaag tgctagacac agcgggccag gagagtaca ccgcactgag agaccagtgg 60 atccgcgatg gtgaagggtt cgttctcgtc tatagcatca catcgcgagc gtcgttcgcc 120 cgcataccca agttctacaa tcagatcaag atggttaaag aatcggcaag ctccgggtca 180 cccgctggag ccagctactt gacgtcgccg atcaattctc cctcgggacc cccgcttcct 240 gtgccggtaa tgttggttgg caaagagc gacaaggcga tggaccgcgc cgtctctgcg 300 caggaaggcc aagctcttgc caaggagctg gggtgcgaat tcgtcgaggc ttccgccaag 360 aactgtatca atgtcgaaaa ggctttctac gacgtcgtga ggatgcttcg gcagcagcga 420 caacagcaac agggaggacg ggcgcaggag cggcgacccg ccgctttcgg atcagggcca 480 atgcgcgatc gggacgccgg tcccgagtac ccaaagtcgt ttcgtccgga tcgatcaagg 540 catcgcaatg gcctcaaatg cgttatccta tgagctcccc ccgatgagtg ttccgatcgg 600 cggatctttc cagcttctga cctccgctta ttcatgaccg ttgctctcta gaatggatgg 660 tgtctagctc cgtgtttctc tttctcggag cgtgtgagcg agcttgagga cagtcgttcc 720 acttgtgccc cctcctatcc gccgcaggcc cttgtcgctg ccgctttgcg gaccgctcgt 780 tttgtctacg ttgtactcga aagcacggcc tctgctttcg tggaagtctc cctttatgcc 840 agctttgggt gcggtggtcg atatgcagat actgtgttct atgctcgctg catgcgattc 900 agaggcgtct tgattccccg tgtcagtatg gggtgttctc gctattcagg gaatcatctg 960 aaaccaattt ttctcatccg ttctgttttt gggaatcgga acacgggggg gatgtctgga 1020 aatctggacc tataactata gaaatgtttc tcaccacctt tctcactcaa ccctcttgat 1080 gaatatccgc ccggcgtctt ctactacttc ctaccgcta ctaccaccaa tctctattct 1140 tcttaccacc caccttctga gccacttctt acacatcatt ctcgtttggt ttgacagcaa 1200 agcggggaga gttcgaagga cagatcccat gcaggattgg aggacgagag gggaagagtc 1260 gaagggagaa aataattaa aaaaaagaaa ggtgcggggg cagaaggagg caggttttggt 1320 tgagagttgc gaatcggtc tgtcgcagtc aagtcccaaa aaagaaaaga tcgcagtcgg 1380 cgcattagca ggcattttga tacgatgata ccctacagcc gagcttcgag ttttgtgtt 1440 ccttttcctt ttttgcaaat gctgatttaa aaaaataaca atagagctac atactgaatg 1500 tggattttt tgacctctca tctttttgtt gcagggatga ccgccaattg gtaaattcat 1560 ccccagtcat aatccgagcg caggatgcat gaactccagt acctcatcat atcgcctgca 1620 cgttcaagtt ccatcaatca ttcggcggcg cctactctgt acgactaagt ctacggagtt 1680 tgttcttgtt gcggggaagg aagcgaaagc cacgactcca acaaacaaac tcagggtgaa 1740 ttgaatcctc agtttctact ctgtagccga agagccatca ttaccattca ggggaagagc 1800 ctaaagagct tgcgaggttg ggctgagctg ctgtgcagtg agcaatatat ttggtcgatg 1860 ttttggatac gttatctgga atgcgcagat gcagtggtta tgcatatcct cacgtactcg 1920 attctgatga ttcacgggac catacggagt cgataccgag actctcgcta caaacctgtc 1980 aattgatatc gtgtacagag taccggagcc gagactggga aatagcacag tctcagtctc 2040 aggtagctat cgatcaattt gacaaggtta gaagtatctc gctagtaatt gccagatgat 2100 tcattcccgg ttgaaaactt ttccattggc cttcttcgct tag 2143 <210> 3 <211> 989 <212> DNA <213> Sistotremastrum niveocremeum <400> 3 atgtcgagag tgagtatttc tgtttattgc ggctctatct tgatctcact cgtcgctagt 60 ctgctgctca ggcttcgttc cttcgcgaat acaagctcgt cgtcgtcggt ggtggtggta 120 tgagccttgt ctctcgttct ctgcaatcaa aatctcactc gcttttctct tgtgctgcct 180 aggtgttggc aaatccgctc tgaccattca attcatccaa agtcatttcg ttgacgagta 240 tgaccctact atcgaaggtc agccgaccgc taggcaacca ttatctgatc aaacagctca 300 tctcgcactc gacagattct tacagaaagc aatgtgtcat cgatgatgaa gttgcccttt 360 tggatgtgtt agataccgct gggcaggaag aatatgggtg agctcgtctc gcagcccgat 420 tcccacgctt attgctaaca cgacatcggc agcgcaatgc gagaacagta tatgcgaacg 480 ggagaaggat tcttgcttgt ctactcgata acgtcgcgga actctttcga agaaatcagc 540 actttccatc agcaaattct tcgagtaaaa gacaaggatg cgttcccggt tatcgtggta 600 gccaacaagt gtgaccttga atatgagcga caagtcggca tgaacggtgc gtttttagtg 660 ttgtttcaat caacattgtg actcatcctt cgtcagaggg ccgtgacctg gccaagcact 720 tcaactgcaa atttatcgag acctcggcga agcagcgaat caacgttgat gaggcctttt 780 cgaaccttgt tcgagagatt cgcaaattca acaaggtatg taagcccaaa cccgacggaa 840 ctcccggcct gatctcttta caggaacaac agaccggacg tcctgcgacc atggctccga 900 gcggccctgt gggtgcattc ggtggtcccc ccggcatgga agatggacct catgacgctg 960 gttgctgctc tggatgtgtc gttgtataa 989 <210> 4 <211> 989 <212> DNA <213> Sistotremastrum suecicum <400> 4 atgtcgagag tgagtatttc tgtttattgc ggctctatct tgatctcact cgtcgctagt 60 ctgctgctca ggcttcgttc cttcgcgaat acaagctcgt cgtcgtcggt ggtggtggta 120 tgagccttgt ctctcgttct ctgcaatcaa aatctcattc gcttttctct tgtgctgcat 180 aggtgttggc aaatccgctc tgaccattca attcatccaa agtcatttcg ttgacgagta 240 tgaccctact atcgaaggtc agccgaccgc taggcaacca ttatctgatc taacagctca 300 tctcgcactc gacagattct tacagaaagc aatgtgtcat cgatgatgaa gttgcccttt 360 tggatgtgtt agataccgct gggcaggaag aatatgggtg agctcgtctc gcagcccgat 420 tcccacgctt attgctaaca cgacatcggc agcgcaatgc gagaacagta tatgcgaacg 480 ggagaaggat tcttgcttgt ctactcgata acgtcgcgga actctttcga agaaatcagc 540 actttccatc agcaaattct tcgagtaaaa gacaaggatg cattccctgt tatcgtggta 600 gccaacaagt gtgaccttga atatgagcga caagttggca tgaacggtgc gattctagtg 660 ttgtttctgt cgatattggg acttatcccc cttcagaggg ccgtgatttg gccaagcact 720 tcaactgcaa atttatcgag acatcggcga agcagcgaat caacgttgat gaggcctttt 780 ccaaccttgt tcgagagatt cgcaaattca acaaggtatg taagcccaaa cccgacggaa 840 ctcccggcct gatctcttta caggaacaac agaccggacg tcctgcgacc atggctccga 900 gcggccctgt gggtgcattc ggtggtcccc ccggcatgga agatggacct catgacgctg 960 I gttgctgctc tggatgtgtc gttgtataa 989 <210> 5 <211> 1127<' <212> DNA !<213> Agaricus bisporus <400> 5 atggcaaaca acgctgcgtc cagagtatgt cctccccaca aaccaccctc agttgcctgg 60 cttatgctct atttcaggct gctcaggccc agttcctgag agaatacaag ctcgtagtgg 120 ; tcggaggagg aggcaagtgc tacccgccct tacaagctag caagtcctaa agtcgtgtac 180 aggtgttgga aaatctgcat tgactatcca attcattcaa agccatttcg tggacgagta 240 It should be noted that there may be some inaccuracies in the above translation because the original text seems to be a DNA sequence with some special tags, and the translation is mainly based on literal conversion. For a more accurate understanding and translation in the context of patent texts, it may need to be combined with relevant biological and patent knowledge.cgacccaact atcgagggtg agcttctttc tcaccaatca atccccttcc aggttatgac 300 atttcggaac atttgtgcta acattctcgt cttaaaacag actcgtacag gaaacaatgc 360 gtcattgatg aagaggtcgc ccttctcgat gtcctggata ccgctggtca agaagaatat 420 gggtcagtgt gctctcctga ataaattccg aagcagtccc cgattttttt tccttcgtc 480 tcgtgattcg actatgaaaa tggtcttcca cgaggcgaag ctttcatttc ccggcataat 540 tcagtatac gaccctggat ctaaccctat atgtacttat tttccagtgc catgcgggag 600 caatacatgc gtactgggga gggatttctt ctcgtctaca gcatcaccgc gcgtagctcc 660 tttgaagaaa tcaaccagtt ttaccagcaa atttgaggg tcaaagatca agattctttc 720 cctgttattg tcgttgcaaa caagtgcgat ttggaatatg aacgccaagt tggtatgaac 780 ggtatgttat caaaccttgg agtatatcag ggccccagta gtgacgcaac ctacagaggg 840 ccgagatctc gcgagacatt ttggctgcaa attcatcgag acgtctgcca aacaacgaat 900 aaacgtggat gaagctttca gcaatctttgt tcgtgaaatc cgaaaatata acaaggtcgg 960 ttttccgcat cacacgcaga gattttacaa actcattggt gcttttatag gaccaacaaa 1020 caggccgccc tctccacggc agcggtggtg gagccggcgg ttatggtggc aaggaccaca 1080 atgacgatgg aggtgctggc tgctgcggcg gttgcgttat tctttaa 1127 <210> 6 <211> 1681 <212> DNA <213> Coprinopsis cinerea <400> 6 atgcctgagg tgatgaatgc tatgtacgcc acgaaaggcg gtatcttcga cgtcagcgag 60 aatgataagg tttggcgttt gcagtgtttc aaagctggcc tctgttgtgt tggagaatac 120 tcggatgctg atatacatat ggtttatcga atacaagtac gcgtataagg aaagctgggg 180 tagacaaggg actatagctg gatcttaact cccaggaggg gacgacatga gagaatgcgg 240 tctacagcaa ttctgatgct cgaaaatcca tcagcagagg tcaaccttg gtttctagcg 300 aaaagaaggg agataggaag cccggaatat caaaacacgc gtcggattgt ggtccaaatt 360 gaaaaatgac cgagagcctc gagctcgtgt cgcgagatgt ttgcacttga gatttaaact 420 ccgctgatga tggcctttga agtgagtttg gttacgatgt ttagaggaac ccagtcgccc 480 cctgctcccg ctcaactccc taaataccct tcctgaccat cttctttctt tcccaaatct 540 ttttcttctc tttcaacaga tttcatttct gaagcatggc tgccagggtc cgtcaaatcc 600 cacagtctgc accgtggaac ctcagcaaac tcacacagcg tccaacaggc tcagttcttg 660 agggagtaca agctcgtcgt cgtaggtggt ggtggtatgt tgcacagctc ttagaacgga 720 atgtagtctc acctgtggtg ccccaggtgt tggaaagtcg gccctgacta ttcagttcat 780 ccaatcccac ttcgtggatg aatatgaccc gactatcgaa ggtccgtata acaaggcctt 840 ctctcgcaag gatgcaatag cttatgctta ttcgacacag actcgtacag aaaacagtgc 900 atcatcgacg acgaggtcgc actcctcgac gttctcgata ccgccggaca ggaagagtat 960 gggtgagtac ccgcgctgca cccctctatt ttccaccgaa tgcttcgtgg acagcccaac 1020 ttttgatcct cgtatcccat accaccgctt tccttgttcc cggaatcttt gcatcaccac 1080 ctctccacct tgccctcttc ttcgggacgt tccgtgatta acacacacct acagagccat 1140 gcgggagcaa tacatgcgca cgggcgaagg cttccttctc gtctaccta tcacctccag 1200 aaactcgttt gaggaaatca gcattttcca ccaacaaatt ttgcgagtca aggaccagga 1260 ttccttcccc gtcattgttg tggctaacaa gtgcgatctc gaatatgaac gtcaagttgg 1320 catgaacggt gtgtagtcca tctttatgtc ccttgccgac atgacatgaa caacgtattg 1380 cagaggggcg tgatctcgcc aaacactttg gttgcaaatt catcgaaacc tcggccaagc 1440 aacgaatcaa cgtcgacgag gcattcagca acctcgttcg ggagattcgc aagtcaacag 1500 ggtgagcaat cctctttcc aaggtattct gactagcatt caaactgtct catgccccca 1560 ggaacaaaa accggtcgtc ctgccatcgc agcaggtgga ggtggtccag ccggctccta 1620 cacccaggac aggcaccacg atgaggcacc tggatgctgt gccggatgtg ttattgccta 1680 to 1681 <210> 7 <211> 961 <212> DNA <213> Colletotrichum higginsianum <400> 7 atggcgtcca aggttcgtcg tcgccacctc ccgtttccct tcatttcttt tgccgcctcg 60 tcgctccatc gctccatcgc cccatcgatc cgttgctaac cagttgccat ctcgcagttt 120 ctgagggagt acaagttggt cgtcgtcggc ggcggtggtg tcggtaaatc ctgcttgacc 180 atccaattga ttcagagcca ctttgtcgac gaatatgacc cgacgatcga aggtgcgtcg 240 tcccgaactt cttgctccac cgttcgatgc gacggcttcg aatcaatcgc atgctaatgt 300 ggatctcacc catttcagat tcctaccgca agcagtgcgt catcgacgag gaggtcgctc 360 tactcgatgt cctcgacacg gccggtcagg aggagtactc cgccatgagg gagcagtaca 420 tgaggacggg agagggtttc cttctggttt actccatcac ttcgcgacag agcttcgagg 480 agatcaccac attccagcag cagattctga gagtaaagga caaggactac ttccccatgg 540 600 ttctgattcc tgctgtgccg cgacaccgca tgaggcggct cctttcgagg cccaggcccg 660 gtgtggattc attgatggaa tgaaaagtag ctgacatcat tcactcgtgc gcgctacaga 720 gggagaggcc cttgccaagt cattcggctg caagttcatc gagacgtcgg ccaagtctcg 780 catcaacgtc gacaaggctt tctatgatat tgtccgagaa atccgtcggt acaaccgcga 840 gatgcagggc tactctaccg gcagtggcgg cgcctcgggc atcaacggcc ccccgaagcc 900 catggacgtc gagaacggcg agcaagaggc aggctgctgc tccaagtgcg tactaatgtg 960 a 961 <210> 8 <211> 326 <212> DNA <213> Gyalolechia flavorubescens <400> 8 atggcttcaa aggtaagtcc atctgtctct ttagagtatt ctcattgctc tttgctaccg 60 agcttctcca tggacgctga cccttacctg ctcaagttcc tacgggaata caagctcgtc 12 gtcgttggcg gaggaggtgt gggcaagtcc tgcttgacca tccagctcat ccagagtcac 180 ttcgtcgacg aatacgatcc caccattgaa ggtaaataga ttcgtcctat ccacccattg 240 cgcttttact gatcgaagcg atttgcaaga ctcctaccgg aagcaatgcg tcatcgacga 300 agaagtcgcc ttactcgatg tactag 326 [[ID=JO]]<210> 9 <2JJ> 809 <212> DNA <213> Bipolaris maydis <400> 9 atgttcttgc ctcaactcta ctccctcaac cctgccttgg ctgccaaaca tgctgatcct 60 cttgctccta cagcccagtt tgtgcaaaac gtgtggatag agagctatga tcccaccatc 120 gaggactcgt accgaaaggt cctcgaagta gacgtgcgta cacgacactc ttatactagccg 180 cgtttttttc actgacccac tctccctccc agggccgtca tgtcattctc gagatcttgg 240 atactgccgg cacagagcag tttagtaagt gattacatac atagccccac cccacgtgga 300 cccaagacta acacgacaat agctgccatg aggtagagtt tcctactacc cccttactcg 360 gtaaacatca aaacttacac ggatgcagag aactgtacat gaaaacgggc caaggattcc 420 ttttggtctt cagcatcaca tcagaatctt ccttttggga gcttgccgag ctgcgtgagc 480 agatacgacg catcaaggaa gacagcaacg tacccatggt tctcattggc aacaagtcgg 540 Your account cgaccgtgcc gtgccgcgcc cacgagcatt tgccatttcg cgtgaatgga 600 acgttcctta tttcgaaacc agtgctcgaa ggagagccaa tgtcgacgaa gcctttgtcg 660 acctctgcag gcaaatcatc cgcaaggatc agaacgaacg aaccgcatg gccccaccgg 720 attccccgag gcctggcggt cccaggagca gaactcacac gggacggcca aagcgcaagg 780 ctcaccggcc ccattgtacc attctttaa 809 <210> 10 <211> 3105 <212> DNA <213> Penicillium vulpinum <400> 10 atggaagttg aaaggccgga tggttggcat catcactgct gaccacgaac agacatcacg 60 aaatatgaca cacctgctct actacacccc ttcctccaag taatacgatc atcctcaacc 120 tctgctgcga tcacctccct cgctctcatc gccatcacga aattcctctc ctacaaaata 180 atctccggtg actccgctcg gctggctgaa gccatgcagc tcctctcatc agctctcacg 240 cactgccggt tcgaggcaag cgattcagca acggacgaaa ttgtactcct gcaggtactg 300 aatctcatgg aaagcattat cttgagtcca ggaggtgaat ctctctgcaa tgagagcgtt 360 tctgagatga tgcaaactgg actgagcatg tgctgccaac ccaggctttc ggaactccta 420 cgacagtctg ctgagattgc catggtctct atttgccaat tggtcttcga gcgatggaag 480 cacctagaag aagaggtggg cgaagagcta ggggccttgg atcaggatgt cagggccgat 540 atgggcacga tgaagctcct tgattcaaaa atgcagacct ccttgaccgg tccaaactcc 600 aagaatctta aatctgagga gaagacacgg tctttgcga gcgtggagaa gctgatcaat 660 gagtccacag ggatgacact gcaaaagggc gacgccacaa ttgatctacc gtcaatgcac 720 gatgaacagg atgaaggcga ggcgctccca atcaagccat actccctgac gttgatacga 780 gagcttcttg tgatcctcat caatatacta gatcctgagg acaagaaaca aacagacaca 840 atgcgtatca cggcactgcg cattctgcat gttgtgttgg aagtagcggg cccatcaatc 900 gtcaaccata ctagtctagc aaccctgaca aggacacgc tatgccgata cttgctccaa 960 ttggttcgct tggataacat gaagattatc agcgagttgc tctgtgtgtg tgttacttta 1020 tttgcaacat gccgaggtgt actcaagcta cagcaggagc tattcctatc gtatgtggtg 1080 acctgtgtgt ttccgacaat ggatattccg ctagagcctg gtatcgaccc ttctctttac 1140 gagggtgtac cgcagtcatt cagcctcctc aagcaatcaa aatcacagtc acctgcgcaa 1200 aaatctacaa gtggcaaatc gacgcccaag tctgccaagg atcgacagaa gctgggaccc 1260 gaagacagca taaggacacc cgatgctcgt gaggcgatat tagagagcgt gagcgccttg 1320 gttaggatcc cctctttat ggtcgagctg ttcgtcaact acgattgcga tattgataga 1380 agcgacctat gttcggatct ggttggactt ctttcgcgga acgctttccc agactcagcc 1440 caggggagta caacaaacct cccaccgcta tgtttggact ctcttctatc ctatgtgcaa 1500 tccattgcag atagactcga tgatgcgccc ctgatagagg gcttccgtga ccccaatgcc 1560 ctacgacagc agcggtcacg taagagtatg attatgaagg gtgcctcgaa attcaatgag 1620 aacccaaagg ctggcatcgc atttctagtc gcccaagggg tcatacaaga gcctgagaat 1680 cctaagaaca ttgcggagtt tatcaaaggc actacgagaa ttgacaagaa gatcctgggg 1740 gagtttattt caaagaaaac aaacgaaaat atattgaacg agttcatgaa gctttttaac 1800 ttcgccggaa aacgaattga cgaggctata cgcgagttac tgggtgcatt ccgcttcct 1860 ggtgagtcgg cacttataga gcgaattgtg gaggtgttcg ctgcacagta tatggacgac 1920 gccaaacccg caggaattgc agactccact gcagcatttg ttctcgtgta tgccaccatc 1980 ttgttaaaca cagatcagca taatcccaat ttcaggggcc agaaacgtat gaccattgag aactttgccc agaatctcag gggtgttaac gatcaggggg actttgattc caacttcctt caggaaatct ttgattctat ccggacacat gagattatcc tgccagagga gcatgatgat aagcatgcct atgattacgc ttggaatgag ctgttgatca aggccgaatc cacttcagac ttggtgtctt gcaacaccaa catttttgat gcggatatgt tcgcggcaac atggaagcca atcgtcggga cactatcata tatgttcata tccgcgactg acgacgctgt gctttcaaaa atagtaaccg gtttcggcca gtgcggtcag attgctgcga agtacagact aagtgatgcc ttggatagga tagtagcctg tctgtcgcat atcagcacgc ttgctccaga agtcacacca agcacgagtc tcaaaatcga ggtccagcat gagaaactta gtgtaatgat ctccgaaacc gccgttcgat ttggggcgtga tgacagagcc cagcttgcaa cagtagtgct gttccgaatt ctcaatggta acgagggtgc aattcgggat ggatgggac aggtaagact tccatcaaca aaagcaattg agatataa gctcacagct gtagattctg cgaatattgc tgaatctttt tataaattct ttgataccct cctcgttctc ctcagctcgc aaatcgcttg aacttccatc 2760 tatcccgcta caaagcccca ctcagatcat caacaaagat gatagagcag cggacaccag 2820 cctgttttat gcctttgctt cttatgtttc gagtttcgcg aacggtgagc caccggaacc 2880 ttcagacgaa gagattgaga acaccttgtg tacaatcgat actatcagcg cttgttcgtt 2940 ggacgaaatc acatccaaca tcttgtaagt cataaaccgc gtggctaatc aggacatgaa 3000 ttaacaagac ctctagcgac atgtccacag aggctttgag acctctgttc atggcgcttt 3060 tgtcacgact acccgaagat acatcgctcc acggtattgc agtaa 3105 <210> 11 <211> 1044 <212> DNA <213> Read more <400> 11 atgcagccgc ggtgagtgtt ggtcgcgctc cttggcaaag gtcaatacta attggacaca 60 ggcgggaata tcatattgtc gtcctgggag ctggtatgtc gagaagagat tcgccacagc 120 ctatcagtcg atatgtgtcc ctaacaatgt tatacaggag gcgtcggggaa gagctgcttg 180 acaggtatgg acgcgatgga ctgcggcgac aacatgcgac cgatggctca ctaacttatc 240 tcatagctca atttgtacaa aatgtttgga ttgagagtta cgacccgaca attgaagatt 300 cctatcgaaa gcagattgaa gtcgatgtga gttcccgtgg cattgatgcg attataccac 360 ctgcttacga tattctattc gcagggtcga caatgcattc tcgagatgta cgtctctctt 420 cagagctgtc gcggagctat ttcatcttac tgatcaccgt gcagtctgga cacagccgga 480 acagagcaat tcagtacgtc ttaacctccc aactccgatg aaaaggacca tccactaacg 540 atgacgacag ctgcgatgag gtattacacg tcaatgcggc gcacatggcc aatgaagttg 600 acatgactgt ccagggaaat ttacatgaaa caagggcagg gattcctgct agtcttctcg 660 atcaccagca tgtcatcgct gaacgagtta tcggaaatcc gggagcagat cctccgcatc 720 aaggacgatg acaaggtccc tatggtgatc gtcggcaaca agtccgatct cgaggaaaac 780 cgagctgtgc ctcgtagcaa agcgtttgcg ctctcgcaga gctggggcaa cgctccttac 840 tacgaaacat ccgctcgacg gcgagcaaac gtcaacgagg tcttcattga cctgtgccga 900 cagatcatcc gcaaggatct gcaagctaca caggcaaagc aagcggaagc cagacaagtt 960 aagcgagagg cgactcctcg caatgacagg agcaagaagg atagaaaatc cacaaggcgt 1020 cggcatcaat gcgcgattat gtga 1044 <210> 12 <211> 1054 <212> DNA <213> Aspergillus rambelli <400> 12 atgctgggaa tagcggtcac tataatgcc tccttcggtg tgaccggtag acggggaatat 60 cacattgtcg tgttgggtgc tggaggagtg ggaaaaagtt gtcttactgg tatgattctc 120 ggtcgcgtcg gcttcgtgct tgcctcggaa ggccgtctct gctctctaga ccaatcagtc 180 gcttacttgt ggcagcgcaa tttgtgcaaa acgtttggat tgaaagctat gatccgacga 240 ttgaagactc ttatcgcaag catatcgagg tagatgtatg tttatcctgc tctcaacttc 300 attcgggt tcattctcaa gtcgctgaca tttctaggg ccgacaatgt attctggaaa 360 tgtatgtcac aaggaacacg gatggtggtt cggaattgcg ctttacgtgt aaacaaacac 420 ggctggctga cccttgacct gtcaacagac ttgatacagc ggggacagaa caatttagtg 480 agttatcttg ctcttgatgc tgggttttct ctccactaac gttttcccag cggccatgag 540 gtaatgaatg ctatatccat ggggtcatcg ggactcacat ctctcagttg ccagatctcg 600 atcgctaaca tgtgaatcct gcagagaact atatatgaag caggccagg gctttttgct 660 tgtattctct atcactagca tgtcgtctct gaacgagctg tccgaattac gagaacaaat 720 tattcgcatt aaagacgacg agaaagttcc catcgtcatt gtgggcaata aatcggattt 780 ggaggaagac cgcgcagtcc cacgtgctcg tgcattgct ctttctcaga gctggggcaa 840 cgctccctac tatgaaacat cggcgcgtcg acgagccaat gttaatgagg tcttcattga 900 cctgtgtcga cagattatac ggaaggacct ccagggaagt tcgaccagcg attatgatgc 960 tgccgcacgt aaacgcgagg gtcaacccg acagaccga aagcgagaga gaaacgaca 1020 agtgcggcga aagggtcctt gtgtcattct ctaa 1054 <210> 13 <211> 1054 <212> DNA <213> Aspergillus ochraceoroseus <400> 13 atgctgggaa tagcggtcac taataatgcc tccttcggtg tgaccggtag acggaatat 60 cacattgtcg tgttgggtgc tggagtg ggaaaagtt gtcttactgg tatgattc 120 ggtcgcgtcg gcttcgtgct tgcctcggaa ggccgtctct gctctctaga ccaatcagtc 180 gcttacttgt ggcagcgcaa tttgtgcaaa acgtttggat tgaaagctat gatccgacga 240 ttgaagactc ttatcgcaag catatcgagg tagatgtatg tttatcctgc tctcaacttc 300 attctcgggt tcattctcaa gtcgctgaca ttttctaggg ccgacaatgt attctggaaa 360 tgtatgtcac aaggaacacg gatggtggtt cggaattgcg ctttacgtgt aaacaaacac 420 ggctggctga cccttgacct gtcaacagac ttgatacagc ggggacagaa caatttagtg 480 agttatcttg ctcttgatgc tgggttttct ctccactaac gttttcccag cggccatgag 540 gtaatgaatg ctatatccat ggggtcatcg ggactcacat ctctcagttg ccagatctcg 600 atcgctaaca tgtgaatcct gcagagaact atatatgaag caaggccagg gctttttgct 660 tgtattctct atcactagca tgtcgtctct gaacgagctg tccgaattac gagaacaaat 720 tattcgcatt aaagacgacg agaaagttcc catcgtcatt gtgggcaata aatcggattt 780 ggaggaagac cgcgcagtcc cacgtgctcg tgcatttgct ctttctcaga gctggggcaa 840 cgctccctac tatgaaacat cggcgcgtcg acgagccaat gttaatgagg tcttcattga 900 cctgtgtcga cagattatac ggaaggacct ccagggaagt tcgaccagcg attatgatgc 960 tgccgcacgt aaacgcgagg gtcaaacccg acaagaccga aagcgagaga gaaaacgaca 1020 agtgcggcga aagggtcctt gtgtcattct ctaa 1054 <210> 14 <211> 1124 <212> DNA <213> Agaricus bisporus <400> 14 atggcaaaca acgctgcgtc cagagtatgt cctccccaca aaccgccctc agtttcttgg 60 cttatgctct atttcaggct gctcaggccc agttcctgag agaatacaag ctcgtagtgg 120 tcggaggagg aggcaagtgc tacccgccct tacaagctag caagtcctaa agtcgtgtac 180 aggtgttgga aagtctgcat tgactatcca attcattaaa agccatttcg tggacgagta 240 cgacccaact attgagggtg agcttctttc tcaccaatca atccccctcc aggttatgac 300 atttcggaac atttgtgcta acattctcgt cttaaaacag actcgtacag gaaacaatgc 360 gtcattgatg aagaggtcgc ccttctcgat gtcctggata ctgctggtca agaagaatat 420 gggtcagtgt gctctcctga ataaattccg aagcagtccc cgattttttt tcctttcgtc 480 tcgtgattcg actatgaaaa tggtcttcca cgaggcgaag ctttcatttc ccggcataat 540 tcagttatac gaccctggat ctaaccctat atgtacttat tttccagtgc catgcgggag 600 caatacatgc gtactgggga gggatttctt ctcgtctaca gcatcaccgc gcgtagctcc 660 tttgaagaaa tcaaccagtt ttaccagcaa attttgaggg tcaaagatca agattctttc 720 cctgttattg tcgttgcaaa caagtgcgat ttggaatatg aacgccaagt tggtatgaac 780 ggtatgttgt taaaccttgg agtatatcag ggcccagtag tgacgcaacc tacagagggc 840 cgagatctcg cgagacactt tggctgcaaa ttcatcgaga cgtctgccaa acaacgaata 900 aacgtggatg aagctttcag caatcttgtt cgtgaaatcc gaaaatataa caaggtcggt 960 tttccacatc acacgcagat tttacaaact cattggtact tttataggac caacaaacag 1020 gccgccctct ccacggcagc ggtggtggag ccggcggtta tggtggcaag gaccacaatg 1080 acgatggagg tgctggctgc tgcggcggtt gcgttattct ttaa 1124 <210> 15 <211> 1031 <212> DNA <213> Catfish (Hypholoma sublateritum) <400> 15 atggctgcta gggtacgtcc cttcacataa ctagccaacg tcgcgtagct catgccctct caggctcagt tcttgcgaga atacaagttg gtggtggtgg gcggaggagg tcagcaaatc ctggcgccat ttcccggtct ttctcctgct cacagtttcc ttcaggtgtc ggaaagtctg 180 ctttgactat tcagttcatt caaagccatt tcgttgacga gtacgatccc accatcgagg gtgagagttt cgtgcttcca gtgccgccgc gacgctgacc gagtcaaga ttcgtaccgt 300 aagcaatgcg taatcgacga ggaggttgct ctcctcgacg ttctggacac tgctggtcag gaggagtacg ggtacgtgtc tgtctttacc attack tcctccccct gttctttttt ggctcgcgcc tcgaggcgcg ttcttgctct ggtgctattc ttatcatggc tgttctctga 480 cggaaatacg tatagtgcta tgcgcgaaca atacatgcgt accggcgagg gtttcttgct cgtctactcc attackccc gcgactcctt cgaggaata agcacattcc accaacagat tctgcgggtc aaggaccagg actcgttccc cgttatcgtt gttgcgaaca agtgcgattt 660 ggagtacgag cgccaggttg gcatgaatgg tacggcagta gaccaccagg ctggagatg 720 ctaatcaact atctctca gaggggccgtg accttgccaa gcactcggt tgcaagttca 780 tcgaaacgtc agccaagcag cggatcaacg tcgatgaggc ttcagcaac cttgttcgcg 840 agattcggaa gtataatag gttagtacgt tatgttattc tacctctccc tatctgacag 900 atattgtcca ccaggaacaa caactggtc gcccggccct tgccggcaat ggaggaagca 960 ctggcgcata cgatgggaaa gaccaccacg atgatactcc tggtgttgt tccggctgtg 1020 ttgtccctcta a 1031 <210> 16 <211> 677 <212> DNA <213> Thermoascus crustaceus (Thermoascus crustaceus) <400> 16 atgacccaac aatcgaaggt tggtcaccgt taagcaacc acgatgggag cgtcccgacc 60 atgatggctc attagatctc ttctctcca gactcgtacc gcaagcagtg tgttattgac 120 gatgaggtcg ccctgttgga cgtcctggat accgccggcc aggaggata ctcagccatg 180 cgagaacagt acatgagaac gggagagggg ttccttctgg tgtactct aacttcgcgt 240 cagtcgttcg aggaaatcat gaccttccaa caacagatct tgcgagtcaa ggacaaggat 300 tatttcccca tcattgtcgt cggcaacaag tgtgatctgg agaaggagag agtggtcacg 360 caagaaggta tgtctttaag ctctccgtcg gcttttgaaa cttggctgga gtgccttgct 420 aatcacatta ccgcttctca acagagggtg aggctctcgc gaagcaattc ggctgcaaat 480 tcctggaaac ctcggcgaag tcgcgtatta atgttgaaaa cgcgttctac gaacttgtgc 540 gtgagatccg ccgctacaac aaagagatgt catcctcgtc cggtggcggt gcgggcgcgc 600atccaattga ttcagagcca ctttgtcgac gaatatgacc cgacgatcga aggtgcgtcg 240 tcccgaactt cttgctccac cgttcgatgc gacggcttcg aatcaatcgc atgctaatgt 300 ggatctcacc catttcagat tcctaccgca agcagtgcgt catcgacgag gaggtcgctc 360 tactcgatgt cctcgacacg gccggtcagg aggagtactc cgccatgagg gagcagtaca 420 tgaggacggg agagggtttc cttctggttt actccatcac ttcgcgacag agcttcgagg 480 agatcaccac attccagcag cagattctga gagtaaagga caaggactac ttccccatgg 540 600 ttctgattcc tgctgtgccg cgacaccgca tgaggcggct cctttcgagg cccaggcccg 660 gtgtggattc attgatggaa tgaaaagtag ctgacatcat tcactcgtgc gcgctacaga 720 gggagaggcc cttgccaagt cattcggctg caagttcatc gagacgtcgg ccaagtctcg 780 catcaacgtc gacaaggctt tctatgatat tgtccgagaa atccgtcggt aaaccggaga 840 gatgcagggc tactctaccg gcagtggcgg cgcctcgggc atcaacggcc ccccgaagcc 900 catggacgtc gagaacggcg agcaagaggc aggctgctgc tccaagtgcg tactaatgtg 960 a 961 <210> 18 <211> 326 <212> DNA <213> Gyalolechia flavorubescens <400> 18 atggcttcaa aggtaagtcc atctgtctct ttagagtatt ctcattgctc tttgctaccg 60 agcttctcca tggacgctga cccttacctg ctcaagttcc tacgggaata caagctcgtc 120 gtcgttggcg gaggaggtgt gggcaagtcc tgcttgacca tccagctcat ccagagtcac 180 ttcgtcgacg aatacgatcc caccattgaa ggtaaataga ttcgtcctat ccacccattg 240 cgcttttact gatcgaagcg atttgcaaga ctcctaccgg aagcaatgcg tcatcgacga 300 agaagtcgcc ttactcgatg tactag(2) + 326 <210> 19 <211> 3190 <212> DNA <213> Lecanosticta acicola <400> 19 atggagctcc ctttcgagaa cccgaccgca acaactgaac caggcccgcg agatcgaaat 60 aatttctttg tccccgacca gacacggcca cctccagagc tgatggcgcg tggctttgag 120 Note: In the translation of line 24, the original text "tactag 326" seems a bit unclear. I translated it as "tactag(2) + 326" based on the context, but it might need further clarification depending on the actual meaning in the original patent content.cgggatgagg acgagtacga tggatctgca tcggaggcag aaggagatc actgatgtta 180 ggctcgcatg actcgatttc tcgccgacgc cagtccgtga tggatggagt atcccctgcc 240 acgtccatgg attccttgta cgccgcag tctaaagatt tcaaaacgcc gcagccgccg 300 agcaagagcc cgcaaaagtc acacagcctc ggcggaaaca gtaccagcac atctgtgacc 360 gaaagctttt ccagaccttt tattcctcc aacccccctc aacactttgt cgacgatggc 420 ttcgcaccgc caatcacctg gcctttgctt gtcgataata tgcggtacgc cgtggaagcc 480 tatcgccagg tgcttttcaa cggtgagcgt ccagagtacg taagaaggc cgaggacata 540 tctgaccatc ttcgcatgct gctggctgct ggatctgaca cgacggataa ccactctggt 600 aacccatcta tcatttccac aaacaaggcg ctatatcctt acttccggga catgatgtct 660 aagttctcga agctggttct ttcatcacat attgccgccg ctgattggcc tggtgccgac 720 tcggccaata aatgtttgca ggaggccgat ggagttatgc aaggcgtgta tggctatgtg 780 caagtggctc aacatcagcg cggcgatgcc atccatcgca tcgtgcctgg cttcgtcagc 840 ggcagctctt cgggtggtag ctggcagaac aacggtgttt ccttgaatac ttcaggcccg 900 acatcattcc tcgttccgga tggaggggac tcgcgagtag agccatcggt ctctcttgac 960 accgcctttt tggattcaat cgacatcctc agaagatctt ttgttggtag tattcggcga 1020 ctagaagaac ggctggttat aaaccggaat atcgttacag tggaggaaca tggagacatt 1080 gccgatgcga tctcagctgc tgcaatcaag gtgattgaac agttccgccc atggatctcc 1140 tcggtggagt cgatgaattt agctccgttg ggaaccagct tccagaaccc ccagctagta 1200 gacttcagct tgcaaaagca gagagtctac gatgctattg gagattttgt cctgagctgt 1260 caagcagtct ctgcccctct ggctgatgag tgggcagagc tccgtggtaa ttctctcgac 1320 gatcgtgtga atgccgtgcg aggcatcgtt agacagttgg agaactatgt ctcccagatt 1380 ggcttctcat tgtcgctgct cctcgagcaa atccccaccg aaccagcatc atctctaaga 1440 cgggatagcc gccaagaagc ggaagatgag tcgtacaaga taatgcatag ccgaggcgag 1500 tccaaggcca agattgccac agagtcaatc gggattccgt cctcctacgc tcctgaaaag 1560 gaaagtggca cagataaagt acgaagaaat atggacaagg cacaacgttt ctttggccag 1620 gcacccccaa cggctatcac ccgagagcca atccgtgagc cagtccgtga gcccgaagaa 1680 actccctggt tcttgaaaat ggcccatgaa ggcgaagtgt tctacgataa caagggagac 1740 ttgcccatcc tcaaatgtgg aacactcgcc ggattggttg aacacctcac ccgccacgat 1800 aagcttgatg catccttcaa caacacattc ctcctcacct atcgctcttt cactactgcc 1860 accgaactat ttgaattgct tgtccagcgg tttaacattc agcctccatt tggcctgaat 1920 caagatgaca tgcaaatgtg gattgaccgg aaacagaagc cgattagatt ccgtgtcgtc 1980 aacattctta agagctggtt cgatcacttc tggatggagc ccaatgatga actgcacatg 2040 gatctcctgc gacgtgtcca tacctttacc agcgactcca tcgctaccac gaagacccca 2100 ggaaccccta cattattggc cgtgatcgaa caacgacttc gaggacaaga taccactgtt 2160 aagcgccttg ttccgactca gagcaccgcc gcaccaacac caatcatccc taagaatatg 2220 aagaaactga agttcctcga cattgatcca acggagtttg ctcggcagtt gaccatcatt 2280 gagtcgcgcc tctactccaa aatccggccc actgagtgtt tgaagaac atggcagaag 2340 aaggtcggcc ctgatgagcc ggaaccatct cccaatgtca aggccttgat tcttcactcg 2400 aaccagctta ccaactgggt cgcggaaatg attctcgccc aaggcgatgt taagaagcgg 2460 gttgtagtca tcaaacactt tgtgaacgtg gctgatgtat gtgtttactc tgcttgcttg 2520 acaaatcccg gcctcactaa ctcaatcata cagaaatgtc gccatctgaa caattatct 2580 accctgactt ccatcatctc ggctcttgga actgcaccca ttcatcgtct aggtagaacg 2640 tggggccagg ttagcggacg cacgtccgca attctggaac agatgcgccg gcttatggct 2700 2760 ccatttttcg gtatgcgtca cggtcatttc aagcagattc aagttgtctt ggagtatctc 2820 acccccttga ctctgtagct aacacatctt aggtgtctat ctcacggatt tgaccttcat 2880 tgaagacggt atcccgtctc taacaccatc agaattgatc aacttcaata agcgggccaa 2940 gaccgcagaa gtcatccggg atatccaaca ataccagaac gtgccttacc ttttgcaacc 3000 cgtcggcgaa cttcaagatt acatcctcag taacctccaa ggtgctggcg atgtacatga 3060 catgtacgac cggagtctgg agatcgagcc tagggagcgc gaggacgaaa agattgcaag 3️⃣1️⃣2️⃣0️⃣ gtatgctgaa gccacaagca gagacaaggg ctccttgtta tttgcatcca ccgtcgctat 3️⃣1️⃣8️⃣0️⃣ cttgcgataa 3️⃣1️⃣9️⃣0️⃣ <210> 2️⃣0️⃣ <211> 3️⃣1️⃣9️⃣0️⃣ <212> DNA <213> Penicillium chrysogenum <400> 2️⃣0️⃣ atggagctcc ctttcgagaa cccgaccgca acaactgaac caggcccgcg agatcgaaat 6️⃣0️⃣ aatttctttg tccccgacca gacacggcca cctccagagc tgatggcgcg tggctttgag 1️⃣🔛0️⃣ cgggatgagg acgagtacga tggatctgca tcggaggcag aaggagagtc actgatgtta 1️⃣8️⃣0️⃣ ggctcgcatg actcgatttc tcgccgacgc cagtccgtga tggatggagt atcccctgcc 2️⃣4️⃣0️⃣ acgtccatgg attccttgta cgccgcagga tctaaagatt tcaaaacgcc gcagccgccg 3️⃣0️⃣0️⃣ agcaagagcc cgcaaaagtc acacagcctc ggcggaaaca gtaccagcac atctgtgacc 3️⃣6️⃣🔛0️⃣ gaaagctttt ccagaccttc tatttcctcc aacccccctc aacactttgt cgacgatggc 4️⃣2️⃣0️⃣ It should be noted that the content you provided seems to be some genetic sequence-related information, and the translation is mainly based on the literal meaning. If there are specific requirements or contexts for these sequences, further interpretation and analysis may be needed.ttcgcaccgc caatcacctg gcctttgctt gtcgataata tgcggtacgc cgtggaagcc 480 tatcgccagg tgcttttcaa cggtgagcgt ccagagtacg taagaaggc cgaggacata 540 tctgaccatc ttcgcatgct gctggctgct ggatctgaca cgacggataa ccactctggt 600 aacccatcta tcatttccac aaacaaggcg ctatatcctt acttccggga catgatgtct 660 aagttctcga agctggttct ttcatcacat attgccgccg ctgattggcc tggtgccgac 720 tcggccaata aatgtttgca ggaggccgat ggagttatgc aaggcgtgta tggctatgtg 780 caagtggctc aacatcagcg cggcgatgcc atccatcgca tcgtgcctgg cttcgtcagc 840 ggcagctctt cgggtggtag ctggcagaac aacggtgttt ccttgaatac ttcaggcccg 900 acatcattcc tcgttccgga tggaggggac tcgcgagtag agccatcggt ctctcttgac 960 accgcctttt tggattcaat cgacatcctc agaagatctt ttgttggtag tattcggcga 1020 ctagaagaac ggctggttat aaaccggaat atcgttacag tggagaaca tggagacatt 1080 gccgatgcga tctcagctgc tgcaatcaag gtgattgaac agttccgccc atggatctcc 1140 tcggtggagt cgatgaattt agctccgttg ggaaccagct tccagaaccc ccagctagta 1200 gacttcagct tgcaaaagca gagagtctac gatgctattg gagattttgt cctgagctgt 1260 caagcagtct ctgcccctct ggctgatgag tgggcagagc tccgtggtaa ttctctcgac 1320 gatcgtgtga atgccgtgcg aggcatcgtt agacagttgg agaactatgt ctcccagatt 1380 ggcttctcat tgtcgctgct cctcgagcaa atccccaccg aaccagcatc atctctaaga 1440 cgggatagcc gccaagaagc ggaagatgag tcgtacaaga taatgcatag ccgaggcgag 1500 tccaaggcca agattgccac agagtcaatc gggattccgt cctcctacgc tcctgaaaag 1560 gaaagtggca cagataaagt acgaagaaat atggacaagg cacaacgttt ctttggccag 1620 gcacccccaa cggctatcac ccgagagcca atccgtgagc cagtccgtga gcccgaagaa 1680 actccctggt tcttgaaaat ggcccatgaa ggcgaagtgt tctacgataa caagggagac 1740 ttgcccatcc tcaaatgtgg aacactcgcc ggattggttg aacacctcac ccgccacgat 1800 aagcttgatg catccttcaa caacacattc ctcctcacct atcgctcttt cactactgcc 1860 accgaactat ttgaattgct tgtccagcgg tttaacattc agcctccatt tggcctgaat 1920 1980 aacattctta agagctggtt cgatcacttc tggatggagc ccaatgatga actgcacatg 2040 gatctcctgc gacgtgtcca tacctttacc agcgactcca tcgctaccac gaagacccca 2100 ggaacccta cattattggc cgtgatcgaa caacgacttc ggagaaga taccactgtt 2160 aagcgccttg ttccgactca gagcaccgcc gcaccaacac caatcatccc taagaatatg 2220 aagaaactga agttcctcga cattgatcca acggagtttg ctcggcagtt gaccatcatt 2280 gagtcgcgcc tctactccaa aatccggccc actgagtgtt tgaagaac atggcagaag 2340 aaggtcggcc ctgatgagcc ggaaccatct cccaatgtca aggccttgat tcttcactcg 2400 aaccagctta ccaactgggt cgcggaaatg attctcgccc aaggcgatgt taagaagcgg 2460 gttgtagtca tcaaacactt tgtgaacgtg gctgatgtat gtgtttactc tgcttgcttg 2520 acaaatcccg gcctcactaa ctcaatcata cagaaatgtc gccatctgaa caattatct 2580 accctgactt ccatcatctc ggctcttgga actgcaccca ttcatcgtct aggtagaacg 2640 tggggccagg ttagcggacg cacgtccgca attctggaac agatgcgccg gcttatggct 2700 agtacgaaga actttggcga ataccgagaa accctgcatc tcgctaaccc gccctgtatt 2760 ccatttttcg gtatgcgtca cggtcatttc aagcagattc aagttgtctt ggagtatctc 2820 acccccttga ctctgtagct aacacatctt aggtgtctat ctcacggatt tgaccttcat 2880 tgaagacggt atcccgtctc taacaccatc agaattgatc aacttcaata agcgggccaa 2940 gaccgcagaa gtcatccggg atatccaaca ataccagaac gtgccttacc ttttgcaacc 3000 cgtcggcgaa cttcaagatt acatcctcag taacctccaa ggtgctggcg atgtacatga 3060 catgtacgac cggagtctgg agatcgagcc tagggagcgc gaggacgaaa agattgcaag 3120 gtatgctgaa gccacaagca gagacaaggg ctccttgtta tttgcatcca ccgtcgctat 3180<00022?9>cttgcgataa 3190 <210> 21 <211> 4030 <212> DNA <213> Magnaporthe oryzae [[ID=?0]]<400> 21 atggtaatgc ccggcgacca ttccatgcag cgggcgagcc ttcaagtggc acccctcgcc 60 atccgtaaca agggctcccg tctcggccac ggctctgaca ccgagaacga tgcttctttc 120 acgtcagtct cgagcaacaa cagtgacgcc accatcacgg actcgaggtc cgacgcaaca 180 aacctcaaca agacaacagc aaccaccacc accaccacaa caacaacgac caccacgagc 240 acaaccaaga aaccaaacgc cgcgatcgat tcgtccaacg gctcccacat gaagtcgtcg 300 tcgcgcaatg gctcgcgaga ggaaccgctg gaggcggatc cggacatggc tccgcccgtc 360 ttccacaact tcttgcgggc cttcttccac ttcaagccga gcttcctcat gacggactcg 420 actgttacac tgccgctggc cgagggcgac gtaatcctgg tgcactcgat acacaccaat 480 ggctgggcag acggcaccct gctggcaacc ggcgccagag gctggctgcc gaccaactac 540 tgcgaaccat acggacccga cgaactcacg aaccttttga acgccttgct taacttttgg 600 gatcttttgc gtagcacgtc ggtcaacgac cacgagatat tcagcaacca ggagttcatg 660 aagggcataa tagccggcgt ccgataccta ctggtaggtt tttgctcttt gtttttcttt 720 tgtcttttta tgactttgct tagccccgag ccttgcgcct ggcgtgggat gaaaaaaaag 780 accaaaaagc cctccgaggc ctgtgcgact gacgctgatt aattgggtgg cacaggaacg 840 cacaaactct cttactagag aggcccctct cattcagcgc cacgagggcc tcagacgcag 900 caggaaatcg ctattgtccg agcttagctc gctggtcaag acggccaaac gcctccagga 960 gcaccagcgt atgattcagc ccattgagga cactaacgat atcattgacg agatgatcct 1020 caaggcattc aagattgtga ccaagggtac tcgctttcta gacattctgg atgaggacag 1080 gaaatctcga gcaccatcag tcacggtcat ggcaaccgtc atggaggagg tgacgccgcc 1140 cgtcgacgga aagcctgcaa atagcgaaca ggcaaaggca ctgcgggcgt tgacggcagg 1200 tgcaggcgaa gactcgtctg ccgtggacga caccacggag cagacggtcg ttgtacgtcc 1260 tactaacagg cgcatgtcga ccatcacatc gccaatttcg gcaaccaaca cgaggagaat 1320 gtcgctgggt agcaaccccc accgggtgtc gacggcaatc tcgcaccgag tctcgcttgt 1380 cccatcacca tccaccaagg cccagaacct catctcacag caattgagcg acagccacga 1440 taccttcctg tcatacctgg gttcgttcat cggccgcctg cacctgcagt cccagtctag 1500 gccgcatttg gcgcttgccg tcaagcagtc ggcaacgtcg ggtggcgagc tgctggtggt 1560 tgtcgatgtg gtgtgcgccc acaaccgcat gagccaggat ttccttgatg cttaccgcga 1620 tgccatgttt gcacgtctcc gagaccttgt cttggcggca caggatgtcc tgaccagccg 1680 cggtcgcgag atggaggacg tcatctttcc ccaggacaac agcagactgc ttcaggcggc 1740 cacgggttgc gtgcgggcca cgggcgagtg tgttgccaag accaaatggt tcctcgaaaa 1800 gattggcgac tttgagtttg agctggaacg gggagcttcg gctctgaaca tggatcttgg 1860 ctttttggag attaaagttg ccgaggacag ggataaggac cagggcatgg acgccaccag 1920 catcgccgag tccaacaaat caggctctac cgaaacctcg acggtaacgg caactacgac 1980 acagtccgcc gcgtcgacaa ccgccacggt gcggccgacg gccctggcca ccaacaagcc 2040 gcttcctgag gtgccccaat ccacaacccc cgacgaggag gccccgcggc ctcaacgatc 2100 ccccgcttcc tcacgaccga cctcgcttgt ggaggagggc cctgccagca tggcttcctc 2160 tgtggcgtcg ctgcgtccta tgctgccgcc tctgcccagg ctttccacct cgcttatgac 2220 gcaggatgag tacagcccgt cggagcactc ggctggccac gacagcgaca actaccatgg 2280 ctcgttccgc tctgagagca tgacagcctc cagctccgga accggcagca catatatcag 2340 ccgcgactcg gagtcaagcc tggtctcaca gtcgtcaacg cgtgcgacaa cgccagacat 2400 tcccttggcg aaccaaaagt cgctctcgga tattagcaac tctggcagcg gagcttgtgt 2460 ggttgaggag gatgacgtcg agtcgaggct gctcgagagg acatatgcgc acgagctcat 2520 gttcaacaag gagggccaag ttaccggcgg ctcactcccc gctctggtcg agaggctgac 2580 cactcacgag tccacccccg acgccatgtt cgtgtcgacc ttttacttga ctttcaggct 2640 cttctgcaca cccgtaaaat tggccgagag cttgatcgac cgattcgact acgttgccga 2700 gtctgctcac atggcaggtc ccgttcgtct gcgtgtctac aacgtcttca agggctggct 2760 cgagtcccac tggagggacg agacggaccg cgaagccctg agtctcatcg agccgtttgc 2820 tactttcaaa cttggcgagg tgcttccctc ggccggcaag cgtatcctcg agcttgtcga 2880 tcgcgtctct gcgtgcggcg gtggtgcatt ggtcccacgc ctggtgtctt cgatgggcaa 2940 gaccaacaca tccatctctc aatacgttcc cgccgacact cccctgccaa acccggtatt 3000 caccaagagc cacgcgcacc tgctggccaa ctggaggaac ggcggcagct gccctagcat 3060 cctcgacctt gatgctctcg agattgcccg gcagcttacc atcaagcaga tgaacatctt 3120 ttgctcgata atgcccgagg agctcctagg ctctcagtgg atgaagaatg gaggtgccga 3180 gtcgcccaac gtcaaggcca tgtcgacctt ttccaacgac ttgtcctcgc tggtgtcgga 3240 cacaatcctg cactacaacg aggtcaagaa gcgtgcagcc gtgctcaagc agtggatcaa 3300 gattgcccac cagtgcctgg acttgaacaa ctatgacgcc ctcatggcga tcatctgcag 3360 tctcaacagc tccaccatca cgcgcctccg gcgcacatgg gaggccgtct cgcctcgtcg 3420 ccgtgagctc ctcaagcagc tccaagccat tgtcgagccg tctcagaaca acaaggtcct 3480 gcgcggtcgc ttggccggcc acgtcccgcc ctgcctgcca ttcctcggca tgttcctcac 3540 cgacctgacc tttgtcgaca ttggcaaccc ggccatcaag cagctccctg gtaacgaggg 3600 cgacggcaag gctccggcca tcaccgtcat caactttgac aagcacgccc gcacggccaa 3660 cgacggcaag gctccggcca tcaccgtcat caactttgac aagcacgccc gcacggccaa 3660 gatcatcggc gagctgcagc gcttccagat tccttaccgg ctgcaggagc ttaccgaggt 3720 gatcatcggc gagctgcagc gcttccagat tccttaccgg ctgcaggagc ttaccgaggt 3720 gcaggagtgg atccaggccc agattgcacg actccgcgag ctcgagacgc ccaacgataa 3780 gcaggagtgg atccaggccc agattgcacg actccgcgag ctcgagacgc ccaacgataa 3780 cgtccaggtc gcctactacc gcaagagtct gctgctcgag ccccgcgagg tcacggccac 3840 cgtccaggtc gcctactacc gcaagagtct gctgctcgag ccccgcgagg tcacggccac 3840 gccccagacg ctacggaact cgtccgagac gttttcctcg tcgtcggcca cgctcgcacc 39?0 gccccagacg ctacggaact cgtccgagac gttttcctcg tcgtcggcca cgctcgcacc 3900 tccaagcgcc agagactcga ccgctgccaa cggcagagca gcagagagaa ctgctcagtc 3960 tccaagcgcc agagactcga ccgctgccaa cggcagagca gcagagagaa ctgctcagtc 3960 gcagaggacg gattattttg gctggatgcg aggatctggg ggcagccaca gagatcatcc 4?20 gcagaggacg gattattttg gctggatgcg aggatctggg ggcagccaca gagatcatcc 4020 tgctgcttga 4030 tgctgcttga 4030 <210> 22<"210"> 22 <211> 3669<211> 3669 <212> DNA <212> DNA <213> 石膏样节皮菌(Arthroderma gypseum) <213> Arthroderma gypseum <400> 22 <400> 22 atggctgctc gcgatggcta ctccagccag ggcgctgctg gtgcggcgaa tgacgatggt 60 atggctgctc gcgatggcta ctccagccag ggcgctgctg gtgcggcgaa tgacgatggt 60 ctgtaccaaa atttacttcc tcttcttccg gttctaccca cgtcgtatta accgcatttc 120 ctgtaccaaa atttacttcc tcttcttccg gttctaccca cgtcgtatta accgcatttc 120 acaggctacg tatcaccaac agaggcgcct ccggctctct atgttagagc tctgtacaag 180 acaggctacg tatcaccaac agaggcgcct ccggctctct atgttagagc tctgtacaag 180 tacacctcag acgaccacac cagccttagc ttcgagcaag gcgacattat tcaggtgctg 240 aatcagctcg agaccggctg gtgggacggt gtgattggtg atgtccgtgg ctggttccca 300 agtaactact gcgctgtcgt tcctgggccc gaggctctca acgagcacgc cggtgatgcc 360 agtgccgaat ctggcgcaga cgatgactac gaggacgacg ttgacggcct tgacactacc 420 ctgagagacg acgacctgcc tattgaaagc aatggagcag acggcggcga gcccgaagag 480 gccgccttct ggatccccca ggccaccgca gacgggcgcc tgttctacta caacacattg 540 accggctaca gcacaatgga acttcccctg gagacgccga cttccgtcaa cgagtctggc 600 cctcgggacc gtacaaacgt ctacgtgccc gaacacacca ggctgccacc tgagatgatg 660 gcccgtggca tcgatcgcta cgaagatgac tatgatggct ctgcctcaga ggctgaaggt 720 gactccctct taatggcatc gcagcgccga cattcgttca tttctgatgg cgtctctcct 780 gctacatcct taggttccgt caatccttca ccaatcacca aacactatga tctcaaatca 840 gctatcctc cccatttcgt tgcaaacggt ggaaacgctg gcatggactc tatccctatc 900 atgggcactc ccatgtccac ccactcgaac gcgactgatc gatctctgcc ctttggcatc 960 tcaacctcta tccctcgcta tttcctggat gactccaccg ctcctcatcc tacctggaac 1020 tcgctcgtca gcaacatgcg agatgcaatt gaggcgtatc gacaggccat catcgaaggt 1080 cggcggtcag agtacgttcg cagggccgag gatgtgtccg atcacctgcg gatgcttctc 1140 gcggcaggct ccgatactac agataaccac tcgggcaacc cgtcaatcat ctctacaaac 1200 aaggcgctat acccgcattt ccgcgatatg atgtccaaat tctccaagct cgtcctatcc 1260 tcacatattg ccgcggctga ctggccggga ccagactctg cgaccaaatg tctccatgaa 1320 gccgagggcg ttctacaggg cgtttacggc tacgtcgaag tggccaagca gcagcgagga 1380 gacgatatcc gccgtctgac acctggcttt gtcgccggca gcacttctgg cggtcactgg 1440 cagaacaaca acctcgctcg aagggatcca acgtctttcc tcgagcatga ctctgagtct 1500 caccgcactc cgtcggtctc gcttgactca aagcttctag agcgaatcga agagcttcgc 1560 aagatgctag ctgtcagctc ccgcaggcta gaagagcagc tctcatcctt caagggtaaa 1620 attgttacgc caaaaagcca tgccgagatt ggcgacgctg tatgtgaagc tggcgtgccg 1680 atagtcgaaa actttcgccc gtgggtggcg ctcatcgagt ctatcgactt gtcacacttt 1740 ggctctgatc tccagaaccc gcaattagcg gacttcagcg ttcagaagca gcgcgtgtac 1800 gacagcatct cggacctcgt tatgagctgc cagcacatct ctgctccgct aggcgacgag 1860 tgggccgaga tcaggggcga ttcgcttgag actcgtctaa ataatacccg catgatgtca 1920 aggcagctca ctaattgcgt tcaacagatt ggattctcgt tgaccttact attggaacaa 1980 gctccacaac aacaaataca aaatggagat ggatataaca aatctgctcc caaggtacgc 2040 aagagtccgc catcatctat tggcatacct tccagctatg gcgtgggcga tgaccatgat 2100 aagccaccac ggtctctgga taaggcgcag cggttctttg gccaacccgt gccgagggag 2160 ccgacttctg ccagagaacc cgaggaaaca ccgtggttcc tgaaactcga ccatgaggcc 2220 gaggtgtttt acgacgtcaa gggtgacgtg cagcagctca agtgcggtac gctggcagga 2280 ctagttgaac agcttacccg ccatgacaag cttgatccct ccttcaagga taccttcctt 2340 ctcacatacc ggtccttcac cacggcttcg gagctttttg agatggtggt acatcgcttc 2400 acactccagc ctccctacgg cctgaccaaa gcagagctac aaatctggac cgaacaaaag 2460 caaataccca tccggatccg tgtcgtcaac atcctcaaga gttggttcga gaacttctgg 2520 atggaaccaa atgatgaggc aaacacacat ttacttggcc gtatacactc cttcgttacc 2580 gaggcagttg catcgactaa gacgcctggc gcgcaacaac tagtcagttt gatagagcaa 2640 cgcctacgtg gagaagaaac taccgccaaa cgcctggtac ccaccattag ctccaatgca 2700 cccactccca tcacacccaa gaacatgagg aggatcaagt tcttggatat cgacccaacg 2760 gagtttgcgc gccagttgac tatcatcgag tcgcggctgt atgctaagat taagcctacg 2820 gagtgtttga ataagacctg gcagaaaaag gctggaccag gcgaggccga gccggcgccg 2880 aacgtcaagg ctcttattct acattctaac cagcttacca actgggtggc tgagatgatt 2940 ttgacccagt cggacgtcag gagacgagtc gtcgttatca aacactttgt ctccgttgct 3000 gatgtaagtt gatttatctt cttaccccct taacacataa aaattatgct aacaaatttg 3060 atagaaatgc cgacaactta acaattattc tactttgaca tctattatct ctgcgcttgg 3120 caccgcgcca atccatcgac tggctcgtac atgggcgcaa gtcagccaga gaaccgctgg 3180 aaccctcgag atgatccgca aactcatggc tagcacaaag aactttggcg aataccgtga 3240 aacccttcac ctagccaatc ccccttgcat tcctttcttc ggtaacgaca atttcctatt 3300 tttttttatc ggcgcagagc cactaacaca cgcacaggtg tctacctaac ggatcttacc 3360 ttcatcgaag acggcattcc ctcactcact caatccgatc taatcaactt caacaaacgc 3420 accaagaccg cggaggtgat ccgcgatatc cagcagtacc agaatgcgcc ttaccagctc 3480 attcccgtgc cggagctgca ggagtacgtg ctgaataata tgcaggctgc aggcgatgtg 3540 cacgacatgt acgaccgcag tcttgaaatc gaaccccgag aaagggaaga cgagaaaatc 3600 gcaaggtatg gtaaacacta ctacgaccca tcggtcgttg cactctccct gacggttggc 3660 attack 3669 <210> 23 <211> 1571 <212> DNA <213> Endocarpon small <400> 23 atggaggaga atgacggaga gagcaggaag cttctcgaca ggatctactc atttgctaaa 60 gactcaattg ccacgaccaa gacaccaggc tcaggacctt tgatggcggt ggttgagcag 120 aggctgaagg gtcaggacac ttctgctaaa agacttgtgc taacattgac gaattctgct 180 cccgccccga tcttgccaaa aaatatgaag aagctcaagt tcctcgacat agacgcaaca 240 gaattcgcac gacagcttac cattatcgag tctaagctct atggaaagat caaaccaact 300 gaatgtttgg gcaagacgtg gcagaaaaag gttggtcctg aggagcccga cccagcaccc 360 aatgtgaagt ccttgatcct ccattccaac cagctcacga actgggttgc ggagatgata 420 ctatcacagt ccgaggttaa gaagcgagta ctcgtcatca agcactttgt ttcgattgca 480 gatgtgagtc cagccgtaaa cgccaattcg caaagactga cccatacaga aatgccgcaa 540 catgaataat ttctcaaccc ttacctctat tgtttctgct ctgggaactg ctccaataca 600 ccggcttaat cgaacatgga cccaagtcag cccaaagacc atgacttctc tgagtgtgat 660 gcgacagctt atggccagca ccaagaactt tggtgaatat cgggagaggc tacgccgggc 720 aaacccgcca tgcataccct tcctaggtgt ttatcttacg gatctgacat tcattgaaga 780 tggaatcgcg tcgatcgtca agaactccaa cctcattaat tttgccaagc ggaccaagac 840 ggccgaggtc attcgtgaca tccagcagta ccagaacgta ccgtactcgc tcaaccctgt 900 tcctgatctt caggagtata tactcagcaa catgagagaa gctggcgatg tacatgagat 960 gtatgataag agcttgcaaa tcgaaccaag ggagcgagag gatgagaaga tcgcaaggtg 1020 agtgtgtaca aggaaatctt cacaccccca acgatgcaga tgggtctgac tcacgtctct 1080 cctcgattat agattgctgt ctgagtctgg tttcctttga tccgtgagca ggactcgcga 1140 ttcgctggtt tctcaacata ccttttgagt tgaatagccg cggggtttgc aggtgccgaa 1200 tctcccttgt ccctaactat gatgtcaatt ctacataagt actggggatg ctacacaagg 1260 ccggtcctac gtaacaagcc attgcatgga tacttggatg gttgggggtt tttctggtag 1320 atatctgatt caggcttggt ggcatggtat tggacgtctg acatgaaatt gcacgagcaa 1380 acgagtcgat gagacactta tctggacatg gtcaaacatc aacgaagctc atggatagga 1440 gcgatactaa ttcaggctga tctcggagct tgtgatgggg aatctgcgat atctagtgct 1500 tttgaatata catttttgtt gctaatgcag aatgagtagc tgcatttttg cgcagtcgat 1560 cggttttcta g 1571 <210> 24 <211> 3012 <212> DNA <213> Fistulina hepatica <400> 24 atgtacgatc tcgtacacga aattgttgcc attgtgtgca agctacttac catcgcggac 60 gctgtaatgc tgcacccaga catcccgcca aacaaagtca agaacctcag ccactcaaag 120 aatgcgctgt acgattcgac gacggcattg atggaatgtg tccagacgtt gacgcaacca 180 ctcgcgccga cggtgacgga agaggatgag aagagcgcgt tactggtcac cgccacatcc 240 gcagtaaagg ttggtgcaga ctgcgtggcc gcgatcaaga tgtgcctgtc gcgttcggtg 300 ggcgaacggc cattcgtact gcagctgccc gacaagaatc accctcctgt ggcagtgcct 360 ttacaacgac caaagctcgg caaggcaaca agccttggct cgttgaatac gtccactaac 420 gtgcctgaag accacgatga caccatccgg ccgcccgtac caccacttcc acaacagtgt 480 tcgcgtgatc tttcttctgg atcggaaaag agcgacgcgt cggcacagag ttcgacgagc 540 tcacgagaca caggctttac gtcgttggac gccttgaagc tggtttcacc gaaagagaag 600 cctctgcccg ctctccttaa gcttacaaaa gccgccgtcg aaaaagatct cccctcgccc 660 acatctcttg ctcccactga agctgcaagt acatgggaag gtgctccatc acattacttg 720 cactcgctcg gtaaatcatc aaccacttca tcaaatgcgg cgctcagctt gcaatctatc 780 ccacaatcct tgccatcact gccggcgctt gtcaatgcac atgactattc cgacgacgag 840 gtggcatgca atagcgaggg acacatcgtc ggtgcaacga tgagcgtgtt agtggcacgg 900 atgacgccac atgacaatct cgtcgatgcc gccttcgccg cagtcttctt catgacgttc 960 cgcttgtttt cgtcgcccga agagctcgtt gacactctga tagctcggta caacatccag 1020 ccgcctgaat tcctgagtca ggcggacaag gagttgtgga tgcatcaaaa gggcatgccc 1080 attcgacttc gtgcggcgaa ccttgtaaag agctgggttg aaagttattg gcgccctggt 1140 gttgacgatg cagtgtcgca gaccatctac gaatttgcag agacttgtgt gcataagacc 1200 1260 gcggtaatca cgccgaaagg cgatcgcaca cgcgaccccg gcatgtcaat taaccctccg 1320 attgtgaatt cgccgtccga aattccgcga ccgatcgtgt ccaaaccatt gtttgcggcg 1380 ttgaggaatc ggaatttctc gtcgatcagt gtgcttgact tcgatgcatt ggaattggcc 1440 cgccaactca cgcttatgga atgcacgctc tattgtgcaa tacggccgga ggaagtgctc 1500 gaacctggcc agccgggaaa gccgaacatg aatgtcaagg cgatgagcac gctgagcact 1560 1620 acgctggtta agttttcgt caaggtcgca gatgtacgtg ttcgttctat gtcaaccgtc 1680 tgtaaagatt ttgaactcct tgtcagagat gtgtctcact gaacaatttc agtacctcgt 1740 ggtcccttct agcggctctc gattcttcta ccatttcacg gcttcatcag acctggaccg 1800 tagtaccaa atttgtctct tgttctctcg ttaaaatatg attcttgctt ccagggcctg 1860 1920 taccgcgagt acagaaataa attgcggaac accgcgccgc cagctgttcc gttcttgggt 1980 ttgtgcactg tttctgctcc ttttcgcgat gagacggatt aacaatggtt ttcaggcctc 2040 tacctgacgg atgtgacatt ttgtcgtgag ggcaatccct ccactaaacc gtcgcctcta 2100 gatcccaata agcagctcat caacttcaat aaattcata agttggcgcg aatcgtgcaa 2160 ggtattttca catgcgcgtg ccgcactatg tcataatgct tgaacttcgg tttgcagata 2220 tgcagcgttt ccaagtgcct tacaatttca aggctatacc tgttatccag gaatatttga 2280 acgtcgcgt cgagacttcg aagaagaaca gcgatcttca agacttgtac cgtcgtaggc 2340 aagtacacag tatgaattct tctgtggcca taatcgctga tgatgtgctc agtcttatga 2400 tcgagccaaa gcggccggtc gacacgccac ccgcgagcgc gagcgatacg cgattgttc 2460 attgggcttc gaagtcccaa acaccatctc agactgtagc tgctccattc tagatgtccc 2520 tgcttatttg tactatttc actatgttta tacatgggtt ccgtgcacac ggttgctcat 2580 atttgtgtct ctttcttttt tttggcggac ccgtcgcgt gttttcgatg actttctctc 2640 tccgtcctcg gtgttcacat attaactcga ctctctgtct ctctgtctct ttcgctcttg 2700 ttattcattt cgctttgttg ttaggttata tatacattat attgtcaagc catctgtacg 2760 ttcacccatc cactgcattg gtttagcctc actacttttg tctgcttgaa tacactggtt 2820 cgtccatcac ccgtgtcgtc tctggccagt agggaaggga gcacgcatcg tttcactata 2880 cataggtcag tcaggtcgag tctttcttct tctgggttca tccctcaggt catgaggtgc 2940 gtcatgcagc gacttgttat tctcaatctg attatggtca gttatatcag cggtgaaact 3000 cttcacgaat ag 3012 <210> 25 <211> 1658 <212> DNA <213> Aureobasidium pullulans <400> 25 atggtgacaa ccccttcagc ctccacaaat ccgcccctcg acattgacac gaatttggac 60 gaatcacgcg acgacattac agactcatct cgttctgagc acgcctcctc atctgaggat 120 ggcggtcttg atgtcgatag tcaatcagag gccccatccg atgagcgggg atactccttt 180 gacaacctcg tcgaccgtct cctggggcta ccgcgatcaa aagccgacac acgattcggc 240 tctgtctttc ttgccctcta tcggaaattc gctgcacccg gacagttgct ggaagctata 300 gttcatcgct tcgaagcctt ggaaaaagaa aactgtcctt tcatgacaaa gactgtctca 360 caattacgct acttatctgt cattgagcaa tggattggaa cataccctgg agactttgca 420 cacacaaaaa cccgccgtcg catgcgcatc ttcgtcgcca agctgtccaa cacacgcatc 480 ttctctgctg ccgctcgtga gatgagctgt gacttggacg ttgtgacaga agacgatgat 540 acaaattggg cttgttgtga catggatcgt gaaaaacgcg gtctcctgag ccccgatctc 600 ggctggtcat cccgtgtgag cacactcctg gacgatcccg aatttgactt tagcgacaac 660 ctgggaagcc tgtctctcga tggcggccag ggtagaaatg cagcccattc cttacatacc 720 gactttggca tgctgcagac cgtggacgca gcacgtagac aaggcccatc tctggttccc 780 gtccccaaga ttccaatcag caagatgcat tggcacatgc ttatggaaac accaacggat 840 cacatcgctt gcgaactgac acgcattgac tggatcatgt tcagtgcagt acgtccgcgt 900 gatttggtgc ggcacgtttc cttatcgcaa actcagaagg cacagtgcaa atccatagta 960 catgtcagcc gcatgatcga ccatttcaat cacattcgag actgggtggc caacttcatc 1020 ttgcttagag agaaggcgaa acaccgtgta ctgatgttgg aaaagctcat gcatgtcgcc 1080 cgtaagctgc gagagatgaa caactacaac tcgctggggg cgttccttgc cggtatcagc 1140 agtgcagccg tacaccgact tgccgctact cgagaactgg tttcacccga gaccggcaag 1200 gattggatga agctggagat attgatgtct cccactcgct cttattctgc ttatcgcctg 1260 gcttgggaga actcaagcgg agagagaatc cctttcctcc ctctacccat ccgagatctt 1320 gtggccgccg aggaaggcaa caagaccttt gtcggcgacg aagtgaatgg cagaatcaac 1380 tggcgcaagt tcgaagtcat gggagaaacg gtcgtcggga ttcaaaaggc gcaaggtctg 1440 ccttatagga actccatgct cggccctagg aatgatgagt tgagagcatt gatcctgaac 1500 agtaacatga tcagagatga cgaggtaagt tcttcagaga tcaaggtatt tgcagacaca 1560 cgactaactg aatcttccag gctctttatg accgtagctg ttctctcgaa tctaccaacg 1620 acagaagggg gctgcgagat atcttcagac gcgcatag 1658 <210> 26 <211> 2510 <212> DNA <213> Acremonium furcatum <400> 26 atgtctgtga tgctgcaagc tccttcccga gcctccactg catcctcctc ctccatccag 60 cccctctccc gacagaacac catgtcttcc tacgatggct cgcggtccgc ccgccagtcg 120 aagcggtact ccatgtccgc gctgtacatg tccatgtcag ccaacgacgg agagctcgag 180 atcgaagacg atctggccaa aggtaggcta catgcattcc cagttgcact acgactggat 240 tctctcgcta acacgctcaa aacacagccc agaaaatcct gcgagaactc aagtccaaga 300 tctcctccca gtccaagaag aacttcgtcc tcgagaagga tgttcgatat ctcgactctc 360 gaatcgccct tctcatccag aaccgcatgg ctctggagga acagaacgaa gtcgccagcc 420 atcttgaaga cgccacagac atgcaagagg gagccttccc gaacgacgac aagacccaga 480 aatatggcaa cctcatgttt ttgctgcagt ccgagccgag gcacatcgcc catctgtgcc 540 gtcttgtgtc catggctgag atcgactcgc tgctccagac cgtcatgttc acgatctatg 600 gaaatcagta cgagagccgc gaagagcacc tgcttcttac catgttccag gtccgcctgc 660 ctacctgcac tatatcagat cattgctaac aggacttcc agtctgttct gacctatcag 720 ttcgacaaca cccctgagta ctcctcgctc ctgcgcgcaa atacccccgt ctctcgcatg 780 atgacgacat acacgaggag aggccctgga cagagtttcc tcaagtctgt tctggccgat 840 aggatcaaca gcctgatcga actgaaggac ctcgaccttg aaatcaaccc cttgaaggtg 900 tatgagcgca tgatcgagca gatcgaagag gacacaggaa gcctacccgc atccctgcca 960 aagggaatca ctgctgagca ggcggcggaa aaccctcaag tccaggccat catcgaaccc 1020 cgtctgacga tgctgacgga tctcgccaat ggcttcttgt cgaccatcat cgaggggctc 1080 gatgaagctc cttatgggat ccgttggatt tgcaagcaga tccgcagctt gaccaagcgc 1140 aagtatcctg atgctaatga tcaggttgtt tgcaccctca tcggcggttt cttcttcctg 1200 cgcttcatca accctgccat tgtcacgccc aagtcctaca tgctcatcga aggccagcct 1260 gccgagcgac ccaggcgcac cttgacctac attgccaaga tgctccagaa cctggccaac 1320 aagccctcgt atgccaagga gccgtacatg gcgaagcttc agcccttcat tcagcacaac 1380 aaggaccggg tcaacaagtt catgctcgac ctctgcgagg tgcaagattt ctacgaaagc 1440 ctcgagatgg acaactacgt ggccctgtcc aagaaggacc tggagctgtc catcacactg 1500 aacgaaatct acgccatgca ctcactgatc gagaagcatc atgatgagct ctgcaaggac 1560 gccaattctc acctggcaat catcatgtct gaactgtctt cggccccggc ccaagtccca 1620 cgcaaggaga acagggtcgt caacttgccc ctattcagtc gctgggagac agccatggat 1680 gacctcactg ccgcacttga cattacgcaa gaggaggtgt tctttatgga agccaagtcc 1740 atcttcgtac agatcatgcg gtccatcccg tccaacagca gcgtttctcg acgccccctg 1800 cgcctcgaga ggatcgctga cgcagcagcc accagccgaa acgatgcggt tatggtccgc 1860 aagggcattc gagccatgga gctgctttca cagcttcagg agctgagggt cattgataag 1920 agcgaccatt tcagtctgct ccgcgatgag gtggagcaag agctgcagca cctggggtcg 1980 ctcaaggaag ccgtcatccg tgagacatcg aagctcgagg aggttttcaa gaccattcgc 2040 gaccataaca cgtacctggt cggccagctc gagacgtaca aaagctatct tcacaacgtc 2100 cgctcgcaga gcgaaggaac gaagaggaag cagcagaagc agcaggtcct tggtccttac 2160 aagttcaccc atcagcagct tgaaaaggag ggcgtcatcc agaagagcaa tgtccccgac 2220 aaccgacggg cgaacattta cttcaacttc acgagccctt tgccgggcac tttcgtcatt 2280 tcccttcact acaagggtga gtattcctca ttgccgcgcc ctcattgatt catgcttaca 2340 actgcgtagg acgcaaccga ggattgctgg agcttgatct caagctggac gaccttctgg 2400 agatgcagaa ggacgggcaa gacgacctcg accttgagta cgtgcagttc aatgtgccca 2460 aggtcctggc gctcttgaac aagcgcttcg cgaggaagaa ggggtggtaa 2510 <210> 27 <211> 5325 <212> DNA <213> Purpureocillium lilacinum <400> 27 atggtcaggg actcagggtc acgtccagga cgggaagacg tctggctgtt ggctgtcttt 60 ctcgtcgccc caagaaaggg tagggagtct tggctgctcc gcagagactt gtatttgcat 120 ggccgccatc ccatcggcgc cgtgtgtggc atggcacggg tgctgctgcg tgcttcgttc 180 gaatggggtc gatgatgccc aattcctcgg ggatcctcgt cgtggtatat taccttacct 240 ttgcatctac ctaatcgacg gcggcggcgc caggcaaatg gaagctggcc gccggcaagg 360. tcaccgccca cggcacccct aatcgtcaat tacgacagac gcccgacca cgacacccac ctctgtggcc ggggctgcac acccgagtgt aggttcaagg ttgatcttgc gtacctcacg 420 gagtacaggc gaggtactgc agcgcgtgcc ttccttgccc gtggccggac cgcgccaaca 480 ccgacgagta cctgccccgt actccgtgga gtgctccgcg ccgtgccttg cttgcctcga 540 gtacaaagcc agggtacttc cgtacattgc ccctcgacca ggccaccacc cagacccgca aaaaagaca cgacaagaca acagcgtcac ccttcccctt gtccgtgctc cagcaagccc 660 gcttgtcctc gaccggtgct gtctcgccgg cgcgcgcgcc tcgataccat accctctctt 720 ttctacaatt gcactcactt tccccatccg tcggggcttt gcgtttttcg gcccagaata 780 gccgtcgacg gtactgtgcg ccattgccag agagcttgat ctgttgttgc ggcgccaaac 840 accgaacgct ctttatactt tcctgtgccc cttgacctga acgcagtcgc agcgttctcg 900 cccttcgacc ccggcatcga gcttggaaac cgacaagcag ctcccccaca atcctggcct 960 gccgctttct ccgcatcgcc tcgtccgcct ttcgtaacga ctgcttcgtc gcccgcgttc 1020 gactccgcct ggcctcgaac ctcgaggcgc ctgcgtgtaa gtcagtcacc ttgcgtgctt 1080 tgatcctgcg gctcagctgc agccccccca ccagcagttt gcccttgctt caggtcctgc 1140 tgtcatgtcg tccacgctct cagcagttct cttgtacttc ttagttcacc ctgcattcct 1200 cgccacgccc gccccgtccc cctctggcct gcattccaga cacgggtcat tggctcttgc 1260 cacaacatcc aggtcgcgct ccggccttgc taacatccaa tctcgcctcc agacaaacga 1320 gcgtcgcgta catctcacaa ctgctggttg cgcccacctg cgttgacctg cgcctcggtg 1380 gtcgcggccg tcgttgtcac atccctgggt cctcgccagc accagcatac cccccctcaa 1440 aaagaacgaa ctgtcacgga acccccccct ggggctcctc cactgcgctc tttggataac 1500 caaagcactt actttggaac cagggcgcgg ctggccgtcg ctctgggacg ggccgtcgac 1560 gctagaccgc gaggcctcga ccagatgatc ttgacactcg tctctacctt ctcacaggcg 1620 caatgctgag cgaccaaccg tcgcgaactg ctttacacgt ggccccgctg gagatacccg 1680 cgtcgcagcc acaggacggt gccaatggct tgtgccatca agaacatcag acgaatctct 1740 actcacagac ccctatgact ccgccggaaa cacctaacgg ctcccaggag gacctgacgc 1800 cggagcctct cgccccgccc gtctttcaca atttccttag ggccttctac ccgtttcacc 1860 ccggctacgc cttgtccgac tcgagcgtca cgctgccact ggacgaaggc gatgtcgtac 1920 ttatacactc tgtacacacc aatggttggg cggacggtac tcttttggca accggcgcca 1980 ggggctggct gccaactaac tactgcgatg cctacgagcc cgaggatatg cggagccttc 2040 tgaaggcgct tctcaacttt tgggacctcc tacgtagcgc atcagtaaac aatgagatct 2100 tcaggaacca agaatttatg aagggcgtca tagctggagt tcggtttctc ttggtaggct 2160 tgcgctatct tctcccccac aagggctcat tgtctttgct aatgacaatg tgctcaggaa 2220 cgcacaaact gcttaaaccg agaatcgacc atcattcagc gcagtgacag tttaaggaga 2280 tgtcgcaaat cattgctctc agaactctca tcattggtca agacagcgaa gaaaacacag 2340 gagtgccaaa aggggacact ccacccaccg caggatgtca acgacatcat tgacgagatg 2400 atactcaagg cattcaagat tgtcaccaaa ggcgtccggt ttctcgatgt tctcgaggac 2460 gaacggaggg ctcgcgcacc agcagctgtc actgtcatgg ccactgtcgc cgaggaatca 2520 tacattccac ctacaccccc tgcggagcgc ttggctttcg acgatcaaag tttgaacaat 2580 ggcagcgaga cggcttcccg cggaacggcc gacagtgtgg ttggcagcag cgccacttcg 2640 gaacccagcg ttgcatcact caatccatgg aacaggcgca tgtcgtctct gggtggatct 2700 caaggcacgg cggcccagaa tcgatggtct caaggaagtc tccaacaagt caaccgtttg 2760 tccacaagta tggcgcacag agtctcgctg gccggcccat ccccgctgtc gaggcctcaa 2820 catttggtat cggagcgcct caaccgcagc catgacaaat tcctctcgca cctcggatct 2880 ttcattgggc gactgcactt gcagtcacac tcgcaaccgg aactggcact cgcgatcaag 2940 caatctgcca catcgggcgg tgaattactg gcagtcatcg acggtgtctg cgagtacaac 3000 agctctagtg ccgcggcgct cgctattgtc cgagatgcca tgtttgagcg cattcagatc 3060 ttggtccact ctgccagaga tattttggcc aatgccgcta ctgaaggggc cgacataatc 3120 ctgccacaag acaatggggt tttgctcatg gcagccactg gttgcgtgaa agccgcagga 3180 gaatgcgtcg ccaaggccaa ggccgccatt gagagggcgg gggacttcga gttcgagctg 3240 gaagagaaca cgctcgggat agacctgagc atcttggaca ttgtcgtgga cgagcgggcg 3300 agaacgccct cggtaacgga tcgatcggac cctatgagca gcgttgcaga atcgttccag 3360 acccccgaat cgactgttca gcctcaaaag cggccgatcg cacccgccgt cgacaagccg 3420 cttccccaag tacccagaat caccatcccc gcagactcgc acagtcgtca aagcaactcc 3480 ccagtgtcct ctcgaccccc gtccctcaac gaggacaatg cttctagcgt cgcgtcgtct 3540 gtttcgtcta ttcgccctgt tctcccgccc ctccctgagg tttccacaac accgcagcct 3600 ctggatcgcg atggttccga cacgacaaca atcgagtcgg acgcccatac ctcgaggttc 3660 gacgccttgg cggcgtccag cgcgggcagc agtaccactt acctcagccg ggactctgaa 3720 acgagcatga tgtcgcagac gtcgacgcga gcgacgacgc cggatcacac cttggtgcct 3780 cgcagccagc cctcgatgtc ggagctgagt acggccggca gcttctccca ggccgaagag 3840 gcggatgacg tcgaaaag acttatggag aggacgtacg ctcacgagct catgttcaat 3900 aaggaaggtc aagtcactgg cggatcgctc caggctctgg tcgaacgtct caccacgcac 3960 gagtcgactc cggacgcggc ttttgtctcg actttctacc tcacattccg actgttttgc 4020 tcaccggtca ggttgacgga agcgctcatc gaacgtttcg attacgttgg agaatcgcct 4080 cacatgtcgg gccccgtgcg tttgaggta tacaatgctt tcaaaggctg gctgggaatcc 4140 4200 aagctggctt cggttctgcc atcagcggga cgccgcctgt ccgagctggc caagcgtgtc 4260 tccggagaag ggtctctggt gccgcggctt gtctcgtcaa tgggaaagac gaccacgtcc 4320 attgctcaat ttgtcccggc tgatagcccc gtgccgcagc ctatcatttc aaaaagccag 4380 cagaatttgc ttacgtcctt caaaattggc agtgggatgc caaccatcct cgactttgac 4440 cctctcgagc tggcacgaca gatcactctg aggcagatgg gcattttctg ctccatccaa 4500 ccggaagagc tgcttgcatc gcagtggatg aagaacggtg gtgtagatgc accacacgtc 4560 aaggctatgt cagcgctgtc gacggacttg tcgaatctgg tggcagagac catccttcag 4620 tacaccgaga tcaagaagcg agccgctgcc atcaagcagt ggattaagat cgcccataaa 4680 tgccacgaac tgcacaacta cgacgggctc atggccataa tttgcagcct gaacagcagc 4740 acgatcagcc gccttcgcaa aacctgggac gcgatttctg caaagcgaaa ggaggtgtta 4800 cgcgcactgc aggagatcgt ggaaccatct cagaacaaca aagttctgcg gacgcgacta 4860 cacgatcacg tacctccttg cctgcccttc ctcggcatgt acctcacgga tctcaccttt 4920 gtggacattg gcaaccccgc gacgaagcag atgtccctgg gcacccagtc ggaagaggac 4980 agcacgggcg gcttgactgt tgtcaacttt gacaagcaca gtcgcactgc caaaatcatt 5040 ggcgagcttc aacgtttcca aatcccgtat cggctggtgg aagtgtctga catgcaggac 5100 tggctggccg ctcaggtgcg gcgtgtgcgc gaaggtgacc aaggcaacgt ccaggtcact 5160 tactatcgca agagcctgct cctggaaccc cgcgagagcg cttcgcgacg cgaagccgag 5220 ccgcctacac ctggttcaac tggtgttggc agctctcgca ccgacttgtt tggctggatg 5280 tcccgcgacc gaagcggaca aaccgctaca ccagcacccg tagag 5325 <210> 28 <211> 1585 <212> DNA <213> Fusarium sp. <400> 28 atgcacaagg gcaccggtgc tgtgcaaaat tgcctcattg cagctgaaag gcagtctaca 60 aagcgtttga cgaccatcga tgaaactagt gacgccccgtc gtccaagctt gagggacgat 120 tcactatccc atccccgact tcatctgaac gagaacgctg aggtgactgg aggcaccctt 180 ccgggccttg tgggccatct cacctctcga caatccgcat ccgacatcat gttcccgtac 240 gctttctttc ttacattccg acaattctgc aagccacgag agctcgcaga acagcttgtc 300 gagagattcg atagtgccaa cgactcttcc tttgccgaag atacgcagtt gagggtctgc 360 gacggtttca agctttggct cgaaatgtac tggcgagtgg agactgacca agaggctcta 420 ccggttatca agccctttat cacatcgagc ttgtcttcta tcatcccagc cgcgagtagg 480 aagctagctc ggttgatcga gcaccttcca gctcgagagc cttgtttgtt gcctctagca 540 gatcatgata aactcataac aactgttttt gactcaccta gagtcaggag acatcgagct 600 cagcctaatg attcagcgac gcatcaatgg ggctttttga ggacgctgag gaacagtaaa 660 agctcgtcga ctttcctcag ctttggctgt ataggtttg cccgacagtt gagcattgag 720 cagacgactc tattctgccg cattcctccc caagagttcc tgggttgtgc gtgggtatgc 780 aaaactggca acatggcgcc tatatcaga gcaatggtgt ctttcactag tcagctttca 840 aaccttgtgg tggaaaccat tctcgaccat caaacggctc gcaagcgggc tgctgccatt 900 aaccactggg tcaacatcgc acaggagtgc tcaaactttc gcaactacga tggccttgtg 960 gccctcctct caggcttggg ccacagtgcc attctccggc tacgtcagac atggaatctg 1020 gtatcaccca agtacataaa caccttacaa ttccttaaga cgcgtatgga ccgctccgat 1080 aatcacaaat cacttcgcgc attattggaa acccatgaca acccatgtct gccctttctt 1140 ggcatgtatc taacagagct ggctttgtg gagatgggtc agtcttggat cgatccgcaa 1200 aatcctcacg acgaaacaac atctgagcag ccctttattg actttgctaa atatgctcgg 1260 acggctaaga ttgtaaggca gcttcagcgt ttccagacgc catccaagtt aacagctcac 1320 cctcgtctac aaaattggtt gtcttttaaa atctcagaac ttgattgcaa taatgaccct 1380 aaactggatg ttagcttttt tgatagaagt gtgtcattgg agccgtacag gataaaaaag 1440 tagttgtggc ccgctctctc taaataaaat aatcgtaatg tctaaagcag tgtttgttta 1500 atccgtgcca gtatatgacc cttatttgcg gattccttgc gctcaaatag ccgtaaacat 1560 gcgttctagt cccccaagct aggcg 1585 <210> 29 <211> 4366 <212> DNA <213> Corynespora cassiicola <400> 29 atggaccaaa caaggcagag tcggcgcaac aggaggaga ttggcgctcc ggaagcaact 60 cagcctctgc cacgcgacca acggcgagac gatcgcggct cgtacaattc ggcgacgatc 120 cgtaccgtca cccccgattc catcccagag gacagcgttg ccagccagac gtacactacc 180 tcccgccca tctcgccgcg ctccaacagt acctcccagt ccgcccgcaa cagcaacccg 240 cgcaccttg ttgcccgcgc aaactcgaac gacgcagaat actcgctcag ggctgcccgt 300 gagactcaga ctcgtccgcg gaccaggacc ctggaggagc gctcaacgcg cgatcgctct 360 ccgccgaatc tattcgtaac cagccgccac cgcatcggct cggtgcacag cgctgcgccc 420 tccaatttcc agagcctcga ggagtcggtc gtcacctcca tcggccatcc ctctaccata 480 tccgccggcc ctaccgcccc tccgcctcga acctcgagca gcaaccgaag ccgcatcata 540 aaacagcagg cccagccgca accgccccag cgtgccacct cgcccacctc ctcgaccgcc 600 gtctctccca atgcccagca gccagactcg tgggtgtcgc ctgtgccggc ttcagacgcc 660 cgcaaggtcc tgaagctcat gcgcgccaca tgcggcaaga tgcagggcat gctggccttc 720 cgcagaggag agtcgaatcc gtggtcgctc tcctactgct acatcaatga ggaggccggg 780 agcttggtgt acgagccaaa gagtgacaca tcgtaccaca ggacgctggt gccggacctg 840 cgcggctgtc gtgtcaagac tgcctacgat gccgagtcgt acaccgccta cattcacgtt 900 ctgggccaca actccaagct cgaggtgttt ttgcgcccgc ccacccaaga agaatttgac 960 tcttggtttg ccgcacttct ctgctggggc cccatccgcc ccaagggcat ccacaacaag 1020 atggcgaagc cccagacgcc aatggtgacg gaacggcgac tcgccgatag caggagacac 1080 tccgaggtgt ctctgctcaa agaggcgccc atcatcaaag tcggaaagat gatctactgg 1140 gataccagcg tgacatatag caacacagga acccccaagg ccactggagt cgccaggccc 1200 caagcctacc ggatacaaag ccatggctcc cgcaggtgga gaagagtatc gtgcaccttg 1260 cgagagaacg gagagctcaa gctatactcc gacactgatg tcactctagt ctcggtcgtt 1320 cagctttccc agctgtcgcg gtgcgccgtc cagcgcctgg acccatctgt tctggataac 1380 gaattctgca tcgctatcta cccgcaatac acctcgacgt cgacgtcatt atcactacta 1440 cgccccattt tcctatcgct ggaatcacga gttctttacg aagtgtggat tgttctgtta 1500 cgagcattta ccattccgca actctacggc ccgaaacagc cgaccctaaa cgacgaaggc 1560 gccctctcgc cttcgttcgg tacacaagac atgttccgca tggagcgttc gctactggtc 1620 agagtcatcg aggcaaggtt gataccaccg ataagcccca aggtctcaga aaacagcggg 1680 cggccgacgt cctcggcgaa tatgaacgcc ggaggttact acgtcgaagt cttgttggat 1740 ggagaagcgc gagcccggac catggccaag aatgagggca acaatccatt ttggcgggag 1800 gaatttgagt ttcttgacct acctgcagtc ctctcaacag cttctttgct gttgaagaag 1860 cgacctccga gccaagcccg caacgacaag aacttttacg agacacagct caactccgaa 1920 tccttcaact cggacggtgc aggtggctat gccggcatct ctttcgatca gacatgcggc 1980 aagacagaca tctatcttga cgacctgggt ccgaatcagg aagttgagaa gtggtggccg 2040 cttgtcaaca tgtacggcaa cagtgtcggc gaagtcctcg tcaaggttag cgctgaagag 2100 tgtgtcattc tcatggctcg agattaccag cccatgtcgg agcttctgca tcgcttctcc 2160 aatggtctga cattgcagat tgcgcagatg atcccgaatg agctcaagaa gctgtcagaa 2220 tacctcctca atatatttca agtttcgggc caggccggcg agtggatcat ggctcttgtt 2280 gaggaagaga ttgatggcac cctcaaggaa agcccggcga gtcgtctgcg tttcagcaag 2340 agactgggat ctagcgagtc tagcgagtcc ttcggctcgt cgagtgaccg cgaactcttt 2400 ttgagagaca tgggcaacaa tgctaagctg gaggcgaact tgttgttccg cggcaacacc 2460 ctcttgacta agtccctgga cttccacatg aaacggctcg gaaaggagta cctggaagag 2520 actcttagcg aaagactgcg agagatcaac gaaaaggacc ccgagtgcga ggtggatcca 2580 aacaagatca catcccaaaa tgagcttgac cgcaactgga ggagactcat caacatcacc 2640 gaggatctct ggcgtgccat ttacaattcc gtctcgcgtt gcccccagga actgaggctg 2700 atctttcgac acattcaagc ttgtgccgag gatcgttatg gcgatttcct caggacggtc 2760 aagtacagca gcgtttcggg ctttcttttc ctccgcttct tcgtcccagc cgtgcttaat 2820 ccgaagctgt tcggcttact gaaaggtatg tggtgacttc ttgccaacag gttgggcgat 2880 actaaataat gcagaccacc cgaaacccag agcacgcaga acatttacac tggtagccaa 2940 gtccctacag ggccttgcca acatgtcatc ttttggtaca aaagaggcat ggatggagcc 3000 gatgaactcc ttcctctcat cgcaccgtca agagttcaag acttacctag acaacatctg 3060 ctccatctcc tcgacaacct cgcctgcccc tcctatacct ccttcgtaca gcacccctct 3120 tgcgattctg cagcgcctac cacccacttc tcgagaaggt tttccttctc ttccgtatct 3180 catcgaccat gcacgcaact ttgctgctct ggtagaccta tggctccaga atacgagaag 3240 cagcgcgccg aatatccagt caacagatgg cgatctttc cgctttcaca acatctgcgt 3300 ggctctacat gaacgcacag atgattgcct gaacagggca gaacgtgccg aacgtcctag 3360 ctcgtcgttg agtgtcaaat gggaagagtt ggtcgagcaa ctgcagggtt ctgcaagctt 3420 tgacagctca aggggcgctg ccacaaggaa tcgaggagca acaatcaaag aaggagag 3480 ggagtatctg ccaatatccc cgggaacgtg cgacgaaatg actagttcct cgtccacgag 3540 cacccctgtg accatgaagc ctgttcgaca acccaagggg cggcatcagc agaacagttc 3600 catatctgcg tctactaatt cagtcgccag caataactca ggcaccatga cctttccaaa 3660 cccctttgca ccaaagactg cccgcagtgc aggttatccg ccctcagtaa acgattcggt 3720 atccgcttcc cagtctgcat cggcctccgc cagcgcatct gcatccgcaa atgaggaaac 3780 gccacctggg agctccgatg gcttgcacat ggcacctgcc cctgcttatc cacagactca 3840 tacccaccca tccgcctcta cgaattcttt cacgtatgcg aaccccaatg cacacattaa 3900 cacggggacg atggcatcag gagcccttac acgccctcct cgtagtgcag gcggccacag 3960 tctggaaaac tccgatgcag gaagcacgca cgaggagag tacactacgg cactccctgc 4020 cttctccaag gactcgcaga aggagaa ggagcgtggc ttccgcggtg tttgccatt 4080 ccaacggcaag cgtaagaca aggaagga taaggacaag aggaagaca aggaagga 4140 aaaaaaaaaaaaaaahaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaackss 4200 gaggggcaaaaaaaaaaaaaaaaaaaaaaaaagggggg agctccgaga 4260 aagggaacga agcgtgaac ggaatgaccg tggtggacac tctgccatgg gcgaatacca 4320 cagccacagt agccttcgggg gtcgagcgca gaacgaagg ttctga 4366 <210> 30 <211> 4768 <212> DNA <213> See also <400> 30 atgtctaaga agcatgagcg cggacagagt ctggacatga gcaactcgc catcttcgaa 60 caagagaaga cggtcacacc cgcaccagcg ccgccgccca ggatggggag cttgcggcca 120 acgtcgatgc tcctcactcg atctgacacg atcaccgag ggggaagtgc cgctgtcgcc 180 catgcccatg gccccgcctc gcctccgccc ttgcaccatc cacaggcagt ccaactgccc 240 ggatccgatc ttgagatcct tggccgctca tccacaaacc aacttcgtac cctctcgaaa 300 cttgcacagt ctggagaggc tgatgagttt gccatcacgt caccggcgca agaggttgtc 360 ggactaaaag gccgccgcag gcttcagaga gccgacaggt caaatgctgg ccgcctcggt 420 cagaagtcta gcggatatgg gtgggaaggg agaaattgga tggacaaaca acgacaattc 480 cttcaggcat acgaatatct ttgtcacatc ggtgaggcaa aagagtggat cgaagatgtt 540 atgaacaaat ctattggtga gattgtgaag ctggaagagg agcttaggaa cggcgagact 600 ctggcggagg tggttcaagc gctgaacccg gaccgaagat atcgtatttt ccggcaccct 660 cgtctacagt accgccattc ggacaacatt gcaatcttct tccggtacct cgacgaggtt 720 gaactacccg acctcttccg atttgaactg atcgatcttt acgaaaagaa aaacatccca 780 aaggtcatct actgcataca tgctctcagc tggctcttat accgcaaagg aattgtcgat 840 ttccgaatcg gaaacctggt tggtcaactt gagtttgagc accacgatct tgaagctatg 900 cagaagggcc tggataagct aggggccagt atgccgacct ttggtgacat gggcgctgac 960 tttggggttg agccggaacc ggaacccgaa gaaacggaag aggagcggat agaaagggaa 1020 cttggagaaa atgaggagtc catcgtcgag ctgcaagctc aagtacgcgg tgcattattg 1080 cgaatgcggc ttggggagac aatgcaggaa ctctgggact cggagaactg gcttgtcgac 1140 cttcaggccc ggattcgtgg tgattttgcg cgccagatca tcgactaccg actaaacatg 1200 aggcgcttcg cagtgaatct acagagtgcc gcacgcgggt tcctcgttcg gtcgcggcaa 1260 gcagagagag agtacatgtg gaagcgctca gagcccgccg ttctgaagct acagagtctc 1320 ttccgggctg caaaggtccg cgatgaggta cgagacgtgc gatctcaatt gtcagaggct 1380 acaggtcctg tacgcgagat ccaagcggtt atgcgaggct ttctcgcccg caagggtgtg 1440 cgcacccagg tgcaagagac gagtcgaacg tcgggagccg caccgggtct ccaggcggcc 1500 attcgtggta tgctgctgag gaacaggctt gatcatgaca gggctattct tgccgaggaa 1560 gctgtttcga tctgcagctt tcaggccgcc tcccgtgcct tgctcacaag aaaacaggtc 1620 gcccttcaac gggagtcact agcaagcttc acggcgcagt gggagggtct tcaatccgcc 1680 tccaggggga tgttcgccag gaacagcatc catgtcacca aggcggagct ccggggacac 1740 tctcctgcca ttggcctcct gcaggctttt tcaagggccg gtgctgtacg acgtgaaacg 1800 acccgggtgt tggacgccat cgctgtacac gagccgcagg tggttgagct tcagggcttg 1860 atccgcggcg ccattcaacg ccaacgtatt gccgccgact accaggacct tgaggaacaa 1920 gtccctcaga ttaccgacct gcagtctcag atccgtggta tgctctgccg caaagagcaa 1980 ggtgagcttc ttgatcagct ccagagcaac gaagagcaaa tcatcacttt gcaggcccag 2040 atcagggcta tgatcctgcg aaacaacttg gatgtagtgc tggccgagct cgaagagcaa 2100 gaagggacga ttgtgcagct gcaggctgcg gccaggggtg tgattgtacg caagaggttc 2160 gaggaagaa agcgtcactt caaggaaac atgtccaagg tcatcaagat ccaaagtttt 2220 gttcgtggaa agctccaagg tgaagcctac aagagcctca caacaggcaa gagcccgccc 2280 gtcagtgccg tcaagaactt tgtccatctg ctgaacgaca gcgattttga cttcaacgag 2340 gaggttgagt ttgagcggat gcgcaagact gtggtacaac aggtgcggca aaacgagatg 2400 ttggagcagt acatcgacca gctggacatc aagatcgctc tgctcgtcaa gaacaagatc 2460 actctggacg aggtagttag gcaccagagc aactttggtg gccacaccag caatctgata 2520 gcgaacagct ccatcgcttc agtgaaccag tatgatctca aggccctgaa caagacgtcg 2580 aggaagaagc tcgagtcata ccagcatctc ttctacaacc tacaaacgca accgcaatat 2640 ctggcacgcc tgttccgcag gatacgtgag caaggcacgg ccgagaagga gtgcaagcgc 2700 atcgagcatc tcatcatggg tctctttggg tatgcacaaa agaggagaga agagtactac 2760 ctcctcaagc taatttctcg ctctatctgg gaggaggttg aagctagcca catggtacaa 2820 gactcactac gtggtaacct cttctggtct aagctcctag gcaactattc gaggtcacct 2880 cgcgacagga agtacctgcg agacctgctc ggccctctga ttcgtgacaa cattatcgag 2940 gaccctgctc tcgaccttga aagcgatcct ctccagatct atcgatccgc catcaacaac 3000 gaggagctgc ggacgggcat gccaagccaa aggccactcg acgtccccag ggaagtagcc 3060 atcaaggatc ccgagacgag ggagctgttc attgatcatc ttcgggatct ccgtgagatt tgcgaccagt tcttgcttgc cctcgaagac ctgcttcctc gactgccata tggcctcaga 3180 3240. tacatatgcc gccagatgtt tgatgccttg tgccaacatt tcaagcgtga gccgcagcac atattgctac agatggtggg caactggttc tggcgctttt acctgcagcc tgccctgacg gctcctgaga acgtcggcgt gatggagag gggttgagcc cgctgcagaa gcgcaacctg ggtgaggttg ccaaggttct cggccaggta gcctctggcc gtccgtttgg cggtgataat atctacctgc agccattaa cgcctttgtc gctgagtcca tggagcgttt aggccatatc 3480. ctgggcgagc tgatctcagt cgccgatgcc gaaagtacat ttgacattga tgagttcaac gacctttacg ccaaaaaccg gcccacgctt fathercaagc ttgcagatat cttcgccata cacaacctga tctcgtcaga ccttcccact atttgtccca accgcgacga catgctccgg gagatcatgc aggagctcgg tagtgccaag aacaacgaga gtgatgac ggctaccggc tcgtccgaca tccagatgtt cctcactccc aagctgcacg atgtcgaaga tcccgaggca 3780. gagatcaagg ctctcttcat ggagacgaag cgctgcatcc tgtacattat tcgtgtccag 3840 tcaggctcaa ccctcctcga gatcctggtc aagcccgtca cgcaagagga cgagcgcaag 3900 tggatggcgg tgctgcacga cgactttagt gacggcgggt ccacaaaggg agcttattcc 3960 gacgttaata tggtcgacgt tacccgtatg tcgtacctcg acctcaagcg cacggcactc 4020 gagaacgtca tgaggctgga gcagcccggc aggatctcca agcacaacca ctaccaagat 4080 atattgaacg ccattgcact cgatatccgg accaaaaagca ggaggagat tcagaggcag 4140 cgcgagctcg acggggtccg catgacgctt tctaatctcc acgagaaggc aaagtaccta 4200 gagcaacagc gcaagagcta cgatgactac attgagcagg ccatggcgac tctgcagaat 4260 aggaaagggt aagtcgacct acgataccaa catgcttctc gtctgaacaa gtaaaggcta 4320 acctcgttgg tttgttaact ttgaacagca agaaacggtt cctgcttcca ttcacaaaagc 4380 agtacaacca ccaacgcgag ctcgagcgta gcggccgggt gcccaagttc ggatcgtaca 4440 agtacagcgc acgccagctc gccgacaagg gcgtacttgt cagctggggcg ggagtgtcgg 4500 agcgcgacct gagccagatc aacctcacta tctcttgcga cgaggtgggc gtatttgtca 4560 tcgagggctc gcgtggccac atccagatcc ccggcgcgag tgccctagtc cctatcgagg 4620 acctgctgca agcccagttt gagcgcatc agttcatgaa cctcttcgag ggcaacctga 4680 ggctcaatgt caatatcttg ctgcatctgc tttataagaa attttatagg acacaataaa 4740 tggtcggggc gaattgggga gggtctaa 4768 <210> 31 <211> 2508 <212> DNA <213> Colletotrichum acutatum <400> 31 atgtccgtca tgctccaaac accatctcga gcttctaccg cctcctcctc ctccttccaa 60 cccctttccc gccaaaacac catgtctttct tacgatggat cgcggtccgc ccgccaatcg 120 aagcgttact ccatgtccgc gctgtacatg tccatgtcag cgaacgagac tgatctggag 180 attgaggatg acttggccaa aggtaggctc tgtaacctcc gcagtttcct tgccctttg 240 ccctactgac gatggatttt acagcccaga agattctcag agagctcaag tccaagatct 300 cttcgcagtc caaaaagaac ttcgtactgg aaaaggatgt acgatatctt gattcacgaa 360 tcgccctcct catccagaac cgaatggctc tcgaggaaca gaacgaagtc gcgagccact 420 tggaagacgc gacggatatg caagaaggcg cctttcctaa cgacgacaag acgcaaaagt 480 atggcaactt gatgttcttg ttgcaatccg agccgaggca tattgcacac ctctgccgtc 540 tggtgtcaat gtcggaaatc gactctctgc tgcagactgt catgttcacc atctacggaa 600 accaatacga gagtcgcgaa gaacatctgc tcttgactat gttccaggtt tgtgacccgt 660 gactatacta cgcgatctgg caagctgact cttgacccat tagtctgttc tgacctacca 720 attcgacaac acccccgaat attcttcgct tctgcgtgcg aacacccccg tctcgagaat 780 gatgaccacg tatacgcgga gaggaccagg acagagcttt ctcaagtcag ttctcgctga 840 tagaatcaac agtctgatcg agttgaagga tctcgacctg gagatcaacc ccctcaaggt 900 ctacgagcgc atgattgagc aaattgagga ggacactggc agtctgcctg catcgcttcc 960 caagggcgtt actgctgagc aggctgcgga gaacccccaa gttcaagcca tcatcgagcc 1020 gcgtctgaca atgctcaccg agattgctaa tggcttcttg acaaccatca ttgacggact 1080 cgacgaagcg ccgtacggta ttcggtggat ttgcaaacag attcgcagct tgacgaagcg 1140 caagtaccct gatgccaatg atcaggtcat ttgcactctt atcggcggat tcttcttctt 1200 gcggttcatc aacccggcaa tcgtgacacc aaagtcatac atgctcattg acggtcagcc 1260 ggctgatcgc ccgagaagaa cgctgacttt gattgcaaag atgctgcaaa accttgctaa 1320 caagccctcc tacgccaagg agccatacat ggccaagctg caacccttca tctaccagaa 1380 caaggagcgt atcaacaaat tcatgcttga cttgtgtgaa gttggcgact tttatgagag 1440 cttggaaatg gataactacg tcgcactctc gaaaggat ttggaactgt ccatcacctt 1500 1560 cgacaactcg cacttgggca tcatcatgtc tgagctagga ggcgcacccc cacaggttcc 1620 tcgcaaggag aaccgcgcca tcaacctccc cctctttagc cgatgggaaa cagcgatcga 1680 cgacttgaca gcggcgctcg acatcacgca agaggaagtg tactttatgg aagccaagtc 1740 cgtctttgtg caaattatgc gatcgatccc atccaacagc agtgttgcac gaagacctct 1800 gcgactcgag cggattgccg acgccgcagc tacgagccga aacgacgccg tgatggtttcg 1860 1920 1980 gctcaaggac ggagtcattc aagagactgg caaattggaa gaagtctaca agaccattcg 2040 cgaccacaac aactatctcg ttggccaatt ggagacgtac aagagctacc tgcacaacgt 2100 gcgttcgcaa agcgagggaa ccaagcgcaa gcagcagaag caacaggtcc tcgggcctta 2160 caaattcact caccagcagc tggagaagga gggtgtcatt caaaaaagca acgtccccgga 2220 taaccggagg gccaacattt acttcaactt taccagccct ttgccgggaa cttttgttat 2280 ctctcttcac tacaagggta cgttgcctcg attggtcatt gcgcaacttt tactgacttt 2340 tgtacaggac gcaacgcgg tcttcttgaa ctggatctca agctcgacga cctgcttgag 2400 atgcagaaag acggcagga cgacctagat ctggagtacg tccagttcaa cgtcaccaag 2460 gttctcactt tgttgaacaa gcgatttgcg agaagaagg ggtggtaa 2508 <210> 32 <211> 1727 <212> DNA <213> Hypoxylon sp. <400> 32 atgactacct acactaggcg tggaccagga cagagcttct tgcgaacggt actggcgcaa 60 agaatcaaca gcctaattga gttgacagat ctagaccttg agatcaaccc cttgaaagtc 120 tatgaacgca tgtgtcaaca aattgaagaa gacaccggta gtcttccccc ctctctacct 180 agaggaatca caggcgaaca agctgccgag aatccccaag tgcaagccat catagacct 240 cgtttaacga tgctaacgga gattgccaat ggcttcctga ccacaattat cgagggcctc 300 gaagaggctc cctatggcat tagatggata tgcaagcaga ttcggagttt gaccaaacga 360 aaatatcctg atgcgaatga ccaggtcatt tgcacactga tcggcggctt tttcttcctg 420 cgctttatca atcctgctat cgttacaccc aagtcctaca tgctcatcga tggagtgcct 480 tctgaacgac cacgccgaac gttaaccctg gttgccaaga tgcttcagaa cttggccaat 540 aaaccatcgt atgctaaaga accgtacatg gcgaagttgc aaccatttat ccagcagaac 600 aaggatcgtg tcaacaagtt tatgcttgat ctctgcgagg tccaggactt ctacgagagt 660 ctcgagatgg acaactatgt tgcactttca aagaaagacc tagagctctc cattacgctg 720 aatgagatat acgccatgca cgccttgatc gaaaagcaca gtggagaact ctgtagggac 780 gagaactccc acttgtcaca aatcatccag gagctcggca aagcacccgc gcaggtacct 840 cggaaggaga atagggcgat taatcttccc ctgtttagcc gatgggagac agctatagat 900 gatttgactg ccgccctaga tatcacgcag gaggaagtgt atttcatgga agcaaagtca 960 atctttgtac aagttatgcg gtccattcct gctaacagct cggttgctcg gcgacctcta 1020 cgcctagaga gaattgctga tgcggctgcc acatcaagga acgacgcagt gatggtccgg 1080 aaaggtatcc gggccatgga gctgcttagt caactacagg agatgaaggt tattgataag 1140 tcagaccagt tcagcctcct gagggatgag gtcgaacaag agttacaaca tctaggttcc 1200 ctgaaggatg gtgtcattgc cgaaaccgcg aagctcgaag aggtttacaa gacgattagg 1260 gatcataact cgtacctcgt cggccagcta gagacttaca agagctatct ccacaacgtg 1320 cgaagtcagt ccgaaggcac gagacggaaa cagcaaaagc agcaagttct cgggccttac 1380 aagtttactc accagcaact agagaaggaa ggcgtcatcc agaagagtaa tgttccggac 1440 aatagaaggg ctaacatcta cttcaatttc acaagtcctt tacctggaac ttttgtgatt 1500 tcattacact acaaaggtca gtcagaaggg acattccact tcagtcacgg gctaacaaat 1560 gaataggacg caatcgtggt cttctagaac tcgaccttaa gttggacgat ctgttagaaa 1620 tgcagaagga caatcaagat gacttggacc tcgaatacgt gcagttcaac gtcacgaagg 1680 tattggcctt gttaaacaag cgctttgcca ggaagaaggg ctggtaa 1727 <210> 33 <211> 2880 <212> DNA <213> Diaporthe ampelina <400> 33 atgtctgtga tgctgcaaac tccttcccgg gcctcaaccg catcctcctc ctcctaccag 60 gccctctccc gccagaacac catgtcttcc tacgatggct cgcggtcagc ccgccaatcg 120 aaacggtact ccatgtcggc attgtacatg tccatgtcgg cacaggaaac cgacttggaa 180 atagaagacg atcttgctaa aggtttgttc ccaacccctc tcatccaggc cgaaatcttg 240 accgaagtcc cattacttac tgtccccagc ccaaaagata ctacgggact tgaagtccaa 300 gattcctcc caatccaaga agaacttcgt gcttgaaaag gacgtgcggt acctcgactc 360 acgtattgca ttgctgattc agaatcgcat ggctttggag gagcagaacg aagtcgccag 420 ccacttagaa gacgcgacag atattcagga aggggtcttt ccaaacgacg acaagacgca 480 gagatatggc aacctcatgt ttctcttgca atcagagccc aggcacattg cgcatctctg 540 ccggctgtg tccatgtccg agatcgactc cctgctgcag acagtcatgt tcaccatcta 600 tggaaaccag tacgagagcc gagaagagca tctgctgcta actatgtttc aggtttgcct 660 accttctatt tcaacgtgag tgctcatgct aactcttgcc accagtccgt tttgacgtat 720 cagttcgata acacgcctga atattcgtcg ctgcttcgcg ccaacacacc agtctcccgg 780 atgatgacaa catacacgag gagaggcccg ggtcagagtt ttttgagatc ggtgcttgcg 840 cacaggatta atggccttat cgagctgcac gatctggatc tcgagatcaa ccccctcaaa 900 gtttacgagc gcatgtgcga acaaatcgag caggacacgg gcagccttcc gccgtctctg 960 ccaaagggca tcactgctga acaggccgcg gagaatgctc aggtccaagc tatcatcgag 1020 ccgagactca ccatgcttac cgagatcgcg aatggctttt tgtcgaccat catcgacggc 1080 ctggacgaag cgccgtacgg aattcgatgg atctgcaaac aaattcgcag cttgacgaag 1140 cggaagtacc ccgatgccaa cgaccaggtc atttgtacac tgatcggagg atttttcttc 1200 ctgcgcttca taaaccctgc catcgttacg ccgaagtcgt acatgctgat agatggaaca 1260 ccggcggatc ggccgaggag gaccttgacg ctgatcgcaa aaatgctgca aaaccttgcg 1320 aacaagccat cctacgcgaa ggagccctac atggccaagc tgcagccgtt tatccaatcg 1380 aacaaagaac ggatcaacaa gttcatgctt gatctttgcg atgtgcaaga cttctacgaa 1440 agtctggaga tggacaacta cgtggcgctt tcaaagaagg atctggagct gtccataaca 1500 ctgaacgaga tctatgccat gcacggcctc attgacaagc accggaatga aatttgcaag 1560 gacgagaact cgcacctaca catcatcatg tccgagcttg gcccttctcc tccgcaggtg 1620 cccaggaagg agaaccgggt gatcaactta ccactgttca gcagatggga gtcggccatg 1680 gatgacttga ccgccgcgct cgatatcacc caggaggaga tttatttcat ggaggccaaa 1740 aacgtatttg tacagatcat gcgttccatt ccatcgaata actcggttca gcgaaggcct 1800 cttcgcctcg agcgtatcgc cgatgcagca gcgacatctc ggaacgacgc ggttatggtc 1860 cgcaaaggta tccgtgctat ggaactgctg agtcaactcc aggagctgcg agtcatagac 1920 aaatccgacc agttcagcct gctacgagat gaggtcgagc aagagctaca gcatctgggc 1980 tctcttaagg atgcggtcct tgtggagact tccaagcttg acgaggtcta caagacaatc 2040 cgcgaccaca acacgtatct ggtcggccag ctggaaacgt acaagagcta tctgcacaat 2100 gtccgcagcc agagtgaggg tacacggcgg aaacagcaga agcagcaggt tctcggtccc 2160 tacaagttca cacaccaaca attggagaag gaaggggtta tccagaagag caatgtgccg 2220 gacaacaggc gggccaacat atatttcaac ttcacaagcc cccttccggg aacattcgtg 2280 atttctctgc actacaaggg caagtatagc agctcgaggg ctgaggacca tccggcatca 2340 ttcgtgaatt tattactgac atccatccca gggcgtaacc gcgggctctt ggagcttgat 2400 ctcaagctcg acgatctcct ggagatgcag aaagacggac aggacgagct ggacctcgag 2460 tatgtccaat tcaatgtgcc gaaagtgctc gcccttctga acaagcggtt cgctcggaag 2520 aagggttggt aaataacgga ttgccacacg ttattctgcc ttgtgtgcat cctagaagat 2580 gaggcaagcg gtggtccatc cacagtggcc tacttcttca caacacatat ctcacgatcg 2640 atcattctgc catctccaac atacaacatc atgtatcgct aagggactgt cggcgttttg 2700 gggccggcgt gcacttttat aatccttgat accatcgtct atacgcaaca tcgtctttca 2760 ggtcgcccgc ctacaccttc acctttcccc actatatatc ccgaccgaga gagccctctc 2820 tcgtactgca cgccccccgc cccctggggc agcacatcaa cagcctctat caatatctaa 2880 <210> 34 <211> 2789 <212> DNA <213> Talaromyces piceae <400> 34 ttcgacaaca cgcccgagta ctcgtcgctt ctccgtcaaa acacccccgt ttcccgcatg 60 atgaccacct acacccgccg cggtcccggt caaagctacc tgaaacatgt cttggctgaa 120 cagatcaata cgctcattga cttgcacgat gtcgatctcg agatcaaccc cttgaaggtg 180 tacgaaagta tggtgcagca gcttcaggaa gacacgggca gtttgcccga ctacctgccc 240 cgagcagtca ccgccgaagt cgctgccgag aacgagcagg tccaggcgat tattgctccg 300 cgcctgaaga tgttgacgga ccttgccaac aattttctca acaccatcat cgaggggctc 360 gaagatgctc cgtacgggat ccgctggatc tgcaaacaaa tccgaagtct ctcccgacgc 420 aagtacccgg acgctcagga ccagaccatc tgcacgctta tcggcggctt ctttttcctt 480 cgcttcatca acccggccat tgtgacgcct cggtcgtaca tgctcattga ggcgaccccg 540 accgacaagc cccgccggac cttgaccctg atcgccaaga tgctgcagaa cttggccaat 600 aagccgtcgt acgccaaaga accgtacatg gccaaattga gcccctttat cgacgagaac 660 aaagaccgcg tgaacaaatt cttgctcgat ctgtgtgaag tccaggactt ttacgagagc 720 ctggagatgg acaactatgt cgccctgacg aagcgggacc tggagctgca gatcacgttg 780 aacgaggtgt atgccacaca cgcgctgctg gagaaacaca gcgccagcct ggcggcttca 840 gaccaacact ctcacttgca agctcttctc caggaactag ggccggcacc gagccaggtt 900 ccccgggaaag aaatcgcgc gatcaacctg ccgctgttta gcaagtggga gacctcggtc 960 gacgatctca cggcggccct ggatatcacc your foot ttttctttat ggaagccaag 1020 tcgacctttg tccagatcct gcgttcgcta ccctccaact cggctgtcat gcggcgtcct 1080 ttgcggctgg atcgcatcgc cgaggccgcg gcaactctga agaacgatgc cgtcatggtt 1140 cgaaagggga ttcgtacgat ggagctcctg agccagctcc aggagctggg cgtgattgat 1200 cgatccgatg agtttgggct gttgcgcgac gaagtcgaac aggaactcgt ccatctgggt 1260 tcgctcaagg agaaagtggt gcaagagact cggcagttgg aggaagtgta caagaccatt 1320 cgcgatcaca acgcctactt ggtcggccag ctcgaaacct acaagtcgta tctgcacaac 1380 gtgcgcagcc agtccgaagg caaatcccgg aacaaaaagg agaagaacca ggagctcggt 1440 ccgtacaagt ttacccacca gcaacttgaa aaggggag tcatccgcaa aagcaacgtg 1500 cccgagaatc ggcgtgccaa catctatttt atgttcaaga gtccgctgcc gggcacattt 1560 gtcatcagtc tacactacaa aggtgagctt tcgtcctttt gttttcctgg tttgctgaag 1620 ccccgcccca aactaactat caccaggacg agcccgcggt cttctcgagc tcgacttgaa 1680 actggacgac cttttggaga tgcaaaaaga caaccaagag gaccttgatc ttgaatacgt 1740 tcaattcaac gtcaccaaag tactgaccct gctgaacaag cgctttgcgc gtaaaaaggg 1800 gtggtaatgg ccccttgacg actttccatg accctggcac cccgttgtgc tttacctaac 1860 ccgtatcctt ttgttttcga aacacagtgc ttgcgttgtc cgtgtgagtt caacagcttg 1920 ccatgatacc ccgctccggc tcgaatttag tctacatctt gattatgcta ttgatcgttt 1980 gcgcataccc ctgttggtttt tttggttcac ctgaattgtt ggtttgattt ttggaaaatg 2040 gattaaaaaa agcacaaaaa aaaagaagaa gaagaagaaa aagaggaaa aaaaaaaaa aagaggaagt caaagtctcc atggggatat cctgttatgg atgtcggggaa atgtggtgaa 2160 ttgcttacat gacttgcgtc caccgttcgc tggctcgaaa ggtgtattgt ttgtgtttgg 2220 tgattgtttg cgtgtcgtgc gctttggtgt ttgagtctct ggaccgtata cagagcggct 2280 tgaagcattt tttgtttgcg tgcgtttcga tggttgggat tgttatgctg atccgaccac 2340 gtgtaataat atatatatat atatatatat atcaatcata ggctttcatg acaatcactt 2400 cttgtctctc ctccccttgg tccatatcgc catatctggt cagaccaggg tggcgagcga 2460 atcagcacag acaaccaaac taggcaaagc tagtcgcaac tttgccgcca actgcaacac 2520 cagccacaac tgccgcccgc cgctgctccc ccagcccact tggcccgatc agcgcacgcc 2580 agacctttgt tttctagttt ctcctcagct gcaatcacac tttacccttc agaccgcaac 2640 ttcagacttg tgtgttctgc aattccttcc cttttcctct tttcctcgct gtgtcctcta 2700 tccgcacctg ccgggcacaa atcgaattga ccgcagtcat catcgactca ccaccagtca 2760 atctcgccgg accccgctat cccgcttga 2789 <210> 35 <211> 2633 <212> DNA <213> Sporothrix insectorum <400> 35 atgtctgtca tgctgcagac gccttctcga gcctccactg cctcttcttc gtcctttcaa 60 cccatctcca gacagaacac catgtcgtcc tacgatggca cgcggtccgc ccgccaatcc 120 aagcggattt ccatgtccgc cctctacatg tccatgtcgg ccaacgaaac cgacttggag 180 attgaggacg agctggccaa aggttggtcc aaagcccctg cctgctgctg ccttctgtgg 240 acgtttttgt ctggtttgca aactgcccat gatgtactaa tgccgtgttc tcttctccct 300 gccacagcac aaaagaagct tcgcgatctc aaagccaaaa tctcgatgca atcgaaacag 360 aactttgtcc tcgagaagga cgtgcggtat ctcgattcga gaattgcctt gctgattcaa 420 aatcgcatgg ccttggaaga ggtatgcatt gaagcggctg cggattacag aaacaacaaa 480 tgacccatac tccgttgttg ttgatgtcgt tcttgttctc ttgttccctt tcctaacgcc 540 aattgtgctt tagcaaaacg aagtggcgag ccgtctcgaa gacgcactcg aattgcaagt 600 cggcgccttt ccgaacgaca tgcaaaccca aaaatacggc aacctgatgt tcctgctaca 660 gtccgagcct cggcacattg cgcatctctg ccgcctggtg tccatgtccg aaatcgactc 720 actgctgcag acggtcatgt tcaccatcta cggcaaccag tacgagagcc gcgaagagca 780 cctgctcctg accatgtttc agtctgtgct cacctaccaa ttcgacaaca cccccgaata 840 ctcctcgctg ctgcgggcca acacccccgt ctcgcgcatg atgacgacgt acacgcgacg 900 cggacccggc cagagctttc tcaagaccat cctcgccgac cggatcaaca gcctcatcga 960 gctccaagac ctcgacctgg aaatcaaccc gctcaaggtc tacgagcgca tggtcgccca 1020 gatcgaagaa gacacgggca gcctccccgc gtccctcccc aagggcatca cggccgaaca 1080 ggccgccgaa aacccacagg tccaggccat catcgagccg cgcctgacca tgctcaacga 1140 gatcgccaac gggttcctcg ccaccatcat tgacggcctg gaggaggcgc cgtacggcat 1200 ccgctggatc tgcaagcaga tccgcagcct cacgaagcgc aagtaccccg acgccaacga 1260 ccaggccatc tgcaccctga tcggcggctt cttcttcctg cgcttcatca acccggccat 1320 tgtcaccccc aagtcgtaca tgctgatcga cggcacgccc gccgaccggc cgcgccggac 1380 cttgacgctg atcgccaaga tgctgcagaa cctggccaac aagccctcgt acgccaagga 1440 gccgtacatg tccaagctgc agcccttcat ccaccacaac aaagaccgtg tcaacaagtt 1500 catgctggac ctgtgcgagg tgcaggattt ctacgagagc ctggagatgg acaactacgt 1560 ggcgctgtcc aagaaggacc tcgagctgtc catcaccctc aacgagatct acgccatgca 1620 cggcctgatc gaaaagcaca gcggcgagct ctgcagcgac gcgaactcgc atctggccgt 1680 catgatggcc gacctcggtg ccgcgcccgc gcagctcccc cgcaaggaaa atcgcgtgat 1740 caacctgccc ctgttcagcc gctgggaagc cgcgctcgac gacctgacgg cggcgctcga 1800 catcacccag gaggaggtgt acttcatgga ggccaagtcc atctttgtgc aaatcatgcg 1860 gtccatcccg cagaactcgt ccgtggcgcg ccggcccctg cgcctcgaac gcatcgccga 1920 cgccgcggcc acgttcaaaa acgacgccgt catggtgcgc aagggcatcc gcgccatgga 1980 gctactgagc cagttgcagg agatgaaggt caccgataag tccgatggct tctctctgtt 2040 gcgcgacgag gtggagcagg agctgcagca cctcggctcg ctgaaggagg gcgtcctcac 2100 cgaaacgaag aagctgtccg aggtgtttgc gaccatcacc gaccacaaca cgtacctgaa 2160 cggccagctc gagacgtaca agagctacct gcacaacgtg cgcagccaga gcgaaggcac 2220 gcgccggaaa ccccagaaac agcaggtact cggcccgtac aagttcacac accagcagct 2280 agagaaggag ggcgtcatcc agaagagcaa tgtccccgac aaccgccgag ccaacatcta 2340 cttcaacttt accagtcctc tgccgggcac ctttgtcatc tccctgcatt ataaaggtaa 2400 ggcgctcctc tgccgcattt gcgttgtccc attatgcatt tgtacggttt cggtactcac 2460 taacgcatgc agggcgcacc cggggcctgt tggagctgga tctcaaactc gacgatctcc 2520 tggagatgca gaaaaacaac ttggacgagc tcgatttgga atacgttcgg ttcaacgtcc 2580 ccaaggtgct ggccctgttg aacaagcgct ttgcaaggaa gaagggctgg tag 2633 <210> 36 <211> 107 <212> PRT <213> Homo sapiens <400> 36 Met Thr Glu Tyr Lys Leu Val Val Val Gly Ala Gly Gly Val Gly Lys 1 5 10 15 Ser Ala Leu Thr Ile Gln Leu Ile Gln Asn His Phe Val Asp Glu Tyr 20 25 30 Asp Pro Thr Ile Glu Asp Ser Tyr Arg Lys Gln Val Val Ile Asp Gly 35 40 45 Glu Thr Cys Leu Leu Asp Ile Leu Asp Thr Ala Gly Gln Glu Glu Tyr 50 55 60 Ser Ala Met Arg Asp Gln Tyr Met Arg Thr Gly Glu Gly Phe Leu Cys 65 70 75 80 Val Phe Ala Ile Asn Asn Thr Lys Ser Phe Glu Asp Ile His His Tyr 85 90 95 Arg Glu Gln Ile Lys Arg Val Lys Asp Ser Glu 100 105 <210> 37 <211> 107 <212> PRT <213> Homo sapiens <400> 37 Met Thr Glu Tyr Lys Leu Val Val Val Gly Ala Gly Gly Val Gly Lys 1 5 10 15 Ser Ala Leu Thr Ile Gln Leu Ile Gln Asn His Phe Val Asp Glu Tyr 20 25 30 Asp Pro Thr Ile Glu Asp Ser Tyr Arg Lys Gln Val Val Ile Asp Gly 35 40 45 Glu Thr Cys Leu Leu Asp Ile Leu Asp Thr Ala Gly Gln Glu Glu Tyr 50 55 60 Ser Ala Met Arg Asp Gln Tyr Met Arg Thr Gly Glu Gly Phe Leu Cys 65 70 75 80 Val Phe Ala Ile Asn Asn Thr Lys Ser Phe Glu Asp Ile His Gln Tyr 85 90 95 Arg Glu Gln Ile Lys Arg Val Lys Asp Ser Asp 100 105 <210> 38 <211> 107 <212> PRT <213> Homo sapiens <400> 38 Met Thr Glu Tyr Lys Leu Val Val Val Gly Ala Gly Gly Val Gly Lys 1 5 10 15 Ser Ala Leu Thr Ile Gln Leu Ile Gln Asn His Phe Val Asp Glu Tyr 20 25 30 Asp Pro Thr Ile Glu Asp Ser Tyr Arg Lys Gln Val Val Ile Asp Gly 35 40 45 Glu Thr Cys Leu Leu Asp Ile Leu Asp Thr Ala Gly Gln Glu Glu Tyr 50 55 60 Ser Ala Met Arg Asp Gln Tyr Met Arg Thr Gly Glu Gly Phe Leu Cys 65 70 75 80<00031 <213> unknown <220> <223> Description of unknown substances: Ras ETaG sequence <400> 39 Met Gln Pro Arg Arg Glu Tyr His Ile Val Val Leu Gly Ala Ala Gln 1 5 10 15 Phe Val Gln Asn Val Trp Ile Glu Ser Tyr Asp Pro Thr Ile Glu Asp 20 25 30 Ser Tyr Arg Lys Gln Ile Glu Val Asp Gly Arg Gln Cys Ile Leu Glu 35 40 45 Ile Leu Asp Thr Ala Gly Thr Glu Gln Phe Ile Phe Ser Ile Thr Ser 50 55 60 Met Ser Ser Leu Asn Glu Leu Ser Glu Ile Arg Glu Gln Ile Leu Arg 65 70 75 80 Ile Lys Asp Asp Asp 85 <210> 40 <211> 105 <212> PRT <213> unknown <220> <223> Description of unknown substances: Ras ETaG sequence <400> 40 Met Ser Gln Arg Glu Tyr His Ile Val Val Leu Gly Ser Gly Gly Val 1 5 10 15 Gly Lys Ser Cys Leu Thr Ala Gln Phe Val Gln Asn Val Trp Ile Glu 20 25 30 Ser Tyr Asp Pro Thr Ile Glu Asp Ser Tyr Arg Lys Val Leu Glu Val 35 40 45 Asp Gly Arg His Val Ile Leu Glu Ile Leu Asp Thr Ala Gly Thr Glu 50 55 60 Gln Phe Lys Leu Tyr Met Lys Thr Gly Gln Gly Phe Leu Leu Val Phe 65 70 75 80 Ser Ile Thr Ser Glu Ser Ser Phe Trp Glu Leu Ala Glu Leu Arg Glu 85 90 95 Gln Ile Arg Arg Ile Lys Glu Asp Ser {100 105} <210> 41 <211> 112Cys Val Ile Asp Glu Glu Val Ala Leu Leu Asp Val Leu Asp Thr Ala 50 55 60 Gly Gln Glu Glu Tyr Ser Ala Met Arg Glu Gln Tyr Met Arg Thr Gly 65 70 75 80 Glu Gly Phe Leu Leu Val Tyr Ser Ile Thr Ser Arg Gln Ser Phe Glu 85 90 95 Glu Ile Thr Thr Phe Gln Gln Gln Ile Leu Arg Val Lys Asp Lys Asp 100 105 110 <210> 42 <211> 123 <212> PRT <213> Unknown <220> <223> Description of unknown substance: Ras ETaG sequence <400> 42 Met Ser Arg Ser Ala Ala Gln Ala Ser Phe Leu Arg Glu Tyr Lys Leu 1 5 10 15 Val Val Val Gly Gly Gly Gly Met Ser Leu Val Ser Arg Val Gly Lys 20 25 30 Ser Ala Leu Thr Ile Gln Phe Ile Gln Ser His Phe Val Asp Glu Tyr 35 40 45 Asp Pro Thr Ile Glu Asp Ser Tyr Arg Lys Gln Cys Val Ile Asp Asp 50 55 60 Glu Val Ala Leu Leu Asp Val Leu Asp Thr Ala Gly Gln Glu Glu Tyr 65 70 75 80 Gly Ala Met Arg Glu Gln Tyr Met Arg Thr Gly Glu Gly Phe Leu Leu 85 90 95 Val Tyr Ser Ile Thr Ser Arg Asn Ser Phe Glu Glu Ile Ser Thr Phe 100 105 110 His Gln Gln Ile Leu Arg Val Lys Asp Lys Asp 115 120 <210> 43 <211> 133 <212> PRT <213> Unknown <220> <223> Description of unknown substance: Ras ETaG sequence <400> 43 Met Pro Glu Val Met Asn Ala Met Tyr Ala Thr Lys Gly Gly Ile Phe 1 5 10 15 Asp Val Ser Glu Asn Asp Lys Ala Gln Phe Leu Arg Glu Tyr Lys Leu 20 25 30 Val Val Val Gly Gly Gly Gly Val Gly Lys Ser Ala Leu Thr Ile Gln 35 40 45 Phe Ile Gln Ser His Phe Val Asp Glu Tyr Asp Pro Thr Ile Glu Asp 50 55 60 Ser Tyr Arg Lys Gln Cys Ile Ile Asp Asp Glu Val Ala Leu Leu Asp 65 70 75 80 Val Leu Asp Thr Ala Gly Gln Glu Glu Tyr Gly Ala Met Arg Glu Gln 85 90 95 Tyr Met Arg Thr Gly Glu Gly Phe Leu Leu Val Tyr Ser Ile Thr Ser 100 105 110 Arg Asn Ser Phe Glu Glu Ile Ser Ile Phe His Gln Gln Ile Leu Arg 115 120 125 Val Lys Asp Gln Asp 130 <210> 44 <211> 121 <212> PRT <213> Unknown <220> <223> Description of unknown substance: Ras ETaG sequence <400> 44 Met Ala Asn Asn Ala Ala Ser Arg Ala Ala Gln Ala Gln Phe Leu Arg 1 5 10 15 Glu Tyr Lys Leu Val Val Val Gly Gly Gly Gly Val Gly Lys Ser Ala 20 25 30 Leu Thr Ile Gln Phe Ile Gln Ser His Phe Val Asp Glu Tyr Asp Pro 35 40 45 Thr Ile Glu Asp Ser Tyr Arg Lys Gln Cys Val Ile Asp Glu Glu Val 50 55 60 Ala Leu Leu Asp Val Leu Asp Thr Ala Gly Gln Glu Glu Tyr Gly Ala 65 70 75 80 Met Arg Glu Gln Tyr Met Arg Thr Gly Glu Gly Phe Leu Leu Val Tyr 85 90 95 Ser Ile Thr Ala Arg Ser Ser Phe Glu Glu Ile Asn Gln Phe Tyr Gln 100 105 110 Gln Ile Leu Arg Val Lys Asp Gln Asp 115 120 <210> 45 <211> 127 <212> PRT <213> Unknown <220> <223> Description of unknown substance: Ras ETaG sequence <400> 45 Met Thr Leu Tyr Lys Leu Val Val Leu Gly Asp Gly Gly Val Gly Lys 1 5 10 15 Thr Ala Leu Thr Ile Gln Leu Cys Leu Asn His Phe Val Glu Thr Tyr 20 25 30 Asp Pro Thr Ile Glu Asp Ser Tyr Arg Lys Gln Val Val Ile Asp Gln 35 40 45 Gln Ser Met Leu Glu Val Leu Asp Thr Ala Gly Gln Glu Glu Tyr Thr 50 55 60 Ala Leu Arg Asp Gln Trp Ile Arg Asp Gly Glu Gly Phe Val Leu Val 65 70 75 80 Tyr Ser Ile Thr Ser Arg Ala Ser Phe Ala Arg Ile Pro Lys Phe Tyr 85 90 95 Asn Gln Ile Lys Met Val Lys Glu Ser Ala Ser Ser Gly Ser Pro Ala 100 105 110 Gly Ala Ser Tyr Leu Thr Ser Pro Ile Asn Ser Pro Ser Gly Pro 115 120 125 <210> 46 <211> 117 <212> PRT <213> unknown <220> <223> unknown description: Ras ETaG sequence <400> 46 Met Asp Ala Asp Pro Tyr Leu Leu Lys Phe Leu Arg Glu Tyr Lys Leu 1 5 10 15 Val Val Val Gly Gly Gly Gly Val Gly Lys Ser Cys Leu Thr Ile Gln 20 25 30 Leu Ile Gln Ser His Phe Val Asp Glu Tyr Asp Pro Thr Ile Glu Asp 35 40 45 Ser Tyr Arg Lys Gln Cys Val Ile Asp Glu Glu Val Ala Leu Leu Asp 50 55 60 Val Leu Asp Thr Ala Gly Gln Glu Glu Tyr Ser Ala Met Arg Glu Gln 65 70 75 80 Tyr Met Arg Thr Gly Glu Gly Phe Leu Leu Val Tyr Ser Ile Thr Ser 85 90 95 Arg Gln Ser Phe Glu Glu Met Leu Thr Phe Gln Gln Gln Ile Leu Arg 100 105 110 Val Lys Asp Arg Asp 115 <210> 47 <211> 203 <212> PRT <213> Homo sapiens <400> 47 Leu Glu Ala Asn Glu Gly Ser Lys Thr Leu Gln Arg Asn Arg Lys Met 1 5 10 15 Ala Met Gly Arg Lys Lys Phe Asn Met Asp Pro Lys Lys Gly Ile Gln 20 25 30 Phe Leu Val Glu Asn Glu Leu Leu Gln Asn Thr Pro Glu Glu Ile Ala 35 40 45 Arg Phe Leu Tyr Lys Gly Glu Gly Leu Asn Lys Thr Ala Ile Gly Asp 50 55 60 Tyr Leu Gly Glu Arg Glu Glu Leu Asn Leu Ala Val Leu His Ala Phe 65 70 75 80 Val Asp Leu His Glu Phe Thr Asp Leu Asn Leu Val Gln Ala Leu Arg 100 105 110 Arg Met Met Glu Ala Phe Ala Gln Arg Tyr Cys Leu Cys Asn Pro Gly 115 120 125 Val Phe Gln Ser Thr Asp Thr Cys Tyr Val Leu Ser Tyr Ser Val Ile 130 135 140 Met Leu Asn Thr Asp Leu His Asn Pro Asn Val Arg Asp Lys Met Gly 145 150 155 160 Leu Glu Arg Phe Val Ala Met Asn Arg Gly Ile Asn Glu Gly Gly Asp 165 170 175 Leu Pro Glu Glu Leu Leu Arg Asn Leu Tyr Asp Ser Ile Arg Asn Glu 180 185 190 Pro Phe Lys Ile Pro Glu Asp Asp Gly Asn Asp 195 200 <210> 48 <211> 203 <212> PRT <213> Homo sapiens <400> 48 Asp Pro Arg Glu Leu Ile Glu Ile Lys Asn Lys Lys Lys Leu Leu Ile 1 5 10 15 Thr Gly Thr Glu Gln Phe Asn Gln Lys Pro Lys Lys Gly Ile Gln Phe 20 25 30 Leu Gln Glu Lys Gly Leu Leu Thr Ile Pro Met Asp Asn Thr Glu Val 35 40 45 Ala Gln Trp Leu Arg Glu Asn Pro Arg Leu Asp Lys Lys Met Ile Gly 50 55 60 Glu Phe Val Ser Asp Arg Lys Asn Ile Asp Leu Leu Glu Ser Phe Val 65 70 75 80 Ser Thr Phe Ser Phe Gln Gly Leu Arg Leu Asp Glu Ala Leu Arg Leu 85 90 95 Tyr Leu Glu Ala Phe Arg Leu Pro Gly Glu Ala Pro Val Ile Gln Arg 100 105 110 Leu Leu Glu Ala Phe Thr Glu Arg Trp Met Asn Cys Asn Gly Ser Pro 115 120 125 Phe Ala Asn Ser Asp Ala Cys Phe Ser Leu Ala Tyr Ala Val Ile Met 130 135 140 Leu Asn Thr Asp Gln His Asn His Asn Val Arg Lys Gln Asn Ala Pro 145 150 155 160 Met Thr Leu Glu Glu Phe Arg Lys Asn Leu Lys Gly Val Asn Gly Gly 165 170 175 Lys Asp Phe Glu Gln Asp Ile Leu Glu Asp Met Tyr His Ala Ile Lys 180 185 190 Asn Glu Glu Ile Val Met Pro Glu Glu Gln Thr 195 200 <210> 49 <211> 203 <212> PRT <213> Homo sapiens <400> 49 Asp Pro Asn Ala Leu Arg Gln Gln Arg Ser Arg Lys Ser Met Ile Met 1 5 10 15 Lys Gly Ala Ser Lys Phe Asn Glu Asn Pro Lys Ala Gly Ile Ala Phe 20 25 30 Leu Val Ala Gln Gly Val Ile Gln Glu Pro Glu Asn Pro Lys Asn Ile 35 40 45 Ala Glu Phe Ile Lys Gly Thr Thr Arg Ile Asp Lys Lys Ile Leu Gly 50 55 60 Glu Phe Ile Ser Lys Lys Thr Asn Glu Asn Ile Leu Asn Glu Phe Met 65 70 75 80 [[ID=2⑨]]Lys Leu Phe Asn Phe Ala Gly Lys Arg Ile Asp Glu Ala Ile Arg Glu 85 90 95 Leu Leu Gly Ala Phe Arg Leu Pro Gly Glu Ser Ala Leu Ile Glu Arg 100 105 110 Ile Val Glu Val Phe Ala Ala Gln Tyr Met Asp Asp Ala Lys Pro Ala 115 120 125 Gly Ile Ala Asp Ser Thr Ala Ala Phe Val Leu Val Tyr Ala Thr Ile 130 135 140 Leu Leu Asn Thr Asp Gln His Asn Pro Asn Phe Arg Gly Gln Lys Arg 145 150 155 160 Met Thr Ile Glu Asn Phe Ala Gln Asn Leu Arg Gly Val Asn Asp Gln 165 170 175 Gly Asp Phe Asp Ser Asn Phe Leu Gln Glu Ile Phe Asp Ser Ile Arg 180 185 190 Thr His Glu Ile Ile Leu Pro Glu Glu His Asp 195 200
Claims
1. A method for identifying a candidate modulator, comprising the steps of: a query set of nucleic acid sequences, each of which is present in a fungal strain and comprises a biosynthetic gene cluster; An embedded target gene sequence is identified within at least one fungal nucleic acid sequence, wherein the embedded target gene sequence is characterized in that: is not required for or involved in the biosynthesis of the product of the biosynthetic gene cluster; within a vicinity relative to at least one gene in the cluster; is homologous to a mammalian nucleic acid sequence; and is co-regulated with at least one biosynthetic gene in the cluster; wherein the product of the biosynthetic gene cluster is a candidate regulator of the mammalian nucleic acid sequence; wherein the adjacent region is no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100 kb upstream or downstream of a biosynthetic gene in the cluster; and The method is implemented by a computer system.
2. The method of claim 1, wherein the embedded target gene sequence is in a proximal region relative to at least one biosynthetic gene in the cluster.
3. The method of any one of the preceding claims, wherein the nucleic acid sequence comprising the biosynthetic gene cluster does not comprise a sequence other than the nucleic acid sequence of the adjacent regions relative to the biosynthetic genes in the biosynthetic gene cluster and the nucleic acid sequence of the biosynthetic gene cluster.
4. The method of claim 1, wherein the adjacent region is no more than 50 kb upstream or downstream of the biosynthetic genes in the cluster.
5. The method of claim 4, wherein the adjacent region is no more than 40 kb upstream or downstream of the biosynthetic genes in the cluster.
6. The method of claim 5, wherein the adjacent region is no more than 30 kb upstream or downstream of the biosynthetic genes in the cluster.
7. The method of claim 6, wherein the adjacent region is no more than 20 kb upstream or downstream of the biosynthetic genes in the cluster.
8. The method of claim 7, wherein the adjacent region is no more than 10 kb upstream or downstream of the biosynthetic genes in the cluster.
9. The method of claim 1, wherein the contiguous region is a region between two biosynthetic genes in a biosynthetic gene cluster.
10. The method of claim 1, wherein the mammalian nucleic acid sequence is an expressed sequence.
11. The method of claim 10, wherein the mammalian nucleic acid sequence is a gene.
12. The method of claim 11, wherein the mammalian nucleic acid sequence is a human nucleic acid sequence.
13. The method of claim 1, wherein the embedded target gene sequence is homologous to an expressed mammalian nucleic acid sequence in that the base sequence of the embedded target gene sequence, or a portion thereof, is at least 50%, 60%, 70%, 80% or 90% identical to the base sequence of the mammalian nucleic acid sequence, or a portion thereof.
14. The method of claim 13, wherein the sequence or a portion thereof is at least 50, 100, 150, or 200 base pairs in length.
15. The method of claim 1, wherein the embedded target gene sequence is homologous to the expressed mammalian nucleic acid sequence in that a portion of the base sequence of the embedded target gene sequence is at least 50%, 60%, 70%, 80% or 90% identical to a portion of the base sequence of the mammalian nucleic acid sequence.
16. The method of claim 15, wherein the homologous portion is at least 50, 100, 150, or 200 base pairs in length.
17. The method of claim 13, wherein the identity is at least 50%.
18. The method of claim 17, wherein the identity is at least 60%.
19. The method of claim 18, wherein the identity is at least 70%.
20. The method of claim 19, wherein the identity is at least 80%.
21. The method of claim 20, wherein the identity is at least 90%.
22. The method of claim 14, wherein the length is at least 100 base pairs.
23. The method of claim 22, wherein the length is at least 150 base pairs.
24. The method of claim 23, wherein the length is at least 200 base pairs.
25. The method of claim 1, wherein the embedded target gene sequence is homologous to the expressed mammalian nucleic acid sequence in that the product encoded by the embedded target gene or a portion thereof is homologous to the product encoded by the mammalian nucleic acid sequence or a portion thereof.
26. The method of claim 25, wherein the product is a protein.
27. The method of claim 25, wherein the protein encoded by the embedded target gene or a portion thereof is at least 50%, 60%, 70%, 80% or 90% similar to the protein encoded by the mammalian nucleic acid sequence or a portion thereof.
28. The method of claim 25, wherein the protein encoded by the embedded target gene or a portion thereof has a 3-dimensional structure similar to the 3-dimensional structure of the protein encoded by the mammalian nucleic acid sequence or a portion thereof.
29. The method of claim 28, wherein the portion of the protein encoded by the embedded target gene has a similar 3-dimensional structure to the portion of the protein encoded by the mammalian nucleic acid sequence.
30. The method of claim 28, wherein the similarity is that the structures have a Ca backbone rmsd (root mean square deviation) within 10 square angstroms and have the same overall fold or core domain.
31. The method of claim 28, wherein the protein encoded by the embedded target gene or a portion thereof has a similar three-dimensional structure to the protein encoded by the mammalian nucleic acid sequence, in that a small molecule that binds to the protein encoded by the embedded target gene or a portion thereof also binds to the protein encoded by the mammalian nucleic acid sequence or a portion thereof.
32. The method of claim 31, wherein the small molecule binds to the protein encoded by the embedded target gene and the mammalian nucleic acid sequence or a portion thereof with a Kd of no more than 100 μM, 50 μM, 10 μM, 5 μM or 1 μM.
33. The method of claim 31, wherein the small molecule is produced by a fungus.
34. The method of claim 33, wherein the small molecule is acyclic.
35. The method of claim 33, wherein the small molecule is cyclic.
36. The method of claim 33, wherein the small molecule is a secondary metabolite molecule produced by a fungus.
37. The method of claim 31, wherein the small molecule is non-ribosomally synthesized.
38. The method of claim 37, wherein the small molecule is a biosynthetic product of a biosynthetic gene cluster.
39. The method of claim 25, wherein a portion of the protein encoded by the embedded target gene is at least 50%, 60%, 70%, 80% or 90% similar to a portion of the protein encoded by the expressed mammalian nucleic acid sequence.
40. The method of claim 39, wherein the similarity is at least 50%.
41. The method of claim 39, wherein the similarity is at least 60%.
42. The method of claim 39, wherein the similarity is at least 70%.
43. The method of claim 39, wherein the similarity is at least 80%.
44. The method of claim 39, wherein the similarity is at least 90%.
45. The method of claim 39, wherein the portion of the protein is a protein domain.
46. The method of claim 45, wherein the portion of the protein is a set of amino acid residues essential for function.
47. The method of claim 46, wherein the function is an enzymatic function.
48. The method of claim 47, wherein the set of amino acid residues contacts a substrate.
49. The method of claim 47, wherein the set of amino acid residues is contacted with an intermediate.
50. The method of claim 47, wherein the collection of amino acid residues is contacted with a product.
51. The method of claim 46, wherein the function is an interaction with another entity.
52. The method of claim 51, wherein the entity is a small molecule.
53. The method of claim 51, wherein the entity is a lipid.
54. The method of claim 51, wherein the entity is a carbohydrate.
55. The method of claim 51, wherein the entity is a nucleic acid.
56. The method of claim 51, wherein the entity is a protein.
57. The method of claim 51, wherein each of said residues in said set of amino acid residues is Inside.
58. The method of claim 39, wherein the portion of the protein comprises at least 2 to 200, 2 to 100, 2 to 50, 2 to 40, 2 to 30, 2 to 20, 2 to 15, 2 to 10, 3 to 200, 3 to 100, 3 to 50, 3 to 40, 3 to 30, 3 to 20, 3 to 15, 3 to 10, 4 to 200, 4 to 100, 4 to 50, 4 to 40, 4 to 30, 4 to 20, 4 to 15, 4 to 10, 5 to 200, 5 to 100, 5 to 50, 5 to 40, 5 to 30, 5 to 20, 5 to 15, or 5 to 10 amino acid residues.
59. The method of claim 39, wherein the portion of the protein comprises at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, or 150 amino acid residues.
60. The method of claim 59, wherein the portion of the protein comprises at least 5 amino acid residues.
61. The method of claim 60, wherein the portion of the protein comprises at least 10 amino acid residues.
62. The method of claim 61, wherein the portion of the protein comprises at least 15 amino acid residues.
63. The method of claim 62, wherein the portion of the protein comprises at least 20 amino acid residues.
64. The method of claim 1, wherein the embedded target gene is co-regulated with at least one gene in the cluster.
65. The method of claim 1, wherein the embedded target gene is absent from 80%, 90%, 95% or 100% of all fungal nucleic acid sequences in the collection from different fungal strains and comprising homologous or identical biosynthetic gene clusters.
66. The method of claim 1, wherein the collection comprises at least 100, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 1,500,000, 2,000,000, or 2,500,000 independent fungal nucleic acid sequences.
67. The method of claim 1, wherein the collection comprises nucleic acid sequences from at least 100, 500, 1,000, 5,000, 10,000, 15,000, 20,000, 22,000, 25,000, or 30,000 independent fungal strains.
68. The method of claim 1, wherein the embedded target gene sequence is not a housekeeping gene.
69. The method of claim 1, wherein the embedded target gene sequence is or comprises a sequence having homology to a second nucleic acid sequence or a portion thereof in the same genome.
70. The method of claim 69, wherein the homology is at least 20%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5%.
71. The method of claim 70, wherein the homology is at least 30%.
72. The method of claim 70, wherein the homology is at least 50%.
73. The method of claim 70, wherein the homology is at least 70%.
74. The method of claim 70, wherein the homology is at least 80%.
75. The method of claim 70, wherein the homology is at least 90%.
76. The method of claim 69, wherein a portion is at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, or 1000 nucleobases in length.
77. The method of claim 76, wherein the portion is at least 100 nucleobases in length.
78. The method of claim 76, wherein the portion is at least 150 nucleobases in length.
79. The method of claim 76, wherein the portion is at least 200 nucleobases in length.
80. The method of claim 1, wherein the embedded target gene sequence is or comprises a sequence encoding a product having homology to a product or a portion thereof encoded by a second nucleic acid sequence in the same genome.
81. The method of claim 70, wherein the homology is at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5%.
82. The method of claim 48, wherein the homology is at least 70%.
83. The method of claim 48, wherein the homology is at least 80%.
84. The method of claim 48, wherein the homology is at least 90%.
85. The method of claim 80, wherein a portion comprises at least 5 amino acid residues.
86. The method of claim 85, wherein a portion comprises at least 10 amino acid residues.
87. The method of claim 86, wherein a portion comprises at least 15 amino acid residues.
88. The method of claim 87, wherein a portion comprises at least 20 amino acid residues.
89. The method of claim 69, wherein the second nucleic acid sequence is or comprises a housekeeping gene.
90. The method of claim 69, wherein the embedded target gene sequence encodes a product that provides resistance to a product of the biosynthetic gene cluster and the second nucleic acid sequence does not.
91. The method of claim 90, wherein the embedded target gene sequence encodes a protein that provides resistance to a small molecule product of the biosynthetic gene cluster, while the protein encoded by the second nucleic acid sequence does not.
92. The method of claim 1, wherein the nucleic acid sequences within the collection comprise biosynthetic gene clusters, the biosynthetic genes of which encode enzymes involved in the synthesis of compounds sharing at least one common chemical attribute.
93. The method of claim 1, wherein the nucleic acid sequences are from multiple fungal strains.
94. The method of claim 92, wherein the common chemical attribute is or comprises a ring system.
95. The method of claim 92, wherein the common chemical attribute is or comprises a macrocycle.
96. The method of claim 92, wherein the common chemical attribute is or comprises an acyclic backbone.
97. The method of claim 92, wherein the compounds sharing at least one common chemical attribute are polyketides.
98. The method of claim 92, wherein the compounds sharing at least one common chemical attribute are non-ribosomal peptides.
99. The method of claim 92, wherein the compounds sharing at least one common chemical attribute are alkaloids.
100. The method of claim 92, wherein the compounds sharing at least one common chemical attribute are terpenes / isoprenes.
101. A method comprising: identifying an embedded target gene that is within a proximal region relative to at least one biosynthetic gene in a biosynthetic gene cluster; The same origin as humans; is not required for or involved in the biosynthesis of a product of the biosynthetic gene cluster; and is co-regulated with at least one biosynthetic gene in the cluster; as well as determining the effect of a product produced by an enzyme encoded by the biosynthetic gene cluster or an analog of said product on said human homolog, wherein the adjacent region is no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100 kb upstream or downstream of a biosynthetic gene in the cluster, and The step of identifying the embedded target gene is performed by a computer system.
102. The method of claim 101, wherein the embedded target gene is an embedded target gene as described in any one of claims 1 to 101.
103. A method for identifying and / or characterizing a modulator of a human target, comprising: Provide human targets; as well as Identifying a product or an analog thereof produced by an enzyme encoded by a biosynthetic gene cluster, wherein an embedded target gene is present in a proximal region relative to at least one gene in the biosynthetic gene cluster, wherein the embedded target gene: is homologous to the human target or a nucleic acid sequence encoding the human target; is not required for or involved in the biosynthesis of the product of the biosynthetic gene cluster, and is co-regulated with at least one biosynthetic gene in the cluster, wherein the adjacent region is no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100 kb upstream or downstream of a biosynthetic gene in the cluster, and The method is implemented by a computer system.
104. The method of claim 103, wherein the embedded target gene is an embedded target gene as described in any one of claims 1 to 103.
105. The method of claim 103, wherein the human target is a Ras protein.
106. The method of claim 105, wherein the human target is KRas, HRas, or NRas.
107. The method of claim 103, wherein the human target comprises a RasGEF domain.
108. The method of claim 103, wherein the human target comprises a RasGAP domain.
109. The method of claim 103, wherein the embedded target gene is an embedded target gene in one of Figures 1 to 39.
110. The method of claim 103, wherein the biosynthetic gene cluster is a biosynthetic gene cluster in one of Figures 1 to 39.
111. A database for use in the method of claim 1, comprising: a collection of nucleic acid sequences, each of which is present in a fungal strain and comprises a biosynthetic gene cluster; wherein the collection of nucleic acid sequences is contained in a computer-readable medium; and wherein one or more embedded target genes in the fungal nucleic acid sequence are indexed.
112. A computer system for use in the method of claim 1, comprising: One or more non-transitory machine-readable storage media storing data representing a set of nucleic acid sequences, each of which is present in a fungal strain and comprises a biosynthetic gene cluster, wherein one or more embedded target genes in the fungal nucleic acid sequences are indexed.
113. A computer system for use in the method of claim 1, comprising: One or more non-transitory machine-readable storage media storing data representing a collection of nucleic acid sequences, each of which is or comprises an embedded target gene sequence, wherein one or more embedded target genes in the fungal nucleic acid sequences are indexed.