Human therapeutic targets and their modulators
By identifying and characterizing ETaGs in eukaryotic biosynthetic gene clusters, the patent addresses the challenge of finding druggable human targets, enhancing the druggability of previously non-druggable proteins through modulators derived from these gene clusters.
Patent Information
- Application Number
- JP2023127703
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-09-14
- Filing Date
- 2023-08-04
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2038-09-14
AI Technical Summary
The challenge of identifying 'druggable' targets within the human proteome is significant, with only about 2% of human proteins successfully targeted by approved drugs, and only 10-15% susceptible to targeting.
The disclosure provides techniques for identifying and characterizing 'embedded target genes' (ETaGs) in eukaryotic biosynthetic gene clusters, which are homologs of human genes of therapeutic interest, and using these to develop modulators for human targets.
This approach identifies novel human targets and improves the druggability of previously non-druggable targets by leveraging the relationship between ETaGs and their related biosynthetic gene clusters, providing effective modulators and therapeutic agents.
Smart Images

Figure 0007721154000071 
Figure 0007721154000072 
Figure 0007721154000073
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 558,744, filed September 14, 2017, which is incorporated herein by reference in its entirety. [Background technology]
[0002] The identification of so-called "druggable" targets within the human proteome has been described as a "significant challenge." See, e.g., Dixon et al. Curr. See Opin. Chem. Biol. 13:549, 2009. As of 2011, reports estimated that only about 2% of human proteins have been successfully targeted by approved drugs, and furthermore, only 10-15% of human proteins are even susceptible to targeting (i.e., "druggable"). See, e.g., Stockwell Sci. Am 305:20, 2011. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Dixon et al., Curr. Opin. Chem. Biol. (2009) 13:549 Summary of the Invention [Means for solving the problem]
[0004] There is emerging evidence that some microbial biosynthetic gene clusters may contain genes (referred to herein as "passenger" genes) that do not appear to be involved in the synthesis of the associated biosynthetic product produced by the enzymes encoded by the cluster. In some cases, such passenger genes have been described as "self-protective" because they apparently encode proteins that can render the host organism resistant to the associated biosynthetic product. For example, in some cases, genes encoding transporters of biosynthetic products, detoxifying enzymes that act on biosynthetic products, or resistant variants of proteins whose activity is targeted by the biosynthetic product have been reported. For example, Cimermancic et al. al Cell 158:412, 2014;Keller Nat. Chem. Biol. 11:671, 2015. Researchers have proposed that identifying such genes and their functions can be useful in determining the role of the biosynthetic products synthesized by the enzymes of the cluster. See, e.g., Yeh et al. ACS Chem. Biol. 11:2275, 2016; Tang et al. ACS Chem. Biol. 10:2841, 2015; Regueira et al. Appl. Environ. Microbiol. 77:3035, 2011; Kennedy et al., Science 284:1368, 1999; Lowther et al., Proc. Natl. Acad. Sci. USA 95:12153, 1998; Abe et al., Mol. Genet. Genomics 268:130, 2002. In particular, the present disclosure provides a different perspective on the non-biosynthetic genes present in the biosynthetic gene clusters described herein or in the proximal zone to the biosynthetic genes of the clusters, and provides new insights into the potential utility of certain such genes in human therapeutics. In some embodiments, the present disclosure provides techniques for utilizing such insights to develop and / or improve human therapeutics.
[0005] Among other things, the present disclosure provides the insight that certain non-biosynthetic genes present in biosynthetic gene clusters or in the proximal zone relative to the biosynthetic genes of the cluster, particularly in eukaryotic (e.g., fungal, as opposed to bacterial) biosynthetic gene clusters, may represent homologs of human genes that represent targets of therapeutic interest. The present disclosure defines parameters that characterize such non-biosynthetic genes of interest, referred to herein as "embedded target genes" or "ETaGs." The present disclosure provides techniques for identifying and / or characterizing ETaGs, databases comprising biosynthetic gene cluster and / or ETaG gene sequences (and, where appropriate, associated annotation), systems for identifying and / or characterizing human target genes corresponding to ETaGs, and methods for making and / or using such human target genes and / or systems containing and / or expressing the same, etc.
[0006] The present disclosure contributes further insight that the relationship between ETaG and its related biosynthetic gene clusters (to which biosynthetic genes within the proximal zone of ETaG reside) informs the identification, design, and / or characterization of effective modulators of corresponding human target genes. The present disclosure provides techniques for such identification, design, and / or characterization, as well as agents that achieve modulation of relevant human target genes, and methods of providing and / or using such agents.
[0007] As noted above, the present disclosure encompasses the insight that ETaGs can function as functional homologs (e.g., orthologs) of human targets of medical (e.g., therapeutic) relevance. In this disclosure, the sequences of passenger (i.e., non-biosynthetic) genes within eukaryotic (e.g., fungal) biosynthetic gene clusters or in the proximal zone relative to biosynthetic genes of the clusters can be compared with the sequences of human genes. Nucleic acid sequence similarity, peptide sequence similarity, and / or phylogenetic relatedness can be determined (e.g., quantitatively assessed and / or by phylogenetic tree visualization) for the compared sequences. Alternatively, or in addition, conservation of known structural and / or protein effector elements can be assessed. In some embodiments, passenger genes with relatively high homology to human sequences and / or conserved structural and / or protein effector elements can be prioritized as ETaGs of interest as human drug targets.
[0008] In some embodiments, the present disclosure provides: interrogating a set of nucleic acid sequences, each of which is found in a fungal strain and which comprises a biosynthetic gene cluster; Among at least one of the fungal nucleic acid sequences: within a proximity zone to at least one gene in the cluster, identifying an embedded target gene (ETaG) sequence that is optionally co-regulated with at least one biosynthetic gene in the cluster; The present invention provides a method comprising:
[0009] Typically, a biosynthetic gene cluster comprises one or more biosynthetic genes. In some embodiments, a biosynthetic gene cluster comprises one or more biosynthetic genes and one or more non-biosynthetic genes. In some embodiments, the non-biosynthetic genes are regulatory, e.g., transcription factors. In some embodiments, in a biosynthetic gene cluster identified by bioinformatics, the non-biosynthetic genes may be hypothetical genes. In some embodiments, the boundaries of the biosynthetic gene cluster are defined by bioinformatics methods, e.g., antiSMASH. In some embodiments, the biosynthetic and non-biosynthetic genes are designated based on bioinformatics. In some embodiments, a non-biosynthetic gene may have a biosynthetic function even if it was identified as a non-biosynthetic gene by bioinformatics methods (and / or designated as a non-biosynthetic gene in this disclosure).
[0010] In some embodiments, the present disclosure provides: interrogating a set of nucleic acid sequences, each of which is found in a fungal strain and which comprises a biosynthetic gene cluster; Among at least one of the fungal nucleic acid sequences, within a proximity zone to at least one biosynthetic gene in the cluster; identifying an embedded target gene (ETaG) sequence that is optionally co-regulated with at least one biosynthetic gene in the cluster; The present invention provides a method comprising:
[0011] In some embodiments, the present disclosure encompasses the recognition that ETaGs from eukaryotic fungi may have more similarity, if any, to mammalian genes than their counterparts in prokaryotes, such as certain bacteria, for example, In some embodiments, fungi contain more therapeutically relevant ETaGs than and / or than organisms that are evolutionarily more distant from humans.
[0012] In some embodiments, the present disclosure provides: interrogating a set of nucleic acid sequences, each of which is found in a fungal strain and which comprises a biosynthetic gene cluster; Among at least one of the fungal nucleic acid sequences, within a proximity zone to at least one gene in the cluster, is homologous to the mammalian nucleic acid sequence to be expressed, identifying an embedded target gene (ETaG) sequence that is optionally co-regulated with at least one biosynthetic gene in the cluster; The present invention provides a method comprising:
[0013] In some embodiments, the present disclosure provides: interrogating a set of nucleic acid sequences, each of which is found in a fungal strain and which comprises a biosynthetic gene cluster; Among at least one of the fungal nucleic acid sequences, within a proximity zone to at least one biosynthetic gene in the cluster; is homologous to the mammalian nucleic acid sequence to be expressed, identifying an embedded target gene (ETaG) sequence that is optionally co-regulated with at least one biosynthetic gene in the cluster; The present invention provides a method comprising:
[0014] In some embodiments, the proximal zone is 1 to 100 or less, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, or 100 kb or less upstream or downstream of a biosynthetic gene in a cluster. In some embodiments, the proximal zone is 1 to 100 or less, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 kb or less upstream or downstream of a biosynthetic gene in a cluster. In some embodiments, ETaG is within a biosynthetic gene cluster. In some embodiments, the proximal zone is between two biosynthetic genes in a biosynthetic gene cluster.
[0015] In some embodiments, the ETaG sequence is homologous to a mammalian nucleic acid sequence. In some embodiments, the mammalian sequence is a human nucleic acid sequence. In some embodiments, the ETaG sequence is homologous to a human nucleic acid sequence. In some embodiments, the ETaG sequence is homologous to an expressed mammalian nucleic acid sequence. In some embodiments, the ETaG sequence is homologous to an expressed human nucleic acid sequence. In some embodiments, the mammalian nucleic acid, e.g., a human nucleic acid sequence, is associated with a human disease, disorder, or condition. In some embodiments, such a human nucleic acid sequence is an existing target of therapeutic interest. In some embodiments, such a human nucleic acid sequence is a novel target of therapeutic interest. In some embodiments, such a human nucleic acid sequence is a target previously deemed insusceptible to targeting, e.g., by small molecules. In some embodiments, a biosynthetic product produced by an enzyme encoded by an associated biosynthetic gene cluster, or an analog thereof, is a modulator (e.g., activator, inhibitor, etc.) of the human target.
[0016] In some embodiments, the ETaG sequence is homologous to the expressed mammalian nucleic acid sequence in that the sequence, or a portion thereof, is at least 50%, 60%, 70%, 80%, or 90% identical to that of the expressed mammalian nucleic acid sequence. In some embodiments, the ETaG sequence is homologous to the mammalian nucleic acid sequence in that the mRNA produced from ETaG, or a portion thereof, is homologous to that of the mammalian nucleic acid sequence. In some embodiments, the homologous portion is at least 50, 100, 150, or 200 base pairs in length. In some embodiments, the homologous portion encodes a protein or a conserved portion of a protein, such as a protein domain, a set of residues involved in a function (e.g., interaction with another molecule (e.g., protein, small molecule, etc.), enzymatic activity, etc.), that is conserved from fungi to mammals.
[0017] In some embodiments, the ETaG sequence is homologous to a mammalian nucleic acid sequence in that the product encoded by ETaG, or a portion thereof, is homologous to that encoded by the mammalian nucleic acid sequence. In some embodiments, the ETaG sequence is homologous to a mammalian nucleic acid sequence in that the protein encoded by ETaG, or a portion thereof, is homologous to that encoded by the mammalian nucleic acid sequence. In some embodiments, the ETaG sequence is homologous to a mammalian nucleic acid sequence in that a portion of the protein encoded by ETaG is homologous to that encoded by the mammalian nucleic acid sequence.
[0018] In some embodiments, the portion of the protein is a protein domain. In some embodiments, the protein domain is an enzyme domain. In some embodiments, the protein domain interacts with one or more agents, such as small molecules, lipids, carbohydrates, nucleic acids, proteins, etc.
[0019] In some embodiments, portions of a protein are functional and / or structural domains that define the protein family to which the protein belongs. Within the specific catalytic or structural domains that define a patent family, there are amino acid residues that can be selected based on predicted subfamily domain architecture and optionally verified by various assays for use in homology alignment analysis.
[0020] In some embodiments, the portion of the protein is a set of key residues, contiguous or non-contiguous, that are important for the function of the protein. In some embodiments, the function is enzymatic activity and the portion of the protein is a set of residues required for the activity. In some embodiments, the function is enzymatic activity and the portion of the protein is a set of residues that interact with a substrate, intermediate, or product. In some embodiments, the set of residues interacts with a substrate. In some embodiments, the set of residues interacts with an intermediate. In some embodiments, the set of residues interacts with a product.
[0021] In some embodiments, the function is interaction with one or more agents, e.g., small molecules, lipids, carbohydrates, nucleic acids, proteins, etc., and the portion of the protein is a set of residues required for the interaction. In some embodiments, each set of residues independently contacts an interacting agent. For example, in some embodiments, each of the residues of the set independently contacts an interacting small molecule. In some embodiments, the protein is a kinase, the interacting small molecule is or includes a nucleobase, and each set of residues independently contacts the nucleobase via, e.g., hydrogen bonding, electrostatic forces, van der Waals forces, aromatic stacking, etc. In some embodiments, the interacting agent is another macromolecule. In some embodiments, the interaction agent is a nucleic acid. In some embodiments, the set of residues is a set that contacts an interacting nucleic acid, e.g., a set in a transcription factor. In some embodiments, the set of residues is a set that contacts an interacting protein.
[0022] In some embodiments, the portion of the protein is or comprises an essential structural element for protein effector recruitment and / or binding, for example, based on the tertiary protein structure of the human target.
[0023] Portions of proteins, such as protein domains, sets of residues responsible for biological function, may be conserved from species to species, e.g., in some embodiments, from fungi to humans, as described in this disclosure.
[0024] In some embodiments, protein homology is measured based on exact identity, e.g., the same amino acid residue at a given position. In some embodiments, homology is measured based on amino acid residues having one or more properties, e.g., one or more identical or similar properties (e.g., polar, non-polar, hydrophobic, hydrophilic, size, acidic, basic, aromatic, etc.). Exemplary methods for assessing homology are widely known in the art and can be utilized in accordance with the present disclosure, e.g., using MUSCLE, TCoffee, ClustalW, etc.
[0025] In some embodiments, the protein encoded by ETaG, or a portion thereof (e.g., as described in this disclosure), is at least 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or 100% (identical in the case of 100%) homologous to that encoded by the mammalian nucleic acid sequence. In some embodiments, the protein encoded by ETaG, or a portion thereof, is at least 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or 100% homologous to that encoded by the expressed mammalian nucleic acid sequence.
[0026] In some embodiments, ETaG is co-regulated with at least one biosynthetic gene in the biosynthetic gene cluster. In some embodiments, ETaG is co-regulated with two or more genes in the biosynthetic gene cluster. In some embodiments, ETaG is co-regulated with the biosynthetic gene cluster in that expression of ETaG is increased or turned on when a biosynthetic product produced by an enzyme encoded by the biosynthetic gene cluster (a biosynthetic product of the biosynthetic gene cluster) is produced. In some embodiments, ETaG is co-regulated with the biosynthetic gene cluster in that expression of ETaG is increased or turned on when the level of a biosynthetic product of the biosynthetic gene cluster is increased.
[0027] In some embodiments, the ETaG gene sequence is optionally greater than about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99% homologous to one or more gene sequences in the same genome. In some embodiments, the ETaG gene sequence is optionally greater than about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99% homologous to two, three, four, five, six, seven, eight, nine, or more gene sequences in the same genome. In some embodiments, the homology is greater than 10%. In some embodiments, the homology is greater than 20%. In some embodiments, the homology is greater than 30%. In some embodiments, the homology is greater than 40%. In some embodiments, the homology is greater than 50%. In some embodiments, the homology is greater than 60%. In some embodiments, the homology is greater than 70%. In some embodiments, the homology is greater than 80%. In some embodiments, the homology is greater than 90%. Certain examples are provided in the figures.
[0028] In some embodiments, the ETaG gene sequence is about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99% identical to any expressed gene sequence in at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% fungal nucleic acid sequences in a set optionally derived from different fungal strains and comprising homologous biosynthetic gene clusters. In some embodiments, the ETaG gene sequence is about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99% identical to at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% fungal gene sequence, as appropriate, within a proximity zone to biosynthetic genes of a homologous biosynthetic gene cluster from a different fungal strain. In some embodiments, the ETaG gene sequence is about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99% identical to at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% fungal gene sequence, as appropriate, within a proximity zone to biosynthetic genes of a homologous biosynthetic gene cluster from a different fungal strain. In some embodiments, the ETaG gene sequence is about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99% or less identical to any expressed gene sequence in any fungal nucleic acid sequence in the set comprising a homologous biosynthetic gene cluster, optionally from a different fungal strain.In some embodiments, the ETaG gene sequence is no more than about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99% identical to any expressed gene sequence within the proximity zone to biosynthetic genes of a homologous biosynthetic gene cluster from a different fungal strain, as appropriate. In some embodiments, it is no more than about 10% identical. In some embodiments, it is no more than about 20% identical. In some embodiments, it is no more than about 30% identical. In some embodiments, it is no more than about 40% identical. In some embodiments, it is no more than about 50% identical. In some embodiments, it is no more than about 60% identical. In some embodiments, it is no more than about 70% identical. In some embodiments, it is no more than about 80% identical. In some embodiments, it is no more than about 90% identical.
[0029] In some embodiments, the human target gene and / or its product is sensitive to modulation by a biosynthetic product of a biosynthetic gene cluster or its analog, and the human target gene has its homologous ETaG within the biosynthetic gene cluster or in a proximal zone to a biosynthetic gene of the cluster. In some embodiments, the protein encoded by the human target gene is sensitive to modulation by a biosynthetic product of a biosynthetic gene cluster or its analog, and the human target gene has its homologous ETaG within the biosynthetic gene cluster or in a proximal zone to a biosynthetic gene of the cluster. Thus, in some embodiments, the present disclosure not only provides novel human targets, but also provides methods and agents for modulating such human targets.
[0030] In some embodiments, the present disclosure provides techniques, e.g., methods, databases, systems, etc., for identifying ETaG and / or its medical relevance, e.g., its therapeutic relevance. In some embodiments, the present disclosure provides databases, optionally with various annotations, structured for efficient identification, retrieval, use, etc. of ETaG, related biosynthetic gene clusters, related biosynthetic products of the biosynthetic gene clusters and / or analogs thereof, related homologous mammalian nucleic acid sequences (e.g., human genes), etc. Among other things, the present disclosure provides databases and / or sequences structured to improve, e.g., computational efficiency and / or accuracy for ETaG identification.
[0031] For example, in some embodiments, the provided database was constructed so that all biosynthetic gene clusters were identified and annotated. The nucleic acid sequences of these clusters were then computationally excised from the rest of the nucleic acid sequences in the fungal genome and databased. The resulting database of biosynthetic gene clusters was then used for ETaG searches. Notably, when hits in the ETaG search were identified using such a database, the hits were ETaGs because only sequences that were present in the biosynthetic clusters (or their adjacent zones) were searched. Separating biosynthetic gene cluster sequences from the entire genome sequence improves the signal-to-noise ratio and greatly speeds up the ETaG search process. Notably, compared to using the provided database, ETaG searches in the entire fungal genome sequence frequently yielded false positives, and the identified hits were "housekeeping" genes located in the genome rather than in the biosynthetic gene clusters or their adjacent zones. In some embodiments, the hits, e.g., ETaGs, identified from the provided technology (e.g., methods, databases, etc.) are not housekeeping genes. In some embodiments, a hit identified from the provided technology, e.g., ETaG, is or comprises a sequence that shares homology with a second nucleic acid sequence (e.g., a gene) or portion thereof in the same genome. The sequence homology of the sequences in the present disclosure can be at least 20%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5%. In some embodiments, the homology is at least 50%; in some embodiments, at least 60%; in some embodiments, at least 70%; in some embodiments, at least 75%; in some embodiments, at least 80%; in some embodiments, at least 85%; in some embodiments, at least 90%; and in some embodiments, at least 95%.A portion of a sequence of the present disclosure can comprise at least 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, or 1000 amino acid residues (in the case of a protein sequence) or nucleobases (in the case of a nucleic acid sequence). In some embodiments, a portion of a nucleic acid sequence is at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, or 1000 nucleobases in length. In some embodiments, the length is at least 20 nucleobases. In some embodiments, the length is at least 30 nucleobases. In some embodiments, the length is at least 40 nucleobases. In some embodiments, the length is at least 50 nucleobases. In some embodiments, the length is at least 100 nucleobases. In some embodiments, the length is at least 150 nucleobases. In some embodiments, the length is at least 200 nucleobases. In some embodiments, the length is at least 300 nucleobases. In some embodiments, the length is at least 400 nucleobases. In some embodiments, the length is at least 500 nucleobases. In some embodiments, a hit identified from the provided techniques, e.g., ETaG, is or comprises a sequence encoding a product, e.g., a protein, that shares homology (e.g., at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5%) with a product, or portion thereof (e.g., a set of key residues of a protein, a protein domain, etc., as described herein) encoded by a second nucleic acid sequence (e.g., a gene) in the same genome. As described herein, homology / similarity can be assessed using a variety of techniques, as recognized by those of skill in the art. In some embodiments, the second nucleic acid sequence is or comprises a housekeeping gene. In some embodiments, the second nucleic acid sequence is shared between two or more species.In some embodiments, ETaG is homologous to but distinct from a second nucleic acid sequence in that ETaG encodes, but does not encode, a product (e.g., a protein) that confers resistance to a product (e.g., a small molecule) of its corresponding biosynthetic cluster.
[0032] In some embodiments, the present disclosure provides: one or more non-transitory machine-readable storage media storing data representing a set of nucleic acid sequences, each found in a fungal strain and comprising a biosynthetic gene cluster; The present invention provides a system including:
[0033] In some embodiments, the present disclosure provides: One or more non-transitory machine-readable storage media storing data representing a set of nucleic acid sequences, each of which is or includes an ETaG sequence. The present invention provides a system including:
[0034] In some embodiments, at least 10, 20, 50, 100, 200, or 500 of the nucleic acid sequences in the set, or at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%, or all, contain indexed and / or annotated ETaGs. In some embodiments, the provided systems are structured to significantly reduce the amount of data to be processed, thereby significantly improving computational efficiency. For example, instead of processing all genome or biosynthetic gene cluster sequence data for one or more (possibly hundreds or thousands or even more) fungal genomes to search for ETaGs, the provided systems search only for genes indexed / marked as ETaGs, thereby saving the time and cost that would be otherwise spent processing sequences that are not indexed as ETaGs. Additionally or alternatively, an ETaG can be independently annotated with information such as its related biosynthetic gene cluster (which contains the biosynthetic gene within its proximal zone), the structure of the biosynthetic product of the related biosynthetic gene cluster, and / or the human homolog of the ETaG. In some embodiments, at least 10, 20, 50, 100, 200, or 500 species or at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% or all of the ETaGs in the set are independently annotated with at least one of the following: the related biosynthetic gene cluster, and the human homolog of the ETaG. In some embodiments, at least 10, 20, 50, 100, 200, or 500, or at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%, or all, of the set of ETaGs are independently annotated with at least one of the following: related biosynthetic gene clusters, biosynthetic products of related biosynthetic gene clusters, and human homologs of ETaGs. In some embodiments, by structuring sequence data with ETaG indexes and annotations, the provided systems can provide a number of advantages.For example, in some embodiments, the systems provided provide rapid access to useful related information, e.g., ETaGs with their related biosynthetic gene clusters and human homologs (and vice versa), while keeping data size and costs low.
[0035] In some embodiments, the provided methods and systems are useful for human target identification and / or characterization, particularly because the provided methods and systems provide a link between the biosynthetic gene cluster, ETaG, and the human target gene. In some embodiments, the present disclosure provides insight into targets that were previously deemed non-druggable, particularly by providing its homologous ETaG and related biosynthetic gene cluster in fungi. In some embodiments, the present disclosure significantly improves the druggability of targets previously deemed non-druggable, in some cases essentially converting them into druggable targets, for example, by the homologous ETaG in fungi, the related biosynthetic gene cluster, or the biosynthetic products of the related biosynthetic gene cluster (which can be used directly as modulators of the human target and / or whose analogs can be used as modulators).
[0036] In some embodiments, the disclosure provides methods for identifying and / or characterizing human targets of biosynthetic products of biosynthetic gene clusters, or analogs of the products.
[0037] In some embodiments, the present disclosure provides: identifying a human homolog of ETaG that is within a proximity zone to at least one gene of a biosynthetic gene cluster or that is within a proximity zone to at least one gene of a second biosynthetic gene cluster, the second biosynthetic gene cluster encoding an enzyme that produces a biosynthetic product produced by an enzyme encoded by the biosynthetic gene cluster; optionally assaying the effect of the biosynthetic products, or analogs of the products, produced by the enzymes encoded by the biosynthetic gene cluster on the human target; The present invention provides a method comprising:
[0038] In some embodiments, the present disclosure provides: identifying a human homolog of ETaG that is within a proximity zone to at least one biosynthetic gene of the biosynthetic gene cluster or that is within a proximity zone to at least one biosynthetic gene of a second biosynthetic gene cluster, the second biosynthetic gene cluster encoding an enzyme that produces a biosynthetic product produced by the enzyme encoded by the biosynthetic gene cluster; optionally assaying the effect of the biosynthetic products, or analogs of the products, produced by the enzymes encoded by the biosynthetic gene cluster on the human target; The present invention provides a method comprising:
[0039] In some embodiments, the present disclosure provides: identifying a human homolog of ETaG within a proximity zone to at least one gene of the biosynthetic gene cluster; optionally assaying the effect of the biosynthetic products, or analogs of the products, produced by the enzymes encoded by the biosynthetic gene cluster on the human target; The present invention provides a method comprising:
[0040] In some embodiments, the present disclosure provides: identifying a human homolog of ETaG within a proximity zone to at least one biosynthetic gene of the biosynthetic gene cluster; and optionally assaying the effect of the biosynthetic products, or analogs of the products, produced by the enzymes encoded by the biosynthetic gene cluster on the human target.
[0041] In some embodiments, for a biosynthetic gene cluster whose proximity zone does not include a biosynthetic gene containing an ETaG, the mammalian target, e.g., human target, of the product (and / or analog thereof) of such biosynthetic gene cluster can be identified via an ETaG present in the proximity zone to a biosynthetic gene of a second biosynthetic gene cluster that encodes an enzyme that produces the same biosynthetic product. In some embodiments, the second biosynthetic gene cluster is present in a different organism. In some embodiments, the second biosynthetic gene cluster is present in a different fungal strain.
[0042] In some embodiments, the disclosure provides a method for identifying and / or characterizing a human target of a biosynthetic product of a biosynthetic gene cluster, or an analog of the product, comprising: identifying a human homolog of ETaG within a proximity zone to at least one biosynthetic gene of a second biosynthetic gene cluster, the second biosynthetic gene cluster encoding an enzyme that produces the same biosynthetic product produced by the enzyme encoded by the biosynthetic gene cluster; optionally assaying the effect of the biosynthetic products, or analogs of the products, produced by the enzymes encoded by the biosynthetic gene cluster on the human target; The present invention provides a method comprising:
[0043] In some embodiments, the provided techniques are useful for assessing the interaction of a compound with a human target. In some embodiments, the present disclosure provides a method for accessing the interaction of a compound with a human target, the method comprising: comparing the nucleic acid sequence of, or encoding, the human target to a set of nucleic acid sequences comprising one or more ETaGs; The present invention provides a method comprising:
[0044] In some embodiments, homology to ETaG (including portions thereof, at the nucleic acid or protein level) leads to a related biosynthetic gene cluster for ETaG and its biosynthetic products. In some embodiments, such an association between a biosynthetic product and a human target indicates an interaction and / or modulation of the human target or the product encoded thereby. In some embodiments, such a biosynthetic product interacts with and / or modulates the human target or the product encoded thereby.
[0045] In some embodiments, the provided technology is useful for designing and / or providing modulators for human targets, since, among other things, the provided technology provides the connection between the biosynthetic gene cluster, ETaG, and human target genes.
[0046] In some embodiments, the present disclosure provides a compound that is a product of an enzyme encoded by a biosynthetic gene cluster, wherein the compound is present within a proximity zone to at least one gene in the biosynthetic gene cluster, the proximity zone being: is homologous to a human target or a nucleic acid sequence encoding a human target, Compounds are provided in which there is at least one biosynthetic gene in a cluster and optionally co-regulated ETaG.
[0047] In some embodiments, provided compounds are products of enzymes encoded by provided biosynthetic gene clusters. In some embodiments, provided compounds are analogs of products of enzymes encoded by provided biosynthetic gene clusters. In some embodiments, provided biosynthetic gene clusters comprise one or more biosynthetic genes presented in one of Figures 5-12 and 20-39. In some embodiments, provided biosynthetic gene clusters are one of Figures 5-12 and 20-39. In some embodiments, provided compounds are products of enzymes encoded by provided biosynthetic gene clusters presented in one of Figures 5-12 and 20-39. In some embodiments, provided compounds are products of enzymes encoded by provided biosynthetic gene clusters presented in one of Figures 5-12 and 20-39, or biosynthetic gene clusters comprising one or more biosynthetic genes presented in one of Figures 5-12 and 20-39. In some embodiments, provided compounds are products of provided biosynthetic gene clusters presented in one of Figures 5-12 and 20-39. In some embodiments, provided compounds are products of provided biosynthetic gene clusters comprising one or more biosynthetic genes presented in one of Figures 5-12 and 20-39. In some embodiments, provided compounds are analogs of products of enzymes encoded by provided biosynthetic gene clusters presented in one of Figures 5-12 and 20-39. In some embodiments, provided compounds are analogs of products of provided biosynthetic gene clusters comprising one or more biosynthetic genes presented in one of Figures 5-12 and 20-39. In some embodiments, provided compounds modulate the function of a human target. In some embodiments, the disclosure provides pharmaceutical compositions of provided compounds. In some embodiments, the disclosure provides pharmaceutical compositions comprising a provided compound, or a pharmaceutically acceptable salt thereof. In some embodiments, the disclosure provides pharmaceutical compositions comprising a provided compound, or a pharmaceutically acceptable salt thereof, and a pharmaceutically acceptable carrier.In some embodiments, the provided compounds in the provided compositions are products of enzymes encoded by biosynthetic gene clusters or salt analogs thereof, hi some embodiments, the provided compounds in the provided compositions are non-naturally occurring salts of products of enzymes encoded by biosynthetic gene clusters.
[0048] In some embodiments, the disclosure provides a method for identifying and / or characterizing a modulator of a human target, comprising: providing a product or analog thereof, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster, and within a proximity zone to at least one gene in the biosynthetic gene cluster: is homologous to a human target or a nucleic acid sequence encoding a human target, The present invention provides a method comprising the steps of: wherein there is at least one biosynthetic gene in a cluster and, optionally, co-regulated ETaG.
[0049] In some embodiments, the disclosure provides a method for identifying and / or characterizing a modulator of a human target, comprising: providing a product or analog thereof, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster, and within a proximity zone to at least one biosynthetic gene in the biosynthetic gene cluster: is homologous to a human target or a nucleic acid sequence encoding a human target, The present invention provides a method comprising the steps of: wherein there is at least one biosynthetic gene in a cluster and, optionally, co-regulated ETaG.
[0050] In some embodiments, the present disclosure provides a method for modulating a human target, comprising: providing a product or analog thereof, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster, and within a proximity zone to at least one gene in the biosynthetic gene cluster: is homologous to a human target or a nucleic acid sequence encoding a human target, The present invention provides a method comprising the steps of: wherein there is at least one biosynthetic gene in a cluster and, optionally, co-regulated ETaG.
[0051] In some embodiments, the present disclosure provides a method for modulating a human target, comprising: providing a product or analog thereof, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster, and within a proximity zone to at least one biosynthetic gene in the biosynthetic gene cluster: is homologous to a human target or a nucleic acid sequence encoding a human target, The present invention provides a method comprising the steps of: wherein there is at least one biosynthetic gene in a cluster and, optionally, co-regulated ETaG.
[0052] In some embodiments, the disclosure provides a method for treating a condition, disorder, or disease associated with a human target, comprising administering to a subject susceptible to or suffering from the same a biosynthetic product or analog thereof; The biosynthetic product is of a biosynthetic gene cluster, and within a proximity zone to at least one gene in the biosynthetic gene cluster: is homologous to a human target or a nucleic acid sequence encoding a human target, The method provides that there is at least one biosynthetic gene in the cluster and optionally co-regulated ETaG.
[0053] In some embodiments, the disclosure provides a method for treating a condition, disorder, or disease associated with a human target, comprising administering to a subject susceptible to or suffering from the same a biosynthetic product or analog thereof; The biosynthetic product is of a biosynthetic gene cluster, and within a proximity zone to at least one biosynthetic gene in the biosynthetic gene cluster, is homologous to a human target or a nucleic acid sequence encoding a human target, The method provides that there is at least one biosynthetic gene in the cluster and optionally co-regulated ETaG.
[0054] In some embodiments, the human target is a Ras protein. In some embodiments, the human target comprises a RasGEF domain. In some embodiments, the human target comprises a RasGAP domain.
[0055] In some embodiments, the ETaG is identified by the methods provided.
[0056] In some embodiments, the product (e.g., a biosynthetic product) is produced by a fungus. In some embodiments, the product is acyclic. In some embodiments, the product is a polyketide. In some embodiments, the product is a terpene compound. In some embodiments, the product is synthesized non-ribosomally.
[0057] In some embodiments, an analog is a substance that shares one or more specific structural features, elements, components, or moieties with a reference substance. Typically, an analog exhibits significant structural similarity with the reference substance, e.g., shares a core or consensus structure, but also differs in certain distinct ways. In some embodiments, an analog is a substance that can be produced from a reference substance, e.g., by chemical manipulation of the reference substance. In some embodiments, an analog is a substance that can be produced by performing a synthetic process that is substantially similar to (e.g., shares multiple steps with) the process that produces the reference substance. In some embodiments, an analog is or can be produced by performing a synthetic process that is different from the process used to produce the reference substance. In some embodiments, an analog of a substance is a substance that is substituted at one or more of its substitutable positions.
[0058] In some embodiments, the analog of the product comprises the structural core of the product. In some embodiments, the biosynthetic product is cyclic, e.g., monocyclic, bicyclic, or polycyclic, and the structural core of the product is or comprises a monocyclic, bicyclic, or polycyclic ring system. In some embodiments, the product is or comprises a polypeptide, and the structural core is the backbone of the polypeptide. In some embodiments, the product is or comprises a polyketide, and the structural core is the backbone of the polyketide.
[0059] In some embodiments, the analog is a substituted biosynthetic product, hi some embodiments, the analog is or comprises a structural core substituted with one or more substituents described herein.
[0060] In some embodiments, the present disclosure provides compositions of biosynthetic products of a provided biosynthetic gene cluster or analog thereof, wherein ETaG is present within a proximity zone relative to at least one gene of the biosynthetic gene cluster. In some embodiments, the provided compositions are pharmaceutical compositions. In some embodiments, the provided pharmaceutical compositions comprise a pharmaceutically acceptable salt of a biosynthetic product of a provided biosynthetic gene cluster or analog thereof and a pharmaceutically acceptable carrier, wherein ETaG is present within a proximity zone relative to at least one gene of the biosynthetic gene cluster.
[0061] In some embodiments, two events or entities are associated with one another if the presence, level, and / or form of one correlates with that of the other. For example, a particular entity (e.g., a polypeptide, genetic signature, metabolite, microorganism, etc.) is considered to be associated with a particular disease, disorder, or condition if its presence, level, and / or form correlates with the incidence and / or susceptibility of the disease, disorder, or condition (e.g., across a relevant population).
[0062] In some embodiments, the disease is cancer. In some embodiments, the disease is an infectious disease. In some embodiments, the disease is heart disease. In some embodiments, the disease is associated with levels of lipids, proteins, human metabolites, etc. [Brief explanation of the drawings]
[0063] [Figure 1] Figure 1 depicts the brefeldin A ETaGs identified in Penicillium vulpinum IBT 29486. The identified example ETaG is the Sec7 guanine nucleotide exchange factor superfamily (pfam01369). Sequence similarity is that of the Sec7 domain, calculated using the MUSCLE alignment algorithm.
[0064] [Figure 2] Figure 2 depicts the lovastatin ETaG identified in Aspergillus terreus ATCC 20542. The identified example ETaG is hydroxymethylglutaryl-coenzyme A reductase (HMG-CoA; pfam00368). Sequence similarity is of the HMG-CoA domain, calculated using the MUSCLE alignment algorithm.
[0065] [Figure 3] Figure 3 depicts the Fellutamide ETaG identified in Aspergillus nidulans FGSC A4. The identified ETaG example is the proteasome 20S beta-subunit (pfam00227). Sequence similarity is calculated using the MUSCLE alignment algorithm for 20S beta-.
[0066] [Figure 4]Figure 4 depicts the cyclosporine ETaG identified in Tolypocladium inflatum NRRL 8044. The identified ETaG example is a cyclophilin-type peptidylprolyl cis-trans isomerase (pfam00160). Sequence similarity is that of the cyclophilin domain, calculated using the MUSCLE alignment algorithm.
[0067] [Figure 5] Figure 5 depicts the Ras ETaGs identified in Thermomyces lanuginosus ATCC 200065 (public). The identified ETaG example is from the Ras family (pfam00071). Sequence similarity is that of the Ras domain, calculated using the MUSCLE alignment algorithm. The ETaGs are presented below the scale.
[0068] [Figure 6] Figure 6 depicts the Ras ETaGs identified in Talaromyces leycettanus strain CBS 398.68. The identified ETaG example is from the Ras family (pfam00071). Sequence similarity is that of the Ras domain, calculated using the MUSCLE alignment algorithm. The ETaGs are presented below the scale.
[0069] [Figure 7] Figure 7 depicts the Ras ETaGs identified in Sistotremastrum niveocremeum HHB9708 or Sistotremastrum suecicum HHB10207 (National Forestry Service). The identified ETaG examples are from the Ras family (pfam00071). Sequence similarity is that of the Ras domain, calculated using the MUSCLE alignment algorithm. The ETaGs are presented below the scale.
[0070] [Figure 8] Figure 8 depicts the Ras ETaGs identified in Agaricus bisporus var. burnettii JB137-S8 (Fungal Genome Stock Center). The identified ETaG example is from the Ras family (pfam00071). Sequence similarity is that of the Ras domain, calculated using the MUSCLE alignment algorithm. The ETaGs are presented below the scale.
[0071] [Figure 9] Figure 9 depicts the Ras ETaGs identified in Coprinopsis cinerea okayama 7#130 (Fungal Genome Stock Center). The identified ETaG example is from the Ras family (pfam00071). Sequence similarity is that of the Ras domain, calculated using the MUSCLE alignment algorithm. The ETaGs are presented below the scale.
[0072] [Figure 10] Figure 10 depicts the Ras ETaGs identified in Colletotrichum higginsianum IMI 349063 (CABI). The identified ETaG example is from the Ras family (pfam00071). Sequence similarity is that of the Ras domain, calculated using the MUSCLE alignment algorithm. The ETaGs are presented below the scale.
[0073] [Figure 11] Figure 11 depicts the Ras ETaGs identified in Gyalolechia flavorubescens KoLRI002931. The identified ETaG example is from the Ras family (pfam00071). Sequence similarity is that of the Ras domain, calculated using the MUSCLE alignment algorithm. The ETaGs are presented below the scale.
[0074] [Figure 12] Figure 12 depicts the Ras ETaGs identified in Bipolaris maydis ATCC 48331. The identified ETaG example is from the Ras family (pfam00071). Sequence similarity is that of the Ras domain, calculated using the MUSCLE alignment algorithm. The ETaGs are presented below the scale.
[0075] [Figure 13] Figure 13 depicts an alignment of the human Ras gene and certain identified Ras ETaGs. As shown, the human Ras gene and the proposed ETaGs share the same amino acid residues at many positions of the KRAS nucleotide binding residues.
[0076] [Figure 14] Figure 14 depicts the alignment of the human Ras gene and certain identified Ras ETaGs. As shown, the human Ras gene and the proposed ETaGs share the same amino acid residues at many positions of KRAS residues that are within 4 Å of BRAF.
[0077] [Figure 15] Figure 15 depicts the alignment of the human Ras gene and certain identified Ras ETaGs. As shown, the human Ras gene and the proposed ETaG share the same amino acid residues at many positions of the KRAS residues that are within 4 Å of the rasGAP.
[0078] [Figure 16] Figure 16 depicts the alignment of the human Ras gene and certain identified Ras ETaGs. As shown, the human Ras gene and the proposed ETaG share the same amino acid residues at many positions of KRAS residues that are within 4 Å of the SOS.
[0079] [Figure 17]FIG. 17 depicts an example sequence in which the ETaG is indexed / marked (dark).
[0080] [Figure 18] FIG. 18 depicts the biosynthetic gene cluster containing the Sec7 homologue in Penicillium vulpinum IBT 29486.
[0081] [Figure 19-1] Figure 19 depicts a sequence alignment of Sec7. (A) Example brefeldin A interacting residues. (B) Example sequence alignment. [Figure 19-2] Figure 19 depicts a sequence alignment of Sec7. (A) Example brefeldin A interacting residues. (B) Example sequence alignment.
[0082] [Figure 20] Figure 20 depicts example biosynthetic gene clusters related to Ras from, for example, Thermomyces lanuginosus ATCC 200065, Aspergillus rambelli, and Aspergillus ochraceoroseus. Illustrated Ras homologs are shown in black.
[0083] [Figure 21] Figure 21 depicts example biosynthetic gene clusters related to Ras from, for example, Agaricus bisporus var. burnettii JB137-S8, Agaricus bisporus H97, Coprinopsis cinerea okayama, and Hypholoma sublateritum FD-334. Illustrated Ras homologs are shown in black.
[0084] [Figure 22]Figure 22 depicts example biosynthetic gene clusters related to Ras from, for example, Sistotremastrum niveocremeum HHB9708 and Sistotremastrum suecicum HHB10207. Illustrated Ras homologs are shown in black.
[0085] [Figure 23] Figure 23 depicts an example biosynthetic gene cluster related to Ras, for example, from Talaromyces leycettanus strain CBS 398.68. The illustrated Ras homolog is shown in black.
[0086] [Figure 24] Figure 24 depicts an example biosynthetic gene cluster related to Ras from, for example, Thermoascus crustaceus. The illustrated Ras homolog is shown in black.
[0087] [Figure 25] Figure 25 depicts an example biosynthetic gene cluster related to Ras, for example, from Bipolaris maydis ATCC 48331. The illustrated Ras homolog is shown in black.
[0088] [Figure 26] Figure 26 depicts an example biosynthetic gene cluster related to Ras, for example, from Colletotrichum higginsianum IMI 349063 (CABI). The illustrated Ras homolog is shown in black.
[0089] [Figure 27] Figure 27 depicts an example biosynthetic gene cluster related to Ras from, for example, Gyalolechia flavorubescens. The illustrated Ras homolog is shown in black.
[0090] [Figure 28] Figure 28 depicts example biosynthetic gene clusters related to RasGEFs from, for example, Penicillium chrysogenum Wisconsin 54-1255 and Lecanosticta acicola CBS 871.95. Illustrated RasGEF homologs are shown in black.
[0091] [Figure 29] Figure 29 depicts an example biosynthetic gene cluster related to RasGEF, for example, from Magnaporthe oryzae 70-15. The illustrated RasGEF homolog is shown in black.
[0092] [Figure 30] Figure 30 depicts an example biosynthetic gene cluster related to RasGEF, for example, from Arthroderma gypseum CBS 118893. The illustrated RasGEF homolog is shown in black.
[0093] [Figure 31] Figure 31 depicts an example biosynthetic gene cluster related to RasGEF, for example, from Endocarpon pusillum strain KoLRI No. LF000583. The illustrated RasGEF homolog is shown in black.
[0094] [Figure 32] Figure 32 depicts an example biosynthetic gene cluster related to RasGEF, for example, from Fistulina hepatica ATCC 64428. The illustrated RasGEF homolog is shown in black.
[0095] [Figure 33]Figure 33 depicts an example biosynthetic gene cluster related to RasGEF from, for example, Aureobasidium pullulans var. pullulans EXF-150. The illustrated RasGEF homolog is shown in black.
[0096] [Figure 34] Figure 34 depicts an example biosynthetic gene cluster related to RasGAP, for example, from Acremonium furcatum var. pullulans EXF-150. The illustrated RasGAP homolog is shown in black.
[0097] [Figure 35] Figure 35 depicts example biosynthetic gene clusters related to RasGEFs from, for example, Purpureocillium lilacinum strain TERIBC 1 and Fusarium sp. JS1030. Illustrated RasGEF homologs are shown in black.
[0098] [Figure 36] Figure 36 depicts example biosynthetic gene clusters related to RasGAPs from, for example, Corynespora cassiicola UM 591 and Magnaporthe oryzae strain SV9610. Illustrated RasGAP homologs are shown in black.
[0099] [Figure 37] Figure 37 depicts an example biosynthetic gene cluster related to RasGAP from, for example, Colletotrichum acutatum strain 1 KC05_01. The illustrated RasGAP homolog is shown in black.
[0100] [Figure 38]Figure 38 depicts example biosynthetic gene clusters related to RasGAPs from, for example, Hypoxylon sp. E7406B and Diaporthe ampelina isolate DA912. Illustrated RasGAP homologs are shown in black.
[0101] [Figure 39] Figure 39 depicts example biosynthetic gene clusters related to RasGAPs from, for example, Talaromyces piceae strain 9-3 and Sporothrix insectorum RCEF 264. Illustrated RasGAP homologs are shown in black. DETAILED DESCRIPTION OF THE INVENTION
[0102] 1.Definition As used herein, the following definitions shall apply unless otherwise indicated: For the purposes of this disclosure, chemical elements are defined as defined in the Periodic Table of the Elements, CAS version, Handbook of Chemistry and Physics, 75 th Furthermore, the general principles of organic chemistry are identified according to Ed. "Organic Chemistry", Thomas Sorrell, University Science Books, Sausalito: 1999 and "March's Advanced Organic Chemistry", 5 th Ed., Ed.: Smith, M.B. and March, J., John Wiley & Sons, New York: 2001.
[0103] Aliphatic: As used herein, "aliphatic" means a straight-chain (i.e., unbranched) or branched, substituted or unsubstituted hydrocarbon chain that is fully saturated or contains one or more units of unsaturation, or a substituted or unsubstituted monocyclic, bicyclic, or polycyclic hydrocarbon ring that is fully saturated or contains one or more units of unsaturation, or a combination thereof. Unless otherwise specified, an aliphatic group contains 1-100 aliphatic carbon atoms. In some embodiments, an aliphatic group contains 1-20 aliphatic carbon atoms. In other embodiments, an aliphatic group contains 1-10 aliphatic carbon atoms. In other embodiments, an aliphatic group contains 1-9 aliphatic carbon atoms. In other embodiments, an aliphatic group contains 1-8 aliphatic carbon atoms. In other embodiments, an aliphatic group contains 1-7 aliphatic carbon atoms. In other embodiments, an aliphatic group contains 1-6 aliphatic carbon atoms. In still other embodiments, aliphatic groups contain 1-5 aliphatic carbon atoms, and in yet other embodiments, aliphatic groups contain 1, 2, 3, or 4 aliphatic carbon atoms. Suitable aliphatic groups include, but are not limited to, linear or branched, substituted or unsubstituted alkyl, alkenyl, alkynyl groups, and hybrids thereof.
[0104] Alkyl: As used herein, the term "alkyl" is given its ordinary meaning in the art and can include saturated aliphatic groups, including straight-chain alkyl groups, branched-chain alkyl groups, cycloalkyl (alicyclic) groups, alkyl-substituted cycloalkyl groups, and cycloalkyl-substituted alkyl groups. In some embodiments, an alkyl has 1 to 100 carbon atoms. In certain embodiments, a straight-chain or branched-chain alkyl has about 1 to 20 carbon atoms in its backbone (e.g., C1 to C6 for a straight chain). 20 , C2 to C for branched chains 20), alternatively having about 1-10 carbon atoms. In some embodiments, cycloalkyl rings have from about 3-10 carbon atoms in their ring structure, and alternatively about 5, 6 or 7 carbons in the ring structure, whether such ring is monocyclic, bicyclic or polycyclic. In some embodiments, alkyl groups can be lower alkyl groups having from 1-4 carbon atoms (e.g., C1-C4 for a straight chain lower alkyl).
[0105] Aryl: The term "aryl," used alone or as part of a larger moiety, as in "aralkyl," "aralkoxy," or "aryloxyalkyl," refers to a monocyclic, bicyclic, or polycyclic ring system having a total of 5 to 30 ring members, wherein at least one ring in the system is aromatic. In some embodiments, an aryl group is a monocyclic, bicyclic, or polycyclic ring system having a total of 5 to 14 ring members, wherein at least one ring in the system is aromatic, and wherein each ring in the system contains 3 to 7 ring members. In some embodiments, an aryl group is a biaryl group. The term "aryl" can be used interchangeably with the term "aryl ring." In certain embodiments of the present disclosure, "aryl" refers to an aromatic ring system, including, but not limited to, phenyl, biphenyl, naphthyl, binaphthyl, anthracyl, and the like, which may bear one or more substituents. In some embodiments, the term "aryl" as used herein also includes within its scope groups in which an aromatic ring is fused to one or more non-aromatic rings, such as indanyl, phthalimidyl, naphthimidyl, phenanthridinyl, or tetrahydronaphthyl, etc., where the radical or point of attachment is on the aryl ring.
[0106] Alicyclic: As used herein, the term "alicyclic" refers to a saturated or partially unsaturated aliphatic monocyclic, bicyclic, or polycyclic ring system, e.g., having 3 to 30 members, where the aliphatic ring system is optionally substituted. Alicyclic groups include, without limitation, cyclopropyl, cyclobutyl, cyclopentyl, cyclopentenyl, cyclohexyl, cyclohexenyl, cycloheptyl, cycloheptenyl, cyclooctyl, cyclooctenyl, norbornyl, adamantyl, and cyclooctadienyl. In some embodiments, cycloalkyl has 3 to 6 carbons. The term "alicyclic" can also include an aliphatic ring fused to one or more aromatic or non-aromatic rings, such as decahydronaphthyl or tetrahydronaphthyl, where the radical or point of attachment is on the aliphatic ring. In some embodiments, a carbocyclic group is bicyclic. In some embodiments, a carbocyclic group is tricyclic. In some embodiments, a carbocyclic group is polycyclic. In some embodiments, "alicyclic" (or "carbocycle" or "cycloalkyl") refers to a monocyclic C3-C6 hydrocarbon, or a C8-C8 ring that is fully saturated or contains one or more units of unsaturation, but is not aromatic. 10 Bicyclic hydrocarbons, or C9-C, which are fully saturated or contain one or more units of unsaturation, but are not aromatic 16 Refers to tricyclic hydrocarbons.
[0107] Halogen: The term "halogen" means F, Cl, Br or I.
[0108] Heteroaliphatic: The term "heteroaliphatic" is given its ordinary meaning in the art and refers to an aliphatic group, as described herein, in which one or more carbon atoms are replaced by one or more heteroatoms (e.g., oxygen, nitrogen, sulfur, silicon, phosphorus, etc.).
[0109] Heteroalkyl: The term "heteroalkyl" is given its ordinary meaning in the art and refers to an alkyl group, as described herein, in which one or more carbon atoms are replaced by a heteroatom (e.g., oxygen, nitrogen, sulfur, silicon, phosphorus, etc.). Examples of heteroalkyl groups include, but are not limited to, alkoxy, poly(ethylene glycol)-, alkyl-substituted amino, tetrahydrofuranyl, piperidinyl, morpholinyl, and the like.
[0110] Heteroaryl: The terms "heteroaryl" and "heteroar-," used alone or as part of a larger moiety, e.g., "heteroaralkyl" or "heteroaralkoxy," refer to a monocyclic, bicyclic, or polycyclic ring system having, e.g., a total of 5 to 30 ring members, in which at least one ring in the system is aromatic and at least one aromatic ring atom is a heteroatom. In some embodiments, the heteroatom is nitrogen, oxygen, or sulfur. In some embodiments, heteroaryl groups are groups having 5 to 10 ring atoms (i.e., monocyclic, bicyclic, or polycyclic), in some embodiments, 5, 6, 9, or 10 ring atoms. In some embodiments, heteroaryl groups have 6, 10, or 14 pi electrons shared in the cyclic array; and in addition to the carbon atoms, have 1 to 5 heteroatoms. Heteroaryl groups include, without limitation, thienyl, furanyl, pyrrolyl, imidazolyl, pyrazolyl, triazolyl, tetrazolyl, oxazolyl, isoxazolyl, oxadiazolyl, thiazolyl, isothiazolyl, thiadiazolyl, pyridyl, pyridazinyl, pyrimidinyl, pyrazinyl, indolizinyl, purinyl, naphthyridinyl, and pteridinyl. In some embodiments, a heteroaryl is a heterobiaryl group, such as bipyridyl, etc. The terms "heteroaryl" and "heteroar-," as used herein, also include groups in which a heteroaromatic ring is fused to one or more aryl, alicyclic, or heterocyclyl rings, where the radical or point of attachment is on the heteroaromatic ring.Non-limiting examples include indolyl, isoindolyl, benzothienyl, benzofuranyl, dibenzofuranyl, indazolyl, benzimidazolyl, benzthiazolyl, quinolyl, isoquinolyl, cinnolinyl, phthalazinyl, quinazolinyl, quinoxalinyl, 4H-quinolizinyl, carbazolyl, acridinyl, phenazinyl, phenothiazinyl, phenoxazinyl, tetrahydroquinolinyl, tetrahydroisoquinolinyl, and pyrido[2,3-b]-1,4-oxazin-3(4H)-one. Heteroaryl groups can be monocyclic, bicyclic, or polycyclic. The term "heteroaryl" can be used interchangeably with the terms "heteroaryl ring," "heteroaryl group," or "heteroaromatic," all of which terms include rings that are optionally substituted. The term "heteroaralkyl" refers to an alkyl group substituted by a heteroaryl group, where the alkyl and heteroaryl portions independently are optionally substituted.
[0111] Heteroatom: The term "heteroatom" refers to an atom that is neither carbon nor hydrogen. In some embodiments, a heteroatom is oxygen, sulfur, nitrogen, phosphorus, boron, or silicon (an oxidized form of any of nitrogen, sulfur, phosphorus, or silicon; a basic or substitutable nitrogen of any of the heterocyclic rings (e.g., N as in 3,4-dihydro-2H-pyrrolyl), NH (as in pyrrolidinyl), or NR + (including quaternized forms of (as in N-substituted pyrrolidinyl); etc.). In some embodiments, the heteroatom is boron, nitrogen, oxygen, silicon, sulfur, or phosphorus. In some embodiments, the heteroatom is nitrogen, oxygen, silicon, sulfur, or phosphorus. In some embodiments, the heteroatom is nitrogen, oxygen, sulfur, or phosphorus. In some embodiments, the heteroatom is nitrogen, oxygen, or sulfur.
[0112] Heterocyclyl: As used herein, the terms "heterocycle," "heterocyclyl," "heterocyclic radical," and "heterocyclic ring" are used interchangeably and refer to a monocyclic, bicyclic, or polycyclic ring moiety (e.g., 3-30 membered) that is saturated or partially unsaturated and has one or more heteroatom ring atoms. In some embodiments, the heteroatom is boron, nitrogen, oxygen, silicon, sulfur, or phosphorus. In some embodiments, the heteroatom is nitrogen, oxygen, silicon, sulfur, or phosphorus. In some embodiments, the heteroatom is nitrogen, oxygen, or sulfur. In some embodiments, the heteroatom is nitrogen, oxygen, or sulfur. In some embodiments, the heterocyclyl group is a stable 5- to 7-membered monocyclic or 7- to 10-membered bicyclic heterocyclic moiety that is saturated or partially unsaturated and has, in addition to carbon atoms, one or more, preferably 1-4, heteroatoms as defined above. When used in reference to a heterocyclic ring atom, the term "nitrogen" includes substituted nitrogen. For example, in a saturated or partially unsaturated ring having 0-3 heteroatoms selected from oxygen, sulfur or nitrogen, the nitrogen can be N (as in 3,4-dihydro-2H-pyrrolyl), NH (as in pyrrolidinyl) or +It can be NR (as in N-substituted pyrrolidinyl). The heterocyclic ring can be attached to its pendant group at any heteroatom or carbon atom that results in a stable structure, and any of the ring atoms can be optionally substituted. Examples of such saturated or partially unsaturated heterocyclic radicals include, but are not limited to, tetrahydrofuranyl, tetrahydrothienyl, pyrrolidinyl, piperidinyl, pyrrolinyl, tetrahydroquinolinyl, tetrahydroisoquinolinyl, decahydroquinolinyl, oxazolidinyl, piperazinyl, dioxanyl, dioxolanyl, diazepinyl, oxazepinyl, thiazepinyl, morpholinyl, and quinuclidinyl. The terms "heterocycle," "heterocyclyl," "heterocyclyl ring," "heterocyclic group," "heterocyclic moiety," and "heterocyclic radical" are used interchangeably herein and also include groups in which a heterocyclyl ring is fused to one or more aryl, heteroaryl, or alicyclic rings, such as indolinyl, 3H-indolyl, chromanyl, phenanthridinyl, or tetrahydroquinolinyl, where the radical or point of attachment is on the heteroaliphatic ring. Heterocyclyl groups can be monocyclic, bicyclic, or polycyclic. The term "heterocyclylalkyl" refers to an alkyl group substituted by a heterocyclyl, where the alkyl and heterocyclyl moieties independently are optionally substituted.
[0113] Partially unsaturated: As used herein, the term "partially unsaturated" refers to a moiety that contains at least one double or triple bond. The term "partially unsaturated" is intended to encompass groups with multiple sites of unsaturation, but is not intended to include aryl or heteroaryl moieties.
[0114] Pharmaceutical composition: As used herein, the term "pharmaceutical composition" refers to an active agent formulated with one or more pharmaceutically acceptable carriers. In some embodiments, the active agent is present in a unit dose suitable for administration in a treatment regimen that, when administered to a relevant population, exhibits a statistically significant probability of achieving a predetermined therapeutic effect. In some embodiments, the pharmaceutical composition can be specially formulated for administration in solid or liquid form, including those adapted for oral administration, e.g., drenches (aqueous or non-aqueous solutions or suspensions), tablets, e.g., tablets targeted for buccal, sublingual, and systemic absorption, boluses, powders, granules, pastes for application to the tongue; parenteral administration, e.g., by subcutaneous, intramuscular, intravenous, or epidural injection, e.g., as a sterile solution or suspension, or sustained-release formulation; topical application, e.g., as a cream, ointment, or controlled-release patch or spray applied to the skin, lungs, or oral cavity; intravaginal or rectal administration, e.g., as a pessary, cream, or foam; sublingual; ophthalmic; transdermal; or intranasal, pulmonary, and other mucosal surface administration.
[0115] Pharmaceutically acceptable: As used herein, the phrase "pharmaceutically acceptable" refers to compounds, materials, compositions and / or dosage forms that are, within the scope of sound medical judgment, suitable for use in contact with the tissues of human beings and animals without excessive toxicity, irritation, allergic response, or other problem or complication commensurate with a reasonable benefit / risk ratio.
[0116] Pharmaceutically acceptable carrier: As used herein, the term "pharmaceutically acceptable carrier" means a pharmaceutically acceptable material, composition, or vehicle, such as a liquid or solid filler, diluent, excipient, or solvent encapsulating material, that is involved in carrying or transporting a compound of interest from one organ or body part to another. Each carrier must be "acceptable" in the sense of being compatible with the other ingredients of the formulation and not injurious to the patient. Some examples of materials that can serve as pharmaceutically acceptable carriers include: sugars such as lactose, glucose, and sucrose; starches such as corn starch and potato starch; cellulose and its derivatives such as sodium carboxymethyl cellulose, ethyl cellulose, and cellulose acetate; powdered tragacanth; malt; gelatin; talc; excipients such as cocoa butter and suppository wax; oils such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil; glycols such as propylene glycol; polyols such as glycerin, sorbitol, mannitol, and polyethylene glycol; esters such as ethyl oleate and ethyl laurate; agar; buffers such as magnesium hydroxide and aluminum hydroxide; alginic acid; pyrogen-free water; isotonic saline; Ringer's solution; ethyl alcohol; pH buffer solutions; polyesters, polycarbonates, and / or polyanhydrides; and other non-toxic, compatible substances used in pharmaceutical formulations.
[0117] Pharmaceutically acceptable salt: The term "pharmaceutically acceptable salt" as used herein refers to a salt of a compound that is suitable for use in a pharmaceutical context, i.e., a salt that, within the scope of sound medical judgment, is suitable for use in contact with the tissues of humans and lower animals without excessive toxicity, irritation, allergic response, etc., commensurate with a reasonable benefit / risk ratio. Pharmaceutically acceptable salts are well known. For example, SM Berge et al., on pharmaceutically acceptable salts, in J. Pharmaceutical Sciences, 66: 1-19 (1977). In some embodiments, pharmaceutically acceptable salts include, but are not limited to, non-toxic acid addition salts of amino groups formed with inorganic acids such as hydrochloric, hydrobromic, phosphoric, sulfuric, and perchloric acids, or organic acids such as acetic, maleic, tartaric, citric, succinic, or malonic acid, or by using other known methods such as ion exchange. In some embodiments, pharmaceutically acceptable salts include adipate, alginate, ascorbate, aspartate, benzenesulfonate, benzoate, bisulfate, borate, butyrate, camphorate, camphorsulfonate, citrate, cyclopentanepropionate, digluconate, dodecyl sulfate, ethanesulfonate, formate, fumarate, glucoheptonate, glycerophosphate, gluconate, hemisulfate, heptanoate, hexanoate, hydroiodide, 2-hydroxy-ethoxybenzoate, benzoic acid salt ... Examples of pharmaceutically acceptable salts include, but are not limited to, toluenesulfonate, lactobionate, lactate, laurate, lauryl sulfate, malate, maleate, malonate, methanesulfonate, 2-naphthalenesulfonate, nicotinate, nitrate, oleate, oxalate, palmitate, pamoate, pectinate, persulfate, 3-phenylpropionate, phosphate, picrate, pivalate, propionate, stearate, succinate, sulfate, tartrate, thiocyanate, p-toluenesulfonate, undecanoate, valerate, and the like. In some embodiments, pharmaceutically acceptable salts include, but are not limited to, non-toxic base addition salts, such as salts formed with acidic groups of provided compounds (e.g., phosphate linkages of oligonucleotides, phosphorothioate linkages of oligonucleotides, etc.) and bases. Representative alkali or alkaline earth metal salts include sodium, lithium, potassium, calcium, magnesium, and the like. In some embodiments, the pharmaceutically acceptable salt is an ammonium salt (e.g., —N(R)3 +) In some embodiments, the pharmaceutically acceptable salt is a sodium salt. In some embodiments, pharmaceutically acceptable salts include amine cations, carboxylates, sulfates, phosphates, nitrates, alkyls having 1 to 6 carbon atoms, sulfonates, and arylsulfonates, formed with non-toxic ammonium, quaternary ammonium, and counterions such as halides, hydroxides, and the like, where appropriate.
[0118] Protecting group: As used herein, the phrase "protecting group" refers to a temporary substituent that protects a potentially reactive functional group from undesired chemical transformations. Examples of such protecting groups include esters of carboxylic acids, silyl ethers of alcohols, and acetals and ketals of aldehydes and ketones, respectively. A "Si-protecting group" is a protecting group containing a Si atom, such as Si-trialkyl (e.g., trimethylsilyl, tributylsilyl, t-butyldimethylsilyl), Si-triaryl, Si-alkyl-diphenyl (e.g., t-butyldiphenylsilyl), or Si-aryl-dialkyl (e.g., Si-phenyldialkyl). Typically, Si-protecting groups are attached to oxygen atoms. The field of protecting group chemistry has been reviewed (Greene, TW; Wuts, PGM Protective Groups in Organic Synthesis, 5th ed.; John Wiley and Sons: Hoboken, NJ, 2014). Exemplary protecting groups (and related protected moieties) are described in detail below.
[0119] Protected hydroxyl groups are well known in the art and are described in Protecting Groups in Organic Synthesis, TW Greene and PGM Wuts, 3, incorporated herein by reference in its entirety. rdedition, John Wiley & Sons, 1999. Examples of suitably protected hydroxyl groups further include, but are not limited to, esters, carbonates, sulfonates, allyl ethers, ethers, silyl ethers, alkyl ethers, arylalkyl ethers, and alkoxyalkyl ethers. Examples of suitable esters include formates, acetates, propionates, pentanoates, crotonates, and benzoates. Specific examples of suitable esters include formates, benzoylformates, chloroacetates, trifluoroacetates, methoxyacetates, triphenylmethoxyacetates, p-chlorophenoxyacetates, 3-phenylpropionates, 4-oxopentanoates, 4,4-(ethylenedithio)pentanoates, pivaloates (trimethylacetates), crotonates, 4-methoxycrotonates, benzoates, p-benzylbenzoates, and 2,4,6-trimethylbenzoates. Examples of suitable carbonates include 9-fluorenylmethyl, ethyl, 2,2,2-trichloroethyl, 2-(trimethylsilyl)ethyl, 2-(phenylsulfonyl)ethyl, vinyl, allyl, and p-nitrobenzyl carbonates. Examples of suitable silyl ethers include trimethylsilyl, triethylsilyl, t-butyldimethylsilyl, t-butyldiphenylsilyl, triisopropylsilyl ether, and other trialkylsilyl ethers. Examples of suitable alkyl ethers include methyl, benzyl, p-methoxybenzyl, 3,4-dimethoxybenzyl, trityl, t-butyl, and allyl ethers, or derivatives thereof. Alkoxyalkyl ethers include acetals such as methoxymethyl, methylthiomethyl, (2-methoxyethoxy)methyl, benzyloxymethyl, beta-(trimethylsilyl)ethoxymethyl, and tetrahydropyran-2-yl ether.Examples of suitable arylalkyl ethers include benzyl, p-methoxybenzyl (MPM), 3,4-dimethoxybenzyl, O-nitrobenzyl, p-nitrobenzyl, p-halobenzyl, 2,6-dichlorobenzyl, p-cyanobenzyl, 2- and 4-picolyl ethers.
[0120] Protected amines are well known in the art and include those described in detail in Greene (1999). Suitable mono-protected amines further include, but are not limited to, aralkylamines, carbamates, allylamines, amides, and the like. Examples of suitable mono-protected amino moieties include t-butyloxycarbonylamino (-NHBOC), ethyloxycarbonylamino, methyloxycarbonylamino, trichloroethyloxycarbonylamino, allyloxycarbonylamino (-NHAlloc), benzyloxocarbonylamino (-NHCBZ), allylamino, benzylamino (-NHBn), fluorenylmethylcarbonyl (-NHFmoc), formamide, acetamide, chloroacetamide, dichloroacetamide, trichloroacetamide, phenylacetamide, trifluoroacetamide, benzamide, t-butyldiphenylsilyl, and the like. Suitable di-protected amines include amines substituted with two substituents independently selected from those listed above for the mono-protected amines, and further include cyclic imides such as phthalimides, maleimides, succinimides, etc. Suitable di-protected amines also include pyrroles, etc., 2,2,5,5-tetramethyl-[1,2,5]azadisilolidine, etc., and azides.
[0121] Protected aldehydes are well known in the art and include those described in detail in Greene (1999). Suitable protected aldehydes further include, but are not limited to, acyclic acetals, cyclic acetals, hydrazones, imines, and the like. Examples of such groups include dimethyl acetal, diethyl acetal, diisopropyl acetal, dibenzyl acetal, bis(2-nitrobenzyl) acetal, 1,3-dioxane, 1,3-dioxolane, semicarbazones, and derivatives thereof.
[0122] Protected carboxylic acids are well known in the art and include those described in detail in Greene (1999). Suitable protected carboxylic acids include optionally substituted C 1~6 Further examples include, but are not limited to, aliphatic esters, optionally substituted aryl esters, silyl esters, activated esters, amides, hydrazides, etc. Examples of such ester groups include methyl, ethyl, propyl, isopropyl, butyl, isobutyl, benzyl, and phenyl esters, each group optionally substituted. Additional suitable protected carboxylic acids include oxazolines and orthoesters.
[0123] Protected thiols are well known in the art and include those described in detail in Greene (1999). Suitable protected thiols further include, but are not limited to, disulfides, thioethers, silyl thioethers, thioesters, thiocarbonates, and thiocarbamates, among others. Examples of such groups include, but are not limited to, alkyl thioethers, benzyl and substituted benzyl thioethers, triphenylmethyl thioethers, and trichloroethoxycarbonyl thioesters, to name just a few.
[0124] Substituted: As described herein, compounds of the present disclosure can be optionally substituted and / or contain substituted moieties. In general, the term "substituted," whether preceded by the term "optionally" or not, means that one or more hydrogens of the specified moiety have been replaced with a suitable substituent. Unless otherwise indicated, an "optionally substituted" group can have a suitable substituent at each substitutable position of the group, and when more than one position in any given structure may be substituted with more than one substituent selected from a specified group, the substituents may be the same or different at all positions. Combinations of substituents contemplated by the present disclosure are preferably those that result in the formation of stable or chemically feasible compounds. The term "stable," as used herein, refers to a compound that is substantially unchanged when subjected to conditions that permit its production, detection, and, in certain embodiments, its recovery, purification, and use for one or more of the purposes disclosed herein. In some embodiments, exemplary substituents are described below.
[0125] Suitable monovalent substituents are halogen; -(CH2) 0~4 R ○ ;-(CH2) 0~4 OR ○ ;-O(CH2) 0~4 R ○ , -O-(CH2) 0~4 C(O)OR ○ ;-(CH2) 0~4 CH(OR ○ )2;R ○ may be substituted with -(CH2) 0~4 Ph;R ○ may be substituted with -(CH2) 0~4 O(CH2) 0~1 Ph;R ○ may be substituted with -CH=CHPh; R ○ may be substituted with -(CH2) 0~4 O(CH2) 0~1 -pyridyl; -NO2; -CN; -N3; -(CH2) 0~4 N(R ○ )2;-(CH2)0~4 N(R ○ )C(O)R ○ ;-N(R ○ )C(S)R ○ ;-(CH2) 0~4 N(R ○ )C(O)N(R ○ )2;-N(R ○ )C(S)N(R ○ )2;-(CH2) 0~4 N(R ○ )C(O)OR ○ ;-N(R ○ )N(R ○ )C(O)R ○ ;-N(R ○ )N(R ○ )C(O)N(R ○ )2;-N(R ○ )N(R ○ )C(O)OR ○ ;-(CH2) 0~4 C(O)R ○ ;-C(S)R ○ ;-(CH2) 0~4 C(O)OR ○ ;-(CH2) 0~4 C(O)SR ○ ;-(CH2) 0~4 C(O)OSi(R ○ )3;-(CH2) 0~4 OC(O)R ○ ;-OC(O)(CH2) 0~4 SR ○ 、-SC(S)SR ○ ;-(CH2) 0~4 SC(O)R ○ ;-(CH2) 0~4 C(O)N(R ○ )2;-C(S)N(R ○ )2;-C(S)SR ○ ;-SC(S)SR ○ 、-(CH2) 0~4 OC(O)N(R ○ )2;-C(O)N(OR ○ )R ○ ;-C(O)C(O)R ○ ;-C(O)CH2C(O)R ○ ;-C(NOR ○ )R ○;-(CH2) 0~4 SSR ○ ;-(CH2) 0~4 S(O)2R ○ ;-(CH2) 0~4 S(O)2OR ○ ;-(CH2) 0~4 OS(O)2R ○ ;-S(O)2N(R ○ )2;-(CH2) 0~4 S(O)R ○ ;-N(R ○ )S(O)2N(R ○ )2;-N(R ○ )S(O)2R ○ ;-N(OR ○ )R ○ ;-C(NH)N(R ○ )2;-Si(R ○ )3;-OSi(R ○ )3;-P(R ○ )2;-P(OR ○ )2;-OP(R ○ )2;-OP(OR ○ )2;-N(R ○ )P(R ○ )2;-B(R ○ )2;-OB(R ○ )2;-P(O)(R ○ )2;-OP(O)(R ○ )2;-N(R ○ )P(O)(R ○ )2;-(C 1~4 Linear or branched alkylene)ON(R ○ )2; or -(C 1~4 Linear or branched alkylene)C(O)ON(R ○ )2, and each R ○ are optionally substituted as defined below and independently represent hydrogen, C 1~20 C having 1 to 5 heteroatoms independently selected from aliphatic, nitrogen, oxygen, sulfur, silicon, and phosphorus 1~20 Heteroaliphatic, -CH2-(C 6~14 aryl), -O(CH2) 0~1 (C 6~14aryl), -CH2- (5-14 membered heteroaryl ring), a 5-20 membered monocyclic, bicyclic or polycyclic, saturated, partially unsaturated, or aryl ring having 0-5 heteroatoms independently selected from nitrogen, oxygen, sulfur, silicon, and phosphorus, or, notwithstanding the above definitions, R ○ two independent occurrences of together with their intervening atom(s) form a 5-20 membered, monocyclic, bicyclic or polycyclic, saturated, partially unsaturated, or aryl ring having 0-5 heteroatoms independently selected from nitrogen, oxygen, sulfur, silicon and phosphorus, which may be substituted as defined below.
[0126] R ○ (or R ○ Suitable monovalent substituents on a ring formed by taking two independent occurrences of -(CH) together with their intervening atoms are independently halogen, -(CH) 0~2 R ● ,-(Halo R ● ), -(CH2) 0~2 OH, -(CH2) 0~2 OR ● , -(CH2) 0~2 CH(OR ● )2;-O(HaloR ● ), -CN, -N3, -(CH2) 0~2 C(O)R ● , -(CH2) 0~2 C(O)OH, -(CH2) 0~2 C(O)OR ● , -(CH2) 0~2 SR ● , -(CH2) 0~2 SH, -(CH2) 0~2 NH2, -(CH2) 0~2 NHR ● , -(CH2) 0~2 NR ● 2, -NO2, -SiR ● 3. -OSiR ● 3. -C(O)SR ● , -(C 1~4 Linear or branched alkylene)C(O)OR ● , or -SSR ●and each R ● is unsubstituted or, if preceded by "halo", substituted with one or more halogens only, and C 1~4 Aliphatic, -CH2Ph, -O(CH2) 0~1 R is independently selected from Ph, or a 5-6 membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, and sulfur. ○ Suitable divalent substituents at a saturated carbon atom of include ═O and ═S.
[0127] Suitable divalent substituents include: =O, =S, =NNR * 2, =NNHC(O)R * , =NNHC(O)OR * , =NNHS(O)2R * , =NR * , =NOR * , -O(C(R * 2)) 2~3 O- or -S(C(R * 2)) 2~3 S- and R * Each independent occurrence of represents hydrogen, C which may be substituted as defined below 1~6 A suitable divalent substituent attached to adjacent substitutable carbon atoms of an "optionally substituted" group is selected from aliphatic or unsubstituted 5-6 membered saturated, partially unsaturated, or aryl rings having 0-4 heteroatoms independently selected from nitrogen, oxygen, and sulfur. * 2) 2~3 Contains O- and R * Each independent occurrence of represents hydrogen, C which may be substituted as defined below 1~6 It is selected from aliphatic or unsubstituted 5-6 membered saturated, partially unsaturated, or aryl rings having 0-4 heteroatoms independently selected from nitrogen, oxygen, and sulfur.
[0128] R * Suitable substituents on the aliphatic groups are halogen, -R ● ,-(Halo R ● ), -OH, -OR● , -O(HaloR ● ), -CN, -C(O)OH, -C(O)OR ● , -NH2, -NHR ● , -NR ● 2 or -NO2, and each R ● is unsubstituted or, if preceded by "halo", substituted with only one or more halogens, and independently represents C 1~4 Aliphatic, -CH2Ph, -O(CH2) 0~1 Ph, or a 5-6 membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, and sulfur.
[0129] In some embodiments, a suitable substituent at a substitutable nitrogen is —R † , -NR † 2. -C(O)R † , -C(O)OR † , -C(O)C(O)R † , -C(O)CHC(O)R † , -S(O)2R † , -S(O)NR † 2. -C(S)NR † 2. -C(NH)NR † 2 or -N(R † )S(O)2R † and each R † are independently hydrogen, optionally substituted C as defined below 1~6 an aliphatic, unsubstituted -OPh, or an unsubstituted 5-6 membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, and sulfur; or, notwithstanding the above definition, R † two independent occurrences of together with their intervening atom(s) form an unsubstituted 3-12 membered saturated, partially unsaturated, or aryl monocyclic or bicyclic ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, and sulfur.
[0130] R †Suitable substituents on the aliphatic group are independently halogen, -R ● ,-(Halo R ● ), -OH, -OR ● , -O(HaloR ● ), -CN, -C(O)OH, -C(O)OR ● , -NH2, -NHR ● , -NR ● 2 or -NO2, and each R ● is unsubstituted or, if preceded by "halo", substituted with only one or more halogens, and independently represents C 1~4 Aliphatic, -CH2Ph, -O(CH2) 0~1 Ph, or a 5-6 membered saturated, partially unsaturated, or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen, and sulfur.
[0131] Unsaturated: The term "unsaturated," as used herein, means that a moiety has one or more units of unsaturation.
[0132] Unless otherwise specified, salts, such as pharmaceutically acceptable acid or base addition salts, stereoisomeric forms, and tautomeric forms of the compounds provided are included.
[0133] 2. Detailed Description of Certain Embodiments Among other things, the present disclosure encompasses the recognition that many products produced by enzymes encoded by fungal biosynthetic gene clusters can be used to develop therapeutics against human targets to treat a variety of diseases. The present disclosure recognizes that one of the challenges of using fungal products is identifying their human targets. In some embodiments, the present disclosure provides techniques for the efficient identification of human targets of biosynthetic products produced by enzymes encoded by fungal biosynthetic gene clusters. In some embodiments, the provided techniques identify embedded target genes (ETaGs) in the proximal zone of the biosynthetic genes of the biosynthetic gene cluster, and, if necessary, further identify the human targets of the biosynthetic products produced by the enzymes encoded by the biosynthetic gene clusters by comparing the ETaG sequences to human nucleic acid sequences, particularly expressed human nucleic acid sequences including human protein-encoding genes. As will be readily recognized by those skilled in the art, once established, the relationship between the biosynthetic products from the biosynthetic gene cluster, the ETaGs, and the human targets can be utilized in a variety of ways. For example, one can start with a biosynthetic product produced by an enzyme encoded by a biosynthetic gene cluster, proceed to ETaG in the proximal zone of the biosynthetic genes of the biosynthetic gene cluster, and then proceed to a human target homologous to ETaG. Once a human target is identified, it can be prioritized (even if it was previously deemed non-druggable), and the biosynthetic product can be used to develop a modulator of the human target, including further optimization of the biosynthetic product for medical use, for example, by preparing and assaying analogs of the product using many methods available to those skilled in the art. Alternatively, one can start with a human target of therapeutic interest, proceed to ETaG homologous to the human target, and then proceed to a biosynthetic gene cluster whose proximal zone contains a biosynthetic gene containing ETaG. Once a biosynthetic gene cluster is identified, the biosynthetic product produced by the enzyme encoded by the biosynthetic gene cluster can be characterized and assayed for modulation of the human target or its product.The biosynthetic products can be used as leads for optimization, using numerous methods in the art in accordance with this disclosure, to provide agents useful for a number of medical, e.g., therapeutic, purposes.
[0134] While not intending to be limited by any theory, in some embodiments, the present disclosure encompasses the recognition that eukaryotic-derived ETaGs and / or products encoded thereby may have more similarity to mammalian genes and / or products encoded thereby than, for example, in any, their counterparts in prokaryotes such as bacteria; in some embodiments, eukaryotic ETaGs may be more therapeutically relevant. In some embodiments, ETaGs in fungi may be particularly useful for the development of human therapeutics, given the relative proximity of mammals to fungi in the phylogenetic tree.
[0135] In some embodiments, the disclosure provides techniques for identifying and / or characterizing ETaG, where ETaG is a non-biosynthetic gene in that it is not necessarily involved in the synthesis of a product produced by an enzyme encoded by a biosynthetic gene cluster containing ETaG, or the gene, in some embodiments, a proximal zone relative to the biosynthetic gene, contains ETaG (the enzymes encoded by the biosynthetic gene cluster can produce a biosynthetic product without ETaG). In some embodiments, ETaG is not required for the synthesis of a product produced by an enzyme encoded by a biosynthetic gene cluster containing ETaG, or the gene, in some embodiments, a proximal zone relative to the biosynthetic gene, contains ETaG (the enzymes encoded by the biosynthetic gene cluster can produce a biosynthetic product without ETaG). In some embodiments, ETaG is not involved in the synthesis of a product produced by an enzyme encoded by a biosynthetic gene cluster containing ETaG, or the gene, in some embodiments, a proximal zone relative to the biosynthetic gene, contains ETaG (the enzymes encoded by the biosynthetic gene cluster can produce a biosynthetic product without ETaG). In some embodiments, the ETaG is homologous to a human gene or comprises a sequence homologous to a human gene, e.g., shares at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90% or 95% homology with a human protein or sequence (e.g., a functional and / or structural unit such as a domain, functional structural feature (helix, sheet, etc.)).
[0136] In some embodiments, ETaG is co-regulated with at least one biosynthetic gene in the biosynthetic gene cluster. In some embodiments, ETaG is co-regulated with the biosynthetic gene cluster in that expression of ETaG correlates with production of a product encoded by an enzyme in the biosynthetic gene cluster. In some embodiments, ETaG provides a self-protective function. In some embodiments, ETaG encodes a transporter for a product produced by an enzyme in the biosynthetic gene cluster. In some embodiments, ETaG encodes a product, e.g., a protein, that can detoxify a product produced by an enzyme in the biosynthetic gene cluster. In some embodiments, ETaG encodes a resistant variant of a protein whose activity is targeted by a product produced by an enzyme in the biosynthetic gene cluster.
[0137] In some embodiments, the present disclosure provides: interrogating a set of nucleic acid sequences, each of which is found in a fungal strain and which comprises a biosynthetic gene cluster; Among at least one of the fungal nucleic acid sequences, not involved in the synthesis of products produced by enzymes encoded by biosynthetic gene clusters, within a proximity zone to at least one biosynthetic gene in a biosynthetic gene cluster; optionally co-regulated with at least one biosynthetic gene in a biosynthetic gene cluster identifying an embedded target gene (ETaG) sequence, The present invention provides a method comprising:
[0138] In some embodiments, the ETaG is homologous to a mammalian nucleic acid sequence. interrogating a set of nucleic acid sequences, each of which is found in a fungal strain and which comprises a biosynthetic gene cluster; Among at least one of the fungal nucleic acid sequences, not involved in the synthesis of products produced by enzymes encoded by biosynthetic gene clusters, within a proximity zone to at least one biosynthetic gene in a biosynthetic gene cluster; is homologous to the mammalian nucleic acid sequence to be expressed, optionally co-regulated with at least one biosynthetic gene in a biosynthetic gene cluster identifying an embedded target gene (ETaG) sequence, The present invention provides a method comprising: Proximity Zone
[0139] In some embodiments, ETaG is typically within the proximal zone of at least one gene in the biosynthetic gene cluster. In some embodiments, ETaG is within the proximal zone of at least one biosynthetic gene in the biosynthetic gene cluster. In some embodiments, the proximal zone is no more than 1-100 kb upstream or downstream of the gene. In some embodiments, the proximal zone is no more than 1-50 kb upstream or downstream of the gene. In some embodiments, the proximal zone is no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, or 90 kb upstream or downstream of the gene. In some embodiments, the proximal zone is no more than 1 kb upstream or downstream of the gene. In some embodiments, the proximal zone is no more than 5 kb upstream or downstream of the gene. In some embodiments, the proximal zone is no more than 10 kb upstream or downstream of the gene. In some embodiments, the proximal zone is no more than 15 kb upstream or downstream of the gene. In some embodiments, the proximal zone is no more than 20 kb upstream or downstream of the gene. In some embodiments, the proximal zone is no more than 25 kb upstream or downstream of the gene. In some embodiments, the proximal zone is no more than 30 kb upstream or downstream of the gene. In some embodiments, the proximal zone is no more than 35 kb upstream or downstream of the gene. In some embodiments, the proximal zone is no more than 40 kb upstream or downstream of the gene. In some embodiments, the proximal zone is no more than 45 kb upstream or downstream of the gene. In some embodiments, the proximal zone is no more than 50 kb upstream or downstream of the gene.
[0140] In some embodiments, the ETaG is within a biosynthetic gene cluster. In some embodiments, the ETaG is not within the region defined by the first and last genes of the biosynthetic gene cluster, but is within a proximal zone relative to the first or last gene of the biosynthetic gene cluster.
[0141] homology In some embodiments, ETaG is homologous to an expressed mammalian nucleic acid sequence. In some embodiments, the mammalian nucleic acid sequence is an expressed mammalian nucleic acid sequence. In some embodiments, the mammalian nucleic acid sequence is a mammalian gene. In some embodiments, the mammalian nucleic acid sequence is an expressed mammalian gene. In some embodiments, the mammalian nucleic acid is a human nucleic acid sequence. In some embodiments, the human nucleic acid sequence is an expressed human nucleic acid sequence. In some embodiments, the human nucleic acid sequence is a human gene. In some embodiments, the human nucleic acid sequence is an expressed human gene. In some embodiments, the human nucleic acid sequence is an existing target of therapeutic interest or encodes a product that is an existing target of therapeutic interest. In some embodiments, the human nucleic acid sequence is a novel target of therapeutic interest or encodes a product that is a novel target of therapeutic interest. In some embodiments, the human nucleic acid sequence is a target that was not considered to be druggable prior to the present disclosure, or encodes a product that is such a target. In some embodiments, the human nucleic acid sequence is a target that was not considered to be druggable by small molecules prior to the present disclosure, or encodes a product that is such a target. In some embodiments, the present disclosure provides the unexpected finding that targets traditionally considered to be non-druggable can be effectively modulated or targeted by small molecules that are biosynthetic products or analogs of biosynthetic products produced by enzymes encoded by a biosynthetic gene cluster, the biosynthetic gene cluster containing biosynthetic genes and a proximal zone to the biosynthetic genes containing an ETaG (or portion thereof, or product and / or portion thereof encoded thereby) that is homologous to the target.
[0142] In some embodiments, the present disclosure provides: interrogating a set of nucleic acid sequences, each of which is found in a fungal strain and which comprises a biosynthetic gene cluster; Among at least one of the fungal nucleic acid sequences, not involved in the synthesis of products produced by enzymes encoded by biosynthetic gene clusters, within a proximity zone to at least one biosynthetic gene in a biosynthetic gene cluster; is homologous to an expressed human nucleic acid sequence, optionally co-regulated with at least one biosynthetic gene in a biosynthetic gene cluster identifying an embedded target gene (ETaG), The present invention provides a method comprising:
[0143] In some embodiments, ETaG and the nucleic acid sequence share nucleic acid sequence homology. In some embodiments, the ETaG sequence is homologous to another nucleic acid sequence (e.g., an expressed human nucleic acid sequence) in that the ETaG nucleic acid sequence, or a portion thereof, shares similarity with the other nucleic acid sequence, or a portion thereof, at the nucleobase sequence level. In some embodiments, the sequence of ETaG shares nucleobase sequence similarity with another nucleic acid sequence. In some embodiments, a portion of the sequence of ETaG shares nucleobase sequence similarity with a portion of another nucleic acid sequence.
[0144] In some embodiments, the homologous portion is at least 50, 100, 150, 200, 300, 400, 500, 600, 70, 800, 900, or 1000 base pairs in length. In some embodiments, the length is at least 50 base pairs. In some embodiments, the length is at least 100 base pairs. In some embodiments, the length is at least 150 base pairs. In some embodiments, the length is at least 200 base pairs. In some embodiments, the length is at least 300 base pairs. In some embodiments, the length is at least 400 base pairs. In some embodiments, the length is at least 500 base pairs.
[0145] In some embodiments, the homologous portion encodes amino acid residues that are part of a particular structural and / or functional unit of the encoded protein. For example, in some embodiments, the homologous portion can encode a protein domain that is enzymatically active, responsible for interaction with an effector, etc., characteristic of the family of the encoded protein, as described in this disclosure.
[0146] Methods for assessing nucleic acid sequence similarity / homology are widely known in the art and can be used in accordance with the present disclosure.
[0147] In some embodiments, ETaG and the nucleic acid sequence share homology in their encoded products, e.g., proteins. In some embodiments, ETaG and the nucleic acid sequence are homologous in that the product encoded by ETaG, or a portion thereof, shares similarity with the product encoded by the nucleic acid sequence, or a portion thereof. In some embodiments, the encoded product is a protein. In some embodiments, the product encoded by ETaG and the nucleic acid sequence shares similarity over their entire length. In some embodiments, the product encoded by ETaG and the nucleic acid sequence shares similarity in certain portions.
[0148] In some embodiments, ETaG and nucleic acids are homologous in that the protein encoded by ETaG or a portion thereof shares similarity with the protein encoded by the nucleic acid or a portion thereof. ETaG and proteins encoded by nucleic acid sequences can share similarity at either their full length or partial level. In some embodiments, all amino acid residues in the homologous portion are contiguous. In some embodiments, not all amino acid residues in the homologous portion are contiguous.
[0149] In some embodiments, the portion of the protein is a protein domain. In some embodiments, the protein domain forms a structure characteristic of a protein family. In some embodiments, the protein domain performs a characteristic function. For example, in some embodiments, the protein domain has an enzymatic function. In some embodiments, such a function is shared by the protein encoded by ETaG and proteins encoded by homologous nucleic acid sequences, e.g., human genes. In some embodiments, the characteristic function is non-enzymatic. In some embodiments, the characteristic function is an interaction with another entity, e.g., a small molecule, nucleic acid, protein, etc.
[0150] In some embodiments, a portion of a protein is a set of contiguous or non-contiguous amino acid residues that are important for the function of the protein. In some embodiments, the function is enzymatic activity. In some embodiments, a portion of a protein is a set of residues required for activity. In some embodiments, a portion is a set of residues that interact with a substrate, intermediate, product, or cofactor. In some embodiments, a portion is a set of residues that interact with a substrate. In some embodiments, a portion is a set of residues that interact with an intermediate. In some embodiments, a portion is a set of residues that interact with a product. In some embodiments, a portion is a set of residues that interact with a cofactor.
[0151] In some embodiments, the function is an interaction with another entity. In some embodiments, the entity is a small molecule. In some embodiments, the entity is a lipid. In some embodiments, the entity is a carbohydrate. In some embodiments, the entity is a nucleic acid. In some embodiments, the entity is a protein. In some embodiments, the moiety is a set of amino acid residues that contact the interacting agent. For example, Figure 13 illustrates the nucleotide-interacting portions (sets of amino acids) of the Ras protein and its homolog ETaG, and Figures 14-16 illustrate the portions involved in protein-protein interactions.
[0152] In some embodiments, the interaction between an interacting entity and an amino acid residue can be assessed by hydrogen bonding, electrostatic forces, van der Waals forces, aromatic stacking, etc. In some embodiments, the interaction can be assessed by the distance of the amino acid residue to the interacting entity (e.g., 4 Å, as used in some cases).
[0153] In some embodiments, similarity means that the two structures have a C alpha backbone rmsd (root mean square deviation) within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, or 50 squared Angstroms and have the same overall fold or core domain. In some embodiments, the C alpha backbone rmsd is within.
[0154] In some embodiments, the portion of the protein is or comprises a structural element essential for protein effector recruitment, hi some embodiments, such a portion can be selected based on structure and / or activity data of proteins encoded by nucleic acid sequences homologous to ETaG, e.g., human genes encoding proteins homologous to ETaG.
[0155] In some embodiments, the portion of the protein comprises at least 2-200, 2-100, 2-50, 2-40, 2-30, 2-20, 2-15, 2-10, 3-200, 3-100, 3-50, 3-40, 3-30, 3-20, 3-15, 3-10, 4-200, 4-100, 4-50, 4-40, 4-30, 4-20, 4-15, 4-10, 5-200, 5-100, 5-50, 5-40, 5-30, 5-20, 5-15, or 5-10 amino acid residues. In some embodiments, a portion of a protein comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, or 150 amino acid residues. In some embodiments, a portion comprises at least 2 amino acid residues. In some embodiments, a portion comprises at least 3 amino acid residues. In some embodiments, a portion comprises at least 4 amino acid residues. In some embodiments, a portion comprises at least 5 amino acid residues. In some embodiments, a portion comprises at least 6 amino acid residues. In some embodiments, a portion comprises at least 7 amino acid residues. In some embodiments, a portion comprises at least 8 amino acid residues. In some embodiments, a portion comprises at least 9 amino acid residues. In some embodiments, a portion comprises at least 10 amino acid residues. In some embodiments, a portion comprises at least 15 amino acid residues. In some embodiments, a portion comprises at least 20 amino acid residues. In some embodiments, a portion comprises at least 25 amino acid residues. In some embodiments, a portion comprises at least 30 amino acid residues.
[0156] In accordance with the present disclosure, the similarity of nucleic acid sequences and protein sequences can be assessed by a number of methods, including those known in the art. For example, MUSCLE for protein sequences. In some embodiments, similarity is measured based on exact identity, e.g., the same amino acid residue at a given position. In some embodiments, similarity is measured based on one or more common properties, e.g., amino acid residues with one or more identical or similar properties (e.g., acidic, basic, aromatic, etc.).
[0157] In some embodiments, ETaG is homologous to a nucleic acid sequence (e.g., an expressed human nucleic acid sequence) in that the similarity between ETaG and the nucleic acid sequence or a portion thereof, or ETaG and the protein encoded by the nucleic acid sequence or a portion thereof, as described herein, is at or above the level based on the nucleic acid sequence. In some embodiments, ETaG is homologous to a nucleic acid sequence in that the similarity between ETaG and the nucleic acid sequence is at or above the level based on the nucleic acid sequence or a portion thereof. In some embodiments, ETaG is homologous to a nucleic acid sequence in that the similarity between ETaG and the nucleic acid sequence is at or above the level based on the protein encoded by the nucleic acid sequence or a portion thereof. In some embodiments, the level is at least 10% to 99%. In some embodiments, the level is at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%. In some embodiments, the level is at least 10%. In some embodiments, the level is at least 20%. In some embodiments, the level is at least 30%. In some embodiments, the level is at least 40%. In some embodiments, the level is at least 50%. In some embodiments, the level is at least 60%. In some embodiments, the level is at least 70%. In some embodiments, the level is at least 80%. In some embodiments, the level is at least 90%. In some embodiments, the level is 100%. In some embodiments, the level is less than 100%. In some embodiments, the level is less than or equal to 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99%.
[0158] In some embodiments, ETaG is homologous to a nucleic acid sequence in that the protein encoded by ETaG, or a portion thereof, has a three-dimensional structure similar to that of the protein encoded by the nucleic acid sequence. In some embodiments, similarity is assessed by, for example, a C alpha backbone rmsd (root mean square deviation) within 1-100, e.g., 5, 10, 20, 30, 40, or 50 squared angstroms. In some embodiments, sequences share similarity, have a C alpha backbone rmsd of 10 squared angstroms or less, and have the same overall fold or core domain. In some embodiments, structural similarity is assessed by interaction with another entity, e.g., a small molecule, nucleic acid, protein, etc. In some embodiments, structural similarity is assessed by small molecule binding. In some embodiments, the protein encoded by the embedded target gene, or a portion thereof, has a three-dimensional structure similar to that of the protein encoded by the nucleic acid sequence, in that a small molecule that binds to the protein encoded by the embedded target gene, or a portion thereof, also binds to the protein encoded by the nucleic acid sequence, or a portion thereof. In some embodiments, the binding has a Kd of 1 to 100 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100) μM or less.
[0159] simultaneous adjustment In some embodiments, ETaG is co-regulated with at least one biosynthetic gene in a biosynthetic gene cluster whose proximal zone contains a biosynthetic gene containing ETaG. In some embodiments, ETaG is co-regulated with a biosynthetic gene cluster whose proximal zone contains a biosynthetic gene containing ETaG. In some embodiments, ETaG is co-regulated with a biosynthetic gene cluster in that expression of ETaG and / or production of a product, e.g., a protein, encoded by ETaG correlates with production of a biosynthetic product produced by an enzyme encoded by the biosynthetic gene cluster. In some embodiments, production of a product, e.g., a protein, encoded by ETaG overlaps in time with production of a biosynthetic product by an enzyme encoded by the biosynthetic gene cluster. In some embodiments, ETaG is co-regulated with a biosynthetic gene cluster in that expression of ETaG is increased or turned on when a biosynthetic product produced by an enzyme encoded by the biosynthetic gene cluster is produced. In some embodiments, ETaG is co-regulated with the biosynthetic gene cluster in that expression of ETaG is increased or turned on when the level of a biosynthetic product produced by an enzyme encoded by the biosynthetic gene cluster is increased.
[0160] In some embodiments, ETaG provides an advantage to its host organism, e.g., a fungus, upon production of a biosynthetic product produced by an enzyme encoded by a co-regulated biosynthetic gene cluster. For example, in some embodiments, the protein encoded by ETaG contributes to the transport of the biosynthetic product out of the cell that produces the product. In some embodiments, the protein encoded by ETaG detoxifies the biosynthetic product so that the biosynthetic product does not harm the organism that produces the biosynthetic product but affects the growth or survival of other organisms.
[0161] In some embodiments, the present disclosure provides various methods for identifying ETaGs. For example, in some embodiments, sets of homologous biosynthetic gene clusters, e.g., biosynthetic gene clusters whose encoded enzymes produce the same biosynthetic product (based on product prediction (e.g., sequence-based prediction) and / or identification), typically from different fungal strains, are compared. Non-biosynthetic genes present in only or a few biosynthetic gene clusters in the set (within the biosynthetic gene cluster or within a proximal zone relative to the biosynthetic genes of the biosynthetic gene cluster), but absent from the majority of the biosynthetic gene clusters, are identified as ETaG candidates and, optionally, further compared with mammalian, e.g., human, nucleic acid sequences to identify homologous mammalian nucleic acid sequences. In some embodiments, as described in the Examples, such methods can be used to identify ETaGs on a genome-scale, e.g., from sequences of many (e.g., hundreds, thousands, or even more) genomes. The identified ETaGs can be prioritized based on the therapeutic importance of their mammalian homologs, particularly human homologs. In some embodiments, the organism containing ETaG contains one or more homologous genes of ETaG, as illustrated in the figures.
[0162] In some embodiments, ETaG is present in no more than 1%, 5%, or 10% of the biosynthetic gene clusters of the set. In some embodiments, ETaG is present in no more than 1%, 5%, or 10% of the homologous biosynthetic gene clusters of the set. In some embodiments, ETaG is present in no more than 1%, 5%, or 10% of the biosynthetic gene clusters of the set, which biosynthetic gene clusters encode enzymes that produce the same biosynthetic product. In some embodiments, the percentage is less than 1%. In some embodiments, the percentage is less than 5%. In some embodiments, the percentage is less than 10%.
[0163] In some embodiments, the present disclosure provides particularly effective and efficient methods for identifying homologous ETaGs for human nucleic acids encoding targets of therapeutic interest by querying a provided set of nucleic acid sequences containing ETaGs within a biosynthetic gene cluster and / or a proximity zone to a biosynthetic gene of the biosynthetic gene cluster.
[0164] In some embodiments, the present disclosure provides a set of nucleic acid sequences described herein. In some embodiments, the present disclosure provides a set of nucleic acid sequences, each found in a fungal strain (stain) and comprising a biosynthetic gene cluster. In some embodiments, the present disclosure provides a set of nucleic acid sequences, each found in a fungal strain (stain) and comprising ETaG. In some embodiments, the present disclosure provides a set of nucleic acid sequences, each found in a fungal strain (stain) and comprising a biosynthetic gene cluster and ETaG within a proximal zone to a biosynthetic gene of the biosynthetic gene cluster. In some embodiments, the nucleic acid sequence comprising the biosynthetic gene cluster does not include sequences outside the proximal zone to a biosynthetic gene of the biosynthetic gene cluster and sequences of the biosynthetic gene cluster. In some embodiments, the present disclosure provides a database comprising the provided set of nucleic acid sequences.
[0165] In some embodiments, the biosynthetic gene clusters of the provided technology include biosynthetic genes encoding enzymes that can be involved in the synthesis of compounds that share at least one common chemical trait. In some embodiments, the common chemical trait is a cyclic core structure. In some embodiments, the common chemical trait is a macrocyclic core structure. In some embodiments, the common chemical trait is a shared acyclic backbone. In some embodiments, the common chemical trait is that the compounds all belong to a particular category, e.g., nonribosomal peptides (NPRS), terpenes, isoprenes, alkaloids, etc. In some embodiments, by identifying individual ETaGs for a biosynthetic gene cluster, the present disclosure can distinguish between compounds that share a common chemical trait, even if they are structurally similar.
[0166] The provided sets can be of various sizes and / or diversity. In some embodiments, it is desirable to have more sequences from more species to increase the number of ETaGs and biosynthetic gene clusters. In some embodiments, the set includes at least 100, 200, 300, 400, 500, 1,000, 1,500, 2,000, 3,000, 5,000, 10,000, 20,000, 50,000, 100,000, 500,000, 1,000,000, 1,500,000, or 2,000,000 nucleic acid sequences comprising biosynthetic gene clusters. In some embodiments, the set includes at least 100, 200, 300, 400, 500, 1,000, 1,500, 2,000, 3,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 1,500,000, or 2,000,000 biosynthetic gene clusters. In some embodiments, the set includes at least 100, 200, 300, 400, 500, 1,000, 1,500, 2,000, 3,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 1,500,000, or 2,000,000 biosynthetic gene clusters related to ETaG (biosynthetic gene clusters whose proximal zones contain biosynthetic genes that contain ETaG). In some embodiments, the set includes at least 100, 200, 300, 400, 500, 1,000, 1,500, 2,000, 3,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 1,500,000, or 2,000,000 ETaGs. In some embodiments, the sequences of the provided set are derived from at least 100, 200, 300, 400, 500, 1,000, 1,500, 2,000, 3,000, 5,000, 10,000, 20,000, 50,000, 100,000 genomes from different species, e.g., different fungal species.
[0167] Among other things, the provided databases and / or provided sets can be used to, for example, identify ETaGs, identify ETaGs related to a given biosynthetic gene cluster, identify biosynthetic gene clusters related to a given ETaG, identify ETaGs homologous to a given mammalian nucleic acid sequence (e.g., a human gene), identify biosynthetic gene clusters related to a given mammalian nucleic acid sequence (e.g., a human gene; optionally via related ETaGs), identify mammalian nucleic acid sequences (e.g., a human gene) homologous to a given ETaG, identify biosynthetic gene clusters related to a given mammalian nucleic acid sequence (e.g., a human gene; optionally via related ETaGs), identify mammalian nucleic acid sequences (e.g., a human gene) homologous to a given ETaG, identify biosynthetic gene clusters (optionally via related ETaGs), identify mammalian nucleic acid sequences (e.g., a human gene) homologous to a given ETaG, identify biosynthetic gene clusters related to a given biosynthetic gene cluster (optionally via related ETaGs), identify ETaGs homologous to a given mammalian nucleic acid sequence (e.g., a human gene), identify biosynthetic gene clusters related ... It is structured to specifically improve the efficiency of identifying homologous human genes, identifying mammalian nucleic acid sequences (e.g., human genes) related to a given product (and / or analogs thereof) produced by enzymes encoded by biosynthetic gene clusters (optionally via related ETaGs and biosynthetic gene clusters), identifying products (and / or analogs thereof) produced by enzymes encoded by biosynthetic gene clusters related to a given mammalian nucleic acid sequence (e.g., human genes; optionally via related biosynthetic gene clusters and ETaGs), etc.
[0168] For example, in some embodiments, ETaGs in the provided sets and / or databases are indexed / marked for searching. For example, Figure 17 (applicant notes that the provided sets and databases may contain hundreds, thousands, or millions of sequences) depicts example sequences from the provided sets and / or databases, where the ETaGs are specifically indexed / marked (darker). Among other things, such structural features can greatly improve query efficiency: rather than searching tens, hundreds, or thousands of genomes for ETaGs homologous to a human gene of interest, provided technology can instead focus the search on the indexed / marked ETaGs (e.g., skipping non-biosynthetic gene cluster sequences and / or non-ETaG sequences (e.g., the blank arrows in Figure 17 and the sequences therebetween)) and quickly locate hits (e.g., the circled ETaGs in Figure 17), thereby saving time and resources from searching through a vast amount of unrelated genomic information.
[0169] Additionally or alternatively, the provided sets and databases of sequences are structured such that an ETaG can be independently annotated with information such as its related biosynthetic gene cluster (a related biosynthetic gene cluster of an ETaG is a biosynthetic gene cluster containing the biosynthetic genes in its proximity zone), the products and analogs produced by the enzymes encoded by the related biosynthetic gene cluster, its homologous mammalian nucleic acid sequence (e.g., human gene), etc. Similarly, a biosynthetic gene cluster can be independently annotated with information such as its related ETaG (a related ETaG of a biosynthetic gene cluster is an etaG that is in the proximity zone to the biosynthetic genes of the biosynthetic gene cluster), the biosynthetic products and analogs produced by the enzymes encoded by the biosynthetic gene cluster, the homologous mammalian nucleic acid sequence of the related ETaG and the products encoded thereby. By structuring sequence data with indexes and annotations, the provided sets and databases can provide numerous advantages. For example, in some embodiments, the provided systems provide rapid access to useful related information, such as ETaGs with their related biosynthetic gene clusters and human homologs (and vice versa), while keeping data sizes and query costs low.
[0170] In some embodiments, at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 2,500, 5,000, or 10,000 species or at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% or all of the ETaGs in the set are independently annotated. In some embodiments, at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 2,500, 5,000, or 10,000 species or at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% or all of the ETaGs in the set are independently annotated with their associated biosynthetic gene clusters and homologous mammalian nucleic acid sequences. In some embodiments, at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 2,500, 5,000, or 10,000, or at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%, or all of the biosynthetic gene clusters in the set are independently annotated. In some embodiments, at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 2,500, 5,000, or 10,000, or at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%, or all of the biosynthetic gene clusters in the set are independently annotated by their associated ETaGs.
[0171] In some embodiments, the provided set of sequences and / or database is embodied in a computer-readable medium.In some embodiments, the present disclosure provides a system comprising one or more non-transitory machine-readable storage media, storing data representing the provided set of sequences and / or database.Non-transitory machine-readable storage media suitable for implementing the provided data include, for example, semiconductor storage devices, such as EPROM, EEPROM and flash storage devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and all forms of non-volatile storage, including CD-ROM and DVD-ROM disks.In particular, the provided system can be particularly efficient due to the provided set and database having the specific structure described herein.
[0172] In some embodiments, the present disclosure provides a computer system capable of performing the provided techniques. In some embodiments, the present disclosure provides a computer system adapted to perform the provided methods. In some embodiments, the present disclosure provides a computer system adapted to query a provided set of sequences. In some embodiments, the present disclosure provides a computer system adapted to query a provided database. In some embodiments, the present disclosure provides a computer system adapted to access a provided database.
[0173] Computer systems that can be used to implement all or a portion of the provided technology can include various forms of digital computers. Examples of digital computers include, but are not limited to, laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, smart televisions, and other suitable computers. Mobile devices can be used to implement all or a portion of the provided technology. Mobile devices include, but are not limited to, tablet computing devices, personal digital assistants, mobile phones, smartphones, digital cameras, digital glasses, and other portable computing devices. The computing devices, their associations and relationships, and their functions described herein are intended to be merely examples and are not intended to limit the implementation of the technology.
[0174] All or part of the technology described herein, and various adaptations thereof, may be implemented, at least in part, via a computer program product, e.g., a computer program tangibly embodied in one or more information carriers, e.g., a data processing device, e.g., a programmable processor, one or more tangible, machine-readable storage media for execution by, or for controlling the operation of, a computer or multiple computers.
[0175] Computer programs for the provided techniques can be written in any form of programming language, including compiled or interpreted languages, and can be implemented in any form, including as a stand-alone program suitable for use in a computer processing environment, or as a module, portion, subroutine, or other unit. The computer program can be implemented to be executed on one computer or on multiple computers at one site or distributed across multiple sites and interconnected by a network.
[0176] For example, acts associated with implementing the programs and techniques may be performed by one or more programmable processors executing one or more computer programs for performing the provided techniques. All or a portion of the processes may be implemented as special purpose logic circuitry, such as an FPGA (field programmable gate array) and / or an ASIC (application-specific integrated circuit).
[0177] Processors suitable for executing a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, a processor will receive instructions and data from a read-only memory or a random-access memory, or both. Elements of a computer (including a server) include one or more processors for executing instructions and one or more storage devices for storing instructions and data. Typically, a computer will also include one or more machine-readable storage media, such as mass storage devices for storing data, e.g., magnetic, magneto-optical, or optical disks, or be operatively coupled to receive data from or transfer data to, or both. Non-transitory machine-readable storage media suitable for embodying computer program instructions and data include, by way of example, semiconductor storage devices, e.g., EPROM, EEPROM, and flash storage devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and all forms of non-volatile storage, including CD-ROM and DVD-ROM disks.
[0178] Each computing device, such as a tablet computer, may include a hard drive for storing data and computer programs, as well as a processing device (e.g., a microprocessor) and memory (e.g., RAM) for executing the computer programs. Each computing device may include an image capture device, such as a still camera or video camera. The image capture device may be built-in or simply accessible to the computing device.
[0179] Each computing device may include a graphics system including a display screen. The display screen, such as an LCD or CRT (cathode ray tube), displays images generated by the graphics system of the computing device to a user. As is well known, a display in a computer display (e.g., a monitor) physically transforms the computer display. For example, if the computer display is LCD-based, the orientation of liquid crystals may be changed by application of a bias voltage in a physical transformation that is visually apparent to a user. As another example, if the computer display is a CRT, the state of a phosphor screen may be changed by the impact of electrons in a similarly visually apparent physical transformation. Each display screen may be touch-sensitive, allowing a user to input information onto the display screen via a virtual keyboard. Some computing devices, such as desktops or smartphones, may provide a physical QWERTY keyboard and scroll wheel for inputting information onto the display screen. Each computing device, and the computer programs embodied therein, may also be configured to accept voice commands and perform functions in response to such commands.
[0180] In particular, the provided technology (methods, sets, databases, systems, etc.) establishes relationships between biosynthetic gene clusters, products produced by enzymes encoded by the biosynthetic gene clusters, ETaG, homologous mammalian nucleic acid sequences of ETaG, e.g., human genes, etc. Thus, the provided technology, in some embodiments, can be particularly powerful for identifying and / or characterizing human targets of products produced by enzymes encoded by biosynthetic gene clusters. The provided technology can also be particularly powerful for identifying and developing modulators of human targets. For example, in some embodiments, the provided technology can be used to rapidly identify ETaG of a human target (or a nucleic acid sequence encoding a human target) along with information about its associated biosynthetic gene cluster and / or the biosynthetic products produced by the enzymes of the biosynthetic gene cluster to develop therapeutics for the human target. The products of the associated biosynthetic gene clusters can be further characterized, and, if necessary, analogs thereof can be prepared, characterized, and assayed to develop therapeutics with improved properties. The provided technology can be particularly useful for human targets that are difficult to target and / or were deemed non-druggable prior to the present disclosure.
[0181] In some embodiments, the present disclosure provides methods for evaluating compounds using the identified ETaGs and products encoded thereby. contacting at least one test compound with a gene product encoded by an embedded target gene of a fungal nucleic acid sequence, wherein the embedded target gene is is not required for or involved in the biosynthesis of the product of the biosynthetic gene cluster, within a proximity zone to at least one biosynthetic gene in the cluster; is homologous to a mammalian nucleic acid sequence, Optionally co-regulated with at least one biosynthetic gene in the cluster the steps of: the level or activity of the gene product is altered in the presence of the test compound compared to its absence; or The level or activity of the gene product is comparable to that observed in the presence of a reference agent that has a known effect on the level or activity. determining that The present invention provides a method comprising:
[0182] In some embodiments, the disclosure provides a method for identifying and / or characterizing a mammalian, e.g., human, target of a product or product analog produced by an enzyme encoded by a biosynthetic gene cluster, comprising: identifying a human homolog of ETaG that is within a proximity zone to at least one biosynthetic gene of the biosynthetic gene cluster or within a proximity zone to at least one biosynthetic gene of a second biosynthetic gene cluster, the second biosynthetic gene cluster encoding an enzyme that produces the same biosynthetic product produced by the enzyme encoded by the biosynthetic gene cluster; Optionally, assaying the effect of the product or product analog produced by the enzyme encoded by the biosynthetic gene cluster on the target. The present invention provides a method comprising:
[0183] In some embodiments, the present disclosure provides methods for evaluating compounds using products encoded by mammalian, e.g., human, nucleic acid sequences that are homologous to ETaG. At least one test compound is is not required for or involved in the biosynthesis of the product of the biosynthetic gene cluster, within a proximity zone to at least one biosynthetic gene in the cluster; is homologous to a mammalian nucleic acid sequence, contacting at least one biosynthetic gene in the cluster with a gene product encoded by a mammalian nucleic acid sequence that is homologous to the embedded target gene, wherein the gene product is optionally co-regulated; the level or activity of the gene product is altered in the presence of the test compound compared to its absence; or The level or activity of the gene product is comparable to that observed in the presence of a reference agent that has a known effect on the level or activity. determining that The present invention provides a method comprising:
[0184] In some embodiments, the disclosure provides a method for identifying and / or characterizing a mammalian, e.g., human, target of a product or product analog produced by an enzyme encoded by a biosynthetic gene cluster, comprising: identifying a human homolog of ETaG within a proximity zone to at least one biosynthetic gene of the biosynthetic gene cluster; Optionally, assaying the effect of the product or product analog produced by the enzyme encoded by the biosynthetic gene cluster on the target. The present invention provides a method comprising:
[0185] In some embodiments, the provided methods and systems are useful for assessing the interaction of a compound with a human target. In some embodiments, the present disclosure provides a method for accessing the interaction of a compound with a human target, the method comprising: comparing the nucleic acid sequence of, or encoding, the human target to a set of nucleic acid sequences comprising one or more ETaGs; The present invention provides a method comprising:
[0186] In some embodiments, a compound produced by an enzyme of the biosynthetic gene cluster interacts with a target encoded by a mammalian, e.g., human, nucleic acid sequence that is homologous to an ETaG associated with the biosynthetic gene cluster.
[0187] In some embodiments, the provided technology is particularly useful for designing and / or providing modulators of human targets because, among other things, the provided technology provides the connection between the biosynthetic gene cluster, ETaG, and human target genes.
[0188] In some embodiments, the disclosure provides a method for identifying and / or characterizing a modulator of a human target, comprising: providing a product or analog thereof, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster, and within a proximity zone to at least one biosynthetic gene in the biosynthetic gene cluster: is homologous to a human target or a nucleic acid sequence encoding a human target, Step 1: There is at least one biosynthetic gene in the cluster and optionally co-regulated ETaG. The present invention provides a method comprising:
[0189] In some embodiments, the human target is a Ras protein. In some embodiments, the Ras protein is an HRas protein. In some embodiments, the Ras protein is a KRas protein. In some embodiments, the Ras protein is an NRas protein. In some embodiments, the human target is a protein comprising a RasGEF domain. In some embodiments, the protein is KNDC1, PLCE1, RALGDS, RALGPS1, RALGPS2, RAPGEF1, RAPGEF2, RAPGEF3, RAPGEF4, RAPGEF5, RAPGEF6, RAPGEFL1, RASGEF1A, RASGEF1B, RASGEF1C, RASGRF1, RASGRF2, RASGRP1, RASGRP2, RASGRP3, RASGRP4, RGL1, RGL2, RGL3, RGL4 / RGR, SOS1, SOS2, or a human guanine nucleotide exchange factor. In some embodiments, the protein is SOS1. In some embodiments, the protein is a human guanine nucleotide exchange factor. In some embodiments, the human target is a protein comprising a RasGAP domain. In some embodiments, the protein is DAB2IP, GAPVD1, IQGAP1, IQGAP2, IQGAP3, NF1, RASA1, RASA2, RASA3, RASA4, RASAL1, RASAL2, or SYNGAP1. In some embodiments, the protein is protein p120. In some embodiments, the protein is a human guanine nucleotide activator.
[0190] In some embodiments, the present disclosure provides a method for identifying and / or characterizing a modulator of a human Ras protein, comprising: Preparing an analog of a compound produced by an enzyme encoded by the biosynthetic gene cluster, Within the proximity zone to at least one biosynthetic gene in the biosynthetic gene cluster, is homologous to a human Ras protein, a RasGEF domain, or a RasGAP domain, or a nucleic acid sequence encoding a human Ras protein, a RasGEF domain, or a RasGAP domain; Step 1: There is at least one biosynthetic gene in the cluster and optionally co-regulated ETaG. The present invention provides a method comprising:
[0191] In some embodiments, proteins comprising a RasGEF domain modulate one or more functions of a human Ras protein. In some embodiments, proteins comprising a RasGAP domain modulate one or more functions of a human Ras protein.
[0192] In some embodiments, the present disclosure provides a method for identifying and / or characterizing a modulator of a human Ras protein, comprising: Preparing an analog of a compound produced by an enzyme encoded by the biosynthetic gene cluster, Within the proximity zone to at least one biosynthetic gene in the biosynthetic gene cluster, is homologous to a human Ras protein or a nucleic acid sequence encoding a human Ras protein, Step 1: There is at least one biosynthetic gene in the cluster and optionally co-regulated ETaG. The present invention provides a method comprising:
[0193] In some embodiments, the disclosure provides a method for identifying and / or characterizing a modulator of a protein comprising a RasGEF domain, the method comprising: Preparing an analog of a compound produced by an enzyme encoded by the biosynthetic gene cluster, Within the proximity zone to at least one biosynthetic gene in the biosynthetic gene cluster, a RasGEF domain or a nucleic acid sequence encoding a RasGEF domain; Step 1: There is at least one biosynthetic gene in the cluster and optionally co-regulated ETaG. The present invention provides a method comprising:
[0194] In some embodiments, the present disclosure provides a method for identifying and / or characterizing a modulator of a protein comprising a RasGAP domain, the method comprising: Preparing an analog of a compound produced by an enzyme encoded by the biosynthetic gene cluster, Within the proximity zone to at least one biosynthetic gene in the biosynthetic gene cluster, a RasGAP domain or a nucleic acid sequence encoding a RasGAP domain; Step 1: There is at least one biosynthetic gene in the cluster and optionally co-regulated ETaG. The present invention provides a method comprising:
[0195] In some embodiments, the biosynthetic gene cluster is an exemplary biosynthetic gene cluster, or a biosynthetic gene cluster comprising one or more biosynthetic genes, illustrated in one of the figures with an ETaG homologous to a Ras protein, e.g., Figures 5-12 and 20-27. In some embodiments, the biosynthetic gene cluster is an exemplary biosynthetic gene cluster, or a biosynthetic gene cluster comprising one or more biosynthetic genes, illustrated in one of the figures with an ETaG homologous to a RasGEF domain, e.g., Figures 28-33 and 35. In some embodiments, the biosynthetic gene cluster is an exemplary biosynthetic gene cluster illustrated in one of the figures with an ETaG homologous to a RasGEF domain, e.g., Figures 28-33 and 35. In some embodiments, the biosynthetic gene cluster is an exemplary biosynthetic gene cluster illustrated in one of the figures with an ETaG homologous to a RasGEF domain, or a biosynthetic gene cluster comprising one or more biosynthetic genes, e.g., Figure 34 and Figures 36-39. In some embodiments, the biosynthetic gene cluster is an exemplary biosynthetic gene cluster illustrated in one of the figures with an ETaG homologous to a RasGEF domain, e.g., Figure 34 and Figures 36-39. Exemplary ETaG sequences are provided in the present disclosure and can be utilized to locate and identify, among other things, biosynthetic gene clusters, biosynthetic genes, and the like.
[0196] In some embodiments, the present disclosure provides a method for modulating a human target, comprising: providing a product or analog thereof, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster, and within a proximity zone to at least one biosynthetic gene in the biosynthetic gene cluster: is homologous to a human target or a nucleic acid sequence encoding a human target, Step 1: There is at least one biosynthetic gene in the cluster and optionally co-regulated ETaG. The present invention provides a method comprising:
[0197] In some embodiments, the present disclosure provides a method for modulating a Ras protein, comprising: Providing a product or analog thereof, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster of one of Figures 5-12 and 20-27. The present invention provides a method comprising:
[0198] In some embodiments, the disclosure provides a method for modulating a RasGEF protein, comprising: Providing a product or analog thereof, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster of one of Figures 28-33 and 35. The present invention provides a method comprising:
[0199] In some embodiments, the present disclosure provides a method for modulating a RasGAP protein, comprising: Providing a product or analog thereof, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster of one of Figures 34 and 36-39. The present invention provides a method comprising:
[0200] In some embodiments, the ETaG is identified by the methods provided.
[0201] In some embodiments, the product is produced by an enzyme encoded by a biosynthetic gene cluster that is a secondary metabolite produced by the biosynthetic gene cluster.
[0202] In some embodiments, the analog of the product comprises the structural core of the product. In some embodiments, the product is cyclic, e.g., monocyclic, bicyclic, or polycyclic. In some embodiments, the structural core of the product is or comprises a monocyclic, bicyclic, or polycyclic ring system. In some embodiments, the structural core of the product comprises one ring of the bicyclic or polycyclic ring system of the product.
[0203] In some embodiments, the product is linear and the structural core is its backbone. In some embodiments, the product is or comprises a polypeptide and the structural core is the backbone of the polypeptide. In some embodiments, the product is or comprises a polyketide and the structural core is the backbone of the polyketide.
[0204] In some embodiments, the analog is a product substituted with one or more suitable substituents described herein. In some embodiments, the analog is a structural core substituted with one or more suitable substituents described herein.
[0205] Among other things, the present disclosure provides the following exemplary embodiments: 1. Querying a set of nucleic acid sequences, each found in a fungal strain and comprising a biosynthetic gene cluster; Among at least one of the fungal nucleic acid sequences, is not required for or involved in the biosynthesis of the product of the biosynthetic gene cluster, within a proximity zone to at least one gene in the cluster, is homologous to a mammalian nucleic acid sequence, identifying an embedded target gene (ETaG) sequence that is optionally co-regulated with at least one biosynthetic gene in the cluster; A method comprising: 2. The method of embodiment 1, wherein the ETaG sequence is within a proximity zone to at least one biosynthetic gene in the cluster. 3. The method of any one of the preceding embodiments, wherein the nucleic acid sequence comprising the biosynthetic gene cluster does not contain any sequence other than nucleic acid sequences of the proximal zone to the biosynthetic genes of the biosynthetic gene cluster and nucleic acid sequences of the biosynthetic gene cluster. 4. The method of any one of the preceding embodiments, wherein the proximal zone is no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 kb upstream or downstream of a biosynthetic gene in the cluster. 5. The method of any one of the preceding embodiments, wherein the proximal zone is no more than 50 kb upstream or downstream of the biosynthetic genes in the cluster. 6. The method of any one of the preceding embodiments, wherein the proximal zone is no more than 40 kb upstream or downstream of the biosynthetic genes in the cluster. 7. The method of any one of the preceding embodiments, wherein the proximal zone is no more than 30 kb upstream or downstream of the biosynthetic genes in the cluster. 8. The method of any one of the preceding embodiments, wherein the proximal zone is no more than 20 kb upstream or downstream of the biosynthetic genes in the cluster. 9. The method of any one of the preceding embodiments, wherein the proximal zone is no more than 10 kb upstream or downstream of the biosynthetic genes in the cluster. 10. The method of any one of the preceding embodiments, wherein the proximal zone is a region between two biosynthetic genes of a biosynthetic gene cluster. 11. The method of any one of the preceding embodiments, wherein the mammalian nucleic acid sequence is an expressed sequence. 12. The method of any one of the preceding embodiments, wherein the mammalian nucleic acid sequence is a gene. 13. The method of any one of the preceding embodiments, wherein the mammalian nucleic acid sequence is a human nucleic acid sequence. 14. The method of any one of the preceding embodiments, wherein the embedded target gene sequence is homologous to the expressed mammalian nucleic acid sequence in that its base sequence or portion thereof is at least 50%, 60%, 70%, 80%, or 90% identical to that of the mammalian nucleic acid sequence. 15. The method of embodiment 14, wherein the sequence or portion thereof is at least 50, 100, 150, or 200 base pairs in length. 16. The method of any one of embodiments 1-13, wherein the embedded target gene sequence is homologous to the mammalian nucleic acid sequence to be expressed in that the product encoded by the embedded target gene or portion thereof is homologous to that of the mammalian nucleic acid sequence or portion thereof. 17. The method of embodiment 16, wherein the product is a protein. 18. The method of embodiment 16, wherein the protein encoded by the embedded target gene or portion thereof is at least 50%, 60%, 70%, 80%, or 90% similar to that encoded by the mammalian nucleic acid sequence or portion thereof. 19. The method of embodiment 16, wherein the protein encoded by the embedded target gene or portion thereof has a three-dimensional structure similar to that of the protein encoded by the mammalian nucleic acid sequence or portion thereof. 20. The method of embodiment 19, wherein the portion of the protein encoded by the embedded target gene has a three-dimensional structure similar to that of the protein encoded by the mammalian nucleic acid sequence. 21. The method of any one of embodiments 19-20, wherein the similarity is such that the structures have a C alpha backbone rmsd (root mean square deviation) within 10 squared Angstroms and have the same overall fold or core domain. 22. The method of any one of embodiments 19-20, wherein the protein encoded by the embedded target gene or portion thereof has a three-dimensional structure similar to the protein encoded by the mammalian nucleic acid sequence, in that small molecules that bind to the protein encoded by the embedded target gene or portion thereof also bind to the protein encoded by the mammalian nucleic acid sequence or portion thereof. 23. The method of embodiment 22, wherein binding of the small molecule to the embedded target gene and the protein encoded by the mammalian nucleic acid sequence or portion thereof has a Kd of 100 μM, 50 μM, 10 μM, 5 μM, or 1 μM or less. 24. The method of any one of embodiments 22-23, wherein the small molecule is produced by a fungus. 25. The method of embodiment 24, wherein the small molecule is acyclic. 26. The method of embodiment 24, wherein the small molecule is cyclic. 27. The method of any one of embodiments 24 to 26, wherein the small molecule is a secondary metabolite molecule produced by a fungus. 28. The method of any one of embodiments 24 to 27, wherein the small molecule is synthesized non-ribosomally. 29. The method of any one of embodiments 24 to 28, wherein the small molecule is a biosynthetic product of a biosynthetic gene cluster. 30. The method of embodiment 16, wherein the portion of the protein encoded by the embedded target gene is at least 50%, 60%, 70%, 80%, or 90% similar to the portion of the protein encoded by the expressed mammalian nucleic acid sequence. 31. The method of embodiment 30, wherein the portion of the protein is a protein domain. 32. The method of any one of embodiments 30-31, wherein the portion of the protein is a set of amino acid residues required for function. 33. The method of embodiment 32, wherein the function is an enzymatic function. 34. The method of embodiment 33, wherein the set of amino acid residues contacts the substrate. 35. The method of embodiment 33, wherein the set of amino acid residues contacts an intermediate. 36. The method of embodiment 33, wherein the set of amino acid residues contacts the product. 37. The method of embodiment 32, wherein the function is an interaction with another entity. 38. The method of embodiment 37, wherein the entity is a small molecule. 39. The method of embodiment 37, wherein the entity is a lipid. 40. The method of embodiment 37, wherein the entity is a carbohydrate. 41. The method of embodiment 37, wherein the entity is a nucleic acid. 42. The method of embodiment 37, wherein the entity is a protein. 43. The method of any one of embodiments 32-42, wherein each of the residues in the set is within 4 Å of the entity. 44. The method of any one of the preceding embodiments, wherein the embedded target gene is co-regulated with at least one gene in the cluster. 45. The method of any one of the preceding embodiments, wherein the embedded target gene is absent from 80%, 90%, 95%, or 100% of all fungal nucleic acid sequences in a set that originates from different fungal strains and that comprises homologous or identical biosynthetic gene clusters. 46. The method of any one of the preceding embodiments, wherein the set comprises at least 100, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 1,500,000, 2,000,000, or 2,500,000 distinct fungal nucleic acid sequences. 47. The method of any one of the preceding embodiments, wherein the set comprises nucleic acid sequences from at least 100, 500, 1,000, 5,000, 10,000, 15,000, 20,000, 22,000, 25,000, or 30,000 distinct fungal strains. 48. The method of any one of the preceding embodiments, wherein the ETaG sequence is not a housekeeping gene. 49. The method of any one of the preceding embodiments, wherein the ETaG sequence is or comprises a sequence that is homologous to a second nucleic acid sequence, or part thereof, in the same genome. 50. The method of any one of the preceding embodiments, wherein the ETaG sequence is, or comprises, a sequence encoding a product that is homologous to a product or portion thereof encoded by a second nucleic acid sequence in the same genome. 51. The method of embodiment 49 or 50, wherein the homology is at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5%. 52. The method of embodiment 49, wherein the homology is at least 70%. 53. The method of embodiment 49, wherein the homology is at least 80%. 54. The method of embodiment 49, wherein the homology is at least 90%. 55. The method of any one of embodiments 48-54, wherein the second nucleic acid sequence is or comprises a housekeeping gene. 56. The method of any one of embodiments 48-55, wherein the ETaG sequence encodes a product that confers resistance to the product of the biosynthetic gene cluster, while the second nucleic acid sequence does not encode such a product. 57. The method of embodiment 56, wherein the ETaG sequence encodes a protein that confers resistance to the small molecule product of the biosynthetic gene cluster, while the protein encoded by the second nucleic acid sequence does not confer such resistance. 58. The method of any one of the preceding embodiments, wherein the nucleic acid sequences in the set comprise a biosynthetic gene cluster whose biosynthetic genes encode enzymes involved in the synthesis of compounds sharing at least one common chemical trait. 59. The method of any one of the preceding embodiments, wherein the nucleic acid sequences are derived from multiple fungal strains. 60. The method of any one of the preceding embodiments, wherein the common chemical feature is or includes a cyclic system. 61. The method of any one of the preceding embodiments, wherein the common chemical feature is or comprises a macrocycle. 62. The method of any one of embodiments 52-61, wherein the common chemical feature is or comprises an acyclic backbone. 63. The method of any one of embodiments 52-62, wherein the compounds sharing at least one common chemical property are polyketides. 64. The method of any one of embodiments 52-62, wherein the compounds sharing at least one common chemical trait are non-ribosomal peptides. 65. The method of any one of embodiments 52-62, wherein the compounds sharing at least one common chemical trait are alkaloids. 66. The method of any one of embodiments 52-62, wherein the compounds sharing at least one common chemical trait are terpenes / isoprenes. 67. Contacting at least one test compound with a gene product encoded by an embedded target gene of a fungal nucleic acid sequence, wherein the embedded target gene (ETaG) is: is not required for or involved in the biosynthesis of the product of the biosynthetic gene cluster, within a proximity zone to at least one biosynthetic gene in the cluster; is homologous to a mammalian nucleic acid sequence, and optionally co-regulating at least one biosynthetic gene in the cluster; the level or activity of the gene product is altered in the presence of the test compound compared to its absence; or The level or activity of the gene product is comparable to that observed in the presence of a reference agent that has a known effect on the level or activity. determining that A method comprising: 68. The method according to embodiment 67, wherein the ETaG is an ETaG according to any one of embodiments 1 to 66. 69. The method of embodiment 67 or 68, wherein the mammalian nucleic acid sequence is a human Ras sequence. 70. The method of embodiment 69, wherein the mammalian nucleic acid sequence is a KRas, HRas or NRas sequence. 71. The method of embodiment 67 or 68, wherein the mammalian nucleic acid sequence is a sequence encoding a RasGEF domain. 72. The method of embodiment 67 or 68, wherein the mammalian nucleic acid sequence is a sequence encoding a RasGAP domain. 73. The method of any one of embodiments 66-72, wherein the ETaG is the ETaG in one of Figures 1-39. 74. The method of any one of embodiments 66-73, wherein the biosynthetic gene cluster is a biosynthetic gene cluster in one of Figures 1-39. 75. The method of any one of embodiments 66-74, wherein the test compound is a biosynthetic product of a biosynthetic gene cluster or an analog thereof. 76. Contacting at least one test compound with a gene product encoded by an expressed mammalian nucleic acid sequence, wherein the sequence is an expressed mammalian nucleic acid sequence that is homologous to an embedded target gene sequence according to any one of embodiments 1 to 75. A method comprising: 77. The method of embodiment 76, wherein the mammalian nucleic acid sequence is a human Ras sequence. 78. The method of embodiment 77, wherein the mammalian nucleic acid sequence is a KRas, HRas or NRas sequence. 79. The method of embodiment 76 or 77, wherein the mammalian nucleic acid sequence is a sequence encoding a RasGEF domain. 80. The method of embodiment 76 or 77, wherein the mammalian nucleic acid sequence is a sequence encoding a RasGAP domain. 81. The method of any one of embodiments 76-80, wherein the ETaG is the ETaG in one of Figures 1-39. 82. The method of any one of embodiments 76-81, wherein the biosynthetic gene cluster is a biosynthetic gene cluster in one of Figures 1-39. 83. The method of any one of embodiments 76-82, wherein the test compound is a biosynthetic product of a biosynthetic gene cluster or an analog thereof. 84. Identifying a human homologue of ETaG within a proximity zone to at least one biosynthetic gene of a biosynthetic gene cluster; Optionally, assaying the effect of the product or analog of the product produced by the enzyme encoded by the biosynthetic gene cluster on the human homolog. A method comprising: 85. The method according to embodiment 77, wherein the ETaG is an ETaG according to any one of embodiments 1 to 66. 86. A method for identifying and / or characterizing a modulator of a human target, comprising: providing a product or analog thereof, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster, and within a proximity zone to at least one gene in the biosynthetic gene cluster: is homologous to a human target or a nucleic acid sequence encoding a human target, Step 1: There is at least one biosynthetic gene in the cluster and optionally co-regulated ETaG. A method comprising: 87. The method according to embodiment 86, wherein the ETaG is an ETaG according to any one of embodiments 1 to 83. 88. The method of embodiment 86, wherein the human target is a Ras protein. 89. The method of embodiment 88, wherein the human target is KRas, HRas or NRas. 90. The method of embodiment 86, wherein the human target comprises a RasGEF domain. 91. The method of embodiment 86, wherein the human target comprises a RasGAP domain. 92. The method of any one of embodiments 86-91, wherein the ETaG is the ETaG in one of Figures 1-39. 93. The method of any one of embodiments 86-92, wherein the biosynthetic gene cluster is a biosynthetic gene cluster in one of Figures 1-39. 94. A method for modulating a human target, comprising: providing a product or analog thereof, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster, and within a proximity zone to at least one biosynthetic gene in the biosynthetic gene cluster: is homologous to a human target or a nucleic acid sequence encoding a human target, Step 1: There is at least one biosynthetic gene in the cluster and optionally co-regulated ETaG. A method comprising: 95. The method of embodiment 94, wherein the human target is a Ras protein. 96. The method of embodiment 94, wherein the human target is KRas, HRas or NRas. 97. The method of embodiment 94, wherein the human target comprises a RasGEF domain. 98. The method of embodiment 94, wherein the human target comprises a RasGAP domain. 99. The method of any one of embodiments 94-98, wherein the ETaG is the ETaG in one of Figures 1-39. 100. The method of any one of embodiments 94-99, wherein the biosynthetic gene cluster is the biosynthetic gene cluster in one of Figures 1-39. 101. The method according to embodiment 94, wherein the ETaG is an ETaG according to any one of embodiments 1 to 93. 102. A set of nucleic acid sequences each found in a fungal strain and comprising a biosynthetic gene cluster. A database comprising: A database in which a set of nucleic acid sequences is embodied in a computer-readable medium. 103. The database according to embodiment 102, in which one or more embedded target genes according to any one of embodiments 1 to 101 are indexed. 104. One or more non-transitory machine-readable storage media storing data representing a set of nucleic acid sequences, each found in a fungal strain and comprising a biosynthetic gene cluster. A system including: 105. One or more non-transitory machine-readable storage media storing data representing a set of nucleic acid sequences, each of which is or includes an ETaG sequence. A system including: 106. The system of embodiment 105, wherein one or more embedded target genes of any one of embodiments 1 to 101 are indexed. 107. A computer system adapted to perform the method of any one of embodiments 1-101. 108. A computer system adapted to access a database described in any one of embodiments 95 to 103. [Example]
[0206] Non-limiting examples of the techniques provided are described below.
[0207] Example 1 Building an example database and its use.
[0208] For example, antiSMASH was used to process approximately 2,000 reported fungal genomes to identify potential biosynthetic gene clusters, and approximately 70,000 identified biosynthetic gene clusters were added to the database. The initial database was queried using human targets of interest. For example, the protein sequence of human Sec7 was used to perform a BLAST search against the initial library to identify ETaGs. Alternatively or additionally, biosynthetic gene clusters can be compared among themselves. For example, in one process, non-biosynthetic genes present in one or several biosynthetic gene clusters (within a proximity zone to at least one biosynthetic gene of the biosynthetic gene cluster) but absent from most other homologous biosynthetic gene clusters for the same biosynthetic product are identified as potential ETaGs and further confirmed by analyzing whether they have homologous mammalian nucleic acid sequences (e.g., human genes) at the nucleic acid level and / or, preferably, at the protein level. The identified ETaGs can be indexed / marked and annotated. The databases are searchable by either nucleotide sequences (eg, BLASTN; tBLASTx) or protein sequences (eg, tBLASTn).
[0209] The results of the BLAST queries of the human targets were, in some embodiments, listed in order of the strength of sequence homology, showing all putative hits in the database. The DNA sequences of all hit biosynthetic gene clusters were then inspected to verify that one or more open reading frame (gene) homologs of the target protein were within the predicted range of the biosynthetic gene cluster.
[0210] In some embodiments, GenBank-formatted sequence files (*.gbk) for each biosynthetic cluster were assembled and curated, from which ETaG protein sequences were obtained using prediction algorithms and / or methods, including, for example, antiSMASH. The protein family (pfam) function of an open reading frame can be predicted, for example, by antiSMASH, and the nucleotide distance between each identified ETaG predicted by antiSMASH and its nearest biosynthetic enzyme can be determined. In some embodiments, the closer a predicted ETaG is to a biosynthetic enzyme, the more likely it is that this open reading frame encodes a legitimate ETaG.
[0211] The applicant has successfully identified many biosynthetic gene clusters with related ETaGs, in addition to several bona fide ETaG-containing biosynthetic gene clusters (biosynthetic gene clusters for cyclosporine, ferutamide, lovastatin, mycophenolic acid and brefeldin).
[0212] In some embodiments, the present disclosure encompasses the recognition that ETaGs can function as functional homologs (orthologs) of putative human target proteins. In some embodiments, the protein sequences of putative ETaG hits were compared to the sequences of the human target orthologs. For example, in a project to find ETaGs for human protein A, n biosynthetic gene clusters containing putative protein A homologs were found, and all n predicted ETaG proteins were aligned with human protein A. In some embodiments, only amino acids within specific catalytic or structural domains (e.g., based on predicted subfamily domain architecture) that define the ETaG / target pfam boundary were used in the alignment analysis. By aligning all ETaGs and the human target protein(s), ETaG sequences were directly compared to their human counterparts, and their phylogenetic relationships were analyzed to generate quantitative correlation data (e.g., peptide sequence similarity and / or phylogenetic tree visualization). Additional analysis can include, for example, the conservation / similarity of essential structural elements for protein effector recruitment / binding based on examination of the tertiary protein structure of the human target. For example, in some embodiments, the aligned sequences were compared to PDB crystal structures corresponding to target protein residues within 4 angstroms of the corresponding engaged protein. Without intending to be limited by any theory, if these structural motifs are conserved within fungal ETaG, this may indicate an increased probability that metabolites produced by ETaG-related biosynthetic gene clusters are effectors of both fungal and human target proteins, and the produced metabolites may be drug candidates or leads for drug development against the human target. In some embodiments, the above-described analysis was used to prioritize ETaG and its related biosynthetic gene clusters, as well as metabolites produced from the biosynthetic gene clusters, for targeting of human targets.
[0213] Example 2 Human target - Modulators of Sec7.
[0214] Among other things, the present disclosure provides techniques for identifying modulators of human targets. In some embodiments, human sequences are used to query provided databases to identify biosynthetic gene clusters within whose proximity zones homologs of the human sequence reside.
[0215] For example, among other things, the present disclosure provides a biosynthetic gene cluster whose biosynthetic product can modulate Sec7 function. To identify modulators of the human Sec7 domain, the Sec7 protein sequence was used to query databases, such as the database provided in Example 1. A Sec7-homologous ETaG instance was identified in Penicillium vulpinum IBT 29486, which has a related biosynthetic gene cluster (ETaG is in the proximal zone to one of the biosynthetic genes of the biosynthetic gene cluster). See Figures 1, 18, and 19. Notably, the identified biosynthetic gene cluster shares homology with the biosynthetic gene cluster of brefeldin A in Eupenicillium brefeldianum and was predicted to produce brefeldin A. Therefore, brefeldin A was identified as a candidate modulator of Sec7 and / or a lead compound for the modulator. If desired, one can use numerous methods available in the art in accordance with this disclosure to express the biosynthetic gene cluster of Penicillium vulpinum IBT 29486, isolate and characterize the product, and then optionally validate the results by assaying the function of the product on Sec7. Because brefeldin A has been reported to target the Sec7 domain of human GBF1, this example illustrates that the provided technology can be successfully utilized to identify modulators of the human target.
[0216] Example 3 ETaG of lovastatin, ferutamide and cyclosporine.
[0217] The provided technology can be used to identify various ETaG entities. For example, as demonstrated herein, the provided technology can be effectively used to identify EtaGs related to lovastatin, ferutamide, and cyclosporine. Example results are presented in Figures 2 to 4.
[0218] Example 4 Human target - Ras modulators
[0219] In particular, the present disclosure provides methods for the production of Ras proteins, the biosynthetic products of which are Ras proteins and / or proteins containing a RasGEF domain (e.g., KNDC1, PLCE1, RALGDS, RALGPS1, RALGPS2, RAPGEF1, RAPGEF2, RAPGEF3, RAPGEF4, RAPGEF5, RAPGEF6, RAPGEFL1, RASGEF1A, RASGEF1B, RASGEF1C, RASGRF1, RASGRF2, RASGRP1, RASGR The present disclosure provides biosynthetic gene clusters that can modulate one or more functions of proteins containing RasGAP domains (e.g., RasP2, RASGRP3, RASGRP4, RGL1, RGL2, RGL3, RGL4 / RGR, SOS1, SOS2, etc.) and / or proteins containing RasGAP domains (e.g., DAB2IP, GAPVD1, IQGAP1, IQGAP2, IQGAP3, NF1, RASA1, RASA2, RASA3, RASA4, RASAL1, RASAL2, SYNGAP1, etc.). Ras proteins, such as HRas, KRas, and NRas, have been linked to many human cancers but are notoriously difficult targets for drug discovery. Among other things, the present disclosure provides techniques for developing Ras modulators, including Ras inhibitors.
[0220] Human Ras sequences were used to query the provided databases, e.g., the database in Example 1. Eight ETaG examples were identified from different strains with varying levels of sequence similarity to the human Ras protein. Related biosynthetic gene clusters encode enzymes for producing different types of compounds. See Figures 5-12 and 20-27. The identified ETaG encoding proteins can be highly homologous to the human Ras protein. See, e.g., Figure 13 for similarity of nucleotide-binding residues, Figure 14 for BRAF-interacting residues, Figure 15 for rasGAP-interacting residues, and Figure 16 for SOS-interacting residues.
[0221] Similarly, biosynthetic gene clusters are identified whose biosynthetic products can modulate RasGEF and RasGAP domains. As demonstrated herein, the exemplary identified biosynthetic gene clusters can contain genes and / or modules involved in the synthesis of various types of moieties / products, e.g., terpenes, PKSs, NRPSs, etc. See, e.g., identified biosynthetic gene clusters and RasGEF and RasGAP homologs, Figures 28-39.
[0222] Examples of identified ETaG sequences are listed below:
[0223] Figure 5: Thermomyces lanuginosus. Ras ETaG sequence: [ka]
[0224] Figure 6: Talaromyces leycettanus CBS 398.68. Ras ETaG sequence: [ka] [ka]
[0225] Figure 7: Sistotremastrum niveocremeum. Ras ETaG sequence: [ka]
[0226] Sistotremastrum suecicum. Ras ETaG sequence: [ka]
[0227] Figure 8: Agaricus bisporus var. burnettii JB137-S8. Ras ETaG sequence: [ka]
[0228] Figure 9: Coprinopsis cinerea okayama. Ras ETaG sequence: [ka]
[0229] Figure 10: Colletotrichum higginsianum. Ras ETaG sequence: [ka]
[0230] Figure 11: Gyalolechia flavorubescens KoLRI002931. Ras ETaG sequence: [ka]
[0231] Figure 12: Bipolaris maydis ATCC 48331. Ras ETaG sequence: [ka]
[0232] Figure 18: Penicillium vulpinum IBT 29486. Sec7 ETaG sequence:
[0233] [ka] [ka] [ka]
[0234] Figure 20: Thermomyces lanuginosus ATCC 200065. Ras ETaG sequence:
[0235] [ka]
[0236] Aspergillus rambelli. Ras ETaG array:
[0237] [ka]
[0238] Aspergillus ochraceoroseus. Ras ETaG sequence:
[0239] [ka]
[0240] Figure 21: Agaricus bisporus var. burnettii JB137-S8. Ras ETaG sequence: [ka]
[0241] Agaricus bisporus H97. Ras ETaG sequence:
[0242] [ka]
[0243] Coprinopsis cinerea okayama. Ras ETaG array: [ka]
[0244] Hypholoma sublateritum FD-334. Ras ETaG array: [ka]
[0245] Figure 22: Sistotremastrum niveocremeum. Ras ETaG sequence: [ka]
[0246] Sistotremastrum suecicum. Ras ETaG sequence: [ka]
[0247] Figure 23: Talaromyces leycettanus CBS 398.68. Ras ETaG sequence: [ka] [ka]
[0248] Figure 24: Thermoascus crustaceus. Ras ETaG sequence: [ka]
[0249] Figure 25: Bipolaris maydis ATCC 48331. Ras ETaG sequence: [ka]
[0250] Figure 26: Colletotrichum higginsianum IMI 349063. Ras ETaG sequence: [ka]
[0251] Figure 27: Gyalolechia flavorubescens. Ras ETaG sequence: [ka]
[0252] Figure 28: Lecanosticta acicola CBS 871.95. RasGEF ETaG sequence:
[0253] [ka] [ka] [ka]
[0254] Penicillium chrysogenum Wisconsin 54-1255. RasGEF ETaG sequence:
[0255] [ka] [ka] [ka]
[0256] Figure 29: Magnaporthe oryzae 70-15. RasGEF ETaG sequence:
[0257] [ka] [ka] [ka]
[0258] Figure 30: Arthroderma gypseum CBS 118893. RasGEF ETaG sequence:
[0259] [ka] [ka] [ka]
[0260] Figure 31: Endocarpon pusillum strain KoLRI No. LF000583. RasGEF ETaG sequence:
[0261] [ka]
[0262] Figure 32: Fistulina hepatica ATCC 64428. RasGEF ETaG sequence:
[0263] [ka] [ka] [ka]
[0264] Figure 33: Aureobasidium pullulans var. pullulans EXF-150. RasGEF ETaG sequence:
[0265] [ka]
[0266] Figure 34: Acremonium furcatum. RasGAP ETaG sequence:
[0267] [ka] [ka]
[0268] Figure 35: Purpureocillium lilacinum strain TERIBC 1. RasGEF ETaG sequence:
[0269] [ka] [ka] [ka] [ka]
[0270] Fusarium sp. JS1030. RasGEF ETaG sequence:
[0271] [ka]
[0272] Figure 36: Corynespora cassiicola UM 591. RasGAP ETaG sequence:
[0273] [ka] [ka] [ka] [ka]
[0274] Magnaporthe oryzae strain SV9610. RasGAP ETaG sequence: [ka] [ka] [ka] [ka]
[0275] Figure 37: Colletotrichum acutatum strain 1 KC05_01. RasGAP ETaG sequence: [ka] [ka]
[0276] Figure 38: Hypoxylon sp. E7406B. RasGAP ETaG sequence: [ka]
[0277] Diaporthe ampelina isolate DA912. RasGAP ETaG sequence: [ka] [ka] [ka]
[0278] Figure 39: Talaromyces piceae strain 9-3. RasGAP ETaG sequence:
[0279] [ka] [ka]
[0280] Sporothrix insectorum RCEF 264. RasGAP ETaG sequence: [ka] [ka]
[0281] The identified biosynthetic gene clusters allow for numerous methods to be utilized to identify and characterize compounds produced by the enzymes of these biosynthetic gene clusters (e.g., as described in Clevenger, et al., Nat. Chem. Bio., 13, 895-901 (2017) and references cited therein) in accordance with the present disclosure. Once identified, compounds can be assayed to assess their ability to modulate human Ras proteins. Additionally and alternatively, compounds can be used, for example, as lead compounds to prepare additional analogs for SAR studies to further improve affinity, efficacy, selectivity, etc. for modulating Ras activity. It is anticipated that useful compounds will be developed from the identified ETaG-related biosynthetic gene clusters.
[0282] While various embodiments have been described and illustrated herein, those skilled in the art will readily envision various other means and / or structures for performing the functions and / or obtaining the results and / or one or more of the advantages described herein, and each such variation and / or modification is intended to be included. More generally, those skilled in the art will readily recognize that any parameters, dimensions, materials, and configurations described herein are meant to be exemplary, and that the actual parameters, dimensions, materials, and / or configurations will depend on the specific application(s) for which the teachings of the present disclosure are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the present disclosure described herein. Accordingly, it should be understood that the foregoing embodiments are presented by way of example only, and that the provided technology, including the claimed technology, may be practiced other than as specifically described and claimed. Additionally, any combination of two or more features, systems, articles, materials, kits, and / or methods is within the scope of the present disclosure, provided that such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent. The present invention provides, for example, the following items. (Item 1) interrogating a set of nucleic acid sequences, each of which is found in a fungal strain and which comprises a biosynthetic gene cluster; Among at least one of the fungal nucleic acid sequences, is not required for or involved in the biosynthesis of the product of said biosynthetic gene cluster, within a proximity zone to at least one gene in the cluster, is homologous to a mammalian nucleic acid sequence, optionally co-regulated with at least one biosynthetic gene in said cluster identifying an embedded target gene (ETaG) sequence, A method comprising: (Item 2) 2. The method of claim 1, wherein the ETaG sequence is within a proximity zone to at least one biosynthetic gene in the cluster. (Item 3) 3. The method of claim 2, wherein the nucleic acid sequence comprising a biosynthetic gene cluster does not contain any sequences other than the nucleic acid sequences of the proximal zone to the biosynthetic genes of the biosynthetic gene cluster and the nucleic acid sequences of the biosynthetic gene cluster. (Item 4) 4. The method of claim 3, wherein the proximal zone is no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 kb upstream or downstream of a biosynthetic gene in the cluster. (Item 5) 5. The method of claim 4, wherein the mammalian nucleic acid sequence is a human nucleic acid sequence. (Item 6) 6. The method of claim 5, wherein the embedded target gene sequence is homologous to the expressed mammalian nucleic acid sequence in that its base sequence or a portion thereof is at least 50%, 60%, 70%, 80%, or 90% identical to that of the mammalian nucleic acid sequence. (Item 7) 7. The method of claim 6, wherein the sequence or portion thereof is at least 50, 100, 150, or 200 base pairs in length. (Item 8) 6. The method of claim 5, wherein the embedded target gene sequence is homologous to the mammalian nucleic acid sequence to be expressed in that the protein encoded by the embedded target gene or portion thereof is homologous to that of the mammalian nucleic acid sequence or portion thereof. (Item 9) 9. The method of claim 8, wherein the protein encoded by the embedded target gene or portion thereof is at least 50%, 60%, 70%, 80%, or 90% similar to a protein encoded by a mammalian nucleic acid sequence or portion thereof. (Item 10) 10. The method of claim 9, wherein the protein encoded by the embedded target gene or portion thereof has a three-dimensional structure similar to the protein encoded by the mammalian nucleic acid sequence, in that small molecules that bind to the protein encoded by the embedded target gene or portion thereof also bind to the protein encoded by the mammalian nucleic acid sequence or portion thereof. (Item 11) 11. The method of claim 10, wherein the binding of the small molecule to the protein encoded by the embedded target gene or portion thereof and the protein encoded by the mammalian nucleic acid sequence or portion thereof has a Kd of 100 μM, 50 μM, 10 μM, 5 μM, or 1 μM or less. (Item 12) Item 13. The method of item 10, wherein the small molecule is a biosynthetic product of a biosynthetic gene cluster. 6. The method of claim 5, wherein the portion of the protein encoded by the embedded target gene is at least 50%, 60%, 70%, 80%, or 90% similar to a portion of the protein encoded by the expressed mammalian nucleic acid sequence, and wherein the portion of the protein is a protein domain. (Item 14) 10. The method of any one of the preceding items, wherein the embedded target gene is absent from 80%, 90%, 95%, or 100% of all fungal nucleic acid sequences in the set that are derived from different fungal strains and that contain homologous or identical biosynthetic gene clusters. (Item 15) 15. The method of claim 14, wherein the set comprises nucleic acid sequences from at least 100, 500, 1,000, 5,000, 10,000, 15,000, 20,000, 22,000, 25,000, or 30,000 distinct fungal strains. (Item 16) contacting at least one test compound with a gene product encoded by an embedded target gene of a fungal nucleic acid sequence, wherein said embedded target gene (ETaG) is: is not required for or involved in the biosynthesis of the product of the biosynthetic gene cluster, within a proximity zone to at least one biosynthetic gene in the cluster; is homologous to a mammalian nucleic acid sequence, optionally co-regulated with at least one biosynthetic gene in said cluster the steps of: the level or activity of the gene product is altered in the presence of the test compound compared to its absence; or the level or activity of said gene product is comparable to that observed in the presence of a reference agent that has a known effect on said level or activity determining that A method comprising: (Item 17) Item 17. The method according to Item 16, wherein the ETaG is the ETaG according to any one of Items 1 to 15. (Item 18) 18. The method of claim 17, wherein the mammalian nucleic acid sequence is a human Ras sequence. (Item 19) Item 17. The method according to item 16, wherein the biosynthetic gene cluster is a biosynthetic gene cluster in one of Figures 1 to 39. (Item 20) 17. The method of claim 16, wherein the test compound is a biosynthetic product of the biosynthetic gene cluster or an analog thereof. (Item 21) contacting at least one test compound with a gene product encoded by an expressed mammalian nucleic acid sequence, wherein the sequence is an expressed mammalian nucleic acid sequence that is homologous to the embedded target gene sequence of any one of items 1 to 15. A method comprising: (Item 22) 22. The method of claim 21, wherein the mammalian nucleic acid sequence is a human Ras sequence. (Item 23) Item 22. The method according to item 21, wherein the ETaG is the ETaG in one of Figures 1 to 39. (Item 24) 22. The method according to item 21, wherein the biosynthetic gene cluster is a biosynthetic gene cluster in one of Figures 1 to 39. (Item 25) 22. The method of claim 21, wherein the test compound is a biosynthetic product of the biosynthetic gene cluster or an analog thereof. (Item 26) identifying a human homolog of ETaG within a proximity zone to at least one biosynthetic gene of the biosynthetic gene cluster; optionally assaying the effect of a product produced by an enzyme encoded by the biosynthetic gene cluster, or an analog of said product, on said human homolog; A method comprising: (Item 27) Item 27. The method according to Item 26, wherein the ETaG is the ETaG according to any one of Items 1 to 15. (Item 28) 1. A method for identifying and / or characterizing a modulator of a human target, comprising: providing a product or analog thereof, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster, and within a proximity zone to at least one gene in the biosynthetic gene cluster: is homologous to said human target or a nucleic acid sequence encoding said human target, optionally co-regulated with at least one biosynthetic gene in said cluster ETaG exists, step A method comprising: (Item 29) Item 29. The method according to Item 28, wherein the ETaG is the ETaG according to any one of Items 1 to 15. (Item 30) 29. The method of claim 28, wherein the human target is a Ras protein. (Item 31) Item 29. The method according to item 28, wherein the ETaG is the ETaG in one of Figures 1 to 39. (Item 32) Item 29. The method according to item 28, wherein the biosynthetic gene cluster is a biosynthetic gene cluster in one of Figures 1 to 39. (Item 33) 1. A method for modulating a human target, comprising: providing a product or analog thereof, wherein the product is produced by an enzyme encoded by a biosynthetic gene cluster, and within a proximity zone to at least one biosynthetic gene in the biosynthetic gene cluster: is homologous to said human target or a nucleic acid sequence encoding said human target, optionally co-regulated with at least one biosynthetic gene in said cluster ETaG exists, step A method comprising: (Item 34) 34. The method of claim 33, wherein the human target is a Ras protein. (Item 35) Item 34. The method according to item 33, wherein the ETaG is the ETaG in one of Figures 1 to 39. (Item 36) Item 34. The method according to item 33, wherein the biosynthetic gene cluster is a biosynthetic gene cluster in one of Figures 1 to 39. (Item 37) Item 34. The method according to Item 33, wherein the ETaG is the ETaG according to any one of Items 1 to 15. (Item 38) A set of nucleic acid sequences each found in a fungal strain and containing a biosynthetic gene cluster. A database comprising: A database, wherein said set of nucleic acid sequences is embodied in a computer-readable medium. (Item 39) 39. The database according to Item 38, wherein one or more embedded target genes according to any one of Items 1 to 37 are indexed. (Item 40) one or more non-transitory machine-readable storage media storing data representing a set of nucleic acid sequences, each found in a fungal strain and comprising a biosynthetic gene cluster; A system including: (Item 41) One or more non-transitory machine-readable storage media storing data representing a set of nucleic acid sequences, each of which is or includes an ETaG sequence. A system including: (Item 42) Item 42. The system according to Item 41, wherein the one or more embedded target genes according to any one of Items 1 to 37 are indexed. (Item 43) A computer system adapted to perform the method according to any one of items 1 to 37 or to access the database according to any one of items 34 to 39. (Item 44) A method, database or system according to any one of exemplary embodiments 1 to 108.
Claims
1. 1. A method for identifying a modulator of a human target, comprising: a) using the human targets to interrogate a set of nucleic acid sequences, each of the nucleic acid sequences being found in a fungal strain and comprising a biosynthetic gene cluster; b) among at least one of said fungal nucleic acid sequences: is not required for or involved in the biosynthesis of the product of said biosynthetic gene cluster; within a proximal zone to at least one biosynthetic gene in the cluster, the proximal zone being no more than 100 kb upstream or downstream of the at least one biosynthetic gene in the biosynthetic gene cluster; and is a homologue of the human target identifying an embedded target gene (ETaG) sequence, thereby identifying a product of said biosynthetic gene cluster as a modulator of said human target.
2. The method of claim 1 , wherein the ETaG sequence is co-regulated with at least one biosynthetic gene in the cluster.
3. 2. The method of claim 1, further comprising the step of c) assaying the effect of the product of the biosynthetic gene cluster or an analog of the product on the human target.
4. 2. The method of claim 1, wherein the ETaG sequence is or comprises a sequence that encodes a product that is homologous to a product or portion thereof encoded by a second nucleic acid sequence in the same genome.
5. 5. The method of claim 4, wherein the homology is at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5%.
6. 5. The method of claim 4, wherein the ETaG sequence encodes a product that provides resistance to a product of the biosynthetic gene cluster, while the second nucleic acid sequence does not encode such a product.
7. 7. The method of claim 6, wherein the ETaG sequence encodes a protein that provides resistance to small molecule products of the biosynthetic gene cluster, while the protein encoded by the second nucleic acid sequence does not provide such resistance.
8. 5. The method of claim 4, wherein the second nucleic acid sequence is or comprises a housekeeping gene.
9. 2. The method of claim 1, wherein the nucleic acid sequences in the set comprise a biosynthetic gene cluster, the biosynthetic genes encoding enzymes involved in the synthesis of compounds that share at least one common chemical trait.
10. 10. The method of claim 1, wherein the at least one common chemical feature is or includes a cyclic system, a macrocyclic system, an acyclic skeleton, or any combination thereof.
11. 11. The method of claim 10, wherein the at least one common chemical signature is a polyketide, a non-ribosomal peptide, an alkaloid, a terpene, or isoprene.
12. The method of claim 1 , wherein steps a) and b) are performed via a computer system.
13. The method of claim 1 , wherein each biosynthetic gene cluster is identified and / or annotated.
14. 14. The method of claim 13, wherein at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95%, or all, of the ETaGs in the set of nucleic acid sequences are each independently identified, indexed, and / or annotated.
15. 15. The method of claim 14, wherein each annotated ETaG is independently annotated with respect to a related biosynthetic gene cluster comprising at least one gene in a proximity zone to said ETaG.
16. 15. The method of claim 14, wherein each annotated ETaG is independently annotated relative to a human homolog of said ETaG.
17. A computer system for executing the method of claim 1, comprising a non-transitory machine-readable storage medium storing data representing the set of nucleic acid sequences and a program for executing the method of claim 1.