Methods and systems for discovery of non-integrated target genes
Patent Information
- Application Number
- JP2024556331
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-11-16
- Filing Date
- 2022-11-15
- Publication Date
- 2025-11-21
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 264,150, filed November 16, 2021, the contents of which are incorporated herein by reference in their entirety.
[0002] Field The present disclosure generally relates to methods and systems for identifying genes associated with (but not integrated into) biosynthetic gene clusters, including predicting the function of secondary metabolites based on the co-occurrence and / or co-evolution of genes encoding secondary metabolites with biosynthetic gene clusters or their core enzymes, and predicting biosynthetic gene clusters that produce secondary metabolites with desired activities, and uses thereof. [Background technology]
[0003] Microorganisms produce a wide variety of small molecule compounds known as secondary metabolites or natural products with diverse chemical structures and functions. Some secondary metabolites enable microorganisms to tolerate hostile environments, while others serve as weapons in inter- and intra-species competition. See, e.g., Piel, J. Nat. Prod. Rep., 26: 338-362, 2009. Many human pharmaceuticals (including, e.g., antibacterial agents, antitumor agents, and insecticides) are derived from secondary metabolites. See, e.g., Newman DJ and Cragg GM, J. Nat. Prod., 79: 629-661, 2016.
[0004] Microorganisms synthesize secondary metabolites using enzyme proteins encoded by clusters of co-located genes called biosynthetic gene clusters (BGCs). Evidence is emerging that some microbial biosynthetic gene clusters contain genes that do not appear to be involved in the synthesis of the associated biosynthetic product produced by the enzymes encoded by the cluster. In some cases, such non-biosynthetic genes have been described as "self-protecting" because they code for proteins that can apparently render the host organism resistant to the associated biosynthetic product. For example, in some cases, resistance mutants of genes encoding transporters of biosynthetic products, detoxifying enzymes that act on biosynthetic products, or proteins whose activity is targeted by biosynthetic products have been reported. See, e.g., Cimermancic, et al., Cell 158:412, 2014; Keller, Nat. Chem. Biol. 11:671, 2015. Researchers have proposed that the identification of such genes and the determination of their function may be useful in determining the role of the biosynthetic products synthesized by the enzymes of the cluster. For example, Yeh,et al.,ACS Chem.Biol.11:2275,2016;Tang,et al.,ACS Chem.Biol.10:2841,2015;Regueira,et al.,Appl,Environ.Microbiol.77:3035,2011;Kennedy,et al.,Science See 284:1368, 1999; Lowther, et al., Proc. Natl. Acad. Sci. USA 95:12153, 1998; Abe, et al., Mol. Genet. Genomics 268:130, 2002. US Patent Application Publication No. 2020 / 0211673 A1 provides insight that certain genes present in biosynthetic gene clusters or located in close proximity to biosynthetic genes of a cluster (particularly in eukaryotic, e.g., fungal, biosynthetic gene clusters, as opposed to bacterial, biosynthetic gene clusters) may represent homologs of human genes that are targets for therapeutic purposes.Such genes that are not involved in the synthesis of secondary metabolites produced by a biosynthetic gene cluster are referred to as "integration target genes" ("ETaGs") or "non-integration target genes" (NETaGs), depending on whether they are located within the cluster of biosynthetic genes or not.
[0005] Traditionally, secondary metabolites have been identified from microbial cultures and screened for therapeutic activity against human targets of interest. However, most microorganisms are not culturable, and even BGCs in culturable microorganisms may remain transcriptionally silent under laboratory conditions. Recent developments in nucleic acid and protein sequencing technologies as well as bioinformatics pipelines have made it possible to rapidly identify a large number of BGCs from environmental microorganisms without the need to culture the microorganisms and test the biological activity of the BGCs. See, for example, Palazzotto E. and Weber T. Curr. Opin. Microbiol., 45:109-116, 2018. However, it remains a challenge to accurately define the genomic boundaries of BGCs using purely computational methods. There are also no computational pipelines available to identify genes associated with (but not embedded within) biosynthetic gene clusters or to predict the function of secondary metabolites and predict biosynthetic gene clusters that produce secondary metabolites with an activity of interest. Summary of the Invention
[0006] Disclosed herein are exemplary methods and systems for identifying resistance genes (e.g., integration target genes (ETaGs) and / or non-integration target genes (NETaGs)) associated with core synthases or other genes from biosynthetic gene clusters (BGCs) that encode biosynthetic pathways for secondary metabolites. The described methods and systems can also be used to predict the function of secondary metabolites based on, for example, co-occurrence and / or co-evolution of resistance genes (e.g., ETaGs or NETaGs) with genes of the biosynthetic gene clusters, as well as to predict biosynthetic gene clusters that produce secondary metabolites with desired activities.
[0007] 1. A computer-implemented method for identifying resistance genes, comprising: receiving a selection of at least one target sequence of interest; receiving a selection of target genomes from a genomics database, the selection of target genomes including a plurality of target genomes from organisms known to produce or likely to produce secondary metabolites; performing a search to identify homologs of at least one target sequence in the plurality of target genomes; creating a phylogenetic tree based on the identified homologs of the at least one target sequence; and classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on the phylogenetic tree, wherein a positive genome is a genome that belongs to a clade in which multiple copies of at least one target sequence homolog are present and a negative genome is a genome that belongs to a clade in which a single copy of at least one target sequence homolog is present, and wherein the target sequence homolog present in multiple copies in the positive genome is a putative resistance gene. and determining, based at least in part on the classification of the positive and negative genomes, at least one genomic parameter selected from: i) one or more scores indicative of co-occurrence of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a biosynthetic gene cluster (BGC); ii) one or more scores indicative of co-evolution of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a BGC; iii) one or more scores indicative of co-regulation of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a BGC; and iv) one or more scores indicative of co-expression of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a BGC; and determining a likelihood that the putative resistance gene is a resistance gene based on the at least one genomic parameter.
[0008] In some embodiments, determining the likelihood that the putative NETaG is a non-integration target gene (NETaG) comprises comparing at least one determined genomic parameter to at least one predetermined threshold.
[0009] In some embodiments, the selection of at least one target sequence of interest is provided as an input by a user of a system configured to perform a computer-implemented method. In some embodiments, the at least one target sequence of interest comprises an amino acid sequence, a nucleotide sequence, or any combination thereof. In some embodiments, the at least one target sequence of interest comprises a peptide sequence or a portion thereof, a protein sequence or a portion thereof, a protein domain sequence or a portion thereof, a gene sequence or a portion thereof, or any combination thereof. In some embodiments, the at least one target sequence of interest comprises a mammalian sequence, a human sequence, a plant sequence, a fungal sequence, a bacterial sequence, an archaeal sequence, a viral sequence, or any combination thereof.
[0010] In some embodiments, the at least one target sequence of interest comprises a primary target sequence and one or more related sequences. In some embodiments, the one or more related sequences comprise a sequence that is functionally related to the primary target sequence. In some embodiments, the one or more related sequences comprise a sequence that is pathway related to the primary target sequence.
[0011] In some embodiments, the selection of the target genome is provided as an input by a user of a system configured to perform the computer-implemented method. In some embodiments, the plurality of target genomes comprises plant genomes, fungal genomes, bacterial genomes, or any combination thereof. In some embodiments, the genomics database comprises a public genomics database. In some embodiments, the genomics database comprises a proprietary genomics database.
[0012] In some embodiments, the search for identifying homologs of at least one target sequence comprises identifying homologs based on a probabilistic sequence alignment model.In some embodiments, the probabilistic sequence alignment model is a profile hidden Markov model (pHMM).In some embodiments, the homologs are identified based on comparing the probabilistic sequence alignment model score with a predefined threshold.
[0013] In some embodiments, the search to identify homologs of at least one target sequence includes identifying homologs based on an alignment of sequences using a local sequence alignment search tool, calculating sequence homology metrics based on the alignment, and comparing the calculated sequence homology metrics to a predefined threshold. In some embodiments, the local sequence alignment search tool includes BLAST, DIAMOND, HMMER, Exonerate, or ggsearch. In some embodiments, the predefined threshold includes a threshold for percent sequence identity, percent sequence coverage, E-value, or bit score value.
[0014] In some embodiments, the search to identify homologs of at least one target sequence comprises identifying homologs based on the use of gene and / or protein domain annotation tools. In some embodiments, the gene and / or protein domain annotation tools comprise InterProScan or EggNOG.
[0015] In some embodiments, creating a phylogenetic tree based on the identified homologs of at least one target sequence includes aligning the homolog sequences using an alignment software tool, trimming the aligned homolog sequences using a sequence trimming software tool, and constructing a phylogenetic tree using a tree construction software tool. In some embodiments, the alignment software tool includes MAFFT, MUSCLE, or ClustalW. In some embodiments, the sequence trimming software tool includes trimAI, GBlocks, or ClipKIT. In some embodiments, the tree construction software tool includes FastTree, IQ-TREE, RAxML, MEGA, MrBayes, BEAST, or PAUP. In some embodiments, the construction of the phylogenetic tree is based on a maximum likelihood algorithm, a parsimony algorithm, a neighbor-joining algorithm, a distance matrix algorithm, or a Bayesian inference algorithm.
[0016] In some embodiments, the one or more scores indicating co-occurrence are determined based on identifying a positive correlation between the presence of multiple copies of the putative resistance gene in the positive genome and the presence of one or more genes of the BGC. In some embodiments, identifying a positive correlation between the presence of multiple copies of the putative resistance gene in the positive genome and the presence of one or more genes of the BGC comprises using a clustering algorithm to cluster aligned protein sequences, aligned nucleotide sequences, aligned protein domain sequences, or aligned pHMMs for a group of BGCs to identify a BGC community in the multiple target genomes. In some embodiments, identifying a positive correlation between the presence of multiple copies of the putative resistance gene in the positive genome and the presence of one or more genes of the BGC comprises using a phylogenetic analysis of protein sequences or protein domains for a group of BGCs to identify a BGC community in the multiple target genomes. In some embodiments, identifying a positive correlation between the presence of multiple copies of the putative resistance gene in the positive genome and the presence of one or more genes of the BGC comprises selecting genomes with a particular taxonomy to identify a BGC community in the multiple target genomes.
[0017] In some embodiments, the one or more scores indicative of co-evolution of the putative resistance gene and the one or more genes associated with the BGC are determined based on a co-evolution correlation score, a co-evolution rank score, a co-evolution slope score, or any combination thereof. In some embodiments, the co-evolution correlation score is based on the correlation between the pairwise sequence identity percentage of the cluster of orthologous groups (COGs) for the putative resistance gene and the pairwise sequence identity percentage of the cluster of orthologous groups (COGs) for one of the one or more genes associated with the BGC. In some embodiments, the co-evolution rank score is based on the ranking of correlation coefficients of COGs that include one of the one or more genes associated with the BGC in ascending order for the COGs that include the putative resistance gene. In some embodiments, in the event of a tie in distance scores, the rank for all COGs in the tie is set equal to the lowest rank in the group. In some embodiments, the co-evolution slope score is based on an orthogonal regression of the pairwise sequence identity percentage of the COGs for the putative resistance gene and the pairwise sequence identity percentage of the COGs for one of the one or more genes associated with the BGC. In some embodiments, only COGs resulting from unique positive genomes with three or more genes remaining after removing the corresponding genes from the negative genome are used to assess the coevolution correlation score, coevolution rank score, or coevolution slope score.
[0018] In some embodiments, the one or more scores indicative of co-regulation are based on DNA motif detection from intergenic sequences of one or more genes associated with BGC and putative resistance genes.
[0019] In some embodiments, the one or more scores indicative of co-expression are based on differential expression and / or clustering analysis of global transcriptomic data.
[0020] In some embodiments, the one or more genes associated with a biosynthetic gene cluster (BGC) include anchor genes, core synthase genes, biosynthetic genes, genes not involved in the biosynthesis of a secondary metabolite produced by the BGC, or any combination thereof.
[0021] In some embodiments, the putative resistance gene is a putative integration target gene (pETaG) or a putative non-integration target gene (pNETaG).
[0022] In some embodiments, the resistance gene is an integration targeted gene (ETaG) or a non-integration targeted gene (NETaG).
[0023] Also provided is a computer-implemented method for predicting a function of a secondary metabolite, comprising: receiving a selection of at least one target sequence of interest, where the at least one target sequence of interest corresponds to a gene sequence associated with a biosynthetic gene cluster (BGC) known to produce a secondary metabolite; receiving a selection of a target genome from a genomics database, where the selection of target genomes includes a plurality of target genomes from organisms known to produce a secondary metabolite; performing a search to identify homologs of the at least one target sequence in the plurality of target genomes; creating a phylogenetic tree based on the identified homologs of the at least one target sequence; and classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on the phylogenetic tree. classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on a phylogenetic tree, where a positive genome is a genome belonging to a clade in which multiple copies of at least one target sequence homolog are present, and a negative genome is a genome belonging to a clade in which a single copy of at least one target sequence homolog is present, and the target sequence homolog present in multiple copies in the positive genome is a putative resistance gene; and based at least in part on the classification of the positive genomes and the negative genomes, determining: i) one or more scores indicative of co-occurrence of at least one target sequence homolog (putative resistance gene) and one or more genes associated with the BGC; ii) one or more scores indicative of co-evolution of at least one target sequence homolog (putative resistance gene) and one or more genes associated with the BGC; iii) one or more scores indicative of co-regulation of at least one target sequence homolog (putative resistance gene) and one or more genes associated with the BGC;and iv) determining at least one genomic parameter selected from one or more scores indicative of co-expression of at least one target sequence homolog (putative resistance gene) with one or more genes associated with the BGC, and determining a likelihood that the putative resistance gene is a resistance gene encoding a protein target acted upon by a secondary metabolite based on the at least one genomic parameter;
[0024] In some embodiments, determining the likelihood that the putative resistance gene is a resistance gene encoding a protein target acted upon by a secondary metabolite comprises comparing at least one determined genomic parameter to at least one predetermined threshold.
[0025] In some embodiments, the selection of at least one target sequence of interest is provided as an input by a user of a system configured to perform a computer-implemented method. In some embodiments, the at least one target sequence of interest comprises an amino acid sequence, a nucleotide sequence, or any combination thereof. In some embodiments, the at least one target sequence of interest comprises a peptide sequence or a portion thereof, a protein sequence or a portion thereof, a protein domain sequence or a portion thereof, a gene sequence or a portion thereof, or any combination thereof.
[0026] In some embodiments, the selection of the target genome is provided as an input by a user of a system configured to perform the computer-implemented method. In some embodiments, the plurality of target genomes comprises plant genomes, fungal genomes, bacterial genomes, or any combination thereof. In some embodiments, the genomics database comprises a public genomics database or a proprietary genomics database.
[0027] In some embodiments, the search for identifying homologs of at least one target sequence comprises identifying homologs based on a probabilistic sequence alignment model.In some embodiments, the probabilistic sequence alignment model is a profile hidden Markov model (pHMM).In some embodiments, the homologs are identified based on comparing the probabilistic sequence alignment model score with a predefined threshold.
[0028] In some embodiments, the search to identify homologs of at least one target sequence comprises identifying homologs based on an alignment of sequences using a local sequence alignment search tool, calculating sequence homology metrics based on the alignment, and comparing the calculated sequence homology metrics to a predefined threshold. In some embodiments, the predefined threshold comprises a threshold for percent sequence identity, percent sequence coverage, E value, or bit score value.
[0029] In some embodiments, the search to identify homologs of at least one target sequence comprises identifying homologs based on the use of gene and / or protein domain annotation tools.
[0030] In some embodiments, creating a phylogenetic tree based on the identified homologs of at least one target sequence comprises aligning the homolog sequences using an alignment software tool, trimming the aligned homolog sequences using a sequence trimming software tool, and constructing a phylogenetic tree using a phylogenetic tree building software tool.
[0031] In some embodiments, at least one target sequence of interest comprises a known NETaG sequence or a core synthase gene sequence.
[0032] Disclosed herein is a computer-implemented method for identifying biosynthetic gene clusters (BGCs) encoding biosynthetic enzymes for producing a secondary metabolite having an activity of interest, the method comprising: receiving a selection of at least one target sequence of interest, wherein the at least one target sequence of interest comprises a sequence encoding a therapeutic target of interest; receiving a selection of a target genome from a genomics database, wherein the selection of target genomes comprises a plurality of target genomes from organisms known to produce secondary metabolites; performing a search to identify homologs of the at least one target sequence in the plurality of target genomes; creating a phylogenetic tree based on the identified homologs of the at least one target sequence; and classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on the phylogenetic tree. classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on a phylogenetic tree, wherein a positive genome is a genome belonging to a clade in which multiple copies of at least one target sequence homolog are present, and a negative genome is a genome belonging to a clade in which a single copy of at least one target sequence homolog is present, and the target sequence homolog present in multiple copies in the positive genome is a putative resistance gene; and based at least in part on the classification of the positive genomes and the negative genomes, determining: i) one or more scores indicative of co-occurrence of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a biosynthetic gene cluster (BGC); ii) one or more scores indicative of co-evolution of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a BGC; iii) one or more scores indicative of co-regulation of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a BGC;and iv) determining at least one genomic parameter selected from one or more scores indicative of co-expression of at least one target sequence homolog (putative resistance gene) with one or more genes associated with the BGC, and determining a likelihood based on the at least one genomic parameter that the putative resistance gene is an actual resistance gene associated with the BGC that produces a secondary metabolite that acts on the protein product encoded by the resistance gene;
[0033] In some embodiments, determining the likelihood that the putative resistance gene is an actual resistance gene associated with the BGC that produces a secondary metabolite comprises comparing the at least one determined genomic parameter to at least one predetermined threshold.
[0034] In some embodiments, the selection of at least one target sequence of interest is provided as an input by a user of a system configured to perform a computer-implemented method. In some embodiments, the at least one target sequence of interest comprises an amino acid sequence, a nucleotide sequence, or any combination thereof. In some embodiments, the at least one target sequence of interest comprises a peptide sequence or a portion thereof, a protein sequence or a portion thereof, a protein domain sequence or a portion thereof, a gene sequence or a portion thereof, or any combination thereof.
[0035] In some embodiments, the selection of the target genome is provided as an input by a user of a system configured to perform the computer-implemented method. In some embodiments, the plurality of target genomes comprises plant genomes, fungal genomes, bacterial genomes, or any combination thereof. In some embodiments, the genomics database comprises a public genomics database or a proprietary genomics database.
[0036] In some embodiments, the search for identifying homologs of at least one target sequence comprises identifying homologs based on a probabilistic sequence alignment model.In some embodiments, the probabilistic sequence alignment model is a profile hidden Markov model (pHMM).In some embodiments, the homologs are identified based on comparing the probabilistic sequence alignment model score with a predefined threshold.
[0037] In some embodiments, the search to identify homologs of at least one target sequence comprises identifying homologs based on an alignment of sequences using a local sequence alignment search tool, calculating sequence homology metrics based on the alignment, and comparing the calculated sequence homology metrics to a predefined threshold. In some embodiments, the predefined threshold comprises a threshold for percent sequence identity, percent sequence coverage, E value, or bit score value.
[0038] In some embodiments, the search to identify homologs of at least one target sequence comprises identifying homologs based on the use of gene and / or protein domain annotation tools.
[0039] In some embodiments, creating a phylogenetic tree based on the identified homologs of at least one target sequence comprises aligning the homolog sequences using an alignment software tool, trimming the aligned homolog sequences using a sequence trimming software tool, and constructing a phylogenetic tree using a phylogenetic tree building software tool.
[0040] In some embodiments, the computer-implemented method further comprises performing an in vitro assay to test secondary metabolites produced by the identified BGC for activity against a therapeutic target of interest.
[0041] In some embodiments, the computer-implemented method further comprises performing an in vivo assay to test secondary metabolites produced by the identified BGC for activity against a therapeutic target of interest.
[0042] Also disclosed herein is a system that includes one or more processors and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform any of the methods described herein.
[0043] Disclosed herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a system, cause the system to perform any of the methods described herein.
[0044] It is to be understood that all combinations of the foregoing concepts and additional concepts described in more detail below (to the extent such concepts are not mutually inconsistent) are considered to be part of the inventive subject matter disclosed herein, and in particular, all combinations of claimed subject matter appearing at the end of this disclosure are considered to be part of the inventive subject matter disclosed herein.
[0045] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety. In the event of a conflict between a term in this specification and a term in an incorporated reference, the term in this specification shall control.
[0046] Various aspects of the disclosed methods, devices, and systems are set forth with particularity in the appended claims. A better understanding of the features and advantages of the disclosed methods, devices, and systems will be obtained by reference to the following detailed description of exemplary embodiments and the accompanying drawings. [Brief description of the drawings]
[0047] [Figure 1] FIG. 1 provides a non-limiting example of a process flow chart for identifying putative resistance genes (e.g., putative integration target genes (pETaG) and / or putative non-integration target genes (pNETaG)) and assessing the likelihood that they are actual resistance genes (e.g., EtaG and / or NETaG).
[0048] [Diagram 2] FIG. 2 is a non-limiting schematic diagram of a computing device in accordance with one or more examples of the present disclosure.
[0049] [Diagram 3] FIG. 3 provides a non-limiting example of a maximum likelihood phylogenetic tree of succinate dehydrogenase complex subunit C (SDHC) homologs.
[0050] [Figure 4] FIG. 4 provides an exemplary illustration of a gene cluster comparison plot. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0051] Disclosed herein are exemplary methods and systems for identifying resistance genes (e.g., integration target genes (ETaGs) and / or non-integration target genes (NETaGs)) associated with core synthases or other genes from biosynthetic gene clusters (BGCs) that encode biosynthetic pathways for secondary metabolites. The described methods and systems can also be used to predict the function of secondary metabolites based on, for example, co-occurrence and / or co-evolution of resistance genes (e.g., ETaGs or NETaGs) with genes of the biosynthetic gene clusters, as well as to predict biosynthetic gene clusters that produce secondary metabolites with desired activities.
[0052] In some cases, for example, the disclosed methods include receiving a selection of at least one target sequence of interest; receiving a selection of target genomes from a genomics database, where the selection of target genomes includes a plurality of target genomes from organisms known to produce or likely to produce secondary metabolites; performing a search to identify homologs of at least one target sequence in the plurality of target genomes; creating a phylogenetic tree based on the identified homologs of the at least one target sequence; and classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on the phylogenetic tree, where a positive genome is a genome that belongs to a clade in which multiple copies of at least one target sequence homolog are present, and a negative genome is a genome that belongs to a clade in which a single copy of at least one target sequence homolog is present, and where the target sequence homolog present in multiple copies in the positive genome is a putative NETaG. The method may include classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on the phylogenetic tree, and determining at least one genomic parameter selected from the following based at least in part on the classification of the positive and negative genomes: i) one or more scores indicative of co-occurrence of at least one target sequence homolog (putative NETaG) and one or more genes associated with a biosynthetic gene cluster (BGC); ii) one or more scores indicative of co-evolution of at least one target sequence homolog (putative NETaG) and one or more genes associated with a BGC; iii) one or more scores indicative of co-regulation of at least one target sequence homolog (putative NETaG) and one or more genes associated with a BGC; and iv) one or more scores indicative of co-expression of at least one target sequence homolog (putative NETaG) and one or more genes associated with a BGC, and determining a likelihood that the putative NETaG is a non-integration target gene (NETaG) based on the at least one genomic parameter.
[0053] definition Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0054] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Any reference herein to "or" is intended to include "and / or" unless specifically stated otherwise and includes any and all possible combinations of one or more of the associated listed items.
[0055] As used herein, the terms "includes," "including," "comprises," and / or "comprising" specify the presence of stated features, integers, steps, operations, elements, components, and / or units, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, units, and / or groups thereof.
[0056] As used herein, the term "about" of a number refers to ±10% of that number. When used in the context of a range, the term "about" refers to the range minus 10% of its lowest value and plus 10% of its highest value.
[0057] As used herein, "secondary metabolite" refers to a small organic molecule compound produced by archaea, bacteria, fungi, or plants that is not directly involved in the normal growth, development, or reproduction of the host organism, but is required for the host organism's interaction with its environment. Secondary metabolites are also known as natural products or genetically encoded small molecules. The term "secondary metabolite" is used interchangeably herein with "biosynthetic product" when referring to the product of a biosynthetic gene cluster.
[0058] The term "biosynthetic gene cluster" or "BGC" is used interchangeably herein and refers to a locally clustered group of one or more genes that together encode a biosynthetic pathway for the production of a secondary metabolite. Exemplary BGCs include, but are not limited to, biosynthetic gene clusters for the synthesis of nonribosomal peptide synthetases (NRPSs), polyketide synthases (PKSs), terpenes, and bacteriocins. See, for example, Keller N, “Fungal secondary metabolism: regulation, function and drug discovery.” Nature Reviews Microbiology 17.3(2019):167-180, and Fischbach M.and Voigt CA, PROKARYOTIC GENE CLUSTERS: A RICH TOOLBOX FOR SYNTHETIC BIOLOGY. In: Institute of Medicine(US)Forum on Microbial Threats.The Science and Applications of Synthetic and Systems Biology:Workshop Summary.Washington(DC):National Academies Press(US);2011.A21. BGCs contain genes encoding signature biosynthetic proteins characteristic of each type of BGC. The longest biosynthetic gene in a BGC is referred to herein as the “core synthase gene” of the BGC. In addition to genes involved in the biosynthesis of secondary metabolites, BGCs may also contain other genes, e.g., genes encoding products not involved in the biosynthesis of secondary metabolites that are interspersed among the biosynthetic genes. These genes are referred to herein as "associated" with a BGC if their products are functionally related to a secondary metabolite of the BGC. Some genes, e.g., genes not involved in the biosynthesis of a secondary metabolite produced by the BGC, are referred to herein as "integrated" with a BGC if their products are functionally related to a secondary metabolite of the BGC and are located in physical proximity to the biosynthetic genes of the cluster.Some genes, e.g., genes not involved in the biosynthesis of a secondary metabolite produced by a BGC, are referred to herein as "non-integrated" if their products are functionally related to the secondary metabolite of the BGC but are not located in physical proximity to the biosynthetic genes of the BGC. An "anchor gene" refers to a biosynthetic gene or a gene not involved in the biosynthesis of a secondary metabolite produced by a BGC that is known to be co-localized and functionally related (i.e., associated) with the BGC.
[0059] The term "co-localized" refers to the presence of two or more genes in close spatial locations, such as about 200 kb or less, about 100 kb or less, about 50 kb or less, about 40 kb or less, about 30 kb or less, about 20 kb or less, about 10 kb or less, about 5 kb or less, or less, within a genome.
[0060] The term "homolog" refers to genes that are part of a gene group that are related by descent from a common ancestor (i.e., the gene sequences (i.e., nucleic acid sequences) of the gene group and / or the sequences of their protein products are inherited through a common origin. Homologs can arise through speciation events (giving rise to "orthologs"), gene duplication events, or horizontal gene transfer events. Homologs can be identified by phylogenetic methods, through the identification of common functional domains in aligned nucleic acid or protein sequences, or through sequence comparison.
[0061] The term "ortholog" refers to a gene that is part of a group of genes predicted to have evolved from a common ancestral gene by speciation.
[0062] The terms "bidirectional best hit" and "BBH" are used interchangeably herein and refer to the relationship between a pair of genes in two genomes (i.e., a first gene in a first genome and a second gene in a second genome), where the first gene or its protein product is identified as having the most similar sequence in the first genome compared to the second gene or its protein product in the second genome, and the second gene or its protein product is identified as having the most similar sequence in the second genome compared to the first gene or its protein product in the first genome. The first gene is the bidirectional best hit (BBH) of the second gene, and the second gene is the bidirectional best hit (BBH) of the first gene. BBH is a commonly used method to infer orthology.
[0063] As used herein, "sequence similarity" between two genes means similarity in either the nucleic acid (e.g., mRNA) sequences encoded by the genes or the amino acid sequences of the gene products.
[0064] "Percent (%) sequence identity" or "percent (%) sequence homology" with respect to a nucleic acid sequence (or protein sequence) described herein is defined as the percentage of nucleotide residues (or amino acid residues) in a candidate sequence that are identical or homologous to the nucleotide residues (or amino acid residues) in an oligonucleotide (or polypeptide) to which the candidate sequence is compared, after aligning the sequences and taking into account any conservative substitutions as part of the sequence identity. Homology between different amino acid residues in a polypeptide is determined based on a substitution matrix such as BLOSUM (BLOcks SUbstitution Matrix). Methods for aligning sequences and determining percent sequence identity or percent sequence homology of nucleic acid or protein sequences are well known to those skilled in the art. Examples of publicly available computer software that can be used include, but are not limited to, BLAST (Basic Local Alignment Search Tool; software for comparing amino acid sequences of proteins or nucleotide sequences of DNA and / or RNA molecules), BLAST-2, ALIGN or Megalign (DNASTAR) software. Any of a variety of appropriate parameters for measuring sequence alignment and determining percent sequence identity or percent sequence homology can be determined by one of skill in the art, including the use of algorithms necessary to achieve maximal alignment over the full length of the sequences being compared.
[0065] Certain aspects of the present disclosure include process steps and instructions described herein in the form of an algorithm. It should be noted that the process steps and instructions of the present disclosure may be implemented in software, firmware, or hardware, and when implemented in software, may be downloaded to reside and operate on different platforms used by various operating systems. Unless otherwise noted, as will be apparent from the following description, throughout the specification, descriptions utilizing terms such as "processing," "calculating," "determining," "displaying," "producing," and the like, are understood to refer to the operations and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities in the memory or registers of the computer system, or other such information storage, transmission, or display device.
[0066] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described. The description is presented to enable one of ordinary skill in the art to make and use the invention and is provided in the context of a patent application and its requirements.
[0067] Methods for identifying resistance genes (e.g., integration target genes and / or non-integration target genes) Microorganisms synthesize secondary metabolites (SMs) using enzymes encoded by clusters of co-located genes called biosynthetic gene clusters (BGCs). Since many SMs target primary metabolic enzymes, BGCs may contain so-called "resistance copies" of primary metabolic genes that have been described as "self-protecting" because they encode proteins that may render the host organism resistant to the SMs produced by the BGC. This "resistance copy" (or "resistance gene") can be referred to as an "integrated target gene" or "ETaG" if it is located in close proximity to one of the biosynthetic genes of the BGC, or as a "non-integrated target gene" or "NETaG" if it is not located in close proximity to one of the biosynthetic genes of the BGC. Identification of ETaGs and NETaGs can be useful in determining the role of SMs synthesized by the enzymes of the cluster. Methods for identifying ETaGs are described in co-pending International Patent Application No. PCT / US2022 / 049016 and International Patent Application No. PCT / US2022 / 049040, the contents of each of which are incorporated by reference in their entirety herein.
[0068] Current BGC prediction tools predict regions in the genome that constitute a BGC (see the section entitled "BGC annotation" below for details). Potential targets of SMs synthesized by the BGC can then be predicted by identifying ETaGs within the genomic boundaries of the biosynthetic cluster. One example is the 3-hydroxy-3-methylglutaryl-CoA (HMG-CoA) gene ETaG, located within the lovastatin gene cluster, which confers resistance to lovastatin. Another example is the inosine-5'-monophosphate dehydrogenase (IMPDH) ETaG (gene name: mpaF), located within the mycophenolic acid gene cluster, which confers resistance to mycophenolic acid. This approach relies on ETaGs embedded within or located in close proximity to the BGC to predict the function of the BGC, and therefore cannot predict the function of BGCs that do not contain an ETaG.
[0069] This disclosure describes (i) a position-independent method for identifying non-integration target genes (NETaGs), (ii) a method for predicting the function of SMs by correlation and / or co-evolution of BGCs and / or core enzymes with NETaGs, and (iii) a method for predicting BGCs responsible for producing SMs with desired activities. The methods described herein are superior to previous approaches because they allow for the detection of BGCs of interest and the prediction of their function independent of the position of the target gene.
[0070] Target of interest: The methods described herein can be performed using a target of interest (or target sequence of interest), which can be any amino acid sequence or nucleotide sequence from the genome of any type of organism, including but not limited to a mammalian genome, a human genome, an avian genome, a reptile genome, a plant genome, a fungal genome, a bacterial genome, an archaeal genome, a viral genome, etc. A target of interest can include any type of biological sequence, such as a gene sequence or a portion thereof, a protein sequence or a portion thereof, a protein domain sequence or a portion thereof, a peptide sequence or a portion thereof, etc.
[0071] Target genome selection: The methods described herein are suitable for identifying ETaGs, NETaGs, and / or BGCs in any type of target genome that contains a BGC or genome for an organism known to produce secondary metabolites. Bacterial, plant, and fungal genomes are known to encode biosynthetic gene clusters. Without wishing to be bound by any theory or hypothesis, fungal genomes are eukaryotic genomes that are phylogenetically related to mammalian genomes than bacterial or plant genomes. Thus, fungal genomes may be preferred for identifying ETaGs or NETaGs that are homologous to human target genes and encode targets of secondary metabolites produced by BGCs.
[0072] Target Homolog Searching: Protein or DNA homologs for a target sequence within a given genome can be detected using, for example: 1) Probabilistic sequence alignment models, including, for example, profile hidden Markov models (pHMMs), by comparing the probabilistic model scores to one or more predefined thresholds (e.g., a confidence cutoff threshold). In some cases, such predefined thresholds may be determined, for example, based on the lowest bit scores of known homologs. 2) Sequence alignment tools including BLAST (basic local alignment search tool), DIAMOND, HMMER, Exonerate, or ggsearch, for example, by comparing sequence alignment or sequence homology metrics such as percent sequence identity, percent sequence coverage, E-value, or bit score to one or more predefined thresholds. 3) Gene sequence / protein domain annotation tools, such as InterProScan or EggNOG.
[0073] Creating a phylogenetic tree based on the identified target homologues: The homologues of the target sequence(s) identified as a result of performing the target homology search can be used to create a phylogenetic tree of the selected target genome. To determine phylogenetic distance, the protein or DNA homologues of the target(s) can be individually aligned using any alignment software (e.g., MAFFT, MUSCLE, or ClustalW, etc.), trimmed to remove gaps using any sequence trimming software (e.g., trimAI, GBlocks, or ClipKIT), and then multiple sequence alignments can be performed using any phylogenetic tree building software known to those skilled in the art (e.g., FastTree, IQ-TREE, RAxML, MEGA, MrBayes, BEAST, or PAUP) to provide a phylogenetic tree of homologous sequences (e.g., homologous gene sequences or protein sequences). The phylogenetic tree can be constructed using any of a variety of different algorithms known to those skilled in the art, including, but not limited to, maximum likelihood algorithms, parsimony algorithms, neighbor-joining algorithms, distance matrix algorithms, or Bayesian inference algorithms.
[0074] Distinguishing house copies and additional copies of candidate NETaG: From the phylogenetic tree, two groups (clades) of genomes containing homologs of the target sequence(s) (e.g., target genes) can be identified. One clade contains genomes containing a single copy of the target gene homolog(s), indicating that the homolog is a "house copy" of the gene (i.e., the single copy of the target gene homolog(s) present in organisms of the first clade is assumed to have only a housekeeping function. The other clade contains genomes containing additional copies of the target gene homolog required for normal functioning of primary metabolism in the presence of BGC products (i.e., multiple copies of the target gene homolog in organisms of the second clade are assumed to be potential resistance-related genes due to their increased copy number). Thus, the target gene homologs present in multiple copies may be candidate (or putative) NETaG.
[0075] In some cases, other targets that may be related or correlated to the primary target of interest can be examined. These relationships can be functional relationships (e.g., genes that share similar functions) or pathway relationships (e.g., genes that are members of the same pathway). For example, if KRAS is the primary target, copy number variations of additional RAS homologs such as HRAS, NRAS, MRAS, ERAS, RRAS2, RRAS, etc. can also be examined. In addition, copy number variations of genes in the RAS pathway such as RAS-GEF, RAS-GAP, RAF, MEK, ERK, PI3K, PDK1, AKT, etc. can also be examined. Thus, genomes with high copy numbers of genes functionally related to the primary target or pathway-related genes may harbor additional candidate NETaG in the form of functionally related or pathway-related genes.
[0076] Positive and negative genome classification: Genomes are classified based on the number of target homologs or genes related to target homologs encoded therein. In genomes that encode multiple copies of a target homolog, one of the copies is assumed to be resistant to a particular BGC product and is therefore required for primary metabolism to function when the BGC product is present. Genomes containing multiple copies of a target homolog are classified as positive genomes, while genomes containing a single copy of a target homolog are classified as negative genomes. Target homologs present in multiple copies may contain putative integrated or non-integrated target genes. The positive and negative genomes can be used to calculate several different genomic metrics (described in the following sections) that can be used to determine whether a putative target gene identified in a phylogenetic tree is an actual ETaG or NETaG.
[0077] BGC annotation: Identification and annotation of biosynthetic gene clusters includes identification of secondary metabolism genes and prediction of secondary metabolism gene groups that compose BGCs. Secondary metabolism genes (or their corresponding proteins or protein domains) are genes or gene products that are not involved in primary metabolism. Examples of secondary metabolism genes include, but are not limited to, genes encoding core enzymes, such as polyketide synthases (PKSs), non-ribosomal peptide synthetases (NRPSs), enzymes containing NRPS or PKS domains (e.g., PKS-like enzymes, NRPS-like enzymes, NRPS-PKSs or PKS-NRPS hybrids), terpene synthases (TPs), enzymes that synthesize isoprenoids, enzymes that synthesize beta-lactones, ribosomal synthesis and post-translationally modified proteins (RIPPSs), or any combination thereof, which are sometimes co-localized with individualized enzymes. Examples of regulatory enzymes include, but are not limited to, cytochrome P450s (CYPs), methyltransferases, glycosyltransferases, and the like.
[0078] BGCs can be predicted using any of a variety of software tools known to those of skill in the art, including, but not limited to, BLAST, pHMM, Antibiotic Secondary Metabolite Analysis Shell (antiSMASH), Secondary Metabolite Unknown Region Finder (SMURF), DeepBGC, or custom BGC prediction tools.
[0079] Correlation Analysis: Clusters of orthologous groups (COGs) are collections of homologous genes that are useful for studying evolutionary relationships. COGs consist of orthologs (homologous genes that have diverged in different species from a common ancestral gene) and paralogs (genes in a single species that have arisen by duplication and divergence). See, for example, Tatusov, et al. (1997), "A Genomic Perspective on Protein Families", Science 278:631-637. COGs for genes or proteins encoded thereby can be identified by performing an all-versus-all protein (amino acid) sequence search (or all-versus-all nucleotide sequence search) of all positive and negative genomes using sequence alignment software such as BLAST, DIAMOND, or ggsearch.
[0080] In some cases, clustering algorithms such as MCL, mmseq, usearch, CD-hit, etc. are used to identify reciprocal best hits (i.e., where the best match of a protein / gene from genome A to genome B is the same as the best match of a protein / gene from genome B to genome A) and clustered into COGs. Alternatively, in some cases, unidirectional search results (rather than reciprocal search results) can be used to identify homologous proteins / genes prior to clustering.
[0081] COGs can also be identified using software tools such as OrthoMCL or OrthoFinder (or other orthogonal group / pangenome identification tools), or using protein or nucleotide clustering tools such as USEARCH, CD-HIT, and MMseq.
[0082] Coevolutionary Analysis: For coevolutionary analysis, all genes from the negative genome are removed from consideration for all COGs. Then, only COGs with more than three remaining genes, each originating from a unique genome, are passed to coevolutionary analysis.
[0083] A multiple protein (amino acid) or DNA (nucleotide) sequence alignment is performed for all remaining COGs, for example using MAFFT or any other sequence alignment software. For each COG, all pairwise alignments can be trimmed based on a specified set of parameters (e.g., removing all gaps, removing gaps larger than a specified threshold (e.g., gaps exceeding 30%, 20% or 10% of sequences in the aligned sequences), retaining all gaps, etc.), followed by calculation of percent sequence identity (e.g., number of identical residues in the alignment). Alternatively, sequence similarity scores can be calculated based on the use of substitution matrices such as BLOSUM or PAM (e.g., when protein sequences are used). The higher the percent sequence identity between two protein sequences, the more likely they are homologs and the more likely they are assigned to the same COG. Coevolution can be identified when changes in percent sequence identity of proteins in one COG correlate with changes in percent sequence identity of proteins in another COG.
[0084] Alternatively, in some cases, a phylogenetic tree can be calculated from the sequences (nucleotide or amino acid sequences) within each COG (e.g., by performing alignment, trimming, and phylogenetic reconstruction). The phylogenetic tree must be constrained to the topology of the species tree of the genomes to which the genes present in both COGs are compared (e.g., after performing a step of removing genes of the negative genome from consideration and analyzing only the remaining COGs with three or more genes). Since the two COG trees are constrained to the species tree topology, they share exactly the same topology. In this sense, all nodes and branches are in exactly the same position, but the branch lengths (indicating the degree of branching and provided as output by phylogenetic software tools such as RAxML, FastTree, IQTREE, PAULP, BEAST, etc.) may differ between the COG trees. For example, the branch length between Node_A and COG_1_genome_x may be 0.05, and the branch length between Node_A and COG_2_genome_x may be 0.075. These pairwise associations are then recorded for use in performing correlation analysis (described below). Alternatively, in some cases, the minutiae may be raw output from a phylogenetic software tool, or may be normalized by the branch lengths of a constrained species tree, or may be normalized using a Z-score transformation or similar transformation metric. This analysis can be performed using custom scripts or using tools such as the Co-Variance algorithm in PhyKIT (https: / / github.com / JLSteenwyk / PhyKIT).
[0085] The pairwise sequence identity percentage, sequence similarity percentage or branch length (between pairs of genomes) of each COG is then used to calculate the degree of correlation of all pairwise COG combinations, for example using Pearson R or any other correlation metric. Correlation is only calculated between pairs of COGs that share at least three genomes.
[0086] Three different coevolutionary correlation metrics can be used to distinguish ETaGs or NETaGs from putative target genes and to identify candidate BGCs of ETaGs or NETaGs. (i) Coevolutionary Correlation: COG x Pairwise sequence identity percentage and COGs y Correlation with percent pairwise sequence identity. (ii) Coevolutionary Rank: Rank of the correlation coefficient of COGs containing core synthases in ascending order relative to COGs containing pETaG or pNETaG. In case of a tie in distance score, all tied COGs are ranked lowest in the group. (iii) Coevolutionary gradient: COG x Pairwise sequence identity percentage and COGs y Orthogonal regression with percent pairwise sequence identity.
[0087] Co-occurrence analysis: In order to correlate the presence of a candidate BGC with additional copies of the target gene homologue in a given genome with stronger statistical power, it is necessary to limit the number of candidate BGCs by creating a BGC "community" in a selected group of genomes. One approach to do this involves the use of a method to group BGCs based on the alignment of all protein or nucleotide sequences within a given BGC (i.e., clustering BGCs into gene cluster families containing orthologous BGCs). Alignment between all protein or nucleotide sequences of a group of BGCs is performed using an alignment search tool, such as one of the programs included in the BLAST+ suite or DIAMOND. The alignments are then summarized by a cluster score that describes the similarity of the BGCs. To create a cluster score, for example, an average sequence identity percentage score for the comparison of the BGC to the BGC can be created by summing the sequence identity percentages of the protein sequence alignments between the BGC proteins and dividing by the total number of biosynthetic proteins in the BGC. A community of BGCs is generated by processing a subset of the cluster scores of the hits (i.e., BGCs that meet a threshold of at least 20%, 30%, 40%, or more than 40% average sequence identity percentage) using a community detection algorithm. Examples of BGC community detection algorithms include, but are not limited to, Cluster Walktrap (from https: / / igraph.org / ) or Markov Clustering (MCL).
[0088] Alternatively, in some cases, instead of using complete protein (or amino acid) sequences, clusteromics can be performed on a set of protein domains (or pHMMs), or phylogenetic analysis of protein domains or protein sequences of BGCs can be used to generate a community of BGCs.
[0089] Taxonomy: In some cases, the number of candidate BGCs can be limited by selecting genomes with specific classification at any level, e.g., species, genus, family, order, class, domain, etc. Genome classifications can be annotated by comparing them to known reference sequences, e.g., based on ribosomal RNA sequences, internal transcribed spacer (ITS) sequences, single copy marker gene sequences, etc.
[0090] Phylogenetic trees: In some cases, specific sequences, such as those of single copy proteins or genes, or ITS regions, can be used to generate phylogenetic trees from a set of genomes. Genomes from specific clades in the phylogenetic tree can be selected to limit the number of genomes used in the co-occurrence analysis.
[0091] Co-occurrence-based candidate BGC detection: To identify relevant candidate BGCs that produce secondary metabolites with activity against the product of the target gene, the presence of predicted BGCs is compared with the presence of single and multi-copy target gene homologs in the genome of the selected organism. Candidate BGCs with a hypothesized function against the target gene product should show a positive correlation with the presence of additional copies of the target gene homolog (ETaG or NETaG clades in the phylogenetic tree), and candidate BGCs should show a negative correlation with the presence of a single copy of the target gene homolog.
[0092] In some cases, the normalized distance can be used to identify top candidate BGC hits, for example for use in drug development. Total positive genomes (TPG) represents the number of genomes in the ETaG or NETaG clades of the phylogenetic tree, and positive genomes (PG) represents the number of positive genomes in the BGC community. Total negative genomes (TNG) represents the number of genomes with only a single "house copy" target gene homolog, and negative genomes (NG) represents the number of negative genomes in the BGC community. The normalized distance is given by the following formula:
number
[0093] Co-regulation: Since functionally related genes are often co-regulated, the determination of co-regulation can serve as an additional layer of information in connecting ETaGs or NETaGs to their associated BGCs. This can be achieved by identifying common regulatory signatures, such as the presence of common putative cis-regulatory elements or transcription factor binding sites (TFBS) in the promoter regions of ETaGs or NETaGs and candidate BGCs. The method for identifying co-regulated ETaGs or NETaGs and BGCs is as follows. 1. To identify co-regulated genes, first extract the intergenic regions (ranging from 100 bp to 5,000 bp) of all genes of the candidate BGCs, or COGs of the candidate core synthase genes. 2. Next, perform de novo DNA motif detection on these intergenic regions using motif detection software such as MEME (Bailey, et al. (2015) “The MEME Suite”, Nucleic Acids Res. 43(W1):W39-49) or HOMER (Heinz, et al. (2010), “Simple Combinations of Lineage-Determining Transcription Factors Prime cis-Regulatory Elements Required for Macrophage and B Cell Identities”, Mol Cell 38(4):576-89). 3. The putative TFBSs, represented as a position weight matrix, identified for each BGC or COG by this analysis can then be used to search the promoter regions of the target ETaGs or NETaGs to assess whether these motifs are conserved in these regions. 4. Alternatively, newly detected motifs from the ETaG or NETaG COGs can be directly compared to motifs detected from candidate BGCs or core synthase COGs to assess the similarity of these motifs. 5. Detection of a BGC or core synthase motif in the promoter region of ETaG or NETaG, or a good match between a BGC / core synthase motif and an ETaG or NETaG motif, provides evidence of an association between the two.
[0094] Co-expression: Functionally related genes are often co-expressed under all or even a subset of conditions, so ETaG or NETaG that functions as a resistance gene to a BGC is expected to be co-expressed with a BGC gene. Thus, transcriptomics analysis can be used to associate ETaG or NETaG with its cognate BGC. Data obtained from transcriptional analyses such as qPCR, microarray, RNA-seq, NanoString, etc., performed under multiple growth conditions (e.g., using different media during fermentation to induce expression of BGC and resistance genes) or over time can be used to evaluate the correlation of expression between ETaG or NETaG and candidate BGC genes. Candidate BGCs that are co-expressed with ETaG or NETaG can be identified as follows: 1. Map all transcriptomic data (e.g., RNA-seq data) from multiple conditions or time points to a reference genome, then calculate read counts, normalize, and perform differential expression analysis using well-established pipelines (such as Bowtie, TopHat, Cufflinks, Cuffdiff, EdgeR, or DESeq) or in-house developed pipelines. 2. A clustering algorithm such as K-means clustering, centroid-based clustering, density-based clustering, or hierarchical clustering is then used to identify genes that are co-expressed with each other under all conditions analyzed, using the normalized read counts for each gene as input for the clustering analysis. 3. Alternatively, a biclustering approach, which clusters based on both genes and conditions, can be used to group genes that are co-expressed with each other under all or a subset of the conditions analyzed. 4. BGCs identified as co-expressed with ETaG or NETaG can be considered as potential candidates.
[0095] FIG. 1 provides a non-limiting example of a flow chart of a process 100 for identifying putative resistance genes (e.g., putative integration target genes (pETaG) and / or putative non-integration target genes (pNETaG)) and assessing the likelihood that they are actual resistance genes (e.g., EtaG and / or NETaG). Process 100 can be implemented as a computer-implemented method using software implemented on one or more processors of one or more electronic devices, computers, or computing platforms, for example. In some examples, process 100 is implemented using a client-server system, with blocks of process 100 being divided in any manner between a server and a client device. In other examples, blocks of process 100 are divided between a server and multiple client devices. Thus, although portions of process 100 are described herein as being implemented by a particular device of a client-server system, it will be understood that process 100 is not so limited. In other examples, process 100 is implemented using only a client device or multiple client devices. Some blocks are optionally combined, the order of some blocks is optionally changed, and some blocks are optionally omitted in process 100. In some examples, additional steps may be performed in combination with process 100. Thus, the operations illustrated (and described in more detail below) are exemplary in nature and, therefore, should not be considered limiting.
[0096] In step 102 of Figure 1, at least one target sequence of interest, e.g., an amino acid sequence or a corresponding nucleotide sequence for a potential therapeutic target, is selected and / or received as an input. In some cases, the selection of the at least one target sequence of interest may be provided as an input by a user of a system configured to perform the computer-implemented method. In some cases, the at least one target sequence may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 100, 1000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, or more than 100,000 target sequences (or any number of target sequences within this range).
[0097] In some cases, the at least one target sequence of interest may comprise an amino acid sequence, a nucleotide sequence, or any combination thereof. In some cases, the at least one target sequence of interest may comprise a peptide sequence or a portion thereof, a protein sequence or a portion thereof, a protein domain sequence or a portion thereof, a gene sequence or a portion thereof, or any combination thereof.
[0098] In some cases, the at least one target sequence of interest may comprise a mammalian sequence, a human sequence, an avian sequence, a reptilian sequence, an amphibian sequence, a plant sequence, a fungal sequence, a bacterial sequence, an archaeal sequence, a viral sequence, or any combination thereof. For example, in some cases, the at least one target sequence of interest may comprise a mammalian target sequence, a human target sequence, an avian target sequence, a reptilian target sequence, an amphibian target sequence, a plant target sequence, a fungal target sequence, a bacterial target sequence, an archaeal target sequence, a viral target sequence, or any combination thereof. In some cases, the target sequence (e.g., a human target sequence) may be a therapeutic target sequence (e.g., a human therapeutic target sequence) or a protein encoded thereby.
[0099] In some cases, the at least one target sequence of interest comprises a primary target sequence and one or more related sequences. In some cases, as described above, the one or more related sequences may comprise sequences that are functionally related to the primary target sequence. In some cases, the one or more related sequences may comprise sequences that are pathway related to the primary target sequence.
[0100] In step 104 of Figure 1, a target genome(s) is selected and / or received as input, where the selection includes multiple target genomes from organisms known to or likely to produce secondary metabolites. In some cases, for example, the multiple target genomes include plant genomes, fungal genomes, bacterial genomes, or any combination thereof.
[0101] In some cases, the selection of the target genome is provided as an input by a user of a system configured to perform the computer-implemented method. The target genome can be selected, for example, from a genomics database, such as a public genomics database or a proprietary genomics database. In some cases, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 500, 1,000, 5,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, or more than 100,000 target genomes (or any number of target genomes within this range) can be selected.
[0102] In step 106 of FIG. 1, a search is performed to identify homologs of at least one target sequence in multiple target genomes.
[0103] In some cases, the search for identifying the homologue of at least one target sequence can include identifying the homologue based on a probabilistic sequence alignment model, such as a profile hidden Markov model (pHMM).In some cases, the homologue of at least one target sequence can be identified based on the comparison of the probabilistic sequence alignment model score with a predetermined threshold.In some cases, such a predetermined threshold can be determined based on, for example, the lowest bit score of known homologues.
[0104] In some cases, the search for identifying homologs of at least one target sequence may include identifying homologs based on sequence alignment using a local sequence alignment search tool, calculating sequence homology metrics based on the alignment, and comparing the calculated sequence homology metrics with a predetermined threshold.In some cases, for example, the local sequence alignment search tool may include BLAST, DIAMOND, HMMER, Exonerate, or ggsearch.In some cases, the predetermined threshold includes a threshold for sequence identity percentage, sequence coverage percentage, E value, or bit score value.
[0105] In some cases, the predetermined threshold for percent sequence identity may be at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more.
[0106] In some cases, the predetermined threshold for percent sequence coverage may be at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more.
[0107] In some cases, the predefined thresholds for E-values are up to 10, 1, 0.1, 0.001, 0.0001, 1e -10 , 1e -20 , 1e -100 , or may be lower.
[0108] In some cases, the predetermined threshold value for the bit score may be at least 5, 10, 25, 50, 100, 250, 500, 1000, 5000, or more.
[0109] In some cases, the search to identify homologs of at least one target sequence may include identifying homologs based on the use of gene and / or protein domain annotation tools. For example, in some cases, the gene and / or protein domain annotation tools may include InterProScan or EggNOG.
[0110] In step 108 of FIG. 1, a phylogenetic tree is generated based on the identified homologs of at least one target sequence, as described elsewhere herein.
[0111] In some cases, creating a phylogenetic tree based on the identified homologs of at least one target sequence may include one or more of: (i) aligning the homolog sequences using an alignment software tool; (ii) trimming the aligned homolog sequences using a sequence trimming software tool; and (iii) constructing a phylogenetic tree using a tree building software tool. In some cases, the alignment software tool may include, for example, MAFFT, MUSCLE, or ClustalW. In some cases, the sequence trimming software tool may include, for example, trimAI, GBlocks, or ClipKIT.
[0112] In some cases, the phylogenetic tree construction software tool may include, for example, FastTree, IQ-TREE, RAxML, MEGA, MrBayes, BEAST, or PAUP. The construction of the phylogenetic tree may be based on any of a variety of algorithms known to those of skill in the art, such as a maximum likelihood algorithm, a parsimony algorithm, a neighbor-joining algorithm, a distance matrix algorithm, or a Bayesian inference algorithm.
[0113] In step 110 of FIG. 1 , genomes of the multiple target genomes are classified as positive genomes or negative genomes based on a phylogenetic tree (as described elsewhere herein), where a positive genome is a genome that belongs to a clade in which multiple copies of at least one target sequence homolog are present, a negative genome is a genome that belongs to a clade in which a single copy of at least one target sequence homolog is present, and the target sequence homolog present in multiple copies in the positive genome is a putative resistance gene (e.g., ETaG or NETaG).
[0114] In step 112 of FIG. 1 , at least one genomic parameter is determined based at least in part on the classification of positive and negative genomes, where the at least one genomic parameter is selected from: i) one or more scores indicative of co-occurrence of at least one target sequence homolog (putative ETaG or NETaG) and one or more genes associated with a biosynthetic gene cluster (BGC), as described elsewhere herein; ii) one or more scores indicative of co-evolution of at least one target sequence homolog (putative ETaG or NETaG) and one or more genes associated with a BGC, as described elsewhere herein; iii) one or more scores indicative of co-regulation of at least one target sequence homolog (putative ETaG or NETaG) and one or more genes associated with a BGC; and iv) one or more scores indicative of co-expression of at least one target sequence homolog (putative ETaG or NETaG) and one or more genes associated with a BGC.
[0115] In some cases, one or more scores indicating co-occurrence are determined based on identifying a positive correlation between the presence of multiple copies of a putative ETaG or NETaG and the presence of one or more genes of the BGC identified in the positive genome.
[0116] In some cases, identifying a positive correlation between the presence of multiple copies of a putative ETaG or NETaG and the presence of one or more genes of an identified BGC in a positive genome may include using a clustering algorithm to cluster aligned protein sequences, aligned nucleotide sequences, aligned protein domain sequences, or aligned pHMMs for a group of BGCs to identify a BGC community within multiple target genomes.
[0117] In some cases, identifying a positive correlation between the presence of multiple copies of a putative ETaG or NETaG and the presence of one or more genes of a BGC identified in a positive genome may involve the use of phylogenetic analysis of protein sequences or protein domains of a group of BGCs to identify a BGC community within multiple target genomes.
[0118] In some cases, identifying a positive correlation between the presence of multiple copies of a putative ETaG or NETaG and the presence of one or more genes of an identified BGC in a positive genome may include selecting genomes with a particular taxonomy to identify a BGC community within multiple target genomes.
[0119] In some cases, one or more scores indicative of coevolution of one or more genes associated with a putative ETaG or NETaG and a BGC may be determined based on a coevolution correlation score, a coevolution rank score, a coevolution slope score, or any combination thereof.
[0120] In some cases, the coevolution correlation score (or coevolution correlation coefficient) may be based on a correlation between the pairwise sequence identity percentage of a cluster of orthologous groups (COG) for a putative ETaG or NETaG (as described elsewhere herein) and the pairwise sequence identity percentage of a cluster of orthologous groups (COG) for one of the one or more genes associated with the BGC. In some cases, the coevolution correlation score (or coevolution correlation coefficient) may range in value from -1.0 to 1.0. In some cases, the coevolution correlation score (or coevolution correlation coefficient) may have a value of -1.0, -0.8, -0.6, -0.4, -0.2, 0, 0.2, 0.4, 0.6, 0.8, 1.0, or any value within this range.
[0121] In some cases, the coevolution rank score (or coevolution rank) may be based on a ranking of correlation coefficients of COGs that include one of the one or more genes associated with the BGC in ascending order to COGs that include a putative ETaG or NETaG (as described elsewhere herein). In some cases, the coevolution rank may range in value from 1 to 10,000. In some cases, the coevolution rank may have a value of 1, 10, 20, 40, 60, 80, 100, 200, 400, 600, 800, 1000, 2000, 4000, 6000, 8000, or 10,000, or any value within this range. In the case of a tie in distance scores, the ranks of all COGs in the tie may be set equal to the lowest rank in the group.
[0122] In some cases, the coevolutionary slope score may be based on an orthogonal regression of the percent pairwise sequence identity of the COG to the putative ETaG or NETaG (as described elsewhere herein) and the percent pairwise sequence identity of the COG to one of the one or more genes associated with the BGC. In some cases, the coevolutionary slope score may range in value from about 0.75 to about 1.25. In some cases, the coevolutionary slope score may have a value of at least 0.75, at least 0.8, at least 0.85, at least 0.9, at least 0.95, at least 1.0, at least 1.05, at least 1.1, at least 1.15, at least 1.20, or at least 1.25. In some cases, the coevolutionary slope score may have a value of at most 1.25, at most 1.20, at most 1.15, at most 1.10, at most 1.10, at most 1.05, at most 1.0, at most 0.95, at most 0.90, at most 0.85, at most 0.80, or at most 0.75. Any of the lower and upper limits listed in this paragraph may be combined to form ranges included within the present disclosure, for example, in some cases, the coevolutionary slope score may range from about 0.80 to about 1.1. One of skill in the art will recognize that the coevolutionary slope score may have any value within this range, for example, about 0.98.
[0123] In some cases, only COGs resulting from unique positive genomes with three or more genes remaining after removing the corresponding genes from the negative genome are used to assess the coevolution correlation score, coevolution rank score, or coevolution slope score.
[0124] In some cases, the score or scores indicating co-regulation may be based on detection of DNA motifs from intergenic sequences of one or more genes associated with BGC and putative resistance genes, for example, as described elsewhere herein.
[0125] In some cases, the score or scores indicative of co-expression may be based on differential expression analysis and / or clustering analysis of global transcriptomic data, e.g., as described elsewhere herein.
[0126] In some cases, the one or more genes associated with a biosynthetic gene cluster (BGC) can include, for example, anchor genes, core synthase genes, biosynthetic genes, genes not involved in the biosynthesis of a secondary metabolite produced by the BGC, or any combination thereof.
[0127] In step 114 of Figure 1, the likelihood that the putative resistance gene (e.g., pETaG or pNETaG) is an actual resistance gene (e.g., ETaG or NETaG) is determined based on at least one genomic parameter determined in step 112. In some cases, the likelihood that the putative resistance gene is an actual resistance gene may be output and / or reported as a probability, for example, a 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99% probability that the putative resistance gene is an actual resistance gene. In some cases, the likelihood that the putative resistance gene is an actual resistance gene may be output and / or reported as a probability having any value within this range.
[0128] In some cases, determining the likelihood that the putative resistance gene (e.g., pETaG or pNETaG) is an actual resistance gene (e.g., ETaG or NETaG) may include outputting or reporting a binary classification of the likelihood (e.g., a yes / no answer) based on comparing the at least one determined genomic parameter to at least one predefined threshold. In some cases, for example, the at least one predefined threshold may include a predefined threshold for a co-occurrence score, a co-evolutionary correlation score, a co-regulation score, and / or a co-expression score.
[0129] In some cases, for example, the predefined threshold of the co-occurrence score may include including the top 20, top 15, top 10, or top 5 co-occurring BGCs ranked by normalized distance. The co-occurrence rank may be used to confirm the association between the BGCs and their putative resistance genes (e.g., pETaG or pNETaG). The normalized distance may be calculated from the occurrence of the BGC genes and resistance genes (e.g., ETaG or NETaG) across the positive and negative genomes. The BGC genes may be ranked by their normalized distance (calculated from the positive and negative genome counts).
[0130] In some cases, the predetermined threshold for the coevolutionary correlation score may include a coevolutionary correlation coefficient of 0.6, 0.7, 0.8, 0.9, 0.95 or more. In some cases, the predetermined threshold for the coevolutionary correlation score may have any value within this range.
[0131] In some cases, the predetermined threshold for the coevolution rank score (or coevolution rank) may include a rank of less than 5, less than 10, less than 20, less than 40, less than 60, less than 80, less than 100, less than 200, less than 400, less than 600, less than 800, less than 1000, less than 2000, less than 4000, less than 6000, less than 8000, or less than 10,000. In some cases, the predetermined threshold for the coevolution rank score (or coevolution rank) may include a rank of any value within this range of values.
[0132] In some cases, the predetermined threshold for the coevolutionary gradient may include a coevolutionary gradient value of about 0.75 to about 1.25. In some cases, the predetermined threshold for the coevolutionary slope score may have a value of at least 0.75, at least 0.8, at least 0.85, at least 0.9, at least 0.95, at least 1.0, at least 1.05, at least 1.1, at least 1.15, at least 1.20, or at least 1.25. In some cases, the predetermined threshold for the coevolutionary slope score may have a value of at most 1.25, at most 1.20, at most 1.15, at most 1.10, at most 1.10, at most 1.05, at most 1.0, at most 0.95, at most 0.90, at most 0.85, at most 0.80, or at most 0.75. Any of the lower and upper limits described in this paragraph may be combined to form a range included within the present disclosure, for example, in some cases, the predetermined threshold coevolutionary slope score may range from about 0.80 to about 1.1. In some cases, the predetermined threshold coevolutionary slope score may have any value within this range, for example, about 1.07.
[0133] In some cases, the pre-defined threshold for the co-regulation score may include detecting a DNA motif in an upstream intergenic sequence of a BGC member and one or more members of the putative resistance gene with a p-value of less than or equal to 0.1, 0.09, 0.08, 0.07, 0.06, or 0.05.
[0134] In some cases, the predetermined threshold for the co-expression score may be based on a value determined for a differential expression analysis metric, such as Spearman correlation coefficient, Kolmogorov-Smirnov distance, Euclidean distance, Kullback-Leibler information, or neighbor difference (see, e.g., Gonzalez-Valbuena, et al. (2017), "Metrics to Estimate Differential Co-Expression Networks", BioData Mining 10:32). In these examples, the predetermined threshold for co-expression may include a co-expression score of 0.6, 0.7, 0.8, 0.9, 0.95 or greater. In some cases, the predetermined threshold for the co-expression score may have any value within this range.
[0135] In some cases, for example, when the co-expression score is based on clustering analysis of global transcriptomic data, a pre-defined threshold for the co-expression score may not be used.
[0136] Method for predicting the function of secondary metabolites and / or for identifying secondary metabolites having a desired activity - Patents.com Also disclosed herein are computer-implemented methods for predicting the function of secondary metabolites and / or for identifying biosynthetic gene clusters (BGCs) that encode biosynthetic enzymes for producing secondary metabolites with a desired activity.
[0137] For example, in some cases, a computer-implemented method for predicting a function of a secondary metabolite includes receiving a selection of at least one target sequence of interest, where the at least one target sequence of interest corresponds to a gene sequence associated with a biosynthetic gene cluster (BGC) known to produce a secondary metabolite; receiving a selection of target genomes from a genomics database, where the selection of target genomes includes a plurality of target genomes from organisms known to produce a secondary metabolite; performing a search to identify homologs of the at least one target sequence in the plurality of target genomes; creating a phylogenetic tree based on the identified homologs of the at least one target sequence; and classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on the phylogenetic tree, where a positive genome is a genome that belongs to a clade in which multiple copies of at least one target sequence homolog are present and a negative genome is a genome that belongs to a clade in which a single copy of at least one target sequence homolog is present. and determining, based at least in part on the classification of the positive and negative genomes, at least one genomic parameter selected from: i) one or more scores indicative of co-occurrence of at least one target sequence homolog (putative resistance gene) and one or more genes associated with the BGC; ii) one or more scores indicative of co-evolution of at least one target sequence homolog (putative resistance gene) and one or more genes associated with the BGC; iii) one or more scores indicative of co-regulation of at least one target sequence homolog (putative resistance gene) and one or more genes associated with the BGC; and iv) one or more scores indicative of co-expression of at least one target sequence homolog (putative resistance gene) and one or more genes associated with the BGC; and determining a likelihood that the putative resistance gene is a resistance gene encoding a protein target acted upon by a secondary metabolite based on the at least one genomic parameter.
[0138] In some cases, the selection of at least one target sequence of interest may be provided as input by a user of a system configured to perform the computer-implemented method.
[0139] In some cases, as described elsewhere herein, the at least one target sequence of interest may comprise an amino acid sequence, a nucleotide sequence, or any combination thereof. In some cases, the at least one target sequence of interest comprises a peptide sequence or a portion thereof, a protein sequence or a portion thereof, a protein domain sequence or a portion thereof, a gene sequence or a portion thereof, or any combination thereof.
[0140] In some cases, the selection of the target genome may be provided as input by a user of a system configured to perform the computer-implemented method.
[0141] In some cases, the multiple target genomes may include plant genomes, fungal genomes, bacterial genomes, or any combination thereof. In some cases, the genomics database includes a public genomics database or a proprietary genomics database.
[0142] In some cases, the search for identifying the homologue of at least one target sequence may include identifying the homologue based on a probabilistic sequence alignment model.In some cases, the probabilistic sequence alignment model is a profile hidden Markov model (pHMM).In some cases, the homologue is identified based on the comparison of the probabilistic sequence alignment model score with a predefined threshold value, as described elsewhere herein.In some cases, for example, such predefined threshold value may be determined based on the lowest bit score of known homologues.
[0143] In some cases, the search to identify homologs of at least one target sequence may include identifying homologs based on an alignment of sequences using a local sequence alignment search tool, calculating sequence homology metrics based on the alignment, and comparing the calculated sequence homology metrics to a predefined threshold. In some cases, the predefined threshold includes a threshold for percent sequence identity, percent sequence coverage, E-value, or bit score value, as described elsewhere herein.
[0144] In some cases, the predetermined threshold for percent sequence identity may be at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more.
[0145] In some cases, the predetermined threshold for percent sequence coverage may be at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more.
[0146] In some cases, the predefined thresholds for E-values are up to 10, 1, 0.1, 0.001, 0.0001, 1e -10 , 1e -20 , 1e -100 , or may be lower.
[0147] In some cases, the predetermined threshold value for the bit score may be at least 5, 10, 25, 50, 100, 250, 500, 1000, 5000, or more.
[0148] In some cases, the search to identify homologs of the at least one target sequence comprises identifying homologs based on the use of gene and / or protein domain annotation tools.
[0149] In some cases, creating a phylogenetic tree based on the identified homologues of at least one target sequence may include aligning the homologue sequences using an alignment software tool, trimming the aligned homologue sequences using a sequence trimming software tool, and constructing a phylogenetic tree using a phylogenetic tree building software tool, as described elsewhere herein.
[0150] In some cases, the at least one target sequence of interest can include, for example, a known ETaG sequence, a known NETaG sequence, or a known core synthase gene sequence.
[0151] In some cases, determining the likelihood that the putative resistance gene is a resistance gene encoding a protein target acted upon by the secondary metabolite may include comparing at least one determined genomic parameter to at least one predefined threshold. In some cases, for example, the at least one predefined threshold may include a predefined threshold for a co-occurrence score, a co-evolution score, a co-regulation score, and / or a co-expression score. Examples of such predefined thresholds are described elsewhere herein.
[0152] As another non-limiting example, a computer-implemented method for identifying biosynthetic gene clusters (BGCs) encoding biosynthetic enzymes for producing a secondary metabolite having a desired activity is also disclosed, the method including: receiving a selection of at least one target sequence of interest, where the at least one target sequence of interest comprises a sequence encoding a therapeutic target of interest; receiving a selection of target genomes from a genomics database, where the selection of target genomes comprises a plurality of target genomes from organisms known to produce secondary metabolites; performing a search to identify homologs of the at least one target sequence in the plurality of target genomes; creating a phylogenetic tree based on the identified homologs of the at least one target sequence; and classifying the genomes of the plurality of target genomes as positive genomes or negative genomes based on the phylogenetic tree. classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on a phylogenetic tree, where a positive genome is a genome belonging to a clade in which multiple copies of at least one target sequence homolog are present and a negative genome is a genome belonging to a clade in which a single copy of at least one target sequence homolog is present, and the target sequence homolog present in multiple copies in the positive genome is a putative resistance gene; and based at least in part on the classification of the positive genomes and the negative genomes, determining: i) one or more scores indicative of co-occurrence of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a biosynthetic gene cluster (BGC); ii) one or more scores indicative of co-evolution of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a BGC; iii) one or more scores indicative of co-regulation of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a BGC;and iv) determining at least one genomic parameter selected from one or more scores indicative of co-expression of at least one target sequence homolog (putative resistance gene) with one or more genes associated with the BGC, and determining a likelihood based on the at least one genomic parameter that the putative resistance gene is an actual resistance gene associated with the BGC that produces a secondary metabolite that acts on the protein product encoded by the resistance gene;
[0153] In some cases, the selection of the at least one target sequence of interest may be provided as an input by a user of a system configured to perform a computer-implemented method. In some cases, as described elsewhere herein, the at least one target sequence of interest comprises an amino acid sequence, a nucleotide sequence, or any combination thereof. In some cases, the at least one target sequence of interest comprises a peptide sequence or a portion thereof, a protein sequence or a portion thereof, a protein domain sequence or a portion thereof, a gene sequence or a portion thereof, or any combination thereof.
[0154] In some cases, the selection of the target genome may be provided as an input by a user of a system configured to perform the computer-implemented method. In some cases, as described elsewhere herein, the multiple target genomes may include plant genomes, fungal genomes, bacterial genomes, or any combination thereof. In some cases, the genomics database may include a public genomics database or a proprietary genomics database.
[0155] In some cases, the search for identifying the homologue of at least one target sequence can include identifying the homologue based on a probabilistic sequence alignment model.In some cases, the probabilistic sequence alignment model is a profile hidden Markov model (pHMM).In some cases, the homologue is identified based on the comparison of the probabilistic sequence alignment model score with a predetermined threshold.In some cases, such a predetermined threshold can be determined based on, for example, the lowest bit score of known homologues.
[0156] In some cases, the search to identify homologs of at least one target sequence may include identifying homologs based on an alignment of sequences using a local sequence alignment search tool, calculating sequence homology metrics based on the alignment, and comparing the calculated sequence homology metrics to a predetermined threshold. In some cases, the predetermined threshold includes a threshold for percent sequence identity, percent sequence coverage, E value, or bit score value.
[0157] In some cases, the predetermined threshold for percent sequence identity may be at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more.
[0158] In some cases, the predetermined threshold for percent sequence coverage may be at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more.
[0159] In some cases, the predefined thresholds for E-values are up to 10, 1, 0.1, 0.001, 0.0001, 1e -10 , 1e -20 , 1e -100 , or may be lower.
[0160] In some cases, the predetermined threshold value for the bit score may be at least 5, 10, 25, 50, 100, 250, 500, 1000, 5000, or more.
[0161] In some cases, the search to identify homologs of at least one target sequence may involve identifying homologs based on the use of gene and / or protein domain annotation tools.
[0162] In some cases, creating a phylogenetic tree based on the identified homologues of at least one target sequence may include aligning the homologue sequences using an alignment software tool, trimming the aligned homologue sequences using a sequence trimming software tool, and constructing a phylogenetic tree using a phylogenetic tree building software tool, as described elsewhere herein.
[0163] In some cases, determining the likelihood that the putative resistance gene is an actual resistance gene associated with the secondary metabolite producing BGC may include comparing at least one determined genomic parameter to at least one predefined threshold. In some cases, for example, the at least one predefined threshold may include a predefined threshold for a co-occurrence score, a co-evolution score, a co-regulation score, and / or a co-expression score. Examples of such predefined thresholds are described elsewhere herein.
[0164] In some cases, as described elsewhere herein, the computer-implemented methods described herein further include performing a computer-implemented in vitro assay to test secondary metabolites produced by the identified BGC for activity against a therapeutic target of interest.
[0165] In some cases, the computer-implemented methods described herein further include performing an in vivo assay to test secondary metabolites produced by the identified BGC for activity against a therapeutic target of interest, as described elsewhere herein.
[0166] Purpose The computer-based methods described herein have a variety of uses, including, for example, identifying homologs or orthologs of one or more target sequences (e.g., gene sequences) of interest in one or more target genomes, identifying resistance genes to secondary metabolites produced by a BGC in a target genome, predicting the function of a secondary metabolite produced by a BGC, and / or identifying a BGC that encodes a biosynthetic enzyme for producing a secondary metabolite having an activity of interest (e.g., a therapeutic activity of interest).
[0167] In some cases, the disclosure includes a method for identifying a genome of the plurality of target genomes as positive or negative genomes based on the phylogenetic tree, wherein a positive genome is a genome that belongs to a clade in which multiple copies of at least one target sequence homolog are present, and a negative genome is a genome that belongs to a clade in which a single copy of at least one target sequence homolog is present, and wherein the target homolog present in multiple copies in the positive genome is a putative ETaG or NETaG. and determining a likelihood that the putative ETaG or NETaG is an integration target gene (ETaG) or a non-integration target gene (NETaG) based on the at least one genomic parameter selected from: i) one or more scores indicative of co-occurrence of at least one target sequence homolog (putative ETaG or NETaG) and one or more genes associated with a biosynthetic gene cluster (BGC); ii) one or more scores indicative of co-evolution of at least one target sequence homolog (putative ETaG or NETaG) and one or more genes associated with a BGC; iii) one or more scores indicative of co-regulation of at least one target sequence homolog (putative ETaG or NETaG) and one or more genes associated with a BGC; and iv) one or more scores indicative of co-expression of at least one target sequence homolog (putative ETaG or NETaG) and one or more genes associated with a BGC; and determining a likelihood that the putative ETaG or NETaG is an integration target gene (ETaG) or a non-integration target gene (NETaG) based on the at least one genomic parameter.
[0168] In some cases, the disclosure includes a method for identifying a genome of the plurality of target genomes from a genomics database, the method comprising: receiving a selection of at least one target sequence of interest, where the at least one target sequence of interest corresponds to a gene sequence associated with a biosynthetic gene cluster (BGC) known to produce a secondary metabolite; receiving a selection of a target genome from a genomics database, where the selection of target genomes includes a plurality of target genomes from organisms known to produce a secondary metabolite; performing a search to identify homologs of the at least one target sequence in the plurality of target genomes; creating a phylogenetic tree based on the identified homologs of the at least one target sequence; and classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on the phylogenetic tree, where a positive genome is a genome that is at least one of the plurality of target genomes ... classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on a phylogenetic tree, wherein the genomes belong to a clade in which at least one target sequence homolog is present in multiple copies, and the negative genomes belong to a clade in which a single copy of at least one target sequence homolog is present, and the target sequence homolog present in multiple copies in the positive genome is a putative resistance gene; and based at least in part on the classification of the positive genomes and the negative genomes, determining: i) one or more scores indicative of co-occurrence of at least one target sequence homolog (putative resistance gene) and one or more genes associated with the BGC; ii) one or more scores indicative of co-evolution of at least one target sequence homolog (putative resistance gene) and one or more genes associated with the BGC; iii) one or more scores indicative of co-regulation of at least one target sequence homolog (putative resistance gene) and one or more genes associated with the BGC;and iv) determining at least one genomic parameter selected from at least one target sequence homolog (putative resistance gene) and one or more scores indicative of co-expression with one or more genes associated with the BGC, and determining a likelihood that the putative resistance gene is a resistance gene encoding a protein target acted upon by the secondary metabolite based on the at least one genomic parameter;
[0169] In some cases, the disclosure provides a method for identifying a genome of the plurality of target genomes, comprising: receiving a selection of at least one target sequence of interest, wherein the at least one target sequence of interest comprises a sequence encoding a therapeutic target of interest; receiving a selection of target genomes from a genomics database, wherein the selection of target genomes comprises a plurality of target genomes from organisms known to produce secondary metabolites; performing a search to identify homologs of the at least one target sequence in the plurality of target genomes; creating a phylogenetic tree based on the identified homologs of the at least one target sequence; and classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on the phylogenetic tree, wherein a positive genome is a genome having multiple copies of the at least one target sequence homolog. classifying genomes of the plurality of target genomes as positive or negative genomes based on a phylogenetic tree, where the genomes belong to a clade in which a single copy of at least one target sequence homolog is present, and the target sequence homolog present in multiple copies in the positive genome is a putative resistance gene; and based at least in part on the classification of the positive and negative genomes, determining: i) one or more scores indicative of co-occurrence of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a biosynthetic gene cluster (BGC); ii) one or more scores indicative of co-evolution of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a BGC; iii) one or more scores indicative of co-regulation of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a BGC;and iv) determining at least one genomic parameter selected from one or more scores indicative of co-expression of at least one target sequence homolog (putative resistance gene) with one or more genes associated with the BGC, and determining a likelihood that the putative resistance gene is an actual resistance gene associated with the BGC that produces a secondary metabolite that acts on the protein product encoded by the resistance gene, based on the at least one genomic parameter;
[0170] In some cases, the disclosed methods (e.g., computer-implemented methods) may further include performing an in vitro assay, e.g., an assay to detect or measure the activity (e.g., receptor binding activity, enzyme activating activity, enzyme inhibitory activity, etc.) of the secondary metabolite (or analog thereof) against a mammalian (e.g., human) protein encoded by a mammalian (e.g., human) gene that is homologous to an ETaG or NETag identified in an organism that comprises a biosynthetic gene cluster (BGC) that produces the secondary metabolite. In some cases, the disclosed methods may further include performing an in vitro assay to detect or measure the activity (e.g., receptor binding activity, enzyme activating activity, enzyme inhibitory activity, etc.) of the secondary metabolite (or analog thereof) against a protein (e.g., reptile, avian, amphibian, plant, fungal, bacterial, or viral protein) encoded by a reptile, avian, amphibian, plant, fungal, bacterial, or viral gene that is homologous to an ETaG or NETag identified in an organism that comprises a biosynthetic gene cluster (BGC) that produces the secondary metabolite.
[0171] In some cases, the methods (e.g., computer-implemented methods) of the present disclosure may further include performing an in vivo assay, e.g., an assay to detect or measure activity (e.g., receptor binding activity, enzyme activating activity, enzyme inhibitory activity, intracellular signaling pathway activity, disease response, etc.) of the secondary metabolite (or an analog thereof) against a mammalian (e.g., human) protein encoded by a mammalian (e.g., human) gene that is homologous to an ETaG or NETag identified in an organism that contains a biosynthetic gene cluster (BGC) that produces the secondary metabolite. In some cases, the method may further include performing an in vivo assay to detect or measure activity (e.g., receptor binding activity, enzyme activating activity, enzyme inhibitory activity, intracellular signaling pathway activity, disease response, etc.) of the secondary metabolite (or analog thereof) against a protein (e.g., a reptilian, avian, amphibian, plant, fungal, bacterial or viral protein) encoded by a reptilian, avian, amphibian, plant, fungal, bacterial or viral gene that is homologous to an ETaG or NETag identified in an organism that contains a biosynthetic gene cluster (BGC) that produces the secondary metabolite.
[0172] In some cases, the methods of the disclosure may be used, for example, to identify and / or characterize mammalian (e.g., human) targets of a secondary metabolite (or analog thereof) produced by a BGC. In some cases, the methods of the disclosure may be used to identify and / or characterize reptile, avian, amphibian, plant, fungal, bacterial, viral targets of a secondary metabolite (or analog thereof) produced by a BGC, or targets from any other organism.
[0173] In some cases, the methods of the present disclosure may be used in drug discovery efforts, for example, to identify small molecule modulators of mammalian (e.g., human) target genes. In some cases, the methods of the present disclosure may be used to identify small molecule modulators of reptilian target genes, avian target genes, amphibian target genes, plant target genes, fungal target genes, bacterial target genes, viral target genes, or target genes from any other organism.
[0174] In some cases, the secondary metabolite is a product of an enzyme encoded by a BGC or a salt thereof, including non-naturally occurring salts. In some cases, the secondary metabolite or an analog thereof is an analog of the product of an enzyme encoded by a BGC, such as a small molecule compound having the same core structure as the secondary metabolite or a salt thereof.
[0175] In some cases, the disclosure provides a method for regulating a human target (or a target from another organism), comprising providing a secondary metabolite or analog thereof produced by an enzyme encoded by a BGC, wherein the human target (or a nucleic acid sequence encoding the human target) is homologous to an ETaG or NETaG associated with the BGC as determined using any one of the methods described herein.
[0176] In some cases, the disclosure provides a method for treating a condition, injury, or disease associated with a human target (or a target from another organism), comprising providing a secondary metabolite or an analog thereof produced by an enzyme encoded by a BGC, wherein the human target (or a nucleic acid sequence encoding the human target) is homologous to an ETaG or NETaG associated with the BGC as determined using any one of the methods described herein.
[0177] In some cases, the secondary metabolite is produced by a fungus. In some cases, the secondary metabolite is acyclic. In some cases, the secondary metabolite is a polyketide. In some cases, the secondary metabolite is a terpene compound. In some cases, the secondary metabolite is a non-ribosomally synthesized peptide.
[0178] In some cases, an analog of a substance (e.g., a secondary metabolite) shares one or more specific structural features, elements, components, or moieties with a reference substance. Typically, an analog exhibits significant structural similarity with a reference substance, for example, sharing a core or consensus structure, but also differs in certain discrete ways. In some cases, an analog is a substance that can be produced from a reference substance, for example, by chemical manipulation of the reference substance. In some cases, an analog is a substance that can be produced by carrying out a synthetic process that is substantially similar (e.g., sharing multiple steps) to that which produces the reference substance. In some cases, an analog is produced or can be produced by carrying out a synthetic process that is different from that used to produce the reference substance. In some cases, an analog of a substance is a substance that is substituted at one or more of its substitutable positions.
[0179] In some cases, the analog of the product comprises the structural core of the product. In some cases, the biosynthetic product is cyclic, e.g., monocyclic, bicyclic, or polycyclic, and the structural core of the product is or comprises a monocyclic, bicyclic, or polycyclic ring system. In some cases, the structural core of the product comprises one ring of the bicyclic or polycyclic ring system of the product. In some cases, the product is or comprises a polypeptide, and the structural core is the backbone of the polypeptide. In some cases, the product is or comprises a polyketide, and the structural core is the backbone of the polyketide. In some cases, the analog is a substituted biosynthetic product that comprises one or more suitable substitutents.
[0180] system Also disclosed herein is a system designed to implement any of the disclosed methods for identifying resistance genes (e.g., ETaG or NETaG). The system may, for example, be communicatively coupled to one or more processors, and when executed by the one or more processors, the system may include: receiving a selection of at least one target sequence of interest; receiving a selection of target genomes from a genomics database, the selection of target genomes from the genomics database including a plurality of target genomes from organisms known to produce or likely to produce secondary metabolites; performing a search to identify homologs of at least one target sequence in the plurality of target genomes; creating a phylogenetic tree based on the identified homologs of at least one target sequence; classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on the phylogenetic tree, where a positive genome is a genome belonging to a clade in which multiple copies of at least one target sequence homolog are present and a negative genome is a genome belonging to a clade in which a single copy of at least one target sequence homolog is present. and a target sequence homolog present in multiple copies in the positive genome is a putative resistance gene (e.g., ETaG or NETaG); determining at least one genomic parameter selected from the following, based at least in part on the classification of the positive genomes and the negative genomes: i) one or more scores indicative of co-occurrence of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a biosynthetic gene cluster (BGC); ii) one or more scores indicative of co-evolution of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a BGC; iii) one or more scores indicative of co-regulation of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a BGC; and iv) one or more scores indicative of co-expression of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a BGC;and a memory unit configured to store instructions to determine the likelihood that the putative resistance gene (e.g., pETaG or pNETaG) is an actual resistance gene (e.g., an integration target gene (ETaG) or a non-integration target gene (NETaG)) based on the at least one genomic parameter. In some cases, determining the likelihood that the putative resistance gene (e.g., pETaG or pNETaG) is a resistance gene (e.g., ETaG or NETaG) comprises comparing the at least one determined genomic parameter to at least one pre-defined threshold. Examples of such pre-defined thresholds are described elsewhere herein;
[0181] Computing Devices and Systems 2 is a diagram illustrating an example of a computing device according to one or more examples of the present disclosure. The device 200 may be a host computer connected to a network. The device 200 may be a client computer or a server. As shown in FIG. 2, the device 200 may be any suitable type of microprocessor-based device, such as a personal computer, a workstation, a server, or a handheld computing device (portable electronic device) such as a phone or tablet. The device may include, for example, one or more of a processor 210, an input device 220, an output device 230, a storage 240, and a communication device 260. The input device 220 and the output device 230 may generally correspond to those described above, and they may be connectable or integrated with the computer.
[0182] The input device 220 may be any suitable device that provides input, such as a touch screen, a keyboard or keypad, a mouse, or a voice recognition device. The output device 230 may be any suitable device that provides output, such as a touch screen, a tactile device, or a speaker.
[0183] Storage 240 may be any suitable device providing storage, such as RAM, cache, electrical, magnetic, or optical memory, including a hard drive, or removable storage disk. Communications device 260 may include any suitable device capable of sending and receiving signals over a network, such as a network interface chip or device. The components of the computer may be connected in any suitable manner, such as via a physical bus 270 or wirelessly.
[0184] Software 250 that may be stored in memory / storage 240 and executed by processor 210 may include, for example, programming that embodies the functions of the present disclosure (e.g., as implemented in the devices described above).
[0185] The software 250 may also be stored and / or transported in any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch and execute instructions associated with the software from the instruction execution system, apparatus, or device. In the context of this disclosure, a computer-readable storage medium may be any medium, such as storage 240, that can contain or store programming for use by or in connection with an instruction execution system, apparatus, or device.
[0186] The software 250 may also be propagated in any transport medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can retrieve and execute instructions associated with the software from the instruction execution system, apparatus, or device. In the context of this disclosure, a transport medium may be any medium that can communicate, propagate, or transport programming for use by or in connection with an instruction execution system, apparatus, or device. Transport readable media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, or infrared wired or wireless propagation media.
[0187] The device 200 may be connected to a network, which may be any suitable type of interconnected communication system. The network may implement any suitable communication protocol and may be protected by any suitable security protocol. The network may include any suitable configuration of network links capable of effecting transmission and reception of network signals, such as wireless network connections, T1 or T3 lines, cable networks, DSL, or telephone lines.
[0188] Device 200 may implement any operating system suitable for operating on a network. Software 250 may be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, application software embodying functionality of the present disclosure may be deployed in different configurations, such as, for example, as a web-based application or web service, in a client / server configuration, or via a web browser. EXAMPLES
[0189] Example 1 - Use of NETaG to identify BGCs with specific targets that may have therapeutic applications Succinate dehydrogenase complex subunit C (SDHC) inhibitors: A collection of protein sequences from a diverse set of fungal genomes from different taxa were annotated using InterProScan and searched for proteins annotated with Interpro ID IPR000701 (succinate dehydrogenase / fumarate reductase type B, transmembrane subunit) to identify succinate dehydrogenase complex subunit C (SDHC) homologs in the set of genomes. Genomes with a single copy of Interpro ID IPR000701 homologs were referred to as negative genomes and genomes with multiple copies of ID IPR000701 homologs were referred to as positive genomes. NETaG is a copy of succinate dehydrogenase complex subunit C (SDHC) that confers resistance to the product of the gene cluster. All SDHC protein sequences were aligned using MAFFT and trimmed using trimAI to remove gaps. The resulting trimmed multiple sequence alignment was processed with IQ-TREE to generate a maximum likelihood phylogeny of SDHC homologs. NETaGs can be identified by their position in the phylogenetic tree. NETaGs from several fungal genera cluster together in one branch or several close branches of the phylogenetic tree, whereas housekeeping copies show greater phylogenetic distance and represent proteins only from a single fungal genus in those branches. Furthermore, NETaG clades contain only proteins from multicopy genomes, whereas housekeeping copies represent proteins together in clades from single and multicopy genomes.
[0190] Figure 3 provides a non-limiting example of a maximum likelihood phylogenetic tree of SDHC homologs from a diverse set of fungal species. NETaG can be identified by the colocalization of homologs from different fungal species, as well as the colocalization of single-copy and multi-copy homologs in other branches of the tree.
[0191] From the phylogenetic tree, we can infer 14 positive genomes (genomes containing NETaG) and 39 negative genomes (single copy; no NETaG). All genomes are annotated using antiSMASH and the resulting gene clusters are classified into gene cluster families using the clusteromics approach described above. The resulting families are classified according to the normalized distance:
number
[0192] Using the number of positive and negative genomes determined by the phylogenetic tree in Figure 3 to perform clusteromics, we hypothesize: the positive genome (containing NETaG) contains a BGC that produces a secondary metabolite with activity against the target gene homolog (or its product), while the negative genome (containing only a housekeeping copy of the target gene homolog) does not contain such a BGC. Clusteromics analysis takes all BGCs from all selected organisms, returns gene cluster families, and then tests for their presence in the positive and negative genomes. We use the normalized distance (see above) as a metric to determine the best scoring gene cluster family. The cluster number indicates the number of clusters in the gene cluster family that are potential targets for SDHC inhibitors in this case.
[0193] Using normalized distance as a metric, gene cluster family 87 is the best scoring candidate for SDHC inhibitors among all gene cluster families. Gene cluster families mainly contain two gene clusters per family, which greatly reduces the number of BGCs that need to be investigated to find BGCs that produce SMs with activity against house copy. For example, a species of the genus Rhizodermea contains 90 BGCs. Using the NETaG method combined with cluster analysis, the number of candidates from the BGC candidates of the top-scoring gene cluster families can be reduced from 90 gene clusters to only two gene clusters. Therefore, only two BGCs need to be investigated for activity against SDHC, demonstrating the strength of the predictive power of the present invention.
[0194] Furthermore, the inventors have determined that gene cluster family 87 (see Table 1) contains a gene cluster similar to the atopenine and harzianopyridone gene clusters (shown in FIG. 4), which are known to be potent inhibitors of SDHC. This provides solid evidence that the disclosed method can successfully predict inhibitors of target genes using NETaG. The use of NETaG is not limited to identifying SDHC inhibitors. The disclosed method can be used to identify BGCs that produce secondary metabolites with functions for any NETaG, and thus can be used to find new bioactive compounds of interest.
[0195] Figure 4 provides a non-limiting example of a BGC comparison of the atpenin BGC (gene clusters extracted using genomic coordinates from Bat-Erdene, et al. (2020), “Iterative Catalysis in the Biosynthesis of Mitochondrial Complex II Inhibitors Harzianopyridone and Atpenin B”, J. Am. Chem. Soc. 142(19):8550-8554) with gene clusters from gene cluster family 87. Each row contains, for each genus, a candidate BGC from the top-scoring gene cluster family (see Table 1). The arrows indicate the genes of the BGC, and the shaded regions between them show the sequence alignment between them generated by the clinker tool (Gilchrist, et al. “Clinker&Clustermap.js: Automatic Generation of Gene Cluster Comparison Figures”, Bioinformatics 37(16):2473-2475). The plot shows a high conservation of the biosynthetic genes across species, supporting our prediction that BGC produces the SDHC inhibitor atpenine. [Table 1]
[0196] Exemplary embodiments Among the embodiments described herein are the following: 1. A computer-implemented method for identifying resistance genes, comprising: receiving a selection of at least one target sequence of interest; receiving a selection of target genomes from a genomics database, the selection of target genomes including a plurality of target genomes from organisms known to produce or likely to produce secondary metabolites; conducting a search to identify homologs of at least one target sequence in a plurality of target genomes; generating a phylogenetic tree based on the identified homologs of at least one target sequence; Classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on a phylogenetic tree, where a positive genome is a genome belonging to a clade in which multiple copies of at least one target sequence homolog are present, and a negative genome is a genome belonging to a clade in which a single copy of at least one target sequence homolog is present, and the target sequence homolog present in multiple copies in the positive genome is a putative resistance gene; and Based at least in part on the classification of positive and negative genomes, i) one or more scores indicating the co-occurrence of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a biosynthetic gene cluster (BGC); ii) one or more scores indicating coevolution of at least one target sequence homolog (putative resistance gene) and one or more genes associated with the BGC; iii) one or more scores indicating co-regulation of at least one target sequence homolog (putative resistance gene) with one or more genes associated with BGC; and iv) one or more scores indicating co-expression of at least one target sequence homolog (putative resistance gene) with one or more genes associated with BGC; determining at least one genomic parameter selected from determining a likelihood that the putative resistance gene is a resistance gene based on the at least one genomic parameter; A method comprising: 2. The computer-implemented method of embodiment 1, wherein determining the likelihood that the putative resistance gene is a resistance gene comprises comparing at least one determined genomic parameter to at least one predetermined threshold. 3. The computer-implemented method of embodiment 1 or embodiment 2, wherein the selection of at least one target sequence of interest is provided as input by a user of a system configured to perform the computer-implemented method. 4. The computer-implemented method of any one of embodiments 1-3, wherein at least one target sequence of interest comprises an amino acid sequence, a nucleotide sequence, or any combination thereof. 5. The computer-implemented method of any one of embodiments 1 to 4, wherein at least one target sequence of interest comprises a peptide sequence or a portion thereof, a protein sequence or a portion thereof, a protein domain sequence or a portion thereof, a gene sequence or a portion thereof, or any combination thereof. 6. The computer-implemented method of any one of embodiments 1-5, wherein at least one target sequence of interest comprises a mammalian sequence, a human sequence, a plant sequence, a fungal sequence, a bacterial sequence, an archaeal sequence, a viral sequence, or any combination thereof. 7. The computer-implemented method of any one of embodiments 1 to 6, wherein the at least one target sequence of interest comprises a primary target sequence and one or more related sequences. 8. The computer-implemented method of embodiment 7, wherein the one or more related sequences comprise sequences functionally related to the primary target sequence. 9. The computer-implemented method of embodiment 8, wherein the one or more associated sequences include sequences that are pathway-associated to the primary target sequence. 10. A computer-implemented method according to any one of embodiments 1 to 9, wherein the selection of the target genome is provided as an input by a user of a system configured to perform the computer-implemented method. 11. The computer-implemented method of any one of embodiments 1-10, wherein the multiple target genomes include plant genomes, fungal genomes, bacterial genomes, or any combination thereof. 12. The computer-implemented method of any one of embodiments 1 to 11, wherein the genomics database comprises a public genomics database. 13. The computer-implemented method of any one of embodiments 1 to 12, wherein the genomics database comprises a proprietary genomics database. 14. The computer-implemented method of any one of embodiments 1 to 13, wherein the search for identifying homologs of at least one target sequence comprises identifying homologs based on a probabilistic sequence alignment model. 15. The computer-implemented method of embodiment 14, wherein the probabilistic sequence alignment model is a profile hidden Markov model (pHMM). 16. The computer-implemented method of embodiment 14 or embodiment 16, wherein homologs are identified based on a comparison of the probabilistic sequence alignment model score to a predefined threshold. 17. The computer-implemented method of any one of embodiments 1 to 16, wherein the search to identify homologs of at least one target sequence comprises identifying homologs based on an alignment of the sequences using a local sequence alignment search tool, calculating a sequence homology metric based on the alignment, and comparing the calculated sequence homology metric with a predefined threshold. 18. The computer-implemented method of embodiment 17, wherein the local sequence alignment search tool comprises BLAST, DIAMOND, HMMER, Exonerate, or ggsearch. 19. The computer-implemented method of embodiment 17 or embodiment 18, wherein the predetermined threshold comprises a threshold for percent sequence identity, percent sequence coverage, E value, or bit score value. 20. A computer-implemented method according to any one of embodiments 1 to 19, wherein the search for identifying homologues of at least one target sequence comprises identifying homologues based on the use of gene and / or protein domain annotation tools. 21. The computer-implemented method of embodiment 20, wherein the gene and / or protein domain annotation tool comprises InterProScan or EggNOG. 22. The computer-implemented method of any one of embodiments 1 to 21, wherein creating a phylogenetic tree based on the identified homologs of at least one target sequence comprises aligning the homolog sequences using an alignment software tool, trimming the aligned homolog sequences using a sequence trimming software tool, and constructing a phylogenetic tree using a phylogenetic tree building software tool. 23. The computer-implemented method of embodiment 22, wherein the alignment software tool comprises MAFFT, MUSCLE, or ClustalW. 24. The computer-implemented method of embodiment 22 or embodiment 23, wherein the sequence trimming software tool comprises trimAI, GBlocks, or ClipKIT. 25. A computer-implemented method according to any one of embodiments 22 to 24, wherein the phylogenetic tree construction software tool comprises FastTree, IQ-TREE, RAxML, MEGA, MrBayes, BEAST, or PAUP. 26. A computer-implemented method according to any one of embodiments 22 to 25, wherein the construction of the phylogenetic tree is based on a maximum likelihood algorithm, a parsimony algorithm, a neighbor-joining algorithm, a distance matrix algorithm, or a Bayesian estimation algorithm. 27. The computer-implemented method of any one of embodiments 1 to 26, wherein the one or more scores indicative of co-occurrence are determined based on identification of a positive correlation between the presence of multiple copies of the putative resistance gene in the positive genome and the presence of one or more genes of the BGC. 28. The computer-implemented method of embodiment 27, wherein identifying a positive correlation between the presence of multiple copies of a putative resistance gene in a positive genome and the presence of one or more genes of a BGC comprises using a clustering algorithm to cluster aligned protein sequences, aligned nucleotide sequences, aligned protein domain sequences, or aligned pHMMs for a group of BGCs to identify a BGC community within multiple target genomes. 29. The computer-implemented method of embodiment 27, wherein identifying a positive correlation between the presence of multiple copies of a putative resistance gene in a positive genome and the presence of one or more genes of a BGC comprises using phylogenetic analysis of protein sequences or protein domains for a group of BGCs to identify a BGC community within multiple target genomes. 30. The computer-implemented method of embodiment 27, wherein identifying a positive correlation between the presence of multiple copies of a putative resistance gene in a positive genome and the presence of one or more genes of a BGC comprises selecting genomes with a specific taxonomy to identify a BGC community within a plurality of target genomes. 31. The computer-implemented method of any one of embodiments 1 to 30, wherein one or more scores indicative of coevolution of the putative resistance gene and one or more genes associated with the BGC are determined based on a coevolution correlation score, a coevolution rank score, a coevolution slope score, or any combination thereof. 32. The computer-implemented method of embodiment 31, wherein the coevolution correlation score is based on the correlation between the pairwise sequence identity percentage of a cluster of orthologous groups (COGs) for the putative resistance gene and the pairwise sequence identity percentage of a cluster of orthologous groups (COGs) for one of the one or more genes associated with the BGC. 33. The computer-implemented method of embodiment 31, wherein the coevolution rank score is based on a ranking of correlation coefficients of COGs that contain a gene among the one or more genes associated with the BGC in ascending order for COGs that contain the putative resistance gene. 34. The computer-implemented method of embodiment 33, wherein in the event of a tie in distance scores, the rank for all COGs in the tie is set equal to the lowest rank in the group. 35. The computer-implemented method of embodiment 31, wherein the coevolutionary slope score is based on an orthogonal regression of the pairwise sequence identity percentage of the COG for the putative resistance gene and the pairwise sequence identity percentage of the COG for one of the one or more genes associated with the BGC. 36. A computer-implemented method according to any one of embodiments 32 to 35, wherein only COGs arising from unique positive genomes having three or more genes remaining after removing the corresponding genes from the negative genome are used to evaluate the coevolution correlation score, the coevolution rank score, or the coevolution slope score. 37. A computer-implemented method according to any one of embodiments 1 to 36, wherein the one or more scores indicative of co-regulation are based on DNA motif detection from intergenic sequences of one or more genes associated with BGC and putative resistance genes. 38. The computer-implemented method of any one of embodiments 1 to 37, wherein the one or more scores indicative of co-expression are based on differential expression analysis and / or clustering analysis of global transcriptomics data. 39. The computer-implemented method of any one of embodiments 1 to 38, wherein the one or more genes associated with a biosynthetic gene cluster (BGC) include anchor genes, core synthase genes, biosynthetic genes, genes not involved in the biosynthesis of a secondary metabolite produced by the BGC, or any combination thereof. 40. The computer-implemented method of any one of embodiments 1 to 39, wherein the putative resistance gene is a putative integration target gene (pETaG) or a putative non-integration target gene (pNETaG). 41. A computer-implemented method according to any one of embodiments 1 to 40, wherein the resistance gene is an integration targeted gene (ETaG) or a non-integration targeted gene (NETaG). 42. A computer-implemented method for predicting the function of a secondary metabolite, comprising: receiving a selection of at least one target sequence of interest, wherein the at least one target sequence of interest corresponds to a gene sequence associated with a biosynthetic gene cluster (BGC) known to produce a secondary metabolite; receiving a selection of target genomes from a genomics database, the selection of target genomes including a plurality of target genomes from organisms known to produce secondary metabolites; conducting a search to identify homologs of at least one target sequence in a plurality of target genomes; generating a phylogenetic tree based on the identified homologs of at least one target sequence; Classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on a phylogenetic tree, where a positive genome is a genome belonging to a clade in which multiple copies of at least one target sequence homolog are present, and a negative genome is a genome belonging to a clade in which a single copy of at least one target sequence homolog is present, and the target sequence homolog present in multiple copies in the positive genome is a putative resistance gene; and Based at least in part on the classification of positive and negative genomes, i) one or more scores indicating the co-occurrence of at least one target sequence homolog (putative resistance gene) and one or more genes associated with BGC; ii) one or more scores indicating coevolution of at least one target sequence homolog (putative resistance gene) and one or more genes associated with the BGC; iii) one or more scores indicating co-regulation of at least one target sequence homolog (putative resistance gene) with one or more genes associated with BGC; and iv) one or more scores indicating co-expression of at least one target sequence homolog (putative resistance gene) with one or more genes associated with BGC; determining at least one genomic parameter selected from determining a likelihood that the putative resistance gene is a resistance gene encoding a protein target acted upon by a secondary metabolite based on the at least one genomic parameter; A method comprising: 43. The computer-implemented method of embodiment 42, wherein determining the likelihood that the putative resistance gene is a resistance gene encoding a protein target acted upon by a secondary metabolite comprises comparing at least one determined genomic parameter to at least one predetermined threshold. 44. The computer-implemented method of embodiment 42 or embodiment 43, wherein the selection of at least one target sequence of interest is provided as an input by a user of a system configured to perform the computer-implemented method. 45. The computer-implemented method of any one of embodiments 42-44, wherein at least one target sequence of interest comprises an amino acid sequence, a nucleotide sequence, or any combination thereof. 46. The computer-implemented method of any one of embodiments 42 to 45, wherein at least one target sequence of interest comprises a peptide sequence or a portion thereof, a protein sequence or a portion thereof, a protein domain sequence or a portion thereof, a gene sequence or a portion thereof, or any combination thereof. 47. A computer-implemented method according to any one of embodiments 42 to 46, wherein the selection of the target genome is provided as an input by a user of a system configured to carry out the computer-implemented method. 48. The computer-implemented method of any one of embodiments 42 to 47, wherein the multiple target genomes include plant genomes, fungal genomes, bacterial genomes, or any combination thereof. 49. The computer-implemented method of any one of embodiments 42 to 48, wherein the genomics database comprises a public genomics database or a proprietary genomics database. 50. A computer-implemented method according to any one of embodiments 42 to 49, wherein the search for identifying homologues of at least one target sequence comprises identifying homologues based on a probabilistic sequence alignment model. 51. The computer-implemented method of embodiment 50, wherein the probabilistic sequence alignment model is a profile hidden Markov model (pHMM). 52. The computer-implemented method of embodiment 50 or embodiment 51, wherein homologs are identified based on a comparison of the probabilistic sequence alignment model score with a predefined threshold. 53. A computer-implemented method according to any one of embodiments 42 to 52, wherein the search for identifying homologues of at least one target sequence comprises identifying homologues based on an alignment of sequences using a local sequence alignment search tool, calculating a sequence homology metric based on the alignment, and comparing the calculated sequence homology metric with a predefined threshold. 54. The computer-implemented method of embodiment 53, wherein the predetermined threshold comprises a threshold for percent sequence identity, percent sequence coverage, E value, or bit score value. 55. A computer-implemented method according to any one of embodiments 42 to 54, wherein the search for identifying homologues of at least one target sequence comprises identifying homologues based on the use of gene and / or protein domain annotation tools. 56. A computer-implemented method according to any one of embodiments 42 to 55, wherein creating a phylogenetic tree based on identified homologues of at least one target sequence comprises aligning homologue sequences using an alignment software tool, trimming the aligned homologue sequences using a sequence trimming software tool, and constructing a phylogenetic tree using a phylogenetic tree building software tool. 57. A computer-implemented method according to any one of embodiments 42 to 56, wherein at least one target sequence of interest comprises a known NETaG sequence or a core synthase gene sequence. 58. A computer-implemented method for identifying biosynthetic gene clusters (BGCs) encoding biosynthetic enzymes for producing a secondary metabolite having a desired activity, comprising: receiving a selection of at least one target sequence of interest, wherein the at least one target sequence of interest comprises a sequence encoding a therapeutic target of interest; receiving a selection of target genomes from a genomics database, the selection of target genomes including a plurality of target genomes from organisms known to produce secondary metabolites; conducting a search to identify homologs of at least one target sequence in a plurality of target genomes; generating a phylogenetic tree based on the identified homologs of at least one target sequence; Classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on a phylogenetic tree, where a positive genome is a genome belonging to a clade in which multiple copies of at least one target sequence homolog are present, and a negative genome is a genome belonging to a clade in which a single copy of at least one target sequence homolog is present, and the target sequence homolog present in multiple copies in the positive genome is a putative resistance gene; and Based at least in part on the classification of positive and negative genomes, i) one or more scores indicating the co-occurrence of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a biosynthetic gene cluster (BGC); ii) one or more scores indicating coevolution of at least one target sequence homolog (putative resistance gene) and one or more genes associated with the BGC; iii) one or more scores indicating co-regulation of at least one target sequence homolog (putative resistance gene) with one or more genes associated with BGC; and iv) one or more scores indicating co-expression of at least one target sequence homolog (putative resistance gene) with one or more genes associated with BGC; determining at least one genomic parameter selected from determining a likelihood that the putative resistance gene is an actual resistance gene associated with a BGC that produces a secondary metabolite that acts on the protein product encoded by the resistance gene based on the at least one genomic parameter; A method comprising: 59. The computer-implemented method of embodiment 58, wherein determining the likelihood that the putative resistance gene is an actual resistance gene associated with the BGC producing a secondary metabolite comprises comparing at least one determined genomic parameter to at least one predetermined threshold. 60. The computer-implemented method of embodiment 58 or embodiment 59, wherein the selection of at least one target sequence of interest is provided as an input by a user of a system configured to perform the computer-implemented method. 61. The computer-implemented method of any one of embodiments 58 to 60, wherein at least one target sequence of interest comprises an amino acid sequence, a nucleotide sequence, or any combination thereof. 62. The computer-implemented method of any one of embodiments 58 to 61, wherein at least one target sequence of interest comprises a peptide sequence or a portion thereof, a protein sequence or a portion thereof, a protein domain sequence or a portion thereof, a gene sequence or a portion thereof, or any combination thereof. 63. A computer-implemented method according to any one of embodiments 58 to 62, wherein the selection of the target genome is provided as an input by a user of a system configured to carry out the computer-implemented method. 64. The computer-implemented method of any one of embodiments 58 to 63, wherein the multiple target genomes include plant genomes, fungal genomes, bacterial genomes, or any combination thereof. 65. The computer-implemented method of any one of embodiments 58 to 64, wherein the genomics database comprises a public genomics database or a proprietary genomics database. 66. A computer-implemented method according to any one of embodiments 58 to 65, wherein the search for identifying homologues of at least one target sequence comprises identifying homologues based on a probabilistic sequence alignment model. 67. The computer-implemented method of embodiment 66, wherein the probabilistic sequence alignment model is a profile hidden Markov model (pHMM). 68. The computer-implemented method of embodiment 66 or embodiment 67, wherein homologs are identified based on a comparison of the probabilistic sequence alignment model score with a predefined threshold. 69. A computer-implemented method according to any one of embodiments 58 to 68, wherein the search for identifying homologues of at least one target sequence comprises identifying homologues based on an alignment of sequences using a local sequence alignment search tool, calculating a sequence homology metric based on the alignment, and comparing the calculated sequence homology metric with a predefined threshold. 70. The computer-implemented method of embodiment 69, wherein the predetermined threshold comprises a threshold for percent sequence identity, percent sequence coverage, E value, or bit score value. 71. A computer-implemented method according to any one of embodiments 58 to 70, wherein the search for identifying homologues of at least one target sequence comprises identifying homologues based on the use of gene and / or protein domain annotation tools. 72. A computer-implemented method according to any one of embodiments 58 to 71, wherein creating a phylogenetic tree based on identified homologues of at least one target sequence comprises aligning homologue sequences using an alignment software tool, trimming the aligned homologue sequences using a sequence trimming software tool, and constructing a phylogenetic tree using a phylogenetic tree building software tool. 73. A computer-implemented method according to any one of embodiments 58 to 72, further comprising performing an in vitro assay to test secondary metabolites produced by the identified BGC for activity against a therapeutic target of interest. 74. A computer-implemented method according to any one of embodiments 58 to 73, further comprising performing an in vivo assay to test secondary metabolites produced by the identified BGC for activity against a therapeutic target of interest. 75. A system comprising: one or more processors; A memory communicatively coupled to one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform a method according to any one of embodiments 1 to 74; Including, the system. 76. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a system, cause the system to perform a method according to any one of embodiments 1 to 74.
[0197] From the above, it should be understood that although specific implementations of the disclosed methods, devices, and systems have been illustrated and described, various modifications can be made and are contemplated herein. It is also not intended that the present invention be limited by the specific examples provided herein. Although the present invention has been described with reference to the above specification, the description and illustration of the preferred embodiments herein are not meant to be construed in a limiting sense. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions described herein, which depend on various conditions and variables. Various modifications in form and details of the embodiments of the present invention will be apparent to those skilled in the art. It is therefore contemplated that the present invention encompasses such modifications, variations, and equivalents.
Claims
1. 1. A computer-implemented method for identifying resistance genes, comprising: receiving a selection of at least one target sequence of interest; receiving a selection of target genomes from a genomics database, the selection of target genomes including a plurality of target genomes from organisms known to produce or likely to produce secondary metabolites; conducting a search to identify homologs of at least one target sequence in a plurality of target genomes; generating a phylogenetic tree based on the identified homologs of at least one target sequence; Classifying genomes of a plurality of target genomes as positive genomes or negative genomes based on a phylogenetic tree, wherein a positive genome is a genome belonging to a clade in which multiple copies of at least one target sequence homolog are present, and a negative genome is a genome belonging to a clade in which a single copy of at least one target sequence homolog is present, and the target sequence homolog present in multiple copies in the positive genome is a putative resistance gene; and Based at least in part on the classification of positive and negative genomes, i) one or more scores indicating the co-occurrence of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a biosynthetic gene cluster (BGC); ii) one or more scores indicating co-evolution of at least one target sequence homologue (putative resistance gene) and one or more genes associated with BGC; iii) one or more scores indicating co-regulation of at least one target sequence homolog (putative resistance gene) with one or more genes associated with BGC; and iv) one or more scores indicating co-expression of at least one target sequence homolog (putative resistance gene) with one or more genes associated with BGC; determining at least one genomic parameter selected from: determining the likelihood that the putative resistance gene is a resistance gene based on at least one genomic parameter; 11. A computer-implemented method comprising:
2. 2. The computer-implemented method of claim 1, wherein determining the likelihood that the putative resistance gene is a resistance gene comprises comparing the at least one determined genomic parameter to at least one predetermined threshold.
3. 10. The computer-implemented method of claim 1, wherein the selection of at least one target sequence of interest is provided as input by a user of a system configured to perform the computer-implemented method.
4. 10. The computer-implemented method of claim 1, wherein the at least one target sequence of interest comprises an amino acid sequence, a nucleotide sequence, or any combination thereof.
5. 5. The computer-implemented method of claim 4, wherein the at least one target sequence of interest comprises a peptide sequence or a portion thereof, a protein sequence or a portion thereof, a protein domain sequence or a portion thereof, a gene sequence or a portion thereof, or any combination thereof.
6. 10. The computer-implemented method of claim 1, wherein the at least one target sequence of interest comprises a mammalian sequence, a human sequence, a plant sequence, a fungal sequence, a bacterial sequence, an archaeal sequence, a viral sequence, or any combination thereof.
7. 10. The computer-implemented method of claim 1, wherein the at least one target sequence of interest comprises a primary target sequence and one or more related sequences.
8. 8. The computer-implemented method of claim 7, wherein the one or more related sequences comprise sequences functionally related to the primary target sequence.
9. 9. The computer-implemented method of claim 8, wherein the one or more associated sequences comprise sequences that are pathway-related to the primary target sequence.
10. The computer-implemented method of claim 1 , wherein the selection of the target genome is provided as input by a user of a system configured to perform the computer-implemented method.
11. The computer-implemented method of claim 1 , wherein the plurality of target genomes comprises plant genomes, fungal genomes, bacterial genomes, or any combination thereof.
12. The computer-implemented method of claim 1 , wherein the genomics database comprises a public genomics database.
13. The computer-implemented method of claim 12 , wherein the genomics database comprises a proprietary genomics database.
14. 10. The computer-implemented method of claim 1, wherein the search to identify homologs of the at least one target sequence comprises identifying homologs based on a probabilistic sequence alignment model.
15. 15. The computer-implemented method of claim 14, wherein the probabilistic sequence alignment model is a profile hidden Markov model (pHMM).
16. 16. The computer-implemented method of claim 15, wherein homologs are identified based on a comparison of the probabilistic sequence alignment model score to a predetermined threshold.
17. 2. The computer-implemented method of claim 1, wherein the search to identify homologs of at least one target sequence comprises identifying homologs based on an alignment of the sequences using a local sequence alignment search tool, calculating sequence homology metrics based on the alignment, and comparing the calculated sequence homology metrics with a predetermined threshold.
18. 18. The computer-implemented method of claim 17, wherein the local sequence alignment search tool comprises BLAST, DIAMOND, HMMER, Exonerate, or ggsearch.
19. 18. The computer-implemented method of claim 17, wherein the predetermined threshold comprises a threshold for percent sequence identity, percent sequence coverage, E-value, or bit score value.
20. 2. The computer-implemented method of claim 1, wherein the search to identify homologs of at least one target sequence comprises identifying homologs based on the use of gene and / or protein domain annotation tools.
21. 21. The computer-implemented method of claim 20, wherein the gene and / or protein domain annotation tool comprises InterProScan or EggNOG.
22. 2. The computer-implemented method of claim 1, wherein creating a phylogenetic tree based on identified homologs of at least one target sequence comprises aligning the homolog sequences using an alignment software tool, trimming the aligned homolog sequences using a sequence trimming software tool, and constructing a phylogenetic tree using a phylogenetic tree building software tool.
23. 23. The computer-implemented method of claim 22, wherein the alignment software tool comprises MAFFT, MUSCLE, or ClustalW.
24. 23. The computer-implemented method of claim 22, wherein the sequence trimming software tool comprises trimAI, GBlocks, or ClipKIT.
25. 23. The computer-implemented method of claim 22, wherein the phylogenetic tree construction software tool comprises FastTree, IQ-TREE, RAxML, MEGA, MrBayes, BEAST, or PAUP.
26. 23. The computer-implemented method of claim 22, wherein the construction of the phylogenetic tree is based on a maximum likelihood algorithm, a parsimony algorithm, a neighbor-joining algorithm, a distance matrix algorithm, or a Bayesian inference algorithm.
27. 2. The computer-implemented method of claim 1, wherein the one or more scores indicating co-occurrence are determined based on identifying a positive correlation between the presence of multiple copies of the putative resistance gene in the positive genome and the presence of one or more genes of the BGC.
28. 28. The computer-implemented method of claim 27, wherein identifying a positive correlation between the presence of multiple copies of a putative resistance gene in a positive genome and the presence of one or more genes of a BGC comprises using a clustering algorithm to cluster aligned protein sequences, aligned nucleotide sequences, aligned protein domain sequences, or aligned pHMMs for groups of BGCs to identify BGC communities within multiple target genomes.
29. 28. The computer-implemented method of claim 27, wherein identifying a positive correlation between the presence of multiple copies of a putative resistance gene in a positive genome and the presence of one or more genes of a BGC comprises using phylogenetic analysis of protein sequences or protein domains for groups of BGCs to identify BGC communities within multiple target genomes.
30. 28. The computer-implemented method of claim 27, wherein identifying a positive correlation between the presence of multiple copies of the putative resistance gene in a positive genome and the presence of one or more genes of the BGC comprises selecting genomes with a specific taxonomy to identify a BGC community within a plurality of target genomes.
31. 2. The computer-implemented method of claim 1, wherein the one or more scores indicating coevolution of the putative resistance gene and the one or more genes associated with the BGC are determined based on a coevolution correlation score, a coevolution rank score, a coevolution slope score, or any combination thereof.
32. 32. The computer-implemented method of claim 31 , wherein the coevolution correlation score is based on the correlation between the pairwise sequence identity percentage of a cluster of orthologous groups (COGs) for the putative resistance gene and the pairwise sequence identity percentage of a cluster of orthologous groups (COGs) for one of the one or more genes associated with the BGC.
33. 32. The computer-implemented method of claim 31, wherein the coevolution rank score is based on a ranking of correlation coefficients of COGs that include one of the one or more genes associated with BGC in ascending order for COGs that include the putative resistance gene.
34. 34. The computer-implemented method of claim 33, wherein in the event of a tie in distance scores, the rank for all COGs in the tie is set equal to the lowest rank in the group.
35. 32. The computer-implemented method of claim 31 , wherein the coevolutionary slope score is based on an orthogonal regression of the pairwise sequence identity percentage of the COG for the putative resistance gene and the pairwise sequence identity percentage of the COG for one of the one or more genes associated with the BGC.
36. 33. The computer-implemented method of claim 32, wherein only COGs arising from unique positive genomes having three or more genes remaining after removing corresponding genes from the negative genome are used to evaluate the coevolution correlation score, coevolution rank score, or coevolution slope score.
37. 2. The computer-implemented method of claim 1, wherein the one or more scores indicative of co-regulation are based on DNA motif detection from intergenic sequences of one or more genes associated with BGC and putative resistance genes.
38. The computer-implemented method of claim 1 , wherein the one or more scores indicative of co-expression are based on differential expression analysis and / or clustering analysis of global transcriptomic data.
39. 2. The computer-implemented method of claim 1, wherein the one or more genes associated with the biosynthetic gene cluster (BGC) include anchor genes, core synthase genes, biosynthetic genes, genes not involved in the biosynthesis of a secondary metabolite produced by the BGC, or any combination thereof.
40. 2. The computer-implemented method of claim 1, wherein the putative resistance gene is a putative integration target gene (pETaG) or a putative non-integration target gene (pNETaG).
41. 2. The computer-implemented method of claim 1, wherein the resistance gene is an integration target gene (ETaG) or a non-integration target gene (NETaG).
42. 40. The computer-implemented method of claim 39, further comprising performing an in vitro or in vivo assay to test secondary metabolites produced by the identified BGC for activity against a therapeutic target of interest.
43. 1. A system comprising: one or more processors; a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform the method of claim 1; Including, the system.
44. 10. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a system, cause the system to perform the method of claim 1.