Deep learning methods for biosynthetic gene cluster discovery
Patent Information
- Application Number
- JP2024531046
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-11-23
- Filing Date
- 2022-11-23
- Publication Date
- 2025-10-30
AI Technical Summary
Current methods struggle to accurately identify and predict biosynthetic gene clusters (BGCs) and their associated genes, particularly those not directly involved in secondary metabolite synthesis, due to challenges in defining genomic boundaries and the lack of computational pipelines for predicting functional genes and metabolites.
A computer-implemented method using advanced machine learning techniques, including new data representations and ensemble learning models, to identify genes associated with BGCs by analyzing genome sequences and protein domain patterns, utilizing trained models to detect patterns in nucleotide and protein domain data.
Enhances the ability to discover and classify BGCs and their encoded molecules, improving the identification of potential drug compounds and therapeutic targets by providing high confidence bounds and accurate gene cluster predictions.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 282,451, filed November 23, 2021, the contents of which are incorporated herein by reference in their entirety.
[0002] The present disclosure relates generally to methods and systems for identifying genes associated with biosynthetic gene clusters and applications thereof, including identifying potential therapeutic targets and drug candidates. [Background technology]
[0003] Microorganisms produce a wide variety of small molecule compounds known as secondary metabolites or natural products with diverse chemical structures and functions. Some secondary metabolites enable microorganisms to tolerate adverse environments, while others serve as weapons in inter- and intra-species competition. See, e.g., Piel, J. Nat. Prod. Rep., 26:338-362, 2009. Many human pharmaceuticals (including, e.g., antibacterial agents, antitumor agents, and insecticides) are derived from secondary metabolites. See, e.g., Newman DJ and Cragg GMJ Nat. Prod., 79:629-661, 2016.
[0004] Microorganisms synthesize secondary metabolites using enzyme proteins encoded by clusters of co-located genes called biosynthetic gene clusters (BGCs). Evidence is emerging that some microbial biosynthetic gene clusters contain genes that do not appear to be involved in the synthesis of the associated biosynthetic product produced by the enzymes encoded by the cluster. In some cases, such non-biosynthetic genes have been described as "self-protecting" because they code for proteins that can apparently confer resistance to the host organism against the associated biosynthetic product. For example, in some cases, resistance mutants of genes encoding transporters of biosynthetic products, detoxifying enzymes that act on biosynthetic products, or proteins whose activity is targeted by biosynthetic products have been reported. See, e.g., Cimermancic, et al., Cell 158:412, 2014; Keller, Nat. Chem. Biol. 11:671, 2015. In some cases, such genes may be referred to as "resistance genes." Researchers have proposed that the identification of such genes and the determination of their functions may be useful in determining the role of the biosynthetic products synthesized by the enzymes of the cluster. For example, Yeh,et al.,ACS Chem.Biol.11:2275,2016;Tang,et al.,ACS Chem.Biol.10:2841,2015;Regueira,et al.,Appl,Environ.Microbiol.77:3035,2011;Kennedy,et al.,Science See 284:1368, 1999; Lowther, et al., Proc. Natl. Acad. Sci. USA 95:12153, 1998; Abe, et al., Mol. Genet. Genomics 268:130, 2002. U.S. Patent Application Publication No. 2020 / 0211673 provides insight that certain genes present in or located in close proximity to biosynthetic genes of a cluster (particularly in eukaryotic, e.g., fungal, biosynthetic gene clusters, as opposed to bacterial biosynthetic gene clusters) may represent homologs of human genes that are targets for therapeutic purposes.Such genes are termed "embedded target genes" ("ETaGs") or "non-embedded target genes" (NETaGs), depending on whether they are located within a cluster of biosynthetic genes or not.
[0005] Traditionally, secondary metabolites have been identified from microbial cultures and screened for therapeutic activity against human targets of interest. However, most microorganisms are not culturable, and even BGCs in culturable microorganisms may remain transcriptionally silent under laboratory conditions. Recent developments in nucleic acid and protein sequencing technologies as well as bioinformatics pipelines have made it possible to rapidly identify a large number of BGCs from environmental microorganisms without the need to culture the microorganisms and test the biological activity of the BGCs. See, for example, Palazzotto E. and Weber T. Curr. Opin. Microbiol., 45:109-116, 2018. However, it remains a challenge to accurately define the genomic boundaries of BGCs using purely computational methods. There are also no computational pipelines available to identify genes associated with (but not embedded within) biosynthetic gene clusters or to predict the function of secondary metabolites and predict biosynthetic gene clusters that produce secondary metabolites with an activity of interest. Summary of the Invention
[0006] Disclosed herein are computer-implemented methods and systems for identifying genes associated with biosynthetic gene clusters (BGCs), which may be used, for example, to identify BGCs encoding potential drug compounds. BGC identification is performed through the application of advanced machine learning techniques. Innovations for computational BGC discovery include novel data representations, novel applications of advanced model architectures, and novel ensemble learning models that include separate computational models. In some examples, training data is generated from a proprietary dataset of BGCs with high-confidence boundaries. Genome-encoded molecule (GEM) compound class and function prediction is performed through a transfer learning framework with novel feature sets.
[0007] Disclosed herein is a computer-implemented method for identifying biosynthetic gene clusters, the method including receiving as input a first representation of at least one first genome, processing the representation of the at least one first genome using a trained machine learning model configured to detect patterns of predicted protein domains encoded by genes belonging to a biosynthetic gene cluster (BGC), and outputting a second representation of the at least one first genome that identifies a set of genes belonging to the BGC based on detection of patterns of predicted protein domains corresponding to the BGC in the first representation of the at least one first genome.
[0008] In some embodiments, the first representation of the at least one first genome comprises a nucleotide sequence for the at least one first genome, in some embodiments, the first representation of the at least one first genome comprises a vector representation of, or an embedding of, the at least one first genome.
[0009] In some embodiments, the first representation of the at least one first genome comprises an array of CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representations, or any combination thereof, of protein domains encoded by genes in the at least one first genome. In some embodiments, the array of protein domain representations is generated by a process that includes searching protein sequences for each gene in the at least one first genome and identifying protein domains in the protein sequences by sequence alignment against a database of CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representations, respectively. In some embodiments, the first representation of the at least one first genome further includes associated Gene Ontology (GO) terms, an identification of any resistance genes present, an identification of additional regulatory elements, or an identification of additional epigenetic elements.
[0010] In some embodiments, the computer-implemented method further comprises encoding each protein domain representation in the sequence of CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representations of the protein domains as a vector representation of the at least one first genome using a representation learning system. In some embodiments, the representation learning system comprises a graphical learning model, a deep autoencoder, Pfam2vec, word2vec, GloVe, or fastText.
[0011] In some embodiments, the first representation of the at least one first genome comprises an array of annotations of genes in the at least one first genome based on a gene function and pathway mapping database, in some embodiments, the pathway mapping database is KEGG.
[0012] In some embodiments, the first representation of the at least one first genome comprises an annotation sequence of the genes in the at least one first genome based on a database comprising data on clusters of orthologous groups (COGs). In some embodiments, the database comprising data on clusters of orthologous groups (COGs) is EggNOG.
[0013] In some embodiments, the computer-implemented method further comprises encoding the sequences of the gene annotations as vector representations of the at least one first genome using a representation learning system, in some embodiments, the representation learning system comprises a graphical learning model, a deep autoencoder, Pfam2vec, word2vec, GloVe, or fastText.
[0014] In some embodiments, the trained machine learning model comprises a deep learning model. In some embodiments, the deep learning model comprises a supervised learning model or an unsupervised learning model. In some embodiments, the deep learning model comprises a convolutional neural network, a long short-term memory network, or a transformer model. In some embodiments, the deep learning model comprises a combination of components from a neural network, a convolutional neural network, a long short-term memory network, or a transformer neural network.
[0015] In some embodiments, the machine learning model is trained using a training dataset that includes data on a plurality of training genomes. In some embodiments, the plurality of training genomes includes a plurality of synthetic training genomes. In some embodiments, one or more of the synthetic training genomes of the plurality of synthetic training genomes each includes a set of gene sequences from an actual BGC randomly inserted into a BGC-negative genome. In some embodiments, one or more of the synthetic training genomes of the plurality of synthetic training genomes each includes a set of gene sequences from a combination of actual positive BGC examples and synthetic negative BGC examples.
[0016] In some embodiments, the second representation of the at least one first genome comprises a vector representation, a graph representation, or a tensor representation of the at least one first genome.
[0017] In some embodiments, the computer-implemented method further comprises evaluating the gene identified as belonging to the BGC to determine whether it is a resistance gene. In some embodiments, the resistance gene is an embedded target gene (ETaG) or a non-embedded target gene (NETaG). In some embodiments, the computer-implemented method further comprises performing an in vitro assay to test a secondary metabolite produced by the BGC in the at least one first genome identified as belonging to the resistance gene for activity against a resistance gene homolog or a protein encoded thereby identified in a second genome different from the at least one first genome. In some embodiments, the computer-implemented method further comprises performing an in vivo assay to test a secondary metabolite produced by the BGC in the at least one first genome identified as belonging to the resistance gene for activity against a resistance gene homolog or a protein encoded thereby identified in a second genome different from the at least one first genome. In some embodiments, the second genome comprises a mammalian genome, a human genome, an avian genome, a reptile genome, an amphibian genome, a plant genome, a fungal genome, a bacterial genome, or a viral genome.
[0018] In some embodiments, the at least one first genome comprises a eukaryotic genome or a prokaryotic genome. In some embodiments, the at least one first genome is a eukaryotic genome, and the eukaryotic genome comprises a plant genome or a fungal genome. In some embodiments, the at least one first genome is a prokaryotic genome, and the prokaryotic genome is a bacterial genome.
[0019] In some embodiments, the first representation of the at least one first genome is input by a user of a system configured to perform the computer-implemented method.
[0020] Disclosed herein is a computer-implemented method that includes receiving as input a sequence for at least one first genome; generating a first representation of the at least one first genome, where the first representation of the at least one first genome includes an array of protein domain representations encoded by genes in the at least one first genome; and encoding each protein domain representation in the array of protein domain representations as a vector representation of the at least one first genome using a representation learning system.
[0021] In some embodiments, the sequence of the protein domain representation encoded by the gene in the at least one first genome comprises a sequence of a CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representation of the protein domain encoded by the gene in the at least one first genome, or any combination thereof. In some embodiments, the sequence of protein domain representations is generated by a process that includes searching protein sequences for each gene in at least one first genome and identifying protein domains in the protein sequences by sequence alignment against a database of CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representations, respectively. In some embodiments, the first representation of the at least one first genome further includes associated Gene Ontology (GO) terms, identification of any resistance genes present, identification of additional regulatory elements, or identification of additional epigenetic elements. In some embodiments, the representation learning system includes a graphical learning model, a deep autoencoder, Pfam2vec, word2vec, GloVe, or fastText. In some embodiments, the representation learning system is trained on a corpus of annotated genomes, each of which includes a sequence of CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representations, or any combination thereof, of protein domains encoded by genes in the genomes of the corpus.
[0022] Also disclosed herein is a system comprising one or more processors and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform any of the methods described herein.
[0023] Also disclosed herein is a non-transitory computer readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a system, cause the system to perform any of the methods described herein.
[0024] It should be understood that all combinations of the foregoing and additional concepts described in more detail below (unless such concepts are mutually inconsistent) are contemplated as part of the inventive subject matter disclosed herein, and in particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as part of the inventive subject matter disclosed herein.
[0025] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety. In the event of a conflict between a term in this specification and a term in an incorporated reference, the term in this specification shall control.
[0026] Various aspects of the disclosed methods, devices, and systems are set forth with particularity in the appended claims. A better understanding of the features and advantages of the disclosed methods, devices, and systems will be obtained by reference to the following detailed description of exemplary embodiments and the accompanying drawings. [Brief description of the drawings]
[0027] [Figure 1] FIG. 1 provides a non-limiting example of a process flow chart for predicting genes belonging to biosynthetic gene clusters (BGCs) in a first representation of a genome.
[0028] [Diagram 2] FIG. 1 provides a non-limiting example of a process flow chart for generating an embedded representation of a genome as an ordered list of vectors, each of which represents a particular protein domain, e.g., a Pfam domain, or annotation.
[0029] [Diagram 3] FIG. 1 provides a non-limiting schematic diagram of a computing device in accordance with one or more examples of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0030] Disclosed herein are computer-implemented methods and systems for identifying genes associated with biosynthetic gene clusters (BGCs), which may be used, for example, to identify BGCs that code for potential drug compounds. BGC identification is performed through the application of advanced machine learning techniques. Innovations for computational BGC discovery include novel data representations, novel applications of advanced model architectures, and novel ensemble learning models that include separate computational models. In some examples, training data is generated from a proprietary dataset of BGCs with high-confidence boundaries. Genome-encoded molecule (GEM) compound class and function prediction is performed through a transfer learning framework with novel feature sets.
[0031] In some examples, for example, a disclosed method (e.g., a computer-implemented method) may include receiving as input a first representation of at least one first genome, processing the representation of the at least one first genome using a trained machine learning model configured to detect patterns of predicted protein domains encoded by genes belonging to a biosynthetic gene cluster (BGC), and outputting a second representation of the at least one first genome that identifies a set of genes belonging to the BGC based on the detection of the patterns of predicted protein domains corresponding to the BGC in the first representation of the at least one first genome.
[0032] definition Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0033] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Any reference herein to "or" is intended to include "and / or" unless specifically stated otherwise and includes any and all possible combinations of one or more of the associated listed items.
[0034] As used herein, the terms "includes," "including," "comprises," and / or "comprising" specify the presence of stated features, integers, steps, operations, elements, components, and / or units, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, units, and / or groups thereof.
[0035] As used herein, the term "about" a number refers to a number that is ±10% of that number. When used in the context of a range, the term "about" refers to that range minus 10% of its minimum value and plus 10% of its maximum value.
[0036] As used herein, "secondary metabolite" refers to a small organic molecule compound produced by archaea, bacteria, fungi, or plants that is not directly involved in the normal growth, development, or reproduction of the host organism, but is required for the host organism's interaction with its environment. Secondary metabolites are also known as natural products or genetically encoded small molecules. The term "secondary metabolite" is used interchangeably herein with "biosynthetic product" when referring to the product of a biosynthetic gene cluster.
[0037] The terms "biosynthetic gene cluster" or "BGC" are used interchangeably herein and refer to a locally clustered group of one or more genes that together encode a biosynthetic pathway for the production of a secondary metabolite. Exemplary BGCs include, but are not limited to, biosynthetic gene clusters for the synthesis of nonribosomal peptide synthetases (NRPS), polyketide synthases (PKS), terpenes, and bacteriocins. See, for example, Keller N, “Fungal secondary metabolism: regulation, function and drug discovery.” Nature Reviews Microbiology 17.3 (2019): 167-180 and Fischbach M. and Voigt CA, PROKARYOTIC GENE CLUSTERS: A RICH TOOLBOX FOR SYNTHETIC BIOLOGY. In: Institute of Medicine (US) Forum on Microbial Threats. The Science and Applications of Synthetic and Systems Biology: Workshop Summary. Washington (DC): National Academies Press (US); 2011. A21. BGCs contain genes encoding signature biosynthetic proteins characteristic of each type of BGC. The longest biosynthetic gene in a BGC is referred to herein as the “core synthase gene” of the BGC. In addition to genes involved in the biosynthesis of secondary metabolites, BGCs may also encompass other genes, e.g., genes encoding products not involved in the biosynthesis of secondary metabolites, interspersed among the biosynthetic genes. These genes are referred to herein as "associated" with a BGC if their products are functionally related to a secondary metabolite of the BGC. Some genes, e.g., genes that are not involved in the biosynthesis of a secondary metabolite produced by the BGC, are referred to herein as "embedded" in a BGC if their products are functionally related to a secondary metabolite of the BGC and are located in physical proximity to the biosynthetic genes of the cluster.Some genes, e.g., genes not involved in the biosynthesis of a secondary metabolite produced by a BGC, are referred to herein as "non-embedded" if their products are functionally related to the secondary metabolite of the BGC but are not located in physical proximity to the biosynthetic genes of the BGC. "Anchor genes" refer to biosynthetic genes or genes not involved in the biosynthesis of a secondary metabolite produced by a BGC that are known to be co-localized and functionally related (i.e., associated) with the BGC.
[0038] The term "co-localization" refers to the presence of two or more genes in close spatial locations, such as about 200 kb or less, about 100 kb or less, about 50 kb or less, about 40 kb or less, about 30 kb or less, about 20 kb or less, about 10 kb or less, about 5 kb or less, or less in the genome.
[0039] The term "homolog" refers to a gene that is part of a gene group related by descent from a common ancestor (i.e., the gene sequences (i.e., nucleic acid sequences) of the gene group and / or the sequences of their protein products are inherited through a common origin). Homologs may arise through speciation events (giving rise to "orthologs"), gene duplication events, or horizontal gene transfer events. Homologs may be identified by phylogenetic methods, through the identification of common functional domains in aligned nucleic acid or protein sequences, or through sequence comparison.
[0040] The term "ortholog" refers to a gene that is part of a group of genes predicted to have evolved from a common ancestral gene by speciation.
[0041] The terms "bidirectional best hit" and "BBH" are used interchangeably herein and refer to the relationship between a pair of genes in two genomes (i.e., a first gene in a first genome and a second gene in a second genome), where the first gene or its protein product is identified as having the most similar sequence in the first genome compared to the second gene or its protein product in the second genome, and the second gene or its protein product is identified as having the most similar sequence in the second genome compared to the first gene or its protein product in the first genome. The first gene is the bidirectional best hit or BBH of the second gene, and the second gene is the bidirectional best hit, or BBH, of the first gene. BBH is a commonly used method for inferring orthology.
[0042] As used herein, "sequence similarity" between two genes means similarity in either the nucleic acid (e.g., mRNA) sequences encoded by the genes or the amino acid sequences of the gene products.
[0043] "Percent sequence identity" or "percent sequence homology" with respect to a nucleic acid sequence (or protein sequence) described herein is defined as the percentage of nucleotide residues (or amino acid residues) in a candidate sequence that are identical or homologous to the nucleotide residues (or amino acid residues) in an oligonucleotide (or polypeptide) to which the candidate sequence is compared, after aligning the sequences and taking into account any conservative substitutions as part of the sequence identity. Homology between different amino acid residues in a polypeptide is determined based on a substitution matrix, such as the BLOSUM (BLOcks SUbstitution Matrix) matrix. Methods for aligning sequences and determining percent sequence identity or percent sequence homology are well known to those skilled in the art. Examples of publicly available computer software that may be used include, but are not limited to, BLAST (Basic Local Alignment Search Tool; software for comparing amino acid sequences of proteins or nucleotide sequences of DNA and / or RNA molecules), BLAST-2, ALIGN or Megalign (DNASTAR) software. Any of a variety of appropriate parameters for measuring sequence alignment and determining percent sequence identity or homology may be determined by those skilled in the art, including the use of algorithms necessary to achieve maximal alignment over the full length of the sequences being compared.
[0044] Certain aspects of the present disclosure encompass process steps and instructions described herein in the form of algorithms. It should be noted that the process steps and instructions of the present disclosure may be embodied in software, firmware, or hardware, and when embodied in software, may be downloaded to reside and operate on different platforms used by various operating systems. Unless otherwise noted, as will be apparent from the following description, throughout the description, descriptions utilizing terms such as "processing," "computing," "calculating," "determining," "displaying," "generating," and the like, are understood to refer to the operations and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities in the memory or registers of the computer system or other such information storage, transmission, or display devices.
[0045] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described. The description is presented to enable one of ordinary skill in the art to make and use the invention and is provided in the context of a patent application and its requirements.
[0046] Deep learning methods for biosynthetic gene cluster (BGC) discovery Genome-encoded molecules (GEMs) are a diverse group of molecules, e.g., natural products or secondary metabolites, that are encoded in genomes and produced by groups of genes known as biosynthetic gene clusters (BGCs) in various eukaryotic and prokaryotic organisms. Due to the inherent challenges associated with discovering BGCs that encode potential drug compounds, most traditional early drug development approaches continue to rely on in vivo or in silico screening methods for lead compound identification. As a result, GEMs remain an underutilized resource in drug discovery.
[0047] Although several computational methods for BGC discovery based on genome sequences already exist, the performance of current computational models is far from sufficient. The difficulties stem from the high dimensionality of genomic data and the relatively small number of high-quality annotated BGCs available for use as training sets.
[0048] To address this challenge, we have acquired a collection of hundreds of thousands of fungal organisms. By fermenting and sequencing the genomes of these organisms, we have created a genomics database that contains a catalog of high-quality annotated fungal genomes. These genomes contain a vast number of BGCs and therefore represent a rich resource that allows the creation of novel computational methods for identifying and classifying BGCs and their associated GEMs. Here, we describe a method for the creation and validation of a novel computational pipeline for the effective discovery of BGCs, taking advantage of recent advances in artificial intelligence and machine learning.
[0049] Generation of high-quality datasets and synthetic training genomes for supervised learning: Annotated genomes (e.g., fungal, bacterial, or plant genomes) are obtained from genomics databases. Examples of suitable genomics databases include, but are not limited to, Brassica.info, Ensembl Plants, EnsemblFungi, National Center for Biotechnology Information (NCBI) Whole Genome Database, Plant Genome Database Japan's DNA Marker and Linkage Database, Phytozome, Plant GDB Genome Browser, FungiDB, MycoCosm 1000 Fungal Genomes Project database, FDBC Fungal Genome Database, Seoul National University Genome Browser (SNUGB) database, AspGD, etc.
[0050] Putative BGC regions are recovered and manually curated using comparative genomics techniques (see, for example, International Patent Application No. PCT / US2022 / 049016, the entire contents of which are incorporated herein) to identify BGCs with reliable boundaries. For each BGC, the nucleotide sequences of the constituent genes are translated into corresponding peptide sequences, and their functional or conserved domains are annotated, for example, using sequence alignment against the Pfam database, or via InterProScan sequence alignment against the InterPro database, or using similar protein domain annotation tools. Thus, each gene in the BGC is represented as a domain architecture. The sequences of the resulting domain architectures are retained as positive BGC examples for use in generating training data for supervised learning.
[0051] The negative BGC examples are created according to the following procedure: Annotated fungal genomes are obtained from genomics databases. Putative BGC regions are removed to create genome-like sequences that lack biosynthetic gene cluster content. The remaining genes are translated into peptide sequences and further processed, as described above, into sequences of, for example, Pfam protein domains or InterPro protein domains. The resulting sequences of domain architectures are called negative genomes. Positive BGC examples and negative genomes are randomly selected. Each domain architecture in the positive BGC examples is replaced with a random domain architecture that contains the same number of Pfam domains from the negative genome to create negative BGC examples for use in generating training data for supervised learning.
[0052] Two sets of synthetic training genomes are created. In the first set, one or more positive BGC examples selected from the subset of positive BGC examples are randomly inserted into each negative genome from the subset of negative genomes to create a training genome. In the second set, all positive and negative BGC examples are combined to create a single training genome. The additional training genomes in this set are created by permuting the ordering of the positive and negative BGC examples in the training genome.
[0053] Representation of training genomes as feature vectors: The representation learning system can be trained on a corpus of, for example, fungal genomes. To create this corpus, annotated fungal genomes are obtained from a genomics database. For each gene in the genome, we retrieve the protein sequence and annotate it according to protein domains, e.g., Pfam domains, as described above. In doing so, we reduce the genome to an ordered list of protein sequences and their constituent protein domains, e.g., Pfam domains. This corpus may then be used to develop fungus-specific embeddings for genome representations via word2vec, GloVe, fastText, or other self-supervised learning algorithms. This embedding may be further refined using the resulting representation, e.g., by training an autoencoder or other unsupervised learning algorithm. The end result is a representation learning system that can accept annotated genome representations as input and generate as output an embedded representation of the genome as an ordered list of vectors, each representing a particular protein domain, e.g., Pfam domain, or annotation.
[0054] In some examples, the representation learning system may be trained on a corpus of plant genomes, fungal genomes, bacterial genomes, or any combination thereof.
[0055] In some examples, the protein sequence for each gene in the training genome may be annotated using, for example, CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, PIRSF, or other representations of protein domains.
[0056] In some examples, each training genome may be represented as a sequence of gene annotations based on, for example, the KEGG or EggNOG databases.
[0057] Generating Class Labels for Supervised Learning: For each synthetic training genome, each annotated protein domain is assigned a class label based on whether it belongs to a positive BGC example in that synthetic training genome. For example, annotated protein domains that belong to a positive BGC example may be assigned a positive class label of 1. Annotated protein domains that do not belong to a positive BGC example may be assigned a negative class label of 0. For each training genome, these labels can be appended in order to their respective annotated protein domains to create a target vector (i.e., a vector that defines a list of dependent variables in the training dataset) for supervised learning. Because the training genomes are composed of positive BGC examples separated by non-biosynthetic or negative BGC example regions, each target vector contains a mixture of positive and negative class labels.
[0058] Training a supervised classification model: Supervised machine learning methods, such as deep learning methods, may be applied to create a computational model that associates training genomic representations with their associated class labels.
[0059] Any of a variety of supervised learning methods known to those skilled in the art may be used, including, but not limited to, deep learning methods based on modern state-of-the-art artificial neural networks such as convolutional neural networks (CNNs), long short-term memory networks (LSTMs), and transformer models.
[0060] Convolutional Neural Network: A convolutional neural network is a specialized deep neural network architecture consisting of alternating convolutional and max-pooling layers that function to learn a feature representation of an input matrix. Each convolutional layer consists of one or more filters that subdivide the input data matrix row-wise to generate genomic region-specific feature maps. These feature maps are summarized by a max-pooling layer to create a condensed representation of the original input matrix. This process can be repeated where the output of the max-pooling layer becomes the input to another pair of convolutional and max-pooling layers. The final max-pooling layer is flattened and serves as the input to a fully connected neural network with an activation function that generates the final classification.
[0061] Long Short-Term Memory Networks: Long short-term memory networks are a special recurrent neural network (RNN) architecture consisting of a collection of sequentially connected memory cells. A basic RNN cell accepts a hidden state from a previous cell, combines it with input data in the form of a single row of an input matrix, modifies it via an activation function, and outputs a new hidden state, both of which are used to calculate the classification of the row, which is then passed to the next RNN cell, which proceeds to process the next row of input data. LSTM cells perform the same basic function, but maintain an additional representation known as the cell state, and contain additional connections and activation functions that allow decisions to retain or forget information. A forget gate takes as input the previous hidden state and new input data, and applies an activation function to determine which information to forget. This is used to modify the previous cell state and overwrite the data to be forgotten with zeros. An input gate takes as input the previous hidden state and new input data, and applies an activation function to determine which information to update. A separate activation function is applied to determine the actual value of the updated information. These two activations are combined and used to update the cell state from the forget gate. Finally, the output gate combines activations from the hidden state and the updated cell state to determine the new hidden state. As with RNN cells, the hidden state is used to compute the classification, and both the new hidden state and the updated cell state are passed to the next cell. LSTMs can be unidirectional, consisting of one array of LSTM cells connected in sequence, or bidirectional, containing two chains of LSTM cells connected in opposite directions.
[0062] Transformer Model: Transformers are a state-of-the-art neural network architecture that enables parallelization and removes the sequential dependency of RNNs through the introduction of a self-attention mechanism. Transformers consist of a stack of encoders and a stack of decoders. The encoder consists of a self-attention layer and a feed-forward neural network. The decoder also contains both of these components, but also an encoder-decoder attention layer to accept and focus the input from the final encoder layer. During training, the entire input matrix is used to determine the self-attention value, while the feed-forward network is evaluated for each row individually. The output of the final decoder layer is used for classification.
[0063] In some examples, the disclosed method for identifying genes belonging to biosynthetic gene clusters in an input genome may be performed using an unsupervised machine learning approach. Unsupervised machine learning is used to identify patterns in a training dataset that includes unclassified or unlabeled data points. Examples of unsupervised machine learning models that may be used include, but are not limited to, generative models such as variational autoencoders, flow-based models, diffusion models, and generative adversarial models, or non-generative methods such as clustering or traditional autoencoders.
[0064] Regardless of the model architecture, each machine learning method generates a computational model that can accept a new encoded representation of a genome, e.g., an encoded Pfam representation, and return a vector that includes annotations that describe whether the encoded protein domain representation belongs to a BGC or not. Depending on the model performance, these solutions may be applied independently, sequentially, or further integrated via ensemble learning techniques such as bagging, boosting, or related methods.
[0065] Training of machine learning models: The weight coefficients, bias values, and thresholds, or other computational parameters of a machine learning model, such as a neural network, may be "taught" or "learned" in a training phase using one or more training datasets and any of a variety of training methods known to those skilled in the art. For example, the parameters of the neural network may be trained using input data from a training dataset and gradient descent or backpropagation techniques such that the output predictions of the trained neural network (e.g., predictions of the presence of biosynthetic gene clusters (BGCs) in a genome) match the examples contained in the training dataset. For example, the adjustable parameters of the neural network model may be obtained from a backpropagation neural network training process that may or may not be performed using the same hardware as that used to process the genomic data during the deployment phase.
[0066] Training Dataset: In some examples, as described above, the training data used to train the machine learning models of the present disclosure may include representations of one or more synthetic training genomes (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 10,000, 100,000 or more than 100,000 synthetic training genomes, or any number of synthetic training genomes within this range). In some examples, e.g., when using a supervised learning approach, the training data may include labeled training data, e.g., labeled representations of one or more synthetic training genomes. In some examples, e.g., when using an unsupervised learning approach, the training data may include unlabeled training data, e.g., unlabeled representations of one or more synthetic training genomes. In some examples, one or more training datasets may be used to train a machine learning algorithm in a training phase that is different from that in the deployment or use phase. In some examples, the training data may be continuously updated and used to update the trained machine learning algorithm in real time. In some cases, the training data may be stored in a training database that resides on a local computer or server. In some cases, the training data may be stored in a training database that resides online or in the cloud.
[0067] Machine Learning Software: Any of a variety of commercially available or open source software packages, languages, or platforms known to those skilled in the art may be used to implement the machine learning algorithms of the disclosed methods and systems. Examples include Shogun (www.shogun-toolbox.org), Mlpack (www.mlpack.org), R (r-project.org), Weka (www.cs.waikato.ac.nz / ml / weka / ), Python (www.python.org), Matlab (MathWorks, Natick, Massachusetts, www.mathworks.com), scikit-learn (www.scikit-learn.org), tensorflow (www.tensorflow.org), pytorch (www.pytorch.org), and / or keras (www.keras.io).
[0068] Distributed Computing Systems and Cloud-Based Training Databases: In some examples, the machine learning-based methods for identifying biosynthetic gene clusters (BGCs) disclosed herein may be used to process genomic data (e.g., sequence data) on one or more computers or computer systems that reside in a single physical or geographic location. In some examples, they may be deployed as part of a distributed system of computers that includes two or more computer systems that reside in two or more physical or geographic locations. Different computer systems, or components or modules thereof, may be physically located in different workspaces and / or different work sites (i.e., in different physical or geographic locations) and may be linked via a local area network (LAN), intranet, extranet, or the Internet such that training data and / or data from processing input genomes may be shared and exchanged between the sites.
[0069] In some embodiments, the training data (e.g., including synthetic training genomic data) may reside in a cloud-based database accessible to local and / or remote computer systems on which the disclosed machine learning based methods are executed. As used herein, the term "cloud-based" refers to shared or shareable storage of electronic data. Cloud-based databases and associated software may be used for archiving electronic data, sharing electronic data, and analyzing electronic data.
[0070] In some embodiments, the locally generated training data (e.g., including synthetic training genomic data) may be uploaded to a cloud-based database, accessed from there, and used to train other machine learning-based systems at the same or different sites. In some examples, the locally generated machine learning-based prediction results (e.g., detection of patterns of predicted protein domains encoded by genes belonging to biosynthetic gene clusters (BGCs), and identification of genes in the clusters) may be uploaded to a cloud-based database and used to update the training dataset in real time for continuous improvement of prediction performance.
[0071] Internal and external validation: Model performance may be evaluated internally, for example, using k-fold cross-validation. In this validation framework, the training dataset is randomly split into k groups, where k can range from 2 to n, and n is the number of training samples. In some examples, for example, n may be 2, 3, 4, 5, 6, 7, 8, 9, 10, or greater than 10. In some examples, a machine learning model is trained on k-1 groups, and then the performance of this trained model is evaluated on the remaining groups. This is done k times, such that each group serves once as a validation group and k-1 times as part of the external training group. Performance metrics across the k folds may then be aggregated.
[0072] External validation is performed, for example, by testing model performance against a gold standard set of fungal genomes that have been manually reviewed and annotated for BGC. In both internal and external validation, model performance is evaluated using typical information retrieval metrics such as precision and recall (where precision quantifies the number of positive class predictions that actually belong to the positive class, and recall quantifies the number of positive class predictions made for all positive examples in the validation data). In some examples, precision and / or recall may be at least 0.5, at least 0.6, at least 0.7, at least 0.75, at least 0.8, at least 0.85, at least 0.9, at least 0.95, at least 0.98, or at least 0.99.
[0073] Prediction of Novel Fungal BGCs: Just as the disclosed machine learning models can be used to annotate well-studied genomes to assess their ability to recover known BGCs, they can also be applied to novel genomes to identify novel BGCs. These novel BGCs can then be classified using additional computational methods, for example, according to the function and class of compounds produced by the BGCs.
[0074] FIG. 1 provides a non-limiting example of a flow chart of a process 100 for predicting genes belonging to a biosynthetic gene cluster (BGC) in a first representation of a genome. The process 100 may be performed, for example, as a computer-implemented method using software executing on one or more processors of one or more electronic devices, computers, or computing platforms. In some examples, the process 100 is performed using a client-server system, and blocks of the process 100 are divided in any manner between a server and a client device. In other examples, blocks of the process 100 are divided between a server and multiple client devices. Thus, although portions of the process 100 are described herein as being performed by a particular device of a client-server system, it will be understood that the process 100 is not so limited. In other examples, the process 100 is performed using only a client device or only multiple client devices. In the process 100, some blocks are combined in any manner, the order of some blocks is changed in any manner, and some blocks are omitted in any manner. In some examples, additional steps may be performed in combination with the process 100. Accordingly, the operations illustrated (and described in more detail below) are exemplary in nature and thus should not be considered as limiting.
[0075] 1, a first representation of at least one first genome is received. For example, the first representation of the at least one first genome may be input by a user of a system configured to perform the computer-implemented methods described herein.
[0076] In some examples, the at least one first genome may comprise a eukaryotic genome or a prokaryotic genome. In some examples, the at least one first genome may be a eukaryotic genome, and the eukaryotic genome may comprise a plant genome or a fungal genome. In some examples, the at least one first genome may be a prokaryotic genome, and the prokaryotic genome may be a bacterial genome.
[0077] In some examples, the at least one first genome may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 400, 600, 800, 1,000, or greater than 1,000 first (or input) genomes, or any number of first (or input) genomes within this range.
[0078] In some examples, the first representation of the at least one first genome may include a nucleotide sequence for the at least one first genome. In some examples, the first representation of the at least one first genome may include a vector representation of, or an embedding of, the at least one first genome.
[0079] In some examples, the first representation of the at least one first genome may include a sequence of a CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representation of a protein domain encoded by a gene in the at least one first genome, or any combination thereof. In some examples, the sequence of protein domain representations is generated by a process that includes searching for protein sequences for each gene in the at least one first genome and identifying protein domains in the protein sequences by sequence alignment to a database of CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representations, respectively. In some examples, the first representation of the at least one first genome may further include associated Gene Ontology (GO) terms, identification of any resistance genes present in the at least one first genome, identification of additional regulatory elements such as promoters, enhancers, or silencers present in the at least one first genome, or identification of additional epigenetic elements such as histone folding, DNA methylation, or acetylation present in the at least one first genome.
[0080] In some examples, the method may further include encoding each protein domain representation in the sequence of CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representations of the protein domains as a vector representation of the at least one first genome using a representation learning system. In some examples, the representation learning system includes a graphical learning model, a deep autoencoder, Pfam2vec, word2vec, GloVe, or fastText.
[0081] In some examples, the first representation of the at least one first genome may include an annotation sequence of the genes in the at least one first genome based on a gene function and pathway mapping database, in some examples, for example, the pathway mapping database may be KEGG.
[0082] In some examples, the first representation of the at least one first genome may include an annotation sequence of the genes in the at least one first genome based on a database that includes data on clusters of orthologous groups (COGs). In some examples, for example, the database that includes data on clusters of orthologous groups (COGs) is EggNOG.
[0083] In step 104 of FIG. 1, the representation of the at least one first genome is processed using a trained machine learning model configured to detect patterns of predicted protein domains encoded by genes belonging to biosynthetic gene clusters (BGCs).
[0084] In some examples, the trained machine learning model may include a deep learning model. For example, the deep learning model may include a supervised learning model (e.g., a supervised deep learning model) or an unsupervised learning model (e.g., an unsupervised deep learning model). In some examples, the deep learning model includes a convolutional neural network, a long short-term memory network, or a transformer model. In some examples, the deep learning model may include a combination of components from a neural network, a convolutional neural network, a long short-term memory network, or a transformer neural network.
[0085] In some examples, the machine learning model may be trained using a training dataset that includes data on a plurality of training genomes (e.g., representations thereof). In some examples, the plurality of training genomes may include a plurality of synthetic training genomes. In some examples, one or more of the synthetic training genomes of the plurality of synthetic training genomes may each include a set of gene sequences from an actual BGC randomly inserted into a BGC-negative genome. In some examples, one or more of the synthetic training genomes of the plurality of synthetic training genomes may each include a set of gene sequences from a combination of an actual BGC and an artificially generated non-BGC.
[0086] In some examples, the training dataset may include data for at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 1000, 10,000, 100,000, or more than 100,000 training genomes (e.g., representations thereof) (e.g., data for at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 1000, 10,000, 100,000, or more than 100,000 synthetic training genomes (e.g., representations thereof)), or any number of training genomes (or synthetic training genomes) within this range.
[0087] 1, a second representation of the at least one first genome is output that identifies a set of genes belonging to the BGC based on detecting the pattern of predicted protein domains corresponding to the BGC in the first representation of the at least one first genome. In some examples, the second representation of the at least one first genome includes a vector representation, a graph representation, or a tensor representation of the at least one first genome.
[0088] In some examples, the computer-implemented method may further include evaluating the gene identified as belonging to the BGC to determine whether it is a resistance gene. In some examples, the resistance gene may be an embedded target gene (ETaG) or a non-embedded target gene (NETaG).
[0089] In some examples, the methods described herein may further include using output of the computer-implemented method (e.g., identification of a resistance gene associated with the BGC) to perform an in vitro assay to test a secondary metabolite produced by the BGC in at least one first genome to which the resistance gene is identified, for activity against a resistance gene homolog or a protein encoded thereby identified in a second genome distinct from the at least one first genome.
[0090] In some examples, the methods described herein may further include using output of the computer-implemented method (e.g., identification of a resistance gene associated with the BGC) to perform an in vivo assay to test a secondary metabolite produced by the BGC in at least one first genome to which the resistance gene was identified, for activity against a resistance gene homolog or a protein encoded thereby identified in a second genome distinct from the at least one first genome.
[0091] In some examples, the second genome may include a mammalian genome, a human genome, an avian genome, a reptilian genome, an amphibian genome, a plant genome, a fungal genome, a bacterial genome, or a viral genome.
[0092] FIG. 2 provides a non-limiting example of a process flow chart for generating an embedded representation of a genome as an ordered list of vectors, each of which represents a particular protein domain, e.g., a Pfam domain, or annotation.
[0093] In step 202 of Figure 2, sequences for at least one first genome are received as input. In some examples, the at least one first genome may be input by a user of a system configured to perform the computer-implemented methods described herein.
[0094] In some examples, the at least one first genome may comprise a eukaryotic genome or a prokaryotic genome as described elsewhere herein. In some examples, the at least one first genome may be a eukaryotic genome, and the eukaryotic genome may comprise a plant genome or a fungal genome. In some examples, the at least one first genome may be a prokaryotic genome, and the prokaryotic genome may be a bacterial genome.
[0095] In some examples, the at least one first genome may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 400, 600, 800, 1,000, or greater than 1,000 first (or input) genomes, or any number of first (or input) genomes within this range.
[0096] In step 204 of FIG. 2, a first representation of the at least one first genome is generated, where the first representation of the at least one first genome includes sequences of protein domain representations, such as Pfam domains or other protein domain representations, encoded by genes in the at least one first genome.
[0097] In some examples, the first representation of the at least one first genome may include a sequence of a CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representation of a protein domain encoded by a gene in the at least one first genome, or any combination thereof, as described elsewhere herein. In some examples, the sequence of protein domain representations is generated by a process that includes searching for protein sequences for each gene in the at least one first genome and identifying protein domains in the protein sequences by sequence alignment to a database of CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representations, respectively. In some examples, the first representation of the at least one first genome may further include associated Gene Ontology (GO) terms, identification of any resistance genes present in the at least one first genome, identification of additional regulatory elements such as promoters, enhancers, or silencers present in the at least one first genome, or identification of additional epigenetic elements such as histone folding, DNA methylation, or acetylation present in the at least one first genome.
[0098] In step 206 of FIG. 2, an encoding of each protein domain representation in the array of protein domain representations as a vector representation of the at least one first genome is output using the representation learning system.
[0099] In some examples, the representation learning system may include a graphical learning model, a deep autoencoder, Pfam2vec, word2vec, GloVe, or fastText, as described elsewhere herein.
[0100] In some examples, as described elsewhere herein, the representation learning system may be trained on a corpus of annotated genomes (e.g., annotated fungal genomes), each of which includes sequences of CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representations, or any combination thereof, of protein domains encoded by genes in the genomes of the corpus.
[0101] Using Random Forests to Predict Compound Classes and Functions in Novel Fungal BGCs Generation of class labels for supervised learning: Biosynthetic product classes and compound activities of bacterial and fungal BGCs for use as training data may be obtained in JSON format, for example, from the MIBiG database (version 2.0; https: / / mibig.secondarymetabolites.org / ). Product classes are obtained from the "biosyn_class" field and compound activities are obtained from the "chem_act" field. Product classes describe the molecular type of the GEM associated with the BGC and include classifications such as polyketides, sugars, or non-ribosomal peptides. Compound activities represent the chemical activity of the GEM associated with the BGC and include annotations such as antibacterial, antifungal, or cytotoxic. BGCs may belong to more than one product class or have more than one type of compound activity. BGCs with unknown activity or product class are omitted from the training set. Each BGC is assigned a label vector, where each element of the vector corresponds to a unique product class or chemical activity. An element of the label vector is marked with a 1 if the BGC produces that product class or the product has the corresponding activity, and marked with a 0 otherwise.
[0102] Representation of BGC as a feature vector: For each gene in the BGC, we translate the nucleotide sequence into the corresponding peptide sequence and identify predicted protein domains. For example, Pfam protein domains may be identified using InterProScan. These genes in the BGC can then be further described as the combination of Pfam domains that they contain in order from beginning to end. These combinations are called domain architectures. Furthermore, we annotate the BGC genes with associated Gene Ontology (GO) terms and the presence of any resistance genes or additional regulatory or epigenetic elements.
[0103] To represent these annotations as input vectors, we create unique input matrices consisting of individual representation schemes or combinations such as, but not limited to: Pfam domains (or other protein domain representations) and copy numbers of resistance genes · Copy number of GO terms Number of copies of the domain architecture Copy number of GO terms and domain architecture
[0104] Training Random Forests: For a given input matrix, each feature vector is mapped to its corresponding label vector. Each input matrix is used to train a separate Random Forest classification model to perform multi-label classification for product classes, multi-label classification for chemical activities, binary classification for each product class, and binary classification for each chemical activity.
[0105] Model performance may be evaluated in a cross-validation framework as described above. Feature selection is performed by recursive feature elimination of features with low or null contribution scores as measured by the GINI criterion. Class imbalance for binary classification is addressed by downsampling of the majority class in the training set, creating an ensemble of models and evaluated on an additional validation set held in reserve.
[0106] This process may be performed, for example, using bacterial BGCs only, fungal BGCs, and fungal+bacterial BGCs to identify the best model for each classification task.
[0107] Identification of feature combinations important for classification: logical rules of the form "if x>1 and y<=3 and z<4, predict class A" are identified for each classification task until the complete training set is explained by a compact set of rules. Rules are identified by traversing the structure of the final tree-based model and greedily reconstructing the minimal set of conditions that associate a BGC with its correct label subject to regularization. For each classification task, an optimal set of association rules is identified using the Certified Optimal Rule Lists (CORELS) algorithm.
[0108] Alternative Methods: Existing methods for BGC identification from genomes include ClusterFinder, antiSMASH, DeepBGC, and TOUCAN. ClusterFinder uses a hidden Markov model (HMM) trained on a collection of bacterial BGCs. DeepBGC utilizes a bidirectional LSTM also trained on bacterial BGCs. Both solutions perform less well when applied to identify BGCs in fungi. antiSMASH consists of a rule-based expert system that integrates data from several different profile hidden Markov models and is the current standard approach for BGC discovery. TOUCAN is a combinatorial framework that utilizes three support vector machines, a multilayer perceptron, a logistic regression, and a random forest algorithm. However, it does not include functionality for combining predictions from these different methods into a single output.
[0109] Purpose The computer-based methods for predicting the presence of BGCs and identifying their associated genes described herein have a variety of uses, including, for example, performing further evaluation of genes predicted to be part of a BGC to (i) identify homologs or orthologs of one or more target sequences (e.g., gene sequences) of interest in one or more target genomes, (ii) identify resistance genes to secondary metabolites produced by the BGC in the target genome, (iii) predict the function of a secondary metabolite produced by the BGC, and / or (iv) identify BGCs that encode biosynthetic enzymes for producing a secondary metabolite having an activity of interest (e.g., a therapeutic activity of interest).
[0110] Methods for evaluating genes embedded within or associated with BGCs to identify resistance genes (e.g., "embedded target genes" (ETaGs) or "non-embedded target genes" (NETaGs)) are described in International Patent Applications PCT / US2022 / 049016, PCT / US2022 / 049040, and PCT / US2022 / 079965, the contents of each of which are incorporated herein in their entirety. In some examples, for example, a method for identifying resistance genes (e.g., embedded target genes (ETaGs) and / or non-embedded target genes (NETaGs)) includes receiving a selection of at least one target sequence of interest; receiving a selection of target genomes from a genomics database, where the selection of target genomes includes a plurality of target genomes from organisms known to produce or likely to produce secondary metabolites; performing a search to identify homologs of the at least one target sequence in the plurality of target genomes; generating a phylogenetic tree based on the identified homologs of the at least one target sequence; and classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on the phylogenetic tree, where a positive genome is a genome that belongs to a clade in which multiple copies of at least one target sequence homolog are present. classifying genomes of the plurality of target genomes as positive or negative genomes, where the negative genome is a genome belonging to a clade in which a single copy of at least one target sequence homolog is present and the target sequence homolog present in multiple copies in the positive genome is a putative resistance gene (e.g., pETaG or pNETaG); and based at least in part on the classification of the positive and negative genomes, determining: i) one or more scores indicative of co-occurrence of at least one target sequence homolog (putative resistance gene (e.g., pETaG or pNETaG)) and one or more genes associated with a biosynthetic gene cluster (BGC); ii) one or more scores indicative of co-evolution of at least one target sequence homolog (putative resistance gene (e.g., pETaG or pNETaG)) and one or more genes associated with a BGC; iii) one or more scores indicative of co-evolution of at least one target sequence homolog (putative resistance gene (e.g.,and iv) one or more scores indicative of co-regulation of at least one target sequence homolog (a putative resistance gene (e.g., pETaG or pNETaG)) with one or more genes associated with the BGC, and iv) one or more scores indicative of co-expression of at least one target sequence homolog (a putative resistance gene (e.g., pETaG or pNETaG)) with one or more genes associated with the BGC; and determining a likelihood that the putative resistance gene (e.g., pETaG or pNETaG) is a resistance gene (e.g., an embedded target gene (ETaG) or a non-embedded target gene (NETaG)) based on the at least one genomic parameter.
[0111] In some examples, determining the likelihood that the putative resistance gene is a resistance gene may comprise comparing the at least one determined genomic parameter to at least one predetermined threshold.
[0112] In some examples, the selection of the at least one target sequence of interest may be provided as an input by a user of a system configured to perform the computer-implemented method. In some examples, for example, the at least one target sequence of interest may include a sequence of a gene identified as belonging to a BGC by any of the methods described elsewhere herein.
[0113] In some examples, the at least one target sequence of interest may comprise an amino acid sequence, a nucleotide sequence, or any combination thereof. In some examples, the at least one target sequence of interest may comprise a peptide sequence or a portion thereof, a protein sequence or a portion thereof, a protein domain sequence or a portion thereof, a gene sequence or a portion thereof, or any combination thereof. In some examples, the at least one target sequence of interest may comprise a mammalian sequence, a human sequence, a plant sequence, a fungal sequence, a bacterial sequence, an archaeal sequence, a viral sequence, or any combination thereof.
[0114] In some instances, the at least one target sequence of interest may include a primary target sequence and one or more related sequences. In some instances, the one or more related sequences may include sequences that are functionally related to the primary target sequence. In some instances, the one or more related sequences may include sequences that are pathway related to the primary target sequence.
[0115] In some examples, the selection of the target genome may be provided as an input by a user of a system configured to perform the computer-implemented method. In some examples, the plurality of target genomes may include plant genomes, fungal genomes, bacterial genomes, or any combination thereof. In some examples, the genomics database may include a public genomics database. In some examples, the genomics database includes a proprietary genomics database.
[0116] In some examples, the search for identifying at least one homolog of the target sequence (e.g., a homolog of the gene sequence identified as belonging to the BGC) may include identifying the homolog based on a probabilistic sequence alignment model. In some examples, the probabilistic sequence alignment model is a profile hidden Markov model (pHMM). In some examples, the homolog is identified based on comparing the probabilistic sequence alignment model score with a predetermined threshold.
[0117] In some examples, the search for identifying homologs of at least one target sequence may include identifying homologs based on an alignment of sequences using a local sequence alignment search tool, calculating sequence homology metrics based on the alignment, and comparing the calculated sequence homology metrics to a predetermined threshold. In some examples, the local sequence alignment search tool includes BLAST, DIAMOND, HMMER, Exonerate, or ggsearch. In some examples, the predetermined threshold may include a threshold for sequence identity percentage, sequence coverage percentage, E value, or bit score value.
[0118] In some examples, the search to identify homologs of at least one target sequence may include identifying homologs based on the use of gene and / or protein domain annotation tools. In some examples, the gene and / or protein domain annotation tools include InterProScan or EggNOG.
[0119] In some examples, generating a phylogenetic tree based on the identified homologs of at least one target sequence may include aligning the homolog sequences using an alignment software tool, trimming the aligned homolog sequences using a sequence trimming software tool, and constructing a phylogenetic tree using a phylogenetic tree construction software tool. In some examples, the alignment software tool includes MAFFT, MUSCLE, or ClustalW. In some examples, the sequence trimming software tool includes trimAI, GBlocks, or ClipKIT. In some examples, the phylogenetic tree construction software tool includes FastTree, IQ-TREE, RAxML, MEGA, MrBayes, BEAST, or PAUP. In some examples, the construction of the phylogenetic tree may be based on a maximum likelihood algorithm, a parsimony algorithm, a neighbor-joining algorithm, a distance matrix algorithm, or a Bayesian estimation algorithm.
[0120] In some examples, the score or scores indicating co-occurrence may be determined based on identifying a positive correlation between the presence of multiple copies of the putative resistance gene in the positive genome and the presence of one or more genes of the BGC. In some examples, identifying a positive correlation between the presence of multiple copies of the putative resistance gene in the positive genome and the presence of one or more genes of the BGC may include using a clustering algorithm to cluster aligned protein sequences, aligned nucleotide sequences, aligned protein domain sequences, or aligned pHMMs for a group of BGCs to identify a BGC community in the multiple target genomes. In some examples, identifying a positive correlation between the presence of multiple copies of the putative resistance gene in the positive genome and the presence of one or more genes of the BGC may include using a phylogenetic analysis of protein sequences or protein domains for a group of BGCs to identify a BGC community in the multiple target genomes. In some examples, identifying a positive correlation between the presence of multiple copies of the putative resistance gene in the positive genome and the presence of one or more genes of the BGC may include selecting a genome with a specific taxonomy to identify a BGC community in the multiple target genomes.
[0121] In some examples, one or more scores indicating the co-evolution of the putative resistance gene and the one or more genes associated with the BGC may be determined based on a co-evolution correlation score, a co-evolution rank score, a co-evolution gradient score, or any combination thereof. In some examples, the co-evolution correlation score may be based on the correlation between the pairwise sequence identity percentage of the cluster of orthologous groups (COGs) for the putative resistance gene and the pairwise sequence identity percentage of the cluster of orthologous groups (COGs) for one of the one or more genes associated with the BGC. In some examples, the co-evolution rank score may be based on the ranking of the correlation coefficients of the COGs including one of the one or more genes associated with the BGC in ascending order with respect to the COG including the putative resistance gene. In some examples, in the event of a tie for the distance score, the rank of all COGs in the tie may be set equal to the lowest rank in the group. In some examples, the co-evolution gradient score may be based on an orthogonal regression of the pairwise sequence identity percentage of the COGs for the putative resistance gene and the pairwise sequence identity percentage of the COGs for one of the one or more genes associated with the BGC. In some instances, only COGs resulting from unique positive genomes with greater than three genes remaining after removing the corresponding genes from the negative genome are used to evaluate the coevolution correlation score, coevolution rank score, or coevolution gradient score.
[0122] In some examples, the score or scores indicative of co-regulation may be based on DNA motif detection from intergenic sequences of one or more genes associated with the BGC and the putative resistance gene.
[0123] In some examples, the score or scores indicative of co-expression may be based on differential expression and / or clustering analysis of global transcriptomic data.
[0124] In some examples, the one or more genes associated with a biosynthetic gene cluster (BGC) may include anchor genes, core synthase genes, biosynthetic genes, genes not involved in the biosynthesis of a secondary metabolite produced by the BGC, or any combination thereof.
[0125] In some instances, the putative resistance gene may be a putative embedded target gene (pETaG) or a putative non-embedded target gene (pNETaG).
[0126] In some instances, the resistance gene may be an implanted targeted gene (ETaG) or a non-implanted targeted gene (NETaG).
[0127] In some examples, a method for predicting a function of a secondary metabolite includes receiving a selection of at least one target sequence of interest, where the at least one target sequence of interest corresponds to a gene sequence associated with a biosynthetic gene cluster (BGC) known to produce a secondary metabolite; receiving a selection of target genomes from a genomics database, where the selection of target genomes includes a plurality of target genomes from organisms known to produce a secondary metabolite; performing a search to identify homologs of the at least one target sequence in the plurality of target genomes; generating a phylogenetic tree based on the identified homologs of the at least one target sequence; and classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on the phylogenetic tree, where a positive genome is a genome that belongs to a clade in which multiple copies of the at least one target sequence homolog are present and a negative ... positive genome is a genome that belongs to a clade in which multiple copies of the at least one target sequence homolog are present and a negative genome is a genome that belongs to a clade in which multiple copies of the at least one target sequence homolog are present and a negative genome is a genome classifying genomes of a plurality of target genomes as positive genomes or negative genomes, the genomes belonging to a clade in which a plurality of copies of a target sequence homolog is present in the positive genome, and the target sequence homolog present in multiple copies in the positive genome is a putative resistance gene; and determining, based at least in part on the classification of the positive genomes and the negative genomes, at least one genomic parameter selected from: i) one or more scores indicative of co-occurrence of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a BGC; ii) one or more scores indicative of co-evolution of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a BGC; iii) one or more scores indicative of co-regulation of at least one target sequence homolog (putative resistance gene) with one or more genes associated with a BGC; and iv) one or more scores indicative of co-expression of at least one target sequence homolog (putative resistance gene) with one or more genes associated with a BGC; and determining, based on the at least one genomic parameter, that the putative resistance gene isand determining the likelihood that the resistance gene encodes a protein target that is acted upon by the secondary metabolite.
[0128] In some examples, a method for identifying a biosynthetic gene cluster (BGC) encoding biosynthetic enzymes for producing a secondary metabolite having an activity of interest includes receiving a selection of at least one target sequence of interest, wherein the at least one target sequence of interest comprises a sequence encoding a therapeutic target of interest; receiving a selection of target genomes from a genomics database, wherein the selection of target genomes comprises a plurality of target genomes from organisms known to produce secondary metabolites; performing a search to identify homologs of the at least one target sequence in the plurality of target genomes; generating a phylogenetic tree based on the identified homologs of the at least one target sequence; and classifying genomes of the plurality of target genomes as positive genomes or negative genomes based on the phylogenetic tree, wherein a positive genome is a genome that belongs to a clade in which multiple copies of at least one target sequence homolog are present and a negative genome is a genome that belongs to a clade in which a single copy of the at least one target sequence homolog is present. and classifying genomes of a plurality of target genomes as positive genomes or negative genomes, the genomes belonging to a clade in which a target sequence homolog present in multiple copies in the positive genome is a putative resistance gene; and determining, based at least in part on the classification of the positive genomes and the negative genomes, at least one genomic parameter selected from: i) one or more scores indicative of co-occurrence of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a biosynthetic gene cluster (BGC); ii) one or more scores indicative of co-evolution of at least one target sequence homolog (putative resistance gene) and one or more genes associated with a BGC; iii) one or more scores indicative of co-regulation of at least one target sequence homolog (putative resistance gene) with one or more genes associated with a BGC; and iv) one or more scores indicative of co-expression of at least one target sequence homolog (putative resistance gene) with one or more genes associated with a BGC; and determining, based on the at least one genomic parameter, that the putative resistance gene isand determining the likelihood that the resistance gene is an actual resistance gene associated with the BGC that produces a secondary metabolite that acts on the protein product encoded by the resistance gene.
[0129] In some examples, the methods of the disclosure may further include performing an in vitro assay, e.g., an assay to detect or measure the activity (e.g., receptor binding activity, enzyme activating activity, enzyme inhibitory activity, etc.) of the secondary metabolite (or analog thereof) against a mammalian (e.g., human) protein encoded by a mammalian (e.g., human) gene that is homologous to ETaG or NETaG identified in an organism that includes a biosynthetic gene cluster (BGC) that produces the secondary metabolite. In some examples, the methods may further include performing an in vitro assay to detect or measure the activity (e.g., receptor binding activity, enzyme activating activity, enzyme inhibitory activity, etc.) of the secondary metabolite (or analog thereof) against a protein (e.g., reptilian, avian, amphibian, plant, fungal, bacterial, or viral protein) encoded by a reptilian, avian, amphibian, plant, fungal, bacterial, or viral gene that is homologous to ETaG or NETaG identified in an organism that includes a biosynthetic gene cluster (BGC) that produces the secondary metabolite.
[0130] In some examples, the methods of the disclosure may further include performing an in vivo assay, e.g., an assay to detect or measure activity (e.g., receptor binding activity, enzyme activating activity, enzyme inhibitory activity, intracellular signaling pathway activity, disease response, etc.) of the secondary metabolite (or analog thereof) against a mammalian (e.g., human) protein encoded by a mammalian (e.g., human) gene that is homologous to ETaG or NETaG identified in an organism that includes a biosynthetic gene cluster (BGC) that produces the secondary metabolite. In some examples, the methods may further include performing an in vivo assay to detect or measure activity (e.g., receptor binding activity, enzyme activating activity, enzyme inhibitory activity, intracellular signaling pathway activity, disease response, etc.) of the secondary metabolite (or analog thereof) against a protein (e.g., reptile, avian, amphibian, plant, fungal, bacterial, or viral protein) encoded by a reptile, avian, amphibian, plant, fungal, bacterial, or viral gene that is homologous to ETaG or NETaG identified in an organism that includes a biosynthetic gene cluster (BGC) that produces the secondary metabolite.
[0131]
[0132]
[0133]
[0134]
[0135]
[0136]
[0137]
[0138]
[0139]
[0140] In some examples, the methods of the disclosure may be used, for example, to identify and / or characterize mammalian (e.g., human) targets of secondary metabolites (or analogs thereof) produced by a BGC. In some examples, the methods of the disclosure may be used to identify and / or characterize reptile, avian, amphibian, plant, fungal, bacterial, viral targets, or targets from any other organism of secondary metabolites (or analogs thereof) produced by a BGC.
[0141] In some examples, the methods of the present disclosure may be used in drug discovery efforts, for example, to identify small molecule modulators of a mammalian (e.g., human) target gene. In some examples, the methods of the present disclosure may be used to identify small molecule modulators of a reptilian target gene, an avian target gene, an amphibian target gene, a plant target gene, a fungal target gene, a bacterial target gene, a viral target gene, or a target gene from any other organism.
[0142] In some instances, the secondary metabolite is a product of an enzyme encoded by a BGC or a salt thereof, including non-naturally occurring salts. In some instances, the secondary metabolite or an analog thereof is an analog of the product of an enzyme encoded by a BGC, such as a small molecule compound having the same core structure as the secondary metabolite or a salt thereof.
[0143] In some examples, the disclosure provides a method of regulating a human target (or a target from another organism), comprising providing a secondary metabolite or an analog thereof produced by an enzyme encoded by a BGC, wherein the human target (or a nucleic acid sequence encoding the human target) is homologous to an ETaG or NETaG associated with the BGC as determined using any one of the methods described herein.
[0144] In some examples, the disclosure provides a method of treating a condition, disorder, or disease associated with a human target (or a target from another organism), comprising administering to a subject susceptible to or suffering from the same a secondary metabolite produced by an enzyme encoded by a BGC, or an analog thereof, wherein the human target (or a nucleic acid sequence encoding the human target) is homologous to an ETaG or NETaG associated with the BGC as determined using any one of the methods described herein.
[0145] In some examples, the secondary metabolite is produced by a fungus. In some examples, the secondary metabolite is acyclic. In some examples, the secondary metabolite is a polyketide. In some examples, the secondary metabolite is a terpene compound. In some examples, the secondary metabolite is a non-ribosomally synthesized peptide.
[0146] In some examples, an analog of a substance (e.g., a secondary metabolite) that shares one or more specific structural features, elements, components, or moieties with a reference substance. Typically, an analog shows significant structural similarity with a reference substance, for example, sharing a core or consensus structure, but also differs in a specific discrete manner. In some examples, an analog is a substance that can be generated from a reference substance, for example, by chemical manipulation of the reference substance. In some examples, an analog is a substance that can be generated by carrying out a synthetic process that is substantially similar (e.g., shares multiple steps) to that which produces the reference substance. In some examples, an analog is generated or can be generated by carrying out a synthetic process that is different from that used to generate the reference substance. In some examples, an analog of a substance is a substance that is substituted at one or more of its substitutable positions.
[0147] In some examples, the analog of the product comprises the structural core of the product. In some examples, the biosynthetic product is cyclic, e.g., monocyclic, bicyclic, or polycyclic, and the structural core of the product is or comprises a monocyclic, bicyclic, or polycyclic ring system. In some examples, the structural core of the product comprises one ring of the bicyclic or polycyclic ring system of the product. In some examples, the product is or comprises a polypeptide, and the structural core is the backbone of the polypeptide. In some examples, the product is or comprises a polyketide, and the structural core is the backbone of the polyketide. In some examples, the analog is a substituted biosynthetic product that comprises one or more suitable substitutents.
[0148] A system for biosynthetic gene cluster (BGC) discovery: Also disclosed herein is a system designed to implement any of the disclosed machine learning based methods for identifying BGCs in a genome. The system may, for example, comprise one or more processors and a memory unit communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to receive as input a first representation of at least one first genome, process the representation of the at least one first genome using a trained machine learning model configured to detect patterns of predicted protein domains encoded by genes belonging to biosynthetic gene clusters (BGCs), and output a second representation of the at least one first genome that identifies a set of genes belonging to a BGC based on detection of patterns of predicted protein domains corresponding to BGCs in the first representation of the at least one first genome.
[0149] Computing devices and systems: FIG. 3 illustrates an example of a computing device according to one or more examples of the present disclosure. The device 300 may be a host computer connected to a network. The device 200 may be a client computer or a server. As shown in FIG. 3, the device 300 may be any suitable type of microprocessor-based device, such as a personal computer, a workstation, a server, or a handheld computing device (portable electronic device) such as a phone or tablet. The device may include, for example, one or more of a processor 310, an input device 320, an output device 330, a storage 340, and a communication device 360. The input device 320 and the output device 330 may generally correspond to those described above, and they may be connectable to or integrated with the computer.
[0150] The input device 320 may be any suitable device that provides input, such as a touch screen, a keyboard or keypad, a mouse, or a voice recognition device. The output device 330 may be any suitable device that provides output, such as a touch screen, a tactile device, or a speaker.
[0151] Storage 340 may be any suitable device providing storage, such as electrical, magnetic, or optical memory, including RAM, cache, hard drives, or removable storage disks. Communications device 360 may include any suitable device capable of sending and receiving signals over a network, such as a network interface chip or device. The components of the computer may be connected in any suitable manner, such as via a physical bus 370 or wirelessly.
[0152] The software 350 that may be stored in the memory / storage 340 and executed by the processor 310 may include, for example, programming that embodies the functionality of the present disclosure (eg, as embodied in the devices described above).
[0153] The software 350 may also be stored and / or transported in any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch and execute instructions associated with the software from the instruction execution system, apparatus, or device. In the context of this disclosure, a computer-readable storage medium may be any medium, such as storage 340, that can contain or store programming for use by or in connection with an instruction execution system, apparatus, or device.
[0154] The software 350 may also be propagated in any transport medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of this disclosure, a transport medium may be any medium that can communicate, propagate, or transport programming for use by or in connection with an instruction execution system, apparatus, or device. Transport-readable media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, or infrared wired or wireless propagation media.
[0155] The device 300 may be connected to a network, which may be any suitable type of interconnected communication system. The network may implement any suitable communication protocol and may be protected by any suitable security protocol. The network may comprise any suitable configuration of network links capable of effecting transmission and reception of network signals, such as wireless network connections, T1 or T3 lines, cable networks, DSL, or telephone lines.
[0156] Device 300 may implement any operating system suitable for operating on a network. Software 350 may be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, application software embodying functionality of the present disclosure may be deployed in different configurations, such as, for example, as a web-based application or web service, in a client / server configuration, or via a web browser.
[0157] Exemplary embodiments Among the embodiments provided are the following: 1. A computer-implemented method for identifying biosynthetic gene clusters, comprising: receiving as input a first representation of at least one first genome; processing the representation of the at least one first genome using a trained machine learning model configured to detect patterns of predicted protein domains encoded by genes belonging to a biosynthetic gene cluster (BGC); outputting a second representation of the at least one first genome that identifies a set of genes belonging to the BGC based on detecting the pattern of predicted protein domains corresponding to the BGC in the first representation of the at least one first genome; A computer-implemented method comprising: 2. The computer-implemented method of embodiment 1, wherein the first representation of the at least one first genome comprises a nucleotide sequence for the at least one first genome. 3. The computer-implemented method of embodiment 1 or embodiment 2, wherein the first representation of the at least one first genome comprises a vector representation of, or an embedding of, the at least one first genome. 4. The computer-implemented method of any one of embodiments 1 to 3, wherein the first representation of the at least one first genome comprises a sequence of a CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representation of protein domains encoded by genes in the at least one first genome, or any combination thereof. 5. The computer-implemented method of embodiment 4, wherein the protein domain representation sequence is generated by a process comprising searching protein sequences for each gene in the at least one first genome and identifying protein domains in the protein sequences by sequence alignment against a database of CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, or TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representations, respectively. 6. The computer-implemented method of embodiment 5, wherein the first representation of the at least one first genome further comprises associated Gene Ontology (GO) terms, an identification of any resistance genes present, an identification of additional regulatory elements, or an identification of additional epigenetic elements. 7. The computer-implemented method of any one of embodiments 4 to 6, further comprising encoding each protein domain representation in the sequence of the CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, or TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representation of the protein domain as a vector representation of the at least one first genome using a representation learning system. 8. The computer-implemented method of embodiment 7, wherein the representation learning system includes a graphical learning model, a deep autoencoder, Pfam2vec, word2vec, GloVe, or fastText. 9. The computer-implemented method of any one of embodiments 1 to 3, wherein the first representation of the at least one first genome comprises an annotation sequence of genes in the at least one first genome based on a gene function and pathway mapping database. 10. The computer-implemented method of embodiment 9, wherein the pathway mapping database is KEGG. 11. The computer-implemented method of any one of embodiments 1 to 3, wherein the first representation of the at least one first genome comprises an annotation sequence of genes in the at least one first genome based on a database containing data on Clusters of Orthologous Groups (COGs). 12. The computer-implemented method of embodiment 11, wherein the database containing data on Clusters of Orthologous Groups (COGs) is EggNOG. 13. The computer-implemented method of any one of embodiments 9 to 12, further comprising encoding the sequence of the gene annotation as a vector representation of the at least one first genome using a representation learning system. 14. The computer-implemented method of embodiment 13, wherein the representation learning system includes a graphical learning model, a deep autoencoder, Pfam2vec, word2vec, GloVe, or fastText. 15. The computer-implemented method of any one of embodiments 1 to 14, wherein the trained machine learning model comprises a deep learning model. 16. The computer-implemented method of embodiment 15, wherein the deep learning model includes a supervised learning model or an unsupervised learning model. 17. The computer-implemented method of embodiment 15 or embodiment 16, wherein the deep learning model comprises a convolutional neural network, a long short-term memory network, or a transformer model. 18. The computer-implemented method of embodiment 15 or embodiment 16, wherein the deep learning model includes a combination of components from a neural network, a convolutional neural network, a long short-term memory network, or a transformer neural network. 19. A computer-implemented method according to any one of the preceding claims, wherein the machine learning model is trained using a training dataset comprising data for multiple training genomes. 20. The computer-implemented method of embodiment 19, wherein the plurality of training genomes comprises a plurality of synthetic training genomes. 21. The computer-implemented method of embodiment 20, wherein one or more of the plurality of synthetic training genomes each comprises a set of gene sequences from an actual BGC randomly inserted into a BGC-negative genome. 22. A computer-implemented method according to embodiment 20 or embodiment 21, wherein one or more of the plurality of synthetic training genomes each comprises a set of gene sequences from a combination of actual positive BGC examples and synthetic negative BGC examples. 23. The computer-implemented method of any one of embodiments 1 to 22, wherein the second representation of the at least one first genome comprises a vector representation, a graph representation, or a tensor representation of the at least one first genome. 24. The computer-implemented method of any one of the preceding claims, further comprising evaluating the gene identified as belonging to the BGC to determine whether it is a resistance gene. 25. The computer-implemented method of embodiment 24, wherein the resistance gene is an embedded target gene (ETaG) or a non-embedded target gene (NETaG). 26. The computer-implemented method of embodiment 24 or embodiment 25, further comprising performing an in vitro assay to test a secondary metabolite produced by a BGC in at least one first genome to which the resistance gene belongs for activity against a resistance gene homolog or a protein encoded thereby identified in a second genome different from the at least one first genome. 27. The computer-implemented method of any one of embodiments 24 to 26, further comprising performing an in vivo assay to test a secondary metabolite produced by a BGC in at least one first genome to which the resistance gene is identified for activity against a resistance gene homolog or a protein encoded thereby identified in a second genome different from the at least one first genome. 28. The computer-implemented method of embodiment 26 or embodiment 27, wherein the second genome comprises a mammalian genome, a human genome, an avian genome, a reptile genome, an amphibian genome, a plant genome, a fungal genome, a bacterial genome, or a viral genome. 29. The computer-implemented method of any one of embodiments 1 to 28, wherein the at least one first genome comprises a eukaryotic genome or a prokaryotic genome. 30. The computer-implemented method of embodiment 29, wherein the at least one first genome is a eukaryotic genome, and the eukaryotic genome comprises a plant genome or a fungal genome. 31. The computer-implemented method of embodiment 29, wherein at least one first genome is a prokaryotic genome, and the prokaryotic genome is a bacterial genome. 32. A computer-implemented method according to any one of the preceding claims, wherein the first representation of at least one first genome is input by a user of a system configured to perform the computer-implemented method. 33. A computer-implemented method comprising: receiving as input a sequence for at least a first genome; generating a first representation of at least one first genome, wherein the first representation of the at least one first genome comprises sequences of protein domain representations encoded by genes in the at least one first genome; encoding each protein domain representation in the array of protein domain representations as a vector representation of at least one first genome using a representation learning system; A computer-implemented method comprising: 34. The computer-implemented method of embodiment 33, wherein the sequence of the protein domain representation encoded by the gene in the at least one first genome comprises a sequence of a CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representation of the protein domain encoded by the gene in the at least one first genome, or any combination thereof. 35. The computer-implemented method of embodiment 34, wherein the protein domain representation sequence is generated by a process comprising searching protein sequences for each gene in at least one first genome and identifying protein domains in the protein sequences by sequence alignment against a database of CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representations, respectively. 36. The computer-implemented method of embodiment 34 or embodiment 35, wherein the first representation of the at least one first genome further comprises associated Gene Ontology (GO) terms, identification of any resistance genes present, identification of additional regulatory elements, or identification of additional epigenetic elements. 37. A computer-implemented method according to any one of embodiments 33 to 36, wherein the representation learning system comprises a graphical learning model, a deep autoencoder, Pfam2vec, word2vec, GloVe, or fastText. 38. A computer-implemented method according to any one of embodiments 33 to 37, wherein the representation learning system is trained on a corpus of annotated genomes, each of which comprises an array of CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representations, or any combination thereof, of protein domains encoded by genes in the genomes of the corpus. 39. A system comprising: one or more processors; A memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform the method according to any one of embodiments 1 to 38. A system comprising: 40. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a system, cause the system to perform a method according to any one of embodiments 1 to 38.
[0158] From the above, it should be understood that although specific implementations of the disclosed method and system have been illustrated and described, various modifications can be made thereto and are contemplated herein. It is also not intended that the present invention be limited by the specific examples provided herein. Although the present invention has been described with reference to the above specification, the description and illustration of the preferred embodiments herein are not meant to be construed in a limiting sense. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations or relative proportions set forth herein, which depend upon various conditions and variables. Various changes in form and details of the embodiments of the present invention will be apparent to those skilled in the art. It is therefore contemplated that the present invention encompasses such modifications, variations and equivalents.
Claims
1. A computer-implemented method for identifying biosynthetic gene clusters, comprising: receiving as input a first representation of at least one first genome; processing the representation of the at least one first genome using a trained machine learning model configured to detect patterns of predicted protein domains encoded by genes belonging to biosynthetic gene clusters (BGCs); and outputting a second representation of the at least one first genome that identifies a set of genes belonging to the BGC based on detecting a pattern of predicted protein domains corresponding to the BGC in the first representation of the at least one first genome.
2. 10. The computer-implemented method of claim 1, wherein the first representation of at least one first genome comprises a nucleotide sequence for the at least one first genome.
3. 3. The computer-implemented method of claim 1 or claim 2, wherein the first representation of the at least one first genome comprises a vector representation of the at least one first genome, or an embedding thereof.
4. 2. The computer-implemented method of claim 1, wherein the first representation of at least one first genome comprises a sequence of a CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representation, or any combination thereof, of protein domains encoded by genes in the at least one first genome.
5. 5. The computer-implemented method of claim 4, wherein the protein domain representation sequences are generated by a process comprising searching protein sequences for each gene in at least one first genome and identifying protein domains in the protein sequences by sequence alignment against a database of CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representations, respectively.
6. 6. The computer-implemented method of claim 5, wherein the first representation of the at least one first genome further comprises associated Gene Ontology (GO) terms, an identification of any resistance genes present, an identification of additional regulatory elements, or an identification of additional epigenetic elements.
7. 7. The computer-implemented method of claim 4, further comprising encoding each protein domain representation in the sequence of the CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representation of the protein domain as a vector representation of at least one first genome using a representation learning system.
8. 8. The computer-implemented method of claim 7, wherein the representation learning system comprises a graphical learning model, a deep autoencoder, Pfam2vec, word2vec, GloVe, or fastText.
9. 2. The computer-implemented method of claim 1, wherein the first representation of at least one first genome comprises an array of annotations of genes in the at least one first genome based on a gene function and pathway mapping database.
10. 10. The computer-implemented method of claim 9, wherein the pathway mapping database is KEGG.
11. 2. The computer-implemented method of claim 1, wherein the first representation of at least one first genome comprises an array of annotations of genes in the at least one first genome based on a database containing data on clusters of orthologous groups (COGs).
12. 12. The computer-implemented method of claim 11, wherein the database containing data on Clusters of Orthologous Groups (COGs) is EggNOG.
13. 13. The computer-implemented method of any one of claims 9 to 12, further comprising encoding sequences of gene annotations as vector representations of the at least one first genome using a representation learning system.
14. 14. The computer-implemented method of claim 13, wherein the representation learning system comprises a graphical learning model, a deep autoencoder, Pfam2vec, word2vec, GloVe, or fastText.
15. The computer-implemented method of claim 1 , wherein the trained machine learning model comprises a deep learning model.
16. 16. The computer-implemented method of claim 15, wherein the deep learning model comprises a supervised learning model or an unsupervised learning model.
17. 17. The computer-implemented method of claim 15 or claim 16, wherein the deep learning model comprises a convolutional neural network, a long short-term memory network, or a transformer model.
18. 17. The computer-implemented method of claim 15 or claim 16, wherein the deep learning model comprises a combination of components from a neural network, a convolutional neural network, a long short-term memory network, or a transformer neural network.
19. 10. The computer-implemented method of claim 1, wherein the machine learning model is trained using a training dataset that includes data for multiple training genomes.
20. 20. The computer-implemented method of claim 19, wherein the plurality of training genomes comprises a plurality of synthetic training genomes.
21. 21. The computer-implemented method of claim 20, wherein one or more synthetic training genomes of the plurality of synthetic training genomes each comprise a set of gene sequences from an actual BGC randomly inserted into a BGC-negative genome.
22. 22. The computer-implemented method of claim 20 or claim 21, wherein one or more of the plurality of synthetic training genomes each comprises a set of gene sequences from a combination of actual positive BGC examples and synthetic negative BGC examples.
23. 10. The computer-implemented method of claim 1, wherein the second representation of the at least one first genome comprises a vector representation, a graph representation, or a tensor representation of the at least one first genome.
24. 10. The computer-implemented method of claim 1, further comprising evaluating a gene identified as belonging to a BGC to determine whether it is a resistance gene.
25. 25. The computer-implemented method of claim 24, wherein the resistance gene is an embedded target gene (ETaG) or a non-embedded target gene (NETaG).
26. 26. The computer-implemented method of claim 24 or claim 25, further comprising performing an in vitro assay to test secondary metabolites produced by a BGC in at least one first genome identified to which the resistance gene belongs for activity against a resistance gene homolog or a protein encoded thereby identified in a second genome different from the at least one first genome.
27. 25. The computer-implemented method of claim 24, further comprising performing an in vivo assay to test secondary metabolites produced by a BGC in at least one first genome identified to which the resistance gene belongs for activity against a resistance gene homolog or a protein encoded thereby identified in a second genome different from the at least one first genome.
28. 27. The computer-implemented method of claim 26, wherein the second genome comprises a mammalian genome, a human genome, an avian genome, a reptile genome, an amphibian genome, a plant genome, a fungal genome, a bacterial genome, or a viral genome.
29. The computer-implemented method of claim 1 , wherein the at least one first genome comprises a eukaryotic genome or a prokaryotic genome.
30. 30. The computer-implemented method of claim 29, wherein the at least one first genome is a eukaryotic genome, wherein the eukaryotic genome comprises a plant genome or a fungal genome.
31. 30. The computer-implemented method of claim 29, wherein the at least one first genome is a prokaryotic genome, and wherein the prokaryotic genome is a bacterial genome.
32. 10. The computer-implemented method of claim 1, wherein the first representation of the at least one first genome is input by a user of a system configured to perform the computer-implemented method.
33. A computer-implemented method comprising: receiving as input sequences for at least one first genome; generating a first representation of the at least one first genome, wherein the first representation of the at least one first genome comprises sequences of protein domain representations encoded by genes in the at least one first genome; encoding each protein domain representation in the array of protein domain representations as a vector representation of said at least one first genome using a representation learning system.
34. 34. The computer-implemented method of claim 33, wherein the sequence of the protein domain representation encoded by the gene in the at least one first genome comprises the sequence of a CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representation, or any combination thereof, of the protein domain encoded by the gene in the at least one first genome.
35. 35. The computer-implemented method of claim 34, wherein the protein domain representation sequences are generated by a process comprising searching protein sequences for each gene in the at least one first genome and identifying protein domains in the protein sequences by sequence alignment against a database of CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representations, respectively.
36. 36. The computer-implemented method of claim 34 or claim 35, wherein the first representation of the at least one first genome further comprises associated Gene Ontology (GO) terms, an identification of any resistance genes present, an identification of additional regulatory elements, or an identification of additional epigenetic elements.
37. 34. The computer-implemented method of claim 33, wherein the representation learning system comprises a graphical learning model, a deep autoencoder, Pfam2vec, word2vec, GloVe, or fastText.
38. 34. The computer-implemented method of claim 33, wherein the representation learning system is trained on a corpus of annotated genomes, each of the genomes comprising sequences of CDD, Gene3D, PANTHER, Pfam, ProSitePatterns, ProSiteProfiles, SUPERFAMILY, SMART, TIGRFAM, SFLD, Hamap, Coils, PRINTS, PIRSR, AntiFam, MobiDBLite, or PIRSF representations, or any combination thereof, of protein domains encoded by genes in the genomes of the corpus.
39. 1. A system comprising: one or more processors; a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform the method of claim 1 or 33; A system comprising:
40. 34. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions that, when executed by one or more processors of a system, cause the system to perform the method of claim 1 or 33.