Systems and methods for ranking molecules
A machine learning model integrates heterogeneous assay data to rank compounds for antibacterial activity and cytotoxicity, improving the efficiency and effectiveness of high-throughput screening for antibiotics by identifying compounds with high efficacy and low toxicity.
Patent Information
- Application Number
- PCT/US2025/042043
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-16
- Filing Date
- 2025-08-14
- Publication Date
- 2026-02-19
AI Technical Summary
Traditional drug discovery methods struggle to effectively combat antibiotic-resistant pathogens, particularly in Gram-negative bacteria, due to challenges in bacterial permeability and efflux, leading to low hit rates and inefficiencies in high-throughput screening.
A machine learning model is trained using a diverse dataset from multiple assays with varying conditions and endpoints to rank compounds based on antibacterial activity and cytotoxicity, leveraging a 'Learn2Rank' algorithm to integrate heterogeneous data and predict the efficacy of compounds, thereby guiding high-throughput screening.
The model enhances the hit rate of effective compounds by efficiently processing larger datasets and predicting compounds with high antibacterial activity and low cytotoxicity, overcoming the limitations of traditional screening methods.
Smart Images

Figure US2025042043_19022026_PF_FP_ABST
Abstract
Description
Docket No.: 60134-711.601SYSTEMS AND METHODS FOR RANKING MOLECULESCROSS-REFERENCE
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 684,276, filed August 16, 2024, which application is incorporated herein by reference in its entirety.BACKGROUND
[0002] The emergence, persistence, and propagation of antibiotic resistant pathogens provide globally-relevant epidemiological problems, which demand continuous development of new antibiotics. An estimated 1.27 million deaths world-wide have been caused by antibiotic resistant pathogens and traditional approaches to drug discovery have yet to deliver a widely- effective antibacterial, owing to challenges in attacking bacteria via bacterial permeability and efflux. This issue is especially prevalent in Gram-negative bacteria.
[0003] Advances in combinatorial chemical synthesis, and information technology revolution, may, in some ways, accelerate the development of drug candidate via high-throughput screening (HTS) for desirable properties. While the breadth of the chemical space that can be explored to find new drug candidates have expanded considerably, the exponential growth of this search space demands efficient strategies for screening compounds in the chemical search space.SUMMARY
[0004] In some aspects, the present disclosure provides a computer-implemented method of training a machine learning model for ranking compounds for efficacy, comprising: (a) providing a training dataset comprising data generated from assays, wherein at least two assays in the assays yield different measures of efficacy of the compounds; and (b) training the machine learning model using the training dataset, wherein the machine learning model yields a ranking of compounds based on the different measures of efficacy. In some aspects, the present disclosure provides a computer-implemented method of ranking compounds for efficacy, comprising ranking compounds based on their measures of efficacy using a model trained using a dataset comprising data generated from assays comprising at least two assays, wherein the at least two assays yield different measures of efficacy of the compounds. In some embodiments, the measures of efficacy comprise antibacterial activity, cytotoxicity, or both.
[0005] In some embodiments, the dataset comprises at least three, four, or five assays, which yield the different measures of efficacy. In some embodiments, the dataset comprises at least 5, 10, 15, 20, 25, 50, 75, or 100 assays, which yield different measures of efficacy.Docket No.: 60134-711.601
[0006] In some embodiments, the assays comprise antimicrobial susceptibility assays. In some embodiments, the assays are selected from the group consisting of agar well diffusion, disk well diffusion, agar dilution, broth dilution, ATP bioluminescence, thin-layer chromatography, antimicrobial susceptibility testing, bioautography, ATP luminescence, a fluorescence assay, and any combination thereof.
[0007] In some embodiments, the at least two assays are performed using different conditions. In some embodiments, the different conditions are selected from the group consisting of temperatures, buffers, media, oxygen concentration, and any combination thereof.
[0008] In some embodiments, the assays comprise at least two assays performed on different species of bacteria. In some embodiments, the assays comprise at least two assays performed on different strains of bacteria. In some embodiments, the different strains of bacteria comprise a wild-type strain. In some embodiments, the different strains of bacteria comprise a mutant-type strain. In some embodiments, the different strains of bacteria comprise an efflux deficient strain. In some embodiments, the different strains of bacteria comprise an efflux competent strain. In some embodiments, the different strains of bacteria comprise a membrane compromised strain. In some embodiments, the different strains of bacteria comprise a membrane intact strain. In some embodiments, the different species of bacteria are selected from the following genera: Acinetobacter, Bacteroides, Citrobacter, Clostridioides, Corynebacterium, Enterobacter, Enterococcus, Escherichia, Eggerthella, Klebsiella, Lactobacillus, Listeria, Moraxella, Morganella, Neisseria, Peptostreptococcus, Prevotella, Propionibacterium, Providencia, Pseudomonas, Salmonella, Shigella, Staphylococcus, Stenotrophomonas, Streptococcus, Mycobacteroides, and Mycobacterium. In some embodiments, the different strains of bacteria are selected from the following genera: Acinetobacter, Bacteroides, Citrobacter, Clostridioides, Corynebacterium, Enterobacter, Enterococcus, Escherichia, Eggerthella, Klebsiella, Lactobacillus, Listeria, Moraxella, Morganella, Neisseria, Peptostreptococcus, Prevotella, Propionibacterium, Providencia, Pseudomonas, Salmonella, Shigella, Staphylococcus, Stenotrophomonas, Streptococcus, Mycobacteroides, and Mycobacterium. In some embodiments, the different species of bacteria are selected from the following species: Acinetobacter baumannii, Acinetobacter haemolyticus, Acinetobacter nosocomialis, Acinetobacter pittii, Acinetobacter radioresistens, Acinetobacter ursingii, Bacteroides fragilis, Citrobacter braakii, Citrobacter freundii, Citrobacter koseri, Citrobacter sedlakii, Clostridioides difficile, Clostridioides perfringens, Corynebacterium jeikeium, Enterobacter aerogenes, Enterobacter cloacae, Enterococcus faecalis, Enterococcus faecium, Eggerthella lenta, Escherichia coli, Klebsiella oxytoca, Klebsiella pneumoniae, Lactobacillus acidophilus, Listeria monocytogenes, Moraxella catarrhalis, Morganella morganii, Neisseria gonorrhoeae,Docket No.: 60134-711.601Peptostreptococcus anaerobius, Prevotella bivia, Propionib acterium acnes, Providencia rettgeri, Providencia stuartii, Pseudomonas aeruginosa, Salmonella typhimurium, Shigella dysenteriae, Staphylococcus aureus, Staphylococcus capitis, Staphylococcus caprae, Staphylococcus epidermidis, Staphylococcus haemolyticus, Staphylococcus hominis, Staphylococcus lugdenensis, Staphylococcus pettenkoferi, Staphylococcus saprophyticus, Staphylococcus simulans, Staphylococcus warnerii, Stenotrophomonas maltophilia, Streptococcus agalactiae, Streptococcus constellatus, Streptococcus pneumoniae, Streptococcus pyogenes, Mycobacterium tuberculosis, Mycobacterium avium, Mycobacterium intracellulare, Mycobacterium smegmatis, and Mycobacteroides abscessus. In some embodiments, the different strains of bacteria are selected from the following species: Acinetobacter baumannii, Acinetobacter haemolyticus, Acinetobacter nosocomialis, Acinetobacter pittii, Acinetobacter radioresistens, Acinetobacter ursingii, Bacteroides fragilis, Citrobacter braakii, Citrobacter freundii, Citrobacter koseri, Citrobacter sedlakii, Clostridioides difficile, Clostridioides perfringens, Corynebacterium jeikeium, Enterobacter aerogenes, Enterobacter cloacae, Enterococcus faecalis, Enterococcus faecium, Eggerthella lenta, Escherichia coli, Klebsiella oxytoca, Klebsiella pneumoniae, Lactobacillus acidophilus, Listeria monocytogenes, Moraxella catarrhalis, Morganella morganii, Neisseria gonorrhoeae, Peptostreptococcus anaerobius, Prevotella bivia, Propionibacterium acnes, Providencia rettgeri, Providencia stuartii, Pseudomonas aeruginosa, Salmonella typhimurium, Shigella dysenteriae, Staphylococcus aureus, Staphylococcus capitis, Staphylococcus caprae, Staphylococcus epidermidis, Staphylococcus haemolyticus, Staphylococcus hominis, Staphylococcus lugdenensis, Staphylococcus pettenkoferi, Staphylococcus saprophyticus, Staphylococcus simulans, Staphylococcus warnerii, Stenotrophomonas maltophilia, Streptococcus agalactiae, Streptococcus constellatus, Streptococcus pneumoniae, Streptococcus pyogenes, Mycobacterium tuberculosis, Mycobacterium avium, Mycobacterium intracellulare, Mycobacterium smegmatis, and Mycobacteroides abscessus.
[0009] In some embodiments, the assays are performed on different cell lines. In some embodiments, the different cell lines are selected from the group consisting of HepG2, HEK293, NIH3T3, U2OS, HaCat, CRL-7250, or any combination thereof.
[0010] In some embodiments, the assays are performed with different endpoints. In some embodiments, the different endpoints are selected from the group consisting of growth rate, total growth, optical density, colony forming units, relative light units. In some embodiments, the training dataset comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 75, or 100 sources of data that are heterogenous.Docket No.: 60134-711.601
[0011] In some embodiments, the compounds comprise at least 2, 5, 10, 25, 50, 75, 1000, 5000, 10000, 50000, or 100000 compounds. In some embodiments, the machine learning model yields a ranking of compounds based on their antibacterial activity. In some embodiments, the machine learning model yields a ranking of compounds based on their cytotoxicity. In some embodiments, the machine learning model yields a ranking of compounds based on their antibacterial activity and cytotoxicity. In some embodiments, the computer-implemented method further comprises selecting a set of candidate compounds based on a threshold criteria. In some embodiments, the threshold criteria is selected from the group consisting of a predetermined number, a predetermined percent of the compounds, an antibacterial activity threshold, a cytotoxicity threshold, and any combination thereof.
[0012] In some embodiments, the computer-implemented method further comprises performing screening assays using the set of candidate compounds. In some embodiments, the screening assays are selected from the group consisting of chemistry assays, biochemistry assays, binding assays, functional assays, cell-based assays, in vivo assays, ADME assays, toxicology assays, and any combination thereof.
[0013] In some aspects, the present disclosure provides a computer-implemented method of screening a set of compounds comprising: (a) conducting at least two assays on each compound of the set of compounds, each assay yielding raw data indicating different measures of efficacy; (b) preparing a training dataset based on the raw data; and (c) training the machine learning model using the training dataset, wherein the machine learning model yields a ranking of compounds based on their measures of efficacy. In some embodiments, the measures of efficacy comprise antibacterial activity, cytotoxicity, or both.
[0014] In some embodiments, the training the machine learning model is performed using a neural network, a random forest, or both. In some embodiments, the ranking of compounds is based on their relative antibacterial activity, cytotoxicity, or both.
[0015] In some aspects, the present disclosure provides a computer-implemented system comprising: a digital processing device comprising: at least one processor, an operating system configured to perform executable instructions, a memory, and a computer program including instructions executable by the digital processing device to implement any one of the computer- implemented methods disclosed herein.
[0016] In some aspects, the present disclosure provides a computer program product comprising a computer-readable medium having computer-executable code encoded therein, the computerexecutable code adapted to be executed to implement any one of the computer-implemented methods disclosed herein.Docket No.: 60134-711.601
[0017] In some aspects, the present disclosure provides a non-transitory computer-readable storage media encoded with a computer program including instructions executable by one or more processors to perform any one of the computer-implemented methods disclosed herein.INCORPORATION BY REFERENCE
[0018] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The novel features of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings of which:
[0020] FIG. 1 schematically illustrates training models on datasets that are generated using different assays.
[0021] FIG. 2 schematically illustrates generation of input features.
[0022] FIG. 3 shows plots evaluating the predictive accuracy of the antibacterial activity ranking models of the present disclosure.
[0023] FIG. 4 shows plots evaluating the early enrichment of the models of the present disclosure.
[0024] FIG. 5 shows plots evaluating the predictive accuracy of the cytotoxicity ranking models of the present disclosure.
[0025] FIG. 6 shows a heat map of compounds that are ranked based on predicted antibacterial activity and predicted cytotoxicity.
[0026] FIG. 7 shows the hit rate and the score of different compound libraries.
[0027] FIG. 8 shows the distribution of measured inhibition for compounds found through cherry-picked led discovery.
[0028] FIG. 9 shows a distribution of cytotoxicity of selected and non-selected hits.
[0029] FIG. 10 shows a t-map projection of the screened chemical space along with the hits, clinical antibiotics, and antibacterials.Docket No.: 60134-711.601
[0030] FIG. 11 shows the sizes of various datasets, and illustrates that models of the present disclosure can identify a short-list of most promising hits.
[0031] FIG. 12 shows the prediction accuracy of two reference compound screening models.
[0032] FIG. 13 shows various datasets available for training the models.
[0033] FIG. 14 shows a computer system.
[0034] FIG. 15 shows examples of compounds that were found using a model of the present disclosure.DETAILED DESCRIPTION
[0035] The emergence, persistence, and propagation of antibiotic resistant pathogens provide globally-relevant epidemiological problems which demand continuous development of new antibiotics. An estimated 1.27 million deaths world-wide have been caused by antibiotic resistant pathogens. Traditional approaches to drug discovery have yet to deliver a widely- effective antibacterial, owing to challenges in attacking bacteria via bacterial permeability and efflux (especially in Gram-negative bacteria).
[0036] Advances in combinatorial chemical synthesis, and information technology revolution, may, in some ways, accelerate the development of drug candidate via high-throughput screening (HTS) for desirable properties. While the breadth of the chemical space that can be explored to find new drug candidates have expanded considerably, the exponential growth of this search space also demands efficient strategies for screening compounds in the chemical search space.
[0037] Whole cell screening for antibacterial compounds is one way to identify compounds that avoid efflux and permeability issues, which may be reasons that compounds fail to show antibacterial efficacy against resistant strains. However, hit rates using solely this method can be low - in an experiment, 400k compounds were screened to create a short-list of about 1.5k confirmed hits. These hits were screened using experiments with a membrane compromised strain (thereby having enhanced permeability and being efflux-competent against the strain).
[0038] Various predictive models based on machine learning (ML) have been used for guiding HTS to identify relevant compounds of clinical relevance. With recent developments in open- source HTS data, building predictive ML models have become more accessible. Nevertheless, open-source data are often heterogeneous not only in their data structure, but at an even more basic level, they can differ in terms of assay protocols, reported metrics, etc. These highly heterogenous sets of data may all elucidate the utility of a compound as an antibacterial, and yet, these sets of data have not yet been used synergistically to support predictions based on the combination of the information that is salient in each data set. In some aspects, the presentDocket No.: 60134-711.601 disclosure provides systems and methods for using these heterogenous datasets in a machine learning model or system to generate predictions of useful compounds. The machine learning model can comprise a learning to rank (which may be referred to as “Learn2Rank” or “L2R”) algorithm or model therein. In some embodiments, the machine learning algorithm is used to ingest heterogeneous datasets for informing and guiding HTS.
[0039] When predicting relevant properties on previously unseen data (e.g., a compound), the models of the present disclosure perform better than other ML models in certain tasks and domains. As shown herein, using the models of the present disclosure in HTS for antibiotics increases the hit rate of useful compounds. Moreover, the models of the present disclosure can be more computationally efficient than other methods. Thus, the models of the present disclosure can be more lightweight and can screen through larger compound datasets more efficiently.
[0040] The models of the present disclosure can be used to generalize to new chemical compounds that are different from existing chemical compounds (e.g., antibiotics). By focusing on “progressible” hits, e.g., those compounds that indicate high activity (e.g., antibacterial activity) and minimum undesired effects (e.g., cytotoxicity towards a human cell), and avoiding compounds that are similar to known chemical compounds (e.g., antibiotics) or those known to have poor medicinal chemistries, the models of the present disclosure can suggest hits that are more likely to be effective chemical compounds (e.g., antibiotics).
[0041] In some aspects, the present disclosure provides a computer-implemented method of training a machine learning (ML) model for ranking compounds for efficacy. The computer- implemented method can comprise providing a training dataset comprising data generated from a plurality of assays. At least two assays of the plurality of assays can yield or report different measures of efficacy of the compounds. For example, a first assay of the at least two assays can measure a first antibacterial activity of a compound, whereas a second assay of the at least two assays can measure a second antibacterial activity of a compound. The computer-implemented method can comprise training the machine learning model using the training dataset. The machine learning model can yield a ranking of compounds based on the different measures of efficacy. In some aspects, the present disclosure provides a computer-implemented method of ranking compounds for efficacy. The computer-implemented method can comprise ranking compounds based on their measures of efficacy using a model. The model can be trained using a dataset comprising data generated from assays comprising at least two assays. The at least two assays can yield or report different measures of efficacy of the compounds. In some embodiments, the measures of efficacy comprise antimicrobial activity (e.g., antibacterial activity), cytotoxicity (e.g., to a non-bacterial cell (e.g., to the subject)), or both. In someDocket No.: 60134-711.601 embodiments, the measure of efficacy is a predictive measure of efficacy. In some embodiments, the ranking of compounds is based on their relative antimicrobial activity (e.g., antibacterial activity), cytotoxicity, or both. In some embodiments, the measure of efficacy comprises a measure of inhibiting the growth of a cell, killing a cell, or a combination thereof. In some embodiments, the cell is a bacterial cell. In some embodiments, the bacterial cell is a gram-negative bacteria. In some embodiments, the bacterial cell is a gram-positive bacteria.
[0042] In some embodiments, the measure of efficacy comprises a measure of antimicrobial activity of the compound. In some embodiments, the measure of antimicrobial activity of the compound comprises an antibacterial activity of the compound. A compound’s antibacterial activity may refer to a compound’s ability to inhibit bacterial growth, kill bacteria, or a combination thereof.
[0043] In some embodiments, the measure of efficacy comprises a measure of cytotoxicity. Cytotoxicity can refer to the toxicity of the compound on a non-target (e.g., non-bacterial) cell. For example, cytotoxicity can refer to the toxicity to mammalian cells (e.g., human cells). Cytotoxicity to non-bacterial cells may be measured through the use of cell lines, such as HepG2, HEK293, NIH3T3, U2OS, HaCat, CRL-7250, or any combination thereof.
[0044] In some aspects, the present disclosure provides a computer-implemented method of screening a compounds (e.g., a set of compounds). The computer-implemented method can comprise conducting at least two assays on each compound of the set of compounds. Each assay can yield or report raw data indicating different measures of efficacy. The computer- implemented method can comprise preparing a training dataset based on the raw data. The computer-implemented method can comprise training the machine learning model using the training dataset. The machine learning model can yield or report a ranking of compounds based on their measures of efficacy. In some embodiments, the measures of efficacy comprise antibacterial activity, cytotoxicity, or both. In some embodiments, the ranking of compounds is based on their relative antibacterial activity, cytotoxicity, or both. Various other measures of efficacy for antibacterials can be used.
[0045] Multitask Ranking Model Classification models can be used to predict a scalar value per virtually screened compound. The scalar can be proportional to the probability of antibiotic activity. When training classification models, the target value to be predicted can be set as a binary number that represents antibiotic activity (e.g., 0 - inactive, 1 - active). A measure of compound activity, however, can be dependent on the assay protocol(s) that were used to obtain the training data. For this reason, training data from multiple, heterogeneous sources cannot typically easily be combined due to contradicting or dissimilar reported activities.Docket No.: 60134-711.601
[0046] Systems and methods of the present disclosure can be used to predict a relative likelihood of efficacy. For example, given a pair of compounds, the relative likelihood of efficacy can provide the likelihood of a first compound being more likely to have activity (e.g., antibiotic activity) than a second compound. This likelihood can be a measure of efficacy that is largely invariant to the assay protocol. Thus, data from multiple, heterogeneous sources can be used to inform a single model. This can lead to a model that is informed by a wider compound space from multiple datasets. As such, the model can perform better across a large search space or latent space of compounds.
[0047] In some embodiments, the relative likelihood can provide a relative measure of efficacy (e.g., antimicrobial efficacy, cytotoxicity, or a combination thereo) between at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, Ik, 2k, 3k, 4k, 5k, 6k, 7k, 8k, 9k, 10k, 20k, 30k, 40k, 50k, 60k, 70k, 80k, 90k, 100k, 200k, 300k, 400k, 500k, 600k, 700k, 800k, 900k, IM, 1.2M, 1.5 M, 1.75 M, or 2M compounds. In some embodiments, the relative likelihood can provide a relative measure of efficacy between at most 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, Ik, 2k, 3k, 4k, 5k, 6k, 7k, 8k, 9k, 10k, 20k, 30k, 40k, 50k, 60k, 70k, 80k, 90k, 100k, 200k, 300k, 400k, 500k, 600k, 700k, 800k, 900k, IM, 1.2M, 1.5 M, 1.75 M, or 2M compounds. The relative likelihood of the compounds can be generated simultaneously, e.g., the machine learning model can process an input data set containing that many compounds and output the entire ranking at once. The relative likelihood of the compounds can be generated piece-meal or step-wise, e.g., the ranking can be generated sorting algorithms, e.g., quicksort, mergesort, heapsort, counting sort, insertion sort, etc.
[0048] The relative measure of efficacy can be a probability. The relative measure of efficacy can be a scalar that represents a probability, e.g., a logit. A probability can indicate a measure such as, for example, compound “A” is X% likely to be more effective than compound “B”. The scalar representation of “A” can be a higher value (or lower, depending on the sign) to indicate that it is better than “B” on a certain metric.
[0049] A machine learning model can be used learn the relative measure of efficacy in the latent space. When a machine learning model, such as a neural network, is trained on one or more datasets of compounds and efficacies, its latent space may be organized such that the latent representations of the compounds may become ordered as a function of the efficacy. Thus, a machine learning model can be used to predict the efficacy of compounds that are not in the training datasets. The latent space can be explored in order to find regions in the chemical space which are novel and more likely to be efficacious.Docket No.: 60134-711.601Assays
[0050] Different types of assays are often carried out in different conditions, against different strains, and with different endpoints. Even within the same type of assay, variance in the environmental conditions, the experimenter, and genetic drift in the bacteria can also cause variances in the assay results. As such, there is a need to effectively and efficiently compare the potentially varying results of the assays. As FIG. 1 schematically illustrates, the models of the present disclosure can be used to train on datasets that are generated using different assays (e.g., including same assays prepared in different manners). The models can then rank compounds within each assay (e.g., pairs of compounds). The models do not require that the compound’s activity is comparable across assays or that compounds are shared between assays.
[0051] In some embodiments, the dataset comprises a plurality of assays. A first assay of the plurality of assays yielding a different measure of efficacy than a second assay of the plurality of assays. In some embodiments, the dataset comprises at least three, four, or five assays, which yield a different measure of efficacy (e.g., antimicrobial efficacy). In some embodiments, the dataset comprises at least 5, 10, 15, 20, 25, 50, 75, or 100 assays (e.g., which yield different measures of efficacy). In some embodiments, the assays comprise an antimicrobial susceptibility assay (e.g., one antimicrobial susceptibility assay, two antimicrobial susceptibility assays, three antimicrobial susceptibility assays, etc.). In some embodiments, the assays are selected from the group consisting of agar well diffusion, disk well diffusion, agar dilution, broth dilution, ATP bioluminescence, thin-layer chromatography, antimicrobial susceptibility testing, bioautography, ATP luminescence, a fluorescence assay, and any combination thereof. In some embodiments, at least two assays are performed using different conditions. In some embodiments, the different conditions are selected from the group consisting of temperatures, buffers, media, oxygen concentration, and any combination thereof.
[0052] In some embodiments, the assays comprise at least two assays performed on different species of bacteria. In some embodiments, the assays comprise at least two assays performed on different species of bacteria. In some embodiments, the assays comprise at least two assays performed on different strains of bacteria. In some embodiments, the different strains of bacteria comprise a wild-type strain. In some embodiments, the different strains of bacteria comprise a mutant-type strain. In some embodiments, the different strains of bacteria comprise an efflux deficient strain. In some embodiments, the different strains of bacteria comprise an efflux competent strain. In some embodiments, the different strains of bacteria comprise a membrane compromised strain. In some embodiments, the different strains of bacteria comprise a membrane intact strain. In some embodiments, the bacteria comprise a gram-positive bacteria. In some embodiments, the bacteria comprise a gram-negative bacteria. In some embodiments, theDocket No.: 60134-711.601 different species of bacteria are selected from the following genera: Acinetobacter, Bacteroides, Citrobacter, Clostridioides, Corynebacterium, Enterobacter, Enterococcus, Escherichia, Eggerthella, Klebsiella, Lactobacillus, Listeria, Moraxella, Morganella, Neisseria, Peptostreptococcus, Prevotella, Propionibacterium, Providencia, Pseudomonas, Salmonella, Shigella, Staphylococcus, Stenotrophomonas, Streptococcus, Mycobacteroides, and Mycobacterium. In some embodiments, the different strains of bacteria are selected from the following genera: Acinetobacter, Bacteroides, Citrobacter, Clostridioides, Corynebacterium, Enterobacter, Enterococcus, Escherichia, Eggerthella, Klebsiella, Lactobacillus, Listeria, Moraxella, Morganella, Neisseria, Peptostreptococcus, Prevotella, Propionibacterium, Providencia, Pseudomonas, Salmonella, Shigella, Staphylococcus, Stenotrophomonas, Streptococcus, Mycobacteroides, and Mycobacterium. In some embodiments, the different species of bacteria are selected from the following species: Acinetobacter baumannii, Acinetobacter haemolyticus, Acinetobacter nosocomialis, Acinetobacter pittii, Acinetobacter radioresistens, Acinetobacter ursingii, Bacteroides fragilis, Citrobacter braakii, Citrobacter freundii, Citrobacter koseri, Citrobacter sedlakii, Clostridioides difficile, Clostridioides perfringens, Corynebacterium jeikeium, Enterobacter aerogenes, Enterobacter cloacae, Enterococcus faecalis, Enterococcus faecium, Eggerthella lenta, Escherichia coli, Klebsiella oxytoca, Klebsiella pneumoniae, Lactobacillus acidophilus, Listeria monocytogenes, Moraxella catarrhalis, Morganella morganii, Neisseria gonorrhoeae, Peptostreptococcus anaerobius, Prevotella bivia, Propionibacterium acnes, Providencia rettgeri, Providencia stuartii, Pseudomonas aeruginosa, Salmonella typhimurium, Shigella dysenteriae, Staphylococcus aureus, Staphylococcus capitis, Staphylococcus caprae, Staphylococcus epidermidis, Staphylococcus haemolyticus, Staphylococcus hominis, Staphylococcus lugdenensis, Staphylococcus pettenkoferi, Staphylococcus saprophyticus, Staphylococcus simulans, Staphylococcus wamerii, Stenotrophomonas maltophilia, Streptococcus agalactiae, Streptococcus constellatus, Streptococcus pneumoniae, Streptococcus pyogenes, Mycobacterium tuberculosis, Mycobacterium avium, Mycobacterium intracellulare, Mycobacterium smegmatis, and Mycobacteroides abscessus. In some embodiments, the different strains of bacteria are selected from the following species: Acinetobacter baumannii, Acinetobacter haemolyticus, Acinetobacter nosocomialis, Acinetobacter pittii, Acinetobacter radioresistens, Acinetobacter ursingii, Bacteroides fragilis, Citrobacter braakii, Citrobacter freundii, Citrobacter koseri, Citrobacter sedlakii, Clostridioides difficile, Clostridioides perfringens, Corynebacterium jeikeium, Enterobacter aerogenes, Enterobacter cloacae, Enterococcus faecalis, Enterococcus faecium, Eggerthella lenta, Escherichia coli, Klebsiella oxytoca, Klebsiella pneumoniae, Lactobacillus acidophilus, Listeria monocytogenes, Moraxella catarrhalis, Morganella morganii,Docket No.: 60134-711.601Neisseria gonorrhoeae, Peptostreptococcus anaerobius, Prevotella bivia, Propionibacterium acnes, Providencia rettgeri, Providencia stuartii, Pseudomonas aeruginosa, Salmonella typhimurium, Shigella dysenteriae, Staphylococcus aureus, Staphylococcus capitis, Staphylococcus caprae, Staphylococcus epidermidis, Staphylococcus haemolyticus, Staphylococcus hominis, Staphylococcus lugdenensis, Staphylococcus pettenkoferi, Staphylococcus saprophyticus, Staphylococcus simulans, Staphylococcus wamerii, Stenotrophomonas maltophilia, Streptococcus agalactiae, Streptococcus constellatus, Streptococcus pneumoniae, Streptococcus pyogenes, Mycobacterium tuberculosis, Mycobacterium avium, Mycobacterium intracellulare, Mycobacterium smegmatis, and Mycobacteroides abscessus.
[0053] In some embodiments, the assays are performed on different cell lines. An assay may be conducted on a non-bacterial cell line to determine the cytotoxic effect(s) of a compound on a non-bacterial cell. Examples of the cell lines include, but not limited to, HepG2, HEK293, NIH3T3, U2OS, HaCat, and CRL-7250. In some embodiments, the different cell lines are selected from the group consisting of HepG2, HEK293, NIH3T3, U2OS, HaCat, CRL-7250, or any combination thereof. Performing cytotoxicity assays on cell lines can provide indirect measures of in vivo cytotoxicity.
[0054] In some embodiments, the assays are performed with different endpoints. In some embodiments, the different endpoints are selected from the group consisting of growth rate, total growth, optical density, colony forming units, relative light units, cell counts, cell density, and any combination thereof. In some embodiments, the training dataset comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 75, or 100 sources of data that are heterogenous. In some embodiments, the training dataset comprises at most 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 75, or 100 sources of data that are heterogenous.
[0055] In some embodiments, the computer-implemented method further comprises performing screening assays using the set of candidate compounds. In some embodiments, the screening assays are selected from the group consisting of chemistry assays, biochemistry assays, binding assays, functional assays, cell-based assays, in vivo assays, ADME assays, toxicology assays, and any combination thereof.
[0056] In some embodiments, a candidate compound can be configured to permeate the outer and / or inner membrane of bacteria. In some embodiments, a candidate compound can be configured to avoid efflux pumps of bacteria. In some embodiments, a candidate compound can be configured to engage with target bacteria.Docket No.: 60134-711.601Input Features
[0057] Various input features can be used. Input features can comprise various cheminformatic features. A cheminformatic feature can comprise a molecular fingerprint. The fingerprint can comprise an Avalon fingerprint, a Morgan fingerprint, a circular fingerprint, a pharmacophore fingerprint, a molecular access system fingerprint (MACCS), an extended-connectivity fingerprint (ECFP4), or any combination thereof. Input features can comprise latent vectors. Latent vectors can be generated using a neural network, e.g., a graph neural network that processes molecular structures or representations. FIG. 2 shows a schematic that illustrates generation of input features. The various input features can be generated from a molecular structure or identifier. The molecular structure or identifier can be used to generate input features, which can be processed by a machine learning model to generate a prediction or ranking.
[0058] A cheminformatic feature can comprise an identifier of an atom, a functional group, or a motif in a molecule. A cheminformatic feature can comprise an electronic configuration, a charge, a size, a bond angle, a dihedral angle, or any combination thereof.
[0059] A cheminformatic feature can comprise an atomistic representation of a molecule. The atomistic representation can comprise relative cartesian coordinates of atoms to each other. In some embodiments, an atomistic representation of a molecule may comprise the relative cartesian coordinates of atoms to an arbitrary point. In some embodiments, an atomistic representation of a molecule may comprise thermodynamic estimations of values such as solvation energy, potential energy of bond lengths, bond angles, dihedral angles, 1-4 intramolecular interaction energies, intramolecular energies among adjacent bond angles, hydrogen bonding energies, and non-bonded interaction energies. In some embodiments, an atomistic representation of a molecule may comprise atom type definitions and generalizations. In some embodiments, an atomistic representation of a molecule may comprise polarizability parameters. In some embodiments, an atomistic representation of a molecule may comprise Lennard-Jones van der Waal parameters. In some embodiments, an atomistic representation of a molecule may comprise electrostatic charge parameters. In some embodiments, an atomistic representation of a molecule may comprise bond length, bond angle, and dihedral force constants. In some embodiments, an atomistic representation of a molecule may comprise bond length, bond angle, and dihedral equilibrium values. In some embodiments, an atomistic representation of a molecule may comprise dihedral phase and periodicity force constants.
[0060] In some embodiments, the compounds comprise at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, Ik, 2k, 3k, 4k, 5k, 6k, 7k, 8k, 9k, 10k, 20k, 30k, 40k, 50k, 60k, 70k, 80k, 90k, 100k, 200k, 300k, 400k, 500k, 600k, 700k,Docket No.: 60134-711.601800k, 900k, IM, 1.2M, 1.5 M, 1.75 M, or 2M compounds. In some embodiments, the compounds comprise at most 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, Ik, 2k, 3k, 4k, 5k, 6k, 7k, 8k, 9k, 10k, 20k, 30k, 40k, 50k, 60k, 70k, 80k, 90k, 100k, 200k, 300k, 400k, 500k, 600k, 700k, 800k, 900k, IM, 1.2M, 1.5 M, 1.75 M, or 2M compounds. In some embodiments, the machine learning model yields a ranking of compounds based on their antibacterial activity. In some embodiments, the machine learning model yields a ranking of compounds based on their cytotoxicity. In some embodiments, the machine learning model yields a ranking of compounds based on their antibacterial activity and cytotoxicity. In some embodiments, the computer-implemented method further comprises selecting a set of candidate compounds based on a threshold criteria. In some embodiments, the threshold criteria is selected from the group consisting of a predetermined number, a predetermined percent of the compounds, an antibacterial activity threshold, a cytotoxicity threshold, anti-inflammatory threshold, and any combination thereof.Machine Learning
[0061] A machine learning model can comprise one or more of various machine learning models. In some embodiments, the machine learning model can comprise one machine learning model. In some embodiments, the machine learning model can comprise a plurality of machine learning models. In some embodiments, the machine learning model can comprise a neural network model. In some embodiments, the machine learning model can comprise a random forest model. In some embodiments, the machine learning model can comprise a manifold learning model. In some embodiments, the machine learning model can comprise a hyperparameter learning model. In some embodiments, the machine learning model can comprise an active learning model.
[0062] A graph, graph model, and graphical model can refer to a method of conceptualizing or organizing information into a representation comprising nodes and edges. In some embodiments, a graph can refer to the principle of conceptualizing or organizing data, wherein the data may be stored in a various and alternative forms such as linked lists, dictionaries, spreadsheets, arrays, in permanent storage, in transient storage, and so on, and is not limited to specific cases disclosed herein. In some embodiments, the machine learning model can comprise a graph model.
[0063] The machine learning model can comprise a neural network comprising various architectures, loss functions, optimization algorithms, priors, and various other neural network design choices. In some embodiments, the machine learning model can comprise a neural network. In some embodiments, the machine learning model can comprise an autoencoder. InDocket No.: 60134-711.601 some embodiments, the machine learning model can comprise a generative model. In some embodiments, the machine learning model can comprise a variational autoencoder. In some embodiments, the machine learning model can comprise a generative adversarial network. In some embodiments, the machine learning model can comprise a flow model. In some embodiments, the machine learning model can comprise an autoregressive model. In some embodiments, the machine learning model can comprise a diffusion model. In some embodiments, the machine learning model can comprise a neural network with one or more layers. In some embodiments, the machine learning model can comprise a neural network with one or more fully connected layers. In some embodiments, the machine learning model can comprise a neural network with one or more convolutional layers. In some embodiments, the machine learning model can comprise a neural network with one or more message-passing layers. In some embodiments, the machine learning model can comprise a neural network with a bottleneck layer. In some embodiments, a layer may comprise an attention mechanism, a generalized message-passing graph neural network, or both. In some embodiments, a generalized message passing graph neural network comprises a graph convolutional neural network.
[0064] In some embodiments, the machine learning model can comprise a neural network with residual blocks. In some embodiments, the machine learning model can comprise a neural network with attention. In some embodiments, the machine learning model can comprise a neural network with one or more non-linearities. In some embodiments, the machine learning model can comprise a neural network with one or more dropout layers. In some embodiments, the machine learning model can comprise a neural network with one or more batch normalization layers. In some embodiments, the machine learning model can comprise a regression loss function. In some embodiments, the machine learning model can comprise a logistic loss function. In some embodiments, the machine learning model can comprise a variational loss. In some embodiments, the machine learning model can comprise a prior. In some embodiments, the machine learning model can comprise a Gaussian prior. In some embodiments, the machine learning model can comprise a non-Gaussian prior. In some embodiments, the machine learning model can comprise an adversarial loss. In some embodiments, the machine learning model can comprise a reconstruction loss. In some embodiments, the machine learning model is trained with the Adam optimizer. In some embodiments, the machine learning model is trained with the stochastic gradient descent optimizer. In some embodiments, the model learning model hyperparameters are optimized with Gaussian Processes. In some embodiments, the machine learning model is trained withDocket No.: 60134-711.601 train / validation / test data splits. In some embodiments, the machine learning model is trained with k-fold data splits, with any positive integer for k.
[0065] The machine learning model can comprise a variety of manifold learning algorithms. In some embodiments, the machine learning model can comprise a manifold learning algorithm. In some embodiments, the manifold learning algorithm comprises principal component analysis. In some embodiments, the manifold learning algorithm comprises a uniform manifold approximation algorithm. In some embodiments, the manifold learning algorithm comprises an isomap algorithm. In some embodiments, the manifold learning algorithm comprises a locally linear embedding algorithm. In some embodiments, the manifold learning algorithm comprises a modified locally linear embedding algorithm. In some embodiments, the manifold learning algorithm comprises a Hessian eigen mapping algorithm. In some embodiments, the manifold learning algorithm comprises a spectral embedding algorithm. In some embodiments, the manifold learning algorithm comprises a local tangent space alignment algorithm. In some embodiments, the manifold learning algorithm comprises a multi-dimensional scaling algorithm. In some embodiments, the manifold learning algorithm comprises a t-distributed stochastic neighbor embedding algorithm (t-SNE). In some embodiments, the manifold learning algorithm comprises a Bames-Hut t-SNE algorithm.
[0066] The types of machine learning models may include deep learning neural network architectures (e.g., convolutional neural networks (CNNs), recursive neural networks (RNNs), long short-term memories (LSTMs), transformers, graph convolution networks (GCNs), graph attention networks (GATs), message-passing neural networks (MPNNs), generative adversarial networks (GANs), other neural network architectures, or any combination thereof) and traditional machine learning (e.g., linear and logistic regression, decision trees, gradient boosting, support vector machines, k-nearest neighbor (kNN), and other predictive, classification, and clustering algorithms).
[0067] In some embodiments, the methods of the disclosure further comprise reducing one or more molecular representations using a machine learning model. The terms “reducing”, “dimensionality reduction”, “projection”, “component analysis”, “feature space reduction”, “latent space engineering”, “feature space engineering”, “representation engineering”, or “latent space embedding”, as used herein, generally refer to a method of transforming a given input data with an initial number of dimensions to another form of data that has fewer dimensions than the initial number of dimensions. In some embodiments, the terms can refer to the principle of reducing a set of input dimensions to a smaller set of output dimensions.
[0068] The term “normalizing”, as used herein, generally refers to a collection of methods for adjusting a dataset to align the dataset to a common scale. In some embodiments, a normalizingDocket No.: 60134-711.601 method can comprise multiplying a portion or the entirety of a dataset by a factor. In some embodiments, a normalizing method can comprise adding or subtracting a constant from a portion or the entirety of a dataset. In some embodiments, a normalizing method can comprise adjusting a portion or the entirety of a dataset to a known statistical distribution. In some embodiments, a normalizing method can comprise adjusting a portion or the entirety of a dataset to a normal distribution. In some embodiments, a normalizing method can comprise adjusting the dataset so that the signal strength of a portion or the entirety of a dataset is about the same.
[0069] Converting can comprise one or more steps of various conversions of data. In some embodiments, converting can comprise normalizing data. In some embodiments, converting can comprise performing a mathematical operation that computes a score based on a distance between 2 points in the data. In some embodiments, the points in the data can comprise a molecular representation. In some embodiments, the distance can comprise a distance between two edges in a graph. In some embodiments, the distance can comprise a number of nodes and / or edges between the two edges. In some embodiments, the distance can comprise a distance between two nodes in a graph. In some embodiments, the distance can comprise a number of nodes and / or edges between the two nodes. In some embodiments, the distance can comprise a distance between a node and an edge in a graph. In some embodiments, the distance can comprise a number of nodes and / or edges between the node and the edge. In some embodiments, the distance can comprise a Euclidean distance. In some embodiments, the distance can comprise a non-Euclidean distance. Hamming distance, Manhattan distance, minkowski distance, cosine similarity, geodesic distance, shortest path distance. In some embodiments, the distance can be computed in a frequency space. In some embodiments, the distance can be computed in Fourier space. In some embodiments, the distance can be computed in Laplacian space. In some embodiments, the distance can be computed in spectral space. In some embodiments, the mathematical operation can be a monotonic function based on the distance. In some embodiments, the mathematical operation can be a non-monotonic function based on the distance. In some embodiments, the mathematical operation can be an exponential decay function. In some embodiments, the mathematical operation can be a learned function.
[0070] In some embodiments, converting can comprise transforming data in one representation to another representation. In some embodiments, converting can comprise transforming data into another form of data with less dimensions. In some embodiments, converting can comprise linearizing one or more curved paths in the data. In some embodiments, converting can be performed on data comprising data in Euclidean space. In some embodiments, converting can be performed on data comprising data in graph space. In some embodiments, converting can be performed on data in a discrete space. In some embodiments, converting can be performed onDocket No.: 60134-711.601 data comprising data in frequency space. In some embodiments, converting can transform data in discrete space to continuous space, continuous space to discrete space, graph space to continuous space, continuous space to graph space, graph space to discrete space, discrete space to graph space, or any combination thereof. In some embodiments, converting can comprise transforming data in discrete space into a frequency domain. In some embodiments, converting can comprise transforming data in continuous space into a frequency domain. In some embodiments, converting can comprise transforming data in graph space into a frequency domain.
[0071] In some embodiments, reducing can comprise transforming a given input data with any initial number of dimensions to another form of data that has any number of dimensions fewer than the initial number of dimensions. In some embodiments, reducing can comprise transforming input data into another form of data with fewer dimensions. In some embodiments, reducing can comprise linearizing one or more curved paths in the input data to the output data. In some embodiments, reducing can be performed on data comprising data in Euclidean space. In some embodiments, reducing can be performed on data comprising data in graph space. In some embodiments, reducing can be performed on data in a discrete space. In some embodiments, reducing can transform data in discrete space to continuous space, continuous space to discrete space, graph space to continuous space, continuous space to graph space, graph space to discrete space, discrete space to graph space, or any combination thereof.
[0072] The terms “clustering”, “cluster analysis”, or “generating modules”, as used herein, generally refer to a method of grouping samples in a dataset by some measure of similarity. Samples can be grouped in a set space, for example, element ‘a’ is in set ‘A’. Samples can be grouped in a continuous space, for example, element ‘a’ is a point in Euclidean space with distance T away from the centroid of elements comprising cluster ‘A’. Samples can be grouped in a graph space, for example, element ‘a’ is highly connected to elements comprising cluster ‘A’. These terms can refer to the principle of organizing a plurality of elements into groups in some mathematical space based on some measure of similarity.
[0073] In some embodiments, the method further comprises clustering a cohort of molecular representations to determine one or more groups of molecular representations with similar structures, properties, or functions. Clustering can comprise grouping any number of samples in a dataset by any quantitative measure of similarity. In some embodiments, clustering can comprise K-means clustering. In some embodiments, clustering can comprise hierarchical clustering. In some embodiments, clustering can comprise using random forest models. In some embodiments, clustering can comprise boosted tree models. In some embodiments, clustering can comprise using support vector machines. In some embodiments, clustering can compriseDocket No.: 60134-711.601 calculating one or more N-l dimensional surfaces in N-dimensional space that partitions a dataset into clusters. In some embodiments, clustering can comprise distribution-based clustering. In some embodiments, clustering can comprise fitting a plurality of prior distributions over the data distributed in N-dimensional space. In some embodiments, clustering can comprise using density-based clustering. In some embodiments, clustering can comprise using fuzzy clustering. In some embodiments, clustering can comprise computing probability values of a data point belonging to a cluster. In some embodiments, clustering can comprise using constraints. In some embodiments, clustering can comprise using supervised learning. In some embodiments, clustering can comprise using unsupervised learning.
[0074] In some embodiments, clustering can comprise grouping samples based on similarity. In some embodiments, clustering can comprise grouping samples based on quantitative similarity. In some embodiments, clustering can comprise grouping samples based on one or more features of each sample. In some embodiments, clustering can comprise grouping samples based on one or more labels of each sample. In some embodiments, clustering can comprise grouping samples based on Euclidean coordinates. In some embodiments, clustering can comprise grouping samples based the features of the nodes and edges of each sample.
[0075] In some embodiments, comparing can comprise comparing between a first group and different second group. In some embodiments, a first or a second group can each independently be a cluster. In some embodiments, a first or a second group can each independently be a group of clusters. In some embodiments, comparing can comprise comparing between one cluster with a group of clusters. In some embodiments, comparing can comprise comparing between a first group of clusters with second group of clusters different than the first group. In some embodiments, one group can be one sample. In some embodiments, one group can be a group of samples. In some embodiments, comparing can comprise comparing between one sample versus a group of samples. In some embodiments, comparing can comprise comparing between a group of samples versus a group of samples.
[0076] Comparing can comprise a variety of analytical methods carried out by a computer or a human. In some embodiments, a statistical test can be carried out to identify one or more molecular representations that are the most different in one group versus a comparison group. In some embodiments, clustering can be carried out on differences in molecular representations, which can lead to the identification of a set of molecular representations that show a high confidence for performing a satisfying a given machine learning task.
[0077] Comparing may comprise use of learned representation in a machine learning model. One or more features of the learned representation can be attributed to a particular prediction. For example, in predicting antimicrobial activity or cytotoxicity, specific features such asDocket No.: 60134-711.601 particular atoms, fingerprints, or any other features disclosed herein may be highlighted to indicate its association with the prediction.
[0078] In some embodiments, the compound structure is obtained using at least a molecular simulation. In some embodiments, the molecular simulation is based at least partially on an electronic structure calculation, a forcefield based calculation, molecular dynamics, a Monte Carlo simulation, or any combination thereof.
[0079] Various molecular modeling techniques may be used to train and / or develop a useful machine learning model. The usefulness of the machine learning model may be dependent on the quality, quantity, and diversity of the molecular properties that serve as inputs for a machine learning algorithm to be trained on. Various experimental and computational methods may be used to generate useful inputs for machine learning, including quantum mechanical calculations, experimental characterizations, and molecular dynamics. In some embodiments, quantum mechanical calculations and / or approximations can yield accurate descriptors of key molecular properties such as charge distribution, optimal molecular conformations, and transition energies between conformations, macroscopic properties (e.g., solvation free energies). Quantum mechanical approximations can range from Hartree-Fock and density functional theory (DFT) to highly expensive calculations such as coupled cluster. Molecular dynamics may be used as a major input source for useful training data for machine learning. Techniques such as thermodynamic integration can provide free energies of solvation.
[0080] In some aspects, the present disclosure provides an autonomous system for drug discovery. The system can generate new compounds for computational screening, perform computational methods to narrow down a number of compounds, suggest compounds for experimental assays, drive an autonomous laboratory to synthesize the compounds, perform experimental assays on the compounds, manage and / or build a compound library, or any combination thereof.
[0081] In some embodiments, a compound structure may be obtained from at least an experiment or a structure database. In some embodiments, the set of labels comprises (i) a physicochemical property, (ii) an absorption, distribution, metabolism, excretion, or toxicity (ADMET) property, (iii) a chemical reaction, (iv) a synthesizability, (v) a solubility, (vi) a chemical stability, (vii) an activity, (viii) a selectivity, (ix) a potency, (x) a pharmacokinetic property, (xi) a pharmacodynamic property, (xii) an in vivo safety property, (xiii) a formulation property, or (xiv) any combination thereof.Docket No.: 60134-711.601Computing System
[0082] In some aspects, the present disclosure describes a computer-implemented system comprising: a digital processing device comprising: at least one processor, an operating system configured to perform executable instructions, a memory, and a computer program including instructions executable by the digital processing device to train a machine learning model or rank compounds for efficacy. In some aspects, the present disclosure describes a computer- implemented method, implementing any one of the methods disclosed herein in a computer system. Referring to FIG. 14, a block diagram is shown depicting an exemplary machine that includes a computer system 1400 (e.g., a processing or computing system) within which a set of instructions can execute for causing a device to perform or execute any one or more of the aspects and / or methodologies for training a machine learning model or ranking compounds for efficacy. The components in FIG. 14 are examples only and do not limit the scope of use or functionality of any hardware, software, embedded logic component, or a combination of two or more such components implementing particular embodiments.
[0083] Computer system 1400 may include one or more processors 1401, a memory 1403, and a storage 1408 that communicate with each other, and with other components, via a bus 1440. The bus 1440 may also link a display 1432, one or more input devices 1433 (which may, for example, include a keypad, a keyboard, a mouse, a stylus, etc.), one or more output devices 1434, one or more storage devices 1435, and various tangible storage media 1436. All of these elements may interface directly or via one or more interfaces or adaptors to the bus 1440. For instance, the various tangible storage media 1436 can interface with the bus 1440 via storage medium interface 1426. Computer system 1400 may have any suitable physical form, including but not limited to one or more integrated circuits (ICs), printed circuit boards (PCBs), mobile handheld devices (such as mobile telephones or PDAs), laptop or notebook computers, distributed computer systems, computing grids, or servers.
[0084] Computer system 1400 includes one or more processor(s) 1401 (e.g., central processing units (CPUs), general purpose graphics processing units (GPGPUs), or quantum processing units (QPUs)) that carry out functions. Computer system 1400 may be one of various high performance computing platforms. For instance, the one or more processor(s) 1401 may form a high performance computing cluster. In some embodiments, the one or more processors 1401 may form a distributed computing system connected by wired and / or wireless networks. In some embodiments, arrays of CPUs, GPUs, QPUs, or any combination thereof may be operably linked to implement any one of the methods disclosed herein. Processor(s) 1401 optionally contains a cache memory unit 1402 for temporary local storage of instructions, data, or computer addresses. Processor(s) 1401 are configured to assist in execution of computer readableDocket No.: 60134-711.601 instructions. Computer system 1400 may provide functionality for the components depicted in FIG. 14 as a result of the processor(s) 1401 executing non-transitory, processor-executable instructions embodied in one or more tangible computer-readable storage media, such as memory 1403, storage 1408, storage devices 1435, and / or storage medium 1436. The computer- readable media may store software that implements particular embodiments, and processor(s) 1401 may execute the software. Memory 1403 may read the software from one or more other computer-readable media (such as mass storage device(s) 1435, 1436) or from one or more other sources through a suitable interface, such as network interface 1420. The software may cause processor(s) 1401 to carry out one or more processes or one or more steps of one or more processes described or illustrated herein. Carrying out such processes or steps may include defining data structures stored in memory 1403 and modifying the data structures as directed by the software.
[0085] The memory 1403 may include various components (e.g., machine readable media) including, but not limited to, a random access memory component (e.g., RAM 1404) (e.g., static RAM (SRAM), dynamic RAM (DRAM), ferroelectric random access memory (FRAM), phasechange random access memory (PRAM), etc.), a read-only memory component (e.g., ROM 1405), and any combinations thereof. ROM 1405 may act to communicate data and instructions unidirectionally to processor(s) 1401, and RAM 1404 may act to communicate data and instructions bidirectionally with processor(s) 1401. ROM 1405 and RAM 1404 may include any suitable tangible computer-readable media described below. In one example, a basic input / output system 1406 (BIOS), including basic routines that help to transfer information between elements within computer system 1400, such as during start-up, may be stored in the memory 1403.
[0086] Fixed storage 1408 is connected bidirectionally to processor(s) 1401, optionally through storage control unit 1407. Fixed storage 1408 provides additional data storage capacity and may also include any suitable tangible computer-readable media described herein. Storage 1408 may be used to store operating system 1409, executable(s) 1410, data 1411, applications 1412 (application programs), and the like. Storage 1408 can also include an optical disk drive, a solid- state memory device (e.g., flash-based systems), or a combination of any of the above. Information in storage 1408 may, in appropriate cases, be incorporated as virtual memory in memory 1403.
[0087] In one example, storage device(s) 1435 may be removably interfaced with computer system 1400 (e.g., via an external port connector (not shown)) via a storage device interface 1425. Particularly, storage device(s) 1435 and an associated machine-readable medium may provide non-volatile and / or volatile storage of machine-readable instructions, data structures,Docket No.: 60134-711.601 program modules, and / or other data for the computer system 1400. In one example, software may reside, completely or partially, within a machine-readable medium on storage device(s) 1435. In another example, software may reside, completely or partially, within processor(s) 1401
[0088] Bus 1440 connects a wide variety of subsystems. Herein, reference to a bus may encompass one or more digital signal lines serving a common function, where appropriate. Bus 1440 may be any of several types of bus structures including, but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combinations thereof, using any of a variety of bus architectures. As an example, and not by way of limitation, such architectures include an Industry Standard Architecture (ISA) bus, an Enhanced ISA (EISA) bus, a Micro Channel Architecture (MCA) bus, a Video Electronics Standards Association local bus (VLB), a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, an Accelerated Graphics Port (AGP) bus, HyperTransport (HTX) bus, serial advanced technology attachment (SATA) bus, and any combinations thereof.
[0089] Computer system 1400 may also include an input device 1433. In one example, a user of computer system 1400 may enter commands and / or other information into computer system 1400 via input device(s) 1433. Examples of an input device(s) 1433 include, but are not limited to, an alpha-numeric input device (e.g., a keyboard), a pointing device (e.g., a mouse or touchpad), a touchpad, a touch screen, a multi-touch screen, a joystick, a stylus, a gamepad, an audio input device (e.g., a microphone, a voice response system, etc.), an optical scanner, a video or still image capture device (e.g., a camera), and any combinations thereof. In some embodiments, the input device is a Kinect, Leap Motion, or the like. Input device(s) 1433 may be interfaced to bus 1440 via any of a variety of input interfaces 1423 (e.g., input interface 1423) including, but not limited to, serial, parallel, game port, USB, FIREWIRE, THUNDERBOLT, or any combination of the above. In some embodiments, an input device 1433 may be used to train a machine learning model or rank compounds for efficacy. In some embodiments, instructions can be received using human inputs through an input device 1433.
[0090] In particular embodiments, when computer system 1400 is connected to network 1430, computer system 1400 may communicate with other devices, specifically mobile devices and enterprise systems, distributed computing systems, cloud storage systems, cloud computing systems, and the like, connected to network 1430. Communications to and from computer system 1400 may be sent through network interface 1420. For example, network interface 1420 may receive incoming communications (such as requests or responses from other devices) in the form of one or more packets (such as Internet Protocol (IP) packets) from network 1430, and computer system 1400 may store the incoming communications in memory 1403 for processing.Docket No.: 60134-711.601Computer system 1400 may similarly store outgoing communications (such as requests or responses to other devices) in the form of one or more packets in memory 1403 and communicated to network 1430 from network interface 1420. Processor(s) 1401 may access these communication packets stored in memory 1403 for processing.
[0091] Examples of the network interface 1420 include, but are not limited to, a network interface card, a modem, and any combination thereof. Examples of a network 1430 or network segment 1430 include, but are not limited to, a distributed computing system, a cloud computing system, a wide area network (WAN) (e.g., the Internet, an enterprise network), a local area network (LAN) (e.g., a network associated with an office, a building, a campus or other relatively small geographic space), a telephone network, a direct connection between two computing devices, a peer-to-peer network, and any combinations thereof. A network, such as network 1430, may employ a wired and / or a wireless mode of communication. In general, any network topology may be used.
[0092] Information and data can be displayed through a display 1432. Examples of a display 1432 include, but are not limited to, a cathode ray tube (CRT), a liquid crystal display (LCD), a thin film transistor liquid crystal display (TFT-LCD), an organic liquid crystal display (OLED) such as a passive-matrix OLED (PMOLED) or active-matrix OLED (AMOLED) display, a plasma display, and any combinations thereof. The display 1432 can interface to the processor(s) 1401, memory 1403, and fixed storage 1408, as well as other devices, such as input device(s) 1433, via the bus 1440. The display 1432 is linked to the bus 1440 via a video interface 1422, and transport of data between the display 1432 and the bus 1440 can be controlled via the graphics control 1421. In some embodiments, the display is a video projector. In some embodiments, the display is a head-mounted display (HMD) such as a VR headset. In further embodiments, suitable VR headsets include, by way of non-limiting examples, HTC Vive, Oculus Rift, Samsung Gear VR, Microsoft HoloLens, Razer OSVR, FOVE VR, Zeiss VR One, Avegant Glyph, Freefly VR headset, and the like. In still further embodiments, the display is a combination of devices such as those disclosed herein.
[0093] In addition to a display 1432, computer system 1400 may include one or more other peripheral output devices 1434 including, but not limited to, an audio speaker, a printer, a storage device, and any combinations thereof. Such peripheral output devices may be connected to the bus 1440 via an output interface 1424. Examples of an output interface 1424 include, but are not limited to, a serial port, a parallel connection, a USB port, a FIREWIRE port, a THUNDERBOLT port, and any combinations thereof.
[0094] In addition, or as an alternative, computer system 1400 may provide functionality as a result of logic hardwired or otherwise embodied in a circuit, which may operate in place of orDocket No.: 60134-711.601 together with software to execute one or more processes or one or more steps of one or more processes described or illustrated herein. Reference to software in this disclosure may encompass logic, and reference to logic may encompass software. Moreover, reference to a computer-readable medium may encompass a circuit (such as an IC) storing software for execution, a circuit embodying logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware, software, or both.
[0095] Those of skill in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality.
[0096] The various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein may be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0097] The steps of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by one or more processor(s), or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.
[0098] In accordance with the description herein, suitable computing devices include, by way of non-limiting examples, server computers, desktop computers, laptop computers, notebookDocket No.: 60134-711.601 computers, sub-notebook computers, netbook computers, netpad computers, set-top computers, media streaming devices, handheld computers, Internet appliances, mobile smartphones, and tablet computers.
[0099] In some embodiments, the computing device includes an operating system configured to perform executable instructions. The operating system is, for example, software, including programs and data, which manages the device’s hardware and provides services for execution of applications. Those of skill in the art will recognize that suitable server operating systems include, by way of non-limiting examples, FreeBSD, OpenBSD, NetBSD®, Linux, Apple® Mac OS X Server®, Oracle® Solaris®, Windows Server®, and Novell® NetWare®. Those of skill in the art will recognize that suitable personal computer operating systems include, by way of nonlimiting examples, Microsoft® Windows®, Apple® Mac OS X®, UNIX®, and UNIX-like operating systems such as GNU / Linux®. In some embodiments, the operating system is provided by cloud computing. Those of skill in the art will also recognize that suitable mobile smartphone operating systems include, by way of non-limiting examples, Nokia® Symbian® OS, Apple® los®, Research In Motion® BlackBerry OS®, Google® Android®, Microsoft® Windows Phone® OS, Microsoft® Windows Mobile® OS, Linux®, and Palm® WebOS®.
[0100] In some embodiments, a computer system 1400 may be accessible through a user terminal to receive user commands. The user commands may include line commands, scripts, programs, etc., and various instructions executable by the computer system 1400. A computer system 1400 may receive instructions to train a machine learning model or rank compounds for efficacy, or schedule a computing job for the computer system 1400 to carry out any instructions.Non-Transitory Computer Readable Storage Medium
[0101] In some aspects, the present disclosure describes a non-transitory computer-readable storage media encoded with a computer program including instructions executable by one or more processors to train a machine learning model or rank compounds for efficacy using any one of the methods disclosed herein. In some embodiments, a non-transitory computer-readable storage media may comprise instructions for training a machine learning model or ranking compounds for efficacy. In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more non-transitory computer readable storage media encoded with a program including instructions executable by the operating system of an optionally networked computing device.
[0102] In further embodiments, a computer readable storage medium is a tangible component of a computing device. In still further embodiments, a computer readable storage medium isDocket No.: 60134-711.601 optionally removable from a computing device. In some embodiments, a computer readable storage medium includes, by way of non-limiting examples, flash memory devices, solid state memory, magnetic disk drives, magnetic tape drives, optical disk drives, distributed computing systems including cloud computing systems and services, and the like. In some embodiments, the program and instructions are permanently, substantially permanently, semi-permanently, or non-transitorily encoded on the media.Computer Program
[0103] In some aspects, the present disclosure describes a computer program product comprising a computer-readable medium having computer-executable code encoded therein, the computer-executable code adapted to be executed to implement any one of the methods disclosed herein. In some embodiments, the platforms, systems, media, and methods disclosed herein include at least one computer program, or use of the same.
[0104] A computer program includes a sequence of instructions, executable by one or more processor(s) of the computing device’s CPU, written to perform a specified task. Computer readable instructions may be implemented as program modules, such as functions, objects, Application Programming Interfaces (APIs), computing data structures, and the like, that perform particular tasks or implement particular abstract data types. In light of the disclosure provided herein, those of skill in the art will recognize that a computer program may be written in various versions of various languages. In some embodiments, APIs may comprise various languages, for example, languages in various releases of TensorFlow, Theano, Keras, PyTorch, or any combination thereof which may be implemented in various releases of Python, Python3, C, C#, C++, MatLab, R, Java, or any combination thereof.
[0105] The functionality of the computer readable instructions may be combined or distributed as desired in various environments. In some embodiments, a computer program comprises one sequence of instructions. In some embodiments, a computer program comprises a plurality of sequences of instructions. In some embodiments, a computer program is provided from one location. In other embodiments, a computer program is provided from a plurality of locations. In various embodiments, a computer program includes one or more software modules. In various embodiments, a computer program includes, in part or in whole, one or more web applications, one or more standalone applications, one or more web browser plug-ins, extensions, add-ins, or add-ons, or combinations thereof.Docket No.: 60134-711.601Web Application
[0106] In some embodiments, a computer program includes a web application. In some embodiments, a user may enter a query for training a machine learning model or ranking compounds for efficacy through a web application. In some embodiments, a user may train a machine learning model or rank compounds for efficacy through a web application. In light of the disclosure provided herein, those of skill in the art will recognize that a web application, in various embodiments, utilizes one or more software frameworks and one or more database systems. In some embodiments, a web application is created upon a software framework such as Microsoft® .NET or Ruby on Rails (RoR). In some embodiments, a web application utilizes one or more database systems including, by way of non-limiting examples, relational, non-relational, object oriented, associative, XML, and document oriented database systems. In further embodiments, suitable relational database systems include, by way of non-limiting examples, Microsoft® SQL Server, mySQL™, and Oracle®. Those of skill in the art will also recognize that a web application, in various embodiments, is written in one or more versions of one or more languages. A web application may be written in one or more markup languages, presentation definition languages, client-side scripting languages, server-side coding languages, database query languages, or combinations thereof. In some embodiments, a web application is written to some extent in a markup language such as Hypertext Markup Language (HTML), Extensible Hypertext Markup Language (XHTML), or extensible Markup Language (XML). In some embodiments, a web application is written to some extent in a presentation definition language such as Cascading Style Sheets (CSS). In some embodiments, a web application is written to some extent in a client-side scripting language such as Asynchronous JavaScript and XML (AJAX), Flash® ActionScript, JavaScript, or Silverlight®. In some embodiments, a web application is written to some extent in a server-side coding language such as Active Server Pages (ASP), ColdFusion®, Perl, Java™, JavaServer Pages (JSP), Hypertext Preprocessor (PHP), Python™, Ruby, Tel, Smalltalk, WebDNA®, or Groovy. In some embodiments, a web application is written to some extent in a database query language such as Structured Query Language (SQL). In some embodiments, a web application integrates enterprise server products such as IBM® Lotus Domino®.Mobile application
[0107] In some embodiments, a computer program includes a mobile application provided to a mobile computing device. In some embodiments, the mobile application is provided to a mobile computing device at the time it is manufactured. In other embodiments, the mobile application is provided to a mobile computing device via the computer network described herein.Docket No.: 60134-711.601
[0108] In view of the disclosure provided herein, a mobile application is created by techniques known to those of skill in the art using hardware, languages, and development environments known to the art. Those of skill in the art will recognize that mobile applications are written in several languages. Suitable programming languages include, by way of non-limiting examples, C, C++, C#, Objective-C, Java™, JavaScript, Pascal, Object Pascal, Python™, Ruby, VB.NET, WML, and XHTML / HTML with or without CSS, or combinations thereof.
[0109] Suitable mobile application development environments are available from several sources. Commercially available development environments include, by way of non-limiting examples, AirplaySDK, alcheMo, Appcelerator®, Celsius, Bedrock, Flash Lite, .NET Compact Framework, Rhomobile, and WorkLight Mobile Platform. Other development environments are available without cost including, by way of non-limiting examples, Lazarus, MobiFlex, MoSync, and Phonegap. Also, mobile device manufacturers distribute software developer kits including, by way of non-limiting examples, iPhone and iPad (los) SDK, Android™ SDK, BlackBerry® SDK, BREW SDK, Palm® OS SDK, Symbian SDK, webOS SDK, and Windows® Mobile SDK.Standalone application
[0110] In some embodiments, a computer program includes a standalone application, which is a program that is run as an independent computer process, not an add-on to an existing process, e.g., not a plug-in. Those of skill in the art will recognize that standalone applications are often compiled. A compiler is a computer program(s) that transforms source code written in a programming language into binary object code such as assembly language or machine code. Suitable compiled programming languages include, by way of non-limiting examples, C, C++, Objective-C, COBOL, Delphi, Eiffel, Java™, Lisp, Python™, Visual Basic, and VB .NET, or combinations thereof. Compilation is often performed, at least in part, to create an executable program. In some embodiments, a computer program includes one or more executable complied applications.Software Modules[OHl] In some embodiments, the platforms, systems, media, and methods disclosed herein include software, server, and / or database modules, or use of the same. In view of the disclosure provided herein, software modules are created by techniques known to those of skill in the art using machines, software, and languages known to the art. The software modules disclosed herein are implemented in a multitude of ways. In various embodiments, a software moduleDocket No.: 60134-711.601 comprises a file, a section of code, a programming object, a programming structure, a distributed computing resource, a cloud computing resource, or combinations thereof. In further various embodiments, a software module comprises a plurality of files, a plurality of sections of code, a plurality of programming objects, a plurality of programming structures, a plurality of distributed computing resources, a plurality of cloud computing resources, or combinations thereof. In various embodiments, the one or more software modules comprise, by way of nonlimiting examples, a web application, a mobile application, a standalone application, and a distributed or cloud computing application. In some embodiments, software modules are in one computer program or application. In other embodiments, software modules are in more than one computer program or application. In some embodiments, software modules are hosted on one machine. In other embodiments, software modules are hosted on more than one machine. In further embodiments, software modules are hosted on a distributed computing platform such as a cloud computing platform. In some embodiments, software modules are hosted on one or more machines in one location. In other embodiments, software modules are hosted on one or more machines in more than one location.Databases
[0112] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more databases, or use of the same. In view of the disclosure provided herein, those of skill in the art will recognize that many databases are suitable for storage and retrieval of information about training a machine learning model or ranking compounds for efficacy, or any combination thereof. In various embodiments, suitable databases include, by way of nonlimiting examples, relational databases, non-relational databases, object oriented databases, object databases, entity-relationship model databases, associative databases, XML databases, document oriented databases, and graph databases. Further non-limiting examples include SQL, PostgreSQL, MySQL, Oracle, DB2, Sybase, and MongoDB. In some embodiments, a database is Internet-based. In further embodiments, a database is web-based. In still further embodiments, a database is cloud computing-based. In a particular embodiment, a database is a distributed database. In other embodiments, a database is based on one or more local computer storage devices.
[0113] Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least” or “greater than” applies to each one of the numerical values in that series of numerical values.Docket No.: 60134-711.601
[0114] Whenever the term “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “no more than” or “less than” applies to each one of the numerical values in that series of numerical values.
[0115] The term “about” or “nearly” as used herein generally refers to within (plus or minus) 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% of a designated value.
[0116] As used herein, the singular forms “a”, “an”, and “the” include plural references unless the context clearly dictates otherwise.
[0117] While preferred embodiments of the present disclosure have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the disclosure. It should be understood that various alternatives to the embodiments of the present disclosure may be employed in practicing the present disclosure. It is intended that the following claims define the scope of the present disclosure and that methods and structures within the scope of these claims and their equivalents be covered thereby.List of Embodiments
[0118] The following list of embodiments of the invention are to be considered as disclosing various features of the invention, which features can be considered to be specific to the particular embodiment under which they are discussed, or which are combinable with the various other features as listed in other embodiments. Thus, simply because a feature is discussed under one particular embodiment does not necessarily limit the use of that feature to that embodiment.
[0119] Embodiment 1. A computer-implemented method of training a machine learning model for ranking compounds for efficacy, comprising: (a) providing a training dataset comprising data generated from assays, wherein at least two assays in the assays yield different measures of efficacy of the compounds; and (b) training the machine learning model using the training dataset, wherein the machine learning model yields a ranking of compounds based on the different measures of efficacy.
[0120] Embodiment 2. A computer-implemented method of ranking compounds for efficacy, comprising ranking compounds based on their measures of efficacy using a model trained using a dataset comprising data generated from assays comprising at least two assays, wherein the at least two assays yield different measures of efficacy of the compounds.
[0121] Embodiment 3. The computer-implemented method of Embodiment 1 or 2, wherein the measures of efficacy comprise antibacterial activity, cytotoxicity, or both.Docket No.: 60134-711.601
[0122] Embodiment 4. The computer-implemented method of any one of Embodiments 1-3, wherein the dataset comprises at least three, four, or five assays, that yield the different measures of efficacy.
[0123] Embodiment 5. The computer-implemented method of Embodiment 4, wherein the dataset comprises at least 5, 10, 15, 20, 25, 50, 75, or 100 assays, that yield different measures of efficacy.
[0124] Embodiment 6. The computer-implemented method of any one of Embodiments 1-5, wherein the assays comprise antimicrobial susceptibility assays.
[0125] Embodiment 7. The computer-implemented method of any one of Embodiments 1-6, wherein the assays are selected from the group consisting of agar well diffusion, disk well diffusion, agar dilution, broth dilution, ATP bioluminescence, thin-layer chromatography, antimicrobial susceptibility testing, bioautography, ATP luminescence, a fluorescence assay, and any combination thereof.
[0126] Embodiment 8. The computer-implemented method of any one of Embodiments 1-7, wherein the at least two assays are performed using different conditions.
[0127] Embodiment 9. The computer-implemented method of Embodiment 8, wherein the different conditions are selected from the group consisting of temperatures, buffers, media, oxygen concentration, and any combination thereof.
[0128] Embodiment 10. The computer-implemented method of any one of Embodiments 1-9, wherein the assays comprise at least two assays performed on different species of bacteria.
[0129] Embodiment 11. The computer-implemented method of Embodiment 10, wherein the different strains of bacteria comprise a wild-type strain.
[0130] Embodiment 12. The computer-implemented method of Embodiments 10 or 11, wherein the different strains of bacteria comprise a mutant-type strain.
[0131] Embodiment 13. The computer-implemented method of any one of Embodiments 10-12, wherein the different strains of bacteria comprise an efflux deficient strain.
[0132] Embodiment 14. The computer-implemented method of any one of Embodiments 10-13, wherein the different strains of bacteria comprise an efflux competent strain.
[0133] Embodiment 15. The computer-implemented method of any one of Embodiments 10-14, wherein the different strains of bacteria comprise a membrane compromised strain.
[0134] Embodiment 16. The computer-implemented method of any one of Embodiments 10-15, wherein the different strains of bacteria comprise a membrane intact strain.
[0135] Embodiment 17. The computer-implemented method of any one of Embodiments 10-16, wherein the different species of bacteria are selected from the following genera: Acinetobacter, Bacteroides, Citrobacter, Clostridioides, Corynebacterium, Enterobacter,Docket No.: 60134-711.601Enterococcus, Escherichia, Eggerthella, Klebsiella, Lactobacillus, Listeria, Moraxella, Morganella, Neisseria, Peptostreptococcus, Prevotella, Propionibacterium, Providencia, Pseudomonas, Salmonella, Shigella, Staphylococcus, Stenotrophomonas, Streptococcus, Mycobacteroides, and Mycobacterium.
[0136] Embodiment 18. The computer-implemented method of any one of Embodiments 10-17, wherein the different strains of bacteria are selected from the following genera: Acinetobacter, Bacteroides, Citrobacter, Clostridioides, Corynebacterium, Enterobacter, Enterococcus, Escherichia, Eggerthella, Klebsiella, Lactobacillus, Listeria, Moraxella, Morganella, Neisseria, Peptostreptococcus, Prevotella, Propionibacterium, Providencia, Pseudomonas, Salmonella, Shigella, Staphylococcus, Stenotrophomonas, Streptococcus, Mycobacteroides, and Mycobacterium.
[0137] Embodiment 19. The computer-implemented method of any one of Embodiments 10-18, wherein the different species of bacteria are selected from the following species: Acinetobacter baumannii, Acinetobacter haemolyticus, Acinetobacter nosocomialis, Acinetobacter pittii, Acinetobacter radioresistens, Acinetobacter ursingii, Bacteroides fragilis, Citrobacter braakii, Citrobacter freundii, Citrobacter koseri, Citrobacter sedlakii, Clostridioides difficile, Clostridioides perfringens, Corynebacterium jeikeium, Enterobacter aerogenes, Enterobacter cloacae, Enterococcus faecalis, Enterococcus faecium, Eggerthella lenta, Escherichia coli, Klebsiella oxytoca, Klebsiella pneumoniae, Lactobacillus acidophilus, Listeria monocytogenes, Moraxella catarrhalis, Morganella morganii, Neisseria gonorrhoeae, Peptostreptococcus anaerobius, Prevotella bivia, Propionibacterium acnes, Providencia rettgeri, Providencia stuartii, Pseudomonas aeruginosa, Salmonella typhimurium, Shigella dysenteriae, Staphylococcus aureus, Staphylococcus capitis, Staphylococcus caprae, Staphylococcus epidermidis, Staphylococcus haemolyticus, Staphylococcus hominis, Staphylococcus lugdenensis, Staphylococcus pettenkoferi, Staphylococcus saprophyticus, Staphylococcus simulans, Staphylococcus warnerii, Stenotrophomonas maltophilia, Streptococcus agalactiae, Streptococcus constellatus, Streptococcus pneumoniae, Streptococcus pyogenes, Mycobacterium tuberculosis, Mycobacterium avium, Mycobacterium intracellulare, Mycobacterium smegmatis, and Mycobacteroides abscessus.
[0138] Embodiment 20. The computer-implemented method of any one of Embodiments 10-19, wherein the different strains of bacteria are selected from the following species: Acinetobacter baumannii, Acinetobacter haemolyticus, Acinetobacter nosocomialis, Acinetobacter pittii, Acinetobacter radioresistens, Acinetobacter ursingii, Bacteroides fragilis, Citrobacter braakii, Citrobacter freundii, Citrobacter koseri, Citrobacter sedlakii, Clostridioides difficile, Clostridioides perfringens, Corynebacterium jeikeium, Enterobacter aerogenes,Docket No.: 60134-711.601Enterobacter cloacae, Enterococcus faecalis, Enterococcus faecium, Eggerthella lenta, Escherichia coli, Klebsiella oxytoca, Klebsiella pneumoniae, Lactobacillus acidophilus, Listeria monocytogenes, Moraxella catarrhalis, Morganella morganii, Neisseria gonorrhoeae, Peptostreptococcus anaerobius, Prevotella bivia, Propionib acterium acnes, Providencia rettgeri, Providencia stuartii, Pseudomonas aeruginosa, Salmonella typhimurium, Shigella dysenteriae, Staphylococcus aureus, Staphylococcus capitis, Staphylococcus caprae, Staphylococcus epidermidis, Staphylococcus haemolyticus, Staphylococcus hominis, Staphylococcus lugdenensis, Staphylococcus pettenkoferi, Staphylococcus saprophyticus, Staphylococcus simulans, Staphylococcus warnerii, Stenotrophomonas maltophilia, Streptococcus agalactiae, Streptococcus constellatus, Streptococcus pneumoniae, Streptococcus pyogenes, Mycobacterium tuberculosis, Mycobacterium avium, Mycobacterium intracellulare, Mycobacterium smegmatis, and Mycobacteroides abscessus.
[0139] Embodiment 21. The computer-implemented method of any one of Embodiments 1-20, wherein the assays are performed on different cell lines.
[0140] Embodiment 22. The computer-implemented method of Embodiment 21, wherein the different cell lines are selected from the group consisting of HepG2, HEK293, NIH3T3, U2OS, HaCat, CRL-7250, or any combination thereof.
[0141] Embodiment 23. The computer-implemented method of any one of Embodiments 1-22, wherein the assays are performed with different endpoints.
[0142] Embodiment 24. The computer-implemented method of Embodiment 23, wherein the different endpoints are selected from the group consisting of growth rate, total growth, optical density, colony forming units, relative light units.
[0143] Embodiment 25. The computer-implemented method of any one of Embodiments 1-24, wherein the training dataset comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 75, or 100 sources of data that are heterogenous.
[0144] Embodiment 26. The computer-implemented method of any one of Embodiments 1-25, wherein the compounds comprise at least 2, 5, 10, 25, 50, 75, 1000, 5000, 10000, 50000, or 100000 compounds.
[0145] Embodiment 27. The computer-implemented method of any one of Embodiments 1-26, wherein the machine learning model yields a ranking of compounds based on their antibacterial activity.
[0146] Embodiment 28. The computer-implemented method of any one of Embodiments 1-27, wherein the machine learning model yields a ranking of compounds based on their cytotoxicity.Docket No.: 60134-711.601
[0147] Embodiment 29. The computer-implemented method of any one of Embodiments 1-28, wherein the machine learning model yields a ranking of compounds based on their antibacterial activity and cytotoxicity.
[0148] Embodiment 30. The computer-implemented method of any one of Embodiments 1-29, further comprising selecting a set of candidate compounds based on a threshold criteria.
[0149] Embodiment 31. The computer-implemented method of Embodiment 30, wherein the threshold criteria is selected from the group consisting of a predetermined number, a predetermined percent of the compounds, an antibacterial activity threshold, a cytotoxicity threshold, and any combination thereof.
[0150] Embodiment 32. The computer-implemented method of any one of Embodiments 1-31, further comprising performing screening assays using the set of candidate compounds.
[0151] Embodiment 33. The computer-implemented method of Embodiment 32, wherein the screening assays are selected from the group consisting of chemistry assays, biochemistry assays, binding assays, functional assays, cell-based assays, in vivo assays, ADME assays, toxicology assays, and any combination thereof .
[0152] Embodiment 34. A computer-implemented method of screening a set of compounds comprising: (a) conducting at least two assays on each compound of the set of compounds, each assay yielding raw data indicating different measures of efficacy; (b) preparing a training dataset based on the raw data; and (c) training the machine learning model using the training dataset, wherein the machine learning model yields a ranking of compounds based on their measures of efficacy.
[0153] Embodiment 35. The computer-implemented method of Embodiment 34, wherein the measures of efficacy comprise antibacterial activity, cytotoxicity, or both.
[0154] Embodiment 36. The computer-implemented method of any one of Embodiments 1-35, wherein the training the machine learning model is performed using a neural network, a random forest, or both.
[0155] Embodiment 37. The computer-implemented method of any one of Embodiments 1-36, wherein the ranking of compounds is based on their relative antibacterial activity, cytotoxicity, or both.
[0156] Embodiment 38. A computer-implemented system comprising: a digital processing device comprising: at least one processor, an operating system configured to perform executable instructions, a memory, and a computer program including instructions executable by the digital processing device to implement any one of the computer-implemented method of Embodiments 1-37.Docket No.: 60134-711.601
[0157] Embodiment 39. A computer program product comprising a computer-readable medium having computer-executable code encoded therein, the computer-executable code adapted to be executed to implement any one of the computer-implemented method of Embodiments 1-37.
[0158] Embodiment 40. A non-transitory computer-readable storage media encoded with a computer program including instructions executable by one or more processors to perform any one of the computer-implemented method of Embodiments 1-37.EXAMPLES
[0159] The following examples are provided to further illustrate some embodiments of the present disclosure, but are not intended to limit the scope of the disclosure; it will be understood by their exemplary nature that other procedures, methodologies, or techniques known to those skilled in the art may alternatively be used.Example 1: Multitask Antibacterial Ranking Model
[0160] This example shows that antibacterial activity ranking models outperform single task classification models. Predictive antibacterial models can be trained from data across multiple bacterial assays, including across strains and media. Screening compounds selected by antibacterial and cytotoxicity models increases antibacterial hit rate by 2x, and reduces cytotoxicity failures by 3x. Screening compounds selected by antibacterial models increases wild type hit rate by 1.5x, indicating the models learn features associated with permeability and efflux avoidance.
[0161] A multitask ranking model was trained to predict antibacterial activity. A random forest (XGBoost) was trained on 1.1M compounds across 9 datasets including WT, efflux deficient, and membrane compromised strains. Input features were crafted by concatenating various molecular fingerprints. l / 5th of each dataset was used as a test set, and held out all scaffolds in test set from the training data. For all other datasets, the entire dataset was held out as the test set. The results of evaluating the models are shown in FIG. 3. The models were evaluated on their ability to find compounds that have known activity. Two metrics were used to evaluate the models. First, “lift@X” indicates the percentage of hits found at 1% screened compared to 1% from a random model (e.g., 20% of hits can be represented by 20x lift). Second, the Gini coefficient was used to measure the function of area under the curve, where 0 represents completely random and 100 represents a perfect model. Single task models were also trained on a permeabilized mutant data or a WT Pubchem screen, and they performed well on the datasets they are trained on but did not generalize as well as the multitask models towards new data.Meanwhile, the ranking model performed equivalently to the single task models on the data they are trained on and better on other datasets.Docket No.: 60134-711.601
[0162] The results show that the ranking model enriches for hits in early selection. Compared to other models, the ranking model identified up to 20x more hits compared to other models earlier in the screening process, as shown in FIG. 4. 20% of actives were identified after screening 1% of the library. As a small percentage of a compound library can selected when cherry-picking compounds in a drug discovery process (e.g., about 1%), early enrichment by the model can be very important.
[0163] The molecule picking was performed by retrieving vendor catalogs to be scored. The vendor catalogs can contain various number of molecules, but often contain from a few hundred thousand to a few million molecules. The compound structures were standardized, and salts and any unwanted elements (e.g. heavy metals) are removed from structures. The molecules were scored with the antibacterial and cytotoxicity ranking models.
[0164] Medicinal chemistry filters were applied to remove reactive or otherwise problematic structures. The Novartis and GSK PAINS (pan-assay interfering compound) filters were used, but additional filters can also be applied (e.g. metal chelators were removed using custom substructure rules). Optionally, molecules could be scored for medicinal chemistry goodness using the QED and / or Molskill models. The molecules were then scored by Tanimoto similarity to approved antibiotics or antibiotics in clinical trials. Finally before ranking, molecules were removed when they met the following criteria from the list: any PAINS filter, Tanimoto similarity > 0.6 to an approved antibiotic.
[0165] The remaining molecules were ranked by distance from the antibiotic activity / cytotoxicity tradeoff line. This line was fit to the antibiotic and cytotoxicity scores using data from compounds that were previously tested in-house so that compounds below the line had a high cytotoxicity / antibacterial activity ratio (and were deprioritized) and compounds above the line had a favorable cytotoxicity / antibacterial activity ratio.
[0166] Then the compounds were picked as follows. The furthest remaining compound from the antibiotic activity / cytotoxicity tradeoff line was picked that is above a minimum antibiotic activity prediction. If there are molecules with Tanimoto similarity > 0.8 to the picked molecule above the tradeoff line, then the top N (typically 2) of these molecules were picked and the remaining similar compounds were removed from the list. This ensured that up to 3 similar molecules were in the list, but no more than three. If the target number of compounds were reached, or there are no compounds remaining, the search stopped. Otherwise, the process was repeated to pick additional compounds. An alternative algorithm could also be used, which does not involve removing compounds based on Tanimoto similarity. Instead, the algorithm can pick molecules until there are no molecules left, then traverse the picked list, removing the lowestDocket No.: 60134-711.601 scoring molecules similar to a picked molecule so that no more than two molecules similar to each pick are chosen.
[0167] The results show that the cytotoxicity ranking models outperform single task classification models. A multitask ranking model was trained to predict cytotoxicity. It was trained on 2M compounds across 16 datasets, including measurements on HepG2, HEK293, and three other cell lines. FIG. 5 shows that the single task models performed well on the datasets they are trained on but did not generalize as well as the multitask models to new data.Meanwhile, the ranking model performed at least as well as the single task models on the data the single task models were trained on and performed better on other datasets.
[0168] Table 1 shows a summary of the various screening models trained.Docket No.: 60134-711.601
[0169] The models can be used to rank the library compounds across multiple relevant factors. FIG. 6 shows a heat map of compounds that are ranked based on predicted antibacterial activity and predicted cytotoxicity. Compounds were selected for further screening based on a decision boundary that includes compounds that are at least 65 percentile in predicted antibacterial activity, and above the linear decision boundary that is a function of predicted cytotoxicity and predicted antibacterial activity. Compounds were chosen by maximizing the distance from the cytotoxicity-antibacterial activity tradeoff boundary. The compounds were limited in structurally similar compounds to maximize diversity of the screened compounds. The compounds were screened in E. coli IptD in M9 media in single replicates, followed by triplicate confirmation of hits. Several pre-plated sets with high fractions of predicted actives were selected.
[0170] Selected compounds were further screened with experiments to verify the hits. Of the ML-guided cherry-picking yielded twice as many hits as previous HTS#1, while selecting for low cytotoxicity and maximizing diversity. The quality of hits improved significantly, and led to fewer weak inhibitors. FIG. 7 and FIG. 8 shows the results of cherry-picking led discovery.
[0171] Table 2 shows reported results of other screening models.
[0172] While also other reported ML models identify large numbers of hits, these results are reported on testing a small number of top-ranked compounds which may increase hit rate. Additionally, they frequently include compounds highly similar to known antibiotics (e.g. 47 / 51 for Stokes). As the ranking model of the present disclosure prioritizes compound novelty, the hit rate may increase if the model was tested only with the top ranked hits.
[0173] By using cytotoxicity as a metric, about 3x reduction in cytotoxic molecules over unguided screening was observed. 13% of ML-selected hits had indications of cytotoxicity (cytotoxicity window <2), whereas 37% of non-selected hits had indications of cytotoxicity. FIG. 9 shows a distribution of cytotoxicity of selected and non-selected hits. Thus, the results show additional enrichment in progressible hits. Combined with the ~2x enrichment in hits, this yields a ~3x efficiency gain in progressible hits from ML-guided screening. Compounds thatDocket No.: 60134-711.601 passed are on average less toxic (median cytotoxicity window 10 vs 6) than those that did not pass. A larger fraction of WT-active compounds were identified in ML-guided screening (67% vs 45%).
[0174] FIG. 10 shows a tmap projection of the screened chemical space along with the hits, clinical antibiotics, and antibacterials. The chemical space covered by the hits are distinct from both clinically used antibiotics and published non-clinical antibacterials.
[0175] In previous models, it was observed that more data from diverse strains outperforms less data from identical strains. Many compounds active in wild-type (WT) E. coli are also active in permeabilized or efflux deficient strains. Table 3 (below) shows the percentage of actives that are also active in:Table 3Most (98%) of compounds are inactive in any strain. Mutant strain data improves WT activity prediction, even though most mutant-active compounds are inactive against WT.
[0177] FIG. 13 shows various datasets available for training the models. The datasets are crafted from assays measured in different strains and assay conditions, thus, it not trivial to compute an MIC or a percentage inhibition. Instead, this example demonstrates training a model that can rank each assay by activity in a manner that is consistent across the heterogenous datasets.
[0178] The trained ranking model was used to screen about 104,779 compounds to yield about 30,000 cherry picked compounds and about 70,000 compounds in pre-plated libraries. 3,376 were primary hits, and 1001 were selected for follow up studies. 287 exhibited activity in inhouse dose-response experiments. FIG. 15 shows examples of compounds that were found via the discovery process.
Claims
Docket No.: 60134-711.601CLAIMSWhat is claimed is:
1. A computer-implemented method of training a machine learning model for ranking compounds for efficacy, comprising:(a) providing a training dataset comprising data generated from assays, wherein at least two assays in the assays yield different measures of efficacy of the compounds; and(b) training the machine learning model using the training dataset, wherein the machine learning model yields a ranking of compounds based on the different measures of efficacy.
2. A computer-implemented method of ranking compounds for efficacy, comprising ranking compounds based on their measures of efficacy using a model trained using a dataset comprising data generated from assays comprising at least two assays, wherein the at least two assays yield different measures of efficacy of the compounds.
3. The computer-implemented method of claim 1, wherein the measures of efficacy comprise antibacterial activity, cytotoxicity, or both.
4. The computer-implemented method of claim 1, wherein the dataset comprises at least three, four, or five assays, that yield the different measures of efficacy.
5. The computer-implemented method of claim 4, wherein the dataset comprises at least 5, 10, 15, 20, 25, 50, 75, or 100 assays, that yield different measures of efficacy.
6. The computer-implemented method of claim 1, wherein the assays comprise antimicrobial susceptibility assays.
7. The computer-implemented method of claim 1, wherein the assays are selected from the group consisting of agar well diffusion, disk well diffusion, agar dilution, broth dilution, ATP bioluminescence, thin-layer chromatography, antimicrobial susceptibility testing, bioautography, ATP luminescence, a fluorescence assay, and any combination thereof.
8. The computer-implemented method of claim 1, wherein the at least two assays are performed using different conditions.
9. The computer-implemented method of claim 8, wherein the different conditions are selected from the group consisting of temperatures, buffers, media, oxygen concentration, and any combination thereof.Docket No.: 60134-711.60110. The computer-implemented method of claim 1, wherein the assays comprise at least two assays performed on different species of bacteria.
11. The computer-implemented method of claim 10, wherein the different strains of bacteria comprise a wild-type strain.
12. The computer-implemented method of claims 10, wherein the different strains of bacteria comprise a mutant-type strain.
13. The computer-implemented method of claim 10, wherein the different strains of bacteria comprise an efflux deficient strain.
14. The computer-implemented method of claim 10, wherein the different strains of bacteria comprise an efflux competent strain.
15. The computer-implemented method of claim 10, wherein the different strains of bacteria comprise a membrane compromised strain.
16. The computer-implemented method of claim 10, wherein the different strains of bacteria comprise a membrane intact strain.
17. The computer-implemented method of claim 10, wherein the different species of bacteria are selected from the following genera: Acinetobacter, Bacteroides, Citrobacter, Clostridioides, Corynebacterium, Enterobacter, Enterococcus, Escherichia, Eggerthella, Klebsiella, Lactobacillus, Listeria, Moraxella, Morganella, Neisseria, Peptostreptococcus, Prevotella, Propionibacterium, Providencia, Pseudomonas, Salmonella, Shigella, Staphylococcus, Stenotrophomonas, Streptococcus, Mycobacteroides, and Mycobacterium.
18. The computer-implemented method of claim 10, wherein the different strains of bacteria are selected from the following genera: Acinetobacter, Bacteroides, Citrobacter, Clostridioides, Corynebacterium, Enterobacter, Enterococcus, Escherichia, Eggerthella, Klebsiella, Lactobacillus, Listeria, Moraxella, Morganella, Neisseria, Peptostreptococcus, Prevotella, Propionibacterium, Providencia, Pseudomonas, Salmonella, Shigella, Staphylococcus, Stenotrophomonas, Streptococcus, Mycobacteroides, and Mycobacterium.
19. The computer-implemented method of claim 10, wherein the different species of bacteria are selected from the following species: Acinetobacter baumannii, Acinetobacter haemolyticus, Acinetobacter nosocomialis, Acinetobacter pittii, Acinetobacter radioresistens, Acinetobacter ursingii, Bacteroides fragilis, Citrobacter braakii,Docket No.: 60134-711.601Citrobacter freundii, Citrobacter koseri, Citrobacter sedlakii, Clostridioides difficile, Clostridioides perfringens, Cory neb acterium jeikeium, Enterobacter aerogenes, Enterobacter cloacae, Enterococcus faecalis, Enterococcus faecium, Eggerthella lenta, Escherichia coli, Klebsiella oxytoca, Klebsiella pneumoniae, Lactobacillus acidophilus, Listeria monocytogenes, Moraxella catarrhalis, Morganella morganii, Neisseria gonorrhoeae, Peptostreptococcus anaerobius, Prevotella bivia, Propionib acterium acnes, Providencia rettgeri, Providencia stuartii, Pseudomonas aeruginosa, Salmonella typhimurium, Shigella dysenteriae, Staphylococcus aureus, Staphylococcus capitis, Staphylococcus caprae, Staphylococcus epidermidis, Staphylococcus haemolyticus, Staphylococcus hominis, Staphylococcus lugdenensis, Staphylococcus pettenkoferi, Staphylococcus saprophyticus, Staphylococcus simulans, Staphylococcus wamerii, Stenotrophomonas maltophilia, Streptococcus agalactiae, Streptococcus constellatus, Streptococcus pneumoniae, Streptococcus pyogenes, Mycobacterium tuberculosis, Mycobacterium avium, Mycobacterium intracellulare, Mycobacterium smegmatis, and Mycobacteroides abscessus.
20. The computer-implemented method of claim 10, wherein the different strains of bacteria are selected from the following species: Acinetobacter baumannii, Acinetobacter haemolyticus, Acinetobacter nosocomialis, Acinetobacter pittii, Acinetobacter radioresistens, Acinetobacter ursingii, Bacteroides fragilis, Citrobacter braakii, Citrobacter freundii, Citrobacter koseri, Citrobacter sedlakii, Clostridioides difficile, Clostridioides perfringens, Cory neb acterium jeikeium, Enterobacter aerogenes, Enterobacter cloacae, Enterococcus faecalis, Enterococcus faecium, Eggerthella lenta, Escherichia coli, Klebsiella oxytoca, Klebsiella pneumoniae, Lactobacillus acidophilus, Listeria monocytogenes, Moraxella catarrhalis, Morganella morganii, Neisseria gonorrhoeae, Peptostreptococcus anaerobius, Prevotella bivia, Propionib acterium acnes, Providencia rettgeri, Providencia stuartii, Pseudomonas aeruginosa, Salmonella typhimurium, Shigella dysenteriae, Staphylococcus aureus, Staphylococcus capitis, Staphylococcus caprae, Staphylococcus epidermidis, Staphylococcus haemolyticus, Staphylococcus hominis, Staphylococcus lugdenensis, Staphylococcus pettenkoferi, Staphylococcus saprophyticus, Staphylococcus simulans, Staphylococcus wamerii, Stenotrophomonas maltophilia, Streptococcus agalactiae, Streptococcus constellatus, Streptococcus pneumoniae, Streptococcus pyogenes, Mycobacterium tuberculosis, Mycobacterium avium, Mycobacterium intracellulare, Mycobacterium smegmatis, and Mycobacteroides abscessus.Docket No.: 60134-711.60121. The computer-implemented method of claim 1, wherein the assays are performed on different cell lines.
22. The computer-implemented method of claim 21, wherein the different cell lines are selected from the group consisting of HepG2, HEK293, NIH3T3, U2OS, HaCat, CRL- 7250, or any combination thereof.
23. The computer-implemented method of claim 1, wherein the assays are performed with different endpoints.
24. The computer-implemented method of claim 23, wherein the different endpoints are selected from the group consisting of growth rate, total growth, optical density, colony forming units, relative light units.
25. The computer-implemented method of claim 1, wherein the training dataset comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 75, or 100 sources of data that are heterogenous.
26. The computer-implemented method of claim 1, wherein the compounds comprise at least 2, 5, 10, 25, 50, 75, 1000, 5000, 10000, 50000, or 100000 compounds.
27. The computer-implemented method of claim 1, wherein the machine learning model yields a ranking of compounds based on their antibacterial activity.
28. The computer-implemented method of claim 1, wherein the machine learning model yields a ranking of compounds based on their cytotoxicity.
29. The computer-implemented method of claim 1, wherein the machine learning model yields a ranking of compounds based on their antibacterial activity and cytotoxicity.
30. The computer-implemented method of claim 1, further comprising selecting a set of candidate compounds based on a threshold criteria.
31. The computer-implemented method of claim 30, wherein the threshold criteria is selected from the group consisting of a predetermined number, a predetermined percent of the compounds, an antibacterial activity threshold, a cytotoxicity threshold, and any combination thereof.
32. The computer-implemented method of claim 1, further comprising performing screening assays using the set of candidate compounds.
33. The computer-implemented method of claim 32, wherein the screening assays are selected from the group consisting of chemistry assays, biochemistry assays, bindingDocket No.: 60134-711.601 assays, functional assays, cell-based assays, in vivo assays, ADME assays, toxicology assays, and any combination thereof .
34. A computer-implemented method of screening a set of compounds comprising:(a) conducting at least two assays on each compound of the set of compounds, each assay yielding raw data indicating different measures of efficacy;(b) preparing a training dataset based on the raw data; and(c) training the machine learning model using the training dataset, wherein the machine learning model yields a ranking of compounds based on their measures of efficacy.
35. The computer-implemented method of claim 34, wherein the measures of efficacy comprise antibacterial activity, cytotoxicity, or both.
36. The computer-implemented method of claim 34, wherein the training the machine learning model is performed using a neural network, a random forest, or both.
37. The computer-implemented method of claim 34, wherein the ranking of compounds is based on their relative antibacterial activity, cytotoxicity, or both.
38. A computer-implemented system comprising: a digital processing device comprising: at least one processor, an operating system configured to perform executable instructions, a memory, and a computer program including instructions executable by the digital processing device to implement the computer-implemented method of claim 1.
39. A computer program product comprising a computer-readable medium having computerexecutable code encoded therein, the computer-executable code adapted to be executed to implement the computer-implemented method of claim 1.
40. A non-transitory computer-readable storage media encoded with a computer program including instructions executable by one or more processors to perform the computer- implemented method of claims 1.
Citation Information
Patent Citations
Antimicrobic susceptibility testing using machine learning
US20210332320A1
Artificial intelligence engine for generating candidate drugs using experimental validation and peptide drug optimization
US20220059196A1