Optimized drug screening using artificial intelligence
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-22
- Publication Date
- 2026-04-01
AI Technical Summary
Current drug development for ultra-rare cancers, such as Fibrolamellar carcinoma, faces challenges due to the lack of effective therapies, with existing methods being costly, time-consuming, and ineffective, necessitating a novel approach for rapid and cost-efficient drug screening that can identify promising candidates within a clinically relevant timeframe.
A computer-implemented method using machine learning models to predict the efficacy of drugs by clustering and analyzing functional and structural information from a diversity drug set, allowing for the selection of candidate drugs that can modulate biological responses in living systems, including cancer cells, through a principal drug set screening process.
This approach enables rapid and accurate identification of effective drugs for ultra-rare cancers, reducing the time and cost associated with traditional drug development, and provides a data-driven platform for personalized oncology and drug repurposing, effectively addressing the limitations of existing methods.
Smart Images

Figure US2024030649_28112024_PF_FP_ABST
Abstract
Description
OPTIMIZED DRUG SCREENING USING ARTIFICIAL INTELLIGENCECROSS-REFERENCE(S) TO RELATED APPLICATION S)
[0001] This application claims the benefit of Provisional Application No. 63 / 503757, filed May 23, 2023, the entire disclosure of which is hereby incorporated by reference herein for all purposes.STATEMENT OF GOVERNMENT LICENSE RIGHTS
[0002] This invention was made with Government support under FD007925 awarded by the National Institutes of Health. The Government has certain rights in the invention.BACKGROUND
[0003] Ultra-rare cancers, which have an annual incidence of fewer than 1000 people in the US, often lack effective therapy due to the many challenges involved in drug development for rare diseases. Agencies such as the FDA recognize this "unmet medical need where there is little economic incentive for commercial entities to conduct the research." Therefore, they are seeking novel yet practical strategies in drug development for these rare and heterogeneous malignancies. Even with the available tools in multi-omics analyses, the path from defining the molecular abnormality to therapy for any one cancer type can take years and millions of dollars. Hence, it is desirable for drug discovery platforms for addressing ultra-rare cancers to have one or more of the following attributes: 1) a disease model that represents the human cancer, including its tumor microenvironment, which is useful in determining response; 2) a screening strategy that makes use of limited unique human tissues yet is suitable for high-throughput analyses; 3) an approach that is agnostic to the underlying genomic alterations and generalizable for any solid tumor; and 4) a protocol to identify candidates within a clinically-relevant timeframe and in a costefficient manner.
[0004] One example of an ultra-rare cancer is Fibrolamellar carcinoma (FLC). FLC is a rare and aggressive liver cancer that primarily affects children and young adults and has ahigh mortality rate. FLC is unique in that it arises in healthy livers without any underlying liver disease or cirrhosis. FLC is characterized by a genetic abnormality, specifically a deletion in chromosome 19 that results in the fusion of DNAJB1 and PRKACA genes, leading to the production of a fusion protein that is thought to be responsible for the development of the tumor.
[0005] Currently, there are no effective treatments for FLC, and patients with the disease have a poor prognosis with a 5-year survival rate of only around 30%. Unfortunately, FLC is often not detected until it has already progressed to an advanced stage, which makes complete surgical resection impossible. In such cases, systemic therapy is the only option, but none of the drugs currently approved for hepatocellular carcinoma in adults or hepatoblastoma in children have been consistently effective against FLC. Clinical trials testing drugs that target AURKA, mTORC, or estrogen receptor have shown uniformly negative results, further highlighting the urgent need for novel and effective therapies for FLC. In this regard, FLC is one of the most refractory cancers, and new approaches are urgently needed to address this unmet clinical need.SUMMARY
[0006] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This summary is not intended to identify key features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.
[0007] In some aspects, the present disclosure provides a computer-implemented method of predicting efficacies of drugs in a diversity drug set. In general, a computing system receives test results for drugs in a principal drug set for samples of a subject. The computing system trains at least one machine learning model to generate functional predicted results of drugs in the diversity drug set based on the test results for drugs in the principal drug set, and predicts functional predicted results of drugs in the diversity drug set by providingfeatures of the drugs in the diversity drug set as input to the at least one machine learning model.
[0008] In some embodiments of the method of predicting efficacies of drugs in a diversity drug set as above, determining the principal drug set from the diversity drug set includes: determining a plurality of rough clusters for the diversity drug set based on functional information for the diversity drug set; and determining a plurality of fine clusters for the diversity drug set based on the plurality of rough clusters, the functional information for the diversity drug set, and structural information for the diversity drug set. In some variations, determining the plurality of rough clusters includes: determining functional values for each drug in the diversity drug set; and organizing the drugs from the diversity drug set into rough clusters based on one or more of a silhouette distance based on the functional values or a within-cluster sum of squares based on the functional values. In some further variations, the functional values include viability data at a single dose of the drugs.
[0009] In certain embodiments of the method of predicting efficacies of drugs in a diversity drug set as above, determining the plurality of fine clusters based on the plurality of rough clusters, the functional information for the diversity drug set, and structural information for the diversity drug set includes: for each rough cluster: dividing the rough cluster into fine clusters based on two or more of an average cosine similarity of functional values of drugs within the rough cluster, an average correlation coefficient of functional values of drugs within the rough cluster, an average binary similarity of functional values of drugs within the rough cluster, or an average Tanimoto similarity of drugs within the rough cluster. In some variations, the method further comprises, for each fine cluster, further dividing the fine cluster into smaller fine clusters until the smaller fine clusters have a combined similarity value greater than a combined similarity threshold, separate similarity values each greater than separate similarity thresholds, and a cluster size less than a cluster size threshold. In further variations, the combined similarity value is a sum of two or more of an average cosine similarity value, an average correlation coefficient value, an average binary similarity value, or an average Tanimoto similarity value. In stillfurther variations, the separate similarity thresholds include two or more of an average cosine similarity value threshold, an average correlation coefficient value threshold, an average binary similarity value threshold, or an average Tanimoto similarity value threshold. In certain embodiments, determining the principal drug set from the diversity drug set includes choosing drugs from the diversity drug set having a highest percentage of average Tanimoto similarity. In certain embodiments, determining the principal drug set from the diversity drug set includes choosing at least one drug from each fine cluster smaller than a small cluster size threshold.
[0010] In some embodiments of the method of predicting efficacies of drugs in a diversity drug set as above, training at least one machine learning model to generate functional predicted results of drugs in the diversity drug set based on the test results for the drugs in the principal drug set includes training the at least one machine learning model to accept functional information and structural information for the principal drug set as input and to generate functional predicted results that match the functional test results as output.
[0011] In some embodiments of the method of predicting efficacies of drugs in a diversity drug set as above, the at least one machine learning model includes at least one of an elastic net regularization model or a deep neural network model.
[0012] In some embodiments, a non-transitory computer-readable medium having computer-executable instructions stored thereon is provided. The instructions, in response to execution by one or more processors of a computing system, cause the computing system to perform actions of a method of predicting efficacies of drugs in a diversity drug set as above.
[0013] In some embodiments, a computing system having at least one processor and a non-transitory computer-readable medium is provided. The non-transitory computer- readable medium has computer-executable instructions stored thereon that, in response to execution by the at least one processor, cause the computing system to perform actions of a method of predicting efficacies of drugs in a diversity drug set as above.
[0014] In some aspects, the present disclosure provides a method for selecting a candidate drug for modulating (e.g., inducing, enhancing, inhibiting, or blocking) a predetermined biological response in a living system. The method generally includes the following steps: (a) testing a principal drug set on a set of samples of a living system, wherein the principal drug set is a subset of a diversity drug set, and wherein, for each drug within the principal drug set, the drug is administered to a different sample is to obtain a test result, wherein the test result is a measurement of a predetermined biological response, thereby obtaining a set of test results for the drugs in the principal drug set; (b) providing the set of test results to a computing system configured to perform a computer-implemented method for predicting efficacies of drugs in a diversity drug set as summarized above; (c) receiving predicted efficacies of drugs in the diversity drug set from the computing system; and (d) selecting, from the diversity drug set, a candidate drug for modulating the biological response in the living system based on the predicted efficacies. In certain embodiments, the method further includes selecting the principal drug set from the drugs in the diversity drug set before the testing step. In some embodiments, the biological response is selected from cell viability, cell death, cellular differentiation, a morphology change, motility, contractility, transcription factor activity, and gene expression.
[0015] In some embodiments of a method for selecting a candidate drug as above, the living system is selected from a tissue, dissociated cells from a tissue, a cell line derived from a tissue, and whole organisms. In some variations comprising a living system selected from a tissue, dissociated cells from a tissue, and a cell line derived from a tissue, the tissue is affected by a disease or disorder; in some such embodiments, the disease or disorder is a cancer and the tissue is a tumor tissue, which may be a solid tumor tissue, a liquid tumor tissue, or a soft tumor tissue. In some variations wherein the living system is a whole organism, the whole organism is selected from the group consisting of zebrafish, Caenorhabditis elegans, Xenopus. Drosophila melanogasler. and mice (Mus miisciilus .
[0016] In certain embodiments of a method for selecting a candidate drug as above, the method further includes testing the selected candidate drug on a sample of the living systemto confirm the predicted efficacy of the drug. In some variations, the selecting step includes selecting, from the diversity drug set, a plurality of candidate drugs modulating the biological response in the living system based on the predicted efficacies; in some such variations, the method includes testing each of the selected candidate drugs on samples of the living system to confirm the predicted efficacies of the drugs.
[0017] In some embodiments of a method for selecting a candidate drug as above, the method further includes identifying, based on the predicted efficacies, a class of drugs efficacious for modulating the biological response in the living system. In other, non- mutually exclusive embodiments, a plurality of drugs are identified as efficacious for modulating the biological response in the living system, and the method further includes identifying at least one signaling pathway or molecular target that is common among the plurality of identified drugs based on functional information for the diversity drug set.
[0018] In some aspects, the present disclosure provides a method for selecting a candidate drug for treatment of a disease or disorder in a subject. The method generally includes the following steps: (a) testing a principal drug set on a set of samples (e.g., tissue samples) from a subject having the disease or disorder, wherein the principal drug set is a subset of a diversity drug set, wherein the samples are affected by the disease or disorder, and wherein, for each drug within the principal drug set, a different sample is cultured in the presence of the drug to obtain a clinically relevant test result, thereby obtaining a set of test results for the drugs in the principal drug set; (b) providing the set of test results to a computing system configured to perform a computer-implemented method for predicting efficacies of drugs in a diversity drug set as summarized above; (c) receiving predicted efficacies of drugs in the diversity drug set from the computing system; and (d) selecting, from the diversity drug set, a candidate drug for treating the disease or disorder in the subject based on the predicted efficacies. In certain embodiments, the method further includes selecting the principal drug set from the drugs in the diversity drug set before the testing step.
[0019] In some embodiments of a method for selecting a candidate drug for treatment of a disease or disorder as above, the disease or disorder is a rare disease or disorder (e.g., an ultra-rare disease or disorder). In some non-mutually exclusive embodiments, the disease or disorder is a cancer, which may be a cancer characterized by a solid tumor, a liquid tumor, or a soft tissue tumor. In certain variations wherein the disease or disorder is a cancer, the samples are from a primary tumor; in some alternative embodiments, the samples are from a metastatic site. Particularly suitable samples include intact tissues samples. For example, in some variations wherein the disease or disorder is a cancer characterized by a solid tumor, the samples are intact tissue samples that preserve the native tumor microenvironment (TME). The samples may be cultured in a complex culture system such as, for example, a three-dimensional (3D) culture system.
[0020] In certain embodiments of a method for selecting a candidate drug for treatment of a disease or disorder as above, the samples are freshly obtained from the subject; in some such embodiments, the principal drug set is tested on the set of samples immediately after the samples are obtained from the subject. In some alternative variations, the samples are cryopreserved samples.
[0021] In some embodiments of a method for selecting a candidate drug for treatment of a disease or disorder as above, the method further includes removing affected tissue from the subject to obtain the set of samples. In some such embodiments, the affected tissue is removed by a needle biopsy. In other embodiments, the affected tissue is removed by resection. The method may further include cutting the removed tissue into separate sections to obtain the set of samples. In certain embodiments, the method further includes cryopreserving part of the removed tissue; for example, in variations comprising cutting the removed tissue into separate sections, the method further includes cryopreserving a subset of the separate tissue sections.
[0022] In certain variations, a method for selecting a candidate drug for treatment of a disease or disorder as above further includes testing the selected candidate drug on a sample from the subject to confirm the predicted efficacy of the drug. In some such variations, thesample used to confirm the predicted efficacy is from cryopreserved tissue. For example, in some embodiments comprising cryopreserving a subset of separate tissue sections, the method further includes testing the selected candidate drug on a sample obtained from the cryopreserved tissue sections to confirm the predicted efficacy of the drug.
[0023] In some embodiments of a method for selecting a candidate drug for treatment of a disease or disorder as above, the method further includes treating the disease or disorder in the subject, wherein the treatment comprises administering an effective amount of the selected candidate drug to the subject. In other embodiments, the subject is already undergoing a treatment for the disease or disorder comprising administration of a drug within the diversity drug set and the method is a method for monitoring treatment in the subject. In some such embodiments, the selected candidate drug is different from the drug being administered to the subject. In certain variations, the method further includes changing the treatment in the subject based on the predicted efficacies of the drugs, wherein changing the treatment comprises at least one of (i) stopping treatment with the drug being administered to the subject and (ii) administering an effective amount of the selected candidate drug to the subject.
[0024] In some embodiments of a method for selecting a candidate drug for treatment of a disease or disorder as above, the selecting step comprises selecting, from the diversity drug set, a plurality of candidate drugs for treating the disease or disorder in the subject based on the predicted efficacies. In some such variations, the method includes testing each of the selected candidate drugs on samples from the subject to confirm the predicted efficacies of the drugs; in such embodiments, the method may further include selecting one of the selected candidate drugs for treating the disease or disorder in the subject.
[0025] In certain embodiments of a method for selecting a candidate drug for treatment of a disease or disorder as above, the method further includes identifying, based on the predicted efficacies, a class of drugs efficacious for treating the disease or disorder. In other, non-mutually exclusive embodiments, a plurality of drugs are identified as efficacious for treating the disease or disorder in the subject, and the method furtherincludes identifying at least one signaling pathway or molecular target that is common among the plurality of identified drugs based on functional information for the diversity drug set.
[0026] In some aspects, the present disclosure provides a method for identifying a signaling pathway associated with a disease or disorder, or identifying a molecular target for treating the disease or disorder. The method generally includes the following steps: (a) testing a principal drug set on a set of samples (e.g., tissue samples) affected by the disease or disorder, wherein the principal drug set is a subset of a diversity drug set, and wherein, for each drug within the principal drug set, a different sample is cultured in the presence of the drug to obtain a clinically relevant test result, thereby obtaining a set of test results for the drugs in the principal drug set; (b) providing the set of test results to a computing system configured to perform a computer-implemented method for predicting efficacies of drugs in a diversity drug set as summarized above; (c) receiving predicted efficacies of drugs in the diversity drug set from the computing system; (d) selecting, from the diversity drug set, a plurality of drugs that are identified as efficacious for treating the disease or disorder based on the predicted efficacies; and (e) identifying at least one signaling pathway or molecular target that is common among the plurality of identified drugs based on functional information for the diversity drug set, thereby identifying a signaling pathway associated with the disease or disorder or identifying a molecular target for treating the disease or disorder. In certain embodiments, the method further includes selecting the principal drug set from the drugs in the diversity drug set before the testing step.
[0027] In some embodiments of a method for identifying a signaling pathway or molecular target as above, the disease or disorder is a rare disease or disorder (e.g. , an ultra- rare disease or disorder). In some non-mutually exclusive embodiments, the disease or disorder is a cancer, which may be a cancer characterized by a solid tumor, a liquid tumor, or a soft tissue tumor. In certain variations wherein the disease or disorder is a cancer, the samples are from a primary tumor; in some alternative embodiments, the samples are from a metastatic site. Particularly suitable samples include intact tissues samples. For example,in some variations wherein the disease or disorder is a cancer characterized by a solid tumor, the tissue samples are intact tissue samples that preserve the native tumor microenvironment (TME). The tissue samples may be cultured in a complex culture system such as, for example, a three-dimensional (3D) culture system.
[0028] In certain embodiments of a method for identifying a signaling pathway or molecular target as above, the samples are freshly obtained from one or more subjects having the disease or disorder; in some such embodiments, the principal drug set is tested on the set of samples immediately after the samples are obtained from the one or more subjects. In some alternative variations, the samples are cryopreserved samples.
[0029] In some embodiments of a method for identifying a signaling pathway or molecular target as above, the method further includes removing affected tissue from one or more subjects to obtain the set of samples. In some such embodiments, the affected tissue is removed by a needle biopsy. In other embodiments, the affected tissue is removed by resection. The method may further include cutting the removed tissue into separate sections to obtain the set of samples. In certain embodiments, the method further includes cryopreserving part of the removed tissue; for example, in variations comprising cutting the removed tissue into separate sections, the method further includes cryopreserving a subset of the separate tissue sections.
[0030] In some embodiments of a method for identifying a signaling pathway or molecular target as above, the method further includes testing the selected drugs on a second set of samples to confirm the predicted efficacies of the drug; in some such embodiments, the samples used to confirm the predicted efficacies are from cryopreserved tissue.
[0031] In certain variations of a method for identifying a signaling pathway or molecular target as above, the method further includes identifying one or more additional drugs targeting the identified signaling pathway or molecular target. In some such variations, the method further includes testing the one or more additional drugs in a physiologically relevant assay or model to evaluate efficacy in treating the disease or disorder.
[0032] In some aspects, the present disclosure provides a method for identifying a drug that modulates a predetermined signaling pathway associated with a disease or disorder. The method generally includes the following steps: (a) testing a principal drug set on a set of cells (e.g., cells from a cell line) capable of signal transduction through a predetermined signaling pathway, wherein the principal drug set is a subset of a diversity drug set, wherein, for each drug within the principal drug set, a different sample of cells is cultured in the presence of the drug, and wherein a cellular phenotype associated with the signaling pathway is detected or measured to obtain a test result, thereby obtaining a set of test results for the drugs in the principal drug set; (b) providing the set of test results to a computing system configured to perform a computer-implemented method for predicting efficacies of drugs in a diversity drug set as summarized above; (c) receiving predicted efficacies of drugs in the diversity drug set from the computing system; and (d) selecting, from the diversity drug set, a drug that is identified as efficacious for modulating the signaling pathway based on the predicted efficacies. In some embodiments, the identified drug is identified as an activator of the signaling pathway. In other embodiments of the method, the identified drug is identified as an inhibitor of the signaling pathway. In yet other, non- mutually exclusive embodiments, the cellular phenotype is selected from the group consisting of cell growth, cell death, cellular differentiation, and a change in transcriptional activity. In certain embodiments, the method further includes selecting the principal drug set from the drugs in the diversity drug set before the testing step.
[0033] In some embodiments of a method for identifying a drug that modulates a predetermined signaling pathway as above, a plurality of drugs are identified as efficacious for modulating the signaling pathway. In other, non-mutually exclusive embodiments, the cells are cultured with the principal drug set in the presence of a natural ligand that initiates signal transduction through the signaling pathway.
[0034] These and other aspects of the invention will become evident upon reference to the following detailed description of the invention.BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The foregoing aspects and many of the attendant advantages of this invention will become more readily appreciated as the same become better understood by reference to the following detailed description, when taken in conjunction with the accompanying drawings, wherein:
[0036] FIG. 1 is a block diagram that illustrates aspects of a non-limiting example embodiment of a prediction computing system according to various aspects of the present disclosure.
[0037] FIG. 2 is a flowchart that illustrates a non-limiting example embodiment of a method of drug screening according to various aspects of the present disclosure.
[0038] FIG. 3 is a flowchart that illustrates a non-limiting example embodiment of a procedure for deriving a principal drug set from a diversity drug set based on functional information and structural information, according to various aspects of the present disclosure.
[0039] FIG. 4 is a flowchart that illustrates a non-limiting example embodiment of a procedure for training a predictive model based on functional test results and features of a principal drug set, according to various aspects of the present disclosure.
[0040] FIG. 5 depicts an embodiment of an Al-based drug screening platform in accordance with the present disclosure. The technology platform utilizes a patient-centric approach using patient tissue samples (e.g., resected tumor or biopsies) to quickly identify potential therapies.
[0041] FIG. 6A and FIG. 6B depict an exemplary diversity set drug library. FIG. 6A is a pie chart showing distribution of drugs in the diversity set library and their clinical or preclinical annotations. FIG. 6B shows STRING cluster maps showing interactions of 1679 proteins which are direct targets of diversity set drug library. The functional enrichment of these 1679 proteins showed significant enrichment of 252 KEGG and 757 Reactome pathways.
[0042] FIG. 7 depicts an exemplary derivation of a principal drug set. The data set consisting of the response to 3897 drugs from a diversity drug set in 475 cancer cell lines was subjected to hierarchical clustering followed by sub-clustering. Both structural information based on the Tanimoto similarity (TS) and functional information based on the cosine similarity (CS) of drug response profile was considered for sub-clustering. The final 239 subclusters mapped to 254 drugs. Distribution plots show the increase in average CS and average TS in the final 239 clusters compared to the initial 31 clusters.
[0043] FIG. 8 A and FIG. 8B depicts UMAP projections of a principal drug set and the diversity drug set from which the principal set was derived. The principal drug set captures the diverse functional responses (FIG. 8A) and chemical diversity (FIG. 8B) of the full diversity set drug library.
[0044] FIG. 9 shows accurate prediction of responses to a full diversity drug set in 26 breast cancer cell lines using a combination of CNN models and screening with a principal drug. Responses to the principal drug set in 449 cancer cell lines (excluding the breast cancer cell lines) were used as the “input,” while the responses in 26 breast cancer cell lines were used as the “output.” Non-linear Convolutional Neural Network (CNN) was applied to predict responses to the full 3897 diversity drug set and accurately predicted the response to the full diversity drug sets with >0.8 Pearson correlation between predicted and observed responses (“Prediction vs Experimental Correlation).
[0045] FIG. 10A and FIG. 10B show accurate prediction of responses to a full diversity drug set in 11 independent cancer cell lines using CNN models. Responses to a 254 principal drug set in 475 cancer cell lines were used as the input, while the responses in 11 cancer cell lines were used as the output. FIG. 10A shows performance of the CNN models in the training and validation sets. Validation in the independent cancer cell lines shows that the CNN models achieved a Pearson correlation of >0.8 between predicted and observed responses.
[0046] FIG. 11A to FIG. HE show identification and validation of inhibitors and activators of the Wnt pathway using a principal drug set to predict cellular responses to afull diversity set. FIG. 11A is a schematic showing the Wnt-dependent activation of P- catenin-TCF. FIG. 11B is an illustration showing the overall experimental design using HEK293 super topflash cells (STF). Cells were treated with the principal drug set in the presence of Wnt3a, and TCF activity was measured 24 hours later. The resulting data was used for CNN modeling. FIG. 11C is a plot showing the accuracy of the CNN model in predicting response to the training dataset (254 drugs). FIG. 1 ID shows predicted TCF activity of the full diversity set library. The drugs that significantly increased or decreased Wnt-dependent TCF activity are indicated with dashed-line boxes to the right (increased activity) and left (decreased activity). FIG. 1 IE shows experimental validation of model- predicted drug which increased or decreased Wnt signaling.
[0047] FIG. 12A to FIG. 12F shows successful identification of effective inhibitors in fibrolamellar carcinoma (FLC) tumor slices using principal drug-based screening. FIG. 12A is a schematic showing the overall experimental strategy for screening in FLC primary and metastatic cells and tissue slices. FIG. 12B are plots showing the accuracy of convolutional neural network (CNN) models for each FLC case. MSE, mean squared error, was used as the measure of accuracy. FIG. 12C shows prediction of the response to the full diversity set library using the optimized CNN models. Black boxes indicate predicted response, while clear (no fill) boxes indicate measured response (based on the training set of 254 drugs). FIG. 12D is an Euler plot showing the overlap of predicted drugs in each of the four FLC cases. Overall, 56 drugs that were effective in all four FLC cases tested were identified. FIG. 12E is a graph showing that all four FLC samples tested had fewer effective drugs compared to the number of effective drugs in all HCC cell lines in the dataset. FIG. 12F is a correlation matrix indicating that the FLC samples form a distinct cluster and differ from the HCC samples.
[0048] FIG. 13 A and FIG. 13B shows identification of several effective drugs for potentially treating FLC. FIG. 13A is a plot showing the predicted response to the top 25 drugs in each of four FLC cases. The plot also lists the “key targets” of these drugs, including PI3K, HDAC, Retinoic acid (RA), HSP90, PLK1, VDAC, Proteasome, TOPI,and BIRC5. FIG. 13B is a plot showing that several inhibitors from the HD AC, HSP90, and PLK1 drug classes are predicted to be effective in all four FLC cases.
[0049] FIG. 14A to FIG. 14D shows selective sensitivity of FLC tumor slices and cells to PLK1 and HSP90 inhibition. FIG. 14A shows viability of human FLC tumor and nontumor liver slices treated with DMSO control and indicated PLK1 inhibitors at 500 nM. Bars represent the mean of two independent slices, and error bars denote SEM. ** p<0.01. FIG. 14B is a dose-response plot showing patient-derived primary FLC cells are sensitive to indicated PLK1 inhibitors. Onvansertib and CYC140 are clinical-grade PLK1 inhibitors. FIG. 14C shows viability of human FLC tumor and non-tumor liver slices treated with DMSO control and indicated HSP90 inhibitors at 1000 nM. Bars represent the mean of two independent slices, and error bars denote SEM. ** p<0.01, * p<0.05. FIG. 14D shows a dose-response plot showing patient-derived FLC cells are sensitive to indicated HSP90 inhibitors.
[0050] FIG. 15 is a plot showing the predicted response to 18 Aurora kinase inhibitors in four indicated FLC cases. ENMD-2076, which failed in the clinic, is also indicated.
[0051] FIG. 16A to FIG. 16D show drug screening in needle biopsy samples. FIG. 16A is a schematic illustrating the process of cutting 18-gauge core needle biopsy samples into 3D cuboids using a tissue chopper and placing them in U-bottom 96-well plates with viability dye. FIG. 16B is a heatmap showing the viability signal after 24 hours, revealing that approximately 92% of wells in two independent 96-well plates exhibit high viability signal. FIG. 16C and FIG. 16D show the response to standard-of-care drugs measured in 3D cuboids generated from liver tumor (FIG. 16C) or cholangiocarcinoma sample (FIG. 16D), with each bar representing the mean of at least 2-3 wells. Each well contained at least three cuboids. Error bars represent SEM.
[0052] FIG. 17A to FIG. 17C shows prediction of responses to a full diversity drug set using an HCC-specific 42 drug set. FIG. 17A is a UMAP showing 300 effective drugs identified from Diversity set screening in 18 HCC cell lines. FIG. 17B is a UMAP showing a smaller set of 42 HCC-specific drugs that capture the “effective drug” pattern. FIG. 17Cis a plot showing the performance of the 254 principal set and 42 HCC-specific drug set in predicting the response to the full 3,897 drugs.
[0053] FIG. 18 is a schematic illustration of a DeepGeneX framework for Al-based prediction of drug responses (see Example 10, infra).
[0054] FIG. 19 shows improvement of model performance in predicting response to 254 drugs using DeepGeneX. The plot shows the prediction accuracy of CNN building from the full gene set (left) vs. 3000 genes (right).
[0055] FIG. 20 is a bar graph showing confirmation of the efficacy of drugs identified for meningioma treatment using Al-based drug screening in accordance with the present disclosure (see Example 11).DETAILED DESCRIPTION
[0056] The present disclosure describes techniques to identify drugs modulating biological responses in living systems, including drugs for treatment of various diseases and disorders, using a data-driven drug response prediction tool. The methods described herein not only transform the treatment of diseases and disorders but also facilitates the discovery of biology (including, e.g., molecular targets and signaling pathways) underlying responses to different drugs. The methods are particularly useful for identifying drugs to treat rare and ultra-rare diseases and interrogating mechanisms of action of such drugs.
[0057] For example, the present disclosure describes techniques to transform the treatment of ultra-rare cancers by providing a clinically implementable, data-driven drug response prediction tool to inform oncologists of candidate drugs that the tumor is likely to respond to. The techniques described herein provide a new approach to tackle the challenges associated with ultra-rare cancers by utilizing an Al-driven systems- pharmacology-based platform. This technology directly tests patient-derived tissues that best represent the original cancer and bypasses the latency for developing organoids or PDX models for drug screening. A 3D organotypic culture protocol allows monitoring of responses to a subset of compounds which captures the functional and structural diversityof a much larger collection of FDA-approved and clinical-grade compounds. The response data for this subset is then used in a machine-learning algorithm that relies on limited drug profiling information to learn and refine drug response predictions. The deep neural networks create accurate and rapid predictions of drug responses. By directly interrogating the cancer of interest, this approach is able to assist in prioritizing drugs or drug combinations from the entire FDA-approved collection, providing a unique advantage in tackling ultra-rare cancers. This platform has the predictive power to provide rapid and accurate treatment predictions for ultra-rare cancers, transforming the field of drug discovery.
[0058] The presently disclosed techniques may be used to find selective vulnerabilities of ultra-rare cancers from a collection of FDA-approved and clinical-grade compounds. Upon validation of candidates in preclinical models, therapeutic discoveries may be quickly transitioned to clinical trials. These techniques may be used to shed light on 'chemoresistant' diseases like FLC by exposing the underlying pathways associated with the known drug targets.
[0059] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art pertinent to the methods and compositions described. As used herein, the following terms and phrases have the meanings ascribed to them unless specified otherwise.
[0060] The terms “a,” “an,” and “the” include plural referents, unless the context clearly indicates otherwise.
[0061] As used herein, the term “drug” refers to an agent, comprising a molecule or complex of molecules, that is being or is to be tested for its efficacy in a biological assay in a living system. A drug may be a therapeutic agent used in the prevention, diagnosis, alleviation, treatment, or cure of a disease or disorder. The drugs can be virtually any agent — whether synthetic, recombinant, or naturally occurring — that may be tested in such biological assays, including, without limitation, small organic drugs, peptides, proteins (e.g., antibodies), lipids, polysaccharides, nucleic acids, and combinations thereof (e.g.,conjugates such as, for example, an antibody-drug conjugate). In typical variations, drugs for use in accordance with the present disclosure are associated with relevant functional information and structural information in a data store. In some such variations, a drug is an FDA approved drug, a clinical trial drug, a clinically withdrawn drug (e.g., with a known clinical safety profile), a preclinical drug, or a drug with reproducible and extensive profiling in biological models (e.g., animal disease models). In some embodiments, a drug is a known drug (e.g., FDA approved, withdrawn, or in any of various stages of clinical or preclinical testing).
[0062] As used herein, the term “diversity drug set” or “diversity set” means any large set of drugs or drug library (e.g., at least about 1,000, at least about 2,000, at least about 3,000 drugs, or at least about 4,000 drugs) from which a subset of the drugs (“principal drug set” or “principal set”) has been selected for testing in accordance with the present disclosure. A “principal drug set” encapsulates sufficient diversity of drugs and mechanisms with the diversity drug set to serve as a basis for combining biological assay screening with machine learning models to predict responses to the full panel of drugs within the diversity set.
[0063] As used herein, “testing,” in reference to a principal drug set or subset thereof, means testing of a plurality of individual drugs within the principal drug set or subset thereof on samples for their efficacy in modulating (e.g., inducing, enhancing, inhibiting, or blocking) a biological result of interest (e.g., efficacy in a tissue sample or on cells to induce a clinically relevant result that indicates efficacy in treating an associated disease or disorder, or activating or inhibiting a known signaling pathway associated with a disease or disorder).
[0064] The term “living system,” as used herein, means any biological system comprising one or more living cells and in which a physiologically relevant biological response can be measured. A living system can be for example, any type of tissue (solid, liquid, or soft, including, e.g., tumor tissue), dissociated cells from any tissue type (e.g., dissociated cellsfrom a tumor tissue), a cell line derived from any tissue type (e.g., a cell line derived from a tumor), or a whole organism.
[0065] The term “disease or disorder,” as used herein, generally refers to an impairment of health or a condition of abnormal functioning. Diseases or disorders may include, for example, cancers, cardiovascular diseases, inflammatory diseases, autoimmune diseases, metabolic disease, neurological (e.g., neurodegenerative) diseases, and infectious diseases, to name a few.
[0066] The term “rare disease or disorder,” as used herein, refers to a disease or disorder with an annual incidence of fewer than 200,000 people in the U.S. In some embodiments, a rare disease or disorder is an “ultra-rare disease or disorder,” which is defined herein as a disease or disorder with an annual incidence of fewer than 1,000 people in the U.S.
[0067] The term “subject” or “patient” are used interchangeably to refer to an individual (e.g., human) from which samples have been removed for testing of drugs in accordance with the present disclosure.
[0068] “Sample,” as used herein, refers to a quantity of cellular material for testing in a biological assay in accordance with the present disclosure. A sample can include, e.g., cells from cell lines or cells removed directly from a subject (whether fresh or cryopreserved), whether cells organized as a tissue or cells dissociated from a tissue. In some variations, a sample is a “tissue sample,” which may include any tissue (e.g., epithelial, muscle, nervous, and / or connective tissue) removed from a subject, including tissues forming organs. In addition to “solid” tissues having distinct structures (e.g., solid tumors), a tissue sample may include “fluid” or “liquid” tissues such as, e.g., blood or lymph, or “soft” tissues such as, e.g., soft tissue tumors. In some variations, a tissue sample is an “intact” tissue sample, i.e., retaining the native structure of the tissue (but allowing for dividing of the tissue into smaller subsections such as tissue slices or cuboids, such as, e.g., using a tissue slicer). In other embodiments, a sample is a whole organism, such as, e.g., zebrafish, Caenorhabditis elegans, Xenopus, Drosophila melanogaster , or a mouse (Mus musculus).
[0069] As used herein, the term “biological response,” in reference to a living system, means a physiologically relevant response to be measured in the living system. A biological response can be any measurable phenotypic readout, whether at, e.g., a molecular, cellular, sub-cellular, inter-cellular, or organismal level.
[0070] As used herein, the term “complex culture system” refers to either (i) a three- dimensional (3D) culture system or (ii) a two-dimensional (2D) culture system that involves the use of complex exogenous protein supplements.
[0071] As used herein, the term “three-dimensional (3D) culture system” refers to a culture system that allows a cultured tissue sample or a plurality of cultured cells to maintain or form three-dimensional multicellular structures. Examples of 3D culture systems include cultures in a concentrated medium or in a gel-like substance (e.g., a hydrogel), cultures on a scaffold, (e.g., an extracellular matrix (ECM) scaffold or a fibrous scaffold derived from a natural or synthetic polymer), and 3D suspension culture systems.
[0072] The term “treat” or “treating” includes abrogating, substantially inhibiting, slowing, or reversing the progression of a disease or disorder, substantially ameliorating one or more symptoms of a disease or disorder, or substantially preventing the appearance of one or more symptoms of a disease or disorder. Treating further refers to accomplishing one or more of the following: (a) reducing the severity of a disease or disorder; (b) limiting the development of symptoms characteristic of a disease or disorder being treated; (c) limiting the worsening of symptoms characteristic of a disease or disorder being treated; (d) limiting the recurrence of a disease or disorder in patients that previously had the disorder; and (e) limiting recurrence of one or more symptoms in patients that were previously symptomatic for the disease or disorder.
[0073] The term “effective amount” refers to the amount necessary or sufficient to realize a desired biologic effect.
[0074] FIG. 1 is a block diagram that illustrates aspects of a non-limiting example embodiment of a prediction computing system according to various aspects of the presentdisclosure. The illustrated prediction computing system 102 may be implemented by any computing device or collection of computing devices, including but not limited to a desktop computing device, a laptop computing device, a mobile computing device, a server computing device, a computing device of a cloud computing system, and / or combinations thereof. In some embodiments, the prediction computing system 102 is configured to support the determination of principal drug sets. In some embodiments, the prediction computing system 102 is also configured (or is instead configured) to train machine learning models based on test results for a principal drug set to allow results to be predicted for a diversity drug set in order to efficiently select one or more candidate drugs for modulating (e.g., inducing, enhancing, inhibiting, or blocking) a predetermined biological response in a living system (e.g., a candidate drug to treat a disease or disorder).
[0075] As shown, the prediction computing system 102 includes one or more processors 104, one or more communication interfaces 106, a drug data store 110, a result data store 118, a principal data store 114 a model data store 116, and a computer-readable medium 108.
[0076] In some embodiments, the processors 104 may include any suitable type of general-purpose computer processor. In some embodiments, the processors 104 may include one or more special-purpose computer processors or Al accelerators optimized for specific computing tasks, including but not limited to graphical processing units (GPUs), vision processing units (VPUs), and tensor processing units (TPUs).
[0077] In some embodiments, the communication interfaces 106 include one or more hardware and or software interfaces suitable for providing communication links between components. The communication interfaces 106 may support one or more wired communication technologies (including but not limited to Ethernet, FireWire, and USB), one or more wireless communication technologies (including but not limited to Wi-Fi, WiMAX, Bluetooth, 2G, 3G, 4G, 5G, and LTE), and / or combinations thereof.
[0078] As shown, the computer-readable medium 108 has stored thereon logic that, in response to execution by the one or more processors 104, cause the prediction computingsystem 102 to provide a result collection engine 112, a principal set determination engine 120, a model training engine 122, and a result prediction engine 124.
[0079] As used herein, "computer-readable medium" refers to a removable or nonremovable device that implements any technology capable of storing information in a volatile or non-volatile manner to be read by a processor of a computing device, including but not limited to: a hard drive; a flash memory; a solid state drive; random-access memory (RAM); read-only memory (ROM); a CD-ROM, a DVD, or other disk storage; a magnetic cassette; a magnetic tape; and a magnetic disk storage.
[0080] In some embodiments, the principal set determination engine 120 is configured to determine a principal drug set that is representative of clusters of drugs in a diversity drug set, and to store the principal drug set in the principal data store 114. In some embodiments, the principal set determination engine 120 may retrieve information from the drug data store 110 in order to determine the principal drug set. In some embodiments, the result collection engine 112 is configured to receive results from tests of a biological response of a living system to the drugs in the principal drug set, and to store the results in the result data store 118. In some embodiments, the result collection engine 112 is configured to use the results of testing the drugs in the principal drug set to train a machine learning model that can predict results of testing the biological response of the living system to drugs in the diversity drug set, and to store the trained model in the model data store 116. In some embodiments, the result prediction engine 124 is configured to use the trained model from the model data store 116 to predict the results of testing the biological response of the living system to the drugs in the diversity drug set, and to thereby allow a drug from the diversity drug set to be selected for further testing or administration.
[0081] Further description of the configuration of each of these components is provided below.
[0082] As used herein, "engine" refers to logic embodied in hardware or software instructions, which can be written in one or more programming languages, including but not limited to C, C++, C#, COBOL, JAVA™, PHP, Perl, HTML, CSS, JavaScript,VBScript, ASPX, Go, and Python. An engine may be compiled into executable programs or written in interpreted programming languages. Software engines may be callable from other engines or from themselves. Generally, the engines described herein refer to logical modules that can be merged with other engines, or can be divided into sub-engines. The engines can be implemented by logic stored in any type of computer-readable medium or computer storage device and be stored on and executed by one or more general purpose computers, thus creating a special purpose computer configured to provide the engine or the functionality thereof. The engines can be implemented by logic programmed into an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or another hardware device.
[0083] As used herein, "data store" refers to any suitable device configured to store data for access by a computing device. One example of a data store is a highly reliable, highspeed relational database management system (DBMS) executing on one or more computing devices and accessible over a high-speed network. Another example of a data store is a key-value store. However, any other suitable storage technique and / or device capable of quickly and reliably providing the stored data in response to queries may be used, and the computing device may be accessible locally instead of over a network, or may be provided as a cloud-based service. A data store may also include data stored in an organized manner on a computer-readable storage medium, such as a hard disk drive, a flash memory, RAM, ROM, or any other type of computer-readable storage medium. One of ordinary skill in the art will recognize that separate data stores described herein may be combined into a single data store, and / or a single data store described herein may be separated into multiple data stores, without departing from the scope of the present disclosure.
[0084] FIG. 2 is a flowchart that illustrates a non-limiting example embodiment of a method of drug screening according to various aspects of the present disclosure. In the method 200, a principal drug set is used to represent clusters of drugs in a larger diversitydrug set, and test results for the principal drug set are used to predict the functional results of testing the drugs in the diversity drug set.
[0085] From a start block, the method 200 proceeds to optional subroutine block 202, where a subroutine is executed wherein a prediction computing system 102 derives a principal drug set from a diversity drug set based on functional information and structural information. The actions of optional subroutine block 202 are illustrates as optional because in some embodiments, the principal drug set may have previously been derived, and the method 200 may simply load the principal drug set from the principal data store 114 instead of re-deriving the principal drug set by executing the subroutine. Any suitable subroutine may be executed at optional subroutine block 202, including but not limited to the subroutine 300 illustrated in FIG. 3 and discussed in further detail below.
[0086] At block 204, a set of samples are obtained. As discussed elsewhere herein, samples may include any type of tissue (e.g., freshly removed from a subject or cryopreserved), dissociated cells from a tissue, cell lines, and / or whole organisms. Samples may be created by removing tissue from one or more subjects (e.g., by surgical biopsy, needle biopsy, etc.), and may be sectioned into tissue slices or cuboids to create the samples. Further details regarding techniques for obtaining samples are discussed elsewhere herein.
[0087] At block 206, one or more samples of the set of samples are tested using the principal drug set to obtain functional test results. Each drug in the principal drug set is administered to one or more of the samples, such that a functional test result may be obtained for each drug. In some embodiments, the functional test results include measurements of a biological response (e.g., cell viability, cell proliferation, cell death, cellular differentiation, chemotaxis, morphology change, motility, contractility, transcription factor activity, gene expression, protein secretion, etc.) of the samples to the drugs of the principal drug set. As discussed elsewhere herein, administering a drug to a sample may include culturing the sample in the presence of the drug, delivering the drugto a target tissue, or any other suitable administration technique from which the biological response to the drug may be determined.
[0088] At block 208, a result collection engine 112 of the prediction computing system 102 receives the functional test results and stores the functional test results in a result data store 118 of the prediction computing system 102. In some embodiments, one or more sensors may be used to collect the functional test results, and the one or more sensors may transmit the functional test results directly to the result collection engine 112. In some embodiments, the result collection engine 112 may present a user interface into which an operator may enter the functional test results. In some embodiments, the result collection engine 112 may ingest one or more data files or query one or more data stores in which the functional test results are stored upon generation in order to receive the functional test results.
[0089] The method 200 then advances to subroutine block 210, where a subroutine is executed wherein the prediction computing system 102 trains a predictive model based on the functional test results and features of the drugs of the principal drug set. In some embodiments, the predictive model is a machine learning model that accepts features of a drug as input and outputs a set of functional predicted results as output. Any suitable features of a drug may be used as the input, including but not limited to functional test results collected for the drugs during previous testing. As a non-limiting example, cell line viability data at a given dosage for a plurality of previously studied cell lines may be used as functional test results to be provided as features of the drugs for input to the predictive model. The predictive model is trained to minimize differences between functional predicted results output by the model and the functional test results received at block 208. Any suitable architecture may be used for the predictive model, including but not limited to an elastic net regularization model or an artificial neural network model (e.g., a nonlinear convolutional neural network model, a deep neural network model, etc.). Any suitable subroutine may be used to train the predictive model, including but not limited to the subroutine 400 illustrated in FIG. 4 and described in further detail below.
[0090] At block 212, the prediction computing system 102 provides features of the diversity drug set as input to the predictive model to generate functional predicted results for the drugs of the diversity drug set. The features of the drugs provided are similar to features of the principal drug set used to train the predictive model, and may include functional test results collected for the drugs of the diversity drug set during previous testing (including but not limited to cell line viability data at a given dosage for the plurality of previously studied cell lines). The functional predicted results include predicted biological responses of the sample to each drug of the diversity drug set. These functional predicted results may be reliable enough to serve in the place of obtaining actual functional test results for all of the drugs of the diversity drug set, thus expanding the scope of the screening to the potentially thousands of drugs in the diversity drug set without requiring that all of the drugs in the diversity drug set actually be tested.
[0091] At block 214, the prediction computing system 102 enables selection of one or more recommended drugs from the diversity drug set based on the functional predicted results. In some embodiments, the prediction computing system 102 may present the functional predicted results to allow an operator to select one or more recommended drugs. In some embodiments, the prediction computing system 102 may automatically select one or more recommended drugs based on the functional predicted results (e.g., one or more of the drugs predicted to have the best functional performance).
[0092] Once the one or more recommended drugs are selected, they may be used for any purpose. The method 200 illustrates three non-limiting and non-exclusive optional actions that may be taken once the one or more recommended drugs are selected in optional block 216, optional block 218, and optional block 220.
[0093] At optional block 216, the prediction computing system 102 presents the one or more recommended drugs to an operator. The operator may use the presented one or more recommended drugs for any purpose, or may investigate characteristics of the one or more recommended drugs to determine any commonalities or other characteristics that suggest directions for further investigation.
[0094] At optional block 218, the one or more recommended drugs are tested on one or more samples of the set of samples. In some embodiments, the functional test results obtained by testing the one or more recommended drugs may be added to the data for principal drug set and used to re-train the predictive model in order to generate a subsequent set of one or more recommended drugs.
[0095] At optional block 220, the one or more recommended drugs are administered. As described elsewhere herein, one or more of the one or more recommended drugs may be administered to a subject as a treatment for a disease or disorder.
[0096] The method 200 then proceeds to an end block and terminates.
[0097] FIG. 3 is a flowchart that illustrates a non-limiting example embodiment of a procedure for deriving a principal drug set from a diversity drug set based on functional information and structural information, according to various aspects of the present disclosure. The subroutine 300 is a non-limiting example of a subroutine suitable for use at optional subroutine block 202 of FIG. 2. In the subroutine 300, clustering techniques are used to determine a principal drug set that optimally represents functional and structural characteristics of a larger diversity drug set.
[0098] From a start block, the subroutine 300 proceeds to block 302, where a principal set determination engine 120 of the prediction computing system 102 retrieves information for the diversity drug set from a drug data store 110 of the prediction computing system 102, the information including functional information and structural information. Any suitable types of functional information and structural information may be used. In some embodiments, the functional information may include cell line viability data previously collected for the drugs of the diversity drug set at a given dosage (e.g., 2.5 pM). As many drugs have been extensively profiled across a variety of cell lines, including but not limited to cancer cell lines, it is likely that such previously collected cell line viability data will be available in the drug data store 110. In some embodiments, the structural information may include information regarding the chemical structures of the drugs. In some embodiments, the structural information may include binarized molecular fingerprints, including but notlimited to fingerprints generated from the SMILES strings of each drug using, for example, the Morgan algorithm (see Morgan, H.L., J. Chem. Doc. 5(2), 107-113, 1965).
[0099] At block 304, the principal set determination engine 120 determines a plurality of rough clusters for the diversity drug set based on the functional information. Any suitable unsupervised hierarchical clustering technique may be used. In some embodiments, two metrics may be used to evaluate the rough clusters: a silhouette distance (SD), which measures the similarity between the functional information for each drug and the other drugs in their assigned cluster compared to drugs in other clusters; and a within-cluster- sum of squares (WSS), which measures a sum of squared distance between the functional information of the drugs in a cluster and the centroid of that cluster. Any suitable technique may be used to determine the number of rough clusters. In some embodiments, an optimal number of rough clusters may be determined based on the SD reaching a local maximum and a slope of the WSS curve becoming level.
[0100] For large diversity drug sets, it is typical that once the rough clusters are determined, the cluster size is still too large for representative drugs to be chosen from each rough cluster. If too many drugs are present in each rough cluster, it is likely that functional differences will remain between the drugs in each cluster. Accordingly, the rough clusters may be further refined into fine clusters. At block 306, the principal set determination engine 120 determines a plurality of fine clusters for the diversity drug set based on the rough clusters, the functional information, and the structural information.
[0101] Any suitable unsupervised hierarchical clustering technique with any suitable metrics may be applied to the rough clusters to sub-cluster the rough clusters into fine clusters. In some embodiments, one or more of an average cosine similarity of functional values of drugs within the rough clusters, an average correlation coefficient of functional values of drugs within the rough clusters, an average binary similarity of functional values of drugs within the rough clusters, or an average Tanimoto similarity may be used to form the fine clusters. The cosine similarity is determined by computing a cosine of the angle between the functional information of the drugs (e.g., vectors of cell line viability profiles),such that drugs with similar functional information will have a higher average cosine similarity. The Tanimoto similarity uses the structural information to measure the similarity between two drugs. For example, in embodiments wherein the structural information includes binarized molecular fingerprints generated from the SMILES strings of each drug using the Morgan algorithm, the Tanimoto similarity calculates the ratio of shared bit positions between the binarized molecular fingerprints of the drugs. Drugs with similar chemical structures will have higher scores based on the Tanimoto similarity. Any suitable binary similarity or correlation coefficient determinations may be used. The averages of these metrics may be defined as average values for the metric between a given drug and all other drugs in its cluster.
[0102] To use these metrics to determine the fine clusters, averages of the metrics across the clusters may be used, which may be defined as the mean values of the average metric values for all drugs within a cluster. An iterative sub-clustering process may be used to form smaller fine clusters from the fine clusters until one or more cluster thresholds are met. In some embodiments, the one or more cluster thresholds may include having a combined similarity value greater than a combined similarity threshold, separate similarity values each greater than separate similarity thresholds, and a cluster size less than a cluster size threshold. The combined similarity value may be a sum of two or more of an average cosine similarity value, an average correlation coefficient value, an average binary similarity value, or an average Tanimoto similarity value. The separate similarity thresholds may include two or more of an average cosine similarity value threshold, an average correlation coefficient value threshold, an average binary similarity value threshold, or an average Tanimoto similarity value threshold.
[0103] At block 308, the principal set determination engine 120 selects the principal drug set based on the plurality of fine clusters. Once the plurality of fine clusters is determined, any suitable technique may be used to select the principal drug set from the plurality of fine clusters. As one non-limiting example, drugs within a top percentile range of average similarity scores (e.g., drugs within a top 5% of the highest average Tanimoto similarity)may be chosen for the principal drug set, as these drugs are likely to be the most representative of their clusters by virtue of being the most structurally similar to the other drugs in their clusters. As another non-limiting example, after choosing drugs within the top percentile range of average similarity scores, additional drugs may be chosen from fine clusters smaller than a threshold cluster size from which no other drugs had yet been chosen for the principal drug set. In such embodiments, the drug from the fine cluster having the highest average Tanimoto similarity may be chosen even if it is not within the top percentile range.
[0104] At block 310, the principal set determination engine 120 stores the principal drug set in a principal data store 114 of the prediction computing system 102. The subroutine 300 then completes execution and returns control to its caller. While FIG. 3 illustrates a general subroutine 300 for determining a principal drug set, various non-limiting examples of determining principal drug sets are provided in the Examples below, including Example 3 and Example 9.
[0105] FIG. 4 is a flowchart that illustrates a non-limiting example embodiment of a procedure for training a predictive model based on functional test results and features of a principal drug set, according to various aspects of the present disclosure. The subroutine 400 is a non-limiting example of a subroutine suitable for use at subroutine block 210 of FIG. 2.
[0106] From a start block, the subroutine 400 advances to block 402, where a model training engine 122 of the prediction computing system 102 retrieves the principal drug set from the principal data store 114. The principal drug set retrieved from the principal data store 114 may include the identities of a subset of drugs from a diversity drug set that are to be used to train the machine learning model, and that were used to generate functional test results as described at block 206 of FIG. 2.
[0107] At block 404, the model training engine 122 retrieves functional information and structural information for the drugs of the principal drug set from the drug data store 110. Any suitable functional information and structural information may be retrieved. In someembodiments, similar functional information and structural information may be retrieved as was described in subroutine 300 for use in determining the principal drug set. For example, the functional information may include cell line viability data from previous testing, and the structural information may include binarized molecular fingerprints and / or other information regarding the chemical structures of the drugs.
[0108] At block 406, the model training engine 122 retrieves the functional test results for the principal drug set from the result data store 118. The functional test results indicate the measured biological responses of living systems (e.g., samples obtained at block 204) upon testing via application of the drugs of the principal drug set (e.g., as generated at block 206 and received at block 208).
[0109] At block 408, the model training engine 122 trains at least one machine learning model to accept the functional information and the structural information as input and to generate functional predicted results that match the functional test results as output. In some embodiments, the machine learning model may be trained to accept the functional information as input without the structural information. Any suitable architecture may be used for the machine learning model, including but not limited to an elastic net regularization model or an artificial neural network model (e.g., a convolutional neural network, a deep neural network, etc.), and any suitable technique may be used to train the machine learning model, including but not limited to gradient descent. In general, the machine learning model is trained by iteratively adjusting weights, biases, and / or other parameters of the machine learning model such that when the functional information and the structural information for a given drug are provided as input, differences between the output functional predicted results and the previously collected functional test results are minimized. A detailed explanation of a non-limiting example embodiment of the training of a machine learning model is provided in Example 6.
[0110] At block 410, the model training engine 122 stores the at least one machine learning model in a model data store 116 of the prediction computing system 102. The subroutine 400 then completes execution and returns control to its caller.[OHl] In another aspect, the present disclosure provides a method for selecting a candidate drug for modulating a biological response in a living system and which utilizes a computer-implemented method for predicting efficacies of drugs as described above. The biological response is “predetermined,” z.e., has already been selected as the biological readout to be detected in the method, and can be any biological response for which identification of a drug modulator is desired. Examples of biological responses for screening of drug modulators include cell viability, cell proliferation, cell death, cellular differentiation, chemotaxis, a morphology change, motility, contractility, transcription factor activity, gene expression, and protein secretion (e.g., growth factor or cytokine secretion), to name a few.
[0112] To select a candidate drug for modulating a predetermined biological response, a principal drug set is tested on a set of samples of a living system. Suitable living systems include any type of tissue (e.g., freshly removed from a subject or cryopreserved), dissociated cells from a tissue, cell lines, and whole organisms. Particularly suitable whole organisms include various model organisms such as, e.g., zebrafish (Danio rerio Caenorhabditis elegans, Xenopus e.g. , Xenopus tropicalis and Xenopus laevis . axolotl (Ambysloma mexicanum . Japanese rice fish (Oryzias talipes . Drosophila melanogasler. chickens (Gallus domesliciis). rats (Rattus norvegicus). rabbits (Oryctolagus ciiniciilus). Guinea pigs (Cavia porcellus). and mice (Mus musculus). Tissues used for testing, or tissues from which dissociated cells or cell lines for testing are derived, may be normal or diseased tissue (also referred to herein as tissue “affected by a disease or disorder”). In some variations, a diseased tissue used in testing (or from which cells for testing are derived) is tissue from a tumor. Similarly, whole organisms used in testing may be normal or models for a disease or disorder, including, e.g, animal models of a cancer, an inflammatory disease, an autoimmune disease, a metabolic disease (e.g, diabetes), a neurodegenerative disease (e.g., Alzheimer’s), or an infectious disease, to name a few. Tissues for testing (or from which cells for testing are derived) can be solid, soft, or liquid tissues. For example, in certain embodiments wherein the tissue is from a tumor, the tumormay be a solid tumor (for example, a tumor from a cancer of the brain, ovary, breast, colon, or liver (e.g., a hepatocellular carcinoma or fibrolamellar carcinoma)), a soft tissue tumor (e.g., a soft tissue sarcoma), or a liquid tumor (e.g., a blood or lymphatic cancer).
[0113] To test the principal drug set, for each drug within the principal set, the drug is administered to at least one different sample from the set of samples to obtain a test result, wherein the test result is a measurement of the biological response, thereby obtaining a set of test results for the drugs in the principal drug set. For samples that are not whole organisms (e.g., tissue or cell samples), “administering” the drug includes culturing the sample in the presence of the drug for a time sufficient to measure induction of or an effect on the biological response. For whole organisms, “administering” the drug can include any appropriate route for delivering the drug to a target tissue. Testing may include the presence of one or more other agents that have a known effect on the biological response being measured (e.g., enhancing or inhibiting the response) to evaluate the effect of the drugs in the principal set in modulating the response. For example, testing may include the presence of a naturally occurring factor (e.g., a ligand to a cellular receptor) known to activate a signaling pathway associated with the biological response. In alternative variations, testing is performed by administering each drug in the absence of other active agents.
[0114] Following testing of the principal drug set and obtaining test results, the test results are provided to a computing system configured to perform a computer-implemented method for predicting efficacies of drugs in a diversity drug set as described herein. Predicted efficacies of drugs in the diversity drug set are then received from the computing system. Based on the predicted efficacies, a candidate drug for modulating the biological response in the living system is selected from the diversity drug set. In certain variations, a plurality of candidate drugs are selected. The method may further include testing the selected candidate drug on a sample of the living system to confirm the predicted efficacy of the drug (or testing the selected plurality of drugs on samples of the living system to confirm the predicted efficacies).
[0115] In certain embodiments, the method further includes, before the testing step, selecting the principal drug set from the drugs in the diversity drug set based on, e.g., functional and / or structural similarities among drugs within the diversity set, such as, for example, described herein.
[0116] In embodiments comprising testing on tissue samples or cells derived directly from a tissue, a method as above may further include removing tissue from one or more subjects to obtain the samples. Removal of tissue may be by any suitable means, including, for example, by resection (surgical biopsy), needle biopsy, endoscopic biopsy, skin biopsy, or bone marrow biopsy, to name a few. Removed tissue may then be further processed to obtain samples, such as, e.g., by cutting the tissue into smaller sections (e.g., tissue slices or cuboids) or by dissociating cells from the tissue. In particular variations, the samples are intact tissue samples (e.g, intact samples that maintain the native microenvironment of the removed tissue). Removed tissue, including dissociated cells derived directly therefrom, may be used for testing immediately or soon after removal from a subject (“freshly removed”) or may be cryopreserved before testing in the method. For embodiments comprising sectioning of tissue into, e.g., tissue slices or cuboids, tissue may be cryopreserved before or after sectioning. Cryopreservation of tissue or cell samples is particularly useful, for example, in embodiments comprising subsequent testing of the selected drug (or drugs) to confirm the predicted efficacy (or efficacies).
[0117] Methods of selecting a drug for modulating a biological response in living systems as described herein are useful, for example, to identify classes of drugs that modulate the biological response and / or interrogate mechanisms of action of active drugs, including, e.g, identification of molecular targets and signaling pathways associated with the biological response. Accordingly, in certain variations, a method of selecting a drug for modulating a biological response in a living system further includes identifying, based on the predicted efficacies, a class of drugs that are efficacious for modulating the response. In such variations, other drugs within the class of drugs (not necessarily within the diversity drug set) may be identified and tested for efficacy in modulating the biological response.In other, non-mutually exclusive embodiments, a plurality of drugs are identified as efficacious, and the method further includes identifying at least one signaling pathway or molecular target that is common among the plurality of identified drugs based on functional information for the diversity drug set, thereby identifying a signaling pathway or molecular target associated with the biological response in the living system. Based on identification of a signaling pathway or molecular target, other drugs (not necessarily with the diversity drug set) may be identified and tested for efficacy. In addition, interrogation of molecular targets and signaling pathways in this manner facilitates the discovery of new biology associated with different biological response in living systems.
[0118] In another aspect, the present disclosure provides a method for selecting a candidate drug for treatment of a disease or disorder in a subject and which utilizes a computer-implemented method for predicting efficacies of drugs as described above. The disease or disorder can be any disease or disorder for which treatment is sought, including, for example, a cancer, an inflammatory disease, an autoimmune disease, a metabolic disease, a neurodegenerative disease, or an infectious disease, to name a few. The method is particularly useful for identification of therapies for rare (e.g., ultra-rare) diseases and disorders and other conditions for which information regarding effective therapies may be limited and / or for which there is significant individual variation in response to treatment. The methods are particularly useful, for example, for selecting candidate drugs for treatment of various cancers, where even within a given cancer type, tumors from different patients can exhibit heterogeneity that limits therapeutic success and is incompletely understood. The methods in accordance with the present disclosure can therefore be used to identify candidate drugs for treatment of various different cancers, whether characterized by solid, soft tissue, or liquid tumors.
[0119] Exemplary cancer types amenable to candidate drug selection in accordance with the present disclosure include breast cancer (e.g., metastatic breast cancer or inflammatory breast cancer); lung cancer (e.g., non-small cell lung cancer, small cell carcinoma, or mesothelioma); cancer of the head and neck (e.g., a cancer of the oral cavity, orophyarynx,nasopharynx, hypopharynx, nasal cavity or paranasal sinuses, larynx, lip, or salivary gland); gastrointestinal tract cancer (e.g., colorectal cancer, gastric cancer, esophageal cancer, or anal cancer); gastrointestinal stromal tumor (GIST); pancreatic adenocarcinoma; pancreatic acinar cell carcinoma; cancer of the small intestine; cancer of the liver or biliary tree (e.g., liver cell adenoma, hepatocellular carcinoma, fibrolamellar carcinoma, hemangiosarcoma, extrahepatic or intrahepatic cholangiosarcoma, cancer of the ampulla of vater, or gallbladder cancer); gynecologic cancer (e.g., cervical cancer, ovarian cancer, fallopian tube cancer, peritoneal carcinoma, vaginal cancer, vulvar cancer, gestational trophoblastic neoplasia, or uterine cancer, including endometrial cancer or uterine sarcoma); cancer of the urinary tract (e.g., prostate cancer; bladder cancer; penile cancer; urethral cancer, or kidney cancer such as, for example, renal cell carcinoma or transitional cell carcinoma, including renal pelvis and ureter); testicular cancer; cancer of the central nervous system (CNS) such as an intracranial tumor (e.g., astrocytoma, anaplastic astrocytoma, glioblastoma, gliosarcoma, oligodendroglioma, anaplastic oligodendroglioma, ependymoma, primary CNS lymphoma, medulloblastoma, germ cell tumor, pineal gland neoplasm, meningioma, pituitary tumor, tumor of the nerve sheath (e.g., schwannoma), chordoma, craniopharyngioma, a chloroid plexus tumor (e.g., chloroid plexus carcinoma), or other intracranial tumor of neuronal or glial origin) or a tumor of the spinal cord (e.g., schwannoma, meningioma); an endocrine neoplasm (e.g., thyroid cancer such as, for example, thyroid carcinoma, medullary cancer, or thyroid lymphoma; a pancreatic endocrine tumor such as, for example, an insulinoma or glucagonoma; an adrenal carcinoma such as, for example, pheochromocytoma; a carcinoid tumor; or a parathyroid carcinoma); skin cancer (e.g., squamous cell carcinoma; basal cell carcinoma; Kaposi’s sarcoma; Merkel cell carcinoma; or a malignant melanoma such as, for example, an intraocular melanoma); bone cancer (e.g., a bone sarcoma such as, for example, osteosarcoma, osteochondroma, or Ewing’s sarcoma); multiple myeloma; a chloroma; a soft tissue sarcoma (e.g., a fibrous tumor or fibrohistiocytic tumor); a tumor of the smooth muscle or skeletal muscle; a blood or lymph vessel perivascular tumor (e.g., Kaposi’ssarcoma); a synovial tumor; a mesothelial tumor; a neural tumor; a paraganglionic tumor; an extraskeletal cartilaginous or osseous tumor; a pluripotential mesenchymal tumor; a non-Hodgkin lymphoma (e.g., B-cell lymphoma, T-cell lymphoma, or undifferentiated lymphoma), a leukemia (e.g., chronic myelogenous leukemia, hairy cell leukemia, chronic lymphocytic leukemia, chronic myelomonocytic leukemia, acute myelocytic leukemia, or acute lymphoblastic leukemia), a myeloproliferative disorder (e.g., multiple myeloma, essential thrombocythemia, myelofibrosis with myeloid metaplasia, hypereosinophilic syndrome, chronic eosinophilic leukemia, or polycythemia vera), and Hodgkin lymphoma.
[0120] To select a candidate drug for treatment of a disease or disorder in a subject, a principal drug set is tested on a set of samples from the subject. Suitable samples include tissues affected by the disease or disorder, or cells derived from such tissues. In some variations, the samples are intact tissue samples. Particularly suitable samples are intact tissue samples that preserve the native environment of disease-affected cells or tissue. For example, in some embodiments wherein the disease or disorder is a cancer, the samples are intact tissue samples that preserve the native tumor microenvironment (TME). Intact tissue samples can include samples from, e.g., needle biopsies or surgical biopsies (resections), wherein the biopsied tissue is dissected into smaller, intact sections (e.g., tissue slices or “cuboids”), such as with a tissue slicer. Samples can be from tissues that have been freshly removed from the subject or from cryopreserved tissues. In some variations, the method further includes the step of removing tissue from the subject to obtain the samples, which can be by any suitable means (e.g., as discussed above in the context of a method for selected a candidate drug for modulating a biological response). Such a step can further include further processing of the removed tissue to obtain the samples (for example, cutting the tissue into smaller sections such as, e.g., tissue slices or cuboids, or dissociating cells from the tissue). As previously discussed in the context of selecting candidate drugs that modulate a biological response, removed tissue, including dissociated cells derived directly therefrom, may be used for testing immediately (e.g., within 24 hours, or within a few to several hours) or soon (e.g, within 24 to 48 hours) after removal from a subject (“freshlyremoved”), or may be cryopreserved before testing in the method. For embodiments comprising sectioning of tissue, such tissue may be cryopreserved before or after sectioning. Cry opreservation of samples is particularly useful, e.g., for subsequent testing of a selected drug to confirm its predicted efficacy for treating the disease or disorder, as discussed further herein.
[0121] To test the principal drug set for selection candidate therapeutic drugs, for each drug within the principal set, at least one different sample is cultured in the presence of the drug to obtain a clinically relevant test result, thereby obtaining a set of test results for the drugs in the principal drug set. For example, each drug can be added to different wells of a multi -well tissue culture plate, where each well contains from one to several (e.g., two or three) samples (e.g., tissue cuboids) in culture media. The samples may be cultured in a complex culture system. In some variations, the complex culture system is a three- dimensional (3D) culture system. The samples are cultured in the presence of drugs (and appropriate controls) for a sufficient time to measure a clinically relevant response (e.g., change in cell viability or expression of disease biomarkers).
[0122] Following testing of the principal drug set and obtaining test results, for selection of candidate therapeutic drugs, the test results are provided to a computing system configured to perform a computer-implemented method for predicting efficacies of drugs in a diversity drug set as described herein. Predicted efficacies of drugs in the diversity drug set are then received from the computing system. Based on the predicted efficacies, a candidate drug for treating the disease or disorder in the subject is selected from the diversity drug set. In certain variations, a plurality of candidate drugs are selected. The method may further include testing the selected candidate drug on a sample from the subject to confirm the predicted efficacy of the drug (or testing the selected plurality of drugs on samples from the subject to confirm the predicted efficacies). Cryopreserved samples from the subject are particularly suitable for testing of selected candidate drugs to confirm the predicted efficacies. In some such embodiments, the cryopreserved samples are from the same tissue removed from the subject for testing of the principal drug set. For example, insome variations comprising testing on tissue samples and wherein tissue removed from the subject is cut into separate sections for the principal drug set screening, a subset of the separate tissue sections have been cryopreserved for subsequent confirmation testing of selected candidate drug(s). In some embodiments comprising selection of a plurality of candidate drugs and wherein the method includes testing the selected candidate drugs on samples from the subject to confirm the predicted efficacies, the method further includes selecting one of the selected candidate drugs (confirmed as efficacious) to treat the disease or disorder in the subject. In other embodiments, a combination of drugs (e.g., a combination of two drugs) is selected for treating the disease or disorder based on the testing data and functional profiling of drugs.
[0123] In certain embodiments, a method for selecting a candidate drug for treating a disease or disorder as above further includes, before the testing step, selecting the principal drug set from the drugs in the diversity drug set based on, e.g., functional and / or structural similarities among drugs within the diversity set, such as, for example, described herein.
[0124] In some embodiments of a method for selecting a candidate drug for treatment of a disease or disorder as above, the method further includes treating the disease or disorder in the subject by administering an effective amount of the selected candidate drug (or combination of drugs) to the subject. For treatment, a drug is formulated according to know methods as a pharmaceutical composition in a mixture with a pharmaceutically acceptable carrier, and the is delivered in a manner consistent with conventional methodologies associated with management of the disease or disorder for which treatment is sought. An effective amount of the drug (or a combination of drugs) is administered to the subject for a time and under conditions sufficient to treat the disease or disorder.
[0125] In some embodiments, the subject is already undergoing a treatment for the disease or disorder comprising administration of a drug within the diversity drug set and the method is a method for monitoring treatment in the subject. In some variations, the selected candidate drug is different from the drug being administered to the subject. In other variations, a plurality of candidate drugs are selected that includes the drug beingadministered and at least one drug that is different; in some such embodiments, the method identifies an efficacious combination therapy comprising the drug being administered and a drug that is different. In certain variations, the method further includes changing the treatment in the subject based on the predicted efficacies of the drugs. For example, changing the treatment may include (i) stopping treatment with the drug being administered to the subject, (ii) administering to the subject an effective amount of a selected candidate drug that is different from the drug be administered, and / or (iii) combining the drug being administered with a second selected drug as a combination therapy. In other embodiments, the selected candidate drug is the same as the drug being administered and the method confirms efficacy of the current treatment.
[0126] In certain embodiments of a method for selecting a candidate drug for treatment of a disease or disorder, the method further includes identifying, based on the predicted efficacies, a class of drugs efficacious for treating the disease or disorder. In other, non- mutually exclusive embodiments, a plurality of drugs are identified as efficacious for treating the disease or disorder in the subject, and the method further includes identifying at least one signaling pathway or molecular target that is common among the plurality of identified drugs based on functional information for the diversity drug set.
[0127] In some related aspects, the present disclosure provides a method for identifying a signaling pathway associated with a disease or disorder, or identifying a molecular target for treating the disease or disorder, and which utilizes a computer-implemented method for predicting efficacies of drugs as described above. The method generally includes the following steps: (a) testing a principal drug set on a set of samples (e.g., tissue samples or cells) affected by the disease or disorder, wherein the principal drug set is a subset of a diversity drug set, and wherein, for each drug within the principal drug set, a different sample is cultured in the presence of the drug to obtain a clinically relevant test result, thereby obtaining a set of test results for the drugs in the principal drug set; (b) providing the set of test results to a computing system configured to perform a computer-implemented method for predicting efficacies of drugs in a diversity drug set as summarized above; (c)receiving predicted efficacies of drugs in the diversity drug set from the computing system; (d) selecting, from the diversity drug set, a plurality of drugs that are identified as efficacious for treating the disease or disorder based on the predicted efficacies; and (e) identifying at least one signaling pathway or molecular target that is common among the plurality of identified drugs based on functional information for the diversity drug set, thereby identifying a signaling pathway associated with the disease or disorder or identifying a molecular target for treating the disease or disorder. In certain embodiments, the method further includes selecting the principal drug set from the drugs in the diversity drug set before the testing step. In certain embodiments, the method further includes, before the testing step, selecting the principal drug set from the drugs in the diversity drug set based on, e.g., functional and / or structural similarities among drugs within the diversity set, such as, for example, described herein. In certain variations, the method further includes identifying one or more additional drugs targeting the identified signaling pathway or molecular target. In some such embodiments, the method further includes testing the one or more additional drugs in a physiologically relevant assay or model to evaluate efficacy in treating the disease or disorder.
[0128] In other related aspects, the present disclosure provides a method for identifying a drug that modulates a predetermined signaling pathway associated with a disease or disorder, and which utilizes a computer-implemented method for predicting efficacies of drugs as described above. The method generally includes the following steps: (a) testing a principal drug set on a set of cells (e.g., cells from a cell line) capable of signal transduction through a predetermined signaling pathway, wherein the principal drug set is a subset of a diversity drug set, wherein, for each drug within the principal drug set, a different sample of cells is cultured in the presence of the drug, and wherein a cellular phenotype associated with the signaling pathway is detected or measured to obtain a test result, thereby obtaining a set of test results for the drugs in the principal drug set; (b) providing the set of test results to a computing system configured to perform a computer-implemented method for predicting efficacies of drugs in a diversity drug set as summarized above; (c) receivingpredicted efficacies of drugs in the diversity drug set from the computing system; and (d) selecting, from the diversity drug set, a drug that is identified as efficacious for modulating the signaling pathway based on the predicted efficacies. In some embodiments, the identified drug is identified as an activator of the signaling pathway. In other embodiments of the method, the identified drug is identified as an inhibitor of the signaling pathway. The cellular phenotype can be any measurable phenotype that can serve as an indicator of signaling pathway activation or inhibition in the cells and under the culture conditions being used. Particularly suitable phenotypes include transcriptional activity as measured, e.g., using transcription factor reporter assays. Other suitable cellular phenotypes include cell growth, cell proliferation, cell death (e.g., apoptosis), cellular differentiation, gene expression, and protein secretion. In certain embodiments, the method further includes, before the testing step, selecting the principal drug set from the drugs in the diversity drug set based on, e.g., functional and / or structural similarities among drugs within the diversity set, such as, for example, described herein. In some embodiments, a plurality of drugs are identified as efficacious for modulating the signaling pathway. In some embodiments, the cells are cultured with the principal drug set in the presence of another drug or factor (e.g., a natural ligand) that initiates signal transduction through the signaling pathway.
[0129] The invention is further illustrated by the following non-limiting examples.Example 1: Overview of Exemplary Al-based Drug Screening Platform Using Principal and Diversity Drug Sets
[0130] The present Al-based drug screening platform was used to enable the screening of -4000 diverse drugs (“diversity set”) for drug repurposing and mechanism interrogation (see FIG. 5). First, using a well-curated and annotated collection of drugs that includes -1800 FDA-approved drugs, -1000 clinical trial drugs, and -1100 preclinical drugs, and their sensitivity in -500 cell lines, functional profiling was applied to computationally derive 254 drugs as a “principal set,” which functionally represent the -4000 diversity set drugs. This reduced “principal set” was used to develop machine learning (ML)-based models that accurately predict responses to the full panel of -4,000 drugs with >85%accuracy in 26 breast cancer cell lines. Thus, despite the limited size of the principal set, they encapsulate sufficient diversity of drugs and mechanisms to serve as the basis for combining ex vivo tumor slices with in silico screening. This Al-based screening platform enables direct use of limited but highly physiologically relevant clinical samples to select drugs for preclinical drug discovery, which can dramatically accelerate the discovery of curative therapies. This patient-centric approach also facilitates personalized oncology. Within a week of receiving fresh tissue from the operation room, actionable recommendations of -1800 FDA-approved drugs can be provided. The extension of this platform to include the use of needle biopsies further strengthens its potential to propel personalized oncology.Example 2: Construction of a Diversity Set for Drug Repurposing and Mechanism Interrogation
[0131] A set of 3897 small molecule inhibitors spanning >250 mechanisms of action (MOA) classes was manually curated. The criteria for drugs in the diversity set library included (1) FDA-approved or drugs in various phases of clinical trials, (2) clinically withdrawn drugs with known clinical safety profiles, and (3) reproducible and extensive profiling in large numbers of cancer models (-500) that may be used as functional information. Supervised curation of various publicly available large-scale profiling efforts (see Barretina etal., Nature 483:603-607, 2012; Garnett etal., Nature 483:570-575, 2012; Corsello et aL, Nature Medicine 23:405-408, 2017; Seashore-Ludlow et aL, Cancer Discovery 5: 1210-1223, 2015; Reinhold etal., Cancer Res. 72:3499-3511, 2012) identified 1772 FDA-approved, 224 Phase I, 32 Phase I / II, 444, Phase II, 19 Phase II / III, 237 Phase III, 1097 preclinical, and 72 withdrawn drugs that satisfied the above criterion (see FIG. 6A). Collectively, this set targeted 1680 primary protein targets (see FIG. 6B). STRING interactome analysis (see Szklarczyk c / aL, Nucleic Acids Res. 43:D447-452, 2015) of these 1680 protein targets found 8,009 experimentally validated protein-protein physical interactions and 16,774 total interactions within the first shell (n=l criteria) (see FIG. 6B). Further, functional enrichment of the primary protein targets revealed significantenrichment in 2862 Gene Ontology (GO) biological processes (FDR<0.05), 545 GO molecular functions (FDR<0.05), 252 KEGG pathways (FDR <0.03), and 757 Reactome pathways (FDR<0.05) (see FIG. 6A and FIG. 6B). Thus, a chemical screening using this collection of diversity set drugs would not only enable probing of broad molecular mechanisms and underlying signaling pathways but also offer an opportunity for drug repurposing in ultra-rare cancers.Example 3: Determination of a Principal Drug Set for Screening
[0132] Due to the limited number of specimens for a given human tumor, tissue slice culture is not amenable to large-scale, systems-based high-throughput approaches. To circumvent this limitation, a strategy was developed that involves screening a selective set of drugs (-200) and using the power of machine learning to predict drug responses for the remaining thousands of drugs in silico. This would reduce the amount of tissue for screening the full diversity set drug library by 20-fold. Therefore, a “principal set” of -200 drugs was identified from the diversity set to represent broad mechanisms of action, which could be used to screen human tumors ex vivo.
[0133] The 3,897 drugs from the diversity set have been extensively profiled to obtain functional information for >500 cancer cell lines spanning more than 20 different tissue / cell-type origins. It was speculated that drugs with similar or overlapping mechanisms of actions would have comparable effects on cell viability profiles across the 500 cell lines and tend to cluster together. For example, PARP inhibitors would be expected to exhibit similar effects in a broad collection of cell lines carrying BRCA mutations, regardless of the cancer type. Thus, rough clusters were used to group drugs with functional similarities using an unbiased approach through unsupervised hierarchical clustering of the 3897 drugs across 475 cancer cell lines based on viability data at a single dose (2.5 pM). The clusters ranging from 2 to 100 were evaluated by two metrics: (1) silhouette distance (SD) (see Saputra et al. , in Sriwijaya International Conference on Information Technology and Its Applications (SICONIAN 2019) 341-346 (Atlantis Press, 2020)), which measures the similarity between samples and their assigned clusters compared to other clusters, and(2) WSS (within-cluster-sum of squares) (see Duong and Vrain, in 2013 IEEE 25th International Conference on Tools with Artificial Intelligence 1060-1067 (IEEE, 2013); Edwards and Cavalli-Sforza, Biometrics 362-375, 1965), which measures the sum of squared distance between data points in a cluster and the centroid of that cluster (see FIG. 7). Thirty-one clusters were selected as the optimal number, based on the SD curve reaching a local maximum and the slope of the WSS curve becoming level (see 8). However, the majority of these clusters contain more than 100 drugs, with some containing over 300 drugs. This large number of drugs makes it challenging to choose representative drugs and may result in differences between drugs in the same cluster being overlooked. To address this challenge, fine clusters were used to further stratify each of the 31 functional clusters and include the similarity in the chemical structure of the 3,897 drugs. Thus, unsupervised hierarchical clustering was applied to each of the 31 initial clusters for further subclustering, guided by the following biological metrics:(1) Cosine similarity (CS). The first metric used was the average cosine similarity (Avg CS) of drug response within a cluster (see Xia el al.. Information Sciences 307:39- 52, 2015). To calculate this, the similarity between the two drugs was measured by computing the cosine of the angle between their vectors of cell line viability profiles. In essence, drugs with similar effects on cancer cell viability would have a higher Avg CS value.(2) Tanimoto similarity (TS). The second metric used was Tanimoto similarity, which is often used to compare the structural similarity between two chemical drugs (see Bajusz et al.. Journal of Cheminformatics ! : 1 -13, 2015; Chung el al.. BMC Bioinformatics 20, 1-11 (2019). The TS measures the similarity between two drugs by calculating the ratio of their shared bit positions based on their binarized molecular fingerprints. These fingerprints were generated from the SMILES strings of each drug using the Morgan algorithm. In essence, drugs with similar chemical structures would have a higher Avg TS value.
[0134] The Avg TS and Avg CS of a cluster were defined as the mean values of the Avg TS and Avg CS of the drugs within a cluster, respectively. In turn, the Avg TS and Avg CS of a drug were defined as the average values of TS or CS between that drug and all other drugs in its cluster. The iterative sub-clustering process continued until three requirements were met:(i) the sum of the Avg CS and Avg TS of the cluster was greater than 0.25;(ii) at least one of the Avg TS and Avg CS of the cluster was greater than 0.11;(iii) the cluster size was less than 100 drugs.
[0135] The above process resulted in 239 final clusters. Most of these clusters contained approximately 10 to 20 drugs, with a few exceptions that contained more than 40 drugs. The cutoff for Avg CS and Avg TS for sub-clustering was set based on the corresponding values from the initial 31 clusters, which filtered out 70% of them. This more compact and specific clustering was justified by the increase in both Avg CS and Avg TS (see FIG. 7). For each of the 239 final clusters, the top 5% of drugs with the highest average Tanimoto similarity were selected to be included in the principal drug set. If a cluster had fewer than 20 drugs, at least one drug was selected to be included in the principal drug set. In total, 254 drugs were selected as a principal set for screening experimentally.
[0136] UMAP (Uniform Manifold Approximation and Projection) is a machine learning technique used for non-linear dimension reduction, often used for visualizing highdimensional data in lower dimensions (see Mclnnes el al.. “Umap: Uniform manifold approximation and projection for dimension reduction,” arXiv preprint arXiv: 1802.03426, 2018). In this case, the cancer cell viability data or molecular fingerprints for all 3897 drugs were projected onto a two-dimensional plot using UMAP, and the drugs from the principal set were overlaid onto these plots to visualize their relationships to each other based on their cell viability profiles or chemical structures. The UMAP projection of the principal drug set shows that the selected drugs capture a large portion of the variability in both the cancer cell screening and chemical space, indicating that they are representative of the functional and chemical diversity of the full drug library of -4000 drugs (see FIG. 8 A andFIG. 8B). Therefore, the principal set of 254 drugs can be used to screen for a broad range of drug activities in cancer tissues, as further exemplified herein.Example 4: Training and Validation of Neural Network Models Using the Cell Line Response Data
[0137] To validate the utility of the principal set, the response data from the 254 drugs in a panel of 26 breast cancer cell lines was modeled. Specifically, responses to the principal drug set in 449 cancer cell lines (excluding the breast cancer cell lines) were used as the functional information “input,” while the responses in 26 breast cancer cell lines were used as the functional test result “output.” Both elastic net regularization and non-linear Convolutional Neural Network (CNN) were applied for predicting responses to the full 3897 diversity drug set. Ten-fold cross-validation was used to evaluate the accuracy of the models. It was found that CNN models outperformed the linear elastic net model, and accurately predicted the response to the full panel of diversity drug sets with >0.8 Pearson correlation between predicted and observed responses (see FIG. 9). Importantly, when a set of 250 drugs were randomly selected for screening, the prediction accuracy for the full 3,897 diversity drug set drops to ~0.5. This indicates that the composition of the principal set is crucial for accurate prediction of the response to the full set of drugs.
[0138] Next, the utility of screening with the principal drug set was further validated in 11 independent cancer cell lines. Again, responses to the 254 principal drug set in 475 cancer cell lines were used as the functional information “input,” while the responses in 11 cancer cell lines were used as the functional test result “output.” Notably, this panel of 11 cell lines was not part of the original 475 cell lines from which the principal drug set was derived, thus representing an independent validation of the approach. Again, the CNN models accurately predicted the response to the full panel of diversity drug sets, achieving a Pearson correlation of >0.8 between predicted and observed responses (see FIG. 10A and FIG. 10B). In certain cell lines, the accuracy was as high as 0.96, indicating that limited screening with the principal set is necessary and sufficient to predict the responses to the full panel of 3,897 diversity drug set with confidence.Example 5: Principal Drug Set Screening to Identify and Validate Inhibitors and Activators of the Wnt Pathway
[0139] To further validate the approach of the present disclosure in exploiting the principal drug set to predict responses to the more encompassing diversity set for phenotypes besides cell viability, the “limited” screening set was applied to identify chemical modulators of the canonical Wnt pathway.
[0140] Wnt proteins are growth factors that play a crucial role in various biological processes, such as proliferation, migration, and differentiation (see Nusse, Cold Spring Harbor Perspectives in Biology 4:a011163, 2012; Steinhart and Angers, Development 145:devl46589, 2018). Abnormal activation of Wnt signaling has been linked to human developmental disorders and several types of cancer, including colon, skin, brain, and prostate cancer (see Zhan et al., Oncogene 36:1461-1473, 2017). The canonical Wnt signaling pathway is initiated by the binding of Wnt ligand to transmembrane receptors, including frizzled (Fzd) and low-density lipoprotein receptor-related protein (LRP5 / 6) (see Steinhart and Angers, supra). This triggers the formation of a macromolecular complex that involves Dishevelled and Fzd, resulting in the sequestering of Axin from the P-catenin destruction complex. As a result, P-catenin accumulates and enters the nucleus, where it acts as a transcriptional co-activator and binds to the TCF transcription factor. The TOPFLASH (TCF binding site) luciferase reporters (see Biechele and Moon, Wnt Signaling: Pathway Methods and Mammalian Models, 99-110, 2008) in HEK293 cells can be used to quantitatively monitor P-catenin-dependent transcription (see FIG. 11 A).
[0141] To identify chemical modulators of Wnt signaling, HEK293 STF cells were treated with Wnt3a in the presence of 254 principal set drugs at a concentration of 2.5 pM (see FIG. 1 IB). A negative control of DMSO was used. TCF activity was measured 24 hours after drug treatment using luciferase assay, and the resulting changes in TCF activity in response to the 254 principal drug set were used as the output. Meanwhile, the responses to the 254 principal drug set in 475 cancer cell lines were used as input to generate CNN models, and hyperparameter tuning with ten-fold cross-validation was used to evaluate theaccuracy of the predictions (see FIG. 11C). The optimized hyperparameters were then used to predict responses to the full 3897 Diversity drug set (see FIG. 1 ID). The principal drug set-based CNN models predicted that 275 out of 3897 drugs significantly affected TCF activity, with 138 drugs expected to enhance Wnt-driven TCF activity by more than 3-fold and 137 drugs predicted to decrease Wnt-driven TCF activity by more than 2-fold (see FIG. 11D). Notably, many previously known regulators of canonical Wnt signaling were identified, including four different GSK3P inhibitors that potently activated (>7000-fold) TCF activity. GSK3P is known to phosphorylate and target P-catenin for proteasomal degradation (see Steinhart and Angers, supra). Therefore, GSK3P inhibition can lead to increased P-catenin levels and enhanced TCF-mediated transcription. Other FDA-approved drugs that activated TCF activity by >1500-fold included Prazosin (an adrenergic receptor inhibitor) and Ranitidine (a histamine receptor inhibitor) (see FIG. 1 IE. Their effects were validated in independent RKO cells (data not shown), but the molecular mechanisms by which these drugs promote Wnt signaling are unclear. The CNN model also predicted several FDA-approved drugs to negatively regulate Wnt signaling, including Aurora kinase inhibitors and ICG001, a known antagonist of Wnt / p-catenin / TCF-mediated transcription (see McMillan and Kahn, Drug Discovery Today 10: 1467-1474, 2005) (see FIG. HE). Together, these data have significant implications for cancers with constitutive activation of Wnt signaling.
[0142] Overall, these results provide evidence for several important conclusions:(1) The screening approach based on the 254-principal set can be applied to various outputs, including cell growth and changes in transcriptional activity, and(2) The CNN models can predict both positive and negative regulators of the phenotype.Example 6: Application of Principal Drug Set Screening in Fibrolamellar Carcinoma (FLC) Cuboids
[0143] In this example, principal drug set screening was applied to the ultra-rare cancer, fibrolamellar carcinoma (FLC), to illustrate the predictive performance of the technology platform of the present disclosure in the setting of resource scarcity.
[0144] As proof-of-principle, two sets of FLC-derived samples were screened, each sample consisting of tissues / cells from the primary tumor (z.e., liver mass) along with their corresponding nodal metastases (referred to as Pirmaryl, Metl, Primary2, and Met2, see FIG. 12A). Primary 1 and Metl are cell lines derived from FLC, whereas Primary2 and Met2 are cuboids derived from FLC tumor slices. Briefly, cuboids were immediately placed into 96-well ultralow-attachment plates and incubated with optimized media containing Williams’ medium, 12 mM nicotinamide, 150 nM ascorbic acid, 2.25 mg / mL sodium bicarbonate, 20 mM HEPES, 50 mg / mL additional glucose, 1 mM sodium pyruvate, 2 mM L-glutamine, 1% (v / v) ITS, 20 ng / mL EGF, 40 lU / mL penicillin, and 40 pg / mL streptomycin and RealTime Gio reagent (Promega) according to the manufacturer's instructions. RealTime Gio reagent allows assessment of the viability (metabolic activity) of the tumors in real-time without disrupting the tissue. After 24 hours, the baseline cell viability of cuboids was measured by RealTime Gio bioluminescence using the IVIS imaging or Synergy H4 instrument (Biotek). Cuboids were then exposed to either DMSO (control), or principal drugs (2.5 pM), and overall tumor tissue viability was measured daily, up to 5 days after treatment. The change in viability relative to baseline was determined, and the resulting data were normalized to DMSO control and subjected to CNN modeling.
[0145] CNN models were generated in Python using the TensorFlow, NumPy, Pandas, and Keras libraries. The principal drug-based phenotypic responses were combined with previously generated functional profiling to build non-linear deep neural network models for predicting the response to the full diversity set library. The neural network model was trained using Bayesian optimization, Optuna, to optimize hyperparameters with 10-foldcross-validation. The best set of hyperparameters was chosen based on the root mean squared error (RMSE) value calculated during cross-validation. For example, the baseline CNN model for the Primary 1 FLC sample, which was trained on the response to 206 (~ 80%) randomly selected inhibitors in the training set and tested on the remaining 52 inhibitors (-20%), performed with >90% accuracy (R=0.97; see FIG. 12B). This model was built using six hidden layers, the SELU activation function, VarianceScaling weight initializer, Adamax optimizer, and 280 epochs. The final model is then trained on the entire dataset with the best set of hyperparameters, which was used to make predictions on the entire diversity set library for each FLC case (see FIG. 12C). By building individual models for each case, more accurate predictions could be made about how each FLC sample would respond to different drugs in the full diversity set library.
[0146] Next, the overlap between CNN model predictions for each sample was compared. Specific criteria were used to define a drug as being effective if it decreases the viability of FLC tissue / cells by 50% or more compared to DMSO control. Using this criterion to the CNN model prediction of effective inhibitors, we found a relatively low percentage of drugs that were effective in all FLC cases tested was found. Specifically, only 209 drugs (5% of all drugs) were effective in Primary 1, 118 drugs (3% of all drugs) were effective in Primary2, 166 drugs (4% of all drugs) were effective in Metl, and 116 drugs (3% of all drugs) in Met2 (see FIG. 12D). These findings suggest that the chemoresistant nature of FLC is due in part to the limited number of drugs that are effective in treating the disease. Consistently, it was observed that the average number of effective drugs in FLC samples was 50% lower than that found in 17 HCC (hepatocellular carcinoma) cell lines (see FIG. 12E). Further, our clustering analysis of the 17 HCC cell lines and four FLC cases revealed that the FLC formed a distinct cluster separate from the HCC samples (FIG. 12F). These findings are consistent with the clinical notion that FLC are “chemoresistant” to the currently available drugs used to treat HCC. Nonetheless, our study identified 56 drugs (1.4% of all the drugs) that are predicted to be effective in all FLC cases tested (see FIG.12D). This is a relatively small number of drugs, but it does offer hope that there may be some common therapeutic options for FLC patients.Example 7: Identification of Effective Classes of Drugs for FLC
[0147] While limited compared to the compendium of the entire diversity set of drugs, 56 drugs identified in the FLC screen of Example 6 is still a large number of candidates to sort through. Rather than testing each one individually, their commonalities in terms of mechanisms of action, targeted pathways, or biological processes were investigated to 1) better understand potential molecular pathways involved in FLC and 2) identify drug candidates that best target the intended pathways. The results may also inform future drug development efforts.
[0148] In this context, several key targets were uncovered that are predicted by the top candidate drugs, including PI3K, HD AC, HSP90, PLK1, and VDAC proteins (see FIG. 13A). Specifically, four HDAC, five HSP90, and three PLK1 inhibitors demonstrate efficacy in all four tested cases of FLC (FIG. 13B). Other pathways inferred from the FLC drug discovery highlight the retinoic acid signaling, proteosomes, and topoisomerases, which will prompt future investigation with specific drugs. In this Example, work validating PLK1 and HSP90 inhibitors in FLC is summarized.
[0149] PLK1 inhibitors: Polo Like Kinases (PLK) are key regulators of centrosome maturation (Lee and Rhee, J. Cell Biol. 195: 1093-1101, 2011), mitosis, and cell division (Lee et al.. Development & Reproduction 18:65, 2014). Using cryopreserved FLC tumor slices and non-tumor liver tissues, the efficacies of two 2nd generation PLK1 inhibitors were measured. We found that treatment of FLC tumor slices with 500 nM CYC140 and 500 nM Onvansertib was found to significantly decrease the viability of FLC tumor slices (see FIG. 14 A). In contrast, the viability of slices prepared from the non-tumor liver from the same patient was not affected by the PLK1 inhibitor treatment. These data are consistent with a previous study showing normal human hepatocytes and untransformed embryonic fibroblasts are less sensitive to PLK1 inhibition (see Lan et al.. Laboratory Investigation 92: 1503-1514, 2012). Whether pharmacological inhibition ofPLKl could inhibit FLC cellgrowth was also investigated. The data revealed that PLK1 inhibitor treatment reduced the growth of patient-derived primary FLC cells in a dose-dependent manner, with clinical- grade Onvansertib (EC50 25 nM) and CYC140 (EC50 8 nM) outperforming the response of BI-2536, a first-generation PLK1 inhibitor (see FIG. 14B). These findings suggest that DNAJ-PKAc fusion may alter the network wiring to create a hypersensitivity towards PLK1 inhibition in FLC cells. Given the encouraging results obtained from our initial experiments, we intend to further investigate the effectiveness of PLK1 inhibitors in a larger sample size of FLC cases.
[0150] HSP90 inhibitors: The Heat shock protein (HSP90) family of ATP-dependent molecular chaperones play a crucial role in tumorigenesis by regulating the stability and function of client proteins involved in the growth, survival, and adaptation of cancer cells to cellular stress (Birbo el al.. hil. J. Mol. Sci. 22: 10317, 2021). However, previous clinical trials of multiple HSP90 inhibitors were unsuccessful due to a lack of clinical efficacy and toxicity (Sanchez el al.. Current Cancer Drug Targets 20:253-270, 2020). In these studies, three HSP90 inhibitors showed significant activities in organotypic tumor slices derived from FLC tumors but not the non-tumor liver tissues (see FIG. 14C). Furthermore, a doseresponse study in primary FLC cells revealed that the next-generation HSP90 inhibitor, XL88879, was more potent than earlier generation clinically abandoned HSP90 inhibitors (see FIG. 14D).
[0151] AURKA inhibitors: A previous clinical trial investigating drugs targeting AURKA has yielded negative results. Using principal drug set screening of the present disclosure to predict responses to Aurora kinase inhibitors in the FLC samples, models predicted highly variable responses to 18 Aurora kinase inhibitors, including ENMD-2076, which previously failed in the clinic (see FIG. 15). Notably, ENMD-2076 was predicted to be ineffective in three out of four FLC cases, suggesting that its failure in the clinic could have been predicted in silico. These findings not only inform effective drug selection but point to the negative predictive power of methods as described herein.Example 8: Application of Principal Drug Set Screening Using Needle Biopsies
[0152] Needle biopsy is a safe, time-tested diagnostic tool that is routinely used in solid cancers to obtain tissues for histology and molecular profiling (see, e.g., Pritzker and Nieminen, Archives of Pathology and Laboratory Medicine 143: 1399-1415, 2019; Pyo et al., Diagnostics 10:717, 2020; Yao et al. , Current Oncology 19: 16-27, 412; Kubo et al. , Medicine 97, 2018; Krishnamurthy, Cancer Cytopathology: Interdisciplinary International Journal of the American Cancer Society 111 : 106-122, 2007). Their application in drug testing is limited by the amount of tissue they provide. The challenge is to develop a needle-biopsy platform that allows for the screening of all FDA-approved drugs. This would be particularly impactful in ultra-rare cancers, where treatment options are limited, and tumor tissues are few. Identifying effective FDA-approved drugs in needle biopsies enables their use as neoadjuvant therapies after the diagnosis and before the surgery, thus immediately impacting patients’ lives. This application also helps monitor treatment response, as well as inform clinical decisions in the case of non-resectable tumors.
[0153] This Example shows that combining primary tumor tissue slices with artificial intelligence can identify potential therapeutics. Principal drug set screening using needle biopsy samples is highly advantageous in repurposing existing FDA-approved or clinically graded drugs for treating and personalizing treatment for ultra-rare cancers.
[0154] Responses to 25-40 drugs using tissues obtained from a single needle biopsy were successfully assessed in less than five days. This approach involves mechanically cutting the needle biopsy samples into -400 pm-wide cuboidal-shaped micro-dissected tissues, resulting in 300-1000 cuboids per sample. These cuboids are placed in hydrogel-coated, 96-well U-bottom plates with Realtime-glo viability detection reagent, with each well containing at least three cuboids (see FIG. 16A). Assessment of untreated cuboid viability after 24 hours revealed a reproducible and high viability signal (above 20,000 RLU) in 92% of wells in a 96-well plate (see FIG. 16B), which is sufficient for screening 30-50 drugs at a single dose in duplicate.
[0155] Responses to standard-of-care chemotherapy drugs in needle biopsies from colorectal cancer liver metastases were also evaluated. Needle biopsy cuboids were treated with chemotherapy drugs as single agents (500nM) or in combination (250nM for each drug). Staurosporine (STS) was used as a positive control. After 48 hours of treatment, a significant decrease in cuboid viability was observed with SN-38 (an active metabolite of irinotecan) and a combination of SN-38 with 5-FU (see FIG. 16C). However, no change in cuboid viability in response to treatment with oxaliplatin or a combination of oxaliplatin and SN-38 was found. In another needle biopsy from an intrahepatic cholangiocarcinoma, greater response was found in cuboids treated with gemcitabine and cisplatin compared to those treated with 5-FU and oxaliplatin (see FIG. 16D). These findings demonstrate the feasibility of using needle biopsy samples to generate cuboids for drug testing.Example 9: Curation of an HCC-specific Principal Drug Set
[0156] As described in Example 3, a set of 254 principal drugs was identified representing broad mechanisms of action that could be used routinely for screening in tissue slices and cuboids from any cancer type. Thus, this set is agnostic to the cancer type. In this Example, a smaller set of principal drugs was identified that are cancer-type-specific, thereby reducing the amount of tissue used to collect the training set without compromising the model performance.
[0157] Data from full diversity set screening in 18 hepatocellular carcinoma (HCC) cell lines was analyzed and 296 drugs were identified that effectively reduced HCC growth by at least 50% in 9 of the 18 cell lines (50% of samples tested). These 296 effective drugs were then mapped onto a UMAP that displayed functional profiling of all 3,897 drugs (see FIG. 17A). As expected, effective drugs clustered together, suggesting overlapping mechanisms of action. Conversely, several clusters had no effective drugs, indicating that drugs in those clusters had no functional role in HCC cell growth. By excluding these “nonfunctional” clusters, the sub-clustering approach described above was then applied to select a representative drug from each “effective” cluster to create a set of 42 HCC-specific drugs.
[0158] The 42 HCC-specific drug set was then used to predict the response to the full diversity set library and compared the performance with that obtained from the original principal set of 254 drugs in terms of correlation between predicted and observed response. Results showed that screening with the 42-drug set could predict the responses to the full diversity set with approximately 80% accuracy in 12 out of 18 cell lines and with 70% accuracy in 15 out of 18 cell lines (see FIG. 17B). Remarkably, no significant loss in predictive power was observed by reducing the drug set from 254 drugs to 42 in 17 out of 18 cell lines. These findings show that it is feasible to build a limited drug set specific to a disease type (e.g., cancer type) that could be used for Al-based drug screening in accordance with the present disclosure.Example 10: Application of DeepGeneX to Principal Drug Set Screening Response Data
[0159] Although testing drug response on needle biopsy provides the most direct evidence, it can be challenging to obtain fresh samples in a timely and efficient manner. A platform that utilizes molecular information from fresh, frozen, or fixed needle biopsy samples further facilitates clinical decision-making. In addition, this approach allows for the determination of treatment options from archived samples when tumors recur. This Example describes principal drug set screening to determine the response to the full diversity set, including 1,770 FDA-approved drugs, by using the molecular information. To achieve this, an orthogonal approach was used that combines the molecular gene signature and previously determined drug response data from tissue samples. Combining this approach with needle-biopsy-based screening, principal drug set screening can be employed at multiple junctions during a cancer patient’s journey. This enables clinicians to better understand how different drugs will interact with a patient’s unique cancer biology, ultimately leading to more effective treatment plans.
[0160] Many large-scale drug screening efforts in cancer cell lines have established that it is feasible to model the relationship between drug sensitivity and the genomic and genetic characteristics of cancer cell lines (see Ghandi et al.. Nature 569:503-508, 2019; Garnettet al., Nature 483, 570-575, 412; Yang c / al., Nucleic Acids Res. 41, D955-961, 2013; Iorio et al., Cell 166; 740-754, 2016; Li et al., BMC Genomics 22:272, 2021; Reinhold et al., Human Genetics 134:3-11, 2015; Guan etal.,Mol. Ther. Nucleic Acids 17: 164-174, 2019; Suphavila et al., Bioinformatics 34:3907-3914, 2018). A framework called DeepGeneX, utilizing RNA-seq data and deep neural network modeling, has been used to predict a patient’s response to immune checkpoint therapy (ICB) (see Kang et al., iScience 25: 104228, 2022). One of the key features of this framework is its use of feature elimination steps (see Chen and Jeong, in Sixth international conference on machine learning and applications (ICMLA 2007) 429-435 (IEEE, 2007); Li and Yang, in Proceedings of the 28th annual international ACM SIGIR conference on Research and development in information retrieval 633-634, 2005) to identify a smaller set of genes that are most predictive of the response. DeepGeneX was used to analyze RNA-seq data from 19 melanoma patients and identified a set of six genes that could accurately predict the response to ICB therapy with 100% accuracy, outperforming linear models see Kang et al., supra).
[0161] In this Example, a similar approach was used to predict responses to a set of 254 drugs in 13 meningioma tissues directly obtained from the operating room. The principal set-based screening of cuboid tissues from these 13 meningioma cases was conducted, as described earlier. Concurrently, RNA expression data from these cases was profiled and utilized as input and drug responses as output to train CNN models see FIG. 18). The performance of the models was assessed through leave-one-out cross-validation (LOOCV) between predicted and observed responses to the principal set of 254 drugs. In LOOCV, a dataset comprising information on all 13 patients was used, where data from 12 patients were used at each step to train the model to predict the response of the remaining patient. The hyperparameters were optimized using Optuna, as described above. The CNN model constructed using the entire gene set (-17000 genes) performed with a mean squared error of 196.1 and R2 of 0.48 (see FIG. 19). However, by restricting the features to 3000 mostvariable genes, the model performance was significantly enhanced, resulting in an MSE of28.2 and R2 of 0.91 (>80% decrease in MSE).Example 11: Application of Principal Drug Set Screening to Meningioma Cases
[0162] In this Example, principal set-based screening as described herein was applied to two patients with meningioma: (1) a 45 -year-old patient with grade 3 meningioma and who had exhausted all therapies, and (2) a patient who had been battling meningioma for 15 years, had gone through seven surgeries, and who had exhausted all systemic therapies. For this study, tumor tissue was removed by surgical resection and tumor slices were cut into -400 pm-wide cuboids. Using a principal drug set of 200 drugs, representing a 4125 diversity drug set, principal set-based screening of these cuboid tissue samples was then conducted, as described earlier. This study identified 86 drugs predicted to have <50% viability relative to control samples, including 22 FDA approved drugs (13 oncology drugs, 3 cardiology drugs, 2 neurology drugs, 1 rheumatology drug, 2 infectious disease drugs, and 1 antibiotic). Plots of experimental vs. predicted responses showed high accuracy of the CNN models in predicting drug efficacies (R2 = 0.88; MSE = 0.01). A graph showing confirmation of the predicted efficacies for several of the identified FDA approved drugs is shown in FIG. 20 (patient samples identified as “Menin 4758” and “Menin 4759”).
[0163] While illustrative embodiments have been illustrated and described, it will be appreciated that various changes can be made therein without departing from the spirit and scope of the invention.
Claims
CLAIMSThe embodiments of the invention in which an exclusive property or privilege is claimed are defined as follows:
1. A computer-implemented method of predicting efficacies of drugs in a diversity drug set, the method comprising: receiving, by a computing system, test results for drugs in a principal drug set for samples of a subject; training, by the computing system, at least one machine learning model to generate functional predicted results of drugs in the diversity drug set based on the test results for drugs in the principal drug set; and predicting, by the computing system, functional predicted results of drugs in the diversity drug set by providing features of the drugs in the diversity drug set as input to the at least one machine learning model.
2. The computer-implemented method of claim 1, wherein determining the principal drug set from the diversity drug set includes: determining a plurality of rough clusters for the diversity drug set based on functional information for the diversity drug set; and determining a plurality of fine clusters for the diversity drug set based on the plurality of rough clusters, the functional information for the diversity drug set, and structural information for the diversity drug set.
3. The computer-implemented method of claim 2, wherein determining the plurality of rough clusters includes: determining functional values for each drug in the diversity drug set; and organizing the drugs from the diversity drug set into rough clusters based on one or more of a silhouette distance based on the functional values or a within-cluster sum of squares based on the functional values.
4. The computer-implemented method of claim 3, wherein the functional values include viability data at a single dose of the drugs.
5. The computer-implemented method of any one of claims 2 to 4, wherein determining the plurality of fine clusters based on the plurality of rough clusters, the functional information for the diversity drug set, and structural information for the diversity drug set includes: for each rough cluster: dividing the rough cluster into fine clusters based on two or more of an average cosine similarity of functional values of drugs within the rough cluster, an average correlation coefficient of functional values of drugs within the rough cluster, an average binary similarity of functional values of drugs within the rough cluster, or an average Tanimoto similarity of drugs within the rough cluster.
6. The computer-implemented method of claim 5, further comprising: for each fine cluster, further dividing the fine cluster into smaller fine clusters until the smaller fine clusters have a combined similarity value greater than a combined similarity threshold, separate similarity values each greater than separate similarity thresholds, and a cluster size less than a cluster size threshold.
7. The computer-implemented method of claim 6, wherein the combined similarity value is a sum of two or more of an average cosine similarity value, an average correlation coefficient value, an average binary similarity value, or an average Tanimoto similarity value.
8. The computer-implemented method of claim 6 or 7, wherein the separate similarity thresholds include two or more of an average cosine similarity value threshold, an average correlation coefficient value threshold, an average binary similarity value threshold, or an average Tanimoto similarity value threshold.
9. The computer-implemented method of any one of claims 5 to 8, wherein determining the principal drug set from the diversity drug set includes: choosing drugs from the diversity drug set having a highest percentage of average Tanimoto similarity.
10. The computer-implemented method of any one of claims 5 to 9, wherein determining the principal drug set from the diversity drug set includes: choosing at least one drug from each fine cluster smaller than a small cluster size threshold.
11. The computer-implemented method of any one of claims 1 to 10, wherein training at least one machine learning model to generate functional predicted results of drugs in the diversity drug set based on the test results for the drugs in the principal drug set includes: training the at least one machine learning model to accept functional information and structural information for the principal drug set as input and to generate functional predicted results that match the functional test results as output.
12. The computer-implemented method of any one of claims 1 to 11, wherein the at least one machine learning model includes at least one of an elastic net regularization model or a deep neural network model.
13. A non-transitory computer-readable medium having computer-executable instructions stored thereon that, in response to execution by one or more processors of a computing system, cause the computing system to perform actions of a method as recited in any one of claims 1 to 12.
14. A computing system having at least one processor and a non-transitory computer-readable medium, wherein the non-transitory computer-readable medium has computer-executable instructions stored thereon that, in response to execution by the atleast one processor, cause the computing system to perform actions of a method as recited in any one of claims 1 to 12.
15. A method for selecting a candidate drug for modulating a predetermined biological response in a living system, the method comprising: testing a principal drug set on a set of samples of a living system, wherein the principal drug set is a subset of a diversity drug set, and wherein, for each drug within the principal drug set, the drug is administered to a different sample to obtain a test result, wherein the test result is a measurement of a predetermined biological response, thereby obtaining a set of test results for the drugs in the principal drug set; providing the set of test results to a computing system configured to perform a method as recited in any one of claims 1 to 12; receiving predicted efficacies of drugs in the diversity drug set from the computing system; and selecting, from the diversity drug set, a candidate drug for modulating the biological response in the living system based on the predicted efficacies.
16. The method of claim 15, further comprising selecting the principal drug set from the drugs in the diversity drug set before the testing step.
17. The method of claim 15 or 16, wherein the living system is selected from the group consisting of a tissue, dissociated cells from a tissue, a cell line derived from a tissue, and whole organisms.
18. The method of claim 17, wherein the living system is a tissue.
19. The method of claim 17, wherein the living system is dissociated cells from a tissue.
20. The method of claim 17, wherein the living system is a cell line derived from a tissue.
21. The method of any one of claims 18 to 20, wherein the tissue is affected by a disease or disorder.
22. The method of claim 21, wherein the disease or disorder is a cancer and the tissue is a tumor tissue.
23. The method of claim 22, wherein the tissue is a solid tumor tissue.
24. The method of claim 22, wherein the tissue is a liquid or soft tumor tissue.
25. The method of claim 17, wherein the living system is a whole organism.
26. The method of claim 25, wherein the whole organism is selected from the group consisting of zebrafish, Caenorhabditis elegans, Xenopus. Drosophila melanogaster , and a mice.
27. The method of any one of claims 15 to 26, wherein the biological response is selected from the group consisting of cell viability, cell death, cellular differentiation, a morphology change, motility, contractility, transcription factor activity, and gene expression.
28. The method of any one of claims 15 to 27, further comprising testing the selected candidate drug on a sample of the living system to confirm the predicted efficacy of the drug.
29. The method of any one of claims 15 to 27, wherein the selecting step comprises selecting, from the diversity drug set, a plurality of candidate drugs for modulating the biological response in the living system based on the predicted efficacies.
30. The method of claim 29, wherein the method comprises testing each of the selected candidate drugs on samples of the living system to confirm the predicted efficacies of the drugs.
31. The method of any one of claims 15 to 30, further comprising identifying, based on the predicted efficacies, a class of drugs efficacious for modulating the biological response in the living system.
32. The method of any one of claims 15 to 31, wherein a plurality of drugs are identified as efficacious for modulating the biological response in the living system, and wherein the method further comprises identifying at least one signaling pathway or molecular target that is common among the plurality of identified drugs based on functional information for the diversity drug set.
33. A method for selecting a candidate drug for treatment of a disease or disorder in a subject, the method comprising: testing a principal drug set on a set of samples from a subject having the disease or disorder, wherein the principal drug set is a subset of a diversity drug set, wherein the samples are affected by the disease or disorder, and wherein, for each drug within the principal drug set, a different sample is cultured in the presence of the drug to obtain a clinically relevant test result, thereby obtaining a set of test results for the drugs in the principal drug set; providing the set of test results to a computing system configured to perform a method as recited in any one of claims 1 to 12; receiving predicted efficacies of drugs in the diversity drug set from the computing system; and selecting, from the diversity drug set, a candidate drug for treating the disease or disorder in the subject based on the predicted efficacies.
34. The method of claim 33, wherein the disease or disorder is a rare disease or disorder.
35. The method of claim 33 or 34, wherein the disease or disorder is a cancer.
36. The method of any one of claims 33 to 35, wherein the cancer is characterized by a solid tumor.
37. The method of any one of claims 33 to 35, wherein the cancer is characterized by a liquid or soft tissue tumor.
38. The method of claim 36 or 37, wherein the samples are from a primary tumor.
39. The method of claim 36 or 37, wherein the samples are from a metastatic site.
40. The method of any one of claims 33 to 39, wherein the samples are intact tissues samples.
41. The method of any one of claims 33 to 40, wherein the samples are cells dissociated from a tissue.
42. The method of any one of claims 36 to 39, wherein the samples are intact tissue samples that preserve the native tumor microenvironment (TME).
43. The method of any one of claims 33 to 42, wherein the samples are freshly obtained from the subject.
44. The method of claim 43, wherein the principal drug set is tested on the set of samples immediately after the samples are obtained from the subject.
45. The method of any one of claims 33 to 42, wherein the samples are cryopreserved samples.
46. The method of any one of claims 33 to 45, wherein the samples are cultured in a complex culture system.
47. The method of claim 46, wherein the complex culture system is a three- dimensional (3D) culture system.
48. The method of any one of claims 33 to 47, further comprising removing affected tissue from the subject to obtain the set of samples.
49. The method of claim 48, wherein the affected tissue is removed by a needle biopsy.
50. The method of claim 48, wherein the affected tissue is removed by resection.
51. The method of any one of claims 48 to 50, further comprising cutting the removed tissue into separate sections to obtain the set of tissue samples.
52. The method of any one of claims 33 to 51, further comprising testing the selected candidate drug on a sample from the subject to confirm the predicted efficacy of the drug.
53. The method of claim 52, wherein the sample used to confirm the predicted efficacy is from cryopreserved tissue.
54. The method of any one of claims 48 to 51, further comprising cryopreserving part of the removed tissue.
55. The method of claim 54, further comprising testing the selected candidate drug on a sample obtained from the cryopreserved tissue to confirm the predicted efficacy of the drug.
56. The method of claim 51, further comprising cryopreserving a subset of the separate tissue sections.
57. The method of claim 56, further comprising testing the selected candidate drug on a sample obtained from the cryopreserved tissue sections to confirm the predicted efficacy of the drug.
58. The method of any one of claims 33 to 57, further comprising treating the disease or disorder in the subject, wherein the treatment comprises administering an effective amount of the selected candidate drug to the subject.
59. The method of any one of claims 33 to 57, wherein the subject is already undergoing a treatment for the disease or disorder comprising administration of a drug within the diversity drug set and the method is a method for monitoring treatment in the subject.
60. The method of claim 59, wherein the selected candidate drug is different from the drug being administered to the subject.
61. The method of claim 60, further comprising changing the treatment in the subject based on the predicted efficacies of the drugs, wherein changing the treatment comprises at least one of (i) stopping treatment with the drug being administered to the subject and (ii) administering an effective amount of the selected candidate drug to the subject.
62. The method of any one of claims 33 to 61, wherein the selecting step comprises selecting, from the diversity drug set, a plurality of candidate drugs for treating the disease or disorder in the subject based on the predicted efficacies.
63. The method of claim 62, wherein the method comprises testing each of the selected candidate drugs on samples from the subject to confirm the predicted efficacies of the drugs.
64. The method of claim 63, wherein the method comprises further selecting one of the selected candidate drugs for treating the disease or disorder in the subject.
65. The method of any one of claims 33 to 64, further comprising identifying, based on the predicted efficacies, a class of drugs efficacious for treating the disease or disorder.
66. The method of any one of claims 33 to 64, wherein a plurality of drugs are identified as efficacious for treating the disease or disorder in the subject, and wherein the method further comprises identifying at least one signaling pathway or molecular target that is common among the plurality of identified drugs based on functional information for the diversity drug set.
67. A method for identifying a signaling pathway associated with a disease or disorder, or identifying a molecular target for treating the disease or disorder, the method comprising: testing a principal drug set on a set of samples affected by the disease or disorder, wherein the principal drug set is a subset of a diversity drug set, and wherein, for each drug within the principal drug set, a different sample is cultured in the presence of the drug to obtain a clinically relevant test result, thereby obtaining a set of test results for the drugs in the principal drug set; providing the set of test results to a computing system configured to perform a method as recited in any one of claims 1 to 12; receiving predicted efficacies of drugs in the diversity drug set from the computing system; selecting, from the diversity drug set, a plurality of drugs that are identified as efficacious for treating the disease or disorder based on the predicted efficacies; and identifying at least one signaling pathway or molecular target that is common among the plurality of identified drugs based on functional information for the diversity drug set, thereby identifying a signaling pathway associated with the disease or disorder or identifying a molecular target for treating the disease or disorder.
68. The method of claim 67, wherein the disease or disorder is a rare disease or disorder.
69. The method of claim 67 or 68, wherein the disease or disorder is a cancer.
70. The method of any one of claims 67 to 69, wherein the cancer is characterized by a solid tumor.
71. The method of any one of claims 67 to 69, wherein the cancer is characterized by a liquid or soft tissue tumor.
72. The method of claim 66 or 67, wherein the samples are from a primary tumor.
73. The method of claim 66 or 67, wherein the samples are from a metastatic site.
74. The method of any one of claims 67 to 73, wherein the samples are intact tissues samples.
75. The method of any one of claims 67 to 73, wherein the samples are cells dissociated from a tissue.
76. The method of any one of claims 70 to 73, wherein the samples are intact tissue samples that preserve the native tumor microenvironment (TME).
77. The method of any one of claims 67 to 76, wherein the samples are freshly obtained from one or more subjects having the disease or disorder.
78. The method of claim 77, wherein the principal drug set is tested on the set of samples immediately after the samples are obtained from the one or more subjects.
79. The method of any one of claims 67 to 76, wherein the samples are cryopreserved tissue samples.
80. The method of any one of claims 67 to 79, wherein the samples are cultured in a complex culture system.
81. The method of claim 80, wherein the complex culture system is a threedimensional (3D) culture system.
82. The method of any one of claims 67 to 81, further comprising removing affected tissue from one or more subjects having the disease or disorder to obtain the set of samples.
83. The method of claim 82, wherein the affected tissue is removed by a needle biopsy.
84. The method of claim 82, wherein the affected tissue is removed by resection.
85. The method of any one of claims 82 to 84, further comprising cutting the removed tissue into separate sections to obtain the set of samples.
86. The method of any one of claims 67 to 85, further comprising testing the selected drugs on a second set of samples to confirm the predicted efficacies of the drugs.
87. The method of claim 86, wherein the samples used to confirm the predicted efficacies are from cryopreserved tissue.
88. The method of any one of claims 82 to 85, further comprising cryopreserving part of the removed tissue.
89. The method of claim 88, further comprising testing the selected drugs on a second set of samples obtained from the cryopreserved tissue to confirm the predicted efficacies of the drugs.
90. The method of claim 85, further comprising cryopreserving a subset of the separate tissue sections.
91. The method of claim 90, further comprising testing the selected drugs on a second set of samples obtained from the cryopreserved tissue sections to confirm the predicted efficacies of the drugs.
92. The method of any one of claims 67 to 90, further comprising identifying one or more additional drugs targeting the identified signaling pathway or molecular target.
93. The method of claim 92, further comprising testing the one or more additional drugs in a physiologically relevant assay or model to evaluate efficacy in treating the disease or disorder.
94. A method for identifying a drug that modulates a predetermined signaling pathway associated with a disease or disorder, the method comprising: testing a principal drug set on a set of cells capable of signal transduction through a predetermined signaling pathway, wherein the principal drug set is a subset of a diversity drug set, wherein, for each drug within the principal drug set, a different sample of cells is cultured in the presence of the drug, and wherein a cellular phenotype associated with the signaling pathway is detected or measured to obtain a test result, thereby obtaining a set of test results for the drugs in the principal drug set; providing the set of test results to a computing system configured to perform a method as recited in any one of claims 1 to 12; receiving predicted efficacies of drugs in the diversity drug set from the computing system; selecting, from the diversity drug set, a drug that is identified as efficacious for modulating the signaling pathway based on the predicted efficacies.
95. The method of claim 94, wherein the identified drug is identified as an activator of the signaling pathway.
96. The method of claim 94, wherein the identified drug is identified as an inhibitor of the signaling pathway.
97. The method of any one of claims 94 to 96, wherein a plurality of drugs are identified as efficacious for modulating the signaling pathway.
98. The method of any one of claims 94 to 97, wherein the cellular phenotype is selected from the group consisting of cell growth, cell death, cellular differentiation, and a change in transcriptional activity.
99. The method of any one of claims 94 to 98, wherein the cells are from a cell line.
100. The method of any one of claims 94 to 99, wherein the cells are cultured with the principal drug set in the presence of a natural ligand that initiates signal transduction through the signaling pathway.