Active learning platform for scalable combination drug screens

BATCHIE addresses the scalability issues in combination drug screening by using adaptive experimental design and probabilistic modeling to efficiently identify effective drug combinations, reducing the experimental burden while maintaining high predictive accuracy.

WO2025117480A1PCT designated stage expired Publication Date: 2025-06-05MEMORIAL SLOAN KETTERING CANCER CENT +2
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/057349
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-11-25
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Current combination drug screening methods face significant challenges due to the vast number of possible experiments, which grows exponentially with the number of conditions, drugs, doses, and combinations, making it impractical to conduct exhaustive pairwise combination screens even with high-throughput technologies.

Method used

The BATCHIE approach employs an adaptive experimental design that dynamically conducts experiments in batches, using information theory and probabilistic modeling to maximize informativeness. This involves fitting a Bayesian tensor factorization model to data from initial experiments, simulating plausible outcomes for candidate experiments, and updating the model based on new data to prioritize the most informative experiments.

Benefits of technology

BATCHIE significantly reduces the number of experiments needed to discover highly effective and synergistic drug combinations, achieving high accuracy in predicting novel combinations and detecting synergies, as demonstrated in a combination screen conducted on a 206 drug library across pediatric cancer cell lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024057349_05062025_PF_FP_ABST
    Figure US2024057349_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is an approach to large-scale combination drug screens. Example versions employ an active learning platform as part of an orthogonal approach that includes conducting experiments dynamically in batches. Each batch may be designed to be maximally informative based on the results of previous experiments. Example designs have been demonstrated to discover highly effective and synergistic combinations of therapeutics. The disclosed approach provides for models that can identify a panel of top therapeutics ("drug") combinations for various diseases. Results demonstrate that adaptive experiments enable large-scale unbiased combination drug screens with a relatively small number of experiments, thereby powering a new wave of combination drug discoveries.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]Atty. Dkt. No.: 115872-3136 ACTIVE LEARNING PLATFORM FOR SCALABLE COMBINATION DRUG SCREENS CROSS-REFERENCE TO RELATED PATENT APPLICATIONS This application claims priority to U.S. Application No. 63 / 603,785, filed November 29, 2023, the content of which is incorporated herein by reference in its entirety. BACKGROUND Single-agent treatment interventions in cancers, viruses, and bacterial infections impose evolutionary selective pressures that can lead to therapeutic resistance and poor outcomes for patients. Combination therapies have the ability to constrain multiple potential avenues of evolutionary escape and thus reduce the likelihood of treatment resistance. Consequently, rational combination therapies have long formed the basis for rapidly evolving pathogens like HIV and are increasingly seen as the future of antibiotics and cancer therapies. Screening for novel effective drug combinations presents the singular challenge of scale. The number of possible experiments in a combination screen grows at a rate of n × md× tdfor n conditions, m drugs, t doses, and d-way combinations. For instance, a single-agent screen of 100 drugs and 50 cell lines at 5 doses would only be 25K experiments whereas a pairwise drug screen on the same libraries would require 12.5M experiments. Given the rapid growth of the experimental design space, even the most efficient high- throughput screening team would struggle to conduct an exhaustive pairwise combination screen of a modest-sized drug library over a modest number of doses and cell lines. Conducting such a screen is currently a substantial undertaking requiring large funding, detailed planning, advanced equipment, and several years of experiments. Thus, even the largest published combination screens have been restricted to libraries of less than 120 drugs. SUMMARY OF THE DISCLOSURE In some aspects, embodiments of the disclosed approach involve an orthogonal approach (referred to herein as “BATCHIE”) that conducts experiments dynamically in batches. BATCHIE seeks to design each batch so as to be maximally informative, or at least more informative, based on the results of previous experiments. -1- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 Information theory and probabilistic modeling may applied to quantify and enhance informativeness. On retrospective experiments from previous large-scale screens, BATCHIE experimental designs rapidly discover highly effective and synergistic combinations of therapeutic agents (herein used interchangeably with “drugs”). To validate BATCHIE prospectively, a combination screen was conducted on a collection of pediatric cancer cell lines using a 206 drug library. After exploring only 4% of the 1.4M experiments possible, the BATCHIE model was highly accurate at predicting novel combinations and detecting synergies. Further, the model identified a panel of top combinations for Ewing sarcomas, all of which were experimentally confirmed to be effective, including the rational and translatable top hit of PARP plus topoisomerase I inhibition. In various embodiments, adaptive experiments enable large-scale unbiased combination drug screens with a relatively small number of experiments, thereby powering a new wave of combination drug discoveries. The disclosed approach may employ one or more biologics (which are capable of being screened in vitro) other than (or in addition to) cell lines. Example biologics include organoids, spheroids, tumoroids, and / or patient-derived cells. Accordingly, references to “cell line” in this disclosure are also applicable to (and, e.g., could be replaceable by) references to any biologics. In other aspects, embodiments of the disclosure relate to a method comprising fitting a model to data from an set of initial experiments. The method may comprise identifying a set of candidate experiments not already performed. The method may comprise simulating plausible outcomes for each experiment in the set of candidate experiments. The method may comprise updating the model based at least on the simulated outcomes and informativeness of each experiment in the set of candidate experiments. The method may comprise updating the model based on new data to generate a trained model. The model may be updated by repeating steps the above steps a number of times. The trained model may be configured to provide predictions of efficacious drug combinations. The method may comprise selecting initial drug combinations to screen. The method may comprise performing the set of initial experiments using the initial drug combinations on a set of biologics. The initial drug combinations may be combinations of drugs in a drug library. In various embodiments, the set of initial experiments are run using a set of candidates designed such that each biologic in the set of biologics and each drug in the drug library is covered by at least one experiment. In various embodiment, fitting the model -2- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 comprises employing probabilistic modeling. The certain embodiments, the probabilistic modeling may be or may comprise Bayesian modeling. In various embodiments, the set of candidate experiments are selected from potential experiments not in the set of initial experiments. In various embodiments, the set of candidate experiments comprises all possible experiments not in the set of initial experiments. In various embodiments, the number of times that the model is iteratively updated is based at least on how much the model changes following updating of the model. In various embodiments, the set of biologics comprises a plurality of cell lines. In various embodiments, the trained model provides predictions of synergistic drug combinations. In various embodiments, the method may comprise validating the model in vitro and / or validating the model in vivo. In yet other aspects, embodiments of the disclosure relate to a method comprising performing a first batch of experiments to obtain first data indicative of effects of a first set of drug combinations on one or more biologics. Each drug combination in the first set of drug combinations may comprise two or more drugs from a drug library. The method may comprise generating or otherwise obtaining a model based at least on the first data. The method may comprise simulating, based at least on the model, a second batch of experiments to obtain second data. The second batch of experiments may correspond to a second set of drug combinations. Each drug combination in the second set of drug combinations comprising two or more drugs from the drug library. The second set of drug combinations may comprise drug combinations not in the first set of drug combinations. The method may comprise updating the model based at least on the second data. The method may comprise scoring each experiment in the second batch of experiments based at least on changes to the model resulting from the updating. The method may comprise performing a third batch of experiments to obtain third data, the third batch of experiments corresponding to a third set of drug combinations. The method may comprise updating the model based at least on the third data to obtain a trained model. The trained model may predict efficacious combinations of drugs from the drug library or from a second drug library with different drugs. In various embodiments, the method may comprise simulating and performing additional experiments to update the trained model. In various embodiments, the first batch of experiments may be designed such that each biologic in the one or more biologics and each drug in the drug library is covered by at least one experiment. In various embodiments, generating or otherwise obtaining the model comprises fitting a probabilistic model to the -3- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 model. In various embodiments, the probabilistic model is or comprises a Bayesian model. In various embodiments, the second batch of experiments are selected from potential experiments not in the first batch of experiments. In various embodiments, the second batch of candidate experiments comprises all possible experiments not in the first batch of experiments. In various embodiments, the trained model provides predictions of synergistic drug combinations. In yet other aspects, embodiments of the disclosure comprise using a model to predict one or more efficacious combinations of drugs from a drug library. The model may have been trained by a training process. The training process may comprise fitting a model to experimental results from a set of initial experiments to quantify uncertainty. The training process may comprise using the uncertainty to simulate plausible outcomes of candidate experiments using candidate drug combinations. The training process may comprise updating the model and quantifying informativeness of each candidate experiment based at least on how much the model changes due to the corresponding candidate experiment. The training process may comprise designing a subsequent batch of experiments based at least on informativeness of candidate experiments. The training process may comprise repeating steps a number of times to obtain a trained model. In various embodiments, the predicted combinations are for treating one or more conditions or diseases in subjects. In yet other aspects, embodiments of the disclosure relate to a method of training a model to predict one or more efficacious combinations of drugs. The method may comprise fitting a model to experimental results from one or more sets of experiments to quantify uncertainty. The method may comprise simulating, based at least on the uncertainty, plausible outcomes of candidate experiments using candidate drug combinations. The method may comprise updating the model based at least on informativeness of candidate experiments. In yet other aspects, embodiments of the disclosure relate to a system comprising one or more processors, the system configured to perform any of the methods disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS Figs.1a – 1c illustrate benefits of adaptive experimental design according to various example embodiments. (1a) Fixed-design approaches to combination entail -4- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 exhaustive enumeration of a small library or random subsampling of a large library. In either case, the resulting models generally are not powerful enough to find clinically-relevant discoveries. (1b) Adaptive experimental design allows exploration of a large library. By focusing on the most informative experiments, one can create a powerful model that is capable of finding clinically-relevant discoveries. Figs. 2a – 2g depict a BATCHIE overview according to various example embodiments. (2a) The BATCHIE workflow begins by specifying a cell line library, a drug library, and an initial ‘seed batch’ of plates to cover every cell line and drug with at least one experiment. (2b) Selected plates are assembled, run, measured, and post-processed to obtain viability scores. Quality control (QC) checks filter out problematic wells. (2c) A Bayesian tensor factorization model is fit to the current data. Posterior samples are drawn via MCMC. (2d) The joint distributions of candidate experiments are estimated using the current set of posterior samples. (2e) The active learning criterion is applied to the joint distribution estimates to score the utility of individual experiments. (2f) The scores of individual experiments are aggregated to define the most informative batch of experiments to run next, possibly subject to design constraints. (2g) After terminating the active learning loop, the most recently fitted Bayesian model is used to predict top hits for individual combinations. These top hits can then be validated in vitro and, potentially, in vivo. Figs. 3a – 3f depict a retrospective study design and results according to various example embodiments. (3a) Retrospective studies are conducted by processing an existing dataset into candidate row / column plates of the kind considered in this work and simulating the choices made by both the Random and BATCHIE data collection procedures. After data collection, the models are trained on the data collected by each of the procedures and evaluated for accuracy. (3b) Statistics for the datasets used in our retrospective studies. (3c) Heatmaps showing the fraction of possible experiments observed for each drug combination. (3d) BATCHIE outperforms Random when both are evaluated after 15 rounds of data collection; p-values derived from a Mann-Whitney U-Test. (3e) The number of additional experiments needed for Random to achieve comparable performance to BATCHIE grows with the number of BATCHIE rounds across all datasets. (3f) The number of additional batches needed for Random to achieve comparable performance to BATCHIE grows with the size of the cell line library; ρspearis Spearman’s rho for nonparametric rank correlation. -5- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 Figs.4a – 4g depict a prospective study design and random validation results according to various example embodiments. (4a) Prospective study cell line library broken down by type. (4b) Prospective study drug library broken down by mechanism of action. (4c) Overview of prospective study: after 15 rounds of BATCHIE data collection, approximately 4% of possible combinations were observed. (4d) Observation breakdown by cell line and drug mechanism of action; Pearson’s chi-square test p < 10–63 where the null hypothesis is that experiments were sampled uniformly at random. Colors in the left panel match the corresponding type colors from panel (4a). (4e) Scatter plot of mean BATCHIE predictions v.s. observed viabilities on random validation data. Orange line indicates regression of predictions onto observations, and black line denotes the identity line. (4f) Pearson correlation between predictions and observed viabilities, broken down by cell line, observation status, concentration and cancer type. All associated p-values satisfy p < 10–25. (4g) ROC curve for synergy identification on random validation data. Synergy is defined here as having an observed Bliss score larger than 0.25. ρ in Figs.4e and 4f is Pearson’s ρ correlation coefficient. Figs. 5a – 5e depict prospective study validation of top hits according to various example embodiments. (5a) Pipeline for top hit selection and validation in the prospective study. The model fit on BATCHIE-collected data is used to simulate outcomes for drug combinations at the low (0.1µM) concentration. Those simulated values are collected into confidence-rated predictions of TI values. The top hits are then selected for further in vitro experimentation, collecting observations over the full dose-response matrix. (5b) Observation status of selected top hits in BATCHIE-collected data. (5c) Mean predictions and observed viabilities for selected top hits at 0.1µM concentration. Pearson’s correlation ρ > 0.92 with p-value p < 10–28. (5d) Histogram of observed single model EWS TI values for combinations at 0.1µM concentration in BATCHIE-collected data (gray) and top predicted hits (orange). Percentiles are drawn with respect to BATCHIE-collected data. (5e) Observed AUCs with respect to the TI dose-response surface for selected top hits and not top hits. p-values for Figs.5d and 5e computed using a Mann-Whitney U Test. Figs. 6a – 6f depict more retrospective results according to various example embodiments. (6a) Example dose-response curve fits for completely unseen pairs of drugs combos and cell lines in each of the datasets after 15 rounds of data collection. Model fit on the entire dataset (Full model) depicted for comparison. (6b, 6c) Batches and experiments -6- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 saved as a function of the target normalized accuracy. (6d) Batches saved as a function of the number of BATCHIE rounds. (6e) Average therapeutic index (TI) of the top 20 picks after 15 rounds of data collection. (6f) Area under the receiver operating characteristic (ROC) curve when selecting picks with therapeutic index in the 99th percentile. p-values for Figs. 6e and 6f computed using a Mann-Whitney U Test. Figs.7a – 7h illustrate more prospective study design information according to various example embodiments. (7a) Illustration of how combination plates are created by overlaying a row plate with a column plate. (7b) Cell line and combo plate breakdown for phase I / II. (7c) Availability of cell lines across BATCHIE rounds. (7d, 7e) Cross-validation accuracy of the model as a function of the rounds of data collection. (7f) Cancer-type breakdown of the high TI predictions at round 10. (7g) TI results for 3 combos tested on a Ewing’s sarcoma cell line at round 10 in comparison with the BATCHIE-collected data observed up to round 10. (7h) Predicted correlation between the 4 Ewing’s sarcoma lines added at round 11 and the other cell lines used in rounds 11-15. Predictions are with respect to the model trained on data collected up to, and including, round 11. Figs. 8a – 8e depict more prospective study results according to various example embodiments. (8a – 8c) Scatter plots of predictions on random validation set broken down by previous observation status, drug concentration, and cell type. ρ is Pearson’s ρ. (8d) Predicted correlation between the 4 Ewing’s sarcoma lines added at round 11 and the other cell lines used in rounds 11-15. Predictions are with respect to the model trained on data collected up to, and including, round 11. (8e) Predicted correlations between all cell lines used in rounds 11-15. Predictions are with respect to the model trained on data collected up to, and including, round 15. Figs. 9a – 9f depict retrospective Bliss results according to various example embodiments. (9a) Retrospective Bliss studies were conducted by first subtracting out the individual drug effects for each combination observation. Datasets were broken into combination plates, and the BATCHIE and Random strategies were simulated over the dataset. (9b) Volcano plots of p-values compared against the ratio of viability predicted by additive model vs the observed viability (cutoff at α = 0.05 threshold using Benjamini- Hochberg correction). (9c) Normalized accuracy on predicting holdout Bliss scores. (9d) Normalized AUC of the ROC curve for predicting synergy. (9e) Normalized AUC of the ROC curve for predicting antagonism. (9f) Excess experiments needed by Random to achieve -7- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 accuracy of BATCHIE as a function of the number of BATCHIE rounds. Models in (9c – 9e) were trained on data collected up to, and including, round 15. Figs. 10a – 10g depict prospective triplet analysis according to various example embodiments. (10a) Triplet analysis was performed by measuring the 3-way dose response tensor at a grid of concentrations spanning 0.1nM - 400nM. Viabilities were smoothed by first applying isotonic regression, and then linearly interpolating the regressed values. TI scores were computed using the interpolated values over a fine-grid, and the combinations optimizing TI were selected. (10b) Observed medians of the Ewing viabilities for the single agents in the triplet analysis, along with the corresponding logistic curve fit and the corresponding IC50values. (10c) TI scores as a function of total concentration for the optimal combination of all three drugs (TZP+TOP+MIT) and the optimal pairwise combinations (TZP+TOP, TZP+MIT, TOP+MIT). (10d) Per-component concentrations for the optimal combination of the three drugs as a function of total concentration. (10e – 10g) Simplex plots of interpolated TI values as the proportion of the three drugs varies. The optimal value is denoted with a star. Fig.11 depicts an overall system that includes a computing system that may interface with other components to enable various aspects of the disclosed approach, according to various example embodiments. Fig.12 depicts a flowchart for generating and updating models for predicting efficacious drug combinations, according to various example embodiments. Fig. 13 depicts a block diagram of a representative server system and client computer system usable to implement certain embodiments of the present disclosure. DETAILED DESCRIPTION Large-scale combination drug screens are generally considered intractable due to the immense number of possible combinations. The intractability of combination drug screens has led to a flurry of machine learning methods for predicting novel drug combinations. But existing approaches use ad hoc fixed experimental designs then train machine learning models to impute novel combinations. The goal of such modeling is to use the predictive model to simulate experiments in silico and filter the list of combinations down to a set of top hits to be validated in vitro. However, predictive models are fundamentally -8- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 limited by the data on which they are trained. Small libraries, which allow for exhaustive enumeration of the combination landscape, are limited to discoveries involving those drugs in the library. If the library is too small, it simply may not contain any useful candidate combinations and thus machine learning models will extract no meaningful signals (Fig.1a). Larger libraries may contain useful combinations, but if a screen is performed on a random or fixed-design subset, then it is unlikely to generate observations of the most surprising and informative combinations. Models fit to pre-designed observations then have poor accuracy and are unable to confidently pinpoint the maximally useful combinations (Fig. 1b). Prior models have thus been unsuccessful as discovery tools for novel rational combinations. The disclosed approach can overcome both the wet lab scalability and predictive modeling challenges of combination screens by rethinking the experimental design approach. Rather than performing a fixed-design experiment and fitting a model post-hoc, various embodiments of the disclosed approach instead draw from Bayesian (or other probabilistic modeling) optimal sequential experimental design. In the sequential setup, experiments are conducted in small batches where each batch is designed adaptively based on the results of the previous batches. When designs are driven by a machine learning model that aims to acquire the most informative training data in each batch, the sequential experimental design task is known as active learning. Active learning has been considered in the drug discovery literature where the objective is to design de novo molecules aimed at better docking on a target. By contrast, various embodiments can employ active learning to discover rational combinations from large libraries of existing drugs. Various embodiments of “BATCHIE” provide a framework for orchestrating large-scale combination drug screens through sequential experimental design. BATCHIE may use a novel optimal Bayesian active learning strategy to design sequential experiments. These sequential designs may be theoretically near-optimal (see Supplement, below, for theory and proofs) and can guarantee that BATCHIE designs are highly efficient. Pragmatically, BATCHIE enables sequential experimental designs that will best improve any user-provided probabilistic (e.g., Bayesian) model. Thus, while example embodiments discussed herein implement an initial model for experiments, BATCHIE can alternatively be paired with any existing and future Bayesian machine learning method for combination drug response. In various embodiments, a result of a BATCHIE screen is a maximally informative -9- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 dataset and an optimal predictive model that enables the discovery of more effective combinations than in a fixed design (Fig.1c). This disclosure validates the empirical performance of BATCHIE using retrospective simulations and through a prospective study. The disclosure first uses data from large-scale combination screens to retrospectively simulate adaptive screens. BATCHIE consistently outperforms fixed designs in these simulations and better prioritizes effective combinations as top hits. The disclosure then implements BATCHIE in a drug screening facility and uses it to conduct a combination drug screen across a 206 drug library over 16 cancer cell lines, focusing on pediatric sarcomas. The BATCHIE screen generates a model with near-optimal predictive accuracy on random novel combinations. Further, The disclosure uses the model to prioritize ten combinations to validate experimentally for a panel of Ewing sarcoma lines; all ten combinations achieve a high therapeutic index score. The top identified hit corresponds to a PARP inhibitor (talazoparib) plus a topoisomerase I inhibitor (topotecan), a biologically rational combination and the subject of two of the five NCI- supported Phase II combination therapy clinical trials currently underway. Adaptive experimental design for combination drug screens: In various example embodiments, BATCHIE uses an active learning algorithm to choose the most informative data to collect in each sequential batch of experiments (Fig. 2). In the initial batch, BATCHIE uses a design of experiments approach to cover the drug and cell line space efficiently (Fig. 2a). The initial batch is then run (Fig. 2b) and used to train a Bayesian predictive model that estimates a distribution over drug combination responses for each cell line (Fig.2c). For subsequent batches, BATCHIE uses the model’s posterior distribution to simulate plausible outcomes of candidate combination experiments along with how they would change the model (Fig.2d). BATCHIE then measures how much each experiment is expected to reduce the posterior uncertainty over the drug responses (Fig. 2e) and uses a submodular approach to design a maximally informative batch (Fig.2f). After designing an optimal batch, the batch is run, the model is updated with the new results, and the next optimal batch is constructed. When the exploratory budget runs out or the model converges to concentrated posterior, for example, the active learning loop ends. The optimally trained model is then used to predict effective novel combinations that are prioritized for experimental validation (Fig.2g). -10- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 BATCHIE is compatible with any Bayesian model capable of modeling combination drug screen data. There have been many models developed for modeling this type of data. Integrating any of these models into BATCHIE would be possible by reformulating them as fully Bayesian models capable of quantifying posterior uncertainty. In an example implementation, the disclosure uses a hierarchical Bayesian tensor factorization model (Fig. 2c). The model includes embeddings for each cell line and each drug-dose, as well as embeddings that capture the effects of drug interactions. In various embodiments, the BATCHIE model assumes that the response of a combination of drugs on a cell line can be decomposed into the individual effects of the drugs and an interaction term. Specifically, themodel posits an embedding ∈ ℝ^^^^ for each cell line k and embeddings ^^^^ ∈ ℝ foreach drug-dose i, which capture the individual and interaction effects of the drug-dose, respectively. When drug-doses i and j are applied to cell line k, the logit-transformed viability is assumed to be normally distributed with mean µijksatisfying where ∈ ℝ is a drug-dose specific offset, ∈ ℝ is a cell line specificoffset, and ^^^^ ∈ ℝ. The variance of the normal distribution is assumed to be a globalparameter. Hierarchical priors are placed on top of the embeddings, offsets, and variance to automatically adapt to the complexity of the data. The priors are chosen to be conditionally conjugate with the likelihoods, allowing the entire model to sampled efficiently using blocked Gibbs sampling (see herein for modeling details). To help make BATCHIE more practical to implement in a high-throughput screening facility, example embodiments consider pools of experiments in the form of, for example, microwell plates (Fig. 7a). The plates may correspond to a set of candidates, and need not be plates but rather any vessel that can be employed with other robotic platforms. Each plate is formed by combining a single cell line with a ‘row’ plate and a ‘column’ plate. Here, a row plate is an n × m plate in which every well of a particular row contains the same drug at the same dose. Similarly, a column plate is of the same size, but constructed so that each column contains the same drug at the same dose. When a row plate and a column plate are overlaid, the resulting combination plate contains all possible combinations of their constituent drug-doses. Row and column plates are constructed with high and low control -11- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 wells so that all associated drug-doses are measured singly and viabilities can be computed. In this way, the number of plates that need to be stamped scales linearly with the drug library size, as opposed to quadratically as the number of possible combinations (further discussed herein). Validation of BATCHIE on existing combination datasets: A core goal of various embodiments of BATCHIE is to reduce the number of experiments needed to make useful discoveries. To benchmark the effectiveness of BATCHIE in pursuit of this goal, the disclosure discusses retrospective simulations that compared the performance of a BATCHIE-trained model after a small number (≤ 15) of rounds against (a) a model that was trained on the entire dataset and (b) a model that was trained on a random subset of the dataset matching the size and constraints of BATCHIE’s. The disclosure benchmarks example embodiments of BATCHIE using three large, publicly-available combination drug screen datasets: the NCI ALMANAC study (ALMANAC), the Genomics of Drug Sensitivity in Cancer combination screen (GDSC2), and Merck’s unbiased combination drug screen (MERCK). The three datasets differed in cell line library size (60 for ALMANAC, 126 for GDSC2, 39 for MERCK) and drug library size (104 for ALMANAC, 66 for GDSC2, 38 for MERCK). The ALMANAC screen covered the NCI-60 panel of cell lines, spanning leukemia, lung, colon, central nervous system, melanoma, ovarian, renal, prostate, and breast cancers. The GDSC2screen covered breast, colon and pancreatic cancer cell lines. The MERCK screen covered a panel of lung, ovarian, melanoma, colon, breast, and prostate cancer cell lines. All three screens used fixed experimental designs but differed in the subset of the combination space to explore. Consequently, the overall sparsity pattern varies substantially between datasets (Figs.3b and 3c). To simulate the behavior of various embodiments of BATCHIE in a realistic adaptive data collection screen, the existing datasets were divided into simulated plates (Fig. 3a). Plates were designed to replicate the statistics of the plates used in the prospective study discussed herein. The data collection processes (BATCHIE and Random) were given batch constraints that approximately 10% of cell lines be selected at every round and three combination plates be selected per chosen cell line. The Random and BATCHIE screens were stopped after 15 batches, the same as in the prospective study. 10% of all experiments were held out as a test set to evaluate predictive performance of the resulting models. Model -12- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 performance was compared relative to a model trained on the full set of experiments conducted, not including the test set. Simulations were repeated 30 times per dataset, with different randomly-organized row and column plates for stamping. At the end of 15 batches, the Random and BATCHIE models observed 1.7%- 20.4% of the full training dataset (1.66%-1.7% for ALMANAC, 10.3%-11.1% for GDSC2, and 19.3% -20.4% for MERCK). Total observed percentages varied due to the difference in the original screen designs. Across the three datasets, BATCHIE produced models with holdout R2accuracy within 5-7% of the models that were fit using all available training data and significantly outperformed the models trained on data collected by the Random strategy with the same number of rounds (Fig.3d). This disclosure also tracked the number of experiments that the Random strategy would need to perform in order to achieve comparable performance to BATCHIE. It can be observed that the number of experiments saved grew as a function of the number of BATCHIE rounds, with the number of experiments saved by round 15 numbering in the 10s of thousands for the ALMANAC and MERCK datasets to over 100 thousand for the GDSC2dataset (Fig. 3b). When looking at the excess number of experiments needed to achieve a certain normalized accuracy, BATCHIE shows an exponential speedup as the model accuracy threshold increases (Fig.6c). When measured by the number of batches needed for Random to achieve BATCHIE-level performance, a similar exponential trend can be seen, but the scale is more consistent across datasets (Fig.6b). Similar trends can be seen when measuring the number of excess batches required for the Random model to reach the equivalent performance of the BATCHIE model at each round (Fig.6d). To test how well example embodiments of BATCHIE scales with the size of the experimental landscape, smaller experimental spaces were simulated. Each dataset was restricted to a random subsample of 20%, 40%, 60%, or 80% of cell lines. In each dataset, a strong positive correlation (ρspear= 0.38, 0.64, 0.48, max(p) = 1.6 × 10–5) was observed between the size of the experimental landscape and the improvement offered by BATCHIE over Random (Fig.3d). These results indicate that BATCHIE screens become increasingly more efficient as the overall landscape becomes larger. While these results confirm that BATCHIE produces accurate models with very few experiments, average predictive accuracy alone does not ensure that the model will -13- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 be able to identify highly effective combinations. This is because desirable properties for treatments, such as high therapeutic index (TI), often correspond to extremal points, which by definition are not average. To evaluate the ability of BATCHIE to discover effective combinations, the BATCHIE-trained model was used to estimate the TI of all drug combinations using all pairs of cell lines as target and control. The top 20 predictions were chosen and their average observed TI was calculated. For all of the datasets, the BATCHIE- trained model selections had significantly higher TI (max(p) = 0.001) than those selected by the Random-trained model (Fig. 6e). It was also observed that BATCHIE was better at identifying high TI combos in terms of area under the curve (AUC) of the receiver operating characteristic (ROC) curve (Fig.6f). BATCHIE was also implemented in the setting where the goal is solely to model synergy or antagonism (Fig.9a). Here, the retrospective setup was adapted so that all single drug data was made available to the models before data collection and models only needed to predict Bliss scores. A synergy-only Bayesian hierarchical model was also implemented, making this benchmark an example of the flexibility of BATCHIE to adapt to new designs and alternative models (further discussed herein). Synergy detection is a challenging task as synergies are rare. In the three benchmark datasets, only 0.12%-1.19% of combinations are synergistic and 0.09%-0.82% are antagonistic (Fig.9b). Significant gains (max(p) = 10−7) were again observed in predictive accuracy on held out data when comparing BATCHIE to the Random design baseline (Fig. 9c). It was also confirmed that the BATCHIE-trained model is better at detecting top synergy hits and top antagonism hits (Figs. 9d, 9e). Across all three datasets, the BATCHIE model saves approximately 25K experiments compared to the Random strategy by batch 15 with similar upward trends in each dataset (Fig. 9f). A prospective pediatric sarcoma study with BATCHIE: Current treatments for many pediatric sarcomas have unacceptably high failure rates, particularly for metastatic and recurrent presentations. Ewing sarcoma (EWS), Rhabdomyosarcoma (RMS), and osteosarcoma (OST) are amongst the most common pediatric sarcomas in need of improved treatments. A a large-scale combination drug screen on pediatric cell lines was conducted, with a focus on pediatric sarcomas. The study covered 16 cell lines: 5 Ewing sarcoma lines (A673, MSKEWS-38338, MSKEWS-66647, SKNEP, TC71), 5 osteosarcoma lines (MG63, MSKOST-11890, SAOS-2, SJSA1 U2OS), 1 rhabdomyosarcoma line (MSKRMS-12808), 3 -14- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 other cancer cell lines (Kelly, MDA-MB-231, Wit49), and 2 non-cancer lines (RPE, BJ) (Fig. 4a). The non-cancer lines were included in the study to allow evaluation of meaningful notions of TI, as high activity in a target cell line alone does not necessarily translate to clinical utility. Mathematically, therapeutic index was defined to be the minimum of the control cell lines (i.e. RPE and BJ) minus the median of the target cell lines (e.g. all EWS lines). Thus, a high TI indicates a drug has high activity in the target lines and not the control lines. The chosen drug library consisted of 206 drugs, both FDA approved and investigational. In order to ensure adequate coverage of complementary mechanisms, the drug library was chosen to span a variety of targets (Fig.4b). The library included inhibitors of the most theoretically promising targets for Ewing sarcoma and osteosarcoma such as PARP, CDK4 / 6, and CD99 as well as commonly-used chemotherapy drugs. Each drug was tested at two concentrations: 0.1µM (the low dose) and 1µM (the high dose). All drugs on a single row or column plate were plated at the same concentration, allowing each drug combination to occur at 4 dose combinations (low-low, low-high, high-low, and high-high). Each active learning batch consisted of choosing 3 cell lines and 3 combination plates for each chosen cell line, resulting in 9 plates in total per batch. The screen was divided into two phases (Fig.7b). Phase I started with a focus on OST, using 4 OST lines and a single line each for EWS and RMS. Phase I also included the three non-sarcoma cancer cell lines and used a single control line (RPE). After 10 rounds of BATCHIE, it was observed that the model was converging, as measured by a diminishing improvement of cross-validation accuracy (Fig. 7d). The screen was thus paused and the model’s predictions on the three sarcoma types was assessed. Five-times as many combinations predicted to have high therapeutic index in the EWS line compared to either the OST lines or the RMS line (Fig. 7f) were observed. As a preliminary test, three drugs (Clofarabine, Eltanexor, and Talazoparib) were selected, each of which was predicted to have high TI on the EWS lines at the low concentration when paired with another of the three at the low concentration. It was experimentally validated that the three hits had TI in the 90th percentile of all observed combinations through round 10 (Fig.7g). Since the phase I results suggested that the drug library may contain effective combinations particularly for EWS, phase II focused the screen on EWS. Four EWS lines were added and the other non-sarcoma cancer cell lines were removed. An additional control -15- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 line (BJ) was added to increase the robustness of the control set in TI estimates; a fifth OST line (MG-63) was also added. After the single initial seed batch for the newly added cell lines, the BATCHIE model quickly learned that the EWS lines were highly related. Three of the four EWS lines introduced in phase II were significantly positively correlated with the phase I EWS line and none of the other phase II lines (Fig.7h). Phase II proceeded for five rounds of BATCHIE until the model’s cross-validation accuracy was observed to be converging (Fig.7e). The adaptive portion of the screen ended and the process moved to the validation phase. In total, BATCHIE was run for 15 rounds of data collection, generating approximately 54K unique cell line, drug-dose pair combinations. As the full design space was approximately 1.4M possible experiments, BATCHIE explored approximately 4% of the total landscape (Fig.4c). This dataset exhibited significant variability (Pearson’s chi-square test, p < 10–63) in the sampling frequencies with respect to both cell lines and drug classes (Fig.4d), indicating that certain cell lines and drugs were more informative than others to the model. The correlations among predictions made by the final model reveal that it clearly learned to separate the EWS cell lines from the OST and control lines (Fig. 8e). The OST lines were less homogeneous in the model predictions, as expected due to OST being a disease marked by chromothripsis which leads to potentially hundreds of random chromosomal translocations and thus high genomic heterogeneity compared to relatively stable EWS cancers. Validation of BATCHIE predictions on random unseen combinations: To evaluate the accuracy of the BATCHIE-trained model, a test set of randomly selected combination plates was constructed from the remaining unexplored experimental space. Example embodiments randomly selected from among the experimental plates that had no overlap in any cell line, drug-dose pair combinations with the BATCHIE-collected data. One plate was selected for every cell line that was under investigation in phase II. One of the test plates (cell line MG-63) did not pass the quality control checks (discussed below) and was excluded from performance measurement. The BATCHIE model predictions were highly accurate on the unseen validation data (Fig.4e, Pearson’s ρ=0.91, p < 10–30). The accuracy of the model was robust to stratification by cell line, previous observation status, drug concentration, and cancer type (Fig.4f). Indeed, for 10 of the 11 cell lines, Pearson’s ρ was above 0.84, and was above 0.91 -16- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 for 9 of the 11. The previous observation status of the drug on the cell line also appeared not to make a large difference, as the correlation remained above 0.91 regardless of whether one, both, or neither of the drugs had previously been observed on the chosen cell lines. It was observed that there was dip in performance to ρ = 0.83 when restricting to the low-low doses. This decrease in performance however is confounded since by chance half of the low-low validation set was on the SJSA-1 cell line, the cell line on which BATCHIE performed worst. SJSA-1 is an OST line that shares little predictive similarity to the other OST lines (Fig.8e), which may explain its relative difficulty in the test set. Similar to previous combination cell line screens very few synergistic combinations were found in the random validation set. Of the 3465 observed combinations, only 13 (0.004%) resulted in a Bliss score larger than 0.25. Nevertheless, the BATCHIE model accurately identified these combinations (Fig. 4g), achieving an AUC of the ROC curve of 0.899 (p < 10–6by permutation test over 100K random permutations). BATCHIE discovers rational drug combinations for Ewing sarcoma: To validate BATCHIE’s ability to discover effective combinations, the BATCHIE model was used to identify combinations with an expected high therapeutic index. The drug combinations were ranked by looking at their predicted viabilities at low concentrations and taking the difference between the minimum predicted viability on the two non-cancer lines and the median viability on the OST and EWS lines. It was found that no combinations were predicted to be robust across all OST lines such that the TI would be high. However, several drug combinations that were predicted to have high TI across a range of EWS lines were identified. Ten of the top-ranked candidates were selected and a fine-grained dose-response matrix spanning 0.006nM - 400nM was collected via four-fold dilution (discussed below), which included the low concentration combination (Fig. 5a). Thirteen negative control combinations that exhibited a large predicted differential effect for at least two cell lines but were not predicted to have a high TI over the five EWS lines were also selected. The top hit selections for EWS had not been observed in the training data at the low concentration for any cell line, and the general observation pattern was sparse (Fig. 5b). Nevertheless, at the low-concentration there was a strong correlation (Pearson’s ρ = 0.92, p < 10–28) between the predicted viabilities and the observed viabilities (Fig.5c). This accuracy at the viability level translated to the observed TIs being large, with the median TI -17- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 score in the top hit predictions being higher than the 99th percentile of observed TI scores in the 54K training observations (Fig.5d). Although the top hits were chosen solely on the basis of their predicted TI at a specific concentration, it was found that they generally exhibited high TI across a wide range of dose pairs. After computing the TI for each entry of the dose-response matrix, the area under the TI surface (3d curve) was calculated and it was observed that the top hits exhibited significantly higher AUC values than the reference combinations (p = 0.0018, Fig. 5e). In the example embodiment, the selected combinations exhibit biologically plausible rationales in Ewing sarcomas. Ewing sarcomas frequently exhibit EWS-FLI1 genomic fusions, which tend to interact with the DNA repair protein PARP-1. Talazoparib is a PARP inhibitor, and combining PARP inhibitors with treatments that induce DNA damage has previously been shown to be lead to cytotoxicity in preclinical Ewing sarcoma studies. These observations have led to clinical trials combining PARP inhibitors with irinotecan (a topoisomerase 1 inhibitor) and temozolomide (an alkylating agent) for Ewing sarcoma and related cancers. Topotecan is a topoisomerase 1 inhibitor, mitomycin is an alkylating agent that cross-links complementary DNA strands, and epirubicin is an anthracycline that blocks the action of topoisomerase 2. Thus, these selected combinations with talazoparib may facilitate the utility of PARP inhibition by accelerating DNA damage. Of the remaining drugs, cytarabine and GSK1324726A, may act by reducing the overall abundance of the EWS-FLI1. This has been directly shown for cytarabine in vitro. On the other hand, GSK1324726A is a BET bromadine inhibitor, and it has been shown that BET bromadine proteins are required for EWS-FLI1 transcription. Clofarabine and cladribine are deamination resistant analogues of deoxyadenosine, and as such interfere with DNA synthesis through incorporation into DNA. Clofarabine and cladribine have also been shown to inhibit Ewing sarcoma growth in vitro by binding to C99, which is overexpressed in Ewing sarcoma. EWS-FLI1 also suppresses SPRY1, a downstream feedback inhibitor of certain Ras-activating receptors. Tipifarnib is a farnesyltransferase inhibitors that interferes with the Ras signaling pathway. Combining tipifarnib with drugs that induce DNA damage can be seen as targeting two separate downstream effects of EWS-FLI1. -18- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 In various embodiments, higher order combinations can be discovered through analysis of the pairwise screens. To demonstrate this, a triplet screen was run on the drugs talazoparib, topotecan, and mitomycin over a fine-grained grid spanning 0.1nM - 400nM (see Methods, Fig.10a). Each pairwise combination of the three drugs appeared in the EWS top hits screen, suggesting that the triplet combination of the three would enable additional efficacy. At the single agent level, talazoparib was observed to have low activity with an IC50 nearly 40x higher than topotecan and 4x higher than mitomycin (Fig. 10b). Pairwise combinations with talazoparib yielded higher TI for both mitomycin and talazoparib. However, isotonic interpolation analysis revealed that for many choices of cumulative concentration, the addition of mitomycin did not lead to substantial improvements over the combination of talazoparib and topotecan (Fig. 10c). Instead, the optimal concentration strategy allocates more towards talazoparib as the total concentration increases but actually reduces the other two drugs (Figs.10d – 10g). This further supports preclinical evidence that PARP inhibitors to sensitize EWS cells to DNA damage with less of the two DNA damaging drugs needed as talazoparib dosing increases. Overall, the inability of the triplet combination to meaningfully improve on the pairwise score indicates that simple “pairwise additivity” is insufficient to detect effective higher order combinations. Methods Bayesian tensor factorization model for predicting combination drug response: Example embodiments use a hierarchical generative model that simultaneously models both single drug and combination drug observations. Example embodiments treat each drug at each dose individually as a single drug-dose. Observations can be modeled as -19- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 where ^^^^(^^^^)^^^^^^^^is the n-th logit-transformed viability measurement of applying drug-dose i to cell line k. Similarly, ^^^^(^^^^)^^^^^^^^^^^^is the logit-transformed viability measurement of applying the combination of drug-doses i and j to cell line k. Example embodiments also have that τ is the global precision of the observations, µikis the mean response of applying i to k, ∆ijkis the combination effect of i and j applied to k, and α is the global mean of all the observations.^^^^(^^^^) ∈ the embedding of cell line k with the following generative process:^^^^(^^^^)^^^^ ~^^^^(0, 1⁄ ^^^^^^^^ ) ^^^^^^^^~Gamma(^^^^^^^^, 1)where ^^^^ ∈ ℝ^^^^ is the vector of precisions for each cell line embedding coordinate, and itfollows a Gamma process prior with elements γs. The hyper-parameters of the Gamma process are chosen as a1= 2 and as= 3 for s ≥ 2. The vector ∈ ℝ^^^^ is the order ℓ (for ℓ ∈ {1, 2}) embedding of drug-dose iwith the following generative model: ∅(^^^^)ℓ,^^^^~Exponential(1), where ^^^^(^^^^)ℓ ∈ ℝ^^^^ is a vector of precisions for the embedding of drug-dose i.Also included in the model are offsets ∈ ℝ (for drug-dose i) ∈(for cell line k). These follow the generative process: -20- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 ∅(^^^^)0~Exponential(1) ^^^^0~Exponential(1), where ^^^^(^^^^)0 and τ0are precisions. To fit this model, example embodiments can utilize Gibbs sampling to sample from the posterior distribution, since all of the relevant priors are conditionally conjugate. Bayesian model for predicting combination drug synergy: For the pure synergy modeling setting, example embodiments use the following model: ^^^^~Gamma(1.1, 1.1)Here, ^^^^(^^^^)^^^^^^^^^^^^is the n-th synergy score between drug-doses i and j on cell line k. It is calculated as ^^^^(^^^^)^^^^^^^^^^^^ = ^̅^^^^^^^^^^^^̅^^^^^^^^^^^ − is the n-th observed viability of applying i and j to k, and ^̅^^^^^^^^^^^is the average observed viability of applying i to k. When ^̅^^^^^^^^^^^is not available because drug-dose i was not tested directly on k, then it is imputed by linear interpolation (in log-concentration space) of neighboring concentrations of the same drug. µijkis the mean synergy value of applying i and j to k and τ is globalobservational precision, ∈ ℝ^^^^ is the embedding of cell line k that follows the sameprior as its counterpart in the previous model: ^^^^^^^^~Gamma(^^^^^^^^, 1),-21- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136where ^^^^ ∈ ℝ^^^^ is the vector of precisions for each cell line embedding coordinate, and itfollows a Gamma process prior with elements γs. The hyper-parameters of the Gamma process are chosen as a1= 2 and as= 3 for s ≥ 2. The drug-dose embeddings ∈ ℝ^^^^ also follow the same prior as thecounterparts in the previous model: where ^^^^(^^^^)ℓ ∈ ℝ^^^^ is a vector of precisions for the embedding of drug-dose i.This example model is also fit via Gibbs sampling. Active learning algorithm: Example embodiments of an active learning procedure is a generalization of an optimal active learning procedure called Diameter-based Active Learning (DBAL). Example embodiments apply to general probabilistic models that include or consist of an experimental space ^^^^, an outcome space ^^^^, and a set of parameters Θ. Example embodiments assume that the likelihoods factorize, so that for a sequence(^^^^1,^^^^1), … , (^^^^^^^^, ^^^^^^^^) ∈ ^^^^ × ^^^^ and parameter ^^^^∈ Θ, example embodiments have Aplate of experiments is a sequence ^^^^ = (^^^^1, … , ^^^^^^^^) of experiments ^^^^^^^^ = ^^^^.For a corresponding set of outcomes (^^^^1, … , ^^^^^^^^), example embodiments can use the shorthandy to denote the sequence and ^^^^^^^^ to denote the likelihood ofoutcome sequence for a given plate. In the combination drug setting, the parameters θ include all of the parameters from the tensor factorization model, i.e. the µijk’s, the , τ, etc. The experiment space includes or consists of triples (i, j, k) and pairs (i, k) where i and j are drug-doses and k is a cell line. Outcomes in this setting are logit-transformed viabilities, and so the outcome spacecorresponds to the reals, i.e. ^^^^ = ℝ.-22- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 Given a prior distribution π over Θ and a set of observations(^^^^1,^^^^1), … , (^^^^^^^^, ^^^^^^^^) ∈ ^^^^ × ^^^^, the posterior distribution over Θ is given by Let d(·,·) be a bounded, non-negative, symmetric distance over Θ. The goal of DBAL-style active learning procedures is to run batches of experiments that will rapidly lead to a posterior πnwith small average diameter: avg − diam Given a plate P and a current posterior distribution πn, example embodiments of the active learning strategy assign an ideal score to each plate: where and Hθ(P) is the Shannon entropy of the plate P under θ, i.e. In the case of the normal likelihood (among others), example embodiments can explicitly compute the functions Lθ*(θ, θ'; P) and Hθ(P). Since computing sn(P) requires integrating over the posterior, it is not done directly in example embodiments. Instead, example embodiments compute a Monte Carlo approximation of sn(P) by sampling θ1, ... θm~ πnand computing This sum can be further approximated by sub-sampling triples (i, j, k) and computing the Monte Carlo average only over the selected triples. -23- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 Armed with the estimator ^̂^^^^^^^(^^^^), example embodiments of the active learning procedure enumerate a set of candidate plates P1,…, PTand select the plate Pi*with the lowest score ^̂^^^^^^^(^^^^^^^^∗). In the Supplement (below), it is shown that this strategy leads to provably near-optimal guarantees. Specifically, it is proved that the convergence rate of example embodiments of BATCHIE is upper-bounded by a function of a problem-specific parameter, called the splitting index, that determines the complexity of the learning problem in the sense that any active learning strategy, regardless of computational power, must have a convergence rate that is lower-bounded by this same splitting index. To select a batch of plates, example embodiments select sequentially. Example embodiments first select ^^^^^^^^1∗as the plate minimizing Having selectedplates ^^^^^^^^1∗ , … ,^^^^^^^^∗^^^^ , example embodiments select plate ^^^^^^^^∗^^^^+1 as the plate minimizing^̂^^^^^^^^^^^^^^^∗1 , … where�^^^^^^^^∗1 , … ,^^^^^^^^∗^^^^+1� is the plate formed by concatenating the constituentplates ^^^^^^^^∗1 , … ,^^^^^^^^∗^^^^+1. Observe that this is equivalent to selecting ^^^^^^^^∗^^^^+1 conditioned on havingalready selected ^^^^^^^^1∗ , … ,^^^^^^^^∗^^^^ . This iterative strategy of optimization has been shown to enjoystrong theoretical guarantees. Retrospective simulations, data retrieval and preparation: The ALMANAC, GDSC2, and MERCK datasets were obtained from their respective sources. For the MERCK dataset, viabilities were provided. For the ALMANAC dataset, PercentGrowth values were provided. These were converted to viability scores using the formula PercentGrowth + 100 . 200 For the GDSC2dataset, well intensity values were provided, along with high control and low control intensities. These were converted to viability scores using the same formula as in the prospective study. For all studies, replicates were averaged to produce a single viability measurement for all recorded cell line, drug-dose 1, drug-dose 2 triplets. Drug-doses that were not present in combination experiments were dropped. Plate and random holdout construction: From the viability measurements, example embodiments constructed synthetic plates to closely match the plates used in the prospective sarcoma study. Drug-doses were randomly divided into groups of 20, and initial plates constructed by considering each cell line c and each pair of groups g, g’ and collecting -24- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 all measurements satisfying that the cell line is in c, one of the drug-doses is in g, and the other drug-dose is in g’. Due to the biased sampling of drug combinations in the three experimental designs, this resulted in plates of unequal sizes. Example embodiments corrected for this by greedily merging the two smallest plates within a cell line until a minimum size threshold was met. Example embodiments then performed a single pass through the plates for each cell line and merged the largest plate with the smallest plate. To make the align the plate sizes, example embodiments selected a threshold and dropped all plates smaller than the threshold and dropped observations from plates larger than the threshold until they had the same number of observations. The threshold was chosen to minimize the total number of observations removed. In each simulation, a random holdout set was created by subsampling 10% of the observations from every plate. Experimental design: For a given dataset and plate construction, example embodiments first created an initial covering plate by randomly selecting observations to greedily cover the cell lines and drug-doses. Given the same dataset, set of plates, and initial covering plate, example embodiments ran both the Random and BATCHIE methods. At each round of data collection, the methods picked K cell lines and 3 plates per cell line. K was selected to be 1 / 10th the number of cell lines in the dataset, rounded down (K = 6 for ALMANAC, K = 12 for GDSC2, and K = 3 for MERCK). Posterior inference: At each round, before selecting plates, 200 MCMC samples were drawn from the current posterior using 5 parallel chains, each with a burn-in period of 2000 steps and a thinning factor of 40. These samples were then used to evaluate accuracy on the holdout set and, for BATCHIE, used to select the next set of plates to observe. For each dataset, example embodiments ran the plate construction and simulation using 25 different random seeds. Cell line restrictions For the cell line restriction simulations, cell lines were subsampled uniformly at random. The number of cell lines chosen per round was adjusted accordingly (1 / 10th of the number of selected cell lines). The number of plates per cell line in each round remained 3. Metrics For each dataset construction, example embodiments trained a model using the full training set, i.e. everything except for the holdout validation set, by drawing 200 MCMC samples from the posterior using 5 parallel chains, each with a burnin period of -25- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 2000 steps and a thinning factor of 40. Given a holdout set of experiment / viability pairs(^^^^1,^^^^1), … , (^^^^^^^^,^^^^^^^^) ∈ ^^^^ × [0,1], a fully trained predictor ^^^f^ull:^^^^ → [0,1], and a candidatepredictor ^^^^:^^^^ → [0,1], the normalized accuracy is given by the ratio of R2 scores:Normalized accuracy where ^^^^2( ) ^^^^ ^^^^ = 1 −∑^^^^^^^^=1 (^^^^^^^^ − ^�^^^)and ^�^^^ = 1^^^^∑^^^^^^^^=1 ^^^^^^^^ is the empirical mean of observed viabilities.Efficiency gains / batches saved and experiments saved were computed by calculating the holdout R2of the BATCHIE-trained model (at round 15, unless otherwise specified) and then searching for the earliest round at which a RANDOM-trained model had comparable performance. BATCHIE and Random models were only compared within the same random seed, and therefore for the same dataset and plate set construction. Pediatric sarcoma combination screen Study design The prospective sarcoma study consisted of 15 rounds of data collection. At each round, different subsets of cell lines were available based on doubling time and the extent to which they had been used in previous rounds (Fig. 7c). Adaptive batches were constrained to 3 combination plates each for 3 available cell lines. In the first round and the eleventh round, unseen cell lines were introduced and the plates were selected to greedily cover unseen drug-doses and drug-dose combinations over the new cell lines. In the remaining rounds, plates were selected using the BATCHIE active learning procedure outlined above with the constraint that 3 separate cell lines be chosen, leading to 9 selected plates in total. In rounds 1-10, each plate was run singly. During the phase I validation, example embodiments identified that BATCHIE was sensitive to undetectable random well failures producing corrupted data. As such, in rounds 11-15, plates were run in and quality control checks flagged wells whose duplicates differed by more than 0.5 in viability. -26- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 Combination plates. Fig. 7a shows the general schematic for our combination plate setup. Plates consisted of 384 wells with 16 rows and 24 columns. For row plates, 15 of the rows consist of a single drug applied to each of the corresponding column wells at a particular concentration, excepting two columns that corresponded to high control and low control. The remaining row of the row plate was filled with Dimethylsulfoxide (DMSO), also with two control columns. For column plates, 21 of the columns consist of a single drug applied to each of the corresponding row wells at a particular concentration. The remaining three columns are the high and low control columns (aligned to spatially match the corresponding high / low columns in the row plate) and a column filled only with DMSO. When a row and column plate are combined, the resulting combination plate contains 15 × 21 = 315 wells that correspond to all combinations of the constituent drug- doses, 15 + 21 = 36 wells that correspond to all single drug-doses, 15 high-control wells, and 15 low-control wells. The drug library consisted of 206 drugs, with four of the drugs duplicated to allow for the resulting 210 drugs to be evenly divided over 14 row plates and 10 column plates. Each row and column plate was constructed at 2 different doses: 0.1 and 1 µM, leading to a total of 28 row plate choices and 20 column plate choices. Finally, a full plate consists of a cell line, a row plate, and a column plate. Viabilities for a non-control well w are calculated as count(well ^^^^) − average low-control countviability(well ^^^^) =average high-control count − average low-control countwhere count(well w) is the reading at well w and average high / low-control counts are the averages of the readings at the corresponding high / low-control wells. Therapeutic index: Example embodiments define the in vitro therapeutic index (TI) to be a differential score that compares two groups of viabilities: one for the target set of cell lines and one for the control set. A high TI corresponds to low viability for most target cells and high viability for all control cells. To calculate TI, example embodiments take the difference between the minimum viability on the control lines and the median viability on the target lines: TI(^^^^, ^^^^) = minimum V(^^^^, ^^^^;^^^^) ^^^^^^^^^^^^ ^^^^ ∈ Controls − median V(^^^^, ^^^^;^^^^) for ^^^^ ∈ Target.-27- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 where V(i, j; N) is the viability of drug-dose pair (i, j) on cell line N. Triplet analysis: For the three drugs talazoparib, topotecan, and mitomycin, example embodiments collected dose-response data on each of the 5 Ewing sarcoma cell lines and the 2 non-cancer cell lines. The drugs were tested on a regular grid spanning 0.1nM - 400nM, and the plates were run in duplicate. For each cell line, example embodiments fit a multi-dimensional isotonic regression to the observed viabilities, restricting the regressed variables to monotonically decrease as a function of dose. Mathematically, example embodiments solved minimize s.t. ^�^^^ 1 2 3 ≥ ^′1, ^^^^′2, ^^^^′′′ ′^^^^ , ^^^^ , ^^^^ ^^�^^^^^ 3 for all ^^^^ 1 ≥ ^^^^1,^^^^ 2 ≥ ^^^^2,^^^^3 ≥ ^^^^ 3,where Dxis the set of concentrations used on drug x, and ^^^^^^^^1,^^^^2,^^^^3is the mean observed viability applying talazoparib at concentration d1, topotecan at concentration d2, and mitomycin at concentration d3. For any candidate set of concentrations, example embodiments linearly (in log-concentration space) interpolated its viability from the smoothed values ^^�^^^^^^1,^^^^2,^^^^3. Discussion: Example embodiments of “BATCHIE” provide an active learning platform that enables large-scale combination drug screens. Various embodiments of the platform derive theory guaranteeing that the batches designed by BATCHIE will always be highly informative. Retrospective simulations on data from previous large-scale combination screens confirmed strong empirical performance of BATCHIE. A prospective study on pediatric cancer cell lines showed BATCHIE screens can enable the rapid discovery of highly efficacious and synergistic drug combinations within libraries of hundreds of drugs. The probabilistic modeling in example embodiments of BATCHIE is modular. The algorithm is able to take any Bayesian model and design optimal batches with respect to that model. Two different hierarchical models were evaluated, one focused on viability and another on synergy prediction. BATCHIE showed performance gains with both models. Any modeling improvements that lead to performance gains would be complementary to the gains from BATCHIE screens. Thus, as new predictive models continue to be developed, they can be readily integrated into BATCHIE for improved screening efficiency. -28- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 Both the retrospective and prospective experiments were conducted on cancer cell lines. Immortalized 2d cell lines have a number of well-understood limitations and better 3d models such as spheroids and organoids are rapidly being developed to replace them in drug screening. Example embodiments of BATCHIE screens transfer seamlessly to the 3d setting and are arguably more useful here since 3d models tend to have longer doubling times and require more expensive equipment and media, exacerbating the need for efficient screens. Example embodiments have focused on pairwise drug viability screens as they are the most common in the combination literature for illustration purposes. However, the BATCHIE approach can be readily adapted to screen the drug interactome for any measurable outcome where experiments are batched. New technologies are emerging that enable a wide range of phenotypic measurement, such as proteome-wide drug effects, but are currently limited to single-agent screens. Multiplexed CRISPR perturbation screens enable combinatorial screens but require specifying a small library of genes. BATCHIE could be used to design optimal libraries in order to efficiently discover synthetic lethal combinations. Drug combinations represent an increasingly important therapeutic strategy in cancer and other diseases. It is expected that the BATCHIE approach could be critical to overcoming the combinatorial explosion in the experimental design space as preclinical screens grow to larger libraries and higher-order combinations like triplets and quadruplets. In doing so, the disclose approach can play an integral role in enabling the discovery of new combination therapies. Supplemental results Let ^^^^ denote a data space and ^^^^ denote a response space. Let ^^^^ denote a marginal distribution over ^^^^. The goal here is to model the data with some parametric probabilistic model Pθ(·; ·), where θ lies in a parameter space Θ and Pθ(y; x) denotes theprobability (or density) of observing ^^^^ ∈ ^^^^ at data point ^^^^ ∈ ^^^^. The notation y ~ Pθ(x) willbe used to denote drawing y from the density Pθ(·; x). Example embodiments consider models that factorize across data points, thatis for ^^^^ , … , ^^^ ^^^^ ^^^^1 ^^^^^ ∈ ^^^^ and ^^^^1, … ,^^^^^^^^ ∈ ^^^^ , we have -29- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 For a data point ^^^^ ∈ ^^^^ and parameter θ ∈ Θ, example embodiments denotethe entropy of the response to x under model θ as Example embodiments take a Bayesian approach to learning. To that end, letπ denote a prior distribution over Θ. Given observations (^^^^1,^^^^1), … , (^^^^^^^^,^^^^^^^^), denote theposterior distribution as where ^^^^ ^^^^^ = ^^^^^^^^′~^^^^[∏^^^^^^^=1 ^^^^^^^^′(^^^^^^^^, ^^^^^^^^) ] is the normalizing constant to make πn integrate to one.Here, it may be assumed that we are in the well-specified Bayesian setting, i.e. there is some ground-truth θ*~ π, and when we query point xi, the observation yiis drawn from Pθ*(·; xi). The notation πn(y; x) will also be used to denote the posterior predictive density and the notation y ~ πn(x) to denote drawing y from the density πn(y; x). A risk-aligned distance is a function d : Θ × Θ → [0, 1] satisfying two properties for all θ, θ' ∈ Θ: ^ Identity. i.e., d(θ, θ) = 0. ^ Symmetry. i.e. d(θ, θ') = d(θ', θ). The requirement that d(θ, θ') ≤ 1 is not onerous - any smooth distance over a bounded space can be transformed into a distance that satisfies this requirement by rescaling. In example setups, a risk-aligned distance encodes the objective: if we committed to the model θ when the true model was θ*, then example embodiments expect to suffer a loss of d(θ, θ*). The goal in our setting is to find a posterior distribution πnwith small average diameter: -30- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 avg-diam To see that this is a reasonable objective, observe that if θ*~ π and(^^^^1,^^^^1), … , (^^^^^^^^, ^^^^^^^^) is generated according to Pθ* , then after observing this data, θ* isdistributed according to πn. If predictions are made by sampling a model from πn, the expected risk of this strategy is exactly the average diameter. Moreover, even without this Bayesian assumption, the risk of this strategy can still be bounded above as a function of the average diameter. Probabilistic DBAL (PDBAL): Example embodiments term the active learning selection criterion Probabilistic DBAL (PDBAL), as it can been as a generalization of the DBAL active learning criterion to the probabilistic setting. Formally, the criterion boilsdown to a score function: given ^^^^ ∈ ^^^^, Constant entropy models: When Θ parameterizes location models with fixed scale parameters, the entropy term is constant. Example embodiments rewrite the objective for these models as For readability, approximations of Eq. (2) will be discussed. Extending these ideas to Eq. (1) can easily be done when we can compute Hθ(x) in closed form (as is the case for Gaussians, Laplacians, t-distributions, and other common likelihoods) or approximate it sufficiently well. Ageneral approach to approximating Eq. (2) is to draw ^^^^1, … ,^^^^^^^^~^^^^^^^^ and Although Eq. (3) is an unbiased estimator, it may have large variance due to the sampling ^^^^^^^^~ ^^^^^^^^^^^^(^^^^). In some cases, this variance can be reduced by avoiding sampling the yi’s whenever we can compute the function -31- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 ^^^^(^^^^;^^^^1;^^^^2;^^^^3) = ^^^^^^^^~^^^^^^^^1(^^^^)�^^^^^^^^2(^^^^; ^^^^)^^^^^^^^3(^^^^; ^^^^)�.This allows example embodiments to compute the alternate approximation .(4) Computing the sums in Eq. (3) or Eq. (4) would take O(m3) time. Example embodiments can approximate these sums via Monte Carlo by subsampling Nmctriples (i, j, k). Thus, to compute an approximation of Eq. (1) for a set of B potential queries takes time . Algorithm 1 (below) presents the full PDBAL selection procedure. The following proposition shows that for the important case of Gaussian likelihoods, example embodiments can compute the function M in closed form. Proposition where ^^^^ = ^^^^2^^^^2 + ^^ 2 2 2 21 2 ^^2^^^^3 + ^^^^1^^^^3.Proofs are provided below. Theoretical results on PDBAL: For the purposes of this section, it will be assumed that all models induce the same entropy for a given x. Assumption 1: Given any ^^^^ ∈ ^^^^ , there is a value H(x) such that Hθ(x) = H(x)for all θ ∈ Θ. Algorithm 1: PDBAL selection Require: Candidate queries ^^^^1, ... , ^^^^^^^^ ∈ ^^^^, posterior distribution πn, Monte Carloparameters m, Nmc. Ensure: Next query xb. if M computable in closed form. then -32- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 Compute else Draw Compute end if end for return arg Assumption 1 is satisfied, for example, whenever Pθis a location model whose scale component is fixed or otherwise assumed to be independent of θ. Assumption 1 is not required to implement PDBAL, but rather only factors into the analysis in this section. For a prior π, observed data (x1:n, y1:n), and value ρ ∈ [0, 1], it can be said thata data point ^^^^ ∈ ^^^^ ρ-splits the posterior πn if^^^^^^^^(^^^^) ≤ (1 − ^^^^)avg-diam(^^^^^^^^), (5)where snis the objective function defined in Eq. (1). Intuitively, Eq. (5) captures the notion that there exists a query x that shrinks the average diameter by at least (1 – ρ) in expectation. It can be said that πnis (ρ, τ)-splittable if Pr^^^^~^^^^(^^^^ ^^^^-splits ^^^^) ≥ ^^^^. (6)Then for parameters ρ, ∊, τ ∈ (0, 1), it can be said that Θ (along with corresponding marginal distribution ^^^^) has splitting index (ρ, ∊, τ) if for every posterior πnover Θ satisfying avg-diam(πn) > ∊, πnis (ρ, τ)-splittable. The definition of splitting in Eq. (5) is similar to those provided in previous diameter-based active learning works. It corresponds to a requirement that posteriors which are not too concentrated (avg-diam > ∊) should have a reasonable number of good queries (at least τ% if sampled from ^^^^). One notable difference in this definition is the entropy term inside sn(x) in Eq. (5). In example, embodiments, without this term, a query with a noisy likelihood will be -33- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 given higher saliency simply because it produces a wide range of possible outcomes. The entropy term balances out this bias by penalizing queries with a low signal-to-noise ratio. A key observation in this analysis is that if a point that ρ-splits the current posterior πtis queried, then in expectation a certain potential function will decrease. Lemma 2: If we query a point xt+1that ρ-splits πt, then ^^^^^^^^^^^^+1[^^^^ 2^^^^+1 avg-diam(^^^^^^^^+1)] ≤ (1 − ^^^^)^^^^2^^^^ avg-diam(^^^^^^^^), In order to prove bounds on the performance of PDBAL, some assumptions about the complexity of the class Θ and the rates at which empirical entropies within this classconcentrate are made. For a sequence of data pairs ^^^^ = denote the projection of Θ onto ωn. That is: Θ|^^^^^^^^ = {^^^^^^^^(^^^^1;^^^^1), … ,^^^^^^^^(^^^^^^^^; ^^^^^^^^) ∶ ^^^^ ∈ Θ}.For a sequence ωnand parameter ∊ > 0, define N (∊, Θ|ωn, dll) as the size of the minimum cover of Θ|ωnwith respect to the distance Here, we consider log 00= 0. The uniform covering number Nll(∊, Θ, n) isgiven by We say that a class Θ is has log-likelihood dimension (c, d) if log^^^^ (^,Θ,^^^^) ≤ ^^^^ log� for n ≥ d. The definition of N (^,Θ, n) is exactly the same as uniform covering numbers, modulo the non-standard distance. In the language of statistical learning theory, the log-likelihood dimension is a bound on the metric entropy log Nll(∊, Θ, n). It will be assumed that the log-likelihood dimension is bounded. Assumption 2: Θ has log-likelihood dimension (c, d), -34- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 As an example of a class with bounded log-likelihood dimension, the following result shows how the complexity of a function class translates to the log-likelihood dimension of the corresponding Gaussian location model. Proposition 3: Fix σ2, d, B > 0. Let Θ denote the class of Gaussian locationmodels parameterized by a set of mean functions ℱ ⊂ {^^^^:^^^^ ^ [−^^^^,^^^^]} such that^^^^^^^^(^^^^; ^^^^) = ^^^^(^^^^|^^^^(^^^^),^^^^2). If the responses y lie in [–B, B] and the pseudo-dimension of ℱis bounded by d, then Θ has log-likelihood dimension (cB2 / σ2, d) for some universal constant c > 0. Recall a mean-zero random variable X is sub-Gamma with variance factor v > 0 and scale parameter c > 0 if for all λ ∈ (0, 1 / c). If we say that the class Θ is entropy sub-Gamma withvariance factor v > 0 and scale parameter c > 0 if the random variable ^^^^ = log ^^^^^^^^(^^^^;^^^^)− ^^^^(^^^^)is sub-Gaussian with variance factor v > 0 for all θ ∈ Θ and ^^^^ ∈ ^^^^, where Y ~ Pθ(x). Forthese bounds we will want Θ to be entropy sub-Gamma. Assumption 3: Θ is entropy sub-Gamma with variance v and scale c', Many likelihoods satisfy Assumption 3 including Gaussians, as illustrated by the following result. Proposition 4: Fix σ2> 0 and let Θ denote the class of Gaussian location models from Proposition 3. Then Θ is entropy sub-Gamma with variance factor 1 and scale parameter 1. Finally, we will require boundedness of both the entropy and the densities of models in Θ. Assumption 4: There are constants c1, c2≥ 0 such that Pθ(y; x) ≤ c1andexp(H(x)) ≤ c2 for all θ ∈ Θ, ^^^^ ∈ ^^^^, and ^^^^ ∈ ^^^^.-35- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 Upper bounds: Given the terminology above, we have the following guarantee on PDBAL. Theorem 5: Pick d ≥ 4 and suppose Assumptions 1 to 4 hold. If at every round t we make a query that ρ-splits πtand terminate when avg-diam(πt) ≤ ∊, then with probability 1 – δ, PDBAL terminates after fewer than ^^^^ ^^^^ 2 ^^^^ ^^^^^^^^ + ^^^^′ (^^^^ + ^^^^′)^^^^1avg-d^�max 1 2 iam(^^^^^^^^ ≤ ^^^)�^^^^ log� 1 2� ,log�� ,log ^^^^ ^^^^^^^^ ^^^^ ^^^^^^^^ ^^^^ ∈ queries with a posterior satisfying avg-diam(πT) ≤ ∊. As discussed above, Assumptions 1 to 4 are conditions on the complexity and form of Θ. In example embodiments, the requirement on the behavior of PDBAL can be guaranteed with high probability across all rounds t with enough unlabeled data and a fine enough Monte Carlo approximation of Eq. (1). Given Lemma 2, the proof of Theorem 5 takes two steps: 1. Showing that ^^^^^^^2^ avg-diam (πt) must decrease exponentially quickly. 2. Showing that Wtcannot decrease too quickly. The only way that both of these can hold is if avg-diam(πt) must also decrease quickly, proving Theorem 5. The first step, showing that ^^^^^^^2^ avg-diam(πt) decreases quickly, is formalized by the following lemma. Lemma 6: Suppose Assumption 4 holds. For any t ≥ 1 and δ > 0, if xiρ-splitsπi–1 for ^^^^ = 1, … , ^^^^, then with probability at least 1 – δ,^^^^2^^^^ avg-diam(^^^^^^^^) ≤ avg-diam The second step, showing that Wtdoes not decrease too quickly, is formalized by the following lemma. -36- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 Lemma 7: Suppose Assumption 2 and Assumption 3 hold. If θ*~ π and yi~Pθ* (·; xi) for ^^^^ = 1, … , ^^^^, then with probability at least 1 – δ Lower bounds: It will now be shown that in some cases, any optimal active learning strategy must have some dependence on the splitting index of the class. The first result along these lines is in the deterministic setting. Theorem 8: Let Θ denote a class of deterministic models that is not (ρ, ∊, τ)- splittable for some ρ, ∊ ∈ (0, 1 / 4) and τ ∈ (0, 1 / 2). Let π be any prior distribution satisfying avg-diam(π) ≥ 4∊ which is not (ρ, τ)-splittable. Then any active learning strategy that, with probability at least 5 / 6 (over the random samples from ^^^^ and the observed responses), finds a posterior distribution satisfying avg-diam data points or make at least12^^^^ queries. We can relax the constraint that our models are deterministic to the case where they have bounded entropy at the expense of a slightly weaker lower bound. Theorem 9: Let Θ denote a class of models that is not (ρ, ∊, τ)-splittable forsome τ, ∊ ∈ (0, 1 / 4) and τ ∈ (0, 1 / 2) such that H(x) = h < ρ3 / 2 / 6 and Pθ(y; x) ≤ 1 for all ^^^^ ∈ ^^^^,θ ∈ Θ and ^^^^ ∈ ^^^^. Let π be any prior distribution satisfying avg-diam(π) ≥ 4∊ which is not(ρ, τ)-splittable. Then any active learning strategy that, with probability at least 5 / 6 (over the random samples from ^^^^ and the observed responses), finds a posterior distribution satisfying avg-diam queries. The key ingredient to proving these lower bounds is demonstrating that splitting values are (approximately) sub-additive. Lemma 10: Let ρ1, ρ2satisfy ρ1+ ρ2< 1. Suppose x1ρ1-splits π, x2ρ2-splits π, and H(x1) = H(x2) = h. Then the following holds: ^ If h = 0, then the combined query (x1, x2) has splitting value at most ρ1+ ρ2. -37- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 ^ 2(ρ1+ ρ2). Proofs Proof of Proposition 1: From the form of the Gaussian likelihood, we have ^^^^2‖2 1 −‖^^^^ − ^ ‖22^^^^32^^^3 � ^^^^^^^^.The bias-variance decomposition of squared error implies that for any^^^^ , ... , ^^^^ ≥ 0 and ^^^^, ^^^^ , ... , ^^^^ ∈ ^^^^1 ^^^^ 1 ^^^^ ℝ , we have where ^^^^sum = ∑^^^^ ^^^^^^^^ and ^�^^^ =1 ^^^^sum ∑^^^^ ^^^^^^^^^^^^^^^^ . Applying this identity to theexponential above with the notation ^^^^we ^^^^^^^^~^^^^�^^^^1,^^^^12�[^^^^(^^^^; ^^^^2,^^^^22)^^^^(^^^^; ^^^^ 23,^^^^3)] -38- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 where the second-to-last line follows from the fact that the integral is exactly the normalizing constant of a spherical Gaussian with mean ^̅^^^ and variance 1 / τ and the last line follows from expanding the definition of τ. Any discrete random variable X taking values xi with probability pi for ^^^^ =1, … ,^^^^ satisfies the following: Applying this to the discrete random variable that takes value µiwith probability +^^^^22‖^^^^1 − ^^^^3‖2).Putting it all together gives us the desired identity. Proof of Lemma 2: Let ^^^^ = so that we have ^^^^^^^^^^^^^^^^^^^^ ∑^^^^^^^^=1 ^^^^(^^^^^^^^) . Observe that Zt is exactly the normalizing constant arising in Bayes’ rule: Thus, for any xt+1, yt+1, we have Moreover, by Bayes’s rule, we also have -39- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 Putting it all together, ^^^^2^^^^+1 avg-diam(^^^^^^^^+1) = ^^^^2^^^^+1 ^^^^^^^^,^^^^′~^^^^^^^^+1[^^^^(^^^^,^^^^′)] Taking expectations over yt+1and applying the definition of splitting finishes the argument. Proof of Proposition 3: Let (^^^^1,^^^^1), … , (^^^^^^^^,^^^^^^^^) be given. For θ, θ' ∈ Θ, Thus, Nll(∊, Θ, n) ≤ N1(σ2∊ / 2B, Θ, n), where N1denotes uniform covering with respect to distance. By known bounds on the covering number in terms of the pseudo- dimension, e.g. Theorem 18.4, we have ^^^^ ^^^^1(∈,Θ,^^^^) ≤ ^^^^(^^^^ + 1)�4^^^^^^^^^^^^ ∈� . Thus, The definition of log-likelihood dimension finishes the argument. -40- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 Proof of Proposition 4: For simplicity, let ^^^^ ∈ ℝ and σ2 > 0. Let ^^^^ (·; µ, σ2)denote the density of a Gaussian with mean µ and variance σ2. Let ^^^^ = 1 12+2log(2^^^^^^^^2)denote the entropy of ^^^^(∙;^^^^,^^^^2).Suppose ^^^^~^^^^(∙; ^^^^,^^^^2) and ^^^^ = log 1^^^^^^^^(^^^^;^^^^)− ^^^^, then( )2⋋^^^^⋋^^^^ − ^^^^⋋ ⋋ ^^^^[^^^^ ] = ^^^^�exp� log(2^^^^^^^^2) +⋋2−log(2^^^^^^^^2) − 2 2^^^^ 2 2 for λ < 1. Here, the last line follows from the fact that is chi-squared with one degree of freedom, and the known form of the chi-squared moment generating function. Thus, for λ ∈ (0, 1), we have log^^^^ where the inequality follows from the bound log(x) ≤ x – 1 for x > 0. Thus, Gaussian location models are entropy sub-Gamma with variance factor 1 and scale parameter 1. Proof of Lemma 6: For t ≥ 1, define ∆ = ^^^^ 1 − ^^^^^^2^^−1 avg-diam(^^^^^^^^−1). Let ℱ^^^^denote the sigma-field of all outcomes up to an including time t. Then if we query point xtwhich ρ- splits πt+1, the definition of splitting implies that ^^^^[∆^^^^+1 | ^^^^^^^^,^^^^^^^^] ≥ ^^^^.Let ^^^^^^^^ = ∑^^^^^^^^=1 (∆^^^^ − ^^^^) . The above implies that St is a submartingale.Moreover, if Pθ(y; x) ≤ c1uniformly for all θ, x, y and ≤ c2for all x, then ^^^^2 ( )0 ≤ ^^^^+1 avg − diam ^^^^+1^^^^^2^^^ avg − diam (^^^^+1) -41- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 This implies |^^^^ − ^^^^ | ≤ ^^^2 2^^^^+1 ^^^^ ^1^^^^2, and thus the Azuma-Hoeffding inequalitytells us that with probability at least 1 – δ, Proof of Lemma 7: To prove Lemma 7, we will use the following lower bound. Lemma 11: covering of Θ|ωn with respect to dll(·, ·). Let Θ1, … ,Θ|^^^^| be the induced Voronoi partition ofΘ (breaking ties arbitrarily). Then for any Θiand any θ*∈ Θi Proof. Fix Θiand let mi∈ M denote the Voronoi ‘center.’ Then for any θ, θ' ∈ Θi, we have ^^^^^^^^^^^^(^^^^,^^^^^^^^) + ^^^^^^^^^^^^(^^^^ ′^^^^,^^^^ ) ≤ 2^^^^,where we have used the fact that M is an ∊-cover. Thus, we have Finally, because Θ1, … ,Θ|^^^^| are disjoint, we have We will also require the following result which follows directly from well- known tail bounds for sub-Gamma random variables. Lemma 12: Suppose Θ is entropy sub-Gamma with variance factor v > 0 andscale parameter c > 0. If yi ~ Pθ(·; xi) for ^^^^ = 1, … , ^^^^, then with probability at least 1 – δ,-42- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 1 1 log + ^^^^ log. ^^^^ ^^^^ Proof. Observe that for independent mean-zero sub-Gamma random variables^^^^1, … ,^^^^^^^^ with variance factor v > 0 and scale parameter c > 0, the random variable ^^^^ = ∑^^^^ ^^^^^^^^satisfies log^^^^[^^^^ ] = log^^^^ for λ ∈ (0, 1 / c). Thus, Z is sub-Gamma with variance factor nv and scale parameter c. The argument is finished via standard concentration results on sub-Gamma random variables, e.g. Chapter 2.4. With Lemmas 11 and 12 in hand, we can turn to proving Lemma 7. Proof of Lemma 7: Recall that we can write Let ^^^^^^^^ =�(^^^^1,^^^^1), ... , (^^^^^^^^,^^^^^^^^)� denote our data. By Lemma 11, we have where Θ1, … ,Θ^^^^ is the Voronoi partition induced by a minimal 1-covering of Θ|ωn, Θi is anyof these partition elements, and θiis any element in Θi. Let S denote the set of indices i such that π(Θi) < δ / (2M ). Then we have Thus, if Θ*is the element of the partition that θ*falls into, we have with probability at least 1 – δ / 2 -43- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 where the last line follows from the fact that M ≤ (ct)d. Finally, observe that with probability at least 1 – δ / 2, Lemma 12 implies A union bound finishes the argument. Proof of Theorem 5: Combining Lemma 6 with a union bound, we have that with probability 1 – δ / 2 ^^^^ avg − diam( avg − diam( ) for all t ≥ 1, simultaneously. Similarly, combining Lemma 7 with a union bound gives us with probability at least 1 ≥ δ, for all t ≥ 1. Thus, with probability 1 ≥ δ, both of these occur simultaneously. Plugging in the value of T from the theorem statement, -44- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 2^^^^(^ )avg-diam(^^^^ ) ≤ avg-diam(^^ ) � ^^^ + 1^^^^ ^^ exp�−^^^^^^^^ + ^^^^1^^^^2 2^^^^ log^^^^+ ^^^^ log 4^^^^(^^^^ + 1) 4^^^^(^^^^ + 1)+2 + log + �2^^^^^^^^ lo ′^^^^g^^^^+�^^^^ + ^^^^ + 1 4^^^^(^^^ )≤ avg-diam(^^^^) exp�−^^^^^^^^ + 4^^^^1^^^^2� ^ + 1^^^^^^^^ log^^^^+ (^^^^ + ^^^^′ + 1) The above is less than ∊ when we have +1) log�3 (^^^^ + ^^^^′ + 1)4^^^^ 3avg-diam∙ � ,(^^^^)log �. ^^^^ ^^^^ ^^^^ ^^^^ Here, we have made use of the fact that if a ≥ 1, b ≥ e and x ≥ 9a log(ab), then x ≥ a log(bx(x + 1)). Proof of Lemma 10: Observe by the product measure assumption of Pθ(·; x1, x2), we have Let U be the random variable that takes on value α1(θ, θ', θ*) and let V denote the random variable that takes on value α2(θ, θ', θ*). Here, θ, θ', θ*occur with probability^^^^(^^^^)^^^^(^^^^′)^^^^(^^^^∗)^^^^(^^^^,^^^^′)avg-diam(^^^^). Then it is not hard to see that ^^^^[^^^^] = (1 − ^^^^1)^^^^−2ℎ and ^^^^[^^^^] =(1 − ^^^^ −2ℎ2)^^^^ . Note that U and V lie in the interval [0, 1] almost surely. Let A = 1 – U andB = 1 – V. Let us first consider the case where h = 0. Then we have ^^^^[^^^^^^^^] = 1 − ^^^^[^^^^] − ^^^^[^^^^] + ^^^^[^^^^^^^^] = 1 − (1 − ^^^^1) − (1 − ^^^^2) + ^^^^[^^^^^^^^]= ^^^^1 + ^^^^2 − 1 + ^^^^[^^^^^^^^].-45- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 Observe that AB ≥ 0 almost surely, and so we have ^^^^[^^^^^^^^] ≥ 1 − ^^^^1 − ^^^^2.Substituting in our definitions of U and V gives us the result. Now consider the case where 0 ≤ ℎ ≤ ^^^^1+^^^^26 . The same argument as before shows that where we have made the substitution ρ = ρ1+ ρ2. To prove the lemma, we will show that the above is greater than (1 – 2ρ)e–4h. This is equivalent to showing (2 − ^^^^)^^^^2ℎ − ^^^^4ℎ − 1 + 2^^^^ ≥ 0.The left-hand side is decreasing for h ≥ 0. Moreover, we also have theinequality Thus, Proof of Theorem 8: Let π be a prior distribution as in the theorem statement. Suppose we draw less than 1 / 2? unlabeled examples, then with probability at least (1 – τ)1 / 2τ≥ 1 / 2 none of these ρ-split π. Let us condition on this event. By induction on Lemma 10, we have that any collection of n ≤ 1 / ρ of these points does not nρ-split π. Suppose that we query n of these points (say ^^^^1, ... , ^^^^^^^^), and receive responses^^^^1, ... ,^^^^^^^^. Let πn denote this posterior. By Lemma 2, we have^^^^^^^^1:^^^^[^^^^2^^^^avg-diam(^^^^^^^^)] ≥ (1 − ^^^^^^^^)avg-diam(^^^^),where ^^^^^^^^ = ^^^^^^^^~^^^^ [^^^^^^^^(^^^^1:^^^^;^^^^1:^^^^)] ≤ 1. Thus,^^^^^^^^1:^^^^ [avg-diam(^^^^^^^^)] ≥ (1 − ^^^^^^^^)avg-diam(^^^^).For a random variable U satisfying U ≥ c almost surely, the reverse Markov inequality gives us -46- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 () ^^^^[^^^^] − ^^^^Pr ^^^^ > ^^^^ ≥^^^^ − ^^^^for any ^^^^ ≤ ^^^^[^^^^]. Applying this to the random variable and assuming ^^^^ ≤ , we have that avg-diam(πn) ≥ ∊ with probability at least 1 / 3. Putting it all together gives us the theorem statement. Proof of Theorem 9: We will use the following lemma. Lemma 13: Let ^^^^ ≤�1 / ^^^^, and let ℎ ∈ ℝ satisfy 0 ≤ h ≤. Suppose ^^^^1, ... , ^^^^^^^^all satisfy H(xi) = h ≤ ρ / 6n and have splitting index ≤ ρ. Then the combined query x1:nhas splitting index less than n2ρ. Proof. We will show the claim for n a power of 2. Extending to other integers is straightforward. The proof is by induction. Where we first observe that any subsequence^^^^1, ^^^^2, ... , ^^^^^^^^ ∈ {1, ... ,^^^^ } satisfies that Now for n = 1, then the claim trivially holds. For n ≥ 2, observe that by our inductive hypothesis, we have x / 2an ^^^^2^^^^ 1:nd xn / 2 + 1: n each have splitting index less than 4 . Applying Lemma 10, completes the argument. Turning to the proof of Theorem 9, let π be a prior distribution as in the theorem statement. Suppose we draw less than 1 / 2τ unlabeled examples, then with probability at least (1 – τ)1 / 2τ≥ 1 / 2 none of these ρ-split π. Let us condition on this event. Now let ^^^^ 12√^^^^. Using the fact that ^^^^ <�1 / ^^^^ and ℎ < ^^^^3 / 2 / 6, we can applyLemma 13, to see that any collection of n of these points does not n2ρ-split π. Lemma 2 then implies that diam(^^^^ )] ≥ (1 − ^^^^2^^^^)avg-diam(^^^^)-47- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 Recall ^^^^^^^^ = ^^^^ ∑^^^^^^^^=1 ^^^^(^^^^^^^^)^^^^^^^^~^^^^[∏^^^^^^^^=1 ^^^^^^^^(^^^^^^^^; ^^^^^^^^) ] . By our assumptions thatWn ≤ 3 / 2 almost surely. Thus, avg-diam(^^^^)For a random variable U satisfying U ≤ c almost surely, the reverse Markov inequality gives us [( ) ^^^^ ^^^^] − ^^^^Pr ^^^^ > ^^^^ ≥^^^^ − ^^^^for any ^^^^ ≤ ^^^^[^^^^]. Applying this to the random variable ^^^^ = avg-diam(^^^^^^^^)avg-diam(^^^^)and threshold ^^^^ =^^^^ avg-diam(^^^^), we have 2^^^^ ^^^^ ^^^^ where we have used the fact that ^^^^ ≤ 12�^^^^and ^^^^(^^^^) = ^^^^^^^^−^^^^^^^^−^^^^ is increasing in x when c ∈ (0, 1). Thus, (πn) ≥ ∊ with probability at least 1 / 3, finishing the argument. Referring to FIG.11, in various embodiments, a system 1100 may include a computing system 1110 (which may be or may include one or more computing devices, co- located or remote to each other), a condition detection system 1160 (capable of, e.g., obtaining data on candidates being evaluated), an information system 1170 (such as a management information system that may record and / or provide information regarding experiments), vessels 1175 (e.g., a multi-well plate or wells thereof), and an experimentation system 1180 (e.g., an automated system that may include robotic components for moving, preparing, and / or manipulating samples, performing tests or other experiments, and / or obtaining certain data needed for controlling the robotic components and / or for evaluating samples or test results). The computing system 1110 (e.g., one or more computing devices) may be used to control and / or exchange signals and / or data with condition detection system 1160, information system 1170, and / or experimentation system 1180, directly (e.g., through -48- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 wireless and / or wired communication) or indirectly via another component of system 1100 (e.g., via any combination of wireless and / or wired communication). In certain embodiments, computing system 1110 may be used to control and / or exchange data or other signals with condition detection system 1160, information system 1170, and / or experimentation system 1180. The computing system 1110 may include one or more processors and one or more volatile and / or non-volatile memories for storing computing code and data that are captured, acquired, recorded, and / or generated. The computing system 1110 may include a controller 112 that is configured to exchange control signals with condition detection system 1160, information system 1170, experimentation system 1180, and / or any components thereof, allowing the computing system 1110 to be used to control, for example, acquisition of test results, capture of images, acquisition of signals by sensors, positioning or repositioning of samples being tested or imaged, recording or obtaining other information, performing experiments, and / or running assays or tests. A transceiver 1114 allows the computing system 1110 to exchange readings, control commands, and / or other data or signals, wirelessly or via wires, directly or indirectly via networking protocols, with, for example, condition detection system 1160, information system 1170, and / or experimentation system 1180, or components thereof. One or more user interfaces 1116 allow the computing device 1110 to receive user inputs (e.g., via a keyboard, touchscreen, microphone, camera, biometric scanners, etc.) and provide outputs (e.g., via a display screen, audio speakers, light emitters, etc.) with users. The computing device 1110 may additionally include one or more databases 1118 for storing, for example, data acquired from one or more systems or devices, signals acquired via one or more sensors, images, biomarker signatures, etc. In some implementations, database 1118 (or portions thereof) may alternatively or additionally be part of another computing device that is co-located or remote (e.g., via “cloud computing”) and in communication with computing device 1110, condition detection system 1160, information system 1170, and / or experimentation system 1180 or components thereof. Condition detection system 1160 may include an imager 1162, which may be or may include, for example, any system or device that is involved in, for example, capturing images during or following experiments (e.g., cells or tissues). Imagers may include detectors for visible light and / or light in any frequencies of interest, such as (but not limited to) the -49- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 spectrum from infrared to ultraviolet. Imagers may employ any suitable optical components (e.g., lenses, mirrors, filters, beam splitters, prisms, diffusers, diffraction gratings, etc.), digital components (e.g., charge-coupled devices (CCDs), as well as an area for placement of samples, computing components (e.g., one or more processors, such as digital signal processors) to process and / or pre-process images, etc. Imagers may be, or may employ, any microscopes or microscopy systems (e.g., confocal microscopy), tools, and / or techniques that will provide the desired imaging data. The imager 162 may have the capability of receiving control signals from a computing device or system to, for example, initiate or cease image capture, and / or return images or other imaging data or status signals to the computing device or system (e.g., computing system 1110). Analytical devices 1164 may include any tools for obtaining additional data about samples and test results (such as spectrometers, chromatographs, etc.). Sensors 1166 may detect, for example, other aspects of samples and / or test results, such as temperature, humidity, location, etc. Experimentation system 1180 may include any components used to automate preparation of samples and running of experiments on samples. Robotics 1182, for example, may include any combination of actuators, stepper motors, servomotors, control system, robot arms, end effectors like grippers and manipulators, etc. Robotics 1182 may be, or may comprise, for example, an automated robotics system with process automation software. In various embodiments, robotics 1182 may employ vessels 1175, such as multi-well plates, that can be moved and manipulated, for example, for positioning of samples to be tested, imaged, or otherwise evaluated using condition detection system 1160 or components thereof. Experimentation system may also include various analytical tools 1184 used to perform tests or otherwise evaluate samples prior to, for example, samples being imaged. Analytical tools 1184 may include any devices needed to carry out experiments (e.g., by processing samples), and may include, for example, centrifuges, light emitters (e.g., radiation sources, lasers, etc.), heaters, fans or other passive, active coolers, etc. Sensors 1184 may be used by experimentation system 1180 to, for example, monitor, guide, and evaluate the progress of experiments (e.g., by detecting temperature, fill level of vessels, level of emitted radiation, or other states or conditions). In various implementations, components of system 1100 may be rearranged or integrated in other configurations. For example, computing system 1110 (or components thereof) may be integrated with one or more of the condition detection system 1160, -50- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 experimentation system 1180, and / or components thereof. The condition detection system 1160, experimentation system 1180, and / or components thereof may be directed to a vessel 1175 on which a sample or other subject can be situated (e.g., so as to test or evaluate biologics). In various embodiments, the vessel 175 may be movable (e.g., using any combination of motors, magnets, etc.) to allow for positioning and repositioning of samples (such as micro-adjustments for positioning of wells to be imaged). It is also noted that not all components of system 1100 are required to implement the disclosed approach, and in various embodiments, only a subset of the components of system 1100 may be employed. For example, in various embodiments, computing system 1110 may obtain and process data that was obtained via an another system that is or is not in direct communication with the computing system 1110. In certain embodiments, data may be obtained for samples that were not prepared in an automated fashion but rather manually by one or more users. If system 1100 includes an experimentation system 1180 (or components thereof), the computing system 1110 may include an experimentation unit 1120 configured to direct, for example, sample preparation, laboratory tests, and / or positioning of vessels 1175 (e.g., via control signals or instructions transmitted directly or indirectly to experimentation system 1180). If system 1100 includes condition detection system 1160 (or components thereof), computing system 1110 may include a condition detector 1122 configured to direct, for example, acquisition of imaging data via imager 1162 or other data from sensors or other components. A data acquisition unit 1124 may retrieve, acquire, or otherwise obtain various data to be used, for example, to run and / or evaluate experiments. The data acquisition unit 1124 may, for example, obtain data stored in database 1118, condition detection system 1160, information system 1170, and / or experimentation system 1180. An interaction unit 1126 may interact (e.g., via user interfaces 1116) with users (e.g., laboratory technicians, data scientists, etc.) to obtain information or commands needed for system 1100 or components thereof to function, or to start operations, repeat processes, and / or stop operations. In certain embodiments, the data acquisition unit 1124 may obtain data from users via interaction unit 1126. An experiment simulator 1128 may simulate plausible outcomes for each experiment in a set of candidate experiments. Experiment simulator 1128 may also determine how simulated experiments would change a model being trained. Machine learning platform -51- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 1130 may be configured to train and update machine learning models, as further discussed herein. Machine learning platform 1130 may, for example, employ certain machine learning techniques and algorithms to train and update predictive models. Machine learning platform 1130 may include a modeler 132 which may, for example, fit models to data from experiments. The modeler 1132 may, for example, obtain experiment data from or via data acquisition unit 1124, condition detection system 1160 or components thereof, and / or information system 1170. Combination predictor 1134 may apply models (e.g., models obtained from machine learning platform 130 or components thereof) to predict efficacious drug combinations as discussed herein. Analyzer 1140 may perform analyses such as computations on experimental data. A reporter 1142 may generate reports that include, for example, information on experiments run (e.g., via experimentation system 1180 or components thereof), imaging or other data obtained (e.g., via condition detection system 1160 or components thereof), data generated (e.g., via analyzer 1140), and / or models trained (e.g., via machine learning platform 1130) and combination prediction (e.g., via combination predictor 1134). For example, reporter 1142 may provide or identify trained and / or updated models, and / or combinations predicted. In certain embodiments, reporter 1142 may obtain information to be reported from, for example, database 1118 and / or information system 1170. Data that may be reported may be, for example, transmitted to another system or device (which may or may not be part of system 1100), saved in non-transitory computer-readable storage media (e.g., database 1118, information system 1170, and / or elsewhere), and / or presented or otherwise provided to users (e.g., via a display device that is part of user interfaces 1116). Referring to FIG. 12, an example process 1200 is illustrated, according to various example embodiments. Various elements of process 1200 may be implemented by or via system 1100 or components thereof. Process 1200 may begin at 1210 with selecting of drug combinations to screen. These may be initial drug combinations for an initial screening or additional combinations for subsequent screenings. The combinations may be combinations of drugs that are in a drug library. At 1220, process 1200 includes performing experiments on biologics using drug combinations. The experiments may be initial experiments using initial drug combinations on a set of biologics, or subsequent experiments. As used herein, biologics are any natural or artificial products isolated for experimentation, -52- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 such as cells, tissues, microorganisms, antibodies, products derived from blood and / or plasma, etc., that can provide an indication of the effectiveness of drugs. At 1230, process 1200 includes fitting a model to data from the experiments performed at 1220. This may include an initial generation of a model or updating a model based on new data. At 1240, a set of candidate experiments not already performed may be identified and plausible outcomes simulated for each candidate experiment. Simulating outcomes may include determining informativeness of simulated experiments. At 1250, the model may be updated based on the simulated outcomes and the informativeness. Process 1200 may return to step 1210 of selecting drug combinations to screen and repeat the experiments and update the model a number of times. When subsequent iterations do not significantly modify the model (e.g., amount of changes to the model start plateauing or result in overfitting), steps 1210 to 1250 can cease and process 1200 may proceed to step 1260. That is, the number of times the steps are repeated is determined based on how much of a contribution additional iterations are expected to make to the model. The final version of the model need not be selected, but rather a recently updated version of the model may be selected if, for example, the model has been overfit. At 1260, the model can be used to make predictions of efficacious drug combinations for one or more diseases or conditions. The potential drug combinations provided by the trained model need not be limited to the drug combinations used in process 1200, and similarly, biologics need not be limited to the biologics used in process 1200. For example, biologics can include any biologic capable of being screened in vitro such as cell lines, organoids, spheroids, tumoroids, patient-derived cells, etc., or any combination thereof. The discussion of “cell lines” herein is intended as an illustrative example and not intended to be limiting, as any biologic (or combination of biologics) may be used in place of (or in combination with) cell lines. Various operations described herein can be implemented on computer systems having various configurations. FIG.13 shows a simplified block diagram of a representative server system 1300 (e.g., computing system 1110 or other components in FIG.11) and client computer system 1314 (e.g., computing system 1110, condition detection system 1160, information system 1170, and / or experimentation system 1180) usable to implement various embodiments of the present disclosure. In various embodiments, server system 1300 or similar systems can implement services or servers described herein or portions thereof. Client computer system 1314 or similar systems can implement clients described herein. -53- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 Server system 1300 can have a modular design that incorporates a number of modules 1302 (e.g., blades in a blade server embodiment); while two modules 1302 are shown, any number can be provided. Each module 1302 can include processing unit(s) 1304 and local storage 1306. Processing unit(s) 1304 can include a single processor, which can have one or more cores, or multiple processors. In some embodiments, processing unit(s) 1304 can include a general-purpose primary processor as well as one or more special-purpose co- processors such as graphics processors, digital signal processors, or the like. In some embodiments, some or all processing units 1304 can be implemented using customized circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions that are stored on the circuit itself. In other embodiments, processing unit(s) 1304 can execute instructions stored in local storage 1306. Any type of processors in any combination can be included in processing unit(s) 1304. Local storage 1306 can include volatile storage media (e.g., conventional DRAM, SRAM, SDRAM, or the like) and / or non-volatile storage media (e.g., magnetic or optical disk, flash memory, or the like). Storage media incorporated in local storage 1306 can be fixed, removable or upgradeable as desired. Local storage 1306 can be physically or logically divided into various subunits such as a system memory, a read-only memory (ROM), and a permanent storage device. The system memory can be a read-and-write memory device or a volatile read-and-write memory, such as dynamic random-access memory. The system memory can store some or all of the instructions and data that processing unit(s) 1304 need at runtime. The ROM can store static data and instructions that are needed by processing unit(s) 1304. The permanent storage device can be a non-volatile read-and- write memory device that can store instructions and data even when module 1302 is powered down. The term “storage medium” as used herein includes any medium in which data can be stored indefinitely (subject to overwriting, electrical disturbance, power loss, or the like) and does not include carrier waves and transitory electronic signals propagating wirelessly or over wired connections. In some embodiments, local storage 1306 can store one or more software programs to be executed by processing unit(s) 1304, such as an operating system and / or programs implementing various server functions or any system or device described herein. -54- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 “Software” refers generally to sequences of instructions that, when executed by processing unit(s) 1304 cause server system 1300 (or portions thereof) to perform various operations, thus defining one or more specific machine embodiments that execute and perform the operations of the software programs. The instructions can be stored as firmware residing in read-only memory and / or program code stored in non-volatile storage media that can be read into volatile working memory for execution by processing unit(s) 1304. Software can be implemented as a single program or a collection of separate programs or program modules that interact as desired. From local storage 1306 (or non-local storage described below), processing unit(s) 1304 can retrieve program instructions to execute and data to process in order to execute various operations described above. In some server systems 1300, multiple modules 1302 can be interconnected via a bus or other interconnect 1308, forming a local area network that supports communication between modules 1302 and other components of server system 1300. Interconnect 1308 can be implemented using various technologies including server racks, hubs, routers, etc. A wide area network (WAN) interface 1310 can provide data communication capability between the local area network (interconnect 1308) and a larger network, such as the Internet. Conventional or other activities technologies can be used, including wired (e.g., Ethernet, IEEE 802.3 standards) and / or wireless technologies (e.g., Wi-Fi, IEEE 802.11 standards). In some embodiments, local storage 1306 is intended to provide working memory for processing unit(s) 1304, providing fast access to programs and / or data to be processed while reducing traffic on interconnect 1308. Storage for larger quantities of data can be provided on the local area network by one or more mass storage subsystems 1312 that can be connected to interconnect 1308. Mass storage subsystem 1312 can be based on magnetic, optical, semiconductor, or other data storage media. Direct attached storage, storage area networks, network-attached storage, and the like can be used. Any data stores or other collections of data described herein as being produced, consumed, or maintained by a service or server can be stored in mass storage subsystem 1312. In some embodiments, additional data storage resources may be accessible via WAN interface 1310 (potentially with increased latency). -55- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 Server system 1300 can operate in response to requests received via WAN interface 1310. For example, one of modules 1302 can implement a supervisory function and assign discrete tasks to other modules 1302 in response to received requests. Conventional work allocation techniques can be used. As requests are processed, results can be returned to the requester via WAN interface 1310. Such operation can generally be automated. Further, in some embodiments, WAN interface 1310 can connect multiple server systems 1300 to each other, providing scalable systems capable of managing high volumes of activity. Conventional or other techniques for managing server systems and server farms (collections of server systems that cooperate) can be used, including dynamic resource allocation and reallocation. Server system 1300 can interact with various user-owned or user-operated devices via a wide-area network such as the Internet. An example of a user-operated device is shown in FIG.13 as client computing system 1314. Client computing system 1314 can be implemented, for example, as a consumer device such as a smartphone, other mobile phone, tablet computer, wearable computing device (e.g., smart watch, eyeglasses), desktop computer, laptop computer, and so on. Client computing system 1314 can communicate via WAN interface 1310. Client computing system 1314 can include conventional computer components such as processing unit(s) 1316, storage device 1318, network interface 1320, user input device 1322, and user output device 1324. Client computing system 1314 can be a computing device implemented in a variety of form factors, such as a desktop computer, laptop computer, tablet computer, smartphone, other mobile computing device, wearable computing device, or the like. Processor 1316 and storage device 1318 can be similar to processing unit(s) 1304 and local storage 1306 described above. Suitable devices can be selected based on the demands to be placed on client computing system 1314; for example, client computing system 1314 can be implemented as a “thin” client with limited processing capability or as a high- powered computing device. Client computing system 1314 can be provisioned with program code executable by processing unit(s) 1316 to enable various interactions with server system 1300 of a message management service such as accessing messages, performing actions on messages, and other interactions described above. Some client computing systems 1314 can also interact with a messaging service independently of the message management service. -56- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 Network interface 1320 can provide a connection to a wide area network (e.g., the Internet) to which WAN interface 1310 of server system 1300 is also connected. In various embodiments, network interface 1320 can include a wired interface (e.g., Ethernet) and / or a wireless interface implementing various RF data communication standards such as Wi-Fi, Bluetooth, or cellular data network standards (e.g., 3G, 4G, 5G, LTE, etc.). User input device 1322 can include any device (or devices) via which a user can provide signals to client computing system 1314; client computing system 1314 can interpret the signals as indicative of particular user requests or information. In various embodiments, user input device 1322 can include any or all of a keyboard, touch pad, touch screen, mouse or other pointing device, scroll wheel, click wheel, dial, button, switch, keypad, microphone, and so on. User output device 1324 can include any device via which client computing system 1314 can provide information to a user. For example, user output device 1324 can include a display to display images generated by or delivered to client computing system 1314. The display can incorporate various image generation technologies, e.g., a liquid crystal display (LCD), light-emitting diode (LED) including organic light-emitting diodes (OLED), projection system, cathode ray tube (CRT), or the like, together with supporting electronics (e.g., digital-to-analog or analog-to-digital converters, signal processors, or the like). Some embodiments can include a device such as a touchscreen that function as both input and output device. In some embodiments, other user output devices 1324 can be provided in addition to or instead of a display. Examples include indicator lights, speakers, tactile “display” devices, printers, and so on. Some embodiments include electronic components, such as microprocessors, storage and memory that store computer program instructions in a computer readable storage medium. Many of the features described in this specification can be implemented as processes that are specified as a set of program instructions encoded on a computer readable storage medium. When these program instructions are executed by one or more processing units, they cause the processing unit(s) to perform various operation indicated in the program instructions. Examples of program instructions or computer code include machine code, such as is produced by a compiler, and files including higher-level code that are executed by a computer, an electronic component, or a microprocessor using an interpreter. Through suitable programming, processing unit(s) 1304 and 1316 can provide various functionality -57- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 for server system 1300 and client computing system 1314, including any of the functionality described herein as being performed by a server or client, or other functionality associated with message management services. It will be appreciated that server system 1300 and client computing system 1314 are illustrative and that variations and modifications are possible. Computer systems used in connection with embodiments of the present disclosure can have other capabilities not specifically described here. Further, while server system 1300 and client computing system 1314 are described with reference to particular blocks, it is to be understood that these blocks are defined for convenience of description and are not intended to imply a particular physical arrangement of component parts. For instance, different blocks can be but need not be located in the same facility, in the same server rack, or on the same motherboard. Further, the blocks need not correspond to physically distinct components. Blocks can be configured to perform various operations, e.g., by programming a processor or providing appropriate control circuitry, and various blocks might or might not be reconfigurable depending on how the initial configuration is obtained. Embodiments of the present disclosure can be realized in a variety of apparatus including electronic devices implemented using any combination of circuitry and software. While the disclosure has been described with respect to specific embodiments, one skilled in the art will recognize that numerous modifications are possible. For instance, although specific examples of rules (including triggering conditions and / or resulting actions) and processes for generating suggested rules are described, other rules and processes can be implemented. Embodiments of the disclosure can be realized using a variety of computer systems and communication technologies including but not limited to specific examples described herein. Embodiments of the present disclosure can be realized using any combination of dedicated components and / or programmable processors and / or other programmable devices. The various processes described herein can be implemented on the same processor or different processors in any combination. Where components are described as being configured to perform certain operations, such configuration can be accomplished, e.g., by designing electronic circuits to perform the operation, by programming programmable electronic circuits (such as microprocessors) to perform the operation, or any combination thereof. Further, while the embodiments described above may make reference to specific -58- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 hardware and software components, those skilled in the art will appreciate that different combinations of hardware and / or software components may also be used and that particular operations described as being implemented in hardware might also be implemented in software or vice versa. Computer programs incorporating various features of the present disclosure may be encoded and stored on various computer readable storage media; suitable media include magnetic disk or tape, optical storage media such as compact disk (CD) or DVD (digital versatile disk), flash memory, and other non-transitory media. Computer readable media encoded with the program code may be packaged with a compatible electronic device, or the program code may be provided separately from electronic devices (e.g., via Internet download or as a separately packaged computer-readable storage medium). Thus, although the disclosure has been described with respect to specific embodiments, it will be appreciated that the disclosure is intended to cover all modifications and equivalents within the scope of the following claims. As utilized herein, the terms “approximately,” “about,” “substantially”, and similar terms are intended to have a broad meaning in harmony with the common and accepted usage by those of ordinary skill in the art to which the subject matter of this disclosure pertains. It should be understood by those of skill in the art who review this disclosure that these terms are intended to allow a description of certain features described and claimed without restricting the scope of these features to the precise numerical ranges provided. Accordingly, these terms should be interpreted as indicating that insubstantial or inconsequential modifications or alterations of the subject matter described and claimed are considered to be within the scope of the disclosure as recited in the appended claims. In some examples, these terms allow for a plus-or-minus deviation of 10 percent. It should be noted that the terms “exemplary,” “example,” “potential,” and variations thereof, as used herein to describe various embodiments, are intended to indicate that such embodiments are possible examples, representations, or illustrations of possible embodiments (and such terms are not intended to connote that such embodiments are necessarily extraordinary or superlative examples). The term “coupled” and variations thereof, as used herein, means the joining of two members directly or indirectly to one another. Such joining may be stationary (e.g., -59- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 permanent or fixed) or moveable (e.g., removable or releasable). Such joining may be achieved with the two members coupled directly to each other, with the two members coupled to each other using a separate intervening member and any additional intermediate members coupled with one another, or with the two members coupled to each other using an intervening member that is integrally formed as a single unitary body with one of the two members. If “coupled” or variations thereof are modified by an additional term (e.g., directly coupled), the generic definition of “coupled” provided above is modified by the plain language meaning of the additional term (e.g., “directly coupled” means the joining of two members without any separate intervening member), resulting in a narrower definition than the generic definition of “coupled” provided above. Such coupling may be mechanical, electrical, or fluidic. The term “or,” as used herein, is used in its inclusive sense (and not in its exclusive sense) so that when used to connect a list of elements, the term “or” means one, some, or all of the elements in the list. Conjunctive language such as the phrase “at least one of X, Y, and Z,” unless specifically stated otherwise, is understood to convey that an element may be either X, Y, Z; X and Y; X and Z; Y and Z; or X, Y, and Z (i.e., any combination of X, Y, and Z). Thus, such conjunctive language is not generally intended to imply that certain embodiments require at least one of X, at least one of Y, and at least one of Z to each be present, unless otherwise indicated. References herein to the positions of elements (e.g., “top,” “bottom,” “above,” “below”) are merely used to describe the orientation of various elements in the Figures. It should be noted that the orientation of various elements may differ according to other exemplary embodiments, and that such variations are intended to be encompassed by the present disclosure. The embodiments described herein have been described with reference to drawings. The drawings illustrate certain details of specific embodiments that implement the systems, methods and programs described herein. However, describing the embodiments with drawings should not be construed as imposing on the disclosure any limitations that may be present in the drawings. It is important to note that the construction and arrangement of the devices, assemblies, and steps as shown in the various exemplary embodiments is illustrative only. -60- 4930-0063-8464.1 Atty. Dkt. No.: 115872-3136 Additionally, any element disclosed in one embodiment may be incorporated or utilized with any other embodiment disclosed herein. Although only one example of an element from one embodiment that can be incorporated or utilized in another embodiment has been described above, it should be appreciated that other elements of the various embodiments may be incorporated or utilized with any of the other embodiments disclosed herein. The foregoing description of embodiments has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the disclosure to the precise form disclosed, and modifications and variations are possible in light of the above teachings or may be acquired from this disclosure. The embodiments were chosen and described in order to explain the principals of the disclosure and its practical application to enable one skilled in the art to utilize the various embodiments and with various modifications as are suited to the particular use contemplated. Other substitutions, modifications, changes and omissions may be made in the design, operating conditions and arrangement of the embodiments without departing from the scope of the present disclosure as expressed in the appended claims. -61- 4930-0063-8464.1

Claims

Atty. Dkt. No.: 115872-3136 WHAT IS CLAIMED IS:

1. A method comprising: (a) selecting initial drug combinations to screen and performing a set of initial experiments using the initial drug combinations on a set of biologics, wherein the initial drug combinations are combinations of drugs in a drug library; (b) fitting a model to data from the set of initial experiments; (c) identifying a set of candidate experiments not already performed and simulating plausible outcomes for each experiment in the set of candidate experiments; (d) updating the model based at least on the simulated outcomes and informativeness of each experiment in the set of candidate experiments; and (e) repeating steps (a) to (d) a number of times to obtain a trained model configured to provide predictions of efficacious drug combinations.

2. The method of claim 1, wherein the set of initial experiments are run using a set of candidates designed such that each biologic in the set of biologics and each drug in the drug library is covered by at least one experiment.

3. The method of claim 1, wherein fitting the model in step (b) comprises employing probabilistic modeling.

4. The method of claim 1, wherein the set of candidate experiments are selected from potential experiments not in the set of initial experiments.

5. The method of claim 1, wherein the set of candidate experiments comprises all possible experiments not in the set of initial experiments.

6. The method of claim 1, wherein the number of times that steps (a) to (d) are repeated is based at least on how much the model changes following updating step (d).

7. The method of claim 1, wherein the set of biologics comprises a plurality of cell lines. -62--0063-8464.1Atty. Dkt. No.: 115872-3136 8. The method of claim 1, wherein the trained model provides predictions of synergistic drug combinations.

9. The method of claim 1, further comprising validating the model in vitro.

10. The method of claim 1, further comprising validating the model in vivo.

11. A method comprising: (a) performing a first batch of experiments to obtain first data indicative of effects of a first set of drug combinations on one or more biologics, each drug combination in the first set of drug combinations comprising two or more drugs from a drug library; (b) obtaining a model based at least on the first data; (c) simulating, based at least on the model, a second batch of experiments to obtain second data, the second batch of experiments corresponding to a second set of drug combinations, each drug combination in the second set of drug combinations comprising two or more drugs from the drug library, wherein the second set of drug combinations comprises drug combinations not in the first set of drug combinations; (d) updating the model based at least on the second data and scoring each experiment in the second batch of experiments based at least on changes to the model resulting from the updating; (e) performing a third batch of experiments to obtain third data, the third batch of experiments corresponding to a third set of drug combinations; and (f) updating the model based at least on the third data to obtain a trained model configured to predict efficacious combinations of drugs from the drug library or from a second drug library with different drugs.

12. The method of claim 11, further comprising simulating and performing additional experiments to update the trained model. -63--0063-8464.1Atty. Dkt. No.: 115872-3136 13. The method of claim 11, wherein the first batch of experiments are designed such that each biologic in the one or more biologics and each drug in the drug library is covered by at least one experiment.

14. The method of claim 11, wherein obtaining the model comprises fitting a probabilistic model to the first data.

15. The method of claim 11, wherein the second batch of experiments are selected from potential experiments not in the first batch of experiments.

16. The method of claim 11, wherein the second batch of candidate experiments comprises all possible experiments not in the first batch of experiments.

17. The method of claim 11, wherein the trained model provides predictions of synergistic drug combinations.

18. A method comprising using a model to predict one or more efficacious combinations of drugs from a drug library, wherein the model was trained by: (a) fitting a model to experimental results from a set of initial experiments to quantify uncertainty; (b) using the uncertainty to simulate plausible outcomes of candidate experiments using candidate drug combinations; (c) updating the model and quantifying informativeness of each candidate experiment based at least on how much the model changes due to the corresponding candidate experiment; (d) designing a subsequent batch of experiments based at least on informativeness of candidate experiments; and (e) repeating steps (a) to (d) a number of times to obtain a trained model.

19. The method of claim 18, wherein the predicted combinations are for treating one or more conditions or diseases in subjects. -64--0063-8464.1Atty. Dkt. No.: 115872-3136 20. A method of training a model to predict one or more efficacious combinations of drugs, comprising: (a) fitting a model to experimental results from one or more sets of experiments to quantify uncertainty; (b) simulating, based at least on the uncertainty, plausible outcomes of candidate experiments using candidate drug combinations; and (c) updating the model based at least on informativeness of candidate experiments. -65--0063-8464.1

Citation Information

Patent Citations

  • Methods for the detection and treatment of leukemias that are responsive to dot1l inhibition

    US20200239963A1

  • Network-based deep learning technology for target identification and drug repurposing

    US20210142173A1

  • Multi-organ "body on a chip" apparatus utilizing a common media

    US20210324311A1

  • Wind power prediction method and system based on deep deterministic policy gradient algorithm

    US20230288607A1

  • Method for filtering and enhancing 2d, 3D and 3D+t ultrasound images

    WO2015169987A1