Methods and model systems for evaluating therapeutic properties of candidate agents, and related computer readable media and systems
Patent Information
- Application Number
- JP2024503912
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-07-23
- Filing Date
- 2022-07-22
- Publication Date
- 2025-07-29
AI Technical Summary
The high cost and inefficiency of drug discovery processes, particularly due to the limitations of high-throughput screening (HTS) methods that rely on in vitro biochemical or cell-based assays, lead to expensive and ineffective drug development.
Development of in vivo and ex vivo model systems for high-throughput screening using equilibrium cell number cultures, where a heterogeneous pool of cells is grown in three dimensions, treated with small molecule compounds, and analyzed through single-cell RNA sequencing to evaluate therapeutic properties.
This approach allows for accurate assessment of drug efficacy and resistance mechanisms, enabling the discovery of effective combination therapies and personalized treatment strategies by analyzing phenotypic and genetic responses at the single-cell level.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 225,209, filed July 23, 2021, which is incorporated herein by reference in its entirety.
[0002] Statement of government support This invention was made with Government support under Grant No. R00 CA194077 awarded by the National Institutes of Health. The Government has certain rights in this invention. [Background technology]
[0003] The recent phenomenon of rising drug discovery costs and diminishing returns in the discovery of therapeutic molecules can be attributed to the fundamentals of pharmaceutical science. Specifically, high-throughput screening (HTS) remains the gold standard for new molecular discovery. However, the major limitation of HTS is its artificial nature, which is only possible with in vitro biochemical or cell-based assays. In vitro biochemical or cell-based assays, followed by preclinical in vivo studies, cannot provide sufficient pharmacological and toxicity data or reliable predictive capabilities to predict therapeutic candidate performance in vivo, making the drug development process expensive and inefficient. The present disclosure provides in vivo and ex vivo model systems and methods for creating such systems to perform scalable HTS screening. Summary of the Invention
[0004] Provided are balanced cell number cultures and methods for making same, for evaluating one or more therapeutic properties of candidate drugs. The method includes growing a heterogeneous pool of cells of different cell types in three dimensions, treating the three-dimensional pool with a small molecule compound, and dissociating the cells of the treated three-dimensional pool into a single-cell suspension with equal representation of cell types suitable for single-cell RNA sequencing. The method further includes performing single-cell ribonucleic acid (RNA) sequencing on the dissociated single cells and dissociated single cells from a control three-dimensional pool that is not treated with the small molecule compound, deconvoluting data from the single-cell RNA sequencing into single-cell transcriptomes sorted by treatment and cell type, and evaluating one or more therapeutic properties of the small molecule compound based on the sorted single-cell transcriptomes. Also provided are computer readable media and systems that find use, for example, in practicing the methods of the present disclosure. [Brief description of the drawings]
[0005] [Figure 1]Figures 1A-1G demonstrate a comparison of cell composition between non-GENEVA and GENEVA cell pools. Figure 1A demonstrates the distribution of cell representation from a non-GENEVA cell pool (pool 1) based on single cell RNA sequencing harvest. Bars indicate the number of cells in the scRNAseq dataset for each cell type. Figure 1B shows the dataset from Figure 1A plotted in two-dimensional transcriptome space using the UMAP clustering visualization algorithm. Figure 1C shows the distribution of cell representation upon single cell RNA sequencing harvest from a GENEVA cell pool (pool 2), allowing for accurate capture of each cell line in the dataset. Figure 1D shows single cell RNA sequencing data plotted as a UMAP plot for GENEVA pool 2. Figure 1E shows the distribution of cell representation upon single cell RNA sequencing harvest from a GENEVA cell pool (pool 3), allowing for accurate capture of each cell line in the dataset. Figure 1F shows single cell RNA sequencing data plotted as a UMAP plot for GENEVA pool 3. Figure 1G shows the extrapolation of pools 1-3. The total number of cells required for single-cell RNA sequencing is significantly higher in pool 1 compared to pool 2 and pool 3. [Diagram 2]Figures 2A-F demonstrate the utilization of GENEVA in multiple in vivo and ex vivo model systems. Figure 2A demonstrates a GENEVA pool grown as ex vivo organoids using as input four different human PDX tumors treated with several doses of ARS1620 (0.4 uM, 1.6 uM, 25.0 uM) or DMSO (vehicle). Figure 2A is plotted in two-dimensional transcriptome space using the UMAP clustering visualization algorithm. Figure 2B shows the dataset from Figure 2A as a table categorized by drug treatment condition and genotype of origin (PDX). Cells in the table are cell counts obtained by single-cell RNA sequencing for each category. Figure 2C demonstrates a GENEVA pool grown as in vivo flank xenografts using as input four different human PDX tumors treated with either ARS1620 (100 mg / kg) or DMSO (vehicle). Figure 2C is plotted in two-dimensional transcriptome space using UMAP clustering visualization algorithm. Figure 2D shows the dataset from Figure 2C as a table categorized by drug treatment condition and genotype of origin (PDX). The cells in the table are the cell counts obtained by single-cell RNA sequencing for each category. Figure 2E demonstrates GENEVA pools grown as flank xenografts in vivo using as input eight sub-different human cancer cell lines treated with either ARS1620 (100 mg / kg) or DMSO (vehicle). Figure 2E is plotted in two-dimensional transcriptome space using UMAP clustering visualization algorithm. Figure 2F shows the dataset from Figure 2E as a table categorized by drug treatment condition and genotype of origin (PDX). The cells in the table are the cell counts obtained by single-cell RNA sequencing for each category. [Diagram 3]Figures 3A-E demonstrate the use of GENEVA for discovery of relative phenotypes to drug compounds, genetic drivers, and IC50 curve reconstructions. Figure 3A demonstrates the relative sensitivity of individual cell types calculated from pre / post drug treatment versus cell counts from GENEVA pools treated with vemurafinib or ARS1620. The most sensitive cell line in the vemurafinib-treated pool is BRAF.V600E mutant-bearing. The most sensitive cell line in the ARS1620-treated pool is KRAS.G12C mutant-bearing. Figure 3B demonstrates discovery of causative driver mutations responsible for changes in relative drug sensitivity in GENEVA pools using a Lasso regression model. BRAF.V600E is predicted by the Lasso algorithm as the causative mutation causing drug sensitivity to vemurafinib. KRAS.G12C is predicted by the Lasso algorithm as the causative mutation causing drug sensitivity to ARS1620. Figure 3C demonstrates the reconstruction of IC50 curves from GENEVA cell pool data after treatment with or without ARS1620 with cell counts and interpolated IC50 logistic regression curves fitted to a scaled measure of relative viability. IC50 curves were constructed from individual cell lines, with non-KRAS.G12C cell lines showing significantly higher viability to ARS1620 than KRAS.G12C cell lines. Figure 3D demonstrates the relative drug sensitivity measurements from different cell lines in the GENEVA pool, thereby recapitulating the discovery of KRAS.G12C as a sensitizing mutation target for ARS1620. Figure 3E demonstrates the calculation of percent cell cycle inhibition from GENEVA performed on PDXs grown as pooled organoids. This reconstruction method recapitulates KRAS.G12C-specific drug sensitivity to ARS-1620 treatment by an alternative method of cell counting using cycle state inference as a phenotypic measurement. [Figure 4]Figure 4A-D demonstrate the utilization of GENEVA for combination therapy and prediction of drug resistance mechanisms. Figure 4A demonstrates that GENEVA finds upregulation of several drug resistance targets indicative of cell survival mechanisms in GENEVA pools treated with ARS1620, specifically KRAS.G12C. Figure 4B demonstrates validation of predicted drug targets by administering the drug targets in combination with i) three ARS1620 inhibitors, and ii) compounds targeting specific drug resistance pathways. Bliss drug synergy is plotted, with several compounds showing significant drug synergy with multiple KRAS.G12C inhibitors. Figure 4C demonstrates relative tumor volumes from an in vivo mouse study of ARS1620 and INK128 in a multi-arm combination therapy using the H1373 KRAS.G12C mutant (n=4-5 mice per condition). FIG. 4D demonstrates that INK128 and ARS1620 synergistically reduce tumor growth in vivo compared to an independent INK128 and ARS1620 or null model of no drug synergy. [Diagram 5] Figure 5A-C demonstrate the utilization of GENEVA for prediction of in vivo specific mechanisms of drug resistance via the endothelial-mesenchymal transition (EMT) pathway. Figure 5A demonstrates that in paired in vivo and in vitro GENEVA experiments of ARS1620 treatment of KRAS.G12C cell lines in cell pools, the EMT gene set was upregulated following drug treatment in vivo but not in vitro. Figure 5B demonstrates relative tumor volumes from an in vivo mouse study of ARS1620 and the EMT inhibitor galunisertib in a multi-arm combination therapy using the H1373 KRAS.G12C mutant line (n=4-5 mice per condition). Figure 5C demonstrates that galunisertib and ARS1620 synergistically reduce tumor growth in vivo compared to a null model of independent or no drug synergy of galunisertib and ARS1620. [Figure 6]Figures 6A-E demonstrate the utilization of GENEVA for the discovery of molecular mechanisms of action of compounds on mitochondrial genes. Figure 6A plots aggregated gene expression across KRAS.G12C lines in the GENEVA pool of mitochondrially encoded and genomically encoded mitochondrially targeted transcripts compared to gene expression of non-mitochondrial gene transcripts following ARS1620 treatment. Mitochondrially encoded genes and genomically encoded mitochondrial resident genes are significantly downregulated in cells that survive following ARS1620 treatment. Figure 6B plots gene expression of mito-encoded transcripts following ARS1620 treatment for each individual KRAS.G12C cell line within the GENEVA cell pool. Figure 6C demonstrates the generation and profiling of long-term ARS1620 tolerant cell lines (30 day treatment, 10 uM) from H2030 (KRAS.G12C). Assay of mitochondrial content using a fluorescent mitochondrial stain (Mitotracker Deep Red FM) between H2030 drug-persistent cell lines and the original parental cell lines shows a decrease in mitochondrial content after chronic drug treatment with ARS1620. Figure 6D demonstrates that the KRAS.G12C inhibitor AMG510 increases mitochondrial respiration and electron transport chain activity as a novel lethality mechanism of KRAS.G12C inhibition using a Seahorse assay measuring oxygen consumption of H2030 (KRAS.G12c) cells after AMG510 treatment (2 hours). Figure 6E demonstrates that the subpopulation structure of KRAS.G12C cell lines shows selection of cell types with a lower number of mitochondrial reads after post-treatment with ARS1620. [Figure 7]Figures 7A-G demonstrate the utilization of GENEVA for the discovery of molecular mechanisms of compound action on ferroptosis genes. Figure 7A plots volcano plots showing aggregated Z-score differences across multiple G12C lines from a GENEVA pool drugged with ARS1620, demonstrating shared upregulation of anti-ferroptosis genes. Figure 7B plots gene expression of anti-ferroptosis genes for each cell line in response to increasing ARS1620 dosages. Figure 7C utilizes experimental investigation of ferroptosis using a fluorescent lipid peroxidation sensor to demonstrate the dose response of lipid peroxidation cells to ARS1620 dosage (48 hours). Figures 7D,E,F demonstrate survival and lipid peroxidation kinetics across the known strong derivative altretamine in Figure 7D, and elastin in Figure 7E, compared to ARS1620 in Figure 7F. Lipid peroxidation and survival kinetics cross over near the IC50 for all compounds indicating that ARS1620 performs as a ferroptosis inducer. Figure 7G demonstrates that multiple KRAS inhibitors specifically induce lipid peroxidation in the KRAS.G12C cell line H2030, but to a lesser extent in the KRAS.G12V cell line H441. [Figure 8]Figures 8A-D demonstrate GENEVA testing of combination therapy incorporating multiple co-administered compounds in cell pools in vivo. Figure 8A demonstrates a GENEVA combination therapy study using CLX pools in KRAS.G12 mutant strains categorized by cell line of origin and drug treatment condition plotted in two-dimensional transcriptome space using the UMAP clustering visualization algorithm. Treatment conditions include antimycin, ARS1620, galunisertib, INK128, DMSO, ARS1620+antimycin, ARS1620+galunisertib, ARS1620+INK128. Figure 8B utilizes the GENEVA data from Figure 8A for synergy calculations performed alone for cell cycle states in each drug condition and in combination estimated for the G12C strain across drug conditions. Figure 8C-D demonstrates the identification of genes driving the synergistic drug effect of galunisertib (Figure 8C) and INK128 (Figure 8D) in combination with ARS1620 using a multifactorial linear model to estimate gene-level synergy, revealing mitochondrial transcripts driving the synergistic drug effect. [Figure 9] Figures 9A-B demonstrate the identification of novel genotypes in patients responding to ARS1620 by the gene demultiplexing improvement algorithm and the GENEVA method. Figure 9A demonstrates the improvement in cell-assigned confidence by the gene demultiplexing denoising algorithm comparison to the standard method of percent representation from a Totalseq labeled single-cell RNA sequencing dataset. High confidence metrics are noted in the dotted lines, or the results of the noise correction algorithm. Figure 9B plots that GENEVA pools drugged with or without ARS1620 demonstrated the sensitivity of EML4-ALK as the most drug sensitive tumor type, with each bar being the relative survival rate of that genotype under ARS1620 treatment or vehicle. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0006] Before the methods, computer-readable media, and systems of the present disclosure are described in more detail, it is to be understood that the methods, computer-readable media, and systems are not limited to particular embodiments described, and as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the methods, computer-readable media, and systems will be limited only by the appended claims.
[0007] Where a range of values is provided, unless the context clearly dictates otherwise, it is understood that each intermediate value between the upper and lower limit of that range, to the tenth of the unit of the lower limit, and any other stated or intermediate value within this stated range, is encompassed by the methods, computer-readable media, and systems. The upper and lower limits of these smaller ranges may be individually included in the smaller ranges, and are also encompassed by the methods, computer-readable media, and systems, subject to any specifically excluded limits in the stated ranges. Where a stated range includes one or both of the limits, ranges excluding either or both of those included limits are also encompassed by the methods, computer-readable media, and systems.
[0008] In this specification, certain ranges are presented with numerical values preceded by the term "about". In this specification, the term "about" is used to provide literal support for the exact number it precedes, as well as a number that is close to or approximately the number it precedes. When determining whether a number is close to or approximately a specifically recited number, the close or approximate unrecited number may be a number that provides a substantial equivalent to the specifically recited number in the context in which it is presented.
[0009] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the methods, computer readable media, and systems belong. Any methods, computer readable media, and systems similar or equivalent to those described herein may also be used in the practice or testing of the methods, computer readable media, and systems, representative exemplary methods, computer readable media, and systems.
[0010] All publications and patents cited herein are incorporated by reference as if each individual publication or patent was specifically and individually indicated to be incorporated by reference, and by incorporation by reference herein discloses and describes the materials and / or methods in connection with which such publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the methods, computer readable media, and systems are not entitled to antedate such publications, since the dates of publication provided may be different from the actual publication dates which may need to be separately confirmed.
[0011] It should be noted that, as used in this specification and the appended claims, the articles "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It should be further noted that the claims may be designed to exclude any optional element. As such, this statement is intended to serve as a predicate to the use of such exclusive terminology, such as "solely," "only," or the use of a "negative" limitation in connection with the recitation of claim elements.
[0012] It should be understood that certain features of the methods, computer-readable media, and systems that are described in the context of separate embodiments may be provided in combination in a single embodiment. Conversely, various features of the methods, computer-readable media, and systems that are described in the context of a single embodiment for brevity may also be provided separately or in any suitable subcombination. All combinations of the embodiments are specifically encompassed by the present disclosure, and to the extent that such combinations encompass operable processes and / or compositions, each and every combination is disclosed herein as if it were individually and expressly disclosed. In addition, all subcombinations listed in the embodiments that describe such variables are also specifically encompassed by the present methods, computer-readable media, and systems, and each and every such subcombination is disclosed herein as if it were individually and expressly disclosed.
[0013] As will be apparent to those skilled in the art upon reading this disclosure, each of the separate embodiments described and illustrated herein has separate components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the invention. Any recited method may be carried out in the order of events recited or in any other order which is logically possible.
[0014] method In one aspect, the disclosure provides a balanced cell number culture and a method for making the balanced cell number culture. In one aspect, the disclosure provides a method for evaluating one or more therapeutic properties of a candidate agent, e.g., a small molecule compound. The method includes growing a heterogeneous pool of cells of different cell types in three dimensions, treating the three-dimensional pool with a small molecule compound, and dissociating the cells of the treated three-dimensional pool into single cells in a manner that allows equal representation of cells from the different cell types. The method further includes performing single-cell ribonucleic acid (RNA) sequencing on the dissociated single cells and dissociated single cells from a control three-dimensional pool that is not treated with the small molecule compound, deconvoluting data from the single-cell RNA sequencing into single-cell transcriptomes sorted by treatment and cell type, and evaluating one or more therapeutic properties of the small molecule compound based on the sorted single-cell transcriptomes.
[0015] The present method addresses this by mixing cells together, drugging them together, and then using single-cell RNA sequencing to read out the cell lines after drug treatment. In this way, by reducing the unit of observation to a single cell and mixing / multiplexing cell lines together, the present method allows the assay of a large number of phenotypically / genotypically distinct cell lines against a large number of small molecules. The resulting single-cell RNA sequencing data is analyzed using different models to discover biological targets, effective synergistic combination therapy targets, disease subtype stratification, etc.
[0016] An embodiment of the disclosed method is provided in FIG. 2 as a single-cell RNA sequencing result. In this example, a large panel of cell types from different patients, organ systems, and / or disease models / subtypes are mixed together to create a pool. The pool is then grown in three dimensions in vivo (e.g., producing xenografts in animal models, e.g., mice) or ex vivo (e.g., producing organoids). The three-dimensional pool is then treated with a research small molecule compound of interest under conditions suitable for the compound to act on the members (cells) of the three-dimensional pool. The drug delivery method varies depending on the type of three-dimensional pool, e.g., systemic injection if the three-dimensional pool is an in vivo xenograft. The treated three-dimensional pool is then harvested, dissociated into single cells, and then subjected to single-cell RNA sequencing. Phenotypic changes are noted by counting the number of individual viable cells for each cell type and comparing it to the number of viable cells for each cell type from the same three-dimensional pool that has not been treated with the research small molecule compound of interest. The single cell sequencing data is then subjected to modeling and / or transcriptomic analysis to evaluate one or more therapeutic properties of the small molecule compounds, non-limiting examples of which include mechanism of action (MOA), combination therapy (drugs that would be effective as clinical combination therapy), and subtype stratification (efficacy in different patient groups or subtypes).
[0017] As summarized above, the methods of the present disclosure include growing a pool of cells of different cell types in three dimensions. In certain embodiments, the pool of cells of different cell types includes no more than 1000, no more than 500, no more than 250, or no more than 100, but more than 2, more than 5, more than 10 (e.g., 10-50), more than 20, more than 30, more than 40, or more than 50 different cell types.
[0018] The cells of different cell types may be selected from any cell type of interest, which may vary depending on the particular small molecule compound of interest, one or more therapeutic properties of the small molecule being evaluated, etc. According to some embodiments, the pool of cells of different cell types includes primary cells obtained from a patient, cells from an organ system, cells from a disease model, or any combination thereof.
[0019] Cells obtained from a patient include, but are not limited to, cells from biopsy tissue obtained from a patient. Biopsy tissue may be obtained from healthy tissue or diseased tissue, including, for example, cancerous tissue. Depending on the type of cancer and / or type of biopsy performed, the cells may be from a solid tissue biopsy or a liquid biopsy. In some examples, cells may be prepared from a surgical biopsy. Any convenient and suitable technique for surgical biopsy may be utilized for collection of cells used in the methods described herein, including, but not limited to, excision biopsy, incision biopsy, wire localization biopsy, and the like. In some examples, a surgical biopsy may be obtained as part of a surgical procedure that has a primary purpose other than obtaining a sample, including, but not limited to, tumor resection, mastectomy, lymph node surgery, axillary lymph node dissection, sentinel lymph node surgery, and the like.
[0020] Various other biopsy techniques can be used to obtain biopsy tissue, and then obtain cells for use in the methods of the present disclosure.As a non-limiting example, sample can be obtained by needle biopsy.Any convenient and suitable technique for needle biopsy can be used to collect sample, including but not limited to fine needle aspiration (FNA), core needle biopsy, stereotactic core biopsy, vacuum-assisted biopsy, etc.
[0021] Cells from an organ system may include, but are not limited to, cells from an organ system selected from skin, brain, heart, kidney, liver, stomach, colon, lung, and the like. According to some embodiments, cells from an organ system may include cells from an organ system selected from adrenal gland, anus, appendix, bladder (urinary), bone, bone marrow, brain, bronchi, diaphragm, ear, esophagus, eye, fallopian tube, gallbladder, reproductive system, heart, hypothalamus, joints, kidney, large intestine, larynx, liver, lung, lymph node, mammary gland, mesentery, mouth, nasal cavity, nose, ovary, pancreas, pineal gland, parathyroid gland, pharynx, pituitary gland, prostate, rectum, salivary gland, skeletal muscle, smooth muscle, skin, small intestine, spinal cord, spleen, stomach, teeth, thymus, thyroid, trachea, tongue, ureter, urethra, ligaments, tendons, hair, vestibular system, placenta, testes, vas deferens, seminal vesicles, bulbourethral gland, parathyroid, thoracic duct, arteries, veins, capillaries, lymphatic vessels, tonsils, neurons, subcutaneous tissue, olfactory epithelium (nose), cerebellum, and any combination thereof.
[0022] Cells from disease models can include, but are not limited to, cells modeling a disease selected from cancer (e.g., cells from one or more different cancer cell lines), cardiovascular disease, cerebrovascular disease (e.g., stroke, transient ischemic attack, subarachnoid hemorrhage, vascular dementia, etc.), respiratory disease, infectious disease, neurodegenerative disease, dementia, Alzheimer's disease, diabetes, kidney disease, liver disease (e.g., cirrhosis, non-alcoholic fatty liver disease (NAFLD), hepatitis A, hepatitis B, hepatitis C, etc.), and any combination thereof.
[0023] According to some embodiments, the cells of different cell types include cells from one or more cancer cell lines. "Cancer cells" refers to cells that exhibit a neoplastic cell phenotype, which may be characterized, for example, by one or more of aberrant cell growth, aberrant cell proliferation, loss of density-dependent growth inhibition, anchorage-independent growth potential, ability to promote tumor growth and / or development in immunocompromised non-human animal models, and / or any suitable indicator of cell transformation. "Cancer cells" may be used interchangeably herein with "tumor cells," "malignant cells," or "cancerous cells," and encompasses cancer cells of solid tumors, semi-solid tumors, hematological malignancies (e.g., leukemia cells, lymphoma cells, myeloma cells, etc.), primary tumors, metastatic tumors, etc.
[0024] Where the cells of the different cell types comprise cells from one or more cancer cell lines, the one or more cancer cell lines may be derived from cancers independently selected from squamous cell carcinoma, small cell lung cancer, non-small cell lung cancer, lung adenocarcinoma, lung squamous cell carcinoma, peritoneal cancer, hepatocellular carcinoma, gastrointestinal cancer, pancreatic cancer, glioblastoma, cervical cancer, ovarian cancer, liver cancer, bile duct cancer, bladder cancer, liver cancer, breast cancer, colon cancer, colorectal cancer, endometrial or uterine cancer, salivary gland cancer, kidney cancer, prostate cancer, vulvar cancer, thyroid cancer, liver cancer, various types of head and neck cancer, and the like. In certain embodiments, the one or more cancer cell lines may be derived from a cancer independently selected from solid tumors, recurrent glioblastoma multiforme (GBM), non-small cell lung cancer, metastatic melanoma, melanoma, peritoneal cancer, epithelial ovarian cancer, glioblastoma multiforme (GBM), metastatic colorectal cancer, colorectal cancer, pancreatic ductal adenocarcinoma, squamous cell carcinoma, esophageal cancer, gastric cancer, neuroblastoma, fallopian tube cancer, bladder cancer, metastatic breast cancer, pancreatic cancer, soft tissue sarcoma, recurrent head and neck cancer squamous cell carcinoma, head and neck cancer, anaplastic astrocytoma, malignant pleural mesothelioma, breast cancer, squamous non-small cell lung cancer, rhabdomyosarcoma, metastatic renal cell carcinoma, basal cell carcinoma (basal cell epithelioma), and gliosarcoma. In certain embodiments, the one or more cancer cell lines may be derived from a cancer independently selected from melanoma, Hodgkin's lymphoma, renal cell carcinoma (RCC), bladder cancer, non-small cell lung cancer (NSCLC), and head and neck squamous cell carcinoma (HNSCC). According to some embodiments, the cells of different cell types include cells from one or more cancer cell lines listed in the Broad Institute Cancer Cell Line Encyclopedia (CCLE), available at portals.broadinstitute.org / ccle.
[0025] According to some embodiments, the cells of the different cell types include cells from one or more different types of stem cells. Non-limiting examples of stem cells that may be included among the cells of the different cell types include embryonic stem (ES) cells, adult stem cells, induced pluripotent stem cells (iPSCs), hematopoietic stem cells (HSCs), mesenchymal stem cells (MSCs), neural stem cells (NSCs), and any combination thereof.
[0026] As summarized above, the method of the present disclosure includes growing a pool of cells of different cell types in three dimensions. In certain embodiments, the pool of cells of different cell types is grown in three dimensions at least partially in vivo. According to one non-limiting example, growing the pool in three dimensions includes producing a xenograft from the pool. As used herein, a "xenograft" is a tissue (including a cell graft, e.g., a cell line graft) from one species transplanted into a recipient of a different species. In certain embodiments, the donor species is human and the recipient animal is a mouse, rat, pig, etc. The recipient animal can be immunodeficient (e.g., athymic nude mice, scid / scid mice, non-obese (NOD)-scid mice, recombination-activating gene 2 (Rag2)-knockout mice, etc.). When the recipient animal is a rodent (e.g., a mouse or rat), production of the xenograft may include parenteral injection of a pool of cells of different cell types into the recipient rodent, for example, by tail vein injection. According to some embodiments, the xenograft is a cell line-derived xenograft (CDX), e.g., a xenograft comprising cells from one or more (e.g., two or more, three or more, four or more, five or more, ten or more, or twenty-five or more) different cell lines, non-exhaustive examples of which include tumor cell lines. In certain embodiments, the xenograft is a patient-derived xenograft (PDX), e.g., a xenograft comprising primary cells (e.g., primary tumor cells) from one or more different patients, e.g., two or more, three or more, four or more, five or more, ten or more, or twenty-five or more different patients. Primary cells may, in some examples, be obtained via biopsy as described elsewhere herein.
[0027] In certain embodiments, a pool of cells of different cell types is grown in three dimensions at least partially ex vivo. The term "ex vivo" is used to refer to a process, experiment, and / or measurement performed in or on a sample (e.g., tissue or cells) obtained from a living organism, which process, experiment, and / or measurement is performed in an environment outside the living organism. Thus, the term "ex vivo manipulation" as applied to cells refers to any processing of cells outside of an organism, including, but not limited to, culturing the cells, performing one or more genetic modifications on the cells, and / or exposing the cells to one or more agents. Thus, ex vivo manipulation may be used herein to refer to processing of cells performed outside an animal, for example, after such cells are obtained from the animal or its organs. In contrast to "ex vivo," the term "in vivo" as used herein refers to cells that are in an animal, e.g., a rodent (e.g., a mouse or rat), a pig, etc.
[0028] In certain embodiments, the pool of cells of different cell types is grown at least partially in vitro in three dimensions. According to some embodiments, the pool of cells of different cell types grown in three dimensions is grown in vitro into organoids. By "organoid" is meant a three-dimensional (3D) multicellular in vitro or ex vivo tissue construct that can mimic a corresponding in vivo organ. Organoids can be created through various types of available 3D cell culture systems, including but not limited to 3D bioprinted scaffolds, organ chips, microfluidics-based 3D cell culture models, and the like. According to some embodiments, the pool of cells of different cell types grown in three dimensions is grown in vitro into spheroids. Organoids can be established for an increasing variety of organs, including, but not limited to, the intestine, stomach, kidney, liver, pancreas, mammary gland, prostate, upper and lower respiratory tract, thyroid, retina, and brain, either from tissue-resident adult stem cells (ASCs) sourced directly from a biopsy sample, or from pluripotent stem cells (PSCs), such as embryonic stem cells (ESCs) or induced PSCs (iPSCs). In certain embodiments, a pool of cells of different cell types grown in three dimensions is grown into tissue-derived organoids, for example, from one or more (e.g., two or more) different biopsy samples. Approaches for producing stem cell-derived and tissue-derived organoids are known and described, for example, in Hofer & Lutolf (2021) Nature Reviews Materials 6:402-420.
[0029] Once the pool of cells of different cell types has grown in three dimensions, the three-dimensional pool is treated with a small molecule compound. By "small molecule" compound is meant a compound (e.g., an organic compound) having a molecular weight of 1000 atomic mass units (amu) or less. In some embodiments, the small molecule is 900 amu or less, 750 amu or less, 500 amu or less, 400 amu or less, 300 amu or less, or 200 amu or less. In certain aspects, the small molecule is not made of repeating molecular units such as those found in polymers. According to some embodiments, the small molecule compound is a known therapeutic agent. By "therapeutic agent" or "drug" is meant a physiologically or pharmacologically active substance capable of producing a desired biological effect at a target site in an animal, such as a mammal or human. Therapeutic agents can be any inorganic or organic compound. Therapeutic agents can reduce, inhibit, attenuate, decrease, stop, or stabilize the onset or progression of a disease, disorder, or cell proliferation in an animal, such as a mammal or human. In some embodiments, the small molecule compounds are approved by the U.S. Food and Drug Administration (FDA) and / or the European Medicines Agency (EMA) for use as therapeutic agents in treating one or more diseases, including, but not limited to, any of the diseases described elsewhere herein, such as cancer, cardiovascular disease, cerebrovascular disease, respiratory disease, infectious disease, neurodegenerative disease, dementia, Alzheimer's disease, diabetes, kidney disease, liver disease, etc.
[0030] In some embodiments, the methods of the disclosure include treating the three-dimensional pool with a small molecule compound from a library of small molecule compounds. For example, the small molecule compounds may be from libraries including, but not limited to, MedChemExpress (a collection of 1280 structurally diverse bioactive cell-permeable compounds approved by the FDA and / or EMA, or a collection of 1600 structurally diverse pharmaceutical active cell-permeable compounds that are or have been in several clinical stages), ChemDiv's Master PPI Library (20,000 diverse computationally selected molecules including seven subsets including natural product-based, 3D mimetics, macrocycles, helix turn mimetics, tripeptide mimetics, 3D diverse natural product-like, and Beyond Plains), the MayBridge collection (a set of 13,000 chemically diverse compounds), ChemBridge DIVERSet-CL (a collection of 50,000 small molecules with enhanced therapeutic development potential), TargetMol (a collection of 3200 structurally diverse pharmaceutical active cell-permeable compounds selected to expand and complement NU-HTA's existing FDA approved and clinical collections), and / or any other small molecule compound library of interest.
[0031] The manner in which the three-dimensional pool is treated with the small molecule compound varies depending on the context of the three-dimensional pool. In the context of a three-dimensional pool that is a xenograft containing cells of different cell types, treating the three-dimensional pool may include administering a small molecule compound to a recipient animal (e.g., mouse, rat, pig, etc.). The small molecule compound may be administered via a route of administration selected from oral (e.g., in tablet form, capsule form, liquid form, etc.), parenteral (e.g., by intravenous, intraarterial, subcutaneous, intramuscular, or epidural injection), topical, intranasal, or intraxenograft administration. In the context of a three-dimensional pool that is an organoid, spheroid, or other 3D multicellular structure maintained and / or grown ex vivo or in vitro, treating the three-dimensional pool may include adding a small molecule compound to the cell culture medium in which the three-dimensional pool is present. Suitable conditions for growing and / or maintaining the three-dimensional pool may vary before, during, and / or after treating the pool with the small molecule compound. Such conditions may include growing and / or maintaining the three-dimensional pool in a suitable container (e.g., a cell culture plate or well thereof), in a suitable medium (e.g., cell culture medium such as DMEM, RPMI, MEM, IMDM, DMEM / F-12, etc.), in an environment having a suitable temperature (e.g., 32° C.-42° C., e.g., 37° C.) and pH (e.g., pH 7.0-7.7, e.g., pH 7.4) and a suitable percentage of CO2, e.g., 3%-10%, e.g., 5%.
[0032] After treating the three-dimensional pool with the small molecule compound, the method includes dissociating the cells of the treated three-dimensional pool into single cells. Various suitable approaches can be used to dissociate the cells of the treated three-dimensional pool into single cells. For example, if the three-dimensional pool is an organoid, spheroid, or other 3D multicellular structure maintained and / or grown ex vivo or in vitro, the cells can be dissociated into single cells by digesting the three-dimensional pool using Liberase™ enzyme blend (Millipore Sigma) in DMEM / F12-based medium and digesting for 1 hour with rotation at 37°C. If the three-dimensional pool is a xenograft, the xenograft (e.g., tumor xenograft) can be dissected from a sacrificed animal (e.g., from the flank of a mouse), cut into pieces using a scalpel, and resuspended in 1X Liberase™ enzyme blend in DMEM / F12-based medium 10 U / uL DNAse I 1 mg / mL collagenase IV and digested for 1 hour with rotation at 37°C.
[0033] The disclosed method further includes performing single-cell ribonucleic acid (RNA) sequencing (sometimes referred to as "single-cell RNA-seq" or "scRNA-seq") on the dissociated single cells and on the dissociated single cells from the control three-dimensional pool that are not treated with the small molecule compound. RNA sequencing (RNA-seq) is a genomic approach for the detection and quantitative analysis of messenger RNA molecules in biological samples, which is useful for studying cellular responses. As a surrogate for studying the proteome, some studies have turned to protein-coding, mRNA molecules (collectively referred to as the "transcriptome"), whose expression correlates well with changes in cell traits and cell states. scRNA-seq allows for the comparison of the transcriptomes of individual cells. A variety of suitable approaches for scRNA-seq are available, non-limiting examples of which include C1 (SMARTer) (see, e.g., Pollen et al. (2014) Nat Biotechnol. 32:1053-8), Smart-seq2 (see, e.g., Picelli et al. (2013) Nat Methods 10:1096-8), MATQ-seq (see, e.g., Sheng et al. (2017) Nat Methods 14:267-70), MARS-seq (see, e.g., Jaitin et al. (2014) Science 343:776-9), CEL-seq (see, e.g., Hashimshony et al. (2012) Cell Rep. 2:666-73), Drop-seq (see, e.g., Macosko et al. (2015) Cell Rep. 2:666-73), and Sigma-Aldrich (see, e.g., Sigma-Aldrich). 161:1202-14), InDrop (see, e.g., Klein et al. (2015) Cell 161:1187-201), Chromium (see, e.g., Zheng et al. (2017) Nat Commun. 8:14049), SEQ-well (see, e.g., Gierahn et al. (2017) Nat Methods 14:395-8), and SPLIT-seq (see, e.g., Rosenberg et al. (2017) BioRxiv doi.org / 10.1101 / 105163).Further details regarding single-cell RNA sequencing can be found, for example, in Haque et al. (2017) Genome Med 9, 75.
[0034] In certain embodiments, performing scRNA-seq on dissociated single cells includes labeling the cells according to the Biolegend TotalSeq™-A protocol (www.biolegend.com / en-us / protocols / totalseq-a-antibodies-and-cell-hashing-with-10x-single-cell-3-reagent-kit-v3-3-1-protocol), performing the 10x 3' Chromium single-cell RNA sequencing protocol (support.10xgenomics.com / single-cell-gene-expression / library-prep / doc / user-guide-chromium-single-cell-3-reagent-kits-user-guide-v31-chemistry), and sequencing at approximately 300-400M reads per 10x library and approximately 25M reads per Biolegend TotalSeq™ library. Data from single-cell RNA sequencing can be deconvoluted, for example, into single-cell transcriptomes sorted by treatment (treated with small molecule compounds vs. untreated) and cell type using barcode sequence information.
[0035] Based on the classified single cell transcriptome, the method further comprises evaluating one or more therapeutic properties of the small molecule compound. The method of the present disclosure is used to evaluate a wide variety of therapeutic properties of the small molecule compound. Non-limiting examples of such therapeutic properties include the candidacy of the small molecule compound for combination therapy with drugs (combination therapy), the mechanism of action (MoA) of the small molecule compound, the candidacy of the small molecule compound for treatment of disease subtypes (e.g., for precision oncology, including novel treatments of cancer / tumor subtypes), toxicity of the small molecule compound, mechanism of resistance / tolerance, drug repurposing for new indications not previously tested in the clinic, etc.
[0036] According to some embodiments, the one or more therapeutic characteristics include the small molecule compound being a candidate for combination therapy with a drug, and such methods include: determining the drug sensitivity of each cell line by counting the number of cells remaining in each condition based on the single cell transcriptome classified by treatment and cell type; and calculating the drug-induced gene expression change of each cell line. Such methods further include providing a weighted score for each gene based on its predicted relevance to drug sensitivity based on the calculated drug-induced gene expression change of each cell line. Such methods further include predicting combination therapy targets based on genes with weighted scores above the false discovery rate, and genes that are anti-correlated with drug sensitivity predict drug resistance and therefore represent candidate targets for combination targeting.
[0037] Combination therapy discovery may include determining which cell types are sensitive to a compound and determining gene expression changes within those cell types before and after treatment. Single cell RNA sequencing may be performed with a cell hash, and then each individual cell is demultiplexed using its single nucleotide polymorphism to give a cell line identity after contribution to a separate reference RNA sequencing dataset used to determine reference SNPs. Cell types that are sensitive to a compound may be determined by counting the number of cells remaining in each condition (drug, non-drug) for each cell line. Then, for each cell line, the gene expression difference may be calculated before and after drug treatment, sometimes referred to herein as "single-strain delta." The aggregated single-strain delta for sensitive cell lines may be compared to the aggregated single-strain delta for non-sensitive cell lines to determine which genes are most upregulated in sensitive cells. These aggregated gene expression changes may then be mapped to an online database and a literature search may be used to determine which of these genes are druggable. Genes that are upregulated in response to the compound across all or a majority of the sensitive strains are identified as candidate combination therapy targets.
[0038] In certain embodiments, the one or more therapeutic properties include a mechanism of action of a small molecule compound, and such methods include determining the drug sensitivity of each cell line by counting the number of cells remaining in each condition based on single cell transcriptomes categorized by treatment and cell type, determining the drug-induced gene expression changes of each cell line, and aggregating the determined drug-induced gene expression changes across the drug-sensitive cell lines. Such methods further include providing a weighted score for each gene based on its predicted relevance to drug sensitivity based on the aggregated calculated drug-induced gene expression changes, identifying genes that correlate with the aggregated drug sensitivity as having a weighted score above the false discovery rate, and predicting the mechanism of action of the compound based on the genes that correlate with the aggregated drug sensitivity.
[0039] The mechanism of action discovery may include experimentally administering a pool of cells using a serial dilution of small molecule compound concentrations, determining which cell types are sensitive to the compound, determining within those cell types the change in gene expression before and after treatment, and modeling the change in gene expression in sensitive versus non-sensitive cell lines as a function of small molecule compound concentration. First, the pool of cells may be experimentally subjected to a serial dilution of a range of small molecule compound concentrations. Then, after single cell RNA sequencing with the cell hash, each individual cell may be demultiplexed using its single nucleotide polymorphism (SNP) to give a cell line identity after contribution to a separate reference RNA sequencing data set used to determine the reference SNP. The cell types that are sensitive to the compound may be determined by counting the number of cells remaining in each condition (drug, non-drug) for each cell line. Then, for each cell line, the difference in gene expression may be calculated before and after drug treatment, sometimes referred to herein as "single line delta." The aggregated single-strain delta for sensitive cell lines can be compared to the aggregated single-strain delta for insensitive cell lines to determine which genes are most upregulated in sensitive cells.The gene expression changes can then be modeled as a function of compound concentration, which is used to determine which genes change as a direct function of drug concentration.These concentration-dependent gene expression changes can then be mapped to a reference gene set database to identify the pathways that these genes fit into.In this model, negatively correlated genes as a function of drug concentration indicate the mechanism of action of the drug.
[0040] According to some embodiments, the one or more therapeutic characteristics include the small molecule compound being a candidate for the treatment of the disease subtype, and such methods include determining the drug sensitivity of each cell line by counting the number of cells remaining in each condition based on single cell transcriptomes classified by treatment and cell type, and each cell line is classified by its genetic mutation and / or transcriptomic signature. Such methods further include aggregating the drug sensitivities determined across cell lines, using a variable selection regression algorithm to provide each mutation and / or transcriptomic signature with a score predictive of its association with the aggregated drug sensitivity, and predicting the efficacy of the compound in the disease subtype based on the disease subtype with a score above the false discovery rate. In certain embodiments, the variable selection regression algorithm is a weighted lasso regression algorithm.
[0041] In recent years, therapeutics have been developed to target specific genetic variants of proteins. These specific genetic variants, or genetic subtypes, are often used to determine which drugs a patient should receive. However, there is currently no simple, high-throughput method to evaluate whether a drug developed for a particular genetic subtype can be used effectively against another subtype. The method described herein can simultaneously determine the relative sensitivity of a molecule to various genetic subtypes in in vivo PDX (patient-derived xenografts), in vitro and ex vivo organoid model systems in one experiment. The method involves pooling cells from multiple genetic subtypes. These mixed genetic subtype pools are then drugged using small molecules. Single-cell RNA sequencing can be performed with a cell hash, and then each individual cell is demultiplexed using its single nucleotide polymorphism to give a cell line identity after contribution to a separate reference RNA sequencing dataset used to determine the reference SNP. The cell types that are sensitive to the compound can be determined by counting the number of cells remaining in each condition (drug, non-drug) for each cell line. The cell lines can then be aggregated by their genetic subtypes and evaluated for common susceptibility or resistance of different groups of strains classified by subtype. A regression model (e.g., Lasso regression model) can be trained on the mutations of each cell type to determine which genetic mutations predict the sensitivity calculated from the single-cell RNA sequencing deconvoluted data. Mutations that effectively predict the susceptibility coefficient derived from the data indicate potential targets, and based on the sign of the coefficient of the model variable in the regression (e.g., Lasso regression) of that mutation, it is possible to determine whether the mutation is a resistance (positive) or sensitization (negative) mutation. The inventors have successfully demonstrated that this assay and modeling can be used to predict the stratification of known genetic subtypes, as well as to discover novel genetic subtypes that are sensitive to molecules developed against different genetic subtypes.Subtype stratification, which may be defined as the ability to rank and quantitatively estimate which genetic subtypes induce sensitivity or resistance to a small molecule, can be achieved in one pooled experiment using this method.
[0042] In one aspect, the disclosure provides a balanced cell number culture comprising two or more different cell types cultured for a period of time, wherein each of the at least two different cell types has a growth rate, and each cell type of the two or more different cell types is combined prior to culturing in an inverse ratio to the growth rate of each of the cell types of the two or more different cell types.
[0043] In one aspect, the disclosure provides a balanced cell number culture comprising at least two or more different cell types, wherein a 0.2% to 10% by volume sample of the balanced cell number culture comprises at least 500 cells of each of the different cell types, the sample being taken from the balanced cell number culture after the two or more cell types have been combined to create a cell pool and inoculated into culture medium to obtain the balanced cell number culture, and the balanced cell number culture has been cultured for a period of between 72 hours and 45 days.
[0044] In one aspect, the disclosure provides a balanced cell number culture comprising at least two or more different cell types, each of the cell types comprising at least 1×10 3 A balanced cell number culture is provided, represented by cells, where at least two of the cell types are derived from different cancer tissues.
[0045] In one aspect, the disclosure provides a balanced cell number culture comprising at least two or more different cell types, each of the cell types comprising at least 1×10 3 Provide a balanced cell number culture, represented by cells, where at least two of the cell types contain different cancer mutations from each other.
[0046] In one aspect, the disclosure provides a balanced cell number culture comprising at least two or more different cell types, each of the cell types comprising at least 1×10 3A balanced cell number culture is provided, represented by cells, where at least two of the cell types contain distinct cancer mutations.
[0047] In some embodiments, each of the two different cell types comprises at least 1 x 10 in a balanced cell number culture. 3In some embodiments, the representation of each cell type in the balanced cell number culture is represented by viable cells. In some embodiments, no cell type of the at least two different cell types in the balanced cell number culture outnumbers the other cell types by more than two orders of magnitude. In some embodiments, the total number of each cell type of the at least two or more different cell types is within two orders of magnitude of each other in the balanced cell number culture. In some embodiments, the balanced cell number culture comprises 2-500 different cell types. In some embodiments, the balanced cell number culture comprises 2-500, 5-400, 6-300, 8-200, 10-100, 10-50, 2-30, 2-25, or 10-30 different cell types. In some embodiments, determining the representation of each cell type in the balanced cell number culture comprising a plurality of cell types comprises UMAP analysis. In some embodiments, the UMAP analysis provides the representation of the different cell types in the balanced cell number culture as one or more clusters. In some embodiments, the balanced cell number culture comprises two or more, at least two or more, at least three or more, at least four or more, at least five or more, at least six or more, at least seven or more, at least eight or more, at least nine or more, at least ten or more, at least eleven or more, at least twelve or more, at least thirteen or more, at least fourteen or more, at least fifteen or more, at least sixteen or more, at least seventeen or more, at least eighteen or more, at least nineteen or more, at least twenty or more different cell types. In some embodiments, the balanced cell number culture comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 100 different cell types. In some embodiments, the balanced cell number culture is cultured for a period of 6 hours to 45 days, 12 hours to 40 days, 24 hours to 35 days, 72 hours to 30 days, 96 hours to 20 days, 120 hours to 15 days. In some embodiments, the balanced cell number culture is cultured for a period of 72 hours. In some embodiments, the balanced cell number culture is cultured for a period of 14 days.In some embodiments, each cell type of the two or more different cell types is combined in step (b) in inverse proportion with the growth rate of each of the cell types as determined by a growth rate determination assay, ii) scaled to the total number of days for growth. In some embodiments, the balanced cell number culture is a growth balance culture (e.g., GENEVA pool). The terms balanced cell number culture, growth balance culture, GENEVA pool, and GENEVA culture are used interchangeably herein. In some embodiments, the growth rate determination assay is a calcein-AM growth assay or a Cell Titer Glo growth assay. In some embodiments, the growth rate is determined by a combination of a calcein-AM growth assay and a Cell Titer Glo growth assay. In some embodiments, the growth rate is determined by the formula (target cell number / Euler's constant^(growth rate*days of growth)) / cell number. In some embodiments, the growth rate is determined by the doubling of the cell number. In some embodiments, the doubling is expressed as Nf / N0, where Nf is the cell number at the end of the culture period and N0 is the cell number at the beginning of the culture period. In some embodiments, the growth rate is determined by r=ln(Nf / N0) / t, where "r" represents the growth rate, "t" represents the time of the assay, and ln represents the natural logarithm. In some embodiments, the cells included in the equilibrium cell culture have a growth rate of 0.01-0.8, 0.05-0.8, 0.07-0.7, 0.9-0.5, or 0.1-0.4. In some embodiments, cells from each cell type are included in a ratio such that the representation from each cell type is inversely proportional to their cell growth rate. In some embodiments, the growth rate is measured when the target cell number is equal to 10 million, 20 million, 30 million, 40 million, 50 million, 60 million, 70 million, 80 million, 90 million, or 100 million. In some embodiments, the growth rate is measured when the target cell number is equal to 100 million.
[0048] In some embodiments, the growth rate is taken from the measurements performed as determined by a cell growth assay and the number of days of growth is equal to 1 day. In some embodiments, the growth rate is taken from the measurements performed as determined by a cell growth assay and the number of days of growth is equal to 2 days. In some embodiments, the growth rate is taken from the measurements performed as determined by a cell growth assay and the number of days of growth is equal to 3 days. In some embodiments, the growth rate is taken from the measurements performed as determined by a cell growth assay and the number of days of growth is equal to 4 days. In some embodiments, the growth rate is taken from the measurements performed as determined by a cell growth assay and the number of days of growth is equal to 5 days. In some embodiments, the growth rate is taken from the measurements performed as determined by a cell growth assay and the number of days of growth is equal to 6 days. In some embodiments, the growth rate is taken from the measurements performed as determined by a cell growth assay and the number of days of growth is equal to 7 days. In some embodiments, the growth rate is taken from the measurements performed as determined by a cell growth assay and the number of days of growth is equal to 10 days. In some embodiments, the different cell types include cells with a cancer mutation, cancer cells from one or more subjects, primary cells from one or more subjects, cells from an organ system, cells from a disease model, cells from various cell lines, or any combination thereof. In some embodiments, the different cell types include cells from a subject with a disease. In some embodiments, the different cell types include cells from one or more subjects with a disease. In some embodiments, the different cell types include cells from a disease model, e.g., an organoid, e.g., a xenograft, e.g., a xenograft derived from a patient. In some embodiments, the disease is a neoplastic disease, e.g., a cancer. In some embodiments, the cancer is selected from one or more of cancers of the head, neck, lung, skin, breast, blood, lymph, bone, soft tissue, brain, eye, reproductive system, circulatory system, digestive system, endocrine system, nervous system, and urinary system. In some embodiments, the cell line is a cancer cell line.In some embodiments, the cancer cell lines may include, but are not limited to, one or more of H358, NCI-H23, H2122, H2030, SW1573, SK-LU-1, H441, CALU-1, H1792, H1373, H23, H358, H1299, H1975, SKMEL2, MEWO, SKMEL28, HTT144, A375, MIAPACA2, or A54.
[0049] In some embodiments, the balanced cell number culture is implanted into a model system, such as an in vitro model system, an in vivo model system, or an ex vivo model system. In some embodiments, the model system is a 2D in vitro system. In some embodiments, the model system is a 3D in vitro model system. In some embodiments, the model system is a 3D scaffold system. In some embodiments, the model system is an ex vivo model system, such as an organoid. In some embodiments, the model system is an in vivo model system, such as an animal, such as a mammal, such as a mouse. In some embodiments, the balanced cell number culture is implanted into a single mouse. In some embodiments, implantation of the balanced cell number culture creates a mosaic tumor in the in vivo system. In some embodiments, implantation of the balanced cell number culture creates a mosaic tumor in the mouse. In embodiments, the present disclosure provides a model system comprising a balanced cell number culture, wherein the balanced cell number culture comprises multiple cell types. In some embodiments, the present disclosure provides a model system comprising a mosaic tumor comprising multiple cell types. In some embodiments, the multiple cell types include cells of different physiological origin, cells from different subjects, cells from different organisms, cells from different tissues of the same organism, cells from the same tissue but from different organisms, cells from different tissues or organs from different subjects. In some embodiments, the different types of cells contain at least one different single nucleotide polymorphism from each other. In some embodiments, the different types of cells contain a cancer mutation. In some embodiments, the cells can contain the same cancer mutation. In some embodiments, the cells can contain different cancer mutations.In some embodiments, the mutation is KRAS.G12C, EML4-ALK, TH21, TP53, PIK3CA, PTEN, APC, VHL, KRAS, MLL3, MLL2, ARID1A, PBRM1, NAV3, EGFR, NF1, PIK3R1, CDKN2A, GATA3, RB1, NOTCH1, FBXW7, CTNNB1, DNMT3A, MAP3K1, FLT3, MALAT1, TSHZ3, KEAP1, CDH1, ARHGAP 35, CTCF, NFE2L2, SETBP1, BAP1, NPM1, RUNX1, NRAS, IDH1, TBX3, MAP2K4, RPL22, STK11, CRIPAK, CEBPA, KDM6A, EPHA3, AKT 1, STAG2, BRAF, AR, AJUBA, EPPK1, TSHZ2, PIK3CG, SOX9, ATM, CDKN1B, WT1, HGF, KDM5C, PRX, ERBB4, MTOR, TLR4, U2AF1, ARI D5B, TET2, ATRX, MLL4, ELF3, BRCA1, LRRK2, POLQ, FOXA1, IDH2, CHEK2, KIT, HIST1H1C, SETD2, PDGFRA, EP300, FGFR2, CCND 1, EPHB6, SMAD4, FOXA2, USP9X, BRCA2, NFE2L3, FGFR3, ASXL1, TGFBR2, SOX17, CDKN1A, B4GALT3, SF3B1, TAF1, PPP2R1A, CB Contains one or more of the following mutations: FB, ATR, SIN3A, VEZF1, HIST1H2BD, EIF4A2, CDK12, PHF6, SMC1A, PTPN11, ACVR1B, MAPK8IP1, H3F3C, NSD1, TBL1XR1, EGR3, ACVR2A, MECOM, LIFR, SMC3, NCOR1, RPL5, SMAD2, SPOP, AXIN2, MIR142, RAD21, ERCC2, CDKN2C, EZH2, PCBP1.
[0050] In one aspect, the disclosure provides a method for preparing a balanced cell number culture having at least two or more different cell types, the method comprising: (a) determining a growth rate for each cell type of two or more different cell types; (b) combining two or more different cell types to create a cell pool, where an initial cell number of each of the two or more different cell types added to the cell pool is determined based on the growth rate of step (a); (c) culturing the cell pool of step (b) over a period of time to create a balanced cell number culture, wherein a 0.2% to 10% by volume sample of the balanced cell number culture comprises at least 500 cells of each of the two or more different cell types.
[0051] In some embodiments, the sample in step (c) comprises between 5,000 and 200,000 cells. In some embodiments, the sample in step (c) comprises less than 200,000, less than 175,000, less than 150,000, less than 140,000, less than 130,000, less than 120,000, less than 110,000, or less than 100,000 cells. In some embodiments, the sample in step (c) comprises 500 or more viable cells of each of two or more distinct cell types. The distinct cell types include cells with a cancer mutation, cancer cells from one or more subjects, primary cells from one or more subjects, cells from an organ system, cells from a disease model, cells from various cell lines, or any combination thereof. In some embodiments, the distinct cell types include cells from a subject with a disease. In some embodiments, the distinct cell types include cells from one or more subjects with a disease. In some embodiments, the different cell types include cells from a disease model, e.g., an organoid, e.g., a xenograft, e.g., a xenograft derived from a patient. In some embodiments, the disease is a neoplastic disease, e.g., a cancer. In some embodiments, the cancer is selected from one or more of cancers of the head, neck, lung, skin, breast, blood, lymph, bone, soft tissue, brain, eye, reproductive system, circulatory system, digestive system, endocrine system, nervous system, and urinary system. In some embodiments, the sample of step (c) is taken at the end of the period. In some embodiments, at least two or more samples of step (c) are taken at different times during the period. In some embodiments, the present disclosure provides a method for correlating cells from a sample of step (c) according to any one of claims 33 to 55 with two or more cells of the cell pool of step (b) from the sample of step (c), the performing step comprising: (i) performing single cell RNA sequencing on one or more cells from the sample to identify single nucleotide polymorphisms in one or more cells from the sample; (ii) comparing the single nucleotide polymorphisms of step (i) with the single nucleotide polymorphisms of two or more cells of the cell pool of step (b), thereby correlating the cells from the sample of step (c) with two or more cells of the cell pool of step (b).
[0052] In some embodiments, each of the two different cell types comprises at least 1 x 10 in a balanced cell number culture. 3 In some embodiments, the cell types of at least two different cell types in the balanced cell number culture do not exceed other cell types by more than two orders of magnitude. In some embodiments, the total number of each cell type of at least two or more different cell types is within two orders of magnitude of each other in the balanced cell number culture.
[0053] In some embodiments, the balanced cell number culture comprises between 2 and 500 different cell types. In some embodiments, the balanced cell number culture comprises between 2 and 500, 5 and 400, 6 and 300, 8 and 200, 10 and 100, 10 and 50, 2 and 30, 2 and 25, or 10 and 30 different cell types. In some embodiments, the balanced cell number culture comprises 2 or more, at least 2 or more, at least 3 or more, at least 4 or more, at least 5 or more, at least 6 or more, at least 7 or more, at least 8 or more, at least 9 or more, at least 10 or more, at least 11 or more, at least 12 or more, at least 13 or more, at least 14 or more, at least 15 or more, at least 16 or more, at least 17 or more, at least 18 or more, at least 19 or more, at least 20 or more different cell types. In some embodiments, a balanced cell number culture comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 100 different cell types.
[0054] In some embodiments, the balanced cell number culture is cultured for a period of 6 hours to 45 days, 12 hours to 40 days, 24 hours to 35 days, 72 hours to 30 days, 96 hours to 20 days, 120 hours to 15 days. In some embodiments, the balanced cell number culture is cultured for a period of 72 hours. In some embodiments, the balanced cell number culture is cultured for a period of 14 days. In some embodiments, the balanced cell number culture comprises two or more different cell types, each of the two or more different cell types being at least 1 x 10 in culture. 3 The mosaic tumor is represented by 100 cells, at least two of the cell types containing different cancer mutations from each other. In some embodiments, the disclosure provides a method of generating a mosaic tumor containing at least two or more different cell types in an in vivo model system. In some embodiments, the mosaic tumor is generated by implanting a balanced cell number culture containing two or more cells from a cancer cell line, a cancer tissue, and / or a subject with cancer, and implanting the balanced cell number culture into an in vivo model system. In some embodiments, the in vivo model system is an animal, e.g., a mammal, e.g., a mouse.
[0055] In some embodiments, the balanced cell population culture is transplanted into a model system, for example, an in vitro model system, an in vivo model system, or an ex vivo model system.
[0056] In one aspect, the disclosure provides a method for evaluating the effect of a candidate agent on two or more cell types, the method comprising preparing a balanced cell number culture; The method includes implanting the balanced cell number culture into a model system, treating the model system with a candidate agent for a duration, and evaluating the balanced cell number culture at the end of the duration to determine the phenotypic, genetic, and transcriptomic effects of the candidate agent on individual cells of the balanced cell number culture. In some embodiments, the disclosure provides a method for evaluating the therapeutic efficacy of a candidate agent on individual cells of a mosaic tumor. In some embodiments, the therapeutic efficacy of the candidate agent is measured by treating the mosaic tumor with the candidate agent for a duration, evaluating the individual cells to determine the phenotypic, genetic, and transcriptional expression of the individual cells of the mosaic tumor at the end of the duration, and comparing the phenotypic, genomic, and transcriptional expression of the individual cells of the mosaic tumor to the phenotypic, genomic, and transcriptional expression of the individual cells of the same mosaic tumor that have not been treated with the candidate agent to determine the therapeutic efficacy of the candidate agent.
[0057] In one aspect, the disclosure provides a method for evaluating the effect of a candidate agent on two or more cell types, the method comprising preparing a balanced cell number culture; The method includes implanting the balanced cell number culture into a model system, treating the model system with a candidate agent for a duration, and evaluating the balanced cell number culture at the end of the duration to determine the phenotypic, genetic, and transcriptomic effects of the candidate agent on individual cells of the balanced cell number culture. In some embodiments, the present disclosure provides a method for simultaneously evaluating the effect of a candidate agent on multiple cell types in an in vivo system. In some embodiments, the present disclosure provides a method for simultaneously evaluating the effect of a candidate agent on multiple cell types in an in vitro system. In some embodiments, the present disclosure provides a method for simultaneously evaluating the effect of a candidate agent on multiple cell types in an ex vivo system. In some embodiments, the method includes preparing a balanced cell number culture, transplanting the balanced cell number culture into a model system, treating the model system with a candidate agent for a duration and evaluating individual cells to determine the phenotypic, genetic, and transcriptomic expression of each individual cell of the plurality of cell types at the end of the duration, and determining the impact of the candidate agent by comparing the phenotypic, genomic, and transcriptomic expression of each individual cell of the plurality of cell types in the model system to the phenotypic, genomic, and transcriptomic expression of each individual cell of the plurality of cell types in the same model system that has not been treated with the candidate agent.
[0058] In some embodiments, the present disclosure provides a method for identifying a candidate drug target in a biological pathway. The method includes preparing a balanced cell number culture, transplanting the balanced cell number culture into a model system, treating the model system with a candidate drug for a duration, evaluating the individual cells to determine the phenotypic, genetic, and transcriptomic expression of each individual cell of the plurality of cell types at the end of the duration, and identifying the candidate drug target by comparing the phenotypic, genomic, and transcriptomic expression of each individual cell of the plurality of cell types in the model system with the phenotypic, genomic, and transcriptomic expression of each individual cell of the plurality of cell types in the same model system that has not been treated with the candidate drug. In some embodiments, the present disclosure provides a method for identifying a subject subpopulation that is sensitive to a candidate drug. The method includes preparing a balanced cell number culture, implanting the balanced cell number culture into a model system, treating the model system with a candidate drug for a duration and evaluating the individual cells to determine the phenotypic, genetic, and transcriptomic expression of the individual cells of each of the plurality of cell types at the end of the duration, and identifying a subject subpopulation that is sensitive to the candidate drug based on the evaluation of the phenotypic, genetic, and transcriptomic effect of the candidate drug on the individual cells of the balanced cell number culture. In some embodiments, the disclosure provides a method for identifying a time point at which a subject subpopulation becomes resistant to a drug by determining the phenotypic, genetic, and transcriptomic expression of individual cells from a subject using the methods described herein. In some embodiments, the disclosure provides a method for determining a time point at which a subject subpopulation becomes resistant to therapeutic treatment with a candidate drug by determining the phenotypic, genetic, and transcriptomic expression of individual cells from a subject using the methods described herein.In some embodiments, the present disclosure provides methods for determining an individualized treatment regimen from a population of subjects by using the methods described herein to determine the effect of one or more therapeutic agents on the phenotypic, genetic, and transcriptomic expression of individual cells from a subject, and determining the treatment regimen based on the phenotypic, genetic, and transcriptomic expression of individual cells from the subject.
[0059] In some embodiments, the disclosure provides a method for identifying the effectiveness of a combination therapy by preparing a balanced cell number culture, implanting the balanced cell number culture into a model system, treating the model system with two or more candidate agents in combination for a duration, and evaluating the individual cells to determine the phenotypic, genetic, and transcriptomic expression of each of the individual cells of the plurality of cell types at the end of the duration, and identifying the effectiveness of the combination treatment by the effect of the combination treatment on the individual cells. In some embodiments, the method includes treating with a first candidate agent and treating with a second candidate agent. In some embodiments, the treatment with the first candidate agent and the second candidate agent is sequential. In some embodiments, the treatment with the first candidate agent and the second candidate agent is continuous. In some embodiments, the method optionally includes treating with a third candidate agent.
[0060] In some embodiments, the model system is an in vitro model system, an in vivo model system, or an ex vivo model system. In some embodiments, the model system is a 2D in vitro system. In some embodiments, the model system is a 3D in vitro model system. In some embodiments, the model system is a 3D scaffold system. In some embodiments, the model system is an ex vivo model system, e.g., an organoid. In some embodiments, the model system is an in vivo model system, e.g., an animal, e.g., a mammal, e.g., a mouse.
[0061] In some embodiments, the duration is about 1 hour, about 2 hours, about 4 hours, about 6 hours, about 10 hours, about 16 hours, about 24 hours, about 36 hours, about 48 hours, about 60 hours, about 72 hours, about 84 hours, about 96 hours, about 120 hours, about 1 day, about 2 days, about 3 days, about 4 days, about 5 days, about 6 days, about 7 days, about 8 days, about 9 days, about 10 days, about 11 days, about 12 days, about 13 days, about 14 days, about 20 days, about 24 days, about 25 days, about 30 days, about 35 days, about 40 days, or about 45 days. In some embodiments, the treatment is intermittent. In some embodiments, the treatment is continuous.
[0062] In some embodiments, the candidate agent is an agent capable of causing a therapeutic perturbation. In some embodiments, the candidate agent is selected from a small molecule, an antibody, a peptide, a gene editor, or a nucleic acid aptamer. In some embodiments, the small molecule is a KRAS.G12C inhibitor, such as ARS-1620, AMG510, or MRTX849. In some embodiments, the candidate agent is an inhibitor of a biological pathway. In some embodiments, the candidate agent is an activator of a biological pathway. In embodiments, the candidate agent is selected from one or more of ARS-1620, AMG510, galunisertib, MRTX849, INK128, and antimycin.
[0063] In some embodiments, assessing the phenotypic change comprises counting the number of viable individual cells of each of the two or more different cell types at the end of the duration. In some embodiments, assessing the transcriptomic effect comprises determining a single-cell transcriptomic profile of the cells in the balanced cell number culture at the end of the duration. In some embodiments, assessing the genetic effect comprises single-cell RNA sequencing of the cells in the balanced cell number culture at the end of the duration. In some embodiments, the effect of the candidate agent on the individual cells of the balanced cell number culture is assessed by calculating the gene expression of the individual cells of the balanced cell number culture treated with the candidate agent, and comparing the gene expression to the gene expression of the individual cells of the identical balanced cell number culture not treated with the candidate agent. In some embodiments, the effect of the candidate agent on the individual cells of the balanced cell number culture is assessed by determining the transcriptomic expression of the individual cells of the balanced cell number culture treated with the candidate agent, and comparing the transcriptomic expression to the gene expression of the individual cells of the identical balanced cell number culture not treated with the candidate agent. In some embodiments, the effect of a candidate agent on individual cells of a balanced cell population culture is evaluated by counting the number of viable individual cells of each of the cell types of two or more different cell types in a balanced cell population culture treated with a candidate agent and comparing the number of viable individual cells of each of the cell types of two or more different cell types in an identical balanced cell population culture that is not treated with a candidate agent. In some embodiments, the evaluation includes determining one or more of a genetic effect, a phenotypic effect, and a transcriptomic effect.
[0064] Computer-readable medium and system Aspects of the present disclosure also include computer-readable media and systems that find use in a variety of contexts, including, but not limited to, in practicing the methods of the present disclosure.
[0065] In certain embodiments, one or more non-transitory computer-readable media are provided, on which instructions are stored. When executed by one or more processors, the instructions cause the one or more processors to deconvolute single-cell RNA sequencing data into single-cell transcriptomes classified by treatment and cell type. The single-cell RNA sequencing data is generated by performing single-cell RNA sequencing on dissociated single cells (e.g., xenografts, organoids, etc.) from a three-dimensional pool of different cell types that are treated with small molecule compounds, and dissociated single cells from a control three-dimensional pool of different cell types that are not treated with small molecule compounds. When executed by one or more processors, the instructions further cause the one or more processors to evaluate one or more therapeutic properties of the small molecule compounds based on the classified single-cell transcriptomes.
[0066] In certain embodiments, the instructions cause the one or more processors to evaluate one or more therapeutic properties of the small molecule compound based on the sorted single-cell transcriptome, the one or more therapeutic properties including the small molecule compound being a candidate for combination therapy with a drug. According to some embodiments, the instructions cause the one or more processors to calculate drug-induced gene expression changes for each cell line based on the sorted single-cell transcriptome by treatment and cell type, provide a weighted score for each gene based on its predicted relevance to drug sensitivity based on the calculated drug-induced gene expression changes for each cell line, and predict combination therapy targets based on genes with weighted scores above the false discovery rate, where genes that are anti-correlated with drug sensitivity predict drug resistance and therefore represent candidate targets for combination targeting.
[0067] In some embodiments, the instructions cause the one or more processors to evaluate one or more therapeutic properties of the small molecule compound based on the sorted single-cell transcriptome, the one or more therapeutic properties including a mechanism of action of the small molecule compound. In certain embodiments, the instructions cause the one or more processors to determine a drug-induced gene expression change for each cell line based on the single-cell transcriptome sorted by treatment and cell type, aggregate the determined drug-induced gene expression changes across the drug-sensitive cell lines, provide a weighted score for each gene based on its predicted relevance to drug sensitivity based on the aggregated calculated drug-induced gene expression changes, identify genes that correlate with the aggregated drug sensitivity as having a weighted score above the false positive rate, and predict a mechanism of action of the compound based on the genes that correlate with the aggregated drug sensitivity.
[0068] In certain embodiments, the instructions cause the one or more processors to evaluate one or more therapeutic properties of the small molecule compound based on the classified single-cell transcriptome, the one or more therapeutic properties including that the small molecule compound is a candidate for treating a disease subtype. According to some embodiments, the instructions cause the one or more processors to aggregate drug sensitivity across cell lines based on the classified single-cell transcriptome by treatment and cell type, where drug sensitivity is determined for each cell line by counting the number of cells remaining in each condition, and each cell line is classified by its genetic mutation and / or transcriptomic signature. The instructions cause the one or more processors to provide a score for each mutation and / or transcriptomic signature that predicts an association to the aggregated drug sensitivity using a variable selection regression algorithm, and to predict the efficacy of the compound in the disease subtype based on the disease subtype with a score above the false positive rate. In certain embodiments, the variable selection regression algorithm is a weighted lasso regression algorithm.
[0069] In certain embodiments, a system for evaluating one or more therapeutic properties of a small molecule compound is provided. Such a system comprises one or more processors and one or more non-transitory computer-readable media having instructions stored thereon. When executed by the one or more processors, the instructions cause the one or more processors to deconvolute single-cell RNA sequencing data into single-cell transcriptomes classified by treatment and cell type. The single-cell RNA sequencing data is generated by performing single-cell RNA sequencing on dissociated single cells (e.g., xenografts, organoids, etc.) from a three-dimensional pool of different cell types that are treated with a small molecule compound, and dissociated single cells from a control three-dimensional pool of different cell types that are not treated with a small molecule compound. When executed by the one or more processors, the instructions further cause the one or more processors to evaluate one or more therapeutic properties of the small molecule compound based on the classified single-cell transcriptomes.
[0070] In certain embodiments, instructions of one or more computer readable media of the disclosed system cause one or more processors to evaluate one or more therapeutic properties of the small molecule compound based on the classified single-cell transcriptome, the one or more therapeutic properties including the candidacy of the small molecule compound for combination therapy with a drug, the mechanism of action of the small molecule compound, the candidacy of the small molecule compound for treatment of a disease subtype, or any combination thereof. Examples of instructions of such non-transitory computer readable media for performing these and other types of evaluations are described herein and will not be repeated herein for the sake of brevity.
[0071] Various processor-based systems can be used to implement embodiments of the present disclosure. Such systems can include a system architecture in which components of the system communicate electronically with each other using a bus. The system architecture can include a processing unit (CPU or processor), as well as a cache, variously coupled to a system bus. The bus couples various system components, including system memory (e.g., read-only memory (ROM) and random access memory (RAM)) to the processor.
[0072] The system architecture may include a cache of high speed memory directly connected to the processor, close to the processor, or integrated as part of the processor. The system architecture may copy data from the memory and / or storage device to the cache for quick access by the processor. In this way, the cache may provide performance improvements that avoid processor delays while waiting for data. These and other modules may control or be configured to control the processor to perform various actions. Other system memories may also be used. The memory may include multiple different types of memory with different performance characteristics. The processor may include any general purpose processor, as well as hardware or software modules, such as first, second, and third modules stored in a storage device, configured to control the processor, as well as special purpose processors where software instructions are incorporated into the actual processor design. The processor may essentially be a fully self-contained computing system, including multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetrical or asymmetrical.
[0073] To enable user interaction with the computing system architecture, the input device may represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, etc. The output device may also be one or more of several output mechanisms. In some examples, a multimodal system may enable a user to provide multiple types of input to communicate with the computing system architecture. The communication interface may generally control and manage user input and system output. There is no restriction to operating on any particular hardware arrangement, and thus the basic features of this specification may be easily replaced with improved hardware or firmware arrangements as they are developed.
[0074] The storage device is typically a non-volatile memory and may be a hard disk or other type of computer-readable medium capable of storing data accessible by a computer, such as a magnetic cassette, a flash memory card, a solid-state memory device, a digital versatile disk, a cartridge, a random access memory (RAM), a read-only memory (ROM), and hybrids thereof.
[0075] The storage device may include a software module for controlling the processor. Other hardware or software modules are contemplated. The storage device may be connected to a system bus. In one aspect, a hardware module performing a specific function may include software components stored in a computer-readable medium in association with necessary hardware components such as a processor, a bus, an output device, etc., to perform various functions of the disclosed technology.
[0076] Embodiments within the scope of the present disclosure may also include tangible and / or non-transitory computer-readable storage media or devices for carrying or having stored thereon computer-executable instructions or data structures. Such tangible computer-readable storage devices may be any available device that can be accessed by a general-purpose computer or a special-purpose computer, including the functional design of any special-purpose processor as described above. By way of example and not limitation, such tangible computer-readable devices may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other device that can be used to carry or store desired program code in the form of computer-executable instructions, data structures, or processor chip designs. When information or instructions are provided to a computer over a network or another communications connection (either hardwired, wireless, or a combination thereof), the computer properly views the connection as a computer-readable medium. Thus, any such connection is properly termed a computer-readable medium. Combinations of the above should also be included within the scope of computer-readable storage devices.
[0077] Computer-executable instructions include, for example, instructions and data that cause a general purpose computer, special purpose computer, or special purpose processing device to perform a particular function or group of functions. Computer-executable instructions also include program modules that are executed by a computer in stand-alone or network environments. Generally, program modules include routines, programs, components, data structures, objects, and other functions specific to the design of a special purpose processor that perform tasks or implement abstract data types. Computer-executable instructions, associated data structures, and program modules represent examples of the program code means for executing steps of the methods disclosed herein. A particular sequence of such executable instructions or associated data structures represents examples of corresponding acts for implementing the functions described in such steps.
[0078] Other embodiments of the present disclosure may be practiced in networked computing environments having many types of computer system configurations including personal computers, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, etc. Embodiments may also be practiced in distributed computing environments where tasks are performed by local and remote processing devices that are linked (either by hardwired links, wireless links, or a combination thereof) through a communications network. In a distributed computing environment, program modules may be located in both local and remote memory storage devices. EXAMPLES
[0079] The following examples are offered by way of illustration and not by way of limitation.
[0080] experiment Example 1 - Creation of 3D heterogeneous cell pools This example is directed to the creation of 3D heterogeneous cell pools. A pool of 11 human cell lines from different people was utilized to create a pool of heterogeneous cell lines that would produce a comparable number of cells at harvest after extended growth. A pool was created in which the number of cells at the time of pooling was comparable for each line. The cells were then subjected to a 7-day growth course in cell culture. After 7 days, the pool was harvested and single-cell RNA sequencing was performed on a 10X Chromium platform to obtain a single-cell RNA sequencing Illumina fragment library. The library was sequenced on an Illumina instrument and the reads were aligned to obtain a single-cell gene expression profile. The single-cell data package was then decomposed to the patient of origin using demultiplexing software (demuxlet, freemuxlet). It was determined that more than 90% of the single cells were from cell line A375, while the other 10 cell lines accounted for less than 10% of the remaining cell number. Due to low / limited numbers of cells, specifically H358, H1975, A549, H1299, SKMEL2, and SKMEL28, accurate transcriptome status could not be obtained for some cell lines (Figures 1A-B).
[0081] Example 2 - Creation of a growth rate balanced 3D heterogeneous cell pool (GENEVA) This example is directed to the creation of growth rate-balanced 3D heterogeneous cell pools. The growth rates of the cell lines constituting the pool were measured individually using cell growth rate assays before cells of each cell line were added to the pool (H23, H358, H1299, H1975, SKMEL2, MEWO, SKMEL28, HTT144, A375, MIAPACA2, A549). These growth rates were then used to equilibrate cell numbers at the time of pooling i) in inverse proportion to the growth rate of the cell line determined by the growth assay, and ii) scaled to days for longitudinal growth. After pooling, growth, harvesting, and single-cell RNA sequencing and analysis, the pool of cell lines thus balanced produced a more even distribution of cell types among the different cell lines, allowing accurate single-cell transcriptome profiles from all cell lines included in the pool compared to the pool obtained in Example 1 (Figure 2A-B). To validate that pools created by reverse growth rate equilibration could produce more evenly equilibrated pools, experiments were performed with different cell lines composing the pools (H358, NCI-H23, H2122, H2030, SW1573, SK-LU-1, H441, CALU-1, H1792, H1373) while retaining the methodology of time-based reverse growth rate equilibration. This third pool was also found to be capable of producing a number of cells more evenly distributed across the different cell lines, allowing for accurate transcriptional inference of expression profiles after extended periods of pooled growth (Figure 3A-B).
[0082] These experiments using inverse growth rate and time equilibration demonstrated that it is essential to measure growth rates and follow a strict equilibration approach to produce cell pools that can grow and be equally represented at defined harvest times. Sampling small samples (<100,000 cell samples) from these pools contained representation from all cell lines of origin, whereas for pools that were not created by growth rate inverse equilibration, 1 million cell samples were required to contain representation from all cell lines of origin (Figure 1G).
[0083] Growth rate measurements for determination of GENEVA inclusion criteria Adherent cell lines were grown in RPMI supplemented with 10% fetal bovine serum (FBS) for two passages after thawing. The cell lines were then dissociated into single cell suspensions using trypsin (0.25%). Cell counts were then obtained using an electronic cell counting instrument, and 5,000 cells of each cell line were seeded individually into the wells of two identical 96-well plates. Two hours after seeding, one plate was assayed for cell viability using cell-titer-glow (CTG) reagent from Promega at 2 hours. Briefly, the medium in the 96-well assay plate was removed by decanting, and 50 μL of CTG reagent was added directly to the plate containing the cells. After incubation at 37 C for 30 minutes, the plate was read at 100 V in a 96-well compatible luminometer to measure cell viability. After 72 hours of growth in RPMI 10% FBS, the medium was removed by decanting and 50 μL of CTG reagent was added directly to the plate containing the cells. After 30 minutes of incubation at 37° C., the plates were read at 100 V in a 96-well compatible luminometer to measure cell viability. To measure growth rate, the raw luminescence signal at 72 hours was divided by the luminescence signal at 2 hours. This ratio was estimated to be the doubling of cell number. Hereafter referred to as Nf / N0. The growth rate of each cell line was then calculated using the formula: r=ln(Nf / N0) / t. where "r" represents the growth rate, "t" represents the time of the assay, and ln represents the natural logarithm. The growth rates of all cell lines that make up the pool were determined in this manner, and only cells with growth rates greater than 0.1 and less than 0.4 were determined to be viable candidates for inclusion in the GENEVA pool. Cell lines with growth rates outside these parameters were omitted from further consideration in the GENEVA pool.
[0084] Creation of GENEVA cell pools by inverse growth rate and time ratio Pools were then created by including cells from each cell type such that the representation from each cell type was inversely proportional to their cell growth rate. Cell counts were taken and the following calculations were performed to determine the volume of cell suspension to seed into the GENEVA pool: (Target cell number / Euler constant^(growth rate*growth days)) / cell number Here, the target cell number was equal to 100 million, the growth rate was taken from the measurements performed determined by cell growth assay, and the number of days of growth was equal to 7 days. Using this formula, the number of cells to be added to the GENEVA cell pool was calculated and the cells were combined into a single suspension. The cells were then grown in cell culture in RPMI, 10% FBS for 7 days. The cells were harvested into a single cell suspension by dissociation with 0.25% trypsin. Following the estimation of cell viability and dilution of the cells to 2000 cells / μL, the single cells were loaded with the "GEM Generation Reagent" as specified in the "10X Chromium v3.0" protocol. Further processing of the single cell suspension was performed as described in the 10X Chromium method. Illumina sequencing was performed to obtain 25,000 reads per cell.
[0085] Generation of single nucleotide reference panels for single cell-to-patient-of-origin deconvolution Continuing from above, cell lines determined to be viable for the GENEVA pool were individually seeded into 6-well cell culture plates for growth as follows: Cell lines were dissociated into a single cell suspension using trypsin (0.25%). Cell counts were then obtained using an electronic cell counter and 200,000 cells of each cell line were individually seeded into wells of a 6-well cell culture plate. 2mL of media was then added to serve as the growth medium. After 2 days of growth, the 6-well plates were then harvested for RNA extraction by decanting the media and adding 400uL of Trizol RNA extraction reagent directly to the cells. Extraction of RNA used the ThermoFisher Trizol RNA extraction method. RNA extracted using this procedure was then assayed for purity by transferring to an RNAse-free microcentrifuge tube and aliquoting 2uL of RNA solution onto a nanodrop instrument. Illumina-compatible DNA libraries were prepared using Lexogen's "Quantseq" kit and sequenced on an Illumina instrument.
[0086] Pooled genetic signature data generation from GENEVA-suitable cell lines Cell lines determined to be viable for the GENEVA pool were then prepared into an evenly distributed pooled mixture of cell lines for GENEVA pooled genetic signature data generation. Cell lines were dissociated into single cell suspensions using trypsin (0.25%). Cell counts were then taken using an electronic cell counter and 500,000 cells of each cell line were individually seeded into one 50 mL conical tube containing 5 mL of 1x phosphate buffered saline (PBS) at 4°C (4C). After all cell lines were added to the tube of pooled GENEVA cell lines, centrifugation was performed at 400g for 10 minutes at 4C. The supernatant was decanted and the cell pellet was resuspended in 5 mL of 1x PBS. The cells were spun again at 400g for 10 minutes at 4C, the supernatant was decanted and the pellet was resuspended in 2.5 mL of 1x PBS. The cells were spun again at 400g for 10 minutes at 4C, the supernatant was decanted, the pellet was resuspended in 0.5mL 1x PBS, and the resulting solution was filtered through a 45 micron filter tube to obtain a single cell suspension of pooled cell lines free of contaminating trypsin and fetal bovine serum. This solution was then counted using an automated cell counter and diluted to 2000 cells / μL. Viable cell estimates were also performed by obtaining a count with Trypan Blue, 1x PBS:GENEVA pool, 1x PBS 1:1. Pool viability if greater than 85% was proceeded for single cell RNA sequencing preparation. Following cell viability estimates and dilution of cells to 2000 cells / μL, single cells were loaded with "GEM Generation Reagents" as specified in the "10X Chromium v3.0" protocol, and the resulting Illumina libraries were sequenced to a depth of 25,000 reads per cell.
[0087] Combining i) individual cell line genetic signature datasets with ii) the pooled GENEVA genetic signature dataset to generate a GENEVA-associated single nucleotide polymorphism list. Sequencing data from single nucleotide reference panels were first computationally decomposed to obtain clean single nucleotide polymorphism calls from RNA sequencing data. Sequencing files (fastq format) were trimmed read by read to remove poly-adenylation and Truseq Illumina sequencing adapter contamination, aligned using the "bwa" whole genome alignment tool, sorted and formatted using the "samtools" tool, and de-duplicated by unique molecular identifiers using the "umi_tools" tool. Individual reads that were cleaned of sequence artifacts and fully aligned to genomic locations were then stacked by location using the tool "samtools mpileup" and then converted to bcf and finally a single merged vcf data structure format using the "bcftools" and "vcf-merge" tools. Sequencing data from section 1.d (above) were computationally decomposed to obtain complete genome aligned data from single cell transcriptomes. Sequencing files (fastq format) were processed into aligned ".bam" format using the "cellranger" tool from 10X Genomics as. The resulting ".bam" files were then used for downstream data integration together with the individual cell line vcf files. The merged ".vcf" files containing all detected SNP variants from the individual cell lines (section 1.di) and the ".bam" files containing all SNP variants from the GENEVA pools created from those same lines generated using single-cell RNA sequencing methods were used as input for data integration for the selection and filtering of relevant SNPs used for downstream GENEVA demultiplexing by SNPs. Data integration was performed with the aim of removing computationally uninformative SNPs that would prevent accurate genotyping of single cells from the GENEVA pool back to their cell line of origin. The merged ".vcf" files were intersected with the ".bam" files with a filtering criterion of >250 reads per locus to allow only high-confidence read mapping between both datasets using the "bedtools intersect" tool.These SNPs were then further filtered using a recursive algorithm that integrated the tool "demuxlet" as a way to measure the improvement of the vcf algorithm. The algorithm individually removed each SNP from the merged vcf file to generate a data-subtracted ".vcf" file as a test subject. This data-subtracted ".vcf" file was then used in conjunction with the ".bam" file to run the demuxlet, which provides the relative singlet ratio, a measure of demultiplexing by SNP fidelity. By iteratively testing the individual contribution of single SNPs to the overall demultiplexing by SNP fidelity, a limited list of high-quality SNPs was arrived at and used as high-quality reference SNPs for downstream demultiplexing and further GENEVA experiments using this particular GENEVA pool.
[0088] Example 3 - Comparison of the ability to perform long-term pool growth and drug treatments on non-growing and growing balanced 3D heterogeneous cell pools This example is directed to growing pools for greater than 72 hours while treating with drug compounds to allow for understanding of long term drug effects on cells.
[0089] Using a long treatment duration of 14 days, pools were balanced as described in Example 2 by setting the "days of growth" variable to 14 in the following formula: (Target cell density / Euler constant^(growth rate*number of days to grow)) / cell number
[0090] We created heterogeneous cell pools that were transplanted into a variety of model systems. A small sample of the treated pool (~1% of the total cells) contained enough cells from all cell types of origin to accurately assess the impact of the drug on that cell line over 14 days of treatment (Figure 1C,D,E,F). In contrast, 90% of the cells in the non-growing equilibrated heterogeneous 3D cell pools were from one cell type (Figure 1A,B).
[0091] Example 4 - Creation of in vivo or ex vivo 3D model systems using GENEVA cell pools This example is directed to evaluating the effects of long-term treatment with drug therapies on complex model systems, such as an in vivo mouse model and an in vitro 3D model system.
[0092] Using growth rate balanced heterogeneous cell pools adjusted for long-term drug treatment, pools were created from four human patient-derived xenograft (PDX) models and implanted as pooled tumors in flank xenograft mouse models. Pooled tumors were drugged by oral gavage with molecule ARS-1620 for 14 days in mice. After the treatment interval, tumors were harvested from mice and single-cell RNA sequencing, gene demultiplexing (as shown in Example 2), and sample hashing using barcoded antibodies were performed. A sufficient number of cells were observed from each PDX genetic background and drug treatment condition (Figure 2C, D) for inference of drug phenotype and drug changes to the transcriptome. Pools from four PDX models were created for implantation for 14 days of drug treatment in the 3D organoid model system. Cell pools were treated with three increasing doses of ARS-1620 (0.4uM, 1.6uM, 25.0uM) and one vehicle condition for 14 days. After the treatment interval, tumors were harvested from the 3D organoids and single-cell RNA sequencing, gene demultiplexing, and sample hashing were performed. Sufficient numbers of cells were observed from each PDX genetic background and drug treatment condition (Figure 2A, B) in the organoid model. Using the methods described above, pools were further created using human cell lines for implantation into in vivo flank xenograft mouse models.
[0093] Implantation, drug administration, and harvesting of 3D organoid GENEVA models Five mL of Matrigel basement membrane reagent was added to the GENEVA pool prepared by growth rate equilibration for a final concentration of 5 M / mL. Fifty microliters of this solution was transferred to a 6-well plate and incubated at 37C for 30 minutes. Organoid medium (Advanced DMEM / F12 Bath, 1x N-2, 1x B-27, 10 mM HEPES, 2 mM L-glutamine, 1x Pen-Strep, 500 ng / mL FGF10, 1% FBS) was added to the cells for overnight recovery. After 16 hours, the organoids were drugged with drug compounds diluted in organoid medium. Organoids were grown for 14 days in drug medium with fresh drug compound medium added every 72 hours. On day 14, organoids were harvested by manual dissociation and resuspended in 10mg / mL Liberase™ Cell Dissociation Reagent in 1:1 DMEM:F12 cell media, DNAse I (10U / uL). Organoids were incubated in a 37C incubator with shaking at 600RPM for 45 minutes for enzymatic dissociation, then spun down at 800g for 5 minutes at 4C. Dissociated organoids were resuspended in 100μL of 1x PBS.
[0094] Implantation, Medication, and Harvesting of the In Vivo GENEVA Model 2mL of Matrigel basement membrane reagent was added to the GENEVA pool prepared by growth rate equilibration for a final concentration of 20M / mL. 100 microliters of this solution was injected into NSG mice via flank xenograft injection. After 24 hours, mice bearing GENEVA tumors were drugged with drug compounds in a vehicle solution of 5% DMSO, 95% Labrasol. Mice were orally gavaged for 14 days, dosed for 5 days, and off for 2 days. On day 14, mice were sacrificed and tumors were harvested by homogenization with surgical shears and resuspended in 5mg / mL Liberase™ cell dissociation reagent in 1:1 DMEM:F12 cell medium, DNAse I (10U / μL). Tumors were incubated in a 37°C incubator with shaking at 600RPM for 45 minutes for enzymatic dissociation, then spun down at 800g for 5 minutes at 4°C. Dissociated tumors were resuspended in 100 μL of 1× PBS.
[0095] Example 5 - Identification of cellular origin by single RNA sequencing and transcriptome analysis This example is directed to the donation of single cells to patients or cell lines of origin. Data from GENEVA experiments were used to obtain fastq sequencing files corresponding to GENEVA pooled mRNA libraries and individual generated reference VCF files with known genotypes. GENEVA mRNA libraries were deconvoluted using the tools freemuxlet and demuxlet (github.com / statgen / popscle). A consensus approach was adopted to match clusters called unique individuals by freemuxlet and known populations from the demuxlet approach. This gene-alone approach was then integrated with transcriptome information. Cells were clustered using GENEVA mRNA library data, and clusters with a Leiden rarefaction factor of 10 or more were called. A maximum likelihood match was then given between each transcriptionally defined Leiden cluster and each gene-alone population to obtain a percent frequency representation of each transcriptome cluster. A cutoff of >70% was implemented to obtain clusters with high accuracy between transcriptome and genes, and clusters below this threshold were removed.
[0096] Deconvolution by genetics Single cells were deconvoluted based on single nucleotide polymorphism calls based on single cell RNA. freemuxlet was run with cluster numbers fixed at the number of cell types used as input to the experiment. VCF files were obtained representing the SNPs fed into each group of genetically distinct cells as determined by freemuxlet. A reference VCF generated from single strain genotyping was then used to run demuxlet. A maximum likelihood approach was then used to cross SNPs between the VCF file used for demuxlet and the VCF file, feeding each unknown freemuxlet cluster into a known reference cell line from the demuxlet VCF file to obtain the final cluster assignment by genotype.
[0097] Integrating transcriptome information through genetic deconvolution Single-cell RNA sequencing formatted gene count matrix data from mRNA reads were overlaid with the single-cell genotype calls given by deconvolution by genetics. In the absence of knowledge of these single-cell genotype calls, overclustering was performed with a Leiden rarefaction factor of >10 or by imposition of the Leiden factor, resulting in a total number of clusters according to the following formula: clusters_number=10*number_of_celltypes_in_pool
[0098] A custom merger of transcriptome-based and gene-based calls was used to arrive at a final list of cell types given based on both data sources. For each Leiden cluster, the percentage of cells in that cluster that belonged to a particular SNP-determined genotype was calculated. If more than 70% of cells from that cluster belonged to one SNP-determined genotype, this cluster was marked as high confidence by both gene and transcriptome methods, and the genotype assignment for that cluster was changed to a population of more than 70% of all cells. All other clusters with confidence less than 70% were marked for deletion as low confidence data points that were not reconciled between both genetic and transcriptome assignments.
[0099] Representative Code for Example 5: format_freemuxvcf: ###Subset only those that uniquely distinguish samples freemux_uniq=freemux_full.apply(lambda row:is_uniq(row),axis=1) format_demuxletvcf: ###Subset only those that uniquely distinguish samples demux_uniq=demux_full[demux_full.index.isin(freemux_uniq.index)] match_demux_freemux(demux_uniq,freemux_uniq): ######Compare two snps between freemux and demux snpsum= compare_snps(demux_uniq.loc[snpid],freemux_uniq.loc[snpid]) ######Get the total number of possible comparisons snptotals=get_snptotals(demux_uniq) ######Divide by the total number of possible comparisons tmp_row=[x / snptotals[i]for x in tmp_row] print(snptotals[i],tmp_row) sums_matrix.append(tmp_row)\ ######Get the final grant dictionary assign_dict=get_assigndict(sums_matrix) Adding demuxlets to single-cell data correct_demuxlet(adata,cutoff=0.7,label='cell_type'): #####This function integrates the demuxlet and transcriptome calls to correct by giving each cluster in transcriptome space a cell type. It uses a cutoff defined in the parameters that refers to the maximum cell type fraction. ###### celldict={} ### About the crusts in the set (adata.obs.leiden): ##First get only the clusters clustdf=adata[adata.obs.leiden==clust].obs ##Calculate the percentage props=clustdf.groupby(label).count()['barcode'] / (len(clustdf)) ##See if there is a majority based on a cutoff majorcell=props[props>cutoff].index.tolist() ## Change all cell_type values to the most abundant value cells = clustdf.index.tolist() if len(majorcell)==1: ##Add to dictionary celldict[clust] = majorcell[0] else: celldict[clust]='delete' ###Replace cell type calls adata.obs[label]=adata.obs.apply(lambda row: replace_celltypes(celldict,row,label),axis=1) ###Remove cells flagged as unconfident calls (putative doublets) adata=adata[adata.obs[label]!='delete']
[0100] Example 6 - Identification of experimental sample origin using a noise-corrected algorithm of sample hashed antibodies This example is directed to attributing single cells to their sample of origin using a noise-corrected sample hashing algorithm. A custom baseline read-adjusted algorithm was used to demultiplex single cells according to antibody-labeled (Totalseq from Biolegend) sub-library data. Each cell was mapped to its cell pool of origin (or equivalently, to the drug it was treated with). A confidence metric associated with each cell's attribution was developed, improving the accuracy of sample origin identification to over 90%, an increase over standard methods of maximum read attribution (Figure 9A).
[0101] Custom Baseline Lead Adjustment Algorithm The steps described here outline the operation of the custom baseline read adjustment algorithm. Single-cell RNA sequencing barcode identifiers were obtained as a whitelist from the 10X Genomics Cellranger Pipeline. A table was generated listing the antibody hash sequences and the corresponding samples associated with those hashes. Raw fastq data was read and constant sequences were removed from the raw reads to isolate only the scRNAseq barcode regions and only the antibody labeled barcode regions. Barcodes were corrected from the scRNAseq barcode reads to the single-cell whitelist within a Hamming distance of 1. Barcodes were corrected from the antibody labeled barcode reads to the antibody hash sequences within a Hamming distance of 1. Reads were filtered by post-hamming the corrected barcodes to be an absolute match to both whitelists. The barcodes were then fed through the median baseline read adjustment algorithm. For each antibody hash, the background noise to subtract was calculated according to the following formula: median_correction_factor*median_reads_per_antibody where the median_correction_factor is typically around 1.6 for best performance. For each cell, the calculated number of reads as background for its specific antibody hash was subtracted to obtain a denoised dataset per custom antibody of raw reads. For each cell, the percentage of reads from each antibody hash was calculated. A final list of single-cell barcodes, and their corresponding highest probability antibody hash, and corresponding confidence interval values were returned. The confidence interval was calculated by taking the percentage of reads from the best identifier and subtracting the percentage of reads from the second best identifier. The best scoring antibody hash from each cell was given as the best identifier. Cells with confidence intervals below 0.4 were removed.
[0102] Representative Code for Example 6 def correct_median(readmatrixmed_factor=1.4): #####Subtract the median multiplied by the coefficient for each barcode (i.e. remove the background) testmed=pivot.transpose()-med_factor*(pivot.transpose().apply(np.median)) #####Set negative numbers to zero testmed[testmed<0]=0 ###Calculate pct for each cell testmed = testmed / testmed.sum() ####Get significance value and call cells new = test.copy() new['call']=test.apply(lambda row:get_sig(list(row),list(test.columns)),axis=1) pymulti(R1,R2,bcs10x,len_10x=16,len_umi=12,len_multi=8,med_factor=1.6,gbc_thresh=None,sampname='pymulti_',split=True,plo ts=True,hamming=False,thresh=False,pct_only=False,median_only=False,huge=False,bcsmulti=None,reads=None,thresh_dict={}): ”””main loop,splits from fastqs and runs through cell calls R1=multiseq / Hash fraction of Read1 fastq R2=multiseq / Hash fraction of Read2 fastq bcsmulti = Whitelist of known multiseq / hash barcodes bcs10x=Whitelist of known cell identifiers len10x=10x the length of the barcode len_umi = length of each pair of reads in umi len_multi = the length of the multiseq barcode sequences the package will split med_factor = if performing a median correction, this is the factor multiplied by the median to estimate the amount of background reads per barcode sampname=This is the handle where the intermediate file will be saved split=True if you haven't split this from the fastq yet, False if you've already split it and want to undo the process. This saves time. hamming=multiseq / True if you want to use multiseq harcode that matches within a Hamming distance of 1 against the hash whitelist thresh=lead, True if you want to use a threshold to gate the different compensation methods pct_only = True if you want to call the barcode using only the raw percentage of the lead median_only=True if you want to use the median correction method huge = True if the file is huge, in which case it will be saved as an h5 file instead of pickled thresh_dict = Barcode dictionary: The threshold used to gate sample reads """ ### if huge==True:print("Assume huge fastqs") ###Split fastqs and pickles os.system('mkdir pymulti') if split==True:reads= split_rawdata(R1,R2,len_10x,len_umi,len_multi,sampname,huge=huge) ###Lead old pickle data readtable=read_pickle(sampname,reads=reads,huge=huge) #####Check for duplicates of multi-rate and 10x rate multirate=check_stats(readtable,bcsmulti,bcs10x) ####Match by Hamming distance if hamming==True: readtable.drop_duplicates(inplace=True) matchby_hamdist(readtable,bcsmulti,bcs10x) #####Check for duplicates of multi-rate and 10x rate multirate=check_stats(readtable,bcsmulti,bcs10x) ####Filter Lead Table filtd=filter_readtable(readtable,bcsmulti,bcs10x,gbc_thresh) ####Implementing multiseq correction within the distribution zscores if median_only==True: correct_median(filtd,sampname,med_factor,plots) else: correct_simple(filtd,sampname,plots,thresh,pct_only,thresh_dict) return(filtd)
[0103] Example 7 - Determining genetic drivers of susceptibility to candidate drugs using GENEVA This example is directed to discovering genetic drivers of sensitivity to drug compounds by simultaneous inference of phenotypes from GENEVA cell pools. Long-term growth equilibrium cell pools were created and subjected to drug treatment. After gene demultiplexing using the method described in Example 5 and assignment to patient of origin and drug treatment condition using sample hash demultiplexing as described in Example 6, a dataset with a distinct number of cells in each drug treatment condition was obtained, as shown in Figure 2A, B, C, D, E, F. For each cell type, the relative number of viable cells remaining in each drug treatment condition was counted, and this drug sensitivity information was used as a response variable in a linear model. Full exome genetic mutation data for all cell lines was obtained, and these features were used as explanatory variables in linear model regressions. A Lasso feature selection model was applied to determine which genetic mutations were involved in drug sensitivity to different compounds. Known compounds that target cell lines with specific genetic mutations were used in the proof of concept experiment, specifically vemurafinib (VEM), which targets cells with BRAF.V600E mutations, or ARS-1620 (ARS), which targets cells with KRAS.G12C mutations. Regression using the Lasso model correctly predicted the mutations inducing sensitivity to the compounds for both vemurafinib and ARS-1620, in line with prior knowledge of these compounds (Figure 3B, 3C). Several doses of ARS1620 were used to simultaneously generate enough information in one pooled experiment to generate survival curves for multiple cell types (Figure 3D), equivalent to IC50 measurements. These relative sensitivities were compared across cell lines to rank cell lines that were more or less sensitive to ARS1620 (Figure 3E).
[0104] Transcriptomic data from the GENEVA assay were also used to determine the drug phenotype. After drug treatment with ARS-1620 in both in vivo and ex vivo systems, xenografts from patients with known KRAS.G12C mutations were shown to be sensitive to the compound (patients 877 and 233) by estimating the percentage of cell cycle inhibition (Figure 3F). In the context of this PDX, the dataset also allowed us to infer the sensitive driver mutation (KRAS.G12C).
[0105] Calculating the fitness of a cell type to a drug The number of cells in each drug condition was counted according to the cell type of origin. These raw numbers were then used as input for a normalization-based function that adjusted for differences in total cell numbers between drug treatment conditions and returned an adjusted fitness calculation for each cell type according to the following formula:
[0106] For each cell type, Ci is given a Dx and D0 for each drug and vehicle pair, respectively, and a matrix of cell numbers with the following data structure: [Table 1] Based on the proportional expression, the fitness of cell type Ci to drug Dx was calculated as follows: F(Ci,Dx)=((#cells Dx,Ci) / SUM(#cells Dx)) / (#cells D0,Ci) / SUM(#cells D0))
[0107] Samples were also adjusted based on dataset size following relative sample normalization by calculating an adjustment factor to account for differences in total cell numbers that may escape proportional calculations.
[0108] For each drug or vehicle Dx, the total number of cells in each condition was counted, the geometric mean of the total cell numbers was calculated, and the total number of cells in each condition was divided by the geometric mean of the cell numbers to obtain a normalized ratio. The following matrix was divided by the corresponding normalization for each condition Dx to obtain a final corrected matrix of cell numbers adjusted for dataset size according to the geometric mean ratio-based correction: [Table 2]
[0109] For each of these fitness calculations, the relative fitness was calculated as follows: Cell line fitness Y = (Fitness of cell line Y) / (Maximum fitness among all cell lines in the pool)
[0110] Lasso regression for discovery of genotypic drivers of drug sensitivity Mutations derived from whole-exome sequencing data for each cell type were downloaded and sorted for crossover across at least two different cell types. These were then formatted into a table of dependent variables suitable for input into the Lasso regression algorithm for feature selection. In a paired manner, the fitness of each cell line was also calculated and formatted as a response variable for the Lasso algorithm. Lasso regression was then run with 50 interval steps, with an alpha between 0.05 and 0.30, to find the most relevant mutations predicting GENEVA cell response. The highest scoring explanatory variable gene mutations were ranked by their covariates and designated as drivers of drug sensitivity.
[0111] Reconstructed IC50 curves based on GENEVA phenotypes across ARS-1620 multiple dose curves Cell counts were obtained across the different drug treatment conditions, specifically across different concentrations of drug. These numbers were annotated at the time of harvest and used to calculate the absolute number of cells for each cell type. The formula was: Absolute number of cells of cell type Cy in drug condition Dx = (Number of counted cells of Cy in Dx) / (Total number of cells counted at Dx)* Number of cells counted from annotation at harvest
[0112] These values were then used as input to a logistic regression fit and IC50 curves were interpolated from multiple dose cell numbers.
[0113] Determination of percentage of cell cycle inhibition as an alternative proxy to cell counting as a phenotypic readout Cell cycle phase was calculated by regressing whole transcriptome readouts and weighting according to specific genes of interest associated with different cell cycle states. Cells were subjected to G1, S, and G2 / M phases. Cell cycle ratios were then calculated as a percentage of cells per population. (G2M #cells) / (Total #cells)
[0114] These ratios were then plotted as a linear function over the experimental dose regimen to which they were subjected. In this case, the dose curve for ARS-1620 (uM). x=ARS-1620(uM) A linear fit with y=(#cells in G2M) / (total #cells) is The resulting function had its slope determined as percent cell cycle inhibition, reflecting an alternative method for measuring phenotype derived solely from transcriptome drug perturbation data over a long time course.
[0115] Representative Code for Example 7 ########## Calculating the fitness of a cell type to a drug ########## ####Get the length of each processed dataset sizedict={} Drugs in the set (adata.obs[treatment]): onedf=adata[adata.obs[treatment]==drug].obs sizedict[drug] = len(onedf) ####deseq normalize If deseq_norm==True: Counting cells by type For cell types in the set (adata.obs[label]): ####Counting onedf=adata[adata.obs[label]==celltype] props=pd.DataFrame(onedf.groupby([treatment]).count()['barcode']) results[celltype]=props['barcode'] #####Create a pseudo reference and perform deseq pseudo=results / gmean(results.iloc[:,:],axis=0) #####Correction results=results.transpose() / ratio.tolist() #####Divide everything with vehicle results=results.div(results[vehicle],axis=0) ####Regular normalization else: ####Calculate the relative percentage per cell line to obtain fitness relative to vehicle ##Save the data in the results For cell types in the set (adata.obs[label]): ####Counting onedf=adata[adata.obs[label]==celltype].obs props=pd.DataFrame(onedf.groupby([treatment]).count()['barcode']) props=pd.DataFrame(props.apply(lambda row:row['barcode'] / sizedict[row.name],axis=1)) ###Divide everything with a vehicle results=results.div(results[vehicle],axis=0) ###Scale everything to the maximum results = results / results.max() ########## Lasso regression for discovery of genotypic drivers of drug sensitivity ########## #######Get only mutations that exist in two or more uniq=pd.DataFrame(mutations_table.sum(axis=0)) uniq=uniq[uniq[0]!=1] uniq = uniq.index.tolist() table=table[uniq] #####Importing response variables vem=fitness['VEM'] ars=fitness['ARS1620'] ########Putting response variables into gene expression data vem_x['y']=vem_x.apply(lambda row:vemurafinib_dict[row.name],axis=1) ############Define the alpha range to search alphas=np.linspace(0.05,0.3,50) alphas=alphas[::-1] coefs=[] ########Calculate alpha Alpha on Alpha: reg=linear_model.Lasso (alpha=alpha) reg.fit(vem_x,vem_y) coefs.append(reg.coef_) ########Visualize the coefficients of all genetic variants log_alphas = -np.log10(alphas) coefs = np.transpose(coefs) coef,c in zip(coefs,colors): plt.plot(log_alphas,coef,c=np.random.rand(3,),linewidth=3.0) ########## Reconstructed IC50 curves based on GENEVA phenotypes across ARS-1620 multiple dose curves ########## ####Scaling cell numbers across cell types and drugs tmp=invitro_clean.obs[['cell_line','sample']] tmp=pd.pivot_table(tmp,values='index',index=['cell_line'],columns=['sample'], aggfunc=lambda x:len(x.unique())) ####Define the dosing regimen conc_dict={ 'C0_1':0.00, 'C0_2':0.00, 'C1_1':0.4, 'C1_2':0.4, 'C2_1':1.6, 'C2_2':1.6, 'C3_2':6.3, 'C4_1':25.0, 'C4_2':25.0 } #####Multiply by cell number from annotation at time of harvest tmp=tmp*[annotated_actual_cell_counts] Logistic regression sns.lmplot(data=tmp,x=”concentration(uM)”,y=”%survival,logistic=True)
[0116] Example 8 - Identification of mitochondrial mechanism of action of candidate drugs during chronic drug treatment using GENEVA This example is directed to identifying the molecular mechanism of action of compound ARS1620 by mitochondrial gene downregulation using GENEVA.
[0117] Long-term GENEVA pools were created to understand how ARS1620 achieved sustained tumor regression in human cells. A 7-day drug treatment regimen was performed and transcriptomic changes in surviving cells were analyzed. Mitochondrially encoded genes were downregulated, indicating the effect of ARS1620 on mitochondria (Figure 6A-B). As validation, generation of experimental lines subjected to long-term treatment under ARS1620 showed that mitochondria were significantly eliminated after long-term ARS1620 treatment (Figure 6C). Using an assay of cellular respiration, significant differences in functional oxygen consumption at mitochondrial sites were also detected as a result of ARS1620 treatment (Figure 6D). Analysis of the subpopulation composition of cells within each cell line constituting the GENEVA pool revealed that following drug treatment with ARS1620, surviving cells exhibited a significantly lower proportion of mitochondrial reads compared to overall reads, indicating a population-level selection pressure for cells with downregulated mitochondrial gene expression (Figure 6E).
[0118] Differential expression of each cell line between drug and non-drug conditions for causal gene set discovery For each cell line, the single-cell dataset was split into two sub-datasets: vehicle-treated and drug-treated. Differential expression was calculated between the two single-cell datasets using a two-sample t-test. A differential expression score for each gene was obtained from the two-sample t-test output. This was repeated until all cell lines were completed, and all differential expression scores for genes were stored in an aggregated differential expression table per cell line. Cells were grouped according to phenotypic sensitivity to the tested compounds. For the tested compounds, genetic drivers of compound sensitivity were discovered as described earlier herein. Using the genetic drivers, cell lines were classified into two categories, sensitive and insensitive, by the presence or absence of the driver. The differential expression matrix was grouped into sensitive and insensitive cell lines. The z-scores of genes within each cell line were calculated, and genes were grouped by gene sets from a biological gene set database curated from the scientific literature, and a two-sample T-test with the grouped z-scores was performed. First, all differential expression scores were normalized to the same scale by the z-scores within each cell line. For all genes in the differential expression matrix, a dictionary of groupings was created taken from each of the mSIGDB databases (https: / / www.gsea-msigdb.org / gsea / msigdb / ). For each grouping of genes from mSIGDB, a two sample T-test was performed for each gene set treating all "sensitive" and "non-sensitive" cell lines separately as replicates. Summary statistics from each gene test were saved and z-score differences were calculated. The biological mechanism of action of the molecule of interest was determined by ranking the median z-score results from each gene set to determine the relative up- or down-regulation of the gene set in response to the compound.
[0119] Representative Code for Example 8 ####Deathcore Matrix de_score_all_lines ####msigdb full database msigdb #### de_zscored_all_lines=de_score_all_lines.apply(zscore, axis=0) genesets_zscored=de_zscored_all_lines.groupby(by=msigdb) ttest_results=genesets_zscored.apply(lambda row:stats.ttest_ind(row['sensitive'],row['insensitive']),axis=1) median_zscore_results=genesets_zscored.apply(lambda row:np.median(row['sensitive'])-np.median(row['insensitive']),axis=1) ####These are the identities of susceptible and less susceptible strains classified by genetic drivers sensitive_lines insensitive_lines ####Deathcore Matrix de_score_all_lines #### For cell_line in de_score_all_lines: if cell_line in sensitive_lines: de_score_all_lines[cell_line].index.replace(“sensitive”) if cell_line in insensitive_lines: de_score_all_lines[cell_line].index.replace(“insensitive”) ###scdata contains all single cell data in the anndata data structure format ###de_score_all_lines contains the final output from this section Import scanpy as sc For cell_line of all_lines: vehicle_data=scdata[scdata.obs.drug=='vehicle'] drug_data=scdata[scdata.obs.drug=='drug'] differential_expression=sc.tl.rank_genes_groups(vehicle_data,drug_data,method='t-test') de_score=scdata.uns['t-test']['scores'] de_score_all_lines.append(de_score)
[0120] Example 9 - Identification of the mechanism of action of candidate drugs in inducing ferroptosis using GENEVA This example is directed to identifying the molecular mechanism of action of compound ARS1620 by upregulating the ferritin gene using the GENEVA cell pool.
[0121] Upregulated genes were analyzed across KRAS.G12C lines in cell pools of cells surviving long-term ARS1620 treatment. Consistently upregulated genes across cell lines were found to be involved in anti-ferroptosis response mechanisms (Fig. 7A,B). Among this group of genes were FTH1 and FTL, two components of the ferritin complex involved in sequestration of labile free iron. Using a lipid peroxidation live cell probe, lipid peroxidation - one of the hallmark phenotypes of ferroptosis - was measured in response to ARS-1620 treatment. Dose curves demonstrated that ARS-1620 induced lipid peroxidation in a dose-dependent manner (Fig. 7C). Furthermore, comparison of cell survival and lipid peroxidation kinetics with known inducers of ferroptosis (erastin and altretamine) showed that ARS-1620 behaved similarly to these ferroptosis inducers in its ferroptosis / survival kinetics (Fig. 7D). Survival and normalized lipid peroxidation curves crossed around the IC50 survival value for all three compounds demonstrating similar pharmacodynamics. In the KRAS.G12C line (H2030), the three KRAS.G12C inhibitors MRTX849, AMG510, and ARS1620 all increased lipid peroxidation, whereas in the non-KRAS.G12C line (H441), this effect was significantly muted (Figure 7E).
[0122] Example 10 - Identification of drug resistance mechanisms using GENEVA This example is directed to identifying multiple targetable drug resistance mechanisms from long-term drug treatment in GENEVA cell pools.
[0123] Gene expression indicators of resistance mechanisms were obtained using the GENEVA dataset for ARS-1620 in a 14-day drug treatment timeline where drug resistance occurred (Figure 4A). Inhibitors of these targets were obtained and tested for drug synergy in combination with three KRAS.G12C inhibitors, ARS1620, AMG510, and MRTX849. Here, drug synergy is defined as the sum of the drug combination exceeding the additive model of the individual compounds alone. Several GENEVA predicted drug targets demonstrated high Bliss synergy scores (scores above 0 indicated significant drug synergy, Figure 4B). The strongest GENEVA prediction, mTOR resistance, was tested in a multi-arm in vivo mouse study and significant in vivo Bliss drug synergy and tumor regression over time were found (Figure 4C,D).
[0124] Differential expression of each cell line between drug and non-drug conditions for identification of drug resistance pathways For each cell line, the single-cell dataset was split into two sub-datasets: vehicle-treated and drug-treated. Differential expression was calculated between the two single-cell datasets using a two-sample t-test. A differential expression score for each gene was obtained from the two-sample t-test output. This was repeated for all cell lines until complete, and all differential expression scores for genes were stored in a differential expression table aggregated by cell type. A z-score was calculated for genes within each cell line, and a two-sample T-test was performed with the grouped z-scores. First, all differential expression scores were normalized to the same scale by the z-score within each cell line. A two-sample T-test was then performed for each gene treating all "sensitive" and "non-sensitive" cell lines individually as replicates. Summary statistics were saved from each gene test, and the z-score difference was calculated. Genes were subsetted by summary statistics and druggability of the gene product. Using the p-value subsetted t-test results for each gene by selecting genes with p-values 0.01 or less. Using the median results subsetted by z-score, the delta between sensitive and non-susceptible strains for each gene is changed by selecting genes with z-score values greater than the top 80 percentile of gene scores. Using DGIDB.org as a resource for druggable targets, cross-referenced genes from summary statistics with druggable targets from DGIDB.org to obtain a final target list of drug resistance targets induced by the compound of interest.
[0125] Representative Code for Example 10 ###scdata contains all single cell data in the anndata data structure format ###de_score_all_lines contains the final output from this section Import scanpy as sc For cell_line of all_lines: vehicle_data=scdata[scdata.obs.drug=='vehicle'] drug_data=scdata[scdata.obs.drug=='drug'] differential_expression=sc.tl.rank_genes_groups(vehicle_data,drug_data,method='t-test') de_score=scdata.uns['t-test']['scores'] de_score_all_lines.append(de_score) ####Deathcore Matrix de_score_all_lines #### de_zscored_all_lines=de_score_all_lines.apply(zscore, axis=0) ttest_results=de_zscored_all_lines.apply(lambda row:stats.ttest_ind(row['sensitive'],row['insensitive']),axis=1) median_zscore_results=de_zscored_all_lines.apply(lambda row:np.median(row['sensitive'])-np.median(row['insensitive']),axis=1) ####These are the identities of susceptible and less susceptible strains classified by genetic drivers sensitive_lines insensitive_lines ####Deathcore Matrix de_score_all_lines #### For cell_line in de_score_all_lines: if cell_line in sensitive_lines: de_score_all_lines[cell_line].index.replace(“sensitive”) if cell_line in insensitive_lines: de_score_all_lines[cell_line].index.replace(“insensitive”) ttest_results median_zscore_results #### ttest_filtered=ttest_results[ttest_results['pvalue']<=0.01] median_zscore_filtered=median_zscore_results[median_zscore_results['zscore']>=np.percentile(median_zscore_results['zscore'],80)] intermediate_genelist=set(ttest_filtered['genes]).intersect(set(median_zscore_filtered['genes'])) ####Intersect the final list with the DGIDB DGIDB_genelist final_genelist=set(intermediate_genelist).intersect(set(DGIDB_genelist))
[0126] Example 11 - Identification of in vivo specific drug resistance mechanisms using GENEVA This example is directed to the identification of induction of endothelial-mesenchymal transition as an in vivo specific mechanism of tumor resistance to ARS1620 using GENEVA.
[0127] We compared GENEVA datasets drugged with the KRAS.G12C inhibitor ARS1620 to GENEVA pools performed both in vitro and in vivo. Specifically, in vitro data was compared to in vivo data to look for differences and similarities due to the context of the model system used. One of the most upregulated gene sets in response to the drug was specific to the in vivo context and showed no differences in vitro (Figure 5A). An endothelial-mesenchymal transition characteristic gene expression signature was increased in vivo, indicating a possible drug adaptability mechanism for ARS1620 KRAS.G12C inhibition. As a validation study, a combination therapy multi-arm in vivo mouse study was designed to test the efficacy of the EMT inhibitor, galunisertib, in combination with ARS1620. The combination therapy was found to be highly effective in suppressing tumor growth (Figure 5B) and acted together with ARS1620 to synergistically reduce growth (Figure 5C).
[0128] Discovery of causal gene sets specific to in vivo model systems The median z-score results determined in Example 8 and the t-test results for the in vivo and in vitro datasets, respectively, were combined and combined on a gene set by gene set paired basis. The combined z-score result table from both in vivo and in vitro was then used to calculate an "in vivo specificity score" as follows: For each gene set, the "in vivo specificity score" = [(in vivo median z score)-(in vitro median z score)] / Mean ([log10(in vivo p-value), log10(in vitro p-value)]
[0129] The ordered gene sets were then ranked by in vivo specificity score to arrive at the gene sets that were specifically significantly up- or down-regulated in vivo in response to ARS1620.
[0130] Example 12 - Testing the Efficacy of Combination Therapy with GENEVA This example is directed to testing combination therapies to identify optimal combination therapies in GENEVA cell pools.
[0131] Utilizing a panel of G12C and non-G12C lines pooled in vivo as xenografts, we performed a combination therapy study comparing ARS1620 with or without galunisertib, INK128, and antimycin for a total of eight different treatment conditions. These compounds were derived from combination therapy discovery and mechanistic understanding of ARS1620's mitochondrial action (Figure 8A). Drug synergy was calculated across several cell lines for each drug using the cell cycle active / inactive percentages in each drug condition. Ink128 and galunisertib demonstrated relative in vivo synergy consistent with previous validation experiments demonstrating in vivo synergy, while antimycin demonstrated antagonistic effects demonstrating a relative rescue of ARS1620 consistent with antagonism of the mitochondrial lethality phenotype (Figure 8B). A linear model built around gene expression and drug treatment + cell line of origin to discover genes that drove the synergistic drug phenotype revealed that galunisertib and INK128 together were able to further increase the synergistic reduction in mitochondrial leads, consistent with the general effect of ARS1620 on mitochondria observed alone (Figure 8C).
[0132] Example 13 - Identification of patient subpopulations sensitive to candidate agents using GENEVA This example is directed to the detection of novel patient subpopulations sensitive to a candidate drug (ARS1620) in PDX models.
[0133] GENEVA pools of PDX models were created for long-term drug treatment assays. Tumors were transplanted in vivo and into organoids, and KRAS.G12C, EML4-ALK, and TH21 lung cancer patient tumors were drugged into GENEVA pools. Significant drug sensitivity was observed in EML4-ALK patients, which was higher than the sensitivity of KRAS.G12C mutant tumors, indicating that EML4-ALK patient tumors responded to KRAS.G12C inhibitors (Figure 9B).
[0134] Thus, the above description merely illustrates the principles of the present disclosure. It will be appreciated that those skilled in the art can devise various configurations that embody the principles of the present invention and are within the spirit and scope of the present invention, although not explicitly described or illustrated herein. Furthermore, all examples and conditional language recited herein are intended primarily to aid the reader in understanding the principles of the present invention and the concepts contributed by the inventors to promote the art, and should not be interpreted as limitations to such specifically recited examples and conditions. Furthermore, all descriptions herein reciting the principles, aspects, and embodiments of the present invention, as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. In addition, such equivalents are intended to include both currently known equivalents and equivalents developed in the future, i.e., any elements developed that perform the same function, regardless of structure. Thus, the scope of the present invention is not intended to be limited to the exemplary embodiments shown and described herein.
Claims
Claim 1: A balanced cell number culture comprising at least two or more different cell types, wherein a sample of 0.2% to 10% by volume of the balanced cell number culture contains at least 500 cells of each of the different cell types, the sample combines two or more cell types to create a cell pool, inoculates the culture medium, and after obtaining the balanced cell number culture, after culturing the balanced cell number culture for a certain period, a balanced cell number culture taken from the balanced cell number culture. Claim 2: Each of the two or more different cell types has a growth rate, and each cell type of the two or more different cell types is combined in inverse proportion to the growth rate of each of the cell types of the two or more different cell types before culturing. The balanced cell number culture according to Claim 1. Claim 3: The period is from 6 hours to 45 days. The balanced cell number culture according to Claim 1. Claim 4: (i) Each of the two different cell types is represented by at least 1×10³ viable cells in the balanced cell number culture, (ii) The cell types of the at least two different cell types in the balanced cell number culture do not exceed the cell number of other cell types by more than two digits, or (iii) In the balanced cell number culture, the total number of each of the at least two or more different cell types is within two digits relative to each other. The balanced cell number culture according to Claim 1. Claim 5: At least two of the different cell types are derived from different cancer tissues. The balanced cell number culture according to Claim 1. Claim 6: At least two of the different cell types contain different cancer mutations. The balanced cell number culture according to Claim 1. Claim 7: The balanced cell number culture contains 2 to 100 different cell types. The balanced cell number culture according to Claim 1. Claim 8: The balanced cell number culture contains 10 to 50 different cell types. The balanced cell number culture according to Claim 7. Claim 9: The different cell types include cancer mutation - having cells, cancer cells from one or more subjects, primary cells from one or more subjects, cells from organ systems, cells from disease models, cells from various cell lines, or any combination thereof. The balanced cell number culture according to Claim 1. Claim 10: The different cell types are selected from one or more xenograft models. The balanced cell number culture according to Claim 9. **Claim 11**: The xenograft model is the balanced cell number culture according to claim 10, derived from one or more subjects with a disease. **Claim 12**: The disease includes cancers of the head, neck, lung, skin, breast, blood, lymph, bone, soft tissue, brain, eye, reproductive system, circulatory system, digestive system, endocrine system, nervous system, or urinary system, and the balanced cell number culture according to claim 11. **Claim 13**: The balanced cell number culture is transplanted into an in vitro model system, an in vivo model system, or an ex vivo model system, and the balanced cell number culture according to claim 1. **Claim 14**: The sample of the balanced cell number culture contains about 5,000 to about 200,000 cells, and the balanced cell number culture according to claim 1. **Claim 15**: More than 500 viable cells of each cell type of the at least two or more different cell types are present in the sample, and the balanced cell number culture according to claim 1. **Claim 16**: A method for preparing a balanced cell number culture having at least two or more different cell types, comprising: (i) determining the growth rate for each cell type of the two or more different cell types; (ii) combining the two or more different cell types to create a cell pool, wherein the initial cell number of each of the two or more different cell types added to the cell pool is determined based on the growth rate of step (i); (iii) culturing the cell pool of step (ii) over a period of time to produce the balanced cell number culture, wherein a sample of 0.2% to 10% by volume of the balanced cell number culture contains at least 500 cells of each cell type of the two or more different cell types; and where the sample of step (iii) may contain 5,000 to 200,000 cells, and the method may further include excluding a cell type in step (ii) if that cell type has a growth rate of 0.2 times per day. **Claim 17**: More than 500 viable cells of each cell type of the two or more different cell types are present in the sample of step (ii), and the method according to claim 16. **Claim 18**: The method according to claim 16 or 17, wherein in step (ii), the proportion from each cell type of the two or more different cell types after being added to the cell pool is inversely proportional to the cell growth rate of that cell type determined in step (i). **Claim 19**: A method for evaluating the therapeutic efficacy of a candidate drug against individual cells of a mosaic tumor comprising at least two different cell types, comprising: (i) treating the mosaic tumor with the candidate drug over a certain duration; (ii) dissociating the cells of the treated mosaic tumor into individual cells; (iii) evaluating the individual cells to determine the phenotypic, genetic, and transcriptomic expression of the individual cells of the mosaic tumor at the end of the duration; (iv) determining the therapeutic efficacy of the candidate drug by comparing the phenotypic, genomic, and transcriptomic expression of the individual cells of the mosaic tumor with the phenotypic, genomic, and transcriptomic expression of individual cells of the same mosaic tumor not treated with the candidate drug. **Claim 20**: A method for simultaneously evaluating the effects of a candidate drug on multiple cell types in an in vivo system, an ex vivo system, or an in vitro system, comprising: (i) preparing a balanced cell number culture comprising at least two cell types, wherein the at least two cell types are different from each other; (ii) transplanting the balanced cell number culture into the in vivo system, ex vivo system, or in vitro system, wherein the in vivo system is a mammal, the ex vivo system is an organoid, and the in vitro system is a 2D or 3D in vitro model system; (iii) treating the in vivo system, ex vivo system, or in vitro system with the candidate drug over a certain duration; (iv) evaluating the at least two cell types to determine the phenotypic, genetic, and transcriptomic expression of the at least two cell types at the end of the duration. A method comprising: (v) determining the effect of the candidate agent by comparing the phenotypic, genomic, and transcriptomic expression of the at least two cell types in the in vivo system, ex vivo system, or in vitro system with the phenotypic, genomic, and transcriptomic expression of the at least two cell types in the same in vivo system, ex vivo system, or in vitro system not treated with the candidate agent. **Claim 21** A method for identifying a subpopulation of subjects sensitive to a candidate agent, comprising: (i) creating a balanced cell number culture comprising a plurality of cell types, wherein the plurality of cell types comprises cells from at least two different subjects; (ii) transplanting the balanced cell number culture into a model system; (iii) treating the model system with the candidate agent for a certain duration; (iv) evaluating the balanced cell number culture at the end of the duration to determine the phenotypic, genetic, and transcriptomic effects of the candidate agent on the individual cells of the balanced cell number culture; (v) identifying the subpopulation of subjects sensitive to the candidate agent based on the evaluation of step (iv). **Claim 22** A method for identifying a candidate agent target in a biological pathway, comprising: (i) preparing a balanced cell number culture comprising a plurality of cell types, wherein the plurality of cell types comprises cells from at least two different subjects; (ii) transplanting the balanced cell number culture into a model system; (iii) treating the model system with the candidate agent for a certain duration; (iv) evaluating the balanced cell number culture at the end of the duration to determine the phenotypic, genetic, and transcriptomic effects of the candidate agent on the individual cells of the balanced cell number culture; (v) identifying the candidate agent target in the biological pathway based on the evaluation of step (iv).