Prioritization of chemical interventions to modulate cell behavior
By employing omics-level deep learning models to associate compound effects with phenotypic changes, the method addresses the challenge of predicting phenotypic activity in drug discovery, enhancing the efficiency of phenotypic screens and improving drug discovery productivity.
Patent Information
- Application Number
- PCT/US2024/055971
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-06
- Filing Date
- 2024-11-14
- Publication Date
- 2025-05-22
AI Technical Summary
Current methods for drug discovery, particularly in phenotypic drug discovery, face challenges in accurately predicting the phenotypic activity of chemical compounds due to a lack of mechanistic understanding and the need for high-throughput screens that often focus on simple phenotypes, leading to compounds that fail to impact complex disease processes effectively.
The development of a system and method that uses omics-level deep learning models to prioritize chemical interventions by associating systems-level compound effects with changes in complex phenotypes, enabling the identification and prioritization of compounds that induce targeted changes in cellular dynamics.
This approach significantly improves the efficiency of phenotypic screens by prioritizing compounds that are likely to induce desired biological activities, thereby bridging the translational gap and enhancing the productivity of drug discovery processes.
Smart Images

Figure US2024055971_22052025_PF_FP_ABST
Abstract
Description
PRIORITIZATION OF CHEMICAL INTERVENTIONS TO MODULATE CELLBEHAVIORTECHNICAL FIELD
[0001] This specification describes, inter alia, technologies relating to prioritizing chemical interventions to induce targeted changes in cellular dynamics.CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and the benefit U.S. Provisional Patent Application No. 63 / 598,864 filed on November 14, 2023, U.S. Provisional Patent Application No. 63 / 548,517 filed on November 14, 2023, and U.S. Provisional Patent Application No. 63 / 606,924 filed on December 6, 2023, the contents of all of which are hereby incorporated by reference in their entireties.BACKGROUND
[0003] Single-cell omics technologies are increasingly contributing to understanding of biological systems, including creating atlases of whole organisms, characterizing disease cell states, and measuring perturbation response. More recently, single-cell multi omics has been combined with machine learning to learn gene regulatory networks, which can be interrogated using CRISPR / Cas9 to causally link inferred regulatory relationships to downstream developmental dynamics. However, to use single-cell omics for drug creation, chemical interventions must be linked to changes in cell state.
[0004] Despite record spending in therapeutics research and development over the past 20 years, overall clinical trial success rates have remained stagnant with the percent of Phase 1 compounds entering the clinic remaining around 10%. Despite the dominance of the target-based discovery model, retrospective analysis of successful drug programs shows that more than 65% of all approved medicines were discovered via phenotypic observations, and further focus on this paradigm might help address the decades-long productivity crisis in drug discovery. The goal of phenotypic drug discovery is to identify a disease-relevant phenotype and to identify compounds that induce a desired change in this phenotype in cells relevant to the disease process. This is challenging because the number of drug-like molecules is large, and without better mechanisticmodels describing the impact of compounds on cells there is no clear way to predict the phenotypic activity of any given compound.
[0005] To overcome this hurdle, many phenotypic discovery programs resort to bruteforce screens of millions of compounds. However, to achieve such throughput, high-throughput screens (HTS) must take a reductionist approach, focusing on simple phenotypes that can easily be assayed in scalable model systems, such as cancer cell lines, which may not accurately represent the disease process of interest. As a result, identified hit compounds often fail to have meaningful impact on the desired complex disease process in more realistic pre-clinical models. However, directly screening using a higher-resolution phenotypic assay in a disease-relevant cellular system is lower throughput and more expensive. Without tools to accurately prioritize compounds that are likely to have a desired biological activity, it is difficult to deploy more realistic phenotypic assays to discover new medicines and thereby bridge the translational gap.
[0006] In target-based discovery, virtual screening offers a path to selecting putative ligands of a target protein by predicting the binding of individual molecules to protein structures. However, generalizing this strategy to phenotypic discovery is not straightforward because there lacks a mechanistic understanding of how different molecules effect changes in most phenotypes. One potential solution is to use machine learning models to directly predict phenotypic activity based on training examples from initial brute-force screens, as has been used to discover a novel antibiotic. However, this requires the development of a purpose-built training dataset and model, neither of which can be reused to predict other phenotypes. Furthermore, this process does not provide a strategy to integrate feedback from validation experiments that generalized across phenotypes.
[0007] Given the above background, improved single-cell omics techniques for drug creation are needed to link chemical interventions to changes in cell state.SUMMARY
[0008] The present disclosure addresses the drawbacks identified above by, inter alia, providing improved systems and methods for prioritizing chemical interventions to induce targeted changes in cellular dynamics.
[0009] In one aspect, the disclosure provides a method for identifying and prioritizing chemical perturbances that induce targeted changes in the state of a cell.
[0010] In one aspect, the disclosure provides a system for identifying and prioritizing chemical perturbances that induce targeted changes in the state of a cell.
[0011] Various embodiments of systems, methods and devices within the scope of the appended claims each have several aspects, no single one of which is solely responsible for the desirable attributes described herein. Without limiting the scope of the appended claims, some prominent features are described herein. After considering this discussion, and particularly after reading the section entitled “Detailed Description” one will understand how the features of various embodiments are used.BRIEF DESCRIPTION OF DRAWINGS
[0012] FIG. 1 depicts a non-limiting example of a framework to enable phenotypic discovery using omics-level deep learning models. (1) The first step in this framework is to identify a target omics signature based on a combination of clinical data and data from a phenotypic assay. (2) To identify compounds for screening, a model trained on perturbation signatures (such as the LINCS Connectivity Map) predicts which compounds will likely induce the target signature. (3) A limited number of prioritized compounds were then experimentally screened, compounds that induce the desired phenotype were identified, and hits were validated in multiple donors. Validated hits were the output and may be used for downstream pre-clinical development. (4) Paired transcriptomic and phenotypic measurements were used to refine the signature, better understand model performance, and better understand the mechanisms by which chemical perturbations altered target phenotypes.
[0013] FIGS. 2A-2C depict a non-limiting example of schematic showing a virtual screening framework enabling performance matching compounds to target signatures. FIG. 2A depicts a schematic showing a virtual screening framework as a neural network classifier trained on a CMap L1000 signatures. All CMap level 4 signatures were first filtered using quality control filters to remove low-quality signatures. The model was trained using cross-entropy focal loss, which emphasizes loss on poorly-classified examples. FIG. 2B depicts the transcriptional structure of a non-limiting example of a 1737-sample scRNA-seq perturbation dataset profiling the impact of 88 compounds from CMap on 10 cell types in duplicate. FIG. 2C depicts a graphof experimental data comparing the virtual screening model to other methods on three different datasets, showing that the virtual screening model outperformed other methods overall. FIG. 2D depicts a graph of experimental data displaying the results of the benchmark split out by dataset and cell type.
[0014] FIGS. 3A-3G depict a non-limiting example of schematic showing a single-cell multi-omics guided phenotypic assay capturing multi-lineage hematopoietic differentiation in human primary cells. FIG. 3A depicts a schematic and graph demonstrating that primary CD34+ hematopoietic stem and progenitor cells (HSPCs) from four healthy donors and sampled singlecell RNA + surface marker measurements (CITE-seq) were taken at 5-time points over a 10-day time course. FIG. 3B and FIG. 3C depict graphs showing major cell populations of the myeloid lineage arising from HSPCs into lineage-committed cell states. FIG. 3D depicts a marker panel for each flow cytometry assay and the average expression of each surface marker in each cell type based on the CITE-seq surface marker measurements. FIG. 3E depicts a schematic showing marker gene expression in HSPC atlas cell types. FIG. 3F depicts a graph comparing cell type density across time points and donors using UMAP. FIG. 3G depicts a schematic and graphs showing the results of the phenotypic assay and controls.
[0015] FIG. 4A-4F depicts non-limiting examples of how Al-enabled virtual screens enhance discovery for complex phenotypes. FIG. 4A depicts a graph showing v-score input signature between the megakaryocytic-erythroid progenitors (MEP) and megakaryocyte progenitors (MPC) cell types. FIG. 4B depicts a graph showing megakaryocyte screening results. FIG. 4C depicts a graph showing erythroid progenitor screening results. FIG. 4D depicts a graph demonstrating megakaryocyte (Mk) hit compounds in additional donors. FIG. 4E depicts a graph demonstrating Ery hit compounds in multiple donors. FIG. 4F depicts graphs showing the results of sampling of the virtual screening model ranks covered by available compounds.
[0016] FIG. 5A-5G depicts result showing that paired phenotypic and transcriptomic screening provides closed-loop feedback into model performance. FIG. 5A shows paired transcriptomic and phenotypic screening of hit and non-hit compounds was performed over a 7- day time course. FIG. 5B shows a graph depicting cosine distance between the 24h CD34+ perturbation signature 4 and each of the 10-nearest neighbor signatures for the same compoundin LINCS. FIG. 5C shows a graph depicting the differential expression of genes known to play a role in Mk differentiation for each compound perturbation at 24h. FIG. 5D shows a graph depicting the association of gene expression with the target Mk-induction phenotype. FIG. 5E shows a graph depicting concordant genes as a refined input signature. FIG. 5F depicts a graph showing the fold-change in Mk was measured via flow cytometry for each compound and the fold-change in the abundance of the various annotated cells in the scRNA data relative to DMSO. FIG. 5G depicts a graph showing the results of variation across chemical perturbations per cell type and time point.
[0017] FIG. 6A-6G depicts the results of evaluating the mechanisms of chemically- induced megakaryopoiesis. FIG. 6A shows a schematic of DE signature of HSPCs at 24h. FIG. 6B shows a graph of trends in gene expression over differentiation. FIG. 6C shows a graph depicting the density of cells per type at various stages of pseudotime. FIG. 6D shows a graph depicting the density of cells at Day 7 per compound class. FIG. 6E shows graphs depicting expression of genes associated with GO Biological Processes per compound class across pseudotime. FIG. 6F shows a graph depicting the results of GO Term enrichment along PC I of Day 1 HSPC DE signatures. FIG. 6G shows a graph depicting the results of GO Term enrichment along PC2 of Day 1 HSPC DE signatures.DETAILED DESCRIPTION
[0018] Reference will now be made in detail to embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0019] In aspects, the disclosure provides a generalizable framework for phenotypic discovery that integrates virtual and experimental screens. In embodiments, the virtual screening strategy uses genomic signatures to associate systems-level compound effects with changes in complex phenotypes. In embodiments, the virtual screening approach of the disclosure provides a new deep learning classifier trained on publicly available datasets to prioritize compounds that induce a desired transcriptional signature. In embodiments, transcriptional signatures are used asthe readout for phenotypic screens of complex disease. In embodiments, disease-associated signatures are used to prioritize compounds for phenotypic screens.
[0020] In a non-limiting embodiment, transcription is the basis for phenotypic virtual screens because there is a link between changes in RNA expression induced by a compound and downstream complex phenotype, providing a mechanism for closed-loop feedback. In embodiments, the results from experimental validation provide refinement of a gene signature towards the omics changes necessary and sufficient to induce a change in phenotype.
[0021] In aspects, the disclosure provides systems and methods for identifying and prioritizing chemical perturbances that induce targeted changes in the state of a cell. In embodiments, the cell is a hematopoietic cell.
[0022] In embodiments, the targeted changes comprise one or more changes in a diseaserelevant phenotype. In embodiments, the one or more changes are measured and / or identified and associated with the disease-relevant phenotype using omics measurements, such as but not limited to transcriptomics. In some embodiments, the omics measurements comprise chromatin accessibility and / or proteomics. In embodiments, the disease-relevant phenotype is gene expression.
[0023] In embodiments, the one or more changes in gene expression is measured and / or calculated by determining a differential expression between each state of the cell to provide a gene expression signature. In embodiments, the gene expression signature is obtained from one or more atlases of cellular data. In embodiments, the cellular data comprises single-cell RNA sequencing data.
[0024] In embodiments, the method further comprises using a deep learning model to virtually screen a plurality of compounds and identify from the plurality of compounds one or more compounds that induces a perturbance that induces changes in the state of the cell that is the same or substantially the as the targeted changes in the state of the cell.
[0025] In embodiments, the method further comprises using the deep learning model to prioritize the one or more compounds, optionally wherein the one or more compounds are prioritized based at least on the similarity of the induced changes in the state of the cell to the targeted changes in the state of the cell.
[0026] In embodiments, the method further comprises experimentally measuring the chemical perturbance of the one or more compounds, optionally using one or more phenotypic assays. In some embodiments, the phenotypic assay comprises a phenotypic screening system such as a patient-derived organoid system and / or tissue-on-a-chip screening systems.
[0027] In embodiments, the experimentally measuring comprises measuring one or more changes in lineage commitment of the cell after treatment with the one or more compounds, optionally wherein the method comprises identifying and / or measuring surface markers associated with each change in lineage.
[0028] In embodiments, the method further comprises inputting data obtained from the experimental measuring to the deep learning model, wherein the inputting further refines the deep learning model and / or the gene expression signature.
[0029] In embodiments, the cell is selected from a megakaryocyte, an erythrocyte, an eosinophil / basophil / mast cell (EBM), a monocyte, and a neutrophil.
[0030] In aspects, the disclosure provides a system and / or framework for identifying and prioritizing chemical perturbances that induce targeted changes in the state of a cell.
[0031] In embodiments, the system and / or framework comprises:(a) desired targeted changes in the state of the cell;(b) a deep learning model; and(c) one or more phenotypic assays.
[0032] In embodiments, the targeted changes comprise one or more changes in a diseaserelevant phenotype, optionally wherein the disease-relevant phenotype is gene expression.
[0033] In embodiments, the one or more changes in gene expression is measured and / or calculated by determining a differential expression between each state of the cell to provide a gene expression signature, optionally wherein the gene expression signature is obtained from one or more atlases of cellular data, optionally wherein the cellular data comprises single-cell RNA sequencing data.
[0034] In embodiments, the deep learning model virtually screens a plurality of compounds and identify from the plurality of compounds one or more compounds that induces aperturbance that induces changes in the state of the cell that is the same or substantially the as the targeted changes in the state of the cell.
[0035] In embodiments, the deep learning model prioritizes the one or more compounds, optionally wherein the one or more compounds are prioritized based at least on the similarity of the induced changes in the state of the cell to the targeted changes in the state of the cell.
[0036] In embodiments, the one or more phenotypic assays measure the chemical perturbance of the one or more compounds.
[0037] In embodiments, the one or more phenotypic assays measure one or more changes in lineage commitment of the cell after treatment with the one or more compounds, optionally wherein the one or more phenotypic assays identify and / or measure surface markers associated with each change in lineage.
[0038] In embodiments, data obtained from the one or more phenotypic assays is inputted into the deep learning model, wherein the inputting further refines the deep learning model, optionally wherein the inputting further refines the gene expression signature.
[0039] In embodiments, the data obtained from the one or more phenotypic assays is obtained from one or more compounds exhibiting and / or inducing a chemical perturbance above a threshold value. In embodiments, the threshold value is at least about one standard deviation, about two standard deviations, about three standard deviations, about four standard deviations, about five standard deviations, about six standard deviations, about seven standard deviations, about eight standard deviations, about nine standard deviations, about ten standard deviations, or about ten or more standard deviations greater than a reference (e.g. a control) sample. In embodiments, the reference sample is a solvent. In embodiments, the solvent is dimethylsulfoxide (DMSO).
[0040] In embodiments, the systems and methods provide a novel experimental and computational framework for drug creation. In embodiments, the systems and methods of the disclosure introduce the first experimentally validated framework for prioritizing chemical interventions to induce targeted changes in cellular dynamics. In embodiments, the systems and methods of the disclosure can be used to identify previously unknown mechanisms of action of chemical perturbagens, including but not limited to small molecule compounds.
[0041] In aspects, the disclosure provides methods and systems useful for machine learning-guided drug creation, including but not limited to methods and systems for validating the performance of a model to prioritize chemical interventions to influence cell behavior. In embodiments, the method includes systematic validation of the accuracy of the model.
[0042] Plural instances may be provided for components, operations or structures described herein as a single instance. Finally, boundaries between various components, operations, and data stores are somewhat arbitrary, and particular operations are illustrated in the context of specific illustrative configurations. Other allocations of functionality are envisioned and may fall within the scope of the implementation(s). In general, structures and functionality presented as separate components in the example configurations may be implemented as a combined structure or component. Similarly, structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the implementation(s).EXAMPLES
[0043] The present technology is further illustrated by the following examples, which should not be construed as limiting in any way.Example 1 : Prioritization of Chemical Interventions using a Machine Learning Framework.
[0044] In this example, it is shown that a machine learning framework integrating transcriptional profiling of chemical perturbations and multimodal single-cell omics can prioritize chemical interventions that induce targeted changes in cell behavior. It is demonstrated that the framework outperforms previously proposed approaches using public data and a novel benchmarking dataset comprising cancer and primary cells. The framework is then applied to drive targeted changes in the development of human primary hematopoietic stem and precursor cells in vitro. Compounds that modulate the abundance of three hematopoietic lineages are identified, including compounds not previously identified as impacting hematopoiesis, and time- resolved changes gene pathways involved in the compounds mechanism of action are identified. These results provide the first experimentally validated framework for using single-cell omics to guide chemical interventions for cell fate control and in aspects along with a comprehensive benchmarking framework, and in aspects, can be useful for creating drugs that target specific cell types or states.
[0045] A machine learning model trained on 1 3M chemical transcriptional measurements from the Connectivity Map using multiple cross-validation strategies was developed. The model outperforms previously described approaches, and accurately prioritizes using public data and a novel IM-cell single-cell RNA-sequencing benchmarking dataset comprising 88 chemical perturbations in six cancer and four primary cell lines. A 300,000-cell multiomics time course of hematopoiesis was generated from four healthy human donors and transcriptional, proteomic, and gene regulatory dynamics underlying blood cell development were identified. The framework’s ability to prioritize chemical perturbations that induce targeted changes in cell state was systematically evaluated across three hematopoietic lineages. For compounds that fail to validate, gaps between the training data and the biological system to which the model is applied are shown to be responsible for the failure of the model to generalize. The mechanism-of-action for hit compounds using single-cell RNA-sequencing was validated and time-resolved transcriptional pathways responsible for phenotypic changes in cell state were identified.
[0046] This Example demonstrates, inter alia, methods useful for machine learning- guided drug creation. Methods include validating the performance of a model to prioritize chemical interventions to influence cell behavior. The approach was systematically validated, and the performance of the model was evaluated. Such systems can be improved by, for example, updating the training data, altering how predictions are made from the model, and identifying biological contexts that are particularly out-of-distribution. The results described above demonstrate, inter alia, the use of chemical interventions in combination with biology.Example 2: Complex phenotypic drug discovery enabled by omics-level deep learning models
[0047] This Example describes, inter alia, the development of a deep learning classifier that enables the nomination of compounds to induce cell phenotypes based on transcriptional signatures, which links chemistry and phenotypic activity using omics-level data. As described herein, implementation of this framework using transcriptional data achieved a 16-19X improvement in hit rate compared to brute-force screening in head-to-head experiments. Feedback from experimental assay results is also utilized to improve predictions.
[0048] FIG. 1 describes a flexible framework for virtual screening that is modular and generalizable and uses omics signatures to link chemical perturbation to phenotypes. Omics datawas widely used to measure cell state, and changes in cell state were intrinsically linked to higher order phenotypes. Furthermore, omics data provided a basis to clinically validate a phenotypic assay by ensuring both systems exhibited matched molecular characteristics and responses to perturbations. An implementation of this framework using deep learning to increase the efficiency of phenotypic screens by an order of magnitude compared to brute-force screening was used in this Example. Further, the value of paired phenotypic and transcriptomic measurements to provide closed-loop feedback into the virtual screen was demonstrated. Causes were identified for why the model worked in some cases and not others, and a strategy to refine the target signature to improve prioritization of molecules that modulate the phenotype of interest was described in this Example. Collectively, the framework enables greater productivity in phenotypic drug discovery, empowering the use of more representative and translatable phenotypes and cellular models.A deep learning virtual screen based on transcriptional signatures
[0049] The core of the phenotypic discovery framework is a deep learning model to identify compounds predicted to induce a change in cellular state linked to a clinically relevant phenotype (FIG. 1). Because transcriptomics data was widely available for various cell types and tissues in perturbational contexts and health and disease, a framework was implemented using transcriptional signatures as a proxy of cell state. A deep learning model was trained on the Connectivity Map (CMap) data set. This data set comprised mRNA perturbation signatures for 978 landmark genes following the treatment of a diversity of compounds (FIG. 2A). Quality control filters were applied to the full CMap dataset as a pre-processing step. This resulted in a training dataset comprising 425,242 transcriptional signatures associated with 9,597 small molecules, measured at multiple doses and in numerous cell lines. The trained model, a virtual screening model, is an ensemble of multi-layer perceptron (MLP) classifiers trained on replicate splits of the training dataset using focal loss. This loss function is a variant of cross-entropy loss that emphasized performance on hard-to-classify labels. It was predicted that the framework would increase the recall for compounds that may be observed in only a few samples or have subtle differential expression patterns.
[0050] FIG. 1 depicts a non-limiting example of a framework to enable phenotypic discovery using omics-level deep learning models. (1) The first step in this framework is toidentify a target omics signature based on a combination of clinical data and data from a phenotypic assay. (2) To identify compounds for screening, a model trained on perturbation signatures (including but not limited to the LINCS Connectivity Map) predicts which compounds will likely induce the target signature. (3) A limited number of prioritized compounds were then experimentally screened, compounds that induce the desired phenotype were identified, and hits were validated in multiple donors. Validated hits were the output and may be used for downstream pre-clinical development. (4) Paired transcriptomic and phenotypic measurements were used to refine the signature, better understand model performance, and better understand the mechanisms by which chemical perturbations altered target phenotypes.
[0051] To evaluate the performance of the model, it was benchmarked against four other algorithms that matched gene signatures to compounds. These model architectures included two baseline machine learning methods: a k-nearest neighbor (kNN) classifier and a logistic regression model. Additionally, two approaches were included that have been used to match query gene signatures to CMap perturbation signatures and for cell-type specific drug repurposing: Gene set enrichment analysis (GSEA; SigCom LINCS implementation) and Dr. Insight.
[0052] Using the benchmark, three independent data sets were covered (FIG. 2A - FIG. 2D). Each method was tested on the CMap Touchstone dataset, comprising 1000 compounds tested in 9 cell lines. Results showed that the virtual screening model outperformed all four algorithms, surpassing Dr. Insight by 80% and GSEA by 15% (FIG. 2C). Second, the five algorithms were compared on the sciPlex3 dataset of 188 compounds measured in three CMap cancer cell lines, where the virtual screening model again outperformed all algorithms and outperformed Dr. Insight and GSEA by 39% and 51% on average, respectively. Finally, to examine extendibility to cell contexts not well-represented in LINCS, a new scRNA-seq dataset profiling 86 compounds was generated from CMap tested in each of 6 cancer cell lines and 4 primary cell lines, resulting in 1737 datasets with a total of 1.26M cells (FIG. 2B). The cancer cell lines were present in CMap, but the primary cell lines were either absent in the dataset or only available for a few compounds. Here, it was found that the virtual screening model outperformed all algorithms, achieving an average 78% increase in recall compared to Dr. Insight and a 27% increase compared to GSEA.
[0053] FIGS. 2A-2C depict a non-limiting example of schematic showing a virtual screening framework enabling performance matching compounds to target signatures. FIG. 2A depicts a schematic showing a virtual screening framework as a neural network classifier trained on a CMap LI 000 signatures. All CMap level 4 signatures were first filtered using quality control filters to remove low-quality signatures. The model was trained using cross-entropy focal loss, which emphasizes loss on poorly-classified examples. FIG. 2B depicts the transcriptional structure of a 1737-sample scRNA-seq perturbation dataset profiling the impact of 88 compounds from CMap on 10 cell types in duplicate. FIG. 2C depicts a graph of experimental data comparing the virtual screening model to other methods on three different datasets, showing that the virtual screening model outperformed other methods overall. FIG. 2D depicts a graph of experimental data displaying the results of the benchmark split out by dataset and cell type.Developing a complex phenotypic assay with a high clinical translatability
[0054] To demonstrate the potential of the framework to identify phenotypically active compounds, the framework was applied to 2 different target phenotypes in human hematopoiesis. Hematopoiesis is an essential developmental process, and aberrant hematopoiesis can lead to numerous proliferative disorders and cytopenias. In particular, the Example focused on screens on human CD34+ hematopoietic stem and progenitor cells (HSPCs) because of their high clinical translatability. Hematopoietic stem cell transplantation treats various blood cancers and other hematological disorders including severe anemias. In addition, HSPCs can be used as a model system to study hematological disorders, including rare disease. The hematopoietic system is also an attractive choice due to an abundance of public human scRNA data from healthy and diseased individuals that can be used to identify target gene signatures associated with various hematopoietic processes.
[0055] As phenotypic targets, the framework aimed to modulate the differentiation of the megakaryocyte and erythroid lineages. To characterize the cell states involved with this process, a CITE-seq36 dataset that was previously generated was analyzed for an international data science competition in 2022. This joint single-cell RNA + surface protein CITE-seq dataset profded primary HSPCs from 4 healthy donors sampled at 5 time points over a 10-day time course (FIG. 3A). The differentiation of major lineages of the myeloid lineage at varying statesat multiple time points was measured, enabling the comprehensive capture of the maturation process. Progenitor and early lineage-committed cell states were identified, including cells at a range of stages of differentiation along the megakaryocyte (Mk), erythroid (Ery), eosinophil / basophil / mast (EBM), monocyte (Mono), and neutrophil (Neu) lineage trajectories (FIG. 3A and FIG. 3E), while observing consistency in cellular differentiation across all four donors (FIG. 3F).
[0056] To design a phenotypic assay to measure changes in lineage differentiation, the joint scRNA and surface marker dataset was combined with literature knowledge to identify surface markers that identify each lineage. To calibrate the phenotypic assay and reference single-cell dataset, it was confirmed that RNA-defined cell types expressed surface markers consistent with the assay for both lineages (FIGs. 3B-3D). For each lineage, positive control compounds were identified and established the assay's dynamic range to facilitate the identification of phenotypically active compounds (FIG. 3G).
[0057] To establish a hit threshold for each of the two cell-type assays, compounds that lead to low cell viability or compounds with an insufficient number of cells were filtered out. A significance cutoff relative to DMSO treatment was then calculated, while also considering the variation of DMSO samples within and across plates. Perturbations that induced the target population abundance at 6 standard deviations above DMSO were considered as hits.
[0058] FIGS. 3A-3G depict a non-limiting example of schematic showing a single-cell multi -omics guided phenotypic assay capturing multi-lineage hematopoietic differentiation in human primary cells. The results shown in FIG. 3A demonstrated that primary CD34+ hematopoietic stem and progenitor cells (HSPCs) from 4 healthy donors and sampled single-cell RNA + surface marker measurements (CITE-seq) were taken at 5-time points over a 10-day time course. In accordance to culture conditions, major cell populations of the myeloid lineage arose from HSPCs into lineage-committed cell states. To calibrate a phenotypic assay measuring lineage abundance to our omics dataset, the expression of protein surface markers for each lineage in conjunction with RNA-defined cell types was examined. In FIG. 3D, the marker panel for each flow cytometry assay and the average expression of each surface marker in each cell type were based on the CITE-seq surface marker measurements. FIG. 3E depicts marker gene expression in HSPC atlas cell types. A dotplot showing the normalized expression of markergenes in each annotated cell state (y-axis). Genes are organized based on prior knowledge associated each gene with distinct cell states (x-axis). FIG. 3F demonstrates consistent differentiation across time and donors. Comparison of cell type density across time points and donors using UMAP. No batch correction was performed on this dataset, aside from the selection of highly variable genes that were consistent across at least 2 donors. The UMAP embedding was calculated once, and then subsets were plotted on each subplot. FIG. 3G depicts phenotypic assay schematic and controls. To validate the reproducibility of the assay, the abundance of each lineage in negative and positive control conditions were measured. A cartoon schematic of the phenotypic assay, in which cells are dosed with compounds on days 0, 2, and 5 is shown in FIG. 3G. Flow cytometry readout was measured at Day 7 post-treatment. As depicted in FIG. 3G, the abundance of each target population in replicate samples of DMSO Vehicle negative control and under positive controls of a representative plate from the screen. For the Mk assay, angiogenesis inhibitor (BRD-K08502430) was the positive control. For the Ery assay, CTL 1 was sirolimus (BRD-K89626439) at lOpM and CTL 2 was erythropoietin (EPO) at 2.5 U / mL.Deep learning enabled phenotypic screening to induce megakaryopoiesis
[0059] To nominate compounds for screening in each lineage, cell state transitions associated with early differentiation into megakaryocytes and erythrocyte progenitors were identified. Those cell state transition were used as input for the model to prioritize compounds. To obtain a signature associated with each transition, a statistic was derived with similar properties to the Z-score, which was the representation of the CMap Level 4 signatures on which the model was trained. The Z-score was used in CMap to estimate the effect size of differential expression in the LI 000 assay in standard deviation units across the other 96 samples on the same plate. To calculate the effect size of the gene expression differences between all cells in two different cell populations, a v-score was computed, an estimate of the difference in the logcounts mean of two distributions in units of the variance of each. Top-ranked compounds were selected from the model’s output to assess their ability to induce the phenotype of interest. For this study, an inventory of 1,635 compounds was relied on from the CMap training set available at the time of study initiation.
[0060] To generate predictions for the megakaryocyte lineage, the framework focused on bipotential megakaryocyte erythroid progenitor (MEP), the earliest cell state associated withlineage decision giving rise to the erythroid (Ery) and megakaryocyte (Mk) lineages. While not being bound by theory, it was reasoned that this was the ideal point to intervene in differentiation because transcriptional and metabolic changes in these cells were associated with commitment to differentiation into either lineage. The framework aimed to alter the MEP cells to adopt a transcriptional state similar to the MPC population, which are progenitors committed to differentiating towards the Mk lineage. V-scores were calculated from the MEP to MPC population and these v-scores were used as input to the virtual screening model to obtain a prioritized list of compounds for screening. It was confirmed that the 1,635 compounds in the inventory were a representative subset of all compounds ranked by the model (FIG. 4F). FIG. 4F depicts graphs showing the results of sampling of virtual screening model r ranks covered by available compounds. A representative sampling of compounds across ranks for both sets of virtual screens was observed. The x-axis shows the rank output of the virtual screening model. The y-axis shows the cumulative number of compounds at that rank or lower out of 1650 compounds available in the inventory at the time of study initiation.
[0061] To experimentally determine which compounds induced the target phenotype, d CD34+ cells were treated with each model-nominated compound under HSPC maintenance conditions (CCIOO / TPO cytokines). On day 7, induction of CD41a+ CD71- CD42b+ Mk population by was evaluated by flow cytometry. 107 compounds with a rank less than 1000 prioritized by the model was tested, and to compare to the brute-force approach, a random selection of 96 compounds were also tested from the same compound inventory.
[0062] Among the 107 highly ranked virtual screening model -nominated compounds, 21 compounds were identified above the 6 standard deviations hit the threshold, with 2 highly active compounds inducing more than a 4-fold increase in Mk progenitors (FIG. 4B), resulting in a 19.6% hit rate. By contrast, only 1 compound was identified from the random selection that passed the hit threshold, resulting in a 1.1% hit rate. These results highlight that the deep learning model enriched the selection of compounds that modulate cell state transitions of interest by over 1,900%.
[0063] To confirm these hits validated in multiple donors, 17 DR-nominated hit compounds were re-tested in 2 additional donors at the dose where maximal induction of the Mk lineage was observed. While 2 compounds did not pass the viability or cell count criteria, 13 outof the remaining 15 hit compounds validated in at least 1 additional donor, demonstrated the robustness in the assay and biological translation of the chemical perturbations across different donors (FIG. 4D). While not wishing to be bound by theory, generally, it was observed that higher phenotypic response in the initial screen was associated with a higher chance of replication in the follow-up validation, suggesting the quantitative nature of the readout.Deep-learning enabled phenotypic screening to induce erythropoiesis
[0064] The flexibility of the framework was demonstrated by aiming to bias the MEP population towards an alternative fate decision: erythroid progenitor cells. Similar to previous transitions, v-scores were calculated between the MEP and Ery erythroid progenitor population to derive an input signature for the virtual screening model. The top 96 compounds were obtained from the virtual screening model-nominated compounds list and a new set of 96 random compounds for screening.
[0065] To experimentally determine which compounds induced Ery progenitors, CD34+ cells were treated with each of the model-nominated compounds and with each of the randomly selected compounds at two doses (IpM and lOpM), dropping the lOOnM dose because few compounds were maximally active at that level. Donor-derived CD34+ HSPCs were cultured and treated with each compound over 7 days and measured Ery lineage abundance using flow cytometry.
[0066] In the virtual screening model-nominated compound screen, after removing samples failing the quality control fdter, it was observed that 13 out of 81 compounds passed the 6-standard deviation above the DMSO hit cutoff, representing a 16% hit rate. In the randomly selected compound set, it was observed that only 1 out of 85 compounds induced Ery above the cutoff, representing a 1.2% hit rate (FIG. 4C). Again, the transcriptomics-based compound prioritization significantly increased the success rate in inducing the desired phenotype (-1600%). Identified hits were evaluated for performance across additional donors as part of the cross-donor validation. Out of 10 compounds passing the quality control filter in the validation experiment, 7 significantly increased Ery frequencies in both donors, and 1 more did so in one of the two (FIG. 4E). This reinforces the capacity of the machine learning model to increase the phenotypic hit rate across multiple experimental settings.
[0067] FIGs. 4A-4F depicts Al-enabled virtual screens enhance discovery for complex phenotypes. Applying the virtual screening model for the virtual screening of Mk-inducing compounds, a v-score input signature between the MEP and MPC cell types were first identified (FIG. 4A). A ranked list of compounds ordered by the probability that each compound matches the input signature was obtained. Comparing 107 available compounds with DR rank <1,000 vs 87 randomly selected compounds passing QC, there was a 19-fold increase in hit rate using the virtual screening algorithm (FIG. 4B). The black dashed line represents the hit significance cutoff for Mk. Applying this same strategy to identify compounds that induced Ery progenitor cells, a 16-fold increase in hit rate was observed using DR to virtually screen compounds (FIG. 4C). FIG. 4D depicts validating Mk hit compounds in additional donors. All but one compound passed the quality control filter induced Mk in at least 2 donors. FIG. 4E depicts validating Ery hit compounds in multiple donors. All but two compounds passed the quality control filters validated in at least 1 donor. FIG. 4F depicts graphs showing the results of sampling of virtual screening model ranks covered by available compounds. A representative sampling of compounds across ranks for both sets of virtual screens was observed. The x-axis shows the rank output of the virtual screening model. The y-axis shows the cumulative number of compounds at that rank or lower out of 1650 compounds available in the inventory at the time of study initiation.Closed-loop signature refinement in the megakaryocyte lineage
[0068] A significant advantage of the phenotypic discovery framework is that transcriptional and phenotypic measurements can be paired to better understand the underlying mechanisms governing the target phenotypic response. This enables refinement of the input target signature based on transcriptional differences between hits and non-hits and to learn the gene regulatory mechanisms by which changes in gene expression following compound perturbation lead to downstream changes in phenotype. To explore the utility of this approach, a scRNA-seq time course was performed on 12 hits, 8 non-hits, a DMSO negative control, and positive control compound with samples collected in duplicate at days 0, 1, 2, 5, and 7 for a total of 192 scRNA datasets. The non-hits were included to explore why the model prioritized these compounds but did not induce target phenotype. Paired phenotypic data on day 7 were also collected from the samples used for scRNA-seq. Across the scRNA samples, 145,157 cells with a median of 754 cells per condition were recovered. This scRNA dataset was integrated with theoriginal time course using Harmony to facilitate comparison with the original time course. In the transcriptional time course, cells from all major cell types were observed (FIG. 5A). A strong correlation between the abundance of Mk cells as determined by the scRNA-seq and phenotypic measurements was also observed (FIG. 5F).
[0069] Using transcriptional validation, it was observed that not all prioritized compounds from the virtual screening model were phenotypically active. While not wishing to be bound by theory, one possible cause was that the compounds have a cell-type specific effect in CD34+ cells that differs from the impact measured in the CMap dataset. 43% of compounds in CMap were previously reported to exhibit cell-type specific effects in the cancer cell lines, and CD34+ HSPCs were absent from the training data. To test this explicitly, each compound was calculated for the distance between the 24h signatures in the follow-up experiment and the 10 most similar signatures for the same compound in LINCS, providing an unbiased estimate of how similar the observed perturbation signatures in CD34+ cells were to signatures for the same compounds in CMap. On average, CD34+ signatures of non-hit compounds were 11% further from their closest neighbors in CMap compared with hit compounds (FIG. 5B, p=0.037, paired t-test). Although this distance can only be calculated using L1000 landmark genes, cell-type- specific perturbation effects span all genes based on ILD1 data (FIG. 5G). Without being bound to a particular theory, this suggests that developing strategies to map transcriptional signatures associated with compounds into new cell types and across all genes will likely improve virtual screening performance.
[0070] Next, the aim was to refine the signature to identify the gene expression patterns sufficient to induce a change in phenotype. It was hypothesized that only a subset of the initial input signature was required to effect the change in phenotype, with the remainder comprising passenger genes, noise, or potentially inhibitory feedback gene expression signals. To test this hypothesis, DE analysis was performed on the 24-hour pseudobulk gene expression counts for each cell type. Limma40 was used to calculate differential expression enabling the modeling of experimental covariates such as library and plate. However, instead of using the compound perturbation label as the explanatory variable, the 7-day fold-change variable was used as the basis for DE (see below “Methods” section). Because the fold-change was a scalar variable, this implementation identified genes that were linearly associated with the induction of the Mk lineage while controlling for various technical confounding variables. 672 landmark genes wereobserved that were significantly associated with Mk induction (adj. p<0.01 ), including genes previously implicated in Mk maturation, like FLU, GATA2, and NFE2 (FIG. 5C).
[0071] The signed -lo lO(p-value) was compared to the original input v-score for each gene. Three patterns were observed: concordantly associated genes (n=366), inversely associated (n=312), and unassociated with the original input v-score (Figure 5D). The concordantly associated genes were tested for whether it could be used as a refined input to the model. When filtering the input v-scores to include only genes concordant between the target transition and transcriptional validation experiment, the median rank of hit compounds improved significantly, from 1,060 to 375. To understand the significance of this shift, 10,000 random gene sets were generated with the same size as the concordant gene set as a background distribution. The median hit rank was measured after filtering to each random set. Concordant genes performed significantly better than random gene sets by median hit rank (p=0.0026). These results confirmed the hypothesis that only part of the original input signature was necessary to prioritize hit compounds and offered a proof-of-concept strategy to identify an essential component of the signature.
[0072] FIGs. 5A-5G depict results showing that paired phenotypic and transcriptomic screening provides closed-loop feedback into model performance. Paired transcriptomic and phenotypic screening of hit and non-hit compounds was performed over a 7-day time course (FIG. 5A). After integrating the data with the reference atlas using Harmony, the presence of all major cell types was observed. The cosine distance between the 24h CD34+ perturbation signature 4 and each of the 10-nearest neighbor signatures for the same compound in LINCS were shown for each compound (FIG. 5B). The differential expression of genes known to play a role in Mk differentiation for each compound perturbation at 24h was calculated (FIG. 5C). The input v-score for each gene was shown. Comparing the association of gene expression with the target Mk-induction phenotype, the data showed that some genes were concordant between the input signature and phenotype DE signature, and others were discordant (FIG. 5D). Using only the concordant genes as a refined input signature, the prioritization of hits vs non-hits across cross-validation folds can be consistently improved (FIG. 5E). FIG. 5F depicts results showing the abundance of cell types in scRNA and flow cytometry following 7 days of chemical perturbation. Differential abundance in scRNA-defined populations aligns with phenotypic assay and confirms lineage-specific induction of the Mk population. On the left of FIG. 5F, the fold-change in Mk was measured via flow cytometry for each compound. On of right of FIG. 5F, the fold-change in the abundance of the various annotated cells in the scRNA data relative to DMSO is displayed. Asterisks denote significance from a permutation test with FDR correction using the Benjamini -Hochberg procedure (adj p val < 0.05). FIG. 5G depicts the results of variation across chemical perturbations per cell type and time point. Cell type-specific variation in differential expression across cell types and time points. Pseudobulked gene expression was used as input to LIMMA to calculate differential expression per cell type and DMSO at each time point. Markers denote post-hoc annotated compound classes.Characterizing mechanisms of chemically-induced megakaryocyte maturation
[0073] In addition to refining the input signature to improve hit prioritization, paired transcriptional and phenotypic data also enabled understanding of how hit compounds induced a change in phenotype. Seeking to understand the variation across compounds better, cell-type specific differential expression analysis relative to DMSO was performed for each condition and time point (FIG. 5G). Focusing on HSPCs at 24h post perturbation, 5 major clusters of perturbations were observed: an inactive cluster that does not induce Mk, a single compound that inhibited Mk differentiation, one cluster of highly active Mk differentiation compounds, and two clusters of moderately active compounds separated along PC2 (FIG. 6A).
[0074] To better understand what drives this variation, but not wishing to be bound by theory, gene set enrichment analysis was performed using GSEAPy looking for biological processes associated with these principal components (FIGs. 6F - 6G). Genes with high PCI loadings were enriched for antigen processing and JAK / STAT signaling, consistent with known stages of megakaryopoiesis. In contrast, genes with high PC2 loadings were enriched for lipid and cholesterol biosynthesis, which has recently been recognized as an important process in megakaryocyte maturation and platelet formation.
[0075] To further understand the trends in chemically-induced megakaryopoiesis, a pseudotime trajectory in the Mk lineage for all cells in the transcriptional validation experiment was identified (FIG. 6D) Induction of Mk differentiation and Mk development gene sets was consistent across pseudotime for all classes of compounds (FIG. 6E). While not wishing to be bound by theory, this suggests a single program of Mk differentiation regardless of chemical perturbation. However, it was observed that positive regulators of Mk differentiation and cellcycling genes, known to play a role in Mk differentiation, were differentially induced by strongly active compounds, especially in MEP and MPC cells. While not wishing to be bound by theory, this suggested that the upregulation of these genes was important for the activity of hit compounds. Finally, it was observed that the compounds that modulate lipid metabolism do this at all time points and in cell types in a pattern consistent with other compounds’ activity but stronger (FIG. 6E).
[0076] FIG. 6A-6G depicts the results of unraveling the mechanisms of chemically- induced megakaryopoiesis. Examining the DE signature of HSPCs at 24h, a strong separation between active and inactive compounds and one group of compounds that separate along PC2 was observed (FIG. 6A). To understand the trends in gene expression over differentiation, a pseudotime axis through the dataset was identified (FIG. 6B). FIG. 6C depicts the density of cells per type at various stages of pseudotime. FIG. 6D depicts the density of cells at Day 7 per compound class. FIG. 6E depicts expression of genes associated with GO Biological Processes per compound class across pseudotime. FIGs. 6F-6G depict the results of GO Term enrichment along PC I and PC2 of Day 1 HSPC DE signatures. Gene sets strongly associated with PC I loading were enriched for antigen presentation and JAK / STAT signaling pathways associated with Mk induction, further supporting the conclusion that hit compounds induced bona fide megakaryopoiesis (FIG. 6F). PC2 genes were enriched for lipid and cholesterol biosynthesis, leading to labeling the three compounds separated by PC2 as lipid-inducing compounds (FIG.6G)MethodsVirtual screening model algorithm overview: Model Architecture
[0077] The virtual screening model classifier was an ensemble of three fully-connected neural networks implemented in PyTorch. Each network has two hidden layers with the same structure. The input layer has 978 nodes (one for each landmark gene), and the output layer has 9597 nodes (one for each target LINCS perturbation). The first hidden layer has 1024 nodes, and the second has 2048 nodes using rectified linear units to compute node activations.
[0078] To generate predictions, the data was split into three folds based on replicate labels from CMap and this model architecture was trained independently on each fold. The three models were then ensembled for inference. The final predicted class probabilities were thesoftmax probabilities of the average score over all three folds. To compute final ranks, the average score was ranked across all three models. Higher scores were ranked lower (i.e., closer to 0).Curating LINCS CMap into a training dataset
[0079] The starting point for training was the LINCS level 4 dataset, which was obtained from https: / / lincsportal.ccs.miami.edu / datasets / view / LDS-1611. The level 4 dataset contained differential expression z-scores for each compound against all values measured on the same platen. Observations according to quality control criteria were filtered out. The criteria were as follows:1. Removed any diversity-oriented synthesis (DOS) compounds.2. Removed any compounds with fewer than 5 observations in total.3. For each compound, removed any observations with a cosine similarity <0.12 to the closest replicate.4. For each compound, selected the most frequently recorded dose between IpM and 20pM.5. Kept only measurements recorded at 6-hour or 24-hour post-treatment.6. After applying the first four filters, removed any compounds measured in fewer than 5 cell lines, more than 40 cell lines, or with fewer than 3 replicates.
[0080] The following chemical filters were then applied:1. Molecular weight must be between 60 and 1000 (inclusive)2. No more than 1 covalent motif (defined by SMARTS60)3. No more than 9 NIBR structure flags614. Pass BRENK critieria625. Must not match 30 SMARTS patterns (exact patterns not disclosed).
[0081] Applying these filters, 425,242 observations were retained comprising 9,597 small molecules measured in 52 cell lines, with a median of 32 observations per compound and 751 compounds measured more than 100 times.Model inputs
[0082] The input to virtual screening model was a representation of a desired transition between two cellular states. Since the model was trained on LINCS, the same principles were used to construct the input representation as LINCS does for perturbation effects: gene log foldchanges were represented in units of standard deviations of log expression. Whereas in LINCS, these standard deviations were obtained from other wells on the same plate, single-cell data enabled the measurement of gene standard deviations from other cells from the same state.
[0083] Using these considerations, the v-score was defined for a gene measured in two states as follows: v-score (where x and y are the transcript counts in each state, normalized to a fixed total count per cell. A pseudocount of 1 was added to each logarithm to avoid the singularity at 0.
[0084] Unlike the t-score, the v-score does not depend on the number of cells in each group in expectation. As in LINCS level 4, differences were measured in units of standard deviation. A high v-score was obtained when genes have different mean log expression between the two states but relatively low standard deviation of log expression within them.Training regime
[0085] The training data was divided randomly into three folds, with perturbation replicates balanced across the folds. Models were independently trained on two of three folds, and the resulting class scores were averaged to give the final class ranks.
[0086] The models were trained using a focal loss function on the softmax probabilities, with a focusing parameter of y=2. To ensure robustness of the trained models, each hidden layer randomly zeroed some of its inputs with a fixed dropout probability of 0.64. Batch normalization was applied during model training using momentum = 0.1. The learning rate was determined by a cosine annealing schedule with warm restarts with 20 epochs before the first restart, an initial learning rate of 0.0139, and a minimum learning rate of 0.00001. Each model was trained for a total of 50 epochs.
[0087] The above dropout probability, initial learning rate, and time to first warm restart were determined via hyperparameter optimization using Optuna. To evaluate a particular set of hyperparameters, the parameterized model was trained on two of the three folds, and measured recall of the held-out compound signatures in the third fold. The hyperparameters with highest average recall were used in training (Table 1).Table 1: Tested and selected hyperparametersModel BenchmarkingCurating the CMap Touchstone Dataset
[0088] The curated CMap level 4 training data was filtered to the 9 cell lines of the CMap touchstone dataset: A375, A549, HA1E, HCC515, HEPG2, HT29, MCF7, PC3, and VCAP. 1000 compounds that have samples in all 9 cell lines based on number of observations per cell line were then selected. For some compounds, there was a long tail in some cell lines, so only the first 30 observations were considered per compound per cell line. The mean number of observations were calculated per cell line and the top 1000 compounds were taken. For this set of 1000 compounds, uniform random was subsampled from each of the 9000 combinations of cell line and compounds to generate a dataset of 9000 samples.
[0089] For this dataset, ensemble models were benchmarked and KNN differently to ensure that test and train data were not mixed.
[0090] Virtual screening model and softmax regression were ensembled, each containing three models. Each model has a unique test fold in the curated CMap dataset, and was trained on the remaining two data folds. When ensemble benchmarks were modeled for this dataset, individual predictions absent in the data folds for each model was also modeled. Then, the recall per compound by cell line was computed. If the rank of the query compound was in the top 1% of the output of the compound from the model, the recall was 1, else 0. The recall across compounds for each cell line was averaged and the recall across cell lines for each model was averaged for final ensemble performance.
[0091] To benchmark k-nearest neighbors, each of the three folds were tested over and the reference dataset to the other two folds were set. The recall per compound by cell line was computed, and then the recall per model was averaged for final ensemble performance.
[0092] For public methods: SigCom LINCS and Dr. Insight, 500 observations were uniformly subsampled due to long runtimes.Curating the sciPlex3 Dataset
[0093] A public GEO (Gene Expression Omnibus) dataset GSE139944 was curated. This dataset tested small molecule inhibitors on A549, MCF7 and K562 cells. The pre-processed version of the dataset was downloaded and v-scores as described above were calculated between each treatment condition and DMSO. The v-scores were used as input to each model.Generating and curating the Intervention Library DatasetHuman cell lines
[0094] A375, A549, HepG2, PC3, and HEK293T cell lines were purchased. The human- embryonic kidney cell line, HA1E, was also used. Cells were cultured in either RPMI or DMEM with fetal bovine serum as per suppliers’ recommendations. Cells were seeded into 24-well dishes and incubated for 24 hours at 37 °C and then treated with compounds for 24 hours at a dose previously determined to be the maxi mum -tolerated dose for the six cell lines. Cells were harvested with 0.05% Trypsin and collected with serum-containing media.Human bronchial epithelial cells
[0095] Normal human bronchial epithelial cells (HBEC) were obtained from Lonza from two separate healthy donors. Cells were thawed and grown in PneumaCult media for three days T 37°C, then plated in 24well dishes for two days. Compounds were then added to cells at a 2x concentration in PneumaCult and incubated at 37°C for 24 hours. Cells were harvested with ACF Enzymatic Dissociation Solution for 7min according to manufacturer’s protocol and collected with media.Human CD8+cytotoxic T cells
[0096] Peripheral blood mononuclear cells (PBMC) from two healthy donors were isolated from leukopaks using MACS Cell Separation kits and frozen at le8 cells per vial. Cellswere thawed and cells were isolated using the CD8+T Cell Isolation Kit. T cells were grown in RPMI supplemented with FBS and IL-2; CD3 / CD28 Dynabeads were added to activate cells for 72 hours at 37°C. Beads were removed from culture by incubating on a magnet for 5 minutes.Cells were resuspended in media containing fresh IL-2 and plated in 96-well plates. Compounds were added at 2x concentration and incubated at 37°C for 24 hours. Cells were harvested by centrifuging for 5 minutes at 300xG.Human CD34+hematopoietic stem cells
[0097] Mobilized human peripheral blood CD34+cells from healthy donors were obtained. Cells were thawed over PBS supplemented with 1% human serum albumin and incubated in StemSpan media containing CC100 and rhTPO at 37°C for 48 hours. Cells were resuspended in fresh media containing CC100 and rhTPO and plated in 96-well plates. Compounds were added at 2x concentration and incubated at 37°C for 24 hours. Cells were harvested by centrifuging for 5 minutes at 300xG.Human preadipocytes
[0098] Preadipocytes from healthy, lean donors were thawed into 24-well plates in plating media. After incubation at 37°C for 24 hours, the media was changed to differentiation media and cells were incubated for 72 hours. Compounds were added to cells at a 2x concentration and then incubated at 37°C for 24 hours. Cells were harvested with 0.05% Trypsin and collected with serum-containing media.Single-cell library generation
[0099] Harvested cells were washed with PBS and labeled with TotalSeq-B hashtag antibodies according to manufacturer’s protocol. Briefly, cells were incubated with 250 ng of TotalSeq-B antibodies (hashtags 1-10) for 30 minutes at room temperature. Cells were washed with PBS a total of three times and then ten samples with different hashtags were pooled and counted on a Luna Cell Counter. Single-cell libraries were then prepared using the Chromium Single Cell 3’ Feature Barcoding Kit targeting 10,000 cells per library, according to manufacturer’s protocols (lOx Genomics; CG000317 Rev B).Single-cell data preprocessing
[0100] Hashed sequencing libraries were filtered to remove cells with too few or too many counts. Specifically, each cell was assigned a score of log(cell library size) - log(mean library size per cell), and cells with a score less than -0.5 or greater than 0.75 were removed. Cells with greater than 18% mitochondrial gene counts were also filtered out. Genes were removed if they were not expressed in at least 0.5% of cells for each plate of data. Finally, raw gene counts were normalized using scanpy.pp.normalize total with the target parameter set to le6, and rescaled using scanpy.loglp. The hashed libraries were then demultiplexed using a multivariate Gaussian mixture model.
[0101] Hashed sequencing libraries were filtered to include libraries ranging from 5000- 80000 counts, with less than 18% mitochondrial gene counts. k-Nearest Neighbors (kNN) implementation
[0102] To construct the model, features were balanced for each compound signature in curated LINCS training dataset. V-scores were scaled to target a standard deviation of 1, and then clip values outside of [-2,2],
[0103] Ski earn’ s pairwise cosine similarity was used to make predictions to compare two datasets: the reference LINCS training dataset as input X and a benchmark test dataset as input Y. The output would contain all pairwise similarities between X and Y. Next, similarities were grouped by compound, as each compound had many signatures in the training dataset. For each observation in Y, test dataset, all similarities with X by compound were grouped. The average was then taken for each group to get a mean similarity for each unique compound. To interpret the similarity results, each compound was ranked by mean similarity for each test observation. Values of cosine similarity range from [-1,1], where values increase as similarity increased. To see if the framework successfully matched a compound label to a test observation, the label was checked to see if it was within the lowest 100 ranked, most similar, compounds.Multinomial Logistic (Softmax) Regression Model
[0104] To construct the model, the curated LINCS training dataset was partitioned into three folds. These folds matched the virtual screening model’s training folds. Each of three sklearn multiclass logistic regression models were trained on a unique combination two of three folds. All models had the same hyperparameters: a regularization penalty of L2, an inverse regularization strength of 1, no class weights, and a limited memory BFGS solver.
[0105] To make predictions for a benchmark observation, each model computes probability estimates of compound classes. The resulting classes were ranked by probability, where lower rank indicates higher probability. Each model then contributed a vote to the ensemble rank; the mean rank was taken across all three models. To finalize the ensemble rank, the mean rank was ranked. An observation was predicted successfully if the compound label was in the lowest 100 rank, highest probability, compounds.SigCom LINCS
[0106] Public method SigCom LINCS is accessible by LoopBack API. It is hosted by the Ma’ayn Laboratory at maayanlab.cloud / sigcom-lincs / .
[0107] To run predictions, all relevant compound signatures were identified in the SigCom Database. Relevant signatures have a clearly identifiable compound that was predictable by the virtual screening model and baseline algorithms. Compound identifiers were maintained by the LINCS consortium. They started with a “BRD-” followed by 9 alphanumeric characters. These identifiers were parsed from the “cmap id” signature metadata field.
[0108] Next, the benchmark data was prepared for signature search. For each observation, genes were sorted by v-score. The highest and lowest 250 gene values were passed into up and down entities of the API “ranktwosided” enrichment query. The server was requested for maximum limit of 10,000 up and down chemical perturbation signatures. The server returned the score, z-sum, and ranked of the top 10,000 mimicker and reverser signatures.
[0109] To convert the signatures scores into compound ranks, the signature was selected with the maximum z-sum of each relevant compound. The compounds were then ranked from low to high with increasing z-sum, increasing similarity. A benchmark observation was predicted successfully if the compound label of the observation is within the top 91 ranks. Note the threshold was at 91, because there were only 8,701 relevant compound classes in SigCom from 9,597 classes in the baseline algorithms.
[0110] Dr. Insight
[0111] An archived version of the software can be obtained from cran.r-project.org / src / contrib / Archive / Drlnsight / Drlnsight 0.1.2.tar.gz.
[0112] This archival version matched signatures to an early CMap dataset comprising 6100 signatures of 1309 compounds at varying concentrations on three cell lines: MCF7, PC3, and HL60. To run Dr. Insight, the following parameters were used; Repurposing unit = “drug”, connectivity = “positive”, and the CEG.threshold to 0.05. Because the Dr. Insight reference dataset only includes 1309 compounds, results were only reported for the intersection in each dataset and the Dr. Insight reference.Generating a single-cell time course of hematopoiesisGenerating CITE-seq data in human CD34+ cells
[0113] Mobilized (Neupogen) peripheral blood CD34+ cells (mPB CD34+) were purchased. mPB34+ cryopreserved cells from four healthy donors were thawed and cultured in StemSpan SFEM supplemented with CC100 and TPO (lOOng / ml) at a density of 300K / ml. Cells were incubated at 37 °C over a period of 12 days with media changes every 2-3 days. Cell collections were done across five time points over a ten-day period (Day 2, 3, 4, 7, and 10). On days of collection, cells were process for CITE staining using Biolend TotalSeq antibody cocktail protocol with minor modifications. Labeled cells were process for single cell RNA sequencing using the lOx Genomics Single Cell Gene Expression with Feature Barcoding technology. Libraries were prepared using the Chromium Single Cell 3’ Reagent Kit v3.1, and sequenced on an Illumina NovaSeq 6000 platform, generating paired-end reads. Raw reads were demultiplexed using bcl2fastq (v2.20.0.422) and processed using Cell Ranger software (v5.0.1, lOx Genomics). Reads were aligned to the human reference genome (GRCh38) using STAR aligner (v2.7.0a).Developing a plate-based flow cytometry assay to measure multiple lineages in CD34+ differentiationHematopoietic Stem Cell in vitro differentiation assay
[0114] Dual Mobilized (Neupogen and Mozobil) peripheral blood CD34+ cells were purchased. Cryopreserved cells were thawed and cultured in flasks in StemSpan SFEM with TPO (lOOng / ml) and lx CC100 supplement. On day 0 (48 hours post-thaw), cells were plated into 96-well plates in the same medium. Plating conditions were optimized for each lineage as follows: megakaryocyte and basophil lineage differentiation 60K / well in round bottom plates,erythroid differentiation 30K cells / well in flat bottom plates. Compound treatment was performed on days 0, 2, and 5 of culture. Cells were passaged at a ratio 1 :4 on day 2, media was refreshed on day 5. On day 7, immunophenotype of differentiated cells was evaluated using flow cytometry. Compound treatment and media changes were performed using Integra Viaflo384.Compound treatment
[0115] Compounds were purchased at lOmM in DMSO and arrayed onto microplates at 0.1, 1, or lOmM in triplicate using Hamilton Microlab Star liquid handler and stored at -80°C. On day of treatment, compound plates were thawed at 37°C for 10 min, diluted with IMDM and then added to cells using Integra Viaflo384.Flow cytometry
[0116] On day 7 of differentiation, cultures were washed and incubated with antibodies in Cell Staining Buffer, washed with DPBS, and incubated in DPBS containing viability dye (1 : 1000), followed by wash with Cell Staining buffer and fixation with Cytofix Fixation buffer. All incubations were performed for 25 min at 4°C in the dark. Cells were then washed and resuspended in Cell Staining Buffer, and analyzed on NovoCyte Quanteon flow cytometer.
[0117] Channel compensations were performed using single stained UltraComp beads or cells. Titrations were performed to assess optimal antibody concentration. Flow cytometry data were analyzed using FlowJo. The following antibody panels were used to define cell populations. Megakaryocytes: CD41a+ CD71- CD42b+, Early erythroid progenitors: CD41a- CD71+ CD36+ CD235a-, late erythroid progenitors: CD41a- CD71+ CD36+ CD235a+.Phenotypic data analysis and hit criteria
[0118] For each set of screening experiments targeting a lineage, this analysis was applied to identify which compounds were hits. For each plate in this set of experiments, the mean percent population of DMSO (N=8 wells) was calculated. This was the mean DMSO value. The percent population value of every well (N=96) was divided by the mean DMSO value. This was the DMSO normalized value. Across all experiments, the DMSO normalized values of the DMSO wells were collected together and the mean and standard deviation were calculated. There were 128 DMSO wells in each lineage effort, except for the megakaryocyte effort, which had 230 DMSO wells. The hit-calling cutoff was equal to the mean + 6*standarddeviation. For a compound to be called a hit, the average of the DMSO normalized values across replicates needed to be greater than the cutoff.
[0119] For validation experiments, significance was determined via a heteroscedastic one-way t-test between normalized DMSO and test sample values.Paired transcriptomic and phenotypic measurements of Mk-inducing compoundsProcessing and integration of perturbational scRNA-seq dataset
[0120] Single-cell DNA sequencing data from one CD34 donor treated at luM with each respective compound or DMSO was collected at 1 day, 2 days, 5 days, and 7 days, in biological duplicates, with paired flow cytometry readouts at Day 7. Libraries were prepared using the Chromium Single Cell 3’ Reagent Kit v3.1, and sequenced on an Illumina NovaSeq 6000 platform, generating paired-end reads. Raw reads were demultiplexed using bcl2fastq (v2.20.0.422) and processed using Cell Ranger software (v5.0.1, lOx Genomics). Reads were aligned to the human reference genome (GRCh38) using STAR aligner (v2.7.0a).
[0121] Hashed sequencing libraries were filtered to include libraries ranging from 5000- 80000 counts, with less than 20% mitochondrial gene counts. Pre-filtered hashed libraries were then demultiplexed using a gaussian mixture model then filtered to singlets. Additional filtering was performed to remove cells with <2500 or >60000 counts, and cells with <1600 or >9000 genes. Total counts per cell were normalized to 10,000 and natural log transformation was applied using functions from Scanpy (vl.9.3). To maintain a consistent embedding of hematopoiesis, SymphonyPy (v0.2.1) was used for reference mapping and label transfer between reference time course CITE-seq dataset and the perturbation time course dataset. In brief, harmony was used to create a batch corrected PC space. The query dataset was then projected into the reference PC space and integrated in the reference’s harmony-corrected PC space. Last, label transfer was conducted using SymphonyPy’s K-nearest neighbors (KNN) classifier, leveraging the shared latent space to transfer cell type annotations from the reference to the new dataset. The robustness of label transfer was validated by examining the weighted Mahalanobis distance of query cells to mapped reference clusters, the cosine similarity across highly variable genes between reference and query cell types, and expression of cell type markers in the query dataset labels achieved from label transfer.Differential abundance testing
[0122] To test whether differences in cell type proportions in the perturbed samples relative to the control DMSO condition were due to random sampling, the python implementation of scProportionTest (vO.1.2) was used. scPropotionTest used a permutation testing framework, appropriate for high dimensional data where standard parametric assumptions may not be suitable. For each comparison, compound vs. DMSO, the proportion of each cell type was calculated. Combined cells for each group were then shuffled to randomize group labels while keeping group size constant and the proportions were recalculated. The process was repeated 1000 times to generate a p-value for significance between the permuted groups.Using transcriptional readout to refine the Megakaryopoiesis signature
[0123] The transcriptional changes in the single-cell time course that were consistently associated with megakaryocyte induction were identified. Using limma [cite], the pseudobulked gene expression of perturbed HSPCs from day 1 of the scRNA-seq time course was regressed against change in megakaryocyte abundance as measured by flow cytometry at day 7. For each gene, the model
[0124] Where ?0is an intercept term,quantifies the relationship between gene expression and the fold change FCmkof late megakaryocytes, and / ?2corrects for library effects. Internally, limma estimated means and variances for each coefficient while correcting for differences in coverage and the inherent sparsity of transcriptomic data.
[0125] For each gene, the model outputs an FDR-adjusted / ?-value indicating the significance of the correlation between that gene’s expression at day 1 and MK induction at day 7. This / 2-value was converted into a score by taking the negative base- 10 logarithm and multiplying by the sign of the association (positive if the gene is correlated with fold change, and negative if it is anti correlated). Genes with FDR-adjusted p-value <0.01 were considered significantly associated with phenotype.
[0126] The genes were divided into three classes (FIG. 5E):Concordant genes were significantly associated with phenotype in the same direction as indicated by the input v-score.Discordant genes were significantly associated with the phenotype in the opposite direction as indicate by the input v-score.The remaining genes had no significant association with phenotype.
[0127] It was hypothesized that the concordant genes drove model performance, whereas the discordant genes reduced it. To test this, the input was modified by setting all but the concordant genes to zero, or all but the discordant genes to zero, and measuring the impact on hit prioritization. Masking the input to only concordant genes improved the prioritization of megakaryocyte inducers as measured by flow cytometry and performed better than masking to random sets of genes of the same size (FIG. 5E).
[0128] Because the compounds used to classify genes as concordant or discordant were the same as those used to test model performance, this strategy may bias the model in favor of these compounds. Therefore, a stratified 5-fold cross-validation was performed to see whether signature refinement can improve recall of unseen hits. The procedure was as follows:1. Divide the profiled compounds randomly into five folds, dividing hits as evenly as possible.2. Designate one-fold as the test set and the remaining four as the training set.3. Using the training set only, identify concordant and discordant genes as described above.4. Mask the input signature to concordant or discordant genes, run the masked signature through the virtual screening model, and record the rank of the test compounds.5. Repeat steps 2-3 four more times, designating each fold as the test set in turn.6. Repeat steps 1-4 ten times with different random seeds, and report for each compound its mean cross-validation rank over all seeds.
[0129] Even with cross-validation, hits were better-prioritized by signatures masked to concordant genes (paired t-test p=0.054). Without being bound to a particular theory, this shows that the signature refinement strategy could enable recovery of novel compounds.Concordance of known megakaryopoiesis markers with measured MK induction
[0130] To better understand the action of hit compounds, their effect on transcription factors involved in megakaryopoiesis, and on marker genes of MKs was observed. A list of nine transcription factors and four MK marker genes was obtained and their differential expressionpatterns in day 1 HSPCs was observed from the transcriptional validation screen. Limma was also used to model their association with MK abundance as described above.
[0131] All nine of the transcription factors were significantly associated with megakaryocyte abundance (p<0.05) and showed significant differential expression in MK inducers, suggesting that the hit compounds bias HSPCs towards the megakaryocyte lineage at an early time point (FIG. 5C). Only one of the MK markers was significantly associated with induction, which was unsurprising given that the sample consisted of HSPCs. In addition, all but two of these genes had positive score in the model input signature, showing that the signature captures at least some of the known biology. The two genes with negative score would be filtered out by signature refinement due to the disagreement between input v-score and observed association with MKs, as described in the previous section.Relating model performance to CD34-relevance
[0132] To quantify the similarity between a compound’s effect in CD34 and in LINCS, the distance between 24h signatures in the experiment and the 10 most similar signatures for the same compound in LINCS was calculated for each compound. Similarities were measured using cosine distance over the 940 landmark genes that were shared between the two assays.Perturbational response in CD34s was represented by a vector of differential expression scores for each shared gene, defined asDES = — log(FDR pvalue) * sign(FC).
[0133] Response in LINCS was represented by the level 4 z-score vector.
[0134] The predicted hits (with MK fold change > 2 at day 7) tend to have smaller differences between CD34 and LINCS response than predicted non-hits. To assess the significance of this result, two-sample independent t-test comparing the mean 10-NN LINCS similarities in hits to those in non-hits was performed, yielding a p-value of 0.02.Calculating cell-type specific pseudobulked differential expression
[0135] To reduce the noise in scRNA-seq, cells were aggregated within a technical replicate of each cell type and day to form pseudobulks. Prior to differential expression test, genes that were expressed by less than 0.5% of cells were removed. The aggregated read counts was analyzed using Limma, with perturbation condition as the main variable, technical replicateas covariate, and the DMSO condition as reference. The differential expression test using Limma on each cell type and day was performed independently to calculate the gene expression logged fold changes (logFC).
[0136] Then, for each cell type and day, principal component analysis was performed on the logFC results of the perturbation conditions.Identifying GO terms associated with PCI and PC2
[0137] Using gseaPy prerank, GO terms most strongly associated with the PCI loadings in the HSPC population at Day 1 were identified using the GO Biological Process 2023 gene set. The terms were sorted on their net enrichment score after filtering the results to terms where the Tag% was more than 0.4 and ranked. This process was repeated for each of PCI and PC2.Inferring pseudotime in the Mk lineage
[0138] A unified developmental pseudotime for cells of the Mk lineage (HSPC, MEMP, MkP, and MPC) were inferred in all perturbation and day conditions, using the ' dpt_pseudotime' function implemented in scanpy (version 3.9).Calculating rolling-window expression of GO terms
[0139] Using the Scanpy score gene function, a gene activity score was calculated for the GO terms of interest in each cell. To find general trends in the change of GO term scores across pseudotime, smoothing for each perturbation was performed by ordering the cells in pseudotime and averaging the scores in rolling windows of 600 cells. The pseudotime values was also transformed using the same rolling windows.
[0140] It will also be understood that, although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first subject could be termed a second subject, and, similarly, a second subject could be termed a first subject, without departing from the scope of the present disclosure. The first subject and the second subject are both subjects, but they are not the same subject.
[0141] The terminology used in the present disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used in the description of the invention and the appended claims, the singular forms “a”, “an” and “the” areintended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0142] As used herein, the term “if’ may be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” depending on the context. Similarly, the phrase “if it is determined” or “if [a stated condition or event] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting (the stated condition or event)” or “in response to detecting (the stated condition or event),” depending on the context.
[0143] The foregoing description included example systems, methods, techniques, instruction sequences, and computing machine program products that embody illustrative implementations. For purposes of explanation, numerous specific details were set forth in order to provide an understanding of various implementations of the inventive subject matter. It will be evident, however, to those skilled in the art that implementations of the inventive subject matter may be practiced without these specific details. In general, well-known instruction instances, protocols, structures and techniques have not been shown in detail.
[0144] The foregoing description, for purpose of explanation, has been described with reference to specific implementations. However, the illustrative discussions above are not intended to be exhaustive or to limit the implementations to the precise forms disclosed. Many modifications and variations are possible in view of the above teachings. The implementations were chosen and described in order to best explain the principles and their practical applications, to thereby enable others skilled in the art to best utilize the implementations and various implementations with various modifications as are suited to the particular use contemplated.INCORPORATION BY REFERENCE
[0145] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference in their entireties to the same extent as if each individualpublication, patent, or patent application was specifically and individually indicated to be incorporated by reference.
[0146] REFERENCES1. Alemany, A., Florescu, M., Baron, C. S., Peterson-Maduro, J. & van Oudenaarden, A. Whole-organism clone tracing using single-cell sequencing. Nature 556, 108-112 (2018).2. Hao, Y. etal. Integrated analysis of multimodal single-cell data. Cell 184, 3573-3587. e29 (2021).3. Burkhardt, D. B. et al. Quantifying the effect of experimental perturbations at single-cell resolution. Nat. Biotechnol. 1-11 (2021) doi: 10.1038 / s41587-020-00803-5.4. Fleck, J. S. et al. Inferring and perturbing cell fate regulomes in human brain organoids. Nature 1-8 (2022) doi:10.1038 / s41586-022-05279-8.5. Kamimoto, K. et al. Dissecting cell identity via network inference and in silico gene perturbation. Nature 614, 742-751 (2023).6. Chan, J., Wang, X., Turner, J. A., Baldwin, N. E. & Gu, J. Breaking the paradigm: Dr Insight empowers signature-free, enhanced drug repurposing. Bioinformatics 35, 2818-2826 (2019).7. Belyaeva, A. et al. Causal network models of SARS-CoV-2 expression and aging to identify candidates for drug repurposing. Nat. Commun. 12, 1024 (2021).8. Hetzel, L. et al. Predicting Cellular Responses to Novel Drug Perturbations at a SingleCell Resolution, in (2022).9. He, B. et al. ASGARD is A Single-cell Guided Pipeline to Aid Repurposing of Drugs. Nat. Commun. 14, 993 (2023).10. Austin, D. & Hayford, T. Research and Development in the Pharmaceutical Industry. (2021).11. Wong, C. H., Siah, K. W. & Lo, A. W. Estimation of clinical trial success rates and related parameters. Biostat. Oxf. Engl. 20, 273-286 (2019).12. Cummings, J. L., Morstorf, T. & Zhong, K. Alzheimer’s disease drug-development pipeline: few candidates, frequent failures. Alzheimers Res. Ther. 6, 37 (2014).13. Sadri, A. Is Target-Based Drug Discovery Efficient? Discovery and “Off-Target” Mechanisms of All Drugs. J. Med. Chem. (2023) doi: 10.1021 / acs.jmedchem.2c01737.14. Scanned, J. W., Blanckley, A., Boldon, H. & Warrington, B. Diagnosing the decline in pharmaceutical R&D efficiency. Nat. Rev. Drug Discov. 11, 191-200 (2012).15. Zheng, W., Thorne, N. & McKew, J. C. Phenotypic screens as a renewed approach for drug discovery. Drug Discov. Today 18, 1067-1073 (2013).16. Moffat, J. G., Vincent, F., Lee, J. A., Eder, J. & Prunotto, M. Opportunities and challenges in phenotypic drug discovery: an industry perspective. Nat. Rev. Drug Discov. 16, 531-543 (2017).17. Cruz-Monteagudo, M. et al. Systemic QSAR and phenotypic virtual screening: chasing butterflies in drug discovery. Drug Discov. Today 22, 994-1007 (2017).18. Stokes, J. M. et al. A Deep Learning Approach to Antibiotic Discovery. Cell 180, 688- 702. el3 (2020).19. Theodoris, C. V. et al. Network-based screen in iPSC-derived cells reveals therapeutic candidate for heart valve disease. Science 371, eabd0724 (2021).20. Zhu, J. et al. Prediction of drug efficacy from transcriptional profiles with deep learning. Nat. Biotechnol. 39, 1444-1452 (2021).21. Subramanian, A. et al. A Next Generation Connectivity Map: L1000 Platform and the First 1,000,000 Profiles. Cell 171, 1437-1452.el7 (2017).22. Srivatsan, S. R. et al. Massively multiplex chemical transcriptomics at single-cell resolution. Science 367, 45-51 (2020).23. Pedregosa, F. et al. Scikit-leam: Machine Learning in Python. J. Mach. Learn. Res. 12, 2825-2830 (2011).24. Evangelista, J. E. et al. SigCom LINCS: data and metadata search engine for a million gene expression signatures. Nucleic Acids Res. 50, W697-W709 (2022).25. He, B. et al. ASGARD is A Single-cell Guided Pipeline to Aid Repurposing of Drugs. Nat. Commun. 14, 993 (2023).26. What is the total gene space accessible by L1000? - CONNECTOPEDIA. clue.io / connectopedia / 11000_gene_space.27. Wechsler, J. et al. Acquired mutations in GATA1 in the megakaryoblastic leukemia of Down syndrome. Nat. Genet. 32, 148-152 (2002).28. Giannoni, P. et al. A High Percentage of CD16+ Monocytes Correlates with the Extent of Bone Erosion in Chronic Lymphocytic Leukemia Patients: The Impact of Leukemic B Cells in Monocyte Differentiation and Osteoclast Maturation. Cancers 14, 5979 (2022).29. Dhawan, A. & Padron, E. Abnormal monocyte differentiation and function in chronic myelomonocytic leukemia. Curr. Opin. Hematol. 29, 20-26 (2022).30. Voit, R. A. et al. A genetic disorder reveals a hematopoietic stem cell regulatory network co-opted in leukemia. Nat. Immunol. 24, 69-83 (2023).31. Drachman, J. G., Jarvik, G. P. & Mehaffey, M. G. Autosomal dominant thrombocytopenia: incomplete megakaryocyte differentiation and linkage to human chromosome 10. Blood 96, 118-125 (2000).32. Wang, N. et al. Targeting of Calbindin 1 rescues erythropoiesis in a human model of Diamond Blackfan anemia. Blood Cells. Mol. Dis. 102, 102759 (2023).33. Germino-Watnick, P. et al. Hematopoietic Stem Cell Gene-Addition / Editing Therapy in Sickle Cell Disease. Cells 11, 1843 (2022).34. Appelbaum, F. R. Hematopoietic-cell transplantation at 50. N. Engl. J. Med. 357, 1472- 1475 (2007).35. Daniel Burkhardt, R. H., Malte Luecken, Andrew Benz, Peter Holderrieth, Jonathan Bloom, Christopher Lance, Ashley Chow. Open Problems - Multimodal Single-Cell Integration Competition. (2022).36. Chan, J., Wang, X., Turner, J. A., Baldwin, N. E. & Gu, J. Breaking the paradigm: Dr Insight empowers signature-free, enhanced drug repurposing. Bioinformatics 35, 2818-2826 (2019).37. McDonald, T. P. & Sullivan, P. S. Megakaryocytic and erythrocytic cell lines share a common precursor cell. Exp. Hematol. 21, 1316-1320 (1993).38. Lu, Y. C. et al. The Molecular Signature of Megakaryocyte-Erythroid Progenitors Reveals a Role for the Cell Cycle in Fate Specification. Cell Rep 25, 2083-2093 (2018).39. Zimmerman, K. D., Espeland, M. A. & Langefeld, C. D. A practical solution to pseudoreplication bias in single-cell studies. Nat. Commun. 12, 738 (2021).40. Huang, H. & Li, Y. Mechanisms Controlling Mast Cell and Basophil Lineage Decisions. Curr. Allergy Asthma Rep. 14, 457 (2014).41. Kurotaki, D. et al. Essential role of the IRF8-KLF4 transcription factor cascade in murine monocyte differentiation. Blood 121, 1839-1849 (2013).42. Zamani, F., Zare Shahneh, F., Aghebati-Maleki, L. & Baradaran, B. Induction of CD14 Expression and Differentiation to Monocytes or Mature Macrophages in Promyelocytic Cell Lines: New Approach. Adv. Pharm. Bull. 3, 329-332 (2013).43. Metzler, K. D. et al. Myeloperoxidase is required for neutrophil extracellular trap formation: implications for innate immunity. Blood 117, 953-959 (2011).44. Korsunsky, I. et al. Fast, sensitive and accurate integration of single-cell data with Harmony. Nat. Methods 16, 1289-1296 (2019).45. Ritchie, M. E. et al. limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res. 43, e47 (2015).46. Jin, S., MacLean, A. L., Peng, T. & Nie, Q. scEpath: energy landscape-based inference of transition probabilities and cellular trajectories from single-cell transcriptomic data.Bioinformatics 34, 2077-2086 (2018).47. Kelly, K. L. et al. De novo lipogenesis is essential for platelet production in humans. Nat. Metab. 2, 1163-1178 (2020).48. de Jonckheere, B. et al. Critical shifts in lipid metabolism promote megakaryocyte differentiation and proplatelet formation. Nat. Cardiovasc. Res. 2, 835-852 (2023).49. Vincent, F. et al. Phenotypic drug discovery: recent successes, lessons learned and new directions. Nat. Rev. Drug Discov. 21, 899-914 (2022).50. Lee, H., Kang, S. & Kim, W. Drug Repositioning for Cancer Therapy Based on Large- Scale Drug-Induced Transcriptional Signatures. PLOS ONE 11, e0150460 (2016).51. Belyaeva, A. et al. Causal network models of SARS-CoV-2 expression and aging to identify candidates for drug repurposing. Nat. Common. 12, 1024 (2021).52. Wang, H. et al. Scientific discovery in the age of artificial intelligence. Nature 620, 47-60 (2023).53. Ye, C. et al. DRUG-seq for miniaturized high-throughput transcriptome profiling in drug discovery. Nat. Commun. 9, 4307 (2018).54. Hetzel, L. et al. Predicting Cellular Responses to Novel Drug Perturbations at a Single-Cell Resolution, in (2022).55. Lange, M. et al. CellRank for directed single-cell fate mapping. Nat. Methods 19, 159-170 (2022).56. Rood, J. E., Maartens, A., Hupalowska, A., Teichmann, S. A. & Regev, A. Impact of the Human Cell Atlas on medicine. Nat. Med. 28, 2486-2496 (2022).57. Dann, E., Teichmann, S. A. & Marioni, J. C. Precise identification of cell states altered in disease with healthy single-cell references. 2022.11.10.515939 Preprint at doi.org / 10.1101 / 2022.11.10.515939 (2022).58. Yang, H., Sun, L., Liu, M. & Mao, Y. Patient-derived organoids: a promising model for personalized cancer treatment. Gastroenterol. Rep. 6, 243-245 (2018).59. Skardal, A., Shupe, T. & Atala, A. Organoid-on-a-chip and body-on-a-chip systems for drug screening and disease modeling. Drug Discov. Today 21, 1399-1411 (2016).
[0147] All references cited herein are incorporated herein by reference in their entirety and for all purposes to the same extent as if each individual publication or patent or patent application was specifically and individually indicated to be incorporated by reference in its entirety for all purposes.
Claims
CLAIMSWhat is claimed:
1. A method for identifying and prioritizing chemical perturbances that induce targeted changes in the state of a cell.
2. The method of claim 1, wherein the targeted changes comprise one or more changes in a disease-relevant phenotype, optionally wherein the one or more changes in a disease-relevant phenotype is obtained from one or more atlases of cellular data, optionally wherein the cellular data comprises single-cell RNA sequencing data.
3. The method of claim 2, wherein the disease-relevant phenotype is gene expression, optionally wherein the one or more changes in gene expression is measured and / or calculated by determining a differential expression between each state of the cell to provide a gene expression signature.
4. The method of any one of claims 1-3, wherein the method further comprises using a deep learning model to virtually screen a plurality of compounds (e.g. a plurality of compounds likely to induce a perturbance that induces the desired changes in the state of the cell) and identify from the plurality of compounds one or more compounds that induces a perturbance that induces changes in the state of the cell that is the same or substantially the as the targeted changes in the state of the cell.
5. The method of claim 4, wherein the method further comprises using the deep learning model to prioritize the one or more compounds, optionally wherein the one or more compounds are prioritized based at least on the similarity of the induced changes in the state of the cell to the targeted changes in the state of the cell.
6. The method of claim 5, wherein the method further comprises experimentally measuring the chemical perturbance of the one or more compounds, optionally using one or more phenotypic assays.
7. The method of claim 6, wherein the experimentally measuring comprises measuring one or more changes in lineage commitment of the cell after treatment with the one or more compounds, optionally wherein the method comprises identifying and / or measuring surface markers associated with each change in lineage.
8. The method of any one of claims 4-7, wherein the method further comprises inputting data obtained from the experimental measuring to the deep learning model, wherein the inputting further refines the deep learning model, optionally wherein the inputting further refines the gene expression signature.
9. The method of any one of claims 1-8, wherein the cell is selected from a megakaryocyte, an erythrocyte, an eosinophil / basophil / mast cell (EBM), a monocyte, and a neutrophil.
10. A system for identifying and prioritizing chemical perturbances that induce targeted changes in the state of a cell.
11. The system of claim 10, comprising:(a) desired targeted changes in the state of the cell;(b) a deep learning model; and(c) one or more phenotypic assays.
12. The system of claim 10 or 11, wherein the targeted changes comprise one or more changes in a disease-relevant phenotype, optionally wherein the disease-relevant phenotype is gene expression.
13. The system of claim 12, wherein the one or more changes in gene expression is measured and / or calculated by determining a differential expression between each state of the cell toprovide a gene expression signature, optionally wherein the gene expression signature is obtained from one or more atlases of cellular data, optionally wherein the cellular data comprises single-cell RNA sequencing data.
14. The system of claims 11-13, wherein the deep learning model virtually screens a plurality of compounds and identify from the plurality of compounds one or more compounds that induces a perturbance that induces changes in the state of the cell that is the same or substantially the same as the targeted changes in the state of the cell.
15. The system of any one of claims 11-14, wherein the deep learning model prioritizes the one or more compounds, optionally wherein the one or more compounds are prioritized based at least on the similarity of the induced changes in the state of the cell to the targeted changes in the state of the cell.
16. The system of any one of claims 11-15, wherein the one or more phenotypic assays measure the chemical perturbance of the one or more compounds.
17. The system of any one of claims 11-16 wherein the one or more phenotypic assays measure one or more changes in lineage commitment of the cell after treatment with the one or more compounds, optionally wherein the one or more phenotypic assays identify and / or measure surface markers associated with each change in lineage.
18. The system of any one of claims 11-17, wherein data obtained from the one or more phenotypic assays is inputted into the deep learning model, wherein the inputting further refines the deep learning model, optionally wherein the inputting further refines the gene expression signature.
19. The system of any one of claims 11-18, wherein the cell is selected from a megakaryocyte, an erythrocyte, an eosinophil / basophil / mast cell (EBM), a monocyte, and a neutrophil.
Citation Information
Patent Citations
Whole cell assays and methods
US20210215673A1
Systems and methods for associating compounds with physiological conditions using fingerprint analysis
US20220403335A1