Methods for drug target deconvolution

The method uses pooled CRISPR screens and single-cell RNA-seq with a Bayesian framework to accurately identify drug targets by addressing non-uniform perturbation responses and cell heterogeneity, enhancing the resolution and reliability of drug target deconvolution.

WO2025171393A1PCT designated stage Publication Date: 2025-08-14NEW YORK GENOME CENT +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/015271
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-10
Filing Date
2025-02-10
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing methods for determining drug targets are limited in resolution and scale, often failing to accurately reflect how drugs function in living cells, and conventional computational approaches struggle with non-uniform perturbation responses, cell heterogeneity, and correlated gene perturbations, leading to misleading inferences.

Method used

A method incorporating pooled CRISPR screens with single-cell RNA-seq and a Bayesian regression framework to deconvolute drug targets, using a multi-condition latent factor model to account for baseline heterogeneity and quantify uncertainty, enabling direct probabilistic statements about likely targets.

Benefits of technology

Enables high-resolution identification of drug targets by disentangling perturbation effects from baseline heterogeneity and noise, providing interpretable and accurate probabilistic assessments of target credibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000015_0001
    Figure IMGF000015_0001
  • Figure IMGF000016_0001
    Figure IMGF000016_0001
  • Figure IMGF000016_0002
    Figure IMGF000016_0002
Patent Text Reader

Abstract

A method for determining one or more targets of a drug is provided. The method includes introducing one or more perturbation constructs encoding for one or more sequence specific perturbations to a plurality of cells in a population of cells; detecting endogenous mRNAs and a sequence that identifies the perturbation for each single cell in the plurality of cells using single cell RNA-seq; and comparing measured differences in the endogenous mRNAs that are correlated to the sequence specific perturbations detected for each single cell with differences in endogenous mRNAs in a cell treated with the drug, thereby identifying one or more targets of the drug.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]NYG-LIPP-215.PCT METHODS FOR DRUG TARGET DECONVOLUTION STATEMENT OF GOVERNMENT SUPPORT This invention was made with government support under R21CA272345 and RM1HG011014-01 awarded by the National Institutes of Health. The government has certain rights in the invention. BACKGROUND Recent studies have suggested that many drugs do not actually function via their purported targets, with critical implications for efficacy and toxicity. Traditional assessments of drug targets are limited in resolution and scale, and often do not recapitulate how a drug will function in a living cell. What is needed are improved methods for drug target deconvolution. SUMMARY OF THE INVENTION In a first aspect, a method for determining one or more targets of a drug is provided. The method includes (a) introducing one or more perturbation constructs encoding for one or more sequence specific perturbations to a plurality of cells in a population of cells, wherein each cell in the plurality of cells receives at least 1 perturbation construct and wherein each perturbation construct comprises a CRISPR guide RNA (gRNA) sequence, a sequence that identifies the perturbation, and a reverse-transcription handle; (b) detecting endogenous mRNAs and the sequence that identifies the perturbation for each single cell in the plurality of cells using single cell RNA-seq; wherein each sequence identifying the perturbation is a unique perturbation barcode that identifies the gRNA introduced to a cell; (c) comparing measured differences in the endogenous mRNAs that are correlated to the sequence specific perturbations detected for each single cell with differences in endogenous mRNAs in a cell treated with the drug, and (d) identifying one or more targets of the drug when a pattern of differences in the endogenous mRNAs in the cell treated with the drug correlates with differences in the endogenous mRNAs that are correlated to the sequence specific perturbations. In certain embodiments, the identifying step further comprises determining the differences in the endogenous mRNAs that are correlated to the sequence specific perturbations on the single cells by applying a NYG-LIPP-215.PCT computational model. In another aspect, a method for determining one or more credible targets of a drug is provided. The method includes (a) introducing one or more perturbation constructs encoding for one or more sequence specific perturbations to a plurality of cells in a population of cells, wherein each cell in the plurality of cells receives at least 1 perturbation construct and wherein each perturbation construct comprises a CRISPR guide RNA (gRNA) sequence, a sequence that identifies the perturbation, and a reverse-transcription handle; (b) detecting endogenous mRNAs and the sequence that identifies the perturbation for each single cell in the plurality of cells using single cell RNA-seq; wherein each sequence identifying the perturbation is a unique perturbation barcode that identifies the gRNA introduced to a cell; and (c) comparing measured differences in the endogenous mRNAs that are correlated to the sequence specific perturbations detected for each single cell with differences in endogenous mRNAs in a cell treated with the drug, thereby identifying one or more credible targets of the drug. Other aspects and advantages of the invention will be readily apparent from the following detailed description of the invention. BRIEF DESCRIPTION OF THE DRAWINGS FIG.1 is an overview of an embodiment of the invention. FIG.2 is an overview of an embodiment of statistical analysis described herein. FIG.3A and FIG.3B show results of a demonstration of the invention, described in Example 1. FIG.3A shows precision vs. sensitivity for classifying cells withheld from a perturbation screen dataset. FIG.3B is a heatmap of normalized gene expression of three selected marker genes for cells receiving either PAF1 or CTR9 perturbations, alongside their estimated posterior inclusion probability under our approach. FIG.4A and FIG.4B show results of a validation of an approach described herein in Example 3. FIG.4A shows the posterior inclusion probabilities of withheld perturbed cells. FIG. 4B is a calibration curve for CPSF1. FIGs.5A and 5B show the results of a target identification experiment from 3,900+ perturbations in positive control drugs as described in Example 3. FIG.2A shows the frequency with which each target was found as the highest-probability perturbation across cells treated with Bcr- Abl inhibitors, after filtering out apoptotic and non-responding cells. FIG.2B shows NYG-LIPP-215.PCT the process for finding the important genes, and illustrates them through a heatmap. FIG.6 shows the results of an experiment as described in Example 4. FIG.7 demonstrates ceconvolving dual-gene perturbations in terms of single-gene perturbation effects. For each dualgene perturbation shown on the x-axis, we show the top two credible sets returned by our sample-level analysis, indicated in blue and red. FIGs.8A and 8B demonstrate deconvolving lncRNAs in terms of protein-coding genes and Genome-Wide Perturb-seq perturbations. (FIG.8A - Left) Expression heatmap comparing knockdowns of the four lncRNAs of interest to non-targeting cells in the original lncRNA screen. (FIG.8A - Right) On the same set of genes, expression heatmap comparing knockdowns of genes related to the PAXT complex to non-targeting cells in Genome-Wide Perturb-seq. (FIG.8B) On the same set of genes, expression heatmap comparing knockdown of HNRNPF to non-targeting cells in the original lncRNA screen. DETAILED DESCRIPTION OF THE INVENTION Pooled CRISPR screens, such as Perturb-seq (Dixit, Atray, et al. Perturb-Seq: dissecting molecular circuits with scalable single-cell RNA profiling of pooled genetic screens. Cell 167.7 (2016): 1853-1866.) and CaRPool-seq (Wessels, HH., et al, Efficient combinatorial targeting of RNA transcripts in single cells with Cas13 RNA Perturb-seq. Nat Methods 20, 86–94 (2023)), coupled with single-cell RNA-sequencing have enabled systematic interrogation of gene function and regulatory networks. The present invention incorporates these massively parallel perturbation profiling techniques with a computational model to elucidate drug molecular targets. Described herein are methods for deconvolution of transcriptional or other molecular patterns from drug screens in terms of patterns from genetic perturbation screens, to identify what target(s) are being inhibited or activated by the drug at the single-cell level. The key idea is that if the expression pattern of drug-treated cells can be deconvolved in terms of one or more perturbation patterns from cells with certain gene(s) knocked down, we can infer those as the drug’s targets. See, FIG.3. However, this high-dimensional deconvolution problem poses several important statistical challenges. First, perturbation responses are typically not uniform, with subpopulations of cells “escaping” the perturbation or exhibiting variable levels of knockdown. NYG-LIPP-215.PCT Second, even for cell lines, there can be substantial baseline heterogeneity in the cell population independent of perturbations or drug treatments, due to variability in cell cycle, differentiation, or other aspects of cell state. Particularly when perturbation and / or drug effects are more subtle, both of these factors can obscure real signals and bias their estimation. Finally, because many genes are closely related to one another (e.g., encoding members of the same complex), their perturbation responses can be highly correlated. Conventional computational approaches would be likely to report mappings to just one of these or to understate the uncertainty among them, even when there is insufficient evidence to draw a strong conclusion. This can yield misleading or even completely incorrect inferences. We address these issues through an interpretable statistical approach with two main components: A multi-condition latent factor model that estimates denoised perturbation effects Λs by accounting for baseline heterogeneity; and a Bayesian regression framework, adapted from fine-mapping G. Wang, A. Sarkar, P. Carbonetto, M. Stephens (2020). A simple new approach to variable selection in regression, with application to genetic fine mapping, J. R. Stat. Soc. Series B Stat. Methodol.82(5), 1273–1300, that deconvolutes drug-treated cells in terms of these Λs and summarizes complex patterns of uncertainty via credible sets. See, FIG.4. Our approach enables direct probabilistic statements about about which sets of related perturbations are likely to contain the true target(s). For example: “The drug inhibits at least one of the closely related targets {1, 2} with relative probabilities of 60% and 40%, and also target 4 with probability 95% (referring to FIG.4)” Pooled Crispr Screen The methods described herein involve matching the effects shown from one or more CRISPR perturbations with the effects of a drug. While existing data can be analyzed using the methods described herein, in certain embodiments, the methods include performing one or more pooled CRISPR screens. Various pooled CRISPR perturbation screens that are useful herein are known in the art and include Perturb-seq (Dixit, Atray, et al. Perturb-Seq: dissecting molecular circuits with scalable single-cell RNA profiling of pooled genetic screens. Cell 167.7 (2016): 1853-1866, incorporated herein by reference) and CaRPool-seq (Wessels, HH., et al, Efficient combinatorial targeting of RNA transcripts in single cells with Cas13 RNA Perturb-seq. Nat Methods 20, 86–94 (2023) incorporated herein by reference), etc., and variations thereon. NYG-LIPP-215.PCT Each of these systems combines pooled, combinatorial CRISPR screens with single-cell RNA sequencing (scRNA-seq) readout. By providing a guide barcode, the identity of the gRNA, and thus the perturbation, can be identified during the sequencing step. Several of these methodologies are described here, in brief. CRISPR-Cas systems refer to non-naturally occurring systems derived from bacterial Clustered Regularly Interspaced Short Palindromic Repeats loci. These systems generally comprise an enzyme (Cas protein, such as Cas9 or Cas13 protein) and one or more RNAs. Said RNA is a CRISPR RNA (crRNA) and may be a single-guide RNA (sgRNA). Said RNA and / or said enzyme may be engineered, for example for optimal use in mammalian cells, for optimal delivery therein, for optimal activity therein, for specific uses in gene editing, etc. CRISPR perturbation screens are known in the art. In certain embodiments PERTURB-seq, such as that described in US Patent No.11,214,797, or a variation thereof, is used. In certain embodiments, CaRPool-seq, such as that described in WO 2022 / 187262 or a variation thereof, is used. These documents are incorporated herein by reference. The methods described herein are useful with both DNA and RNA targets, e.g., such as using Perturb-seq or CaRPool-seq. For each embodiment described using DNA as a target, a similar embodiment is meant using RNA as a target, and vice versa. As used herein, the term “perturbation” refers to the effects on one more transcripts, genes, or gene products (including protein) as a result of a mutation or modification of a target sequence. Mutation or modifications include, e.g. small nucleotide insertions or deletions (indels) or a larger deletion, insertion, or inversion. In certain embodiments, the introduction a mutation or modification is referred to as “editing” or “gene editing”. As used herein, “perturbation” also includes CRISPR-mediated activation or inactivation (CRISPRa / i) systems, which may be used to activate or inactivate gene transcription. Briefly, a nuclease-dead (deactivated) Cas9 RNA-guided DNA binding domain (dCas9) (Qi, L. S., et al, Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression. Cell.152, 1173-1183, doi:10.1016 / j.cell.2013.02.022 (2013). PMCID:3664290) tethered to transcriptional repressor domains that promote epigenetic silencing (e.g., KRAB) forms a “CRISPRi” (Gilbert, L. A., et al, CRISPR-mediated modular RNA-guided regulation of transcription in eukaryotes. Cell.154, 442-451, doi:10.1016 / j.cell.2013.06.044 (2013). PMCID:3770145; Konermann, S., et al, F. Optical control of mammalian endogenous NYG-LIPP-215.PCT transcription and epigenetic states. Nature.500, 472-476, doi:10.1038 / nature12466 (2013). PMCID:3856241) that represses transcription. To use dCas9 as an activator (CRISPRa), a guide RNA may be engineered to carry RNA binding motifs (e.g., MS2) that recruit effector domains fused to RNA-motif binding proteins, increasing transcription (Konermann, S., et al., Genome- scale transcriptional activation by an engineered CRISPR-Cas9 complex. Nature.517, 583-588, doi:10.1038 / nature14136 (2015). PMCID:4420636). Provided herein are methods for determining one or more targets of a drug. In certain embodiments, the method includes introducing one or more perturbation constructs encoding for one or more sequence specific perturbations to a plurality of cells in a population of cells, wherein each cell in the plurality of cells receives at least 1 perturbation construct and wherein each perturbation construct comprises a CRISPR guide RNA (gRNA) sequence, a sequence that identifies the perturbation, and a reverse-transcription handle. The method further includes detecting endogenous mRNAs and the sequence that identifies the perturbation for each single cell in the plurality of cells using single cell RNA-seq; wherein each sequence identifying the perturbation is a unique perturbation barcode that identifies the gRNA introduced to a cell. In certain embodiments, one or more perturbation constructs is introduced into a cell or plurality of cells in a population of cells. The perturbation constructs perturb a polynucleotide sequence in the cells, and each distinct sequence-perturbation construct comprises a CRISPR guide RNA (gRNA) sequence, a sequence that identifies the perturbation, and a reverse- transcription handle. The gRNA allows for sequence-specific perturbations. The terms “guide RNA” or “gRNA” or “single guide RNA” or “sgRNA” are used interchangeably herein and refer to a nucleic acid sequence which can hybridize to a sequence (hybridization region or target region) of a target nucleic acid (or target sequence), e.g., a target RNA or DNA. The sequence of the gRNA determines the target sequence for gene editing, knock-down, knock-out, insertion, etc. For genome-wide approaches, it is possible to design and construct suitable gRNA libraries. Such gRNAs may be delivered to cells using vector delivery such as viral vector delivery. Combination of CRISPR-Cas-mediated perturbations may be obtained by delivering multiple gRNAs within a single cell. This may be achieved in pooled format. In the case of gRNA viral vector delivery, combined perturbation may be obtained by delivering several gRNA vectors to the same cell. This may also be achieved in pooled format, and number of combined perturbations in a cell then corresponds to the MOI (multiplicity of NYG-LIPP-215.PCT infection). Using CRISPR-Cas systems, one may generally implement MOI values of up to 10, 12 or 15. As used herein, a “target sequence” refers to a sequence to which a guide (spacer) sequence is designed to have reverse complementarity, where hybridization between a target sequence and a spacer sequence promotes the formation of a CRISPR complex. A target sequence may comprise RNA polynucleotides or DNA polynucleotides. In certain embodiments, CARP-seq, such as that described in WO 2022 / 187262 or a variation thereof, is used. Briefly, CARP-seq utilizes CRISPR arrays that include one or more crRNA(s) and a barcode guide RNA (bcgRNA), as well as nucleic acids, expression cassettes, and vectors comprising sequences encoding for the same. A “CRISPR array” refers to an arrangement of crRNA(s) and a barcode guide RNA (bcgRNA) capable of introducing one or more perturbations in a cell transcriptome. The presence of the bcgRNA facilitates detection of the CRISPR array during downstream analysis, including single cell readouts. In some embodiments, the CRISPR array comprises, from 5’ to 3’ one, two, three, or more crRNAs. As with CARP-seq and PERTURB-seq, methods provided herein require the use of barcode sequences which allows for detection and / or identification of the specific gRNA in use to allow for correlation with the specific perturbation. As used herein, the term “host cell” may refer to any target cell having a target RNA or DNA or suspected of having a target RNA or DNA. Thus, a “host cell,” refers to a prokaryotic or eukaryotic cell that contains a perturbation construct described herein, that has been introduced into the cell by any means, e.g., electroporation, nucleofection, calcium phosphate precipitation, microinjection, transformation, viral infection, transfection, liposome delivery, membrane fusion techniques, high velocity DNA-coated pellets, viral infection and protoplast fusion. In certain embodiments herein, the term “host cell” refers to a cultured cell of any mammalian species for in vitro assessment of the compositions described herein. The term “host cell” may also refer to the packaging cell line that contains a production plasmid to generate a viral or non-viral vector describe herein. In one embodiment, the host cell is a primary cell. The term “primary cell” refers to a cell isolated directly from a multicellular organism. Primary cells typically have undergone very few population doublings and are therefore more representative of the main functional component of the tissue from which they are derived in comparison to continuous (tumor or NYG-LIPP-215.PCT artificially immortalized) cell lines. In some cases, primary cells are cells that have been isolated and then used immediately. In other cases, primary cells cannot divide indefinitely and thus cannot be cultured for long periods of time in vitro. In other embodiments, the cell is a cultured cell, e.g., a mammalian cell or bacterial cell. In certain embodiments, the host cell is an immortalized cell line. In certain instances, the primary cell is a stem cell or an immune cell. Non-limiting examples of stem cells include hematopoietic stem and progenitor cells (HSPCs) such as CD34+ HSPCs, mesenchymal stem cells, neural stem cells, organ stem cells, and combinations thereof. Non-limiting examples of immune cells include T cells (e.g., CD3+ T cells, CD4+ T cells, CD8+ T cells, tumor infiltrating cells (TILs), memory T cells, memory stem T cells, effector T cells), natural killer cells, monocytes, peripheral blood mononuclear cells (PBMCs), peripheral blood lymphocytes (PBLs), and combinations thereof. A “nucleic acid” or a “nucleotide”, as described herein, can be RNA, DNA, or a modification thereof, and can be selected, for example, from a group including: nucleic acid encoding a protein of interest, oligonucleotides, nucleic acid analogues, for example peptide- nucleic acid (PNA), pseudocomplementary PNA (pc-PNA), locked nucleic acid (LNA) etc. In certain embodiments, the terms “nucleotide” “nucleic acid” “nucleotide residue” and “nucleic acid residue” are used interchangeably, referring to a nucleotide in a nucleic acid polymer. In a further embodiment, consecutive nucleotide residues refer to nucleotide residues in a contiguous region of a nucleic acid polymer. As used herein, an “expression cassette” refers to a nucleic acid molecule which encodes one or more elements of a gene editing system, e.g. perturbation constructs as described herein and / or a Cas enzyme. An expression cassette also contains a promoter and may contain additional regulatory elements that control expression of one or more elements of a gene editing system in a host cell. In one embodiment, the expression cassette may be packaged into the capsid of a viral vector (e.g., a viral particle). In one embodiment, such an expression cassette for generating a viral vector as described herein is flanked by packaging signals of the viral genome and other expression control sequences such as those described herein. The term “expression” is used herein in its broadest meaning and comprises the production of RNA, of protein, or of both RNA and protein. Expression may be transient or may be stable. Expression cassettes can be delivered to a host cell via any suitable delivery system. NYG-LIPP-215.PCT Suitable non-viral delivery systems are known in the art (see, e.g., Ramamoorth and Narvekar. J Clin Diagn Res.2015 Jan; 9(1):GE01-GE06, which is incorporated herein by reference) and can be readily selected by one of skill in the art and may include, e.g., naked DNA, naked RNA, dendrimers, PLGA, polymethacrylate, an inorganic particle, a lipid particle (e.g., a lipid nanoparticle or LNP), or a chitosan-based formulation. In certain embodiments, one or more elements of gene editing system are encoded by a nucleic acid sequence that is delivered to a host cell by a vector or a viral vector, of which many are known and available in the art. In one embodiment, provided is a vector comprising an expression cassette as described herein. In one embodiment, the vector is a non-viral vector. In another embodiment, the vector is a viral vector. A “viral vector” refers to a synthetic or artificial viral particle in which an expression cassette containing a nucleic acid sequence of interest is packaged in a viral capsid or envelope. Examples of viral vectors include but are not limited to lentivirus, adenoviruses (Ads), retroviruses (γ-retroviruses and lentiviruses), poxviruses, adeno-associated viruses (AAV), baculoviruses, herpes simplex viruses. In one embodiment, the viral vector is replication defective. A “replication-defective virus” refers to a viral vector, wherein any viral genomic sequences also packaged within the viral capsid or envelope are replication-deficient, i.e., they cannot generate progeny virions but retain the ability to infect cells. In one embodiment, the vector is a non-viral plasmid that comprises an expression cassette described herein, e.g., naked DNA, naked plasmid DNA, RNA, and mRNA; coupled with various compositions and nano particles, including, e.g., micelles, liposomes, cationic lipid - nucleic acid compositions, poly-glycan compositions and other polymers, lipid and / or cholesterol-based - nucleic acid conjugates, and other constructs such as are described herein. See, e.g., X. Su et al, Mol. Pharmaceutics, 2011, 8 (3), pp 774–787; web publication: March 21, 2011; WO2013 / 182683, WO 2010 / 053572 and WO 2012 / 170930, all of which are incorporated herein by reference. The method includes providing a host cell comprising a perturbation construct as described herein and a Cas enzyme, wherein the CRISPR-Cas enzyme introduces one or more perturbations in the cell transcriptome. The method further includes isolating RNA from the cell and performing reverse-transcription comprising contacting an RNA sample with a primer specific for the reverse-transcription handle and identifying the barcode sequence. The method NYG-LIPP-215.PCT further includes detecting expression of one or more transcripts or gene products. In some embodiments, the method further includes the use of one or more of flow cytometric analysis, cell-hashing, single-cell sequencing analysis, single cell RNA sequencing (scRNA-seq), Perturb- seq, CROP-seq, CRISP-seq, ECCITE-seq., and cellular indexing of transcriptomes and epitopes (CITE-seq). In some embodiments, the method includes amplifying the barcode sequence using a primer specific for the PCR handle. In another embodiment, the method of performing gene perturbation profiling includes providing a host cell comprising a CRISPR array as described herein and a Cas enzyme, wherein the CRISPR-Cas enzyme introduces one or more perturbations in the cell transcriptome, and the host cell is labeled with one or fluorophore-conjugated antibodies, and sorted using flow cytometry. The method further includes isolating RNA from the isolate cell and performing reverse-transcription comprising contacting an RNA sample with a primer specific for the reverse-transcription handle and identifying the barcode sequence. The method further includes detecting expression of one or more transcripts or gene products. In some embodiments, the method further includes the use of one or more of flow cytometric analysis, cell-hashing, single- cell sequencing analysis, single cell RNA sequencing (scRNA-seq), Perturb-seq, CROP-seq, CRISP-seq, ECCITE-seq., and cellular indexing of transcriptomes and epitopes (CITE-seq). In some embodiments, the method includes amplifying the barcode sequence using a primer specific for the PCR handle. To use the CRISPR arrays described herein, it may be desirable to provide them in conjunction with a nucleic acid that encodes a Cas13d protein. The nucleic acid encoding the Cas13d may be cloned into an intermediate vector for transformation into prokaryotic or eukaryotic cells for replication and / or expression. Intermediate vectors are typically prokaryote vectors, e.g., plasmids, or shuttle vectors, or insect vectors, for storage or manipulation of the nucleic acid encoding the Cas13 protein variant for production of the same. The nucleic acid encoding the Cas13d protein can also be cloned into an expression vector, for administration to a plant cell, animal cell, preferably a mammalian cell or a human cell, fungal cell, bacterial cell, or protozoan cell. The CRISPR arrays as described herein and Cas13 polypeptide, mRNA encoding a Cas13d polypeptide, and / or recombinant expression vector comprising a nucleotide sequence encoding a Cas13d polypeptide variant or fragment thereof are introduced into the cell using any NYG-LIPP-215.PCT suitable method such as by electroporation. In certain instances, the CRISPR array is complexed with a Cas nuclease (e.g., Cas13d polypeptide) or a variant or fragment thereof to form a ribonucleoprotein (RNP)-based delivery system for introduction into a cell (e.g., an in vitro cell for transcriptome perturbation analysis). In other instances, the CRISPR array is introduced into a cell (e.g., an in vitro cell for transcriptome perturbation analysis) with an mRNA encoding a Cas nuclease (e.g., Cas13d polypeptide) or a variant or fragment thereof. In yet other instances, the CRISPR array is introduced into a cell (e.g., an in vitro cell for transcriptome perturbation analysis) with a recombinant expression vector comprising a nucleotide sequence encoding a Cas nuclease (e.g., Cas13 polypeptide) or a variant or fragment thereof. In some embodiments, the RNA or DNA sequencing occurs by methods that include, without limitation, whole transcriptome analysis, whole genome analysis, barcoded sequencing of whole or targeted regions of the genome, and combinations thereof. In some embodiments, the method comprises detection of cell surface proteins using, e.g. flow cytometry. In certain embodiments, the methods, comprise detection or identification of a CRISPR array as provide herein in combination with profiling additional molecular modalities using methods described in the art, including for example single-cell sequencing analysis (e.g.10X Genomics Multiome platform), single-cell RNA-sequencing (scRNA-seq) (See, e.g., Haque et al. A practical guide to single-cell RNA-sequencing for biomedical research and clinical applications, Genome Medicine, 9, Article number: 75 (2017); Hwang et al. Single-cell RNA sequencing technologies and bioinformatics pipelines. Exp Mol Med.2018 Aug 7;50(8):96), cell-hashing (See, e.g., Stoeckius et al. Cell Hashing with barcoded antibodies enables multiplexing and doublet detection for single cell genomics. Genome Biol.2018; 19: 224), Perturb-Seq. (See, e.g., Dixit et al. Perturb-seq: Dissecting molecular circuits with scalable single cell RNA profiling of pooled genetic screens. Cell.2016 Dec 15; 167(7): 1853–1866.e17), CROP-seq (See, e.g., Datlinger et al. Pooled CRISPR screening with single-cell transcriptome readout Nat Methods.2017 Mar;14(3):297-301), CRISP-seq (See, e.g., Jaitin et al. Dissecting Immune Circuits by Linking CRISPR-Pooled Screens with Single-Cell RNA-Seq Cell.2016 Dec 15;167(7):1883-1896.e15), Expanded CRISPR-compatible CITE-seq (ECCITE-seq) (See, e.g., Mimitou et al. Multiplexed detection of proteins, transcriptomes, clonotypes and CRISPR perturbations in single cells. Nat Methods.2019 May;16(5):409-41), and cellular indexing of transcriptomes and epitopes-seq (CITE-seq) (See, e.g., Stoeckius et al. Simultaneous epitope and transcriptome measurement in NYG-LIPP-215.PCT single cells. Nat Methods.2017 Sep;14(9):865-868). As used herein, “reverse transcription” refers to the process of copying the nucleotide sequence of an RNA molecule into a DNA molecule. Reverse transcription can be done by contacting an RNA template with an RNA-dependent DNA polymerase, also known as a reverse transcriptase. A reverse transcriptase is a DNA polymerase that transcribes single-stranded RNA into single-stranded DNA. Depending on the polymerase used, the reverse transcriptase can also have RNase H activity for subsequent degradation of the RNA template. As used herein, “complementary DNA” or “cDNA” can refer to a synthetic DNA reverse transcribed from RNA through the action of a reverse transcriptase. The cDNA may be single- stranded or double-stranded and can include strands that have either or both of a sequence that is substantially identical to a part of the RNA sequence or a complement to a part of the RNA sequence. In one embodiment, the target polynucleotide is DNA or RNA and the modulation or perturbation comprises inhibition, cutting, editing or modulation of protein expression of the target RNA. Target RNAs include, without limitation, mRNA, viral RNAs, noncoding RNAs, miRNA, piRNAs, circRNAs, synthetic RNA and other RNAs. Drug Screen The methods described herein include determining a pattern of differences in the endogenous mRNAs in a cell treated with a drug and correlating them with differences in the endogenous mRNAs due to the sequence specific perturbations performed in the pooled CRISPR screen described above. Thus, in addition to scRNA-seq data from the pooled CRISPR screen, the methods also include analysis of scRNA-seq data from a drug screen. In certain embodiments, the method includes treating one or more cells in a population of cells with a drug. As used herein, the term “drug” refers to small molecules or biomolecules that alter, inhibit, activate, or otherwise affect biological or chemical events of one or more targets. The drug may also be referred to as a compound, for example, when multiple compounds (such as a library) are being screened for target elucidation. The drug is preferably contacted first with the cell for a predetermined time period. Any suitable amount of drug may be used, depending on the potency of the test compound and toxicity. NYG-LIPP-215.PCT For compound library screening, the concentration of test compounds used typically ranges from about 1 to about 1000 micromolar, usually about 10-100 micromolar. Any suitable incubation time may be used. Generally, the time period for incubation of a test compound with cells typically ranges from about 1 minute to about 96 hours, usually about 24 hours to about 72 hours. For the incubation period, the suitable temperature and pH ranges will depend on the requirements of the cell. Generally, most cells are cultured at pH 7.4 and at 37° C., but most cells are viable, at least temporarily, over a range of temperature of 4 to 42° C. and at pH ranges of about 6 to 9. Thereafter, scRNA-seq is performed on cells treated with the drug, and control cells that have not been treated. Single-cell RNA sequencing techniques are known in the art. See, Jovic D, Liang X, Zeng H, Lin L, Xu F, Luo Y. Single-cell RNA sequencing technologies and applications: A brief overview. Clin Transl Med.2022 Mar;12(3):e694. doi: 10.1002 / ctm2.694. PMID: 35352511; PMCID: PMC8964935, which is incorporated herein by reference. Using the single-cell readouts, the drug perturbation signature – or pattern of differences in the endogenous mRNAs in the drug-treated cell versus a non-treated cell – is learned. Computational Model The methods further include comparing measured differences in the endogenous mRNAs that are correlated to the sequence specific perturbations detected for each single cell with differences in endogenous mRNAs in a cell treated with the drug. One or more targets, or potential targets, of the drug is identified when a pattern of differences in the endogenous mRNAs in the cell treated with the drug correlates with differences in the endogenous mRNAs that are correlated to the sequence specific perturbations. In certain embodiments, the method includes determining the differences in the endogenous mRNAs that are correlated to the sequence specific perturbations on the single cells by applying a computational model. The computational model may account for one or more of i) mRNA copy number in each single cell, ii) whether the perturbation actually perturbed the cell, iii) the presence of subpopulations of either different cells or cell states, iv) analysis of cells in the population of cells without any perturbation, and v) noise. In one aspect, is provided a computational model. In certain embodiments, the framework has three main components: a multi-dataset latent factor model that infers robust representations of perturbation responses; a high-dimensional Bayesian regression approach, NYG-LIPP-215.PCT adapted from statistical genetics [1], that deconvolutes treatment signals in terms of these perturbation representations; and a calibration procedure that quantifies and reports meaningful summaries of uncertainty in these results. Hence, given appropriate drug and perturbation screen data, our framework outputs so-called credible sets that describe groups of targets likely to be affected by the drug, and the associated levels of certainty. Below, we detail each of the three components further. Multi-dataset latent factor model. The first component is a latent factor model that jointly describes shared and unique sources of variation, building on general-purpose multi- study factor models that we and others have previously developed [2, 3]. In particular, in the perturbation screen data, let Y0represent a g × n0expression matrix (or other read-out) of g genes in the n0cells that did not receive a perturbation, and let Ykrepresent g × nkexpression matrices for those cells that received each perturbation k in 1, ... , K. We model these data as Y0= Φf0+ E0Yk= Φfk+ Λkℓk+ Ek.Here, Φ is a g × m matrix of latent factor loadings encoding shared variation, including cell cycle and other aspects of baseline heterogeneity; Λkis a g × 1 latent factor loading describing the effect of perturbation k; and E0, Ekrepresent noise. Under this model, the perturbation factors Λkcan thus be disentangled from both baseline heterogeneity and random noise, and variation in response is accounted for via cell- specific factor scores ℓk. Each Λkcan consequently be interpreted as a denoised representation of how perturbation k affects the expression of each gene. We fit this model by first performing principal components analysis (PCA) on Y0in order to estimate the principal components ^�^^^ . We then reconstruct each Ykusing ^�^^^ , residualize out this reconstruction, and perform a rank-1 PCA on the resulting residual matrix to estimatethe perturbation effect Λ̂ k .For the drug screen data, we use an analogous model capturing shared variation between untreated cells ^^^^^^^^^^^^and treated cells YD, i.e. ii) Yk= Φfk+ Λkℓk+ Ek,This allows us to estimate the expression variation in the treated cells that is attributable NYG-LIPP-215.PCT to factors ΦDobserved in the untreated cells. The remaining residual expression r can then be interpreted as the expression component attributable to treatment and random noise. Similar toabove, we fit this model by performing PCA on YDto estimate then subtract out the dataexplained by from YDto obtain the residualized expression rˆ.High-dimensional Bayesian regression. After obtaining the residualized treatmenteffect rˆ and the perturbation effects Λˆ1, . . . , ΛˆK , we employ a high-dimensional Bayesianregression approach to explicitly deconvolute the treatment effect into contributions from one or more perturbation effects. We can then infer these as the drug’s targets. However, a naive implementation of this regression would not be able to accommodate the potential sparsity in these effects, where only one or a few perturbations might be relevant out of tens, hundreds, or even, at the genome-wide scale, thousands of assessed perturbations. Moreover, standard regression techniques fail in the presence of highly correlated features (here, Λs) and cannot quantify uncertainty with the nuance required in this application. To address this, we adapt a recently developed Bayesian regression approach from statistical genetics, the “sum of single effects” model or SuSiE [1], that encodes sparsity and summarizes uncertainty among correlated effects through credible sets. Under this approach, when the exact target cannot be identified, our framework can make direct probabilistic statements about which sets of related perturbations are likely to contain the true target(s). This is done by decomposing the coefficient vector βi, which describes the contributions of each perturbation effect Λkto cell i, into a sum of L “single-effect” coefficient vectors βilthat each only have one non-zero entry, i.e.: ^^^^^^^^^^^^ = ^^^^^^^^^^^^^^^^with priors specified as: Multinomial(1, π) bl∼ Normal(0, σ0l) ^i∼ Normal(0, σ). NYG-LIPP-215.PCT As a result, we can report up to L credible sets of any size characterizing the uncertainty with which correlated perturbations are likely to be responsible for each non-zero entry, as well as the overall probability that each credible set contains at least one true effect. For example, we can make statements such as: “The drug inhibits at least one of the closely related targets { A, B, C } with relative probabilities of 80%, 10%, and 5%, and at least one of the closely related targets { D, E } with indistinguishable probabilities of 50% and 50%.” Uncertainty calibration procedure. The probabilities produced by the Bayesian regression approach are likely to not be appropriately calibrated. This is because the estimated probabilities are computed under the assumption that the predictors Λ are known exactly; while this is reasonably true for the analogous predictors in the original statistical genetics context, in our case we often have large uncertainties around the estimates Λ̂ due to small numbers of available cells and / or heterogeneity in knock-down strength. To address this problem, we model each Λkas arising from a multivariate normal distribution with diagonal covariance. We fit the parameters of this distribution by casting the PCA decomposition in a linear regression framework. In particular, after performing the rank-1 PCA for each perturbation k to estimate Λk, we hold the PC scores ℓkfixed and estimate the means and variances of each entry of Λk by treating them as coefficients in feature-wise linear regressions. This yields simple and closed-form expressions that can be easily computed. We then update the uncertainty estimates and credible sets reported by SuSiE by marginalizing out the Λs according to these distributions from the posterior inclusion probabilities (PIPs). We do so empirically through a sampling approach, in which we repeatedly draw the Λs from their distributions, re-compute the PIPs, and report the mean across all samples. Our approach offers two critical advantages over existing technology. First, our approach is the first to leverage single-cell drug and perturbation screens in order to deconvolute drug targets based on high- resolution, comprehensive transcriptional data. Existing technology instead uses simpler measures of binding or read-outs of proxy markers, which only offer limited characterizations of drug response and specificity. Second, our computational framework uses statistical modeling in innovative ways to address challenges of noise and heterogeneity, and to produce direct, interpretable quantifications of uncertainty. Unless defined otherwise in this specification, technical and scientific terms used herein NYG-LIPP-215.PCT have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs and by reference to published texts, which provide one skilled in the art with a general guide to many of the terms used in the present application. As used throughout this specification and the claims, the terms “comprising”, “containing”, “including”, and its variants are inclusive of other components, elements, integers, steps and the like. Conversely, the term “consisting” and its variants are exclusive of other components, elements, integers, steps and the like. It is to be noted that the term “a” or “an”, refers to one or more, for example, “RNA target”, is understood to represent one or more RNA target(s). As such, the terms “a” (or “an”), “one or more,” and “at least one” is used interchangeably herein. As used herein, the term “about” means a variability of plus or minus 10% from the reference given, unless otherwise specified. Specific Embodiments 1. A method for determining one or more targets of a drug, the method comprising: (a) introducing one or more perturbation constructs encoding for one or more sequence specific perturbations to a plurality of cells in a population of cells, wherein each cell in the plurality of cells receives at least 1 perturbation construct and wherein each perturbation construct comprises a CRISPR guide RNA (gRNA) sequence, a sequence that identifies the perturbation, and a reverse-transcription handle; (b) detecting endogenous mRNAs and the sequence that identifies the perturbation for each single cell in the plurality of cells using single cell RNA-seq; wherein each sequence identifying the perturbation is a unique perturbation barcode that identifies the gRNA introduced to a cell; and (c) comparing measured differences in the endogenous mRNAs that are correlated to the sequence specific perturbations detected for each single cell with differences in endogenous mRNAs in a cell treated with the drug, (d) identifying one or more targets of the drug when a pattern of differences in the endogenous mRNAs in the cell treated with the drug correlates with differences in the endogenous mRNAs that are correlated to the sequence specific perturbations. NYG-LIPP-215.PCT 2. The method of embodiment 1, wherein the guide RNA is an sgRNA and wherein step (a) further comprises introducing a CRISPR-Cas9 enzyme. 3. The method of embodiment 1, wherein the guide RNA is a crRNA and wherein step (a) further comprises introducing a CRISPR-Cas13 enzyme. 4. The method of any one of embodiments 1 to 3, wherein the identifying step further comprises determining the differences in the endogenous mRNAs that are correlated to the sequence specific perturbations on the single cells by applying a computational model. 5. The method of embodiment 4, wherein the computational model accounts for one or more of i) mRNA copy number in each single cell, ii) whether the perturbation actually perturbed the cell, iii) the presence of subpopulations of either different cells or cell states, iv) analysis of cells in the population of cells without any perturbation, and v) noise. 6. The method of embodiment 4 or 5, wherein the computational model utilizes matrix decomposition or dimensional reduction or both. 7. The method of any one of embodiments 4 to 6, wherein the computational model incorporates the formula:i) Y0= Φf0+ E0ii) Yk= Φfk+ Λkℓk+ Ek,where Y0represents a g × n0expression matrix of g genes in the n0cells that did not receive a perturbation; Ykrepresents g × nkexpression matrices for those cells that received each perturbation k in 1, ... , K; Φ is a g × m matrix of latent factor loadings encoding shared variation, including cell cycle and other aspects of baseline heterogeneity; Λkis a g × 1 latent factor loading describing the effect of perturbation k; ℓkis a cell-specific factor score; and E0, Ekrepresent noise. NYG-LIPP-215.PCT 8. The method of any one of embodiments 4 to 7, wherein the model is fit by i) performing principal components analysis (PCA) on Y0to estimate the principal components ^�^^^ ; ii) reconstructing Ykusing ^�^^^ , residualizing out the reconstruction, and performing a rank-1 PCA on the resulting residual matrix to estimate the perturbation effect Λ̂ k .9. The method of any one of embodiments 4 to 8, wherein differences in endogenous mRNAs in a cell treated with the drug are determined using the computational model. 10. The method of any one of embodiments 4 to 9, wherein the computational model incorporates the formula:^^^^^^^^ = Φ^^^^^^^^^^^ ^^^^0 0^ + ^^^^0^^^^^^^^ = Φ^^^^^^^^^^^^ + ^^^^where ^^^0^^^^^represents a g × n0expression matrix of g genes in the n0cells that did not receive the drug, g × n expression matrix for those cells received the drug, Φ^^^^is a g × m matrix of latent factor loadings encoding shared variation in untreated cells, residual expression r is expression attributable to treatment and random noise. 11. The method of embodiment 10, wherein the model is fit by i) performing principal components analysis (PCA) on ^^^^^^^^to estimate principal components ii) subtracting ^^�^^^^^^from ^^^^^^^^to generate residualized expression ^�^^^.12. The method of any one of embodiments 4 to 11, wherein ^�^^^ ,^�^^^, ^�^^^^^^^ , are used todeconvolute the treatment effect of the drug into contributions from one or more perturbation effects, thereby identifying the one or more targets of the drug. 13. The method of embodiment 12, wherein the deconvolution is performed using Bayesian regression. NYG-LIPP-215.PCT 14. A method for determining one or more credible targets of a drug, the method comprising: (a) introducing one or more perturbation constructs encoding for one or more sequence specific perturbations to a plurality of cells in a population of cells, wherein each cell in the plurality of cells receives at least 1 perturbation construct and wherein each perturbation construct comprises a CRISPR guide RNA (gRNA) sequence, a sequence that identifies the perturbation, and a reverse-transcription handle; (b) detecting endogenous mRNAs and the sequence that identifies the perturbation for each single cell in the plurality of cells using single cell RNA-seq; wherein each sequence identifying the perturbation is a unique perturbation barcode that identifies the gRNA introduced to a cell; and (c) comparing measured differences in the endogenous mRNAs that are correlated to the sequence specific perturbations detected for each single cell with differences in endogenous mRNAs in a cell treated with the drug, (d) identifying one or more credible targets of the drug. 15. The method of embodiment 14, further wherein (d) identifying one or more credible targets L of the drug comprises using the following formula: ^^^^^^^^^^^^ = ^^^^^^^^^^^^^^^^wherein βidescribes the contributions of each perturbation effect Λkto cell i, into a sum of L “single-effect” coefficient vectors βil,with priors specified as:γl∼ Multinomial(1, π)bl∼ Normal(0, σ0l) ^i∼ Normal(0, σ). NYG-LIPP-215.PCT EXAMPLES The following examples disclose both general and specific embodiments of the disclosed compositions and methods described herein, which should be construed to encompass any and all variations that become evident as a result of the teaching provided herein. Example 1: Application of Computational Model We show two demonstrations of our approach. First, we learned perturbation effects from an existing perturbation screen dataset [M.H. Kowalski, et al (2023). CPA-Perturb-seq: Multiplexed single-cell characterization of alternative polyadenylation regulators, bioRxiv], then classified a set of withheld cells from these data to evaluate our approach in a setting with known ground truth. We conducted this analysis after filtering out “escaping” cells that did not exhibit knock-down, and used a subset of 18 perturbations that produced strong transcriptional effects and for which we had at least 50 reference cells. We assessed performance in terms of sensitivity (i.e. out of withheld cells receiving each perturbation, what proportion were assigned the correct perturbation within their credible set) and precision (i.e. out of all cells assigned a single-member credible set with each perturbation, what proportion truly received that perturbation). The results are shown in FIG.3A with a median sensitivity of 92% and a median precision of 99%. Furthermore, we demonstrate meaningful quantifications of uncertainty (FIG. 3B) by producing PIP estimates that track with expression gradients across closely related perturbations. Example 2: Validation on perturbation screen data To assess our approach in a setting with known ground truth, we estimated perturbation effects from one replicate of a perturbation screen [M.H. Kowalski, et al (2023). CPA-Perturb- seq: Multiplexed single-cell characterization of alternative polyadenylation regulators, bioRxiv] and successfully made classifications in the other (median sensitivity 89%, median precision 96%). We further demonstrated well-calibrated assessments of uncertainty among closely related perturbations, e.g. complex members. FIG.4A and FIG.4B. Example 3: Target identification from 3,900+ perturbations in positive control drugs We applied our approach to a large-scale target deconvolution setting by estimating perturbation effects from 3,900+ perturbations in Genome-Wide Perturb-Seq (GWPS) [J.M. NYG-LIPP-215.PCT Replogle, et al (2022). Mapping information-rich genotype-phenotype landscapes with genome-scale Perturb-seq, Cell.185(14), 2559–2575], and classifying cells treated with one of four positive control drugs from sci-plex [S.R. Srivatsan, et al (2020). Massively multiplex chemical transcriptomics at single-cell resolution, Science.367(6473), 45–51.]. For each drug, the most frequent highest-probability match was the true target. This classification is particularly challenging due to previously characterized discrepancies between 10X transcriptional readouts (with which the GWPS data is profiled) and sci-RNA-seq read-outs (with which the sciplex data is profiled) [S. Lechner, et al (2022). Target deconvolution of HDAC pharmacopoeia reveals MBLAC2 as common off-target, Nat. Chem. Biol.18(8), 812–820.]. However, without using any prior knowledge, our framework correctly identified BCR as the highest-probability target out of all of these perturbed genes (FIG.5A). FIG.5B is a heatmap showing the genes driving the classification. These data show that, under our approach, drug and perturbation screen data can be leveraged together to find the correct drug targets even in this ultra-high-dimensional, cross- technology setting. Example 4: Ambiguity in HDAC inhibitors motivates further experiments The mechanisms of HDAC inhibitors have been recently called into question [A. Lin, et al (2019). Off-target toxicity is a common mechanism of action of cancer drugs undergoing clinical trials, Sci. Transl. Med.11(509), eaaw8412, S. Lechner, et al (2022). Target deconvolution of HDAC pharmacopoeia reveals MBLAC2 as common off-target, Nat. Chem. Biol.18(8), 812–820.] When classifying with 3,900+ GWPS perturbation effects, we found that some pan- HDACi-treated cells matched to known HDAC regulators or complex members, but not to any HDACs themselves. Very few matches were found for class II selective HDACis. FIG.6. Our next step is to conduct a targeted HDAC perturbation screen, including combinatorial knockdowns and potential off-targets, and a comprehensive HDACi drug screen covering a wider dose range. By applying our framework to these data, we hope to elucidate mechanisms across classes of HDAC inhibitors. Example 5: Identifying multiple effects in a combinatorial perturbation screen NYG-LIPP-215.PCT We illustrate our approach’s ability to correctly identify multiple simultaneous effects by applying it to a combinatorial perturbation screen (Wessels et al., 2022). In particular, we learned perturbation effects (Λs) using the subset of cells that received a single-gene perturbation, then we deconvolved cells that received dual-gene perturbations in terms of these effects. We carried out this analysis at both the single-cell level and at the sample-level; in sample-level analysis, we perform one joint deconvolution for each group of cells of interest (here, for example, all cells that received the same dual-gene perturbation). In both analyses, we were able to correctly recover both effects of the dual-gene perturbations for the cells or groups of cells showing dual-gene phenotypes (FIG.7). For example, when considering cells that simultaneously received an EP300 knockdown and a FLI1 knockdown, our approach correctly returned one credible set containing EP300 and one credible set containing FLI1. See, Ref.11. Example 6: Identifying multiple effects in a lncRNA perturbation screen We illustrate our approach’s ability to correctly identify multiple simultaneous effects in a biological application by mapping data from a lncRNA perturbation screen (Liang et al., 2024) to a set of perturbation effects comprising protein-coding genes in the same screen, as well as thousands of perturbations from GenomeWide Perturb-Seq (Replogle et al., 2022). When applied at the sample-level, there were four lncRNAs that were deconvolved in terms of multiple effects: one credible set containing HNRNPF, and one credible set containing one or more genes related to the PAXT complex (FIG.8A and 8B). This suggests that this set of lncRNAs may regulate the PAXT complex, and provides insight into their mechanism. See, Refs.12 and 13. REFERENCES 1. Gao Wang, Abhishek Sarkar, Peter Carbonetto, and Matthew Stephens. A simple new approach to variable selection in regression, with application to genetic fine mapping. Journal of the Royal Statistical Society Series B: Statistical Methodology, 82(5):1273– 1300, 2020. 2. Roberta De Vito, Ruggero Bellio, Lorenzo Trippa, and Giovanni Parmigiani. Bayesian multistudy factor analysis for high-throughput biological data. The annals of applied statistics, 15(4):1723–1741, 2021. 3. Isabella N Grabski, Roberta De Vito, Lorenzo Trippa, and Giovanni Parmigiani. NYG-LIPP-215.PCT Bayesian combinatorial multistudy factor analysis. The annals of applied statistics, 17(3):2212, 2023. 4. Madeline H Kowalski, Hans-Hermann Wessels, Johannes Linder, Saket Choudhary, Austin Hartman, Yuhan Hao, Isabella Mascio, Carol Dalgarno, Anshul Kundaje, and Rahul Satija. Cpa-perturb-seq: Multiplexed single-cell characterization of alternative polyadenylation regulators. bioRxiv, 2023. 5. Joseph M Replogle, Reuben A Saunders, Angela N Pogson, Jeffrey A Hussmann, Alexander Lenail, Alina Guna, Lauren Mascibroda, Eric J Wagner, Karen Adelman, Gila Lithwick-Yanai, et al. Mapping information-rich genotype-phenotype landscapes with genome-scale perturb-seq. Cell, 185(14):2559– 2575, 2022. 6. Sanjay R Srivatsan, Jose L McFaline-Figueroa, Vijay Ramani, Lauren Saunders, Junyue Cao, Jonathan Packer, Hannah A Pliner, Dana L Jackson, Riza M Daza, Lena Christiansen, et al. Massively multiplex chemical transcriptomics at single-cell resolution. Science, 367(6473):45–51, 2020. 7. Jiarui Ding, Xian Adiconis, Sean K Simmons, Monika S Kowalczyk, Cynthia C Hession, Nemanja D Mar- janovic, Travis K Hughes, Marc H Wadsworth, Tyler Burks, Lan T Nguyen, et al. Systematic comparison of single-cell and single-nucleus rna-sequencing methods. Nature biotechnology, 38(6):737–746, 2020. 8. Lin, C.J. Giuliano, [...], J.M. Sheltzer (2019). Off-target toxicity is a common mechanism of action of cancer drugs undergoing clinical trials, Sci. Transl. Med.11(509), eaaw8412. 9. Lin, C.J. Giuliano, N.M. Sayles, J.M. Sheltzer (2017). CRISPR / Cas9 mutagenesis invalidates a putative cancer dependency targeted in on-going clinical trials, Elife.6, e24179. 10. S. Lechner, M.I.P. Malgapo, [...], G. Médard (2022). Target deconvolution of HDAC pharmacopoeia reveals MBLAC2 as common off-target, Nat. Chem. Biol.18(8), 812– 820. 11. Wessels, Hans-Hermann, et al. ”Efficient combinatorial targeting of RNA transcripts in single cells with Cas13 RNA Perturb-seq.” Nature methods 20.1 (2023): 86-94. 12. Liang, Wen-Wei, et al. “Transcriptome-scale RNA-targeting CRISPR screens reveal essential lncRNAs in human cells.” Cell 187.26 (2024): 7637-7654. NYG-LIPP-215.PCT 13. Replogle, Joseph M., et al. ”Mapping information-rich genotype-phenotype landscapes with genome-scale Perturb-seq.” Cell 185.14 (2022): 2559-2575. All patent and non-patent publications cited in this specification are incorporated herein by reference in their entireties. While the invention has been described with reference to particular embodiments, it will be appreciated that modifications can be made without departing from the spirit of the invention. Such modifications are intended to fall within the scope of the appended claims.

Claims

NYG-LIPP-215.PCT What is claimed is:

1. A method for determining one or more targets of a drug, the method comprising: (a) introducing one or more perturbation constructs encoding for one or more sequence specific perturbations to a plurality of cells in a population of cells, wherein each cell in the plurality of cells receives at least 1 perturbation construct and wherein each perturbation construct comprises a CRISPR guide RNA (gRNA) sequence, a sequence that identifies the perturbation, and a reverse-transcription handle; (b) detecting endogenous mRNAs and the sequence that identifies the perturbation for each single cell in the plurality of cells using single cell RNA-seq; wherein each sequence identifying the perturbation is a unique perturbation barcode that identifies the gRNA introduced to a cell; and (c) comparing measured differences in the endogenous mRNAs that are correlated to the sequence specific perturbations detected for each single cell with differences in endogenous mRNAs in a cell treated with the drug, (d) identifying one or more targets of the drug when a pattern of differences in the endogenous mRNAs in the cell treated with the drug correlates with differences in the endogenous mRNAs that are correlated to the sequence specific perturbations.

2. The method of claim 1, wherein the guide RNA is an sgRNA and wherein step (a) further comprises introducing a CRISPR-Cas9 enzyme.

3. The method of claim 1, wherein the guide RNA is a crRNA and wherein step (a) further comprises introducing a CRISPR-Cas13 enzyme.

4. The method of any one of claims 1 to 3, wherein the identifying step further comprises determining the differences in the endogenous mRNAs that are correlated to the sequence specific perturbations on the single cells by applying a computational model.

5. The method of claim 4, wherein the computational model accounts for one or more of i) mRNA copy number in each single cell, ii) whether the perturbation actually perturbed the cell, iii) the presence of subpopulations of either different cells or cell states, iv) analysis of cells inNYG-LIPP-215.PCT the population of cells without any perturbation, and v) noise.

6. The method of claim 4, wherein the computational model utilizes matrix decomposition or dimensional reduction or both.

7. The method of claim 4, wherein the computational model incorporates the formula: i) Y0= Φf0+ E0ii) Yk= Φfk+ Λkℓk+ Ek,where Y0represents a g × n0expression matrix of g genes in the n0cells that did not receive a perturbation; Ykrepresents g × nkexpression matrices for those cells that received each perturbation k in 1, ... , K; Φ is a g × m matrix of latent factor loadings encoding shared variation, including cell cycle and other aspects of baseline heterogeneity; Λkis a g × 1 latent factor loading describing the effect of perturbation k; ℓkis a cell-specific factor score; and E0, Ekrepresent noise.

8. The method of claim 4, wherein the model is fit by i) performing principal components analysis (PCA) on Y0to estimate the principal components ^�^^^ ; ii) reconstructing Ykusing ^�^^^ , residualizing out the reconstruction, and performing arank-1 PCA on the resulting residual matrix to estimate the perturbation effect Λ̂ k .

9. The method of claim 4, wherein differences in endogenous mRNAs in a cell treated with the drug are determined using the computational model.

10. The method of claim 4, wherein the computational model incorporates the formula: ^^^^^^^^ = Φ^^^^^^^^^^^^ + ^^^ ^^^^0 0 ^0^^^^^^^^ = Φ^^^^^^^^^^^^ + ^^^^where ^^^0^^^^^represents a g × n0expression matrix of g genes in the n0cells that did not receive the drug, ^^^^^^^^represents g × n expression matrix for those cells received the drug,is a g × m matrix of latent factor loadings encoding shared variation in untreated cells, residual expression r is expression attributable to treatment and random noise.NYG-LIPP-215.PCT 11. The method of claim 10, wherein the model is fit by i) performing principal components analysis (PCA) on ^^^^^^^^to estimate principal components ^^�^^^^^^; ii) subtracting ^^�^^^^^^from ^^^^^^^^to generate residualized expression ^�^^^.

12. The method of claim 4, wherein ^�^^^ ,^�^^^, ^�^^^^^^^ , are used to deconvolute the treatment effect ofthe drug into contributions from one or more perturbation effects, thereby identifying the one or more targets of the drug.

13. The method of claim 12, wherein the deconvolution is performed using Bayesian regression.

14. A method for determining one or more credible targets of a drug, the method comprising: (a) introducing one or more perturbation constructs encoding for one or more sequence specific perturbations to a plurality of cells in a population of cells, wherein each cell in the plurality of cells receives at least 1 perturbation construct and wherein each perturbation construct comprises a CRISPR guide RNA (gRNA) sequence, a sequence that identifies the perturbation, and a reverse-transcription handle; (b) detecting endogenous mRNAs and the sequence that identifies the perturbation for each single cell in the plurality of cells using single cell RNA-seq; wherein each sequence identifying the perturbation is a unique perturbation barcode that identifies the gRNA introduced to a cell; and (c) comparing measured differences in the endogenous mRNAs that are correlated to the sequence specific perturbations detected for each single cell with differences in endogenous mRNAs in a cell treated with the drug, (d) identifying one or more credible targets of the drug.

15. The method of claim 14, further wherein (d) identifying one or more credible targets L of the drug comprises using the following formula:NYG-LIPP-215.PCT^^^^^^^^^^^^ = ^^^^^^^^^^^^^^^^wherein βidescribes the contributions of each perturbation effect Λkto cell i, into asum of L “single-effect” coefficient vectors βil,with priors specified as:γl∼ Multinomial(1, π)bl∼ Normal(0, σ0l) ^i∼ Normal(0, σ).

Citation Information

Patent Citations

  • Exploiting genomics in the search for new drugs

    US20030180774A1

  • High-throughput screening system for identification of novel drugs and drug targets

    US20230083853A1

  • Methods for drug target screening

    US7122312B1

  • Methods and compositions for combinatorial targeting of the cell transcriptome

    WO2022187262A1