Multi-reporter system for chemical screening
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- OCTANT INC
- Filing Date
- 2024-07-18
- Publication Date
- 2026-05-27
AI Technical Summary
Current pharmaceutical screening methods are limited in their ability to simultaneously assess the activity of test agents against multiple signaling pathways, target variants, and off-target effects, leading to inefficiencies in drug discovery.
The development of a multi-reporter system using engineered cell lines with different reporter constructs, allowing for the simultaneous determination of test agent activity against specific targets, variants, and downstream effectors, while also identifying off-target effects.
This system enhances the efficiency of pharmaceutical screening by enabling the parallel assessment of multiple biological activities, improving statistical significance, and reducing the risk of nonspecific or off-target effects.
Smart Images

Figure US2024038606_23012025_PF_FP_ABST
Abstract
Description
MULTI-REPORTER SYSTEM FOR CHEMICAL SCREENINGCROSS-REFERENCE
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 528,030 filed on July 20, 2023, which is incorporated herein by reference in its entirety.BACKGROUND
[0002] Understanding signaling pathways and cellular functions and identifying test agents that modulate a given signaling pathway or cellular function is an important goal in pharmaceutical sciences and drug discovery.SUMMARY
[0003] The methods and systems described herein provide cell-based screening assays. These systems and methods include several improvements over previous screening methods. The systems and methods described herein may allow for multiplexing to screen for multiple signaling pathways or cellular functions in parallel. These systems and methods allow identification of off-target effects or identification of compounds that agonize or antagonize specific target variants, signaling pathways, and / or biological functions. The systems and methods described herein include multiple indexes to allow for improvements in statistical significance. The methods and systems also allow for identification or avoidance of nonspecific or off-target effects by allowing screening of molecules that agonize or antagonize the same target but signal through different downstream effectors. Overall, the methods and systems described herein increase the efficiency of pharmaceutical screening and understanding of biological system. The methods and systems, described herein, allow, simultaneously or substantially simultaneously, the determination of a certain test agents activity against a target (e.g., a heterologous polypeptide), variants of the target, different downstream promoters that may be activated by the target, and off-target effects of the test agent simultaneously in a single reaction vessel (e g., a well of a multi-well plate). The methods described herein are capable of such determination for thousands of compounds simultaneously. Described herein are methods and systems of screening and identifying test agents that regulate target activity, either positively (as an agonist) or negatively (as an antagonist).
[0004] Described herein is a system comprising a plurality of engineered cell lines in a partition, wherein the plurality of engineered cell lines comprises a first engineered cell line and a second engineered cell line, wherein the first engineered cell line comprises a first reporter construct and the second engineered cell line comprises a second reporter construct, wherein, the first reporterconstruct and the second reporter construct are selected from: (a) a constitutive promoter operatively coupled to a reporter gene; (b) a first inducible promoter operatively coupled to a reporter gene; or (c) a second inducible promoter operatively coupled to a reporter gene; wherein the first reporter construct is different from the second reporter construct; and wherein the first reporter construct and the second reporter construct are independently readable. In certain embodiments, the plurality of engineered cell lines further comprises a third engineered cell line, wherein the third engineered cell line comprises a third reporter construct that is different from the first reporter construct and the second reporter construct, wherein the third reporter construct is selected from: (a) a constitutive promoter operatively coupled to a reporter gene; (b) a first inducible promoter operatively coupled to a reporter gene; or (c) a second inducible promoter operatively coupled to a reporter gene. In certain embodiments, one or more of the first engineered cell line, second engineered cell line, third engineered cell line, or combinations thereof further comprises a first heterologous polypeptide. In certain embodiments, one or more of the first engineered cell line, second engineered cell line, third engineered cell line, or combinations thereof further comprise a second heterologous polypeptide, wherein the second heterologous polypeptide comprises at least one amino acid alteration relative to the first heterologous polypeptide. In certain embodiments, the second heterologous polypeptide comprises less than 10, less than 5, less than 3, or less than 2 amino acid alterations relative to the first heterologous polypeptide. In certain embodiments, the second heterologous polypeptide comprises more than 10, more than 20, more than 50, more than 100, more than 500, or more than 1000 amino acid alterations relative to the first heterologous polypeptide. In certain embodiments, the first heterologous polypeptide, the second heterologous polypeptide, or both are coupled to a transcription factor. In certain embodiments, the transcription factor comprises one or more of aGal4, PPR1, Lac9, zinc finger, or LexA DNA binding domain. In certain embodiments, the transcription factor comprises one or more of aVP64, p65, RoTev, or Rta DNA activating domain. In certain embodiments, the first reporter construct and the second reporter construct are independently readable. In certain embodiments, the third reporter construct is independently readable from the first reporter construct and / or the second reporter construct reporter. In certain embodiments, the first inducible promoter, the second inducible promoter, or both are configured to be activated by the first heterologous polypeptide, a signal from the first heterologous polypeptide, a transcription factor coupled to the first heterologous polypeptide, or any combination thereof. In certain embodiments, the first engineered cell line, second engineered cell line, or third engineered cell line comprises an additional reporter constructselected from: (a) a constitutive promoter operatively coupled to a reporter gene; (b) a first inducible promoter operatively coupled to a reporter gene; or (c) a second inducible promoter operatively coupled to a reporter gene, wherein the additional reporter construct is different from the first reporter construct and the second reporter construct. In certain embodiments, one or more of the first reporter construct, the second reporter construct, or the third reporter construct are integrated into the genome of the cell line. In certain embodiments, any one or more of the first engineered cell line, the second engineered cell line, or the third engineered cell line is a eukaryotic cell line. In certain embodiments, the eukaryotic cell line is a mammalian cell line. In certain embodiments, the mammalian cell line is a human cell line. In certain embodiments, the reporter gene encodes a fluorescent protein or a luciferase protein. In certain embodiments, the reporter gene encodes a barcode RNA sequence. In certain embodiments, the reporter gene encodes a fluorescent protein and a barcode RNA sequence or a luciferase protein and a barcode RNA sequence. In certain embodiments, the first heterologous polypeptide or the second heterologous polypeptide is a cell-surface protein. In certain embodiments, the cell-surface protein is a G-protein coupled receptor, a receptor tyrosine kinase, an ion channel, a cytokine receptor, a chemokine receptor, a growth factor receptor, or a cellular adhesion molecule. In certain embodiments, the cell-surface protein is expressed by any one or more of the plurality of engineered cell lines. In certain embodiments, the first heterologous polypeptide or the second heterologous polypeptide is an intracellular protein. In certain embodiments, the intracellular protein is an enzyme, ER transporter, nuclear transporter, intracellular signaling protein, a chaperone, or a transcription factor. In certain embodiments, the plurality of engineered cell lines further comprises a cell that comprises a barcode sequence but does not express the barcode sequence. In certain embodiments, the plurality of engineered cell lines comprises mammalian cells. In certain embodiments, the mammalian cells are human cells. In certain embodiments, the partition is a well of an rc-well plate. In certain embodiments, the / / -well plate is a 96-well plate. In certain embodiments, the n-well plate is a 384-well plate. In certain embodiments, the n-well plate is a 1536-well plate. In certain embodiments, the first reporter construct, the second reporter construct, or the third reporter construct comprises a constitutive promoter operatively coupled to a reporter gene. In certain embodiments, the constitutive promoter is selected from a SV40 promoter, a CMV promoter, an Efl A promoter, aPGKl promoter, an Ubc promoter, a beta actin promoter, a CAG promoter, an Ac5 promoter, a polyhedrin promoter, a TEF1 promoter, a GDS promoter, a CaMV355 promoter, an Ubi protomer, or any combination thereof. In certain embodiments, the first inducible promoter comprises an NF AT promoter, a CRE promoter, a p53promoter, an ISRE promoter, a Gal4-UAS promoter, a Lex A promoter. In certain embodiments, the second inducible promoter is selected from an NF AT promoter, a CRE promoter, a p53 promoter, an ISRE promoter, a Gal4-UAS promoter, a Lex A promoter. In certain embodiments, the first inducible promoter and second inducible promoter are different promoters that mediate signaling through the same intracellular protein or cell-surface protein. In certain embodiments described herein is a method of screening for a compound that regulates a biological activity of any one or more of the plurality of engineered cell lines of any one of the preceding claims, the method comprising contacting the plurality of engineered cell lines with a test agent and measuring an activity of the reporter constructs of the first engineered cell line, the second engineered cell line, third engineered cell line, or any combination thereof. In certain embodiments, a plurality of w-test agents is contacted to the plurality of engineered cell lines that has been divided into at least / / -partitions. In certain embodiments, the plurality of engineered cell lines comprises at least 100, at 1,000, or at least 10,000 different heterologous polypeptides. In certain embodiments, the biological assay is performed in a 96-well, 384-well or a 1536-well plate. In certain embodiments, the first reporter construct and the second reporter construct are present in different engineered cells in the same well of a 96-well, 384-well, or a 1536-well plate. In certain embodiments, the plurality of n-test agents is prepared as a plurality of individual reaction mixtures by reacting a core fragment A comprising a reactive functionality x, wherein x is an amine, aldehyde, boronate, imide, isothiocyanate, carboxylic acid, halide, hydroxyamidine, or thiourea; with a plurality of naive test fragments (y-Ti, y-T2, . . . y-Tn), each test fragment comprising a reactive functionality y, wherein y is a amine, boronate, imide, isothiocyanate, halide, hydroxyamide, thiourea, carboxylic acid or aldehyde and one of a plurality of naive test moieties (T T2, . . . Tn) under reaction conditions sufficient to form a plurality of n test agents (A-Ti, A-T2, . . . A-Tn), wherein each of A-Ti, A-T2, . . . A-Tnis prepared in a well of a test plate. In certain embodiments, x is an amine. In certain embodiments, the amine is a primary or secondary amine. In certain embodiments, x is an aldehyde or carboxylic acid. In certain embodiments, y is an amine. In certain embodiments, the amide is a primary or secondary amine. In certain embodiments, y is a carboxylic acid or aldehyde. In certain embodiments, the reaction conditions comprise a base. In certain embodiments, the reaction conditions comprise a Lewis acid or Bronstead acid. In certain embodiments, the reaction conditions comprise an amide coupling reagent. In certain embodiments, the reaction conditions comprise a palladium reagent. In certain embodiments, the reaction conditions comprise room temperature. In certain embodiments, the reaction conditions comprise a reaction temperature between about roomtemperature and about 80 °C. In certain embodiments, the reacting is performed from about 1 to about 24 hours. In certain embodiments, the reacting is performed from about 6 to about 18 hours. In certain embodiments, the reacting is reversible or irreversible. In certain embodiments, the method comprises an amide coupling, reductive amination, or oxidative addition. In certain embodiments, the method comprises a Buchwald or Suzuki coupling. In certain embodiments, the core moiety A has a mass of from about 150 Da to about 800 Da. In certain embodiments, each test moiety (Tl, T2, . . . Tn) has a mass of from about 80 Da to about 500 Da. In certain embodiments, each test agent (A-Tl , A-T2, . . . A-Tn) has a mass of less than about 1500 Da. In certain embodiments, each test agent (A-Tl, A-T2, . . . A-Tn) has a mass of from about 350 Da to about 800 Da. In certain embodiments, each reaction is performed on a nanoscale. In certain embodiments, the nanoscale is performed in a volume of 50 nL to 500 nL. In certain embodiments, the test plate is a 96-well, 384-well, or 1536-well test plate. In certain embodiments, each naive test moieties (Tl, T2, . . . Tn) is different. In certain embodiments, n is from 2 to 100,000. In certain embodiments, n is from 2 to 25,000. In certain embodiments, n is from 2 to 2,500. In certain embodiments, n is from 2 to 2,000.INCORPORATION BY REFERENCE
[0005] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Various aspects of the disclosure are set forth with particularity in the appended claims. A better understanding of the featuresand advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings below.
[0007] Fig. l shows an exemplary cell interacting with a test agent.
[0008] Fig. 2 shows an exemplary cell comprising at least two reporter constructs interacting with a test agent.
[0009] Fig. 3 shows an exemplary cell comprising at least two reporter constructs interacting with a test agent, where one of the at least two reporter constructs is activated when a heterologous polypeptide is not activated.
[0010] Fig. 4 shows an example well containing one or more cells, where the cells comprise different heterologous polypeptides, response elements and / or reporter constructs that provideunique information on the biology of the heterologous polypeptides, response elements, and or reporter constructs in response to a test agent.
[0011] Figs 5A-5D shows an example configuration for heterologous polypeptides and / or reporter constructs for interrogating the biological function of one or more test agents. Fig. 5A illustrates integrating multiple response elements. Fig. 5B illustrates integrating multiple heterologous polypeptide variants with a response element. Fig. 5C illustrates integrating multiple heterologous polypeptides with multiple response elements and multiple indexes. Fig. 5D illustrates integrating heterologous polypeptides or heterologous polypeptide variants with multiple indexes.
[0012] Fig. 6 shows LCMS peak area distributions for 96 reactions in the benzimidazole-forming exemplary libraries. The pink / red bars show the peak area for the expected benzimidazole product, while the green bars are peak areas of the aldehyde fragment, and the blue bars are the peak areas of the o-phenylenediamine core.
[0013] Fig. 7 displays QC scores for 96 reactions measured by LCMS for benzimidazole formation show that more than 70 of the 96 measured reactions have QC scores >0.5, indicating that about 80% of the reactions that were attempted formed large amounts of the desired benzimidazoles.
[0014] Fig. 8 illustrates certain examples of assays to detect different biological activities, which may be adapted for or performed in addition with the methods and systems described herein.
[0015] Figs. 9A-9G depicts an overview of the multiplexed MAHDS screening platform. Fig. 9A, 9C and 9E depicts the general design of the engineered GPCR signaling circuit. An inducible promoter drives the expression of the receptor of interest. Receptor activation by an agonist stimulates G-protein activation. Fig. 9B, 9D and 9F depicts representative dose response curves to agonists: The activity (depicted as log2 fold change over basal) was fitted to a 4 parameter log-logistic function using the drc package from Ritz et al. (available on GitHub). The EC50 values are provided + / - their standard error. Fig. 9A depicts activated Gs protein stimulates adenylate cyclase, leading to production of cAMP and activation of the CRE promoter by the CREB transcription factor. Fig. 9B depicts the results of the assay in Fig. 9A detected as an increase in barcode count. Fig. 9C depicts activated Gi protein inhibits the activity of (forskolin stimulated) adenylate cyclase, leading to lower production of cAMP and lower activation of the CRE promoter. Fig. 9D depicts the results of Fig. 9C detected as a decrease in barcode count. Fig. 9E depicts activated Gq protein stimulates PLCB, leading to an increase in intracellular calcium concentrations and activation of an engineered NFAT promoter. Fig. 9F depicts theresults of Fig. 9E detected as an increase in barcode count. Fig. 9G depicts the ability of the assay to detect the activity of each receptor to each G-protein reporter is indicated by shading. Reported primary coupling of each receptor is denoted by a single asterisk [*], while the reported secondary coupling is denoted by a double asterisk [**]. Reported couplings are from IUPHARS.
[0016] Fig. 10A-10E depicts the activity (depicted as log2 fold change over basal) of each class of receptors in each library is reported against it's cognate endogenous agonist. Each drugreceptor interaction was fitted to a 4 parameter log-logistic function using the drc package (available on GitHub). The EC50 values are provided + / - their standard error. The reported canonical couplings of each GPCR is listed in parentheses. Responses are broken out into their respective G-protein reporter classes (Gs / CRE in blue, Gi / InvCRE in brown, Gq / NFAT in orange). Fig. 10A depicts acetylcholine activity against the muscarinic receptors. Fig. 10B depicts noradrenaline activity against the adrenergic receptors. Fig. 10C depicts histamine activity against the histaminergic receptors. Fig. 10D depicts dopamine activity against the dopaminergic receptors. Fig. 10E depicts serotonin activity against the serotonergic receptors.
[0017] Fig. 11A-11B depicts the Performance of the MAHDS receptors against endogenous neurotransmitters. FIG. 11A depicts activity of the MAHDS receptors against the endogenous agonists. Shading indicates pEC50 of the dose response. White indicates no response detected. Fig. 11B depicts activity dose responses of ADRB1 and DRD1 in the CRE-Gs reporter against noradrenaline and dopamine in the Gs reporter. EC50's are reported and indicated by dashed line.
[0018] Fig. 12A-12C depicts measuring and validating the activity of 4 antipsychotics. FIG. 12A depicts representative antagonist activity dose response plot of Gs, Gi, and Gq. Fig. 12B depicts the comparison of available reported binding affinities (pKi's) with those calculated from the data using the Cheng-Prusoff equation. Fig. 12C depicts the summary of interactions of each antipsychotic with each receptor. "Hits" indicate when a reported interaction is recapitulated in our assay, "divergent hits" indicate when we observed an interaction that is not reported, "misses" indicate when an interaction is reported but we observed none.
[0019] FIG. 13A-13B depicts antipsychotic activity and receptor selectivity at DRD2 and HTR2A. Fig. 13 A depicts the antagonist dose response curves of each antipsychotic are plotted for DRD2 (Gi) and HTR2A (Gq). Fig. 13B depicts pIC50s for DRD2, HTR2A, and the ratio between HTR2A and DRD2 are listed.
[0020] Fig. 14 depicts the results of screening 34 compounds against MclR, MC3R, MC4R and MC4R.
[0021] Fig. 15A depicts the engineering of activity and potency versus selectivity along the MC4R / 1R axis. Fig. 15B depicts the bias between Gs and Gq signaling.
[0022] FIG. 16 shows a heatmap of the amount of mutated RHO trafficked to the extracellular domain relative to wildtype RHO.DETAILED DESCRIPTION
[0023] Determining what test agents affect certain biological activities but not others, and consequently, which test agents will treat certain diseases while others do not can be a difficult and time consuming process. The methods described herein include a method for assessing the regulation of a reporter construct that reports on a particular biological activity by analyzing a plurality of barcodes expressed from the reporter constructs when performing a biological assay. In doing so, a test agent may be tested against multiple reporter constructs at the same time, and the barcodes that result indicate which reporter constructs the test agent did or did not activate. For example, a test agent may be added to a well to react with plurality of different engineered cell lines comprising a plurality of different reporter constructs, and when the biological assay is performed, expression of the barcodes show that the test agent often not only activated a relevant promoter, but also, for example, activated other reporters that may indicate off-target effects, for example, on different signaling pathways.
[0024] Thus, by performingthe biological assays in well plates or chips with larger numbers of partitions (e.g., wells), the interactions between test agents and biological activity may be studied more efficiently and in greater detail by showing a full array of the interactions, including whether the test agent binds to a specific heterologous polypeptide, whether the test agent binds to a variant of a heterologous polypeptide, whether certain promoters are activated through certain pathways, whether the test agent promotes toxicity in the well, or whether the test agent does not affect any biological activity. In doing so, it can be determined which test agents help treat conditions or diseases more efficiently. Further, this can allow for multiplexed testing of many test agents in parallel. This can be combined with fragment based chemistry to further iterate and develop test agents into candidate molecules (e.g., small molecules, peptides or biologies) that have increased potency, greater specify to agonize or antagonize specific biological functions, or decreased side effects.
[0025] Described herein is a system comprising a plurality of engineered cell lines in a partition, wherein the plurality of engineered cell lines comprises a first engineered cell line and a second engineered cell line, wherein the first engineered cell line comprises a first reporter construct and the second engineered cell line comprises a second reporter construct, wherein, the first reporterconstruct and the second reporter construct are selected from: (a) a constitutive promoter operatively coupled to a reporter gene; (b) a first inducible promoter operatively coupled to a reporter gene; or (c) a second inducible promoter operatively coupled to a reporter gene; wherein the first reporter construct is different from the second reporter construct; and wherein the first reporter construct and the second reporter construct are independently readable.
[0026] Described herein, in some aspects, is a system comprising a plurality of engineered cell lines in a partition, wherein the plurality of engineered cell lines comprises a first engineered cell line and a second engineered cell line, wherein the first engineered cell line comprises a first reporter construct and the second engineered cell line comprises a second reporter construct. In some embodiments, the system comprises the first reporter construct and the second report construct, where the first reporter construct and the second reporter construct are selected from: a constitutive promoter operatively coupled to a reporter gene; a first inducible promoter operatively coupled to a reporter gene; ora second inducible promoter operatively coupled to a reporter gene; and the first reporter construct is different from the second reporter construct. In some embodiments, the first reporter construct and the second reporter construct are independently readable. In some embodiments, the plurality of engineered cell lines further comprises a third engineered cell line, where the third engineered cell line comprises a third reporter construct different from the first reporter construct and the second reporter construct, wherein the third reporter construct is selected from: a constitutive promoter operatively coupled to a reporter gene; a first inducible promoter operatively coupled to a reporter gene; or a second inducible promoter operatively coupled to a reporter gene. In some embodiments, the first engineered cell line, second engineered cell line, or third engineered cell line comprises an additional reporter construct.
[0027] Described herein, in some aspects is a method for screening for a compound (e.g., a ligand described herein) that regulates a biological activity of any one or more of the plurality of engineered cell lines descirbed herein. In some embodiments, the method comprises contacting the plurality of engineered cell lines with a test agent and measuring the activity of the reporter constructs of the first engineered cell line, the second engineered cell line, third engineered cell line, or any combination thereof. In some emboidments, the plurality of engineered cell lines comprises a first heterlogous polypeptide. In some mebodiments, the first heterologus polypeptide can be complexed or contacted with the test agent. In some embodiments, the first heterologous polypeptide is operatively coupled to a transcription factoer, where the contacting between the test agent and the first heterologous polypeptide leads to transcription or expressionof the report gene (e.g., as shown in Fig. 1 and Fig. 5A). In some mbodiments, the plurality of engineered cell lines comprises reporter genes for screening for interaction between a heterolgous polypeptide and a test agent (e g., as shown in Fig. 2, Fig. 3, or Fig. 5B-D).Engineered Cell Lines
[0028] Described herein are systems, methods, and assays utilizing a plurality of engineered cell lines comprising one or more of a reporter construct, heterologous polypeptide, or variant of a heterologous polypeptide. The engineered cell lines can be comprised within one or more partitions to facilitate screening of a plurality of compounds (e.g., one compound per partition). Such compounds canbe for example synthesized de novo as described herein or procured from chemical suppliers as known to those in the art. The plurality of engineered cell lines may comprise 2, 3, 4, 5, 6, 7, 8, 9, 10, 50, 75, 100, 1,000, 2,0003,0004,000 5,000 or 10,000 or more different engineered cell lines. Within the partition each engineered cell line is present in a particular amount (e.g., 1, 2, 3, 5, 10, 20, 30, 40 50, 75, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 cells or more). Within the partition each engineered cell line is present in a particular amount or less (e.g., 2, 3, 5, 10, 20, 30, 40 50, 75, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000, 1,100 cells or less). As described herein an “engineered cell line” refers to a cell that comprises one or more exogenous nucleic acids that encode or comprise, a heterologous polypeptide, a variant of a heterologous polypeptide, or a reporter construct. In certain embodiments, the exogenous nucleic acid is integrated into a genomic location of the cell (e.g., a chromosome). Such engineered cell lines can be constructed using various methods to introduce nucleic acids into cells (e g. CaC , cationic lipid transfection reagents, electroporation, viral transduction, etc.).
[0029] In the embodiments described herein, each well or partition may comprise a plurality of engineered cell lines. Engineered cell lines may comprise one or more nucleic acids. The one or more nucleic acids may encode one or more heterologous polypeptides described herein and / or comprise the reporter constructs described herein. The one or more nucleic acids may comprise a promoter operatively coupled to a reporter gene (e.g., a reporter construct). The reporter gene may further comprise a barcode (e.g., index sequence) which is unique and identifiable, and, as described below, may indicate one or more aspects regarding regulation of a reporter construct by one or more test agents. Such barcodes may also be included on reporter constructs with constitutively active promoters driving reporter genes to provide information of cell viability or toxicity. In certain embodiments, the barcode may not be under any type of transcription regulation to allow normalizing of overall cell numbers in an assay. In certain embodiments thebarcode is uniquely paired to indicate regulation of a specific promoter or a specific heterologous polypeptide or variant thereof. Such regulation includes activation and inhibition of the biological by a test ligand.
[0030] The reporter genes of the engineered cell lines are independently readable, meaning that the reporter genes are not dependent on each other for expression and are able to provide information on different biological activities in the same assay. Such activities include, but are not limited to, two or more different intracellular signaling pathways, one or more signaling pathways and cell viability, one or more signaling pathways and protein trafficking or abundance.
[0031] As described further below, when test agents are introduced to wells or partition, the test agents may interact with the plurality of cells within each well. The interactions between the test agent and the cells may include activation of one or more biological events, which may further lead to the expression or suppression of reporter genes containing barcodes, informing a user of the systems described herein on the biological effects of the one or more test agents.
[0032] In some embodiments, the cell described herein is obtained from a cell line described herein. In some embodiments, the cell is an eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is obtained from a subject. In some embodiments, the cell is obtained from a subject, who has a disease or condition. In some embodiments, the cell obtained from the subject can be screened for a test agentfortreatingthe disease or condition. In some embodiments, the cell is associated with a disease or condition. In some embodiments, the test agent can be used to diagnose, treat, or prevent the disease or condition. In some embodiments, the cell can be further engineered to express a heterologous polypeptide, a reporter construct, or a combination thereof.
[0033] In some embodiments, the methods and systems described herein comprise using n- different engineered cell lines. In some embodiments the ^-different engineered cell lines are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 or more different cell lines. In certain embodiments, these n-different engineered cell lines may be comprised with in a single partition.
[0034] The engineered cell line may be nay cell line useful for screening molecules or test agents. In some aspects the cell or cells are cancer cells, tumor cells, or otherwise immortalized cells. In further aspects, the cells represent a disease-model cell. In certain aspects the cells can be A549, B-cells, Bl 6, BHK-21, C2C12, C6, CaCo-2, CAP / , CAP-T, CHO, CHO2, CHO-DG44, CHO-K1, COS-1, Cos-7, CV-1, Dendritic cells, DLD-1, Embryonic Stem (ES) Cell orderivative, H1299, HEK293, 293T, 293FT, Hep G2, Hematopoietic Stem Cells, HOS, Huh-7, Induced Pluripotent Stem (iPS) Cell or derivative, Jurkat, K562, L5278Y, LNCaP, MCF7, MDA-MB-231, MDCK, Mesenchymal Cells, Min-6, Monocytic cell, Neuro2a, NIH 3T3, NIH3T3L1, K562, NK-cells, NSO, Panc-1, PC12, PC-3, Peripheral blood cells, Plasma cells, Primary Fibroblasts, RBL, Renca, RLE, SF21, SF9, SH-SY5Y, SK-MES-1, SK-N-SH, SL3, SW403, Stimulus-triggered Acquisition of Pluripotency (STAP) cell or derivate SW403, T- cells, THP-1, Tumor cells, U2O5, U937, peripheral blood lymphocytes, expanded T cells, hematopoietic stem cells, or Vero cells. In some embodiments, the cells are HEK293T cells. Heterologous polypeptides
[0035] Heterologous polypeptides are any polypeptide exogenous to the engineered cell line and may be encoded by an exogenous nucleic acid transfected into or added to the engineered cell line. The heterologous polypeptide may be a recombinant form of a polypeptide already present in the cell, or a recombinant form of a polypeptide not expressed in the engineered cell line. Heterologous polypeptides may comprise one or more amino acid variants (e.g., a substitution, deletion, addition, or truncation). Such exogenous nucleic acids maybe stably integrated into the genome of the engineered cell line, in some cases creating a clonal cell line expressing a heterologous polypeptide The nucleic acids may be single or double stranded. In certain embodiments, the nucleic acids are DNA. Such heterologous polypeptides may be included on a plasmid, a viral vector, or a linearized DNA to facilitate transfer of the nucleic acid to the engineered cell line.
[0036] In certain embodiments heterologous polypeptides may be expressed in the same cell as a reporter construct comprising a barcode, where barcodes are uniquely identifiable sequences on a reporter gene that can be used to identify the reporter gene and associated promoters and / or heterologous polypeptides (as described further below with respect to Figs. 1-3). In some embodiments, heterologous polypeptides may be activated or inhibited when a test agent binds to the heterologous polypeptide. Further, in some embodiments, the activation of a heterologous polypeptide may send a downstream signal that results in activation of a promoter or response element. One or more reporter constructs, as described below, may further be activated by the activation of the heterologous polypeptide or by the downstream signal initiated by the heterologous polypeptides. The promoter may further be operatively coupled to a reporter gene comprising a barcode. Variants of heterologous polypeptides may similarly be activated and / or expressed in a different plurality of engineered cell lines.
[0037] As described herein the heterologous polypeptides may comprise a polypeptide or protein that participates in signaling pathway in a cell. Heterologous polypeptides may belong to a class of proteins or protein family. Such heterologous polypeptides include G-coupled protein receptors (GPCRs), Receptor tyrosine kinases (RTK’s), intracellular signaling molecules, intracellular trafficking molecules, chaperones / heat shock proteins, transcriptional activators, enhancers or repressors, secreted proteins (e.g., growth factors, cytokines, chemokines, etc.) cell adhesion molecules, cytoskeletal components, DNA replication proteins, histones, or nuclear hormone receptors.
[0038] The systems, methods and assays described herein may use variants of heterologous polypeptides to determine the structural and functional basis for potential therapeutic intervention. As used herein, a “variant” of a polypeptide includes a polypeptide with an amino acid sequence that is different from a first heterologous polypeptide. Some variants of polypeptides useful for this disclosure are those that have a reduced ability to participate in the normal biological activity of the cell (e.g., cell signaling, trafficking, enzymatic, or metabolic activity). In general, amino acid sequences of variants of polypeptides useful for this disclosure may vary from amino acid sequences of the wild-type of the polypeptide by one or more amino acids. In some embodiments, an amino acid sequence of a variant of a polypeptide varies by one amino acid. In some embodiments, an amino acid sequence of a variant of a polypeptide varies by at least one amino acid, at least two amino acids, at least three amino acids, at least four amino acids, at least five amino acids, at least six amino acids, at least seven amino acids, at least eight amino acids, at least nine amino acids, or at least ten amino acids. In some embodiments, an amino acid sequence an amino acid sequence of a variant of a polypeptide varies by at most one amino acid, at most two amino acids, at most three amino acids, at most four amino acids, at most five amino acids, at most six amino acids, at most seven amino acids, at most eight amino acids, atmostnine amino acids, or atmostten amino acids. In some embodiments, the methods and assays use a plurality of heterologous polypeptides comprising single amino acid variants that cover at least about 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or 100% of amino acid residues of a heterologous polypeptide of interest. In some embodiments, plurality of heterologous polypeptides comprising single amino acid variants comprise 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000 or more heterologous polypeptide variants.
[0039] In some embodiments, the heterologous polypeptide is expressed by an engineered cell line described herein. For example, the heterologous polypeptide is expressed by a firstengineered cell line, a second engineered cell line, a third engineered cell line, or any additional cell line. In some embodiments, the heterologous polypeptide described herein can be a first heterologous polypeptide or a second heterologous polypeptide, where the first heterologous polypeptide and a second heterologous polypeptide can differ by at least one amino acid. In such case, the first heterologous polypeptide and the second heterologous polypeptide can be used to screen for interaction with a test agent, where the interaction is specific to the at least one amino acid alteration. In some embodiments, the second heterologous polypeptide comprises at least one amino acid alteration relative to the first heterologous polypeptide. In some embodiments, the second heterologous polypeptide is another polypeptide in another polypeptide in a common class of polypeptides, such as GPCRs, receptor tyrosine kinases, nuclear hormone receptors, etc. The systems described herein may comprise n different cell lines, each with a different heterologous polypeptide. In some embodiments, the / / -different cell lines are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 or more different cell lines. In this way many variants of a single heterologous polypeptide or of a biologically or structurally related class of heterologous polypeptides may be simultaneously analyzed.
[0040] In some embodiments, each heterologous polypeptide or each combination of the heterologous polypeptide can be expressed in an engineered cell line, or a combination of the engineered cell lines. In some embodiments, the expression of the heterologous polypeptide or a combination of the heterologous polypeptides in each engineered cell line or in a combination of the engineered cell lines can be portioned by a system described herein, where each partition of the system comprises an unique combination of cell lines expressing one or more of the heterologous polypeptides.
[0041] In some embodiments, the heterologous polypeptide is encoded by a gene that is operatively coupled to a promoter that controls its expression. In some embodiments, the promoter is constitutive. In some embodiments, the promoter is conditional or inducible (e.g., tetracycline inducible).
[0042] In some embodiments, the heterologous polypeptide may be fused to a transcription factor by a protease sensitive linker, that can be cleaved upon activation of or proximity to a protease with specificity for the linker
[0043] In some embodiments, the heterologous polypeptide is a cell-surface protein. For example, the heterologous polypeptide is a cell-surface protein of a G-protein coupled receptor, a receptor tyrosine kinase, an ion channel, a cytokine receptor, a chemokine receptor, a growthfactor receptor, a cellular adhesion molecule, a fragment thereof, or a combination thereof. In some embodiments, the heterologous polypeptide is an intracellular protein. For example, the heterologous polypeptide is an intracellular protein of an enzyme, ER transporter, nuclear transporter, intracellular signaling protein, a trafficking protein, a chaperone, a transcription factor, a fragment thereof, or a combination thereof.Reporter Constructs
[0044] In some embodiments, the engineered cell lines within a partition comprise a reporter construct. In certain embodiments, the reporter construct is encoded by an exogenous nucleic acid. In certain embodiments, the reporter construct comprises a promoter or response element operatively coupled to a reporter gene. A promoter may include any genomic element that can be bound by a transcription factor, enhancer, or the like, as well as be able to initiate transcription of a reporter gene downstream of that transcription factor, enhancer, or the like. Promoters described herein maybe inducible (e.g., promotersmay be activated by a particular transcription factor activated under specific signaling conditions). Promoters described herein may be constitutive active. Examples of constitutive active promoter can include from a SV40 promoter, a CMV promoter, an EflA promoter, aPGKl promoter, an Ubcpromoter, abeta actin promoter, a CAG promoter, an Ac5 promoter, a polyhedrin promoter, a TEF1 promoter, a GDS promoter, a CaMV355 promoter, an Ubi protomer, or any combination thereof. In some embodiments, the promoter can be an inducible promoter. Examples of inducible promoter can include a TRE promoter, a GALI promoter, a GAL10 promoter, or any combination thereof.
[0045] In some embodiments, the reporter gene comprises a barcode, wherein the barcode is a uniquely identifiable sequence. In some embodiments, the barcode may be used to identify the reporter gene or an associated promoter when expressed. In some embodiments, the barcode may be used to identify a heterologous receptor gene or a variant thereof. The reporter gene may comprise a gene encoding a fluorescent protein, a luciferase protein, a beta-galactosidase, a betaglucuronidase, a chloramphenicol acetyltransferase, or a secreted placental alkaline phosphatase, or a combination thereof. In some embodiments, the reporter gene comprises both the barcode and a gene encoding a reporter protein such as a fluorescent protein, a luciferase protein, a betagalactosidase, a beta-glucuronidase, a chloramphenicol acetyltransferase, or a secreted placental alkaline phosphatase, or a combination thereof.
[0046] Reporter constructs may be activated based on heterologous protein activation, a signal generated by heterologous polypeptide activation, or some other biological activity of the cell. Activation of a reporter construct may comprise a promoter of the reporter construct beingactivated by recruitment of the necessary transcription factors and transcriptional activators. Additionally, activation of the reporter construct may cause the reporter gene operatively coupled to the promoter to be expressed, wherein the expression of the reporter gene allows for the barcode on the reporter gene to be identified in one or more subsequent assays. The barcode may indicate one or more aspects associated with the activation of the reporter construct, such as what promoter was activated, by what pathway the promoter was activated, if a certain heterologous polypeptide or a variant of the heterologous polypeptide was activated, if there is toxicity due to one or more test agents in the well, or if a test agent within the well did not regulate any biological activity within the well. In some embodiments, the reporter gene comprising the barcode may not be operatively coupled to a promoter so that the barcodes may be used to assess proliferation and viability of a cell.
[0047] The barcodes that are expressed may indicate what reporter constructs were activated, and therefore, may further indicate what heterologous polypeptides or variants of heterologous polypeptides were activated, what pathway s certain reporter constructs were activated through, if there is toxicity in a partition, or if test agents did not bind to anything within the partition or produce a biological effect. Thus, when a test agent is introduced to a partition, the resultant expressed barcodes indicate what reporter constructs were activated as well as how the reporter constructs were activated, which may allow for determinations of what test agents are useful for modulating a particular biological activity, including diseases or conditions of interest.
[0048] In some embodiments, the reporter construct is integrated into the genome of the cell. In some embodiments, the reporter construct is not integrated into the genome of the cell.
[0049] In some embodiments, the system and the method described herein utilize at least one reporter construct. In some embodiments, each of the reporter construct can be regulated by a different promoter. For example, a first reporter construct can be regulated by a first promoter, and a second reporter construct can be regulated by a second promoter, and so on. In some embodiments, each partition of the system can include an unique or a combination of promoters regulating the expression of an unique or a combination of reporter constructs. In some embodiments, the promoters can be constitutively active promoters, inducible promoters, or a combination thereof. The systems described herein may comprise n different cell lines, each with a different reporter construct. In some embodiments, the n -different cell lines are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 or more different cell lines. In this way many different promoters may be simultaneously analyzed in response to a test agent.
[0050] As described above, the plurality of reporter constructs may include one or more promoters. In some embodiments, the plurality of reporter constructs may include a plurality of promoters. In some embodiments, the plurality of promoters may include a plurality of types of promoters. In some embodiments, the plurality of promoters may include a promoter and an alternative promoter. In some embodiments, the promoter may be a first type of promoter, and the alternative promoter may be a second type of promoter. In some embodiments, the alternative promoter may be a constitutive promoter. In some embodiments, the plurality of promoters may include the promoter, the alternative promoter, and a constitutive promoter. While plurality a promoters with a promoter, an alternative promoter, a constitutive promoter, or combinations thereof are described, this plurality of promoters is exemplary and other plurality of promoters may be used. For example, while pluralities of promoters with two or three promoters are described, a largernumber ofpromoters maybe used, such as four, five, six, seven, eight, nine, or ten or more promoters.
[0051] In some embodiments, one or both of a promoter and a second, third, fourth, etc., promoter are activated. In some embodiments, the promoter and the second, third, fourth, etc., promoter are activated via different signaling pathways. For example, the promoter may be activated by a first type of test agent, while the second, third, fourth, etc., promoter may not be activated by the first type of test agent. Additionally, promoter and the second, third, fourth, etc., promoter may both be activated to different extents. For example, the promoter may be activated greater than 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold or more compared to the second, third, fourth, etc , promoter; orthe second, third, fourth, etc., promoter may be activated greater than 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold or more compared to the promoter. Thus, with one or more reporter constructs comprised within an engineered cell line of a partition, a promoter and / or second, third, fourth, etc., promoter may be operatively coupled to one or more reporter genes which may be expressed upon activation of the promoter and / or alternative promoter. Each reporter gene may further include a barcode uniquely paired with the promoter. In some embodiments, the barcode may be used to uniquely identify the reporter gene, and by extension, the promoter operatively coupled to the reporter gene when the reporter gene is expressed. Thus, in a cell including a first reporter construct including, for example, a CREB promoter operatively coupled to a first reporter gene comprising a first barcode, where the cell also includes a second reporter construct including, for example, an NFAT promoter (e.g., an alternative promoter to the first promoter) operatively coupled to a second reporter gene comprising a second barcode, identification of the first barcode indicates that the first promoterwas activated, while identification of the second barcode indicates that the second promoter was activated. Accordingly, when a test agent
[0052] In some embodiments, one or both a promoter and a constitutive promoter present in an different engineered cell lines in a partition are activated. In some embodiments, the constitutive promoter detects toxicity (by a reduction in activation) or an off-target reduction in signaling. Constitutive promoters may include CMV, RSV, SV40, and the like.Biological assays
[0053] The engineered cell lines described herein may be used in one or more biological assays to acquire data as to the regulation of biological activities when contacted by a test agent. As described herein, biological assays comprise contacting or culturing the cells in the wells with one or more test agents.
[0054] As used herein “Biological activity” refers to any activity of the cell necessary for normal cell growth and development. Such activities include DNA synthesis, protein synthesis, folding and trafficking, intracellular singling activity, gene transcription, enzymatic activities for the cell, synthesis and secretion of growth factorsand other cell-to-cell signaling molecules, cell division, apoptosis, etc. Certain examples of assays to detect different biological activities are depicted in FIG. 8
[0055] When cells are contacted or cultured with the one or more test agents, one or more biological activities of the cells may be regulated by the test agent. The activation of a target may further lead to the expression of one or more reporter genes and the ability to identify one or more barcodes of those reporter genes through sequencing, for example. The identification of barcodes allows for determinations of what test agents regulate specific reporter genes, and provide insight into which biological activities are regulated by which test agents.
[0056] In some embodiments, one or more test agents may be assayed. In some embodiments, the amount of test agents that may be used in an assay of the systems described herein is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 100, 1,000, 10,000 or more.
[0057] In some embodiments, one or more test agents may be assayed. In some embodiments, the amount of test agents that may be used in assays is about 1 test agent to about 1,536 test agents. In some embodiments, the amount of test agents that may be used in assays is about 1 test agentto about48 test agents, about 1 test agent to about 96 test agents, about 1 test agentto about 192 test agents, about 1 test agent to about 288 test agents, about 1 test agent to about 384 test agents, about 1 test agent to about 768 test agents, about 1 test agent to about 1,152 test agents, about 1 testagentto about 1,536 test agents, about48 test agents to about 96 test agents, about 48test agents to about 1 2 test agents, about 48 test agents to about 288 test agents, about 48 test agents to about 384 test agents, about 48 test agents to about 768 test agents, about 48 test agents to about 1,152 test agents, about 48 test agents to about 1,536 test agents, about 96 test agents to about 192 test agents, about 96 test agents to about 288 test agents, about 96 test agents to about 384 test agents, about 96 test agents to about 768 test agents, about 96 test agents to about 1,152 test agents, about 96 test agents to about 1 ,536 test agents, about 192 test agents to about 288 test agents, about 192 test agents to about 384 test agents, about 192 test agents to about 768 test agents, about 192 test agents to about 1,152 test agents, about 192 test agents to about 1,536 test agents, about 288 test agents to about 384 test agents, about 288 test agents to about 768 test agents, about288 test agents to about 1,152 test agents, about 288 test agents to about 1,536 test agents, about 384 test agents to about 768 test agents, about 384 test agents to about 1,152 test agents, about 384 test agents to about 1,536 test agents, about 768 test agents to about 1,152 test agents, about768 test agents to about 1,536 test agents, or about 1,152 test agents to about 1,536 test agents. In some embodiments, the amount of test agents that may be used in assays is about 1 test agent, about 48 test agents, about 96 test agents, about 192 test agents, about 288 test agents, about 384 test agents, about 768 test agents, about 1,152 test agents, or about 1,536 test agents. In some embodiments, the amount of test agents that may be used in assays is at least about 1 test agent, about 48 test agents, about 96 test agents, about 192 test agents, about 288 test agents, about 384 test agents, about 768 test agents, or about 1, 152 test agents. In some embodiments, the amount of test agents that may be used in assays is at most about 48 test agents, about 96 test agents, about 192 test agents, about 288 test agents, about 384 test agents, about 768 test agents, about 1,152 test agents, or about 1,536 test agents.
[0058] The one or more test agents may come into contact with the plurality of cells comprising one or more heterologous polypeptides or variants of heterologous polypeptides. In some embodiments, the amount of heterologous polypeptides that may be assayed is about 1 heterologous polypeptide to about 10,000 heterologous polypeptides. In some embodiments, the amount of heterologous polypeptides that may be used in assays is about 1 heterologous polypeptide to about 100 heterologous polypeptides, about 1 heterologous polypeptide to about 500 heterologous polypeptides, about 1 heterologous polypeptide to about 1,000 heterologous polypeptides, about 1 heterologous polypeptide to about 1,500 heterologous polypeptides, about 1 heterologous polypeptide to about 2,000 heterologous polypeptides, about 1 heterologous polypeptide to about 4,000 heterologous polypeptides, about 1 heterologous polypeptide to about 6,000 heterologous polypeptides, about 1 heterologous polypeptide to about 8,000 heterologouspolypeptides, about 1 heterologous polypeptide to about 10,000 heterologous polypeptides, about 100 heterologous polypeptides to about 500 heterologous polypeptides, about 100 heterologous polypeptides to about 1,000 heterologous polypeptides, about 100 heterologous polypeptides to about 1,500 heterologous polypeptides, about 100 heterologous polypeptides to about 2,000 heterologous polypeptides, about 100 heterologous polypeptides to about 4,000 heterologous polypeptides, about 100 heterologous polypeptides to about 6,000 heterologous polypeptides, about 100 heterologous polypeptides to about 8,000 heterologous polypeptides, about 100 heterologous polypeptides to about 10,000 heterologous polypeptides, about 500 heterologous polypeptides to about 1,000 heterologous polypeptides, about 500 heterologous polypeptides to about 1,500 heterologous polypeptides, about 500 heterologous polypeptides to about 2,000 heterologous polypeptides, about 500 heterologous polypeptides to about 4,000 heterologous polypeptides, about 500 heterologous polypeptides to about 6,000 heterologous polypeptides, about 500 heterologous polypeptides to about 8,000 heterologous polypeptides, about 500 heterologous polypeptides to about 10,000 heterologous polypeptides, about 1,000 heterologous polypeptides to about 1,500 heterologous polypeptides, about 1,000 heterologous polypeptides to about 2,000 heterologous polypeptides, about 1,000 heterologous polypeptides to about 4,000 heterologous polypeptides, about 1,000 heterologous polypeptides to about 6,000 heterologous polypeptides, about 1,000 heterologous polypeptides to about 8,000 heterologous polypeptides, about 1,000 heterologous polypeptides to about 10,000 heterologous polypeptides, about 1,500 heterologous polypeptides to about2,000 heterologous polypeptides, about 1,500 heterologous polypeptides to about 4,000 heterologous polypeptides, about 1,500 heterologous polypeptides to about 6,000 heterologous polypeptides, about 1,500 heterologous polypeptides to about 8,000 heterologous polypeptides, about 1,500 heterologous polypeptides to about 10,000 heterologous polypeptides, about 2,000 heterologous polypeptides to about 4,000 heterologous polypeptides, about 2,000 heterologous polypeptides to about 6,000 heterologous polypeptides, about 2,000 heterologous polypeptides to about 8,000 heterologous polypeptides, about 2,000 heterologous polypeptides to about 10,000 heterologous polypeptides, about 4,000 heterologous polypeptides to about 6,000 heterologous polypeptides, about 4,000 heterologous polypeptides to about 8,000 heterologous polypeptides, about 4,000 heterologous polypeptides to about 10,000 heterologous polypeptides, about 6,000 heterologous polypeptides to about 8,000 heterologous polypeptides, about 6,000 heterologous polypeptides to about 10,000 heterologous polypeptides, or about 8,000 heterologous polypeptides to about 10,000 heterologous polypeptides. In some embodiments, the amount of heterologous polypeptides that may be used in assays is about 1 heterologouspolypeptide, about 100 heterologous polypeptides, about 500 heterologous polypeptides, about 1,000 heterologous polypeptides, about 1,500 heterologous polypeptides, about 2,000 heterologous polypeptides, about 4,000 heterologous polypeptides, about 6,000 heterologous polypeptides, about 8,000 heterologous polypeptides, or about 10,000 heterologous polypeptides. In some embodiments, the amount of heterologous polypeptides that may be used in assays is at least about 1 heterologous polypeptide, about 100 heterologous polypeptides, about 500 heterologous polypeptides, about 1,000 heterologous polypeptides, about 1,500 heterologous polypeptides, about 2,000 heterologous polypeptides, about 4,000 heterologous polypeptides, about 6,000 heterologous polypeptides, or about 8,000 heterologous polypeptides. In some embodiments, the amount of heterologous polypeptides that may be used in assays is at most about 100 heterologous polypeptides, about 500 heterologous polypeptides, about 1,000 heterologous polypeptides, about 1,500 heterologous polypeptides, about 2,000 heterologous polypeptides, about 4,000 heterologous polypeptides, about 6,000 heterologous polypeptides, about 8,000 heterologous polypeptides, or about 10,000 heterologous polypeptides.
[0059] In some embodiments, assessing the regulation of a specific biological activity may include using a biological assay to test which test agents effected a given biological activity. In some embodiments, the test agents may bind to the target at one or more binding sites based on the ligands binding. In some embodiments, the biological assay may include information from two or more reporter constructs of one or more pluralities of engineered cell lines. Reporter constructs may include one or more promoters. In some embodiments, the one or more promoters may include promoters activated by the target or a variant of the target (directly or indirectly through signaling intermediates, or lack thereof). In
[0060] Biological assays described herein comprise contacting a test agent to one or more wells comprising engineered cell lines. The biological assay may further comprise incubating the test agentwith the cells for a time sufficient to allow for expression of a barcode operatively coupled to a promoter or alternative promoter activated by the target. A test agent may be added to each well. In some embodiments, a different test agent is added to each respective well. The biological assay may be carried out in any particular multi-vessel format such as 96-, 384-, or 1536 well plates, microfluidic chips comprising microwells, or in emulsions.
[0061] Fig. 1 depicts a cell 100 interacting with a test agent 102. Cell 100 comprises heterologous polypeptide 104, which is encoded by nucleic acid 106, and a reporter construct 108 also encoded by a nucleic acid. The target may be expressed on the cell surface (e.g., in the case of a G-protein coupled receptor or receptor tyrosine kinase) or intracellularly (e.g., in the case ofa nuclear hormone receptor). Reporter construct 108 includes promoter 110 operatively coupled to reporter gene 112. Reporter gene 112 may further include a uniquely identifiable barcode 114. In some embodiments, the target 104 may be a heterologous polypeptide.
[0062] In this depicted embodiment, test agent 102 is added to a well containing cell 100, where test agent 102 then interacts with the cell (e.g., via heterologous polypeptide 104). In some embodiments, test agent 102 interacts with the cell by bindin to heterologous polypeptide 104 or affecting a signaling pathway downstream from the heterologous polypeptide. The interaction of test agent 102 affects signaling through heterologous polypeptide 104. In response to the test agent 102 the promoter 110 of reporter construct 108 is activated. In some embodiments, the activation of the promoter 110 may cause reporter gene 112, which is operatively coupled to the promoter 110, to be expressed (e.g., as expression 118). The reporter gene 112, and by extension, barcode 114, may be expressed as expression 118, which may be analyzed or quantified by sequencing (e.g., using a next-generation sequencing assay) in order to identify barcode 114. Barcode 114 may be associated with one or more promoters (e.g., promoter 110) and may indicate certain aspects about the assay (e.g., that test agent 102 activated reporter construct 108 through a first pathway).
[0063] Thus, a well or a partition within a well containing multiple cells may be assayed in order to determine expression of barcodes that uniquely identify heterologous polypeptides and / or response elements, modulated by test agents. In some cases, there may be more than one test agent assayed, and there may be more than one target (e.g., a target and its variant, or a plurality of completely different targets) that may be assayed. Further, in some embodiments there may be more reporter construct that may be activated within a cell, as described below. In some embodiments, there may only be one reporter construct within a single cell, but different reporter construct may be present in other cells.
[0064] In embodiments where more than one promoter is present, whether within the same cell or different cells within a well or a partition, a biological assay may indicate which test agents have modulated biological activity based on the one or more reporter constructs. For example, when a ligand binds or activates the target, one or more promoters may be activated. When the one or more promoters are activated, the one or more promoters may express one or more reporter genes (e g., a barcode mRNA) operatively coupled to the promoter. Each respective reporter gene of the one or more reporter constructs may indicate that a test agent has bound to heterologous polypeptide or a variant of the heterologous polypeptide, that the test agent promotes toxicity, or that a test agent has not produced any significant biological effect.
[0065] Fig. 2 depicts an exemplary system described herein, in this example cell 200 interacting with test agent 202. In this depicted example, test agent 202 interacts with cell 200 by binding to heterologous polypeptide 204 that is encoded by a nucleic acid 206, which causes activation of signaling through the heterologous polypeptide 204. In this depicted example, the activation of heterologous polypeptide 204 causes promoter 212 of reporter construct 210 to activate via pathway 208a. Further, in this depicted example, the activation of the target could have caused promoter 222 of reporter construct 220 (comprising barcode 226) to activate via pathway 208b. The activation of promoter 212 causes reporter gene 214, which is operatively coupled to promoter 212, to be expressed 218. By using sequencing techniques, barcode 216 can be identified as being expressed 218 while barcode 226 is not activated, which indicates one or more aspects about the interaction of test agent 202 with the cell and / or heterologous polypeptide 204 or signaling pathway thereof, for example in this figure, that test agent may act in a pathway specific manner with respect to pathway 208a. Other outcomes are potentially available indicating that only pathway 208b or that both pathways 208a and 208b are activated. This analysis can be scaled to interrogate the biological activities of multiple different test agents on multiple different biological activities or signaling pathways. The systems can be deployed in simultaneous or near simultaneous assays allowing for the high-throughput detection of biological differences between test agents
[0066] For example, barcode 216 may indicate that promoter 212 was activated via pathway 208a. Barcode 226 could indicate that promoter 222 was activated via pathway 208b. In some embodiments, a plurality of barcodes identified through a plurality of expressions, including expression 218, may indicate that promoter 212 was activated in multiple cells, and that promoter 222 was also activated in multiple cells. In some embodiments, promoter 222 may be an alternative promoter of promoter 212. In some embodiments, promoter 212 or promoter 222 may be activated based on a test agent 202 affecting a variant of target 204. In some embodiments, the activation of promoter 212 or promoter 222 may indicate toxicity within the well in which cell 200 interacts with test agent 202.
[0067] Fig. 3 depicts example cell 300 interacting with test agent 302. In this depicted example, test agent 302 interacts with a signaling pathway of cell 300. In this depicted example, test agent 302 leads to the activation of promoter 322 though an interaction with the signaling machinery of the cell. Depending upon the experimental design this activation could represent a promoter intended to interrogate an off-target effect or an on-target effect. Further, in this depicted example, promoter 312 of reporter construct 310 was not activated by the test agent. Theactivation of promoter 322 causes reporter gene 324, which is operatively coupled to promoter 322, to be expressed 328. By using sequencing techniques, barcode 326 of reporter gene 324 can be identified as being expressed by expression 328, which indicates one or more aspects about the interaction of test agent 302 with cell 300. For example, barcode 326 may indicate that promoter 322 was activated as an off-target effect of test agent 302.
[0068] Thus, by using a plurality of cell lines engineered to include one or more reporter constructs, the activities of test agents biological activities of a cell can be identified based on how they interact with certain promoters through identification of resultant barcodes.
[0069] These examples are not limited to assays which use heterologous polypeptides or variants thereof. Cell lines lacking heterologous polypeptides and comprising reporter constructs may also be used to understand the effect a given test agent may have on a biological activity of a cell. For example, a plurality of wells may include a plurality of cell lines, where the plurality of cell lines includes at least a first cell line and second cell line. In some embodiments, the first cell line may include a first reporter construct including a first promoter operatively coupled to a first reporter gene including a first barcode and the second cell line may include a second reporter construct including a second promoter operatively coupled to a second reporter gene including a second barcode. In some embodiments, the first promoter may be inducible or constitutive. In some embodiments, the second promoter may be inducible or constitutive. In some embodiments, the first promoter and second promoter may be inducible. In some embodiments, the first promoter may be a first inducible promoter and the second promoter may be a second inducible promoter different from the first promoter In some embodiments, the plurality of cell lines may include a third engineered cell line with a third reporter construct including a third promoter operatively coupled to a third reporter gene including a third barcode. In some embodiments, the third promoter may be inducible. In some embodiments, the third promoter may be different from this first and second promoters. In some embodiments, the third promoter is constitutive if the first promoter and the second promoter are inducible.
[0070] The plurality of cell lines may include one or more heterologous polypeptides. In some embodiments, each of the first, second, and third cell lines includes a same heterologous polypeptide. In some embodiments, each of first, second, andthird cell lines includes a different heterologous polypeptide. In some embodiments, any combination of the first, second, and third cell lines have a same heterologous polypeptide, while the remainder have a different polypeptide. In some embodiments, a heterologous polypeptide with a cell line of the plurality of cells lines may be coupled to a transcription factor. In some embodiments, the transcription factorcomprises one or more of aGal4, PPR1, Lac9, zinc finger, or LexA DNA binding domain. In some embodiments, the transcription factor comprises oneor more of aVP64, p65, RoTev, orRta DNA activating domain.
[0071] When a test agent is added to a partition including one or more of the plurality of cell lines, the first promoter, second promoter, or third promoter may be activated, leading to the expression of the first reporter gene, second reporter gene, or third promoter gene, respectively, and which allows for the identification of the first barcode, the second barcode, or the third barcode, respectively. Upon identification of the first barcode, the second barcode, or the third barcode, the activation of the respective promoter can thus be identified, and therefore, the ability of a test agent to regulate a heterologous polypeptide that activates a respective promoter. In doing so, the ability of various test agents to regulate particular heterologous polypeptides can be studied.
[0072] Fig. 4 depicts an example of a partition (e.g., a well of a multi-well plate) containing one or more engineered cell lines (e g., cells 100, 200, and / or 300 of Figs. 1, 2, and / or 3, respectively). Test agents may be introduced to the well during a biological assay, leading to the activation of one or more reporter constructs. Sequencing may then be performed to identify barcodes related to those reporter constructs. Thus, multiple reporters of the current disclosure may yield information simultaneously in response to one or more test agents.
[0073] Figs. 5A-5D depict an example nucleic acids encoding heterologous polypeptides and / or comprising reporter constructs. While depicted as a single nucleic acid for simplicity, the skilled artisan will appreciate that target encoding nucleic acids and reporter construct comprising nucleic acids may be supplied on different nucleic acids, within the well. Fig. 5A depicts two reporter constructs for the analysis of the expression of reporter genes with different response elements, allowing the determination of different signaling pathways activated by a single molecule. Additionally, a reporter is included without a target allowing for the assessment of off target activation. Fig. 5B depicts two different heterologous polypeptides, signaling though the same response element, allowing for the analysis of whether a test agent activates one target or a variant of that target. Additionally, a reporter is included without a heterologous polypeptide allowing for the assessment of off target activation. Figs. 5C-5D depict systems to assess the effect of a test agent on interaction with multiple targets, through multiple response elements. While depicted as a single nucleic acid for simplicity, a heterologous peptide and response element may suitably be supplied on different nucleic acids.Test Agents
[0074] In some embodiments, described herein are methods utilizing the biological assays described herein orthe systems described herein. In some embodiments, the method comprises contacting the plurality of engineered cell lines with a test agent such as a compound and measuring an activity of the reporter construct of the first engineered cell line, the second engineered cell line, third engineered cell line, or any combination thereof. In some embodiments, the activity of the reporter construct comprises a readout of the reporter construct. In some embodiments, the activity of the reporter construct comprises detecting the barcode encoded by the reporter construct. In some embodiments the activity of the reporter construct comprises detecting a signal of a protein encoded by the reporter construct. For example, the activity of the reporter construct can include signal generated by a fluorescent protein or a luciferase protein encoded by the reporter construct. In some embodiments, the method comprises screening the test agent in one or more partitions, where the one or more partitions comprise one or more of the cell lines described herein harboring one or more of the heterologous polypeptides described herein and one or more of the reporter constructs described herein. In some embodiments, the method comprises utilizing a plurality of / / -test agents, where the test agents are contacted to the plurality of engineered cell lines that has been divided into at least / / -partitions. In some embodiments, the plurality of engineered cell lines comprises at least 100, at 1,000, or at least 10,000 different heterologous polypeptides. In some embodiments, the method is performed in a 96-well, 384-well ora 1536-well plate. In some embodiments, the first reporter construct and the second reporter construct are present in different engineered cells in the same well of a 96-well, 384-well, or a 1536-well plate. In some embodiments, the plurality of n-test agents is prepared as a plurality of individual reaction mixtures by reacting a core fragment A comprising a reactive functionality x, wherein x is an amine, aldehyde, boronate, imide, isothiocyanate, carboxylic acid, halide, hydroxy amidine, orthiourea; with a plurality of naive test fragments (y-Ti, y-T2, . . . y-Tn), each test fragment comprising a reactive functionality y, wherein y is a amine, boronate, imide, isothiocyanate, halide, hydroxyamide, thiourea, carboxylic acid or aldehyde and one of a plurality of naive test moieties (T T2, . . . Tn) under reaction conditions sufficient to form a plurality of n test agents (A-T A-T2, . . . A-Tn), wherein each of A-Ti, A-T2, . . A-Tnis prepared in a well of a test plate.
[0075] In some embodiments, the method comprises contacting the test agent and the heterologous polypeptide under a reaction condition. In some embodiments, the reaction condition comprises a basic environment. In some embodiments, the reaction conditioncomprises a Lewis acid or Bronstead acid condition. In some embodiments, the reaction condition comprises contacting the cell line, the test agent, the heterologous polypeptide, or a combination thereof with an amide coupling reagent. In some embodiments, the reaction condition comprises contacting the cell line, the test agent, the heterologous polypeptide, or a combination thereof with a palladium reagent. In some embodiments, the reaction condition comprises contacting the cell line, the test agent, the heterologous polypeptide, or a combination thereof at room temperature. In some embodiments, the reaction condition comprises contacting the cell line, the test agent, the heterologous polypeptide, or a combination thereof at a temperature between about room temperature to about 80 °C. In some embodiments, the reaction condition comprises contacting the cell line, the test agent, the heterologous polypeptide, or a combination thereof for a duration from about 1 to about 24 hours. In some embodiments, the reaction condition comprises contacting the cell line, the test agent, the heterologous polypeptide, or a combination thereof for a duration from about 6 to about 18 hours. In some embodiments, the reaction of the test agent contacting the heterologous polypeptide is reversible or irreversible. In some embodiments, the reaction condition comprises contacting the cell line, the test agent, the heterologous polypeptide, or a combination thereof on a nanoscale. In some embodiments, the nanoscale is performed in a volume of 50 nL to 500 nL. In some embodiments, the reaction condition is partition in a 96-well, 384-well, or 1536-well test plate.
[0076] Described herein is a method for screening a test agent that regulates a biological activity of a cell, comprising performing a first screening procedure comprising:
[0077] (a) preparing a plurality of individual reaction mixtures by reacting a core fragment A comprising a reactive functionality x, wherein x is an amine, aldehyde, boronate, imide, isothiocyanate, carboxylic acid, halide, hydroxy amidine, or thiourea; with a plurality of naive test fragments (y-Tby-T2, . . . y-Tn), each test fragment comprising a reactive functionality y, wherein y is a amine, boronate, imide, isothiocyanate, halide, hydroxyamide, thiourea, carboxylic acid or aldehyde and one of a plurality of naive test moieties (Ti, T2, . . . Tn) under reaction conditions sufficient to form a plurality of n test agents (A-Ti, A-T2, . . . A-Tn), wherein each of A-Ti, A-T2, . . . A-Tnis prepared in a well of a test plate; (b) contacting a plurality of engineered cell lines with one or more of the reaction mixtures of (a) under conditions that permit regulation of the biological activity of the cell by any one or more of the plurality of n test agents.
[0078] As used herein “core fragment” means a chemical compound selected for use in the method disclosed herein, and is also denoted as “A”. “Selected” means that a skilled person willhave identified the core as having potential utility as a reference or anchor for the development of potential ligands for a target.
[0079] In some embodiments, the core fragment A comprises a reactive functionality “x”. The reactive functionality x may be a chemical functional group in the structure of a core fragment. In some embodiments, the core fragment comprising a reactive function x is denoted by “(A-x)”.
[0080] In some embodiments, x is an amine, aldehyde, boronate, imide, isothiocyanate, carboxylic acid, halide, hydroxyamidine, or thiourea. In some embodiments, x is a boronate. In some embodiments, x is an imide or isothiocyanate. In some embodiments, x is a halide. In some embodiments, x is a hydroxyamidine. In some embodiments, x is a thiourea.
[0081] In some embodiments, x is an amine. In some embodiments, the amine is a primary or secondary amine. In some embodiments, x is a primary amine. In some embodiments, x is a secondary amine.
[0082] In some embodiments, x is an aldehyde or carboxylic acid. In some embodiments, x is an aldehyde. In some embodiments, x is a carboxylic acid.
[0083] As used herein “naive test moiety” means a compound that may have an intrinsic binding affinity for the target and is a component of a test agent, and may be denoted as“T”; e.g. TbT2, .• • T ^n-
[0084] As used herein “naive test fragment “(y-T) refers to a test moiety T bound to a reactive functionality “y”.
[0085] In some embodiments, each test fragment comprising a reactive functionality y. In some embodiments, y is an amine, boronate, imide, isothiocyanate, halide, hydroxy amide, thiourea, carboxylic acid or aldehyde. In some embodiments, y is a boronate. In some embodiments, y is an imide or isothiocyanate. In some embodiments, y is a halide. In some embodiments, y is a hydroxyamide or thiourea.
[0086] In some embodiments, y is an amine. In some embodiments, the amide is a primary or secondary amine. In some embodiments, y is a primary amine. In some embodiments, y is a secondary amine
[0087] In some embodiments, y is a carboxylic acid or aldehyde. In some embodiments, y is an aldehyde. In some embodiments, y is a carboxylic acid.
[0088] A test fragment “(y-T)” is selected to permit reaction with the core fragment A to yield a test agent, denoted “(A-T)”, in which the reaction is between the two reactive functionalities x and y.
[0089] In some embodiments, the reaction conditions of step (a) comprise a base. Without being bound by theory, bases include for example a Hunig’s base.
[0090] In some embodiments, the reaction conditions of step (a) comprise a Lewis acid or Bronstead acid. In some embodiments, step (a) comprises a Lewis acid. In some embodiments, step (a) comprises a Bronstead acid.
[0091] In some embodiments, the reaction conditions of step (a) comprise an amide coupling reagent. Without being bound by theory, in some embodiments, exemplary coupling agents include but are not limited to propanephosphonic acid anhydride (T3P), HATU, HBTU, COMU, and DMTMM.
[0092] In some embodiments, the reaction conditions of step (a) comprises a palladium reagent. In some embodiments, the reaction conditions of step (a) compares a stoichiometric palladium oxidating addition complex. In some embodiments, the reaction conditions of step (a) comprises a stoichiometric G3 Pd catalyst.
[0093] In some embodiments, the reaction conditions comprise a reductive amination.
[0094] In some embodiments, the reactive conditions comprise a heterocycle formation, for example an oxadiazole formation from closure of hydroxyamidines and carboxylic acids, thiourea formation by in-situ conversion of primary amines to isothiocyanates, followed by coupling with other amines.
[0095] In some embodiments, the reaction conditions comprise Buchwald coupling-type chemistries (e.g. with stoichiometric palladium oxidative addition complexes or with stoichiometric G3 Pd catalysts) to form secondary or tertiary aryl amines from aryl halides and primary or secondary amines.
[0096] In some embodiments, the reactive conditions comprise Suzuki-like coupling where a Pd reagent is combined with boronic acids or boronates to give rise to sp2-sp3C-C coupling product.
[0097] In some embodiments, the reaction conditions comprise room temperature.
[0098] In some embodiments, the reaction conditions comprise a reaction temperature between about room temperature and about 80 °C, or any temperature therein.
[0099] In some embodiments, the reaction conditions comprise a reaction temperature of about room temperature. In some embodiments, the reaction conditions comprise a reaction temperature of about 25 °C, 30 °C, 35 °C, 40 °C, 45 °C, 50 °C, 55 °C, 60 °C, 65 °C, 70 °C, 75 °C, or 80 °C.
[0100] In some embodiments, the reacting is performed from about 1 to about 24 hours. In some embodiments, the reacting is performed from about 6 to about 18 hours.
[0101] In some embodiments, the reacting is performed for about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 hours. In some embodiments, the reacting is performed for about 24 hours. In some embodiments, the reacting is performed for about 22 hours. In some embodiments, the reacting is performed for about 20 hours. In some embodiments, the reacting is performed for about 18 hours. In some embodiments, the reacting is performed for about 16 hours. In some embodiments, the reacting is performed for about 14 hours. In some embodiments, the reacting is performed for about 12 hours. In some embodiments, the reacting is performed for about 10 hours. In some embodiments, the reacting is performed for about 8 hours. In some embodiments, the reacting is performed for about 6 hours. In some embodiments, the reacting is performed for about 4 hours.
[0102] In some embodiments, the reacting is reversible or irreversible. In some embodiments, the reacting is reversible. In some embodiments, the reacting is irreversible.
[0103] In some embodiments, step (a) comprises an amide coupling, reductive amination, oxidative addition, or heterocycle formation. In some embodiments, step (a) comprises an amide coupling. In some embodiments, step (a) comprises a reductive amination. In some embodiments, step (a) comprises an oxidative addition. In some embodiments, step (a) comprises a heterocycle formation.
[0104] In some embodiments, step (a) comprises a Buchwald or Suzuki coupling. In some embodiments, step (a) comprises a Buchwald coupling. In some embodiments, step (a) comprises a Suzuki coupling.
[0105] In some embodiments, the core moiety A has a mass of from about 150 Da to about 800 Da. In some embodiments, the core moiety A has a mass of from about 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, or 800 Da.
[0106] In some embodiments, the core moiety A has a mass of from about 150 Da. In some embodiments, the core moiety A has a mass of from about 200 Da. In some embodiments, the core moiety A has a mass of from about 300 Da. In some embodiments, the core moiety A has a mass of from about 400 Da. In some embodiments, the core moiety A has a mass of from about 500 Da. In some embodiments, the core moiety A has a mass of from about 600 Da. In some embodiments, the core moiety A has a mass of from about 700 Da. In some embodiments, the core moiety A has a mass of from about 800 Da.
[0107] In some embodiments, each test moiety (T T2, . . . Tn) has a mass of from about 80 Da to about 500 Da.
[0108] In some embodiments, each test moiety (TbT2, . . . Tn) has a mass of about 80, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, 320, 340, 360, 380, 400, 420, 440, 460, 480, or 500 Da. In some embodiments, each test moiety (TbT2, . . . Tn) has a mass of about 80 Da. In some embodiments, each test moiety (TbT2, . . . Tn) has a mass of about 100 Da. In some embodiments, each test moiety (TbT2, . . . Tn) has a mass of about 200 Da. In some embodiments, each test moiety (TbT2, . . . Tn) has a mass of about 300 Da. In some embodiments, each test moiety (TbT2, . . . Tn) has a mass of about 400 Da. In some embodiments, each test moiety (TbT2, . . . Tn) has a mass of about 500 Da.
[0109] In some embodiments, each test agent (A-TbA-T2, . . . A-Tn) has a mass of less than about 1500 Da. In some embodiments, each test agent (A-TbA-T2, . . . A-Tn) has a mass of less than about 1400 Da. In some embodiments, each test agent (A-TbA-T2, . . . A-Tn) has a mass of less than about 1300 Da. In some embodiments, each test agent (A-TbA-T2, . . . A-Tn) has a mass of less than about 1200 Da. In some embodiments, each test agent (A-TbA-T2, . . . A-Tn) has a mass of less than about 1100 Da. In some embodiments, each test agent (A-TbA-T2, . . . A- Tn) has a mass of less than about 1000 Da. In some embodiments, each test agent (A-TbA-T2, . . . A-Tn) has a mass of less than about 900 Da. In some embodiments, each test agent (A-TbA-T2, . . . A-Tn) has a mass of less than about 800 Da.
[0110] In some embodiments, each test agent (A-TbA-T2, . . . A-Tn) has a mass of from about 350 Da to about 800 Da.
[0111] In some embodiments, each test agent (A-TbA-T2, . . . A-Tn) has a mass of about 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, or 850 Da. In some embodiments, each test agent (A-TbA-T2, . . . A-Tn) has a mass of about 300 Da. In some embodiments, each test agent (A-TbA-T2, . . . A-Tn) has a mass of about 400 Da. In some embodiments, each test agent (A-TbA-T2, . . . A-Tn) has a mass of about 500 Da. In some embodiments, each test agent (A-TbA-T2, . . . A-Tn) has a mass of about 600 Da. In some embodiments, each test agent (A-TbA-T2, . . . A-Tn) has a mass of about 700 Da. In some embodiments, each test agent (A-TbA-T2, . . . A- Tn) has a mass of about 800 Da.
[0112] In some embodiments, each reaction is performed on a nanoscale. In some embodiments, the nanoscale is performed in a volume of about 50 nL to about 500 nL. In some embodiments, the nanoscale is performed in a volume of about 50, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nL. In some embodiments, the nanoscale is performed in a volume of about 50 nL. In some embodiments, the nanoscale is performed in a volume of about 100 nL. In some embodiments, the nanoscale is performed in a volume of about 150 nL. In some embodiments,the nanoscale is performed in a volume of about 200 nL. In some embodiments, the nanoscale is performed in a volume of about 250 nL. In some embodiments, the nanoscale is performed in a volume of about 300 nL. In some embodiments, the nanoscale is performed in a volume of about350 nL. In some embodiments, the nanoscale is performed in a volume of about 400 nL. In some embodiments, the nanoscale is performed in a volume of about 450 nL. In some embodiments, the nanoscale is performed in a volume of about 500 nL.
[0113] In some embodiments, the test plate is a 96-well, 384-well, or 1536-well test plate. In some embodiments, the test plate is a 96-well test plate. In some embodiments, the test plate is a 384-well test plate. In some embodiments, the test plate is a 1536-well test plate.
[0114] In some embodiments, each naive test moieties (TbT2, . . . Tn) is different.
[0115] In some embodiments, n is from 2 to 100,000, or any integer therein. In some embodiments,from 2 to 90,000. In some embodiments, n is from 2 to 80,000. In some embodiments,from 2 to 75,000. In some embodiments, n is from 2 to 60,000. In some embodiments,from 2 to 50,000. In some embodiments, n is from 2 to 40,000. In some embodiments, n is from 2 to 30,000. In some embodiments, n is from 2 to 25,000. In some embodiments, n is from 2 to 20,000. In some embodiments, n is from 2 to 15,000. In some embodiments, n is from 2 to 10,000. In some embodiments, n is from 2 to 8,000. In some embodiments, n is from 2 to 6,000. In some embodiments, n is from 2 to 5,000. In some embodiments, n is from 2 to 4,000. In some embodiments, n is from 2 to 3,000. In some embodiments, n is from 2 to 2,500. In some embodiments, n is from 2 to 2,000. In some embodiments, n is from 2 to 1,800 In some embodiments, n is from 2 to 1,600.
[0116] In some embodiments, step (a) is performed in the absence of target (e.g., in absence of a heterologous polypeptide described herein).
[0117] In some embodiments, said method does not utilize mass spectrometric detection. In some embodiments, said method utilizes mass spectrometric detection. Methods can include but are not limited to LCMS.
[0118] In some embodiments, the reaction mixtures of step (a) are not purified before contacting the target with one or more of the reaction mixtures.
[0119] The fragment based technology of the disclosure combines the advantages of fragment based approaches with the power and speed of high throughput screening (HTS). The technology is based on a functional screen of target-directed “made-to-order” libraries of test agents, i.e., the libraries are constructed based on target specific information. Test agents within such libraries are constructed between a target-directed fragment and a library of naive testfragments. The assembly process can be fully automated with low reagent cost and no need for purification of the assembled fragments.Barcodes
[0120] Variable nucleotide sequences (barcodes or “unique molecular identifiers”) that serve as an index can be included as a reporter gene described herein. Additionally, barcodes may be added in a separate library preparation reaction. The variable nucleotide sequences described herein can be used as a sample index in order to deconvolve results obtained from a sequencing reaction used herein. Barcodes may include an index region that is uniquely identifiable to a heterologous receptor gene, and may be used to identify the heterologous receptor that is activated in the same cell. Barcodes may also or alternatively uniquely identify a signaling pathway, tests agent, heterologous polypeptide, a location of an engineered cell-line (e.g., a well of a plate), or any combination thereof. An index region of a barcode may be continuous along the length of the barcode sequence and the barcode may include stretches of nucleic acid sequence that is not unique to any one barcode. In one embodiment, a barcode may have more than one unique identifier. In those embodiments, the index regions may be separate by a stretch of nucleic acids removedby cellular machinery during transcription into mRNA (e.g., an intron).
[0121] The term “heterologous”, in the context of polynucleotides, refers to a gene or polynucleotide or polypeptide that has been transferred to a cell by gene transfer methods known in the art or described herein; progeny of such cells may also be referred to as containing the heterologous nucleic acid sequence if the exogenously derived sequence remains in the descendant cells. The cell may already contain an endogenous gene that is identical to the heterologous receptor gene or the cell may lack any endogenous genes that are related or identical to the heterologous gene. The term "heterologous cell" or "host cell" refers to a cell intentionally containing a heterologous nucleic acid sequence.
[0122] The index region of a barcode is a polynucleotide sequence that can be used to identify the target that is activated and / or expressed in the same cell as the barcode because it is unique to a particular heterologous receptor in the context of the screen being utilized. In particular, the include ofbarcodes facilitates the determination of the activity of specific nucleic acid regulatory elements (i.e., receptor-responsive elements such as unique identifiers), which may indicate activated receptors.
[0123] Once the contents of the cells are released into their respective partitions by a lysis agent, the macromolecular components (e.g., macromolecular constituents of samples, such as RNA, DNA, or proteins) contained therein may be further processed within the partitions. Inaccordance with the methods and systems described herein, the macromolecular component contents of individual samples can be provided with unique identifiers such that, upon characterization of those macromolecular components they may be attributed as having been derived from the same sample or particles. The ability to attribute characteristics to individual samples or groups of samplesis providedby the assignment of unique identifiers specifically to an individual sample or groups of samples. Barcodesthat can be uniquely identified by a unique molecular identifier sequence (UMI or “unique identifier”) can be assigned or associated with individual samples or populations of samples, in order to tag or label the sample's macromolecular components (and as a result, its characteristics) with the unique identifiers. These unique identifiers can thenbe used to attribute the sample's components and characteristics to an individual sample or group of samples.
[0124] In some aspects, this is performed by co-partitioning the individual sample or groups of samples with the unique identifiers or barcodes comprising a UMI. In some aspects, the unique identifiers are provided in the form of nucleic acid molecules (e g., oligonucleotides) that comprise nucleic acid barcode sequences that may be attached to or otherwise associated with the nucleic acid contents of individual sample, or to other components of the sample, and particularly to fragments of those nucleic acids. The nucleic acid molecules are partitioned such that as between nucleic acid molecules in a given partition, the nucleic acid barcode sequences contained therein are the same, but as between different partitions, the nucleic acid molecule can, and do have differing barcode sequences, or at least represent a large number of different barcode sequences across all of the partitions in a given analysis. In some aspects, only one nucleic acid barcode sequence can be associated with a given partition, although in some embodiments, two or more different barcode sequences may be present.
[0125] The nucleic acid barcode sequences can include from about 6 to about 20 or more nucleotides within the sequence of the nucleic acid molecules (e.g., oligonucleotides). The nucleic acid barcode sequences can include from about 6 to about 20, 30, 40, 50, 60, 70, 80, 90, 100 or more nucleotides. In some embodiments, the length of a barcode sequence may be about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or longer. In some embodiments, the length of a barcode sequence may be at least about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or longer. In some embodiments, the length of a barcode sequence may be at most about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or shorter. These nucleotides may be completely contiguous, i.e., in a single stretch of adjacent nucleotides, or they may be separated into two or more separate subsequences that are separated by 1 or morenucleotides. In some embodiments, separated barcode subsequences can be from about 4 to about 16 nucleotides in length. In some embodiments, the barcode subsequence may be about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or longer. In some embodiments, the barcode subsequence maybe at least about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or longer. In some embodiments, the barcode subsequence may be at most about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or shorter.
[0126] The co-partitioned nucleic acid molecules can also comprise other functional sequences useful in the processing of the nucleic acids from the co-partitioned samples. These sequences include, e.g., targeted, or random / universal amplification primer sequences for amplifying the genomic DNA from the individual samples within the partitions while attaching the associated barcode sequences, sequencing primers, or primer recognition sites, hybridization, or probing sequences, e.g., for identification of presence of the sequences or for pulling down barcoded nucleic acids, or any of a number of other potential functional sequences. Other mechanisms of co-partitioning oligonucleotides may also be employed, including, e.g., coalescence of two or more partitions, where one partition contains oligonucleotides, or microdispensing of oligonucleotides into partitions, e.g., partitions within microfluidic systems. In some embodiments, a primer comprises a barcode oligonucleotide. In some embodiments the primer sequence is a targeted primer sequence complementary to a sequence in the template nucleic acid molecule. In some embodiments, the first nucleic acid molecule further comprises one or more functional sequences and wherein the second nucleic acid molecule comprises the one or more functional sequences. In some embodiments, the one or more functional sequences are selected from the group consisting of an adapter sequence, an additional primer sequence, a primer annealing sequence, a sequencing primer sequence, a sequence configured to attach to a flow cell of a sequencer, and a unique molecular identifier sequence.
[0127] For example, the above described barcoded nucleic acid molecules (e.g., barcoded oligonucleotides) are added to a sample. In some embodiments, a partition comprises barcoded oligonucleotides having the same barcode sequence. In some embodiments, a partition among a plurality of partitions comprises barcoded oligonucleotides having an identical barcode sequence, wherein each partition among within the plurality of partitions comprises a unique barcode sequence In some embodiments, the population of barcoded oligonucleotides provides a diverse barcode sequence library that includes at least about 1,000 different barcode sequences, at least about 5,000 different barcode sequences, at least about 10,000 different barcode sequences, at least about 50,000 different barcode sequences, at least about 100,000 different barcodesequences, at least about 1,000,000 different barcode sequences, at least about 5,000,000 different barcode sequences, or at least about 10,000,000 different barcode sequences, or more. Additionally, each barcoded oligonucleotide can be provided with large numbers of nucleic acid (e.g., oligonucleotide) molecules attached. In particular, the number of molecules of nucleic acid molecules including the barcode sequence on an individual barcoded oligonucleotide can be at least about 1,000 nucleic acid molecules, at least about 5,000 nucleic acid molecules, at least about 10,000 nucleic acid molecules, at least about 50,000 nucleic acid molecules, at least about 100,000 nucleic acid molecules, at least about 500,000 nucleic acids, at least about 1,000,000 nucleic acid molecules, at least about 5,000,000 nucleic acid molecules, at least about 10,000,000 nucleic acid molecules, at least about 50,000,000 nucleic acid molecules, at least about 100,000,000 nucleic acid molecules, at least about 250,000,000 nucleic acid molecules and in some embodiments at least about 1 billion nucleic acid molecules, or more. Nucleic acid molecules of a given barcoded oligonucleotide can include identical (or common) barcode sequences, different barcode sequences, or a combination of both. Nucleic acid molecules of a given barcoded oligonucleotide can include multiple sets of nucleic acid molecules. Nucleic acid molecules of a given set can include identical barcode sequences. The identical barcode sequences can be different from barcode sequences of nucleic acid molecules of another set
[0128] Moreover, when the population of barcoded oligonucleotides is partitioned, the resulting population of partitions can also include a diverse barcode library that includes at least about 1,000 different barcode sequences, at least about 5,000 different barcode sequences, at least about 10,000 different barcode sequences, at least at least about 50,000 different barcode sequences, at least about 100,000 different barcode sequences, at least about 1,000,000 different barcode sequences, at least about 5,000,000 different barcode sequences, or at least about 10,000,000 different barcode sequences. Additionally, each partition of the population can include at least about 1,000 nucleic acid molecules, at least about 5,000 nucleic acid molecules, at least about 10,000 nucleic acid molecules, at least about 50,000 nucleic acid molecules, at least about 100,000 nucleic acid molecules, at least about 500,000 nucleic acids, at least about 1,000,000 nucleic acid molecules, at least about 5,000,000 nucleic acid molecules, at least about 10,000,000 nucleic acid molecules, at least about 50,000,000 nucleic acid molecules, at least about 100,000,000 nucleic acid molecules, at least about 250,000,000 nucleic acid molecules and in some embodiments at least about 1 billion nucleic acid molecules.
[0129] In some embodiments, it may be desirable to incorporate multiple different barcodes within a given partition. For example, in some embodiments, a barcoded oligonucleotide within apartition can comprise (1) a common barcode sequence shared by all barcoded oligonucleotides within the partition and (2) a unique identifier or additional barcode sequence that is different among each barcoded oligonucleotide. The common barcode sequences may provide greater assurance of identification in the subsequent processing, e.g., by providing a stronger address or attribution of the barcodes to a given partition, as a duplicate or independent confirmation of the output from a given partition.
[0130] In some embodiments, the barcoded oligonucleotides are attached to the beads, where all of the nucleic acid molecules attached to a particular bead will include the same nucleic acid barcode sequence, but where a large number of diverse barcode sequences are represented across the population of beads used. In some embodiments, hydrogel beads, e.g., comprising polyacrylamide polymer matrices, are used as a solid support and delivery vehicle for the nucleic acid molecules into the partitions, as they are capable of carrying large numbers of nucleic acid molecules, and may be configured to release those nucleic acid molecules upon exposure to a particular stimulus, as described elsewhere herein.
[0131] The nucleic acid molecules (e.g., oligonucleotides) can be releasable from the beads upon the application of a particular stimulus to the beads. In some embodiments, the stimulus may be a photo-stimulus, e.g., through cleavage of a photo-labile linkage that releases the nucleic acid molecules. In other embodiments, a thermal stimulus may be used, where elevation of the temperature of the beads environment will result in cleavage of a linkage or other release of the nucleic acid molecules form the beads. In still other embodiments, a chemical stimulus can be used that cleaves a linkage of the nucleic acid molecules to the beads, or otherwise results in release of the nucleic acid molecules from the beads. In one embodiment, such compositions include the polyacrylamide matrices described above for encapsulation of samples, and may be degraded for release of the attached nucleic acid molecules through exposure to a reducing agent, such as DTT.
[0132] A support can be contemplated for use in a method of the present disclosure may be, for example, a well, matrix, rod, container, or bead(s). A support may have any useful features and characteristics, such as any useful size, surface chemistry, fluidity, solidity, density, porosity, and composition. In some embodiments, a support is a surface of a well on a plate. In some embodiments, a support may be a bead such as a gel bead. A bead may be solid or semi-solid. Additional details of beads are provided elsewhere herein.
[0133] A support (e.g., a bead) may comprise an anchor sequence functionalized thereto (e.g., as described herein). An anchor sequence may be attached to the support via, for example, adisulfide linkage. An anchor sequence may comprise a partial read sequence and / or flow cell functional sequence. Such a sequence may permit sequencing of nucleic acid molecules attached to the sequence by a sequencer (e.g., an Illumina sequencer). Different anchor sequences may be useful for different sequencing applications. An anchor sequence may comprise, for example, a TruSeq orNextera sequence. An anchor sequence may have any useful characteristics such as any useful length and nucleotide composition. For example, an anchor sequence may comprise 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides. In some embodiments, an anchor sequence may comprise 15 nucleotides. Nucleotides of an anchor sequence may be naturally occurring or non-naturally occurring (e.g., as described herein). A bead may comprise a plurality of anchor sequences attached thereto. For example, a bead may comprise a plurality of first anchor sequences attached thereto. In some embodiments, a bead may comprise two or more different anchor sequences attached thereto. For example, a bead may comprise both a plurality of first anchor sequences (e.g., Nextera sequences) and a plurality of second anchor sequences (e.g., TruSeq sequences) attached thereto. For a bead comprising two or more different anchor sequences attached thereto, the sequence of each different anchor sequence may be distinguishable from the sequence of each other anchor sequence at an end distal to the bead. For example, the different anchor sequences may comprise one or more nucleotide differences in the 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides furthest from the bead.
[0134] In some embodiments, multiple different barcode molecules (e.g., nucleic acid barcode molecules) may be generated on the same support (e.g., bead). For example, two different barcode molecules may be generated on the same support Alternatively, three or more different barcode molecules may be generated on the same support. Different barcode molecules attached to the same support may comprise one or more different sequences. For example, different barcode molecules may comprise one or more different barcode sequences, and / or other sequences (e.g., starter sequences). In some embodiments, different barcode molecules attached to the same support may comprise the same barcode sequences. Different barcode molecules attached to the same support may comprise barcode sequences that are the same or different. Similarly, different barcode molecules may comprise UMIs that are the same or different.
[0135] It is contemplated that in some embodiments, the sensitivity of sequencing a barcode allows for expression levels that are lower than what is needed for less sensitive assays. In some embodiments, the level of RNA transcripts is, is atleast, oris at mostabout 10, IO2, 103, 104,105, 106, 107', 108, 109, or 1010or any range derivable therein.
[0136] As used herein, the terms “cell,” “cell line,” and “cell culture” may be used interchangeably. All of these terms also include both freshly isolated cells and in vitro cultured or expanded cells. All of these terms also include their progeny, which is any and all subsequent generations. It is understood that all progeny may not be identical due to deliberate or inadvertent mutations. In the context of expressing a heterologous nucleic acid sequence, a "host cell" or simply a "cell" refers to a prokaryotic or eukaryotic cell, and it includes any transformable organism that is capable of replicating a vector or expressing a heterologous gene encoded by a vector or integrated nucleic acid. A host cell can, and has been, used as a recipient for vectors, viruses, and nucleic acids. A host cell may be "transfected" or "transformed," which refers to a process by which exogenous nucleic acid, such as a recombinant protein-encoding sequence, is transferred or introduced into the host cell. A transformed cell includes the primary subject cell and its progeny.
[0137] In certain embodiments the nucleic acid transfer can be carried out on any prokaryotic or eukaryotic cell. In some aspects the cells of the disclosure are human cells. In other aspects the cells of the disclosure are an animal cell. In some aspects the cell or cells are cancer cells, tumor cells or immortalized cells. In further aspects, the cells represent a disease-model cell. In certain aspects the cells canbe A549, B-cells, B16,BHK-21, C2C12, C6, CaCo-2, CAP / , CAP-T, CHO, CHO2, CHO-DG44, CHO-K1, COS-1, Cos-7, CV-1, Dendritic cells, DLD-1, Embryonic Stem (ES) Cell or derivative, H1299, HEK293, 293T, 293FT, Hep G2, Hematopoietic Stem Cells, HOS, Huh-7, Induced Pluripotent Stem (iPS) Cell or derivative, Jurkat, K562, L5278Y, LNCaP, MCF7, MDA-MB-231, MDCK, Mesenchymal Cells, Min-6, Monocytic cell, Neuro2a, NIH 3T3, NIH3T3L1, K562, NK-cells, NSO, Panc-1, PC12, PC-3, Peripheral blood cells, Plasma cells, Primary Fibroblasts, RBL, Renca, RLE, SF21, SF9, SH-SY5Y, SK-MES-1, SK-N-SH, SL3, SW403, Stimulus-triggered Acquisition of Pluripotency (STAP) cell or derivate SW403, T-cells, THP-1, Tumor cells, U2O5, U937, peripheral blood lymphocytes, expanded T cells, hematopoietic stem cells, or Vero cells. In some embodiments, the cells are HEK293T cells.
[0138] The term "passaged," as used herein, is intended to refer to the process of splitting cells in order to produce large number of cells from pre-existing ones. Cells may be passaged multiple times prior to or after any step described herein. Passaging involves splitting the cells and transferring a small number into each new vessel. For adherent cultures, cells first need to be detached, commonly done with a mixture of trypsin-EDTA. A small number of detached cells can then be used to seed a new culture, while the rest is discarded. Also, the amount of culturedcells can easily be enlarged by distributing all cells to fresh flasks. Cells may be kept in culture and incubated under conditions to allow cell replication. In some embodiments, the cells are kept in culture conditions that allow the cells to under 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more rounds of cell division.
[0139] In some embodiments, cells may be subjected to limiting dilution methods to enable the expansion of clonal populations of cells. The methods of limiting dilution cloning are well known to those of skill in the art. Such methods have been described, for example for hybridomas but canbe applied to any cell. Such methods are described in (Cloning hybridoma cells by limiting dilution, Journal of tissue culture methods, 1985, Volume 9, Issue 3, pp 175177, by Joan C. Rener, Bruce L. Brown, and Roland M. Nardone) which is incorporated by reference herein.
[0140] Methods of the disclosure include the culturing of cells. Methods of culturing suspension and adherent cells are well-known to those skilled in the art. In some embodiments, cells are cultured in suspension, using commercially available cell-culture vessels and cell culture media. Examples of commercially available culturing vessels that may be used in some embodiments including ADME / TOX Plates, Cell Chamber Slides and Coverslips, Cell Counting Equipment, Cell Culture Surfaces, Coming HYPERFlask Cell Culture Vessels, Coated Cultureware, Nalgene Cryoware, Culture Chamber, Culture Dishes, Glass Culture Flasks, Plastic Culture Flasks, 3D Culture Formats, Culture Multiwell Plates, Culture Plate Inserts, Glass Culture Tubes, Plastic Culture Tubes, Stackable Cell Culture Vessels, Hypoxic Culture Chamber, Petri dish and flask carriers, Quickfit culture vessels, Scale-Up Cell Culture using Roller Bottles, Spinner Flasks, 3D Cell Culture, or cell culture bags.
[0141] In other embodiments, media may be formulated using components well-known to those skilled in the art. Formulations and methods of culturing cells are described in detail in the following references: Short Protocols in Cell Biology J. Bonifacino, et al., ed., John Wiley & Sons, 2003, 826 pp; Live Cell Imaging: A Laboratory Manual D. Spector & R. Goldman, ed., Cold Spring Harbor Laboratory Press, 2004, 450 pp.; Stem Cells Handbook S. Sell, ed., Humana Press, 2003, 528 pp.; Animal Cell Culture: Essential Methods, John M. Davis, John Wiley & Sons, Mar 16, 2011 ; Basic Cell Culture Protocols, Cheryl D. Helgason, Cindy Miller, Humana Press, 2005; Human Cell Culture Protocols, Series: Methods in Molecular Biology, Vol. 806, Mitry, Ragai R.; Hughes, Robin D. (Eds.), 3rd ed. 2012, XIV, 435 p. 89, Humana Press; Cancer Cell Culture: Method and Protocols, Cheryl D. Helgason, Cindy Miller, Humana Press, 2005; Human Cell Culture Protocols, Series: Methods in Molecular Biology, Vol. 806, Mitry, Ragai R.;Hughes, Robin D. (Eds.), 3rd ed. 2012, XIV, 435 p. 89, Humana Press; Cancer Cell Culture: Method and Protocols, Simon P. Langdon, Springer, 2004; Molecular Cell Biology. 4th edition., Lodish H, Berk A, Zipursky SL, et al., New York: W. H. Freeman; 2000., Section 6.2Growth of Animal Cells in Culture, all of which are incorporated herein by reference.Sequencing Methods to Detect BarcodesMassively Parallel Signature Sequencing (MPSS)
[0142] The first of the next-generation sequencing technologies, massively parallel signature sequencing (or MPSS), was developed in the 1990s at Lynx Therapeutics. MPSS was a beadbased method thatused a complex approach of adapter ligation followed by adapter decoding, reading the sequence in increments of four nucleotides. This method made it susceptible to sequence-specific bias or loss of specific sequences. Because the technology was so complex, MPSS was only performed 'in-house' by Lynx Therapeutics and no DNA sequencing machines were sold to independent laboratories. Lynx Therapeutics merged with Solexa (later acquired by Illumina) in 2004, leading to the development of sequencing-by-synthesis, a simpler approach acquired from Manteia Predictive Medicine, which rendered MPSS obsolete. However, the essential properties of the MPSS output were typical of later "next-generation" data types, including hundreds of thousands of short DNA sequences. In the case of MPSS, these were typically used for sequencing cDNA for measurements of gene expression levels. Indeed, the powerful Illumina HiSeq2000, HiSeq2500 and MiSeq systems are based on MPSS.Polony Sequencing
[0143] The Polony sequencing method, developed in the laboratory of George M. Church at Harvard, was among the first next-generation sequencing systems and was used to sequence a full genome in 2005. It combined an in vitro paired-tag library with emulsion PCR, an automated microscope, and ligation-based sequencing chemistry to sequence an E. coli genome at an accuracy of >99.9999% and a cost approximately 1 / 9 that of Sanger sequencing. The technology was licensed to Agencourt Biosciences, subsequently spun out into Agencourt Personal Genomics, and eventually incorporated into the Applied Biosystems SOLID platform, which is now owned by Life Technologies.454 pyrosequencing.
[0144] A parallelized version of pyrosequencing was developed by 454 Life Sciences, which has since been acquired by Roche Diagnostics. The method amplifies DNA inside water droplets in an oil solution (emulsion PCR), with each droplet containing a single DNA template attached to a single primer-coated bead that then forms a clonal colony. The sequencing machine containsmany picoliter-volume wells each containing a single bead and sequencing enzymes. Pyrosequencing uses luciferase to generate light for detection of the individual nucleotides added to the nascent DNA, and the combined data are used to generate sequence read-outs. This technology provides intermediate read length and price per base compared to Sanger sequencing on one end and Solexa and SOLiD on the other.Illumina (Solexa) Sequencing
[0145] Solexa, now part of Illumina, developed a sequencing method based on reversible dye- terminators technology, and engineered polymerases, that it developed internally. The terminated chemistry was developed internally at Solexa and the concept of the Solexa system was invented by Balasubramanian and Klennerman from Cambridge University's chemistry department. In 2004, Solexa acquired the company Manteia Predictive Medicine in order to gain a massively parallel sequencing technology based on "DNA Clusters", which involves the clonal amplification of DNA on a surface. The cluster technology was coacquired with Lynx Therapeutics of California. Solexa Ltd. later merged with Lynx to form Solexa Inc.
[0146] In this method, DNA molecules and primers are first attached on a slide and amplified with polymerase so that local clonal DNA colonies, later coined "DNA clusters", are formed. To determine the sequence, four types of reversible terminator bases (RT-bases) are added and nonincorporated nucleotides are washed away. A camera takes images of the fluorescently labeled nucleotides, then the dye, along with the terminal 3' blocker, is chemically removed from the DNA, allowing for the next cycle to begin. Unlike pyrosequencing, the DNA chains are extended one nucleotide at a time and image acquisition can be performed at a delayed moment, allowing for very large arrays of DNA colonies to be captured by sequential images taken from a single camera.
[0147] Decoupling the enzymatic reaction and the image capture allows for optimal throughput and theoretically unlimited sequencing capacity. With an optimal configuration, the ultimately reachable instmment throughput is thus dictated solely by the analog-to-digital conversion rate of the camera, multiplied by the number of cameras and divided by the number of pixels per DNA colony required for visualizing them optimally (approximately 10 pixels / colony). In 2012, with cameras operating at more than 10 MHz A / D conversion rates and available optics, fluidics and enzymatics, throughput can be multiples of 1 million nucleotides / second, corresponding roughly to one human genome equivalent at lx coverage per SOLiD sequencing
[0148] Applied Biosystems' (now a Life Technologies brand) SOLiD technology employs sequencing by ligation. Here, a pool of all possible oligonucleotides of a fixed length are labeled according to the sequenced position. Oligonucleotides are annealed and ligated; the preferential ligation by DNA ligase for matching sequences results in a signal informative of the nucleotide at that position. Before sequencing, the DNA is amplified by emulsion PCR. The resulting beads, each containing single copies of the same DNA molecule, are deposited on a glass slide. The result is sequences of quantities and lengths comparable to Illumina sequencing. This sequencing by ligation method has been reported to have some issue sequencing palindromic sequences. Ion Torrent Semiconductor Sequencing
[0149] Ion Torrent Systems Inc. (now owned by Life Technologies) developed a system based on using standard sequencing chemistry, but with a novel, semiconductor based detection system. This method of sequencing is based on the detection of hydrogen ions that are released during the polymerization of DNA, as opposed to the optical methods used in other sequencing systems. A microwell containing a template DNA strand to be sequenced is flooded with a single type of nucleotide. If the introduced nucleotide is complementary to the leading template nucleotide it is incorporated into the growing complementary strand. This causes the release of a hydrogen ion that triggers a hypersensitive ion sensor, which indicates that a reaction has occurred. If homopolymer repeats are present in the template sequence multiple nucleotides will be incorporated in a single cycle. This leads to a corresponding number of released hydrogens and a proportionally higher electronic signal.DNA Nanoball Sequencing
[0150] DNA nanoball sequencing is a type of high throughput sequencing technology used to determine the entire genomic sequence of an organism. The company Complete Genomics uses this technology to sequence samples submitted by independent researchers. The method uses rolling circle replication to amplify small fragments of genomic DNA into DNA nanoballs. Unchained sequencing by ligation is then used to determine the nucleotide sequence. This method of DNA sequencing allows large numbers of DNA nanoballs to be sequenced per run and at low reagent costs compared to other next generation sequencing platforms. However, only short sequences of DNA are determined from each DNA nanoball which makes mapping the short reads to a reference genome difficult. This technology has been used for multiple genome sequencing projects and is scheduled to be used for more.Heliscope Single Molecule Sequencing
[0151] Heliscope sequencing is a method of single-molecule sequencing developed by Helicos Biosciences. It uses DNA fragments with added poly -A tail adapters which are attached to the flow cell surface. The next steps involve extension-based sequencing with cyclic washes of the flow cell with fluorescently labeled nucleotides (one nucleotide type at a time, as with the Sanger method). The reads are performed by the Heliscope sequencer. The reads are short, up to 55 bases per run, but recent improvements allow for more accurate reads of stretches of one type of nucleotides. This sequencing method and equipment were used to sequence the genome of the Ml 3 bacteriophage.Single Molecule Real Time (SMRT) Sequencing
[0152] SMRT sequencing is based on the sequencing by synthesis approach. The DNA is synthesized in zero-mode wave-guides (ZMWs) — small well-like containers with the capturing tools located at the bottom of the well. The sequencing is performed with use of unmodified polymerase (attached to the ZMW bottom) and fluorescently labelled nucleotides flowing freely in the solution. The wells are constructed in a way that only the fluorescence occurring by the bottom of the well is detected. The fluorescent label is detached from the nucleotide at its incorporation into the DNA strand, leaving an unmodified DNA strand. According to Pacific Biosciences, the SMRT technology developer, this methodology allows detection of nucleotide modifications (such as cytosine methylation). This happens through the observation of polymerase kinetics. This approach allows reads of 20,000 nucleotides or more, with average read lengths of 5 kilobases.Next generation sequencing
[0153] As described in the methods disclosed herein, the sequencing of nucleic acid molecules is used and is useful for the detection of biological effect by a test agent against a cell comprising a cell based assay . Generally, sequencing refers to methods and technologies for determining the sequence of nucleotide bases in one or more polynucleotides. The polynucleotides can be, for example, deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), including variants or derivatives thereof (e.g., single stranded DNA). Sequencing can be performed by various systems currently available, such as, without limitation, a sequencing system by Illumina, Pacific Biosciences, Oxford Nanopore, or Life Technologies (Ion Torrent). Such devices may provide a plurality of raw genetic data corresponding to the genetic information of a subject (e.g., human), as generated by the device from a sample provided by the subject. In some situations, systems and methods provided herein may be used with proteomic information. Alternatively, or in addition, sequencing may be performed using nucleic acidamplification, polymerase chain reaction (PCR) (e g., digital PCR, quantitative PCR, or real time PCR), or isothermal amplification. Such systems may provide a plurality of raw genetic data corresponding to the genetic information of a subject (e.g., human), as generated by the systems from a sample provided by the subject. In some examples, such systems provide sequencing reads (also “reads” herein). A read may include a string of nucleic acid bases corresponding to a sequence of a nucleic acid molecule that has been sequenced. In some situations, systems and methods provided herein may be used with proteomic information.
[0154] Next generation sequencing includes many technologies capable of generating large amounts of sequence information and excluding Sanger sequencing or Maxam-Gilbert sequencing. Generally, next generation sequencing encompasses single molecule real-time sequencing, sequencing-by-synthesis, ion semiconductor sequencing and the like. Exemplary next-generation sequencing machines may comprise the MiniSeq, the iSeqlOO, the NextSeq 1000, the NextSeq 2000, the NovaSeq 6000, the NextSeq 550 series and the like from Illumina, Inc; Ion Torrent machines from Thermo Fisher Scientific; or the Sequel systems from Pacific Biosciences.
[0155] Next generation sequencing machines used with the method herein can generate at least 1, 5, 10, 15, 25, 50, 75, 100, 200, 300 gigabases of data ormore in a 24 hour period from a single machine.
[0156] Next generation sequencing machines used with the method herein can generate at least 1, 1, 4, 10, 15, 25, 50, 75, 100, 200, 300, 500, or 1,000 million sequence reads of data or more in a 24 hour period from a single machine.
[0157] Also included is a computer program, computing device, or analysis platform / system to receive and analyze sequencing data, and output one or more reports that can be transmitted or accessed electronically via a server, an analysis portal, or by e-mail. The computing device or analysis platform can operate according to the algorithms and methods described herein.Reaction Mixtures
[0158] Also provided herein are reaction mixtures for determining the expression level of a reporter in a sample by sequencing. In some embodiments, the reaction mixture comprises a control nucleic acid provided herein, at least a portion of said biological sample, and one or more enzyme or reagents sufficient to amplify a barcode in a sample, if present.
[0159] The control nucleic acid may be any one or more of a single stranded DNA, a double stranded DNA, a single stranded RNA, or a double stranded RNA. In some embodiments, said control nucleic acid is present at a concentration of about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100,110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 320, 340, 360, 380, 400, 420, 440, 460, 480, or 500 copies per rection mixture.
[0160] In some embodiments, the enzymes or reagents comprise a reverse transcriptase enzyme, dNTPs, a primer pair specific for a barcode or control nucleic acid sequence, a magnesium salt, or combinations thereof.
[0161] In some embodiments, the reaction mixture comprises one or more enzymes which can be used to amplify or replicate a control nucleic acid or a barcode nucleic acid. In some embodiments, the enzyme is a reverse transcriptase. Non-limiting examples of reversetranscriptase enzymes include Avian Myeloblastosis Virus (AMV) Reverse Transcriptase and Moloney Murine Leukemia Virus (M-MuLV, MMLV), and variants thereof.
[0162] In some embodiments, the reaction comprises deoxynucleotide triphosphates (dNTPs). In some embodiments, the kit comprises a mixture of each of the dNTPs necessary for amplification of nucleic acids, as well as any other desired nucleic acids (e.g., dATG, dCTP, dTTP, dGTP).
[0163] In some embodiments, the reaction mixture comprises a magnesium salt. In some embodiments, the magnesium salt is included in a sufficient quantity to allow the enzymes of the reaction (e.g., the reverse transcriptase enzyme) to function and to amplify targeted nucleic acids. In some embodiments, the magnesium salt is magnesium chloride. In some embodiments, the reaction mixture comprises a concentration of magnesium ions of about 0.1 mMto about 50 mM In some embodiments, the concentration of magnesium ion is from about 1 mM to about 10 mM.
[0164] In some embodiments, the volume of said reaction mixture is from about 10 microliters to about 100 microliters. In some embodiments, the volume of said reaction mixture is from about 20 microliters to about 90 microliters, from about 30 microliters to about 80 microliters, or from about 40 microliters to about 60 microliters. In some embodiments, the volume of said reaction mixture is about 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 microliters.Further Forms of Compounds
[0165] In some aspects, a compound disclosed herein possesses one or more stereocenters and each stereocenter exists independently in either the R or S configuration. The compounds presented herein include all diastereomeric, enantiomeric, and epimeric forms as well as the appropriate mixtures thereof. The compounds and methods provided herein include all cis, trans, syn, anti, entgegen (E), and zusammen (Z) isomers as well as the appropriate mixtures thereof. In certain embodiments, compounds described herein are prepared as their individual stereoisomers by reacting a racemic mixture of the compound with an optically active resolving agent to form apair of diastereoisomeric compounds / salts, separating the diastereomers and recovering the optically pure enantiomers. In some embodiments, resolution of enantiomers is carried out using covalent diastereomeric derivativesof the compounds described herein. In another embodiment, diastereomers are separated by separation / resolution techniques based upon differences in solubility. In other embodiments, separation of stereoisomers is performed by chromatography or by the forming diastereomeric salts and separation by recrystallization, or chromatography, or any combination thereof. Jean Jacques, Andre Collet, Samuel H. Wilen, “Enantiomers, Racemates and Resolutions”, John Wiley And Sons, Inc., 1981. In one aspect, stereoisomers are obtained by stereoselective synthesis.
[0166] In some embodiments, compounds described herein are prepared as prodrugs. A “prodrug” refers to an agent that is converted into the parent drug in vivo. Prodrugs are often useful because, in some situations, they may be easier to administer than the parent drug. They may, for instance, be bioavailableby oral administration whereas the parent is not. The prodrug may also have improved solubility in pharmaceutical compositions over the parent drug. In some embodiments, the design of a prodrug increases the effective water solubility. An example, without limitation, of a prodrug is a compound described herein, which is administered as an ester (the “prodrug”) to facilitate transmittal across a cell membrane where water solubility is detrimental to mobility but which then is metabolically hydrolyzed to the carboxylic acid, the active entity, once inside the cell where water-solubility is beneficial. A further example of a prodrug might be a shortpeptide (polyaminoacid) bonded to an acid group where the peptide is metabolized to reveal the active moiety. In certain embodiments, upon in vivo administration, a prodrug is chemically converted to the biologically, pharmaceutically or therapeutically active form of the compound. In certain embodiments, a prodrug is enzymatically metabolized by one or more steps or processes to the biologically, pharmaceutically or therapeutically active form of the compound.
[0167] In one aspect, prodrugs are designed to alter the metabolic stability or the transport characteristics of a drug, to mask side effects or toxicity, to improve the flavor of a drug or to alter other characteristics or properties of a drug. By virtue of knowledge of pharmacokinetic, pharmacodynamic processes and drug metabolism in vivo, once a pharmaceutically active compound is known, the design of prodrugs of the compound is possible.
[0168] In some embodiments, some of the herein-described compounds may be a prodrug for another derivative or active compound.
[0169] In some embodiments, sites on the aromatic ring portion of compounds described herein are susceptible to various metabolic reactions Therefore incorporation of appropriate substituents on the aromatic ring structures will reduce, minimize or eliminate this metabolic pathway. In specific embodiments, the appropriate substituent to decrease or eliminate the susceptibility of the aromatic ring to metabolic reactions is, by way of example only, a halogen, or an alkyl group.
[0170] In another embodiment, the compounds described herein are labeled isotopically (e.g, with a radioisotope) or by another other means, including, but not limited to, the use of chromophores or fluorescent moieties, bioluminescent labels, or chemiluminescent labels.
[0171] Compounds described herein include isotopically-labeled compounds, which are identical to those recited in the various formulae and structures presented herein, but for the fact that one or more atoms are replaced by an atom having an atomic mass or mass number different from the atomic mass or mass number usually found in nature. Examples of isotopes that can be incorporated into the present compounds include isotopes of hydrogen, carbon, nitrogen, oxygen, sulfur, fluorine, chlorine, and iodine such as, for example,2H,3H,13C,14C,15N,180,17O,35S,18F,36C1, and123I. In one aspect, isotopically-labeled compounds described herein, for example those into which radioactive isotopes such as3H and14C are incorporated, are useful in drug and / or substrate tissue distribution assays. In one aspect, substitution with isotopes such as deuterium affords certain therapeutic advantages resulting from greater metabolic stability, such as, for example, increased in vivo half-life or reduced dosage requirements.
[0172] In additional or further embodiments, the compounds described herein are metabolized upon administration to an organism in need to produce a metabolite that is then used to produce a desired effect, including a desired therapeutic effect.
[0173] “Pharmaceutically acceptable” as used herein, refers a material, such as a carrier or diluent, which does not abrogate the biological activity or properties of the compound, and is relatively nontoxic, i.e., the material may be administered to an individual without causing undesirable biological effects or interacting in a deleterious manner with any of the components of the composition in which it is contained.
[0174] The term “pharmaceutically acceptable salt” refers to a formulation of a compound that does not cause significant irritation to an organism to which it is administered and does not abrogate the biological activity and properties of the compound. In some embodiments, pharmaceutically acceptable salts are obtained by reacting a compound disclosed herein withacids. Pharmaceutically acceptable salts are also obtained by reacting a compound disclosed herein with a base to form a salt.
[0175] Compounds described herein may be formed as, and / or used as, pharmaceutically acceptable salts. The type of pharmaceutical acceptable salts, include, but are not limited to: (1) acid addition salts, formed by reacting the free base form of the compound with a pharmaceutically acceptable: inorganic acid, such as, for example, hydrochloric acid, hydrobromic acid, sulfuric acid, phosphoric acid, metaphosphoric acid, and the like; or with an organic acid, such as, for example, acetic acid, propionic acid, hexanoic acid, cyclopentanepropionic acid, glycolic acid, pyruvic acid, lactic acid, malonic acid, succinic acid, malic acid, maleic acid, fumaric acid, trifluoroacetic acid, tartaric acid, citric acid, benzoic acid, 3-(4-hydroxybenzoyl)benzoic acid, cinnamic acid, mandelic acid, methanesulfonic acid, ethanesulfonic acid, 1,2-ethanedisulfonic acid, 2 -hydroxy ethanesulfonic acid, benzenesulfonic acid, toluenesulfonic acid, 2-naphthalenesulfonic acid, 4-methylbicyclo-[2.2.2]oct-2-ene-l- carboxylic acid, glucoheptonic acid, 4,4’-methylenebis-(3-hydroxy-2-ene-l -carboxylic acid), 3- phenylpropionic acid, trimethylacetic acid, tertiary butylacetic acid, lauryl sulfuric acid, gluconic acid, glutamic acid, hydroxynaphthoic acid, salicylic acid, stearic acid, muconic acid, butyric acid, phenylacetic acid, phenylbutyric acid, valproic acid, and the like; (2) salts formed when an acidic proton present in the parent compound is replaced by a metal ion, e.g., an alkali metal ion (e.g., lithium, sodium, potassium), an alkaline earth ion (e.g., magnesium, or calcium), or an aluminum ion. In some cases, compounds described herein may coordinate with an organic base, such as, but not limited to, ethanolamine, diethanolamine, triethanolamine, tromethamine, N- methylglucamine, dicyclohexylamine, tris(hydroxymethyl)methylamine. In other cases, compounds described herein may form salts with amino acids such as, but not limited to, arginine, lysine, and the like. Acceptable inorganic bases used to form salts with compounds that include an acidic proton, include, but are not limited to, aluminum hydroxide, calcium hydroxide, potassium hydroxide, sodium carbonate, sodium hydroxide, and the like.
[0176] It should be understood that a reference to a pharmaceutically acceptable salt includes the solvent addition forms, particularly solvates. Solvates contain either stoichiometric or non- stoichiometric amounts of a solvent, and may be formed during the process of crystallization with pharmaceutically acceptable solvents such as water, ethanol, and the like. Hydrates are formed when the solvent is water, or alcoholates are formed when the solvent is alcohol. Solvates of compounds described herein can be conveniently prepared or formed during the processes described herein. In addition, the compounds provided herein can exist in unsolvated as well assolvated forms. In general, the solvated forms are considered equivalent to the unsolvated forms for the purposes of the compounds and methods provided herein.Definitions
[0177] Unless specifically noted otherwise herein, the definitions of the terms used are standard definitions used in the art of organic and peptide synthesis, medicinal chemistry, and pharmaceutical sciences.
[0178] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the content clearly dictates otherwise. It should also be noted that the term “or” is generally employed in its sense including “and / or” unless the content clearly dictates otherwise. Further, headings provided herein are for convenience only and do not interpret the scope or meaning of the claimed invention.
[0179] As used herein the term “test agent” refers to a molecular compound of any size and includes small molecule compounds, peptides, polypeptides, antibodies, nucleic acid molecules and the like.
[0180] As used herein, the term “biological activity” “or cellular function” refers to any change in biological state of a cell including but not limited to gene expression, cell growth and proliferation, metabolic activity of the cell, and protein RNA or DNA synthesis by the cell.
[0181] The term “target” means a chemical or biological entity for which a ligand or molecule has intrinsic binding affinity. The target can be a molecule, a portion of a molecule, or an aggregate of molecules. Specific examples of targets include polypeptides, proteins, ligands for receptors, allosteric enzyme regulators, immunoglobulins, polynucleotides, carbohydrates, glycolipids, and other macromolecules, such as protein complexes, nucleic acid-protein complexes, chromatin, ribosomes, lipid bilayer-containing structures, such as membranes, or structures derived from membranes, such as vesicle. Proteins and protein complexes include, but are not limited to, cell-surface receptors, nuclear receptors, enzymes, receptor tyrosine kinases, cell signaling proteins, cytokines, chemokines, structural or organellular proteins.
[0182] As used herein, “protein” means any molecule comprising two or more peptide units, each comprising an amino acid residue, arranged in a linear chain and joined together by peptide bonds. Protein chains comprising more than 30 amino acid residues may be referred to as polypeptides. Protein chains of 30 amino acid residues or fewer may be referred to as oligopeptides. Proteins include, but are not limited to, enzymes (e.g., cysteine protease, serine protease, and aspartyl proteases), receptors, transcription factors, growth factors, cytokines,immunoglobulins, nuclear proteins, signal transduction components (e.g., kinases, phosphatases), and glycoproteins.
[0183] A “ligand” as defined herein is a molecule that has an intrinsic binding affinity for the target. Ligands are typically small organic molecules that have an intrinsic binding affinity for the target, but may also be other sequence-specific binding molecules, such as peptides (D-, L-, or a mixture of D- and L-), peptidomimetics, complex carbohydrates, antibodies, or other oligomeric molecules with the capacity to bind specifically to the target.
[0184] A “promoter” is a control sequence. The promoter is typically a region of a nucleic acid sequence at which initiation and rate of transcription are controlled. It may contain genetic elements at which regulatory proteins and molecules may bind such as RNA polymerase and other transcription factors. The phrases “operatively positioned,” “operatively linked,” “under control,” and “under transcriptional control” mean that a promoter is in a correct functional location and / or orientation in relation to a nucleic acid sequence to control transcriptional initiation and expression of that sequence. A promoter may or may not be used in conjunction with an “enhancer,” which refers to a cis-acting regulatory sequence involved in the transcriptional activation of a nucleic acid sequence. The term promoter is used interchangeably with the term “response element.” Response elements useful in the methods described herein include: cAMP response element (CRE), a nuclear factor of activated T-cells response element (NFAT-RE), serum response element (SRE), and serum response factor response element (SRF- RE), androgen response element (ARE), glucocorticoid response element (GRE), HRE (Hormone Response Element), ERE (Estrogen Response Element), and combinations thereof As used herein “measurable” in reference to binding affinity or other affinity parameter means that a value for the affinity parameter is reliably detectable for a ligand of the target. The skilled person will understand that different affinity parameters may be measured with different degrees of precision and accuracy. Ideally, the, precision, accuracy, and dynamic range of an assay will easily accommodate a range of values, so that ligands exhibiting a wide range of measured values for the affinity parameter can be studied. The skilled person will often establish thresholds against which a given test result may be said to be meaningful. For example, in an assay of enzyme inhibition, the concentration of a putative inhibitor of the enzyme may be required to be below a preselected concentration to be considered to be meaningful. To illustrate, an IC50 threshold may be established for a biochemical assay of a target protein. A target moiety may be preselected that does not itself meet the threshold, but which shows a weaker IC50. Then, pursuant to a screening procedure according to the invention, one or more test agents may beassessed as being more potent and meeting the IC50threshold. In other cases, a degree of improvement in potency over that of the bait moiety, e.g., a 10-fold improvement, may be the threshold chosen for the assessment.
[0185] The term “monophore” as used herein means a monomeric unit of a test agent. The term “diaphore” denotes two monophores covalently linked to form a unit, i.e., a test agent, that, ideally, has a higher affinity for the target insofar as the two constituent monophores bind to two separate but nearby sites on the target. The binding affinity of a diaphore (test agent), which is a product of the affinities of the individual monophores, maybe referred to as “avidity.” The term “diaphore” is used irrespective of whether the unit is covalently bound to the target or exists separately after its release from the target.
[0186] “ Small molecules” are usually about 2,000 Da molecular weight or less, and include but are not limited to synthetic organic or inorganic compounds, peptides, (poly)nucleotides, (oligo)saccharides and the like. Small molecules specifically include inter alia small non- polymeric (e.g., not peptide or polypeptide) organic and inorganic molecules. Many pharmaceutical companies have extensive libraries of such molecules, which can be conveniently used in the methods of the invention. In one embodiment, small molecules have molecular weights of up to about 1,000 Da. In another embodiment small molecules have molecular weights of less than about 650 Da. In one embodiment, small molecules have molecular weights of up to about 300 Da. Included within this definition are small organic (including non- polymeric) molecules containing metals such as Zn, Hg, Fe, Cd, and As which may form a bond with nucleophiles.
[0187] A “site” on a target refers to a site to which a specific ligand binds, which may include a specific sequence of monomeric subunits, e g., amino acid residues, or nucleotides, and may have a characterized three-dimensional structure. Typically, the molecular interactions between the ligand and the site of interest on the target are non-covalent, and include hydrogen bonds, van der Waals interactions and electrostatic interactions. In the case of polypeptides a site of interest broadly includes the amino acid residues involved in binding of the target to a molecule with which it forms a natural complex in vivo or in vitro.
[0188] When, for example, the target is a protein that exerts its biological effect through binding to another protein, such as with hormones, cytokines or other proteins involved in signaling, it may form a natural complex in vivo with one or more other proteins. In this case, the site of interest is defined as the critical contact residues involved in a particular protein :protein binding interface. Critical contact residues are defined as those amino acids on a first protein thatmake direct contact with amino acids on a second protein, and when mutated to alanine decrease the binding affinity by at least 10-fold, alternately at least 20-fold, as measured with a direct binding or competition assay (e g. ELISA).
[0189] The term “antagonist” is used in the broadest sense and includes any ligand that partially or fully blocks, inhibits or neutralizes a biological activity exhibited by a target.
[0190] The term “agonist” is used in the broadest sense and includes any ligand that mimics a biological activity exhibited by a target, such as a target, for example, by specifically changing the function or expression of such target, or the efficiency of signaling through such target, thereby altering (increasing or inhibiting) an already existing biological activity or triggering a new biological activity.
[0191] The phrase “adjustingthe conditions” as usedherein means subjecting a target to any individual, combination, or series of reaction conditions or reagents necessary to cause a covalent bond to form between the ligand and the target, or to break a covalent bond already formed.
[0192] “ Active” or “activity” means a measurable, quantitative biological and / or immunological property. Examples of biological activities for cells include protein-protein binding, transcriptional activity, cell growth and division, protein synthesis and folding, and catalytic activity of enzymes.
[0193] “Derivative” as used herein means a compound obtained from another compound (i.e., a “parent” compound) and containing essential elements of the parent compound, or is a compound related structurally to such parent compound. “Derivative” encompasses compounds that may be obtained directly from the parent compound, or that may be obtained from a common intermediate thereto using analogous chemical methods. For example, adenine is a derivative of purine.
[0194] The terms below, as used herein, have the following meanings, unless indicated otherwise:
[0195] “ Oxo” refers to the =0 substituent.
[0196] “Alkyl” refers to a straight or branched hydrocarbon chain radical, having from one to twenty carbon atoms, and which is attached to the rest of the molecule by a single bond. An alkyl comprising up to 10 carbon atoms is referred to as a Ci-Cio alkyl, likewise, for example, an alkyl comprising up to 6 carbon atoms is a Ci-Cg alkyl. Alkyls (and other moieties defined herein) comprising other numbers of carbon atoms are represented similarly. Alkyl groups include, but are not limited to, Ci-Cio alkyl, C1-C9 alkyl, Ci-C8alkyl, C1-C7 alkyl, Ci-Cg alkyl, C C5alkyl, C C4alkyl, C1-C3 alkyl, Ci-C2alkyl, C2-C8alkyl, C3-C8alkyl and C4-C8alkyl.Representative alkyl groups include, but are not limited to, methyl, ethyl, / / -propyl, 1 -methylethyl ( / -propyl), / / -butyl, z-butyl, -butyl, / / -pentyl, 1,1 -dim ethylethyl ( / -butyl), 3 -methylhexyl, 2-methylhexyl, 1 -ethyl -propyl, and the like. In some embodiments, the alkyl is methyl or ethyl. Unless stated otherwise specifically in the specification, an alkyl group may be optionally substituted as described below.
[0197] “Alkylene” refers to a straight or branched divalent hydrocarbon chain linking the rest of the molecule to a radical group. In some embodiments, the alkylene is -CH2-, -CH2CH2-, or - CH2CH2CH2-. In some embodiments, the alkylene is -CH2-. In some embodiments, the alkylene is -CH2CH2-. In some embodiments, the alkylene is -CH2CH2CH2-.
[0198] “Alkoxy” refers to a radical of the formula -OR where R is an alkyl radical as defined. Unless stated otherwise specifically in the specification, an alkoxy group may be optionally substituted as described below. Representative alkoxy groups include, but are not limited to, methoxy, ethoxy, propoxy, butoxy, pentoxy. In some embodiments, the alkoxy is methoxy. In some embodiments, the alkoxy is ethoxy.
[0199] “Heteroalkyl” refers to an alkyl radical as described above where one or more carbon atoms of the alkyl is replaced with a O, N (i.e., NH, N-alkyl) or S atom. “Heteroalkylene” refers to a straight or branched divalent heteroalkyl chain linking the rest of the molecule to a radical group. Unless stated otherwise specifically in the specification, the heteroalkyl or heteroalkylene group may be optionally substituted as described below. Representative heteroalkyl groups include, but are not limited to -OCH2OMe, -OCH2CH2OMe, or -OCH2CH2OCH2CH2NH2.Representative heteroalkylene groups include, but are not limited to -OCH2CH2O-, - OCH2CH2OCH2CH2O-, or -OCH2CH2OCH2CH2OCH2CH2O-.
[0200] “Alkylamino” refers to a radical of the formula -NHR or -NRR where each R is, independently, an alkyl radical as defined above. Unless stated otherwise specifically in the specification, an alkylamino group may be optionally substituted as described below.
[0201] The term “aromatic” refers to a planar ring having a delocalized K-electron system containing 4n+2 z electrons, where n is an integer. Aromatics can be optionally substituted. The term “aromatic” includes both aryl groups (e g., phenyl, naphthalenyl) and heteroaryl groups (e.g., pyridinyl, quinolinyl).
[0202] “Aryl” refers to an aromatic ring wherein each of the atoms forming the ring is a carbon atom. Aryl groups can be optionally substituted. Examples of aryl groups include, but are not limited to phenyl, and naphthyl. In some embodiments, the aryl is phenyl. Depending on the structure, an aryl group can be a monoradical or a diradical (i.e., an arylene group). Unless statedotherwise specifically in the specification, the term “aryl” or the prefix “ar-” (such as in “aralkyl”) is meant to include aryl radicals that are optionally substituted.
[0203] “Carboxy” refers to -CO2H. In some embodiments, carboxy moieties may be replaced with a “carboxylic acid bioisostere”, which refers to a functional group or moiety that exhibits similar physical and / or chemical properties as a carboxylic acid moiety. A carboxylic acid bioisostere has similar biological properties to that of a carboxylic acid group. A compound with a carboxylic acid moiety canhavethe carboxylic acid moiety exchanged with a carboxylic acid bioisostere and have similar physical and / or biological properties when compared to the carboxylic acid-containing compound. For example, in one embodiment, a carboxylic acid bioisostere would ionize at physiological pH to roughly the same extent as a carboxylic acid group. Examples of bioisosteres of a carboxylic acid include, but are not limited to:
[0204] “Cycloalkyl” refers to a monocyclic or polycyclic non-aromatic radical, wherein each of the atoms forming the ring (i.e., skeletal atoms) is a carbon atom. Cycloalkyls may be saturated, or partially unsaturated. Cycloalkyls maybe fused with an aromatic ring (in which case the cycloalkyl is bonded through a non-aromatic ring carbon atom). Cycloalkyl groups include groups having from 3 to 10 ring atoms. Representative cycloalkyls include, but are not limited to, cycloalkyls having from three to ten carbon atoms, from three to eight carbon atoms, from three to six carbon atoms, or from three to five carbon atoms. In some embodiments, a cycloalkyl is a Cs-Cecycloalkyl. In some embodiments, the cycloalkyl is monocyclic, bicyclic or polycyclic. In some embodiments, cycloalkyl groups are selected from among cyclopropyl, cyclobutyl, cyclopentyl, cyclopentenyl, cyclohexyl, cyclohexenyl, cycloheptyl, cyclooctyl, spiro[2.2]pentyl, bicyclo[l.l.l]pentyl, bicyclo[3.3.0]octane, bicyclo[4.3.0]nonane, bicyclo[2. 1. l]hexane, bicyclo[2.2.1]heptane, bicyclo[2.2.2]octane, bicyclo[3.2.2]nonane, bicyclo[3.3.2]decane, norbornyl, decalinyl and adamantyl. In some embodiments, the cycloalkyl is monocyclic. Monocyclic cyclcoalkyl radicals include, for example, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl, and cyclooctyl. In some embodiments, the monocyclic cyclcoalkyl is cyclopropyl, cyclobutyl, cyclopentyl or cyclohexyl. In some embodiments, the cycloalkyl isbicyclic. Bicyclic cycloalkyl groups include fused bicyclic cycloalkyl groups, spiro bicyclic cycloalkyl groups, and bridged bicyclic cycloalkyl groups. In some embodiments, cycloalkyl groups are selected from among spiro[2.2]pentyl, bicyclo[l . l . l]pentyl, bicyclo[3.3.0]octane, bicyclo[4.3.0]nonane, bicyclo[2.1 ,l]hexane, bicyclo[2.2.1]heptane, bicyclo[2.2.2]octane, bicyclo[3.2.2]nonane, bicyclo[3.3.2]decane, norbornyl, 3,4-dihydronaphthalen-l(2H)-one and decalinyl. In some embodiments, the cycloalkyl is polycyclic. Polycyclic radicals include, for example, adamantyl, and. In some embodiments, the polycyclic cycloalkyl is adamantyl. Unless otherwise stated specifically in the specification, a cycloalkyl group may be optionally substituted.
[0205] “Fused” refers to any ring structure described herein which is fused to an existing ring structure. When the fused ring is a heterocyclyl ring or a heteroaryl ring, any carbon atom on the existing ring structure which becomes part of the fused heterocyclyl ring or the fused heteroaryl ring may be replaced with a nitrogen atom.
[0206] “Halo” or “halogen” refers to bromo, chloro, fluoro or iodo.
[0207] “Haloalkyl” refers to an alkyl radical, as defined above, that is substituted by one or more halo radicals, as defined above, e.g., trifluoromethyl, difluoromethyl, fluoromethyl, trichloromethyl, 2,2,2-trifluoroethyl, 1,2-difluoroethyl, 3-bromo-2-fluoropropyl,1.2-dibromoethyl, and the like. Unless stated otherwise specifically in the specification, a haloalkyl group may be optionally substituted.
[0208] “Haloalkoxy” refers to an alkoxy radical, as defined above, that is substituted by one or more halo radicals, as defined above, e g., trifluoromethoxy, difluoromethoxy, fluoromethoxy, trichloromethoxy, 2,2,2-trifluoroethoxy, 1,2-difluoroethoxy, 3-bromo-2-fluoropropoxy,1.2-dibromoethoxy, and the like. Unless stated otherwise specifically in the specification, a haloalkoxy group may be optionally substituted.
[0209] “Heterocycloalkyl” or “heterocyclyl” or “heterocyclic ring” refers to a stable 3- to 14-membered non-aromatic ring radical comprising 2 to 10 carbon atoms and from one to 4 heteroatoms selected from the group consisting of nitrogen, oxygen, and sulfur. Unless stated otherwise specifically in the specification, the heterocycloalkyl radical may be a monocyclic, bicyclic ring (which may include a fused bicyclic heterocycloalkyl (when fused with an aryl or a heteroaryl ring, the heterocycloalkyl is bonded through a non-aromatic ring atom), bridged heterocycloalkyl or spiro heterocycloalkyl), or polycyclic. In some embodiments, the heterocycloalkyl is monocyclic or bicyclic. In some embodiments, the heterocycloalkyl is monocyclic. In some embodiments, the heterocycloalkyl is bicyclic. The nitrogen, carbon orsulfur atoms in the heterocyclyl radical may be optionally oxidized. The nitrogen atom may be optionally quaternized. The heterocycloalkyl radical is partially or fully saturated. Examples of such heterocycloalkyl radicals include, but are not limited to, dioxolanyl, thienyl[l,3]dithianyl, decahydroisoquinolyl, imidazolinyl, imidazolidinyl, isothiazolidinyl, isoxazolidinyl, morpholinyl, octahydroindolyl, octahydroisoindolyl, 2-oxopiperazinyl, 2-oxopiperidinyl, 2-oxopyrrolidinyl, oxazolidinyl, piperidinyl, piperazinyl, 4-piperidonyl, pyrrolidinyl, pyrazolidinyl, quinuclidinyl, thiazolidinyl, tetrahydrofuryl, trithianyl, tetrahydropyranyl, thiomorpholinyl, thiamorpholinyl, 1 -oxo-thiomorph olinyl, 1, 1-dioxo-thiomorpholinyl. The term heterocycloalkyl also includes all ring forms of carbohydrates, including but not limited to monosaccharides, disaccharides and oligosaccharides. Unless otherwise noted, heterocycloalkyls have from 2 to 10 carbons in the ring. In some embodiments, heterocycloalkyls have from 2 to 8 carbons in the ring. In some embodiments, heterocycloalkyls have from 2 to 8 carbons in the ring and 1 or 2 N atoms. In some embodiments, heterocycloalkyls have from 2 to 10 carbons, 0-2 N atoms, 0-2 O atoms, and 0-1 S atoms in the ring. In some embodiments, heterocycloalkyls have from 2 to 10 carbons, 1-2N atoms, 0-1 O atoms, and 0-1 S atoms in the ring. It is understood that when referring to the number of carbon atoms in a heterocycloalkyl, the number of carbon atoms in the heterocycloalkyl is not the same as the total number of atoms (including the heteroatoms) that make up the heterocycloalkyl (i.e., skeletal atoms of the heterocycloalkyl ring). Unless stated otherwise specifically in the specification, a heterocycloalkyl group may be optionally substituted.
[0210] “Heteroaryl” refers to an aryl group that includes one or more ring heteroatoms selected from nitrogen, oxygen and sulfur. The heteroaryl is monocyclic or bicyclic. Illustrative examples of monocyclic heteroaryls include pyridinyl, imidazolyl, pyrimidinyl, pyrazolyl, triazolyl, pyrazinyl, tetrazolyl, furyl, thienyl, isoxazolyl, thiazolyl, oxazolyl, isothiazolyl, pyrrolyl, pyridazinyl, triazinyl, oxadiazolyl, thiadiazolyl, furazanyl, indolizine, indole, benzofuran, benzothiophene, indazole, benzimidazole, purine, quinolizine, quinoline, isoquinoline, cinnoline, phthalazine, quinazoline, quinoxaline, 1,8-naphthyridine, and pteridine. Illustrative examples of monocyclic heteroaryls include pyridinyl, imidazolyl, pyrimidinyl, pyrazolyl, triazolyl, pyrazinyl, tetrazolyl, furyl, thienyl, isoxazolyl, thiazolyl, oxazolyl, isothiazolyl, pyrrolyl, pyridazinyl, triazinyl, oxadiazolyl, thiadiazolyl, and furazanyl. Illustrative examples of bicyclic heteroaryls include indolizine, indole, benzofuran, benzothiophene, indazole, benzimidazole, purine, quinolizine, quinoline, isoquinoline, cinnoline, phthalazine, quinazoline, quinoxaline, 1,8-naphthyridine, and pteridine. In some embodiments, heteroaryl ispyridinyl, pyrazinyl, pyrimidinyl, thiazolyl, thienyl, thiadiazolyl or furyl. In some embodiments, a heteroaryl contains 0-4 N atoms in the ring. In some embodiments, a heteroaryl contains 1-4 N atoms in the ring. In some embodiments, a heteroaryl contains 0-4N atoms, 0-1 O atoms, and 0-1 S atoms in the ring. In some embodiments, a heteroaryl contains 1-4 N atoms, 0-1 O atoms, and 0-1 S atoms in the ring. In some embodiments, heteroaryl is a Ci-Cgheteroaryl. In some embodiments, monocyclic heteroaryl is a Ci-C5heteroaryl. In some embodiments, monocyclic heteroaryl is a 5-membered or 6-membered heteroaryl. In some embodiments, a bicyclic heteroaryl is a Ce-Cgheteroaryl.
[0211] The term “optionally substituted” or “substituted” means that the referenced group may be substituted with one or more additional group(s) individually and independently selected from alkyl, haloalkyl, cycloalkyl, aryl, heteroaryl, heterocycloalkyl, -OH, alkoxy, aryloxy, alkylthio, arylthio, alkylsulfoxide, arylsulfoxide, alkylsulfone, arylsulfone, -CN, alkyne, Ci- Cgalkylalkyne, halogen, acyl, acyloxy, -CO2H, -CO2alkyl, nitro, and amino, including mono- and di-substituted amino groups (e.g., -NH2, -NHR, -NR2), and the protected derivatives thereof. In some embodiments, optional substituents are independently selected from alkyl, alkoxy, haloalkyl, cycloalkyl, halogen, -CN, -NH2, -NH(CH3), -N(CH3)2, -OH, -CO2H, and -CO2alkyl. In some embodiments, optional substituents are independently selected from fluoro, chloro, bromo, iodo, -CH3, -CH2CH3, -CF3, -OCH3, and -OCF3. In some embodiments, substituted groups are substituted with one or two of the preceding groups. In some embodiments, an optional substituent on an aliphatic carbon atom (acyclic or cyclic) includes oxo (=0).
[0212] A "tautomer" refers to a proton shift from one atom of a molecule to another atom of the same molecule. The compounds presented herein may exist as tautomers. Tautomers are compounds that are interconvertible by migration of a hydrogen atom, accompanied by a switch of a single bond and adjacent double bond. In bonding arrangements where tautomerization is possible, a chemical equilibrium of the tautomers will exist. All tautomeric forms of the compounds disclosed herein are contemplated. The exact ratio of the tautomers depends on several factors, including temperature, solvent, and pH. Some examples of tautomeric interconversions include:Example 1. Preparation of Compound Libraries
[0213] In some embodiments, the syntheses of compounds described herein are accomplished using means described in the chemical literature, using the methods described herein, or by a combination thereof. In addition, solvents, temperatures and other reaction conditions presented herein may vary.
[0214] In other embodiments, the starting materials and reagents used for the synthesis of the compounds described herein are synthesized or are obtained from commercial sources, such as, but not limited to, Sigma-Aldrich, Fisher Scientific (Fisher Chemicals), and Acros Organics.
[0215] In further embodiments, the compounds described herein, and other related compounds having different substituents are synthesized using techniques and materials described herein as well as those that are recognized in the field, such as described, for example, in Fieser and Fieser’s Reagents for Organic Synthesis, Volumes 1-17 (John Wiley and Sons, 1991); Rodd’s Chemistry of Carbon Compounds, Volumes 1-5 and Suppiementals (Elsevier Science Publishers, 1989); Organic Reactions, Volumes 1-40 (John Wiley and Sons, 1991), Larock’s Comprehensive Organic Transformations (VCH Publishers Inc., 1989), March, Advanced Organic Chemistry 4th Ed., (Wiley 1992); Carey and Sundberg, Advanced Organic Chemistry 4th Ed., Vols. A and B (Plenum 2000, 2001), and Green and Wuts, Protective Groups in Organic Synthesis 3rd Ed., (Wiley 1999) (all of which are incorporated by reference for such disclosure). General methods for the preparation of compounds as disclosed herein may be derived from reactions and the reactions may be modified by the use of appropriate reagents and conditions, for the introduction of the various moieties found in the formulae as provided herein. As a guide the following synthetic methods may be utilized.
[0216] In the reactions described, it may be necessary to protect reactive functional groups, for example hydroxy, amino, imino, thio or carboxy groups, where these are desired in the final product, in order to avoid their unwanted participation in reactions. A detailed description oftechniques applicable to the creation of protecting groups and their removal are described in Greene and Wuts, Protective Groups in Organic Synthesis, 3rd Ed., John Wiley & Sons, New York, NY, 1999, and Kocienski, Protective Groups, Thieme Verlag, New York, NY, 1994, which are incorporated herein by reference for such disclosure).
[0217] It is understood that other analogous procedures and reagents could be used, and that these Schemes are only meant as non-limiting examples.AbbreviationsDMA: dimethylacetamideDMTMM: 4-(4,6-dimeth oxy-1, 3, 5 -triazin-2 -yl)-4-methyl-morpholinium chlorideDMSO: dimethyl sulfoxideHATU: l-[Bis(dimethylamino)methylene]-lH-l,2,3-triazolo[4,5-b]pyridinium 3-oxid hexafluorophosphateHPLC: high performance liquid chromatographyHRMS: high resolution mass spectrometry h or hr(s): hour(s) min(s): minutes m / z: mass-to-charge ratioExample 2. Synthesis of Exemplary Amide LibraryScheme 1.Example Core 1 Library 1
[0218] In an example that illustrates the ability to couple high-throughput chemical synthesis of one compound per well, a seconary amide “reactive core” was constructed and coupled it to 8639 distinct carboxylic acids. These conditions are establieshed in for example Chem. Soc. Rev., 2009, 38, 606-631.
[0219] To a 1536-well plate 100 nL of each carboxylic acid "fragment" was addedfrom a 50 mM stock solution in dimethylacetamide (DMA) [5 nmol; 2 equiv]. This was followed by addition of 100 nL ofDMTMMfrom a 60 mM stock [6 nmol; 2.4 equiv], followed by addition of 100 nL of 150 mM of Hunig's base [15 nmol, 6 equiv], and followed by 50 nL of the secondary amine core (example core 1) from a stock of 50 mM [2.5 nmol; 1 equiv]. The resulting mixtures were incubated at room temperature for approximately 12 hours, after which time DMSO was added to each well in a total volume of 3000 nL.
[0220] LCMS analysis revealed that approximately 81% of the 8639 reactions yielded measurble amide products.Example 3. Synthesis of Exemplary Benzimidazole LibraryScheme 2.
[0221] To a 1536-well plate 50 nL of 50 mM of aldehyde (2.5 nmol; 1 equiv) from a 2- methoxy ethanol (2ME) solution was added. This was followed by additional of 50 nL of 50 mM of a freshly prepared solution of reactive exemplary core 2 [2.5 nmol, 1 equiv] also from a 2ME solution, and then followed by 50 nL of 50 mM lanthanum (III) chloride [2.5 nmol, 1 equiv]. The 1536-well plate was then incubated at room temperature for approximately 12 hours, after which time reactions were quenched by the addition of 3000 nL of DMSO.
[0222] A fraction of the library was analyzed by LCMS to evaluate overall performance of the library synthesis. A 1 -minute LCMS method enabled analysis of large fractions of thelibraries. Generally about 10-15% of a library was characterized, shown in FIG. 6 for reaction Example 3.
[0223] Using this data, and other LCMS derived information, including information on the isotopic abundance patterns of expected products, and the m / z value the products were detected. A formulate a QC score was calculated for each measured reaction, e.g. shown in FIG. 7. This score encapsulates both the certainty that the desicred product was made as well as an estimate of the yeild of the desired reaction.Example 4. Performing a biological assay and analyzing barcodes
[0224] A six well segment of a 1536 well plate is set up, with each well containing a set of heterologous polypeptides. Each heterologous polypeptide in each set of heterologous polypeptides contains one or more binding sites that a test agent can bind to and modulate activy . Six different test agents are prepared in solution, and each of the different test agents are contacted to cells in the respective well. After a period of one hour has passed, a biological assay is performed to analyze resultant barcodes of the biological assay.
[0225] In the first well with a first test agent of the six different test agents added, a first count of barcodes of a reporter gene associated with a first reporter construct and a second count of barcodes of a reporter gene associated with a second reporter construct are determined from the biological assay. The first reporter construct comprises a CRE promoter and is activated by a signal from the heterologous polypeptide from a first pathway while the second reporter construct comprises aan NF AT promoter and is activated by a signal from the heterologous polypeptide from a second pathway. The first count of barcodes is 10,000 barcodes, while the second count of barcodes is 500 barcodes, indicating that when the first test agent is added to the well, it is very likely to activate the CRE promoter through the first signaling pathway as opposed to the NF AT through the second signaling pathway.
[0226] In the second well with a second test agent of the six different test agents added, a third count of barcodes of a reporter gene associated with a third reporter construct and a fourth count of barcodes of a reporter gene associated with a fourth reporter construct are determined from the biological assay. The third reporter construct comprises a CRE promoter and is activated by a signal from the heterologous polypeptide while the fourth reporter construct comprises a CRE promoter and is activated by a signal from a variant of the heterologous polypeptide. The first count of barcodes is 8,000 barcodes, while the second count of barcodes is 8,000 barcodes, indicating that when the second test agent activates CRE regardless of the variant present.
[0227] In the third well with a third test agent of the six different test agents added, a fifth count of barcodes of a reporter gene associated with a fifth reporter construct and a sixth count of barcodes of a reporter gene associated with a sixth reporter construct are determined from the biological assay. The fifth reporter construct comprises a CRE promoter and is activated by a signal from the heterologous polypeptide while the sixth reporter construct comprises a constituitively active promoter activated from toxicity in the well. The fifth count of barcodes is 300 barcodes, while the third count of barcodes is 10,000 barcodes, indicating that there is toxicity associated with the test agent.
[0228] In the fourth well with a fourth test agent of the six different test agents added, a seventh count of barcodes of a reporter gene associated with a seventh reporter construct and an eighth count of barcodes of a reporter gene associated with an eighth reporter construct are determined from the biological assay. The seventh reporter construct comprises a CRE promoter and is activated by a variant of the heterologous polypeptide while the eighth reporter construct comprieses an NF AT promoter activated by a variant of the heterologous polypeptide. The first count of barcodes is 12,000 barcodes, while the second count of barcodes is 6,000 barcodes, indicating that when the fourth test agentis added to the well, it is twice as likely to activate CRE through a pathway associated with the CRE as it is to activate NFAT through a pathway associated with the NFAT.
[0229] In the fifth well with a fifth test agent of the six different test agents added, a ninth count of barcodes of a reporter gene associated with a ninth reporter construct and a tenth count of barcodes of a reporter gene associated with an tenth reporter construct are determined from the biological assay. The ninth reporter construct comprises a CRE promoter activated by signaling through the heterologous polypeptide while the tenth reporter construct is a promoter only activated when the test agent does not bind to the heterologous polypeptide. The first count of barcodes is 50 barcodes, while the second count of barcodes is 15,000 barcodes, indicating that when the fifth test agent is added to the well, it is very unlikely to affect signaling through the heterologous polypeptide.
[0230] In the sixth well with a sixth test agent of the six different test agents added, an eleventh count of barcodes of a reporter gene associated with an eleventh reporter construct and a twelfth count of barcodes of a reporter gene associated with an twelfth reporter construct are determined from the biological assay. The eleventh reporter construct comprises a CRE promoter activated by an activation of the heterologous polypeptide while the twelfth reporter construct is a promoter only activated when the test agent does not affect singling through the heterologouspolypeptide. The first count of barcodes is 10,000 barcodes, while the second count of barcodes is 1,000 barcodes, indicating that when the sixth test agent is added to the well, it is very likely to affect signaling through the heterologous polypeptide.Example 5. Screening aminiergic GPCRs
[0231] A system was developed to measure the activities of the Muscarinic, Adrenergic, Histaminergic, Dopaminergic and Serotonergic (MAHDS) GPCRs with the three major G- protein pathways on our multiplexed platform. Gs activity was measured through the activity of a CRE reporter (Fig. 9A-9B), Gi activity was measured through the suppression of forskolin stimulated activity of a CRE reporter (Fig. 9C-9D), and Gq activity was measured through the activity of an engineered NFAT reporter (Fig 9E-9F). Reporters and GPCR expression cassettes were stably integrated into HEK293 lines engineered to have lower endogenous response, rTTA expression, and higher reporter signal. Individual cell-lines were engineered to express a specific MAHDS GPCR and G-protein reporter each linked to a unique barcode. Cell-lines were mixed together into a single library for each G-protein reporter class into cell-libraries. Here the validation of the assay and cell-libraries are described, and the activities of the MAHDS receptors are reported against their cognate broad agonists (Acetylcholine, Noradrenaline, Histamine, Dopamine, Serotonin) and 4 antipsychotics (Aripiprazole, Clozapine, Haloperidol, Lurasidone).
[0232] To validate the multiplexed reporter libraries, dose responses of the expected interactions between The MAHDS receptors and their endogenous cognate agonists were measured (Fig. 10). The concentrations of the drug inducing half of the maximum effect (EC50) of each agonist to their cognate receptors in the assay were typically in the 1 - 1000 nM range (Fig. 10 ) which generally agrees with other GPCR activity assays as reported by CHEMBL. The overall performance of the libraries is summarized in Fig. 9G.
[0233] Currently the assay is able to robustly detect 33 out of 35 of the reported MAHDS GPCR primary couplings (with the exception of HTR1E and HTR2A), 6 out of 12 reported secondary couplings, and 17 "unexpected" couplings (Fig, 9G and Fig. 10).
[0234] For the apparent missing primary coupling to HTR1E, a weak signal was detected and a dose response was calculated in the Gi pathway (Fig. 10), but the signal was not statistically significant. However, a strong Gs coupling was observed for this receptor. For HTR2A, no Gi signaling was observed, but its reported secondary coupling to Gq by NFAT signaling was observed.
[0235] Of the six reported secondary couplings that were not observed, four (ADRA2C, ADRB1, ADRB2, ADRB3) were from receptors that are primary coupled to Gi or Gs, and secondary coupled to the inverse pathway, Gs or Gi. Since these pathways are convoluted - one is the inverse activity of the other - it was not always possible to optimize our cell-lines and conditions to reliably detect both. For HRH2, though could actually detect Gs activity, it was removed from the library due to it's propensity to cause high paracrine signaling. For HTR5A, no appreciable Gq signaling was detected on the NFAT reporter.Receptor cross reactivity to non-cognate agonists
[0236] Because the assay collects all receptor responses in multiplex, the response of each MAHDS receptor was measured not just to it's cognate agonist, but to all of the endogenous agonists simultaneously. With the exception of the dopaminergic and adrenergic receptors, the MAHDS responded only to their cognate endogenous agonists (Fig. 11A). In the case of the adrenergic and dopaminergic receptors, the highest affinity interactions were to their respective cognate agonists.Receptor selectivity in the cross reactive adrenergic and dopaminergic receptors
[0237] To verify that the assay was capable of measuring the receptor selectivity of different compounds, the activity profiles of ADRB1 and DRD1 were measured against noradrenaline and dopamine. These agonists are known to activate both the dopaminergic class and the adrenergic class in vivo and in vitro, each with higher affinity for their cognate receptors. The expected receptor selectivity of these agonists were observed, with noradrenaline exhibiting ~60 fold higher selectivity for ADRB1 than DRD1, and conversely, with dopamine exhibiting ~80 fold higher selectivity forDRDl than ADRB1 (Fig. 11B). As expected, higher affinity of DRD1 was observed for dopamine vs. noradrenaline, and higher affinity of ADRB1 for noradrenaline vs. dopamine. These results are generalizable to all of the receptors where additional reactivity to their non-cognate agonists was observed; the highest affinity agonists are cognate for each receptor (Fig. 11 A).Detecting unexpected GPCR couplings
[0238] Interestingly, many of our GPCRs exhibit activity for reporter pathways not reported in the IUPHAR database. For example, within the muscarinic (CHRM) family receptors, only CHRM2 is reported to have Gs coupling (Fig. 9G). However, in the assay Gs reporter activity was observed for all five CHRM receptors, suggesting that these receptors couple to Gs in our cell-lines (Fig. 9G). With the exception of the dopaminergic class of receptors, which only respond to their expected primary couplings, there are instances of these "unexpected" couplingsin the other receptor classes as well (Fig. 9G). There are many possible reasons for these discrepancies. Likely, most of these "unexpected" receptor couplings are not properly annotated in the IUPHAR database or are understudied / underappreciated. Other recent broad studies of GPCR coupling with different approaches have yielded similar results to what was observed . The general approach here is to be agnostic to these considerations, only taking note when a primary G-protein coupling is expected but not observed in our assay (described above).Measuring the activity of 4 model Antipsychotics
[0239] To assess the ability of the multiplexed reporter assay to report the profiles of complex polypharmacological drugs, the activities of 4 well known antipsychotics for agonism, inverse agonism, and antagonism activity were measured against the MAHDS receptors [Aripiprazole (Abilify), clozapine (Clozaril), haloperidol (Haldol), and lurasidone (Latuda)]. Clozapine and haloperidol were chosen as model Atypical and Typical antipsychotics respectively. Aripiprazole and lurasidone are newer drugs, classified as atypical, that have recently gained popularity for treatment of schizophrenia, bi-polar, and depression.
[0240] Agonism and inverse agonism activity were collected by simply applying drug and measuring any apparent change in reporter output. To measure antagonism, a cocktail of the endogenous agonists was developed that simultaneously agonized all of the receptors in the library. The ability of each antipsychotic to antagonize those interactions was observed, detected as an inverse of the original interaction (Figure 4a). Conditions are described in detail in the supplement. All plots of all antipsychotic interactions against all receptors, in agonism or antagonism can be found in the supplement.
[0241] To validate the assay for antagonism, theoretical binding affinities from the antagonist interactions were calculated using the Cheng-Prusoff equation, which takes into account the EC50 of the agonist, the agonist concentration, andthe IC50 of the antagonist. Reported binding affinities are usually measured with radioligand binding assays and that theoretical calculated affinities with the Cheng-Prusoff method should be treated as approximations. Despite this caveat, the vast majority of the calculated binding affinities are within an order of magnitude of the expected values , validating both the approach and assay (Fig. 12B).Majority of expected antipsychotic interactions recapitulatedThe data was compared to what is reported for these antipsychotics (Fig. 12C). An interaction is "expected" when it is reported to interact with a human receptor at sub 1000 nM affinity (it is a general rule of thumb that affinity values higher than 200 nM are not physiologically relevant). An interaction was a "hit" when the expected mode of action was reproduced: agonism,antagonism, or inverse agonism in our assay. An interaction a "miss" when an activity was reported and none was observed. Many of the reported drug interactions were listed as "not determined" underthe mode of action despite listing a high affinity. This suggests antagonism as the mode of action. The assay identified 53 / 62 (85%) of the "expected" hits. Of the 9 "misses", most were antagonist interactions reported with low potency or were cases where the assay measured agonist activity but not antagonist.
[0242] A large number (75) of the interactions detected were classified as "divergent hits"; instances where our assay reported an interaction that differed from what was reported. These are mostly cases where 1) no interaction was reported in lUPHAR / Wikipedia but was identified in our assay or 2) the interaction was reported as simply an antagonist but inverse or partial agonist activity was observed as well as the antagonist activity (Fig. 12A-12C). Inverse agonists and partial agonists will appear to have antagonist activity in most conventional antagonist assays.
[0243] While some of these can be explained by difference in assay type (ie. different celllines), it's likely that most of these "divergent hits" are either under reported / misannotated in the lUPHARS / Wikipedia databases. For example, lurasidone had few interactions reported for the MAHDS. This atypical antipsychotic was not reported to have activity against the D2 class of receptors (DRD2, DRD3, DRD4) in the databases. A closer inspection revealed activity here was reported, but only against rat homologs.Receptor selectivity of antipsychotics: DRD2 vs. HTR2A
[0244] It has been proposed that a key differentiation of atypical antipsychotic's ability to treat psychoses while having lower risk of extrapyramidal side-effects (EPS) vs. typical antipsychotics lies in their high HTR2A antagonism potency relative to that of DRD2. When the relative IC50s of clozapine, lurasidone, and haloperidol for HTR2A and DRD2 antagonism was compared, clozapine possessed the highest potency ratio of HTR A to DRD2, while haloperidol had the lowest (Fig. 13A-13B) (Aripiprazole was not included in this analyses, as it is often considered to be a third class of antipsychotic, possessing partial agonist activity at both HTR2A and DRD2). This is consistent with the expectation, given that clozapine is considered the model atypical antipsychotic with the highest clinical value for treatment of positive symptoms, while haloperidol is a model typical antipsychotic with a higher risk of EPS. These data demonstrate that the multiplexed assay is sensitive to trends in receptor selectivity of these compounds that can potentially explain differences in clinical outcomes.Example 6. Screening compounds for interactions with MC4R
[0245] 34 compounds were screened for interactions with MC4Rusingthe systems described herein. The tested compounds included both small molecules and peptides. Constructs with reporters were made for for 4 receptors (MC4R1, MC4R2, MC4R3, and MC4R4) with signaling read outs for both the Gq pathway and the Gs pathway. Each well contained cells comprising each reporter construct. One compound was screened for each well. Results are depicted in Fig. 14.
[0246] Each compound was profiled against multiple GPCR receptors and multiple signaling pathways. Fig. 15A depicts the selectivity of each compound between the different MC4R receptors. Fig. 15B depicts the Bias seen in the MC4R Gq / Gs axis. Different compounds bias signaling towards either the Gs or Gq axis.Example 7: High-resolution DMS for identification of Retinitis Pigmentosa mutations
[0247] Described in this example is an in vitro diagnostic assay to determine a class for a Retinitis Pigmentosa disease variant and potential for response to a specific rhodopsin corrector molecule. A gene expressing a variant RHO with a single missense mutation is fused to gene expressing a transcription factor (TF) and integrated into a cell. Each cell has a single variant RHO gene and a unique barcode downstream of a response element (RE). A mutated RHO- transcription factor fusion that is properly folded and trafficked to the plasma membrane will encounter a plasma membrane specific protease. The protease will cleave the mutated RHO- transcription factor and the transcription factor will be released from the membrane. The transcription factor binds to the response element and will allow for transcription of the variant specific barcode, which can be read by next generation sequencing.
[0248] Barcodes were isolated, and quantified by next generation sequencing.
[0249] As shown in FIG. 16, a heatmap was generated by deep mutational screening of the trafficking score of each missense mutation against wildtype RHO. Each individual amino acid of RHO was mutated to one of twenty amino acids. All 348 amino acids of RHO were tested with twenty amino acid changes, including a stop codon Each mutant RHO was encoded with a unique barcode for identification. The assay was performed as described above and the barcode reads for each unique mutation was quantified and normalized against wildtype. The normalized trafficking was evaluated for each variant of RHO. The library was then treated with a rhodopsin corrector and the assay was performed to identify which mutations respond to a specific rhodopsin corrector.
[0250] It is understood that the examples and embodiments described herein are for illustrative purposes only and that various modifications or changes in light thereof will be suggested to persons skilled in the art and are to be included within the spirit and purview of this application and scope of the appended claims. All publications, patents, and patent applications cited herein are hereby incorporated by reference in their entirety for all purposes.
Claims
CLAIMSWHAT IS CLAIMED IS:
1. A system comprising a plurality of engineered cell lines in a partition, wherein the plurality of engineered cell lines comprises a first engineered cell line and a second engineered cell line, wherein the first engineered cell line comprises a first reporter construct and the second engineered cell line comprises a second reporter construct, wherein, the first reporter construct and the second reporter construct are selected from: a) a constitutive promoter operatively coupled to a reporter gene; b) a first inducible promoter operatively coupled to a reporter gene; or c) a second inducible promoter operatively coupled to a reporter gene; wherein the first reporter construct is different from the second reporter construct; and wherein the first reporter construct and the second reporter construct are independently readable.
2. The system of claim 1 , wherein the plurality of engineered cell lines further comprises a third engineered cell line, wherein the third engineered cell line comprises a third reporter construct that is different from the first reporter construct and the second reporter construct, wherein the third reporter construct is selected from: a) a constitutive promoter operatively coupled to a reporter gene; b) a first inducible promoter operatively coupled to a reporter gene; or c) a second inducible promoter operatively coupled to a reporter gene.
3. The system of claim 1 or 2, wherein one or more of the first engineered cell line, second engineered cell line, third engineered cell line, or combinations thereof further comprises a first heterologous polypeptide.
4. The system of claim 3, wherein one or more of the first engineered cell line, second engineered cell line, third engineered cell line, or combinations thereof further comprise a second heterologous polypeptide, wherein the second heterologous polypeptide comprises at least one amino acid alteration relative to the first heterologous polypeptide.
5. The system of claim 4, wherein the second heterologous polypeptide comprises less than 10, less than 5, less than 3, or less than 2 amino acid alterations relative to the first heterologous polypeptide.
6. The system of claim 4, wherein the second heterologous polypeptide comprises more than 10, more than 20, more than 50, more than 100, more than 500, or more than 1000 amino acid alterations relative to the first heterologous polypeptide.
7. The system of any one of claims 3 to 6, wherein the first heterologous polypeptide, the second heterologous polypeptide, or both are coupled to a transcription factor.
8. The system of claim 7, wherein the transcription factor comprises one or more of aGal4, PPR1, Lac9, zinc finger, or LexA DNA binding domain.
9. The system of claim 7 or 8, wherein the transcription factor comprises one or more of aVP64, p65, RoTev, or Rta DNA activating domain.
10. The system of any one of claims 1 to 9, wherein the first reporter construct and the second reporter construct are independently readable.
11. The system of any one of claims 2 to 10, wherein the third reporter construct is independently readable from the first reporter construct and / or the second reporter construct reporter.
12. The system of any one of claims 3 to 11, wherein the first inducible promoter, the second inducible promoter, or both are configured to be activated by the first heterologous polypeptide, a signal from the first heterologous polypeptide, a transcription factor coupled to the first heterologous polypeptide, or any combination thereof.
13. The system of any one of claims 1 to 12, wherein the first engineered cell line, second engineered cell line, or third engineered cell line comprises an additional reporter construct selected from: a) a constitutive promoter operatively coupled to a reporter gene; b) a first inducible promoter operatively coupled to a reporter gene; or c) a second inducible promoter operatively coupled to a reporter gene, wherein the additional reporter construct is different from the first reporter construct and the second reporter construct.
14. The system of any one of claims 1 to 13, wherein one or more of the first reporter construct, the second reporter construct, or the third reporter construct are integrated into the genome of the cell line.
15. The system of any one of claims 1 to 14, wherein any one or more of the first engineered cell line, the second engineered cell line, or the third engineered cell line is a eukaryotic cell line.
16. The system of claim 15, wherein the eukaryotic cell line is a mammalian cell line.
17. The system of claim 16, wherein the mammalian cell line is a human cell line.
18. The system of any one of claims 1 to 17, wherein the reporter gene encodes a fluorescent protein or a luciferase protein.
19. The system of any one of claims 1 to 17, wherein the reporter gene encodes a barcode RNA sequence.
20. The system of any one of claims 1 to 19, wherein the reporter gene encodes a fluorescent protein and a barcode RNA sequence or a luciferase protein and a barcode RNA sequence.21 . The system of any one of claims 1 to 20, wherein the first heterologous polypeptide or the second heterologous polypeptide is a cell-surface protein.
22. The system of claim 21, wherein the cell-surface protein is a G-protein coupled receptor, a receptor tyrosine kinase, an ion channel, a cytokine receptor, a chemokine receptor, a growth factor receptor, or a cellular adhesion molecule.
23. The system of any one of claims 21 to 22, wherein the cell-surface protein is expressed by any one or more of the plurality of engineered cell lines.
24. The system of any one of claims 1 to 23, wherein the first heterologous polypeptide or the second heterologous polypeptide is an intracellular protein.
25. The system of claim 24, wherein the intracellular protein is an enzyme, ER transporter, nuclear transporter, intracellular signaling protein, a chaperone, or a transcription factor.
26. The system of any one of claims 1 to 25, wherein the plurality of engineered cell lines further comprises a cell that comprises a barcode sequence but does not express the barcode sequence.
27. The system of claim 26, wherein the plurality of engineered cell lines comprises mammalian cells.
28. The system of claim 27, wherein the mammalian cells are human cells.
29. The system of any one of claims 1 to 28, wherein the partition is a well of an rc-well plate.
30. The system of claim 29, wherein the n-well plate is a 96-well plate.31 . The system of claim 29, wherein the n-well plate is a 384-well plate.
32. The system of claim 29, wherein the n-well plate is a 1536-well plate.
33. The system of any one of claims 1 to 32, wherein the first reporter construct, the second reporter construct, or the third reporter construct comprises a constitutive promoter operatively coupled to a reporter gene.
34. The system of any one of claims 1 to 33, wherein the constitutive promoter is selected from a SV40 promoter, a CMV promoter, an Efl A promoter, a PGK1 promoter, an Ubc promoter, abeta actin promoter, a CAG promoter, an Ac5 promoter, a polyhedrin promoter, a TEF1 promoter, a GDS promoter, a CaMV355 promoter, an Ubi protomer, or any combination thereof.
35. The system of any one of claims 1 to 34, wherein the first inducible promoter comprises an NFAT promoter, a CRE promoter, a p53 promoter, an ISRE promoter, a Gal4-UAS promoter, a Lex A promoter.
36. The system of any one of claims 1 to 35, wherein the second inducible promoter is selected from an NFAT promoter, a CRE promoter, a p53 promoter, an ISRE promoter, a Gal4-UAS promoter, a Lex A promoter.
37. The system of any one of claims 1 to 36, wherein the first inducible promoter and second inducible promoter are different promoters that mediate signaling through the same intracellular protein or cell-surface protein.
38. A method of screening for a compound that regulates a biological activity of any one or more of the plurality of engineered cell lines of any one of the preceding claims, the method comprising contacting the plurality of engineered cell lines with a test agent and measuring an activity of the reporter constructs of the first engineered cell line, the second engineered cell line, third engineered cell line, or any combination thereof.
39. The method of claim 38, wherein a plurality of n-test agents is contacted to the plurality of engineered cell lines that has been divided into at least n-partitions.
40. The method of claims 38 or 39, wherein the plurality of engineered cell lines comprises at least 100, at 1,000, or at least 10,000 different heterologous polypeptides.41 . The method of claims 39 to 40, wherein the biological assay is performed in a 96-well, 384- well or a 1536-well plate.
42. The method of any one of claims 38 to 41, wherein the first reporter construct and the second reporter construct are present in different engineered cells in the same well of a 96-well, 384- well, or a 1536-well plate.
43. The method of any one of claims 38 to 42, wherein the plurality of n-test agents is prepared as a plurality of individual reaction mixtures by reacting a core fragment A comprising a reactive functionality x, wherein x is an amine, aldehyde, boronate, imide, isothiocyanate, carboxylic acid, halide, hydroxy amidine, or thiourea; with a plurality of naive test fragments (y-Ti, y-T2, . . . y-Tn), each test fragment comprising a reactive functionality y, wherein y is a amine, boronate, imide, isothiocyanate, halide, hydroxyamide, thiourea, carboxylic acid oraldehyde and one of a plurality of naive test moieties (T T2, . . . Tn) under reaction conditions sufficientto form a plurality of n test agents (A-Ti, A-T2, . . . A-Tn), wherein each of A-TbA-T2, . . . A-Tnis prepared in a well of a test plate.
44. The method of claim 43, wherein x is an amine.
45. The method of claim 44, wherein the amine is a primary or secondary amine.
46. The method of claim 43, wherein x is an aldehyde or carboxylic acid.
47. The method of claim 43, wherein y is an amine.
48. The method of claim 47, wherein the amide is a primary or secondary amine.
49. The method of claim 43, wherein y is a carboxylic acid or aldehyde.
50. The method of any one of claims 43 to 49, wherein the reaction conditions comprise a base.51 . The method of any one of claims 43 to 49, wherein the reaction conditions comprise a Lewis acid or Bronstead acid.
52. The method of any one of claims 43 to 49, wherein the reaction conditions comprise an amide coupling reagent.
53. The method of any one of claims 43 to 49, wherein the reaction conditions comprise a palladium reagent.
54. The method of any one of claims 43 to 53, wherein the reaction conditions comprise room temperature.
55. The method of any one of claims 43 to 53, wherein the reaction conditions comprise a reaction temperature between about room temperature and about 80 °C.
56. The method of any one of claims 43 to 55, wherein the reactingis performed from about 1 to about 24 hours.
57. The method of claim 56, wherein the reacting is performed from about 6 to about 18 hours.
58. The method of any one of claims 43 to 57, wherein the reacting is reversible or irreversible.
59. The method of any one of claims 43 to 58, comprising an amide coupling, reductive amination, or oxidative addition.
60. The method of any one of claims 43 to 58, comprising a Buchwald or Suzuki coupling.61 . The method of any one of claims 43 to 60 wherein the core moiety A has a mass of from about 150 Da to about 800 Da.
62. The method of any one of claims 43 to 61, wherein each test moiety (Tl, T2, . . . Tn) has a mass of from about 80 Da to about 500 Da.
63. The method of any one of claims 43 to 62, wherein each test agent (A-Tl, A-T2, . . . A-Tn) has a mass of less than about 1500 Da.
64. The method of claim 63, wherein each test agent (A-Tl, A-T2, . . A-Tn) has a mass of from about 350 Da to about 800 Da.
65. The method of any one of claims 43 to 64, wherein each reaction is performed on a nanoscale.
66. The method of claim 65, wherein the nanoscale is performed in a volume of 50nLto 500 nL.
67. The method of any one of claims 43 to 66, wherein the test plate is a 96-well, 384-well, or 1536-well test plate.
68. The method of any one of claims 43 to 67, wherein each naive test moieties (Tl, T2, . . . Tn) is different.
69. The method of any one of claims 43 to 68, wherein n is from 2 to 100,000.
70. The method of claim 69, wherein n is from 2 to 25,000.
71. The method of claim 69, wherein n is from 2 to 2,500.
72. The method of claim 69, wherein n is from 2 to 2,000.