Multi-reporter system for chemical screening
By using various engineered cell lines and reporter gene constructs coupled with different promoters, the problem of difficulty in screening multiple signal transduction pathways and cell functions in parallel in existing technologies has been solved, achieving efficient screening and identification of the specificity and efficacy of compounds and improving the efficiency of drug discovery.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- OCTANT INC
- Filing Date
- 2024-07-18
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies struggle to efficiently screen multiple signal transduction pathways or cellular functions in parallel, and are unable to identify off-target effects or compounds that agonize or antagonize specific target variants, resulting in low drug screening efficiency and insufficient understanding of biological systems.
Using a system containing multiple engineered cell lines and reporter gene constructs coupled with different promoters, the target activity and off-target effects can be determined simultaneously or substantially simultaneously by measuring the bioactivity of multiple barcode reporters, and compounds can be screened and identified in parallel using multi-well plates.
It improves the efficiency of drug screening and the understanding of biological systems, enabling the simultaneous screening of thousands of compounds, reducing nonspecificity or off-target effects, and enhancing the specificity and effectiveness of compound screening.
Smart Images

Figure CN122029294A_ABST
Abstract
Description
Cross-referencing
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 528,030, filed July 20, 2023, which is incorporated herein by reference in its entirety. Background Technology
[0002] Understanding signal transduction pathways and cellular function, and identifying assays that regulate given signal transduction pathways or cellular function, are important goals of pharmaceutical science and drug discovery. Summary of the Invention
[0003] The methods and systems described herein provide cell-based screening assays. These systems and methods include several improvements over previous screening methods. The systems and methods described herein allow for multiplexing to screen multiple signaling pathways or cellular functions in parallel. These systems and methods allow for the identification of off-target effects or the identification of compounds that activate or antagonize specific target variants, signaling pathways, and / or biological functions. The systems and methods described herein include multiple indexes to allow for improved statistical significance. The methods and systems also allow for the identification or avoidance of nonspecific or off-target effects by allowing the screening of molecules that activate or antagonize the same target but transduce signals through different downstream effectors. Overall, the methods and systems described herein increase the efficiency of drug screening and the understanding of biological systems. The methods and systems described herein allow for the simultaneous or substantially simultaneous determination of the activity of a certain assay agent against a target (e.g., a heterologous peptide), a variant of the target, different downstream promoters that can be activated by the target, and the simultaneous off-target effects of the assay agent in a single reaction vessel (e.g., the wells of a multi-well plate). The methods described herein are capable of simultaneously determining thousands of compounds in this way. This article describes methods and systems for screening and identifying assay agents that regulate target activity, either positively (as agonists) or negatively (as antagonists).
[0004] This document describes a system comprising multiple engineered cell lines in a partition, wherein the multiple engineered cell lines include a first engineered cell line and a second engineered cell line, wherein the first engineered cell line includes a first reporter gene construct and the second engineered cell line includes a second reporter gene construct, wherein the first and second reporter gene constructs are selected from: (a) a constitutive promoter operatively coupled to a reporter gene; (b) a first inducible promoter operatively coupled to a reporter gene; or (c) a second inducible promoter operatively coupled to a reporter gene; wherein the first reporter gene construct is different from the second reporter gene construct; and wherein the first and second reporter gene constructs are independently readable. In some embodiments, the multiple engineered cell lines further include a third engineered cell line, wherein the third engineered cell line includes a third reporter gene construct different from the first and second reporter gene constructs, wherein the third reporter gene construct is selected from: (a) a constitutive promoter operatively coupled to a reporter gene; (b) a first inducible promoter operatively coupled to a reporter gene; or (c) a second inducible promoter operatively coupled to a reporter gene. In some embodiments, one or more of the first engineered cell line, the second engineered cell line, the third engineered cell line, or combinations thereof further comprise a first heterologous polypeptide. In some embodiments, one or more of the first engineered cell line, the second engineered cell line, the third engineered cell line, or combinations thereof further comprise a second heterologous polypeptide, wherein the second heterologous polypeptide comprises at least one amino acid change relative to the first heterologous polypeptide. In some embodiments, the second heterologous polypeptide comprises less than 10, less than 5, less than 3, or less than 2 amino acid changes relative to the first heterologous polypeptide. In some embodiments, the second heterologous polypeptide comprises more than 10, more than 20, more than 50, more than 100, more than 500, or more than 1000 amino acid changes relative to the first heterologous polypeptide. In some embodiments, the first heterologous polypeptide, the second heterologous polypeptide, or both are coupled to a transcription factor. In some embodiments, the transcription factor comprises one or more of aGal4, PPR1, Lac9, a zinc finger, or a LexA DNA-binding domain. In some embodiments, the transcription factor comprises one or more of aVP64, p65, RoTev, or Rta DNA activation domains. In some embodiments, the first and second reporter gene constructs are independently readable. In some embodiments, the third reporter gene construct is independently readable as a reporter gene alongside the first and / or second reporter gene constructs. In some embodiments, the first, second, or both inducible promoters are configured to be activated by a first heteropeptide, a signal from the first heteropeptide, a transcription factor coupled to the first heteropeptide, or any combination thereof.In some embodiments, the first, second, or third engineered cell line comprises an additional reporter gene construct selected from: (a) a constitutive promoter operatively coupled to the reporter gene; (b) a first inducible promoter operatively coupled to the reporter gene; or (c) a second inducible promoter operatively coupled to the reporter gene, wherein the additional reporter gene construct differs from the first and second reporter gene constructs. In some embodiments, one or more of the first, second, or third reporter gene constructs are integrated into the genome of the cell line. In some embodiments, any one or more of the first, second, or third engineered cell line is a eukaryotic cell line. In some embodiments, the eukaryotic cell line is a mammalian cell line. In some embodiments, the mammalian cell line is a human cell line. In some embodiments, the reporter gene encodes a fluorescent protein or a luciferase protein. In some embodiments, the reporter gene encodes a barcoded RNA sequence. In some embodiments, the reporter gene encodes a fluorescent protein and a barcoded RNA sequence or a luciferase protein and a barcoded RNA sequence. In some embodiments, the first or second heteropeptide is a cell surface protein. In some embodiments, the cell surface protein is a G protein-coupled receptor, a receptor tyrosine kinase, an ion channel, a cytokine receptor, a chemokine receptor, a growth factor receptor, or a cell adhesion molecule. In some embodiments, the cell surface protein is expressed by any one or more of a plurality of engineered cell lines. In some embodiments, the first or second heteropeptide is an intracellular protein. In some embodiments, the intracellular protein is an enzyme, an ER transporter, a nuclear transporter, an intracellular signaling protein, a chaperone molecule, or a transcription factor. In some embodiments, the plurality of engineered cell lines further comprises cells containing a barcode sequence but not expressing the barcode sequence. In some embodiments, the plurality of engineered cell lines comprise mammalian cells. In some embodiments, the mammalian cells are human cells. In some embodiments, the partitions are wells of an n-well plate. In some embodiments, the n-well plate is a 96-well plate. In some embodiments, the n-well plate is a 384-well plate. In some embodiments, the n-well plate is a 1536-well plate. In some embodiments, the first, second, or third reporter gene construct includes a constitutive promoter operatively coupled to the reporter gene. In some embodiments, the constitutive promoter is selected from the SV40 promoter, CMV promoter, Ef1A promoter, PGK1 promoter, Ubc promoter, β-actin promoter, CAG promoter, Ac5 promoter, polyhedrin promoter, TEF1 promoter, GDS promoter, CaMV355 promoter, Ubi promoter, or any combination thereof.In some embodiments, the first inducible promoter comprises the NFAT promoter, CRE promoter, p53 promoter, ISRE promoter, Gal4-UAS promoter, and Lex A promoter. In some embodiments, the second inducible promoter is selected from the NFAT promoter, CRE promoter, p53 promoter, ISRE promoter, Gal4-UAS promoter, and Lex A promoter. In some embodiments, the first and second inducible promoters are different promoters that mediate signal transduction through the same intracellular or cell surface protein. In some embodiments described herein, a method for screening compounds that regulate the biological activity of one or more of a plurality of engineered cell lines according to any one of the preceding claims is described, the method comprising contacting a plurality of engineered cell lines with a test agent and measuring the activity of a reporter gene construct of a first engineered cell line, a second engineered cell line, a third engineered cell line, or any combination thereof. In some embodiments, a plurality of n test agents are contacted with a plurality of engineered cell lines that have been divided into at least n partitions. In some embodiments, multiple engineered cell lines contain at least 100, 1,000, or at least 10,000 different heterologous peptides. In some embodiments, bioassays are performed in 96-well, 384-well, or 1536-well plates. In some embodiments, a first reporter gene construct and a second reporter gene construct are present in different engineered cells within the same wells of a 96-well, 384-well, or 1536-well plate. In some embodiments, multiple n assay reagents are prepared by reacting a core fragment A containing a reactive functional group x with multiple naïve assay fragments (y-T1, y-T2, ... yT). n ) sufficient to form multiple n test reagents (A-T1, A-T2, ...AT) n The reaction is carried out under the following reaction conditions to prepare multiple separate reaction mixtures, where x is an amine, aldehyde, borate ester, imide, isothiocyanate, carboxylic acid, halide, hydroxyamidine, or thiourea; each test fragment contains a reactive functional group y and multiple initial test fractions (T1, T2, ... T). n One of ), where y is an amine, borate ester, imide, isothiocyanate, halide, hydroxyamide, thiourea, carboxylic acid, or aldehyde, wherein A-T1, A-T2, ... AT nEach of the components is prepared in a well of a test plate. In some embodiments, x is an amine. In some embodiments, the amine is a primary or secondary amine. In some embodiments, x is an aldehyde or carboxylic acid. In some embodiments, y is an amine. In some embodiments, the amide is a primary or secondary amine. In some embodiments, y is a carboxylic acid or aldehyde. In some embodiments, the reaction conditions include a base. In some embodiments, the reaction conditions include a Lewis acid or a Bronstead acid. In some embodiments, the reaction conditions include an amide coupling reagent. In some embodiments, the reaction conditions include a palladium reagent. In some embodiments, the reaction conditions include room temperature. In some embodiments, the reaction conditions include a reaction temperature between about room temperature and about 80°C. In some embodiments, the reaction proceeds for about 1 to about 24 hours. In some embodiments, the reaction proceeds for about 6 to about 18 hours. In some embodiments, the reaction is reversible or irreversible. In some embodiments, the method includes amide coupling, reductive amination, or oxidative addition. In some embodiments, the method includes Buchwald or Suzuki coupling. In some embodiments, the mass of core portion A is about 150 Da to about 800 Da. In some embodiments, the mass of each test portion (T1, T2, ... Tn) is from about 80 Da to about 500 Da. In some embodiments, the mass of each reagent (A-T1, A-T2, ... A-Tn) is less than about 1500 Da. In some embodiments, the mass of each reagent (A-T1, A-T2, ... A-Tn) is from about 350 Da to about 800 Da. In some embodiments, each reaction is carried out at the nanoscale. In some embodiments, the nanoscale reaction is carried out in volumes of 50 nL to 500 nL. In some embodiments, the test plate is a 96-well, 384-well, or 1536-well test plate. In some embodiments, each initial test portion (T1, T2, ... Tn) is different. In some embodiments, n is from 2 to 100,000. In some embodiments, n is from 2 to 25,000. In some embodiments, n is from 2 to 2,500. In some embodiments, n is from 2 to 2,000. Incorporation
[0005] All publications, patents and patent applications mentioned in this specification are incorporated herein by reference to the extent that each individual publication, patent or patent application is specifically and individually indicated to be incorporated by reference. Attached Figure Description
[0006] Various aspects of this disclosure are set forth in the appended claims. A better understanding of the features and advantages of this disclosure will be obtained by referring to the following detailed description and the accompanying drawings, which illustrate illustrative embodiments utilizing the principles of this disclosure.
[0007] Figure 1 Exemplary cells showing interaction with the test reagent.
[0008] Figure 2 Exemplary cells containing at least two reporter gene constructs that interact with the test reagent are shown.
[0009] Figure 3 The illustration shows an exemplary cell containing at least two reporter gene constructs that interact with the test reagent, wherein one of the at least two reporter gene constructs is activated when the heterologous peptide is not activated.
[0010] Figure 4 The example wells show one or more cells containing different heterologous peptides, response elements, and / or reporter gene constructs that provide unique biological information about the response of the heterologous peptides, response elements, and / or reporter gene constructs to the test reagent.
[0011] Figures 5A-5D This shows an example configuration of a heterologous peptide and / or reporter gene construct used to inquire about the biological function of one or more assay reagents. Figure 5A The diagram shows the integration of multiple response elements. Figure 5B The illustration shows the integration of multiple heterologous peptide variants with response elements. Figure 5C The diagram shows the integration of multiple heterologous peptides with multiple response elements and multiple indexes. Figure 5D The diagram integrates heterologous peptides or heterologous peptide variants with multiple indexes.
[0012] Figure 6 This chart shows the LCMS peak area distribution for 96 reactions in an exemplary library that form benzimidazole. The pink / red bars represent the peak areas of the expected benzimidazole products, while the green bars represent the peak areas of the aldehyde fragments and the blue bars represent the peak areas of the o-phenylenediamine cores.
[0013] Figure 7 The QC scores of 96 reactions for benzimidazole formation, measured by LCMS, are shown. More than 70 of the 96 measured reactions had a QC score ≥0.5, indicating that approximately 80% of the reactions successfully formed a large quantity of the desired benzimidazole.
[0014] Figure 8 The illustrations show some examples of assays for detecting different biological activities, which can be applied to or performed in conjunction with the methods and systems described herein.
[0015] Figures 9A-9G Describe an overview of the multiplexed MAHDS screening platform. Figure 9A , 9CThe overall design of the engineered GPCR signaling circuit is depicted in Figure 9E. An inducible promoter drives the expression of the receptor of interest. G protein activation is stimulated by receptor activation via an agonist. Figure 9B , 9D Representative dose-response curves for agonists were plotted using 9F: the activity (plotted as a log2 fold change relative to baseline) was fitted to a 4-parameter log-log function using the drc package from Ritz et al. (available on GitHub). EC50 values are provided plus or minus their standard errors. Figure 9A The activated Gs protein stimulates adenylate cyclase, leading to cAMP production and activation of the CRE promoter by the CREB transcription factor. Figure 9B Depicting Figure 9A The result of the test was an increase in the barcode count. Figure 9C The activated Gi protein was shown to inhibit the activity of (forsocrine-stimulated) adenylate cyclase, resulting in lower cAMP production and lower activation of the CRE promoter. Figure 9D The detection method is used to reduce the number of barcodes. Figure 9C The result. Figure 9E The activation of the Gq protein stimulates PLCB, leading to an increase in intracellular calcium concentration and activation of the engineered NFAT promoter. Figure 9F The depiction detection increases the barcode count. Figure 9E The result. Figure 9G The ability of assays to detect the activity of each receptor for each G protein reporter gene is depicted by shading. Major reporting conjugates for each receptor are indicated by a single asterisk [*], while minor reporting conjugates are indicated by double asterisks [**]. Reporting conjugates are derived from IUPHARS.
[0016] Figures 10A-10E The activity of each receptor class in each library relative to its homologous endogenous agonist is depicted (described as a log2 fold change relative to baseline). Each drug-receptor interaction is fitted to a 4-parameter log-log function using the drc package (available on GitHub). EC50 values are provided plus or minus their standard errors. The canonical couplings of each GPCR are listed in parentheses. Responses are categorized into their respective G protein reporter gene classes (blue: Gs / CRE, brown: Gi / InvCRE, orange: Gq / NFAT). Figure 10A Describe the acetylcholine activity against muscarinic receptors. Figure 10B Describe norepinephrine activity against adrenergic receptors. Figure 10C Describe histamine activity against histaminergic receptors. Figure 10D Describe dopamine activity against dopaminergic receptors. Figure 10E Describe serotonin activity against serotonergic receptors.
[0017] Figures 11A-11B Describe the effects of MAHDS receptors on endogenous neurotransmitters. Figure 11A Depicts the activity of MAHDS receptors against endogenous agonists. Shaded areas represent pEC50 of the dose response. White areas indicate no response was detected. Figure 11B The active dose-response of ADRB1 and DRD1 in the CRE-Gs reporter genes to norepinephrine and dopamine in the Gs reporter genes is depicted. EC50 is reported and indicated by dashed lines.
[0018] Figures 12A-12C The activity of four antipsychotic drugs was measured and verified. Figure 12A Plot the dose-response diagrams of representative antagonist activities for Gs, Gi, and Gq. Figure 12B The binding affinity (pKi) of available reports is depicted and compared with the affinity calculated from the data using the Cheng-Prusoff equation. Figure 12C A summary describing the interactions of each antipsychotic drug with each receptor. “Hit” indicates when the reported interaction was reproduced in our assays, “Difference hit” indicates when we observed an unreported interaction, and “Miss hit” indicates when the interaction was reported but we did not observe it.
[0019] Figures 13A-13B Describe the antipsychotic activity and receptor selectivity of DRD2 and HTR2A. Figure 13A Plot the dose-response curves of each antipsychotic drug against DRD2 (Gi) and HTR2A (Gq). Figure 13B The pIC50 of DRD2 and HTR2A is depicted, and the ratio between HTR2A and DRD2 is listed.
[0020] Figure 14 The results of screening 34 compounds for Mc1R, MC3R, MC4R, and MC4R are presented.
[0021] Figure 15A The engineering of activity and potency along the MC4R / 1R axis with selectivity is described. Figure 15B Describe the bias between Gs and Gq signal conduction.
[0022] Figure 16 A heatmap showing the amount of mutant RHO transported to the extracellular domain relative to wild-type RHO. Detailed Implementation
[0023] Determining which assays affect certain biological activities without affecting others, and therefore which assays will treat certain diseases while others will not, can be a difficult and time-consuming process. The methods described herein include approaches for evaluating reporter gene construct regulation that report specific biological activities by analyzing multiple barcodes expressed from reporter gene constructs when a bioassay is performed. In doing so, an assay can be tested against multiple reporter gene constructs simultaneously, and the resulting barcodes indicate which reporter gene construct the assay activates or does not activate. For example, an assay can be added to wells to react with multiple engineered cell lines containing multiple different reporter gene constructs, and when a bioassay is performed, barcode expression shows that the assay typically activates not only the relevant promoter but also, for example, other reporter genes that may indicate off-target effects, such as in different signaling pathways.
[0024] Therefore, by performing bioassays in well plates or chips with numerous partitions (e.g., wells), the interactions between the assay agent and biological activity can be studied more effectively and in more detail by displaying an array of complete interactions. This includes whether the assay agent binds to a specific heteropeptide, whether the assay agent binds to a variant of the heteropeptide, whether certain promoters are activated through certain pathways, whether the assay agent promotes toxicity in the well, or whether the assay agent does not affect any biological activity. By doing so, it is possible to identify which assay agents contribute to more effective treatment of conditions or symptoms. Furthermore, this allows for the parallel multiplexing of many assay agents. This can be combined with fragment-based chemistry to further iterate and develop assay agents into candidate molecules (e.g., small molecules, peptides, or biologics) with increased potency, greater specificity to activate or antagonize specific biological functions, or reduced side effects.
[0025] This document describes a system comprising multiple engineered cell lines in a partition, wherein the multiple engineered cell lines include a first engineered cell line and a second engineered cell line, wherein the first engineered cell line includes a first reporter gene construct and the second engineered cell line includes a second reporter gene construct, wherein the first reporter gene construct and the second reporter gene construct are selected from: (a) a constitutive promoter operatively coupled to a reporter gene; (b) a first inducible promoter operatively coupled to a reporter gene; or (c) a second inducible promoter operatively coupled to a reporter gene; wherein the first reporter gene construct is different from the second reporter gene construct; and wherein the first reporter gene construct and the second reporter gene construct are independently readable.
[0026] In some aspects, this document describes a system comprising multiple engineered cell lines in a partition, wherein the multiple engineered cell lines include a first engineered cell line and a second engineered cell line, wherein the first engineered cell line includes a first reporter gene construct and the second engineered cell line includes a second reporter gene construct. In some embodiments, the system includes a first reporter gene construct and a second reporter gene construct, wherein the first reporter gene construct and the second reporter gene construct are selected from: a constitutive promoter operatively coupled to a reporter gene; a first inducible promoter operatively coupled to a reporter gene; or a second inducible promoter operatively coupled to a reporter gene; and the first reporter gene construct is different from the second reporter gene construct. In some embodiments, the first reporter gene construct and the second reporter gene construct are independently readable. In some embodiments, the multiple engineered cell lines further include a third engineered cell line, wherein the third engineered cell line contains a third reporter gene construct that is different from the first and second reporter gene constructs, wherein the third reporter gene construct is selected from: a constitutive promoter operatively coupled to a reporter gene; a first inducible promoter operatively coupled to a reporter gene; or a second inducible promoter operatively coupled to a reporter gene. In some embodiments, the first, second, or third engineered cell line includes an additional reporter gene construct.
[0027] In some aspects, this document describes a method for screening compounds (e.g., ligands described herein) that regulate the biological activities of any one or more of the various engineered cell lines described herein. In some embodiments, the method includes contacting the various engineered cell lines with a test agent and measuring the activity of a reporter gene construct of a first engineered cell line, a second engineered cell line, a third engineered cell line, or any combination thereof. In some embodiments, the various engineered cell lines comprise a first heterologous polypeptide. In some embodiments, the first heterologous polypeptide may be complexed with or contacted with the test agent. In some embodiments, the first heterologous polypeptide is operatively coupled to a transcription factor, wherein contact between the test agent and the first heterologous polypeptide results in transcription or expression of the reporter gene (e.g., as described herein). Figure 1 and Figure 5A (As shown). In some embodiments, multiple engineered cell lines contain reporter genes (e.g., such as...) for screening interactions between heterologous peptides and assay reagents. Figure 2 , Figure 3 or Figure 5B -D is shown.
[0028] engineered cell lines This document describes systems, methods, and assays utilizing multiple engineered cell lines comprising one or more of a reporter gene construct, a heterologous peptide, or a variant of a heterologous peptide. The engineered cell lines may be contained within one or more partitions to facilitate screening of multiple compounds (e.g., one compound per partition). Such compounds may be synthesized de novo, as described herein, or sourced from chemical suppliers known in the art. Multiple engineered cell lines may comprise 2, 3, 4, 5, 6, 7, 8, 9, 10, 50, 75, 100, 1,000, 2,000, 3,000, 4,000, 5,000, or 10,000 or more different engineered cell lines. Within a partition, each engineered cell line is present in a specific quantity (e.g., 1, 2, 3, 5, 10, 20, 30, 40, 50, 75, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 cells or more). Within a partition, each engineered cell line is present in a specific quantity or fewer (e.g., 2, 3, 5, 10, 20, 30, 40, 50, 75, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000, 1,100 cells or fewer). As described herein, an "engineered cell line" refers to a cell containing one or more exogenous nucleic acids encoding or containing a heterologous polypeptide, a variant of a heterologous polypeptide, or a reporter gene construct. In some embodiments, the exogenous nucleic acid is integrated into a genomic location of the cell (e.g., a chromosome). Such engineered cell lines can be constructed using various methods to introduce nucleic acids into the cells (e.g., CaCl2, cationic lipid transfection reagents, electroporation, viral transduction, etc.).
[0029] In the embodiments described herein, each well or partition may contain multiple engineered cell lines. The engineered cell lines may contain one or more nucleic acids. One or more nucleic acids may encode one or more heterologous polypeptides described herein and / or contain a reporter gene construct described herein. One or more nucleic acids may contain a promoter operatively coupled to a reporter gene (e.g., a reporter gene construct). The reporter gene may also contain a unique and identifiable barcode (e.g., an index sequence) and, as described below, may indicate one or more aspects of regulation of the reporter gene construct by one or more assay agents. Such barcodes may also be included on reporter gene constructs having a constitutively active promoter that drives the reporter gene to provide information on cell viability or toxicity. In some embodiments, the barcode may be unaffected by any type of transcriptional regulation to allow for standardization of total cell number in an assay. In some embodiments, the barcodes are uniquely paired to indicate regulation by a specific promoter or a specific heterologous polypeptide or variant thereof. Such regulation includes activation and inhibition of the organism via test ligands.
[0030] The reporter genes in engineered cell lines are independently readable, meaning that reporter genes are not dependent on each other's expression and can provide information about different biological activities in the same assay. Such activities include, but are not limited to, two or more different intracellular signaling pathways, one or more signaling pathways and cell viability, and one or more signaling pathways and protein transport or abundance.
[0031] As further described below, when the test reagent is introduced into a well or partition, the test reagent can interact with multiple cells within each well. The interaction between the test reagent and the cells can include the activation of one or more biological events, which can further lead to the expression or repression of a reporter gene containing a barcode, thereby informing the user of the system described herein about the biological effects of one or more test reagents.
[0032] In some embodiments, the cells described herein are obtained from the cell lines described herein. In some embodiments, the cells are eukaryotic cells. In some embodiments, the cells are mammalian cells. In some embodiments, the cells are human cells. In some embodiments, the cells are obtained from a subject. In some embodiments, the cells are obtained from a subject suffering from a disease or condition. In some embodiments, the cells obtained from the subject can be screened against a test kit for treating the disease or condition. In some embodiments, the cells are associated with a disease or condition. In some embodiments, the test kit can be used to diagnose, treat, or prevent a disease or condition. In some embodiments, the cells can be further engineered to express heterologous peptides, reporter gene constructs, or combinations thereof.
[0033] In some embodiments, the methods and systems described herein include the use of n different engineered cell lines. In some embodiments, the n different engineered cell lines are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 or more different cell lines. In some embodiments, these n different engineered cell lines may be contained in a single partition.
[0034] Engineered cell lines can be nay cell lines that can be used to screen molecules or assays. In some respects, the cells are cancer cells, tumor cells, or other immortalized cells. In further respects, the cells represent disease model cells. In some respects, the cells can be A549, B cells, B16, BHK-21, C2C12, C6, CaCo-2, CAP / , CAP-T, CHO, CHO2, CHO-DG44, CHO-K1, COS-1, Cos-7, CV-1, dendritic cells, DLD-1, embryonic stem (ES) cells or derivatives, H1299, HEK293, 293T, 293FT, Hep G2, hematopoietic stem cells, HOS, Huh-7, induced pluripotent stem (iPS) cells or derivatives, Jurkat, K562, L5278Y, LNCaP, MCF7, MDA-MB-231, MDCK, mesenchymal cells, Min-6, monocytes, Neuro2a, NIH 3T3, NIH3T3L1, K562, NK cells, NSO, Panc-1, PC12, PC-3, peripheral blood cells, plasma cells, primary fibroblasts, RBL, Renca, RLE, SF21, SF9, SH-SY5Y, SK-MES-1, SK-N-SH, SL3, SW403, stimulation-triggered pluripotency gain (STAP) cells or derivatives of SW403, T cells, THP-1, tumor cells, U2O5, U937, peripheral blood lymphocytes, expanded T cells, hematopoietic stem cells, or Vero cells. In some implementations, the cells are HEK293T cells.
[0035] Heteropeptides A heterologous polypeptide is any polypeptide foreign to an engineered cell line and may be encoded by a foreign nucleic acid transfected into or added to the engineered cell line. The heterologous polypeptide may be a recombinant form of a polypeptide already present in the cell or a recombinant form of a polypeptide not expressed in the engineered cell line. A heterologous polypeptide may contain one or more amino acid variants (e.g., substitution, deletion, addition, or truncation). Such foreign nucleic acids can be stably integrated into the genome of the engineered cell line, creating, in some cases, clonal cell lines expressing the heterologous polypeptide. The nucleic acid may be single-stranded or double-stranded. In some embodiments, the nucleic acid is DNA. Such heterologous polypeptides may be included on plasmids, viral vectors, or linearized DNA to facilitate nucleic acid transfer to the engineered cell line.
[0036] In some embodiments, the heterologous peptide can be expressed in the same cells as a reporter gene construct containing a barcode, wherein the barcode is a unique, identifiable sequence on the reporter gene that can be used to identify the reporter gene and associated promoters and / or the heterologous peptide (as described below). Figure 1-3(As described). In some embodiments, the heteropeptide can be activated or inhibited when the assay reagent binds to it. Furthermore, in some embodiments, activation of the heteropeptide can send downstream signals leading to activation of a promoter or response element. As described below, one or more reporter gene constructs can be further activated by activation of the heteropeptide or by downstream signals induced by the heteropeptide. The promoter can be further operatively coupled to a barcode-containing reporter gene. Variants of the heteropeptide can similarly be activated and / or expressed in various engineered cell lines.
[0037] As described herein, heteropeptides can comprise peptides or proteins involved in signal transduction pathways within cells. Heteropeptides can belong to a class of proteins or a family of proteins. Such heteropeptides include G-coupled protein receptors (GPCRs), receptor tyrosine kinases (RTKs), intracellular signaling molecules, intracellular transport molecules, chaperone molecules / heat shock proteins, transcription activators, enhancers or repressors, secretory proteins (e.g., growth factors, cytokines, chemokines, etc.), cell adhesion molecules, cytoskeleton components, DNA replication proteins, histone or nuclear hormone receptors.
[0038] The systems, methods, and assays described herein can be used to determine the structural and functional basis of potential therapeutic interventions using variants of heterologous peptides. As used herein, a “variant” of a peptide includes a peptide having an amino acid sequence different from that of a first heterologous peptide. Some variants of peptides that can be used with this disclosure are those that have the ability to participate in a reduction of normal cellular biological activities, such as cell signaling, transport, enzymatic, or metabolic activities. Generally, the amino acid sequence of peptide variants that can be used with this disclosure may differ from the amino acid sequence of a wild-type peptide by one or more amino acids. In some embodiments, the amino acid sequence of a peptide variant differs by one amino acid. In some embodiments, the amino acid sequence of a peptide variant differs by at least one amino acid, at least two amino acids, at least three amino acids, at least four amino acids, at least five amino acids, at least six amino acids, at least seven amino acids, at least eight amino acids, at least nine amino acids, or at least ten amino acids. In some embodiments, the amino acid sequence of a peptide variant differs by at most one amino acid, at most two amino acids, at most three amino acids, at most four amino acids, at most five amino acids, at most six amino acids, at most seven amino acids, at most eight amino acids, at most nine amino acids, or at most ten amino acids. In some embodiments, the methods and assays use multiple heteropeptides comprising a single amino acid variant, the single amino acid variant covering at least about 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or 100% of the amino acid residues of the heteropeptide of interest. In some embodiments, the multiple heteropeptides comprising a single amino acid variant comprise 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1,000, or more heteropeptide variants.
[0039] In some embodiments, the heterologous peptide is expressed by the engineered cell lines described herein. For example, the heterologous peptide is expressed by a first engineered cell line, a second engineered cell line, a third engineered cell line, or any other cell line. In some embodiments, the heterologous peptide described herein can be a first heterologous peptide or a second heterologous peptide, wherein the first heterologous peptide and the second heterologous peptide may differ from each other by at least one amino acid. In such cases, the first and second heterologous peptides can be used to screen for interactions with a test reagent, wherein the interaction is specific to at least one amino acid change. In some embodiments, the second heterologous peptide comprises at least one amino acid change relative to the first heterologous peptide. In some embodiments, the second heterologous peptide is another peptide among another peptides in a class of common peptides, such as GPCRs, receptor tyrosine kinases, nuclear hormone receptors, etc. The system described herein may comprise n different cell lines, each with a different heterologous peptide. In some implementations, the n different cell lines can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 or more different cell lines. In this way, multiple variants of a single heteropeptide or a heteropeptide of a biologically or structurally related class can be analyzed simultaneously.
[0040] In some embodiments, each heteropeptide or combination of heteropeptides may be expressed in an engineered cell line or combination of engineered cell lines. In some embodiments, the expression of heteropeptides or combinations of heteropeptides in each engineered cell line or combination of engineered cell lines may be partitioned by a system described herein, wherein each partition of the system comprises a unique combination of cell lines expressing one or more of the heteropeptides.
[0041] In some embodiments, the heterologous polypeptide is operatively coupled to a gene encoding a promoter that controls its expression. In some embodiments, the promoter is constitutive. In some embodiments, the promoter is conditional or inducible (e.g., tetracycline-inducible).
[0042] In some implementations, the heterologous polypeptide can be fused to a transcription factor via a protease-sensitive linker that can be cleaved upon activation or proximity to a protease specific to the linker.
[0043] In some embodiments, the heterologous peptide is a cell surface protein. For example, the heterologous peptide is a cell surface protein of a G protein-coupled receptor, a receptor tyrosine kinase, an ion channel, a cytokine receptor, a chemokine receptor, a growth factor receptor, a cell adhesion molecule, a fragment thereof, or a combination thereof. In some embodiments, the heterologous peptide is an intracellular protein. For example, the heterologous peptide is an enzyme, an ER transporter, a nuclear transporter, an intracellular signaling protein, a transport protein, a chaperone molecule, a transcription factor, a fragment thereof, or a combination thereof.
[0044] Reporting gene constructs In some embodiments, the engineered cell line within the partition contains a reporter gene construct. In some embodiments, the reporter gene construct is encoded by a foreign nucleic acid. In some embodiments, the reporter gene construct contains a promoter or response element operatively coupled to a reporter gene. The promoter may include any genomic element that can be bound by a transcription factor, enhancer, etc., and that genomic element is capable of initiating transcription of a reporter gene downstream of that transcription factor, enhancer, etc. The promoters described herein may be inducible (e.g., the promoter may be activated by a specific transcription factor activated under specific signaling conditions). The promoters described herein may be constitutively active. Examples of constitutively active promoters may include those from SV40 promoter, CMV promoter, Ef1A promoter, PGK1 promoter, Ubc promoter, β-actin promoter, CAG promoter, Ac5 promoter, polyhedrin promoter, TEF1 promoter, GDS promoter, CaMV355 promoter, Ubi promoter, or any combination thereof. In some embodiments, the promoter may be an inducible promoter. Examples of inducible promoters may include the TRE promoter, the GAL1 promoter, the GAL10 promoter, or any combination thereof.
[0045] In some embodiments, the reporter gene includes a barcode, wherein the barcode is a uniquely identifiable sequence. In some embodiments, the barcode can be used to identify the reporter gene or associated promoter when expressed. In some embodiments, the barcode can be used to identify a heterologous receptor gene or a variant thereof. The reporter gene may include a gene encoding a fluorescent protein, luciferase protein, β-galactosidase, β-glucuronidase, chloramphenicol acetyltransferase, or secretory placental alkaline phosphatase, or a combination thereof. In some embodiments, the reporter gene includes both a barcode and a gene encoding a reporter gene protein, such as a fluorescent protein, luciferase protein, β-galactosidase, β-glucuronidase, chloramphenicol acetyltransferase, or secretory placental alkaline phosphatase, or a combination thereof.
[0046] Reporter gene constructs can be activated based on heterologous protein activation, signals generated by heterologous peptide activation, or other cellular biological activities. Activation of the reporter gene construct can include activating the promoter of the reporter gene construct by recruiting essential transcription factors and transcription activators. Additionally, activation of the reporter gene construct can lead to the expression of a reporter gene operatively coupled to the promoter, wherein the expression of the reporter gene allows for the identification of a barcode on the reporter gene in one or more subsequent assays. The barcode can indicate one or more aspects associated with the activation of the reporter gene construct, such as which promoter was activated, through which pathway the promoter was activated, whether a heterologous peptide or a variant of the heterologous peptide was activated, whether there is toxicity due to one or more assays in the well, or whether the assays in the well do not regulate any biological activity within the well. In some embodiments, the reporter gene containing the barcode may not be operatively coupled to the promoter, allowing the barcode to be used to assess cell proliferation and viability.
[0047] The expressed barcode can indicate which reporter gene construct is activated, and thus further indicate which heterologous peptide or variant of heterologous peptide is activated, through which pathway certain reporter gene constructs are activated, whether toxicity is present in the partition, or whether the assay does not bind to or produce a biological effect with any substance within the partition. Therefore, when an assay is introduced into a partition, the resulting expressed barcode indicates which reporter gene construct is activated and how it is activated, which can allow for the determination of which assay can be used to modulate specific biological activities, including diseases or conditions of interest.
[0048] In some implementations, the gene construct is reported to be integrated into the cell's genome. In other implementations, the gene construct is reported not to be integrated into the cell's genome.
[0049] In some embodiments, the systems and methods described herein utilize at least one reporter gene construct. In some embodiments, each reporter gene construct may be regulated by a different promoter. For example, a first reporter gene construct may be regulated by a first promoter, and a second reporter gene construct may be regulated by a second promoter, and so on. In some embodiments, each partition of the system may include a unique or combined promoter that regulates the expression of a unique or combined combination of reporter gene constructs. In some embodiments, the promoter may be a constitutively active promoter, an inducible promoter, or a combination thereof. The systems described herein may comprise n different cell lines, each with a different reporter gene construct. In some embodiments, the n different cell lines are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 or more different cell lines. This method allows for the simultaneous analysis of many different promoters responding to the test reagent.
[0050] As described above, multiple reporter gene constructs may include one or more promoters. In some embodiments, multiple reporter gene constructs may include multiple promoters. In some embodiments, multiple promoters may include multiple types of promoters. In some embodiments, multiple promoters may include a promoter and an alternative promoter. In some embodiments, the promoter may be a first type of promoter, and the alternative promoter may be a second type of promoter. In some embodiments, the alternative promoter may be a constitutive promoter. In some embodiments, multiple promoters may include a promoter, an alternative promoter, and a constitutive promoter. Although multiple promoters having a promoter, an alternative promoter, a constitutive promoter, or a combination thereof have been described, this multiple promoter is exemplary and other multiple promoters may be used. For example, although multiple promoters having two or three promoters have been described, a much larger number of promoters, such as four, five, six, seven, eight, nine, or ten or more promoters, may be used.
[0051] In some implementations, one or both of the promoter and the second, third, and fourth promoters are activated. In some implementations, the promoter and the second, third, and fourth promoters are activated via different signal transduction pathways. For example, the promoter may be activated by a first-type assay agent, while the second, third, and fourth promoters may not be activated by the first-type assay agent. Furthermore, the promoter and the second, third, and fourth promoters may be activated to varying degrees. For example, the promoter may be activated by more than 1.5, 2, 3, 4, 5, 10, or more times compared to the second, third, and fourth promoters; or the second, third, and fourth promoters may be activated by more than 1.5, 2, 3, 4, 5, 10, or more times compared to the first promoter. Therefore, for one or more reporter gene constructs contained within a partitioned engineered cell line, the promoter and / or the second, third, and fourth promoters may be operatively coupled to one or more reporter genes, which may be expressed upon activation of the promoter and / or alternative promoters. Each reporter gene may further include a barcode uniquely paired with the promoter. In some implementations, the barcode can be used to uniquely identify a reporter gene, and by extension, the promoter is operatively coupled to the reporter gene when it is expressed. Thus, in a cell comprising a first reporter gene construct, which includes, for example, a CREB promoter operatively coupled to a first reporter gene containing a first barcode, wherein the cell also includes a second reporter gene construct, which includes, for example, an NFAT promoter operatively coupled to a second reporter gene containing a second barcode (e.g., an alternative promoter to the first promoter), identification of the first barcode indicates activation of the first promoter, while identification of the second barcode indicates activation of the second promoter. Therefore, when the test kit...
[0052] In some implementations, one or two promoters and constitutive promoters present in different engineered cell lines within the partition are activated. In some implementations, the constitutive promoter detects toxicity (through reduced activation) or off-target reduction in signal transduction. Constitutive promoters may include CMV, RSV, SV40, etc.
[0053] Bioassay The engineered cell lines described herein can be used in one or more bioassays to obtain data related to the regulation of biological activity when exposed to a test reagent. As described herein, bioassays involve contacting cells in culture wells with one or more test reagents.
[0054] As used in this article, “bioactivity” refers to any cellular activity necessary for normal cell growth and development. These activities include DNA synthesis, protein synthesis, folding and transport, intracellular signal transduction, gene transcription, cellular enzymatic activity, synthesis and secretion of growth factors and other intercellular signaling molecules, cell division, apoptosis, etc. Figure 8 Some examples of assays used to detect different biological activities are described.
[0055] When cells are exposed to or cultured with one or more assay agents, one or more cellular biological activities can be regulated by the assay agents. Activation of the target can further lead to the expression of one or more reporter genes, and the ability to identify one or more barcodes of these reporter genes, for example, through sequencing. Barcode identification allows determination of which assay agent regulates a specific reporter gene and provides insights into which assay agents regulate which biological activities.
[0056] In some embodiments, one or more test reagents may be measured. In some embodiments, the amount of test reagent that may be used for measurement in the system described herein is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 100, 1,000, 10,000 or more.
[0057] In some embodiments, one or more test reagents can be measured. In some embodiments, the amount of test reagents that can be used for measurement is from about 1 test reagent to about 1,536 test reagents. In some embodiments, the amount of test reagents that can be used for measurement is from about 1 test reagent to about 48 test reagents, from about 1 test reagent to about 96 test reagents, from about 1 test reagent to about 192 test reagents, from about 1 test reagent to about 288 test reagents, from about 1 test reagent to about 384 test reagents, from about 1 test reagent to about 768 test reagents, from about 1 test reagent to about 1,152 test reagents, from about 1 test reagent to about 1,536 test reagents, or from about 48 test reagents to about 96 test reagents. Approximately 48 test reagents to approximately 192 test reagents, approximately 48 test reagents to approximately 288 test reagents, approximately 48 test reagents to approximately 384 test reagents, approximately 48 test reagents to approximately 768 test reagents, approximately 48 test reagents to approximately 1,152 test reagents, approximately 48 test reagents to approximately 1,536 test reagents, approximately 96 test reagents to approximately 192 test reagents, approximately 96 test reagents to approximately 288 test reagents, approximately 96 test reagents to approximately 384 test reagents, approximately 96 test reagents to approximately 7 68 test reagents, approximately 96 test reagents to approximately 1,152 test reagents, approximately 96 test reagents to approximately 1,536 test reagents, approximately 192 test reagents to approximately 288 test reagents, approximately 192 test reagents to approximately 384 test reagents, approximately 192 test reagents to approximately 768 test reagents, approximately 192 test reagents to approximately 1,152 test reagents, approximately 192 test reagents to approximately 1,536 test reagents, approximately 288 test reagents to approximately 384 test reagents, approximately 288 test reagents to approximately The quantities of test reagents may be approximately 768, approximately 288 to approximately 1,152, approximately 288 to approximately 1,536, approximately 384 to approximately 768, approximately 384 to approximately 1,152, approximately 384 to approximately 1,536, approximately 768 to approximately 1,152, approximately 768 to approximately 1,536, or approximately 1,152 to approximately 1,536. In some embodiments, the amount of test reagents that can be used for determination may be approximately 1, approximately 48, approximately 96, approximately 192, approximately 288, approximately 384, approximately 768, approximately 1,152, or approximately 1,536. In some implementations, the amount of test reagent that can be used for determination is about 1 test reagent, about 48 test reagents, about 96 test reagents, about 192 test reagents, about 288 test reagents, about 384 test reagents, about 768 test reagents, or about 1,152 test reagents.In some implementations, the amount of test reagents that can be used for determination is up to about 48 test reagents, about 96 test reagents, about 192 test reagents, about 288 test reagents, about 384 test reagents, about 768 test reagents, about 1,152 test reagents, or about 1,536 test reagents.
[0058] One or more assay reagents can contact multiple cells containing one or more heterologous peptides or variants of heterologous peptides. In some embodiments, the amount of heterologous peptides that can be measured is from about 1 heterologous peptide to about 10,000 heterologous peptides. In some embodiments, the amount of heterologous peptides that can be used for measurement is from about 1 heterologous peptide to about 100 heterologous peptides, from about 1 heterologous peptide to about 500 heterologous peptides, from about 1 heterologous peptide to about 1,000 heterologous peptides, from about 1 heterologous peptide to about 1,500 heterologous peptides, from about 1 heterologous peptide to about 2,000 heterologous peptides, from about 1 heterologous peptide to about 4,000 heterologous peptides, from about 1 heterologous peptide to about 6,000 heterologous peptides, from about 1 heterologous peptide to about 8,000 heterologous peptides, from about 1 heterologous peptide to about 10,000 heterologous peptides, from about 100 heterologous peptides to about 500 heterologous peptides, from about 100 heterologous peptides to... Approximately 1,000 heterologous peptides; approximately 100 heterologous peptides to approximately 1,500 heterologous peptides; approximately 100 heterologous peptides to approximately 2,000 heterologous peptides; approximately 100 heterologous peptides to approximately 4,000 heterologous peptides; approximately 100 heterologous peptides to approximately 6,000 heterologous peptides; approximately 100 heterologous peptides to approximately 8,000 heterologous peptides; approximately 100 heterologous peptides to approximately 10,000 heterologous peptides; approximately 500 heterologous peptides to approximately 1,000 heterologous peptides; approximately 500 heterologous peptides to approximately 1,500 heterologous peptides; approximately 500 heterologous peptides to approximately 2,000 heterologous peptides; approximately 500 heterologous peptides to approximately 4,000 heterologous peptides. Polypeptides, approximately 500 to 6,000 heteropeptides, approximately 500 to 8,000 heteropeptides, approximately 500 to 10,000 heteropeptides, approximately 1,000 to 1,500 heteropeptides, approximately 1,000 to 2,000 heteropeptides, approximately 1,000 to 4,000 heteropeptides, approximately 1,000 to 6,000 heteropeptides, approximately 1,000 to 8,000 heteropeptides, approximately 1,000 to 10,000 heteropeptides, approximately 1,500 to 2,000 heteropeptides. 000 heteropeptides, approximately 1,500 heteropeptides to approximately 4,000 heteropeptides, approximately 1,500 heteropeptides to approximately 6,000 heteropeptides, approximately 1,500 heteropeptides to approximately 8,000 heteropeptides, approximately 1,500 heteropeptides to approximately 10,000 heteropeptides, approximately 2,000 heteropeptides to approximately 4,000 heteropeptides, approximately 2,000 heteropeptides to approximately 6,000 heteropeptides, approximately 2,000 heteropeptides to approximately 8,000 heteropeptides, approximately 2,000 heteropeptides to approximately 10,000 heteropeptides, approximately 4,000 heteropeptides to approximately 6,000 heteropeptides, approximately 4,The amount of heteropeptides that can be measured is approximately 1 heteropeptide, approximately 100 heteropeptides, approximately 4,000 heteropeptides, approximately 10,000 heteropeptides, approximately 6,000 heteropeptides, approximately 6,000 heteropeptides, approximately 10,000 heteropeptides, or approximately 8,000 heteropeptides, or approximately 10,000 heteropeptides. In some embodiments, the amount of heteropeptides that can be measured is approximately 1 heteropeptide, approximately 100 heteropeptides, approximately 500 heteropeptides, approximately 1,000 heteropeptides, approximately 1,500 heteropeptides, approximately 2,000 heteropeptides, approximately 4,000 heteropeptides, approximately 6,000 heteropeptides, approximately 8,000 heteropeptides, or approximately 10,000 heteropeptides. In some embodiments, the amount of heterologous peptides that can be used for determination is at least about 1 heterologous peptide, about 100 heterologous peptides, about 500 heterologous peptides, about 1,000 heterologous peptides, about 1,500 heterologous peptides, about 2,000 heterologous peptides, about 4,000 heterologous peptides, about 6,000 heterologous peptides, or about 8,000 heterologous peptides. In some embodiments, the amount of heterologous peptides that can be used for determination is at most about 100 heterologous peptides, about 500 heterologous peptides, about 1,000 heterologous peptides, about 1,500 heterologous peptides, about 2,000 heterologous peptides, about 4,000 heterologous peptides, about 6,000 heterologous peptides, about 8,000 heterologous peptides, or about 10,000 heterologous peptides.
[0059] In some embodiments, assessing the regulation of a specific biological activity may include using bioassays to test which assays affect a given biological activity. In some embodiments, the assay may be based on ligand binding to a target at one or more binding sites. In some embodiments, the bioassay may include information from two or more reporter gene constructs from one or more engineered cell lines. A reporter gene construct may include one or more promoters. In some embodiments, one or more promoters may include promoters activated by a target or a target variant (directly or indirectly via a signaling intermediate or their absence). The bioassays described herein involve contacting a test reagent with one or more wells containing an engineered cell line. The bioassay may further include incubating the test reagent with the cells for a sufficient time to allow expression to be operatively coupled to a barcode activated by a target promoter or alternative promoter. The test reagent may be added to each well. In some embodiments, different test reagents are added to each respective well. The bioassay can be performed in any particular multi-container format, such as a 96-well, 384-well, or 1536-well plate, a microfluidic chip containing micropores, or in an emulsion.
[0060] Figure 1A cell 100 interacting with test reagent 102 is depicted. Cell 100 contains a heterologous polypeptide 104 encoded by nucleic acid 106 and a reporter gene construct 108 also encoded by nucleic acid. The target may be expressed on the cell surface (e.g., in the case of a G protein-coupled receptor or receptor tyrosine kinase) or intracellularly (e.g., in the case of a nuclear hormone receptor). Reporter gene construct 108 includes a promoter 110 operatively coupled to reporter gene 112. Reporter gene 112 may further include a uniquely identifiable barcode 114. In some embodiments, target 104 may be a heterologous polypeptide.
[0061] In the embodiments depicted herein, a test reagent 102 is added to a well containing cells 100, wherein the test reagent 102 then interacts with the cells (e.g., via a heterologous polypeptide 104). In some embodiments, the test reagent 102 interacts with the cells by binding to the heterologous polypeptide 104 or by influencing signaling pathways downstream of the heterologous polypeptide. The interaction of the test reagent 102 affects signaling via the heterologous polypeptide 104. In response to the test reagent 102, the promoter 110 of the reporter gene construct 108 is activated. In some embodiments, activation of the promoter 110 may result in the expression of a reporter gene 112 operatively coupled to the promoter 110 (e.g., as expression 118). The reporter gene 112, as well as expression 118 via an extended barcode 114, can be expressed and analyzed or quantified by sequencing (e.g., using next-generation sequencing assays) to identify the barcode 114. Barcode 114 may be associated with one or more promoters (e.g., promoter 110) and may indicate certain aspects of the assay (e.g., assay 102 activates reporter gene construct 108 via a first pathway).
[0062] Therefore, wells containing multiple cells or partitions within wells can be measured to determine the expression of barcodes that uniquely identify heterologous peptides and / or response elements regulated by the assay reagent. In some cases, more than one assay reagent can be measured, and more than one target can be measured (e.g., a target and its variants or multiple completely different targets). Furthermore, as described below, in some embodiments, there can be more reporter gene constructs that can be activated intracellularly. In some embodiments, only one reporter gene construct may be present in a single cell, but different reporter gene constructs may be present in other cells.
[0063] In embodiments containing more than one promoter, whether within the same cell or different cells within a well or partition, the bioassay can indicate which assays modulate biological activity based on one or more reporter gene constructs. For example, one or more promoters can be activated when a ligand binds to or activates a target. When one or more promoters are activated, they can express one or more reporter genes (e.g., barcoded mRNA) operatively coupled to the promoter. Each corresponding reporter gene of one or more reporter gene constructs can indicate that the assay has bound to a heterologous peptide or a variant of the heterologous peptide, that the assay promotes toxicity, or that the assay has not produced any significant biological effect.
[0064] Figure 2 An exemplary system described herein is depicted, in which cell 200 interacts with test reagent 202. In this depicted example, test reagent 202 interacts with cell 200 by binding to a heterologous polypeptide 204 encoded by nucleic acid 206, which results in activation of signal transduction via heterologous polypeptide 204. In this depicted example, activation of heterologous polypeptide 204 leads to activation of promoter 212 of reporter gene construct 210 via pathway 208a. Furthermore, in this depicted example, activation of the target can lead to activation of promoter 222 of reporter gene construct 220 (containing barcode 226) via pathway 208b. Activation of promoter 212 results in the expression 218 of reporter gene 214 operatively coupled to promoter 212. Using sequencing technology, barcode 216 can be identified as expression 218, while barcode 226 is not activated. This indicates one or more aspects of the interaction between assay 202 and the cell and / or heterologous peptide 204 or its signaling pathway; for example, in this figure, the assay may act in a pathway-specific manner against pathway 208a. Other indicators, such as whether only pathway 208b is activated or both pathways 208a and 208b are activated, may be available. This analysis can be scaled up to inquire about the biological activity of multiple different assays against multiple different biological activities or signaling pathways. The system can be deployed in simultaneous or near-simultaneous assays, allowing for high-throughput detection of biological differences between assays. For example, barcode 216 may indicate that promoter 212 is activated via pathway 208a. Barcode 226 may indicate that promoter 222 is activated via pathway 208b. In some embodiments, multiple barcodes identified by multiple expressions (including expression 218) may indicate that promoter 212 is activated in multiple cells, and that promoter 222 is also activated in multiple cells. In some embodiments, promoter 222 may be an alternative promoter to promoter 212. In some embodiments, promoter 212 or promoter 222 may be activated based on a test agent 202 that affects a variant of target 204. In some embodiments, activation of promoter 212 or promoter 222 may indicate toxicity within wells where cells 200 interact with test agent 202.
[0065] Figure 3 An example cell 300 interacting with test agent 302 is depicted. In this depicted example, test agent 302 interacts with the signaling pathways of cell 300. In this depicted example, test agent 302 leads to the activation of promoter 322 through its interaction with the cell's signaling mechanisms. Depending on the experimental design, this activation may represent a promoter intended to inquire about off-target or targeted effects. Furthermore, in this depicted example, promoter 312 of reporter gene construct 310 is not activated by the test agent. Activation of promoter 322 leads to the expression 328 of reporter gene 324 operatively coupled to promoter 322. Using sequencing technology, barcode 326 of reporter gene 324 can be identified as being expressed by expression 328, indicating one or more aspects of the interaction between test agent 302 and cell 300. For example, barcode 326 may indicate that promoter 322 is activated as an off-target effect of test agent 302.
[0066] Therefore, by using multiple cell lines engineered to include one or more reporter gene constructs, the bioactivity of the active cells of the assay can be identified based on how they interact with certain promoters by identifying the resulting barcode.
[0067] These examples are not limited to assays using heterologous peptides or variants thereof. Cell lines lacking heterologous peptides but containing reporter gene constructs can also be used to understand the effect of a given assay on cellular bioactivity. For example, multiple wells may include multiple cell lines, wherein the multiple cell lines include at least a first cell line and a second cell line. In some embodiments, the first cell line may include a first reporter gene construct comprising a first promoter operatively coupled to a first reporter gene including a first barcode, and the second cell line may include a second reporter gene construct comprising a second promoter operatively coupled to a second reporter gene including a second barcode. In some embodiments, the first promoter may be inducible or constitutive. In some embodiments, the second promoter may be inducible or constitutive. In some embodiments, both the first and second promoters may be inducible. In some embodiments, the first promoter may be a first inducible promoter, and the second promoter may be a second inducible promoter different from the first promoter. In some embodiments, multiple cell lines may include a third engineered cell line having a third reporter gene construct comprising a third promoter operatively coupled to a third reporter gene including a third barcode. In some implementations, the third promoter may be inductive. In some implementations, the third promoter may differ from the first and second promoters. In some implementations, if the first and second promoters are inductive, then the third promoter is constitutive.
[0068] Multiple cell lines may include one or more heterologous polypeptides. In some embodiments, each of the first, second, and third cell lines includes the same heterologous polypeptide. In some embodiments, each of the first, second, and third cell lines includes a different heterologous polypeptide. In some embodiments, any combination of the first, second, and third cell lines has the same heterologous polypeptide, while the remainder has different polypeptides. In some embodiments, the heterologous polypeptides of the cell lines in the multiple cell lines may be coupled to a transcription factor. In some embodiments, the transcription factor comprises one or more of aGal4, PPR1, Lac9, zinc finger, or LexA DNA-binding domains. In some embodiments, the transcription factor comprises one or more of aVP64, p65, RoTev, or Rta DNA-activating domains.
[0069] When an assay reagent is added to one or more partitions comprising multiple cell lines, it can activate a first, second, or third promoter, leading to the expression of a first, second, or third reporter gene, respectively, and allowing the identification of a first, second, or third barcode. Identification of the first, second, or third barcode thus identifies the activation of the corresponding promoter, and consequently, the ability of the assay reagent to regulate heterologous peptides that activate the corresponding promoter. In doing so, the ability of various assay reagents to regulate specific heterologous peptides can be studied.
[0070] Figure 4 Depicting cells containing one or more engineered cell lines (e.g., respectively) Figure 1 , Figure 2 and / or Figure 3 Examples of cell partitions (e.g., wells of a multiwell plate) of 100, 200, and / or 300 cells. During bioassay, a test reagent can be introduced into the well, resulting in the activation of one or more reporter gene constructs. Sequencing can then be performed to identify barcodes associated with these reporter gene constructs. Thus, multiple reporter genes of this disclosure can simultaneously generate information about responses to one or more test reagents.
[0071] Figures 5A-5D Example nucleic acids encoding heterologous polypeptides and / or containing reporter gene constructs are depicted. Although described as single nucleic acids for simplicity, those skilled in the art will understand that targets encoding nucleic acids and reporter gene constructs containing nucleic acids can be provided on different nucleic acids within the wells. Figure 5A Two reporter gene constructs were described for analyzing the expression of reporter genes with different response elements, allowing for the identification of different signaling pathways activated by a single molecule. Furthermore, including reporter genes in the absence of a target allows for the assessment of off-target activation. Figure 5B Two distinct heteropeptides were depicted that transmit signals through the same response element, allowing analysis of whether an assay activates a target or a variant of that target. Furthermore, the inclusion of a reporter gene in the absence of a heteropeptide allows for the assessment of off-target activation. Figures 5C-5D A system for evaluating the effects of an assay pair on the interaction with multiple targets using multiple response elements is described. Although described as a single nucleic acid for simplicity, heteropeptides and response elements can be appropriately provided on different nucleic acids.
[0072] Test reagents In some embodiments, methods are described herein using the bioassays or systems described herein. In some embodiments, the method includes contacting multiple engineered cell lines with an assay agent (such as a compound) and measuring the activity of a reporter gene construct of a first engineered cell line, a second engineered cell line, a third engineered cell line, or any combination thereof. In some embodiments, the activity of the reporter gene construct includes readings of the reporter gene construct. In some embodiments, the activity of the reporter gene construct includes detecting a barcode encoded by the reporter gene construct. In some embodiments, the activity of the reporter gene construct includes detecting a signal of a protein encoded by the reporter gene construct. For example, the activity of the reporter gene construct may include a signal generated by a fluorescent protein or luciferase protein encoded by the reporter gene construct. In some embodiments, the method includes screening assay agents in one or more partitions, wherein the one or more partitions contain one or more cell lines described herein containing one or more heterologous peptides and one or more reporter gene constructs described herein. In some embodiments, the method includes utilizing multiple n assay agents, wherein the assay agents are contacted with multiple engineered cell lines that have been divided into at least n partitions. In some embodiments, multiple engineered cell lines contain at least 100, 1,000, or at least 10,000 different heterologous peptides. In some embodiments, the method is performed in 96-well, 384-well, or 1536-well plates. In some embodiments, a first reporter gene construct and a second reporter gene construct are present in different engineered cells in the same wells of a 96-well, 384-well, or 1536-well plate. In some embodiments, multiple n assay reagents are prepared by reacting a core fragment A containing a reactive functional group x with multiple initial test fragments (y-T1, y-T2, ... yT...). n ) sufficient to form multiple n test reagents (A-T1, A-T2, ...AT) n The reaction is carried out under the following reaction conditions to prepare multiple separate reaction mixtures, where x is an amine, aldehyde, borate ester, imide, isothiocyanate, carboxylic acid, halide, hydroxyamidine, or thiourea; each test fragment contains a reactive functional group y and multiple initial test fractions (T1, T2, ... T). n One of ), where y is an amine, borate ester, imide, isothiocyanate, halide, hydroxyamide, thiourea, carboxylic acid, or aldehyde, wherein A-T1, A-T2, ... AT n Each of them is prepared in a well of the test plate.
[0073] In some embodiments, the method includes contacting the test reagent and the heterologous peptide under reaction conditions. In some embodiments, the reaction conditions include an alkaline environment. In some embodiments, the reaction conditions include Lewis acid or Brønsted acid conditions. In some embodiments, the reaction conditions include contacting the cell line, test reagent, heterologous peptide, or combination thereof with an amide coupling reagent. In some embodiments, the reaction conditions include contacting the cell line, test reagent, heterologous peptide, or combination thereof with a palladium reagent. In some embodiments, the reaction conditions include contacting the cell line, test reagent, heterologous peptide, or combination thereof at room temperature. In some embodiments, the reaction conditions include contacting the cell line, test reagent, heterologous peptide, or combination thereof at a temperature between about room temperature and about 80°C. In some embodiments, the reaction conditions include contacting the cell line, test reagent, heterologous peptide, or combination thereof for a duration of about 1 to about 24 hours. In some embodiments, the reaction conditions include contacting the cell line, test reagent, heterologous peptide, or combination thereof for a duration of about 6 to about 18 hours. In some embodiments, the reaction of contacting the test reagent with the heterologous peptide is reversible or irreversible. In some embodiments, the reaction conditions include contacting the cell line, test reagent, heterologous peptide, or combination thereof at the nanoscale. In some implementations, the nanoscale reaction is carried out in volumes of 50 nL to 500 nL. In some implementations, the reaction conditions are partitioned in 96-well, 384-well, or 1536-well test plates.
[0074] This article describes a method for screening assays that regulate cellular biological activity, including performing a first screening procedure, comprising: (a) By making the core fragment A containing the reactive functional group x with multiple initial test fragments (y-T1, y-T2, ... yT) n ) sufficient to form multiple n test reagents (A-T1, A-T2, ...AT) n The reaction is carried out under the following conditions to prepare multiple separate reaction mixtures, where x is an amine, aldehyde, borate ester, imide, isothiocyanate, carboxylic acid, halide, hydroxymidamine, or thiourea; each test fragment contains a reactive functional group y and multiple initial test fractions (T1, T2, ... T...). n One of ), where y is an amine, borate ester, imide, isothiocyanate, halide, hydroxyamide, thiourea, carboxylic acid, or aldehyde, wherein A-T1, A-T2, ... AT n (a) Each of the n test reagents is prepared in a well of a test plate; (b) a variety of engineered cell lines are contacted with one or more reaction mixtures of (a) under conditions that allow any one or more of the n test reagents to regulate the biological activity of the cells.
[0075] As used herein, “core fragment” means a compound selected for use in the methods disclosed herein, and is also denoted as “A”. “Selected” means that a person skilled in the art has identified the core as having potential utility as a reference or anchor for potential ligands for development targets.
[0076] In some embodiments, core fragment A contains a reactive functional group "x". The reactive functional group x can be a chemical functional group in the core fragment structure. In some embodiments, the core fragment containing the reactive functional group x is denoted by "(Ax)".
[0077] In some embodiments, x is an amine, aldehyde, borate, imide, isothiocyanate, carboxylic acid, halide, hydroxymidamine, or thiourea. In some embodiments, x is a borate. In some embodiments, x is an imide or isothiocyanate. In some embodiments, x is a halide. In some embodiments, x is hydroxymidamine. In some embodiments, x is thiourea.
[0078] In some embodiments, x is an amine. In some embodiments, the amine is a primary or secondary amine. In some embodiments, x is a primary amine. In some embodiments, x is a secondary amine.
[0079] In some embodiments, x is an aldehyde or a carboxylic acid. In some embodiments, x is an aldehyde. In some embodiments, x is a carboxylic acid.
[0080] As used in this article, "initial test component" refers to a compound that has an inherent binding affinity for the target and is a component of the test reagent, and can be represented as "T"; for example, T1, T2, ... T n .
[0081] As used in this paper, the “initial test fragment” (yT) refers to the test portion T that binds to the reactive functional group “y”.
[0082] In some embodiments, each test fragment contains a reactive functional group y. In some embodiments, y is an amine, borate ester, imide, isothiocyanate, halide, hydroxyamide, thiourea, carboxylic acid, or aldehyde. In some embodiments, y is a borate ester. In some embodiments, y is an imide or isothiocyanate. In some embodiments, y is a halide. In some embodiments, y is a hydroxyamide or thiourea.
[0083] In some embodiments, y is an amine. In some embodiments, the amide is a primary or secondary amine. In some embodiments, y is a primary amine. In some embodiments, y is a secondary amine. In some embodiments, y is a carboxylic acid or an aldehyde. In some embodiments, y is an aldehyde. In some embodiments, y is a carboxylic acid.
[0084] The test fragment “(yT)” is selected to allow the reaction with the core fragment A to produce the test reagent, denoted as “(AT)”, where the reaction takes place between the two reactive functional groups x and y.
[0085] In some implementations, the reaction conditions in step (a) include a base. Without being theoretically constrained, the base includes, for example, Hunig's base.
[0086] In some embodiments, the reaction conditions of step (a) include Lewis acids or Brønsted acids. In some embodiments, step (a) comprises Lewis acids. In some embodiments, step (a) comprises Brønsted acids.
[0087] In some embodiments, the reaction conditions of step (a) include an amide coupling agent. Without being bound by theory, exemplary coupling agents in some embodiments include, but are not limited to, propylphosphonic anhydride (T3P), HATU, HBTU, COMU, and DMTMM.
[0088] In some embodiments, the reaction conditions of step (a) include a palladium reagent. In some embodiments, the reaction conditions of step (a) are compared to a stoichiometric palladium oxidative addition complex. In some embodiments, the reaction conditions of step (a) include a stoichiometric G3 Pd catalyst.
[0089] In some implementations, the reaction conditions include reductive amination.
[0090] In some embodiments, the reaction conditions include heterocycle formation, such as the closure of hydroxyamidine and carboxylic acid to form oxadiazole, which is then converted in situ to isothiocyanate via primary amine and subsequently coupled with other amines to form thiourea.
[0091] In some embodiments, the reaction conditions include Buchwald coupling-type chemical reactions (e.g., with a stoichiometric palladium oxidative addition complex or with a stoichiometric G3 Pd catalyst) to form secondary or tertiary aromatic amines from aryl halides and primary or secondary amines.
[0092] In some embodiments, the reaction conditions include Suzuki-like coupling, wherein the Pd reagent is combined with boric acid or borate ester to produce sp 2 -sp 3 CC-coupled products.
[0093] In some implementations, the reaction conditions include room temperature.
[0094] In some implementations, the reaction conditions include a reaction temperature between about room temperature and about 80°C, or any temperature thereof.
[0095] In some embodiments, the reaction conditions include a reaction temperature of approximately room temperature. In some embodiments, the reaction conditions include reaction temperatures of approximately 25°C, 30°C, 35°C, 40°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C, 75°C, or 80°C.
[0096] In some embodiments, the reaction takes about 1 to about 24 hours. In some embodiments, the reaction takes about 6 to about 18 hours.
[0097] In some embodiments, the reaction proceeds for approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 hours. In some embodiments, the reaction proceeds for approximately 24 hours. In some embodiments, the reaction proceeds for approximately 22 hours. In some embodiments, the reaction proceeds for approximately 20 hours. In some embodiments, the reaction proceeds for approximately 18 hours. In some embodiments, the reaction proceeds for approximately 16 hours. In some embodiments, the reaction proceeds for approximately 14 hours. In some embodiments, the reaction proceeds for approximately 12 hours. In some embodiments, the reaction proceeds for approximately 10 hours. In some embodiments, the reaction proceeds for approximately 8 hours. In some embodiments, the reaction proceeds for approximately 6 hours. In some embodiments, the reaction proceeds for approximately 4 hours.
[0098] In some implementations, the reaction is reversible or irreversible. In some implementations, the reaction is reversible. In some implementations, the reaction is irreversible.
[0099] In some embodiments, step (a) includes amide coupling, reductive amination, oxidative addition, or heterocycle formation. In some embodiments, step (a) includes amide coupling. In some embodiments, step (a) includes reductive amination. In some embodiments, step (a) includes oxidative addition. In some embodiments, step (a) includes heterocycle formation.
[0100] In some embodiments, step (a) includes Buchwald or Suzuki coupling. In some embodiments, step (a) includes Buchwald coupling. In some embodiments, step (a) includes Suzuki coupling.
[0101] In some embodiments, the mass of core part A is about 150 Da to about 800 Da. In some embodiments, the mass of core part A is about 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750 or 800 Da.
[0102] In some embodiments, the mass of core component A is approximately 150 Da. In some embodiments, the mass of core component A is approximately 200 Da. In some embodiments, the mass of core component A is approximately 300 Da. In some embodiments, the mass of core component A is approximately 400 Da. In some embodiments, the mass of core component A is approximately 500 Da. In some embodiments, the mass of core component A is approximately 600 Da. In some embodiments, the mass of core component A is approximately 700 Da. In some embodiments, the mass of core component A is approximately 800 Da.
[0103] In some implementations, each test section (T1, T2, ... T) n The mass of ) is approximately 80 Da to approximately 500 Da.
[0104] In some implementations, each test section (T1, T2, ... T) n The mass is approximately 80, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, 320, 340, 360, 380, 400, 420, 440, 460, 480, or 500 Da. In some implementations, each test section (T1, T2, ... T... n The mass is approximately 80 Da. In some implementations, each test section (T1, T2, ... T...) has a mass of approximately 80 Da. n The mass is approximately 100 Da. In some implementations, each test section (T1, T2, ... T...) has a mass of approximately 100 Da. n The mass is approximately 200 Da. In some implementations, each test section (T1, T2, ... T...) has a mass of approximately 200 Da. n The mass is approximately 300 Da. In some implementations, each test section (T1, T2, ... T...) has a mass of approximately 300 Da. n The mass is approximately 400 Da. In some implementations, each test section (T1, T2, ... T...) has a mass of approximately 400 Da. n The mass is approximately 500 Da.
[0105] In some implementations, each test reagent (A-T1, A-T2, ... AT) n The mass of each test reagent is less than approximately 1500 Da. In some embodiments, each test reagent (A-T1, A-T2, ... AT) n The mass of each test reagent is less than approximately 1400 Da. In some embodiments, each test reagent (A-T1, A-T2, ... AT) has a mass of less than approximately 1400 Da. n The mass of each test reagent is less than approximately 1300 Da. In some embodiments, each test reagent (A-T1, A-T2, ... AT) has a mass of less than approximately 1300 Da. nThe mass of each test reagent is less than approximately 1200 Da. In some embodiments, each test reagent (A-T1, A-T2, ... AT) has a mass of less than approximately 1200 Da. n The mass of each test reagent is less than approximately 1100 Da. In some embodiments, each test reagent (A-T1, A-T2, ... AT) has a mass of less than approximately 1100 Da. n The mass of each test reagent is less than approximately 1000 Da. In some embodiments, each test reagent (A-T1, A-T2, ... AT) n The mass of each test reagent is less than approximately 900 Da. In some embodiments, each test reagent (A-T1, A-T2, ... AT) n The mass is less than approximately 800 Da.
[0106] In some implementations, each test reagent (A-T1, A-T2, ... AT) n The mass of the ) is approximately 350 Da to approximately 800 Da.
[0107] In some implementations, each test reagent (A-T1, A-T2, ... AT) n The mass of each test reagent is approximately 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, or 850 Da. In some embodiments, each test reagent (A-T1, A-T2, ... AT) has a mass of approximately 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, or 850 Da. n The mass of each test reagent is approximately 300 Da. In some embodiments, each test reagent (A-T1, A-T2, ... AT) n The mass of each test reagent is approximately 400 Da. In some embodiments, each test reagent (A-T1, A-T2, ... AT) n The mass of each test reagent is approximately 500 Da. In some embodiments, each test reagent (A-T1, A-T2, ... AT) n The mass of each test reagent is approximately 600 Da. In some embodiments, each test reagent (A-T1, A-T2, ... AT) has a mass of approximately 600 Da. n The mass of each test reagent is approximately 700 Da. In some embodiments, each test reagent (A-T1, A-T2, ... AT) n The mass is approximately 800 Da.
[0108] In some embodiments, each reaction is carried out at the nanoscale. In some embodiments, the nanoscale reaction is carried out in a volume of about 50 nL to about 500 nL. In some embodiments, the nanoscale reaction is carried out in a volume of about 50, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nL. In some embodiments, the nanoscale reaction is carried out in a volume of about 50 nL. In some embodiments, the nanoscale reaction is carried out in a volume of about 100 nL. In some embodiments, the nanoscale reaction is carried out in a volume of about 150 nL. In some embodiments, the nanoscale reaction is carried out in a volume of about 200 nL. In some embodiments, the nanoscale reaction is carried out in a volume of about 250 nL. In some embodiments, the nanoscale reaction is carried out in a volume of about 300 nL. In some embodiments, the nanoscale reaction is carried out in a volume of about 350 nL. In some embodiments, the nanoscale reaction is carried out in a volume of about 400 nL. In some embodiments, the nanoscale reaction is carried out in a volume of about 450 nL. In some implementations, the nanoscale is carried out in a volume of approximately 500 nL.
[0109] In some implementations, the test board is a 96-well, 384-well, or 1536-well test board. In some implementations, the test board is a 96-well test board. In some implementations, the test board is a 384-well test board. In some implementations, the test board is a 1536-well test board.
[0110] In some implementations, each initial test section (T1, T2, ... T) n They are different.
[0111] In some embodiments, n is 2 to 100,000 or any integer thereof. In some embodiments, n is 2 to 90,000. In some embodiments, n is 2 to 80,000. In some embodiments, n is 2 to 75,000. In some embodiments, n is 2 to 60,000. In some embodiments, n is 2 to 50,000. In some embodiments, n is 2 to 40,000. In some embodiments, n is 2 to 30,000. In some embodiments, n is 2 to 25,000. In some embodiments, n is 2 to 20,000. In some embodiments, n is 2 to 15,000. In some embodiments, n is 2 to 10,000. In some embodiments, n is 2 to 8,000. In some embodiments, n is 2 to 6,000. In some embodiments, n is 2 to 5,000. In some embodiments, n is 2 to 4,000. In some implementations, n is 2 to 3,000. In some implementations, n is 2 to 2,500. In some implementations, n is 2 to 2,000. In some implementations, n is 2 to 1,800. In some implementations, n is 2 to 1,600.
[0112] In some implementations, step (a) is performed in the absence of a target (e.g., in the absence of the heterologous peptides described herein).
[0113] In some embodiments, the method does not utilize mass spectrometry detection. In other embodiments, the method utilizes mass spectrometry detection. The method may include, but is not limited to, LCMS.
[0114] In some embodiments, the reaction mixture of step (a) is not purified before the target is contacted with one or more reaction mixtures.
[0115] This disclosed fragment-based technique combines the advantages of fragment-based methods with the power and speed of high-throughput screening (HTS). The technique is based on functional screening of target-oriented "custom" libraries for assay reagents, meaning the libraries are constructed based on target-specific information. Assay reagents in such libraries are constructed between target-oriented fragments and an initial assay fragment library. The assembly process can be fully automated, reagent costs are low, and there is no need to purify the assembled fragments.
[0116] barcode Variable nucleotide sequences (barcodes or “unique molecular identifiers”) used as indexes can be included as reporter genes as described herein. Additionally, barcodes can be added in separate library preparation reactions. The variable nucleotide sequences described herein can be used as sample indexes for deconvolution of results obtained from the sequencing reactions used herein. Barcodes can include index regions uniquely identifiable to heterologous receptor genes and can be used to identify heterologous receptors activated in the same cells. Barcodes can also, or alternatively, uniquely identify the location of signaling pathways, assays, heterologous peptides, engineered cell lines (e.g., wells of a plate), or any combination thereof. The index regions of a barcode can be continuous along the length of the barcode sequence, and barcodes can include segments of nucleic acid sequences not unique to any single barcode. In one embodiment, a barcode can have more than one unique identifier. In these embodiments, the index regions can be isolated by a segment of nucleic acid removed by cellular mechanisms during transcription into mRNA (e.g., introns).
[0117] In the context of polynucleotides, the term "heterologous" refers to a gene, polynucleotide, or polypeptide transferred to a cell by gene transfer methods known in the art or described herein; if the exogenously derived sequence is retained in progeny cells, the progeny of such cells may also be referred to as containing the heterologous nucleic acid sequence. A cell may already contain an endogenous gene identical to the heterologous recipient gene, or a cell may lack any endogenous gene related to or identical to the heterologous gene. The terms "heterologous cell" or "host cell" refer to a cell intentionally containing a heterologous nucleic acid sequence.
[0118] The index region of the barcode is a polynucleotide sequence that can be used to identify targets that are activated and / or expressed in the same cells as the barcode, because it is unique to a specific heterologous receptor in the context of the screening utilized. In particular, the inclusion of the barcode helps to determine the activity of specific nucleic acid regulatory elements (i.e., receptor response elements such as unique identifiers), which can indicate the activated receptor.
[0119] Once the contents of the cells are released into their respective compartments by a lysis agent, the macromolecular components contained therein (e.g., macromolecular components of the sample, such as RNA, DNA, or proteins) can be further processed within the compartments. According to the methods and systems described herein, the macromolecular component contents of individual samples can have unique identifiers, allowing them to be attributed to the same sample or particle when characterizing those macromolecular components. The ability to attribute characteristics to individual samples or sample groups is provided by specifically assigning unique identifiers to individual samples or sample groups. Barcodes that can be uniquely identified by a unique molecular identifier sequence (UMI or “unique identifier”) or associated with individual samples or sample groups can be assigned to tag or label the macromolecular components (and therefore their characteristics) of the sample with unique identifiers. These unique identifiers can then be used to attribute the components and characteristics of the sample to individual samples or sample groups.
[0120] In some aspects, this is done by partitioning individual samples or groups of samples together with a unique identifier or a barcode containing a UMI. In some aspects, the unique identifier is provided in the form of nucleic acid molecules (e.g., oligonucleotides) containing nucleic acid barcode sequences that can be attached to or otherwise associated with the nucleic acid contents of an individual sample, or to other components of the sample, and particularly fragments of those nucleic acids. Nucleic acid molecules are partitioned such that the nucleic acid barcode sequences contained therein are identical between nucleic acid molecules within a given partition, but between different partitions, nucleic acid molecules can and do have different barcode sequences, or at least represent a large number of different barcode sequences across all partitions in a given analysis. In some aspects, only one nucleic acid barcode sequence can be associated with a given partition, but in some embodiments, two or more different barcode sequences can be present.
[0121] Nucleic acid barcode sequences may include about 6 to about 20 or more nucleotides within the sequence of a nucleic acid molecule (e.g., an oligonucleotide). Nucleic acid barcode sequences may include about 6 to about 20, 30, 40, 50, 60, 70, 80, 90, 100 or more nucleotides. In some embodiments, the length of the barcode sequence may be about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or longer. In some embodiments, the length of the barcode sequence may be at least about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or longer. In some embodiments, the length of the barcode sequence may be at most about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or shorter. These nucleotides can be completely continuous, i.e., within a single segment of adjacent nucleotides, or they can be divided into two or more separate subsequences separated by one or more nucleotides. In some embodiments, the length of the separate barcode subsequences can be from about 4 to about 16 nucleotides. In some embodiments, the barcode subsequences can be about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or longer. In some embodiments, the barcode subsequences can be at least about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or longer. In some embodiments, the barcode subsequences can be at most about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or shorter.
[0122] The nucleic acid molecules of the co-partitions may also contain other functional sequences of nucleic acids that can be used to process samples from the co-partitions. These sequences include, for example, targeted or random / universal amplification primer sequences for amplifying genomic DNA from individual samples within the partition, while attaching associated barcode sequences, sequencing primers or primer recognition sites, hybridization or probe sequences, such as those for identifying the presence of sequences or for pulling down barcode nucleic acids, or any of many other potential functional sequences. Other mechanisms using co-partition oligonucleotides may also be employed, including, for example, the aggregation of two or more partitions, one of which contains oligonucleotides, or the micro-distribution of oligonucleotides into partitions, such as partitions within a microfluidic system. In some embodiments, the primers contain barcode oligonucleotides. In some embodiments, the primer sequence is a targeted primer sequence complementary to a sequence in the template nucleic acid molecule. In some embodiments, the first nucleic acid molecule also contains one or more functional sequences, and the second nucleic acid molecule contains one or more functional sequences. In some embodiments, the one or more functional sequences are selected from adaptor sequences, additional primer sequences, primer annealing sequences, sequencing primer sequences, sequences configured to be attached to a flow cell of a sequencer, and unique molecular identifier sequences.
[0123] For example, the barcoded nucleic acid molecules described above (e.g., barcoded oligonucleotides) are added to a sample. In some embodiments, partitions contain barcoded oligonucleotides having the same barcode sequence. In some embodiments, partitions within a plurality of partitions contain barcoded oligonucleotides having the same barcode sequence, wherein each partition within the plurality of partitions contains a unique barcode sequence. In some embodiments, the barcoded oligonucleotide population provides a variety of barcode sequence libraries, including at least about 1,000 different barcode sequences, at least about 5,000 different barcode sequences, at least about 10,000 different barcode sequences, at least about 50,000 different barcode sequences, at least about 100,000 different barcode sequences, at least about 1,000,000 different barcode sequences, at least about 5,000,000 different barcode sequences, or at least about 10,000,000 different barcode sequences or more. Additionally, each barcoded oligonucleotide can have a large number of attached nucleic acid (e.g., oligonucleotide) molecules. Specifically, the number of nucleic acid molecules including a barcoded sequence on a single barcoded oligonucleotide can be at least about 1,000 nucleic acid molecules, at least about 5,000 nucleic acid molecules, at least about 10,000 nucleic acid molecules, at least about 50,000 nucleic acid molecules, at least about 100,000 nucleic acid molecules, at least about 500,000 nucleic acid molecules, at least about 1,000,000 nucleic acid molecules, at least about 5,000,000 nucleic acid molecules, at least about 10,000,000 nucleic acid molecules, at least about 50,000,000 nucleic acid molecules, at least about 100,000,000 nucleic acid molecules, at least about 250,000,000 nucleic acid molecules, and in some embodiments at least about 1 billion nucleic acid molecules or more. A given set of barcoded oligonucleotides may include the same (or common) barcode sequence, different barcode sequences, or a combination of both. A given set of barcoded oligonucleotides may include multiple sets of nucleic acid molecules. Nucleic acid molecules in a given set may include the same barcode sequence. The same barcode sequence may differ from the barcode sequence of another set of nucleic acid molecules. Furthermore, when partitioning the barcoded oligonucleotide population, the resulting partitioned population may also include various barcode libraries, including at least about 1,000 different barcode sequences, at least about 5,000 different barcode sequences, at least about 10,000 different barcode sequences, at least about 50,000 different barcode sequences, at least about 100,000 different barcode sequences, at least about 1,000,000 different barcode sequences, at least about 5,000,000 different barcode sequences, or at least about 10,000,000 different barcode sequences. Additionally, each partition of the population may include at least approximately 1,000 nucleic acid molecules, at least approximately 5,000 nucleic acid molecules, at least approximately 10,000 nucleic acid molecules, at least approximately 50,000 nucleic acid molecules, at least approximately 100,000 nucleic acid molecules, at least approximately 500,000 nucleic acid molecules, at least approximately 1,000,000 nucleic acid molecules, at least approximately 5,000,000 nucleic acid molecules, at least approximately 10,000,000 nucleic acid molecules, at least approximately 50,000,000 nucleic acid molecules, at least approximately 100,000,000 nucleic acid molecules, at least approximately 250,000,000 nucleic acid molecules, and in some implementations at least approximately 1 billion nucleic acid molecules.
[0124] In some implementations, it may be desirable to incorporate multiple different barcodes within a given partition. For example, in some implementations, the barcoded oligonucleotides within a partition may include (1) a common barcode sequence shared by all barcoded oligonucleotides within the partition and (2) a unique identifier or additional barcode sequence that differs within each barcoded oligonucleotide. The common barcode sequence can provide greater assurance of identification in subsequent processing, for example, by providing a strong barcode address or attribution to a given partition as a confirmation of duplication or independence of output from that partition.
[0125] In some embodiments, barcoded oligonucleotides are attached to beads, wherein all nucleic acid molecules attached to a particular bead will include the same nucleic acid barcode sequence, but represent a large number of different barcode sequences within the population of beads used. In some embodiments, hydrogel beads (e.g., containing a polyacrylamide polymer matrix) are used as solid supports and delivery carriers for nucleic acid molecule entry into partitions because they are capable of carrying a large number of nucleic acid molecules and can be configured to release those nucleic acid molecules upon exposure to a specific stimulus, as described elsewhere herein.
[0126] Nucleic acid molecules (e.g., oligonucleotides) can be released from beads upon application of a specific stimulus. In some embodiments, the stimulus may be a photostimulation, such as by cleaving photoinstantaneous bonds that release nucleic acid molecules. In other embodiments, a thermal stimulus may be used, wherein an increase in temperature in the bead environment will cause bond cleavage or other release of nucleic acid molecules from the beads. In yet another embodiment, a chemical stimulus may be used, which cleaves the bonds of nucleic acid molecules to the beads, or otherwise causes the release of nucleic acid molecules from the beads. In one embodiment, such a composition comprises the polyacrylamide matrix described above for encapsulating samples and can be degraded by exposure to a reducing agent such as DTT to release attached nucleic acid molecules.
[0127] Supports, such as pores, substrates, rods, containers, or beads, may be used in the methods disclosed herein. Supports can have any available characteristics and properties, such as any available size, surface chemistry, flowability, robustness, density, porosity, and composition. In some embodiments, the support is the surface of a pore on a plate. In some embodiments, the support can be beads, such as gel beads. Beads can be solid or semi-solid. Further details about the beads are provided elsewhere herein.
[0128] Supports (e.g., beads) may contain anchoring sequences functionalized thereto (e.g., as described herein). Anchoring sequences can be attached to the support via, for example, disulfide bonds. Anchoring sequences may include partial read sequences and / or flow cell functional sequences. Such sequences can allow sequencing of nucleic acid molecules with the attached sequences using a sequencer (e.g., an Illumina sequencer). Different anchoring sequences can be used for different sequencing applications. Anchoring sequences may include, for example, TruSeq or Nextera sequences. Anchoring sequences can have any available characteristics, such as any available length and nucleotide composition. For example, anchoring sequences may contain 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides. In some embodiments, anchoring sequences may contain 15 nucleotides. The nucleotides of the anchoring sequences may be naturally occurring or non-naturally occurring (e.g., as described herein). Beads may contain multiple anchoring sequences attached thereto. For example, a bead may contain multiple first anchoring sequences attached thereto. In some embodiments, a bead may contain two or more different anchoring sequences attached thereto. For example, a bead may contain multiple first anchoring sequences (e.g., Nextera sequences) and multiple second anchoring sequences (e.g., TruSeq sequences) attached thereto. For a bead containing two or more different anchoring sequences attached thereto, the sequence of each different anchoring sequence may be distinguishable from the sequence of each other anchoring sequence at the distal end of the bead. For example, the different anchoring sequences may contain one or more nucleotide differences in the 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotides furthest from the bead.
[0129] In some embodiments, multiple different barcode molecules (e.g., nucleic acid barcode molecules) can be generated on the same support (e.g., beads). For example, two different barcode molecules can be generated on the same support. Alternatively, three or more different barcode molecules can be generated on the same support. Different barcode molecules attached to the same support can contain one or more different sequences. For example, different barcode molecules can contain one or more different barcode sequences and / or other sequences (e.g., start sequences). In some embodiments, different barcode molecules attached to the same support can contain the same barcode sequence. Different barcode molecules attached to the same support can contain the same or different barcode sequences. Similarly, different barcode molecules can contain the same or different UMIs.
[0130] It is anticipated that in some implementations, the sensitivity of barcode sequencing will allow expression levels lower than those required for less sensitive assays. In some implementations, the RNA transcript level is at least or at most about 10, 10⁻⁶. 210 3 10 4 10 5 10 6 10 7 '、10 8 10 9 Or 10 10 Or any range that may be derived from it.
[0131] As used herein, the terms “cell,” “cell line,” and “cell culture” are used interchangeably. All these terms also include freshly isolated cells and cells cultured or expanded in vitro. All these terms also include their progeny, which are any and all subsequent generations. It is understood that all progeny can be different due to intentional or unintentional mutations. In the context of expressing a heterologous nucleic acid sequence, “host cell” or simply “cell” refers to a prokaryotic or eukaryotic cell, and it includes any transformable organism capable of replicating a vector or expressing a heterologous gene encoded by the vector or integrated nucleic acid. Host cells can and have been used as recipients of vectors, viruses, and nucleic acids. Host cells can be “transfected” or “transformed,” which refers to the method by which a foreign nucleic acid, such as a recombinant protein-coding sequence, is transferred or introduced into a host cell. Transformed cells include primary target cells and their progeny.
[0132] In some embodiments, nucleic acid transfer can occur on any prokaryotic or eukaryotic cell. In some aspects, the cells of this disclosure are human cells. In other aspects, the cells of this disclosure are animal cells. In some aspects, the cells are cancer cells, tumor cells, or immortalized cells. In a further aspect, the cells represent disease model cells. In some respects, the cells can be A549, B cells, B16, BHK-21, C2C12, C6, CaCo-2, CAP / , CAP-T, CHO, CHO2, CHO-DG44, CHO-K1, COS-1, Cos-7, CV-1, dendritic cells, DLD-1, embryonic stem (ES) cells or derivatives, H1299, HEK293, 293T, 293FT, Hep G2, hematopoietic stem cells, HOS, Huh-7, induced pluripotent stem (iPS) cells or derivatives, Jurkat, K562, L5278Y, LNCaP, MCF7, MDA-MB-231, MDCK, mesenchymal cells, Min-6, monocytes, Neuro2a, NIH 3T3, NIH3T3L1, K562, NK cells, NSO, Panc-1, PC12, PC-3, peripheral blood cells, plasma cells, primary fibroblasts, RBL, Renca, RLE, SF21, SF9, SH-SY5Y, SK-MES-1, SK-N-SH, SL3, SW403, stimulation-triggered pluripotency gain (STAP) cells or derivatives of SW403, T cells, THP-1, tumor cells, U2O5, U937, peripheral blood lymphocytes, expanded T cells, hematopoietic stem cells, or Vero cells. In some implementations, the cells are HEK293T cells.
[0133] As used herein, the term "passage" is intended to refer to the process of dividing cells to produce a large number of cells from pre-existing cells. Cells can be passaged multiple times before or after any of the steps described herein. Passage involves dividing cells and transferring a small number of cells to each new container. For adherent culture, the cells must first be isolated, typically using a mixture of trypsin and EDTA. A small number of isolated cells can then be used to inoculate new cultures, while the remainder is discarded. Furthermore, the amount of cultured cells can be easily scaled up by distributing all cells into fresh flasks. Cells can be held in the culture and incubated under conditions that allow cell replication. In some embodiments, cells are held in culture conditions that allow cells to undergo 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more rounds of cell division.
[0134] In some implementations, limiting dilution methods can be used to allow the clonal cell population to expand. Methods of limiting dilution of clones are well known to those skilled in the art. Such methods have been described, for example, for hybridomas, but can be applied to any cell type. Such methods are described in: (Cloning hybridoma cells by limiting dilution, Journal of tissue culture methods, 1985, Vol. 9, No. 3, pp. 175-177, by Joan C. Rener, Bruce L. Brown, and Roland M. Nardone), which is incorporated herein by reference.
[0135] The methods disclosed herein include cell culture. Methods for culturing suspension cells and adherent cells are well known to those skilled in the art. In some embodiments, cells are cultured in suspension using commercially available cell culture containers and cell culture media. Examples of commercially available culture containers that may be used in some embodiments include ADME / TOX plates, cell chamber slides and coverslips, cell counting devices, cell culture surfaces, Corning HYPERFlask cell culture containers, coated cultureware, Nalgene Cryoware, culture chambers, culture dishes, glass culture flasks, plastic culture flasks, 3D culture formats, multi-well culture plates, culture plate inserts, glass culture tubes, plastic culture tubes, stackable cell culture containers, hypoxic culture chambers, Petri dish and flask carriers, Quickfit culture containers, and scaled-up cell culture using roller flasks, spin flasks, 3D cell culture, or cell culture bags.
[0136] In other embodiments, the culture medium may be formulated using components well known to those skilled in the art. The cell culture formulations and methods are described in detail in the following references: Short Protocols in Cell Biology J. Bonifacino, et al., eds., John Wiley & Sons, 2003, p. 826; Live Cell Imaging: A Laboratory Manual D. Spector & R. Goldman, eds., Cold Spring Harbor Laboratory Press, 2004, p. 450; Stem Cells Handbook S. Sell, ed., Humana Press, 2003, p. 528; Animal Cell Culture: Essential Methods, John M. Davis, John Wiley & Sons, Mar 16, 2011; Basic Cell Culture Protocols, Cheryl D. Helgason, Cindy Miller, Humana Press, 2005; Human Cell Culture Protocols, Series: Methods in Molecular Biology, Vol. 806, Mitry, Ragai R.; Hughes, Robin D. (eds.). 3rd edition.2012, XIV, 435 page 89, Humana Press; Cancer Cell Culture: Method and Protocols, Simon P. Langdon, Springer, 2004; Molecular Cell Biology.4th edition, Lodish H, Berk A, Zipursky SL, et al., New York: WH Freeman; 2000., Section 6.2 Growth of Animal Cells in Culture, all of which are incorporated herein by reference.
[0137] Sequencing methods for detecting barcodes Massive parallel signature sequencing (MPSS) The first of the next-generation sequencing technologies, massively parallel signature sequencing (or MPSS), was developed at Lynx Therapeutics in the 1990s. MPSS is a bead-based approach that uses a complex adaptor ligation method followed by adaptor decoding to read sequences in four-nucleotide increments. This method makes it susceptible to sequence-specific bias or the loss of specific sequences. Due to the complexity of the technology, MPSS was performed only “internally” at Lynx Therapeutics, and no DNA sequencing machines were sold to independent laboratories. Lynx Therapeutics’ merger with Solexa (later acquired by Illumina) in 2004 led to the development of synthetic sequencing, a simpler method derived from Manteia Predictive Medicine, which rendered MPSS obsolete. However, the fundamental properties of MPSS output are typical of later “next-generation” data types, comprising hundreds of thousands of short DNA sequences. In the case of MPSS, these are often used for cDNA sequencing to measure gene expression levels. In fact, the powerful Illumina HiSeq2000, HiSeq2500, and MiSeq systems are all based on MPSS.
[0138] Polymerase cloning and sequencing (Polony Sequencing) The polymerase chain cloning sequencing method, developed in George M. Church's lab at Harvard, was one of the first next-generation sequencing systems and was used for whole-genome sequencing in 2005. It combines in vitro paired-tag libraries with emulsion PCR, automated microscopy, and ligation-based sequencing chemistry to sequence the *E. coli* genome with >99.9999% accuracy at approximately one-ninth the cost of Sanger sequencing. The technology was licensed to Agencourt Biosciences, subsequently spun out as Agencourt Personal Genomics, and eventually incorporated into the Applied Biosystems SOLiD platform, which is currently owned by Life Technologies.
[0139] 454 pyrosequencing A parallel version of pyrosequencing was developed by 454 Life Sciences, which has been acquired by Roche Diagnostics. This method amplifies DNA within droplets in an oil solution (emulsion PCR), where each droplet contains a single DNA template, which attaches to a single primer-coated bead, forming a cloning colony. The sequencing machine contains numerous picoliter-volume wells, each containing a single bead and a sequencing enzyme. Pyrosequencing uses luciferase to generate light for detecting single nucleotides added to the nascent DNA and combines the data to generate sequence reads. Compared to Sanger sequencing at one end and Solexa and SOLiD at the other, this technology offers moderate read lengths and cost per base.
[0140] Illumina (Solexa) sequencing Solexa, now part of Illumina, developed a sequencing method based on reversible dye terminator technology and internally developed engineered polymerases. Termination chemistry was developed in-house at Solexa, and the concept for the Solexa system was coined by Balasubramanian and Klennerman of the Department of Chemistry at the University of Cambridge. In 2004, Solexa acquired Manteia Predictive Medicine to gain access to massively parallel sequencing technology based on “DNA clusters,” which involves the clonal amplification of DNA on a surface. This cluster technology was acquired in conjunction with Lynx Therapeutics in California. Solexa Ltd. later merged with Lynx to form Solexa Inc.
[0141] In this method, DNA molecules and primers are first attached to a glass slide and amplified using polymerase to form localized clonal DNA colonies, later referred to as “DNA clusters.” To determine the sequence, four types of reversible terminator bases (RT bases) are added, and unincorporated nucleotides are washed away. A camera images the fluorescently labeled nucleotides, and then the dye, along with a 3' terminator, is chemically removed from the DNA, allowing the start of the next cycle. Unlike pyrosequencing, the DNA strand is extended one nucleotide at a time, and images can be acquired at delayed moments, allowing for the capture of very large arrays of DNA colonies from a series of images taken from a single camera.
[0142] Uncoupling enzymatic reactions and image capture allow for optimal throughput and theoretically unlimited sequencing capabilities. Under optimal configuration, the final achievable instrument throughput is determined solely by the camera's analog-to-digital conversion rate, multiplied by the number of cameras, and divided by the number of pixels per DNA colony required for optimal visualization (approximately 10 pixels per colony). In 2012, with cameras operating at A / D conversion rates exceeding 10 MHz, and with available optical, fluidic, and enzymatic capabilities, throughput could be multiples of one million nucleotides per second, roughly equivalent to one human genome per lx coverage. SOLiD sequencing Applied Biosystems' (now Life Technologies) SOLiD technology uses ligation sequencing. Here, all possible pools of fixed-length oligonucleotides are labeled according to the sequencing location. The oligonucleotides are annealed and ligated; the preferential ligation of matching sequences by DNA ligase results in signaling information from the nucleotides at that location. Prior to sequencing, the DNA is amplified by emulsion PCR. The resulting beads (each containing a single copy of the same DNA molecule) are deposited on a glass slide. The result is a sequence with a quantity and length comparable to Illumina sequencing. This ligation-based sequencing method has reportedly encountered some problems when sequencing palindromic sequences.
[0143] Ion Torrent semiconductor sequencing Ion Torrent Systems Inc. (now owned by Life Technologies) has developed a novel semiconductor-based detection system that utilizes standard sequencing chemistry. This sequencing method is based on detecting hydrogen ions released during DNA polymerization, rather than the optical methods used in other sequencing systems. Microwells containing the template DNA strand to be sequenced are flooded with a single type of nucleotide. If an introduced nucleotide is complementary to the leader template nucleotide, it is incorporated into the growing complementary strand. This results in the release of hydrogen ions, triggering a hypersensitive ion sensor, indicating a reaction has occurred. If a homopolymer repeat sequence is present in the template sequence, multiple nucleotides are incorporated in a single cycle. This results in a corresponding number of hydrogen ions released and a proportionally higher electron signal.
[0144] DNA nanosphere sequencing DNA nanosphere sequencing is a high-throughput sequencing technology used to determine the entire genome sequence of an organism. Complete Genomics uses this technology to sequence samples submitted by independent researchers. The method uses rolling circle replication to amplify small fragments of genomic DNA into DNA nanospheres. Unchained sequencing by ligation is then used to determine the nucleotide sequence. Compared to other next-generation sequencing platforms, this DNA sequencing method allows sequencing a large number of DNA nanospheres per run and is inexpensive in terms of reagents. However, only short DNA sequences are determined from each DNA nanosphere, making it difficult to map these short reads to a reference genome. This technology has been used in several genome sequencing projects and is planned for use in more projects.
[0145] Heliscope single-molecule sequencing Heliscope sequencing is a single-molecule sequencing method developed by Helicos Biosciences. It uses DNA fragments with poly-A tail adaptors attached to the surface of a flow cell. Subsequent steps involve extension-based sequencing, where the flow cell is cycled and washed with fluorescently labeled nucleotides (one nucleotide type at a time, such as using the Sanger method). Reads are performed using the Heliscope sequencer. Reads are short, up to 55 bases per run, but recent improvements allow for more accurate readings of one nucleotide segment at a time. This sequencing method and instrument were used for genome sequencing of M13 bacteriophage.
[0146] Single-molecule real-time (SMRT) sequencing SMRT sequencing is based on a synthetic sequencing approach. DNA is synthesized in a zero-mode waveguide (ZMW)—a small, porous container with a trapping device located at the bottom. Sequencing is performed using an unmodified polymerase (attached to the bottom of the ZMW) and fluorescently labeled nucleotides flowing freely in solution. The pores are constructed in a way that detects fluorescence only at the bottom. The fluorescent label separates from the nucleotides as it is incorporated into the DNA strand, leaving the unmodified DNA strand. According to Pacific Biosciences, the developer of SMRT technology, this method allows the detection of nucleotide modifications, such as cytosine methylation. This is achieved by observing polymerase kinetics. This method allows reads of 20,000 nucleotides or more, with an average read length of 5 kilobases.
[0147] Next-generation sequencing As described in the methods disclosed herein, sequencing of nucleic acid molecules is used, and this can be used to detect biological effects via cell-based assays. Generally, sequencing refers to methods and techniques for determining the sequence of nucleotide bases in one or more polynucleotides. Polynucleotides can be, for example, deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), including variants or derivatives thereof (e.g., single-stranded DNA). Sequencing can be performed using a variety of currently available systems, such as, but not limited to, sequencing systems from Illumina, Pacific Biosciences, Oxford Nanopore, or Life Technologies (Ion Torrent). Such devices can provide multiple raw genetic data corresponding to the genetic information of an object (e.g., a human), such as data generated from a sample provided by the object through the device. In some cases, the systems and methods described herein can be used in conjunction with proteomic information. Alternatively or additionally, sequencing can be performed using nucleic acid amplification, polymerase chain reaction (PCR) (e.g., digital PCR, quantitative PCR, or real-time PCR), or isothermal amplification. Such systems can provide multiple raw genetic data corresponding to the genetic information of an object (e.g., a human), such as data generated from a sample provided by the object through the system. In some examples, such systems provide sequencing reads (also referred to as “reads” in this document). A read can comprise a string of nucleic acid bases that corresponds to the sequence of a sequenced nucleic acid molecule. In some cases, the systems and methods presented herein can be used in conjunction with proteomic information.
[0148] Next-generation sequencing encompasses many technologies capable of generating large amounts of sequence information, but excludes Sanger sequencing or Maxam-Gilbert sequencing. Generally, next-generation sequencing covers single-molecule real-time sequencing, sequencing-on-synthesis, and ion semiconductor sequencing. Exemplary next-generation sequencing machines may include the MiniSeq, iSeq100, NextSeq 1000, NextSeq 2000, NovaSeq 6000, and NextSeq 550 series from Illumina, Inc.; the Ion Torrent machine from Thermo Fisher Scientific; or the Sequel system from Pacific Biosciences.
[0149] Next-generation sequencing machines used with the methods described in this paper can generate at least 1, 5, 10, 15, 25, 50, 75, 100, 200, 300 gigabases of data or more from a single machine within a 24-hour timeframe.
[0150] Next-generation sequencing machines used with the methods described in this paper can generate at least 100, 100, 400, 1,000, 1,500, 2,500, 5,000, 7,500, 10,000, 20,000, 30,000, 50,000, or 100,000 reads or more from a single machine within a 24-hour timeframe.
[0151] It also includes computer programs, computing devices, or analysis platforms / systems that receive and analyze sequencing data and output one or more reports, which can be transmitted or accessed electronically via a server, analysis portal, or email. The computing devices or analysis platforms may operate according to the algorithms and methods described herein.
[0152] reaction mixture This document also provides a reaction mixture for determining the expression levels of reporter genes in a sample by sequencing. In some embodiments, the reaction mixture comprises a control nucleic acid provided herein, at least a portion of the biological sample, and one or more enzymes or reagents sufficient to amplify barcodes in the sample (if present).
[0153] The control nucleic acid can be any one or more of single-stranded DNA, double-stranded DNA, single-stranded RNA, or double-stranded RNA. In some embodiments, the control nucleic acid is present at a concentration of about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 320, 340, 360, 380, 400, 420, 440, 460, 480, or 500 copies / reaction mixture.
[0154] In some implementations, the enzyme or reagent includes reverse transcriptase, dNTPs, primer pairs specific to barcode or control nucleic acid sequences, magnesium salts, or combinations thereof.
[0155] In some embodiments, the reaction mixture contains one or more enzymes that can be used to amplify or replicate control nucleic acid or barcode nucleic acid. In some embodiments, the enzyme is a reverse transcriptase. Non-limiting examples of reverse transcriptases include avian myeloblastic leukemia virus (AMV) reverse transcriptase and Moloney mouse leukemia virus (M-MuLV, MMLV) and its variants.
[0156] In some embodiments, the reaction includes deoxynucleotide triphosphates (dNTPs). In some embodiments, the kit contains a mixture of each of the dNTPs necessary for amplifying nucleic acids, as well as any other desired nucleic acids (e.g., dATG, dCTP, dTTP, dGTP).
[0157] In some embodiments, the reaction mixture contains a magnesium salt. In some embodiments, the magnesium salt is included in an amount sufficient to enable the enzyme (e.g., reverse transcriptase) of the reaction to function and amplify the target nucleic acid. In some embodiments, the magnesium salt is magnesium chloride. In some embodiments, the reaction mixture contains magnesium ions at a concentration of about 0.1 mM to about 50 mM. In some embodiments, the concentration of magnesium ions is about 1 mM to about 10 mM.
[0158] In some embodiments, the volume of the reaction mixture is from about 10 μL to about 100 μL. In some embodiments, the volume of the reaction mixture is from about 20 μL to about 90 μL, from about 30 μL to about 80 μL, or from about 40 μL to about 60 μL. In some embodiments, the volume of the reaction mixture is about 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 μL.
[0159] Other forms of compounds In some aspects, the compounds disclosed herein have one or more stereocenters, each of which exists independently in an R or S configuration. The compounds presented herein include all diastereomers, enantiomers, and epimers, as well as suitable mixtures thereof. The compounds and methods provided herein include all cis, trans, syn, anti, engegen (E), and zusammen (Z) isomers, as well as suitable mixtures thereof. In some embodiments, the compounds described herein are prepared as their respective stereoisomers by reacting a racemic mixture of the compounds with an optically active resolving agent to form diastereomer compound / salt pairs, separating the diastereomers, and recovering the optically pure enantiomers. In some embodiments, enantiomer resolution is performed using covalent diastereomer derivatives of the compounds described herein. In another embodiment, diastereomers are separated by a separation / resolution technique based on differences in solubility. In other embodiments, the separation of stereoisomers is carried out by chromatography or by forming diastereomer salts, and by recrystallization or chromatography, or any combination thereof. Jean Jacques, Andre Collet, Samuel H. Wilen, “Enantiomers, Racemates and Resolutions”, John Wiley and Sons, Inc., 1981. In one respect, stereoisomers are obtained through stereoselective synthesis.
[0160] In some embodiments, the compounds described herein are prepared as prodrugs. A “prodrug” is an agent that is converted into a parent drug in vivo. Prodrugs are often useful because, in some cases, they can be more readily administered than the parent drug. For example, they may be bioavailable by oral administration, while the parent drug may not be. Prodrugs may also have improved solubility in pharmaceutical compositions compared to the parent drug. In some embodiments, prodrugs are designed to enhance effective water solubility. An example (but not limited to) of a prodrug is a compound described herein administered as an ester (“prodrug”) to facilitate transmembrane transport, where water solubility is unfavorable for migration, but is subsequently metabolized and hydrolyzed into a carboxylic acid, i.e., the active entity, once inside the cell where water solubility is beneficial. A further example of a prodrug may be a short peptide (polyamino acid) bonded to an acid group, wherein the peptide is metabolized to reveal the active moiety. In some embodiments, upon administration in vivo, the prodrug is chemically converted into a biologically, pharmaceutically, or therapeutically active form of the compound. In some embodiments, the prodrug is enzymatically metabolized through one or more steps or processes to the biologically, pharmaceutically, or therapeutically active form of the compound.
[0161] In one respect, prodrugs are designed to alter the metabolic stability or transport characteristics of a drug, mask side effects or toxicity, improve the flavor of a drug, or change other characteristics or properties of a drug. Based on knowledge of pharmacokinetic, pharmacodynamic processes, and in vivo drug metabolism, prodrugs of a known active drug compound can be designed.
[0162] In some implementations, some of the compounds described herein may be prodrugs of another derivative or active compound.
[0163] In some embodiments, sites on the aromatic ring moiety of the compounds described herein are susceptible to various metabolic reactions. Therefore, incorporating appropriate substituents into the aromatic ring structure will reduce, minimize, or eliminate this metabolic pathway. In specific embodiments, by way of example only, suitable substituents for reducing or eliminating the susceptibility of the aromatic ring to metabolic reactions are halogen or alkyl groups.
[0164] In another embodiment, the compounds described herein are labeled by isotopes (e.g., with radioactive isotopes) or by another other means, including but not limited to the use of chromophores or fluorescent portions, bioluminescent labeling, or chemiluminescent labeling.
[0165] The compounds described herein include isotopically labeled compounds that are identical to those shown in the various formulas and structures presented herein, but in fact, one or more atoms are replaced by atoms with atomic masses or mass numbers different from those commonly found in nature. Examples of isotopes that may be incorporated into these compounds include isotopes of hydrogen, carbon, nitrogen, oxygen, sulfur, fluorine, chlorine, and iodine, such as, for example,2 H, 3 H, 13 C 14 C 15 N、 18 O、 17 O、 35 S, 18 F, 36 Cl and 125 I. In one aspect, the isotope-labeled compounds described herein, such as those doped with radioactive isotopes, are... 3 H and 14 Compounds containing C can be used for drug and / or substrate tissue distribution assays. In one respect, substitution with isotopes such as deuterium provides certain therapeutic advantages resulting from greater metabolic stability, such as, for example, increased in vivo half-life or reduced dose requirements.
[0166] In other or further embodiments, the compounds described herein are metabolized when administered to an organism that requires the production of metabolites, and then used to produce desired effects, including desired therapeutic effects.
[0167] As used herein, “pharmaceutically acceptable” refers to materials such as carriers or diluents that do not eliminate the biological activity or properties of a compound and are relatively non-toxic, meaning that the material can be administered to an individual without causing undesirable biological effects or interacting in a harmful manner with any component of the composition containing it.
[0168] The term "pharmaceutically acceptable salt" refers to a formulation of a compound that does not cause significant irritation to the organism to which it is administered, and does not eliminate the biological activity and properties of the compound. In some embodiments, a pharmaceutically acceptable salt is obtained by reacting the compound disclosed herein with an acid. A pharmaceutically acceptable salt is also obtained by reacting the compound disclosed herein with a base to form a salt.
[0169] The compounds described herein can be formed as and / or used as pharmaceutically acceptable salts. Types of pharmaceutically acceptable salts include, but are not limited to: (1) acid addition salts formed by reacting the free base form of the compound with pharmaceutically acceptable substances such as: inorganic acids, such as hydrochloric acid, hydrobromic acid, sulfuric acid, phosphoric acid, metaphosphoric acid, etc.; or with organic acids, such as acetic acid, propionic acid, hexanoic acid, cyclopentylpropionic acid, glycolic acid, pyruvic acid, lactic acid, malonic acid, succinic acid, malic acid, maleic acid, fumaric acid, trifluoroacetic acid, tartaric acid, citric acid, etc. Citric acid, benzoic acid, 3-(4-hydroxybenzoyl)benzoic acid, cinnamic acid, mandelic acid, methanesulfonic acid, ethanesulfonic acid, 1,2-ethanedisulfonic acid, 2-hydroxyethanesulfonic acid, benzenesulfonic acid, toluenesulfonic acid, 2-naphthalenesulfonic acid, 4-methylbicyclo-[2.2.2]oct-2-en-1-carboxylic acid, glucoheponic acid, 4,4'-methylenebis-(3-hydroxy-2-en-1-carboxylic acid), 3-phenylpropionic acid, trimethylacetic acid, tert-butylacetic acid, lauryl (1) Acids such as hydroxysulfuric acid, gluconic acid, glutamic acid, hydroxynaphthylcarboxylic acid, salicylic acid, stearic acid, mucoconic acid, butyric acid, phenylacetic acid, phenylbutyric acid, valproic acid, etc.; (2) Salts formed when the acidic protons present in the parent compound are replaced by metal ions (e.g., alkali metal ions (e.g., lithium, sodium, potassium), alkaline earth metal ions (e.g., magnesium or calcium) or aluminum ions. In some cases, the compounds described herein may coordinate with organic bases such as, but not limited to, ethanolamine, diethanolamine, triethanolamine, tromethamine, N-methylglucosamine, dicyclohexylamine, tri(hydroxymethyl)methylamine. In other cases, the compounds described herein may form salts with amino acids such as, but not limited to, arginine, lysine, etc. Acceptable inorganic bases for forming salts with compounds including acidic protons include, but are not limited to, aluminum hydroxide, calcium hydroxide, potassium hydroxide, sodium carbonate, sodium hydroxide, etc.
[0170] It should be understood that references to pharmaceutically acceptable salts include solvation forms, particularly solvates. Solvates contain stoichiometric or non-stoichiometric amounts of solvent and can be formed by crystallization using pharmaceutically acceptable solvents such as water, ethanol, etc. When the solvent is water, a hydrate is formed, or when the solvent is an alcohol, an alcoholic compound is formed. Solvates of the compounds described herein can be conveniently prepared or formed using the methods described herein. Furthermore, the compounds provided herein can exist in both unsolvated and solvated forms. Generally, for the purposes of the compounds and methods provided herein, the solvated form is considered equivalent to the unsolvated form.
[0171] definition Unless otherwise specified herein, the definitions of the terms used are the standard definitions used in the fields of organic and peptide synthesis, medicinal chemistry, and pharmaceutical science.
[0172] Unless otherwise expressly stated, the singular forms “a,” “an,” and “the” used in this specification and the appended claims include plural indicators. It should also be noted that, unless otherwise expressly stated, the term “or” is generally used to include the meaning of “and / or.” Furthermore, the headings provided herein are for convenience only and do not constitute an explanation of the scope or meaning of the claimed invention.
[0173] As used in this article, the term "testing reagent" refers to a molecular compound of any size and includes small molecule compounds, peptides, polypeptides, antibodies, nucleic acid molecules, etc.
[0174] As used herein, the terms “bioactivity” or “cellular function” refer to any change in the biological state of a cell, including but not limited to gene expression, cell growth and proliferation, cellular metabolic activity, and the synthesis of proteins, RNA, or DNA by the cell.
[0175] The term "target" refers to a chemical or biological entity to which a ligand or molecule has an inherent binding affinity. A target can be a molecule, a portion of a molecule, or an aggregate of molecules. Specific examples of targets include peptides, proteins, receptor ligands, allosteric enzyme regulators, immunoglobulins, polynucleotides, carbohydrates, glycolipids, and other macromolecules such as protein complexes, nucleic acid-protein complexes, chromatin, ribosomes, lipid bilayer structures such as membranes, or membrane-derived structures such as vesicles. Proteins and protein complexes include, but are not limited to, cell surface receptors, nuclear receptors, enzymes, receptor tyrosine kinases, cell signaling proteins, cytokines, chemokines, structural or organelle proteins.
[0176] As used herein, “protein” means any molecule comprising two or more peptide units, each peptide unit containing amino acid residues arranged in a linear chain and linked together by peptide bonds. A protein chain containing more than 30 amino acid residues may be called a polypeptide. A protein chain with 30 or fewer amino acid residues may be called an oligopeptide. Proteins include, but are not limited to, enzymes (e.g., cysteine proteases, serine proteases, and aspartic proteases), receptors, transcription factors, growth factors, cytokines, immunoglobulins, nucleoproteins, signal transduction components (e.g., kinases, phosphatases), and glycoproteins.
[0177] As defined in this article, a “ligand” is a molecule that has an intrinsic binding affinity to a target. Ligands are typically small organic molecules with an intrinsic binding affinity to a target, but can also be other sequence-specific binding molecules, such as peptides (D-, L-, or a mixture of D- and L-), peptide mimics, complex carbohydrates, antibodies, or other oligomers that have the ability to specifically bind to a target.
[0178] A promoter is a control sequence. A promoter is typically a region of a nucleic acid sequence that controls the initiation and rate of transcription. It may contain genetic elements that regulate proteins and molecules, such as RNA polymerases and other transcription factors, to which they can bind. The phrases “operationally located,” “operationally linked,” “controlled,” and “transcriptionally controlled” mean that the promoter is in the correct functional position and / or orientation relative to the nucleic acid sequence to control the initiation and expression of that sequence. A promoter may or may not be used with the term “enhancer,” which refers to a cis-regulatory sequence involved in the transcriptional activation of a nucleic acid sequence. The term promoter is used interchangeably with the term “response element.” Response elements that can be used in the methods described herein include: cAMP response elements (CREs), nuclear factor response elements that activate T cells (NFAT-REs), serum response elements (SREs) and serum response factor response elements (SRF-REs), androgen response elements (AREs), glucocorticoid response elements (GREs), HREs (hormone response elements), EREs (estrogen response elements), and combinations thereof. As used herein, the phrase "measurable" in connection with binding affinity or other affinity parameters means that the value of the affinity parameter is reliably detectable for the ligand of the target. Those skilled in the art will understand that different affinity parameters can be measured with varying degrees of precision and accuracy. Ideally, the precision, accuracy, and dynamic range of the determination will readily adapt to a range of values, allowing the study of ligands exhibiting a wide range of measurements for the affinity parameter. Technicians will typically establish thresholds by which test results can be considered meaningful. For example, in enzyme inhibition assays, it may be necessary to establish an IC50 value below a preselected concentration for the enzyme inhibitor to be considered meaningful. To illustrate, an IC50 value can be established for the biochemical assay of a target protein. 50 Threshold. You can pre-select options that do not meet the threshold but still display IC. 50 The weaker target portion. Then, according to the screening procedure of the present invention, one or more test reagents can be evaluated as more potent and compliant with IC. 50 Threshold. In other cases, the degree of improvement in the effectiveness of the decoy portion, for example, a 10-fold improvement, can be used to evaluate the chosen threshold.
[0179] As used herein, the term "monophore" refers to a monomeric unit of the assay. The term "diaphore" refers to a unit formed by the covalent linkage of two monophasores, i.e., the assay. Ideally, these units have a higher affinity for the target because the two constituent monophasores bind to two independent but adjacent sites on the target. The binding affinity of a diaphore (assay), which is a product of the affinity of a single monophasore, can be referred to as "avidity." The use of the term "diaphore" is independent of whether the unit is covalently bound to the target or exists independently after its release from the target.
[0180] "Small molecules" typically have a molecular weight of about 2,000 Da or less, and include, but are not limited to, synthetic organic or inorganic compounds, peptides, (poly)nucleotides, (oligo)saccharides, etc. Small molecules particularly include small non-polymeric (e.g., non-peptide or polypeptide) organic and inorganic molecules. Many pharmaceutical companies possess large libraries of such molecules, which can be readily used in the methods of this invention. In one embodiment, the molecular weight of the small molecule is up to about 1,000 Da. In another embodiment, the molecular weight of the small molecule is less than about 650 Da. In one embodiment, the molecular weight of the small molecule is up to about 300 Da. This definition includes small organic (including non-polymeric) molecules containing metals (such as Zn, Hg, Fe, Cd, and As) that can form bonds with nucleophiles.
[0181] A "site" on a target refers to the location where a specific ligand binds. This site can include a specific sequence of the monomeric subunit, such as an amino acid residue or nucleotide, and can have a characterized three-dimensional structure. Typically, the molecular interactions between the ligand and the site of interest on the target are non-covalent and include hydrogen bonds, van der Waals interactions, and electrostatic interactions. In the case of peptides, the site of interest broadly includes amino acid residues involved in the binding of the target to the molecule, forming a native complex with the molecule in vivo or in vitro.
[0182] For example, when a target is a protein that exerts its biological effect by binding to another protein, such as a hormone, cytokine, or other protein involved in signal transduction, it can form a natural complex with one or more other proteins in vivo. In this case, the site of interest is defined as a specific protein: a key contact residue involved in the protein-protein binding interface. A key contact residue is defined as those amino acids on the first protein that are in direct contact with an amino acid on the second protein, and when mutated to alanine, the binding affinity is reduced to at least 1 / 10, or alternatively at least 1 / 20, as measured by a direct binding or competitive assay (e.g., ELISA).
[0183] The term “antagonist” is used in the broadest sense and includes any ligand that partially or completely blocks, inhibits or neutralizes the biological activity exhibited by the target.
[0184] The term “agonist” is used in the broadest sense and includes any ligand that mimics the biological activity exhibited by a target (such as a target) by, for example, altering (increasing or inhibiting) existing biological activity or triggering new biological activity by specifically altering the function or expression of such a target, or by the efficiency of signal transduction through such a target.
[0185] As used in this article, “adjustment conditions” means subjecting the target to any single, combined, or series of necessary reaction conditions or reagents to induce the formation or disruption of covalent bonds between the ligand and the target.
[0186] "Active" or "active" means measurable, quantitative biological and / or immunological properties. Examples of cellular biological activity include protein-protein binding, transcriptional activity, cell growth and division, protein synthesis and folding, and the catalytic activity of enzymes.
[0187] As used herein, “derivative” means a compound obtained from another compound (i.e., the “parent” compound) and containing an essential element of the parent compound, or a compound whose structure is related to that of such a parent compound. “Derivative” encompasses compounds that can be obtained directly from the parent compound, or compounds that can be obtained from its common intermediates using similar chemical methods. For example, adenine is a derivative of purine.
[0188] Unless otherwise stated, the following terms as used herein have the following meanings: "Oxyto" refers to the =O substituent.
[0189] "Alkyl" refers to a straight-chain or branched hydrocarbon chain group having one to twenty carbon atoms, attached to the rest of the molecule by a single bond. Alkyl groups containing up to 10 carbon atoms are called C1-C alkyl groups. 10 Alkyl groups, for example, are C1-C6 alkyl groups containing up to six carbon atoms. Alkyl groups containing other numbers of carbon atoms (and other parts defined herein) are similarly represented. Alkyl groups include, but are not limited to, C1-C6 alkyl groups. 10Alkyl groups include C1-C9 alkyl, C1-C8 alkyl, C1-C7 alkyl, C1-C6 alkyl, C1-C5 alkyl, C1-C4 alkyl, C1-C3 alkyl, C1-C2 alkyl, C2-C8 alkyl, C3-C8 alkyl, and C4-C8 alkyl. Representative alkyl groups include, but are not limited to, methyl, ethyl, n-propyl, 1-methylethyl (isopropyl), n-butyl, isobutyl, sec-butyl, n-pentyl, 1,1-dimethylethyl (tert-butyl), 3-methylhexyl, 2-methylhexyl, 1-ethyl-propyl, etc. In some embodiments, the alkyl group is methyl or ethyl. Unless otherwise specified in this specification, the alkyl group may be optionally substituted as described below.
[0190] "Alkylene" refers to a straight-chain or branched divalent hydrocarbon chain in which the remainder of the molecule is attached to a group. In some embodiments, the alkylene is -CH2-, -CH2CH2-, or -CH2CH2CH2-. In some embodiments, the alkylene is -CH2-. In some embodiments, the alkylene is -CH2CH2-. In some embodiments, the alkylene is -CH2CH2CH2-.
[0191] "Alkoxy" refers to a group of formula -OR, where R is an alkyl group as defined. Unless otherwise specified in this specification, the alkoxy group may be optionally substituted as described below. Representative alkoxy groups include, but are not limited to, methoxy, ethoxy, propoxy, butoxy, and pentoxy. In some embodiments, the alkoxy group is methoxy. In some embodiments, the alkoxy group is ethoxy.
[0192] "Heteroalkyl" refers to an alkyl group as described above, wherein one or more carbon atoms of the alkyl group are substituted with O, N (i.e., NH, N-alkyl), or S atoms. "Heteroalkylene" refers to a straight-chain or branched divalent heteroalkyl chain to which the remainder of the molecule is attached. Unless otherwise specified in this specification, heteroalkyl or heteroalkylene groups may be optionally substituted as described below. Representative heteroalkyl groups include, but are not limited to, -OCH2OMe, -OCH2CH2OMe, or -OCH2CH2OCH2CH2NH2. Representative heteroalkylene groups include, but are not limited to, -OCH2CH2O-, -OCH2CH2OCH2CH2O-, or -OCH2CH2OCH2CH2OCH2CH2O-.
[0193] "alkylamino" refers to a group of the formula -NHR or -NRR, wherein each R is independently an alkyl group as defined above. Unless otherwise specified in this specification, the alkylamino group may be optionally substituted as described below.
[0194] The term "aromatic" refers to a planar ring having a delocalized π-electron system containing 4n + 2 π electrons, where n is an integer. Aromatic compounds can be optionally substituted. The term "aromatic" includes aryl groups (e.g., phenyl, naphthyl) and heteroaryl groups (e.g., pyridyl, quinolinyl).
[0195] “Aryl” refers to an aromatic ring in which each atom forming the ring is a carbon atom. The aryl group may be optionally substituted. Examples of aryl groups include, but are not limited to, phenyl and naphthyl. In some embodiments, the aryl group is phenyl. Depending on the structure, the aryl group may be a mono- or di-group (i.e., an arylene group). Unless otherwise specified in this specification, the term “aryl” or the prefix “ar” (such as in “aralkyl”) means to include optionally substituted aryl groups.
[0196] "Carboxyl group" refers to -CO2H. In some embodiments, the carboxyl moiety may be replaced by a "carboxylic acid bioisostere," which refers to a functional group or moiety exhibiting similar physical and / or chemical properties to the carboxylic acid moiety. Carboxylic acid bioisosteres have similar biological properties to the carboxylic acid group. Compounds having a carboxylic acid moiety can exchange the carboxylic acid moiety with a carboxylic acid bioisostere and have similar physical and / or biological properties when compared to carboxylic acid-containing compounds. For example, in one embodiment, a carboxylic acid bioisostere will ionize at physiological pH to approximately the same extent as a carboxylic acid group. Examples of carboxylic acid bioisosteres include, but are not limited to: , , , , , , , , , wait.
[0197] "Cycloalkyl" refers to a monocyclic or polycyclic non-aromatic group in which each atom forming the ring (i.e., the skeleton atom) is a carbon atom. Cycloalkyl groups can be saturated or partially unsaturated. Cycloalkyl groups can be fused with aromatic rings (in which case the cycloalkyl group is bonded through non-aromatic carbon atoms). Cycloalkyl groups include groups having 3 to 10 ring atoms. Representative cycloalkyl groups include, but are not limited to, cycloalkyl groups having three to ten carbon atoms, three to eight carbon atoms, three to six carbon atoms, or three to five carbon atoms. In some embodiments, the cycloalkyl group is a C3-C6 cycloalkyl group. In some embodiments, the cycloalkyl group is monocyclic, bicyclic, or polycyclic. In some embodiments, the cycloalkyl group is selected from cyclopropyl, cyclobutyl, cyclopentyl, cyclopentenyl, cyclohexyl, cyclohexenyl, cycloheptyl, cyclooctyl, spiro[2.2]pentyl, bicyclo[1.1.1]pentyl, bicyclo[3.3.0]octane, bicyclo[4.3.0]nonane, bicyclo[2.1.1]hexane, bicyclo[2.2.1]heptane, bicyclo[2.2.2]octane, bicyclo[3.2.2]nonane, bicyclo[3.3.2]decane, norbornyl, decahydronaphthyl, and adamantyl. In some embodiments, the cycloalkyl group is monocyclic. Monocyclic cycloalkyl groups include, for example, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl, and cyclooctyl. In some embodiments, the monocyclic cycloalkyl group is cyclopropyl, cyclobutyl, cyclopentyl, or cyclohexyl. In some embodiments, the cycloalkyl group is bicyclic. Bicyclic cycloalkyl groups include fused bicyclic cycloalkyl groups, spirobicyclic cycloalkyl groups, and bridged bicyclic cycloalkyl groups. In some embodiments, the cycloalkyl group is selected from spiro[2.2]pentyl, bicyclic[1.1.1]pentyl, bicyclic[3.3.0]octane, bicyclic[4.3.0]nonane, bicyclic[2.1.1]hexane, bicyclic[2.2.1]heptane, bicyclic[2.2.2]octane, bicyclic[3.2.2]nonane, bicyclic[3.3.2]decane, norbornyl, 3,4-dihydronaphthyl-1(2H)-one, and decahydronaphthyl. In some embodiments, the cycloalkyl group is polycyclic. Polycyclic groups include, for example, adamantyl and . In some embodiments, the polycyclic cycloalkyl group is adamantyl. Unless otherwise specifically stated in this specification, the cycloalkyl group may be optionally substituted.
[0198] "Fused" refers to any ring structure fused with an existing ring structure as described herein. When the fused ring is a heterocyclic or heteroaryl ring, any carbon atom in the existing ring structure that is part of the fused heterocyclic or heteroaryl ring can be replaced by a nitrogen atom.
[0199] "Halogen" or "halogen" refers to bromine, chlorine, fluorine, or iodine.
[0200] "Halogenated alkyl" refers to an alkyl group as defined above that has been substituted with one or more halogenated groups as defined above, such as trifluoromethyl, difluoromethyl, fluoromethyl, trichloromethyl, 2,2,2-trifluoroethyl, 1,2-difluoroethyl, 3-bromo-2-fluoropropyl, 1,2-dibromoethyl, etc. Unless otherwise specified in this specification, the halogenated alkyl group may be optionally substituted.
[0201] "Haloalkoxy" refers to an alkoxy group as defined above, substituted with one or more halogenated groups as defined above, such as trifluoromethoxy, difluoromethoxy, fluoromethoxy, trichloromethoxy, 2,2,2-trifluoroethoxy, 1,2-difluoroethoxy, 3-bromo-2-fluoropropoxy, 1,2-dibromoethoxy, etc. Unless otherwise specified in this specification, the haloalkoxy group may be optionally substituted.
[0202] "Heterocyclic alkyl," "heterocyclic group," or "heterocycle" refers to a stable 3- to 14-membered non-aromatic ring group comprising 2 to 10 carbon atoms and one to four heteroatoms selected from nitrogen, oxygen, and sulfur. Unless otherwise specified in this specification, a heterocyclic alkyl group can be monocyclic, bicyclic (which may include fused bicyclic heterocyclic alkyl groups (where the heterocyclic alkyl group is bonded by non-aromatic atoms when fused with an aryl or heteroaryl ring), bridged heterocyclic alkyl groups, or spirocyclic alkyl groups), or polycyclic. In some embodiments, the heterocyclic alkyl group is monocyclic or bicyclic. In some embodiments, the heterocyclic alkyl group is monocyclic. In some embodiments, the heterocyclic alkyl group is bicyclic. The nitrogen, carbon, or sulfur atom in the heterocyclic group may optionally be oxidized. The nitrogen atom may optionally be quaternized. The heterocyclic alkyl group is partially or fully saturated. Examples of such heterocyclic alkyl groups include, but are not limited to, dioxolane, thienyl[1,3]dithiaalkyl, decahydroisoquinolinyl, imidazolinyl, imidazoalkyl, isothiazolinyl, isoxazolinyl, morpholinyl, octahydroindolyl, octahydroisoindolyl, 2-oxopiperazinyl, 2-oxopiperidinyl, 2-oxopiperidinyl, 2-oxopiperylalkyl, oxazolinyl, piperidinyl, piperazinyl, 4-piperidinoneyl, pyrrolylalkyl, pyrazolyl, quininecycloyl, thiazoalkyl, tetrahydrofuranyl, trithiaalkyl, tetrahydropyranyl, thiomorpholinyl, thio-morpholinyl, 1-oxo-thiomorpholinyl, 1,1-dioxo-thiomorpholinyl. The term heterocyclic alkyl also includes all cyclic forms of carbohydrates, including but not limited to monosaccharides, disaccharides, and oligosaccharides. Unless otherwise stated, heterocyclic alkyl groups have 2 to 10 carbons in the ring. In some embodiments, heterocyclic alkyl groups have 2 to 8 carbons in the ring. In some embodiments, the heterocyclic alkyl group has 2 to 8 carbon atoms and 1 or 2 nitrogen atoms in the ring. In some embodiments, the heterocyclic alkyl group has 2 to 10 carbon atoms, 0-2 nitrogen atoms, 0-2 oxygen atoms, and 0-1 sulfur atoms in the ring. In some embodiments, the heterocyclic alkyl group has 2 to 10 carbon atoms, 1-2 nitrogen atoms, 0-1 oxygen atoms, and 0-1 sulfur atoms in the ring. It is understood that when referring to the number of carbon atoms in a heterocyclic alkyl group, the number of carbon atoms in the heterocyclic alkyl group is not the same as the total number of atoms constituting the heterocyclic alkyl group (including heteroatoms) (i.e., the skeletal atoms of the heterocyclic alkyl ring). Unless otherwise specifically stated in this specification, the heterocyclic alkyl group may be optionally substituted.
[0203] "Heteroaryl" refers to an aryl group comprising one or more cyclic heteroatoms selected from nitrogen, oxygen, and sulfur. Heteroaryl groups can be monocyclic or bicyclic. Illustrative examples of monocyclic heteroaryl groups include pyridinyl, imidazolyl, pyrimidinyl, pyrazolyl, triazolyl, pyrazinyl, tetrazolyl, furanyl, thiophene, isoxazolyl, thiazolyl, oxazolyl, isothiazolyl, pyrroleyl, pyridazinyl, triazinyl, oxadiazolyl, thiazolyl, furazolyl, indazine, indole, benzofuran, benzothiophene, indazole, benzimidazole, purine, quinazine, quinoline, isoquinoline, cinnamoline, phthalazine, quinazoline, quinoxaline, 1,8-naphthidine, and pteridine. Illustrative examples of monocyclic heteroaryl groups include pyridinyl, imidazolyl, pyrimidinyl, pyrazolyl, triazolyl, pyrazinyl, tetrazolyl, furanyl, thiopheneyl, isoxazolyl, thiazolyl, oxazolyl, isothiazolyl, pyrroleyl, pyridazinyl, triazinyl, oxadiazolyl, thiazolyl, and furazonyl. Illustrative examples of bicyclic heteroaryl groups include indazine, indole, benzofuran, benzothiophene, indazole, benzimidazole, purine, quinazine, quinoline, isoquinoline, cyclophosphine, quinazoline, quinoxaline, 1,8-naphthidine, and pteridine. In some embodiments, the heteroaryl group is pyridinyl, pyrazinyl, pyrimidinyl, thiazolyl, thiopheneyl, thiazolyl, or furanyl. In some embodiments, the heteroaryl group contains 0-4 nitrogen atoms in the ring. In some embodiments, the heteroaryl group contains 1-4 nitrogen atoms in the ring. In some embodiments, the heteroaryl group contains 0-4 nitrogen atoms, 0-1 oxygen atoms, and 0-1 sulfur atoms in the ring. In some embodiments, the heteroaryl group contains 1-4 nitrogen atoms, 0-1 oxygen atoms, and 0-1 sulfur atoms in the ring. In some embodiments, the heteroaryl group is a C1-C9 heteroaryl group. In some embodiments, the monocyclic heteroaryl group is a C1-C5 heteroaryl group. In some embodiments, the monocyclic heteroaryl group is a 5- or 6-membered heteroaryl group. In some embodiments, the bicyclic heteroaryl group is a C6-C9 heteroaryl group.
[0204] The term "optional substitution" or "substitution" means that the referenced group can be substituted by one or more additional groups, individually and independently selected from alkyl, haloalkyl, cycloalkyl, aryl, heteroaryl, heterocycloalkyl, -OH, alkoxy, aryloxy, alkylthio, arylthio, alkyl sulfoxide, aryl sulfoxide, alkyl sulfone, aryl sulfone, -CN, alkyne, C1-C6 alkylalkyne, halogen, acyl, acyloxy, -CO2H, -CO2alkyl, nitro, and amino, including monosubstituted and disubstituted amino groups (e.g., -NH2, -NHR, -NR2) and their protected derivatives. In some embodiments, the optional substituents are independently selected from alkyl, alkoxy, haloalkyl, cycloalkyl, halogen, -CN, -NH2, -NH(CH3), -N(CH3)2, -OH, -CO2H, and -CO2alkyl. In some embodiments, the optional substituents are independently selected from fluorine, chlorine, bromine, iodine, CH3, -CH2CH3, -CF3, -OCH3, and -OCF3. In some embodiments, the substituted group is replaced by one or two of the aforementioned groups. In some embodiments, the optional substituents on the aliphatic carbon atom (acyclic or cyclic) include oxo (=O).
[0205] "Tautomerism" refers to the transfer of a proton from one atom of a molecule to another atom within the same molecule. The compounds presented herein can exist in tautomer form. Tautomers are compounds that can interconvert through the migration of hydrogen atoms, accompanied by the conversion of single bonds and adjacent double bonds. A chemical equilibrium of tautomers will exist in the possible tautomer bond arrangements. All tautomer forms of the compounds disclosed herein are expected. The exact proportions of tautomers depend on several factors, including temperature, solvent, and pH. Some examples of tautomer interconversion include: .
[0206] Example Example 1. Preparation of compound libraries In some embodiments, the synthesis of the compounds described herein is accomplished using the means described in the chemical literature, the methods described herein, or a combination thereof. Furthermore, the solvents, temperatures, and other reaction conditions presented herein may vary.
[0207] In other embodiments, the starting materials and reagents used to synthesize the compounds described herein are synthetic or obtained from commercial sources, such as, but not limited to, Sigma-Aldrich, Fisher Scientific (FisherChemicals), and Acros Organics.
[0208] In further embodiments, the compounds described herein and other related compounds with different substituents are synthesized using the techniques and materials described herein as well as techniques and materials recognized in the art, such as those described, for example, in: Fieser and Fieser's Reagents for Organic Synthesis, Volumes 1–17 (John Wiley and Sons, 1991); Rodd's Chemistry of Carbon Compounds, Volumes 1–5 and Supplementals (Elsevier Science Publishers, 1989); Organic Reactions, Volumes 1–40 (John Wiley and Sons, 1991), Larock's Comprehensive Organic Transformations (VCH Publishers Inc., 1989), March, Advanced Organic Chemistry, 4th Edition (Wiley, 1992); Carey and Sundberg, Advanced Organic Chemistry, 4th Edition, Volumes A and B (Plenum, 2000, 2001); and Green and Wuts, Protective Groups in Organic Synthesis, 3rd Edition (Wiley, 1992). (1999) (all of which are incorporated herein by reference). General methods for preparing compounds as disclosed herein can be derived from reactions, and reactions can be modified by using appropriate reagents and conditions to introduce various motifs found in the formulas provided herein. As guidance, the following synthetic methods can be utilized.
[0209] In the described reactions, it may be necessary to protect desired reactive functional groups in the final product, such as hydroxyl, amino, imino, thio, or carboxyl groups, to prevent them from unnecessarily participating in the reaction. Detailed descriptions of techniques suitable for creating and removing protecting groups are provided in: Greene and Wuts, Protective Groups in Organic Synthesis, 3rd Edition, John Wiley & Sons, New York, NY, 1999; and Kocienski, Protective Groups, Thieme Verlag, New York, NY, 1994, which are incorporated herein by reference (and are used herein by reference).
[0210] It is understood that other similar procedures and reagents may be used, and these schemes are meant only as non-limiting examples.
[0211] abbreviation DMA: Dimethylacetamide DMTMM: 4-(4,6-dimethoxy-1,3,5-triazin-2-yl)-4-methyl-morpholinium chloride DMSO: Dimethyl sulfoxide HATU: 1-[bis(dimethylamino)methylene]-1H-1,2,3-triazolo[4,5-b]pyridinium 3-oxide hexafluorophosphate HPLC: High Performance Liquid Chromatography HRMS: High Resolution Mass Spectrometry h or hr(s): hours min(s): minutes m / z: Mass-to-charge ratio Example 2. Synthesis of an exemplary amide library Option 1.
[0212] In an example illustrating the ability of high-throughput chemical synthesis to couple one compound per pore, a secondary amide “reactive core” was constructed and coupled with 8639 different carboxylic acids. These conditions were established, for example, in: Chem. Soc. Rev., 2009, 38, 606-631.
[0213] Add 100 nL of each carboxylic acid “fragment” [5 nmol; 2 equivalents] from a 50 mM stock solution of dimethylacetamide (DMA) to a 1536-well plate. Then add 100 nL of DMTMM [6 nmol; 2.4 equivalents] from a 60 mM stock solution, followed by 100 nL of 150 mM Hunig's base [15 nmol; 6 equivalents], and then 50 nL of a secondary amine core (Example Core 1) from a 50 mM stock solution [2.5 nmol; 1 equivalent]. Incubate the resulting mixture at room temperature for approximately 12 hours, after which DMSO is added to each well in a total volume of 3000 nL.
[0214] LCMS analysis revealed that approximately 81% of the 8,639 reactions produced measurable amide products.
[0215] Example 3. Synthesis of an exemplary benzimidazole library Option 2.
[0216] Add 50 nL of 50 mM aldehyde (2.5 nmol; 1 equivalent) from a 2-methoxyethanol (2ME) solution to a 1536-well plate. Then add another 50 nL of a freshly prepared solution of 50 mM reactive exemplary core 2 [2.5 nmol, 1 equivalent], also from a 2ME solution, and then add 50 nL of 50 mM lanthanum(III) chloride [2.5 nmol, 1 equivalent]. The 1536-well plate was then incubated at room temperature for approximately 12 hours, after which the reaction was quenched by adding 3000 nL of DMSO.
[0217] A small portion of the library is analyzed using LCMS to evaluate the overall performance of library synthesis. A 1-minute LCMS method can analyze most of the library. Approximately 10-15% of the library is characterized overall. Figure 6 The reaction is shown in Example 3.
[0218] Using this data and other LCMS-derived information, including isotopic abundance pattern information and m / z values of the expected products, the products are detected. A QC score is calculated for each measurement reaction, for example... Figure 7 As shown in the figure. This score summarizes both the certainty of the desired product manufacturing and the yield estimate of the desired reaction.
[0219] Example 4. Performing bioassays and analyzing barcodes A six-well section of a 1536-well plate was constructed, with each well containing a set of heterologous peptides. Each heterologous peptide in each set contained one or more binding sites that could be used by an assay reagent to modulate activity. Six different assay reagents were prepared in solution, and each different assay reagent was contacted with cells in the corresponding well. After a one-hour period, a bioassay was performed to analyze the resulting barcode.
[0220] In the first wells containing the first reagent from six different assays, a first count of barcodes for reporter genes associated with the first reporter gene construct and a second count of barcodes for reporter genes associated with the second reporter gene construct were determined from bioassays. The first reporter gene construct contains the CRE promoter and is activated by signaling from a heteropeptide from the first pathway, while the second reporter gene construct contains the NFAT promoter and is activated by signaling from a heteropeptide from the second pathway. The first barcode count was 10,000 barcodes, and the second barcode count was 500 barcodes, indicating that when the first reagent was added to the wells, it was more likely to activate the CRE promoter via the first signaling pathway than to activate NFAT via the second signaling pathway.
[0221] In the second well, to which the second reagent from six different assays was added, a third count of barcodes for reporter genes associated with a third reporter gene construct and a fourth count of barcodes for reporter genes associated with a fourth reporter gene construct were determined from bioassays. The third reporter gene construct contains a CRE promoter and is activated by a signal from a heterologous peptide, while the fourth reporter gene construct contains a CRE promoter and is activated by a signal from a variant of the heterologous peptide. The first barcode count was 8,000 barcodes, and the second barcode count was also 8,000 barcodes, indicating that the presence of a variant was not required when the second reagent activated the CRE.
[0222] In the third well containing the third of six different test reagents, a fifth barcode count for the reporter gene associated with the fifth reporter gene construct and a sixth barcode count for the reporter gene associated with the sixth reporter gene construct were determined from bioassays. The fifth reporter gene construct contains a CRE promoter and is activated by a signal from a heteropeptide, while the sixth reporter gene construct contains a constitutively active promoter activated by toxicity in the well. The fifth barcode count was 300 bars, while the third barcode count was 10,000 bars, indicating the presence of toxicity associated with the test reagent.
[0223] In the fourth well containing the fourth of six different assay reagents, a seventh count of barcodes for reporter genes associated with the seventh reporter gene construct and an eighth count of barcodes for reporter genes associated with the eighth reporter gene construct were determined from bioassays. The seventh reporter gene construct contains a CRE promoter activated by a variant of the heterologous peptide, while the eighth reporter gene construct contains an NFAT promoter activated by a variant of the heterologous peptide. A first barcode count of 12,000 barcodes and a second barcode count of 6,000 barcodes indicate that when the fourth assay reagent is added to the well, the probability of CRE activation via the CRE-associated pathway is twice that of NFAT activation via the NFAT-associated pathway.
[0224] In the fifth well containing the fifth of six different assay reagents, a ninth count of barcodes for reporter genes associated with the ninth reporter gene construct and a tenth count of barcodes for reporter genes associated with the tenth reporter gene construct were determined from bioassays. The ninth reporter gene construct contains a CRE promoter activated via heteropeptide signal transduction, while the tenth reporter gene construct is a promoter activated only when the assay reagent does not bind to the heteropeptide. The first barcode count was 50 barcodes, and the second barcode count was 15,000 barcodes, indicating that the addition of the fifth assay reagent to the well is highly unlikely to affect signal transduction via the heteropeptide.
[0225] In the sixth well containing the sixth of six different assay reagents, an eleventh count of barcodes for the reporter gene associated with the eleventh reporter gene construct and a twelfth count of barcodes for the reporter gene associated with the twelfth reporter gene construct were determined from bioassays. The eleventh reporter gene construct contains a CRE promoter activated by the heteropeptide, while the twelfth reporter gene construct is a promoter activated only if the assay reagent does not affect signal transduction via the heteropeptide. The first barcode count was 10,000 barcodes, and the second barcode count was 1,000 barcodes, indicating that when the sixth assay reagent was added to the well, it was likely to affect signal transduction via the heteropeptide.
[0226] Example 5. Screening for aminiergic GPCRs A system was developed to measure the activity of muscarinic, adrenergic, histaminergic, dopaminergic, and serotonergic (MAHDS) GPCRs through three major G protein pathways on our multiplexing platform. Gs activity was measured by the activity of CRE reporter genes. Figures 9A-9B Gi activity was measured by inhibiting the activity of the CRE reporter gene stimulated by forskolin. Figure 9C-9D Gq activity was measured using engineered NFAT to report gene activity. Figure 9E-9F Reporter genes and GPCR expression cassettes were stably integrated into the HEK293 cell line, which was engineered to exhibit low endogenous response, rTTA expression, and high reporter gene signaling. Individual cell lines were engineered to express specific MAHDS GPCRs and G protein reporter genes, each reporter gene linked to a unique barcode. Cell lines were pooled together into single libraries to classify each G protein reporter gene into a cell library. Assays and validation of the cell libraries are described here, and the activity of the MAHDS receptor against its broad-spectrum homologous agonists (acetylcholine, norepinephrine, histamine, dopamine, serotonin) and four antipsychotics (aripiprazole, clozapine, haloperidol, and lurasidone) is reported.
[0227] To validate the multiplexed reporter gene library, the dose-response relationship of the expected interaction between the MAHDS receptor and its endogenous homologous agonist was measured (Figure 10). In the assays, the drug concentrations that induced half (EC50) of the maximum effect of each agonist on its homologous receptor were typically in the range of 1–1000 nM (Figure 10), which is generally consistent with other GPCR activity assays reported by CHEMBL. The overall performance of the library is summarized in […]. Figure 9G middle.
[0228] Currently, this assay can robustly detect 33 of the 35 reported major MAHDS GPCR conjugates (excluding HTR1E and HTR2A), 6 of the 12 reported minor conjugates, and 17 "accidental" conjugates. Figure 9G (and Figure 10).
[0229] For the primary coupling with HTR1E, which was clearly missed, a weak signal was detected in the Gi pathway and a dose-response was calculated (Figure 10), but the signal was not statistically significant. However, a strong Gs coupling was observed for this receptor. For HTR2A, no Gi signaling was observed, but a secondary coupling with Gq via NFAT signaling, as reported therein, was observed.
[0230] Of the six unobserved minor couplings reported, four (ADRA2C, ADRB1, ADRB2, ADRB3) originated from receptors primarily coupled to Gi or Gs and secondarily coupled to the reverse pathway Gs or Gi. Because these pathways are tortuous—one is the reverse activity of the other—it is not always possible to optimize our cell lines and conditions to reliably detect both. For HRH2, although Gs activity was actually detectable, it was removed from the library because it tends to induce high paracrine signaling. For HTR5A, no significant Gq signaling was detected on the NFAT reporter gene.
[0231] Receptor cross-reactivity to non-homologous agonists Because this assay collects responses from all receptors multiplexed, it measures not only the response of each MAHDS receptor to its homologous agonist, but also its response to all endogenous agonists simultaneously. Except for dopaminergic and adrenergic receptors, MAHDS receptors only respond to their homologous endogenous agonists. Figure 11A In the case of adrenergic and dopaminergic receptors, the highest affinity interaction is between their respective homologous agonists.
[0232] Receptor selectivity of cross-reactive adrenergic and dopaminergic receptors To verify that this assay can measure the receptor selectivity of different compounds, the activity profiles of ADRB1 and DRD1 against norepinephrine and dopamine were measured. These agonists are known to activate dopaminergic and adrenergic classes in vivo and in vitro, with each agonist exhibiting a higher affinity for its homologous receptor. The expected receptor selectivity of these agonists was observed; norepinephrine showed approximately 60-fold higher selectivity for ADRB1 than for DRD1, while dopamine showed approximately 80-fold higher selectivity for DRD1 than for ADRB1. Figure 11BAs expected, DRD1 was observed to have a higher affinity for dopamine than for norepinephrine, and ADRB1 had a higher affinity for dopamine than for norepinephrine. These results can be generalized to all receptors, where additional responsiveness to their non-homologous agonists was observed; the highest affinity agonists were homologs of each receptor. Figure 11A ).
[0233] Detection of unexpected GPCR conjugation Interestingly, many of our GPCRs showed activity against reporter gene pathways not reported in the IUPHAR database. For example, among muscarinic (CHRM) family receptors, only CHRM2 has been reported to have Gs coupling (…). Figure 9G However, in the assay, Gs reporter gene activity was observed for all five CHRM receptors, indicating that these receptors are coupled to Gs in our cell lines. Figure 9G Besides dopaminergic receptors, which only respond to their expected primary couplings, these "unexpected" couplings also exist in other receptor classes. Figure 9G These discrepancies could have many causes. It is likely that most of these “accidental” receptor conjugations are not properly annotated in the IUPHAR database, or are not adequately studied / understood. Recent extensive studies of GPCR conjugation with other methods have yielded similar results to those observed. The general approach described here is agnostic to these considerations and only notices when major G protein conjugation is expected but not observed in our assays (as described above).
[0234] Measuring the activity of four model antipsychotic drugs To evaluate the ability of multiplexed reporter gene assays to report complex multiphasic drug profiles, the agonist, anti-agonist, and antagonist activities of four well-known antipsychotics [aripiprazole (Abilify), clozapine (Clozaril), haloperidol (Haldol), and lurasidone (Latuda)] on MAHDS receptors were measured. Clozapine and haloperidol were selected as atypical and typical antipsychotic models, respectively. Aripiprazole and lurasidone are atypical new drugs that have recently gained popularity in the treatment of schizophrenia, bipolar disorder, and depression.
[0235] Agonistaltic and anti-agonist activities were collected by simply applying the drug and measuring any apparent changes in reporter gene output. To measure antagonism, a mixture of endogenous agonists was developed that simultaneously activates all receptors in the library. The ability of each antipsychotic drug to antagonize these interactions was observed, detected as the inverse of the original interaction. Figure 4a) The conditions are described in detail in the supplementary material. All graphs of the interactions of all antipsychotic drugs with all receptors in terms of agonist or antagonist effects can be found in the supplementary material.
[0236] To validate the determination of antagonistic effects, the theoretical binding affinity from the antagonist interaction was calculated using the Cheng-Prusoff equation, which considers the EC50 of the agonist, the agonist concentration, and the IC50 of the antagonist. Reported binding affinities are typically measured using radioligand binding assays, and the affinity theoretically calculated using the Cheng-Prusoff method should be considered an approximation. Despite this caution, the vast majority of calculated binding affinities are within orders of magnitude of the expected values, thus validating both the method and the assay. Figure 12B ).
[0237] Most of the anticipated antipsychotic drug interactions were reproduced. The data were compared with reported data for these antipsychotic drugs. Figure 12C When an interaction is reported to interact with a human receptor with an affinity below 1000 nM, the interaction is considered “expected” (a general rule of thumb is that affinity values above 200 nM are physiologically irrelevant). An interaction is considered a “hit” when the expected pattern of action—agonist, antagonist, or inverse agonist—is reproduced in our assays. An interaction is considered a “miss” when activity is reported but not observed. Despite listing high affinity, many reported drug interactions were listed as “undetermined” in terms of the mode of action. This indicates that the mode of action is antagonist. The assays identified 53 / 62 (85%) “expected” hits. Of the nine “misses,” most were reported low-potency antagonist interactions or cases where the assay measured agonist activity but not antagonist activity.
[0238] A large number of (75) interactions detected were classified as “differential hits”; that is, interactions reported in our assays that differed from previously reported ones. These were the majority of cases where 1) the interaction was not reported in IUPHAR / Wikipedia but was identified in our assays, or 2) the interaction was reported only as an antagonist but reverse or partial agonist activity was observed as well as antagonist activity. Figures 12A-12C In most conventional antagonist assays, inverse agonists and partial agonists appear to possess antagonist activity.
[0239] While some of these could be explained by differences in assay type (i.e., different cell lines), most of these “difference hits” are likely underreported / incorrectly annotated in the IUPHARS / Wikipedia database. For example, there are few reports of lurasidone’s interaction with MAHDS. No activity against D2 receptors (DRD2, DRD3, DRD4) has been reported in the database. A closer examination reveals that activity has been reported here, but only against the rat homolog.
[0240] Receptor selectivity of antipsychotics: DRD2 vs. HTR2A Atypical antipsychotics have been shown to have a lower risk of extrapyramidal side effects (EPS) while treating psychosis. A key difference between them and typical antipsychotics lies in their higher potency in antagonizing HTR2A relative to DRD2. When comparing the relative IC50 of clozapine, lurasidone, and haloperidol for antagonizing HTR2A and DRD2, clozapine has the highest HTR2A to DRD2 potency ratio, while haloperidol has the lowest. Figures 13A-13B (Aripiprazole was not included in this analysis because it is generally considered a class III antipsychotic with partial agonist activity against both HTR2A and DRD2.) This is consistent with expectations, as clozapine is considered the atypical antipsychotic model with the highest clinical value in treating positive symptoms, while haloperidol is a typical antipsychotic model with a higher risk of EPS. These data suggest that multiplexing assays are sensitive to receptor selectivity trends for these compounds, which could potentially explain the differences in clinical outcomes.
[0241] Example 6. Screening for compounds interacting with MC4R The system described in this paper was used to screen for the interaction of 34 compounds with MC4R. The compounds tested included both small molecules and peptides. Constructs with reporter genes were prepared for four receptors (MC4R1, MC4R2, MC4R3, and MC4R4), and signal transduction readouts of the Gq and Gs pathways were performed. Each well contained cells with a construct containing each reporter gene. One compound was screened for each well. Results are presented in... Figure 14 Described in the text.
[0242] Each compound was analyzed for multiple GPCR receptors and multiple signal transduction pathways. Figure 15A The selectivity of each compound among different MC4R receptors was described. Figure 15B The bias observed on the Gq / Gs axes of MC4R is depicted. Different compounds cause signal conduction to be biased towards the Gs or Gq axes.
[0243] Example 7: High-resolution DMS used to identify retinitis pigmentosa mutations This embodiment describes an in vitro diagnostic assay to determine the class of a retinitis pigmentosa variant and its response potential to a specific rhodopsin corrector molecule. A gene expressing a variant RHO with a single missense mutation is fused with a gene expressing a transcription factor (TF) and integrated into the cell. Each cell has a single variant RHO gene and a unique barcode downstream of the response element (RE). The mutant RHO transcription factor fusion, properly folded and transported to the plasma membrane, encounters a plasma membrane-specific protease. The protease cleaves the mutant RHO transcription factor, and the transcription factor is released from the membrane. The transcription factor binds to the response element and allows the transcription variant-specific barcode to be read via next-generation sequencing.
[0244] The barcodes were separated and quantified using next-generation sequencing.
[0245] like Figure 16 As shown, the heatmap was generated through deep mutation screening using transport fractions for each missense mutation against wild-type RHO. Each individual amino acid in RHO was mutated to one of twenty amino acids. Twenty amino acid variations (including stop codons) were tested for all 348 amino acids in RHO. Each mutant RHO was identified using a unique barcode. Assays were performed as described above, and the barcode reads for each unique mutation were quantified and normalized against the wild type. Normalized transport was evaluated for each variant of RHO. The library was then treated with a rhodopsin corrector, and assays were performed to identify which mutations responded to a specific rhodopsin corrector.
[0246] It should be understood that the examples and embodiments described herein are for illustrative purposes only, and those skilled in the art will make various modifications or changes based on them, and such modifications or changes will be included within the spirit and scope of this application and the appended claims. For all purposes, all publications, patents and patent applications cited herein are hereby incorporated herein by reference in their entirety.
Claims
1. A system comprising multiple engineered cell lines in a partition, wherein the multiple engineered cell lines include a first engineered cell line and a second engineered cell line, wherein the first engineered cell line includes a first reporter gene construct and the second engineered cell line includes a second reporter gene construct, wherein, The first reporter gene construct and the second reporter gene construct are selected from: a) Operant coupling to the constitutive promoter of the reporter gene; b) Operant coupling to the first inducible promoter of the reporter gene; or c) Operant coupling to the second inducible promoter of the reporter gene; The first reporter gene construct is different from the second reporter gene construct; and the first reporter gene construct and the second reporter gene construct are independently readable.
2. The system of claim 1, wherein the plurality of engineered cell lines further comprises a third engineered cell line, wherein the third engineered cell line comprises a third reporter gene construct different from the first reporter gene construct and the second reporter gene construct, wherein the third reporter gene construct is selected from: a) Operant coupling to the constitutive promoter of the reporter gene; b) Operant coupling to the first inducible promoter of the reporter gene; or c) Operant coupling to the second inducible promoter of the reporter gene.
3. The system of claim 1 or 2, wherein one or more of the first engineered cell line, the second engineered cell line, the third engineered cell line, or combinations thereof further comprise a first heterologous polypeptide.
4. The system of claim 3, wherein one or more of the first engineered cell line, the second engineered cell line, the third engineered cell line, or combinations thereof further comprise a second heterologous polypeptide, wherein the second heterologous polypeptide comprises at least one amino acid change relative to the first heterologous polypeptide.
5. The system of claim 4, wherein the second heteropeptide comprises less than 10, less than 5, less than 3, or less than 2 amino acid changes relative to the first heteropeptide.
6. The system of claim 4, wherein the second heteropeptide comprises more than 10, more than 20, more than 50, more than 100, more than 500, or more than 1000 amino acid changes relative to the first heteropeptide.
7. The system of any one of claims 3 to 6, wherein the first heteropeptide, the second heteropeptide, or both are coupled to a transcription factor.
8. The system of claim 7, wherein the transcription factor comprises one or more of aGal4, PPR1, Lac9, zinc finger, or LexA DNA-binding domain.
9. The system of claim 7 or 8, wherein the transcription factor comprises one or more of aVP64, p65, RoTev, or Rta DNA activation domains.
10. The system of any one of claims 1 to 9, wherein the first reporter gene construct and the second reporter gene construct are independently readable.
11. The system of any one of claims 2 to 10, wherein the third reporter gene construct is independently readable as a reporter gene of the first reporter gene construct and / or the second reporter gene construct.
12. The system of any one of claims 3 to 11, wherein the first inducible promoter, the second inducible promoter, or both are configured to be activated by the first heteropeptide, a signal from the first heteropeptide, a transcription factor coupled to the first heteropeptide, or any combination thereof.
13. The system of any one of claims 1 to 12, wherein the first engineered cell line, the second engineered cell line, or the third engineered cell line comprises an additional reporter gene construct selected from: a) Operant coupling to the constitutive promoter of the reporter gene; b) Operant coupling to the first inducible promoter of the reporter gene; or c) Operant coupling to the second inducible promoter of the reporter gene. The additional reporter gene construct therein is different from the first reporter gene construct and the second reporter gene construct.
14. The system of any one of claims 1 to 13, wherein one or more of the first reporter gene construct, the second reporter gene construct, or the third reporter gene construct are integrated into the genome of the cell line.
15. The system of any one of claims 1 to 14, wherein one or more of the first engineered cell line, the second engineered cell line, or the third engineered cell line are eukaryotic cell lines.
16. The system of claim 15, wherein the eukaryotic cell line is a mammalian cell line.
17. The system of claim 16, wherein the mammalian cell line is a human cell line.
18. The system of any one of claims 1 to 17, wherein the reporter gene encodes a fluorescent protein or a luciferase protein.
19. The system of any one of claims 1 to 17, wherein the reporter gene encodes a barcode RNA sequence.
20. The system of any one of claims 1 to 19, wherein the reporter gene encodes a fluorescent protein and a barcode RNA sequence or a luciferase protein and a barcode RNA sequence.
21. The system of any one of claims 1 to 20, wherein the first heteropeptide or the second heteropeptide is a cell surface protein.
22. The system of claim 21, wherein the cell surface protein is a G protein-coupled receptor, a receptor tyrosine kinase, an ion channel, a cytokine receptor, a chemokine receptor, a growth factor receptor, or a cell adhesion molecule.
23. The system of any one of claims 21 to 22, wherein the cell surface protein is expressed by any one or more of the plurality of engineered cell lines.
24. The system of any one of claims 1 to 23, wherein the first heteropeptide or the second heteropeptide is an intracellular protein.
25. The system of claim 24, wherein the intracellular protein is an enzyme, an ER transporter, a nuclear transporter, an intracellular signaling protein, a chaperone molecule, or a transcription factor.
26. The system of any one of claims 1 to 25, wherein the plurality of engineered cell lines further comprises cells that contain a barcode sequence but do not express the barcode sequence.
27. The system of claim 26, wherein the plurality of engineered cell lines comprise mammalian cells.
28. The system of claim 27, wherein the mammalian cell is a human cell.
29. The system of any one of claims 1 to 28, wherein the partition is a hole in an n-hole plate.
30. The system of claim 29, wherein the n-well plate is a 96-well plate.
31. The system of claim 29, wherein the n-well plate is a 384-well plate.
32. The system of claim 29, wherein the n-well plate is a 1536-well plate.
33. The system of any one of claims 1 to 32, wherein the first reporter gene construct, the second reporter gene construct, or the third reporter gene construct comprises a constitutive promoter operatively coupled to a reporter gene.
34. The system of any one of claims 1 to 33, wherein the constitutive promoter is selected from the SV40 promoter, CMV promoter, Ef1A promoter, PGK1 promoter, Ubc promoter, β-actin promoter, CAG promoter, Ac5 promoter, polyhedrone protein promoter, TEF1 promoter, GDS promoter, CaMV355 promoter, Ubi promoter, or any combination thereof.
35. The system of any one of claims 1 to 34, wherein the first inducible promoter comprises the NFAT promoter, the CRE promoter, the p53 promoter, the ISRE promoter, the Gal4-UAS promoter, and the Lex A promoter.
36. The system of any one of claims 1 to 35, wherein the second inducible promoter is selected from the NFAT promoter, CRE promoter, p53 promoter, ISRE promoter, Gal4-UAS promoter, and Lex A promoter.
37. The system of any one of claims 1 to 36, wherein the first inducible promoter and the second inducible promoter are different promoters that mediate signal transduction through the same intracellular protein or cell surface protein.
38. A method for screening compounds that regulate the bioactivity of one or more of a plurality of engineered cell lines according to any one of the preceding claims, the method comprising contacting the plurality of engineered cell lines with a test agent and measuring the activity of the reporter gene construct of the first engineered cell line, the second engineered cell line, the third engineered cell line, or any combination thereof.
39. The method of claim 38, wherein a plurality of n test agents are contacted with the plurality of engineered cell lines that have been divided into at least n partitions.
40. The method of claim 38 or 39, wherein the plurality of engineered cell lines comprises at least 100, 1,000, or at least 10,000 different heterologous polypeptides.
41. The method of claims 39 to 40, wherein the bioassay is performed in a 96-well, 384-well, or 1536-well plate.
42. The method of any one of claims 38 to 41, wherein the first reporter gene construct and the second reporter gene construct are present in different engineered cells in the same wells of a 96-well, 384-well, or 1536-well plate.
43. The method of any one of claims 38 to 42, wherein the plurality of n test reagents are prepared by reacting a core fragment A containing a reactive functional group x with a plurality of initial test fragments (y-T1, y-T2, ... yT). n ) sufficient to form multiple n test reagents (A-T1, A-T2, ...AT) n The reaction is carried out under the following reaction conditions to prepare multiple separate reaction mixtures, where x is an amine, aldehyde, borate ester, imide, isothiocyanate, carboxylic acid, halide, hydroxyamidine, or thiourea; each test fragment contains a reactive functional group y and multiple initial test fractions (T1, T2, ... T). n One of ), where y is an amine, borate ester, imide, isothiocyanate, halide, hydroxyamide, thiourea, carboxylic acid, or aldehyde, wherein A-T1, A-T2, ... AT n Each of the components is prepared in a well of the test plate.
44. The method of claim 43, wherein x is an amine.
45. The method of claim 44, wherein the amine is a primary amine or a secondary amine.
46. The method of claim 43, wherein x is an aldehyde or a carboxylic acid.
47. The method of claim 43, wherein y is an amine.
48. The method of claim 47, wherein the amide is a primary amine or a secondary amine.
49. The method of claim 43, wherein y is a carboxylic acid or an aldehyde.
50. The method of any one of claims 43 to 49, wherein the reaction conditions include a base.
51. The method of any one of claims 43 to 49, wherein the reaction conditions comprise Lewis acids or Brønsted acids.
52. The method of any one of claims 43 to 49, wherein the reaction conditions include an amide coupling agent.
53. The method of any one of claims 43 to 49, wherein the reaction conditions include a palladium reagent.
54. The method of any one of claims 43 to 53, wherein the reaction conditions include room temperature.
55. The method of any one of claims 43 to 53, wherein the reaction conditions include a reaction temperature between about room temperature and about 80°C.
56. The method of any one of claims 43 to 55, wherein the reaction is carried out for about 1 to about 24 hours.
57. The method of claim 56, wherein the reaction is carried out for about 6 to about 18 hours.
58. The method of any one of claims 43 to 57, wherein the reaction is reversible or irreversible.
59. The method according to any one of claims 43 to 58, comprising amide coupling, reductive amination, or oxidative addition.
60. The method of any one of claims 43 to 58, comprising Buchwald or Suzuki coupling.
61. The method of any one of claims 43 to 60, wherein the mass of the core portion A is from about 150 Da to about 800 Da.
62. The method of any one of claims 43 to 61, wherein the mass of each test section (T1, T2, ... Tn) is about 80 Da to about 500 Da.
63. The method of any one of claims 43 to 62, wherein the mass of each test reagent (A-T1, A-T2, ... A-Tn) is less than about 1500 Da.
64. The method of claim 63, wherein the mass of each test reagent (A-T1, A-T2, ... A-Tn) is from about 350 Da to about 800 Da.
65. The method of any one of claims 43 to 64, wherein each reaction is carried out at the nanoscale.
66. The method of claim 65, wherein the nanoscale is carried out in a volume of 50 nL to 500 nL.
67. The method of any one of claims 43 to 66, wherein the test plate is a 96-well, 384-well, or 1536-well test plate.
68. The method of any one of claims 43 to 67, wherein each initial test portion (T1, T2, ... Tn) is different.
69. The method of any one of claims 43 to 68, wherein n is 2 to 100,000.
70. The method of claim 69, wherein n is 2 to 25,000.
71. The method of claim 69, wherein n is 2 to 2,500.
72. The method of claim 69, wherein n is 2 to 2,000.