Transcriptional relay system
By using a transcription relay system with binding sites for non-endogenous synthetic transcription factors, the problems of high background signal and high coefficient of variation in endogenous response element regulation methods are solved, improving the signal-to-noise ratio and screening accuracy, and enhancing the screening effect of small molecules or biological agonists.
Patent Information
- Application Number
- CN202080054299.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-28
- Filing Date
- 2020-05-27
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2040-05-27
AI Technical Summary
In the prior art, promoter methods that utilize endogenous response elements encoding reporter molecules are hampered by high-intensity background signals and high coefficients of variation, resulting in low signal-to-noise ratios and low absolute values of reporter signal activation.
By employing a transcription relay system with highly selective binding sites for non-endogenous synthetic transcription factors, and by using promoters regulated by response elements and synthetic transcription factors, we can reduce the level of biological variation, improve the signal-to-noise ratio of the reporter signal, and reduce background signal.
This technology reduces background signals and improves the signal-to-noise ratio in cell signaling pathways, thereby enhancing the accuracy and efficiency of screening small molecules or biological agonists or antagonists for these pathways.
Smart Images

Figure CN114585741B_ABST
Abstract
Description
[0001] Cross-referencing
[0002] This application claims the benefit of U.S. Provisional Application No. 62 / 853,637, filed May 28, 2019, which is incorporated herein by reference in its entirety. Summary of the Invention
[0003] This article describes nucleic acids, systems, and methods for probing cell signaling pathway responses, screening antagonists or agonists of cell signaling pathways, or discovering novel cell signaling pathways. Previously known methods in the art utilize promoters regulated by endogenous responsive elements proximal to nucleic acids encoding reporter molecules. These methods are plagued by high-intensity background signals from reporter molecules due to the “leakage” nature of endogenous responsive elements binding to promoters. Furthermore, these methods suffer from high coefficients of variation. Finally, such methods are also affected by the low absolute value of reporter activation, leading to low signal-to-noise ratios. The nucleic acids and systems disclosed herein reduce biological variability, improve the signal-to-noise ratio of the reporter signal, and reduce background signals by using non-endogenous synthetic transcription factors with high selectivity for binding sites to synthetic transcription factors. Therefore, transcription of the reporter molecule is not initiated by endogenous transcription factors, which helps reduce background signals and improve the signal-to-noise ratio of the reporter. These nucleic acids and systems can be used to screen small molecules or biological agonists or antagonists of signaling pathways, such as G protein-coupled receptors, receptor tyrosine kinases, ion channels, and nuclear receptors. In a broad aspect, the system comprises a promoter encoding: a) a response element located proximal to the 5' end of the synthetic transcription factor's reading frame; and b) a promoter element capable of being bound by the synthetic transcription factor, the promoter element being located proximal to the 5' end of the reporter gene's reading frame. In this system, the reporter gene may include a unique molecular identifier (UMI) to allow for multiplexing of reporter assays.
[0004] On one hand, this document describes a transcription relay system comprising: a transcription factor nucleic acid containing a promoter nucleotide sequence regulated by a response element and a nucleotide sequence encoding a synthetic transcription factor, wherein the promoter nucleotide sequence regulated by the response element is located 5' to the nucleotide sequence encoding the synthetic transcription factor; and a reporter nucleic acid containing a promoter nucleotide sequence of the synthetic transcription factor and a nucleotide sequence encoding a reporter, wherein the promoter nucleotide sequence of the synthetic transcription factor is located 5' to the nucleotide sequence encoding the reporter, and wherein the promoter nucleotide sequence of the synthetic transcription factor is capable of being bound by the synthetic transcription factor. In some embodiments, the promoter nucleotide sequence regulated by the response element comprises a cAMP response element nucleotide sequence, an NFAT transcription factor response element nucleotide sequence, a FOS promoter nucleotide sequence, or a serum response element nucleotide sequence. In some embodiments, the synthetic transcription factor comprises a DNA-binding domain from a first transcription factor and a transcription activation domain from a second transcription factor. In some embodiments, the DNA-binding domain is derived from Gal4, PPR1, Lac9, or LexA. In some embodiments, the DNA-binding domain comprises an amino acid sequence having at least about 90% identity with the sequence shown in SEQ ID NO:1. In some embodiments, the DNA-binding domain comprises an amino acid sequence having at least about 95% identity with the sequence shown in SEQ ID NO:1. In some embodiments, the DNA-binding domain comprises the amino acid sequence shown in SEQ ID NO:1. In some embodiments, the DNA-binding domain comprises a variant of the amino acid sequence of SEQ ID NO:1. In some embodiments, the transcriptional activation domain comprises VP64, p65, and Rta. In some embodiments, the transcriptional activation domain comprises an amino acid sequence having at least about 90% identity with the sequence shown in SEQ ID NO:14. In some embodiments, the transcriptional activation domain comprises an amino acid sequence having at least about 95% identity with the sequence shown in SEQ ID NO:14. In some embodiments, the transcriptional activation domain comprises the amino acid sequence shown in SEQ ID NO:14. In some embodiments, the transcriptional activation domain comprises a variant of the amino acid sequence of SEQ ID NO:14, wherein the sequence variant increases or decreases transcriptional activation. In some embodiments, the synthetic transcription factor comprises a variant of the amino acid sequence shown in SEQ ID NO:10. In some embodiments, the synthetic transcription factor comprises a polypeptide sequence that destabilizes the synthetic transcription factor. In some embodiments, the polypeptide sequence that destabilizes the synthetic transcription factor comprises a PEST or CL1 polypeptide sequence.In some embodiments, the synthetic transcription factor promoter nucleotide sequence comprises a nucleotide sequence capable of binding to Gal4, PPR1, Lac9, or LexA. In some embodiments, the reporter includes a fluorescent protein, luciferase protein, β-galactosidase, β-glucuronidase, chloramphenicol acetyltransferase, secretory placental alkaline phosphatase, or a unique molecular identifier. In some embodiments, the reporter includes a fluorescent protein, luciferase protein, β-galactosidase, β-glucuronidase, chloramphenicol acetyltransferase, or secretory placental alkaline phosphatase, and UMI. In some embodiments, the unique molecular identifier is specific to a test peptide encoded by the reporter nucleic acid. In some embodiments, the transcription factor nucleic acid comprises a nucleotide sequence located proximal to the promoter nucleotide sequence regulated by the response element, the nucleotide sequence being capable of binding to a transcriptional repressor. In some embodiments, the transcription factor nucleic acid comprises a nucleotide sequence located proximal to the promoter nucleotide sequence regulated by the response element, the nucleotide sequence extending the 5' untranslated region of the mRNA encoded by the nucleotide sequence encoding the synthetic transcription factor. In some embodiments, the 5' untranslated region of the mRNA encoded by the nucleotide sequence encoding the synthetic transcription factor comprises one or more sequences that reduce the translation of the synthetic transcription factor. In some embodiments, the transcription factor nucleic acid and the reporter nucleic acid are components of a single nucleic acid. In some embodiments, as described herein, the cell comprises the relay system. In some embodiments, the cell comprises a eukaryotic cell. In some embodiments, the cell comprises a mammalian cell. In some embodiments, the transcription factor nucleic acid, the reporter nucleic acid, or both the transcription factor nucleic acid and the reporter nucleic acid are integrated as a single copy into the cell genome. In some embodiments, as described herein, the cell population comprises the relay system. In some embodiments, the cell population comprises a eukaryotic cell population. In some embodiments, the cell population comprises a mammalian cell population. In some embodiments, the cell or cell population contains high basal reporter activity. In some embodiments, the cells or cell populations include those with a high baseline reporter activity at least about 30-fold higher than the background, where the background is the level of reporter activity observed against parental cells or cell lines that do not contain the reporter. In some embodiments, the cells or cell populations include those with a low coefficient of biotic variation for reporter activity. In some embodiments, the cells or cell populations include those with a low coefficient of biotic variation for reporter activity of less than about 0.5.
[0005] In some embodiments, as described herein, a method for detecting the effect of a test reagent on the activity of a promoter regulating a response element includes contacting a cell or cell population with the test substance. In some embodiments, the test reagent is a chemical. Attached Figure Description
[0006] Figure 1A A schematic diagram of the transcription relay system is shown, illustrating transcription factor nucleic acids (left) and reporter nucleic acids (right).
[0007] Figure 1B The nucleic acid sequence encoding a reporter was depicted, wherein the reporter contains a unique RNA sequence.
[0008] Figure 2 Reported outputs are shown for cells carrying a single integrated CRE-luciferase (gray) and cells carrying a single integrated UAS-luciferase accompanied by multiple copies of semi-randomly integrated CRE-Gal4-VPR (black).
[0009] Figure 3 It shows Figure 2 The coefficient of variation for each sample is depicted in the figure, which is the result of three repeated operations.
[0010] Figure 4 The effect of the destabilization sequence tag (degron tag) on the nucleotide sequence of the Gal4-VPR promoter on the fold induction of the transcription relay system is shown.
[0011] Figure 5 Cell libraries derived from isoclonal NFAT relay cell lines are shown. Cell lines were screened using positive control compounds to determine their ability to detect the NFAT relay reporter activity of Gq-coupled GPCRs. Receptor-compound combinations that produced a false discovery rate (FDR) below 0.001 or a maximum Q value above 3 were considered significant hits. In this screening, libraries cb29 and cb37 produced the most significant hits.
[0012] Figure 6 The variation and basic activity of isoclonal cell lines used to generate cell libraries are shown. Detailed Implementation
[0013] On one hand, this article describes a transcription relay system comprising: (a) a transcription factor nucleic acid containing a promoter nucleotide sequence regulated by a response element and a nucleotide sequence encoding a synthetic transcription factor, wherein the promoter nucleotide sequence regulated by the response element is located 5' to the nucleotide sequence encoding the synthetic transcription factor; and (b) a reporter nucleic acid containing a synthetic transcription factor promoter nucleotide sequence and a nucleotide sequence encoding a reporter, wherein the synthetic transcription factor promoter nucleotide sequence is located 5' to the nucleotide sequence encoding the reporter, and wherein the synthetic transcription factor promoter nucleotide sequence is capable of being bound by the synthetic transcription factor.
[0014] On the other hand, this article describes a method for determining the effect of a test substance on the activity of a promoter regulated by a response element, comprising: (a) contacting a cell with the test substance, the cell comprising (i) a transcription factor nucleic acid containing a promoter nucleotide sequence regulated by a response element and a nucleotide sequence encoding a synthetic transcription factor, wherein the promoter nucleotide sequence regulated by the response element is located 5' to the nucleotide sequence encoding the synthetic transcription factor; and (ii) a reporter nucleic acid containing a promoter nucleotide sequence regulated by a synthetic transcription factor and a nucleotide sequence encoding a reporter, wherein the promoter nucleotide sequence regulated by the synthetic transcription factor is located 5' to the nucleotide sequence encoding the reporter, and wherein the promoter nucleotide sequence regulated by the synthetic transcription factor is capable of being bound by the synthetic transcription factor; and (b) performing at least one assay measuring the transcription of the reporter.
[0015] In the following description, certain specific details are set forth to provide a thorough understanding of the various embodiments. However, those skilled in the art will understand that the embodiments provided can be practiced without these details. Unless the context otherwise requires, throughout this specification and the following claims, the word “comprising” and its variations, such as “including,” should be interpreted as open-ended and inclusive, i.e., “including but not limited to.” Unless the context otherwise expressly specifies, the singular forms “a,” “an,” and “the” as used in this specification and the appended claims include plural references. It should also be noted that unless the context otherwise expressly specifies, the term “or” generally takes on the meaning of “and / or.” Furthermore, the headings provided herein are for convenience only and do not explain the scope or meaning of the claimed embodiments.
[0016] As used in this article, the term “about” refers to a quantity that differs from the stated quantity by no more than 10%.
[0017] The terms “peptide” and “protein” are used interchangeably to refer to polymers of amino acid residues and are not limited to a minimum length. Peptides (including the provided polypeptide chain and other peptides, such as linkers and binding peptides) can include amino acid residues, including native and / or non-native amino acid residues. The term also includes post-expression modifications of the peptide, such as glycosylation, sialylation, acetylation, phosphorylation, etc. In some respects, peptides may contain modifications to native or natural sequences, provided the protein retains the desired activity. These modifications can be intentional (e.g., through site-directed mutagenesis) or accidental (e.g., through mutations in the host producing the protein or errors resulting from PCR amplification).
[0018] The percentage of sequence identity (%) relative to a reference polypeptide sequence is the percentage of amino acid residues in the candidate sequence that are identical to those in the reference polypeptide sequence after sequence alignment and the introduction of gaps (if necessary) to achieve the maximum percentage of sequence identity, without considering any conserved substitutions as part of the sequence identity. Alignments for determining the percentage of amino acid sequence identity can be performed in various known ways, such as using publicly available computer software like BLAST, BLAST-2, ALIGN, or Megalign (DNASTAR) software. Appropriate parameters for sequence alignment can be determined, including the algorithm required to achieve maximum alignment across the full length of the sequences being compared. However, for the purposes of this paper, the amino acid sequence identity value % is generated using the sequence alignment computer program ALIGN-2. The ALIGN-2 sequence alignment computer program was written by Genentech, Inc., and the source code has been submitted with the U.S. Copyright Office (Washington DC, 20559) under U.S. Copyright Registration No. TXU510087, along with user documentation. The ALIGN-2 program is publicly available from Genentech, Inc. (South San Francisco, Calif.) or can be compiled from source code. The ALIGN-2 program should be compiled for use on UNIX operating systems (including digital UNIX V4.0D). All sequence comparison parameters are set by the ALIGN-2 program and do not change.
[0019] When using ALIGN-2 for amino acid sequence comparison, the percentage of amino acid sequence identity of a given amino acid sequence A with respect to a given amino acid sequence B (or, as can be expressed as a given amino acid sequence A having / containing a certain percentage of amino acid sequence identity with respect to a given amino acid sequence B) is calculated as follows: 100 multiplied by the fraction X / Y, where X is the number of amino acid residues that ALIGN-2 scores as identical matches in the A and B alignments, and Y is the total number of amino acid residues in B. It should be understood that when the length of amino acid sequence A is not equal to the length of amino acid sequence B, the percentage of amino acid sequence identity of A with respect to B will not be equal to the percentage of amino acid sequence identity of B with respect to A. Unless otherwise specified, all amino acid sequence identity percentage values used herein were obtained using the ALIGN-2 computer program as described in the preceding paragraph.
[0020] In this document, when describing nucleic acid sequences relative to a reference sequence, the terms "identity," "sameness," or "percentage of identity" are determined using the formula described by Karlin and Altschul (with improvements in Proc. Natl. Acad. Sci. USA 87:2264-2268, 1990; Proc. Natl. Acad. Sci. USA 90:5873-5877, 1993). Such a formula is incorporated into the Basic Local Alignment Search Tool (BLAST) procedure of Altschul et al. (J. Mol. Biol. 215:403-410, 1990). As of the filing date of this application, the percentage of sequence identity can be determined using the latest version of BLAST.
[0021] The polypeptides of the systems described herein can be encoded by nucleic acids. A nucleic acid is a polynucleotide containing two or more nucleotide bases. In some embodiments, the nucleic acid is a component of a vector that can be used to transport the polynucleotide encoding the polypeptide into the cell. As used herein, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it is linked. One type of vector is a genome integration vector, or "integration vector," which can be integrated into the chromosomal DNA of a host cell. Another type of vector is an "attachment" vector, for example, a nucleic acid capable of extrachromosomal replication. A vector capable of directing the expression of a gene operatively linked to it is referred to herein as an "expression vector." Suitable vectors include plasmids, bacterial artificial chromosomes, yeast artificial chromosomes, viral vectors, etc. In expression vectors, regulatory elements such as promoters, enhancers, and polyadenylation signals used to control transcription can be derived from genes of mammals, microorganisms, viruses, or insects. Additionally, the ability to replicate in the host (usually conferred by the origin of replication) and selection genes that facilitate transformant recognition can be incorporated. Vectors derived from viruses such as lentiviruses, retroviruses, adenoviruses, and adeno-associated viruses can be used. Plasmid vectors can be linearized for integration into chromosomal locations. The vector may contain a sequence that guides site-specific integration into a defined set of restricted sites in the genome (e.g., AttP-AttB recombination). Additionally, the vector may contain a sequence derived from a transposon element used for integration.
[0022] As used herein, the term "transfection" or "transfected" refers to the intentional introduction of exogenous nucleic acids into cells using methods commonly used in the laboratory. Transfection can be achieved, for example, through lipid transfection, calcium phosphate precipitation, viral transduction, or electroporation. Transfection can be transient or stable.
[0023] As used herein, the term "transfection efficiency" refers to the extent or degree to which a cell population incorporates exogenous nucleic acids. Transfection efficiency can be measured as the percentage (%) of cells in a given population that have incorporated exogenous nucleic acids, relative to the total number of cells in the system. Transfection efficiency can be measured in both transiently and stably transfected cells.
[0024] As used herein, the term "bioactivated peptide" refers to a peptide expressed by cells that regulate gene expression. Bioactivated peptides can directly regulate gene expression through signal transduction in response to stimuli via one or more intermediate molecules or peptides, or through any other mechanism. Bioactivated peptides can be transmembrane peptides (e.g., receptors or channel proteins), intracellular peptides (e.g., signal transduction intermediates), extracellular peptides, or secreted peptides.
[0025] As used herein, “reporter activity” refers to the experimental reading of a reporter. For example, a luciferase reporter will exhibit a luminescent reading when incubated with a suitable substrate. Other reporters, such as fluorescent proteins, may not require a substrate and can be measured, for example, by a microscope or a fluorescence plate reader.
[0026] System Overview
[0027] The systems, nucleic acids, and methods described herein can be used to screen for the presence and / or activation levels of response element-binding promoters. The nucleic acids, systems, and methods described herein allow for activation of transcription at lower levels of background signal than conventional reporter systems. In some embodiments, the response element-binding promoter is activated at the end of a cell signaling cascade. In some embodiments, the presence of the response element-binding promoter can be measured before and after an external stimulus, such as a physical or chemical stimulus, or the presence of the response element-binding promoter can be compared to a control condition operated in parallel. The chemical stimulus can be an agonist or antagonist small molecule or biomolecule. In some embodiments, the system can be used for screening for drug discovery purposes. The system comprises at least a nucleic acid containing a response element-regulated promoter, a synthetic transcription factor promoter, a synthetic transcription factor, and a reporter. The response element-regulated promoter is located 5' to the synthetic transcription factor and, when the response element-binding promoter is present, activates transcription of the synthetic transcription factor. After translation, the synthetic transcription factor can subsequently bind to the synthetic transcription factor promoter, which is located 5' to the nucleic acid sequence encoding the reporter. Upon binding, the synthetic transcription factor promoter activates transcription of the nucleic acid sequence encoding the reporter. In some embodiments, the reporter is a polypeptide. In some embodiments, the reporter is a UMI. Other optional features of the system include a nucleotide sequence located proximal to the promoter nucleotide sequence regulated by the response element, which can be bound by a transcriptional repressor. In some embodiments, this nucleotide sequence proximal to the promoter nucleotide sequence regulated by the response element extends the 5' untranslated region of the mRNA encoded by the nucleotide sequence encoding a synthetic transcription factor. In some embodiments, the 5' untranslated region of the mRNA encoded by the nucleotide sequence encoding a synthetic transcription factor has one or more sequences that reduce the translation of the synthetic transcription factor.
[0028] Figure 1A The diagram illustrates a non-limiting embodiment of the invention. The left figure shows transcription factor nucleic acid 100. Transcription factor nucleic acid 100 has a promoter nucleic acid 102 regulated by a response element located at the 5' position of the nucleotide sequence encoding synthetic transcription factor 104. The right figure shows reporter nucleic acid 110, which contains a synthetic transcription factor promoter nucleotide sequence 112 located to the 5' side of the nucleotide sequence encoding reporter 114. In some embodiments, the transcription factor nucleic acid and the reporter nucleic acid are present on different nucleic acid molecules, such as different plasmids or viral vectors. In some embodiments, the transcription factor nucleic acid and the reporter nucleic acid are linear. In some embodiments, the transcription factor nucleic acid and the reporter nucleic acid are present on the same nucleic acid, which may be a plasmid, a viral vector, a linear, or any other conformation.
[0029] Figure 1BA non-limiting embodiment of the nucleotide sequence encoding a reporter peptide is shown. The nucleotide sequence encoding reporter peptide 114 comprises a nucleic acid sequence encoding reporter peptide 122 and a nucleic acid sequence encoding UMI 124. Sequence 124 is also referred to as a unique molecular identifier (UMI). The UMI can identify a specific bioactivating peptide that leads to activation of a promoter nucleic acid regulated by a response element at position 102. As a non-limiting example, the bioactivating peptide can contain a specific G protein-coupled receptor, of which hundreds are known. Thus, the UMI element allows for easy and rapid detection of signal transduction of a variety of different bioactivating peptides in a multiplexed manner. Furthermore, the provided relay system reduces background signal transduction via a promoter regulated by a response element. In any multiplexed screening of compounds that can activate bioactivating peptides, this can make quantification more accurate and reduce the number of false positive test compounds. In some embodiments, the nucleic acid sequence encoding the reporter peptide is absent. In some embodiments, the nucleic acid sequence encoding the UMI is absent. In some embodiments, the nucleic acid sequence encoding the UMI is located 5' to the side of the nucleic acid sequence encoding the reporter peptide. In some implementations, the nucleic acid sequence encoding the reporter polypeptide is located 5' to the side of the nucleic acid sequence encoding the UMI.
[0030] In some embodiments, the nucleic acid encoding the reporter encodes a reporter polypeptide. In some embodiments, the reporter polypeptide can be directly detected. In some embodiments, the reporter polypeptide generates a detectable signal based on the protein's enzymatic activity against the substrate. In some embodiments, the detection of the reporter polypeptide can be quantitatively performed. In some embodiments, the reporter polypeptide comprises a luciferase protein, β-galactosidase, β-glucuronidase, chloramphenicol acetyltransferase, secretory placental alkaline phosphatase, or a combination thereof. In some embodiments, wherein the reporter polypeptide is a luciferase protein, and non-limiting examples of substrates include firefly luciferin, latia luciferin, bacterial luciferin, coelenterate luciferin, dinoflagellate luciferin, vargulin, and 3-hydroxyhispidin.
[0031] In some implementations, the nucleic acid encoding the reporter encodes a UMI. The UMI comprises a short nucleotide sequence specific to the nucleic acid. The length of the UMI can be 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides. The UMI can be detected in any suitable manner that allows for the determination of the UMI sequence, such as by next-generation sequencing methods. Methods for detecting the UMI can be quantitative and include next-generation sequencing methods.
[0032] In some embodiments, this document describes methods for deploying a system for drug discovery, the system comprising nucleic acids encoding transcription factor nucleic acids and reporter nucleic acids. In some embodiments, the method includes contacting the nucleic acid with cells or cell populations under conditions sufficient to allow the nucleic acid to be internalized and expressed in cells (e.g., transfection); contacting the cells with physical or chemical stimuli; and determining activation of the reporter element by one or more assays. In some embodiments, the method includes contacting cells or cell populations containing nucleic acids encoding transcription factor nucleic acids and reporter nucleic acids; and determining activation of the reporter element by one or more assays.
[0033] Promoter regulated by response element
[0034] A response element is a short DNA sequence within the promoter region of a gene that binds to specific transcription factors and regulates gene transcription. Some response elements are specific to certain promoters. Some response elements can be bound by endogenous transcription factors. Multiple copies of the same response element can be located in different parts of the nucleotide sequence, responding to the same stimulus and activating different genes. Non-limiting examples of response elements that can be incorporated into the systems described herein include cAMP response elements (CREs), B recognition elements, AhR-, dioxin or biological heterologous substance response elements, HIF response elements, hormone response elements, serum response elements, retinoic acid response elements, peroxisome proliferator-hormone response elements, metal response elements, DNA damage response elements, IFN stimulation response elements, ROR response elements, glucocorticoid response elements, calcium response elements (CaRE1), antioxidant response elements, p53 response elements, thyroid hormone response elements, growth hormone response elements, sterol response elements, polycomb response elements, and vitamin D response elements.
[0035] A promoter nucleotide sequence regulated by a response element is a nucleic acid region containing one or more response elements that facilitate the recruitment of promoters and other molecules to regulate gene transcription. Cells contain numerous nucleotide sequences regulated by response elements that utilize endogenous proteins to regulate gene transcription. In cases where promoter nucleotide sequences regulated by endogenous response elements directly regulate reporter transcription, a high level of background signal is present due to the presence of the endogenous promoter. Systems that use non-endogenous transcription factors to regulate reporter transcription (where non-endogenous refers to the cells containing the system) have advantages over systems that use endogenous transcription factors to regulate reporter transcription. One advantage of such systems is the reduced background noise in reporter generation.
[0036] In some embodiments, the transcription relay system of the present invention comprises a transcription factor nucleic acid comprising a promoter nucleotide sequence regulated by a response element and a nucleotide sequence encoding a synthetic transcription factor, wherein the promoter nucleotide sequence regulated by the response element is located 5' to the nucleotide sequence encoding the synthetic transcription factor. The promoter nucleotide sequence regulated by the response element controls the expression of the synthetic transcription factor encoded by the synthetic transcription factor nucleotide sequence. In some embodiments, the promoter nucleotide sequence regulated by the response element comprises a cAMP response element nucleotide sequence, an NFAT transcription factor response element nucleotide sequence, a FOS promoter nucleotide sequence, a serum response element nucleotide sequence, or a combination thereof. In some embodiments, the promoter nucleotide sequence regulated by the response element comprises a cAMP response element nucleotide sequence. In some embodiments, the promoter nucleotide sequence regulated by the response element comprises an NFAT transcription factor response element nucleotide sequence. In some embodiments, the promoter nucleotide sequence regulated by the response element comprises a FOS promoter nucleotide sequence. In some embodiments, the promoter nucleotide sequence regulated by the response element comprises a serum response element nucleotide sequence. In some embodiments, the promoter nucleotide sequence regulated by the response element includes any combination of cAMP response element nucleotide sequences, NFAT transcription factor response element nucleotide sequences, FOS promoter nucleotide sequences, and / or serum response element nucleotide sequences.
[0037] In some embodiments, the promoter regulated by the response element can be bound by transcription factors. Non-limiting examples of common transcription factors include LexA, Gal4, VP16 (from herpes simplex virus), heat shock factor (HSF), NFAT, CREB, or combinations thereof. The system described herein is compatible with any transcription factor or any combination thereof that is commonly or potentially available in reported assays.
[0038] In some embodiments, the promoter regulated by the responsive element is bound by an endogenous transcription factor. Endogenous transcription factors are naturally occurring transcription factors in an organism, tissue, or cell. The presence of endogenous transcription factors depends on the presence of a transcription relay system therein. In some embodiments, the endogenous transcription factor promotes the transcription of synthetic transcription factors at a background rate.
[0039] In some embodiments, the transcription factor nucleic acid comprises a nucleotide sequence located proximal to the promoter nucleic acid sequence regulated by the response element, which can be bound by a transcriptional repressor. The transcriptional repressor inhibits transcription of the distal nucleotide sequence. Non-limiting examples of common transcriptional repressors include TetR, lac repressor, KRAB repressor, and combinations thereof. The system described herein is compatible with any repressor or combination thereof that is commonly or potentially available in reported assays.
[0040] In some embodiments, the transcription factor nucleic acid comprises a nucleotide sequence located proximal to the promoter nucleotide sequence regulated by the response element, the nucleotide sequence extending into the 5' untranslated region of the mRNA encoded by the nucleotide sequence encoding the synthetic transcription factor. In some embodiments, the 5' untranslated region of the mRNA encoded by the nucleotide sequence encoding the synthetic transcription factor comprises one or more sequences that reduce the translation of the synthetic transcription factor. In some embodiments, the one or more sequences that reduce the translation of the synthetic transcription factor comprise secondary structures that reduce the translation of the synthetic transcription factor. In some embodiments, the one or more sequences that reduce the translation of the synthetic transcription factor comprise sequences that affect the binding of RNA-binding proteins. In some embodiments, the one or more sequences that reduce the translation of the synthetic transcription factor comprise an upstream open reading frame.
[0041] Determination methods
[0042] The system described above can be utilized effectively using a variety of methods. It can be used to detect the activity of cell signaling pathways in steady state and in response to physical or chemical stimuli. When the reporter element contains a UMI sequence paired with a specific reporter element, the system can be deployed in multiplex assays.
[0043] In a non-limiting illustrative example, multiple cells are incubated in one well of a multi-well plate. Multiple cells are transfected with a reporter nucleic acid containing a promoter nucleotide sequence of a synthetic transcription factor and a nucleotide sequence encoding a reporter gene. The cells may already contain the transcription factor nucleic acid, which contains a promoter nucleotide sequence regulated by a response element and a nucleotide sequence encoding a synthetic transcription factor, or the cells may be transfected with the transcription factor nucleic acid. The transfected cells are then exposed to a chemical stimulus. After allowing sufficient time for reporter gene expression, cell lysates are harvested and the activation of the reporter gene is quantified. In this example, the presence of an increased reporter gene indicates that the chemical stimulus leads to enhanced activity of the transcription factor binding to the promoter regulated by the response element. In some embodiments, the activity of the transcription factor binding to the promoter regulated by the response element is enhanced following a cell signaling cascade.
[0044] In embodiments where the reporter gene comprises an enzyme that generates a detectable signal upon interaction with a substrate, the activation of the reporter gene can be quantified using standard assays known in the art. In embodiments where the reporter gene comprises a fluorescent molecule, the activation of the reporter gene can be measured by fluorescence microscopy or a fluorescence plate reader, and may not require cell lysis. The fluorescent molecule can be used to measure reporter activation in living cells. In embodiments where the reporter gene comprises a UMI, the mRNA is reverse transcribed, and the UMI is sequenced using next-generation sequencing technology.
[0045] In some embodiments, the assay is performed in a multi-well format such as 6, 12, 24, 48, 96, or 384 wells. In some embodiments, a different test chemical is provided to each well, or the test chemical is provided in two, three, or four-well configurations. The assay may also include one or more positive or negative control wells.
[0046] Synthetic transcription factors
[0047] Synthetic transcription factors are artificial proteins capable of targeting and regulating gene expression. Some synthetic transcription factors are chimeric proteins containing domains from multiple different genes. In some embodiments, a synthetic transcription factor contains a DNA-binding domain from one gene and a transcriptional regulatory domain from another gene.
[0048] In the methods, nucleic acids, and systems described herein, the transcriptional activation polypeptide is encoded on a transcription factor nucleic acid. In some embodiments, the transcriptional activation polypeptide is a synthetic transcription factor. In some embodiments, the synthetic transcription factor is a chimeric protein. In some embodiments, the synthetic transcription factor includes a DNA-binding domain from a first transcription factor. In some embodiments, the synthetic transcription factor includes a transcriptional activation domain from a second transcription factor. In some embodiments, the first transcription factor is different from the second transcription factor.
[0049] In some embodiments, the synthetic transcription factor exhibits higher specificity for the promoter nucleotide sequence of a synthetic transcription factor compared to any endogenous transcription factor. In some embodiments, the synthetic transcription factor binds to the promoter nucleotide sequence of a synthetic transcription factor that cannot be bound by an endogenous promoter. In some embodiments, the synthetic transcription factor results in less background reporter generation compared to the use of an endogenous transcription factor.
[0050] In some embodiments, the DNA-binding domain is non-endogenous to cells containing the transcription relay system of the present invention. In some embodiments, the DNA-binding domain from the first transcription factor is derived from Gal4, PPR1, LexA, Lac9, or a combination thereof. In some embodiments, the DNA-binding domain comprises the following amino acid sequence:
[0051] MKLLSSIEQACDICRLKKLKCSKEKPKCAKCLKNNWECRYSPKTKRSPLTRAHLTEVESRLERLEQLFLLIFPREDLDMILKMDSLQDIKALLTGLFVQDNVNKDAVTDRLASVETDMPLTLRQHRISATSSSEESSNKGQRQLTVS, SEQ ID NO:1. In some embodiments, the DNA-binding domain comprises the following amino acid sequence:
[0052] MKKKNSKKSNRTDSKRGDSNGSKSRTACKRCRKKKCDSCKRCAKVCVSDATGKDVRSYVDRAVMMRVKYGVDTKRGNATSDDDKKYSSVSS, SEQ ID NO:2. In some embodiments, the DNA-binding domain comprises the following amino acid sequence:
[0053] MKSRTACKRCRLKKIKCDQEFPSCKRCAKLEVPCYSPKTKRSPLTRAHLTEVESRLERLEQLFLLIFPREDLDMILKMDSLQDIKALLTGLFVQDNVNKDAVTDRLASVETDMPLTLRQHRISATSSSEESSNKGQRQLTVS, SEQ ID NO:3. In some embodiments, the DNA-binding domain comprises the following amino acid sequence:
[0054] MKSRTACKRCRLKKIKCDQEFPSCKRCAKLEVPCVSSPKTKRSPLTRAHLTEVESRLERLEQLFLLIFPREDLDMILKMDSLQDIKALLTGLFVQDNVNKDAVTDRLASVETDMPLTLRQHRISATSSSEESSNKGQRQLTVS, SEQ ID NO:4. In some embodiments, the DNA-binding domain comprises the following amino acid sequence:
[0055] MNKKSSEVMHQACDACRKKKWKCSKTVPTCTNCLKYNLDCVYSPQVVRTPLTRAHLTEMENRVAELEQFLKELFPVWDIDRLLQQKDTYRIRELLTMGSTNTVPGLASNNIDSSLEQPVAFGTAQPAQSLSTDPAVQSQAYPMQPV, SEQ ID NO:5. In some embodiments, the DNA-binding domain comprises the following amino acid sequence:
[0056] MNKKSSEVMHQACVECRQQKSKCDAHERAPEPCTKCAKKNVPCIVYSPQVVRTPLTRAHLTEMENRVAELEQFLKELFPVWDIDRLLQQKDTYRIRELLTMGSTNTVPGLASNNIDSSLEQPVAFGTAQPAQSLSTDPAVQSQAYPMQPV, SEQ ID NO:6. In some embodiments, the DNA-binding domain comprises the following amino acid sequence:
[0057] MNKKSSEVMHQACKRCRLKKIKCDQEFPSCKRCLKYNLDCVYSPQVVRTPLTRAHLTEMENRVAELEQFLKELFPVWDIDRLLQQKDTYRIRELLTMGSTNTVPGLASNNIDSSLEQPVAFGTAQPAQSLSTDPAVQSQAYPMQPV, SEQ ID NO:7. In some embodiments, the DNA-binding domain comprises the following amino acid sequence:
[0058] SEQ ID NO:8.
[0059] In some embodiments, the DNA binding domain comprises an amino acid sequence variant of SEQ ID NO:1. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is R15W, K23P, K23T, K23W, K23M, K23N, F68R, F68Q, L69P, L70P, Q9E, Q9A, Q9N, R15K, R15A, R15M, K18R, K18A, K18M, K23R, K23A, K23M, or a combination thereof. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is R15W. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is K23P. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is K23T. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is K23W. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is K23M. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is K23N. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is F68R. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is F68Q. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is L69P. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is L70P. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is Q9E. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is Q9A. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is Q9N. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is R15K. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is R15A. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is R15M. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is K18R. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is K18A. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is K18M. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is K23R. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is K23A. In some embodiments, the amino acid sequence variant of SEQ ID NO:1 is K23M.
[0060] In some embodiments, the transcriptional activation domain from the second transcription factor is derived from VP64, p65, and Rta, or combinations thereof. In some embodiments, the transcriptional activation domain comprises the following amino acid sequence:
[0061] , SEQ ID NO:14.
[0062] In some embodiments, the nucleic acid described herein encodes a transcription factor having a VPR amino acid sequence that has at least 90%, 95%, 97%, 98%, 99%, or 100% identity with the amino acid sequence shown in SEQ ID NO:14. In some embodiments, the nucleic acid described herein encodes a transcription factor having a VPR amino acid sequence that has at least 90% identity with the amino acid sequence shown in SEQ ID NO:14. In some embodiments, the nucleic acid described herein encodes a transcription factor having a VPR amino acid sequence that has at least 95% identity with the amino acid sequence shown in SEQ ID NO:14. In some embodiments, the nucleic acid described herein encodes a transcription factor having a VPR amino acid sequence that has at least 97% identity with the amino acid sequence shown in SEQ ID NO:14. In some embodiments, the nucleic acid described herein encodes a transcription factor having a VPR amino acid sequence that has at least 98% identity with the amino acid sequence shown in SEQ ID NO:14. In some embodiments, the nucleic acid described herein encodes a transcription factor having a VPR amino acid sequence that is at least 99% identical to the amino acid sequence shown in SEQ ID NO:10. In some embodiments, the nucleic acid described herein encodes a transcription factor having a VPR amino acid sequence that is 100% identical to the amino acid sequence shown in SEQ ID NO:14.
[0063] In some embodiments, the transcriptional activation domain on the synthetic transcription factor contains amino acid sequence variants that increase or decrease transcriptional activation. In some embodiments, the transcriptional activation domain containing amino acid sequence variants that increase or decrease transcriptional activation is a sequence variant of SEQ ID NO:14.
[0064] In some embodiments, the synthetic transcription factor encoded by the nucleic acid sequence of the transcription factor nucleic acid includes a polypeptide sequence, also known as a "degradation determinant," that destabilizes the synthetic transcription factor. In some embodiments, the polypeptide sequence that destabilizes the transcription factor includes a PEST polypeptide sequence. A PEST polypeptide sequence is a polypeptide sequence containing multiple amino acids, wherein the polypeptide sequence is rich in proline, glutamic acid, serine, and / or threonine. In some embodiments, the polypeptide sequence that destabilizes the transcription factor includes a CL1 polypeptide sequence. The CL1 polypeptide sequence can act as a degradation signal, resulting in a shortened half-life of the resulting synthetic transcription factor. In some embodiments, the polypeptide sequence that destabilizes the synthetic transcription factor helps reduce the background signal of the reporter.
[0065] In some embodiments, the synthetic transcription factor comprises a GAL4-VP16 chimeric transcription factor. In some embodiments, the transcription factor comprises a GAL4-VPR chimeric transcription factor. The sequence of the Gal4-VPR chimeric transcription factor is given below:
[0066] SEQ ID NO:10. In some embodiments, the nucleic acid encoding amino acid sequence described herein has at least 90%, 95%, 97%, 98%, 99%, or 100% identity with the amino acid sequence shown in SEQ ID NO:10. In some embodiments, the nucleic acid encoding amino acid sequence described herein has at least 90% identity with the amino acid sequence shown in SEQ ID NO:10. In some embodiments, the nucleic acid encoding amino acid sequence described herein has at least 95% identity with the amino acid sequence shown in SEQ ID NO:10. In some embodiments, the nucleic acid encoding amino acid sequence described herein has at least 97% identity with the amino acid sequence shown in SEQ ID NO:10.In some embodiments, the nucleic acid-encoding amino acid sequence described herein has at least 98% identity with the amino acid sequence shown in SEQ ID NO:10. In some embodiments, the nucleic acid-encoding amino acid sequence described herein has at least 99% identity with the amino acid sequence shown in SEQ ID NO:10. In some embodiments, the nucleic acid-encoding amino acid sequence described herein has 100% identity with the amino acid sequence shown in SEQ ID NO:10.
[0067] In some embodiments, the synthetic transcription factor comprises a Gal4 DNA-binding domain given by the amino acid sequence listed in SEQ ID NO:1. In some embodiments, the synthetic transcription factor comprises a DNA-binding domain having at least 90%, 95%, 97%, 98%, 99%, or 100% identity with the amino acid sequence shown in SEQ ID NO:1. In some embodiments, the synthetic transcription factor comprises a DNA-binding domain having at least 90% identity with the amino acid sequence shown in SEQ ID NO:1. In some embodiments, the synthetic transcription factor comprises a DNA-binding domain having at least 95% identity with the amino acid sequence shown in SEQ ID NO:1. In some embodiments, the synthetic transcription factor comprises a DNA-binding domain having at least 97% identity with the amino acid sequence shown in SEQ ID NO:1. In some embodiments, the synthetic transcription factor comprises a DNA-binding domain having at least 98% identity with the amino acid sequence shown in SEQ ID NO:1. In some embodiments, the synthetic transcription factor comprises a DNA-binding domain having at least 99% identity with the amino acid sequence shown in SEQ ID NO:1. In some embodiments, the synthetic transcription factor comprises a DNA-binding domain having an amino acid sequence that is 100% identical to the amino acid sequence shown in SEQ ID NO:1.
[0068] In some embodiments, the synthetic transcription factor comprises a transcriptional activation domain from VP64, given by the amino acid sequence listed below:
[0069] RAGKPIPNPLLGLDSTDALDDFDLDMLGSDALDDFDLDMLGSDALDDFDLDMLGSDALDDFDLDMLGSPKKKRKV, SEQ ID NO:11. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having at least 90%, 95%, 97%, 98%, 99%, or 100% identity with the amino acid sequence shown in SEQ ID NO:11. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having at least 90% identity with the amino acid sequence shown in SEQ ID NO:11. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having at least 95% identity with the amino acid sequence shown in SEQ ID NO:11. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having at least 97% identity with the amino acid sequence shown in SEQ ID NO:11. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having at least 98% identity with the amino acid sequence shown in SEQ ID NO:11. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having at least 99% identity with the amino acid sequence shown in SEQ ID NO:11. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having 100% identity with the amino acid sequence shown in SEQ ID NO:11.
[0070] In some embodiments, the synthetic transcription factor comprises a transcriptional activation domain from p65, given by the amino acid sequence listed below:
[0071] QYLPDTDDRHRIEEKRKRTYETFKSIMKKSPFSGPTDPRPPPRRIAVPSRSSASVPKPAPQPYPFTSSLSTINYDEFPTMVFPSGQISQASALAPAPPQVLPQAPAPAPAPAMVSALAQAPAPVPVLAPGPPQAVAPPAPKPTQAGEGTLSEALLQLQFDDEDLGALLGNSTDPAVFTDLASVDNSEFQQLLNQGIPVAPHTTEPMLMEYPEAITRLVTGAQRPPDPAPAPLGAPGLPNGLLSGDEDFSSIADMDFSALLSQISS, SEQ ID NO:12. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having at least 90%, 95%, 97%, 98%, 99%, or 100% identity with the amino acid sequence shown in SEQ ID NO:12. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having at least 90% identity with the amino acid sequence shown in SEQ ID NO:12. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having at least 95% identity with the amino acid sequence shown in SEQ ID NO:12. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having at least 97% identity with the amino acid sequence shown in SEQ ID NO:12. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having at least 98% identity with the amino acid sequence shown in SEQ ID NO:12. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having at least 99% identity with the amino acid sequence shown in SEQ ID NO:12. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having 100% identity with the amino acid sequence shown in SEQ ID NO:12.
[0072] In some embodiments, the synthetic transcription factor comprises a transcriptional activation domain from Rta, given by the amino acid sequence listed below:
[0073] RDSREGMFLPKPEAGSAISDVFEGREVCQPKRIRPFHPPGSPWANRPLPASLAPTPTGPVHEPVGSLTPAPVPQPLDPAPAVTPEASHLLEDPDEETSQAVKALREMADTVIPQKEEAAICGQMDLSHPPPRGHLDELTTTLESMTEDLNLDSPLTPELNEILDTFLNDECLLHAMHISTGLSIFDTSLF, SEQ ID NO:13. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having at least 90%, 95%, 97%, 98%, 99%, or 100% identity with the amino acid sequence shown in SEQ ID NO:13. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having at least 90% identity with the amino acid sequence shown in SEQ ID NO:13. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having at least 95% identity with the amino acid sequence shown in SEQ ID NO:13. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having at least 97% identity with the amino acid sequence shown in SEQ ID NO:13. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having at least 98% identity with the amino acid sequence shown in SEQ ID NO:13. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having at least 99% identity with the amino acid sequence shown in SEQ ID NO:13. In some embodiments, the synthetic transcription factor comprises a transcription activation domain having 100% identity with the amino acid sequence shown in SEQ ID NO:13.
[0074] Synthetic transcription factor promoter nucleotide sequence
[0075] A synthetic transcription factor promoter nucleotide sequence is a nucleic acid sequence that can be bound by a synthetic transcription factor. In some embodiments, the synthetic transcription factor nucleotide sequence is not bound by endogenous transcription factors. The synthetic transcription factor promoter nucleotide sequence facilitates the recruitment of the synthetic transcription factor to activate the transcription of a reporter molecule. The reporter molecule is encoded on a nucleic acid located on the 3' side of the synthetic transcription factor promoter nucleotide sequence.
[0076] In the methods, nucleic acids, and systems described herein, the synthetic transcription factor promoter nucleotide sequence is encoded on a reporter nucleic acid. The synthetic transcription factor promoter nucleotide sequence is capable of being bound by the synthetic transcription factor encoded on the transcription factor nucleic acid. The synthetic transcription factor promoter nucleotide sequence is located 5' to the side of the nucleotide sequence encoding the reporter. In some embodiments, the synthetic transcription factor promoter nucleotide sequence is not bound by endogenous transcription factors. In some embodiments, the synthetic transcription factor exhibits high specificity for the synthetic transcription factor promoter nucleotide sequence.
[0077] In some embodiments, the synthetic transcription factor promoter nucleotide sequence is capable of being bound by Gal4, PPR1, Lac9, or LexA. In some embodiments, the synthetic transcription factor is capable of being bound by a polypeptide comprising the amino acid sequence shown in SEQ ID NO:1.
[0078] In some embodiments, the synthetic transcription factor promoter nucleotide sequence is capable of being bound by amino acid sequence variants of Gal4, PPR1, Lac9, or LexA. In some embodiments, the synthetic transcription factor promoter nucleotide sequence is capable of being bound by an amino acid sequence variant of SEQ ID NO:1.
[0079] Reporting elements
[0080] The reporter nucleic acid contains at least a regulatory element capable of being bound by a synthetic transcription factor and a nucleotide sequence encoding a reporter. The nucleotide sequence encoding the reporter is downstream of the regulatory element capable of being bound by the synthetic transcription factor. The synthetic transcription factor regulates the expression of the reporter.
[0081] In some embodiments, the nucleotide sequence encoding the reporter protein comprises a reporter gene. In some embodiments, the reporter gene encodes a reporter protein selected from fluorescent proteins, luciferase proteins, β-galactosidase, β-glucuronidase, chloramphenicol acetyltransferase, and secretory placental alkaline phosphatase. Specific enzymatic activities of these reporter proteins can be measured, or, in the case of a fluorescent reporter protein, fluorescence emission can be measured. In some embodiments, the fluorescent protein includes green fluorescent protein (GFP), red fluorescent protein (RFP), yellow fluorescent protein (YFP), or cyan fluorescent protein (CFP).
[0082] In some embodiments, the nucleotide sequence encoding the reporter gene includes a nucleotide sequence encoding a unique sequence identifier (UMI). In some embodiments, the UMI is specific to a test peptide encoded by the reporter nucleic acid. Generally, the UMI is between 8 and 20 nucleotides in length, but may be longer. In some embodiments, the UMI is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleotides in length. In some embodiments, the UMI is 8 nucleotides in length. In some embodiments, the UMI is 9 nucleotides in length. In some embodiments, the UMI is 10 nucleotides in length. In some embodiments, the UMI is 11 nucleotides in length. In some embodiments, the UMI is 12 nucleotides in length. In some embodiments, the UMI is 13 nucleotides in length. In some embodiments, the UMI is 14 nucleotides in length. In some embodiments, the UMI is 15 nucleotides in length. In some embodiments, the UMI is 16 nucleotides in length. In some embodiments, the UMI is 17 nucleotides long. In some embodiments, the UMI is 18 nucleotides long. In some embodiments, the UMI is 19 nucleotides long. In some embodiments, the UMI is 20 nucleotides long. In some embodiments, the UMI is more than 20 nucleotides long.
[0083] The system described herein can utilize many different regulatory sequences that control reporter gene activation by binding to synthetic transcription factors. Regulatory sequences are sequences that can be bound by synthetic transcription factor peptides. Typically, the regulatory sequence is configured such that it is located on the 5' side of the UMI, the reporter gene, or both. In some embodiments, the regulatory sequence comprises Gal4, PPR1-, or LexA-UAS, which can be bound by synthetic transcription factors.
[0084] In some embodiments, the reporter nucleus includes a fluorescent protein, a luciferase protein, β-galactosidase, β-glucuronidase, chloramphenicol acetyltransferase, or secretory placental alkaline phosphatase, and a UMI. In some embodiments, the UMI is encoded on a reporter nucleic acid located on the 5' side of the fluorescent protein, luciferase protein, β-galactosidase, β-glucuronidase, chloramphenicol acetyltransferase, or secretory placental alkaline phosphatase. In some embodiments, the nucleotide sequence encoding the fluorescent protein, luciferase protein, β-galactosidase, β-glucuronidase, chloramphenicol acetyltransferase, or secretory placental alkaline phosphatase is located on the 5' side of the UMI.
[0085] UMIs allow for multiplexing of different transcriptional relay systems in the same assay because the transcription of the UMI will indicate the association of a specific relay system with a reporter. UMIs can be of any length that allows for sufficient diversity to permit multiplexing of different transcriptional relay systems in the same assay. The length should be sufficient to distinguish at least 100, 500, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, or 10,000 transcriptional relay targets. In some embodiments, the different transcriptional relay systems may be present in different cells. In some embodiments, the different transcriptional relay systems may be present in the same cell.
[0086] The reporting element may also include a 5' UTR, a 3' UTR, or both. The UTR may be heterogeneous from the reporting element.
[0087] Report Sub-Activation
[0088] Activation of reporter molecules can be determined by detecting luciferase, β-galactosidase, β-glucuronidase, chloramphenicol acetyltransferase, and secretory placental alkaline phosphatase proteins using standard assays. Typically, these are enzyme assays where the protein's enzymatic activity on its substrate produces a detectable signal. For example, luciferase expression can be measured by a photometer in the presence of the luciferase substrate. Fluorescent reporter molecules do not require a substrate and their signals can be measured using fluorescence microscopy or a fluorescence plate reader. Fluorescent reporter molecules are particularly useful for measuring reporter activation in living cells.
[0089] In embodiments where the reporter molecule contains a unique RNA sequence, the activation of the reporter can be measured in any suitable manner, provided that the manner allows for sequencing of the unique RNA sequence, preferably in a multiplexed manner. Such methods include high-throughput sequencing methods that can generate information on at least about 100,000, 1,000,000, 10,000,000, or 100,000,000 DNA or RNA bases within 24 hours. In some embodiments, next-generation sequencing technologies are used to determine the sequence of the unique RNA sequence. Next-generation sequencing includes many types of sequencing, such as pyrosequencing, sequencing by synthesis, single-molecule sequencing, second-generation sequencing, nanopore sequencing, ligation sequencing, or hybridization sequencing. Next-generation sequencing platforms include those commercially available from Illumina (RNA-Seq) and Helicos (digital gene expression or “DGE”).Next-generation sequencing methods include, but are not limited to, those commercialized by the following companies: 1) 454 / Roche Lifesciences, including but not limited to the methods and apparatus described in Margulies et al., Nature (2005) 437:376-380 (2005); and U.S. Patent Nos. 7,244,559; 7,335,762; 7,211,390; 7,244,567; 7,264,929; 7,323,305; 2) Helicos Biosciences Corporation (Cambridge, MA), such as U.S. Application Serial No. 11 / 167046 and U.S. Patent Nos. 7,501,245; 7,491,498; 7,276,720; and U.S. Patent Application Publication Nos. US20090061439; US20080087826; US20060286566; US2006002471. 1) As described in US20060024678; US20080213770; and US20080103058; 3) Applied Biosystems (e.g., SOLiD sequencing); 4) Dover Systems (e.g., Polonator G.007 sequencing); 5) Illumina, Inc., as described in US Patent Nos. 5,750,341; 6,306,597; and 5,969,119; and 6) Pacific Biosciences, as described in U.S. Patent Nos. 7,462,452; 7,476,504; 7,405,281; 7,170,050; 7,462,468; 7,476,503; 7,315,019; 7,302,146; 7,313,308; and U.S. Application Publications Nos. US20090029385; US20090068655; US20090024331; and US20080206764. Such methods and apparatus are provided herein by way of example and are not intended to be limiting.
[0090] markers
[0091] In some embodiments, the nucleic acid described herein further comprises one or more additional genes encoding a selectable or labeled polypeptide. In some embodiments, the nucleic acid described herein further comprises one or more additional genes encoding a polypeptide that confers antibiotic resistance to transfected cells. For example, the nucleic acid may contain a selectable marker, such as an antibiotic resistance gene conferring neomycin / G418 resistance, puromycin resistance, bleomycin resistance, or blast fungicide resistance. In some embodiments, the nucleic acid described herein further comprises one or more additional genes encoding a polypeptide that includes an epitope tag expressed on the cell surface. This enables affinity purification or cell sorting to collect cells transfected with said nucleic acid. In some embodiments, the epitope tag includes a c-Myc tag, a hemagglutinin (HA) tag, a histidine tag, a V5 tag, or a FLAG tag. In some embodiments, the nucleic acid described herein further comprises one or more additional promoterless genes encoding a fluorescent polypeptide. Such genes are useful when transfection is intended to induce integration and target a specific site or landing pad. In these cases, the “landing region” in the cell genome contains a promoter that complements the absence of a promoter in a promoterless gene and leads to expression of the promoterless gene only when integrated into the intended genomic location. Cells with correct integration can be selected by flow cytometry and cell sorting. This type of marker also ensures that only a single copy of the intended nucleic acid is integrated into the genome and helps avoid ectopic overexpression. In some embodiments, the nucleic acid encoding the decoy peptide comprises: a gene encoding a peptide that confers antibiotic resistance to transfected cells; a gene encoding a peptide including an epitope tag expressed on the cell surface; or a promoterless gene encoding a fluorescent peptide.
[0092] cell
[0093] Cells suitable for use in the methods described herein are typically transgenic cells that can be readily modified with exogenous nucleic acids encoding synthetic transcription factors and reporter elements. Systematic nucleic acids encoding synthetic transcription factors and reporter elements can be transfected or transduced into suitable cell lines using methods known in the art, such as calcium phosphate transfection, liposome-mediated transfection (e.g., or HD), electroporation, or viral transduction. Cells can also be grown into confluent or nearly confluent populations of the same type in suitable tissue culture containers.
[0094] In some embodiments, the cells used contain nucleic acids encoding synthetic transcription factors, nucleic acids containing reporter elements, or stable integration of both. Stable cell lines can be prepared using random integration of linearized plasmids, directed integration of viruses or transposons, or directed integration (e.g., using site-specific recombination between AttP and AttB sites). In some embodiments, either of these nucleic acids is encoded at a safe landing site, such as the AAVS1 site.
[0095] In some embodiments, the cells or cell populations used in the system are eukaryotic cells. In some embodiments, the cells or cell populations are mammalian cells. In some embodiments, the cells or cell populations are human cells. In some implementations, the cells or cell populations are SH-SY5Y, human neuroblastoma; Hep G2, Caucasian hepatocellular carcinoma; 293 (also known as HEK 293), human embryonic kidney; RAW 264.7, mouse mononuclear macrophages; HeLa, human cervical epithelioid carcinoma; MRC-5 (PD19), human fetal lung; A2780, human ovarian cancer; CACO-2, Caucasian colonic adenocarcinoma; THP 1, human monocytic leukemia; A549, Caucasian lung cancer; MRC-5 (PD 30), human fetal lung; MCF7, Caucasian breast cancer; SNL 76 / 7, mouse SIM strain embryonic fibroblasts; C2C12, mouse C3H myoblasts; Jurkat E6.1, human leukemia T-cell lymphoblasts; U937, Caucasian histiocytic lymphoma; L929, mouse C3H / An connective tissue; 3T3 L1, mouse embryo; HL60, promyelocytic leukemia in Caucasians; PC-12, pheochromocytoma of the adrenal gland in rats; HT29, colonic adenocarcinoma in Caucasians; OE33, esophageal cancer in Caucasians; OE19, esophageal cancer in Caucasians; NIH 3T3, Swiss mouse NIH embryo; MDA-MB-231, breast cancer in Caucasians; K562, chronic myeloid leukemia in Caucasians; U-87MG, glioblastoma-astrocytoma in humans; MRC-5 (PD 25), fetal lung in humans; A2780cis, ovarian cancer in humans; B9, mouse B-cell hybridoma; CHO-K1, ovary of a Chinese hamster; MDCK, kidney in a Cocker Spaniel; 1321N1, astrocytoma of the human brain; A431, squamous cell carcinoma in humans; ATDC5, AT805 derivative of teratoma of mouse 129; RCC4 PLUSVECTOR ALONE, RCC4 renal cell carcinoma cell line stably transfected with the neomycin-resistant empty expression vector pcDNA3; HUVEC (S200-05n), pre-selected human umbilical vein endothelial cells (HUVEC); neonate; Vero, African green monkey kidney; RCC4PLUS VHL, RCC4 renal cell carcinoma cell line stably transfected with pcDNA3-VHL; Fao, rat liver cancer; J774A.1, mouse BALB / c mononuclear macrophages; MC3T3-E1, mouse C57BL / 6 cranial tectum; J774.2. Mouse BALB / c mononuclear macrophages; PNT1A, human normal post-pubertal prostate, immortalized with SV40; U-2OS, human osteosarcoma; HCT 116, human colon cancer; MA104, African green monkey kidney; BEAS-2B, human normal bronchial epithelial cells; NB2-11, rat lymphoma; BHK 21 (clone 13), Syrian hamster kidney; NS0, mouse myeloma; Neuro 2a, mouse albino neuroblastoma; SP2 / 0-Ag14, mouse x mouse myeloma, nonproductive; T47D, human breast tumor; 1301, human T-cell leukemia; MDCK-II, Cocker Spaniel kidney; PNT2, human normal prostate, immortalized with SV40; PC-3, Caucasian prostate cancer; TF1, human erythroleukemia; COS-7, African green monkey kidney, SV40 transformed; MDCK, Cocker Spaniel kidney; HUVEC (200 -05n), Human umbilical vein endothelial cells (HUVEC); Neonate; NCI-H322, Caucasian bronchioloalveolar carcinoma; SK.N.SH, Caucasian neuroblastoma; LNCaP.FGC, Caucasian prostate cancer; OE21, Caucasian esophageal squamous cell carcinoma; PSN1, Human pancreatic cancer; ISHIKAWA, Asian endometrial adenocarcinoma; MFE-280, Caucasian endometrial adenocarcinoma; MG-63, Human osteosarcoma; RK 13. Rabbit kidney, BVDV negative; EoL-1 cells, human eosinophilic leukemia; VCaP, human prostate cancer metastasis; tsA201, human embryonic kidney, SV40 transformed; CHO, Chinese hamster ovary; HT1080, human fibrosarcoma; PANC-1, Caucasian pancreas; Saos-2, human primary osteoblastic sarcoma; fibroblast growth medium (116K-500), fibroblast growth medium kit; ND7 / 23, mouse neuroblastoma x rat neuron heterozygote; SK-OV-3, Caucasian ovarian adenocarcinoma; COV434, human ovarian granulosa cell tumor; Hep 3B, human hepatocellular carcinoma; Vero (WHO), African green monkey kidney; Nthy-ori 3-1, human thyroid follicular epithelial cells; U373 MG (Uppsala), human glioblastoma astrocytoma; A375, human malignant melanoma; AGS, Caucasian gastric adenocarcinoma; CAKI 2, Caucasian renal cancer; COLO205, Caucasian colonic adenocarcinoma; COR-L23, Caucasian large cell lung cancer; IMR 32, Caucasian neuroblastoma; QT 35, Japanese quail fibrosarcoma; WI 38, Caucasian fetal lung; HMVII, human vaginal malignant melanoma; HT55, human colon cancer; TK6, human lymphoblast, thymidine kinase heterozygote; SP2 / 0-AG14 (AC-FREE), mouse x mouse hybridoma, non-secreting, serum-free, animal component-free (AC); AR42J, or rat pancreatic exocrine tumor, or any combination thereof.
[0096] This document describes cells and cell lines containing transcription factor nucleic acids, which comprise a promoter nucleotide sequence regulated by a response element and a nucleotide sequence encoding a synthetic transcription factor, wherein the promoter nucleotide sequence regulated by the response element is located 5' to the nucleotide sequence encoding the synthetic transcription factor. In some embodiments, the cell line is a mammalian cell line. In some embodiments, the promoter regulated by the response element is a cAMP response element nucleotide sequence, an NFAT transcription factor response element nucleotide sequence, a FOS promoter nucleotide sequence, or a serum response element nucleotide sequence. In some embodiments, the promoter regulated by the response element is an NFAT response element-regulated promoter. In some embodiments, the cell line comprises a reporter nucleic acid comprising a synthetic transcription factor promoter nucleotide sequence and a nucleotide sequence encoding a reporter, wherein the synthetic transcription factor promoter nucleotide sequence is located 5' to the nucleotide sequence encoding the reporter, and wherein the synthetic transcription factor promoter nucleotide sequence is capable of being bound by the synthetic transcription factor.
[0097] In some embodiments, the cell line contains high baseline reporter activity. In some embodiments, the high baseline reporter activity is at least about 5%, 10%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, or 500% higher than the background, where the background is the level of reporter activity observed against cells or cell lines that do not contain the reporter. For such comparisons, the cell or cell line typically used as a reference would be a parent cell line containing the reporter (e.g., HEK293 containing the reporter versus HEK293 without the reporter).
[0098] In some embodiments, the cell line contains high baseline reporter activity. In some embodiments, the high baseline reporter activity is at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 32, 50, 75, 100, 200, 500, 750, 1,000, 2,000, 5,000, 10,000, or 20,000 times higher than the background, where the background is the reporter activity level observed for cells or cell lines that do not contain the reporter. In some embodiments, the cell line contains high baseline reporter activity. In some embodiments, the high baseline reporter activity is at least about 30 times higher than the background, where the background is the reporter activity level observed for cells or cell lines that do not contain the reporter. In some embodiments, the high baseline reporter activity is at least about 32 times higher than the background, where the background is the reporter activity level observed for cells or cell lines that do not contain the reporter. For such comparisons, the cells or cell lines typically used as references are parental cell lines containing reporters (e.g., HEK293 containing reporters versus HEK293 without reporters).
[0099] In some embodiments, the cell line contains low variability in basic reporter activity. In some embodiments, low variability in basic reporter activity means a coefficient of variation of less than about 0.6. In some embodiments, low variability in basic reporter activity means a coefficient of variation of less than about 0.5. In some embodiments, low variability in basic reporter activity means a coefficient of variation of less than about 0.4. In some embodiments, low variability in basic reporter activity means a coefficient of variation of less than about 0.3. In some embodiments, low variability in basic reporter activity means a coefficient of variation of less than about 0.2. In some embodiments, low variability in basic reporter activity means a coefficient of variation of less than about 0.1.
[0100] Without being bound by theory, reduced variability and high levels of basal activity can be obtained by selecting clonal cell lines comprising at least 2, 3, 4, 5, or more copies of a transcription factor nucleic acid containing a promoter nucleotide sequence regulated by a response element and a nucleotide sequence encoding a synthetic transcription factor, wherein the promoter nucleotide sequence regulated by the response element is located 5' to the side of the nucleotide sequence encoding the synthetic transcription factor. In some embodiments, the promoter regulated by the response element is a cAMP response element nucleotide sequence, an NFAT transcription factor response element nucleotide sequence, a FOS promoter nucleotide sequence, or a serum response element nucleotide sequence. In some embodiments, the promoter regulated by the response element is an NFAT response element-regulated promoter. In some embodiments, the cell line contains only 1 copy of a reporter nucleic acid containing a synthetic transcription factor promoter nucleotide sequence and a nucleotide sequence encoding a reporter. In some embodiments, the cell line contains only 2 copies of a reporter nucleic acid containing a synthetic transcription factor promoter nucleotide sequence and a nucleotide sequence encoding a reporter. In some embodiments, the cell line contains a reporter nucleic acid comprising a nucleotide sequence for a synthetic transcription factor promoter and a nucleotide sequence encoding a reporter that remains in an unintegrated or episome state. In some embodiments, the cell line also contains cDNA or other intronless forms of nucleic acid encoding a cell signaling protein. In some embodiments, the cell signaling protein is a GPCR or a GPCR subunit.
[0101] In some embodiments, the cell contains nucleic acids encoding members of the G protein-coupled receptor (GPCR) family. GPCRs, also known as seven-transmembrane receptors, are ligand-binding cell surface signaling proteins. When a ligand binds to a GPCR, it causes a conformational change in the GPCR, allowing it to act as a guanine nucleotide exchange factor (GEF). The GPCR can then activate the associated G protein by exchanging GDP bound to the G protein for GTP. The α subunit of the G protein can then dissociate from the β and γ subunits along with the bound GTP to further influence intracellular signaling proteins or target functional proteins (Gαs, Gαi / o, Gαq / 11, Gα12 / 13) that are directly dependent on the α subunit type. At least approximately 800 GPCRs are encoded in the human genome, broadly classified into classes A, B, and C, all of which can be used with the systems described herein. In some embodiments, nucleic acids encoding members of the GPCR family can be integrated into the genome. In some embodiments, nucleic acids encoding members of the GPCR family can be kept in an episodic state.
[0102] In some implementations, the cell contains nucleic acid encoding a member of the receptor tyrosine kinase family. Receptor tyrosine kinases (RTKs) are cell surface receptors with high affinity for many polypeptide growth factors, cytokines, and hormones. Receptor tyrosine kinases have been shown to play a crucial role not only in the regulation of normal cellular processes but also in the development and progression of many types of cancer. Many types of RTKs exist, and any member can be used in the system described herein. In some implementations, RTKs include members of class I RTKs (EGF receptor family) (ErbB family); class II RTKs (insulin receptor family); class III RTKs (PDGF receptor family); class IV RTKs (VEGF receptor family); class V RTKs (FGF receptor family); class VI RTKs (CCK receptor family); class VII RTKs (NGF receptor family); class VIII RTKs (HGF receptor family); class IX RTKs (Eph receptor family); class X RTKs (AXL receptor family); class XI RTKs (TIE receptor family); class XII RTKs (RYK receptor family); class XIII RTKs (DDR receptor family); class XIV RTKs (RET receptor family); class XV RTKs (ROS receptor family); class XVI RTKs (LTK receptor family); class XVII RTKs (ROR receptor family); class XVIII RTKs (MuSK receptor family); class XIX RTKs (LMR receptor); or class XX RTKs (undetermined). In some implementations, nucleic acids encoding RTK family members can be integrated into the genome. In other implementations, nucleic acids encoding RTK family members can be kept in an episodic state.
[0103] This document also describes mammalian cell lines containing NFAT response elements. In some embodiments, mammalian cell lines containing NFAT response elements include cb29.
[0104] This document also describes mammalian cell lines containing NFAT response elements. In some embodiments, mammalian cell lines containing NFAT response elements include cb37.
[0105] Using a systematic approach
[0106] The polynucleotide sequences of this invention can be used when transfected into cells. Transfection can be performed using a variety of transfection agents, including but not limited to lipid transfection, calcium phosphate precipitation, viral transduction, or electroporation. Transfection can be transient or stable. In a stable transfection embodiment, stably transfected cells can be frozen or stored for later use.
[0107] In some embodiments, a single nucleic acid relay system is transfected into the cell population. In some embodiments, 1, 2, 3, 4, 5, 10, 100, or more nucleic acid relay systems are transfected into the cell population. In some embodiments, two nucleic acid relay systems are transfected into the cell population. In some embodiments, three nucleic acid relay systems are transfected into the cell population. In some embodiments, four nucleic acid relay systems are transfected into the cell population. In some embodiments, five nucleic acid relay systems are transfected into the cell population. In some embodiments of transfecting the cell population with multiple nucleic acid relay systems, the multiple nucleic acid relay systems comprise promoters regulated by different response elements. In some embodiments where the multiple nucleic acid relay systems comprise promoters regulated by different response elements, the multiple nucleic acid relay systems comprise different reporter units. In some embodiments, the different reporter units comprise UMIs.
[0108] The cell population transfected with the nucleic acids of the present invention can be of any size. In some embodiments, the cell population comprises 1,000, 10,000, 100,000, 1,000,000, 10,000,000 or more cells. In some embodiments, at least about 1,000 cells are transfected using one or more transcription relay systems. In some embodiments, at least about 10,000 cells are transfected using one or more transcription relay systems. In some embodiments, at least about 100,000 cells are transfected using one or more transcription relay systems. In some embodiments, at least about 1,000,000 cells are transfected using one or more transcription relay systems. In some embodiments, at least about 10,000,000 cells are transfected using one or more transcription relay systems.
[0109] In some embodiments, the nucleic acid system of the present invention can be used in multi-well plate experiments. Non-limiting examples of multi-well plates compatible with the nucleic acid relay system of the present invention include 6, 12, 24, 48, 96, 384, or 1,536-well plates. In some embodiments, each well of the multi-well plate contains a cell population transfected with a single transcription relay system. In some embodiments, each well of the multi-well plate contains a cell population transfected with multiple transcription relay systems. In some embodiments, each well contains multiple cell populations, each cell population transfected with a single nucleic acid relay system. In some embodiments, each well contains multiple cell populations, each cell population transfected with multiple nucleic acid relay systems.
[0110] In some embodiments, the test reagent is applied to cells transfected with the transcription relay system of the present invention. In some embodiments, after contacting the cells with the test reagent, the activation level of reporter molecule transcription is measured. In some embodiments, the test reagent is a chemical, small molecule, biomolecule, peptide, polynucleotide, aptamer, or any combination thereof. In some embodiments, a single test reagent is applied to a cell population. In some embodiments, multiple test reagents are applied to a cell population.
[0111] In some embodiments, the transcription relay system of the present invention is suitable for measuring the response of a GPCR to a test reagent. The nucleic acid system of the present invention is suitable for use with any GPCR receptor. In some embodiments, the transcription relay system is suitable for use with a GPCR receptor by utilizing a promoter regulated by a cAMP response element. Non-restrictive examples of GPCRs include serotonin receptors, acetylcholine receptors, adenosine receptors, adrenaline receptors, angiotensin receptors, apelin receptors, bile acid receptors, bufotenoid receptors, bradykinin receptors, cannabinoid receptors, chemerin receptors, chemokine receptors, cholecystokinin receptors, dopamine receptors, endothelin receptors, formylate receptors, free fatty acid receptors, glycopeptide receptors, ghrelin receptors, glycoprotein hormone receptors, gonadotropin-releasing hormone receptors, GPR18, GPR55, GPR119, G protein-coupled estrogen receptors, histamine receptors, hydroxycarboxylic acid receptors, kisspeptin receptors, leukotriene receptors, LPA receptors, S1P receptors, melanin-gathering hormone receptors, melanocortin receptors, and melatonin receptors. Motilin receptor, neuropeptide U receptor, neuropeptide FF / neuropeptide AF receptor, neuropeptide S receptor, neuropeptide W / neuropeptide B receptor, neuropeptide Y receptor, neurotensin receptor, opioid receptor, opsin receptor, orexin receptor, ketoglutarate receptor, P2Y receptor, platelet-activating factor receptor, prodylin receptor, prolactin-releasing peptide receptor, prostaglandin receptor, protease-activating receptor, QRFP receptor, relaxin family peptide receptor, somatostatin receptor, succinate receptor, tachykinin receptor, thyrotropin-releasing hormone receptor, trace amine receptor, ostrich teratosine receptor, angiotensin and oxytocin receptors, calcitonin receptor, corticotropin-releasing factor receptor, glucagon receptor family, parathyroid hormone receptor, VIP and PACAP receptors, calcium-sensitive receptor, GABA B Receptors, metabolized glutamate receptors, the first family of taste receptors, coiled receptors, adhesive GPCRs, orphan receptors, and any combination thereof.
[0112] The nucleic acids of the present invention are compatible with many vectors common in the art. Non-limiting examples of vectors include genome integration vectors, augmentation vectors, plasmids, viral vectors, granules, bacterial artificial chromosomes, and yeast artificial chromosomes. Non-limiting examples of viral vectors compatible with the nucleic acids of the present invention include vectors derived from lentiviruses, retroviruses, adenoviruses, and adeno-associated viruses. In some embodiments, the nucleic acids of the present invention are present on a vector comprising a sequence that specifically integrates into a fixed location or a set of defined sites in the genome (e.g., AttP-AttB recombination).
[0113] In some embodiments, the transcription relay system described herein is incorporated into a single vector. In some embodiments, the single vector is transiently transfected into cells. In some embodiments, the single vector is stably transfected into cells.
[0114] In some embodiments, the transcription relay system is divided into two vectors. In some embodiments, a transcription factor nucleic acid containing a promoter nucleotide sequence regulating a response element and a nucleotide sequence encoding a synthetic transcription factor is incorporated into a first vector, and a reporter nucleic acid containing a promoter nucleotide sequence of a synthetic transcription factor and a nucleotide sequence encoding a reporter is incorporated into a second vector. In some embodiments, the first and second vectors are transiently transfected into cells. In some embodiments, the first and second vectors are stably transfected into cells. In some embodiments, the first vector is stably transfected into cells, and the second vector is transiently transfected into cells. In some embodiments, the first vector is transiently transfected into cells, and the second vector is stably transfected into cells.
[0115] Many well-known molecular biology techniques can be used to construct vectors containing the transcription relay system or parts thereof described herein. Detailed protocols for many such procedures (including amplification, cloning, mutagenesis, transformation, etc.) are described, for example, in Ausubel et al., Current Protocols in Molecular Biology (supplemented through 2012), John Wiley & Sons, New York 10 (“Ausubel”); Sambrook et al., Molecular Cloning – A Laboratory Manual (4th Ed.), Vol. 1-3, Cold Spring Harbor Laboratory, Cold Spring Harbor, New York, 2012 (“Sambrook”); and Abelson et al., Guide to Molecular Cloning Techniques (Methods in Enzymology), Volume 152, Academic Press, Inc., San Diego, CA (“Abelson”).
[0116] Example
[0117] The following illustrative examples represent embodiments of the compositions and methods described herein and are not intended to be limiting in any way.
[0118] Example 1 - Screening of Example GPCR Receptors for CRE Activation
[0119] In this embodiment, as Figure 1A and 1B The configured transcription relay system, containing nucleic acids, is used to screen for potential compounds that induce GPCR signaling. For this embodiment, Figure 1A The nucleic acid includes cAMP response element (CRE) activation, which leads to the expression of the synthetic transcription factor Gal4-VPR (containing the Gal4 DNA binding domain and the chimeric activation domain VP64-p65-Rta). Figure 1B The nucleic acids contained promoters that could be bound and activated by the Gal4-VPR synthetic transcription factor, leading to the expression of reporter elements including the luciferase gene and the gene encoding UMI. The cells used contained... Figure 1A and 1B The system stably integrates nucleic acids and a given GPCR. Each UMI is associated with a given GPCR, allowing CRE expression to map to a specific GPCR. This enables multiplexing of assays.
[0120] On day 1, cells were seeded at 35,000 cells / well in DMEM in 96-well assay plates. On day 2, the medium was replaced with 0.5% FBS + DMEM. On day 3, the medium was removed, and the test compound was added to 25 μL of Opti-membrane at the desired concentration. Approximately 4 hours later, the medium was removed and replaced with lysis buffer for RNA extraction. RNA was extracted using standard methods or kits and subsequently quantified using standard assays. After sequencing library preparation, RNAseq was performed on an Illumina MiSeq instrument.
[0121] Example 2 - Screening of Example GPCR Receptors for NFAT Activation
[0122] In this embodiment, as Figure 1A and 1B The configured transcription relay system, containing nucleic acids, is used to screen for potential compounds that induce GPCR signaling. For this embodiment, Figure 1A The nucleic acids include activation of the nuclear factor response element (NFAT) of activated T cells, which leads to the expression of the synthetic transcription factor Gal4-VPR (containing the Gal4 DNA-binding domain and the chimeric activation domain VP64-p65-Rta). Figure 1B The nucleic acids contained promoters that could be bound and activated by the Gal4-VPR synthetic transcription factor, leading to the expression of reporter elements including the luciferase gene and the gene encoding UMI. The cells used contained... Figure 1A and 1B The system stably integrates nucleic acids and a given GPCR. Each UMI is associated with a given GPCR, allowing CRE expression to map to a specific GPCR. This enables multiplexing of assays.
[0123] On day 1, cells were seeded at 35,000 cells / well in DMEM in 96-well assay plates. On day 2, the medium was replaced with 0.5% FBS + DMEM. On day 3, the medium was removed, and the test compound was added to 25 μL of Opti-membrane at the desired concentration. Approximately 4 hours later, the medium was removed and replaced with lysis buffer for RNA extraction. RNA was extracted using standard methods or kits and subsequently quantified using standard assays. After sequencing library preparation, RNAseq was performed on an Illumina MiSeq instrument.
[0124] Example 3 - Screening of example GPCR receptors for CRE activation targeting multiple GPCRs
[0125] In this embodiment, each as follows Figure 1A and 1BThe configured transcription relay system, comprising 100 or more nucleic acids, is used to screen for potential compounds that induce GPCR signaling. For this embodiment, Figure 1A Each of the nucleic acids includes cAMP response element (CRE) activation, which leads to the expression of the synthetic transcription factor Gal4-VPR (containing the Gal4 DNA-binding domain and the chimeric activation domain VP64-p65-Rta). Figure 1B Each nucleic acid contained a promoter that could be bound and activated by the Gal4-VPR synthetic transcription factor, leading to the expression of reporter elements including the luciferase gene and the gene encoding UMI. The cell populations used each contained genes encoding... Figure 1A and 1B The system stably integrates nucleic acids, as well as a given single GPCR. Multiple cell populations of 100 or more are mixed together to form a hybrid cell population, each encoding a single unique GPCR. Each UMI is associated with a given GPCR, allowing CRE expression to map to a specific GPCR. This enables multiplexing of assays.
[0126] On day 1, the mixed cell population was seeded at 35,000 cells / well in DMEM in a 96-well assay plate. On day 2, the medium was replaced with 0.5% FBS + DMEM. On day 3, the medium was removed, and the test compound was added in 25 μL Opti-membrane at the desired concentration. After approximately 4 hours, the medium was removed and replaced with lysis buffer for RNA extraction. RNA was extracted using standard methods or kits and subsequently quantified using standard assays. After sequencing library preparation, RNAseq was performed on an Illumina MiSeq instrument.
[0127] Example 4 - Using transcription relay to amplify report output
[0128] The experiments in this example demonstrate that, compared to systems without transcriptional relays, the use of a transcriptional relay system increases luciferase signal and reduces the coefficient of variation of the luciferase signal. HEK293-derived cells carrying a single integrated CRE-luciferase or cells carrying a single integrated UAS-luciferase with multiple copies of semi-randomly integrated CRE-Gal4-VPR were seeded at 30,000 cells / well in 100 μL of DMEM + 10% FBS in white-walled poly-L-lysine-coated 96-well plates. 50 μL of Opti-mem containing 45 ng of doxycycline was added to the top of the cells. After 24 hours, DMSO was added. Cells were treated with DMSO for specified time periods. After the specified incubation time, the medium was aspirated and replaced with 35 μL of DMEM, and the cells were then assayed using the Bright-Glo luciferase assay kit [Promega] according to the manufacturer's instructions. Figure 2 Cells carrying a single integrated CRE-luciferase (gray) and cells carrying a single integrated UAS-luciferase accompanied by multiple copies of semi-randomly integrated CRE-Gal4-VPR (black) are shown, illustrating the luciferase activity expressed therefrom. Experiments were performed technically in triplicate, and the coefficient of variation for each sample was calculated as follows: Figure 3 As shown.
[0129] Example 5 - Enhancing Fold Induction of Transcription Relays Using Degradation Determinant Tags on Gal4-VPR
[0130] The experiments in this example demonstrate that the fold-induction of luciferase signaling increases when the Gal4-VPR in the transcription relay system contains a degradation determinant tag. HEK293-derived cells carrying a single integrated TRE-CHRM3::UAS-luciferase dual gene cassette and multiple semi-randomly integrated FOS-Gal4-VPR-CP (degradation determinant) or FOS-Gal4-VPR (without degradation determinant) were seeded at 30,000 cells / well in 100 μL LMEM + 10% FBS in white-walled poly-L-lysine-coated 96-well plates. 50 μL of Opti-mem containing 45 ng doxycycline was added to the top of the cells. After 24 hours, the cells were treated with DMSO or 1 M carbacholine for 8 hours. After the specified incubation time, the medium was aspirated and replaced with 35 μL LMEM, and the cells were then assayed using the Bright-Glo luciferase assay kit [Promega] according to the manufacturer's instructions. The ratio of luciferase activity in carbacholine to luciferase activity in DMSO is plotted on... Figure 4 middle.
[0131] Example 6 - Cell lines containing NFAT response elements
[0132] The cell lines described in this embodiment possess integrated copies of the NFAT response element transcription relay (NFAT promoter that drives the transcription of synthetic transcription factors). These cell lines were generated as a genetic heterologous pool in terms of copy number and integration site. Single-cell clones were isolated from this pool and amplified. These cell lines were further used to integrate GPCRs and UAS-luciferase-barcode reporter molecules to test their ability to detect NFAT signaling in multiplexed configurations. From these 10 cell libraries, two cell libraries capable of detecting the highest numerical values of different GPCR hits against control agonists were identified: cb29 (constructed from clone c713) and cb37 (constructed from clone c708), as shown below. Figure 5 As shown.
[0133] Importantly, we found that the isoclonal cell lines that generated these two cell libraries shared two common characteristics. First, these cell lines exhibited the highest levels of reporter expression in the unstimulated state (see...). Figure 6 (“basal activity-reverse transfection”). Secondly, the two corresponding cell libraries may show minimal levels of variation in a dependent manner (see “basal activity-reverse transfection”). Figure 6 (“BCV”).
[0134] While preferred embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided by way of example only. Many variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein can be used to practice the invention.
[0135] All publications, patent applications, granted patents, and other documents mentioned in this specification are incorporated herein by reference, as if each individual publication, patent application, granted patent, or other document were expressly and separately incorporated herein by reference in its entirety. Definitions contained in the text incorporated by reference that contradict the definitions in this disclosure are excluded.
Claims
1. A population of mammalian cells comprising a transcriptional relay system, comprising: a) a transcription factor nucleic acid comprising a response element-regulated promoter nucleotide sequence and a nucleotide sequence encoding an exogenous synthetic transcription factor, wherein the response element-regulated promoter nucleotide sequence comprises a cAMP response element nucleotide sequence, a NFAT transcription factor response element nucleotide sequence, a FOS promoter nucleotide sequence, or a serum response element nucleotide sequence located 5' to the nucleotide sequence encoding the exogenous synthetic transcription factor; and b) a reporter nucleic acid comprising an exogenous synthetic transcription factor promoter nucleotide sequence and a nucleotide sequence encoding a reporter, wherein the reporter nucleic acid further comprises a unique molecular identifier, wherein the exogenous synthetic transcription factor promoter nucleotide sequence is located 5' to the nucleotide sequence encoding the reporter, and wherein the exogenous synthetic transcription factor promoter nucleotide sequence is capable of being bound by the exogenous synthetic transcription factor, wherein the population of mammalian cells has a basal reporter activity that is at least 2-fold greater than a background, wherein the background is the level of reporter activity observed for a parental cell or cell line that does not comprise the transcriptional relay system, wherein the population of mammalian cells is incapable of developing into an animal individual.
2. The population of mammalian cells of claim 1, wherein the response element-regulated promoter nucleotide sequence comprises the cAMP response element nucleotide sequence.
3. The population of mammalian cells of claim 1, wherein the exogenous synthetic transcription factor comprises a DNA binding domain from a first transcription factor and a transcriptional activation domain from a second transcription factor.
4. The population of mammalian cells of claim 3, wherein the DNA binding domain is from Gal4, PPR1, Lac9, or LexA.
5. The population of mammalian cells of claim 4, wherein the DNA binding domain comprises an amino acid sequence that is at least 90% identical to the sequence set forth in SEQ ID NO:
1.
6. The population of mammalian cells of claim 4, wherein the DNA binding domain comprises an amino acid sequence that is at least 95% identical to the sequence set forth in SEQ ID NO:
1.
7. The population of mammalian cells of claim 4, wherein the DNA binding domain comprises an amino acid sequence that is identical to the sequence set forth in SEQ ID NO:
1.
8. The population of mammalian cells of claim 5, wherein the DNA binding domain comprises a variant of the amino acid sequence of SEQ ID NO:
1.
9. The population of mammalian cells of claim 3, wherein the transcriptional activation domain comprises VP64, p65, or Rta.
10. The population of mammalian cells of claim 9, wherein the transcriptional activation domain comprises an amino acid sequence that is at least 90% identical to the sequence set forth in SEQ ID NO:
14.
11. The population of mammalian cells of claim 9, wherein the transcriptional activation domain comprises an amino acid sequence that is at least 95% identical to the sequence set forth in SEQ ID NO:
14.
12. The population of mammalian cells of claim 9, wherein the transcriptional activation domain comprises an amino acid sequence that is identical to the sequence set forth in SEQ ID NO:
14.
13. The population of mammalian cells of claim 10, wherein the transcriptional activation domain comprises an amino acid sequence variant of SEQ ID NO: 14, wherein the sequence variant increases or decreases transcriptional activation.
14. The population of mammalian cells of claim 1, wherein the exogenous synthetic transcription factor comprises an amino acid sequence that is at least 90% identical to the sequence set forth in SEQ ID NO:
10.
15. The population of mammalian cells of claim 14, wherein the exogenous synthetic transcription factor comprises an amino acid sequence that is at least 95% identical to the sequence set forth in SEQ ID NO:
10.
16. The population of mammalian cells of claim 15, wherein the exogenous synthetic transcription factor comprises an amino acid sequence that is identical to the sequence set forth in SEQ ID NO:
10.
17. The population of mammalian cells of claim 1, wherein the exogenous synthetic transcription factor comprises a polypeptide sequence that destabilizes the exogenous synthetic transcription factor.
18. The population of mammalian cells of claim 17, wherein the polypeptide sequence that destabilizes the exogenous synthetic transcription factor comprises a PEST or CL1 polypeptide sequence.
19. The population of mammalian cells of claim 1, wherein the exogenous synthetic transcription factor promoter nucleotide sequence comprises a nucleotide sequence that can be bound by Gal4, PPR1, Lac9, or LexA.
20. The population of mammalian cells of claim 1, wherein the reporter comprises a fluorescent protein, a luciferase protein, a beta-galactosidase, a beta-glucuronidase, a chloramphenicol acetyltransferase, or a placental alkaline phosphatase.
21. The population of mammalian cells of claim 1, wherein the unique molecular identifier is unique to a test polypeptide, wherein the test polypeptide is encoded by the reporter nucleic acid.
22. The population of mammalian cells of claim 1, wherein the transcription factor nucleic acid comprises a nucleotide sequence proximal to the response element-regulated promoter nucleotide sequence that can be bound by a transcriptional repressor.
23. The population of mammalian cells of claim 22, wherein the transcription factor nucleic acid comprises a nucleotide sequence proximal to the response element-regulated promoter nucleotide sequence that extends the 5' untranslated region of the mRNA encoded by the nucleotide sequence encoding the exogenous synthetic transcription factor.
24. The population of mammalian cells of claim 23, wherein the 5' untranslated region of the mRNA encoded by the nucleotide sequence encoding the exogenous synthetic transcription factor comprises one or more sequences that reduce translation of the exogenous synthetic transcription factor.
25. The population of mammalian cells of claim 1, wherein the transcription factor nucleic acid and the reporter nucleic acid are components of a single nucleic acid.
26. The population of mammalian cells of any one of claims 1-25, wherein the transcription factor nucleic acid, the reporter nucleic acid, or both the transcription factor nucleic acid and the reporter nucleic acid are integrated as a single copy into the genome of the population of cells.
27. The population of mammalian cells of claim 26, wherein the population of mammalian cells has a basal reporter activity that is at least 5-fold higher than background.
28. The population of mammalian cells of claim 26, wherein the population of mammalian cells has a basal reporter activity that is at least 30-fold higher than background.
29. The population of mammalian cells of claim 26, wherein the population of mammalian cells has a low coefficient of biological variation in reporter activity.
30. The population of mammalian cells of claim 29, wherein the low coefficient of biological variation in reporter activity is less than 0.
5.
31. A method for detecting the effect of a test agent on the activity of a promoter regulated by a response element, comprising contacting a population of mammalian cells according to any one of claims 1-30 with a test substance.
32. The method of claim 31, wherein the test agent is a small molecule chemical.
Citation Information
Patent Citations
Use of single-stranded nucleic acid binding proteins in sequencing
US20060024678A1
Methods for nucleic acid amplification and sequence determination
US20060024711A1
Detecting apparent mutations in nucleic acid sequences
US20060286566A1
Optical train and method for TIRF single molecule detection and analysis
US20080087826A1
Molecules and methods for nucleic acid sequencing
US20080103058A1