Systems and methods for utilizing plasma membrane-anchored proteases
By utilizing the plasma membrane-adjacent reporter system and activating reporter gene expression through plasma membrane-anchored protease cleavage linkers, the problem of intracellular peptide transport regulation has been solved, enabling high-throughput screening and identification of assays that regulate peptide transport and improving the therapeutic effects of degenerative diseases.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-10
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies are insufficient to effectively regulate intracellular protein processing, folding, and transport, resulting in the lack of effective treatment for various human diseases and conditions, such as Wilson's disease, retinitis pigmentosa, and cystic fibrosis. In particular, the therapeutic needs for regulating intracellular peptide transport-related diseases remain unmet.
A system incorporating a plasma membrane-adjacent reporter is designed to activate reporter gene expression by anchoring a protease cleavage linker to the plasma membrane. This system utilizes the ability of fluorescent proteins or luciferase proteins to express unique molecular identifiers, enabling the screening of assays that regulate intracellular processing, folding, or transport of peptides. The system also employs the ability to screen for reporter constructs through linker cleavage and laser-mediated linker cleavage, thereby activating reporter gene expression.
This invention enables the identification of intracellular processing, folding, or transport of regulatory peptides in high-throughput assays, assesses the plasma membrane transport outcomes of mutant peptides, screens for assays that can alleviate degenerative diseases, provides a method for screening and identifying regulatory peptide transport, and improves the effectiveness of disease treatment.
Smart Images

Figure CN122422757A_ABST
Abstract
Description
[0001] Cross-referencing This application claims the benefit of U.S. Provisional Patent Application No. 63 / 537,762, filed September 11, 2023, the contents of which are incorporated herein by reference in their entirety. Background Technology
[0002] Many human diseases and conditions are associated with suboptimal protein processing, folding, and transport. These include, among others, Wilson's disease, Menke's disease, or degenerative diseases such as retinitis pigmentosa, Alzheimer's disease, Parkinson's disease, Huntington's disease, cystic fibrosis, alpha-1 antitrypsinemia, or amyotrophic lateral sclerosis (ALS). There remains a need for compounds that can modulate cellular protein processing pathways for the treatment of diseases. Summary of the Invention
[0003] This document describes systems incorporating plasma membrane-near reporters and their uses for screening and identifying assays that regulate or influence the intracellular processing, folding, or transport of peptides intended for use on the plasma membrane or for secretion. In some embodiments, the systems and methods described herein relate to plasma membrane-near reporters and their activity in response to doses of various assays. In particular, the systems and methods described herein relate to unique molecular identifiers that induce peptide expression or are associated with plasma membrane-near reporters in response to doses of one or more assays. The methods described herein are also suitable for identifying assays that regulate the intracellular processing, folding, or transport of peptides in the context of high-throughput assays. The systems and methods described herein can also be used for compliance assays to assess the effect of assays on plasma membrane transport outcomes of mutant peptides, including pathogenic mutant peptides identified in patients with certain conditions or diseases (e.g., retinitis pigmentosa, GLP-1R downregulation, cystic fibrosis, etc.).
[0004] One aspect of the systems and methods described herein envisions a system comprising eukaryotic cells, wherein the eukaryotic cells comprise: a plasma membrane construct comprising a plasma membrane polypeptide coupled to a transcription factor via a linker, wherein the linker comprises a protease cleavage site; and a plasma membrane-anchored protease; wherein the plasma membrane-anchored protease is capable of cleaving the linker. In some embodiments, the system further comprises a reporter construct comprising a promoter and a reporter gene. In some embodiments, the promoter is bound by a transcription factor upon linker cleavage, and wherein the reporter gene is expressed upon linker cleavage. In some embodiments, the plasma membrane construct is encoded by a foreign nucleic acid. In some embodiments, the reporter construct is encoded by a foreign nucleic acid. In some embodiments, the reporter gene comprises a unique molecular identifier. In some embodiments, the reporter gene comprises a fluorescent protein or a luciferase protein. In some embodiments, the reporter gene comprises a fluorescent protein and a unique molecular identifier or a luciferase protein and a unique molecular identifier. In some embodiments, the plasma membrane-anchored protease is a component of the plasma membrane of the eukaryotic cell. In some embodiments, the plasma membrane-anchored protease comprises a membrane-tethering protease. In some embodiments, the membrane-tethering protease comprises a pleckstrin homology domain, a platelet-derived growth factor receptor, or a Lyn anchor domain. In some embodiments, the transcription factor comprises a DNA-binding domain and a transcription activation domain. In some embodiments, the DNA-binding domain comprises a Gal4, PPR1, Lac9, zinc finger, or LexA DNA-binding domain. In some embodiments, the transcription activation domain comprises a VP64, VPr, p65, Rta, or VP16 activation domain. In some embodiments, the linker comprises a flexible amino acid linker. In some embodiments, the linker is about 2 to about 31 amino acids in length. In some embodiments, the plasma membrane-anchoring protease comprises a tobacco etch virus, aspartic acid, glutamate, metalloproteinase, cysteine, serine, or threonine protease. In some embodiments, the plasma membrane polypeptide comprises rhodopsin. In some embodiments, the expression of the plasma membrane construct is inducible. In some embodiments, the expression of the plasma membrane construct is induced by a response to doxycycline. In some embodiments, the plasma membrane construct is localized to the plasma membrane of eukaryotic cells after expression. In some embodiments, the plasma membrane polypeptide comprises an amino acid sequence having at least about 90%, 95%, 97%, 98%, 99%, or 100% identity with any of SEQ ID 1-6. In some embodiments, expression of the reporter construct indicates the ability of the assay agent to alleviate the condition. In some embodiments, the condition is a degenerative disease. In some embodiments, the eukaryotic cells are mammalian cells. In some embodiments, the mammalian cells are human cells. In some embodiments, a population of eukaryotic cells contains the system.Some implementations envision a method for screening test reagents that involves contacting a population of eukaryotic cells with the test reagent. In some implementations, the test reagent comprises a small molecule compound.
[0005] One aspect of the systems and methods described herein envisions a system comprising eukaryotic cells, wherein the eukaryotic cells comprise: a plasma membrane construct comprising a plasma membrane polypeptide coupled to a transcription factor via a linker, wherein the linker is cleavable; and a reporter construct comprising a promoter and a reporter gene comprising a unique molecular identifier; wherein the promoter is bound by the transcription factor upon linker cleavage, and wherein the reporter gene is expressed upon linker cleavage. In some embodiments, the system further comprises a plasma membrane-anchoring protease capable of cleaving the linker. In some embodiments, the plasma membrane construct is encoded by a foreign nucleic acid. In some embodiments, the reporter construct is encoded by a foreign nucleic acid. In some embodiments, the reporter gene further encodes a fluorescent protein or luciferase protein. In some embodiments, the plasma membrane-anchoring protease is a component of the plasma membrane of the eukaryotic cell. In some embodiments, the membrane-tethering protease comprises a Plek substrate protein homology domain, a platelet-derived growth factor receptor, or a Lyn anchor domain. In some embodiments, the transcription factor comprises a DNA-binding domain and a transcription activation domain. In some embodiments, the DNA-binding domain comprises a Gal4, PPR1, Lac9, zinc finger, or LexA DNA-binding domain. In some embodiments, the transcriptional activation domain comprises a VP64, VPRr, p65, Rta, or VP16 activation domain. In some embodiments, the linker comprises a flexible amino acid linker. In some embodiments, the linker is about 2 to about 31 amino acids in length. In some embodiments, the plasma membrane anchoring protease comprises a tobacco etch virus, aspartic acid, glutamic acid, metalloid, cysteine, serine, or threonine protease. In some embodiments, the plasma membrane polypeptide comprises rhodopsin. In some embodiments, the expression of the plasma membrane construct is inducible. In some embodiments, the expression of the plasma membrane construct is induced by doxycycline. In some embodiments, the plasma membrane polypeptide is localized to the plasma membrane of a eukaryotic cell after expression. In some embodiments, the plasma membrane polypeptide comprises an amino acid sequence having at least about 90%, 95%, 97%, 98%, 99%, or 100% identity with any of SEQ ID 1-6. In some embodiments, the expression of the report sub-construct indicates the ability of the test agent to alleviate the condition. In some embodiments, the condition is a degenerative disease. In some embodiments, the eukaryotic cells are mammalian cells. In some embodiments, the mammalian cells are human cells. In some embodiments, a population of eukaryotic cells contains the system. Some embodiments envision a method for screening test agents, including contacting a population of eukaryotic cells with the test agent. In some embodiments, the test agent comprises a small molecule compound.
[0006] One aspect of the systems and methods described herein envisions a system comprising eukaryotic cells, wherein the eukaryotic cells comprise: a plasma membrane construct (PMC), the PMC comprising a plasma membrane polypeptide coupled to a transcription factor via a PMC linker, wherein the PMC linker comprises a protease cleavage site; and a plasma membrane-anchored protease; wherein the plasma membrane-anchored protease is capable of cleaving the PMC linker. In some embodiments, the system further comprises a reporter construct (RC), the reporter construct (RC) comprising an RC promoter and a reporter gene. In some embodiments, the RC promoter is bound by a transcription factor upon cleavage of the PMC linker. In some embodiments, the RC promoter comprises a synthetic DNA-binding domain-responsive promoter. In some embodiments, the RC promoter comprises SEQ ID NO: 43. In some embodiments, the RC promoter comprises a zinc finger binding site. In some embodiments, the RC promoter comprises 2 to 12 zinc finger binding sites. In some embodiments, the plasma membrane construct is encoded by a foreign nucleic acid. In some embodiments, the reporter construct is encoded by a foreign nucleic acid. In some embodiments, the reporter gene comprises a unique molecular identifier. In some embodiments, the reporter gene encodes a fluorescent protein or a luciferase protein. In some embodiments, the reporter gene encodes a fluorescent protein or a luciferase protein and further includes a unique molecular identifier. In some embodiments, the plasma membrane-anchored protease is encoded by a foreign nucleic acid, optionally wherein the expression of the plasma membrane-anchored protease is driven by a constitutive promoter. In some embodiments, the plasma membrane-anchored protease comprises a membrane-tethered protease. In some embodiments, the plasma membrane-anchored protease comprises a plasma membrane anchor linked to the protease via a protease tether. In some embodiments, the plasma membrane anchor comprises any one of SEQ ID NO: 35-38. In some embodiments, the protease tether comprises SEQ ID NO: 39 or 40. In some embodiments, the membrane-tethered protease comprises tobacco etch virus (TEV), aspartic, glutamic, metalloid, cysteine, serine, or threonine protease. In some embodiments, the membrane-tethered protease comprises a TEV protease or a variant of a TEV protease, or a functional fragment thereof. In some embodiments, the membrane-tethered protease comprises a sequence having at least 90% identity with SEQ ID NO: 41. In some embodiments, the membrane-tethering protease comprises SEQ ID NO: 42. In some embodiments, the transcription factor comprises a DNA-binding domain and a transcription activation domain. In some embodiments, the DNA-binding domain comprises a Gal4, PPR1, Lac9, zinc finger, or LexA DNA-binding domain. In some embodiments, the DNA-binding domain comprises a sequence having at least 90% sequence identity with any of SEQ ID NO: 25-29. In some embodiments, the transcription activation domain comprises a VP64, VPR, p65, Rta, or VP16 activation domain.In some embodiments, the transcriptional activation domain comprises a sequence having at least 90% sequence identity with any of SEQ ID NO: 30-34. In some embodiments, the PMC linker comprises a flexible amino acid linker. In some embodiments, the PMC linker is about 2 to about 31 amino acids in length. In some embodiments, the PMC linker comprises a TEV-cleavable sequence. In some embodiments, the PMC linker comprises a sequence having at least 90% sequence identity with SEQ ID NO: 20 or 21. In some embodiments, the PMC linker comprises a protease cleavage site comprising at least one of SEQ ID NO: 22-24. In some embodiments, the plasma membrane polypeptide comprises rhodopsin or a variant thereof. In some embodiments, the plasma membrane polypeptide comprises a sequence having at least 90% sequence identity with any of SEQ ID NO: 1-6. In some embodiments, the plasma membrane polypeptide comprises cystic fibrosis transmembrane conduction regulator (CFTR) or a variant of CFTR. In some embodiments, the plasma membrane polypeptide comprises a sequence having at least 90% sequence identity with SEQ ID NO: 7. In some embodiments, the plasma membrane polypeptide comprises a G protein-coupled receptor. In some embodiments, the plasma membrane polypeptide comprises a glucagon-like peptide-1 receptor (GLP-1R) or a variant of GLP-1R. In some embodiments, the plasma membrane polypeptide comprises a sequence having at least 90% sequence identity with SEQ ID NO: 8. In some embodiments, the plasma membrane polypeptide comprises a sequence having at least 90% sequence identity with any of SEQ ID NO: 9-19. In some embodiments, expression of the plasma membrane construct is inducible. In some embodiments, expression of the plasma membrane construct is induced in response to doxycycline. In some embodiments, the plasma membrane construct is localized to the plasma membrane of a eukaryotic cell after expression. In some embodiments, expression of the reporter construct indicates the ability of an assay agent to alleviate the condition. In some embodiments, the condition is a degenerative disease. In some embodiments, the eukaryotic cells are mammalian cells. In some embodiments, the mammalian cells are human cells. In some embodiments, the eukaryotic cell population comprises the system provided herein. The methods further provided herein include a method for screening test reagents, which involves contacting a population of eukaryotic cells comprising the system provided herein with the test reagent. In some embodiments, the test reagent comprises a small molecule compound.
[0007] One aspect of the systems and methods described herein envisions a system comprising eukaryotic cells, wherein the eukaryotic cells comprise: a plasma membrane construct (PMC), the PMC comprising a plasma membrane polypeptide coupled to a transcription factor via a PMC linker, wherein the PMC linker is cleavable; and a reporter construct (RC), the RC comprising an RC promoter and a reporter gene comprising a unique molecular identifier; wherein the RC promoter is bound by a transcription factor upon cleavage of the PMC linker. In some embodiments, the system further comprises a plasma membrane-anchoring protease capable of cleaving the linker, optionally wherein the plasma membrane-anchoring protease is encoded by a foreign nucleic acid. In some embodiments, the plasma membrane construct is encoded by a foreign nucleic acid. In some embodiments, the plasma membrane construct is encoded by a foreign nucleic acid. In some embodiments, the reporter construct is encoded by a foreign nucleic acid. In some embodiments, the reporter gene further encodes a fluorescent protein or luciferase protein. In some embodiments, the plasma membrane-anchoring protease is a component of the plasma membrane of the eukaryotic cell. In some embodiments, the plasma membrane-anchoring protease comprises a plasma membrane anchor linked to the protease via a protease tether. In some embodiments, the plasma membrane anchor comprises any one of SEQ ID NO: 35-38. In some embodiments, the protease tether comprises SEQ ID NO: 39 or 40. In some embodiments, the membrane-tethering protease comprises tobacco etch virus (TEV), aspartic acid, glutamic acid, metalloid, cysteine, serine, or threonine protease. In some embodiments, the membrane-tethering protease comprises a TEV protease or a variant of the TEV protease, or a functional fragment thereof. In some embodiments, the membrane-tethering protease comprises a sequence having at least 90% identity with SEQ ID NO: 41. In some embodiments, the membrane-tethering protease comprises SEQ ID NO: 42. In some embodiments, the transcription factor comprises a DNA-binding domain and a transcription activation domain. In some embodiments, the DNA-binding domain comprises a Gal4, PPR1, Lac9, zinc finger, or LexA DNA-binding domain. In some embodiments, the DNA-binding domain comprises a sequence having at least 90% sequence identity with any one of SEQ ID NO: 25-29. In some embodiments, the transcriptional activation domain comprises a VP64, VPr, p65, Rta, or VP16 activation domain. In some embodiments, the transcriptional activation domain comprises a sequence having at least 90% sequence identity with any of SEQ ID NO: 30-34. In some embodiments, the PMC adapter comprises a flexible amino acid adapter. In some embodiments, the PMC adapter is about 2 to about 31 amino acids in length. In some embodiments, the PMC adapter comprises a TEV-cleavable sequence. In some embodiments, the PMC adapter comprises a sequence having at least 90% sequence identity with SEQ ID NO: 20 or 21.In some embodiments, the PMC linker includes a protease cleavage site comprising at least one of SEQ ID NO: 22-24. In some embodiments, the plasma membrane polypeptide comprises rhodopsin or a variant thereof. In some embodiments, the plasma membrane polypeptide comprises a sequence having at least 90% sequence identity with any of SEQ ID NO: 1-6. In some embodiments, the plasma membrane polypeptide comprises cystic fibrosis transmembrane conduction regulator (CFTR) or a variant of CFTR. In some embodiments, the plasma membrane polypeptide comprises a sequence having at least 90% sequence identity with SEQ ID NO: 7. In some embodiments, the plasma membrane polypeptide comprises a G protein-coupled receptor. In some embodiments, the plasma membrane polypeptide comprises glucagon-like peptide-1 receptor (GLP-1R) or a variant of GLP-1R. In some embodiments, the plasma membrane polypeptide comprises a sequence having at least 90% sequence identity with SEQ ID NO: 8. In some embodiments, the plasma membrane polypeptide comprises a sequence having at least 90% sequence identity with any of SEQ ID NO: 9-19. In some embodiments, expression of the plasma membrane construct is inducible. In some embodiments, expression of the plasma membrane construct is induced by doxycycline. In some embodiments, the plasma membrane peptide is localized to the plasma membrane of eukaryotic cells after expression. In some embodiments, expression of the reporter construct indicates the ability of the assay agent to alleviate the condition. In some embodiments, the condition is a degenerative disease. In some embodiments, the eukaryotic cells are mammalian cells. In some embodiments, the mammalian cells are human cells. In some embodiments, the eukaryotic cell population comprises the system provided herein. A further method provided herein includes a method for screening assay agents, comprising contacting a eukaryotic cell population comprising the system provided herein with the assay agent. In some embodiments, the assay agent comprises a small molecule compound.
[0008] One aspect of the systems and methods described herein envisions a method for determining compliance with defective plasma membrane transport of a variant plasma membrane protein of interest in the rescue of an experimental agent, the method comprising: (a) expressing a recombinant form of the variant plasma membrane protein in a host cell and contacting the host cell with the experimental agent; (b) measuring the transport of the variant plasma membrane protein into the plasma membrane of the host cell using the system described herein; (c) comparing the transport determined in (b) with the transport in the host cell when not contacted with the experimental agent; and (d) if the transport of the plasma membrane protein in the host cell contacted with the experimental agent is increased compared with the transport in the host cell not contacted with the experimental agent, then identifying a patient who has or is susceptible to a disease associated with the variant plasma membrane protein as a candidate for treatment with the experimental agent. In some embodiments, step (d) includes identifying a patient as a candidate for treatment with the experimental agent if, in step (c), the transport in the host cell contacted with the experimental agent increases by at least 1.3 to 40 times compared with the transport in the host cell not contacted with the experimental agent. In some embodiments, step (d) includes identifying the patient as a candidate for treatment with the test agent if the transport in the host cell is at least 2% to about 100% of the non-mutated plasma membrane protein. In some embodiments, the test agent is a pharmacological school corrector. In some embodiments, the variant plasma membrane protein or the gene encoding the variant plasma membrane protein has been identified from patients who have or are susceptible to diseases associated with the variant plasma membrane protein.
[0009] Incorporation All publications, patents and patent applications mentioned in this specification are incorporated herein by reference to the extent that each individual publication, patent or patent application is specifically and individually indicated to be incorporated by reference. Attached Figure Description
[0010] The novel features of this disclosure are set forth in the appended claims. A better understanding of the features and advantages of this disclosure will be obtained by referring to the following detailed description and accompanying drawings, which illustrate illustrative embodiments utilizing the principles of this disclosure, and in the drawings: Figure 1 The illustration shows an exemplary embodiment described herein, in which a cell comprises a plasma membrane construct including plasma membrane peptides, adapters and transcription factors, and a plasma membrane proximity reporter encoded by one or more nucleic acid sequences.
[0011] Figure 2 An exemplary embodiment described herein is illustrated, in which a plasma membrane construct is expressed and transported to the plasma membrane of a cell.
[0012] Figure 3 The illustration shows an exemplary embodiment described herein, in which the plasma membrane construct reaches the plasma membrane of the cell and is localized to the plasma membrane at the polypeptide of the plasma membrane construct.
[0013] Figure 4 The exemplary embodiment described herein is illustrated, wherein the plasma membrane construct is cleaved at the linker by a plasma membrane-anchored protease, releasing a transcription factor that drives the expression of a unique molecular identifier by inducing a plasma membrane-adjacent promoter operatively coupled to the unique molecular identifier.
[0014] Figure 5 The exemplary methods described herein are illustrated, involving the localization of plasma membrane constructs, cleavage of plasma membrane constructs, expression of unique molecular identifier read counts, and expression of luciferase genes.
[0015] Figure 6 The luciferase readout data are shown, in which different cleavage sites were tested.
[0016] Figure 7 The luciferase readout data are shown, in which the promoter and plasma membrane anchor that drive the expression of the plasma membrane-anchored protease were tested.
[0017] Figure 8 The data displayed are luciferase readout data, in which the DNA-binding domains of transcription factors were tested.
[0018] Figure 9 Extensive target scans were depicted to measure the effectiveness of different test agents across numerous pathogenic autosomal dominant retinitis pigmentosa variants.
[0019] Figure 10 A deep mutation scan was depicted to measure the transport consequences of every single possible amino acid change in rhodopsin. Detailed Implementation
[0020] Although preferred embodiments of this disclosure have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided by way of example only. Many changes, modifications, and substitutions will now occur to those skilled in the art without departing from this disclosure. It should be understood that various alternatives to the embodiments of this disclosure described herein may be employed in the practice of this disclosure. The following claims define the scope of this disclosure, and methods and structures within the scope of these claims and their equivalents are thereby covered.
[0021] The use of absolute or sequential terms, such as “will,” “will not,” “should,” “should not,” “must,” “must not,” “first,” “initial,” “next,” “following,” “before,” “after,” “finally,” and “ultimately,” is not intended to limit the scope of the embodiments disclosed herein, but is intended as exemplary embodiments.
[0022] As used herein, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the / described” should also include the plural forms. Furthermore, where the terms “comprising,” “including,” “having,” “having,” “with,” or variations thereof are used in the specification and / or claims, these terms are intended to be inclusive in a manner similar to the term “comprising.”
[0023] As used in this article, the phrases “at least one,” “one or more,” and “and / or” are open-ended expressions that are both conjunction and disjunctive in operation. For example, each of the expressions “at least one of A, B, and C,” “at least one of A, B, or C,” “one or more of A, B, and C,” “one or more of A, B, or C,” and “A, B, and / or C” means A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together.
[0024] As used herein, “or” can mean “and,” “or,” “and / or,” and can be used exclusively or inclusively. For example, the term “A or B” can mean “A or B,” “A but not B,” “B but not A,” and “A and B.” In some cases, the context may dictate a specific meaning.
[0025] Any systems and methods described herein are modular and not limited to sequential steps. Therefore, terms such as “first” and “second” do not necessarily imply priority, order of importance, or order of actions.
[0026] The terms “about” or “approximate” refer to an acceptable margin of error for a particular value, as determined by a person skilled in the art, which will depend in part on how the value is measured or determined, such as limitations of the measurement system. For example, “about” can mean within one or more standard deviations based on practice for a given value. In some cases, “about” or “approximate” refers to a quantity that is 10% away from (plus or minus) a specified quantity.
[0027] Where specific values are described in this application and claims, the term “about” shall be assumed to be an acceptable range of error for the specific value unless otherwise stated.
[0028] As used herein, the terms “increased,” “gradual,” or “increase” generally refer to an increase of a statistically significant amount. In some cases, the terms “increased” or “increase” mean an increase of at least 10% compared to a reference level, such as an increase of at least about 10%, at least about 20%, or at least about 30%, or at least about 40%, or at least about 50%, or at least about 60%, or at least about 70%, or at least about 80%, or at least about 90%, or up to and including an increase of 100%, or any increase between 10% and 100%. Other examples of “increase” include increases of at least 2 times, at least 5 times, at least 10 times, at least 20 times, at least 50 times, at least 100 times, at least 1000 times, or more compared to a reference level.
[0029] As used herein, the terms “reduced,” “gradually decreasing,” or “reduction” generally refer to a statistically significant reduction. In some cases, “reduced” or “reduction” means a reduction of at least 10% compared to a reference level, such as a reduction of at least about 20%, or at least about 30%, or at least about 40%, or at least about 50%, or at least about 60%, or at least about 70%, or at least about 80%, or at least about 90%, or up to and including a reduction of 100% (e.g., no level or undetectable level compared to the reference level) or any reduction between 10% and 100%. In the context of a biomarker or symptom, these terms mean that the level is statistically significant. This reduction can be, for example, at least 10%, at least 20%, at least 30%, at least 40%, or more, and preferably a reduction to a level within the normal range acceptable to an individual without a given disease.
[0030] As used herein, “cell” generally refers to a biological cell. A cell is the basic structural, functional, and / or biological unit of a living organism. Cells can originate from any organism having one or more cells. Some non-limiting examples include: prokaryotic cells, eukaryotic cells, bacterial cells, archaea cells, cells of unicellular eukaryotes, protozoan cells, cells from plants, fungal cells (e.g., yeast cells, cells from mushrooms), animal cells, cells from invertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), or cells from mammals (e.g., pigs, cattle, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.). Sometimes cells are not derived from natural organisms (e.g., cells are artificially synthesized, sometimes called artificial cells). In some cases, cells are primary cells. In some cases, cells are derived from cell lines.
[0031] As used herein, the term "nucleotide" generally refers to a base-sugar-phosphate combination. Nucleotides include synthetic nucleotides. Nucleotides include synthetic nucleotide analogs. A nucleotide is a monomeric unit of a nucleic acid sequence, such as deoxyribonucleic acid (DNA) and ribonucleic acid (RNA). The term nucleotide may include adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleoside triphosphates such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives may include, for example, [αS]dATP, 7-denitro-dGTP, and 7-denitro-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used herein, the term nucleotide may refer to dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Illustrative examples of dideoxyribonucleoside triphosphates may include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP.
[0032] The terms "polynucleotide," "oligonucleotide," and "nucleic acid" are used interchangeably to refer to polymeric forms (deoxyribonucleotides or ribonucleotides) of any length, in single-stranded, double-stranded, or multi-stranded form, or analogs thereof. In some cases, polynucleotides are exogenous (e.g., heteropolynucleotides). In some cases, polynucleotides are endogenous within cells. In some cases, polynucleotides can exist in cell-free environments. In some cases, polynucleotides are genes or fragments thereof. In some cases, polynucleotides are DNA. In some cases, polynucleotides are RNA. Polynucleotides can have any three-dimensional structure and can perform any known or unknown function. In some cases, polynucleotides contain one or more analogs (e.g., altered backbones, sugars, or nucleobases). If present, the nucleotide structure can be modified before or after polymer assembly. Some non-limiting examples of analogues include: 5-bromouracil, peptide nucleic acid, xenonucleic acid, morpholino, locked nucleic acid, glycerol nucleic acid, threonine nucleic acid, dideoxynucleotide, cordycepin, 7-denitro-GTP, fluorophores (e.g., rhodamine or fluorescein linked to sugars), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogues, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, piracetin, and woyoside. Non-restrictive examples of polynucleotides include coding or non-coding regions of genes or gene fragments, loci defined by linkage analysis (locuses), exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), guide RNA (gRNA), microRNA (miRNA), non-coding RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. In some cases, the nucleotide sequence is interrupted by non-nucleotide components.
[0033] The terms “peptide” and “protein” are used interchangeably to refer to polymers of amino acid residues and are not limited to a minimum length. A peptide, including the provided polypeptide chain and other peptides such as linker and binding peptides, may include amino acid residues, including native and / or non-native amino acid residues. The term also includes post-expression modifications of the peptide, such as glycosylation, sialylation, acetylation, phosphorylation, etc. In some aspects, a peptide may contain modifications relative to its native or natural sequence, provided that the protein retains the desired activity. These modifications may be intentional, such as through site-directed mutagenesis, or they may be accidental, such as mutations in the host producing the protein or errors due to PCR amplification. In some cases, the peptide encodes a gene or transgene as described herein.
[0034] The terms “plasma membrane protein” or “plasma membrane polypeptide” include, but are not limited to, transmembrane polypeptides or polypeptides that are otherwise targeted to the plasma membrane (e.g., by tethering or anchoring), and include polypeptides that are close to the plasma membrane such that the linker attached to the polypeptide can be cleaved by a plasma membrane-anchored protease.
[0035] As used herein, the term "assay reagent" refers to any one or more compounds, such as small molecules, peptides, antibodies, nucleic acids, siRNAs, sequence-guided nuclease constructs, or gene constructs, introduced into the system described herein to determine the effect of the assay reagent on the reporter output of the system. Not all assay reagents can introduce detectable changes in the reporter output. Assay reagents also include environmental conditions, such as pH, temperature, and medium tension.
[0036] As used herein, the term "gene" or "transgenic" refers to a segment of nucleic acid (also called a "coding sequence" or "coding region") that encodes a single protein or RNA, optionally accompanied by an associated regulatory element (such as a promoter, operator, terminator, etc.) located upstream or downstream of the coding sequence. In some cases, the promoter is an inducible promoter. In some embodiments, the regulatory element contains at least one open reading frame (ORF) that does not encode the transgenic. Instead, the ORF in the regulatory element can upregulate the transgenic. In some embodiments, the ORF in the regulatory element can downregulate the transgenic. In some embodiments, the ORF in the regulatory element is located 5' upstream of the transgenic. In some embodiments, the ORF in the regulatory element is located 3' downstream of the transgenic. The terms "gene" or "transgenic" should be interpreted broadly and may encompass the mRNA, cDNA, and genomic DNA forms of the gene. In some uses, the term "gene" encompasses the transcribed sequence, including the 5' and 3' untranslated regions (5'-UTR and 3'-UTR), exons, and introns. In some genes, the transcribed region will contain an "open reading frame" encoding a polypeptide. In some uses of the term, "gene" or "transgenic" contains only the coding sequence necessary to encode a polypeptide (e.g., "reading frame" or "coding region"). In some respects, a gene or transgenic does not encode a polypeptide, such as ribosomal RNA genes (rRNA) and transfer RNA (tRNA) genes. In other respects, the term "gene" or "transgenic" includes not only the transcribed sequence but also additionally includes non-transcribed regions, including upstream and downstream regulatory elements such as regulatory regions, enhancers, and promoters. The term "gene" or "transgenic" can encompass the mRNA, cDNA, and genomic forms of a gene.
[0037] The term "expression" generally refers to one or more processes of transcription (such as transcription into mRNA or other RNA transcripts) of a polynucleotide from a DNA template and / or the subsequent translation of the transcribed mRNA into a peptide, polypeptide, or protein. Transcripts and encoded polypeptides can be collectively referred to as "gene products." If the polynucleotide is derived from genomic DNA, expression may include splicing of mRNA in eukaryotic cells. In some embodiments, expression may include the biological activity of a polypeptide encoded by the polynucleotide described herein. "Upregulation," in relation to expression, generally refers to an increase in the expression level of a polynucleotide (e.g., RNA, such as mRNA) and / or polypeptide sequence compared to its expression level in the wild-type state. For example, the expression of a gene or transgene may be upregulated by at least 0.1-fold, 0.2-fold, 0.3-fold, 0.4-fold, 0.5-fold, 1-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold, 20-fold, 50-fold, or more times by the system described herein compared to the expression of a gene or transgene in the wild-type state (e.g., in a system described herein that does not upregulate gene or transgene expression). "Downregulation" generally refers to a reduction in the expression level of a polynucleotide (e.g., RNA, such as mRNA) and / or polypeptide sequence compared to its expression in the wild-type state. For example, gene or transgene expression may be downregulated by the systems described herein by at least 0.1-fold, 0.2-fold, 0.3-fold, 0.4-fold, 0.5-fold, 1-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold, 20-fold, 50-fold, or more times compared to the expression of a gene or transgene in the wild-type state (e.g., in systems without upregulated gene or transgene expression). "Wild-type" or "wild-type state" can refer to the phenotypic or biological measurement or observation of expression occurring in nature without reducing the expression of the target through expression vectors or nucleic acid manipulation (e.g., expression as a product of a normal allele, rather than as a product of a mutant or engineered gene or through siRNA or CRISPR / Cas9 systems).
[0038] The percentage of sequence identity (%) relative to a reference polypeptide sequence refers to the percentage of amino acid residues in the candidate sequence that are identical to those in the reference polypeptide sequence after sequence alignment and the introduction of vacancies (if necessary) to achieve maximum sequence identity, without considering any conserved substitutions as part of the sequence identity. Alignments used to determine the percentage of amino acid sequence identity can be performed in various known ways, such as using publicly available computer software like BLAST, BLAST-2, ALIGN, or Megalign (DNASTAR) software. Appropriate parameters for sequence alignment can be determined, including the algorithms required to achieve maximum alignment across the full length of the sequences being compared.
[0039] When used in this paper to describe nucleic acid sequences, the terms “identity,” “sameness,” or “percentage of identity” compared to a reference sequence may be determined using the formula described by Karlin and Altschul (Proc. Natl. Acad. Sci. USA 87:2264-2268, 1990, modified to Proc. Natl. Acad. Sci. USA 90:5873-5877, 1993). This formula is incorporated into the Basic Local Alignment Search Tool (BLAST) procedure of Altschul et al. (J. Mol. Biol. 215: 403-410, 1990). The percentage of sequence identity may be determined using the latest version of BLAST as of the date of this application.
[0040] The polypeptides of the systems described herein can be encoded by nucleic acids. Nucleic acids are a type of polynucleotide containing two or more nucleotide bases. In some embodiments, the nucleic acid is a component of a vector that can be used to transfer the polynucleotide encoding the polypeptide into a cell. As used herein, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it is linked. One type of vector is a genome integration vector, or "integration vector," which can be integrated into the chromosomal DNA of a host cell. Another type of vector is an "attachment" vector, such as a nucleic acid capable of extrachromosomal replication. Vectors capable of directing the expression of genes operatively linked to them are referred to herein as "expression vectors." Suitable vectors include plasmids, bacterial artificial chromosomes, yeast artificial chromosomes, viral vectors, etc. In some cases, the vector contains regulatory elements for controlling transcription, such as promoters, enhancers, and polyadenylation signals. Regulatory elements can be derived from mammalian, microbial, viral, or insect genes. In embodiments, the vector includes the gene expression cassette described herein. Additional genes, such as the ability to replicate in the host typically conferred by the origin of replication and selection genes that facilitate the recognition of transformants, can be incorporated. Virus-derived vectors, such as lentiviruses, retroviruses, adenoviruses, adeno-associated viruses, etc., can be used. Plasmid vectors can be linearized to integrate into chromosomal locations. Vectors can contain sequences that guide site-specific integration into site-defined locations or restriction sets within the genome (e.g., AttP-AttB recombination). Furthermore, vectors can include sequences derived from transposon elements used for integration.
[0041] As used herein, the term "transfection" or "transfected" refers to a method of intentionally introducing exogenous nucleic acids into cells using procedures commonly used in the laboratory. Transfection can be achieved, for example, by lipid transfection, calcium phosphate precipitation, viral transduction, or electroporation. Transfection can be transient or stable.
[0042] As used herein, the term "transfection efficiency" refers to the extent or degree to which a cell population incorporates exogenous nucleic acids. Transfection efficiency can be measured as the percentage (%) of cells in a given population incorporating exogenous nucleic acids compared to the total cell population in the system. Transfection efficiency can be measured in both transiently and stably transfected cells.
[0043] As used herein, “reporter gene,” “reporter construct,” or equivalent refers to one or more genetic elements in a cell that, when expressed in the cell, can be detected using laboratory methods. Reporter genes include, but are not limited to, luciferase genes, genes encoding fluorescent proteins, genes encoding enzymes (which can act on certain substrates, resulting in a detectable signal), or unique molecular identifiers or barcode sequences. Reporter genes are generally coupled to promoter elements, response elements, or transcription factor binding sites, enabling them to be used to understand cellular signaling events.
[0044] As used herein, “reporter activity” refers to the empirical readout of a reporter. For example, a luciferase reporter will exhibit luminescent readout when incubated with a suitable substrate. Other reporters, such as fluorescent proteins, may not require a substrate but can be measured via, for example, microscopy or a fluorescent plate reader. Unique molecular identifiers, for example, can be identified by sequencing or amplification reactions. In some embodiments, the reporter is encoded by a reporter nucleic acid. In some cases, reporter expression is driven by regulatory elements or promoters as described herein. In some embodiments, reporter expression is driven by cAMP response elements such as cAMP response element binding protein (CREB).
[0045] As used herein, “heterologous” or “exogenous” can describe a component in a cell that, after introduction, may be present in or expressed by the cell, but is originally foreign to the cell. For example, heterologous or exogenous expression of a protein can occur after the introduction of complementary DNA or RNA encoding a protein of interest into a cell, thus allowing the cell to express a foreign (now heterologous) protein. Heterologous or exogenous also refers to components (e.g., nucleic acids or proteins) expressed or present at levels higher or lower than naturally occurring in the cell, and includes components in mutant and wild-type forms. Heterologous or exogenous also refers to components (e.g., nucleic acids or proteins) integrated or expressed at a cellular or genomic location different from their natural presence.
[0046] As used herein, “operably coupled” or “operably linked” components of a cell have one or more activities of the linked component, for example, a second activity of the second component occurs when a first activity of the first component occurs. Thus, in practice, when the first activity of the first component occurs, the second activity of the second component follows, making them “operably coupled” or “operably linked.” For example, for transcription factor binding sites or response elements, activation of the response element will result in the expression of a reporter gene downstream of the binding site or response element. As used herein, “variants” of peptides include peptides having an amino acid sequence different from that of peptides in human species, as shown in SEQ ID NO: 1-6. Generally, peptide variants usable in this disclosure are those with reduced ability to be localized at or transported to the cell’s plasma membrane. Generally, the amino acid sequence of peptide variants usable in this disclosure may differ from the amino acid sequence of wild-type peptides due to one or more amino acids. In some embodiments, the amino acid sequence of the peptide variant differs due to one amino acid. In some embodiments, the amino acid sequence of the peptide variant differs due to at least one amino acid. In some embodiments, the amino acid sequence of the polypeptide variant differs due to at least one amino acid, at least two amino acids, at least three amino acids, at least four amino acids, at least five amino acids, at least six amino acids, at least seven amino acids, at least eight amino acids, at least nine amino acids, or ten amino acids. In some embodiments, the amino acid sequence of the polypeptide variant differs due to one amino acid. In some embodiments, the amino acid sequence of the polypeptide variant differs due to one amino acid, two amino acids, three amino acids, four amino acids, five amino acids, six amino acids, seven amino acids, eight amino acids, nine amino acids, or ten amino acids.
[0047] Overview The systems and methods described herein can measure gene regulatory activity by assessing the expression of one or more peptides and / or the activation levels of membrane-near promoters associated with peptide activity. In this system, a model peptide (e.g., rhodopsin) acts as a proxy for intracellular peptide transport. This peptide may have one or more mutations that result in misprocessing, misfolding, or mistransport. For example, an assay agent that increases or decreases peptide expression will affect signal transduction via a membrane-near reporter, which contains a membrane-near promoter and can be detected by a reporter gene induced by transcription factors. Systems and methods for assessing the effects of an assay agent on the expression or transport of one or more peptides are also described herein. In some embodiments, the assay agent may be a pharmacologically schooled positron (e.g., a molecular chaperone). In some embodiments, the assay agent comprises a small molecule compound. The systems and methods described herein can be used to assess the effects of the assay agent on the treatment or alleviation of certain conditions or diseases, such as Menke's disease, Wilson's disease, or degenerative diseases such as retinitis pigmentosa, Alzheimer's disease, Parkinson's disease, Huntington's disease, cystic fibrosis, α-1 antitrypsinemia, or amyotrophic lateral sclerosis. The systems and methods described herein can also be used to evaluate the effects of assays on increased plasma membrane expression of downregulated peptides (e.g., glucagon-like peptide-1 receptor (GLP-1R)). The systems and methods described herein can also be used for compliance assays to assess the effects of specific mutations on plasma membrane transport or to evaluate the effects of assays on the plasma membrane transport outcomes of mutant peptides, including pathogenic mutations identified in patients with certain conditions or diseases (e.g., retinitis pigmentosa, GLP-1R downregulation, cystic fibrosis, etc.). As described herein, degenerative diseases can include inherited conditions that affect protein transport in the plasma membrane. Such degenerative diseases can be chronic.
[0048] Plasma membrane activity and related peptides This document describes systems and uses comprising plasma membrane proximity reporters for screening and identifying assays that regulate or influence intracellular transport of peptides intended for use on the plasma membrane or for secretion. In some embodiments, the systems and methods described herein relate to plasma membrane proximity reporters and their activity in response to doses of one or more assays. In particular, the systems and methods described herein relate to peptide expression or the expression of unique molecular identifiers associated with plasma membrane proximity reporters that respond to one or more assays. As described herein, the term "proximity" indicates that the expression of a reporter is related to the transport of peptides to the cell's plasma membrane.
[0049] Figure 1An example cell 102 within a cell population 100 is illustrated. In this example, cell 102 includes nucleic acids 110, 120, 130, and a protease 140. In some embodiments, cell 102 may include only one of nucleic acids 110, 120, or 130, or may include a combination thereof. In some embodiments, one or more of nucleic acids 110, 120, and 130 are exogenous to cell 102. In some embodiments, cell 102 is a eukaryotic cell. In the embodiment depicted here, protease 140 is anchored to the cell's plasma membrane.
[0050] Nucleic acid 110 further includes a doxycycline-inducible promoter 112 operatively coupled to the coding region of plasma membrane polypeptide 114. In some embodiments, the doxycycline-inducible promoter 112 may be replaced by a constitutive promoter (e.g., GAPDH or efla). In some embodiments, plasma membrane polypeptide 114 may be transported to the cell's plasma membrane as part of a plasma membrane construct after expression, the plasma membrane construct including one or more components (e.g., Figure 3 (as described above). In some embodiments, the components may include one or more plasma membrane-localizing peptides, transcription factors, and plasma membrane construct (PMC) linkers (containing, for example, DNA-binding domains and / or transcription activation domains) that connect the plasma membrane-localizing peptides and transcription factors.
[0051] In some embodiments, the PMC linker is a flexible amino acid linker. In some embodiments, the PMC linker has a length. In some embodiments, the length of the PMC linker is from about 2 amino acids to about 31 amino acids. In some embodiments, the PMC linker comprises the sequence GSENLYFQSGS (SEQ ID NO: 20) or GSENLYFQYGS (SEQ ID NO: 21). In some embodiments, the PMC linker comprises a sequence having at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 20 or 21.
[0052] In some embodiments, the PMC linker includes a cleavage site. In some embodiments, the cleavage site is protease-cleavable (e.g., a protease capable of cleaving the transcription factor from the remainder of the plasma membrane construct including polypeptide 114 at the cleavage site). In some embodiments, the cleavage site may be cleaved by a TEV protease or a variant thereof (e.g., TEV-S219V protease). In some embodiments, the cleavage site includes the sequence ENLYFQ(x) (SEQ ID NO: 22). In some embodiments, the cleavage site includes a sequence having at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 22. In some embodiments, the cleavage site includes the sequences ENLYFQS (SEQ ID NO: 23) or ENLYFQY (SEQ ID NO: 24). In some embodiments, the cleavage site comprises a sequence having at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 23 or 24.
[0053] In some embodiments, the transcription factor comprises a DNA-binding domain (DBD) and a transcription activator (also referred to herein as an "activation domain"). In some embodiments, the DNA-binding domain of the transcription factor includes a Gal4, PPR1, Lac9, zinc finger, or LexA DNA-binding domain. In some embodiments, the DBD comprises at least one of SEQ ID NO: 25-29. In some embodiments, the DBD comprises a sequence having at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of SEQ ID NO: 25-29. In some cases, the DBD binds to a synthetically produced DBD-responsive promoter (e.g., YB_tata). In some implementations, DBD binds to a sequence having approximately 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 43.
[0054] In some embodiments, the transcription activator (i.e., the transactivator) comprises an activation domain of VP64, VPr, p65, Rta, or VP16. In some embodiments, the transactivator comprises at least one of SEQ ID NO: 30-34. In some embodiments, the transactivator comprises a sequence having at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of SEQ ID NO: 30-34.
[0055] In some embodiments, peptide 114 is rhodopsin or a variant of rhodopsin (e.g., rhodopsin-P23H). In some embodiments, peptide 114 comprises a sequence having about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any of SEQ ID NO: 1-6. In some embodiments, peptide 114 is a cystic fibrosis transmembrane conduction regulator (CFTR) or a variant thereof. In some embodiments, peptide 114 comprises a sequence having about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 7. In some embodiments, peptide 114 is a G protein-coupled receptor (e.g., GLP-1R) or a variant of a G protein-coupled receptor. In some embodiments, peptide 114 comprises a sequence having about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any of SEQ ID NO: 8-19. The plasma membrane construct comprising peptide 114 may be a variant of a wild-type plasma membrane peptide. In some embodiments, the plasma membrane construct comprising peptide 114 comprises a variant or mutant of any of SEQ ID NO: 1-19.
[0056] Nucleic acid 120 further includes a plasma membrane-adjacent reporter, which contains a plasma membrane-adjacent promoter 122 that can be bound and activated by a transcription factor linked to a plasma membrane polypeptide 114 operatively coupled to the reporter gene. In the embodiments depicted herein, the reporter gene includes a luciferase gene 124 and a unique molecular identifier 126. As described herein, the unique molecular identifier may be referred to as a “UMI” or a “barcode.” In some embodiments, the reporter gene, and consequently the nucleic acid 120, does not include the luciferase gene 124 and the plasma membrane-adjacent promoter is operatively coupled only to the unique molecular identifier 126. In some embodiments, the reporter gene encodes a fluorescent protein instead of the luciferase gene 124. In some embodiments, the reporter gene encodes both a fluorescent protein and the luciferase gene 124. In some embodiments, the plasma membrane-adjacent promoter 122 may be an inducible promoter.
[0057] In some embodiments, the plasma membrane-proximity promoter comprises a synthetic DBD-responsive promoter (e.g., YB_tata). In some embodiments, the plasma membrane-proximity promoter comprises SEQ ID NO: 43. In some embodiments, the plasma membrane-proximity promoter comprises a sequence having at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 43. In some embodiments, the plasma membrane-proximity promoter comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 zinc finger binding sites. In some embodiments, the plasma membrane-proximity promoter comprises about 1 to 12 zinc finger binding sites. In some embodiments, the plasma membrane-proximity promoter is configured to bind a sequence comprising SEQ ID NO: 25. In some embodiments, the plasma membrane-adjacent promoter is configured to bind a sequence having at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 25.
[0058] Nucleic acid 130 encodes a plasma membrane-anchored protease. Nucleic acid 130 further includes a promoter 132 operatively coupled to a gene encoding a plasma membrane anchor 134 and a protease 136. In some embodiments, promoter 132 may be a constitutive promoter (e.g., GAPDH or ef1a) or an inducible promoter. In some embodiments, when promoter 132 is activated, plasma membrane anchor 134 and protease 136 may be expressed. In some embodiments, the expressed protease is a protease of the same type as protease 140. In some embodiments, the expressed plasma membrane anchor 134 may anchor the expressed protease to the plasma membrane. In some embodiments, the expressed plasma membrane anchor 134 comprises a transmembrane or plasma membrane localization domain. In some embodiments, plasma membrane anchor 134 encodes a platelet-derived growth factor receptor, Lyn (e.g., Lyn11), or a Plekluck substrate protein homology (PH) domain polypeptide or a variant thereof (e.g., PHDelta1, PHDelta3). In some embodiments, the plasma membrane anchor comprises a sequence having about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any of SEQ ID NO: 35-38. In some embodiments, nucleic acid 130 further comprises a sequence encoding a protease tether between plasma membrane anchor 134 and protease 136. The protease tether may comprise a flexible amino acid linker. In some embodiments, the protease tether is about 2 to about 31 amino acids in length. In some embodiments, the protease tether comprises the sequence ASPSNPGASNGS (SEQ ID NO: 39) or GGGGSGGGGS (SEQ ID NO: 40). In some embodiments, the protease tether comprises a sequence having at least 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 39 or 40. In some embodiments, protease 136 comprises a sequence having about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 41 or 42. When targeted or localized to the plasma membrane, the protease can cleave the linker (e.g., the membrane-localized polypeptide that binds to the transcription factor) at the cleavage site. Figure 3-4 (as further described).
[0059] In some embodiments, the cells may be present in an environment with increased doxycycline concentrations. In some embodiments, doxycycline may be introduced into the cells, thereby increasing the activation rate of doxycycline-inducible promoter 112 and increasing the expression of plasma membrane polypeptide 114 bound to the doxycycline-inducible promoter. In some embodiments, doxycycline-inducible promoter 112 comprises a sequence having approximately 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 44. In some embodiments, upon introduction of doxycycline, the expression of plasma membrane polypeptide 114 increases within a cell population 100, resulting in the plasma membrane polypeptide 114 reaching the cell's plasma membrane. In some embodiments, the plasma membrane peptide may not reach the plasma membrane or may reach the plasma membrane at a reduced rate compared to the wild-type version in wild-type cells (e.g., the plasma membrane peptide contains one or more variants that disrupt folding, processing, or transport). In some embodiments, the presence of the assay reagent within cell 102 or contact between the assay reagent and cell 102 may result in an increased ability of the expressed plasma membrane peptide 114 to localize to the plasma membrane of cell 102.
[0060] In the example depicted herein, cell 102 further includes protease 140. In the embodiments depicted herein, protease 140 is anchored to the cell's plasma membrane. In some embodiments, the protease is capable of cleaving a linker of a plasma membrane construct including plasma membrane polypeptide 114. In some embodiments, protease 140 is a component of the cell's plasma membrane. For example, the protease may enter the cell via a cellular secretory pathway or may reside within a cellular secretory pathway. In some embodiments, protease 140 includes a transmembrane domain. In some embodiments, protease 140 is anchored or located near the plasma membrane via a plasma membrane anchor. In some embodiments, the plasma membrane anchor includes a domain derived from platelet-derived growth factor receptor, Lyn (e.g., Lyn11), Plek substrate protein homology domain (e.g., PHDelta1 or PHDelta3 polypeptide), or a variant thereof. In some embodiments, protease 140 is anchored or located near the plasma membrane via platelet-derived growth factor receptor, Lyn (e.g., Lyn11), PHDelta1 or PHDelta3 polypeptide, a functional fragment thereof, or a variant thereof. In some embodiments, the plasma membrane anchor comprises a sequence having about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any of SEQ ID NO: 35-38. In some embodiments, protease 140 comprises tobacco etch virus, aspartic acid, glutamic acid, metalloid, cysteine, serine, or threonine protease. In some embodiments, protease 140 comprises a sequence having about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any of SEQ ID NO: 41 or 42. Optional protease tethers between the plasma membrane anchor and the protease may be included. In some embodiments, the protease tether comprises the sequence ASPSNPGASNGS (SEQ ID NO: 39) or GGGGSGGGGS (SEQ ID NO: 40). In some embodiments, the protease tether comprises a sequence having about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 39 or 40.
[0061] As described above, although in the embodiments depicted herein, nucleic acid 110 comprises the coding regions of a doxycycline-inducible promoter 112 and a plasma membrane polypeptide 114, nucleic acid 120 comprises a plasma membrane-adjacent promoter 122 operatively coupled to a luciferase gene 124 and / or UMI 126, and nucleic acid 130 comprises a promoter 132 operatively coupled to a gene encoding a plasma membrane anchor 134 and a protease 136. Other embodiments of nucleic acids 110, 120, and 130 may have different constructs or may be absent. For example, in some embodiments, nucleic acid 120 may include the luciferase gene 124 but not UMI 126. In other embodiments, nucleic acid 120 may include UMI 126 but may not include the luciferase gene 124. In some embodiments, nucleic acid 110 may be absent. In some embodiments, nucleic acid 130 may be absent. In some embodiments, both nucleic acids 110 and 130 may be absent. In some embodiments, different reporter genes, such as β-galactosidase, β-lactamase, alkaline phosphatase, RFP (red fluorescent protein), and / or GFP (green fluorescent protein), may be operatively coupled to a membrane-adjacent promoter 122 instead of a luciferase gene 124. While certain constructs have been described above, these constructs are exemplary and other constructs may be used. Furthermore, in other embodiments, various features of nucleic acids 110, 120, and 130 from the examples depicted herein (e.g., the coding regions of the doxycycline-inducible promoter 112 and / or the membrane polypeptide 114 of nucleic acid 110, the membrane-adjacent promoter 122 of nucleic acid 120, the luciferase gene 124 and / or UMI 126, and the promoter 132, membrane anchor 134, and protease 136 of nucleic acid 130) may be included on multiple nucleic acids instead of a single nucleic acid (e.g., each of nucleic acids 110, 120, and / or 130). Alternatively, in other embodiments, various features of nucleic acids 110, 120, and 130 may be included on a single nucleic acid (e.g., the coding regions of the doxycycline-inducible promoter 112 and / or the plasma membrane construct 114 of nucleic acid 110, the plasma membrane-adjacent promoter 122 of nucleic acid 120, the luciferase gene 124 and / or UMI 126, and the promoter 132, the plasma membrane anchor 134, and the protease 136 of nucleic acid 130 are all included on a single nucleic acid).
[0062] In some embodiments, the doxycycline promoter 112 can be induced by a dose of doxycycline. In some embodiments, the plasma membrane-adjacent promoter 122 can be induced by a transcription factor cleaved from plasma membrane polypeptide 114.
[0063] Figure 2An example cell 102 expressing plasma membrane polypeptide 114 is illustrated. In the embodiment depicted here, cell 102 has been contacted with doxycycline. In other embodiments, doxycycline may not be introduced into cell 102. When doxycycline is present in the cell population, the presence of the doxycycline-inducible promoter 112 in the nucleic acids 102 of the cells in cell population 100 leads to increased expression of plasma membrane polypeptide 114, which in turn can lead to increased presence of plasma membrane polypeptide 114 at the plasma membrane of cell 102. In embodiments without the doxycycline-inducible promoter 112, plasma membrane polypeptide 114 may still be expressed, but the expression of plasma membrane polypeptide 114 may be reduced compared to cells containing the doxycycline-inducible promoter 112. Therefore, in this depicted embodiment where cell 102 has been contacted with doxycycline, the doxycycline-inducible promoter 112 is induced and plasma membrane polypeptide 114 is expressed.
[0064] In some embodiments, expression of plasma membrane peptide 114 is associated with one or more diseases, such as retinitis pigmentosa, Alzheimer's disease, Parkinson's disease, Huntington's disease, cystic fibrosis, α-1 antitrypsinemia, or amyotrophic lateral sclerosis. Specifically, one or more diseases can occur when the plasma membrane of a human cell (e.g., cell 102) is defective, including one or more peptides comprising plasma membrane peptide 114. In some cases, even when the peptide is present in a human cell, the peptide does not translocate to the cell's plasma membrane due to defects in the peptide or peptide variants or defects in the cell's associated transport pathways. In some cases, even when plasma membrane peptide 114 is expressed, plasma membrane peptide 114 comprising the peptide may not reach the cell's plasma membrane (e.g., due to mutations in the plasma membrane peptide or defects in the cellular processing pathways of the plasma membrane peptide).
[0065] Therefore, in some embodiments, the test reagent may be present in the cell to facilitate the transport of plasma membrane peptide 114 to the plasma membrane of cell 102. In some embodiments, the test reagent may be exogenously added to cell 102. In some embodiments, the test reagent may contact cell population 100. In some embodiments, the presence of the test reagent in cell 102 may increase the rate at which plasma membrane construct 114 is transported to the cell's plasma membrane. In some embodiments, the presence of the test reagent in cell 102 may increase the rate at which one or more peptides are present at the cell's plasma membrane after expression.
[0066] Therefore, by both expressing plasma membrane peptide 114 and increasing the rate at which plasma membrane peptide 114 reaches the cell's plasma membrane, the systems and methods described above can be used to identify test agents that can be used to treat diseases associated with misfolded or mistransported peptides.
[0067] Figure 3The illustration shows an example cell 102 in which a plasma membrane construct 150 containing an expressed plasma membrane polypeptide 152 (e.g., expressed by nucleic acid 110) has reached the plasma membrane of cell 102. In the example depicted here, as described above, the expressed plasma membrane construct 150 includes the plasma membrane polypeptide 152, a linker 154, and a transcription factor containing a DNA-binding domain 156 and a transcription activator 158. In some embodiments, polypeptide 152 is rhodopsin or a variant of rhodopsin. In some embodiments, the rhodopsin variant may be a RhoP23H variant, an hRHO variant, a p.G90D variant, a p.T941 variant, a p.E113K variant, a p.A292E variant, a p.A295V variant, or other variants of rhodopsin.
[0068] In the embodiment depicted herein, linker 154 includes a protease-cleavable cleavage site. Furthermore, protease 140 is capable of cleaving the linker at the cleavage site. Upon cleavage of the linker at the cleavage site, the transcription factor is released into the cell, while peptide 152 remains at the cell's plasma membrane. Following cleavage of linker 154, peptide 152 of the plasma membrane construct 150 can remain at the plasma membrane of cell 102.
[0069] Figure 4 The illustration shows an example cell 102, in which a transcription factor, including a DNA-binding domain 156 and a transcription activator 158, binds to a nucleic acid 120, while a polypeptide 152 remains at the plasma membrane of cell 102 after cleavage by a linker 154. (See above regarding...) Figure 3 The linker 154 is cleaved, allowing the polypeptide 152 to remain localized at the cell membrane, while the DNA-binding domain 156 and transcription activator 158 are released into the cell 102. In the example depicted here, the DNA-binding domain 156 and transcription activator 158 bind to nucleic acid 120 at a site near promoter 122 on the plasma membrane and activate transcription. In some embodiments, the DNA-binding domain 156 and transcription activator 158 form a chimeric transcription factor (e.g., derived from different naturally occurring transcription factors). In some embodiments, the DNA-binding domain 156 and transcription activator 158 are synthetic transcription factors (e.g., the DNA-binding domain and / or transcription activator contain non-naturally occurring sequences).
[0070] When the DNA-binding domain 156 binds to nucleic acid 120 at a promoter 122 near the plasma membrane, expression of one or more components of nucleic acid 120 is induced. In some embodiments, expression of UMI 126 is induced. In some embodiments, expression of luciferase gene 124 is induced. In some embodiments, expression of both UMI 126 and luciferase gene 124 is induced. Following expression of UMI 126 and / or luciferase gene 124, read counts of UMI 126 or luciferase gene 124 for a plurality of cells 100 can be determined (i.e., by sequencing UMI in a downstream process).
[0071] Therefore, as through Figure 1-4 The process described above, involving similar cells in cells 102 and multiple cells 100, helps determine how these diseases or conditions can be alleviated through the use of test reagents. For example, encoding... Figure 1-4 plasma membrane peptide 114 and Figure 3 Genes of the plasma membrane construct 150 include polypeptides that are prone to misfolding or mistransport (e.g., polypeptide 152) and may be associated with certain diseases or conditions (e.g., rhodopsin, which is associated with degenerative diseases such as retinitis pigmentosa). In many cases, certain diseases or conditions are caused by polypeptides that cannot reach the cell's plasma membrane. Other polypeptides may be encoded by genes encoding plasma membrane polypeptide 114, including, for example, cystic fibrosis transmembrane transport regulators or α-1 antitrypsin.
[0072] However, when an assay agent includes a peptide that increases plasma membrane transport, the expressed peptide has an increased ability to reach the cell's plasma membrane. This increased ability of the assay agent to transport the plasma membrane peptide can then be measured by counting reads of the expressed UMI or the activity of the expressed luciferase gene in multiple cells (e.g., ). For example, after the plasma membrane peptide reaches the cell's plasma membrane, a plasma membrane construct containing the peptide can then be used at a connector (e.g., ). Figure 3 The plasma membrane construct linker (154) is cleaved at the plasma membrane by a protease anchored to the cell membrane, which may be capable of cleaving the linker. In some embodiments, the cell may express the protease and / or the plasma membrane anchor to endogenously localize at the plasma membrane. In other embodiments, the protease may be supplied on an exogenous nucleic acid, and upon cleavage as depicted, the polypeptide may remain localized at the plasma membrane while a transcription factor comprising the DNA-binding domain of the plasma membrane construct and a transcription activator is released into the cell.
[0073] Transcription factors can then bind to promoters near the plasma membrane, which in turn induces the expression of UMI and / or luciferase genes by activating their domains. For example, the read count of UMI in multiple cells indicates the level of UMI expression, which is related to the activation of reporter genes via transcription factors in multiple cells. Furthermore, the activation of reporter genes (e.g., UMI, luciferase, fluorescent proteins, etc.) in multiple cells is related to the rate at which plasma membrane-anchored proteases cleave plasma membrane constructs, including plasma membrane peptides, after they have localized to the cell's plasma membrane. The rate at which plasma membrane-anchored proteases cleave plasma membrane constructs further indicates the rate at which plasma membrane peptides localize to the cell's plasma membrane (where peptides alleviating certain diseases or conditions can bind).
[0074] Therefore, by measuring UMI read counts in downstream sequencing assays, the rate at which plasma membrane peptides localize to the plasma membrane can be measured, indicating the ability of an assay agent or test condition applied to cells to alleviate or prevent certain diseases or conditions. Furthermore, multiple cells can be contacted with an assay agent to determine the agent's ability to increase or decrease the transport of plasma membrane constructs including peptides, and thus determine whether the agent can help alleviate or prevent certain diseases or conditions. By running multiple assays with different assay agents contacting multiple cells, the ideal assay agent for alleviating or preventing certain diseases or conditions can then be determined.
[0075] Figure 5 The illustration shows an example process 500 displaying steps 502, 512, 522, and 532, which indicates the localization rate 514 of the plasma membrane construct (e.g., Figure 1-2 The rate at which the expressed plasma membrane polypeptide 114 localizes to the cell's plasma membrane), and plasma membrane constructs (e.g., which may contain...) Figure 3 peptide 152 Figure 3 The plasma membrane polypeptide 150) is generated by an anchored protease (e.g., Figure 1-4 The rate of cleavage by the anchoring membrane protease 140, UMI (e.g., Figure 1-4 The expression of UMI 126) 536 and / or luciferase gene (e.g., Figure 1-4 The relationship between the expression of the luciferase gene 124 and 534.
[0076] For example, the process begins at step 502, where the promoter expresses a plasma membrane peptide. In some embodiments, the plasma membrane peptide is localized to the plasma membrane after expression (e.g., Figure 3-4 Peptide 152 localized to the plasma membrane. The expressed plasma membrane peptide may be part of a plasma membrane construct that further includes a linker with cleavage sites (e.g., Figure 3 The plasma membrane construct linker 154), and transcription factors including DNA-binding domains (e.g., Figure 3-4DNA-binding domain 156) and / or transcription activators (e.g., Figure 3-4 The transcription activator 158). In some embodiments, the promoter may be inducible or constitutive. If the promoter is inducible, it may be a doxycycline-inducible promoter (e.g., Figure 1-4 (doxycycline-inducible promoter 112). In some embodiments, the doxycycline-inducible promoter can be induced by introducing a certain dose of doxycycline into multiple cells, wherein the cells of the multiple cells include nucleic acids (e.g., nucleic acid 110) that include a doxycycline-inducible promoter operatively coupled to a plasma membrane construct.
[0077] In step 512, the expressed plasma membrane polypeptide is localized to the cell's plasma membrane at a rate 514. In some embodiments, the polypeptide is rhodopsin or a variant of rhodopsin. In some embodiments, the localization of the polypeptide to the cell's plasma membrane is associated with alleviating or preventing certain diseases (e.g., by restoring some or all of the activity of the plasma membrane-localized polypeptide). In some embodiments, the localization rate 514 indicates the ability to alleviate or prevent certain diseases. In some embodiments, a test reagent is added to multiple cells. In these embodiments, the test reagent can affect the ability of the plasma membrane polypeptide to localize to the plasma membrane. In some embodiments, multiple cells can be in wells. In some embodiments, the wells can be part of a well set, such as as part of a multiwell plate or other container comprising multiple partitions. In some embodiments, each well in the well set can have multiple cells that are the same as or similar to the multiple cells described above. In some embodiments, a different test reagent can be added to each well.
[0078] In step 522, the plasma membrane construct is cleaved by a protease at a cleavage rate 524. In some embodiments, the plasma membrane construct is cleaved by a protease anchored to the plasma membrane (e.g., Figure 1-4 The protease 140 cleaves the peptide. In some embodiments, the protease cleaves the plasma membrane construct at the linker. In some embodiments, the peptide remains localized to the plasma membrane. In some embodiments, the transcription factor of the plasma membrane construct is released into the cell. In the embodiments depicted herein, the cleavage rate 524 is equal to the binding rate multiplied by a coefficient "a".
[0079] In step 532, the transcription factor binds to a nucleic acid containing a promoter operatively coupled to a reporter gene (e.g., UMI, luciferase, etc.) in the cell. In some embodiments, the nucleic acid includes the UMI and / or luciferase gene. In some embodiments, a DNA-binding domain binds to the promoter of the nucleic acid (e.g., a plasma membrane-adjacent promoter 122). In some embodiments, the promoter is activated upon transcription factor binding, which induces expression of the UMI and / or luciferase gene, resulting in a UMI read count 536 (as determined by downstream sequencing) and / or an expression rate 534 of the luciferase gene (as determined by measuring the enzymatic activity of luciferase in the cell). In some embodiments, the UMI read count 536 is equal to the cleavage rate 524 multiplied by a coefficient "b". Therefore, the UMI read count 536 is further equal to the binding rate 514 multiplied by both coefficients "b" and "a". In some embodiments, the expression rate 534 is equal to the cleavage rate 524 multiplied by a coefficient "c". Therefore, the expression rate 534 is further equal to the binding rate 514 multiplied by the coefficients "c" and "a". In some embodiments, the coefficient "b" is equal to the coefficient "c". In some embodiments, the coefficient "b" is different from the coefficient "c".
[0080] Therefore, by determining the relationship between UMI expression 536 and expression rate 534 at the end of the process using one or more next-generation sequencing technologies, the localization rate 514 of plasma membrane peptides can be determined or inferred, indicating the ability to transport plasma membrane peptides to the cell's plasma membrane and thus alleviate or prevent certain diseases or conditions. Furthermore, by using multiple assays with pore sets, the localization rate 514 of plasma membrane peptides associated with each assay can be compared, and therefore, the assays can be compared based on their ability to assist in alleviating or preventing certain diseases or conditions.
[0081] barcode Variable nucleotide sequences (unique molecular identifiers, also referred to herein as barcodes) acting as indexes can be included on any reporter genes described herein. Furthermore, barcodes can be added in separate library preparation reactions. The variable nucleotide sequences described herein can be used as sample indexes to deconvolve the results obtained from the sequencing reactions used herein.
[0082] Once the contents of a cell are released into their respective compartments by a lysis agent, the macromolecular components contained therein (e.g., macromolecular components of the sample, such as RNA, DNA, or proteins) can be further processed within the compartments. According to the methods and systems described herein, unique identifiers can be provided for the macromolecular component contents of individual samples, allowing them to be attributed to the same sample or particle when characterizing these macromolecular components. By specifically assigning unique identifiers to individual samples or groups of samples, the ability to attribute characteristics to individual samples or groups of samples is provided. Unique identifiers, for example in the form of nucleic acid barcodes, can be assigned or associated to individual samples or groups of samples to tag or label the macromolecular components of the sample (and, consequently, their characteristics) with unique identifiers. These unique identifiers can then be used to attribute the components and characteristics of the sample to individual samples or groups of samples.
[0083] In some aspects, this is done by co-partitioning individual samples or groups of samples with a unique identifier or a barcode containing a unique molecular identifier (UMI) sequence. In some aspects, the unique identifier is provided in the form of nucleic acid molecules (e.g., oligonucleotides) containing nucleic acid barcode sequences that can be attached to or otherwise associated with the nucleic acid contents of an individual sample, or attached to other components of the sample, and particularly to fragments of those nucleic acids. Nucleic acid molecules are partitioned such that the nucleic acid barcode sequences contained therein are identical among nucleic acid molecules in a given partition, but as different partitions, nucleic acid molecules can and do have different barcode sequences, or at least represent a large number of different barcode sequences across all partitions in a given analysis. In some aspects, only one nucleic acid barcode sequence can be associated with a given partition, although in some embodiments, two or more different barcode sequences can be present.
[0084] Nucleic acid barcode sequences may comprise about 6 to about 20 or more nucleotides (e.g., oligonucleotides) within a nucleic acid molecule sequence. Nucleic acid barcode sequences may comprise about 6 to about 20, 30, 40, 50, 60, 70, 80, 90, 100 or more nucleotides. In some embodiments, the length of the barcode sequence may be about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or longer. In some embodiments, the length of the barcode sequence may be at least about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or longer. In some embodiments, the length of the barcode sequence may be at most about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or shorter. These nucleotides can be completely continuous, i.e., within a single segment of adjacent nucleotides, or they can be separated into two or more dividing subsequences separated by one or more nucleotides. In some embodiments, the length of the dividing barcode subsequences can be from about 4 to about 16 nucleotides. In some embodiments, the barcode subsequences can be about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or longer. In some embodiments, the barcode subsequences can be at least about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or longer. In some embodiments, the barcode subsequences can be at most about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 nucleotides or shorter.
[0085] The co-partitioned nucleic acid molecules may also contain other functional sequences that can be used to process nucleic acids from the co-partitioned sample. These sequences include, for example, targeted or random / universal amplification primer sequences for simultaneously attaching associated barcode sequences to genomic DNA from individual samples within the partition, sequencing primers or primer recognition sites, hybridization or probe sequences, such as for identifying the presence of sequences or for pulling down barcode-coded nucleic acids, or any of many other potential functional sequences. Other mechanisms for co-partitioned oligonucleotides may also be employed, including, for example, the merging of two or more partitions, one of which contains oligonucleotides, or the microdispersion of oligonucleotides into partitions, such as partitions within a microfluidic system. In some embodiments, the primers contain barcode oligonucleotides. In some embodiments, the primer sequence is a targeted primer sequence complementary to a sequence in the template nucleic acid molecule. In some embodiments, the first nucleic acid molecule further contains one or more functional sequences, and the second nucleic acid molecule contains one or more functional sequences. In some embodiments, the one or more functional sequences are selected from adaptor sequences, additional primer sequences, primer annealing sequences, sequencing primer sequences, sequences configured to attach to a flow cell of a sequencer, and unique molecular identifier sequences.
[0086] For example, the barcoded nucleic acid molecules described above (e.g., barcoded oligonucleotides) are added to a sample. In some embodiments, partitions contain barcoded oligonucleotides having the same barcode sequence. In some embodiments, partitions within a plurality of partitions contain barcoded oligonucleotides having the same barcode sequence, wherein each partition within the plurality of partitions contains a unique barcode sequence. In some embodiments, the barcoded oligonucleotide population provides a diverse library of barcode sequences comprising at least about 1,000 different barcode sequences, at least about 5,000 different barcode sequences, at least about 10,000 different barcode sequences, at least about 50,000 different barcode sequences, at least about 100,000 different barcode sequences, at least about 1,000,000 different barcode sequences, at least about 5,000,000 different barcode sequences, or at least about 10,000,000 different barcode sequences, or more. Furthermore, each barcoded oligonucleotide can be provided using a large number of attached nucleic acid (e.g., oligonucleotide) molecules. Specifically, the number of nucleic acid molecules comprising the barcode sequence on a single barcoded oligonucleotide can be at least about 1,000 nucleic acid molecules, at least about 5,000 nucleic acid molecules, at least about 10,000 nucleic acid molecules, at least about 50,000 nucleic acid molecules, at least about 100,000 nucleic acid molecules, at least about 500,000 nucleic acid molecules, at least about 1,000,000 nucleic acid molecules, at least about 5,000,000 nucleic acid molecules, at least about 10,000,000 nucleic acid molecules, at least about 50,000,000 nucleic acid molecules, at least about 100,000,000 nucleic acid molecules, at least about 250,000,000 nucleic acid molecules, and in some embodiments at least about 1 billion nucleic acid molecules or more. A given set of barcoded oligonucleotides may include the same (or common) barcode sequence, different barcode sequences, or a combination of both. A given set of barcoded oligonucleotides may include multiple sets of nucleic acid molecules. A given set of nucleic acid molecules may include the same barcode sequence. The same barcode sequence may be different from the barcode sequence of another set of nucleic acid molecules.
[0087] Furthermore, when the population of barcoded oligonucleotides is partitioned, the resulting partitioned population may also include a diverse barcoded library comprising at least about 1,000 different barcoded sequences, at least about 5,000 different barcoded sequences, at least about 10,000 different barcoded sequences, at least about 50,000 different barcoded sequences, at least about 100,000 different barcoded sequences, at least about 1,000,000 different barcoded sequences, at least about 5,000,000 different barcoded sequences, or at least about 10,000,000 different barcoded sequences. In addition, each partition of the population may include at least about 1,000 nucleic acid molecules, at least about 5,000 nucleic acid molecules, at least about 10,000 nucleic acid molecules, at least about 50,000 nucleic acid molecules, at least about 100,000 nucleic acid molecules, at least about 500,000 nucleic acid molecules, at least about 1,000,000 nucleic acid molecules, at least about 5,000,000 nucleic acid molecules, at least about 10,000,000 nucleic acid molecules, at least about 50,000,000 nucleic acid molecules, at least about 100,000,000 nucleic acid molecules, at least about 250,000,000 nucleic acid molecules, and in some implementations at least about 1 billion nucleic acid molecules.
[0088] In some implementations, it may be desirable to incorporate multiple different barcodes within a given partition. For example, in some implementations, the barcoded oligonucleotides within a partition may include (1) a common barcode sequence shared by all barcoded oligonucleotides within the partition and (2) a unique molecular identifier or additional barcode sequence that differs within each barcoded oligonucleotide. The common barcode sequence can provide greater assurance of identification in subsequent processing, for example, by providing a stronger address or attribute of the barcode to a given partition as a duplicate or independent confirmation of the output from the given partition.
[0089] In some embodiments, barcoded oligonucleotides are attached to beads, wherein all nucleic acid molecules attached to a particular bead will include the same nucleic acid barcode sequence, but the group of beads used represents a large variety of barcode sequences. In some embodiments, hydrogel beads, for example comprising a polyacrylamide polymer matrix, serve as a solid support and delivery medium for nucleic acid molecule entry into the partition, because they are capable of carrying a large number of nucleic acid molecules and can be configured to release these nucleic acid molecules upon exposure to a specific stimulus, as described elsewhere herein.
[0090] Nucleic acid molecules (e.g., oligonucleotides) can be released from the beads upon application of a specific stimulus. In some embodiments, the stimulus may be a photostimulation, for example, by cleaving photoinstantaneous linkages that release the nucleic acid molecules. In other embodiments, a thermal stimulus may be used, wherein an increase in the ambient temperature of the beads will result in the cleavage of linkages or other release of the nucleic acid molecules forming the beads. In yet another embodiment, a chemical stimulus may be used to cleave the linkages of the nucleic acid molecules to the beads or otherwise cause the nucleic acid molecules to be released from the beads. In one embodiment, such a composition comprises the polyacrylamide matrix described above for encapsulating samples and can be degraded by exposure to a reducing agent (such as DTT) to release the attached nucleic acid molecules.
[0091] The supports that can be used in the methods disclosed herein may be, for example, pores, substrates, rods, containers, or beads. Supports can have any useful characteristics and properties, such as any useful size, surface chemistry, flowability, robustness, density, porosity, and composition. In some embodiments, the support is the surface of pores in a plate. In some embodiments, the support can be beads, such as gel beads. Beads can be solid or semi-solid. Additional details about beads are provided elsewhere herein.
[0092] Supports (e.g., beads) may contain anchor sequences (e.g., as described herein) for their functionalization. Anchor sequences may be attached to the support via, for example, disulfide bonds. Anchor sequences may contain partial read sequences and / or flow cell functional sequences. Such sequences can allow sequencing of nucleic acid molecules to which the sequences are attached using a sequencer (e.g., an Illumina sequencer). Different anchor sequences can be used for different sequencing applications. Anchor sequences may contain, for example, TruSeq or Nextera sequences. Anchor sequences can have any useful characteristics, such as any useful length and nucleotide composition. For example, anchor sequences may contain 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides. In some embodiments, anchor sequences may contain 15 nucleotides. The nucleotides of the anchor sequences may be naturally occurring or non-naturally occurring (e.g., as described herein). Beads may contain multiple anchor sequences attached thereto. For example, beads may contain multiple first anchor sequences attached thereto. In some implementations, the bead may contain two or more distinct anchor sequences attached thereto. For example, the bead may contain both multiple first anchor sequences (e.g., Nextera sequences) and multiple second anchor sequences (e.g., TruSeq sequences) attached thereto. For a bead containing two or more distinct anchor sequences, the sequence of each distinct anchor sequence may be distinguishable from the sequence of each other anchor sequence at the distal end of the bead. For example, the distinct anchor sequences may contain one or more nucleotide differences in the 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, 10th, or more nucleotides furthest from the bead.
[0093] In some embodiments, multiple different barcode molecules (e.g., nucleic acid barcode molecules) can be generated on the same support (e.g., beads). For example, two different barcode molecules can be generated on the same support. Alternatively, three or more different barcode molecules can be generated on the same support. Different barcode molecules attached to the same support can contain one or more different sequences. For example, different barcode molecules can contain one or more different barcode sequences and / or other sequences (e.g., start sequences). In some embodiments, different barcode molecules attached to the same support can contain the same barcode sequence. Different barcode molecules attached to the same support can contain the same or different barcode sequences. Similarly, different barcode molecules can contain the same or different unique molecular identifiers (UMIs).
[0094] cell Cells that can be used in the systems and methods described herein are generally those that can be readily transgenic using one or more of the nucleic acids described herein. Methods known in the art, such as calcium phosphate transfection, lipid-based transfection (e.g., Lipofectamine), are employed. TM Lipofectamine-2000 TM Lipofectamine-3000 TM or Fugene HD), electroporation, or viral transduction can be used to transfect or transduce systemic nucleic acids encoding regulatory elements, effectors, and / or reporter elements into suitable cell lines. Cells can also be populations of the same type that have grown to confluence or near confluence in appropriate tissue culture containers.
[0095] In some embodiments, the cells used herein contain a stable integration of nucleic acids encoding regulatory elements, nucleic acids encoding effectors, nucleic acids containing reporter elements, or combinations thereof. Stable cell lines can be prepared from the cells described herein by means of random integration of linearized plasmids, directed or targeted integration of viruses or transposons, for example, using site-specific recombination between AttP and AttB sites. In some embodiments, either nucleic acid is integrated at a secure landing site such as the AAVS1 site.
[0096] In some embodiments, the cells described herein contain nucleic acids stably integrated into the cell genome. In some embodiments, the cells described herein contain nucleic acids encoding regulatory elements stably integrated into the cell genome. In some embodiments, the cells described herein contain nucleic acids encoding at least one effector described herein stably integrated into the cell genome. In some embodiments, the cells contain stably integrated nucleic acids encoding regulatory elements for regulating effector expression. In some cases, the cells contain stably integrated nucleic acids encoding regulatory elements for upregulating the effector described herein. In some cases, the cells contain stably integrated nucleic acids encoding regulatory elements for upregulating ADCY6. In some cases, the cells contain stably integrated nucleic acids encoding ADCY6. In some cases, the cells contain stably integrated nucleic acids encoding regulatory elements for downregulating the effector described herein. In some cases, the cells contain stably integrated nucleic acids encoding regulatory elements for downregulating ADCY3.
[0097] In some embodiments, the cells or cell populations used in the system are eukaryotic cells. In some embodiments, the cells or cell populations are mammalian cells. In some embodiments, the cells or cell populations are human cells. In some embodiments, the cells or cell populations are SH-SY5Y, human neuroblastoma; Hep G2, Caucasian human hepatocellular carcinoma; 293 (also known as HEK 293), human embryonic kidney; RAW 264.7, mouse mononuclear macrophages; HeLa, human cervical epithelioid carcinoma; MRC-5 (PD 19), human fetal lung; A2780, human ovarian cancer; CACO-2, Caucasian human colonic adenocarcinoma; THP1, human monocytic leukemia; A549, Caucasian human lung cancer; MRC-5 (PD 30), human fetal lung; MCF7, Caucasian human breast cancer; SNL 76 / 7, mouse SIM strain embryonic fibroblasts; C2C12, mouse C3H myoblasts; Jurkat E6.1, human leukemia T-cell lymphoblasts; U937, Caucasian histiocytic lymphoma; L929, mouse C3H / An connective tissue; 3T3L1, mouse embryo; HL60, Caucasian promyelocytic leukemia; PC-12, rat adrenal pheochromocytoma; HT29, Caucasian colonic adenocarcinoma; OE33, Caucasian esophageal cancer; OE19, Caucasian esophageal cancer; NIH 3T3, Swiss mouse NIH embryo; MDA-MB-231, Caucasian breast cancer; K562, Caucasian chronic myeloid leukemia; U-87MG, human glioblastoma astrocytoma; MRC-5 (PD 25), human fetal lung; A2780cis, human ovarian cancer; B9, mouse B-cell hybridoma; CHO-K1, U2OS, Chinese hamster ovary; MDCK, canine Cocker Spaniel kidney; 1321N1, human brain astrocytoma; A431, human squamous cell carcinoma; ATDC5, derived mouse 129 teratoma AT805; RCC4 PLUS VECTOR ALONE, renal cell carcinoma line RCC4 stably transfected with empty expression vector pcDNA3, conferring neomycin resistance; HUVEC (S200-05n), human pre-selected umbilical vein endothelial cells (HUVEC); neonate; Vero, African monkey green kidney; RCC4 PLUS VHL, RCC4 renal cell carcinoma cell line stably transfected with pcDNA3-VHL; Fao, rat liver cancer; J774A.1, mouse BALB / c monocytes / macrophages; MC3T3-E1, mouse C57BL / 6 skull; J774.2. Mouse BALB / c mononuclear macrophages; PNT1A, normal human post-pubertal prostate, immortalized with SV40; U-2 OS, human osteosarcoma; HCT 116, human colon cancer; MA104, African monkey green kidney; BEAS-2B, human bronchial epithelium, normal; NB2-11, rat lymphoma; BHK 21 (clone 13), Syrian hamster kidney; NS0, mouse myeloma; Neuro 2a, mouse albino neuroblastoma; SP2 / 0-Ag14, mouse x mouse myeloma, nonproductive; T47D, human mammary tumor; 1301, human T-cell leukemia; MDCK-II, canine cocker Spaniel kidney; PNT2, human prostate normal, immortalized with SV40; PC-3, Caucasian human prostate adenocarcinoma; TF1, human erythroleukemia; COS-7, African monkey green kidney, SV40 transformed; MDCK, canine cocker Spaniel kidney; HUVEC (200-05n), human umbilical vein endothelial cells (HUVEC); neonate; NCI-H322, Caucasian human bronchiolar carcinoma; SK.N.SH, Caucasian human neuroblastoma; LNCaP.FGC, Caucasian prostate cancer; OE21, Caucasian esophageal squamous cell carcinoma; PSN1, human pancreatic adenocarcinoma; ISHICHAWA, Asian human endometrial adenocarcinoma; MFE-280, Caucasian endometrial adenocarcinoma; MG-63, human osteosarcoma; RK 13, rabbit kidney, BVDV negative; EoL-1 cells, human eosinophilic leukemia; VCaP, human prostate cancer metastasis; tsA201, human embryonic kidney, SV40 transformation; CHO, Chinese hamster ovary; HT 1080, Human Fibrosarcoma; PANC-1, Caucasian Pancreas; Saos-2, Human Primary Osteosarcoma; Fibroblast Growth Matrix (116K-500), Fibroblast Growth Matrix Kit; ND7 / 23, Mouse Neuroblastoma x Rat Neuronal Hybridoma; SK-OV-3, Caucasian Ovarian Adenocarcinoma; COV434, Human Ovarian Granulosarcoma; Hep 3B, Human Hepatocellular Carcinoma; Vero (WHO), African Monkey Green Kidney; Nthy-ori3-1, Human Thyroid Follicular Epithelium; U373 MG (Uppsala), Human Glioblastoma Astrocytoma; A375, Human Malignant Melanoma; AGS, Caucasian Gastric Adenocarcinoma; CAKI 2, Caucasian Renal Carcinoma; COLO 205, Caucasian Colonic Adenocarcinoma; COR-L23, Caucasian Large Cell Lung Carcinoma; IMR32, Caucasian Neuroblastoma; QT 35, Japanese quail fibrosarcoma; WI 38, Caucasian human fetal lung; HMVII, human vaginal malignant melanoma; HT55, human colon cancer; TK6, human lymphoblastic thymidine kinase heterozygote; SP2 / 0-AG14 (AC-FREE), mouse x mouse hybridoma, non-secreting, serum-free, animal component-free (AC); AR42J, or rat exocrine pancreatic tumor, or any combination thereof.
[0098] A method for measuring plasma membrane neighbor reporter activity using unique molecular identifiers The following describes an example of a method for measuring plasma membrane activity by using a unique molecular identifier after exposure to a test reagent.
[0099] In a non-limiting example, multiple cells are incubated in one well of a multi-well plate. The multiple cells are incubated with at least one nucleic acid described herein (e.g., ...). Figure 1-4The nucleic acid is used for transfection with nucleic acid 110, nucleic acid 120, and / or nucleic acid 130. In some embodiments, at least one nucleic acid encodes a regulatory element described herein (e.g., any combination of doxycycline-inducible promoter 112, plasma membrane polypeptide 114, plasma membrane adjacent promoter 122, luciferase gene 124, UMI 126, promoter 132, and / or protease 134). In some embodiments, at least one nucleic acid encodes a reporter gene described herein (e.g., including luciferase gene 124 and UMI 126, or including only unique molecular identifier 126). In some embodiments, at least one nucleic acid encodes a doxycycline-inducible promoter operably combinable with a plasma membrane construct. In some embodiments, at least one nucleic acid encodes a promoter operably combinable with a protease. In this example, multiple cells are introduced with doxycycline, which can increase the activation of the doxycycline-inducible promoter in multiple cells, which can lead to increased expression of plasma membrane polypeptides (e.g., rhodopsin), linkers with cleavage sites, and transcription factors.
[0100] Multiple cells can then be exposed to the test reagent. Exposure to a specific dose of the test reagent can improve the transport rate of plasma membrane peptides to the plasma membranes of multiple cells. The plasma membrane peptides can then localize to the plasma membranes of multiple cells. The plasma membrane construct can then be cleaved at the cleavage site of the linker, leaving the peptides that remain localized to the plasma membrane, while transcription factors are released.
[0101] Once released, transcription factors can bind to cell membrane-near promoters of nucleic acids. This binding can induce the expression of a unique molecular identifier. The ability of a cell membrane construct (e.g., at the peptide site) to bind to the cell membrane can then be used to assess the cell's capacity to alleviate certain diseases or conditions (e.g., retinitis pigmentosa, Alzheimer's disease, Parkinson's disease, Huntington's disease, cystic fibrosis, α-1 antitrypsinemia, or amyotrophic lateral sclerosis).
[0102] Furthermore, as described above, exposure of multiple cells to the test agent can additionally allow the test agent to interact with cellular components. Specifically, the test agent can interact with peptides or cellular components important for peptide biosynthesis and proper transport, thereby increasing or decreasing the peptide's ability to localize at the cell's plasma membrane. Therefore, due to changes in the peptide's ability to localize at the cell's plasma membrane (e.g., increase or decrease), transcription factors may or may not be released (e.g., because linkers may or may not be cleaved), which can directly affect (e.g., increase or decrease) the expression of unique molecular identifiers. Thus, by observing changes in the expression of unique molecular identifiers resulting from exposure to the test agent, the effect of the test agent on the peptide's ability to localize at the cell's plasma membrane can also be observed, and consequently, the effect of the test agent on the cell's ability to alleviate certain diseases or conditions.
[0103] Next-generation sequencing As described in the methods disclosed herein, sequencing of nucleic acid molecules is used and can be used to detect the biological effects of assays on cells, including cell-based assays. Generally, sequencing refers to methods and techniques for determining the nucleotide base sequence in one or more polynucleotides. Polynucleotides can be, for example, deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), including variants or derivatives thereof (e.g., single-stranded DNA). Sequencing can be performed using a variety of currently available systems, such as, but not limited to, sequencing systems from Illumina, Pacific Biosciences, Oxford Nanopore, or Life Technologies (Ion Torrent). Such devices can provide multiple raw genetic data corresponding to the genetic information of an object (e.g., a human), such as data generated by the device from a sample provided by the object. In some cases, the systems and methods described herein can be used in conjunction with proteomics information. Alternatively or additionally, sequencing can be performed using nucleic acid amplification, polymerase chain reaction (PCR) (e.g., digital PCR, quantitative PCR, or real-time PCR), or isothermal amplification. Such systems can provide multiple raw genetic data corresponding to the genetic information of an object (e.g., a human), such as data generated by the system from a sample provided by the object. In some examples, such systems provide sequencing reads (also referred to herein as "reads"). A read can include a string of nucleic acid bases corresponding to the sequenced nucleic acid molecule. In some cases, the systems and methods presented herein can be used in conjunction with proteomics information.
[0104] Next-generation sequencing encompasses a wide range of technologies capable of generating large amounts of sequence information and excludes Sanger sequencing or Maxam-Gilbert sequencing. In general, next-generation sequencing covers single-molecule real-time sequencing, synthetic sequencing, ion semiconductor sequencing, and more. Exemplary next-generation sequencing machines may include the MiniSeq, iSeq100, NextSeq1000, NextSeq 2000, NovaSeq 6000, and NextSeq 550 series from Illumina, Inc.; the Ion Torrent machine from Thermo Fisher Scientific; or the Sequel system from Pacific Biosciences.
[0105] Next-generation sequencing machines used with the methods described in this paper can generate at least 1, 5, 10, 15, 25, 50, 75, 100, 200, 300 gigabases of data or more from a single machine within a 24-hour timeframe.
[0106] Next-generation sequencing machines used with the methods described in this paper can generate at least 1, 1, 4, 10, 15, 25, 50, 75, 100, 200, 300, 500 or 1 billion sequence reads or more from a single machine within a 24-hour timeframe.
[0107] It also includes computer programs, computing devices, or analysis platforms / systems for receiving and analyzing sequencing data and outputting one or more reports, which may be transmitted or accessed electronically via a server, analysis portal, or email. The computing devices or analysis platforms may operate according to the algorithms and methods described herein.
[0108] The nucleic acids disclosed herein are compatible with many vectors common in the art. Non-limiting examples of vectors include genome integration vectors, free vectors, plasmids, viral vectors, granules, bacterial artificial chromosomes, and yeast artificial chromosomes. Non-limiting examples of viral vectors compatible with the nucleic acids of this disclosure include vectors derived from lentiviruses, retroviruses, adenoviruses, and adeno-associated viruses. In some embodiments, the nucleic acids of this disclosure are present on a vector containing a sequence that specifically integrates into a site in the genome at a defined location or restriction set (e.g., AttP-AttB recombination).
[0109] In some embodiments, the system described herein is incorporated into a single vector. In some embodiments, the single vector is transiently transfected into cells. In some embodiments, the single vector is stably transfected into cells.
[0110] In some embodiments, the system is divided into two vectors. In some embodiments, the first vector contains a first regulatory element and a first effector, while the second vector contains a second regulatory element for regulating the expression of a second effector. In some embodiments, the first and second vectors are transiently transfected into cells. In some embodiments, the first and second vectors are stably transfected into cells. In some embodiments, the first vector is stably transfected into cells, and the second vector is transiently transfected into cells. In some embodiments, the first vector is transiently transfected into cells, and the second vector is stably transfected into cells. In some embodiments, a single vector containing a reporter can be transfected into cells. In some embodiments, cells transfected with the first or second vector already contain a reporter.
[0111] Vectors or portions thereof containing the systems described herein can be constructed using a number of well-known molecular biology techniques. Detailed protocols for many such procedures, including amplification, cloning, mutation formation, transformation, etc., are described below, for example, Ausubel et al., Current protocols in Molecular Biology (2012 supplement), John Wiley & Sons, New York 10 (“Ausubel”); Sambrook et al., Molecular Cloning - A Laboratory Manual (4th edition), Volumes 1–3, Cold Spring Harbor Laboratory, Cold Spring Harbor, New York, 2012 (“Sambrook”); and Abelson et al., Guide to Molecular Cloning Techniques (Methods in Enzymology), Volume 152, Academic Press, Inc., San Diego, CA (“Abelson”).
[0112] Compliance test This document describes a method for determining compliance with defective plasma membrane transport of a variant plasma membrane protein of interest in a test agent. The document provides a method comprising: (a) expressing a recombinant form of the variant plasma membrane protein in a host cell and contacting the host cell with the test agent; (b) measuring the transport of the variant plasma membrane protein to the plasma membrane of the host cell using any of the systems provided herein; (c) comparing the transport determined in (b) with the transport in the host cell when not contacted with the test agent; and (d) if the transport of the plasma membrane protein in the host cell contacted with the test agent is increased compared with the transport in the host cell not contacted with the test agent, then identifying a patient who has or is susceptible to a disease associated with the variant plasma membrane protein as a candidate for treatment with the test agent. In some embodiments, step (d) includes identifying a patient as a candidate for treatment with the test agent if, in step (c), the transport in the host cell contacted with the test agent increases by at least 1.3 to 40-fold compared with the transport in the host cell not contacted with the test agent. In some implementations, the increase in transport in host cells exposed to the test reagent, compared to transport in host cells not exposed to the test reagent, can be at least about 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.2, 2.4, 2.6, 2.8, 3.0, 3.2, 3.4, 3.6, 3.8, 4.0, 4.2, 4.4, 4.6, 4.8, 5.0, 5.5, 6.0, 6.5, 7.0, 7.5, 8.0, 8.5, 9.0, 9.5. The values are 10.0, 10.5, 11.0, 11.5, 12.0, 12.5, 13.0, 13.5, 14.0, 14.5, 15.0, 15.5, 16.0, 16.5, 17.0, 17.5, 18.0, 18.5, 19.0, 19.5, 20.0, 25.0, 30.0, 35.0, 40.0, 45.0, 50.0, 55.0, 60.0, 65.0, 70.0, 75.0, 80.0, 85.0, 90.0, 95.0, or 100.0 times greater. In some embodiments, step (d) includes determining the patient as a candidate for treatment with the test agent if the transport in the host cell is at least 2% to about 100% of the non-mutated plasma membrane protein. In some implementations, transport within the host cell can be at least about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 22%, 23%, 24%, 26%, 28%, 30%, 32%, 34%, 36%, 38%, 40%, 42%, 44%, 46%, 48%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the non-mutant plasma membrane protein.
[0113] In some implementations, variant plasma membrane proteins or genes encoding variant plasma membrane proteins have been identified from patients who have or are susceptible to diseases associated with variant plasma membrane proteins.
[0114] In some implementations, the test reagent is a pharmacologically derived positron.
[0115] Reagent test kit In some embodiments, the kit comprises the system described herein, which can be used to perform the methods described herein. The kit comprises a collection of materials or compositions, including at least one composition of the system. In other embodiments, the kit contains all compositions necessary and / or sufficient to perform the methods described herein, including all controls and indicators.
[0116] In some cases, instructions for use may be included in the kit. Optionally, the kit may also contain other useful components such as diluents, buffers, pharmaceutically acceptable carriers, syringes, catheters, applicators, pipettes or measuring tools, bandaging materials, or other useful equipment. Materials or components assembled in the kit may be provided to the practitioner in any convenient and suitable manner to maintain their operability and usability. For example, components may be in dissolved, dehydrated, or lyophilized form; they may be provided at room temperature, refrigerated, or frozen temperatures. Components are typically contained in suitable packaging materials. As used herein, the phrase “packaging material” refers to one or more physical structures used to contain the contents of the kit, such as compositions. Packaging materials are constructed using well-known methods to preferably provide a sterile, contaminant-free environment. The packaging materials used in the kit are those conventionally used for gene expression assays and therapeutic administration. As used herein, the term “packaging” refers to a suitable solid matrix or material, such as glass, plastic, paper, foil, etc., capable of containing individual kit components. Thus, for example, packaging may be a glass vial or a pre-filled syringe for containing a suitable amount of pharmaceutical composition. The packaging material has an external label that indicates the contents and / or use of the kit and its components.
[0117] Example Specialized processes have been developed to inform the performance of the workflow described herein. These processes include directing input to multiple cells (e.g., Figure 1-4 Multiple cells (100) are given doxycycline and / or a certain dose of one or more assay reagents, and certain peptides (e.g., Figure 3-4 The polypeptide 152) and / or unique molecular identifiers (e.g., Figure 1-4The expression of unique molecular identifiers (126) was determined. At the end of the multiplexing assay, these peptides and unique molecular identifiers were sequenced and readouts were obtained, and the results provided crucial insights into how one or more components introduced into a cell population affect the peptide's ability to localize at the cell's plasma membrane. Specific embodiments provided herein include, but are not limited to, assays identifying pharmacologically schooled positrons (e.g., molecular chaperones) capable of restoring plasma membrane expression of specific mutant peptides that cannot be properly transported to the plasma membrane, and compliance assays identifying plasma membrane peptide mutations that can be treated with specific pharmacologically schooled positrons.
[0118] Example 1: Assessment of rhodopsin expression using unique molecular identifier output The HTS assay uses cell runs in the wells of a 384-well plate, with each well containing approximately 20,000 cells. The cells comprise a first nucleic acid sequence encoding a plasma membrane construct containing a doxycycline-inducible promoter operably coupled to a nucleic acid sequence encoding a plasma membrane polypeptide (rhodopsin polypeptide in this example), and a second nucleic acid sequence containing a reporter construct promoter operably coupled to a luciferase gene and a unique molecular identifier (i.e., a plasma membrane-adjacent promoter).
[0119] A specific dose of doxycycline was added to each well of a 384-well plate to induce the expression of rhodopsin peptides. The expression of rhodopsin peptides allows the rhodopsin peptides to be transported to the cell membrane.
[0120] After the rhodopsin peptide localizes to the cell's plasma membrane, the linker of the plasma membrane construct, which includes the plasma membrane peptide, is cleaved at the cleavage site by a protease anchored to the plasma membrane. Consequently, the transcription factors of the plasma membrane construct, which include DNA-binding domains and transcription activators, are released into the cell.
[0121] The DNA-binding domain of a transcription factor activates the promoter of a second nucleic acid to induce the expression of a luciferase gene and a unique molecular identifier, the expression of which is related to the localization rate of the rhodopsin polypeptide at the cell membrane.
[0122] The photometer output after the addition of luciferase reagent is determined to indicate the expression of unique molecular identifiers.
[0123] Example 2: Assessment of rhodopsin expression using unique molecular identifiers output from test reagent responses To begin, the first HTS measurement according to Example 1 is run.
[0124] Next, a second HTS assay was performed using cells in the wells of a 384-well plate, with each well containing approximately 20,000 cells. The cells comprised a first nucleic acid sequence encoding a plasma membrane construct containing a doxycycline-inducible promoter operably coupled to a nucleic acid sequence encoding a plasma membrane polypeptide (in this example, rhodopsin polypeptide), and a second nucleic acid sequence containing a reporter construct promoter operably coupled to a luciferase gene and a unique molecular identifier (i.e., a plasma membrane-adjacent promoter). A dose of doxycycline was added to each well of the 384-well plate to induce rhodopsin polypeptide expression. Rhodopsin polypeptide expression allows the rhodopsin peptide to be transported to the cell's plasma membrane.
[0125] A specific dose of the assay reagent was then added to each well of a 384-well plate to assess the effect of the assay reagent on the expression of rhodopsin peptides and the related transport of rhodopsin peptides to the cell plasma membrane.
[0126] After the rhodopsin peptide localizes to the plasma membrane, the plasma membrane construct, including the expressed rhodopsin peptide, the adaptor, and the transcription factor, is cleaved at the adaptor. More specifically, the adaptor is cleaved at the cleavage site by a protease anchored to the cell's plasma membrane. The rhodopsin peptide is then retained at the plasma membrane while the transcription factor, including the DNA-binding domain and the transcription activator, is released.
[0127] The DNA-binding domain of a transcription factor activates the promoter of a second nucleic acid to induce the expression of a luciferase gene and a unique molecular identifier, the expression of which is related to the localization rate of the rhodopsin polypeptide at the cell membrane.
[0128] The photometer output after the addition of the luciferase agent is determined to indicate the expression of the unique molecular identifier. The photometer output associated with the wells in this example is then compared to the photometer output associated with the wells in Example 1. The expression of the unique molecular identifier in the wells of the second assay is compared to the expression of the unique molecular identifier in the wells of the first assay to determine whether the assay agent caused an increase in the unique molecular identifier readout. This allows for the assessment of the transport of rhodopsin peptides to the plasma membrane in the second assay compared to the first assay, thus indicating whether the added assay agent increased or decreased the transport of rhodopsin peptides.
[0129] Example 3: Compliance Measurement The systems and methods described herein can also be used to perform compliance assays to assess the transport consequences of each individual possible amino acid change in a specific plasma membrane peptide. The transport consequences of each individual possible amino acid change in a specific plasma membrane peptide, or at least of each pathogenic mutation identified in patients with certain conditions or diseases (e.g., retinitis pigmentosa, GLP-1R downregulation, cystic fibrosis, etc.), can also be assessed in the presence of a test kit to determine whether treatment with the test kit is likely to effectively correct plasma membrane transport of the pathogenic mutant peptide.
[0130] A DNA construct encoding (i) a plasma membrane polypeptide construct containing a plasma membrane polypeptide variant, (ii) a plasma membrane-anchored protease, and (iii) a reporter with a variant-specific barcode can be stably integrated into HEK293T cells as a single copy. Cell lines can be prepared for each variant of the plasma membrane polypeptide, each with a unique molecular identifier that can be associated with each variant. Cells from each cell line can be pooled and plated in 15-cm wells of DMEM 10% FBS. Expression of the plasma membrane polypeptide construct can be driven by a doxycycline-inducible promoter. Expression of the plasma membrane-anchored protease can be driven by a constitutive promoter. Assay reagents and / or doxycycline can be applied to the cells immediately after plating. Cells can be incubated for 24 hours in a cell culture incubator at 37°C and 5% CO2. Cells can be lysed after incubation. The RNA barcode can be selectively reverse transcribed (from the primer upstream of the barcode) to generate cDNA, which can be amplified to prepare an NGS library. Barcodes can be sequenced, and barcode counts can be modeled using a negative binomial generalized linear model to determine variant effects relative to the wild type. If the linker between transcription factors and variant plasma membrane peptides is cleaved by plasma membrane-anchored proteases, the barcodes associated with the variant plasma membrane peptides will be sequenced.
[0131] Using the methods and systems provided herein, the transport consequences of each individual possible amino acid change in proteins such as rhodopsin or GLP-1R, or any other plasma membrane-targeting proteins, can be evaluated in the presence or absence of a test reagent (e.g., a molecular chaperone or a pharmacological school positron).
[0132] Example 4: Optimization of the protease cleavage site of the adapter The plasma membrane construct linker between plasma membrane peptides and transcription factors can be cleaved at the cleavage site by a protease anchored to the cell's plasma membrane. A variety of proteases can be used in the system described herein. Experiments were performed to determine the optimal cleavage site of the tobacco etch virus (“TEV”)-S219V variant protease, which cleaves the sequence ENLYFQ(X) (SEQ ID NO: 22). Figure 6The optimized cleavage site sequence for TEV cleavage was depicted, showing a comparison between serine and tyrosine substitutions within amino acid (X). In this embodiment, rhodopsin is a plasma membrane construct. RHO-WT is wild-type rhodopsin, and RHO-P23H is a mutant with poor expression at the plasma membrane. Rhodopsin expression is induced by doxycycline. At low expression, the difference between TCS(S) (i.e., ENLYFQ / S) and TCS(Y) (i.e., ENLYFQ / Y) is minimal. Surprisingly, at higher expression, TCS(Y) appears to be a better protease cleavage site for TEV, and the reporter is better able to distinguish RHO-WT from RHO-P23H expression at the plasma membrane.
[0133] Example 5: Optimization of plasma membrane protease anchors and promoters The plasma membrane-anchored protease can be encoded by an exogenous nucleic acid. A promoter is operatively coupled to a nucleic acid sequence encoding a plasma membrane anchor (e.g., Lyn11, PHDelta1, PHDelta3) and a protease (e.g., TEV) that binds via a linker (e.g., ASPSNPGASNGS (“ASPS”; SEQ ID NO: 39) or GGGGSGGGGS (“2xGS”; SEQ ID NO: 40)). When the promoter is activated, the plasma membrane anchor and protease can be expressed. When targeted or localized to the plasma membrane, the protease can be able to cleave the plasma membrane construct linker that binds the plasma membrane polypeptide to the transcription factor at a cleavage site. Experiments are performed to determine the optimal protease anchor and / or promoter. Figure 7 A comparison between constitutive promoters using GAPDH and ef1a to drive membrane-anchored TEVs is depicted. Different membrane anchors were also tested to indicate the effectiveness of Lyn11 versus the PHDelta variant in detecting reagent rescue for RHO-P23H membrane transport. In this embodiment, the 2xGS adapter was used unless ASPS was specified. Surprisingly, the Lyn11-ASPS anchor-connector worked best for reagent rescue in detecting RHO-P23H membrane transport. Lyn11-2xGS showed better performance when using a constitutive ef1a promoter compared to GAPDH.
[0134] Example 6: Optimization of DNA Binding Domain (DBD) Transcription factors containing a deep brain stimulation (DBD) and a trans-activator can link to plasma membrane peptides. When plasma membrane-anchored proteases cleave the linker, the transcription factor can bind to the reporter construct promoter and drive the expression of the reporter construct. Experiments were conducted to determine the optimal DBD for detecting plasma membrane construct transport rescue. Figure 8A comparison of Gal4 and ZF (i.e., RARR) DBD in detecting the rescue of RHO-P23H transport by the assay reagent was presented. Surprisingly, ZF performed best in detecting the rescue of RHO-P23H at higher concentrations of the assay reagent. DBD-responsive promoters can be further optimized for use with ZF; for example, DBD-responsive promoters can be prepared to have up to 12 ZF binding sites.
[0135] Example 7: Luciferase Protocol DNA constructs encoding RHO-TF, PM-TEV, and a luciferase reporter can be stably integrated into HEK293T cells. RHO-TEV is a rhodopsin tethering transcription factor. PM-TEV is a plasma membrane-anchored TEV protease capable of cleaving the linker between rhodopsin and the transcription factor. The transcription factor binds to the promoter that drives luciferase reporter expression. Cells can be plated at 20,000 cells / well in 384-well plates in DMEM 10% FBS. RHO-TF expression can be driven by a doxycycline-inducible promoter. PM-TEV expression can be driven by a constitutive promoter. When doxycycline is applied, assay reagents or other assay conditions can be added to determine the effect on RHO-TF plasma membrane transport. Cells can be incubated for 24 hours in a cell culture incubator at 37°C and 5% CO2. (Based on Promega Bright-Globe) TM The assay system allows for the acquisition of luciferase readouts on a photometer.
[0136] Example 8: Wide Target Scan (BTS) Protocol DNA constructs encoding RHO-TF, PM-TEV, and a reporter can be stably integrated into HEK293T cells. In this embodiment, the reporter is a variant-specific barcode. Cellular libraries can be generated by incorporating the integrated cells. Cells can be plated in 384-well plates at 20,000 cells / well in DMEM 10% FBS. RHO-TF expression can be driven by a doxycycline-inducible promoter. PM-TEV expression can be driven by a constitutive promoter. Assay reagents and / or doxycycline can be applied to cells immediately after plating. Cells can be incubated for 24 hours in a cell culture incubator at 37°C and 5% CO2. Cells can be lysed after incubation. The RNA barcode can be selectively reverse transcribed (from the primer upstream of the barcode) to generate cDNA, which can be amplified to prepare an NGS library. Barcode sequencing can be performed, and barcode counting can be modeled using a negative binomial generalized linear model to identify variants and effects of the test compounds.
[0137] Example 9: Multiplexable transport determination via broad target scan (BTS) A 384-well plate can be set up to simultaneously measure the transport consequences of 50-100 variants. Dose-response curves for plasma transport salvage can be collected simultaneously for dozens of variants, equivalent to 264 dose-response curves (17,000 data points) collected from a single 384-well plate. Figure 9 The specificity of six different assays using BTS across 32 different pathogenic autosomal dominant retinitis pigmentosa (ADRP) variants is depicted. Grayscale reflects the potency level against the mutation.
[0138] Example 10: Deep Mutation Scan (DMS) Protocol DNA constructs encoding RHO-TF, PM-TEV, and a reporter with variant-specific barcodes can be stably integrated into HEK293T cells in a single-copy, pooled format. Cells can be plated in 15-cm wells in DMEM 10% FBS. RHO-TF expression can be driven by a doxycycline-inducible promoter. PM-TEV expression can be driven by a constitutive promoter. Assay reagents and / or doxycycline can be applied to cells immediately after platening. Cells can be incubated at 37°C and 5% CO2 for 24 hours. Cells can be lysed after incubation. RNA barcodes can be selectively reverse transcribed (upstream of the barcode primer) to generate cDNA, which can be amplified to prepare NGS libraries. Barcode sequencing is possible, and barcode counting can be modeled using a negative binomial generalized linear model to determine variant effects relative to wild-type.
[0139] Figure 10 This paper describes how the transport consequences of every single possible amino acid change in rhodopsin protein can be measured via DMS using multiplexable transport assays. Transport scores are normalized to wild-type (“WT”) transport with a value of 1.0.
[0140] Although preferred embodiments of the systems and methods have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided by way of example only. The systems and methods are not limited to the specific examples provided in the specification. While the systems and methods have been described with reference to the foregoing specification, the descriptions and illustrations of the embodiments herein are not intended to be construed as limiting. Many variations, alterations, and substitutions will now occur to those skilled in the art without departing from this disclosure. Furthermore, it should be understood that all aspects of the systems and methods are not limited to the specific descriptions, configurations, or relative proportions described herein, and depend on variations in various conditions and variables. It should be understood that various alternatives to the embodiments of the systems and methods described herein may be employed in the practice of this disclosure. Therefore, the systems and methods described herein should also cover any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the systems and methods described herein, and the systems and methods described herein are within the scope of these claims, and their equivalents are thereby covered.
[0141] Numbering Implementation Plan 1. A system comprising eukaryotic cells, wherein the eukaryotic cells comprise: (a) a plasma membrane construct (PMC) comprising a plasma membrane polypeptide coupled to a transcription factor via a PMC linker, wherein the linker comprises a protease cleavage site; and (b) a plasma membrane anchoring protease capable of cleaving the linker.
[0142] 2. The system as described in Implementation Scheme 1 further includes a reporter construct, the reporter construct comprising a promoter and a reporter gene.
[0143] 3. The system as described in embodiment 2, wherein the promoter is bound by the transcription factor upon cleavage of the adapter.
[0144] 4. The system of any one of embodiments 1 to 3, wherein the plasma membrane construct is encoded by an exogenous nucleic acid.
[0145] 5. The system of any one of embodiments 2 to 4, wherein the report sub-construct is encoded by an exogenous nucleic acid.
[0146] 6. The system of any one of embodiments 2 to 5, wherein the reporter gene contains a unique molecular identifier.
[0147] 7. The system of any one of embodiments 2 to 6, wherein the reporter gene comprises a fluorescent protein or a luciferase protein.
[0148] 8. The system of any one of embodiments 2 to 7, wherein the reporter gene comprises a fluorescent protein and a unique molecular identifier or a luciferase protein and a unique molecular identifier.
[0149] 9. The system of any one of embodiments 1 to 8, wherein the plasma membrane anchoring protease is a component of the plasma membrane of the eukaryotic cell.
[0150] 10. The system of any one of embodiments 1 to 9, wherein the plasma membrane anchoring protease comprises a membrane tethering protease.
[0151] 11. The system of embodiment 10, wherein the membrane tethering protease comprises a Plek substrate protein homology domain, a platelet-derived growth factor receptor, or a Lyn anchor domain.
[0152] 12. The system of any one of embodiments 1 to 12, wherein the transcription factor comprises a DNA-binding domain and a transcription activation domain.
[0153] 13. The system of embodiment 12, wherein the DNA binding domain comprises a Gal4, PPR1, Lac9, zinc finger, or LexA DNA binding domain.
[0154] 14. The system as described in embodiment 12 or 13, wherein the transcriptional activation domain comprises a VP64, VPr, p65, Rta, or VP16 activation domain.
[0155] 15. The system of any one of embodiments 1 to 14, wherein the connector comprises a flexible amino acid connector.
[0156] 16. The system of any one of embodiments 1 to 15, wherein the length of the connector is about 2 to about 31 amino acids.
[0157] 17. The system of any one of embodiments 1 to 16, wherein the membrane-anchoring protease comprises tobacco etch virus, aspartic acid, glutamic acid, metalloid, cysteine, serine, or threonine protease.
[0158] 18. The system of any one of embodiments 1 to 17, wherein the plasma membrane polypeptide comprises rhodopsin.
[0159] 19. The system of any one of embodiments 1 to 18, wherein the expression of the plasma membrane construct is inducible.
[0160] 20. The system of any one of embodiments 1 to 19, wherein the expression of the plasma membrane construct is induced in response to doxycycline.
[0161] 21. The system of any one of embodiments 1 to 20, wherein the plasma membrane construct is located at the plasma membrane of the eukaryotic cell after expression of the plasma membrane construct.
[0162] 22. The system of any one of embodiments 1 to 21, wherein the plasma membrane polypeptide comprises an amino acid sequence having at least about 90%, 95%, 97%, 98%, 99% or 100% identity with any one of SEQ ID 1-6.
[0163] 23. The system of any one of embodiments 1 to 22, wherein the expression of the report sub-construct indicates the ability of the test agent to alleviate the condition.
[0164] 24. The system as described in embodiment 23, wherein the condition is a degenerative disease.
[0165] 25. The system according to any one of embodiments 1 to 24, wherein the eukaryotic cell is a mammalian cell.
[0166] 26. The system as described in embodiment 25, wherein the mammalian cell is a human cell.
[0167] 27. A population of eukaryotic cells comprising the system described in any one of embodiments 1 to 26.
[0168] 28. A method for screening a test reagent, the method comprising contacting a population of eukaryotic cells as described in embodiment 27 with the test reagent.
[0169] 29. The method of embodiment 28, wherein the test reagent comprises a small molecule compound.
[0170] 30. A system comprising eukaryotic cells, wherein the eukaryotic cells comprise: (a) a plasma membrane construct comprising a plasma membrane polypeptide coupled to a transcription factor via a linker, wherein the linker is cleavable; and (b) a reporter construct comprising a promoter and a reporter gene comprising a unique molecular identifier; wherein the promoter is bound by the transcription factor upon cleavage of the linker.
[0171] 31. The system of embodiment 30 further comprises a plasma membrane anchoring protease capable of cleaving the connector.
[0172] 32. The system as described in embodiment 30 or 31, wherein the plasma membrane construct is encoded by an exogenous nucleic acid.
[0173] 33. The system of any one of embodiments 30 to 32, wherein the report sub-construct is encoded by an exogenous nucleic acid.
[0174] 34. The system of any one of embodiments 30 to 33, wherein the reporter gene further encodes a fluorescent protein or a luciferase protein.
[0175] 35. The system of any one of embodiments 31 to 34, wherein the plasma membrane anchoring protease is a component of the plasma membrane of the eukaryotic cell.
[0176] 36. The system of any one of embodiments 31 to 34, wherein the plasma membrane anchoring protease comprises a membrane tethering protease.
[0177] 37. The system of embodiment 36, wherein the membrane tethering protease comprises a pleleck substrate protein homology domain, a platelet-derived growth factor receptor, or a Lyn anchor domain.
[0178] 38. The system of any one of embodiments 30 to 37, wherein the transcription factor comprises a DNA-binding domain and a transcription activation domain.
[0179] 39. The system of embodiment 38, wherein the DNA binding domain comprises a Gal4, PPR1, Lac9, zinc finger, or LexA DNA binding domain.
[0180] 40. The system of embodiment 38 or 39, wherein the transcriptional activation domain comprises a VP64, VPr, p65, Rta, or VP16 activation domain.
[0181] 41. The system of any one of embodiments 30 to 40, wherein the connector comprises a flexible amino acid connector.
[0182] 42. The system of any one of embodiments 30 to 41, wherein the length of the connector is about 2 to about 31 amino acids.
[0183] 43. The system of any one of embodiments 30 to 42, wherein the membrane-anchoring protease comprises tobacco etch virus, aspartic acid, glutamic acid, metalloid, cysteine, serine, or threonine protease.
[0184] 44. The system of any one of embodiments 30 to 43, wherein the plasma membrane polypeptide comprises rhodopsin.
[0185] 45. The system of any one of embodiments 30 to 44, wherein the expression of the plasma membrane construct is inducible.
[0186] 46. The system of any one of embodiments 30 to 45, wherein the expression of the plasma membrane construct is induced by doxycycline.
[0187] 47. The system as described in embodiment 45 or 46, wherein the plasma membrane polypeptide is localized to the plasma membrane of the eukaryotic cell after expression of the plasma membrane polypeptide.
[0188] 48. The system of any one of embodiments 30 to 47, wherein the plasma membrane polypeptide comprises an amino acid sequence having at least about 90%, 95%, 97%, 98%, 99% or 100% identity with any one of SEQ ID 1-6.
[0189] 49. The system of any one of embodiments 30 to 48, wherein the expression of the report sub-construct indicates the ability of the test agent to alleviate the condition.
[0190] 50. The system as described in embodiment 49, wherein the condition is a degenerative disease.
[0191] 51. The system as described in any one of embodiments 30 to 50, wherein the eukaryotic cell is a mammalian cell.
[0192] 52. The system as described in embodiment 51, wherein the mammalian cell is a human cell.
[0193] 53. A population of eukaryotic cells comprising the system described in any one of embodiments 30 to 52.
[0194] 54. A method for screening a test reagent, the method comprising contacting a population of eukaryotic cells as described in embodiment 53 with the test reagent.
[0195] 55. The method of embodiment 54, wherein the test reagent comprises a small molecule compound.
[0196] Table 1 - Sequences of plasma membrane peptides Table 2 - Sequence of membrane construct linkers and their domains Table 3 - Sequences of transcription factors and their domains Table 4 - Sequences of plasma membrane-anchored proteases and their domains Table 5 - Sequences of Synthetic Promoters
Claims
1. A system comprising eukaryotic cells, wherein the eukaryotic cells comprise: a. A plasma membrane construct (PMC) comprising a plasma membrane polypeptide coupled to a transcription factor via a PMC linker, wherein the PMC linker comprises a protease cleavage site; and b. Membrane-anchored proteases; The plasma membrane-anchored protease is capable of cleaving the PMC connector.
2. The system of claim 1, further comprising a reporter construct (RC), the reporter construct (RC) comprising an RC promoter and a reporter gene.
3. The system of claim 2, wherein the RC promoter is bound by the transcription factor upon cleavage of the PMC linker.
4. The system of claim 3, wherein the RC promoter comprises a synthetic DNA-binding domain-responsive promoter.
5. The system of any one of claims 2 to 4, wherein the RC promoter comprises SEQ ID NO:
43.
6. The system of any one of claims 2 to 5, wherein the RC promoter comprises a zinc finger binding site.
7. The system of claim 6, wherein the RC promoter comprises 2 to 12 zinc finger binding sites.
8. The system of any one of claims 1 to 7, wherein the plasma membrane construct is encoded by an exogenous nucleic acid.
9. The system of any one of claims 2 to 8, wherein the reporter construct is encoded by an exogenous nucleic acid.
10. The system of any one of claims 2 to 9, wherein the reporter gene contains a unique molecular identifier.
11. The system of any one of claims 2 to 10, wherein the reporter gene encodes a fluorescent protein or a luciferase protein.
12. The system of any one of claims 2 to 11, wherein the reporter gene encodes a fluorescent protein or luciferase protein, and further comprises a unique molecular identifier.
13. The system of any one of claims 1 to 12, wherein the plasma membrane anchoring protease is encoded by an exogenous nucleic acid, and optionally wherein the expression of the plasma membrane anchoring protease is driven by a constitutive promoter.
14. The system of any one of claims 1 to 13, wherein the plasma membrane anchoring protease comprises a membrane tethering protease.
15. The system of claim 14, wherein the membrane-anchored protease comprises a membrane anchor connected to the protease via a protease tether.
16. The system of claim 15, wherein the plasma membrane anchor comprises any one of SEQ ID NO: 35-38.
17. The system of claim 15 or 16, wherein the protease tether comprises SEQ ID NO: 39 or 40.
18. The system of any one of claims 15 to 17, wherein the membrane tethering protease comprises tobacco etch virus (TEV), aspartic acid, glutamic acid, metalloid, cysteine, serine, or threonine protease.
19. The system of claim 18, wherein the membrane tethering protease comprises a TEV protease or a variant of the TEV protease, or a functional fragment thereof.
20. The system of claim 19, wherein the membrane tethering protease comprises a sequence having at least 90% identity with SEQ ID NO:
41.
21. The system of claim 20, wherein the membrane tethering protease comprises SEQ ID NO:
42.
22. The system of any one of claims 1 to 21, wherein the transcription factor comprises a DNA-binding domain and a transcription activation domain.
23. The system of claim 22, wherein the DNA binding domain comprises a Gal4, PPR1, Lac9, zinc finger, or LexA DNA binding domain.
24. The system of claim 23, wherein the DNA binding domain comprises a sequence having at least 90% sequence identity with any of SEQ ID NO: 25-29.
25. The system of any one of claims 22 to 24, wherein the transcriptional activation domain comprises a VP64, VPr, p65, Rta, or VP16 activation domain.
26. The system of claim 25, wherein the transcriptional activation domain comprises a sequence having at least 90% sequence identity with any of SEQ ID NO: 30-34.
27. The system of any one of claims 1 to 26, wherein the PMC connector comprises a flexible amino acid connector.
28. The system of any one of claims 1 to 27, wherein the length of the PMC connector is about 2 to about 31 amino acids.
29. The system of any one of claims 1 to 28, wherein the PMC connector comprises a TEV cuttable sequence.
30. The system of any one of claims 1 to 29, wherein the PMC connector comprises a sequence having at least 90% sequence identity with SEQ ID NO: 20 or 21.
31. The system of any one of claims 1 to 30, wherein the PMC connector comprises a protease cleavage site, the protease cleavage site comprising at least one of SEQ ID NO: 22-24.
32. The system of any one of claims 1 to 31, wherein the plasma membrane polypeptide comprises rhodopsin or a variant thereof.
33. The system of claim 32, wherein the plasma membrane polypeptide comprises a sequence having at least 90% sequence identity with any one of SEQ ID NO: 1-6.
34. The system of any one of claims 1 to 31, wherein the plasma membrane polypeptide comprises cystic fibrosis transmembrane conduction regulator (CFTR) or a variant of CFTR.
35. The system of claim 34, wherein the plasma membrane polypeptide comprises a sequence having at least 90% sequence identity with SEQ ID NO:
7.
36. The system of any one of claims 1 to 31, wherein the plasma membrane polypeptide comprises a G protein-coupled receptor.
37. The system of claim 36, wherein the plasma membrane polypeptide comprises glucagon-like peptide-1 receptor (GLP-1R) or a variant of GLP-1R.
38. The system of claim 37, wherein the plasma membrane polypeptide comprises a sequence having at least 90% sequence identity with SEQ ID NO:
8.
39. The system of claim 36, wherein the plasma membrane polypeptide comprises a sequence having at least 90% sequence identity with any one of SEQ ID NO: 9-19.
40. The system of any one of claims 1 to 39, wherein the expression of the plasma membrane construct is inducible.
41. The system of any one of claims 1 to 40, wherein the expression of the plasma membrane construct is induced in response to doxycycline.
42. The system of any one of claims 1 to 41, wherein the plasma membrane construct is located at the plasma membrane of the eukaryotic cell after expression of the plasma membrane construct.
43. The system of any one of claims 1 to 42, wherein the expression of the report sub-construct indicates the ability of the test agent to alleviate the condition.
44. The system of claim 43, wherein the condition is a degenerative disease.
45. The system of any one of claims 1 to 44, wherein the eukaryotic cell is a mammalian cell.
46. The system of claim 45, wherein the mammalian cell is a human cell.
47. A population of eukaryotic cells comprising the system as described in any one of claims 1 to 46.
48. A method for screening a test reagent, the method comprising contacting a population of eukaryotic cells as described in claim 47 with the test reagent.
49. The method of claim 48, wherein the test reagent comprises a small molecule compound.
50. A system comprising eukaryotic cells, wherein the eukaryotic cells comprise: a. A plasma membrane construct (PMC) comprising a plasma membrane polypeptide coupled to a transcription factor via a PMC linker, wherein the PMC linker is cleavable; and b. A reporter construct (RC), which contains an RC promoter and a reporter gene containing a unique molecular identifier; The RC promoter is bound by the transcription factor when the PMC linker is cleaved.
51. The system of claim 50, further comprising a membrane-anchoring protease capable of cleaving the connector, optionally wherein the membrane-anchoring protease is encoded by an exogenous nucleic acid.
52. The system of claim 50 or 51, wherein the plasma membrane construct is encoded by an exogenous nucleic acid.
53. The system of any one of claims 50 to 52, wherein the reporter construct is encoded by an exogenous nucleic acid.
54. The system of any one of claims 50 to 53, wherein the reporter gene further encodes a fluorescent protein or a luciferase protein.
55. The system of any one of claims 51 to 54, wherein the plasma membrane anchoring protease is a component of the plasma membrane of the eukaryotic cell.
56. The system of any one of claims 51 to 54, wherein the membrane-anchored protease comprises a membrane anchor attached to the protease via a protease tether.
57. The system of claim 56, wherein the plasma membrane anchor comprises any one of SEQ ID NO: 35-38.
58. The system of claim 56 or 57, wherein the protease tether comprises SEQ ID NO: 39 or 40.
59. The system of any one of claims 56 to 58, wherein the membrane tethering protease comprises tobacco etch virus (TEV), aspartic acid, glutamic acid, metalloid, cysteine, serine, or threonine protease.
60. The system of claim 59, wherein the membrane tethering protease comprises a TEV protease or a variant of the TEV protease, or a functional fragment thereof.
61. The system of claim 60, wherein the membrane tethering protease comprises a sequence having at least 90% identity with SEQ ID NO:
41.
62. The system of claim 61, wherein the membrane tethering protease comprises SEQ ID NO:
42.
63. The system of any one of claims 50 to 62, wherein the transcription factor comprises a DNA-binding domain and a transcription activation domain.
64. The system of claim 63, wherein the DNA binding domain comprises a Gal4, PPR1, Lac9, zinc finger, or LexA DNA binding domain.
65. The system of claim 64, wherein the DNA binding domain comprises a sequence having at least 90% sequence identity with any of SEQ ID NO: 25-29.
66. The system of any one of claims 63 to 65, wherein the transcriptional activation domain comprises a VP64, VPr, p65, Rta, or VP16 activation domain.
67. The system of claim 66, wherein the transcriptional activation domain comprises a sequence having at least 90% sequence identity with any of SEQ ID NO: 30-34.
68. The system of any one of claims 50 to 67, wherein the PMC connector comprises a flexible amino acid connector.
69. The system of any one of claims 50 to 68, wherein the length of the PMC connector is about 2 to about 31 amino acids.
70. The system of any one of claims 50 to 69, wherein the PMC connector comprises a TEV cuttable sequence.
71. The system of any one of claims 50 to 70, wherein the PMC connector comprises a sequence having at least 90% sequence identity with SEQ ID NO: 20 or 21.
72. The system of any one of claims 50 to 71, wherein the PMC connector comprises a protease cleavage site, the protease cleavage site comprising at least one of SEQ ID NO: 22-24.
73. The system of any one of claims 50 to 72, wherein the plasma membrane polypeptide comprises rhodopsin or a variant thereof.
74. The system of claim 73, wherein the plasma membrane polypeptide comprises a sequence having at least 90% sequence identity with any one of SEQ ID NO: 1-6.
75. The system of any one of claims 50 to 72, wherein the plasma membrane polypeptide comprises cystic fibrosis transmembrane conduction regulator (CFTR) or a variant of CFTR.
76. The system of claim 75, wherein the plasma membrane polypeptide comprises a sequence having at least 90% sequence identity with SEQ ID NO:
7.
77. The system of any one of claims 50 to 72, wherein the plasma membrane polypeptide comprises a G protein-coupled receptor.
78. The system of claim 77, wherein the plasma membrane polypeptide comprises glucagon-like peptide-1 receptor (GLP-1R) or a variant of GLP-1R.
79. The system of claim 78, wherein the plasma membrane polypeptide comprises a sequence having at least 90% sequence identity with SEQ ID NO:
8.
80. The system of claim 77, wherein the plasma membrane polypeptide comprises a sequence having at least 90% sequence identity with any one of SEQ ID NO: 9-19.
81. The system of any one of claims 50 to 80, wherein the expression of the plasma membrane construct is inducible.
82. The system of any one of claims 50 to 81, wherein the expression of the plasma membrane construct is induced by doxycycline.
83. The system of claim 81 or 82, wherein the plasma membrane polypeptide is localized to the plasma membrane of the eukaryotic cell after expression of the plasma membrane polypeptide.
84. The system of any one of claims 50 to 83, wherein the expression of the report sub-construct indicates the ability of the test agent to alleviate the condition.
85. The system of claim 84, wherein the condition is a degenerative disease.
86. The system of any one of claims 50 to 85, wherein the eukaryotic cell is a mammalian cell.
87. The system of claim 86, wherein the mammalian cell is a human cell.
88. A population of eukaryotic cells comprising the system as described in any one of claims 50 to 87.
89. A method for screening a test reagent, the method comprising contacting a population of eukaryotic cells as described in claim 88 with the test reagent.
90. The method of claim 89, wherein the test reagent comprises a small molecule compound.
91. A method for determining compliance of a test reagent in rescuing defective plasma membrane transport of a variant plasma membrane protein of interest, the method comprising: (a) Expressing the recombinant form of the variant plasma membrane protein in a host cell and contacting the host cell with the test reagent; (b) Measuring the transport of the variant plasma membrane protein into the plasma membrane of the host cell using the system as described in any one of claims 1-90; (c) Compare the transport identified in (b) with the transport in the host cells when not in contact with the test agent, and (d) If the transport of the plasma membrane protein in the host cells in contact with the test agent is increased compared with the transport in the host cells that are not in contact with the test agent, then a patient who has or is susceptible to a disease associated with the variant plasma membrane protein is identified as a candidate to be treated with the test agent.
92. The method of claim 91, wherein step (d) comprises determining the patient as a candidate for treatment with the test agent if, in step (c), the transport in the host cells in contact with the test agent increases by at least 1.3 to 40 times compared to the transport in the host cells not in contact with the test agent.
93. The method of claim 91 or 92, wherein step (d) comprises determining the patient as a candidate to be treated with the test agent if the transport in the host cell is at least 2% to about 100% of a non-mutated plasma membrane protein.
94. The method of any one of claims 91 to 93, wherein the test reagent is a pharmacological positron.
95. The method of any one of claims 91 to 94, wherein the variant plasma membrane protein or the gene encoding the variant plasma membrane protein has been identified from the patient who has or is susceptible to a disease associated with the variant plasma membrane protein.