Spliced RNA barcodes for multiplexed cell-based high throughput screening
Recombinant cells with engineered splicing factors and reporter unspliced mRNA constructs address the challenges of multiplexing in HTS by enabling efficient and accurate screening of multiple target proteins without relying on transcription-based reporter methods.
Patent Information
- Application Number
- PCT/US2024/060125
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-12-13
- Publication Date
- 2025-06-19
AI Technical Summary
Current high-throughput screening (HTS) methods face challenges in multiplexing cell-based assays, particularly in identifying all molecular targets of a chemical compound due to the combinatorial explosion of protein targets and compounds, and the limitations of transcription-based reporter methods.
The development of recombinant cells expressing fusion proteins with engineered splicing factors (ESFs) and protease recognition sequences, along with reporter unspliced mRNA (RUM) constructs, allows for multiplexed cell-based assays without requiring test compound-induced activation of reporter gene transcription.
This approach enables simultaneous screening of multiple target proteins with reduced false negatives and false positives, improving the efficiency and accuracy of HTS by using nucleic acid barcodes to distinguish between different splicing forms of the reporter mRNA.
Smart Images

Figure IMGF000017_0001 
Figure IMGF000018_0001 
Figure IMGF000019_0001
Abstract
Description
[0001] Spliced RNA Barcodes for Multiplexed Cell-Based High Throughput Screening
[0002] Sequence Listing Statement
[0003] A computer readable form of the Sequence Listing is filed with this application by electronic submission and is incorporated into this application by reference in its entirety. The Sequence Listing is contained in the file created on December 11, 2024 having the file name “23-1522-WO” and is 119,072 bytes in size.
[0004] Background
[0005] High-throughput screening (HTS), in which chemical libraries of anywhere from tens to millions of distinct molecules are screened for their biological activity, remains prominent in drug discovery. HTS may utilize biochemical, biophysical, or biological (e.g., live cell- based) assays, typically testing a single test compound per well in multi-well plates. Phenotypic cell-based assays are typically mechanistically agnostic (i.e., they report complex cellular physiological responses to a test compound rather than being restricted to measuring a single molecular target's response). Conversely, target-directed assays employing an exogenous target typically report a test compound’s effect (e.g., agonism or antagonism) at only that single specific molecular target.
[0006] Increasingly, the concerns of polypharmacology, toxicology, and machine learning- based drug discovery all argue for deeper knowledge of the identities of all of a chemical compound’s gene product targets (including those with deleterious, neutral, or therapeutic effects), and including all compounds of interest (such as, for example, all FDA-approved small-molecule drugs, all marketed chemicals, etc.). Yet with 20,000+ protein-encoding genes in the human genome, with most of their protein products occurring in multiple splice forms and with multiple post-translational modifications, plus about 1,600 FDA-approved small-molecule drugs and an even much larger number of non-drug marketed chemicals to which humans are or may be exposed, the task of screening such a potentially large number of compounds against such a large number of molecular targets (termed ‘mapping the pharmome’) presents a combinatorial explosion problem of seemingly intractable scale when screening is performed via the conventional target-directed approach of testing one compound against one target in one well. Phenotypic screening as a potential alternative largely fails to address this combinatorial problem because of its inability to reliably identify all or even any of a compound’s molecular targets.
[0007] In the case of target protein molecules that bind to other protein molecules upon activation, including heterodimerization (such as an activated G protein-coupled receptor binding a G protein or arrestin) or homodimerization (such as a receptor tyrosine kinase), it has become popular to use cell-based assays that indirectly report such binding events via dimerization-induced transcription of a reporter gene encoding a fluorescent or chemiluminescent protein. An example is the Tango assay of U.S. Patent 7,049,076, which provides a transcription factor fused to a first member of the pair of interacting proteins via a protease recognition sequence that can be cleaved via an exogenous protease fused to a second member of the pair when the two proteins bind. The released transcription factor then translocates to the cell nucleus where it activates an exogenous gene to express a fluorescent or chemiluminescent protein in the cell, which in turn is detected via photometric methods.
[0008] However, for multiplexed applications such as pharmome mapping, such a transcription-based reporter method presents serious limitations, since the number of suitable fluorescent or chemiluminescent reporter proteins with well-separated excitation and / or emission spectra is low. Further, the processes involved in gene transcription are complex, depending on a host of interdependent molecular / biochemical reactions / interactions, any one of which might be interfered with by the test compound, and are further tightly regulated by equally complex pathways, presenting myriad opportunities for the assay to generate a false negative or false positive result in response to a test compound that modulates unintended targets involved in transcription of the reporter gene. Finally, depending upon the assay conditions and reporter cell type employed, reporter gene expression can be a slow process requiring extended (overnight or longer) incubation of the cells with test compound to allow sufficient accumulation of the reporter protein, presenting difficulties for the screening of more labile compounds.
[0009] Thus there exists a need within the field of HTS for cell-based assay technologies that are inherently capable of being multiplexed, with respect to targets, without requiring test compound-induced activation of reporter gene transcription.
[0010] Summary
[0011] In a first aspect, the disclosure provides recombinant cells comprising: (a) a first recombinant nucleic acid molecule comprising a nucleotide sequence encoding a first fusion protein operatively linked to a promoter, wherein the first fusion protein comprises:
[0012] (i) a target protein;
[0013] (ii) a protease recognition sequence and cleavage site;
[0014] (iii) an engineered splicing factor (ESF), comprising:
[0015] (1) a splicing exclusion sequence;
[0016] (2) a nuclear localization signal; and
[0017] (3) a sequence-specific RNA-binding domain, wherein the sequence- specific RNA-binding domain is capable of binding specifically to a nucleotide sequence within a test exon to be excluded; wherein the protease recognition sequence and cleavage site is located between the target protein and the ESF;
[0018] (b) a second recombinant nucleic acid molecule, comprising a nucleotide sequence encoding a second fusion protein operatively linked to a promoter, wherein the second fusion protein comprises:
[0019] (i) an accessory protein, wherein the accessory protein is capable of binding to the target protein as a function of the target protein’s activity state; and
[0020] (ii) a protease capable of specifically recognizing the protease recognition sequence and cleaving the cleavage site; and
[0021] (c) a third recombinant nucleic acid molecule, comprising a nucleotide sequence encoding a reporter unspliced mRNA (RUM) operatively linked to a promoter, wherein the RUM comprises in 5’ to 3’ order
[0022] (i) a nucleotide sequence encoding a 5 ’ portion of a reporter protein gene, terminating in an exon of the reporter gene;
[0023] (ii) a first intron;
[0024] (iii) a nucleotide sequence comprising a test exon;
[0025] (iv) a second intron;
[0026] (v) a nucleotide sequence coding for the remaining 3’ portion of the reporter protein gene, beginning with an exon of the reporter gene; wherein expression of the test exon in the reporter protein disrupts function of the reporter protein, wherein the test exon comprises the nucleotide sequence bound specifically by the sequence-specific RNA binding domain. In one embodiment, the first fusion protein comprises, in amino terminal to carboxy terminal order, the target protein, the protease recognition sequence and cleavage site, and the ESF. In another embodiment, the protease recognition sequence and cleavage site is not recognized by proteases endogenously present in the cell, and wherein the protease is not endogenously expressed in the cell. In a further embodiment, the second fusion protein comprises, in amino terminal to carboxy terminal order, the accessory protein and the protease. In one embodiment, the test exon is not present in a native sequence of the reporter protein gene. In another embodiment, the target protein is a plasma membrane-associated protein, including but not limited to a transmembrane protein. In a further embodiment, the target protein is a cell surface receptor protein.
[0027] In one embodiment, the protease recognition sequence and cleavage site is bound and cleaved by tobacco etch virus protease, and wherein the protease comprises tobacco etch virus protease. In another embodiment, the nucleotide sequence encoding the protease recognition and cleavage site encodes the amino acid sequence selected from the group consisting of SEQ ID NO:5-7, or variants thereof. In a further embodiment, the nucleotide sequence encoding the protease recognition and cleavage site comprises the nucleotide sequence of SEQ ID NO:3 or 4, or variants thereof. In on embodiment, the nucleotide sequence encoding the protease encodes the amino acid sequence of SEQ ID NO:2, 81 , or variants thereof. In another embodiment, the nucleotide sequence encoding the protease comprises the nucleotide sequence of SEQ ID NO: 1, 62, or variants thereof. In a further embodiment, the cell is a U2OS cell, an human embryonic kidney (HEK)293 cell, or a Chinese hamster ovary (CHO) cells.
[0028] In one embodiment, the nucleotide sequence encoding the splicing exclusion sequence encodes the amino acid sequence selected from the group consisting of SEQ ID NO:9, 18, 20, and 2, or variants thereof. In another embodiment, the nucleotide sequence encoding the splicing exclusion sequence comprises the nucleotide sequence selected from the group consisting of SEQ ID NO:8, 17, 19, and 21, or variants thereof. In a further embodiment, the nucleotide sequence encoding the nuclear localization signal encodes the amino acid sequence selected from the group consisting of SEQ ID NO: 11 , 24, 26, and 28, or variants thereof. In a still further embodiment, the nucleotide sequence encoding the nuclear localization signal comprises the nucleotide sequence selected from the group consisting of SEQ ID NO: 10, 25, 23, and 27, or variants thereof.
[0029] In one embodiment, the nucleotide sequence encoding the sequence-specific RNA- binding domain encodes the amino acid sequence selected from the group consisting of SEQ ID NO: 13, 69, or variants thereof. In another embodiment the nucleotide sequence encoding the sequence-specific RNA-binding domain comprises the nucleotide sequence selected from the group consisting of SEQ ID NO: 12, 50 or variants thereof. In a further embodiment the nucleotide sequence bound specifically by the sequence-specific RNA-binding domain comprises the nucleotide sequence of SEQ ID NO:29, or variants thereof.
[0030] In one embodiment the second fusion protein comprises, in amino terminal to carboxy terminal order, (i) the accessory protein, and (ii) the protease. In another embodiment, the second fusion protein comprises, in amino terminal to carboxy terminal order, (i) the protease, and (ii) the accessory protein. In a further embodiment the nucleotide sequence encoding the test exon comprises the nucleotide sequence selected from the group consisting of SEQ ID NO: 16 substituted or inserted with a sequence (including but not limited to TGTATATA or TGTATATAX where X is any nucleotide) to which the sequence- specific RNA-binding domain binds, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:51. SEQ ID NO:53, SEQ ID NO:55, or variants thereof.
[0031] In one embodiment the 5’ portion of the reporter protein gene and the 3’ portion of the reporter protein gene, combined, encode a reporter protein. In another embodiment, the nucleotide sequence encoding the introns are selected from the group consisting of SEQ ID NO: 14-15, or variants thereof. In a further embodiment one, two, or all three of the first recombinant nucleic acid molecule, the second recombinant nucleic acid molecule, and the third recombinant nucleic acid molecule are stably integrated in the recombinant cell genome. In a still further embodiment the first recombinant nucleic acid molecule, the second recombinant nucleic acid molecule, and the third recombinant nucleic acid molecule are comprised in one or more expression vectors. In one embodiment the first recombinant nucleic acid molecule, the second recombinant nucleic acid molecule, and the third recombinant nucleic acid molecule are comprised in a single expression vector. In another embodiment the first recombinant nucleic acid molecule, the second recombinant nucleic acid molecule, and the third recombinant nucleic acid molecule are comprised in two or three expression vectors. In a further embodiment the first recombinant nucleic acid molecule, the second recombinant nucleic acid molecule, and the third recombinant nucleic acid molecule are each comprised in a separate expression vector. In another embodiment the recombinant cells is a mammalian cell.
[0032] In a second aspect, the disclosure provides compositions, comprising a plurality of the recombinant cells of any embodiment or combination of embodiments of the first aspect. In one embodiment, all of the recombinant cells comprise a first recombinant nucleic acid molecule encoding the same target protein and a second nucleic acid molecule encoding the same accessory protein. In another embodiment, all of the recombinant cells comprise the same first recombinant nucleic acid molecule, second recombinant nucleic acid molecule, and third recombinant nucleic acid molecule. In a further embodiment, the plurality of recombinant cells comprises sub-populations of recombinant cells, wherein each sub- population comprises a different first recombinant nucleic acid molecule encoding a different target protein. In one embodiment, each sub-population comprises a different second recombinant nucleic acid molecule encoding a different accessory protein. In another embodiment, each sub-population comprises (i) a first recombinant nucleic acid molecule that encodes the same protease recognition sequence and cleavage site; and (ii) a second recombinant nucleic acid molecule that encodes the same protease. In a further embodiment, one or more sub-populations comprise (i) a first recombinant nucleic acid molecule that encodes a protease recognition sequence and cleavage site that differs from the protease recognition sequence and cleavage site encoded by a first recombinant nucleic acid molecule in other sub-populations, and (ii) a second recombinant nucleic acid molecule that encodes a protease that differs from the protease encoded by a second recombinant nucleic acid molecule in other sub-populations. In another embodiment each sub-population comprises a first recombinant nucleic acid molecule that encodes the same ESF. In one embodiment, one or more sub-populations comprise a first recombinant nucleic acid molecule that encodes an ESF that differs from the ESF encoded by a first recombinant nucleic acid molecule in other sub-populations. In another embodiment, each sub-population comprises a third recombinant nucleic acid molecule that encodes the same RUM. In a further embodiment, one or more sub-populations comprise a third recombinant nucleic acid molecule that encodes a RUM that differs from the RUM encoded by a third recombinant nucleic acid molecule in other sub- populations.
[0033] In one embodiment of the compositions, the plurality of recombinant cells comprises sub-populations of recombinant cells, wherein each sub-population comprises:
[0034] (a) a different first recombinant nucleic add molecule encoding (i) a different target protein, (ii) an identical protease recognition sequence and cleavage site, and (iii) an identical ESF;
[0035] (b) a different second recombinant nucleic acid molecule encoding (i) a different or identical accessory protein, and (ii) an identical protease;
[0036] (c) a third recombinant nucleic acid molecule encoding a RUM, wherein the RUM comprises, in 5’ to 3’ order: (i) the nucleotide sequence encoding the 5’ portion of a reporter protein gene, terminating in an exon of the reporter gene;
[0037] (ii) an identical first intron;
[0038] (iii) a nucleotide sequence comprising an identical test exon;
[0039] (iv) an identical second intron; and
[0040] (v) the nucleotide sequence coding for the remaining 3’ portion of the reporter protein gene, beginning with an exon of the reporter gene.
[0041] In another embodiment, each sub-population comprises a different third recombinant nucleic acid molecule, wherein:
[0042] (a) a first region is a portion of the nucleotide sequence encoding the 5’ portion of the reporter protein gene and a second region is a portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene, wherein one or both of the first region and the second region differ in each sub-population through use of degenerate codons, wherein upon splicing of an mRNA expression product of the third recombinant nucleic acid molecule that excises the test exon the first region and the second region are directly adjacent and forma first nucleic acid barcode; and
[0043] (b) the reporter protein encoded by the combination of the 5’ portion of the reporter protein gene and the nucleotide sequence of the 3’ portion of the reporter protein gene is the same in each sub-population.
[0044] In one such embodiment, both of the first region and the second region differ in each sub-population through use of degenerate codons. In another embodiment, a third region comprises a portion of the test exon, wherein upon splicing of an mRNA expression product of the third recombinant nucleic acid molecule that does not excise the test exon, the third region and either (a) a portion of the nucleotide sequence encoding the 5’ portion of the reporter protein gene, or (b) a portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene, are directly adjacent and form a second nucleic acid barcode, wherein one or both of (i) the third region and (ii) the portion of the nucleotide sequence encoding the 5’ portion of the reporter protein gene, or (b) a portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene differ in each sub-population through use of degenerate codons. In one such embodiment, both of the (i) the third region and (ii) the portion of the nucleotide sequence encoding the 5 ’ portion of the reporter protein gene, or (b) a portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene differ in each sub-population through use of degenerate codons. In another embodiment, (a) the portion of the nucleotide sequence encoding the 5’ portion of the reporter protein gene comprises the first region, or (b) the portion of the nucleotide sequence encoding the 3’ portion of the reporter protein comprises the second region.
[0045] In one embodiment, a third region comprises a portion of the test exon, wherein the third region differs in each sub-population through use of degenerate codons, wherein upon splicing of an mRNA expression product of the third recombinant nucleic acid molecule that does not excise the test exon, the third region and either the first region or the second region are directly adjacent and form a second nucleic acid barcode.
[0046] In another embodiment of the compositions, a fourth region comprises:
[0047] (a) a portion of the first intron,
[0048] (b) a portion of the second intron,
[0049] (c) a portion of the first intron, and (ii) a portion of the test exon, or the third region,
[0050] (d) or a portion of second intron, and (ii) a portion of the test exon, or the third region,
[0051] (e) a portion of the first intron, and (ii) a portion of the nucleotide sequence encoding the 5’ portion of the reporter protein gene, or the first region,
[0052] (f) a portion of second intron, and (ii) a portion of the nucleotide sequence encoding the 5 ’ portion of the reporter protein gene, or the first region
[0053] (g) a portion of the first intron, and (ii) a portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene, or the second region; or
[0054] (h) a portion of second intron, and (ii) a portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene, or the second region; wherein the fourth region differs in each sub-population through use of degenerate codons, wherein the fourth region is a third nucleic acid barcode. In some embodiments where the fourth region includes a portion of more than one domain, one or both domain portions may differs in each sub-population through use of degenerate codons. In one embodiment where the fourth region includes a portion of more than one domain, both domain portions may differs in each sub-population through use of degenerate codons
[0055] In one embodiment, a fourth region comprises
[0056] (a) a portion of the first intron,
[0057] (b) a portion of the second intron,
[0058] (c) (i) a portion of the first intron or a portion of second intron, and (ii) the third region,
[0059] (d) (i) a portion of the first intron or a portion of second intron, and (ii) the first region, or (e) (i) a portion of the first intron or a portion of second intron, and (ii) the second region, wherein the fourth region differs in each sub-population through use of degenerate codons, wherein the fourth region is a third nucleic acid barcode.
[0060] In another embodiment, each nucleic acid barcode is at least 6 nucleotides in length. In a further embodiment, each nucleic acid barcode is between 6 nucleotides and 2000 nucleotides in length. In a further embodiment, in each sub-population, (i) the first recombinant nucleic acid comprises the same promoter; (ii) the second recombinant nucleic acid comprises the same promoter; and (iii) the third recombinant nucleic acid comprises the same promoter.
[0061] In a third aspect, the disclosure provides methods for screening of a target proteins, comprising:
[0062] (a) contacting the recombinant cell of any embodiment of the first aspect of the disclosure with a test compound or test stimulus;
[0063] (b) detecting reporter protein activity and / or abundance of mRNA expressed from the third recombinant nucleic acid molecule.
[0064] In one embodiment, the method is for multiplexed screening of two or more target proteins, comprising:
[0065] (a) contacting the composition of any embodiment of the second aspect of the disclosure with a test compound or test stimulus;
[0066] (b) detecting reporter protein activity' and / or abundance of mRNA expressed from the third recombinant nucleic acid molecule.
[0067] In some embodiments, the methods further comprise
[0068] (c) measuring an amount of at least the first nucleic acid barcode to separately determine activities of each of the target proteins in the composition.
[0069] In one such embodiment, step (c) comprises measuring an amount the first nucleic acid barcode, the second nucleic acid barcode, and the third nucleic acid barcode, to separately determine activities of each of the target proteins in the composition. In another embodiment, the mRNA expressed from the third recombinant nucleic acid molecule comprises 1, 2, or all 3 of (i) unspliced mRNA reported on by the third nucleic acid barcode, (ii) spliced mRNA in which the test exon is excised (“short-form mRNA”) reported on by the first nucleic acid barcode, and (iii) spliced mRNA in which the test exon is not excised (“long-form mRNA”) reported on by the second nucleic acid barcode. In one embodiment of any of the methods of the disclosure, detecting an abundance of mRNA expressed from the third recombinant nucleic acid molecule comprises detecting a relative abundance of the unspliced mRNA, the short-form mRNA, and the long-form mRNA. In another embodiment, a sum of the abundance of the unspliced mRNA, the short- form mRNA, and the long-form mRNA is determined to normalize individual abundance detected for the unspliced mRNA, the short-form mRNA, and the long-form mRNA, and / or is used as a housekeeping signal to provide a relative measure of cell health, cell number, or reporter gene transcription activity and / or mRNA turnover.
[0070] Description of the Figures
[0071] Figure 1. Reporter cell 100 line of the disclosure. A: target protein 101 is a fusion protein comprising (from top to bottom) the native target protein of interest, a linker sequence bearing a protease recognition sequence, and an engineered splicing factor (ESF). In the absence of interaction with a suitable test compound 107, accessory protein 102 (a fusion protein comprising the native accessory protein and a protease) does not interact with the target protein. The ESF is thus unable to enter the nucleus 103. In the absence of nuclear ESF, the constitutively expressed reporter unspliced form mRNA 104 (RUM; see Fig 2 for definition of its sub-sequences) is spliced to include an exogenous (non-native) exon disrupting the reporter protein’s native structure, such that when the spliced mRNA (also known as the reporter gene product’s ‘long form’) is translated it yields an inactive version of the reporter protein 106. B: When a test compound 107 capable of activating the target protein binds to the target protein, accessory protein binds to the target protein, enabling its protease component to cleave the test protein’s linker sequence, thus freeing the unlinked ESF 108 to translocate to the nucleus and bind the ESF binding sequence of the reporter protein mRNA. Splicing of this complex produces the natively spliced form (also known as the ‘short form’) of the mRNA excluding the exogenous exon. The natively spliced mRNA in turn is translated into active detectable reporter protein 110, whose detected abundance is proportional to target protein activation by the test compound.
[0072] Figure 2. Nucleotide sequence barcodes that are unique to each of the three splice forms of the reporter mRNA (the unspliced RUM, the short-form spliced mRNA containing only reporter protein exons encoding active protein and lacking the test exon, and the long- form mRNA encoding inactive reporter protein due to the presence of an interfering test exon sequence) are either present in the RUM (Barcode A in the figure) or are formed by the splicing-mediated concatenation of two exonic segments of RUM mRNA previously separated by two introns flanking the test exon (Barcodes B and C in the figure). In one embodiment of the disclosure the test exon, comprising an ESF-binding sequence and flanked by two introns, is interposed between the 5’ and 3’ exons of the reporter protein. During splicing, ESF binding to the test exon results in its exclusion from the spliced mRNA, thus concatenating the reporter protein’s 5’ and 3’ exons, which creates an uninterrupted nucleotide sequence (Barcode B) that is unique to the short-form mRNA. In contrast, in the absence of ESF binding, splicing concatenates the reporter protein’s 3’ exons with the test exon, creating a unique uninterrupted nucleotide sequence, Barcode C, which is unique to the long-form mRNA. In either case (+ or - ESF binding) Barcode A, which comprises at least some intronic RUM sequence, is unique to the unspliced RUM
[0073] Figure 3. Method of encoding a single species of reporter protein via two or more distinctly barcoded spliced mRNAs (see Fig. 2 for definitions of the sequence shadings used here). Reporter unspliced mRNAs (RUMs) 1 (top) and 2 (bottom) both encode the same reporter protein, but employ different degenerate codon sequences (A is degenerate with respect to C, and B is degenerate with respect to D) at feeing ends of the exons flanking the central exogenous exon. Exclusion of the exogenous exon during splicing in the presence of the cognate engineered splicing factor creates distinct mRNAs, both coding for the same active reporter protein but bearing distinct internal barcodes 1 (AB) and 2 (CD) (for example without limitation), which may be detected and quantified via methods such as nucleic acid hybridization.
[0074] Figure 4. RT-PCR gel electrophoresis results. Large bands (795 bp) correspond to SRs without exon exclusion, while small bands (720 bp) correspond to SRs with exon exclusion. Notably, exon exclusion is promoted in conditions where ESF variant and SR variant are expected to bind with high affinity (e.g., Lane 6 and Lane 10). Note that, in all conditions, introns are removed from the sequence encoded by the plasmid (2523 bp). Lane 1: 50 bp Ladder, Lane 2: SR-No-BS, Lane 3: SR-No-BS + ESF-WT, Lane 4: SR-No-BS + ESF-Mut, Lane 5: SR-WT-BS, Lane 6: SR-WT-BS + ESF-WT, Lane 7: SR-WT-BS + ESF- Mut, Lane 8: SR-Mut-BS, Lane 9: SR-Mut-BS + ESF-WT, Lane 10: SR-Mut-BS + ESF-Mut.
[0075] Figure 5. Gel electrophoresis results. Large bands (795 bp) correspond to SRs without exon exclusion, while small bands (720 bp) correspond to SRs with exon exclusion. Notably, exon exclusion is promoted in conditions where agonist activates receptor (e.g., clonidine activation of ADRA2C) (e.g., Lane 3, Lane 6). Lane 1: 50 bp Ladder, Lane 2: DMSO treatment, ADRA2C-ESF, Lane 3: DMSO treatment, MTNR1A-ESF, Lane 4: Clonidine treatment, ADRA2C-ESF, Lane 5: Clonidine treatment, MTNR1A-ESF, Lane 6: Melatonin treatment, ADRA2C-ESF, Lane 7: Melatonin treatment, MTNR1A-ESF
[0076] Figure 6. Gel electrophoresis results. Large bands (795 bp) correspond to SRs without exon exclusion, while small bands (720 bp) correspond to SRs with exon exclusion. Notably, exon exclusion is promoted in conditions where agonist are expected to activate receptor (e.g., clonidine activation of ADRA2C) and the receptor-SR primer pairing is used (e.g., Lane 3). Lane 1: 50 bp Ladder, Lane 2: DMSO treatment, ADRA2C SR Primers, Lane 3: DMSO treatment, MTNR1 A SR Primers, Lane 4: Clonidine treatment, ADRA2C SR Primers, Lane 5: Clonidine treatment, MTNR1A SR Primers, Lane 6: Melatonin treatment, ADRA2C SR Primers, Lane 7: Melatonin treatment, MTNR1A SR Primers.
[0077] Detailed Description
[0078] All references cited are herein incorporated by reference in their entirety. Within this application, unless otherwise stated, the techniques utilized may be found in any of several well-known references such as: Molecular Cloning: A Laboratory Manual (Sambrook, et al.,
[0079] 1989, Cold Spring Harbor Laboratory Press), Gene Expression Technology (Methods in Enzymology, Vol. 185, edited by D. Goeddel, 1991. Academic Press, San Diego, CA), “Guide to Protein Purification” in Methods in Enzymology (M.P. Deutshcer, ed., (1990) Academic Press, Inc.); PCR Protocols: A Guide to Methods and Applications (Innis, et al.
[0080] 1990. Academic Press, San Diego, CA), Culture of Animal Cells: A Manual of Basic Technique, 2ndEd. (R.I. Freshney. 1987. Liss, Inc. New York, NY), Gene Transfer and Expression Protocols, pp. 109-128, ed. EJ. Murray, The Humana Press Inc., Clifton, N.J.), Dang, B. et al. SNAC-tag for sequence-specific chemical protein cleavage. Nat. Methods 16, 319-322 (2019), and the Ambion 1998 Catalog (Ambion, Austin, TX).
[0081] As used herein, the singular forms "a", "an" and "the" include plural referents unless the context clearly dictates otherwise.
[0082] As used herein, “about” means + / - 5% of the recited parameter.
[0083] All embodiments of any aspect of the disclosure can be used in combination, unless the context clearly dictates otherwise.
[0084] Unless the context clearly requires otherwise, throughout the description and the claims, the words ‘comprise’, ‘comprising’, and the like are to be construed in an inclusive sense as opposed to an exclusive or exhaustive sense; that is to say, in the sense of “including, but not limited to”. Words using the singular or plural number also include the plural and singular number, respectively. Additionally, the words ‘ herein,” “above,” and “below” and words of similar import, when used in this application, shall refer to this application as a whole and not to any particular portions of the application.
[0085] As used herein, “intron” and “intronic” have their familiar meaning in the art, referring to a non-protein-coding segment of pre-mRNA sequence spliced out from surrounding exonic sequences to yield the mature spliced protein-coding mRNA sequence. In humans, introns are typically characterized by three short conserved sequences, including (where N is any RNA base and Y is any RNA pyrimidine base) a 5’ GURAGN or GUNNNN, an internal branch point sequence YNYURAY, and a 3’ YAG.
[0086] As used herein, “exon” and “exonic” have their familiar meaning in the art, referring to a protein-coding segment of pre-mRNA frequently (but not always) preserved in the mature spliced mRNA sequence after introns are removed by splicing. In some alternatively spliced forms of an mRNA one or more exons may be spliced out of the mature mRNA along with their surrounding introns.
[0087] As used herein, a “housekeeping signal” means a constitutively expressed RNA’s measured abundance or the sum of a constitutively expressed RNAs’ various splice forms’ measured abundances, which is sufficiently constant and independent of a healthy cell’s state to serve as a usefill index of cell number and / or cell viability.
[0088] In one aspect, the disclosure provides a recombinant cell comprising:
[0089] (a) a first recombinant nucleic acid molecule comprising a nucleotide sequence encoding a first fusion protein operatively linked to a promoter, wherein the first fusion protein comprises:
[0090] (i) a target protein;
[0091] (ii) a protease recognition sequence and cleavage site;
[0092] (iii) an engineered splicing factor (ESF), comprising:
[0093] (1) a splicing exclusion sequence;
[0094] (2) a nuclear localization signal; and
[0095] (3) a sequence-specific RNA-binding domain, wherein the sequence- specific RNA-binding domain is capable of binding with high specificity to a nucleotide sequence within a test exon to be excluded; wherein the protease recognition sequence and cleavage site is located between the target protein and the ESF;
[0096] (b) a second recombinant nucleic acid molecule, comprising a nucleotide sequence encoding a second fusion protein operatively linked to a promoter, wherein the second fusion protein comprises: (i) an accessory protein, wherein the accessory protein is capable of binding to the target protein as a function of the target protein’s activity state; and
[0097] (ii) a protease capable of specifically recognizing the protease recognition sequence and cleaving the cleavage site; and
[0098] (c) a third recombinant nucleic acid molecule, comprising a nucleotide sequence encoding a reporter unspliced mRNA (RUM) operatively linked to a promoter, wherein the RUM comprises in 5’ to 3’ order
[0099] (i) a nucleotide sequence encoding a 5’ portion of a reporter protein gene, terminating in an exon of the reporter gene;
[0100] (ii) a first intron;
[0101] (iii) the nucleotide sequence comprising the test exon;
[0102] (iv) a second intron; and
[0103] (v) a nucleotide sequence coding for the remaining 3’ portion of the reporter protein gene, beginning with an exon of the reporter gene; wherein expression of the test exon in the reporter protein disrupts function of the reporter protein, wherein the test exon comprises the nucleotide sequence bound specifically by the sequence-specific RNA -binding domain.
[0104] An exemplary embodiment of a recombinant cell of the disclosure is provided in Figure 1. The cells can be used, for example, in the methods of the disclosure as detailed below. By way of non-limiting example, the cells can be used in multiplexed cell-based assay technologies, with respect to targets, without requiring test compound-induced activation of reporter gene transcription.
[0105] In one embodiment, the first fusion protein comprises, in amino terminal to carboxy terminal order, the target protein, the protease recognition sequence and cleavage site, and the ESF. In another embodiment, the first fusion protein comprises, in amino terminal to carboxy terminal order, the ESF, the protease recognition sequence and cleavage site, and the target protein.
[0106] As will be understood by those of skill in the art, the target protein may be any target protein as determined by an end user, for which a protein capable of binding to the target protein as a function of the target protein’s activity state is available, and whose fusion protein with the ESF either does not enter the nucleus or else does not productively participate in intra-nuclear mRNA splicing. In one embodiment, the target protein is a plasma membrane-associated protein, including but not limited to a transmembrane protein. In another embodiment, the target protein is a cell surface receptor protein. In a further embodiment, the target protein may be an intracellular protein. In various non-limiting embodiments, the target protein may be selected from the group consisting of G protein- coupled receptors, nuclear hormone receptors, receptor tyrosine kinases, receptor serine- threonine kinases, cytosolic protein kinases, transcription fectors, ATP binding cassette family member proteins, adenylate cyclases, ATP-dependent membrane transporters, B cell leukemia / lymphoma 2 (BCL2), BCL2 -related proteins, BCL2-like proteins, basic helix-loop- helix family proteins, caveolins, CCAAT / enhancer binding proteins, centromere proteins, diacylglycerol kinases, Fc receptor peptides, frizzled class receptors, histones, Jun dimerization proteins, mitogen activated protein kinases, neuronal PAS domain proteins, phosphoinositide-3-kinase subunits, DNA polymerases, Rai GTPase activating protein subunits, solute carriers, synaptotagmins, TATA-box binding protein associated fectors, teneurins, tropomyosins, tubulins, proteases, RNA splicing factors, cell adhesion proteins, cytoskeletal proteins, ionotropic receptors, voltage-gated ion channels, membrane transporters, enzymes of intermediary' metabolism, mitochondrial proteins, nuclear proteins, endoplasmic reticulum proteins, lysosomal proteins, endocytotic vesicular proteins, exocytotic vesicular proteins, Golgi proteins, synaptic proteins, dendritic proteins, cytosolic proteins, antiporters, symporters, phosphoprotein binding proteins, phospholipases, and sortilin.
[0107] Any suitable protease recognition sequence and cleavage site may be used. In one embodiment, the protease recognition sequence and cleavage site is not recognized by proteases endogenously present in the cell, and wherein the protease is not endogenously expressed in the cell (i.e., not present in the cell before the cell is modified to recombinantly express the protease). In non-limiting embodiments, the protease recognition sequence and cleavage site may be bound and cleaved by tobacco etch virus protease (TEVP), and the protease comprises TEVP. In one such embodiment, a nucleotide sequence encoding fee protease recognition sequence and cleavage site comprises a nucleotide sequence encoding fee amino acid sequence selected from fee group consisting of SEQ ID NO:5-7 (see Table 1). In these embodiments, cleavage occurs between fee “Q” residue and fee final residue (which may be any amino acid). For example, fee wild type sequence has an “S” residue as fee final residue (SEQ ID NO:7), but in order to decrease basal cleavage substitutions such as “L” among others can be used (SEQ ID NO:6). In another embodiment, a nucleotide sequence encoding fee protease recognition sequence and cleavage site comprises fee nucleotide sequence of SEQ ID NO:3 or 4. In some embodiments, fee nucleotide sequence encoding fee protease encodes fee amino acid sequence of SEQ ID NO:2 or 81. In another embodiment, the nucleotide sequence encoding the protease comprises the nucleotide sequence of SEQ ID
[0108] NO: l or 62.
[0109] Table 1
[0110] The last codon in SEQ ID NO:4, here shown as tct encoding a serine, is the consensus sequence for this cleavage site; however, alternative amino acids are tolerated in the peptide, and may be desirable in assays for purposes such as lowering basal protease cleavage levels.
[0111] Thus, SEQ ID NO:3 permits any nucleotides for encoding the final amino acid. For example, in embodiments where SEQ ID NO:4 encodes the amino acid sequence of SEQ ID NO:6, the
[0112] “xxx” nucleotides may be selected from TTA, TTG, CTA, CTG, CTT, and CTC.
[0113] Non-limiting examples of mammalian cell lines in which TEVP and its cognate protease recognition and cleavage site may be used without evidence of endogenous protease interference or TEVP proteolysis of endogenous proteins include U2OS, HEK293, and CHO cells.
[0114] The engineered splicing factor (ESF) is a protein that promotes the test exon’s exclusion during mRNA splicing by binding to a nucleotide sequence located within the test exon. The ESF comprises (1) a splicing exclusion sequence; (2) a nuclear localization signal; and (3) a sequence-specific RNA-binding domain, wherein the sequence-specific
[0115] RNA-binding domain is capable of binding specifically to a nucleotide sequence within a test exon to be excluded. The ESF that is released upon protease cleavage (i.e., when carrying out the methods of the disclosure) comprises a splicing exclusion sequence that directs the exclusion of the test exon (described below) from the mature RNA molecule, fused to a sequence-specific RNA-binding domain via a nuclear localization signal.
[0116] A splicing exclusion sequence is an amino acid sequence designating the test exon to which it is bound for exclusion from the spliced mRNA (and thus from its translated protein).
[0117] Any suitable splicing exclusion sequence may be used. Those skilled in the art will recognize that suitable splicing exclusion sequences include the glycine-rich domains of the heterogeneous nuclear ribonucleoprotein (hnRNP) family. In one non-limiting embodiment, the nucleotide sequence encoding the splicing exclusion sequence encodes the amino acid sequence of SEQ ID NO:9. In another embodiment the amino acid sequence is that of SEQ ID NO: 18. In still another embodiment the amino acid sequence is that of SEQ ID NO:20. Notably, even some arbitrary synthetic glycine-rich sequences are known in the art to serve as splicing exclusion sequences, as demonstrated, for example, in Y. Wang et al. (2009) Nat Methods. 6: 825 (SEQ ID NO:22). Exemplary nucleic acids encoding the domains are provided as SEQ ID NO: 17, 19, and 21.
[0118] Those skilled in the art will recognize that any of a wide variety of nuclear localization signals (NLS) may be used in the ESF to ensure its translocation to the nucleus following its proteolytic release from the target protein. In one non-limiting embodiment, the amino acid sequence of the SV40 NLS may be used (SEQ ID NO: 11). In another embodiment, the amino acid sequence of the vasopressin activated calcium mobilizing receptor-like protein NLS may be used (SEQ ID NO:24. In still another embodiment, the amino acid sequence of the chicken anemia virus NLS may be used (SEQ ID NO:28). Exemplary nucleic acids encoding these NLS are provided as SEQ ID NO: 10, 23 and 27)
[0119] The sequence-specific RNA-binding domain is capable of binding to a nucleotide sequence within a test exon to be excluded. In one embodiment, the nucleotide sequence encoding the sequence-specific RNA-binding domain encodes the amino acid sequence of the PUF domain of Pumilio homolog 1 (SEQ ID NO: 13), and the corresponding nucleotide sequence within the test exon to which it binds is the cognate sequence of the Nanos response element (SEQ ID NO:29). Other man-made RNA-binding domain amino acid sequences (derived from modifications to the native PUF domain sequence), and their cognate RNA nucleotide binding sequences, are well known in the art (Y. Wang et al. 2009, Nat Methods. 6: 825; US Patent 9,499,805 B2; US Patent 10,330,674 B2, US Patent 11,275,081 B2, US Patent Application 17 / 029,666). Exemplary nucleic acids encoding these domains are provided as SEQ ID NO: 12, 29, and 50. In some embodiments the sequence-specific RNA- binding domain binds to its cognate RNA sequence with nanomolar Kd. In some other embodiments it binds with sub-nanomolar Kd.
[0120] Table 2
[0121] The second recombinant nucleic acid molecule comprises a nucleotide sequence encoding a second fusion protein comprising (i) an accessory protein, wherein the accessory protein is capable of binding to the target protein as a function of the target protein’s activity state; and (ii) a protease capable of specifically recognizing the protease recognition sequence and cleaving the cleavage site. In some embodiments, the second fusion protein comprises, in amino terminal to carboxy terminal order, (i) the accessory protein, and (ii) the protease. In other embodiments, the second fusion protein comprises, in amino terminal to carboxy terminal order, (i) the protease, and (ii) the accessory' protein.
[0122] The protease and exemplary embodiments thereof are discussed above. The accessory protein may be any accessory protein that binds to target protein being used as a function of the target protein’s activity' state. In various non-limiting embodiments, (a) the accessory protein may bind the target protein only when the target protein is bound to ligand;
[0123] (b) the accessory protein may bind the target protein only when the target protein is not bound to ligand; (c) the accessory protein may bind the target protein only when the target protein is multimerized; (d) the accessory protein may bind the target protein only when the target protein is not multimerized (e) the accessory protein may bind the target protein only when the target protein is phosphorylated; (f) the accessory protein may bind the target protein only when the target protein is not phosphorylated; (g) the accessory protein and the target protein may be two or more molecules of the same protein that reversibly dimerize or multimerize when one or more of them is bound to ligand; or (h) the accessory protein and the target protein may be two or more molecules of the same protein that reversibly dimerize or multimerize when one or more of them is not bound to ligand.
[0124] In various non-limiting embodiments, a target protein-accessory protein pair may include, but are not limited to those shown in Table 3. Those of skill in the art will be able to identify other such target protein-accessory protein pairs based on the teachings herein.
[0125] Table 3
[0126] The third recombinant nucleic acid molecule comprises a nucleotide sequence encoding a reporter unspliced mRNA (RUM) operatively linked to a promoter, wherein the RUM comprises in 5’ to 3’ order:
[0127] (i) a nucleotide sequence encoding a 5 ’ portion of a reporter protein gene, terminating in an exon of the reporter gene;
[0128] (ii) a first intron; (iii) the nucleotide sequence comprising the test exon;
[0129] (iv) a second intron; and
[0130] (v) a nucleotide sequence coding for the remaining 3’ portion of the reporter protein gene, beginning with an exon of the reporter gene; wherein expression of the test exon in the reporter protein disrupts function of the reporter protein, wherein the test exon comprises the nucleotide sequence bound specifically by the sequence-specific RNA-binding domain
[0131] In these embodiments, the nucleotide sequence encoding the 5’ portion of a reporter protein gene, terminating in an exon of the reporter gene and the remaining 3’ portion of the reporter protein gene, beginning with an exon of the reporter gene, when joined in the absence of the test exon on a mature RNA (after splicing of the RUM), encode a functional reporter protein. Any protein reporter gene may be used that, upon expression, is detectable with a quantifiable signal that parallels the protein’s abundance. Exemplary such reporter proteins include but are not limited to a Beta-galactosidase; a luciferase; a green fluorescent protein, yellow fluorescent protein, red fluorescent protein, ECFP, DsRed2FP, EGFP, mTurquoise™, mVenus™, mCherry™, or any fluorescent protein disclosed in www.microscopyu.com / techniques / fluorescence / introduction-to-fluorescent-proteins.
[0132] Introns are as defined above. Any intron may be used, as suitable for an intended purpose. In one embodiment, each intron is from the same gene. In other embodiments, the intron(s) are not from the same gene. In one non-limiting embodiment, the introns are selected from the group consisting of SEQ ID NO: 14-15, or portions thereof. Each intron may be identical, or may be different. The intron lengths may be modified (increased or decreased) as appropriate for an intended use, so long as they meet the definition of introns provided above. See Table 4 for the sequences.
[0133] The test exon is chosen so that its expression in the reporter protein will disrupt the function of the reporter protein (for example without exclusion, rendering the normally fluorescent reporter protein non-fluorescent). Examples of fluorescent or luminescent proteins bearing an interpolated exogenous exon (a test exon) which, when not spliced out, disrupts the protein’s ability to generate a fluorescent or luminescent signal are familiar to those skilled in the art as reporters used to study alternative splicing. Examples, without limitation, of such disruptive test exons inserted into green fluorescent protein or into luciferase include, along with IGF2BP1 exon 12, vascular endothelial growth factor exon 6A [R Wang et al (2009). Identification of an exonic splicing silencer in exon 6A of the human VEGF gene. BMC Molecular Biology 10: 103], alpha tropomyosin exon 2 or exon 3 [PD Ellis et al (2004). Regulated tissue-specific alternative splicing of enhanced green fluorescent protein transgenes conferred by alpha-tropomyosin regulatory elements in transgenic mice. J Biol Chem 279:36660], fibroblast growth factor receptor 2 exon 8 or exon 9 [JA Somarelli et al (2013). Fluorescence-based alternative splicing reporters for the study of epithelial plasticity in vivo. RNA 19: 116], dihydrofolate reductase exon 2 [Z Wang et al (2004). Systematic identification and analysis of exonic splicing silencers. Cell 119:831], and oncogene MDM2 exons 4, 10, and 11 [Y Shi et al (2016). A triple exon-skipping luciferase reporter assay identifies a new CLK inhibitor pharmacophore. Bioorg Med Chem Lett 27:406] . The test exon also contains the nucleotide sequence recognized by the sequence- specific RNA-binding domain. Any suitable such test exon may be used. In one embodiment, the test exon is not part of the native sequence of the reporter protein gene. In another non-limiting embodiment, the test exon comprises the nucleotide sequence of SEQ ID NO: 16, modified to include a sequence-specific RNA binding domain. See Table 4 for the sequence. The precise location, within the test exon, of the nucleotide sequence recognized by the sequence-specific RNA-binding domain may be decided by the user, and is usually chosen to be within 10 - 50 nucleotides from the alternative splice site. The sequence- specific RNA-binding domain may be added to or substituted for a portion of the test exon. In one non-limiting embodiment, the RNA binding domain is encoded by (TGTATATA, bolded). In this embodiment, the RNA binding sequence is eight nucleotides in length, and may substitute for eight nucleotides in the test exon nucleotide sequence (such as SEQ ID NO: 16), or be inserted with any additional flanking 5’ or 3’ nucleotide to maintain the test exon (such as SEQ ID NO: 16) open reading frame, and thus the integrity of the test exon’s 3’ exon / intron junction. SEQ ID NO:31 illustrates one example without limitation of a sequence of a test exon (IGF2BP1 exon 12) with an RNA-binding domain’s (Puf homology domain 1) cognate binding sequence inserted 20 nucleotides downstream of the test exon’s 5’ end, SEQ ID NO:31 is an exemplary “insertion” of the ESF binding sequence (plus an additional nt to maintain the ORF). SEQ ID NO:32 represents another example, in which the ESF binding sequence “substitutes” an 8 NA strand of the wild type test exon sequence. These are just two examples of how to incorporate the ESF binding sequence.
[0134] Table 4
[0135] Any suitable promoter may be used in the first, second, and third recombinant nucleic acid molecules. In one embodiment, the first, second, and / or third recombinant nucleic acid molecules’ expression in a test cell line may be under the control of an inducible promoter such as the tetracycline response element. In another embodiment, the first, second, and / or third recombinant nucleic acid molecules’ expression may be under the control of a constitutively active promoter such as the cytomegalovirus minimal promoter. In one embodiment, the first, second, and third recombinant nucleic acid molecules’ expressions are all under the control of the same promoter and / or response element. In another embodiment their expressions are under the control of two or more different promoters and / or response elements. In one embodiment, the third recombinant nucleic acid molecules’ expression may be under the control of a strong constitutive promoter, including but not limited to CMV, or an inducible promoter (e.g., Tet-dependent expression via TRE). See Table 5.
[0136] Non-limiting examples of 5’ and 3’ untranslated region components, response elements, and promoters whose sequences might be included in any or all of the first, second, and / or third recombinant nucleic acid molecules, are shown in Table 5. Also shown are non- limiting examples of antibiotic resistance genes that may be used in the disclosure (i.e., for selecting for and maintain cell lines stably expressing the desired recombinant nucleic acid molecules. Those of skill in the art will understand that variations to these sequences are permissible, and additional components, such as Kozak sequences, etc. may also be included.
[0137] Table 5 •These lengths can vary for reasons such as repeats, spacer sequences between repeats, and other optimizing conditions. BP = base pairs
[0138] In one embodiment, one, two, or all three of the first recombinant nucleic acid molecule, the second recombinant nucleic acid molecule, and the third recombinant nucleic acid molecule are stably integrated in the recombinant cell genome. In another embodiment, the first recombinant nucleic acid molecule, the second recombinant nucleic acid molecule, and the third recombinant nucleic acid molecule are comprised in one or more expression vectors. In another embodiment, the first recombinant nucleic acid molecule, the second recombinant nucleic acid molecule, and the third recombinant nucleic acid molecule are comprised in a single expression vector. In a further embodiment, the first recombinant nucleic acid molecule, the second recombinant nucleic acid molecule, and the third recombinant nucleic acid molecule are comprised in two or three expression vectors. In a still further embodiment, the first recombinant nucleic acid molecule, the second recombinant nucleic acid molecule, and the third recombinant nucleic acid molecule are each comprised in a separate expression vector. Any suitable expression vector may be used as appropriate for expression in a recombinant cell being used. Those of skill in the art will understand what expression vectors may be used in light of the teachings herein. In some embodiments, the expression vector is a plasmid vector. In other embodiments, the expression vector is a viral expression vector.
[0139] The recombinant cells of the disclosure may be of any type suitable for an intended use. In one embodiment, the recombinant cells is a mammahan cell. In various non-limiting embodiments, the mammalian cell may be selected from the group consisting of U-2 OS cells, HEK-293 cells, A549 cells, A431 cells, CHO cells, cancer cell lines, donor-derived or cultured neoplastic cells, primary cells of human or animal origin including without limitation hepatocytes, renal cells, fibroblasts, neurons, glial cells, myocytes, monocytes, lymphocytes, dendritic cells, endothelial cells, epithelial cells, adipose cells, retinal cells, keratinocytes, multipotent stem cells, pluripotent stem cells, progenitor cells, embryonic cells, and cells approximating such primary cell types derived via the ex vivo differentiation of multipotent or pluripotent stem cells or progenitor cells
[0140] In another aspect, the disclosure provides a composition, comprising a plurality' of the recombinant cells of any embodiment or combination of embodiments herein. The compositions may be used, for example, in multiplexed cell-based assay technologies. In one embodiment all of the recombinant cells in the composition comprise a first recombinant nucleic acid molecule encoding the same target protein and a second nucleic acid molecule encoding the same accessory protein. In another embodiment, all of the recombinant cells in the composition comprise the same first recombinant nucleic acid molecule, second recombinant nucleic acid molecule, and third recombinant nucleic acid molecule.
[0141] In another embodiment, the plurality of recombinant cells comprises sub-populations of recombinant cells, wherein each sub-population comprises a different first recombinant nucleic acid molecule encoding a different target protein. This embodiment is particularly useful, for example, in methods of the disclosure that measure activation or inhibition of the target proteins of the composition, as described below. In one embodiment, each sub- population comprises a different second recombinant nucleic acid molecule encoding a different accessory- protein. In some embodiments, the second recombinant nucleic acid molecule of all cells in the composition may encode the same accessory protein (i.e., some accessory proteins can bind to more than one target protein).
[0142] In another embodiment, each sub-population comprises (i) a first recombinant nucleic acid molecule that encodes the same protease recognition sequence and cleavage site; and (ii) a second recombinant nucleic acid molecule that encodes the same protease. In a further embodiment, one or more sub-populations comprise (i) a first recombinant nucleic acid molecule that encodes a protease recognition sequence and cleavage site that differs from the protease recognition sequence and cleavage site encoded by a first recombinant nucleic acid molecule in other sub-populations, and (ii) a second recombinant nucleic acid molecule that encodes a protease that differs from the protease encoded by a second recombinant nucleic acid molecule in other sub-populations.
[0143] In one embodiment, each sub-population comprises a first recombinant nucleic acid molecule that encodes the same ESF. In other embodiments, one or more sub-populations comprise a first recombinant nucleic acid molecule that encodes an ESF that differs from the ESF encoded by a first recombinant nucleic acid molecule in other sub-populations.
[0144] In one embodiment, each sub-population comprises a third recombinant nucleic acid molecule that encodes the same RUM. In another embodiment, one or more sub-populations comprise a third recombinant nucleic acid molecule that encodes a RUM that differs from the RUM encoded by a third recombinant nucleic acid molecule in other sub-populations.
[0145] In another embodiment, the plurality of recombinant cells comprises sub-populations of recombinant cells, wherein each sub-population comprises: (a) a different first recombinant nucleic add molecule encoding (i) a different target protein, (ii) an identical protease recognition sequence and cleavage site, and (iii) an identical ESF;
[0146] (b) a different second recombinant nucleic acid molecule encoding (i) a different or identical accessory protein, and (ii) an identical protease;
[0147] (c) a third recombinant nucleic acid molecule encoding a RUM, wherein the RUM comprises, in 5’ to 3’ order:
[0148] (i) the nucleotide sequence encoding the 5’ portion of a reporter protein gene, terminating in an exon of the reporter gene;
[0149] (ii) an identical first intron;
[0150] (iii) a nucleotide sequence comprising an identical test exon;
[0151] (iv) an identical second intron; and
[0152] (v) the nucleotide sequence coding for the remaining 3’ portion of the reporter protein gene, beginning with an exon of the reporter gene.
[0153] In one embodiment, each sub-population comprises a different third recombinant nucleic acid molecule, wherein:
[0154] (a) a first region is a portion of the nucleotide sequence encoding the 5’ portion of the reporter protein gene and a second region is a portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene, wherein one or both of the first region and the second region differ in each sub-population through use of degenerate codons, wherein upon splicing of an mRNA expression product of the third recombinant nucleic acid molecule that excises the test exon the first region and the second region are directly adjacent and form a first nucleic acid barcode; and
[0155] (b) the reporter protein encoded by the combination of the 5’ portion of the reporter protein gene and the nucleotide sequence of the 3’ portion of the reporter protein gene is the same in each sub-population. In one embodiment, the first region is a portion of the 3’ end of the nucleotide sequence encoding the 5 ’ portion of the reporter protein gene, and the second region is a portion of the 5’ end of the nucleotide sequence encoding the 3’ portion of the reporter protein gene. Figure 2 shows an example of a splicing event that leads to formation of the short-form mRNA (test exon and introns spliced out), wherein the first region and the second region are directly adjacent and form the first nucleic acid barcode (labeled as “Barcode B” in Figure 2). In some embodiments, just one of the first region and the second region differ in each sub-population through use of degenerate codons, and the splicing event leads to nucleic acid barcodes that are distinguishable between different sub-populations due to the splicing product (i.e., the short-form mRNA) including (i) a nucleic acid barcode in the (for example) 5’ reporter exon, and (ii) the now adjacent sequence via splicing of a non- barcoded portion of the 3’ exon. In other embodiments, the first region and the second region differ in each sub-population through use of degenerate codons.
[0156] The compositions of this embodiment provide a dual-functioning reporter protein activity / RNA barcode that overcomes technical concerns related to hybridization assays and affords flexibility between assay development and multiplex HTS. Abundances of the RUM (unspliced mRNA), short-form mRNA (spliced mRNA lacking the test exon), and long-form mRNA (spliced mRNA including the test exon) additionally provides a housekeeping signal for quality control of pooled cells that, additionally, streamlines transcript abundance measurement, reducing the likelihood of handling error in assay methods such as qPCR or hybridization. One purpose of this housekeeping signal is to track quality control (e.g., cell viability, cell count) across wells and plates in a manner where the use of traditional housekeeping genes is unable to distinguish between different test cell populations of the same cell background. One utility of this approach is to enable attribution of a lack of response to a test compound to its lack of effect on a test cell population’s exogenous target, rather than to a lack of the cells (or their viability) themselves; thus, this approach reduces false negatives. A second utility of this approach is that it minimizes the number of probes needed for hybridization assays. In the absence of housekeeping barcodes, it could be imagined that branched DNA probes against exogenous target RNA transcripts could be used. This would confirm expression of these targets (or at least successfill transformation and transcription of their sequences) in cells, while serving as a method for assessing cell viability and (roughly) count. However, when tens to hundreds of different targets are assessed in HTS, this necessitates having tens to hundreds of different branched DNA probes, which must be appropriately added to the wells where the relevant exogenous targets are expressed. In contrast, the proposed housekeeping barcodes limit the number of probes needed for housekeeping purposes to the number of test cell populations in a given well (i.e., one probe set per housekeeping signal), and since these cells can be used throughout the plate, whether expressing the same or different exogenous targets, additionally limits the number of probes needed per plate (and by extension the entirety of the HTS campaign), with the same probe sets being used for all wells. This greatly reduces the likelihood of sample handling error (i.e., adding the wrong probes to the wrong well), and overall simplifies the hybridization assay workflow. In embodiments where random sequence nucleotides are used, there is also greater ability to design housekeeping signal-branched DNA probe pairs with similar binding efficiencies across the different cell lines, as opposed to being reliant upon existing exogenous target sequences that are variable in their sequences (e.g., GC content), which may influence probe binding and consequent signal detection sensitivity.
[0157] Table 6 illustrates the sets of degenerate codons that may be used to generate the unique barcodes comprising in part or whole the reporter protein component of the third recombinant nucleic acid molecule of this disclosure.
[0158] Table 6 By way of example in order to demonstrate the utility of this degenerate codon approach, the first ten amino acids of green fluorescent protein are shown in Table 7, with two different nucleotide sequences degenerately encoding this amino acid sequence. Note that, for this specific 10 amino acid sequence alone, with consideration for the number of degenerate codons for each amino acid, Sequences 1 and 2 represent just two of the 36,864 potential barcodes that may be generated while still encoding the same protein.
[0159] Table 7
[0160] When these sequences are compared to one another, they are found to have 53% sequence identity (Table 8). Thus, it can be appreciated how this approach readily facilitates the design of probes specific to degenerate nucleotide sequences encoding the same protein product.
[0161] Table 8
[0162] In another embodiment, a third region of the third recombinant nucleic acid comprises a portion of the test exon, wherein the third region differs in each sub-population through use of degenerate codons, wherein upon splicing of an mRNA expression product of the third recombinant nucleic acid molecule that does not excise the test exon, the third region and either the first region or the second region are directly adjacent and form a second nucleic acid barcode. Figure 2 shows an example of a splicing event that leads to formation of the long-form mRNA (test exon not spliced out), wherein the third region is at the 3’ terminus of the test exon, and splicing of the introns leads the third region and the second region to be directly adjacent to form the second barcode (labeled as “Barcode C” in Figure 2). In one alternative embodiment, the third region is at the 5’ terminus of the test exon, and splicing of the introns leads the third region and the first region to be directly adjacent to form the second nucleic acid barcode.
[0163] In another embodiment, a third region comprises a portion of the test exon, wherein upon splicing of an mRNA expression product of the third recombinant nucleic acid molecule that does not excise the test exon, the third region and either (a) a portion of the nucleotide sequence encoding the 5’ portion of the reporter protein gene, or (b) a portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene, are directly adjacent and form a second nucleic acid barcode, wherein one or both of (i) the third region and (ii) the portion of the nucleotide sequence encoding the 5 ’ portion of the reporter protein gene, or (b) a portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene differ in each sub-population through use of degenerate codons. In some embodiments, just one of the third region and (a) the portion of the nucleotide sequence encoding the 5’ portion of the reporter protein gene, or (b) the portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene differ in each sub-population through use of degenerate codons, and the splicing event leads to nucleic acid barcodes that are distinguishable between different sub-populations due to the splicing product including (i) a nucleic acid barcode in the (for example) third region, and (ii) the now adjacent sequence via splicing of a non- barcoded portion of the 3’ exon. In other embodiments, the third region and (a) the portion of the nucleotide sequence encoding the 5 ’ portion of the reporter protein gene, or (b) the portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene differ in each sub-population through use of degenerate codons. In these embodiments, the portion of the nucleotide sequence encoding the 5’ portion of the reporter protein gene, or the portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene may differ from the first region and the second region that form the nucleic acid barcode for the short- form mRNA, or may comprise the first region and the second region that form the nucleic acid barcode for the short-form mRNA. In another embodiment, a fourth region of the third recombinant nucleic acid comprises
[0164] (a) a portion of the first intron,
[0165] (b) a portion of the second intron,
[0166] (c) a portion of the first intron, and (ii) a portion of the test exon, or the third region,
[0167] (d) or a portion of second intron, and (ii) a portion of the test exon, or the third region,
[0168] (e) a portion of the first intron, and (ii) a portion of the nucleotide sequence encoding the 5’ portion of the reporter protein gene, or the first region,
[0169] (f) a portion of second intron, and (ii) a portion of the nucleotide sequence encoding the 5’ portion of the reporter protein gene, or the first region
[0170] (g) a portion of the first intron, and (ii) a portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene, or the second region; or
[0171] (h) a portion of second intron, and (ii) a portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene, or the second region; wherein the fourth region differs in each sub-population through use of degenerate codons, wherein the fourth region is a third nucleic acid barcode. In some embodiments where the fourth region includes a portion of more than one domain, one or both domain portions may differs in each sub-population through use of degenerate codons. In one embodiment where the fourth region includes a portion of more than one domain, both domain portions may differs in each sub-population through use of degenerate codons
[0172] Figure 2 shows an example of the fourth region that includes a portion of the first intron at the 3’ terminus of the first intron and portion of the third region at the 5’ terminus of the test exon which to form the third nucleic acid barcode (“Barcode A” in Figure 2). The third barcode is present in unspliced mRNA expressed from the third recombinant nucleic acid.
[0173] In one embodiment, the first, second, and third nucleic acid barcodes are at least 6 nucleotides in length. In another embodiment, the first, second, and third nucleic acid barcodes are between 6 and about 2000 nucleotides in length. In various further embodiments the first, second, and third nucleic acid barcodes are between 6-2000 nucleotides in length, or between 6-1500, 6-1000, 6-750, 6-500, 6400, 6-300, 6-200, 6-100, 6-50, or 6-25 nucleotides in length.
[0174] In another embodiment of all of the compositions of the disclosure, in each sub- population, (i) the first recombinant nucleic acid comprises the same promoter; (ii) the second recombinant nucleic acid comprises the same promoter; and (iii) the third recombinant nucleic acid comprises the same promoter.
[0175] In another aspect, the disclosure provides method for screening of a target proteins, comprising:
[0176] (a) contacting the recombinant cell of any embodiment herein with a test compound or test stimulus;
[0177] (b) detecting reporter protein activity and / or abundance of mRNA expressed from the third recombinant nucleic acid molecule.
[0178] In another embodiment, the disclosure provides methods for multiplexed screening of two or more target proteins, comprising:
[0179] (a) contacting the composition of any embodiment herein with a test compound or test stimulus;
[0180] (b) detecting reporter protein activity and / or abundance of mRNA expressed from the third recombinant nucleic acid molecule.
[0181] In a further embodiment, the disclosure provides methods for multiplexed screening of two or more target proteins, comprising:
[0182] (a) contacting a test compound or test stimulus to the composition of any embodiment herein wherein each sub-population comprises a different third recombinant nucleic acid molecule, and the third recombinant nucleic acid molecules form the first or second nucleic acid barcode upon splicing;
[0183] (b) detecting reporter protein activity and / or abundance of mRNA expressed from the third recombinant nucleic acid molecule; and
[0184] (c) measuring an amount of at least the first nucleic acid barcode to separately determine activities of each of the target proteins in the composition.
[0185] As used herein, a test stimulus comprises but is not limited to application or withdrawal of a physical condition, including shear stress, heat, cold, pH changes, vibration, light, electric or magnetic field etc., or application or withdrawal of a contact with a physical structure, including but not limited to a surface, a medium, or another cell, or application or withdrawal of contact with an environmental sample.
[0186] In one embodiment of this aspect of the disclosure, a chemical compound is tested for its ability to activate, inhibit, or otherwise modulate said target protein by contacting said test cell population with the compound or stimulus and measuring the resulting change in activity (for example, fluorescence) of said reporter protein, which is translated in functional form from the normally-spliced mRNA (short-form mRNA) (Figure 2 and Figure 2). In one embodiment, step (c) (the measuring step) comprises measuring an amount of the first nucleic acid barcode, the second nucleic acid barcode, and the third nucleic acid barcode, to separately determine activities of each of the target proteins in the composition. In one embodiment, the mRNA expressed from the third recombinant nucleic acid molecule comprises 1, 2, or all 3 of (i) unspliced mRNA reported on by the third nucleic acid barcode, (ii) spliced mRNA in which the test exon is excised (“short-form mRNA”) reported on by the first nucleic acid barcode, and (iii) spliced mRNA in which the test exon is not excised (“long-form mRNA”) reported on by the second nucleic acid barcode. In a further embodiment, detecting an abundance of mRNA expressed from the third recombinant nucleic acid molecule comprises detecting a relative abundance of the unspliced mRNA, the short- form mRNA, and the long-form mRNA.
[0187] In another embodiment, activation or inhibition of the target protein is measured not by the reporter protein’s activity, but instead by the abundance of the short-form mRNA as measured by methods such as (without limitation) qPCR, RNA-Seq, branched DNA hybridization, or other methods of RNA quantification known to those skilled in the art. In still another, embodiment, the abundances of both short-form mRNA and long-form mRNA are separately measured and expressed as a ratio reflecting the activation or inhibition of the target protein (Figure 2).
[0188] In one embodiment, the measured abundance of the short-form mRNA is divided by (a) the sum of the unspliced RNA’s, the short-form’s, and long-form’s abundances, (b) the long-form’s abundance, or (c) the sum of the unspliced RNA’s and the long-form’s abundances, in order to normalize the readout, wherein the denominator of this ratio serves as a ‘housekeeping’ signal correcting for a variety of otherwise difficult-to-control extraneous variables such as cell number, cell health, transcription rate, mRNA degradation, RNA recovery efficiency, and mRNA half-life.
[0189] In another embodiment, pooled test cell populations are contacted with test compounds, then each pool is analyzed for reporter protein activity (such as fluorescence or bioluminescence), following which only the pools displaying reporter activity above some threshold (which threshold value may be zero) are analyzed for the abundances of each of the first, second, and third nucleic acid barcodes in that pool, in order to separately determine activities of each of the target proteins in the pool. Such selective barcode analysis, based on pooled reporter protein activity', significantly improves the throughput, speed, and cost- effectiveness of barcode-based multiplex screening. In another aspect, multiplexed screening of two or more target proteins may be accomplished using a single test cell population, wherein the test cell population comprises (a) two or more sets of cognate nucleic acid molecules, each set comprising
[0190] (i) one of said first fusion protein-encoding nucleic acid molecules, comprising
[0191] (1) a nucleotide sequence encoding a target protein that is unique to that set, and
[0192] (2) a nucleotide sequence encoding an engineered splicing factor comprising a splicing exclusion sequence, a nuclear localization signal, and a sequence- specific RNA-binding domain that recognizes an RNA sequence unique to that set,
[0193] (ii) one of said second fusion protein-encoding nucleic acid molecules, which may be either unique to that set or shared by all sets, and
[0194] (iii) one of said third fusion protein-encoding nucleic acid molecules, encoding a RUM whose exogenous exon contains the RNA sequence that is recognized by the set’s engineered splicing factor and that is unique to the set. Said RUM also comprises degenerate codon sequences comprising distinct short form and long form barcodes that are unique to the set.
[0195] In this aspect, the ESF fused to each of the two or more target proteins would recognize a distinct binding sequence, and so would bind a distinct RUM comprising a degenerate codon sequence cognate to that ESF and target protein.
[0196] Examples
[0197] Example 1. Test Exon Exdusion by an Engineered Splicing Factor
[0198] The ability of engineered splicing factors (ESFs) to promote exon exclusion from splicing reporters (SRs) was tested. Here, two ESF variants were utilized, both comprised of (beginning at the N-terminus) an optional FLAG™ tag, the glycine-rich domain from heterogeneous nuclear ribonucleoprotein 1A (hnRNPlA), a nuclear localization signal, and Pumilio homology domain 1 (PUM1). The first variant (ESF-WT; SEQ ID NO: 82-83; Table 11) comprised the wildtype PUM1 sequence that binds to the wildtype PUM1 binding sequence (BS) with high affinity, whereas the second variant (ESF-Mut; SEQ ID NO: 84-85) contained four point mutations within the BS (N1043S, Q1047E, S1079N, E1083Q; residue numbers correspond to those in wildtype PUM1) that binds to the mutant PUM1 BS, but not the wildtype BS, with high affinity. Additionally, three SR variants comprised of a 5’ green fluorescent protein GFP2 exon and a 3" GFP2 exon flanking intron 11-12, exon 12 (“test exon”), and intron 12-13 from human insulin-like growth factor 2 mRNA-binding protein 1(IGF2BP1) wrere utilized. The first variant (SR-No-BS) (SEQ ID NO:86; Table 12) comprised the wildtype sequences of all components but lacked a PUM1 BS. The second variant (SR-WT-BS) (SEQ ID NO:87) comprised an 8-nucleotide BS (TGTATATA) substituted within the text exon that could be bound by wildtype PUM1 with high affinity. The third variant (SR-Mut-BS) (SEQ ID
[0199] NO: 88) comprised an 8-nucleotide BS (TTGATATA) substituted within the test exon that could be bound by the mutant PUM1. Nucleotide sequences and encoded protein sequences for ESFs and SRs, and components thereof, used in the example are provided in Tables 9 and
[0200] 10, respectively, with full length sequences in Tables 11-12. Following expression, cells were lysed and their RNA was purified. RT-PCR was performed on the purified RNA, and gel electrophoresis was performed to demonstrate 1) removal of introns in all mRNA products and 2) promotion of test exon exclusion with co-transfection of high affinity partners (e.g.,
[0201] ESF-WT plus SR-WT-BS, ESF-Mut plus SR-Mut-BS) compared to low affinity partners or no ESF co-transfection.
[0202] Table 9. Engineered Splicing Factor Variants Used in Examples
[0203] Table 10. Splicing Reporter Variants Used in Examples EGYVQERT I FFKDDGNYK TRAE (SEQ ID NO: 70 )
[0204] Table 11. Engineered Splicing Factors Used in Examples
[0205] Splicing Reporters Used in Examples
[0206] Cell Culture
[0207] 1. HEK 293T cells were plated in 6-well plates in Dulbecco’s Modified Eagle Medium (DMEM) supplemented with 10% fetal bovine serum (FBS) at a density of 750,000 cells per well and allowed to recover and expand overnight in an incubator (37° C, 5% carbon dioxide) overnight.
[0208] 2. At 80% confluency, cells were transfected with either SR-No-BS, SR-WT-BS, or SR- Mut-BS alone or co-transfected with either ESF-WT or ESF-Mut (e.g., SR-No-BS + ESF-WT, SR-No-BS + ESF-Mut, SR-WT-BS + ESF-WT, etc.), and maintained in an incubator for 36 hours.
[0209] Cell Lysis and RNA Purification
[0210] Cell lysis and purification were performed using a Qiagen RNeasy™ kit according to the manufacturer’s instructions. In brief:
[0211] 1. Cell media were aspirated from each well and replaced with 350 μL Buffer RTL1 supplemented with 20μM dithiothreitol.
[0212] 2. Cells were monitored for lysis using a brightfield microscope, and lysates were harvested into 1.5 mL tubes.
[0213] 3. Lysate tubes were vortexed for 30 seconds, and then briefly spun using a mini centrifuge to remove bubbles.
[0214] 4. Lysates were added to a gDNA eliminator spin column, and the column was centrifuged for 30 seconds at 10,000 rpm remove cellular genomic DNA.
[0215] 5. The spin column was discarded, 350 μL 70% ethanol were added to the flow-through in the collection tube, and the resulting solution was mixing by pipetting.
[0216] 6. The solution was added to a RNeasy spin column, and the column was centrifuged for 15 seconds at 10,000 rpm; the flow-through was discarded.
[0217] 7. 700 μL Buffer RW1 were added to the column, and the column was centrifuged for 15 seconds at 10,000 rpm; the flow-through was discarded, and the step was repeated.
[0218] 8. The column was spun for 1 minute at 10,000 rpm to dry the membrane of any remnant Buffer RW1; the collection tube was discarded and replaced with a 1.5 mL tube. 9. 30 μL of RNAse-free water were added to the spin column, and the column was spun for 1 minute at 10,000 rpm; the column was discarded.
[0219] 10. RNA concentration and purity was measured using a NanoDrop™ Lite spectrophotometer (Thermo Fisher Scientific).
[0220] Reverse transcription-polymerase chain reaction (RT-PCR)
[0221] RT-PCR was performed using the SuperScript™ IV One-Step PCR System (Invitrogen) and oligonucleotide primer pairs targeting the SR variants. Each reaction was 50 μL, consisting of: 25 μL 2X Platinum™ SuperFi RT-PCR Master Mix, 18.5 μL nuclease-free water, 2.5 μL forward primer (10μM), 2.5 μL reverse primer (10 μM), 1 μL of RNA (1 μg), and 0.5 μL SuperScript™ IV RT Mix.
[0222] The thermocycler reaction steps and conditions were as follows: reverse transcription (45° C for 10 minutes), reverse transcriptase inactivate (98° C for 2 minutes), deoxyribonucleic acid (DNA) denaturation (98° C for 10 seconds), DNA primer annealing (55° C for 10 seconds), polymerase extension (72° C for 30 seconds) (the denaturation through extension steps were repeated for 35 cycles), and final extension (72° C for 1 minute).
[0223] Gel Electrophoresis
[0224] 1. A 3% agarose gel containing ethidium bromide was prepared.
[0225] 2. Gel loading dye was added to RT-PCR products.
[0226] 3. A 50 bp ladder and dyed RT-PCR products were loaded into the gel, and the gel was run until clear resolution between the ladder bands was obtained.
[0227] 4. The gel was inspected with an ultraviolet light transilluminator, and bands were cut from the gel for subsequent gel extraction and PCR clean-up.
[0228] DNA Gel Extraction / PCR Clean-up and Linear DNA Sequencing
[0229] PCR clean-up was performed using the NucleoSpin™ Gel and PCR Clean-up kit (Macherey-Nagel) according to the manufacturer’s instructions. In brief:
[0230] 1. Cut out gels containing DNA bands were weighed and placed in tubes to which an appropriate volume of NTI buffer was added (200 μL buffer per 100 g gel). The tubes were heated at 50° C and briefly vortexed intermittently until the gel was completely dissolved. 2. The dissolved gel / NTI buffer solution was added to NucleoSpin™ Gel and PCR Clean-up columns, and the column was centrifuged for 30 seconds at 11,000 x g; the flow-through was discarded.
[0231] 3. 700 μL of NT3 buffer was added to the column, and the column was centrifuged for 30 seconds at 11,000 x g; the-flow through was discarded, and the step was
[0232] 4. The column was spun for 1 minute at 11 ,000 x g to dry' the membrane of any remnant Buffer RW1; the collection tube was discarded and replaced with a 1.5 mLtube.
[0233] 5. 20 μL of NE buffer were added to the spin column, and the column was spun for 1 minute at 11,000 x g; the column was discarded.
[0234] 6. DNA concentration and purity was measured using a Qubit™ 4 fluorophotometer (Invitrogen) and prepared appropriately for linear DNA sequencing (performed by Plasmidsaurus).
[0235] Results
[0236] Figure 4 demonstrates that a SR-No-BS exhibits a basal level of exon exclusion (i.e., presence of both 795 bp and 720 bp bands) that is not further influenced by the co-expression of either ESF-WT or ESF-Mut. SR-WT-BS and SR-Mut-BS similarly exhibited a basal level of exon exclusion. Additionally, co-expression of SR-WT-BS with ESF-WT or SR-Mut-BS with ESF-Mut (but not vice versa) promoted exon exclusion by increasing the ratio of the 720 bp:795 bp band ratio. Notably, the SR-WT-BS / ESF-WT pairing more strongly promoted exon exclusion than the SR-Mut-BS / ESF-Mut pairing in this experiment.
[0237] Example 2. Singleplexed Assay
[0238] An assay is performed individually on two cell sets to demonstrate singpleplexing capabilities of ESF-Mut in a receptor activity-dependent manner. One set of cells co- expresses the alpha 2C adrenergic receptor (ADRA2C) fused to ESF-Mut via a fragment of the vasopressin receptor 2 (V2) C-terminus and a tobacco etch virus (TEV) protease cleavage sequence (ADRA2C-ESF), beta-arrestin2 fused to a TEV protease (Barr2-TEV), and SR- Mut-BS. The second set of cells co-expresses the melatonin receptor type 1A (MTNR1A) fused to ESF-Mut via a fragment of the V2 receptor C-terminus and TEV cleavage sequence (MTNR1A-ESF), Barr2-TEV and SR-Mut-BS. Sequences for receptor-ESFs and Barr2-TEV, and components thereof, used in the example are provided in Tables 13 and 14, respectively. Wells containing the individual cell sets are treated either with dimethylsulfoxide (DMSO) as a vehicle control, 10μM clonidine (an ADRA2C agonist), or 10μM melatonin (an MTNR1A agonist). Following stimulation, cells are lysed and their RNA are purified. RT-PCR is performed on the purified RNA, and gel electrophoresis is performed to demonstrate that SRs in cells expressing receptor responsive to ligand promote test exon exclusion compared to
[0239] DMSO or unproductive ligand pairing (e.g., ADRA2C treatment with melatonin) conditions.
[0240] Table 13. Receptor-ESF Fusions
[0241] Table 14. Barr2-TEV Fusion
[0242] Cell Culture
[0243] 3. HEK 293T cells are plated in 6-well plates in Dulbecco’s Modified Eagle Medium
[0244] (DMEM) supplemented with 10% fetal bovine serum (FBS) at a density of 750,000 cells per well and allowed to recover and expand overnight in an incubator (37° C,
[0245] 5% carbon dioxide) overnight.
[0246] 4. At 80% confluency, cells were co-transfected with either ADRA2C-ESF, Barr2-TEV, and SR or MTNR1A, Barr2-TEV, and SR, and maintained in an incubator overnight.
[0247] Compound Treatment
[0248] 1. After 24 hours, wells are treated with either 10 μM clonidine, 10μM melatonin, or
[0249] DMSO, and maintained in an incubator overnight.
[0250] Cell Lysis and RNA Purification
[0251] RT-PCR is performed using oligonucleotide primer pairs targeting the SR-Mut-BS.
[0252] The remaining reaction steps and conditioned are performed the same as described in
[0253] Example 1.
[0254] RT-PCR
[0255] RT-PCR is performed as described in Example 1,
[0256] Gel Electrophoresis
[0257] Gel electrophoresis is performed as described in Example 1.
[0258] Results Figure 5 demonstrates the results of treating either cell line with the agonist for to its co-transfected ESF-tagged receptor, as revealed by gel electrophoresis of the RT-PCR- amplified cell products. For cells expressing ADRA2C-ESF, test exon exclusion is promoted above basal levels only when cells are treated with clonidine (Lane 4), while for cells expressing MTNR1A-ESF, test exon exclusion promotion above basal is observed only when cells are treated with melatonin (Lane 6).
[0259] Example 3. Multiplexed Assay
[0260] An assay is performed on a pool of two cell sets to demonstrate multiplexing capabilities. The same transfection conditions and plasmids are used as in Example 2, except that the SRs in the two cell sets differed by substitution of degenerate N-terminal and C- terminal tag sequences (i.e., degenerate barcodes) recognized by unique sets of RT-PCR oligonucleotide primers. Additionally, following initial transfection and expression, but prior to stimulation, the two separate cell sets are pooled into a single well. Following stimulation, cells are lysed and their RNA are purified. RT-PCR is performed on the purified RNA, and gel electrophoresis is performed to demonstrate that SRs in cells expressing receptor responsive to ligand promote test exon exclusion compared to DMSO or unproductive ligand pairing (e.g., ADRA2C treatment with melatonin) conditions.
[0261] Cell Culture
[0262] 1. HEK 293T cells are plated in 6-well plates in Dulbecco’s Modified Eagle Medium (DMEM) supplemented with 10% fetal bovine serum (FBS) at a density of 750,000 cells per well and allowed to recover and expand overnight in an incubator (37° C, 5% carbon dioxide) overnight.
[0263] 2. At 80% confluency, cells are co-transfected with either ADRA2C-ESF, Barr2-TEV, and SR or MTNR1A, Barr2-TEV, and SR, and maintained in an incubator overnight.
[0264] Compound Treatment
[0265] Compound treatments are performed as described in Example 2.
[0266] Cell Lysis and RNA Purification
[0267] Cell lysis and RNA purification are performed as described in Example 1.
[0268] RT-PCR RT-PCR is performed using oligonucleotide primer pairs specific for the degenerate SR-Mut-BS. The remaining reaction steps and conditioned are performed the same as described in Example 1.
[0269] Gel Electrophoresis
[0270] Gel electrophoresis is performed as described in Example 1.
[0271] Results
[0272] Figure 6 demonstrates the results of treating the two pooled cell lines co-transfected with ADRA2C-ESF or MTNR1 A-ESF, Barr2-TEV, plus their corresponding degenerately barcoded SRs, as revealed by gel electrophoresis of the RT-PCR-amplified pooled cell products. For pooled cells treated with vehicle (DMSO; lanes 2 and 3) the amplified mRNA reveals basal levels of test exon exclusion for both the ADRA2C- and MTNRIA-associated degenerate SRs. In clonidine-treated pooled cells, RT-PCR product amplified with primers specific for ADRA2C condition SR (lane 4) reveals an increase in the 720 bp:795 bp bands ratio, indicating promotion of exclusion of the test exon in the presence of receptor agonist. Similarly, PCR product amplified with primers specific for MTNR1 A condition SR (lane 7) likewise reveals an increase in the 720 bp:795 bp bands ratio. In contrast, when the pooled cells are treated with ADRA2C agonist but amplified with MTNR1A condition-specific primers (lane 5), or with MTNR1A agonist but amplified with ADRA2C condition-specific primers (lane 6) only basal levels of test exon exclusion are observed, indicating the absence of cross-talk between the two pooled cell lines in this multiplexed configuration.
[0273] Figure 6 demonstrates the results of treating either cell line with the agonist for to its co-transfected ESF-tagged receptor, as revealed by gel electrophoresis of the RT-PCR- amplified cell products. For cells expressing ADRA2C-ESF, test exon exclusion is promoted above basal only when cells are treated with clonidine (Lane 4), while for cells expressing MTNR1A-ESF, test exon exclusion promotion above basal is observed only when cells are treated with melatonin (Lane 6).
[0274] Table 15. Receptor-ESF Fusions
Claims
We claim1. A recombinant cell comprising:(a) a first recombinant nucleic acid molecule comprising a nucleotide sequence encoding a first fusion protein operatively linked to a promoter, wherein the first fusion protein comprises:(i) a target protein;(ii) a protease recognition sequence and cleavage site;(iii) an engineered splicing factor (ESF), comprising:(1) a splicing exclusion sequence;(2) a nuclear localization signal; and(3) a sequence-specific RNA-binding domain, wherein the sequence- specific RNA-binding domain is capable of binding specifically to a nucleotide sequence within a test exon to be excluded; wherein the protease recognition sequence and cleavage site is located between the target protein and the ESF;(b) a second recombinant nucleic acid molecule, comprising a nucleotide sequence encoding a second fusion protein operatively linked to a promoter, wherein the second fusion protein comprises:(i) an accessory protein, wherein the accessory protein is capable of binding to the target protein as a function of the target protein's activity state; and(ii) a protease capable of specifically recognizing the protease recognition sequence and cleaving the cleavage site; and(c) a third recombinant nucleic acid molecule, comprising a nucleotide sequence encoding a reporter unspliced mRNA (RUM) operatively linked to a promoter, wherein the RUM comprises in 5’ to 3’ order(i) a nucleotide sequence encoding a 5 ’ portion of a reporter protein gene, terminating in an exon of the reporter gene;(ii) a first intron;(iii) a nucleotide sequence comprising a test exon;(iv) a second intron;(v) a nucleotide sequence coding for the remaining 3’ portion of the reporter protein gene, beginning with an exon of the reporter gene;wherein expression of the test exon in the reporter protein disrupts function of the reporter protein, wherein the test exon comprises the nucleotide sequence bound specifically by the sequence-specific RNA binding domain.
2. The recombinant cell of claim 1, wherein the first fusion protein comprises, in amino terminal to carboxy terminal order, the target protein, the protease recognition sequence and cleavage site, and the ESF.
3. The recombinant cell of any one of claims 1-2, wherein the protease recognition sequence and cleavage site is not recognized by proteases endogenously present in the cell, and wherein the protease is not endogenously expressed in the cell.
4. The recombinant cell of any one of claims 1-3, wherein the second fusion protein comprises, in amino terminal to carboxy terminal order, the accessory' protein and the protease.
5. The recombinant cell of any one of claims 1-4, wherein the test exon is not present in a native sequence of the reporter protein gene.
6. The recombinant cell of any one of claims 1-5, wherein the target protein is a plasma membrane-associated protein, including but not limited to a transmembrane protein.
7. The recombinant cell of any one of claims 1-6, wherein the target protein is a cell surface receptor protein.
8. The recombinant cell of any one of claims 1-7, wherein the target protein is selected from the group consisting of G protein-coupled receptors, nuclear hormone receptors, receptor tyrosine kinases, receptor serine-threonine kinases, cytosolic protein kinases, transcription factors, ATP binding cassette family member proteins, adenylate cyclases, ATP- dependent membrane transporters, B cell leukemia / lymphoma 2 (BCL2), BCL2-related proteins, BCL2-like proteins, basic helix-loop-helix family proteins, caveolins, CCAAT / enhancer binding proteins, centromere proteins, diacylglycerol kinases, Fc receptor peptides, frizzled class receptors, histones, Jun dimerization proteins, mitogen activated protein kinases, neuronal PAS domain proteins, phosphoinositide-3-kinase subunits, DNApolymerases, Rai GTPase activating protein subunits, solute carriers, synaptotagmins, TATA- box binding protein associated factors, teneurins, tropomyosins, tubulins, proteases, RNA splicing factors, cell adhesion proteins, cytoskeletal proteins, ionotropic receptors, voltage- gated ion channels, membrane transporters, enzymes of intermediary metabolism, mitochondrial proteins, nuclear proteins, endoplasmic reticulum proteins, lysosomal proteins, endocytotic vesicular proteins, exocytotic vesicular proteins, Golgi proteins, synaptic proteins, dendritic proteins, cytosolic proteins, antiporters, symporters, phosphoprotein binding proteins, phospholipases, and sortilin.
9. The recombinant cell of any one of claims 1-8, wherein the protease recognition sequence and cleavage site is bound and cleaved by tobacco etch virus protease, and wherein the protease comprises tobacco etch virus protease.
10. The recombinant cell of claim 9, wherein the nucleotide sequence encoding the protease recognition and cleavage site encodes the amino acid sequence selected from the group consisting of SEQ ID NO:5-7, or variants thereof.
11. The recombinant cell of claim 9 or 10, wherein the nucleotide sequence encoding the protease recognition and cleavage site comprises the nucleotide sequence of SEQ ID NO:3 or 4, or variants thereof.
12. The recombinant cell of any one of claims 9-11, wherein the nucleotide sequence encoding the protease encodes the amino acid sequence of SEQ ID NO:2, 81, or variants thereof.
13. The recombinant cell of any one of claims 9-12, wherein the nucleotide sequence encoding the protease comprises the nucleotide sequence of SEQ ID NO: 1, 62, or variants thereof.
14. The recombinant cell of any one of claims 9-13, wherein the cell is a U2OS cell, an human embryonic kidney (HEK)293 cell, or a Chinese hamster ovary (CHO) cells.
15. The recombinant cell of any one of claims 1-14, wherein the nucleotide sequence encoding the splicing exclusion sequence encodes the amino acid sequence selected from the group consisting of SEQ ID NO:9, 18, 20, and 2, or variants thereof.
16. The recombinant cell of any one of claims 1-15, wherein the nucleotide sequence encoding the splicing exclusion sequence comprises the nucleotide sequence selected from the group consisting of SEQ ID NO:8, 17, 19, and 21, or variants thereof.
17. The recombinant cell of any one of claims 1-16, wherein the nucleotide sequence encoding the nuclear localization signal encodes the amino acid sequence selected from the group consisting of SEQ ID NO: 11, 24, 26, and 28, or variants thereof.
18. The recombinant cell of any one of claims 1-17, wherein the nucleotide sequence encoding the nuclear localization signal comprises the nucleotide sequence selected from the group consisting of SEQ ID NO: 10, 25, 23, and 27, or variants thereof.
19. The recombinant cell of any one of claims 1-18, wherein the nucleotide sequence encoding the sequence-specific RNA-binding domain encodes the amino acid sequence selected from the group consisting of SEQ ID NO: 13, 69, or variants thereof.
20. The recombinant cell of any one of claims 1-19, wherein the nucleotide sequence encoding the sequence-specific RNA-binding domain comprises the nucleotide sequence selected from the group consisting of SEQ ID NO: 12, 50 or variants thereof.
21. The recombinant cell of any one of claims 1 -20, wherein the nucleotide sequence bound specifically by the sequence-specific RNA-binding domain comprises the nucleotide sequence of SEQ ID NO:29, or variants thereof.
22. The recombinant cell of any one of claims 1-21, wherein the second fusion protein comprises, in amino terminal to carboxy terminal order, (i) the accessory protein, and (ii) the protease.
23. The recombinant cell of any one of claims 1-21, wherein the second fusion protein comprises, in amino terminal to carboxy terminal order, (i) the protease, and (ii) the accessory protein.
24. The recombinant cell of any one of claims 1-23, wherein the nucleotide sequence encoding the test exon comprises the nucleotide sequence selected from the group consisting of SEQ ID NO: 16 substituted or inserted with a sequence (including but not limited to TGTATATA or TGTATATAX where X is any nucleotide) to which the sequence-specific RNA-binding domain binds, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:
51. SEQ ID NO:53, SEQ ID NO:55, or variants thereof.
25. The recombinant cell of any one of claims 1-24, wherein the 5’ portion of the reporter protein gene and the 3’ portion of the reporter protein gene, combined, encode a reporter protein.
26. The recombinant cell of any one of claims 1-25, wherein the nucleotide sequence encoding the introns are selected from the group consisting of SEQ ID NO: 14-15, or variants thereof.
27. The recombinant cell of any one of claims 1-26, wherein one, two, or all three of the first recombinant nucleic acid molecule, the second recombinant nucleic acid molecule, and the third recombinant nucleic acid molecule are stably integrated in the recombinant cell genome.
28. The recombinant cell of any one of claims 1 -26, wherein the first recombinant nucleic acid molecule, the second recombinant nucleic acid molecule, and the third recombinant nucleic acid molecule are comprised in one or more expression vectors.
29. The recombinant cell of any one of claims 1-26, wherein the first recombinant nucleic acid molecule, the second recombinant nucleic acid molecule, and the third recombinant nucleic acid molecule are comprised in a single expression vector.
30. The recombinant cell of any one of claims 1-26, wherein the first recombinant nucleic acid molecule, the second recombinant nucleic acid molecule, and the third recombinant nucleic acid molecule are comprised in two or three expression vectors.
31. The recombinant cell of any one of claims 1-26, wherein the first recombinant nucleic acid molecule, the second recombinant nucleic acid molecule, and the third recombinant nucleic acid molecule are each comprised in a separate expression vector.
32. The recombinant cell of any one of claims 1-31, wherein the recombinant cells is a mammalian cell.
33. The recombinant cell of claim 32, wherein the mammalian cell is selected from the group consisting of U-2 OS cells, HEK-293 cells, A549 cells, A431 cells, CHO cells, cancer cell lines, donor-derived or cultured neoplastic cells, primary cells of human or animal origin including without limitation hepatocytes, renal cells, fibroblasts, neurons, glial cells, myocytes, monocytes, lymphocytes, dendritic cells, endothelial cells, epithelial cells, adipose cells, retinal cells, keratinocytes, multipotent stem cells, pluripotent stem cells, progenitor cells, embryonic cells, and cells approximating such primary cell types derived via the ex vivo differentiation of multipotent or pluripotent stem cells or progenitor cells.
34. A composition, comprising a plurality of the recombinant cells of any one of claims 1-33.
35. The composition of claim 34, wherein all of the recombinant cells comprise a first recombinant nucleic acid molecule encoding the same target protein and a second nucleic acid molecule encoding the same accessory protein.
36. The composition of claim 34, wherein all of the recombinant cells comprise the same first recombinant nucleic acid molecule, second recombinant nucleic acid molecule, and third recombinant nucleic acid molecule.
37. The composition of claim 34, wherein the plurality of recombinant cells comprises sub-populations of recombinant cells, wherein each sub-population comprises a different first recombinant nucleic acid molecule encoding a different target protein.
38. The composition of claim 37, wherein each sub-population comprises a different second recombinant nucleic acid molecule encoding a different accessory protein.
39. The composition of any one of claims 37-38, wherein each sub-population comprises (i) a first recombinant nucleic acid molecule that encodes the same protease recognition sequence and cleavage site; and (ii) a second recombinant nucleic acid molecule that encodes the same protease.
40. The composition of any one of claims 37-38, wherein one or more sub-populations comprise (i) a first recombinant nucleic acid molecule that encodes a protease recognition sequence and cleavage site that differs from the protease recognition sequence and cleavage site encoded by a first recombinant nucleic acid molecule in other sub-populations, and (ii) a second recombinant nucleic acid molecule that encodes a protease that differs from the protease encoded by a second recombinant nucleic acid molecule in other sub-populations.
41. The composition of any one of claims 37-40, wherein each sub-population comprises a first recombinant nucleic acid molecule that encodes the same ESF.
42. The composition of any one of claims 37-40, wherein one or more sub-populations comprise a first recombinant nucleic acid molecule that encodes an ESF that differs from the ESF encoded by a first recombinant nucleic acid molecule in other sub-populations.
43. The composition of any one of claims 37-42, wherein each sub-population comprises a third recombinant nucleic acid molecule that encodes the same RUM.
44. The composition of any one of claims 37-42, wherein one or more sub-populations comprise a third recombinant nucleic acid molecule that encodes a RUM that differs from the RUM encoded by a third recombinant nucleic acid molecule in other sub-populations.
45. The composition of claim 44 wherein the plurality of recombinant cells comprises sub-populations of recombinant cells, wherein each sub-population comprises:(a) a different first recombinant nucleic add molecule encoding (i) a different target protein, (ii) an identical protease recognition sequence and cleavage site, and (iii) an identical ESF;(b) a different second recombinant nucleic acid molecule encoding (i) a different or identical accessory protein, and (ii) an identical protease;(c) a third recombinant nucleic acid molecule encoding a RUM, wherein the RUM comprises, in 5’ to 3’ order:(i) the nucleotide sequence encoding the 5’ portion of a reporter protein gene, terminating in an exon of the reporter gene;(ii) an identical first intron;(iii) a nucleotide sequence comprising an identical test exon;(iv) an identical second intron; and(v) the nucleotide sequence coding for the remaining 3’ portion of the reporter protein gene, beginning with an exon of the reporter gene.
46. The composition of any one of claims 37-45, wherein each sub-population comprises a different third recombinant nucleic acid molecule, wherein:(a) a first region is a portion of the nucleotide sequence encoding the 5’ portion of the reporter protein gene and a second region is a portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene, wherein one or both of the first region and the second region differ in each sub-population through use of degenerate codons, wherein upon splicing of an mRNA expression product of the third recombinant nucleic acid molecule that excises the test exon the first region and the second region are directly adjacent and forma first nucleic acid barcode; and(b) the reporter protein encoded by the combination of the 5 ’ portion of the reporter protein gene and the nucleotide sequence of the 3’ portion of the reporter protein gene is the same in each sub-population.
47. The composition of claim 46, wherein both of the first region and the second region differ in each sub-population through use of degenerate codons.
48. The composition of claim 46 or 47, wherein a third region comprises a portion of the test exon, wherein upon splicing of an mRNA expression product of the third recombinant nucleic acid molecule that does not excise the test exon, the third region and either (a) aportion of the nucleotide sequence encoding the 5’ portion of the reporter protein gene, or (b) a portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene, are directly adjacent and form a second nucleic acid barcode, wherein one or both of (i) the third region and (ii) the portion of the nucleotide sequence encoding the 5’ portion of the reporter protein gene, or (b) a portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene differ in each sub-population through use of degenerate codons.
49. The composition of claim 48, wherein both of the (i) the third region and (ii) the portion of the nucleotide sequence encoding the 5 ’ portion of the reporter protein gene, or (b) a portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene differ in each sub-population through use of degenerate codons.
50. The composition of claim 48 or 49, wherein (a) the portion of the nucleotide sequence encoding the 5’ portion of the reporter protein gene comprises the first region, or (b) the portion of the nucleotide sequence encoding the 3’ portion of the reporter protein comprises the second region.
51. The composition of claim 46 or 47, wherein a third region comprises a portion of the test exon, wherein the third region differs in each sub-population through use of degenerate codons, wherein upon splicing of an mRNA expression product of the third recombinant nucleic acid molecule that does not excise the test exon, the third region and either the first region or the second region are directly adjacent and form a second nucleic acid barcode.
52. The composition of any one of claims 46-51, wherein a fourth region comprises:(a) a portion of the first intron,(b) a portion of the second intron,(c) a portion of the first intron, and (ii) a portion of the test exon, or the third region,(d) or a portion of second intron, and (ii) a portion of the test exon, or the third region,(e) a portion of the first intron, and (ii) a portion of the nucleotide sequence encoding the 5 ’ portion of the reporter protein gene, or the first region,(f) a portion of second intron, and (ii) a portion of the nucleotide sequence encoding the 5’ portion of the reporter protein gene, or the first region(g) a portion of the first intron, and (ii) a portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene, or the second region; or(h) a portion of second intron, and (ii) a portion of the nucleotide sequence encoding the 3’ portion of the reporter protein gene, or the second region; wherein the fourth region differs in each sub-population through use of degenerate codons, wherein the fourth region is a third nucleic acid barcode. In some embodiments where the fourth region includes a portion of more than one domain, one or both domain portions may differs in each sub-population through use of degenerate codons. In one embodiment where the fourth region includes a portion of more than one domain, both domain portions may differs in each sub-population through use of degenerate codons53. The composition of any one of claims 46-51, wherein a fourth region comprises(a) a portion of the first intron,(b) a portion of the second intron,(c) (i) a portion of the first intron or a portion of second intron, and (ii) the third region,(d) (i) a portion of the first intron or a portion of second intron, and (ii) the first region, or(e) (i) a portion of the first intron or a portion of second intron, and (ii) the second region, wherein the fourth region differs in each sub-population through use of degenerate codons, wherein the fourth region is a third nucleic acid barcode.
54. The composition of any one of claims 46-53, wherein each nucleic acid barcode is at least 6 nucleotides in length.
55. The composition of any one of claims 46-53, wherein each nucleic acid barcode is between 6 nucleotides and 2000 nucleotides in length.
56. The composition of any one of claims 37-55, wherein in each sub-population, (i) the first recombinant nucleic acid comprises the same promoter; (ii) the second recombinant nucleic acid comprises the same promoter; and (iii) the third recombinant nucleic acid comprises the same promoter.
57. A method for screening of a target proteins, comprising:(a) contacting the recombinant cell of any one of claims 1-33 with a test compound or test stimulus;(b) detecting reporter protein activity and / or abundance of mRNA expressed from the third recombinant nucleic acid molecule.
58. A method for multiplexed screening of two or more target proteins, comprising:(a) contacting the composition of any one of claims 34-55 with a test compound or test stimulus;(b) detecting reporter protein activity and / or abundance of mRNA expressed from the third recombinant nucleic acid molecule.
59. A method for multiplexed screening of two or more target proteins, comprising:(a) contacting the composition of any one of claims 46-55 with a test compound or test stimulus;(b) detecting reporter protein activity and / or abundance of mRNA expressed from the third recombinant nucleic acid molecule; and(c) measuring an amount of at least the first nucleic acid barcode to separately determine activities of each of the target proteins in the composition.
60. The method of claim 59, wherein step (c) comprises measuring an amount the first nucleic acid barcode, the second nucleic acid barcode, and the third nucleic acid barcode, to separately determine activities of each of the target proteins in the composition.
61. The method of claim 59 or 60, wherein the mRNA expressed from the third recombinant nucleic acid molecule comprises 1, 2, or all 3 of (i) unspliced mRNA reported on by the third nucleic acid barcode, (ii) spliced mRNA in which the test exon is excised (“short-form mRNA”) reported on by the first nucleic acid barcode, and (iii) spliced mRNA in which the test exon is not excised (‘long-form mRNA”) reported on by the second nucleic acid barcode.
62. The method of any one of claims 57-61, wherein detecting an abundance of mRNA expressed from the third recombinant nucleic acid molecule comprises detecting a relative abundance of the unspliced mRNA, the short-form mRNA, and the long-form mRNA.
63. The method of claim 62, wherein a sum of the abundance of the unspliced mRNA, the short-form mRNA, and the long-form mRNA is determined to normalize individual abundance detected for the unspliced mRNA, the short-form mRNA, and the long-form mRNA, and / or is used as a housekeeping signal to provide a relative measure of cell health, cell number, or reporter gene transcription activity and / or mRNA turnover.
Citation Information
Patent Citations
Methods and compositions for synthetic RNA endonucleases
US20190153429A1
Binding domains directed against GPCR:g protein complexes and uses derived thereof
US20200239534A1