Methods and systems for identifying functional sites
A high-resolution method using a variant library and index sequences in protein expression systems identifies functional sites with single amino acid precision, addressing the limitations of existing mapping techniques and improving drug development.
Patent Information
- Application Number
- PCT/US2025/023634
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-31
- Filing Date
- 2025-04-08
- Publication Date
- 2025-10-16
AI Technical Summary
Existing methods struggle to accurately map functional sites of proteins with high resolution, particularly in identifying alterations that affect biological functions such as signal transduction, protein expression, and folding, which are crucial for understanding protein behavior and drug development.
A method involving a variant library of nucleic acids encoding protein sequences with alterations to at least 75% of amino acid residues, expressed in cells and subjected to biological assays with unique index sequences for high-resolution mapping of functional sites, using systems that incorporate mixed effect negative binomial generalized linear models to analyze barcode data.
Enables precise identification of functional sites in proteins with resolutions down to single amino acids, enhancing our understanding of protein behavior and facilitating drug development by quantifying subtle differences in protein variants.
Smart Images

Figure US2025023634_16102025_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEMS FOR IDENTIFYING FUNCTIONAL SITESCROSS-REFERENCE
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 631,944, filed on April 9, 2024, U.S. Provisional Application No. 63 / 654,529, filed on May 31, 2024, and U.S. Provisional Application No. 63 / 677,984, filed on July 31, 2024, which are incorporated herein by reference in their entireties.SUMMARY
[0002] In certain aspects, disclosed herein is a method of high resolution mapping a functional site of a protein that influences a biological function of the protein, the method comprising: providing a variant library, wherein the variant library comprises a plurality of nucleic acids, wherein individual members of the plurality of nucleic acids encode the protein or a variant of the protein, wherein the plurality of nucleic acids encode a plurality of protein sequences comprising a plurality of alterations to at least 75% of the amino acid residues of the protein across the plurality of the protein sequences; expressing the variant library in a plurality of cells; and performing a biological assay on the plurality of cells, wherein the biological assay comprises detection of an index sequence unique to the protein or the variant of the protein. In some embodiments, the biological assay comprises a measure of protein abundance. In some embodiments, the plurality of cells is comprised within a partition. In some embodiments, the partition comprises a tissue culture flask, plate, or dish. In some embodiments, the partition comprises a well of a well-plate. In some embodiments, the well-plate comprises a 6-well, 12- well, 24-well, 48-well, 96-well or 384-well plate. In some embodiments, the plurality of cells comprises mammalian cells. In some embodiments, the plurality of cells comprises human cells. In some embodiments, each cell of the plurality of cells comprises only one of the plurality of nucleic acids. In some embodiments, the protein comprises at least 100 amino acids. In some embodiments, the protein is a human protein. In some embodiments, the protein comprises a GPCR, a receptor tyrosine kinase, a nuclear hormone receptor, an ion channel, a transcription factor, an integrin. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 80% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 90% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 95% of the amino acid residues of the protein across the plurality of proteinsequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 98% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 99% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to 100% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of alterations to the amino acid residues of the protein comprises amino acid substitutions to at least 10 different amino acids. In some embodiments, the plurality of alterations to the amino acid residues of the protein comprises amino acid substitutions to at least 15 different amino acids. In some embodiments, the plurality of alterations to the amino acid residues of the protein comprises amino acid substitutions to at least 19 different amino acids. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to 80% of the codons encoding the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to 90% of the codons encoding the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to 95% of the codons encoding the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to 98% of the codons encoding the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to 99% of the codons encoding the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to 100% of the codons encoding the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the protein or the variant of the protein is represented by at least 10 different unique index sequences unique to the protein or the variant of the protein. In some embodiments, the protein or the variant of the protein is represented by at least 20 different unique index sequences unique to the protein or the variant of the protein. In some embodiments, the protein or the variant of the protein is represented by at least 30 different unique index sequences unique to the protein or the variant of the protein. In some embodiments, the biological assay comprises activation of a reporter gene, wherein the reporter gene comprises an index sequence unique to the protein or the variant of the protein. In some embodiments, the reporter gene is operatively coupled to apromoter. In some embodiments, the functional site comprises an allosteric regulatory site. In some embodiments, the functional site is a site controlling differential signaling by the protein. In some embodiments, the method further comprises identifying the functional site to a resolution of five amino acids. In some embodiments, the method further comprises identifying the functional site to single amino acid resolution. In some embodiments, the biological function comprises signal transduction of the protein, a protein expression level, protein folding, intracellular trafficking, or cell surface expression. In some embodiments, individual members of the plurality of nucleic acids of the variant library encode the protein and a variant of the protein. In some embodiments, the method further comprises contacting the plurality of cells with a test agent. In some embodiments, contacting the plurality of cells with a test agent occurs before performing the biological assay on the plurality of cells. In some embodiments, the biological assay that comprises a measure of protein abundance measures steady-state abundance, cell surface expression, protein trafficking, or folding of the protein. In some embodiments, the method further comprises performing a second biological assay. In some embodiments, the second biological assay comprises a measure of signaling by the protein.BRIEF DESCRIPTION OF THE DRAWINGS
[0003] FIG. 1 depicts an example of a system used to identify functional sites at high resolution.
[0004] FIG. 2 illustrates a schematic of healthy and disease states of autosomal dominant retinitis pigmentosa (adRP).
[0005] FIG. 3 illustrates a schematic of the in vitro diagnostic assay for the quantification of RHO trafficking.
[0006] FIG. 4 shows a heatmap of the amount of mutated RHO trafficked to the extracellular domain relative to wildtype RHO.
[0007] FIG. 5 shows a heatmap of the amount of mutated RHO trafficked to the extracellular domain when rescued by a rhodopsin corrector molecule relative to wildtype RHO.
[0008] FIGs. 6A-6B shows a graph of the amount of select mutated RHO categorized by class trafficked to the extracellular domain relative to wildtype RHO.
[0009] FIGs. 7A-7B shows a graph of the amount of select mutated RHO categorized by class trafficked to the extracellular domain when rescued by a rhodopsin corrector molecule relative to wildtype RHO.
[0010] FIG. 8 depicts the effects of amino acid substitution at different locations on interferon-alpha signaling and protein stability.
[0011] FIG. 9 depicts the relationship between protein expression and interferon alpha signaling of different TYK2 variants.
[0012] FIG. 10 depicts significant variants detected at the kinase active site and a known allosteric regulatory site.
[0013] FIG. 11 depicts the location of variants with significant loss of signaling without loss of expression on their location across the TYK2 structure.
[0014] FIG. 12 depicts an experimental schematic.
[0015] FIG. 13 depicts TYK2 variants that contribute to drug resistance.
[0016] FIG. 14 depicts an experimental schematic.
[0017] FIG. 15 depicts TYK2 variants that increase resistance to Drug 1, Drug 2, or both Drug 1 and Drug 2.
[0018] FIG. 16 depicts TYK2 variants that increase potency of Drug 1, Drug 2, or both Drug1 and Drug 2.
[0019] FIG. 17 depicts Highly quantitative Deep Mutational Scanning (DMS) methods for drug discovery applications. Top panel: DMS experimental and analysis methods improvements reported in this paper. Bottom: Questions commonly encountered in drug discovery and development that can be addressed by improved DMS methodology.
[0020] FIG. 18 depicts a schematic for analysis. DMS data can be represented as a tensor, with dimensions for Position in the protein, Amino Acid substitution, and Condition + Replicate. Within the Condition + Replicate dimension, there are multiple drugs (i.e. a-MSH and THIQ) and multiple signaling pathways (i.e. Gs and Gq). Furthermore, each amino acid substitution has multiple barcodes associated with it. Our negative binomial mixed-effect model takes all of this structure into account to produce summary statistics that are the basis of our comparisons.
[0021] FIG. 19A depicts a distribution of Z-statistics for stop codons versus all other variant effects from published P2AR and MC4R (this study) DMS assays assessing Gs signaling activity.
[0022] FIG. 19B depicts downsampling simulation from the data showing the effect of number of barcodes per variant on the estimation of log2 fold change relative to wild type. Two representative positions in the a-MSH low condition were used as a starting point, and a fixed number of barcodes associated with that variant were randomly downsampled. That process was repeated for 5 independent samplings, and the resulting data was run through the model resulting in parameter estimates. Increasing the number of barcodes per variant decreased the magnitude of the error bars (+ / - 2 standard errors).
[0023] FIGs. 20A-20G depicts an analysis of the effects of all single amino acid substitutions (including nonsense variants) on MC4R’s GPCR signaling. FIG. 20A depicts heatmaps showingthe functional effects (z-scores) for all possible amino acid substitutions on MC4R activity for two GPCR signaling functions under a variety of conditions. Heatmaps showing results (both z- score and log2[fold change of variant activity over wild-type]) for all experimental conditions. The results of the Gs assay with low a-MSH stimulation are highlighted on the left (and in FIGs. 20B-20F). TM: transmembrane domain; GoF: gain-of-function; LoF: loss-of-function; WT: wild-type activity. FIG. 20B depicts a snake plot showing the sensitivity of each MC4R residue to mutation. FIG. 20C Z-scores for each variant (point), broken out by variant type and clinical (ClinVar) classification. Dark gray indicates statistically significant LOF, white is significant GOF (FDR < 1%). VUS: variant of uncertain significance. FIG. 20D depicts a functional effect (Log2[fold change of variant activity over wild-type], x-axis) for all human variants relative to the allelic frequency in the gnomAD global population (y-axis). FIG. 20E depicts DMS results for human MC4R variants (y-axis) relative to previous functional classifications of human variants in the literature. FIG. 20F depicts a DMS results (z-scores, x- axis) compared to change in a-MSH potency (relative to WT) of 25 MC4R variants made to the orthosteric binding site for a-MSH and as measured by cAMP accumulation assay (y-axis). FIG. 20G depicts fractions of MC4R variants that result in LOF, GOF, or WT activity for eight unique experimental conditions.
[0024] FIGs. 21A-21D depicts systematic identification of variants that have biased effects on MC4R signaling. FIG. 21A Principal component analysis of eight MC4R DMS conditions (Gs and Gq signaling each at four a-MSH stimulation doses: zero, low, medium, high). Each point indicates a unique variant, with those exhibiting extreme Gq or Gs bias labelled. 223L, 152R, 158R, 150G, 166S, 137L, 204V, 254P, 79S, 228R, 77R, 230S, and 79 G exhibited extreme Gq bias. 250A, 145S, 156P and 164L exhibited extreme Gs bias. FIG. 21B Ribbon structure of MC4R showing the maximum absolute PC2 value for each residue position. FIG. 21C Top: Closeup of MC4R structure shown in bottom left box in FIG. 21B, highlighting positions where variants result in Gs (open circles) or Gq (closed circles) bias. Bottom: MC4R signaling activity (log2[fold change of variant activity relative to wild-type]) for three selected variants across a- MSH doses. Error bars are + / - 2 standard errors. Med: medium; WT: wild-type FIG. 21D Left: Closeup of MC4R structure shown in bottom right box in FIG. 21B, highlighting positions where variants result in Gs or Gq bias as in FIG. 21C. Bottom: MC4R signaling activity for select variants as in FIG. 21C.
[0025] FIG. 22 depicts Identification of variants that respond to corrector treatment. Scatter plot showing is variant activity upon stimulation with 1 pM a-MSH (>EC99) in the absence of corrector (x-axis) versus is variant activity upon treatment with 1 pM Ipsen 17 followed by stimulation with 1 pM a-MSH (>EC99) (y-axis).
[0026] FIG. 23 depicts corrector therapy rescues the activity of a subset of human MC4R variants. Gs signaling activity of 21 selected MC4R variant alleles (of 6,633 tested) with (dark gray) and without (gray) Ipsen 17, a small molecule corrector that has been shown to restore the activity of misfolded MC4R. Bars represent the activity of the variant allele normalized to that of WT MC4R in the no corrector condition, and error bars are + / - two standard errors.
[0027] FIGs. 24A-24E depicts systematic identification of functional protein-ligand interactions. FIG. 24A depicts Bayesian meta-regression of the a-MSH and THIQ datasets at low agonist concentration for Gs reveals variants that specifically affect MC4R activation by each agonist. Statistically significant effects are colored by which agonist condition they impaired MC4R activation in (a-MSH, dark gray; THIQ, light gray, FDR < 0.05). The most extreme variants (Z-statistic) at positions 48, 104, and 129 are labeled. FIG. 24B depicts sideviews of MC4R structure in surface view (left) and ribbon view (right) show enrichment of a- MSH-resistant variants in the extracellular orthosteric binding site. Positions are colored by whether variants at that position perturb activation by a-MSH (dark gray) or have no effect (gray). FIG. 24C depicts the effect of all possible variants at protein positions 48, 104, and 129 at low concentration of a-MSH (dark gray) and THIQ (light gray). Variants that disproportionately affect activation by only one agonist are boxed by respective color. FIG. 24D depicts top-down surface views of MC4R (PDB: 7f58) with bound a-MSH (; PDB: 7f53) or THIQ (right; PDB: 7f58). Positions are colored by whether variants at that position perturb activation by a-MSH (dark), THIQ (dark), or neither (light gray). FIG. 24E depicts the zoomed view of the binding pocket with a-MSH bound (left; PDB: 7f53) or THIQ bound (right; PDB: 7f58). . Residues that form the HFRW motif of a-MSH and functional groups Rl, R2, and R3 of THIQ are labeled in bold. Residues 48, 104, and 129 are shown in stick form.
[0028] FIGs. 25A-25C depict the results of two deep mutational screens of GLP1R with different agonist drugs. FIG. 25A depicts loss of function variant effects for GLP1R residues within 5 angstroms of either drug (based on resolved structures of GLP1R with said drugs), classifying variants as leading to loss of function for drug 1 only, drug 2 only, both drugs 1 and 2, or neither drug. FIG. 25B depicts log2 fold change of signaling relative to wild-type for each amino acid substitution at R380 of GLP1R. FIG. 25C depicts how different mutations at R380 can improve or reduce the affinity of Drug 1 or Drug 2 to bind the target site.DETAILED DESCRIPTION
[0029] Deep mutational scanning can be used to measure the effects of amino acid variants on a target protein’s signaling and protein expression. The ability to quantify subtle differences in variant effects enabled mapping of core interactions responsible to drug binding, in addition toperipheral protein residues that contribute indirectly to inhibition. Described herein are methods and systems for high resolution mapping of a protein, including mapping of a functional site of a protein.I. METHODS AND USES
[0030] In certain aspects, described herein is a method of high resolution mapping a functional site of a protein that influences a biological function of the protein. The method may comprise providing a variant library comprising a plurality of nucleic acids. Individual members of the plurality of nucleic acids may encode the protein or a variant of the protein. The plurality of nucleic acids may encode a plurality of protein sequences comprising a plurality of alterations to at least 75% of the amino acid residues of the protein across the plurality of protein sequences. For example, the method may comprise making single amino acid substitutions to independent members of the plurality such that in aggregate mutations to at least 75% of the amino acids of the protein are captured. The method may comprise expressing the variant library in a plurality of cells. The method may comprise performing a biological assay on the plurality of cells.A. System
[0031] Described herein is a system for identifying functional variants of a protein. One example of the system is described in FIG. 1. A nucleic acid library is made to comprise multiple different variants of a protein of interest. Each variant is encoded on a nucleic acid. The nucleic acid may be linked to a unique barcode (also referred to herein as an “index sequence” or a unique molecular identifier “UMI”), which is used to identify the variant in downstream assays. The barcode may be activated as a result of a downstream reporter assay, or the barcode may be used to mark a cell expressing a particular variant that is selected in a downstream stream reporter assay (e.g., flow cytometry) The library may have multiple variants for each amino acid position of the protein (e.g., all different potential amino acid substitutions). Alternatively, the cells may comprise a barcoded reporter under the control of a constitutive promoter or may be present in the cell to mark a cell as expressing a particular variant protein. The nucleic acid library is transformed or transfected into cells and an assay is performed (with each cell on average comprising a single different variant). Then the results of the assay for each amino acid change at each position is compared (by using the barcodes to identify variants that affect biological function). This may allow for identification of functional sites in the protein. Advantages of this system include the ability to model different sources of variation and leverage the multiple layers of replication to identify functional sites with high resolution. For instance, increasing the number of barcodes per variant decreases the magnitude of the standard error of the estimate with increasing sample size.
[0032] The analysis may comprise use of a mixed effect negative binomial generalized linear model. The model may contain a random effect. The model may share barcode information between replicate and conditions. The model may incorporate sample-specific off-sets. The mean shift in barcode count may be calculated. The standard error may be calculated.
[0033] Another example of a system for high resolution protein mapping is shown in FIG. 3. In order to identify adRP mutations as class 2 or class 3, an in vitro assay was developed to determine the mutation’s effects on RHO trafficking to the plasma membrane. As shown in FIG. 3, a gene expressing a variant RHO with a single missense mutation is fused to gene expressing a transcription factor (TF) and integrated into a cell. Each cell has a single variant RHO gene and a unique barcode downstream of a response element (RE). A mutated RHO- transcription factor fusion that is properly folded and trafficked to the plasma membrane will encounter a plasma membrane specific protease. The protease will cleave the mutated RHO- transcription factor and the transcription factor will be released from the membrane. The transcription factor binds to the response element and will allow for transcription of the variant specific barcode, which can be read by next generation sequencing. By this method, a RHO mutation that prevents trafficking to the plasma membrane, i.e. a class 2 or class 3 RP disease, will have fewer barcode reads than wildtype or a mutation that does not prevent RHO trafficking as the mutated RHO-transcription factor fusion would not reach the plasma membrane to release the TF. The variant RHO barcode quantification is normalized against wildtype RHO as a trafficking score.B. Library
[0034] The systems, compositions and methods described herein comprise a nucleic acid library or a plurality of libraries where the library or plurality libraries comprising plurality of nucleic acids wherein nucleic acids of the plurality encodes a plurality of protein sequences comprising one or more amino acid alterations to a particular protein of interest. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 80% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 85% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 90% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 95% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprisingalterations to at least 96% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 97% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 98% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 99% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 99.5% of the amino acid residues of the protein across the plurality of protein sequences.
[0035] The library may comprise at least 100, 1,000, 10,000, 100,000, 1,000,000, 10,000,000 or more nucleic acids, wherein each nucleic acid encodes the protein with a single amino acid variant. In some embodiments, the library comprises at least 100, 1,000, 10,000, 100,000, 1,000,000, 10,000,000 or more nucleic acids, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 80% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the library comprises at least 100, 1,000, 10,000, 100,000, 1,000,000, 10,000,000 or more nucleic acids, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 85% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the library comprises at least 100, 1,000, 10,000, 100,000, 1,000,000, 10,000,000 or more nucleic acids, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 90% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the library comprises at least 100, 1,000, 10,000, 100,000, 1,000,000, 10,000,000 or more nucleic acids, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 95% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the library comprises at least 100, 1,000, 10,000, 100,000, 1,000,000, 10,000,000 or more nucleic acids, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 96% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the library comprises at least 100, 1,000, 10,000, 100,000, 1,000,000, 10,000,000 or more nucleic acids, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 97% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the library comprises at least 100, 1,000, 10,000, 100,000,1,000,000, 10,000,000 or more nucleic acids, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 98% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the library comprises at least 100, 1,000, 10,000, 100,000, 1,000,000, 10,000,000 or more nucleic acids, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 99% of the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the library comprises at least 100, 1,000, 10,000, 100,000, 1,000,000, 10,000,000 or more nucleic acids, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 99.5% of the amino acid residues of the protein across the plurality of protein sequences.
[0036] The plurality of alterations may comprise amino acid substitutions to a plurality of amino acids at each residue of the protein. In some embodiments, the plurality of alterations to the amino acid residues of the protein comprises amino acid substitutions to at least about 10 amino acids to about 19 amino acids. In some embodiments, the plurality of alterations to the amino acid residues of the protein comprise amino acid substitutions to at least about 10 amino acids to about 11 amino acids, about 10 amino acids to about 12 amino acids, about 10 amino acids to about 13 amino acids, about 10 amino acids to about 14 amino acids, about 10 amino acids to about 15 amino acids, about 10 amino acids to about 16 amino acids, about 10 amino acids to about 17 amino acids, about 10 amino acids to about 18 amino acids, about 10 amino acids to about 19 amino acids, about 11 amino acids to about 12 amino acids, about 11 amino acids to about 13 amino acids, about 11 amino acids to about 14 amino acids, about 11 amino acids to about 15 amino acids, about 11 amino acids to about 16 amino acids, about 11 amino acids to about 17 amino acids, about 11 amino acids to about 18 amino acids, about 11 amino acids to about 19 amino acids, about 12 amino acids to about 13 amino acids, about 12 amino acids to about 14 amino acids, about 12 amino acids to about 15 amino acids, about 12 amino acids to about 16 amino acids, about 12 amino acids to about 17 amino acids, about 12 amino acids to about 18 amino acids, about 12 amino acids to about 19 amino acids, about 13 amino acids to about 14 amino acids, about 13 amino acids to about 15 amino acids, about 13 amino acids to about 16 amino acids, about 13 amino acids to about 17 amino acids, about 13 amino acids to about 18 amino acids, about 13 amino acids to about 19 amino acids, about 14 amino acids to about 15 amino acids, about 14 amino acids to about 16 amino acids, about 14 amino acids to about 17 amino acids, about 14 amino acids to about 18 amino acids, about 14 amino acids to about 19 amino acids, about 15 amino acids to about 16 amino acids, about 15 amino acids to about 17 amino acids, about 15 amino acids to about 18 amino acids, about 15 amino acids to about 19 amino acids, about 16 amino acids to about 17 amino acids, about 16 aminoacids to about 18 amino acids, about 16 amino acids to about 19 amino acids, about 17 amino acids to about 18 amino acids, about 17 amino acids to about 19 amino acids, or about 18 amino acids to about 19 amino acids. In some embodiments, the plurality of alterations to the amino acid residues of the protein comprises amino acid substitutions to at least about 10 amino acids, about 11 amino acids, about 12 amino acids, about 13 amino acids, about 14 amino acids, about 15 amino acids, about 16 amino acids, about 17 amino acids, about 18 amino acids, or about 19 amino acids. In some embodiments, the plurality of alterations to the amino acid residues of the protein comprises amino acid substitutions to at least at least about 10 amino acids, about 11 amino acids, about 12 amino acids, about 13 amino acids, about 14 amino acids, about 15 amino acids, about 16 amino acids, about 17 amino acids, or about 18 amino acids. In some embodiments, the plurality of alterations to the amino acid residues of the protein comprises amino acid substitutions to at least at most about 11 amino acids, about 12 amino acids, about 13 amino acids, about 14 amino acids, about 15 amino acids, about 16 amino acids, about 17 amino acids, about 18 amino acids, or about 19 amino acids. In some embodiments, the protein is a human protein.
[0037] In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 80% of the codons encoding the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 90% of the codons encoding the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 95% of the codons encoding the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 98%of the codons encoding the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 99% of the codons encoding the amino acid residues of the protein across the plurality of protein sequences. In some embodiments, the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to 100% of the codons encoding the amino acid residues of the protein across the plurality of protein sequences.
[0038] The protein includes, but is not limited to, a GPCR, a receptor tyrosine kinase, a nuclear hormone receptor, an ion channel, a transcription factor, an integrin, a cell adhesion molecule, and a secreted protein. In some embodiments, the protein comprises a GPCR. In some embodiments, the protein comprises a receptor tyrosine kinase. In some embodiments, the protein comprises a nuclear hormone receptor. In some embodiments, the protein comprises anion channel. In some embodiments, the protein comprises a transcription factor. In some embodiments, the protein comprises an integrin. In some embodiments, the protein comprises a cell adhesion molecule. In some embodiments, the protein comprises a secreted protein.
[0039] The exact amino acid composition or length of a protein that can be mapped in a high resolution way by the methods and systems described herein. In some embodiments, the protein comprises at least about 10, 20, 30, 40, 50, 100, 150, 200, 250, 300, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900 or 2000 amino acids. In some embodiments, the protein comprises no more than about 50, 100, 150, 200, 250, 300, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900 or 2000 amino acids.1. Index sequences
[0040] The protein or variant of the protein described herein may be represented by a plurality of unique index sequences unique to the protein or the variant of the protein. Without being limited by theory, the resolution may be increased by the increasing number of unique index sequences. For instance, as the number of barcodes per variant is increased, the magnitude of the standard error of the estimate decreases in a manner consistent with increasing sample size. This may enable the quantification of subtle differences within a standard hypothesis testing framework. It may also provide an experimental parameter that one can vary to improve the power to detect effect at variants of particular interest.
[0041] In some embodiments, the protein or variant of the protein is represented by at least 10 different unique index sequences unique to the protein or the variant of the protein. In some embodiments, the protein or the variant of the protein is represented by at least 20 different unique index sequences unique to the protein or the variant of the protein. In some embodiments, the protein or the variant of the protein is represented by at least 30 different unique index sequences unique to the protein or the variant of the protein. In some embodiments, the protein or variant of the protein is represented by at least 10 different unique index sequences to 50 different unique index sequences. In some embodiments, the protein or variant of the protein is represented by at least 10 different unique index sequences to 15 different unique index sequences, 10 different unique index sequences to 20 different unique index sequences, 10 different unique index sequences to 25 different unique index sequences, 10 different unique index sequences to 30 different unique index sequences, 10 different unique index sequences to 35 different unique index sequences, 10 different unique index sequences to 40 different unique index sequences, 10 different unique index sequences to 45 different unique index sequences, 10 different unique index sequences to 50 different unique index sequences, 15different unique index sequences to 20 different unique index sequences, 15 different unique index sequences to 25 different unique index sequences, 15 different unique index sequences to 30 different unique index sequences, 15 different unique index sequences to 35 different unique index sequences, 15 different unique index sequences to 40 different unique index sequences, 15 different unique index sequences to 45 different unique index sequences, 15 different unique index sequences to 50 different unique index sequences, 20 different unique index sequences to 25 different unique index sequences, 20 different unique index sequences to 30 different unique index sequences, 20 different unique index sequences to 35 different unique index sequences, 20 different unique index sequences to 40 different unique index sequences, 20 different unique index sequences to 45 different unique index sequences, 20 different unique index sequences to 50 different unique index sequences, 25 different unique index sequences to 30 different unique index sequences, 25 different unique index sequences to 35 different unique index sequences, 25 different unique index sequences to 40 different unique index sequences, 25 different unique index sequences to 45 different unique index sequences, 25 different unique index sequences to 50 different unique index sequences, 30 different unique index sequences to 35 different unique index sequences, 30 different unique index sequences to 40 different unique index sequences, 30 different unique index sequences to 45 different unique index sequences, 30 different unique index sequences to 50 different unique index sequences, 35 different unique index sequences to 40 different unique index sequences, 35 different unique index sequences to 45 different unique index sequences, 35 different unique index sequences to 50 different unique index sequences, 40 different unique index sequences to 45 different unique index sequences, 40 different unique index sequences to 50 different unique index sequences, or 45 different unique index sequences to 50 different unique index sequences. In some embodiments, the protein or variant of the protein is represented by at least 10 different unique index sequences, 15 different unique index sequences, 20 different unique index sequences, 25 different unique index sequences, 30 different unique index sequences, 35 different unique index sequences, 40 different unique index sequences, 45 different unique index sequences, or 50 different unique index sequences. In some embodiments, the protein or variant of the protein is represented by at least at least 10 different unique index sequences, 15 different unique index sequences, 20 different unique index sequences, 25 different unique index sequences, 30 different unique index sequences, 35 different unique index sequences, 40 different unique index sequences, or 45 different unique index sequences. In some embodiments, the protein or variant of the protein is represented by at least at most 15 different unique index sequences, 20 different unique index sequences, 25 different unique index sequences, 30 different unique index sequences, 35 different unique indexsequences, 40 different unique index sequences, 45 different unique index sequences, or 50 different unique index sequences.C. Assays
[0042] In certain aspects, the methods described herein comprise at least one biological assay. In certain aspects, the methods described herein comprise a plurality of biological assays. In some embodiments, the biological assay comprises detection of an index sequence unique to the protein or the variant of the protein. In some embodiments, the biological assay comprises a measure of protein abundance. In some embodiments, the biological assay comprises activation of a reporter gene. In some embodiments, the biological assay comprises measurement of a biological function.
[0043] In some embodiments, the biological assay comprises activation of a reporter gene, wherein the reporter gene comprises an index sequence unique to the protein or the variant of the protein. The reporter may be operatively coupled to a promoter.
[0044] In some embodiments, the biological function comprises signal transduction of the protein, a protein expression level, protein folding, intracellular trafficking, cell surface expression, or a regulatory site. In some embodiments, the regulatory site comprises an allosteric regulatory site. The functional site may include, without limitations, a phosphorylation site, a catalytic site, a dimerization domain, a ligand binding site, a nucleic acid binding site, and a protein binding site.
[0045] This method may comprise identifying the functional site of the protein described herein. In some embodiments, the method further comprises identifying the functional site to a resolution of 1 amino acid to 10 amino acids. In some embodiments, the method further comprises identifying the functional site to a resolution of 1 amino acid to 2 amino acids, 1 amino acid to 3 amino acids, 1 amino acid to 4 amino acids, 1 amino acid to 5 amino acids, 1 amino acid to 6 amino acids, 1 amino acid to 7 amino acids, 1 amino acid to 8 amino acids, 1 amino acid to 9 amino acids, 1 amino acid to 10 amino acids, 2 amino acids to 3 amino acids, 2 amino acids to 4 amino acids, 2 amino acids to 5 amino acids, 2 amino acids to 6 amino acids, 2 amino acids to 7 amino acids, 2 amino acids to 8 amino acids, 2 amino acids to 9 amino acids, 2 amino acids to 10 amino acids, 3 amino acids to 4 amino acids, 3 amino acids to 5 amino acids, 3 amino acids to 6 amino acids, 3 amino acids to 7 amino acids, 3 amino acids to 8 amino acids, 3 amino acids to 9 amino acids, 3 amino acids to 10 amino acids, 4 amino acids to 5 amino acids, 4 amino acids to 6 amino acids, 4 amino acids to 7 amino acids, 4 amino acids to 8 amino acids, 4 amino acids to 9 amino acids, 4 amino acids to 10 amino acids, 5 amino acids to 6 amino acids, 5 amino acids to 7 amino acids, 5 amino acids to 8 amino acids, 5 amino acids to 9 amino acids, 5 amino acids to 10 amino acids, 6 amino acids to 7 amino acids, 6 amino acids to8 amino acids, 6 amino acids to 9 amino acids, 6 amino acids to 10 amino acids, 7 amino acids to 8 amino acids, 7 amino acids to 9 amino acids, 7 amino acids to 10 amino acids, 8 amino acids to 9 amino acids, 8 amino acids to 10 amino acids, or 9 amino acids to 10 amino acids. In some embodiments, the method further comprises identifying the functional site to a resolution of 1 amino acid, 2 amino acids, 3 amino acids, 4 amino acids, 5 amino acids, 6 amino acids, 7 amino acids, 8 amino acids, 9 amino acids, or 10 amino acids. In some embodiments, the method further comprises identifying the functional site to a resolution of at least 1 amino acid, 2 amino acids, 3 amino acids, 4 amino acids, 5 amino acids, 6 amino acids, 7 amino acids, 8 amino acids, or 9 amino acids. In some embodiments, the method further comprises identifying the functional site to a resolution of at most 2 amino acids, 3 amino acids, 4 amino acids, 5 amino acids, 6 amino acids, 7 amino acids, 8 amino acids, 9 amino acids, or 10 amino acids. In some embodiments, individual members of the plurality of nucleic acids of the variant library encode the protein and a variant of the protein. In some embodiments, the functional site is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15,16, 17, 18, 19, or 20 amino acids and is identified to a resolution of between 1 amino acid and 10 amino acids.
[0046] The method may further comprise contacting the plurality of cells with a test agent. Contacting the plurality of cells with a test agent may occur before performing the biological assay on the plurality of cells. In certain embodiments, test agents are applied to cells transfected with at least one of the plurality of nucleic acids of the present invention. In certain embodiments, level of activation of transcription of a reporter molecule is measured after said cells are contacted by said test agent. In certain embodiments, said test agent is a chemical, small-molecule, biological molecule, polypeptide, polynucleotide, aptamer, or any combination thereof. In certain embodiments, a single test agent is applied to a population of cells. In certain embodiments, a plurality of test agents is applied to a population of cells. In some embodiments, the biological assay that comprises a measure of protein abundance measures steady-state abundance, cell surface expression, protein trafficking, or folding of the protein.
[0047] In some embodiments, the method further comprises performing a second biological assay. The second biological assay may be any of the assays described herein. In some embodiments, the second biological assay comprises a measure of signaling by the protein.D. Cells
[0048] Cells useful in the method described herein are generally those that are able to be easily rendered transgenic with nucleic acids encoding the variant proteins and / or comprising reporter systems or barcodes. The system nucleic acid(s) encoding a synthetic transcription factor and a reporter element can be transfected or transduced into suitable cell line using methods known in the art, such as calcium phosphate transfection, lipid based transfection (e.g.,Lipofectamine™, Lipofectamine-2000™, Lipofectamine-3000™, or Fugene® HD), electroporation, or viral transduction. The cell can also be a population of cells of the same type grown to confluency or near confluency in an appropriate tissue culture vessel.
[0049] In certain embodiments, the cell used comprises a stable integration of either the nucleic acid encoding the synthetic transcription factor, the nucleic acid comprising the reporter element, or both. Stable cell lines can be made using random integration of a linearized plasmid, virally or transposon directed integration, or directed integration, for example using site specific recombination between an AttP and an AttB site. In certain embodiments, either of the nucleic acids are encoded at a safe landing site such as the AAVS1 site.
[0050] In certain embodiments, the cell or cell population used in the system is a eukaryotic cell. In certain embodiments, the cell or cell population is a mammalian cell. In certain embodiments, the cell or cell population is a human cell. In certain embodiments, the cell or cell population is SH-SY5Y, Human neuroblastoma; Hep G2, Human Caucasian hepatocyte carcinoma; 293 (also known as HEK 293), Human Embryo Kidney; RAW 264.7, Mouse monocyte macrophage; HeLa, Human cervix epitheloid carcinoma; MRC-5 (PD 19), Human fetal lung; A2780, Human ovarian carcinoma; CACO-2, Human Caucasian colon adenocarcinoma; THP 1, Human monocytic leukemia; A549, Human Caucasian lung carcinoma; MRC-5 (PD 30), Human fetal lung; MCF7, Human Caucasian breast adenocarcinoma; SNL 76 / 7, Mouse SIM strain embryonic fibroblast; C2C12, Mouse C3H muscle myoblast; Jurkat E6.1, Human leukemic T cell lymphoblast; U937, Human Caucasian histiocytic lymphoma; L929, Mouse C3H / An connective tissue; 3T3 LI, Mouse Embryo; HL60, Human Caucasian promyelocytic leukaemia; PC- 12, Rat adrenal phaeochromocytoma; HT29, Human Caucasian colon adenocarcinoma; OE33, Human Caucasian oesophageal carcinoma; OE19, Human Caucasian oesophageal carcinoma; NTH 3T3, Mouse Swiss NIH embryo; MDA- MB-231, Human Caucasian breast adenocarcinoma; K562, Human Caucasian chronic myelogenous leukemia; U-87 MG, Human glioblastoma astrocytoma; MRC-5 (PD 25), Human fetal lung; A2780cis, Human ovarian carcinoma; B9, Mouse B cell hybridoma; CHO-K1, Hamster Chinese ovary; MDCK, Canine Cocker Spaniel kidney; 132 INI, Human brain astrocytoma; A431, Human squamous carcinoma; ATDC5, Mouse 129 teratocarcinoma AT805 derived; RCC4 PLUS VECTOR ALONE, Renal cell carcinoma cell line RCC4 stably transfected with an empty expression vector, pcDNA3, conferring neomycin resistance.;HUVEC (S200-05n), Human Pre-screened Umbilical Vein Endothelial Cells (HUVEC); neonatal; Vero, Monkey African Green kidney; RCC4 PLUS VHL, Renal cell carcinoma cell line RCC4 stably transfected with pcDNA3-VHL; Fao, Rat hepatoma; J774A.1, Mouse BALB / c monocyte macrophage; MC3T3-E1, Mouse C57BL / 6 calvaria; J774.2, Mouse BALB / cmonocyte macrophage; PNT1 A, Human post pubertal prostate normal, immortalised with SV40; U-2 OS, Human Osteosarcoma; HCT 116, Human colon carcinoma; MAI 04, Monkey African Green kidney; BEAS-2B, Human bronchial epithelium, normal; NB2-11, Rat lymphoma; BHK 21 (clone 13), Hamster Syrian kidney; NSO, Mouse myeloma; Neuro 2a, Mouse Albino neuroblastoma; SP2 / 0-Agl4, Mouse x Mouse myeloma, non-producing; T47D, Human breast tumor; 1301, Human T-cell leukemia; MDCK-II, Canine Cocker Spaniel Kidney; PNT2, Human prostate normal, immortalized with SV40; PC-3, Human Caucasian prostate adenocarcinoma; TF1, Human erythroleukaemia; COS-7, Monkey African green kidney, SV40 transformed; MDCK, Canine Cocker Spaniel kidney; HUVEC (200-05n), Human Umbilical Vein Endothelial Cells (HUVEC); neonatal; NCI-H322, Human Caucasian bronchioalveolar carcinoma; SK.N.SH, Human Caucasian neuroblastoma; LNCaPEGC, Human Caucasian prostate carcinoma; OE21, Human Caucasian oesophageal squamous cell carcinoma; PSN1, Human pancreatic adenocarcinoma; ISHIKAWA, Human Asian endometrial adenocarcinoma; MFE- 280, Human Caucasian endometrial adenocarcinoma; MG-63, Human osteosarcoma; RK 13, Rabbit kidney, BVDV negative; EoL-1 cell, Human eosinophilic leukemia; VCaP, Human Prostate Cancer Metastasis; tsA201, Human embryonal kidney, SV40 transformed; CHO, Hamster Chinese ovary; HT 1080, Human fibrosarcoma; PANC-1, Human Caucasian pancreas; Saos-2, Human primary osteogenic sarcoma; Fibroblast Growth Medium (116K-500), Fibroblast Growth Medium Kit; ND7 / 23, Mouse neuroblastoma x Rat neuron hybrid; SK-OV-3, Human Caucasian ovary adenocarcinoma; COV434, Human ovarian granulosa tumor; Hep 3B, Human hepatocyte carcinoma; Vero (WHO), Monkey African Green kidney; Nthy-ori 3-1, Human thyroid follicular epithelial; U373 MG (Uppsala), Human glioblastoma astrocytoma; A375, Human malignant melanoma; AGS, Human Caucasian gastric adenocarcinoma; CAKI 2, Human Caucasian kidney carcinoma; COLO 205, Human Caucasian colon adenocarcinoma;COR-L23, Human Caucasian lung large cell carcinoma; IMR 32, Human Caucasian neuroblastoma; QT 35, Quail Japanese fibrosarcoma; WI 38, Human Caucasian fetal lung; HMVII, Human vaginal malignant melanoma; HT55, Human colon carcinoma; TK6, Human lymphoblast, thymidine kinase heterozygote; SP2 / 0-AG14 (AC -FREE), Mouse x mouse hybridoma non-secreting, serum-free, animal component (AC) free; AR42J, or Rat exocrine pancreatic tumor, or any combination thereof.E. Partitions
[0051] In some embodiments, the plurality of cells described herein are comprised within a partition. The partition may comprise a tissue culture flask, plate, or dish. In some embodiments, the partition comprises a well of a well-plate. In certain embodiments, the nucleic acid systems of the present invention can be utilized in multiwell plate experiments. Non-limiting examplesof multiwell plates compatible with the nucleic acid relay systems of the present invention include 6, 12, 24, 48, 96, 384, or 1,536 well plates. In certain embodiments, each well of a multiwell plate comprises a cell population transfected with one of the plurality of nucleic acids as described herein. In certain embodiments, each well of a multiwell plate comprises a cell population transfected with a plurality of nucleic acids. In certain embodiments, each well comprises multiple cell populations, each cell population transfected with a single nucleic acid relay. In some embodiments each cell of the plurality of cells comprises only one of the plurality of nucleic acids.
[0052] Cell populations transfected with nucleic acids of the present invention can be any size. In certain embodiments, cell populations comprise 1,000, 10,000, 100,000, 1,000,000, 10,000,000 or more cells. In certain embodiments, at least about 1,000 or more cells are transfected with one or more nucleic acids. In certain embodiments, at least about 10,000 or more cells are transfected with one or more nucleic acids. In certain embodiments, at least about 100,000 or more cells are transfected with one or more of the plurality of nucleic acids. In certain embodiments, at least about 1,000,000 or more cells are transfected with one or more nucleic acids. In certain embodiments, at least about 10,000,000 or more cells are transfected with one or more nucleic acids. In some embodiments each cell of the plurality of cells comprises only one of the plurality of nucleic acids.F. Vectors
[0053] The nucleic acids of the present invention are compatible with many vectors common in the art. Non-limiting examples of vectors include genomic integrated vectors, episomal vectors, plasmids, viral vectors, cosmids, bacterial artificial chromosomes, and yeast artificial chromosomes. Non-limiting examples of viral vectors compatible with the nucleic acids of the present invention include vectors derived from lentiviruses, retroviruses, adenoviruses, and adeno-associated viruses. In certain embodiments, the nucleic acids of the present invention are present on vectors comprising sequences that direct site specific integration into a defined location or a restricted set of sites in the genome (e.g. AttP-AttB recombination).
[0054] In certain embodiments, one of the plurality of nucleic acids as described herein is incorporated into a single vector. In certain embodiments, said single vector is transfected into a cell transiently. In certain embodiments, said single vector is transfected into a cell stably.
[0055] Vectors comprising the plurality of nucleic acids described herein or portions thereof may be constructed using many well-known molecular biology techniques. Detailed protocols for numerous such procedures, including amplification, cloning, mutagenesis, transformation, and the like, are described in, e.g., in Ausubel et al. Current Protocols in Molecular Biology (supplemented through 2012) John Wiley & Sons, New York 10 (“Ausubel”); Sambrook et al.Molecular Cloning - A Laboratory Manual (4th Ed.), Vol. 1-3, Cold Spring Harbor Laboratory, Cold Spring Harbor, New York, 2012 (“Sambrook”); and Abelson et al. Guide to Molecular Cloning Techniques (Methods in Enzymology) volume 152 Academic Press, Inc., San Diego, CA (“Abelson”).II. DEFINITIONS
[0056] Unless defined otherwise, all terms of art, notations and other technical and scientific terms or terminology used herein are intended to have the same meaning as is commonly understood by one of ordinary skill in the art to which the claimed subject matter pertains. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a substantial difference over what is generally understood in the art.
[0057] Throughout this application, various embodiments may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the disclosure. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.
[0058] As used in the specification and claims, the singular forms “a”, “an” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a sample” includes a plurality of samples, including mixtures thereof.
[0059] The terms “determining,” “measuring,” “evaluating,” “assessing,” “assaying,” and “analyzing” are often used interchangeably herein to refer to forms of measurement. The terms include determining if an element is present or not (for example, detection). These terms can include quantitative, qualitative or quantitative and qualitative determinations. Assessing can be relative or absolute. “Detecting the presence of’ can include determining the amount of something present in addition to determining whether it is present or absent depending on the context.
[0060] The terms “subject,” and “patient” may be used interchangeably herein. A “subject” can be a biological entity containing expressed genetic materials. The biological entity can be a plant, animal, or microorganism, including, for example, bacteria, viruses, fungi, and protozoa. The subject can be a mammal. The mammal can be a human. The subject may be diagnosed orsuspected of being at high risk for a disease. In some cases, the subject is not necessarily diagnosed or suspected of being at high risk for the disease.
[0061] As used herein, the term “about” a number refers to that number plus or minus 10% of that number. The term “about” a range refers to that range minus 10% of its lowest value and plus 10% of its greatest value.
[0062] As used herein the term “about” refers to an amount that is near the stated amount by 10%.
[0063] The terms “polypeptide” and “protein” are used interchangeably to refer to a polymer of amino acid residues, and are not limited to a minimum length. Polypeptides, including the provided polypeptide chains and other peptides, e.g., linkers and binding peptides, may include amino acid residues including natural and / or non-natural amino acid residues. The terms also include post-expression modifications of the polypeptide, for example, glycosylation, sialylation, acetylation, phosphorylation, and the like. In some aspects, the polypeptides may contain modifications with respect to a native or natural sequence, as long as the protein maintains the desired activity. These modifications may be deliberate, as through site-directed mutagenesis, or may be accidental, such as through mutations of hosts which produce the proteins or errors due to PCR amplification.
[0064] Percent (%) sequence identity with respect to a reference polypeptide sequence is the percentage of amino acid residues in a candidate sequence that are identical with the amino acid residues in the reference polypeptide sequence, after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent sequence identity, and not considering any conservative substitutions as part of the sequence identity. Alignment for purposes of determining percent amino acid sequence identity can be achieved in various ways that are known for instance, using publicly available computer software such as BLAST, BLAST-2, ALIGN or Megalign (DNASTAR) software. Appropriate parameters for aligning sequences are able to be determined, including algorithms needed to achieve maximal alignment over the full length of the sequences being compared. For purposes herein, however, % amino acid sequence identity values are generated using the sequence comparison computer program ALIGN-2. The ALIGN-2 sequence comparison computer program was authored by Genentech, Inc., and the source code has been filed with user documentation in the U.S. Copyright Office, Washington D.C., 20559, where it is registered under U.S. Copyright Registration No. TXU510087. The ALIGN-2 program is publicly available from Genentech, Inc., South San Francisco, Calif., or may be compiled from the source code. The ALIGN-2 program should be compiled for use on a UNIX operating system, including digital UNIX V4.0D. All sequence comparison parameters are set by the ALIGN-2 program and do not vary.
[0065] In situations where ALIGN-2 is employed for amino acid sequence comparisons, the % amino acid sequence identity of a given amino acid sequence A to, with, or against a given amino acid sequence B (which can alternatively be phrased as a given amino acid sequence A that has or comprises a certain % amino acid sequence identity to, with, or against a given amino acid sequence B) is calculated as follows: 100 times the fraction X / Y, where X is the number of amino acid residues scored as identical matches by the sequence alignment program ALIGN-2 in that program's alignment of A and B, and where Y is the total number of amino acid residues in B. It will be appreciated that where the length of amino acid sequence A is not equal to the length of amino acid sequence B, the % amino acid sequence identity of A to B will not equal the % amino acid sequence identity of B to A. Unless specifically stated otherwise, all % amino acid sequence identity values used herein are obtained as described in the immediately preceding paragraph using the ALIGN-2 computer program.
[0066] The terms “identity,” “identical,” or “percent identical” when used herein to describe to a nucleic acid sequence, relative to a reference sequence, can be determined using the formula described by Karlin and Altschul (Proc. Natl. Acad. Sci. USA 87: 2264-2268, 1990, modified as in Proc. Natl. Acad. Sci. USA 90:5873-5877, 1993). Such a formula is incorporated into the basic local alignment search tool (BLAST) programs of Altschul et al. (J. Mol. Biol. 215: 403- 410, 1990). Percent identity of sequences can be determined using the most recent version of BLAST, as of the filing date of this application.
[0067] The polypeptides of the systems described herein can be encoded by a nucleic acid. A nucleic acid is a type of polynucleotide comprising two or more nucleotide bases. In certain embodiments, the nucleic acid is a component of a vector that can be used to transfer the polypeptide encoding polynucleotide into a cell. As used herein, the term “vector” refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. One type of vector is a genomic integrated vector, or “integrated vector,” which can become integrated into the chromosomal DNA of the host cell. Another type of vector is an “episomal” vector, e.g., a nucleic acid capable of extra-chromosomal replication. Vectors capable of directing the expression of genes to which they are operatively linked are referred to herein as “expression vectors.” Suitable vectors comprise plasmids, bacterial artificial chromosomes, yeast artificial chromosomes, viral vectors and the like. In the expression vectors regulatory elements such as promoters, enhancers, polyadenylation signals for use in controlling transcription can be derived from mammalian, microbial, viral or insect genes. The ability to replicate in a host, usually conferred by an origin of replication, and a selection gene to facilitate recognition of transformants may additionally be incorporated. Vectors derived from viruses, such as lentiviruses, retroviruses, adenoviruses, adeno-associated viruses, and the like, may beemployed. Plasmid vectors can be linearized for integration into a chromosomal location. Vectors can comprise sequences that direct site-specific integration into a defined location or restricted set of sites in the genome (e.g., AttP-AttB recombination). Additionally, vectors can comprise sequences derived from transposable elements for integration.
[0068] As used herein the term “transfection” or “transfected” refers to methods that intentionally introduce an exogenous nucleic acid into a cell through a process commonly used in laboratories. Transfection can be effected by, for example, lipofection, calcium phosphate precipitation, viral transduction, or electroporation. Transfection can be either transient or stable.
[0069] Differential signaling as described herein refers to the ability of a protein in a signal transduction cascade to directly or indirectly activate of inhibit signaling of different signaling pathways through the protein and effect different biological processes. Such different biological signaling may comprise opposite biological effects (transcription of inhibitory genes vs. transcription of inducing genes for a certain biological process), or a different scope or magnitude of the same biological effect (e.g., transcription or repression of different subsets of genes that achieve the same or a similar biological effect, or stronger induction or repression of a biological effect).
[0070] As used herein “reporter activity” refers to the empirical readout from the reporter. For example, a luciferase reporter will have a luminescent readout when incubated with an appropriate substrate. Other reporters like a fluorescent protein may not require a substrate but can be measured via microscopy or a fluorescence plate reader for example.III. EXAMPLESExample 1: High-resolution DMS for identification of Retinitis Pigmentosa mutations that respond to a specific rhodopsin corrector
[0071] Described in this example is an in vitro diagnostic assay to determine a class for a Retinitis Pigmentosa disease variant and potential for response to a specific rhodopsin corrector molecule. An assay as depicted in FIG. 3 was utilized for examples 1 and 2.
[0072] The following assay protocol was followed:
[0073] Day 1 - Seed cells1. Cells were grown in culture.2. Cells were trypsinized, collected, spun down, and counted.3. Cells were diluted to desired amount in a new tube.4. Cells were resuspended in media with serum.5. Cells were seeded into a dish.6. Dishes were incubated overnight at 37° C.
[0074] Day 2 - Doxycycline inducing cells1. Prepared 6x Doxycycline media in Opti-MEM.2. Dispensed Doxycycline Opti-MEM mixture into the dish.3. Mixed di shes thoroughly .4. Incubated plates overnight at 37° C.
[0075] Day 3 - Read signal output1. Barcodes were isolated, and quantified by next generation sequencing.
[0076] As shown in FIG. 4, a heatmap was generated by deep mutational screening of the trafficking score of each missense mutation against wildtype RHO. Each individual amino acid of RHO was mutated to one of twenty amino acids. All 348 amino acids of RHO were tested with twenty amino acid changes, including a stop codon. Each mutant RHO was encoded with a unique barcode for identification. The assay was performed as described above and the barcode reads for each unique mutation was quantified and normalized against wildtype. The normalized trafficking was evaluated for each variant of RHO. Variant RHO with trafficking defects of less than about 50% compared to wildtype were considered to be class 2 or class 3 RP disease mutations. Overall, about 15% of missense mutations caused a <50% trafficking defect.Mutations throughout RHO were shown to affect trafficking, and the most sensitive regions were the N-terminus and second extracellular domain. Additionally, substantially all of the stop codon nonsense mutations had trafficking defects. Mutations identified in this assay as being class 2 or class 3 are listed in Table 1.Table 1: Class 2 and Class 3 mistrafficking variants of Rho (1,026)G51F, G51K, G51N, G51P, G51R, G51W, G51Y, F52D, F52E, F52K, F52N, F52P, F52Q, F52R, P53E, P53R, I54D, I54H, I54K, I54R, N55D, N55E, N55I, N55M, N55P, N55R, N55T, N55V, N55W, N55Y, F56D, F56E, F56K, F56P, F56R, L57D, L57K, L57P, L57R, T58D, L59D, L59E, L59H, L59K, L59N, L59R, V61D, V61P, V61R, I75D, I75R, L76D, N78A, N78F, N78I, N78K, N78L, N78P, N78R, N78V, N78W, N78Y, V81D, V81K, V81P, V81R, L84P, L84R, F85R, V87D, V87H, V87K, V87N, V87P, V87R, V87W, L88R, G89D, G89R, G90F, G90K, G90R, G90W, G90Y, F91R, T92D, T92E, T92K, T92R, T94K, T94R, T94Y, L95D, L95K, L95N, L95P, L95R, T97E, T97K, T97P, T97R, T97W, S98D, S98P, S98R, L99D, H100P, G101D, G1O1I, G1O1P, G101V, Y102D, Y102E, Y102I, Y102N, Y102P, Y102V, F1O3K, F1O3P, F1O3R, G106C, G106D, G106E, G106F, G106H, G106I, G106L, G106M, G106P, G106R, G106V, G106W, G106 Y, C 110 A, C 11 OD, C 11 OE, C 11 OF, C110G, C110H, CHOI, C110K, C110L, C110M, CHON, CHOP, C110Q, C110R, Cl IOS, Cl 10T, Cl 10V, Cl 10W, CHOY, LI 12D, LI 12P, El 13G, El 13K, El 13P, El 13R, El 13W, G114C, G114D, G114E, G114F, G114H, G1141, G114K, G114L, G114M, G114N, G114P, G114Q, G114R, G114S, G114T, G114V, G114W, G114Y, Fl 15D, Fl 15E, Fl 15H, Fl 15K, F115N, F115P, F115Q, F115R, F116K, F116R, T118P, L119D, L119K, L119P, L119R, I123P, I123R, L125D, W126D, W126I, W126K, W126P, W126R, W126V, S127F, S127I, S127K, S127L, S127P, S127R, S127W, S127Y, L128P, V129H, V129K, V129R, V129Y, V130K, V130R, L131D, L131P, L131R, A132P, A132R, I133D, E134P, R135F, R135G, R135I, R135L, R135M, R135P, R135W, R135Y, A153V, V157H, V157K, V157R, A158D, F159D, F159E, F159K, F159P, F159R, T160D, T160F, T160H, T160K, T160R, T160W, T160Y, W161K, W161P, W161R, V162D, V162R, M163D, M163K, M163R, M163Y, A164D, A164E, A164F, A164H, Al 641, A164K, A164L, A164N, A164P, A164Q, A164R, A164W, A164Y, L165D, L165E, L165K, L165P, L165R, A166K, A166P, A166R, C167P, C167W, A168D, A168E, A168F, A168H, A168I, A168K, A168L, A168M, A168N, A168P, A168Q, A168R, A168W, A168Y, A169D, A169K, A169P, A169R, P170D, P170E, P170H, P170I, P170K, P170N, P170R, P170T, P170V, P170Y, P171A, P171C, P171D, P171E, P171F, P171G, P171H, P171I, P171K, P171L, P171M, P171N, P171Q, P171R, P171S, P171T, P171V, P171W, P171Y, L172D, L172E, L172H, L172K, L172P, L172Q, L172R, A173P, G174C, G174P, W175C, W175D, W175E, W175G, W175K, W175N, W175P, W175Q, W175R, W175S, W175T, S176D, S176E, S176F, S176H, S176I, S176K, S176L, S176M, S176P, S176Q, S176R, S176T, SI 76V, S176W, S176Y, R177C, R177D, R177E, R177F, R177G, R177I, R177K, R177L, R177N, R177P, R177T, R177V, R177W, R177Y, Y178A, Y178C, Y178D, Y178E, Y178G, Y178H, Y178I, Y178K, Y178L, Y178M, Y178N, Y178P, Y178Q, Y178R, Y178S, Y178T, Y178V, Y178W, I179A, I179C, I179D, I179E, I179F, I179G, I179H, I179K, I179N, I179P, I179Q, I179R, I179S, I179W, I179Y, P180A, P180C, P180D, P180E, P180F, P180G, P180H, Pl 801, P180K, P180L, P180M, P180N, P180Q, P180R, P180S, P180T, P180V, P180W, P180Y, E181C, E181G, E181K, E181P, E181R, G182C, G182D, G182E, G182F, G182H, G182I, G182K, G182L, G182M, G182N, G182P, G182Q, G182R, G182S, G182T, G182V, G182W, G182Y, L183D, L183E, L183G, L183K, L183P, L183R, QI 84 A, Q184D, Q184E, Q184F, Q184H, QI 841, Q184K, Q184N, Q184P, Q184R, Q184S, Q184T, Q184V, Q184W, Q184Y, C185D, C185E, C185K, C185P, C185R, C185W, C185Y, S186K, S186P, S186R, C187A, C187D, C187E, C187F, C187G, C187H, Cl 871, C187K, C187L, C187M, C187N, C187P, C187Q, C187R, C187S, C187T, C187V, C187W, C187Y, G188C, G188E, G188H, G188I, G188K, G188L, G188M, G188P, G188Q, G188R, G188V, G188Y, I189C, I189D, I189E, I189G, I189K, I189N, I189Q, I189R, I189S, I189T, I189W, D190A, D190C, D190E, D190F, D190G, D190H, D190I, D190K, D190L, D190M, D190N, D190P, D190Q, D190R, D190S, D190T, D190V, D190W, D190Y, Y191D, Y191E, Y191G, Y191N, Y191P, Y191Q, Y191R, Y191S, Y192D, Y192G, Y192K, Y192P, T193E, T193P, T193W, N200E, N200F, N200L, N200W, F203C, F203D, F203E, F203G, F203H, F203I, F203K, F203L, F203M, F203N, F203P,
[0077] As shown in FIGS. 6A-6B, a select number of known mutations were examined. Using the data from the heatmap, select mutations of interest were highlighted to demonstrate the effect of those mutations in this assay. Class 1-7 variants putatively identified in Athanasiou, et al. were selected to demonstrate efficacy of the assay for selectively identifying class 2 and class 3 disease states. Using this assay, substantially all of the putative class 2 and class 3 mutations were identified as such by having <50% trafficking score, such as N15S, T17M, V20G, P23A, P23H, P23L, Q28H, G51R, P53R, V87D, G89D, G106R G106W, Cl 10F, Cl 10R, Cl IOS, CHOY, E113K, W161R, A164E, C167W, P171Q, P171L, P171S, Y178N, Y178D, Y178C, E181K, G182S, G182V, C185R, C187G, C187Y, G188R, G188E, D190N, D190G, D190Y, H211R, H21 IP, C222R, P267R, P267L, R135G, R135L, R135P, and R135W.Example 2: High resolution DMS assay for evaluating variant responsiveness to a rhodopsin corrector molecule
[0078] Herein provides a high-resolution DMS assay to determine whether a particular variant responds to a rhodopsin corrector molecule based on the mutation / class of RP disease.
[0079] The following assay protocol was followed:
[0080] Day 1 - Seed cells1. Cells were grown in culture.2. Cells were trypsinized, collected, spun down, and counted.3. Cells were diluted to desired amount in a new tube.4. Cells were resuspended in media with serum.5. Cells were seeded into a dish.6. Dishes were incubated overnight at 37° C.
[0081] Day 2 - Doxycycline inducing cells and treatment1. Prepared 6x Doxycycline media in Opti-MEM.2. Dispensed Doxycycline Opti-MEM mixture into the dish.3. Cells were treated with DMSO or a rhodopsin corrector molecule.4. Mixed dishes thoroughly.5. Incubated plates overnight at 37° C.
[0082] Day 3 - Read signal output1. Barcodes were isolated, and quantified by next generation sequencing.
[0083] As shown in FIG. 5, a heatmap was generated by deep mutational screening of the trafficking score of each missense mutation against wildtype RHO. Cells were additionally treated with a rhodopsin corrector molecule before the RHO trafficking was measured. The assay was performed, except with addition of a RHO corrector molecule, as described above and the barcode reads for each unique mutation was quantified and normalized against wildtype. The normalized trafficking was evaluated for each variant RHO. Variant RHO with trafficking defects of less than about 50% compared to wildtype were considered to be class 2 or class 3 RP disease mutations. If a variant RHO had <50% trafficking score without treatment, but had >50% trafficking score with treatment, the variant is considered rescued. Compared against no treatment as shown in FIG. 4, the corrector molecule rescued -70% of RHO variants.Additionally, substantially all of the stop codon nonsense mutation variants were not rescued by the corrector molecule. Mi straffi eking mutations identified and rescued in this assay are listed in Table 2.Table 2: Mistrafficking mutants rescued (815)L47D, L47K, L47P, L47R, I48D, I48E, I48H, I48K, I48P, V49P, L50D, L50E, L50K, G51D, G51F, G51K, G51N, G51P, G51R, G51W, G51Y, F52E, F52K, F52N, F52P, F52Q, P53E, I54D, I54H, I54K, N55D, N55E, N55I, N55M, N55P, N55R, N55T, N55V, N55W, N55Y, F56E, F56K, L57D, L57K, L57P, L57R, T58D, L59D, L59E, L59H, L59N, V61D, V61P, V61R, I75D, I75R, L76D, N78A, N78I, N78K, N78L, N78P, N78V, N78Y, V81D, V81K, V81R, L84P, L84R, F85R, V87D, V87H, V87N, V87P, V87W, L88R, G89D, G89R, G90F, G90K, G90R, G90W, G90Y, F91R, T92D, T92E, T92K, T92R, T94K, T94R, T94Y, L95D, L95K, L95N, L95P, L95R, T97E, T97P, T97R, T97W, S98D, S98P, S98R, L99D, H100P, G101D, G1O1I, G1O1P, G101V, Y102D, Y102E, Y102I, Y102N, Y102P, Y102V, F1O3K, F1O3P, F1O3R, G106C, G106D, G106E, G106F, G106H, G106I, G106L, G106M, G106P, G106R, G106V, G106W, G106Y, C110D, CHOE, C11OF, C110K, C11OL, C110M, Cl IOS, CHOW, L112D, L112P, E113G, E113K, E113R, E113W, G114C, G114E, G114I, G114K, G114P, G114R, G114S, G114T, F115H, F115N, F115P, F116K, F116R, T118P, L119K, L119R, L125D, W126D, W126I, W126P, W126V, S127F, S127I, S127L, V129H, V129K, V129Y, V130K, V130R, L131D, L131R, A132P, A132R, I133D, E134P, R135F, R135G, R135I, R135L, R135M, R135P, R135W, R135Y, A153V, V157H, V157K, V157R, A158D, F159D, F159E, F159K, F159P, T160F, T160H, T160W, T160Y, W161K, W161R, V162D, V162R, M163D, M163K, M163Y, A164D, A164F, A164P, A164W, L165D, L165E, L165K, L165P, L165R, A166K, A166P, C167W, A168E, A168I, A168K, A168N, A168P, A168Q, A168R, A168W, A169D, A169K, A169P, P170D, P170E, P170I, P170K, P170N, P170T, P170V, P170Y, P171A, P171G, P171I, P171L, P171N, P171Q, P171R, P171S, P171T, P171V, P171W, P171Y, L172D, L172E, L172H, L172K, L172P, L172Q, L172R, A173P, G174C, G174P, W175C, W175E, W175G, W175K, W175N, W175Q, W175S, W175T, S176D, S176E, S176F, S176H, SI 761, S176P, S176T, R177C, R177D, R177E, R177F, R177G, R177I, R177K, R177L, R177N, R177T, R177W, R177Y, Y178C, Y178G, Y178H, Y178I, Y178L, Y178M, Y178N, Y178W, 1179 A, I179C, I179D, I179E, I179F, I179G, I179H, I179K, I179N, I179P, I179Q, I179R, I179S, I179W, I179Y, P180A, P18OC, P180D, P18OE, P18OF, P18OG, P180H, P18OI, P180K, P18OL, P180M, P180N, P18OQ, P180R, P18OS, P18OT, P180V, P180W, P180Y, E181C, E181G, E181K, E181P, E181R, G182C, G182D, G182E, G182F, G182H, G182I, G182K, G182L, G182M, G182N, G182P, G182Q, G182R, G182S, G182T, G182V, G182W, G182Y, L183D, L183E, L183G, L183K, L183P, L183R, Q184A, Q184D, Q184E, Q184F, Q184H, Q184I, Q184K, Q184N, Q184P, Q184R, Q184S, Q184T, Q184V, Q184W, Q184Y, C185D, C185E, C185K, C185P, C185R, C185W, C185Y, S186K, S186P, S186R, C187A, C187E, Cl 871, C187K, C187L, C187M, C187N, C187Q, C187S, C187T, C187V, C187Y, G188C, G188E, G188H, G188I, G188K, G188M, G188Q, G188R, G188V, I189C, I189G, I189N, I189Q, I189R, I189S, I189T, DI 90 A, D190C, D190E, D190F, D190G, D190H, DI 901, D190K, D190M, D190N, D190P, D190Q, D190R, D190S, D190T, D190V, D190W, D190Y, Y191D, Y191E, Y191G, Y191N, Y191P, Y191Q, Y191R, Y191S, Y192D, Y192G, Y192K, Y192P, T193E, T193P, T193W, N200E, N200F, N200L, N200W, F203C, F203E, F203G, F203H, F203L, F203M, F203P, F203S, F203V, V204D, V204E, Y206E, Y206G, Y206L, Y206N, Y206P, F208D, F208K, F208R, V209D, V209K, V210D, V210P, H211L, H211R, H211V, H211W, I214D, I214K, I214P, P215A, P215K, P215L, I218E, I218H, I218N, I218Y, I219D, I219E, I219P, F221D, C222D, C222R, M253K, V254K, M257D, M257G, M257P, A260R, V266D, V266H, V266K, V266R, P267C, P267D, P267F, P267H, P267I, P267L, P267M, P267N, P267V, Y268I, Y268P, S270P, V271P, A272F, F273D, F273P, F276D, T277P, H278P, P285D, M288D, M288K, T289K, T289P, T289R, I290D, I290E, I290K, P291D, P291K, F293P, F293R, F294D, F294R, K296T, K296V, S297D, S297R, A299D, I300D, I300E, BOOH, I300Q, N302Q, P303E, P303F, P303W, P303Y, V304D, V304E, V304K, V304Q, V304R, I305E, I305H, I305K, I305Q, I305Y, I307D, I307E, I307K,
[0084] As shown in FIGs. 7A-7B, the select number of known mutations in the art were examined. Using the data from the heatmap, select mutations of interest were highlighted to demonstrate the effect of the rhodopsin corrector molecule in rescuing RHO variants in this assay. Many known class 2 and class 3 mutations were rescued by the treatment, such as N15S, T17M, V20G, P23A, P23H, P23L, Q28H, G51R, V87D, G89D, G106R G106W, E113K, P171S, E181K, G182S, G182V, C185R, G188E, D190N, D190G, D190Y, C222R, P267L, R135G, R135L, R135P, and R135W. Surprisingly, some of the rescued mutations had recovered trafficking scores equivalent to wildtype.
[0085] Different Rho corrector compounds were also tested. A second Rho corrector compound rescued the following mi straffi eking mutations listed in Table 3. A third Rho corrector compound rescued the following mistrafficking mutations listed in Table 4.Table 3: Mistrafficking mutants rescued (827)Table 4: Mistrafficking mutants rescued (508)Example 3: High-Resolution DMS allows for the determination of proteins regions that affect IFN-alpha protein expression and signaling with greater precision
[0086] Deep mutational scanning was used to measure the effects of more than 20,000 amino acid variants on TYK2’s interferon-alpha (INF-a) signaling and protein expression. Morevariants and barcodes were used than in previous experiments. Table 5 depicts a comparison with VAMP-seq. Additional details on the VAMP-seq process can be found in Matreyek et al, Multiplex assessment of protein variant abundance by massively parallel sequencing. Nat Genet 50, 874-882 (2018) and Boyle et al, Deep mutational scanning of CYP2C19 reveals a substrate specificity-abundance tradeoff bioRxiv (2023).Table 5
[0087] IFNa signaling and protein stability are depicted in FIG. 8. FIG. 9 depicts protein expression against IFNa signaling.
[0088] Positions of variants that affect IFNa signaling but are not significantly destabilizing are depicted in FIG. 10. Variants that impacted signaling but not stability were particularly frequent in the pseudokinase domain. This includes both the kinase active site and a known allosteric regulatory site on the pseuodokinase. This also allows for capture of areas of previously uncharacterized regulatory function.
[0089] These results illustrate that DMS can be used to identify novel target sites and locations for allosteric regulation.Example 4: TYK2 Protein-drug interactions
[0090] Deep mutational scanning was used to measure the effects of more than 20,000 amino acid variants on TYK2’s interferon-alpha (IFN-a) signaling and protein expression. The ability to quantify subtle differences in variant effects enabled mapping of core interactions responsible to drug binding, in addition to peripheral protein residues that contribute indirectly to inhibition. The IFNa assay was performed as depicted in FIG. 12 to identify drug resistant variants.
[0091] A first inhibitor of TYK2 (Drug 1) was tested as depicted in FIG. 11. The variants with the most significance were located in the known drug binding pocket.
[0092] Drug resistance of Drug 1 was then tested at varying concentrations. An experimental schematic is depicted in FIG. 12. The amount of drug resistant positions increased as the concentration of Drug 1 was lowered, as depicted in Table 6.Table 6
[0093] The DMS library was then tested against a second TYK2 inhibitor (Drug 2). There were distinct variants that were either resistant to Drug 1, Drug 2, or both Drug 1 and 2, as depicted in FIG. 15.
[0094] Next, lower drug concentrations were used to identify variants that increased the potency of Drug 1 and Drug 2. Results are depicted in FIG. 16. Variants were identified that improve inhibitor activities.
[0095] Subtle differences in binding interactions between two inhibitors that target the same site were distinguished and variants that increase the potency of these compounds were captured. The drug-protein interactions define structure-activity relationships for how the inhibitors functionally interact with TYK2 and point to where compounds could be optimized to increase potency. DMS can be used in the functional characterization and optimization of drug candidates.Example 5: High Resolution DMS for identification of Mc4R variants Assays for diseaserelevant mechanisms
[0096] Stimulation of MC4R with its agonist, alpha melanocyte-stimulating hormone (a- MSH), results in signaling through multiple canonical GPCR pathways, including Ga, -coupled cyclic adenosine monophosphate signaling (hereafter referred to as Gs) and Ga,-coupled calcium signaling. Therefore, first multiplexed reporter assays were developed for these two critical MC4R G-protein signaling functions (FIG. 17) These methods harness high-throughput DNA synthesis to construct every possible single amino acid variant, and each variant is then linked to a transcriptional reporter containing a sequence barcode unique to that variant. Reporter constructs were then integrated into cells using a site-specific recombination-based landing pad system and drug selection to ensure that each cell contains a single variant-barcode combination. Activation of the receptor turns on a response element for the signaling pathway, leading to the expression of the barcoded reporter, which is then quantified using RNA sequencing.Analysis model for statistically robust comparisons
[0097] One aspect of designing a potential drug is making iterative chemical changes to the molecule and testing these changes for their effects on the function of the target protein. The accumulation of these changes forms the basis of the structure-activity relationship (SAR). Bycomparing the functional consequences of thousands of different protein-ligand interactions at once, interactions may be identified that could be reverse-engineered into a more potent compound. However, the effects of chemical changes are typically small, so detecting them requires a much higher degree of statistical power than exists in typical DMS assays.
[0098] To explore this and many other hypotheses that arise in drug discovery applications, it is valuable to assay a DMS library using experimental replication and under a variety of conditions, such as different drugs and / or pathways of interest. This requires an analysis framework that enables direct and statistically robust comparison of different DMS datasets. However, other methods for DMS analysis do not leverage barcode or other replicate information, nor do they support hypothesis testing between conditions . An alternative modeling framework to enable this (FIG. 18). Borrowing from approaches for inferring differential expression from RNA-seq data, a mixed effect negative binomial generalized linear model (GLM) was applied to raw barcode counts directly. The model contains a random effect across barcodes to share barcode information between replicates and conditions, and incorporates sample-specific offsets to account for technical covariates like sequencing depth, as is common for RNA-seq. For each variant, the mean shift in barcode count and associated standard error is estimated for each treatment condition, relative to wild-type. Using percondition summary statistics, it was either directly tested whether each variant barcode mean is significantly different from wild-type (zero) in each treatment, or it was defined more complex linear contrasts on variant effects across multiple treatments.
[0099] In brief, variant segments of MC4R cDNA were amplified from DNA microarrays and cloned into base vectors through a multi-step process to yield pooled libraries oiMC4R variants with fully intact expression and reporter gene cassettes. Fully assembled plasmid libraries were then co-transfected with a plasmid encoding Bxbl recombinase into HEK293T cells containing a landing pad at the Hl 1 safe harbor locus to achieve single copy integration per cell. In a deviation from the previously published method, two independent replicates of each sub-library were cloned and pooled together post-cellular integration in order to maximize library coverage and the number of barcodes per variant.
[0100] Variant-Barcode Mapping: Illumina 2x150 BCL files were demultiplexed with bcl2fastq2 into R1 and R2 FASTQ files, which were merged into single fragments using Flash2 requiring a minimum 5 bp of overlap. The first 21 bp of each fragment corresponding to the barcode sequence were extracted into the read name, and the remaining fragment was adaptertrimmed using umi tools and cutadapt, respectively. The remaining fragments were mapped against a custom reference composed of the designed oligonucleotide library using STAR with default parameters except requiring that alignments be strictly unique to be reported. Takingeach alignment as an oligo-barcode pair, the read counts per unique oligo-barcode pair were computed for each replicate and joined by barcode. Finally, the resulting maps were filtered to require each oligo-barcode pair to pass three requirements in both replicates: correct barcode length, total read depth > 10, and purity > 0.75. The purity of an oligo-barcode pair was defined as the read count of that pair divided by the total number of reads containing that barcode. Postprocessing after STAR was performed using samtools for BAM manipulation and custom R code otherwise.
[0101] Running DMS Assays: The DMS assay protocol was adapted from a previously described method. HEK293T single-copy variant cell libraries were seeded at a density of -17x10 cells per 150 mm tissue-culture treated dish in DMEM + 10% fetal bovine serum (FBS). Four dishes were seeded for each experimental condition, with each dish being treated as an independent biological replicate (4 replicates per condition). Twenty -four hours after seeding, media was exchanged with DMEM + 0.5% FBS + / - 10 ng / ml Doxycycline. For chaperone experiments, all conditions were additionally replicated + / - 1 pM Ipsen-17. 24 hrs after Doxycycline induction, media was exchanged with Opti-MEM + DMSO, Forskolin, or MC4R agonist (a-MSH or THIQ). Forskolin bypasses MC4R to constitutively activate cAMP signaling, so this condition was used as a variant-independent measurement of library composition. For chaperone experiments, cells were washed 3x with 10 mL DMEM to remove Ipsen-17 prior to agonist stimulation. Six hours after agonist stimulation, cells were harvested by scraping in 4 ml lysis buffer (RLT buffer (Qiagen) + 143 mM P-ME). Lysis was performed by passing the cell slurry 6x through a sterile 18G needle and then spinning through QIAshredder (Qiagen) columns. RNA was extracted from 1 ml of the homogenized lysate with the RNeasy Plus Mini kit (Qiagen), including optional on-column DNAse digestion, and eluted into 100 pl H2O. Eight reverse transcriptase reactions per sample were performed with the SuperScript IV kit (Thermo Fisher). cDNA from each sample was treated with 1 pl RNase A (100 pg / ml, Thermo Fisher) and 3.2 pl RNase H (5,000 U / ml, NEB) at 37°C for 30 min. RNase-treated samples were concentrated to -55 pl by spinning through Amicon Ultra lOKDa concentrators (EMD Millipore) for -8 min. To determine the necessary cycle numbers for equivalent amplification of each sample library, qPCR reactions were performed on 1 pl cDNA (diluted 1 :8 in H2O) with Q5 polymerase (NEB), SYBR Green (Thermo Fisher), and library amplification primers. Final amplification cycles for each sample were chosen by adding 3 cycles to the Cq values generated from each respective qPCR reaction. Illumina sequencing libraries were prepared by amplifying 50 pl of each cDNA sample with sequencing adapters (500nM each library amplification primer) using the NEBNext Q5 High Fidelity 2x PCR Kit (NEB) under the following cycling conditions: 98°C for 30 s, X cycles of 98°C for 8 s, 65°C for 20 s, and 72°C for 10 s, followedby an extension of 72°C for 2 min. 3 pl of each DNA library sample was run on a 4% E-Gel (Thermo Fisher) and densitometry was performed with Fiji to account for differences in library yields. Samples were mixed at equal amounts into a single pool and then purified into 200 pl IDTE (Qiagen) with AxyPrep Magnetic beads (Fisher Scientific). The purified library was quantified with the DeNovix High-Sensitivity Fluorescence kit and prepared for sequencing with a 10% PhiX spike-in. Final library mixture was sequenced using custom read and index primers on an Illumina NextSeq 550 with the High Output 75 cycle kit.
[0102] Sequence Processing for Barcode Expression: Illumina 1x26 BCL files were demultiplexed with bcl2fastq2 and processed to remove the last 5 bp using basic bash commands. The resulting sequences were counted for each sample, and the resulting barcodes were joined with the appropriate oligo-barcode map. The resulting barcodes were joined with sample and MC4R variant metadata and returned for regression analysis. All processing after demultiplexing was performed with custom R code.
[0103] Negative Binomial Regression Analysis Pipeline: A mixed effects negative binomial general linear model (GLM) was developed to analyze the data. These models have been widely deployed to model count data, and in particular identify differential expression, in bulk and single-cell RNA-seq analysis. Maximum likelihood estimation was implemented for this model using glmmTMB, which can accommodate the potentially large scale of multiplex count data. For each position, all variants located at that position along with all wild-type variants were considered in the same chunk and the following model was applied:
[0107] For the zth condition, the / th variant, the Ath barcode, and the mth sample. Consequently, the first two terms in the last equation above correspond to a global mean term for each condition and a term for the variant-specific deviation from wild-type in each condition. The last two terms are the random effect for barcode k, and the sample-specific technical offset for sample m. The definition of the offset is often context specific, and here the log of the sum of barcode counts was used derived from stops, reasoning they should be constant across conditions and replicates.
[0108] The model for each MC4R position was fitted independently and extract coefficients for the additive shift in the mean of each variant relative to wild-type. Using the per-condition summary statistics, Wald test statistics were obtained by dividing the effect size by the standarderror and computed p-values against the normal distribution. P-values were adjusted for multiple testing using the Benjamini -Hochberg method and thresholded to 1% or 5% where indicated.
[0109] To define more complex null hypotheses like chaperone rescue, marginal means were extracted for each variant under each treatment using the emmeans package. The chaperone rescue contrast were defined as the additive shift of each variant in each treatment condition to the wild-type mean specifically in the untreated condition. Since this quantity is a linear contrast across marginal means, the associated contrast estimates and standard errors were computed.Increasing power through barcoding
[0110] As described above, the assay design and analysis framework harness DNA barcodes that are uniquely associated with a particular variant and provide multiple independent measures of a variant’s effect. The library cloning and cellular integration protocols were scaled and optimized to target ~30 barcodes per variant in building DMS libraries for MC4R. This increased the power to detect variant effects. For example, the separation between the activity of alleles that are clearly deleterious (i.e., a stop codon at any position in the protein) and all other alleles (i.e., wild-type or missense variants) was drastically increased in the MC4R assay relative to the same assay for p2AR (FIG. 19A). To further test the effect of the number of barcodes for a given variant, the barcodes were down-sampled for representative positions and ran the resulting data through the pipeline. As expected, by increasing the number of barcodes per variant, the magnitude of the standard error of the estimated variant effect decreases in a manner consistent with increased sample size (FIG. 19B). This confirms that increasing the number of barcodes per variant enabled the quantification of subtle differences within a standard hypothesis testing framework, and provides an experimental parameter that one can vary to improve the power to detect the effects of sequence variants of particular interest.
[0111] By treating barcodes as independent replicates for their associated variants, a statistical model that leverages this added replication, accounts for various technical covariates, enables hypothesis testing for individual variant effects (e.g., the difference between a variant and wildtype), and even more complicated linear contrasts (e.g. the difference between a variant effect in a-MSH versus THIQ) was developed. Having this level of statistical precision is a major step in advancing DMS from a descriptive method to a quantitative one, and opens up further applications in fields like drug discovery where addressing these problems are paramount to progress.Comprehensive deep mutational scanning of MC4R
[0112] With these methods in hand, a comprehensive assessment of the effects of all single amino acid substitutions (including nonsense variants) on MC4R’s Gs and Gq signaling activities under a variety of experimental conditions was performed (FIG. 20A). Experimentalconditions were selected that would inform on highly-relevant aspects of drug discovery and development programs, such as elucidating protein structure-function relationships, identifying regions of the protein that bias activity towards or away from a specific function, classifying the effects of human variants in the presence and absence of potential therapies, and uncovering functional differences in protein-ligand interactions. In total, 18 unique conditions were tested, each performed in quadruplicate, including: basal activity (i.e., no stimulation) of MC4R, stimulation of MC4R with a range of doses of the native peptide agonist alpha-melanocyte- stimulating hormone (a-MSH), stimulation with a range of doses of a small molecule agonist (THIQ), treatment with a small molecule corrector (Ipsen-17), and library composition normalization controls (forskolin). The resulting DMS assays had extraordinary variant coverage, with >99.89% (6,633 / 6,640) of all possible single amino acid substitutions present in all experimental conditions. Each variant was represented by an average of 56 and 28 barcodes for the Gs and Gq signaling pathways, respectively. Between both assays, this translates to more than 557,000 uniquely engineered human cells, each containing a distinctive variant-reporterbarcode combination. When factoring in the number of experimental conditions (18 unique), replicates (four per condition), amino acid variants tested (>99.89% of 6,640 possible), and the mean barcodes per variant (56 and 28 for Gs and Gq, respectively), this equates to >21,500,000 independent measurements across all datasets.
[0113] Multiple lines of evidence support the high quality and utility of these data (FIG. 20A- 20F). Focusing on one dataset as a representative example (Gs signaling using a low dose of a- MSH stimulation), variants that introduce stop codons or fall within transmembrane domains and buried surfaces disproportionately lead to significant loss of MC4R function (FIG. 20A- 20C). The results also correlate well with expectations from human genetics data and variant effect prediction algorithms (FIG. 20C-20D). For example, the majority of human MC4R variants classified as pathogenic or likely pathogenic in ClinVar lead to a significant reduction of Gs signaling under low a-MSH stimulation conditions (FIG. 20C). Variants that are significantly loss-of-function in this assay are rarer in the human population, and more common human variants have no significant effect on MC4R function (FIG. 20D). Loss-of-function variants by DMS are also typically predicted to be deleterious by commonly used variant effect predictors like AlphaMissense and popEVE.
[0114] Because of the sensitivity of the reporter system and the statistical power gained by testing dozens of unique barcodes per variant, it was anticipated that these assays would capture subtle quantitative, rather than just qualitative, effects on MC4R function. To assess this, the results were benchmarked against previous quantitative characterizations of MC4R variants from the literature (FIG. 20E-20F). For example, many MC4R variants that have been observedin the human population have been previously tested for their effects on MC4R function, and a review summarizing this work systematically classified >70 variants according to whether they result in “full”, “partial”, or “no / mild” loss of MC4R activity. The results were consistent with these classifications: variants classified previously as “full” loss-of-function typically have very low MC4R activity in our assay, “partial” variants have intermediate effects, and “no / mild” effect variants have near-normal activity (FIG. 20E, median log2[fold change of variant activity over wild-type activity] of -1.4, -0.8, and -0.3, respectively, for each group). Finally, the results show a high degree of correlation (Pearson correlation = 0.84, R2= 0.71) with quantitative effect measurements reported for 25 variants individually introduced into the orthosteric site of MC4R. Collectively, this demonstrates the our high quality MC4R DMS data accurately and quantitatively assessed the effects of variants on MC4R’s function.Systematic Human Variant Interpretation of MC4R
[0115] Looking across multiple data sets gives a comprehensive picture of variant effects, as the various experimental conditions tested have disparate power to detect loss-of-function versus gain-of-function activities (FIG. 20G). For example, unstimulated conditions (i.e., zero a-MSH) highlight variants that lead to constitutive activation of MC4R, but they have less power to detect loss-of-function variants. In contrast, conditions with agonist stimulation are much better powered to identify loss of MC4R function. Considering individual a-MSH stimulation conditions (zero, low, medium, and high) for both the Gs and Gq assays, each condition identifies 6.6 - 39.3% of variants as loss-of-function and 0.02 - 1.1% as gain-of-function (FIG. 20G). Collectively across all a-MSH stimulation conditions, 3,370 variants (50.8%) are loss-of- function in at least one condition, 347 variants (5.2%) are gain-of-function in at least one condition, and 2,996 (45.2%) always show wild-type activity. Interestingly, 80 (1.2%) variants are classified as both loss- and gain-of-function, depending upon the condition.
[0116] To aid in clinical variant interpretation, detailed functional effect classifications were provided for 220 human variants reported in ClinVar or the literature from patient sequencing studies. In total, 130 of these human variants (59.1%) are LOF in at least one condition, consistent with being pathogenic for obesity -related phenotypes. This includes 83.9% (26 / 31) of the variants that are reported in ClinVar as pathogenic or likely pathogenic, 53.3% (32 / 60) of those that are unclassified or have conflicting interpretation, and 0% (0 / 3) of those classified as benign or likely benign. A small number of reported human variants (V103I, H158R, I251L) result in significant increases in Gs and / or Gq signaling, consistent with having a protective effect for obesity, and are generally classified as benign in ClinVar or as having wild-type activity in the literature. These results highlight the utility of systematic deep scans across multiple experimental conditions for facilitating human variant interpretation.Variants that bias MC4R function
[0117] MC4R signals through multiple G protein pathways. To gain a better understanding of how MC4R structure relates to its various functions, variants that differentially impact Gs versus Gq signaling were searched for. Principal Component Analysis (PCA) was applied across eight total DMS datasets (four each for Gs and Gq: with zero, low, medium, and high a-MSH stimulation). The first two principal components explained 66% and 12% of the variance, respectively (FIG. 21A). Through inspection, Principal Component 1 (PCI) separated variants that impact both signaling functions, with variants that are loss-of-function for both Gs and Gq having higher PCI values. PC2 separates variants that affect Gs and Gq signaling (FIGS. 21 A- 21D) Variants with higher PC2 values exhibited Gq bias by having greater than wild-type levels of Gq signaling activity, while retaining wild-type levels of Gs activity (FIG. 21C,DD). In contrast, variants with more negative PC2 values are Gs-biased variants that typically have wildtype levels of Gs signaling and reduced Gq signaling (FIF. 21C).
[0118] Overall, more variants increase Gq-bias than Gs-bias (FIG. 21A). Gs is the primary G protein coupling for MC4R, and the data suggested that there is little room for further improving MC4R’s robust Gs signaling activity. Biased variants were positionally diverse. For example, the 14 variants that displayed the most extreme Gq bias (FIG. 21A) were found at 12 different residue positions, with position 79 unique in having several variants that result in Gq bias. Many of the variants with extreme Gq or Gs bias were located within the regions of MC4R that interact with G proteins, with some scattered throughout the transmembrane domains and far fewer in the vicinity of the agonist binding site (FIG. 21-21D).
[0119] Comparing these results with existing structural information could provide additional detailed insights into MC4R’s signaling functions. It is thought that ligand binding in GPCR orthosteric sites is communicated to the intra-cellular G-protein binding domain through a series of conserved residues or “microswitches”. Structural studies comparing the inactive and active state structures have confirmed that MC4R shares a similar signaling. Upon ligand binding, W258 (W258«8in GPCRdb nomenclature) of the conserved CWxP motif undergoes a conformational rearrangement that is translated to L I 33 and 1137 , of the conserved PIF motif (MIF in melanocortin receptors). This causes F254“ in the PIF motif to rearrange, which in turn, disrupts the packing of three different interactions: 1) L I 40 and 1143 , 2) 125 F- and L247“ and 3) R.147 and N240 . These, amongst other rearrangements, culminate in the receptor being able to bind a G-protein. This interaction with the G-protein is primarily mediated through R.147 in the conserved DRY motif,
[0120] A number of mutations at residues throughout this signaling cascade had extremely positive PC2 values, implicating them as Gq-biasing mutations. Within the core of the cascade,I l 37 L, F254 P, and LI 40 1 were identified as Gq-biased (FIG. 21A). A number of Gq- biasing mutations were identified within the G-protein binding pocket, specifically Tl 50 G and Hl 58 R (FIG. 21A,C). Hl 58 R is found in the human population and has previously been shown to preferentially signal through Gq. Hl 58 is also co-located near two other mutations (KI 64 L, Fl 52 R) in intracellular loop 2 (ICL2) and near the ends of the third and fourth transmembrane domains (TM3 and TM4), that display bias (FIG. 21C). K 164 L exhibits a Gs bias in that it drives loss of function through Gq.
[0121] The data also point to a number of potentially novel interactions. For example, M79 packs against residues H387°“ ’ and Q390 of the Gs alpha subunit. This position has multiple different variants that result in Gq bias, with M79 R, M79 S, and M79 G being the most extreme (FIG. 21A). The most extreme signal in our PCA analysis came from I223 L (FIG. 21 A,D), which interfaces with the C-terminus of ,a, near position L394. Further down towards the intracellular side of TM5, V228R (FIG. 21A,D) was identified, which interfaces proximal to E323“‘M I3in a.
[0122] Collectively, these results highlight the power of DMS to identify variants that result in biased signaling, which could be harnessed for designing drugs that precisely modulate specific cellular functions. Furthermore, the combination of DMS data and structural information is a fruitful avenue for generating and validating protein structure-function hypotheses.Example 6: High resolution DMS for identification systematic prediction of treatment response to a MC4R corrector
[0123] Many variants of MC4R disrupt signaling by causing protein misfolding, which ultimately inhibits proper localization of MC4R to the cell membrane. Correctors are small molecule drugs that facilitate protein folding and trafficking. Identifying variants that respond to corrector therapy is typically done by rigorously testing the effect of a compound on a single variant at a time. DMS offers an attractive avenue to systematically test the treatment response of thousands of patient variants in a single assay.
[0124] To this end, Ipsen 17 was tested to see if it was able to restore the Gs signaling function of the MC4R variants in our DMS library. Out of all 6,633 tested variants, 290 (4.4%) showed disrupted Gs signaling in the absence of treatment that was partially or fully rescued by the addition of Ipsen 17 (FIG. 22). This includes a number of variants that have been classified as pathogenic in ClinVar. Other reported patient variants showed no functional improvement in response to corrector therapy (FIG. 23). Collectively, these data support that performing DMS in the presence of a small molecule corrector can be used to predict which patients are likely to benefit from such treatment options.Mapping protein-ligand interactions
[0125] Substantial work has been done to characterize how peptide agonists interact structurally with MC4R, but similar work on small-molecule agonists with comparable activity and selectivity remains relatively limited. DMS experiments can be used to define “drug- resistant” variants within MC4R that disrupt the activity of different types of ligands, providing functional insight into protein-ligand interactions that are key for understanding the mechanisms underlying protein agonism. Such functional information would be a valuable addition to structural methods and has the potential to streamline the lengthy and iterative cycle of compound optimization in drug discovery. To begin to explore this application, DMS of MC4R was performed using both native peptide agonist stimulation (a-MSH) and small molecule agonist stimulation (THIQ) to characterize the functional interaction landscapes of each ligand and to better understand what distinguishes peptide and small-molecule MC4R pharmacophores. Bayesian meta-regression analysis of the lowest dose concentrations a-MSH and THIQ revealed a set of variants that uniquely disrupt activation by one ligand but not the other (FIG. 24A). These variants cluster exclusively within close proximity of the orthosteric binding pocket (FIG. 24B). This includes all known binding interactions of each molecule, in addition to uncharacterized proximal positions, confirming their direct effect on ligand binding.
[0126] These results revealed that some of the variants with the most significant reductions in signaling activity occur at positions that harbor multiple variants that can have differentially deleterious effects under different ligand conditions. For example, MC4R is much less tolerant of mutations at P48“ under a-MSH stimulation than THIQ stimulation, but the negatively charged P48 D variant uniquely ablates THIQ activation. Conversely, multiple polar uncharged variants at 1129 alter THIQ activation, while H29M3V uniquely inhibits a-MSH activation (FIG. 23C-23E). The HFRW motif (His6-Phe7-Arg8-Trp9) of a-MSH represents a conserved pharmacophore critical for activation of MC4R by peptides, and the tri -branched THIQ molecule (R1-R2-R3) mimics the HFRW conformational architecture (FIG. 24E). The observation that a-MSH is more sensitive to variants at position P48Mtis consistent with how the His6 of a-MSH forms more interactions with this hydrophobic pocket, while the analogous R3 group of THIQ is more flexible and forms non-specific interactions in this region (FIG. 24C- 24E). The pattern of variant effects observed at I129«2is more complex - substitutions to histidine and polar uncharged residues uniquely abrogate MC4R activity upon THIQ stimulation, while substitution to valine only disrupts MC4R activity upon a-MSH activation (FIG. 24C-24E). The Phe7 of a-MSH and the analogous R2 group of THIQ form key interactions with Ca2+ and the core hydrophobic pocket formed by 1129 of MC4R (FIG. 24E). The enrichment of variants at 1129 that uniquely disrupt MC4R activation by THIQ points to astronger dependency of the small-molecule on interactions with this residue. A puzzling exception to this pattern is the fact that the relatively minor I I 29 V variant reduces MC4R activation under a-MSH stimulation but not THIQ stimulation. Conversely, 1104 is an example of the many detected as more important for a-MSH stimulation than THIQ stimulation (FIG.24C-24E)
[0127] In summary, by assessing the functional consequences of all mutations within the MC4R orthosteric site, known binding interactions were confirmed but also a network of underlying interdependencies were revealed that distinguish peptide and small-molecule activation. These relationships provide additional functional insight into the structural mechanism of MC4R ligand binding that could be harnessed for drug design.Example 7: Classifying GLP1R and GIPR variants for drug development
[0128] Deep mutational scanning was used to measure the effects of multiple amino acid variants on GLP1R and GIPR signaling and protein expression. The library was treated with multiple compounds to identify which mutations affect signaling pathways through GLP1R or GIPR. FIG. 25A depicts loss of function variant effects for GLP1R residues within 5 angstroms of either drug (based on resolved structures of GLP1R with said drugs), classifying variants as leading to loss of function for drug 1 only, drug 2 only, both drugs 1 and 2, or neither drug. FIG. 25B depicts log2 fold change of signaling relative to wild-type for each amino acid substitution at R380 of GLP1R. FIG. 25C depicts how different mutations at R380 can improve or reduce the affinity of Drug 1 or Drug 2 to bind the target site. These insights can be used to develop a drug that targets both GLP1R and GIPR.
[0129] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is intended that the following claims define the scope of the invention and that methods and compositions within the scope of these claims and their equivalents be covered thereby.
Claims
CLAIMS1. A method of high resolution mapping a functional site of a protein that influences a biological function of the protein, the method comprising: a. providing a variant library, wherein the variant library comprises a plurality of nucleic acids, wherein individual members of the plurality of nucleic acids encode the protein or a variant of the protein, wherein the plurality of nucleic acids encode a plurality of protein sequences comprising a plurality of alterations to at least 75% of the amino acid residues of the protein across the plurality of the protein sequences; b. expressing the variant library in a plurality of cells; and c. performing a biological assay on the plurality of cells, wherein the biological assay comprises detection of an index sequence unique to the protein or the variant of the protein.
2. The method of claim 1, wherein the biological assay comprises a measure of protein abundance.
3. The method of claim 1 or 2, wherein the plurality of cells is comprised within a partition.
4. The method of claim 3, wherein the partition comprises a tissue culture flask, plate, or dish.
5. The method of claim 3, wherein the partition comprises a well of a well-plate.
6. The method of claim 4, wherein the well-plate comprises a 6-well, 12-well, 24-well, 48- well, 96-well or 384-well plate.
7. The method of any one of claims 1 to 6, where the plurality of cells comprises mammalian cells.
8. The method of any one of claims 1 to 6, wherein the plurality of cells comprises human cells.
9. The method of any one of claims 1 to 8, wherein each cell of the plurality of cells comprises only one of the plurality of nucleic acids.
10. The method of any one of claims 1 to 9, wherein the protein comprises at least 100 amino acids.
11. The method of any one of claims 1 to 10, wherein the protein is a human protein.
12. The method of any one of claims 1 to 11, wherein the protein comprises a GPCR, a receptor tyrosine kinase, a nuclear hormone receptor, an ion channel, a transcription factor, or an integrin.
13. The method of any one of claims 1 to 12, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 80% of the amino acid residues of the protein across the plurality of protein sequences.
14. The method of any one of claims 1 to 12, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 90% of the amino acid residues of the protein across the plurality of protein sequences.
15. The method of any one of claims 1 to 12, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 95% of the amino acid residues of the protein across the plurality of protein sequences.
16. The method of any one of claims 1 to 12, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 98% of the amino acid residues of the protein across the plurality of protein sequences.
17. The method of any one of claims 1 to 12, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to at least 99% of the amino acid residues of the protein across the plurality of protein sequences.
18. The method of any one of claims 1 to 12, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to 100% of the amino acid residues of the protein across the plurality of protein sequences.
19. The method of any one of claims 1 to 18, wherein the plurality of alterations to the amino acid residues of the protein comprises amino acid substitutions to at least 10 different amino acids.
20. The method of any one of claims 1 to 18, wherein the plurality of alterations to the amino acid residues of the protein comprises amino acid substitutions to at least 15 different amino acids.
21. The method of any one of claims 1 to 18, wherein the plurality of alterations to the amino acid residues of the protein comprises amino acid substitutions to at least 19 different amino acids.
22. The method of any one of claims 1 to 21, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to 80% of the codons encoding the amino acid residues of the protein across the plurality of protein sequences.
23. The method of any one of claims 1 to 21, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to 90% of the codons encoding the amino acid residues of the protein across the plurality of protein sequences.
24. The method of any one of claims 1 to 21, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to 95% of the codons encoding the amino acid residues of the protein across the plurality of protein sequences.
25. The method of any one of claims 1 to 21, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to 98% of the codons encoding the amino acid residues of the protein across the plurality of protein sequences.
26. The method of any one of claims 1 to 21, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to 99% of the codons encoding the amino acid residues of the protein across the plurality of protein sequences.
27. The method of any one of claims 1 to 21, wherein the plurality of nucleic acids encodes a plurality of protein sequences comprising alterations to 100% of the codons encoding the amino acid residues of the protein across the plurality of protein sequences.
28. The method of any one of claims 1 to 27, wherein the protein or the variant of the protein is represented by at least 10 different unique index sequences unique to the protein or the variant of the protein.
29. The method of any one of claims 1 to 27, wherein the protein or the variant of the protein is represented by at least 20 different unique index sequences unique to the protein or the variant of the protein.
30. The method of any one of claims 1 to 27, wherein the protein or the variant of the protein is represented by at least 30 different unique index sequences unique to the protein or the variant of the protein.
31. The method of any one of claims 1 to 30, wherein the biological assay comprises activation of a reporter gene, wherein the reporter gene comprises an index sequence unique to the protein or the variant of the protein.
32. The method of any one of claims 1 to 31, wherein the reporter gene is operatively coupled to a promoter.
33. The method of any one of claims 1 to 32, wherein the functional site comprises an allosteric regulatory site.
34. The method of any one of claims 1 to 32, wherein the functional site is a site controlling differential signaling by the protein.
35. The method of any one of claims 1 to 34, further comprising identifying the functional site to a resolution of five amino acids.
36. The method of any one of claims 1 to 34, further comprising identifying the functional site to single amino acid resolution.
37. The method of any one of claims 1 to 36, wherein the biological function comprises signal transduction of the protein, a protein expression level, protein folding, intracellular trafficking, or cell surface expression.
38. The method of any one of claims 1 to 37, wherein individual members of the plurality of nucleic acids of the variant library encode the protein and a variant of the protein.
39. The method of any one of claims 1 to 38, the method further comprising contacting the plurality of cells with a test agent.
40. The method of claim 39, wherein contacting the plurality of cells with a test agent occurs before performing the biological assay on the plurality of cells.
41. The method of any one of claims 1 to 40, wherein the biological assay that comprises a measure of protein abundance measures steady-state abundance, cell surface expression, protein trafficking, or folding of the protein.
42. The method of any one of claims 1 to 41, wherein the method further comprises performing a second biological assay.
43. The method of claim 42, wherein the second biological assay comprises a measure of signaling by the protein.
Citation Information
Patent Citations
De novo synthesized combinatorial nucleic acid libraries
WO2018170164A1
Cell-stored barcoded deep mutational scanning libraries and uses of the same
WO2020006494A1
Method for evaluating clinical relevance of genetic variance
WO2024036234A1